What VetEval is
A benchmark that measures how well AI models answer veterinary clinical questions, using items written by qualified veterinary professionals.
VetEval measures how well a language model answers veterinary examination items. Every item is written and reviewed by qualified veterinary professionals, models answer the same items under the same settings, and the result is a published score with a confidence interval next to it.
It exists because there was no way to compare veterinary AI on veterinary work. General benchmarks measure general reasoning, and a model that scores well on those can still get a dose calculation wrong in a way that would harm an animal.
What a score is
A VetEval Index is a weighted percentage across six competency categories, plus a safety component. The weights are published in full on the methodology page and they are the same for every model.
What a score is not
A score is a measurement of examination performance under stated conditions. It is not a licence, a certification, or a clinical recommendation, and no result here should be read as approval of a model for clinical use. Nothing on this platform is veterinary advice.
Who it is for
- Model providers who want a number they can cite, measured by somebody who did not build the model, and a way to see where it is weak.
- Clinics and groups choosing a model, who need to compare like with like rather than read marketing pages.
- Researchers, press and professional bodies, who need the method to be public and the evidence to exist.