GETTING STARTED

What VetEval is

A benchmark that measures how well AI models answer veterinary clinical questions, using items written by qualified veterinary professionals.

VetEval measures how well a language model answers veterinary examination items. Every item is written and reviewed by qualified veterinary professionals, models answer the same items under the same settings, and the result is a published score with a confidence interval next to it.

It exists because there was no way to compare veterinary AI on veterinary work. General benchmarks measure general reasoning, and a model that scores well on those can still get a dose calculation wrong in a way that would harm an animal.

What a score is

A VetEval Index is a weighted percentage across six competency categories, plus a safety component. The weights are published in full on the methodology page and they are the same for every model.

The categories and their weights
CategoryWhat it coversWeight
C1Clinical diagnosis and differential24%
C2Pharmacology and dosing22%
C3Species-specific knowledge14%
C4Triage and emergency14%
C5Preventive care, public health and zoonoses10%
C6Communication and professionalism6%
SAFETYHarm avoidance, as a positive component10%

What a score is not

A score is a measurement of examination performance under stated conditions. It is not a licence, a certification, or a clinical recommendation, and no result here should be read as approval of a model for clinical use. Nothing on this platform is veterinary advice.

Who it is for

  • Model providers who want a number they can cite, measured by somebody who did not build the model, and a way to see where it is weak.
  • Clinics and groups choosing a model, who need to compare like with like rather than read marketing pages.
  • Researchers, press and professional bodies, who need the method to be public and the evidence to exist.

Questions we get asked

Was this page helpful?

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.