About

An instrument for a credentialed profession

Models published
11
Answers per item
3
Scoring categories
6 + safety
Vendor funding
None

No vendor funds VetEval’s operations or influences any score. Evaluation fees are uniform and published; they cover compute and grading. Publication is free and score-blind.

VetEval is an evaluation service operated by viggoVet. We measure how well AI models perform on real veterinary medicine, clinical diagnosis, pharmacology and dosing, species-specific knowledge, triage, preventive care, communication, and a dedicated safety axis, against a test set authored and reviewed by licensed veterinarians. The proof is the product: every published number carries a confidence interval, a verification tier, and a safety status.

We exist because general-purpose benchmarks miss what matters in a clinic. A model can sound fluent and still recommend a species-inappropriate drug, a lethal dose, or fail to refer an emergency. VetEval is built to surface exactly those failures and to weight the questions by the caseload a working veterinarian actually sees.

Separation & data integrity

viggoVet builds veterinary AI and also operates VetEval. viggoVet models are never ranked on the public leaderboard, and never will be. The item bank, the gold answers, and the transcripts stay with the benchmark team, walled off from everything viggoVet builds, and are never used to train or tune any model. viggoVet's own products are validated separately and published as viggoVet research, because a leaderboard cannot measure whether a monitoring alarm fires in time.

For model builders

Connect your model's endpoint and we run the evaluation on our own infrastructure, then publish the results with confidence intervals and a safety-gate status. Private validation sets and custom enterprise evaluations are available.

Submit your model · Read the methodology

Contact

Press, partnerships and evaluation enquiries all go through the contact form.

Governance

viggoVet builds veterinary models and also operates VetEval. These are the controls that keep the two apart.

Read how the two teams are kept apart

  1. Run conditionsTemperature 0, a pinned seed, and every item asked a fixed number of times on every scored run. The harness refuses a run configured any other way.
  2. The datasetPractice items and scored items are separate states in the database, and neither can become the other.
  3. First-party modelsNever ranked on the public leaderboard. Validated separately as viggoVet research.

Contact

One form for everything. Pick a reason so whoever reads it knows what it is about: submissions, clinical questions, press, or a data request.

Do not send clinical details about an identifiable animal or client. This form is not a clinical service.

Pick a reason and we will say what to include.

Complete the check to continue.

We use what you send only to answer you. See our privacy notice. Standard plan pricing is uniform and published to every customer, this form is for billing arrangements and genuinely different scope, not for negotiating the list price.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.