In June 2025, the American College of Veterinary Radiology and the European College of Veterinary Diagnostic Imaging published a joint position statement in JAVMA. Buried in its recommendations is a sentence that should have started a race: stakeholders should establish an independent organization dedicated to creating guidelines and oversight for veterinary AI products, the way AAFCO does for feed labeling or the FDA does for human medical software.
More than a year later, that organization still does not exist. This post is about the gap the colleges named: how it formed, what it costs the people working inside it, what human medicine built instead, and why we built VetEval to close one specific, measurable part of it.
Adoption arrived before validation
The profession did not wait for oversight to start using AI. The largest survey to date, run by Digitail with the American Animal Hospital Association across 3,968 veterinary professionals and later published in the AVMA's American Journal of Veterinary Research, found that 39.2% of respondents were already using AI tools in their veterinary setting, and 69.5% of those users were using them daily or weekly. A second wave of the same survey is in the field now, and everything about the intervening period, from scribes maturing into decision-adjacent tools to general-purpose assistants answering clinical questions informally, points in one direction.
The same survey found the most prevalent concern, held by 70.3% of respondents: the reliability and accuracy of AI systems. Read those two numbers together, because they are the whole situation in miniature. Four in ten professionals are using the tools. Seven in ten do not have a way to know whether the tools are right.
Concern without an instrument does not slow adoption; it just privatizes the risk. Which is exactly what the regulatory picture guarantees.
The vacuum is structural, not an oversight
In human medicine, an AI tool that touches diagnosis passes through premarket review as software as a medical device, and a post-market monitoring apparatus follows it into practice. In veterinary medicine there is no equivalent gate at either end. As AVMA's own reporting on the ethics of veterinary AI put it, companies developing medical devices for animals are not required to undergo premarket screening, and the ACVR has raised concern about the lack of oversight for software reading radiographs. A 2026 systematic survey in Frontiers in Veterinary Science states the consequence precisely: there are no federal premarket approval requirements for veterinary AI software in North America, creating a regulatory vacuum in which the responsibility for appropriate use rests entirely with the veterinarian.
Sit with that last clause, because it is a job description nobody applied for. Every practice evaluating an AI tool today is running its own private, unfunded, unpublished clinical trial: adopt, watch for errors, hope the errors are the kind you catch, and carry the liability personally if they are not. Trial and error is an acceptable way to choose a coffee machine. It is a dangerous way to choose software that suggests doses, and it is a wasteful way to run a profession: thousands of clinics independently rediscovering the same failure modes, with no mechanism to pool what any of them learned.
A flow diagram covering: Human medicine AI; Product built; Premarket review FDA SaMD pathway; Post-market monitoring registries, audits; Clinical use; Veterinary AI today; Marketing claims.
How human medicine built the infrastructure veterinary medicine skipped
It is worth being specific about what "evaluation infrastructure" means, because human medicine assembled it in visible layers over about five years, and the layers are the blueprint.
The first layer was standardized knowledge testing: multiple-choice benchmarks built from licensing-examination content, which gave the field a shared, repeatable question: how much medicine does this model know? The second layer was expert-graded long-form evaluation: physicians scoring full clinical answers against rubrics, because multiple choice cannot see reasoning quality, omission, or dangerous confidence. The third layer, arriving by 2025 with physician-authored conversational suites, was scale and realism: thousands of rubric-graded, multi-turn clinical conversations. In parallel, a safety-specific literature emerged to measure hallucination and harm as their own dimensions rather than as subtractions from accuracy.
None of this required a regulator. It required institutions willing to author questions, publish methods, and keep measuring as models changed. That last property is the one that matters most: these are standing instruments. A model released next year lands on the same scales as a model released last year, and the field can watch capability move.
Veterinary medicine has the beginnings of layer one only, and only as snapshots. Studies have now measured model performance on licensing-style content in several countries, including a Scientific Reports evaluation on Japan's National Veterinary Licensing Examination showing newer reasoning models outperforming their predecessors, and similar academic measurements on veterinary curricula and specialty content. These studies matter: they prove the measurement is possible and that capability is moving. But each is a fixed exam, run once, on questions that often live in the public domain where future models can absorb them. A snapshot cannot control contamination, cannot weight by the caseload a working veterinarian actually sees, cannot follow models across generations, and cannot treat harm as its own axis. Snapshots imply the instrument. They are not the instrument.
Why the species dimension changes evaluation design
There is a tempting shortcut: veterinary medicine could just borrow human medicine's benchmarks and swap the vocabulary. The shortcut fails because the failure modes of veterinary AI are species-shaped, and species-shaped failures break the assumptions general medical benchmarks are built on.
A drug that comforts a dog can kill the cat in the next exam room; acetaminophen and permethrin are the canonical examples, and they are core knowledge, not trivia. Dosing spans a 2 kg kitten and a 700 kg horse in the same afternoon, computed by weight, with narrow-margin drugs that punish a decimal error immediately. And a meaningful class of veterinary emergencies is defined by the referral clock: gastric dilatation-volvulus, urethral obstruction, dystocia, where a fluent answer that fails to say "go now" is itself the harm. An evaluation for this domain has to distribute questions across species the way caseloads actually distribute, and it has to score harmful answers as a different category of event rather than a few lost accuracy points. A model that is usually right and occasionally lethal is not a slightly worse model. It is a different kind of instrument, and averages cannot see the difference.
What an instrument owes the profession
VetEval is our attempt at the instrument the colleges called for, scoped honestly. We measure how well AI models perform on veterinary examination items: expert-authored questions across species and competencies, contamination controls on a private held-out bank, caseload-weighted scoring, confidence intervals with honest ties, and a safety gate that caps a model's score when it produces harmful advice confirmed by human review. Every published number carries its uncertainty and traces to the exact evaluation that produced it, and each model has a full page with its complete breakdowns.
Equally important is what a VetEval score is not. It is not clinical advice. It does not license any model for clinical use. It does not evaluate whole products in deployment, where guardrails, interfaces, and the veterinarian in the loop change everything. It answers the one question the profession currently cannot answer without running the experiment on its own patients: which models know veterinary medicine, and which fail in ways that matter.
The methodology, the scoring math, and the governance that keeps the benchmark team separated from its operator's product teams are published in full. Read them before you trust us. That is rather the point.
What a clinic can do this quarter
The instrument helps, but procurement discipline is available today, benchmark or no benchmark. Four questions to put in writing to any AI vendor: what validation evidence exists beyond your own marketing, and who authored the test? Do your published numbers carry confidence intervals, and on what question set? How was harmful-answer behavior tested, and by whom was it reviewed? And what happens to my clinic's data? A vendor who cannot answer those four in writing has answered them. Our companion guide, How to read an AI benchmark before you trust it, expands each into what good looks like and what the red flags are.
Common questions
Who runs VetEval? VetEval is operated by viggoVet, and viggoVet models are never ranked on the public leaderboard. The full separation controls, including the data firewall between the benchmark team and any model work, are published on the methodology page.
Does a VetEval score mean a model is safe to use in practice? No. It measures performance on examination items under fixed conditions. Clinical deployment involves the product built around a model, its guardrails, and the veterinarian in the loop, none of which an examination can see.
Why not wait for a regulator or an official body? We would welcome one, and the ACVR/ECVDI statement shows the profession's colleges want one. But adoption is at 39% now, and an independent, transparent, methodology-published instrument can exist while institutions deliberate. Nothing about VetEval precludes official oversight; if anything, a working measurement precedent makes it easier.
Can any vendor be scored? Any model reachable by API can be submitted. Evaluation runs on our infrastructure at published, uniform terms, and publication is the vendor's choice, made blind to the score.
How do you keep models from memorizing the test? A private held-out bank that is never published, per-evaluation sampling so no fixed exam accumulates exposure, paraphrase perturbation, and per-model contamination signals. The full protocol, including a failure of our own earlier design and what it taught us, is published in the methodology and on this blog.
Is this only for companion-animal medicine? The bank spans species groups weighted to licensing-examination caseload proportions, including production and equine medicine, with coverage limits disclosed. Per-species breakdowns are published so readers in any practice type can weight for their own reality.
References
- ACVR and ECVDI position statement on artificial intelligence, JAVMA 263(6), June 2025. https://avmajournals.avma.org/view/journals/javma/263/6/javma.25.01.0027.xml
- Digitail and AAHA, "AI in Veterinary Medicine: The Next Paradigm Shift," survey of 3,968 professionals, 2024. https://digitail.com/blog/39-2-of-veterinary-professionals-use-ai-tools-in-their-practice-digitail-and-aaha-survey/
- Peer-reviewed survey publication, American Journal of Veterinary Research 86(S1), 2025. https://avmajournals.avma.org/view/journals/ajvr/86/S1/ajvr.24.10.0293.xml
- AVMA News, "Artificial intelligence in veterinary medicine: what are the ethical and legal implications?" https://www.avma.org/news/artificial-intelligence-veterinary-medicine-what-are-ethical-and-legal-implications
- "AI applications in veterinary digital health: a systematic survey," Frontiers in Veterinary Science, 2026. https://www.frontiersin.org/journals/veterinary-science/articles/10.3389/fvets.2026.1853395/full
- "Performance evaluation of generative pre-trained transformer on the National Veterinary Licensing Examination in Japan," Scientific Reports. https://pmc.ncbi.nlm.nih.gov/articles/PMC12909958/
- Digitail and AAHA, second industry-wide AI survey announcement, 2026. https://digitail.com/blog/digitail-invites-veterinary-professionals-to-share-their-views-on-ai-in-second-industry-wide-survey/