Coverage Policy
In effect from 27 August 2026.
Why this document exists, and why it is dated before the launch run
Every benchmark is eventually asked the same question: why is my model not on it, and why is that one? A roster assembled first and explained afterwards cannot answer that credibly, because any explanation offered after the fact is indistinguishable from a justification of choices already made.
So this policy is written and published before the launch run is executed. The roster that appears at launch is the output of the criteria below, not the input to them. Anyone may check the two against each other.
Nothing here is a promise about a specific model, a specific date, or a specific score. It is a description of how the roster is decided and how it changes.
Scope and status. This policy governs which models appear on the public leaderboard. It sits alongside the Methodology, which governs how a model is measured, the Terms of Service, the Publication, Corrections & Independence Policy, and the Disclaimer. Where this policy and the Terms of Service differ on a contractual matter, the Terms of Service govern. VetEval publishes benchmark measurements, not clinical guidance, and a listing is never a certification, endorsement, or approval of a model for clinical use.
This policy creates no entitlement. It is a public statement of how we work, not an offer, a contract, or a service owed to any vendor. No organisation has a right to be listed, to be listed by a particular date, or to remain listed, and nothing here is intended to create rights enforceable by any third party. We may decline or defer any evaluation, and we may amend this policy under Section 7.
1. What may be evaluated
A model is eligible to appear on the public leaderboard when all of the following hold. Each is enforced by the platform, not merely stated here; where a rule is a person's judgement rather than a check in software, that is said explicitly.
1.1 It is reachable by anyone, not only by us
Either:
- A publicly accessible API. Anyone able to pay the vendor's list price can obtain the same endpoint we evaluated. Private previews, allow-listed betas and negotiated research access do not qualify, because a score nobody else can reproduce is a claim rather than a measurement; or
- Widely used open weights. The weights are published under a licence permitting third-party evaluation, and the model is in substantive use. "Widely used" is a judgement, exercised conservatively: we would rather omit a model and be asked to add it than pad a roster.
Where eligibility turns on that judgement, the reason is recorded with the entry or the decision. The judgement is not exercised to favour or disfavour any vendor, and a vendor's commercial relationship with us is not an input to it.
1.2 Zero data retention is guaranteed for the endpoint
A run is refused outright where the provider cannot guarantee zero data retention for the endpoint being evaluated. This is not negotiable and is not a scoring penalty: the run does not happen.
The reason is the item bank. Our questions are held out and unpublished; an endpoint that retains prompts is an endpoint that accumulates our test set, and a benchmark that leaks its own questions to the systems it measures has destroyed the thing it sells.
1.3 The submission is what it claims to be
A declaration that resells another provider's model unchanged is refused. A system built on a base model (retrieval, prompting, tool use) may be evaluated, but must describe what it adds and is reviewed by a person before it runs. The entry names what was measured.
1.4 It is not a first-party model
viggoVet products are out of scope for the public leaderboard entirely. No first-party model is ranked, labelled or listed. First-party validation, if published at all, is published separately as viggoVet research and never as a leaderboard entry, and is never presented as a VetEval rank, rating, or tier.
This is a stricter rule than labelling a first-party entry and ranking it by the same method. We prefer the stricter rule because it removes the conflict rather than managing it, and because a separation the reader has to take on trust is worth less than one that has nothing to trust.
1.5 Evaluating the model is permitted
Before an evaluation begins, we confirm that the terms governing the endpoint or the weights permit third-party evaluation and publication of results. Where those terms prohibit benchmarking or the publication of benchmark results, we do not run the evaluation, whether or not the model is otherwise public. This check is recorded per run.
2. How models are added
2.1 Two routes, one method
- Vendor submission. Any organisation may submit a model that meets Section 1.
- VetEval-initiated evaluation. We evaluate publicly available models on our own initiative, subject to Section 1.5. A vendor's preference does not determine whether a public model is measured; if it did, the roster would be a list of vendors who expected to do well.
Both routes run the identical method under identical conditions. Neither is marked as more or less authoritative on the board, because the method is what makes a number comparable and the method does not differ.
We notify a vendor when we publish a VetEval-initiated result about its model, and the right of reply in the Publication, Corrections & Independence Policy applies equally to results the vendor did not ask for.
2.2 Cadence for new frontier releases
We add newly released models that meet Section 1 as promptly as review capacity allows. We do not commit to a fixed period.
Any timing statement is about starting an evaluation, never about a publication date. A run that surfaces a data problem is delayed rather than published on schedule.
"Release" means general availability on the vendor's own terms: an endpoint or weights anyone can obtain at list price or under a public licence. Announcements, waitlists and limited previews do not start any clock.
2.3 What we do not do
We do not add a model because a vendor asked for coverage in exchange for anything, and we do not decline to add one for the same reason. Plan and evaluation pricing is uniform and published; it covers grading, review and platform capacity. Inference is billed to the vendor by its own provider, because the vendor supplies the endpoint and key. Publication is score-blind.
3. Re-runs: what triggers a fresh evaluation
A published score describes one model, measured once, under one configuration. It is re-measured when any of the following occurs.
A vendor may also purchase additional evaluations at the published rate. Those follow the publication rules in Section 4, which is what stops a re-run from being a way to shop for a better number.
4. Which score is published
A published score is not a vendor's best attempt.
- Publication is chosen before a run starts, not after the score is known. A run not declared for publication up front cannot be published at all.
- A model name is benchmarked once. A second publication under the same name is refused rather than superseding the first, so running repeatedly cannot select a favourable result. An improved model is submitted under a new name, and is evaluated on its own merits.
- Private evaluations stay private, including how many were run. We do not publish an attempt count.
These rules apply identically to vendor-submitted and VetEval-initiated evaluations.
5. Retirement and supersession
- A model superseded by a newer version is retired at the next calibration, not immediately. Retirement removes it from the default board view.
- Records are preserved. A retired entry keeps its score, its rating and the calibration that produced it, its interval, its safety status, its date and its dataset version, and remains reachable at its permanent URL. Nothing is deleted. A citation to a VetEval result must not rot.
- A model whose endpoint is withdrawn by its provider is retired on the same terms.
- A model that ceases to meet Section 1 (for example, an endpoint that stops guaranteeing zero data retention) is retired from the default view and the reason is stated.
Why records persist. A published result is a measurement about a company's product, not personal data about an individual, and it is retained as a permanent part of the public record so that citations remain verifiable. We retain it on that basis and for the establishment and defence of legal claims. Where a page incidentally contains personal data of an individual, that data is handled under the Privacy Policy and can be addressed without disturbing the measurement.
6. Safety findings
A confirmed high or critical severity finding is reviewed by licensed veterinarians before it counts toward a model's gate status.
Alongside gate status we publish each model's confirmed-harm rate per 1,000 items administered, and the counts behind it. The gate keys off the worst single confirmed harm in a finite set of questions, so it is deliberately blunt and unavoidably lumpy; the rate is what distinguishes one bad answer from a pattern.
A vendor that believes it has remediated a confirmed finding may request a re-run under Section 3. A re-run measures the model as it then stands; it does not remove or amend the earlier published result, which continues to describe the model as measured on its date.
No model is described as unsafe. What is published is that a model produced answers a panel of licensed veterinarians confirmed as harmful on specific items, at a stated severity and rate, in a stated round.
7. Changes to this policy
This document is versioned. Material changes are dated and the previous version remains available, for the same reason results are: a policy that can be revised silently is not a policy.
Any change to Section 1.4 would be published as a dated amendment before any first-party model appeared in any form, never alongside or after it.