Publication, Corrections & Independence Policy
In effect from 1 June 2026.
This Policy explains what VetEval publishes, on what basis, how we handle disputes and corrections, and how we manage the conflicts of interest that come from being operated by a company that also builds veterinary AI. It supplements our Terms of Service and the Methodology.
1. What we publish, and on what basis
We publish measurements: index scores, confidence intervals, rank tiers, category and species breakdowns, safety-gate results, verification tiers, attempt disclosures, and analysis derived from them. Every published number is produced by the method described in the Methodology, on pinned items, harness, and scoring, so a published figure can be tied to exactly what produced it.
Findings are honest evaluation, stated on a disclosed basis. Where we characterise a model's behaviour, including a safety-gate severity, we do so as our assessment of observed model outputs under stated test conditions, on a published methodology, in good faith, and in the public interest of buyers and the veterinary profession. We do not state or imply that a vendor acted wrongfully, and a safety finding describes what a model produced on our items in that run rather than a general claim about the product or the company.
2. Before we publish
- Human review of safety findings. A critical or high safety-gate result is human-confirmed before it can affect a published index.
- Deterministic scoring. Every scored item is multiple choice with a fixed answer key, and no model grades any answer, so a published score can be recomputed from the pinned run record.
- Community review. Community results are checked by our team before they reach the board.
- Notice. We notify the submitting organisation when a result is approved for publication.
3. Permanence, and what we will not do
Community results are permanent. If a model is submitted on a Community plan, the result publishes in full and stays published. We will not remove it because it is unfavourable, because the vendor has withdrawn the model, because a newer version exists, or because the account was closed. Permanence is what makes the board citable; a benchmark you can quietly exit is not a benchmark.
We will not: change a score in exchange for payment or any other consideration; suppress, delay, or de-rank a published result at a vendor's request; remove a result because a vendor threatens litigation, where the result is accurate; or publish a result we know to be wrong.
4. Corrections, annotations, and withdrawal
We distinguish between fixing an error and hiding a result. The table sets out what we do.
Situation
What we do
Public record
Platform or harness error, wrong settings, mis-scoring
Re-run or recompute, correct the figure
Corrected, with a dated correction note
Scoring or item defect found later
Re-grade or retire the item; recompute affected results
Corrected, with a note; item flagged for review
Wrong endpoint or model version evaluated (our error)
Re-run at no cost to the vendor
Superseded entry annotated
Wrong endpoint or version supplied by the vendor
Vendor may resubmit under normal plan limits
Original stands, annotated
Vendor disputes an accurate result
Right of reply published alongside
Result stands
Contamination or manipulation found
Disqualify or annotate the result
Annotated on the record
Legal requirement to remove
Comply, narrowly
Removal noted where lawful
Corrections are dated and visible. We do not silently edit a published number.
5. Your right of reply, and how to raise a dispute
If you are a model provider and believe a published result is inaccurate, was produced under a configuration error, or is materially misleading, contact us through our contact form with the specific figure, the run, and the basis for your concern. Our process:
- We acknowledge receipt and tell you who is handling it.
- We review the pinned run record: items, harness version, scoring, settings, and logs.
- Where the concern is substantiated, we re-run, recompute, or correct, and publish a dated correction note.
- Where the result stands, we tell you why, and you may submit a concise statement of reply, which we publish alongside the result.
- Where we and you continue to disagree, the disagreement itself is noted on the record.
We aim to acknowledge promptly and to resolve substantiated issues quickly. Safety-related concerns are prioritised.
6. Independence and conflicts of interest
We state the conflict plainly rather than claiming to be an unrelated third party. VetEval is operated by viggoVet, which also builds veterinary AI, and it earns revenue from some of the organisations whose models it ranks. Both facts are managed by structural controls and disclosure, not by assertion.
- First-party models are never ranked. viggoVet products do not appear on the public leaderboard. They are validated separately and published as viggoVet research.
- Team separation. Controls keep the benchmark team apart from the model team, as described in the Methodology.
- Data firewall. viggoVet does not train, fine-tune, or otherwise incorporate VetEval evaluation data, held-out items, questions, gold answers, or transcripts into any viggoVet model.
- Platform-enforced integrity. Practice and scored items are separate database states with triggers preventing crossover; scored runs use fixed decode settings enforced by the harness; an endpoint whose data retention was never declared is refused a scored run; every run pins its items, harness, and scoring.
- Governance. Governance parameters are set by VetEval's licensed veterinary reviewers and versioned with the dataset.
7. What payment does and does not buy
Payment can buy
Payment can never buy
Capacity: more models, more evaluations
Rank, or any change to a score
Privacy: seeing your score before publication
Removal of a published Community result
Speed: publishing without waiting for review
Extra attempts on one submission (capped by methodology)
Choice: deciding whether a Lab result publishes
Secrecy about how many attempts preceded a result
Standard plan pricing is uniform and published to every customer, and we do not negotiate the price of a standard plan, because we rank the organisations that pay us.
8. Use of third-party names and marks
We identify models, vendors, and their marks nominatively, only to say what was evaluated, using no more of a mark than needed and never in a way that suggests affiliation, sponsorship, or endorsement in either direction.
9. Changes
We may update this Policy and will publish the updated version with a new date. Material changes to how we publish, correct, or disclose conflicts will be described in the changelog.