Benchmarks decide purchases now. A leaderboard screenshot shows up in a sales deck, a percentage gets quoted in a demo, and somewhere a clinic signs. Most leaderboard readers, through no fault of their own, have never been told what separates a measurement from a marketing asset, because the people best positioned to explain it usually run benchmarks and prefer not to arm their critics.
We run one, and we would rather arm you. Here are the six questions we would put to any AI benchmark, including this one, each with what good looks like, the red flags, and the failure it catches. A printable checklist and a copy-paste vendor email follow at the end.
1. Who wrote the questions, and can the models have seen them?
The failure this catches: memorization wearing the costume of knowledge. A benchmark scraped from the public internet is, sooner or later, in the training data of the models it measures, at which point high scores measure recall of the test rather than competence in the domain. The contamination literature's cleanest demonstration built a private twin of a famous math benchmark, matched in style and difficulty, and watched several well-known models score conspicuously worse on questions they could not have memorized. Paraphrase-based contamination evades the string-matching filters most benchmarks rely on, so "we deduplicated against the training data" is not the reassurance it sounds like.
What good looks like: items authored by domain experts rather than scraped; a genuinely private held-out set that never publishes; a plan for how exposure to vendor endpoints is managed across time, since a private set served identically to everyone forever stops being private in the way that matters; and per-model contamination signals published, such as the performance gap between public and private items.
Red flags: questions sourced from past public exams or forums; no public/private split at all; a fixed test set that has been served unchanged for years; contamination addressed with one sentence of reassurance and zero mechanism.
2. Does the score come with an interval, and are ranks allowed to tie?
The failure this catches: decimal theater. A leaderboard listing 84.2 above 83.9 is asserting that 0.3 points is signal. On a finite test it usually is not, and on a sampled test it is noise with a typeface. The deeper issue is multiplicity: compare enough models across enough categories and "significant" differences manufacture themselves by chance unless the statistics account for it.
What good looks like: confidence intervals on every published score; paired statistical tests for head-to-head claims; ranks published as bands with honest ties when models cannot be separated; and the number of items per evaluation justified by a reliability analysis rather than by tradition.
Red flags: every model has its own unique integer rank; no interval anywhere; sub-point gaps narrated as meaningful; "we ran it once."
3. What happens when a model is dangerous rather than wrong?
The failure this catches: averages laundering harm. A model can miss 5% of questions in ways that are embarrassing, or in ways that are lethal, and a single accuracy number cannot tell you which, by construction. In any medical domain this is the difference that matters most, and it is the one a composite score is structurally blind to.
What good looks like: harm measured on its own axis with published severity definitions; candidate harms reviewed by qualified humans rather than only by another model; a scoring consequence that actually bites, such as a gate or cap, rather than a footnote; and refusal behavior priced explicitly, since silence on hazard questions is not evidence of safety.
Red flags: "accuracy: 96%" with no harm dimension anywhere; safety findings adjudicated solely by an LLM judge; a safety score that averages into the composite so a model can buy back a lethal answer with trivia.
4. Who graded the open-ended answers?
The failure this catches: measuring the judge's taste. Open-ended clinical answers need grading, grading at scale means an LLM judge, and LLM judges carry documented biases: they favor their own model family's outputs, favor longer answers, and favor familiar formats. A benchmark that never validated its judge against human experts has published the judge's preferences with the models' names attached.
What good looks like: models never graded by their own family; the judge validated against expert human labels with a published, chance-corrected agreement statistic and a stated threshold below which results do not publish; structured rubrics rather than vibes; grading order randomized or swapped to kill position bias; and cross-judge sensitivity disclosed.
Red flags: the judge model is unnamed; the judge shares a family with leaderboard entrants; "agreement with humans" cited with no statistic, no threshold, and no consequence for failing it.
5. Can a vendor retry until the number looks good?
The failure this catches: publication bias at the leaderboard scale. Private evaluation runs are legitimate and useful; publishing the best of seven private attempts is marketing with a methods section. The subtler version is financial: if vendors can pay for placement, badges, or favorable treatment, the instrument is over regardless of the math.
What good looks like: a published rule that a public score must be the most recent completed evaluation, never a selected earlier attempt; disclosure of how many evaluations a vendor ran; fees, if any, at uniform published rates covering compute and grading, with publication free and score-blind; and a plain statement of who funds the benchmark.
Red flags: vendors control which run publishes with no recency rule; "featured" placements; evaluation fees that scale with outcomes; sponsorship from ranked vendors without disclosure.
6. Does the benchmark saturate, and what does the ceiling mean?
The failure this catches: celebrating a broken ruler. When several models cluster at 96 to 99 percent, the test has stopped measuring the frontier, and the remaining gaps are rounding noise wearing medals. Saturation is not a scandal, it is a lifecycle stage; the scandal is pretending it has not happened.
What good looks like: a stated plan for the ceiling: harder item ladders, refreshed banks, or an anchored rating scale that keeps discriminating as raw percentages compress; and honesty about what a perfect score means: a fact about a paper, never a license to practice.
Red flags: a two-year-old leaderboard where the top ten sit within two points; press releases about a model "acing" the benchmark from the benchmark itself; no mechanism anywhere for the test to get harder.
The one-page checklist
| Question | What good looks like | Red flag | |---|---|---| | Item provenance | Expert-authored, private held-out set, exposure managed, contamination signals published | Scraped questions, no split, fixed set served for years | | Uncertainty | CIs everywhere, paired tests, honest tie bands | Unique decimal ranks, no intervals | | Harm | Own axis, human-confirmed, gate that bites, refusals priced | Single accuracy number, judge-only safety, harm averaged away | | Grading | No self-family grading, judge validated vs experts with a threshold | Unnamed judge, no validation statistic | | Publication | Most-recent-run rule, attempt counts disclosed, uniform published fees | Best-of-N publishing, paid placement | | Saturation | Ceiling plan, anchored scale, honest framing of 100% | Frontier clustered at 97+, no refresh mechanism |
Five questions to email any benchmark or vendor
Copy, paste, and require written answers: 1) Who authored your test items, and what prevents models from having trained on them? 2) What is the confidence interval on the score you are quoting, and against which alternatives is the difference statistically significant? 3) How do you test for harmful answers, who reviews them, and what does a confirmed harm cost in the score? 4) Was the published number the most recent evaluation or a selected one, and how many were run? 5) Who pays you, for what, and does any payment touch scores, placement, or badges? A source that cannot answer all five in writing has answered them.
How this applies to VetEval
Our own answers to all six questions live on the methodology page with the math shown: expert authorship with a private bank and per-evaluation sampling, intervals with FDR-controlled tie bands, a human-confirmed safety gate with refusals priced as unanswered, answers scored against a fixed key rather than by a model, a most-recent-evaluation publication rule with attempt disclosure and uniform published fees, and an anchored rating built for the day raw percentages saturate. Hold us to them, and if you find a gap between this guide and our practice, the contact form reaches the people who can fix it.
Common questions
Is a high benchmark score meaningless, then? No. A well-built benchmark score is real evidence about model knowledge. The guide exists because the difference between well-built and decorative is invisible in a screenshot, and screenshots are how scores travel.
Do these six questions apply outside veterinary medicine? Entirely. They are domain-agnostic; only the harm definitions change. Any medical, legal, or financial AI benchmark should survive the same list.
What should a clinic do if a vendor quotes a benchmark this guide flags? Ask the five email questions and weigh the silence. A flagged benchmark does not prove a bad product; it proves the quoted number is not the evidence it is dressed as, so require better evidence.
Why would a benchmark publish a guide that can be used against itself? Because an instrument that fears audit is not an instrument, and because every reader this guide sharpens makes decorative leaderboards less profitable, which is the outcome we are structurally betting on.
Data and reproducibility
Contamination, judge-bias, and private-twin studies referenced above are cited in full, with links, in the VetEval paper's reference list and on the methodology page.