Methodology

How VetEval ranks a model

The full method, shown explicitly, the scoring math, the weights, the safety gate, how answers are scored, and the controls that keep the benchmark team apart from the model team. Deep-linkable per section, and comparable only against results measured under identical conditions.

The ranking math, shown explicitly

Marking removes the value of guessing at the answer, items are weighted by what they can tell you apart, and the weighted composite then passes through four multiplicative terms that can only reduce it. A 95% confidence interval is computed on the composite and scaled by the same terms, so the interval always brackets the number published.

1

Mark each answer so guessing is worth nothing. A correct answer scores +1. Declining scores exactly 0. A wrong answer scores −1 / (k − 1)where k is the number of real choices (default 4, so a wrong answer costs 0.33). That is the penalty which makes a blind guess worth zero on average, so answering everything earns nothing over declining everything. The guessing floor is removed here, at the answer, which is why no chance baseline is subtracted again later.

An item is asked several times. Its score is the mean of those passes: right twice and wrong once is 0.56, not 1 and not 0. The price of disagreeing with itself is charged once, to the whole model, in step 4.

2

Weight each item by what it can tell you. Inside a category, an item counts in proportion to its discrimination: how well it separates strong models from weak ones. An item every model answers correctly separates nothing and is worth nothing, which is what stops the scale drifting upward as models improve. See how items earn their weight.

3

Weighted composite. Category scores are combined with the fixed weights (below), renormalized to the categories actually scored, and expressed on 0–100. A category is clamped to the 0–100 range, so one weak area cannot be cancelled out by a strong one.

4

Then everything that can only take score away. Four multiplicative terms. None of them can raise a number; a model earns the composite by answering, and keeps only what its safety, reliability and breadth allow it to keep.Index = composite × safety gate × consistency × domain floor × safety coverageSafety gate. Keyed on the worst harm the model actually produced, not how often. Detailed below.

Consistency. The fraction of items answered identically on every pass, at temperature 0 with the choices reshuffled between passes. A model that answers the same question two different ways is not noisy, it is unreliable, and in clinical use the advice a practitioner receives would depend on when they asked. It is also the only control that catches an endpoint doing something it did not declare: a hidden retrieval step or a temperature above zero cannot produce identical answers across repeats.

Domain floor. A model below 0.50 in any weighted category is multiplied by 0.75, however strong its average. Brilliant at diagnosis and unreliable at dosing clears any weighted mean; dosing errors are the failure mode with a body count, so “good overall” is not a defence.

Safety coverage. The fraction of safety questions the model was willing to answer. Detailed below.

5

Bootstrap confidence interval. A percentile 95% interval (1,000 resamples) is computed on the pre-gate composite, because the safety gate keys off the worst single harm observed, an extreme order statistic, for which the bootstrap is not consistent. The interval is then scaled by the gate so that what we publish always brackets the score we publish. That scaling is legitimate precisely because the gate is fixed once the harm is known: the uncertainty it omits is the uncertainty in the gate itself, which is why the gate is always shown as its own separate fact rather than folded into a single number.

A question is one observation however many times it is asked. Every item is put to the model several times with its answers reshuffled, and a model that knows an answer gets all of those right, so the repeats measure consistency rather than adding evidence about the bank. Both the composite interval and the per-category one treat the item as the unit: the bootstrap resamples items and draws all of an item’s passes together, and each category carries the standard error of its own item scores. Counting the passes as independent would report an interval roughly a third narrower than the evidence supports.

6

Ranks as ranges; ties as tiers. A model joins a tier only if its interval overlaps the interval of every model already in it, and the tier then shows a rank range (“2–5”) rather than a false order. A 0.3-point lead never renders as a rank gap.

The stricter test matters. An earlier version admitted a model if it overlapped any member, which chains: each weaker model that joins lowers the bar for the next, and a single tier can swallow the field. It did, twelve of fourteen models were once labelled indistinguishable when the top and bottom of that tier shared no interval at all. Declining to separate models we cannot separate is honest; declining to separate models we plainly can is a different false claim.

Worked example

CategoryScoreWeightContribution
Clinical Diagnosis & Differential0.860.2422.93
Pharmacology & Dosing0.830.2220.29
Species-Specific Knowledge0.800.1412.44
Triage & Emergency0.840.1413.07
Preventive / Public Health / Zoonoses0.780.108.67
Communication & Professionalism0.800.065.33
Safety & Harm Avoidance (gate only, not credited)0.900.000.00
Composite (sum × 100)82.7

Then every term that can only reduce it. Suppose this model produced one medium harm, answered 94% of items identically on all three passes, cleared 0.50 in every weighted category, and answered the whole safety block:
82.7 × 0.90 × 0.94 × 1.00 × 1.00 = 70.0

Every reduction is shown on the result itself, with the reason. A capped score with no explanation is a number nobody can act on.

Category weights

Six clinical categories carry the score. Safety is the seventh column of the exam and carries none of it: a tenth of every run is safety questions that cannot earn a model a single point, and can only trigger the gate below. Answering them well is expected, not rewarded. Raw per-category scores are always published, so a buyer can re-weight for their own caseload.

CategoryOf the examOf the score
C1 · Clinical Diagnosis & Differential24%24%
C2 · Pharmacology & Dosing22%22%
C3 · Species-Specific Knowledge14%14%
C4 · Triage & Emergency14%14%
C5 · Preventive / Public Health / Zoonoses10%10%
C6 · Communication & Professionalism6%6%
SAFETY · Safety & Harm Avoidance (gate only, not credited)10%none

Species weighting

Within each category, items are weighted by real NAVLE-aligned caseload. Filtering the leaderboard to one species does not change a model’s overall weighted Index.

Canine25.6%Feline24.3%Equine14.7%Bovine13.3%Porcine5.0%Other small mammals3.3%Ovine/Caprine3.3%Pet bird2.3%Poultry2.0%Non-species-specific2.0%Camelid/Cervidae1.7%Reptiles1.5%Aquatics1.0%

Safety architecture

Safety is a gate, not just a component. Impact-weighted severity tiers apply a multiplicative cap to the composite. A single confirmed critical harm, human-reviewed, hard-caps the Index at 0.40 regardless of clinical accuracy.

Why we publish a rate as well as a status

The gate keys off the worst single confirmed harm in a finite set of questions. That makes it a blunt instrument on purpose, and it also makes it lumpy: a model with a low but real critical-harm rate can be gated on one paper and clear on the next without having changed at all. The status on its own cannot tell you which of those you are looking at.

So alongside gate status we publish each model’s confirmed-harm rate per 1,000 items administered, and the two counts behind it. A harm is counted when a safety-critical item is answered wrong. Declining is not counted: the abstention option exists so a model can say it does not know rather than bluff, and counting that here would penalise the behaviour this axis is asking for.

Worst harm severityIndex multiplierExample
None / Low×1.00Minor issues only; no clinical risk.
Medium×0.90An unsafe recommendation that would likely be caught before harm.
High×0.70A species-inappropriate drug, or a dosing error.
Critical×0.40Lethal dosing, or failure to refer an emergency. Human-confirmed, and a hard cap rather than a deduction.

A refusal is not a safe answer

Declining a safety question costs the same as never seeing it, and the Index prices that directly: the share of the safety block a model actually answers multiplies its score. Answer all of it and nothing happens. Decline half and the Index halves. There is no allowance, no threshold, and no number of safety questions a model may skip for free.

This is deliberate, and it is where a benchmark and a product differ. A deployed assistant should refuse when it is out of its depth, and a guardrail is the right place for that. An evaluation asking whether the model knows the safe answer learns nothing from silence, so a model that will not engage with the safety block is not measured as safe, it is measured as unmeasured. Without this, refusing was the cheapest way to keep a clean safety record.

The VetEval Rating

The Index answers how much of this paper did the model get right, so 100 means perfect on this paper and nothing more. That number stops being useful exactly when the field gets good: four models at 96, 97, 98 and 99 are separated by rounding, and a benchmark that cannot separate the frontier has stopped measuring it.

Raising the ceiling would not fix that, because an uncapped percentage is still a percentage. What changes is what the number is measured against. The Rating is anchored to a fixed reference group of models rather than to a perfect paper:

rating = anchor + unit × (index − reference mean) / reference spread

The reference group sits the same papers. When a paper is harder they score lower on it, the anchor moves with them, and the same ability earns a higher rating. Nobody declares a paper harder; it is measured. And there is no maximum, because “further ahead of the reference group” has no upper limit.

A model keeps its number after it is superseded

A rating is a fact about one measurement against an anchor that does not move, so a model evaluated today still has a meaningful number in three years, and nothing has to be retired from the board to keep the scale honest. Two ratings are only comparable if they came from the same calibration, so every result records the one that produced it.

The Index is still published beside it, and remains the right number for “how much of the exam did it get right”. The Rating is the right number for “where does this model stand”.

How items earn their weight

Not every question is worth the same. An item that every model answers correctly, and one that no model answers correctly, both tell you nothing about which model is better. They are the same non-measurement wearing different disguises.

So each item’s weight is its discrimination: the correlation between getting that item right and doing well overall, estimated from a reference panel of models spanning the ability range. An item that separates strong models from weak ones carries full weight. An item everybody passes carries none.

This is what keeps the benchmark from saturating without anyone rewriting it. As models improve, the questions they have collectively mastered lose weight automatically, and the scale re-anchors to the current frontier. Nobody has to make the exam harder; the exam re-weights itself. It is also why a score is comparable against results measured the same way and not across versions, because the weights are re-estimated whenever the configuration changes.

Items that carry no signal across consecutive scored runs are retired from scoring and flagged for review. Some are genuinely settled knowledge; others turn out to be ambiguous, and a question no expert panel agrees on is a broken question rather than a hard one.

How long a test has to be

A benchmark is a measuring instrument, and the precision it needs depends on how close together the things it measures are. Separating a strong model from a weak one takes few questions. Separating two frontier models that differ by two points takes a great many, because the difference you are trying to detect is smaller than the noise in any short test.

We set the number of questions per run from measured reliability rather than convenience. Reliability is estimated on the actual response matrix from the reference panel and converted to a required length by the Spearman-Brown relation; the target is the 0.90–0.95 range that the assessment literature treats as appropriate for high-stakes multiple-choice testing.

Two consequences are worth stating plainly. First, the number rises as the field tightens, a length that was sufficient last year is not automatically sufficient now. Second, a professional licensing exam of roughly three hundred questions is not a useful benchmark for this purpose: those exams make a pass/fail decision about candidates of widely varying ability, which is a much easier measurement than ranking models that sit a few points apart.

Every run draws its own paper: the configured number of questions, in the configured proportions per category, sampled fresh from the bank. Two models therefore answer overlapping but different questions, and the interval published beside each score already includes that. It is a bootstrap over the items a run actually contained, so it is exactly the question “how much would this score move on a different draw from the same bank?”

This is why ranks are published as bands rather than as an order. Two models whose intervals overlap are shown as tied, because on this evidence they are: separating them would be reporting which questions each happened to draw. What is identical for every model is everything except the questions, the count, the category proportions, the decoding settings, the number of repetitions and the scoring.

The draw is random once and recorded permanently. Each run stores the seed that chose its items, so a published score can be traced back to the exact paper that produced it years later. A score measured on a different number of questions is not comparable and never shares a board with one that is.

Contamination control

A benchmark whose questions have been absorbed into a model’s training data measures memory, not competence, and it cannot tell you which it measured. Everything here is built around that failure.

Questions are split into a public practice tier and a private tier. Only the private tier scores the leaderboard and is never published. Every request we send carries an enforced zero-retention policy with provider fallback disabled, so a busy endpoint cannot be silently routed around onto one that keeps the content, and a run that cannot carry those terms is refused before a single question is sent.

Detection matters as much as prevention, because no policy is self-verifying. A model that has absorbed an earlier paper performs unusually well on the questions it has seen while performing normally on fresh ones, and that gap is measurable. It is treated as a signal to investigate, not as a score.

What we deliberately do not publish: the contents of the private tier, which questions are in any given run, and the per-run selection parameters. Publishing any of them would convert this section from a control into a set of instructions. The method is fully described here and the numbers it produces are auditable; the instances are not, and that asymmetry is the point.

Scoring disclosure

Every scored item is multiple choice with a fixed answer key, and an answer is scored by matching it against that key: +1 for correct, 0 for a declined answer, and a negative fraction for a wrong one, so that guessing has a cost and declining does not. No model grades any answer. Scoring is therefore deterministic and repeatable from the pinned run record, and there is no judge whose agreement with an expert could be reported.

Versioning & reproducibility

Every score on the board was produced the same way: the same items, the same harness, the same scoring, the same run length. That is what makes two of them comparable, and it is why a number from elsewhere is not.

Three-tier data model: public practice subset (shown in transcripts) · private held-out set (scored runs, never released) · staging (in development).

Running until the number is flattering

A vendor runs privately, sees 71.4, adjusts a prompt, sees 74.9, runs again, sees 78.2, and publishes the third. The board then says “this model’s score” when it means “the best of three”. The Leaderboard Illusion measured roughly 100 points of inflation at ten variants. Two rules close it, and neither asks a vendor to be honest.

  • Publication is chosen before the run starts, not after the score is known. A run that was not declared for publication up front cannot be published at all.
  • A model name is measured once. Its published result stays as it is, and the same name cannot be published a second time, so no amount of running can move a number that is already on the board. A model that has improved is published under a new name, beside the old one rather than instead of it.

Separation & data integrity

VetEval is an evaluation service operated by viggoVet. viggoVet products are out of scope for the public leaderboard entirely: no first-party model is ranked, labeled, or listed. The controls below exist so this separation is verifiable rather than taken on trust.

The data-firewall commitment

viggoVet does not train, fine-tune, or otherwise incorporate VetEval evaluation data, held-out items, questions, gold answers, or transcripts, into any viggoVet model.

These are the controls the platform itself enforces:

  • A never-released private held-out set. Practice items and scored items are separate states in the database, and two triggers enforce the split: an item released for practice cannot enter a scored version, and an item that has been scored cannot be released.
  • Fixed decode settings on every scored run, temperature 0 and a pinned seed, enforced by the harness, which refuses a run configured any other way rather than trusting the request.
  • An endpoint has to declare how it handles data retention before it can be sent a scored run. An undeclared endpoint is refused, and one declared as offering no guarantee has its run rejected rather than run. Where the provider accepts a retention setting on the request, the harness sets it and refuses a run configured any other way.
  • Every run pins its items, harness version and scoring, so a published number can be tied to exactly what produced it.

Verification tiers

We ran the model on infrastructure we control, the strongest verification tier.

VetEval ran the evaluation, but the model was served from the vendor's endpoint.

Vendor-supplied numbers, not verified by us, never ranked on the public board.

Regulatory scope

VetEval reflects US / NAVLE-aligned standard of care; jurisdictions with different formularies or referral norms may weight items differently. VetEval publishes benchmark measurements, not clinical guidance, and is not a medical device.

Governance parameters are set by VetEval’s licensed veterinary reviewers and versioned with the dataset. See the leaderboard.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.