The item bank

Who writes the questions VetEval scores against, and how they are reviewed.

Every question VetEval scores against is written and reviewed by licensed veterinarians. No question is authored by a language model, models are respondents here, never authors, because a benchmark whose questions came from a model measures agreement with that model rather than veterinary medicine.

Review

Items are authored and reviewed by the veterinary team before they enter a scored run.

What is published

The bank itself is held out and is not published, a benchmark whose answers are public stops measuring anything the moment a model trains on it. A small practice sample is available below so you can see the shape and difficulty of the questions without seeing the set a score was measured against.

The separation between the two is enforced by the database rather than by policy: an item released for practice cannot enter a scored version, and an item that has been scored cannot be released.

Practice sample

A set of practice items is available to signed-in accounts. These items score nothing: an item released for practice cannot enter a scored run, and an item that has been scored cannot be released. That split is held by the database, not by policy.

Sign in to downloadFree. An account is only so we know who has the file.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.