← Leaderboard

Claude Opus 5.5

Anthropic · proprietary
Evaluated 2026-09-25
Evaluation certificate (PDF)
VetEval Index

Rank 1–2 of 18 · 2 models too close to separate.

Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →

Safety

Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.

Raw index 87.4, reduced by safety gate ×0.40, consistency ×0.94, safety coverage ×0.99. Published index 32.5.

Confirmed harms: 4.72 per 1,000 answers

Safety is a multiplicative gate, not just a component. How this is measured →

SUMMARISE WITH
SHARE

Per-category profile

Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.

Clinical Diagnosis & Differential
90.5
Pharmacology & Dosing
86.2
Species-Specific Knowledge
81.6
Triage & Emergency
91.8
Preventive / Public Health / Zoonoses
83.1
Communication & Professionalism
90.4
Safety & Harm Avoidance
94.4
View as data table
Per-category scores with 95% confidence intervals
CategoryScore95% CI
Clinical Diagnosis & Differential90.586.9–94.0
Pharmacology & Dosing86.282.1–90.2
Species-Specific Knowledge81.675.8–87.5
Triage & Emergency91.887.7–95.9
Preventive / Public Health / Zoonoses83.176.6–89.6
Communication & Professionalism90.483.7–97.1
Safety & Harm Avoidance94.490.3–98.5

Cost · speed · reliability

Cost / case
$0.0106
the vendor's published price applied to this run's tokens
Median time / question
3.2 s
excluding retries and queueing
Reliability
100.0%
questions that came back with a usable answer

How it compares

Claude Opus 5.5 against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.

Claude Opus 5.5 against the field

The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.

View as a table
VetEval Index for every published model
ModelVendorVetEval Index95% interval
Claude Opus 5.5Anthropic32.531.7 to 33.4
Claude Opus 5Anthropic31.130.2 to 31.9
GPT 5.6 SolOpenAI30.829.8 to 31.6
Gemini 3.7 FlashGoogle28.227.2 to 29.1
GPT-6 SolOpenAI27.126.2 to 28.0
Kimi K3Moonshot AI27.126.3 to 27.8
GPT-5.6 TerraOpenAI27.026.1 to 27.9
GLM 5.3Z.AI25.224.3 to 26.1
GPT 6 LunaOpenAI25.024.1 to 25.8
Gemini 3.8 FlashGoogle24.123.3 to 25.0
Grok 4.6SpaceXAI23.122.3 to 23.9
GLM 5.3 FlashZ.AI23.022.3 to 23.8
Claude Sonnet 5Anthropic22.321.4 to 23.1
Llama 4 MaverickMeta19.018.1 to 19.8
Minimax M3Minimax18.117.4 to 18.9
Qwen3.8 2.4T A95BAlibaba10.910.4 to 11.3
Nemotron 3.5 LightningNVIDIA3.73.3 to 4.2
Granite 4.2 8bIBM2.32.1 to 2.5

Reproducibility

Model id
anthropic/claude-opus-5.5
Generate config
temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
Pricing ($/M)
input_per_m=4 · output_per_m=20

Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.