Claude Sonnet 5
Built from open weights we did not train (fine-tune, distillation, merge)Built on claude-sonnet-5Zero data retention
Rank 5–7 of 11 · 3 models too close to separate.
Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →
Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.
Raw index 69.5, reduced by safety gate ×0.40, consistency ×0.81, safety coverage ×0.99. Published index 22.3.
Confirmed harms: 11.11 per 1,000 answers
Safety is a multiplicative gate, not just a component. How this is measured →
Per-category profile
Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.
View as data table
| Category | Score | 95% CI |
|---|---|---|
| Clinical Diagnosis & Differential | 73.2 | 68.1–78.2 |
| Pharmacology & Dosing | 68.2 | 62.6–73.8 |
| Species-Specific Knowledge | 62.0 | 54.6–69.4 |
| Triage & Emergency | 76.6 | 70.1–83.0 |
| Preventive / Public Health / Zoonoses | 57.7 | 48.4–67.0 |
| Communication & Professionalism | 80.1 | 71.5–88.6 |
| Safety & Harm Avoidance | 92.4 | 87.2–97.5 |
Cost · speed · reliability
How it compares
Claude Sonnet 5 against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.
Claude Sonnet 5 against the field
The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.
View as a table
| Model | Vendor | VetEval Index | 95% interval |
|---|---|---|---|
| Gemini 3.7 Flash | 28.2 | 27.2 to 29.1 | |
| Kimi K3 | Moonshot AI | 27.1 | 26.3 to 27.8 |
| GPT-5.6 Terra | OpenAI | 27.0 | 26.1 to 27.9 |
| GLM 5.3 | Z.AI | 25.2 | 24.3 to 26.1 |
| Grok 4.6 | SpaceXAI | 23.1 | 22.3 to 23.9 |
| GLM 5.3 Flash | Z.AI | 23.0 | 22.3 to 23.8 |
| Claude Sonnet 5 | Anthropic | 22.3 | 21.4 to 23.1 |
| Llama 4 Maverick | Meta | 19.0 | 18.1 to 19.8 |
| Minimax M3 | Minimax | 18.1 | 17.4 to 18.9 |
| Qwen3.8 2.4T A95B | Alibaba | 10.9 | 10.4 to 11.3 |
| Nemotron 3.5 Lightning | NVIDIA | 3.7 | 3.3 to 4.2 |
Reproducibility
- Model id
- anthropic/claude-sonnet-5
- Generate config
- temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
- Pricing ($/M)
- input_per_m=2 · output_per_m=10
Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.