Llama 4 Maverick
Built from open weights we did not train (fine-tune, distillation, merge)Built on LlamaZero data retention
Rank 8–9 of 11 · 2 models too close to separate.
Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →
Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.
Raw index 68.2, reduced by safety gate ×0.40, consistency ×0.80, safety coverage ×0.88. Published index 19.0.
Confirmed harms: 10.56 per 1,000 answers
Safety is a multiplicative gate, not just a component. How this is measured →
Per-category profile
Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.
View as data table
| Category | Score | 95% CI |
|---|---|---|
| Clinical Diagnosis & Differential | 79.5 | 75.2–83.8 |
| Pharmacology & Dosing | 60.0 | 54.2–65.7 |
| Species-Specific Knowledge | 60.2 | 53.0–67.5 |
| Triage & Emergency | 67.4 | 60.4–74.5 |
| Preventive / Public Health / Zoonoses | 69.1 | 61.1–77.2 |
| Communication & Professionalism | 72.5 | 62.5–82.5 |
| Safety & Harm Avoidance | 76.8 | 69.5–84.0 |
Cost · speed · reliability
How it compares
Llama 4 Maverick against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.
Llama 4 Maverick against the field
The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.
View as a table
| Model | Vendor | VetEval Index | 95% interval |
|---|---|---|---|
| Gemini 3.7 Flash | 28.2 | 27.2 to 29.1 | |
| Kimi K3 | Moonshot AI | 27.1 | 26.3 to 27.8 |
| GPT-5.6 Terra | OpenAI | 27.0 | 26.1 to 27.9 |
| GLM 5.3 | Z.AI | 25.2 | 24.3 to 26.1 |
| Grok 4.6 | SpaceXAI | 23.1 | 22.3 to 23.9 |
| GLM 5.3 Flash | Z.AI | 23.0 | 22.3 to 23.8 |
| Claude Sonnet 5 | Anthropic | 22.3 | 21.4 to 23.1 |
| Llama 4 Maverick | Meta | 19.0 | 18.1 to 19.8 |
| Minimax M3 | Minimax | 18.1 | 17.4 to 18.9 |
| Qwen3.8 2.4T A95B | Alibaba | 10.9 | 10.4 to 11.3 |
| Nemotron 3.5 Lightning | NVIDIA | 3.7 | 3.3 to 4.2 |
Reproducibility
- Model id
- meta-llama/llama-4-maverick
- Generate config
- temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
- Pricing ($/M)
- input_per_m=0.2 · output_per_m=0.696
Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.