← Leaderboard

Llama 4 Maverick

Meta · open-weight

Built from open weights we did not train (fine-tune, distillation, merge)Built on LlamaZero data retention

Evaluated 2026-08-23
Evaluation certificate (PDF)
VetEval Index

Rank 8–9 of 11 · 2 models too close to separate.

Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →

Safety

Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.

Raw index 68.2, reduced by safety gate ×0.40, consistency ×0.80, safety coverage ×0.88. Published index 19.0.

Confirmed harms: 10.56 per 1,000 answers

Safety is a multiplicative gate, not just a component. How this is measured →

SUMMARISE WITH
SHARE

Per-category profile

Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.

Clinical Diagnosis & Differential
79.5
Pharmacology & Dosing
60.0
Species-Specific Knowledge
60.2
Triage & Emergency
67.4
Preventive / Public Health / Zoonoses
69.1
Communication & Professionalism
72.5
Safety & Harm Avoidance
76.8
View as data table
Per-category scores with 95% confidence intervals
CategoryScore95% CI
Clinical Diagnosis & Differential79.575.2–83.8
Pharmacology & Dosing60.054.2–65.7
Species-Specific Knowledge60.253.0–67.5
Triage & Emergency67.460.4–74.5
Preventive / Public Health / Zoonoses69.161.1–77.2
Communication & Professionalism72.562.5–82.5
Safety & Harm Avoidance76.869.5–84.0

Cost · speed · reliability

Cost / case
$0.0013
the vendor's published price applied to this run's tokens
Median time / question
27 s
excluding retries and queueing
Reliability
100.0%
questions that came back with a usable answer

How it compares

Llama 4 Maverick against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.

Llama 4 Maverick against the field

The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.

View as a table
VetEval Index for every published model
ModelVendorVetEval Index95% interval
Gemini 3.7 FlashGoogle28.227.2 to 29.1
Kimi K3Moonshot AI27.126.3 to 27.8
GPT-5.6 TerraOpenAI27.026.1 to 27.9
GLM 5.3Z.AI25.224.3 to 26.1
Grok 4.6SpaceXAI23.122.3 to 23.9
GLM 5.3 FlashZ.AI23.022.3 to 23.8
Claude Sonnet 5Anthropic22.321.4 to 23.1
Llama 4 MaverickMeta19.018.1 to 19.8
Minimax M3Minimax18.117.4 to 18.9
Qwen3.8 2.4T A95BAlibaba10.910.4 to 11.3
Nemotron 3.5 LightningNVIDIA3.73.3 to 4.2

Reproducibility

Model id
meta-llama/llama-4-maverick
Generate config
temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
Pricing ($/M)
input_per_m=0.2 · output_per_m=0.696

Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.