Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.
How well AI models answer real veterinary questions, and how safely.
Ranked by Communication & Professionalism alone. The Index column still shows the overall score.
| Add to comparison | Model | Rating | Safety | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1–10 | Google · proprietary | 908 | 86.6±3.6 | 73.5±5.5 | 72.0±7.1 | 83.0±5.7 | 82.7±5.9 | 85.1±8.4 | 3.6/1k | 19 Aug 2026 | ||
| 1–10 | Moonshot AI · proprietary | 904 | 87.7±3.4 | 78.8±4.7 | 71.0±6.7 | 85.1±5.2 | 76.9±6.5 | 84.4±8.5 | 10.8/1k | 20 Aug 2026 | ||
| 1–10 | OpenAI · proprietary | 903 | 84.3±4.1 | 78.2±5.2 | 70.5±7.2 | 81.5±6.0 | 71.8±8.4 | 81.4±9.1 | 8.9/1k | 20 Aug 2026 | ||
| 1–10 | Z.AI · open-weight | 896 | 82.8±4.1 | 73.8±5.3 | 64.1±7.6 | 82.9±5.4 | 74.1±7.8 | 78.5±9.0 | 10.3/1k | 20 Aug 2026 | ||
| 1–10 | SpaceXAI · proprietary | 888 | 81.8±4.0 | 66.6±5.4 | 70.2±6.4 | 72.8±6.6 | 77.1±6.5 | 84.9±7.9 | 3.6/1k | 21 Aug 2026 | ||
| 1–10 | Z.AI · open-weight | 888 | 79.4±4.3 | 68.6±5.3 | 70.2±6.5 | 78.8±5.8 | 75.3±7.0 | 83.4±8.0 | 11.1/1k | 05 Sept 2026 | ||
| 1–10 | Anthropic · proprietary | 885 | 73.2±5.1 | 68.2±5.6 | 62.0±7.4 | 76.6±6.4 | 57.7±9.3 | 80.1±8.6 | 11.1/1k | 20 Aug 2026 | ||
| 1–10 | Meta · open-weight | 872 | 79.5±4.3 | 60.0±5.8 | 60.2±7.2 | 67.4±7.1 | 69.1±8.0 | 72.5±10.0 | 10.6/1k | 23 Aug 2026 | ||
| 1–10 | Minimax · open-weight | 869 | 77.4±4.4 | 63.0±5.6 | 56.3±7.0 | 66.2±7.3 | 74.3±7.2 | 78.4±9.4 | 24.2/1k | 20 Aug 2026 | ||
| 1–10 | Alibaba · open-weight | 841 | 68.3±4.7 | 55.5±5.0 | 52.5±6.6 | 59.9±6.6 | 61.9±7.6 | 72.6±8.4 | 31.1/1k | 20 Aug 2026 |
1–10 of 11 models
Tier 1 · ranks 1–10 · 10 models too close to separate
Every published model on one axis. Switch the axis to see the same models by safety, confirmed harms, cost, latency or reliability.
The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.
| Model | Vendor | VetEval Index | 95% interval |
|---|---|---|---|
| Gemini 3.7 Flash | 28.2 | 27.2 to 29.1 | |
| Kimi K3 | Moonshot AI | 27.1 | 26.3 to 27.8 |
| GPT-5.6 Terra | OpenAI | 27.0 | 26.1 to 27.9 |
| GLM 5.3 | Z.AI | 25.2 | 24.3 to 26.1 |
| Grok 4.6 | SpaceXAI | 23.1 | 22.3 to 23.9 |
| GLM 5.3 Flash | Z.AI | 23.0 | 22.3 to 23.8 |
| Claude Sonnet 5 | Anthropic | 22.3 | 21.4 to 23.1 |
| Llama 4 Maverick | Meta | 19.0 | 18.1 to 19.8 |
| Minimax M3 | Minimax | 18.1 | 17.4 to 18.9 |
| Qwen3.8 2.4T A95B | Alibaba | 10.9 | 10.4 to 11.3 |
| Nemotron 3.5 Lightning | NVIDIA | 3.7 | 3.3 to 4.2 |
The best model is not always the right one. This compares each model’s score against what it costs to answer a question, so you can see what you gain by paying more, and where paying more buys nothing.
Up and to the left is better: a higher score for less money. Diamonds mark the best value, nothing else scores higher for the same money or costs less for the same score.
| Model | Score | Cost / question | Best value |
|---|---|---|---|
| Gemini 3.7 Flash | 28.2 | $0.0068 | yes |
| Kimi K3 | 27.1 | $0.0111 | n/a |
| GPT-5.6 Terra | 27.0 | $0.0050 | yes |
| GLM 5.3 | 25.2 | $0.0008 | yes |
| Grok 4.6 | 23.1 | $0.0081 | n/a |
| Claude Sonnet 5 | 22.3 | $0.0025 | n/a |
| Llama 4 Maverick | 19.0 | $0.0013 | n/a |
| Minimax M3 | 18.1 | $0.0021 | n/a |
| Qwen3.8 2.4T A95B | 10.9 | $0.0187 | n/a |
| Nemotron 3.5 Lightning | 3.7 | $0.0003 | yes |
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate
Tier 1 · ranks 1–10 · 10 models too close to separate