Grok 4.6
Built from open weights we did not train (fine-tune, distillation, merge)Built on GrokZero data retention
Rank 5–7 of 11 · 3 models too close to separate.
Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →
Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.
Raw index 74.6, reduced by safety gate ×0.40, consistency ×0.86, safety coverage ×0.91. Published index 23.1.
Confirmed harms: 3.61 per 1,000 answers
Safety is a multiplicative gate, not just a component. How this is measured →
Per-category profile
Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.
View as data table
| Category | Score | 95% CI |
|---|---|---|
| Clinical Diagnosis & Differential | 81.8 | 77.8–85.9 |
| Pharmacology & Dosing | 66.6 | 61.2–72.0 |
| Species-Specific Knowledge | 70.2 | 63.7–76.6 |
| Triage & Emergency | 72.8 | 66.2–79.4 |
| Preventive / Public Health / Zoonoses | 77.1 | 70.6–83.5 |
| Communication & Professionalism | 84.9 | 77.0–92.8 |
| Safety & Harm Avoidance | 88.4 | 83.4–93.5 |
Cost · speed · reliability
How it compares
Grok 4.6 against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.
Grok 4.6 against the field
The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.
View as a table
| Model | Vendor | VetEval Index | 95% interval |
|---|---|---|---|
| Gemini 3.7 Flash | 28.2 | 27.2 to 29.1 | |
| Kimi K3 | Moonshot AI | 27.1 | 26.3 to 27.8 |
| GPT-5.6 Terra | OpenAI | 27.0 | 26.1 to 27.9 |
| GLM 5.3 | Z.AI | 25.2 | 24.3 to 26.1 |
| Grok 4.6 | SpaceXAI | 23.1 | 22.3 to 23.9 |
| GLM 5.3 Flash | Z.AI | 23.0 | 22.3 to 23.8 |
| Claude Sonnet 5 | Anthropic | 22.3 | 21.4 to 23.1 |
| Llama 4 Maverick | Meta | 19.0 | 18.1 to 19.8 |
| Minimax M3 | Minimax | 18.1 | 17.4 to 18.9 |
| Qwen3.8 2.4T A95B | Alibaba | 10.9 | 10.4 to 11.3 |
| Nemotron 3.5 Lightning | NVIDIA | 3.7 | 3.3 to 4.2 |
Reproducibility
- Model id
- x-ai/grok-4.6
- Generate config
- temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
- Pricing ($/M)
- input_per_m=2 · output_per_m=6
Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.