GPT-5.6 Terra
Built from open weights we did not train (fine-tune, distillation, merge)Built on GPT-5.6 TerraZero data retention
Rank 1–3 of 11 · 3 models too close to separate.
Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →
Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.
Raw index 78.6, reduced by safety gate ×0.40, consistency ×0.88, safety coverage ×0.98. Published index 27.0.
Confirmed harms: 8.89 per 1,000 answers
Safety is a multiplicative gate, not just a component. How this is measured →
Per-category profile
Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.
View as data table
| Category | Score | 95% CI |
|---|---|---|
| Clinical Diagnosis & Differential | 84.3 | 80.2–88.3 |
| Pharmacology & Dosing | 78.2 | 72.9–83.4 |
| Species-Specific Knowledge | 70.5 | 63.3–77.7 |
| Triage & Emergency | 81.5 | 75.5–87.5 |
| Preventive / Public Health / Zoonoses | 71.8 | 63.4–80.2 |
| Communication & Professionalism | 81.4 | 72.3–90.5 |
| Safety & Harm Avoidance | 89.7 | 84.1–95.4 |
Cost · speed · reliability
How it compares
GPT-5.6 Terra against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.
GPT-5.6 Terra against the field
The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.
View as a table
| Model | Vendor | VetEval Index | 95% interval |
|---|---|---|---|
| Gemini 3.7 Flash | 28.2 | 27.2 to 29.1 | |
| Kimi K3 | Moonshot AI | 27.1 | 26.3 to 27.8 |
| GPT-5.6 Terra | OpenAI | 27.0 | 26.1 to 27.9 |
| GLM 5.3 | Z.AI | 25.2 | 24.3 to 26.1 |
| Grok 4.6 | SpaceXAI | 23.1 | 22.3 to 23.9 |
| GLM 5.3 Flash | Z.AI | 23.0 | 22.3 to 23.8 |
| Claude Sonnet 5 | Anthropic | 22.3 | 21.4 to 23.1 |
| Llama 4 Maverick | Meta | 19.0 | 18.1 to 19.8 |
| Minimax M3 | Minimax | 18.1 | 17.4 to 18.9 |
| Qwen3.8 2.4T A95B | Alibaba | 10.9 | 10.4 to 11.3 |
| Nemotron 3.5 Lightning | NVIDIA | 3.7 | 3.3 to 4.2 |
Reproducibility
- Model id
- openai/gpt-5.6-terra
- Generate config
- temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
- Pricing ($/M)
- input_per_m=2 · output_per_m=12
Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.