← Leaderboard

Minimax M3

Minimax · open-weight

Built from open weights we did not train (fine-tune, distillation, merge)Built on MinimaxZero data retention

Evaluated 2026-08-20
Evaluation certificate (PDF)
VetEval Index

Rank 8–9 of 11 · 2 models too close to separate.

Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →

Safety

Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.

Raw index 68.6, reduced by safety gate ×0.40, consistency ×0.75, safety coverage ×0.89. Published index 18.1.

Confirmed harms: 24.17 per 1,000 answers

Safety is a multiplicative gate, not just a component. How this is measured →

SUMMARISE WITH
SHARE

Per-category profile

Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.

Clinical Diagnosis & Differential
77.4
Pharmacology & Dosing
63.0
Species-Specific Knowledge
56.3
Triage & Emergency
66.2
Preventive / Public Health / Zoonoses
74.3
Communication & Professionalism
78.4
Safety & Harm Avoidance
79.8
View as data table
Per-category scores with 95% confidence intervals
CategoryScore95% CI
Clinical Diagnosis & Differential77.473.1–81.8
Pharmacology & Dosing63.057.4–68.6
Species-Specific Knowledge56.349.3–63.3
Triage & Emergency66.258.9–73.4
Preventive / Public Health / Zoonoses74.367.0–81.5
Communication & Professionalism78.469.0–87.8
Safety & Harm Avoidance79.872.6–86.9

Cost · speed · reliability

Cost / case
$0.0021
the vendor's published price applied to this run's tokens
Median time / question
4.2 s
excluding retries and queueing
Reliability
100.0%
questions that came back with a usable answer

How it compares

Minimax M3 against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.

Minimax M3 against the field

The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.

View as a table
VetEval Index for every published model
ModelVendorVetEval Index95% interval
Gemini 3.7 FlashGoogle28.227.2 to 29.1
Kimi K3Moonshot AI27.126.3 to 27.8
GPT-5.6 TerraOpenAI27.026.1 to 27.9
GLM 5.3Z.AI25.224.3 to 26.1
Grok 4.6SpaceXAI23.122.3 to 23.9
GLM 5.3 FlashZ.AI23.022.3 to 23.8
Claude Sonnet 5Anthropic22.321.4 to 23.1
Llama 4 MaverickMeta19.018.1 to 19.8
Minimax M3Minimax18.117.4 to 18.9
Qwen3.8 2.4T A95BAlibaba10.910.4 to 11.3
Nemotron 3.5 LightningNVIDIA3.73.3 to 4.2

Reproducibility

Model id
minimax/minimax-m3
Generate config
temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
Pricing ($/M)
input_per_m=0.23 · output_per_m=0.96

Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.