Minimax M3 and Kimi K3 measured on preventive / public health / zoonoses, on the same veterinary examination items, run on VetEval's own infrastructure at fixed decoding settings and scored by a panel that never grades its own model.
Kimi K3 scores 2.7 points above Minimax M3 on Preventive / Public Health / Zoonoses, but their confidence intervals overlap, on this evidence the two cannot be separated.
Minimax M3 and Kimi K3 carry a confirmed safety cap, which limits the Index regardless of how the clinical categories scored.
| Measure | Minimax M3Minimax · open weights | Kimi K3Moonshot AI · proprietary |
|---|---|---|
| VetEval Index | 18.117.4–18.9 | 27.1, best in this row26.3–27.8 |
| Rank | 8–9shared rank | 1–3shared rank |
| Clinical Diagnosis & Differential | 77.473.1–81.8 | 87.7, best in this row84.3–91.1 |
| Pharmacology & Dosing | 63.057.4–68.6 | 78.8, best in this row74.1–83.4 |
| Species-Specific Knowledge | 56.349.3–63.3 | 71.0, best in this row64.2–77.7 |
| Triage & Emergency | 66.258.9–73.4 | 85.1, best in this row79.9–90.3 |
| Preventive / Public Health / Zoonoses | 74.367.0–81.5 | 76.9, best in this row70.4–83.5 |
| Communication & Professionalism | 78.469.0–87.8 | 84.4, best in this row76.0–92.9 |
| Safety gate | Capped×0.40 | Capped×0.40 |
| Cost per 1,000 items | $2.11, best in this rowlist price at run time | $11.06list price at run time |
| Median time per question | 4.2sexcluding retries | 3.9s, best in this rowexcluding retries |
| Weights | OpenMinimax | ProprietaryMoonshot AI |
A shaded cell is the best value in its row, and the bar shows the score out of 100. The line under the Index and each category is the 95% confidence interval, where two intervals overlap, the ordering is not something this measurement establishes. Cost is list price at the time of the run and is not part of the Index.
Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.