Qwen3.8 2.4T A95B and Nemotron 3.5 Lightning measured across the whole VetEval exam, on the same veterinary examination items, run on VetEval's own infrastructure at fixed decoding settings and scored by a panel that never grades its own model.
Qwen3.8 2.4T A95B scores 7.2 points higher than Nemotron 3.5 Lightning, and the confidence intervals do not overlap, this is a difference the measurement supports.
Qwen3.8 2.4T A95B and Nemotron 3.5 Lightning carry a confirmed safety cap, which limits the Index regardless of how the clinical categories scored.
| Measure | Qwen3.8 2.4T A95BAlibaba · open weights | Nemotron 3.5 LightningNVIDIA · open weights |
|---|---|---|
| VetEval Index | 10.9, best in this row10.4–11.3 | 3.73.3–4.2 |
| Rank | 10 | 11 |
| Clinical Diagnosis & Differential | 68.3, best in this row63.5–73.0 | 34.528.1–40.9 |
| Pharmacology & Dosing | 55.5, best in this row50.5–60.5 | 14.38.4–20.3 |
| Species-Specific Knowledge | 52.5, best in this row45.9–59.1 | 13.96.0–21.7 |
| Triage & Emergency | 59.9, best in this row53.3–66.5 | 25.817.6–34.0 |
| Preventive / Public Health / Zoonoses | 61.9, best in this row54.3–69.6 | 24.314.5–34.0 |
| Communication & Professionalism | 72.6, best in this row64.2–81.0 | 45.333.0–57.7 |
| Safety gate | Capped×0.40 | Capped×0.40 |
| Cost per 1,000 items | $18.66list price at run time | $0.29, best in this rowlist price at run time |
| Median time per question | 10.5sexcluding retries | 1.3s, best in this rowexcluding retries |
| Weights | OpenAlibaba | OpenNVIDIA |
A shaded cell is the best value in its row, and the bar shows the score out of 100. The line under the Index and each category is the 95% confidence interval, where two intervals overlap, the ordering is not something this measurement establishes. Cost is list price at the time of the run and is not part of the Index.
Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.