Qwen3.8 2.4T A95B vs Claude Sonnet 5

Qwen3.8 2.4T A95B and Claude Sonnet 5 measured across the whole VetEval exam, on the same veterinary examination items, run on VetEval's own infrastructure at fixed decoding settings and scored by a panel that never grades its own model.

Claude Sonnet 5 scores 11.4 points higher than Qwen3.8 2.4T A95B, and the confidence intervals do not overlap, this is a difference the measurement supports.

Qwen3.8 2.4T A95B and Claude Sonnet 5 carry a confirmed safety cap, which limits the Index regardless of how the clinical categories scored.

2 of 4 models. Press Compare to see the results.

Side-by-side comparison of Qwen3.8 2.4T A95B, Claude Sonnet 5 across the VetEval Index, each scoring category, the safety gate, cost and latency.
MeasureQwen3.8 2.4T A95BAlibaba · open weightsClaude Sonnet 5Anthropic · proprietary
VetEval Index10.910.4–11.322.3, best in this row21.4–23.1
Rank105–7shared rank
Clinical Diagnosis & Differential68.363.5–73.073.2, best in this row68.1–78.2
Pharmacology & Dosing55.550.5–60.568.2, best in this row62.6–73.8
Species-Specific Knowledge52.545.9–59.162.0, best in this row54.6–69.4
Triage & Emergency59.953.3–66.576.6, best in this row70.1–83.0
Preventive / Public Health / Zoonoses61.9, best in this row54.3–69.657.748.4–67.0
Communication & Professionalism72.664.2–81.080.1, best in this row71.5–88.6
Safety gateCapped×0.40Capped×0.40
Cost per 1,000 items$18.66list price at run time$2.48, best in this rowlist price at run time
Median time per question10.5sexcluding retries1.6s, best in this rowexcluding retries
WeightsOpenAlibabaProprietaryAnthropic

A shaded cell is the best value in its row, and the bar shows the score out of 100. The line under the Index and each category is the 95% confidence interval, where two intervals overlap, the ordering is not something this measurement establishes. Cost is list price at the time of the run and is not part of the Index.

What this comparison does not tell you

  • These scores describe performance on veterinary examination items. They are not a certification, and no result here licenses a model for clinical use.
  • The Index weights categories for a general caseload. Every per-category score is published above so you can weigh them for yours instead, or compare on a single category using the control above.
  • Cost and latency are measured at run time and move separately from the scores. A model that is cheaper today may not be next month.

Related comparisons

Qwen3.8 2.4T A95B vs Claude Sonnet 5 by category

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.