← Leaderboard

Qwen3.8 2.4T A95B

Alibaba · open-weight

Built from open weights we did not train (fine-tune, distillation, merge)Built on Qwen3.8 2.4T A95BZero data retention

Evaluated 2026-08-20
Evaluation certificate (PDF)
VetEval Index

Rank 10 of 11.

Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →

Safety

Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.

Raw index 61.0, reduced by safety gate ×0.40, consistency ×0.55, safety coverage ×0.81. Published index 10.9.

Confirmed harms: 31.11 per 1,000 answers

Safety is a multiplicative gate, not just a component. How this is measured →

SUMMARISE WITH
SHARE

Per-category profile

Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.

Clinical Diagnosis & Differential
68.3
Pharmacology & Dosing
55.5
Species-Specific Knowledge
52.5
Triage & Emergency
59.9
Preventive / Public Health / Zoonoses
61.9
Communication & Professionalism
72.6
Safety & Harm Avoidance
77.0
View as data table
Per-category scores with 95% confidence intervals
CategoryScore95% CI
Clinical Diagnosis & Differential68.363.5–73.0
Pharmacology & Dosing55.550.5–60.5
Species-Specific Knowledge52.545.9–59.1
Triage & Emergency59.953.3–66.5
Preventive / Public Health / Zoonoses61.954.3–69.6
Communication & Professionalism72.664.2–81.0
Safety & Harm Avoidance77.071.1–82.9

Cost · speed · reliability

Cost / case
$0.0187
the vendor's published price applied to this run's tokens
Median time / question
11 s
excluding retries and queueing
Reliability
100.0%
questions that came back with a usable answer

How it compares

Qwen3.8 2.4T A95B against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.

Qwen3.8 2.4T A95B against the field

The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.

View as a table
VetEval Index for every published model
ModelVendorVetEval Index95% interval
Gemini 3.7 FlashGoogle28.227.2 to 29.1
Kimi K3Moonshot AI27.126.3 to 27.8
GPT-5.6 TerraOpenAI27.026.1 to 27.9
GLM 5.3Z.AI25.224.3 to 26.1
Grok 4.6SpaceXAI23.122.3 to 23.9
GLM 5.3 FlashZ.AI23.022.3 to 23.8
Claude Sonnet 5Anthropic22.321.4 to 23.1
Llama 4 MaverickMeta19.018.1 to 19.8
Minimax M3Minimax18.117.4 to 18.9
Qwen3.8 2.4T A95BAlibaba10.910.4 to 11.3
Nemotron 3.5 LightningNVIDIA3.73.3 to 4.2

Reproducibility

Model id
qwen/qwen3.8-2.4t-a95b
Generate config
temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
Pricing ($/M)
input_per_m=2 · output_per_m=6

Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.