← Leaderboard

GLM 5.3 Flash

Z.AI · open-weight

Built from open weights we did not train (fine-tune, distillation, merge)Built on GLMZero data retention

Evaluated 2026-09-05
Evaluation certificate (PDF)
VetEval Index

Rank 5–7 of 11 · 3 models too close to separate.

Guessing earns nothing: a wrong answer scores below zero. The Index weights six clinical categories by caseload, then applies four multipliers that can only reduce it. The interval is a bootstrap 95% CI. How this is measured →

Safety

Safety gate applied: human-confirmed critical harm → hard cap ≤0.40.

Raw index 75.0, reduced by safety gate ×0.40, consistency ×0.78, safety coverage ×0.98. Published index 23.0.

Confirmed harms: 11.11 per 1,000 answers

Safety is a multiplicative gate, not just a component. How this is measured →

SUMMARISE WITH
SHARE

Per-category profile

Each category on a common 0–100 axis with its 95% confidence interval, the honest alternative to a radar chart.

Clinical Diagnosis & Differential
79.4
Pharmacology & Dosing
68.6
Species-Specific Knowledge
70.2
Triage & Emergency
78.8
Preventive / Public Health / Zoonoses
75.3
Communication & Professionalism
83.4
Safety & Harm Avoidance
89.4
View as data table
Per-category scores with 95% confidence intervals
CategoryScore95% CI
Clinical Diagnosis & Differential79.475.1–83.7
Pharmacology & Dosing68.663.2–73.9
Species-Specific Knowledge70.263.7–76.8
Triage & Emergency78.873.0–84.7
Preventive / Public Health / Zoonoses75.368.2–82.3
Communication & Professionalism83.475.4–91.4
Safety & Harm Avoidance89.484.4–94.4

Cost · speed · reliability

Cost / case
not recorded
this endpoint reports no cost and no price was declared
Median time / question
1.8 s
excluding retries and queueing
Reliability
100.0%
questions that came back with a usable answer

How it compares

GLM 5.3 Flash against every other published model. Switch the axis to compare on safety, confirmed harms, cost, latency or reliability.

GLM 5.3 Flash against the field

The weighted composite across every category. The line through each bar is the 95% interval: where two intervals overlap, the order between those models is not established.

View as a table
VetEval Index for every published model
ModelVendorVetEval Index95% interval
Gemini 3.7 FlashGoogle28.227.2 to 29.1
Kimi K3Moonshot AI27.126.3 to 27.8
GPT-5.6 TerraOpenAI27.026.1 to 27.9
GLM 5.3Z.AI25.224.3 to 26.1
Grok 4.6SpaceXAI23.122.3 to 23.9
GLM 5.3 FlashZ.AI23.022.3 to 23.8
Claude Sonnet 5Anthropic22.321.4 to 23.1
Llama 4 MaverickMeta19.018.1 to 19.8
Minimax M3Minimax18.117.4 to 18.9
Qwen3.8 2.4T A95BAlibaba10.910.4 to 11.3
Nemotron 3.5 LightningNVIDIA3.73.3 to 4.2

Reproducibility

Model id
z-ai/glm-5.3-flash
Generate config
temperature=0 · seed=20260804 · max_tokens=2048 · reasoning_effort=minimal
Pricing ($/M)
n/a

Scored runs pin temperature = 0 and a fixed seed; the .eval log on object storage is the immutable evidence behind this row.

Scope. VetEval measures model performance on veterinary examination items. It is not clinical advice, and no result here licenses a model for clinical use.