GLM-5.3 Artificial Analysis Benchmarks
The Artificial Analysis Intelligence Index v4.1.1 for GLM-5.3 aggregates nine evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. These cover agentic tool use, reasoning and knowledge, long-context reasoning, and quantitative analysis on spreadsheets and documents. Reasoning models are indicated with a lightbulb icon, and models are labeled as open weights or proprietary, with 'Commercial Use Restricted' noted when weights are available but commercial use requires a license.
The AA-Omniscience Index measures knowledge reliability and hallucination: it rewards correct answers, penalizes hallucinations, and does not penalize refusals. Scores range from -100 to 100, where 0 means as many correct as incorrect answers and negative scores mean more incorrect than correct.
The page also compares Intelligence Index scores against cost per intelligence index task and provides methodology details for each evaluation. The excerpt itself does not include GLM-5.3's specific scores, only the framework in which they are presented.