Mercury 2.5 LLM hits 770 tokens per second
The item is a Hacker News post about Mercury 2.5, an LLM claimed to reach 770 tokens per second. Throughput at that level matters because generation speed, rather than raw capability alone, is often the limiting factor for interactive agents and real-time applications.
The excerpt itself is mostly Artificial Analysis methodology documentation rather than Mercury 2.5 results. It describes the Artificial Analysis Intelligence Index v4.3.2, which incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. The index is broken out by open weights versus proprietary models, with weights labelled "Commercial Use Restricted" when commercial use is limited by conditions and "Non-commercial" when the license prohibits it.
Capability indexes are reported across several domains — agentic knowledge work and agentic real-world work tasks (both on (Elo-500)/2000 scales), agentic SaaS workflows, agentic coding and terminal use, professional document reasoning (all-pass), physics reasoning, and medical long-context reasoning. The methodology notes that while model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.
Two evaluations are described in more detail. AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis, where the AA-Briefcase Elo combines rubric pass rate, analytical quality Elo and presentation Elo, with rubric performance converted into Elo through synthetic head-to-head matches and both Elo and 95% confidence interval bounds clamped at 0. AA-Omniscience Index is a knowledge-reliability measure that rewards correct answers and penalizes hallucination.
Notably, the excerpt contains no Mercury 2.5 scores on any of these benchmarks, so its measured standing relative to other models is not established by this text — the 770 tokens/second figure remains the headline claim.