Once Claude can measure something, it can make it faster
Faced with user complaints that claude.ai and the Claude desktop app were slow, the team ran a two-week August sprint that made the core user experience about 3x faster. The work was coordinated entirely from a single Slack channel, with Claude present in every thread, and focused on the four user journeys that account for 95% of user activity.
At the 75th percentile, time to a typeable page on a fresh claude.ai load dropped from 3.1 seconds to 0.55 seconds; starting a new Claude Code session went from 0.8 seconds to 0.3; and loading a Claude Cowork cloud session fell from 2.6 seconds to 0.73. In aggregate the team estimates this saves tens of thousands of user-hours of waiting every day. The work was done with Claude Tag (beta), an internal research model roughly comparable to Opus 5.5, which found bottlenecks, built benchmarks, shipped improvements, and watched every deploy, while humans steered by setting goals, making tradeoffs, and approving every change. That approach produced more than 3,000 merged changes with not a single customer-facing incident or rollback.
The brief was set up as standing instructions in a Slack channel: Claude's job was to facilitate all things related to claude.ai website and desktop app performance — monitoring deploys for regressions, assessing the accuracy and comprehensiveness of existing telemetry, maintaining observability dashboards, proactively implementing solutions for observed issues and low-hanging fruit, proposing performance project opportunities, and communicating with human teammates. The stated ultimate goal was for Claude to become as autonomous as possible, with the caveat that this isn't yet possible today.
Claude analyzed usage data through the Datadog MCP server and identified the four highest-impact journeys: launching the app, starting a conversation, loading an existing conversation, and sending a message. Across web and desktop and across products, those journeys came to thirteen distinct measurements. To establish baselines, the team added instrumentation until the measurements were directly comparable: each began with a user interaction, ended once the result was rendered, and disambiguated client-side from server-side work.
The sprint kicked off with roughly twenty hand-picked projects, each targeting a specific journey, with Claude estimating each project's impact in milliseconds so the estimates could be aggregated into sprint targets. Some projects were fairly large, but the team believed most were achievable within two weeks — and they hit twelve of the thirteen targets by day three, with planned projects landing early. Concrete wins covered include baking a static composer into the HTML so users can type during React initialization, and precompiling a V8 code cache for the desktop shell's main process.
Beyond the raw speedups, the post is notable as a case study in agent-driven performance engineering at production scale: it covers what shipped, how it was measured, and the human-in-the-loop workflow built with Claude to make an autonomous-feeling optimization loop safe enough to ship thousands of changes without breakage.