Product Updates Hacker News (Claude)

Maximizing the value of your Claude Code sessions

Claude Codetoken pricingprompt cachingeffort levels

Claude Code sessions incur costs through two token phases: prefill (input) and decode (output). Input tokens include the system prompt, CLAUDE.md, your message, and everything added to the conversation. Output tokens are generated one at a time and keep the GPU busy longer, making output roughly 5x the price of input. Most output tokens are thinking tokens; the /effort level controls how much thinking occurs per turn and persists as the default for future sessions.

The article recommends deliberately setting /model and /effort at the start of a session, since both remember previous choices. For grunt work, MAX_THINKING_TOKENS=0 disables thinking for that session (except on Fable 5), going below /effort low. Prompt caching automatically stores the state of a shared token prefix; reading from cache costs 0.1x input price, while writing costs up to 2x, but the write happens once and reads happen on every subsequent turn.

Claude Code manages caching on every request, but it's easy to break. For example, asking to 'fix the failing test in utils.test.ts' sends an initial request with system prompt, CLAUDE.md, and message, which is fully prefilled and cached. The model then issues a Read call for the test file, which gets appended and resent; the first request's tokens are read from cache at 0.1x, and only the new Read call and file are prefilled at full price. When the model requests the file under test, the cycle repeats, with the first two requests served from cache and only the new file costing full input price.

Read original →

← Back to home