Product Updates Hacker News (Claude)

Portal by Spotify cut my Claude Code token usage by 90%

Claude CodePortaltoken usagecost optimization

Most of an AI coding agent's work is I/O, not reasoning—reading multiple files, generating repetitive test patterns, and updating docs consume thousands of tokens with almost no cognitive demand. Yet these tokens are often fed to frontier models that are wildly overqualified for the task. The waste is significant: by 2028 AI coding costs are projected to exceed the average developer's salary, and a quarter of engineering leaders already spend $200–$500 per developer per month on tokens, with some past $2,000.

The fix described here requires no platform team and no new subscription—just two 'modes' built with Portal by Spotify's AiKA Modes feature. A mode is a declarative agent on an ephemeral runtime (like AWS Lambda for agents), where you define instructions, select a model, set parameters such as temperature, and attach MCP tools. Modes are callable from the Portal CLI/API and can be shared company-wide or kept private. The author created two public modes: bulk-reader, which analyzes files and answers questions with structured bullets (temperature 0.2), and code-writer, which generates boilerplate code exactly matching existing patterns (temperature 0.2). Both use Gemini 2.5 Flash as the worker model, but the model field accepts any model you have in your Portal instance.

The result, as the title states, is a 90% reduction in Claude Code token usage—achieved with two mode definitions and zero custom code. The broader lesson is that intelligently routing token-heavy grunt work to cheaper models can drastically cut costs, leaving frontier models for the problems that actually need their reasoning ability.

Read original →

← Back to home