Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Magnitude, launched on Hacker News as part of YC S25, is an open source inference engine for agents that optimizes itself for the user's exact hardware. It compiles and tunes its kernels on-device before a model runs, which lets open models run up to 2x faster than llama.cpp. It ships as a desktop app for macOS, Windows, and Linux, includes the magnitude CLI without separate installation, and is free, private, and Apache 2.0 licensed: there are no token costs, nothing leaves the machine, and no internet is needed once a model is downloaded. The launch asks developers to star the repo to grow the community.
For performance, Magnitude claims up to 2x faster performance than llama.cpp, with 92% faster decode on Metal and 19% on CUDA. It writes hand-optimized kernels for popular open-weight families and tunes them on the actual device rather than shipping kernels precompiled for broad hardware classes like llama.cpp, Ollama, or LM Studio. It also flexes memory: 27% less memory per agent, freed when agents stop, and fast concurrent sessions share prefix caches to prevent slowdown. It runs on any Apple Silicon, NVIDIA, or AMD GPU, or on nothing but a CPU, with no fixed minimum; smaller machines run smaller models, and more memory allows larger ones. The supported model list is at magnitude.dev/models.
For integration, the app connects with one click to Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, while anything else works through an OpenAI-compatible API. To get started, users download and install Magnitude, open the app, choose a recommended model in Discover and download it, then connect their agent in Connections and start using it.
The FAQ reiterates that Magnitude is an open source inference engine that optimizes itself for the user's hardware, ships as a desktop app, and runs open models connected to the agent the user already uses. Its differentiator is on-device kernel compilation and tuning, aimed at making local agents faster and more memory-efficient without sending prompts, files, or models off the machine.