Other Hacker News (AI)

How to Build an AI Software Factory: Agents That Open, Review, and Merge PRs

AI coding agentssoftware factorycode reviewMCP

The piece opens with a concrete example: on January 6, 2026, Stephen Toub opened nine pull requests from his phone at 35,000 feet, and seven of them merged. He works on dotnet/runtime, and he wrote up what the experience told him: AI changes the economics of code production, and one person with good judgment and a phone can generate PRs faster than a team can review them. That sentence is the entire subject—a single engineer with a coding agent can now saturate a team's review capacity from an airplane seat. The interesting question stopped being how to make agents write code and became how to absorb the output. An AI software factory is the answer companies have converged on, and it is what turns autonomous coding agents from a demo into throughput a team can absorb. The guide breaks it into five stages, each with the published architecture behind it and the configuration to build it; for the broader picture of how AI agents reason, call tools, and pull in web context, it points to a primer, and assumes you have already picked an agent from a roundup of the best AI coding agents. The software factory is everything around that agent.

An AI software factory is defined as the system around a coding agent rather than the agent itself. Work arrives from a queue, agents run in isolated workspaces, verification happens automatically, and a human sits at an explicit merge gate. It is also called an agentic software factory. The useful distinction is between an agent and a software factory: running a coding agent on your laptop is an agent—you choose the task, you watch it work, you read the diff, you merge—so everything except the typing is still you, and your attention is the limit. A software factory moves those steps into infrastructure: nobody decides which issue an agent picks up because intake rules do, nobody sets up a workspace because isolation is provisioned, and the excerpt cuts off as it begins to describe the next step.

The factory is five stages with a gate at each one, and the agent is the cheap part. Intake decides which work is worth starting; Sentry's Seer scores every incoming issue for actionability first. Isolation decides where the agent runs without colliding; Stripe boots pre-warmed devboxes in about 10 seconds. Tools decide what the agent can reach; Stripe's Toolshed exposes roughly 500 internal tools over MCP. Verification decides whether the change is right; Spotify's LLM judge vetoes about 25% of agent sessions. Merge gate decides who is accountable; Faire requires two human reviews on agent-authored PRs. The short version is that every company that made this work built the gates before the fleet. Spotify's Fleetshift shipped in 2023, two years before it had an agent to put in it. Generation scales with spend, review does not, and that asymmetry is the whole design problem.

Where Firecrawl fits: agents need live web context the repo does not carry. Search and scrape cover the open web, and a curated developer index covers code, both behind one MCP block, with prompt injection detection on every fetch.

Read original →

← Back to home