The CPU is back: Rethinking the CPU-GPU split for LLM inference
For the past three years, GPUs have dominated LLM inference due to their massive parallelism, which excels at the dense matrix multiplications of transformer forward passes. In traditional chatbot applications, GPUs deliver higher tokens-per-second and better throughput per dollar under high-concurrency batched inference. However, inference is no longer a single model answering a single question; growing reliance on tool calls, multistep reasoning, and orchestration across small specialized models changes where compute should live. Intel has highlighted that the CPU-to-GPU ratio is shifting from 1:8 in training workloads to 1:1, and in some cases 4:1, in agentic deployments.
CPUs and GPUs are fundamentally different: GPUs have tens of thousands of cores for same-operation-on-many-data parallelism, while CPUs have one to hundreds of cores optimized for sequential, conditional, and branching logic. CPUs have direct access to main system memory and naturally host the orchestration layer: tool dispatch, code execution, Python runtimes, sandboxes, I/O, and the agent loop control flow. The key distinction is FLOPS versus instruction latency — GPUs maximize floating-point operations per second, while CPUs minimize the latency of executing unpredictable, diverse command chains. Forcing a GPU to run a chaotic Python runtime leaves its FLOP capacity idle due to constant task-switching.
The two architectures are not competitors but complementary; the real question is workload placement. As agentic applications grow, CPUs are becoming essential for the parts of inference that involve logic-heavy orchestration, while GPUs remain ideal for dense batched computation. The data from Intel suggests the industry is moving toward a more balanced split, renegotiating the assumption that GPUs are the obvious choice for all LLM inference.