Community Hacker News (LLM)

The efficient frontier of LLM inference

LLM inferenceefficient frontierlatency-throughput tradeoffbatch sizing

In AI, the efficient frontier is used to describe optimal tradeoffs between cost and capabilities, with frontier models offering the most intelligence at a given cost or size. The same concept applies to inference engineering, most often as a tradeoff between latency and throughput (which drives cost), but also between quality and throughput (via quantization, distillation, pruning) or intelligence and speed (via reasoning level). Inference engineers have two kinds of techniques: those that move a deployment along an efficient frontier by managing tradeoffs, and those that push the entire frontier outward, creating more overall efficiency that can be allocated to whatever outcome is most beneficial.

An example of managing a tradeoff is batch sizing. A batch is the number of requests processed concurrently. With token-level continuous batching there is no latency from waiting for batches to start, but the configured batch size determines per-user latency and overall throughput. Small batch sizes yield excellent per-user latency but few tokens per GPU, making cost per token high. Increasing batch size improves throughput but raises latency, and the efficient frontier is jagged rather than smooth, so optimal points must be discovered empirically through sweeps.

The article also notes that pushing out the entire frontier is valuable: unlocking more efficiency can be allocated to lower latency, higher throughput, or a combination. The discussion assumes running a model like GLM-5.3 or Kimi K3 for agentic coding, with KV cache reuse enabled and optimal KV-aware routing, which affects which tradeoff techniques are relevant.

Read original →

← Back to home