Community Hacker News (LLM)

“Next-token predictor” is the wrong mental model for LLMs

LLMsnext-token predictionRLVRpost-training

The author acknowledges that LLMs do emit tokens autoregressively, making 'next-token predictor' a reasonable zeroth-order approximation. But they contend the label is incomplete, as it only captures the base model's pre-training, where the model learns to make the actual next token from existing training sequences more likely.

Modern post-training includes RLVR, where the model generates new sequences and learns from their outcomes. The training loop there makes explored tokens more likely if the sequence earns a high reward, which is fundamentally different from pre-training's simple prediction of existing text.

The piece uses a chess analogy to make the distinction clearer, contrasting a system trained on grandmaster games with one that explores new moves—though the excerpt cuts off mid-explanation.

Read original →

← Back to home