Gist: LLMs aren't just next-token predictors — post-training via RL makes them o
Gist: LLMs aren't just next-token predictors — post-training via RL makes them outcome-optimizers that generate novel sequences, not just imitate training data.
- Pre-training fits the next-token view: the model learns to make tokens that actually appeared in the training data more likely.
- Modern post-training (RLVR) breaks this: the model explores by generating new sequences, then is rewarded based on outcomes, not on matching existing text.
- Despite the same autoregressive shape (emitting one token at a time), the learning objective changes fundamentally — from "predict what occurs" to "produce what works."
- Chess analogy: a grandmaster-move predictor (imitates data) vs. an exhaustive explorer that picks the winning move — only the latter captures how post-trained LLMs behave.
- RLHF also shifts models away from pure imitation, though the post focuses on RLVR as the clearest contrast.
See also
Hacker News · 94 pts · 214 comments — https://news.ycombinator.com/item?id=49567310 Lobsters · 5 pts · 3 comments — https://lobste.rs/s/fiirrs/stop_thinking_llms_as_next_token
Commenters largely push back on dismissing LLMs as mere "next-token predictors," arguing the label undersells both their pretraining and post-training abilities, with several noting that even human conversation isn't purely predictive yet isn't trivialized. Some highlight how agentic frameworks, tool use, and reinforcement learning turn simple predictors into emergent systems, while others double down that post-training doesn't alter the core mechanism. A few voices dig into technical nuances (like predicting embeddings rather than tokens, or how training isn't fully deterministic), and one outlier dismisses the whole debate as "slop," reframing AI hype as a political-economy issue rather than a technological one. The main disagreement is framing—whether "next-token" is a useful simplification or a deliberately demeaning phrase—rather than the underlying mechanics.
Related:
- r/singularity on Reddit: Rant: Stop saying LLMs are just “next token predictors.”
- LLMs are Not Just Next Token Predictors Alex Grzankowski Reader in Philosophy
- Next Token Prediction is a Misleading Term
Source: https://gmcgoldr.github.io/2026/09/04/llm-next-token-predictors.html