Large Language Models and Emergence: A Complex Systems Perspective
Can a language model acquire a new capability suddenly, or can a scoring rule make gradual improvement look sudden? In Wei et al.'s 2022 study, some benchmark scores remained near chance across smaller models before rising at larger scales. Those observations motivate the emergence debate. They do not, by themselves, identify its mechanism.
The distinction matters when you decide whether another increase in training compute will improve a system predictably or produce an unexpected result. This article separates three questions: what the benchmark measures, what changes inside the model, and what scaling laws actually predict.
What Emergence Means Here
There is no universally accepted definition of emergence. Krakauer, Krakauer, and Mitchell emphasize collective organization that permits a simpler, predictive description of a system. Following that approach, this article uses mechanistic emergence for behavior explained by an interacting system's organization, with evidence for a useful reduced description of that organization.
Here is an operational test you can apply: specify the components, the behavior to explain, and a baseline that combines component contributions independently. Show where that baseline fails, then test whether a description of their interactions predicts the behavior more effectively. Failure of a linear superposition is a useful starting point, not a sufficient definition: a nonlinear function alone does not demonstrate a new collective organization.
Water illustrates the distinction. A description of a single isolated molecule omits the interactions that produce a crystal. The point is not that microscopic physics becomes invalid, but that collective variables can provide a more useful account. Anderson's “More Is Different” argues for taking those levels of description seriously.
A scale threshold also needs interpretation. Think of the Sorites question: how many grains make a heap? A label can acquire a boundary even when the underlying quantity changes gradually. Similarly, defining “usable” as exceeding 90% task success creates a threshold by decision. That threshold alone is not evidence of a physical phase transition. The analogy is about classification, not a physical model of freezing.
For a model experiment, record the model family, training data, compute, prompts, metric, and success criterion. If you change several of those together, you cannot attribute a sharp score change to parameter count alone.
The Benchmark Observation
Wei et al. use a narrower, behavioral meaning: an ability absent in smaller models and present in larger ones, with its appearance not predicted by extrapolating the smaller models' performance. That definition depends on the task and evaluation. Paper and definition.
Figure 1: Benchmark results reproduced from Wei et al., Figure 2. The vertical metrics differ: accuracy, exact match, or BLEU. Each point represents a separate model. Sharp increases between sampled points do not establish a mathematical discontinuity. Source.
Read panel A carefully. It shows modular-arithmetic results for GPT-3 and LaMDA, not a near-perfect PaLM result. The plotted scores reach roughly 33% and 16%, respectively. A comparison across a small set of trained models leaves the behavior between those points unmeasured.
The Mirage Problem
Schaeffer, Miranda, and Koyejo's “Are Emergent Abilities of Large Language Models a Mirage?” tests an alternative explanation: nonlinear or discontinuous scoring can turn smooth changes in model outputs into apparently abrupt benchmark gains. Their evidence includes arithmetic experiments and a BIG-Bench meta-analysis.
Consider an illustrative calculation, not a result from the paper. Suppose each of ten output tokens is independently correct with probability p. Whole-answer accuracy is then p raised to the tenth power. Increasing p from 0.8 to 0.9 increases this exact-match probability from about 0.107 to 0.349. Token-level improvement and whole-answer success tell different stories even though they describe the same hypothetical system.
Partial-credit measures, such as token edit distance, can expose progress hidden by exact match. But edit distance is not itself “percentage accuracy,” and changing metrics does not guarantee a power law. The study challenges particular emergence claims; it does not prove that every capability or internal mechanism improves smoothly. Source.
For your own evaluation, retain task success as well as a more informative diagnostic metric. A correct complete answer may be what users need, even when its apparent sudden arrival has a measurement explanation.
The Complexity Science Reframe
Krakauer and colleagues distinguish knowledge-out systems, with simple components and interactions, from knowledge-in systems shaped by learning, evolution, or structured environments. LLMs belong to the latter category. The distinction concerns where structure comes from; it does not separate memorized answers from genuinely new reasoning. Section 4.
Figure 2: Krakauer et al.'s examples of knowledge-in emergence include flocking, visual receptive fields, and content-addressable memory. The table does not contain an LLM row. Source, Table 1.
Their framework also distinguishes emergent capability from intelligence: they ask whether a system can use abstractions efficiently and generalize them. “Less is more” describes that efficiency criterion, not a claim that larger models merely retrieve stored answers. Conclusions.
Mechanistic Evidence: Induction Heads and Grokking
Induction heads implement a simple copying pattern: after seeing A followed by B, another A can prompt the model to predict B. Olsson et al. connect their development during training with improved in-context learning. They report strong causal evidence in small attention-only models and correlational evidence in larger models with multilayer perceptrons.
That result directly contradicts a universal claim that induction heads or in-context learning require about 100 billion parameters. It also separates two experimental axes: a change during training is not necessarily a threshold encountered when comparing parameter counts. The paper's in-context-learning measure concerns prediction improving later in a sequence; it should not automatically be equated with every form of learning a task from prompt examples.
Grokking is delayed generalization after a model has already fitted its training data. Power et al. demonstrated it on small algorithmic datasets. Nanda et al. then analyzed a modular-addition model and found gradual circuit development underlying a sharp improvement in generalization.
These are reasons to investigate mechanisms. They are not a license to label every abrupt validation curve a discontinuous reorganization. A mechanism can develop gradually before a task metric reveals its practical effect.
What the Scaling Laws Actually Predict
Kaplan et al., “Scaling Laws for Neural Language Models”, Section 1.2, fits test cross-entropy loss, not a general measure of intelligence. Its single-bottleneck regimes are:
| Limiting resource | Approximate fitted relationship | Conditions |
|---|---|---|
| Non-embedding parameter count N | L ∝ N⁻⁰·⁰⁷⁶ | Convergence with sufficiently abundant data |
| Dataset size D, in tokens | L ∝ D⁻⁰·⁰⁹⁵ | Large models, limited data, early stopping |
| Optimally allocated compute C_min | L ∝ C_min⁻⁰·⁰⁵⁰ | Sufficient data, suitable model size and small batch size |
These are separate empirical regimes, not three factors to multiply together. A power law is straight on log–log axes: log L = constant − α log N. It is not the claim that performance is proportional to log(parameters).
The fits are not mathematical bounds or promises of unlimited improvement. Kaplan et al. explicitly discuss eventual breakdown in Section 6.3. Later, Hoffmann et al. obtained a different compute-optimal allocation, recommending that model size and training tokens grow in roughly equal proportion.
There is therefore no necessary paradox between a smooth average loss trend and an unexpected task result. An average cannot specify every component of that average, and a loss fit does not identify the circuitry producing it. Predicting a deployment outcome requires evaluating that outcome.
Why the Distinction Matters
Instead of assigning whole capabilities permanently to “smooth,” “emergent,” or “artifact,” ask what each observation warrants:
| Observation | What you can conclude | What you still need |
|---|---|---|
| Exact-match score rises sharply | Complete answers became more frequent | Partial-credit metrics and denser scale measurements |
| Average loss follows a fitted curve | The fit describes the measured regime | Held-out task validation and checks outside that regime |
| Intervening on a circuit changes behavior | The circuit contributes causally in that experiment | Tests of generality across models and tasks |
| Generalization improves late in training | Training and test behavior evolved differently | Analysis of when the responsible mechanism developed |
This is a proposed evaluation checklist, not a taxonomy of proven phase transitions. For a system you plan to deploy, specify which successful behaviors matter, measure them directly, and examine enough intermediate checkpoints or model sizes to test your explanation.
Open Questions
Can a mechanistic account predict a capability before the benchmark reveals it? Does the same explanation survive a change in training data or architecture? Can a small model obtain the ability through a different training procedure? Those experiments are more informative than assuming a universal parameter threshold.
The immediate practical question is narrower: which conclusions survive changing the metric while holding the model outputs fixed? Start there, then investigate whether a circuit-level explanation accounts for the remaining behavior.
The Honest Conclusion
LLM emergence is a family of claims that require different evidence. A benchmark jump, a useful-performance threshold, and a new internal organization are not interchangeable observations. Scaling laws help forecast loss within studied regimes; they do not settle every question about capability or mechanism.
Pick one task you care about and plot both complete-answer success and a partial-credit measure across the same models or checkpoints. If the curves suggest different stories, inspect the outputs before claiming a phase transition. That is a concrete way to turn the emergence debate into a testable question.
Primary References
- Kaplan et al. (2020), Scaling Laws for Neural Language Models: arXiv, arXiv DOI.
- Wei et al. (2022), Emergent Abilities of Large Language Models: arXiv, arXiv DOI.
- Schaeffer et al. (2023), Are Emergent Abilities of Large Language Models a Mirage?: arXiv, arXiv DOI.
- Krakauer et al. (2025), Large Language Models and Emergence: A Complex Systems Perspective: arXiv, arXiv DOI.
- Olsson et al. (2022), In-context Learning and Induction Heads: arXiv, arXiv DOI.
- Power et al. (2022), Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets: arXiv, arXiv DOI.
- Nanda et al. (2023), Progress Measures for Grokking via Mechanistic Interpretability: arXiv, arXiv DOI.
- Hoffmann et al. (2022), Training Compute-Optimal Large Language Models: arXiv, arXiv DOI.
- Anderson (1972), More Is Different: publisher DOI.
Help improve this article
Found an error or a better source? Leave a note here, or highlight a passage to comment on it.