1. The bottleneck shift: from I/O throughput to gradient information density
There was an era when pre-training performance meant systems performance. Dataloaders starved accelerators; sharding, prefetching, zero-copy token streams, and fused kernels were the difference between a run that finished and one that didn't. Those problems are now largely solved — modern training stacks stream tokens at line rate and keep devices saturated.
That victory relocates the bottleneck rather than removing it. When hardware idles at zero, the binding constraint on capability-per-dollar becomes gradient information density: how much usable structural signal each parameter update absorbs. A batch of near-duplicate web pages and a batch of formally verified proofs both cost the same FLOPs, but they do not buy the same learning.
Two facts sharpen this into an emergency rather than a preference:
- Modern pre-training still treats all tokens equally. The default recipe — scrape everything, shuffle uniformly, train once — assigns identical marginal value to noise, redundancy, and structure.
- The commodity itself has a ceiling. Villalobos et al. project that public human-generated text will be exhausted as a scaling resource somewhere between 2026 and 2032, with the median estimate landing in 2028 (Villalobos et al., 2024). Simply re-reading the pile doesn't dodge the wall: Muennighoff et al. show repeated data retains full value for roughly four epochs, then decays toward zero return (Muennighoff et al., 2023).
When you cannot scale the dataset, you must scale what the dataset teaches. Compute efficiency has to come from data selection and curriculum design — deciding which tokens enter the next batch — rather than unconstrained dataset growth. That decision procedure needs a quantity to optimize, which brings us back to information theory and to why the classical version isn't good enough.
2. Beyond Shannon: information for compute-bounded learners
Shannon entropy measures the surprise of a source for an observer with unlimited computational power. Kolmogorov complexity measures the shortest program that reproduces a string, again with no runtime budget. Both are the right tools for communication channels and randomness extraction, but both are quietly wrong for learning, because learners are compute-bounded observers.
The failure mode is precise: pseudo-random noise generated by a short program has tiny Kolmogorov complexity but is unlearnable within any realistic budget, while rich, compressible regularities — grammar, arithmetic, multi-step dependencies — may carry modest Shannon entropy yet dominate what a bounded learner can extract. Classical measures cannot distinguish “unpredictable because random” from “unpredictable until you find the circuit.”
As we argued in 003, Epiplexity (\(S_T\), Finzi et al., 2026) names the missing decomposition: the structural information a time-bounded model can actually extract and store, separate from residual time-bounded entropy \(H_T\). Models don't memorize tokens; they allocate parameter capacity to internalize reusable subprograms — copying mechanisms, counting routines, algorithmic circuits. \(S_T\) counts that allocation, and noise contributes nothing to it no matter how surprising.
What makes this operational rather than philosophical is a much older idea: the MDL prequential view (Rissanen, 1996; Grünwald, 2007). Treat the learner as a sequential coder. Its cumulative codelength on unseen data, \(\sum_i -\log_2 q(x_i \mid x_{is its running validation loss in different units. Learning progress is coding progress: the prequential coding gap between a trained model and a trivial baseline measures exactly how much description work the model has taken over.
This reframes learning as a process, not a property:
- Information content is observer-dependent — it depends on the compute available to the observer.
- It is measurable online, from loss trajectories, with no labels and no ground truth about “true structure.”
- It decomposes: the part of the gap attributable to reusable structure, as opposed to domain-specific fitting, is precisely the transferable component.
A training loop already computes the raw material — losses over time, per domain. The engineering question is what to build on top of that signal.
3. Dynamic curricula and cross-domain scaling laws
The operational layer we want is a controller: observe per-domain learning dynamics, forecast each domain's marginal return on more exposure, and adjust the sampling mix accordingly.
The forecast is possible because per-domain loss trajectories follow predictable parametric laws. For domain \(k\), fit
$$L_k(n) = \epsilon_k + \beta_k \, n^{-\alpha_k}$$
where \(n\) is cumulative tokens seen and \(\alpha_k, \beta_k, \epsilon_k\) capture decay steepness, scale, and irreducible floor. The derivative gives the marginal value of the next token:
$$\frac{\partial L_k}{\partial n} = -\alpha_k \beta_k \, n^{-\alpha_k - 1}, \qquad \frac{\partial \hat{S}}{\partial n_k} \propto -\frac{\partial L_k}{\partial n_k}$$
A scheduler that tracks these fits online can rank domains by \(\partial \hat{S}/\partial n_k\) and spend the next window of batches where structure is being absorbed fastest. This is the design of ADO, which fits per-domain laws during training and steers the mixture without proxy models (Jiang et al., 2024); its predecessor DoReMi showed a 280M-parameter proxy could produce mixture weights that train an 8B model to baseline accuracy 2.6× faster (Xie et al., 2023). EpiSelect extends the family by making the law cross-domain: a mixing vector \(\gamma_{m,k}\) attributes each domain's improvement to exposures across all domains, turning a ranking heuristic into a transfer map (Su et al., 2026).
Two properties of that cross-domain matrix matter enormously:
- Directional: code helps natural-language reasoning far more than the reverse. Exposure is not symmetric.
- Sparse: most pairs interact weakly; a handful of edges (code → reasoning, formal math → code) carry most of the transfer.
Targeted exposure to highly structured domains therefore acts as an accelerator for unrelated downstream tasks. Muennighoff et al. observed a clean instance: adding code data to a data-constrained run left natural-language perplexity intact while tasks like WebNLG and bAbI jumped — attributed to long-range state-tracking abilities learned from code (Muennighoff et al., 2023). Uniform shuffling buys none of this, because it spends its budget proportionally to dataset size rather than to interaction strength.
4. Mechanistic foundations: induction heads and circuit formation
Why does structure transfer? Because transformers don't store training text — they compile it into circuits, and circuits are portable.
The canonical example is the induction head: a two-layer attention motif that finds previous occurrences of the current token and boosts the tokens that followed it. Its formation marks a sharp phase transition in in-context learning ability across scales (Olsson et al., 2022). Edelman et al. give the cleanest controlled account: transformers trained to in-context-learn Markov chains pass through discrete phases — uniform predictions, then a suboptimal unigram circuit, then a rapid transition to a bigram-computing statistical induction head approaching Bayes-optimal prediction (Edelman et al., 2024). Each phase is a distinct subnetwork, built by the data and recycled by later learning.
The data side of the story matters just as much. Zhang et al. trained GPT-2 models on elementary cellular automata spanning Wolfram's complexity classes and evaluated them downstream: models trained on trivially ordered rules learned trivial solutions and transferred poorly; models trained on maximally chaotic rules were effectively fitting noise and also transferred poorly; models trained on Class IV rules — complex but structured — performed best on reasoning and chess-move prediction (Zhang et al., 2024). There is a sweet spot, and it is not “more data” but “data at the edge of chaos.”
This is the mechanistic justification for epiplexity-guided curricula:
- Structured, multi-step dependencies force the formation of dedicated subnetworks — induction heads, counting routines, state-tracking loops.
- These circuits are reusable: they get recruited by downstream zero-shot reasoning and fine-tuning, including on distributions the pre-training loss never measured.
- In-distribution loss is blind to this inventory. Two checkpoints with identical validation loss can differ widely in which circuits exist inside them — which is why a curriculum optimized purely on aggregate loss can plateau while one optimized on where loss drops come from does not.
High-\(S_T\) domains are, almost by definition, the ones that make circuit formation unavoidable: they contain dependencies shallow statistics cannot absorb, so the only path to lower loss is building another mechanism.
5. Systems architecture: engineering the feedback seam
None of the above survives contact with production unless the control loop can be bolted onto a training engine without slowing it down. Our Rust pre-training stack (nanogpt-rs) implements this seam in three layers, and the constraints it enforces are the interesting part.
5.1 One trait, one seam
The autodiff engine knows nothing about curricula. Everything flows through a single observer interface:
// src/train.rs — the ONLY seam between the training loop and later phases.
pub trait TrainObserver: Send {
fn on_step(
&mut self,
step: u64,
domain_id: usize,
loss: f32,
tokens_delta: u64,
// Cumulative tokens seen per domain as of this step, length K.
domain_tokens: &[u64],
);
}
The dependency rule is absolute: train.rs imports no curriculum modules. The observer is injected as Box<dyn TrainObserver>, so the engine ships with a NoopObserver and pays nothing when telemetry is off. This modularity is what lets the numerical core stay frozen while the data-policy layer iterates weekly.
5.2 Non-blocking orchestration: the scaling-law sidecar
Fitting nonlinear least squares with L-BFGS-B belongs nowhere near the accelerator. The tracker accumulates per-domain (tokens, loss) observations in a fixed-size ring buffer and, every \(\nu = 500\) steps, hands windows to a Python sidecar over one-JSON-object-per-line IPC:
Rust trainer —stdin→ sidecar.py —stdout→ Rust trainer
| L(n)=ε+β·n^(-α) |
| {"domain_id":2, L-BFGS-B, numpy/scipy | {"domain_id":2,
| "history":[...], SingleDomain | | "alpha":..,"beta":..,
| "phase":"Cross"<--- CrossDomain γ-mixing ---> | "gamma":[...],"ok":true}
The contract is failure-tolerant by construction:
- Malformed input or internal errors produce
{"ok": false}responses — the sidecar never exits, so the connection stays warm. - The Rust wrapper enforces a timeout by killing and respawning the child; every fallible path returns a typed
SidecarErrorandfit()never panics. - Fits are gated before use: at least 20 observations, second half of the history strictly below the first, non-trivial variance and range. A domain failing five consecutive cycles is marked ineligible for the rest of the run with a forced-zero gradient — bad fits degrade to no steering, never to wrong steering.
Because the solver lives out-of-process, the GPU pipeline stays saturated: the fit cycle costs milliseconds of CPU time amortized over hundreds of steps.
5.3 From fits to weights — and back to the sampler
The fitted parameters become analytic marginal gradients, computed from the post-cycle snapshot so a fresh fit immediately influences its own weight:
// Analytic ∂&Shat;/∂n_j under the fitted multi-domain law:
// Σ_m α_m β_m γ_{m,j} (Σ_k γ_{m,k} n_k)^(-α_m - 1) · t_total
fn compute_marginal_grad(params, domain_tokens, j) -> f64 { /* closed form */ }
// score -> stable softmax -> momentum blend -> exact floor projection.
let new_w = project_floor(
&blend(softmax(scores / tau), prev_weights, mu),
pi_min,
);
debug_assert!((new_w.iter().sum<f32>() - 1.0).abs() < 1e-4);
Three details earn their place in production code:
- Momentum (\(\mu\)) on the weights prevents oscillation from noisy per-window fits.
- An exact simplex floor (\(\pi_{min}\)) via water-filling guarantees no domain is ever fully starved — clip-then-renormalize would silently break the floor, so the projection is exact and asserted.
- Phase discipline: the current release is observational — weights are written to
weights.jsonlbut sampling remains uniform, and the loop-closing path is a separateWeightProvidertrait behind a mutex. Ship the measurement layer, validate its outputs offline, and only then let it drive the loader.
That last point generalizes into the validation principle: verify information-theoretic schedulers against controlled synthetic baselines before deploying at scale. Multi-domain mixtures with known ground-truth dynamics — one fast-decaying domain, one pure-noise domain, one redundant domain — turn the theory's predictions (noise domains get down-weighted, structured domains win share) into unit assertions. Only after the controller passes there does it earn a slot in a real run.
┌──────────────┐ batches ┌────────────────┐ Box<dyn TrainObserver>
│ ShardedData │────────────▶│ Training Loop │───────────────┐
│ Loader │ │ (autodiff + │ ▼
│ (weights ◀──┐│ │ kernels) │ ┌─────────────────┐
└─────────────┼┘ └────────────────┘ │ Epiplexity │
│ WeightProvider │ Tracker │
│ (Phase 3 seam) │ ring buffer + │
└───────────────────────────────────────┤ epi_accum │
└───────┬─────────┘
ν-step fit cycle│JSONL IPC
▼
┌───────────────────┐
│ sidecar.py │
│ scaling-law fits │
└───────────────────┘
6. Current literature to follow
| Paper | Contribution | Link |
|---|---|---|
| Villalobos et al., 2024 | The Data Wall — human-generated public text exhausted as a scaling resource between 2026–2032 | arXiv:2211.04325 |
| Muennighoff et al., 2023 | Data-constrained scaling laws; repetition value decays past ~4 epochs; code adds ~2× effective data | arXiv:2305.16264 |
| Finzi et al., 2026 | Epiplexity (\(S_T\)) — information for computationally bounded observers | arXiv:2601.03220 |
| Su et al., 2026 | EpiSelect/EpiGen — cross-domain scaling laws and epiplexity-guided selection for OOD generalization | arXiv:2608.11746 |
| Jiang et al., 2024 | ADO — online per-domain scaling laws steer the data mixture with no proxy model | arXiv:2410.11820 |
| Xie et al., 2023 | DoReMi — proxy-model mixture optimization reaches baseline 2.6× faster | arXiv:2305.10429 |
| Olsson et al., 2022 | Induction heads — circuit-level account of the in-context learning phase transition | Transformer Circuits |
| Edelman et al., 2024 | Statistical induction heads — staged circuit formation on Markov chains | arXiv:2402.11004 |
| Zhang et al., 2024 | Intelligence at the Edge of Chaos — downstream gains peak for structured-complex training data | arXiv:2410.02536 |
| Rissanen, 1996 | Stochastic complexity — Fisher-information codelengths; the prequential foundation of MDL | IEEE DOI |
| Grünwald, 2007 | The MDL Principle — prequential analysis as the bridge from compression to learning | MIT Press |
7. Closing idea
The arc of the series so far: retrieval infrastructure (001), uncertainty depletion (002), the epiplexity decomposition (003) — and now the engine that acts on it.
The cleanest way to state this post's thesis: once I/O is solved, pre-training becomes a control problem. The plant is the model; the sensors are per-domain loss trajectories; the reference signal is maximal structural information absorbed per FLOP; the actuators are sampling weights. Every piece of that loop exists today — parametric scaling laws for sensing, cross-domain gradients for planning, observer interfaces and sidecars for actuation.
Uniform shuffling is not a neutral default. It is a controller stuck at its initialization, spending compute where the data is plentiful instead of where the structure is dense. The engineering task is to close the loop.