Status: Research note. This post presents a proposed training objective and system architecture. Code blocks are illustrative pseudocode unless stated otherwise, and benchmark numbers are reported results from the cited EpiGen paper (Su et al., 2026) — treat them as directional pending independent replication.
1. The Ceiling of Data Selection and the Failure Mode of Naive Synthesis
Pre-training has run on a simple premise: scrape trillions of tokens, shuffle uniformly, scale compute. That era is ending. Villalobos et al. project public human text exhausted between 2026–2032 (median 2028) (Villalobos et al., 2024). Re-reading buys roughly four epochs before returns decay to zero (Muennighoff et al., 2023). Selection methods — ADO (Jiang et al., 2024), DoReMi (Xie et al., 2023), EpiSelect (Su et al., 2026) — help choose which existing tokens see the next batch. They are powerful but bounded: they can only redistribute a finite pool.
The natural next step is to generate new data. But recursive generation without verification collapses. Gerstgrasser et al. (2024) show that models trained on their own outputs suffer distribution tail erosion: rare but structurally critical patterns vanish, diversity collapses, and the model converges to a degenerate mode. AlphaZero avoids this via an external rules engine (chess/go); SPIRAL uses a code interpreter. Natural language has no compiler — no executable ground truth. The reward must come from within the learning process.
# Illustrative pseudocode — what goes wrong
def naive_self_generate(model, steps, prompt):
"""Generate -> add to training -> repeat. No verification."""
for step in range(steps):
synthetic = model.generate(prompt, max_new_tokens=512)
model.train_on(synthetic) # No reward signal, no quality filter
prompt = synthetic # Drift accumulates
return model
# Failure modes observed empirically:
# 1. Mode collapse: output diversity drops exponentially
# 2. Tail erosion: rare but critical patterns (multi-step reasoning,
# formal syntax, algorithmic structure) disappear first
# 3. Reward hacking: model learns to produce "easy" tokens that
# lower loss without adding structure
The fix is not a bigger generator. It is a dual-agent loop where the learner's own loss reduction on an evaluation buffer serves directly as the generator's reward signal.
2. Theoretical Formulation — Reward as the Epiplexity Gradient
EpiGen formalizes this as a coupled two-agent system (Su et al., 2026):
- Generator $P_{\theta^g}$: produces candidate synthetic token batches $\mathcal{X}_t$
- Learner $P_{\theta}$: evaluates pre-update loss on reference buffer $\mathcal{B}$, executes $K$ gradient steps on $\mathcal{X}_t$, and re-evaluates post-update loss on $\mathcal{B}$
The key question is simple: after the learner trains on a generated batch, does it become better at data it did not generate? The reward at step $t$ is the loss reduction measured by the learner over the evaluation buffer:
$$r_t = \Delta \hat{S}(t) = L^{\text{pre}}_t(\mathcal{B}) - L^{\text{post}}_t(\mathcal{B}) = \frac{1}{|\mathcal{B}|} \sum_{x \in \mathcal{B}} \left( \mathcal{L}(\theta_{t-1}, x) - \mathcal{L}(\theta_t, x) \right)$$
In this article, we use epiplexity gradient as an operational name for the improvement a learner obtains on a fixed evaluation buffer after an update. It is a measurable proxy for whether a synthetic batch teaches transferable structure rather than merely reducing loss on itself.
The prequential view (Rissanen, 1996; Grünwald, 2007) motivates this signal: the learner acts as a sequential coder, and its cumulative codelength
$$\sum_i -\log q(x_i \mid x_{<i})$$
equals validation loss in different units. A held-out-buffer loss delta after a finite update is related to, but not identical with, a full prequential description-length measure — we use it here as a practical, measurable stand-in. Learning progress, in this view, is compression progress.
Why this reward avoids collapse:
| Input type | $L^{\text{pre}}$ | $L^{\text{post}}$ | $r_t = \Delta \hat{S}$ |
|---|---|---|---|
| Pure noise | High | High (no learning) | $\approx 0$ |
| Trivial redundancy | Low | Low (already known) | $\approx 0$ |
| Structured, learnable | High | Lower | > 0 |
The generator only earns reward for data that teaches the learner something new and reusable. This is the epiplexity gradient: structural information extractable per compute cycle.
A concrete illustration: a generator proposes two synthetic batches for a small language learner. One consists of repeated familiar sentence templates; the other contains short symbolic transformations that share structure with held-out examples. Both may be fluent. Only the latter should earn a high reward — illustrative numbers:
| Batch | $L^{\text{pre}}$ | $L^{\text{post}}$ | $r_t$ | Verdict |
|---|---|---|---|---|
| A: familiar sentence templates | 2.10 | 2.09 | 0.01 | Already mastered — no reward |
| B: symbolic transformations sharing structure with held-out probes | 2.10 | 1.85 | 0.25 | Transferable structure — high reward |
Illustrative pseudocode — proposed training objective, not a verified implementation:
import copy
import torch
import torch.nn as nn
class EpiGen(nn.Module):
"""
Policy gradient over synthetic data generator.
Reward = learner validation loss reduction on evaluation buffer B.
"""
def __init__(self, generator, learner, buffer, k_steps=10, ema_decay=0.99):
super().__init__()
self.generator = generator # P_{\theta^g} — trainable policy
self.learner = learner # P_{\theta} — updated online
self.buffer = buffer # Reference evaluation buffer B
self.k_steps = k_steps # Learner gradient steps per generator step
self.ema_baseline = 0.0
self.ema_decay = ema_decay
def compute_reward(self, synthetic_batch):
"""r_t = L_pre(B) - L_post(B) after K learner steps on synthetic_batch."""
# 1. Pre-update evaluation of the learner on the buffer
with torch.no_grad():
l_pre = self.learner.evaluate(self.buffer)
# 2. Simulate/execute K gradient steps on detached synthetic batch
learner_sim = self._clone_learner()
for _ in range(self.k_steps):
learner_sim.train_step(synthetic_batch.detach())
# 3. Post-update evaluation of the updated learner on the same buffer
with torch.no_grad():
l_post = learner_sim.evaluate(self.buffer)
reward = l_pre - l_post # Positive reward = learner improved on the buffer
return reward, learner_sim
def forward(self, batch_size):
# Generate synthetic batch with log-probabilities
synthetic_batch, log_probs = self.generator.generate_with_logprobs(batch_size)
# Compute epiplexity-gradient reward
reward, updated_learner = self.compute_reward(synthetic_batch)
# Advance true learner state
self.learner.load_state_dict(updated_learner.state_dict())
# EMA baseline for variance reduction
self.ema_baseline = self.ema_decay * self.ema_baseline + (1 - self.ema_decay) * reward
advantage = reward - self.ema_baseline
# REINFORCE loss: maximize expected advantage
pg_loss = -(advantage * log_probs).mean()
return pg_loss, {"reward": reward, "advantage": advantage}
def _clone_learner(self):
"""Fast learner clone for step simulation."""
clone = copy.deepcopy(self.learner)
clone.train()
return clone
The policy gradient uses an EMA baseline for variance reduction — standard REINFORCE with a moving-average control variate. This setup avoids an external critic or learned reward model, the usual collapse points — but it does not eliminate failure modes: the generator can overfit the buffer, exploit shallow transfer signals, or destabilize the learner through non-stationary updates. We return to these risks in Section 6.
3. Systems Architecture — The Replay Buffer and Tensor Lifecycle
The evaluation buffer $\mathcal{B}$ is the linchpin. It serves two roles simultaneously:
- Reference probe: measures the learner's current capability (pre-update loss)
- Transfer anchor: prevents mode collapse by mixing real data
Proposed buffer interface — illustrative architecture, not compiling repository code:
// Proposed buffer lifecycle (illustrative)
use burn::tensor::Tensor;
use burn::backend::Autodiff;
use std::sync::Arc;
pub struct EvaluationBuffer<B: Autodiff> {
/// Fixed probe set — never updated during training
probe_tokens: Tensor<B, 2, Int>,
probe_labels: Tensor<B, 2, Int>,
/// Real-data anchor slices (25% OpenWebText)
anchor_tokens: Tensor<B, 2, Int>,
anchor_labels: Tensor<B, 2, Int>,
/// Metadata for stratified sampling
domain_ids: Vec<usize>,
strata_bounds: Vec<usize>,
}
impl<B: Autodiff> EvaluationBuffer<B> {
pub fn new(
probe_data: &[TokenizedSequence],
anchor_ratio: f32, // 0.25 default
device: &B::Device,
) -> Self {
// Probe: diverse held-out domains (reasoning, code, math, etc.)
// Anchor: fixed OpenWebText slices, never regenerated
let (probe_tokens, probe_labels) = Self::pack_probe(probe_data, device);
let (anchor_tokens, anchor_labels) = Self::pack_anchor(anchor_ratio, device);
Self {
probe_tokens,
probe_labels,
anchor_tokens,
anchor_labels,
domain_ids: Self::compute_domains(probe_data),
strata_bounds: Self::compute_strata(probe_data),
}
}
/// Combined buffer for validation: probe + anchor
pub fn validation_batch(&self) -> (Tensor<B, 2, Int>, Tensor<B, 2, Int>) {
let tokens = Tensor::cat(vec![self.probe_tokens.clone(), self.anchor_tokens.clone()], 0);
let labels = Tensor::cat(vec![self.probe_labels.clone(), self.anchor_labels.clone()], 0);
(tokens, labels)
}
/// Asynchronous non-blocking sampling for generator training
pub async fn sample_async(&self, batch_size: usize) -> BufferBatch<B> {
// Stratified sampling across domains to maintain coverage
let indices = self.stratified_sample(batch_size).await;
BufferBatch {
tokens: self.probe_tokens.index_select(0, indices.clone()),
labels: self.probe_labels.index_select(0, indices),
domain_ids: self.domain_ids[indices].to_vec(),
}
}
}
Memory and systems constraints:
- Buffer size: typically 4K–16K tokens (small enough for fast validation, large enough for signal)
- Validation runs every $N$ generator steps (e.g., $N=10$) — amortized cost
- Gradient checkpointing on the $K$-step learner simulation: only final loss retained, intermediates recomputed
- Non-blocking IPC: generator runs on GPU; validation buffer evaluation runs on CPU sidecar (Python) via JSONL streaming — zero GPU stall
Cold-start constraint: The learner must start from a pretrained checkpoint. Random initialization yields flat loss landscapes — $\nabla_\theta L \approx 0$ everywhere — so the reward signal is pure noise. Pretrained weights provide the linguistic structure that makes loss reduction a meaningful gradient. Without them, there's no signal for the policy gradient to find structure in token space.
4. Empirical Transfer and Circuit Formation
The payoff is not lower pre-training perplexity — it is out-of-distribution transfer.
GLUE Fine-Tuning Transfer
The following table adapts results reported in the EpiGen paper (Su et al., 2026); it is reproduced here as directional evidence, not as an independent reproduction. All tasks report accuracy except CoLA (Matthews Correlation Coefficient due to class imbalance):
| Method | CoLA | SST-2 | MRPC | QQP | MNLI | QNLI | RTE | Average |
|---|---|---|---|---|---|---|---|---|
| Pretrained GPT-2 Base | 0.264 | 0.930 | 0.779 | 0.891 | 0.814 | 0.884 | 0.639 | 0.743 |
| FrozenGen (Stationary Synth) | 0.388 | 0.924 | 0.774 | 0.893 | 0.817 | 0.882 | 0.632 | 0.759 |
| Perplexity-Reward (PPL) | 0.380 | 0.911 | 0.770 | 0.892 | 0.815 | 0.882 | 0.643 | 0.756 |
| No-Buffer Reward | 0.367 | 0.919 | 0.767 | 0.892 | 0.815 | 0.881 | 0.661 | 0.757 |
| EpiGen (Epiplexity Guided) | 0.422 | 0.920 | 0.777 | 0.892 | 0.817 | 0.886 | 0.675 | 0.770 |
Key takeaway (as reported in Su et al., 2026): without consuming additional real tokens, the epiplexity-guided condition improved the reported mean GLUE score from 0.743 to 0.770 while avoiding the representation collapse of the perplexity/no-buffer objectives. Treat these numbers as directional pending multi-seed replication.
Zero-Shot Reasoning
Same source and same caveat (Su et al., 2026):
| Benchmark | Frozen | Perplexity-Reward | EpiGen |
|---|---|---|---|
| GSM8K (math) | 12.4% | 14.7% | 18.9% |
| BBH (reasoning) | 28.1% | 30.2% | 34.6% |
| CodeEval (pass@1) | 15.3% | 17.1% | 21.4% |
Mechanistic Hypotheses: Circuits Under Compression Pressure
Why might epiplexity-gradient synthesis transfer? A testable hypothesis is that synthetic batches rewarded for transferable loss reduction preferentially train mechanisms associated with longer-range dependency tracking, including induction-like attention patterns. The evidence below is consistent with this, but verifying it requires causal interventions and controlled comparisons — not attention maps alone.
- Induction heads: Olsson et al. (2022) showed these two-layer attention motifs emerge when the model must copy previous context to predict next tokens. EpiGen data — rich in multi-step dependencies — is the kind of data under which such motifs are known to form; whether it speeds their formation in this system is an open empirical question.
- Multi-step state tracking: Edelman et al. (2024) demonstrated staged circuit formation on Markov chains: uniform → unigram → statistical induction head. EpiGen's reward favors data that demands these circuits, if the hypothesis holds.
- Edge of chaos: Zhang et al. (2024) found downstream reasoning peaks for Class IV (complex-structured) cellular automata training data. A well-calibrated epiplexity reward would discover this regime automatically — the reward points toward structured complexity by construction.
Example diagnostic, not a complete circuit test:
# Attention-pattern probe (from Su et al., 2026)
def verify_induction_heads(model, layer, head):
"""Test if attention head implements copying mechanism."""
# Input: [A, B, C, A, B, ?] -> should attend to previous A,B and predict C
test_seq = torch.tensor([[1, 2, 3, 1, 2, 0]]) # 0 = predict
with torch.no_grad():
attn_weights = model.get_attention(layer, head, test_seq)
# Induction pattern: position 5 attends to positions 1,2 (previous A,B)
induction_score = attn_weights[0, 5, 1:3].sum().item()
return induction_score > 0.5 # Threshold from Olsson et al.
# EpiGen runs reportedly show induction-like patterns forming faster than
# uniform baselines at matched compute (Su et al., 2026). Verifying this
# requires causal interventions, not attention maps alone.
5. Synthesis & Architectural Outlook
EpiSelect and EpiGen are not separate systems — they are two phases of the same feedback loop:
┌─────────────────────────────────────────────────────────────────┐
│ EPIPLEXITY FEEDBACK LOOP │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ EpiSelect │────▶│ Training │────▶│ EpiGen │ │
│ │ (measure) │ │ (learn) │ │ (generate) │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Evaluation Buffer B (probe + anchor) │ │
│ │ Measures: ΔL = L_pre - L_post = Epiplexity gain │ │
│ └─────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────┐ │
│ │ Cross-domain │ │
│ │ scaling law fits │ │
│ │ (γ-mixing matrix) │ │
│ └─────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────┘
EpiSelect (004) measures per-domain epiplexity via cross-domain scaling laws and steers the sampler toward the highest-structure sources.
EpiGen (this post) generates new data where natural sources fall short, using the same epiplexity signal — loss reduction on $\mathcal{B}$ — as its reward.
The unification: both optimize $\frac{\partial \hat{S}}{\partial n_k}$ — marginal structural information per token. EpiSelect does it by selection (reallocating existing tokens); EpiGen does it by synthesis (creating new tokens where the gradient is highest).
6. What Can Go Wrong: Failure Modes and Open Risks
The proposal has failure modes of its own. Naming them is part of the design, not a disclaimer:
- Reward hacking and buffer overfitting: the generator optimizes $\Delta L$ on $\mathcal{B}$, not on the world. Goodhart's law applies — batches can be crafted to move buffer loss without teaching transferable structure. Rotating and held-out buffer splits are the first-line defense.
- Non-stationarity: the reward is defined against a learner that changes after every step. The same batch yields different rewards at different times, which complicates credit assignment and variance control for the policy gradient.
- Evaluation leakage and contamination: synthetic data that memorizes or closely paraphrases buffer content inflates the reward. The real-data anchor (25% of $\mathcal{B}$) mitigates but does not eliminate this; contamination checks remain necessary.
- Computational cost: each generator step costs $K$ learner updates plus two buffer evaluations. Gradient checkpointing and amortized validation (every $N$ generator steps) bound the overhead but do not remove it.
- Residual drift: a better reward reduces distribution tail erosion; it does not abolish it. Diversity monitoring on generated batches remains necessary.
7. Closing Perspective
Pre-training is shifting from static scraping to continuous, self-directed information extraction.
As high-quality public text becomes more constrained, synthetic and interactive data are likely to become increasingly important complements to curation, reuse, licensing, and improved training efficiency. Among these, generation has a property the others lack: it can target the learner's current blind spots directly. The reward signal this proposal relies on is already computed in every training run — validation loss. It needs no external compiler, no human labels, and no learned reward model, though it introduces the failure modes outlined above.
This is self-play via epiplexity: the generator proposes, the learner disposes, the buffer judges.
Next validation milestone: multi-seed replication of the transfer results above and publication of the evaluation configuration. Until then, treat this note as a research proposal with directional evidence.