How we build spectral intelligence.
Notes on systems, algorithms, and the engineering choices behind ArrowSpace. Written for the teams shipping embedding-heavy software in production.
Synthetic Data That Teaches: A Feedback Loop for Training Beyond Static Corpora
Data-selection methods can make better use of a finite corpus, but they cannot create new training signal once data becomes scarce. Our proposal: reward a generator for one thing only — whether its batch improves a learner on a fixed, held-out reference buffer.
Genefold First Principles: The Curriculum Engine
Once the pipes are fast enough, the bottleneck stops moving bytes and starts moving meaning. Build the engine that chases epiplexity online: measure what each data domain teaches, forecast its marginal return, and steer the sampler.
Genefold First Principles: From Entropy to Epiplexity
Entropy measures uncertainty. But not all uncertainty is learnable. When compute is finite, the question shifts from how much uncertainty exists to how much structure we can extract per FLOP.
Genefold First Principles: Depletion of Uncertainty
Entropy is the irreducible uncertainty of the source. Cross-entropy measures how close the model's distribution is to the truth — training depletes the gap: the KL divergence.
Designing ArrowSpace and graph wiring
How graph wiring in the feature-space Laplacian powers spectral indexing for vector retrieval, and what the experimental evidence shows on CVE and TREC-COVID.