Spectral retrieval glossary
Canonical, short definitions of the terms used across Genefold documentation. Each entry is self-contained and links to a page covering the concept in depth.
Spectral intelligence
Spectral intelligence is the use of eigenvalues, eigenvectors, and related graph operators to describe structure in embedding spaces for retrieval, monitoring, and analysis. It differs from purely geometric similarity by capturing corpus-level structure such as manifold shape, redundancy, and connectivity.
See the full definition.
Spectral vector search
Spectral vector search ranks embedding candidates using the structure of a graph built over the embedding space, combined with direct geometric similarity. In ArrowSpace this means blending cosine similarity with a bounded spectral score difference at query time.
Graph Laplacian
The matrix L = D − A formed from a graph’s degree matrix D and adjacency matrix A. Its quadratic form measures how strongly a vector disagrees with its neighbourhood, which retrieval methods use as smoothness energy. Distinct from the Laplacians used in graph neural networks, because the graph is built over the embedding geometry itself.
Laplacian spectrum
The set of eigenvalues and eigenvectors of a graph Laplacian. The spectrum summarises global graph properties: connectivity, cluster separation, and the frequency structure of the data. It is the source of the per-item scores used in spectral retrieval.
Eigenvalue distribution
The distribution of a Laplacian’s eigenvalues over a corpus. Used as a compact structural summary of the dataset: changes in this distribution between time windows are a spectral drift signal.
Embedding manifold
The region of high-dimensional space where a corpus’ embeddings concentrate. Retrieval and monitoring quality depend on whether queries and indexed items lie on the same manifold. See also out-of-distribution retrieval.
Manifold-aware retrieval
Retrieval that ranks candidates using the geometry of the embedding manifold, not only pairwise distances. Spectral retrieval is a manifold-aware method: its scores derive from the graph over the data. Contrast with pure nearest-neighbour ranking, which uses only local proximity.
Long-tail retrieval
Retrieval performance on infrequent or peripheral items, away from dense cluster centres. Standard similarity signals degrade on the tail because local neighbourhoods are sparse; spectral ranking is evaluated specifically on tail behaviour (Engineering 001 reports tail-shape metrics on CVE).
See Engineering 001 for the protocol.
Out-of-distribution retrieval
Retrieval behaviour when a query falls outside the support of the corpus manifold. Standard indexes still return nearest items, so the failure is silent; spectral diagnostics expose structural symptoms that can trigger abstention or rerouting.
See OOD vector retrieval.
Embedding drift
A change over time in the distribution of embeddings for the same input population, caused by model updates, preprocessing changes, or data shift. Monitored with embedding-level statistics and with spectral signals such as eigenvalue-distribution change.
Spectral drift
A change over time in Laplacian-derived quantities (eigenvalue distributions, bounded spectral scores) of a corpus graph. A structural counterpart to embedding-level drift indicators.
Tau-modulation
The runtime adjustment of the parameter blending geometric similarity and spectral signal during retrieval. In the SPIN formulation, τ = 1.0 is pure cosine retrieval and τ → 0 is fully spectral retrieval. Implementations may parameterise the blend differently.
Lambda-tau (λτ, taumode)
The bounded per-item spectral score in ArrowSpace, computed from graph Laplacian smoothness (Rayleigh quotient) and edge-wise dispersion, and normalised to [0, 1). Boundedness keeps scores comparable across collections, windows, and model updates.
SPIN (SPectral INdexing)
The task-specific instantiation of spectral indexing for vector search in ArrowSpace: items indexed with λτ scores, queries answered with a blend of cosine and spectral signals. Introduced in the Energy Dispersion Networks paper.
Graph wiring
The parameterised construction of edges in a feature-space graph (neighbour selection and weighting) before computing the Laplacian. Wiring choices determine which structure the spectral signal sees; studied empirically in the Epiplexity and Graph Wiring paper.
Epiplexity
The structural information a dataset generates, as defined in Genefold’s research line on first principles. Used to reason about how much learnable structure a corpus yields per unit of compute, and to guide graph construction. See Engineering 003 for the framing.
HNSW
Hierarchical Navigable Small World: a standard approximate nearest-neighbour index built on a hierarchical proximity graph, optimised for low-latency top-k lookup. It does not expose corpus-structure diagnostics by itself. See spectral search vs HNSW.
Cosine similarity
The cosine of the angle between two vectors: a pairwise, local similarity signal insensitive to vector magnitude. The baseline signal in embedding retrieval; spectral retrieval uses it as one of two signals rather than replacing it.
Approximate nearest neighbour search
The family of algorithms (HNSW, IVF, LSH and relatives) that trade exact recall for query latency on large vector collections. Their indexes encode proximity for lookup, not corpus structure for diagnostics.