Spectral vector search for embeddings

How spectral vector search ranks candidates, what the bounded spectral score contributes, and how tau-modulation switches retrieval behaviour at runtime.

Published 2026-09-21 by Genefold AI

Spectral vector search ranks embedding candidates using the structure of a graph built over the embedding space and the spectrum of its graph Laplacian, in addition to direct geometric similarity. It is designed to preserve structural information that pure cosine ranking does not represent.

Key properties

  • Bounded spectral score. Each item carries a score lambda-tau (λτ) in [0, 1), computed from graph Laplacian smoothness energy and edge-wise dispersion. Bounded scores stay comparable across collections and model updates.
  • Two signals at query time. Retrieval combines cosine similarity with the spectral score difference; the blend is a runtime parameter, not a fixed index property.
  • Runtime adaptability. Tau-modulation changes the blend during a session, supporting pipelines that need different behaviour for exploration and for focused evidence gathering.
  • Structural diagnostics. Because scores derive from corpus structure, the same artifacts expose OOD behaviour and drift signals.

How it works

  1. Embed the corpus.
  2. Construct a graph over the vectors with a parameterised wiring strategy (see graph Laplacian retrieval).
  3. Compute each item’s bounded spectral score λτ from the Laplacian.
  4. At query time, score each candidate with a blend of cosine similarity and spectral difference.
  5. Rank and return; optionally modulate τ between query phases.

Mathematical formulation

The spectral difference between query and item is the difference of their bounded scores (Engineering 001):

sim_λ(q, i) = λτq − λτi

where λτx = R(x) / (R(x) + ε) ∈ [0, 1) and R(x) is the Rayleigh quotient of the item on the graph Laplacian. The combined distance is

distSPIN = τ · simcos(q, i) + (1 − τ) · simλ(q, i)

In this parameterisation τ = 1.0 is pure cosine retrieval, τ → 0 is fully spectral retrieval, and intermediate values mix the two. Other implementations may parameterise the same blend differently; check the library documentation.

Tau-modulation

The blend parameter is a runtime control, not a one-off hyperparameter. In long-context retrieval pipelines, an early phase with τ close to 1.0 leans on geometric similarity to establish context; later phases with lower τ embed more spectral signal to filter against corpus structure. Engineering 001 reports τ = 0.42 as the strongest taumode configuration on the CVE dataset under its protocol, and gains on TREC-COVID concentrated around hybrid settings τ ≈ 0.65–0.80.

Comparison with cosine-only ranking

PropertyCosine similaritySpectral retrieval (SPIN)
Primary signalPairwise vector angleGraph and spectral structure plus vector angle
Index structureVector indexGraph-derived spectral artifacts
Score rangeUnbounded depends on modelBounded [0, 1) for λτ
Runtime controlNoneTau-modulation blend
Extra costNoneGraph build and spectral artifacts

Evidence

Engineering 001 evaluates SPIN against cosine (tau = 1.0), hybrid (τ = 0.72), and taumode (τ = 0.42) configurations on the CVE corpus (347,113 reports) and TREC-COVID (about 171,000 scientific articles). On CVE, taumode achieved 29.9% lower tail coefficient of variation at the evaluated setting; on TREC-COVID, intermediate τ values improved Relevance Recall and NDCG at k = 10 over cosine. Consult the full note for the evaluation protocol, tables, and the exact configurations before reusing these numbers.

A live implementation over the NVD corpus is available in the CVE search demo.

Implementation

Limitations

Spectral retrieval adds build and storage cost, and its benefit depends on graph construction, data distribution, evaluation protocol, and query type. Published evidence compares it against cosine baselines on specific datasets; comparisons with HNSW-class indexes are qualitative (see spectral search vs HNSW).