Spectral vector search for embeddings
How spectral vector search ranks candidates, what the bounded spectral score contributes, and how tau-modulation switches retrieval behaviour at runtime.
Spectral vector search ranks embedding candidates using the structure of a graph built over the embedding space and the spectrum of its graph Laplacian, in addition to direct geometric similarity. It is designed to preserve structural information that pure cosine ranking does not represent.
Key properties
- Bounded spectral score. Each item carries a score lambda-tau (λτ) in [0, 1), computed from graph Laplacian smoothness energy and edge-wise dispersion. Bounded scores stay comparable across collections and model updates.
- Two signals at query time. Retrieval combines cosine similarity with the spectral score difference; the blend is a runtime parameter, not a fixed index property.
- Runtime adaptability. Tau-modulation changes the blend during a session, supporting pipelines that need different behaviour for exploration and for focused evidence gathering.
- Structural diagnostics. Because scores derive from corpus structure, the same artifacts expose OOD behaviour and drift signals.
How it works
- Embed the corpus.
- Construct a graph over the vectors with a parameterised wiring strategy (see graph Laplacian retrieval).
- Compute each item’s bounded spectral score λτ from the Laplacian.
- At query time, score each candidate with a blend of cosine similarity and spectral difference.
- Rank and return; optionally modulate τ between query phases.
Mathematical formulation
The spectral difference between query and item is the difference of their bounded scores (Engineering 001):
where λτx = R(x) / (R(x) + ε) ∈ [0, 1) and R(x) is the Rayleigh quotient of the item on the graph Laplacian. The combined distance is
In this parameterisation τ = 1.0 is pure cosine retrieval, τ → 0 is fully spectral retrieval, and intermediate values mix the two. Other implementations may parameterise the same blend differently; check the library documentation.
Tau-modulation
The blend parameter is a runtime control, not a one-off hyperparameter. In long-context retrieval pipelines, an early phase with τ close to 1.0 leans on geometric similarity to establish context; later phases with lower τ embed more spectral signal to filter against corpus structure. Engineering 001 reports τ = 0.42 as the strongest taumode configuration on the CVE dataset under its protocol, and gains on TREC-COVID concentrated around hybrid settings τ ≈ 0.65–0.80.
Comparison with cosine-only ranking
| Property | Cosine similarity | Spectral retrieval (SPIN) |
|---|---|---|
| Primary signal | Pairwise vector angle | Graph and spectral structure plus vector angle |
| Index structure | Vector index | Graph-derived spectral artifacts |
| Score range | Unbounded depends on model | Bounded [0, 1) for λτ |
| Runtime control | None | Tau-modulation blend |
| Extra cost | None | Graph build and spectral artifacts |
Evidence
Engineering 001 evaluates SPIN against cosine (tau = 1.0), hybrid (τ = 0.72), and taumode (τ = 0.42) configurations on the CVE corpus (347,113 reports) and TREC-COVID (about 171,000 scientific articles). On CVE, taumode achieved 29.9% lower tail coefficient of variation at the evaluated setting; on TREC-COVID, intermediate τ values improved Relevance Recall and NDCG at k = 10 over cosine. Consult the full note for the evaluation protocol, tables, and the exact configurations before reusing these numbers.
A live implementation over the NVD corpus is available in the CVE search demo.
Implementation
- Python: pyarrowspace (
pip install arrowspace). - Rust: arrowspace-rs (
cargo add arrowspace). - LLM agents: arrowspace-mcp exposes spectral search to models.
Limitations
Spectral retrieval adds build and storage cost, and its benefit depends on graph construction, data distribution, evaluation protocol, and query type. Published evidence compares it against cosine baselines on specific datasets; comparisons with HNSW-class indexes are qualitative (see spectral search vs HNSW).