Spectral search vs HNSW
A comparison of spectral retrieval and hierarchical navigable small-world (HNSW) indexes: what each signal captures, what each costs, and where they fit in a production stack.
HNSW is an approximate nearest-neighbour index that traverses a hierarchical proximity graph for fast top-k lookup. Spectral search (ArrowSpace/SPIN) combines geometric similarity with graph-Laplacian spectral scores. They solve overlapping but different problems: HNSW optimises fast approximate recall; spectral retrieval adds corpus-structure signals and diagnostics.
Comparison
| Property | Cosine similarity | HNSW | ArrowSpace spectral retrieval |
|---|---|---|---|
| Primary signal | Pairwise vector angle | Approximate graph traversal over a proximity graph | Graph and spectral structure plus vector angle |
| Index structure | Flat vector index | Navigable small-world graph with hierarchy | Graph-derived spectral artifacts |
| Query-time behaviour | Exhaustive scan (small corpora) | Approximate traversal, very low latency | Blend of cosine and spectral signal; runtime τ control |
| Long-tail behaviour | Depends on local similarity | Depends on graph connectivity in sparse regions | Depends on graph and spectral ranking; Engineering 001 reports flatter score tails on CVE |
| OOD handling | Requires external monitoring | Requires external monitoring | Can expose spectral diagnostics per query |
| Drift/monitoring support | None built in | None built in | Spectral artifacts support distribution monitoring |
| Main trade-off | Simple and fast | Fast approximate search at scale | Additional build and storage cost; smaller published track record |
| Evidence required | Dataset-specific | Extensive, dataset-specific | Dataset-specific; published vs cosine baselines |
How to read the comparison honestly
- Evidence status. The published evaluation (Engineering 001) compares spectral retrieval against cosine configurations on CVE and TREC-COVID under a stated protocol. No published head-to-head benchmark against HNSW is cited here; the HNSW column describes its documented behaviour as an ANN index, not a measured comparison.
- Preserved structure. No retrieval method preserves every property of an embedding space. Spectral retrieval preserves information captured by the Laplacian of the specific graph built by its wiring parameters, measurable as smoothness and dispersion; HNSW preserves proximity reachability. Which preservation matters is a per-dataset question.
- Complementarity. The approaches can coexist: HNSW or flat cosine for high-throughput top-k, with spectral artifacts for diagnostics, re-ranking, and monitoring. ArrowSpace does not implement HNSW traversal internally.
Choosing between them
- Need maximum query throughput on large corpora with well-documented ANN behaviour? Start with an established HNSW implementation.
- Need domain-tuned ranking, bounded comparable scores, or structural diagnostics (OOD, drift)? Add spectral retrieval.
- Need both? Keep the baseline for recall and use spectral signals for monitoring and re-ranking; measure before committing either to a full rebuild.