Embedding drift detection
What embedding drift is, why embedding-level statistics alone miss part of the picture, and how spectral signals support production monitoring.
Embedding drift is a change over time in the distribution of embeddings produced for the same input population, caused by model updates, upstream preprocessing changes, or shifts in the incoming data. It degrades retrieval quality and downstream decisions before it is visible in raw metrics.
Why drift matters for retrieval systems
Embedding search quality depends on the compatibility of three layers: the embedding model, the corpus, and the query distribution. When any layer moves, distances that used to separate relevant from irrelevant items stop doing so. Typical triggers include:
- Model updates. Re-embedding with a new model version changes the geometry of the whole space; indexes built on the old geometry silently decay.
- Corpus growth. New documents change manifold density and can invalidate calibration done on the original dataset.
- Upstream changes. Text preprocessing, chunking, or tokenisation changes alter what the same model produces for identical inputs.
- Domain shift. The user population or topics genuinely move away from the indexed content.
Signals: embedding-level and spectral
Common monitoring tracks embedding-level statistics: mean pooling distances, cluster assignments, or nearest-neighbour distance distributions between time windows. These catch coarse shifts but say little about how the structure of the corpus changes.
Spectral drift signals compare distributions of Laplacian-derived quantities between windows: the eigenvalue distribution of the corpus graph, and per-item bounded spectral scores (λτ, see graph Laplacian retrieval). Because these scores are bounded and comparable across model updates, they can flag structural change that raw embedding statistics miss. The capability is part of the spectral intelligence platform: one Laplacian artifact serving both search and monitoring.
A monitoring policy
- Compute spectral descriptors per window (eigenvalue distribution, λτ histograms) on a fixed sampling scheme.
- Compare each new window against a calibrated baseline using the same wiring parameters.
- Alert on sustained distributional change, not single-window noise.
- On alert: inspect which regions of the manifold moved, then decide between recalibration, partial reindex, or full rebuild.
Qualifications
Spectral drift signals are indicative, not conclusive: thresholds and sensitivity depend on sampling, graph wiring, and workload. Validate any alerting policy against labelled incidents on your own deployment. For handling individual queries that fall outside the corpus, see out-of-distribution vector retrieval.