* fix(embeddings): spill cached vectors to a Float32 temp file
Keep restore metadata in RAM and write embeddings once the in-memory
row limit is exceeded so incremental analyze can survive large caches
without a full-table number[] heap (#3306).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(lbug): stream CodeEmbedding cache under the connection lock
Spill vectors once the in-memory limit is crossed and fail the load
instead of adopting an empty snapshot, so incremental analyze cannot
OOM or quietly drop the restore cache (#3306).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(analyze): restore cached embeddings from a streamed spill snapshot
Hold row metadata across wipe, materialize 200-row batches, and treat
cache-load failures as warn-and-continue so incremental analyze can
preserve vectors without a full-table heap (#3306).
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3310)
Loop spill writes until the full vector lands, keep materialize failures out of the insert catch and the Phase 4 hash skip-set, and assert spilled restore subsets by node id instead of scan order.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3310)
Discard only this analyze run's embedding spills so a concurrent analyze on another index keeps its restore file, and isolate the default in-memory limit test from inherited env.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Address PR review feedback (#3310)
Mark a node stale when any restore batch fails so leftover chunks are deleted and rembedded, and exercise a full-length bad-magic spill header.
---------
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>