Skip to main content
The first-party ingestion plugin: it cuts every Record Version into windows of tokens and embeds each window with intfloat/multilingual-e5-small. The engine did this itself until THE-777; it now segments and embeds nothing, and the API and worker refuse to start without an ingestion plugin. Every stack pins this one unless another ingestion plugin is pinned (scripts/connector_plugin.py FIRST_PARTY, Railway CONNECTORS in deploy/railway/core-entrypoint.py).

What it does

  • Segments. Title and body text Parts only. Each body Part is cut into windows of at most 384 tokens of the pinned tokenizer that prefer paragraph, line and sentence ends, overlapping by 48 tokens. Recipe values are in profile.json. A request with no space answers the segments alone (Plugin API 0.8), which is how a Version becomes searchable by keyword while the embedding service is down.
  • Embeds. One vector per window in space core.ingest.e5-small@1 (384 dimensions, cosine): passage: <title, at most 64 tokens>\n\n<window>, one request per window to the deployment’s TEI. embed_query encodes query: <query> and refuses a query over 256 tokens.
  • Refuses (segmentation_limit, the Version is blocked as ingestion_refused) text over 256 KiB, more than 64 Parts or 256 windows, a window over 4,096 code points or a model input over 512 tokens.
  • Provenance. Each segment records its token range, overlap, hard cuts, title and model-input token counts and the SHA-256 of its model input.
The pinned tokenizer (tokenizers 0.23.2 and the checksummed tokenizer.json, prepared by scripts/prepare_tokenizer.py) runs as one helper process per plugin process, loaded once (tokenizer.py).

Configuration

Parity and certification

testdata/golden.json holds what the engine produced before the move, for testdata/parity-input.json: segments, offsets, derivations, refusals and the SHA-256 of every float32 vector. parity_test.go holds the plugin to it bit for bit. TEI’s CPU kernels round differently across processor families, so on another processor than the capture’s the vectors are compared with the engine’s former TEI request, sent to the same TEI in the same run. It needs TEI and the tokenizer, so the verify stack runs it and certifies the plugin with quivr plugin test (scripts/core_ingest_plugin.py); make check only vets and unit-tests it.

Moving an existing deployment

Corpora built before this plugin keep being served by the engine’s former E5 space: search works in every mode, and new Versions are searchable by keyword at once, but their vectors wait. Rebuild each Corpus (POST /v0/corpora/{corpus_id}/rebuilds); the rebuild re-embeds with the same TEI and model, so results do not change, and the waiting vectors attach.