intfloat/multilingual-e5-small. The engine
did this itself until THE-777; it now segments and embeds nothing, and the API
and worker refuse to start without an ingestion plugin. Every stack pins this
one unless another ingestion plugin is pinned (scripts/connector_plugin.py
FIRST_PARTY, Railway CONNECTORS in deploy/railway/core-entrypoint.py).
What it does
- Segments. Title and body text Parts only. Each body Part is cut into
windows of at most 384 tokens of the pinned tokenizer that prefer paragraph,
line and sentence ends, overlapping by 48 tokens. Recipe values are in
profile.json. A request with no space answers the segments alone (Plugin API 0.8), which is how a Version becomes searchable by keyword while the embedding service is down. - Embeds. One vector per window in space
core.ingest.e5-small@1(384 dimensions, cosine):passage: <title, at most 64 tokens>\n\n<window>, one request per window to the deployment’s TEI.embed_queryencodesquery: <query>and refuses a query over 256 tokens. - Refuses (
segmentation_limit, the Version is blocked asingestion_refused) text over 256 KiB, more than 64 Parts or 256 windows, a window over 4,096 code points or a model input over 512 tokens. - Provenance. Each segment records its token range, overlap, hard cuts, title and model-input token counts and the SHA-256 of its model input.
tokenizers 0.23.2 and the checksummed tokenizer.json,
prepared by scripts/prepare_tokenizer.py) runs as one helper process per
plugin process, loaded once (tokenizer.py).
Configuration
Parity and certification
testdata/golden.json holds what the engine produced before the move, for
testdata/parity-input.json: segments, offsets, derivations, refusals and the
SHA-256 of every float32 vector. parity_test.go holds the plugin to it bit
for bit. TEI’s CPU kernels round differently across processor families, so on
another processor than the capture’s the vectors are compared with the engine’s
former TEI request, sent to the same TEI in the same run. It needs TEI and the tokenizer, so the verify stack runs it and
certifies the plugin with quivr plugin test (scripts/core_ingest_plugin.py);
make check only vets and unit-tests it.
Moving an existing deployment
Corpora built before this plugin keep being served by the engine’s former E5 space: search works in every mode, and new Versions are searchable by keyword at once, but their vectors wait. Rebuild each Corpus (POST /v0/corpora/{corpus_id}/rebuilds); the rebuild re-embeds with the same
TEI and model, so results do not change, and the waiting vectors attach.