jev.rerank when you want to evaluate Jev over hybrid candidates. It is
optional; installing it does not change core.retrieve or make deep the
default. Jev 1.0.0 serves only deep and requires core.retrieve’s
default profile (version >=1.1.0 <2.0.0) and Plugin API 0.13. Run both as separate sidecars.
Install and pin
From the repository checkout, install both packages in a Python 3.12+ virtual environment:TYPESAFE_API_KEY to this process through your secret manager or
environment before starting it. Do not put the key in the plugin configuration.
Without it, deep works in fallback mode, without a paid call.
Keep core.retrieve running and add Jev to your engine’s plugin pins.
This illustrative configuration maps each short name to its provider:
Search and inspect
Send your ordinary search request withprofile: deep.
Jev asks Quivr for core.retrieve/default’s hybrid results, whatever mode the
search sent, then scores them with the jev-1.13.0 model. Each re-ranked hit’s
explanation shows Jev’s probability that the passage answers the query.
usage.paid_calls and usage.cost_cents report the paid work, and
usage.profiles lists both profiles with their own share. default follows
the requested search mode and never calls Jev.
re-ranker unavailable: <reason> means hybrid order was returned, without a
Jev probability. Missing credentials, provider errors, invalid answers and
the paid call’s 2-second deadline or the cost left all fall back this way.
Search profiles explains how the two profiles
share one time and cost budget.
GET /v0/search/profiles lists core.retrieve/default with alias default
and jev.rerank/deep with alias deep, alongside other installed profiles.
Settings
Trimming counts tokens with a local tokenizer file,
tokenizer_path, checked
against tokenizer_sha256; by default the digest is core.ingest’s pinned
tokenizer. These are tokenizer tokens, not provider billing tokens. Scores
from a failed response never enter the cache.
Watch relevance, p95 latency, input tokens, cost and the fallback rate on
your own workload before relying on deep. The plugin’s
README
describes the cache, cost reservation, provider limits and how it was measured.