Skip to main content
Use jev.rerank when you want to evaluate Jev over hybrid candidates. It is optional; installing it does not change core.retrieve or make deep the default. Jev 1.0.0 serves only deep and requires core.retrieve’s default profile (version >=1.1.0 <2.0.0) and Plugin API 0.13. Run both as separate sidecars.

Install and pin

From the repository checkout, install both packages in a Python 3.12+ virtual environment:
Supply TYPESAFE_API_KEY to this process through your secret manager or environment before starting it. Do not put the key in the plugin configuration. Without it, deep works in fallback mode, without a paid call. Keep core.retrieve running and add Jev to your engine’s plugin pins. This illustrative configuration maps each short name to its provider:
Keep the ingestion plugin pinned. Search profile aliases are required when several retrieval plugins are installed. Follow the plugin activation guide when changing an existing installation.

Search and inspect

Send your ordinary search request with profile: deep. Jev asks Quivr for core.retrieve/default’s hybrid results, whatever mode the search sent, then scores them with the jev-1.13.0 model. Each re-ranked hit’s explanation shows Jev’s probability that the passage answers the query. usage.paid_calls and usage.cost_cents report the paid work, and usage.profiles lists both profiles with their own share. default follows the requested search mode and never calls Jev. re-ranker unavailable: <reason> means hybrid order was returned, without a Jev probability. Missing credentials, provider errors, invalid answers and the paid call’s 2-second deadline or the cost left all fall back this way. Search profiles explains how the two profiles share one time and cost budget. GET /v0/search/profiles lists core.retrieve/default with alias default and jev.rerank/deep with alias deep, alongside other installed profiles.

Settings

Trimming counts tokens with a local tokenizer file, tokenizer_path, checked against tokenizer_sha256; by default the digest is core.ingest’s pinned tokenizer. These are tokenizer tokens, not provider billing tokens. Scores from a failed response never enter the cache. Watch relevance, p95 latency, input tokens, cost and the fallback rate on your own workload before relying on deep. The plugin’s README describes the cache, cost reservation, provider limits and how it was measured.