> ## Documentation Index
> Fetch the complete documentation index at: https://docs.quivr.thevibecompany.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Re-rank deep searches with Jev

> Add Jev, an optional paid re-ranker, to deep searches, and fall back to the hybrid order when Jev is unavailable.

Use `jev.rerank` when you want to evaluate Jev over hybrid candidates. It is
optional; installing it does not change `core.retrieve` or make `deep` the
default. Jev 1.0.0 serves only `deep` and requires `core.retrieve`'s
`default` profile (version `>=1.1.0 <2.0.0`) and Plugin API 0.13. Run both as separate sidecars.

## Install and pin

From the repository checkout, install both packages in a Python 3.12+ virtual
environment:

```bash theme={null}
python3 -m pip install ./sdks/python ./plugins/jev-rerank
export QUIVR_PLUGIN_MANIFEST="$PWD/plugins/jev-rerank/quivr-plugin.yaml"
export QUIVR_PLUGIN_HOST=127.0.0.1
export QUIVR_PLUGIN_PORT=8083
python3 -m jev_rerank
```

Supply `TYPESAFE_API_KEY` to this process through your secret manager or
environment before starting it. Do not put the key in the plugin configuration.
Without it, `deep` works in fallback mode, without a paid call.

Keep `core.retrieve` running and add Jev to your engine's [plugin pins](/run-quivr/pin).
This illustrative configuration maps each short name to its provider:

```yaml theme={null}
plugins:
  - manifest: /path/to/quivr/plugins/core-retrieve/quivr-plugin.yaml
    endpoint: http://127.0.0.1:8084
  - manifest: /path/to/quivr/plugins/jev-rerank/quivr-plugin.yaml
    endpoint: http://127.0.0.1:8083
    configuration:
      candidate_count: 30
      trim_tokens: full
      ranking: noul
      cache_entries: 4096
retrieval:
  profiles:
    default: core.retrieve/default
    deep: jev.rerank/deep
```

Keep the ingestion plugin pinned. [Search profile aliases](/reference/configuration#search-profiles)
are required when several retrieval plugins are installed. Follow the [plugin activation guide](/run-quivr/upgrade-a-plugin)
when changing an existing installation.

## Search and inspect

Send your ordinary [search request](/guides/search) with `profile: deep`.
Jev asks Quivr for `core.retrieve/default`'s hybrid results, whatever mode the
search sent, then scores them with the `jev-1.13.0` model. Each re-ranked hit's
explanation shows Jev's probability that the passage answers the query.
`usage.paid_calls` and `usage.cost_cents` report the paid work, and
`usage.profiles` lists both profiles with their own share. `default` follows
the requested search mode and never calls Jev.

`re-ranker unavailable: <reason>` means hybrid order was returned, without a
Jev probability. Missing credentials, provider errors, invalid answers and
the paid call's 2-second deadline or the cost left all fall back this way.
[Search profiles](/run-quivr/search-profiles) explains how the two profiles
share one time and cost budget.

`GET /v0/search/profiles` lists `core.retrieve/default` with alias `default`
and `jev.rerank/deep` with alias `deep`, alongside other installed profiles.

## Settings

| Setting | Effect |
| - | - |
| `candidate_count` | How many hybrid results Jev scores: 20, 30 (default) or 50 |
| `trim_tokens` | Cuts each passage before scoring: `"128"`, `"256"` or `full` (default, no cut), as strings |
| `ranking` | `noul` (default) sorts by Jev's probability; `rrf` fuses it with the hybrid rank |
| `cache_entries` | How many scores to keep for repeated queries, up to 16384; default 4096, `0` turns it off |

Trimming counts tokens with a local tokenizer file, `tokenizer_path`, checked
against `tokenizer_sha256`; by default the digest is `core.ingest`'s pinned
tokenizer. These are tokenizer tokens, not provider billing tokens. Scores
from a failed response never enter the cache.

Watch relevance, p95 latency, input tokens, cost and the fallback rate on
your own workload before relying on `deep`. The plugin's
[README](https://github.com/The-Vibe-Company/quivr/tree/main/plugins/jev-rerank)
describes the cache, cost reservation, provider limits and how it was measured.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.