Skip to main content
You run the new version next to the old one, switch to it in one call, watch the old one drain, and roll back in one call if needed. If the new version brings an embedding model, you then fill its vector space for past articles. No step restarts the api or the worker. Each step links to the page that details it.

Prerequisites

  • An operator key in QUIVR_OPERATOR_KEY, with the actions plugins:admin, observability:read, operations:read and operations:write on every Corpus of its Organization. Registering, activating and rolling back apply to the whole deployment; the plugin stats, the quarantine list and backfills cover only the key’s Organization, so watch each Organization with its own key.
  • The new version certified by quivr plugin test, deployed at its own address and reachable from the api and the worker. The old version keeps running.
  • For a new embedding model: the new version’s manifest declares the new vector space, and it cuts articles into the same segments as the old version.

Steps

1

Note the plan in service

This is the plan a rollback returns to.
2

Register the new version and wait for its check

Follow the “Register it” and “Wait for the check” steps of Switch plugins without restarting. Keep the vector space in service as served. Register a new embedding model’s space as evaluation, so search does not use it before it is filled:
Go on once the registration is validated, and keep its id:
3

Activate it

Every api and worker process follows within plugin_plan_poll (2 seconds by default). From then on, new work and every search use the new version.
4

Watch the old version drain

Work that started before the switch finishes on the old version. The old version reads draining until that work is done, then inactive:
While it drains, compare the two versions’ calls and errors, and list the articles quarantined since the switch, which should stay empty:
If the new version’s errors climb or articles are quarantined, roll back.
5

Stop the old version

Stop its process once it reads inactive. Stopping it earlier makes its pinned work retry, then stop with pinned_plugin_unavailable; the work never moves to the new version (see Work finishes on the version it started with).

Roll back

One call returns to the plan you noted in the first step. Choose what happens to the work already pinned to the new version: The request, its answer and its errors are in Roll back. After a stop, reprocess the quarantined articles through the plan now active.

Fill a new vector space

When the new version brings an embedding model, a Corpus created after the switch carries its evaluation space from the start. An existing Corpus gets vectors in it only once a backfill of that Corpus starts: from then on live ingestion fills it for new articles, and the backfill fills the ones before. So fill it for each existing Corpus without accepted_before: run the dry run, start the backfill, follow it to succeeded, then promote the space so search uses it. Promoting the former space goes back to the previous model at once.

Check it worked

  • GET /v0/admin/plugins/plan names the new version for its roles, and the old version reads inactive.
  • No article was quarantined since the switch.
  • A semantic search’s hits name the vector space you promoted, in vector_space_id.

Restarts during an upgrade

Every step survives a restart of the api or the worker:
  • an activation or a rollback is kept in the database, and the configuration applies at startup only what changed in it;
  • pinned work keeps its plan across worker restarts;
  • a backfill resumes after its checkpoint.
A single api process does not answer while it restarts. Run two or more behind a load balancer if you restart the api during an upgrade.

Troubleshooting