Skip to main content
Run a campaign to compare public search settings over several days. It keeps the best quality, cost and latency trade-offs while enforcing daily and total spending caps. The lead submits bounded proposals and exact usage receipts; daily summaries go to Linear and Slack. Promotion requires trusted full-engine confirmation.

Prerequisites

Use Python 3.12 and a committed Quivr checkout. Install scripts/eval/requirements-campaign.txt. Your operator supplies Modal access, measurement secrets, aggregate MLflow credentials and these environment variables:
  • EVAL_CONTROL_DATABASE_URL: the evaluation control PostgreSQL database.
  • EVAL_STUDY_DATABASE_URL: PostgreSQL for Optuna’s separate eval_optuna schema.
  • EVAL_CONTROL_CA_PEM when the remote database needs a private certificate authority.
Remote database URLs require sslmode=verify-full. A database owner applies deploy/mlflow/eval-control.sql and deploy/mlflow/optuna.sql first. Keep secrets outside campaign files. See the repository’s campaign guide for roles and runtime requirements.

Validate and start

Copy scripts/eval/examples/search-campaign.yaml to an ignored operator file. Set an actual end date, deployment revision, opaque ticket reference and spending caps. The baseline is Cohere-Embed-V5-Pro, 1024 dimensions, hybrid weight 0.5. Run this validation without credentials or network:
It prints the normalized goal, test sets, search ranges and budgets. Exploration accepts public development splits only. It maximizes weighted nDCG@10 and minimizes worst-set serving cost and p95 latency. A Pareto front lists candidates whose scores cannot all be improved by another candidate. Gate failures remain visible but cannot qualify for promotion. Examples requiring operator credentials and rented compute; not run during development:
start launches the supervisor. Run it and the watchdog as separate managed processes on a machine with this frozen checkout, Modal access and both databases. After registration, configure resume as the supervisor restart command. A running supervisor resumes automatically on the next UTC day after a daily cap; resume recovers a lost supervisor after the watchdog reconciles its compute. Terminal stops require a new campaign. Restart the watchdog automatically too. Daily exhaustion cancels compute and pauses until the next UTC day. Total caps, the end date and the trial limit stop the campaign. Conservative reservations include uncertain charges; Modal builds/storage and other account charges need separate budgeting. No measurements run in CI.

Check and stop

status returns aggregate results, Pareto objectives, spending and held-out reads remaining. It exports no query IDs or raw records. agent_token_usage: null means unknown when no exact receipts have been ingested; usage is never estimated. Exploration does not consume held-out reads. Configure EVAL_LINEAR_TOKEN, EVAL_SLACK_BOT_TOKEN and EVAL_SLACK_CHANNEL outside the repository for daily summaries. The supervisor freezes one aggregate snapshot per UTC day and retries unacknowledged destinations. Slack timeouts can cause duplicate messages. See the reporting guide for proposals, receipt format and delivery recovery. Example requiring Modal and database credentials; not run during development:
Stop closes paid admission and terminates only this campaign’s compute. Check cleanup_pending: false before treating cleanup as complete. If termination cannot be acknowledged, the command returns a nonzero exit code; retry stop or let the managed watchdog retry. Uncertain charges and completed evidence are retained. SIGTERM and Ctrl-C stop the campaign; hard kills are recovered by the watchdog.

Next

Confirm finalists on the full stack and frozen holdout before changing production. The trusted confirmation runner is unavailable; the engine smoke command cannot qualify a candidate. Campaigns report confirmation_available: false and open no promotion PRs until that integration lands. The promotion adapter maps hosted model/dimensions and retrieval weight, depth and fusion only after confirmation. Record and compare measurements.