Prerequisites
Use Python 3.12 and a committed Quivr checkout. Installscripts/eval/requirements-campaign.txt. Your operator supplies Modal access,
measurement secrets, aggregate MLflow credentials and these environment variables:
EVAL_CONTROL_DATABASE_URL: the evaluation control PostgreSQL database.EVAL_STUDY_DATABASE_URL: PostgreSQL for Optuna’s separateeval_optunaschema.EVAL_CONTROL_CA_PEMwhen the remote database needs a private certificate authority.
sslmode=verify-full. A database owner applies
deploy/mlflow/eval-control.sql and deploy/mlflow/optuna.sql first. Keep secrets
outside campaign files. See the repository’s
campaign guide
for roles and runtime requirements.
Validate and start
Copyscripts/eval/examples/search-campaign.yaml to an ignored operator file. Set
an actual end date, deployment revision, opaque ticket reference and spending caps.
The baseline is Cohere-Embed-V5-Pro, 1024 dimensions, hybrid weight 0.5.
Run this validation without credentials or network:
start launches the supervisor. Run it and the watchdog as separate managed
processes on a machine with this frozen checkout, Modal access and both databases.
After registration, configure resume as the supervisor restart command. A running
supervisor resumes automatically on the next UTC day after a daily cap; resume
recovers a lost supervisor after the watchdog reconciles its compute. Terminal
stops require a new campaign. Restart the watchdog automatically too.
Daily exhaustion cancels compute and pauses until the next UTC day. Total caps,
the end date and the trial limit stop the campaign. Conservative reservations
include uncertain charges; Modal builds/storage and other account charges need
separate budgeting. No measurements run in CI.
Check and stop
status returns aggregate results, Pareto objectives, spending and held-out reads
remaining. It exports no query IDs or raw records. agent_token_usage: null means
unknown when no exact receipts have been ingested; usage is never estimated.
Exploration does not consume held-out reads. Configure EVAL_LINEAR_TOKEN,
EVAL_SLACK_BOT_TOKEN and EVAL_SLACK_CHANNEL outside the repository for daily
summaries. The supervisor freezes one aggregate snapshot per UTC day and retries
unacknowledged destinations. Slack timeouts can cause duplicate messages.
See the reporting guide
for proposals, receipt format and delivery recovery.
Example requiring Modal and database credentials; not run during development:
cleanup_pending: false before treating cleanup as complete. If termination cannot
be acknowledged, the command returns a nonzero exit code; retry stop or let the
managed watchdog retry. Uncertain charges and completed evidence are retained.
SIGTERM and Ctrl-C stop the campaign; hard kills are recovered by the watchdog.
Next
Confirm finalists on the full stack and frozen holdout before changing production. The trusted confirmation runner is unavailable; the engine smoke command cannot qualify a candidate. Campaigns reportconfirmation_available: false and open no
promotion PRs until that integration lands. The promotion adapter maps hosted
model/dimensions and retrieval weight, depth and fusion only after confirmation.
Record and compare measurements.