Prerequisites
Use Python 3.9+ in a Quivr repository checkout. Your operator supplies an HTTPS tracking URL and HTTP Basic credentials throughMLFLOW_TRACKING_URI,
MLFLOW_TRACKING_USERNAME and MLFLOW_TRACKING_PASSWORD. Keep them in a secret
manager. Paired comparisons need scipy from scripts/eval/requirements.txt.
Measurements need the code commit, plugin digest, full configuration, dataset
name/version/split/fingerprint, tier (direct or engine), machine, duration,
provider cost, aggregate metrics and optional per-query scores. Use null for
unmeasured values. Configuration contains no credentials; query scores contain
opaque ids and numbers, never query or document text.
The complete JSON fields and Python helper are in the repository’s
results guide.
Log and compare
For your measurement files, examples requiring your tracking credentials:result_key, status (synced or pending)
and, when uploaded, the MLflow run_id. Use --baseline before the candidate’s
first log to store paired statistics. Listing returns aggregates and lineage;
query arrays are artifacts. Comparisons pair identical query ids on the same
dataset fingerprint, split and tier. The paired Student t-test is two-sided, alpha 0.05, with no
multiple-comparison correction.
For a completed make eval report, use log-engine report.json --experiment public/engine --lineage lineage.json --public-set miracl-fr. Repeat --public-set
for each known public set; all others default to private. The optional lineage
JSON supplies machine, plugin_digest and full config settings absent from
older reports. Direct bakeoff files use log-direct <file> --experiment <name>.
Private rows follow the separate credential rules.
Both --compare-to sides import; lineage.baseline supplies the baseline’s own settings.
Check the leaderboard
These commands were exercised locally with the imported public evidence:Recover pending uploads
Once credentials and the server are available, replay your outbox:.scratch/eval/results; set --directory before the command to
choose another public directory. Preserve it when retiring a worker. Listing
uses cached aggregates when offline and marks them source=local. Authorization
and input errors fail the command. Repeated logging retains the first result
with the same experiment/config/dataset/code/plugin/tier key; incomplete uploads
resume. Simultaneous writers can race on MLflow’s search-before-create, and MLflow
records remain mutable.
Restrict private query scores
Setdataset.private=true. Aggregates stay in the ordinary experiment. Query
arrays require a pre-provisioned private/<experiment> and separate credentials:
MLFLOW_PRIVATE_TRACKING_URI, MLFLOW_PRIVATE_TRACKING_USERNAME,
MLFLOW_PRIVATE_TRACKING_PASSWORD. Its owner grants ordinary agents access only
to aggregates. The wrapper never creates private experiments.
Private offline arrays stay outside the repository, by default in
~/.local/share/quivr/eval-private. Keep that location on a private measurement
machine; agents sharing its OS account can read its files. Without private
credentials, sync refuses private arrays and compare shows aggregate deltas only.