Prerequisites
- A ready Quivr deployment, its ingestion and normalization plugins, and operator access to worker, database and model-server settings.
- A collection of documents (Corpus) and a connector supported by your source. Import archives covers immutable object-storage archives and source credentials.
- A key with
connectors:writeandconnectors:readon the Corpus. Rebuilds needprojections:rebuild; reading their Operations needsoperations:read. - A separate installation-wide queue key with only
queues:readandcorpora: ["*"]. Keep this key out of application clients. - For hosted inference, the provider’s credentials in the plugin environment, its concurrency and rate allowances, and permission to send document text there.
Steps
1. Finish model changes before importing
Choose an embedding model and pin its exact configuration first. Replacing a servedhosted.embed configuration with a different vector-space identity requires a Corpus rebuild. Until the replacement activates, new documents can become searchable by keywords while their vector enrichment reports rebuild_required; their vectors serve after the rebuild and subsequent processing complete.
For a plugin that adds a new evaluation space with compatible stored cuts, use backfill and space promotion instead. Complete the applicable model-switch procedure, wait for rebuild activation or backfill and promotion, and verify vector coverage and semantic search. Plan this before a large import: model preparation and importing together consume the same provider, storage and indexing capacity. Keep previous registrations reachable while their pinned work drains.
2. Measure a representative sample
Import a bounded sample with the same document lengths, revisions, media types and model as the full collection. Record accepted and searchable documents per second, passages per document, provider latency and peak memory. Acceptance means the source was durably submitted; it does not mean it is searchable. Measure storage before and after the sample with the read-only storage helper. It needs Python 3,psql, a protected Quivr configuration and database read access. For example, not run; use your deployment configuration and an existing private report directory:
public_table_bytes from both reports for incremental table storage, or database_bytes from both for the whole database. Estimate bytes per new document as (after_bytes - before_bytes) / (after_documents - before_documents), using the report’s documents counts on an otherwise idle sample. For imports adding revisions, use the change in versions as the denominator instead. A zero denominator cannot produce this estimate. Physical allocation includes dead tuples, retained audit detail and empty tables; an aggregate average is not a fixed document cost.
Measure object-storage and Weaviate disk growth separately using your platform’s bucket-size and volume-usage metrics over the same sample interval. Count physical retained objects, including versions if bucket versioning is enabled. The helper’s referenced-file counts do not measure all physical object bytes or search-index storage. Allow room for source blobs, vectors, retained generations, WAL, temporary rebuild data and backups. Keep the reports private and compare matched sample sizes when changing settings.
3. Budget PostgreSQL and Weaviate
Setpostgres.max_connections for each engine process. The server’s usable connection limit must cover the sum of pools across all replicas, other clients and rolling-deployment headroom. One API, one live worker and eight bulk workers with pools of 16 can use 160 connections. Activity slots do not set the pool size.
Use the shipped PostgreSQL connection, memory and WAL settings, shared by Compose and Railway. Set a container memory limit or an explicit memory budget; leave room for concurrent sort/hash operations, connections, autovacuum and the OS cache. Monitor each process’s pool saturation and database I/O before adding workers.
Watch changes in pg_stat_checkpointer.num_requested and num_timed over the same interval. Requested checkpoints far above timed ones, together with WAL pressure, can indicate a checkpoint storm. PostgreSQL’s stock 1 GB max_wal_size can be too small for a sustained import: repeated checkpoints generate more full-page-image WAL. Raise max_wal_size, lengthen checkpoint_timeout and enable wal_compression rather than increasing worker load into that bottleneck.
The shared startup settings budget maximum WAL at 10% of the data filesystem, capped at 32 GiB, with checkpoint_timeout=15min and wal_compression=lz4. Set QUIVR_POSTGRES_VOLUME_MB to the allocated quota in integer MiB if the mount reports host capacity. The linked guide lists size overrides and floors. max_wal_size is a soft limit; leave disk headroom. Larger WAL budgets can lengthen crash recovery; compression costs CPU. Keep synchronous_commit, fsync and full_page_writes on. See PostgreSQL WAL configuration.
Restart PostgreSQL after changing the template’s startup budgets, then verify the effective settings. Check WAL growth, checkpoint rates and write latency during the sample before increasing import load.
For Weaviate, size by passage count, vector dimensions and the selected index compression, then measure RSS and disk under import and search load. Quivr defaults to rq-8; rq-1 uses less vector memory and none uses full precision. Compressed indexes retain full-precision vectors on disk for rescoring. Evaluate search quality before changing compression; index configuration applies to existing Corpora only after a rebuild.
Set a container memory budget and leave runtime and OS headroom. For example, GOMEMLIMIT=3GiB on a 4 GiB service is a starting soft limit; it does not bound RSS or make an oversized index fit. Follow Weaviate resource planning and measure your sample. Give index loading enough startup time and watch disk free space, indexing throughput and search latency.
4. Put collection on the bulk queue
When creating the Connector Instance, setwork_queue: "bulk" at the request’s top level, alongside kind, config and schedule. It is an engine field, outside the connector’s config. The queue follows each run and accepted document through processing. Rebuilds, backfills and quarantine reprocessing also use bulk; direct submissions and pushes use live.
Run separate live and bulk worker processes against the same dependencies and plugin plan. For example, these fragments select one class per process:
Live worker
Bulk worker
worker.slots.bulk × rebuild.concurrency Versions when every slot runs a rebuild. Backfills and other bulk work share those slots.
Separate queues reserve worker capacity. Live and bulk still share databases, provider requests and search indexing. See worker queues and rebuild settings for defaults and bounds.
5. Scale the embedding provider before workers
For a plugin sidecar per bulk worker, size provider admission for at leastbulk replicas × max_concurrent_requests, plus live workers and query traffic. Eight bulk replicas at four provider requests per process can offer 32 concurrent requests before that additional load. If workers share one plugin service, use the number of plugin processes instead: its request cap is process-local.
Above a server’s capacity, it can answer 503. Documents then repeat their baseline processing retries and rebuilds stall, even with idle worker slots. Scale the provider first, or lower plugin concurrency and bulk replicas. Check completed throughput before raising workers again. Concurrency does not enforce a provider’s requests-per-minute or tokens-per-minute allowance.
hosted.embed batches document inputs within each Organization and plugin process. batch_size limits inputs per request; max_batch_tokens limits the summed token-cost estimate, which can differ from the provider’s actual or billed count. Queries bypass the collection window but share remote provider admission. A 429 or 503 shares Retry-After cooldown across that plugin process; other replicas still have their own limits and cooldowns. Follow the execution-only tuning procedure when changing these controls.
Lower rebuild.concurrency if provider, PostgreSQL or Weaviate pressure increases. Add replicas only after the sample shows spare downstream capacity. Scale workers on backlog covers KEDA and quivr-autoscaler on Kubernetes or Railway. Set its maximum replicas within your provider and connection budgets; retain at least one bulk worker to finish in-flight work.
6. Watch backlog and pause acquisition when needed
ReadGET /v0/admin/queues with the dedicated queue key. For each of queues.bulk and queues.live, watch:
The bulk autoscaler counts rebuild work in its waiting signal, alongside import, backfill and quarantine work. Rebuild/backfill counts are estimates and overlapping operation work can count more than once. A larger backlog can therefore be a rebuild, not additional source documents.
Observations refresh every 15 seconds by default. Missing, stale or unavailable observations return
503; treat that as unknown capacity, not zero work. Replica metrics expose the same installation-wide snapshot: aggregate with max, not sum. See queue observation rules.
Compare queue progress with connector health, searchable-document counts, provider errors and database/index load. Pause the connector to stop further acquisition when downstream work cannot keep up. Pause retains its last saved checkpoint; already accepted work keeps processing. A source request already running can finish but cannot advance the checkpoint after pause. Resume makes collection due immediately unless the source’s Retry-After delays it. A longer schedule does not stop a continuing import.