Skip to main content
The reference Quivr normalizer. It turns an application/pdf Blob into a Manifest with one text Part per page, so a PDF is found by the text of its pages. It is built with the Python Plugin SDK, and it is also the worked example of Write a normalizer.

Output

  • The Version extension pdf-text.document (schema version 1) holds page_count and text_pages, the number of pages with text. Retrieval mappings may point at /extensions/pdf-text.document/data/page_count.
  • Quivr indexes text Parts with the title and body roles, so every page is searchable, and a search hit names its page through part_key.
  • The Record Version also keeps the input PDF in provenance.source_blob_ids, and provenance.normalization names pdf-text and its version.
  • Text is cleaned so Quivr can index it: NUL characters and invalid UTF-8 are removed, trailing spaces are trimmed, and runs of blank lines are collapsed to one.
  • There is no OCR: a scanned page has no extractable text.

Warnings and errors

A warning keeps the output publishable. An error quarantines the Record Version: none of the plugin’s output is published, and the Version read lists a normalizer_failed diagnostic that names pdf-text and carries the code below. PDFs protected only by an owner password (an empty user password) are read normally, whether they use RC4 or AES encryption. Warnings are bounded: at most one per kind, and each message is at most 1024 characters.

Configuration

Set in the plugin pin’s configuration (see the guide). Every field is optional.

Limits

Quivr’s built-in indexing reads, per Record Version, at most 64 title/body text Parts, 256 KiB of text, and 256 segments of about 384 tokens. A Version beyond any of these is published, but it is not searchable (segmentation_limit). The plugin therefore merges pages and truncates text instead of failing:
  • max_page_parts covers the Part limit.
  • The default max_text_bytes of 128 KiB keeps ordinary prose well under the segment limit.
Text that tokenizes densely, such as long tables of numbers or some scripts, can still exceed 256 segments. For such PDFs, lower max_text_bytes. Other bounds that apply to a PDF:
  • Upload Sessions accept up to 1 GiB.
  • The plugin reads the whole PDF into memory.
  • One invocation must finish within the declared timeout_ms (60 s), which the engine caps at 2 minutes.
  • The response is bounded by max_response_bytes (4 MiB by default), and the stored Manifest by 2 MiB.
A PDF that keeps exceeding timeout_ms is invoked up to retry.max_attempts (3) times in total, and each attempt parses it again. It is then quarantined with normalizer_timeout.

Develop

Python 3.12 or later, from this directory:
fixtures/sample.pdf is generated by python3 scripts/make_fixtures.py, and the unit tests check that the committed file matches the script. The script writes the PDF objects directly, so it needs no PDF library and gives the same bytes on every run. It has three pages: text on pages 1 and 2, none on page 3. In the Quivr repository, make test runs these checks, and CI publishes the Contract Runner report as the pdf-text-contract-report artifact. make dev pins this plugin for application/pdf by default: see the local harness.

Dependencies and licences

The SDK brings PyYAML (MIT) and jsonschema (MIT). pypdf is pinned so that every installation extracts the same text: a new pypdf version is a new plugin version. This plugin is MIT-licensed, like the rest of the repository.