> ## Documentation Index
> Fetch the complete documentation index at: https://docs.quivr.thevibecompany.co/llms.txt
> Use this file to discover all available pages before exploring further.

# pdf-text

> The reference PDF normalizer, one Part per page

The reference Quivr normalizer. It turns an `application/pdf` Blob into a
Manifest with one text Part per page, so a PDF is found by the text of its
pages. It is built with the [Python Plugin SDK](/sdks/python),
and it is also the worked example of [Write a normalizer](/plugins/write-a-normalizer).

<h2 id="output">
  Output
</h2>

| Part key | Role | Content |
| - | - | - |
| `page-<n>` | `body` | The extracted text of page *n* (1-based), for every page that has text |
| `source` | `source` | A Blob Part referencing the input PDF itself; omitted when `include_source` is `false` |

* The Version extension `pdf-text.document` (schema version `1`) holds
  `page_count` and `text_pages`, the number of pages with text. Retrieval
  mappings may point at `/extensions/pdf-text.document/data/page_count`.
* Quivr indexes text Parts with the `title` and `body` roles, so every page is
  searchable, and a search hit names its page through `part_key`.
* The Record Version also keeps the input PDF in `provenance.source_blob_ids`,
  and `provenance.normalization` names `pdf-text` and its version.
* Text is cleaned so Quivr can index it: NUL characters and invalid UTF-8 are
  removed, trailing spaces are trimmed, and runs of blank lines are collapsed to one.
* There is no OCR: a scanned page has no extractable text.

<h2 id="warnings-and-errors">
  Warnings and errors
</h2>

A warning keeps the output publishable. An error quarantines the Record Version:
none of the plugin's output is published, and the Version read lists a
`normalizer_failed` diagnostic that names `pdf-text` and carries the code below.

| Case | Result |
| - | - |
| Pages without text (blank or scanned) | Skipped. One `empty_pages` warning lists them, e.g. `3, 5-9` |
| A page that cannot be read | Skipped. One `unreadable_pages` warning lists them |
| More pages with text than `max_page_parts` | Later pages are merged into the last Part. Warning `pages_merged` |
| More text than `max_text_bytes` | Later text is dropped. Warning `text_truncated` |
| No text at all | Only the `source` Part is kept, with a `no_text` warning: the Record exists but is not searchable. With `include_source: false` this is a terminal `no_text` error instead |
| Encrypted PDF that an empty password does not open | Terminal error `encrypted_pdf` |
| Damaged or non-PDF bytes | Terminal error `corrupt_pdf` |

PDFs protected only by an owner password (an empty user password) are read
normally, whether they use RC4 or AES encryption. Warnings are bounded: at most one per kind, and each message is at
most 1024 characters.

<h2 id="configuration">
  Configuration
</h2>

Set in the plugin pin's `configuration` (see the guide). Every field is optional.

| Field | Default | Range | Meaning |
| - | - | - | - |
| `max_page_parts` | 64 | 1–64 | Page Parts kept; later pages join the last one |
| `max_text_bytes` | 131072 | 1024–262144 | UTF-8 bytes of page text kept in total |
| `include_source` | `true` | | Add the `source` Blob Part |

<h2 id="limits">
  Limits
</h2>

Quivr's core.ingest plugin indexes, per Record Version, at most 64 `title`/`body`
text Parts, 256 KiB of text, and 256 segments of about 384 tokens. A Version
beyond any of these is published, but it is not searchable (`ingestion_refused`).
The plugin therefore merges pages and truncates text instead of failing:

* `max_page_parts` covers the Part limit.
* The default `max_text_bytes` of 128 KiB keeps ordinary prose well under the
  segment limit.

Text that tokenizes densely, such as long tables of numbers or some scripts, can
still exceed 256 segments. For such PDFs, lower `max_text_bytes`.

Other bounds that apply to a PDF:

* Upload Sessions accept up to 1 GiB.
* The plugin reads the whole PDF into memory.
* One invocation must finish within the declared `timeout_ms` (60 s), which the
  engine caps at 2 minutes.
* The response is bounded by `max_response_bytes` (4 MiB by default), and the
  stored Manifest by 2 MiB.

A PDF that keeps exceeding `timeout_ms` is invoked up to `retry.max_attempts`
(3) times in total, and each attempt parses it again. It is then quarantined with
`normalizer_timeout`.

<h2 id="develop">
  Develop
</h2>

Python 3.12 or later, from this directory:

```bash theme={null}
python3 -m venv .venv && . .venv/bin/activate
pip install -e ../../sdks/python -e .
python3 -m unittest discover -s tests                 # unit tests
quivr plugin inspect .                                # validate quivr-plugin.yaml
quivr plugin dev --fixture fixtures/sample.json       # run once on the sample PDF
quivr plugin test --report contract-report.json       # certify with the Contract Runner
```

`fixtures/sample.pdf` is generated by `python3 scripts/make_fixtures.py`, and
the unit tests check that the committed file matches the script. The script
writes the PDF objects directly, so it needs no PDF library and gives the same
bytes on every run. It has three pages: text on pages 1 and 2, none on page 3.

In the Quivr repository, `make test` runs these checks, and CI publishes the
Contract Runner report as the `pdf-text-contract-report` artifact. `make dev`
pins this plugin for `application/pdf` by default: see the
[local harness](/quivr-v2-local-harness#plugin-substitution-and-handoff).

<h2 id="dependencies-and-licences">
  Dependencies and licences
</h2>

| Dependency | Version | Licence | Why |
| - | - | - | - |
| [pypdf](https://github.com/py-pdf/pypdf) | 6.19.0 (pinned) | BSD-3-Clause | PDF parsing and text extraction; pure Python |
| [cryptography](https://github.com/pyca/cryptography) | 50.0.1 (pinned) | Apache-2.0 OR BSD-3-Clause | Lets pypdf open AES-encrypted PDFs; brings cffi (MIT-0) and pycparser (BSD-3-Clause) |
| Quivr Plugin SDK (`sdks/python`) | 0.1 | MIT | Protocol models, HTTP server, errors |

The SDK brings PyYAML (MIT) and jsonschema (MIT). pypdf is pinned so that every
installation extracts the same text: a new pypdf version is a new plugin version.
This plugin is MIT-licensed, like the rest of the repository.
