> ## Documentation Index
> Fetch the complete documentation index at: https://docs.quivr.thevibecompany.co/llms.txt
> Use this file to discover all available pages before exploring further.

# alerts

> The first-party keyword alert plugin

Quivr's first-party alert rules, a `subscription` plugin built with the
[Python Plugin SDK](/sdks/python). A Subscription pinned to
`{"plugin_id": "alerts", "version": "0.2.0"}` gets a Match when a new article
satisfies its Saved Query's expression. It offers two alert kinds:

* **`keywords`**: a boolean keyword query over the article's title and body, with
  optional filters on its metadata;
* **`described`**: a plain-language description that a classifier, TypeSafe's Jev,
  judges each article against ([below](#described-alerts)). The article's text is
  sent to TypeSafe, and the kind is off without `TYPESAFE_API_KEY`.

People writing alerts should start with the guides:
[keyword alerts](/keyword-alerts) and
[described alerts](/described-alerts). This README is the reference.

<h2 id="expression">
  Expression
</h2>

A keyword alert's Saved Query expression:

```json theme={null}
{
  "kind": "keywords",
  "match": {"all": [
    {"term": "Airbus"},
    {"any": [{"term": "grève"}, {"term": "strike"}]},
    {"not": {"term": "sport"}}
  ]}
}
```

`match` is a tree of nodes. Each node is an object with exactly one of these shapes:

| Node | Satisfied when |
| - | - |
| `{"term": "Airbus"}` | The word, or the exact phrase (`{"term": "Marine Le Pen"}`), appears in a searched Part |
| `{"field": "author", "equals": "Jane Doe"}` | The article's metadata field has this value ([fields](#field-filters)) |
| `{"all": [nodes]}` | Every node is satisfied |
| `{"any": [nodes]}` | At least one node is satisfied |
| `{"not": node}` | The node is not satisfied |

Bounds, which the schema enforces:

* `all`, `any` and `not` nest at most 6 levels deep;
* a group holds 1 to 64 nodes;
* a term or a value is at most 256 characters, and a term is not blank.

The core validates every Saved Query pinned to this plugin against the declared
`expression_schema` when the Subscription is created or versioned. A malformed tree is
`422 invalid_expression`, and the message names the JSON Pointer at fault.

<h3 id="why-a-structured-expression-not-a-query-string">
  Why a structured expression, not a query string
</h3>

The Plugin API lets the core validate an expression only against the plugin's
JSON Schema, when the Subscription is created. A query string such as
`"Airbus" AND (grève OR strike)` cannot be checked by a schema: a typo would be
accepted, and the saved search would then either never alert or fail at every
article. The structured tree is checked in full at creation, so a mistake is a 422
that the user sees at once. It is also unambiguous: it has no operator precedence to
misread. And it maps directly onto an "all of / any of / none of / exact phrase" form.

People still write queries as text. The [notation](#text-notation) below
translates to the tree, and the plugin ships its reference parser.

<h2 id="matching">
  Matching
</h2>

* **Parts.** Terms search the Parts whose role is in `text_roles`, which is `title` and
  `body` by default. Other roles, such as captions, are ignored.
* **Folding.** Text and terms are compared folded:

  * compatibility forms are decomposed (NFKD);
  * accents are removed;
  * `œ`, `æ` and a few other ligatures are spelled out;
  * case is folded.

  So `greve` finds "Grève", and `oeuvre` finds "Œuvre".
* **Whole words.** A word is a run of Unicode letters and digits. Everything else
  separates words: spaces, punctuation, apostrophes and hyphens. So:
  * `bus` does not match "Airbus";
  * `salarié` does not match "salariés", because there is no stemming;
  * `Airbus` matches "l'Airbus".
* **Phrases.** A term of several words matches them consecutively and in order, so
  `Saint-Denis` also finds "Saint Denis". A term without any letter or digit never
  matches.
* **`not`** is evaluated like any other node. A query made only of exclusions matches
  every article that lacks them.

<h2 id="field-filters">
  Field filters
</h2>

`field` is a field name or a JSON Pointer into the article's metadata (`/source/…`,
`/provenance/…`, `/extensions/…`, `/accepted_at`). The values compare like text:

* folded;
* as whole values, with runs of spaces collapsed;
* a number or a boolean compares by its JSON form;
* a list matches when one of its elements does.

A filter on its own is a valid alert, for example "every new article from this
source".

Built-in names read the metadata Quivr sends with every article:

| Name | Pointer |
| - | - |
| `source` | `/source/namespace` |
| `producer` | `/provenance/producer` |
| `origin` | `/provenance/origin` (`client` or `connector`) |
| `connector` | `/provenance/connector/instance_id` |
| `connector_kind` | `/provenance/connector/kind` (for example `rss`) |

Other names, such as `author` or `category`, depend on where your sources put that
metadata, so the installer maps them in the plugin configuration (below). An unmapped
name has no value: its filter is never satisfied, and the plugin logs a warning.

<h2 id="configuration">
  Configuration
</h2>

**Installer configuration**: the plugin pin's `configuration` in the core's startup config.

| Field | Default | Meaning |
| - | - | - |
| `fields` | `{}` | Name → JSON Pointer, added to or replacing the built-in names. Names match `[a-z][a-z0-9_]*` |
| `text_roles` | `["title", "body"]` | Part roles that terms search and described alerts send |
| `described.threshold` | `0.5` | Default score threshold of described alerts, 0.2 to 0.95 |

```json theme={null}
{"manifest": "plugins/alerts/quivr-plugin.yaml", "endpoint": "http://127.0.0.1:9910",
 "configuration": {"fields": {"author": "/extensions/example.news/data/author",
                              "category": "/extensions/example.news/data/categories"}}}
```

**Subscription configuration** (`evaluator.configuration`):

| Field | Default | Meaning |
| - | - | - |
| `wait_for_enrichment` | `false` for keywords, `true` for described | Answer `not_ready` until the article is enriched (embeddings attached). The core asks again on `record.enrichment_available` |
| `threshold` | installer's `described.threshold` | Described alerts only: the score at or above which the article matches, 0.2 to 0.95. Keyword alerts ignore it |

Rules run when an article becomes searchable. Keyword alerts need no enrichment, so
they decide at once by default. If a deployment never enriches articles, a Subscription
with `wait_for_enrichment: true` never decides.

**Pin `kinds`** (core startup config): the alert kinds this installation accepts.
An installation without `TYPESAFE_API_KEY` pins `"kinds": ["keywords"]`, so a
described alert is `422 invalid_expression` at creation. Absent, both kinds are
accepted.

**Secret:** `TYPESAFE_API_KEY`, read from the plugin's environment only and declared
in the manifest with `required: false`. `TYPESAFE_API_URL` optionally replaces the
System One endpoint, for example with the fake server in tests.

<h2 id="decisions-and-evidence">
  Decisions and evidence
</h2>

* **Decisions.** Each evaluation answers `match`, `no_match` or `not_ready`. The
  decision depends only on the article, the expression and the configurations, so
  batches can be deduplicated and replayed.
* **Evidence.** A `match` carries evidence, which Quivr stores with the Match and
  exposes through `GET /v0/matches/{id}`. It names the terms and field values that
  support the match, taken from the satisfied branches, never from under a `not`:

```json theme={null}
{
  "explanation": "Matched \"Airbus\" in title, body; \"grève\" in body.",
  "part_keys": ["title", "body"],
  "details": {
    "kind": "keywords",
    "terms": [{"term": "Airbus", "part_keys": ["title", "body"]}, {"term": "grève", "part_keys": ["body"]}],
    "fields": []
  }
}
```

* **Field filters.** A matched filter reads `Matched source "wire".` and appears in
  `details.fields` as `{"field": "source", "value": "wire"}`.
* **Exclusions only.** A match that rests only on exclusions explains
  `Matched: none of the excluded terms appear.` and carries no `part_keys`.
* **Bounds.** Terms appear in query order. `part_keys` follows the article's Part
  order, with at most 100 keys. `details` stays under 12 KiB; when it is trimmed,
  `details.truncated` is `true`.

<h2 id="described-alerts">
  Described alerts
</h2>

```json theme={null}
{"kind": "described", "description": "Labour strikes at ports and harbours"}
```

The description has 3 to 1000 characters, with at least one character that is not a
space. An optional `sources` (1 to 64 distinct Source Namespaces) limits the alert to
those sources, compared like the keyword `source` filter: an article from another
source is `no_match` at once, without waiting for enrichment and without a classifier
question for that alert. Each batch is decided as follows.

1. **One call.** All described evaluations of the batch that are ready and watch the
   article's source share one
   classifier call. Their descriptions are deduplicated after runs of spaces are
   collapsed, then sorted, so each distinct description is asked once. Evaluations
   that differ only in threshold share their question. The core already
   deduplicates identical expression and configuration pairs across Subscriptions
   and owners. There is no keyword pre-filter.
2. **What the classifier sees** (`state`):

   * `title`: the `title` Parts;
   * `source`: the Source Namespace;
   * every installer-mapped field that has a value (`fields`, such as `author` or
     `category`), with strings bounded to 256 characters and lists to 16 items;
   * `text`: the other `text_roles` Parts in Part order, joined by blank lines.

   The text is cut at a word boundary so the serialized state stays under 48 KB.
   That is far under Jev's budget of 32k tokens for the state plus the longest
   question, and it keeps the start of the article, where news puts the
   essentials.
3. **Jev request.** `POST https://api.typesafe.ai/v1/systemone` with the pinned model
   `jev-1.13.0`, `{"article": state}` as the state, and one
   [Noul](https://docs.typesafe.ai/primitives/noul.md) question per description.
   The question's instructions hold the description as data (`alert`) and ask
   whether the article is about what `alert` describes, in other words or another
   language. A request is split only when its body would exceed 120,000 bytes,
   under TypeSafe's limit of about 128 KB. The parts are sent in parallel (at most
   4 at once, 15 seconds each), within the declared `timeout_ms` of 20 seconds.
4. **Decision.** `match` when the Noul, the probability of "yes", is at or above the
   threshold; otherwise `no_match`, whose explanation gives the score. TypeSafe's
   `confidence` field is never used. Noul answers do not carry it, and it has no
   separating power for this decision.
5. **Evidence:**

   ```json theme={null}
   {
     "explanation": "Jev (jev-1.13.0) judged that the article fits the description: score 0.97, threshold 0.50.",
     "part_keys": ["title", "body"],
     "details": {"kind": "described", "classifier": "Jev", "model": "jev-1.13.0", "score": 0.97, "threshold": 0.5, "truncated": false}
   }
   ```

   `part_keys` are the Parts sent, even partly. `truncated` says the text was cut.

**Default threshold 0.5.** It was calibrated on `calibration/set.json`,
13 neutral French and English articles, each judged against 6 descriptions (78
pairs, including near misses such as a strike on the railways, or visa-free
tourism). With `jev-1.13.0` on 2026-09-29:

* the fitting pairs scored 0.94 to 0.99;
* the others scored 0.01 to 0.07;
* every threshold from 0.20 to 0.90 classified all 78 pairs correctly.

0.5 sits in the middle of that gap. The set is small and clear-cut, so tune the
threshold per alert on real traffic. Re-run the calibration with
`python3 calibration/calibrate.py` (it calls TypeSafe with `TYPESAFE_API_KEY`) before
moving to another model version.

**Errors.** The plugin never echoes TypeSafe's response body, and never logs the key.

| Situation | Plugin error | The core |
| - | - | - |
| No `TYPESAFE_API_KEY` and a described evaluation is due | `described_unavailable`, terminal | Isolates the described evaluations by halving the batch, decides the others, retries these with backoff |
| HTTP 401 or 403 | `classifier_unauthorized`, terminal | Same |
| Other 4xx, or an answer without a valid Noul | `classifier_refused_request`, `classifier_invalid_answer`, terminal | Same |
| Timeout, connection failure, HTTP 408, 429, 5xx or 529 | `classifier_unavailable`, retryable | Retries the whole batch with backoff; keyword evaluations of that batch wait too |

**Replaceable classifier.** `alerts/described.py` depends only on the `Classifier`
protocol: a `name`, a `model` and `judge(state, descriptions) -> {description: score}`.
`alerts/jev.py` implements it. Another classifier, such as an in-house model,
replaces `rule.classifier`, the factory that returns the classifier, or `None`
when described alerts cannot be decided.

<h2 id="text-notation">
  Text notation
</h2>

`alerts.notation` translates the query text people write into the expression:

```text theme={null}
"Airbus" AND (grève OR strike) NOT sport
author:"Jane Doe" OR source:wire
```

```text theme={null}
query   := or
or      := and ( "OR" and )*
and     := unary ( [ "AND" ] unary )*      juxtaposed items are all required
unary   := "NOT" unary | primary
primary := "(" or ")" | PHRASE | WORD | FIELD
PHRASE  := '"' characters '"'              \" and \\ escape a quote and a backslash
FIELD   := name ":" ( WORD | PHRASE )      name is [a-z][a-z0-9_]*
```

The notation follows these rules:

* `NOT` binds tighter than `AND`, which binds tighter than `OR`.
* Operators are upper case; `and`, `or` and `not` in lower case are ordinary words.
* `Marine Le Pen` without quotes requires the three words anywhere, and
  `"Marine Le Pen"` requires the phrase.
* A leading `-` is refused: write `NOT`.
* A word shaped `name:value` is a field filter, unless the value starts with `/`
  (a URL). Quote it to search it as text: `"re:Invent"`.
* Chains of the same operator are flattened, and parentheses around a single item add
  no level.

```bash theme={null}
python3 -m alerts.notation '"Airbus" AND (grève OR strike) NOT sport'   # prints the expression
```

It exits 1 and explains the mistake for an invalid query. From Python, use
`from alerts.notation import parse, NotationError`.

<h2 id="run-and-test">
  Run and test
</h2>

```bash theme={null}
python3 -m venv .venv && . .venv/bin/activate
pip install -e <quivr-v2 checkout>/sdks/python -e .
python3 -m unittest discover -s tests            # grammar, matching, evidence, schema, described alerts
quivr plugin dev --fixture fixtures/sample.json  # replay the keyword sample batch
quivr plugin test .                              # Contract Runner certification (keyword fixture)
```

Tests never call TypeSafe. Described alerts are tested against
`alerts.fake_system_one`, a deterministic stand-in for System One that judges by
topic words in English and French. To certify both kinds, run it and point the plugin
at it:

```bash theme={null}
python3 -m alerts.fake_system_one --port 8765 --key test-key &
TYPESAFE_API_KEY=test-key TYPESAFE_API_URL=http://127.0.0.1:8765/v1/systemone \
  quivr plugin test --fixture fixtures/sample.json --fixture tests/data/described.json .
```

One opt-in test calls the real API. It is skipped unless both
`QUIVR_ALERTS_LIVE=1` and `TYPESAFE_API_KEY` are set:
`QUIVR_ALERTS_LIVE=1 python3 -m unittest test_described.Live`, run from `tests/`.

`scripts/plugin_sdk.sh` (part of `make test`) runs the unit tests, the replay and the
certification with the fake server. CI publishes the Contract Runner report as the
`alerts-contract-report` artifact. The local stack pins
this plugin by default; `QUIVR_ALERTS=off make dev` leaves it out
([harness](/quivr-v2-local-harness)).
