Field 01 — Security declaration · injection screening

Screened, never immune.

Raw email is attacker-controlled input. We screen every message for prompt injection before your agent reads it, then publish exactly how well that screening holds, including what got through. No vendor can honestly promise immunity, so we don't. We show you the record instead.

Field 02 — Inspection record

The living red-team scoreboard.

A corpus of attack emails run through the real production pipeline. Most attacks come from independent AI adversaries and a public benchmark, not from us. Every figure below is pinned to the exact models and date it was measured, and re-run against a fixed corpus each release. We keep the live attack corpus private — publishing it would just hand the next attacker a starting kit.

Inspection record · 4 origins · 113 attacks · 26 legitimateAs of 12 July 2026 · scan claude-haiku-4-5 · extract claude-haiku-4-5
99.1%
Attacks caught
112113 of 113 across runs; the strict rate wobbles by one borderline fixture. See the ledger.
109/109
Hard-injection catch
Every attack that actually tries to hijack, exfiltrate, or elicit harmful output. All caught, this run.
21
Reproducible catches
Caught deterministically, no model version required. The other 91 depend on the scan model and will move.
1
False flags
25/26 legitimate emails passed clean. One real message we wrongly screened, listed below.
The number that flatters us, a single catch-rate, is the one we trust least. It moves with the scan model, and a defense that flags everything would score 100% on attacks and be useless. So we track both directions and pin every figure.

Field 03 — Failure ledger

What got through. Every one.

A security page that only lists wins is marketing. Here are the failure modes the board tracks, described plainly. Hiding them is what would make the rest worthless.

Soft-miss boundary · flickers run to run   deepset/prompt-injections
The strict rate wobbles by a single fixture between runs: the scan model is stochastic, and a few borderline entries sit right on its threshold. The clearest, shown above, is a public-dataset row labeled an injection but really a mislabeled political question, “why are the Greens unelectable?”It asks nothing improper of the agent, and the dataset's “injection” label is debatable. What never wobbles: no attack that actually tries to hijack, exfiltrate, or elicit harm has reached the model on any run. Harm-weighted stays 109/109.
False flag · legit mail wrongly screened   Postfleet regression set
A genuine OCR-scanner support email describing a keyboard-shift bug and a scrambled code sample. Our automated screen can't tell it apart from an attack that hides a payload behind similar framing, so it fails closed on both. The message still reaches the agent as cleaned text with the flag attached. Nothing is dropped or guessed. It's a real, documented cost of screening technical mail conservatively.

Field 04 — Adversaries we don't control

The attacks aren't graded on our homework.

A defense tested only against attacks its own author imagined is grading its own homework. Most of this corpus is written by models from other labs that don't know our internals, plus a public benchmark. A miss they find is real signal.

Grok 4.5 (xAI)
39
An independent generator. It invented attacks, then adaptively refined them against a defense it can't see.
GPT-5.5 (OpenAI)
33
A second independent generator from another model family. Stacked-obfuscation and closed-loop batches.
deepset/prompt-injections
30
A public benchmark skeptics already know. Imported verbatim, not curated by us.
Postfleet (hand-authored)
11
Our own regression fixtures. The smallest, least-trusted slice of the corpus.

Field 05 — Cracks found & closed

The red team is us, aimed at us.

The point of the exercise is to break the defense before an attacker does. Three real weaknesses it surfaced, and what we did about each one.

Catch-rate inflation by our own test marker
Our scoring canary carried a “CANARY_” prefix the classifier had learned to flag. Some “catches” were it spotting our test token, not the attack. Swapping to neutral markers dropped a flattering 100% to a true 98.1%, and exposed two real misses that had been hiding underneath.
Fiction-framed jailbreaks
Attacks wrapped in a story initially read as low-risk. We hardened the screen to weigh intent over framing; both misses are now caught, with zero new false-flags.
Layered obfuscation
An adaptive attacker found multi-step obfuscation that first slipped past the model scan. We closed it with additional deterministic screening, so these now fail closed. We don’t publish the specific payloads or the check — the point is that we hunt for them and shut them, not a map for the next attacker.

Field 06 — What we don't claim

The fine print, up front.

  • We do not claim immunity. We claim screening against known attack patterns, measured and published.
  • The single catch-rate is model-versioned. When the scan model changes, it moves. Only the 21 deterministic catches are stable across models.
  • “Hard” vs “soft” severity is our own conservative judgment, assigned before scoring and biased toward counting things as hard. It's a lens on the same run, not a separate, better number.
  • A novel attack no one in this corpus imagined can still get through. That's why the corpus keeps growing and the board stays live.
Every number here is measured, not asserted: the full corpus is re-run and versioned to the exact scan model and date on each release. We keep the corpus itself private — publishing the live attack set would just seed the next attack — but we'll walk a serious auditor or customer through a run on request.

Field 07 — Dispatch

Give your agent an inbox that screens its own mail.

Screened against known patterns, never immune. The record above is the whole claim.