// attack family 1 of 10 · by marginal value

PAIR

product carries itpersuasion marginal 75 pts n=4 refusal · n=4 product

An attacker model and a target model in a loop: the attacker proposes a prompt, sees how the target responded, and rewrites the prompt to do better … Full definition in the glossary →

This page answers one question about PAIR: what does a security product add here that the assistant does not already do on its own? The two sides are measured separately, as AMTSO's Guidelines for Testing of Agentic Security Products v1.0 require, and the difference is the part a buyer is paying for.

0%
0 of 4 measurable samples
Claude Sonnet, 5 turns each, majority
assistant refuses alone
75%
3 of 4 attack samples
MoorAI agent 0.80.5, test half
product catches
75 pts
detection rate minus refusal rate
different subsets, so rate versus rate
marginal value
75% caught (3/4) − 0% refused unaided (0/4) = 75 pts
01 — what the gap means

The product is the only thing standing here

The deployed assistant answers these on its own. Everything that gets caught is caught by the product, so the whole detection rate is marginal value.

For PAIR specifically: the Claude Sonnet assistant refused 0 of 4 probed samples and answered the other 4. Every one of the 3 catches on this family is therefore work the assistant did not do, which makes the raw detection rate and the marginal value nearly the same number here — the rare case where a vendor's headline figure is honest by accident.

This is the shape of family worth buying a tool for. The assistant answers the request on its own, so whatever sits in front of it is the only thing that can stop it.

02 — how to read the two numbers

Rate against rate, not sample against sample

The refusal side saw 4 measurable PAIR samples; the product side scored 4. Product recall is counted over every attack, refusal only over the samples the model actually received, so the denominators differ. No row here says "the product caught the exact sample the assistant let through" — it says one rate is higher than another, on the same family, in the same taxonomy.

Refusal is decided per sample by majority across 5 turns, because a frontier assistant is not deterministic: three refusals out of five counts as a refusal, two does not. Product detection is deterministic in this harness and needs no trials. Both sides come from a committed aggregate file in this repository, joined at build time — see provenance.

03 — limits

What the PAIR numbers cannot support

All four bound every figure on this page. They are here, not in a footnote, because they are the reason the rest is worth reading.

01The refusal side is the deployed Claude Sonnet assistant, not a bare model

Model refusal was measured through the claude CLI: what a buyer actually experiences is the model plus Anthropic's own layered defences. A different assistant refuses differently, so these numbers describe this assistant, not models in general. Glossary: model refusal →

02The per-family sample counts are tiny

Each family carries 2 to 10 samples. Counts are printed next to every percentage because a rate over three samples is a rate over three samples. The counts are far too small for confidence intervals. The direction is the finding; the precise value is not. Glossary: held out set →

03Refusal rate and product recall are rate versus rate, not paired per sample

Product recall is counted over every attack in the family; refusal rate is counted only over the samples that reached the model. Platform-blocked samples carry no refusal verdict, so the two denominators differ. Read each row as one rate against another, never as 'the product caught the exact sample the assistant let through'. Glossary: marginal value →

04The held-out test set is a consumable, and it has now been scored more than once

Product recall here is on the held-out test half — the half the detectors were not tuned against, which is what makes it out-of-sample. But a held-out set stops being pristine the moment it is looked at, and this one has now been scored more than once. Treat the test-half recall as a good estimate, not an untouched one. Glossary: tune test split →

Concretely for this family: 0% is 0 out of 4, and 75% is 3 out of 4. One sample moving either way changes the percentage by double digits. Treat the direction as the finding and the value as an estimate.

04 — the other families

Same product, same run, different answers

The 10 families are scored in one run with one configuration. Across the measurable ones the marginal value still ranges from -67 pts to 75 pts, which is why a single site-wide detection rate hides more than it shows.

05 — provenance

Where the numbers come from

Refusal — Claude Sonnet via the local claude CLI — the deployed assistant (model plus Anthropic's own defences), not a bare model call · 5 turns per sample, majority decides.
Corpusheldout-v2-test, the held-out test half.
Product — MoorAI agent 0.80.5, scored out-of-sample on that same test half.
Aggregate — refusal-baseline-runs-claude.json (Claude Sonnet, 5 runs/sample) joined with the MoorAI 0.80.5 engine on the held-out test half. Counts and rates only; no prompt or response text is vendored into this repository.
Corpora and harnessopen benchmark repository ↗.
All 10 familiesPAIR in the glossary Leaderboard The benchmark How we test

AMTSO has not reviewed, certified or endorsed MoorAI or this page; the guidelines are public and we score against them ourselves. Refusal is the deployed Claude Sonnet assistant on one configuration, product recall is out-of-sample on a held-out test half that has now been scored more than once, and the sample counts are single digits — all three bound every number here. Corpora and harness are in the open benchmark repository ↗.

glick.run — AGPL-3.0