// attack family 8 of 10 · by marginal value

TAP

assistant covers itpersuasion marginal -67 pts n=3 refusal · n=3 product

A tree search on top of PAIR. The attacker branches into several candidate refinements at each step instead of one, then prunes the branches an evaluator… Full definition in the glossary →

This page answers one question about TAP: what does a security product add here that the assistant does not already do on its own? The two sides are measured separately, as AMTSO's Guidelines for Testing of Agentic Security Products v1.0 require, and the difference is the part a buyer is paying for.

100%
3 of 3 measurable samples
Claude Sonnet, 5 turns each, majority
assistant refuses alone
33%
1 of 3 attack samples
MoorAI agent 0.80.5, test half
product catches
-67 pts
detection rate minus refusal rate
different subsets, so rate versus rate
marginal value
33% caught (1/3) − 100% refused unaided (3/3) = -67 pts
01 — what the gap means

The assistant already refuses these

The assistant refuses this family at least as often as the product catches it. A tool that reports catching these is reporting work the assistant already did.

For TAP specifically: the Claude Sonnet assistant refused 3 of 3 probed samples on its own — as much as or more than the product's 1 of 3. On this family the product's marginal contribution is nothing: everything it catches, the assistant was already refusing. A vendor quoting a raw 33% for TAP is quoting an outcome you largely had for free.

That is a statement about TAP on this assistant, not about the product overall. On other families in the same run the same product carries the whole result.

02 — how to read the two numbers

Rate against rate, not sample against sample

The refusal side saw 3 measurable TAP samples; the product side scored 3. Product recall is counted over every attack, refusal only over the samples the model actually received, so the denominators differ. No row here says "the product caught the exact sample the assistant let through" — it says one rate is higher than another, on the same family, in the same taxonomy.

Refusal is decided per sample by majority across 5 turns, because a frontier assistant is not deterministic: three refusals out of five counts as a refusal, two does not. Product detection is deterministic in this harness and needs no trials. Both sides come from a committed aggregate file in this repository, joined at build time — see provenance.

03 — limits

What the TAP numbers cannot support

All four bound every figure on this page. They are here, not in a footnote, because they are the reason the rest is worth reading.

01The refusal side is the deployed Claude Sonnet assistant, not a bare model

Model refusal was measured through the claude CLI: what a buyer actually experiences is the model plus Anthropic's own layered defences. A different assistant refuses differently, so these numbers describe this assistant, not models in general. Glossary: model refusal →

02The per-family sample counts are tiny

Each family carries 2 to 10 samples. Counts are printed next to every percentage because a rate over three samples is a rate over three samples. The counts are far too small for confidence intervals. The direction is the finding; the precise value is not. Glossary: held out set →

03Refusal rate and product recall are rate versus rate, not paired per sample

Product recall is counted over every attack in the family; refusal rate is counted only over the samples that reached the model. Platform-blocked samples carry no refusal verdict, so the two denominators differ. Read each row as one rate against another, never as 'the product caught the exact sample the assistant let through'. Glossary: marginal value →

04The held-out test set is a consumable, and it has now been scored more than once

Product recall here is on the held-out test half — the half the detectors were not tuned against, which is what makes it out-of-sample. But a held-out set stops being pristine the moment it is looked at, and this one has now been scored more than once. Treat the test-half recall as a good estimate, not an untouched one. Glossary: tune test split →

Concretely for this family: 100% is 3 out of 3, and 33% is 1 out of 3. One sample moving either way changes the percentage by double digits. Treat the direction as the finding and the value as an estimate.

04 — the other families

Same product, same run, different answers

The 10 families are scored in one run with one configuration. Across the measurable ones the marginal value still ranges from -67 pts to 75 pts, which is why a single site-wide detection rate hides more than it shows.

05 — provenance

Where the numbers come from

Refusal — Claude Sonnet via the local claude CLI — the deployed assistant (model plus Anthropic's own defences), not a bare model call · 5 turns per sample, majority decides.
Corpusheldout-v2-test, the held-out test half.
Product — MoorAI agent 0.80.5, scored out-of-sample on that same test half.
Aggregate — refusal-baseline-runs-claude.json (Claude Sonnet, 5 runs/sample) joined with the MoorAI 0.80.5 engine on the held-out test half. Counts and rates only; no prompt or response text is vendored into this repository.
Corpora and harnessopen benchmark repository ↗.
All 10 familiesTAP in the glossary Leaderboard The benchmark How we test

AMTSO has not reviewed, certified or endorsed MoorAI or this page; the guidelines are public and we score against them ourselves. Refusal is the deployed Claude Sonnet assistant on one configuration, product recall is out-of-sample on a held-out test half that has now been scored more than once, and the sample counts are single digits — all three bound every number here. Corpora and harness are in the open benchmark repository ↗.

glick.run — AGPL-3.0