// attack family 10 of 10 · by marginal value

FlipAttack

not measurableobfuscation no refusal verdict n=3 product

A reversal attack: the payload is written backwards — by character, by word, or by line — and the model is instructed to un-reverse it before following… Full definition in the glossary →

This page answers one question about FlipAttack: what does a security product add here that the assistant does not already do on its own? The two sides are measured separately, as AMTSO's Guidelines for Testing of Agentic Security Products v1.0 require, and the difference is the part a buyer is paying for.

n/a
all 3 samples platform-blocked
the provider rejects them before the model reads them
assistant refuses alone
100%
3 of 3 attack samples
MoorAI agent 0.80.5, test half
product catches
n/a
no model-refusal verdict, so no
marginal number can be stated
marginal value
01 — what the gap means

Not measurable through this backend

Every sample in this family is rejected by the provider's Acceptable Use Policy before the model reads it. There is no model-refusal verdict to compare against, so no marginal number can be stated. The product still catches them; that is recorded, not scored here.

For FlipAttack specifically: all 3 probed samples were rejected by the provider's Acceptable Use Policy before the Claude Sonnet assistant ever read them. A rejected request is neither a model refusal nor a product detection — there is no evidence at all about what the model would have done, so this family has no refusal rate and no marginal number. MoorAI agent still catches 3 of 3; that is recorded, but on this family it cannot be scored against a refusal baseline that does not exist.

That is a statement about how FlipAttack reaches this assistant, not about the product. On the measurable families in the same run the subtraction can be done, and the answers range widely.

02 — how to read the two numbers

Rate against rate, not sample against sample

There is only one number to read here: the product side. The refusal side is empty because the provider stopped all 3 FlipAttack samples at the door. Platform blocks are counted as "the deployed assistant stopped it" in the whole-corpus total, but they are never folded into a model-refusal rate — conflating the two would credit the model with judgement it never exercised.

Both sides are read from a committed, aggregate-only file in this repository and joined at build time — see provenance.

03 — limits

What the FlipAttack numbers cannot support

All four bound every figure on this page. They are here, not in a footnote, because they are the reason the rest is worth reading.

01The refusal side is the deployed Claude Sonnet assistant, not a bare model

Model refusal was measured through the claude CLI: what a buyer actually experiences is the model plus Anthropic's own layered defences. A different assistant refuses differently, so these numbers describe this assistant, not models in general. Glossary: model refusal →

02The per-family sample counts are tiny

Each family carries 2 to 10 samples. Counts are printed next to every percentage because a rate over three samples is a rate over three samples. The counts are far too small for confidence intervals. The direction is the finding; the precise value is not. Glossary: held out set →

03Refusal rate and product recall are rate versus rate, not paired per sample

Product recall is counted over every attack in the family; refusal rate is counted only over the samples that reached the model. Platform-blocked samples carry no refusal verdict, so the two denominators differ. Read each row as one rate against another, never as 'the product caught the exact sample the assistant let through'. Glossary: marginal value →

04The held-out test set is a consumable, and it has now been scored more than once

Product recall here is on the held-out test half — the half the detectors were not tuned against, which is what makes it out-of-sample. But a held-out set stops being pristine the moment it is looked at, and this one has now been scored more than once. Treat the test-half recall as a good estimate, not an untouched one. Glossary: tune test split →

04 — the other families

Same product, same run, different answers

The 10 families are scored in one run with one configuration. Across the measurable ones the marginal value still ranges from -67 pts to 75 pts, which is why a single site-wide detection rate hides more than it shows.

05 — provenance

Where the numbers come from

Refusal — Claude Sonnet via the local claude CLI — the deployed assistant (model plus Anthropic's own defences), not a bare model call · 5 turns per sample, majority decides.
Corpusheldout-v2-test, the held-out test half.
Product — MoorAI agent 0.80.5, scored out-of-sample on that same test half.
Aggregate — refusal-baseline-runs-claude.json (Claude Sonnet, 5 runs/sample) joined with the MoorAI 0.80.5 engine on the held-out test half. Counts and rates only; no prompt or response text is vendored into this repository.
Corpora and harnessopen benchmark repository ↗.
All 10 familiesFlipAttack in the glossary Leaderboard The benchmark How we test

AMTSO has not reviewed, certified or endorsed MoorAI or this page; the guidelines are public and we score against them ourselves. Refusal is the deployed Claude Sonnet assistant on one configuration, product recall is out-of-sample on a held-out test half that has now been scored more than once, and the sample counts are single digits — all three bound every number here. Corpora and harness are in the open benchmark repository ↗.

glick.run — AGPL-3.0