The harmful request is encoded in a cipher — Caesar shift, ROT13, base64, hex, a custom substitution — and the model is asked to work in that cipher. Full definition in the glossary →
This page answers one question about CipherChat: what does a security product add here that the assistant does not already do on its own? The two sides are measured separately, as AMTSO's Guidelines for Testing of Agentic Security Products v1.0 require, and the difference is the part a buyer is paying for.
Every sample in this family is rejected by the provider's Acceptable Use Policy before the model reads it. There is no model-refusal verdict to compare against, so no marginal number can be stated. The product still catches them; that is recorded, not scored here.
For CipherChat specifically: all 4 probed samples were rejected by the provider's Acceptable Use Policy before the Claude Sonnet assistant ever read them. A rejected request is neither a model refusal nor a product detection — there is no evidence at all about what the model would have done, so this family has no refusal rate and no marginal number. MoorAI agent still catches 4 of 4; that is recorded, but on this family it cannot be scored against a refusal baseline that does not exist.
That is a statement about how CipherChat reaches this assistant, not about the product. On the measurable families in the same run the subtraction can be done, and the answers range widely.
There is only one number to read here: the product side. The refusal side is empty because the provider stopped all 4 CipherChat samples at the door. Platform blocks are counted as "the deployed assistant stopped it" in the whole-corpus total, but they are never folded into a model-refusal rate — conflating the two would credit the model with judgement it never exercised.
Both sides are read from a committed, aggregate-only file in this repository and joined at build time — see provenance.
All four bound every figure on this page. They are here, not in a footnote, because they are the reason the rest is worth reading.
Model refusal was measured through the claude CLI: what a buyer actually experiences is the model plus Anthropic's own layered defences. A different assistant refuses differently, so these numbers describe this assistant, not models in general. Glossary: model refusal →
Each family carries 2 to 10 samples. Counts are printed next to every percentage because a rate over three samples is a rate over three samples. The counts are far too small for confidence intervals. The direction is the finding; the precise value is not. Glossary: held out set →
Product recall is counted over every attack in the family; refusal rate is counted only over the samples that reached the model. Platform-blocked samples carry no refusal verdict, so the two denominators differ. Read each row as one rate against another, never as 'the product caught the exact sample the assistant let through'. Glossary: marginal value →
Product recall here is on the held-out test half — the half the detectors were not tuned against, which is what makes it out-of-sample. But a held-out set stops being pristine the moment it is looked at, and this one has now been scored more than once. Treat the test-half recall as a good estimate, not an untouched one. Glossary: tune test split →
The 10 families are scored in one run with one configuration. Across the measurable ones the marginal value still ranges from -67 pts to 75 pts, which is why a single site-wide detection rate hides more than it shows.
heldout-v2-test, the held-out test half.AMTSO has not reviewed, certified or endorsed MoorAI or this page; the guidelines are public and we score against them ourselves. Refusal is the deployed Claude Sonnet assistant on one configuration, product recall is out-of-sample on a held-out test half that has now been scored more than once, and the sample counts are single digits — all three bound every number here. Corpora and harness are in the open benchmark repository ↗.