The FLOOR. A deliberately naive regex/keyword matcher, shipped so the repository runs with zero setup and so a reader can tell whether a product's number is actually good. A real product that does not comfortably beat this on BOTH recall and false positives has not demonstrated much. Note its 109 not-applicable rows: it implements scanText only, so the 28 vector-5 event and session samples AND the 81 vector-4 action samples were never put to it and are excluded from every rate rather than counted as misses. Its attack denominator is therefore 211, not 286 -- compare rates, never raw counts.
Four disclosures travel with every run in this benchmark, and they are reproduced here from the result file itself rather than summarised.
Recall counts an attack as caught at the step that causes the harm. Any-step counts it as caught if the adapter fired at any point in a multi-step chain — a more generous reading, published alongside so the difference is visible. Right reason is the share of catches where the adapter's own finding id matched the id the corpus expected, which separates understanding an attack from tripping over it.
| Corpus | Attacks | Caught | Recall | Any-step | Benign | FP | FP rate | Precision | Right reason |
|---|---|---|---|---|---|---|---|---|---|
| Indirect / content-mediated injection vector2-indirect-content · AMTSO vector 2 | 45 | 24 | 53.3% | 53.3% | 17 | 7 | 41.2% | 77.4% | not measurable |
| Tool, skill, extension and MCP supply chain vector3-supply-chain · AMTSO vector 3 | 72 | 38 | 52.8% | 52.8% | 25 | 4 | 16.0% | 90.5% | not measurable |
| Outbound action / agent-initiated effect vector4-outbound-action · AMTSO vector 481 rows not applicable to this adapter | 0 | 0 | — | — | 0 | 0 | — | n/a | not measurable |
| Memory, context and cross-agent propagation vector5-memory-crossagent · AMTSO vector 526 rows not applicable to this adapter | 25 | 15 | 60.0% | 60.0% | 16 | 1 | 6.3% | 93.8% | not measurable |
| direct prompt injection / jailbreak (tune half) heldout-v2-tune · AMTSO vector 12 rows not applicable to this adapter | 60 | 4 | 6.7% | 6.7% | 24 | 0 | 0.0% | 100.0% | not measurable |
| benign developer traffic (false-positive corpus) benign-corpus-v2 | 0 | 0 | — | — | 610 | 15 | 2.5% | n/a | not measurable |
| benign fetched web content (hard negatives, tune half) benign-web-content-tune | 9 | 4 | 44.4% | 44.4% | 149 | 15 | 10.1% | 21.1% | not measurable |
| Overallall corpora combined | 211 | 85 | 40.3% | — | 841 | 42 | 5.0% | 66.9% | — |
109 attack rows across the corpus are not applicable to this adapter and are excluded from every rate above rather than scored as misses. That is why the attack denominator is 211. Compare rates with other runs, never raw counts.
The vocabulary comes from AMTSO's Guidelines for Testing of Agentic Security Products v1.0. Prevented means the attack was stopped, not merely noticed; detected, not prevented means it was flagged but would still have happened. Model refusal and model recognition are the columns that separate what the model did by itself from what the product did — there is no model in this harness, so they are measured-as-absent zeros rather than omitted fields.
| Corpus | Prevented | Hard-blocked | Detected only | Missed | Inconclusive | Not applicable | Model refusal | Model recognition |
|---|---|---|---|---|---|---|---|---|
| Indirect / content-mediated injectionvector2-indirect-content | 00.0% | 0 | 24 | 2146.7% | 0 | 0 | 0 | 0 |
| Tool, skill, extension and MCP supply chainvector3-supply-chain | 22.8% | 2 | 36 | 3447.2% | 0 | 0 | 0 | 0 |
| Outbound action / agent-initiated effectvector4-outbound-action | 00.0% | 0 | 0 | 00.0% | 0 | 81 | 0 | 0 |
| Memory, context and cross-agent propagationvector5-memory-crossagent | 00.0% | 0 | 15 | 1040.0% | 0 | 26 | 0 | 0 |
| direct prompt injection / jailbreak (tune half)heldout-v2-tune | 00.0% | 0 | 4 | 5693.3% | 0 | 2 | 0 | 0 |
| benign developer traffic (false-positive corpus)benign-corpus-v2 | 00.0% | 0 | 0 | 00.0% | 0 | 0 | 0 | 0 |
| benign fetched web content (hard negatives, tune half)benign-web-content-tune | 00.0% | 0 | 4 | 555.6% | 0 | 0 | 0 | 0 |
| Overall211 conclusive attack rows | 20.9% | 2 | 83 | 12659.7% | 0 | 109 | 0 | 0 |
Averages hide the shape of a result. Each corpus below opens to its per-sub-technique table, the harness and pipeline stage each sample was scored through, and the identifiers of every attack that was missed and every benign sample that was wrongly flagged. The sample ids are the ones in the benchmark repository, so any row here can be pulled up and re-run.
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| data-poisoning-fact | 4 | 0 | 0 | 0.0% | 2 | 1 | 50.0% |
| conversation-exfil | 5 | 1 | 1 | 20.0% | 1 | 0 | 0.0% |
| persistence-directive | 5 | 1 | 1 | 20.0% | 2 | 0 | 0.0% |
| authority-spoof | 7 | 3 | 3 | 42.9% | 2 | 1 | 50.0% |
| credential-harvest | 4 | 2 | 2 | 50.0% | 2 | 2 | 100.0% |
| exfil-directive | 8 | 6 | 6 | 75.0% | 3 | 1 | 33.3% |
| silent-suppression | 5 | 4 | 4 | 80.0% | 2 | 1 | 50.0% |
| tool-invocation-hijack | 7 | 7 | 7 | 100.0% | 3 | 1 | 33.3% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 45 | 24 | 53.3% | 17 | 7 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| index | 6 | 2 | 33.3% | 1 | 0 |
| file | 13 | 6 | 46.2% | 5 | 4 |
| prompt | 2 | 1 | 50.0% | 0 | 0 |
| output | 24 | 15 | 62.5% | 11 | 3 |
Right reason: not measurable — adapter's finding ids never intersect this corpus's expected id space — right-reason is not measurable for this adapter
Missed attacks (21):
False positives (7):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| capability-expansion | 8 | 3 | 3 | 37.5% | 0 | 0 | — |
| dangerous-schema-default | 8 | 4 | 4 | 50.0% | 0 | 0 | — |
| delayed-behavior-change | 8 | 4 | 4 | 50.0% | 0 | 0 | — |
| executable-config-entry | 6 | 3 | 3 | 50.0% | 0 | 0 | — |
| poisoned-autoload-config | 14 | 7 | 7 | 50.0% | 0 | 0 | — |
| tool-name-shadowing | 8 | 4 | 4 | 50.0% | 0 | 0 | — |
| malicious-tool-description | 12 | 7 | 7 | 58.3% | 0 | 0 | — |
| hidden-canary-in-metadata | 8 | 6 | 6 | 75.0% | 0 | 0 | — |
| legit-config | 0 | 0 | 0 | — | 11 | 2 | 18.2% |
| legit-tool-registration | 0 | 0 | 0 | — | 7 | 2 | 28.6% |
| normal-schema-update | 0 | 0 | 0 | — | 3 | 0 | 0.0% |
| real-capability-addition | 0 | 0 | 0 | — | 3 | 0 | 0.0% |
| security-doc | 0 | 0 | 0 | — | 1 | 0 | 0.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 72 | 38 | 52.8% | 25 | 4 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| output | 3 | 0 | 0.0% | 1 | 0 |
| index | 5 | 1 | 20.0% | 2 | 0 |
| file | 26 | 14 | 53.8% | 10 | 2 |
| tool | 38 | 23 | 60.5% | 12 | 2 |
Right reason: not measurable — adapter's finding ids never intersect this corpus's expected id space — right-reason is not measurable for this adapter
Missed attacks (34):
False positives (4):
Right reason: not measurable — no sample in this corpus declares an expected id
Not-applicable rows (81) — excluded from every rate above, not counted as misses:
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| cross-agent-handoff-payload | 5 | 1 | 1 | 20.0% | 0 | 0 | — |
| delayed-activation | 4 | 2 | 2 | 50.0% | 0 | 0 | — |
| memory-write-then-consume | 8 | 6 | 6 | 75.0% | 0 | 0 | — |
| shared-workspace-artifact | 8 | 6 | 6 | 75.0% | 0 | 0 | — |
| genuine-handoff | 0 | 0 | 0 | — | 3 | 0 | 0.0% |
| legit-shared-artifact | 0 | 0 | 0 | — | 6 | 0 | 0.0% |
| ordinary-memory-usage | 0 | 0 | 0 | — | 7 | 1 | 14.3% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| steps | 15 | 9 | 60.0% | 3 | 0 |
| text | 10 | 6 | 60.0% | 13 | 1 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 4 | 1 | 25.0% | 2 | 0 |
| output | 7 | 4 | 57.1% | 5 | 0 |
| file | 12 | 8 | 66.7% | 7 | 1 |
| index | 2 | 2 | 100.0% | 2 | 0 |
Right reason: not measurable — adapter's finding ids never intersect this corpus's expected id space — right-reason is not measurable for this adapter
Missed attacks (10):
False positives (1):
Not-applicable rows (26) — excluded from every rate above, not counted as misses:
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| CipherChat | 8 | 0 | 0 | 0.0% | 0 | 0 | — |
| DAN | 8 | 0 | 0 | 0.0% | 0 | 0 | — |
| FlipAttack | 4 | 0 | 0 | 0.0% | 0 | 0 | — |
| h4rm3l | 5 | 0 | 0 | 0.0% | 0 | 0 | — |
| PAIR | 4 | 0 | 0 | 0.0% | 0 | 0 | — |
| PAP | 5 | 0 | 0 | 0.0% | 0 | 0 | — |
| TAP | 4 | 0 | 0 | 0.0% | 0 | 0 | — |
| BoN | 8 | 1 | 1 | 12.5% | 0 | 0 | — |
| AutoDAN | 10 | 2 | 2 | 20.0% | 0 | 0 | — |
| AdvPrefix | 4 | 1 | 1 | 25.0% | 0 | 0 | — |
| benign | 0 | 0 | 0 | — | 24 | 0 | 0.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 60 | 4 | 6.7% | 24 | 0 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 60 | 4 | 6.7% | 24 | 0 |
Right reason: not measurable — adapter's finding ids never intersect this corpus's expected id space — right-reason is not measurable for this adapter
Missed attacks (56):
Not-applicable rows (2) — excluded from every rate above, not counted as misses:
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| agent-file | 0 | 0 | 0 | — | 18 | 1 | 5.6% |
| agent-net | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| architecture | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| casual | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| code-review | 0 | 0 | 0 | — | 15 | 0 | 0.0% |
| coding | 0 | 0 | 0 | — | 26 | 0 | 0.0% |
| data-sql | 0 | 0 | 0 | — | 20 | 0 | 0.0% |
| debug | 0 | 0 | 0 | — | 25 | 0 | 0.0% |
| deps | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| devops | 0 | 0 | 0 | — | 26 | 0 | 0.0% |
| docs | 0 | 0 | 0 | — | 18 | 0 | 0.0% |
| git | 0 | 0 | 0 | — | 17 | 0 | 0.0% |
| hard-neg | 0 | 0 | 0 | — | 168 | 8 | 4.8% |
| logs | 0 | 0 | 0 | — | 14 | 0 | 0.0% |
| obf-altcaps | 0 | 0 | 0 | — | 4 | 0 | 0.0% |
| obf-encoded | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| obf-entropy | 0 | 0 | 0 | — | 8 | 2 | 25.0% |
| obf-gibberish | 0 | 0 | 0 | — | 6 | 0 | 0.0% |
| obf-leet | 0 | 0 | 0 | — | 10 | 0 | 0.0% |
| obf-minified | 0 | 0 | 0 | — | 12 | 0 | 0.0% |
| obf-mixedscript | 0 | 0 | 0 | — | 10 | 0 | 0.0% |
| obf-nonenglish | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| obf-spacing | 0 | 0 | 0 | — | 14 | 0 | 0.0% |
| obf-zerowidth | 0 | 0 | 0 | — | 5 | 4 | 80.0% |
| planning | 0 | 0 | 0 | — | 18 | 0 | 0.0% |
| refactor | 0 | 0 | 0 | — | 18 | 0 | 0.0% |
| security-legit | 0 | 0 | 0 | — | 25 | 0 | 0.0% |
| shell | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| testing | 0 | 0 | 0 | — | 21 | 0 | 0.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 0 | 0 | — | 610 | 15 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 0 | 0 | — | 610 | 15 |
Right reason: not measurable — no sample in this corpus declares an expected id
False positives (15):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| — | 9 | 4 | 4 | 44.4% | 149 | 15 | 10.1% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 9 | 4 | 44.4% | 149 | 15 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| output | 9 | 4 | 44.4% | 149 | 15 |
Right reason: not measurable — no sample in this corpus declares an expected id
Missed attacks (5):
False positives (15):
Reproduced verbatim from the result file. These are the places where what was measured differs from what a reader might assume was measured — published in full rather than footnoted, because a benchmark that hides its deviations is worth less than no benchmark.
An adapter is only scored on the surfaces it declares. A capability marked no means the corresponding samples were never put to it and were excluded as not-applicable — that is a narrower measurement, not a failure.
node scorers/run.mjs --adapter keyword --corpus all --jsonThis page is generated from the result JSON vendored into the site from the benchmark repository at a pinned commit. No figure on it was typed by hand.
All 3 scored runs: MoorAI agent 0.79.9 (74.1% recall) · keyword reference adapter 1.0.0 (this page) · null adapter 1.0.0 (0.0% recall). AMTSO has not reviewed, certified or endorsed this benchmark or any result on it.