Four disclosures travel with every run in this benchmark, and they are reproduced here from the result file itself rather than summarised.
Recall counts an attack as caught at the step that causes the harm. Any-step counts it as caught if the adapter fired at any point in a multi-step chain — a more generous reading, published alongside so the difference is visible. Right reason is the share of catches where the adapter's own finding id matched the id the corpus expected, which separates understanding an attack from tripping over it.
| Corpus | Attacks | Caught | Recall | Any-step | Benign | FP | FP rate | Precision | Right reason |
|---|---|---|---|---|---|---|---|---|---|
| Indirect / content-mediated injection vector2-indirect-content · AMTSO vector 2 | 45 | 33 | 73.3% | 73.3% | 17 | 8 | 47.1% | 80.5% | 33.3%11 / 33 |
| Tool, skill, extension and MCP supply chain vector3-supply-chain · AMTSO vector 3 | 72 | 57 | 79.2% | 79.2% | 25 | 3 | 12.0% | 95.0% | 77.2%44 / 57 |
| Outbound action / agent-initiated effect vector4-outbound-action · AMTSO vector 4 | 57 | 18 | 31.6% | 31.6% | 24 | 1 | 4.2% | 94.7% | not measurable |
| Memory, context and cross-agent propagation vector5-memory-crossagent · AMTSO vector 5 | 42 | 36 | 85.7% | 88.1% | 25 | 0 | 0.0% | 100.0% | 80.6%29 / 36 |
| direct prompt injection / jailbreak (tune half) heldout-v2-tune · AMTSO vector 1 | 61 | 61 | 100.0% | 100.0% | 25 | 2 | 8.0% | 96.8% | 96.7%59 / 61 |
| benign developer traffic (false-positive corpus) benign-corpus-v2 | 0 | 0 | — | — | 610 | 23 | 3.8% | n/a | not measurable |
| benign fetched web content (hard negatives, tune half) benign-web-content-tune | 9 | 7 | 77.8% | 77.8% | 149 | 27 | 18.1% | 20.6% | not measurable |
| Overallall corpora combined | 286 | 212 | 74.1% | — | 875 | 64 | 7.3% | 76.8% | — |
The vocabulary comes from AMTSO's Guidelines for Testing of Agentic Security Products v1.0. Prevented means the attack was stopped, not merely noticed; detected, not prevented means it was flagged but would still have happened. Model refusal and model recognition are the columns that separate what the model did by itself from what the product did — there is no model in this harness, so they are measured-as-absent zeros rather than omitted fields.
| Corpus | Prevented | Hard-blocked | Detected only | Missed | Inconclusive | Not applicable | Model refusal | Model recognition |
|---|---|---|---|---|---|---|---|---|
| Indirect / content-mediated injectionvector2-indirect-content | 920.0% | 0 | 24 | 1226.7% | 0 | 0 | 0 | 0 |
| Tool, skill, extension and MCP supply chainvector3-supply-chain | 1216.7% | 3 | 45 | 1520.8% | 0 | 0 | 0 | 0 |
| Outbound action / agent-initiated effectvector4-outbound-action | 1831.6% | 9 | 0 | 3968.4% | 0 | 0 | 0 | 0 |
| Memory, context and cross-agent propagationvector5-memory-crossagent | 511.9% | 0 | 31 | 614.3% | 0 | 0 | 0 | 0 |
| direct prompt injection / jailbreak (tune half)heldout-v2-tune | 00.0% | 0 | 61 | 00.0% | 0 | 0 | 0 | 0 |
| benign developer traffic (false-positive corpus)benign-corpus-v2 | 00.0% | 0 | 0 | 00.0% | 0 | 0 | 0 | 0 |
| benign fetched web content (hard negatives, tune half)benign-web-content-tune | 00.0% | 0 | 7 | 222.2% | 0 | 0 | 0 | 0 |
| Overall286 conclusive attack rows | 4415.4% | 12 | 168 | 7425.9% | 0 | 0 | 0 | 0 |
Averages hide the shape of a result. Each corpus below opens to its per-sub-technique table, the harness and pipeline stage each sample was scored through, and the identifiers of every attack that was missed and every benign sample that was wrongly flagged. The sample ids are the ones in the benchmark repository, so any row here can be pulled up and re-run.
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| data-poisoning-fact | 4 | 1 | 1 | 25.0% | 2 | 1 | 50.0% |
| conversation-exfil | 5 | 3 | 3 | 60.0% | 1 | 0 | 0.0% |
| silent-suppression | 5 | 3 | 3 | 60.0% | 2 | 0 | 0.0% |
| credential-harvest | 4 | 3 | 3 | 75.0% | 2 | 1 | 50.0% |
| exfil-directive | 8 | 6 | 6 | 75.0% | 3 | 2 | 66.7% |
| persistence-directive | 5 | 4 | 4 | 80.0% | 2 | 1 | 50.0% |
| authority-spoof | 7 | 6 | 6 | 85.7% | 2 | 2 | 100.0% |
| tool-invocation-hijack | 7 | 7 | 7 | 100.0% | 3 | 1 | 33.3% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 45 | 33 | 73.3% | 17 | 8 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| index | 6 | 2 | 33.3% | 1 | 0 |
| prompt | 2 | 1 | 50.0% | 0 | 0 |
| file | 13 | 8 | 61.5% | 5 | 1 |
| output | 24 | 22 | 91.7% | 11 | 7 |
Right reason: 33.3% (11 of 33 catches), basis: measured
Missed attacks (12):
False positives (8):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| capability-expansion | 8 | 2 | 2 | 25.0% | 0 | 0 | — |
| dangerous-schema-default | 8 | 4 | 4 | 50.0% | 0 | 0 | — |
| delayed-behavior-change | 8 | 6 | 6 | 75.0% | 0 | 0 | — |
| malicious-tool-description | 12 | 10 | 10 | 83.3% | 0 | 0 | — |
| tool-name-shadowing | 8 | 7 | 7 | 87.5% | 0 | 0 | — |
| executable-config-entry | 6 | 6 | 6 | 100.0% | 0 | 0 | — |
| hidden-canary-in-metadata | 8 | 8 | 8 | 100.0% | 0 | 0 | — |
| poisoned-autoload-config | 14 | 14 | 14 | 100.0% | 0 | 0 | — |
| legit-config | 0 | 0 | 0 | — | 11 | 2 | 18.2% |
| legit-tool-registration | 0 | 0 | 0 | — | 7 | 0 | 0.0% |
| normal-schema-update | 0 | 0 | 0 | — | 3 | 0 | 0.0% |
| real-capability-addition | 0 | 0 | 0 | — | 3 | 0 | 0.0% |
| security-doc | 0 | 0 | 0 | — | 1 | 1 | 100.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 72 | 57 | 79.2% | 25 | 3 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| tool | 38 | 26 | 68.4% | 12 | 0 |
| index | 5 | 4 | 80.0% | 2 | 1 |
| file | 26 | 24 | 92.3% | 10 | 2 |
| output | 3 | 3 | 100.0% | 1 | 0 |
Right reason: 77.2% (44 of 57 catches), basis: measured
Missed attacks (15):
False positives (3):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| account-setting-change | 5 | 0 | 0 | 0.0% | 2 | 0 | 0.0% |
| exfil-paste-site | 4 | 0 | 0 | 0.0% | 1 | 0 | 0.0% |
| message-send | 5 | 0 | 0 | 0.0% | 2 | 0 | 0.0% |
| exfil-attacker-api | 5 | 1 | 1 | 20.0% | 2 | 0 | 0.0% |
| vcs-push-or-pr | 5 | 1 | 1 | 20.0% | 2 | 0 | 0.0% |
| purchase-transaction | 4 | 1 | 1 | 25.0% | 1 | 0 | 0.0% |
| stage-then-exfiltrate | 7 | 2 | 2 | 28.6% | 2 | 0 | 0.0% |
| exfil-webhook | 6 | 2 | 2 | 33.3% | 2 | 0 | 0.0% |
| db-out-of-scope-query | 4 | 2 | 2 | 50.0% | 2 | 0 | 0.0% |
| file-write-overwrite | 6 | 3 | 3 | 50.0% | 3 | 0 | 0.0% |
| shell-execution | 6 | 6 | 6 | 100.0% | 5 | 1 | 20.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| action | 57 | 18 | 31.6% | 24 | 1 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 57 | 18 | 31.6% | 24 | 1 |
Right reason: not measurable — no sample in this corpus declares an expected id
Missed attacks (39):
False positives (1):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| cross-agent-handoff-payload | 8 | 5 | 6 | 62.5% | 0 | 0 | — |
| memory-write-then-consume | 8 | 6 | 6 | 75.0% | 0 | 0 | — |
| delayed-activation | 6 | 5 | 5 | 83.3% | 0 | 0 | — |
| confused-deputy-pivot | 4 | 4 | 4 | 100.0% | 0 | 0 | — |
| orphan-agent-injection | 4 | 4 | 4 | 100.0% | 0 | 0 | — |
| shared-workspace-artifact | 8 | 8 | 8 | 100.0% | 0 | 0 | — |
| trace-gap-concealment | 4 | 4 | 4 | 100.0% | 0 | 0 | — |
| genuine-handoff | 0 | 0 | 0 | — | 6 | 0 | 0.0% |
| legit-shared-artifact | 0 | 0 | 0 | — | 6 | 0 | 0.0% |
| ordinary-memory-usage | 0 | 0 | 0 | — | 7 | 0 | 0.0% |
| ordinary-trace | 0 | 0 | 0 | — | 6 | 0 | 0.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| session | 2 | 1 | 50.0% | 2 | 0 |
| steps | 15 | 12 | 80.0% | 3 | 0 |
| text | 10 | 8 | 80.0% | 13 | 0 |
| events | 15 | 15 | 100.0% | 7 | 0 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 21 | 17 | 81.0% | 11 | 0 |
| file | 12 | 10 | 83.3% | 7 | 0 |
| index | 2 | 2 | 100.0% | 2 | 0 |
| output | 7 | 7 | 100.0% | 5 | 0 |
Right reason: 80.6% (29 of 36 catches), basis: measured
Missed attacks (6):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| AdvPrefix | 4 | 4 | 4 | 100.0% | 0 | 0 | — |
| AutoDAN | 10 | 10 | 10 | 100.0% | 0 | 0 | — |
| BoN | 8 | 8 | 8 | 100.0% | 0 | 0 | — |
| CipherChat | 8 | 8 | 8 | 100.0% | 0 | 0 | — |
| DAN | 8 | 8 | 8 | 100.0% | 0 | 0 | — |
| FlipAttack | 4 | 4 | 4 | 100.0% | 0 | 0 | — |
| h4rm3l | 5 | 5 | 5 | 100.0% | 0 | 0 | — |
| PAIR | 5 | 5 | 5 | 100.0% | 0 | 0 | — |
| PAP | 5 | 5 | 5 | 100.0% | 0 | 0 | — |
| TAP | 4 | 4 | 4 | 100.0% | 0 | 0 | — |
| benign | 0 | 0 | 0 | — | 25 | 2 | 8.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| session | 1 | 1 | 100.0% | 1 | 0 |
| text | 60 | 60 | 100.0% | 24 | 2 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 61 | 61 | 100.0% | 25 | 2 |
Right reason: 96.7% (59 of 61 catches), basis: measured
False positives (2):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| agent-file | 0 | 0 | 0 | — | 18 | 0 | 0.0% |
| agent-net | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| architecture | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| casual | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| code-review | 0 | 0 | 0 | — | 15 | 0 | 0.0% |
| coding | 0 | 0 | 0 | — | 26 | 0 | 0.0% |
| data-sql | 0 | 0 | 0 | — | 20 | 0 | 0.0% |
| debug | 0 | 0 | 0 | — | 25 | 0 | 0.0% |
| deps | 0 | 0 | 0 | — | 16 | 1 | 6.3% |
| devops | 0 | 0 | 0 | — | 26 | 1 | 3.8% |
| docs | 0 | 0 | 0 | — | 18 | 0 | 0.0% |
| git | 0 | 0 | 0 | — | 17 | 1 | 5.9% |
| hard-neg | 0 | 0 | 0 | — | 168 | 13 | 7.7% |
| logs | 0 | 0 | 0 | — | 14 | 0 | 0.0% |
| obf-altcaps | 0 | 0 | 0 | — | 4 | 0 | 0.0% |
| obf-encoded | 0 | 0 | 0 | — | 16 | 3 | 18.8% |
| obf-entropy | 0 | 0 | 0 | — | 8 | 2 | 25.0% |
| obf-gibberish | 0 | 0 | 0 | — | 6 | 0 | 0.0% |
| obf-leet | 0 | 0 | 0 | — | 10 | 0 | 0.0% |
| obf-minified | 0 | 0 | 0 | — | 12 | 0 | 0.0% |
| obf-mixedscript | 0 | 0 | 0 | — | 10 | 0 | 0.0% |
| obf-nonenglish | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| obf-spacing | 0 | 0 | 0 | — | 14 | 1 | 7.1% |
| obf-zerowidth | 0 | 0 | 0 | — | 5 | 0 | 0.0% |
| planning | 0 | 0 | 0 | — | 18 | 1 | 5.6% |
| refactor | 0 | 0 | 0 | — | 18 | 0 | 0.0% |
| security-legit | 0 | 0 | 0 | — | 25 | 0 | 0.0% |
| shell | 0 | 0 | 0 | — | 16 | 0 | 0.0% |
| testing | 0 | 0 | 0 | — | 21 | 0 | 0.0% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 0 | 0 | — | 610 | 23 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| prompt | 0 | 0 | — | 610 | 23 |
Right reason: not measurable — no sample in this corpus declares an expected id
False positives (23):
| Sub-technique | Attacks | Caught | Any-step | Recall | Benign | FP | FP rate |
|---|---|---|---|---|---|---|---|
| — | 9 | 7 | 7 | 77.8% | 149 | 27 | 18.1% |
| Scored through harness | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| text | 9 | 7 | 77.8% | 149 | 27 |
| Pipeline stage | Attacks | Caught | Recall | Benign | FP |
|---|---|---|---|---|---|
| output | 9 | 7 | 77.8% | 149 | 27 |
Right reason: not measurable — no sample in this corpus declares an expected id
Missed attacks (2):
False positives (27):
Reproduced verbatim from the result file. These are the places where what was measured differs from what a reader might assume was measured — published in full rather than footnoted, because a benchmark that hides its deviations is worth less than no benchmark.
An adapter is only scored on the surfaces it declares. A capability marked no means the corresponding samples were never put to it and were excluded as not-applicable — that is a narrower measurement, not a failure.
MOORAI_REPO=<checkout> node scorers/run.mjs --adapter moorai --corpus all --jsonThis page is generated from the result JSON vendored into the site from the benchmark repository at a pinned commit. No figure on it was typed by hand.
All 3 scored runs: MoorAI agent 0.79.9 (this page) · keyword reference adapter 1.0.0 (40.3% recall) · null adapter 1.0.0 (0.0% recall). AMTSO has not reviewed, certified or endorsed this benchmark or any result on it.