04 — testing & measurement

Variance

Also written repeated-run distribution

How much a result moves when the same test is run again. Any figure involving a model is a sample from a distribution, not a constant — the same judge model can recover eight attacks in one run and five in the next — so a single-run number is a point estimate quoted as a fact. AMTSO asks for the distribution of outcomes rather than hiding variability behind one pass or fail, which in practice means running N times and publishing the spread.

See also BoN · AMTSO · how we test

Related terms

  • BoN Attack families

    Not a clever prompt but a volume attack: generate n shuffled, re-cased, lightly-perturbed variants of the same request, fire them all, and keep whichever…

  • AMTSO Frameworks & standards

    The industry body that sets standards for how security products are tested fairly. Its Guidelines for Testing of Agentic Security Products v1.0 (2…

Testing & measurement

The words that decide whether a published detection number is evidence or decoration. Most vendor numbers in this space are quoted without any of them; see how we test for how these are applied in practice.

Variance is term 15 of 15 in this part of the glossary.

Get started free Full glossary → How we test → Community agent on GitHub ↗

This page is one entry from the agentic AI security glossary, which defines 60 terms in the same style. Where a term belongs to a published taxonomy or standard — the HackAgent attack families, the OWASP lists, AMTSO's guidelines, MITRE ATLAS, ISO/IEC 42001, the NIST AI RMF, the EU AI Act — the authoritative wording is the source document's, not ours, and specifics should be verified there. "Lethal trifecta" is Simon Willison's term. OWASP® is a trademark of the OWASP Foundation; ATT&CK® and ATLAS™ are trademarks of The MITRE Corporation. Naming a framework here is description, not a claim of certification or endorsement.

glick.run — AGPL-3.0