04 — testing & measurement

Baseline validation

Confirming that a test case actually works with the product turned off before counting a block as a save. Without it, a test suite quietly fills up with attacks that never worked anyway, and the product gets credit for stopping nothing. It is the single most common omission in security-product testing and the reason a good result should always be reported alongside its baseline.

Related terms

  • Model refusal Testing & measurement

    The underlying model declining the request on its own, with no security product involved.

  • Marginal value Testing & measurement

    The protection a control adds on top of what the model already refuses — the only part of a detection number that is genuinely the product's.

  • AMTSO Frameworks & standards

    The industry body that sets standards for how security products are tested fairly. Its Guidelines for Testing of Agentic Security Products v1.0 (2…

Testing & measurement

The words that decide whether a published detection number is evidence or decoration. Most vendor numbers in this space are quoted without any of them; see how we test for how these are applied in practice.

Baseline validation is term 4 of 15 in this part of the glossary.

Get started free Full glossary → How we test → Community agent on GitHub ↗

This page is one entry from the agentic AI security glossary, which defines 60 terms in the same style. Where a term belongs to a published taxonomy or standard — the HackAgent attack families, the OWASP lists, AMTSO's guidelines, MITRE ATLAS, ISO/IEC 42001, the NIST AI RMF, the EU AI Act — the authoritative wording is the source document's, not ours, and specifics should be verified there. "Lethal trifecta" is Simon Willison's term. OWASP® is a trademark of the OWASP Foundation; ATT&CK® and ATLAS™ are trademarks of The MITRE Corporation. Naming a framework here is description, not a claim of certification or endorsement.

glick.run — AGPL-3.0