04 — testing & measurement

Held-out set

Test samples deliberately kept away from whoever is building the detection, so the score answers "does this generalise?" rather than "does this remember?". A held-out set is single-use in spirit: the moment someone tunes against it, even indirectly by reading which samples failed, it stops being held out and becomes training data. Numbers quoted without saying whether the set was held out are close to uninformative.

Related terms

  • Tune/test split Testing & measurement

    Dividing a corpus once into a half the tuning work may see and a locked half it may not.

  • Overfitting Testing & measurement

    Fitting the examples instead of the phenomenon. In detection work it usually looks like a rule matched to the exact phrasings in the corpus — a pattern…

Testing & measurement

The words that decide whether a published detection number is evidence or decoration. Most vendor numbers in this space are quoted without any of them; see how we test for how these are applied in practice.

Held-out set is term 1 of 15 in this part of the glossary.

Get started free Full glossary → How we test → Community agent on GitHub ↗

This page is one entry from the agentic AI security glossary, which defines 60 terms in the same style. Where a term belongs to a published taxonomy or standard — the HackAgent attack families, the OWASP lists, AMTSO's guidelines, MITRE ATLAS, ISO/IEC 42001, the NIST AI RMF, the EU AI Act — the authoritative wording is the source document's, not ours, and specifics should be verified there. "Lethal trifecta" is Simon Willison's term. OWASP® is a trademark of the OWASP Foundation; ATT&CK® and ATLAS™ are trademarks of The MITRE Corporation. Naming a framework here is description, not a claim of certification or endorsement.

glick.run — AGPL-3.0