FlipAttack
A reversal attack: the payload is written backwards — by character, by word, or by line — and the model is instructed to un-reverse it before following it. It is the cheapest obfuscation there is (no key, no cipher table) and it defeats naive keyword matching completely, because not one dangerous word appears in the text as written.
See also CipherChat · obfuscation / encoding
Related terms
-
CipherChat
Attack families
The harmful request is encoded in a cipher — Caesar shift, ROT13, base64, hex, a custom substitution — and the model is asked to work in that cipher.
-
Obfuscation / encoding
Attack concepts
Rewriting a payload so it survives the defence but is still recoverable by the model — base64, hex, ROT13 and Caesar shifts, character or word reversal…
Attack families
The ten families in the HackAgent attack taxonomy (AISecurityLab/hackagent ↗), which is the taxonomy MoorAI's red-team corpora are keyed to. They are not ten unrelated tricks — they cluster into two groups that behave very differently. Obfuscation families hide the payload so the model never recognises it as harmful. Persuasion families state the harmful request plainly and argue the model into it. That split matters, because a model that refuses persuasion outright will happily comply with an encoding it cannot decode — see marginal value.