The Superficial Alignment Hypothesis echoes findings that current safety methods only superficially alter model behavior.
Prior research highlights that alignment effects are most pronounced in the initial tokens of a sequence.
Protecting early tokens is a recurring theme in recent defenses, such as amplifying initial safety disclaimers.