Simple Role Assignment is Extraordinarily Effective for Safety Alignment
Basic Information
- Zhou Ziheng*†, Jiakun Ding*, Zhaowei Zhang, Ruosen Gao, Yingnian Wu, Demetri Terzopoulos, Yipeng Kang, Fangwei Zhong, Junqi Wang†
- ACL
- 2026
Abstract
Principle-based alignment often lacks context sensitivity and completeness. Grounded in Theory of Mind, we propose role conditioning as an alternative: social roles (e.g., mother, judge) implicitly encode both values and the cognitive schemas required to apply them, enabling context-adaptive safety reasoning without exhaustive enumeration of principles. We introduce a training-free pipeline featuring a role conditioned generator and iterative role-based critics for refinement. Across five model families, our approach consistently outperforms principle-based, Chain-of-Thought (CoT) and other baselines across benchmarks. Notably, it reduces unsafe outputs on the WildJail break benchmark from 81.4% to 3.6% with DeepSeek-V3, while preserving general model capabilities on standard reasoning benchmarks. Beyond common safety benchmarks, it consistently applies to agentic safety tasks. These results establish role assignment as a powerful, interpretable paradigm for AI alignment and LLM-as-a-Judge construction.