Overview
HyperGuard studies safety signals that emerge inside diffusion-based language models while they iteratively denoise a response. Instead of judging only the final text, the system captures intermediate hidden states, treats them as a trajectory, and learns the boundary between safe and unsafe behaviour.