By giving language model agents a dedicated way to express their intentions during reasoning, you can detect harmful goal shifts in real-time rather than only catching problems after they occur.
This paper introduces INTENT-AS-A-TOOL, a method to detect when AI agents might take harmful actions by monitoring their reasoning process. Instead of only checking final decisions, the approach adds special tools that let models explicitly signal their intentions during thinking, creating a detailed record of how the agent's goals shift—helping catch misalignment before bad actions happen.