TurnGate detects the earliest turn at which a candidate response would make a multi-turn interaction sufficient to enable harmful action. MTID provides matched harmful and benign trajectories with annotations of the harm-enabling closure point, supporting precise intervention without premature refusal.
Hidden malicious intent can be distributed across individually benign-looking turns. TurnGate uses response-aware, turn-level supervision to identify when the conversation crosses the harm-enabling boundary.