
As single-turn safety alignment improves, risk increasingly moves into the composition of knowledge and capabilities across turns. This talk introduces distributed hidden intent, presents CKA-Agent as feedback-driven tree search over a reasoning DAG, and explains how TurnGate identifies the first turn at which a conversation becomes harm-enabling.
The talk connects two sides of multi-turn agent safety: searching for harmful capability assembled from locally benign exchanges, and intervening at the earliest turn where that capability becomes actionable.
It covers The Trojan Knowledge, CKA-Agent, and TurnGate.