Distributed Hidden Malice: Attacks and Defenses

Abstract

As single-turn safety alignment improves, risk increasingly moves into the composition of knowledge and capabilities across turns. This talk introduces distributed hidden intent, presents CKA-Agent as feedback-driven tree search over a reasoning DAG, and explains how TurnGate identifies the first turn at which a conversation becomes harm-enabling.

Date
Aug 19, 2026 8:00 PM
Location
Online

The talk connects two sides of multi-turn agent safety: searching for harmful capability assembled from locally benign exchanges, and intervening at the earliest turn where that capability becomes actionable.

It covers The Trojan Knowledge, CKA-Agent, and TurnGate.

Xinjie Shen 沈鑫杰
Xinjie Shen 沈鑫杰
PhD Student @ Georgia Tech

I study how to train capable agents and make them safe and reliable in open-ended environments.