SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

Abstract

Language-model agents increasingly use tools to act on external systems, where earlier actions can change the state that determines whether a later action is harmful. SEAD formulates attack and defense as partially observed state control. DART uses execution feedback to search for harmful tool trajectories, while SAGE investigates relevant state through read-only queries before allowing or blocking each action.

Publication
arXiv preprint

SEAD studies when apparently routine tool actions become harmful because of state changes accumulated across an agent trajectory.

Xinjie Shen 沈鑫杰
Xinjie Shen 沈鑫杰
PhD Student @ Georgia Tech

I study how to train capable agents and make them safe and reliable in open-ended environments.