Language-model agents increasingly use tools to act on external systems, where earlier actions can change the state that determines whether a later action is harmful. SEAD formulates attack and defense as partially observed state control. DART uses execution feedback to search for harmful tool trajectories, while SAGE investigates relevant state through read-only queries before allowing or blocking each action.
SEAD studies when apparently routine tool actions become harmful because of state changes accumulated across an agent trajectory.