Xinjie Shen
Xinjie Shen
Home
News & Talks
Writing
Research
Experience
CV
Contact
Light
Dark
Automatic
3
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
A reproducible 365-day benchmark for agents operating online stores under negotiation, delayed outcomes, market shocks, and persistent memory demands.
Wei Fan
,
Xinjie Shen 沈鑫杰
,
Xudong Guo
,
Jianhong Tu
,
Yang Su
,
Yinger Zhang
,
Lianghao Deng
,
Fengyu Wang
,
Baohua Dong
,
Yangqiu Song
,
Dayiheng Liu
PDF
Code
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
Introduce CKA-Agent, a framework that reformulates jailbreaking as an adaptive tree search over the target LLM’s correlated knowledge, achieving 96-99% attack success rates against state-of-the-art commercial LLMs.
Rongzhe Wei
,
Peizhi Niu
,
Xinjie Shen 沈鑫杰
,
Tong Tu
,
Yifan Li
,
Ruihan Wu
,
Eli Chien
,
Pin-Yu Chen
,
Olgica Milenkovic
,
Pan Li
PDF
Project Page
Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
A comprehensive evaluation benchmark for assessing privacy awareness of large language models in physical environments, revealing significant gaps when privacy is grounded in real-world contexts across four evaluation tiers.
Xinjie Shen 沈鑫杰
PDF
Code
Cite
×