I am a Ph.D. student in Machine Learning at Georgia Tech, advised by Prof. Pan Li. My research connects advances in long-horizon agent post-training with safety and privacy. I study how increasingly capable agents can learn from real feedback, infer hidden intent and social context, and act reliably in open-ended real-world environments.
At Qwen, I work on long-horizon agentic post-training and verifiable training environments. My recent work includes VHD-Play, E-Commerce Bench, and research on agent safety and physical-world privacy.
Before Georgia Tech, I worked at Amazon Web Services on efficient sparse retrieval and at Microsoft Research Asia on autonomous R&D agents. I enjoy turning research ideas into systems, benchmarks, and visual explanations that other people can use.
PhD in Machine Learning, 2025 - 2029
Georgia Institute of Technology
BS in Artificial Intelligence, 2021 - 2025
South China University of Technology
Recent research, releases, and invited talks
Notes on research, agents, and how the work gets made
Agent post-training, safety, privacy, and retrieval

Generate diverse, stateful agentic RL environments with verifiable rewards by deriving both dynamics and evaluation from solved mechanisms.

Introduce CKA-Agent, a framework that reformulates jailbreaking as an adaptive tree search over the target LLM’s correlated knowledge, achieving 96-99% attack success rates against state-of-the-art commercial LLMs.

A comprehensive evaluation benchmark for assessing privacy awareness of large language models in physical environments, revealing significant gaps when privacy is grounded in real-world contexts across four evaluation tiers.

With increasing demands for efficiency, information retrieval has developed a branch of sparse retrieval, further advancing towards inference-free retrieval where the documents are encoded during indexing time and there is no model-inference for queries. Existing sparse retrieval models rely on FLOPS regularization for sparsification, while this mechanism was originally designed for Siamese encoders, it is considered to be suboptimal in inference-free scenarios which is asymmetric. Previous attempts to adapt FLOPS for inference-free scenarios have been limited to rule-based methods, leaving the potential of sparsification approaches for inference-free retrieval models largely unexplored. In this paper, we explore ℓ0 inspired sparsification manner for inference-free retrievers. Through comprehensive out-of-domain evaluation on the BEIR benchmark, our method achieves state-of-the-art performance among inference-free sparse retrieval models and is comparable to leading Siamese sparse retrieval models. Furthermore, we provide insights into the trade-off between retrieval effectiveness and computational efficiency, demonstrating practical value for real-world applications.
Research conversations and collaborations are welcome
Email is the best way to reach me.