About Me
I joined Computer Science & Engineering (CSE) Department at Washington University in St. Louis (WashU) in 2024 as an Assistant Professor. I received my Ph.D. degree in Computer Science Department, UIUC, advised by Prof. Jiawei Han. After that, I visited UW as a researcher and worked with Prof. Hanna Hajishirzi. Prior to UIUC, I received my Bachelor Degree in Electronic Engineering in Tsinghua University in 2018. My research interest broadly lies in the intersection of natural language processing and machine learning, and I am especially interested in understanding the properties of language models as well as improving their trustworthiness and efficiency.
📢 I am looking for PhD students and interns! If you are interested in working with me, please fill in this form. Check out this page for more details!
News
- [Aug 2026] Two papers accepted to EMNLP 2026: RelayLLM and Nonsense Helps (LoPE).
- [Jun 2026] SPEED accepted to COLM 2026. See you in San Francisco!
- [May 2026] Received the NSF CAREER Award!
- [May 2026] Received the WashU AI for Health Grant!
- [May 2026] Three papers accepted to ICML 2026: Training Data Efficiency in Multimodal Process Reward Models, Parallel-Probe, and Rethinking the Reranker.
- [Feb 2026] VisPlay accepted to CVPR 2026.
- [Jan 2026] Selected for the AAAI New Faculty Highlight Program 2026.
- [Jan 2026] Two papers accepted to ICLR 2026: R-Zero and Self-Calibration.
Recent Research Interests
My group studies how to make language models self-improving, efficient, and trustworthy. Selected recent work:
Self-Evolving Models: Learning without Human-Curated Data
Training models that generate their own tasks, data, and environments.
- ICLR 2026 R-Zero: Self-Evolving Reasoning LLM from Zero Data
- CVPR 2026 VisPlay: Self-Evolving Vision-Language Models from Images
- Preprint 2026 G-Zero: Self-Play for Open-Ended Generation from Zero Data
- Preprint 2026 EnvHarness: Awakening Static Worlds for Agent Learning
- EMNLP 2023 Large Language Models Can Self-Improve
Reinforcement Learning and Exploration for Reasoning
How reasoning improves under RL, and how to keep exploration broad.
- Preprint 2026 You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
- EMNLP 2026 Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration
- COLM 2025 CrossWordBench: Evaluating the Reasoning Capabilities of LLMs with Controllable Puzzle Generation
Efficient Reasoning and Inference
Reducing the cost of long reasoning chains at training and inference time.
- EMNLP 2026 RelayLLM: Efficient Reasoning via Collaborative Decoding
- COLM 2026 SPEED: Position Specialist Generates Better Draft for Speculative Decoding
- ICML 2026 Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing
- EACL 2026 Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning
Reward Modeling and Calibration
Making reward signals and model confidence reliable enough to train on.
- ICLR 2026 Efficient Test-Time Scaling via Self-Calibration
- ICML 2026 Training Data Efficiency in Multimodal Process Reward Models
- ICLR 2025 Taming Overconfidence in LLMs: Reward Calibration in RLHF
- Preprint 2026 Process Rewards with Learned Reliability
Honors and Awards
NSF CAREER Award 2026
AAAI New Faculty Highlight Program 2026
Microsoft Research PhD Fellowship 2021-2023
C.W. Gear Outstanding Graduate Award
Chirag Foundation Graduate Fellowship
Outstanding Graduates, Tsinghua University 2018
Academic Excellence Scholarship, Tsinghua University 2015-2017
China National Scholarship (Top 1%) 2016
Samsung Scholarship 2015
