PhD Student @ USTC · Research Intern @ BIGAI
Reward Modeling · LLM Alignment · Continual Learning · Agent Memory
I work on building language models that keep improving from experience through reliable reward signals, robust alignment, continual adaptation, and memory.
I am currently a PhD student at the University of Science and Technology of China and a research intern at the Beijing Institute for General Artificial Intelligence (BIGAI). I am open to research collaboration.
- SAVE — On-policy feedback for reward model self-supervised improvement. [Paper] [Project]
- UAPO — Uncertainty-aware preference optimization. EMNLP 2025. [Paper] [Data]
- PoliCon — Evaluating LLMs on diverse political consensus objectives. ICLR 2026. [Paper] [Code]
- ICE — Learning knowledge from self-induced contextual distributions. ICLR 2025. [Paper] [Code]
- RAM — An ever-improving memory system that learns from communication. [Paper] [Code]
- RSIHub — A research framework for verifiable agent self-improvement.
- Belief Context Graph — Confidence-aware context for long-horizon agents.
- Simple Long-Horizon Agent — A lightweight agent for learning, experimentation, and long-horizon work.

