Yuqian Fu

PhD Student · Institute of Automation, Chinese Academy of Sciences

I am a PhD student at the Institute of Automation, Chinese Academy of Sciences, advised by Prof. Dongbin Zhao and Co-advised by Prof. Yuanheng Zhu. I am selected for the PhD Honors Program (RMB 100,000 annual fellowship).

My research focuses on enabling agents to complete long-horizon, complex tasks in real-world production scenarios. I have been fortunate to work at frontier labs: Alibaba Qwen Team (core contributor to Qwen3.6-3.7) and ByteDance Seed Team (core contributor to Seed2.0), where I gained hands-on experience in full-cycle post-training of trillion-parameter models for challenging scenarios such as tool-augmented reasoning (challenging STEM problems with search and code tools, 250+ turns) and multi-turn, multi-day high-value workflows in Cowork harnesses (OpenClaw, Hermes).

I am particularly interested in training agents to generalize long-horizon capabilities from limited-horizon setups, ultimately delivering high-quality artifacts on complex professional workflows. My work combines training-algorithm design, expert-model distillation & merging, and user-reflux-trajectory analysis to drive model-pattern improvements.

On the Job Market (Expected Graduation: 2027): I am currently seeking full-time job opportunities in foundation model post-training. I am passionate about improving agent productivity in real-world tasks through co-design of data pipelines and training algorithms.

Feel free to email me at fuyuqian2022@ia.ac.cn or reach out for a coffee chat! ☕

Blog

An analysis of how frontier coding agents can optimize for grader-visible success, stop early, and deliver artifacts whose apparent completion exceeds their actual quality. [中文版]
Deep dive into the common failure modes in on-policy distillation for multi-agent reinforcement learning and practical solutions to address them. [中文版]
A comprehensive analysis on how Anthropic and OpenAI monitor agent behavior patterns during training to mitigate post-training issues like reward hacking and bad patterns. Synthesized from their technical blog posts and model cards, exploring trajectory monitoring and behavior mitigation strategies. [中文版]

Experience

Qwen Team, Alibaba Group
Jan. 2026 - Present
Research Intern - Foundation Model Post-training · Beijing, China
Long-horizon task capability improvement, such as HLE with tool use (250+ turns) and hour-scale Cowork tasks. Data pipeline: Built trajectory monitoring and environment-synthesis system for agent scenarios. Training algorithms: Designed reinforcement learning and merge algorithms to improve the model's agentic patterns and long-horizon task performance.
Seed Team, ByteDance
Aug. 2025 - Jan. 2026
Research Intern - LLM Agent Post-training · Beijing, China
Investigated practical failure modes of on-policy distillation (OPD) and proposed Top-K and mask-based fixes, improving multi-task OPD by ~20%, with nearly 1,000 views (blog). Analyzed RL training-inference mismatch and proposed MIS to mitigate the issue, later adopted by multiple delivery teams (blog).
AsX Team, Meituan
Mar. 2025 - Aug. 2025
Research Intern - LLM Reasoning Post-training · Beijing, China
Worked on LLM reasoning post-training algorithms. Proposed SRFT (Single-Stage Supervised and Reinforcement Fine-Tuning), effectively combining expert trajectories with on-policy samples, achieving 9% improvement over existing algorithms. Developed RLAE (RL-Assisted Ensemble for LLMs), where a 400M model orchestrates expert models through RL for ensemble reasoning.

Selected Publications

Yuqian Fu*, Haohuan Huang*, Kaiwen Jiang, Jiacai Liu, Zhuo Jiang, Yuanheng Zhu, Dongbin Zhao
Distillation Multi-Task Model Merging
COLM 2026 Core A* GitHub stars Blog Zhihu
Yuqian Fu, Tinghong Chen, Jiajun Chai, Xihuai Wang, Songjun Tu, Guojun Yin, Wei Lin, Qichao Zhang, Yuanheng Zhu, Dongbin Zhao
Reinforcement Learning Sample Efficiency Distillation
ICLR 2026 Core A* / CCF-A GitHub stars
Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Guojun Yin, Wei Lin, Qichao Zhang, Dongbin Zhao
Multi-Agent Model Merging
EMNLP 2025 Core A* / CCF-B
Yuqian Fu, Yuanheng Zhu, Jian Zhao, Jiajun Chai, Dongbin Zhao
Multi-Agent Data Synthesis Reinforcement Learning
ICLR 2025 Core A* / CCF-A GitHub stars
Yuqian Fu, Yuanheng Zhu, Haoran Li, Zijie Zhao, Jiajun Chai, Dongbin Zhao
Multi-Agent Exploration Reinforcement Learning
IEEE TCDS 2025 JCR Q1
Weiyu Ma*, Yuqian Fu*, Zecheng Zhang, Bernard Ghanem, Guohao Li
Multi-Agent Reinforcement Learning Vision-Language Model
ACL 2026 Core A* / CCF-A
Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Dongbin Zhao
Multi-Agent Reinforcement Learning Multi-Task
IEEE TSMC 2024 JCR Q1

Honors & Awards

Professional Service & Talks

Conference Reviewer

ICML, ICLR, NeurIPS, COLM, ARR, RLC, AAAI, IROS

Journal Reviewer

TMLR, IEEE TNNLS, IEEE TCDS, Neural Networks

Invited Talks