Portrait of Xiaohan Yan

Xiaohan Yan   颜yán小xiǎo涵hán

Algorithm engineer at AgiBot

I build vision-language-action models for robot manipulation at AgiBot. The part I work on is post-training with reinforcement learning: residual policies that take human corrections, one policy that keeps earlier tasks while it learns new ones, and action-level evaluation on real robots.

Before AgiBot I was at NIO, and a research intern at AIR, Tsinghua University. My master's at Tongji, with Gang Wei, was 3D scene understanding: zero-shot instance segmentation, semantic Gaussian splatting, and a lightweight 3D-LLM. The thread through all of that is still multimodal models, 3D vision, and reinforcement learning.

Education & Experience

Experience
Algorithm Engineer / Large Model
2026 – Present
Algorithm Engineer / Large Model
2024 – 2025
Research Intern
Dec 2023 – Apr 2024
Education
M.S. in Computer Science
2022 – 2025
B.S. in Computer Science
Hohai University · ACM Team Captain
2018 – 2022

News

Sep 2026 Our paper Bee is released on arXiv.
Aug 2026 Our paper CIDER is released on arXiv.
Jul 2026 Our paper VINE is released on arXiv.
Feb 2026 Our paper ALOE is released on arXiv.

Publications

* indicates equal contribution. † indicates corresponding author.

2026
Bee method overview RoboticsRL
Bee: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models
arXiv 2026
Authors
Weihui Zhao, Xiaohan Yan†, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao
Overview

A residual policy on a frozen VLA. Human corrections set a per-dimension constraint, with tighter bounds where people agree and looser ones where they don't.

CIDER continual learning overview RoboticsRL
CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning
arXiv 2026
Authors
Houlin Li, Minghui Xu, Guo Xu, Xuan Du, Xiaohan Yan, Chun Wang, Yuxiang Yan, Shukai Yang, Yongcheng Liu, Wei Shan, Maoqing Yao
Overview

One policy learns six real-robot tasks in sequence. Each new task takes 10–20 minutes, and the earlier ones stay.

VINE sampling overview RoboticsRL
VINE: Taming Generative Control Policies for Reinforcement Learning
arXiv 2026
Authors
Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao
Overview

Rebuilds the interpolation state at each denoising step, so a flow-matching policy can take value gradients without the training blowing up.

ALOE framework overview RoboticsRL
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
arXiv 2026
Authors
Rushuai Yang*, Hecheng Wang*, Zhichao Wu*, Chiming Liu*, Xiaohan Yan, Xuan Du, Shuoyu Yue, Chuheng Zhang, Yunlong Wang, Yongcheng Liu, Lizhe Qi, Yi Chen, Wei Shan, Maoqing Yao
Overview

Scores the current policy on action chunks, then uses that to post-train a VLA on real robots.

HOLO pipeline overview 3D Scene UnderstandingVLM
HOLO: Holistic Lightweight Optimization for Scene Understanding with Auto-Annotation and Multimodal Learning
WACV 2026
Authors
Xiaoyun Hu*, Xiaohan Yan*, Nan Wang, Xiaowei Song, Gang Wei, Zhicheng Wang
Overview

A large-scale scene description dataset together with a lightweight 3D-LLM.

Projects

Pokémon Trading Card Game cards Pokémon TCGAI
Pokémon TCG AI Battle
Kaggle 2026
Team
Xiaohan Yan · 宝糕手 (Poké Messters)
Overview

A 60-card deck and a decision agent for the Pokémon Trading Card Game AI Battle. Team 宝糕手 (Poké Messters): 108th of 6,807, score 1011.9, Silver Medal.

LLM Science Exam competition Science QALLM
LLM Science Exam — Use LLMs to Answer Difficult Science Questions
Kaggle 2023
Authors
Xiaohan Yan, Nan Wang, Xiaowei Song, Jinyu He
Overview

Fine-tuning large language models on private datasets. Score 0.905, top 3% worldwide, Silver Medal.

eScape game screenshot GeometryGame
eScape — A Geometry Storm Game
GameJam 2023
Authors
Origami-hui, Xiaohan Yan
Overview

Scale your device and escape from this geometry storm. 1st in innovation and 2nd in theme interpretation.

Awards

Most of the awards I won during my student years 2018 – 2024
The 2019 ICPC Asia-East Continent Final — Bronze Medal 2019 – 2020
Jiangsu Collegiate Programming Contest — Silver Medal, 2nd place 2019 – 2020

Misc