Hello, I'm

Haotian Xu 许皓天

RL Algorithms & RL-Infra · Agentic-RL

Apodex

2,000+ Citations
16 h-index
22 i10-index
Scroll

About Me

I work at Apodex on RL algorithms and Agentic-RL, backed by the RL-Infra I build. I develop the miles-based RL training framework behind apodex-v1.0 and v1.1 — a fully-asynchronous stack that decouples training from rollout, evaluates without pausing training, and reclaims idle GPUs for rollout via hybrid-colocate — plus miles-values, a dedicated SAO (Async Single-rollout PPO) framework. Earlier I built async_openrlhf, a fully-async RLHF pipeline extending OpenRLHF. My research spans LLM reasoning, MCTS, agentic scaling laws, and reasoning compression.

Research Interests

RL Algorithms Agentic-RL LLM Reasoning MCTS RLHF RL-Infra Scaling Laws NLP LLM Safety

Work Experience

2025.11 – Present

Apodex (formerly MiroMind)

RL Algorithms & RL-Infra (Agentic-RL)

2024 – 2025.8

Xiaohongshu (小红书) — HiLab

Math-Reasoning & Agentic-RL

2023

Douyin (抖音 / ByteDance)

AIGC Safety

2018 – 2023

Alibaba (阿里巴巴)

Content Safety, Public Opinion Analysis, Personal Privacy, LLM Safety

Education

2013 – 2016

Tsinghua University (清华大学)

Department of Electronic Engineering

2009 – 2013

UESTC (电子科技大学)

Bachelor's Degree

Featured Work · RL Algorithms & Infra

apodex-v1.0 & v1.1 · RL Training System

apodex-v1.0 and apodex-v1.1, Apodex's flagship agentic models, are trained with a fully-asynchronous, high-throughput Agentic-RL framework I develop on top of miles. On the same stack I also built miles-values, a dedicated SAO (Async Single-rollout PPO) framework.

Fully-Async Train / Inference Separation

Training workers and rollout engines run on independent loops, communicating only through a shared data buffer and output queue, so neither ever blocks the other. Off-policy staleness is bounded by two-level (cross-group + per-slot) filtering and a configurable staleness threshold, with cooperative-abort weight sync over NCCL between iterations.

Async Evaluation — No Training Stalls

A dedicated eval pool hot-reloads the rolling training checkpoints and evaluates in the background, so evaluation never interrupts the training loop (“training-data roll”). Eval failures are isolated and can never abort training, keeping throughput steady while metrics stream continuously.

Hybrid-Colocate GPU Reclamation

When the buffer has no consumable data, colocated training GPUs onload as generation engines and join the rollout routing pool — then offload again for the optimizer step. Idle training cards become extra rollout throughput, driving near-full cluster utilization with dedicated remote engines handling steady-state generation.

hl-gauss-sao-agentic-task

Eval accuracy over RL steps — HL-Gauss value loss + SAO on an agentic task. Raw (light) vs. EMA-smoothed (bold); dataset names withheld.

Benchmark A
0.6 0.62 0.64 0.66 25 50 75 100 125 150 175 RL step
Benchmark B
0.38 0.4 0.42 0.44 25 50 75 100 125 150 175 RL step
Benchmark C
0.35 0.4 0.45 0.5 0.55 25 50 75 100 125 150 175 RL step

2025 Highlights

Top-venue publications accepted in 2025–2026

arXiv 2026
13 citations

MiroThinker: Towards Heavy-Duty Research Agents via Verification & Model-Context Scaling

MiroMind Team (incl. H Xu)

A family of open research agents (MiroThinker-1.7 & h1) that scale long-horizon, verification-centric reasoning, trained with our fully-asynchronous RL stack.

arXiv 2026
1 citation

Argus: Evidence Assembly for Scalable Deep Research Agents

Z Zhang, L Su, Z Chen, X Lin, H Xu, et al.

An evidence-assembly framework that lets deep-research agents accumulate and verify evidence at scale for reliable long-horizon research.

NeurIPS 2025
71 citations

Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving

X Mai, H Xu, W Wang, Y Zhang, W Zhang

Discovering scaling laws for agentic reinforcement learning where LLMs spontaneously learn to execute code for mathematical reasoning.

ICLR 2026
44 citations

IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling

G Chen, Z Qiao, X Chen, D Yu, H Xu, WX Zhao, R Song, W Yin, H Yin, et al.

Rethinking how to scale agent interactions for long-horizon research tasks.

ICLR 2026
17 citations

GEM: A Gym for Generalist LLMs

Z Liu, A Sims, K Duan, C Chen, S Yu, X Zhou, H Xu, S Xiong, B Liu, C Tan, et al.

A comprehensive gym environment for training and evaluating generalist large language models.

ICLR 2026
25 citations

Search Self-Play: Pushing the Frontier of Agent Capability without Supervision

H Lu, Y Wen, P Cheng, R Ding, J Guo, H Xu, C Wang, H Chen, X Jiang, et al.

A novel self-play framework that advances agent capabilities through search without external supervision.

ICLR 2026
9 citations

Omni-DPO: A Dual-Perspective Paradigm for Dynamic Preference Learning of LLMs

S Peng, W Wang, Z Tian, S Yang, X Wu, H Xu, C Zhang, T Isobe, B Hu, M Zhang

A dual-perspective dynamic preference learning framework that adaptively reweights samples by jointly considering data quality and model performance evolution during training.

IEEE TPAMI 2025
529 citations

From System 1 to System 2: A Survey of Reasoning Large Language Models

D Zhang, ZZ Li, ML Zhang, J Zhang, Z Liu, Y Yao, H Xu, J Zheng, X Chen, et al.

A comprehensive survey of reasoning LLMs, tracing the evolution from fast intuitive (System 1) to slow deliberate (System 2) reasoning paradigms.

EMNLP 2025
371 citations

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

J Hu, X Wu, W Shen, JK Liu, W Wang, S Jiang, H Wang, H Chen, B Chen, H Xu, et al.

A Ray-based open-source framework enabling scalable RLHF training, widely adopted by the research community.

ACL 2024
14 citations

ITD: Large Language Models Can Teach Themselves Induction Through Deduction

W Sun, H Xu, X Yu, P Chen, S He, J Zhao, K Liu

Showing that LLMs can bootstrap inductive reasoning capability through deductive self-teaching.

ACL 2026
10 citations

TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression

ZZ Li, X Liang, Z Tang, L Ji, P Wang, H Xu, H Huang, W Deng, Y Gong, et al.

An efficient method to compress long chain-of-thought reasoning while preserving accuracy.

Selected Publications

Full list on Google Scholar

2026

MiroThinker: Towards Heavy-Duty Research Agents via Verification & Model-Context Scaling

MiroMind Team (incl. H Xu)

arXiv preprint 13 citations
2026

Argus: Evidence Assembly for Scalable Deep Research Agents

Z Zhang, L Su, Z Chen, X Lin, H Xu, et al.

arXiv preprint 1 citation
2026

VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning

Z Gao, S Shen, T Chai, W Wang, H Xu, et al.

ECCV 2026 1 citation
2025

Beyond Single-Task: Robust Multi-Task Length Generalization for LLMs

Y Hu, S Kang, H Yang, H Xu, M Zhang

NeurIPS 2025 11 citations
2025

MedKGent: A Large Language Model Agent Framework for Constructing Medical Knowledge Graphs

D Zhang, Z Wang, ZZ Li, Y Yu, H Xu, et al.

arXiv preprint 10 citations
2026

LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs

J Jia, X Wu, C Gao, Z Chen, H Xu, et al.

AAAI 2026 1 citation
2026

Shuttle Between the Instructions and the Parameters of Large Language Models

W Sun, H Xu, H Liao, X Yu, Z Jiang, et al.

ACL 2026
2025

From System 1 to System 2: A Survey of Reasoning Large Language Models

D Zhang, ZZ Li, ML Zhang, J Zhang, Z Liu, Y Yao, H Xu, et al.

IEEE Transactions on Pattern Analysis and Machine Intelligence 529 citations
2025

Reinforce++: Stabilizing Critic-free Policy Optimization with Global Advantage Normalization

J Hu, JK Liu, H Xu, W Shen

arXiv preprint 484 citations
2025

OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework

J Hu, X Wu, W Shen, JK Liu, W Wang, S Jiang, H Wang, H Chen, B Chen, H Xu, et al.

EMNLP 2025 371 citations
2023

CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

G Xu, J Liu, M Yan, H Xu, J Si, Z Zhou, P Yi, X Gao, J Sang, R Zhang, et al.

arXiv preprint 104 citations
2025

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

H Xu, X Wu, W Wang, Z Li, D Zheng, B Chen, Y Hu, S Kang, J Ji, Y Zhang, et al.

arXiv preprint 72 citations
2024

Interpretable Contrastive Monte Carlo Tree Search Reasoning

Z Gao, B Niu, X He, H Xu, H Liu, A Liu, X Hu, L Wen

arXiv preprint 59 citations
2025

Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving

X Mai, H Xu, W Wang, Y Zhang, W Zhang

NeurIPS 2025 71 citations
2025

Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

W Sun, X Yu, Z Huang, H Xu, S He, J Zhao, K Liu

CCL 2025 37 citations
2023

No Train Still Gain. Unleash Mathematical Reasoning of LLMs with Monte Carlo Tree Search Guided by Energy Function

H Xu

arXiv preprint 26 citations
2022

Generating Disentangled Arguments with Prompts: A Simple Event Extraction Framework that Works

J Si, X Peng, C Li, H Xu, J Li

ICASSP 2022 22 citations
2026

IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling

G Chen, Z Qiao, X Chen, D Yu, H Xu, WX Zhao, R Song, W Yin, H Yin, et al.

ICLR 2026 44 citations
2024

ITD: Large Language Models Can Teach Themselves Induction Through Deduction

W Sun, H Xu, X Yu, P Chen, S He, J Zhao, K Liu

ACL 2024 14 citations
2026

GEM: A Gym for Generalist LLMs

Z Liu, A Sims, K Duan, C Chen, S Yu, X Zhou, H Xu, S Xiong, B Liu, C Tan, et al.

ICLR 2026 17 citations
2026

Search Self-Play: Pushing the Frontier of Agent Capability without Supervision

H Lu, Y Wen, P Cheng, R Ding, J Guo, H Xu, C Wang, H Chen, X Jiang, et al.

ICLR 2026 25 citations
2026

Omni-DPO: A Dual-Perspective Paradigm for Dynamic Preference Learning of LLMs

S Peng, W Wang, Z Tian, S Yang, X Wu, H Xu, C Zhang, T Isobe, B Hu, M Zhang

ICLR 2026 9 citations
2025

Probabilistic Uncertain Reward Model

W Sun, X Cheng, X Yu, H Xu, Z Yang, S He, J Zhao, K Liu

arXiv preprint 15 citations
2026

TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression

ZZ Li, X Liang, Z Tang, L Ji, P Wang, H Xu, et al.

ACL 2026 10 citations
2020

End-to-end Latent-variable Task-oriented Dialogue System with Exact Log-likelihood Optimization

H Xu, H Peng, H Xie, E Cambria, L Zhou, W Zheng

World Wide Web 23(3) 45 citations

Zhihu Contributions

Technical writing on LLM reasoning, MCTS, and reinforcement learning

haotian

Research interests: Probabilistic Graphical Models, Reinforcement Learning, NLP

9,251 Upvotes 3,813 Likes 17,292 Collects 12+ Articles
Visit my Zhihu Profile

Get in Touch