Heng Zhou

NeoteAI
Portrait of Heng Zhou

I am a research intern at NeoteAI, working on the N0 series of tactile-centric foundation models for robot manipulation.

Before that I spent a year and a quarter at the Shanghai Artificial Intelligence Laboratory with Lei Bai, working on multi-agent systems, agentic reinforcement learning, and embodied spatial intelligence.

Embodied Foundation Models ยท World Models ยท Agents

Education

Beijing University of Posts and Telecommunications
School of Internet of Things Engineering
B.Eng. โ€” GPA 3.74/4.0, Rank 1/183
Sep. 2021 - Jun. 2025

Experience

NeoteAI
Research Intern
Jan. 2026 - Aug. 2026
Shanghai Artificial Intelligence Laboratory
Research Intern, with Lei Bai
Oct. 2024 - Jan. 2026
Westlake University
CAIRI Lab
Research Intern
Apr. 2024 - Aug. 2024
Tsinghua University
THU-LYJ Lab
Research Intern
Jul. 2023 - Dec. 2023

Honours & Awards

SAC Highlight Award (top 1%), EMNLP 2025
2025
Outstanding Graduate of Beijing Municipality
2025
Beijing Merit Student
2025
National Scholarship (top 0.2% nationwide)
2023
First Prize, 14th National College Student Mathematics Competition
2023
Honorable Mention, Mathematical Contest in Modeling (MCM/ICM)
2023

News

2026
Released the N0 series of tactile-centric foundation models โ€” N0-VTLA and N0-TWAM ๐Ÿฅณ
Jul
Scaling Behaviors of LLM RL Post-Training was accepted to ACL 2026 as an Oral ๐ŸŽ‰
Apr
LFQA-E was accepted to ICLR 2026 ๐ŸŽ‰
Jan
Started a research internship at NeoteAI, working on tactile foundation models for manipulation
Jan
2025
VIKI-R was accepted to NeurIPS 2025 ๐ŸŽ‰
Sep
ReSo was accepted to EMNLP 2025 as an Oral and received the SAC Highlight Award (top 1%) ๐ŸŽ‰
Aug
Graduated from Beijing University of Posts and Telecommunications, ranked 1st of 183, and was named an Outstanding Graduate of Beijing Municipality ๐ŸŽ“
Jun
2024
SS3DM was accepted to the NeurIPS 2024 Datasets and Benchmarks Track ๐ŸŽ‰
Sep

Selected Publications view all →

Building agents that reason in language, coordinate as teams, and act in the physical world.

Agents that Reason

Search, self-organize, and learn to think. 3

LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge

Heng Zhou, Ao Yu, Yuchen Fan, Jianing Shi, Li Kang, Hejia Geng, Yongting Zhang, Yutao Fan, Yuhao Wu, Tiancheng He, Yiran Qin, Lei Bai#, Zhenfei Yin# (# corresponding author)

preprint

An automatically constructed benchmark that mines deltas between successive Wikidata snapshots into SPARQL-verified questions at three reasoning depths, exposing how sharply models fail on post-pretraining facts.

SSRL: Self-Search Reinforcement Learning

Yuchen Fan*, Kaiyan Zhang*, Heng Zhou*, Yuxin Zuo, Yanxu Chen, Yu Fu, Xinwei Long, Xuekai Zhu, Che Jiang, Yuchen Zhang, Li Kang, Gang Chen, Cheng Huang, Zhizhou He, Bingning Wang, Lei Bai#, Ning Ding#, Bowen Zhou# (* equal contribution, # corresponding author)

preprint

A reinforcement learning method that trains language models to simulate search internally through format- and rule-based rewards, cutting reliance on external search engines while still transferring to real retrieval.

ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks

Heng Zhou*, Hejia Geng*, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin#, Lei Bai# (* equal contribution, # corresponding author)

EMNLP 2025 main Oral paper, SAC Highlight Award, (Top 1%)

A reward-driven multi-agent framework that pairs task-graph generation with two-stage agent selection guided by a Collaborative Reward Model, plus an annotation-free pipeline for synthesizing multi-agent benchmarks.

Agents that Collaborate

From one policy to many bodies in one world. 3

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment

Li Kang*, Yutao Fan*, Rui Li*, Heng Zhou*, Yiran Qin, Zhemeng Zhang, Songtao Huang, Xiufeng Song, Zaibin Zhang, Bruno N.Y. Chen, Zhenfei Yin, Dongzhan Zhou, Wangmeng Zuo, Lei Bai# (* equal contribution, # corresponding author)

preprint

A compositional environment framework that pairs real-to-sim scene reconstruction with VLM-driven action synthesis and collision-checked sim-to-real transfer for safe multi-arm robotic collaboration.

Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning

Heng Zhou, Li Kang, Yiran Qin, Xiufeng Song, Ao Yu, Zilu Zhang, Haoming Song, Kaixin Xu, Yuchen Fan, Dongzhan Zhou, Xiaohong Liu, Ruimao Zhang, Philip Torr, Lei Bai#, Zhenfei Yin# (# corresponding author)

preprint

A cross-view benchmark plus a two-stage supervised-then-RL framework whose Cross-View Spatial Reward ties reasoning steps to visual evidence, fusing ego-centric observations into world-centric scene understanding.

VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning

Li Kang*, Xiufeng Song*, Heng Zhou*, Yiran Qin#, Jie Yang, Xiaohong Liu, Philip Torr, Lei Bai#, Zhenfei Yin# (* equal contribution, # corresponding author)

Annual Conference on Neural Information Processing Systems (NeurIPS) 2025

A hierarchical benchmark for embodied multi-agent cooperation spanning agent activation, task planning, and trajectory perception, paired with a chain-of-thought fine-tuning plus multi-level reinforcement learning framework.

Agents that Act

Grounding intelligence in perception and contact. 2

N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

NeoteAI Team and Fudan TEAI Team

Technical Report

A tactile-native world-action model that predicts future vision, contact, and action using a unified force-based tactile representation and an asymmetric Mixture-of-Transformers for real-time manipulation.

N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

NeoteAI Team and Fudan TEAI Team

Technical Report

A vision-tactile-language-action foundation model pretrained on large-scale visuo-tactile robot data, adding a predictive tactile pathway and advantage-conditioned offline reinforcement learning for contact-rich manipulation.