👋 I am a first-year Ph.D. student and a member of the Self-Evolving Embodied Agents Group at MMLab@SIGS, Tsinghua Shenzhen International Graduate School, advised by Prof. Zhi Wang.

🤔 My research focuses on Embodied AI, 4D World Modeling, Multimodal Dynamic Reasoning, and Vision Foundation Models, toward open-world embodied intelligence.

🙋‍♂️ If you are seeking any form of academic cooperation, please feel free to email me at yz_huang13@163.com.

🎓 Before that, I received my master's degree from the School of Informatics at Xiamen University, where I was a member of the SmartDSP Lab and was advised by Prof. Xinghao Ding and Prof. Yue Huang. I also collaborate with Dr. Chenxin Li from the AIM group at The Chinese University of Hong Kong.

📌 Research Interests

My research focuses on building embodied agents that can perceive, remember, reason, and act in dynamic physical environments, including: (i) VLM-based robotic agents with persistent spatio-temporal memory for long-horizon manipulation, (ii) physical-scale and streaming 4D world modeling, (iii) multimodal evaluation and enhancement for localized dynamics perception and 4D reasoning, and (iv) foundation-model-based perception under anomalous or ambiguous visual conditions. These directions connect dynamic perception, multimodal reasoning, persistent memory, and grounded action toward open-world embodied intelligence.

Embodied AI & Robotic Agents

  • Long-horizon robotic manipulation and VLM-based planning: RoboStream

4D World Modeling & Streaming Perception

  • Physical-scale multimodal 4D modeling: DynamicVerse
  • Object-centric streaming perception and memory compression: Zip-VGGT

Multimodal Dynamic Reasoning & Evaluation

Vision Foundation Models for Robust Perception

✉️ Welcome to contact me for discussions and collaborations on embodied AI, 4D world understanding, and multimodal reasoning.

🔥 News

  • 2026.07:  🎉 Two papers were accepted at ACM MM 2026!
  • 2026.06:  🎉 Thinking in Dynamics was selected for an oral presentation at the CVPR 2026 DataMFM Workshop and was also accepted to the CVPR 2026 4D Vision Workshop.
  • 2026.06:  🎉 Awarded the Tsinghua Future Scholar Scholarship, a highly selective university-level fellowship recognizing academic potential; only four students from Tsinghua SIGS were selected in 2026.
  • 2026.06:  🎉 RoboStream was accepted at ECCV 2026.
  • 2026.02:  🎉 Thinking in Dynamics (Dyn-Bench) was accepted to the CVPR 2026 Main Conference.
  • 2025.09:  🎉 DynamicVerse was accepted at NeurIPS 2025.
  • 2025.02:  🎉 Track Any Anomalous Object was accepted at CVPR 2025.
  • 2024.09:  🎉 Flaws can be Applause was accepted at NeurIPS 2024.
  • 2024.07:  🎉 P²SAM was accepted at ACM MM 2024.

📝 Publications

* denotes equal contribution, † denotes project lead.

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs

Rongxin Gao*, Yuzhi Huang*†, Dongxuan Liu*, Chu Li, Zhenye Wang, Jie Wu, Shuzhao Xie, Jingyan Jiang, Xinghao Ding, Xiaotong Tu, Yue Huang

ACM MM 2026

- We introduce DynTrace, a training-free framework that enables MLLMs to continuously track dynamic object evidence for coherent 4D spatio-temporal reasoning. It combines geometry-informed Dynamic Trajectory Visualization with a structured Dynamic Trace Graph to disentangle genuine object dynamics from camera-induced motion and preserve evolving object-level cues across time.

RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics

Yuzhi Huang*†, Jie Wu*, Weijue Bu, Ziyi Xiong, Gaoyang Jiang, Ye Li, Kangye Ji, Shuzhao Xie, Yue Huang, Chenglei Wu, Jingyan Jiang, Zhi Wang

ECCV 2026

- We propose RoboStream, a training-free framework that equips VLM-based robotic planners with persistent spatio-temporal memory for long-horizon manipulation. Spatio-Temporal Fusion Tokens maintain object grounding, while a Causal Spatio-Temporal Graph records action-triggered state transitions, achieving 90.5% on long-horizon RLBench and 44.4% on challenging real-world block-building tasks.

Zip-VGGT: Object-Centric Spatiotemporal KV Compression for Streaming Vision Transformers

Kairun Wen†, Runyu Chen, Wenyan Cong, Peiwei Lin, Tao Lu, Lihan Jiang, Weiguang Zhao, Junting Dong, Yunlong Lin, Yuzhi Huang, Xinghao Ding, Hongsheng Li, Linning Xu, Mulin Yu

Preprint

- Object-centric cache compression for efficient long-term 4D reconstruction.

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

Yuzhi Huang*†, Kairun Wen*, Rongxin Gao*, Dongxuan Liu, Yibin Lou, Jie Wu, Jing Xu, Jian Zhang, Zheng Yang, Yunlong Lin, Chenxin Li, Panwang Pan, Junbin Lu, Jingyan Jiang, Xinghao Ding, Yue Huang, Zhi Wang

CVPR 2026 Main Conference CVPR 2026 DataMFM Workshop — Oral Presentation CVPR 2026 4D Vision Workshop

- We introduce Dyn-Bench, a large-scale benchmark for systematically evaluating how MLLMs perceive, track, and reason about dynamics in the physical 4D world. It contains 1K diverse videos, 7K visual question-answering pairs, and 3K dynamic-object grounding pairs, revealing that existing models struggle to maintain both strong spatio-temporal reasoning and precise localized grounding. We further study structured integration strategies, including Mask-Guided Fusion and Spatio-Temporal Textual Cognitive Maps, to improve dynamic-world understanding.

DynamicVerse: Physically-Aware Multimodal Modeling for Dynamic 4D Worlds

Kairun Wen*, Yuzhi Huang*, Runyu Chen, Hui Zheng, Yunlong Lin, Panwang Pan, Chenxin Li, Wenyan Cong, Jian Zhang, Junbin Lu, Chenguo Lin, Dilin Wang, Zhicheng Yan, Hongyu Xu, Justin Theiss, Yue Huang, Xinghao Ding, Rakesh Ranjan, Zhiwen Fan

NeurIPS 2025

- We present DynamicVerse, a physically-aware multimodal 4D modeling framework that recovers metric-scale geometry, real-world motion, instance masks, and descriptive captions from monocular Internet videos. DynamicVerse provides a large-scale resource of 100K+ videos, 800K+ annotated masks, and 10M+ frames for learning and evaluating dynamic-world understanding.

Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline

Yuzhi Huang*, Chenxin Li*, Haitao Zhang, Zixu Lin, Yunlong Lin, Hengyu Liu, Wuyang Li, Xinyu Liu, Jiechao Gao, Yue Huang, Xinghao Ding, Yixuan Yuan

CVPR 2025

- A unified pipeline for detecting and tracking multiple fine-grained anomalous objects.

P²SAM: Probabilistically Prompted SAMs Are Efficient Segmentators for Ambiguous Medical Images

Yuzhi Huang*, Chenxin Li*, Zixu Lin, Hengyu Liu, Haote Xu, Yifan Liu, Yue Huang, Xinghao Ding, Yixuan Yuan

ACM MM 2024

- Probabilistic prompting turns SAM into an efficient model for ambiguous medical segmentation.

🏆 Honors and Awards

  • Tsinghua Future Scholar Scholarship, 2026 Selective university-level fellowship for academic Ph.D. students; four recipients from Tsinghua SIGS in 2026.
  • National Scholarship, 2026
  • Outstanding Graduate, 2026
  • Huang Xilie University-Level Scholarship, 2025
  • University Level Outstanding Student Award, 2024
  • Meritorious Winner in the Mathematical Contest in Modeling (MCM), USA, 2024

🎓 Educations

  • 2026 - (now), PhD, MMLab@SIGS, Tsinghua Shenzhen International Graduate School, Shenzhen.
  • 2023 - 2026, Master, School of Informatics, Xiamen University, Xiamen.
  • 2019 - 2023, Undergraduate, Fujian Normal University, Fuzhou.

💼 Internships

YXGN Robotics (远行光年)

Core Algorithm Member (Intern) · Embodied-AI startup incubated by our research group

📍 Shenzhen, China

📚 Academic Services

  • AAAI Conference on Artificial Intelligence (AAAI), Program Committee Member, 2026
  • International Conference on Machine Learning (ICML), 2025
  • International Conference on Learning Representations (ICLR), 2025
  • Conference on Neural Information Processing Systems (NeurIPS), 2025, 2024