One paper was accepted by IJCV 2026. ↗
Mingzhe Zheng郑名哲 · HKUSTProfile
PhD Student @ HKUST
Research Intern @ Tencent Hunyuan
Mingzhe Zheng · 郑名哲
Building agentic and multimodal AI towards physical intelligence
I am a second-year PhD student at the Hong Kong University of Science and Technology, advised by Prof. Harry Yang and Prof. Qifeng Chen. I received my B.Eng. in Artificial Intelligence with the Outstanding Graduate Award from Northwestern Polytechnical University in 2024.
I am currently a Qingyun Research Intern (Top Talent Internship) at Tencent Hunyuan, working on pre-training and post-training for large-scale video foundation models. Previously, I was fortunate to work with Everlyn Labs and HeyGen.
I aim to study efficient multimodal visual generation algorithms from both theoretical and practical perspectives while exploring meaningful exploration of multimodal AI and agentic AI. I believe multimodal AI is naturally progressing from end-to-end generation toward interactive systems. In parallel, advances in long-horizon agentic reinforcement learning will extend LLM-based agents beyond human-designed office tasks into interactive physical environments. These two research directions will ultimately converge toward Physical AI, leading to a still underexplored area that I am particularly interested in: Agentic World Models.
My long-term research goal is to create a visual-centric agentic system with a self-evolving "brain" that can comprehend the environment and generate reactions to real-life interactions, then produce cinematic-level visual content similar to how human directors make films, titled "AI Movie Industry."
News
ICDepth was accepted to ECCV 2026. ↗
Twins was accepted to ICML 2026. ↗
We released SAGE-GRPO and Group Editing. ↗
Group Editing was accepted to CVPR 2026. ↗
Two papers were accepted to ICLR 2026. ↗
HunyuanVideo 1.5 Technical Report was released. ↗
We released CML-Bench for LLM-powered movie script generation. ↗
VideoGen-of-Thought was selected as a NeurIPS NextVid Workshop oral. ↗
We released our survey on controllable video generation. ↗
Two papers were accepted to ICCV 2025. ↗
FineCLIPER was accepted to ACM Multimedia 2024. ↗
Projects & Selected Papers

Complete record · newest first
Google Scholar ↗Industrial Experience
Qingyun Research Intern · Top Talent Internship
Mentor: Dr. Yue Wu ↗Large-scale video foundation model pre-training and post-training, SAGE-GRPO, agentic data recipes, emotion data, agentic RL, and reward systems for image and video generation.
Researcher & Algorithm Engineer
Mentor: Prof. Ser-Nam Lim ↗Developed autoencoders and experimental pipelines for image and video generation, and studied frontier video-model architectures and training paradigms.
Computer Vision Algorithm Intern
Mentor: Dr. Zhibin (Charly) Hong ↗Developed multilingual forced alignment across 26 languages for HeyGen’s first-generation avatar pipeline and researched audio- and video-driven talking-head generation.
Machine Learning Algorithm Intern
Core Algorithm Department ↗Built a millimeter-wave-radar sleep-status prediction system with LightGBM and improved accuracy from 72% to 86%.
Academic Experience
Hong Kong University of Science and Technology
PhD Student & Researcher · Generative Models, Multimodality, Movie Generation
Advisors: Prof. Harry Yang & Prof. Qifeng Chen ↗Research on agentic and multimodal visual generation, foundation-model post-training, and machine-assisted movie creation.
DianLab · Northwestern Polytechnical University
Research Assistant · Generative Models & Multimodality
Advisor: Prof. Dian Shao ↗Worked on LLM-assisted visual generation, FineCLIPER, and high-fidelity text-guided 3D face generation with diffusion models and neural radiance fields.
StatsLE Lab · University of Toronto
Research Assistant · Generative Models & Text-to-3D
Advisor: Prof. Qiang Sun ↗Studied text-to-3D architectures and acceleration strategies, combining efficient generation with statistical machine learning.
National Engineering Lab · NPU
Research Assistant · Generative Models & Super Resolution
Advisor: Prof. Wei Wei ↗Developed a latent-diffusion and VQGAN system for super resolution and explored modified latent diffusion for visual classification.
Northwestern Polytechnical University
Undergraduate Researcher · HCI, Computer Vision & EEG
Advisor: Prof. XinBo Zhao ↗Built EEG- and gaze-driven interaction prototypes and led a deep-learning project on cognitive-perception fatigue-driving detection.
Appendix
China Generative AI Conference
Video Generation Venue
Conference ↗Meshy Finalist Fellowship
Recognized as a 2026 finalist fellow.
Fellowship ↗Review Service
ECCV · CVPR · NeurIPS · ICLR · ICML
ICCV · CVPR · NeurIPS · ICLR · ICML
Visitor Map
Live around the world
A global research conversation.
The live counter tracks unique visits while the dotted globe reflects a worldwide research audience.



















