Mingzhe Zheng郑名哲 · HKUSTProfile

Mingzhe Zheng · 郑名哲

Building agentic and multimodal AI towards physical intelligence

I am a second-year PhD student at the Hong Kong University of Science and Technology, advised by Prof. Harry Yang and Prof. Qifeng Chen. I received my B.Eng. in Artificial Intelligence with the Outstanding Graduate Award from Northwestern Polytechnical University in 2024.

I am currently a Qingyun Research Intern (Top Talent Internship) at Tencent Hunyuan, working on pre-training and post-training for large-scale video foundation models. Previously, I was fortunate to work with Everlyn Labs and HeyGen.

I aim to study efficient multimodal visual generation algorithms from both theoretical and practical perspectives while exploring meaningful exploration of multimodal AI and agentic AI. I believe multimodal AI is naturally progressing from end-to-end generation toward interactive systems. In parallel, advances in long-horizon agentic reinforcement learning will extend LLM-based agents beyond human-designed office tasks into interactive physical environments. These two research directions will ultimately converge toward Physical AI, leading to a still underexplored area that I am particularly interested in: Agentic World Models.

My long-term research goal is to create a visual-centric agentic system with a self-evolving "brain" that can comprehend the environment and generate reactions to real-life interactions, then produce cinematic-level visual content similar to how human directors make films, titled "AI Movie Industry."

01

News

One paper was accepted by IJCV 2026.

ICDepth was accepted to ECCV 2026.

Twins was accepted to ICML 2026.

We released SAGE-GRPO and Group Editing.

Group Editing was accepted to CVPR 2026.

Two papers were accepted to ICLR 2026.

HunyuanVideo 1.5 Technical Report was released.

We released CML-Bench for LLM-powered movie script generation.

VideoGen-of-Thought was selected as a NeurIPS NextVid Workshop oral.

We released our survey on controllable video generation.

Two papers were accepted to ICCV 2025.

FineCLIPER was accepted to ACM Multimedia 2024.

02

Projects & Selected Papers

Complete record · newest first

Google Scholar ↗
First figure from Twins: Learn to Predict Unified Representations with Focal Loss2026

Twins: Learn to Predict Unified Representations with Focal Loss

Kaixiong Gong*, Xin Cai*, Bin Lin, Hao Wang, Yunlong Lin, Mingzhe Zheng, Bohao Li, Jian-Wei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Xiangyu Yue

ICML 2026

First figure from ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning2026

ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning

Xuanhua He, Jiaxin Xie, Mingzhe Zheng, Qifeng Chen

ECCV 2026

First figure from D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models2026

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven Hoi

arXiv 2026

First figure from Group Editing: Edit Multiple Images in One Go2026

Group Editing: Edit Multiple Images in One Go

Yue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang, Mingzhe Zheng, Xiangpeng Yang, Hao Li, Chongbo Zhao, Jixuan Ying, Harry Yang, Hongyu Liu, Qifeng Chen

CVPR 2026

First figure from Manifold-Aware Exploration for Reinforcement Learning in Video Generation2026

Manifold-Aware Exploration for Reinforcement Learning in Video Generation

Mingzhe Zheng*, Weijie Kong*, Yue Wu, Dengyang Jiang, Yue Ma, Xuanhua He, Bin Lin, Kaixiong Gong, Zhao Zhong, Liefeng Bo, Qifeng Chen, Harry Yang

arXiv 2026 · SAGE-GRPO

First figure from SCALAR++: Efficient Controllable Generation via Scale-wise Visual Autoregressive Learning2026

SCALAR++: Efficient Controllable Generation via Scale-wise Visual Autoregressive Learning

Ryan Xu*, Dongyang Jin*, Shawn Chen*, Yancheng Bai, Jingzhe Ma, Rui Lan, Jianhao Zeng, Yunyang Ge, Mingzhe Zheng, Lei Sun, Xiangxiang Chu

International Journal of Computer Vision, 2026

First figure from FastVMT: Eliminating Redundancy in Video Motion Transfer2026

FastVMT: Eliminating Redundancy in Video Motion Transfer

Yue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng, Hongyu Liu, Jiayi Guo, Kunyu Feng, Yuxuan Xue, Zixiang Zhao, Konrad Schindler, Qifeng Chen, Linfeng Zhang

ICLR 2026

First figure from iFSQ: Improving FSQ for Image Generation with 1 Line of Code2026

iFSQ: Improving FSQ for Image Generation with 1 Line of Code

Bin Lin, Zongjian Li, Yuwei Niu, Kaixiong Gong, Yunyang Ge, Yunlong Lin, Mingzhe Zheng, Jian-Wei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Li Yuan

arXiv 2026

First figure from CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation2025

CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation

Mingzhe Zheng, Dingjie Song, Guanyu Zhou, Jun You, Jiahao Zhan, Xuran Ma, Xinyuan Song, Ser-Nam Lim, Qifeng Chen, Harry Yang

arXiv 2025

First figure from Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control2025

Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control

Zeqian Long*, Mingzhe Zheng* (co-first), Kunyu Feng*, Xinhua Zhang, Hongyu Liu, Harry Yang, Linfeng Zhang, Qifeng Chen, Yue Ma

ICLR 2026

First figure from Controllable Video Generation: A Survey2025

Controllable Video Generation: A Survey

Yue Ma, Kunyu Feng, Zhongyuan Hu, Xinyu Wang, Yucheng Wang, Mingzhe Zheng, Bingyuan Wang, Qinghe Wang, Xuanhua He, Hongfa Wang, Chenyang Zhu, Hongyu Liu, Yingqing He, Zeyu Wang, Zhifeng Li, Xiu Li, Sirui Han, Yike Guo, Wei Liu, Dan Xu, Linfeng Zhang, Qifeng Chen

arXiv Survey, 2025

First figure from Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models2025

Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models

Xuran Ma, Yexin Liu, Yaofu Liu, Xianfeng Wu, Mingzhe Zheng, Zihao Wang, Ser-Nam Lim, Harry Yang

ICCV 2025

First figure from VideoGen-of-Thought: Step-by-step Generating Multi-shot Video with Minimal Manual Intervention2025

VideoGen-of-Thought: Step-by-step Generating Multi-shot Video with Minimal Manual Intervention

Mingzhe Zheng, Yongqi Xu, Haojian Huang, Xuran Ma, Yexin Liu, Wenjie Shu, Yatian Pang, Feilong Tang, Qifeng Chen, Harry Yang, Ser-Nam Lim

NeurIPS 2025 NextVid Workshop · Oral

First figure from DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses2025

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

Yatian Pang, Bin Zhu, Bin Lin, Mingzhe Zheng, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan

ICCV 2025

First figure from FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs2024

FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs

Haodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng, Dian Shao

ACM Multimedia 2024

03

Industrial Experience

Tencent Hunyuan

Qingyun Research Intern · Top Talent Internship

Mentor: Dr. Yue Wu

Large-scale video foundation model pre-training and post-training, SAGE-GRPO, agentic data recipes, emotion data, agentic RL, and reward systems for image and video generation.

Everlyn Labs

Researcher & Algorithm Engineer

Mentor: Prof. Ser-Nam Lim

Developed autoencoders and experimental pipelines for image and video generation, and studied frontier video-model architectures and training paradigms.

04

Academic Experience

Hong Kong University of Science and Technology

PhD Student & Researcher · Generative Models, Multimodality, Movie Generation

Advisors: Prof. Harry Yang & Prof. Qifeng Chen

Research on agentic and multimodal visual generation, foundation-model post-training, and machine-assisted movie creation.

DianLab · Northwestern Polytechnical University

Research Assistant · Generative Models & Multimodality

Advisor: Prof. Dian Shao

Worked on LLM-assisted visual generation, FineCLIPER, and high-fidelity text-guided 3D face generation with diffusion models and neural radiance fields.

StatsLE Lab · University of Toronto

Research Assistant · Generative Models & Text-to-3D

Advisor: Prof. Qiang Sun

Studied text-to-3D architectures and acceleration strategies, combining efficient generation with statistical machine learning.

National Engineering Lab · NPU

Research Assistant · Generative Models & Super Resolution

Advisor: Prof. Wei Wei

Developed a latent-diffusion and VQGAN system for super resolution and explored modified latent diffusion for visual classification.

Northwestern Polytechnical University

Undergraduate Researcher · HCI, Computer Vision & EEG

Advisor: Prof. XinBo Zhao

Built EEG- and gaze-driven interaction prototypes and led a deep-learning project on cognitive-perception fatigue-driving detection.

05

Appendix

Invited Talk · 2026

China Generative AI Conference

Video Generation Venue

Conference ↗
Award · 2026

Meshy Finalist Fellowship

Recognized as a 2026 finalist fellow.

Fellowship ↗

Review Service

ECCV · CVPR · NeurIPS · ICLR · ICML

ICCV · CVPR · NeurIPS · ICLR · ICML

06

Visitor Map

Live around the world

A global research conversation.

The live counter tracks unique visits while the dotted globe reflects a worldwide research audience.