Research · Shanghai Jiao Tong University
Spatial Intelligence
Neural Rendering3D Gaussian Splatting
Spatial AgentsVision-Language Models
I am Zhihang Zhong (钟志航), an associate professor at the School of Artificial Intelligence, Shanghai Jiao Tong University, where I lead Visionary Laboratory (空间多媒体实验室). Previously, I was a researcher at the Shanghai AI Laboratory.
I received my PhD in Computer Science and ME in Precision Engineering from the University of Tokyo, and my BE in Mechatronics from Chu Kochen Honors College, Zhejiang University.
News
- I am appointed as an Area Chair for ICLR.
- I am appointed as a Senior Program Committee member for AAAI.
- Two papers are accepted to ACM MM 2026.
- Two papers (one Spotlight) are accepted to ECCV 2026!
- We release SpaceDG, the first benchmark for Spatial Intelligence under visual degradations!
- Two papers (one Oral: Holi-Spatial) are accepted to ICML 2026!
Less news
- We release CourtSI, the first benchmark for Sports Spatial Intelligence!
- We release Holi-Spatial, a data creation engine that transforms video into spatial intelligence!
- Two papers are accepted to CVPR 2026 (one Best Paper Candidate : Proxy-GS)
- InterpAny is accepted to TPAMI!
- We are thrilled to release Visionary, the World Model Carrier!!
- One paper (Oral) is accepted to AAAI 2026.
- Three papers are accepted to ICCV 2025.
- MaskGaussian is accepted to CVPR 2025.
- Glad to receive the 2023 Chinese Government Award for Outstanding
Self-financed Students Abroad (Group B, Global Top 50)! - Three papers (one Oral) are accepted to ECCV 2024!
- Glad to receive the Dean's Award for Academic Achievement from the UTokyo!
- Two papers are accepted to CVPR 2024.
- I give a talk at OpenMMLab about temporal super-resolution.
- Glad to release InterpAny-Clearer project!
- Two papers are accepted to ICCV 2023.
- One paper is accepted to ACM MM 2023.
- I am accepted to CVPR 2023's Doctoral Consortium.
- Two papers are accepted to CVPR 2023.
- I give a talk at MIPI Workshop 2022.
- One paper is accepted to IJCV.
- I become a JSPS DC fellow!
- Three papers (one Oral) are accepted to ECCV 2022!
- I become a JEM intern at Microsoft.
- One paper is accepted to CVPR 2022.
- I become a research intern in the Visual Computing group at MSRA.
- I become a IIW fellow of UTokyo!
- One paper is accepted to IoTJ.
- One paper is accepted to CVPR 2021.
- I become a MSRA D-CORE fellow!
- I obtain my M.E. degree from UTokyo with an outstanding thesis award!
- One paper (Spotlight) is accepted to ECCV 2020!
- One paper is accepted to IUI 2020.
Projects
Publications
denotes corresponding author2026
| Intern-S2-Preview: Scientific Agentic Foundation Model arXiv, 2026 arXiv / code |
| SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation arXiv, 2026 project / arXiv / code |
| PhotoFlow: Agentic 3D Virtual Photography Missions arXiv, 2026 project / arXiv / code |
| Segment and Select: Vision-Language Segmentation in 3D Scenarios arXiv, 2026 arXiv |
| Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence ICML, 2026, Oral project / arXiv / code |
| Perceptual Flow Network for Visually Grounded Reasoning ICML, 2026 arXiv |
| Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports arXiv, 2026 project / arXiv / code |
| InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing Technical Report, 2026 arXiv / code |
| GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing ECCV, 2026, Spotlight project / arXiv / code |
| Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation ECCV, 2026 |
| Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting CVPR, 2026, Oral, Best Paper Candidate 🎖️ project / arXiv / code |
| Motion-Aware Animatable Gaussian Avatars Deblurring CVPR, 2026 arXiv / code |
| Velocity Disambiguation for Video Frame Interpolation TPAMI, 2026 paper / arXiv |
| RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis AAAI, 2026, Oral arXiv / code |
| AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models ACM MM, 2026 project / arXiv / code |
| Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark ACM MM, 2026 |
2025
| Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform Technical Report, 2025 project / arXiv / code / editor |
| CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction ICCV, 2025 project / arXiv / code |
| Sequential Gaussian Avatars with Hierarchical Motion Context ICCV, 2025 project / arXiv / code |
| Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars ICCV, 2025 arXiv / code |
| MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks CVPR, 2025 project / arXiv / code |
| DiffBody: Human Body Image Restoration with Generative Diffusion Prior ICCP, 2025 arXiv |
2024
| Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence ECCV, 2024 project / arXiv / code |
| Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation ECCV, 2024, Oral project / arXiv / code |
| KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter ECCV, 2024 arXiv / code |
| IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame Interpolation CVPR, 2024 paper |
| Fooling Polarization-based Vision using Locally Controllable Polarizing Projection CVPR, 2024 paper / arXiv |
2023
| NIR-assisted Video Enhancement via Unpaired 24-hour Data ICCV, 2023 paper / code |
| Rethinking Video Frame Interpolation from Shutter Mode Induced Degradation ICCV, 2023 paper / code |
| Event-guided Frame Interpolation and Dynamic Range Expansion of Single Rolling Shutter Image ACM MM, 2023 paper |
| Blur Interpolation Transformer for Real-World Motion from Blur CVPR, 2023 project / paper / arXiv / code / zhihu |
| Visibility Constrained Wide-band Illumination Spectrum Design for Seeing-in-the-Dark CVPR, 2023 paper / arXiv / code |
2022
| Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance ECCV, 2022 project / paper / arXiv / code / zhihu |
| Bringing Rolling Shutter Images Alive with Dual Reversed Distortion ECCV, 2022, Oral project / paper / arXiv / code |
| Efficient Video Deblurring Guided by Motion Magnitude ECCV, 2022 paper / arXiv / code |
| Learning Adaptive Warping for Real-World Rolling Shutter Correction CVPR, 2022 paper / arXiv / code |
| Real-world Video Deblurring: A Benchmark Dataset and An Efficient Recurrent Neural Network International Journal of Computer Vision (IJCV), 2022 paper / arXiv / code |
2021
| Towards Rolling Shutter Correction and Deblurring in Dynamic Scenes CVPR, 2021 paper / arXiv / code |
| Multistream Temporal Convolutional Network for Correct/Incorrect Patient Transfer Action Detection using Body Sensor Network IEEE Internet of Things Journal (IoTJ), 2021 paper / code |
2020
| Efficient Spatio-Temporal Recurrent Neural Network for Video Deblurring ECCV, 2020, Spotlight paper / arXiv / code |
| Multi-attention Deep Recurrent Neural Network for Nursing Action Evaluation using Wearable Sensor IUI, 2020 paper |
Teaching
Fall 2026
Parallel Computing and Operator Programming (并行计算与算子编程)
Shanghai Jiao Tong University
AI Engineering (AI工程学)
Shanghai Innovation Institute



