avatar_circle.jpeg

Yongtao Ge ่‘›ๆถŒๆถ›

I am a Researcher at AlayaLab, working on world models and game agents. Previously, I was a Research Scientist at SpreeAI, working with Dr. Minh Vo. I obtained my Ph.D. in Computer Science from the University of Adelaide, advised by Prof. Chunhua Shen, and my Masterโ€™s degree from Southeast University.

My current research focuses on multimodal video generation, spatial intelligence and game agents. Previously, I worked on 3D human reconstruction and 2D perception.

News

Oct 2025 HumanWild is accepted by TPAMI. An updated demo is available on Huggingface.
Jun 2025 POMATO is accepted by ICCV 2025, with pointmap representation for dynamic 3D reconstruction.
Mar 2025 GVM, a generative video matting framework, has been accepted by SIGGRAPH 2025. ๐ŸŽ‰
Jan 2025 GenPercept is accepted by ICLR 2025. ๐ŸŽ‰
Jun 2024 Release GeoBench, a monocular geometry benchmark for analyzing SOTA monocular geometry estimation models. ๐ŸŽ‰
Mar 2024 Release HumanWild, feel free to try our Huggingface Demo. ๐ŸŽ‰
Jul 2023 Zolly is accepted by ICCV 2023, selected as oral (top 1.8%). :sparkles:
Mar 2023 Release HumanWild, focusing on perspective-distorted 3D human pose and shape estimation. :sparkles:
Nov 2022 Point-Teaching is accepted by AAAI 2023. :sparkles:
Jul 2022 Poseur is accepted by ECCV 2022. :sparkles:

Selected Publications

  1. WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
    Xiaojie Xu*, Zhengyuan Lin*, Runyi Li, Yihao Liu, Kaipeng Zhang, and Yongtao Ge
    arXiv preprint arXiv:2608.15659, 2026
  2. AlayaWorld: Long-Horizon and Playable Video World Generation
    AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu, Mingliang Zhai, and others
    arXiv preprint arXiv:2608.13492, 2026
  3. Sekai2: From World Exploration to Interactive World Modeling
    Kang He, Wenshuo Peng, Zihui Gao, Jiaming Tan, Kaipeng Zhang, and Yongtao Ge
    arXiv preprint arXiv:2608.09449, 2026
  4. worldmark.jpg
    WorldMark: A Unified Benchmark Suite for Interactive Video World Models
    Xiaojie Xu, Zhengyuan Lin, Kang He, Yukang Feng, Xiaofeng Mao, Yuanyang Yin, Yongtao Ge, and Kaipeng Zhang
    arXiv preprint arXiv:2604.21686, 2026
  5. WorldSculpt: Generating Compositional Worlds from Grounded Videos
    Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, and Zhixiang Wang
    arXiv preprint arXiv:2609.05416, 2026
  6. Generative Video Matting
    Yongtao Ge, Kangyang Xie, Guangkai Xu, Li Ke, Mingyu Liu, Longtao Huang, Hui Xue, Hao Chen, and Chunhua Shen
    In ACM SIGGRAPH Conference Papers, 2025
  7. pomato_teaser.png
    POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction
    Songyan Zhang*, Yongtao Ge*, Jinyuan Tian*, Guangkai Xu, Chen Lv, Hao Chen, and Chunhua Shen
    In Proc. of the IEEE International Conf. on Computer Vision, 2025
  8. humanwild.gif
    3D Human Reconstruction in the Wild with Synthetic Data Using Generative Models
    Yongtao Ge, Wenjia Wang, Yongfan Chen, Hao Chen, and Chunhua Shen
    In IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 2025
  9. genpercept_pipeline.png
    What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
    Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen, and Chunhua Shen
    In Proc. of the IEEE International Conf. on Learning Representations, 2025
  10. geobench_logo.png
    GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
    Yongtao Ge, Guangkai Xu, Zhiyue Zhao, Zheng Huang, Libo Sun, Yanlong Sun, Hao Chen, and Chunhua Shen
    arXiv preprint arXiv:2406.12671, 2024
  11. zolly.png
    Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh Reconstruction
    Wenjia Wang, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Yanjun Wang, Chunhua Shen, Lei Yang, and Komura Taku
    In Proc. of the IEEE International Conf. on Computer Vision (ICCV Oral), 2023
  12. poseur_arch.jpg
    Poseur: Direct Human Pose Regression with Transformers
    Weian Mao*, Yongtao Ge*, Chunhua Shen, Zhi Tian, Xinlong Wang, Zhibin Wang, and Anton van den Hengel
    In Proc. of the European Conf. on Computer Vision (ECCV), 2022
  13. point_teaching_arch.jpg
    Point-Teaching: Weakly Semi-Supervised Object Detection with Point Annotations
    Yongtao Ge*, Qiang Zhou*, Xinlong Wang, Chunhua Shen, Zhibin Wang, and Hao Li
    In AAAI Conference on Artificial Intelligence (AAAI), 2023