|
Ziyang Xu (徐子扬)
I am a Ph.D. student in Information and Communication Engineering at the School of Electronic Information and Communications (EIC), Huazhong University of Science and Technology (HUST, 华中科技大学), advised by Professors Xinggang Wang and Wenyu Liu. I began my Ph.D. studies in September 2025. Prior to this, I received my M.E. degree in Information and Communication Engineering from EIC, HUST (华中科技大学) in 2025, and my B.E. degree in Information Engineering from the School of Information Engineering (IE), Wuhan University of Technology (WHUT, 武汉理工大学) in 2022. My research spans visual generation, efficient multimodal models, visual perception and memory, and AI for science, with a current focus on generative AI.
Email |
Google Scholar |
GitHub
|
ORCID |
X
Latest Update: September 1, 2026
|
|
Highlight
PixelHacker and Moebius, my first/co-first-authored image inpainting projects, both achieved the No. 1 daily ranking on Hugging Face.
An early quantized on-device image inpainting model that I developed during my internship at vivo AI Lab—predating PixelHacker and Moebius, and related to the later PixelHacker line of work—has been rolled out across vivo’s full smartphone lineup, serving 400 million active users worldwide. You can find it in Gallery under AI Retouching → AI Eraser.
My research on multi-frame generative low-dose DSA (Digital Subtraction Angiography) imaging spans MoSt-DSA, GaraMoSt, GenDSA, and GenDSA-V2, with results published in ECAI, AAAI, Med, and Nature Medicine. This line of work contributes to a National Key R&D Program of China and has advanced from algorithm development to clinical validation and medical-device integration. More than 100,000 patients worldwide undergo DSA-guided procedures every day. In a completed prospective randomized controlled trial involving 1,068 patients, GenDSA-V2 reduced radiation exposure for patients and healthcare professionals by approximately two-thirds. Clinical trial registration: ChiCTR2400084789.
As of September 1, 2026, my open-source research projects have received 1,150+ GitHub stars and 110+ forks in total, with a single project reaching 580+ stars.
|
Research
My research spans visual generation, efficient multimodal models, visual perception and memory, and AI for science, with a particular emphasis on generative AI and medical imaging. Publications are grouped by research theme; my name is bolded, * denotes equal contribution, and † denotes project leadership.
|
1/3 Efficient Visual Generation and Image Inpainting
I develop image generation systems that balance high-quality synthesis and structural-semantic consistency with practical model efficiency.
|
|
|
Representative Work
* Equal contribution; † project leadership
Role: Co-first Author · Project Lead
This work was completed while Ziyang Xu and Kangsheng Duan were research interns at vivo AI Lab.
Moebius extends the PixelHacker line of work by asking whether high-quality image inpainting requires continued model scaling. It combines a compact architecture with knowledge distillation, using only 0.22B parameters—less than 2% of FLUX.1-Fill-Dev—while matching or surpassing it across six natural and portrait benchmarks and reducing total inference time by more than 15×.
|
|
|
Representative Work
† Project leadership
Role: First Author · Project Lead
This work was completed while Ziyang Xu and Kangsheng Duan were research interns at vivo AI Lab.
PixelHacker introduces Latent Categories Guidance (LCG), which separately models foreground and background semantic distributions and injects them into diffusion-based denoising. Trained with 14 million image-mask pairs, it improves structural and semantic consistency across natural and portrait inpainting benchmarks while remaining substantially smaller and faster than large industrial baselines.
|
2/3 Generative Low-Dose DSA Imaging
This research line progresses from direct multi-frame interpolation and efficient generation to large-scale pretraining, multi-center validation, and prospective randomized clinical evaluation.
|
|
|
Representative Work
Huangxuan Zhao*, Yaowei Bai*, Lei Chen*, Jinqiang Ma*, Yu Lei, Tao Sun, Linxia Wu, Ruiheng Zhang,
Ziyang Xu, Xiaoyun Liang, Yi Li, Yan Huang, Yun Feng, Cheng Hong, Zhongrong Miao,
Lin Long, Haidong Zhu, Jiahe Zheng, Lin Fan, Zhuting Fang, Peng Dong, Lefei Zhang, Xiaoyu Han,
Bin Wang, Bin Liang, Xiangwen Xia, Xuefeng Kan, Chengcheng Zhu, Bo Du,
Xinggang Wang,
Chuansheng Zheng
* Equal contribution
Role: Algorithm Lead
GenDSA-V2 iterates our previously developed GenDSA system using data from 46,829 patients, more than 5 million DSA images, and 70 centers. In a completed prospective randomized controlled trial of 1,068 patients, it reduced radiation exposure for patients and healthcare professionals by approximately two-thirds while maintaining noninferior operation time and similar complication rates.
|
|
|
Representative Work
Huangxuan Zhao*,
Ziyang Xu*, Lei Chen*, Linxia Wu*,
Ziwei Cui, Jinqiang Ma, Tao Sun, Yu Lei, Nan Wang,
Hongyao Hu, Yiqing Tan, Wei Lu, Wenzhong Yang, Kaibing Liao, Gaojun Teng, Xiaoyun Liang, Yi Li,
Congcong Feng, Tong Nie, Xiaoyu Han, Dongqiao Xiang, Charles B.L.M. Majoie, Wim H. van Zwam,
Aad van der Lugt, P. Matthijs van der Sluijs, Theo van Walsum, Yun Feng, Guoli Liu, Yan Huang,
Wenyu Liu, Xuefeng Kan, Ruisheng Su,
Weihua Zhang, Xinggang Wang,
Chuansheng Zheng
* Equal contribution
Role: Co-first Author · Algorithm Lead
GenDSA is a large-scale pretrained multi-frame generative system for real-time low-dose DSA imaging. It was developed on approximately 3 million DSA images from 27,117 patients at 10 hospitals and evaluated on two additional datasets from 25 hospitals; using one-third of the clinical radiation dose, it generated frames in 0.07 seconds and achieved image quality comparable to full-sampled videos.
|
|
|
Role: First Author
GaraMoSt extends MoSt-DSA with parallel multi-granularity motion and structural feature extraction. The design improves noise suppression, accuracy, and visual quality within the same computational time scale, outperforming MoSt-DSA and natural-scene video frame interpolation baselines on DSA interpolation.
|
|
|
Role: First Author
MoSt-DSA is the first deep learning method designed specifically for DSA frame interpolation. It models motion-structure interactions and directly generates an arbitrary number of intermediate frames at arbitrary time steps in a single forward pass, establishing the algorithmic foundation for the subsequent GaraMoSt and GenDSA systems.
|
3/3 Multimodal Generation and Visual Perception
I also study multimodal world generation and robust visual perception, including joint video-LiDAR generation and extremely small object detection in videos.
|
|
|
Xiangyu Guo*, Zhanqian Wu*, Kaixin Xiong*, Ziyang Xu, Lijun Zhou, Gangwei Xu,
Shaoqing Xu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye,
Wenyu Liu,
Xinggang Wang
* Equal contribution
Genesis jointly generates multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. It couples a DiT-based video diffusion model and a BEV-aware LiDAR generator through a shared latent space, achieving strong generation results on nuScenes and improving downstream segmentation and 3D detection.
|
|
|
XS-VID introduces a diverse aerial video benchmark for extremely small object detection, covering eight categories and objects below 32² pixels together with track and motion annotations. The accompanying YOLOFT baseline strengthens local feature association and temporal motion modeling, improving the accuracy and stability of small video object detection.
|
Internship
Research Intern, vivo AI Lab
October 2024 – August 2025 · Hangzhou, China
Worked on efficient on-device image inpainting and multimodal content generation, with a focus on model distillation and deployment-oriented optimization. This work led to PixelHacker and Moebius.
|
Selected Honors & Awards
Outstanding Master’s Graduate, Huazhong University of Science and Technology (HUST) (优秀硕士毕业生), 2025.
National Scholarship for Graduate Students (Master’s Level) (研究生国家奖学金,硕士研究生), 2024.
Awarded to only 0.2% of candidates nationwide.
Outstanding Graduate Student, HUST (三好硕士生), 2024.
Outstanding Graduate, Wuhan University of Technology (WHUT) (优秀本科毕业生), 2022.
Provincial Undergraduate Innovation and Entrepreneurship Training Program — Excellent Completion (优秀结题), 2021.
Team Lead “Research and Implementation of a Video-Based Vehicle-Assisted Driving Device.”
Outstanding Undergraduate Student, WHUT (三好本科生), 2021.
|
Selected Competitions
Team Lead Huawei Cup 19th China Graduate Mathematical Modeling Competition (华为杯第十九届中国研究生数学建模竞赛), National Second Prize (国家级二等奖), 2022.
Top 13% (前13%) of all teams.
Team Lead Artificial Intelligence Track of the 14th China College Student Computer Design Competition (中国大学生计算机设计大赛人工智能赛道), National Third Prize (国家级三等奖), 2021.
Among the 4,886 entries that advanced from the provincial round to the national competition, 8.5% received this award.
National College Student Integrated Circuit Innovation and Entrepreneurship Competition (全国大学生集成电路创新创业大赛), Third Prize in Hubei Province (湖北省三等奖), 2021.
Team Lead The 20th Innovation Cup College Students Extracurricular Academic Science and Technology Works Competition (创新杯大学生课外学术科技作品竞赛), Special Prize (特等奖) in the Mathematical Information Group (数理信息组), 2020.
I led the team to win the sole Special Prize (唯一特等奖) in the Mathematical Information category.
Team Lead National College Student Mathematical Modeling Competition (全国大学生数学建模竞赛), Second Prize in Hubei Province (湖北省二等奖), 2020.
|
Academic Service
Conference Reviewer: CVPR 2026; ECCV 2024 and 2026.
Journal Reviewer: IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) and Image and Vision Computing, 2026; International Journal of Computer Vision (IJCV), 2024.
|
|