About Me
My research focuses on distributed computing theory and machine learning systems. I have published over 40 papers in top conferences and journals, been granted more than 40 invention patents, and led the release of 3 ITU international standards. My work has received the World Internet Conference Leading Science and Technology Award, Zhejiang Provincial Science and Technology Progress Award (First and Second Class), and has achieved significant application results in the construction of domestic AI computing clusters.
Biography
Associate Professor
Research Contributions
1. Distributed Computing Theory and System Architecture Innovation
Addressing the challenge of decoupling logical representation from physical resources in distributed training, I proposed the SBP (Split-Batch-Parallel) abstraction for distributed training, which serves as the core mathematical foundation for the domestic deep learning framework OneFlow. This theoretical innovation enables unified expression of operator-level parallel strategies and automatic derivation of data routing, mathematically solving the separation problem between distributed logical views and physical device views. It enabled OneFlow to become a performance-leading distributed framework and a representative domestic alternative to PyTorch at the time. This work provides an important theoretical basis for the formal study of distributed training systems, and its design ideas have been adopted by multiple domestic frameworks.
2. Heterogeneous Cluster Training Framework and Automatic Parallelism
Addressing the challenges of complex heterogeneous network environments and difficult parallel strategy adaptation in domestic clusters, I developed Holmes, the first distributed training framework adapted for heterogeneous network environments. We broke through non-uniform partitioning of 3D parallel strategies and automatic routing technologies for heterogeneous networks, achieving 1.4x performance improvement over mainstream frameworks like Megatron in hybrid network environments. This work became a pioneering study in heterogeneous hybrid training on domestic clusters, published at ICPP, the top conference in distributed computing. Furthermore, we proposed AutoHAAP, an automatic parallel framework for heterogeneous clusters, achieving automatic parallel strategy optimization based on cost models and heuristic search, supporting minute-level search at the 10,000-card scale. This work was published at HPCA, the top conference in computer architecture, and is at the leading level in the field of heterogeneous distributed training.
3. Computational Efficiency Optimization and Domestic Chip Adaptation
Focusing on the core challenges of diverse numerical formats and complex distribution characteristics of domestic chips, I systematically conducted research on hardware-aware compression techniques. We proposed a series of tensor distribution-sensitive compression methods including FlattenQuant (for large language models), ADFQ-ViT (for vision models), and PTQ4PLM (for protein language models), effectively solving the problem of outlier suppression and quantization accuracy loss. Our results have been published in top conferences and journals including ICML, Neural Networks, and LREC-Coling. Through cross-team collaboration, we applied these technologies to the adaptation and optimization of 236B MoE language models, 72B geoscience models, 70B astronomy models, and 13B protein models on domestic chips, achieving 40% MFU for MoE model training on a domestic 1,000-card cluster, which reaches the industry-leading level. As subsystem leader, our work "Aggregated Heterogeneous Computing Power Intelligent Computing System" won the Second Prize of Zhejiang Provincial Science and Technology Progress Award.
4. Large-Scale System Integration and Engineering Practice
As technical lead, I led the development of the large model training and inference system in the Nanhu Computing Framework, leading the team to breakthrough multiple key efficiency optimization technologies including mixed-precision training, asynchronous checkpointing, data splicing compression, and deep bottleneck analysis. This supported the construction and validation of the basic 10,000-card and hybrid 10,000-card clusters at Zhejiang's New Computing Power Center. The research results have been applied to multiple scientific large model training tasks at Zhejiang Laboratory, including 021, geoscience, astronomy, and biology, and the related technologies have saved nearly 100 million yuan in computing costs, achieving significant social and economic benefits. The "Nanhu Computing Framework" I participated in won the World Internet Conference Leading Science and Technology Award, and "Key Technologies and Applications of Intelligent Computing for Large Language Model Training" won the First Prize of Zhejiang Provincial Science and Technology Progress Award. We are currently advancing the training of a trillion-parameter large model on a domestic 10,000-card cluster, which is expected to become a benchmark achievement on domestic computing power and an independent ecosystem.
5. Standardization
Led the release of 3 international standards for AI cloud platforms at the International Telecommunication Union (ITU), forming a complete standard system covering technical architecture, training processes, and performance evaluation:
- ITU-T F.748.38: Technical Specification for Artificial Intelligence Cloud Platform – General Architecture
- ITU-T F.748.17: Technical Specification for Artificial Intelligence Cloud Platform – AI Model Development
- ITU-T F.748.26: Technical Specification for Artificial Intelligence Cloud Platform – Performance Evaluation
Awards & Honors
Leading Science and Technology Award
First Class
Second Class