杨非

Fei Yang

Ph.D. | Associate Professor

Deputy Director
High-Performance Computing Facility Research Center
Zhejiang Lab

About Me

My research focuses on distributed computing theory and machine learning systems. I have published over 40 papers in top conferences and journals, been granted more than 40 invention patents, and led the release of 3 ITU international standards. My work has received the World Internet Conference Leading Science and Technology Award, Zhejiang Provincial Science and Technology Progress Award (First and Second Class), and has achieved significant application results in the construction of domestic AI computing clusters.

Distributed Computing ML Systems Large Model Training Heterogeneous Clusters Quantization & Compression Domestic Chip Adaptation Standardization

Biography

2020 - Present
Deputy Director, High-Performance Computing Facility Research Center
Associate Professor
2019 - 2020
Research Fellow
Singapore
Advisor: Prof. Yang Liu
2014 - 2018
Ph.D.
Netherlands
Ph.D. Advisors: Prof. Jos Baeten, Prof. Bas Luttik
2011 - 2014
M.S.
Shanghai, China
M.S. Advisor: Prof. Yuxi Fu
2007 - 2011
B.Eng.
Shanghai, China

Research Contributions

1. Distributed Computing Theory and System Architecture Innovation

Addressing the challenge of decoupling logical representation from physical resources in distributed training, I proposed the SBP (Split-Batch-Parallel) abstraction for distributed training, which serves as the core mathematical foundation for the domestic deep learning framework OneFlow. This theoretical innovation enables unified expression of operator-level parallel strategies and automatic derivation of data routing, mathematically solving the separation problem between distributed logical views and physical device views. It enabled OneFlow to become a performance-leading distributed framework and a representative domestic alternative to PyTorch at the time. This work provides an important theoretical basis for the formal study of distributed training systems, and its design ideas have been adopted by multiple domestic frameworks.

2. Heterogeneous Cluster Training Framework and Automatic Parallelism

Addressing the challenges of complex heterogeneous network environments and difficult parallel strategy adaptation in domestic clusters, I developed Holmes, the first distributed training framework adapted for heterogeneous network environments. We broke through non-uniform partitioning of 3D parallel strategies and automatic routing technologies for heterogeneous networks, achieving 1.4x performance improvement over mainstream frameworks like Megatron in hybrid network environments. This work became a pioneering study in heterogeneous hybrid training on domestic clusters, published at ICPP, the top conference in distributed computing. Furthermore, we proposed AutoHAAP, an automatic parallel framework for heterogeneous clusters, achieving automatic parallel strategy optimization based on cost models and heuristic search, supporting minute-level search at the 10,000-card scale. This work was published at HPCA, the top conference in computer architecture, and is at the leading level in the field of heterogeneous distributed training.

3. Computational Efficiency Optimization and Domestic Chip Adaptation

Focusing on the core challenges of diverse numerical formats and complex distribution characteristics of domestic chips, I systematically conducted research on hardware-aware compression techniques. We proposed a series of tensor distribution-sensitive compression methods including FlattenQuant (for large language models), ADFQ-ViT (for vision models), and PTQ4PLM (for protein language models), effectively solving the problem of outlier suppression and quantization accuracy loss. Our results have been published in top conferences and journals including ICML, Neural Networks, and LREC-Coling. Through cross-team collaboration, we applied these technologies to the adaptation and optimization of 236B MoE language models, 72B geoscience models, 70B astronomy models, and 13B protein models on domestic chips, achieving 40% MFU for MoE model training on a domestic 1,000-card cluster, which reaches the industry-leading level. As subsystem leader, our work "Aggregated Heterogeneous Computing Power Intelligent Computing System" won the Second Prize of Zhejiang Provincial Science and Technology Progress Award.

4. Large-Scale System Integration and Engineering Practice

As technical lead, I led the development of the large model training and inference system in the Nanhu Computing Framework, leading the team to breakthrough multiple key efficiency optimization technologies including mixed-precision training, asynchronous checkpointing, data splicing compression, and deep bottleneck analysis. This supported the construction and validation of the basic 10,000-card and hybrid 10,000-card clusters at Zhejiang's New Computing Power Center. The research results have been applied to multiple scientific large model training tasks at Zhejiang Laboratory, including 021, geoscience, astronomy, and biology, and the related technologies have saved nearly 100 million yuan in computing costs, achieving significant social and economic benefits. The "Nanhu Computing Framework" I participated in won the World Internet Conference Leading Science and Technology Award, and "Key Technologies and Applications of Intelligent Computing for Large Language Model Training" won the First Prize of Zhejiang Provincial Science and Technology Progress Award. We are currently advancing the training of a trillion-parameter large model on a domestic 10,000-card cluster, which is expected to become a benchmark achievement on domestic computing power and an independent ecosystem.

5. Standardization

Led the release of 3 international standards for AI cloud platforms at the International Telecommunication Union (ITU), forming a complete standard system covering technical architecture, training processes, and performance evaluation:

  • ITU-T F.748.38: Technical Specification for Artificial Intelligence Cloud Platform – General Architecture
  • ITU-T F.748.17: Technical Specification for Artificial Intelligence Cloud Platform – AI Model Development
  • ITU-T F.748.26: Technical Specification for Artificial Intelligence Cloud Platform – Performance Evaluation

Awards & Honors

World Internet Conference
Leading Science and Technology Award
Nanhu Computing Framework
Zhejiang Provincial Science and Technology Progress Award
First Class
Key Technologies and Applications of Intelligent Computing for Large Language Model Training
Zhejiang Provincial Science and Technology Progress Award
Second Class
Aggregated Heterogeneous Computing Power Intelligent Computing System

Selected Publications

View Full Publication List →

OptiCo: Adaptive Distributed Training Optimization via Collaborative Agent Reasoning
Sheng Chen, Zhe Tang, Weixing Zhang, Fei Yang, Yuanyuan Wang, Tianlin Li, Yang Liu
ACL26 CCF A Corresponding
AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training
Yuanyuan Wang, Nana Tang, Yuyang Wang, Shu Pan, Dingding Yu,Zeyue Wang,Mou Sun, Kejie Fu, Fangyu Wang, Yunchuan Chen, Ning Sun, Fei Yang
HPCA26 CCF A Corresponding
LaRA: Layer-wise Rank Allocation for Efficient Fine-tuning of Pruned Large Language Models
Yuhua Zhou, Changhai Zhou, Shiyang Zhang, Fei Yang, Yi Zhang, Aimin Pan
IPM26 CCF A Corresponding
Zip Your Data: Length-Adaptive Visual Token Optimization for Efficient Multi-Modal Training
Ning Sun, Fangwen Wu, Yi Zhang, Si Chen, Xiuting Tao, Fei Yang
ICASSP26 CCF B Corresponding
HOC: Hierarchical Overlapped Communication Optimization for Parallelism in Distributed Training
Zeyue Wang, Yuanyuan Wang, Shu Pan, Yuyang Wang, Nana Tang, Fei Yang
ISCAS26 CCF B Corresponding
Deputy: Accelerating Large Language Model Inference with Dynamic Low-Rank Substitution
Yuhua Zhou, Shichao Weng, Changhai Zhou, Yuhan Wu, Qian Qiao, Jun Gao, Fei Yang, Aimin Pan
ACL Findings 26 Corresponding
Parameter-Efficient Fine-Tuning in Large Models: A Survey of Methodologies
Luping Wang, Sheng Chen, Linnan Jiang, Shu Pan, Runze Cai, Sen Yang, Fei Yang
Artificial Intelligence Review 中科院一区 Corresponding
BSLoRA: Enhancing the Parameter Efficiency of LoRA with Intra-Layer and Inter-Layer Sharing
Yuhua Zhou, Ruifeng Li, Changhai Zhou, Fei Yang, Aimin Pan
ICML25 CCF A Corresponding
On the performance and memory footprint of distributed training: An empirical study on transformers
Zhengxian Lu, Fangyu Wang, Zhiwei Xu , Fei Yang, Tao Li.
Software:Practice and Experience. CCF B Corresponding
ADFQ-ViT: Activation-Distribution-Friendly post-training Quantization for Vision Transformers
Yanfeng Jiang, Ning Sun, Xueshuo Xie, Fei Yang, Tao Li.
Neural Networks 中科院二区 Corresponding
Robust and Privacy-Preserving Collaborative Learning: A Comprehensive Survey
Fei Yang, Xu Zhang, Shangwei Guo, Daiyuan Chen, Tianwei Zhang, Yan Gan, Tao Xiang, Yang Liu.
Artificial Intelligence Review 中科院一区 First Author
Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment
Fei Yang, Shuang Peng, Ning Sun, Fangyu Wang, Yuanyuan Wang, Fu Wu, Jiezhong Qiu, Aimn Pan
ICPP24 CCF B First Author
Wi-Fi based Gait Recognition using Spectrogram and Phase
Sheng Chen, Fei Yang, Aimin Pan, Zhewei Mei.
ICME24 CCF B Corresponding
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
Yi Zhang, Fei Yang, Shuang Peng, Fangyu Wang, Aimin Pan.
LREC-Coling24 CCF B Corresponding
Exploring Post-Training Quantization of Protein Language Models
Shuang Peng, Fei Yang, Ning Sun, Sheng Chen, Yanfeng Jiang, Aimin Pan.
BIBM23 CCF B Corresponding