Zhiwei Bai (白志威)

alt text 

Ph.D. Student,
School of Mathematical Sciences,
Institute of Natural Sciences,
Shanghai Jiao Tong University,
Shanghai, China
E-mail: bai299@sjtu.edu.cn

About Me

I am currently a doctoral candidate at the School of Mathematical Sciences and the Institute of Natural Sciences, Shanghai Jiao Tong University, where I am pursuing a Ph.D. in Applied Mathematics. I received my B.S. in Mathematics and Applied Mathematics from Shanghai Jiao Tong University in 2022, where I studied in the Wu Wenjun Honors Program in Mathematics.

Under the mentorship of Professors Yaoyu Zhang and Zhi-Qin John Xu, my research focuses primarily on the theoretical foundations of machine learning and deep learning. I am driven by a deep curiosity about the mathematical principles behind modern artificial intelligence, with a particular interest in using a phenomenon-driven approach to understand the complexity that emerges during the actual training of deep learning models.

I am especially interested in developing a bottom-up, first-principles understanding of interesting phenomena in deep learning. My research spans neural network loss landscapes, condensation phenomena, implicit regularization, generalization in overparameterized regimes, Adam dynamics, and the reasoning preferences of large language models.

In one sentence: I try to understand how data, architecture, and optimization dynamics shape what deep neural networks actually learn.

Research Interests

My research areas include:

  • Machine Learning and Deep Learning Theory

  • Large Language Models

  • Brain-Inspired Computing

  • Algebra

Current Works

  • Critical Points Analysis of Loss Landscapes in Deep Neural Networks

  • Training Dynamics and Induced Implicit Regularization in Matrix Factorization Models

  • Theory of Optimistic Sample Estimation for Nonlinear Models

  • Loss Spike Analysis in Adam Training

  • Complexity Control for Improving Reasoning Ability in LLMs

  • Self-Adjoint Functors in the Category of Nilpotent Morphisms

Recent Publications

  1. Zhiwei Bai, Tao Luo, Zhi-Qin John Xu*, Yaoyu Zhang*, "Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks", CSIAM Transactions on Applied Mathematics, 5(2):350-389, 2024. ISSN 2708-0579. doi: https://doi.org/10.4208/csiam-am.SO-2023-0020. [pdf] and on [arXiv]

    TL;DR: Adding layers does not erase a shallower neural network's critical points: each can be lifted to a deeper network without changing its outputs. Theory and experiments show these lifted degenerate points can shape training across depths and suggest a principled route to layer pruning.

  2. Zhiwei Bai, Jiajie Zhao, Yaoyu Zhang*, "Connectivity shapes implicit regularization in matrix factorization models for matrix completion", In The Thirty-eighth Annual Conference on Neural Information Processing Systems, volume 37, pages 45914-45955, 2024. [pdf] and on [arXiv]

    TL;DR: The pattern of observed entries changes which completion a matrix-factorization model prefers: connected data tends to produce the lowest-rank solution, while some disconnected patterns produce the minimum nuclear-norm solution. A hierarchy of invariant manifolds explains how training moves through solutions of increasing rank.

  3. Zhiwei Bai#, Jiajie Zhao#, Zhangchen Zhou, Zhi-Qin John Xu*, Yaoyu Zhang*, "Towards Understanding Adam Convergence on Highly Degenerate Polynomials", Forty-Third International Conference on Machine Learning, 2026 [pdf] and on [arXiv]

    TL;DR: For highly degenerate polynomials (i.e., with vanishing curvature), Adam can converge at a linear rate under constant learning rates, unlike the sub-linear rates of gradient descent and momentum. Decoupling between second moments and squared gradients compensates for the vanishing curvature, while hyperparameters separate stable, spiking, and oscillatory regimes.

  4. Zhiwei Bai#, Zhangchen Zhou#, Jiajie Zhao, Xiaolong Li, Zhiyu Li, Feiyu Xiong, Hongkang Yang, Yaoyu Zhang*, Zhi-Qin John Xu*, "Adaptive Preconditioners Trigger Loss Spikes in Adam", Forty-Third International Conference on Machine Learning, 2026 [pdf] and on [arXiv]

    TL;DR: Why do Adam runs suddenly lose stability? We show that adaptive preconditioners can decay out of sync with gradients, pushing the effective curvature past the stability threshold; alignment with the gradient direction then produces a spike. Experiments confirm this mechanism across MLP, convolutional, and Transformer networks.

  5. Yaoyu Zhang*, Leyang Zhang, Zhongwang Zhang, Zhiwei Bai, "Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization", Journal of Machine Learning Research 26 (69), 1-30, 2025. [pdf] and on [arXiv]

    TL;DR: How many samples does an overparameterized network necessarily need to recover a target? We introduce local linear recovery theory, a tractable best-case guarantee, and prove narrower-network targets can meet it with fewer samples than a wider model's parameter count. The best-case scenario provides a theoretical foundation for understanding condensation benefits in deep learning.

  6. Yaoyu Zhang*, Zhongwang Zhang, Leyang Zhang, Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, "Optimistic Estimate Uncovers the Potential of Nonlinear Models", Journal of Machine Learning, 2025. [pdf] and on [arXiv]

    TL;DR: How can we estimate what an overparameterized nonlinear model could fit in the best case? We define an optimistic sample size tied to a target's effective model rank. Experiments support two network patterns: width need not raise it, while unnecessary connections are costly.

  7. Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang*, "Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization", Forty-Third International Conference on Machine Learning, 2026. [pdf]

    TL;DR: How does infinitesimal initialization shape learning in deep diagonal linear networks? We show gradient-flow training moves through successive saddle points reproduced by a recursive algorithm, selecting a solution of a modified L1 problem. Structural invariant manifolds explain this incremental feature selection.

  8. Yaoyu Zhang*, Zhongwang Zhang, Leyang Zhang, Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, "Linear Stability Hypothesis and Rank Stratification for Nonlinear Models", arXiv:2211.11623, 2022. [pdf] and on [arXiv]

    TL;DR: Why can nonlinear models recover targets with fewer samples than parameters? We define model rank as a function's effective parameter count, prove linear stability there, and hypothesize training favors stable interpolations. Experiments on matrix factorization and neural networks show recovery transitions near this rank.

  9. Zhiwei Bai, Xiang Cao, Songtao Mao, Han Zhang, Yuehui Zhang*, "Nilpotent Category of Abelian Categories and Self-Adjoint Functors", Frontiers of Mathematics 18 (6), 1363-1377, 2023. [pdf] and on [arXiv]

    TL;DR: What does a category of nilpotent operators remember, and which functors are self-adjoint? We show Nil(C) is abelian when C is, and two abelian categories are equivalent exactly when their nilpotent categories are. Over finite-dimensional vector spaces, self-adjoint functors are Hom or Tensor.

  10. Jiajie Zhao, Zhiwei Bai, Yaoyu Zhang*, "Disentangling Sample Size and Initialization Effects on Perfect Generalization for Single-Neuron Targets", arXiv:2405.13787, 2024. [pdf] and on [arXiv]

    TL;DR: Can a network recover a single-neuron target from few samples? Smaller initialization helps, while an initial imbalance ratio shapes the path. Below the optimistic threshold recovery is impossible; at it, only exceptional initializations work, while separation can make recovery possible for a positive-measure set.

  11. Liangkai Hang#, Junjie Yao#, Zhiwei Bai, Tianyi Chen, Yang Chen, Rongjie Diao, Hezhou Li, Pengxiao Lin, Zhiwei Wang, Cheng Xu, Zhongwang Zhang*, Zhangchen Zhou, Zhiyu Li, Zehao Lin, Kai Chen, Feiyu Xiong*, Yaoyu Zhang*, Weinan E*, Hongkang Yang*, Zhi-Qin John Xu*, "Scalable Complexity Control Facilitates Reasoning Ability of LLMs", arXiv:2505.23013, 2025 [pdf] and on [arXiv]

    TL;DR: Can lower model complexity help language models reason? Across model and data scales we find controlling initialization rate and weight decay improves scaling curves and many benchmark scores, especially reasoning tasks. Analyses link lower complexity to condensed, structured representations.

Note: * indicates the corresponding author; # indicates equal contribution.

Full list of publications on Google Scholar.

Projects

  • Book (Chinese)《深度学习现象导论:从感知机到大模型》

    • Collaborating with supervisors Zhi-Qin John Xu and Yaoyu Zhang.

    • This book introduces fundamental concepts of deep learning with a phenomenon-driven approach. Please see the project on GitHub.

    • Your feedback is welcome!

Academic Service

Academic Conferences

  • Presented a oral talk titled "Understanding Implicit Regularization in Matrix Factorization via Data Connectivity and Hierarchical Invariant Manifold Dynamics" at the Shanghai SIAM Artificial Intelligence Fundamentals Workshop, Shanghai, China, September 2025.

  • Presented a oral talk titled "Connectivity shapes implicit regularization in matrix factorization models for matrix completion" at the Shanghai Jiao Tong University AI for math Workshop, Shanghai, China, May 2025.

  • Presented a poster titled "Connectivity shapes implicit regularization in matrix factorization models for matrix completion" at the Thirty-eighth Annual Conference on Neural Information Processing Systems, Vancouver, Canada, December 2024.

  • Presented a talk titled "Connectivity shapes implicit regularization in matrix factorization models for matrix completion" at the NUS-SJTU PhD Forum, Singapore, Singapore, November 2024.

  • Presented a talk titled "Connectivity shapes implicit regularization in matrix factorization models for matrix completion" at the Peking University PhD Forum, Beijing, China, November 2024.

  • Presented a poster titled "Connectivity shapes implicit regularization in matrix factorization models for matrix completion" at the The 22-nd Annual Meeting of the Chinese Society for Industrial and Applied Mathematics, Student Forum, Nanjing, China, October 2024.

  • Presented a talk titled "Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks" at the 2024 Scientific Machine Learning Conference (CSML2024), Shanghai, China, August 2024.

  • Volunteer Manager for the 2024 Scientific Machine Learning Conference (CSML2024) [CSML2024].

  • Presented a talk titled "Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks" at the 2023 Machine Learning and Materials Science Workshop, Shanghai, China, July 2023.

Teaching

  • Fall 2025: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

  • Spring 2025: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

  • Fall 2024: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

  • Spring 2024: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

  • Fall 2023: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

  • Spring 2023: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

  • Fall 2022: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors Program.

Education

2022–present, Ph.D. in Mathematics, School of Mathematical Sciences, Shanghai Jiao Tong University, China.

2018–2022, B.S. in Mathematics and Applied Mathematics (Wu Wenjun Honors Class), School of Mathematical Sciences, Shanghai Jiao Tong University, China.

2021–2022, Minor in AI+X, a collaborative program offered by: Shanghai Jiao Tong University, Zhejiang University, Fudan University, University of Science and Technology of China, Nanjing University, and Tongji University, China.

Activities

  • Peking University Graduate Applied Mathematics Workshop and Summer School, July–August 2023, Peking University, Beijing, China.
  • Deep Learning Theory and Application Summer School, July 2023, Shanghai Jiao Tong University, Shanghai, China.

Competitions and Awards

  • National Scholarship of PhD, Ministry of Education of the People's Republic of China, 2025

  • Special PhD Fellowship of the Young Talents Support Program, Chinese Society for Science and Technology(CAST), 2024

  • Best Paper Award, Shanghai Jiao Tong University AI for Math Workshop, 2025

  • Outstanding Poster Award, The 22-nd Annual Meeting of the Chinese Society for Industrial and Applied Mathematics, 2024

  • National Scholarship, Ministry of Education of the People's Republic of China, 2021

  • Shanghai Outstanding Graduate, Shanghai Municipal Education Commission, 2022

  • Outstanding Bachelor's Degree Thesis(top1%), Shanghai Jiao Tong University, 2022

  • First Prize, National Undergraduate Mathematical Contest in Modeling, Chinese Society of Industrial and Applied Mathematics, 2020

  • Second Prize, National Postgraduate Mathematical Contest in Modeling, Chinese Society for Academic Degrees and Postgraduate Education, 2023

  • Huatai Securities Technology Scholarship, Shanghai Jiao Tong University, 2023

  • Samsung Scholarship, Shanghai Jiao Tong University, 2020

  • Third Prize, National College Mathematics Competition, Chinese Mathematical Society, 2019

  • President, Mathematical Modeling Association, Shanghai Jiao Tong University, 2020–2023

Contact

I am always open to connecting with fellow researchers and enthusiasts in related fields. Feel free to reach out—let’s explore new ideas together!