|
Zhiwei Bai (白志威)
|
Ph.D. Student,
School of Mathematical Sciences,
Institute of Natural Sciences,
Shanghai Jiao Tong University,
Shanghai, China
E-mail: bai299@sjtu.edu.cn
|
About Me
I am currently a doctoral candidate at the School of Mathematical Sciences and the Institute of Natural
Sciences, Shanghai Jiao Tong University, where I am pursuing a Ph.D. in Applied Mathematics. I received my B.S. in Mathematics and Applied Mathematics from Shanghai Jiao Tong University in 2022, where I studied in the Wu Wenjun Honors Program in Mathematics.
Under the mentorship of Professors
Yaoyu Zhang and
Zhi-Qin John Xu,
my research focuses primarily on the theoretical foundations of machine learning and deep learning.
I am driven by a deep curiosity about the mathematical principles behind modern artificial intelligence,
with a particular interest in using a phenomenon-driven approach to understand the complexity
that emerges during the actual training of deep learning models.
I am especially interested in developing a bottom-up, first-principles understanding of interesting phenomena
in deep learning. My research spans neural network loss landscapes, condensation phenomena, implicit
regularization, generalization in overparameterized regimes, Adam dynamics, and the reasoning preferences
of large language models.
In one sentence: I try to understand how data, architecture, and optimization dynamics shape what deep neural networks actually learn.
Research Interests
My research areas include:
Current Works
-
Critical Points Analysis of Loss Landscapes in Deep Neural Networks
-
Training Dynamics and Induced Implicit Regularization in Matrix Factorization Models
-
Theory of Optimistic Sample Estimation for Nonlinear Models
-
Loss Spike Analysis in Adam Training
-
Complexity Control for Improving Reasoning Ability in LLMs
-
Self-Adjoint Functors in the Category of Nilpotent Morphisms
Recent Publications
-
Zhiwei Bai, Tao Luo, Zhi-Qin John Xu*, Yaoyu Zhang*, "Embedding Principle in Depth for the Loss
Landscape Analysis of Deep Neural Networks", CSIAM Transactions on Applied Mathematics, 5(2):350-389,
2024. ISSN 2708-0579. doi: https://doi.org/10.4208/csiam-am.SO-2023-0020.
[pdf] and on [arXiv]
TL;DR: Adding layers does not erase a shallower neural network's critical points:
each can be lifted to a deeper network without changing its outputs. Theory and experiments show these
lifted degenerate points can shape training across depths and suggest a principled route to layer
pruning.
-
Zhiwei Bai, Jiajie Zhao, Yaoyu Zhang*, "Connectivity shapes implicit regularization in matrix
factorization models for matrix completion", In The Thirty-eighth Annual Conference on Neural
Information Processing Systems, volume 37, pages 45914-45955, 2024. [pdf] and on [arXiv]
TL;DR: The pattern of observed entries changes which completion a matrix-factorization
model prefers: connected data tends to produce the lowest-rank solution, while some disconnected
patterns produce the minimum nuclear-norm solution. A hierarchy of invariant manifolds explains how
training moves through solutions of increasing rank.
-
Zhiwei Bai#, Jiajie Zhao#, Zhangchen Zhou, Zhi-Qin John Xu*, Yaoyu Zhang*, "Towards Understanding
Adam Convergence on Highly Degenerate Polynomials", Forty-Third International Conference on Machine
Learning, 2026 [pdf] and on [arXiv]
TL;DR: For highly degenerate polynomials (i.e., with vanishing curvature), Adam can converge at a linear rate under constant
learning rates, unlike the sub-linear rates of gradient descent and momentum.
Decoupling between second moments and squared gradients compensates for the vanishing curvature, while hyperparameters
separate stable, spiking, and oscillatory regimes.
-
Zhiwei Bai#, Zhangchen Zhou#, Jiajie Zhao, Xiaolong Li, Zhiyu Li, Feiyu Xiong, Hongkang Yang,
Yaoyu Zhang*, Zhi-Qin John Xu*, "Adaptive Preconditioners Trigger Loss Spikes in Adam",
Forty-Third International Conference on Machine Learning, 2026 [pdf] and on [arXiv]
TL;DR: Why do Adam runs suddenly lose stability? We show that adaptive
preconditioners can decay out of sync with gradients, pushing the effective curvature past the stability threshold;
alignment with the gradient direction then produces a spike. Experiments confirm this mechanism across
MLP, convolutional, and Transformer networks.
-
Yaoyu Zhang*, Leyang Zhang, Zhongwang Zhang, Zhiwei Bai, "Local Linear Recovery Guarantee of Deep
Neural Networks at Overparameterization", Journal of Machine Learning Research 26 (69), 1-30,
2025. [pdf] and on [arXiv]
TL;DR: How many samples does an overparameterized network necessarily need to recover a target?
We introduce local linear recovery theory, a tractable best-case guarantee, and prove narrower-network targets
can meet it with fewer samples than a wider model's parameter count. The best-case scenario provides a theoretical
foundation for understanding condensation benefits in deep learning.
-
Yaoyu Zhang*, Zhongwang Zhang, Leyang Zhang, Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, "Optimistic
Estimate Uncovers the Potential of Nonlinear Models", Journal of Machine Learning, 2025. [pdf] and on [arXiv]
TL;DR: How can we estimate what an overparameterized nonlinear model could fit in
the best case? We define an optimistic sample size tied to a target's effective model rank. Experiments
support two network patterns: width need not raise it, while unnecessary connections are costly.
-
Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang*, "Gradient Flow Dynamics and
Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization", Forty-Third
International Conference on Machine Learning, 2026. [pdf]
TL;DR: How does infinitesimal initialization shape learning in deep diagonal linear
networks? We show gradient-flow training moves through successive saddle points reproduced by a recursive
algorithm, selecting a solution of a modified L1 problem. Structural invariant manifolds explain this
incremental feature selection.
-
Yaoyu Zhang*, Zhongwang Zhang, Leyang Zhang, Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, "Linear
Stability Hypothesis and Rank Stratification for Nonlinear Models", arXiv:2211.11623, 2022. [pdf] and on [arXiv]
TL;DR: Why can nonlinear models recover targets with fewer samples than parameters?
We define model rank as a function's effective parameter count, prove linear stability there, and
hypothesize training favors stable interpolations. Experiments on matrix factorization and neural
networks show recovery transitions near this rank.
-
Zhiwei Bai, Xiang Cao, Songtao Mao, Han Zhang, Yuehui Zhang*, "Nilpotent Category of Abelian
Categories and Self-Adjoint Functors", Frontiers of Mathematics 18 (6), 1363-1377, 2023. [pdf] and on [arXiv]
TL;DR: What does a category of nilpotent operators remember, and which functors are
self-adjoint? We show Nil(C) is abelian when C is, and two abelian categories are equivalent exactly when
their nilpotent categories are. Over finite-dimensional vector spaces, self-adjoint functors are Hom or
Tensor.
-
Jiajie Zhao, Zhiwei Bai, Yaoyu Zhang*, "Disentangling Sample Size and Initialization Effects on
Perfect Generalization for Single-Neuron Targets", arXiv:2405.13787, 2024. [pdf] and on [arXiv]
TL;DR: Can a network recover a single-neuron target from few samples? Smaller
initialization helps, while an initial imbalance ratio shapes the path. Below the optimistic threshold
recovery is impossible; at it, only exceptional initializations work, while separation can make recovery
possible for a positive-measure set.
-
Liangkai Hang#, Junjie Yao#, Zhiwei Bai, Tianyi Chen, Yang Chen, Rongjie Diao, Hezhou Li, Pengxiao
Lin, Zhiwei Wang, Cheng Xu, Zhongwang Zhang*, Zhangchen Zhou, Zhiyu Li, Zehao Lin, Kai Chen, Feiyu Xiong*,
Yaoyu Zhang*, Weinan E*, Hongkang Yang*, Zhi-Qin John Xu*, "Scalable Complexity Control Facilitates
Reasoning Ability of LLMs", arXiv:2505.23013, 2025 [pdf] and on [arXiv]
TL;DR: Can lower model complexity help language models reason? Across model and
data scales we find controlling initialization rate and weight decay improves scaling curves and many
benchmark scores, especially reasoning tasks. Analyses link lower complexity to condensed, structured
representations.
Note: * indicates the corresponding author; # indicates equal contribution.
Full list of publications on Google
Scholar.
Projects
Academic Service
Academic Conferences
-
Presented a oral talk titled "Understanding Implicit Regularization in Matrix Factorization via Data
Connectivity and Hierarchical Invariant Manifold Dynamics" at the Shanghai SIAM Artificial Intelligence
Fundamentals Workshop, Shanghai, China, September
2025.
-
Presented a oral talk titled "Connectivity shapes implicit regularization in matrix factorization models
for matrix completion" at the Shanghai Jiao Tong University AI for math Workshop, Shanghai, China, May
2025.
-
Presented a poster titled "Connectivity shapes implicit regularization in matrix factorization models for
matrix completion" at the Thirty-eighth Annual Conference on Neural Information Processing Systems,
Vancouver, Canada, December 2024.
-
Presented a talk titled "Connectivity shapes implicit regularization in matrix factorization models for
matrix completion" at the NUS-SJTU PhD Forum, Singapore, Singapore, November 2024.
-
Presented a talk titled "Connectivity shapes implicit regularization in matrix factorization models for
matrix completion" at the Peking University PhD Forum, Beijing, China, November 2024.
-
Presented a poster titled "Connectivity shapes implicit regularization in matrix factorization models for
matrix completion" at the The 22-nd Annual Meeting of the Chinese Society for Industrial and Applied
Mathematics, Student Forum, Nanjing, China, October 2024.
-
Presented a talk titled "Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural
Networks" at the 2024 Scientific Machine Learning Conference (CSML2024), Shanghai, China, August 2024.
-
Volunteer Manager for the 2024 Scientific Machine Learning Conference (CSML2024) [CSML2024].
-
Presented a talk titled "Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural
Networks" at the 2023 Machine Learning and Materials Science Workshop, Shanghai, China, July 2023.
Teaching
-
Fall 2025: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan
Honors Program.
-
Spring 2025: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan
Honors Program.
-
Fall 2024: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors
Program.
-
Spring 2024: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan
Honors Program.
-
Fall 2023: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors
Program.
-
Spring 2023: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan
Honors Program.
-
Fall 2022: Teaching Assistant, Mathematical Analysis, Shanghai Jiao Tong University Zhiyuan Honors
Program.
Education
2022–present, Ph.D. in Mathematics, School of Mathematical Sciences, Shanghai Jiao Tong University, China.
2018–2022, B.S. in Mathematics and Applied Mathematics (Wu Wenjun Honors Class), School of Mathematical
Sciences, Shanghai Jiao Tong University, China.
2021–2022, Minor in AI+X, a collaborative
program offered by: Shanghai Jiao Tong University, Zhejiang University, Fudan University, University of
Science and Technology of China, Nanjing University, and Tongji University, China.
Activities
- Peking University Graduate Applied Mathematics Workshop and Summer School, July–August 2023, Peking
University, Beijing, China.
- Deep Learning Theory and Application Summer School, July 2023, Shanghai Jiao Tong University, Shanghai,
China.
Competitions and Awards
-
National Scholarship of PhD, Ministry of Education of the People's Republic of China, 2025
-
Special PhD Fellowship of the Young Talents Support Program, Chinese Society for Science and
Technology(CAST), 2024
-
Best Paper Award, Shanghai Jiao Tong University AI for Math Workshop, 2025
-
Outstanding Poster Award, The 22-nd Annual Meeting of the Chinese Society for Industrial and Applied
Mathematics, 2024
-
National Scholarship, Ministry of Education of the People's Republic of China, 2021
-
Shanghai Outstanding Graduate, Shanghai Municipal Education Commission, 2022
-
Outstanding Bachelor's Degree Thesis(top1%), Shanghai Jiao Tong University, 2022
-
First Prize, National Undergraduate Mathematical Contest in Modeling, Chinese Society of Industrial and
Applied Mathematics, 2020
-
Second Prize, National Postgraduate Mathematical Contest in Modeling, Chinese Society for Academic
Degrees and Postgraduate Education, 2023
-
Huatai Securities Technology Scholarship, Shanghai Jiao Tong University, 2023
-
Samsung Scholarship, Shanghai Jiao Tong University, 2020
-
Third Prize, National College Mathematics Competition, Chinese Mathematical Society, 2019
-
President, Mathematical Modeling Association, Shanghai Jiao Tong University, 2020–2023
Contact
I am always open to connecting with fellow researchers and enthusiasts in related fields. Feel free to reach
out—let’s explore new ideas together!
|