References
Abadi, Martı́n, Paul Barham, Jianmin Chen, et
al. 2016. “TensorFlow: A System for
Large-Scale Machine Learning.” 12th USENIX
Symposium on Operating Systems
Design and Implementation (OSDI
16), 265–83. https://www.usenix.org/conference/osdi16/technical-sessions/presentation/abadi.
Abdel-Hamid, Ossama, Abdel-Rahman Mohamed, Hui Jiang, Li Deng, Gerald
Penn, and Dong Yu. 2014. “Convolutional Neural Networks for Speech
Recognition.” IEEE/ACM Transactions
on Audio, Speech, and Language
Processing 22 (10): 1533–45. https://doi.org/10.1109/taslp.2014.2339736.
Ainslie, Joshua, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy,
Federico Lebrón, and Sumit Sanghai. 2023. “GQA:
Training Generalized Multi-Query Transformer Models from Multi-Head
Checkpoints.” Proceedings of the 2023 Conference on Empirical
Methods in Natural Language Processing, 4895–901.
Akiba, T., S. Sano, T. Yanase, T. Ohta, and M. Koyama. 2019.
“Optuna: A Next-Generation Hyperparameter
Optimization Framework.” Proceedings of the 25th
ACM SIGKDD International
Conference on Knowledge Discovery
& Data Mining. https://doi.org/10.1145/3292500.3330701.
Alayrac, Jean-Baptiste, Jeff Donahue, Pauline Luc,
et al. 2022. “Flamingo: A Visual Language Model for
Few-Shot Learning.” ArXiv:2204.14198. https://arxiv.org/abs/2204.14198.
Albergo, Michael S., Nicholas M. Boffi, and Eric Vanden-Eijnden. 2023.
“Stochastic Interpolants: A Unifying Framework for Flows and
Diffusions.” arXiv Preprint arXiv:2303.08797.
Alemi, Alexander A., Ian Fischer, Joshua V. Dillon, and Kevin Murphy.
2017. “Deep Variational Information Bottleneck.”
International Conference on Learning Representations.
Alsallakh, Bilal, Narine Kokhlikyan, Vivek Miglani, Jun Yuan, and Orion
Reblitz-Richardson. 2020. “Mind the PAD –
CNNs Can Develop Blind Spots.”
ArXiv:2010.02178. https://arxiv.org/abs/2010.02178.
Amari, Shun-ichi. 1998. “Natural Gradient Works Efficiently in
Learning.” Neural Computation 10 (2): 251–76.
Amari, Shun-ichi. 2016. Information Geometry and Its
Applications. Springer.
Anderson, Brian D. O. 1982. “Reverse-Time Diffusion Equation
Models.” Stochastic Processes and Their Applications 12
(3): 313–26.
Anil, Rohan, Andrew M Dai, Orhan Firat, et
al. 2023. “PaLM 2 Technical
Report.” ArXiv:2305.10403. https://arxiv.org/abs/2305.10403.
Anil, Rohan, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer.
2020. “Scalable Second-Order Optimization for Deep
Learning.” ArXiv:2002.09018. https://arxiv.org/abs/2002.09018.
Ansel, Jason, Edward Yang, Horace He, et al.
2024. “PyTorch 2: Faster Machine
Learning Through Dynamic Python Bytecode Transformation and
Graph Compilation.” 29th ACM
International Conference on
Architectural Support for
Programming Languages and
Operating Systems (ASPLOS),
929–47. https://doi.org/10.1145/3620665.3640366.
Arjevani, Yossi, Yair Carmon, John C. Duchi, Dylan J. Foster, Nathan
Srebro, and Blake Woodworth. 2023. “Lower Bounds for Non-Convex
Stochastic Optimization.” Mathematical Programming 199:
165–214.
Arjovsky, Martin, Soumith Chintala, and Léon Bottou. 2017.
“Wasserstein Generative Adversarial Networks.”
Proceedings of the 34th International Conference on Machine
Learning, 214–23.
Armijo, Larry. 1966. “Minimization of Functions Having
Lipschitz Continuous First Partial Derivatives.”
Pacific Journal of Mathematics 16 (1): 1–3.
Aronszajn, Nachman. 1950. “Theory of reproducing kernels.” Transactions of
the American Mathematical
Society 68 (3): 337–404. https://doi.org/10.1090/s0002-9947-1950-0051437-7.
Arora, Simran, Sabri Eyuboglu, Aman Timalsina, et al. 2024.
“Zoology: Measuring and Improving Recall in Efficient
Language Models.” International Conference on Learning
Representations. https://arxiv.org/abs/2312.04927.
Arora, Simran, Sabri Eyuboglu, Michael Zhang, et al. 2024. “Simple
Linear Attention Language Models Balance the Recall-Throughput
Tradeoff.” International Conference on Machine Learning.
Arpit, Devansh, Stanisław Jastrzȩbski, Nicolas Ballas, et al. 2017.
“A Closer Look at Memorization in Deep Networks.”
International Conference on Machine
Learning, 233–42.
Åström, Karl Johan, and Richard M. Murray. 2021. Feedback Systems:
An Introduction for Scientists and Engineers. 2nd ed. Princeton
University Press.
Austin, Jacob, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne
van den Berg. 2021. “Structured Denoising Diffusion Models in
Discrete State-Spaces.” Advances in Neural Information
Processing Systems 34.
Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016.
“Layer Normalization.”
ArXiv:1607.06450. https://arxiv.org/abs/1607.06450.
Bae, Sangmin, Bilge Acun, Chien-Yu Lin, et al. 2025. “Hybrid
Architectures for Language Models: Systematic Analysis and Design
Insights.” arXiv Preprint arXiv:2510.04800.
Baevski, Alexei, and Michael Auli. 2018. “Adaptive Input
Representations for Neural Language Modeling.”
International Conference on Learning
Representations. https://openreview.net/forum?id=ByxZX20qFQ.
Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio. 2015. “Neural
Machine Translation by Jointly Learning to Align and Translate.”
International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1409.0473.
Bai, Shaojie, J. Zico Kolter, and Vladlen Koltun. 2019. “Deep
Equilibrium Models.” Advances in Neural Information
Processing Systems 32.
Baptista, R., and M. Poloczek. 2018. “Bayesian
Optimization of Combinatorial Structures.” Proceedings of the
35th International Conference on
Machine Learning. https://proceedings.mlr.press/v80/baptista18a.html.
Barber, David, and Felix Agakov. 2003. “The IM
Algorithm: A Variational Approach to Information Maximization.”
Advances in Neural Information Processing Systems 16.
Bardenet, R., M. Brendel, B. Kégl, and M. Sebag. 2013.
“Collaborative Hyperparameter Tuning.” Proceedings of
the 30th International Conference on
Machine Learning (ICML’13).
https://proceedings.mlr.press/v28/bardenet13.html.
Bartlett, Peter L., Philip M. Long, Gábor Lugosi, and Alexander Tsigler.
2020. “Benign Overfitting in Linear Regression.”
Proceedings of the National Academy of Sciences 117 (48):
30063–70.
Bartlett, Peter L., and Shahar Mendelson. 2002. “Rademacher and
Gaussian Complexities: Risk Bounds and Structural
Results.” Journal of Machine Learning Research 3:
463–82.
Bartlett, Peter L, Andrea Montanari, and Alexander Rakhlin. 2021.
“Deep Learning: A Statistical Viewpoint.”
ArXiv:2103.09177. https://arxiv.org/abs/2103.09177.
Baur, Walter, and Volker Strassen. 1983. “The Complexity of
Partial Derivatives.” Theoretical Computer Science 22
(3): 317–30.
Bay, Herbert, Tinne Tuytelaars, and Luc Van Gool. 2006.
“SURF: Speeded up Robust
Features.” European Conference on
Computer Vision, 404–17. https://doi.org/10.1007/11744023_32.
Baydin, Atilim Gunes, Barak A Pearlmutter, Alexey Andreyevich Radul, and
Jeffrey Mark Siskind. 2018. “Automatic Differentiation in Machine
Learning: A Survey.” Journal of Machine
Learning Research 18 (153): 1–43. https://www.jmlr.org/papers/volume18/17-468/17-468.pdf.
Bayes, Thomas. 1763. “An Essay Towards Solving a Problem in the
Doctrine of Chances.” Philosophical Transactions of the Royal
Society of London 53: 370–418.
Beck, Amir, and Marc Teboulle. 2009. “A Fast Iterative
Shrinkage-Thresholding Algorithm for Linear Inverse Problems.”
SIAM Journal on Imaging Sciences 2 (1): 183–202.
Beck, Maximilian, Korbinian Pöppel, Markus Spanring, et al. 2024.
“xLSTM: Extended Long Short-Term
Memory.” Advances in Neural Information Processing
Systems 37. https://arxiv.org/abs/2405.04517.
Behrouz, Ali, Peilin Zhong, and Vahab Mirrokni. 2025. “Titans:
Learning to Memorize at Test Time.” arXiv Preprint
arXiv:2501.00663.
Belghazi, Mohamed Ishmael, Aristide Baratin, Sai Rajeshwar, et al. 2018.
“Mutual Information Neural Estimation.” Proceedings of
the 35th International Conference on Machine Learning, 531–40.
Belkin, Mikhail, Daniel Hsu, Siyuan Ma, and Soumik Mandal. 2019.
“Reconciling Modern Machine-Learning Practice and the Classical
Bias–Variance Trade-Off.” Proceedings of the National Academy
of Sciences 116 (32): 15849–54.
Bellman, R. 1966. “Dynamic Programming.” Science
153: 34–37. https://doi.org/10.1126/science.153.3731.34.
Bellman, Richard. 1952. “On the Theory of Dynamic
Programming.” Proceedings of the National
Academy of Sciences 38 (8): 716–19. https://doi.org/10.1073/pnas.38.8.716.
Bellman, Richard. 1957a. “A Markovian Decision
Process.” Journal of Mathematics and
Mechanics 6 (5): 679–84. http://www.jstor.org/stable/24900506.
Bellman, Richard. 1957b. Dynamic Programming.
Dover. Dover Publications. https://doi.org/10.1515/9781400835386.
Beltagy, Iz, Matthew E Peters, and Arman Cohan. 2020. “Longformer:
The Long-Document Transformer.”
ArXiv:2004.05150. https://arxiv.org/abs/2004.05150.
Benamou, Jean-David, and Yann Brenier. 2000. “A Computational
Fluid Mechanics Solution to the
Monge–Kantorovich Mass Transfer
Problem.” Numerische Mathematik 84 (3): 375–93.
Bengio, Yoshua, Réjean Ducharme, Pascal Vincent, and Christian Jauvin.
2003. “A Neural Probabilistic Language Model.”
Journal of Machine Learning
Research 3 (Feb): 1137–55. https://jmlr.org/papers/v3/bengio03a.html.
Bengio, Yoshua, and Yves Grandvalet. 2004. “No Unbiased Estimator
of the Variance of k-Fold
Cross-Validation.” Journal of Machine
Learning Research 5: 1089–105.
Bengio, Yoshua, Nicholas Léonard, and Aaron Courville. 2013.
“Estimating or Propagating Gradients Through Stochastic Neurons
for Conditional Computation.” arXiv Preprint
arXiv:1308.3432.
Bengio, Yoshua, Patrice Simard, and Paolo Frasconi. 1994.
“Learning Long-Term Dependencies with Gradient Descent Is
Difficult.” IEEE Transactions on
Neural Networks 5 (2): 157–66. https://doi.org/10.1109/72.279181.
Benjamini, Yoav, and Yosef Hochberg. 1995. “Controlling the False
Discovery Rate: A Practical and Powerful Approach to Multiple
Testing.” Journal of the Royal Statistical Society: Series B
(Methodological) 57 (1): 289–300.
Bergsma, Shane, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva,
and Joel Hestness. 2025a. “Power Lines: Scaling Laws for Weight
Decay and Batch Size in LLM Pre-Training.”
Advances in Neural Information Processing Systems. https://arxiv.org/abs/2505.13738.
Bergsma, Shane, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva,
and Joel Hestness. 2025b. “Straight to Zero: Why Linearly Decaying
the Learning Rate to Zero Works Best for LLM
Pretraining.” International Conference on Learning
Representations. https://arxiv.org/abs/2502.15938.
Bergstra, James, and Yoshua Bengio. 2012. “Random Search for
Hyper-Parameter Optimization.” Journal of Machine Learning
Research 13: 281–305.
Bergstra, James, Olivier Breuleux, Frédéric Bastien, et al. 2010.
“Theano: A CPU and GPU Math Compiler in
Python.” Proc. 9th Python in
Science Conference 1: 3–10. https://www.iro.umontreal.ca/~lisa/pointeurs/theano_scipy2010.pdf.
Bernstein, Jeremy. 2025. Deriving Muon. Https://jeremybernste.in/writing/deriving-muon.
Bernstein, Jeremy, and Laker Newhouse. 2024. “Old Optimizer, New
Norm: An Anatomy.” arXiv Preprint arXiv:2409.20325.
Beutel, Alex, Kenton Murray, Christos Faloutsos, and Alexander J Smola.
2014. “CoBaFi: Collaborative
Bayesian Filtering.” Proceedings of the 23rd
International Conference on World
Wide Web, 97–108. https://doi.org/10.1145/2566486.2567980.
Beyer, Lucas, Olivier J. Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and
Aäron van den Oord. 2020. “Are We Done with
ImageNet?” ArXiv:2006.07159.
https://arxiv.org/abs/2006.07159.
Bick, Aviv, Kevin Y. Li, Eric P. Xing, J. Zico Kolter, and Albert Gu.
2024. “Transformers to SSMs: Distilling Quadratic
Knowledge to Subquadratic Models.” Advances in Neural
Information Processing Systems.
Bishop, Chris M. 1995. “Training with Noise Is Equivalent to
Tikhonov Regularization.” Neural
Computation 7 (1): 108–16. https://doi.org/10.1162/neco.1995.7.1.108.
Bishop, Christopher M. 2006. Pattern Recognition and
Machine Learning. Springer. https://doi.org/10.1007/978-0-387-45528-0.
Black, Fischer, and Myron Scholes. 1973. “The Pricing of Options
and Corporate Liabilities.” Journal of Political
Economy 81: 637–54. https://doi.org/10.1086/260062.
Blelloch, Guy E. 1990. Prefix Sums and Their Applications.
CMU-CS-90-190. School of Computer Science, Carnegie Mellon University.
https://www.cs.cmu.edu/~scandal/papers/CMU-CS-90-190.html.
Bodla, Navaneeth, Bharat Singh, Rama Chellappa, and Larry S Davis. 2017.
“Soft-NMS-Improving Object Detection with One Line of
Code.” Proceedings of the IEEE
International Conference on
Computer Vision, 5561–69. https://doi.org/10.1109/iccv.2017.593.
Bojanowski, Piotr, Edouard Grave, Armand Joulin, and Tomas Mikolov.
2017. “Enriching Word Vectors with Subword Information.”
Transactions of the Association for
Computational Linguistics 5: 135–46. https://doi.org/10.1162/tacl_a_00051.
Bollobás, B. 1999. Linear Analysis. Cambridge
University Press. https://doi.org/10.1017/CBO9780511626296.
Bolte, Jérôme, and Edouard Pauwels. 2021. “Conservative Set Valued
Fields, Automatic Differentiation, Stochastic Gradient Methods and Deep
Learning.” Mathematical Programming 188 (1): 19–51.
Bommasani, Rishi, Drew A Hudson, Ehsan Adeli, et
al. 2021. “On the Opportunities and Risks of Foundation
Models.” ArXiv:2108.07258. https://arxiv.org/abs/2108.07258.
Bonferroni, Carlo E. 1936. “Teoria Statistica Delle Classi e
Calcolo Delle Probabilità.” Pubblicazioni Del R.
Istituto Superiore Di Scienze Economiche e Commerciali Di Firenze
8: 3–62.
Bottou, Léon. 2010. “Large-Scale Machine Learning with Stochastic
Gradient Descent.” In Proceedings of
COMPSTAT’2010. Springer. https://doi.org/10.1201/b11429-4.
Bottou, Léon, Frank E. Curtis, and Jorge Nocedal. 2018.
“Optimization Methods for Large-Scale Machine Learning.”
SIAM Review 60 (2): 223–311.
Bottou, Léon, and Yann Le Cun. 1988. “SN: A Simulator
for Connectionist Models.” Proceedings of
NeuroNimes 88 (Nimes, France), 371–82. http://leon.bottou.org/papers/bottou-lecun-88.
Boucheron, Stéphane, Olivier Bousquet, and Gábor Lugosi. 2005.
“Theory of Classification: A Survey of Some Recent
Advances.” ESAIM: Probability and
Statistics 9: 323–75. https://doi.org/10.1214/154957805100000046.
Boucheron, Stéphane, Gábor Lugosi, and Pascal Massart. 2013.
Concentration Inequalities: A Nonasymptotic Theory of
Independence. Oxford University Press.
Bowman, Samuel R, Gabor Angeli, Christopher Potts, and Christopher D
Manning. 2015. “A Large Annotated Corpus for Learning Natural
Language Inference.” Proceedings of the 2015 Conference on
Empirical Methods in Natural Language Processing, 632–42. https://arxiv.org/abs/1508.05326.
Boyd, Stephen, and Lieven Vandenberghe. 2004. Convex
Optimization. Cambridge University
Press. https://web.stanford.edu/~boyd/cvxbook/.
Bradley, Ralph Allan, and Milton E Terry. 1952. “Rank Analysis of
Incomplete Block Designs: I. The Method of
Paired Comparisons.” Biometrika 39 (3/4): 324–45. https://doi.org/10.1093/biomet/41.3-4.502.
Bray, Alan J., and David S. Dean. 2007. “Statistics of Critical
Points of Gaussian Fields on Large-Dimensional
Spaces.” Physical Review Letters 98 (15): 150201.
Brenier, Yann. 1991. “Polar Factorization and Monotone
Rearrangement of Vector-Valued Functions.” Communications on
Pure and Applied Mathematics 44 (4): 375–417.
Brin, Sergey, and Lawrence Page. 1998. “The Anatomy of a
Large-Scale Hypertextual Web Search Engine.” Computer
Networks and ISDN Systems 30 (1–7): 107–17.
Brock, Andrew, Soham De, Samuel L. Smith, and Karen Simonyan. 2021.
“High-Performance Large-Scale Image Recognition Without
Normalization.” International
Conference on Machine
Learning. https://arxiv.org/abs/2102.06171.
Brown, Noam, and Tuomas Sandholm. 2017. “Libratus: The Superhuman
AI for No-Limit Poker.” IJCAI,
5226–28. https://doi.org/10.24963/ijcai.2017/772.
Brown, Peter F, John Cocke, Stephen A Della Pietra, et al. 1990.
“A Statistical Approach to Machine Translation.”
Computational Linguistics 16 (2):
79–85. https://doi.org/10.1162/coli.1990.16.2.79.
Brown, Tom, Benjamin Mann, Nick Ryder, et
al. 2020. “Language Models Are Few-Shot Learners.”
Advances in Neural Information
Processing Systems 33: 1877–901. https://arxiv.org/abs/2005.14165.
Buslaev, Alexander, Vladimir I Iglovikov, Eugene Khvedchenya, Alex
Parinov, Mikhail Druzhinin, and Alexandr A Kalinin. 2020.
“Albumentations: Fast and Flexible Image
Augmentations.” Information 11 (2): 125. https://doi.org/10.3390/info11020125.
Campbell, Murray, A Joseph Hoane Jr, and Feng-hsiung Hsu. 2002.
“Deep Blue.” Artificial Intelligence
134 (1-2): 57–83. https://doi.org/10.1016/s0004-3702(01)00129-1.
Candès, Emmanuel J., and Benjamin Recht. 2009. “Exact Matrix
Completion via Convex Optimization.” Foundations of
Computational Mathematics 9 (6): 717–72.
Canny, John. 1987. “A Computational Approach to Edge
Detection.” In Readings in Computer
Vision. Elsevier. https://doi.org/10.1109/tpami.1986.4767851.
Carion, Nicolas, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier,
Alexander Kirillov, and Sergey Zagoruyko. 2020. “End-to-End Object
Detection with Transformers.” European Conference on Computer
Vision, 213–29.
Cer, Daniel, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia
Specia. 2017. “SemEval-2017 Task 1:
Semantic Textual Similarity Multilingual and Crosslingual Focused
Evaluation.” Proceedings of the 11th
International Workshop on
Semantic Evaluation
(SemEval-2017), 1–14. https://doi.org/10.18653/v1/s17-2001.
Chebyshev, Pafnuty L. 1867. “Des Valeurs Moyennes.”
Journal de Mathématiques Pures Et
Appliquées 12 (2): 177–84.
Chechik, Gal, Amir Globerson, Naftali Tishby, and Yair Weiss. 2005.
“Information Bottleneck for Gaussian
Variables.” Journal of Machine Learning Research 6:
165–88.
Chen, Lili, Kevin Lu, Aravind Rajeswaran, et al. 2021. “Decision
Transformer: Reinforcement Learning via Sequence Modeling.”
Advances in Neural Information
Processing Systems 34: 15084–97. https://arxiv.org/abs/2106.01345.
Chen, Ricky T. Q., Yulia Rubanova, Jesse Bettencourt, and David
Duvenaud. 2018. “Neural Ordinary Differential Equations.”
Advances in Neural Information Processing Systems 31.
Chen, Shouyuan, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023.
“Extending Context Window of Large Language Models via Positional
Interpolation.” arXiv Preprint arXiv:2306.15595.
Chen, Tianqi, Mu Li, Yutian Li, et al. 2015. “MXNET:
A Flexible and Efficient Machine Learning Library for Heterogeneous
Distributed Systems.” ArXiv:1512.01274. https://arxiv.org/abs/1512.01274.
Chen, Tianqi, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016.
“Training Deep Nets with Sublinear Memory Cost.” arXiv
Preprint arXiv:1604.06174.
Chen, Ting, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton.
2020. “A Simple Framework for Contrastive Learning of Visual
Representations.” Proceedings of the 37th International
Conference on Machine Learning, 1597–607.
Chen, Xiangning, Chen Liang, Da Huang, Esteban
Real, et al. 2023. “Symbolic Discovery of Optimization
Algorithms.” Advances in Neural Information Processing
Systems 36. https://arxiv.org/abs/2302.06675.
Chetlur, Sharan, Cliff Woolley, Philippe Vandermersch, et al. 2014.
“CuDNN: Efficient Primitives for Deep
Learning.” ArXiv:1410.0759. https://arxiv.org/abs/1410.0759.
Child, Rewon, Scott Gray, Alec Radford, and Ilya Sutskever. 2019.
“Generating Long Sequences with Sparse Transformers.”
ArXiv:1904.10509. https://arxiv.org/abs/1904.10509.
Cho, Kyunghyun, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua
Bengio. 2014. “On the Properties of Neural Machine Translation:
Encoder–Decoder Approaches.”
ArXiv:1409.1259. https://arxiv.org/abs/1409.1259.
Cho, Kyunghyun, Bart Van Merriënboer, Caglar Gulcehre, et al. 2014.
“Learning Phrase Representations Using RNN
Encoder-Decoder for Statistical Machine Translation.”
Proceedings of the 2014 Conference on Empirical Methods in Natural
Language Processing, 1724–34. https://arxiv.org/abs/1406.1078.
Chollet, François. 2017. “Xception: Deep Learning with Depthwise
Separable Convolutions.” Proceedings of the IEEE
Conference on Computer Vision and
Pattern Recognition. https://arxiv.org/abs/1610.02357.
Choromanski, Krzysztof, Valerii Likhosherstov,
David Dohan, et al. 2021. “Rethinking Attention with
Performers.” International Conference on Learning
Representations.
Chowdhery, Aakanksha, Sharan Narang, Jacob Devlin,
et al. 2022. “PaLM: Scaling Language Modeling
with Pathways.” ArXiv:2204.02311. https://arxiv.org/abs/2204.02311.
Chu, Xiangxiang, Liang Li, and Bo Zhang. 2024. “Make
RepVGG Greater Again: A Quantization-Aware
Approach.” AAAI Conference on
Artificial Intelligence. https://arxiv.org/abs/2212.01593.
Chung, Junyoung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio.
2014. “Empirical Evaluation of Gated Recurrent Neural Networks on
Sequence Modeling.” ArXiv:1412.3555. https://arxiv.org/abs/1412.3555.
Church, Kenneth Ward, and Patrick Hanks. 1990. “Word Association
Norms, Mutual Information, and Lexicography.” Computational
Linguistics 16 (1): 22–29.
Chwialkowski, Kacper, Heiko Strathmann, and Arthur Gretton. 2016.
“A Kernel Test of Goodness of Fit.” Proceedings of the
33rd International Conference on Machine Learning, 2606–15.
Coddington, Earl A., and Norman Levinson. 1955. Theory of Ordinary
Differential Equations. McGraw-Hill.
Cohen, Jeremy M., Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet
Talwalkar. 2021. “Gradient Descent on Neural Networks Typically
Occurs at the Edge of Stability.” International Conference on
Learning Representations.
Collobert, Ronan, Jason Weston, Léon Bottou, Michael Karlen, Koray
Kavukcuoglu, and Pavel Kuksa. 2011. “Natural Language Processing
(Almost) from Scratch.” Journal of
Machine Learning Research
12: 2493–537. https://jmlr.org/papers/v12/collobert11a.html.
Cordonnier, Jean-Baptiste, Andreas Loukas, and Martin Jaggi. 2020.
“On the Relationship Between Self-Attention and Convolutional
Layers.” International Conference
on Learning Representations. https://openreview.net/forum?id=HJlnC1rKPB.
Cortes, Corinna, and Vladimir Vapnik. 1995. “Support-Vector
Networks.” Machine Learning 20 (3): 273–97.
Cover, Thomas M., and Peter E. Hart. 1967. “Nearest Neighbor
Pattern Classification.” IEEE Transactions on Information
Theory 13 (1): 21–27.
Cover, T, and JM Thomas. 1999. Elements of Information
Theory. John Wiley &
Sons. https://doi.org/10.1002/047174882X.
Cramér, H. 1946. Mathematical Methods of
Statistics. Princeton University
Press.
Csiszár, Imre. 1967. “Information-Type Measures of Difference of
Probability Distributions and Indirect Observations.” Studia
Scientiarum Mathematicarum Hungarica 2: 299–318.
Csiszár, Imre. 2008. “Axiomatic Characterizations of Information
Measures.” Entropy 10 (3): 261–73. https://doi.org/10.3390/e10030261.
Cubuk, Ekin D., Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2020.
“RandAugment: Practical Automated Data Augmentation
with a Reduced Search Space.” Proceedings of the
IEEE/CVF Conference on
Computer Vision and Pattern
Recognition Workshops. https://arxiv.org/abs/1909.13719.
Cuturi, Marco. 2013. “Sinkhorn Distances: Lightspeed Computation
of Optimal Transport.” Advances in Neural Information
Processing Systems 26.
Cybenko, George. 1989. “Approximation by Superpositions of a
Sigmoidal Function.” Mathematics of Control,
Signals and Systems 2 (4): 303–14. https://doi.org/10.1007/bf02134016.
D’Angelo, Francesco, Maksym Andriushchenko, Aditya Varre, and Nicolas
Flammarion. 2024. “Why Do We Need Weight Decay in Modern Deep
Learning?” Advances in Neural Information Processing
Systems 37. https://arxiv.org/abs/2310.04415.
Dahl, George E., Frank Schneider, Zachary Nado, et
al. 2023. “Benchmarking Neural Network Training
Algorithms.” arXiv Preprint arXiv:2306.07179.
Dahlquist, Germund G. 1963. “A Special Stability Problem for
Linear Multistep Methods.” BIT Numerical Mathematics 3:
27–43.
Dai, Damai, Chengqi Deng, Chenggang Zhao, et
al. 2024. “DeepSeekMoE: Towards Ultimate
Expert Specialization in Mixture-of-Experts Language Models.”
arXiv Preprint arXiv:2401.06066.
Dalal, Navneet, and Bill Triggs. 2005. “Histograms of Oriented
Gradients for Human Detection.” 2005 IEEE
Computer Society Conference on
Computer Vision and Pattern
Recognition (CVPR’05) 1: 886–93. https://doi.org/10.1109/cvpr.2005.177.
Dao, Tri, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.
2022. “FlashAttention: Fast and Memory-Efficient
Exact Attention with IO-Awareness.” Advances in
Neural Information Processing Systems.
Dao, Tri, and Albert Gu. 2024. “Transformers Are
SSMs: Generalized Models and Efficient Algorithms Through
Structured State Space Duality.” International Conference on
Machine Learning, 10041–71. https://arxiv.org/abs/2405.21060.
Dauphin, Yann N., Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya
Ganguli, and Yoshua Bengio. 2014. “Identifying and Attacking the
Saddle Point Problem in High-Dimensional Non-Convex
Optimization.” Advances in Neural Information Processing
Systems 27: 2933–41.
De Cock, Dean. 2011. “Ames, Iowa: Alternative to the
Boston Housing Data as an End of Semester Regression
Project.” Journal of Statistics
Education 19 (3). https://doi.org/10.1080/10691898.2011.11889627.
De, Soham, Samuel L. Smith, Anushan Fernando, et al. 2024.
“Griffin: Mixing Gated Linear Recurrences with Local
Attention for Efficient Language Models.” arXiv Preprint
arXiv:2402.19427.
Dean, Jeffrey, Greg S Corrado, Rajat Monga, et
al. 2012. “Large Scale Distributed Deep Networks.”
Proceedings of the 25th International
Conference on Neural Information
Processing Systems, Volume
1, 1223–31. https://doi.org/10.5555/2999134.2999271.
DeepSeek-AI. 2024. “DeepSeek-V2: A Strong,
Economical, and Efficient Mixture-of-Experts Language Model.”
arXiv Preprint arXiv:2405.04434.
DeepSeek-AI, Xiao Bi, Deli Chen, Guanting Chen, et
al. 2024. “DeepSeek LLM: Scaling Open-Source
Language Models with Longtermism.” arXiv Preprint
arXiv:2401.02954.
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, et
al. 2024. “DeepSeek-V3 Technical
Report.” arXiv Preprint arXiv:2412.19437.
Defazio, Aaron, Ashok Cutkosky, Harsh Mehta, and Konstantin Mishchenko.
2023. “Optimal Linear Decay Learning Rate Schedules and Further
Refinements.” arXiv Preprint arXiv:2310.07831.
Defazio, Aaron, Xingyu Alice Yang, Harsh Mehta, Konstantin Mishchenko,
Ahmed Khaled, and Ashok Cutkosky. 2024. “The Road Less
Scheduled.” Advances in Neural Information Processing
Systems 37. https://arxiv.org/abs/2405.15682.
Dehghani, Mostafa, Josip Djolonga, Basil Mustafa,
et al. 2023. “Scaling Vision Transformers to 22 Billion
Parameters.” International Conference on Machine
Learning, 7480–512.
Deisenroth, Marc Peter, A. Aldo Faisal, and Cheng Soon Ong. 2020.
Mathematics for Machine Learning. Cambridge University Press.
https://mml-book.github.io/.
Delétang, Grégoire, Anian Ruoss, Paul-Ambroise Duquenne, et al. 2023.
“Language Modeling Is Compression.”
ArXiv:2309.10668. https://arxiv.org/abs/2309.10668.
Dempster, Arthur P, Nan M Laird, and Donald B Rubin. 1977.
“Maximum Likelihood from Incomplete Data via the EM
Algorithm.” Journal of the Royal
Statistical Society: Series
B 39 (1): 1–38.
Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei.
2009. “Imagenet: A Large-Scale Hierarchical Image
Database.” 2009 IEEE Conference on
Computer Vision and Pattern
Recognition, 248–55. https://doi.org/10.1109/cvpr.2009.5206848.
Der Kiureghian, Armen, and Ove Ditlevsen. 2009. “Aleatory or
Epistemic? Does It Matter?” Structural
Safety 31 (2): 105–12. https://doi.org/10.1016/j.strusafe.2008.06.020.
Dettmers, Tim, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2022.
“8-Bit Optimizers via Block-Wise Quantization.”
International Conference on Learning Representations. https://arxiv.org/abs/2110.02861.
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018.
“BERT: Pre-Training of Deep
Bidirectional Transformers for Language Understanding.”
ArXiv:1810.04805. https://arxiv.org/abs/1810.04805.
Dey, Nolan, Quentin Anthony, and Joel Hestness. 2024. The
Practitioner’s Guide to the Maximal Update Parameterization. Https://blog.eleuther.ai/mutransfer/.
Dey, Nolan, Gurpreet Gosal, Zhiming Chen, et al. 2023.
“Cerebras-GPT: Open Compute-Optimal Language Models
Trained on the Cerebras Wafer-Scale Cluster.”
arXiv Preprint arXiv:2304.03208.
Dhariwal, Prafulla, and Alex Nichol. 2021. “Diffusion Models Beat
GANs on Image Synthesis.” Advances in Neural
Information Processing Systems 34.
Ding, Xiaohan, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding,
and Jian Sun. 2021. “RepVGG: Making
VGG-Style ConvNets Great
Again.” Proceedings of the IEEE/CVF
Conference on Computer Vision and
Pattern Recognition. https://arxiv.org/abs/2101.03697.
Ding, Xiaohan, Xiangyu Zhang, Yizhuang Zhou, Jungong Han, Guiguang Ding,
and Jian Sun. 2022. “Scaling up Your Kernels to 31x31: Revisiting
Large Kernel Design in CNNs.” Proceedings of the
IEEE/CVF Conference on
Computer Vision and Pattern
Recognition. https://arxiv.org/abs/2203.06717.
Ding, Xiaohan, Yiyuan Zhang, Yixiao Ge, et al. 2024.
“UniRepLKNet: A Universal Perception Large-Kernel
ConvNet for Audio, Video, Point Cloud,
Time-Series and Image Recognition.” Proceedings of the
IEEE/CVF Conference on
Computer Vision and Pattern
Recognition. https://arxiv.org/abs/2311.15599.
Dinh, Laurent, David Krueger, and Yoshua Bengio. 2014.
“NICE: Non-Linear Independent Components
Estimation.” ArXiv:1410.8516. https://arxiv.org/abs/1410.8516.
Dinh, Laurent, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017.
“Sharp Minima Can Generalize for Deep Nets.”
Proceedings of the 34th International Conference on Machine
Learning, Proceedings of machine learning research, vol. 70:
1019–28. https://proceedings.mlr.press/v70/dinh17b.html.
Dinh, Laurent, Jascha Sohl-Dickstein, and Samy Bengio. 2017.
“Density Estimation Using Real NVP.”
International Conference on
Learning Representations. https://openreview.net/forum?id=HkpbnH9lx.
Doersch, Carl, Abhinav Gupta, and Alexei A Efros. 2015.
“Unsupervised Visual Representation Learning by Context
Prediction.” Proceedings of the IEEE
International Conference on
Computer Vision, 1422–30. https://doi.org/10.1109/iccv.2015.167.
Domingos, Pedro, and Michael Pazzani. 1997. “On the Optimality of
the Simple Bayesian Classifier Under Zero-One Loss.”
Machine Learning 29: 103–30.
Dong, Xin, Yonggan Fu, Shizhe Diao, et al. 2024. “Hymba: A
Hybrid-Head Architecture for Small Language Models.” arXiv
Preprint arXiv:2411.13676.
Dong, Yihe, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021.
“Attention Is Not All You Need: Pure Attention Loses Rank Doubly
Exponentially with Depth.” International Conference on
Machine Learning, 2793–803.
Donsker, Monroe D., and S. R. Srinivasa Varadhan. 1983.
“Asymptotic Evaluation of Certain Markov Process
Expectations for Large Time. IV.” Communications
on Pure and Applied Mathematics 36 (2): 183–212.
Dormand, John R., and Peter J. Prince. 1980. “A Family of Embedded
Runge–Kutta Formulae.” Journal of
Computational and Applied Mathematics 6 (1): 19–26.
Dosovitskiy, Alexey, Lucas Beyer, Alexander
Kolesnikov, et al. 2021. “An Image Is Worth 16 x 16 Words:
Transformers for Image Recognition at Scale.”
International Conference on
Learning Representations. https://openreview.net/forum?id=YicbFdNTTy.
Duchi, John, Elad Hazan, and Yoram Singer. 2011. “Adaptive
Subgradient Methods for Online Learning and Stochastic
Optimization.” Journal of Machine
Learning Research 12: 2121–59. https://jmlr.org/papers/v12/duchi11a.html.
Duchi, John, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra.
2008. “Efficient Projections onto the ℓ1-Ball for Learning in
High Dimensions.” Proceedings of the 25th International
Conference on Machine Learning (ICML), 272–79.
Dumoulin, Vincent, and Francesco Visin. 2016. “A Guide to
Convolution Arithmetic for Deep Learning.”
ArXiv:1603.07285. https://arxiv.org/abs/1603.07285.
Dwivedi, Vijay Prakash, and Xavier Bresson. 2020. “A
Generalization of Transformer Networks to Graphs.”
ArXiv:2012.09699. https://arxiv.org/abs/2012.09699.
Dwork, Cynthia, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer
Reingold, and Aaron Roth. 2015. “Preserving Statistical Validity
in Adaptive Data Analysis.” ACM Symposium on Theory of
Computing, 117–26.
Dwork, Cynthia, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer
Reingold, and Aaron Leon Roth. 2015. “Preserving Statistical
Validity in Adaptive Data Analysis.” Proceedings of the 47th
Annual ACM Symposium on
Theory of Computing, 117–26. https://doi.org/10.1145/2746539.2746580.
E, Weinan. 2017. “A Proposal on Machine Learning via Dynamical
Systems.” Communications in Mathematics and Statistics 5
(1): 1–11.
Eckart, Carl, and Gale Young. 1936. “The Approximation of One
Matrix by Another of Lower Rank.” Psychometrika 1 (3):
211–18.
Efron, Bradley. 1979. “Bootstrap Methods: Another Look at the
Jackknife.” Annals of Statistics 7 (1): 1–26.
Efron, Bradley. 2011. “Tweedie’s Formula and Selection
Bias.” Journal of the American Statistical Association
106 (496): 1602–14.
Efron, Bradley, and Trevor Hastie. 2016. Computer Age Statistical
Inference: Algorithms, Evidence, and Data Science. Cambridge
University Press. https://hastie.su.domains/CASI/.
Elhage, Nelson, Neel Nanda, Catherine Olsson, et
al. 2021. “A Mathematical Framework for Transformer
Circuits.” Transformer Circuits Thread.
Elsken, T., J. H. Metzen, and F. Hutter. 2018. “Neural
Architecture Search: A Survey.” ArXiv:1808.05377
[Stat.ML]. https://arxiv.org/abs/1808.05377.
Endres, Dominik M., and Johannes E. Schindelin. 2003. “A New
Metric for Probability Distributions.” IEEE Transactions on
Information Theory 49 (7): 1858–60.
Erven, Tim van, and Peter Harremoës. 2014. “Rényi
Divergence and Kullback–Leibler
Divergence.” IEEE Transactions on Information Theory 60
(7): 3797–820.
Fan, Angela, Mike Lewis, and Yann Dauphin. 2018. “Hierarchical
Neural Story Generation.” Proceedings of the 56th Annual
Meeting of the Association for Computational Linguistics, 889–98.
Fechner, Gustav Theodor. 1860. Elemente Der
Psychophysik. Vol. 2. Breitkopf u.
Härtel.
Fedus, William, Barret Zoph, and Noam Shazeer. 2022. “Switch
Transformers: Scaling to Trillion Parameter Models with Simple and
Efficient Sparsity.” Journal of
Machine Learning Research 23
(120): 1–39. https://arxiv.org/abs/2101.03961.
Feng, Leo, Frederick Tung, Mohamed Osama Ahmed, Yoshua Bengio, and
Hossein Hajimirsadeghi. 2024. “Were RNNs All We
Needed?” arXiv Preprint arXiv:2410.01201.
Fernando, Randima. 2004. GPU Gems:
Programming Techniques, Tips, and
Tricks for Real-Time
Graphics. Addison-Wesley.
Feurer, M., and F. Hutter. 2019. “Hyperparameter
Optimization.” In Automated Machine
Learning: Methods, Systems,
Challenges. Springer. https://doi.org/10.1007/978-3-030-05318-5_1.
Feurer, M., B. Letham, F. Hutter, and E. Bakshy. 2022. “Practical
Transfer Learning for Bayesian Optimization.”
ArXiv:1802.02219 [Stat.ML]. https://arxiv.org/abs/1802.02219.
Field, David J. 1987. “Relations Between the Statistics of Natural
Images and the Response Properties of Cortical Cells.”
JOSA A 4 (12): 2379–94. https://doi.org/10.1364/josaa.4.002379.
Fisher, R A. 1925. Statistical Methods for
Research Workers. Oliver &
Boyd.
Fisher, Ronald A. 1935. The Design of Experiments. Oliver;
Boyd.
Fokker, Adriaan D. 1914. “Die Mittlere Energie
Rotierender Elektrischer Dipole Im
Strahlungsfeld.” Annalen Der Physik 348
(5): 810–20.
Folland, Gerald B. 1999. Real Analysis: Modern Techniques and Their
Applications. 2nd ed. Wiley.
Foret, Pierre, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur.
2021. “Sharpness-Aware Minimization for Efficiently Improving
Generalization.” International Conference on Learning
Representations. https://arxiv.org/abs/2010.01412.
Forrester, Alexander IJ, András Sóbester, and Andy J Keane. 2007.
“Multi-Fidelity Optimization via Surrogate Modelling.”
Proceedings of the Royal Society
A: Mathematical, Physical and
Engineering Sciences 463 (2088): 3251–69.
https://doi.org/10.1098/rspa.2007.1900.
Franceschi, L., M. Donini, P. Frasconi, and M. Pontil. 2017.
“Forward and Reverse Gradient-Based Hyperparameter
Optimization.” Proceedings of the 34th
International Conference on
Machine Learning (ICML’17).
https://proceedings.mlr.press/v70/franceschi17a.html.
Frankle, Jonathan, and Michael Carbin. 2019. “The Lottery Ticket
Hypothesis: Finding Sparse, Trainable Neural Networks.”
International Conference on Learning Representations. https://arxiv.org/abs/1803.03635.
Frazier, Peter I. 2018. “A Tutorial on Bayesian
Optimization.” ArXiv:1807.02811. https://arxiv.org/abs/1807.02811.
Freund, Yoav, and Robert E Schapire. 1996. “Experiments with a New
Boosting Algorithm.” Proceedings of the
International Conference on
Machine Learning 96: 148–56. https://dl.acm.org/doi/10.5555/3091696.3091715.
Friedman, Jerome H. 1987. “Exploratory Projection Pursuit.”
Journal of the American Statistical
Association 82 (397): 249–66. https://doi.org/10.1080/01621459.1987.10478427.
Frostig, Roy, Matthew James Johnson, and Chris Leary. 2018.
“Compiling Machine Learning Programs via High-Level
Tracing.” In Proceedings of Systems for Machine
Learning. https://mlsys.org/Conferences/2018/doc/19.pdf.
Fukushima, Kunihiko. 1982. “Neocognitron: A Self-Organizing Neural
Network Model for a Mechanism of Visual Pattern Recognition.” In
Competition and Cooperation in Neural
Nets. Springer. https://doi.org/10.1007/978-3-642-46466-9_18.
Gal, Yarin, and Zoubin Ghahramani. 2016. “Dropout as a
Bayesian Approximation: Representing Model Uncertainty in
Deep Learning.” International Conference on Machine
Learning, 1050–59. https://arxiv.org/abs/1506.02142.
Gardner, Jacob, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and
Andrew G Wilson. 2018. “GPyTorch: Blackbox
Matrix–Matrix Gaussian Process Inference with
GPU Acceleration.” Advances in
Neural Information Processing
Systems 31. https://doi.org/10.5555/3327345.3327419.
Garg, Saurabh, Sivaraman Balakrishnan, Zico Kolter, and Zachary Lipton.
2021. “RATT: Leveraging Unlabeled Data to Guarantee
Generalization.” International
Conference on Machine
Learning, 3598–609. https://proceedings.mlr.press/v139/garg21a.html.
Gatys, Leon A, Alexander S Ecker, and Matthias Bethge. 2016.
“Image Style Transfer Using Convolutional Neural Networks.”
Proceedings of the IEEE Conference on
Computer Vision and Pattern
Recognition, 2414–23. https://doi.org/10.1109/cvpr.2016.265.
Gauss, Carl Friedrich. 1809. “Theoria Motus Corporum
Coelestium.” In Werke. Königlich
Preussische Akademie der
Wissenschaften. https://doi.org/10.1007/978-3-642-92478-1.
Gavish, Matan, and David L. Donoho. 2014. “The Optimal Hard
Threshold for Singular Values Is 4/√3.” IEEE Transactions on
Information Theory 60 (8): 5040–53.
Geman, Stuart, Elie Bienenstock, and René Doursat. 1992. “Neural
Networks and the Bias/Variance Dilemma.” Neural
Computation 4 (1): 1–58.
Gemma Team. 2025. “Gemma 3 Technical Report.” arXiv
Preprint arXiv:2503.19786.
Gershgorin, Semyon A. 1931. “Über Die
Abgrenzung Der Eigenwerte Einer
Matrix.” Izvestiya Akademii Nauk SSSR, Seriya
Matematicheskaya, no. 6: 749–54.
Geshkovski, Borjan, Cyril Letrouit, Yury Polyanskiy, and Philippe
Rigollet. 2023. “A Mathematical Perspective on
Transformers.” arXiv Preprint arXiv:2312.10794.
Ghadimi, Saeed, and Guanghui Lan. 2013. “Stochastic First- and
Zeroth-Order Methods for Nonconvex Stochastic Programming.”
SIAM Journal on Optimization 23 (4): 2341–68. https://arxiv.org/abs/1309.5549.
Gibbs, Josiah Willard. 1902. Elementary Principles of
Statistical Mechanics. Scribner’s.
Ginibre, Jean. 1965. “Statistical Ensembles of Complex,
Quaternion, and Real Matrices.” Journal of
Mathematical Physics 6 (3): 440–49. https://doi.org/10.1063/1.1704292.
Girshick, Ross. 2015. “Fast
R-CNN.” Proceedings of the
IEEE International Conference on
Computer Vision, 1440–48. https://doi.org/10.1109/iccv.2015.169.
Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014.
“Rich Feature Hierarchies for Accurate Object Detection and
Semantic Segmentation.” Proceedings of the IEEE
Conference on Computer Vision and
Pattern Recognition, 580–87. https://doi.org/10.18127/j00338486-202109-11.
Glorioso, Paolo, Quentin Anthony, Yury Tokpanov, Anna Golubeva, et al.
2024. “The Zamba2 Suite: Technical Report.”
arXiv Preprint arXiv:2411.15242.
Glorioso, Paolo, Quentin Anthony, Yury Tokpanov, James Whittington, et
al. 2024. “Zamba: A Compact 7B SSM
Hybrid Model.” arXiv Preprint arXiv:2405.16712.
Glorot, Xavier, and Yoshua Bengio. 2010. “Understanding the
Difficulty of Training Deep Feedforward Neural Networks.”
Proceedings of the 13th International
Conference on Artificial
Intelligence and Statistics, 249–56. https://proceedings.mlr.press/v9/glorot10a.html.
Gneiting, Tilmann, and Adrian E. Raftery. 2007. “Strictly Proper
Scoring Rules, Prediction, and Estimation.” Journal of the
American Statistical
Association 102 (477): 359–78. https://doi.org/10.1198/016214506000001437.
Godbole, Varun, George E. Dahl, Justin Gilmer, Christopher J. Shallue,
and Zachary Nado. 2023. Deep Learning Tuning Playbook. Https://github.com/google-research/tuning_playbook.
Goh, Gabriel. 2017. “Why Momentum Really Works.”
Distill. http://distill.pub/2017/momentum.
Goldberg, David. 1991. “What Every Computer Scientist Should Know
about Floating-Point Arithmetic.” ACM Computing Surveys
23 (1): 5–48.
Goldberg, David, David Nichols, Brian M Oki, and Douglas Terry. 1992.
“Using Collaborative Filtering to Weave an Information
Tapestry.” Communications of the ACM 35
(12): 61–71. https://doi.org/10.1145/138859.138867.
Golub, Gene H, and Charles F Van Loan. 1996. Matrix
Computations. Johns Hopkins
University Press. https://jhupbooks.press.jhu.edu/title/matrix-computations.
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep
Learning. MIT Press.
Goodfellow, Ian, Jean Pouget-Abadie, Mehdi Mirza, et al. 2014.
“Generative Adversarial Nets.” Advances in
Neural Information Processing
Systems, 2672–80. https://doi.org/10.5555/2969033.2969125.
Gotmare, Akhilesh, Nitish Shirish Keskar, Caiming Xiong, and Richard
Socher. 2018. “A Closer Look at Deep Learning Heuristics: Learning
Rate Restarts, Warmup and Distillation.”
ArXiv:1810.13243. https://arxiv.org/abs/1810.13243.
Goyal, Priya, Piotr Dollár, Ross Girshick, et al. 2017. “Accurate,
Large Minibatch SGD: Training
ImageNet in 1 Hour.”
ArXiv:1706.02677. https://arxiv.org/abs/1706.02677.
Grathwohl, Will, Ricky T. Q. Chen, Jesse Bettencourt, Ilya Sutskever,
and David Duvenaud. 2019. “FFJORD: Free-Form
Continuous Dynamics for Scalable Reversible Generative Models.”
International Conference on Learning Representations.
Grattafiori, Aaron, Abhimanyu Dubey, Abhinav
Jauhri, et al. 2024. “The Llama 3 Herd of
Models.” arXiv Preprint arXiv:2407.21783.
Graves, Alex. 2013. “Generating Sequences with Recurrent Neural
Networks.” ArXiv:1308.0850. https://arxiv.org/abs/1308.0850.
Graves, Alex, and Jürgen Schmidhuber. 2005. “Framewise Phoneme
Classification with Bidirectional LSTM and Other Neural
Network Architectures.” Neural Networks 18 (5-6):
602–10. https://doi.org/10.1109/ijcnn.2005.1556215.
Grazzi, Riccardo, Julien Siems, Jörg K. H. Franke, Arber Zela, Frank
Hutter, and Massimiliano Pontil. 2025. “Unlocking State-Tracking
in Linear RNNs Through Negative Eigenvalues.”
International Conference on Learning Representations. https://arxiv.org/abs/2411.12537.
Gretton, Arthur, Karsten M. Borgwardt, Malte J. Rasch, Bernhard
Schölkopf, and Alexander Smola. 2012. “A Kernel Two-Sample
Test.” Journal of Machine Learning Research 13: 723–73.
Griewank, Andreas. 1992. “Achieving Logarithmic Growth of Temporal
and Spatial Complexity in Reverse Automatic Differentiation.”
Optimization Methods and Software 1 (1): 35–54.
Griewank, Andreas, and Andrea Walther. 2008. Evaluating Derivatives:
Principles and Techniques of Algorithmic Differentiation. 2nd ed.
Society for Industrial; Applied Mathematics (SIAM). https://doi.org/10.1137/1.9780898717761.
Grinsztajn, Léo, Edouard Oyallon, and Gaël Varoquaux. 2022. “Why
Do Tree-Based Models Still Outperform Deep Learning on Tabular
Data?” Advances in Neural Information Processing Systems
35: 507–20. https://arxiv.org/abs/2207.08815.
Grobman, David M. 1959. “Homeomorphism of Systems of Differential
Equations.” Doklady Akademii Nauk SSSR 128: 880–81.
Gu, Albert. 2023. “Modeling Sequences with Structured State
Spaces.” PhD thesis, Stanford University.
Gu, Albert. 2025. On the Tradeoffs of SSMs and
Transformers. Https://goombalab.github.io/blog/2025/tradeoffs/.
Gu, Albert, and Tri Dao. 2023. “Mamba: Linear-Time
Sequence Modeling with Selective State Spaces.” arXiv
Preprint arXiv:2312.00752.
Gu, Albert, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré.
2020. “HiPPO: Recurrent Memory with Optimal
Polynomial Projections.” Advances in Neural Information
Processing Systems 33: 1474–87.
Gu, Albert, Karan Goel, and Christopher Ré. 2022. “Efficiently
Modeling Long Sequences with Structured State Spaces.”
International Conference on Learning Representations. https://arxiv.org/abs/2111.00396.
Gu, Albert, Ankit Gupta, Karan Goel, and Christopher Ré. 2022. “On
the Parameterization and Initialization of Diagonal State Space
Models.” Advances in Neural Information Processing
Systems 35: 35971–83.
Gulati, Anmol, James Qin, Chung-Cheng Chiu, et
al. 2020. “Conformer: Convolution-Augmented Transformer for
Speech Recognition.” Proc. Interspeech
2020, 5036–40. https://doi.org/10.21437/interspeech.2020-3015.
Gulrajani, Ishaan, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and
Aaron Courville. 2017. “Improved Training of
Wasserstein GANs.” Advances in
Neural Information Processing Systems 30.
Gumbel, Emil J. 1954. Statistical Theory of Extreme Values and Some
Practical Applications. U.S. National Bureau of Standards.
Gunawardana, Asela, and Guy Shani. 2015. “Evaluating Recommender
Systems.” In Recommender Systems
Handbook. Springer. https://doi.org/10.1007/978-1-4899-7637-6_8.
Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017.
“On Calibration of Modern Neural Networks.”
International Conference on Machine Learning, 1321–30. https://arxiv.org/abs/1706.04599.
Guo, Huifeng, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He.
2017. “DeepFM: A Factorization-Machine Based Neural Network for
CTR Prediction.” Proceedings of the 26th
International Joint Conference on
Artificial Intelligence, 1725–31. https://doi.org/10.24963/ijcai.2017/239.
Gupta, Suyog, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish
Narayanan. 2015. “Deep Learning with Limited Numerical
Precision.” International Conference on Machine
Learning, 1737–46.
Gupta, Vineet, Tomer Koren, and Yoram Singer. 2018. “Shampoo:
Preconditioned Stochastic Tensor Optimization.” Proceedings
of the 35th International Conference on Machine Learning (ICML),
1842–50.
Guyon, Isabelle, Steve Gunn, Masoud Nikravesh, and Lotfi A Zadeh. 2008.
Feature Extraction: Foundations and
Applications. Springer. https://doi.org/10.1007/978-3-540-35488-8.
Hägele, Alexander, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro
von Werra, and Martin Jaggi. 2024. “Scaling Laws and
Compute-Optimal Training Beyond Fixed Training Durations.”
Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2405.18392.
Hairer, Ernst, Syvert P. Nørsett, and Gerhard Wanner. 1993. Solving
Ordinary Differential Equations I: Nonstiff Problems.
2nd ed. Springer.
Halko, Nathan, Per-Gunnar Martinsson, and Joel A. Tropp. 2011.
“Finding Structure with Randomness: Probabilistic Algorithms for
Constructing Approximate Matrix Decompositions.” SIAM
Review 53 (2): 217–88.
Hanson, Stephen José, and Lorien Y Pratt. 1988. “Comparing Biases
for Minimal Network Construction with Back-Propagation.”
Advances in Neural Information
Processing Systems 1.
Harris, Charles R., K. Jarrod Millman, Stéfan J.
van der Walt, et al. 2020. “Array Programming with
NumPy.” Nature 585 (7825): 357–62. https://doi.org/10.1038/s41586-020-2649-2.
Hartley, Richard, and Andrew Zisserman. 2000. Multiple
View Geometry in Computer
Vision. Cambridge University
Press. https://doi.org/10.1017/cbo9780511811685.
Hartman, Philip. 1960. “A Lemma in the Theory of Structural
Stability of Differential Equations.” Proceedings of the
American Mathematical Society 11 (4): 610–20.
Haveliwala, Taher H., and Sepandar D. Kamvar. 2003. The Second
Eigenvalue of the Google Matrix. Stanford University.
http://ilpubs.stanford.edu:8090/582/.
He, Horace. 2022. Making Deep Learning Go Brrrr from First
Principles. Https://horace.io/brrr_intro.html.
He, Kaiming, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and
Ross Girshick. 2022. “Masked Autoencoders Are Scalable Vision
Learners.” Proceedings of the
IEEE/CVF Conference on
Computer Vision and Pattern
Recognition, 16000–16009. https://doi.org/10.1109/cvpr52688.2022.01553.
He, Kaiming, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017.
“Mask R-CNN.” Proceedings of
the IEEE International Conference
on Computer Vision, 2961–69. https://doi.org/10.1109/iccv.2017.322.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015.
“Delving Deep into Rectifiers: Surpassing Human-Level Performance
on ImageNet Classification.”
Proceedings of the IEEE International
Conference on Computer
Vision, 1026–34. https://doi.org/10.1109/iccv.2015.123.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016.
“Deep Residual Learning for Image Recognition.”
Proceedings of the IEEE Conference on
Computer Vision and Pattern
Recognition, 770–78. https://doi.org/10.1109/cvpr.2016.90.
He, Tong, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li.
2019. “Bag of Tricks for Image Classification with Convolutional
Neural Networks.” IEEE/CVF
Conference on Computer Vision and
Pattern Recognition, 558–67. https://arxiv.org/abs/1812.01187.
He, Xiangnan, and Tat-Seng Chua. 2017. “Neural Factorization
Machines for Sparse Predictive Analytics.” Proceedings of the
40th International ACM SIGIR
Conference on Research and
Development in Information
Retrieval, 355–64. https://doi.org/10.1145/3077136.3080777.
He, Xiangnan, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and
Tat-Seng Chua. 2017. “Neural Collaborative Filtering.”
Proceedings of the 26th International
Conference on World Wide
Web, 173–82. https://doi.org/10.1145/3038912.3052569.
Hebb, Donald Olding. 1949. The Organization of
Behavior. Wiley.
Held, Michael, Philip Wolfe, and Harlan P. Crowder. 1974.
“Validation of Subgradient Optimization.” Mathematical
Programming 6: 62–88.
Hendrycks, Dan, and Kevin Gimpel. 2016. “Gaussian Error Linear
Units (GELUs).”
ArXiv:1606.08415. https://arxiv.org/abs/1606.08415.
Hennessy, John L, and David A Patterson. 2011. Computer
Architecture: A Quantitative
Approach. Elsevier. https://www.elsevier.com/books/computer-architecture/hennessy/978-0-12-383872-8.
Henry, Alex, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan
Chen. 2020. “Query-Key Normalization for Transformers.”
Findings of the Association for Computational Linguistics: EMNLP
2020.
Herlocker, Jonathan L, Joseph A Konstan, Al Borchers, and John Riedl.
1999. “An Algorithmic Framework for Performing Collaborative
Filtering.” 22nd Annual
International ACM Conference on
Research and Development in
Information Retrieval, SIGIR
1999, 230–37. https://doi.org/10.1145/312624.312682.
Hidasi, Balázs, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos
Tikk. 2015. “Session-Based Recommendations with Recurrent Neural
Networks.” ArXiv:1511.06939. https://arxiv.org/abs/1511.06939.
Higham, Nicholas J. 2002. Accuracy and Stability of Numerical
Algorithms. 2nd ed. SIAM.
Hinton, Geoffrey, Oriol Vinyals, and Jeff Dean. 2015. “Distilling
the Knowledge in a Neural Network.” arXiv Preprint
arXiv:1503.02531.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising
Diffusion Probabilistic Models.” Advances in
Neural Information Processing
Systems 33: 6840–51. https://arxiv.org/abs/2006.11239.
Ho, Jonathan, and Tim Salimans. 2022. “Classifier-Free Diffusion
Guidance.” arXiv Preprint arXiv:2207.12598.
Hochreiter, Sepp, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber.
2001. “Gradient Flow in Recurrent Nets: The Difficulty of Learning
Long-Term Dependencies.” In A Field
Guide to Dynamical Recurrent
Neural Networks. IEEE
Press. https://doi.org/10.18034/ei.v8i2.570.
Hochreiter, Sepp, and Jürgen Schmidhuber. 1997. “Long Short-Term
Memory.” Neural Computation 9 (8): 1735–80.
https://doi.org/10.1162/neco.1997.9.8.1735.
Hoeffding, Wassily. 1963. “Probability Inequalities for Sums of
Bounded Random Variables.” Journal of the American
Statistical Association 58 (301): 13–30.
Hoerl, Arthur E, and Robert W Kennard. 1970. “Ridge Regression:
Biased Estimation for Nonorthogonal Problems.”
Technometrics 12 (1): 55–67.
Hoffmann, Jordan, Sebastian Borgeaud, Arthur
Mensch, et al. 2022. “Training Compute-Optimal Large
Language Models.” ArXiv:2203.15556. https://arxiv.org/abs/2203.15556.
Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020.
“The Curious Case of Neural Text Degeneration.”
International Conference on Learning Representations. https://arxiv.org/abs/1904.09751.
Horn, Roger A., and Charles R. Johnson. 2012. Matrix Analysis.
2nd ed. Cambridge University Press.
Hornik, Kurt. 1991. “Approximation Capabilities of Multilayer
Feedforward Networks.” Neural Networks 4 (2): 251–57. https://doi.org/10.1016/0893-6080(91)90009-T.
Horowitz, Mark. 2014. “1.1 Computing’s Energy Problem (and What We
Can Do about It).” IEEE
International Solid-State
Circuits Conference (ISSCC),
10–14. https://doi.org/10.1109/ISSCC.2014.6757323.
Howard, Andrew G., Menglong Zhu, Bo Chen, et al. 2017.
“MobileNets: Efficient Convolutional Neural Networks
for Mobile Vision Applications.”
ArXiv:1704.04861. https://arxiv.org/abs/1704.04861.
Hu, Edward J., Yelong Shen, Phillip Wallis, et al. 2021.
“LoRA: Low-Rank Adaptation of Large Language
Models.” arXiv Preprint arXiv:2106.09685.
Hu, Jie, Li Shen, and Gang Sun. 2018. “Squeeze-and-Excitation
Networks.” Proceedings of the IEEE
Conference on Computer Vision and
Pattern Recognition, 7132–41. https://doi.org/10.1109/cvpr.2018.00745.
Hu, Shengding, Yuge Tu, Xu Han, et al. 2024.
“MiniCPM: Unveiling the Potential of Small Language
Models with Scalable Training Strategies.”
ArXiv:2404.06395. https://arxiv.org/abs/2404.06395.
Hu, Yifan, Yehuda Koren, and Chris Volinsky. 2008. “Collaborative
Filtering for Implicit Feedback Datasets.” 2008 8th
IEEE International Conference on
Data Mining, 263–72. https://doi.org/10.1109/icdm.2008.22.
Hu, Zhiqiang, Roy Ka-Wei Lee, Charu C. Aggarwal, and Aston Zhang. 2022.
“Text Style Transfer: A Review and Experimental
Evaluation.” SIGKDD Explor.
Newsl. 24 (1). https://doi.org/10.1145/3544903.3544906.
Huang, Cheng-Zhi Anna, Ashish Vaswani, Jakob Uszkoreit, et al. 2018.
“Music Transformer: Generating Music with Long-Term
Structure.” International
Conference on Learning
Representations. https://arxiv.org/abs/1809.04281.
Huang, Gao, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger.
2017. “Densely Connected Convolutional Networks.”
Proceedings of the IEEE Conference on
Computer Vision and Pattern
Recognition, 4700–4708. https://doi.org/10.1109/cvpr.2017.243.
Huang, Gao, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Weinberger.
2016. “Deep Networks with Stochastic Depth.” European
Conference on Computer
Vision. https://arxiv.org/abs/1603.09382.
Huang, Zhiheng, Wei Xu, and Kai Yu. 2015. “Bidirectional
LSTM–CRF Models for Sequence Tagging.”
ArXiv:1508.01991. https://arxiv.org/abs/1508.01991.
Hubel, David H, and Torsten N Wiesel. 1959. “Receptive Fields of
Single Neurones in the Cat’s Striate Cortex.” Journal of
Physiology 148 (3): 574–91. https://doi.org/10.1113/jphysiol.1959.sp006308.
Hubel, David H, and Torsten N Wiesel. 1962. “Receptive Fields,
Binocular Interaction and Functional Architecture in the Cat’s Visual
Cortex.” Journal of Physiology 160 (1):
106–54. https://doi.org/10.1113/jphysiol.1962.sp006837.
Hubel, David H, and Torsten N Wiesel. 1968. “Receptive Fields and
Functional Architecture of Monkey Striate Cortex.” Journal of
Physiology 195 (1): 215–43. https://doi.org/10.1113/jphysiol.1968.sp008455.
Hutchinson, Michael F. 1989. “A Stochastic Estimator of the Trace
of the Influence Matrix for Laplacian Smoothing
Splines.” Communications in Statistics—Simulation and
Computation 18 (3): 1059–76.
Hutter, F., H. Hoos, and K. Leyton-Brown. 2011. “Sequential
Model-Based Optimization for General Algorithm Configuration.”
Proceedings of the Fifth International
Conference on Learning and
Intelligent Optimization
(LION’11). https://doi.org/10.1007/978-3-642-25566-3_40.
Hutter, F., L. Kotthoff, and J. Vanschoren, eds. 2019. Automated
Machine Learning: Methods,
Systems, Challenges. Springer. https://doi.org/10.1007/978-3-030-05318-5.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical
Models by Score Matching.” Journal of Machine Learning
Research 6: 695–709.
IBM Granite Team. 2025. IBM Granite 4.0:
Hyper-Efficient, High-Performance Hybrid Models. Https://www.ibm.com/new/announcements/ibm-granite-4-0.
IEEE. 2019. IEEE Standard for Floating-Point
Arithmetic. IEEE Std 754-2019.
Inan, Hakan, Khashayar Khosravi, and Richard Socher. 2017. “Tying
Word Vectors and Word Classifiers: A Loss Framework for Language
Modeling.” International Conference on Learning
Representations.
Ioffe, Sergey. 2017. “Batch Renormalization: Towards Reducing
Minibatch Dependence in Batch-Normalized Models.” Advances in
Neural Information Processing
Systems, 1945–53. https://proceedings.mlr.press/v70/ioffe17a.html.
Ioffe, Sergey, and Christian Szegedy. 2015. “Batch Normalization:
Accelerating Deep Network Training by Reducing Internal Covariate
Shift.” International Conference on Machine Learning,
448–56. https://arxiv.org/abs/1502.03167.
Isensee, Fabian, Paul F. Jaeger, Simon A. A. Kohl, Jens Petersen, and
Klaus H. Maier-Hein. 2021. “NnU-Net: A
Self-Configuring Method for Deep Learning-Based Biomedical Image
Segmentation.” Nature Methods 18: 203–11.
https://doi.org/10.1038/s41592-020-01008-z.
Izmailov, Pavel, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and
Andrew Gordon Wilson. 2018. “Averaging Weights Leads to Wider
Optima and Better Generalization.” Uncertainty in Artificial
Intelligence, 876–85. https://arxiv.org/abs/1803.05407.
Jacobs, Robert A., Michael I. Jordan, Steven J. Nowlan, and Geoffrey E.
Hinton. 1991. “Adaptive Mixtures of Local Experts.”
Neural Computation 3 (1): 79–87.
Jacot, Arthur, Franck Gabriel, and Clément Hongler. 2018. “Neural
Tangent Kernel: Convergence and Generalization in Neural
Networks.” Advances in Neural
Information Processing
Systems 31. https://doi.org/10.1145/3406325.3465355.
Jaeger, Herbert. 2002. Tutorial on Training Recurrent Neural
Networks, Covering BPTT, RTRL,
EKF and the “Echo State Network”
Approach. GMD-Forschungszentrum
Informationstechnik Bonn.
Jaegle, Andrew, Sebastian Borgeaud, Jean-Baptiste
Alayrac, et al. 2022. “Perceiver IO: A General
Architecture for Structured Inputs & Outputs.”
International Conference on Learning Representations.
Jaegle, Andrew, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew
Zisserman, and Joao Carreira. 2021. “Perceiver: General Perception
with Iterative Attention.” International Conference on
Machine Learning, 4651–64.
Jain, Sarthak, and Byron C. Wallace. 2019. “Attention Is Not
Explanation.” Proceedings of the 2019 Conference of the North
American Chapter of the Association for Computational Linguistics,
3543–56.
Jamieson, K., and A. Talwalkar. 2016. “Non-Stochastic Best Arm
Identification and Hyperparameter Optimization.” Proceedings
of the 19th International Conference on
Artificial Intelligence and
Statistics. https://proceedings.mlr.press/v51/jamieson16.html.
Jang, Eric, Shixiang Gu, and Ben Poole. 2017. “Categorical
Reparameterization with Gumbel-Softmax.”
International Conference on Learning Representations.
Jelassi, Samy, David Brandfonbrener, Sham M. Kakade, and Eran Malach.
2024. “Repeat After Me: Transformers Are Better Than State Space
Models at Copying.” International Conference on Machine
Learning, 21502–21. https://arxiv.org/abs/2402.01032.
Jelinek, Frederick, Robert L. Mercer, Lalit R. Bahl, and James K. Baker.
1977. “Perplexity—a Measure of the Difficulty of Speech
Recognition Tasks.” Journal of the Acoustical Society of
America 62 (S1): S63.
Jenatton, R., C. Archambeau, J. González, and M. Seeger. 2017.
“Bayesian Optimization with tree-Structured dependencies.” Proceedings of the 34th
International Conference on
Machine Learning (ICML’17).
https://proceedings.mlr.press/v70/jenatton17a.html.
Jia, Xianyan, Shutao Song, Wei He, et al.
2018. “Highly Scalable Deep Learning Training System with
Mixed-Precision: Training ImageNet in Four
Minutes.” ArXiv:1807.11205. https://arxiv.org/abs/1807.11205.
Jia, Yangqing, Evan Shelhamer, Jeff Donahue, et al. 2014. “Caffe:
Convolutional Architecture for Fast Feature Embedding.”
Proceedings of the 22nd ACM International
Conference on Multimedia, 675–78. https://doi.org/10.1145/2647868.2654889.
Jiang, Albert Q., Alexandre Sablayrolles, Arthur
Mensch, et al. 2023. “Mistral 7B.”
arXiv Preprint arXiv:2310.06825.
Jiang, Albert Q., Alexandre Sablayrolles, Antoine
Roux, et al. 2024. “Mixtral of Experts.” arXiv
Preprint arXiv:2401.04088.
Jordan, Keller, and contributors. 2024.
Modded-Nanogpt: Speedrunning the NanoGPT Baseline.
Https://github.com/KellerJordan/modded-nanogpt.
Jordan, Keller, Yuchen Jin, Vlado Boza, et al. 2024. Muon: An
Optimizer for Hidden Layers in Neural Networks. Https://kellerjordan.github.io/posts/muon/.
Jordan, Richard, David Kinderlehrer, and Felix Otto. 1998. “The
Variational Formulation of the Fokker–Planck
Equation.” SIAM Journal on Mathematical Analysis 29 (1):
1–17.
Jouppi, Norman P, Cliff Young, Nishant Patil, et
al. 2017. “In-Datacenter Performance Analysis of a Tensor
Processing Unit.” 2017 ACM/IEEE 44th
Annual International Symposium on
Computer Architecture
(ISCA), 1–12. https://doi.org/10.1145/3140659.3080246.
Kaddour, Jean. 2022. “Stop Wasting My Time! Saving
Days of ImageNet and BERT
Training with Latest Weight Averaging.” arXiv Preprint
arXiv:2209.14981.
Kahan, William. 1965. “Further Remarks on Reducing Truncation
Errors.” Communications of the ACM 8 (1): 40.
Kalchbrenner, Nal, Edward Grefenstette, and Phil Blunsom. 2014. “A
Convolutional Neural Network for Modelling Sentences.”
ArXiv:1404.2188. https://arxiv.org/abs/1404.2188.
Kalman, Rudolf E. 1960. “A New Approach to Linear Filtering and
Prediction Problems.” Journal of Basic Engineering 82
(1): 35–45.
Kalra, Dayal Singh, and Maissam Barkeshli. 2024. “Why Warmup the
Learning Rate? Underlying Mechanisms and
Improvements.” Advances in Neural Information Processing
Systems 37. https://arxiv.org/abs/2406.09405.
Kaplan, Jared, Sam McCandlish, Tom Henighan, et al. 2020. “Scaling
Laws for Neural Language Models.”
ArXiv:2001.08361. https://arxiv.org/abs/2001.08361.
Karimi, Hamed, Julie Nutini, and Mark Schmidt. 2016. “Linear
Convergence of Gradient and Proximal-Gradient Methods Under the
Polyak–Łojasiewicz Condition.”
Joint European Conference on Machine Learning and Knowledge
Discovery in Databases (ECML PKDD), 795–811.
Karnin, Z., T. Koren, and O. Somekh. 2013. “Almost Optimal
Exploration in Multi-Armed Bandits.” Proceedings of the 30th
International Conference on
Machine Learning (ICML’13).
https://proceedings.mlr.press/v28/karnin13.html.
Karras, Tero, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018.
“Progressive Growing of GANs for Improved Quality,
Stability, and Variation.” International Conference on
Learning Representations. https://arxiv.org/abs/1710.10196.
Karras, Tero, Miika Aittala, Timo Aila, and Samuli Laine. 2022.
“Elucidating the Design Space of Diffusion-Based Generative
Models.” Advances in Neural Information Processing
Systems 35.
Karras, Tero, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila,
and Samuli Laine. 2024. “Analyzing and Improving the Training
Dynamics of Diffusion Models.” Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition.
Karush, William. 1939. “Minima of Functions of Several Variables
with Inequalities as Side Constraints.” Master’s thesis,
Department of Mathematics, University of Chicago.
Kasimbeg, Priya, Frank Schneider, Runa Eschenhagen,
et al. 2025. “Accelerating Neural Network Training: An
Analysis of the AlgoPerf Competition.”
International Conference on Learning Representations. https://arxiv.org/abs/2502.15015.
Katharopoulos, Angelos, Apoorv Vyas, Nikolaos Pappas, and François
Fleuret. 2020. “Transformers Are RNNs: Fast
Autoregressive Transformers with Linear Attention.”
International Conference on Machine Learning.
Kazemnejad, Amirhossein, Inkit Padhi, Karthikeyan Natesan Ramamurthy,
Payel Das, and Siva Reddy. 2023. “The Impact of Positional
Encoding on Length Generalization in Transformers.” Advances
in Neural Information Processing Systems 36. https://arxiv.org/abs/2305.19466.
Keskar, Nitish Shirish, Dheevatsa Mudigere, Jorge Nocedal, Mikhail
Smelyanskiy, and Ping Tak Peter Tang. 2017. “On Large-Batch
Training for Deep Learning: Generalization Gap and Sharp Minima.”
International Conference on Learning Representations. https://arxiv.org/abs/1609.04836.
Kidger, Patrick. 2022. “On Neural Differential Equations.”
arXiv Preprint arXiv:2202.02435.
Kim, Jaeyoung, Mostafa El-Khamy, and Jungwon Lee. 2017. “Residual
LSTM: Design of a Deep Recurrent Architecture for Distant
Speech Recognition.” ArXiv:1701.03360. https://arxiv.org/abs/1701.03360.
Kim, Yoon. 2014. “Convolutional Neural Networks for Sentence
Classification.” ArXiv:1408.5882. https://arxiv.org/abs/1408.5882.
Kimeldorf, G. S., and G. Wahba. 1971. “Some Results on
Tchebycheffian Spline Functions.” J.
Math. Anal. Appl. 33: 82–95.
https://doi.org/10.1214/aoms/1177693054.
Kimi Team. 2025a. “Kimi K2: Open Agentic
Intelligence.” arXiv Preprint arXiv:2507.20534.
Kimi Team. 2025b. “Kimi Linear: An Expressive, Efficient Attention
Architecture.” arXiv Preprint arXiv:2510.26692.
Kingma, Diederik P, and Jimmy Ba. 2015. “Adam: A Method for
Stochastic Optimization.” International Conference on
Learning Representations. https://arxiv.org/abs/1412.6980.
Kingma, Diederik P., Tim Salimans, Ben Poole, and Jonathan Ho. 2021.
“Variational Diffusion Models.” Advances in Neural
Information Processing Systems 34.
Kingma, Diederik P., and Max Welling. 2014. “Auto-Encoding
Variational Bayes.” International
Conference on Learning
Representations (ICLR). https://arxiv.org/abs/1312.6114.
Kipf, Thomas N, and Max Welling. 2017. “Semi-Supervised
Classification with Graph Convolutional Networks.”
International Conference on Learning Representations. https://arxiv.org/abs/1609.02907.
Kitaev, Nikita, Lukasz Kaiser, and Anselm Levskaya. 2020.
“Reformer: The Efficient Transformer.” International
Conference on Learning Representations.
Kloeden, Peter E., and Eckhard Platen. 1992. Numerical Solution of
Stochastic Differential Equations. Springer.
Koh, Pang Wei, Shiori Sagawa, Henrik Marklund, et al. 2021.
“WILDS: A Benchmark of in-the-Wild Distribution
Shifts.” International Conference on Machine Learning,
5637–64. https://arxiv.org/abs/2012.07421.
Kohavi, Ron. 1995. “A Study of Cross-Validation and Bootstrap for
Accuracy Estimation and Model Selection.” International
Joint Conference on Artificial
Intelligence 14: 1137–45.
Koller, Daphne, and Nir Friedman. 2009. Probabilistic
Graphical Models: Principles and
Techniques. MIT Press. https://doi.org/10.7551/mitpress/7432.001.0001.
Kolmogorov, Andrey. 1933. “Sulla Determinazione Empirica Di Una
Legge Di Distribuzione.” Inst. Ital.
Attuari, Giorn. 4: 83–91. https://doi.org/10.1007/BF03017337.
Kolmogorov, Andrey N. 1931. “Über Die Analytischen
Methoden in Der
Wahrscheinlichkeitsrechnung.” Mathematische
Annalen 104: 415–58.
Kolter, Zico. 2008. “Linear Algebra Review and Reference.”
Available Online:
Http://Cs229.stanford.edu/Section/Cs229-Linalg.pdf. http://cs229.stanford.edu/section/cs229-linalg.pdf.
Koren, Yehuda, Robert Bell, and Chris Volinsky. 2009. “Matrix
Factorization Techniques for Recommender Systems.”
Computer 42 (8): 30–37. https://doi.org/10.1109/mc.2009.263.
Kosson, Atli, Bettina Messmer, and Martin Jaggi. 2024. “Rotational
Equilibrium: How Weight Decay Balances Learning Across Neural
Networks.” International Conference on Machine Learning.
https://arxiv.org/abs/2305.17212.
Kosson, Atli, Jeremy Welborn, Yang Liu, Martin Jaggi, and Xi Chen. 2025.
“Weight Decay May Matter More Than muP
for Learning Rate Transfer in Practice.” arXiv Preprint
arXiv:2510.19093.
Kraft, Leon G. 1949. “A Device for Quantizing, Grouping, and
Coding Amplitude-Modulated Pulses.” Master’s thesis,
Massachusetts Institute of Technology.
Kraskov, Alexander, Harald Stögbauer, and Peter Grassberger. 2004.
“Estimating Mutual Information.” Physical Review E
69 (6): 066138.
Krizhevsky, Alex, Ilya Sutskever, and Geoffrey E Hinton. 2012.
“ImageNet Classification with Deep Convolutional
Neural Networks.” Advances in Neural
Information Processing
Systems, 1097–105. https://doi.org/10.5555/2999134.2999257.
Krogh, Anders, and John A Hertz. 1992. “A Simple Weight Decay Can
Improve Generalization.” Advances in Neural
Information Processing
Systems, 950–57. https://doi.org/10.5555/2986916.2987033.
Kuhn, Harold W., and Albert W. Tucker. 1951. “Nonlinear
Programming.” Proceedings of the Second Berkeley Symposium on
Mathematical Statistics and Probability, 481–92.
Kullback, Solomon, and Richard A. Leibler. 1951. “On Information
and Sufficiency.” Annals of Mathematical Statistics 22
(1): 79–86.
Kung, Sun Yuan. 1988. “VLSI Array
Processors.” Prentice Hall.
Kunstner, Frederik, Jacques Chen, Jonathan Wilder Lavington, and Mark
Schmidt. 2023. “Noise Is Not the Main Factor Behind the Gap
Between SGD and Adam on Transformers, but Sign
Descent Might Be.” International Conference on Learning
Representations. https://arxiv.org/abs/2304.13960.
Kunstner, Frederik, Robin Yadav, Alan Milligan, Mark Schmidt, and
Alberto Bietti. 2024. “Heavy-Tailed Class Imbalance and Why
Adam Outperforms Gradient Descent on Language
Models.” Advances in Neural Information Processing
Systems 37. https://arxiv.org/abs/2402.19449.
Kuzovkin, Ilya, Raul Vicente, Mathilde Petton, et al. 2018.
“Activations of Deep Convolutional Neural Networks Are Aligned
with Gamma Band Activity of Human Visual Cortex.”
Communications Biology 1 (1): 1–12. https://doi.org/10.1038/s42003-018-0110-y.
Lahoti, Aakash, Kevin Y. Li, Berlin Chen, et al. 2026. “Mamba-3:
Improved Sequence Modeling Using State Space Principles.”
arXiv Preprint arXiv:2603.15569.
Laplace, Pierre-Simon. 1814. Essai Philosophique Sur Les
Probabilités. Courcier.
Lavin, Andrew, and Scott Gray. 2016. “Fast Algorithms for
Convolutional Neural Networks.” Proceedings of the
IEEE Conference on Computer
Vision and Pattern
Recognition, 4013–21. https://doi.org/10.1109/cvpr.2016.435.
Le, Quoc V. 2013. “Building High-Level Features Using Large Scale
Unsupervised Learning.” Proceedings of the IEEE
International Conference on
Acoustics, Speech and Signal
Processing, 8595–98. https://doi.org/10.1109/icassp.2013.6639343.
LeCun, Yann, Yoshua Bengio, and et al. 1995.
“Convolutional Networks for Images, Speech, and Time
Series.” In The Handbook of Brain
Theory and Neural Networks.
MIT Press. http://yann.lecun.com/exdb/publis/pdf/lecun-bengio-95a.pdf.
LeCun, Yann, Bernhard Boser, John S Denker, et al. 1989.
“Backpropagation Applied to Handwritten Zip Code
Recognition.” Neural Computation 1 (4):
541–51. https://doi.org/10.1162/neco.1989.1.4.541.
LeCun, Yann, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998.
“Gradient-Based Learning Applied to Document Recognition.”
Proceedings of the IEEE 86 (11): 2278–324. https://doi.org/10.1109/5.726791.
LeCun, Yann, Leon Bottou, G Orr, and Klaus-Robert Muller. 1998.
“Efficient Backprop.” In Neural Networks:
Tricks of the Trade. Springer. https://doi.org/10.1007/3-540-49430-8_2.
LeCun, Yann, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu
Jie Huang. 2006. “A Tutorial on Energy-Based Learning.” In
Predicting Structured Data. MIT Press.
LeCun, Yann, LD Jackel, Leon Bottou, et al.
1995. “Comparison of Learning Algorithms for Handwritten Digit
Recognition.” International
Conference on Artificial Neural
Networks, 53–60.
Lee, Hyunji, Wenhao Yu, Hongming Zhang, et al. 2025.
“Understanding and Enhancing Mamba-Transformer
Hybrids for Memory Recall and Language Modeling.” arXiv
Preprint arXiv:2510.26912.
Legendre, Adrien Marie. 1805. Mémoire Sur Les
Opérations
Trigonométriques: Dont Les
Résultats Dépendent
de La Figure de La Terre. F.
Didot.
Lepikhin, Dmitry, HyoukJoong Lee, Yuanzhong Xu, et al. 2021.
“GShard: Scaling Giant Models with Conditional
Computation and Automatic Sharding.” International Conference
on Learning Representations.
Leshno, Moshe, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken.
1993. “Multilayer Feedforward Networks with a Nonpolynomial
Activation Function Can Approximate Any Function.” Neural
Networks 6 (6): 861–67.
Lessard, Laurent, Benjamin Recht, and Andrew Packard. 2016.
“Analysis and Design of Optimization Algorithms via Integral
Quadratic Constraints.” SIAM Journal on Optimization 26
(1): 57–95.
Leviathan, Yaniv, Matan Kalman, and Yossi Matias. 2023. “Fast
Inference from Transformers via Speculative Decoding.”
International Conference on Machine Learning, 19274–86.
Levy, Omer, and Yoav Goldberg. 2014. “Neural Word Embedding as
Implicit Matrix Factorization.” Advances in Neural
Information Processing Systems 27.
Lewis, Mike, Yinhan Liu, Naman Goyal, et al. 2019.
“BART: Denoising Sequence-to-Sequence Pre-Training
for Natural Language Generation, Translation, and Comprehension.”
ArXiv:1910.13461. https://arxiv.org/abs/1910.13461.
Li, Chun-Liang, Wei-Cheng Chang, Yu Cheng, Yiming Yang, and Barnabás
Póczos. 2017. “MMD GAN: Towards Deeper
Understanding of Moment Matching Network.” Advances in Neural
Information Processing Systems 30.
Li, Junnan, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023.
“BLIP-2: Bootstrapping Language-Image Pre-Training
with Frozen Image Encoders and Large Language Models.”
International Conference on Machine Learning, 19730–42.
Li, L., K. Jamieson, A. Rostamizadeh, et al. 2018. “Massively
Parallel Hyperparameter Tuning.”
ArXiv:1810.05934. https://arxiv.org/abs/1810.05934.
Li, Mu. 2017. “Scaling Distributed
Machine Learning with System and
Algorithm Co-Design.” PhD thesis,
PhD Thesis, CMU. https://www.cs.cmu.edu/~muli/file/mu-thesis.pdf.
Li, Mu, David G Andersen, Jun Woo Park, et al. 2014. “Scaling
Distributed Machine Learning with the Parameter Server.” 11th
Symposium on Operating Systems
Design and Implementation (OSDI
14), 583–98. https://doi.org/10.1145/2640087.2644155.
Li, Mu, Tong Zhang, Yuqiang Chen, and Alexander J Smola. 2014.
“Efficient Mini-Batch Training for Stochastic
Optimization.” Proceedings of the 20th ACM
SIGKDD International Conference
on Knowledge Discovery and Data
Mining, 661–70. https://doi.org/10.1145/2623330.2623612.
Li, Shen, Yanli Zhao, Rohan Varma, et al. 2020.
“PyTorch Distributed: Experiences on
Accelerating Data Parallel Training.” Proceedings of the
VLDB Endowment 13 (12): 3005–18. https://doi.org/10.14778/3415478.3415530.
Li, Xiang, Shuo Chen, Xiaolin Hu, and Jian Yang. 2019.
“Understanding the Disharmony Between Dropout and Batch
Normalization by Variance Shift.” Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition, 2682–90. https://arxiv.org/abs/1801.05134.
Li, Yujia, Kevin Swersky, and Richard Zemel. 2015. “Generative
Moment Matching Networks.” Proceedings of the 32nd
International Conference on Machine Learning, 1718–27.
Liaw, R., E. Liang, R. Nishihara, P. Moritz, J. Gonzalez, and I. Stoica.
2018. “Tune: A Research Platform for Distributed
Model Selection and Training.”
ArXiv:1807.05118. https://arxiv.org/abs/1807.05118.
Lieber, Opher, Barak Lenz, Hofit Bata, et al. 2024.
“Jamba: A Hybrid Transformer-Mamba
Language Model.” arXiv Preprint arXiv:2403.19887.
Lin, Jianhua. 1991. “Divergence Measures Based on the
Shannon Entropy.” IEEE Transactions on
Information Theory 37 (1): 145–51.
Lin, Min, Qiang Chen, and Shuicheng Yan. 2013. “Network in
Network.” ArXiv:1312.4400. https://arxiv.org/abs/1312.4400.
Lin, Tsung-Yi, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár.
2017. “Focal Loss for Dense Object Detection.”
Proceedings of the IEEE International
Conference on Computer
Vision, 2980–88. https://doi.org/10.1109/iccv.2017.324.
Lin, Yuanqing, F Lv, S Zhu, et al. 2010.
“ImageNet Classification: Fast Descriptor Coding and
Large-Scale SVM Training.”
Large Scale Visual Recognition Challenge, ahead of print. https://doi.org/10.1109/cvpr.2010.5539970.
Lin, Zhouhan, Minwei Feng, Cicero Nogueira dos Santos, et al. 2017.
“A Structured Self-Attentive Sentence Embedding.”
ArXiv:1703.03130. https://arxiv.org/abs/1703.03130.
Linnainmaa, Seppo. 1970. “The Representation of the Cumulative
Rounding Error of an Algorithm as a Taylor Expansion of the
Local Rounding Errors.” Master’s thesis, University of Helsinki.
Lipman, Yaron, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and
Matt Le. 2023. “Flow Matching for Generative Modeling.”
International Conference on Learning Representations.
Lipton, Zachary C, and Jacob Steinhardt. 2018. “Troubling Trends
in Machine Learning Scholarship.” Communications of the
ACM 17: 45–77. https://doi.org/10.1145/3317287.3328534.
Lipton, Zachary, Yu-Xiang Wang, and Alexander Smola. 2018.
“Detecting and Correcting for Label Shift with Black Box
Predictors.” International Conference on Machine
Learning, 3122–30. https://arxiv.org/abs/1802.03916.
Liu, Bo, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, and Qiang Liu.
2024. “Longhorn: State Space Models Are Amortized Online
Learners.” arXiv Preprint arXiv:2407.14207.
Liu, Chaoyue, Libin Zhu, and Mikhail Belkin. 2022. “Loss
Landscapes and Optimization in over-Parameterized Non-Linear Systems and
Neural Networks.” Applied and Computational Harmonic
Analysis 59: 85–116.
Liu, Dong C, and Jorge Nocedal. 1989. “On the Limited Memory
BFGS Method for Large Scale Optimization.”
Mathematical Programming 45 (1): 503–28. https://doi.org/10.1007/bf01589116.
Liu, Hanxiao, Karen Simonyan, and Yiming Yang. 2018.
“DARTS: Differentiable Architecture Search.”
ArXiv:1806.09055. https://arxiv.org/abs/1806.09055.
Liu, Hong, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. 2024.
“Sophia: A Scalable Stochastic Second-Order Optimizer for Language
Model Pre-Training.” International Conference on Learning
Representations. https://arxiv.org/abs/2305.14342.
Liu, Jingyuan, Jianlin Su, Xingcheng Yao, et
al. 2025. “Muon Is Scalable for LLM
Training.” arXiv Preprint arXiv:2502.16982.
Liu, Qiang, Jason D. Lee, and Michael I. Jordan. 2016. “A
Kernelized Stein Discrepancy for Goodness-of-Fit
Tests.” Proceedings of the 33rd International Conference on
Machine Learning, 276–84.
Liu, Qiang, and Dilin Wang. 2016. “Stein Variational Gradient
Descent: A General Purpose Bayesian Inference
Algorithm.” Advances in Neural Information Processing
Systems 29.
Liu, Shiwei, Tianlong Chen, Xiaohan Chen, et al. 2022. “More
ConvNets in the 2020s: Scaling up Kernels
Beyond 51x51 Using Sparsity.”
ArXiv:2207.03620. https://arxiv.org/abs/2207.03620.
Liu, Wei, Dragomir Anguelov, Dumitru Erhan, et al. 2016.
“SSD: Single Shot Multibox Detector.”
European Conference on Computer
Vision, 21–37. https://doi.org/10.1007/978-3-319-46448-0_2.
Liu, Xingchao, Chengyue Gong, and Qiang Liu. 2023. “Flow Straight
and Fast: Learning to Generate and Transfer Data with Rectified
Flow.” International Conference on Learning
Representations.
Liu, Yinhan, Myle Ott, Naman Goyal, et al. 2019.
“RoBERTa: A Robustly Optimized BERT
Pretraining Approach.” ArXiv:1907.11692. https://arxiv.org/abs/1907.11692.
Liu, Ze, Yutong Lin, Yue Cao, et al. 2021. “Swin Transformer:
Hierarchical Vision Transformer Using Shifted Windows.”
Proceedings of the IEEE/CVF
International Conference on
Computer Vision, 10012–22. https://doi.org/10.1109/iccv48922.2021.00986.
Liu, Zhuang, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor
Darrell, and Saining Xie. 2022. “A ConvNet for the
2020s.” ArXiv:2201.03545. https://arxiv.org/abs/2201.03545.
Łojasiewicz, Stanisław. 1963. “Une
Propriété Topologique Des Sous-Ensembles
Analytiques réels.” In Les
Équations Aux dérivées
Partielles. Éditions du CNRS.
Long, Jonathan, Evan Shelhamer, and Trevor Darrell. 2015. “Fully
Convolutional Networks for Semantic Segmentation.”
Proceedings of the IEEE Conference on
Computer Vision and Pattern
Recognition, 3431–40. https://doi.org/10.1109/tpami.2016.2572683.
Loshchilov, Ilya, and Frank Hutter. 2016. “SGDR:
Stochastic Gradient Descent with Warm Restarts.”
ArXiv:1608.03983. https://arxiv.org/abs/1608.03983.
Loshchilov, Ilya, and Frank Hutter. 2019. “Decoupled Weight Decay
Regularization.” International Conference on Learning
Representations. https://arxiv.org/abs/1711.05101.
Lowe, David G. 2004. “Distinctive Image Features from
Scale-Invariant Keypoints.” International
Journal of Computer Vision
60 (2): 91–110. https://doi.org/10.1023/b:visi.0000029664.99615.94.
Luo, Calvin. 2022. “Understanding Diffusion Models: A Unified
Perspective.” arXiv Preprint arXiv:2208.11970.
Luo, Ping, Xinjiang Wang, Wenqi Shao, and Zhanglin Peng. 2018.
“Towards Understanding Regularization in Batch
Normalization.” ArXiv:1809.00846. https://arxiv.org/abs/1809.00846.
Luo, Wenjie, Yujia Li, Raquel Urtasun, and Richard Zemel. 2016.
“Understanding the Effective Receptive Field in Deep Convolutional
Neural Networks.” Advances in Neural
Information Processing
Systems. https://arxiv.org/abs/1701.04128.
Maas, Andrew L, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng,
and Christopher Potts. 2011. “Learning Word Vectors for Sentiment
Analysis.” Proceedings of the 49th Annual
Meeting of the Association for
Computational Linguistics: Human
Language Technologies, Volume
1, 142–50. https://aclanthology.org/P11-1015.
Mack, Yue-Pok, and Bernard W Silverman. 1982. “Weak and Strong
Uniform Consistency of Kernel Regression Estimates.”
Zeitschrift für Wahrscheinlichkeitstheorie
Und Verwandte Gebiete 61 (3): 405–15. https://doi.org/10.1007/bf00539840.
MacKay, David JC. 2003. Information Theory,
Inference and Learning
Algorithms. Cambridge University
Press. https://www.inference.org.uk/mackay/itila/book.html.
Maclaurin, D., D. Duvenaud, and R. Adams. 2015. “Gradient-Based
Hyperparameter Optimization Through Reversible Learning.”
Proceedings of the 32nd International
Conference on Machine Learning
(ICML’15). https://proceedings.mlr.press/v37/maclaurin15.html.
Maddison, Chris J., Andriy Mnih, and Yee Whye Teh. 2017. “The
Concrete Distribution: A Continuous Relaxation of Discrete Random
Variables.” International Conference on Learning
Representations.
Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras,
and Adrian Vladu. 2018. “Towards Deep Learning Models Resistant to
Adversarial Attacks.” International Conference on Learning
Representations. https://arxiv.org/abs/1706.06083.
Malladi, Sadhika, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora.
2022. “On the SDEs and Scaling Rules for Adaptive
Gradient Algorithms.” Advances in Neural Information
Processing Systems 35. https://arxiv.org/abs/2205.10287.
Mangram, Myles E. 2013. “A Simplified Perspective of the
Markowitz Portfolio Theory.” Global
Journal of Business Research
7 (1): 59–70. https://scholarworks.iu.edu/journals/index.php/jiuspa/article/view/4517.
Manning, Christopher D., Prabhakar Raghavan, and Hinrich Schütze. 2008.
Introduction to Information
Retrieval. Cambridge University
Press. https://doi.org/10.1017/CBO9780511809071.
Marchenko, Vladimir A., and Leonid A. Pastur. 1967. “Distribution
of Eigenvalues for Some Sets of Random Matrices.” Mathematics
of the USSR-Sbornik 1 (4): 457–83.
Maron, Melvin E. 1961. “Automatic Indexing: An Experimental
Inquiry.” Journal of the ACM 8 (3): 404–17.
Martens, James, and Roger Grosse. 2015. “Optimizing Neural
Networks with Kronecker-Factored Approximate
Curvature.” Proceedings of the 32nd International Conference
on Machine Learning (ICML), 2408–17.
Martin, Charles H., and Michael W. Mahoney. 2021. “Implicit
Self-Regularization in Deep Neural Networks: Evidence from Random Matrix
Theory and Implications for Learning.” Journal of Machine
Learning Research 22 (165): 1–73.
Martins, André F. T., and Ramón F. Astudillo. 2016. “From Softmax
to Sparsemax: A Sparse Model of Attention and Multi-Label
Classification.” Proceedings of the 33rd International
Conference on Machine Learning (ICML), 1614–23.
Maruyama, Gisiro. 1955. “Continuous Markov Processes
and Stochastic Equations.” Rendiconti Del Circolo Matematico
Di Palermo 4: 48–90.
Matthews, Alexander G de G, Mark Rowland, Jiri Hron, Richard E Turner,
and Zoubin Ghahramani. 2018. “Gaussian Process Behaviour in Wide
Deep Neural Networks.” ArXiv:1804.11271. https://arxiv.org/abs/1804.11271.
McAllester, David, and Karl Stratos. 2020. “Formal Limitations on
the Measurement of Mutual Information.” International
Conference on Artificial Intelligence and Statistics, 875–84.
McCandlish, Sam, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. 2018.
“An Empirical Model of Large-Batch Training.” arXiv
Preprint arXiv:1812.06162.
McCann, Bryan, James Bradbury, Caiming Xiong, and Richard Socher. 2017.
“Learned in Translation: Contextualized Word
Vectors.” Advances in Neural
Information Processing
Systems, 6294–305. https://doi.org/10.5555/3294996.3295037.
McCulloch, Warren S, and Walter Pitts. 1943. “A Logical Calculus
of the Ideas Immanent in Nervous Activity.” Bulletin of
Mathematical Biophysics 5 (4): 115–33. https://doi.org/10.1016/s0092-8240(05)80006-0.
McDiarmid, Colin. 1989. “On the Method of Bounded
Differences.” In Surveys in Combinatorics. Cambridge
University Press.
McKinney, Wes. 2010. “Data Structures for Statistical Computing in
Python.” Proceedings of the 9th Python in
Science Conference, 56–61. https://doi.org/10.25080/Majora-92bf1922-00a.
McMahan, H Brendan, Gary Holt, David Sculley, et
al. 2013. “Ad Click Prediction: A View from the
Trenches.” Proceedings of the 19th ACM SIGKDD
International Conference on
Knowledge Discovery and Data
Mining, 1222–30. https://doi.org/10.1145/2487575.2488200.
McMillan, Brockway. 1956. “Two Inequalities Implied by Unique
Decipherability.” IRE Transactions on Information Theory
2 (4): 115–16.
Mead, Carver, and Lynn Conway. 1980. Introduction to
VLSI Systems. Addison-Wesley.
Meng, Fanxu, Zhaohui Wang, and Muhan Zhang. 2024.
“PiSSA: Principal Singular Values and Singular
Vectors Adaptation of Large Language Models.” Advances in
Neural Information Processing Systems.
Merity, Stephen, Caiming Xiong, James Bradbury, and Richard Socher.
2016. “Pointer Sentinel Mixture Models.”
ArXiv:1609.07843. https://arxiv.org/abs/1609.07843.
Merrill, William, Shane Arora, Dirk Groeneveld, and Hannaneh Hajishirzi.
2025. “Critical Batch Size Revisited: A Simple Empirical Approach
to Large-Batch Language Model Training.” arXiv Preprint
arXiv:2505.23971.
Merrill, William, Jackson Petty, and Ashish Sabharwal. 2024. “The
Illusion of State in State-Space Models.” International
Conference on Machine Learning. https://arxiv.org/abs/2404.08819.
Meta AI. 2025. The Llama 4 Herd: The Beginning of a New
Era of Natively Multimodal AI Innovation. Https://ai.meta.com/blog/llama-4-multimodal-intelligence/.
Metropolis, Nicholas, and Stanislaw Ulam. 1949. “The
Monte Carlo Method.” Journal of the
American Statistical Association 44 (247): 335–41.
Micchelli, Charles A. 1984. “Interpolation of Scattered Data:
Distance Matrices and Conditionally Positive Definite Functions.”
In Approximation Theory and Spline
Functions. Springer. https://doi.org/10.1007/bf01893414.
Micikevicius, Paulius, Sharan Narang, Jonah Alben, et al. 2018.
“Mixed Precision Training.” International Conference on
Learning Representations.
Micikevicius, Paulius, Dusan Stosic, Neil Burgess, et al. 2022.
“FP8 Formats for Deep Learning.” arXiv
Preprint arXiv:2209.05433.
Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013.
“Efficient Estimation of Word Representations in Vector
Space.” ArXiv:1301.3781. https://arxiv.org/abs/1301.3781.
Mikolov, Tomas, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean.
2013. “Distributed Representations of Words and Phrases and Their
Compositionality.” Advances in Neural
Information Processing
Systems, 3111–19. https://doi.org/10.5555/2999792.2999959.
Milakov, Maxim, and Natalia Gimelshein. 2018. “Online Normalizer
Calculation for Softmax.” arXiv Preprint
arXiv:1805.02867.
Miller, George A. 1995. “WordNet: A Lexical Database
for English.” Communications of the
ACM 38 (11): 39–41. https://doi.org/10.1145/219717.219748.
Milstein, Grigori N. 1975. “Approximate Integration of Stochastic
Differential Equations.” Theory of Probability and Its
Applications 19 (3): 557–62.
MiniMax. 2025. “MiniMax-01: Scaling Foundation Models
with Lightning Attention.” arXiv Preprint
arXiv:2501.08313.
Mironov, Ilya. 2017. “Rényi Differential
Privacy.” 2017 IEEE 30th Computer Security Foundations
Symposium (CSF), 263–75.
Mirsky, Leon. 1960. “Symmetric Gauge Functions and Unitarily
Invariant Norms.” The Quarterly Journal of Mathematics
11 (1): 50–59.
Miyato, Takeru, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida.
2018. “Spectral Normalization for Generative Adversarial
Networks.” International Conference on Learning
Representations.
Mnih, Volodymyr, Nicolas Heess, Alex Graves, et
al. 2014. “Recurrent Models of Visual Attention.”
Advances in Neural Information
Processing Systems, 2204–12. https://doi.org/10.5555/2969033.2969073.
Mnih, Volodymyr, Koray Kavukcuoglu, David Silver, et al. 2013.
“Playing Atari with Deep Reinforcement
Learning.” ArXiv:1312.5602.
Mnih, Volodymyr, Koray Kavukcuoglu, David Silver,
et al. 2015. “Human-Level Control Through Deep
Reinforcement Learning.” Nature 518 (7540):
529–33. https://doi.org/10.1038/nature14236.
Montúfar, Guido, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio. 2014.
“On the Number of Linear Regions of Deep Neural Networks.”
Advances in Neural Information Processing Systems 27.
Moon, Taesup, Alex Smola, Yi Chang, and Zhaohui Zheng. 2010.
“Intervalrank: Isotonic Regression with Listwise and Pairwise
Constraints.” Proceedings of the 3rd ACM
International Conference on Web
Search and Data Mining,
151–60. https://doi.org/10.1145/1718487.1718520.
Morozov, Vladimir Alekseevich. 1984. Methods for
Solving Incorrectly Posed
Problems. Springer.
Muennighoff, Niklas, Alexander M. Rush, Boaz Barak, et al. 2023.
“Scaling Data-Constrained Language Models.” Advances in
Neural Information Processing Systems.
Müller, Alfred. 1997. “Integral Probability Metrics and Their
Generating Classes of Functions.” Advances in Applied
Probability 29 (2): 429–43.
Murphy, Kevin P. 2022. Probabilistic Machine
Learning: An Introduction.
MIT Press. https://probml.github.io/pml-book/book1.html.
Nadaraya, Elizbar A. 1964. “On Estimating Regression.”
Theory of Probability & Its
Applications 9 (1): 141–42. https://doi.org/10.1137/1109020.
Nado, Zachary, Justin M. Gilmer, Christopher J. Shallue, Rohan Anil, and
George E. Dahl. 2021. “A Large Batch Optimizer Reality Check:
Traditional, Generic Optimizers Suffice Across Batch Sizes.”
arXiv Preprint arXiv:2102.06356.
Nair, Vinod, and Geoffrey E Hinton. 2010. “Rectified Linear Units
Improve Restricted Boltzmann Machines.”
ICML, 807–14. https://dl.acm.org/doi/10.5555/3104322.3104425.
Nakkiran, Preetum, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak,
and Ilya Sutskever. 2021. “Deep Double Descent: Where Bigger
Models and More Data Hurt.” Journal of
Statistical Mechanics: Theory and
Experiment 2021 (12): 124003. https://doi.org/10.1088/1742-5468/ac3a74.
Naor, Moni, and Omer Reingold. 1999. “On the Construction of
Pseudorandom Permutations: Luby–Rackoff
Revisited.” Journal of Cryptology 12 (1):
29–66. https://doi.org/10.1007/s001459900037.
Naumann, Uwe. 2008. “Optimal Jacobian Accumulation Is
NP-Complete.” Mathematical Programming 112
(2): 427–41.
Neal, Radford M. 1996. Bayesian Learning for
Neural Networks. Springer. https://doi.org/10.1007/978-1-4612-0745-0.
Nemirovski, Arkadi, and David Yudin. 1983. Problem Complexity and
Method Efficiency in Optimization. Wiley.
Nesterov, Yu. 2018. Lectures on Convex
Optimization. Springer. https://doi.org/10.1007/978-3-319-91578-4.
Nesterov, Yurii. 1983. “A Method of Solving a Convex Programming
Problem with Convergence Rate O(1/k2).”
Soviet Mathematics Doklady 27 (2): 372–76.
Neyman, Jerzy. 1937. “Outline of a Theory of Statistical
Estimation Based on the Classical Theory of Probability.”
Philosophical Transactions of the Royal
Society of London. Series
A, Mathematical and Physical
Sciences 236 (767): 333–80. https://doi.org/10.2307/jj.8501421.24.
Ng, Andrew Y., and Michael I. Jordan. 2002. “On Discriminative Vs.
Generative Classifiers: A Comparison of Logistic Regression and Naive
Bayes.” Advances in Neural
Information Processing
Systems 14: 841–48.
Nguyen, Minh Nhat, Andrew Baker, Clement Neo, Allen Roush, Andreas
Kirsch, and Ravid Shwartz-Ziv. 2025. “Turning up the Heat: Min-p
Sampling for Creative and Coherent LLM Outputs.”
International Conference on Learning Representations. https://arxiv.org/abs/2407.01082.
Nguyen, XuanLong, Martin J. Wainwright, and Michael I. Jordan. 2010.
“Estimating Divergence Functionals and the Likelihood Ratio by
Convex Risk Minimization.” IEEE Transactions on Information
Theory 56 (11): 5847–61.
Niculae, Vlad, and Mathieu Blondel. 2017. “A Regularized Framework
for Sparse and Structured Neural Attention.” Advances in
Neural Information Processing Systems 30.
Nocedal, Jorge, and Stephen J. Wright. 2006. Numerical
Optimization. 2nd ed. Springer.
Norelli, Antonio, Marco Fumero, Valentino Maiorca, Luca Moschella,
Emanuele Rodolà, and Francesco Locatello. 2022.
“ASIF: Coupled Data Turns Unimodal Models to
Multimodal Without Training.”
ArXiv:2210.01738. https://arxiv.org/abs/2210.01738.
Novak, Roman, Lechao Xiao, Jaehoon Lee, et al. 2018. “Bayesian
Deep Convolutional Networks with Many Channels Are Gaussian
Processes.” ArXiv:1810.05148. https://arxiv.org/abs/1810.05148.
Novikoff, A. B. J. 1962. “On Convergence Proofs for
Perceptrons.” Proceedings of the Symposium on
the Mathematical Theory of
Automata, 615–22. https://cs.nyu.edu/~mohri/pub/nov62.pdf.
Nowozin, Sebastian, Botond Cseke, and Ryota Tomioka. 2016.
“F-GAN: Training Generative Neural Samplers Using
Variational Divergence Minimization.” Advances in Neural
Information Processing Systems 29.
NVIDIA. 2025. “Nemotron-H: A Family of Accurate and
Efficient Hybrid Mamba-Transformer Models.”
arXiv Preprint arXiv:2504.03624.
Øksendal, Bernt. 2003. Stochastic Differential Equations: An
Introduction with Applications. 6th ed. Springer.
Olshausen, Bruno A, and David J Field. 1996. “Emergence of
Simple-Cell Receptive Field Properties by Learning a Sparse Code for
Natural Images.” Nature 381 (6583): 607–9. https://doi.org/10.1038/381607a0.
Olsson, Catherine, Nelson Elhage, Neel Nanda, et
al. 2022. “In-Context Learning and Induction Heads.”
Transformer Circuits Thread.
Ong, Cheng Soon, Alexander Smola, and Robert Williamson. 2005.
“Learning the Kernel with Hyperkernels.”
Journal of Machine Learning
Research 6: 1043–71. https://doi.org/10.1109/jcss.2005.25.2.
Oord, Aaron van den, Yazhe Li, and Oriol Vinyals. 2018.
“Representation Learning with Contrastive Predictive
Coding.” arXiv Preprint arXiv:1807.03748.
OpenAI. 2025. “Gpt-Oss-120b & Gpt-Oss-20b Model Card.”
arXiv Preprint arXiv:2508.10925.
Orvieto, Antonio, Samuel L. Smith, Albert Gu, et al. 2023.
“Resurrecting Recurrent Neural Networks for Long
Sequences.” International Conference on Machine
Learning, 26670–98. https://arxiv.org/abs/2303.06349.
Ouyang, Long, Jeff Wu, Xu Jiang, et al.
2022. “Training Language Models to Follow Instructions with Human
Feedback.” ArXiv:2203.02155. https://arxiv.org/abs/2203.02155.
Paley, Raymond E. A. C., Norbert Wiener, and Antoni Zygmund. 1933.
“Notes on Random Functions.” Mathematische
Zeitschrift 37: 647–68.
Papineni, Kishore, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002.
“BLEU: A Method for Automatic Evaluation of Machine
Translation.” Proceedings of the 40th Annual
Meeting of the Association for
Computational Linguistics, 311–18. https://doi.org/10.3115/1073083.1073135.
Parikh, Ankur P, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit.
2016. “A Decomposable Attention Model for Natural Language
Inference.” ArXiv:1606.01933. https://arxiv.org/abs/1606.01933.
Park, Taesung, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019.
“Semantic Image Synthesis with Spatially-Adaptive
Normalization.” Proceedings of the IEEE
Conference on Computer Vision and
Pattern Recognition, 2337–46. https://doi.org/10.1109/cvpr.2019.00244.
Parr, Terence, and Jeremy Howard. 2018. “The Matrix Calculus You
Need for Deep Learning.” arXiv Preprint
arXiv:1802.01528.
Parzen, Emanuel. 1957. “On Consistent Estimates of the Spectrum of
a Stationary Time Series.” Annals of
Mathematical Statistics 28: 329–48. https://doi.org/10.1214/aoms/1177706962.
Pascanu, Razvan, Tomas Mikolov, and Yoshua Bengio. 2013. “On the
Difficulty of Training Recurrent Neural Networks.”
International Conference on Machine Learning, 1310–18.
Paszke, Adam, Sam Gross, Francisco Massa, et
al. 2019. “PyTorch: An Imperative Style,
High-Performance Deep Learning Library.” Advances in
Neural Information Processing
Systems 32: 8026–37. https://proceedings.neurips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html.
Patarasuk, Pitch, and Xin Yuan. 2009. “Bandwidth Optimal
All-Reduce Algorithms for Clusters of Workstations.” Journal
of Parallel and Distributed
Computing 69 (2): 117–24. https://doi.org/10.1016/j.jpdc.2008.09.002.
Pearlmutter, Barak A. 1994. “Fast Exact Multiplication by the
Hessian.” Neural Computation 6 (1): 147–60.
https://doi.org/10.1162/neco.1994.6.1.147.
Peebles, William, and Saining Xie. 2023. “Scalable Diffusion
Models with Transformers.” Proceedings of the
IEEE/CVF International
Conference on Computer
Vision. https://arxiv.org/abs/2212.09748.
Peng, Bo, Daniel Goldstein, Quentin Anthony, et
al. 2024. “Eagle and Finch: RWKV
with Matrix-Valued States and Dynamic Recurrence.” First
Conference on Language Modeling. https://arxiv.org/abs/2404.05892.
Peng, Bowen, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024.
“YaRN: Efficient Context Window Extension of Large
Language Models.” International Conference on Learning
Representations. https://arxiv.org/abs/2309.00071.
Peng, Bo, Ruichong Zhang, Daniel Goldstein, et al. 2025.
“RWKV-7 "Goose" with Expressive Dynamic
State Evolution.” arXiv Preprint arXiv:2503.14456.
Pennington, Jeffrey, Samuel Schoenholz, and Surya Ganguli. 2017.
“Resurrecting the Sigmoid in Deep Learning Through Dynamical
Isometry: Theory and Practice.” Advances in
Neural Information Processing
Systems, 4785–95. https://proceedings.neurips.cc/paper/2017/hash/a9fc2d3b0c721b5b3f0b3f11bb24a75c-Abstract.html.
Pennington, Jeffrey, Richard Socher, and Christopher Manning. 2014.
“GloVe: Global Vectors for Word
Representation.” Proceedings of the 2014
Conference on Empirical Methods
in Natural Language Processing
(EMNLP), 1532–43. https://doi.org/10.3115/v1/d14-1162.
Peters, Jonas, Dominik Janzing, and Bernhard Schölkopf. 2017.
Elements of Causal Inference:
Foundations and Learning
Algorithms. MIT Press. https://doi.org/10.7551/mitpress/11283.001.0001.
Peters, Matthew, Waleed Ammar, Chandra Bhagavatula, and Russell Power.
2017. “Semi-Supervised Sequence Tagging with Bidirectional
Language Models.” Proceedings of the 55th Annual
Meeting of the Association for
Computational Linguistics, Volume
1, 1756–65. https://doi.org/10.18653/v1/p17-1161.
Peters, Matthew, Mark Neumann, Mohit Iyyer, et al. 2018. “Deep
Contextualized Word Representations.” Proceedings of the 2018
Conference of the North American
Chapter of the Association for
Computational Linguistics: Human
Language Technologies, Volume
1, 2227–37. https://doi.org/10.18653/v1/n18-1202.
Petersen, Kaare Brandt, and Michael Syskind Pedersen. 2008. The
Matrix Cookbook. Technical University of Denmark. https://www.math.uwaterloo.ca/~hwolkowi/matrixcookbook.pdf.
Peyré, Gabriel, and Marco Cuturi. 2019. “Computational Optimal
Transport.” Foundations and Trends in Machine Learning
11 (5–6): 355–607.
Pinsker, Mark S. 1964. Information and Information Stability of
Random Variables and Processes. Holden-Day.
Planck, Max. 1917. “Über Einen Satz Der
Statistischen Dynamik Und Seine Erweiterung in
Der Quantentheorie.” Sitzungsberichte Der
Preussischen Akademie Der Wissenschaften, 324–41.
Pleiss, Geoff, Danlu Chen, Gao Huang, Tongcheng Li, Laurens Van Der
Maaten, and Kilian Q Weinberger. 2017. “Memory-Efficient
Implementation of Densenets.”
ArXiv:1707.06990. https://arxiv.org/abs/1707.06990.
Polyak, Boris T. 1964. “Some Methods of Speeding up the
Convergence of Iteration Methods.” USSR
Computational Mathematics and
Mathematical Physics 4 (5): 1–17. https://doi.org/10.1016/0041-5553(64)90137-5.
Polyak, Boris T. 1963. “Gradient Methods for the Minimisation of
Functionals.” USSR Computational Mathematics and Mathematical
Physics 3 (4): 864–78.
Polyak, Boris T., and Anatoli B. Juditsky. 1992. “Acceleration of
Stochastic Approximation by Averaging.” SIAM Journal on
Control and Optimization 30 (4): 838–55.
Pontryagin, Lev S., Vladimir G. Boltyanskii, Revaz V. Gamkrelidze, and
Evgenii F. Mishchenko. 1962. The Mathematical Theory of Optimal
Processes. Interscience.
Pooladian, Aram-Alexandre, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon
Amos, Yaron Lipman, and Ricky T. Q. Chen. 2023. “Multisample Flow
Matching: Straightening Flows with Minibatch Couplings.”
Proceedings of the 40th International Conference on Machine
Learning.
Poole, Ben, Sherjil Ozair, Aaron van den Oord, Alexander A. Alemi, and
George Tucker. 2019. “On Variational Bounds of Mutual
Information.” Proceedings of the 36th International
Conference on Machine Learning, 5171–80.
Pope, Reiner, Sholto Douglas, Aakanksha Chowdhery, et al. 2023.
“Efficiently Scaling Transformer Inference.”
Proceedings of Machine Learning and Systems 5.
Popović, Maja. 2015. “ChrF: Character n-Gram
F-Score for Automatic MT Evaluation.”
Proceedings of the Tenth Workshop on Statistical Machine
Translation, 392–95. https://doi.org/10.18653/v1/W15-3049.
Popper, Karl. 2005. The Logic of
Scientific Discovery. Routledge. https://doi.org/10.4324/9780203994627.
Power, Alethea, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant
Misra. 2022. “Grokking: Generalization Beyond Overfitting on Small
Algorithmic Datasets.” ArXiv:2201.02177. https://arxiv.org/abs/2201.02177.
Prakash, Aaditya, Sadid A Hasan, Kathy Lee, et al. 2016. “Neural
Paraphrase Generation with Stacked Residual LSTM
Networks.” ArXiv:1610.03098. https://arxiv.org/abs/1610.03098.
Press, Ofir, Noah A. Smith, and Mike Lewis. 2022. “Train Short,
Test Long: Attention with Linear Biases Enables Input Length
Extrapolation.” International Conference on Learning
Representations. https://arxiv.org/abs/2108.12409.
Press, Ofir, and Lior Wolf. 2017. “Using the Output Embedding to
Improve Language Models.” Proceedings of the 15th Conference
of the European Chapter of the Association for Computational
Linguistics, 157–63.
Qin, Danfeng, Chas Leichner, Manolis Delakis, et al. 2024.
“MobileNetV4: Universal Models for the Mobile
Ecosystem.” European Conference on
Computer Vision. https://arxiv.org/abs/2404.10518.
Quadrana, Massimo, Paolo Cremonesi, and Dietmar Jannach. 2018.
“Sequence-Aware Recommender Systems.” ACM
Computing Surveys 51 (4): 66. https://doi.org/10.1145/3190616.
Quinlan, J Ross. 1993. C4.5: Programs for
Machine Learning. Elsevier. https://doi.org/10.1016/c2009-0-27846-9.
Qwen Team. 2025. Qwen3-Next: Towards Ultimate Training
and Inference Efficiency. Model release, https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct.
Rabe, Markus N., and Charles Staats. 2021. “Self-Attention Does
Not Need O(n2)
Memory.” arXiv Preprint arXiv:2112.05682.
Rabiner, Lawrence, and Biing-Hwang Juang. 1993. Fundamentals of
Speech Recognition.
Prentice-Hall.
Rademacher, Hans. 1919. “Über Partielle Und Totale
Differenzierbarkeit von Funktionen Mehrerer
Variabeln Und über Die
Transformation Der Doppelintegrale.”
Mathematische Annalen 79 (4): 340–59.
Radford, Alec, Jong Wook Kim, Chris Hallacy, et
al. 2021. “Learning Transferable Visual Models from Natural
Language Supervision.” International
Conference on Machine Learning, 8748–63. https://proceedings.mlr.press/v139/radford21a.html.
Radford, Alec, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey,
and Ilya Sutskever. 2023. “Robust Speech Recognition via
Large-Scale Weak Supervision.” International
Conference on Machine
Learning. https://arxiv.org/abs/2212.04356.
Radford, Alec, Luke Metz, and Soumith Chintala. 2015.
“Unsupervised Representation Learning with Deep Convolutional
Generative Adversarial Networks.”
ArXiv:1511.06434. https://arxiv.org/abs/1511.06434.
Radford, Alec, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever.
2018. “Improving Language Understanding by Generative
Pre-Training.” OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf.
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and
Ilya Sutskever. 2019. “Language Models Are Unsupervised Multitask
Learners.” OpenAI Blog 1 (8):
9. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
Radhakrishna Rao, C. 1945. “Information and Accuracy Attainable in
the Estimation of Statistical Parameters.” Bulletin of the
Calcutta Mathematical
Society 37 (3): 81–91. https://doi.org/10.1007/BF02908227.
Radon, Johann. 1921. “Mengen Konvexer
Körper, Die Einen Gemeinsamen
Punkt Enthalten.” Mathematische Annalen 83
(1–2): 113–15.
Radosavovic, Ilija, Justin Johnson, Saining Xie, Wan-Yen Lo, and Piotr
Dollár. 2019. “On Network Design Spaces for Visual
Recognition.” Proceedings of the
IEEE/CVF International
Conference on Computer
Vision, 1882–90. https://doi.org/10.1109/iccv.2019.00052.
Radosavovic, Ilija, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and
Piotr Dollár. 2020. “Designing Network Design Spaces.”
Proceedings of the IEEE/CVF
Conference on Computer Vision and
Pattern Recognition, 10428–36. https://doi.org/10.1109/cvpr42600.2020.01044.
Rae, Jack W, Sebastian Borgeaud, Trevor Cai, et
al. 2021. “Scaling Language Models: Methods, Analysis &
Insights from Training Gopher.”
ArXiv:2112.11446. https://arxiv.org/abs/2112.11446.
Raffel, Colin, Noam Shazeer, Adam Roberts, et al. 2020. “Exploring
the Limits of Transfer Learning with a Unified Text-to-Text
Transformer.” Journal of Machine
Learning Research 21: 1–67. https://arxiv.org/abs/1910.10683.
Rahimi, Ali, and Benjamin Recht. 2007. “Random Features for
Large-Scale Kernel Machines.” Advances in Neural Information
Processing Systems 20.
Rajbhandari, Samyam, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020.
“ZeRO: Memory Optimizations Toward Training Trillion
Parameter Models.” SC20: International Conference for High
Performance Computing, Networking, Storage and Analysis, 1–16.
Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang.
2016. “SQuAD:
100,000+ Questions for Machine Comprehension of Text.”
ArXiv:1606.05250. https://arxiv.org/abs/1606.05250.
Ramachandran, Prajit, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm
Levskaya, and Jon Shlens. 2019. “Stand-Alone Self-Attention in
Vision Models.” Advances in Neural
Information Processing
Systems 32. https://arxiv.org/abs/1906.05909.
Ramachandran, Prajit, Barret Zoph, and Quoc V Le. 2017. “Searching
for Activation Functions.”
ArXiv:1710.05941. https://arxiv.org/abs/1710.05941.
Ramesh, Aditya, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark
Chen. 2022. “Hierarchical Text-Conditional Image Generation with
Clip Latents.” ArXiv:2204.06125. https://arxiv.org/abs/2204.06125.
Ramón y Cajal, Santiago, and L. Azoulay. 1894. Les
Nouvelles Idées Sur La
Structure Du Système
Nerveux Chez l’Homme Et Chez Les
Vertébrés. Paris, C.
Reinwald & Cie.
Ranzato, Marc-Aurelio, Y-Lan Boureau, Sumit Chopra, and Yann LeCun.
2007. “A Unified Energy-Based Framework for Unsupervised
Learning.” Artificial Intelligence and
Statistics, 371–79. https://proceedings.mlr.press/v2/ranzato07a.html.
Rasmussen, Carl Edward, and Christopher KI Williams. 2006. Gaussian
Processes for Machine
Learning. MIT Press. https://gaussianprocess.org/gpml/.
Recht, Benjamin, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar.
2019. “Do ImageNet Classifiers Generalize to
ImageNet?” International Conference on Machine
Learning, 5389–400. https://arxiv.org/abs/1902.10811.
Reddi, Sashank J, Satyen Kale, and Sanjiv Kumar. 2018. “On the
Convergence of Adam and Beyond.” International
Conference on Learning Representations. https://arxiv.org/abs/1904.09237.
Redmon, Joseph, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016.
“You Only Look Once: Unified, Real-Time Object Detection.”
Proceedings of the IEEE Conference on
Computer Vision and Pattern
Recognition, 779–88. https://doi.org/10.1109/cvpr.2016.91.
Redmon, Joseph, and Ali Farhadi. 2018. “YOLOv3: An
Incremental Improvement.” ArXiv:1804.02767.
https://arxiv.org/abs/1804.02767.
Reed, Scott, and Nando De Freitas. 2015. “Neural
Programmer-Interpreters.” ArXiv:1511.06279.
https://arxiv.org/abs/1511.06279.
Reed, Scott, Konrad Zolna, Emilio Parisotto, et
al. 2022. “A Generalist Agent.”
ArXiv:2205.06175. https://arxiv.org/abs/2205.06175.
Ren, Liliang, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu
Chen. 2024. “Samba: Simple Hybrid State Space Models for Efficient
Unlimited Context Language Modeling.” arXiv Preprint
arXiv:2406.07522.
Ren, Shaoqing, Kaiming He, Ross Girshick, and Jian Sun. 2015.
“Faster R-CNN: Towards Real-Time Object
Detection with Region Proposal Networks.” Advances in
Neural Information Processing
Systems, 91–99. https://doi.org/10.5555/2969239.2969250.
Rendle, Steffen. 2010. “Factorization Machines.” 2010
IEEE International Conference on
Data Mining, 995–1000. https://doi.org/10.1109/icdm.2010.127.
Rendle, Steffen, Christoph Freudenthaler, Zeno Gantner, and Lars
Schmidt-Thieme. 2009. “BPR: Bayesian
Personalized Ranking from Implicit Feedback.” Proceedings of
the 25th Conference on Uncertainty in
Artificial Intelligence, 452–61. https://arxiv.org/abs/1205.2618.
Rényi, Alfréd. 1961. “On Measures of Entropy and
Information.” Proceedings of the Fourth Berkeley Symposium on
Mathematical Statistics and Probability, Volume 1: Contributions to the
Theory of Statistics, 547–61.
Revels, Jarrett, Miles Lubin, and Theodore Papamarkou. 2016.
“Forward-Mode Automatic Differentiation in
Julia.” ArXiv:1607.07892. https://arxiv.org/abs/1607.07892.
Rezende, Danilo Jimenez, Shakir Mohamed, and Daan Wierstra. 2014.
“Stochastic Backpropagation and Approximate Inference in Deep
Generative Models.” International
Conference on Machine
Learning, 1278–86. https://proceedings.mlr.press/v32/rezende14.html.
Rezende, Danilo, and Shakir Mohamed. 2015. “Variational Inference
with Normalizing Flows.” International
Conference on Machine
Learning, 1530–38. https://proceedings.mlr.press/v37/rezende15.html.
Riesenhuber, Maximilian, and Tomaso Poggio. 1999. “Hierarchical
Models of Object Recognition in Cortex.” Nature
Neuroscience 2 (11): 1019–25. https://doi.org/10.1038/14819.
Risken, Hannes. 1996. The Fokker–Planck
Equation: Methods of Solution and Applications. 2nd ed. Springer.
Rissanen, Jorma J. 1976. “Generalized Kraft
Inequality and Arithmetic Coding.” IBM Journal of Research
and Development 20 (3): 198–203.
Robbins, Herbert, and Sutton Monro. 1951. “A Stochastic
Approximation Method.” The Annals of Mathematical
Statistics 22 (3): 400–407.
Rockafellar, R. T. 1970. Convex Analysis.
Princeton University Press. https://doi.org/10.1515/9781400873173.
Rolnick, David, Andreas Veit, Serge Belongie, and Nir Shavit. 2017.
“Deep Learning Is Robust to Massive Label Noise.”
ArXiv:1705.10694. https://arxiv.org/abs/1705.10694.
Rombach, Robin, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and
Björn Ommer. 2022. “High-Resolution Image Synthesis with Latent
Diffusion Models.” IEEE/CVF Conference on Computer Vision and
Pattern Recognition, 10684–95.
Rudin, W. 1973. Functional Analysis. McGraw-Hill.
Rudin, Walter. 1976. Principles of Mathematical Analysis. 3rd
ed. McGraw-Hill.
Rumelhart, David E, Geoffrey E Hinton, and Ronald J Williams. 1986.
“Learning Representations by Back-Propagating Errors.”
Nature 323 (6088): 533–36. https://doi.org/10.1038/323533a0.
Russakovsky, Olga, Jia Deng, Zhiheng Huang, Alexander C. Berg, and Li
Fei-Fei. 2013. “Detecting Avocados to Zucchinis: What Have We
Done, and Where Are We Going?” International
Conference on Computer Vision
(ICCV). https://doi.org/10.1109/iccv.2013.258.
Russakovsky, Olga, Jia Deng, Hao Su, et al.
2015. “ImageNet Large Scale Visual Recognition
Challenge.” International Journal of
Computer Vision 115 (3): 211–52. https://doi.org/10.1007/s11263-015-0816-y.
Russell, Stuart J, and Peter Norvig. 2016. Artificial
Intelligence: A Modern
Approach. Pearson Education
Limited.
Saerens, Marco, Patrice Latinne, and Christine Decaestecker. 2002.
“Adjusting the Outputs of a Classifier to New a Priori
Probabilities: A Simple Procedure.” Neural Computation
14 (1): 21–41. https://doi.org/10.1162/089976602753284446.
Sahami, Mehran, Susan Dumais, David Heckerman, and Eric Horvitz. 1998.
“A Bayesian Approach to Filtering Junk
e-Mail.” AAAI Workshop on Learning for Text
Categorization.
Saharia, Chitwan, William Chan, Saurabh Saxena, et
al. 2022. “Photorealistic Text-to-Image Diffusion Models
with Deep Language Understanding.”
ArXiv:2205.11487. https://arxiv.org/abs/2205.11487.
Salimans, Tim, and Jonathan Ho. 2022. “Progressive Distillation
for Fast Sampling of Diffusion Models.” International
Conference on Learning Representations.
Salinas, D., M. Seeger, A. Klein, V. Perrone, M. Wistuba, and C.
Archambeau. 2022. “Syne Tune: A Library for Large
Scale Hyperparameter Tuning and Reproducible Research.” First
Conference on Automated Machine
Learning. https://proceedings.mlr.press/v188/salinas22a.html.
Sandler, Mark, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and
Liang-Chieh Chen. 2018. “MobileNetV2: Inverted
Residuals and Linear Bottlenecks.” Proceedings of the
IEEE Conference on Computer
Vision and Pattern
Recognition. https://arxiv.org/abs/1801.04381.
Santurkar, Shibani, Dimitris Tsipras, Andrew Ilyas, and Aleksander
Madry. 2018. “How Does Batch Normalization Help
Optimization?” Advances in Neural
Information Processing
Systems, 2483–93. https://doi.org/10.5555/3327345.3327508.
Särkkä, Simo, and Arno Solin. 2019. Applied Stochastic Differential
Equations. Cambridge University Press.
Sarwar, Badrul Munir, George Karypis, Joseph A Konstan, and John Riedl.
2001. “Item-Based Collaborative Filtering Recommendation
Algorithms.” Proceedings of 10th International
Conference on World Wide
Web, 285–95. https://doi.org/10.1145/371920.372071.
Saxe, Andrew M., Yamini Bansal, Joel Dapello, et al. 2018. “On the
Information Bottleneck Theory of Deep Learning.”
International Conference on Learning Representations.
Saxe, Andrew M., James L. McClelland, and Surya Ganguli. 2014.
“Exact Solutions to the Nonlinear Dynamics of Learning in Deep
Linear Neural Networks.” International Conference on Learning
Representations. https://arxiv.org/abs/1312.6120.
Schein, Andrew I, Alexandrin Popescul, Lyle H Ungar, and David M
Pennock. 2002. “Methods and Metrics for Cold-Start
Recommendations.” Proceedings of the 25th Annual
International ACM SIGIR
Conference on Research and
Development in Information
Retrieval, 253–60. https://doi.org/10.1145/564376.564421.
Schlag, Imanol, Kazuki Irie, and Jürgen Schmidhuber. 2021. “Linear
Transformers Are Secretly Fast Weight Programmers.”
International Conference on Machine Learning.
Schmidt, Robin M., Frank Schneider, and Philipp Hennig. 2021.
“Descending Through a Crowded Valley — Benchmarking Deep Learning
Optimizers.” Proceedings of the 38th International Conference
on Machine Learning, 9367–76. https://arxiv.org/abs/2007.01547.
Schölkopf, Bernhard, and Alexander J Smola. 2002. Learning with
Kernels: Support Vector
Machines, Regularization,
Optimization, and Beyond.
MIT Press. https://doi.org/10.7551/mitpress/4175.001.0001.
Schölkopf, B., R. Herbrich, and A. J. Smola. 2001. “A Generalized
Representer Theorem.” In Proceedings of the
Annual Conference on
Computational Learning
Theory, edited by D. P. Helmbold and B. Williamson.
Springer-Verlag. https://doi.org/10.1007/3-540-44581-1_27.
Schuster, Mike, and Kuldip K Paliwal. 1997. “Bidirectional
Recurrent Neural Networks.” IEEE Transactions on
Signal Processing 45 (11): 2673–81. https://doi.org/10.1109/78.650093.
Sedhain, Suvash, Aditya Krishna Menon, Scott Sanner, and Lexing Xie.
2015. “AutoRec: Autoencoders Meet Collaborative
Filtering.” Proceedings of the 24th
International Conference on World
Wide Web, 111–12. https://doi.org/10.1145/2740908.2742726.
Sennrich, Rico, Barry Haddow, and Alexandra Birch. 2016. “Neural
Machine Translation of Rare Words with Subword Units.”
Proceedings of the 54th Annual Meeting of the Association for
Computational Linguistics, 1715–25. https://doi.org/10.18653/v1/P16-1162.
Shah, Ishaan, Anthony M. Polloreno, Karl Stratos,
et al. 2025. “Practical Efficiency of Muon for
Pretraining.” arXiv Preprint arXiv:2505.02222.
Shallue, Christopher J., Jaehoon Lee, Joseph Antognini, Jascha
Sohl-Dickstein, Roy Frostig, and George E. Dahl. 2019. “Measuring
the Effects of Data Parallelism on Neural Network Training.”
Journal of Machine Learning Research 20 (112): 1–49. https://arxiv.org/abs/1811.03600.
Shannon, Claude E. 1959. “Coding Theorems for a Discrete Source
with a Fidelity Criterion.” IRE National Convention
Record 7 (4): 142–63.
Shannon, Claude Elwood. 1948. “A Mathematical Theory of
Communication.” The Bell System
Technical Journal 27 (3): 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x.
Shannon, Claude Elwood. 1951. “Prediction and Entropy of Printed
English.” The Bell
System Technical Journal 30
(1): 50–64. https://doi.org/10.1002/j.1538-7305.1951.tb01366.x.
Shaw, Peter, Jakob Uszkoreit, and Ashish Vaswani. 2018.
“Self-Attention with Relative Position Representations.”
ArXiv:1803.02155. https://arxiv.org/abs/1803.02155.
Shazeer, Noam. 2019. “Fast Transformer Decoding: One Write-Head Is
All You Need.” arXiv Preprint arXiv:1911.02150.
Shazeer, Noam. 2020. “GLU Variants Improve
Transformer.” ArXiv:2002.05202. https://arxiv.org/abs/2002.05202.
Shazeer, Noam, Azalia Mirhoseini, Krzysztof Maziarz, et al. 2017.
“Outrageously Large Neural Networks: The Sparsely-Gated
Mixture-of-Experts Layer.” International Conference on
Learning Representations.
Shazeer, Noam, and Mitchell Stern. 2018. “Adafactor: Adaptive
Learning Rates with Sublinear Memory Cost.” International
Conference on Machine Learning. https://arxiv.org/abs/1804.04235.
Shimodaira, Hidetoshi. 2000. “Improving Predictive Inference Under
Covariate Shift by Weighting the Log-Likelihood Function.”
Journal of Statistical Planning and
Inference 90 (2): 227–44.
Shwartz-Ziv, Ravid, and Amitai Armon. 2022. “Tabular Data: Deep
Learning Is Not All You Need.” Information Fusion 81:
84–90. https://doi.org/10.1016/j.inffus.2021.11.011.
Shwartz-Ziv, Ravid, and Naftali Tishby. 2017. “Opening the Black
Box of Deep Neural Networks via Information.” arXiv Preprint
arXiv:1703.00810.
Siems, Julien, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano
Pontil, and Riccardo Grazzi. 2025. “DeltaProduct:
Improving State-Tracking in Linear RNNs via
Householder Products.” arXiv Preprint
arXiv:2502.10297.
Silver, David, Aja Huang, Chris J Maddison, et
al. 2016. “Mastering the Game of Go with Deep
Neural Networks and Tree Search.” Nature 529 (7587):
484–89. https://doi.org/10.1038/nature16961.
Silverman, B. W. 1986. Density Estimation for
Statistical and Data
Analysis. Chapman; Hall.
Simonyan, Karen, and Andrew Zisserman. 2015. “Very Deep
Convolutional Networks for Large-Scale Image Recognition.”
International Conference on Learning Representations. https://arxiv.org/abs/1409.1556.
Sindhwani, Vikas, Tara N Sainath, and Sanjiv Kumar. 2015.
“Structured Transforms for Small-Footprint Deep Learning.”
ArXiv:1510.01722. https://arxiv.org/abs/1510.01722.
Sivic, Josef, and Andrew Zisserman. 2003. “Video
Google: A Text Retrieval Approach to Object Matching in
Videos.” Proceedings of the IEEE
International Conference on
Computer Vision 3: 1470–77. https://doi.org/10.1109/iccv.2003.1238663.
Slater, Morton. 1950. Lagrange Multipliers Revisited.
Discussion Paper Mathematics 403. Cowles Commission for Research in
Economics.
Smith, Jimmy T. H., Andrew Warrington, and Scott W. Linderman. 2023.
“Simplified State Space Layers for Sequence Modeling.”
International Conference on Learning Representations. https://arxiv.org/abs/2208.04933.
Smith, Samuel L., Andrew Brock, Leonard Berrada, and Soham De. 2023.
“ConvNets Match Vision Transformers at Scale.”
ArXiv:2310.16764. https://arxiv.org/abs/2310.16764.
Smith, Samuel L., Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le.
2018. “Don’t Decay the Learning Rate, Increase the Batch
Size.” International Conference on Learning
Representations. https://arxiv.org/abs/1711.00489.
Snoek, J., H. Larochelle, and R. Adams. 2012. “Practical
Bayesian Optimization of Machine Learning
Algorithms.” Advances in Neural
Information Processing Systems
25, 2951–59. https://doi.org/10.5555/2999325.2999464.
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya
Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium
Thermodynamics.” International
Conference on Machine
Learning, 2256–65. https://proceedings.mlr.press/v37/sohl-dickstein15.html.
Song, Jiaming, Chenlin Meng, and Stefano Ermon. 2021. “Denoising
Diffusion Implicit Models.” International Conference on
Learning Representations.
Song, Yang, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. 2023.
“Consistency Models.” Proceedings of the 40th
International Conference on Machine Learning.
Song, Yang, Conor Durkan, Iain Murray, and Stefano Ermon. 2021.
“Maximum Likelihood Training of Score-Based Diffusion
Models.” Advances in Neural Information Processing
Systems 34.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by
Estimating Gradients of the Data Distribution.” Advances in
Neural Information Processing
Systems 32. https://arxiv.org/abs/1907.05600.
Song, Yang, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar,
Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative
Modeling Through Stochastic Differential Equations.”
International Conference on
Learning Representations. https://doi.org/10.52202/075280-1645.
Soudry, Daniel, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and
Nathan Srebro. 2018. “The Implicit Bias of Gradient Descent on
Separable Data.” Journal of Machine Learning Research 19
(70): 1–57.
Speelpenning, Bert. 1980. “Compiling Fast Partial Derivatives of
Functions Given by Algorithms.” PhD thesis, University of
Illinois at Urbana-Champaign. https://doi.org/10.2172/5254402.
Srivastava, Nitish, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever,
and Ruslan Salakhutdinov. 2014. “Dropout: A Simple Way to Prevent
Neural Networks from Overfitting.” Journal of
Machine Learning Research 15
(1): 1929–58. https://doi.org/10.5555/2627435.2670313.
Srivastava, Rupesh Kumar, Klaus Greff, and Jürgen Schmidhuber. 2015.
“Highway Networks.” ArXiv:1505.00387.
https://arxiv.org/abs/1505.00387.
Stein, Charles M. 1981. “Estimation of the Mean of a Multivariate
Normal Distribution.” Annals of Statistics 9 (6):
1135–51.
Sterbenz, Pat H. 1974. Floating-Point Computation.
Prentice-Hall.
Strang, Gilbert. 1993. Introduction to Linear
Algebra. Wellesley–Cambridge
Press. https://math.mit.edu/~gs/linearalgebra/.
Stratonovich, Ruslan L. 1966. “A New Representation for Stochastic
Integrals and Equations.” SIAM Journal on Control 4 (2):
362–71.
Student (Gosset, William S.). 1908. “The Probable Error of a
Mean.” Biometrika 6 (1): 1–25.
Su, Jianlin, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng
Liu. 2024. “RoFormer: Enhanced Transformer with
Rotary Position Embedding.” Neurocomputing 568: 127063.
https://arxiv.org/abs/2104.09864.
Su, Xiaoyuan, and Taghi M Khoshgoftaar. 2009. “A Survey of
Collaborative Filtering Techniques.” Advances in
Artificial Intelligence 2009. https://doi.org/10.1155/2009/421425.
Sukhbaatar, Sainbayar, Jason Weston, and Rob Fergus. 2015.
“End-to-End Memory Networks.” Advances in
Neural Information Processing
Systems, 2440–48. https://doi.org/10.5555/2969239.2969426.
Sun, Yu, Xinhao Li, Karan Dalal, et al. 2024. “Learning to (Learn
at Test Time): RNNs with Expressive Hidden States.”
arXiv Preprint arXiv:2407.04620.
Sun, Yutao, Li Dong, Shaohan Huang, et al. 2023. “Retentive
Network: A Successor to Transformer for Large Language Models.”
arXiv Preprint arXiv:2307.08621.
Sutskever, Ilya, James Martens, George Dahl, and Geoffrey Hinton. 2013.
“On the Importance of Initialization and Momentum in Deep
Learning.” International Conference
on Machine Learning, 1139–47. https://proceedings.mlr.press/v28/sutskever13.html.
Sutskever, Ilya, Oriol Vinyals, and Quoc V Le. 2014. “Sequence to
Sequence Learning with Neural Networks.” Advances in
Neural Information Processing
Systems, 3104–12. https://doi.org/10.5555/2969033.2969173.
Szegedy, Christian, Sergey Ioffe, Vincent Vanhoucke, and Alexander A
Alemi. 2017. “Inception-V4,
Inception-ResNet and the Impact
of Residual Connections on Learning.” 31st AAAI
Conference on Artificial
Intelligence. https://doi.org/10.1609/aaai.v31i1.11231.
Szegedy, Christian, Wei Liu, Yangqing Jia, et al. 2015. “Going
Deeper with Convolutions.” Proceedings of the
IEEE Conference on Computer
Vision and Pattern
Recognition, 1–9. https://doi.org/10.1109/cvpr.2015.7298594.
Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and
Zbigniew Wojna. 2016. “Rethinking the Inception
Architecture for Computer Vision.” Proceedings of the
IEEE Conference on Computer
Vision and Pattern
Recognition, 2818–26. https://doi.org/10.1109/cvpr.2016.308.
Székely, Gábor J., and Maria L. Rizzo. 2013. “Energy Statistics: A
Class of Statistics Based on Distances.” Journal of
Statistical Planning and Inference 143 (8): 1249–72.
Tallec, Corentin, and Yann Ollivier. 2017. “Unbiasing Truncated
Backpropagation Through Time.”
ArXiv:1705.08209. https://arxiv.org/abs/1705.08209.
Tan, Mingxing, and Quoc Le. 2019. “EfficientNet:
Rethinking Model Scaling for Convolutional Neural Networks.”
International Conference on
Machine Learning, 6105–14. https://proceedings.mlr.press/v97/tan19a.html.
Tan, Mingxing, and Quoc V. Le. 2021.
“EfficientNetV2: Smaller Models and
Faster Training.” International
Conference on Machine
Learning. https://arxiv.org/abs/2104.00298.
Tang, Jiaxi, and Ke Wang. 2018. “Personalized Top-n Sequential
Recommendation via Convolutional Sequence Embedding.”
Proceedings of the Eleventh ACM
International Conference on Web
Search and Data Mining,
565–73. https://doi.org/10.1145/3159652.3159656.
Tao, Terence, and Van Vu. 2010. “Random Matrices: Universality of
ESDs and the Circular Law.” Annals of
Probability 38 (5): 2023–65.
Taskar, Ben, Carlos Guestrin, and Daphne Koller. 2004. “Max-Margin
Markov Networks.” Advances in
Neural Information Processing
Systems 16: 25.
Tay, Yi, Mostafa Dehghani, Samira Abnar, et al. 2021. “Long Range
Arena: A Benchmark for Efficient Transformers.” International
Conference on Learning Representations. https://arxiv.org/abs/2011.04006.
Tay, Yi, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020.
“Efficient Transformers: A Survey.”
ArXiv:2009.06732. https://arxiv.org/abs/2009.06732.
Team OLMo. 2025. “Olmo 3.” arXiv Preprint
arXiv:2512.13961.
Team OLMo, Pete Walsh, Luca Soldaini, et al.
2025. “2 OLMo 2 Furious.” arXiv Preprint
arXiv:2501.00656.
Telgarsky, Matus. 2016. “Benefits of Depth in Neural
Networks.” Conference on Learning Theory, 1517–39.
Tencent Hunyuan Team. 2025. “Hunyuan-TurboS:
Advancing Large Language Models Through Mamba-Transformer
Synergy and Adaptive Chain-of-Thought.” arXiv Preprint
arXiv:2505.15431.
Teye, Mattias, Hossein Azizpour, and Kevin Smith. 2018. “Bayesian
Uncertainty Estimation for Batch Normalized Deep Networks.”
ArXiv:1802.06455. https://arxiv.org/abs/1802.06455.
Thomee, Bart, David A Shamma, Gerald Friedland, et al. 2016.
“YFCC100M: The New Data in Multimedia Research.”
Communications of the ACM 59 (2): 64–73. https://doi.org/10.1145/2812802.
Tibshirani, Robert. 1996. “Regression Shrinkage and Selection via
the Lasso.” Journal of the Royal Statistical Society: Series
B 58 (1): 267–88.
Tieleman, Tijmen, and Geoffrey Hinton. 2012. “Divide the Gradient
by a Running Average of Its Recent Magnitude.” In
COURSERA: Neural Networks for
Machine Learning, Lecture 6.5-Rmsprop. https://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf.
Tikhonov, A. N., and V. Y. Arsenin. 1977. Solutions of
Ill-Posed Problems.
W.H. Winston. https://doi.org/10.1137/1.9780898719741.
Tillet, Philippe, H. T. Kung, and David Cox. 2019.
“Triton: An Intermediate Language and Compiler for
Tiled Neural Network Computations.” 3rd ACM
SIGPLAN International Workshop on
Machine Learning and Programming
Languages (MAPL), 10–19. https://doi.org/10.1145/3315508.3329973.
Tishby, Naftali, Fernando C. Pereira, and William Bialek. 1999.
“The Information Bottleneck Method.” Proceedings of the
37th Annual Allerton Conference on Communication, Control, and
Computing, 368–77.
Tong, Alexander, Kilian Fatras, Nikolay Malkin, et al. 2024.
“Improving and Generalizing Flow-Based Generative Models with
Minibatch Optimal Transport.” Transactions on Machine
Learning Research.
Torralba, Antonio, Rob Fergus, and William T Freeman. 2008. “80
Million Tiny Images: A Large Data Set for Nonparametric Object and Scene
Recognition.” IEEE Transactions on
Pattern Analysis and Machine
Intelligence 30 (11): 1958–70. https://doi.org/10.1109/tpami.2008.128.
Töscher, Andreas, Michael Jahrer, and Robert M Bell. 2009. The
Bigchaos Solution to the Netflix Grand Prize. https://www.netflixprize.com/assets/GrandPrize2009_BPC_BigChaos.pdf.
Touvron, Hugo, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre
Sablayrolles, and Hervé Jégou. 2021. “Training Data-Efficient
Image Transformers & Distillation Through Attention.”
International Conference on
Machine Learning, 10347–57. https://proceedings.mlr.press/v139/touvron21a.html.
Touvron, Hugo, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve,
and Hervé Jégou. 2021. “Going Deeper with Image
Transformers.” IEEE/CVF
International Conference on
Computer Vision, 32–42. https://arxiv.org/abs/2103.17239.
Touvron, Hugo, Thibaut Lavril, Gautier Izacard, et
al. 2023a. “LLaMA: Open and
Efficient Foundation Language Models.”
ArXiv:2302.13971, 2023a. https://arxiv.org/abs/2302.13971.
Touvron, Hugo, Louis Martin, Kevin Stone, et
al. 2023b. “LLaMA 2: Open
Foundation and Fine-Tuned Chat Models.”
ArXiv:2307.09288, 2023b. https://arxiv.org/abs/2307.09288.
Tschannen, Michael, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly,
and Mario Lucic. 2020. “On Mutual Information Maximization for
Representation Learning.” International Conference on
Learning Representations.
Tsoumakas, Grigorios, and Ioannis Katakis. 2007. “Multi-Label
Classification: An Overview.” International
Journal of Data Warehousing and
Mining 3 (3): 1–13. https://doi.org/10.4018/jdwm.2007070101.
Turing, Alan. 1950. “Computing Machinery and Intelligence.”
Mind 59 (236): 433–60. https://doi.org/10.1093/mind/lix.236.433.
Uhlenbeck, George E., and Leonard S. Ornstein. 1930. “On the
Theory of the Brownian Motion.” Physical
Review 36 (5): 823–41.
Uijlings, Jasper RR, Koen EA Van De Sande, Theo Gevers, and Arnold WM
Smeulders. 2013. “Selective Search for Object Recognition.”
International Journal of Computer
Vision 104 (2): 154–71. https://doi.org/10.1007/s11263-013-0620-5.
Vapnik, V. 1995. The Nature of Statistical
Learning Theory. Springer.
Vapnik, V. 1998. Statistical Learning Theory. John Wiley; Sons.
Vapnik, V. N., and A. Y. Chervonenkis. 1974. “Ordered Risk
Minimization.” Automation and Remote Control 35:
1226–35, 1403–12.
Vapnik, V., and A. Chervonenkis. 1964. “A Note on One Class of
Perceptrons.” Automation and Remote Control 25.
Vapnik, V., and A. Chervonenkis. 1968. “Uniform Convergence of
Frequencies of Occurence of Events to Their Probabilities.”
Dokl. Akad. Nauk SSSR 181: 915–18.
Vapnik, V., and A. Chervonenkis. 1971b. “On the Uniform
Convergence of Relative Frequencies of Events to Their
Probabilities.” Theory Probab.
Appl. 16 (2): 264–81. https://doi.org/10.1137/1116025.
Vapnik, V., and A. Chervonenkis. 1971a. “On the Uniform
Convergence of Relative Frequencies of Events to Their
Probabilities.” Theory Probab. Appl. 16 (2): 264–81.
Vapnik, V., and A. Chervonenkis. 1981. “The Necessary and
Sufficient Conditions for the Uniform Convergence of Averages to Their
Expected Values.” Teoriya Veroyatnostei i Ee Primeneniya
26 (3): 543–64.
Vapnik, V., and A. Chervonenkis. 1991. “The Necessary and
Sufficient Conditions for Consistency in the Empirical Risk Minimization
Method.” Pattern Recognition and
Image Analysis 1 (3): 283–305.
Vapnik, Vladimir. 1992. “Principles of Risk Minimization for
Learning Theory.” Advances in Neural
Information Processing
Systems, 831–38. https://doi.org/10.5555/2986916.2987019.
Vapnik, Vladimir, Esther Levin, and Yann Le Cun. 1994. “Measuring
the VC-Dimension of a Learning Machine.” Neural
Computation 6 (5): 851–76. https://doi.org/10.1162/neco.1994.6.5.851.
Vasu, Pavan Kumar Anasosalu, James Gabriel, Jeff Zhu, Oncel Tuzel, and
Anurag Ranjan. 2023a. “FastViT: A Fast Hybrid Vision
Transformer Using Structural Reparameterization.” Proceedings
of the IEEE/CVF International
Conference on Computer
Vision. https://arxiv.org/abs/2303.14189.
Vasu, Pavan Kumar Anasosalu, James Gabriel, Jeff Zhu, Oncel Tuzel, and
Anurag Ranjan. 2023b. “MobileOne: An Improved One
Millisecond Mobile Backbone.” Proceedings of the
IEEE/CVF Conference on
Computer Vision and Pattern
Recognition. https://arxiv.org/abs/2206.04040.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. 2017.
“Attention Is All You Need.” Advances in
Neural Information Processing
Systems, 5998–6008. https://doi.org/10.5555/3295222.3295349.
Vershynin, Roman. 2018. High-Dimensional Probability: An
Introduction with Applications in Data Science. Cambridge
University Press.
Vincent, Pascal. 2011. “A Connection Between Score Matching and
Denoising Autoencoders.” Neural Computation 23 (7):
1661–74.
Vyas, Nikhil, Depen Morwani, Rosie Zhao, et al. 2024.
“SOAP: Improving and Stabilizing Shampoo
Using Adam.” arXiv Preprint
arXiv:2409.11321.
Wahba, Grace. 1990. Spline Models for
Observational Data. SIAM. https://doi.org/10.1137/1.9781611970128.
Waibel, Alex, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and
Kevin J Lang. 1989. “Phoneme Recognition Using Time-Delay Neural
Networks.” IEEE Transactions on
Acoustics, Speech, and Signal
Processing 37 (3): 328–39. https://doi.org/10.1016/b978-0-08-051584-7.50037-1.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical
Models, Exponential Families, and Variational Inference.”
Foundations and Trends in Machine Learning 1 (1–2): 1–305.
Waleffe, Roger, Wonmin Byeon, Duncan Riach, et al. 2024. “An
Empirical Study of Mamba-Based Language Models.”
arXiv Preprint arXiv:2406.07887.
Wan, Li, Matthew Zeiler, Sixin Zhang, Yann LeCun, and Rob Fergus. 2013.
“Regularization of Neural Networks Using Dropconnect.”
International Conference on Machine
Learning, 1058–66.
Wang, Dustin, Rui-Jie Zhu, Steven Abreu, et al. 2025. “A
Systematic Analysis of Hybrid Linear Attention.” arXiv
Preprint arXiv:2507.06457.
Wang, Haotao, Aston Zhang, Shuai Zheng, Xingjian Shi, Mu Li, and
Zhangyang Wang. 2022. “Removing Batch Normalization Boosts
Adversarial Training.” International
Conference on Machine
Learning, 23433–45. https://openreview.net/forum?id=2J8bBfGCPi.
Wang, Junxiong, Daniele Paliotta, Avner May, Alexander M. Rush, and Tri
Dao. 2024. “The Mamba in the Llama:
Distilling and Accelerating Hybrid Models.” Advances in
Neural Information Processing Systems.
Wang, Ke Alexander, Jiaxin Shi, and Emily B. Fox. 2025. “Test-Time
Regression: A Unifying Framework for Designing Sequence Models with
Associative Memory.” arXiv Preprint arXiv:2501.12352.
Wang, Lean, Huazuo Gao, Chenggang Zhao, Xu Sun, and Damai Dai. 2024.
“Auxiliary-Loss-Free Load Balancing Strategy for
Mixture-of-Experts.” arXiv Preprint arXiv:2408.15664.
Wang, Qiang, Bei Li, Tong Xiao, et al. 2019. “Learning Deep
Transformer Models for Machine Translation.” Proceedings of
the 57th Annual Meeting of the
Association for Computational
Linguistics, 1810–22. https://doi.org/10.18653/v1/p19-1176.
Wang, Sinong, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020.
“Linformer: Self-Attention with Linear Complexity.”
arXiv Preprint arXiv:2006.04768.
Wang, Wenhai, Jifeng Dai, Zhe Chen, et al. 2023.
“InternImage: Exploring Large-Scale Vision Foundation
Models with Deformable Convolutions.” Proceedings of the
IEEE/CVF Conference on
Computer Vision and Pattern
Recognition. https://arxiv.org/abs/2211.05778.
Warner, Benjamin, Antoine Chaffin, Benjamin Clavié,
et al. 2024. “Smarter, Better, Faster, Longer: A Modern
Bidirectional Encoder for Fast, Memory Efficient, and Long Context
Finetuning and Inference.” arXiv Preprint
arXiv:2412.13663.
Warstadt, Alex, Amanpreet Singh, and Samuel R Bowman. 2019.
“Neural Network Acceptability Judgments.”
Transactions of the Association for
Computational Linguistics 7: 625–41. https://doi.org/10.1162/tacl_a_00290.
Wasserman, Larry. 2013. All of Statistics:
A Concise Course in
Statistical Inference. Springer. https://link.springer.com/book/10.1007/978-0-387-21736-9.
Wasserstein, Ronald L, and Nicole A Lazar. 2016. “The
ASA Statement on p-Values: Context, Process, and
Purpose.” The American
Statistician 70 (2): 129–33. https://doi.org/10.1080/00031305.2016.1154108.
Watkins, Christopher JCH, and Peter Dayan. 1992.
“Q-Learning.” Machine Learning 8
(3–4): 279–92. https://doi.org/10.1007/bf00992698.
Watson, Geoffrey S. 1964. “Smooth Regression Analysis.”
Sankhyā: The Indian Journal
of Statistics, Series A,
359–72. https://doi.org/10.1007/bf02868765.
Wei, Jason, Yi Tay, Rishi Bommasani, et al.
2022. “Emergent Abilities of Large Language Models.”
ArXiv:2206.07682. https://arxiv.org/abs/2206.07682.
Welford, B. P. 1962. “Note on a Method for Calculating Corrected
Sums of Squares and Products.” Technometrics 4 (3):
419–20.
Welling, Max, and Yee W Teh. 2011. “Bayesian Learning via
Stochastic Gradient Langevin Dynamics.”
Proceedings of the 28th International
Conference on Machine Learning
(ICML-11), 681–88. https://dl.acm.org/doi/10.5555/3104482.3104568.
Wen, Kaiyue, David Hall, Tengyu Ma, and Percy Liang. 2025.
“Fantastic Pretraining Optimizers and Where to Find Them.”
arXiv Preprint arXiv:2509.02046.
Wen, Kaiyue, Zhiyuan Li, Jason Wang, David Hall, Percy Liang, and Tengyu
Ma. 2024. “Understanding Warmup-Stable-Decay Learning Rates: A
River Valley Loss Landscape Perspective.” arXiv Preprint
arXiv:2410.05192.
Wengert, Robert Edwin. 1964. “A Simple Automatic Derivative
Evaluation Program.” Communications of the
ACM 7 (8): 463–64. https://doi.org/10.1145/355588.365726.
Werbos, Paul J. 1990. “Backpropagation Through Time: What It Does
and How to Do It.” Proceedings of the IEEE
78 (10): 1550–60. https://doi.org/10.1109/5.58337.
White, Halbert. 1982. “Maximum Likelihood Estimation of
Misspecified Models.” Econometrica 50 (1): 1–25.
Widrow, Bernard, and Marcian E. Hoff. 1960. “Adaptive Switching
Circuits.” IRE WESCON Convention
Record 4: 96–104.
Wiegreffe, Sarah, and Yuval Pinter. 2019. “Attention Is Not Not
Explanation.” Proceedings of the 2019 Conference on Empirical
Methods in Natural Language Processing, 11–20.
Wiener, Norbert. 1923. “Differential-Space.” Journal of
Mathematics and Physics 2: 131–74.
Wightman, Ross, Hugo Touvron, and Hervé Jégou. 2021.
“ResNet Strikes Back: An Improved Training Procedure
in Timm.” ArXiv:2110.00476. https://arxiv.org/abs/2110.00476.
Wigner, Eugene P. 1958. “On the Distribution of the Roots of
Certain Symmetric Matrices.” Ann. Math.,
325–27. https://doi.org/10.2307/1970079.
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following
Algorithms for Connectionist Reinforcement Learning.” Machine
Learning 8: 229–56.
Williams, Samuel, Andrew Waterman, and David Patterson. 2009.
Roofline: An Insightful Visual Performance Model for Floating-Point
Programs and Multicore Architectures. Lawrence
Berkeley National Lab. https://doi.org/10.2172/1407078.
Wilson, Andrew G, and Pavel Izmailov. 2020. “Bayesian Deep
Learning and a Probabilistic Perspective of Generalization.”
Advances in Neural Information
Processing Systems 33: 4697–708. https://arxiv.org/abs/2002.08791.
Wistuba, M., A. Rawat, and T. Pedapati. 2019. “A Survey on Neural
Architecture Search.” ArXiv:1905.01392
[Cs.LG]. https://arxiv.org/abs/1905.01392.
Wistuba, M., N. Schilling, and L. Schmidt-Thieme. 2018. “Scalable
Gaussian Process-Based Transfer Surrogates for
Hyperparameter Optimization.” Machine
Learning 108: 43–78. https://doi.org/10.1007/s10994-017-5684-y.
Witten, Ian H., Radford M. Neal, and John G. Cleary. 1987.
“Arithmetic Coding for Data Compression.”
Communications of the ACM 30 (6): 520–40.
Wolpert, David H. 1996. “The Lack of a Priori Distinctions Between
Learning Algorithms.” Neural Computation 8
(7): 1341–90.
Woo, Sanghyun, Shoubhik Debnath, Ronghang Hu, et al. 2023.
“ConvNeXt V2: Co-Designing and Scaling
ConvNets with Masked Autoencoders.”
Proceedings of the IEEE/CVF
Conference on Computer Vision and
Pattern Recognition. https://arxiv.org/abs/2301.00808.
Wood, Frank, Jan Gasthaus, Cédric Archambeau, Lancelot James, and Yee
Whye Teh. 2011. “The Sequence Memoizer.” Communications
of the ACM 54 (2): 91–98. https://doi.org/10.1162/neco_a_00154.
Wortsman, Mitchell, Gabriel Ilharco, Samir Yitzhak Gadre, et al. 2022.
“Model Soups: Averaging Weights of Multiple Fine-Tuned Models
Improves Accuracy Without Increasing Inference Time.”
International Conference on Machine Learning.
Wu, Bichen, Alvin Wan, Xiangyu Yue, et al. 2018. “Shift: A Zero
Flop, Zero Parameter Alternative to Spatial Convolutions.”
Proceedings of the IEEE Conference on
Computer Vision and Pattern
Recognition, 9127–35. https://doi.org/10.1109/cvpr.2018.00951.
Wu, C. F. Jeff. 1983. “On the Convergence Properties of the
EM Algorithm.” Annals of Statistics 11 (1):
95–103.
Wu, Yonghui, Mike Schuster, Zhifeng Chen, et
al. 2016. “Google’s Neural Machine Translation System:
Bridging the Gap Between Human and Machine Translation.”
ArXiv:1609.08144. https://arxiv.org/abs/1609.08144.
Wu, Yuxin, and Kaiming He. 2018. “Group Normalization.”
European Conference on Computer
Vision. https://arxiv.org/abs/1803.08494.
Xiao, Guangxuan, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis.
2024. “Efficient Streaming Language Models with Attention
Sinks.” International Conference on Learning
Representations.
Xiao, Han, Kashif Rasul, and Roland Vollgraf. 2017.
“Fashion-MNIST: A Novel Image Dataset for
Benchmarking Machine Learning Algorithms.”
ArXiv:1708.07747. https://arxiv.org/abs/1708.07747.
Xiao, Lechao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz,
and Jeffrey Pennington. 2018. “Dynamical Isometry and a Mean Field
Theory of CNNs: How to Train 10,000-Layer Vanilla
Convolutional Neural Networks.” International
Conference on Machine
Learning, 5393–402. https://proceedings.neurips.cc/paper/2018/hash/d76b67bcd3ec823de384ef62dc7e4c7e-Abstract.html.
Xie, Saining, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He.
2017. “Aggregated Residual Transformations for Deep Neural
Networks.” Proceedings of the IEEE
Conference on Computer Vision and
Pattern Recognition, 1492–500. https://doi.org/10.1109/cvpr.2017.634.
Xiong, Ruibin, Yunchang Yang, Di He, et al. 2020. “On Layer
Normalization in the Transformer Architecture.”
International Conference on
Machine Learning, 10524–33. https://proceedings.mlr.press/v119/xiong20b.html.
Xiong, Wayne, Lingfeng Wu, Fil Alleva, Jasha Droppo, Xuedong Huang, and
Andreas Stolcke. 2018. “The Microsoft 2017
Conversational Speech Recognition System.” 2018
IEEE International Conference on
Acoustics, Speech and Signal
Processing (ICASSP), 5934–38. https://doi.org/10.1109/TASLP.2018.2876459.
Xu, Yuanzhong, HyoukJoong Lee, Dehao Chen, et
al. 2021. “GSPMD: General and Scalable
Parallelization for ML Computation Graphs.”
ArXiv:2105.04663. https://arxiv.org/abs/2105.04663.
Yamaguchi, Kouichi, Kenji Sakamoto, Toshio Akabane, and Yoshiji
Fujimoto. 1990. “A Neural Network for Speaker-Independent Isolated
Word Recognition.” First International
Conference on Spoken Language
Processing. https://doi.org/10.21437/icslp.1990-282.
Yang, An, Anfeng Li, Baosong Yang, et al.
2025. “Qwen3 Technical Report.” arXiv
Preprint arXiv:2505.09388.
Yang, Greg, and Edward J. Hu. 2021. “Tensor Programs
IV: Feature Learning in Infinite-Width Neural
Networks.” International Conference on Machine Learning,
11727–37.
Yang, Greg, Edward J. Hu, Igor Babuschkin, et al. 2022. “Tensor
Programs V: Tuning Large Neural Networks via Zero-Shot
Hyperparameter Transfer.” arXiv Preprint
arXiv:2203.03466.
Yang, Greg, James B. Simon, and Jeremy Bernstein. 2023. “A
Spectral Condition for Feature Learning.” arXiv Preprint
arXiv:2310.17813.
Yang, Songlin, Jan Kautz, and Ali Hatamizadeh. 2025. “Gated Delta
Networks: Improving Mamba2 with Delta Rule.”
International Conference on Learning Representations. https://arxiv.org/abs/2412.06464.
Yang, Songlin, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim.
2024. “Gated Linear Attention Transformers with Hardware-Efficient
Training.” International Conference on Machine Learning,
56501–23. https://arxiv.org/abs/2312.06635.
Yang, Songlin, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. 2024.
“Parallelizing Linear Transformers with the Delta Rule over
Sequence Length.” Advances in Neural Information Processing
Systems 37. https://arxiv.org/abs/2406.06484.
Yang, Zichao, Zhiting Hu, Yuntian Deng, Chris Dyer, and Alex Smola.
2016. “Neural Machine Translation with Recurrent Attention
Modeling.” ArXiv:1607.05108. https://arxiv.org/abs/1607.05108.
Yang, Zichao, Marcin Moczulski, Misha Denil, et al. 2015. “Deep
Fried Convnets.” Proceedings of the IEEE
International Conference on
Computer Vision, 1476–83. https://doi.org/10.1109/iccv.2015.173.
Ye, Mao, Peifeng Yin, Wang-Chien Lee, and Dik-Lun Lee. 2011.
“Exploiting Geographical Influence for Collaborative
Point-of-Interest Recommendation.” Proceedings of the 34th
International ACM SIGIR
Conference on Research and
Development in Information
Retrieval, 325–34. https://doi.org/10.1145/2009916.2009962.
You, Yang, Igor Gitman, and Boris Ginsburg. 2017. “Large Batch
Training of Convolutional Networks.”
ArXiv:1708.03888. https://arxiv.org/abs/1708.03888.
Yu, Fisher, and Vladlen Koltun. 2016. “Multi-Scale Context
Aggregation by Dilated Convolutions.”
International Conference on
Learning Representations. https://arxiv.org/abs/1511.07122.
Yuan, Jingyang, Huazuo Gao, Damai Dai, et
al. 2025. “Native Sparse Attention: Hardware-Aligned and
Natively Trainable Sparse Attention.” arXiv Preprint
arXiv:2502.11089.
Yun, Sangdoo, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe,
and Youngjoon Yoo. 2019. “CutMix: Regularization
Strategy to Train Strong Classifiers with Localizable Features.”
Proceedings of the IEEE/CVF
International Conference on
Computer Vision. https://arxiv.org/abs/1905.04899.
Zaheer, Manzil, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv
Kumar. 2018. “Adaptive Methods for Nonconvex Optimization.”
Advances in Neural Information
Processing Systems, 9793–803. https://proceedings.neurips.cc/paper/2018/hash/90365351ccc7437a1309dc64e4db32a3-Abstract.html.
Zeiler, Matthew D. 2012. “ADADELTA: An Adaptive
Learning Rate Method.” ArXiv:1212.5701. https://arxiv.org/abs/1212.5701.
Zeiler, Matthew D, and Rob Fergus. 2013. “Stochastic Pooling for
Regularization of Deep Convolutional Neural Networks.”
ArXiv:1301.3557. https://arxiv.org/abs/1301.3557.
Zeng, Aohan, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin
Chen, et al. 2025. “GLM-4.5: Agentic,
Reasoning, and Coding (ARC) Foundation Models.”
arXiv Preprint arXiv:2508.06471.
Zhang, Aston, Yi Tay, Shuai Zhang, et al. 2021. “Beyond
Fully-Connected Layers with Quaternions: Parameterization of
Hypercomplex Multiplications with 1/n Parameters.”
International Conference on
Learning Representations. https://openreview.net/forum?id=rcQdycl0zyk.
Zhang, Biao, and Rico Sennrich. 2019. “Root Mean Square Layer
Normalization.” Advances in Neural Information Processing
Systems 32.
Zhang, Chiyuan, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol
Vinyals. 2021. “Understanding Deep Learning (Still) Requires
Rethinking Generalization.” Communications of the
ACM 64 (3): 107–15. https://doi.org/10.1145/3446776.
Zhang, Hanlin, Depen Morwani, Nikhil Vyas, et al. 2024. “How Does
Critical Batch Size Scale in Pre-Training?” arXiv Preprint
arXiv:2410.21676.
Zhang, Hongyi, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz.
2018. “mixup: Beyond Empirical Risk
Minimization.” International
Conference on Learning
Representations. https://arxiv.org/abs/1710.09412.
Zhang, Jingzhao, Sai Praneeth Karimireddy, Andreas Veit, et al. 2020.
“Why Are Adaptive Methods Good for Attention Models?”
Advances in Neural Information Processing Systems 33.
Zhang, Richard. 2019. “Making Convolutional Networks
Shift-Invariant Again.” International
Conference on Machine
Learning. https://arxiv.org/abs/1904.11486.
Zhang, Shuai, Lina Yao, Aixin Sun, and Yi Tay. 2019. “Deep
Learning Based Recommender System: A Survey and New
Perspectives.” ACM Computing
Surveys 52 (1): 5. https://doi.org/10.1145/3285029.
Zhang, Susan, Stephen Roller, Naman Goyal, et
al. 2022. “OPT: Open Pre-Trained Transformer
Language Models.” ArXiv:2205.01068. https://arxiv.org/abs/2205.01068.
Zhang, Wei, Jun Tanida, Kazuyoshi Itoh, and Yoshiki Ichioka. 1988.
“Shift-Invariant Pattern Recognition Neural Network and Its
Optical Architecture.” Proceedings of Annual
Conference of the Japan Society
of Applied Physics.
Zhang, Yushun, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and
Zhi-Quan Luo. 2024. “Why Transformers Need Adam: A
Hessian Perspective.” Advances in Neural
Information Processing Systems 37. https://arxiv.org/abs/2402.16788.
Zhao, Yanli, Andrew Gu, Rohan Varma, et al.
2023. “PyTorch FSDP:
Experiences on Scaling Fully Sharded Data Parallel.”
Proceedings of the VLDB Endowment 16
(12): 3848–60. https://doi.org/10.14778/3611540.3611569.
Zhao, Zhong-Qiu, Peng Zheng, Shou-tao Xu, and Xindong Wu. 2019.
“Object Detection with Deep Learning: A Review.”
IEEE Transactions on Neural
Networks and Learning
Systems 30 (11): 3212–32. https://doi.org/10.1016/j.neucom.2018.09.013.
Zhu, Jun-Yan, Taesung Park, Phillip Isola, and Alexei A Efros. 2017.
“Unpaired Image-to-Image Translation Using Cycle-Consistent
Adversarial Networks.” Proceedings of the IEEE
International Conference on
Computer Vision, 2223–32. https://doi.org/10.1109/iccv.2017.244.
Zhu, Yukun, Ryan Kiros, Rich Zemel, et al. 2015. “Aligning Books
and Movies: Towards Story-Like Visual Explanations by Watching Movies
and Reading Books.” Proceedings of the IEEE
International Conference on
Computer Vision, 19–27. https://doi.org/10.1109/iccv.2015.11.
Zoph, Barret, and Quoc V Le. 2016. “Neural Architecture Search
with Reinforcement Learning.”
ArXiv:1611.01578. https://arxiv.org/abs/1611.01578.
Zuo, Jingwei, Maksim Velikanov, Ilyas Chahed, et
al. 2025. “Falcon-H1: A Family of Hybrid-Head
Language Models Redefining Efficiency and Performance.” arXiv
Preprint arXiv:2507.22448.