References

Abadi, Martı́n, Paul Barham, Jianmin Chen, et al. 2016. “TensorFlow: A System for Large-Scale Machine Learning.” 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 265–83. https://www.usenix.org/conference/osdi16/technical-sessions/presentation/abadi.
Abdel-Hamid, Ossama, Abdel-Rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu. 2014. “Convolutional Neural Networks for Speech Recognition.” IEEE/ACM Transactions on Audio, Speech, and Language Processing 22 (10): 1533–45. https://doi.org/10.1109/taslp.2014.2339736.
Ainslie, Joshua, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4895–901.
Akiba, T., S. Sano, T. Yanase, T. Ohta, and M. Koyama. 2019. Optuna: A Next-Generation Hyperparameter Optimization Framework.” Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. https://doi.org/10.1145/3292500.3330701.
Alayrac, Jean-Baptiste, Jeff Donahue, Pauline Luc, et al. 2022. “Flamingo: A Visual Language Model for Few-Shot Learning.” ArXiv:2204.14198. https://arxiv.org/abs/2204.14198.
Albergo, Michael S., Nicholas M. Boffi, and Eric Vanden-Eijnden. 2023. “Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.” arXiv Preprint arXiv:2303.08797.
Alemi, Alexander A., Ian Fischer, Joshua V. Dillon, and Kevin Murphy. 2017. “Deep Variational Information Bottleneck.” International Conference on Learning Representations.
Alsallakh, Bilal, Narine Kokhlikyan, Vivek Miglani, Jun Yuan, and Orion Reblitz-Richardson. 2020. “Mind the PADCNNs Can Develop Blind Spots.” ArXiv:2010.02178. https://arxiv.org/abs/2010.02178.
Amari, Shun-ichi. 1998. “Natural Gradient Works Efficiently in Learning.” Neural Computation 10 (2): 251–76.
Amari, Shun-ichi. 2016. Information Geometry and Its Applications. Springer.
Anderson, Brian D. O. 1982. “Reverse-Time Diffusion Equation Models.” Stochastic Processes and Their Applications 12 (3): 313–26.
Anil, Rohan, Andrew M Dai, Orhan Firat, et al. 2023. “PaLM 2 Technical Report.” ArXiv:2305.10403. https://arxiv.org/abs/2305.10403.
Anil, Rohan, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. 2020. “Scalable Second-Order Optimization for Deep Learning.” ArXiv:2002.09018. https://arxiv.org/abs/2002.09018.
Ansel, Jason, Edward Yang, Horace He, et al. 2024. PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation.” 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 929–47. https://doi.org/10.1145/3620665.3640366.
Arjevani, Yossi, Yair Carmon, John C. Duchi, Dylan J. Foster, Nathan Srebro, and Blake Woodworth. 2023. “Lower Bounds for Non-Convex Stochastic Optimization.” Mathematical Programming 199: 165–214.
Arjovsky, Martin, Soumith Chintala, and Léon Bottou. 2017. Wasserstein Generative Adversarial Networks.” Proceedings of the 34th International Conference on Machine Learning, 214–23.
Armijo, Larry. 1966. “Minimization of Functions Having Lipschitz Continuous First Partial Derivatives.” Pacific Journal of Mathematics 16 (1): 1–3.
Aronszajn, Nachman. 1950. “Theory of reproducing kernels.” Transactions of the American Mathematical Society 68 (3): 337–404. https://doi.org/10.1090/s0002-9947-1950-0051437-7.
Arora, Simran, Sabri Eyuboglu, Aman Timalsina, et al. 2024. Zoology: Measuring and Improving Recall in Efficient Language Models.” International Conference on Learning Representations. https://arxiv.org/abs/2312.04927.
Arora, Simran, Sabri Eyuboglu, Michael Zhang, et al. 2024. “Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff.” International Conference on Machine Learning.
Arpit, Devansh, Stanisław Jastrzȩbski, Nicolas Ballas, et al. 2017. “A Closer Look at Memorization in Deep Networks.” International Conference on Machine Learning, 233–42.
Åström, Karl Johan, and Richard M. Murray. 2021. Feedback Systems: An Introduction for Scientists and Engineers. 2nd ed. Princeton University Press.
Austin, Jacob, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. 2021. “Structured Denoising Diffusion Models in Discrete State-Spaces.” Advances in Neural Information Processing Systems 34.
Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. “Layer Normalization.” ArXiv:1607.06450. https://arxiv.org/abs/1607.06450.
Bae, Sangmin, Bilge Acun, Chien-Yu Lin, et al. 2025. “Hybrid Architectures for Language Models: Systematic Analysis and Design Insights.” arXiv Preprint arXiv:2510.04800.
Baevski, Alexei, and Michael Auli. 2018. “Adaptive Input Representations for Neural Language Modeling.” International Conference on Learning Representations. https://openreview.net/forum?id=ByxZX20qFQ.
Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio. 2015. “Neural Machine Translation by Jointly Learning to Align and Translate.” International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1409.0473.
Bai, Shaojie, J. Zico Kolter, and Vladlen Koltun. 2019. “Deep Equilibrium Models.” Advances in Neural Information Processing Systems 32.
Baptista, R., and M. Poloczek. 2018. Bayesian Optimization of Combinatorial Structures.” Proceedings of the 35th International Conference on Machine Learning. https://proceedings.mlr.press/v80/baptista18a.html.
Barber, David, and Felix Agakov. 2003. “The IM Algorithm: A Variational Approach to Information Maximization.” Advances in Neural Information Processing Systems 16.
Bardenet, R., M. Brendel, B. Kégl, and M. Sebag. 2013. “Collaborative Hyperparameter Tuning.” Proceedings of the 30th International Conference on Machine Learning (ICML’13). https://proceedings.mlr.press/v28/bardenet13.html.
Bartlett, Peter L., Philip M. Long, Gábor Lugosi, and Alexander Tsigler. 2020. “Benign Overfitting in Linear Regression.” Proceedings of the National Academy of Sciences 117 (48): 30063–70.
Bartlett, Peter L., and Shahar Mendelson. 2002. “Rademacher and Gaussian Complexities: Risk Bounds and Structural Results.” Journal of Machine Learning Research 3: 463–82.
Bartlett, Peter L, Andrea Montanari, and Alexander Rakhlin. 2021. “Deep Learning: A Statistical Viewpoint.” ArXiv:2103.09177. https://arxiv.org/abs/2103.09177.
Baur, Walter, and Volker Strassen. 1983. “The Complexity of Partial Derivatives.” Theoretical Computer Science 22 (3): 317–30.
Bay, Herbert, Tinne Tuytelaars, and Luc Van Gool. 2006. SURF: Speeded up Robust Features.” European Conference on Computer Vision, 404–17. https://doi.org/10.1007/11744023_32.
Baydin, Atilim Gunes, Barak A Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind. 2018. “Automatic Differentiation in Machine Learning: A Survey.” Journal of Machine Learning Research 18 (153): 1–43. https://www.jmlr.org/papers/volume18/17-468/17-468.pdf.
Bayes, Thomas. 1763. “An Essay Towards Solving a Problem in the Doctrine of Chances.” Philosophical Transactions of the Royal Society of London 53: 370–418.
Beck, Amir, and Marc Teboulle. 2009. “A Fast Iterative Shrinkage-Thresholding Algorithm for Linear Inverse Problems.” SIAM Journal on Imaging Sciences 2 (1): 183–202.
Beck, Maximilian, Korbinian Pöppel, Markus Spanring, et al. 2024. xLSTM: Extended Long Short-Term Memory.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2405.04517.
Behrouz, Ali, Peilin Zhong, and Vahab Mirrokni. 2025. “Titans: Learning to Memorize at Test Time.” arXiv Preprint arXiv:2501.00663.
Belghazi, Mohamed Ishmael, Aristide Baratin, Sai Rajeshwar, et al. 2018. “Mutual Information Neural Estimation.” Proceedings of the 35th International Conference on Machine Learning, 531–40.
Belkin, Mikhail, Daniel Hsu, Siyuan Ma, and Soumik Mandal. 2019. “Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-Off.” Proceedings of the National Academy of Sciences 116 (32): 15849–54.
Bellman, R. 1966. “Dynamic Programming.” Science 153: 34–37. https://doi.org/10.1126/science.153.3731.34.
Bellman, Richard. 1952. “On the Theory of Dynamic Programming.” Proceedings of the National Academy of Sciences 38 (8): 716–19. https://doi.org/10.1073/pnas.38.8.716.
Bellman, Richard. 1957a. “A Markovian Decision Process.” Journal of Mathematics and Mechanics 6 (5): 679–84. http://www.jstor.org/stable/24900506.
Bellman, Richard. 1957b. Dynamic Programming. Dover. Dover Publications. https://doi.org/10.1515/9781400835386.
Beltagy, Iz, Matthew E Peters, and Arman Cohan. 2020. “Longformer: The Long-Document Transformer.” ArXiv:2004.05150. https://arxiv.org/abs/2004.05150.
Benamou, Jean-David, and Yann Brenier. 2000. “A Computational Fluid Mechanics Solution to the Monge–Kantorovich Mass Transfer Problem.” Numerische Mathematik 84 (3): 375–93.
Bengio, Yoshua, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003. “A Neural Probabilistic Language Model.” Journal of Machine Learning Research 3 (Feb): 1137–55. https://jmlr.org/papers/v3/bengio03a.html.
Bengio, Yoshua, and Yves Grandvalet. 2004. “No Unbiased Estimator of the Variance of k-Fold Cross-Validation.” Journal of Machine Learning Research 5: 1089–105.
Bengio, Yoshua, Nicholas Léonard, and Aaron Courville. 2013. “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.” arXiv Preprint arXiv:1308.3432.
Bengio, Yoshua, Patrice Simard, and Paolo Frasconi. 1994. “Learning Long-Term Dependencies with Gradient Descent Is Difficult.” IEEE Transactions on Neural Networks 5 (2): 157–66. https://doi.org/10.1109/72.279181.
Benjamini, Yoav, and Yosef Hochberg. 1995. “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing.” Journal of the Royal Statistical Society: Series B (Methodological) 57 (1): 289–300.
Bergsma, Shane, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, and Joel Hestness. 2025a. “Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-Training.” Advances in Neural Information Processing Systems. https://arxiv.org/abs/2505.13738.
Bergsma, Shane, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, and Joel Hestness. 2025b. “Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLM Pretraining.” International Conference on Learning Representations. https://arxiv.org/abs/2502.15938.
Bergstra, James, and Yoshua Bengio. 2012. “Random Search for Hyper-Parameter Optimization.” Journal of Machine Learning Research 13: 281–305.
Bergstra, James, Olivier Breuleux, Frédéric Bastien, et al. 2010. “Theano: A CPU and GPU Math Compiler in Python.” Proc. 9th Python in Science Conference 1: 3–10. https://www.iro.umontreal.ca/~lisa/pointeurs/theano_scipy2010.pdf.
Bernstein, Jeremy. 2025. Deriving Muon. Https://jeremybernste.in/writing/deriving-muon.
Bernstein, Jeremy, and Laker Newhouse. 2024. “Old Optimizer, New Norm: An Anatomy.” arXiv Preprint arXiv:2409.20325.
Beutel, Alex, Kenton Murray, Christos Faloutsos, and Alexander J Smola. 2014. “CoBaFi: Collaborative Bayesian Filtering.” Proceedings of the 23rd International Conference on World Wide Web, 97–108. https://doi.org/10.1145/2566486.2567980.
Beyer, Lucas, Olivier J. Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord. 2020. “Are We Done with ImageNet?” ArXiv:2006.07159. https://arxiv.org/abs/2006.07159.
Bick, Aviv, Kevin Y. Li, Eric P. Xing, J. Zico Kolter, and Albert Gu. 2024. “Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models.” Advances in Neural Information Processing Systems.
Bishop, Chris M. 1995. “Training with Noise Is Equivalent to Tikhonov Regularization.” Neural Computation 7 (1): 108–16. https://doi.org/10.1162/neco.1995.7.1.108.
Bishop, Christopher M. 2006. Pattern Recognition and Machine Learning. Springer. https://doi.org/10.1007/978-0-387-45528-0.
Black, Fischer, and Myron Scholes. 1973. “The Pricing of Options and Corporate Liabilities.” Journal of Political Economy 81: 637–54. https://doi.org/10.1086/260062.
Blelloch, Guy E. 1990. Prefix Sums and Their Applications. CMU-CS-90-190. School of Computer Science, Carnegie Mellon University. https://www.cs.cmu.edu/~scandal/papers/CMU-CS-90-190.html.
Bodla, Navaneeth, Bharat Singh, Rama Chellappa, and Larry S Davis. 2017. “Soft-NMS-Improving Object Detection with One Line of Code.” Proceedings of the IEEE International Conference on Computer Vision, 5561–69. https://doi.org/10.1109/iccv.2017.593.
Bojanowski, Piotr, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. “Enriching Word Vectors with Subword Information.” Transactions of the Association for Computational Linguistics 5: 135–46. https://doi.org/10.1162/tacl_a_00051.
Bollobás, B. 1999. Linear Analysis. Cambridge University Press. https://doi.org/10.1017/CBO9780511626296.
Bolte, Jérôme, and Edouard Pauwels. 2021. “Conservative Set Valued Fields, Automatic Differentiation, Stochastic Gradient Methods and Deep Learning.” Mathematical Programming 188 (1): 19–51.
Bommasani, Rishi, Drew A Hudson, Ehsan Adeli, et al. 2021. “On the Opportunities and Risks of Foundation Models.” ArXiv:2108.07258. https://arxiv.org/abs/2108.07258.
Bonferroni, Carlo E. 1936. “Teoria Statistica Delle Classi e Calcolo Delle Probabilità.” Pubblicazioni Del R. Istituto Superiore Di Scienze Economiche e Commerciali Di Firenze 8: 3–62.
Bottou, Léon. 2010. “Large-Scale Machine Learning with Stochastic Gradient Descent.” In Proceedings of COMPSTAT’2010. Springer. https://doi.org/10.1201/b11429-4.
Bottou, Léon, Frank E. Curtis, and Jorge Nocedal. 2018. “Optimization Methods for Large-Scale Machine Learning.” SIAM Review 60 (2): 223–311.
Bottou, Léon, and Yann Le Cun. 1988. SN: A Simulator for Connectionist Models.” Proceedings of NeuroNimes 88 (Nimes, France), 371–82. http://leon.bottou.org/papers/bottou-lecun-88.
Boucheron, Stéphane, Olivier Bousquet, and Gábor Lugosi. 2005. “Theory of Classification: A Survey of Some Recent Advances.” ESAIM: Probability and Statistics 9: 323–75. https://doi.org/10.1214/154957805100000046.
Boucheron, Stéphane, Gábor Lugosi, and Pascal Massart. 2013. Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press.
Bowman, Samuel R, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015. “A Large Annotated Corpus for Learning Natural Language Inference.” Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 632–42. https://arxiv.org/abs/1508.05326.
Boyd, Stephen, and Lieven Vandenberghe. 2004. Convex Optimization. Cambridge University Press. https://web.stanford.edu/~boyd/cvxbook/.
Bradley, Ralph Allan, and Milton E Terry. 1952. “Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons.” Biometrika 39 (3/4): 324–45. https://doi.org/10.1093/biomet/41.3-4.502.
Bray, Alan J., and David S. Dean. 2007. “Statistics of Critical Points of Gaussian Fields on Large-Dimensional Spaces.” Physical Review Letters 98 (15): 150201.
Brenier, Yann. 1991. “Polar Factorization and Monotone Rearrangement of Vector-Valued Functions.” Communications on Pure and Applied Mathematics 44 (4): 375–417.
Brin, Sergey, and Lawrence Page. 1998. “The Anatomy of a Large-Scale Hypertextual Web Search Engine.” Computer Networks and ISDN Systems 30 (1–7): 107–17.
Brock, Andrew, Soham De, Samuel L. Smith, and Karen Simonyan. 2021. “High-Performance Large-Scale Image Recognition Without Normalization.” International Conference on Machine Learning. https://arxiv.org/abs/2102.06171.
Brown, Noam, and Tuomas Sandholm. 2017. “Libratus: The Superhuman AI for No-Limit Poker.” IJCAI, 5226–28. https://doi.org/10.24963/ijcai.2017/772.
Brown, Peter F, John Cocke, Stephen A Della Pietra, et al. 1990. “A Statistical Approach to Machine Translation.” Computational Linguistics 16 (2): 79–85. https://doi.org/10.1162/coli.1990.16.2.79.
Brown, Tom, Benjamin Mann, Nick Ryder, et al. 2020. “Language Models Are Few-Shot Learners.” Advances in Neural Information Processing Systems 33: 1877–901. https://arxiv.org/abs/2005.14165.
Buslaev, Alexander, Vladimir I Iglovikov, Eugene Khvedchenya, Alex Parinov, Mikhail Druzhinin, and Alexandr A Kalinin. 2020. “Albumentations: Fast and Flexible Image Augmentations.” Information 11 (2): 125. https://doi.org/10.3390/info11020125.
Campbell, Murray, A Joseph Hoane Jr, and Feng-hsiung Hsu. 2002. “Deep Blue.” Artificial Intelligence 134 (1-2): 57–83. https://doi.org/10.1016/s0004-3702(01)00129-1.
Candès, Emmanuel J., and Benjamin Recht. 2009. “Exact Matrix Completion via Convex Optimization.” Foundations of Computational Mathematics 9 (6): 717–72.
Canny, John. 1987. “A Computational Approach to Edge Detection.” In Readings in Computer Vision. Elsevier. https://doi.org/10.1109/tpami.1986.4767851.
Carion, Nicolas, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. “End-to-End Object Detection with Transformers.” European Conference on Computer Vision, 213–29.
Cer, Daniel, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. “SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation.” Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), 1–14. https://doi.org/10.18653/v1/s17-2001.
Chebyshev, Pafnuty L. 1867. “Des Valeurs Moyennes.” Journal de Mathématiques Pures Et Appliquées 12 (2): 177–84.
Chechik, Gal, Amir Globerson, Naftali Tishby, and Yair Weiss. 2005. “Information Bottleneck for Gaussian Variables.” Journal of Machine Learning Research 6: 165–88.
Chen, Lili, Kevin Lu, Aravind Rajeswaran, et al. 2021. “Decision Transformer: Reinforcement Learning via Sequence Modeling.” Advances in Neural Information Processing Systems 34: 15084–97. https://arxiv.org/abs/2106.01345.
Chen, Ricky T. Q., Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. 2018. “Neural Ordinary Differential Equations.” Advances in Neural Information Processing Systems 31.
Chen, Shouyuan, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. “Extending Context Window of Large Language Models via Positional Interpolation.” arXiv Preprint arXiv:2306.15595.
Chen, Tianqi, Mu Li, Yutian Li, et al. 2015. MXNET: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems.” ArXiv:1512.01274. https://arxiv.org/abs/1512.01274.
Chen, Tianqi, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. “Training Deep Nets with Sublinear Memory Cost.” arXiv Preprint arXiv:1604.06174.
Chen, Ting, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. “A Simple Framework for Contrastive Learning of Visual Representations.” Proceedings of the 37th International Conference on Machine Learning, 1597–607.
Chen, Xiangning, Chen Liang, Da Huang, Esteban Real, et al. 2023. “Symbolic Discovery of Optimization Algorithms.” Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2302.06675.
Chetlur, Sharan, Cliff Woolley, Philippe Vandermersch, et al. 2014. “CuDNN: Efficient Primitives for Deep Learning.” ArXiv:1410.0759. https://arxiv.org/abs/1410.0759.
Child, Rewon, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. “Generating Long Sequences with Sparse Transformers.” ArXiv:1904.10509. https://arxiv.org/abs/1904.10509.
Cho, Kyunghyun, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014. “On the Properties of Neural Machine Translation: Encoder–Decoder Approaches.” ArXiv:1409.1259. https://arxiv.org/abs/1409.1259.
Cho, Kyunghyun, Bart Van Merriënboer, Caglar Gulcehre, et al. 2014. “Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation.” Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, 1724–34. https://arxiv.org/abs/1406.1078.
Chollet, François. 2017. “Xception: Deep Learning with Depthwise Separable Convolutions.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/1610.02357.
Choromanski, Krzysztof, Valerii Likhosherstov, David Dohan, et al. 2021. “Rethinking Attention with Performers.” International Conference on Learning Representations.
Chowdhery, Aakanksha, Sharan Narang, Jacob Devlin, et al. 2022. “PaLM: Scaling Language Modeling with Pathways.” ArXiv:2204.02311. https://arxiv.org/abs/2204.02311.
Chu, Xiangxiang, Liang Li, and Bo Zhang. 2024. “Make RepVGG Greater Again: A Quantization-Aware Approach.” AAAI Conference on Artificial Intelligence. https://arxiv.org/abs/2212.01593.
Chung, Junyoung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. “Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.” ArXiv:1412.3555. https://arxiv.org/abs/1412.3555.
Church, Kenneth Ward, and Patrick Hanks. 1990. “Word Association Norms, Mutual Information, and Lexicography.” Computational Linguistics 16 (1): 22–29.
Chwialkowski, Kacper, Heiko Strathmann, and Arthur Gretton. 2016. “A Kernel Test of Goodness of Fit.” Proceedings of the 33rd International Conference on Machine Learning, 2606–15.
Coddington, Earl A., and Norman Levinson. 1955. Theory of Ordinary Differential Equations. McGraw-Hill.
Cohen, Jeremy M., Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar. 2021. “Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.” International Conference on Learning Representations.
Collobert, Ronan, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. “Natural Language Processing (Almost) from Scratch.” Journal of Machine Learning Research 12: 2493–537. https://jmlr.org/papers/v12/collobert11a.html.
Cordonnier, Jean-Baptiste, Andreas Loukas, and Martin Jaggi. 2020. “On the Relationship Between Self-Attention and Convolutional Layers.” International Conference on Learning Representations. https://openreview.net/forum?id=HJlnC1rKPB.
Cortes, Corinna, and Vladimir Vapnik. 1995. “Support-Vector Networks.” Machine Learning 20 (3): 273–97.
Cover, Thomas M., and Peter E. Hart. 1967. “Nearest Neighbor Pattern Classification.” IEEE Transactions on Information Theory 13 (1): 21–27.
Cover, T, and JM Thomas. 1999. Elements of Information Theory. John Wiley & Sons. https://doi.org/10.1002/047174882X.
Cramér, H. 1946. Mathematical Methods of Statistics. Princeton University Press.
Csiszár, Imre. 1967. “Information-Type Measures of Difference of Probability Distributions and Indirect Observations.” Studia Scientiarum Mathematicarum Hungarica 2: 299–318.
Csiszár, Imre. 2008. “Axiomatic Characterizations of Information Measures.” Entropy 10 (3): 261–73. https://doi.org/10.3390/e10030261.
Cubuk, Ekin D., Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2020. RandAugment: Practical Automated Data Augmentation with a Reduced Search Space.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. https://arxiv.org/abs/1909.13719.
Cuturi, Marco. 2013. “Sinkhorn Distances: Lightspeed Computation of Optimal Transport.” Advances in Neural Information Processing Systems 26.
Cybenko, George. 1989. “Approximation by Superpositions of a Sigmoidal Function.” Mathematics of Control, Signals and Systems 2 (4): 303–14. https://doi.org/10.1007/bf02134016.
D’Angelo, Francesco, Maksym Andriushchenko, Aditya Varre, and Nicolas Flammarion. 2024. “Why Do We Need Weight Decay in Modern Deep Learning?” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2310.04415.
Dahl, George E., Frank Schneider, Zachary Nado, et al. 2023. “Benchmarking Neural Network Training Algorithms.” arXiv Preprint arXiv:2306.07179.
Dahlquist, Germund G. 1963. “A Special Stability Problem for Linear Multistep Methods.” BIT Numerical Mathematics 3: 27–43.
Dai, Damai, Chengqi Deng, Chenggang Zhao, et al. 2024. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.” arXiv Preprint arXiv:2401.06066.
Dalal, Navneet, and Bill Triggs. 2005. “Histograms of Oriented Gradients for Human Detection.” 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05) 1: 886–93. https://doi.org/10.1109/cvpr.2005.177.
Dao, Tri, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” Advances in Neural Information Processing Systems.
Dao, Tri, and Albert Gu. 2024. “Transformers Are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.” International Conference on Machine Learning, 10041–71. https://arxiv.org/abs/2405.21060.
Dauphin, Yann N., Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio. 2014. “Identifying and Attacking the Saddle Point Problem in High-Dimensional Non-Convex Optimization.” Advances in Neural Information Processing Systems 27: 2933–41.
De Cock, Dean. 2011. “Ames, Iowa: Alternative to the Boston Housing Data as an End of Semester Regression Project.” Journal of Statistics Education 19 (3). https://doi.org/10.1080/10691898.2011.11889627.
De, Soham, Samuel L. Smith, Anushan Fernando, et al. 2024. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.” arXiv Preprint arXiv:2402.19427.
Dean, Jeffrey, Greg S Corrado, Rajat Monga, et al. 2012. “Large Scale Distributed Deep Networks.” Proceedings of the 25th International Conference on Neural Information Processing Systems, Volume 1, 1223–31. https://doi.org/10.5555/2999134.2999271.
DeepSeek-AI. 2024. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.” arXiv Preprint arXiv:2405.04434.
DeepSeek-AI, Xiao Bi, Deli Chen, Guanting Chen, et al. 2024. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.” arXiv Preprint arXiv:2401.02954.
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, et al. 2024. DeepSeek-V3 Technical Report.” arXiv Preprint arXiv:2412.19437.
Defazio, Aaron, Ashok Cutkosky, Harsh Mehta, and Konstantin Mishchenko. 2023. “Optimal Linear Decay Learning Rate Schedules and Further Refinements.” arXiv Preprint arXiv:2310.07831.
Defazio, Aaron, Xingyu Alice Yang, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled, and Ashok Cutkosky. 2024. “The Road Less Scheduled.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2405.15682.
Dehghani, Mostafa, Josip Djolonga, Basil Mustafa, et al. 2023. “Scaling Vision Transformers to 22 Billion Parameters.” International Conference on Machine Learning, 7480–512.
Deisenroth, Marc Peter, A. Aldo Faisal, and Cheng Soon Ong. 2020. Mathematics for Machine Learning. Cambridge University Press. https://mml-book.github.io/.
Delétang, Grégoire, Anian Ruoss, Paul-Ambroise Duquenne, et al. 2023. “Language Modeling Is Compression.” ArXiv:2309.10668. https://arxiv.org/abs/2309.10668.
Dempster, Arthur P, Nan M Laird, and Donald B Rubin. 1977. “Maximum Likelihood from Incomplete Data via the EM Algorithm.” Journal of the Royal Statistical Society: Series B 39 (1): 1–38.
Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. “Imagenet: A Large-Scale Hierarchical Image Database.” 2009 IEEE Conference on Computer Vision and Pattern Recognition, 248–55. https://doi.org/10.1109/cvpr.2009.5206848.
Der Kiureghian, Armen, and Ove Ditlevsen. 2009. “Aleatory or Epistemic? Does It Matter?” Structural Safety 31 (2): 105–12. https://doi.org/10.1016/j.strusafe.2008.06.020.
Dettmers, Tim, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2022. “8-Bit Optimizers via Block-Wise Quantization.” International Conference on Learning Representations. https://arxiv.org/abs/2110.02861.
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding.” ArXiv:1810.04805. https://arxiv.org/abs/1810.04805.
Dey, Nolan, Quentin Anthony, and Joel Hestness. 2024. The Practitioner’s Guide to the Maximal Update Parameterization. Https://blog.eleuther.ai/mutransfer/.
Dey, Nolan, Gurpreet Gosal, Zhiming Chen, et al. 2023. Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster.” arXiv Preprint arXiv:2304.03208.
Dhariwal, Prafulla, and Alex Nichol. 2021. “Diffusion Models Beat GANs on Image Synthesis.” Advances in Neural Information Processing Systems 34.
Ding, Xiaohan, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun. 2021. RepVGG: Making VGG-Style ConvNets Great Again.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2101.03697.
Ding, Xiaohan, Xiangyu Zhang, Yizhuang Zhou, Jungong Han, Guiguang Ding, and Jian Sun. 2022. “Scaling up Your Kernels to 31x31: Revisiting Large Kernel Design in CNNs.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2203.06717.
Ding, Xiaohan, Yiyuan Zhang, Yixiao Ge, et al. 2024. UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2311.15599.
Dinh, Laurent, David Krueger, and Yoshua Bengio. 2014. NICE: Non-Linear Independent Components Estimation.” ArXiv:1410.8516. https://arxiv.org/abs/1410.8516.
Dinh, Laurent, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017. “Sharp Minima Can Generalize for Deep Nets.” Proceedings of the 34th International Conference on Machine Learning, Proceedings of machine learning research, vol. 70: 1019–28. https://proceedings.mlr.press/v70/dinh17b.html.
Dinh, Laurent, Jascha Sohl-Dickstein, and Samy Bengio. 2017. “Density Estimation Using Real NVP.” International Conference on Learning Representations. https://openreview.net/forum?id=HkpbnH9lx.
Doersch, Carl, Abhinav Gupta, and Alexei A Efros. 2015. “Unsupervised Visual Representation Learning by Context Prediction.” Proceedings of the IEEE International Conference on Computer Vision, 1422–30. https://doi.org/10.1109/iccv.2015.167.
Domingos, Pedro, and Michael Pazzani. 1997. “On the Optimality of the Simple Bayesian Classifier Under Zero-One Loss.” Machine Learning 29: 103–30.
Dong, Xin, Yonggan Fu, Shizhe Diao, et al. 2024. “Hymba: A Hybrid-Head Architecture for Small Language Models.” arXiv Preprint arXiv:2411.13676.
Dong, Yihe, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021. “Attention Is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth.” International Conference on Machine Learning, 2793–803.
Donsker, Monroe D., and S. R. Srinivasa Varadhan. 1983. “Asymptotic Evaluation of Certain Markov Process Expectations for Large Time. IV.” Communications on Pure and Applied Mathematics 36 (2): 183–212.
Dormand, John R., and Peter J. Prince. 1980. “A Family of Embedded Runge–Kutta Formulae.” Journal of Computational and Applied Mathematics 6 (1): 19–26.
Dosovitskiy, Alexey, Lucas Beyer, Alexander Kolesnikov, et al. 2021. “An Image Is Worth 16 x 16 Words: Transformers for Image Recognition at Scale.” International Conference on Learning Representations. https://openreview.net/forum?id=YicbFdNTTy.
Duchi, John, Elad Hazan, and Yoram Singer. 2011. “Adaptive Subgradient Methods for Online Learning and Stochastic Optimization.” Journal of Machine Learning Research 12: 2121–59. https://jmlr.org/papers/v12/duchi11a.html.
Duchi, John, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra. 2008. “Efficient Projections onto the 1-Ball for Learning in High Dimensions.” Proceedings of the 25th International Conference on Machine Learning (ICML), 272–79.
Dumoulin, Vincent, and Francesco Visin. 2016. “A Guide to Convolution Arithmetic for Deep Learning.” ArXiv:1603.07285. https://arxiv.org/abs/1603.07285.
Dwivedi, Vijay Prakash, and Xavier Bresson. 2020. “A Generalization of Transformer Networks to Graphs.” ArXiv:2012.09699. https://arxiv.org/abs/2012.09699.
Dwork, Cynthia, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth. 2015. “Preserving Statistical Validity in Adaptive Data Analysis.” ACM Symposium on Theory of Computing, 117–26.
Dwork, Cynthia, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Leon Roth. 2015. “Preserving Statistical Validity in Adaptive Data Analysis.” Proceedings of the 47th Annual ACM Symposium on Theory of Computing, 117–26. https://doi.org/10.1145/2746539.2746580.
E, Weinan. 2017. “A Proposal on Machine Learning via Dynamical Systems.” Communications in Mathematics and Statistics 5 (1): 1–11.
Eckart, Carl, and Gale Young. 1936. “The Approximation of One Matrix by Another of Lower Rank.” Psychometrika 1 (3): 211–18.
Efron, Bradley. 1979. “Bootstrap Methods: Another Look at the Jackknife.” Annals of Statistics 7 (1): 1–26.
Efron, Bradley. 2011. “Tweedie’s Formula and Selection Bias.” Journal of the American Statistical Association 106 (496): 1602–14.
Efron, Bradley, and Trevor Hastie. 2016. Computer Age Statistical Inference: Algorithms, Evidence, and Data Science. Cambridge University Press. https://hastie.su.domains/CASI/.
Elhage, Nelson, Neel Nanda, Catherine Olsson, et al. 2021. “A Mathematical Framework for Transformer Circuits.” Transformer Circuits Thread.
Elsken, T., J. H. Metzen, and F. Hutter. 2018. “Neural Architecture Search: A Survey.” ArXiv:1808.05377 [Stat.ML]. https://arxiv.org/abs/1808.05377.
Endres, Dominik M., and Johannes E. Schindelin. 2003. “A New Metric for Probability Distributions.” IEEE Transactions on Information Theory 49 (7): 1858–60.
Erven, Tim van, and Peter Harremoës. 2014. “Rényi Divergence and Kullback–Leibler Divergence.” IEEE Transactions on Information Theory 60 (7): 3797–820.
Fan, Angela, Mike Lewis, and Yann Dauphin. 2018. “Hierarchical Neural Story Generation.” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 889–98.
Fechner, Gustav Theodor. 1860. Elemente Der Psychophysik. Vol. 2. Breitkopf u. Härtel.
Fedus, William, Barret Zoph, and Noam Shazeer. 2022. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” Journal of Machine Learning Research 23 (120): 1–39. https://arxiv.org/abs/2101.03961.
Feng, Leo, Frederick Tung, Mohamed Osama Ahmed, Yoshua Bengio, and Hossein Hajimirsadeghi. 2024. “Were RNNs All We Needed?” arXiv Preprint arXiv:2410.01201.
Fernando, Randima. 2004. GPU Gems: Programming Techniques, Tips, and Tricks for Real-Time Graphics. Addison-Wesley.
Feurer, M., and F. Hutter. 2019. “Hyperparameter Optimization.” In Automated Machine Learning: Methods, Systems, Challenges. Springer. https://doi.org/10.1007/978-3-030-05318-5_1.
Feurer, M., B. Letham, F. Hutter, and E. Bakshy. 2022. “Practical Transfer Learning for Bayesian Optimization.” ArXiv:1802.02219 [Stat.ML]. https://arxiv.org/abs/1802.02219.
Field, David J. 1987. “Relations Between the Statistics of Natural Images and the Response Properties of Cortical Cells.” JOSA A 4 (12): 2379–94. https://doi.org/10.1364/josaa.4.002379.
Fisher, R A. 1925. Statistical Methods for Research Workers. Oliver & Boyd.
Fisher, Ronald A. 1935. The Design of Experiments. Oliver; Boyd.
Fokker, Adriaan D. 1914. “Die Mittlere Energie Rotierender Elektrischer Dipole Im Strahlungsfeld.” Annalen Der Physik 348 (5): 810–20.
Folland, Gerald B. 1999. Real Analysis: Modern Techniques and Their Applications. 2nd ed. Wiley.
Foret, Pierre, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021. “Sharpness-Aware Minimization for Efficiently Improving Generalization.” International Conference on Learning Representations. https://arxiv.org/abs/2010.01412.
Forrester, Alexander IJ, András Sóbester, and Andy J Keane. 2007. “Multi-Fidelity Optimization via Surrogate Modelling.” Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences 463 (2088): 3251–69. https://doi.org/10.1098/rspa.2007.1900.
Franceschi, L., M. Donini, P. Frasconi, and M. Pontil. 2017. “Forward and Reverse Gradient-Based Hyperparameter Optimization.” Proceedings of the 34th International Conference on Machine Learning (ICML’17). https://proceedings.mlr.press/v70/franceschi17a.html.
Frankle, Jonathan, and Michael Carbin. 2019. “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.” International Conference on Learning Representations. https://arxiv.org/abs/1803.03635.
Frazier, Peter I. 2018. “A Tutorial on Bayesian Optimization.” ArXiv:1807.02811. https://arxiv.org/abs/1807.02811.
Freund, Yoav, and Robert E Schapire. 1996. “Experiments with a New Boosting Algorithm.” Proceedings of the International Conference on Machine Learning 96: 148–56. https://dl.acm.org/doi/10.5555/3091696.3091715.
Friedman, Jerome H. 1987. “Exploratory Projection Pursuit.” Journal of the American Statistical Association 82 (397): 249–66. https://doi.org/10.1080/01621459.1987.10478427.
Frostig, Roy, Matthew James Johnson, and Chris Leary. 2018. “Compiling Machine Learning Programs via High-Level Tracing.” In Proceedings of Systems for Machine Learning. https://mlsys.org/Conferences/2018/doc/19.pdf.
Fukushima, Kunihiko. 1982. “Neocognitron: A Self-Organizing Neural Network Model for a Mechanism of Visual Pattern Recognition.” In Competition and Cooperation in Neural Nets. Springer. https://doi.org/10.1007/978-3-642-46466-9_18.
Gal, Yarin, and Zoubin Ghahramani. 2016. “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning.” International Conference on Machine Learning, 1050–59. https://arxiv.org/abs/1506.02142.
Gardner, Jacob, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and Andrew G Wilson. 2018. GPyTorch: Blackbox Matrix–Matrix Gaussian Process Inference with GPU Acceleration.” Advances in Neural Information Processing Systems 31. https://doi.org/10.5555/3327345.3327419.
Garg, Saurabh, Sivaraman Balakrishnan, Zico Kolter, and Zachary Lipton. 2021. RATT: Leveraging Unlabeled Data to Guarantee Generalization.” International Conference on Machine Learning, 3598–609. https://proceedings.mlr.press/v139/garg21a.html.
Gatys, Leon A, Alexander S Ecker, and Matthias Bethge. 2016. “Image Style Transfer Using Convolutional Neural Networks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2414–23. https://doi.org/10.1109/cvpr.2016.265.
Gauss, Carl Friedrich. 1809. “Theoria Motus Corporum Coelestium.” In Werke. Königlich Preussische Akademie der Wissenschaften. https://doi.org/10.1007/978-3-642-92478-1.
Gavish, Matan, and David L. Donoho. 2014. “The Optimal Hard Threshold for Singular Values Is 4/√3.” IEEE Transactions on Information Theory 60 (8): 5040–53.
Geman, Stuart, Elie Bienenstock, and René Doursat. 1992. “Neural Networks and the Bias/Variance Dilemma.” Neural Computation 4 (1): 1–58.
Gemma Team. 2025. “Gemma 3 Technical Report.” arXiv Preprint arXiv:2503.19786.
Gershgorin, Semyon A. 1931. Über Die Abgrenzung Der Eigenwerte Einer Matrix.” Izvestiya Akademii Nauk SSSR, Seriya Matematicheskaya, no. 6: 749–54.
Geshkovski, Borjan, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. 2023. “A Mathematical Perspective on Transformers.” arXiv Preprint arXiv:2312.10794.
Ghadimi, Saeed, and Guanghui Lan. 2013. “Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming.” SIAM Journal on Optimization 23 (4): 2341–68. https://arxiv.org/abs/1309.5549.
Gibbs, Josiah Willard. 1902. Elementary Principles of Statistical Mechanics. Scribner’s.
Ginibre, Jean. 1965. “Statistical Ensembles of Complex, Quaternion, and Real Matrices.” Journal of Mathematical Physics 6 (3): 440–49. https://doi.org/10.1063/1.1704292.
Girshick, Ross. 2015. “Fast R-CNN.” Proceedings of the IEEE International Conference on Computer Vision, 1440–48. https://doi.org/10.1109/iccv.2015.169.
Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014. “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 580–87. https://doi.org/10.18127/j00338486-202109-11.
Glorioso, Paolo, Quentin Anthony, Yury Tokpanov, Anna Golubeva, et al. 2024. “The Zamba2 Suite: Technical Report.” arXiv Preprint arXiv:2411.15242.
Glorioso, Paolo, Quentin Anthony, Yury Tokpanov, James Whittington, et al. 2024. “Zamba: A Compact 7B SSM Hybrid Model.” arXiv Preprint arXiv:2405.16712.
Glorot, Xavier, and Yoshua Bengio. 2010. “Understanding the Difficulty of Training Deep Feedforward Neural Networks.” Proceedings of the 13th International Conference on Artificial Intelligence and Statistics, 249–56. https://proceedings.mlr.press/v9/glorot10a.html.
Gneiting, Tilmann, and Adrian E. Raftery. 2007. “Strictly Proper Scoring Rules, Prediction, and Estimation.” Journal of the American Statistical Association 102 (477): 359–78. https://doi.org/10.1198/016214506000001437.
Godbole, Varun, George E. Dahl, Justin Gilmer, Christopher J. Shallue, and Zachary Nado. 2023. Deep Learning Tuning Playbook. Https://github.com/google-research/tuning_playbook.
Goh, Gabriel. 2017. “Why Momentum Really Works.” Distill. http://distill.pub/2017/momentum.
Goldberg, David. 1991. “What Every Computer Scientist Should Know about Floating-Point Arithmetic.” ACM Computing Surveys 23 (1): 5–48.
Goldberg, David, David Nichols, Brian M Oki, and Douglas Terry. 1992. “Using Collaborative Filtering to Weave an Information Tapestry.” Communications of the ACM 35 (12): 61–71. https://doi.org/10.1145/138859.138867.
Golub, Gene H, and Charles F Van Loan. 1996. Matrix Computations. Johns Hopkins University Press. https://jhupbooks.press.jhu.edu/title/matrix-computations.
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press.
Goodfellow, Ian, Jean Pouget-Abadie, Mehdi Mirza, et al. 2014. “Generative Adversarial Nets.” Advances in Neural Information Processing Systems, 2672–80. https://doi.org/10.5555/2969033.2969125.
Gotmare, Akhilesh, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018. “A Closer Look at Deep Learning Heuristics: Learning Rate Restarts, Warmup and Distillation.” ArXiv:1810.13243. https://arxiv.org/abs/1810.13243.
Goyal, Priya, Piotr Dollár, Ross Girshick, et al. 2017. “Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.” ArXiv:1706.02677. https://arxiv.org/abs/1706.02677.
Graham, Benjamin. 2014. “Fractional Max-Pooling.” ArXiv:1412.6071. https://arxiv.org/abs/1412.6071.
Grathwohl, Will, Ricky T. Q. Chen, Jesse Bettencourt, Ilya Sutskever, and David Duvenaud. 2019. FFJORD: Free-Form Continuous Dynamics for Scalable Reversible Generative Models.” International Conference on Learning Representations.
Grattafiori, Aaron, Abhimanyu Dubey, Abhinav Jauhri, et al. 2024. “The Llama 3 Herd of Models.” arXiv Preprint arXiv:2407.21783.
Graves, Alex. 2013. “Generating Sequences with Recurrent Neural Networks.” ArXiv:1308.0850. https://arxiv.org/abs/1308.0850.
Graves, Alex, and Jürgen Schmidhuber. 2005. “Framewise Phoneme Classification with Bidirectional LSTM and Other Neural Network Architectures.” Neural Networks 18 (5-6): 602–10. https://doi.org/10.1109/ijcnn.2005.1556215.
Grazzi, Riccardo, Julien Siems, Jörg K. H. Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil. 2025. “Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues.” International Conference on Learning Representations. https://arxiv.org/abs/2411.12537.
Gretton, Arthur, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. 2012. “A Kernel Two-Sample Test.” Journal of Machine Learning Research 13: 723–73.
Griewank, Andreas. 1992. “Achieving Logarithmic Growth of Temporal and Spatial Complexity in Reverse Automatic Differentiation.” Optimization Methods and Software 1 (1): 35–54.
Griewank, Andreas, and Andrea Walther. 2008. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation. 2nd ed. Society for Industrial; Applied Mathematics (SIAM). https://doi.org/10.1137/1.9780898717761.
Grinsztajn, Léo, Edouard Oyallon, and Gaël Varoquaux. 2022. “Why Do Tree-Based Models Still Outperform Deep Learning on Tabular Data?” Advances in Neural Information Processing Systems 35: 507–20. https://arxiv.org/abs/2207.08815.
Grobman, David M. 1959. “Homeomorphism of Systems of Differential Equations.” Doklady Akademii Nauk SSSR 128: 880–81.
Gu, Albert. 2023. “Modeling Sequences with Structured State Spaces.” PhD thesis, Stanford University.
Gu, Albert. 2025. On the Tradeoffs of SSMs and Transformers. Https://goombalab.github.io/blog/2025/tradeoffs/.
Gu, Albert, and Tri Dao. 2023. Mamba: Linear-Time Sequence Modeling with Selective State Spaces.” arXiv Preprint arXiv:2312.00752.
Gu, Albert, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. 2020. HiPPO: Recurrent Memory with Optimal Polynomial Projections.” Advances in Neural Information Processing Systems 33: 1474–87.
Gu, Albert, Karan Goel, and Christopher Ré. 2022. “Efficiently Modeling Long Sequences with Structured State Spaces.” International Conference on Learning Representations. https://arxiv.org/abs/2111.00396.
Gu, Albert, Ankit Gupta, Karan Goel, and Christopher Ré. 2022. “On the Parameterization and Initialization of Diagonal State Space Models.” Advances in Neural Information Processing Systems 35: 35971–83.
Gulati, Anmol, James Qin, Chung-Cheng Chiu, et al. 2020. “Conformer: Convolution-Augmented Transformer for Speech Recognition.” Proc. Interspeech 2020, 5036–40. https://doi.org/10.21437/interspeech.2020-3015.
Gulrajani, Ishaan, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville. 2017. “Improved Training of Wasserstein GANs.” Advances in Neural Information Processing Systems 30.
Gumbel, Emil J. 1954. Statistical Theory of Extreme Values and Some Practical Applications. U.S. National Bureau of Standards.
Gunawardana, Asela, and Guy Shani. 2015. “Evaluating Recommender Systems.” In Recommender Systems Handbook. Springer. https://doi.org/10.1007/978-1-4899-7637-6_8.
Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. “On Calibration of Modern Neural Networks.” International Conference on Machine Learning, 1321–30. https://arxiv.org/abs/1706.04599.
Guo, Huifeng, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. “DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction.” Proceedings of the 26th International Joint Conference on Artificial Intelligence, 1725–31. https://doi.org/10.24963/ijcai.2017/239.
Gupta, Suyog, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. “Deep Learning with Limited Numerical Precision.” International Conference on Machine Learning, 1737–46.
Gupta, Vineet, Tomer Koren, and Yoram Singer. 2018. “Shampoo: Preconditioned Stochastic Tensor Optimization.” Proceedings of the 35th International Conference on Machine Learning (ICML), 1842–50.
Guyon, Isabelle, Steve Gunn, Masoud Nikravesh, and Lotfi A Zadeh. 2008. Feature Extraction: Foundations and Applications. Springer. https://doi.org/10.1007/978-3-540-35488-8.
Hägele, Alexander, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro von Werra, and Martin Jaggi. 2024. “Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2405.18392.
Hairer, Ernst, Syvert P. Nørsett, and Gerhard Wanner. 1993. Solving Ordinary Differential Equations I: Nonstiff Problems. 2nd ed. Springer.
Halko, Nathan, Per-Gunnar Martinsson, and Joel A. Tropp. 2011. “Finding Structure with Randomness: Probabilistic Algorithms for Constructing Approximate Matrix Decompositions.” SIAM Review 53 (2): 217–88.
Hanson, Stephen José, and Lorien Y Pratt. 1988. “Comparing Biases for Minimal Network Construction with Back-Propagation.” Advances in Neural Information Processing Systems 1.
Harris, Charles R., K. Jarrod Millman, Stéfan J. van der Walt, et al. 2020. “Array Programming with NumPy.” Nature 585 (7825): 357–62. https://doi.org/10.1038/s41586-020-2649-2.
Hartley, Richard, and Andrew Zisserman. 2000. Multiple View Geometry in Computer Vision. Cambridge University Press. https://doi.org/10.1017/cbo9780511811685.
Hartman, Philip. 1960. “A Lemma in the Theory of Structural Stability of Differential Equations.” Proceedings of the American Mathematical Society 11 (4): 610–20.
Haveliwala, Taher H., and Sepandar D. Kamvar. 2003. The Second Eigenvalue of the Google Matrix. Stanford University. http://ilpubs.stanford.edu:8090/582/.
He, Horace. 2022. Making Deep Learning Go Brrrr from First Principles. Https://horace.io/brrr_intro.html.
He, Kaiming, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022. “Masked Autoencoders Are Scalable Vision Learners.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16000–16009. https://doi.org/10.1109/cvpr52688.2022.01553.
He, Kaiming, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017. “Mask R-CNN.” Proceedings of the IEEE International Conference on Computer Vision, 2961–69. https://doi.org/10.1109/iccv.2017.322.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.” Proceedings of the IEEE International Conference on Computer Vision, 1026–34. https://doi.org/10.1109/iccv.2015.123.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. “Deep Residual Learning for Image Recognition.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–78. https://doi.org/10.1109/cvpr.2016.90.
He, Tong, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li. 2019. “Bag of Tricks for Image Classification with Convolutional Neural Networks.” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 558–67. https://arxiv.org/abs/1812.01187.
He, Xiangnan, and Tat-Seng Chua. 2017. “Neural Factorization Machines for Sparse Predictive Analytics.” Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 355–64. https://doi.org/10.1145/3077136.3080777.
He, Xiangnan, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. “Neural Collaborative Filtering.” Proceedings of the 26th International Conference on World Wide Web, 173–82. https://doi.org/10.1145/3038912.3052569.
Hebb, Donald Olding. 1949. The Organization of Behavior. Wiley.
Held, Michael, Philip Wolfe, and Harlan P. Crowder. 1974. “Validation of Subgradient Optimization.” Mathematical Programming 6: 62–88.
Hendrycks, Dan, and Kevin Gimpel. 2016. “Gaussian Error Linear Units (GELUs).” ArXiv:1606.08415. https://arxiv.org/abs/1606.08415.
Hennessy, John L, and David A Patterson. 2011. Computer Architecture: A Quantitative Approach. Elsevier. https://www.elsevier.com/books/computer-architecture/hennessy/978-0-12-383872-8.
Henry, Alex, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. 2020. “Query-Key Normalization for Transformers.” Findings of the Association for Computational Linguistics: EMNLP 2020.
Herlocker, Jonathan L, Joseph A Konstan, Al Borchers, and John Riedl. 1999. “An Algorithmic Framework for Performing Collaborative Filtering.” 22nd Annual International ACM Conference on Research and Development in Information Retrieval, SIGIR 1999, 230–37. https://doi.org/10.1145/312624.312682.
Hidasi, Balázs, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2015. “Session-Based Recommendations with Recurrent Neural Networks.” ArXiv:1511.06939. https://arxiv.org/abs/1511.06939.
Higham, Nicholas J. 2002. Accuracy and Stability of Numerical Algorithms. 2nd ed. SIAM.
Hinton, Geoffrey, Oriol Vinyals, and Jeff Dean. 2015. “Distilling the Knowledge in a Neural Network.” arXiv Preprint arXiv:1503.02531.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising Diffusion Probabilistic Models.” Advances in Neural Information Processing Systems 33: 6840–51. https://arxiv.org/abs/2006.11239.
Ho, Jonathan, and Tim Salimans. 2022. “Classifier-Free Diffusion Guidance.” arXiv Preprint arXiv:2207.12598.
Hochreiter, Sepp, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber. 2001. “Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-Term Dependencies.” In A Field Guide to Dynamical Recurrent Neural Networks. IEEE Press. https://doi.org/10.18034/ei.v8i2.570.
Hochreiter, Sepp, and Jürgen Schmidhuber. 1997. “Long Short-Term Memory.” Neural Computation 9 (8): 1735–80. https://doi.org/10.1162/neco.1997.9.8.1735.
Hoeffding, Wassily. 1963. “Probability Inequalities for Sums of Bounded Random Variables.” Journal of the American Statistical Association 58 (301): 13–30.
Hoerl, Arthur E, and Robert W Kennard. 1970. “Ridge Regression: Biased Estimation for Nonorthogonal Problems.” Technometrics 12 (1): 55–67.
Hoffmann, Jordan, Sebastian Borgeaud, Arthur Mensch, et al. 2022. “Training Compute-Optimal Large Language Models.” ArXiv:2203.15556. https://arxiv.org/abs/2203.15556.
Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. “The Curious Case of Neural Text Degeneration.” International Conference on Learning Representations. https://arxiv.org/abs/1904.09751.
Horn, Roger A., and Charles R. Johnson. 2012. Matrix Analysis. 2nd ed. Cambridge University Press.
Hornik, Kurt. 1991. “Approximation Capabilities of Multilayer Feedforward Networks.” Neural Networks 4 (2): 251–57. https://doi.org/10.1016/0893-6080(91)90009-T.
Horowitz, Mark. 2014. “1.1 Computing’s Energy Problem (and What We Can Do about It).” IEEE International Solid-State Circuits Conference (ISSCC), 10–14. https://doi.org/10.1109/ISSCC.2014.6757323.
Howard, Andrew G., Menglong Zhu, Bo Chen, et al. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.” ArXiv:1704.04861. https://arxiv.org/abs/1704.04861.
Hu, Edward J., Yelong Shen, Phillip Wallis, et al. 2021. LoRA: Low-Rank Adaptation of Large Language Models.” arXiv Preprint arXiv:2106.09685.
Hu, Jie, Li Shen, and Gang Sun. 2018. “Squeeze-and-Excitation Networks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7132–41. https://doi.org/10.1109/cvpr.2018.00745.
Hu, Shengding, Yuge Tu, Xu Han, et al. 2024. “MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.” ArXiv:2404.06395. https://arxiv.org/abs/2404.06395.
Hu, Yifan, Yehuda Koren, and Chris Volinsky. 2008. “Collaborative Filtering for Implicit Feedback Datasets.” 2008 8th IEEE International Conference on Data Mining, 263–72. https://doi.org/10.1109/icdm.2008.22.
Hu, Zhiqiang, Roy Ka-Wei Lee, Charu C. Aggarwal, and Aston Zhang. 2022. “Text Style Transfer: A Review and Experimental Evaluation.” SIGKDD Explor. Newsl. 24 (1). https://doi.org/10.1145/3544903.3544906.
Huang, Cheng-Zhi Anna, Ashish Vaswani, Jakob Uszkoreit, et al. 2018. “Music Transformer: Generating Music with Long-Term Structure.” International Conference on Learning Representations. https://arxiv.org/abs/1809.04281.
Huang, Gao, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017. “Densely Connected Convolutional Networks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4700–4708. https://doi.org/10.1109/cvpr.2017.243.
Huang, Gao, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Weinberger. 2016. “Deep Networks with Stochastic Depth.” European Conference on Computer Vision. https://arxiv.org/abs/1603.09382.
Huang, Zhiheng, Wei Xu, and Kai Yu. 2015. “Bidirectional LSTMCRF Models for Sequence Tagging.” ArXiv:1508.01991. https://arxiv.org/abs/1508.01991.
Hubel, David H, and Torsten N Wiesel. 1959. “Receptive Fields of Single Neurones in the Cat’s Striate Cortex.” Journal of Physiology 148 (3): 574–91. https://doi.org/10.1113/jphysiol.1959.sp006308.
Hubel, David H, and Torsten N Wiesel. 1962. “Receptive Fields, Binocular Interaction and Functional Architecture in the Cat’s Visual Cortex.” Journal of Physiology 160 (1): 106–54. https://doi.org/10.1113/jphysiol.1962.sp006837.
Hubel, David H, and Torsten N Wiesel. 1968. “Receptive Fields and Functional Architecture of Monkey Striate Cortex.” Journal of Physiology 195 (1): 215–43. https://doi.org/10.1113/jphysiol.1968.sp008455.
Hutchinson, Michael F. 1989. “A Stochastic Estimator of the Trace of the Influence Matrix for Laplacian Smoothing Splines.” Communications in Statistics—Simulation and Computation 18 (3): 1059–76.
Hutter, F., H. Hoos, and K. Leyton-Brown. 2011. “Sequential Model-Based Optimization for General Algorithm Configuration.” Proceedings of the Fifth International Conference on Learning and Intelligent Optimization (LION’11). https://doi.org/10.1007/978-3-642-25566-3_40.
Hutter, F., L. Kotthoff, and J. Vanschoren, eds. 2019. Automated Machine Learning: Methods, Systems, Challenges. Springer. https://doi.org/10.1007/978-3-030-05318-5.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical Models by Score Matching.” Journal of Machine Learning Research 6: 695–709.
IBM Granite Team. 2025. IBM Granite 4.0: Hyper-Efficient, High-Performance Hybrid Models. Https://www.ibm.com/new/announcements/ibm-granite-4-0.
IEEE. 2019. IEEE Standard for Floating-Point Arithmetic. IEEE Std 754-2019.
Inan, Hakan, Khashayar Khosravi, and Richard Socher. 2017. “Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling.” International Conference on Learning Representations.
Ioffe, Sergey. 2017. “Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models.” Advances in Neural Information Processing Systems, 1945–53. https://proceedings.mlr.press/v70/ioffe17a.html.
Ioffe, Sergey, and Christian Szegedy. 2015. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.” International Conference on Machine Learning, 448–56. https://arxiv.org/abs/1502.03167.
Isensee, Fabian, Paul F. Jaeger, Simon A. A. Kohl, Jens Petersen, and Klaus H. Maier-Hein. 2021. “NnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation.” Nature Methods 18: 203–11. https://doi.org/10.1038/s41592-020-01008-z.
Izmailov, Pavel, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. 2018. “Averaging Weights Leads to Wider Optima and Better Generalization.” Uncertainty in Artificial Intelligence, 876–85. https://arxiv.org/abs/1803.05407.
Jacobs, Robert A., Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. 1991. “Adaptive Mixtures of Local Experts.” Neural Computation 3 (1): 79–87.
Jacot, Arthur, Franck Gabriel, and Clément Hongler. 2018. “Neural Tangent Kernel: Convergence and Generalization in Neural Networks.” Advances in Neural Information Processing Systems 31. https://doi.org/10.1145/3406325.3465355.
Jaeger, Herbert. 2002. Tutorial on Training Recurrent Neural Networks, Covering BPTT, RTRL, EKF and the “Echo State Network” Approach. GMD-Forschungszentrum Informationstechnik Bonn.
Jaegle, Andrew, Sebastian Borgeaud, Jean-Baptiste Alayrac, et al. 2022. “Perceiver IO: A General Architecture for Structured Inputs & Outputs.” International Conference on Learning Representations.
Jaegle, Andrew, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021. “Perceiver: General Perception with Iterative Attention.” International Conference on Machine Learning, 4651–64.
Jain, Sarthak, and Byron C. Wallace. 2019. “Attention Is Not Explanation.” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, 3543–56.
Jamieson, K., and A. Talwalkar. 2016. “Non-Stochastic Best Arm Identification and Hyperparameter Optimization.” Proceedings of the 19th International Conference on Artificial Intelligence and Statistics. https://proceedings.mlr.press/v51/jamieson16.html.
Jang, Eric, Shixiang Gu, and Ben Poole. 2017. “Categorical Reparameterization with Gumbel-Softmax.” International Conference on Learning Representations.
Jelassi, Samy, David Brandfonbrener, Sham M. Kakade, and Eran Malach. 2024. “Repeat After Me: Transformers Are Better Than State Space Models at Copying.” International Conference on Machine Learning, 21502–21. https://arxiv.org/abs/2402.01032.
Jelinek, Frederick, Robert L. Mercer, Lalit R. Bahl, and James K. Baker. 1977. “Perplexity—a Measure of the Difficulty of Speech Recognition Tasks.” Journal of the Acoustical Society of America 62 (S1): S63.
Jenatton, R., C. Archambeau, J. González, and M. Seeger. 2017. “Bayesian Optimization with tree-Structured dependencies.” Proceedings of the 34th International Conference on Machine Learning (ICML’17). https://proceedings.mlr.press/v70/jenatton17a.html.
Jia, Xianyan, Shutao Song, Wei He, et al. 2018. “Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes.” ArXiv:1807.11205. https://arxiv.org/abs/1807.11205.
Jia, Yangqing, Evan Shelhamer, Jeff Donahue, et al. 2014. “Caffe: Convolutional Architecture for Fast Feature Embedding.” Proceedings of the 22nd ACM International Conference on Multimedia, 675–78. https://doi.org/10.1145/2647868.2654889.
Jiang, Albert Q., Alexandre Sablayrolles, Arthur Mensch, et al. 2023. “Mistral 7B.” arXiv Preprint arXiv:2310.06825.
Jiang, Albert Q., Alexandre Sablayrolles, Antoine Roux, et al. 2024. “Mixtral of Experts.” arXiv Preprint arXiv:2401.04088.
Jordan, Keller, and contributors. 2024. Modded-Nanogpt: Speedrunning the NanoGPT Baseline. Https://github.com/KellerJordan/modded-nanogpt.
Jordan, Keller, Yuchen Jin, Vlado Boza, et al. 2024. Muon: An Optimizer for Hidden Layers in Neural Networks. Https://kellerjordan.github.io/posts/muon/.
Jordan, Richard, David Kinderlehrer, and Felix Otto. 1998. “The Variational Formulation of the Fokker–Planck Equation.” SIAM Journal on Mathematical Analysis 29 (1): 1–17.
Jouppi, Norman P, Cliff Young, Nishant Patil, et al. 2017. “In-Datacenter Performance Analysis of a Tensor Processing Unit.” 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), 1–12. https://doi.org/10.1145/3140659.3080246.
Kaddour, Jean. 2022. “Stop Wasting My Time! Saving Days of ImageNet and BERT Training with Latest Weight Averaging.” arXiv Preprint arXiv:2209.14981.
Kahan, William. 1965. “Further Remarks on Reducing Truncation Errors.” Communications of the ACM 8 (1): 40.
Kalchbrenner, Nal, Edward Grefenstette, and Phil Blunsom. 2014. “A Convolutional Neural Network for Modelling Sentences.” ArXiv:1404.2188. https://arxiv.org/abs/1404.2188.
Kalman, Rudolf E. 1960. “A New Approach to Linear Filtering and Prediction Problems.” Journal of Basic Engineering 82 (1): 35–45.
Kalra, Dayal Singh, and Maissam Barkeshli. 2024. “Why Warmup the Learning Rate? Underlying Mechanisms and Improvements.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2406.09405.
Kaplan, Jared, Sam McCandlish, Tom Henighan, et al. 2020. “Scaling Laws for Neural Language Models.” ArXiv:2001.08361. https://arxiv.org/abs/2001.08361.
Karimi, Hamed, Julie Nutini, and Mark Schmidt. 2016. “Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak–Łojasiewicz Condition.” Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD), 795–811.
Karnin, Z., T. Koren, and O. Somekh. 2013. “Almost Optimal Exploration in Multi-Armed Bandits.” Proceedings of the 30th International Conference on Machine Learning (ICML’13). https://proceedings.mlr.press/v28/karnin13.html.
Karras, Tero, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018. “Progressive Growing of GANs for Improved Quality, Stability, and Variation.” International Conference on Learning Representations. https://arxiv.org/abs/1710.10196.
Karras, Tero, Miika Aittala, Timo Aila, and Samuli Laine. 2022. “Elucidating the Design Space of Diffusion-Based Generative Models.” Advances in Neural Information Processing Systems 35.
Karras, Tero, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. 2024. “Analyzing and Improving the Training Dynamics of Diffusion Models.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
Karush, William. 1939. “Minima of Functions of Several Variables with Inequalities as Side Constraints.” Master’s thesis, Department of Mathematics, University of Chicago.
Kasimbeg, Priya, Frank Schneider, Runa Eschenhagen, et al. 2025. “Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition.” International Conference on Learning Representations. https://arxiv.org/abs/2502.15015.
Katharopoulos, Angelos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020. “Transformers Are RNNs: Fast Autoregressive Transformers with Linear Attention.” International Conference on Machine Learning.
Kazemnejad, Amirhossein, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023. “The Impact of Positional Encoding on Length Generalization in Transformers.” Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2305.19466.
Keskar, Nitish Shirish, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017. “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.” International Conference on Learning Representations. https://arxiv.org/abs/1609.04836.
Kidger, Patrick. 2022. “On Neural Differential Equations.” arXiv Preprint arXiv:2202.02435.
Kim, Jaeyoung, Mostafa El-Khamy, and Jungwon Lee. 2017. “Residual LSTM: Design of a Deep Recurrent Architecture for Distant Speech Recognition.” ArXiv:1701.03360. https://arxiv.org/abs/1701.03360.
Kim, Yoon. 2014. “Convolutional Neural Networks for Sentence Classification.” ArXiv:1408.5882. https://arxiv.org/abs/1408.5882.
Kimeldorf, G. S., and G. Wahba. 1971. “Some Results on Tchebycheffian Spline Functions.” J. Math. Anal. Appl. 33: 82–95. https://doi.org/10.1214/aoms/1177693054.
Kimi Team. 2025a. “Kimi K2: Open Agentic Intelligence.” arXiv Preprint arXiv:2507.20534.
Kimi Team. 2025b. “Kimi Linear: An Expressive, Efficient Attention Architecture.” arXiv Preprint arXiv:2510.26692.
Kingma, Diederik P, and Jimmy Ba. 2015. “Adam: A Method for Stochastic Optimization.” International Conference on Learning Representations. https://arxiv.org/abs/1412.6980.
Kingma, Diederik P., Tim Salimans, Ben Poole, and Jonathan Ho. 2021. “Variational Diffusion Models.” Advances in Neural Information Processing Systems 34.
Kingma, Diederik P., and Max Welling. 2014. “Auto-Encoding Variational Bayes.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/1312.6114.
Kipf, Thomas N, and Max Welling. 2017. “Semi-Supervised Classification with Graph Convolutional Networks.” International Conference on Learning Representations. https://arxiv.org/abs/1609.02907.
Kitaev, Nikita, Lukasz Kaiser, and Anselm Levskaya. 2020. “Reformer: The Efficient Transformer.” International Conference on Learning Representations.
Kloeden, Peter E., and Eckhard Platen. 1992. Numerical Solution of Stochastic Differential Equations. Springer.
Koh, Pang Wei, Shiori Sagawa, Henrik Marklund, et al. 2021. WILDS: A Benchmark of in-the-Wild Distribution Shifts.” International Conference on Machine Learning, 5637–64. https://arxiv.org/abs/2012.07421.
Kohavi, Ron. 1995. “A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection.” International Joint Conference on Artificial Intelligence 14: 1137–45.
Koller, Daphne, and Nir Friedman. 2009. Probabilistic Graphical Models: Principles and Techniques. MIT Press. https://doi.org/10.7551/mitpress/7432.001.0001.
Kolmogorov, Andrey. 1933. “Sulla Determinazione Empirica Di Una Legge Di Distribuzione.” Inst. Ital. Attuari, Giorn. 4: 83–91. https://doi.org/10.1007/BF03017337.
Kolmogorov, Andrey N. 1931. Über Die Analytischen Methoden in Der Wahrscheinlichkeitsrechnung.” Mathematische Annalen 104: 415–58.
Kolter, Zico. 2008. “Linear Algebra Review and Reference.” Available Online: Http://Cs229.stanford.edu/Section/Cs229-Linalg.pdf. http://cs229.stanford.edu/section/cs229-linalg.pdf.
Koren, Yehuda, Robert Bell, and Chris Volinsky. 2009. “Matrix Factorization Techniques for Recommender Systems.” Computer 42 (8): 30–37. https://doi.org/10.1109/mc.2009.263.
Kosson, Atli, Bettina Messmer, and Martin Jaggi. 2024. “Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.” International Conference on Machine Learning. https://arxiv.org/abs/2305.17212.
Kosson, Atli, Jeremy Welborn, Yang Liu, Martin Jaggi, and Xi Chen. 2025. “Weight Decay May Matter More Than muP for Learning Rate Transfer in Practice.” arXiv Preprint arXiv:2510.19093.
Kraft, Leon G. 1949. “A Device for Quantizing, Grouping, and Coding Amplitude-Modulated Pulses.” Master’s thesis, Massachusetts Institute of Technology.
Kraskov, Alexander, Harald Stögbauer, and Peter Grassberger. 2004. “Estimating Mutual Information.” Physical Review E 69 (6): 066138.
Krizhevsky, Alex, Ilya Sutskever, and Geoffrey E Hinton. 2012. “ImageNet Classification with Deep Convolutional Neural Networks.” Advances in Neural Information Processing Systems, 1097–105. https://doi.org/10.5555/2999134.2999257.
Krogh, Anders, and John A Hertz. 1992. “A Simple Weight Decay Can Improve Generalization.” Advances in Neural Information Processing Systems, 950–57. https://doi.org/10.5555/2986916.2987033.
Kuhn, Harold W., and Albert W. Tucker. 1951. “Nonlinear Programming.” Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, 481–92.
Kullback, Solomon, and Richard A. Leibler. 1951. “On Information and Sufficiency.” Annals of Mathematical Statistics 22 (1): 79–86.
Kung, Sun Yuan. 1988. VLSI Array Processors.” Prentice Hall.
Kunstner, Frederik, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt. 2023. “Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.” International Conference on Learning Representations. https://arxiv.org/abs/2304.13960.
Kunstner, Frederik, Robin Yadav, Alan Milligan, Mark Schmidt, and Alberto Bietti. 2024. “Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2402.19449.
Kuzovkin, Ilya, Raul Vicente, Mathilde Petton, et al. 2018. “Activations of Deep Convolutional Neural Networks Are Aligned with Gamma Band Activity of Human Visual Cortex.” Communications Biology 1 (1): 1–12. https://doi.org/10.1038/s42003-018-0110-y.
Lahoti, Aakash, Kevin Y. Li, Berlin Chen, et al. 2026. “Mamba-3: Improved Sequence Modeling Using State Space Principles.” arXiv Preprint arXiv:2603.15569.
Laplace, Pierre-Simon. 1814. Essai Philosophique Sur Les Probabilités. Courcier.
Lavin, Andrew, and Scott Gray. 2016. “Fast Algorithms for Convolutional Neural Networks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4013–21. https://doi.org/10.1109/cvpr.2016.435.
Le, Quoc V. 2013. “Building High-Level Features Using Large Scale Unsupervised Learning.” Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, 8595–98. https://doi.org/10.1109/icassp.2013.6639343.
LeCun, Yann, Yoshua Bengio, and et al. 1995. “Convolutional Networks for Images, Speech, and Time Series.” In The Handbook of Brain Theory and Neural Networks. MIT Press. http://yann.lecun.com/exdb/publis/pdf/lecun-bengio-95a.pdf.
LeCun, Yann, Bernhard Boser, John S Denker, et al. 1989. “Backpropagation Applied to Handwritten Zip Code Recognition.” Neural Computation 1 (4): 541–51. https://doi.org/10.1162/neco.1989.1.4.541.
LeCun, Yann, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. “Gradient-Based Learning Applied to Document Recognition.” Proceedings of the IEEE 86 (11): 2278–324. https://doi.org/10.1109/5.726791.
LeCun, Yann, Leon Bottou, G Orr, and Klaus-Robert Muller. 1998. “Efficient Backprop.” In Neural Networks: Tricks of the Trade. Springer. https://doi.org/10.1007/3-540-49430-8_2.
LeCun, Yann, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu Jie Huang. 2006. “A Tutorial on Energy-Based Learning.” In Predicting Structured Data. MIT Press.
LeCun, Yann, LD Jackel, Leon Bottou, et al. 1995. “Comparison of Learning Algorithms for Handwritten Digit Recognition.” International Conference on Artificial Neural Networks, 53–60.
Lee, Hyunji, Wenhao Yu, Hongming Zhang, et al. 2025. “Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling.” arXiv Preprint arXiv:2510.26912.
Legendre, Adrien Marie. 1805. Mémoire Sur Les Opérations Trigonométriques: Dont Les Résultats Dépendent de La Figure de La Terre. F. Didot.
Lepikhin, Dmitry, HyoukJoong Lee, Yuanzhong Xu, et al. 2021. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.” International Conference on Learning Representations.
Leshno, Moshe, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken. 1993. “Multilayer Feedforward Networks with a Nonpolynomial Activation Function Can Approximate Any Function.” Neural Networks 6 (6): 861–67.
Lessard, Laurent, Benjamin Recht, and Andrew Packard. 2016. “Analysis and Design of Optimization Algorithms via Integral Quadratic Constraints.” SIAM Journal on Optimization 26 (1): 57–95.
Leviathan, Yaniv, Matan Kalman, and Yossi Matias. 2023. “Fast Inference from Transformers via Speculative Decoding.” International Conference on Machine Learning, 19274–86.
Levy, Omer, and Yoav Goldberg. 2014. “Neural Word Embedding as Implicit Matrix Factorization.” Advances in Neural Information Processing Systems 27.
Lewis, Mike, Yinhan Liu, Naman Goyal, et al. 2019. BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension.” ArXiv:1910.13461. https://arxiv.org/abs/1910.13461.
Li, Chun-Liang, Wei-Cheng Chang, Yu Cheng, Yiming Yang, and Barnabás Póczos. 2017. MMD GAN: Towards Deeper Understanding of Moment Matching Network.” Advances in Neural Information Processing Systems 30.
Li, Junnan, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models.” International Conference on Machine Learning, 19730–42.
Li, L., K. Jamieson, A. Rostamizadeh, et al. 2018. “Massively Parallel Hyperparameter Tuning.” ArXiv:1810.05934. https://arxiv.org/abs/1810.05934.
Li, Mu. 2017. “Scaling Distributed Machine Learning with System and Algorithm Co-Design.” PhD thesis, PhD Thesis, CMU. https://www.cs.cmu.edu/~muli/file/mu-thesis.pdf.
Li, Mu, David G Andersen, Jun Woo Park, et al. 2014. “Scaling Distributed Machine Learning with the Parameter Server.” 11th Symposium on Operating Systems Design and Implementation (OSDI 14), 583–98. https://doi.org/10.1145/2640087.2644155.
Li, Mu, Tong Zhang, Yuqiang Chen, and Alexander J Smola. 2014. “Efficient Mini-Batch Training for Stochastic Optimization.” Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 661–70. https://doi.org/10.1145/2623330.2623612.
Li, Shen, Yanli Zhao, Rohan Varma, et al. 2020. PyTorch Distributed: Experiences on Accelerating Data Parallel Training.” Proceedings of the VLDB Endowment 13 (12): 3005–18. https://doi.org/10.14778/3415478.3415530.
Li, Xiang, Shuo Chen, Xiaolin Hu, and Jian Yang. 2019. “Understanding the Disharmony Between Dropout and Batch Normalization by Variance Shift.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2682–90. https://arxiv.org/abs/1801.05134.
Li, Yujia, Kevin Swersky, and Richard Zemel. 2015. “Generative Moment Matching Networks.” Proceedings of the 32nd International Conference on Machine Learning, 1718–27.
Liaw, R., E. Liang, R. Nishihara, P. Moritz, J. Gonzalez, and I. Stoica. 2018. Tune: A Research Platform for Distributed Model Selection and Training.” ArXiv:1807.05118. https://arxiv.org/abs/1807.05118.
Lieber, Opher, Barak Lenz, Hofit Bata, et al. 2024. Jamba: A Hybrid Transformer-Mamba Language Model.” arXiv Preprint arXiv:2403.19887.
Lin, Jianhua. 1991. “Divergence Measures Based on the Shannon Entropy.” IEEE Transactions on Information Theory 37 (1): 145–51.
Lin, Min, Qiang Chen, and Shuicheng Yan. 2013. “Network in Network.” ArXiv:1312.4400. https://arxiv.org/abs/1312.4400.
Lin, Tsung-Yi, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017. “Focal Loss for Dense Object Detection.” Proceedings of the IEEE International Conference on Computer Vision, 2980–88. https://doi.org/10.1109/iccv.2017.324.
Lin, Yuanqing, F Lv, S Zhu, et al. 2010. “ImageNet Classification: Fast Descriptor Coding and Large-Scale SVM Training.” Large Scale Visual Recognition Challenge, ahead of print. https://doi.org/10.1109/cvpr.2010.5539970.
Lin, Zhouhan, Minwei Feng, Cicero Nogueira dos Santos, et al. 2017. “A Structured Self-Attentive Sentence Embedding.” ArXiv:1703.03130. https://arxiv.org/abs/1703.03130.
Linnainmaa, Seppo. 1970. “The Representation of the Cumulative Rounding Error of an Algorithm as a Taylor Expansion of the Local Rounding Errors.” Master’s thesis, University of Helsinki.
Lipman, Yaron, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2023. “Flow Matching for Generative Modeling.” International Conference on Learning Representations.
Lipton, Zachary C, and Jacob Steinhardt. 2018. “Troubling Trends in Machine Learning Scholarship.” Communications of the ACM 17: 45–77. https://doi.org/10.1145/3317287.3328534.
Lipton, Zachary, Yu-Xiang Wang, and Alexander Smola. 2018. “Detecting and Correcting for Label Shift with Black Box Predictors.” International Conference on Machine Learning, 3122–30. https://arxiv.org/abs/1802.03916.
Liu, Bo, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, and Qiang Liu. 2024. “Longhorn: State Space Models Are Amortized Online Learners.” arXiv Preprint arXiv:2407.14207.
Liu, Chaoyue, Libin Zhu, and Mikhail Belkin. 2022. “Loss Landscapes and Optimization in over-Parameterized Non-Linear Systems and Neural Networks.” Applied and Computational Harmonic Analysis 59: 85–116.
Liu, Dong C, and Jorge Nocedal. 1989. “On the Limited Memory BFGS Method for Large Scale Optimization.” Mathematical Programming 45 (1): 503–28. https://doi.org/10.1007/bf01589116.
Liu, Hanxiao, Karen Simonyan, and Yiming Yang. 2018. DARTS: Differentiable Architecture Search.” ArXiv:1806.09055. https://arxiv.org/abs/1806.09055.
Liu, Hong, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. 2024. “Sophia: A Scalable Stochastic Second-Order Optimizer for Language Model Pre-Training.” International Conference on Learning Representations. https://arxiv.org/abs/2305.14342.
Liu, Jingyuan, Jianlin Su, Xingcheng Yao, et al. 2025. “Muon Is Scalable for LLM Training.” arXiv Preprint arXiv:2502.16982.
Liu, Qiang, Jason D. Lee, and Michael I. Jordan. 2016. “A Kernelized Stein Discrepancy for Goodness-of-Fit Tests.” Proceedings of the 33rd International Conference on Machine Learning, 276–84.
Liu, Qiang, and Dilin Wang. 2016. “Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm.” Advances in Neural Information Processing Systems 29.
Liu, Shiwei, Tianlong Chen, Xiaohan Chen, et al. 2022. “More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 Using Sparsity.” ArXiv:2207.03620. https://arxiv.org/abs/2207.03620.
Liu, Wei, Dragomir Anguelov, Dumitru Erhan, et al. 2016. SSD: Single Shot Multibox Detector.” European Conference on Computer Vision, 21–37. https://doi.org/10.1007/978-3-319-46448-0_2.
Liu, Xingchao, Chengyue Gong, and Qiang Liu. 2023. “Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.” International Conference on Learning Representations.
Liu, Yinhan, Myle Ott, Naman Goyal, et al. 2019. “RoBERTa: A Robustly Optimized BERT Pretraining Approach.” ArXiv:1907.11692. https://arxiv.org/abs/1907.11692.
Liu, Ze, Yutong Lin, Yue Cao, et al. 2021. “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows.” Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–22. https://doi.org/10.1109/iccv48922.2021.00986.
Liu, Zhuang, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. “A ConvNet for the 2020s.” ArXiv:2201.03545. https://arxiv.org/abs/2201.03545.
Łojasiewicz, Stanisław. 1963. “Une Propriété Topologique Des Sous-Ensembles Analytiques réels.” In Les Équations Aux dérivées Partielles. Éditions du CNRS.
Long, Jonathan, Evan Shelhamer, and Trevor Darrell. 2015. “Fully Convolutional Networks for Semantic Segmentation.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3431–40. https://doi.org/10.1109/tpami.2016.2572683.
Loshchilov, Ilya, and Frank Hutter. 2016. SGDR: Stochastic Gradient Descent with Warm Restarts.” ArXiv:1608.03983. https://arxiv.org/abs/1608.03983.
Loshchilov, Ilya, and Frank Hutter. 2019. “Decoupled Weight Decay Regularization.” International Conference on Learning Representations. https://arxiv.org/abs/1711.05101.
Lowe, David G. 2004. “Distinctive Image Features from Scale-Invariant Keypoints.” International Journal of Computer Vision 60 (2): 91–110. https://doi.org/10.1023/b:visi.0000029664.99615.94.
Luo, Calvin. 2022. “Understanding Diffusion Models: A Unified Perspective.” arXiv Preprint arXiv:2208.11970.
Luo, Ping, Xinjiang Wang, Wenqi Shao, and Zhanglin Peng. 2018. “Towards Understanding Regularization in Batch Normalization.” ArXiv:1809.00846. https://arxiv.org/abs/1809.00846.
Luo, Wenjie, Yujia Li, Raquel Urtasun, and Richard Zemel. 2016. “Understanding the Effective Receptive Field in Deep Convolutional Neural Networks.” Advances in Neural Information Processing Systems. https://arxiv.org/abs/1701.04128.
Maas, Andrew L, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. “Learning Word Vectors for Sentiment Analysis.” Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Volume 1, 142–50. https://aclanthology.org/P11-1015.
Mack, Yue-Pok, and Bernard W Silverman. 1982. “Weak and Strong Uniform Consistency of Kernel Regression Estimates.” Zeitschrift für Wahrscheinlichkeitstheorie Und Verwandte Gebiete 61 (3): 405–15. https://doi.org/10.1007/bf00539840.
MacKay, David JC. 2003. Information Theory, Inference and Learning Algorithms. Cambridge University Press. https://www.inference.org.uk/mackay/itila/book.html.
Maclaurin, D., D. Duvenaud, and R. Adams. 2015. “Gradient-Based Hyperparameter Optimization Through Reversible Learning.” Proceedings of the 32nd International Conference on Machine Learning (ICML’15). https://proceedings.mlr.press/v37/maclaurin15.html.
Maddison, Chris J., Andriy Mnih, and Yee Whye Teh. 2017. “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables.” International Conference on Learning Representations.
Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. “Towards Deep Learning Models Resistant to Adversarial Attacks.” International Conference on Learning Representations. https://arxiv.org/abs/1706.06083.
Malladi, Sadhika, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora. 2022. “On the SDEs and Scaling Rules for Adaptive Gradient Algorithms.” Advances in Neural Information Processing Systems 35. https://arxiv.org/abs/2205.10287.
Mangram, Myles E. 2013. “A Simplified Perspective of the Markowitz Portfolio Theory.” Global Journal of Business Research 7 (1): 59–70. https://scholarworks.iu.edu/journals/index.php/jiuspa/article/view/4517.
Manning, Christopher D., Prabhakar Raghavan, and Hinrich Schütze. 2008. Introduction to Information Retrieval. Cambridge University Press. https://doi.org/10.1017/CBO9780511809071.
Marchenko, Vladimir A., and Leonid A. Pastur. 1967. “Distribution of Eigenvalues for Some Sets of Random Matrices.” Mathematics of the USSR-Sbornik 1 (4): 457–83.
Maron, Melvin E. 1961. “Automatic Indexing: An Experimental Inquiry.” Journal of the ACM 8 (3): 404–17.
Martens, James, and Roger Grosse. 2015. “Optimizing Neural Networks with Kronecker-Factored Approximate Curvature.” Proceedings of the 32nd International Conference on Machine Learning (ICML), 2408–17.
Martin, Charles H., and Michael W. Mahoney. 2021. “Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning.” Journal of Machine Learning Research 22 (165): 1–73.
Martins, André F. T., and Ramón F. Astudillo. 2016. “From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification.” Proceedings of the 33rd International Conference on Machine Learning (ICML), 1614–23.
Maruyama, Gisiro. 1955. “Continuous Markov Processes and Stochastic Equations.” Rendiconti Del Circolo Matematico Di Palermo 4: 48–90.
Matthews, Alexander G de G, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani. 2018. “Gaussian Process Behaviour in Wide Deep Neural Networks.” ArXiv:1804.11271. https://arxiv.org/abs/1804.11271.
McAllester, David, and Karl Stratos. 2020. “Formal Limitations on the Measurement of Mutual Information.” International Conference on Artificial Intelligence and Statistics, 875–84.
McCandlish, Sam, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. 2018. “An Empirical Model of Large-Batch Training.” arXiv Preprint arXiv:1812.06162.
McCann, Bryan, James Bradbury, Caiming Xiong, and Richard Socher. 2017. “Learned in Translation: Contextualized Word Vectors.” Advances in Neural Information Processing Systems, 6294–305. https://doi.org/10.5555/3294996.3295037.
McCulloch, Warren S, and Walter Pitts. 1943. “A Logical Calculus of the Ideas Immanent in Nervous Activity.” Bulletin of Mathematical Biophysics 5 (4): 115–33. https://doi.org/10.1016/s0092-8240(05)80006-0.
McDiarmid, Colin. 1989. “On the Method of Bounded Differences.” In Surveys in Combinatorics. Cambridge University Press.
McKinney, Wes. 2010. “Data Structures for Statistical Computing in Python.” Proceedings of the 9th Python in Science Conference, 56–61. https://doi.org/10.25080/Majora-92bf1922-00a.
McMahan, H Brendan, Gary Holt, David Sculley, et al. 2013. “Ad Click Prediction: A View from the Trenches.” Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1222–30. https://doi.org/10.1145/2487575.2488200.
McMillan, Brockway. 1956. “Two Inequalities Implied by Unique Decipherability.” IRE Transactions on Information Theory 2 (4): 115–16.
Mead, Carver, and Lynn Conway. 1980. Introduction to VLSI Systems. Addison-Wesley.
Meng, Fanxu, Zhaohui Wang, and Muhan Zhang. 2024. PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models.” Advances in Neural Information Processing Systems.
Merity, Stephen, Caiming Xiong, James Bradbury, and Richard Socher. 2016. “Pointer Sentinel Mixture Models.” ArXiv:1609.07843. https://arxiv.org/abs/1609.07843.
Merrill, William, Shane Arora, Dirk Groeneveld, and Hannaneh Hajishirzi. 2025. “Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training.” arXiv Preprint arXiv:2505.23971.
Merrill, William, Jackson Petty, and Ashish Sabharwal. 2024. “The Illusion of State in State-Space Models.” International Conference on Machine Learning. https://arxiv.org/abs/2404.08819.
Meta AI. 2025. The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation. Https://ai.meta.com/blog/llama-4-multimodal-intelligence/.
Metropolis, Nicholas, and Stanislaw Ulam. 1949. “The Monte Carlo Method.” Journal of the American Statistical Association 44 (247): 335–41.
Micchelli, Charles A. 1984. “Interpolation of Scattered Data: Distance Matrices and Conditionally Positive Definite Functions.” In Approximation Theory and Spline Functions. Springer. https://doi.org/10.1007/bf01893414.
Micikevicius, Paulius, Sharan Narang, Jonah Alben, et al. 2018. “Mixed Precision Training.” International Conference on Learning Representations.
Micikevicius, Paulius, Dusan Stosic, Neil Burgess, et al. 2022. FP8 Formats for Deep Learning.” arXiv Preprint arXiv:2209.05433.
Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. “Efficient Estimation of Word Representations in Vector Space.” ArXiv:1301.3781. https://arxiv.org/abs/1301.3781.
Mikolov, Tomas, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. “Distributed Representations of Words and Phrases and Their Compositionality.” Advances in Neural Information Processing Systems, 3111–19. https://doi.org/10.5555/2999792.2999959.
Milakov, Maxim, and Natalia Gimelshein. 2018. “Online Normalizer Calculation for Softmax.” arXiv Preprint arXiv:1805.02867.
Miller, George A. 1995. “WordNet: A Lexical Database for English.” Communications of the ACM 38 (11): 39–41. https://doi.org/10.1145/219717.219748.
Milstein, Grigori N. 1975. “Approximate Integration of Stochastic Differential Equations.” Theory of Probability and Its Applications 19 (3): 557–62.
MiniMax. 2025. MiniMax-01: Scaling Foundation Models with Lightning Attention.” arXiv Preprint arXiv:2501.08313.
Mironov, Ilya. 2017. “Rényi Differential Privacy.” 2017 IEEE 30th Computer Security Foundations Symposium (CSF), 263–75.
Mirsky, Leon. 1960. “Symmetric Gauge Functions and Unitarily Invariant Norms.” The Quarterly Journal of Mathematics 11 (1): 50–59.
Miyato, Takeru, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018. “Spectral Normalization for Generative Adversarial Networks.” International Conference on Learning Representations.
Mnih, Volodymyr, Nicolas Heess, Alex Graves, et al. 2014. “Recurrent Models of Visual Attention.” Advances in Neural Information Processing Systems, 2204–12. https://doi.org/10.5555/2969033.2969073.
Mnih, Volodymyr, Koray Kavukcuoglu, David Silver, et al. 2013. “Playing Atari with Deep Reinforcement Learning.” ArXiv:1312.5602.
Mnih, Volodymyr, Koray Kavukcuoglu, David Silver, et al. 2015. “Human-Level Control Through Deep Reinforcement Learning.” Nature 518 (7540): 529–33. https://doi.org/10.1038/nature14236.
Montúfar, Guido, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio. 2014. “On the Number of Linear Regions of Deep Neural Networks.” Advances in Neural Information Processing Systems 27.
Moon, Taesup, Alex Smola, Yi Chang, and Zhaohui Zheng. 2010. “Intervalrank: Isotonic Regression with Listwise and Pairwise Constraints.” Proceedings of the 3rd ACM International Conference on Web Search and Data Mining, 151–60. https://doi.org/10.1145/1718487.1718520.
Morozov, Vladimir Alekseevich. 1984. Methods for Solving Incorrectly Posed Problems. Springer.
Muennighoff, Niklas, Alexander M. Rush, Boaz Barak, et al. 2023. “Scaling Data-Constrained Language Models.” Advances in Neural Information Processing Systems.
Müller, Alfred. 1997. “Integral Probability Metrics and Their Generating Classes of Functions.” Advances in Applied Probability 29 (2): 429–43.
Murphy, Kevin P. 2022. Probabilistic Machine Learning: An Introduction. MIT Press. https://probml.github.io/pml-book/book1.html.
Nadaraya, Elizbar A. 1964. “On Estimating Regression.” Theory of Probability & Its Applications 9 (1): 141–42. https://doi.org/10.1137/1109020.
Nado, Zachary, Justin M. Gilmer, Christopher J. Shallue, Rohan Anil, and George E. Dahl. 2021. “A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes.” arXiv Preprint arXiv:2102.06356.
Nair, Vinod, and Geoffrey E Hinton. 2010. “Rectified Linear Units Improve Restricted Boltzmann Machines.” ICML, 807–14. https://dl.acm.org/doi/10.5555/3104322.3104425.
Nakkiran, Preetum, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. 2021. “Deep Double Descent: Where Bigger Models and More Data Hurt.” Journal of Statistical Mechanics: Theory and Experiment 2021 (12): 124003. https://doi.org/10.1088/1742-5468/ac3a74.
Naor, Moni, and Omer Reingold. 1999. “On the Construction of Pseudorandom Permutations: Luby–Rackoff Revisited.” Journal of Cryptology 12 (1): 29–66. https://doi.org/10.1007/s001459900037.
Naumann, Uwe. 2008. “Optimal Jacobian Accumulation Is NP-Complete.” Mathematical Programming 112 (2): 427–41.
Neal, Radford M. 1996. Bayesian Learning for Neural Networks. Springer. https://doi.org/10.1007/978-1-4612-0745-0.
Nemirovski, Arkadi, and David Yudin. 1983. Problem Complexity and Method Efficiency in Optimization. Wiley.
Nesterov, Yu. 2018. Lectures on Convex Optimization. Springer. https://doi.org/10.1007/978-3-319-91578-4.
Nesterov, Yurii. 1983. “A Method of Solving a Convex Programming Problem with Convergence Rate O(1/k2).” Soviet Mathematics Doklady 27 (2): 372–76.
Neyman, Jerzy. 1937. “Outline of a Theory of Statistical Estimation Based on the Classical Theory of Probability.” Philosophical Transactions of the Royal Society of London. Series A, Mathematical and Physical Sciences 236 (767): 333–80. https://doi.org/10.2307/jj.8501421.24.
Ng, Andrew Y., and Michael I. Jordan. 2002. “On Discriminative Vs. Generative Classifiers: A Comparison of Logistic Regression and Naive Bayes.” Advances in Neural Information Processing Systems 14: 841–48.
Nguyen, Minh Nhat, Andrew Baker, Clement Neo, Allen Roush, Andreas Kirsch, and Ravid Shwartz-Ziv. 2025. “Turning up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs.” International Conference on Learning Representations. https://arxiv.org/abs/2407.01082.
Nguyen, XuanLong, Martin J. Wainwright, and Michael I. Jordan. 2010. “Estimating Divergence Functionals and the Likelihood Ratio by Convex Risk Minimization.” IEEE Transactions on Information Theory 56 (11): 5847–61.
Niculae, Vlad, and Mathieu Blondel. 2017. “A Regularized Framework for Sparse and Structured Neural Attention.” Advances in Neural Information Processing Systems 30.
Nocedal, Jorge, and Stephen J. Wright. 2006. Numerical Optimization. 2nd ed. Springer.
Norelli, Antonio, Marco Fumero, Valentino Maiorca, Luca Moschella, Emanuele Rodolà, and Francesco Locatello. 2022. ASIF: Coupled Data Turns Unimodal Models to Multimodal Without Training.” ArXiv:2210.01738. https://arxiv.org/abs/2210.01738.
Novak, Roman, Lechao Xiao, Jaehoon Lee, et al. 2018. “Bayesian Deep Convolutional Networks with Many Channels Are Gaussian Processes.” ArXiv:1810.05148. https://arxiv.org/abs/1810.05148.
Novikoff, A. B. J. 1962. “On Convergence Proofs for Perceptrons.” Proceedings of the Symposium on the Mathematical Theory of Automata, 615–22. https://cs.nyu.edu/~mohri/pub/nov62.pdf.
Nowozin, Sebastian, Botond Cseke, and Ryota Tomioka. 2016. “F-GAN: Training Generative Neural Samplers Using Variational Divergence Minimization.” Advances in Neural Information Processing Systems 29.
NVIDIA. 2025. “Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.” arXiv Preprint arXiv:2504.03624.
Øksendal, Bernt. 2003. Stochastic Differential Equations: An Introduction with Applications. 6th ed. Springer.
Olshausen, Bruno A, and David J Field. 1996. “Emergence of Simple-Cell Receptive Field Properties by Learning a Sparse Code for Natural Images.” Nature 381 (6583): 607–9. https://doi.org/10.1038/381607a0.
Olsson, Catherine, Nelson Elhage, Neel Nanda, et al. 2022. “In-Context Learning and Induction Heads.” Transformer Circuits Thread.
Ong, Cheng Soon, Alexander Smola, and Robert Williamson. 2005. “Learning the Kernel with Hyperkernels.” Journal of Machine Learning Research 6: 1043–71. https://doi.org/10.1109/jcss.2005.25.2.
Oord, Aaron van den, Yazhe Li, and Oriol Vinyals. 2018. “Representation Learning with Contrastive Predictive Coding.” arXiv Preprint arXiv:1807.03748.
OpenAI. 2023. GPT-4 Technical Report.” ArXiv:2303.08774. https://arxiv.org/abs/2303.08774.
OpenAI. 2025. “Gpt-Oss-120b & Gpt-Oss-20b Model Card.” arXiv Preprint arXiv:2508.10925.
Orvieto, Antonio, Samuel L. Smith, Albert Gu, et al. 2023. “Resurrecting Recurrent Neural Networks for Long Sequences.” International Conference on Machine Learning, 26670–98. https://arxiv.org/abs/2303.06349.
Ouyang, Long, Jeff Wu, Xu Jiang, et al. 2022. “Training Language Models to Follow Instructions with Human Feedback.” ArXiv:2203.02155. https://arxiv.org/abs/2203.02155.
Paley, Raymond E. A. C., Norbert Wiener, and Antoni Zygmund. 1933. “Notes on Random Functions.” Mathematische Zeitschrift 37: 647–68.
Papineni, Kishore, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation.” Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311–18. https://doi.org/10.3115/1073083.1073135.
Parikh, Ankur P, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016. “A Decomposable Attention Model for Natural Language Inference.” ArXiv:1606.01933. https://arxiv.org/abs/1606.01933.
Park, Taesung, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019. “Semantic Image Synthesis with Spatially-Adaptive Normalization.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2337–46. https://doi.org/10.1109/cvpr.2019.00244.
Parr, Terence, and Jeremy Howard. 2018. “The Matrix Calculus You Need for Deep Learning.” arXiv Preprint arXiv:1802.01528.
Parzen, Emanuel. 1957. “On Consistent Estimates of the Spectrum of a Stationary Time Series.” Annals of Mathematical Statistics 28: 329–48. https://doi.org/10.1214/aoms/1177706962.
Pascanu, Razvan, Tomas Mikolov, and Yoshua Bengio. 2013. “On the Difficulty of Training Recurrent Neural Networks.” International Conference on Machine Learning, 1310–18.
Paszke, Adam, Sam Gross, Francisco Massa, et al. 2019. “PyTorch: An Imperative Style, High-Performance Deep Learning Library.” Advances in Neural Information Processing Systems 32: 8026–37. https://proceedings.neurips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html.
Patarasuk, Pitch, and Xin Yuan. 2009. “Bandwidth Optimal All-Reduce Algorithms for Clusters of Workstations.” Journal of Parallel and Distributed Computing 69 (2): 117–24. https://doi.org/10.1016/j.jpdc.2008.09.002.
Pearlmutter, Barak A. 1994. “Fast Exact Multiplication by the Hessian.” Neural Computation 6 (1): 147–60. https://doi.org/10.1162/neco.1994.6.1.147.
Peebles, William, and Saining Xie. 2023. “Scalable Diffusion Models with Transformers.” Proceedings of the IEEE/CVF International Conference on Computer Vision. https://arxiv.org/abs/2212.09748.
Peng, Bo, Daniel Goldstein, Quentin Anthony, et al. 2024. “Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.” First Conference on Language Modeling. https://arxiv.org/abs/2404.05892.
Peng, Bowen, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024. YaRN: Efficient Context Window Extension of Large Language Models.” International Conference on Learning Representations. https://arxiv.org/abs/2309.00071.
Peng, Bo, Ruichong Zhang, Daniel Goldstein, et al. 2025. RWKV-7 "Goose" with Expressive Dynamic State Evolution.” arXiv Preprint arXiv:2503.14456.
Pennington, Jeffrey, Samuel Schoenholz, and Surya Ganguli. 2017. “Resurrecting the Sigmoid in Deep Learning Through Dynamical Isometry: Theory and Practice.” Advances in Neural Information Processing Systems, 4785–95. https://proceedings.neurips.cc/paper/2017/hash/a9fc2d3b0c721b5b3f0b3f11bb24a75c-Abstract.html.
Pennington, Jeffrey, Richard Socher, and Christopher Manning. 2014. “GloVe: Global Vectors for Word Representation.” Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1532–43. https://doi.org/10.3115/v1/d14-1162.
Peters, Jonas, Dominik Janzing, and Bernhard Schölkopf. 2017. Elements of Causal Inference: Foundations and Learning Algorithms. MIT Press. https://doi.org/10.7551/mitpress/11283.001.0001.
Peters, Matthew, Waleed Ammar, Chandra Bhagavatula, and Russell Power. 2017. “Semi-Supervised Sequence Tagging with Bidirectional Language Models.” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Volume 1, 1756–65. https://doi.org/10.18653/v1/p17-1161.
Peters, Matthew, Mark Neumann, Mohit Iyyer, et al. 2018. “Deep Contextualized Word Representations.” Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, 2227–37. https://doi.org/10.18653/v1/n18-1202.
Petersen, Kaare Brandt, and Michael Syskind Pedersen. 2008. The Matrix Cookbook. Technical University of Denmark. https://www.math.uwaterloo.ca/~hwolkowi/matrixcookbook.pdf.
Peyré, Gabriel, and Marco Cuturi. 2019. “Computational Optimal Transport.” Foundations and Trends in Machine Learning 11 (5–6): 355–607.
Pinsker, Mark S. 1964. Information and Information Stability of Random Variables and Processes. Holden-Day.
Planck, Max. 1917. Über Einen Satz Der Statistischen Dynamik Und Seine Erweiterung in Der Quantentheorie.” Sitzungsberichte Der Preussischen Akademie Der Wissenschaften, 324–41.
Pleiss, Geoff, Danlu Chen, Gao Huang, Tongcheng Li, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017. “Memory-Efficient Implementation of Densenets.” ArXiv:1707.06990. https://arxiv.org/abs/1707.06990.
Polyak, Boris T. 1964. “Some Methods of Speeding up the Convergence of Iteration Methods.” USSR Computational Mathematics and Mathematical Physics 4 (5): 1–17. https://doi.org/10.1016/0041-5553(64)90137-5.
Polyak, Boris T. 1963. “Gradient Methods for the Minimisation of Functionals.” USSR Computational Mathematics and Mathematical Physics 3 (4): 864–78.
Polyak, Boris T., and Anatoli B. Juditsky. 1992. “Acceleration of Stochastic Approximation by Averaging.” SIAM Journal on Control and Optimization 30 (4): 838–55.
Pontryagin, Lev S., Vladimir G. Boltyanskii, Revaz V. Gamkrelidze, and Evgenii F. Mishchenko. 1962. The Mathematical Theory of Optimal Processes. Interscience.
Pooladian, Aram-Alexandre, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon Amos, Yaron Lipman, and Ricky T. Q. Chen. 2023. “Multisample Flow Matching: Straightening Flows with Minibatch Couplings.” Proceedings of the 40th International Conference on Machine Learning.
Poole, Ben, Sherjil Ozair, Aaron van den Oord, Alexander A. Alemi, and George Tucker. 2019. “On Variational Bounds of Mutual Information.” Proceedings of the 36th International Conference on Machine Learning, 5171–80.
Pope, Reiner, Sholto Douglas, Aakanksha Chowdhery, et al. 2023. “Efficiently Scaling Transformer Inference.” Proceedings of Machine Learning and Systems 5.
Popović, Maja. 2015. “ChrF: Character n-Gram F-Score for Automatic MT Evaluation.” Proceedings of the Tenth Workshop on Statistical Machine Translation, 392–95. https://doi.org/10.18653/v1/W15-3049.
Popper, Karl. 2005. The Logic of Scientific Discovery. Routledge. https://doi.org/10.4324/9780203994627.
Power, Alethea, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. 2022. “Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.” ArXiv:2201.02177. https://arxiv.org/abs/2201.02177.
Prakash, Aaditya, Sadid A Hasan, Kathy Lee, et al. 2016. “Neural Paraphrase Generation with Stacked Residual LSTM Networks.” ArXiv:1610.03098. https://arxiv.org/abs/1610.03098.
Press, Ofir, Noah A. Smith, and Mike Lewis. 2022. “Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.” International Conference on Learning Representations. https://arxiv.org/abs/2108.12409.
Press, Ofir, and Lior Wolf. 2017. “Using the Output Embedding to Improve Language Models.” Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, 157–63.
Qin, Danfeng, Chas Leichner, Manolis Delakis, et al. 2024. MobileNetV4: Universal Models for the Mobile Ecosystem.” European Conference on Computer Vision. https://arxiv.org/abs/2404.10518.
Quadrana, Massimo, Paolo Cremonesi, and Dietmar Jannach. 2018. “Sequence-Aware Recommender Systems.” ACM Computing Surveys 51 (4): 66. https://doi.org/10.1145/3190616.
Quinlan, J Ross. 1993. C4.5: Programs for Machine Learning. Elsevier. https://doi.org/10.1016/c2009-0-27846-9.
Qwen Team. 2025. Qwen3-Next: Towards Ultimate Training and Inference Efficiency. Model release, https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct.
Rabe, Markus N., and Charles Staats. 2021. “Self-Attention Does Not Need O(n2) Memory.” arXiv Preprint arXiv:2112.05682.
Rabiner, Lawrence, and Biing-Hwang Juang. 1993. Fundamentals of Speech Recognition. Prentice-Hall.
Rademacher, Hans. 1919. Über Partielle Und Totale Differenzierbarkeit von Funktionen Mehrerer Variabeln Und über Die Transformation Der Doppelintegrale.” Mathematische Annalen 79 (4): 340–59.
Radford, Alec, Jong Wook Kim, Chris Hallacy, et al. 2021. “Learning Transferable Visual Models from Natural Language Supervision.” International Conference on Machine Learning, 8748–63. https://proceedings.mlr.press/v139/radford21a.html.
Radford, Alec, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. “Robust Speech Recognition via Large-Scale Weak Supervision.” International Conference on Machine Learning. https://arxiv.org/abs/2212.04356.
Radford, Alec, Luke Metz, and Soumith Chintala. 2015. “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks.” ArXiv:1511.06434. https://arxiv.org/abs/1511.06434.
Radford, Alec, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. “Improving Language Understanding by Generative Pre-Training.” OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf.
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. “Language Models Are Unsupervised Multitask Learners.” OpenAI Blog 1 (8): 9. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
Radhakrishna Rao, C. 1945. “Information and Accuracy Attainable in the Estimation of Statistical Parameters.” Bulletin of the Calcutta Mathematical Society 37 (3): 81–91. https://doi.org/10.1007/BF02908227.
Radon, Johann. 1921. “Mengen Konvexer Körper, Die Einen Gemeinsamen Punkt Enthalten.” Mathematische Annalen 83 (1–2): 113–15.
Radosavovic, Ilija, Justin Johnson, Saining Xie, Wan-Yen Lo, and Piotr Dollár. 2019. “On Network Design Spaces for Visual Recognition.” Proceedings of the IEEE/CVF International Conference on Computer Vision, 1882–90. https://doi.org/10.1109/iccv.2019.00052.
Radosavovic, Ilija, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. 2020. “Designing Network Design Spaces.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10428–36. https://doi.org/10.1109/cvpr42600.2020.01044.
Rae, Jack W, Sebastian Borgeaud, Trevor Cai, et al. 2021. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher.” ArXiv:2112.11446. https://arxiv.org/abs/2112.11446.
Raffel, Colin, Noam Shazeer, Adam Roberts, et al. 2020. “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.” Journal of Machine Learning Research 21: 1–67. https://arxiv.org/abs/1910.10683.
Rahimi, Ali, and Benjamin Recht. 2007. “Random Features for Large-Scale Kernel Machines.” Advances in Neural Information Processing Systems 20.
Rajbhandari, Samyam, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.” SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, 1–16.
Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text.” ArXiv:1606.05250. https://arxiv.org/abs/1606.05250.
Ramachandran, Prajit, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. 2019. “Stand-Alone Self-Attention in Vision Models.” Advances in Neural Information Processing Systems 32. https://arxiv.org/abs/1906.05909.
Ramachandran, Prajit, Barret Zoph, and Quoc V Le. 2017. “Searching for Activation Functions.” ArXiv:1710.05941. https://arxiv.org/abs/1710.05941.
Ramesh, Aditya, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. “Hierarchical Text-Conditional Image Generation with Clip Latents.” ArXiv:2204.06125. https://arxiv.org/abs/2204.06125.
Ramón y Cajal, Santiago, and L. Azoulay. 1894. Les Nouvelles Idées Sur La Structure Du Système Nerveux Chez l’Homme Et Chez Les Vertébrés. Paris, C. Reinwald & Cie.
Ranzato, Marc-Aurelio, Y-Lan Boureau, Sumit Chopra, and Yann LeCun. 2007. “A Unified Energy-Based Framework for Unsupervised Learning.” Artificial Intelligence and Statistics, 371–79. https://proceedings.mlr.press/v2/ranzato07a.html.
Rasmussen, Carl Edward, and Christopher KI Williams. 2006. Gaussian Processes for Machine Learning. MIT Press. https://gaussianprocess.org/gpml/.
Recht, Benjamin, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019. “Do ImageNet Classifiers Generalize to ImageNet?” International Conference on Machine Learning, 5389–400. https://arxiv.org/abs/1902.10811.
Reddi, Sashank J, Satyen Kale, and Sanjiv Kumar. 2018. “On the Convergence of Adam and Beyond.” International Conference on Learning Representations. https://arxiv.org/abs/1904.09237.
Redmon, Joseph, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. “You Only Look Once: Unified, Real-Time Object Detection.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 779–88. https://doi.org/10.1109/cvpr.2016.91.
Redmon, Joseph, and Ali Farhadi. 2018. “YOLOv3: An Incremental Improvement.” ArXiv:1804.02767. https://arxiv.org/abs/1804.02767.
Reed, Scott, and Nando De Freitas. 2015. “Neural Programmer-Interpreters.” ArXiv:1511.06279. https://arxiv.org/abs/1511.06279.
Reed, Scott, Konrad Zolna, Emilio Parisotto, et al. 2022. “A Generalist Agent.” ArXiv:2205.06175. https://arxiv.org/abs/2205.06175.
Ren, Liliang, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu Chen. 2024. “Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.” arXiv Preprint arXiv:2406.07522.
Ren, Shaoqing, Kaiming He, Ross Girshick, and Jian Sun. 2015. “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.” Advances in Neural Information Processing Systems, 91–99. https://doi.org/10.5555/2969239.2969250.
Rendle, Steffen. 2010. “Factorization Machines.” 2010 IEEE International Conference on Data Mining, 995–1000. https://doi.org/10.1109/icdm.2010.127.
Rendle, Steffen, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian Personalized Ranking from Implicit Feedback.” Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, 452–61. https://arxiv.org/abs/1205.2618.
Rényi, Alfréd. 1961. “On Measures of Entropy and Information.” Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics, 547–61.
Revels, Jarrett, Miles Lubin, and Theodore Papamarkou. 2016. “Forward-Mode Automatic Differentiation in Julia.” ArXiv:1607.07892. https://arxiv.org/abs/1607.07892.
Rezende, Danilo Jimenez, Shakir Mohamed, and Daan Wierstra. 2014. “Stochastic Backpropagation and Approximate Inference in Deep Generative Models.” International Conference on Machine Learning, 1278–86. https://proceedings.mlr.press/v32/rezende14.html.
Rezende, Danilo, and Shakir Mohamed. 2015. “Variational Inference with Normalizing Flows.” International Conference on Machine Learning, 1530–38. https://proceedings.mlr.press/v37/rezende15.html.
Riesenhuber, Maximilian, and Tomaso Poggio. 1999. “Hierarchical Models of Object Recognition in Cortex.” Nature Neuroscience 2 (11): 1019–25. https://doi.org/10.1038/14819.
Risken, Hannes. 1996. The Fokker–Planck Equation: Methods of Solution and Applications. 2nd ed. Springer.
Rissanen, Jorma J. 1976. “Generalized Kraft Inequality and Arithmetic Coding.” IBM Journal of Research and Development 20 (3): 198–203.
Robbins, Herbert, and Sutton Monro. 1951. “A Stochastic Approximation Method.” The Annals of Mathematical Statistics 22 (3): 400–407.
Rockafellar, R. T. 1970. Convex Analysis. Princeton University Press. https://doi.org/10.1515/9781400873173.
Rolnick, David, Andreas Veit, Serge Belongie, and Nir Shavit. 2017. “Deep Learning Is Robust to Massive Label Noise.” ArXiv:1705.10694. https://arxiv.org/abs/1705.10694.
Rombach, Robin, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. “High-Resolution Image Synthesis with Latent Diffusion Models.” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10684–95.
Rudin, W. 1973. Functional Analysis. McGraw-Hill.
Rudin, Walter. 1976. Principles of Mathematical Analysis. 3rd ed. McGraw-Hill.
Rumelhart, David E, Geoffrey E Hinton, and Ronald J Williams. 1986. “Learning Representations by Back-Propagating Errors.” Nature 323 (6088): 533–36. https://doi.org/10.1038/323533a0.
Russakovsky, Olga, Jia Deng, Zhiheng Huang, Alexander C. Berg, and Li Fei-Fei. 2013. “Detecting Avocados to Zucchinis: What Have We Done, and Where Are We Going?” International Conference on Computer Vision (ICCV). https://doi.org/10.1109/iccv.2013.258.
Russakovsky, Olga, Jia Deng, Hao Su, et al. 2015. “ImageNet Large Scale Visual Recognition Challenge.” International Journal of Computer Vision 115 (3): 211–52. https://doi.org/10.1007/s11263-015-0816-y.
Russell, Stuart J, and Peter Norvig. 2016. Artificial Intelligence: A Modern Approach. Pearson Education Limited.
Saerens, Marco, Patrice Latinne, and Christine Decaestecker. 2002. “Adjusting the Outputs of a Classifier to New a Priori Probabilities: A Simple Procedure.” Neural Computation 14 (1): 21–41. https://doi.org/10.1162/089976602753284446.
Sahami, Mehran, Susan Dumais, David Heckerman, and Eric Horvitz. 1998. “A Bayesian Approach to Filtering Junk e-Mail.” AAAI Workshop on Learning for Text Categorization.
Saharia, Chitwan, William Chan, Saurabh Saxena, et al. 2022. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.” ArXiv:2205.11487. https://arxiv.org/abs/2205.11487.
Salimans, Tim, and Jonathan Ho. 2022. “Progressive Distillation for Fast Sampling of Diffusion Models.” International Conference on Learning Representations.
Salinas, D., M. Seeger, A. Klein, V. Perrone, M. Wistuba, and C. Archambeau. 2022. “Syne Tune: A Library for Large Scale Hyperparameter Tuning and Reproducible Research.” First Conference on Automated Machine Learning. https://proceedings.mlr.press/v188/salinas22a.html.
Sandler, Mark, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/1801.04381.
Santurkar, Shibani, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. 2018. “How Does Batch Normalization Help Optimization?” Advances in Neural Information Processing Systems, 2483–93. https://doi.org/10.5555/3327345.3327508.
Särkkä, Simo, and Arno Solin. 2019. Applied Stochastic Differential Equations. Cambridge University Press.
Sarwar, Badrul Munir, George Karypis, Joseph A Konstan, and John Riedl. 2001. “Item-Based Collaborative Filtering Recommendation Algorithms.” Proceedings of 10th International Conference on World Wide Web, 285–95. https://doi.org/10.1145/371920.372071.
Saxe, Andrew M., Yamini Bansal, Joel Dapello, et al. 2018. “On the Information Bottleneck Theory of Deep Learning.” International Conference on Learning Representations.
Saxe, Andrew M., James L. McClelland, and Surya Ganguli. 2014. “Exact Solutions to the Nonlinear Dynamics of Learning in Deep Linear Neural Networks.” International Conference on Learning Representations. https://arxiv.org/abs/1312.6120.
Schein, Andrew I, Alexandrin Popescul, Lyle H Ungar, and David M Pennock. 2002. “Methods and Metrics for Cold-Start Recommendations.” Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 253–60. https://doi.org/10.1145/564376.564421.
Schlag, Imanol, Kazuki Irie, and Jürgen Schmidhuber. 2021. “Linear Transformers Are Secretly Fast Weight Programmers.” International Conference on Machine Learning.
Schmidt, Robin M., Frank Schneider, and Philipp Hennig. 2021. “Descending Through a Crowded Valley — Benchmarking Deep Learning Optimizers.” Proceedings of the 38th International Conference on Machine Learning, 9367–76. https://arxiv.org/abs/2007.01547.
Schölkopf, Bernhard, and Alexander J Smola. 2002. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. MIT Press. https://doi.org/10.7551/mitpress/4175.001.0001.
Schölkopf, B., R. Herbrich, and A. J. Smola. 2001. “A Generalized Representer Theorem.” In Proceedings of the Annual Conference on Computational Learning Theory, edited by D. P. Helmbold and B. Williamson. Springer-Verlag. https://doi.org/10.1007/3-540-44581-1_27.
Schuster, Mike, and Kuldip K Paliwal. 1997. “Bidirectional Recurrent Neural Networks.” IEEE Transactions on Signal Processing 45 (11): 2673–81. https://doi.org/10.1109/78.650093.
Sedhain, Suvash, Aditya Krishna Menon, Scott Sanner, and Lexing Xie. 2015. “AutoRec: Autoencoders Meet Collaborative Filtering.” Proceedings of the 24th International Conference on World Wide Web, 111–12. https://doi.org/10.1145/2740908.2742726.
Sennrich, Rico, Barry Haddow, and Alexandra Birch. 2016. “Neural Machine Translation of Rare Words with Subword Units.” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1715–25. https://doi.org/10.18653/v1/P16-1162.
Shah, Ishaan, Anthony M. Polloreno, Karl Stratos, et al. 2025. “Practical Efficiency of Muon for Pretraining.” arXiv Preprint arXiv:2505.02222.
Shallue, Christopher J., Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. 2019. “Measuring the Effects of Data Parallelism on Neural Network Training.” Journal of Machine Learning Research 20 (112): 1–49. https://arxiv.org/abs/1811.03600.
Shannon, Claude E. 1959. “Coding Theorems for a Discrete Source with a Fidelity Criterion.” IRE National Convention Record 7 (4): 142–63.
Shannon, Claude Elwood. 1948. “A Mathematical Theory of Communication.” The Bell System Technical Journal 27 (3): 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x.
Shannon, Claude Elwood. 1951. “Prediction and Entropy of Printed English.” The Bell System Technical Journal 30 (1): 50–64. https://doi.org/10.1002/j.1538-7305.1951.tb01366.x.
Shaw, Peter, Jakob Uszkoreit, and Ashish Vaswani. 2018. “Self-Attention with Relative Position Representations.” ArXiv:1803.02155. https://arxiv.org/abs/1803.02155.
Shazeer, Noam. 2019. “Fast Transformer Decoding: One Write-Head Is All You Need.” arXiv Preprint arXiv:1911.02150.
Shazeer, Noam. 2020. GLU Variants Improve Transformer.” ArXiv:2002.05202. https://arxiv.org/abs/2002.05202.
Shazeer, Noam, Azalia Mirhoseini, Krzysztof Maziarz, et al. 2017. “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” International Conference on Learning Representations.
Shazeer, Noam, and Mitchell Stern. 2018. “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.” International Conference on Machine Learning. https://arxiv.org/abs/1804.04235.
Shimodaira, Hidetoshi. 2000. “Improving Predictive Inference Under Covariate Shift by Weighting the Log-Likelihood Function.” Journal of Statistical Planning and Inference 90 (2): 227–44.
Shwartz-Ziv, Ravid, and Amitai Armon. 2022. “Tabular Data: Deep Learning Is Not All You Need.” Information Fusion 81: 84–90. https://doi.org/10.1016/j.inffus.2021.11.011.
Shwartz-Ziv, Ravid, and Naftali Tishby. 2017. “Opening the Black Box of Deep Neural Networks via Information.” arXiv Preprint arXiv:1703.00810.
Siems, Julien, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi. 2025. DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products.” arXiv Preprint arXiv:2502.10297.
Silver, David, Aja Huang, Chris J Maddison, et al. 2016. “Mastering the Game of Go with Deep Neural Networks and Tree Search.” Nature 529 (7587): 484–89. https://doi.org/10.1038/nature16961.
Silverman, B. W. 1986. Density Estimation for Statistical and Data Analysis. Chapman; Hall.
Simonyan, Karen, and Andrew Zisserman. 2015. “Very Deep Convolutional Networks for Large-Scale Image Recognition.” International Conference on Learning Representations. https://arxiv.org/abs/1409.1556.
Sindhwani, Vikas, Tara N Sainath, and Sanjiv Kumar. 2015. “Structured Transforms for Small-Footprint Deep Learning.” ArXiv:1510.01722. https://arxiv.org/abs/1510.01722.
Sivic, Josef, and Andrew Zisserman. 2003. “Video Google: A Text Retrieval Approach to Object Matching in Videos.” Proceedings of the IEEE International Conference on Computer Vision 3: 1470–77. https://doi.org/10.1109/iccv.2003.1238663.
Slater, Morton. 1950. Lagrange Multipliers Revisited. Discussion Paper Mathematics 403. Cowles Commission for Research in Economics.
Smith, Jimmy T. H., Andrew Warrington, and Scott W. Linderman. 2023. “Simplified State Space Layers for Sequence Modeling.” International Conference on Learning Representations. https://arxiv.org/abs/2208.04933.
Smith, Samuel L., Andrew Brock, Leonard Berrada, and Soham De. 2023. “ConvNets Match Vision Transformers at Scale.” ArXiv:2310.16764. https://arxiv.org/abs/2310.16764.
Smith, Samuel L., Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le. 2018. “Don’t Decay the Learning Rate, Increase the Batch Size.” International Conference on Learning Representations. https://arxiv.org/abs/1711.00489.
Snoek, J., H. Larochelle, and R. Adams. 2012. “Practical Bayesian Optimization of Machine Learning Algorithms.” Advances in Neural Information Processing Systems 25, 2951–59. https://doi.org/10.5555/2999325.2999464.
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.” International Conference on Machine Learning, 2256–65. https://proceedings.mlr.press/v37/sohl-dickstein15.html.
Song, Jiaming, Chenlin Meng, and Stefano Ermon. 2021. “Denoising Diffusion Implicit Models.” International Conference on Learning Representations.
Song, Yang, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. 2023. “Consistency Models.” Proceedings of the 40th International Conference on Machine Learning.
Song, Yang, Conor Durkan, Iain Murray, and Stefano Ermon. 2021. “Maximum Likelihood Training of Score-Based Diffusion Models.” Advances in Neural Information Processing Systems 34.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by Estimating Gradients of the Data Distribution.” Advances in Neural Information Processing Systems 32. https://arxiv.org/abs/1907.05600.
Song, Yang, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative Modeling Through Stochastic Differential Equations.” International Conference on Learning Representations. https://doi.org/10.52202/075280-1645.
Soudry, Daniel, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. 2018. “The Implicit Bias of Gradient Descent on Separable Data.” Journal of Machine Learning Research 19 (70): 1–57.
Speelpenning, Bert. 1980. “Compiling Fast Partial Derivatives of Functions Given by Algorithms.” PhD thesis, University of Illinois at Urbana-Champaign. https://doi.org/10.2172/5254402.
Srivastava, Nitish, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. “Dropout: A Simple Way to Prevent Neural Networks from Overfitting.” Journal of Machine Learning Research 15 (1): 1929–58. https://doi.org/10.5555/2627435.2670313.
Srivastava, Rupesh Kumar, Klaus Greff, and Jürgen Schmidhuber. 2015. “Highway Networks.” ArXiv:1505.00387. https://arxiv.org/abs/1505.00387.
Stein, Charles M. 1981. “Estimation of the Mean of a Multivariate Normal Distribution.” Annals of Statistics 9 (6): 1135–51.
Sterbenz, Pat H. 1974. Floating-Point Computation. Prentice-Hall.
Strang, Gilbert. 1993. Introduction to Linear Algebra. Wellesley–Cambridge Press. https://math.mit.edu/~gs/linearalgebra/.
Stratonovich, Ruslan L. 1966. “A New Representation for Stochastic Integrals and Equations.” SIAM Journal on Control 4 (2): 362–71.
Student (Gosset, William S.). 1908. “The Probable Error of a Mean.” Biometrika 6 (1): 1–25.
Su, Jianlin, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2024. RoFormer: Enhanced Transformer with Rotary Position Embedding.” Neurocomputing 568: 127063. https://arxiv.org/abs/2104.09864.
Su, Xiaoyuan, and Taghi M Khoshgoftaar. 2009. “A Survey of Collaborative Filtering Techniques.” Advances in Artificial Intelligence 2009. https://doi.org/10.1155/2009/421425.
Sukhbaatar, Sainbayar, Jason Weston, and Rob Fergus. 2015. “End-to-End Memory Networks.” Advances in Neural Information Processing Systems, 2440–48. https://doi.org/10.5555/2969239.2969426.
Sun, Yu, Xinhao Li, Karan Dalal, et al. 2024. “Learning to (Learn at Test Time): RNNs with Expressive Hidden States.” arXiv Preprint arXiv:2407.04620.
Sun, Yutao, Li Dong, Shaohan Huang, et al. 2023. “Retentive Network: A Successor to Transformer for Large Language Models.” arXiv Preprint arXiv:2307.08621.
Sutskever, Ilya, James Martens, George Dahl, and Geoffrey Hinton. 2013. “On the Importance of Initialization and Momentum in Deep Learning.” International Conference on Machine Learning, 1139–47. https://proceedings.mlr.press/v28/sutskever13.html.
Sutskever, Ilya, Oriol Vinyals, and Quoc V Le. 2014. “Sequence to Sequence Learning with Neural Networks.” Advances in Neural Information Processing Systems, 3104–12. https://doi.org/10.5555/2969033.2969173.
Szegedy, Christian, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017. “Inception-V4, Inception-ResNet and the Impact of Residual Connections on Learning.” 31st AAAI Conference on Artificial Intelligence. https://doi.org/10.1609/aaai.v31i1.11231.
Szegedy, Christian, Wei Liu, Yangqing Jia, et al. 2015. “Going Deeper with Convolutions.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1–9. https://doi.org/10.1109/cvpr.2015.7298594.
Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. “Rethinking the Inception Architecture for Computer Vision.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2818–26. https://doi.org/10.1109/cvpr.2016.308.
Székely, Gábor J., and Maria L. Rizzo. 2013. “Energy Statistics: A Class of Statistics Based on Distances.” Journal of Statistical Planning and Inference 143 (8): 1249–72.
Tallec, Corentin, and Yann Ollivier. 2017. “Unbiasing Truncated Backpropagation Through Time.” ArXiv:1705.08209. https://arxiv.org/abs/1705.08209.
Tan, Mingxing, and Quoc Le. 2019. “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.” International Conference on Machine Learning, 6105–14. https://proceedings.mlr.press/v97/tan19a.html.
Tan, Mingxing, and Quoc V. Le. 2021. “EfficientNetV2: Smaller Models and Faster Training.” International Conference on Machine Learning. https://arxiv.org/abs/2104.00298.
Tang, Jiaxi, and Ke Wang. 2018. “Personalized Top-n Sequential Recommendation via Convolutional Sequence Embedding.” Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, 565–73. https://doi.org/10.1145/3159652.3159656.
Tao, Terence, and Van Vu. 2010. “Random Matrices: Universality of ESDs and the Circular Law.” Annals of Probability 38 (5): 2023–65.
Taskar, Ben, Carlos Guestrin, and Daphne Koller. 2004. “Max-Margin Markov Networks.” Advances in Neural Information Processing Systems 16: 25.
Tay, Yi, Mostafa Dehghani, Samira Abnar, et al. 2021. “Long Range Arena: A Benchmark for Efficient Transformers.” International Conference on Learning Representations. https://arxiv.org/abs/2011.04006.
Tay, Yi, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020. “Efficient Transformers: A Survey.” ArXiv:2009.06732. https://arxiv.org/abs/2009.06732.
Team OLMo. 2025. “Olmo 3.” arXiv Preprint arXiv:2512.13961.
Team OLMo, Pete Walsh, Luca Soldaini, et al. 2025. “2 OLMo 2 Furious.” arXiv Preprint arXiv:2501.00656.
Telgarsky, Matus. 2016. “Benefits of Depth in Neural Networks.” Conference on Learning Theory, 1517–39.
Tencent Hunyuan Team. 2025. “Hunyuan-TurboS: Advancing Large Language Models Through Mamba-Transformer Synergy and Adaptive Chain-of-Thought.” arXiv Preprint arXiv:2505.15431.
Teye, Mattias, Hossein Azizpour, and Kevin Smith. 2018. “Bayesian Uncertainty Estimation for Batch Normalized Deep Networks.” ArXiv:1802.06455. https://arxiv.org/abs/1802.06455.
Thomee, Bart, David A Shamma, Gerald Friedland, et al. 2016. “YFCC100M: The New Data in Multimedia Research.” Communications of the ACM 59 (2): 64–73. https://doi.org/10.1145/2812802.
Tibshirani, Robert. 1996. “Regression Shrinkage and Selection via the Lasso.” Journal of the Royal Statistical Society: Series B 58 (1): 267–88.
Tieleman, Tijmen, and Geoffrey Hinton. 2012. “Divide the Gradient by a Running Average of Its Recent Magnitude.” In COURSERA: Neural Networks for Machine Learning, Lecture 6.5-Rmsprop. https://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf.
Tikhonov, A. N., and V. Y. Arsenin. 1977. Solutions of Ill-Posed Problems. W.HWinston. https://doi.org/10.1137/1.9780898719741.
Tillet, Philippe, H. T. Kung, and David Cox. 2019. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations.” 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (MAPL), 10–19. https://doi.org/10.1145/3315508.3329973.
Tishby, Naftali, Fernando C. Pereira, and William Bialek. 1999. “The Information Bottleneck Method.” Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing, 368–77.
Tong, Alexander, Kilian Fatras, Nikolay Malkin, et al. 2024. “Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport.” Transactions on Machine Learning Research.
Torralba, Antonio, Rob Fergus, and William T Freeman. 2008. “80 Million Tiny Images: A Large Data Set for Nonparametric Object and Scene Recognition.” IEEE Transactions on Pattern Analysis and Machine Intelligence 30 (11): 1958–70. https://doi.org/10.1109/tpami.2008.128.
Töscher, Andreas, Michael Jahrer, and Robert M Bell. 2009. The Bigchaos Solution to the Netflix Grand Prize. https://www.netflixprize.com/assets/GrandPrize2009_BPC_BigChaos.pdf.
Touvron, Hugo, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021. “Training Data-Efficient Image Transformers & Distillation Through Attention.” International Conference on Machine Learning, 10347–57. https://proceedings.mlr.press/v139/touvron21a.html.
Touvron, Hugo, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. 2021. “Going Deeper with Image Transformers.” IEEE/CVF International Conference on Computer Vision, 32–42. https://arxiv.org/abs/2103.17239.
Touvron, Hugo, Thibaut Lavril, Gautier Izacard, et al. 2023a. “LLaMA: Open and Efficient Foundation Language Models.” ArXiv:2302.13971, 2023a. https://arxiv.org/abs/2302.13971.
Touvron, Hugo, Louis Martin, Kevin Stone, et al. 2023b. “LLaMA 2: Open Foundation and Fine-Tuned Chat Models.” ArXiv:2307.09288, 2023b. https://arxiv.org/abs/2307.09288.
Tschannen, Michael, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly, and Mario Lucic. 2020. “On Mutual Information Maximization for Representation Learning.” International Conference on Learning Representations.
Tsoumakas, Grigorios, and Ioannis Katakis. 2007. “Multi-Label Classification: An Overview.” International Journal of Data Warehousing and Mining 3 (3): 1–13. https://doi.org/10.4018/jdwm.2007070101.
Turing, Alan. 1950. “Computing Machinery and Intelligence.” Mind 59 (236): 433–60. https://doi.org/10.1093/mind/lix.236.433.
Uhlenbeck, George E., and Leonard S. Ornstein. 1930. “On the Theory of the Brownian Motion.” Physical Review 36 (5): 823–41.
Uijlings, Jasper RR, Koen EA Van De Sande, Theo Gevers, and Arnold WM Smeulders. 2013. “Selective Search for Object Recognition.” International Journal of Computer Vision 104 (2): 154–71. https://doi.org/10.1007/s11263-013-0620-5.
Vapnik, V. 1995. The Nature of Statistical Learning Theory. Springer.
Vapnik, V. 1998. Statistical Learning Theory. John Wiley; Sons.
Vapnik, V. N., and A. Y. Chervonenkis. 1974. “Ordered Risk Minimization.” Automation and Remote Control 35: 1226–35, 1403–12.
Vapnik, V., and A. Chervonenkis. 1964. “A Note on One Class of Perceptrons.” Automation and Remote Control 25.
Vapnik, V., and A. Chervonenkis. 1968. “Uniform Convergence of Frequencies of Occurence of Events to Their Probabilities.” Dokl. Akad. Nauk SSSR 181: 915–18.
Vapnik, V., and A. Chervonenkis. 1971b. “On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities.” Theory Probab. Appl. 16 (2): 264–81. https://doi.org/10.1137/1116025.
Vapnik, V., and A. Chervonenkis. 1971a. “On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities.” Theory Probab. Appl. 16 (2): 264–81.
Vapnik, V., and A. Chervonenkis. 1981. “The Necessary and Sufficient Conditions for the Uniform Convergence of Averages to Their Expected Values.” Teoriya Veroyatnostei i Ee Primeneniya 26 (3): 543–64.
Vapnik, V., and A. Chervonenkis. 1991. “The Necessary and Sufficient Conditions for Consistency in the Empirical Risk Minimization Method.” Pattern Recognition and Image Analysis 1 (3): 283–305.
Vapnik, Vladimir. 1992. “Principles of Risk Minimization for Learning Theory.” Advances in Neural Information Processing Systems, 831–38. https://doi.org/10.5555/2986916.2987019.
Vapnik, Vladimir, Esther Levin, and Yann Le Cun. 1994. “Measuring the VC-Dimension of a Learning Machine.” Neural Computation 6 (5): 851–76. https://doi.org/10.1162/neco.1994.6.5.851.
Vasu, Pavan Kumar Anasosalu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan. 2023a. FastViT: A Fast Hybrid Vision Transformer Using Structural Reparameterization.” Proceedings of the IEEE/CVF International Conference on Computer Vision. https://arxiv.org/abs/2303.14189.
Vasu, Pavan Kumar Anasosalu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan. 2023b. MobileOne: An Improved One Millisecond Mobile Backbone.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2206.04040.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. 2017. “Attention Is All You Need.” Advances in Neural Information Processing Systems, 5998–6008. https://doi.org/10.5555/3295222.3295349.
Vershynin, Roman. 2018. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press.
Vincent, Pascal. 2011. “A Connection Between Score Matching and Denoising Autoencoders.” Neural Computation 23 (7): 1661–74.
Vyas, Nikhil, Depen Morwani, Rosie Zhao, et al. 2024. SOAP: Improving and Stabilizing Shampoo Using Adam.” arXiv Preprint arXiv:2409.11321.
Wahba, Grace. 1990. Spline Models for Observational Data. SIAM. https://doi.org/10.1137/1.9781611970128.
Waibel, Alex, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J Lang. 1989. “Phoneme Recognition Using Time-Delay Neural Networks.” IEEE Transactions on Acoustics, Speech, and Signal Processing 37 (3): 328–39. https://doi.org/10.1016/b978-0-08-051584-7.50037-1.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical Models, Exponential Families, and Variational Inference.” Foundations and Trends in Machine Learning 1 (1–2): 1–305.
Waleffe, Roger, Wonmin Byeon, Duncan Riach, et al. 2024. “An Empirical Study of Mamba-Based Language Models.” arXiv Preprint arXiv:2406.07887.
Wan, Li, Matthew Zeiler, Sixin Zhang, Yann LeCun, and Rob Fergus. 2013. “Regularization of Neural Networks Using Dropconnect.” International Conference on Machine Learning, 1058–66.
Wang, Dustin, Rui-Jie Zhu, Steven Abreu, et al. 2025. “A Systematic Analysis of Hybrid Linear Attention.” arXiv Preprint arXiv:2507.06457.
Wang, Haotao, Aston Zhang, Shuai Zheng, Xingjian Shi, Mu Li, and Zhangyang Wang. 2022. “Removing Batch Normalization Boosts Adversarial Training.” International Conference on Machine Learning, 23433–45. https://openreview.net/forum?id=2J8bBfGCPi.
Wang, Junxiong, Daniele Paliotta, Avner May, Alexander M. Rush, and Tri Dao. 2024. “The Mamba in the Llama: Distilling and Accelerating Hybrid Models.” Advances in Neural Information Processing Systems.
Wang, Ke Alexander, Jiaxin Shi, and Emily B. Fox. 2025. “Test-Time Regression: A Unifying Framework for Designing Sequence Models with Associative Memory.” arXiv Preprint arXiv:2501.12352.
Wang, Lean, Huazuo Gao, Chenggang Zhao, Xu Sun, and Damai Dai. 2024. “Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.” arXiv Preprint arXiv:2408.15664.
Wang, Qiang, Bei Li, Tong Xiao, et al. 2019. “Learning Deep Transformer Models for Machine Translation.” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 1810–22. https://doi.org/10.18653/v1/p19-1176.
Wang, Sinong, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. “Linformer: Self-Attention with Linear Complexity.” arXiv Preprint arXiv:2006.04768.
Wang, Wenhai, Jifeng Dai, Zhe Chen, et al. 2023. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2211.05778.
Warner, Benjamin, Antoine Chaffin, Benjamin Clavié, et al. 2024. “Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.” arXiv Preprint arXiv:2412.13663.
Warstadt, Alex, Amanpreet Singh, and Samuel R Bowman. 2019. “Neural Network Acceptability Judgments.” Transactions of the Association for Computational Linguistics 7: 625–41. https://doi.org/10.1162/tacl_a_00290.
Wasserman, Larry. 2013. All of Statistics: A Concise Course in Statistical Inference. Springer. https://link.springer.com/book/10.1007/978-0-387-21736-9.
Wasserstein, Ronald L, and Nicole A Lazar. 2016. “The ASA Statement on p-Values: Context, Process, and Purpose.” The American Statistician 70 (2): 129–33. https://doi.org/10.1080/00031305.2016.1154108.
Watkins, Christopher JCH, and Peter Dayan. 1992. “Q-Learning.” Machine Learning 8 (3–4): 279–92. https://doi.org/10.1007/bf00992698.
Watson, Geoffrey S. 1964. “Smooth Regression Analysis.” Sankhyā: The Indian Journal of Statistics, Series A, 359–72. https://doi.org/10.1007/bf02868765.
Wei, Jason, Yi Tay, Rishi Bommasani, et al. 2022. “Emergent Abilities of Large Language Models.” ArXiv:2206.07682. https://arxiv.org/abs/2206.07682.
Welford, B. P. 1962. “Note on a Method for Calculating Corrected Sums of Squares and Products.” Technometrics 4 (3): 419–20.
Welling, Max, and Yee W Teh. 2011. “Bayesian Learning via Stochastic Gradient Langevin Dynamics.” Proceedings of the 28th International Conference on Machine Learning (ICML-11), 681–88. https://dl.acm.org/doi/10.5555/3104482.3104568.
Wen, Kaiyue, David Hall, Tengyu Ma, and Percy Liang. 2025. “Fantastic Pretraining Optimizers and Where to Find Them.” arXiv Preprint arXiv:2509.02046.
Wen, Kaiyue, Zhiyuan Li, Jason Wang, David Hall, Percy Liang, and Tengyu Ma. 2024. “Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.” arXiv Preprint arXiv:2410.05192.
Wengert, Robert Edwin. 1964. “A Simple Automatic Derivative Evaluation Program.” Communications of the ACM 7 (8): 463–64. https://doi.org/10.1145/355588.365726.
Werbos, Paul J. 1990. “Backpropagation Through Time: What It Does and How to Do It.” Proceedings of the IEEE 78 (10): 1550–60. https://doi.org/10.1109/5.58337.
White, Halbert. 1982. “Maximum Likelihood Estimation of Misspecified Models.” Econometrica 50 (1): 1–25.
Widrow, Bernard, and Marcian E. Hoff. 1960. “Adaptive Switching Circuits.” IRE WESCON Convention Record 4: 96–104.
Wiegreffe, Sarah, and Yuval Pinter. 2019. “Attention Is Not Not Explanation.” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 11–20.
Wiener, Norbert. 1923. “Differential-Space.” Journal of Mathematics and Physics 2: 131–74.
Wightman, Ross, Hugo Touvron, and Hervé Jégou. 2021. “ResNet Strikes Back: An Improved Training Procedure in Timm.” ArXiv:2110.00476. https://arxiv.org/abs/2110.00476.
Wigner, Eugene P. 1958. “On the Distribution of the Roots of Certain Symmetric Matrices.” Ann. Math., 325–27. https://doi.org/10.2307/1970079.
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning.” Machine Learning 8: 229–56.
Williams, Samuel, Andrew Waterman, and David Patterson. 2009. Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures. Lawrence Berkeley National Lab. https://doi.org/10.2172/1407078.
Wilson, Andrew G, and Pavel Izmailov. 2020. “Bayesian Deep Learning and a Probabilistic Perspective of Generalization.” Advances in Neural Information Processing Systems 33: 4697–708. https://arxiv.org/abs/2002.08791.
Wistuba, M., A. Rawat, and T. Pedapati. 2019. “A Survey on Neural Architecture Search.” ArXiv:1905.01392 [Cs.LG]. https://arxiv.org/abs/1905.01392.
Wistuba, M., N. Schilling, and L. Schmidt-Thieme. 2018. “Scalable Gaussian Process-Based Transfer Surrogates for Hyperparameter Optimization.” Machine Learning 108: 43–78. https://doi.org/10.1007/s10994-017-5684-y.
Witten, Ian H., Radford M. Neal, and John G. Cleary. 1987. “Arithmetic Coding for Data Compression.” Communications of the ACM 30 (6): 520–40.
Wolpert, David H. 1996. “The Lack of a Priori Distinctions Between Learning Algorithms.” Neural Computation 8 (7): 1341–90.
Woo, Sanghyun, Shoubhik Debnath, Ronghang Hu, et al. 2023. ConvNeXt V2: Co-Designing and Scaling ConvNets with Masked Autoencoders.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2301.00808.
Wood, Frank, Jan Gasthaus, Cédric Archambeau, Lancelot James, and Yee Whye Teh. 2011. “The Sequence Memoizer.” Communications of the ACM 54 (2): 91–98. https://doi.org/10.1162/neco_a_00154.
Wortsman, Mitchell, Gabriel Ilharco, Samir Yitzhak Gadre, et al. 2022. “Model Soups: Averaging Weights of Multiple Fine-Tuned Models Improves Accuracy Without Increasing Inference Time.” International Conference on Machine Learning.
Wu, Bichen, Alvin Wan, Xiangyu Yue, et al. 2018. “Shift: A Zero Flop, Zero Parameter Alternative to Spatial Convolutions.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 9127–35. https://doi.org/10.1109/cvpr.2018.00951.
Wu, C. F. Jeff. 1983. “On the Convergence Properties of the EM Algorithm.” Annals of Statistics 11 (1): 95–103.
Wu, Yonghui, Mike Schuster, Zhifeng Chen, et al. 2016. “Google’s Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation.” ArXiv:1609.08144. https://arxiv.org/abs/1609.08144.
Wu, Yuxin, and Kaiming He. 2018. “Group Normalization.” European Conference on Computer Vision. https://arxiv.org/abs/1803.08494.
Xiao, Guangxuan, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. “Efficient Streaming Language Models with Attention Sinks.” International Conference on Learning Representations.
Xiao, Han, Kashif Rasul, and Roland Vollgraf. 2017. “Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms.” ArXiv:1708.07747. https://arxiv.org/abs/1708.07747.
Xiao, Lechao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington. 2018. “Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000-Layer Vanilla Convolutional Neural Networks.” International Conference on Machine Learning, 5393–402. https://proceedings.neurips.cc/paper/2018/hash/d76b67bcd3ec823de384ef62dc7e4c7e-Abstract.html.
Xie, Saining, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017. “Aggregated Residual Transformations for Deep Neural Networks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1492–500. https://doi.org/10.1109/cvpr.2017.634.
Xiong, Ruibin, Yunchang Yang, Di He, et al. 2020. “On Layer Normalization in the Transformer Architecture.” International Conference on Machine Learning, 10524–33. https://proceedings.mlr.press/v119/xiong20b.html.
Xiong, Wayne, Lingfeng Wu, Fil Alleva, Jasha Droppo, Xuedong Huang, and Andreas Stolcke. 2018. “The Microsoft 2017 Conversational Speech Recognition System.” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5934–38. https://doi.org/10.1109/TASLP.2018.2876459.
Xu, Yuanzhong, HyoukJoong Lee, Dehao Chen, et al. 2021. GSPMD: General and Scalable Parallelization for ML Computation Graphs.” ArXiv:2105.04663. https://arxiv.org/abs/2105.04663.
Yamaguchi, Kouichi, Kenji Sakamoto, Toshio Akabane, and Yoshiji Fujimoto. 1990. “A Neural Network for Speaker-Independent Isolated Word Recognition.” First International Conference on Spoken Language Processing. https://doi.org/10.21437/icslp.1990-282.
Yang, An, Anfeng Li, Baosong Yang, et al. 2025. Qwen3 Technical Report.” arXiv Preprint arXiv:2505.09388.
Yang, Greg, and Edward J. Hu. 2021. “Tensor Programs IV: Feature Learning in Infinite-Width Neural Networks.” International Conference on Machine Learning, 11727–37.
Yang, Greg, Edward J. Hu, Igor Babuschkin, et al. 2022. “Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.” arXiv Preprint arXiv:2203.03466.
Yang, Greg, James B. Simon, and Jeremy Bernstein. 2023. “A Spectral Condition for Feature Learning.” arXiv Preprint arXiv:2310.17813.
Yang, Songlin, Jan Kautz, and Ali Hatamizadeh. 2025. “Gated Delta Networks: Improving Mamba2 with Delta Rule.” International Conference on Learning Representations. https://arxiv.org/abs/2412.06464.
Yang, Songlin, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. 2024. “Gated Linear Attention Transformers with Hardware-Efficient Training.” International Conference on Machine Learning, 56501–23. https://arxiv.org/abs/2312.06635.
Yang, Songlin, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. 2024. “Parallelizing Linear Transformers with the Delta Rule over Sequence Length.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2406.06484.
Yang, Zichao, Zhiting Hu, Yuntian Deng, Chris Dyer, and Alex Smola. 2016. “Neural Machine Translation with Recurrent Attention Modeling.” ArXiv:1607.05108. https://arxiv.org/abs/1607.05108.
Yang, Zichao, Marcin Moczulski, Misha Denil, et al. 2015. “Deep Fried Convnets.” Proceedings of the IEEE International Conference on Computer Vision, 1476–83. https://doi.org/10.1109/iccv.2015.173.
Ye, Mao, Peifeng Yin, Wang-Chien Lee, and Dik-Lun Lee. 2011. “Exploiting Geographical Influence for Collaborative Point-of-Interest Recommendation.” Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, 325–34. https://doi.org/10.1145/2009916.2009962.
You, Yang, Igor Gitman, and Boris Ginsburg. 2017. “Large Batch Training of Convolutional Networks.” ArXiv:1708.03888. https://arxiv.org/abs/1708.03888.
Yu, Fisher, and Vladlen Koltun. 2016. “Multi-Scale Context Aggregation by Dilated Convolutions.” International Conference on Learning Representations. https://arxiv.org/abs/1511.07122.
Yuan, Jingyang, Huazuo Gao, Damai Dai, et al. 2025. “Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.” arXiv Preprint arXiv:2502.11089.
Yun, Sangdoo, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. 2019. CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features.” Proceedings of the IEEE/CVF International Conference on Computer Vision. https://arxiv.org/abs/1905.04899.
Zaheer, Manzil, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar. 2018. “Adaptive Methods for Nonconvex Optimization.” Advances in Neural Information Processing Systems, 9793–803. https://proceedings.neurips.cc/paper/2018/hash/90365351ccc7437a1309dc64e4db32a3-Abstract.html.
Zeiler, Matthew D. 2012. ADADELTA: An Adaptive Learning Rate Method.” ArXiv:1212.5701. https://arxiv.org/abs/1212.5701.
Zeiler, Matthew D, and Rob Fergus. 2013. “Stochastic Pooling for Regularization of Deep Convolutional Neural Networks.” ArXiv:1301.3557. https://arxiv.org/abs/1301.3557.
Zeng, Aohan, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, et al. 2025. GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.” arXiv Preprint arXiv:2508.06471.
Zhang, Aston, Yi Tay, Shuai Zhang, et al. 2021. “Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with 1/n Parameters.” International Conference on Learning Representations. https://openreview.net/forum?id=rcQdycl0zyk.
Zhang, Biao, and Rico Sennrich. 2019. “Root Mean Square Layer Normalization.” Advances in Neural Information Processing Systems 32.
Zhang, Chiyuan, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2021. “Understanding Deep Learning (Still) Requires Rethinking Generalization.” Communications of the ACM 64 (3): 107–15. https://doi.org/10.1145/3446776.
Zhang, Hanlin, Depen Morwani, Nikhil Vyas, et al. 2024. “How Does Critical Batch Size Scale in Pre-Training?” arXiv Preprint arXiv:2410.21676.
Zhang, Hongyi, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018. mixup: Beyond Empirical Risk Minimization.” International Conference on Learning Representations. https://arxiv.org/abs/1710.09412.
Zhang, Jingzhao, Sai Praneeth Karimireddy, Andreas Veit, et al. 2020. “Why Are Adaptive Methods Good for Attention Models?” Advances in Neural Information Processing Systems 33.
Zhang, Richard. 2019. “Making Convolutional Networks Shift-Invariant Again.” International Conference on Machine Learning. https://arxiv.org/abs/1904.11486.
Zhang, Shuai, Lina Yao, Aixin Sun, and Yi Tay. 2019. “Deep Learning Based Recommender System: A Survey and New Perspectives.” ACM Computing Surveys 52 (1): 5. https://doi.org/10.1145/3285029.
Zhang, Susan, Stephen Roller, Naman Goyal, et al. 2022. OPT: Open Pre-Trained Transformer Language Models.” ArXiv:2205.01068. https://arxiv.org/abs/2205.01068.
Zhang, Wei, Jun Tanida, Kazuyoshi Itoh, and Yoshiki Ichioka. 1988. “Shift-Invariant Pattern Recognition Neural Network and Its Optical Architecture.” Proceedings of Annual Conference of the Japan Society of Applied Physics.
Zhang, Yushun, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo. 2024. “Why Transformers Need Adam: A Hessian Perspective.” Advances in Neural Information Processing Systems 37. https://arxiv.org/abs/2402.16788.
Zhao, Yanli, Andrew Gu, Rohan Varma, et al. 2023. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.” Proceedings of the VLDB Endowment 16 (12): 3848–60. https://doi.org/10.14778/3611540.3611569.
Zhao, Zhong-Qiu, Peng Zheng, Shou-tao Xu, and Xindong Wu. 2019. “Object Detection with Deep Learning: A Review.” IEEE Transactions on Neural Networks and Learning Systems 30 (11): 3212–32. https://doi.org/10.1016/j.neucom.2018.09.013.
Zhu, Jun-Yan, Taesung Park, Phillip Isola, and Alexei A Efros. 2017. “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks.” Proceedings of the IEEE International Conference on Computer Vision, 2223–32. https://doi.org/10.1109/iccv.2017.244.
Zhu, Yukun, Ryan Kiros, Rich Zemel, et al. 2015. “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books.” Proceedings of the IEEE International Conference on Computer Vision, 19–27. https://doi.org/10.1109/iccv.2015.11.
Zoph, Barret, and Quoc V Le. 2016. “Neural Architecture Search with Reinforcement Learning.” ArXiv:1611.01578. https://arxiv.org/abs/1611.01578.
Zuo, Jingwei, Maksim Velikanov, Ilyas Chahed, et al. 2025. “Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance.” arXiv Preprint arXiv:2507.22448.