Papers
References and working notes around ML systems, evaluation, reliability, artificial life, uncertainty, and applied modeling.
Recent
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Andrey A. Dukhovny, Andrey M. Lange · 2026
PDF
A unified perspective of Gaussian process approximation for differential equations
Mengwu Guo · 2026
PDF
Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston · 2026
PDF
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Hong Zhao · 2026
PDF
Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Giovanni Pezzulo, Michael Levin · 2026
PDF alifeintelligence
Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 2026
PDF
Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Kuo Gai, Shihua Zhang · 2026
PDF
E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
Ben Chugg, Aaditya Ramdas, Peter Grünwald · 2026
PDF
E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
Alexandra Sokolova, Vadim Sokolov · 2026
PDF
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou · 2026
PDF
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · 2026
PDF
Recent Reading
A tighter selection from my paper library, with links out to the source paper when available.
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Andrey A. Dukhovny, Andrey M. Lange · 2026
PDF
A unified perspective of Gaussian process approximation for differential equations
Mengwu Guo · 2026
PDF
Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston · 2026
PDF
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Hong Zhao · 2026
PDF
Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 2026
PDF
Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Kuo Gai, Shihua Zhang · 2026
PDF
E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
Ben Chugg, Aaditya Ramdas, Peter Grünwald · 2026
PDF
E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
Alexandra Sokolova, Vadim Sokolov · 2026
PDF
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou · 2026
PDF
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · 2026
PDF
Fast Score-Based Sampling via Log-Concave Reductions
M. J. Wainwright · 2026
PDF
From Approximation to Emergence: A Theory of Deep Learning
Zhilin Zhao · 2026
PDF
Functional Attention: From Pairwise Affinities to Functional Correspondences
Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers · 2026
PDF
Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Peilin Liu, Ding-Xuan Zhou · 2026
PDF
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová · 2026
PDF
Lecture notes on random matrix theory: the results, the applications, and the analytical tools
Joseph W. Baron · 2026
PDF
Predictable GRPO: A Closed-Form Model of Training Dynamics
Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta · 2026
PDF
Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
Zhenyu Liao, Michael W. Mahoney · 2026
PDF
Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Daniela Schkoda, Dominik Janzing · 2026
PDF
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
Pratyaksh Rao, Wancong Zhang, Randall Balestriero, Yann LeCun, Giuseppe Loianno · 2026
PDF
Statistical Properties of Training & Generalization
Itay Lavie, Noam Levi, Yonatan Kahn · 2026
PDF
Still: Amortized KV Cache Compaction in a Single Forward Pass
Charles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby · 2026
PDF
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
Haggai Roitman · 2026
PDF
Basic Category Theory
Tom Leinster · 2025
PDF
E-Values Expand the Scope of Conformal Prediction
Etienne Gauthier, Francis Bach, Michael I. Jordan · 2025
PDF
Nonparametric Treatment Effect Identification in School Choice
Jiafeng Chen · 2025
PDF
Prediction-Powered E-Values
Daniel Csillag, Claudio José Struchiner, Guilherme Tegoni Goedert · 2025
PDF
Table Foundation Models: on knowledge pre-training for tabular learning
Myung Jun Kim, Félix Lefebvre, Gaëtan Brison, Alexandre Perez-Lebel, Gaël Varoquaux · 2025
PDF
Tensor Logic: The Language of AI
Pedro Domingos · 2025
PDF
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas · 2025
PDF
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
DINOv2: Learning Robust Visual Features without Supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski · 2024
PDF
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · 2024
PDF
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković · 2024
PDF
Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
Kexin Zhang, Qingsong Wen, Chaoli Zhang, Rongyao Cai, Ming Jin, Yong Liu, James Y. Zhang, Yuxuan Liang, Guansong Pang, Dongjin Song, Shirui Pan · 2024
Paper
SGLang: Efficient Execution of Structured Language Model Programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng · 2024
PDF
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou · 2023
PDF
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica · 2023
PDF
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu · 2023
PDF
Statistical Foundations of Prior-Data Fitted Networks
Thomas Nagler · 2023
PDF
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie · 2022
PDF
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · 2022
PDF
Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features
Florian Pargent, Florian Pfisterer, Janek Thomas, Bernd Bischl · 2022
PDF
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei · 2022
PDF
Scheduling Techniques for Liver Segmentation: ReduceLRonPlateau Vs OneCycleLR
Ayman Al-Kababji, Faycal Bensaali, Sarada Prasad Dakua · 2022
PDF
Survival Regression with Accelerated Failure Time Model in XGBoost
Avinash Barnwal, Hyunsu Cho, Toby Hocking · 2022
Paper
A Bayesian take on option pricing with Gaussian processes
Martin Tegner, Stephen Roberts · 2021
PDF
A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk · 2021
PDF
Assessing Developer Expertise from the Statistical Distribution of Programming Syntax Patterns
Arghavan Moradi Dakhel, Michel C. Desmarais, Foutse Khomh · 2021
Paper
Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Michael M. Bronstein, Joan Bruna, Taco Cohen, Petar Veličković · 2021
PDF
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen · 2021
PDF
Make your database system dream of electric sheep: towards self-driving operation
Andrew Pavlo, Matthew Butrovich, Lin Ma, Prashanth Menon, Wan Shen Lim, Dana Van Aken, William Zhang · 2021
Paper
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell · 2021
Paper
RecSysOps: Best Practices for Operating a Large-Scale Recommender System
Mohammad Saberian, Justin Basilico · 2021
Paper
Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
Ramaravind Kommiya Mothilal, Divyat Mahajan, Chenhao Tan, Amit Sharma · 2021
PDF
TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands
Bradley T. Shapiro, Günter J. Hitsch, Anna E. Tuchman · 2021
Paper
What are the most important statistical ideas of the past 50 years?
Andrew Gelman, Aki Vehtari · 2021
PDF
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Alex Wang, Kyunghyun Cho, Mike Lewis · 2020
PDF
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi · 2020
PDF
Controlled division of cell-sized vesicles by low densities of membrane-bound proteins
Jan Steinkühler, Roland L. Knorr, Ziliang Zhao, Tripta Bhatia, Solveig M. Bartelt, Seraphine Wegner, Rumiana Dimova, Reinhard Lipowsky · 2020
Paper
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei · 2020
PDF
Learning from positive and unlabeled data: a survey
Jessa Bekker, Jesse Davis · 2020
Paper
Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task
Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, Ankur P. Parikh · 2020
PDF
Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics
Guido W. Imbens · 2020
PDF
The Future of Origin of Life Research: Bridging Decades-Old Divisions
Martina Preiner, Silke Asche, Sidney Becker, Holly C. Betts, Adrien Boniface, Eloi Camprubi, Kuhan Chandru, Valentina Erastova, Sriram G. Garg, Nozair Khawaja, Gladys Kostyrka, Rainer Machné, Giacomo Moggioli, Kamila B. Muchowska, Sinje Neukirchen, Benedikt Peter, Edith Pichlhöfer, Ádám Radványi, Daniele Rossetto, Annalena Salditt, Nicolas M. Schmelling, Filipa L. Sousa, Fernando D. K. Tria, Dániel Vörös, Joana C. Xavier · 2020
Paper
When Does Label Smoothing Help?
Rafael Müller, Simon Kornblith, Geoffrey Hinton · 2020
PDF
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Backprop as Functor: A compositional perspective on supervised learning
Brendan Fong, David I. Spivak, Rémy Tuyéras · 2019
PDF
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, Luke Zettlemoyer · 2019
PDF
Evaluating the Factual Consistency of Abstractive Text Summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher · 2019
PDF
Lenia - Biology of Artificial Life
Bert Wang-Chak Chan · 2019
PDF
Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, Bin Yu · 2019
PDF
Recommending what video to watch next: a multitask ranking system
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed Chi · 2019
Paper
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, Samuel Bowman · 2018
Paper
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network
Jared L. Katzman, Uri Shaham, Alexander Cloninger, Jonathan Bates, Tingting Jiang, Yuval Kluger · 2018
Paper
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger · 2018
PDF
Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Kyle Hundman, Valentino Constantinou, Christopher Laporte, Ian Colwell, Tom Soderstrom · 2018
Paper
Systems protobiology: origin of life in lipid catalytic networks
Doron Lancet, Raphael Zidovetzki, Omer Markovitch · 2018
Paper
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J. Foster, Matus Telgarsky · 2017
PDF
A Field Guide to Forward-Backward Splitting with a FASTA Implementation
Tom Goldstein, Christoph Studer, Richard Baraniuk · 2016
PDF
Group Equivariant Convolutional Networks
Taco S. Cohen, Max Welling · 2016
PDF
How a life-like system emerges from a simplistic particle motion law
Thomas Schmickl, Martin Stefanec, Karl Crailsheim · 2016
Paper
Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, Alan L. Yuille · 2016
PDF
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning · 2015
Paper
A Strategy for Origins of Life Research
Caleb Scharf, Nathaniel Virgo, H. James Cleaves, Masashi Aono, Nathanael Aubert-Kato, Arsev Aydinoglu, Ana Barahona, Laura M. Barge, Steven A. Benner, Martin Biehl, Ramon Brasser, Christopher J. Butch, Kuhan Chandru, Leroy Cronin, Sebastian Danielache, Jakob Fischer, John Hernlund, Piet Hut, Takashi Ikegami, Jun Kimura, Kensei Kobayashi, Carlos Mariscal, Shawn McGlynn, Brice Menard, Norman Packard, Robert Pascal, Juli Pereto, Sudha Rajamani, Lana Sinapayen, Eric Smith, Christopher Switzer, Ken Takai, Feng Tian, Yuichiro Ueno, Mary Voytek, Olaf Witkowski, Hikaru Yabuta · 2015
Paper
FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, James Philbin · 2015
PDF
Towards Open Set Deep Networks
Abhijit Bendale, Terrance Boult · 2015
PDF
Practical Lessons from Predicting Clicks on Ads at Facebook
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, Joaquin Quiñonero Candela · 2014
Paper
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
Solving big data challenges for enterprise application performance management
Tilmann Rabl, Sergio Gómez-Villamor, Mohammad Sadoghi, Victor Muntés-Mulero, Hans-Arno Jacobsen, Serge Mankovskii · 2012
Paper
Feature Selection with the Boruta Package
Miron B. Kursa, Witold R. Rudnicki · 2010
Paper
Prevolutionary dynamics and the origin of evolution
Martin A. Nowak, Hisashi Ohtsuki · 2008
Paper
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu · 2001
Paper
The use of the area under the ROC curve in the evaluation of machine learning algorithms
Andrew P. Bradley · 1997
Paper
Self-reproduction in cellular automata
Christopher G. Langton · 1984
Paper
Information Theory and Statistical Mechanics
E. T. Jaynes · 1957
Paper
Sutskever / Carmack
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
PDF scalinglanguage-models
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
The Annotated Transformer
Alexander Rush · 2018
Paper transformersattention
Relational Recurrent Neural Networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap · 2018
PDF relational-reasoningsequence-models
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · 2017
PDF transformerssequence-modeling
CS231n: Convolutional Neural Networks for Visual Recognition
Fei-Fei Li, Andrej Karpathy, Justin Johnson · 2017
Paper visioncnn
Kolmogorov Complexity and Algorithmic Randomness
A. Shen, V. A. Uspensky, N. Vereshchagin · 2017
PDF kolmogorov-complexityalgorithmic-randomness
Neural Message Passing for Quantum Chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, George E. Dahl · 2017
PDF graph-neural-networksmessage-passing
Pointer Networks
Oriol Vinyals, Meire Fortunato, Navdeep Jaitly · 2017
PDF sequence-modelingneural-networks
A Simple Neural Network Module for Relational Reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap · 2017
PDF relational-reasoningrepresentation
Variational Lossy Autoencoder
Ilya Sutskever, Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Pieter Abbeel · 2017
PDF generative-modelsrepresentation-learning
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, Zhenyao Zhu · 2016
PDF sequence-modelsspeech
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Order Matters: Sequence to Sequence for Sets
Oriol Vinyals, Samy Bengio, Manjunath Kudlur · 2016
PDF sequence-modelssets
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2015
PDF attentionsequence-models
Understanding LSTM Networks
Christopher Olah · 2015
Paper sequence-modelsrnn
The Unreasonable Effectiveness of Recurrent Neural Networks
Andrej Karpathy · 2015
Paper sequence-modelsrnn
Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton
Scott Aaronson, Sean M. Carroll, Lauren Ouellette · 2014
PDF complexitykolmogorov-complexity
Neural Turing Machines
Alex Graves, Greg Wayne, Ivo Danihelka · 2014
PDF memorysequence-models
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals · 2014
PDF sequence-modelsrnn
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2012
PDF visioncnn
The First Law of Complexodynamics
Scott Aaronson · 2011
Paper complexitythermodynamics
Machine Super Intelligence
Shane Legg · 2008
PDF artificial-general-intelligencetheory
A Tutorial Introduction to the Minimum Description Length Principle
Peter Grunwald · 2004
PDF mdlcompression
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
Geoffrey E. Hinton, Drew van Camp · 1993
PDF mdlcompression
Core AI
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2025
PDF tabularfoundation-models
World Modeling with Probabilistic Structure Integration
Klemen Kotar, Wanhee Lee, Rahul Venkatesh, Honglin Chen, Daniel Bear, Jared Watrous, Simon Kim, Khai Loong Aw, Lilian Naing Chen, Stefan Stojanov, Kevin Feigelis, Imran Thobani, Alex Durango, Khaled Jedoui, Atlas Kazemian, Dan Yamins · 2025
PDF world-modelsprobabilistic-deep-learning
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng Long · 2022
PDF time-seriesanomaly-detection
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, James Glass · 2021
PDF audiotransformers
RealFormer: Transformer Likes Residual Attention
Ruining He, Anirudh Ravula, Bhargav Kanagal, Joshua Ainslie · 2021
PDF transformersattention
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
PDF scalinglanguage-models
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
The Annotated Transformer
Alexander Rush · 2018
Paper transformersattention
Deep One-Class Classification
Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Muller, Marius Kloft · 2018
PDF anomaly-detectiondeep-svdd
Relational Recurrent Neural Networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap · 2018
PDF relational-reasoningsequence-models
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · 2017
PDF transformerssequence-modeling
Pointer Networks
Oriol Vinyals, Meire Fortunato, Navdeep Jaitly · 2017
PDF sequence-modelingneural-networks
A Simple Neural Network Module for Relational Reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap · 2017
PDF relational-reasoningrepresentation
Variational Lossy Autoencoder
Ilya Sutskever, Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Pieter Abbeel · 2017
PDF generative-modelsrepresentation-learning
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, Zhenyao Zhu · 2016
PDF sequence-modelsspeech
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Order Matters: Sequence to Sequence for Sets
Oriol Vinyals, Samy Bengio, Manjunath Kudlur · 2016
PDF sequence-modelssets
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2015
PDF attentionsequence-models
Understanding LSTM Networks
Christopher Olah · 2015
Paper sequence-modelsrnn
The Unreasonable Effectiveness of Recurrent Neural Networks
Andrej Karpathy · 2015
Paper sequence-modelsrnn
Neural Turing Machines
Alex Graves, Greg Wayne, Ivo Danihelka · 2014
PDF memorysequence-models
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals · 2014
PDF sequence-modelsrnn
Uncertainty & Evidence
Generative Neural Operators through Diffusion Last Layer
Sungwon Park, Anthony Zhou, Hongjoong Kim, Amir Barati Farimani · 2026
PDF neural-operatorsdiffusion
Hypothesis Testing with E-Values
Aaditya Ramdas, Ruodu Wang · 2025
PDF statisticse-values
Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
Haoyi Song, Ruihan Ji, Naichen Shi, Fan Lai, Raed Al Kontar · 2025
PDF uncertaintyprobabilistic-deep-learning
Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil · 2025
PDF uncertaintylanguage-models
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
Variational Inference: A Review for Statisticians
David M. Blei, Alp Kucukelbir, Jon D. McAuliffe · 2017
PDF variational-inferencebayesian
How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
Christopher C. Strelioff, James P. Crutchfield · 2006
PDF bayesian-inferencedynamical-systems
Agency & RL
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
Maximum Likelihood Reinforcement Learning
Yiding Jiang, Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette · 2026
PDF reinforcement-learningmaximum-likelihood
Time, Identity and Consciousness in Language Model Agents
Michael Timothy Bennett, Elija Perrier · 2026
PDF consciousnesslanguage-models
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
Vicente Raja, Dinesh Valluri, Edward Baggs, Anthony Chemero, Michael L. Anderson · 2021
Paper active-inferencesystems
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
PDF reinforcement-learningprobabilistic-inference
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine · 2018
PDF reinforcement-learningmaximum-entropy
Active Inference: A Process Theory
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, Giovanni Pezzulo · 2017
PDF active-inferencefree-energy
VIME: Variational Information Maximizing Exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, Pieter Abbeel · 2016
PDF reinforcement-learningexploration
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
Efficient Computation of Optimal Actions
Emanuel Todorov · 2009
PDF controlreinforcement-learning
AI Consciousness
A Mind Cannot Be Smeared Across Time
Michael Timothy Bennett · 2026
PDF consciousnessai-agency
Time, Identity and Consciousness in Language Model Agents
Michael Timothy Bennett, Elija Perrier · 2026
PDF consciousnesslanguage-models
Why Is Anything Conscious?
Michael Timothy Bennett, Sean Welsh, Anna Ciaunica · 2026
PDF consciousnessai-agency
Learning
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Muon: An optimizer for hidden layers in neural networks
Keller Jordan · 2024
Paper optimizationtraining
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Cyclical Learning Rates for Training Neural Networks
Leslie N. Smith · 2017
PDF optimizationtraining
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
Applied Modeling
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2025
PDF tabularfoundation-models
WTNN: Weibull-Tailored Neural Networks for Survival Analysis
Gabrielle Rives, Olivier Lopez, Nicolas Bousquet · 2025
PDF survivalwtte
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Scholkopf, Thomas Brox, Peter Gehler · 2022
Paper industrial-visionanomaly-detection
Deep Cox Mixtures for Survival Regression
Chirag Nagpal, Steve Yadlowsky, Negar Rostamzadeh, Katherine Heller · 2021
PDF survivaltime-to-event
Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing Risks
Chirag Nagpal, Xinyu Li, Artur Dubrawski · 2021
Paper survivalcompeting-risks
Tabular Data: Deep Learning Is Not All You Need
Ravid Shwartz-Ziv, Amitai Armon · 2021
PDF tabularbaselines
Estimation of Conditional Mixture Weibull Distribution with Right-Censored Data Using Neural Network for Time-to-Event Analysis
Achraf Bennis, Sandrine Mouysset, Mathieu Serrurier · 2020
PDF survivalwtte
Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Changhee Lee, Jinsung Yoon, Mihaela van der Schaar · 2020
Paper survivalcompeting-risks
Underspecification Presents Challenges for Credibility in Modern Machine Learning
D. Sculley, Alex Beutel, Zachary Nado, Xuezhi Wang, Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai · 2020
PDF systemsevaluation
Anomaly Detection Using One-Class Neural Networks
Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla · 2019
PDF anomaly-detectionvision
Reliability-Equivalent Field Reference Usage Level When Both Field Usage and Usage to Failure Are Random
Fengbin Sun · 2019
Paper reliabilityusage-modeling
Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, Dan Pei · 2019
Paper time-seriesanomaly-detection
DeepHit: A Deep Learning Approach to Survival Analysis With Competing Risks
Changhee Lee, William Zame, Jinsung Yoon, Mihaela van der Schaar · 2018
PDF survivalcompeting-risks
Deep Learning for Patient-Specific Kidney Graft Survival Analysis
Margaux Luck, Tristan Sylvain, Heloise Cardinal, Andrea Lodi, Yoshua Bengio · 2017
PDF survivalmedical-modeling
WTTE-RNN: Weibull Time To Event Recurrent Neural Network
Egil Martinsson · 2017
PDF survivalwtte
Random Survival Forests
Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, Michael S. Lauer · 2008
Paper survivalreliability
Vision & Anomaly
EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
Kilian Batzner, Lars Heckler, Rebecca König · 2024
PDF industrial-visionanomaly-detection
Multimodal Industrial Anomaly Detection via Hybrid Fusion
Yue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi, Yabiao Wang, Chengjie Wang · 2023
Paper industrial-visionanomaly-detection
WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, Onkar Dabeer · 2023
PDF industrial-visionanomaly-detection
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng Long · 2022
PDF time-seriesanomaly-detection
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Scholkopf, Thomas Brox, Peter Gehler · 2022
Paper industrial-visionanomaly-detection
DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj · 2021
PDF industrial-visionanomaly-detection
PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization
Thomas Defard, Aleksandr Setkov, Angelique Loesch, Romaric Audigier · 2021
Paper industrial-visionanomaly-detection
Anomaly Detection Using One-Class Neural Networks
Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla · 2019
PDF anomaly-detectionvision
MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, Carsten Steger · 2019
Paper industrial-visionanomaly-detection
Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, Dan Pei · 2019
Paper time-seriesanomaly-detection
Deep One-Class Classification
Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Muller, Marius Kloft · 2018
PDF anomaly-detectiondeep-svdd
Automatic Liver and Tumor Segmentation of CT and MRI Volumes Using Cascaded Fully Convolutional Neural Networks
Patrick Ferdinand Christ, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Felix Hofmann, Seyed-Ahmad Ahmadi, Felix Grun, Mohamed Ezzeldin A. Elshaera, Jana Lipkova, Sebastian Schlecht, Freba Ahmaddy, Melvin D. Anastasi, Georgios Kaissis, Julian Holch, Wieland Sommer, Rickmer Braren, Volker Heinemann, Bjoern Menze · 2017
PDF visionsegmentation
CS231n: Convolutional Neural Networks for Visual Recognition
Fei-Fei Li, Andrej Karpathy, Justin Johnson · 2017
Paper visioncnn
Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Patrick Ferdinand Christ, Mohamed Ezzeldin A. Elshaer, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Marco Armbruster, Felix Hofmann, Melvin D'Anastasi, Wieland H. Sommer, Seyed-Ahmad Ahmadi, Bjoern H. Menze · 2016
PDF visionsegmentation
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2012
PDF visioncnn
Suggested Next
EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
Kilian Batzner, Lars Heckler, Rebecca König · 2024
PDF industrial-visionanomaly-detection
WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, Onkar Dabeer · 2023
PDF industrial-visionanomaly-detection
Deep Cox Mixtures for Survival Regression
Chirag Nagpal, Steve Yadlowsky, Negar Rostamzadeh, Katherine Heller · 2021
PDF survivaltime-to-event
DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj · 2021
PDF industrial-visionanomaly-detection
Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Changhee Lee, Jinsung Yoon, Mihaela van der Schaar · 2020
Paper survivalcompeting-risks
WTTE-RNN: Weibull Time To Event Recurrent Neural Network
Egil Martinsson · 2017
PDF survivalwtte
Practice
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, Lora M. Aroyo · 2021
Paper systemsdata
The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
Vicente Raja, Dinesh Valluri, Edward Baggs, Anthony Chemero, Michael L. Anderson · 2021
Paper active-inferencesystems
Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Stella Biderman, Walter J. Scheirer · 2021
PDF evaluationresearch-practice
Tabular Data: Deep Learning Is Not All You Need
Ravid Shwartz-Ziv, Amitai Armon · 2021
PDF tabularbaselines
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Underspecification Presents Challenges for Credibility in Modern Machine Learning
D. Sculley, Alex Beutel, Zachary Nado, Xuezhi Wang, Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai · 2020
PDF systemsevaluation
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, Carsten Steger · 2019
Paper industrial-visionanomaly-detection
Wide & Deep Learning for Recommender Systems
Lichan Hong, Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Vihan Jain, Xiaobing Liu, Hemal Shah · 2016
PDF recommender-systemsproduction-ml
How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
Christopher C. Strelioff, James P. Crutchfield · 2006
PDF bayesian-inferencedynamical-systems
All
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Andrey A. Dukhovny, Andrey M. Lange · 2026
PDF
A unified perspective of Gaussian process approximation for differential equations
Mengwu Guo · 2026
PDF
Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston · 2026
PDF
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Hong Zhao · 2026
PDF
Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Giovanni Pezzulo, Michael Levin · 2026
PDF alifeintelligence
Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 2026
PDF
Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Kuo Gai, Shihua Zhang · 2026
PDF
E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
Ben Chugg, Aaditya Ramdas, Peter Grünwald · 2026
PDF
E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
Alexandra Sokolova, Vadim Sokolov · 2026
PDF
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou · 2026
PDF
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · 2026
PDF
Fast Score-Based Sampling via Log-Concave Reductions
M. J. Wainwright · 2026
PDF
From Approximation to Emergence: A Theory of Deep Learning
Zhilin Zhao · 2026
PDF
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
Marc Finzi, Shikai Qiu, Yiding Jiang, Pavel Izmailov, J. Zico Kolter, Andrew Gordon Wilson · 2026
PDF information-theoryintelligence
Functional Attention: From Pairwise Affinities to Functional Correspondences
Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers · 2026
PDF
Generative Neural Operators through Diffusion Last Layer
Sungwon Park, Anthony Zhou, Hongjoong Kim, Amir Barati Farimani · 2026
PDF neural-operatorsdiffusion
Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Peilin Liu, Ding-Xuan Zhou · 2026
PDF
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová · 2026
PDF
Lecture notes on random matrix theory: the results, the applications, and the analytical tools
Joseph W. Baron · 2026
PDF
Maximum Likelihood Reinforcement Learning
Yiding Jiang, Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette · 2026
PDF reinforcement-learningmaximum-likelihood
A Mind Cannot Be Smeared Across Time
Michael Timothy Bennett · 2026
PDF consciousnessai-agency
Predictable GRPO: A Closed-Form Model of Training Dynamics
Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta · 2026
PDF
Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
Zhenyu Liao, Michael W. Mahoney · 2026
PDF
Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Daniela Schkoda, Dominik Janzing · 2026
PDF
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
Pratyaksh Rao, Wancong Zhang, Randall Balestriero, Yann LeCun, Giuseppe Loianno · 2026
PDF
Statistical Properties of Training & Generalization
Itay Lavie, Noam Levi, Yonatan Kahn · 2026
PDF
Still: Amortized KV Cache Compaction in a Single Forward Pass
Charles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby · 2026
PDF
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
Haggai Roitman · 2026
PDF
Time, Identity and Consciousness in Language Model Agents
Michael Timothy Bennett, Elija Perrier · 2026
PDF consciousnesslanguage-models
Training Language Models via Neural Cellular Automata
Dan Lee, Seungwook Han, Akarsh Kumar, Pulkit Agrawal · 2026
PDF language-modelscellular-automata
Why Is Anything Conscious?
Michael Timothy Bennett, Sean Welsh, Anna Ciaunica · 2026
PDF consciousnessai-agency
Basic Category Theory
Tom Leinster · 2025
PDF
E-Values Expand the Scope of Conformal Prediction
Etienne Gauthier, Francis Bach, Michael I. Jordan · 2025
PDF
Hypothesis Testing with E-Values
Aaditya Ramdas, Ruodu Wang · 2025
PDF statisticse-values
Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
Haoyi Song, Ruihan Ji, Naichen Shi, Fan Lai, Raed Al Kontar · 2025
PDF uncertaintyprobabilistic-deep-learning
Nonparametric Treatment Effect Identification in School Choice
Jiafeng Chen · 2025
PDF
Prediction-Powered E-Values
Daniel Csillag, Claudio José Struchiner, Guilherme Tegoni Goedert · 2025
PDF
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2025
PDF tabularfoundation-models
Table Foundation Models: on knowledge pre-training for tabular learning
Myung Jun Kim, Félix Lefebvre, Gaëtan Brison, Alexandre Perez-Lebel, Gaël Varoquaux · 2025
PDF
Tensor Logic: The Language of AI
Pedro Domingos · 2025
PDF
Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil · 2025
PDF uncertaintylanguage-models
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas · 2025
PDF
World Modeling with Probabilistic Structure Integration
Klemen Kotar, Wanhee Lee, Rahul Venkatesh, Honglin Chen, Daniel Bear, Jared Watrous, Simon Kim, Khai Loong Aw, Lilian Naing Chen, Stefan Stojanov, Kevin Feigelis, Imran Thobani, Alex Durango, Khaled Jedoui, Atlas Kazemian, Dan Yamins · 2025
PDF world-modelsprobabilistic-deep-learning
WTNN: Weibull-Tailored Neural Networks for Survival Analysis
Gabrielle Rives, Olivier Lopez, Nicolas Bousquet · 2025
PDF survivalwtte
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Computational Life: How Well-formed, Self-replicating Programs Emerge from Simple Interaction
Blaise Aguera y Arcas, Jyrki Alakuijala, James Evans, Ben Laurie, Alexander Mordvintsev, Eyvind Niklasson, Ettore Randazzo, Luca Versari · 2024
PDF alifesimulation
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
DINOv2: Learning Robust Visual Features without Supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski · 2024
PDF
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · 2024
PDF
EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
Kilian Batzner, Lars Heckler, Rebecca König · 2024
PDF industrial-visionanomaly-detection
Muon: An optimizer for hidden layers in neural networks
Keller Jordan · 2024
Paper optimizationtraining
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković · 2024
PDF
Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
Kexin Zhang, Qingsong Wen, Chaoli Zhang, Rongyao Cai, Ming Jin, Yong Liu, James Y. Zhang, Yuxuan Liang, Guansong Pang, Dongjin Song, Shirui Pan · 2024
Paper
SGLang: Efficient Execution of Structured Language Model Programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng · 2024
PDF
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou · 2023
PDF
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica · 2023
PDF
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu · 2023
PDF
Multimodal Industrial Anomaly Detection via Hybrid Fusion
Yue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi, Yabiao Wang, Chengjie Wang · 2023
Paper industrial-visionanomaly-detection
Statistical Foundations of Prior-Data Fitted Networks
Thomas Nagler · 2023
PDF
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, Onkar Dabeer · 2023
PDF industrial-visionanomaly-detection
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie · 2022
PDF
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng Long · 2022
PDF time-seriesanomaly-detection
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · 2022
PDF
Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features
Florian Pargent, Florian Pfisterer, Janek Thomas, Bernd Bischl · 2022
PDF
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei · 2022
PDF
Scheduling Techniques for Liver Segmentation: ReduceLRonPlateau Vs OneCycleLR
Ayman Al-Kababji, Faycal Bensaali, Sarada Prasad Dakua · 2022
PDF
Survival Regression with Accelerated Failure Time Model in XGBoost
Avinash Barnwal, Hyunsu Cho, Toby Hocking · 2022
Paper
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Scholkopf, Thomas Brox, Peter Gehler · 2022
Paper industrial-visionanomaly-detection
A Bayesian take on option pricing with Gaussian processes
Martin Tegner, Stephen Roberts · 2021
PDF
A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk · 2021
PDF
Assessing Developer Expertise from the Statistical Distribution of Programming Syntax Patterns
Arghavan Moradi Dakhel, Michel C. Desmarais, Foutse Khomh · 2021
Paper
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, James Glass · 2021
PDF audiotransformers
Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, Lora M. Aroyo · 2021
Paper systemsdata
Deep Cox Mixtures for Survival Regression
Chirag Nagpal, Steve Yadlowsky, Negar Rostamzadeh, Katherine Heller · 2021
PDF survivaltime-to-event
Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing Risks
Chirag Nagpal, Xinyu Li, Artur Dubrawski · 2021
Paper survivalcompeting-risks
DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj · 2021
PDF industrial-visionanomaly-detection
Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Michael M. Bronstein, Joan Bruna, Taco Cohen, Petar Veličković · 2021
PDF
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen · 2021
PDF
Make your database system dream of electric sheep: towards self-driving operation
Andrew Pavlo, Matthew Butrovich, Lin Ma, Prashanth Menon, Wan Shen Lim, Dana Van Aken, William Zhang · 2021
Paper
The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
Vicente Raja, Dinesh Valluri, Edward Baggs, Anthony Chemero, Michael L. Anderson · 2021
Paper active-inferencesystems
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell · 2021
Paper
PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization
Thomas Defard, Aleksandr Setkov, Angelique Loesch, Romaric Audigier · 2021
Paper industrial-visionanomaly-detection
Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Stella Biderman, Walter J. Scheirer · 2021
PDF evaluationresearch-practice
RealFormer: Transformer Likes Residual Attention
Ruining He, Anirudh Ravula, Bhargav Kanagal, Joshua Ainslie · 2021
PDF transformersattention
RecSysOps: Best Practices for Operating a Large-Scale Recommender System
Mohammad Saberian, Justin Basilico · 2021
Paper
Tabular Data: Deep Learning Is Not All You Need
Ravid Shwartz-Ziv, Amitai Armon · 2021
PDF tabularbaselines
Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
Ramaravind Kommiya Mothilal, Divyat Mahajan, Chenhao Tan, Amit Sharma · 2021
PDF
TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands
Bradley T. Shapiro, Günter J. Hitsch, Anna E. Tuchman · 2021
Paper
What are the most important statistical ideas of the past 50 years?
Andrew Gelman, Aki Vehtari · 2021
PDF
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Alex Wang, Kyunghyun Cho, Mike Lewis · 2020
PDF
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi · 2020
PDF
Estimation of Conditional Mixture Weibull Distribution with Right-Censored Data Using Neural Network for Time-to-Event Analysis
Achraf Bennis, Sandrine Mouysset, Mathieu Serrurier · 2020
PDF survivalwtte
Controlled division of cell-sized vesicles by low densities of membrane-bound proteins
Jan Steinkühler, Roland L. Knorr, Ziliang Zhao, Tripta Bhatia, Solveig M. Bartelt, Seraphine Wegner, Rumiana Dimova, Reinhard Lipowsky · 2020
Paper
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Changhee Lee, Jinsung Yoon, Mihaela van der Schaar · 2020
Paper survivalcompeting-risks
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei · 2020
PDF
Learning from positive and unlabeled data: a survey
Jessa Bekker, Jesse Davis · 2020
Paper
Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task
Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, Ankur P. Parikh · 2020
PDF
Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics
Guido W. Imbens · 2020
PDF
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
PDF scalinglanguage-models
The Future of Origin of Life Research: Bridging Decades-Old Divisions
Martina Preiner, Silke Asche, Sidney Becker, Holly C. Betts, Adrien Boniface, Eloi Camprubi, Kuhan Chandru, Valentina Erastova, Sriram G. Garg, Nozair Khawaja, Gladys Kostyrka, Rainer Machné, Giacomo Moggioli, Kamila B. Muchowska, Sinje Neukirchen, Benedikt Peter, Edith Pichlhöfer, Ádám Radványi, Daniele Rossetto, Annalena Salditt, Nicolas M. Schmelling, Filipa L. Sousa, Fernando D. K. Tria, Dániel Vörös, Joana C. Xavier · 2020
Paper
Underspecification Presents Challenges for Credibility in Modern Machine Learning
D. Sculley, Alex Beutel, Zachary Nado, Xuezhi Wang, Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai · 2020
PDF systemsevaluation
When Does Label Smoothing Help?
Rafael Müller, Simon Kornblith, Geoffrey Hinton · 2020
PDF
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Anomaly Detection Using One-Class Neural Networks
Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla · 2019
PDF anomaly-detectionvision
Backprop as Functor: A compositional perspective on supervised learning
Brendan Fong, David I. Spivak, Rémy Tuyéras · 2019
PDF
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, Luke Zettlemoyer · 2019
PDF
Evaluating the Factual Consistency of Abstractive Text Summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher · 2019
PDF
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Lenia - Biology of Artificial Life
Bert Wang-Chak Chan · 2019
PDF
Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, Bin Yu · 2019
PDF
MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, Carsten Steger · 2019
Paper industrial-visionanomaly-detection
Recommending what video to watch next: a multitask ranking system
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed Chi · 2019
Paper
Reliability-Equivalent Field Reference Usage Level When Both Field Usage and Usage to Failure Are Random
Fengbin Sun · 2019
Paper reliabilityusage-modeling
Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, Dan Pei · 2019
Paper time-seriesanomaly-detection
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, Samuel Bowman · 2018
Paper
The Annotated Transformer
Alexander Rush · 2018
Paper transformersattention
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Deep One-Class Classification
Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Muller, Marius Kloft · 2018
PDF anomaly-detectiondeep-svdd
DeepHit: A Deep Learning Approach to Survival Analysis With Competing Risks
Changhee Lee, William Zame, Jinsung Yoon, Mihaela van der Schaar · 2018
PDF survivalcompeting-risks
DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network
Jared L. Katzman, Uri Shaham, Alexander Cloninger, Jonathan Bates, Tingting Jiang, Yuval Kluger · 2018
Paper
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger · 2018
PDF
Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Kyle Hundman, Valentino Constantinou, Christopher Laporte, Ian Colwell, Tom Soderstrom · 2018
Paper
Relational Recurrent Neural Networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap · 2018
PDF relational-reasoningsequence-models
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
PDF reinforcement-learningprobabilistic-inference
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine · 2018
PDF reinforcement-learningmaximum-entropy
Systems protobiology: origin of life in lipid catalytic networks
Doron Lancet, Raphael Zidovetzki, Omer Markovitch · 2018
Paper
Active Inference: A Process Theory
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, Giovanni Pezzulo · 2017
PDF active-inferencefree-energy
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · 2017
PDF transformerssequence-modeling
Automatic Liver and Tumor Segmentation of CT and MRI Volumes Using Cascaded Fully Convolutional Neural Networks
Patrick Ferdinand Christ, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Felix Hofmann, Seyed-Ahmad Ahmadi, Felix Grun, Mohamed Ezzeldin A. Elshaera, Jana Lipkova, Sebastian Schlecht, Freba Ahmaddy, Melvin D. Anastasi, Georgios Kaissis, Julian Holch, Wieland Sommer, Rickmer Braren, Volker Heinemann, Bjoern Menze · 2017
PDF visionsegmentation
CS231n: Convolutional Neural Networks for Visual Recognition
Fei-Fei Li, Andrej Karpathy, Justin Johnson · 2017
Paper visioncnn
Cyclical Learning Rates for Training Neural Networks
Leslie N. Smith · 2017
PDF optimizationtraining
Deep Learning for Patient-Specific Kidney Graft Survival Analysis
Margaux Luck, Tristan Sylvain, Heloise Cardinal, Andrea Lodi, Yoshua Bengio · 2017
PDF survivalmedical-modeling
Kolmogorov Complexity and Algorithmic Randomness
A. Shen, V. A. Uspensky, N. Vereshchagin · 2017
PDF kolmogorov-complexityalgorithmic-randomness
Neural Message Passing for Quantum Chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, George E. Dahl · 2017
PDF graph-neural-networksmessage-passing
Pointer Networks
Oriol Vinyals, Meire Fortunato, Navdeep Jaitly · 2017
PDF sequence-modelingneural-networks
A Simple Neural Network Module for Relational Reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap · 2017
PDF relational-reasoningrepresentation
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J. Foster, Matus Telgarsky · 2017
PDF
Variational Inference: A Review for Statisticians
David M. Blei, Alp Kucukelbir, Jon D. McAuliffe · 2017
PDF variational-inferencebayesian
Variational Lossy Autoencoder
Ilya Sutskever, Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Pieter Abbeel · 2017
PDF generative-modelsrepresentation-learning
WTTE-RNN: Weibull Time To Event Recurrent Neural Network
Egil Martinsson · 2017
PDF survivalwtte
A Field Guide to Forward-Backward Splitting with a FASTA Implementation
Tom Goldstein, Christoph Studer, Richard Baraniuk · 2016
PDF
Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Patrick Ferdinand Christ, Mohamed Ezzeldin A. Elshaer, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Marco Armbruster, Felix Hofmann, Melvin D'Anastasi, Wieland H. Sommer, Seyed-Ahmad Ahmadi, Bjoern H. Menze · 2016
PDF visionsegmentation
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, Zhenyao Zhu · 2016
PDF sequence-modelsspeech
Group Equivariant Convolutional Networks
Taco S. Cohen, Max Welling · 2016
PDF
How a life-like system emerges from a simplistic particle motion law
Thomas Schmickl, Martin Stefanec, Karl Crailsheim · 2016
Paper
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Order Matters: Sequence to Sequence for Sets
Oriol Vinyals, Samy Bengio, Manjunath Kudlur · 2016
PDF sequence-modelssets
Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, Alan L. Yuille · 2016
PDF
VIME: Variational Information Maximizing Exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, Pieter Abbeel · 2016
PDF reinforcement-learningexploration
Wide & Deep Learning for Recommender Systems
Lichan Hong, Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Vihan Jain, Xiaobing Liu, Hemal Shah · 2016
PDF recommender-systemsproduction-ml
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning · 2015
Paper
A Strategy for Origins of Life Research
Caleb Scharf, Nathaniel Virgo, H. James Cleaves, Masashi Aono, Nathanael Aubert-Kato, Arsev Aydinoglu, Ana Barahona, Laura M. Barge, Steven A. Benner, Martin Biehl, Ramon Brasser, Christopher J. Butch, Kuhan Chandru, Leroy Cronin, Sebastian Danielache, Jakob Fischer, John Hernlund, Piet Hut, Takashi Ikegami, Jun Kimura, Kensei Kobayashi, Carlos Mariscal, Shawn McGlynn, Brice Menard, Norman Packard, Robert Pascal, Juli Pereto, Sudha Rajamani, Lana Sinapayen, Eric Smith, Christopher Switzer, Ken Takai, Feng Tian, Yuichiro Ueno, Mary Voytek, Olaf Witkowski, Hikaru Yabuta · 2015
Paper
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, James Philbin · 2015
PDF
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2015
PDF attentionsequence-models
Towards Open Set Deep Networks
Abhijit Bendale, Terrance Boult · 2015
PDF
Understanding LSTM Networks
Christopher Olah · 2015
Paper sequence-modelsrnn
The Unreasonable Effectiveness of Recurrent Neural Networks
Andrej Karpathy · 2015
Paper sequence-modelsrnn
Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton
Scott Aaronson, Sean M. Carroll, Lauren Ouellette · 2014
PDF complexitykolmogorov-complexity
Neural Turing Machines
Alex Graves, Greg Wayne, Ivo Danihelka · 2014
PDF memorysequence-models
Practical Lessons from Predicting Clicks on Ads at Facebook
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, Joaquin Quiñonero Candela · 2014
Paper
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals · 2014
PDF sequence-modelsrnn
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2012
PDF visioncnn
Solving big data challenges for enterprise application performance management
Tilmann Rabl, Sergio Gómez-Villamor, Mohammad Sadoghi, Victor Muntés-Mulero, Hans-Arno Jacobsen, Serge Mankovskii · 2012
Paper
The First Law of Complexodynamics
Scott Aaronson · 2011
Paper complexitythermodynamics
Feature Selection with the Boruta Package
Miron B. Kursa, Witold R. Rudnicki · 2010
Paper
Efficient Computation of Optimal Actions
Emanuel Todorov · 2009
PDF controlreinforcement-learning
Machine Super Intelligence
Shane Legg · 2008
PDF artificial-general-intelligencetheory
Prevolutionary dynamics and the origin of evolution
Martin A. Nowak, Hisashi Ohtsuki · 2008
Paper
Random Survival Forests
Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, Michael S. Lauer · 2008
Paper survivalreliability
How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
Christopher C. Strelioff, James P. Crutchfield · 2006
PDF bayesian-inferencedynamical-systems
A Tutorial Introduction to the Minimum Description Length Principle
Peter Grunwald · 2004
PDF mdlcompression
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu · 2001
Paper
The use of the area under the ROC curve in the evaluation of machine learning algorithms
Andrew P. Bradley · 1997
Paper
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
Geoffrey E. Hinton, Drew van Camp · 1993
PDF mdlcompression
Self-reproduction in cellular automata
Christopher G. Langton · 1984
Paper
Information Theory and Statistical Mechanics
E. T. Jaynes · 1957
Paper