Papers

References and working notes around ML systems, evaluation, reliability, artificial life, uncertainty, and applied modeling.

Recent

Thumbnail: A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
Thumbnail: A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Andrey A. Dukhovny, Andrey M. Lange · 2026
PDF
Thumbnail: A unified perspective of Gaussian process approximation for differential equations
A unified perspective of Gaussian process approximation for differential equations
Mengwu Guo · 2026
PDF
Thumbnail: Autodata: An agentic data scientist to create high quality synthetic data
Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston · 2026
PDF
Thumbnail: Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Hong Zhao · 2026
PDF
Thumbnail: Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Giovanni Pezzulo, Michael Levin · 2026
PDF alifeintelligence
Thumbnail: Conditional Inference Trees and Forests for Feature Selection
Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 2026
PDF
Thumbnail: Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Kuo Gai, Shihua Zhang · 2026
PDF
Thumbnail: E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
Ben Chugg, Aaditya Ramdas, Peter Grünwald · 2026
PDF
Thumbnail: E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
Alexandra Sokolova, Vadim Sokolov · 2026
PDF
Thumbnail: Experience Augmented Policy Optimization for LLM Reasoning
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou · 2026
PDF
Thumbnail: Fast KV Compaction via Attention Matching
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · 2026
PDF

Recent Reading

A tighter selection from my paper library, with links out to the source paper when available.

Thumbnail: A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
Thumbnail: A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Andrey A. Dukhovny, Andrey M. Lange · 2026
PDF
Thumbnail: A unified perspective of Gaussian process approximation for differential equations
A unified perspective of Gaussian process approximation for differential equations
Mengwu Guo · 2026
PDF
Thumbnail: Autodata: An agentic data scientist to create high quality synthetic data
Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston · 2026
PDF
Thumbnail: Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Hong Zhao · 2026
PDF
Thumbnail: Conditional Inference Trees and Forests for Feature Selection
Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 2026
PDF
Thumbnail: Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Kuo Gai, Shihua Zhang · 2026
PDF
Thumbnail: E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
Ben Chugg, Aaditya Ramdas, Peter Grünwald · 2026
PDF
Thumbnail: E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
Alexandra Sokolova, Vadim Sokolov · 2026
PDF
Thumbnail: Experience Augmented Policy Optimization for LLM Reasoning
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou · 2026
PDF
Thumbnail: Fast KV Compaction via Attention Matching
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · 2026
PDF
Thumbnail: Fast Score-Based Sampling via Log-Concave Reductions
Fast Score-Based Sampling via Log-Concave Reductions
M. J. Wainwright · 2026
PDF
Thumbnail: From Approximation to Emergence: A Theory of Deep Learning
From Approximation to Emergence: A Theory of Deep Learning
Zhilin Zhao · 2026
PDF
Thumbnail: Functional Attention: From Pairwise Affinities to Functional Correspondences
Functional Attention: From Pairwise Affinities to Functional Correspondences
Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers · 2026
PDF
Thumbnail: Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Peilin Liu, Ding-Xuan Zhou · 2026
PDF
Thumbnail: How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová · 2026
PDF
Thumbnail: Lecture notes on random matrix theory: the results, the applications, and the analytical tools
Lecture notes on random matrix theory: the results, the applications, and the analytical tools
Joseph W. Baron · 2026
PDF
Thumbnail: Predictable GRPO: A Closed-Form Model of Training Dynamics
Predictable GRPO: A Closed-Form Model of Training Dynamics
Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta · 2026
PDF
Thumbnail: Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
Zhenyu Liao, Michael W. Mahoney · 2026
PDF
Thumbnail: Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Daniela Schkoda, Dominik Janzing · 2026
PDF
Thumbnail: SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
Pratyaksh Rao, Wancong Zhang, Randall Balestriero, Yann LeCun, Giuseppe Loianno · 2026
PDF
Thumbnail: Statistical Properties of Training & Generalization
Statistical Properties of Training & Generalization
Itay Lavie, Noam Levi, Yonatan Kahn · 2026
PDF
Thumbnail: Still: Amortized KV Cache Compaction in a Single Forward Pass
Still: Amortized KV Cache Compaction in a Single Forward Pass
Charles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby · 2026
PDF
Thumbnail: The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
Haggai Roitman · 2026
PDF
Thumbnail: Basic Category Theory
Basic Category Theory
Tom Leinster · 2025
PDF
Thumbnail: E-Values Expand the Scope of Conformal Prediction
E-Values Expand the Scope of Conformal Prediction
Etienne Gauthier, Francis Bach, Michael I. Jordan · 2025
PDF
Thumbnail: Nonparametric Treatment Effect Identification in School Choice
Nonparametric Treatment Effect Identification in School Choice
Jiafeng Chen · 2025
PDF
Thumbnail: Prediction-Powered E-Values
Prediction-Powered E-Values
Daniel Csillag, Claudio José Struchiner, Guilherme Tegoni Goedert · 2025
PDF
Thumbnail: Table Foundation Models: on knowledge pre-training for tabular learning
Table Foundation Models: on knowledge pre-training for tabular learning
Myung Jun Kim, Félix Lefebvre, Gaëtan Brison, Alexandre Perez-Lebel, Gaël Varoquaux · 2025
PDF
Thumbnail: Tensor Logic: The Language of AI
Tensor Logic: The Language of AI
Pedro Domingos · 2025
PDF
Thumbnail: V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas · 2025
PDF
Thumbnail: An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: DINOv2: Learning Robust Visual Features without Supervision
DINOv2: Learning Robust Visual Features without Supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski · 2024
PDF
Thumbnail: Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · 2024
PDF
Thumbnail: Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković · 2024
PDF
Thumbnail: Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
Kexin Zhang, Qingsong Wen, Chaoli Zhang, Rongyao Cai, Ming Jin, Yong Liu, James Y. Zhang, Yuxuan Liang, Guansong Pang, Dongjin Song, Shirui Pan · 2024
Paper
Thumbnail: SGLang: Efficient Execution of Structured Language Model Programs
SGLang: Efficient Execution of Structured Language Model Programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng · 2024
PDF
Thumbnail: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou · 2023
PDF
Thumbnail: Efficient Memory Management for Large Language Model Serving with PagedAttention
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica · 2023
PDF
Thumbnail: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu · 2023
PDF
Thumbnail: Statistical Foundations of Prior-Data Fitted Networks
Statistical Foundations of Prior-Data Fitted Networks
Thomas Nagler · 2023
PDF
Thumbnail: A ConvNet for the 2020s
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie · 2022
PDF
Thumbnail: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · 2022
PDF
Thumbnail: Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features
Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features
Florian Pargent, Florian Pfisterer, Janek Thomas, Bernd Bischl · 2022
PDF
Thumbnail: Scaling Instruction-Finetuned Language Models
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei · 2022
PDF
Thumbnail: Scheduling Techniques for Liver Segmentation: ReduceLRonPlateau Vs OneCycleLR
Scheduling Techniques for Liver Segmentation: ReduceLRonPlateau Vs OneCycleLR
Ayman Al-Kababji, Faycal Bensaali, Sarada Prasad Dakua · 2022
PDF
Thumbnail: Survival Regression with Accelerated Failure Time Model in XGBoost
Survival Regression with Accelerated Failure Time Model in XGBoost
Avinash Barnwal, Hyunsu Cho, Toby Hocking · 2022
Paper
Thumbnail: A Bayesian take on option pricing with Gaussian processes
A Bayesian take on option pricing with Gaussian processes
Martin Tegner, Stephen Roberts · 2021
PDF
Thumbnail: A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk · 2021
PDF
Thumbnail: Assessing Developer Expertise from the Statistical Distribution of Programming Syntax Patterns
Assessing Developer Expertise from the Statistical Distribution of Programming Syntax Patterns
Arghavan Moradi Dakhel, Michel C. Desmarais, Foutse Khomh · 2021
Paper
Thumbnail: Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Michael M. Bronstein, Joan Bruna, Taco Cohen, Petar Veličković · 2021
PDF
Thumbnail: LoRA: Low-Rank Adaptation of Large Language Models
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen · 2021
PDF
Thumbnail: Make your database system dream of electric sheep: towards self-driving operation
Make your database system dream of electric sheep: towards self-driving operation
Andrew Pavlo, Matthew Butrovich, Lin Ma, Prashanth Menon, Wan Shen Lim, Dana Van Aken, William Zhang · 2021
Paper
Thumbnail: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell · 2021
Paper
Thumbnail: RecSysOps: Best Practices for Operating a Large-Scale Recommender System
RecSysOps: Best Practices for Operating a Large-Scale Recommender System
Mohammad Saberian, Justin Basilico · 2021
Paper
Thumbnail: Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
Ramaravind Kommiya Mothilal, Divyat Mahajan, Chenhao Tan, Amit Sharma · 2021
PDF
Thumbnail: TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands
TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands
Bradley T. Shapiro, Günter J. Hitsch, Anna E. Tuchman · 2021
Paper
Thumbnail: What are the most important statistical ideas of the past 50 years?
What are the most important statistical ideas of the past 50 years?
Andrew Gelman, Aki Vehtari · 2021
PDF
Thumbnail: Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Alex Wang, Kyunghyun Cho, Mike Lewis · 2020
PDF
Thumbnail: BERTScore: Evaluating Text Generation with BERT
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi · 2020
PDF
Thumbnail: Controlled division of cell-sized vesicles by low densities of membrane-bound proteins
Controlled division of cell-sized vesicles by low densities of membrane-bound proteins
Jan Steinkühler, Roland L. Knorr, Ziliang Zhao, Tripta Bhatia, Solveig M. Bartelt, Seraphine Wegner, Rumiana Dimova, Reinhard Lipowsky · 2020
Paper
Thumbnail: Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Thumbnail: Language Models are Few-Shot Learners
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei · 2020
PDF
Thumbnail: Learning from positive and unlabeled data: a survey
Learning from positive and unlabeled data: a survey
Jessa Bekker, Jesse Davis · 2020
Paper
Thumbnail: Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task
Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task
Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, Ankur P. Parikh · 2020
PDF
Thumbnail: Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics
Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics
Guido W. Imbens · 2020
PDF
Thumbnail: The Future of Origin of Life Research: Bridging Decades-Old Divisions
The Future of Origin of Life Research: Bridging Decades-Old Divisions
Martina Preiner, Silke Asche, Sidney Becker, Holly C. Betts, Adrien Boniface, Eloi Camprubi, Kuhan Chandru, Valentina Erastova, Sriram G. Garg, Nozair Khawaja, Gladys Kostyrka, Rainer Machné, Giacomo Moggioli, Kamila B. Muchowska, Sinje Neukirchen, Benedikt Peter, Edith Pichlhöfer, Ádám Radványi, Daniele Rossetto, Annalena Salditt, Nicolas M. Schmelling, Filipa L. Sousa, Fernando D. K. Tria, Dániel Vörös, Joana C. Xavier · 2020
Paper
Thumbnail: When Does Label Smoothing Help?
When Does Label Smoothing Help?
Rafael Müller, Simon Kornblith, Geoffrey Hinton · 2020
PDF
Thumbnail: A general system of differential equations to model first order adaptive algorithms
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Thumbnail: Backprop as Functor: A compositional perspective on supervised learning
Backprop as Functor: A compositional perspective on supervised learning
Brendan Fong, David I. Spivak, Rémy Tuyéras · 2019
PDF
Thumbnail: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, Luke Zettlemoyer · 2019
PDF
Thumbnail: Evaluating the Factual Consistency of Abstractive Text Summarization
Evaluating the Factual Consistency of Abstractive Text Summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher · 2019
PDF
Thumbnail: Lenia - Biology of Artificial Life
Lenia - Biology of Artificial Life
Bert Wang-Chak Chan · 2019
PDF
Thumbnail: Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, Bin Yu · 2019
PDF
Thumbnail: Recommending what video to watch next: a multitask ranking system
Recommending what video to watch next: a multitask ranking system
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed Chi · 2019
Paper
Thumbnail: A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, Samuel Bowman · 2018
Paper
Thumbnail: Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Thumbnail: DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network
DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network
Jared L. Katzman, Uri Shaham, Alexander Cloninger, Jonathan Bates, Tingting Jiang, Yuval Kluger · 2018
Paper
Thumbnail: Densely Connected Convolutional Networks
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger · 2018
PDF
Thumbnail: Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Kyle Hundman, Valentino Constantinou, Christopher Laporte, Ian Colwell, Tom Soderstrom · 2018
Paper
Thumbnail: Systems protobiology: origin of life in lipid catalytic networks
Systems protobiology: origin of life in lipid catalytic networks
Doron Lancet, Raphael Zidovetzki, Omer Markovitch · 2018
Paper
Thumbnail: Spectrally-normalized margin bounds for neural networks
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J. Foster, Matus Telgarsky · 2017
PDF
Thumbnail: A Field Guide to Forward-Backward Splitting with a FASTA Implementation
A Field Guide to Forward-Backward Splitting with a FASTA Implementation
Tom Goldstein, Christoph Studer, Richard Baraniuk · 2016
PDF
Thumbnail: Group Equivariant Convolutional Networks
Group Equivariant Convolutional Networks
Taco S. Cohen, Max Welling · 2016
PDF
Thumbnail: How a life-like system emerges from a simplistic particle motion law
How a life-like system emerges from a simplistic particle motion law
Thomas Schmickl, Martin Stefanec, Karl Crailsheim · 2016
Paper
Thumbnail: Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, Alan L. Yuille · 2016
PDF
Thumbnail: A large annotated corpus for learning natural language inference
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning · 2015
Paper
Thumbnail: A Strategy for Origins of Life Research
A Strategy for Origins of Life Research
Caleb Scharf, Nathaniel Virgo, H. James Cleaves, Masashi Aono, Nathanael Aubert-Kato, Arsev Aydinoglu, Ana Barahona, Laura M. Barge, Steven A. Benner, Martin Biehl, Ramon Brasser, Christopher J. Butch, Kuhan Chandru, Leroy Cronin, Sebastian Danielache, Jakob Fischer, John Hernlund, Piet Hut, Takashi Ikegami, Jun Kimura, Kensei Kobayashi, Carlos Mariscal, Shawn McGlynn, Brice Menard, Norman Packard, Robert Pascal, Juli Pereto, Sudha Rajamani, Lana Sinapayen, Eric Smith, Christopher Switzer, Ken Takai, Feng Tian, Yuichiro Ueno, Mary Voytek, Olaf Witkowski, Hikaru Yabuta · 2015
Paper
Thumbnail: FaceNet: A Unified Embedding for Face Recognition and Clustering
FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, James Philbin · 2015
PDF
Thumbnail: Towards Open Set Deep Networks
Towards Open Set Deep Networks
Abhijit Bendale, Terrance Boult · 2015
PDF
Thumbnail: Practical Lessons from Predicting Clicks on Ads at Facebook
Practical Lessons from Predicting Clicks on Ads at Facebook
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, Joaquin Quiñonero Candela · 2014
Paper
Thumbnail: Thermodynamics as a theory of decision-making with information processing costs
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
Thumbnail: Solving big data challenges for enterprise application performance management
Solving big data challenges for enterprise application performance management
Tilmann Rabl, Sergio Gómez-Villamor, Mohammad Sadoghi, Victor Muntés-Mulero, Hans-Arno Jacobsen, Serge Mankovskii · 2012
Paper
Thumbnail: Feature Selection with the Boruta Package
Feature Selection with the Boruta Package
Miron B. Kursa, Witold R. Rudnicki · 2010
Paper
Thumbnail: Prevolutionary dynamics and the origin of evolution
Prevolutionary dynamics and the origin of evolution
Martin A. Nowak, Hisashi Ohtsuki · 2008
Paper
Thumbnail: BLEU: a method for automatic evaluation of machine translation
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu · 2001
Paper
Thumbnail: The use of the area under the ROC curve in the evaluation of machine learning algorithms
The use of the area under the ROC curve in the evaluation of machine learning algorithms
Andrew P. Bradley · 1997
Paper
Thumbnail: Self-reproduction in cellular automata
Self-reproduction in cellular automata
Christopher G. Langton · 1984
Paper
Thumbnail: Information Theory and Statistical Mechanics
Information Theory and Statistical Mechanics
E. T. Jaynes · 1957
Paper

Sutskever / Carmack

Thumbnail: Scaling Laws for Neural Language Models
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
PDF scalinglanguage-models
Thumbnail: GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Thumbnail: The Annotated Transformer
The Annotated Transformer
Alexander Rush · 2018
Paper transformersattention
Thumbnail: Relational Recurrent Neural Networks
Relational Recurrent Neural Networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap · 2018
PDF relational-reasoningsequence-models
Thumbnail: Attention Is All You Need
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · 2017
PDF transformerssequence-modeling
Thumbnail: CS231n: Convolutional Neural Networks for Visual Recognition
CS231n: Convolutional Neural Networks for Visual Recognition
Fei-Fei Li, Andrej Karpathy, Justin Johnson · 2017
Paper visioncnn
Thumbnail: Kolmogorov Complexity and Algorithmic Randomness
Kolmogorov Complexity and Algorithmic Randomness
A. Shen, V. A. Uspensky, N. Vereshchagin · 2017
PDF kolmogorov-complexityalgorithmic-randomness
Thumbnail: Neural Message Passing for Quantum Chemistry
Neural Message Passing for Quantum Chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, George E. Dahl · 2017
PDF graph-neural-networksmessage-passing
Thumbnail: Pointer Networks
Pointer Networks
Oriol Vinyals, Meire Fortunato, Navdeep Jaitly · 2017
PDF sequence-modelingneural-networks
Thumbnail: A Simple Neural Network Module for Relational Reasoning
A Simple Neural Network Module for Relational Reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap · 2017
PDF relational-reasoningrepresentation
Thumbnail: Variational Lossy Autoencoder
Variational Lossy Autoencoder
Ilya Sutskever, Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Pieter Abbeel · 2017
PDF generative-modelsrepresentation-learning
Thumbnail: Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, Zhenyao Zhu · 2016
PDF sequence-modelsspeech
Thumbnail: Identity Mappings in Deep Residual Networks
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Thumbnail: Multi-Scale Context Aggregation by Dilated Convolutions
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Thumbnail: Order Matters: Sequence to Sequence for Sets
Order Matters: Sequence to Sequence for Sets
Oriol Vinyals, Samy Bengio, Manjunath Kudlur · 2016
PDF sequence-modelssets
Thumbnail: Deep Residual Learning for Image Recognition
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
Thumbnail: Neural Machine Translation by Jointly Learning to Align and Translate
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2015
PDF attentionsequence-models
Thumbnail: Understanding LSTM Networks
Understanding LSTM Networks
Christopher Olah · 2015
Paper sequence-modelsrnn
Thumbnail: The Unreasonable Effectiveness of Recurrent Neural Networks
The Unreasonable Effectiveness of Recurrent Neural Networks
Andrej Karpathy · 2015
Paper sequence-modelsrnn
Thumbnail: Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton
Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton
Scott Aaronson, Sean M. Carroll, Lauren Ouellette · 2014
PDF complexitykolmogorov-complexity
Thumbnail: Neural Turing Machines
Neural Turing Machines
Alex Graves, Greg Wayne, Ivo Danihelka · 2014
PDF memorysequence-models
Thumbnail: Recurrent Neural Network Regularization
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals · 2014
PDF sequence-modelsrnn
Thumbnail: ImageNet Classification with Deep Convolutional Neural Networks
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2012
PDF visioncnn
Thumbnail: The First Law of Complexodynamics
The First Law of Complexodynamics
Scott Aaronson · 2011
Paper complexitythermodynamics
Thumbnail: Machine Super Intelligence
Machine Super Intelligence
Shane Legg · 2008
PDF artificial-general-intelligencetheory
Thumbnail: A Tutorial Introduction to the Minimum Description Length Principle
A Tutorial Introduction to the Minimum Description Length Principle
Peter Grunwald · 2004
PDF mdlcompression
Thumbnail: Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
Geoffrey E. Hinton, Drew van Camp · 1993
PDF mdlcompression

Core AI

Thumbnail: TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
Thumbnail: TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2025
PDF tabularfoundation-models
Thumbnail: World Modeling with Probabilistic Structure Integration
World Modeling with Probabilistic Structure Integration
Klemen Kotar, Wanhee Lee, Rahul Venkatesh, Honglin Chen, Daniel Bear, Jared Watrous, Simon Kim, Khai Loong Aw, Lilian Naing Chen, Stefan Stojanov, Kevin Feigelis, Imran Thobani, Alex Durango, Khaled Jedoui, Atlas Kazemian, Dan Yamins · 2025
PDF world-modelsprobabilistic-deep-learning
Thumbnail: TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
Thumbnail: Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng Long · 2022
PDF time-seriesanomaly-detection
Thumbnail: AST: Audio Spectrogram Transformer
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, James Glass · 2021
PDF audiotransformers
Thumbnail: RealFormer: Transformer Likes Residual Attention
RealFormer: Transformer Likes Residual Attention
Ruining He, Anirudh Ravula, Bhargav Kanagal, Joshua Ainslie · 2021
PDF transformersattention
Thumbnail: Scaling Laws for Neural Language Models
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
PDF scalinglanguage-models
Thumbnail: GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Thumbnail: The Annotated Transformer
The Annotated Transformer
Alexander Rush · 2018
Paper transformersattention
Thumbnail: Deep One-Class Classification
Deep One-Class Classification
Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Muller, Marius Kloft · 2018
PDF anomaly-detectiondeep-svdd
Thumbnail: Relational Recurrent Neural Networks
Relational Recurrent Neural Networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap · 2018
PDF relational-reasoningsequence-models
Thumbnail: Attention Is All You Need
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · 2017
PDF transformerssequence-modeling
Thumbnail: Pointer Networks
Pointer Networks
Oriol Vinyals, Meire Fortunato, Navdeep Jaitly · 2017
PDF sequence-modelingneural-networks
Thumbnail: A Simple Neural Network Module for Relational Reasoning
A Simple Neural Network Module for Relational Reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap · 2017
PDF relational-reasoningrepresentation
Thumbnail: Variational Lossy Autoencoder
Variational Lossy Autoencoder
Ilya Sutskever, Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Pieter Abbeel · 2017
PDF generative-modelsrepresentation-learning
Thumbnail: Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, Zhenyao Zhu · 2016
PDF sequence-modelsspeech
Thumbnail: Identity Mappings in Deep Residual Networks
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Thumbnail: Multi-Scale Context Aggregation by Dilated Convolutions
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Thumbnail: Order Matters: Sequence to Sequence for Sets
Order Matters: Sequence to Sequence for Sets
Oriol Vinyals, Samy Bengio, Manjunath Kudlur · 2016
PDF sequence-modelssets
Thumbnail: Neural Machine Translation by Jointly Learning to Align and Translate
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2015
PDF attentionsequence-models
Thumbnail: Understanding LSTM Networks
Understanding LSTM Networks
Christopher Olah · 2015
Paper sequence-modelsrnn
Thumbnail: The Unreasonable Effectiveness of Recurrent Neural Networks
The Unreasonable Effectiveness of Recurrent Neural Networks
Andrej Karpathy · 2015
Paper sequence-modelsrnn
Thumbnail: Neural Turing Machines
Neural Turing Machines
Alex Graves, Greg Wayne, Ivo Danihelka · 2014
PDF memorysequence-models
Thumbnail: Recurrent Neural Network Regularization
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals · 2014
PDF sequence-modelsrnn

Uncertainty & Evidence

Agency & RL

Thumbnail: A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
Thumbnail: Maximum Likelihood Reinforcement Learning
Maximum Likelihood Reinforcement Learning
Yiding Jiang, Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette · 2026
PDF reinforcement-learningmaximum-likelihood
Thumbnail: Time, Identity and Consciousness in Language Model Agents
Time, Identity and Consciousness in Language Model Agents
Michael Timothy Bennett, Elija Perrier · 2026
PDF consciousnesslanguage-models
Thumbnail: An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
Vicente Raja, Dinesh Valluri, Edward Baggs, Anthony Chemero, Michael L. Anderson · 2021
Paper active-inferencesystems
Thumbnail: Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Thumbnail: A general system of differential equations to model first order adaptive algorithms
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Thumbnail: Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Thumbnail: Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
PDF reinforcement-learningprobabilistic-inference
Thumbnail: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine · 2018
PDF reinforcement-learningmaximum-entropy
Thumbnail: Active Inference: A Process Theory
Active Inference: A Process Theory
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, Giovanni Pezzulo · 2017
PDF active-inferencefree-energy
Thumbnail: VIME: Variational Information Maximizing Exploration
VIME: Variational Information Maximizing Exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, Pieter Abbeel · 2016
PDF reinforcement-learningexploration
Thumbnail: Thermodynamics as a theory of decision-making with information processing costs
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
Thumbnail: Efficient Computation of Optimal Actions
Efficient Computation of Optimal Actions
Emanuel Todorov · 2009
PDF controlreinforcement-learning

AI Consciousness

Learning

Thumbnail: A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
Thumbnail: An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: Muon: An optimizer for hidden layers in neural networks
Muon: An optimizer for hidden layers in neural networks
Keller Jordan · 2024
Paper optimizationtraining
Thumbnail: Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Thumbnail: A general system of differential equations to model first order adaptive algorithms
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Thumbnail: GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Thumbnail: Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Thumbnail: Cyclical Learning Rates for Training Neural Networks
Cyclical Learning Rates for Training Neural Networks
Leslie N. Smith · 2017
PDF optimizationtraining
Thumbnail: Identity Mappings in Deep Residual Networks
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Thumbnail: Deep Residual Learning for Image Recognition
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
Thumbnail: Thermodynamics as a theory of decision-making with information processing costs
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF

Applied Modeling

Thumbnail: TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
Thumbnail: TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2025
PDF tabularfoundation-models
Thumbnail: WTNN: Weibull-Tailored Neural Networks for Survival Analysis
WTNN: Weibull-Tailored Neural Networks for Survival Analysis
Gabrielle Rives, Olivier Lopez, Nicolas Bousquet · 2025
PDF survivalwtte
Thumbnail: TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
Thumbnail: Towards Total Recall in Industrial Anomaly Detection
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Scholkopf, Thomas Brox, Peter Gehler · 2022
Paper industrial-visionanomaly-detection
Thumbnail: Deep Cox Mixtures for Survival Regression
Deep Cox Mixtures for Survival Regression
Chirag Nagpal, Steve Yadlowsky, Negar Rostamzadeh, Katherine Heller · 2021
PDF survivaltime-to-event
Thumbnail: Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing Risks
Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing Risks
Chirag Nagpal, Xinyu Li, Artur Dubrawski · 2021
Paper survivalcompeting-risks
Thumbnail: Tabular Data: Deep Learning Is Not All You Need
Tabular Data: Deep Learning Is Not All You Need
Ravid Shwartz-Ziv, Amitai Armon · 2021
PDF tabularbaselines
Thumbnail: Estimation of Conditional Mixture Weibull Distribution with Right-Censored Data Using Neural Network for Time-to-Event Analysis
Estimation of Conditional Mixture Weibull Distribution with Right-Censored Data Using Neural Network for Time-to-Event Analysis
Achraf Bennis, Sandrine Mouysset, Mathieu Serrurier · 2020
PDF survivalwtte
Thumbnail: Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Changhee Lee, Jinsung Yoon, Mihaela van der Schaar · 2020
Paper survivalcompeting-risks
Thumbnail: Underspecification Presents Challenges for Credibility in Modern Machine Learning
Underspecification Presents Challenges for Credibility in Modern Machine Learning
D. Sculley, Alex Beutel, Zachary Nado, Xuezhi Wang, Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai · 2020
PDF systemsevaluation
Thumbnail: Anomaly Detection Using One-Class Neural Networks
Anomaly Detection Using One-Class Neural Networks
Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla · 2019
PDF anomaly-detectionvision
Thumbnail: Reliability-Equivalent Field Reference Usage Level When Both Field Usage and Usage to Failure Are Random
Reliability-Equivalent Field Reference Usage Level When Both Field Usage and Usage to Failure Are Random
Fengbin Sun · 2019
Paper reliabilityusage-modeling
Thumbnail: Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, Dan Pei · 2019
Paper time-seriesanomaly-detection
Thumbnail: DeepHit: A Deep Learning Approach to Survival Analysis With Competing Risks
DeepHit: A Deep Learning Approach to Survival Analysis With Competing Risks
Changhee Lee, William Zame, Jinsung Yoon, Mihaela van der Schaar · 2018
PDF survivalcompeting-risks
Thumbnail: Deep Learning for Patient-Specific Kidney Graft Survival Analysis
Deep Learning for Patient-Specific Kidney Graft Survival Analysis
Margaux Luck, Tristan Sylvain, Heloise Cardinal, Andrea Lodi, Yoshua Bengio · 2017
PDF survivalmedical-modeling
Thumbnail: WTTE-RNN: Weibull Time To Event Recurrent Neural Network
WTTE-RNN: Weibull Time To Event Recurrent Neural Network
Egil Martinsson · 2017
PDF survivalwtte
Thumbnail: Random Survival Forests
Random Survival Forests
Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, Michael S. Lauer · 2008
Paper survivalreliability

Vision & Anomaly

Thumbnail: EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
Kilian Batzner, Lars Heckler, Rebecca König · 2024
PDF industrial-visionanomaly-detection
Thumbnail: Multimodal Industrial Anomaly Detection via Hybrid Fusion
Multimodal Industrial Anomaly Detection via Hybrid Fusion
Yue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi, Yabiao Wang, Chengjie Wang · 2023
Paper industrial-visionanomaly-detection
Thumbnail: WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, Onkar Dabeer · 2023
PDF industrial-visionanomaly-detection
Thumbnail: Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng Long · 2022
PDF time-seriesanomaly-detection
Thumbnail: Towards Total Recall in Industrial Anomaly Detection
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Scholkopf, Thomas Brox, Peter Gehler · 2022
Paper industrial-visionanomaly-detection
Thumbnail: DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj · 2021
PDF industrial-visionanomaly-detection
Thumbnail: PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization
PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization
Thomas Defard, Aleksandr Setkov, Angelique Loesch, Romaric Audigier · 2021
Paper industrial-visionanomaly-detection
Thumbnail: Anomaly Detection Using One-Class Neural Networks
Anomaly Detection Using One-Class Neural Networks
Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla · 2019
PDF anomaly-detectionvision
Thumbnail: MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, Carsten Steger · 2019
Paper industrial-visionanomaly-detection
Thumbnail: Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, Dan Pei · 2019
Paper time-seriesanomaly-detection
Thumbnail: Deep One-Class Classification
Deep One-Class Classification
Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Muller, Marius Kloft · 2018
PDF anomaly-detectiondeep-svdd
Thumbnail: Automatic Liver and Tumor Segmentation of CT and MRI Volumes Using Cascaded Fully Convolutional Neural Networks
Automatic Liver and Tumor Segmentation of CT and MRI Volumes Using Cascaded Fully Convolutional Neural Networks
Patrick Ferdinand Christ, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Felix Hofmann, Seyed-Ahmad Ahmadi, Felix Grun, Mohamed Ezzeldin A. Elshaera, Jana Lipkova, Sebastian Schlecht, Freba Ahmaddy, Melvin D. Anastasi, Georgios Kaissis, Julian Holch, Wieland Sommer, Rickmer Braren, Volker Heinemann, Bjoern Menze · 2017
PDF visionsegmentation
Thumbnail: CS231n: Convolutional Neural Networks for Visual Recognition
CS231n: Convolutional Neural Networks for Visual Recognition
Fei-Fei Li, Andrej Karpathy, Justin Johnson · 2017
Paper visioncnn
Thumbnail: Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Patrick Ferdinand Christ, Mohamed Ezzeldin A. Elshaer, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Marco Armbruster, Felix Hofmann, Melvin D'Anastasi, Wieland H. Sommer, Seyed-Ahmad Ahmadi, Bjoern H. Menze · 2016
PDF visionsegmentation
Thumbnail: Multi-Scale Context Aggregation by Dilated Convolutions
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Thumbnail: Deep Residual Learning for Image Recognition
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
Thumbnail: ImageNet Classification with Deep Convolutional Neural Networks
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2012
PDF visioncnn

Suggested Next

Practice

Thumbnail: TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
Thumbnail: Data Cascades in High-Stakes AI
Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, Lora M. Aroyo · 2021
Paper systemsdata
Thumbnail: The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
Vicente Raja, Dinesh Valluri, Edward Baggs, Anthony Chemero, Michael L. Anderson · 2021
Paper active-inferencesystems
Thumbnail: Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Stella Biderman, Walter J. Scheirer · 2021
PDF evaluationresearch-practice
Thumbnail: Tabular Data: Deep Learning Is Not All You Need
Tabular Data: Deep Learning Is Not All You Need
Ravid Shwartz-Ziv, Amitai Armon · 2021
PDF tabularbaselines
Thumbnail: Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Thumbnail: Underspecification Presents Challenges for Credibility in Modern Machine Learning
Underspecification Presents Challenges for Credibility in Modern Machine Learning
D. Sculley, Alex Beutel, Zachary Nado, Xuezhi Wang, Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai · 2020
PDF systemsevaluation
Thumbnail: A general system of differential equations to model first order adaptive algorithms
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Thumbnail: GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Thumbnail: MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, Carsten Steger · 2019
Paper industrial-visionanomaly-detection
Thumbnail: Wide & Deep Learning for Recommender Systems
Wide & Deep Learning for Recommender Systems
Lichan Hong, Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Vihan Jain, Xiaobing Liu, Hemal Shah · 2016
PDF recommender-systemsproduction-ml
Thumbnail: How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
Christopher C. Strelioff, James P. Crutchfield · 2006
PDF bayesian-inferencedynamical-systems

All

Thumbnail: A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna · 2026
PDF
Thumbnail: A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Andrey A. Dukhovny, Andrey M. Lange · 2026
PDF
Thumbnail: A unified perspective of Gaussian process approximation for differential equations
A unified perspective of Gaussian process approximation for differential equations
Mengwu Guo · 2026
PDF
Thumbnail: Autodata: An agentic data scientist to create high quality synthetic data
Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston · 2026
PDF
Thumbnail: Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Hong Zhao · 2026
PDF
Thumbnail: Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Giovanni Pezzulo, Michael Levin · 2026
PDF alifeintelligence
Thumbnail: Conditional Inference Trees and Forests for Feature Selection
Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 2026
PDF
Thumbnail: Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Kuo Gai, Shihua Zhang · 2026
PDF
Thumbnail: E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
Ben Chugg, Aaditya Ramdas, Peter Grünwald · 2026
PDF
Thumbnail: E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
E-values for Adaptive Clinical Trials: Anytime-Valid Monitoring in Practice
Alexandra Sokolova, Vadim Sokolov · 2026
PDF
Thumbnail: Experience Augmented Policy Optimization for LLM Reasoning
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou · 2026
PDF
Thumbnail: Fast KV Compaction via Attention Matching
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · 2026
PDF
Thumbnail: Fast Score-Based Sampling via Log-Concave Reductions
Fast Score-Based Sampling via Log-Concave Reductions
M. J. Wainwright · 2026
PDF
Thumbnail: From Approximation to Emergence: A Theory of Deep Learning
From Approximation to Emergence: A Theory of Deep Learning
Zhilin Zhao · 2026
PDF
Thumbnail: From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
Marc Finzi, Shikai Qiu, Yiding Jiang, Pavel Izmailov, J. Zico Kolter, Andrew Gordon Wilson · 2026
PDF information-theoryintelligence
Thumbnail: Functional Attention: From Pairwise Affinities to Functional Correspondences
Functional Attention: From Pairwise Affinities to Functional Correspondences
Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers · 2026
PDF
Thumbnail: Generative Neural Operators through Diffusion Last Layer
Generative Neural Operators through Diffusion Last Layer
Sungwon Park, Anthony Zhou, Hongjoong Kim, Amir Barati Farimani · 2026
PDF neural-operatorsdiffusion
Thumbnail: Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Peilin Liu, Ding-Xuan Zhou · 2026
PDF
Thumbnail: How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová · 2026
PDF
Thumbnail: Lecture notes on random matrix theory: the results, the applications, and the analytical tools
Lecture notes on random matrix theory: the results, the applications, and the analytical tools
Joseph W. Baron · 2026
PDF
Thumbnail: Maximum Likelihood Reinforcement Learning
Maximum Likelihood Reinforcement Learning
Yiding Jiang, Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette · 2026
PDF reinforcement-learningmaximum-likelihood
Thumbnail: A Mind Cannot Be Smeared Across Time
A Mind Cannot Be Smeared Across Time
Michael Timothy Bennett · 2026
PDF consciousnessai-agency
Thumbnail: Predictable GRPO: A Closed-Form Model of Training Dynamics
Predictable GRPO: A Closed-Form Model of Training Dynamics
Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta · 2026
PDF
Thumbnail: Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
Zhenyu Liao, Michael W. Mahoney · 2026
PDF
Thumbnail: Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Daniela Schkoda, Dominik Janzing · 2026
PDF
Thumbnail: SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
Pratyaksh Rao, Wancong Zhang, Randall Balestriero, Yann LeCun, Giuseppe Loianno · 2026
PDF
Thumbnail: Statistical Properties of Training & Generalization
Statistical Properties of Training & Generalization
Itay Lavie, Noam Levi, Yonatan Kahn · 2026
PDF
Thumbnail: Still: Amortized KV Cache Compaction in a Single Forward Pass
Still: Amortized KV Cache Compaction in a Single Forward Pass
Charles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby · 2026
PDF
Thumbnail: TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2026
PDF tabularfoundation-models
Thumbnail: The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
Haggai Roitman · 2026
PDF
Thumbnail: Time, Identity and Consciousness in Language Model Agents
Time, Identity and Consciousness in Language Model Agents
Michael Timothy Bennett, Elija Perrier · 2026
PDF consciousnesslanguage-models
Thumbnail: Training Language Models via Neural Cellular Automata
Training Language Models via Neural Cellular Automata
Dan Lee, Seungwook Han, Akarsh Kumar, Pulkit Agrawal · 2026
PDF language-modelscellular-automata
Thumbnail: Why Is Anything Conscious?
Why Is Anything Conscious?
Michael Timothy Bennett, Sean Welsh, Anna Ciaunica · 2026
PDF consciousnessai-agency
Thumbnail: Basic Category Theory
Basic Category Theory
Tom Leinster · 2025
PDF
Thumbnail: E-Values Expand the Scope of Conformal Prediction
E-Values Expand the Scope of Conformal Prediction
Etienne Gauthier, Francis Bach, Michael I. Jordan · 2025
PDF
Thumbnail: Hypothesis Testing with E-Values
Hypothesis Testing with E-Values
Aaditya Ramdas, Ruodu Wang · 2025
PDF statisticse-values
Thumbnail: Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
Haoyi Song, Ruihan Ji, Naichen Shi, Fan Lai, Raed Al Kontar · 2025
PDF uncertaintyprobabilistic-deep-learning
Thumbnail: Nonparametric Treatment Effect Identification in School Choice
Nonparametric Treatment Effect Identification in School Choice
Jiafeng Chen · 2025
PDF
Thumbnail: Prediction-Powered E-Values
Prediction-Powered E-Values
Daniel Csillag, Claudio José Struchiner, Guilherme Tegoni Goedert · 2025
PDF
Thumbnail: TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Gael Varoquaux, Jingang Qu, David Holzmuller, Marine Le Morvan · 2025
PDF tabularfoundation-models
Thumbnail: Table Foundation Models: on knowledge pre-training for tabular learning
Table Foundation Models: on knowledge pre-training for tabular learning
Myung Jun Kim, Félix Lefebvre, Gaëtan Brison, Alexandre Perez-Lebel, Gaël Varoquaux · 2025
PDF
Thumbnail: Tensor Logic: The Language of AI
Tensor Logic: The Language of AI
Pedro Domingos · 2025
PDF
Thumbnail: Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil · 2025
PDF uncertaintylanguage-models
Thumbnail: V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas · 2025
PDF
Thumbnail: World Modeling with Probabilistic Structure Integration
World Modeling with Probabilistic Structure Integration
Klemen Kotar, Wanhee Lee, Rahul Venkatesh, Honglin Chen, Daniel Bear, Jared Watrous, Simon Kim, Khai Loong Aw, Lilian Naing Chen, Stefan Stojanov, Kevin Feigelis, Imran Thobani, Alex Durango, Khaled Jedoui, Atlas Kazemian, Dan Yamins · 2025
PDF world-modelsprobabilistic-deep-learning
Thumbnail: WTNN: Weibull-Tailored Neural Networks for Survival Analysis
WTNN: Weibull-Tailored Neural Networks for Survival Analysis
Gabrielle Rives, Olivier Lopez, Nicolas Bousquet · 2025
PDF survivalwtte
Thumbnail: An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
An Abstract Lyapunov Control Optimizer: Local Stabilization and Global Convergence
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: Computational Life: How Well-formed, Self-replicating Programs Emerge from Simple Interaction
Computational Life: How Well-formed, Self-replicating Programs Emerge from Simple Interaction
Blaise Aguera y Arcas, Jyrki Alakuijala, James Evans, Ben Laurie, Alexander Mordvintsev, Eyvind Niklasson, Ettore Randazzo, Luca Versari · 2024
PDF alifesimulation
Thumbnail: Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
Bilel Bensaid, Gaël Poëtte, Rodolphe Turpault · 2024
PDF
Thumbnail: DINOv2: Learning Robust Visual Features without Supervision
DINOv2: Learning Robust Visual Features without Supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski · 2024
PDF
Thumbnail: Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · 2024
PDF
Thumbnail: EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
Kilian Batzner, Lars Heckler, Rebecca König · 2024
PDF industrial-visionanomaly-detection
Thumbnail: Muon: An optimizer for hidden layers in neural networks
Muon: An optimizer for hidden layers in neural networks
Keller Jordan · 2024
Paper optimizationtraining
Thumbnail: Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković · 2024
PDF
Thumbnail: Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
Kexin Zhang, Qingsong Wen, Chaoli Zhang, Rongyao Cai, Ming Jin, Yong Liu, James Y. Zhang, Yuxuan Liang, Guansong Pang, Dongjin Song, Shirui Pan · 2024
Paper
Thumbnail: SGLang: Efficient Execution of Structured Language Model Programs
SGLang: Efficient Execution of Structured Language Model Programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng · 2024
PDF
Thumbnail: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou · 2023
PDF
Thumbnail: Efficient Memory Management for Large Language Model Serving with PagedAttention
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica · 2023
PDF
Thumbnail: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu · 2023
PDF
Thumbnail: Multimodal Industrial Anomaly Detection via Hybrid Fusion
Multimodal Industrial Anomaly Detection via Hybrid Fusion
Yue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi, Yabiao Wang, Chengjie Wang · 2023
Paper industrial-visionanomaly-detection
Thumbnail: Statistical Foundations of Prior-Data Fitted Networks
Statistical Foundations of Prior-Data Fitted Networks
Thomas Nagler · 2023
PDF
Thumbnail: TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter · 2023
PDF tabularfoundation-models
Thumbnail: WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, Onkar Dabeer · 2023
PDF industrial-visionanomaly-detection
Thumbnail: A ConvNet for the 2020s
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie · 2022
PDF
Thumbnail: Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng Long · 2022
PDF time-seriesanomaly-detection
Thumbnail: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · 2022
PDF
Thumbnail: Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features
Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features
Florian Pargent, Florian Pfisterer, Janek Thomas, Bernd Bischl · 2022
PDF
Thumbnail: Scaling Instruction-Finetuned Language Models
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei · 2022
PDF
Thumbnail: Scheduling Techniques for Liver Segmentation: ReduceLRonPlateau Vs OneCycleLR
Scheduling Techniques for Liver Segmentation: ReduceLRonPlateau Vs OneCycleLR
Ayman Al-Kababji, Faycal Bensaali, Sarada Prasad Dakua · 2022
PDF
Thumbnail: Survival Regression with Accelerated Failure Time Model in XGBoost
Survival Regression with Accelerated Failure Time Model in XGBoost
Avinash Barnwal, Hyunsu Cho, Toby Hocking · 2022
Paper
Thumbnail: Towards Total Recall in Industrial Anomaly Detection
Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Scholkopf, Thomas Brox, Peter Gehler · 2022
Paper industrial-visionanomaly-detection
Thumbnail: A Bayesian take on option pricing with Gaussian processes
A Bayesian take on option pricing with Gaussian processes
Martin Tegner, Stephen Roberts · 2021
PDF
Thumbnail: A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk · 2021
PDF
Thumbnail: Assessing Developer Expertise from the Statistical Distribution of Programming Syntax Patterns
Assessing Developer Expertise from the Statistical Distribution of Programming Syntax Patterns
Arghavan Moradi Dakhel, Michel C. Desmarais, Foutse Khomh · 2021
Paper
Thumbnail: AST: Audio Spectrogram Transformer
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, James Glass · 2021
PDF audiotransformers
Thumbnail: Data Cascades in High-Stakes AI
Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, Lora M. Aroyo · 2021
Paper systemsdata
Thumbnail: Deep Cox Mixtures for Survival Regression
Deep Cox Mixtures for Survival Regression
Chirag Nagpal, Steve Yadlowsky, Negar Rostamzadeh, Katherine Heller · 2021
PDF survivaltime-to-event
Thumbnail: Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing Risks
Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing Risks
Chirag Nagpal, Xinyu Li, Artur Dubrawski · 2021
Paper survivalcompeting-risks
Thumbnail: DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection
Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj · 2021
PDF industrial-visionanomaly-detection
Thumbnail: Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Michael M. Bronstein, Joan Bruna, Taco Cohen, Petar Veličković · 2021
PDF
Thumbnail: LoRA: Low-Rank Adaptation of Large Language Models
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen · 2021
PDF
Thumbnail: Make your database system dream of electric sheep: towards self-driving operation
Make your database system dream of electric sheep: towards self-driving operation
Andrew Pavlo, Matthew Butrovich, Lin Ma, Prashanth Menon, Wan Shen Lim, Dana Van Aken, William Zhang · 2021
Paper
Thumbnail: The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
The Markov Blanket Trick: On the Scope of the Free Energy Principle and Active Inference
Vicente Raja, Dinesh Valluri, Edward Baggs, Anthony Chemero, Michael L. Anderson · 2021
Paper active-inferencesystems
Thumbnail: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell · 2021
Paper
Thumbnail: PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization
PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization
Thomas Defard, Aleksandr Setkov, Angelique Loesch, Romaric Audigier · 2021
Paper industrial-visionanomaly-detection
Thumbnail: Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Stella Biderman, Walter J. Scheirer · 2021
PDF evaluationresearch-practice
Thumbnail: RealFormer: Transformer Likes Residual Attention
RealFormer: Transformer Likes Residual Attention
Ruining He, Anirudh Ravula, Bhargav Kanagal, Joshua Ainslie · 2021
PDF transformersattention
Thumbnail: RecSysOps: Best Practices for Operating a Large-Scale Recommender System
RecSysOps: Best Practices for Operating a Large-Scale Recommender System
Mohammad Saberian, Justin Basilico · 2021
Paper
Thumbnail: Tabular Data: Deep Learning Is Not All You Need
Tabular Data: Deep Learning Is Not All You Need
Ravid Shwartz-Ziv, Amitai Armon · 2021
PDF tabularbaselines
Thumbnail: Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
Ramaravind Kommiya Mothilal, Divyat Mahajan, Chenhao Tan, Amit Sharma · 2021
PDF
Thumbnail: TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands
TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands
Bradley T. Shapiro, Günter J. Hitsch, Anna E. Tuchman · 2021
Paper
Thumbnail: What are the most important statistical ideas of the past 50 years?
What are the most important statistical ideas of the past 50 years?
Andrew Gelman, Aki Vehtari · 2021
PDF
Thumbnail: Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Alex Wang, Kyunghyun Cho, Mike Lewis · 2020
PDF
Thumbnail: BERTScore: Evaluating Text Generation with BERT
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi · 2020
PDF
Thumbnail: Estimation of Conditional Mixture Weibull Distribution with Right-Censored Data Using Neural Network for Time-to-Event Analysis
Estimation of Conditional Mixture Weibull Distribution with Right-Censored Data Using Neural Network for Time-to-Event Analysis
Achraf Bennis, Sandrine Mouysset, Mathieu Serrurier · 2020
PDF survivalwtte
Thumbnail: Controlled division of cell-sized vesicles by low densities of membrane-bound proteins
Controlled division of cell-sized vesicles by low densities of membrane-bound proteins
Jan Steinkühler, Roland L. Knorr, Ziliang Zhao, Tripta Bhatia, Solveig M. Bartelt, Seraphine Wegner, Rumiana Dimova, Reinhard Lipowsky · 2020
Paper
Thumbnail: Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Convergence and Dynamical Behavior of the ADAM Algorithm for Non-Convex Stochastic Optimization
Anas Barakat, Pascal Bianchi · 2020
PDF
Thumbnail: Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Dynamic-DeepHit: A Deep Learning Approach for Dynamic Survival Analysis With Competing Risks Based on Longitudinal Data
Changhee Lee, Jinsung Yoon, Mihaela van der Schaar · 2020
Paper survivalcompeting-risks
Thumbnail: Language Models are Few-Shot Learners
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei · 2020
PDF
Thumbnail: Learning from positive and unlabeled data: a survey
Learning from positive and unlabeled data: a survey
Jessa Bekker, Jesse Davis · 2020
Paper
Thumbnail: Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task
Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task
Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, Ankur P. Parikh · 2020
PDF
Thumbnail: Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics
Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics
Guido W. Imbens · 2020
PDF
Thumbnail: Scaling Laws for Neural Language Models
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
PDF scalinglanguage-models
Thumbnail: The Future of Origin of Life Research: Bridging Decades-Old Divisions
The Future of Origin of Life Research: Bridging Decades-Old Divisions
Martina Preiner, Silke Asche, Sidney Becker, Holly C. Betts, Adrien Boniface, Eloi Camprubi, Kuhan Chandru, Valentina Erastova, Sriram G. Garg, Nozair Khawaja, Gladys Kostyrka, Rainer Machné, Giacomo Moggioli, Kamila B. Muchowska, Sinje Neukirchen, Benedikt Peter, Edith Pichlhöfer, Ádám Radványi, Daniele Rossetto, Annalena Salditt, Nicolas M. Schmelling, Filipa L. Sousa, Fernando D. K. Tria, Dániel Vörös, Joana C. Xavier · 2020
Paper
Thumbnail: Underspecification Presents Challenges for Credibility in Modern Machine Learning
Underspecification Presents Challenges for Credibility in Modern Machine Learning
D. Sculley, Alex Beutel, Zachary Nado, Xuezhi Wang, Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai · 2020
PDF systemsevaluation
Thumbnail: When Does Label Smoothing Help?
When Does Label Smoothing Help?
Rafael Müller, Simon Kornblith, Geoffrey Hinton · 2020
PDF
Thumbnail: A general system of differential equations to model first order adaptive algorithms
A general system of differential equations to model first order adaptive algorithms
André Belotto da Silva, Maxime Gazeau · 2019
PDF
Thumbnail: Anomaly Detection Using One-Class Neural Networks
Anomaly Detection Using One-Class Neural Networks
Raghavendra Chalapathy, Aditya Krishna Menon, Sanjay Chawla · 2019
PDF anomaly-detectionvision
Thumbnail: Backprop as Functor: A compositional perspective on supervised learning
Backprop as Functor: A compositional perspective on supervised learning
Brendan Fong, David I. Spivak, Rémy Tuyéras · 2019
PDF
Thumbnail: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, Luke Zettlemoyer · 2019
PDF
Thumbnail: Evaluating the Factual Consistency of Abstractive Text Summarization
Evaluating the Factual Consistency of Abstractive Text Summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher · 2019
PDF
Thumbnail: GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, Zhifeng Chen · 2019
PDF systemsscaling
Thumbnail: Lenia - Biology of Artificial Life
Lenia - Biology of Artificial Life
Bert Wang-Chak Chan · 2019
PDF
Thumbnail: Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, Bin Yu · 2019
PDF
Thumbnail: MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, Carsten Steger · 2019
Paper industrial-visionanomaly-detection
Thumbnail: Recommending what video to watch next: a multitask ranking system
Recommending what video to watch next: a multitask ranking system
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed Chi · 2019
Paper
Thumbnail: Reliability-Equivalent Field Reference Usage Level When Both Field Usage and Usage to Failure Are Random
Reliability-Equivalent Field Reference Usage Level When Both Field Usage and Usage to Failure Are Random
Fengbin Sun · 2019
Paper reliabilityusage-modeling
Thumbnail: Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, Dan Pei · 2019
Paper time-seriesanomaly-detection
Thumbnail: A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, Samuel Bowman · 2018
Paper
Thumbnail: The Annotated Transformer
The Annotated Transformer
Alexander Rush · 2018
Paper transformersattention
Thumbnail: Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah · 2018
PDF
Thumbnail: Deep One-Class Classification
Deep One-Class Classification
Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Muller, Marius Kloft · 2018
PDF anomaly-detectiondeep-svdd
Thumbnail: DeepHit: A Deep Learning Approach to Survival Analysis With Competing Risks
DeepHit: A Deep Learning Approach to Survival Analysis With Competing Risks
Changhee Lee, William Zame, Jinsung Yoon, Mihaela van der Schaar · 2018
PDF survivalcompeting-risks
Thumbnail: DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network
DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network
Jared L. Katzman, Uri Shaham, Alexander Cloninger, Jonathan Bates, Tingting Jiang, Yuval Kluger · 2018
Paper
Thumbnail: Densely Connected Convolutional Networks
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger · 2018
PDF
Thumbnail: Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Kyle Hundman, Valentino Constantinou, Christopher Laporte, Ian Colwell, Tom Soderstrom · 2018
Paper
Thumbnail: Relational Recurrent Neural Networks
Relational Recurrent Neural Networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap · 2018
PDF relational-reasoningsequence-models
Thumbnail: Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
PDF reinforcement-learningprobabilistic-inference
Thumbnail: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine · 2018
PDF reinforcement-learningmaximum-entropy
Thumbnail: Systems protobiology: origin of life in lipid catalytic networks
Systems protobiology: origin of life in lipid catalytic networks
Doron Lancet, Raphael Zidovetzki, Omer Markovitch · 2018
Paper
Thumbnail: Active Inference: A Process Theory
Active Inference: A Process Theory
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, Giovanni Pezzulo · 2017
PDF active-inferencefree-energy
Thumbnail: Attention Is All You Need
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · 2017
PDF transformerssequence-modeling
Thumbnail: Automatic Liver and Tumor Segmentation of CT and MRI Volumes Using Cascaded Fully Convolutional Neural Networks
Automatic Liver and Tumor Segmentation of CT and MRI Volumes Using Cascaded Fully Convolutional Neural Networks
Patrick Ferdinand Christ, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Felix Hofmann, Seyed-Ahmad Ahmadi, Felix Grun, Mohamed Ezzeldin A. Elshaera, Jana Lipkova, Sebastian Schlecht, Freba Ahmaddy, Melvin D. Anastasi, Georgios Kaissis, Julian Holch, Wieland Sommer, Rickmer Braren, Volker Heinemann, Bjoern Menze · 2017
PDF visionsegmentation
Thumbnail: CS231n: Convolutional Neural Networks for Visual Recognition
CS231n: Convolutional Neural Networks for Visual Recognition
Fei-Fei Li, Andrej Karpathy, Justin Johnson · 2017
Paper visioncnn
Thumbnail: Cyclical Learning Rates for Training Neural Networks
Cyclical Learning Rates for Training Neural Networks
Leslie N. Smith · 2017
PDF optimizationtraining
Thumbnail: Deep Learning for Patient-Specific Kidney Graft Survival Analysis
Deep Learning for Patient-Specific Kidney Graft Survival Analysis
Margaux Luck, Tristan Sylvain, Heloise Cardinal, Andrea Lodi, Yoshua Bengio · 2017
PDF survivalmedical-modeling
Thumbnail: Kolmogorov Complexity and Algorithmic Randomness
Kolmogorov Complexity and Algorithmic Randomness
A. Shen, V. A. Uspensky, N. Vereshchagin · 2017
PDF kolmogorov-complexityalgorithmic-randomness
Thumbnail: Neural Message Passing for Quantum Chemistry
Neural Message Passing for Quantum Chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, George E. Dahl · 2017
PDF graph-neural-networksmessage-passing
Thumbnail: Pointer Networks
Pointer Networks
Oriol Vinyals, Meire Fortunato, Navdeep Jaitly · 2017
PDF sequence-modelingneural-networks
Thumbnail: A Simple Neural Network Module for Relational Reasoning
A Simple Neural Network Module for Relational Reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap · 2017
PDF relational-reasoningrepresentation
Thumbnail: Spectrally-normalized margin bounds for neural networks
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J. Foster, Matus Telgarsky · 2017
PDF
Thumbnail: Variational Inference: A Review for Statisticians
Variational Inference: A Review for Statisticians
David M. Blei, Alp Kucukelbir, Jon D. McAuliffe · 2017
PDF variational-inferencebayesian
Thumbnail: Variational Lossy Autoencoder
Variational Lossy Autoencoder
Ilya Sutskever, Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Pieter Abbeel · 2017
PDF generative-modelsrepresentation-learning
Thumbnail: WTTE-RNN: Weibull Time To Event Recurrent Neural Network
WTTE-RNN: Weibull Time To Event Recurrent Neural Network
Egil Martinsson · 2017
PDF survivalwtte
Thumbnail: A Field Guide to Forward-Backward Splitting with a FASTA Implementation
A Field Guide to Forward-Backward Splitting with a FASTA Implementation
Tom Goldstein, Christoph Studer, Richard Baraniuk · 2016
PDF
Thumbnail: Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Patrick Ferdinand Christ, Mohamed Ezzeldin A. Elshaer, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Marco Armbruster, Felix Hofmann, Melvin D'Anastasi, Wieland H. Sommer, Seyed-Ahmad Ahmadi, Bjoern H. Menze · 2016
PDF visionsegmentation
Thumbnail: Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, Zhenyao Zhu · 2016
PDF sequence-modelsspeech
Thumbnail: Group Equivariant Convolutional Networks
Group Equivariant Convolutional Networks
Taco S. Cohen, Max Welling · 2016
PDF
Thumbnail: How a life-like system emerges from a simplistic particle motion law
How a life-like system emerges from a simplistic particle motion law
Thomas Schmickl, Martin Stefanec, Karl Crailsheim · 2016
Paper
Thumbnail: Identity Mappings in Deep Residual Networks
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2016
PDF deep-learningoptimization
Thumbnail: Multi-Scale Context Aggregation by Dilated Convolutions
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu, Vladlen Koltun · 2016
PDF visioncnn
Thumbnail: Order Matters: Sequence to Sequence for Sets
Order Matters: Sequence to Sequence for Sets
Oriol Vinyals, Samy Bengio, Manjunath Kudlur · 2016
PDF sequence-modelssets
Thumbnail: Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, Alan L. Yuille · 2016
PDF
Thumbnail: VIME: Variational Information Maximizing Exploration
VIME: Variational Information Maximizing Exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, Pieter Abbeel · 2016
PDF reinforcement-learningexploration
Thumbnail: Wide & Deep Learning for Recommender Systems
Wide & Deep Learning for Recommender Systems
Lichan Hong, Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Vihan Jain, Xiaobing Liu, Hemal Shah · 2016
PDF recommender-systemsproduction-ml
Thumbnail: A large annotated corpus for learning natural language inference
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning · 2015
Paper
Thumbnail: A Strategy for Origins of Life Research
A Strategy for Origins of Life Research
Caleb Scharf, Nathaniel Virgo, H. James Cleaves, Masashi Aono, Nathanael Aubert-Kato, Arsev Aydinoglu, Ana Barahona, Laura M. Barge, Steven A. Benner, Martin Biehl, Ramon Brasser, Christopher J. Butch, Kuhan Chandru, Leroy Cronin, Sebastian Danielache, Jakob Fischer, John Hernlund, Piet Hut, Takashi Ikegami, Jun Kimura, Kensei Kobayashi, Carlos Mariscal, Shawn McGlynn, Brice Menard, Norman Packard, Robert Pascal, Juli Pereto, Sudha Rajamani, Lana Sinapayen, Eric Smith, Christopher Switzer, Ken Takai, Feng Tian, Yuichiro Ueno, Mary Voytek, Olaf Witkowski, Hikaru Yabuta · 2015
Paper
Thumbnail: Deep Residual Learning for Image Recognition
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
PDF deep-learningcomputer-vision
Thumbnail: FaceNet: A Unified Embedding for Face Recognition and Clustering
FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, James Philbin · 2015
PDF
Thumbnail: Neural Machine Translation by Jointly Learning to Align and Translate
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2015
PDF attentionsequence-models
Thumbnail: Towards Open Set Deep Networks
Towards Open Set Deep Networks
Abhijit Bendale, Terrance Boult · 2015
PDF
Thumbnail: Understanding LSTM Networks
Understanding LSTM Networks
Christopher Olah · 2015
Paper sequence-modelsrnn
Thumbnail: The Unreasonable Effectiveness of Recurrent Neural Networks
The Unreasonable Effectiveness of Recurrent Neural Networks
Andrej Karpathy · 2015
Paper sequence-modelsrnn
Thumbnail: Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton
Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton
Scott Aaronson, Sean M. Carroll, Lauren Ouellette · 2014
PDF complexitykolmogorov-complexity
Thumbnail: Neural Turing Machines
Neural Turing Machines
Alex Graves, Greg Wayne, Ivo Danihelka · 2014
PDF memorysequence-models
Thumbnail: Practical Lessons from Predicting Clicks on Ads at Facebook
Practical Lessons from Predicting Clicks on Ads at Facebook
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, Joaquin Quiñonero Candela · 2014
Paper
Thumbnail: Recurrent Neural Network Regularization
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals · 2014
PDF sequence-modelsrnn
Thumbnail: Thermodynamics as a theory of decision-making with information processing costs
Thermodynamics as a theory of decision-making with information processing costs
Pedro A. Ortega, Daniel A. Braun · 2013
PDF
Thumbnail: ImageNet Classification with Deep Convolutional Neural Networks
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2012
PDF visioncnn
Thumbnail: Solving big data challenges for enterprise application performance management
Solving big data challenges for enterprise application performance management
Tilmann Rabl, Sergio Gómez-Villamor, Mohammad Sadoghi, Victor Muntés-Mulero, Hans-Arno Jacobsen, Serge Mankovskii · 2012
Paper
Thumbnail: The First Law of Complexodynamics
The First Law of Complexodynamics
Scott Aaronson · 2011
Paper complexitythermodynamics
Thumbnail: Feature Selection with the Boruta Package
Feature Selection with the Boruta Package
Miron B. Kursa, Witold R. Rudnicki · 2010
Paper
Thumbnail: Efficient Computation of Optimal Actions
Efficient Computation of Optimal Actions
Emanuel Todorov · 2009
PDF controlreinforcement-learning
Thumbnail: Machine Super Intelligence
Machine Super Intelligence
Shane Legg · 2008
PDF artificial-general-intelligencetheory
Thumbnail: Prevolutionary dynamics and the origin of evolution
Prevolutionary dynamics and the origin of evolution
Martin A. Nowak, Hisashi Ohtsuki · 2008
Paper
Thumbnail: Random Survival Forests
Random Survival Forests
Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, Michael S. Lauer · 2008
Paper survivalreliability
Thumbnail: How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
How Random Is a Coin Toss? Bayesian Inference and the Symbolic Dynamics of Deterministic Chaos
Christopher C. Strelioff, James P. Crutchfield · 2006
PDF bayesian-inferencedynamical-systems
Thumbnail: A Tutorial Introduction to the Minimum Description Length Principle
A Tutorial Introduction to the Minimum Description Length Principle
Peter Grunwald · 2004
PDF mdlcompression
Thumbnail: BLEU: a method for automatic evaluation of machine translation
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu · 2001
Paper
Thumbnail: The use of the area under the ROC curve in the evaluation of machine learning algorithms
The use of the area under the ROC curve in the evaluation of machine learning algorithms
Andrew P. Bradley · 1997
Paper
Thumbnail: Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
Geoffrey E. Hinton, Drew van Camp · 1993
PDF mdlcompression
Thumbnail: Self-reproduction in cellular automata
Self-reproduction in cellular automata
Christopher G. Langton · 1984
Paper
Thumbnail: Information Theory and Statistical Mechanics
Information Theory and Statistical Mechanics
E. T. Jaynes · 1957
Paper