Paper Digest: 100 Must-Read Machine Learning Papers of the Past 10 Years (2016-2025)
Machine learning has changed faster in the past decade than in any period before it. The Transformer replaced recurrence with attention. GPT-3 showed what scale alone could do. Diffusion models became the standard way to generate images. InstructGPT, LoRA, and DPO made it practical to adapt and align large language models. PyTorch and AdamW became default tools in nearly every lab. Keeping up is hard: NeurIPS alone now accepts thousands of papers a year. This reading list is a starting point: 100 machine learning papers from 2016 to 2025 that the research community has built on most.
How the papers were selected
The papers were selected by citation count from the three flagship machine learning conferences: NeurIPS, ICML, and ICLR. For each year from 2016 to 2025, we took the ten most-cited papers across that year’s conferences. Selecting by year keeps recent work from being crowded out by older papers that have simply had more time to collect citations. Our ICLR data begins in 2018, so the 2016 and 2017 picks come from NeurIPS and ICML. These venues cover the whole field, so the list includes plenty of vision, language, and speech papers alongside core machine learning.
The list is ordered newest year first. Read from the bottom up, it doubles as a short history of how the field moved:
- GANs, deep reinforcement learning, and graph neural networks
- transformers and large-scale pretraining
- self-supervised learning and diffusion models
- today’s work on aligning large language models and teaching them to reason
A note on scope: this list reflects our selection criteria, not a definitive ranking of the best machine learning research. There may well be better or more important papers elsewhere. Some landmark work appeared outside these three conferences: BERT at NAACL, AlphaFold in Nature, and many influential LLM technical reports, such as LLaMA, only on arXiv. Citation counts also favor popular topics and lag behind the newest work. Treat the list as a well-grounded map of the field rather than the final word. For the full year-by-year lists, see the Most Influential NeurIPS, ICML, and ICLR papers pages. Lists for other venues are on the Best Paper Digest page.
Keeping up beyond this list
If you work across fields, the companion lists follow the same method: 100 must-read computer vision papers of the past 10 years (CVPR, ICCV, ECCV) and 100 must-read natural language processing papers of the past 10 years (ACL, EMNLP, NAACL).
A citation-based list looks backward. It tells you what mattered, not what is emerging this month. Every paper on Paper Digest links to related papers, patents, grants, and experts, so you can explore the research around any of the work below. You can also run a literature review on a specific topic. If you would like new machine learning papers matched to your interests each morning, you can sign up and set up a daily digest.
TABLE 1: Paper Digest: 100 Must-Read Machine Learning Papers of the Past 10 Years (2016-2025)
| Year | Rank | Paper | Author(s) |
|---|---|---|---|
| 2025 | 1 | SAM 2: Segment Anything in Images and Videos IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. |
NIKHILA RAVI et. al. |
| 2025 | 2 | CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos that align seamlessly with text prompts, with a frame rate of 16 fps and resolution of 768 x 1360 pixels. |
ZHUOYI YANG et. al. |
| 2025 | 3 | KAN: Kolmogorov–Arnold Networks IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the Kolmogorov-Arnold representation theorem, we propose Kolmogorov-Arnold Networks (KANs) as promising alternatives to Multi-Layer Perceptrons (MLPs). |
ZIMING LIU et. al. |
| 2025 | 4 | DAPO: An Open-Source LLM Reinforcement Learning System at Scale IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. |
QIYING YU et. al. |
| 2025 | 5 | LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which collects new problems over time from contests across three competition platforms, Leetcode, Atcoder, and Codeforces. |
NAMAN JAIN et. al. |
| 2025 | 6 | YOLOv12: Attention-Centric Real-Time Object Detectors IF:7 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes an attention-centric YOLO framework, namely YOLOv12, that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms. |
Yunjie Tian; Qixiang Ye; David Doermann; |
| 2025 | 7 | WorldSimBench: Towards Video Generation Models As World Simulators IF:7 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we classify the functionalities of predictive models into a hierarchy and take the first step in evaluating World Simulators by proposing a dual evaluation framework called WorldSimBench.In the Explicit Perceptual Evaluation, we introduce the HF-Embodied Dataset, a video assessment dataset based on fine-grained human feedback, which we use to train a Human Preference Evaluator that aligns with human perception and explicitly assesses the visual fidelity of World Simulater. |
YIRAN QIN et. al. |
| 2025 | 8 | WizardMath: Empowering Mathematical Reasoning for Large Language Models Via Reinforced Evol-Instruct IF:7 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present WizardMath, which enhances the mathematical reasoning abilities of LLMs, by applying our proposed Reinforcement Learning from Evol-Instruct Feedback (RLEIF) method to the domain of math. |
HAIPENG LUO et. al. |
| 2025 | 9 | Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond The Base Model? IF:7 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we take a critical look at \textit{the current state of RLVR} by systematically probing the reasoning capability boundaries of RLVR-trained LLMs across diverse model families, RL algorithms, and math/coding/visual reasoning benchmarks, using pass@\textit{k} at large \textit{k} values as the evaluation metric. |
YANG YUE et. al. |
| 2025 | 10 | Show-o: One Single Transformer to Unify Multimodal Understanding and Generation IF:6 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. |
JINHENG XIE et. al. |
| 2024 | 1 | SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Stable Diffusion XL (SDXL), a latent diffusion model for text-to-image synthesis. |
DUSTIN PODELL et. al. |
| 2024 | 2 | YOLOv10: Real-Time End-to-End Object Detection IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we aim to further advance the performance-efficiency boundary of YOLOs from both the post-processing and the model architecture. |
AO WANG et. al. |
| 2024 | 3 | Scaling Rectified Flow Transformers for High-Resolution Image Synthesis IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales. |
PATRICK ESSER et. al. |
| 2024 | 4 | MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We believe that the enhanced multi-modal generation capabilities of GPT-4 stem from the utilization of sophisticated large language models (LLM). To examine this phenomenon, we present MiniGPT-4, which aligns a frozen visual encoder with a frozen advanced LLM, Vicuna, using one projection layer. |
Deyao Zhu; Jun Chen; Xiaoqian Shen; Xiang Li; Mohamed Elhoseiny; |
| 2024 | 5 | Let’s Verify Step By Step IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We conduct our own investigation, finding that process supervision significantly outperforms outcome supervision for training models to solve problems from the challenging MATH dataset. |
HUNTER LIGHTMAN et. al. |
| 2024 | 6 | FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that the inefficiency is due to suboptimal work partitioning between different thread blocks and warps on the GPU, causing either low-occupancy or unnecessary shared memory reads/writes. We propose FlashAttention-2, with better work partitioning to address these issues. |
Tri Dao; |
| 2024 | 7 | VMamba: Visual State Space Model IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we transplant Mamba, a state-space language model, into VMamba, a vision backbone that works in linear time complexity. |
LIU YUE et. al. |
| 2024 | 8 | Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we show that the reliance on self-attention for visual representation learning is not necessary and propose a new generic vision backbone with bidirectional Mamba blocks (Vim), which marks the image sequences with position embeddings and compresses the visual representation with bidirectional state space models. |
LIANGHUI ZHU et. al. |
| 2024 | 9 | SWE-bench: Can Language Models Resolve Real-world Github Issues? IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories. |
CARLOS E JIMENEZ et. al. |
| 2024 | 10 | Self-RAG: Learning to Retrieve, Generate, and Critique Through Self-Reflection IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new framework called **Self-Reflective Retrieval-Augmented Generation (Self-RAG)** that enhances an LM’s quality and factuality through retrieval and self-reflection. |
Akari Asai; Zeqiu Wu; Yizhong Wang; Avirup Sil; Hannaneh Hajishirzi; |
| 2023 | 1 | Visual Instruction Tuning IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first attempt to use language-only GPT-4 to generate multimodal language-image instruction-following data. By instruction tuning on such generated data, we introduce LLaVA: Large Language and Vision Assistant, an end-to-end trained large multimodal model that connects a vision encoder and an LLM for general-purpose visual and language understanding. |
Haotian Liu; Chunyuan Li; Qingyang Wu; Yong Jae Lee; |
| 2023 | 2 | Direct Preference Optimization: Your Language Model Is Secretly A Reward Model IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, RLHF is a complex and often unstable procedure, first fitting a reward model that reflects the human preferences, and then fine-tuning the large unsupervised LM using reinforcement learning to maximize this estimated reward without drifting too far from the original model. In this paper, we leverage a mapping between reward functions and optimal policies to show that this constrained reward maximization problem can be optimized exactly with a single stage of policy training, essentially solving a classification problem on the human preference data. |
RAFAEL RAFAILOV et. al. |
| 2023 | 3 | BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps vision-language pre-training from off-the-shelf frozen pre-trained image encoders and frozen large language models. |
Junnan Li; Dongxu Li; Silvio Savarese; Steven Hoi; |
| 2023 | 4 | Robust Speech Recognition Via Large-Scale Weak Supervision IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. |
ALEC RADFORD et. al. |
| 2023 | 5 | Self-Consistency Improves Chain of Thought Reasoning in Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a new decoding strategy, self-consistency, to replace the naive greedy decoding used in chain-of-thought prompting. |
XUEZHI WANG et. al. |
| 2023 | 6 | ReAct: Synergizing Reasoning and Acting in Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore the use of LLMs to generate both reasoning traces and task-specific actions in an interleaved manner, allowing for greater synergy between the two: reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources, such as knowledge bases or environments, to gather additional information. |
SHUNYU YAO et. al. |
| 2023 | 7 | QLoRA: Efficient Finetuning of Quantized LLMs IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present QLoRA, an efficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance. |
Tim Dettmers; Artidoro Pagnoni; Ari Holtzman; Luke Zettlemoyer; |
| 2023 | 8 | Tree of Thoughts: Deliberate Problem Solving with Large Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. |
SHUNYU YAO et. al. |
| 2023 | 9 | DreamFusion: Text-to-3D Using 2D Diffusion IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D or multiview data and efficient architectures for denoising 3D data, neither of which currently exist. In this work, we circumvent these limitations by using a pretrained 2D text-to-image diffusion model to perform text-to-3D synthesis. |
Ben Poole; Ajay Jain; Jonathan T. Barron; Ben Mildenhall; |
| 2023 | 10 | Flow Matching for Generative Modeling IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. |
Yaron Lipman; Ricky T. Q. Chen; Heli Ben-Hamu; Maximilian Nickel; Matthew Le; |
| 2022 | 1 | Training Language Models to Follow Instructions with Human Feedback IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We fine-tune GPT-3 using data collected from human labelers. The resulting model, called InstructGPT, outperforms GPT-3 on a range of NLP tasks. |
LONG OUYANG et. al. |
| 2022 | 2 | LoRA: Low-Rank Adaptation of Large Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Finetuning updates have a low intrinsic rank which allows us to train only the rank decomposition matrices of certain weights, yielding better performance and practical benefits. |
EDWARD J HU et. al. |
| 2022 | 3 | Chain of Thought Prompting Elicits Reasoning in Large Language Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We explore how generating a chain of thought—a series of intermediate reasoning steps—significantly improves the ability of large language models to perform complex reasoning. |
JASON WEI et. al. |
| 2022 | 4 | Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. |
CHITWAN SAHARIA et. al. |
| 2022 | 5 | Large Language Models Are Zero-Shot Reasoners IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a single zero-shot prompt that elicits effective chain of thought reasoning across diverse benchmarks that require multi-step thinking. |
Takeshi Kojima; Shixiang (Shane) Gu; Machel Reid; Yutaka Matsuo; Yusuke Iwasawa; |
| 2022 | 6 | BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. |
Junnan Li; Dongxu Li; Caiming Xiong; Steven Hoi; |
| 2022 | 7 | Flamingo: A Visual Language Model for Few-Shot Learning IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. |
JEAN-BAPTISTE ALAYRAC et. al. |
| 2022 | 8 | LAION-5B: An Open Large-scale Dataset for Training Next Generation Image-text Models IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present LAION-5B, an open, publically available dataset of 5.8B image-text pairs and validate it by reproducing results of training state-of-the-art CLIP models of different scale. |
CHRISTOPH SCHUHMANN et. al. |
| 2022 | 9 | GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We explore diffusion models for the problem of text-conditional image synthesis and compare two different guidance strategies: CLIP guidance and classifier-free guidance. |
ALEXANDER QUINN NICHOL et. al. |
| 2022 | 10 | FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a fast and memory-efficient exact attention algorithm by accounting for GPU memory reads/writes, yielding faster end-to-end training time and higher quality models with longer sequences. |
Tri Dao; Dan Fu; Stefano Ermon; Atri Rudra; Christopher Ré; |
| 2021 | 1 | An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Transformers applied directly to image patches and pre-trained on large datasets work really well on image classification. |
ALEXEY DOSOVITSKIY et. al. |
| 2021 | 2 | Learning Transferable Visual Models From Natural Language Supervision IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. |
ALEC RADFORD et. al. |
| 2021 | 3 | Diffusion Models Beat GANs on Image Synthesis IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. |
Prafulla Dhariwal; Alexander Nichol; |
| 2021 | 4 | Denoising Diffusion Implicit Models IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show and justify a GAN-like iterative generative model with relatively fast sampling, high sample quality and without any adversarial training. |
Jiaming Song; Chenlin Meng; Stefano Ermon; |
| 2021 | 5 | Score-Based Generative Modeling Through Stochastic Differential Equations IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A general framework for training and sampling from score-based models that unifies and generalizes previous methods, allows likelihood computation, and enables controllable generation. |
YANG SONG et. al. |
| 2021 | 6 | Training Data-efficient Image Transformers & Distillation Through Attention IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we produce competitive convolution-free transformers trained on ImageNet only using a single computer in less than 3 days. |
HUGO TOUVRON et. al. |
| 2021 | 7 | SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SegFormer, a simple, efficient yet powerful semantic segmentation framework which unifies Transformers with lightweight multilayer perceptron (MLP) decoders. |
ENZE XIE et. al. |
| 2021 | 8 | Measuring Massive Multitask Language Understanding IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We test language models on 57 different multiple-choice tasks. |
DAN HENDRYCKS et. al. |
| 2021 | 9 | Deformable DETR: Deformable Transformers for End-to-End Object Detection IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Deformable DETR is an efficient and fast-converging end-to-end object detector. It mitigates the high complexity and slow convergence issues of DETR via a novel sampling-based efficient attention mechanism. |
XIZHOU ZHU et. al. |
| 2021 | 10 | Zero-Shot Text-to-Image Generation IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. |
ADITYA RAMESH et. al. |
| 2020 | 1 | Language Models Are Few-Shot Learners IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. |
TOM BROWN et. al. |
| 2020 | 2 | Denoising Diffusion Probabilistic Models IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. |
Jonathan Ho; Ajay Jain; Pieter Abbeel; |
| 2020 | 3 | A Simple Framework for Contrastive Learning of Visual Representations IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a simple framework for contrastive representation learning. |
Ting Chen; Simon Kornblith; Mohammad Norouzi; Geoffrey Hinton; |
| 2020 | 4 | Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. |
PATRICK LEWIS et. al. |
| 2020 | 5 | ELECTRA: Pre-training Text Encoders As Discriminators Rather Than Generators IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A text encoder trained to distinguish real input tokens from plausible fakes efficiently learns effective language representations. |
Kevin Clark; Minh-Thang Luong; Quoc V. Le; Christopher D. Manning; |
| 2020 | 6 | BERTScore: Evaluating Text Generation With BERT IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose BERTScore, an automatic evaluation metric for text generation, which correlates better with human judgments and provides stronger model selection performance than existing metrics. |
Tianyi Zhang*; Varsha Kishore*; Felix Wu*; Kilian Q. Weinberger; Yoav Artzi; |
| 2020 | 7 | Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler. |
Alexei Baevski; Yuhao Zhou; Abdel-rahman Mohamed; Michael Auli; |
| 2020 | 8 | ALBERT: A Lite BERT For Self-supervised Learning Of Language Representations IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A new pretraining method that establishes new state-of-the-art results on the GLUE, RACE, and SQuAD benchmarks while having fewer parameters compared to BERT-large. |
ZHENZHONG LAN et. al. |
| 2020 | 9 | Supervised Contrastive Learning IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we extend the self-supervised batch contrastive approach to the fully-supervised setting, allowing us to effectively leverage label information. |
PRANNAY KHOSLA et. al. |
| 2020 | 10 | Unsupervised Learning of Visual Features By Contrasting Cluster Assignments IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose an online algorithm, SwAV, that takes advantage of contrastive methods without requiring to compute pairwise comparisons. |
MATHILDE CARON et. al. |
| 2019 | 1 | PyTorch: An Imperative Style, High-Performance Deep Learning Library IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. |
ADAM PASZKE et. al. |
| 2019 | 2 | Decoupled Weight Decay Regularization IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Novel variants of optimization methods that combine the benefits of both adaptive and non-adaptive methods. |
Ilya Loshchilov; Frank Hutter; |
| 2019 | 3 | EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically study model scaling and identify that carefully balancing network depth, width, and resolution can lead to better performance. |
Mingxing Tan; Quoc Le; |
| 2019 | 4 | How Powerful Are Graph Neural Networks? IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop theoretical foundations for the expressive power of GNNs and design a provably most powerful GNN. |
Keyulu Xu*; Weihua Hu*; Jure Leskovec; Stefanie Jegelka; |
| 2019 | 5 | XLNet: Generalized Autoregressive Pretraining for Language Understanding IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. |
ZHILIN YANG et. al. |
| 2019 | 6 | GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a multi-task benchmark and analysis platform for evaluating generalization in natural language understanding systems. |
ALEX WANG et. al. |
| 2019 | 7 | Large Scale GAN Training for High Fidelity Natural Image Synthesis IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: GANs benefit from scaling up. |
Andrew Brock; Jeff Donahue; Karen Simonyan; |
| 2019 | 8 | Parameter-Efficient Transfer Learning for NLP IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As an alternative, we propose transfer with adapter modules. |
NEIL HOULSBY et. al. |
| 2019 | 9 | Generative Modeling By Estimating Gradients of The Data Distribution IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new generative model where samples are produced via Langevin dynamics using gradients of the data distribution estimated with score matching. |
Yang Song; Stefano Ermon; |
| 2019 | 10 | DARTS: Differentiable Architecture Search IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a differentiable architecture search algorithm for both convolutional and recurrent networks, achieving competitive performance with the state of the art using orders of magnitude less computation resources. |
Hanxiao Liu; Karen Simonyan; Yiming Yang; |
| 2018 | 1 | Graph Attention Networks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A novel approach to processing graph-structured data by neural networks, leveraging attention over a node’s neighborhood. Achieves state-of-the-art results on transductive citation network tasks and an inductive protein-protein interaction task. |
PETAR VELICKOVIC et. al. |
| 2018 | 2 | Towards Deep Learning Models Resistant to Adversarial Attacks IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide a principled, optimization-based re-look at the notion of adversarial examples, and develop methods that produce models that are adversarially robust against a wide range of adversaries. |
Aleksander Madry; Aleksandar Makelov; Ludwig Schmidt; Dimitris Tsipras; Adrian Vladu; |
| 2018 | 3 | Mixup: Beyond Empirical Risk Minimization IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Training on convex combinations between random training examples and their labels improves generalization in deep neural networks |
Hongyi Zhang; Moustapha Cisse; Yann N. Dauphin; David Lopez-Paz; |
| 2018 | 4 | Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with A Stochastic Actor IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose soft actor-critic, an off-policy actor-critic deep RL algorithm based on the maximum entropy reinforcement learning framework. |
Tuomas Haarnoja; Aurick Zhou; Pieter Abbeel; Sergey Levine; |
| 2018 | 5 | Progressive Growing of GANs for Improved Quality, Stability, and Variation IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We train generative adversarial networks in a progressive fashion, enabling us to generate high-resolution images with high quality. |
Tero Karras; Timo Aila; Samuli Laine; Jaakko Lehtinen; |
| 2018 | 6 | Addressing Function Approximation Error in Actor-Critic Methods IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluate our method on the suite of OpenAI gym tasks, outperforming the state of the art in every environment tested. |
Scott Fujimoto; Herke Hoof; David Meger; |
| 2018 | 7 | Neural Ordinary Differential Equations IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new family of deep neural network models. |
Tian Qi Chen; Yulia Rubanova; Jesse Bettencourt; David K. Duvenaud; |
| 2018 | 8 | Spectral Normalization for Generative Adversarial Networks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel weight normalization technique called spectral normalization to stabilize the training of the discriminator of GANs. |
Takeru Miyato; Toshiki Kataoka; Masanori Koyama; Yuichi Yoshida; |
| 2018 | 9 | Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A neural sequence model that learns to forecast on a directed graph. |
Yaguang Li; Rose Yu; Cyrus Shahabi; Yan Liu; |
| 2018 | 10 | Neural Tangent Kernel: Convergence and Generalization in Neural Networks IF:8 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function (which maps input vectors to output vectors) follows the so-called kernel gradient associated with a new object, which we call the Neural Tangent Kernel (NTK). |
Arthur Jacot; Franck Gabriel; Clement Hongler; |
| 2017 | 1 | Attention Is All You Need IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel, simple network architecture based solely onan attention mechanism, dispensing with recurrence and convolutions entirely.Experiments on two machine translation tasks show these models to be superiorin quality while being more parallelizable and requiring significantly less timeto train. |
ASHISH VASWANI et. al. |
| 2017 | 2 | A Unified Approach to Interpreting Model Predictions IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this problem, we present a unified framework for interpreting predictions, SHAP (SHapley Additive exPlanations). |
Scott M. Lundberg; Su-In Lee; |
| 2017 | 3 | Inductive Representation Learning on Large Graphs IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here we present GraphSAGE, a general, inductive framework that leverages node feature information (e.g., text attributes) to efficiently generate node embeddings. |
Will Hamilton; Zhitao Ying; Jure Leskovec; |
| 2017 | 4 | GANs Trained By A Two Time-Scale Update Rule Converge to A Local Nash Equilibrium IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a two time-scale update rule (TTUR) for training GANs with stochastic gradient descent on arbitrary GAN loss functions. |
Martin Heusel; Hubert Ramsauer; Thomas Unterthiner; Bernhard Nessler; Sepp Hochreiter; |
| 2017 | 5 | Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an algorithm for meta-learning that is model-agnostic, in the sense that it is compatible with any model trained with gradient descent and applicable to a variety of different learning problems, including classification, regression, and reinforcement learning. |
Chelsea Finn; Pieter Abbeel; Sergey Levine; |
| 2017 | 6 | PointNet++: Deep Hierarchical Feature Learning on Point Sets in A Metric Space IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a hierarchical neural network that applies PointNet recursively on a nested partitioning of the input point set. With further observation that point sets are usually sampled with varying densities, which results in greatly decreased performance for networks trained on uniform densities, we propose novel set learning layers to adaptively combine features from multiple scales. |
Charles Ruizhongtai Qi; Li Yi; Hao Su; Leonidas J. Guibas; |
| 2017 | 7 | LightGBM: A Highly Efficient Gradient Boosting Decision Tree IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To tackle this problem, we propose two novel techniques: \emph{Gradient-based One-Side Sampling} (GOSS) and \emph{Exclusive Feature Bundling} (EFB). |
GUOLIN KE et. al. |
| 2017 | 8 | Improved Training of Wasserstein GANs IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an alternative to clipping weights: penalize the norm of gradient of the critic with respect to its input. |
Ishaan Gulrajani; Faruk Ahmed; Martin Arjovsky; Vincent Dumoulin; Aaron C. Courville; |
| 2017 | 9 | Prototypical Networks for Few-shot Learning IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Prototypical Networks for the problem of few-shot classification, where a classifier must generalize to new classes not seen in the training set, given only a small number of examples of each new class. |
Jake Snell; Kevin Swersky; Richard Zemel; |
| 2017 | 10 | Wasserstein Generative Adversarial Networks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a new algorithm named WGAN, an alternative to traditional GAN training. |
Martin Arjovsky; Soumith Chintala; L�on Bottou; |
| 2016 | 1 | Dropout As A Bayesian Approximation: Representing Model Uncertainty In Deep Learning IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper we develop a new theoretical framework casting dropout training in deep neural networks (NNs) as approximate Bayesian inference in deep Gaussian processes. |
Yarin Gal; Zoubin Ghahramani; |
| 2016 | 2 | Improved Techniques for Training GANs IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a variety of new architectural features and training procedures that we apply to the generative adversarial networks (GANs) framework. |
TIM SALIMANS et. al. |
| 2016 | 3 | Asynchronous Methods For Deep Reinforcement Learning IF:10 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a conceptually simple and lightweight framework for deep reinforcement learning that uses asynchronous gradient descent for optimization of deep neural network controllers. |
VOLODYMYR MNIH et. al. |
| 2016 | 4 | Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we are interested in generalizing convolutional neural networks (CNNs) from low-dimensional regular grids, where image, video and speech are represented, to high-dimensional irregular domains, such as social networks, brain connectomes or words’ embedding, represented by graphs. |
Micha�l Defferrard; Xavier Bresson; Pierre Vandergheynst; |
| 2016 | 5 | Matching Networks for One Shot Learning IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we employ ideas from metric learning based on deep neural features and from recent advances that augment neural networks with external memories. |
Oriol Vinyals; Charles Blundell; Timothy Lillicrap; koray kavukcuoglu; Daan Wierstra; |
| 2016 | 6 | R-FCN: Object Detection Via Region-based Fully Convolutional Networks IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present region-based, fully convolutional networks for accurate and efficient object detection. |
Jifeng Dai; Yi Li; Kaiming He; Jian Sun; |
| 2016 | 7 | Equality of Opportunity in Supervised Learning IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a criterion for discrimination against a specified sensitive attribute in supervised learning, where the goal is to predict some target based on available features. |
Moritz Hardt; Eric Price; Nati Srebro; |
| 2016 | 8 | InfoGAN: Interpretable Representation Learning By Information Maximizing Generative Adversarial Nets IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper describes InfoGAN, an information-theoretic extension to the Generative Adversarial Network that is able to learn disentangled representations in a completely unsupervised manner. |
XI CHEN et. al. |
| 2016 | 9 | Dueling Network Architectures For Deep Reinforcement Learning IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present a new neural network architecture for model-free reinforcement learning. |
ZIYU WANG et. al. |
| 2016 | 10 | Man Is to Computer Programmer As Woman Is to Homemaker? Debiasing Word Embeddings IF:9 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This raises concerns because their widespread use, as we describe, often tends to amplify these biases. |
Tolga Bolukbasi; Kai-Wei Chang; James Y. Zou; Venkatesh Saligrama; Adam T. Kalai; |