原文: http://zhuanlan.zhihu.com/p/499010769
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Self-supervised Learning is More Robust to Dataset Imbalance
- Looking Back on Learned Experiences For Class/task Incremental Learning
- Path Auxiliary Proposal for MCMC in Discrete Space
- On Bridging Generic and Personalized Federated Learning for Image Classification
- Improved deterministic l2 robustness on CIFAR-10 and CIFAR-100
- Learning Strides in Convolutional Neural Networks
- How to Robustify Black-Box ML Models? A Zeroth-Order Optimization Perspective
- NASPY: Automated Extraction of Automated Machine Learning Models
- Scalable Sampling for Nonsymmetric Determinantal Point Processes
- Strength of Minibatch Noise in SGD
- A Fine-Grained Analysis on Distribution Shift
- Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and Beyond
- Continual Learning with Recursive Gradient Optimization
- Learning meta-features for AutoML
- Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit Differentiation
- Exploring the Limits of Large Scale Pre-training
- SUMNAS: Supernet with Unbiased Meta-Features for Neural Architecture Search
- On Redundancy and Diversity in Cell-based Neural Architecture Search
- NASI: Label- and Data-agnostic Neural Architecture Search at Initialization
- NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet Training
- Surrogate NAS Benchmarks: Going Beyond the Limited Search Spaces of Tabular NAS Benchmarks
- Automatic Loss Function Search for Predict-Then-Optimize Problems with Strong Ranking Property