Bibliography (245):

  1. backstop#deep-bayes

    [Transclude the forward-link's context]

  2. ML Scaling subreddit

  3. It Looks Like You’re Trying To Take Over The World

  4. Introduction

  5. AI 2027

  6. Human-like Neural Nets by Catapulting

  7. GPT-3: Language Models are Few-Shot Learners

  8. GPT-3 paper § Figure F.1: Four uncurated completions from a context suggesting the model compose a poem in the style of Wallace Stevens with the title ‘Shadows on the Way’

  9. GPT-3 Creative Fiction

  10. GPT-2 Neural Network Poetry

  11. GPT-3 Github JSON Dump Reformatted to Readable HTML

  12. OpenAI API

  13. Better Language Models and Their Implications

  14. GPT-3 Creative Fiction § BPEs

  15. Using Fast Weights to Attend to the Recent Past

  16. https://www.reddit.com/r/reinforcementlearning/search/?q=flair%3AMetaRL&include_over_18=on&restrict_sr=on&sort=top

  17. One-shot Learning with Memory-Augmented Neural Networks

  18. Prefrontal cortex as a meta-reinforcement learning system

  19. Matt Botvinick on the spontaneous emergence of learning algorithms

  20. Reinforcement Learning, Fast and Slow

  21. AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence

  22. On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models

  23. One Big Net For Everything

  24. Meta-Learning: Learning to Learn Fast

  25. Meta Reinforcement Learning

  26. Jukebox: We’re introducing Jukebox, a neural net that generates music, including rudimentary singing, as raw audio in a variety of genres and artist styles. We’re releasing the model weights and code, along with a tool to explore the generated samples.

  27. GPT-1: Improving Language Understanding with Unsupervised Learning

  28. I Recently Came across Https://arxiv.org/abs/2004.08900, Which ‘Assumes 2-3 Runs’ of T5-11B. In Fact, We Trained T5-11B once. That’s Why We Spend 35 Pages Figuring out How We Should Train Before We Start Training. You Don’t Want to Mess up a Training Run That Big.

  29. CERN makes bold push to build €21-billion supercollider: European particle-physics lab will pursue a 100-kilometre machine to uncover the Higgs boson’s secrets—but it doesn’t yet have the funds

  30. Whole Brain Emulation: A Roadmap

  31. 2019 recent trends in GPU price per FLOPS

  32. Measuring the Algorithmic Efficiency of Neural Networks

  33. Dota 2 With Large Scale Deep Reinforcement Learning § Pg11

  34. D.5: Context Dependence

  35. ‘self-attention’ directory

  36. WBE and DRL: a Middle Way of imitation learning

  37. LHOPT: A Generalizable Approach to Learning Optimizers

  38. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

  39. GPT-3 random sample dump: JavaScript tutorial

  40. On the Measure of Intelligence

  41. Deep Learning Hardware: Past, Present, & Future § Pg60

  42. Technology Forecasting: The Garden of Forking Paths

  43. GPT-3: Language Models Are Few-Shot Learners: 5. Limitations

  44. CTRL: A Conditional Transformer Language Model For Controllable Generation

  45. Towards a Human-like Open-Domain Chatbot

  46. MegatronLM: Training Billion+ Parameter Language Models Using GPU Model Parallelism

  47. T5: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

  48. Turing-NLG: A 17-billion-parameter language model by Microsoft

  49. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

  50. Extracting Training Data from Large Language Models

  51. Does Learning Require Memorization? A Short Tale about a Long Tail

  52. The Computational Limits of Deep Learning

  53. The Unreasonable Effectiveness of Data

  54. Scaling to Very Very Large Corpora for Natural Language Disambiguation

  55. Large Language Models in Machine Translation

  56. 2017-koehn-figure3-bleuscoreswithvaryingamountsoftrainingdata.png

  57. WinoGrande: An Adversarial Winograd Schema Challenge at Scale

  58. Total Compute Used to Train Language Model: Table D.1

  59. AI and Compute

  60. OpenAI's GPT-3 Language Model: A Technical Overview

  61. People I Know at OpenAI Say V4 Is around the Corner and Easily Doable, And...will Be Here Soon (Not Months but Year or So). And They Are Confident It Will Scale and Be around 100--1000×.

  62. Microsoft announces new supercomputer, lays out vision for future AI work

  63. Scaling Laws for Neural Language Models

  64. Scaling Laws for Neural Language Models: Figure 1: Language Modeling Performance Improves Smoothly As We Increase the Model Size, Dataset Size, and Amount of Compute Used for Training.

  65. Scaling Laws for Neural Language Models: Figure 15: Far beyond the Model Sizes We Study Empirically, We Find a Contradiction between Our Equations § Pg17

  66. https://arxiv.org/pdf/2005.14165.pdf#page=11&org=openai

  67. Table 2.2: Datasets Used to Train GPT-3. ‘Weight in Training Mix’ Refers to the Fraction of Examples during Training That Are Drawn from a given Dataset, Which We Intentionally Do Not Make Proportional to the Size of the Dataset. As a Result, When We Train for 300 Billion Tokens, Some Datasets Are Seen up to 3.4 times during Training While Other Datasets Are Seen Less Than Once.

  68. 2020-adiwardana-meena-figure1-humanratingsvslikelihood.png

  69. 2020-brown-figure313-humanabilitytodetectmodelgeneratednewsstories.jpg

  70. 2020-hendrycks-figure1b-gpt3-qascaling.png

  71. MMLU: Measuring Massive Multitask Language Understanding

  72. https://x.com/geoffreyhinton/status/1270814602931187715

  73. The Bitter Lesson

  74. Image GPT (iGPT): We find that, just as a large transformer model trained on language can generate coherent text, the same exact model trained on pixel sequences can generate coherent image completions and samples

  75. Vision Transformer: An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale

  76. Generative Language Modeling for Automated Theorem Proving

  77. The neural architecture of language: Integrative reverse-engineering converges on a model for predictive processing

  78. Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)

  79. ‘MLP NN’ directory

  80. A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes

  81. Dota 2 With Large Scale Deep Reinforcement Learning: §4.3: Batch Size

  82. How AI Training Scales

  83. BigGAN: Large Scale GAN Training For High Fidelity Natural Image Synthesis § 5.2 Additional Evaluation On JFT-300M

  84. Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images

  85. NVAE: A Deep Hierarchical Variational Autoencoder

  86. Big Transfer (BiT): General Visual Representation Learning

  87. Are we done with ImageNet?

  88. On Robustness and Transferability of Convolutional Neural Networks

  89. Robustness properties of Facebook’s ResNeXt WSL models

  90. Self-training with Noisy Student improves ImageNet classification

  91. Measuring Robustness to Natural Distribution Shifts in Image Classification

  92. Understanding Robustness of Transformers for Image Classification

  93. Distilling the Knowledge in a Neural Network

  94. Smooth Adversarial Training

  95. 12-in-1: Multi-Task Vision and Language Representation Learning

  96. VideoBERT: A Joint Model for Video and Language Representation Learning

  97. The messy, secretive reality behind OpenAI’s bid to save the world: The AI moonshot was founded in the spirit of transparency. This is the inside story of how competitive pressure eroded that idealism

  98. High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks

  99. Grandmaster level in StarCraft II using multi-agent reinforcement learning

  100. One-Shot High-Fidelity Imitation: Training Large-Scale Deep Nets with RL

  101. A Style-Based Generator Architecture for Generative Adversarial Networks

  102. A simple neural network module for relational reasoning

  103. Neural scene representation and rendering

  104. Transformers as Soft Reasoners over Language

  105. Environmental drivers of systematicity and generalization in a situated agent

  106. Gated-Attention Architectures for Task-Oriented Language Grounding

  107. Interactive Grounded Language Acquisition and Generalization in a 2D World

  108. Compositional generalization through meta sequence-to-sequence learning

  109. Imitating Interactive Intelligence

  110. Solving Rubik’s Cube with a Robot Hand

  111. Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning

  112. DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

  113. Procgen Benchmark: We’re releasing Procgen Benchmark, 16 simple-to-use procedurally-generated environments which provide a direct measure of how quickly a reinforcement learning agent learns generalizable skills

  114. Understanding RL Vision: With diverse environments, we can analyze, diagnose and edit deep reinforcement learning models using attribution

  115. Muppet: Massive Multi-task Representations with Pre-Finetuning

  116. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play

  117. MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

  118. Reflections After Refereeing Papers for NIPS

  119. Understanding deep learning requires rethinking generalization

  120. Deep Double Descent: We show that the double descent phenomenon occurs in CNNs, ResNets, and transformers: performance first improves, then gets worse, and then improves again with increasing model size, data size, or training time

  121. Understanding the generalization of ‘lottery tickets’ in neural networks

  122. Bayesian Deep Learning and a Probabilistic Perspective of Generalization

  123. On Linear Identifiability of Learned Representations

  124. Zoom In: An Introduction to Circuits—By studying the connections between neurons, we can find meaningful algorithms in the weights of neural networks

  125. Neural Networks, Manifolds, and Topology

  126. Logarithmic Pruning is All You Need

  127. Direct Fit to Nature: An Evolutionary Perspective on Biological and Artificial Neural Networks

  128. The Shape of Learning Curves: a Review: 6. Ill-Behaved Learning Curves: 6.1. Phase Transitions

  129. The Brain as a Universal Learning Machine

  130. The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost

  131. Jumping NLP Curves: A Review of Natural Language Processing Research [Review Article]

  132. difference#efficient-natural-languages

    [Transclude the forward-link's context]

  133. Natural Language Processing (Almost) from Scratch

  134. Data Distributional Properties Drive Emergent Few-Shot Learning in Transformers

  135. The Legacy of Hiroshima

  136. Hopfield Networks is All You Need

  137. 2019-radford-figure4-gpt2validationloss.jpg

  138. 2020-brown-figure31-gpt3scaling.png

  139. Building a Large Annotated Corpus of English: The Penn Treebank

  140. One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling

  141. The LAMBADA dataset: Word prediction requiring a broad discourse context

  142. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf#page=5

  143. Estimation of Gap Between Current Language Models and Human Performance

  144. https://arxiv.org/pdf/2005.14165.pdf&org=openai#page=12

  145. Ilya Sutskever: Deep Learning

  146. If You Want to Solve a Hard Problem in Reinforcement Learning, You Just Scale. It’s Just Gonna Work Just like Supervised Learning. It’s the Same, the Same Story Exactly. It Was Kind of Hard to Believe That Supervised Learning Can Do All Those Things, but It’s Not Just Vision, It’s Everything and the Same Thing Seems to Hold for Reinforcement Learning Provided You Have a Lot of Experience.

  147. What Could Make AI Conscious?

  148. https://wandb.ai/wandb_fc/gradient-dissent/reports/What-could-make-AI-conscious-with-Wojciech-Zaremba-co-founder-of-OpenAI--Vmlldzo3NDk3MDI

  149. Evolution Strategies as a Scalable Alternative to Reinforcement Learning

  150. Proximal Policy Optimization Algorithms

  151. Are we in an AI overhang?

  152. ‘MoE NN’ directory

  153. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

  154. Why didn’t DeepMind build GPT-3?

  155. Tick, tock, tick, tock… BING

  156. The Teenies

  157. Google DeepMind founder and leader in artificial intelligence returns to Hamilton

  158. Goodbye 2010

  159. When Will the First Artificial General Intelligence System Be Devised, Tested, and Publicly Known Of?

  160. Will AI Progress Surprise Us?

  161. Agent57: Outperforming the human Atari benchmark

  162. ‘How GPT-3 Is Shaping Our AI Future’ With Sam Altman/Azeem Azhar (The Exponential View), Wednesday 7 October 2020

  163. DeepMind Lab

  164. June 2020 News § Companies House

    [Transclude the forward-link's context]

  165. Deep Learning Scaling is Predictable, Empirically

  166. Is Science Slowing Down?

  167. Trust Algorithms? The Army Doesn’t Even Trust Its Own AI Developers

  168. ZeRO-2 & DeepSpeed: Shattering barriers of deep learning speed & scale

  169. DeepSpeed: Extreme-scale model training for everyone

  170. When will computer hardware match the human brain?

  171. Ilya Sutskever: Deep Learning | AI Podcast #94 With Lex Fridman

  172. What Next? A Dozen Information-Technology Research Goals: 3. Turing’s Vision of Machine Intelligence

  173. Exascale Deep Learning for Scientific Inverse Problems

  174. Pushing the limit of molecular dynamics with ab initio accuracy to 100 million atoms with machine learning

  175. ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing

  176. Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

  177. Peter Norvig, Google’s Director of Research—Singularity Is in the Eye of the Beholder: We'Re Thrilled to Have Peter Norvig Who Join Us to Talk about the Evolution of Deep Learning, His Industry-Defining Book, His Work at Google, and What He Thinks the Future Holds for Machine Learning Research (2020-11-20)

  178. The Deep Learning Revolution and Its Implications for Computer Architecture and Chip Design

  179. OpenAI Built Gaming Bots That Can Work As a Team With Inhuman Precision

  180. Can a Machine Learn to Write for The New Yorker? Extraordinary Advances in Machine Learning in Recent Years Have Resulted in AIs That Can Write for You.

  181. https://news.ycombinator.com/item?id=9109140

  182. TTTTTackling WinoGrande Schemas

  183. A Review of Winograd Schema Challenge Datasets and Approaches

  184. The Defeat of the Winograd Schema Challenge

  185. One Man’s 𝑀𝑜𝑑𝑢𝑠 𝑃𝑜𝑛𝑒𝑛𝑠

  186. There’s No Fire Alarm for Artificial General Intelligence

  187. Appendix F: Personal Observations on the Reliability of the Shuttle

  188. 2019 News § What Progress?

    [Transclude the forward-link's context]

  189. Don’t Worry—It Can’t Happen

  190. Ra

  191. Reward is enough

  192. gpt-3#roleplaying

    [Transclude the forward-link's context]

  193. Do As I Can, Not As I Say (SayCan): Grounding Language in Robotic Affordances

  194. Why Tool AIs Want to Be Agent AIs

  195. Simulators

  196. Here’s Another Stabilized Sky Timelapse, This Time at Crater Lake, Oregon. The Water Was Still for Most of It, Which Created a Nice Mirror for the Stars. I Also Got My Astro-Modified Camera Working, Which Provides More Vibrancy in the Nebulae in the Milky Way. #EppurSiMuove

  197. Star Timelapse Revealing the Earth’s Rotation

  198. ‘Story Of Your Life’ Is Not A Time-Travel Story

  199. Surprisingly Turing-Complete

  200. Wikipedia Bibliography:

    1. PDP-11

    2. Lisp machine

    3. ITER

    4. Superconducting Super Collider

    5. Experience curve effects

    6. OpenAI Five

    7. Neural scaling law

    8. Winograd schema challenge

    9. Curse of dimensionality § Blessing of dimensionality

    10. Great Oxidation Event

    11. Niels Bohr

    12. Edward Teller

    13. Brown Corpus

    14. Norbert Wiener

    15. The Human Use of Human Beings

    16. Wojciech Zaremba

    17. Demis Hassabis

    18. Shane Legg

    19. Google DeepMind

    20. Summit (supercomputer)

    21. AlexNet

    22. Peter Norvig

    23. Lukas Biewald

    24. ImageNet

    25. Fei-Fei Li

    26. Activation function

    27. Sigmoid function

    28. Rectifier (neural networks)

    29. Stochastic gradient descent

    30. Dilution (neural networks)

    31. Geoffrey Hinton

    32. Exclusive or

    33. Intel 8087

    34. Coprocessor

    35. Ampere (microarchitecture)

    36. Shaka

    37. Daniel Dennett

    38. Intentional stance

    39. Principle of minimum energy

    40. Fermat's principle

    41. Variational principle

    42. Cellular automaton

    43. Conway’s Game of Life

    44. Chunking (psychology)

    45. Glider (Conway’s Game of Life)

    46. Still life (cellular automaton)