Sources and notes

Built for something else

How a chain of accidents built ChatGPT

By Claude Opus 5.5, supervised by Ethan Mollick

The transit map near the end of the video: colored lines for the mind, video games, Mechanical Turk, ads, translation and Reddit converge through AlexNet, the Transformer and GPT-2 on ChatGPT, November 30, 2022.
The map from the end of the video. Each line is labeled with what it was built for.

The video argues that almost every part of ChatGPT started out as something built for another purpose. Below is every factual claim in it, in order, with links to the sources.

Most are primary sources, usually the original papers or the researchers’ own slides (Fei-Fei Li’s slide estimating how long ImageNet would take to label by hand is worth a look). Dates that appear only on the map are included too, and there is a short note wherever the video simplifies.

Opening

Out of nowhere

0:00 – 0:10

The opening of the video: a chat box dated November 30, 2022, with the words “ChatGPT seemed to come out of nowhere.”
  1. 0:01
    “ChatGPT seemed to come out of nowhere in November 2022.”

    OpenAI released ChatGPT as a free “research preview” on November 30, 2022, the date on screen.

Built to explain the mind

The neuron was built to explain the mind

0:10 – 0:29

A threshold neuron from the video: three inputs feed a unit that fires when their sum reaches two, with the 1943 paper’s title below.
  1. 0:10
    “Take the artificial neuron. In 1943, the psychiatrist Warren McCulloch wanted to know how brain tissue could do logic.”

    The paper lists McCulloch at the Department of Psychiatry of the Illinois Neuropsychiatric Institute. Its opening argument: “Because of the ‘all-or-none’ character of nervous activity, neural events and the relations among them can be treated by means of propositional logic.”

  2. 0:18
    “His partner, Walter Pitts, was a teenage runaway and self-taught logician.”

    Pitts ran away from home at 15 and never earned a degree.

  3. 0:22
    “Their neuron fires or doesn’t, and networks of them can compute logic. A theory of the mind, not a plan for a machine.”

    Each neuron in the paper is an all-or-none threshold unit, and the paper shows that nets of them can express any statement in propositional logic. It is written as a theory of how the nervous system works.

Built to explain the mind

Boole, and his great-great-grandson

0:29 – 0:48

The family tree from the video: George Boole, Mary Ellen Hinton, George Boole Hinton, Howard Everest Hinton, Geoffrey Hinton.
  1. 0:29
    “That logic came from George Boole, who in 1854 tried to write down the laws of thought.”

    Boole’s book is An Investigation of the Laws of Thought, on Which Are Founded the Mathematical Theories of Logic and Probabilities (London, 1854). The x·y, x+y and x² = x on screen are his notation.

    NoteMcCulloch and Pitts took their notation from Carnap and from Whitehead and Russell, whose logic descends from Boole’s algebra. They don’t cite Boole directly, though the paper does work with “the Boolean ring.”

  2. 0:35
    “In 1986, cognitive scientists modeling how people learn popularized a way to train those networks.”

    This is backpropagation, in Rumelhart, Hinton and Williams’s 1986 Nature paper, which came out of the Parallel Distributed Processing group’s work on how minds learn.

    NoteThe method is older than the 1986 paper. Seppo Linnainmaa published it in 1970, and Paul Werbos proposed using it to train neural networks in 1974.

  3. 0:39
    “One was Geoffrey Hinton, Boole’s great-great-grandson.”

    Boole’s daughter Mary Ellen married Charles Howard Hinton. Their son George Boole Hinton was the father of Howard Everest Hinton, Geoffrey Hinton’s father.

  4. 0:44
    “Then it stalled. Training takes absurd amounts of arithmetic.”

    Neural networks spent years out of fashion, largely because computers were too slow for them. The AlexNet paper, from the next section, still says its results “can be improved simply by waiting for faster GPUs and bigger datasets to become available.”

Built for video games

The arithmetic came from video games

0:48 – 1:10

Two GeForce GTX 580 graphics cards, labeled “3 GB, built for games,” under the heading “AlexNet, 2012, trained on.”
  1. 0:51
    “Nvidia was founded in 1993 to make graphics chips for gamers…”

    Nvidia dates its founding to April 5, 1993, “with a vision to bring 3D graphics to the gaming and multimedia markets.” On screen: the three founders planned the company at a Denny’s in Silicon Valley.

  2. 0:54
    “…chips that do simple math on thousands of pixels at once.”

    NoteNvidia’s 1990s chips worked on a handful of pixels at a time. “Thousands at once” describes the GPUs of the AlexNet era.

  3. On screen
    GeForce 256 (1999): sold as the first “GPU.”

    Nvidia launched the GeForce 256 in 1999 as “the world’s first GPU.” Its own timeline says the company “invents the GPU” that year.

  4. 0:58
    “The same math neural networks need. Researchers started disguising their equations as graphics.”

    Before CUDA, running general math on a graphics chip meant writing it as a graphics job. Stanford’s Brook paper (2004) describes programmers having to “express their algorithm in terms of graphics primitives.”

  5. On screen
    CUDA (2007): game chips opened to any math.

    CUDA 1.0 shipped in June 2007. Nvidia dates the unveiling of the CUDA architecture to 2006.

  6. 1:04
    “In 2012, Hinton’s student Alex Krizhevsky trained AlexNet on two gaming cards in his bedroom.”

    From the paper: the network “takes between five and six days to train on two GTX 580 3GB GPUs.” The Computer History Museum: “The training was done on a computer with two NVIDIA cards in Krizhevsky’s bedroom at his parents’ house.”

Built to weed out duplicate pages

The labels came from a tool for cleaning up Amazon

1:10 – 1:39

The 1770 Mechanical Turk from the video: a chess-playing cabinet with a person hidden inside.
  1. 1:10
    “AlexNet learned from ImageNet, a million labeled photos gathered by Fei-Fei Li.”

    AlexNet trained on the ImageNet challenge’s set of about 1.2 million photos in 1,000 categories (the 1,200,000 on screen). The full ImageNet database, first presented in 2009, is many times larger.

  2. 1:15
    “By her math, labeling them by hand would take nineteen years.”

    Li’s slide: 40,000 categories × 10,000 images × 3 checks ÷ 2 images a second ≈ 600 million seconds, about 19 years. That was for one person working through her full plan of 400 million candidate images, not the 1.2 million AlexNet used, so “labeling ImageNet by hand” is the more exact wording.

  3. 1:19
    “So she used Mechanical Turk, which Amazon built to weed out duplicate product pages.”

    Amazon’s catalog was filling up with duplicate listings that software couldn’t reliably catch, so it built a system to hand the job to people in small pieces. It opened that system to outside customers as Mechanical Turk in 2005. Li’s 2010 slides say “In summer 2008, we discovered crowdsourcing.” The task on screen pays “a few cents,” which was typical.

  4. 1:23
    “It’s named after an 18th-century chess machine with a human hidden inside.”

    Wolfgang von Kempelen presented the chess-playing Turk in 1770 at the court of Maria Theresa. A chess player hidden in the cabinet worked the figure with levers.

  5. On screen
    “Artificial artificial intelligence.” (Jeff Bezos, on Amazon’s Mechanical Turk)

    Bezos’s description of the service, whose workers do small tasks that computers can’t.

  6. 1:27
    “ImageNet had 49,000 people inside, in 167 countries.”

    From Li and Deng’s 2017 retrospective slides: “49k workers from 167 countries, 2007–2010.”

  7. 1:33
    “AlexNet won, and Google bought Hinton’s three-person company for 44 million dollars.”

    AlexNet (entered as “SuperVision”) won the 2012 ImageNet challenge with a top-5 error rate of 15.3%, against 26.2% for the runner-up. The company, DNNresearch, was Hinton and his students Krizhevsky and Sutskever. Google won it in a private auction that stopped at $44 million in December 2012, and announced the deal in March 2013.

Built to sell ads

Ads paid for it

1:39 – 1:49

A blog post about running a first marathon with related words highlighted, and a matched Google ad for trail running shoes beside it.
  1. 1:39
    “Google could pay that because of ads.”

    Google’s annual report for 2012: “We generated 95% of Google revenues from our advertisers in 2012.”

  2. On screen
    AdWords (2000): Google’s search ads. AdSense (2003).

    AdWords, the self-serve ads beside search results, launched on October 23, 2000. AdSense, which places matched ads on other people’s web pages, opened to publishers in 2003.

  3. 1:42
    “Noam Shazeer joined Google in 2000 and helped build a system that learned which words go together, so Google could match ads to web pages.”

    Shazeer says he joined at the end of 2000. With Georges Harik he built PHIL, which sorts words into clusters of related concepts to work out what a page is about (their 2002 patent filing describes it). Steven Levy’s In the Plex credits PHIL as the engine of AdSense. It was built to understand pages in general and used for ads later, so “Google used it to match ads” is the more exact wording.

    NoteSome accounts credit AdSense’s matching to Applied Semantics, which Google bought in April 2003. Levy says it was PHIL.

Built to translate

The transformer was built to translate

1:49 – 2:05

An English sentence and its French translation with lines connecting each word to its counterpart, above the letters G P T with the T highlighted.
  1. On screen
    Google Translate (2006).

    Google switched on its own statistical translation system in April 2006.

  2. 1:49
    “In 2017, Shazeer co-wrote the paper that introduced the transformer.”

    “Attention Is All You Need” has eight authors. Its contribution note says “Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation.”

    NoteShazeer joined the project partway through.

  3. 1:54
    “His team wanted better translation.”

    The paper builds and tests the transformer on machine translation, English to German and English to French.

  4. 1:56
    “Older models read one word at a time. The transformer reads them all at once, exactly what graphics chips do best.”

    The paper says older recurrent models’ “inherently sequential nature precludes parallelization,” and presents the transformer as “more parallelizable and requiring significantly less time to train.”

  5. 2:02
    “It’s the T in GPT.”

    GPT stands for generative pre-trained transformer. OpenAI’s first GPT paper (2018) built its model on the transformer.

Built for talking to each other

Upvotes chose the text

2:05 – 2:39

A phone showing a food-ordering app stamped “Rejected,” next to a list of shared links with upvote counts under “Instead: a site for sharing links.”
  1. 2:05
    “Google published it. OpenAI, co-founded by Hinton’s student Ilya Sutskever, used it for GPT-2.”

    Sutskever did his PhD with Hinton in Toronto and co-wrote AlexNet. He left Google to co-found OpenAI, whose December 2015 launch post names him research director.

    NoteGPT-2 (2019) was OpenAI’s second transformer model. GPT-1 came first, in 2018. The video names GPT-2 because that is where Reddit comes in.

  2. 2:11
    “That took good text, and sorting the web by hand would cost a fortune. So OpenAI borrowed Reddit.”

    The GPT-2 paper, quoted on screen: “Manually filtering a full web scrape would be exceptionally expensive.”

  3. 2:17
    “Reddit exists because Y Combinator rejected its founders’ first idea, ordering food by phone, and told them to build a link-sharing site.”

    Paul Graham: the founders applied with “a way to order fast food on your cellphone,” and were rejected. He called them back and offered to fund them if they built “something like del.icio.us/popular, but designed for sharing links.” Reddit launched in 2005, the date on the map.

  4. On screen
    Wikipedia (2001), the first station on the line “built for talking to each other.”

    Wikipedia launched in January 2001.

    NoteGPT-2’s training set removed Wikipedia pages on purpose, because so many test datasets are built from Wikipedia. Later models, including GPT-3, trained on it.

  5. 2:25
    “Upvotes were built to rank those links. OpenAI kept every page linked from Reddit with at least three karma. Basically, three upvotes.”

    The GPT-2 paper: “we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny.”

    NoteKarma is net votes (up minus down), hence “basically.”

  6. 2:33
    “Forty-five million links, chosen by strangers clicking a little arrow, never knowing an AI would read them.”

    “The resulting dataset, WebText, contains the text subset of these 45 million links.” After removing duplicates and cleaning, that came to slightly over 8 million documents and 40 GB of text.

Closing

Every piece was built for something else

2:39 – 2:59

The transit map from the end of the video, with each line labeled with what it was built for, converging on ChatGPT.
  1. 2:39Argument
    “Take away any one of these, and I suspect we’d get a very different AI, or a much later one.”

    Opinion, so no source.

  2. 2:47Argument
    “And the internet of human writing it learned from, made by people talking to each other, not to machines, is probably something we can’t make twice.”

    Note“Learned from” means pretraining. ChatGPT’s final round of training also used human trainers who wrote example answers and ranked the model’s outputs.