Where does generative AI come from? 10 dates to understand ChatGPT (and what it means for your SME)

Why should an SME owner care about history?

For most people, generative AI starts on 30 November 2022, with ChatGPT. For researchers, ChatGPT was no surprise: it was the result of about ten years of work, much of it published openly.

Knowing that path is more than a curiosity. It explains three things every manager should know before handing work to an AI: why it works, why it gets things wrong and why prices are falling so fast.

The video: Monsieur Phi on the early days of machines that talk

Thibaut Giraud, a philosopher and science communicator known as Monsieur Phi, published an hour-and-twenty-minute video (in French) on this story on 10 October 2026, "Pourquoi le langage semblait inaccessible aux machines (et comment on le leur a donné par accident)" ("Why language seemed out of reach for machines, and how we gave it to them by accident"), starting from GPT-1 and even earlier. It is clear, well sourced and free of needless jargon: ideal for a train journey or a lunch break.

Who is Monsieur Phi? Thibaut Giraud, who holds a doctorate in philosophy and is a former secondary-school philosophy teacher, has been publishing videos on YouTube since 2016: logic, philosophy of science, ethics, with humour and pop-culture references. In recent years he has taken a close interest in large language models and devoted a book of nearly 500 pages to them, “La Parole aux machines” (Grasset, 2025, in French). The review La Vie des idées calls it a teaching success while questioning some of its theses. His videos are long, well argued and sourced: a good antidote to both over-enthusiastic and over-sceptical takes.

His thesis: between 2019 and 2020, humans lost their monopoly on language. Not because machines "think", but because they now produce language that is new, coherent and suited to a wide variety of situations, which until then was ours alone. The thesis is debated: for some philosophers, speaking requires an intention and an experience of the world that these models do not have.

Two notes while watching: the full version of GPT-2 was released in November 2019 (not 2018, a slip of the tongue). And a few recent AI successes mentioned in passing (maths problems solved, an award-level novel…) have not been checked here: take them as examples, not established facts.

The 10 dates to remember

Date Step Key point
2011 A neural network writes letter by letter About 5 million parameters, 100 MB of Wikipedia: grammatical sentences in places, but no logical thread
2012 AlexNet 60 million parameters trained on graphics cards: deep learning takes over image recognition
January 2017 Google's "mixture of experts" A 137-billion-parameter model: the race for size begins
June 2017 The Transformer ("Attention Is All You Need") The architecture today's models are still built on
2018 GPT-1 (OpenAI) Pre-training on about 7,000 books, then fine-tuning for each task
October 2018 BERT (Google) A score of 80.5 on the GLUE benchmark: pre-training becomes the norm
14 February 2019 GPT-2 1.5 billion parameters, 40 GB of web text; OpenAI delays the full release for fear of misuse
13 March 2019 Rich Sutton's "The Bitter Lesson" General methods that leverage computing power end up winning, by a large margin
May 2020 GPT-3 175 billion parameters; a few examples in the request are enough for it to perform a new task
30 November 2022 ChatGPT A GPT trained for dialogue using feedback from human reviewers; GPT-4 follows on 14 March 2023

2011-2012: modest beginnings

In 2011, Ilya Sutskever (later a co-founder of OpenAI), James Martens and Geoffrey Hinton trained a network that reads text character by character and learns to guess the next one. From a distance the result looks like English, but it makes no sense. The following year, AlexNet, by Sutskever, Hinton and Alex Krizhevsky, showed that a large network trained on graphics cards crushes classic methods in image recognition.

2017: size and the Transformer

In January 2017, a Google team (including Noam Shazeer and Geoffrey Hinton) built a 137-billion-parameter model that activates only a small part of the network each time. In June, eight researchers, most of them at Google, published the Transformer, which drops word-by-word reading in favour of an "attention" mechanism: each word is related to all the others, and training parallelises much better. It is the T in GPT, which keeps only the part that generates text.

2018-2019: pre-train first, specialise later

GPT-1 introduced a simple idea: first train the model to predict how ordinary texts (books) continue, then adjust it to a specific task. BERT, at Google, pushed the same logic and topped the benchmarks. With GPT-2, OpenAI simply scaled up the model and the data. The model began to summarise, translate or answer questions without being trained to do so, still very unevenly, but this was new. OpenAI first released only a small version, then the full version on 5 November 2019.

At the same time, Rich Sutton, a pioneer of reinforcement learning, published a short text that became famous: over 70 years of research, general methods that benefit from growing computing power have ended up beating, by a large margin, approaches that hand-code human knowledge.

2020-2023: from the lab to the general public

GPT-3, more than a hundred times larger than GPT-2, can perform a new task from a few examples given in the request. ChatGPT added a decisive step: human reviewers wrote examples of good answers and then ranked the model's answers, and those judgements were used to train it to respond like a helpful conversation partner. The result: a tool that holds a conversation, open to everyone. You know the rest.

What this history means for your SME

Why it works: predicting what comes next, at very large scale

A language model does not look things up in a database of facts: it predicts the most plausible continuation of a text. At very large scale, that prediction captures grammar, style, a lot of knowledge and some reasoning. That is why it excels at drafting, rewording, summarising and translating, the tasks we recommend for getting started (see AI in SMEs: where to start, in French).

Why it gets things wrong: it prefers to guess

Predicting a plausible continuation is not the same as telling the truth. OpenAI acknowledged this in 2025: spelling and grammar appear everywhere in texts and are learned well, but a rare, isolated fact (a date, a number, a reference) cannot be deduced from any pattern. And the tests used to score models reward a random guess over an "I don't know".

In practice: always check the figures, names, dates and references produced by an AI, especially in a document going to a client or a public authority.

Why prices are falling: Sutton's lesson, wallet edition

According to Stanford University's AI Index 2025, the price of a model that achieves the same score as GPT-3.5 (the model behind the first ChatGPT) on a general-knowledge test fell from about 20 dollars per million tokens (pieces of words) in November 2022 to 0.07 dollars in October 2024: more than 280 times cheaper in under two years.

For an SME, this means:

  • avoid long commitments to a tool or a price: review at least once a year;
  • today's "mid-range" model is often as good as the best model of two years ago: paying for the most expensive one is not always worth it;
  • fairly capable models now run on your own machines, without sending your data to the cloud (see local AI, in French).

History is not a promise: nothing guarantees that progress will continue at the same pace. Judge a tool on your own tasks, not on announcements.

Book La Parole aux machines by Thibaut Giraud Further reading: La Parole aux machines, by Thibaut Giraud (Monsieur Phi), in French.
A philosophy of large language models: the book in which he develops his thesis that humans have lost their monopoly on language. Grasset, 2025.
See the book on Amazon Affiliate link

Test yourself

Which neural network architecture, published by Google in 2017, underpins ChatGPT and today's language models?

Show the answer

Answer: The Transformer

The paper “Attention Is All You Need” (June 2017) introduced the Transformer, the T in GPT.

Why did OpenAI not immediately release the full version of its GPT-2 language model in February 2019?

Show the answer

Answer: For fear of malicious use

OpenAI cited the risk of misuse (disinformation, spam) and only released the full 1.5-billion-parameter version in November 2019.

According to Stanford's AI Index 2025, by how much did the cost of an AI model as capable as the first ChatGPT fall between late 2022 and late 2024?

Show the answer

Answer: More than 280 times cheaper

The price fell from about 20 dollars to 0.07 dollars per million tokens between November 2022 and October 2024.

At ExsIT

We help SMEs and freelancers choose the AI tool that fits their tasks, at the right price, and train the team to check what it produces. And when your data must not leave your premises, we install local models.

Sources

Anselin, G. (2026, May 7). Robots parleurs. La Vie des idées. https://laviedesidees.fr/Robots-parleurs

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020). Language models are few-shot learners. arXiv. https://doi.org/10.48550/arXiv.2005.14165

Centre national du cinéma et de l'image animée. (2021, September 24). Monsieur Phi, à la découverte de la philosophie. CNC. https://www.cnc.fr/creation-numerique/actualites/monsieur-phi-a-la-decouverte-de-la-philosophie_1357404

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. https://doi.org/10.48550/arXiv.1810.04805

Giraud, T. (2025). La parole aux machines : philosophie des grands modèles de langage [Words to the machines: philosophy of large language models]. Grasset.

Heaven, W. D. (2022, November 30). ChatGPT is OpenAI's latest fix for GPT-3. It's slick but still spews nonsense. MIT Technology Review. https://www.technologyreview.com/2022/11/30/1063878/openai-still-fixing-gpt3-ai-large-language-model/

Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv. https://doi.org/10.48550/arXiv.2509.04664

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25. https://papers.nips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html

Maslej, N. (2025, April 7). AI Index 2025: State of AI in 10 charts. Stanford HAI. https://hai.stanford.edu/news/ai-index-2025-state-of-ai-in-10-charts

Monsieur Phi [Giraud, T.]. (2026, October 10). Pourquoi le langage semblait inaccessible aux machines (et comment on le leur a donné par accident) [Video]. YouTube. https://youtu.be/PUxgpy81Cnc

OpenAI. (2019, February 14). Better language models and their implications. https://openai.com/index/better-language-models/

OpenAI. (2023, March 14). GPT-4. https://openai.com/index/gpt-4/

OpenAI. (2025, September 5). Why language models hallucinate. https://openai.com/index/why-language-models-hallucinate/

Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf

Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., & Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv. https://doi.org/10.48550/arXiv.1701.06538

Sutskever, I., Martens, J., & Hinton, G. E. (2011). Generating text with recurrent neural networks. Proceedings of the 28th International Conference on Machine Learning (ICML 2011). https://www.cs.utoronto.ca/~ilya/pubs/2011/LANG-RNN.pdf

Sutton, R. (2019, March 13). The bitter lesson. Incomplete Ideas. http://www.incompleteideas.net/IncIdeas/BitterLesson.html

Synced. (2019, November 5). OpenAI releases 1.5 billion parameter GPT-2 model. https://syncedreview.com/2019/11/05/openai-releases-1-5-billion-parameter-gpt-2-model/

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. arXiv. https://doi.org/10.48550/arXiv.1706.03762

← Back to articles