Where does generative AI come from? 10 dates to understand ChatGPT (and what it means for your SME)
Why should an SME owner care about history?
For most people, generative AI starts on 30 November 2022, with ChatGPT. For researchers, ChatGPT was no surprise: it was the result of about ten years of work, much of it published openly.
Knowing that path is more than a curiosity. It explains three things every manager should know before handing work to an AI: why it works, why it gets things wrong and why prices are falling so fast.
The video: Monsieur Phi on the early days of machines that talk
Thibaut Giraud, a philosopher and science communicator known as Monsieur Phi, published an hour-and-twenty-minute video (in French) on this story on 10 October 2026, "Pourquoi le langage semblait inaccessible aux machines (et comment on le leur a donné par accident)" ("Why language seemed out of reach for machines, and how we gave it to them by accident"), starting from GPT-1 and even earlier. It is clear, well sourced and free of needless jargon: ideal for a train journey or a lunch break.
His thesis: between 2019 and 2020, humans lost their monopoly on language. Not because machines "think", but because they now produce language that is new, coherent and suited to a wide variety of situations, which until then was ours alone. The thesis is debated: for some philosophers, speaking requires an intention and an experience of the world that these models do not have.
The 10 dates to remember
| Date | Step | Key point |
|---|---|---|
| 2011 | A neural network writes letter by letter | About 5 million parameters, 100 MB of Wikipedia: grammatical sentences in places, but no logical thread |
| 2012 | AlexNet | 60 million parameters trained on graphics cards: deep learning takes over image recognition |
| January 2017 | Google's "mixture of experts" | A 137-billion-parameter model: the race for size begins |
| June 2017 | The Transformer ("Attention Is All You Need") | The architecture today's models are still built on |
| 2018 | GPT-1 (OpenAI) | Pre-training on about 7,000 books, then fine-tuning for each task |
| October 2018 | BERT (Google) | A score of 80.5 on the GLUE benchmark: pre-training becomes the norm |
| 14 February 2019 | GPT-2 | 1.5 billion parameters, 40 GB of web text; OpenAI delays the full release for fear of misuse |
| 13 March 2019 | Rich Sutton's "The Bitter Lesson" | General methods that leverage computing power end up winning, by a large margin |
| May 2020 | GPT-3 | 175 billion parameters; a few examples in the request are enough for it to perform a new task |
| 30 November 2022 | ChatGPT | A GPT trained for dialogue using feedback from human reviewers; GPT-4 follows on 14 March 2023 |
2011-2012: modest beginnings
In 2011, Ilya Sutskever (later a co-founder of OpenAI), James Martens and Geoffrey Hinton trained a network that reads text character by character and learns to guess the next one. From a distance the result looks like English, but it makes no sense. The following year, AlexNet, by Sutskever, Hinton and Alex Krizhevsky, showed that a large network trained on graphics cards crushes classic methods in image recognition.
2017: size and the Transformer
In January 2017, a Google team (including Noam Shazeer and Geoffrey Hinton) built a 137-billion-parameter model that activates only a small part of the network each time. In June, eight researchers, most of them at Google, published the Transformer, which drops word-by-word reading in favour of an "attention" mechanism: each word is related to all the others, and training parallelises much better. It is the T in GPT, which keeps only the part that generates text.
2018-2019: pre-train first, specialise later
GPT-1 introduced a simple idea: first train the model to predict how ordinary texts (books) continue, then adjust it to a specific task. BERT, at Google, pushed the same logic and topped the benchmarks. With GPT-2, OpenAI simply scaled up the model and the data. The model began to summarise, translate or answer questions without being trained to do so, still very unevenly, but this was new. OpenAI first released only a small version, then the full version on 5 November 2019.
At the same time, Rich Sutton, a pioneer of reinforcement learning, published a short text that became famous: over 70 years of research, general methods that benefit from growing computing power have ended up beating, by a large margin, approaches that hand-code human knowledge.
2020-2023: from the lab to the general public
GPT-3, more than a hundred times larger than GPT-2, can perform a new task from a few examples given in the request. ChatGPT added a decisive step: human reviewers wrote examples of good answers and then ranked the model's answers, and those judgements were used to train it to respond like a helpful conversation partner. The result: a tool that holds a conversation, open to everyone. You know the rest.
What this history means for your SME
Why it works: predicting what comes next, at very large scale
A language model does not look things up in a database of facts: it predicts the most plausible continuation of a text. At very large scale, that prediction captures grammar, style, a lot of knowledge and some reasoning. That is why it excels at drafting, rewording, summarising and translating, the tasks we recommend for getting started (see AI in SMEs: where to start, in French).
Why it gets things wrong: it prefers to guess
Predicting a plausible continuation is not the same as telling the truth. OpenAI acknowledged this in 2025: spelling and grammar appear everywhere in texts and are learned well, but a rare, isolated fact (a date, a number, a reference) cannot be deduced from any pattern. And the tests used to score models reward a random guess over an "I don't know".
In practice: always check the figures, names, dates and references produced by an AI, especially in a document going to a client or a public authority.
Why prices are falling: Sutton's lesson, wallet edition
According to Stanford University's AI Index 2025, the price of a model that achieves the same score as GPT-3.5 (the model behind the first ChatGPT) on a general-knowledge test fell from about 20 dollars per million tokens (pieces of words) in November 2022 to 0.07 dollars in October 2024: more than 280 times cheaper in under two years.
For an SME, this means:
- avoid long commitments to a tool or a price: review at least once a year;
- today's "mid-range" model is often as good as the best model of two years ago: paying for the most expensive one is not always worth it;
- fairly capable models now run on your own machines, without sending your data to the cloud (see local AI, in French).
Test yourself
At ExsIT
We help SMEs and freelancers choose the AI tool that fits their tasks, at the right price, and train the team to check what it produces. And when your data must not leave your premises, we install local models.
Sources
Anselin, G. (2026, May 7). Robots parleurs. La Vie des idées. https://laviedesidees.fr/Robots-parleurs
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020). Language models are few-shot learners. arXiv. https://doi.org/10.48550/arXiv.2005.14165
Centre national du cinéma et de l'image animée. (2021, September 24). Monsieur Phi, à la découverte de la philosophie. CNC. https://www.cnc.fr/creation-numerique/actualites/monsieur-phi-a-la-decouverte-de-la-philosophie_1357404
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. https://doi.org/10.48550/arXiv.1810.04805
Giraud, T. (2025). La parole aux machines : philosophie des grands modèles de langage [Words to the machines: philosophy of large language models]. Grasset.
Heaven, W. D. (2022, November 30). ChatGPT is OpenAI's latest fix for GPT-3. It's slick but still spews nonsense. MIT Technology Review. https://www.technologyreview.com/2022/11/30/1063878/openai-still-fixing-gpt3-ai-large-language-model/
Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv. https://doi.org/10.48550/arXiv.2509.04664
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25. https://papers.nips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html
Maslej, N. (2025, April 7). AI Index 2025: State of AI in 10 charts. Stanford HAI. https://hai.stanford.edu/news/ai-index-2025-state-of-ai-in-10-charts
Monsieur Phi [Giraud, T.]. (2026, October 10). Pourquoi le langage semblait inaccessible aux machines (et comment on le leur a donné par accident) [Video]. YouTube. https://youtu.be/PUxgpy81Cnc
OpenAI. (2019, February 14). Better language models and their implications. https://openai.com/index/better-language-models/
OpenAI. (2023, March 14). GPT-4. https://openai.com/index/gpt-4/
OpenAI. (2025, September 5). Why language models hallucinate. https://openai.com/index/why-language-models-hallucinate/
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., & Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv. https://doi.org/10.48550/arXiv.1701.06538
Sutskever, I., Martens, J., & Hinton, G. E. (2011). Generating text with recurrent neural networks. Proceedings of the 28th International Conference on Machine Learning (ICML 2011). https://www.cs.utoronto.ca/~ilya/pubs/2011/LANG-RNN.pdf
Sutton, R. (2019, March 13). The bitter lesson. Incomplete Ideas. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Synced. (2019, November 5). OpenAI releases 1.5 billion parameter GPT-2 model. https://syncedreview.com/2019/11/05/openai-releases-1-5-billion-parameter-gpt-2-model/
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. arXiv. https://doi.org/10.48550/arXiv.1706.03762