The history of artificial intelligence

Summarise this article with AI

Updated 9 July 2026

The history of artificial intelligence is fascinating and spans several decades – and when you think about it, perhaps even centuries. From the first philosophical questions about the nature of thinking and the possibility of machines thinking, to the sophisticated algorithms and applications that today help us in our everyday routine. AI has become a key element of modern scientific and technological progress. Its development is a story not only of technology, but also of the human spirit, of the desire to understand, to innovate and to keep searching for new possibilities. Whether we look at attempts to imitate human behaviour in early automata, or at today’s breakthroughs in deep learning, AI forces us to reconsider what it means to be human and how we define intelligence in ourselves. In this article we will explore how this field developed from its beginnings to the present day. The text draws on primary scientific works and standard survey publications of the discipline (Nilsson, 2010; Russell & Norvig, 2021), so that it can also serve as a reliable starting point for academic and student work.

Antiquity and the Middle Ages

From the earliest times of human history, philosophers have been occupied with the question of the nature of thinking. What does it mean to “think”? And is it even possible for a non-human or inanimate object to imitate this mysterious process?

In ancient Greece, Plato in his Republic (Politeia) touched on the question of the real versus the apparent – a theme that resonates in today’s discussions about the simulation of thinking. Aristotle, in his treatise On the Soul (De anima), examined the nature of the human soul and its faculties, including thinking; for him reason represented the highest and most divine faculty of the soul, but the idea that a machine could possess such a faculty was completely alien to him. Aristotle’s syllogistic – the doctrine of formally valid inference – is nevertheless considered one of the first attempts at a mechanisable description of thinking and is cited as the oldest intellectual root of artificial intelligence (Russell & Norvig, 2021).

If we leave aside the convictions of animists, who ascribed spirit and soul even to natural objects, we find in modern philosophy thinkers whose reflections anticipated today’s debate about machine intelligence with surprising precision. Thomas Hobbes in Leviathan (1651) declared that “reasoning is nothing but reckoning” – he thus conceived of thinking as a symbolic operation that can in principle be mechanised (Russell & Norvig, 2021). René Descartes, by contrast, in his Discourse on the Method (1637) argued that a machine could never use language meaningfully in an arbitrary situation – thereby inadvertently formulating a criterion that three centuries later Alan Turing turned into his famous test. Baruch Spinoza then, in his Ethics developed a monistic metaphysics in which thought and extension represent two attributes of a single substance; some interpreters derive from this a panpsychist reading according to which the mental aspect belongs to a certain degree to all things, and therefore potentially to machines as well – this is, however, an interpretation that is a matter of philosophical dispute.

Medieval scholasticism brought further perspectives. Thomas Aquinas, one of the most important medieval theologians and philosophers, adopted many of Aristotle’s ideas and examined them in the context of Christian theology (Summa theologiae). He was convinced that the intellect is not a faculty separate from the soul, but a part of it – for him, to have a soul meant to have reason and intelligence. The idea that a machine could imitate human thinking would have been unacceptable to him, because nothing without a soul can possess intelligence.

The question of whether it is possible to imitate human intelligence nevertheless remained provocative. The philosophical questions posed in antiquity and the Middle Ages thus created the intellectual bedrock on which modern artificial intelligence grew up centuries later.

Do you want to learn much more about AI tools? Exclusively and first? Join my community on Patreon. A regular dose of tips, tricks and how-tos is waiting for you…

Automata and mechanical predecessors: the dream of an artificial human

The dream of an artificial being is far older than computers. Greek mythology knows the bronze giant Talos, said to have been made by Hephaestus himself to guard Crete. Jewish tradition tells of the golem – an artificial man brought to life from clay – and its most famous version is linked directly to Prague and Rabbi Judah Loew at the court of Rudolf II. These myths are not mere curiosities: they formulate the ancient intuition that intelligent behaviour might perhaps be created artificially, and at the same time they warn against losing control over one’s own creation – a theme that keeps returning in debates about AI to this day (Nilsson, 2010).

Alongside the myths there was also a real mechanical tradition. Heron of Alexandria built programmable theatre automata driven by weights and ropes in the first century AD, the Arab engineer al-Jazari described dozens of ingenious automata including mechanical musicians in the 13th century, and in the 18th century Vaucanson’s mechanical figures astonished Europe. A special chapter is Wolfgang von Kempelen’s “Mechanical Turk” of 1770 – an alleged chess automaton that defeated human opponents for decades until it turned out that a hidden human player was sitting inside it. It was therefore a hoax, but a culturally significant one: the public was willing to believe that a machine could think, two centuries before the first computer. The name of the crowdsourcing service Amazon Mechanical Turk, where the real work “for the machine” is again done by people, ironically refers to this story today.

In parallel, the idea of mechanising calculation itself was developing. Blaise Pascal built a mechanical adding machine in the 1640s, Gottfried Wilhelm Leibniz improved it to handle multiplication and moreover dreamed of a characteristica universalis – a universal formal language in which disputes could be settled by calculation. The culmination of this line was Charles Babbage’s design of the Analytical Engine in the 19th century, the first universal programmable computer, even though it was never completed. The mathematician Ada Lovelace, who wrote the first published algorithm for the Analytical Engine, presciently noted that the machine “has no pretensions whatever to originate anything” and can do only what a human instructs it to do. This “Lady Lovelace objection” became a classic argument against machine intelligence – and a hundred years later Alan Turing devoted a separate passage of his famous article to disputing it (Turing, 1950).

At the threshold of the twentieth century, this dream was given the name by which the world still calls it today – and it happened in Czech. Karel Čapek, in his drama R.U.R. (Rossum’s Universal Robots) of 1920, used the word “robot” for the first time, derived from robota, meaning forced labour; according to his own later testimony, the word itself was suggested to him by his brother Josef (Čapek, 1920). The play premiered in 1921 at the National Theatre in Prague, quickly travelled around the world’s stages, and the word robot passed practically unchanged into dozens of languages – making it probably the most successful Czech contribution to the world’s vocabulary. Čapek’s robots are not mechanical machines, but artificially created organic beings indistinguishable from humans, and the plot of the play culminates in their revolt and the extermination of humanity. A hundred years before today’s debates about the risks of artificial intelligence, Čapek thus formulated their central motif: what happens when a creation surpasses its creator and slips out of his control. The Prague of the golem and Čapek’s imagination thus stand at the very linguistic and intellectual birth of a field that was to be formally founded only thirty-six years later at Dartmouth.

The twentieth century and the birth of artificial intelligence

In the 1930s the British mathematician and logician Alan Turing introduced the revolutionary concept known today as the Turing machine (Turing, 1937). This abstract computational model, designed to simulate the logical operations of any algorithm, laid the foundations of the modern theory of computability: it showed that if a problem can be solved by an algorithm, it can be solved by a Turing machine. This opened the way to universal computing machines – computers.

An important milestone came as early as 1943, when the neurophysiologist Warren McCulloch and the logician Walter Pitts published the first mathematical model of an artificial neuron and showed that networks of such simplified neurons can in principle compute any logical function (McCulloch & Pitts, 1943). It is precisely this work that is regarded as the theoretical beginning of artificial neural networks – the direction that seven decades later came to dominate the whole field.

In 1950 Turing entered the field of philosophy and cognitive science with the question: “Can machines think?” Instead of a direct answer he proposed, in the article “Computing Machinery and Intelligence” published in the journal Mind, the experiment known today as the Turing test (Turing, 1950). If a computer program managed to conduct a conversation with a human in such a way that the judge could not decide whether they were communicating with a machine or with another person, such a machine could be considered “intelligent”.

The end of the 1940s also brought the broader intellectual ferment out of which artificial intelligence grew. Norbert Wiener founded cybernetics – the science of control and communication in machines and living organisms, which for the first time systematically studied feedback and goal-directed behaviour (Wiener, 1948). In the same year Claude Shannon laid the foundations of information theory and showed that information can be measured and transmitted independently of its meaning. The first electronic computers were being built, and with them a community of scientists convinced that thinking could be studied as information processing.

Dartmouth 1956

The summer research workshop at Dartmouth College in 1956 is conventionally regarded as the official birth of artificial intelligence as a scientific discipline. The term “artificial intelligence” itself, however, first appeared as early as in the preparatory project proposal of 31 August 1955, whose authors were John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon (McCarthy et al., 1955/2006). The proposal was based on the conjecture that every aspect of learning or any other feature of intelligence can in principle be described so precisely that a machine can be made to simulate it. Besides the organisers, other pioneers of the field took part in the workshop, among them Allen Newell, Herbert Simon, Arthur Samuel and Oliver Selfridge; the historic centres of AI research at Carnegie Mellon, MIT and Stanford subsequently took shape out of Dartmouth (Cordeschi, 2007). The conference gave rise to a wave of enthusiasm and optimism for future research.

The optimism had real foundations. As early as 1956 Newell and Simon demonstrated the Logic Theorist program, which could prove mathematical theorems on its own and is often called the first artificial intelligence program (Newell & Simon, 1956). In 1958 John McCarthy designed the LISP programming language, which for decades became the main tool of AI research (McCarthy, 1960). In 1958 Frank Rosenblatt presented the perceptron – the first learning neural network implemented in a real machine, which could classify simple patterns (Rosenblatt, 1958). And at IBM, Arthur Samuel developed a checkers-playing program that improved by playing against itself; in his 1959 article he used the term “machine learning” in its present-day meaning for the first time (Samuel, 1959).

In the 1960s the first well-known applications began to appear. Joseph Weizenbaum created the ELIZA program, which simulated a conversation with a psychotherapist using simple text-rewriting rules (Weizenbaum, 1966); Weizenbaum himself was later taken aback by how easily people ascribed understanding to the program. Terry Winograd designed the SHRDLU system, which could communicate in natural language about a simplified “world of blocks” (Winograd, 1972). These systems represented the first steps towards the practical realisation of ideas that had until then been purely theoretical – and the scientific community, the media and the public began to believe that machines would soon reach human-level intelligence.

The AI winters

After the wave of enthusiasm of the 1960s came periods of disappointment and scepticism, today referred to as the “AI winters”. Historians of the field distinguish two separate episodes: the first in the mid-1970s and the second at the turn of the 1980s and 1990s (Nilsson, 2010; Russell & Norvig, 2021).

The first winter was brought about in part by two influential documents. The book Perceptrons by Marvin Minsky and Seymour Papert (1969) mathematically demonstrated fundamental limitations of single-layer perceptrons – for example the inability to learn the logical XOR function – and because it was not then known how to train multi-layer networks efficiently, both interest in and funding for neural network research fell sharply for more than a decade. Then in 1973 the British Science Research Council published a report by the mathematician James Lighthill, which sharply criticised the field’s failure to fulfil its own grand goals and drew attention in particular to the problem of combinatorial explosion – the fact that methods which work on toy tasks collapse computationally when scaled to real problems (Lighthill, 1973). The result was a drastic reduction of government funding for AI in the United Kingdom and subsequently in the United States as well.

The 1980s brought a temporary revival thanks to expert systems – programs that captured the knowledge of human specialists in the form of “if–then” rules. The MYCIN system for diagnosing bacterial infections (Shortliffe, 1976) and the commercially successful XCON from Digital Equipment Corporation showed the practical value of AI and launched a billion-dollar industry, which the ambitious Japanese fifth-generation computer project also joined. The market for expert systems collapsed at the end of the 1980s, however – the systems turned out to be brittle, expensive to maintain and unable to learn from data – and with it came the second AI winter, accompanied by another deep round of cuts to research budgets (Nilsson, 2010).

Besides exaggerated expectations, a persistent problem was the lack of computing power and data. The computers of the 1970s and 1980s had negligible memory capacity by today’s standards, so methods requiring large volumes of training data were practically impossible to implement. At the same time, however, the AI winters were a period of self-reflection and quiet progress: in 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams published in the journal Nature the backpropagation algorithm, which finally made it possible to train multi-layer neural networks efficiently (Rumelhart et al., 1986). This laid the technical foundation for the renaissance that was to come a quarter of a century later.

The same period also saw the most influential philosophical critique of the whole project. In 1980 John Searle published the thought experiment known as the Chinese room: a person shut in a room mechanically manipulates Chinese characters according to the rules in a manual so successfully that from the outside it looks as though he understands Chinese – even though he does not understand a word of the language. Searle argued from this that mere manipulation of symbols according to syntactic rules can never create genuine understanding, and he introduced the distinction, still used today, between “weak AI” (the machine merely simulates intelligence) and “strong AI” (the machine really thinks and understands) (Searle, 1980). The debate this argument unleashed has taken on new urgency with the arrival of large language models – the question of whether a system generating convincing text also “understands” its output is precisely Searle’s question.

The end of the twentieth century: the revival of AI

After winter came spring – a period of new, this time more sober optimism. The 1990s and the beginning of the new millennium witnessed a shift from hand-programmed rules to machine learning from data, grounded in statistics and rigorous mathematical foundations (Russell & Norvig, 2021).

Algorithms still in use today were created. Support Vector Machines found application in handwriting recognition as well as image classification (Cortes & Vapnik, 1995). Decision trees provided an intuitive way of modelling decision processes, used for example in medical diagnostics or credit scoring. The first recommendation systems appeared, pioneered in e-commerce by Amazon with personalised recommendations based on a combination of a user’s purchase history and the behaviour of other customers.

In 1997 the chess computer IBM Deep Blue defeated the reigning world champion Garry Kasparov in a rematch by 3.5 : 2.5 – a year earlier Kasparov had still won (Campbell et al., 2002). It was the first defeat of a reigning world chess champion by a computer in a match under standard tournament conditions, and a symbolic signal of the growing capabilities of computing systems; it is only fair to add, however, that Deep Blue relied above all on brute search power and hand-tuned chess heuristics, not on learning.

Users of button mobile phones remember the groundbreaking T9 function (Text on 9 keys), a predictive typing system that guessed the intended word from the sequence of pressed keys on the basis of a dictionary and statistics. It was one of the first mass-adopted applications of statistical language processing in the pocket of an ordinary user. Applications such as Dragon NaturallySpeaking made speech-to-text conversion accessible and laid the foundation for later voice assistants of the Siri, Google Assistant or Alexa type. As unsolicited mail increased, statistical (especially Bayesian) spam filters gained ground, one of the first examples of machine learning deployed on a mass scale.

Google, founded in 1998, built its search on the PageRank algorithm – which is, however, a graph algorithm, not machine learning. Machine learning was deployed by Google gradually, first in supporting tasks such as spam filtering or ad targeting, and it was only later that it was involved to a greater extent in the very core of ranking search results (the RankBrain system was introduced in 2015).

During the 1990s and the beginning of the new millennium, computing technology made an enormous leap. Processors got faster, memory got cheaper and the expansion of the internet began to generate previously unimaginable volumes of data. It was precisely the convergence of three factors – large data sets, computing power (especially graphics cards) and algorithmic advances – that later made the deep learning revolution possible (LeCun et al., 2015). The turn of the century was thus crucial for the transformation of artificial intelligence from a purely academic field into a practical technology with a real impact on society and industry.

A harbinger of the following era was the performance of the IBM Watson system in the American quiz show Jeopardy! in 2011. Watson, built on the DeepQA architecture combining natural language processing, searching in knowledge sources and statistical evaluation of hypotheses, defeated the two most successful human champions in the history of the show (Ferrucci et al., 2010). It thereby showed that machines can work with open-ended questions in natural language – that is, precisely the type of task on which symbolic AI had been failing for decades.

The era of deep learning

The period from roughly 2012 onwards can safely be called the era of deep learning. Machine learning is a method that allows systems to learn and improve from experience without explicit programming; machine learning models predict outputs on the basis of historical data. Deep learning (deep learning) is a subset of machine learning that uses neural networks with a large number of layers (so-called deep neural networks) for learning and is able to create ever more abstract representations of data by itself – layer by layer (LeCun et al., 2015).

The symbolic beginning of this era is 30 September 2012, when the convolutional neural network AlexNet by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton won the ImageNet image recognition competition with an error rate of 15.3 %, while the best traditional method achieved 26.1 % (Krizhevsky et al., 2012). Such a marked lead convinced the field that deep neural networks – trained on large data using graphics cards – outperform previous approaches by a wide margin. An indispensable precondition was the existence of the ImageNet data set itself, with millions of annotated images (Deng et al., 2009). The survey article by the founding trio of deep learning – Yann LeCun, Yoshua Bengio and Geoffrey Hinton – in the journal Nature then canonised this turn (LeCun et al., 2015); all three later received the Turing Award for 2018.

In 2013 a Google team introduced the Word2Vec method for creating vector representations of words on the basis of context (Mikolov et al., 2013). Thanks to it, machines were able to capture the meaning of and relationships between words better – a famous example being vector arithmetic of the type “king − man + woman ≈ queen” – which significantly advanced the whole of natural language processing. The main author of the concept was the Czech researcher Tomáš Mikolov, who likes to illustrate his intuition about how the meaning of words can be derived from their surroundings with his own experience from a border region and his understanding of the related Polish language. Distributed word representations became one of the building blocks on which large language models later grew.

In 2014 Ian Goodfellow and colleagues introduced the concept of generative adversarial networks (GANs), in which two neural networks – a generator and a discriminator – train each other in competition (Goodfellow et al., 2014). This opened the way to generating images, music and text almost indistinguishable from human creation and foreshadowed today’s generative AI.

The program AlphaGo from DeepMind, combining deep learning, reinforcement learning and tree search, defeated the three-time European champion Fan Hui in October 2015 – as the first program to beat a professional Go player without a handicap – and in March 2016 it won 4 : 1 against Lee Sedol, one of the best players in the world (Silver et al., 2016). In 2017 it also defeated the world number one Ke Jie. This was a breakthrough, because Go, with its enormous state space, had long resisted the methods that had succeeded in chess.

DeepMind subsequently turned the same principles to a scientific problem with enormous practical reach. The AlphaFold 2 system won the CASP14 competition in predicting the three-dimensional structure of proteins from their amino acid sequence by a wide margin in 2020 – a task that had resisted biologists for half a century – and its description published in Nature became one of the most cited scientific articles of the decade (Jumper et al., 2021). The freely accessible database of predicted structures of practically all known proteins now serves researchers all over the world and represents what is probably the most convincing example so far of AI as a tool accelerating basic science.

The era of transformers and large language models

The turning point for natural language processing was 2017, when Google researchers led by Ashish Vaswani published the article “Attention Is All You Need” and with it the transformerarchitecture, built on the attention mechanism (Vaswani et al., 2017). Transformers removed key limitations of recurrent networks – they made parallel training and the capturing of relationships between distant parts of a text possible – and became the basis of practically all of today’s large language models.

Two influential families of models grew up on the transformer architecture. BERT from Google (Devlin et al., 2019) and its variants (RoBERTa, DistilBERT and others) became the standard for text understanding. The second line was founded by OpenAI with the GPT (Generative Pre-trained Transformer) model, which combined generative pre-training on a large corpus with subsequent fine-tuning for specific tasks (Radford et al., 2018). Its successor GPT-2 of 2019 generated text so convincing that OpenAI opted for a gradual, staged release of the model with reference to the risk of misuse (Radford et al., 2019) – which triggered one of the first major public debates about the responsible release of AI models.

In 2020 OpenAI introduced the GPT-3 model with 175 billion parameters, which demonstrated a surprising ability to learn new tasks from just a few examples given directly in the prompt, without further training (Brown et al., 2020). It thus surpassed by an order of magnitude the largest language model until then, Microsoft’s Turing-NLG with 17 billion parameters. The bet on size was not blind: in the same year OpenAI researchers empirically described the so-called scaling laws, according to which the capabilities of language models improve predictably with a growing number of parameters, volume of data and computing power (Kaplan et al., 2020). It was precisely this finding that legitimised the race to enlarge models that followed. A parameter is a variable that is adjusted during the training of a model – in neural networks these are the weights determining how strongly one neuron influences another. A model with more parameters can capture more complex patterns in data; imagine it as painting with a richer palette of colours. The number of parameters of top models grew exponentially at that time: in April 2022 Google introduced the

In parallel with text, a revolution was taking place in image generation. Diffusion models, which learn to gradually “denoise” an image out of random noise (Ho et al., 2020), surpassed the older generative adversarial networks and became the basis of a new generation of tools. In 2021 OpenAI introduced the DALL-Emodel, generating images from a text description (Ramesh et al., 2021), and in 2022 came Midjourney and the open Stable Diffusion built on efficient latent diffusion (Rombach et al., 2022). Within a single year, the generation of photorealistic images from text thus turned from a research curiosity into a mass-available tool – with consequences for the creative industries, copyright and the trustworthiness of visual information that society is still absorbing today.

Generative AI entered wide public awareness on 30 November 2022, when OpenAI made ChatGPT available – a conversational interface on top of a model of the GPT-3.5 line, fine-tuned with reinforcement learning from human feedback (RLHF), which teaches the model to answer helpfully and in line with human preferences (Ouyang et al., 2022). The service gained an estimated one hundred million users within two months and thus became one of the fastest-growing consumer applications in history. It was followed by the multimodal GPT-4 (March 2023; OpenAI, 2023) and competing models from other laboratories – Claude from Anthropic, Gemini from Google, the open Llama models from Meta and the models of the company Mistral, which made top language models available for local operation and research as well. From 2024 onwards the field’s attention has been shifting to so-called reasoning models, which perform explicit multi-step reasoning before answering, and to agentic systems capable of independently carrying out longer tasks.

With the growing social impact of the technology came regulation as well. In 2024 the European Union adopted, as the first in the world, a comprehensive legal framework for artificial intelligence – the regulation known as the AI Act, which classifies AI systems according to the degree of risk and lays down obligations for their providers and users (EU Regulation 2024/1689). The history of artificial intelligence has thus entered a phase in which it is no longer written only by scientists and engineers, but also by legislators.

A symbolic full stop after this period – and at the same time a recognition of the whole journey from McCulloch and Pitts to the present – is the year 2024, when the Royal Swedish Academy of Sciences awarded the Nobel Prize in Physics to John Hopfield and Geoffrey Hinton for foundational discoveries that enable machine learning with artificial neural networks (The Royal Swedish Academy of Sciences, 2024a). The same academy simultaneously honoured Demis Hassabis and John Jumper of DeepMind with the Nobel Prize in Chemistry for AlphaFold (The Royal Swedish Academy of Sciences, 2024b) – in a single year artificial intelligence thus became the subject of two Nobel Prizes, something the pioneers from Dartmouth in 1956 could hardly have imagined.

The era of deep learning and large language models has meant a breakthrough in many sectors and has created possibilities that were still considered science fiction fifteen years ago. And with continuing progress in hardware, algorithms and data volumes, we can expect that the pace of development will not slow down any time soon.

Do you want to learn much more about AI tools? Exclusively and first? Join my community on Patreon. A regular dose of tips, tricks and how-tos is waiting for you…

References

  • Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
  • Campbell, M., Hoane, A. J., & Hsu, F. (2002). Deep Blue. Artificial Intelligence, 134(1–2), 57–83. https://doi.org/10.1016/S0004-3702(01)00129-1
  • Cordeschi, R. (2007). AI turns fifty: Revisiting its origins. Applied Artificial Intelligence, 21(4–5), 259–279. https://doi.org/10.1080/08839510701252304
  • Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. https://doi.org/10.1007/BF00994018
  • Čapek, K. (1920). R.U.R. (Rossum’s Universal Robots): A Collective Drama in a Comic Prologue and Three Acts. Aventinum.
  • Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (pp. 248–255). IEEE. https://doi.org/10.1109/CVPR.2009.5206848
  • Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 4171–4186). ACL. https://doi.org/10.18653/v1/N19-1423
  • Ferrucci, D., Brown, E., Chu-Carroll, J., Fan, J., Gondek, D., Kalyanpur, A. A., Lally, A., Murdock, J. W., Nyberg, E., Prager, J., Schlaefer, N., & Welty, C. (2010). Building Watson: An overview of the DeepQA project. AI Magazine, 31(3), 59–79. https://doi.org/10.1609/aimag.v31i3.2303
  • Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27, 2672–2680.
  • Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.
  • Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., … Fiedel, N. (2023). PaLM: Scaling language modeling with Pathways. Journal of Machine Learning Research, 24(240), 1–113.
  • Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., … Hassabis, D. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583–589. https://doi.org/10.1038/s41586-021-03819-2
  • Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv. https://doi.org/10.48550/arXiv.2001.08361
  • Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25, 1097–1105.
  • LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539
  • Lighthill, J. (1973). Artificial intelligence: A general survey. In Artificial intelligence: A paper symposium. Science Research Council.
  • McCarthy, J. (1960). Recursive functions of symbolic expressions and their computation by machine, Part I. Communications of the ACM, 3(4), 184–195. https://doi.org/10.1145/367177.367199
  • McCarthy, J., Minsky, M. L., Rochester, N., & Shannon, C. E. (2006). A proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955. AI Magazine, 27(4), 12–14. https://doi.org/10.1609/aimag.v27i4.1904 (Original document from 1955)
  • McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5(4), 115–133. https://doi.org/10.1007/BF02478259
  • Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv. https://doi.org/10.48550/arXiv.1301.3781
  • Minsky, M., & Papert, S. (1969). Perceptrons: An introduction to computational geometry. MIT Press.
  • Newell, A., & Simon, H. A. (1956). The logic theory machine: A complex information processing system. IRE Transactions on Information Theory, 2(3), 61–79. https://doi.org/10.1109/TIT.1956.1056797
  • Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). (2024). Official Journal of the European Union, L 2024/1689. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
  • Nilsson, N. J. (2010). The quest for artificial intelligence: A history of ideas and achievements. Cambridge University Press.
  • OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774
  • Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
  • Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI. https://openai.com/research/language-unsupervised
  • Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI. https://openai.com/research/better-language-models
  • Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., & Sutskever, I. (2021). Zero-shot text-to-image generation. In Proceedings of the 38th International Conference on Machine Learning (pp. 8821–8831). PMLR.
  • Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684–10695). IEEE. https://doi.org/10.1109/CVPR52688.2022.01042
  • Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386–408. https://doi.org/10.1037/h0042519
  • Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. https://doi.org/10.1038/323533a0
  • Russell, S., & Norvig, P. (2021). Artificial intelligence: A modern approach (4th ed.). Pearson.
  • Samuel, A. L. (1959). Some studies in machine learning using the game of checkers. IBM Journal of Research and Development, 3(3), 210–229. https://doi.org/10.1147/rd.33.0210
  • Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–424. https://doi.org/10.1017/S0140525X00005756
  • Shortliffe, E. H. (1976). Computer-based medical consultations: MYCIN. Elsevier.
  • Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489. https://doi.org/10.1038/nature16961
  • The Royal Swedish Academy of Sciences. (2024a, October 8). The Nobel Prize in Physics 2024 [press release]. https://www.nobelprize.org/prizes/physics/2024/press-release/
  • The Royal Swedish Academy of Sciences. (2024b, October 9). The Nobel Prize in Chemistry 2024 [press release]. https://www.nobelprize.org/prizes/chemistry/2024/press-release/
  • Turing, A. M. (1937). On computable numbers, with an application to the Entscheidungsproblem. Proceedings of the London Mathematical Society, s2-42(1), 230–265. https://doi.org/10.1112/plms/s2-42.1.230
  • Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433–460. https://doi.org/10.1093/mind/LIX.236.433
  • Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
  • Weizenbaum, J. (1966). ELIZA – A computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36–45. https://doi.org/10.1145/365153.365168
  • Wiener, N. (1948). Cybernetics: Or control and communication in the animal and the machine. MIT Press.
  • Winograd, T. (1972). Understanding natural language. Cognitive Psychology, 3(1), 1–191. https://doi.org/10.1016/0010-0285(72)90002-3

Is this article useful to you and are you citing it? Copy the citation