When the first electronic computers were finished in the 1940s, they were used to calculate: artillery tables, scientific computations, data processing. A question that some researchers began to ask was whether machines could also perform tasks that in people we associate with intelligence, such as proving theorems, playing chess, translating text or recognizing images. The difficulty was that these tasks were not defined by known formulas; nobody could describe, step by step, how to recognize a face. The history of computational intelligence, also called artificial intelligence (AI), is the history of different answers to that difficulty, among them rules written by experts, search through spaces of possibilities, and learning from examples.
The term covers a heterogeneous set of fields, and its definition is contested. Here it means the methods that let computer systems perform tasks of perception, reasoning, language or decision based on data, rules or both. This does not imply that such machines understand, think or are conscious, which remain open questions in philosophy and cognitive science.
How it works
There are two broad families of approach, which coexist and blend. In the symbolic approach, knowledge is represented by explicit rules and symbols: if the temperature is high and there is a cough, consider a certain diagnosis. The program manipulates those rules through logic or search. In the learning-based approach, the program is not given the rules; it adjusts internal parameters so that, given many examples of inputs and expected outputs, it produces similar answers for new inputs.
Artificial neural networks are the best-known learning method. A network is made of simple units arranged in layers; each unit adds up its inputs, multiplied by weights, and applies a function to produce an output. Training consists of changing the weights to reduce the error measured on examples, done today with gradient-based optimization algorithms, with backpropagation, a procedure that calculates how each weight contributes to the error, at the center. The result is a statistical model: its performance depends on the data, the architecture and the conditions of use, and it can fail in ways that are hard to predict.
Antecedents and first steps
In 1943 the neurophysiologist Warren McCulloch and the mathematician Walter Pitts published a model of the neuron as a simple logical unit, which would inspire neural networks. In 1950 Alan Turing published Computing Machinery and Intelligence in the journal Mind, proposing to replace the question of whether machines think with an imitation game conducted through written conversation, and discussing objections to the possibility of intelligent machines. The test, often called the "Turing test," is still discussed as a criterion, including for its shortcomings. This work grew out of the context of the first computers.
The field received its name and identity through a proposal for a summer workshop, written in 1955 by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon and held at Dartmouth College in 1956. The proposal used the phrase "artificial intelligence" and rested on the conjecture that aspects of intelligence could be described precisely enough to be simulated. The workshop drew fewer participants than is often imagined and did not yield a single result, but it is treated as an institutional starting point. In the same period Allen Newell, Herbert Simon and Cliff Shaw developed the Logic Theorist, a program that proved theorems from a well-known logic textbook and is usually presented as one of the first AI programs.
In 1958 Frank Rosenblatt, working at the Cornell Aeronautical Laboratory with funding from the U.S. Office of Naval Research, presented the perceptron, an artificial neuron model with a learning rule, implemented first on a computer and later in dedicated hardware. Press coverage at the time included inflated expectations. In 1969 Marvin Minsky and Seymour Papert published the book Perceptrons, analyzing mathematical limits of single-layer networks; historians debate how far that critique reduced funding for neural network research, and how much of it was later misread or amplified.
In 1966 Joseph Weizenbaum at MIT published ELIZA, a conversational program that answered the user's sentences with simple text-substitution patterns, imitating a therapist. Although the method was elementary, many users attributed understanding to the program, which led its author to reflect critically on the phenomenon. In the same years the LISP language, created by McCarthy in the late 1950s, was widely used in research, a topic covered in programming languages.
Historical context
In the 1970s and 1980s, expert systems tried to encode professionals' knowledge in rule bases. DENDRAL, begun in the mid-1960s at Stanford by Edward Feigenbaum, Joshua Lederberg, Bruce Buchanan and colleagues, helped identify chemical structures, and MYCIN, developed at Stanford in the early 1970s, suggested diagnoses and treatments for bacterial infections from rules. MYCIN was not adopted in clinical practice but influenced the design of later systems. Companies invested in products of this kind, and in 1982 Japan launched its Fifth Generation Computer Systems project, a ten-year program of great ambition led by the government's industry ministry, which fell short of its stated goals.
The periods of difficulty and funding cuts, around the mid-1970s and the late 1980s, are known as the AI winters. A report commissioned in the United Kingdom by the Science Research Council and published in 1973, known as the Lighthill report, criticized the results of the field. The expression is debated, though: some historians note that research continued under other names, such as pattern recognition, robotics and machine learning, and that the intensity of the cycles varied by country and by type of funding.
In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams published a paper in Nature that popularized backpropagation for training multilayer networks; the method already had antecedents in the work of other authors. In the 1990s statistical machine-learning methods consolidated, with larger data sets and techniques such as support vector machines. In May 1997 IBM's Deep Blue defeated world chess champion Garry Kasparov in a six-game match by 3.5 to 2.5, having lost a first match to him in 1996. The system relied on massive search and evaluation functions designed by experts, not on learning as it is understood now, and the event was accompanied by controversy over the match conditions.
From the 2000s, three factors combined: very large volumes of data, often drawn from the Web; graphics processors repurposed for parallel computation; and improvements in deep neural network techniques, meaning networks with many layers. In 2009 researchers led by Fei-Fei Li released ImageNet, a large set of labeled images. In 2012 the AlexNet network, by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton, scored far better than its competitors in the ImageNet-associated competition, spurring the use of deep learning in computer vision. In March 2016 DeepMind's AlphaGo defeated the professional player Lee Sedol four games to one in a match of Go, a game whose search space is particularly difficult. In 2017 Google researchers published the paper Attention Is All You Need, introducing the Transformer architecture, based on attention mechanisms and now used in many language models.
Large language models, trained on enormous quantities of text to predict sequences of words, began to be offered to the public in conversational applications in late 2022. As of this article's latest review, their capabilities, their errors and their best uses are being studied and contested by researchers, and this article makes no statements about future developments.
Impact and limitations
Computational intelligence methods are used in translation, speech and image recognition, content recommendation, support for medical image analysis and other areas, with results that depend on context and validation. Documented limitations and risks include the following.
- Data and bias: systems trained on data that reflect inequalities can reproduce them, as studies of facial recognition and of résumé screening have shown.
- Interpretability: in deep networks it is hard to explain why a given output was produced, a concern in fields such as health and justice.
- Plausible errors: language models can produce false statements that sound confident.
- Energy and resources: training and running large models requires data centers with significant electricity use; estimates vary with methodology.
- Labor and rights: there are debates about employment, copyright in training data and privacy.
On the regulatory side, different jurisdictions have debated or adopted specific rules; the European Union, for example, adopted a regulation on the subject in 2024, entering into force that year with phased application. Some caution about hype is warranted: throughout the field's history, predictions of rapid advances often failed to come true on the announced schedules, and claims about current capabilities should rest on evaluations and peer-reviewed publications.
Connections to other technologies
Computational intelligence depends on the hardware described in first computers and its successors, on the abundance of data from the Web, and on advances in programming languages. In robotics it seeks to give machines perception and planning, with large challenges remaining outside controlled settings.
