Version M10-AI-EN-001 · derived from M2-AI-REV-004
Artificial intelligence: questions, learning, and use
A selective six-section route through technical events up to 2025, with the measurement boundaries of a 2026 report. This is not a complete history of AI and does not claim worldwide coverage through September 2026.
Can machines think?
In 1950, Alan Turing began with the question “Can machines think?” and redirected it toward the imitation game, in which an interrogator tries to identify unseen participants. Machine intelligence was thereby placed inside a specific task of judgment, with participants, a mode of communication, and a contestable standard. What any one performance demonstrates still has to be discussed in relation to that arrangement.
On 31 August 1955, the proposal for the Dartmouth summer research project listed learning, language, abstraction, and self-improvement among the research problems of machine intelligence, and imagined organizing the project the following summer. What survives is a research agenda: which problems were grouped together and who was expected to work on them. The date of the proposal, the planned project, and later results are not the same event.
Turing’s article made the question “Can machines think?” more specific. The Dartmouth proposal separated language, learning, and other problems and arranged them as research tasks. Together, the two documents show how a question became an agenda while preserving the choices and disputes within each document. They cannot, by themselves, account for the origins of every method developed over the following decades.
Knowledge, evaluation, and learning
Feigenbaum’s Stanford author-archive bibliography lists the 1968 work “Heuristic DENDRAL: A Program for Generating Explanatory Hypotheses in Organic Chemistry,” attributed to Feigenbaum, Sutherland, and Buchanan. A participant’s retrospective in the archive presents the knowledge base separately from the processes used to generate and evaluate candidates. This offers one route into how knowledge systems organized domain knowledge inside programs; the implementation described at the time still needs comparison with the contemporary paper.
Who evaluates a research direction is also part of this history. The surviving Lighthill report is dated July 1972. Its introduction says that the assessment was commissioned to inform the UK Science Research Council’s consideration of funding applications for AI research. The author described the view as a personal one formed after limited study and consultation. This is a particular funding assessment, not a sufficient account of a single worldwide rise and fall in research.
In 1986, Rumelhart, Hinton, and Williams described a learning procedure that adjusted the weights of network connections to reduce the difference between actual and target outputs. They argued that hidden units could develop representations relevant to the task during this process. Training therefore changed not only the final answer but also how inputs were represented inside the network. The paper concerns layered networks and a specific method; it does not establish that machines had come to understand the world.
Making results comparable
A chess match in 1997 produced a clear result. IBM’s later historical account says that Deep Blue defeated Garry Kasparov in a six-game series. This was a rule-bound chess contest. It left a result and a boundary: evaluating the machine’s performance in other activities still requires other evidence. The source used here is a later corporate retrospective; an independent match record remains to be added.
The 2009 ImageNet paper first brings the organization of materials into view. Its authors arranged images through the category structure of WordNet and described a collection process involving Amazon Mechanical Turk. Before any algorithmic score, the choice of categories, images, and labels already shaped the terms of comparison. The paper’s first page opens these questions but cannot replace an examination of collection and labeling in practice; the scale reported then also does not describe every later version of the dataset.
The abstract of the 2012 AlexNet paper reports experiments with a deep convolutional neural network for ImageNet classification and describes the use of GPUs in training. The network, training material, and computing equipment jointly formed the conditions of the experiment. How far the classification results extend to other images and tasks requires separate tests. The paper reports a method and an evaluation, not an answer to every problem in vision.
The AlphaGo paper published in January 2016 combined policy networks, value networks, and Monte Carlo tree search, with training that used expert games and self-play. Its authors reported that the program defeated the European Go champion by five games to zero. This is the contest recorded in that paper, not the later match against Lee Sedol, and it cannot be understood apart from its date and opponent. The way the Go system combined methods and training material still has to be described component by component.
A shared foundation for language tasks
The Transformer paper first submitted in June 2017 proposed an architecture based entirely on attention mechanisms, removing the recurrent and convolutional structures described in the paper, and reported experiments on two machine-translation tasks. How a method for language processing was organized thus became a question that could be stated and compared more precisely. The abstract of the first submission covers those tasks and cannot be extended on its own to other uses.
The BERT paper first submitted in October 2018 described pre-training as forming deep bidirectional representations from both left and right context, then adding an output layer and fine-tuning for particular tasks. It connected the creation of a reusable representation with adaptation to a specific task. How well each task was performed still has to be read within the paper’s experimental setting; no conclusion follows directly from the word “understanding.”
The 2019 FLORES paper makes differences in language materials visible. Its authors prepared evaluation data translated from Wikipedia sentences for Nepali–English and Sinhala–English, supplied baselines under different supervision settings, and reported that then-current leading methods still performed poorly on this low-resource evaluation. Two specified language pairs and one set of results from that time are not enough to make judgments about other languages.
From continuation to human feedback
The GPT-3 paper first submitted in May 2020 gave the model task instructions and a small number of examples, and described evaluations that did not apply a new gradient update or fine-tuning for each task. Requirements and examples placed in the prompt became one way to adapt the model to a task. The authors also identified difficult tasks and methodological problems associated with training on web material. The reliability of any particular answer still has to be tested separately.
The InstructGPT paper first submitted in March 2022 brought demonstrations and rankings of outputs into the training process: it first fine-tuned on human demonstrations, then used output rankings in reinforcement learning. Human judgment thereby helped change how the model answered. The authors still reported simple mistakes. Being preferred cannot be equated with every answer being correct. The tasks from which feedback comes and the way it is formed also determine how broad a claim it can support.
The DeepSeek-R1 paper first submitted in January 2025 distinguishes two training paths. According to its abstract, R1-Zero used reinforcement learning without preliminary supervised fine-tuning and developed problems including poor readability and language mixing; R1 introduced cold-start data and multi-stage training. The paper therefore places training steps and observed behavior within the same account. This reading is fixed to the first version: it does not write later revisions back into the initial submission or treat the authors’ comparisons as a context-free capability ranking.
Who uses AI, and what changes
In May 2025, the International Labour Organization published an occupational-exposure index that estimated, task by task, which jobs generative AI might affect. The report estimated that about one quarter of workers worldwide were in occupations with some degree of exposure, while noting that most occupations still required human input and that transformation was more likely than direct replacement. It measures potential impact: an occupation containing tasks that might change does not mean its workers have already lost their jobs.
Actual use leaves a different kind of record. According to the abstract published by Johns Hopkins University, the 2025 study “How People Use ChatGPT” examined use of the consumer product from November 2022 through July 2025 and used an automated process to classify sampled conversations; practical guidance, seeking information, and writing were listed among the main uses. Some authors were OpenAI researchers, and only the abstract was read here, so sampling and classification methods remain unchecked. The result applies to that product sample, not to all uses of AI.
To observe how people use AI, we must first ask how the numbers were produced. The economy chapter of the 2026 AI Index places comparisons of technology adoption beside cross-national telemetry of use. One figure distinguishes samples of people aged 18–64 from samples of all ages; the other uses separate Microsoft telemetry. The two figures offer different views and cannot be joined into one percentage of the world’s population. The report year must also be separated from the years represented by the underlying data.
Looking back along this route, the question moves from how machine intelligence was framed to how systems were trained and compared, and then to how people participated and used them. FLORES preserves specific differences in language materials; the labor study distinguishes potential exposure from actual outcomes. Conditions easily hidden by a single aggregate score become visible again. Future records still need to follow how those conditions change and revise earlier judgments when new evidence arrives.