History
Games became the field’s yardstick because they could be scored, not because they were the right test
For half a century progress was measured against board games, and each victory turned out to demonstrate something narrower than the anticipation surrounding it had suggested.
By Daniel Okonkwo3 min read

Why games, of all things
Board games have properties that make them almost irresistible to researchers. The rules are complete and unambiguous, the state is fully visible, the outcome is unarguable, and there is an existing population of skilled humans to measure against. Almost nothing else offers all four.
They also carried cultural weight. Chess in particular was widely regarded as a distilled form of reasoning, so a machine that played it well would be demonstrating something important about thought. That assumption was rarely examined and it did an enormous amount of work in motivating the research.
Practical considerations reinforced the choice. Progress was measurable, results were comparable across groups, and success was legible to funders and to the public in a way that incremental work on representation or search was not. A programme that can point to a rating against human players has an argument for its next grant that a programme reporting improved knowledge representation does not.
Checkers came first and taught the more interesting lesson
Work on draughts in the late 1950s produced a program that improved through repeated play against itself, adjusting the weights it used to evaluate positions. It eventually played respectably, which was remarkable given the hardware available.
The historically significant part is the method. Self-play, an evaluation function tuned by experience, and improvement without a human specifying the improvements — these are recognisably the ingredients of approaches that would dominate much later, arrived at decades before the computation existed to exploit them.
That programme received far less attention than the chess work that followed, partly because the game carried less prestige. The habit of judging results by the status of the domain rather than by the method is one the field has never entirely shaken.
Chess fell to search and hardware more than to insight
The system that defeated a reigning world champion in 1997 relied on examining an enormous number of positions per second, using specialised hardware, guided by an evaluation function shaped with extensive input from human players. It was a formidable engineering achievement and a narrow one.
The reaction split immediately and the split has never healed. One view held that a hallmark of human intellect had been matched. Another held that the machine had demonstrated brute enumeration, that the approach generalised to nothing, and that the reasoning people admired in chess players was not what the machine was doing.
Both readings have merit. The system genuinely played chess better than any human, and it could not do anything else at all, including explain a move. What the episode actually established was that a domain long treated as a proxy for general intelligence was in fact susceptible to search at sufficient scale.
Go was chosen because search alone would not do
Go was the standing counter-example. Its branching factor makes exhaustive search hopeless, positions resist simple evaluation, and strong players describe their judgement in terms of shape and intuition rather than calculation. It was widely expected to remain out of reach for a long time.
The systems that succeeded in the middle of the 2010s combined learned evaluation with a sampling-based search and, in later versions, learned largely through self-play from the rules alone. That combination is more interesting than the chess result because the domain knowledge was mostly acquired rather than supplied.
It remains a single game with fixed rules, complete information and a definite outcome. Later work extended similar methods to games with hidden information and to real-time settings, which is a real broadening, and the environments are still environments with defined objectives and unlimited practice.
What the whole tradition selected for
Fifty years of measuring progress by games pushed the field towards problems with clear objectives, cheap repetition and unambiguous feedback, because those are the problems the methods could be developed against. That is not a criticism; it is how research programmes work, and it produced enormous progress.
The cost was a systematic understatement of everything games leave out. Ill-defined goals, situations where the rules are contested, consequences that arrive years later, and the impossibility of practising a thousand times — these describe most consequential human activity and none of it is present in a board game.
The pattern also explains a recurring disappointment. Each victory was received as evidence of general capability and each turned out to be a demonstration that the domain had been misjudged, which should probably temper how the next benchmark result is read.
Common questions
Did solving chess advance the wider field?
It advanced search techniques, hardware design and evaluation methods, and it demonstrated what specialised engineering could achieve. Its transfer to other problems was limited, which is precisely why the later game results, based on learning rather than hand-built evaluation, are generally regarded as the more significant ones.
Are games still used as research environments?
Yes, extensively, because they remain cheap, fast, reproducible and safe places to test methods that would be expensive or dangerous to try elsewhere. The change is in the claims made about the results: performance in an environment is now more often described as performance in that environment.
Why is a game so much easier than an everyday task?
Because a game supplies what real tasks withhold: complete rules, a defined objective, immediate scoring and unlimited repetition without consequence. Most of the difficulty in ordinary tasks lies in the absence of those things, which is exactly the part a game cannot exercise.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





