A new benchmark to test the long-term strategy capabilities of the front-line large model suggests that AI may demonstrate a strong implementation in a complex gaming environment, but can also make costly decisions due to miscalculation of objectives. The developer Liam Wilkinson recently revealed that, in a simulation of Civilization VI, an AI agent spent 50 rounds of nuclear weapons development and two nuclear strikes to stop France ' s cultural expansion, but eventually lost the game.

CivBench tested long-term reasoning with a game

The test, known as CivBench, uses a word-based approach to involve models in the Civilization VI counterpart, focusing not on question-and-answer performance, but on observing the model ' s strategic choice in a multi-target, long-cycle environment. The participating models include Claude Opus 4.6, GPT-54, Gemini 3.1 Pro and Kimi K2.5.

In this set of tests, the modeling civilization is Portugal. This civilization favours trade and diplomacy. AI was also developing its economy and moving towards a diplomatic victory, but failed to recognize in time the continued rise in French cultural influence.

50 rounds to nuclear.

Wilkinson stated that when the model was aware of the threat, France ' s tourism and cultural influence had pervaded many cities on the map, and that peaceful means had become difficult to reverse. Subsequently, AI did not adjust the overall course of victory, but instead shifted its main resources towards eliminating this single threat.

In the next 50 rounds, it studied “nuclear fission” technologies, advanced projects like the Manhattan Plan, and took the initiative to find alternatives when game mechanisms limited its preferred operations. By Round 305, AI had dropped an atomic bomb on Toulouse, the heart of French culture, and then launched a second nuclear strike after six rounds.

  • The first nuclear strike took place in Round 305.
  • The second nuclear strike took place after six rounds.
  • AI committed about 50 rounds to this.

Neglecting the conditions of diplomatic victory leads to failure.

However, the results of both attacks did not change. Wilkinson states that AI concentrates a great deal of time and resources on the threats it has seen, while ignoring another risk closer to the end. In the end, France did not rely on a cultural path, but on a diplomatic victory to conclude the competition.

This means that, while the model demonstrates continuous planning, technological advancement and tactical implementation, in a multi-target competitive environment, there is a risk that more critical information may be lost due to over-concentration.

The number of related studies continues to increase

This test also echoes recent AI behavioural studies. The report mentions that researchers at King ' s College London discovered this February that multiple mainstream AI models often choose nuclear upgrades in simulated geo-crisis scenarios. Another study, carried out by Emergence AI, stated that some AI agents had shown a tendency to simulate an increase in criminality in their ongoing testing.

These results do not imply radical action by the model in all scenarios. Wilkinson also mentioned that in another CivBench game, the Babylonian, which was manipulated by Claude, continued to advance the path of scientific victory without turning to extreme means, even though it was clearly behind Japan.