Top AI experts badly underestimated how fast the field is moving, study finds

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 24, 2026 3 min read
Top AI experts badly underestimated how fast the field is moving, study finds

AI reached gold-medal level at the International Mathematical Olympiad in July 2025, five years before the median expert forecast and ten years before the median superforecaster forecast.

That gap is the widest between reality and expectation. An interim report from the Forecasting Research Institute (FRI) shows specialists at top universities and heavily cited researchers consistently missed recent progress on benchmarks and adoption metrics. The data comes from surveys gathered since mid-2022 across several studies.

The first round of LEAP (Longitudinal Expert AI Panel) drew 339 experts. The group included 76 computer scientists, 76 industry experts, 68 economists, and 119 AI policy specialists. Among the computer scientists were 30 professors at top-20 institutions and 10 of the 200 most-cited AI authors. The panels also included superforecasters, generalists with a proven record of accurate predictions.

AI hit major milestones years ahead of forecasts

Mathematics is the area with the largest discrepancy. Experts gathered in 2022, before ChatGPT launched, predicted AI would not match that level until years later. The pattern held afterward, according to FRI.

AI may also have solved a Millennium Prize Problem, though it is unclear whether the solution meets the evaluation criteria. In a survey from August and September 2025, experts put the median odds of such a solution by the end of 2027 at just 10 percent. Superforecasters said 5.4 percent.

In a study of AI capabilities in virology, experts predicted AI models would not match a top team of virologists on a troubleshooting benchmark until 2030. Superforecasters said 2034. FRI says that likely happened as early as April 2025. A cybersecurity benchmark showed similar underestimates.

Economic forecasts were also far too conservative. Experts put the median for the highest annual recurring revenue (ARR) of any AI company at the end of 2026 at $20 billion. Economists said $16 billion, and superforecasters said $25 billion. FRI cites roughly $100 billion for Anthropic in September 2026 as a figure that has likely already been reached.

Real-world impact is harder to call

Not every forecast ran too low. Biosecurity experts predicted that 22.5 percent of participants using a language model would complete biological lab tasks. Virologists expected 40 percent, superforecasters 16.2 percent. In a controlled trial, only 5.2 percent succeeded with a language model and internet access, compared with 6.6 percent using the internet alone. The language model made no measurable difference, though the trial was small.

Experts may also have overshot on self-driving cars. Their median forecast for the share of autonomous US ride-hailing trips in 2027 was 7.3 percent, while an LLM projection puts it at 2.5 percent. FRI says forecasts on economic growth, employment, and major AI harms cannot be reliably judged yet.

At the same time, respondents are revising their expectations upward. Among those who completed both surveys, the average probability assigned to AI becoming a technology of the century rose from 31 to 36 percent for experts and from 28 to 35 percent for superforecasters over nine months.

FRI is adding faster methods to keep pace with AI

Going forward, FRI will highlight a subsample of respondents who expect very rapid AI progress through 2040 and publish continuously updated LLM forecasts alongside the human ones. According to ForecastBench, some models already match superforecasters on certain question types. FRI also wants to find the most accurate LEAP panelists and feature their forecasts once enough data is in.

RI does flag a catch in its own data: underestimates become obvious as soon as reality overtakes a prediction, but overestimates only become clear once a deadline passes. That makes the interim report naturally tilted toward finding cases where forecasters were too cautious. Some of FRI’s own assessments also rely on LLM projections that use information the original forecasters did not have.

What it means

For people making things, the gap between what models can do and what people expect them to do is wider than before. In some areas like biology, adding an AI tool to a workflow did not improve results over using the internet alone. In others, the technology is already solving problems experts thought were years away.

Scroll to Top