自2022年中以来,FRI(预测研究所)已在多项研究和项目中收集了关于AI进展的预测。其样本包括资深专家。第一轮LEAP(纵向专家AI小组)吸引了339名专家,其中包括76名计算机科学家、76名行业专家、68名经济学家和119名AI政策专家。其中的计算机科学家包括20所顶尖机构的30名教授以及被引用次数最多的200位AI作者中的10人。这些小组还包括“超级预测者”,即那些拥有准确预测记录的通才。
AI实现重大里程碑的时间远早于预测
最大的差距出现在数学领域。AI在2025年7月达到了国际数学奥林匹克竞赛的金牌水平,比中位数专家预测提前了五年,比中位数超级预测者预测提前了十年。FRI表示,这些预测是在ChatGPT发布之前的2022年收集的,但在此之后这一模式依然成立。
根据他们的预测,专家对实际发生的基准结果赋予的平均概率为24.6%,而超级预测者仅赋予9.7%。对于在国际数学奥林匹克竞赛中获得金牌的表现,这两个数字分别降至8.6%和2.3%。| 图片来源:Forecasting Research Institute
AI可能已经解决了一个千禧年大奖难题,但目前尚不清楚该解决方案是否符合评估标准。在2025年8月和9月的一项调查中,专家对到2027年底出现此类解决方案的中位数概率估计仅为10%,而超级预测者的估计为5.4%。
在一项关于AI在病毒学领域能力的研究中,专家预测AI模型要到2030年才能在故障排除基准测试中与顶级病毒学家团队相媲美。超级预测者则认为要等到2034年。FRI表示,这一情况很可能早在2025年4月就已发生。一项网络安全基准测试也显示出类似的低估现象。
中位数显示,专家预计AI将在2030年前在病毒学能力测试中与顶级团队持平,而超级预测者认为要等到2034年。FRI表示,这一里程碑很可能发生在2025年4月。受访者将这一里程碑与更高的预期生物风险联系起来,而非任何实际危害的有据可查的增加。| 图片来源:Forecasting Research Institute
经济预测也过于保守。专家对任何AI公司在2026年底的最高年度经常性收入(ARR)的中位数预测为200亿美元。经济学家预测为160亿美元,超级预测者预测为250亿美元。FRI引用Anthropic在2026年9月约1000亿美元的营收数据,认为这一数字可能已经达成。
Anthropic和OpenAI的年化收入(左图)远远超过了所有受访群体的中位数预测(右图)。即使是最高组的估计值250亿美元,也远低于2026年7月报告的650亿美元。| 图片来源:Forecasting Research Institute
现实世界的影响更难判断
并非所有预测都偏低。生物安全专家预测,使用语言模型的参与者中有22.5%将完成生物实验室任务。病毒学家预计为40%,超级预测者预计为16.2%。在一项受控试验中,仅5.2%的人在使用语言模型和互联网访问的情况下成功完成任务,而仅使用互联网的成功率为6.6%。尽管试验规模较小,但语言模型并未产生可衡量的差异。
专家在自动驾驶汽车方面的预测也可能偏高。他们对2027年美国网约车中自动驾驶行程占比的中位数预测为7.3%,而LLM(大型语言模型)的预测则为2.5%。FRI表示,目前还无法可靠地评估关于经济增长、就业和重大AI危害的预测。
与此同时,受访者正在上调他们的预期。在完成了两项调查的人群中,专家对AI成为“世纪技术”赋予的平均概率从31%上升至36%,超级预测者则从28%上升至35%,这一变化发生在九个月的时间里。
专家和超级预测者现在认为,人工智能将带来比九个月前预期更大的社会影响。这两个群体都将最高平均概率赋予“世纪技术”,其地位与电力相当。| 图片来源:Forecasting Research Institute(预测研究所)
FRI 正在增加更快的方法以保持与人工智能发展的同步
展望未来,FRI 将突出显示那些预计在 2040 年之前人工智能进展极其迅速的受访者子样本,并持续发布不断更新的大语言模型(LLM)预测,与人类预测并列。根据 ForecastBench 的数据,某些模型在特定类型的问题上已经能够与超级预测者相媲美。FRI 还希望找到最准确的 LEAP 小组成员,并在积累足够数据后展示他们的预测结果。
FRI 确实指出了其自身数据中的一个陷阱:当现实超越预测时,低估的情况会立刻显现;而高估则只有在截止日期过后才会变得清晰。这使得中期报告自然倾向于发现预测者过于谨慎的案例。FRI 的一些自身评估也依赖于使用原始预测者所不具备的信息的 LLM 投影。
Since mid-2022, FRI has collected forecasts on AI progress across several studies and projects . Its samples include senior specialists. The first round of LEAP (Longitudinal Expert AI Panel) drew 339 experts, including 76 computer scientists, 76 industry experts, 68 economists, and 119 AI policy specialists. The computer scientists included 30 professors at top-20 institutions and 10 of the 200 most-cited AI authors. The panels also included superforecasters, generalists with a proven record of accurate predictions.
AI hit major milestones years ahead of forecasts
The widest gap involves math . AI reached gold-medal level at the International Mathematical Olympiad in July 2025 , five years before the median expert forecast and ten years before the median superforecaster forecast. Those predictions were gathered in 2022, before ChatGPT launched, but the pattern held afterward too, according to FRI.
Based on their forecasts, experts assigned an average probability of 24.6 percent to the benchmark results that actually happened, while superforecasters assigned just 9.7 percent. For gold-medal performance at the Math Olympiad, the figures dropped to 8.6 and 2.3 percent. | Image: Forecasting Research Institute
AI may also have solved a Millennium Prize Problem , though it's still unclear whether the solution meets the evaluation criteria. In a survey from August and September 2025, experts had put the median odds of such a solution by the end of 2027 at just 10 percent, and superforecasters at 5.4 percent.
In a study of AI capabilities in virology, experts predicted AI models wouldn't match a top team of virologists on a troubleshooting benchmark until 2030. Superforecasters said 2034. FRI says that likely happened as early as April 2025 . A cybersecurity benchmark showed similar underestimates.
At the median, experts expected AI to match a top team on the Virology Capabilities Test by 2030, and superforecasters by 2034. FRI says it likely happened in April 2025. Respondents tied this milestone to higher expected biorisk, not to any documented rise in actual harm. | Image: Forecasting Research Institute
Economic forecasts were also far too conservative. Experts put the median for the highest annual recurring revenue (ARR) of any AI company at the end of 2026 at $20 billion. Economists said $16 billion, and superforecasters said $25 billion. FRI cites roughly $100 billion for Anthropic in September 2026 as a figure that has likely already been reached.
Annualized revenue at Anthropic and OpenAI (left) far exceeded the median forecasts of every surveyed group (right). Even the highest group estimate of $25 billion fell well short of the $65 billion reported in July 2026. | Image: Forecasting Research Institute
Real-world impact is harder to call
Not every forecast ran too low. Biosecurity experts predicted that 22.5 percent of participants using a language model would complete biological lab tasks. Virologists expected 40 percent, superforecasters 16.2 percent. In a controlled trial, only 5.2 percent succeeded with a language model and internet access, compared with 6.6 percent using the internet alone. The language model made no measurable difference, though the trial was small.
Experts may also have overshot on self-driving cars. Their median forecast for the share of autonomous US ride-hailing trips in 2027 was 7.3 percent, while an LLM projection puts it at 2.5 percent. FRI says forecasts on economic growth, employment, and major AI harms can't be reliably judged yet.
At the same time, respondents are revising their expectations upward. Among those who completed both surveys, the average probability assigned to AI becoming a "technology of the century" rose from 31 to 36 percent for experts and from 28 to 35 percent for superforecasters over nine months.
Experts and superforecasters now expect bigger societal effects from AI than they did nine months earlier. Both groups assign their highest average probability to "technology of the century," on par with electricity. | Image: Forecasting Research Institute
FRI is adding faster methods to keep pace with AI
Going forward, FRI will highlight a subsample of respondents who expect very rapid AI progress through 2040 and publish continuously updated LLM forecasts alongside the human ones. According to ForecastBench , some models already match superforecasters on certain question types. FRI also wants to find the most accurate LEAP panelists and feature their forecasts once enough data is in.
RI does flag a catch in its own data: underestimates become obvious as soon as reality overtakes a prediction, but overestimates only become clear once a deadline passes. That makes the interim report naturally tilted toward finding cases where forecasters were too cautious. Some of FRI's own assessments also rely on LLM projections that use information the original forecasters didn't have.
首次收录 · 2026-09-25 · 10.82 分