an 现在使用 OpenAI 的 GPT-5.6 Sol Pro(https://the-decoder.com/openai-staffer-maps-out-which-of-gpt-5-6-sols-five-reasoning-levels-fits-which-task-complexity/)推翻了这一假设。在他的预印本中(https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf),他利用 AI 构建了一个统计模型,其中实际假发现率可证明超过了目标水平。模拟也证实了这一结果。Dobriban 还发布了相关的代码(https://github.com/dobriban/BH)。
这个结果对统计学家仍然具有重要意义,因为在人工未能解决问题后,人工智能快速地解决了它。多布里班(Dobriban)表示,GPT-5.6 Sol Pro 大约用了 90 分钟,而 GPT-5.5 即使在几个代理共同工作约 20 小时后也无法找到解答。“所以能力的提升是真实存在的。生活在这样激动人心的时代!”他写道。完整的聊天记录和提示可在此查看:https://chatgpt.com/share/6a541c6f-a2d0-83ea-bb2f-782271a103ca。
A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 to solve one of the central open questions in his field.
When researchers test thousands of hypotheses at once, like scanning the human genome for disease-linked genes, they run into a problem: The more tests you run, the more false positives slip through.
In 1995, statisticians Yoav Benjamini and Yosef Hochberg developed a method to limit these false positives. It controls the false discovery rate, or FDR, which is the share of reported significant results that are actually false alarms.
The Benjamini-Hochberg procedure:https://en.wikipedia.org/wiki/False_discovery_rate, or BH, is now widely used in modern statistics and across many scientific fields. According to Edgar Dobriban:https://x.com/EdgarDobriban, an associate professor at the University of Pennsylvania's Wharton School, the original paper has received more than 130,000 citations.
Benjamini and Hochberg originally showed that their method works with independent data. Real-world data points, however, are often linked. Genetic variants can be correlated, for example, when certain locations in the genome are frequently inherited together.
For years, experts assumed the BH procedure would also work reliably with correlated, normally distributed data, specifically when testing for deviations in both directions. But nobody had ever proved it.
an has now disproven that assumption using OpenAI's GPT-5.6 Sol Pro:https://the-decoder.com/openai-staffer-maps-out-which-of-gpt-5-6-sols-five-reasoning-levels-fits-which-task-complexity/. In his preprint:https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf, he uses the AI to construct a statistical model where the actual false discovery rate provably exceeds the target level. Simulations confirm the result. Dobriban also published the accompanying code:https://github.com/dobriban/BH.
Dobriban writes that the gap above the target level is "relatively small (0.104 vs 0.1)," so the result mainly matters for theory at this point. Practical effects still need further study, and the finding doesn't mean the BH procedure is generally unusable.
The result is still significant for statisticians because AI solved the problem quickly after humans had failed. Dobriban says GPT-5.6 Sol Pro took about 90 minutes. GPT-5.5 couldn't find a solution even after roughly 20 hours of work with several agents. "So the capability improvement is quite real. Exciting times to live in!" he writes. The full chat and prompt are available here:https://chatgpt.com/share/6a541c6f-a2d0-83ea-bb2f-782271a103ca.
Berkeley statistician Will Fithian:https://x.com/wfithian/status/2077218361398964684 called the disproved conjecture "the most interesting open problem in my area of statistics" and the result "another marker of advancing AI capabilities whose consequences will reach far beyond math."
Fithian also hinted at how much these results are shaking experts' sense of their own work. "I can't help but mourn the bygone days when a key result always meant a colleague to celebrate; a human insight to admire; a human achievement to be inspired by."
As with similar cases in mathematics:https://the-decoder.com/openais-gpt-5-6-sol-ultra-reportedly-solves-a-50-year-old-math-problem-in-under-an-hour/, the solution appears to combine existing approaches rather than produce something entirely new. Dobriban said the combination was unusual, but the result was ultimately "not especially surprising." The challenge was finding the right way to connect known methods, and the newer model managed to do that.
This leaves a broader question unanswered. Can models trained on human data reason their way to genuinely new knowledge, or can they "only" recombine what they learned during training:https://the-decoder.com/so-called-reasoning-models-are-more-efficient-but-not-more-capable-than-regular-llms-study-finds/? Even if recombination is all these systems can do, they already prove useful as everyday tools built into human workflows. Dobriban's result adds to a growing list of examples:https://the-decoder.com/terence-tao-argues-ai-could-bring-division-of-labor-to-math-for-the-first-time-in-history/.
But more ambitious goals, like building self-improving AI that can generalize:https://the-decoder.com/deepmind-ceo-hassabis-says-nobody-in-the-world-knows-what-happens-next-so-cautious-optimism-means-building-guardrails-now/, may demand something beyond recombination. Deep learning pioneer Richard Sutton is among those who think so, having recently founded a startup to tackle exactly that problem:https://the-decoder.com/turing-award-winner-rich-sutton-founds-oak-lab-to-build-ai-agents-that-learn-on-their-own/.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:宾夕法尼亚大学沃顿商学院教授 Edgar Dobriban 使用 OpenAI 的 GPT-5.6 Sol Pro 在约 90 分钟内推翻了一个存在 30 年的统计学猜想,证明 Aioga 将其归入「论文研究」方向,重点关注它对真实使用和行业竞争的影响。