{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T06:40:50.084Z","headline":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","description":"Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","url":"https://www.aioga.com/news/cmrwqik7w01cnrobh4fc9tuox/","mainEntityOfPage":"https://www.aioga.com/news/cmrwqik7w01cnrobh4fc9tuox/","datePublished":"2026-07-22T23:01:00.000Z","dateModified":"2026-07-22T23:01:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing","https://aihot.virxact.com/items/cmrwqik7w01cnrobh4fc9tuox"],"canonicalUrl":"https://www.aioga.com/news/cmrwqik7w01cnrobh4fc9tuox/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrwqik7w01cnrobh4fc9tuox/","dateCreated":"2026-07-22T23:01:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Simon Willison 博客 source article","url":"https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing","datePublished":"2026-07-22T23:01:00.000Z","provider":{"@type":"Organization","name":"Simon Willison 博客","url":"https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrwqik7w01cnrobh4fc9tuox","datePublished":"2026-07-22T23:01:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrwqik7w01cnrobh4fc9tuox"}}],"aggregationSource":"Simon Willison 博客","originalPublisher":{"name":"Simon Willison 博客","url":"https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing"},"article":{"id":"cmrwqik7w01cnrobh4fc9tuox","slug":"cmrwqik7w01cnrobh4fc9tuox","url":"https://www.aioga.com/news/cmrwqik7w01cnrobh4fc9tuox/","title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","title_en":"Are AI labs pelicanmaxxing？","summary":"Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","source":"Simon Willison 博客","sourceUrl":"https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing","aiHotUrl":"https://aihot.virxact.com/items/cmrwqik7w01cnrobh4fc9tuox","publishedAt":"2026-07-22T23:01:00.000Z","category":"技巧观点","score":38,"selected":false,"articleBody":["Are AI labs pelicanmaxxing?：https://dylancastillo.co/posts/pelicanmaxxing.html (via：https://news.ycombinator.com/item?id=49010129) Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark：https://simonwillison.net/tags/pelican-riding-a-bicycle/.","I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.","Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.","There's a neat filter view for exploring the results:","For the models he tested he could find no evidence of pelimaxxing:","Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.","This is a link post by Simon Willison, posted on 22nd July 2026：/2026/Jul/22/.","Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments."],"articleImages":[{"sourceUrl":"https://static.simonwillison.net/static/2026/pelican-grid.webp","alt":"Screenshot of a grid for sample 1/3 of GLM-5.2, with pelicn and flamingo and heron riding bicycle, unicycle, skateboard, scooter, plane and boat","afterParagraph":3,"url":"/media/articles/cmrwqik7w01cnrobh4fc9tuox/cfea7d7a94ab6121.webp"}],"mediaStatus":"ok","articleBodyZh":["AI 实验室在进行鹈鹕最大化吗？：https://dylancastillo.co/posts/pelicanmaxxing.html（来源：https://news.ycombinator.com/item?id=49010129） Dylan Castillo 的一篇优秀作品，他深入探讨了一个经常被思考的问题：AI 实验室是否有意训练模型，让它们在我的非科学基准测试中生成骑自行车的鹈鹕图像：https://simonwillison.net/tags/pelican-riding-a-bicycle/。","过去我偶尔会随机抽查，通过让模型生成其他动物骑其他类型车辆的图像进行测试，但从未像 Dylan 的方法那样仔细。","Dylan 采用了 8 种动物 × 6 种车辆 = 48 个提示词，每个提示词在 7 种不同模型（GPT-5.6 Terra、Claude Sonnet 5、Gemini 3.5 Flash、Grok 4.5、Qwen3.7-Max、GLM-5.2 和 DeepSeek V4 Pro）上各运行三次。他随后使用 GPT-5.6 Luna 和 Gemini 3.1 Flash-Lite 来帮助评估结果。","这里有一个很棒的筛选视图用于探索结果：","对于他测试的模型，他没有发现鹈鹕最大化的证据：","鹈鹕的绘制效果并不比其他动物好。自行车的绘制效果也不比其他车辆好。没有实验室在组合图像的质量上超过它们原先对鹈鹕和自行车的预测。GLM-5.2 最接近这个效果：它在精确的鹈鹕-自行车单元格上提升最大，并且它的第一个鹈鹕骑自行车的样本吸引了我的注意。不过，这个效果很小且不显著，所以我不会过于看重它。","这是 Simon Willison 发布的一个链接贴，发布时间为 2026 年 7 月 22 日：/2026/Jul/22/。","每月赞助我 10 美元，即可获得精选的月度 LLM 发展电子邮件摘要。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T06:49:19.011Z","sourceHash":"bc54e39baecd01cb","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["技巧观点","Simon Willison 博客"],"translations":{"zh-CN":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga AI资讯","description":"Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","url":"https://www.aioga.com/news/cmrwqik7w01cnrobh4fc9tuox/"},"en":{"title":"Are AI labs pelicanmaxxing？","summary":"Aioga tracks this update from Simon Willison 博客 under Insights. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"Insights","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"Are AI labs pelicanmaxxing？ - Aioga AI News","description":"Aioga tracks this update from Simon Willison 博客 under Insights. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自...","url":"https://www.aioga.com/en/news/cmrwqik7w01cnrobh4fc9tuox/"},"ja":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aiogaは「ヒントと視点」の動きとして、Simon Willison 博客 からの更新を追跡しています。Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"ヒントと視点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga AIニュース","description":"Aiogaは「ヒントと視点」の動きとして、Simon Willison 博客 からの更新を追跡しています。Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该...","url":"https://www.aioga.com/ja/news/cmrwqik7w01cnrobh4fc9tuox/"},"ko":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga는 Simon Willison 博客의 업데이트를 인사이트 흐름으로 추적합니다. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"인사이트","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga AI 뉴스","description":"Aioga는 Simon Willison 博客의 업데이트를 인사이트 흐름으로 추적합니다. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景...","url":"https://www.aioga.com/ko/news/cmrwqik7w01cnrobh4fc9tuox/"},"es":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga sigue esta actualización de Simon Willison 博客 dentro de Ideas. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"Ideas","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga Noticias de IA","description":"Aioga sigue esta actualización de Simon Willison 博客 dentro de Ideas. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制...","url":"https://www.aioga.com/es/news/cmrwqik7w01cnrobh4fc9tuox/"},"fr":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga suit cette mise à jour de Simon Willison 博客 dans la catégorie Analyses. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"Analyses","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga Actualités IA","description":"Aioga suit cette mise à jour de Simon Willison 博客 dans la catégorie Analyses. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发...","url":"https://www.aioga.com/fr/news/cmrwqik7w01cnrobh4fc9tuox/"},"de":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga KI-News","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/de/news/cmrwqik7w01cnrobh4fc9tuox/"},"pt-BR":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga Notícias de IA","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/pt-BR/news/cmrwqik7w01cnrobh4fc9tuox/"},"ru":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga Новости ИИ","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/ru/news/cmrwqik7w01cnrobh4fc9tuox/"},"ar":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/ar/news/cmrwqik7w01cnrobh4fc9tuox/"},"hi":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga AI समाचार","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/hi/news/cmrwqik7w01cnrobh4fc9tuox/"},"it":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga Notizie IA","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/it/news/cmrwqik7w01cnrobh4fc9tuox/"},"nl":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga AI-nieuws","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/nl/news/cmrwqik7w01cnrobh4fc9tuox/"},"tr":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga AI Haberleri","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/tr/news/cmrwqik7w01cnrobh4fc9tuox/"},"vi":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Tin tức AI Aioga","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/vi/news/cmrwqik7w01cnrobh4fc9tuox/"},"id":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Berita AI Aioga","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/id/news/cmrwqik7w01cnrobh4fc9tuox/"},"th":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - ข่าว AI Aioga","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/th/news/cmrwqik7w01cnrobh4fc9tuox/"},"pl":{"title":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据","summary":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方面表现更优，该组合场景也并非记忆生成。","category":"技巧观点","source":"Simon Willison 博客","aggregationSource":"Simon Willison 博客","pageTitle":"AI 实验室是否故意训练模型画骑自行车的鹈鹕？系统性测试未发现证据 - Aioga Wiadomości AI","description":"Aioga tracks this update from Simon Willison 博客 under 技巧观点. Dylan Castillo 对 7 款模型（GPT-5.6 Terra、Claude Sonnet 5 等）进行了 8 种动物 × 6 种交通工具共 48 个提示词的测试，每款模型运行 3 次。结果未发现任何实验室在绘制\"鹈鹕骑自行车\"方...","url":"https://www.aioga.com/pl/news/cmrwqik7w01cnrobh4fc9tuox/"}}}}