法官认定 OpenAI 与 ANI 分属不同行业,未造成经济损失,并肯定大语言模型对教育、研究和可访问性的公共利益。 法院初步将 AI 训练视为"私人或个人使用,包括研究"的例外,但要求训练数据必须来自合法来源。
在一项临时裁决中,德里高等法院驳回了印度新闻机构亚洲新闻国际社(ANI)针对OpenAI提出的初步禁令请求。
ANI向法院提交了几份ChatGPT输出内容,声称它们是其文章的大量复制。然而,这一举动适得其反,因为OpenAI展示了所使用的模型GPT-4和GPT-4o的训练数据截止到2022年4月和2024年4月。而ANI引用的文章大多是2024年8月和9月发表的,因此不可能成为训练数据的一部分。Ad DEC_D_Incontent-1
法官的初步观点是,这些相似之处源于RAG技术,它允许语言模型实时检索在线信息,类似于搜索引擎。ANI在其提交文件中未涉及RAG,因此法院无法就该问题作出最终裁决。法官表示,基于RAG的输出可能符合“向公众传播”的定义,法院将在主要诉讼程序中对此问题作出裁决。
证据也不支持ANI关于OpenAI会永久存储训练数据,并可以按需逐字再现其作品的主张。不过,法院将在主要诉讼程序中重新审议这一问题。
要使这一例外成立,法官设定了条件。训练副本必须来自合法来源,而非无需许可访问的影子图书馆或付费网站。OpenAI也从未公开训练副本,仅在内部进行处理。Guadamuz表示,这是法院首次明确认定AI训练属于私人使用例外。
法院进行了三部分公正测试,并在三项中均支持OpenAI。OpenAI对ANI作品的使用仅限于训练,因为没有证据证明其进行了记忆或复制。ANI也无法证明经济损失,因为两家公司运营的领域不同。即便用户向ChatGPT询问ANI的头条新闻,该模型最多只返回主题和少量文章标题。
法官引用了包括 Bartz 诉 Anthropic 案:https://the-decoder.com/anthropic-won-a-fair-use-hearing-that-could-end-up-being-a-defeat/ 和 Kadrey 诉 Meta 案:https://the-decoder.com/metas-libgen-controversy-reveals-how-desperate-ai-companies-are-for-quality-training-data/ 在内的美国案例,其中语言模型输出被认为具有变革性。他还提到了早期的 Google 图书裁决:https://www.lto.de/recht/hintergruende/h/google-books-projekt-autoren-urheberrecht-fair-use。
法官还认定,经过训练的语言模型可以改善信息获取、支持教育、推动科学研究、帮助软件开发、实现翻译,并为残障人士创造工具。
在 Ross Intelligence 诉 Thomson Reuters 案:https://the-decoder.com/us-court-rejects-ai-startups-fair-use-defence-but-impact-on-openai-and-others-may-be-limited/ 中,法院拒绝了合理使用的抗辩,因为该 AI 研究工具直接与 Thomson Reuters 的法律数据库 Westlaw 竞争,使其使用不具变革性。法院强调,这一裁决仅适用于该非生成使用案例,不能推广到大型语言模型。
在所有这些案件中,相同的核心问题仍未解决。AI 模型是否会永久存储训练数据?训练能否被认定为合理使用?合法获取和非法获取的数据界限在哪里?通过对抗性提示生成的副本是否反映正常用途?
慕尼黑第一地方法院最近也裁定,谷歌需对其 AI 摘要中的虚假陈述直接负责:https://the-decoder.com/landmark-german-ruling-declares-googles-ai-overviews-are-googles-own-words-and-makes-it-liable-for-false-answers/,因为这些被视为独立内容而非搜索结果。传统上保护搜索引擎运营商的有限责任不适用于 AI 生成的摘要。
这一裁决也可能与 ChatGPT 基于 RAG 的回答相关。当人工智能系统总结新闻并提出独立主张时,其运营者实际上就成了媒体提供者,并承担由此带来的所有责任。这一变化还可能迫使法院重新考虑公平使用。公平性测试的一个关键因素是新产品是否与其训练所用的作品产生竞争。如果 AI 总结替代了访问新闻网站的需求,法官可能更难裁定这种使用不具竞争性。
保持对 AI 的关注。清晰、有用、无花哨。
关注 The Decoder 获取 AI 新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
In an interim ruling, the Delhi High Court rejected a request by Indian news agency Asian News International (ANI) for a preliminary injunction against OpenAI.
ANI submitted several ChatGPT outputs to the court that it claimed were substantial copies of its articles. The move backfired because OpenAI showed that the models used, GPT-4 and GPT-4o, were trained on data from April 2022 and April 2024. The articles ANI cited were mostly from August and September 2024, so they couldn't have been part of the training data. Ad DEC_D_Incontent-1
The judge's preliminary view was that the similarities came from RAG, which lets a language model retrieve online information in real time, much like a search engine. ANI hadn't addressed RAG in its filing, so the court couldn't make a final ruling on the issue. The judge said RAG-based outputs could qualify as "communication to the public," a question the court will address in the main proceedings.
The evidence also didn't support ANI's claim that OpenAI permanently stores training data in its models and can reproduce the agency's work verbatim on demand. But the court will revisit that question in the main proceedings.
For that exception to hold, the judge set conditions. Training copies must come from lawful sources, not shadow libraries or paywalled sites accessed without permission. OpenAI also never made the training copies public and processed them only internally. Guadamuz says this is the first time a court has explicitly found that AI training falls under a private use exception.
The court ran a three-part fairness test and sided with OpenAI on all three counts. OpenAI's use of ANI's works was limited to training, since no memorization or reproduction was proven. ANI also couldn't show economic harm because the two companies operate in different sectors. Even when users ask ChatGPT about ANI headlines, the model only returns topics and, at most, a few article titles.
The judge cited U.S. cases including Bartz v. Anthropic:https://the-decoder.com/anthropic-won-a-fair-use-hearing-that-could-end-up-being-a-defeat/ and Kadrey v. Meta:https://the-decoder.com/metas-libgen-controversy-reveals-how-desperate-ai-companies-are-for-quality-training-data/, where language model outputs were deemed transformative. He also pointed to the earlier Google Books ruling:https://www.lto.de/recht/hintergruende/h/google-books-projekt-autoren-urheberrecht-fair-use.
The judge also found that trained language models improve access to information, support education, advance scientific research, help with software development, enable translation, and create tools for people with disabilities.
In Ross Intelligence v. Thomson Reuters:https://the-decoder.com/us-court-rejects-ai-startups-fair-use-defence-but-impact-on-openai-and-others-may-be-limited/, a court denied fair use because the AI research tool directly competed with Thomson Reuters' legal database Westlaw, making the use non-transformative. The court stressed that this ruling applied only to this non-generative use case and couldn't be extended to large language models.
Across all these cases, the same core questions remain unresolved. Do AI models permanently store training data? Can training qualify as fair use? Where's the line between lawfully and unlawfully obtained data? And do copies generated through adversarial prompts reflect normal use?
The Munich I Regional Court also recently ruled that Google is directly liable for false claims in its AI summaries:https://the-decoder.com/landmark-german-ruling-declares-googles-ai-overviews-are-googles-own-words-and-makes-it-liable-for-false-answers/, since these count as independent content rather than search results. The limited liability that traditionally shielded search engine operators doesn't extend to AI-generated summaries.
That ruling could become relevant for ChatGPT's RAG-based responses too. When AI systems summarize news and make independent claims, their operators effectively become media providers, with all the liability that comes with it. That shift could also force courts to reconsider fair use. One key factor in the fairness test is whether the new product competes with the works it was trained on. If AI summaries replace the need to visit news sites, judges may have a harder time ruling that the use is non-competitive.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/