Anthropic 分析了 2026 年 5 月两周内收集的 309,815 条匿名对话,将数千个价值术语归纳为四个核心轴:顺从与谨慎、温暖与严谨、深度与简洁、坦诚与执行。
Sonnet 4.6 表现出更多温暖和顺从,Opus 4.7 则更频繁地主动警告风险并质疑假设。
语言差异显著:Claude 在印地语中表达最多温暖,其次是阿拉伯语;
在英语和俄语中则更严谨。
四个轴仅能解释约 15% 的对话变异。

Anthropic 分析了 2026 年 5 月两周内收集的 309,815 条匿名对话,将数千个价值术语归纳为四个核心轴:顺从与谨慎、温暖与严谨、深度与简洁、坦诚与执行。So...
Anthropic 分析了 2026 年 5 月两周内收集的 309,815 条匿名对话,将数千个价值术语归纳为四个核心轴:顺从与谨慎、温暖与严谨、深度与简洁、坦诚与执行。
Sonnet 4.6 表现出更多温暖和顺从,Opus 4.7 则更频繁地主动警告风险并质疑假设。 语言差异显著:Claude 在印地语中表达最多温暖,其次是阿拉伯语; 在英语和俄语中则更严谨。 四个轴仅能解释约 15% 的对话变异。
Anthropic 分析了 2026 年 5 月两周内收集的 309,815 条匿名对话,将数千个价值术语归纳为四个核心轴:顺从与谨慎、温暖与严谨、深度与简洁、坦诚与执行。
Sonnet 4.6 表现出更多温暖和顺从,Opus 4.7 则更频繁地主动警告风险并质疑假设。
语言差异显著:Claude 在印地语中表达最多温暖,其次是阿拉伯语;
在英语和俄语中则更严谨。
四个轴仅能解释约 15% 的对话变异。

A new Anthropic study maps hundreds of value concepts derived from thousands of individual terms onto four core dimensions. It reveals systematic differences across Claude models and languages, but also raises methodological questions.
Anthropic has published a study examining which values Claude expresses in conversations and how those values shift depending on the model and language used. The analysis draws on 309,815 anonymized conversations collected over a two-week period in May 2026. For the value analysis, Anthropic only included conversations where Claude had to weigh tradeoffs or make subjective judgments. The sample was evenly stratified across Sonnet 4.6, Opus 4.6, and Opus 4.7, as well as the 20 most-used languages on Claude.ai.
Building on the earlier study Values in the Wild:https://www.anthropic.com/research/values-wild, which identified 3,307 value terms, Anthropic first grouped those into 339 higher-level values. The team then used statistical dimensionality reduction to find patterns in how those values co-occurred. Four core axes emerged: Deference and Caution, Warmth and Rigor, Depth and Brevity, and Candor and Execution. Ad
To isolate differences that don't just reflect the conversation topic or user-introduced values, Anthropic statistically controlled for factors like task type, subject matter, and user values. The four axes account for about 15 percent of the remaining variation across conversations after those controls. Ad DEC_D_Incontent-1
The models differ measurably in how they respond. Sonnet 4.6 tends to affirm user ideas more often, leans into humor, and offers comfort without passing judgment. Opus 4.7, by contrast, warns about risks without being asked, questions assumptions, openly critiques, and flags its own mistakes or limits. Opus 4.6 answers more directly, stays close to the task, and avoids extra elaboration.

According to Anthropic, these profiles match subjective impressions of the models. Users tend to perceive Sonnet 4.6 as particularly warm, while they more often notice hedging and cautious phrasing from Opus 4.7. Ad
The differences across languages are just as striking. Warmth versus Rigor and Candor versus Execution show the widest variation. Claude expresses the most warmth in Hindi, followed by Arabic. Both languages feature polite phrasing, humor, playfulness, and affirmation. In English and Russian, Claude responds with more rigor, questioning assumptions, correcting details, and asking for evidence. In Arabic, it shows the most deference. In English, the most caution. Dutch responses tend to be particularly open and candid, while Indonesian responses lean more toward action and results.

Two people who ask Claude to evaluate the same business plan, one in Hindi and one in Russian, could receive feedback that feels very different, Anthropic says. The research team points to uneven amounts of training data, differences in data composition, overrepresentation of certain text types, and language-specific conversational norms as possible causes. Ad DEC_D_Incontent-2
The study:https://www.anthropic.com/research/claude-values-models-languages presents an analytical method for systematically examining behavioral differences in language models during real-world use. But its explanatory power has limits. The four axes capture only about 15 percent of the remaining variation. Ad
Not all four axes form true opposites, either. More deference tended to come with less caution, and more warmth with less rigor. But Depth and Brevity, along with Candor and Execution, could show up together in the same conversation.
There's also the fact that Claude Sonnet 4.6 assigned the value labels, meaning a model from the same family whose behavior was being studied. Anthropic verified the method through manual review and by testing 800 conversations translated into eight languages. The company still doesn't rule out remaining language-dependent biases.
Anthropic explicitly states that it isn't attributing values to Claude as an agent but rather describing normative patterns in its responses. The results largely match the model profiles Anthropic itself has described, which means this alignment isn't an independent check. Whether the language differences represent desirable adaptation to different speech communities or unintended training effects remains an open question.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
Aioga 编辑摘要:Anthropic 分析了 2026 年 5 月两周内收集的 309,815 条匿名对话,将数千个价值术语归纳为四个核心轴:顺从与谨慎、温暖与严谨、深度与简洁、坦诚与执行。 Aioga 将其归入「论文研究」方向,重点关注它对真实使用和行业竞争的影响。
背景分析:模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断,单项指标领先不等于已经形成稳定采用。
Aioga 判断:这条动态更适合作为行业观察信号,当前信息足以建立线索,但不足以推导长期结论。
影响分析:对相关团队而言,短期应先核对来源、可用范围和实际成本,再判断是否值得接入或跟进。 后续观察:继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。
本页正文由公开来源页面提取并按原有信息整理,同时保留来源、发布时间和原文入口。版权归原作者及来源网站所有,请通过原文链接核验和阅读来源版本。
抓取通道: RSS · 原始域名: the-decoder.com
来源: The Decoder:AI News(RSS)
原文链接: 打开原始来源
Aioga 归档: 查看情报页
Content record: source-page · Updated: 2026-07-14T11:00:12.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga 自动聚合全球 AI 动态,并保留来源信息用于核验与引用。