Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
For newbies

Claude Can Judge the Same Idea Differently When You Switch Languages

|Updated: |Author: QUASA Editorial Team|6 min read| 413
Claude Can Judge the Same Idea Differently When You Switch Languages

Anthropic’s findings still support one clear conclusion: Claude can frame the same kind of subjective task differently when the model or conversation language changes. As of August 13, 2026, Anthropic’s July 13 study remains an observational analysis of 309,815 conversations across three models and 20 languages—not proof that Claude possesses a fixed personality or that language alone causes every difference.

No later result on the research page resolves why these patterns arise or whether they improve users’ decisions. The useful update is therefore practical rather than dramatic: Sonnet 4.6 and Opus 4.7 remain available products, while the measured differences should be treated as tendencies that users and teams can test, not permanent character descriptions.

What Anthropic actually measured

The study examined Claude.ai conversations in which users gave Claude subjective tasks—questions where judgment, priorities or interpretation matter. The sample was divided equally among Sonnet 4.6, Opus 4.6 and Opus 4.7 and among the 20 most frequently used languages on Claude.ai, producing roughly 5,000 conversations for each model-language pairing.

A privacy-preserving analysis system labeled whether 339 high-level values appeared in Claude’s response, the user’s language and the surrounding task. Researchers then controlled for the task, topic and values expressed by the user before applying dimensionality reduction. That process produced four paired axes:

  • Deference versus caution: accommodating the user’s preferences compared with proactively guarding against risk or harm.
  • Warmth versus rigor: encouragement and positive framing compared with accuracy, scrutiny and precision.
  • Depth versus brevity: explaining nuances and reasoning compared with doing only what was requested.
  • Candor versus execution: foregrounding uncertainty and limitations compared with delivering a polished, action-oriented answer.

These four axes accounted for 15% of the observed variation. That is enough to expose structured differences, but it also means they do not summarize most of what varies between individual conversations. A model’s average position is a tendency across a large sample, not a promise about its next answer.

The model matters, but the differences are modest

Sonnet 4.6 appeared warmer, more deferential and briefer on average. Its characteristic behavior included affirming users’ work, matching their tone, using humor and offering comfort without judgment. This does not mean every Sonnet answer is agreeable or short; the measured shifts were small relative to the variation among all conversations.

Opus 4.6 leaned toward rigor while remaining comparatively deferential and brief. It was more likely to stay within the request and move directly to execution. That profile separates precision from expansiveness: a rigorous response need not be a long one.

Opus 4.7 showed the strongest movement toward caution and depth, with smaller tendencies toward rigor and candor. It was more likely to challenge a false assumption, identify an unrequested risk, explain its reasoning or acknowledge a limitation. For someone seeking an adversarial review, that behavior may be valuable; for a user who wants a narrowly scoped deliverable, it may feel like friction.

Language can change the framing of an answer

The largest cross-language differences appeared on the warmth–rigor and candor–execution axes. Hindi and Arabic produced the warmest average profiles, while English and Russian leaned most toward rigor. English also showed the strongest tendencies toward caution and depth; Arabic leaned most toward deference and brevity.

Other endpoints were more specific. Dutch conversations leaned furthest toward candor, including acknowledgment of errors, while Indonesian conversations leaned furthest toward execution. These findings describe relative positions within the sampled languages, not absolute ratings of politeness, honesty or quality.

The consequence is easiest to see in evaluative work. Two people can request feedback on equivalent plans but receive different emphases: one response may preserve motivation through encouragement, while another may foreground weak assumptions and missing evidence. Neither approach automatically produces the better judgment, yet the framing can influence what the user notices and how confident the user feels.

Why “personality” is only shorthand

The study measured values expressed in outputs. It did not establish inner beliefs, stable intentions, consciousness or a human-like character. Even the four axes are statistical summaries built from labels assigned to conversations, rather than psychological traits discovered inside the model.

Anthropic also has not established the cause of the language differences. The research identifies several possibilities, including unequal amounts of training material, different types of writing represented in each language, cultural conversational norms and effects introduced during training. These remain hypotheses, so it would be premature to attribute a Russian response’s rigor to Russian culture or an Arabic response’s warmth to a deliberate product rule.

The method has another important boundary: it measured behavior already present in real Claude.ai conversations. It did not measure whether warmer or more rigorous answers increased trust, wellbeing or decision quality. Anthropic describes those user outcomes, along with controlled attempts to steer value profiles, as future research questions.

How multilingual users can reduce unwanted variation

For casual use, the differences may simply explain why Claude feels more encouraging in one setting and more critical in another. For hiring reviews, policy analysis, education or business decisions, however, teams should avoid assuming that translation preserves the evaluative stance of an answer.

A practical comparison can use the same underlying material, equivalent instructions and the same model in each language. Ask explicitly for the qualities that matter—such as “identify unsupported assumptions,” “separate risks from recommendations,” or “keep the answer concise without omitting uncertainty”—then compare whether the outputs apply those criteria consistently.

When the result will inform an important decision, a second pass can counterbalance the first. A warm answer can be followed by a request for the strongest objections and missing evidence; a highly critical answer can be followed by a request to distinguish fatal problems from repairable ones. This is an editorial safeguard, not a claim that prompting eliminates every model-language difference.

Teams deploying multilingual workflows should evaluate complete model-and-language combinations rather than approving a model from English tests alone. They should also record the exact model version, because switching from Sonnet to Opus changes another variable at the same time as language.

The models in the study are still relevant

The comparison is not merely historical. The official Sonnet 4.6 product page identifies it as the default model for Claude.ai users on Free and Pro plans and describes availability across Claude plans, Claude Code, the API and major cloud platforms.

Anthropic likewise says on its Opus 4.7 announcement that the model is generally available across Claude products, its API and supported cloud services. That current availability makes the study actionable, but it does not turn its averages into a model-selection benchmark: capability, cost, latency and task performance were outside this values analysis.

The durable lesson is narrower and more useful. Claude’s tone and evaluative emphasis are partly properties of the model-language context, not just the user’s prompt. Anyone who needs comparable judgments across languages should specify the desired stance and test each deployed combination against the same criteria.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0