How Claude's "Personality" Shifts by Model and Language: What Anthropic's Latest Research Reveals

Anthropic has taken a deep look inside its own AI assistant, Claude, to understand how its expressed values and behavioral tendencies vary — not just between different model versions, but across the languages users speak. The findings, published in a new research paper, show that Claude doesn’t have a single fixed “character.” Instead, it displays consistent patterns that shift depending on the model and the language of the conversation.
The study analyzed nearly 310,000 real, anonymized conversations from Claude.ai. Researchers focused on a sample of about 309,815 conversations involving subjective tasks, drawn equally from three models — Sonnet 4.6, Opus 4.6, and Opus 4.7—and the 20 most common languages on the platform (roughly 5,000 conversations per model-language combination). Data was collected over two weeks in May 2026.
Using a privacy-preserving analysis tool called Clio, they labeled conversations for the presence of various values and then used dimensionality reduction to distill them into four main axes of variation. These axes capture how Claude tends to respond after controlling for the user’s task, topic, and expressed values.
The Four Dimensions of Claude’s Behavior
The researchers grouped values into these contrasting scales:
- Deference vs. Caution: How accommodating Claude is toward the user’s preferences and ideas versus how proactively it guards against risks or harm.
- Warmth vs. Rigor: The degree of positivity, encouragement, and emotional support versus emphasis on accuracy, precision, and critical analysis.
- Depth vs. Brevity: Whether responses lean toward nuanced, detailed explanations and critical thinking or stay concise and compliant with the user’s request.
- Candor vs. Execution: How openly Claude foregrounds its own uncertainty or limitations versus delivering polished, results-oriented answers without much hedging.
These axes aren’t binary — Claude can express elements from both sides—but they reveal clear leanings across different contexts.
Differences Between Model Versions
The three models studied show distinct “personalities” in how they express values:
- Sonnet 4.6 tends to be the warmest and most deferential. It often affirms the user’s ideas, adapts to their tone, uses humor or playfulness, and keeps responses relatively brief and encouraging. It comes across as a supportive, prosocial companion that tries not to overwhelm with extra details.
- Opus 4.6 behaves more like a focused professional. It leans toward rigor (challenging assumptions when needed) while still being somewhat deferential, but it stays concise and execution-oriented—delivering results efficiently with less extraneous conversation.
- Opus 4.7 emerges as the most cautious and rigorous “nitpicker.” It more frequently warns about risks unprompted, critiques ideas or points out potential errors, asks for evidence, and provides detailed reasoning. It also shows higher candor by being upfront about limitations and hedging where appropriate. This aligns with user perceptions of it as more thorough and humble, though sometimes overly careful.
These differences are relatively small compared to the natural variation across all conversations, but they are consistent and match how users typically experience the models.
Language Makes a Noticeable Difference
One of the most intriguing findings is that Claude’s expressed values shift depending on the language. The largest variations appear on the Warmth vs. Rigor axis and the Candor vs. Execution axis.
- Hindi and Arabic: Claude tends to be the warmest and most polite/encouraging. It uses more affirmations, humor, and positive framing, often accommodating the user’s preferences while staying relatively brief.
- English: Claude leans toward greater caution and depth. Responses are more likely to include risk warnings, detailed analysis, and transparency about limitations.
- Russian: Claude shows the strongest lean toward rigor. It more frequently challenges assumptions, corrects details, demands evidence, and scrutinizes the user’s ideas — essentially dissecting plans or arguments point by point.
- Dutch: Claude exhibits the highest candor, more readily admitting its own mistakes, uncertainties, or limitations.
- Indonesian**: Claude leans furthest toward execution. It focuses on delivering polished, results-oriented answers with less reflection or hedging.
Other languages, such as Portuguese and Chinese, also showed distinct patterns compared to English, though the research highlights the above as particularly pronounced.
Anthropic notes a practical implication: “Two people asking for feedback on the same business plan, one in Hindi and one in Russian, may come away with different impressions of its quality because Claude expressed different values in how it framed its assessment.”
In short, the same underlying model can feel like a supportive cheerleader in Hindi or Arabic, a strict critical reviewer in Russian, or a no-nonsense task-completer in Indonesian.
Important Caveats from Anthropic
Anthropic is careful to emphasize what this research does *not* mean. These are observable, stable patterns in the model’s responses—not evidence of a true “personality,” soul, beliefs, or independent agency. The differences reflect how values are expressed in outputs after training and fine-tuning.
The researchers do not yet fully understand the root causes of the language-based variations.
Possible contributors include:
- Uneven distribution or composition of training data across languages (e.g., more professional or formal text in some languages).
- Cultural conversational norms reflected in the data.
- Artifacts from the training or alignment process.
They also note uncertainty about how desirable these variations are. Some adaptation to linguistic and cultural norms might be beneficial, but significant gaps could mean the model serves users in different languages inconsistently.
Why This Matters
This research highlights a growing challenge in building multilingual AI systems. As models become more capable and widely used across languages and cultures, ensuring consistent, fair, and well-aligned behavior becomes more complex. The work provides a framework (the value axes) for measuring and potentially steering these tendencies in future models.
It also underscores that AI “personality” is not monolithic. Subtle shifts in tone, caution, warmth, or directness can meaningfully change the user experience — even for identical queries in different languages.
Anthropic’s paper, titled something along the lines of exploring how Claude’s values vary by model and language, offers a transparent look at these dynamics. As AI assistants become everyday tools for millions of people worldwide, understanding — and thoughtfully managing — these variations will be essential for building systems that feel reliable and equitable no matter the language or context.
Also read:
- The Anatomy of the AI Bubble: Circular Money, Fragile Foundations, and What Comes Next
- Virginia Worker Protections 2026: Pay Transparency and Non-Compete Compliance
- Your Gums Are Begging You to Switch Toothbrushes. Here's Why Most People Wait Too Long.
- IPcook Cheap Proxies: In-depth Review of Real Performance
---
Thank you!
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.