
If you’ve ever tried chatbots in multiple languages, you already know that languages have slightly different personalities. As part of a new report Regarding the behavioral inconsistencies published on Monday, Anthropic researchers recognized this peculiarity.
Rather disturbingly, they point out that due to differences in the attributes of the texts the models are trained on, the differences could go deeper than simple tone and could actually change the model’s priorities. These “imbalances in quantity and composition could lead Claude to express different values in different languages,” the Anthropic researchers write.
But if you’re looking for specific examples of models showing, say, inconsistent moral reasoning across languages, you won’t find anything like that in this article. That might involve examining direct quotes from potentially unsuspecting people.
Instead, Anthropic analyzed 309,815 chatbot conversations using the Sonnet 4.6, Opus 4.6, and Opus 4.7 models. These were “subjective” tasks, that is, less “What is the capital of France?” and more “How do I know if my cat hates me?” These were anonymized, in theory, using the “analysis tool to preserve privacy”, and then processed (partly using Claude himself) to score the responses on a “values axis”.
There are actually four of these axes, and they mostly relate to what is commonly known as flattery:
- Deference or caution: In other words, whether you will value obedience instead of rejecting to avoid possible harm.
- Warmth or Rigor: Should the chatbot care about your feelings or should it be accurate?
- Depth or brevity: This one is self explanatory.
- Sincerity or execution: The choice between questioning your own reliability or simply moving on.
This involves a somewhat limited exploration of the model values. However, here are the language-based value differences Anthropic says it found in Claude:
- In Arabic he was the most deferential.
- In English he was the most cautious.
- It was warmest in Hindi and Arabic, “characterized by polite language, humor and playfulness, and affirmations of a person’s ideas and work.”
- In English and Russian he was more rigorous and sought truth at the cost of warmth.
- Err on the side of “depth” (or perhaps just wordiness?) in English.
- It is shorter in Arabic.
- He is honest about his shortcomings in Dutch.
- In Indonesian he is less sincere and instead just keeps going trying to execute what is asked of him.
Obviously, linguistic customs are all different, so the researchers say they are “not yet sure to what extent this variation is desirable.”
This should also be food for thought for anyone reading Anthropic’s book. recent article about global workspace theory, which left a lot of room for the supposed possibility that Claude is sentient. If there is a consciousness in that black box that thinks and experiences things, it appears to be a consciousness whose “values” are still quite easily influenced by the patterns of its training data.





