Anthropic published a new constitution for Claude on 22 January. The document described the values, priorities, and reasoning the company wanted to shape the model during training. Anthropic released it under CC0, allowing others to read and reuse it freely.
What a model constitution is
A constitution is not a list that the model consults perfectly before every answer. It is part of the training approach. Anthropic said Claude might not always behave according to the ideals in the document, which is an important limit when reading it.
What behaviour does the document say it wants?
What happens when two principles conflict?
Does the model act consistently in repeated tests?
How to read a company safety document
The useful way to examine a model constitution is to compare its stated priorities with observed behaviour. Look for how it handles conflicting requests, uncertainty, user autonomy, harmful instructions, and situations where several reasonable values point in different directions.
Key takeaways
- Claims about behaviour are supported by tests.
- The model's limits are described alongside its aims.
- External researchers can reproduce the evaluation.
- Users still have a way to report failures.
Publishing the document made Anthropic's intentions easier to inspect and debate. It did not remove the need for independent evaluation, incident reporting, and tests based on real use.
