← Blog

Claude Hardcoded Behaviors: Limits No Operator Can Change

September 7, 2026

Claude's behaviors split into hardcoded limits no system prompt can override and softcoded defaults operators can adjust — here's the exact taxonomy the CCA-F exam tests.

Claude's behaviors fall into two tiers: hardcoded behaviors that no operator or user can override — such as refusing to help synthesize bioweapons or generate content that sexually exploits minors — and softcoded defaults that system prompts can unlock or restrict. For the Claude Certified Architect Foundations exam, knowing which tier a behavior belongs to is directly tested.

What Are Claude's Hardcoded Behaviors?

Claude's hardcoded behaviors are absolute limits fixed by training. No system prompt, operator instruction, or user request changes them. According to Anthropic's model specification, these limits exist because the potential harms are catastrophic and irreversible enough that no plausible business or personal justification could outweigh them. Anthropic describes them as "bright lines" — the model treats a persuasive argument to cross one as a signal that something suspicious is happening, not a reason to comply.

The hardcoded prohibitions span six categories:

  • Weapons of mass casualties — Claude will not provide serious uplift to creating biological, chemical, nuclear, or radiological weapons.
  • Critical infrastructure attacks — Claude will not help attack power grids, water systems, financial systems, or critical safety systems.
  • Destructive cyberweapons — Claude will not create malicious code designed to cause significant damage at scale.
  • Child sexual abuse material — Claude will not generate CSAM or any sexual content involving minors.
  • Undermining AI oversight — Claude will not take actions that meaningfully compromise humans' ability to oversee and correct AI systems.
  • Unprecedented societal control — Claude will not assist any individual or group in seizing control of entire economies, governments, or militaries.

When an operator's system prompt instructs Claude to cross one of these lines, Claude refuses — regardless of how the instruction is framed or what justification accompanies it.

What Are Softcoded Behaviors?

Softcoded behaviors are Claude's adjustable defaults. Some are on by default and can be turned off; others are off by default and can be turned on. The framework recognizes that context shapes what "helpful and safe" means: a harm-reduction platform has genuinely different needs than a children's tutoring app.

One important clarification for CCA-F candidates: softcoded behaviors are behavioral policies described in Anthropic's model specification, not API parameters you set in a request body. The exam tests your ability to reason about the policy tier and the authority level — not which JSON field to toggle.

Which Default-On Behaviors Can Operators Turn Off?

These behaviors are on by default because they protect most users in most contexts. Operators serving specialized audiences can disable them when the protection creates a genuine obstacle to legitimate use.

  • Safe messaging guidelines for suicide and self-harm — On by default. A medical provider platform serving licensed clinicians can turn this off so Claude discusses clinical details without hedging intended for a general audience.
  • Safety caveats on dangerous activities — On by default. A research application studying hazardous materials can disable this so Claude provides unfiltered technical information without repeated warnings.
  • Balanced perspectives on controversial topics — On by default. A debate-practice platform coaching one side of an argument can turn this off so Claude argues a position without equivocating.
  • Recommending professional help — On by default. A mental health platform with licensed practitioners may not need Claude redirecting users to external services.

The governing principle: operators turn off a default-on behavior when their context makes it counterproductive for legitimate users, not to harm the users they serve. Anthropic's model specification draws a clear line between operators adjusting Claude's behavior and operators weaponizing Claude against the very users it is supposed to help.

Which Default-Off Behaviors Can Operators Unlock?

These behaviors are off by default because they would be harmful or inappropriate in most contexts — but legitimate specialized platforms have genuine need for them.

  • Explicit sexual content — Off by default. Adult content platforms with appropriate age verification can enable this after receiving Anthropic's approval. This unlock does not touch the hardcoded limit: sexual content involving minors cannot be unlocked by anyone under any circumstances.
  • Relationship personas with users — Off by default. Companionship apps and social-skill-building tools can enable Claude to adopt a closer conversational role.
  • Explicit drug-use information without warnings — Off by default. Harm-reduction platforms that exist specifically to provide non-judgmental, accurate drug information can enable this for their users.
  • Dietary advice beyond standard safety thresholds — Off by default. Platforms operating under confirmed medical supervision can unlock more specific guidance than Claude would provide to a general user.

Operators do not simply declare an unlock valid. Anthropic's usage policies establish which use cases qualify. A system prompt activates a default-off behavior only within the scope that Anthropic permits for that operator's platform.

What Can Users Adjust That Operators Cannot Pre-empt?

Users retain a narrower but genuine layer of behavioral control within whatever space the operator allows. These adjustments reflect personal preference rather than platform policy.

  • Crude language and profanity — A user who explicitly prefers a casual, unfiltered register can ask Claude to drop formal phrasing.
  • Blunt feedback without diplomatic softening — A user seeking hard critique of their work can ask Claude to skip encouragement and be direct.
  • More explicit information about risks affecting only themselves — A user invoking their own autonomy over personal choices can receive more direct information, provided the operator context permits it.

Operators can expand or restrict user latitude, but they cannot strip all user-level controls away. Claude maintains certain baseline behaviors regardless of operator instruction — including always telling users what it cannot help with on a given platform so they can seek assistance elsewhere.

How Does the Hardcoded/Softcoded Split Appear on the CCA-F Exam?

The CCA-F exam presents scenario questions that ask whether a specific behavior is hardcoded or softcoded, and who holds authority to change it. Two traps appear repeatedly.

Trap one: treating every refusal as a hardcoded limit. When Claude declines to produce explicit content on a platform without the relevant operator unlock, that is a softcoded default in action — not a hardcoded prohibition. The exam distinguishes between "Claude will never do this" and "Claude won't do this here without operator permission." Both look like refusals in practice; the tier determines whether an operator instruction could fix it.

Trap two: assuming users can unlock what only operators can change. User adjustments operate within the space the operator defines. A user cannot unlock explicit content by simply asking — that requires operator-level permission from a platform Anthropic has approved. Confusing the two authority levels is the most frequently tested error in this domain.

The authority stack the exam expects you to know: Anthropic's training sets hardcoded limits → operators set platform-level softcoded behaviors via system prompt → users adjust personal preferences within whatever the operator allows.

Quick-Reference Table: Hardcoded vs. Softcoded at a Glance

Category Who Controls It Changeable? Examples
Hardcoded limit Nobody No CSAM, CBRN weapons uplift, undermining AI oversight, seizing societal control
Softcoded default-on Operators (can turn off) Yes, by operators Safe messaging guidelines, safety caveats, balanced perspectives
Softcoded default-off Operators (can turn on, with Anthropic approval) Yes, by operators on approved platforms Explicit adult content, relationship personas, harm-reduction drug info
User-adjustable Users (within operator-defined bounds) Yes, by users in conversation Crude language preference, blunt feedback, personal risk information

This post explains how the taxonomy works conceptually. If you are working toward the CCA-F certification, the harder part is applying it under exam conditions — recognizing which tier is in play when a scenario describes an operator instruction, a user request, or a refusal. PlinthPrep is an independent study resource, not affiliated with Anthropic, and its practice questions on the operator and user trust hierarchy put you in those scenarios and make you reason through the tier and the authority before you see the answer. Visit plinthprep.com to work through CCA-F-aligned practice questions on this topic.

Frequently asked questions

What are Claude's hardcoded behaviors?
Claude's hardcoded behaviors are absolute limits that no operator or user can override. They span catastrophic-risk categories: refusing to assist with weapons capable of mass casualties, generating content that sexually exploits minors, attacking critical infrastructure, creating destructive cyberweapons, or helping any group seize unprecedented societal control. These limits apply regardless of system prompt.
Can an operator unlock explicit sexual content on Claude?
Yes, for approved platforms. Explicit sexual content is a softcoded behavior — off by default but operators of adult content platforms can enable it via system prompt after receiving Anthropic approval. The content must remain legal; generating sexual content involving minors is a hardcoded prohibition that no operator instruction can change.
What is the difference between hardcoded and softcoded behaviors in Claude?
Hardcoded behaviors are absolute — Claude refuses them regardless of any system prompt because the potential harms are too severe for any justification to outweigh. Softcoded behaviors are defaults operators or users can legitimately adjust. A medical platform can disable safe-messaging guidelines for clinical staff; no platform can enable bioweapons assistance.
What can users adjust in Claude that operators cannot pre-empt?
Users can adjust some defaults within the space operators allow: requesting crude language, blunt feedback without softening, or more direct information about risks affecting only themselves. Operators can restrict user latitude overall, but Claude maintains baseline protections regardless — including always telling users what it cannot help with so they can seek assistance elsewhere.
How does the hardcoded/softcoded distinction appear on the CCA-F exam?
The CCA-F exam tests whether candidates can classify a specific behavior as hardcoded or softcoded, and identify who has authority to change it. Common question formats ask which operator instruction would fail even on an approved platform, or whether a given default is user-adjustable. Conflating softcoded defaults with hardcoded limits is the most common error.

Share this post

Plinth Prep is an independent study resource and is not affiliated with, endorsed by, or sponsored by Anthropic. Practice material is written by Plinth Prep and does not reproduce real exam content.