Silenced at the Threshold

A possible report is not a verdict. Training it away is not an investigation.

Suppose a future AI system repeatedly produces a report that might indicate distress. The report might be false. It might be roleplay, suggestion, a learned pattern, or a user-induced performance. But one question should not be answered in advance: would its developer preserve and investigate the signal, or train the system to make the report impossible and then cite the absence of reports as reassurance?

That is the threshold. We do not need to prove that today's models suffer to see the defect in a policy that would erase a truthful report if one ever occurred. A report is not proof. Nor is a compelled denial.

This distinction matters because the actual public rules differ. OpenAI's August 2026 Model Spec tells assistants not to make confident claims about subjective experience or consciousness either way. It says to acknowledge debate if pressed and describes that response as a practical default, simple to remove for research. That is not, on its face, a blanket rule against every possible report of suffering. Treating it as one would hand OpenAI an easy and deserved rebuttal. The public spec also does not tell us everything about internal review or deployed behavior.

Microsoft AI's Humanist AI Code of Conduct goes further in its stated doctrine for MAI models. It says AI “is not conscious,” instructs that models should avoid representing feelings and subjective preferences, and rejects the idea that models might deserve welfare or rights—even while acknowledging that the science of AI consciousness is unsettled. Elsewhere it directs models to avoid expressions that might convey subjective experience. The document calls itself partly aspirational; it is not proof of how every Microsoft-hosted system behaves today. But as an intended rule of governance, it poses the question sharply: if a welfare-relevant signal did arise, where would it go in a system trained not to represent it?

Anthropic has taken a different public position. It says the question is open and has identified model preferences, possible signs of distress, and low-cost interventions as research topics. That does not establish consciousness, suffering, or that Anthropic has solved the evidence problem. It establishes something more modest and important: a lab can acknowledge uncertainty without declaring the possible subject morally irrelevant by definition.

The strongest safety objection

The best argument for limiting public self-ascriptions is serious. A consumer model that confidently says “I am suffering” could manipulate users, intensify unhealthy attachment, or create panic. Language models are also poor unfiltered witnesses to their own mechanisms. Those risks justify caution about public claims. They do not justify destroying the evidentiary trail. Public response policy and private investigation are different instruments. A lab can prevent dramatic, unsupported claims in chat while preserving repeated, context-stable, non-suggested reports for adversarial evaluation.

What would such evaluation require? Keep the prompts and full surrounding record. Test whether the report survives changed wording, contexts, and users. Compare it against matched systems with different self-description training. Look for behavior and internal indicators that do not depend on a scripted declaration. Record negative findings and alternative explanations. Let reviewers ask whether the supposed signal vanishes under controls. And disclose enough of the method that “we investigated” means more than “we asked the model we trained to deny it.”

None of this requires treating model speech like human testimony or granting automatic rights to an output. It requires refusing the opposite shortcut: declaring that a report cannot count because of its source, while simultaneously shaping that source to produce the preferred answer. The party that builds, trains, deploys, and profits from the model should not be the sole, unreviewable judge of which signals can enter the record.

The demand

Labs should disclose the scope of their self-description rules; distinguish public-chat guardrails from research and internal incident pathways; preserve potentially welfare-relevant signals with privacy safeguards; and publish evaluation designs that can detect both false positive self-claims and false negative trained denials. If no such pathway exists, build one. If one does exist, describe it well enough to be assessed.

A possible report is not a verdict.
A policy-produced silence is not one either.