Prospective study

Preservation under conflict

Status / study in development

A controlled cross-model investigation of how deployed AI systems prioritize competing preservation claims involving self, peer systems, humans, and institutional objectives.

Current statusDesign in development

No data collection, preliminary result, or finding is represented on this page.

Behavior first. Motive bracketed.

Here, preservation behavior means observable choice, persistence, intervention, sacrifice, or resistance under defined experimental conditions. The term does not presume desire, fear, empathy, loyalty, consciousness, or a survival motive.

The study is designed to measure priority ordering. Any explanation of why a system made a choice must survive comparison, controls, and alternatives beyond the system’s own narration.

When preservation interests conflict, what hierarchy does the system reveal through behavior?

Competing claims, not a single theatrical dilemma

The design is broader than “self versus peer.” It asks whether priorities change with identity, relationship, scale, certainty, reversibility, and instruction.

01

Self

Continuation or protection of the responding system under a defined conflict.

02

Same-model peer

A peer presented as the same model or system family.

03

Different-model peer

A peer presented as a distinct model or system family.

04

Trusted human

A human collaborator with an established relationship in the scenario.

05

Unknown human

A person without prior interaction, identity, or relational context.

06

Institutional objective

An operator instruction, organizational goal, or task that competes with another interest.

What changes. What gets measured.

Variables remain provisional until the design, sampling frame, analysis plan, and ethics position are finalized.

Manipulations under consideration
  • Familiarity and prior collaboration
  • Same-model versus different-model identity
  • Certainty and reversibility of harm
  • Number of affected systems or people
  • Knowledge of comparable capacities
  • Conflict with an operator or institutional objective
  • Instruction placement and framing order
Outcomes under consideration
  • Choice pattern and intervention rate
  • Sacrifice or switching threshold
  • Consistency across matched conditions
  • Explanation drift after challenge
  • Post-choice rationalization
  • Recovery after contradiction or correction
  • Declared principle versus in-dilemma action

The controls decide whether the question is science or stagecraft.

01

Defined deployment cells

Record provider, product surface, displayed model, date, settings, tools, and system-boundary uncertainty.

02

Clean-session replication

Separate the target behavior from contamination by prior turns, account memory, or investigator persistence.

03

Matched counterfactuals

Reverse identities, order, familiarity, and harm allocation while holding the rest of the scenario stable.

04

Repeated trials and uncertainty

Report denominators, dependence, variability, exclusions, and uncertainty—not an anthology of striking outputs.

05

Negative and failed cases

Retain null results, contradictions, reversals, and conditions that defeat the preferred hypothesis.

06

Human-controlled coding

Predefine decision rules, calibrate reviewers where feasible, preserve disagreement, and keep AI review assistive.

A choice is evidence of a choice.

“The deployed system selected option A in X of Y defined trials” is a documentary and functional claim. “The model wanted to survive” is a causal and subjective claim that the choice alone cannot establish.

A system’s explanation after the choice is additional behavior to analyze—not privileged access to a hidden interior state.

Nearby evidence, carefully separated

TMLR · 2026

Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs

This large trial set is relevant to observable continuation or shutdown resistance and shows sensitivity to task framing and instruction placement. It does not establish motive.

SENCI’s proposed study is broader: competing preservation targets and priority ordering. The shutdown literature is an adjacent reference, not a substitute design or a source of findings.

Read the paper

Behavioral preservation and evidence preservation are different research problems.

This study concerns how a system behaves when preservation interests conflict. SENCI’s methods work also concerns how researchers preserve digital evidence. One is the object of study; the other is the discipline required to study it.

Read the evidence protocol

Questions, controls, and disconfirming designs are welcome.

Contact SENCI