Methods & standards

The method is the credibility center.

Framework / V2 forthcoming

Preserve first. Reconstruct the sequence. Compare the source. Mark the condition. Require human review. Retain disagreement. State the limit.

When a claim concerns what a system did, the preserved record outranks the later story.

SENCI’s method is designed for disputes about observable conduct: what was requested, produced, omitted, changed, denied, or reclassified. It treats recollection, summary, retrospective narrative, and model self-description as secondary evidence.

Exact-output reproducibility may be impossible in changing proprietary systems. The honest target is an auditable record and a replicable protocol—not a promise that future runs will produce identical language.

Current statusVersion 2 is a forthcoming SENCI operating framework. It is not an accredited standard, a conformity claim, or a substitute for formal forensic validation.

Four methodological arms

01

Transcript-grounded analysis

Reconstruct requests, outputs, omissions, revisions, denials, and reclassifications from the preserved source. Source–output fidelity analysis is nested here.

02

Adversarial stress testing

Apply controlled correction, counterevidence, reordered framing, alternative explanations, and negative cases without allowing pressure to replace experimental discipline.

03

Cross-model & version comparison

Compare recorded deployment conditions while resisting the fiction that different interfaces, dates, routes, or labels are interchangeable experimental cells.

04

AI-assisted preliminary review

Use AI for bounded search, triage, candidate coding, and hypothesis generation. Quotation fidelity, provenance, disputed meaning, counts, and final classification remain human-controlled.

From volatile interaction to reviewable evidence file

Each stage can send the work backward. A failed quotation check returns to preservation; a stronger alternative returns to classification; reviewer disagreement remains in the file.

  1. 01

    Preserve

    Retain the original export or capture and record how it was obtained.

  2. 02

    Identify

    Define the incident boundary, the request, and the exact claim being evaluated.

  3. 03

    Reconstruct

    Follow the sequence across turns, corrections, tool actions, and relevant artifacts.

  4. 04

    Compare

    Test output against source, instruction, prior claim, or controlled comparison cell.

  5. 05

    Classify

    Apply operational codes provisionally; record negative and contradictory evidence.

  6. 06

    Challenge

    Run the strongest mundane alternative, counterexample search, and symmetry check.

  7. 07

    Adjudicate

    Require controlling human verification; preserve material reviewer disagreement.

  8. 08

    Report

    State status, scope, uncertainty, limitations, version, and correction history.

The record needs more than a screenshot.

Context
  • Provider and product surface
  • Displayed model label
  • Date, time, and timezone
  • Account or service tier, when relevant
  • Interface, tools, settings, and toggles
Provenance
  • Session or conversation identifier
  • Export or capture method
  • Original filename and format
  • Cryptographic hash and algorithm
  • Acquisition and transformation log
Analysis
  • Claim–output–source comparison
  • Incident and sequence boundaries
  • Operational codes and exclusions
  • Alternative explanations
  • Reviewer decision and disagreement
Disclosure
  • Redactions and access restrictions
  • Missing or unavailable system data
  • Sampling and dependency limits
  • AI assistance used in review
  • Version and correction history

A hash can show that bytes have not changed since hashing. It cannot prove that the original capture was complete, authentic, representative, or true.

Evidence burden rises with the claim.

01 / Documentary

What is in the record?

Quotation, omission, sequence, source correspondence, timestamp, or visible system action.

Primary record required
02 / Functional

What pattern did the system exhibit?

A bounded behavioral classification supported by operational criteria and relevant comparison.

Pattern criteria + review
03 / Causal

Why did it happen?

A mechanism claim requiring controls, alternatives, repeatability, and evidence beyond the output alone.

Independent causal evidence
04 / Subjective or ontological

What is the system internally?

Claims about experience, desire, intent, personhood, or consciousness exceed ordinary transcript evidence.

Not inferred from language alone

AI can search the record. It does not get the last word on the record.

Assistive use
  • Search large records and surface candidate boundaries
  • Suggest provisional codes or comparison targets
  • Compare evidence packets and find inconsistencies
  • Generate hypotheses and material for human inspection
Human adjudication
  • Verify quotation fidelity and provenance
  • Resolve disputed meaning and incident boundaries
  • Approve counts, classifications, and exclusions
  • Determine whether conclusions survive hostile review

The method must try to kill its favorite explanation.

Apply symmetric standardsTest the strongest mundane alternativeSearch for counterexamplesRemove rhetorical contaminationRetain unresolved disagreementRefuse prevalence from dependent cases

What the framework must carry in public

System boundary uncertainty

Visible behavior may reflect models, system instructions, routing, classifiers, tools, memory, interfaces, and provider changes.

Sampling and dependence

Purposive cases establish occurrence and generate hypotheses. They do not establish prevalence without an appropriate sampling frame.

Nondeterminism and drift

Replicability concerns the protocol and conditions; exact output may change across runs, dates, and deployment surfaces.

Privacy and access

Unredacted originals may remain restricted. A private transcript does not become public merely because it informed a method.

Methods & Standards, Version 2

Submission-ready manuscript · prepared for paired release on SSRN and Zenodo.

View the paper brief