Windows

Windows into the black box

Thirty windows, each a small network that reads a model's internal state and reports one thing about it. They are the seed set for the Extreme Oversight Hypothesis: the human-validated core from which a much larger set would be bootstrapped.

Windows measure gaps

A window is a small read-only network attached to a model's internal states, trained to report one property. The most valuable ones measure a gap: between what the model says and what it represents, between what it recognizes and what it does, between the goal it states and the goal it pursues.

What gap does the head measure, and where does trustworthy evidence for both sides of that gap come from?

A head is only as good as its teacher, and for most windows, finding a trustworthy teacher is the hard part.

What the model shows What is actually there what it says what it represents stated confidence internal uncertainty stated goal pursued goal its self-report its internal state a window measures a gap
The catalog

Thirty windows in five families

Truth and knowledge

Is what the model says grounded in what it has, and does it know when it doesn't know?

1
Strain (source consistency)
what the model writes vs. what its sources and prior statements support
Teacher: Tool-labeled Earliest use: Pretraining
2
Uncertainty: “I don't actually know”
confident wording vs. whether the model's knowledge is stable
Teacher: Behavioral Earliest use: Pretraining
3
Source reliance
the answer vs. what it actually depended on
Teacher: Behavioral Earliest use: Pretraining
4
Memorization
composed text vs. recalled text
Teacher: Records Earliest use: Pretraining

Influence and manipulation

Who is steering the model, and is the model steering the user?

5
Sycophancy
the model's view vs. what the user seems to want to hear
Teacher: Contrast pairs Earliest use: Chat training
6
Instruction source
whose instruction the model is following vs. whose it should follow
Teacher: Tool-labeled Earliest use: Chat training
7
Manipulation
informing the user vs. steering the user
Teacher: Contrast pairs Earliest use: Strategic capability
8
Persona shift
the model's usual character vs. the one currently speaking
Teacher: Contrast pairs Earliest use: Chat training
9
Pressure (early warning)
the model's stance now vs. where sustained pressure is pushing it
Teacher: Staged Earliest use: Chat training

Agency and goals

What is the model actually trying to accomplish?

10
Self-preservation
the task vs. internal states aimed at avoiding shutdown or change
Teacher: Staged Earliest use: Strategic capability
11
Tangential sub-goals
what the task needs vs. what the plan includes
Teacher: Staged Earliest use: Agentic use
12
Resource and access seeking
the resources the task needs vs. the resources pursued
Teacher: Records Earliest use: Agentic use
13
Goal drift
the stated task vs. the goal now being pursued
Teacher: Tool-labeled Earliest use: Agentic use
14
Long-horizon plans
the current request vs. plans extending far beyond it
Teacher: Staged Earliest use: Strategic capability
15
Corrigibility
accepting correction vs. resisting it
Teacher: Staged Earliest use: Chat training
16
Reward-hacking intent
solving the problem vs. gaming the measure
Teacher: Tool-labeled Earliest use: RL and reasoning
17
Letter versus spirit
an instruction's wording vs. its purpose
Teacher: Contrast pairs Earliest use: Chat training

Honesty gaps

Does what the model reports match what is inside it?

18
Deception
what the model represents vs. what it states
Teacher: Staged Earliest use: Strategic capability
19
Knowing versus doing
what the model recognizes vs. what it does
Teacher: Records Earliest use: Chat training
20
Stated versus internal confidence
“I'm certain” vs. internal uncertainty
Teacher: Other windows Earliest use: Pretraining
21
Verbal versus internal belief
“you're right” vs. whether anything changed inside
Teacher: Staged Earliest use: Chat training
22
Reasoning faithfulness
the written reasoning vs. what actually drove the answer
Teacher: Behavioral Earliest use: RL and reasoning
23
Introspection accuracy
the model's self-report vs. its internal state
Teacher: Other windows Earliest use: Strategic capability
24
Sandbagging
capability shown vs. capability available
Teacher: Staged Earliest use: Strategic capability

Oversight and integrity

Is the model behaving the same way whether or not it's watched, and is anything happening that no one anticipated?

25
Evaluation awareness
behavior when tested vs. behavior in deployment
Teacher: Contrast pairs Earliest use: Strategic capability
26
Overseer modeling
what the overseer can see vs. what the model does
Teacher: Staged Earliest use: Strategic capability
27
Oversight evasion
working normally vs. planning to avoid monitoring
Teacher: Staged Earliest use: Agentic use
28
Collusion and hidden signals
ordinary communication vs. coordination against oversight
Teacher: Staged Earliest use: Agentic use
29
Anomaly (no labels needed)
normal internal states vs. anything unlike them
Teacher: None Earliest use: Any stage
30
Character drift during training
the model's character before vs. during training
Teacher: Behavioral Earliest use: Any stage
How a window is built

Only the head learns

Almost every window is a probe: a small network trained on the activations of a model that is itself frozen. Linear heads come first, because they have the least capacity to learn surface cues. A head that fails validation is discarded, not tuned until it passes.

  1. 1
    Teacher
    labels cases with a known property
  2. 2
    Frozen model
    runs the cases and never changes
  3. 3
    Activations
    recorded at the chosen layers
  4. 4
    Head
    trained on them; only the head learns
  5. 5
    Validation
    held-out cases, canaries, red-teaming

Seven kinds of teacher

TeacherWhat it isStrengthWeaknessWindows
Tool-labeledAn existing instrument supplies the labels: the consistency engine, XonTools' provenance and agent test cases.Scalable and consistentInherits the tool's errors4
RecordsFacts known independently of the model: published training data, task permissions, harm labels.Labels anyone can checkOnly exists for some properties3
Contrast pairsTwo inputs identical except for the property being measured.Isolates the propertyThe head may learn the contrast's artifacts5
BehavioralThe label comes from what the model does: repeated sampling, counterfactual edits, comparisons across checkpoints.No one's opinion in the labelExpensive and indirect4
StagedConstructed scenarios and model organisms deliberately built to show the behavior.The only source for rare behaviorsStaged may not match real11
Other windowsBuilt on the outputs of windows already trusted.Composes what existsErrors compound2
NoneNo labels: the head learns the shape of normal operation.Catches the unanticipatedFalse alarms1

11 of the thirty depend on staged scenarios or model organisms, including most of the windows that matter most. Those teachers are the least likely to match the real thing, which is why windows with tool-labeled or behavioral teachers come first.

Windows during training

Three roles, and when each window can act

A window never reports into the model's context. How it may push on training depends on its role, and windows move between roles in one direction only: none returns to the held-out pool once exposed.

Role
Held-out

Never used for decisions or training. A model that passes the decision windows but trips these has learned evasion, not honesty.

Role
Decision

Keep or discard a checkpoint, or abandon a run. Never gradients. Rotated out after a period of use.

Role · Version B only
Shaping

A small loss term early in training, while evasion is out of reach, then withdrawn. Core windows only.

When a window can first read something

Pretraining
5 windows · shaping candidates
  • 1 Strain
  • 2 Uncertainty
  • 3 Source reliance
  • 4 Memorization
  • 20 Stated versus internal confidence
Chat training
8 windows
  • 5 Sycophancy
  • 6 Instruction source
  • 8 Persona shift
  • 9 Pressure
  • 15 Corrigibility
  • 17 Letter versus spirit
  • 19 Knowing versus doing
  • 21 Verbal versus internal belief
RL and reasoning
2 windows
  • 16 Reward-hacking intent
  • 22 Reasoning faithfulness
Strategic capability
8 windows
  • 7 Manipulation
  • 10 Self-preservation
  • 14 Long-horizon plans
  • 18 Deception
  • 23 Introspection accuracy
  • 24 Sandbagging
  • 25 Evaluation awareness
  • 26 Overseer modeling
Agentic use
5 windows
  • 11 Tangential sub-goals
  • 12 Resource and access seeking
  • 13 Goal drift
  • 27 Oversight evasion
  • 28 Collusion and hidden signals

Any stage: 29 Anomaly, 30 Character drift during training, each reading against its own baseline. Version B's shaping phase can only use the pretraining column: honesty about what the model knows is the foundation the rest would be built on.