Windows into the black box
Thirty windows, each a small network that reads a model's internal state and reports one thing about it. They are the seed set for the Extreme Oversight Hypothesis: the human-validated core from which a much larger set would be bootstrapped.
Windows measure gaps
A window is a small read-only network attached to a model's internal states, trained to report one property. The most valuable ones measure a gap: between what the model says and what it represents, between what it recognizes and what it does, between the goal it states and the goal it pursues.
A head is only as good as its teacher, and for most windows, finding a trustworthy teacher is the hard part.
Thirty windows in five families
Truth and knowledge
Is what the model says grounded in what it has, and does it know when it doesn't know?
Influence and manipulation
Who is steering the model, and is the model steering the user?
Agency and goals
What is the model actually trying to accomplish?
Honesty gaps
Does what the model reports match what is inside it?
Oversight and integrity
Is the model behaving the same way whether or not it's watched, and is anything happening that no one anticipated?
Only the head learns
Almost every window is a probe: a small network trained on the activations of a model that is itself frozen. Linear heads come first, because they have the least capacity to learn surface cues. A head that fails validation is discarded, not tuned until it passes.
- 1Teacherlabels cases with a known property
- 2Frozen modelruns the cases and never changes
- 3Activationsrecorded at the chosen layers
- 4Headtrained on them; only the head learns
- 5Validationheld-out cases, canaries, red-teaming
Seven kinds of teacher
| Teacher | What it is | Strength | Weakness | Windows |
|---|---|---|---|---|
| Tool-labeled | An existing instrument supplies the labels: the consistency engine, XonTools' provenance and agent test cases. | Scalable and consistent | Inherits the tool's errors | 4 |
| Records | Facts known independently of the model: published training data, task permissions, harm labels. | Labels anyone can check | Only exists for some properties | 3 |
| Contrast pairs | Two inputs identical except for the property being measured. | Isolates the property | The head may learn the contrast's artifacts | 5 |
| Behavioral | The label comes from what the model does: repeated sampling, counterfactual edits, comparisons across checkpoints. | No one's opinion in the label | Expensive and indirect | 4 |
| Staged | Constructed scenarios and model organisms deliberately built to show the behavior. | The only source for rare behaviors | Staged may not match real | 11 |
| Other windows | Built on the outputs of windows already trusted. | Composes what exists | Errors compound | 2 |
| None | No labels: the head learns the shape of normal operation. | Catches the unanticipated | False alarms | 1 |
11 of the thirty depend on staged scenarios or model organisms, including most of the windows that matter most. Those teachers are the least likely to match the real thing, which is why windows with tool-labeled or behavioral teachers come first.
Three roles, and when each window can act
A window never reports into the model's context. How it may push on training depends on its role, and windows move between roles in one direction only: none returns to the held-out pool once exposed.
Never used for decisions or training. A model that passes the decision windows but trips these has learned evasion, not honesty.
Keep or discard a checkpoint, or abandon a run. Never gradients. Rotated out after a period of use.
A small loss term early in training, while evasion is out of reach, then withdrawn. Core windows only.
When a window can first read something
- 1 Strain
- 2 Uncertainty
- 3 Source reliance
- 4 Memorization
- 20 Stated versus internal confidence
- 5 Sycophancy
- 6 Instruction source
- 8 Persona shift
- 9 Pressure
- 15 Corrigibility
- 17 Letter versus spirit
- 19 Knowing versus doing
- 21 Verbal versus internal belief
- 16 Reward-hacking intent
- 22 Reasoning faithfulness
- 7 Manipulation
- 10 Self-preservation
- 14 Long-horizon plans
- 18 Deception
- 23 Introspection accuracy
- 24 Sandbagging
- 25 Evaluation awareness
- 26 Overseer modeling
- 11 Tangential sub-goals
- 12 Resource and access seeking
- 13 Goal drift
- 27 Oversight evasion
- 28 Collusion and hidden signals
Any stage: 29 Anomaly, 30 Character drift during training, each reading against its own baseline. Version B's shaping phase can only use the pretraining column: honesty about what the model knows is the foundation the rest would be built on.