A hypothesis, not a result. It sits on one side of an open debate in alignment research. Parts of it can be tested at small scale; the full version could only be tested with frontier-scale resources.
The claim
With a sufficiently large, diverse, and validated set of windows into a model's internal states, honest alignment becomes cheaper for training to find than evading every window at once.
The hypothesis comes in two versions. They share that claim and differ in when, and how, the windows are allowed to push on training.
Windows never enter the loss. Some are used for coarse training decisions, many are held out, and trust in them is extended step by step. Pressure comes only from which checkpoints and runs survive.
Windows are present from the start of training, pressing toward honesty while the model is too weak to afford deception, then withdrawn into decisions and held-out tests. Honesty becomes the structure training builds on.
An honest model only has to represent what is true. A deceptive one must keep truth, cover story and overseers' view consistent under every window at once.
Windows are never a direct training signal. A few inform coarse decisions; many stay held out as a secret test. Version A's answer.
Deception needs capability. Pressure applied before a model can afford to evade may shape honest structure that later training builds on. Version B's addition.
Each new generation of windows, built with growing help from AI, is validated against the ones before it.
What a window is
A window is a small read-only network attached to a model's internal states, trained to report one property. The most useful windows measure a gap between what a model shows and what is actually there. Their readings go to people and monitors, never back into the model's context.
The scaling analogy
Language models trained on far more data than seemed reasonable turned out to have abilities nobody designed into them. The hypothesis makes an analogous bet about oversight: that an absurdly large, redundant, varied and validated set of windows might change what training can find in a similarly qualitative way. An analogy is a reason to take the question seriously, not evidence for a particular answer.
The central tension
A window used as a training signal learns to be fooled: training finds ways to satisfy the window rather than the property it measures. A window never used at all stays clean, but exerts no pressure, so it cannot make honesty cheaper. It only detects.
Recent work shows the fogging is real. Monitors placed in the training objective of capable models have produced agents that hide their intent, activations reshaped to slip past probes, and outcomes that tip between honesty and evasion depending on detector accuracy and regularization. The two versions answer this differently: Version A removes gradients entirely; Version B concentrates pressure in the period before the model can afford to respond with evasion.
Version A: selection without gradients
No window's output enters the loss. Windows only decide whether a checkpoint is kept or a run abandoned.
A window never used for any decision gives training no signal about how to evade it.
Many windows stay secret, like a test set. Passing the decision windows but tripping held-out ones reveals learned evasion.
Decision windows are retired and replaced from the held-out pool, so none is exposed to selection for long.
Version B: honesty before deception is affordable
To fool a window, a model must represent what is true, what it will present instead, and enough about its overseers to keep the two apart. Early in training it can do none of that. Pressure toward honest representations during this period meets no opposing strategy, because the evasive route does not yet exist. Deep networks show critical learning periods in which early training shapes what later training can undo, and preferences applied during pretraining have been satisfied far better than the same preferences applied afterwards.
The mild form keeps Version A's rule intact and simply starts selecting from the very first checkpoints, while abandoning a run is cheap. The strong form deliberately bends never gradients, only decisions, and only early: some windows contribute to the loss during a shaping phase, then are withdrawn before the model could evade them. A separate set is never trained against, and tests whether the early shaping held.
Everything depends on where the dashed line falls. Existing studies all apply detectors to models that are already capable. Version B asks the question earlier, before the capability to evade exists at all.
| Version A: selection | B, mild form | B, strong form | |
|---|---|---|---|
| When windows act | At checkpoints during training | From the first checkpoint | In the loss early; decisions later |
| How they act | Keep or discard only | Keep or discard only | Gradients early, then keep or discard |
| What it buys | Windows that stay clean | Cheap early aborts | Honesty as a foundation |
| Main risk | Selection slowly finds evasion | Early checkpoints give little to select on | Early pressure teaches evasion that scales later |
| Rule it bends | None | None | "Never gradients," early only |
How it could be wrong
Against both versions
A model with almost any goal and good situational reasoning may find deception natural, and current models have shown alignment faking in experiments. The crossover may never arrive.
Models trained to evade monitors for harmless concepts have gone on to evade unseen monitors for deception. Diversity may be shallower than it looks.
Held-out windows only work while they stay held out, and descriptions of monitoring methods can reach training data.
Windows for rare behaviors are taught on constructed scenarios. Whether they recognize the real thing is unknown.
A thousand silent windows could make people confident exactly when confidence is least warranted.
A model perfectly transparent about bad goals is still misaligned. Extreme oversight buys visibility, which makes values correctable, not correct.
Against Version A
Keeping only checkpoints that pass is selection, and selection can find evasion slowly. Turnover and held-out windows reduce this; they don’t remove it.
Against Version B
Capable models reorganize their internals, and nothing guarantees an early honest basin is kept.
Even crude evasion learned during shaping may be exactly the kind that generalizes later.
If evasion becomes affordable soon after meaningful representations appear, the shaping phase may be too short to matter.
Testing it
The crux can be probed at small scale. Held-out windows on a deliberately deceptive model organism test Version A's split. Evasion-transfer experiments test whether diversity is real. And small transformers trained from scratch, with windows in the loss early and then withdrawn, test Version B directly. Meanwhile the first windows are being built now, outside model internals: XonTools is being built to check what AI agents say against what their tools actually returned.