The Extreme Oversight Hypothesis

Enough windows might make honesty the cheapest path.

Build a large, diverse, validated set of read-only windows into a model's internal states. Past some scale, evading every window at once may cost training more than simply being honest. XonTools is the open research program testing that bet, starting with tools built to check what AI agents say against what actually happened.

ILLUSTRATIVE — the shape the hypothesis predicts, not data
honest alignment evading every window crossover? cost to training number, diversity and validation of windows →
Why this exists

AI systems are gaining capability faster than our ability to check them. XonTools exists to help close that gap by building instruments that show where an AI system's account of its work, or of itself, doesn't fit the evidence. Everything here is open: the code, the specifications, the results and the failures, free for anyone to use, test and improve.

What this is for →
Research program

From one window to many

Each stage builds a window that measures a gap: between what a system shows and what is actually there.

Full roadmap →
TESTED
Consistency engine

Claim and entity graphs that find contradictions no single sentence reveals.

SPECIFIED
Agent transcript monitor

Agent claims to be checked against hash-chained tool evidence, with a three-way verdict.

IN PROGRESS
Harder test corpora

Generated, cross-screened consistency corpora at any size (XonForge).

DESIGNED
Read-only heads

The strain head first; a catalog of thirty candidate windows into model internals.

LONGER TERM
Xon learning core

A settling architecture to work alongside standard neural networks.

First result

Contradictions that pairwise checking misses

Three statements can each look harmless and still be jointly impossible. On 180 test documents with planted cycles, the consistency engine found them where checking claims two at a time mostly did not. A strong language-model judge did about as well; the engine's case is structure and traceable reasons, not raw accuracy.

All results, including what failed →
0.92
Cycle F1, consistency engine
0.33
Cycle F1, pairwise checks only
0.93
Cycle F1, strong LLM judge, for comparison
Origins

Where the Xon came from

A Xon is the settling network architecture at the heart of this project. It grew out of several years of thinking about consciousness and global workspace theory, shaped by work in electrophysiology signal analysis: many rhythms, settling into one coherent whole.

Read the origins

Get involved

Collaborate

Working on oversight, interpretability, or agent evaluation? I'd like to compare notes and test ideas together.

Get in touch →

Support the work

Grant applications are under way. If you fund independent safety research, get in touch.

Contribute

XonTools is open source under Apache 2.0, with every result and pre-registration in the repository.

View on GitHub →