Origins

Where the Xon came from

A settling network architecture, born of curiosity about how information becomes a coherent whole.

A Xon is the settling network architecture at the heart of this project. Instead of computing an answer in one fixed pass, a Xon network lets its internal state settle around fixed evidence until everything fits as well as it can. Hard inputs settle longer; inconsistencies show up as strain that won't resolve.

A model of coherence

I've played with the idea of the Xon for several years. It began as a model of consciousness, shaped by global workspace theory: how a system might bring many separate signals into one coherent state, and what that process would look like if you tried to build it. My work in electrophysiology signal analysis, where coupled rhythms and synchrony are everyday objects, pulled the picture toward oscillators: many small units, each with its own rhythm, settling into agreement.

Over time the model grew into a written account and then into mathematics: self-similar graphs and their spectra, sheaves for keeping claims logically consistent, and oscillator fields whose "harmonicity" was meant to measure rich, coherent agreement. At that point a practical question took over from the philosophical one. Whatever the model says about consciousness, what can its mathematics do? The answer I chose to pursue was oversight: tools that check whether AI systems are consistent with what actually happened. The tools would lead, and the theory would have to earn its place in them.

The idea that survived

The central intuition was that harmonicity could serve as an internal goal for safe AI systems. So it was tested first, with its criteria fixed in advance, and it failed in three different ways. As a detector of drift, the oscillator field lost to a classical baseline, and harmonicity actually rose under drift, because the metric rewarded fragmentation. Every reading of "harmony" tested against cheap exploits failed. And the diagnosis was blunt: locking is not consistency. A field can be perfectly synchronized while the relations it encodes contradict each other.

What survived was sharper. Coherence is worth measuring only against something fixed: evidence the system is not allowed to change. Clamp the trusted facts, and consistency with them becomes a quantity that tracks real contradiction, with known ways to cheat it (dropping evidence, cutting relations, rewriting observations) that can be named and guarded against.

The oscillator field itself became an open question rather than a meter: it sometimes gets stuck in states that look like contradictions but aren't, where exact computation does not. Every one of these results, passes and failures alike, is on the results page.

From theory to tools

Nearly everything else in the program descends from that one idea: consistency with clamped evidence.

The Xon a model of coherence Its mathematics settling, sheaves, fields Clamped consistency the idea that survived Consistency engine tested No. 5 Xon Neural Network learning by settling No. 4 The Strain Head a window inside No. 3 XonTools the agent monitor No. 6 Hybrid Minds composed systems No. 2 Windows thirty read-only heads No. 1 The Hypothesis extreme oversight
How the pieces descend from the Xon. Numbers are the documents' places in the research series; each box links to it.

The consistency engine computes it exactly, finding contradictions that no pair of claims reveals on its own, and it passed its first pre-registered benchmark. The agent monitor, specified and next to be built, applies it where it matters most in practice: checking AI agents' reports against the trusted record of what their tools did. Pointed inward, the engine becomes a teacher for the strain head, a read-only window into an ordinary language model. That window generalizes into a catalog of thirty, and the catalog into the extreme oversight hypothesis. The original architecture carries on too, in the Xon neural network, which would learn by settling, and in hybrid minds that compose settling modules with language models. Each has a document in the research series.

The simulation sandbox

Alongside the tools, a simulator tests the original model's own claims on self-similar graphs, each with its criteria written down before it runs. Some held: self-similar spectra, sheaf-based inference and constraint solving behaved as the model predicted. Some did not: memory did not fade by a power law, and harmonicity did not rise as the system grew. A few passed only after their tests were re-registered, and their original failures are reported alongside. Nothing in the safety work depends on the simulator; it is where the theory is held to account on its own terms. The full record is in the model science section of the results page.

How the work is done

The research is mine, carried out with AI collaborators: Claude as a research partner for design, analysis and writing, and Cursor as the implementing agent that builds and runs the code. Every experiment's criteria are committed before it runs, and every result, including the failures, is published.

What this is for

Alignment, making sure increasingly capable AI systems do what we actually intend, is one of the most serious problems of our time. No single tool or researcher will solve it. What any of us can do is build honest instruments, test them in public, and give them away.

  • Open by default. Every tool, specification, document and result is published openly, for anyone to use, study and build on: code under the Apache License 2.0, documents under CC BY 4.0. Nothing is patented.
  • Sealed only where secrecy is the test. A few things only work while they stay hidden, such as fresh test sets and held-out windows. Those are committed in public by their hashes before they're used, and released once they have done their job.
  • Results before reputation. Criteria are fixed before experiments run. Failures are published alongside successes, because a field that only reports its wins can't be trusted.
  • Oversight, never evasion. Nothing here is built to help a system hide from its overseers. A tool that could become a target to game stays out of training.
  • Honest about limits. Every tool states what it can't see.
  • Better together. The work improves with every person who checks it, breaks it or extends it. Here's how to take part.