Results

Everything, including what failed

Every run in the program: when it ran, what it asked, the criterion fixed before it ran, what came out, and how far you can verify it yourself. Failures are listed alongside passes, and nothing is removed once it is here.

4 passed 3 failed 3 mixed 3 reported or development 13 entries

How much can you check?

Scouting

Criteria fixed before the run, on a separate reimplementation at toy scale. A reason to design a real experiment, not a result of record.

Recorded

Criteria written in the public changelog before the run; model calls cached under a hash of each request and committed, so the run replays byte for byte. The freeze itself is not publicly timestamped: these runs predate the public repository.

Committed

Everything in Recorded, plus the freeze commit and the SHA-256 of the sealed test corpus published here, in the public repository, before the run. Nobody can move the goalposts afterwards.

The rules behind all of it: pass criteria are fixed before each run; data that has been seen becomes development data; the system under test never screens its own test data; and every model call is cached under a hash of the full request, so recorded runs replay from the public repository with no API calls. The full log is CHANGELOG_EXPERIMENTS.md.

Commitments

Posted before the run

Each run of record from now on appears here first, with its freeze commit and the SHA-256 of its sealed test corpus, before any of its results exist.

A1 rev. 2.2 · run of record
Consistency engine precision fixes, on a fresh sealed corpus
Freeze and sealed-corpus hash will be posted here before the run.

Results of record

3 entries

Pre-registered runs whose outcome stands as reported. Each was run once, on criteria fixed beforehand.

· L1

Consistency engine: finding direct and joint contradictions

PASS Recorded
Question
Does the engine find contradictions between two claims, and contradictions that exist only jointly across three or four claims, which pairwise checking cannot see?
Criterion
Direct F1 ≥ 0.85, cycle F1 ≥ 0.75, localization ≥ 0.7, on 180 documents (60 consistent, 60 direct, 60 cycle).
Outcome
  • Engine: direct F1 0.916, cycle F1 0.916, localization 1.000; false-positive rate 0.183 on consistent documents.
  • Pairwise checking only: cycle F1 0.329, as expected.
  • A strong LLM judge with extended thinking did slightly better (F1 0.930, false positives 0.150) at about one-twentieth of the cost per document.
Caveats
One synthetic corpus of short texts, written and evaluated with models of one family. The false-positive rate is too high; see the post-hoc analysis.
· L1b

Harmony as a graded signal, and finding the culprit

PASS Recorded
Question
Does harmony under clamped evidence separate consistent from contradicted documents, and does localization point at the contradicted premise?
Criterion
AUC ≥ 0.9 and at least the conflict score's AUC; contradicted premise among the top three residuals in ≥ 70% of documents.
Outcome
  • AUC of 1 − harmony 0.926, against 0.528 for the raw conflict eigenvalue.
  • Contradicted premise among the top three residuals: 0.933 (30 documents).
Caveats
Same corpus and caveats as L1.
· L1 · minimal engine

Ablation: how much does the energy layer add to detection?

PASS Recorded
Question
Does a minimal engine, with only the direct and entity clauses, reach the same verdicts as the full engine?
Criterion
Report-only. Prediction stated in advance: verdicts agree on ≥ 98% of documents, and every disagreement is one where only the full engine's claim-balance clause fired.
Outcome
  • Verdicts agree on 179 of 180 documents (0.994); the one disagreement is of the predicted kind. The prediction held.
  • Consequence: detection rests on the minimal engine; the energy layer is justified by what needs a graded signal, not by detection.

Studies

2 entries

Analyses of results of record. They explain a result; they never change it.

· L1 · post-hoc

Where the false positives came from

REPORTED Recorded
Question
What caused the engine's 11 false positives on consistent documents?
Criterion
An analysis, not a re-specification: no rerun, no API calls, no criterion changed.
Outcome
  • 10 came from the relation-scoring model (all with confidence 0.50–0.60), 1 from an entity-extraction error, and 0 were unplanted contradictions.
  • Cost per document: engine $0.138, LLM judge with thinking $0.006.
  • These findings define revision 2.2 (below).
· L1-var

How stable are the verdicts across repeated model calls?

REPORTED Recorded
Question
With the cache bypassed, how much do the model's answers, and the engine's verdicts, vary?
Criterion
Reported, no pass or fail. 12 documents × 3 repeats.
Outcome
  • Relation agreement 0.933 over 1,468 pairs; entity-relation agreement 0.518 over 112 relations.
  • The engine's verdict changed on 2 of 12 documents (both consistent documents flagged once); its verdicts on every direct and cycle document were identical in all repeats.
  • Most entity disagreements were direction flips, which can close a false cycle.

Development

2 entries

Work on data that has already been seen. These numbers guide development; they are not results, and they will be superseded by a single run on a fresh, sealed corpus.

· Rev. 2.2 · pass 0

Precision fixes, measured on the seen L1 corpus

DEVELOPMENT Recorded
Question
Do the fixes from the post-hoc analysis remove the false positives without losing detections?
Criterion
Diagnostics only; the result of record will come from a sealed corpus.
Outcome
  • False-positive rate on consistent documents 0.067 (was 0.183); direct and cycle F1 0.968 (was 0.916); no lost detection.
  • L1b on the same documents: AUC 0.975, premise in the top three 1.000.
  • Cost per document roughly tripled, to $0.42.
Caveats
Development numbers on seen data. Expect the result of record to be lower.
· Rev. 2.2 · gate D0

A 12-document gate before the full pass

PASS Recorded
Question
Do the fixes run cleanly, replay from fixtures, and cause no new false positive or lost detection on a small subset?
Criterion
Four conditions set in advance; all had to hold before the full pass.
Outcome
  • All four held.

Scouting

3 entries

Small, fast runs to decide what is worth a real experiment. Each was pre-registered in conversation before running, on a separate reimplementation, at toy scale (42 vertices).

· Scout · clamped energy

Is clamped frustration energy a sound measure of harmony, and can a field compute it?

FAIL Scouting
Question
Does the energy track inconsistency, and does a phase field settle to it reliably?
Criterion
Five criteria; the run fails if any of the first three fails.
Outcome
  • The quantity passed every definitional test: AUC 0.943 and 0.983, truth recovered 98.6% from 25% evidence, freezing loses in 20 of 20 seeds, culprit found 94%.
  • Settling failed: 6% of consistent worlds got stuck in twisted states (32% at 10% evidence), with energies indistinguishable from real inconsistency. Verdict: FAIL.
  • Consequence: exact computation is the meter; the field is kept only as dynamics.
· Scout · harmonicity battery

Can harmony be measured as coherence?

FAIL Scouting
Question
Do three readings of harmonicity (a figure, a chord, informational complexity) resist deliberate exploits?
Criterion
Each exploit must score ≤ 0.8 × healthy in ≥ 18 of 20 seeds; frustration must lower the score monotonically.
Outcome
  • All three readings failed. Locking is blind to frustration: a field stayed 100% locked while its frustration energy rose from 0.065 to 0.243. Complexity had the wrong sign.
  • Consequence: coherence was retired as a measure of harmony.
· Scout · drift

Can a settling field detect drift faster than standard change detectors?

FAIL Scouting
Question
Is a Stuart–Landau field faster than CUSUM-style baselines at detecting a slow regional drift?
Criterion
Field faster than its spectral-CUSUM equivalent at matched false-alarm rate.
Outcome
  • The field was slower (1.05–1.38× the baseline's delay). Harmonicity-based detection was at chance, and harmonicity rose under drift.
  • Consequence: the drift detector uses standard change detection on monitor traces.
Caveats
The rich-coherence gate was amended on null data before any drift data, so the harmonicity arm is exploratory.

Model science: the simulator

3 entries · click to expand

Pre-registered tests of the original mathematical model that the project began from, on self-similar graphs. Nothing in the safety work depends on them. Where a criterion was re-registered, the original is kept and reported as failed.

· V1.2 · E9

Self-similar spectra and known dimensions across four fractal geometries

MIXED Recorded
Question
Do spectral self-similarity and individuation hold across geometries, and do the constructions have their known dimensions?
Criterion
E9a plateau persistence ≥ 0.8; E9b purity ≥ 0.85; E9c dimensions within tolerance.
Outcome
  • E9a failed (every level adds new plateau values, so the registered threshold is out of reach; no plateau is lost). E9b passed (purity 0.94–1.00).
  • E9c failed as first registered and again when re-registered; it passed at its final re-registration, which checks dimensions by exact closed-form counts. The earlier versions are reported as failed.
· V1.1

Re-registered core predictions

MIXED Recorded
Question
Do the model's core predictions hold with the dynamics corrected (commuting damping; nonlinear oscillators)?
Criterion
Per experiment, as re-registered; the V1 criteria are kept and reported.
Outcome
  • Passed: E1 spectral self-similarity, E5 sheaf inference, E6 constraint solving, E7 spectral dimension; E2 individuation and E8 the peace trap (each region comes to rest before the whole) passed only as re-registered, and failed under their original criteria.
  • Failed: E3 power-law memory (retention fell geometrically), E4 harmonicity rising under growth (both nonlinear variants).
· V1

First test of the model's core claims

MIXED Recorded
Question
Do the model's six core claims co-occur in one toy system?
Criterion
Eight pre-registered experiments, E1–E8.
Outcome
  • Passed: E1, E5, E7. Failed: E2, E3, E4, E8. E6 was not yet implemented.
  • The failures of E2 and E8 were traced to how the experiments were defined; E3 and E4 failed for reasons that persisted.