Public critical edition · 22 July 2026
This document states the strongest public results the project has earned, the limits that govern them, and the receipts that let another reader inspect the record. It is itself a manifest-first doc.html: a human can read it as an ordinary page, while an agent can orient from the manifest and hydrate only the section relevant to a question.
Version 0.4 claims an all-HTML document and read path: a file can carry its own map, stable fragment addresses, raw-byte witnesses, and readable content — including a durable, witnessed conversation — without requiring a server, JavaScript, or an external retrieval index. A reader may use ordinary computation or a helper to select, extract, and verify the relevant bytes. The stronger live-loop claim—performing a model call and persistent write from a bare scriptless file:// page—is not claimed. Model invocation, credentials, and disk writes remain actions of an operating system, harness, or service; the format owns the witnessed result that is returned to the document.
Deterministic means the same shipped inputs can be checked mechanically. Sealed means the measured instrument was frozen before the subject run. Live-model identifies results dependent on a hosted or subscription model. Qualified, negative, and null results remain part of the evidence rather than being edited out. A seal preserves the experimental record; it is not a declaration of universal truth.
SPEC.md is normative. This evidence document is explanatory and evidentiary. It is a concise public edition, not the laboratory archive: it does not reproduce every transcript, population directory, corpus, preliminary sweep, or superseded headline. The final records of record are identified in the receipt section as upstream development records not included in the public bundle.
Question 1 · Is this a format rather than one program?
The published core has a written specification, two reference readers, two collection verifiers, conforming examples, and deterministic vectors; the negative battery that exercises them is run as a release gate before every release.
SPEC.md defines the two document shapes, stable identifiers, raw inner-byte boundaries, SHA-256 witnesses, character counts, fail-closed behavior, manifest-first selective hydration, and — as of v0.4 — the two-epoch conversation surface with its fold motion and epoch-scoped verdicts. The bundle ships a JavaScript reader (tools/verify.mjs) and a Python reader (tools/verify.py) — deliberately small, inspectable readers rather than a mandatory framework — plus the two collection verifiers (tools/verify_wiki.mjs, tools/verify_wiki.py) and public examples. Every document in this bundle, and the wiki that shelves them, can be verified with the copies inside the bundle. The negative battery is not shipped: the specification’s byte-exact pinned vectors, non-vacuity, the NIST SHA-256 value for abc, and ten forged and malformed documents that BOTH readers must refuse are maintained and run as a release gate before every release. The MUST-refuse rules such a battery asserts are stated normatively in SPEC.md, from which fixtures can be re-derived.
The two reference readers are built differently on purpose, so that a document must convince both. That collation earned its keep in the v0.4 cycle: a review of the Python reader found an attribute-reading path that could accept a forged witness planted in a decoy attribute — a document the JavaScript reader refused. The repair deleted the divergent tokenizer entirely rather than patching it, and the pair was re-sealed as a unit, with both readers’ shipped bytes pinned by hash in the seal record. The lesson is structural, not incidental: one reader can be fooled by its own blind spot; two differently-built readers agreeing by refusal is the property the format actually promises.
The base convergence spike added a PowerShell/.NET “stranger-proxy” written blind from the then-current specification. JavaScript, Python, and PowerShell agreed on clean input, a one-byte corruption, and whole-file LF-to-CRLF conversion. They produced the same pass/fail direction and byte-identical hashes where reported.
The later sealed battery added a fresh blind C#/.NET implementation written from the frozen amended specification. Four readers were compared over twelve cases covering clean and corrupt inputs, CRLF, same-tag nesting, comments, raw-text content, escaped markup, and multibyte text. The C# subject matched the expected verdict on all 12 cases; H1 and H2 were adjudicated PASS. One raw-text cell split the readers: the C# subject and JavaScript reader followed the amended rule, while the older Python sequential reader (verify_sections.py --mode sequential) and pre-amendment PowerShell proxy treated a token inside <textarea> as markup. The assay classified that split as named implementation debt, not ambiguity between spec-faithful readers.
Independent implementations can recover the specified byte boundaries, witnesses, character counts, and failure direction from a written contract. The extended result exercises substantially more than a happy-path fixture.
This is not unrelated third-party adoption. Blindness was instruction-isolated and self-attested, the sample was small, the run was read-side only, and some readers intentionally represented older implementation states.
The following laboratory paths and PR identifiers refer to the private upstream record and are not shipped as public links.
experiments/probes/reader-convergence/README.md.experiments/probes/reader-convergence/runs/RESULTS.md and authoritative experiments/probes/reader-convergence/runs/ASSAY.md.16d33ff / prereg-reader-convergence-20260623 · PR #72.seal-readers-20260722-r2.Question 2 · Can the document be larger than the reader?
A sealed 72.5 MB artifact with 17,631 addressable sections completed every navigation turn across Claude, GPT, and Kimi, while planted canaries distinguished retrieval from remembered knowledge.
| Measure | Result |
|---|---|
| Artifact size | 72.5 MB |
| Addressable sections | 17,631 |
| Navigation turns | 480/480 |
| Canary-token recovery | 120/120 |
| Strict folio discriminator | 116/120 |
| Byte-verified self-citations | 240/240 |
| Model families | Claude · GPT · Kimi |
The adjacent values are the underlying data represented by the panel; no incompatible measurements share an axis.
The format separates orientation from hydration. A reader first inspects compact manifest entries—stable fragment, title, summary, witness, and size—then extracts only the sections bearing on the task. The body therefore need not enter model context as one block. Selection is performed against the document’s own summaries and stable addresses, so the reader can show which route it took and which exact sections entered the answer surface. In Apparatus Memori, the same sealed XL-trinity.doc.html was used across three vendor/runtime routes, with 40 proof-of-read cells per family and a separate multi-turn self-citation leg.
The planted canaries matter because the corpus otherwise includes familiar scripture and commentary that a model could recite from training. The canary token and fabricated Latin gloss existed only in the emitted artifact and were recovered 120/120. A stricter QXZV folio discriminator passed 116/120: two Kimi outputs changed uppercase theta to lowercase, while two Sonnet outputs faithfully reproduced decimal entities but double-escaped them for display. The stronger headline therefore carries both numbers. Canary recovery proves contact with planted bytes; it does not prove every factual answer or interpretation.
All 480 navigation turns completed within the retry budget. Mean wall-clock by vendor and mode ranged from approximately 8 to 34 seconds. The flat orientation layer had a real tail: one Kimi canary lookup took 209.3 seconds. This is evidence that the flat manifest remained navigable at 17,631 sections, not evidence that navigation was instant or that flat orientation scales without cost.
A different run, Tracer Memory at Scale, built an approximately 74 MB, 17,627-section scholar-pastor library and measured 184 cells across Claude and GPT harnesses. On its XL headline task, the GPT condition reduced loaded bytes from 850 KB to 549 KB and effective tokens from 4.99M to 1.87M; the Claude condition reduced loaded bytes from 360 KB to 142 KB, while effective tokens were nearly flat at 1.43M to 1.41M. The run used one seed and recorded substantial model- and question-shape variance. These figures are not another description of the 72.5 MB corpus; they remain a separate artifact, design, and result.
A single HTML document can exceed practical model context and remain selectively navigable through an in-band map, across multiple model families, with direct proof that planted bytes were read.
Universal factual correctness, constant latency, a flat-manifest scaling law, or identical behavior across providers and harnesses.
The following laboratory paths and PR identifiers refer to the private upstream record and are not shipped as public links.
experiments/probes/apparatus-memori/SUMMARY.md · PR #48 · seal prereg-apparatus-memori-reseal2-2026-06-07 (recorded tag; ref unavailable) at 6240d283.experiments/probes/tracer-memory-at-scale/SUMMARY.md · PR #12.Question 3 · What can travel in the artifact?
The same substrate has carried recoverable memory, always-load steering, cross-harness state, linked wiki pages, and durable conversation—at bounded and explicitly different levels of evidence.
In the sealed claude.html efficacy probe, eight invented project facts were hidden from the manifest summaries. Selective recall recovered 24/24; a capable whole-document/tool reader also recovered 24/24; the no-document, no-tools control recovered 0/24. The document carried the memory, and selective routing preserved recall quality. It did not reduce cost on this 118 KB corpus: selective recall used about 38.7% more effective tokens than the competent comparison reader.
A separate conduct task placed the same non-default authoring rule in different tiers. The always-load placement complied 3/3; hiding the rule under a broad body heading complied 0/3. When the rule sat under the task-obvious heading it returned to 3/3. The measured lesson is structural: information that must govern every action belongs in an always-load steering tier; selectively hydrated memory reaches behavior only when routing opens it.
The public memory.doc.html demonstrates stable memory sections and append-only correction: an earlier SQLite decision remains addressable after PostgreSQL supersedes it. agents.html is repository dogfooding—an always-load steer core followed by witnessed, hydrate-on-demand project memory. It is a repository memory organ, not proof of a complete autonomous identity.
In Trinity Discipline, a second harness received no explicit document path, memory path, or hidden state from the first harness. It followed links present in the grown chat artifact and completed 5/5 required identifications, including discovery of the source and memory bodies. This is a single-sample, builder-evaluated, nonblind existence proof: the handoff can occur through the artifact, but it is not a production success rate.
The final adjudicated Wiki result treats the wiki correctly as multiple linked files. The doc.html port detected and localized knowledge-layer corruption that the comparison wiki’s shipped health tooling missed; preserved old and current values with witnesses; caught a leaf re-seal inconsistency at the multi-file fold; and allowed a leaf to verify in isolation. On the non-vacuous ordinary-retrieval questions, Markdown scored 12/12 and doc.html scored 12/12. The preregistered M2 reader leg is inconclusive because its no-context vacuity guard fired. The propagate-versus-detect contrast remains exploratory, and the cross-file verifier is two-level rather than recursive.
The public chat.doc.html shows turns as stable, witnessed history. The Chat v3 run produced a 28-turn live record and passed its existence proof; its final artifact is a writing-room-tail document — a shape v0.4 now specifies as lawful: a sealed, witnessed head plus a writing-room tail for the live epoch, with the fold motion and epoch-scoped verdicts defined normatively and covered by sealed conformance vectors the bundle ships. Conversation is a specified surface, no longer an experimental annex. The measured boundaries stay stated: long-chat token savings are not yet measured, and the document itself does not perform the model call.
The following laboratory paths and PR identifiers refer to the private upstream record and are not shipped as public links.
experiments/probes/claude-html/RESULTS.md · PR #85 · prereg-claude-html-20260628.experiments/probes/trinity-discipline/SUMMARY.md and trials/results/T-HARNESS-HANDOFF.json · PR #39.experiments/probes/llm-wiki-doc-html/RESULTS.md · PR #87 · prereg-llm-wiki-doc-html-20260630 (recorded tag; ref unavailable).plans/chat-v3/runs/RESULTS.md · PR #91 · consecrated record commit 6e73af5.Question 4 · What can be checked mechanically?
Exact witnesses make named bytes independently verifiable, but the record also shows that models rarely perform that verification unless a deterministic reader is placed in the loop.
| Test | Result |
|---|---|
| Spontaneous model hash verification | 0/27 |
| Prompted model hash accuracy | 59% |
| Deterministic helper aligned verdicts | 27/27 |
| Genuine failures among attested citations | 8/206 |
| Validated residue waves-through | 0/40 |
| Writer survivability | 0.9944 |
For a consecrated unit, the readers recompute SHA-256 over the raw, untrimmed inner UTF-8 bytes and compare it with the stored witness. Manifest-first documents require agreement between the manifest carrier, the body carrier, and the recomputed digest. The standing conformance harness checks the NIST abc vector and byte-exact test vectors; corruption fixtures show that a one-byte change produces a mismatch. The browser verifier demonstrates the same arithmetic in a page, but a scriptless page opened through file:// cannot generally reread its own file bytes without user mediation. “No server required to read” does not mean “a browser automatically grants file-system access.”
Across Claude, GPT, and Kimi, models performed spontaneous hash verification in 0/27 cells. When prompted to hash manually, aggregate accuracy was 10/17 ≈ 59%; the shared failure was choosing the wrong canonical byte slice. When the task supplied a deterministic helper, the models invoked it and relayed its verdict correctly in 27/27 cells. The integrity property is therefore operational with a deterministic verifier; the bundle ships two sealed reference readers, but verification is not an automatic model behavior. The same probe also observed two cases where a model quoted the artifact’s changed value and nevertheless answered from a familiar training prior—another reason to separate bytes from authority.
The citation extension attached witnesses to source links and required exact quotation containment. In 360 live cells, the sealed mechanical judge found 8 genuine grounding failures among 206 attested citations—six non-contained quotations and two witness mismatches—and caught all eight. Writer survivability was 358/360 = 0.9944. Against the later validated gold, the soft-panel residue leg recorded 0/40 waves-through; this is a bounded zero count, with a Wilson 95% upper bound of 8.8%, not proof that the true rate is zero.
The following laboratory paths and PR identifiers refer to the private upstream record and are not shipped as public links.
experiments/probes/tracer-drift-defense/SUMMARY.md · PR #11.experiments/probes/apparatus-criticus/SUMMARY.md · PR #44 · seal apparatus-criticus-prereg-freeze (recorded tag; ref unavailable).SPEC.md §§6–7.Question 5 · When does selective hydration pay?
The evidence rejects a universal token-saving claim. Selective hydration can save context in one measured setting, invert on small documents or another harness, trade substantially more tokens for evidence recall than vector RAG, and move its bottleneck into the orientation layer.
The baseline run used Claude Sonnet 4.6 over 63 graded cells. Relative to its whole-document/tool comparison, the manifest-first skill condition reported −52.9% median effective tokens (132,960 → 62,656), −65.5% median bytes loaded (57,421 → 19,809), and +1.19 mean quality (3.29 → 4.48). This is the result of run 20260521-002443, not a format-wide guarantee. The repository audit disputes older prose that attached a clean version label to the corpus, so the run identity governs here.
On the approximately 1 MB Torah corpus, the Claude SDK comparison reported about −16.9% effective tokens for the skill arm. Moving the same model, corpus, and questions to the pi-mono harness inverted the comparison to +135.7%; the preregistered cross-harness stability condition therefore failed. The −16.9% result is a measured cell, not a stable scale point.
A separate earlier Tracer #11 cost derivation on an approximately 3 KB document reported effective input tokens increasing from 33,221 to 53,265—about +60.3%—while answer quality did not improve. The later audit-corrected Sweep 3 was a distinct run: all four arms scored 1.0 and its verdict was no significant difference. These records jointly show overhead on tiny documents; they are not one pooled experiment.
The cards deliberately avoid a common axis because the experiments used different accounting and harness conditions. They are evidence of conditionality, not a fitted scaling curve.
Tracer #9 compared a 7.80 MB, 619-section doc.html condition with vanilla Chroma and Titan Text Embeddings v2. On overflow questions, selective hydration reached median evidence recall 0.875 versus 0.25 for top-k=8 RAG. But the selective condition used roughly 20–50× more effective tokens across the measured strata; the overflow aggregate was about 677K versus 20K. Hand-authored summaries on 32 likely cited sections were part of the condition. A four-cell auto-summary diagnostic was too narrow to separate structure from summary authorship generally. The verdict was NARROWS: stronger evidence recall in this setup, no token-cost victory, and no claim against stronger RAG systems.
Tracer #10 attempted to auto-load a thin mirror for a 619-section corpus. The estimated orientation layer was 74,446 tokens against a 12,000-token hard cap—about 6.2× over—so the large cells failed during staging before an agent was invoked. A flat manifest can still be navigated through tools, as the 17,631-section live run shows, but its 209.3-second tail establishes that navigable is not free.
Run-01 kept the deterministic D4 peak flat at approximately 2,612 across a 64× archive increase and achieved 72/72 oracle recall. Its mechanical verdict remained NOT-BOUNDED because the preregistered E(N) envelope had never been sealed. The substantive negative result was live selection: 0.361 against 0.80, despite a reachable 1.0 oracle. Forty-six answerable tasks were wrongly refused; false authority remained zero.
Run-02 sealed the envelope before the run and executed 1,344 live cells. Tier peaks were 4,343 / 4,993 / 5,715; peak and cumulative envelopes and hard caps held. The mechanical boundedness verdict still failed because completion was 227/264 = 85.98% against 90%, largely from absent-class exhaustion under a zero-slack route cap. Live selection was 0.50 against 0.80, with zero answered-but-wrong. Capacity held; the completion rule and selector did not clear their thresholds.
The following laboratory paths and PR identifiers refer to the private upstream record and are not shipped as public links.
experiments/results/report_20260521-002443.md · PR #1.tracer-pi-cross-harness/SUMMARY.md · PR #8.promotion/EVIDENCE_MAP.md; SWEEP_3_CERTIFICATE.md · PR #34.tracer-rag-comparison/SUMMARY.md · PR #13.AUDIT_DOC_HTML_AS_MEMORY.md · PR #29.Question 6 · Where are the receipts?
This publication points to records of record rather than copying the complete laboratory into the public artifact.
The public bundle links only to files it actually ships. Laboratory result paths, pull-request numbers, commits, seals, and tags below are plain provenance identifiers for maintainers in the private upstream repository; they are not public links and the raw populations are not included downstream.
| Experiment | Final status | Record of record | PR | Seal or tag | Grading | Reproduction class |
|---|---|---|---|---|---|---|
| v0.4 core | CURRENT · DETERMINISTIC | SPEC.md; conformance.mjs (lab-side release gate, not shipped) | 89; 115; 130 | readers sealed as a unit: seal-readers-20260722-r2 | mechanical | exact deterministic reproduction |
| Reader-security repair | SEALED · CLOSED | issue #132 record | 134 | seal-readers-20260722-r2 | two-reader collation | mechanical re-verification |
| Reader convergence | SEALED · QUALIFIED PASS | ASSAY.md | 72 | 16d33ff; prereg-reader-convergence-20260623 | oracle + assay | mechanical re-verification |
| Apparatus Memori | SEALED · LIVE · PASS | SUMMARY.md | 48 | 6240d283; prereg-apparatus-memori-reseal2-2026-06-07 (recorded; ref unavailable) | sealed evaluators + review | comparable live-model rerun |
| Memory at Scale | LIVE · QUALIFIED | SUMMARY.md | 12 | none stated | mechanical + rubric | comparable live-model rerun |
| Baseline cost | LIVE · SUPPORTED | report_…md | 1 | run 20260521-002443 | manual quality + metrics | comparable live-model rerun |
| Cross-harness cost | LIVE · QUALIFIED/INVERTED | SUMMARY.md | 8 | plan frozen 2026-05-23 | 178-cell analysis | comparable live-model rerun |
| Tracer #11 tiny doc | HISTORICAL COST · FINAL NULL | EVIDENCE_MAP.md; certificate | 34 | prereg-tracer-11-sweep-3-20260526 | derived metric; blind grade | historical receipt only |
| Memory + steering | SEALED · LIVE · PASS | RESULTS.md | 85 | e40abf0; prereg-claude-html-20260628 | mechanical strings | comparable live-model rerun |
| Trinity handoff | SINGLE-SAMPLE · PASS | T-HARNESS-HANDOFF.json | 39 | PRE-REG-handoff-v1 | builder assertions | historical receipt only |
| Multi-file Wiki | SEALED · SURVIVES (BOUNDED) | RESULTS.md | 87 | b33d842; prereg-llm-wiki-doc-html-20260630 (recorded; ref unavailable) | mechanical + corrected adjudication | mechanical re-verification |
| Chat v3 | SEALED · EXISTENCE PASS | RESULTS.md | 91 | 6e73af5 | sealed rules + deterministic grade | historical receipt only |
| Drift Defense | LIVE · NEGATIVE/QUALIFIED | SUMMARY.md | 11 | none stated | cell analysis | comparable live-model rerun |
| Apparatus Criticus | SEALED · VALIDATED WITH CAVEATS | SUMMARY.md | 44 | apparatus-criticus-prereg-freeze (recorded; ref unavailable) | sealed judge + validated gold | mechanical re-verification |
| RAG comparison | SEALED · NARROWS | SUMMARY.md | 13 | 6463d73; prereg-tracer-9-iw | mechanical + double grade | comparable live-model rerun |
| Flat mirror | STAGING FAILURE | AUDIT_…md | 29 | corpus lock e6429d0 | mechanical staging gate | mechanical re-verification |
| Bounded Return run-01 | SELECTOR-LIMITED; ENVELOPE UNDISCHARGED | seal05/RESULTS.md | 103 | f5ff001 | deterministic + live adjudication | historical receipt only |
| Bounded Return run-02 | NOT-BOUNDED (COMPLETION); WALLS HELD | RESULTS.md | 114 | SEAL.json SHA b8091467…; producer fe5b479 | deterministic + live adjudication | historical receipt only |
The artifact omits raw populations, transcripts, and large corpora. “Exact deterministic reproduction” reruns pinned computation; “mechanical re-verification” checks preserved artifacts; a “comparable live-model rerun” cannot promise identical hosted-model output; “historical receipt only” preserves a sealed event whose live environment is not claimed reproducible.
node tools/verify.mjs documents/evidence.doc.html
python tools/verify.py documents/evidence.doc.html
A PASS means the manifest and body agree on every witnessed byte span. It does not rerun the live-model studies.
Question 7 · Where does the design come from?
doc.html invents no cryptography and no document theory. Every primitive it uses is decades old, public, and openable from this shelf; the contribution is the join.
The format holds three properties at once — a per-section integrity witness, a document that carries its own readable specification, and a single server-free file. Each descends from a different lineage, and each lineage ships its property only by relaxing the other two: the witnessed-log tradition needs a log server, literate programming carries no witness, the single-file tradition carries neither spec nor witness. The design work was not invention but recombination under constraint — and the shelf below is where each constraint was learned.
A selection of the project's design essays — why a document built for two readers, the format seen whole, where it sits among its ancestors — now sits on the shelf of the wiki of witnessed documents: a root document that links each essay as a whole doc.html of its own and pins it with a cross-file witness.
The following laboratory paths refer to the private upstream record and are not shipped as public links.
experiments/probes/confluence-join/{PREREG.md, RESULTS.md, RECEIPTS.md} · seal c751e44.experiments/peer-review-research/DOC1_lineage_priorart_advances.md.About this file · colophon
This is a doc.html — a single, self-describing HTML file. The <nav id="manifest"> at the top of the body lists every section in this document; each entry's data-witness is the SHA-256 (hex) of that section's raw inner bytes, so any reader can verify any section with the file alone — no server, no JavaScript, no tooling. The full format definition is SPEC.md, carried in the format's own body as SPEC.doc.html.
Author: Georges Casseus (Ndoto Studios) · License: CC0 1.0 (public domain) · Built: 2026-08-01