The async/future measurements were produced by a C stub in /tmp, and that artifact was destroyed when the session worktrees were removed. The log then asserted results with nothing behind them -- a claim inside an evidence record, which is exactly what turns a chain of custody into a pile. Rerun, not reconstructed. Rebuilding the missing file would have been a fabrication with a fresh timestamp; rerunning produces new evidence with its own. lang/tests/integration/fixtures/future.c the future, as a tagged heap object lang/tests/integration/async_future.sh the harness, 6/6 ok unbound: synchronous, correct result ok unbound: el_await on a non-future passes through, no crash ok bound: does not crash ok bound: the awaited result is correct ok bound: the caller continues BEFORE the body finishes ok bound: wrap returns in <10ms while the body takes 50ms LABELLED AS A REPLICATION. The outcomes were already known when this harness was written, so its expectations are NOT predictions committed in advance. Its evidentiary value is that a third party can reproduce it, not that it was called ahead of time. Recording it as anything stronger would corrupt the record it is meant to repair. The fixture also carries the P5 defect and its fix in a comment: the first el_await dereferenced ->magic off an unvalidated slot and SIGSEGV'd on the unbound path, sixty seconds after the same defect was diagnosed elsewhere in the runtime.
v1 — Experiments
Every change to El on iteration-1 was produced by one loop, run repeatedly:
Ishikawa → scientific method → Six Sigma → repeat
- Ishikawa — name the root cause, not the symptom. Why is this table here? never why is this table ugly?
- Scientific method — state a hypothesis, commit predictions before running, then run it in an isolated worktree and grade every prediction including the ones that failed.
- Six Sigma — eliminate the defect class, then add a control so it cannot silently return.
The organising finding
Predictions that came back FALSE were worth more than the ones that held.
Nineteen cycles, sixty-one predictions. The eleven that failed produced every significant result:
| Failed prediction | What it found |
|---|---|
| "the arity table has drifted from the header" | Zero drift — but 110 functions had no entry at all. The table was not wrong, it was 40% incomplete. |
| "codegen drops below baseline" (×4) | The traversal is irreducible. Walking an AST to find calls does not move no matter who decides. Only the rule and the judgment leave. |
| "guards cannot refuse through the seam" | One line, and refusal works. Six compile-time kinds were unnecessary. |
| "C forbids the struct redefinition" | C allows shadowing — and a different defect surfaced: an exit injection emitted with an empty target. |
| "routing el_bin_lookup through the gate fixes the SIGSEGV" | It did not. The fallback was the hazard: strlen() on an integer. I would have shipped the wrong fix and called it verified. |
A prediction that only ever confirms is a demonstration, not a test. One cycle
was run without committing predictions first — async-half-expressible —
and it produced a rigged result: pthread_join immediately after
pthread_create, with the word DEFERRED printed by the test itself. It had to
be discarded and re-run.
Layout
cycles/ one file per loop, numbered in order, named for the DEFECT
findings/ what the cycles produced, cross-cut by kind
Scoreboard
cycles run 19
predictions committed 61
predictions FALSE 11 ← the useful ones
silent miscompilations found 4
security-relevant defects 2
architecture questions closed 5
defects in my own measurement 4
Every cycle verified the same three things before landing: the compiler self-hosts byte-identically (gen2 == gen3), the native suite passes, and the integration harnesses pass. A cycle that could not show all three did not land.