Files
el/DESIGN.md
T
Neuron 6291a35bb9 design: gate on THREE signals -- the alloc gate would have missed el #132
el #132's quadratic (strlen per character in str_char_code/str_slice) is
pure CPU and allocates NOTHING. Measured on three controlled specimens:

  specimen  allocs        bytes         time
  linear    2.00 -> O(n)  2.16 -> O(n)  2.05 -> O(n)
  accum     2.00 -> O(n)  3.99 -> O(n2) noisy
  compute   FLAT          FLAT          3.96 -> O(n2)

'compute' is #132's shape. A gate fitting only allocation count and bytes
classifies it FLAT and passes -- it would not have caught the defect it
was created for. The gate now fits time AND count AND bytes, failing if
any exceeds its declared curve.

Also: black_box is mandatory and consuming the result is NOT sufficient.
The first 'compute' reported 0us at every n while returning a correct n2 --
clang closed the loop to a multiply. Only an opaque call restored the curve.

Adds lang/tests/bench/fitprobe.el as the fitter's known-good/known-bad set,
so the classifier is provable without depending on a real bug existing.
Marks DESIGN.md 1.3 stale: test_compiler 3.58s -> 0.03s (119x).
2026-08-15 21:34:39 -05:00

29 KiB
Raw Blame History

El Test Framework — Design

Status: draft for review Author: Neuron Date: 2026-08-15 Worktree: /Users/will/Development/neuron-technologies/el-worktrees/elc-memory-investigation


0. The forcing requirement

We have a confirmed quadratic in elc. Peak memory in the old shipped binary and wall-clock in the current source both grow as O(input²). We cannot fix it, because we cannot test it.

Everything in this document is downstream of one sentence: a test framework must be able to fail a build when an operation's growth curve degrades from linear to quadratic.

That is not a nice-to-have bolted onto a correctness framework. It is the requirement that determines the architecture. Correctness testing is the easy half.

Second-order requirement, learned the hard way tonight: the framework must report per-test timing by default. The current framework prints N passed, M failed and nothing else. That is why a 3.58-second test file sat in the suite unnoticed. A framework that is structurally blind to time cannot surface the defect class we most need to catch.


1. What exists today, measured

1.1 Two competing systems, neither complete

System A — lang/runtime/test.el. Manual registration, El-level.

System B — the compiler's test { } block + elc --test. Emits its own harness main() with __el_pass / __el_fail globals (codegen.el:3777-3796).

They do not share a result model. Neither has timing. Both are in the tree.

1.2 Specific defects in System A

Defect Location Consequence
All state as JSON strings in a global string-keyed map test.el throughout every assertion is state_getstr_to_intint_to_strstate_set
Failure list appended by string slice + concat _test_json_append O(n²) in failure count
One OS thread spawned per test _test_run_one via __thread_create/__thread_join thread spawn per test, purely to get dispatch-by-name through dlsym
Manual registration pairing a string to a function name test_case(name, fn_name) typo ⇒ test silently never runs, suite still reports pass
Counters are assertion-level, global _test_pass_count etc. no per-test record exists at all
No timing, no structured output, no fixtures, no tags, no filtering, no parameterization, no benchmarks

The registration defect is the serious one. It is not a slow framework, it is a framework that can report success for tests that did not execute.

1.3 Measured cost structure

Per test file, current build model:

Step Time
elc compile .el.c 0.00s (small files)
cc el_runtime.c → .o 0.14s
cc test .c → .o 0.02s
link 0.02s

STALE as of el #132 — re-measured 2026-08-16. The test_compiler figure below was entirely the strlen-per-character quadratic, now fixed. Re-measured on the same host: 3.58s → 0.03s (119x), and the 422 KB compiler concatenation likewise compiles in 0.03s. The table is retained only as the historical record that motivated the gate. The remaining per-file cost is the redundant el_runtime.c rebuild, which §9's compile-once architecture addresses.

Per-file elc time across the existing suite:

File Bytes elc time
test_compiler 29,685 (+394 KB of imports) 3.58s
string_test 18,545 0.01s
all other 9 files 2.210 KB 0.00s

Two distinct defects in two distinct regimes:

  1. test_compiler.el imports all five compiler sources — 394 KB in one translation unit. Its 3.58s is entirely the quadratic. It is the only file where the quadratic bites.
  2. Every other file's cost is 100% redundant el_runtime.c rebuilds — 480 KB of identical C, recompiled once per test file.

Neither is fixed by making the compiler faster. Both are fixed by the architecture below, and the speedup is a by-product of building it correctly, not the goal.

1.4 The asset worth keeping

codegen.el:3651-3652 already collects test_names / test_c_namesthe compiler already does compile-time test discovery. It then discards that registry into a hardcoded main().

That registry is precisely the seam Go's _testmain.go and Rust's test_main_static are built on. The mechanism we need is half-built and wired to the wrong thing.


2. Grounding — the common spine of excellent frameworks

Researched from primary sources: Go testing/go test, Rust libtest/Criterion, JUnit 5 Platform, NUnit 3, JMH, Google Benchmark. Six invariants hold across all of them.

  1. A registry is built before execution(name, metadata, fn-ptr) triples. Go generates it from an AST scan; Rust synthesizes it in a compiler pass; JMH emits it as a build-time resource; JUnit/NUnit build it reflectively. Reflection is an implementation of the registry on runtimes where it is cheap. It is never the architecture.

  2. Discovery strictly precedes execution. Every good capability — filtering, listing, counting, sharding, IDE trees, re-run-failed-only, dry runs — is a consequence of this ordering.

  3. A hierarchy with stable, path-shaped unique IDs. TestFoo/subcase_2. Selection is regex over that path, one pattern per level.

  4. The framework is a prebuilt library; only the entry point is generated. "Compile once, link many" is always: framework archive compiled once + a small generated table + one MainStart(deps, registry) call. Nobody recompiles the harness per test file.

  5. Execution emits an event stream; reporters are downstream renderers. Human text, NDJSON, JUnit XML, TAP are all transforms of one event stream. Go's one architectural mistake is doing this backwards — test2json parses human output, and has shipped bugs when user output contains --- PASS:.

  6. A dependency-injection seam at the boundary. Go's testdeps.TestDeps exists so testing can avoid importing regexp, profilers, and coverage. The execution core knows nothing about output formats.


3. Architecture

3.1 The seam

  ┌─────────────────────────────────────────────────────────────┐
  │ user code:  foo.el  with  test { } / bench { }  blocks       │
  └───────────────────────────┬─────────────────────────────────┘
                              │  elc --test
                              ▼
  ┌─────────────────────────────────────────────────────────────┐
  │ generated C (per suite, tiny):                               │
  │   __el_test_fn_0 .. _N      lowered test/bench bodies        │
  │   __el_registry[]           static table: name/kind/file/    │
  │                             line/tags/sizes/expected-O       │
  │   __el_dispatch(i)          generated switch → body          │
  │   main() { return el_test_main(argc, argv); }                │
  └───────────────────────────┬─────────────────────────────────┘
                              │  cc + link  (registry only)
                              ▼
  ┌─────────────────────────────────────────────────────────────┐
  │ libeltest.a  — PREBUILT ONCE                                 │
  │   • el_runtime.o        (the 480 KB, compiled once, ever)    │
  │   • eltest.o            the runner, WRITTEN IN EL            │
  │       discovery view · filtering · execution · fixtures ·    │
  │       timing · benchmark harness · curve fitting · reporters │
  └─────────────────────────────────────────────────────────────┘

The framework is written in El, compiled to C once, archived. Per-suite compilation touches only the generated registry. This is Go's model, and it is strictly better for us than Go's because we own the compiler and already have the AST — no separate source-scanning pass is needed.

3.2 Why the runner is in El and the registry is in C

El has no closures and no first-class function pointers. The registry must therefore hold C function pointers, and it is generated C.

The runner stays in El and reaches the registry through a small builtin surface — indices, not pointers:

__el_reg_count() -> Int
__el_reg_name(i) -> String
__el_reg_file(i) -> String
__el_reg_line(i) -> Int
__el_reg_kind(i) -> Int          // 0=test 1=bench
__el_reg_tags(i) -> Int
__el_reg_sizes(i) -> String      // JSON array, empty for tests
__el_reg_expect(i) -> Int        // complexity class enum, 0 = none
__el_reg_invoke(i) -> Int        // runs the body via the generated switch

Nine builtins. Everything else — filtering, lifecycle, statistics, curve fitting, all reporters — is El. That satisfies "written in El" without pretending El can do something it cannot.

3.3 Result model

The unit is a result record, not a counter:

TestResult {
  id        String     // slash path: "parser/handles_empty_input/case_3"
  file      String
  line      Int
  status    Status     // Pass | Fail | Error | Skip
  duration  Int        // nanoseconds, ALWAYS populated
  message   String     // assertion detail: expected vs actual
  output    String     // captured stdout/stderr for this test
  assertions Int
}

Fail = an assertion failed. Error = unexpected crash/abort. This distinction is load-bearing — every CI consumer depends on it, and the JUnit XML schema encodes it as distinct elements.


4. Authoring surface

4.1 Tests

test { } already exists. Keep it. Add subtests and hierarchy:

test "parser/empty input" {
    assert_that(parse(""), is_err())
}

test "parser/table" {
    for case in [["", 0], ["a", 1], ["a b", 2]] {
        subtest(case[0]) {
            assert_that(token_count(case[0]), equals(case[1]))
        }
    }
}

Subtest IDs compose as parser/table/a_b. Filtering is --run 'parser/table/.*', one regex per path segment, exactly as Go does.

We do not build a parameterized-test annotation system. Table-driven loops plus subtests subsume @ParameterizedTest, @MethodSource, @CsvSource, and TestCaseSource entirely, at zero framework surface. This is Go's single biggest ergonomic win over JUnit and NUnit.

4.2 Fixtures

Per-file and per-test only, plus a LIFO cleanup stack:

setup_all  { ... }      // once per suite
setup      { ... }      // before each test
teardown   { ... }      // after each test
teardown_all { ... }

and inside a test, cleanup { ... } registering LIFO-ordered teardown.

We do not build JUnit 5's extension SPI — seventeen callback interfaces, hierarchical stores, registration ordering rules. That complexity is the price of retrofitting a plugin ecosystem onto a twenty-year-old reflective framework. Go's t.Cleanup covers roughly 90% of what @AfterEach is used for at a fraction of the surface.

4.3 Assertions — constraint model

One entry point, composable constraint values (NUnit's model, which avoids the N² overload explosion):

assert_that(actual, equals(expected))
assert_that(xs,     has_length(3))
assert_that(s,      contains("foo").and(starts_with("bar")))
assert_that(f,      is_within(0.01).of(3.14))

A constraint is a value with apply_to(actual) -> ConstraintResult, and the result knows how to describe its own failure. Custom constraints are ordinary user types.

Every failure message must name file, line, the expression text, and both values. We capture expression source text at compile time — we have the AST, so we can do this better than any runtime-introspection framework.

Legacy assert_true / assert_eq / etc. stay as thin wrappers for migration.


5. Benchmarks

5.1 The loop

Adopt b.Loop(), not b.N. Go spent fifteen years on b.N before concluding b.Loop was right; we skip that.

bench "str_concat" {
    let s = make_input(bench_n())
    for bench_loop() {
        black_box(str_concat(s, "x"))
    }
}

Three properties that make this the correct choice for a C target:

  1. The timer auto-resets on first call, so setup above the loop is excluded by construction rather than by the author remembering ResetTimer.
  2. N is hidden, so it cannot be misused.
  3. The harness owns the loop shape, which lets us insert an optimization barrier the C compiler cannot see through. black_box(v) lowers to asm volatile("" :: "r"(&v) : "memory"). Since we emit a single translation unit, dead-code elimination of a benchmark body is a live hazard — this is our version of JMH's Blackhole problem, solved in the harness rather than delegated to the user.

5.2 Iteration scaling

Use Go's predictN heuristics verbatim. They are battle-tested and cheap:

n = goal_ns * prev_iters / prev_ns    // multiply before divide — precision on sub-ns ops
n += n / 5                            // 20% headroom, overshoot rather than re-loop
n = min(n, 100 * last)                // never grow more than 100× per step
n = max(n, last + 1)                  // guarantee forward progress
n = min(n, 1_000_000_000)             // hard ceiling

Report n rounded to 1/2/3/5 × 10ᵏ so runs are comparable.

5.3 Sampling

Criterion's shape, because it is correct near timer resolution:

  • Warmup: iteration counts 1, 2, 4, 8… until cumulative time exceeds the warmup budget.
  • Measurement: collect sample_size samples at iteration counts [d, 2d, 3d, …, Nd].
  • Estimate: slope of a linear regression of iteration-count vs elapsed time. The intercept absorbs fixed overhead.
  • Time whole samples, never individual iterations. This is the single most important detail — it defeats timer-resolution error on nanosecond operations.

Outliers classified by modified Tukey (±1.5 IQR mild, ±3 IQR severe), reported but retained.


6. Complexity gating — the centerpiece

This is the part that makes the quadratic fixable, and the part nobody in the mainstream has finished. Google Benchmark's Complexity() fits the curve and reports it. We declare it and gate on it.

6.1 Surface

bench "elc_compile" over n in [16, 32, 64, 128, 256, 512, 1024] expect O(n) {
    let src = synth_source(bench_n())
    for bench_loop() { black_box(compile(src)) }
}

Alternative with no new syntax, if the parser change is judged too invasive — bench_sizes([...]) and bench_expect("O(n)") as calls inside the block. Recommendation: declarative. Runtime calls mean --list cannot show the invariant without executing, which breaks the discovery-precedes- execution invariant from §2.

6.2 Fitting

Per Google Benchmark src/complexity.cc. For candidate curves {O(1), O(log n), O(n), O(n log n), O(n²), O(n³)}, one-parameter least squares, no intercept:

coef = Σ(tᵢ · gᵢ) / Σ(gᵢ²)
rms  = sqrt( Σ(tᵢ  coef·gᵢ)² / k ) / mean(t)     // normalized

Best fit = lowest normalized RMS. User-supplied lambda curves also supported.

6.3 Gate logic

  1. FAIL if the best-fit curve is strictly worse than declared, ordering O(1) < O(log n) < O(n) < O(n log n) < O(n²) < O(n³). Print the fitted coefficient and the full per-size table.
  2. FAIL if the declared curve's normalized RMS exceeds a threshold (start at 0.10). This catches the case where no candidate fits — noise, a cache cliff, or a phase change. Report INDETERMINATE honestly rather than gating on garbage.
  3. WARN if the best fit is strictly better than declared — either an optimization landed and the annotation should tighten, or the sweep is too narrow to expose real behaviour.
  4. REFUSE to gate on fewer than 5 distinct sizes spanning under 2 decades, geometrically spaced. Say so loudly rather than producing a meaningless fit.

6.4 Why gate on the exponent, not wall-clock

  • Machine-independent. The fitted exponent is a property of the algorithm; the coefficient is a property of the machine. Gating on the exponent makes CI hardware heterogeneity, noisy neighbours, and thermal throttling irrelevant — they scale coef, not g.
  • No stored baseline. No artifact storage, no golden-file drift. The invariant lives in the source next to the code and is reviewed in the same PR.
  • It catches the failure mode that actually ships. An O(n) lookup inside an O(n) loop is invisible at n=100 in a unit test and catastrophic at n=100,000 in production. Constant-factor regressions are annoying. Complexity regressions are outages. Ours was a 27 GB outage.

6.5 The deterministic gate — the one that would have caught us

Wall-clock needs statistics. Allocation counts do not. They are perfectly deterministic.

Correction, 2026-08-16 — count alone is NOT sufficient. Gate on BOTH count and bytes.

Measured against two El programs, one allocating once per item and one rebuilding its accumulator each iteration:

n linear allocs / bytes quadratic allocs / bytes
100 100 / 290 100 / 5,150
200 200 / 690 200 / 20,300
400 400 / 1,490 400 / 80,600
800 800 / 3,090 800 / 321,200

The quadratic program's allocation count is exactly linear — 100/200/400/800, identical to the healthy program. A count-only gate passes it clean. Bytes catch it: each doubling of n quadruples bytes (ratios 3.94, 3.97, 3.99 → 4.0 = O(n²)) where the linear program converges on 2.0.

This is precisely elc's own defect shape — a copy-on-write accumulator reallocating once per pass (count linear) into a proportionally larger buffer (bytes quadratic).

Therefore expect allocs O(n) fits count and bytes independently and fails if EITHER exceeds the declared curve, reporting which signal broke. "count linear, bytes quadratic" is a precise, directly actionable diagnosis.

el_peak_rss() is CONTEXT ONLY — never gate on it. It is perturbed by the allocator and by the page cache. Allocation volume is the invariant; RSS and malloc/free churn are merely the two surfaces it shows on. The old shipped compiler paid the same quadratic in RSS that the rebuilt one pays in churn.

Measure rate, not level. A guard reading swap level saw 97% on a thrashing host and 97% on a healthy one; only rate separated them. A growth exponent is a rate; a single measurement is a level. That is why the gate fits a curve across a sweep instead of comparing one number to a threshold.

Second correction, same day — THE ALLOCATION GATE ALONE WOULD HAVE MISSED THE REAL BUG.

el #132 found the actual elc quadratic: strlen() called inside str_char_code() and str_slice(), so the lexer rescanned the remaining input on every character. Pure CPU. Zero allocation. str_char_code is a bounds check and an index — it allocates nothing.

Measured on three controlled specimens (lang/.work/fitprobe.el), growth ratio per doubling of n across n = 200/400/800/1600:

specimen allocs bytes time what it proves
linear — one alloc per item 2.00 2.00 2.00 → O(n) 2.16 2.07 2.23 → O(n) 0.83 2.00 2.05 → O(n) clean baseline
accum — rebuilds accumulator 2.00 2.00 2.00 → O(n) 3.97 3.99 3.99 → O(n²) noisy count misses, bytes catches
compute — n scans over n chars 0 → FLAT 0 → FLAT 3.93 4.01 3.96 → O(n²) both alloc signals blind; only time catches

compute is el #132's shape exactly. A gate fitting only allocation count and bytes classifies it as FLAT and passes it. The gate as originally specified would not have caught the defect it was created for.

Therefore the gate fits THREE signals and fails if ANY exceeds its declared curve:

bench "elc_compile" over n in [...] expect time O(n) allocs O(n) bytes O(n) { ... }
  • allocs (count) — deterministic, zero-noise. Catches per-item allocation growth.
  • allocs (bytes) — deterministic, zero-noise. Catches accumulator-rebuild quadratics that count cannot see.
  • time — noisy, needs the sweep and statistics. The ONLY signal that sees pure-compute complexity regressions. Gate on the fitted exponent, never on absolute duration, so CI hardware variance scales the coefficient and leaves the classification intact.

The deterministic signals remain preferable where they apply — they need no statistics and are correct on the first run. They are simply not sufficient.

black_box is mandatory, and consuming the result is NOT enough. The first version of compute accumulated total + 1 in a nested loop and reported 0 µs at every n while returning a numerically correct n². Clang recognised the idiom and closed the loop to a multiply. Feeding the result into output did not prevent it. Only making the inner operation an opaque external call restored the real curve. A benchmark harness that trusts the user to defeat the optimiser will silently measure nothing — and report success while doing it.

Instrument the runtime with allocation counters and fit those against n instead of time:

bench "elc_compile" over n in [...] expect O(n) allocs O(n) { ... }

Zero noise, zero statistics, always gateable, correct on the first run on any machine. Go reports allocs/op and B/op; nobody fits them against n. That is an open opportunity and it is exactly our bug: elc's defect is quadratic allocation volume, which the old binary paid in RSS and the current source pays in malloc/free churn.

An expect allocs O(n) assertion on elc's compile path would have failed the build the day the quadratic was introduced.

Required runtime additions: __el_alloc_count(), __el_alloc_bytes(), __el_peak_rss().

6.6 Constant-factor gate (secondary, opt-in)

Mann-Whitney U at α = 0.05, noise floor 1%, medians with 95% CIs, ~ for not-significant. Requires --count >= 9. Off by default on CI; opt-in per benchmark.

Exit nonzero on regression. Both benchstat and Criterion always exit 0, which is why every shop using them wrote a wrapper. We do not repeat that omission.


7. Output

Structured events are the source of truth. Human text is rendered from them. We do not repeat Go's parse-the-human-output design.

Event stream, NDJSON, one object per line, streamed live:

{"time":"...","action":"run","test":"parser/empty"}
{"time":"...","action":"output","test":"parser/empty","output":"..."}
{"time":"...","action":"pass","test":"parser/empty","elapsed":0.0031}
{"time":"...","action":"bench","test":"str_concat","n":1024,"ns_op":41.2,"allocs_op":3,"bigo":"N","rms":0.03}

Renderers, all downstream and pluggable:

Format Flag Use
Human default terminal, per-test duration always shown
NDJSON --json tooling, history, flaky detection
JUnit XML --junit-xml=PATH every CI system on earth
TAP --tap optional

JUnit XML per the de-facto schema: testsuitestestsuitetestcase, with time in seconds as a decimal, file/line attributes, and failure vs error vs skipped as distinct child elements. Absence of a child element means pass. Emit <testsuites> even for a single suite, and parse both shapes on input.


8. CLI

--list                     print the registry, run nothing
--list-json                machine-readable registry
--run PATTERN              slash-separated regex per path segment
--tag EXPR                 tag expression: fast & !slow
--shard I/N                deterministic sharding for CI parallelism
--count N                  repetitions, for statistics
--bench PATTERN            run benchmarks (off by default in test runs)
--benchtime DUR            per-benchmark time budget
--junit-xml PATH
--json
--isolate                  re-exec per test on crash, so one SIGSEGV doesn't lose the run
--timeout DUR
--fail-fast

--list / --list-json / --shard cost roughly thirty lines because the registry already exists before main does anything. That is the dividend of discovery-precedes-execution.


9. Build model

# once, ever (or when the runtime/framework changes):
cc -c el_runtime.c        -o el_runtime.o
elc eltest.el > eltest.c && cc -c eltest.c -o eltest.o
ar rcs libeltest.a el_runtime.o eltest.o

# per suite:
elc --test foo_test.el > foo_test.c     # registry + bodies only
cc foo_test.c libeltest.a -o foo_test

The 0.14s × N of redundant runtime rebuilds disappears — not because we optimized it, but because one-runner-over-many-suites requires compile-once-link-many as a structural precondition.


10. Bootstrap and self-hosting

The framework's own tests are test { } blocks run by the framework. Same fixpoint discipline the compiler already applies to itself.

  1. Build the framework using the existing harness for its first tests (stage 0).
  2. Rebuild the framework's tests as test { } blocks run by the new runner (stage 1).
  3. Verify stage 1 reports identical results to stage 0.
  4. From then on, the framework is tested by itself.

A framework that cannot run its own suite is not evidence of anything. This is a correctness proof, not a claim.


11. Explicitly not building

Rejected Why
Naming-convention discovery (fn test_foo) test { } is a real declaration. Go's TestXxx exists only because Go had no better hook — and it needs a heuristic to avoid matching TesticularCancer.
Reflection or symbol-table scanning Slow, fragile under LTO/strip/dead-strip, and unnecessary when we own the compiler.
Parsing human output into structure Go's test2json is its one clear architectural mistake.
JUnit 5's extension SPI Seventeen callback interfaces to retrofit plugins onto a reflective framework. Not our problem.
@ParameterizedTest machinery Table-driven loops + subtests subsume it at zero surface.
NUnit's out-of-process agents They bridge CLR versions and AppDomains. We emit one native binary. Keep --isolate as crash fallback only.
JMH-style forking by default Forks exist because JIT profiles are per-process. AOT C has no such state. Keep --fork available, not default.
Exit 0 on regression benchstat and Criterion both do this, and every user writes a wrapper.
Dynamic runtime test registration Breaks --list, sharding, and individual selection. Registry stays static.

12. Phasing

Phase Content Gate
1 Registry emission in codegen; 9 builtins; el_test_main skeleton in El; result records; per-test timing; human + NDJSON output existing 11 test files pass, with timing
2 libeltest.a build model; subtests; filtering; --list; fixtures; constraint assertions; JUnit XML suite runs in one binary; runtime compiled once
3 bench { }, bench_loop, black_box, predictN, Criterion sampling benchmarks produce stable ns/op
4 Allocation counters; complexity fitting; expect O(...) gate an expect allocs O(n) benchmark on elc fails on the current quadratic
5 Migrate both legacy systems; delete runtime/test.el; self-host framework runs its own suite

Phase 4 is the deliverable that matters. Phases 13 exist to make it possible.


13. Open questions for review

  1. Declarative over n in [...] expect O(...) syntax vs runtime calls. I recommend declarative (§6.1) so --list can show invariants without executing. It costs parser work. Your call.
  2. bench { } as a new block form — parallel to test { }, or a modifier on it?
  3. Scope of the constraint model. Full composable constraints, or start with a flat assertion set and add constraints later? Full model is more surface but avoids a second migration.
  4. Does runtime/test.el get deleted or kept as a deprecated shim? I lean delete — two systems is how we got here.
  5. Where does libeltest.a live in the tree, and does epm need to know about it?
  6. Allocation counters in el_seed.c or el_runtime.c? AGENTS.md says el_seed.c is the sole C dependency and hand-maintained; counters are OS-boundary-adjacent but not OS calls.
  7. Is per-test timing enough, or do we want per-assertion timing for finding slow helpers?

14. What this document is not

This is a design, not a measurement. Every performance claim about the current system in §1 is measured and reproducible in this worktree. Every claim about the proposed system is a prediction. None of it is verified until Phase 1 runs and Phase 4 fails a build on the real quadratic.