Reversal runbook for the ENGRAM_GEOMETRY_PRIMING cutover (default OFF, reversible flag flip; exact rollback) and the A/B perf profile: default-OFF binary GO (byte-identical to M8), enabling the flag NO-GO on latency (3.2x/13x) with no demonstrated recall benefit; safety/sanitizer clean.
5.0 KiB
Perf Profile — M9 Geometry Priming (ENGRAM_GEOMETRY_PRIMING)
Date: 2026-08-12
Branch: engram-tiered-storage
Change: ENGRAM_GEOMETRY_PRIMING (default OFF) in el_runtime.c engram_activate + engram_geometry.c
Method: A/B over 15 representative queries against a copy of the recovered store
(~/.neuron/engram/.neuron.egm.disabled, ~4190 embedded nodes, 768-d nomic-embed-text),
throwaway HOME, ports 48799/48800. Live :8742 never touched. engram.c (folded from
server.el) reused byte-identical across M8 and M9, so the only variable is el_runtime.c.
Three configs: A = M9 flag OFF · B = M9 flag ON (=1) · C = pre-M9 M8 baseline binary.
Build
| Artifact | Result |
|---|---|
M9 -O2 link (… engram_geometry.c … -lssl -lcrypto -lcurl -lpthread -lm) |
rc=0, 499,720 B arm64 |
ASan/UBSan link (-fsanitize=address,undefined -O1) |
rc=0, 1,945,616 B |
Warnings from el_runtime.c / engram_geometry.c |
0 (3 pre-existing -Wparentheses-equality in generated engram.c only) |
nm: engram_geo_mean_build, engram_geometry_descriptor |
present (T); eg_geometry_priming_on inlined (static-local .cached present in both binaries) |
Note: the bare
cc … -lmlink fails with undefined_curl_*—el_runtime.cuses libcurl for the ollama embedder. The canonical link must include-lssl -lcrypto -lcurl(perlink.sh).
Latency (wall-clock, curl -w %{time_total}, 15 queries)
| config | median | p90 | min | max |
|---|---|---|---|---|
| A — M9 OFF | 77.8 ms | 80.5 ms | 71.1 | 84.2 |
| C — M8 baseline | 76.0 ms | 81.2 ms | 71.4 | 91.4 |
| B — M9 ON | 249.6 ms | 1039.2 ms | 169.2 | 1256.3 |
- OFF adds zero cost: 77.8 ms vs M8 76.0 ms — within noise. The flag is free when unset.
- ON regresses hard: 3.21x median (+171.8 ms), ~13x p90 (80 → 1039 ms), max 1.26 s.
- The warm-cache path (global mean already built) is ~0.5 s; the cold path pays the full
engram_geo_mean_buildscan (O(N·dim) over ~4190 × 768). The persistent per-query cost is the descriptor itself — covariance eigensolve over up tomax_members(400) × 768-d plus onestore_get_nodepaged read per member — run on every activation while the flag is ON.
Retrieval quality (the win it was supposed to buy)
Coherence — mean pairwise cosine in centered space, top-20 by activation strength (node embeddings re-derived via nomic-embed-text; centered against the mean of the gathered result set — the true store-wide mean is not exposed by the API, flagged as an approximation):
| OFF | ON | Δ | |
|---|---|---|---|
| mean over 15 queries | 0.1067 | 0.1114 | +0.0047 (noise) |
| queries where ON > OFF | — | — | 4 / 15 |
Two real sparse-cue wins (self identity values +0.118, hebbian learning edges +0.064), but the
polysemous cues — the disambiguation target — are mostly flat or down.
Disambiguation — no clean "scope to one sense" pattern on polysemous cues. Additions/drops are
small (±2..8 of 300-item sets) and not sense-coherent (e.g. memory gains some on-domain nodes but
also infra items; core similar).
Count shift: ON adds sub-threshold neighbors to sparse cues (+3..+4) and trims a few from dense polysemous cues (−1..−3) — consistent with priming warming sparse neighborhoods and damping off-domain seeds on dense ones, but the net does not move measured coherence.
Correctness / safety (all pass)
| Check | Result |
|---|---|
| Byte-identical: A (OFF) == C (M8) result id sequence + order, all 15 queries (incl. 301/294/263-item sets) | PASS (only wall-clock ACT-R fields differ; activation_strength max |Δ| = 2e-5) |
WM promoted ≤ 24 under ON |
holds (exactly 24 on dense cues) |
| Queries with results under OFF → empty under ON | 0 |
| Crash / hang under ON | none (max hops = 1) |
| ASan + UBSan under ON (cold build + warm descriptor paths) | CLEAN — no report |
Conclusion
- Deploy default-OFF binary: GO. Byte-identical to M8, zero cost off, clean build, sanitizer clean.
- Enable flag: NO-GO (for now). 3.21x median / ~13x p90 latency for no reliable quality gain (coherence +0.0047 mean = noise; no clean disambiguation). Correctness/safety are fine — it simply does not earn its cost. This is a cost/benefit NO-GO, not a defect.
Prerequisites before re-evaluating the flag
- Amortize the descriptor cost. The per-query geo-mean build + eigensolve + paged reads
dominate. Cache the neighborhood descriptor (it is the M10 cell-assembly cache's job) and/or
compute geometry periodically/off-hot-path rather than on every
engram_activate. - Center against the true store-wide mean (the
GeoMeanCachealready computes it) rather than a per-query gathered-set approximation, and re-measure coherence — the current signal may be understated by the approximation. - Re-tune
ENGRAM_GEO_SEED_LO/PRIME_SCALE/PRIME_MAXand re-measure only after (1), so tuning is not chasing latency noise.