# Perf Profile — M9 Geometry Priming (ENGRAM_GEOMETRY_PRIMING) **Date:** 2026-08-12 **Branch:** `engram-tiered-storage` **Change:** `ENGRAM_GEOMETRY_PRIMING` (default OFF) in `el_runtime.c` `engram_activate` + `engram_geometry.c` **Method:** A/B over 15 representative queries against a **copy** of the recovered store (`~/.neuron/engram/.neuron.egm.disabled`, ~4190 embedded nodes, 768-d nomic-embed-text), throwaway HOME, ports 48799/48800. **Live `:8742` never touched.** `engram.c` (folded from `server.el`) reused byte-identical across M8 and M9, so the only variable is `el_runtime.c`. Three configs: **A** = M9 flag OFF · **B** = M9 flag ON (`=1`) · **C** = pre-M9 M8 baseline binary. --- ## Build | Artifact | Result | |---|---| | M9 `-O2` link (`… engram_geometry.c … -lssl -lcrypto -lcurl -lpthread -lm`) | rc=0, 499,720 B arm64 | | ASan/UBSan link (`-fsanitize=address,undefined -O1`) | rc=0, 1,945,616 B | | Warnings from `el_runtime.c` / `engram_geometry.c` | **0** (3 pre-existing `-Wparentheses-equality` in generated `engram.c` only) | | `nm`: `engram_geo_mean_build`, `engram_geometry_descriptor` | present (T); `eg_geometry_priming_on` inlined (static-local `.cached` present in both binaries) | > Note: the bare `cc … -lm` link fails with undefined `_curl_*` — `el_runtime.c` uses libcurl for > the ollama embedder. The canonical link must include `-lssl -lcrypto -lcurl` (per `link.sh`). --- ## Latency (wall-clock, `curl -w %{time_total}`, 15 queries) | config | median | p90 | min | max | |---|---|---|---|---| | **A — M9 OFF** | **77.8 ms** | 80.5 ms | 71.1 | 84.2 | | C — M8 baseline | 76.0 ms | 81.2 ms | 71.4 | 91.4 | | **B — M9 ON** | **249.6 ms** | **1039.2 ms** | 169.2 | **1256.3** | - **OFF adds zero cost:** 77.8 ms vs M8 76.0 ms — within noise. The flag is free when unset. - **ON regresses hard:** **3.21x median** (+171.8 ms), **~13x p90** (80 → 1039 ms), max **1.26 s**. - The warm-cache path (global mean already built) is ~0.5 s; the cold path pays the full `engram_geo_mean_build` scan (O(N·dim) over ~4190 × 768). The persistent per-query cost is the **descriptor** itself — covariance eigensolve over up to `max_members` (400) × 768-d plus one `store_get_node` **paged read per member** — run on *every* activation while the flag is ON. --- ## Retrieval quality (the win it was supposed to buy) **Coherence** — mean pairwise cosine in centered space, top-20 by activation strength (node embeddings re-derived via nomic-embed-text; centered against the mean of the gathered result set — the *true* store-wide mean is not exposed by the API, flagged as an approximation): | | OFF | ON | Δ | |---|---|---|---| | mean over 15 queries | 0.1067 | 0.1114 | **+0.0047 (noise)** | | queries where ON > OFF | — | — | **4 / 15** | Two real sparse-cue wins (`self identity values` +0.118, `hebbian learning edges` +0.064), but the **polysemous cues — the disambiguation target — are mostly flat or down.** **Disambiguation** — no clean "scope to one sense" pattern on polysemous cues. Additions/drops are small (±2..8 of 300-item sets) and not sense-coherent (e.g. `memory` gains some on-domain nodes but also infra items; `core` similar). **Count shift:** ON adds sub-threshold neighbors to sparse cues (+3..+4) and trims a few from dense polysemous cues (−1..−3) — consistent with priming warming sparse neighborhoods and damping off-domain seeds on dense ones, but the net does not move measured coherence. --- ## Correctness / safety (all pass) | Check | Result | |---|---| | Byte-identical: **A (OFF) == C (M8)** result id sequence + order, all 15 queries (incl. 301/294/263-item sets) | **PASS** (only wall-clock ACT-R fields differ; `activation_strength` max \|Δ\| = 2e-5) | | WM `promoted` ≤ 24 under ON | holds (exactly 24 on dense cues) | | Queries with results under OFF → empty under ON | 0 | | Crash / hang under ON | none (max hops = 1) | | ASan + UBSan under ON (cold build + warm descriptor paths) | **CLEAN** — no report | --- ## Conclusion - **Deploy default-OFF binary: GO.** Byte-identical to M8, zero cost off, clean build, sanitizer clean. - **Enable flag: NO-GO (for now).** 3.21x median / ~13x p90 latency for no reliable quality gain (coherence +0.0047 mean = noise; no clean disambiguation). Correctness/safety are fine — it simply does not earn its cost. **This is a cost/benefit NO-GO, not a defect.** ### Prerequisites before re-evaluating the flag 1. **Amortize the descriptor cost.** The per-query geo-mean build + eigensolve + paged reads dominate. Cache the neighborhood descriptor (it is the M10 cell-assembly cache's job) and/or compute geometry periodically/off-hot-path rather than on every `engram_activate`. 2. **Center against the true store-wide mean** (the `GeoMeanCache` already computes it) rather than a per-query gathered-set approximation, and re-measure coherence — the current signal may be understated by the approximation. 3. **Re-tune** `ENGRAM_GEO_SEED_LO` / `PRIME_SCALE` / `PRIME_MAX` and re-measure only after (1), so tuning is not chasing latency noise.