8.4 KiB
Engram Tiered Storage — M8 Performance Profile (milestone-0 sample)
Milestone: M8 (ANN wired into engram_activate seed selection). Trunk =
worktree /tmp/engram-tiered-wt, branch engram-tiered-storage, HEAD 1507614.
Date: 2026-08-12. Author: first full-binary build + profile of the tiered trunk.
This is sample zero of an accumulating per-milestone profile series (see the BUILD-LEDGER "keep performance telemetry CONTINUOUSLY" decision). Compare future milestones (M9, M10, …) against these numbers.
Methodology (read this before trusting a number)
- On a COPY, never live. All runtime measurements used a read-only copy of the
live store booted on a non-live port (:8798) with a throwaway
$HOME. The live engram service (:8742,~/.neuron/engram) was never touched. - Two store substrates were used:
- Live-egm copy (
neuron.egm458 MB +neuron.wal44 MB, copied read-only) — this substrate crashes both the M8 binary and the live binary on boot (see Integration Findings). Unusable for runtime measurement. - Clean import — a dir seeded with only
snapshot.json(65 MB, stable 06:10), which the binary imported into a fresh 59.9 MBneuron.egm. All healthy runtime numbers below are from this substrate (real graph content, healthy store).
- Live-egm copy (
- Build machine: Apple Silicon (arm64), macOS. Native
cc -O2compile; fold done in a memory-capped (--memory=3g --memory-swap=3g, no swap)linux/amd64container runningelc-linux-amd64. - Hardware/thermals uncontrolled; single run per metric unless noted. Treat as order-of-magnitude, not benchmark-grade, except the module benchmark (Test suite).
Build
| Metric | Value |
|---|---|
| Fold input | engram/src/server.el (44,273 B El, no imports) |
| Fold output | engram.c (30,855 B, 645 lines C), ELC_EXIT=0, 0 fold warnings |
| Fold time (pure elc) | sub-second (server.el is small, importless) |
| Fold container wall | ~38 s (dominated by one-time apt-get install libcurl4 in the throwaway container; the elc invocation itself is <1 s) |
| Compile | cc -O2 -DHAVE_CURL engram.c el_runtime.c engram_store.c engram_vindex.c -lssl -lcrypto -lcurl -lpthread -lm |
| Compile time | 1.83 s wall |
| Compile warnings | 3, all -Wparentheses-equality in the folded engram.c (El if-expr codegen emits if ((x == 0))); cosmetic. el_runtime.c / engram_store.c / engram_vindex.c compiled 0 warnings — notably none around the M8 deferred free(e_eff) or the vindex integration. |
| Binary | 482,008 B (471 KB) Mach-O arm64 executable |
| ANN linkage verified | nm: vindex_search, vindex_build_from_store, eg_vindex_sync, _eg_vindex, store_scan_nodes, engram_store_boot, engram_activate all present |
Boot & footprint (clean-import substrate)
| Metric | Value |
|---|---|
Boot from healthy neuron.egm |
~2 s to listening |
Boot from snapshot.json (one-time import + fresh egm) |
~9 s |
| node_count | 13,036 (matches ledger import-dedup: 13,038 snapshot − 2 dup-id entries) |
| edge_count / layer_count | 43,402 / 5 |
| embedded_count | 4,190 |
| Fresh egm size | 59.9 MB (vs the live egm's bloated 458 MB — see Findings) |
| RSS after boot | 126.2 MB (whole 13k-node/43k-edge/4,190-emb graph resident + fresh egm) |
Activation latency — q="bullshit" (clean-import substrate)
50 sequential GET /api/activate?q=bullshit&limit=10&depth=3:
| Metric | Value |
|---|---|
| p50 | 34.35 ms |
| p95 | 35.26 ms |
| min / max | 33.50 ms / 9,886 ms |
| Sample | n=50 |
- The max = 9.9 s is the first call only — a cold query-embedding fetch
(
eg_embed_fetch→ Ollamanomic-embed-text, cold model load). All subsequent calls hit the single-slot query-embedding cache (_eg_qcache) → 34 ms steady state. - This 34 ms is the lexical/spread path, NOT the ANN seed path. The store's 4,190
embeddings were generated by the live neuron's native embedding model; the harness's
nomic-embed-textquery vectors are a different vector space, so no candidate clearedENGRAM_EMBED_SEED_MIN=0.60(act-stats:dup_seeds:0,ctx_cos:-2.000sentinel) → results empty, ANN discovery produced no admitted seeds. Environmental (embedding provenance), not an M8 defect. The ANN wiring still executed (embed fetch succeeded,embed_breaker_open:0;eg_vindex_sync+ seed block +free(e_eff)all ran) without crashing on the real store.
KEY M8 METRIC — ANN vs O(n) seed selection (authoritative)
Because the HTTP path can't exercise ANN seeding without embedding-space parity, the
authoritative ANN-vs-exact-scan numbers come from the module benchmark
(engram/test/run_vindex_tests.sh, PASS 1, optimised), which measures the exact
vindex_search code the M8 wiring calls, at full size:
| N (768-dim vectors) | Brute-force (O(n)) | ANN (HNSW) | Speedup |
|---|---|---|---|
| 5,000 | 3.475 ms/query | 0.353 ms/query | 9.8× |
| 20,000 | 13.809 ms/query | 0.698 ms/query | 19.8× |
- recall@10 = 0.9365 at
ef_search=128(gate ≥0.90 — PASS). Lower ef trades recall for latency: ef=64→0.844, ef=32→0.750, ef=10→0.625. - Determinism: two independent seeded builds give byte-identical query results.
vindex_build_from_storeover a realengram_store: inserts exactly the embedded nodes, top-1 resolves to the correct node id at ~0 distance.- HNSW build cost (single-threaded, note for boot/index-build budgeting):
5,000 vectors ≈ 15–35 s, 20,000 vectors ≈ 75 s. In-process the index is built
lazily on first activation (
eg_vindex_sync) and grown incrementally; the M8 seed block only fires oncevindex_size ≥ ENGRAM_EMBED_SEED_K. At the real store's 4,190 embedded nodes this is a one-time few-second first-activation cost — worth watching as the embedded set grows (a future milestone may want to build the index at boot or persist it viavindex_save/vindex_load).
Seed-set parity note (why there is no runtime A/B toggle)
M8 has no ANN on/off env flag by design (ENGRAM_EMBED_SEED_K is a compile-time
constant). The exact O(n) cosine scan is preserved verbatim and "tops up" any seed
slot the ANN leaves unfilled; every ANN candidate is admitted through the identical
cosine/dedup/threshold gate the exact scan uses. So ANN changes only which nodes are
discovered and how fast, never the final seed set — parity is structural, not
A/B-tested via flag. Recovered-behaviour-when-index-absent is the pre-M8 exact scan.
Integration Findings (first full build of the tiered trunk)
- CRITICAL / pre-existing (NOT M8):
btree_insertstack-buffer-overflow on opening the live 458 MBneuron.egm.SIGABRT(__stack_chk_fail) viabtree_insert ← btree_put ← edge_place/apply_edge_put ← engram_open ← engram_store_boot. Reproduces on the M8 binary and the deployed live binary — the live engram :8742 was crash-looping (18 crash reports 16:18→17:44 on this date; service refusing connections). A clean import into a fresh 59.9 MB egm does not crash (43k edges load fine), so the trigger is the specific pathological live store: 458 MB (8× the healthy 59.9 MB) from a day of churn/tombstones with no M5 compaction + a 44 MB un-checkpointed WAL replayed on open. The bug is inengram_store.c(the edge B-tree / WAL-redo path), upstream of everything M8 touched (M8 lives inel_runtime.c::engram_activate). Fix required before any re-cutover;btree_insertmust bound-check regardless of on-disk content. - M8 deferred
free(e_eff)(the flagged memory-management concern): compiled warning-free, and the wired path executed end-to-end over HTTP on the real store (with a real query embedding) with no crash / no new crash report — no double-free or use-after-free observed. Freed on all early-return paths and exactly once post-seed-selection. - Write durability + retrieval-fix dedup: create → checkpoint → clean SIGTERM → restart → node found by id and by search (node_count 13,036→13,037 preserved).
Caveats
- All on a copy; healthy-substrate numbers are from a re-imported store, not the live paged store (which is currently un-bootable — Finding 1).
- Single-run metrics; no thermal control.
- Embedding-space mismatch prevented a real semantic
q=bullshitactivation in this harness; the ANN speedup number is the module benchmark, which is the correct gauge.