docs(engram): cognitive architecture design, prior-art scan, M10 reification, M8 perf profile, cutover-reversal runbooks, and 2026-08-12 session record
This commit is contained in:
@@ -0,0 +1,140 @@
|
||||
# Engram Tiered Storage — M8 Performance Profile (milestone-0 sample)
|
||||
|
||||
**Milestone:** M8 (ANN wired into `engram_activate` seed selection). Trunk =
|
||||
worktree `/tmp/engram-tiered-wt`, branch `engram-tiered-storage`, HEAD `1507614`.
|
||||
**Date:** 2026-08-12. **Author:** first full-binary build + profile of the tiered trunk.
|
||||
|
||||
This is **sample zero** of an accumulating per-milestone profile series (see the
|
||||
BUILD-LEDGER "keep performance telemetry CONTINUOUSLY" decision). Compare future
|
||||
milestones (M9, M10, …) against these numbers.
|
||||
|
||||
## Methodology (read this before trusting a number)
|
||||
|
||||
- **On a COPY, never live.** All runtime measurements used a read-only copy of the
|
||||
live store booted on a **non-live port (:8798)** with a **throwaway `$HOME`**. The
|
||||
live engram service (:8742, `~/.neuron/engram`) was never touched.
|
||||
- **Two store substrates were used:**
|
||||
1. *Live-egm copy* (`neuron.egm` 458 MB + `neuron.wal` 44 MB, copied read-only) —
|
||||
**this substrate crashes both the M8 binary and the live binary on boot** (see
|
||||
Integration Findings). Unusable for runtime measurement.
|
||||
2. *Clean import* — a dir seeded with only `snapshot.json` (65 MB, stable 06:10),
|
||||
which the binary imported into a **fresh 59.9 MB `neuron.egm`**. All healthy
|
||||
runtime numbers below are from this substrate (real graph content, healthy store).
|
||||
- **Build machine:** Apple Silicon (arm64), macOS. Native `cc -O2` compile; fold done
|
||||
in a memory-capped (`--memory=3g --memory-swap=3g`, no swap) `linux/amd64` container
|
||||
running `elc-linux-amd64`.
|
||||
- Hardware/thermals uncontrolled; single run per metric unless noted. Treat as
|
||||
order-of-magnitude, not benchmark-grade, except the module benchmark (Test suite).
|
||||
|
||||
## Build
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| Fold input | `engram/src/server.el` (44,273 B El, no imports) |
|
||||
| Fold output | `engram.c` (30,855 B, 645 lines C), `ELC_EXIT=0`, **0 fold warnings** |
|
||||
| Fold time (pure elc) | sub-second (server.el is small, importless) |
|
||||
| Fold container wall | ~38 s (dominated by one-time `apt-get install libcurl4` in the throwaway container; the elc invocation itself is <1 s) |
|
||||
| Compile | `cc -O2 -DHAVE_CURL engram.c el_runtime.c engram_store.c engram_vindex.c -lssl -lcrypto -lcurl -lpthread -lm` |
|
||||
| Compile time | **1.83 s** wall |
|
||||
| Compile warnings | **3**, all `-Wparentheses-equality` in the *folded* `engram.c` (El if-expr codegen emits `if ((x == 0))`); cosmetic. `el_runtime.c` / `engram_store.c` / `engram_vindex.c` compiled **0 warnings** — notably none around the M8 deferred `free(e_eff)` or the vindex integration. |
|
||||
| Binary | **482,008 B (471 KB)** Mach-O arm64 executable |
|
||||
| ANN linkage verified | `nm`: `vindex_search`, `vindex_build_from_store`, `eg_vindex_sync`, `_eg_vindex`, `store_scan_nodes`, `engram_store_boot`, `engram_activate` all present |
|
||||
|
||||
## Boot & footprint (clean-import substrate)
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| Boot from healthy `neuron.egm` | **~2 s** to listening |
|
||||
| Boot from `snapshot.json` (one-time import + fresh egm) | **~9 s** |
|
||||
| node_count | **13,036** (matches ledger import-dedup: 13,038 snapshot − 2 dup-id entries) |
|
||||
| edge_count / layer_count | 43,402 / 5 |
|
||||
| embedded_count | 4,190 |
|
||||
| Fresh egm size | **59.9 MB** (vs the live egm's bloated 458 MB — see Findings) |
|
||||
| RSS after boot | **126.2 MB** (whole 13k-node/43k-edge/4,190-emb graph resident + fresh egm) |
|
||||
|
||||
## Activation latency — q="bullshit" (clean-import substrate)
|
||||
|
||||
50 sequential `GET /api/activate?q=bullshit&limit=10&depth=3`:
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| p50 | **34.35 ms** |
|
||||
| p95 | **35.26 ms** |
|
||||
| min / max | 33.50 ms / 9,886 ms |
|
||||
| Sample | n=50 |
|
||||
|
||||
- The **max = 9.9 s is the first call only** — a cold query-embedding fetch
|
||||
(`eg_embed_fetch` → Ollama `nomic-embed-text`, cold model load). All subsequent
|
||||
calls hit the single-slot query-embedding cache (`_eg_qcache`) → **34 ms steady state**.
|
||||
- **This 34 ms is the lexical/spread path, NOT the ANN seed path.** The store's 4,190
|
||||
embeddings were generated by the live neuron's native embedding model; the harness's
|
||||
`nomic-embed-text` query vectors are a **different vector space**, so no candidate
|
||||
cleared `ENGRAM_EMBED_SEED_MIN=0.60` (`act-stats`: `dup_seeds:0`, `ctx_cos:-2.000`
|
||||
sentinel) → results empty, ANN discovery produced no admitted seeds. Environmental
|
||||
(embedding provenance), **not** an M8 defect. The ANN wiring still *executed*
|
||||
(embed fetch succeeded, `embed_breaker_open:0`; `eg_vindex_sync` + seed block +
|
||||
`free(e_eff)` all ran) **without crashing** on the real store.
|
||||
|
||||
## KEY M8 METRIC — ANN vs O(n) seed selection (authoritative)
|
||||
|
||||
Because the HTTP path can't exercise ANN seeding without embedding-space parity, the
|
||||
authoritative ANN-vs-exact-scan numbers come from the **module benchmark**
|
||||
(`engram/test/run_vindex_tests.sh`, PASS 1, optimised), which measures the exact
|
||||
`vindex_search` code the M8 wiring calls, at full size:
|
||||
|
||||
| N (768-dim vectors) | Brute-force (O(n)) | ANN (HNSW) | **Speedup** |
|
||||
|---|---|---|---|
|
||||
| 5,000 | 3.475 ms/query | 0.353 ms/query | **9.8×** |
|
||||
| 20,000 | 13.809 ms/query | 0.698 ms/query | **19.8×** |
|
||||
|
||||
- **recall@10 = 0.9365** at `ef_search=128` (gate ≥0.90 — **PASS**). Lower ef trades
|
||||
recall for latency: ef=64→0.844, ef=32→0.750, ef=10→0.625.
|
||||
- **Determinism:** two independent seeded builds give byte-identical query results.
|
||||
- **`vindex_build_from_store`** over a real `engram_store`: inserts exactly the
|
||||
embedded nodes, top-1 resolves to the correct node id at ~0 distance.
|
||||
- **HNSW build cost (single-threaded, note for boot/index-build budgeting):**
|
||||
5,000 vectors ≈ 15–35 s, 20,000 vectors ≈ 75 s. In-process the index is built
|
||||
**lazily on first activation** (`eg_vindex_sync`) and grown incrementally; the
|
||||
M8 seed block only fires once `vindex_size ≥ ENGRAM_EMBED_SEED_K`. At the real
|
||||
store's 4,190 embedded nodes this is a **one-time few-second first-activation
|
||||
cost** — worth watching as the embedded set grows (a future milestone may want
|
||||
to build the index at boot or persist it via `vindex_save`/`vindex_load`).
|
||||
|
||||
## Seed-set parity note (why there is no runtime A/B toggle)
|
||||
|
||||
M8 has **no ANN on/off env flag** by design (`ENGRAM_EMBED_SEED_K` is a compile-time
|
||||
constant). The exact O(n) cosine scan is **preserved verbatim** and "tops up" any seed
|
||||
slot the ANN leaves unfilled; every ANN candidate is admitted through the *identical*
|
||||
cosine/dedup/threshold gate the exact scan uses. So ANN changes only *which nodes are
|
||||
discovered and how fast*, never the final seed set — parity is **structural**, not
|
||||
A/B-tested via flag. Recovered-behaviour-when-index-absent is the pre-M8 exact scan.
|
||||
|
||||
## Integration Findings (first full build of the tiered trunk)
|
||||
|
||||
1. **CRITICAL / pre-existing (NOT M8): `btree_insert` stack-buffer-overflow on
|
||||
opening the live 458 MB `neuron.egm`.** `SIGABRT` (`__stack_chk_fail`) via
|
||||
`btree_insert ← btree_put ← edge_place/apply_edge_put ← engram_open ←
|
||||
engram_store_boot`. Reproduces on the M8 binary **and** the deployed live binary —
|
||||
**the live engram :8742 was crash-looping (18 crash reports 16:18→17:44 on this
|
||||
date; service refusing connections).** A clean import into a fresh 59.9 MB egm does
|
||||
**not** crash (43k edges load fine), so the trigger is the specific pathological
|
||||
live store: 458 MB (8× the healthy 59.9 MB) from a day of churn/tombstones with no
|
||||
M5 compaction + a 44 MB un-checkpointed WAL replayed on open. The bug is in
|
||||
`engram_store.c` (the edge B-tree / WAL-redo path), **upstream of everything M8
|
||||
touched** (M8 lives in `el_runtime.c::engram_activate`). Fix required before any
|
||||
re-cutover; `btree_insert` must bound-check regardless of on-disk content.
|
||||
2. **M8 deferred `free(e_eff)`** (the flagged memory-management concern): compiled
|
||||
warning-free, and the wired path executed end-to-end over HTTP on the real store
|
||||
(with a real query embedding) with **no crash / no new crash report** — no
|
||||
double-free or use-after-free observed. Freed on all early-return paths and exactly
|
||||
once post-seed-selection.
|
||||
3. **Write durability + retrieval-fix dedup**: create → checkpoint → clean SIGTERM →
|
||||
restart → node found **by id and by search** (node_count 13,036→13,037 preserved).
|
||||
|
||||
## Caveats
|
||||
|
||||
- All on a copy; healthy-substrate numbers are from a re-imported store, not the live
|
||||
paged store (which is currently un-bootable — Finding 1).
|
||||
- Single-run metrics; no thermal control.
|
||||
- Embedding-space mismatch prevented a real semantic `q=bullshit` activation in this
|
||||
harness; the ANN speedup number is the module benchmark, which is the correct gauge.
|
||||
Reference in New Issue
Block a user