High-fanout identity anchors (voice, writing-imprint, self-root) have ~670KB
neighborhoods. inspect_graph returned the full traversal, which overflowed the
MCP client's context and socket-closed the wrapper mid self-load -- the soul
could not traverse its own identity graph.
handle_api_inspect_graph gains an opt-in `compact` projection (compact=1|true):
the neighborhood is relevance-ranked, the top K (default 12) keep a UTF-8-safe
content snippet (default snip=600), and the remainder collapse to lightweight
{id,label,node_type,tier,edge,pointer:true} stubs. This bounds the voice node
from 669,799B -> 25,353B (HTTP 200, valid JSON) and the wrapper's soul-load no
longer socket-closes. New helpers: api_compact_neighbors, api_neigh_full,
api_neigh_pointer, api_neigh_rank, api_neigh_better, api_float_or.
The flag is gated: ABSENT it, the response is byte-identical to the old plain
traversal, so the studio app (which never sends it) is unaffected. The MCP
wrapper (mcp-wrapper/src/main.el) appends &compact=1 on its inspectGraph and
fetch-by-id paths.
dist/soul.c is REGENERATED so CI ships the fix: CI compiles the committed
single-TU dist/soul.c directly (running elb/elc on the Linux runner OOM-kills
it), so an .el-only change would build the OLD behavior. Regenerated and verified
on macOS -- compiles with the CI cc line (0 errors) and, on a throwaway soul over
a copy of the live snapshot, serves compact ~25KB / non-compact ~670KB. The regen
also syncs the amalgamation to this branch's .el sources, which had drifted
several self-review commits ahead of the previously-committed soul.c.
Docs: docs/architecture/00-05 added; 01/02/05 corrected so the relevance-ranked
inspect_graph projection reads as committed source, not an in-flight concern.
13 KiB
Neuron — Data & Memory (the Engram Graph Model)
The engram is neuron's durable substrate. This document describes the graph model: node/edge structure, the consciousness layers, the two distinct tier systems, write-protection, the tombstone/supersede immutability model, and persistence. Sources: the runtime
el_runtime.c(where the graph engine physically lives — "the runtime IS the database",foundation/el/engram/src/server.el:1-6), the engram HTTP faceserver.el, and the neuron-layer semantics inmemory.el/neuron-api.el.Runtime path analyzed:
foundation/el/lang/releases/v1.0.0-20260501/el_runtime.c.
Where the model lives
The engram is not a database library. The graph, the activation math, and
Hebbian learning are compiled C in el_runtime.c; server.el is a thin HTTP
server that exposes them on :8742; the storage format is a single JSON
snapshot. There is no SQL, no SQLite, no append log. Keep this in mind: the
"schema" below is C structs, not tables.
Design-doc caveat.
engram/README.mddescribes a Rust/sled/bincodeEngramDbwith aNodeType::Conceptenum. That is aspirational/legacy narrative — it does not match the shipped C engine. Treat the README as design story, not as the implementation. (unverified against runtime)
Nodes
EngramNode — el_runtime.c:5958-6018+. Every node carries:
| Field group | Fields | Notes |
|---|---|---|
| Identity/content | id, content, node_type, label, tier, tags, metadata |
all char* (:5959-5965) |
| Epistemic weights | salience, importance, confidence (double), temporal_decay_rate |
per-node decay λ override; 0 = use global (:5966-5969) |
| Access history | activation_count, last_activated, created_at, updated_at |
:5970-5973 |
| Two-layer activation | background_activation (Layer 1, BFS fan-out), working_memory_weight (Layer 2, executive filter), suppression_count |
context compilation uses only working_memory_weight (:5974-5991) |
| Consciousness layer | layer_id |
default 1 = CORE_IDENTITY (:5996) |
| ACT-R learning | access_ts[K] ring buffer, access_head, access_filled, wm_anchor |
base-level learning (:5997-6008) |
| Semantics | emb (768-dim nomic-embed-text vector, lazily backfilled), emb_dim |
:6009-6016 |
| Hebbian | eligibility trace | :6017+ |
Node types are strings, not an enum
node_type is a free char*, defaulting to "Memory" when unset
(el_runtime.c:7401, server.el:159). There is no closed node-type enum in
the shipped engine. Two consequences:
- The runtime special-cases a handful of type strings for activation
thresholds (
engram_type_threshold,:5933-5955):DharmaSelf/Safety(0.05, fire easily),Belief/Entity(0.30),Knowledge(0.20), everything elseNote/Memory/Working(0.40).InternalStateEventandTagare excluded from working-memory promotion (:6674-6676,:7368-7370). - Type strings the neuron layer actually writes:
Memory(default),Knowledge(server.el:549),InternalStateEvent(server.el:493),Tombstone(memory.el:55),Conversation(session nodes,sessions.el),Persona(soul.el:250-292), plus identity/valueKnowledgenodes.
The types the MCP surface names — Self, BacklogItem, SessionSummary,
Artifact, Process, ConfigEntry — are node_type string conventions set
by higher neuron/Axon layers, not runtime-known types. Where BacklogItem /
Artifact are set was not in the files read (they route to the Axon backend, doc
02) — flag as unverified/TODO for a human pass.
Edges
EngramEdge — el_runtime.c:6701-6730+. Directed, typed, weighted:
| Field | Meaning |
|---|---|
id, from_id, to_id, relation |
typed relation string |
weight (double) |
authored strength — never mutated by activation |
hebb (double) |
learned co-activation potentiation — the fraction of recent activations in which both endpoints were in working memory together; strictly separate from weight |
inhibitory (int flag) |
if set, activating the source suppresses the target's WM weight instead of exciting it |
confidence, created_at, updated_at, last_fired, layer |
— |
The hebb field is the co-activation weight — the Hebbian/LTP channel — kept
deliberately separate from the static authored weight. Edges are created via
engram_connect(from, to, weight, relation) (server.el:253).
Relation strings observed: associates (default, server.el:248),
identity, co-value, birthday-twin, canonical-self (soul.el:37-108),
supersedes, tombstones, contains, tagged (neuron-api.el,
el_runtime.c:6168).
Consciousness layers
Orthogonal to memory tiers, the engram has five canonical layers
(el_runtime.c:5919-5924):
| id | Name | activation_priority | Role |
|---|---|---|---|
| 0 | SAFETY | 0 (fires earliest) | deepest / limbic |
| 1 | CORE_IDENTITY | — | default for all nodes (ENGRAM_LAYER_DEFAULT, :7423) |
| 2 | DOMAIN | — | domain knowledge |
| 3 | IMPRINT | — | persona overlay |
| 4 | SUIT | — | outermost |
EngramLayer (:6731-6738) carries activation_priority (lower fires first),
suppressible (can higher layers suppress it?), transparent (invisible to
introspection?), and injectable (add/remove at runtime?). Layers are managed
via engram_add_layer / engram_node_layered / engram_list_layers. This is
the identity-vs-domain-knowledge stratification, independent of the tier system
below.
Two tier systems — do not conflate them
This is the single most important clarification in the data model, and the source of the vocabulary mismatch flagged throughout this set.
A. Cognitive memory tiers — the tier field
Working / Episodic / Semantic / Procedural (and Canonical in use).
Runtime default "Working" (el_runtime.c:7408; README.md:41-49). Nodes
migrate between these by salience decay/reinforcement, driven by the runtime.
Salience decays as importance × 1/(1 + days_since) × ln(count + 1)
(README.md:57-62). memory.el exposes tier_working/episodic/canonical
helpers (memory.el:1-3); soul.el writes Semantic-tier persona nodes
(:267, :282). So the live tier set is {Working, Episodic, Semantic,
Procedural, Canonical} with continuous salience/importance/confidence floats.
B. Epistemic tiers & disposition — tags, not runtime concepts
The MCP-facing vocabulary — tiers note → lesson → canonical, disposition
experimental → provisional → stable → deprecated — is not enforced anywhere
in el_runtime.c. It is stored as tags:
- Knowledge capture preserves the incoming epistemic tier as a
tier:<x>tag rather than mapping onto a cognitive tier — deliberately, to avoid a lossy mapping (server.el:519-522, 544). promote_knowledgewrites a canonical node tagged["Knowledge","tier:canonical","disposition:stable"](neuron-api.el:533).
There is no state machine validating experimental → … → deprecated.
Disposition and epistemic tier are convention-by-tag. (Flag: not structurally
guarded. The exact MCP-enum → tag/float mapping is not fully traced in the files
read — unverified/TODO.)
Write-protection
is_protected_node(id) (neuron-api.el:20-37) is a hard-coded allowlist of 15
identity/value node IDs — the self root, the values hub, intellectual-dna,
memory-philosophy, voice, and the 8 value nodes. Handlers that could mutate the
graph (tombstone / supersede / evolve / connect) check it and return HTTP 403
api_err_protected (:39-41) for a protected target (checked at :384, 511, 692, 705, 746, 768). Edges into a protected node are also blocked
(handle_api_link_entities).
The one sanctioned override is POST /api/neuron/cultivate
(neuron-api.el:781-816) — it performs the same ops with the protection check
skipped, gated by convention to Will's explicit cultivation sessions. The self
layer is writable, but only through a deliberate door.
Immutability — tombstone, never delete
Engram nodes are immutable (memory.el:64-69). The model is:
- Tombstone —
mem_tombstone(node_id)(memory.el:46-71) keeps the node and all its edges, creates aTombstonemarker node (content = target id,label = "tombstone:<id>") and wires atombstonesedge (weight 1.0). It never callsengram_forget. This is the one canonical delete — every user-facing forget path routes through it. Default bounded reads hide tombstoned nodes (memory_hide_tombstoned,neuron-api.el:239-249);?include_deleted=1recovers them. - Supersede — updates/evolves (
neuron-api.el:394-428, 506-541, 715-734) create a new node with the new content, wire asupersedesedge new→old (weight 0.9, or 0.95 for promote), and keep the original. The response returns both ids so the caller re-points. This is thesupersedes_idpattern: new node linked, old preserved, full audit trail.
The hole to know about. The raw runtime
engram_forgetdoes hard-delete (frees node + edges,el_runtime.c:7647), and the engram HTTP routeDELETE /api/nodes/:idcalls it directly (server.el:322-328). Immutability is therefore an invariant of the neuron-api / MCP layer routing, not of the store. A client that hits engram HTTP directly can bypass it. (flag)
engram_forget is also used internally for genuine GC: boot-counter pruning
(memory.el:184), session-summary/telemetry pruning (soul.el:369,
sessions.el). Those are bounded housekeeping, not user deletes.
Persistence, snapshots, backups
- Storage: a single JSON snapshot
snapshot.jsonunderENGRAM_DATA_DIR, written byengram_save/ read byengram_load(el_runtime.c:9660+; format{"nodes":[...],"edges":[...]}). In prod that dir is the RWO PVC mount/data(doc 04). - Write policy:
persist_canonical()writes the full snapshot after every durable write (server.el:133-141). The batch-edge route snapshots once per batch to avoid ~150 GB/day of writes from Hebbian edge churn (server.el:258-305) — this is whyhebb_consolidatebatches (doc 02). - Boot safety: on load, engram writes
snapshot.boot-backup.json(good load) orsnapshot.failed-load.json(a non-empty file that parsed to 0 nodes) (server.el:718-734). Read routes export to scratch paths (.scan-export.json,.sync-export.json) and never touch the canonical (server.el:207-223, 418-437) — a guard added after a read-route corrupted the snapshot. - Off-cluster backup: a Kubernetes CronJob (
engram-backup) tars/dataevery 15 minutes togs://neuron-db-backup/gke/neuron-prod/and keeps the last 96 (24h) (infrastructure/platform/k8s/neuron-mcp/backup-cronjob.yaml). - Retention: InternalStateEvent telemetry pruned at 48h
(
ENGRAM_ISE_RETENTION_MS,server.el:485-499).
Data-dir mismatch to flag: the
server.elheader comment says the default is~/.neuron/engram(:16) but the code defaults to/tmp/engram(:135, 717). Prod overrides both viaENGRAM_DATA_DIR=/data. (unverified — which default is intended)
The engram HTTP surface (:8742)
Dispatcher handle_request (server.el:592-707). Auth: ENGRAM_API_KEY; GETs
always allowed, mutations require "_auth":"<key>" in the JSON body
(server.el:578-588).
| Endpoint | Purpose |
|---|---|
GET /health, GET / |
health + live node/edge counts |
POST /api/nodes, GET /api/nodes, GET /api/nodes/:id, DELETE /api/nodes/:id |
node CRUD (DELETE = hard engram_forget) |
GET /api/edges, POST /api/edges, POST /api/edges/batch, GET /api/neighbors/:id?depth |
edge ops + traversal |
POST|GET /api/activate?q&depth, POST|GET /api/search |
spreading activation vs lexical search |
POST /api/strengthen |
Hebbian potentiation |
POST /api/save, /api/load, /api/load-merge |
snapshot control |
GET /api/sync |
soul daemon periodic pull |
GET /api/embed-backfill, GET /api/similarity?a&b |
embeddings + cosine |
POST /api/neuron/state-events (auth-exempt), POST /api/neuron/knowledge/capture |
neuron-layer helpers |
GET /api/stats, /api/act-stats, /api/text-health |
telemetry |
Retrieval model (summary)
Retrieval is spreading activation, not query matching:
strength = parent_strength × edge_weight × target_salience × cosine(query, target) — multiplicative, top-N, with the two-layer
background → working-memory promotion (README.md:27-36; el_runtime.c:5892+, 6094+). mem_recall / /api/activate fire this and mutate WM; mem_search /
/api/search are passive lexical scans. The cognitive API's begin_session and
compile_ctx return a bounded projection of the activated set, never the raw
graph (doc 02, §2).