Files
neuron/docs/architecture/03-data-and-memory.md
T
will.anderson 4bff40fa4a
Neuron Soul CI / build (pull_request) Failing after 14m5s
Neuron Soul CI / deploy (pull_request) Has been skipped
fix(api): bound inspect_graph with relevance-ranked projection; regen soul.c
High-fanout identity anchors (voice, writing-imprint, self-root) have ~670KB
neighborhoods. inspect_graph returned the full traversal, which overflowed the
MCP client's context and socket-closed the wrapper mid self-load -- the soul
could not traverse its own identity graph.

handle_api_inspect_graph gains an opt-in `compact` projection (compact=1|true):
the neighborhood is relevance-ranked, the top K (default 12) keep a UTF-8-safe
content snippet (default snip=600), and the remainder collapse to lightweight
{id,label,node_type,tier,edge,pointer:true} stubs. This bounds the voice node
from 669,799B -> 25,353B (HTTP 200, valid JSON) and the wrapper's soul-load no
longer socket-closes. New helpers: api_compact_neighbors, api_neigh_full,
api_neigh_pointer, api_neigh_rank, api_neigh_better, api_float_or.

The flag is gated: ABSENT it, the response is byte-identical to the old plain
traversal, so the studio app (which never sends it) is unaffected. The MCP
wrapper (mcp-wrapper/src/main.el) appends &compact=1 on its inspectGraph and
fetch-by-id paths.

dist/soul.c is REGENERATED so CI ships the fix: CI compiles the committed
single-TU dist/soul.c directly (running elb/elc on the Linux runner OOM-kills
it), so an .el-only change would build the OLD behavior. Regenerated and verified
on macOS -- compiles with the CI cc line (0 errors) and, on a throwaway soul over
a copy of the live snapshot, serves compact ~25KB / non-compact ~670KB. The regen
also syncs the amalgamation to this branch's .el sources, which had drifted
several self-review commits ahead of the previously-committed soul.c.

Docs: docs/architecture/00-05 added; 01/02/05 corrected so the relevance-ranked
inspect_graph projection reads as committed source, not an in-flight concern.
2026-08-10 10:28:50 -05:00

13 KiB
Raw Blame History

Neuron — Data & Memory (the Engram Graph Model)

The engram is neuron's durable substrate. This document describes the graph model: node/edge structure, the consciousness layers, the two distinct tier systems, write-protection, the tombstone/supersede immutability model, and persistence. Sources: the runtime el_runtime.c (where the graph engine physically lives — "the runtime IS the database", foundation/el/engram/src/server.el:1-6), the engram HTTP face server.el, and the neuron-layer semantics in memory.el / neuron-api.el.

Runtime path analyzed: foundation/el/lang/releases/v1.0.0-20260501/el_runtime.c.

Where the model lives

The engram is not a database library. The graph, the activation math, and Hebbian learning are compiled C in el_runtime.c; server.el is a thin HTTP server that exposes them on :8742; the storage format is a single JSON snapshot. There is no SQL, no SQLite, no append log. Keep this in mind: the "schema" below is C structs, not tables.

Design-doc caveat. engram/README.md describes a Rust/sled/bincode EngramDb with a NodeType::Concept enum. That is aspirational/legacy narrative — it does not match the shipped C engine. Treat the README as design story, not as the implementation. (unverified against runtime)

Nodes

EngramNodeel_runtime.c:5958-6018+. Every node carries:

Field group Fields Notes
Identity/content id, content, node_type, label, tier, tags, metadata all char* (:5959-5965)
Epistemic weights salience, importance, confidence (double), temporal_decay_rate per-node decay λ override; 0 = use global (:5966-5969)
Access history activation_count, last_activated, created_at, updated_at :5970-5973
Two-layer activation background_activation (Layer 1, BFS fan-out), working_memory_weight (Layer 2, executive filter), suppression_count context compilation uses only working_memory_weight (:5974-5991)
Consciousness layer layer_id default 1 = CORE_IDENTITY (:5996)
ACT-R learning access_ts[K] ring buffer, access_head, access_filled, wm_anchor base-level learning (:5997-6008)
Semantics emb (768-dim nomic-embed-text vector, lazily backfilled), emb_dim :6009-6016
Hebbian eligibility trace :6017+

Node types are strings, not an enum

node_type is a free char*, defaulting to "Memory" when unset (el_runtime.c:7401, server.el:159). There is no closed node-type enum in the shipped engine. Two consequences:

  • The runtime special-cases a handful of type strings for activation thresholds (engram_type_threshold, :5933-5955): DharmaSelf/Safety (0.05, fire easily), Belief/Entity (0.30), Knowledge (0.20), everything else Note/Memory/Working (0.40). InternalStateEvent and Tag are excluded from working-memory promotion (:6674-6676, :7368-7370).
  • Type strings the neuron layer actually writes: Memory (default), Knowledge (server.el:549), InternalStateEvent (server.el:493), Tombstone (memory.el:55), Conversation (session nodes, sessions.el), Persona (soul.el:250-292), plus identity/value Knowledge nodes.

The types the MCP surface names — Self, BacklogItem, SessionSummary, Artifact, Process, ConfigEntry — are node_type string conventions set by higher neuron/Axon layers, not runtime-known types. Where BacklogItem / Artifact are set was not in the files read (they route to the Axon backend, doc 02) — flag as unverified/TODO for a human pass.

Edges

EngramEdgeel_runtime.c:6701-6730+. Directed, typed, weighted:

Field Meaning
id, from_id, to_id, relation typed relation string
weight (double) authored strength — never mutated by activation
hebb (double) learned co-activation potentiation — the fraction of recent activations in which both endpoints were in working memory together; strictly separate from weight
inhibitory (int flag) if set, activating the source suppresses the target's WM weight instead of exciting it
confidence, created_at, updated_at, last_fired, layer

The hebb field is the co-activation weight — the Hebbian/LTP channel — kept deliberately separate from the static authored weight. Edges are created via engram_connect(from, to, weight, relation) (server.el:253).

Relation strings observed: associates (default, server.el:248), identity, co-value, birthday-twin, canonical-self (soul.el:37-108), supersedes, tombstones, contains, tagged (neuron-api.el, el_runtime.c:6168).

Consciousness layers

Orthogonal to memory tiers, the engram has five canonical layers (el_runtime.c:5919-5924):

id Name activation_priority Role
0 SAFETY 0 (fires earliest) deepest / limbic
1 CORE_IDENTITY default for all nodes (ENGRAM_LAYER_DEFAULT, :7423)
2 DOMAIN domain knowledge
3 IMPRINT persona overlay
4 SUIT outermost

EngramLayer (:6731-6738) carries activation_priority (lower fires first), suppressible (can higher layers suppress it?), transparent (invisible to introspection?), and injectable (add/remove at runtime?). Layers are managed via engram_add_layer / engram_node_layered / engram_list_layers. This is the identity-vs-domain-knowledge stratification, independent of the tier system below.

Two tier systems — do not conflate them

This is the single most important clarification in the data model, and the source of the vocabulary mismatch flagged throughout this set.

A. Cognitive memory tiers — the tier field

Working / Episodic / Semantic / Procedural (and Canonical in use). Runtime default "Working" (el_runtime.c:7408; README.md:41-49). Nodes migrate between these by salience decay/reinforcement, driven by the runtime. Salience decays as importance × 1/(1 + days_since) × ln(count + 1) (README.md:57-62). memory.el exposes tier_working/episodic/canonical helpers (memory.el:1-3); soul.el writes Semantic-tier persona nodes (:267, :282). So the live tier set is {Working, Episodic, Semantic, Procedural, Canonical} with continuous salience/importance/confidence floats.

B. Epistemic tiers & disposition — tags, not runtime concepts

The MCP-facing vocabulary — tiers note → lesson → canonical, disposition experimental → provisional → stable → deprecated — is not enforced anywhere in el_runtime.c. It is stored as tags:

  • Knowledge capture preserves the incoming epistemic tier as a tier:<x> tag rather than mapping onto a cognitive tier — deliberately, to avoid a lossy mapping (server.el:519-522, 544).
  • promote_knowledge writes a canonical node tagged ["Knowledge","tier:canonical","disposition:stable"] (neuron-api.el:533).

There is no state machine validating experimental → … → deprecated. Disposition and epistemic tier are convention-by-tag. (Flag: not structurally guarded. The exact MCP-enum → tag/float mapping is not fully traced in the files read — unverified/TODO.)

Write-protection

is_protected_node(id) (neuron-api.el:20-37) is a hard-coded allowlist of 15 identity/value node IDs — the self root, the values hub, intellectual-dna, memory-philosophy, voice, and the 8 value nodes. Handlers that could mutate the graph (tombstone / supersede / evolve / connect) check it and return HTTP 403 api_err_protected (:39-41) for a protected target (checked at :384, 511, 692, 705, 746, 768). Edges into a protected node are also blocked (handle_api_link_entities).

The one sanctioned override is POST /api/neuron/cultivate (neuron-api.el:781-816) — it performs the same ops with the protection check skipped, gated by convention to Will's explicit cultivation sessions. The self layer is writable, but only through a deliberate door.

Immutability — tombstone, never delete

Engram nodes are immutable (memory.el:64-69). The model is:

  • Tombstonemem_tombstone(node_id) (memory.el:46-71) keeps the node and all its edges, creates a Tombstone marker node (content = target id, label = "tombstone:<id>") and wires a tombstones edge (weight 1.0). It never calls engram_forget. This is the one canonical delete — every user-facing forget path routes through it. Default bounded reads hide tombstoned nodes (memory_hide_tombstoned, neuron-api.el:239-249); ?include_deleted=1 recovers them.
  • Supersede — updates/evolves (neuron-api.el:394-428, 506-541, 715-734) create a new node with the new content, wire a supersedes edge new→old (weight 0.9, or 0.95 for promote), and keep the original. The response returns both ids so the caller re-points. This is the supersedes_id pattern: new node linked, old preserved, full audit trail.

The hole to know about. The raw runtime engram_forget does hard-delete (frees node + edges, el_runtime.c:7647), and the engram HTTP route DELETE /api/nodes/:id calls it directly (server.el:322-328). Immutability is therefore an invariant of the neuron-api / MCP layer routing, not of the store. A client that hits engram HTTP directly can bypass it. (flag)

engram_forget is also used internally for genuine GC: boot-counter pruning (memory.el:184), session-summary/telemetry pruning (soul.el:369, sessions.el). Those are bounded housekeeping, not user deletes.

Persistence, snapshots, backups

  • Storage: a single JSON snapshot snapshot.json under ENGRAM_DATA_DIR, written by engram_save / read by engram_load (el_runtime.c:9660+; format {"nodes":[...],"edges":[...]}). In prod that dir is the RWO PVC mount /data (doc 04).
  • Write policy: persist_canonical() writes the full snapshot after every durable write (server.el:133-141). The batch-edge route snapshots once per batch to avoid ~150 GB/day of writes from Hebbian edge churn (server.el:258-305) — this is why hebb_consolidate batches (doc 02).
  • Boot safety: on load, engram writes snapshot.boot-backup.json (good load) or snapshot.failed-load.json (a non-empty file that parsed to 0 nodes) (server.el:718-734). Read routes export to scratch paths (.scan-export.json, .sync-export.json) and never touch the canonical (server.el:207-223, 418-437) — a guard added after a read-route corrupted the snapshot.
  • Off-cluster backup: a Kubernetes CronJob (engram-backup) tars /data every 15 minutes to gs://neuron-db-backup/gke/neuron-prod/ and keeps the last 96 (24h) (infrastructure/platform/k8s/neuron-mcp/backup-cronjob.yaml).
  • Retention: InternalStateEvent telemetry pruned at 48h (ENGRAM_ISE_RETENTION_MS, server.el:485-499).

Data-dir mismatch to flag: the server.el header comment says the default is ~/.neuron/engram (:16) but the code defaults to /tmp/engram (:135, 717). Prod overrides both via ENGRAM_DATA_DIR=/data. (unverified — which default is intended)

The engram HTTP surface (:8742)

Dispatcher handle_request (server.el:592-707). Auth: ENGRAM_API_KEY; GETs always allowed, mutations require "_auth":"<key>" in the JSON body (server.el:578-588).

Endpoint Purpose
GET /health, GET / health + live node/edge counts
POST /api/nodes, GET /api/nodes, GET /api/nodes/:id, DELETE /api/nodes/:id node CRUD (DELETE = hard engram_forget)
GET /api/edges, POST /api/edges, POST /api/edges/batch, GET /api/neighbors/:id?depth edge ops + traversal
POST|GET /api/activate?q&depth, POST|GET /api/search spreading activation vs lexical search
POST /api/strengthen Hebbian potentiation
POST /api/save, /api/load, /api/load-merge snapshot control
GET /api/sync soul daemon periodic pull
GET /api/embed-backfill, GET /api/similarity?a&b embeddings + cosine
POST /api/neuron/state-events (auth-exempt), POST /api/neuron/knowledge/capture neuron-layer helpers
GET /api/stats, /api/act-stats, /api/text-health telemetry

Retrieval model (summary)

Retrieval is spreading activation, not query matching: strength = parent_strength × edge_weight × target_salience × cosine(query, target) — multiplicative, top-N, with the two-layer background → working-memory promotion (README.md:27-36; el_runtime.c:5892+, 6094+). mem_recall / /api/activate fire this and mutate WM; mem_search / /api/search are passive lexical scans. The cognitive API's begin_session and compile_ctx return a bounded projection of the activated set, never the raw graph (doc 02, §2).