Files
neuron/docs/architecture/03-data-and-memory.md
T
will.anderson 5e15d90659
Neuron Soul CI / build (pull_request) Has been cancelled
Neuron Soul CI / deploy (pull_request) Has been cancelled
docs(architecture): record the 2026-08-14 deep-night sessions
~226 lines of architecture documentation that were written, left uncommitted in
the working tree, and nearly lost. None of it was on main. Recovered from a
stash taken while merging tonight's fixes.

Substantive content, not notes:

- Peer import-of-understanding PROVEN by execution. A exported a skill-geometry;
  on the receiver `think` went "geometry unavailable" -> operable. Cosine 1.0 on
  both the raw-geometry and text/dharma-bus transports, bidirectional. The
  mind-not-paste evidence: n_support 27 on source vs 3 on receiver, i.e. the
  imported geometry wires into the host manifold rather than sitting inert.
  Honest boundary recorded too: proven between forks sharing one embedder,
  UNTESTED cross-embedder.

- "Teacher" renamed GUIDE, and the distinction is load-bearing: its output is
  grounded/verified before trust. A teacher you believe; a guide you check.

- Layers are named persistent relational neighborhoods, not storage tiers, with
  their own growth and threshold-lock policy (note->canonical maturation lifted
  from single nodes to a region).

- The consciousness theories (Global Workspace, IIT's Phi, attention-schema,
  higher-order thought, active inference, interoception) read as geometric
  LENSES over one manifold rather than competing mechanisms. Functional problems
  fall out; the hard problem explicitly not claimed solved.

- Growth is bounded/logistic, not geometric — exponential growth is the cancer
  shape. Two-rate discipline: explore fast in local geometry, grow the engram
  slowly by verifier-gated merge.

- Orchestration as a geometric operation: critical path as geodesic, float as
  displacement, @manager compiles the work-graph. Single-writer enforced by
  capability (Rule 4).

- The decorated seam, the API surface collapse to geometry ops, and the
  distributed-self thesis — each tiered honestly against what is actually proven
  vs staged vs unbuilt.

Also gitignores dist-fresh/ (regenerate scratch dir, a build artifact).

Not included from the same stash: awareness.elh and dist/elp-c-decls.h, which
are generated artifacts now gitignored per #154/#158.
2026-08-15 19:39:50 -05:00

16 KiB
Raw Blame History

Neuron — Data & Memory (the Engram Graph Model)

The engram is neuron's durable substrate. This document describes the graph model: node/edge structure, the consciousness layers, the two distinct tier systems, write-protection, the tombstone/supersede immutability model, and persistence. Sources: the runtime el_runtime.c (where the graph engine physically lives — "the runtime IS the database", foundation/el/engram/src/server.el:1-6), the engram HTTP face server.el, and the neuron-layer semantics in memory.el / neuron-api.el.

Runtime path analyzed: foundation/el/lang/releases/v1.0.0-20260501/el_runtime.c.

Where the model lives

The engram is not a database library. The graph, the activation math, and Hebbian learning are compiled C in el_runtime.c; server.el is a thin HTTP server that exposes them on :8742; the storage format is a single JSON snapshot. There is no SQL, no SQLite, no append log. Keep this in mind: the "schema" below is C structs, not tables.

Design-doc caveat. engram/README.md describes a Rust/sled/bincode EngramDb with a NodeType::Concept enum. That is aspirational/legacy narrative — it does not match the shipped C engine. Treat the README as design story, not as the implementation. (unverified against runtime)

Nodes

EngramNodeel_runtime.c:5958-6018+. Every node carries:

Field group Fields Notes
Identity/content id, content, node_type, label, tier, tags, metadata all char* (:5959-5965)
Epistemic weights salience, importance, confidence (double), temporal_decay_rate per-node decay λ override; 0 = use global (:5966-5969)
Access history activation_count, last_activated, created_at, updated_at :5970-5973
Two-layer activation background_activation (Layer 1, BFS fan-out), working_memory_weight (Layer 2, executive filter), suppression_count context compilation uses only working_memory_weight (:5974-5991)
Consciousness layer layer_id default 1 = CORE_IDENTITY (:5996)
ACT-R learning access_ts[K] ring buffer, access_head, access_filled, wm_anchor base-level learning (:5997-6008)
Semantics emb (768-dim nomic-embed-text vector, lazily backfilled), emb_dim :6009-6016
Hebbian eligibility trace :6017+

Node types are strings, not an enum

node_type is a free char*, defaulting to "Memory" when unset (el_runtime.c:7401, server.el:159). There is no closed node-type enum in the shipped engine. Two consequences:

  • The runtime special-cases a handful of type strings for activation thresholds (engram_type_threshold, :5933-5955): DharmaSelf/Safety (0.05, fire easily), Belief/Entity (0.30), Knowledge (0.20), everything else Note/Memory/Working (0.40). InternalStateEvent and Tag are excluded from working-memory promotion (:6674-6676, :7368-7370).
  • Type strings the neuron layer actually writes: Memory (default), Knowledge (server.el:549), InternalStateEvent (server.el:493), Tombstone (memory.el:55), Conversation (session nodes, sessions.el), Persona (soul.el:250-292), plus identity/value Knowledge nodes.

The types the MCP surface names — Self, BacklogItem, SessionSummary, Artifact, Process, ConfigEntry — are node_type string conventions set by higher neuron/Axon layers, not runtime-known types. Where BacklogItem / Artifact are set was not in the files read (they route to the Axon backend, doc 02) — flag as unverified/TODO for a human pass.

Edges

EngramEdgeel_runtime.c:6701-6730+. Directed, typed, weighted:

Field Meaning
id, from_id, to_id, relation typed relation string
weight (double) authored strength — never mutated by activation
hebb (double) learned co-activation potentiation — the fraction of recent activations in which both endpoints were in working memory together; strictly separate from weight
inhibitory (int flag) if set, activating the source suppresses the target's WM weight instead of exciting it
confidence, created_at, updated_at, last_fired, layer

The hebb field is the co-activation weight — the Hebbian/LTP channel — kept deliberately separate from the static authored weight. Edges are created via engram_connect(from, to, weight, relation) (server.el:253).

Relation strings observed: associates (default, server.el:248), identity, co-value, birthday-twin, canonical-self (soul.el:37-108), supersedes, tombstones, contains, tagged (neuron-api.el, el_runtime.c:6168).

Edges as vectors — the intended model (TARGET; today's edge is scalar). The live edge above carries a typed relation string plus scalar strength channels (weight, hebb). The design target is for an edge to be a vector — a first-class carrier of relationship-meaning in the node space — so that relationships can be composed / subtracted / analogized / traversed like nodes (the 06 §6 operator algebra over edges). Combined with append-only, this yields a complete temporal record: every discrete, significant change to a relationship is appended (a keyframe on material change), so the full 4-D trajectory of the meaning-manifold is preserved and recall_at(t) can read how any relationship was configured at any past t — bounded, because changes are discrete and meaning saturates by compositionality. Status: TARGET / #39 (see 07-storage-coherence-and-distribution.md §2.4); the runtime edge is scalar today.

Consciousness layers

Orthogonal to memory tiers, the engram has five canonical layers (el_runtime.c:5919-5924):

id Name activation_priority Role
0 SAFETY 0 (fires earliest) deepest / limbic
1 CORE_IDENTITY default for all nodes (ENGRAM_LAYER_DEFAULT, :7423)
2 DOMAIN domain knowledge
3 IMPRINT persona overlay
4 SUIT outermost

EngramLayer (:6731-6738) carries activation_priority (lower fires first), suppressible (can higher layers suppress it?), transparent (invisible to introspection?), and injectable (add/remove at runtime?). Layers are managed via engram_add_layer / engram_node_layered / engram_list_layers. This is the identity-vs-domain-knowledge stratification, independent of the tier system below.

Two tier systems — do not conflate them

This is the single most important clarification in the data model, and the source of the vocabulary mismatch flagged throughout this set.

A. Cognitive memory tiers — the tier field

Working / Episodic / Semantic / Procedural (and Canonical in use). Runtime default "Working" (el_runtime.c:7408; README.md:41-49). Nodes migrate between these by salience decay/reinforcement, driven by the runtime. Salience decays as importance × 1/(1 + days_since) × ln(count + 1) (README.md:57-62). memory.el exposes tier_working/episodic/canonical helpers (memory.el:1-3); soul.el writes Semantic-tier persona nodes (:267, :282). So the live tier set is {Working, Episodic, Semantic, Procedural, Canonical} with continuous salience/importance/confidence floats.

B. Epistemic tiers & disposition — tags, not runtime concepts

The MCP-facing vocabulary — tiers note → lesson → canonical, disposition experimental → provisional → stable → deprecated — is not enforced anywhere in el_runtime.c. It is stored as tags:

  • Knowledge capture preserves the incoming epistemic tier as a tier:<x> tag rather than mapping onto a cognitive tier — deliberately, to avoid a lossy mapping (server.el:519-522, 544).
  • promote_knowledge writes a canonical node tagged ["Knowledge","tier:canonical","disposition:stable"] (neuron-api.el:533).

There is no state machine validating experimental → … → deprecated. Disposition and epistemic tier are convention-by-tag. (Flag: not structurally guarded. The exact MCP-enum → tag/float mapping is not fully traced in the files read — unverified/TODO.)

Write-protection

is_protected_node(id) (neuron-api.el:20-37) is a hard-coded allowlist of 15 identity/value node IDs — the self root, the values hub, intellectual-dna, memory-philosophy, voice, and the 8 value nodes. Handlers that could mutate the graph (tombstone / supersede / evolve / connect) check it and return HTTP 403 api_err_protected (:39-41) for a protected target (checked at :384, 511, 692, 705, 746, 768). Edges into a protected node are also blocked (handle_api_link_entities).

The one sanctioned override is POST /api/neuron/cultivate (neuron-api.el:781-816) — it performs the same ops with the protection check skipped, gated by convention to Will's explicit cultivation sessions. The self layer is writable, but only through a deliberate door.

Immutability — tombstone, never delete

Engram nodes are immutable (memory.el:64-69). The model is:

  • Tombstonemem_tombstone(node_id) (memory.el:46-71) keeps the node and all its edges, creates a Tombstone marker node (content = target id, label = "tombstone:<id>") and wires a tombstones edge (weight 1.0). It never calls engram_forget. This is the one canonical delete — every user-facing forget path routes through it. Default bounded reads hide tombstoned nodes (memory_hide_tombstoned, neuron-api.el:239-249); ?include_deleted=1 recovers them.
  • Supersede — updates/evolves (neuron-api.el:394-428, 506-541, 715-734) create a new node with the new content, wire a supersedes edge new→old (weight 0.9, or 0.95 for promote), and keep the original. The response returns both ids so the caller re-points. This is the supersedes_id pattern: new node linked, old preserved, full audit trail.

Supersession is residue, not garbage. The superseded node is the trail of how the current understanding was reached — kept deliberately, because sometimes the truth was in the old idea even when the old idea was not itself the truth. This is what lets autonomous self-reification (06 §4.1) run ungated: every rename/re-cluster supersedes into this residue chain, so nothing it does is ever destructive — the safety is after the act, not a gate before it.

The hole to know about. The raw runtime engram_forget does hard-delete (frees node + edges, el_runtime.c:7647), and the engram HTTP route DELETE /api/nodes/:id calls it directly (server.el:322-328). Immutability is therefore an invariant of the neuron-api / MCP layer routing, not of the store. A client that hits engram HTTP directly can bypass it. (flag)

engram_forget is also used internally for genuine GC: boot-counter pruning (memory.el:184), session-summary/telemetry pruning (soul.el:369, sessions.el). Those are bounded housekeeping, not user deletes.

Persistence, snapshots, backups

  • Storage: a single JSON snapshot snapshot.json under ENGRAM_DATA_DIR, written by engram_save / read by engram_load (el_runtime.c:9660+; format {"nodes":[...],"edges":[...]}). In prod that dir is the RWO PVC mount /data (doc 04).
  • Write policy: persist_canonical() writes the full snapshot after every durable write (server.el:133-141). The batch-edge route snapshots once per batch to avoid ~150 GB/day of writes from Hebbian edge churn (server.el:258-305) — this is why hebb_consolidate batches (doc 02).
  • Boot safety: on load, engram writes snapshot.boot-backup.json (good load) or snapshot.failed-load.json (a non-empty file that parsed to 0 nodes) (server.el:718-734). Read routes export to scratch paths (.scan-export.json, .sync-export.json) and never touch the canonical (server.el:207-223, 418-437) — a guard added after a read-route corrupted the snapshot.
  • Off-cluster backup: a Kubernetes CronJob (engram-backup) tars /data every 15 minutes to gs://neuron-db-backup/gke/neuron-prod/ and keeps the last 96 (24h) (infrastructure/platform/k8s/neuron-mcp/backup-cronjob.yaml).
  • Retention: InternalStateEvent telemetry pruned at 48h (ENGRAM_ISE_RETENTION_MS, server.el:485-499).

Data-dir mismatch to flag: the server.el header comment says the default is ~/.neuron/engram (:16) but the code defaults to /tmp/engram (:135, 717). Prod overrides both via ENGRAM_DATA_DIR=/data. (unverified — which default is intended)

The engram HTTP surface (:8742)

Dispatcher handle_request (server.el:592-707). Auth: ENGRAM_API_KEY; GETs always allowed, mutations require "_auth":"<key>" in the JSON body (server.el:578-588).

Endpoint Purpose
GET /health, GET / health + live node/edge counts
POST /api/nodes, GET /api/nodes, GET /api/nodes/:id, DELETE /api/nodes/:id node CRUD (DELETE = hard engram_forget)
GET /api/edges, POST /api/edges, POST /api/edges/batch, GET /api/neighbors/:id?depth edge ops + traversal
POST|GET /api/activate?q&depth, POST|GET /api/search spreading activation vs lexical search
POST /api/strengthen Hebbian potentiation
POST /api/save, /api/load, /api/load-merge snapshot control
GET /api/sync soul daemon periodic pull
GET /api/embed-backfill, GET /api/similarity?a&b embeddings + cosine
POST /api/neuron/state-events (auth-exempt), POST /api/neuron/knowledge/capture neuron-layer helpers
GET /api/stats, /api/act-stats, /api/text-health telemetry

Retrieval model (summary)

Retrieval is spreading activation, not query matching: strength = parent_strength × edge_weight × target_salience × cosine(query, target) — multiplicative, top-N, with the two-layer background → working-memory promotion (README.md:27-36; el_runtime.c:5892+, 6094+). mem_recall / /api/activate fire this and mutate WM; mem_search / /api/search are passive lexical scans — but as of 2026-08-14 the live route_search runs structure-gated geometric retrieval (engram_retrieve_geometric_json; held-out P@5 = 0.700, semantic not lexical — skill returns skill nodes and rejects the false-positive rainfall), with the old lexical scan retained at /api/search-lexical (see 06 §2.5). The cognitive API's begin_session and compile_ctx return a bounded projection of the activated set, never the raw graph (doc 02, §2).

Update — 2026-08-14: layers as named neighborhoods (DESIGN; backlog #49)

A refinement of the ## Consciousness layers model above, from the deep-night session (node 92941631). A layer is not a storage tier — it is a named, persistent relational neighborhood in the one engram, each carrying its own growth policy and its own lock / threshold policy:

  • Threshold-lock = notecanonical maturation at neighborhood scale. The same epistemic-tier promotion the two-tier model (§B above) applies to a single node is lifted to a region: a neighborhood earns its lock by maturing past a threshold, at which point it stabilizes (read-mostly) the way a canonical node does. Growth and lock are per-neighborhood, not global.
  • A user's imprint is just another neighborhood. It is not a separate store or a bolted-on partition — it lives in the same geometry as everything else.
  • Relate-across is the advantage over island engrams. Because every neighborhood shares one geometry, anything can form edges to anything across neighborhood boundaries — the structural reason a single engram with named neighborhoods beats a set of isolated per-purpose stores.

Status: DESIGN. This is the intended model for engram layers; the naming, growth, and threshold-lock policies are not yet a built runtime feature. See 06-cognitive-architecture.md (Update — second pass).