Files
neuron/docs/architecture/03-data-and-memory.md
T
will.anderson 65dd2cf097
Neuron Soul CI / build (pull_request) Failing after 4m9s
Neuron Soul CI / deploy (pull_request) Has been skipped
docs: record the correspondence corrections — grounding, faculties, wonder, consolidation
The architecture docs describe four things the design spec has since ruled out,
and each one is a supervisor invented for something that should be a property of
the substrate: grounding modelled as a subsystem rather than as the edge weight
it already is; faculties modelled as parameters of a read when abduce is a write;
wonder materialized as a maintained manifest when it is the boundary of the
structure; and consolidation implemented eleven times behind tickers when a brain
has no cron job.

Left standing rather than deleted, per the repo's own supersession discipline —
the trail of how the understanding matured is the point. Each stale passage is
marked inline and points at a new 06 §12 that transcribes the corrections and
records the measured consolidation inventory.

Authority: foundation/el, branch design/correspondence-and-censorship,
lang/spec/correspondence-and-censorship.md.
2026-08-16 13:31:27 -05:00

329 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Neuron — Data & Memory (the Engram Graph Model)
> The engram is neuron's durable substrate. This document describes the graph
> model: node/edge structure, the consciousness layers, the two distinct tier
> systems, write-protection, the tombstone/supersede immutability model, and
> persistence. Sources: the runtime `el_runtime.c` (where the graph engine
> physically lives — "the runtime IS the database",
> `foundation/el/engram/src/server.el:1-6`), the engram HTTP face
> `server.el`, and the neuron-layer semantics in `memory.el` / `neuron-api.el`.
>
> Runtime path analyzed:
> `foundation/el/lang/releases/v1.0.0-20260501/el_runtime.c`.
## Where the model lives
The engram is **not** a database library. The graph, the activation math, and
Hebbian learning are compiled C in `el_runtime.c`; `server.el` is a thin HTTP
server that exposes them on `:8742`; the storage format is a single JSON
snapshot. There is no SQL, no SQLite, no append log. Keep this in mind: the
"schema" below is C structs, not tables.
> **Design-doc caveat.** `engram/README.md` describes a Rust/`sled`/`bincode`
> `EngramDb` with a `NodeType::Concept` enum. That is **aspirational/legacy
> narrative** — it does not match the shipped C engine. Treat the README as
> design story, not as the implementation. *(unverified against runtime)*
## Nodes
`EngramNode``el_runtime.c:5958-6018+`. Every node carries:
| Field group | Fields | Notes |
|---|---|---|
| Identity/content | `id`, `content`, `node_type`, `label`, `tier`, `tags`, `metadata` | all `char*` (`:5959-5965`) |
| Epistemic weights | `salience`, `importance`, `confidence` (double), `temporal_decay_rate` | per-node decay λ override; 0 = use global (`:5966-5969`) |
| Access history | `activation_count`, `last_activated`, `created_at`, `updated_at` | `:5970-5973` |
| Two-layer activation | `background_activation` (Layer 1, BFS fan-out), `working_memory_weight` (Layer 2, executive filter), `suppression_count` | context compilation uses **only** `working_memory_weight` (`:5974-5991`) |
| Consciousness layer | `layer_id` | default 1 = CORE_IDENTITY (`:5996`) |
| ACT-R learning | `access_ts[K]` ring buffer, `access_head`, `access_filled`, `wm_anchor` | base-level learning (`:5997-6008`) |
| Semantics | `emb` (768-dim nomic-embed-text vector, lazily backfilled), `emb_dim` | `:6009-6016` |
| Hebbian | eligibility trace | `:6017+` |
### Node types are strings, not an enum
`node_type` is a free `char*`, defaulting to `"Memory"` when unset
(`el_runtime.c:7401`, `server.el:159`). There is **no closed node-type enum** in
the shipped engine. Two consequences:
- The runtime *special-cases* a handful of type strings for activation
thresholds (`engram_type_threshold`, `:5933-5955`): `DharmaSelf`/`Safety`
(0.05, fire easily), `Belief`/`Entity` (0.30), `Knowledge` (0.20), everything
else `Note`/`Memory`/`Working` (0.40). `InternalStateEvent` and `Tag` are
**excluded from working-memory promotion** (`:6674-6676`, `:7368-7370`).
- Type strings the neuron layer actually writes: `Memory` (default), `Knowledge`
(`server.el:549`), `InternalStateEvent` (`server.el:493`), `Tombstone`
(`memory.el:55`), `Conversation` (session nodes, `sessions.el`), `Persona`
(`soul.el:250-292`), plus identity/value `Knowledge` nodes.
The types the MCP surface names — `Self`, `BacklogItem`, `SessionSummary`,
`Artifact`, `Process`, `ConfigEntry` — are **`node_type` string conventions set
by higher neuron/Axon layers**, not runtime-known types. Where `BacklogItem` /
`Artifact` are set was not in the files read (they route to the Axon backend, doc
02) — **flag as unverified/TODO** for a human pass.
## Edges
`EngramEdge``el_runtime.c:6701-6730+`. Directed, typed, weighted:
| Field | Meaning |
|---|---|
| `id`, `from_id`, `to_id`, `relation` | typed relation string |
| `weight` (double) | **authored** strength — never mutated by activation |
| `hebb` (double) | **learned** co-activation potentiation — the fraction of recent activations in which both endpoints were in working memory together; strictly separate from `weight` |
| `inhibitory` (int flag) | if set, activating the source **suppresses** the target's WM weight instead of exciting it |
| `confidence`, `created_at`, `updated_at`, `last_fired`, `layer` | — |
The **`hebb` field is the co-activation weight** — the Hebbian/LTP channel — kept
deliberately separate from the static authored `weight`. Edges are created via
`engram_connect(from, to, weight, relation)` (`server.el:253`).
**Relation strings observed:** `associates` (default, `server.el:248`),
`identity`, `co-value`, `birthday-twin`, `canonical-self` (`soul.el:37-108`),
`supersedes`, `tombstones`, `contains`, `tagged` (`neuron-api.el`,
`el_runtime.c:6168`).
> **Edges as vectors — the intended model (TARGET; today's edge is scalar).** The
> live edge above carries a typed `relation` string plus **scalar** strength
> channels (`weight`, `hebb`). The design target is for an edge to be a **vector**
> — a first-class carrier of relationship-*meaning* in the node space — so that
> relationships can be **composed / subtracted / analogized / traversed** like
> nodes (the `06` §6 operator algebra over edges). Combined with append-only, this
> yields a **complete temporal record**: every discrete, significant change to a
> relationship is appended (a keyframe on material change), so the **full 4-D
> trajectory** of the meaning-manifold is preserved and `recall_at(t)` can read
> how any relationship was configured at any past `t` — bounded, because changes
> are discrete and meaning saturates by compositionality. **Status: TARGET / #39**
> (see `07-storage-coherence-and-distribution.md` §2.4); the runtime edge is scalar
> today.
## Consciousness layers
Orthogonal to memory tiers, the engram has five canonical **layers**
(`el_runtime.c:5919-5924`):
| id | Name | activation_priority | Role |
|---|---|---|---|
| 0 | SAFETY | 0 (fires earliest) | deepest / limbic |
| 1 | CORE_IDENTITY | — | **default** for all nodes (`ENGRAM_LAYER_DEFAULT`, `:7423`) |
| 2 | DOMAIN | — | domain knowledge |
| 3 | IMPRINT | — | persona overlay |
| 4 | SUIT | — | outermost |
`EngramLayer` (`:6731-6738`) carries `activation_priority` (lower fires first),
`suppressible` (can higher layers suppress it?), `transparent` (invisible to
introspection?), and `injectable` (add/remove at runtime?). Layers are managed
via `engram_add_layer` / `engram_node_layered` / `engram_list_layers`. This is
the identity-vs-domain-knowledge stratification, independent of the tier system
below.
## Two tier systems — do not conflate them
This is the single most important clarification in the data model, and the source
of the vocabulary mismatch flagged throughout this set.
### A. Cognitive memory tiers — the `tier` field
`Working` / `Episodic` / `Semantic` / `Procedural` (and `Canonical` in use).
Runtime default `"Working"` (`el_runtime.c:7408`; `README.md:41-49`). Nodes
**migrate between these by salience decay/reinforcement**, driven by the runtime.
Salience decays as `importance × 1/(1 + days_since) × ln(count + 1)`
(`README.md:57-62`). `memory.el` exposes `tier_working`/`episodic`/`canonical`
helpers (`memory.el:1-3`); `soul.el` writes `Semantic`-tier persona nodes
(`:267`, `:282`). So the live tier set is **{Working, Episodic, Semantic,
Procedural, Canonical}** with continuous salience/importance/confidence floats.
### B. Epistemic tiers & disposition — tags, not runtime concepts
The MCP-facing vocabulary — tiers `note → lesson → canonical`, disposition
`experimental → provisional → stable → deprecated` — is **not enforced anywhere
in `el_runtime.c`.** It is stored as **tags**:
- Knowledge capture preserves the incoming epistemic tier as a `tier:<x>` tag
rather than mapping onto a cognitive tier — deliberately, to avoid a lossy
mapping (`server.el:519-522, 544`).
- `promote_knowledge` writes a canonical node tagged
`["Knowledge","tier:canonical","disposition:stable"]` (`neuron-api.el:533`).
There is **no state machine** validating `experimental → … → deprecated`.
Disposition and epistemic tier are convention-by-tag. *(Flag: not structurally
guarded. The exact MCP-enum → tag/float mapping is not fully traced in the files
read — unverified/TODO.)*
## Write-protection
> **Superseded (2026-08-16) — see `06-cognitive-architecture.md` §12.5.**
> Authority: `foundation/el/lang/spec/correspondence-and-censorship.md` §6 (branch
> `design/correspondence-and-censorship`).
>
> > **In an immutable substrate, any mechanism that refuses a write is either
> > redundant with immutability, or an epistemic constraint misfiled as a
> > protective one.**
>
> The requirement this gate was built for was never stated. It is
> **non-circularity of the reference frame** — a reference fitted to its own
> readings reports perfect correspondence forever while drift becomes
> undetectable from inside. That requirement is satisfied by *when* (the frame
> updates while activation is internally seeded, not while it is being used to
> act), not by *what*, so **the gate becomes unnecessary rather than removed, and
> nothing takes its place.** Corruption requires mutation, and the engram does not
> mutate: recoverability, governance, evidence quality, and rate are already
> properties of the substrate. Only **authorization** is residue, and it is
> bounded — an unauthorized writer can *propose*, never erase.
>
> **This section contradicts itself.** Thirty-five lines below, under
> *Immutability*, this same document states the conclusion in its own words:
> *"nothing it does is ever destructive — the safety is **after** the act, not a
> gate before it"* (`:185-187`). The 403 gate documented here **is** the
> before-the-act gate that sentence says is unnecessary. The design spec §6 names
> exactly this redundancy.
>
> The mechanism below is **still LIVE in code** and is described accurately; what
> is superseded is the claim that it is needed.
`is_protected_node(id)` (`neuron-api.el:20-37`) is a **hard-coded allowlist of 15
identity/value node IDs** — the self root, the values hub, intellectual-dna,
memory-philosophy, voice, and the 8 value nodes. Handlers that could mutate the
graph (tombstone / supersede / evolve / connect) check it and return HTTP 403
`api_err_protected` (`:39-41`) for a protected target (checked at `:384, 511,
692, 705, 746, 768`). Edges *into* a protected node are also blocked
(`handle_api_link_entities`).
**The one sanctioned override** is `POST /api/neuron/cultivate`
(`handle_api_cultivate`, `neuron-api.el:960` — **the `:781-816` cited here is
stale; verified 2026-08-16**) — it performs the same ops with the protection check
skipped, gated by convention to Will's explicit cultivation sessions. The self
layer is writable, but only through a deliberate door.
> **Superseded (2026-08-16).** A door built for a wall that need not stand. Per
> §12.5 of `06`, the gate above is redundant with immutability, so the override
> for it is redundant too. Neither is deleted here — this is a documentation
> branch; the change is sequenced in `correspondence-and-censorship.md` §11.
## Immutability — tombstone, never delete
Engram nodes are immutable (`memory.el:64-69`). The model is:
- **Tombstone** — `mem_tombstone(node_id)` (`memory.el:46-71`) **keeps the node
and all its edges**, creates a `Tombstone` marker node
(`content = target id`, `label = "tombstone:<id>"`) and wires a `tombstones`
edge (weight 1.0). It never calls `engram_forget`. This is *the* one canonical
delete — every user-facing forget path routes through it. Default bounded reads
hide tombstoned nodes (`memory_hide_tombstoned`, `neuron-api.el:239-249`);
`?include_deleted=1` recovers them.
- **Supersede** — updates/evolves (`neuron-api.el:394-428, 506-541, 715-734`)
create a **new** node with the new content, wire a `supersedes` edge new→old
(weight 0.9, or 0.95 for promote), and **keep the original**. The response
returns both ids so the caller re-points. This is the `supersedes_id`
pattern: new node linked, old preserved, full audit trail.
> **Supersession is residue, not garbage.** The superseded node is the *trail of
> how the current understanding was reached* — kept deliberately, because sometimes
> the truth was in the **old** idea even when the old idea was not itself the truth.
> This is what lets autonomous self-reification (`06` §4.1) run ungated: every
> rename/re-cluster supersedes into this residue chain, so nothing it does is ever
> destructive — the safety is *after* the act, not a gate before it.
> **The hole to know about.** The raw runtime `engram_forget` **does** hard-delete
> (frees node + edges, `el_runtime.c:7647`), and the engram HTTP route
> `DELETE /api/nodes/:id` calls it directly (`server.el:322-328`). Immutability
> is therefore an invariant of the **neuron-api / MCP layer routing**, not of the
> store. A client that hits engram HTTP directly can bypass it. *(flag)*
`engram_forget` is also used *internally* for genuine GC: boot-counter pruning
(`memory.el:184`), session-summary/telemetry pruning (`soul.el:369`,
`sessions.el`). Those are bounded housekeeping, not user deletes.
## Persistence, snapshots, backups
- **Storage:** a single JSON snapshot `snapshot.json` under `ENGRAM_DATA_DIR`,
written by `engram_save` / read by `engram_load` (`el_runtime.c:9660+`; format
`{"nodes":[...],"edges":[...]}`). In prod that dir is the RWO PVC mount `/data`
(doc 04).
- **Write policy:** `persist_canonical()` writes the **full** snapshot after every
durable write (`server.el:133-141`). The batch-edge route snapshots **once per
batch** to avoid ~150 GB/day of writes from Hebbian edge churn
(`server.el:258-305`) — this is why `hebb_consolidate` batches (doc 02).
- **Boot safety:** on load, engram writes `snapshot.boot-backup.json` (good load)
or `snapshot.failed-load.json` (a non-empty file that parsed to 0 nodes)
(`server.el:718-734`). Read routes export to scratch paths
(`.scan-export.json`, `.sync-export.json`) and **never** touch the canonical
(`server.el:207-223, 418-437`) — a guard added after a read-route corrupted the
snapshot.
- **Off-cluster backup:** a Kubernetes CronJob (`engram-backup`) tars `/data`
every 15 minutes to `gs://neuron-db-backup/gke/neuron-prod/` and keeps the last
96 (24h) (`infrastructure/platform/k8s/neuron-mcp/backup-cronjob.yaml`).
> **Ticker, but not consolidation (2026-08-16).** Flagged because
> `06` §12.4's sequencing item is *"no tickers, no cron"* and an auditor
> sweeping for tickers will land here. This one is **ops/backup, not
> cognition** — it does not consolidate and must not be folded into the
> dreamer. Its local counterpart is the `ai.neuron.engram-backup` launch agent
> (`StartInterval = 3600`, measured 2026-08-16); a separate
> `ai.neuron.snapshot-backup` runs at `StartInterval = 900`. Note the
> **discrepancy**: this doc says the backup interval is 15 min, which matches
> `snapshot-backup` (900 s) rather than the local `engram-backup` (3600 s).
> The cluster manifest was not read on this branch — treat the 15-min figure as
> unverified here.
- **Retention:** InternalStateEvent telemetry pruned at 48h
(`ENGRAM_ISE_RETENTION_MS`, `server.el:485-499`).
> **Data-dir mismatch to flag:** the `server.el` header comment says the default
> is `~/.neuron/engram` (`:16`) but the code defaults to `/tmp/engram`
> (`:135, 717`). Prod overrides both via `ENGRAM_DATA_DIR=/data`. *(unverified —
> which default is intended)*
## The engram HTTP surface (`:8742`)
Dispatcher `handle_request` (`server.el:592-707`). Auth: `ENGRAM_API_KEY`; GETs
always allowed, mutations require `"_auth":"<key>"` in the JSON body
(`server.el:578-588`).
| Endpoint | Purpose |
|---|---|
| `GET /health`, `GET /` | health + live node/edge counts |
| `POST /api/nodes`, `GET /api/nodes`, `GET /api/nodes/:id`, `DELETE /api/nodes/:id` | node CRUD (DELETE = hard `engram_forget`) |
| `GET /api/edges`, `POST /api/edges`, `POST /api/edges/batch`, `GET /api/neighbors/:id?depth` | edge ops + traversal |
| `POST\|GET /api/activate?q&depth`, `POST\|GET /api/search` | spreading activation vs lexical search |
| `POST /api/strengthen` | Hebbian potentiation |
| `POST /api/save`, `/api/load`, `/api/load-merge` | snapshot control |
| `GET /api/sync` | soul daemon periodic pull |
| `GET /api/embed-backfill`, `GET /api/similarity?a&b` | embeddings + cosine |
| `POST /api/neuron/state-events` (auth-exempt), `POST /api/neuron/knowledge/capture` | neuron-layer helpers |
| `GET /api/stats`, `/api/act-stats`, `/api/text-health` | telemetry |
## Retrieval model (summary)
Retrieval is **spreading activation, not query matching**:
`strength = parent_strength × edge_weight × target_salience ×
cosine(query, target)` — multiplicative, top-N, with the two-layer
background → working-memory promotion (`README.md:27-36`; `el_runtime.c:5892+,
6094+`). `mem_recall` / `/api/activate` fire this and mutate WM; `mem_search` /
`/api/search` are passive lexical scans — **but as of 2026-08-14 the live
`route_search` runs structure-gated *geometric* retrieval**
(`engram_retrieve_geometric_json`; held-out **P@5 = 0.700**, semantic not lexical —
`skill` returns skill nodes and *rejects* the false-positive `rainfall`), with the
old lexical scan retained at `/api/search-lexical` (see `06` §2.5). The cognitive
API's `begin_session` and `compile_ctx` return a **bounded projection** of the
activated set, never the raw
graph (doc 02, §2).
## Update — 2026-08-14: layers as named neighborhoods (DESIGN; backlog #49)
A refinement of the `## Consciousness layers` model above, from the deep-night
session (node `92941631`). A **layer is not a storage tier — it is a named,
persistent relational neighborhood** in the one engram, each carrying its own
**growth policy** and its own **lock / threshold policy**:
- **Threshold-lock = `note``canonical` maturation at neighborhood scale.** The
same epistemic-tier promotion the two-tier model (§B above) applies to a single
node is lifted to a *region*: a neighborhood **earns its lock** by maturing past
a threshold, at which point it stabilizes (read-mostly) the way a canonical node
does. Growth and lock are per-neighborhood, not global.
- **A user's imprint is just another neighborhood.** It is not a separate store or
a bolted-on partition — it lives in the same geometry as everything else.
- **Relate-across is the advantage over island engrams.** Because every
neighborhood shares one geometry, anything can form edges to anything across
neighborhood boundaries — the structural reason a single engram with named
neighborhoods beats a set of isolated per-purpose stores.
**Status: DESIGN.** This is the intended model for engram layers; the naming,
growth, and threshold-lock policies are not yet a built runtime feature. See
`06-cognitive-architecture.md` (Update — second pass).