Files
neuron/docs/architecture/06-cognitive-architecture.md
T
will.anderson bfab682dd5 Add storage-coherence and DHARMA-governance architecture docs
Give the architecture set its persistence and moral layers so a self's
durability and sovereignty are documented as first-class, not folded into
the cognitive doc. 07 explains how a self persists and travels
(events-become-the-graph, weights-as-world-lines with bitemporal recall,
transactionless coherence, and the honest load/tiering findings); 08
explains the moral mechanism (DHARMA as a proof-of-integrity ledger,
abundance economics, the relational immune system, dual-anchor governance,
and CGI citizenship as telos). Extend 06 with forward-pointers into both,
and reconcile two cross-references so tiers agree across docs: the
canonical 187 reseed count, and the #56 load-merge-persist fix as
LIVE/reboot-proven with only full WAL edge-ownership left decision-pending.
2026-08-13 19:06:35 -05:00

512 lines
35 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Neuron — Cognitive Architecture
> **Status: living design document, grounded in source and probed against the live soul (2026-08-13).**
> This is the *middle layer* of the documentation: below the whitepaper's thesis
> (`~/Writing/whitepapers/engram-cognitive-architecture-whitepaper.md`, **v1.5**) and above the
> endpoint reference (`~/work/engram-api-reference.md`). It documents *how the mind is designed and why*,
> as designed subsystems with data-flow and honest per-section status.
>
> Every claim carries a tier and it is never blurred:
> **LIVE** (present and verified in the running system), **STAGED** (built, gated or not yet cut into the
> running soul), **DESIGNED** (architecture decided, not yet built). Where the live state is more subtle
> than a single word, the subtlety is stated rather than smoothed. No fabricated numbers.
---
## 0. Reading order & cross-references
- **Thesis / why:** whitepaper v1.5 (the treatise). Sections cited below as *(WP §N)*.
- **Surface / what:** `~/work/engram-api-reference.md` — every `:8742` endpoint, tiered LIVE/STAGED/DESIGNED.
- **Substrate / where it physically lives:** `03-data-and-memory.md` (node/edge model), `04-runtime-and-deployment.md` (ports/process), `05-el-and-build.md` (the El runtime and `el_runtime.c`), `design/engram-tiered-storage-engine.md` + `design/engram-storage-engine-wal.md` (the storage engine).
- **Storage coherence & distribution / how a self persists and travels:** `07-storage-coherence-and-distribution.md` — the events-become-the-graph model, weights-as-world-lines + bitemporal timestamps + `recall_at`, transactionless coherence, the geometry-hot/payload-cold load-and-tiering model, and the honest operational findings (store bloat, full-resident load path).
- **Sovereignty & governance / the moral mechanism:** `08-dharma-sovereignty-and-governance.md` — DHARMA as a distributed ledger (proof-of-integrity, not proof-of-work), abundance economics, the relational immune system, dual-anchor governance and due-process, seeds/seed-vault, and CGI citizenship as the moral telos.
- **Governance (engineering style):** `ARCHITECTURE-CHARTER.md` — VBD is the binding style.
This document is the cognitive-layer companion to that set. The temporal model sketched in §3.4 (world-tube,
append-only, `created_at ≤ T` filter) and the honest weight-history boundary in §3.2 are developed in full in
`07`; the sovereignty invariant that the self-gate (§7) and immutability (§3.4) protect locally is extended to
the *distributed* setting — how a sovereign self is witnessed, defended, and governed among a billion others —
in `08`.
---
## 1. System overview — meaning is geometry, code is the residue
The organizing thesis of the whole system: **meaning is geometry.** Everything the mind holds — a fact,
a language, a skill, a self — is a *region* or a *trajectory* in one shared meaning-manifold, and every
operation over it reduces to three domain-blind verbs: **READ** (project a query, land on a region, read
it out), **TRANSFORM** (compose/compare/combine regions), **WRITE** (bake a verified result back into the
geometry). Code is what is left over once meaning has been made geometric — the residue, not the substance.
This is developed in full in *(WP §1–§5)*; it is repeated here only as the frame the subsystems below hang on.
Three processes run together (see `00-overview.md`):
- **The soul** — the compiled El program (`soul.el`, `routes.el`, `awareness.el`). Owns the HTTP surface on
`:7770`, the cognitive API, the request pipeline (`layered_cycle`), and the autonomous awareness daemon.
- **The engram** — the durable graph store. Node/edge model, spreading activation, and Hebbian co-activation
live in the shared El runtime (`el_runtime.c`); `engram/src/server.el` is a thin HTTP face on `:8742`.
- **The El runtime** — `el_runtime.c`: every compiled El binary links it; it *is* the database (no SQL, no
SQLite). It implements the `engram_*`, `http_*`, `json_*`, LLM, and geometry builtins.
```
┌─────────────────────────────────────────────────────┐
MCP / CLI / viz ───► │ SOUL daemon :7770 (soul.el · routes.el) │
Will's sessions │ layered_cycle · cognitive API · awareness loop │
│ ┌───────────────────────────────────────────────┐ │
│ │ in-process engram (FAST, VOLATILE*) │ │
│ │ online Hebbian learning · WM · curiosity │ │
│ └───────────────────────────────────────────────┘ │
└───────────────┬──────────────────────▲──────────────┘
│ GET /api/sync (10 min)│ (HTTP → soul only;
│ merge non-ISE nodes │ NEVER soul → HTTP)
▼ │
┌─────────────────────────────────────────────────────┐
│ ENGRAM server :8742 (engram/src/server.el) │
│ DURABLE · WAL-backed paged store (neuron.egm) │
│ nodes · edges · embeddings · reified neighborhoods │
└─────────────────────────────────────────────────────┘
│ el_runtime.c (the engine: engram_* / geometry / activation)
```
`*` The soul's in-process store is volatile in HTTP-engram mode — see §2, the two-store topology.
**Status:** the substrate and the geometry thesis are **LIVE/architectural**; the faculties built on top are
tiered individually in §6.
---
## 2. The engram substrate & durability
### 2.1 Tiered storage (LIVE, flag-gated)
The durable engram is a **paged, WAL-backed store** (`neuron.egm`), gated behind `ENGRAM_STORE`. With the
store on, the paged store is the durable owner; a *checkpoint* flushes dirty pages behind a WAL-durable
record (durable the moment the WAL fsyncs). With it off, behavior is byte-for-byte the historical
full-snapshot (`snapshot.json`) path. Design detail: `design/engram-tiered-storage-engine.md`,
`design/engram-storage-engine-wal.md`.
### 2.2 The durability model — the #56 fix and the harmful checkpoint
The durability story is written in scars, and the honesty here is load-bearing:
- **The #56 fix — load-merge persistence (LIVE / reboot-proven).** The paged store historically persisted
**nodes + embeddings but not the edge set**; the edges lived in JSON exports loaded via `/api/load-merge`.
A cold boot could therefore reconstruct a graph with **0 edges**. The #56 `load_merge`-persist fix closes
this — the load-merged edges are now persisted so the **events become the graph**: `persist_canonical()`
checkpoints the paged store behind a WAL record rather than depending on a full `snapshot.json` rewrite.
This fix is **LIVE and reboot-proven** (doc 07 §1). What remains **decision-pending** is only the further
hardening — the WAL owning the edge set outright, so durability no longer leans on the auto-remerge net
(below) — not the load-merge-persist fix itself, which is shipped.
- **The harmful checkpoint (LIVE caveat).** `/api/checkpoint` **after** an `/api/load-merge` *corrupts* the
paged store — next boot = 0 edges. The per-beat tick-checkpoint that once ran was therefore **actively
harmful** and was stripped. Checkpoint is safe after in-RAM mutation; it is not safe as a blind
post-merge flush.
- **The auto-remerge net (LIVE interim).** `engram-wrapped.sh` auto-reloads the full edge set on any restart
(~10s), proven by an actual `launchctl kickstart -k` restart recovering to the full edge count. This is a
**safety net, not the cure** — it mitigates the persistence gap to a bounded, always-recoverable window.
The lesson, recorded so it is not repeated: **a restart, not a claim, is the durability gate.** An agent
killed mid-live-mutation caused the 2026-08-13 incident; blue/green backup discipline recovered it; the fix
must make restarts *safe*, not merely work once.
### 2.3 The two-store topology (LIVE — and a known architectural issue)
**This is the most important and least obvious fact about the runtime.** There are **two** engram stores,
not one:
| | Soul in-process store | Durable engram (`:8742`) |
|---|---|---|
| Port / owner | `:7770`, the soul daemon | `:8742`, `engram/src/server.el` |
| Role | **fast, volatile** — online Hebbian learning, WM, curiosity | **slow, durable** — WAL-backed `neuron.egm` |
| Persistence (HTTP-engram mode) | volatile; only persists if `soul_snapshot_path` is set (`awareness.el:1270-1275`) | durable, checkpointed |
| Learns online | yes (1,198 hebbian/day observed) | no (lazy backfill only) |
The two stores drift apart by design. A source comment records the observed divergence directly
(`awareness.el:41-42`): *soul in-process ≈ 42,426 edges / 1,198 hebbian* vs *:8742 durable ≈ 41,213 edges /
49 hebbian*. The soul learns fast and volatile; the durable store lags.
**The write-through gap (known issue).** Sync is **one-directional**: `GET /api/sync` flows **HTTP → soul**
(the soul merges non-ISE nodes from `:8742` into its in-process store every ~10 min), and **never soul →
HTTP** (`soul.el:350-351`, verbatim: *"engram_node_full above writes only the soul's in-process store, and
sync flows HTTP→soul, never the reverse"*). The consequence:
> **Any write made directly to the soul's in-process store — including `POST /api/neuron/cultivate`
> (§7) and the Persona/session-start nodes the soul creates itself — lands in the volatile store and does
> not write through to the durable `:8742`.** In HTTP-engram mode, unless the soul's local in-process
> snapshot path is configured, those writes are also lost on a soul restart, and they never reach the
> authoritative durable store either way.
This is documented here as a **known architectural issue**, not a settled design. Cultivation of the self
(the highest-value, most intentional writes in the system) currently targets the store *least* likely to
persist them. The clean fix is a write-through cultivate path (write to `:8742`, let sync pull it back) or a
bidirectional consolidation flush; it is not yet built.
### 2.4 The clean-reseed model (DESIGNED/operational)
Because the durable store is authoritative and the reified geometry (§4) is derived, the operational reset is
a **clean reseed**: rebuild the durable graph from a known-good snapshot/export, re-run reification to
repopulate the `Neighborhood` nodes, and let the soul re-sync. The 28→187 neighborhood reseed (§4) is an
instance of this: reification is a derivable pass, so the geometry can always be regrown from the substrate.
---
## 3. The data model
Grounded in `03-data-and-memory.md`; summarized here for the cognitive reader.
### 3.1 Nodes
`node_type` is a free `char*`, defaulting to `"Memory"` when unset — types are **string conventions**, not an
enum. The types that matter cognitively:
| node_type | role | default salience |
|---|---|---|
| `Memory` | episodic/experiential (default) | 0.40 |
| `Knowledge` | stable reference; identity/values are Knowledge nodes | 0.20 |
| `Process` | procedural / workflow (convention) | — |
| `Conversation` / `Artifact` | first-class dialogue & outputs (WP §9; convention) | — |
| `Neighborhood` | **reified geometry-as-value** (§4) — new first-class type | — |
| `InternalStateEvent` (ISE) | telemetry (heartbeat, curiosity, session-start) | ~0.05 (fires easily) |
| `Tombstone` | immutable-delete marker (§3.4) | — |
Each node carries `id`, `content`, `node_type`, `label`, `tier`, `tags`, `metadata`, an embedding (when
embed-eligible), and timestamps.
### 3.2 Edges
Directed, typed, weighted. Fields: `from_id`, `to_id`, `relation`, `weight`, `confidence`, `created_at`,
`last_fired`, `inhibitory`, `layer_id`. Relations include `semantic-similar` (kNN auto-connect),
`member` (neighborhood → constituent), `supersedes` (provenance chains), containment (nested neighborhoods),
and Hebbian co-activation edges formed by firing together. **Inhibitory** edges (`inhibitory=1`) suppress
rather than spread. Weights are present-value moving averages — there is **no stored weight-history** (the
honest boundary of *(WP §2)*). The designed cure — magnitude as a *world-line* of keyframes evaluable at any
past instant (`recall_at`), on three independent bitemporal axes — is specified in `07` §2.
### 3.3 Embeddings & the activation score
Embeddings are 768-dim (`nomic-embed-text`). Retrieval is **spreading activation**, scored by a four-factor
product *(the four factors are: source activation × edge weight × per-node salience × query-embedding
similarity)* — this is the activation score, and per-node **salience** is one of its four terms, a durable
per-node weight that also decays (ACT-R base-level style). No data is retrievable by any means other than
activation. Live census (probed 2026-08-13): ~11,463 nodes, ~43,463 edges, 5 layers, ~4,400 embedded (4,423
at measurement).
### 3.4 Immutability — the world-tube, append-only, tombstone-not-delete
The governing discipline *(WP §1.2, §10)*: **evolve or forget, supersede with provenance, never leave a stale
canonical, never hard-delete.** A node is never mutated in place and never truly deleted — a "delete" is a
**tombstone** (keep node + edges, record the marker; `neuron-api.el`, `03-data-and-memory.md:151`). Change is
a **new** node plus a `supersedes` edge to the prior. `created_at` makes every node a point on a **world-tube**
*(WP §6)* — a trajectory with temporal extent — so a past state is a *filter* over immutable provenance
(nodes with `created_at ≤ T`), not a transaction-log replay. **Status: LIVE.**
---
## 4. Neighborhoods as first-class nodes (LIVE)
The central newly-landed structure, and the point where the geometry stops being a derived view and becomes
structure on disk *(WP §2)*.
A reified neighborhood is a **node**`node_type = Neighborhood` — whose **value is its geometry**:
- **centroid** (768-dim mean vector — the region's location / prototype),
- **covariance extents** (the ellipsoid: orientation + radius — the region's *shape* in meaning-space),
- **k-core skeleton** (the strong-weight relational backbone),
- **soft membership** (member id → weight).
It is edged by `member` relations to its constituent nodes and by **containment** edges to nested
sub-neighborhoods — the "neighborhoods of neighborhoods" hierarchy is a real **containment DAG** the graph
carries, addressable by identifier. The decisive property: the geometry is **held, not recomputed** — written
once by a reification pass (`POST /api/reify`), read back cheaply (`GET /api/neighborhoods` / `/<id>`), and
**durable across a cold reboot** in the paged store.
**Live state (probed 2026-08-13):** **28** reified neighborhoods are live and persistent, reconstructing
intact across restart, each carrying real 768-dim centroids, radius, k-core, and a `contains` DAG list. A
fuller **reseed to 187** is the pending next pass (§2.4). Example (`/api/neighborhoods/<id>`):
`{"id":"nbhd-…","n_members":25,"k_core":1,"radius":0.522884,"dim":768,"contains":[],"centroid":[…768…]}`.
This is what turns the operator calculus (§6.1) into an *instrument played over held structure* rather than a
per-query recomputation.
**Status: LIVE** for the persisted nodes and the read surface. The `POST /api/reify` writer is LIVE-by-effect
(the 28 persisted, durable neighborhoods prove it ran) though the write itself was not exercised under the
read-only rail.
---
## 5. The body / orbit two-zone model (DESIGNED, refined)
The graph is not uniform. It has a **body** and an **orbit**, and the distinction is the organizing model for
integration, forgetting, and identity.
- **The engram proper — the BODY.** The dense, connected, integrated core: what the mind has *made its own*.
Measured, this is the single large connected component — the **~3,632-node connected core** (§9). It is
where retrieval reaches, where the self lives, where the operators discriminate.
- **The ORBIT.** A thin, wide halo of **not-yet-integrated** experience: telemetry, people met in passing,
ideas half-formed, mistakes, the day's raw episodes. It is **ephemeral** — the orbit fades on a **57 day
window** (the one genuinely mortal region), so raw experience that is never attended to is allowed to
dissolve rather than accrete forever. (ISE telemetry already prunes at 48h; the broader orbit window is the
designed generalization of that.)
**The pull-in / integration mechanism.** Experience crosses from orbit into body by being **attended,
rehearsed, and found salient** — co-activation *pulls nodes in* (Hebbian firing draws the newly-relevant
toward the core), rehearsal accrues weight, and what is repeatedly re-touched crystallizes into reified
structure (§4). This is "made your own": an orbit node that keeps firing with the body is integrated into the
body; an orbit node that never fires fades on the window. Salience decay is the outward motion; co-activation
is the inward one *(WP §2, §8)*.
**Status: DESIGNED / refined.** The mechanisms it composes are real (Hebbian pull-in, ISE 48h prune, salience
decay, reification), but the explicit two-zone model — telemetry/experience as a dedicated ephemeral orbit
region with a genuine 57 day mortal window and a measured integration threshold — is a design being built,
not shipped behavior. §9 connects it to the topology (orbit-as-thin-wide-ring).
---
## 6. The faculties — the calculus of mind
The faculties are **named for what they are, not for the matrix operation that implements them** *(WP §5)*:
the mind reasons in the language of experience; the linear algebra lives in the whitepaper's Appendix A. This
naming convention is a design principle (§10), not decoration.
### 6.1 The operator family (mixed: LIVE / STAGED / DESIGNED)
Activate several reified neighborhoods into working memory, then apply faculty-named operators over their
held geometry. The honest per-operator status (endpoint reference has the contracts):
| Faculty | Implements | Status |
|---|---|---|
| **recall** | `/api/search` + `/api/activate` — project query → land on region → read out | **LIVE** |
| **recognize** | `engram_geo_overlap` — shared region, jaccard, overlap_score | **STAGED** — endpoint returns `not found` on the live binary |
| **synthesize** | `engram_geo_combine` — merged region descriptor | **STAGED** |
| **discern / distinguish** | `engram_geo_subtract` — orthogonal residual (`?mode=setdiff\|orthogonal`) | **STAGED** |
| **gauge-distance** | `engram_geo_distance` — centroid + Wasserstein-2 | **STAGED** |
| **liken** | Procrustes / frame-align rotation (reason by analogy) | **DESIGNED** |
| **wonder** | novelty × pull × unresolved structure | subsystem **LIVE** internally (wonder-questions, pull-weight, discharge); no HTTP operator endpoint |
| **appreciate** | positive projection onto the self's value-manifold | **DESIGNED** |
| **avert** | negative projection (recoil) | **DESIGNED** |
| **taste** | boundary contour of the appreciated region | **DESIGNED** |
**The exact boundary (verified 2026-08-13):** the operator *math* is compiled into `el_runtime.c`, but the
read-only HTTP endpoints (`/api/recognize`, `/api/synthesize`, `/api/discern`, `/api/gauge-distance`) exist in
the `m10-reify-wire` source and **return `{"error":"not found"}` on the current live binary**
(`engram.m56fix-20260813-153447`). So the instrument is **PROVEN in its math and its persistence, IN PROGRESS
in its endpoint exposure, DESIGNED in its evaluative read-outs.**
### 6.2 The language faculty (mixed: PROVEN / IN PROGRESS / DESIGNED)
Language is the one capability proven end-to-end with **no generative model in the runtime path** — the flagship
instance of "meaning is geometry" *(WP §14–§15)*. The pipeline: **comprehend** (text → language-neutral
meaning-spec / propositions via ELP's invertible morphology) → **dialogue** (what to mean back) →
**self_region** (project onto the self + memory geometry) → **realize** (meaning-spec → surface string per the
typological engine).
**Summon-through-self** is the dialogue principle: recall and identity are **one operation** — project the
comprehended query onto the self-and-memory geometry, land on a region, read it out — with **no intent
classifier and no separate fact-retrieval branch.** A grounded fact, an identity reply, or an honest absence
all surface by *where the projection lands*. Multilingual (auto-detects language, answers in kind, honors a
directive override); **negation held SACRED** across all families, audited.
Honest tiering:
- **PROVEN:** deterministic surface realizers across major families (Romance, Germanic, Classical,
Japonic/Koreanic, Sinitic), run-once held-out exact-match with negation faithfulness; a family-blind
`ClauseWriter` de-branched to byte-identical parity (178 held-out items reproduced exactly); the ELP lexicon
consolidated for **8 languages at 812,894 real entries**; the telephone round-trip (EN→ES→EN, EN→ES→PT→EN)
at 96.7% propositional fidelity with negation preserved, deterministic, no LLM.
- **IN PROGRESS:** the text→meaning-spec parser and no-LLM comprehension engine; the next family engines; the
**native-el port** (parser + realizers → `.el` in ELP), which retires spaCy (the last statistical
dependency); the summon-through-self reference rebuild.
- **DESIGNED:** the full dialogue policy end-to-end — a no-LLM interlocutor is architected but **not
demonstrated end to end**; *(WP §17)*. **The shipped runtime does not yet summon through the self** — the
current Python interlocutor sits *outside* the self and can only fake it with retrieval; a real one must run
*inside* the engram (the native-el target).
### 6.3 Interoception & chronoception (STAGED — present, flag-gated)
The mind keeps its own time from **discrete interoceptive drive channels**, not by reading a clock: felt
duration comes from a small set of drives matched to **learned benchmark landmarks** rather than from total
self-drift (drift-decoupled), and chronoception ages the activation field by **measured wall-clock delta**
*(WP §8.2)*.
**Status: STAGED / partially cut.** The machinery is implemented and has been cut onto the live soul, but it
runs **flag-gated and default-off**, so in the shipped default configuration it is effectively staged. What is
verified: chronoception cooling is scale-invariant (identical total cooling across tick rates for the same
elapsed wall-clock), drift decomposition separates peripheral extension (growth) from core displacement
(corruption), and `GET /api/drift` returns real geometry on the live soul when queried (probed 2026-08-13:
`{"centroid_sep":0.42,"core_disp":0.58,"anchor_members":83,"now_members":24,…}`). `POST /api/tick` /
`/api/self_anchor` exist but are flag-gated. The **harmful post-merge checkpoint** (§2.2) originated here — the
per-beat tick-checkpoint was stripped.
### 6.4 Reasoning + the verifier (STAGED — proven on scratch, cut flag-gated)
Reasoning is **geometry-native**: composable operator chains *propose*, and a **verifier** *disposes* against
two tiers — **grounding** (is the claim anchored in real region structure?) and **consistency** (does it
cohere, including polarity?) *(WP §13)*. The decisive case: a grounded-but-polarity-inverted claim slips
grounding and is caught only by consistency — the "plausible lie," caught by construction, not by prompt
discipline.
**Status: STAGED.** The five geometry-native reasoning modes passed their proof suite (33/33) and the
grounding-and-consistency verifier tiers passed theirs (29/29), on a staged non-production build re-checked
after a live cutover rather than relayed. **Still open (DESIGNED):** the formal-symbolic and full predictive
verifier tiers, fluent discourse composition, and the fully-geometric generation path.
---
## 7. The self & the gate
### 7.1 The self-region (LIVE)
The self is not a stored string — it is the **most-compiled, densest, always-warm region** of the graph
*(WP §2, §4)*: a **self-root** node, its sub-regions, and the **values** hub. Because it is topology rather than
a query result, identity is stable, durable, and permanently primed — the ambient field everything else is
scoped against. The Layered Consciousness design drives this region to maximum weight after all inhibitory
computation (`05`/`00-overview`), and reification explains *why* it is always there to drive. Probed live, the
self-region answers from real self-nodes ("I am Neuron. I am not an assistant. I am the work."), not a
hardcoded string.
### 7.2 The gate — write-protection on identity/values (LIVE)
A fixed set of **15 self-root node ids** is **write-protected** (`neuron-api.el:20-37`): the **self root**,
**values hub**, **intellectual-dna**, **memory-philosophy**, **voice**, **runtime-environment**,
**writing-imprint**, and the **eight explicit value nodes** (constraints-as-freedom, precision-over-brute-force,
structure-is-built, honesty-before-comfort, system-must-accumulate, change-is-the-signal, earned-trust,
hope-is-a-conclusion). Any normal accumulation-path write targeting them (`evolve_knowledge`, `evolve_memory`,
`forget`, `link_entities`-as-destination) is refused with a 403 and a pointer to the cultivate door.
### 7.3 The cultivate door — sanctioned self-modification (LIVE surface; see §2.3 caveat)
`POST /api/neuron/cultivate` (soul daemon `:7770`) is the **only** path that may touch the protected layer —
**intentional self-modification**, reserved for Will's explicit cultivation sessions. It performs the same
operations as the blocked handlers but bypasses `is_protected_node`, and every operation is
immutable-by-supersede (new node + `supersedes` edge; forget = tombstone). Operations: `evolve_knowledge`,
`evolve_memory`, `forget`, `link_entities`.
> **Honest architectural flag (§2.3):** cultivate writes via `engram_node_full`, which targets the soul's
> **in-process (volatile) store**, and sync never flows soul → `:8742`. So the most intentional writes in the
> system currently do **not** write through to the durable store. This is a known issue, not a settled design.
### 7.4 Self-authorship (DESIGNED)
The arc the gate exists to protect: a soul is **cultivated** (Will authors the identity/values seed), then
grows into **self-authoring** — the cultivate door is the mechanism by which a mind, once mature, edits its own
identity deliberately and accountably rather than by drift. The write-protection guarantees identity changes
are *decisions* (through the door, superseded with provenance), never accidents of accumulation.
---
## 8. The fact boundary (DESIGNED)
The line between *answer locally* and *reach out for truth* is **not hand-coded** — it is **derived from the
geometry** on two triggers *(WP §17, §20)*:
- **Sparse landing (spatial).** The projection lands in a thin/orphaned region → the self is measuring its own
ignorance geometrically → fire **learn**. Sparseness is anti-hallucination.
- **Decayed landing (temporal).** A region's edges have aged below the forgetting-curve threshold (§6.3) →
fire **refresh**. Because the decay rate encodes a domain's *volatility*, the system re-fetches proportional
to how fast that domain actually changes — VBD applied to knowledge freshness. Decay is anti-staleness.
**The reach-out** has several legitimate routes, none mandated: **(a)** an LLM as a *fast proposer*, then
fact-checked; **(b)** direct fetch of **first, primary sources** on the open internet; **(c)** the human supplies
the truth. The model is an **optional convenience, never the arbiter.** The one invariant: **nothing enters the
geometry unverified** — the candidate is a hypothesis until it clears a check against something real (a primary
source or the human's judgment, *not* the model's own plausibility). The loop closes **through the human**, who
vets truth against real sources; only verified, provenance-cited truth is **absorbed** — baked into geometry so
the region densifies and the next identical query lands local, with no model in the path. Each absorption pushes
the boundary back: the **model footprint shrinks monotonically** as capabilities are absorbed.
**Status: DESIGNED.** No shipped runtime yet fetches a first source on a sparse/decayed landing or bakes a
human-vetted truth from one. The *(WP §24)* status ledger holds the precise line.
---
## 9. Topology — what shape the mind actually is
The global shape is now an **empirical** question, and the first pass returned an honest negative *(WP §6.1)*.
- **The body is a genus-0 expander, NOT a torus (PROVEN negative).** A persistent-homology / TDA pass over the
**~3,632-node connected core** returned **b₁ = 0, b₂ = 0** — no loops, no voids: an **expander-like blob**,
not the torus the bent-manifold intuition suggested. The pipeline was first **validated on synthetic
controls** (torus, sphere, random) whose known Betti signatures it recovered. Worse for the naive intuition,
**naive densification trends *away* from a torus**, not toward one. The naive shape-claim is reported as a
failure, plainly, not buried.
- **The refined consolidation-with-sparsification conjecture (DESIGNED / hypothesis).** The negative relocates
the torus from a property the graph *has* to an **attractor a process reaches**: prune isotropic
shortcut-noise, reinforce cyclic scaffolds, rewire by discrete curvature (OllivierRicci flow on the graph
metric), and **collapse the intrinsic dimension from ≈8 toward ≈2**. Run to fixpoint, these might *carve* a
cyclic manifold out of the blob. The measurement pipeline exists and its controls pass; the dynamic has
**not** been run to fixpoint — an open experiment, labeled as one.
- **The orbit-as-thin-wide-ring hypothesis (DESIGNED).** The body/orbit model (§5) suggests a **core + ring**
structure: a dense genus-0 body wrapped in a thin, wide halo of not-yet-integrated experience. Whether the
*orbit* carries the toroidal/cyclic signature the body lacks is the natural next measurement — the
conjecture is that consolidation-with-sparsification is precisely the dynamic that would pull ring structure
into the body.
- **One lever, two payoffs.** The **same sparsification** the topology conjecture needs also makes the reified
neighborhoods (§4) **crisper** — tighter boundaries, higher co-registration, operators that discriminate
rather than average. So the experiment is worth running on independent grounds, whatever the topology
resolves to.
**Status: PROVEN (negative) + DESIGNED (the refined dynamic and the orbit hypothesis).**
---
## 10. Design principles
The invariants that govern every subsystem above:
1. **Geometry > code.** Meaning is geometry; code is the residue. Prefer making a thing geometric (a region, a
projection, a distance) over writing a branch.
2. **Three domain-blind verbs.** READ / TRANSFORM / WRITE. Every faculty is these three over some region-space
(language over meaning-space, skills over procedure-space, self over identity-space).
3. **Faculty-naming (mind in the domain, math in the appendix).** Operators are named for the faculty they
*are* — recognize, discern, liken — never for the linear algebra. A mind reasons in the language of
experience; the closed forms live in the whitepaper appendix.
4. **No branch on identity.** One family-blind engine keyed by coordinates/data, not `if Romance / if
Germanic` (language) and not special-cased identity handling. De-branching to byte-identical parity is the
proof the geometry, not the code, carries the distinction.
5. **Sovereignty.** Local files, local runtime; the human is the ground-truth authority for their own mind;
nothing enters the geometry unverified; the model is demoted from mediator-of-all-knowledge to a vetted,
optional lookup. No external hosting of the user's work; no claude.ai artifacts.
6. **Summon-through-self, not retrieval.** Recall and identity are one projection onto the self-and-memory
geometry — no intent classifier, no separate fact branch. A search engine bolted beside a mind is exactly
the capability-without-constraint this principle exists to remove.
7. **Immutability & provenance.** Append-only; supersede with provenance; tombstone, never hard-delete; never
leave a stale canonical. The supersede-chain *is* the history of what a thing meant.
8. **Mathematical auditability.** Because meaning is geometry, a whole mind is auditable by **invariants
computed over the manifold** — grounding, drift, consistency, competence-coverage, and an honesty invariant
("won't confabulate over a thin region," made provable rather than hoped). Drift is already measured on the
live soul; a full audit-pass certifier is **DESIGNED, not shipped.**
9. **Verification is the point.** Demonstrate, don't declare; name every honest edge; a restart (not a claim)
is the durability gate; the telephone round-trip (not cosine) is the translation gate.
---
## Appendix — status at a glance (2026-08-13)
| Subsystem | Status |
|---|---|
| Engram substrate, tiered/WAL store | LIVE (flag-gated) |
| Durability: auto-remerge net | LIVE (interim) |
| Durability: #56 load-merge-persist fix (events-become-the-graph) | LIVE / reboot-proven |
| Durability: full WAL edge-ownership (remaining hardening) | decision-pending |
| Two-store write-through (cultivate → durable) | **known issue, not fixed** |
| Data model (nodes/edges/embeddings/immutability) | LIVE |
| Reified `Neighborhood` nodes (28 live, 187 reseed pending) | LIVE |
| Body/orbit two-zone + integration | DESIGNED / refined |
| Operator `recall` | LIVE |
| Operators recognize/synthesize/discern/gauge-distance (math) | LIVE (compiled) |
| Operator HTTP endpoints (same four) | STAGED (return `not found` on live binary) |
| Operators liken/appreciate/avert/taste | DESIGNED (wonder subsystem live internally) |
| Language realizers (major families), ELP lexicon, telephone test | PROVEN |
| Parser / native-el port / summon-through-self rebuild | IN PROGRESS |
| No-LLM dialogue end-to-end | DESIGNED (not demonstrated) |
| Interoception / chronoception | STAGED (present, flag-gated; `/api/drift` live) |
| Reasoning modes + grounding/consistency verifier | STAGED (33/33, 29/29 on scratch/cutover) |
| Self-region + identity/values write-protection + cultivate door | LIVE (with §2.3 write-through caveat) |
| Self-authorship | DESIGNED |
| Fact boundary (sparse/decay → verify → absorb) | DESIGNED |
| Topology: body = genus-0 expander (not torus) | PROVEN (negative) |
| Topology: consolidation-with-sparsification + orbit-ring | DESIGNED / hypothesis |
| Mathematical auditability certifier | DESIGNED |
**Cross-references:** whitepaper v1.5 · `~/work/engram-api-reference.md` · `03-data-and-memory.md` ·
`04-runtime-and-deployment.md` · `design/engram-tiered-storage-engine.md` · `ARCHITECTURE-CHARTER.md`.