Files
neuron/docs/architecture/04-runtime-and-deployment.md
T
will.anderson 4bff40fa4a
Neuron Soul CI / build (pull_request) Failing after 14m5s
Neuron Soul CI / deploy (pull_request) Has been skipped
fix(api): bound inspect_graph with relevance-ranked projection; regen soul.c
High-fanout identity anchors (voice, writing-imprint, self-root) have ~670KB
neighborhoods. inspect_graph returned the full traversal, which overflowed the
MCP client's context and socket-closed the wrapper mid self-load -- the soul
could not traverse its own identity graph.

handle_api_inspect_graph gains an opt-in `compact` projection (compact=1|true):
the neighborhood is relevance-ranked, the top K (default 12) keep a UTF-8-safe
content snippet (default snip=600), and the remainder collapse to lightweight
{id,label,node_type,tier,edge,pointer:true} stubs. This bounds the voice node
from 669,799B -> 25,353B (HTTP 200, valid JSON) and the wrapper's soul-load no
longer socket-closes. New helpers: api_compact_neighbors, api_neigh_full,
api_neigh_pointer, api_neigh_rank, api_neigh_better, api_float_or.

The flag is gated: ABSENT it, the response is byte-identical to the old plain
traversal, so the studio app (which never sends it) is unaffected. The MCP
wrapper (mcp-wrapper/src/main.el) appends &compact=1 on its inspectGraph and
fetch-by-id paths.

dist/soul.c is REGENERATED so CI ships the fix: CI compiles the committed
single-TU dist/soul.c directly (running elb/elc on the Linux runner OOM-kills
it), so an .el-only change would build the OLD behavior. Regenerated and verified
on macOS -- compiles with the CI cc line (0 errors) and, on a throwaway soul over
a copy of the live snapshot, serves compact ~25KB / non-compact ~670KB. The regen
also syncs the amalgamation to this branch's .el sources, which had drifted
several self-review commits ahead of the previously-committed soul.c.

Docs: docs/architecture/00-05 added; 01/02/05 corrected so the relevance-ranked
inspect_graph projection reads as committed source, not an in-flight concern.
2026-08-10 10:28:50 -05:00

10 KiB

Neuron — Runtime & Deployment

Process/port topology, the end-to-end MCP request path, local vs GKE blue/green production, and a high-level view of secrets/config. Grounded in entrypoint.sh, scripts/blue-green-deploy.sh, the k8s manifests under infrastructure/platform/k8s/neuron-mcp/, and .gitea/workflows/.

Process & port topology

A running neuron is two processes in one container: the soul and the engram, started by entrypoint.sh.

 container (one pod)
 ┌──────────────────────────────────────────────────────────┐
 │  entrypoint.sh                                             │
 │    1. start engram  (background)   ── listens :8742       │
 │    2. wait /health up to 60s                              │
 │    3. exec soul (PID 1 foreground)  ── listens :7770      │
 │                                                           │
 │   soul  :7770  ──HTTP──►  engram  :8742                    │
 │   (ENGRAM_URL=http://localhost:8742, HTTP mode)           │
 │                                                           │
 │   /data  (PVC mount) ◄── engram snapshot.json             │
 └──────────────────────────────────────────────────────────┘
  • entrypoint.sh starts engram with ENGRAM_BIND=:8742 and ENGRAM_DATA_DIR=/data, polls http://localhost:8742/health (up to 60s; Autopilot cold starts are slow), then execs the soul. SOUL_ENGRAM_PATH is deliberately unset so ENGRAM_URL triggers HTTP mode (soul talks to engram over localhost HTTP, not an in-process embed).
  • EL HTTP runtime is tuned down for co-located calls: EL_HTTP_TIMEOUT_MS=10000, EL_HTTP_CONNECT_TIMEOUT_MS=3000 (entrypoint.sh).

Full port map

Port Process Role Source
7779 mcp-proxy MCP client entry; byte-forward + retry mcp-proxy/src/main.el
17779 mcp-wrapper MCP JSON-RPC ⇄ soul REST; tool catalog mcp-wrapper/src/main.el
7770 soul HTTP cognitive API + handle_request NEURON_PORT, deployment-blue.yaml
8742 engram graph store HTTP entrypoint.sh, server.el:711
7771 neuron-connectd MCP connector bridges routes.el connectd_*

Local vs prod, an important distinction. The proxy → wrapper chain is the local developer adapter: a stdio MCP client (Claude Code) needs to reach an HTTP soul, so the proxy/wrapper translate and add resilience. In production, the neuron-mcp Kubernetes Service is a ClusterIP that targets the soul's :7770 directly (service.yaml) — external access is "to be wired via Cloudflare Tunnel later" (annotation, same file). So in prod the MCP/HTTP boundary is the soul's own HTTP surface; the proxy/wrapper are not (yet) in the cluster path. (inference from the ClusterIP-only Service + the local-only proxy/wrapper binaries.)

The MCP request path (end to end)

A single tools/call from an MCP client, local topology:

 client                proxy :7779         wrapper :17779        soul :7770            engram :8742
   │  JSON-RPC             │                    │                    │                     │
   │ tools/call ─────────► │  forward+retry     │                    │                     │
   │                       │ ─────────────────► │  map tool→REST     │                     │
   │                       │                    │ ─── HTTP POST ────► │  handle_request     │
   │                       │                    │   /api/neuron/...   │   → handle_api_*    │
   │                       │                    │                    │  engram_* builtin   │
   │                       │                    │                    │ ── (HTTP mode) ───► │  activate/search/
   │                       │                    │                    │                     │  save
   │                       │                    │                    │ ◄─── nodes/edges ── │
   │                       │                    │ ◄── JSON result ── │                     │
   │                       │                    │  fire_activation    │                     │
   │                       │                    │  /recall warm-up ─► soul (side effect)    │
   │ ◄──── result ──────── │ ◄───────────────── │                    │                     │

Responsibilities per hop, and the volatility each isolates (VBD reading):

  1. proxy — transport resilience. Isolates client connection volatility (drops, retries, health) from everything above. No MCP semantics.
  2. wrapper — protocol translation. Isolates the MCP protocol from the soul: owns initialize/tools/list/tools/call, the ~90-tool catalog, and dispatch_tool_call. Also fires the fire_activation /recall side effect so tool use warms working memory.
  3. soul — cognition. handle_request dispatch → handle_api_* → engram builtins. In HTTP mode it reaches engram over localhost; otherwise embedded.
  4. engram — the graph. Spreading activation, Hebbian edges, snapshot persistence.

For the user-facing chat pipeline (not tool calls), /api/chat enters layered_cycle (soul.el) — L1 safety → L2 stewardship → L3 imprint — described in 02-components.md §3c.

Production: GKE blue/green

Neuron prod runs on GKE cluster neuron-platform (Autopilot, us-central1), namespace neuron-prod. Two Deployments, neuron-mcp-blue and neuron-mcp-green, share one Service selector that names the active slot.

  • Deployments (deployment-blue.yaml / deployment-green.yaml): one container soul, image pinned by digest (not :latest) so Argo CD can't drift the active slot to an untested build (see the pin comment in deployment-blue.yaml). strategy: Recreate — the PVC is RWO so only one pod can hold it at a time. Probes hit /health on :7770.
  • Service (service.yaml): ClusterIP neuron-mcp, port 7770 → 7770, selector: {app: neuron-mcp, slot: blue}. The blue/green script patches slot.
  • Storage (pvc.yaml): neuron-engram-data, standard-rwo (pd-balanced), 10Gi, RWO. Engram data is the single snapshot.json (~8MB active).
  • The swap (scripts/blue-green-deploy.sh): (1) set image on the target slot; (2) scale target to 1, wait for rollout; (3) patch the Service selector to the new slot (traffic flip); (4) scale the old slot to 0. Imperative kubectl for the live swap, then git-update the Argo manifests so a sync doesn't revert replica counts.
  • Backup (backup-cronjob.yaml): every 15 min, tar /data → GCS, keep 96.

Resource sizing (learned the hard way)

deployment-blue.yaml documents the memory history in comments: idle soul RSS ~860Mi; the beginSession call (loads memories + backlog + preferences) spikes past 1Gi and OOM-killed the pod mid-request (client socket closed). Current setting: requests = limits = 2Gi, cpu 250m/1000m. This is why the cognitive API projects/compacts payloads so aggressively (doc 02 §2, doc 01 axis 1) — the memory ceiling is real and close.

CI/CD

Two Gitea Actions workflows (.gitea/workflows/), serialized on a single GCE runner (concurrency: neuron-runner).

  • ci.yaml (push/PR to main):
    • build: free disk → checkout → install gcc/libcurl/gcloud → download el-runtime-c/el-runtime-h from Artifact Registry foundation-prod (elc/elb intentionally not downloaded) → compile the committed dist/soul.c directly: cc -O2 -DHAVE_CURL dist/soul.c el_runtime.c -lssl -lcrypto -lcurl -lpthread -lm -o dist/neuronstrip -s → smoke test dist/neuron --help → publish neuron-soul@<sha8> to AR (push only).
    • deploy (push-to-main only): auth GCP → get-credentials neuron-platformdetermine idle slot (the deployment at 0 replicas) → prepare artifacts (soul binary + elc + runtime for the Docker build) → clone the engram repo into ./engram/ (Dockerfile builds engram from source) → docker build+push neuron-soul:<sha>scripts/blue-green-deploy.sh --image --slot → git-push updated infra manifests → kubectl rollout status → verify neuron-mcp endpoints.
  • deploy-gke.yaml (workflow_dispatch only, slot default green) — manual rollback / forced-slot deploy without a rebuild; same auth → slot → docker → blue-green → manifest-sync → verify steps.

The Docker image (Dockerfile) is a two-stage build: stage 1 compiles engram/src/server.elengram.cengram binary via elc + cc; stage 2 is an Ubuntu 24.04 runtime (GLIBC 2.39 satisfies both binaries) with soul + engram + entrypoint.sh.

Config & secrets (high level)

Runtime configuration is injected as environment, sourced from a Kubernetes Secret neuron-soul-secrets via ExternalSecret (ESO → GCP Secret Manager, Workload Identity — no key files). From deployment-blue.yaml:

Env Meaning
NEURON_PORT soul HTTP port (7770)
NEURON_LLM_0_URL / _KEY / _FORMAT primary LLM endpoint (Anthropic format)
SOUL_CGI_ID / SOUL_IDENTITY CGI id + identity seed (→ seed_persona_from_env, soul.el:250)
NEURON_TOKEN auth token (present in env; note the HTTP dispatch does not currently check it — doc 01 Divergence 3)
NEURON_API_URL self-callback URL (http://neuron-mcp.neuron-prod.svc.cluster.local:7770)
ENGRAM_URL / ENGRAM_DATA_DIR http://localhost:8742 / /data

There is also an in-graph config surface: ConfigEntry nodes read/written by inspect_config / tune_config (neuron-api.el:616-653) — runtime-tunable persona/behavior keys stored in the engram rather than the environment.

Operational note to flag. The deployment-blue.yaml image pin comment (dated Jul 2026) records that :latest resolved to an untested build lacking a mem_save/genesis-SIGSEGV fix, which is why the active slot is pinned to a digest. Any promotion must (a) rebuild a good soul and (b) update the digest in git so Argo CD and blue-green-deploy.sh agree. (state as-of the manifests read; verify current slot before deploying.)