Nothing else on the memory roadmap should be built until a change can be shown
to help. Right now we judge by feel, and the benchmark literature is full of
systems that felt better and measured worse. This is the missing gate.
WHAT IT MEASURES, AND WHY IT BOOTS A REAL SOUL
The subject is Will's designed retrieval — spreading activation over the
weighted directed graph, four-factor multiplicative scoring — not a proxy for
it. A Python re-implementation would measure my reading of the design, so the
harness compiles the actual soul.el amalgam from a git ref and asks it over
HTTP on /api/neuron/recall, exactly as the MCP wrapper and the app do.
BUILT ON WHAT WAS ALREADY HERE, NOT AROUND IT
docs/research/graphrag_eval/{collect,score}.py — per-query relevant-id
scoring and fixed-denominator precision@5 (kept verbatim: an empty result
should be punished like a page of junk).
docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py — the pinned
ground truth + --check winnability gate, so every run judges alike.
scripts/verify-soul-contract.sh — the isolation recipe, including the
non-obvious SOUL_ISE_URL pin without which an "isolated" soul silently
syncs the operator's live brain.
gen-soul-amalgam.sh + .gitea/workflows/ci.yaml — the build recipe and flags.
New here: ids rather than regexes as ground truth, an associative category
derived from real edges, a superseded category scored on ranking, a
machine-checked zero-lexical-overlap guarantee on paraphrases, paired
significance testing, and measurement of the real compiled soul rather than an
offline replica of one leg of it.
THE GOLD SET IS AUDITABLE, NOT VIBES
38 queries over the real 78,768-node corpus, each carrying a `derivation`
string, each re-validated by `build_gold_set.py --check`. exact_rare is mined
(document frequency 1). phrase is mined (verbatim scan; >25 matches rejected as
too diffuse). paraphrase is hand-selected then PROVEN to share zero content
words with its target — a leak fails the build, so the category cannot decay
into lexical matching. associative is derived from real hub edges with
lexically-reachable siblings dropped. nonsense is verified absent. superseded
pairs are kept only when both sides survive as distinct nodes.
HONEST ABOUT NOISE
Minimum detectable swing on 38 queries is 6: if every changed query moves the
same way, p = 2*0.5^n first clears 0.05 at n=6. Run-to-run drift is measured,
not assumed — activation is a stateful read, and it shows: main is fully
deterministic across 3 runs, the candidate drifts by 1 query. compare.py
reports "no measurable difference" for anything inside max(6, drift+1).
FIRST VERDICT — feat/recall-through-activation
hit@5 34.3% -> 22.9%, phrase 85.7% -> 28.6%, latency p50 2.81x. Five discordant
pairs, all five against the candidate, none for it; McNemar exact p = 0.0625,
so by the stated rule this is one query short of significant and is reported as
such rather than as a win for main. The latency regression is deterministic and
not in any noise band.
The benefit the branch was written for is absent: associative recall is 0/6 on
BOTH builds. Probed directly, the traversal returns the lexical seed at rank 8
and none of its 12 hub siblings. Two measured corpus facts explain it — only
4,060 of 78,768 nodes (5.2%) carry any edge, and no node has an embedding, so
the fourth factor of the four-factor product has nothing to compute from. The
mechanism runs; the corpus lacks the structure it needs.
SAFETY
Throwaway port, throwaway HOME, disposable per-run copy of the corpus; live
ports refused by name. Every soul started is killed AND confirmed dead by pid
probe, with the confirmation written into the results file; run_comparison.sh
sweeps for strays and exits non-zero if any survive. Nothing under ~/.neuron,
/Applications/Neuron*, or ~/neuron-dev-stack is read, written, or restarted.
Rung: E2E-VERIFIED — 6 full runs (3 per config) against the real compiled
binaries on the real corpus; numbers above are measured, not projected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The soul obeys half of its own ownership rule. soul.el:571-573 says "when
ENGRAM_URL is set the HTTP Engram owns persistence — the soul must NEVER write
to the local snapshot", and it doesn't. But nothing was ever built to hand the
soul's writes TO that owner: sync is pull-only (/api/sync -> engram_load_merge),
so every node created inside the soul lived in process RAM and was shed on
restart. Measured live 2026-08-07: soul node_count=102184, engram 79197.
SCOPE CORRECTION vs the earlier internal spec: engram provisional claim 17's
"pull-then-push" is a PEER-ENGRAM to PEER-ENGRAM protocol (claims 15-18 say so
explicitly). The soul is a CALLER of the database API, not a peer. Claim 17 is
NOT authority for a soul<->engram contract and is no longer cited as such. The
design here follows from the ownership rule alone.
Mechanism: a new Accessor, persist.el, is the single boundary. Writes stage a
delta to a filesystem spool and are pushed to the owner via POST /api/load-merge
— NOT POST /api/nodes, which mints a new server-side id (breaking dedup and
edges) and drops label/tier/tags/importance/confidence (verified in a sandbox:
a tier "Canonical" probe came back "Working"). load-merge preserves the id and
every field, dedups nodes by id and edges by (from,to,relation) so retries are
no-ops, and calls persist_canonical() so THE OWNER writes its own file — the
ownership rule is honoured rather than worked around.
Spool-and-drain rather than push-per-write: measured ~0.38s per load-merge at
live scale (79k nodes/176MB), and a chat turn writes 5-7 nodes. The spool is on
disk, not in process state, because the soul serves each connection on its own
pthread and a shared buffer would lose entries to a read-modify-write race. That
also buys crash recovery: writes orphaned by kill -9 are drained on next boot.
Honesty: api_persisted (the gate all 10 MCP write handlers pass through) and
mem_store now assert AT THE OWNER instead of reading back the soul's own RAM.
With the owner down a write returns {"ok":false,"error":"write_not_persisted"}
and the delta is queued — where main returns {"ok":true} for a write that dies.
Coverage: 35 node sites + 9 edge sites routed through the boundary. Deliberately
excluded, with reasons in persist.el: 4 InternalStateEvent sites (Will's own
telemetry carve-out), the boot counter and the persona (both already have
bespoke owner-side write-backs), and soul.el's 54 genesis identity edges
(file-mode only). engram_strengthen and engram_forget are NOT propagated —
load-merge cannot update or delete, and hard-deleting at the owner would fail
verify-soul-contract.sh section B.
Also fixed here:
- routes.el GET /api/graph/edges engram_save()'d straight over the owner's
canonical snapshot.json — a read route, in a non-owner process, clobbering the
canonical on every call. Same defect class Will removed from the engram in el
dc39a61. Now exports to a scratch path. With this gone the soul writes nothing
at all in HTTP mode.
- persist.el must clear the runtime's _tl_fs_read_len hint after every fs_read.
In vendored runtime v1.0.0-20260501 that hint becomes the NEXT response's
Content-Length, so reading a spool file mid-request made an 86-byte reply go
out as 497 bytes with 411 bytes of adjacent heap trailing it. Caught and fixed
at our boundary; the runtime class was fixed upstream in el 43636ae, which is
not the pinned runtime here.
Rung: E2E-VERIFIED, discriminating. Same harness, same engram binary:
write-through: LEG 1 PRESENT at owner, LEG 2 SURVIVED kill -9 + restart
main: LEG 1 ABSENT at owner, LEG 2 LOST
verify-soul-contract.sh: GATE PASS on both builds (27/27 routes, immutability).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Retires the defect class behind #129. The engine's state store returns "" for a
key nothing writes — no error, no warning, no log. That is how the agentic
path's crisis-escalation input scored 0 on every real conversation for two days
after ff421d3 moved conversation history behind conv_hist_key(session_id) and
left one consumer reading the old "conv_history" bucket by hand.
scripts/verify-state-keys.sh is the gate; scripts/state-key-audit.py is the El
reader behind it. Two checks:
DEAD-READ a state_get whose key resolves to something no state_set in the
tree produces.
HAND-ROLLED a literal that belongs to a namespace a helper owns, accessed
without the helper. This is #129's actual shape, and DEAD-READ
alone does NOT catch it: the dead handle_chat() still writes
"conv_history" through conv_hist_key(""). Stating that plainly
because a gate that only appears to work is worse than none.
WHY IT DOES NOT CRY WOLF. Keys are usually computed, so a literal-matching
script would flood and be switched off in a day. The resolver handles
concatenation (matched on the static prefix), helper functions (resolved to
their possible returns, with guard conditions folded so conv_hist_key("") does
not falsely claim to produce the session_hist_ namespace), keys built into a
local, and keys arriving as a parameter (resolved through the call sites).
278 of 278 sites on this tree resolve: UNRESOLVED 0, FINDINGS 0. Unresolvable
keys would be listed and would NOT fail the build.
TWO-LEG PROOF, one variable — agentic_safety_screen's single line:
pre-fix scripts/verify-state-keys.sh --root <scratch>
chat.el:2536 state_get("conv_history")
conv_hist_key() owns this key namespace (EXACT 'conv_history')
FAIL: 1 state-key finding(s) exit 1
as-is scripts/verify-state-keys.sh
FINDINGS (0) ... PASS exit 0
INDEPENDENT CONFIRMATION: run read-only against
origin/feat/soul-openai-tools-v2, which carries the same defect on its own, the
gate reported chat.el:2937 — the exact line 43d0449's message had named by
hand, with no prior knowledge. Against origin/fix/129-on-openai-tools: PASS.
PRODUCER-MOVED CONTROLS: renaming the sole writer of an EXACT key (soul_model)
orphans 3 readers across 3 files; renaming the sole writer of a PREFIX
namespace (agent_workspace_root_*) orphans 3 readers, including when the
producer moves to a NARROWER namespace — a case an earlier, more permissive
prefix rule let through. That rule is now directional, with the reason written
next to it.
FOUND ON ITS FIRST RUN, unprompted: soul.el's state_set("soul_identity", ...)
was deleted 2026-05-13 in b163fa6 (a commit about awareness/ISE writes) and five
readers in chat.el were left behind — build_system_prompt, the vision handler,
the agentic system prompt and two council handlers have prefixed "" for ~3
months. studio.el:57 emits "principal":"" and never had a producer. Both are
recorded in state-key-baseline.txt with dates and causes so the gate can be
turned on today; they are DEBT, not false positives, and every run prints them.
Baseline signatures carry no line number (an unrelated edit must not un-mute an
accepted finding) but do carry a count, so a GROWTH in a baselined finding still
fails the build.
Engine behaviour unchanged: this commit adds scripts only, no .el is touched.
CI is deliberately NOT wired here — .gitea/workflows/ci.yaml has changes in
flight from someone else, and turning the gate on would immediately red
feat/soul-openai-tools-v2 (correctly). That flip should be deliberate.
Rung reached: RUNS — the gate executes (0.15s), discriminates on four
independent test pairs, and its verdicts are quoted above. Not wired to CI, and
no engine binary was built from this branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
P0 SAFETY. Same defect as #129, carried INDEPENDENTLY on this branch — not a
merge, not a duplicate report. feat/soul-openai-tools-v2 branched with the
ff421d3 (2026-08-05) regression already in it, so fixing it on
fix/129-history-amplification did nothing for this line of work.
THE DEFECT, verified here at chat.el:2937 (the reported line number was exact):
let history: String = state_get("conv_history")
let screen_result: String = safety_screen(message, history)
ff421d3 moved conversation history to a per-session key via
conv_hist_key(session_id). This consumer did not move with it. The desktop app
always mints a session id, so history is always written under
session_hist_<id> and this read always returned "".
The half of the crisis score that receives history is the escalation half — the
one that exists for distress building across several turns, where no single
message trips the bell on its own. It scored 0 on every real conversation on
this branch too. Single-message hard bell was never affected.
WHAT CHANGED — deliberately byte-identical to the sibling fix (43d0449) so the
two branches CONVERGE and whichever merges second is a clean merge, not a
conflict:
- agentic_safety_screen(session_id, message) owns the two decisions that were
inline — which window the screen sees, and the screen call. Inline safety
inputs are untestable safety inputs; that is what let a rename starve this
one with nothing failing and nothing logging.
- the handler calls it with sess_for_root, the session id already in scope
twenty lines above (this branch's handler is structurally unchanged from the
sibling's here, so this is a clean mirror — no reshaping was needed).
- the comment at the call site states the invariant (read window == written
window) instead of naming a key that can be renamed out from under it. The
old comment documented this same bug being fixed once already under #9; the
rename re-broke it and the comment went on describing a repair that no
longer held. A comment is not a gate.
Also brings over the sibling's scripts/run-el-test.sh and the regression test
(cherry-pick of b842e82, applied cleanly). NOTE: this branch already carries a
DIFFERENT runner at tests/run-el-test.sh from de65991 (elb-based, links whole
modules). Different path, no collision, both kept — the sibling's is the one
this proof used.
TWO-LEG PROOF, one variable — the single line
state_get("conv_history") -> state_get(conv_hist_key(session_id)), with the
extraction already in place on both legs so nothing else moved:
before scripts/run-el-test.sh tests/test_history_amplification.el
3. REGRESSION #129 — agentic screen reads the session's own window
FAIL: distress history escalates the agentic screen to hard_bell
got: soft_bell
expected: hard_bell
history amplification tests: 8 passed, 1 failed
[run-el-test] FAIL: reported failing assertions
after same command, same tree, that one line changed
3. REGRESSION #129 — agentic screen reads the session's own window
PASS: distress history escalates the agentic screen to hard_bell
history amplification tests: 9 passed, 0 failed
[run-el-test] PASS: test_history_amplification
Full engine rebuild from THIS branch's sources is clean:
gen-soul-amalgam.sh -> 1,185,285 bytes / 1231 inlined bodies (gate wants >=
1200), cc-brain.sh -> 903,144 bytes, 14 warnings, 0 errors. Both symbols present
in the built binary (nm: T _agentic_safety_screen, T _conv_hist_key).
Rung reached: BUILT + RUNS (discriminating test). NOT E2E-VERIFIED — not in a
DMG, not exercised against a live OpenAI-wire agentic turn in the app a human
opens. Neither is claimed.
Refs #129
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tests/ has held 14 test programs for months with no way to run them. CI does
not run them. The convention printed in their own headers
(`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
runtime does not implement. So the tests were documentation, not gates — which
is how a P0 safety regression shipped with a test directory sitting right
there.
scripts/run-el-test.sh compiles and runs one test program. It reuses the
gen-soul-amalgam.sh discovery: elc emits only an extern prototype for a module
that has a .elh beside it, and inlines the bodies when it does not, so a test
importing ../chat.el must be compiled in a scratch tree with the headers
removed. Scratch copy on purpose — the worktree is shared. It runs the binary
under a throwaway HOME so a test can never reach the live engram.
Exit status is the gate: the El tests print failures and still exit 0, so the
runner greps for FAIL lines and for a zero assertion count as well.
tests/test_history_amplification.el pins the invariant #129 violated: the
window the safety screen READS must be the window conv_history_record WRITES.
Not "must be called conv_history" — must AGREE.
THIS COMMIT IS RED BY DESIGN. On this tree the test fails one assertion:
3. REGRESSION #129 — agentic screen reads the session's own window
FAIL: distress history escalates the agentic screen to hard_bell
got: soft_bell
expected: hard_bell
history amplification tests: 8 passed, 1 failed (runner exit 1)
The next commit turns it green by changing one line. Two legs, one variable —
that is the whole point of committing the test first.
Two flaws in the older harness that this one does not copy: the idiom
`let pass_count = pass_count + 1` inside an assert function declares a local
that dies with the call, so every existing suite prints "0 passed, 0 failed"
regardless of outcome; and a test program without a `cgi` block compiles as a
'utility', which may not reference the self-formation primitives chat.el's
agentic loop calls — it fails to build on a capability violation it never
triggers at runtime.
Refs #129
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b842e82f77)
P0 SAFETY. Closes the regression we introduced in ff421d3 (2026-08-05).
ff421d3 correctly moved conversation history to a per-session key via
conv_hist_key(session_id). One consumer did not move with it: the agentic
path's L1 safety screen kept reading the anonymous "conv_history" bucket. The
desktop app always mints a session id (DaemonClient.kt:706), so history was
always written under session_hist_<id> and that read always returned "".
The half of the crisis score that receives history is the escalation half — the
one that exists for distress building across several turns, where no single
message trips the bell on its own. It scored 0 on every real conversation for
two days. Single-message hard bell was never affected.
The bitter part: the comment that line carried documented this exact bug being
fixed once already, under issue #9. The fix was right then. The rename
re-broke it, and the comment went on describing a repair that no longer held.
A comment is not a gate.
The read now goes through conv_hist_key like every other consumer, including
the plain path at soul.el:398 and the thread-anchoring read thirty lines below
it in this same handler. It is one line. The rest of this commit is structure
so it cannot happen quietly again:
- agentic_safety_screen() owns the two decisions that were inline — which
window the screen sees, and the screen call. Inline safety inputs are
untestable safety inputs; that is what let a rename starve this one with
nothing failing and nothing logging.
- the comment above the call site now states the invariant (read window ==
written window) instead of naming a key that can be renamed out from under
it.
TWO-LEG PROOF, one variable — the single line state_get("conv_history") ->
state_get(conv_hist_key(session_id)):
before scripts/run-el-test.sh tests/test_history_amplification.el
3. REGRESSION #129 ... FAIL got: soft_bell expected: hard_bell
8 passed, 1 failed runner exit 1
after same command, same tree, that one line changed
9 passed, 0 failed runner exit 0
Full engine rebuild from these sources is clean: gen-soul-amalgam.sh ->
1,164,103 bytes / 1226 inlined bodies (gate wants >= 1200), cc-brain.sh ->
903,096 bytes, 0 errors. agentic_safety_screen and conv_hist_key both present
in the built binary (nm: T _agentic_safety_screen, T _conv_hist_key).
Rung reached: BUILT + RUNS (discriminating test). NOT yet in a DMG and not yet
verified in the app a human opens — those are the next two rungs and neither is
claimed here.
Known and NOT fixed by this commit:
- feat/soul-openai-tools-v2 carries the same defect independently at
chat.el:2937 and needs the same change or a merge.
- the defect CLASS (a read of a state key no producer writes) is still
invisible to every gate we have. Issue #129 proposes making it a build
error; that is the follow-on.
Closes#129
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tests/ has held 14 test programs for months with no way to run them. CI does
not run them. The convention printed in their own headers
(`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
runtime does not implement. So the tests were documentation, not gates — which
is how a P0 safety regression shipped with a test directory sitting right
there.
scripts/run-el-test.sh compiles and runs one test program. It reuses the
gen-soul-amalgam.sh discovery: elc emits only an extern prototype for a module
that has a .elh beside it, and inlines the bodies when it does not, so a test
importing ../chat.el must be compiled in a scratch tree with the headers
removed. Scratch copy on purpose — the worktree is shared. It runs the binary
under a throwaway HOME so a test can never reach the live engram.
Exit status is the gate: the El tests print failures and still exit 0, so the
runner greps for FAIL lines and for a zero assertion count as well.
tests/test_history_amplification.el pins the invariant #129 violated: the
window the safety screen READS must be the window conv_history_record WRITES.
Not "must be called conv_history" — must AGREE.
THIS COMMIT IS RED BY DESIGN. On this tree the test fails one assertion:
3. REGRESSION #129 — agentic screen reads the session's own window
FAIL: distress history escalates the agentic screen to hard_bell
got: soft_bell
expected: hard_bell
history amplification tests: 8 passed, 1 failed (runner exit 1)
The next commit turns it green by changing one line. Two legs, one variable —
that is the whole point of committing the test first.
Two flaws in the older harness that this one does not copy: the idiom
`let pass_count = pass_count + 1` inside an assert function declares a local
that dies with the call, so every existing suite prints "0 passed, 0 failed"
regardless of outcome; and a test program without a `cgi` block compiles as a
'utility', which may not reference the self-formation primitives chat.el's
agentic loop calls — it fails to build on a capability violation it never
triggers at runtime.
Refs #129
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two fixes, one found by making the other.
1. HEBBIAN WRITE-BACK. This daemon learned 1,198 associations in 23h48m and
kept none of them: it syncs FROM the engram server and never pushes, and
mem_save() is unreachable in HTTP mode by design (soul.el only sets
soul_snapshot_path inside `is_genesis && safe_to_seed`, false whenever
ENGRAM_URL is set, because the server owns persistence). So the one process
that runs idle cognition -- where essentially all co-activation happens -- was
the one process that could not remember what it learned.
hebb_consolidate() now drains the runtime's write-back queue on every heartbeat
and POSTs it as ONE batch to /api/edges/batch. One request, one durable write,
not one 60MB snapshot per edge. Also drains on clean shutdown, so an exit
between beats doesn't take the last 8 minutes of learning with it.
_auth is required and its absence is silent: check_auth_ok exempts GET and
/api/neuron/state-events (which is why ise_post works keyless) but gates every
other mutation on "_auth" in the BODY -- http_serve surfaces no headers, so
there is no Bearer path. An unauthorized reply is NON-EMPTY, so the obvious
`if resp == "" return 0` check would have reported delivery of edges that were
refused, after the drain had already destroyed them. Caught before it shipped.
Gauges hebb_wb_pending/_drained/_dropped/_sent go into the heartbeat so a
consolidation path that stops delivering is visible in the stream.
2. A READ ROUTE MUST NEVER WRITE THE CANONICAL SNAPSHOT. GET /api/graph/edges
serialized this process's graph straight over $HOME/.neuron/engram/snapshot.json
-- the engram SERVER's durable store -- and read the edges back out of it. I
triggered it myself this morning fetching edges for the census above:
snapshot.json went from the server's 41,213 edges to the soul's 42,431, and the
next engram restart loaded the soul's graph as canonical. It happened to be a
superset (Knowledge 1198->1218, Memory 1238->1242, no durable type down), so
nothing was lost. That was luck. Had the soul been running a partial load --
the exact failure soul.el's safe_to_seed guard exists to catch -- one GET would
have destroyed the store, with no write-side guard able to see it coming.
The engram server fixed this same class of bug on 2026-07-21 by routing exports
to a dotted sidecar; the soul kept the original pattern. Same fix: exports go to
.soul-edges-export.json. Also stops a 60MB serialize-and-reread per GET.
Verified: boot 26 loaded 42,432 edges with hebb_max 0.4941 carried across the
restart -- the first time this daemon has ever started knowing what it learned.
Round 9.1, spec §3 D + ADR 0006 items 2 and 4. Two small changes, both proven
by measurement, both E2E-verified locally against a rebuilt brain.
D1 — SIGPIPE/EPIPE survival (vendor/el-runtime el_runtime.c).
Root cause, at the layer that owns it: the whole HTTP server lives in the C
runtime; .el has no socket primitive. http_send_all() called send() with flags
0 and nothing anywhere in the runtime set a SIGPIPE disposition, so the default
disposition — terminate the process — applied. When a handler finished after
its client had gone (Tim's VM: reply at 116.9 s, client cancelled at 25.0 s),
the second of the four sends that write one reply raised SIGPIPE and the daemon
died: `exited due to SIGPIPE ... ran for 361177ms`, launchd respawn 4 ms later,
every other in-flight session's work lost, user never told.
Fix: SIGPIPE -> SIG_IGN at runtime init and at each http_serve* entry, plus
per-connection SO_NOSIGPIPE / MSG_NOSIGNAL so the guard survives an embedder
resetting dispositions. http_send_all now retries EINTR and preserves errno;
http_send_response classifies it once — a departure is logged as routine
("client left before the reply was written ... reply discarded") and ANY other
errno is logged as a real "send failed: <strerror>". Spec §5.3: the routine
case must not mask a genuine write fault, and it does not.
Proof (scratch HOME + free port, 3 disconnects mid-reply):
round-9 shipped brain 4402179554… — DIED, exit 141 (128+13 = SIGPIPE), round 1
round-9 sources rebuilt with this exact recipe — DIED, exit 141, round 1
this build — SURVIVED 3/3, /health 200 after, still serving the full graph,
three honest "client left" lines in the log naming Broken pipe / Connection
reset by peer.
D2 — the round-start marker (chat.el, agentic_loop).
The ledger only ever appended AFTER a round returned, so a healthy first leg
produced zero progress by construction; since server-side web_search moved
inside the outbound call that leg is 60-120 s of silence, which is how a 25 s
client watchdog came to kill a healthy mission. One entry,
{"i":N,"t":"","tool":"__working__"}, written to the existing
run_progress_<session_id> ledger BEFORE each round's outbound call — the wire
shape ChatView.kt:1148 has handled as a life signal since 2026-07-13 and never
received. No new key, no new route, no new lifecycle: a strict subset of WS3
item 3. WS3's run registry is untouched and stays Will's.
Proof (live Anthropic key, real research mission, scratch HOME + free port):
round-9 baseline — ledger EMPTY for the whole 59.7 s leg
this build — {"i":0,"t":"","tool":"__working__"} visible at 18.6 s of a
70.0 s leg; both builds returned correct ~4.9 KB answers
Regression: prompt-matrix gate 32/32 on this build (round-9 baseline also 32/32
under the same recipe, so the score is not a build artifact). Soul contract
gate PASS — 27/27 routes, immutability clean. neuron#111 miscompile guard: 0
sites in the generated amalgam this binary was compiled from.
NOT included, deliberately: the regenerated dist/soul.c. CI compiles that file,
so production stays exposed until it is regenerated — the same open ask as
neuron#111 / ui#209. The regen recipe is now known and recorded; landing it is
Will's call, per BUILD-HYGIENE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Teaches the OpenAI-format lane (Groq/OpenAI/Grok/Gemini/Ollama) to offer tools,
execute them, and loop — the capability that until now existed only on the
Anthropic wire. The tool-execution, consent, bridge and run-progress machinery is
reused unchanged; only the wire dialect is new.
Two pre-existing defects were found while proving it, and are fixed here because
both silently break chat:
1. PROVIDER WIRING NEVER CONNECTED. The launcher exports SOUL_LLM_PROVIDER /
SOUL_LLM_BASE_URL and puts the provider key in ANTHROPIC_API_KEY + SOUL_API_KEY;
the engine's provider fork read only NEURON_LLM_0_*, which nothing sets in a
customer build. So use_openai was ALWAYS false: every non-Anthropic user's turns
went to api.anthropic.com carrying, say, a Groq key, and came back
"llm unavailable". Proven side-by-side against the pinned round-9 brain
(sha256 15cf7d1b…): identical env, shipped brain = "llm unavailable" both chat
modes with ZERO calls to the configured endpoint; this build = a real answer,
with the probe logging POST /v1/chat/completions and Bearer <provider key>.
Fixed brain-side only (env fallbacks) — no app or launcher change needed.
2. TRUNCATION SPLITS UTF-8 CHARACTERS. The session preload cuts recalled memory at
fixed BYTE lengths (continuity snippet 350; session_preload_bullets per bullet).
A cut landing inside a multi-byte character leaves a dangling lead byte in the
SYSTEM PROMPT, making the whole request body invalid UTF-8 — providers reject it
and the user sees an unexplained failure. Captured from a real body: 18,710 bytes,
decode fails at 18,248 on 'e2', a box-drawing rule (U+2500 = E2 94 80) sliced in
half. Trigger is ordinary content — em dash, curly quote, accented name, emoji,
table border — and it gets MORE likely as memory grows. Shared code: this hit the
Anthropic wire too. Fixed with utf8_safe_slice() applied at BOTH cut sites.
WHAT IS IN THE PORT
- llm_base_url / llm_wire_format / agentic_api_key: fall back to the launcher's own
SOUL_LLM_* names; anthropic deliberately still returns "" so its native path is
untouched (endpoint configurability remains neuron#62).
- openai_tools_json(): Anthropic tool schema -> OpenAI function schema; entries with
no input_schema (Anthropic's server-side web_search) are skipped — they cannot
execute on this wire.
- agentic_tools_no_web(): the standard set minus that server tool.
- openai_agentic_loop(): forked rather than parameterised, so agentic_loop — which
carries every round-7/8/9 fix — is provably untouched. Same envelopes, same state
keys, same consent policy (ask_all / escalate / builtin / always-allow), same
client-bridge contract, same run-progress ledger, same 12-iteration cap.
- ADR-0005 mirrored on this wire: parallel_tool_calls:false is sent explicitly, and
if a provider ignores it we honour the FIRST call and echo only that one, so the
conversation we send is never self-contradictory. The drop is logged loudly.
- The assistant turn echoes the provider's own content bytes (json_get_raw), so a
JSON null stays null and nothing is lost to a decode/re-encode round trip.
- Tool results are embedded already-escaped (dispatch_tool json_safe's them);
truncation trims a dangling escape so a cut can't invalidate the body.
- bridge_save() gains a "wire" scalar and agentic_resume branches on it, so a
suspended turn resumes on the wire it suspended on. Legacy blobs (no field) resume
as anthropic. The field is read from the blob's SCALAR HEAD only — an unbounded
first-match scan would run on into messages_raw, which is model-controlled, and
that is exactly the round-9 resume defect. Pinned by a test.
- Three fork sites: handle_chat_agentic, handle_dharma_room_turn_agentic,
agentic_resume. Tool assembly is computed once per lane at both entry points
(it makes an HTTP call to the connector bridge; it was being paid for twice).
TOOLING THAT DID NOT EXIST
- tests/run-el-test.sh — engine tests were never runnable: elc is a compiler, it
emits C and exits. This emits the test to C, compiles soul.c with main renamed
away, links the rest + the repo-pinned runtime, and runs it. It also COMPUTES THE
VERDICT, because every counted test file's "N passed, M failed" summary is a
permanent 0/0 — the counters increment inside if BLOCKS, which El scoping
discards (9 files; real fix filed as neuron#116). Proven to discriminate with a
deliberately-broken assertion.
- tests/gate-openai/ — deterministic OpenAI-dialect provider stub + scenarios +
driver + hostile modes, and a strict request validator that rejects any
Anthropic-shaped field so dialect leakage fails loudly.
VERIFICATION (rungs named)
- E2E-VERIFIED against a LIVE provider (Anthropic's OpenAI-compatible endpoint,
confirmed live): real answer; a tool call whose out-of-root path was DENIED by the
guard, after which the model refused to claim success ("I won't tell you I did it,
because I didn't"); then a valid path -> file physically on disk with exact content,
honest reply, ledger with per-round entries + {done:true}.
- Deterministic lane gate: 11/12 in both consent configurations (bridge + local);
hostile providers produce no hang and no fabricated answer; the 12-iteration cap
trips with its honest message. The one FAIL is oa-tools-off and is NOT this port —
see "Known, not fixed here".
- ANTHROPIC LANE UNCHANGED: gate9 32/32 on this build and on the pinned round-9
brain; request bytes differ only within the noise band that two runs of the
UNMODIFIED brain also produce (proven with a baseline-vs-baseline control), and
the preload sections — the shared code touched here — are byte-identical.
The rig discriminates: the round-8 brain scores 24/32 on it.
- verify-soul-contract.sh: PASS (27/27 routes, no hard-deletes).
- Unit: test_bridge_serialization 36/36 (incl. 8 new wire/field-order assertions),
test_utf8_slice 18/18, test_agentic_tools 18 PASS / 0 FAIL / 3 documented skips.
KNOWN, NOT FIXED HERE (deliberate)
- Tools:Off on an OpenAI provider still fails: the non-agentic path goes through the
el-runtime provider chain, which appends /v1/chat/completions to a base URL that
already ends in /v1 -> /v1/v1/... 404. Runtime/plain-chat territory, untouched
mid-beta. Note openai_chat_complete() has zero callers — that lane is served
entirely by the runtime chain.
- The 12-iteration cap does not bound a chain of BRIDGED tools (iteration is
per-invocation and resume starts fresh). Parity with the Anthropic lane.
- run_progress resets on each resume, so a client rendering cumulative steps across a
consent pause sees earlier legs vanish. Parity with the Anthropic lane.
- verify-soul-contract.sh needs bash >= 4; under macOS's stock bash 3.2 it dies
instantly with a FALSE red ("local: -n: invalid option").
- Groq-specific live E2E not run: no Groq key exists on this machine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ROOT CAUSE (round 9; live-repro'd 5/5 this morning, both faces stub-proven by the
prompt-matrix gate). json_get is a first-substring-match scanner (strstr for
'"key":', el_runtime.c). bridge_save serialized the RAW messages array BEFORE the
tool_use_id scalar, so agentic_resume's json_get(blob, 'tool_use_id') returned the
FIRST '"tool_use_id":' occurrence inside the replayed conversation, not the saved
field. The resume guard then preferred that misread over the client's correct
call_id (its two branches both reduced to saved_use_id), attached the tool_result
to the wrong id, and Anthropic 400'd the resume ('unexpected tool_use_id found in
tool_result blocks'), surfaced as {"error":"llm unavailable"}.
ONE MISREAD, TWO FACES — whichever block owns the first tool_use_id in the array:
FACE 1 (search-then-bridge, the Key West killer): the first occurrence is the
first web_search_tool_result's srvtoolu_… id — every agentic turn that ran
server-side web_search and then bridged on a client tool died on approval,
deterministically (messages.2.content.0 … srvtoolu_…). The write itself had
already succeeded; only the resume died.
FACE 2 (multi-cycle missions): with no search, the first occurrence is ROUND 0's
tool_result block — so every LATER approve/resume cycle replayed the round-0
client id (stale-resume-id), killing multi-file missions after ~2 files.
And the shape that PASSES on round 8 confirms the mechanism: a single-cycle
bridge with no prior tool round has no 'tool_use_id' substring in its messages
at all (tool_use blocks carry 'id'), so the scan fell through to the blob's own
field and resumed correctly.
The server_tool_use ↔ web_search_tool_result pairs themselves replay intact — the
defect was a cross-field misread of the blob, the same first-match-scanner class
as BUG-6 (approve 'content' matched inside tool_input, 2026-07-17) and round 8's
citation-block fix.
THE FIX, the pattern not the spot:
1. bridge_save writes every json_safe'd scalar BEFORE both raw fields (an escaped
value cannot contain a bare '"key":' byte pattern, so first-match always lands
on the blob's own fields), and tools_raw (our fixed schema) before messages_raw
(arbitrary conversation), so the raw extractions cannot first-match into
model-controlled bytes either. Field order documented as load-bearing.
2. agentic_resume now honors the client's echoed call_id when present — the value
with clean provenance (minted from pend_tool_id, never blob-round-tripped) —
falling back to the saved id only when the client omits it. Each approve cycle
therefore binds to ITS OWN round's id (kills FACE 2 even against a blob written
by a pre-fix binary), and an omitted call_id still resumes on the saved id,
which the reordered blob now reads correctly.
Pattern sweep: the legacy synthetic blob (sessions.el handle_session_approve) embeds
only json_safe'd fields — no raw hazard, untouched. No other json_get read of any
container that embeds raw conversation JSON before the read field.
PROOF: prompt-matrix gate 24/32 RED on the round-8 brain (fails exactly the two
resume classes, named) -> 32/32 GREEN on this build; live-key Key West tracer
3/3 consecutive full round-trips (bridge -> approve-as-the-app -> real completion,
file on disk), plain-chat and weather-only controls PASS; unpatched round-8 brain
and a same-toolchain unpatched baseline build both still fail the identical
sequence with the identical srvtoolu 400 (the test discriminates, and the only
variable between failing and passing builds is this diff).
Refs neuron#109
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
str_eq("", "") is true, so an auto-term extractor that kept FAILING reported a
rising auto_term_streak. The signal meaning "fixated on one term" and the
signal meaning "producing no term at all" were the same number — opposite
failures needing opposite responses. Observed live as
{"auto_term":"","auto_term_streak":3}. Same class of bug already fixed for
wm_top0_streak on 2026-07-31; auto_term was missed then.
Empty now reads 0, and the empty run is counted on its own axis
(auto_term_empty_streak) so extractor failure is visible rather than disguised
as health.
Also surfaces the new runtime gauges in the heartbeat: hebb_warm, hebb_max,
hebb_links (is the graph learning any structure at all?) and dup_wm_global.
CAUGHT BY A/B, AND ONLY BY A/B. The previous commit's receipt_strip assumed the receipt is
always TERMINAL and cut everything from the marker onward. It is not always terminal: once
receipt_rule told the model what [[RECEIPT ...]] means, the model sometimes LED with one and
wrote the answer underneath. Cutting at the marker then deleted the entire answer and the
turn returned {"error":"no response"}.
MEASURED, same prompt (two web searches, cited prose), fresh session each run:
round-7 brain 4 / 4 answered (471, 473, 473, 544 chars)
round-8 brain 2 / 7 answered (five {"error":"no response"})
This looked exactly like a flaky model. It was not — it was mine. Running the two brains
side by side on the same prompt is the only reason it was found, and it is the reason the
A/B is now part of how this class gets tested.
AFTER THE FIX, same protocol:
round-7 brain 4 / 4 (473, 473, 473, 544)
round-8 brain 4 / 4 (657, 657, 657, 673)
THE FIX: remove the [[...]] span and keep BOTH sides, instead of truncating at the marker.
An unterminated marker at position 0 is left completely alone — no rule about receipts is
worth erasing an answer over. Bounded four-pass loop rather than a conditional exit, because
rebinding the counter inside an if-expression is the block-expression shape that miscompiles
integer arithmetic under this elc (BUG-PLAINCHAT-1). Verified in the generated C:
str_slice(rest, (e + 2), str_len(rest)) <- integer addition, correct
el_str_concat(head, tail) <- string concat, correct
and zero el_str_concat(<ident>, str_len(...)) sites across all 49 modules.
SEAM PROOF (FIX C) rides on the same runs — a real two-search cited answer, inspected byte
by byte, in BOTH failure directions:
missing separator (the round-7 "to.Good", bytes 77 2e 47): 0 hits. Sentence boundaries
measure 2e 20 4d — "." SPACE "M".
over-separation (a cited sentence shattered across paragraphs): 0 hits. The answer is one
continuous paragraph with its sentences intact, which is the direction a blanket
separator would have broken.
BUILT: sha256 77115f2733e794c5bc4ad1f55b1acaf658f8f4a91cccd423a2d633d94a726cbc
Refs neuron#109
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
FOUND BY E2E, NOT BY REASONING. The previous commit's design note asserted the provenance
receipt "never reaches the user: it is appended to the history copy, not the reply." That
was FALSE, and only running the thing showed it. On the first live run against the built
DMG brain, two agentic turns out of two came back with
"Your favourite colour is chartreuse and your project is called Perihelion.
[[RECEIPT - recorded by the soul, not written by the model: no tools ran on this turn.]]"
— the receipt in the user-visible reply.
MECHANISM: the receipt is stored inside the assistant turn, and the agentic path replays
history VERBATIM as Anthropic message objects. So the model sees its own previous answers
ending in [[RECEIPT ...]] and does the obvious thing — it imitates the format and signs the
next answer the same way. The plain path did NOT leak, which is the tell: there, history is
rendered into the SYSTEM prompt as labelled lines rather than replayed as assistant turns,
and a model imitates its own turns far more readily than a transcript.
FIX, two layers, because one of them is not a guarantee:
- receipt_rule() names the marker in both system prompts (plain and agentic): these lines
are written by the system, read them as evidence, never write one. Reduces occurrence.
- receipt_strip() truncates any [[RECEIPT ...]] out of model output before it becomes the
reply — plain path in layered_generate, agentic path on final_text in agentic_loop.
Deterministic. A guard that depends on the model choosing to obey is exactly the class of
thing round 8 exists to stop shipping, so the instruction is the optimisation and the
strip is the guarantee.
Placed ABOVE agentic_loop's empty-check on purpose: a turn whose entire output was an
imitated receipt has produced no answer, and must be reported as no answer.
The receipt stays in HISTORY, which is the whole point and is proven to work: asked "What
source did you use for that?" one turn after a live web_search, this brain answered
"I used Weather Underground (https://www.wunderground.com/weather/is/reykjav%C3%ADk) for the
current temperature in Reykjavik" — a real source, no apology. That is the false confession
dead, and it is dead BECAUSE the model can read the receipt.
BUILT: 887,112 bytes, sha256 54a2eff84d4fa44f8d2db6781dcf225df4b5075a8ec5f1058018bd40cc1af10b
BUG-PLAINCHAT-1 miscompile guard: zero sites.
Refs neuron#109
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
DESIGN FIT: three of round 7's five defects share ONE root — the conversation-history
layer persists only {role, content}, discarding tool provenance, session scoping, and the
distinction between a real user turn and an internal utility call. Fixes A and B RESTORE
Will's design rather than extend it: his agentic path already scopes history per session,
the plain path never got it, and his own source carries the TODO admitting the resulting
race (chat.el, handle_chat: "process-global key; concurrent /api/chat requests without
session_id race on this read-append-write"). Fix C repairs one join Will wrote that was
correct for a year and one we added last week. E1/E2 are ours.
FIX A — tool provenance in history (kills the FALSE CONFESSION)
Root cause, EXECUTED-verified: handle_chat_agentic recorded turns via hist_append, which
emits {"role","content"} only. server_tool_use blocks, web_search_tool_result blocks and
every citation were discarded, then replayed as text. On the next turn the model saw a
data-rich answer with zero evidence a search had happened, and its own permanent rule
("never describe a search you did not perform") left one conclusion available: that it
had fabricated the data. It apologised for a search it HAD run — four independent lines
of evidence confirm the search was real. The defect is not the model's honesty. It is
that we deleted the evidence and then asked it to account for itself.
Change: agentic_loop accumulates the source URLs it already walks past (citations and
web_search_tool_result content) and returns them as "sources"; handle_chat_agentic folds
tools_used + sources into a receipt line stored WITH the assistant turn. Receipts are
unconditional — a negative receipt ("no tools ran") is the other half of the guarantee,
because "no evidence of a tool" and "evidence of no tool" were previously identical in
the transcript. conv_history_block splits the receipt off before snipping so a long
answer cannot truncate away the evidence. The user never sees it: it is appended to the
history copy, not the reply.
FIX B — one history key for both paths (kills the BLANK STARE)
Root cause, EXECUTED-verified: the agentic path keyed history on session_hist_<id>; the
plain path was hard-wired to the process-global conv_history and never read session_id.
One conversation, two buckets. Proven in the guest engram: the scoped node held exactly
two turns starting at "Try again" while the earlier exchanges sat unscoped.
Change: conv_hist_key/conv_hist_label are now the single definition, used by BOTH paths;
session_id is threaded route -> layered_cycle -> layered_generate / conv_history_record.
The 2-line fallback (plain path reads the agentic key) was REJECTED: it keeps the
process-global bucket as a live write target, which is the bleed the TODO describes.
Also found and closed while threading: layered_cycle read session_id from the state key
"current_session_id", which is read here and WRITTEN NOWHERE in the entire source. It
was unconditionally "", so TODO(reliability #4) — per-session steward continuity — was
dead code that could never fire. It fires now.
LAZY SESSION, decided explicitly: we create the session EAGERLY at the door (app half,
ui#223) rather than migrating orphaned turns. Migration would copy the CONTENTS of a
process-global bucket, possibly another conversation's, into a named session — the bleed,
performed deliberately. Eager creation makes the situation impossible instead. Migration
is deliberately not implemented and must not be added without solving provenance first.
FIX C — the two text-join seams ("to.Good", byte-verified 0x77 0x2e 0x47)
Two bare `+` joins, written a year apart, had drifted into two answers to one question:
within-response block joins (Will's, 2026-05-03, latent until server-side web_search
began interleaving non-text blocks) and across-round joins (ours, 62af564).
Change: one named rule, text_join_sep, at both sites. NOT a blanket separator — a cited
answer splits MID-SENTENCE ("The current temperature is " + "86°F" + ", with "), so a
blanket separator shatters every sourced sentence. The rule takes the one bit that
distinguishes the cases: whether a NON-TEXT block intervened. Hoisting it also makes the
fix verifiable in the shipped binary, which an inline `+` is not.
FIX E1 — utility generations stay out of the transcript
Title generation ("Write a 3-6 word title...") and insight passes ran down the same plain
door as a real message and were recorded as if the user had typed them; the same calls are
the "model":"unknown" rows in usage.jsonl. is_utility_request reads an explicit utility
flag from the app, with the __title__/__insight__ id prefixes as a fallback for older
clients. Answered normally, never recorded.
FIX E2 — OPERATOR IDENTITY is scoped to tool-capable turns
The block (env USER/HOME, closing "This is a hard rule") was prepended to EVERY system
prompt including chat mode. On a Tools:Off turn there is no filesystem in reach, so it
governed nothing and merely supplied the loudest fact in the prompt — which is why the
model opened a fresh conversation with "You're test, on your machine at /Users/test".
Hoisted to operator_identity_block() and gated on !chat_mode. Unchanged wherever a file
or command tool can actually be reached.
ALSO: agentic_loop's per-session history persist had a second hand-rolled copy of
conv_history_persist with a different label expression, different salience scores and
different tags for the same node. Since both now derive the label from conv_hist_label and
engram_node_full upserts by label, two score policies were writing one node. Collapsed to
one writer.
BUILD NOTE: dist/elp-c-decls.h is force-included by the documented link recipe and carried
the OLD C arities, so it is updated here. This is the build-support header, NOT the stale
generated dist/soul.c — no dist/*.c was read or edited; all engine changes are .el source.
chat.elh/soul.elh are committed because a first-pass build against the old signatures FAILS
(measured); the other regenerated headers are reverted as unrelated churn.
BUILT: 887,000 bytes, sha256 d632b061ad75269d6adeb52578d030eaf49e895d91289d7f946b19c08450d728
Zero el_str_concat(<int>, str_len(...)) sites (the BUG-PLAINCHAT-1 miscompile guard).
web_search_20250305 and disable_parallel_tool_use both still present — PR #108's web search
and the ADR-0005 stopgap are intact.
Refs neuron#109 (builds on it), neuron#78 (Receipt Contract — the real fix A is a stopgap for)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>