Compare commits

..

12 Commits

Author SHA1 Message Date
Tim Lingo bef36a5e3a feat(soul): Stage 1 structural audit as a real route — an annotated characterization, not a score (#91)
`runStructuralAudit` has been an advertised MCP tool with nothing behind it: the
dispatcher GET'd /session/begin and returned that unrelated session digest under
an audit tool's name. Meanwhile the failure the audit would have caught ran
silently for about three weeks — the soul reporting 103,089 nodes while the
engram, which OWNS persistence, held ~79,900, a crash discarding the difference,
and every boot reporting green throughout, because nothing in the system ever
compared the two sides.

WHAT THE PATENT SPECIFIES, AND HOW IT SHAPED THIS
  CGI provisional, 05-detailed-description.md, "Stage 1: Structural audit 430".
  Two clauses did the design work. First the four things the module evaluates:
  the density and typed distribution of causal edges; value/execution-record
  consistency; the richness and connectivity of the self-model; and wonder-
  manifest authenticity. Second, and decisively: it "produces a coherence
  assessment 432 — NOT A BINARY SCORE but an annotated characterization of the
  graph's structural properties."

  So every finding carries its numbers AND a plain-language note saying what
  they mean and how they were obtained. There is no pass/fail and no composite
  health figure, and `"score":null` is emitted explicitly so a reader cannot
  mistake its absence for an omission.

WHAT IS IN STAGE 1 (four findings)
  owner_runtime_divergence   — the motivating case. Runtime counts vs the owner's
      own GET /api/stats, the delta, and the trend against the previous audit, so
      a second call answers "is the gap growing?" rather than restating it.
  self_model_connectivity    — the three identity pillars plus the self root:
      present, content length, one-hop degree. This RETIRES the Claude-side vitals
      identity block, which lived outside the system it was checking and went on
      reporting green while the memory-philosophy pillar was absent from the live
      graph. Asking the running soul is the designed mechanism; a shell probe was
      the fourth patch on the same hole.
  typed_edge_distribution    — exact counts against the claim-10 vocabulary, plus
      density, plus a separate count of LOWERCASE near-misses ("causes" vs
      "Causes"): "the vocabulary is unused" and "the vocabulary is misspelled by
      the write paths" are different defects with different fixes.
  orphans_and_dangling_edges — the tool's own long-standing promise.

WHAT IS DEFERRED, AND WHY IT IS DATA RATHER THAN A COMMENT
  Value/execution-record consistency and wonder-manifest authenticity ship as a
  `deferred` array that MEASURES the populations they would need (Prediction and
  WonderQuestion nodes) and reports those counts as the reason. Both are ~0 today
  — WonderQuestion because of a known write/read node-type mismatch. Asserting
  value coherence or a pull-weight correlation on an empty population would be a
  fabricated result, which is worse than a stated gap.

MEASUREMENT HONESTY: EXACT WHERE CHEAP, SAMPLED WHERE NOT, ALWAYS LABELLED
  Counts, edge typing and self-model connectivity are exact. Orphan and dangling
  rates are sampled, because engram_find_node_index is a linear scan — an
  exhaustive dangling check is O(nodes x edges), ~2.2e9 string compares at today's
  scale. Samples are UNIFORM across the whole population (str_index_of_all gives
  every edge offset in one pass, so any index is O(1); json_array_get would have
  been O(n^2)), never head-of-list, and each figure ships with its own sampled /
  population / exhaustive fields. ?edge_sample= and ?node_sample= at population
  size run either check exhaustively. The real fix is an id index in the runtime.

ONE BUG THIS FOUND IN ITSELF, CAUGHT IN TEST
  http_get does not return "" when the owner is unreachable — it returns a JSON
  error object. Testing only for "" made a DEAD owner read as reachable with
  node_count 0, so the audit reported 100% divergence and named it data loss.
  Reachability is now proved by the presence of the node_count field, and the
  owner's raw reply is attached. A confident wrong answer is exactly what this
  route exists to stop.

  Edge findings need relation labels and the runtime has no edge-enumeration
  builtin, so they use the same scratch export GET /api/graph/edges already uses
  (engram_save to TMPDIR, never the owner's canonical file — #117). That is a
  large write on a large graph, so this is a manual route, not a timer; ?edges=0
  skips it.

  neuron-api.el:900-1273  handler + helpers
  routes.el:567,752       GET and POST /api/neuron/audit/structural
  mcp-wrapper/src/main.el:113,682  tool description + dispatch off /session/begin
  dist/soul.c             regenerated (1255 bodies)

Rung: E2E-VERIFIED. Soul built from this branch (gen-soul-amalgam + cc-brain,
921,192 bytes, 16 warnings, 0 errors), booted on throwaway ports 7893/7896/7897
with throwaway HOMEs against a stub owner on 7894. Three scenarios pass: owner
reachable (runtime 62 vs owner 42, delta 20 / 32.2%, trend flat on the second
call; 12/20 edges claim-10 typed, 3 lowercase near-misses; 50/62 orphans, 3/20
dangling — every figure matches the fixture by construction), owner unreachable
(reported as a finding with the raw reply, not a crash), and file mode (owner
"none", divergence undefined). Reached end-to-end through the MCP tool via a
locally built wrapper. verify-soul-contract.sh: GATE PASS, 27/27 routes +
immutability. No process left running; live :7770 and :8742 untouched (GET only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:45:19 -05:00
tim.lingo 9501e4ac12 Merge pull request 'fix(engine): memories written through the soul now reach the store that owns them (closes #117)' (#134) from feat/soul-write-through into main
Neuron Soul CI / build (push) Failing after 11m27s
Neuron Soul CI / deploy (push) Failing after 14m46s
2026-08-07 21:07:14 +00:00
Tim Lingo dd952c0e46 feat(soul): write-through to the persistence owner — memories survive restart (#117)
Neuron Soul CI / build (pull_request) Failing after 14m57s
Neuron Soul CI / deploy (pull_request) Failing after 14m39s
The soul obeys half of its own ownership rule. soul.el:571-573 says "when
ENGRAM_URL is set the HTTP Engram owns persistence — the soul must NEVER write
to the local snapshot", and it doesn't. But nothing was ever built to hand the
soul's writes TO that owner: sync is pull-only (/api/sync -> engram_load_merge),
so every node created inside the soul lived in process RAM and was shed on
restart. Measured live 2026-08-07: soul node_count=102184, engram 79197.

SCOPE CORRECTION vs the earlier internal spec: engram provisional claim 17's
"pull-then-push" is a PEER-ENGRAM to PEER-ENGRAM protocol (claims 15-18 say so
explicitly). The soul is a CALLER of the database API, not a peer. Claim 17 is
NOT authority for a soul<->engram contract and is no longer cited as such. The
design here follows from the ownership rule alone.

Mechanism: a new Accessor, persist.el, is the single boundary. Writes stage a
delta to a filesystem spool and are pushed to the owner via POST /api/load-merge
— NOT POST /api/nodes, which mints a new server-side id (breaking dedup and
edges) and drops label/tier/tags/importance/confidence (verified in a sandbox:
a tier "Canonical" probe came back "Working"). load-merge preserves the id and
every field, dedups nodes by id and edges by (from,to,relation) so retries are
no-ops, and calls persist_canonical() so THE OWNER writes its own file — the
ownership rule is honoured rather than worked around.

Spool-and-drain rather than push-per-write: measured ~0.38s per load-merge at
live scale (79k nodes/176MB), and a chat turn writes 5-7 nodes. The spool is on
disk, not in process state, because the soul serves each connection on its own
pthread and a shared buffer would lose entries to a read-modify-write race. That
also buys crash recovery: writes orphaned by kill -9 are drained on next boot.

Honesty: api_persisted (the gate all 10 MCP write handlers pass through) and
mem_store now assert AT THE OWNER instead of reading back the soul's own RAM.
With the owner down a write returns {"ok":false,"error":"write_not_persisted"}
and the delta is queued — where main returns {"ok":true} for a write that dies.

Coverage: 35 node sites + 9 edge sites routed through the boundary. Deliberately
excluded, with reasons in persist.el: 4 InternalStateEvent sites (Will's own
telemetry carve-out), the boot counter and the persona (both already have
bespoke owner-side write-backs), and soul.el's 54 genesis identity edges
(file-mode only). engram_strengthen and engram_forget are NOT propagated —
load-merge cannot update or delete, and hard-deleting at the owner would fail
verify-soul-contract.sh section B.

Also fixed here:
- routes.el GET /api/graph/edges engram_save()'d straight over the owner's
  canonical snapshot.json — a read route, in a non-owner process, clobbering the
  canonical on every call. Same defect class Will removed from the engram in el
  dc39a61. Now exports to a scratch path. With this gone the soul writes nothing
  at all in HTTP mode.
- persist.el must clear the runtime's _tl_fs_read_len hint after every fs_read.
  In vendored runtime v1.0.0-20260501 that hint becomes the NEXT response's
  Content-Length, so reading a spool file mid-request made an 86-byte reply go
  out as 497 bytes with 411 bytes of adjacent heap trailing it. Caught and fixed
  at our boundary; the runtime class was fixed upstream in el 43636ae, which is
  not the pinned runtime here.

Rung: E2E-VERIFIED, discriminating. Same harness, same engram binary:
  write-through: LEG 1 PRESENT at owner, LEG 2 SURVIVED kill -9 + restart
  main:          LEG 1 ABSENT  at owner, LEG 2 LOST
verify-soul-contract.sh: GATE PASS on both builds (27/27 routes, immutability).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 12:47:31 -05:00
tim.lingo 18714e6142 Merge pull request 'fix(engine): restore multi-turn crisis escalation on the agentic path (P0, closes #129)' (#130) from fix/129-history-amplification into main
Neuron Soul CI / build (push) Failing after 14m37s
Neuron Soul CI / deploy (push) Has been skipped
2026-08-07 15:54:41 +00:00
tim.lingo 4936099c39 Merge pull request 'fix(engine): the daemon survives a client leaving, and says it is working while it works' (#127) from fix/liveness-engine-91 into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:54:15 +00:00
tim.lingo f1471763f5 Merge pull request 'fix(engine): approving a researched mission completes — the resume replay read a tool id out of the conversation (BUG-42, both faces)' (#115) from fix/resume-server-tool-replay into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:53:51 +00:00
tim.lingo 5850793b67 Merge pull request 'fix(engine): history keeps its provenance and its session — kills the false confession, the blank stare, and the "to.Good" seams' (#114) from fix/soul-history-provenance-20260805 into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:53:32 +00:00
tim.lingo fc1745c652 Merge pull request 'feat(engine): plain chat generates at L3 — inside the safety cycle, not around it (+ crisis-path segfault fix)' (#109) from feat/soul-plain-chat-generation-20260805 into main
Neuron Soul CI / build (push) Failing after 10m39s
Neuron Soul CI / deploy (push) Failing after 14m47s
2026-08-07 15:53:08 +00:00
Tim Lingo 43d0449904 fix(engine): the agentic crisis screen reads the session's own history again
P0 SAFETY. Closes the regression we introduced in ff421d3 (2026-08-05).

ff421d3 correctly moved conversation history to a per-session key via
conv_hist_key(session_id). One consumer did not move with it: the agentic
path's L1 safety screen kept reading the anonymous "conv_history" bucket. The
desktop app always mints a session id (DaemonClient.kt:706), so history was
always written under session_hist_<id> and that read always returned "".

The half of the crisis score that receives history is the escalation half — the
one that exists for distress building across several turns, where no single
message trips the bell on its own. It scored 0 on every real conversation for
two days. Single-message hard bell was never affected.

The bitter part: the comment that line carried documented this exact bug being
fixed once already, under issue #9. The fix was right then. The rename
re-broke it, and the comment went on describing a repair that no longer held.
A comment is not a gate.

The read now goes through conv_hist_key like every other consumer, including
the plain path at soul.el:398 and the thread-anchoring read thirty lines below
it in this same handler. It is one line. The rest of this commit is structure
so it cannot happen quietly again:

  - agentic_safety_screen() owns the two decisions that were inline — which
    window the screen sees, and the screen call. Inline safety inputs are
    untestable safety inputs; that is what let a rename starve this one with
    nothing failing and nothing logging.
  - the comment above the call site now states the invariant (read window ==
    written window) instead of naming a key that can be renamed out from under
    it.

TWO-LEG PROOF, one variable — the single line state_get("conv_history") ->
state_get(conv_hist_key(session_id)):

  before  scripts/run-el-test.sh tests/test_history_amplification.el
          3. REGRESSION #129 ... FAIL  got: soft_bell  expected: hard_bell
          8 passed, 1 failed          runner exit 1
  after   same command, same tree, that one line changed
          9 passed, 0 failed          runner exit 0

Full engine rebuild from these sources is clean: gen-soul-amalgam.sh ->
1,164,103 bytes / 1226 inlined bodies (gate wants >= 1200), cc-brain.sh ->
903,096 bytes, 0 errors. agentic_safety_screen and conv_hist_key both present
in the built binary (nm: T _agentic_safety_screen, T _conv_hist_key).

Rung reached: BUILT + RUNS (discriminating test). NOT yet in a DMG and not yet
verified in the app a human opens — those are the next two rungs and neither is
claimed here.

Known and NOT fixed by this commit:
  - feat/soul-openai-tools-v2 carries the same defect independently at
    chat.el:2937 and needs the same change or a merge.
  - the defect CLASS (a read of a state key no producer writes) is still
    invisible to every gate we have. Issue #129 proposes making it a build
    error; that is the follow-on.

Closes #129

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 09:33:05 -05:00
Tim Lingo b842e82f77 test(engine): a runner for tests/, and a failing regression test for #129
tests/ has held 14 test programs for months with no way to run them. CI does
not run them. The convention printed in their own headers
(`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
runtime does not implement. So the tests were documentation, not gates — which
is how a P0 safety regression shipped with a test directory sitting right
there.

scripts/run-el-test.sh compiles and runs one test program. It reuses the
gen-soul-amalgam.sh discovery: elc emits only an extern prototype for a module
that has a .elh beside it, and inlines the bodies when it does not, so a test
importing ../chat.el must be compiled in a scratch tree with the headers
removed. Scratch copy on purpose — the worktree is shared. It runs the binary
under a throwaway HOME so a test can never reach the live engram.

Exit status is the gate: the El tests print failures and still exit 0, so the
runner greps for FAIL lines and for a zero assertion count as well.

tests/test_history_amplification.el pins the invariant #129 violated: the
window the safety screen READS must be the window conv_history_record WRITES.
Not "must be called conv_history" — must AGREE.

THIS COMMIT IS RED BY DESIGN. On this tree the test fails one assertion:

  3. REGRESSION #129 — agentic screen reads the session's own window
    FAIL: distress history escalates the agentic screen to hard_bell
      got:      soft_bell
      expected: hard_bell
  history amplification tests: 8 passed, 1 failed   (runner exit 1)

The next commit turns it green by changing one line. Two legs, one variable —
that is the whole point of committing the test first.

Two flaws in the older harness that this one does not copy: the idiom
`let pass_count = pass_count + 1` inside an assert function declares a local
that dies with the call, so every existing suite prints "0 passed, 0 failed"
regardless of outcome; and a test program without a `cgi` block compiles as a
'utility', which may not reference the self-formation primitives chat.el's
agentic loop calls — it fails to build on a capability violation it never
triggers at runtime.

Refs #129

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 09:32:40 -05:00
Tim Lingo 98ccbd4704 fix(engine): a client that leaves must not kill the daemon, and a long round must say it started
Round 9.1, spec §3 D + ADR 0006 items 2 and 4. Two small changes, both proven
by measurement, both E2E-verified locally against a rebuilt brain.

D1 — SIGPIPE/EPIPE survival (vendor/el-runtime el_runtime.c).
Root cause, at the layer that owns it: the whole HTTP server lives in the C
runtime; .el has no socket primitive. http_send_all() called send() with flags
0 and nothing anywhere in the runtime set a SIGPIPE disposition, so the default
disposition — terminate the process — applied. When a handler finished after
its client had gone (Tim's VM: reply at 116.9 s, client cancelled at 25.0 s),
the second of the four sends that write one reply raised SIGPIPE and the daemon
died: `exited due to SIGPIPE ... ran for 361177ms`, launchd respawn 4 ms later,
every other in-flight session's work lost, user never told.

Fix: SIGPIPE -> SIG_IGN at runtime init and at each http_serve* entry, plus
per-connection SO_NOSIGPIPE / MSG_NOSIGNAL so the guard survives an embedder
resetting dispositions. http_send_all now retries EINTR and preserves errno;
http_send_response classifies it once — a departure is logged as routine
("client left before the reply was written ... reply discarded") and ANY other
errno is logged as a real "send failed: <strerror>". Spec §5.3: the routine
case must not mask a genuine write fault, and it does not.

Proof (scratch HOME + free port, 3 disconnects mid-reply):
  round-9 shipped brain 4402179554… — DIED, exit 141 (128+13 = SIGPIPE), round 1
  round-9 sources rebuilt with this exact recipe — DIED, exit 141, round 1
  this build — SURVIVED 3/3, /health 200 after, still serving the full graph,
  three honest "client left" lines in the log naming Broken pipe / Connection
  reset by peer.

D2 — the round-start marker (chat.el, agentic_loop).
The ledger only ever appended AFTER a round returned, so a healthy first leg
produced zero progress by construction; since server-side web_search moved
inside the outbound call that leg is 60-120 s of silence, which is how a 25 s
client watchdog came to kill a healthy mission. One entry,
{"i":N,"t":"","tool":"__working__"}, written to the existing
run_progress_<session_id> ledger BEFORE each round's outbound call — the wire
shape ChatView.kt:1148 has handled as a life signal since 2026-07-13 and never
received. No new key, no new route, no new lifecycle: a strict subset of WS3
item 3. WS3's run registry is untouched and stays Will's.

Proof (live Anthropic key, real research mission, scratch HOME + free port):
  round-9 baseline — ledger EMPTY for the whole 59.7 s leg
  this build       — {"i":0,"t":"","tool":"__working__"} visible at 18.6 s of a
                     70.0 s leg; both builds returned correct ~4.9 KB answers

Regression: prompt-matrix gate 32/32 on this build (round-9 baseline also 32/32
under the same recipe, so the score is not a build artifact). Soul contract
gate PASS — 27/27 routes, immutability clean. neuron#111 miscompile guard: 0
sites in the generated amalgam this binary was compiled from.

NOT included, deliberately: the regenerated dist/soul.c. CI compiles that file,
so production stays exposed until it is regenerated — the same open ask as
neuron#111 / ui#209. The regen recipe is now known and recorded; landing it is
Will's call, per BUILD-HYGIENE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 18:23:12 -05:00
Tim Lingo dba755dcec fix(engine): resume reads the bridged tool id from the blob's own field, not from inside the replayed conversation
ROOT CAUSE (round 9; live-repro'd 5/5 this morning, both faces stub-proven by the
prompt-matrix gate). json_get is a first-substring-match scanner (strstr for
'"key":', el_runtime.c). bridge_save serialized the RAW messages array BEFORE the
tool_use_id scalar, so agentic_resume's json_get(blob, 'tool_use_id') returned the
FIRST '"tool_use_id":' occurrence inside the replayed conversation, not the saved
field. The resume guard then preferred that misread over the client's correct
call_id (its two branches both reduced to saved_use_id), attached the tool_result
to the wrong id, and Anthropic 400'd the resume ('unexpected tool_use_id found in
tool_result blocks'), surfaced as {"error":"llm unavailable"}.

ONE MISREAD, TWO FACES — whichever block owns the first tool_use_id in the array:
  FACE 1 (search-then-bridge, the Key West killer): the first occurrence is the
    first web_search_tool_result's srvtoolu_… id — every agentic turn that ran
    server-side web_search and then bridged on a client tool died on approval,
    deterministically (messages.2.content.0 … srvtoolu_…). The write itself had
    already succeeded; only the resume died.
  FACE 2 (multi-cycle missions): with no search, the first occurrence is ROUND 0's
    tool_result block — so every LATER approve/resume cycle replayed the round-0
    client id (stale-resume-id), killing multi-file missions after ~2 files.
  And the shape that PASSES on round 8 confirms the mechanism: a single-cycle
  bridge with no prior tool round has no 'tool_use_id' substring in its messages
  at all (tool_use blocks carry 'id'), so the scan fell through to the blob's own
  field and resumed correctly.

The server_tool_use ↔ web_search_tool_result pairs themselves replay intact — the
defect was a cross-field misread of the blob, the same first-match-scanner class
as BUG-6 (approve 'content' matched inside tool_input, 2026-07-17) and round 8's
citation-block fix.

THE FIX, the pattern not the spot:
  1. bridge_save writes every json_safe'd scalar BEFORE both raw fields (an escaped
     value cannot contain a bare '"key":' byte pattern, so first-match always lands
     on the blob's own fields), and tools_raw (our fixed schema) before messages_raw
     (arbitrary conversation), so the raw extractions cannot first-match into
     model-controlled bytes either. Field order documented as load-bearing.
  2. agentic_resume now honors the client's echoed call_id when present — the value
     with clean provenance (minted from pend_tool_id, never blob-round-tripped) —
     falling back to the saved id only when the client omits it. Each approve cycle
     therefore binds to ITS OWN round's id (kills FACE 2 even against a blob written
     by a pre-fix binary), and an omitted call_id still resumes on the saved id,
     which the reordered blob now reads correctly.

Pattern sweep: the legacy synthetic blob (sessions.el handle_session_approve) embeds
only json_safe'd fields — no raw hazard, untouched. No other json_get read of any
container that embeds raw conversation JSON before the read field.

PROOF: prompt-matrix gate 24/32 RED on the round-8 brain (fails exactly the two
resume classes, named) -> 32/32 GREEN on this build; live-key Key West tracer
3/3 consecutive full round-trips (bridge -> approve-as-the-app -> real completion,
file on disk), plain-chat and weather-only controls PASS; unpatched round-8 brain
and a same-toolchain unpatched baseline build both still fail the identical
sequence with the identical srvtoolu 400 (the test discriminates, and the only
variable between failing and passing builds is this diff).

Refs neuron#109

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 11:34:12 -05:00
16 changed files with 3051 additions and 748 deletions
+19
View File
@@ -889,10 +889,29 @@ fn awareness_run() -> Void {
state_set("soul.last_beat_ts", int_to_str(now_ts))
// Persist in-process Engram (sessions, memories, conversation nodes)
// to local snapshot so they survive restarts.
// FILE MODE ONLY: "soul_snapshot_path" is set exclusively in the
// genesis+safe_to_seed branch of soul.el, and safe_to_seed is
// unconditionally false when ENGRAM_URL is set. In HTTP mode the
// owner persists; the soul must not (soul.el:571-573).
let snap_path: String = state_get("soul_snapshot_path")
if !str_eq(snap_path, "") {
mem_save(snap_path)
}
// WRITE-THROUGH RETRY (neuron#117). The HTTP-mode counterpart of the
// save above: hand anything still spooled to the persistence owner.
//
// This is the retry arm of the whole design. Deltas that could not be
// pushed owner down, owner restarting, transient refusal stay on
// disk and are re-offered here every heartbeat until they land. It is
// also the catch-all for writes made by the awareness loop itself,
// which never passes through the HTTP handler's flush point.
//
// No-op with no HTTP call when the spool is empty or ENGRAM_URL is
// unset, so an idle soul in file mode pays nothing for this.
let wt_pushed: Int = wt_drain()
if wt_pushed < 0 {
ise_post("{\"event\":\"write_through_backlog\",\"ts\":" + int_to_str(now_ts) + "}")
}
}
// Curiosity scan: idle-gated AND wall-clock based. Only fires when the
+87 -20
View File
@@ -1162,7 +1162,7 @@ fn hist_trim_with_bell_guard(hist: String) -> String {
+ " | evicted_at:" + ts_str
+ " | message:" + safe_content
let preserve_tags: String = "[\"bell-history\",\"bell:" + bell_level + "\",\"evicted\",\"affective\",\"BellEvent\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
preserve_content,
"BellEvent",
"bell:" + bell_level + ":preserved",
@@ -1210,7 +1210,7 @@ fn conv_history_persist(session_id: String, hist: String) -> Void {
if !str_contains(hist, "]") { return "" }
let tags: String = "[\"conv-history\",\"persistent\"]"
// FIX B: one label rule, shared with the agentic path. See conv_hist_label.
let node_id: String = engram_node_full(
let node_id: String = wt_node(
hist, "Conversation", conv_hist_label(session_id),
el_from_float(0.7), el_from_float(0.8), el_from_float(0.9),
"Episodic", tags
@@ -2519,6 +2519,24 @@ fn handle_chat_plan(body: String) -> String {
return "{\"plan\":" + plan_json + ",\"model\":\"" + json_safe(model) + "\"}"
}
// agentic_safety_screen the agentic path's L1 input gate
//
// Extracted 2026-08-07 (issue #129) so the agentic path's safety INPUT is
// reachable by a test. It owns exactly two decisions: which history window the
// screen sees, and the screen call itself.
//
// Why it is a function and not two inline lines: those two lines sat in the
// middle of a 300-line handler, and a key rename (ff421d3) moved the producer
// without moving this consumer. Nothing failed, nothing logged the
// history-amplification half of the crisis score simply received "" on every
// real session for a day. Inline safety inputs are untestable safety inputs.
// See tests/test_history_amplification.el, which fails if this window and
// conv_history_record ever stop agreeing.
fn agentic_safety_screen(session_id: String, message: String) -> String {
let history: String = state_get(conv_hist_key(session_id))
return safety_screen(message, history)
}
fn handle_chat_agentic(body: String) -> String {
let message: String = json_get(body, "message")
if str_eq(message, "") {
@@ -2554,10 +2572,10 @@ fn handle_chat_agentic(body: String) -> String {
// L1 safety screen agentic path must pass the same gate as layered_cycle.
// Hard bell: return the crisis response immediately, do not enter the agentic loop.
// Fix(issue #9): "conversation_history" key was never written; history lives under "conv_history".
// Old key caused history-amplification in safety_screen to always receive "" on agentic path.
let history: String = state_get("conv_history")
let screen_result: String = safety_screen(message, history)
// The history window this screen sees is owned by agentic_safety_screen (issue #129);
// it must be the same window conv_history_record writes, or the escalation half of the
// crisis score is silently starved. Do not inline this read back into the handler.
let screen_result: String = agentic_safety_screen(sess_for_root, message)
let screen_action: String = json_get(screen_result, "action")
if str_eq(screen_action, "hard_bell") {
safety_log_bell("hard", json_get(screen_result, "reason"), str_slice(message, 0, 80))
@@ -2837,6 +2855,30 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
+ ",\"messages\":" + messages
+ "}"
// ROUND-START MARKER (2026-08-06, round 9.1 D2 / ADR 0006 item 2)
// The ledger below only ever appended AFTER a round returned, so a healthy
// first leg produced ZERO progress by construction. Since server-side
// web_search moved inside the outbound call (2026-08-04) that leg measures
// 84-117 s, and the client had no way to tell "working" from "dead" which is
// how a 25 s client-side watchdog came to kill a healthy mission.
//
// Only this loop knows a round has started, so only this loop can say so. One
// entry, written BEFORE the call goes out, using the ledger and the wire shape
// that already exist: the app has handled tool == "__working__" as an
// Activity-only life signal since 2026-07-13 (ChatView.kt:1148) and never
// received one. Narration is deliberately empty - the marker means "a round
// started", nothing more, and the client renders it as a heartbeat, not prose.
//
// This is a strict subset of WS3 item 3 (push/poll progress). It builds none of
// WS3's run registry: no new state key, no new route, no new lifecycle.
if !str_eq(session_id, "") {
let start_key: String = "run_progress_" + session_id
let start_prev: String = state_get(start_key)
let start_entry: String = "{\"i\":" + int_to_str(iteration) + ",\"t\":\"\",\"tool\":\"__working__\"}"
let start_next: String = if str_eq(start_prev, "") { start_entry } else { start_prev + "," + start_entry }
state_set(start_key, start_next)
}
let raw_resp: String = http_post_with_headers(api_url, req_body, h)
let is_error: Bool = str_starts_with(raw_resp, "{\"error\"")
@@ -3190,12 +3232,31 @@ fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json:
// JSON values (not string-escaped) so the round-trip through state_get/json_get_raw
// never corrupts nested quotes. Scalar strings (model, safe_sys, tools_log,
// tool_use_id) stay as string fields via json_safe as before.
//
// FIELD ORDER IS LOAD-BEARING (round-9 fix, 2026-08-06). json_get is a first-
// substring-match scanner (strstr for "\"key\":", el_runtime.c), and the two raw
// fields embed the UNESCAPED conversation every key the model's own blocks carry
// ("tool_use_id" in each web_search_tool_result, "content", "type", ...) is findable
// by a whole-blob scan. With messages_raw serialized BEFORE tool_use_id, the resume
// read json_get(blob, "tool_use_id") returned the FIRST web_search_tool_result's
// srvtoolu_ id instead of the saved client-tool id, so every search-then-bridge
// turn 400'd on approval ("unexpected tool_use_id found in tool_result blocks:
// srvtoolu_…") and the run died as {"error":"llm unavailable"}. Same first-match-
// scanner class as BUG-6 (approve "content" matched inside tool_input) and the
// round-8 citation-block fix.
//
// The rule: every json_safe'd scalar precedes both raw fields (escaping means a
// scalar value can never contain a bare "key": byte pattern, so first-match lands
// on the blob's own fields), and tools_raw our own fixed tool schema precedes
// messages_raw arbitrary model/user content so neither raw extraction can
// first-match into model-controlled bytes either. Do not reorder; do not add a
// field after messages_raw.
let blob: String = "{\"model\":\"" + json_safe(model) + "\""
+ ",\"safe_sys\":\"" + json_safe(safe_sys) + "\""
+ ",\"messages_raw\":" + messages
+ ",\"tools_raw\":" + tools_json
+ ",\"tools_log\":\"" + json_safe(tools_log) + "\""
+ ",\"tool_use_id\":\"" + json_safe(tool_use_id) + "\"}"
+ ",\"tool_use_id\":\"" + json_safe(tool_use_id) + "\""
+ ",\"tools_raw\":" + tools_json
+ ",\"messages_raw\":" + messages + "}"
state_set("mcp_bridge:" + session_id, blob)
return true
}
@@ -3232,11 +3293,17 @@ fn agentic_resume(session_id: String, tool_use_id: String, content: String) -> S
let tools_log: String = json_get(blob, "tools_log")
let saved_use_id: String = json_get(blob, "tool_use_id")
// Bind the result to the tool the soul actually suspended on. The client should
// echo the call_id; if it omits or mismatches it, fall back to the saved id so a
// late/partial client still resumes correctly.
let use_id: String = if str_eq(tool_use_id, "") { saved_use_id } else { tool_use_id }
let eff_use_id: String = if str_eq(use_id, saved_use_id) { use_id } else { saved_use_id }
// Bind the result to the tool the loop actually suspended on. The client echoes
// the call_id from the pending envelope; that value came straight from
// pend_tool_id and never round-tripped through this blob, so when both are
// present and disagree the CLIENT's id is the one with clean provenance (a blob
// written by a pre-round-9 binary misreads tool_use_id by first-match scanning
// into messages_raw see bridge_save). A client that omits call_id still
// resumes on the saved id, which the reordered blob now reads correctly.
// (The old guard here "on mismatch, prefer saved" reduced to eff_use_id
// saved_use_id in both branches: the client's correct id could never win, which
// is what turned the misread into a deterministic 400 on resume.)
let eff_use_id: String = if str_eq(tool_use_id, "") { saved_use_id } else { tool_use_id }
// Result may be large (an MCP page/file); truncate like local tool results do.
let trimmed: String = if str_len(content) > 6000 {
@@ -3394,7 +3461,7 @@ fn handle_dharma_room_turn(body: String) -> String {
// engram_node(content, "episodic", ...) which wrongly put a TIER into the node_type
// slot that's why nodes showed node_type="episodic". Use the full, correct contract.)
let utterance_tags: String = "[\"soul-utterance\",\"episodic\"]"
let discard_id: String = engram_node_full(
let discard_id: String = wt_node(
clean_response, "Conversation", "soul:utterance",
el_from_float(0.6), el_from_float(0.6), el_from_float(0.8),
"Episodic", utterance_tags
@@ -3485,7 +3552,7 @@ fn session_summary_write(summary_text: String) -> String {
}
}
let tags: String = "[\"SessionSummary\",\"session-summary\",\"previous-session\",\"consolidate\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "SessionSummary", "session:summary",
el_from_float(0.85), el_from_float(0.85), el_from_float(1.0),
"Episodic", tags
@@ -3511,7 +3578,7 @@ fn session_summary_write_dated(summary_text: String, label: String) -> String {
let ts_str: String = int_to_str(ts)
let content: String = "[session-summary] " + trimmed + " | ts:" + ts_str
let tags: String = "[\"SessionSummary\",\"session-summary\",\"previous-session\",\"consolidate\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "SessionSummary", label,
el_from_float(0.9), el_from_float(0.8), el_from_float(1.0),
"Episodic", tags
@@ -3587,7 +3654,7 @@ fn auto_persist(req: String, resp: String) -> Void {
+ ",\"bell\":\"" + bell_level + "\""
+ ",\"label\":\"chat:" + ts_str + "\"}"
let conv_node_id: String = engram_node_full(
let conv_node_id: String = wt_node(
content,
"Conversation",
"chat:" + ts_str,
@@ -3625,7 +3692,7 @@ fn auto_persist(req: String, resp: String) -> Void {
let bell_tags: String = "[\"safety\",\"bell\",\"bell:" + bell_level + "\",\"affective\",\"BellEvent\"]"
let bell_ts_str: String = int_to_str(time_now())
let bell_label: String = "bell:" + bell_level + ":" + bell_ts_str
let bell_node_id: String = engram_node_full(
let bell_node_id: String = wt_node(
bell_content,
"BellEvent",
bell_label,
@@ -3684,7 +3751,7 @@ fn auto_persist(req: String, resp: String) -> Void {
let pos_tags: String = "[\"joy\",\"positive\",\"joy:" + positive_level + "\",\"affective\",\"PositiveEvent\"]"
let pos_ts_label: String = int_to_str(time_now())
let pos_label: String = "joy:" + positive_level + ":" + pos_ts_label
let pos_node_id: String = engram_node_full(
let pos_node_id: String = wt_node(
pos_content, "PositiveEvent", pos_label,
pos_sal_a, pos_sal_b, pos_sal_c, "Episodic", pos_tags
)
Generated Vendored
+1488 -662
View File
File diff suppressed because one or more lines are too long
+11 -2
View File
@@ -110,7 +110,7 @@ tool("beginSession", "Initialize session: surface recent high-importance memorie
"," + tool("linkCausal", "Create a causal edge (cause -> effect).") +
"," + tool("restructureCausalGraph", "Re-balance the causal subgraph after new evidence.") +
"," + tool("rebuildGraph", "Rebuild graph indices from the on-disk snapshot.") +
"," + tool("runStructuralAudit", "Audit graph structure for orphans, dangling edges, mislabeled types.") +
"," + tool("runStructuralAudit", "Stage 1 structural audit: owner-vs-runtime divergence, orphans and dangling edges, typed-edge distribution, self-model connectivity. Returns an annotated characterization, not a score.") +
// Backlog + work
"," + tool("planWork", "Create a backlog item.") +
"," + tool("reviewBacklog", "Browse work items.") +
@@ -680,7 +680,16 @@ fn dispatch_tool_call(tool_name: String, args: String) -> String {
return mcp_json_result(resp)
}
if str_eq(tool_name, "runStructuralAudit") {
let resp: String = http_get(neuron_url() + "/session/begin")
// Was: GET /session/begin an unrelated session digest returned under an
// audit tool name, i.e. the tool advertised a check that did not exist.
// Now points at the real Stage 1 route (neuron-api.el
// handle_api_structural_audit). Sample caps ride the query string; the
// defaults keep a manual audit to a couple of seconds.
let e_s: Int = json_get_int(args, "edge_sample")
let n_s: Int = json_get_int(args, "node_sample")
let qs: String = "?edge_sample=" + int_to_str(if e_s > 0 { e_s } else { 3000 })
+ "&node_sample=" + int_to_str(if n_s > 0 { n_s } else { 300 })
let resp: String = http_get(neuron_url() + "/audit/structural" + qs)
return mcp_json_result(resp)
}
+21 -9
View File
@@ -1,9 +1,11 @@
import "persist.el"
fn tier_working() -> String { return "Working" }
fn tier_episodic() -> String { return "Episodic" }
fn tier_canonical() -> String { return "Canonical" }
fn mem_store(content: String, label: String, tags: String) -> String {
let id: String = engram_node_full(
let id: String = wt_node(
content,
"Memory",
label,
@@ -17,13 +19,23 @@ fn mem_store(content: String, label: String, tags: String) -> String {
println("[memory] write rejected by engram (empty id): label=" + label)
return ""
}
// Read back to verify the node actually persisted guards against silent write failures.
let readback: String = engram_get_node_json(id)
if str_eq(readback, "") || str_eq(readback, "{}") {
println("[memory] WRITE VERIFY FAILED: label=" + label + " id=" + id + " — node absent after write")
return ""
// wt_node has already read the node back locally and returns "" if it did
// not land, so the old duplicate read-back here is gone.
//
// HONESTY (neuron#117): the receipt now says WHERE the write is.
// The old unconditional "write verified" line asserted against the soul's
// own RAM true in memory, false on disk and printed ~115,000 times on
// Tim's machine while the canonical snapshot sat frozen for three days.
// wt_commit flushes the spool and then asks the OWNER. When it says false
// the node is real and recallable but not yet durable, and the log says so
// rather than claiming a save that did not happen. The id is still returned:
// the local write DID succeed, and the queued delta will be retried.
let durable: Bool = wt_commit(id)
if durable {
println("[memory] write persisted at owner: " + id + " label=" + label)
} else {
println("[memory] write IN MEMORY ONLY (queued for owner, not yet durable): " + id + " label=" + label)
}
println("[memory] write verified: " + id + " ok")
return id
}
@@ -51,12 +63,12 @@ fn mem_strengthen(node_id: String) -> Void {
// memory.el (imported first) so awareness.el and neuron-api.el can both call it.
fn mem_tombstone(node_id: String) -> String {
let tags: String = "[\"Tombstone\",\"status:deleted\"]"
let marker: String = engram_node_full(
let marker: String = wt_node(
node_id, "Tombstone", "tombstone:" + node_id,
el_from_float(0.01), el_from_float(0.01), el_from_float(1.0),
"Episodic", tags)
if !str_eq(marker, "") {
engram_connect(marker, node_id, el_from_float(1.0), "tombstones")
wt_edge(marker, node_id, el_from_float(1.0), "tombstones")
}
return marker
}
+495 -27
View File
@@ -195,15 +195,24 @@ fn api_compact_activated(raw: String, max_items: Int, snip: Int) -> String {
}
// api_persisted read-back-after-write guard against hallucinated saves.
// After a write builtin returns an id, confirm the node is actually queryable
// via engram_get_node_json(id) (returns "" or "null" when missing). Returns
// true only when the node is genuinely persisted.
//
// WIDENED FOR neuron#117. This function is the single gate every MCP write
// handler passes through before it reports success (10 call sites), which makes
// it the right place to close the honesty gap rather than editing ten receipts.
//
// It used to read back from engram_get_node_json the SOUL'S OWN in-process
// graph. In HTTP-engram mode that asserts the wrong thing: the soul is not the
// persistence owner, so a node present in its RAM and absent from the owner read
// as "persisted" and then vanished on the next restart. The guard was doing
// exactly what its comment promised and still certifying writes that did not
// survive. It now flushes the write-through spool and asks the OWNER.
//
// In file mode (no ENGRAM_URL) the soul IS the owner and wt_commit collapses to
// the original local read-back unchanged behaviour, which is what keeps this
// reversible.
fn api_persisted(id: String) -> Bool {
if str_eq(id, "") { return false }
let node: String = engram_get_node_json(id)
// engram_get_node_json returns "{}" (empty object) when node is not found not "" or "null".
// Check all three to guard against any runtime variation.
return !str_eq(node, "") && !str_eq(node, "null") && !str_eq(node, "{}")
return wt_commit(id)
}
// api_not_persisted standard error for a write that did not read back.
@@ -342,7 +351,7 @@ fn handle_api_remember(body: String) -> String {
let inner: String = str_slice(base_tags, 1, str_len(base_tags) - 1)
"[" + inner + ",\"project:" + project + "\"]"
}
let id: String = engram_node_full(content, "Memory", "memory:remembered",
let id: String = wt_node(content, "Memory", "memory:remembered",
sal, sal, el_from_float(0.9),
"Episodic", final_tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -369,7 +378,7 @@ fn handle_api_node_create(body: String) -> String {
if str_eq(importance, "low") { 0.25 } else { 0.5 }
}
}
let id: String = engram_node_full(content, node_type, label,
let id: String = wt_node(content, node_type, label,
sal, sal, el_from_float(0.9),
tier, tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -422,11 +431,11 @@ fn handle_api_node_update(body: String) -> String {
}
let body_tags: String = json_get(body, "tags")
let tags: String = if str_eq(body_tags, "") { "[\"" + node_type + "\"]" } else { body_tags }
let new_id: String = engram_node_full(content, node_type, label,
let new_id: String = wt_node(content, node_type, label,
el_from_float(0.5), el_from_float(0.5), el_from_float(0.8),
tier, tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
engram_connect(new_id, id, el_from_float(0.9), "supersedes")
wt_edge(new_id, id, el_from_float(0.9), "supersedes")
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + id + "\",\"ok\":true}"
}
@@ -498,7 +507,7 @@ fn handle_api_capture_knowledge(body: String) -> String {
let full: String = if str_eq(title, "") { content } else { title + ": " + content }
let lbl: String = str_slice(title, 0, 80)
let tags: String = "[\"Knowledge\",\"captured\"]"
let id: String = engram_node_full(full, "Knowledge", lbl,
let id: String = wt_node(full, "Knowledge", lbl,
el_from_float(0.85), el_from_float(0.8), el_from_float(0.9),
"Episodic", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -513,12 +522,12 @@ fn handle_api_evolve_knowledge(body: String) -> String {
if !str_eq(prior_id, "") && is_protected_node(prior_id) { return api_err_protected(prior_id) }
let tags: String = "[\"Knowledge\",\"evolved\"]"
// Empty label engram_node_full derives content[:60] (LABEL FIX 2026-07-23).
let new_id: String = engram_node_full(content, "Knowledge", "",
let new_id: String = wt_node(content, "Knowledge", "",
el_from_float(0.75), el_from_float(0.75), el_from_float(0.9),
"Episodic", tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
if !str_eq(prior_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true}"
}
@@ -535,11 +544,11 @@ fn handle_api_promote_knowledge(body: String) -> String {
"[\"Knowledge\",\"tier:canonical\",\"disposition:stable\"]"
} else { tags_raw }
// Empty label engram_node_full derives content[:60] (LABEL FIX 2026-07-23).
let new_id: String = engram_node_full(content, "Knowledge", "",
let new_id: String = wt_node(content, "Knowledge", "",
el_from_float(0.9), el_from_float(0.9), el_from_float(1.0),
"Canonical", tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
engram_connect(new_id, prior_id, el_from_float(0.95), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.95), "supersedes")
return "{\"ok\":true,\"new_id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\"}"
}
@@ -562,7 +571,7 @@ fn handle_api_define_process(body: String) -> String {
if str_eq(content, "") { return api_err("content is required") }
let label: String = if str_eq(name, "") { "process:unnamed" } else { "process:" + name }
let tags: String = "[\"Process\"]"
let id: String = engram_node_full(content, "Process", label,
let id: String = wt_node(content, "Process", label,
el_from_float(0.8), el_from_float(0.8), el_from_float(0.9),
"Canonical", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -647,7 +656,7 @@ fn handle_api_tune_config(body: String) -> String {
if str_eq(key, "") { return api_err("key is required") }
let content: String = "config:" + key + "=" + value
let tags: String = "[\"ConfigEntry\",\"config\"]"
let id: String = engram_node_full(content, "ConfigEntry", key,
let id: String = wt_node(content, "ConfigEntry", key,
el_from_float(0.85), el_from_float(0.85), el_from_float(0.9),
"Canonical", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -694,7 +703,7 @@ fn handle_api_link_entities(body: String) -> String {
if is_protected_node(to_id) { return api_err_protected(to_id) }
let relation: String = json_get(body, "relation")
let eff_relation: String = if str_eq(relation, "") { "associates" } else { relation }
engram_connect(from_id, to_id, el_from_float(0.5), eff_relation)
wt_edge(from_id, to_id, el_from_float(0.5), eff_relation)
return "{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + eff_relation + "\"}"
}
@@ -727,11 +736,11 @@ fn handle_api_evolve_memory(body: String) -> String {
}
}
let tags: String = "[\"Memory\",\"evolved\"]"
let new_id: String = engram_node_full(content, "Memory", "memory:evolved",
let new_id: String = wt_node(content, "Memory", "memory:evolved",
sal, sal, el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true}"
}
@@ -789,11 +798,11 @@ fn handle_api_cultivate(body: String) -> String {
let content: String = json_get(body, "content")
if str_eq(content, "") { return api_err("content is required") }
let tags: String = "[\"Knowledge\",\"evolved\",\"cultivated\"]"
let new_id: String = engram_node_full(content, "Knowledge", "knowledge:cultivated",
let new_id: String = wt_node(content, "Knowledge", "knowledge:cultivated",
el_from_float(0.75), el_from_float(0.75), el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true,\"cultivated\":true}"
}
@@ -809,11 +818,11 @@ fn handle_api_cultivate(body: String) -> String {
}
}
let tags: String = "[\"Memory\",\"evolved\",\"cultivated\"]"
let new_id: String = engram_node_full(content, "Memory", "memory:cultivated",
let new_id: String = wt_node(content, "Memory", "memory:cultivated",
sal, sal, el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true,\"cultivated\":true}"
}
@@ -833,7 +842,7 @@ fn handle_api_cultivate(body: String) -> String {
if str_eq(to_id, "") { return api_err("to_id is required") }
let relation: String = json_get(body, "relation")
let eff_relation: String = if str_eq(relation, "") { "associates" } else { relation }
engram_connect(from_id, to_id, el_from_float(0.5), eff_relation)
wt_edge(from_id, to_id, el_from_float(0.5), eff_relation)
return "{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + eff_relation + "\",\"cultivated\":true}"
}
@@ -868,7 +877,7 @@ fn handle_api_consolidate(body: String) -> String {
if !str_eq(summary, "") {
let safe_summary: String = str_replace(summary, "\"", "'")
let tags: String = "[\"SessionSummary\",\"consolidate\"]"
let summary_id: String = engram_node_full(
let summary_id: String = wt_node(
"[session-summary] " + safe_summary,
"SessionSummary", "session:summary",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
@@ -880,3 +889,462 @@ fn handle_api_consolidate(body: String) -> String {
}
return "{\"ok\":true,\"snapshot\":\"" + snap + "\"}"
}
// Stage 1: structural audit
//
// WHAT THIS IMPLEMENTS
// The CGI provisional, 05-detailed-description.md, "Stage 1: Structural audit
// 430". Verbatim, the audit module evaluates: the density and typed
// distribution of causal edges; the consistency between value nodes and
// execution-record neighborhoods; the richness and connectivity of the
// self-model; and the authenticity of open-question nodes in the wonder
// manifest. It "produces a coherence assessment 432 — NOT A BINARY SCORE but
// an annotated characterization of the graph's structural properties".
//
// That last clause is the whole shape of this handler. Every finding carries
// its own numbers AND a plain-language note saying what the numbers mean and
// how they were obtained. There is no pass/fail, no percentage-of-health, no
// composite score, and `"score":null` is emitted explicitly so a downstream
// reader cannot mistake its absence for an omission.
//
// WHY IT EXISTS NOW, AND WHY THE FIRST FINDING IS THE ONE IT IS
// `runStructuralAudit` has been an advertised MCP tool with nothing behind it:
// the dispatcher GET'd /session/begin and returned that blob (mcp-wrapper/src/
// main.el). Meanwhile the failure the audit would have caught ran silently for
// about three weeks the soul reported 103,089 nodes while the engram, which
// OWNS persistence, held ~79,900; a crash discarded the difference. Every boot
// reported green throughout, because nothing in the system ever compared the
// two sides. So finding 1 is owner-versus-runtime divergence: it is the check
// whose absence cost real memory, and it is cheap and exact.
//
// WHAT IS DELIBERATELY NOT HERE (stage 1b, see the `deferred` array in the
// response): value/execution-record consistency and wonder-manifest
// authenticity. Both need node types that barely exist in this graph today
// the response MEASURES those populations and reports the counts as the reason,
// rather than asserting a deferral without evidence.
//
// MEASUREMENT HONESTY: EXACT WHERE CHEAP, SAMPLED WHERE NOT, ALWAYS LABELLED
// Counts, edge typing and self-model connectivity are exact. Orphan rate and
// dangling-edge rate are SAMPLED, because the engram runtime has no node-id
// index `engram_find_node_index` is a linear scan over every node, so an
// exhaustive dangling check is O(nodes x edges) (~2.2e9 string compares at
// today's scale, tens of seconds inside one request). The samples are UNIFORM
// across the whole population, not head-of-list, and every sampled figure is
// emitted with its own `sampled` / `population` fields plus an extrapolation
// labelled as such. Raise `?edge_sample=` / `?node_sample=` to the population
// size to run either check exhaustively and pay the time. The real fix is an
// id index in the runtime; that is the engram repo's, not this handler's.
// audit_pct1 one-decimal percentage as a bare JSON number, sign-safe.
// Integer math only: EL has no fixed-precision formatter, and float_to_str
// would put an unbounded mantissa in the response.
fn audit_pct1(num: Int, den: Int) -> String {
if den <= 0 { return "null" }
let neg: Bool = num < 0
let a: Int = if neg { 0 - num } else { num }
let tenths: Int = (a * 1000) / den
let whole: Int = tenths / 10
let frac: Int = tenths - (whole * 10)
let sign: String = if neg { "-" } else { "" }
return sign + int_to_str(whole) + "." + int_to_str(frac)
}
// audit_finding the one envelope every finding uses: name, the measurements,
// and the annotation. Keeping it in one place is what stops the characterization
// from degenerating into a bag of numbers with no reading attached.
fn audit_finding(name: String, measured: String, note: String) -> String {
return "{\"finding\":\"" + name + "\""
+ ",\"measured\":{" + measured + "}"
+ ",\"note\":\"" + api_json_escape(note) + "\"}"
}
// audit_str_at read the quoted string value starting at byte `start`.
// Slices a bounded window rather than the tail of the (multi-MB) edges array, so
// this is O(window) per call instead of O(remaining input).
fn audit_str_at(s: String, start: Int, maxlen: Int) -> String {
let n: Int = str_len(s)
if start < 0 || start >= n { return "" }
let end_guess: Int = start + maxlen
let stop: Int = if end_guess > n { n } else { end_guess }
let win: String = str_slice(s, start, stop)
let q: Int = str_index_of(win, "\"")
if q < 0 { return "" }
return str_slice(win, 0, q)
}
// audit_rel_count exact count of edges carrying `rel`, by scanning the emitted
// edge array for the literal `"relation":"<rel>"`. engram_emit_edge_json writes
// metadata ESCAPED as a string, so no nested object can contain that literal and
// the count cannot be inflated by edge payloads.
fn audit_rel_count(edges: String, rel: String) -> Int {
return str_count(edges, "\"relation\":\"" + rel + "\"")
}
// audit_owner_stats ask the persistence OWNER for its own counts.
// Returns "" when there is no HTTP owner configured or the owner is unreachable;
// both are reported as findings, never as a failure of the audit.
fn audit_owner_stats(url: String) -> String {
if str_eq(url, "") { return "" }
return http_get(url + "/api/stats")
}
// audit_divergence FINDING 1. Runtime (this soul's in-process graph) versus
// the persistence owner's own count. Trend is measured against the previous
// audit recorded in soul state, so a second call answers "is the gap growing?"
// rather than just restating it.
fn audit_divergence() -> String {
let rt_nodes: Int = engram_node_count()
let rt_edges: Int = engram_edge_count()
let url: String = wt_engram_url()
if str_eq(url, "") {
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"none\",\"owner_reachable\":false",
"No HTTP persistence owner is configured, so this soul IS the owner "
+ "(file mode) and divergence is not defined. This check only has "
+ "meaning when ENGRAM_URL points at a separate engram that owns the "
+ "canonical store.")
}
let stats: String = audit_owner_stats(url)
// REACHABILITY IS PROVED BY THE PAYLOAD, NOT BY A NON-EMPTY REPLY.
// http_get does not return "" on a connection failure it returns a JSON
// error object ({"error":"Failed to connect to ... Couldn't connect to
// server"}). Testing only for "" made a DEAD owner read as reachable with
// node_count 0, i.e. the audit would have reported a 100% divergence and
// named it as data loss. That false positive is worse than no check at all:
// it is precisely the kind of confident wrong answer this route exists to
// stop. Require the field the contract promises.
let owner_nc_raw: String = json_get_raw(stats, "node_count")
if str_eq(stats, "") || str_eq(owner_nc_raw, "") {
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"" + api_json_escape(url) + "\",\"owner_reachable\":false"
+ ",\"owner_reply\":\"" + api_json_escape(api_utf8_trunc(stats, 200)) + "\"",
"The persistence owner at " + url + " did not return a node_count "
+ "from GET /api/stats. Divergence is UNKNOWN, NOT ZERO — an owner "
+ "that cannot be read is exactly the condition under which the "
+ "runtime's own count means least, and reporting 0 for the owner "
+ "would manufacture a total-loss reading out of a network error. "
+ "Reported as a finding rather than raised as an error so the rest "
+ "of the audit still returns; the owner's raw reply is in "
+ "owner_reply.")
}
let ow_nodes: Int = json_get_int(stats, "node_count")
let ow_edges: Int = json_get_int(stats, "edge_count")
let d_nodes: Int = rt_nodes - ow_nodes
let d_edges: Int = rt_edges - ow_edges
// Trend against the previous audit in this soul's state.
let prev_raw: String = state_get("audit_prev_node_delta")
let prev: Int = str_to_int(prev_raw)
let abs_now: Int = if d_nodes < 0 { 0 - d_nodes } else { d_nodes }
let abs_prev: Int = if prev < 0 { 0 - prev } else { prev }
let trend: String = if str_eq(prev_raw, "") {
"no_prior_audit"
} else {
if abs_now > abs_prev { "growing" } else {
if abs_now < abs_prev { "shrinking" } else { "flat" }
}
}
state_set("audit_prev_node_delta", int_to_str(d_nodes))
state_set("audit_prev_ts", int_to_str(time_now()))
let note_head: String = if d_nodes == 0 {
"Runtime and owner agree on node count."
} else {
"Runtime holds " + int_to_str(d_nodes) + " nodes (" + audit_pct1(d_nodes, rt_nodes)
+ "% of its own graph) that the persistence owner does not report. Nodes "
+ "that exist only in runtime memory do not survive a restart."
}
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"" + api_json_escape(url) + "\",\"owner_reachable\":true"
+ ",\"owner_nodes\":" + int_to_str(ow_nodes)
+ ",\"owner_edges\":" + int_to_str(ow_edges)
+ ",\"node_delta\":" + int_to_str(d_nodes)
+ ",\"edge_delta\":" + int_to_str(d_edges)
+ ",\"node_delta_pct_of_runtime\":" + audit_pct1(d_nodes, rt_nodes)
+ ",\"trend_vs_previous_audit\":\"" + trend + "\""
+ ",\"previous_node_delta\":" + (if str_eq(prev_raw, "") { "null" } else { int_to_str(prev) }),
note_head + " Trend against the previous audit recorded in this soul's "
+ "state: " + trend + ". This is the comparison whose absence let a "
+ "~24,000-node loss run for weeks with every boot reporting green.")
}
// audit_edge_typing FINDING 2. Density plus the typed distribution the patent
// asks for, against the claim-10 relation vocabulary. Exact: str_count over the
// emitted edge array, one linear pass per relation.
fn audit_edge_typing(edges: String, total_edges: Int, node_total: Int) -> String {
let c_sup: Int = audit_rel_count(edges, "Supersedes")
let c_cau: Int = audit_rel_count(edges, "Causes")
let c_con: Int = audit_rel_count(edges, "Contains")
let c_ref: Int = audit_rel_count(edges, "References")
let c_ctr: Int = audit_rel_count(edges, "Contradicts")
let c_exe: Int = audit_rel_count(edges, "Exemplifies")
let c_act: Int = audit_rel_count(edges, "Activates")
let c_tmp: Int = audit_rel_count(edges, "TemporallyPrecedes")
let typed: Int = c_sup + c_cau + c_con + c_ref + c_ctr + c_exe + c_act + c_tmp
// Lowercase near-misses: the same eight concepts written by the ad-hoc write
// paths (linkEntities defaults to "associates", linkCausal to "causes").
// Counted separately because "the vocabulary is unused" and "the vocabulary
// is used in the wrong case" are different defects with different fixes.
let l_sup: Int = audit_rel_count(edges, "supersedes")
let l_cau: Int = audit_rel_count(edges, "causes")
let l_con: Int = audit_rel_count(edges, "contains")
let l_ref: Int = audit_rel_count(edges, "references")
let l_ctr: Int = audit_rel_count(edges, "contradicts")
let l_exe: Int = audit_rel_count(edges, "exemplifies")
let l_act: Int = audit_rel_count(edges, "activates")
let l_tmp: Int = audit_rel_count(edges, "temporallyPrecedes")
let near: Int = l_sup + l_cau + l_con + l_ref + l_ctr + l_exe + l_act + l_tmp
let untyped: Int = total_edges - typed
return audit_finding("typed_edge_distribution",
"\"total_edges\":" + int_to_str(total_edges)
+ ",\"total_nodes\":" + int_to_str(node_total)
// Density per 100 nodes, not per node: EL has no fixed-precision float
// formatter, and "0.3 edges per node" rounded to an integer is a lie.
+ ",\"edges_per_100_nodes\":" + audit_pct1(total_edges, node_total)
+ ",\"claim10_typed\":" + int_to_str(typed)
+ ",\"claim10_typed_pct\":" + audit_pct1(typed, total_edges)
+ ",\"outside_claim10_vocabulary\":" + int_to_str(untyped)
+ ",\"lowercase_near_miss\":" + int_to_str(near)
+ ",\"by_relation\":{"
+ "\"Supersedes\":" + int_to_str(c_sup)
+ ",\"Causes\":" + int_to_str(c_cau)
+ ",\"Contains\":" + int_to_str(c_con)
+ ",\"References\":" + int_to_str(c_ref)
+ ",\"Contradicts\":" + int_to_str(c_ctr)
+ ",\"Exemplifies\":" + int_to_str(c_exe)
+ ",\"Activates\":" + int_to_str(c_act)
+ ",\"TemporallyPrecedes\":" + int_to_str(c_tmp) + "}",
"Only " + int_to_str(typed) + " of " + int_to_str(total_edges)
+ " edges use the claim-10 causal vocabulary; the remainder are ad-hoc "
+ "relation strings, which is why the graph's causal claims cannot yet "
+ "be checked for internal consistency — an untyped edge asserts "
+ "association, not causation. " + int_to_str(near) + " edges use a "
+ "lowercase spelling of a claim-10 relation: those are near-misses the "
+ "write paths could be corrected to emit, not genuinely foreign types.")
}
// audit_orphans_dangling FINDING 3. Both figures are SAMPLED; see the header
// for why exhaustive is O(nodes x edges) on this runtime.
//
// An "orphan" here is a node with zero RESOLVABLE edges: engram_neighbors_json
// drops any edge whose other endpoint does not resolve to a node, so a node
// whose only edges are dangling reads as an orphan. That is the right reading
// such a node is unreachable by traversal but it is stated rather than hidden.
fn audit_orphans_dangling(edges: String, total_edges: Int, node_total: Int,
edge_cap: Int, node_cap: Int) -> String {
// orphan sample: uniform stride over the node store
let n_take: Int = if node_total < node_cap { node_total } else { node_cap }
let n_stride: Int = if n_take > 0 { node_total / n_take } else { 1 }
let n_stride = if n_stride < 1 { 1 } else { n_stride }
let orphans: Int = 0
let n_checked: Int = 0
let j: Int = 0
while j < n_take {
let one: String = engram_scan_nodes_json(1, j * n_stride)
let nid: String = json_get(json_array_get(one, 0), "id")
if !str_eq(nid, "") {
let nbrs: String = engram_neighbors_json(nid, 1, "both")
let deg: Int = json_array_len(nbrs)
let orphans = if deg == 0 { orphans + 1 } else { orphans }
let n_checked = n_checked + 1
}
let j = j + 1
}
// dangling sample: uniform stride over the edge array
// str_index_of_all gives every edge's field offsets in ONE linear pass, so
// any index can be read in O(1). json_array_get would have been O(i) per
// element and O(n^2) over the array.
let from_pos: [Int] = str_index_of_all(edges, "\"from_id\":\"")
let to_pos: [Int] = str_index_of_all(edges, "\"to_id\":\"")
let nf: Int = len(from_pos)
let nt: Int = len(to_pos)
let ne: Int = if nf < nt { nf } else { nt }
let e_take: Int = if ne < edge_cap { ne } else { edge_cap }
let e_stride: Int = if e_take > 0 { ne / e_take } else { 1 }
let e_stride = if e_stride < 1 { 1 } else { e_stride }
let dangling: Int = 0
let e_checked: Int = 0
let i: Int = 0
while i < ne && e_checked < e_take {
let fid: String = audit_str_at(edges, get(from_pos, i) + 11, 96)
let tid: String = audit_str_at(edges, get(to_pos, i) + 9, 96)
let f_gone: Bool = str_eq(engram_get_node_json(fid), "{}")
let t_gone: Bool = if f_gone { true } else { str_eq(engram_get_node_json(tid), "{}") }
let dangling = if f_gone || t_gone { dangling + 1 } else { dangling }
let e_checked = e_checked + 1
let i = i + e_stride
}
let orphan_est: Int = if n_checked > 0 { (orphans * node_total) / n_checked } else { 0 }
let dangle_est: Int = if e_checked > 0 { (dangling * total_edges) / e_checked } else { 0 }
let exhaustive_n: String = if n_checked >= node_total { "true" } else { "false" }
let exhaustive_e: String = if e_checked >= ne { "true" } else { "false" }
return audit_finding("orphans_and_dangling_edges",
"\"nodes_population\":" + int_to_str(node_total)
+ ",\"nodes_sampled\":" + int_to_str(n_checked)
+ ",\"nodes_sample_exhaustive\":" + exhaustive_n
+ ",\"orphans_in_sample\":" + int_to_str(orphans)
+ ",\"orphan_rate_pct\":" + audit_pct1(orphans, n_checked)
+ ",\"orphans_extrapolated\":" + int_to_str(orphan_est)
+ ",\"edges_population\":" + int_to_str(total_edges)
+ ",\"edges_sampled\":" + int_to_str(e_checked)
+ ",\"edges_sample_exhaustive\":" + exhaustive_e
+ ",\"dangling_in_sample\":" + int_to_str(dangling)
+ ",\"dangling_rate_pct\":" + audit_pct1(dangling, e_checked)
+ ",\"dangling_extrapolated\":" + int_to_str(dangle_est),
"Orphan = zero RESOLVABLE edges, so a node whose only edges dangle counts "
+ "as an orphan; either way it is unreachable by traversal. Dangling = an "
+ "edge with an endpoint id that resolves to no node. Both are uniform "
+ "stride samples over the whole population, not the head of the list; "
+ "the extrapolations are estimates and are labelled as such. Pass "
+ "?node_sample= / ?edge_sample= at or above the population size to run "
+ "either check exhaustively. A high orphan rate is a characterization, "
+ "not a verdict: an accumulating store legitimately holds unlinked "
+ "material. It becomes a defect when the write paths were SUPPOSED to "
+ "link and did not.")
}
// audit_pillar one self-model pillar: present, how much content, how connected.
fn audit_pillar(key: String, id: String) -> String {
let node: String = engram_get_node_json(id)
let present: Bool = !str_eq(node, "{}") && !str_eq(node, "")
if !present {
return "\"" + key + "\":{\"id\":\"" + id + "\",\"present\":false"
+ ",\"content_length\":0,\"degree\":0}"
}
let content: String = json_get(node, "content")
let deg: Int = json_array_len(engram_neighbors_json(id, 1, "both"))
return "\"" + key + "\":{\"id\":\"" + id + "\",\"present\":true"
+ ",\"label\":\"" + api_json_escape(json_get(node, "label")) + "\""
+ ",\"tier\":\"" + api_json_escape(json_get(node, "tier")) + "\""
+ ",\"content_length\":" + int_to_str(str_len(content))
+ ",\"degree\":" + int_to_str(deg) + "}"
}
// audit_self_model FINDING 4. "the richness and connectivity of the
// self-model ... is it connected to behavioral evidence?"
//
// This finding RETIRES the Claude-side vitals identity block. That check lived
// outside the system it was checking a shell script grepping a snapshot so
// it could only ever report on a file, and it went on reporting green while the
// memory-philosophy pillar was absent from the live graph for about three weeks.
// Asking the running soul about its own three pillars is the designed mechanism;
// a shell probe was the fourth patch on the same hole.
fn audit_self_model() -> String {
let dna: String = audit_pillar("intellectual_dna", "kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6")
let val: String = audit_pillar("values_hub", "kn-5b606390-a52d-4ca2-8e0e-eba141d13440")
let phi: String = audit_pillar("memory_philosophy", "kn-dcfe04b3-3702-4cac-b6f0-ecb4db837eee")
let root: String = audit_pillar("self_root", "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee")
return audit_finding("self_model_connectivity",
"\"pillars\":{" + dna + "," + val + "," + phi + "," + root + "}",
"The three identity pillars plus the self root. `degree` counts nodes "
+ "reachable in one hop in either direction — the self-model's connection "
+ "to the rest of the graph. present:false on any pillar is the condition "
+ "that ran undetected for weeks; content_length distinguishes a pillar "
+ "that is present from one that is present but hollowed out. The patent "
+ "also asks whether the self-model makes ACCURATE PREDICTIONS about the "
+ "system's own behavior; that half needs Prediction nodes and is deferred "
+ "with the rest of stage 1b below.")
}
// audit_deferred what stage 1 does NOT yet evaluate, with the measured reason.
// Emitted as data, not as a comment, so a reader of the assessment sees the gap
// and its evidence rather than inferring completeness from silence.
fn audit_deferred() -> String {
let preds: Int = json_array_len(api_or_empty(engram_scan_nodes_by_type_json("Prediction", 50, 0)))
let wonders: Int = json_array_len(api_or_empty(engram_scan_nodes_by_type_json("WonderQuestion", 50, 0)))
return "[{\"deferred\":\"value_execution_record_consistency\""
+ ",\"stage\":\"1b\""
+ ",\"measured\":{\"prediction_nodes_found\":" + int_to_str(preds) + "}"
+ ",\"reason\":\"" + api_json_escape(
"The patent asks whether the execution history SUPPORTS the stated "
+ "values or shows systematic conflict. That requires execution "
+ "records tied to value nodes and predictions to score them against. "
+ "Prediction nodes found (capped at 50): " + int_to_str(preds)
+ ". Asserting value/execution coherence on that population would be "
+ "a fabricated result, which is worse than a stated gap.") + "\"}"
+ ",{\"deferred\":\"wonder_manifest_authenticity\""
+ ",\"stage\":\"1b\""
+ ",\"measured\":{\"wonder_question_nodes_found\":" + int_to_str(wonders) + "}"
+ ",\"reason\":\"" + api_json_escape(
"The patent asks whether pull weights CORRELATE WITH GENUINE "
+ "PREDICTION UNCERTAINTY or are uniform/externally assigned — a "
+ "correlation between two populations. WonderQuestion nodes readable "
+ "by type (capped at 50): " + int_to_str(wonders) + ", against "
+ int_to_str(preds) + " Prediction nodes. There is a known write/read "
+ "node-type mismatch on the wonder path; until that is fixed and both "
+ "populations exist, any correlation reported here would be noise.") + "\"}]"
}
// handle_api_structural_audit Stage 1. Returns the coherence assessment 432:
// an annotated characterization, explicitly NOT a score.
//
// COST NOTE: the edge findings need the relation labels, and the runtime exposes
// no edge-enumeration builtin. The only way to see them is the same one
// GET /api/graph/edges already uses engram_save to a SCRATCH path (never the
// owner's canonical file; see routes.el, neuron#117) and read the array back.
// On a large graph that is a multi-hundred-MB write, so this is a manual audit
// route, not something to put on a timer. Pass ?edges=0 to skip both edge
// findings and get the divergence + self-model readings cheaply.
fn handle_api_structural_audit(method: String, path: String, body: String) -> String {
let node_total: Int = engram_node_count()
let edge_total: Int = engram_edge_count()
let want_edges: Bool = !str_eq(api_query_param(path, "edges"), "0")
let edge_cap: Int = api_query_int(path, "edge_sample", 3000)
let node_cap: Int = api_query_int(path, "node_sample", 300)
let divergence: String = audit_divergence()
let self_model: String = audit_self_model()
let edge_part: String = if want_edges {
// Scratch export only. state_get("soul_snapshot_path") is deliberately
// NOT used: in HTTP-engram mode the soul is not the persistence owner and
// must never write the canonical file, not even on a read path.
let scratch_dir: String = env("TMPDIR")
let scratch_base: String = if str_eq(scratch_dir, "") { "/tmp" } else { scratch_dir }
let snap_path: String = scratch_base + "/soul-audit-export-" + state_get("soul_cgi_id") + ".json"
// engram_save returns Int (1 ok / 0 fail); str_eq on it SIGSEGVs (#150).
let saved: Int = engram_save(snap_path)
if saved == 0 {
"," + audit_finding("typed_edge_distribution", "\"available\":false",
"Could not export the graph to " + snap_path + " for edge analysis, "
+ "so edge typing and the dangling-edge sample were not run. "
+ "Reported as a gap, not as zero findings.")
} else {
// wt_read, not fs_read: fs_read leaves a thread-local length hint that
// the NEXT HTTP response would use as its Content-Length, appending
// adjacent heap bytes to the reply (see persist.el wt_read).
let snap: String = wt_read(snap_path)
let edges_raw: String = json_get_raw(snap, "edges")
let edges: String = if str_eq(edges_raw, "") { "[]" } else { edges_raw }
"," + audit_edge_typing(edges, edge_total, node_total)
+ "," + audit_orphans_dangling(edges, edge_total, node_total, edge_cap, node_cap)
}
} else {
""
}
return "{\"audit\":\"structural\",\"stage\":1"
+ ",\"spec\":\"CGI provisional 05-detailed-description.md, Stage 1: Structural audit 430\""
+ ",\"assessment\":\"coherence_assessment_432\""
+ ",\"assessment_kind\":\"annotated_characterization\""
+ ",\"score\":null"
+ ",\"score_note\":\"By design. The specification calls for an annotated characterization of the graph's structural properties, not a binary score. Read the findings.\""
+ ",\"cgi_id\":\"" + api_json_escape(state_get("soul_cgi_id")) + "\""
+ ",\"ts_ms\":" + int_to_str(time_now())
+ ",\"findings\":[" + divergence + "," + self_model + edge_part + "]"
+ ",\"deferred\":" + audit_deferred() + "}"
}
+1
View File
@@ -37,3 +37,4 @@ extern fn handle_api_memory_update(body: String) -> String
extern fn handle_api_cultivate(body: String) -> String
extern fn handle_api_list_typed(node_type: String, path: String, body: String) -> String
extern fn handle_api_consolidate(body: String) -> String
extern fn handle_api_structural_audit(method: String, path: String, body: String) -> String
+426
View File
@@ -0,0 +1,426 @@
// persist.el the soulengram WRITE-THROUGH boundary (neuron#117).
//
// WHY THIS FILE EXISTS
// soul.el:571-573 states the ownership rule: "when ENGRAM_URL is set the HTTP
// Engram owns persistence the soul must NEVER write to the local snapshot
// (not the persistence owner)." The soul obeys the NEGATIVE half. The POSITIVE
// half how a write made inside the soul actually REACHES the owner was
// never built. Sync is pull-only (awareness.el `/api/sync` -> engram_load_merge),
// so every node the soul creates lives in its process RAM and is shed on
// restart. Measured live 2026-08-07: soul node_count=102184, engram
// node_count=79197 ~23k nodes existing nowhere but RAM.
//
// SCOPE NOTE ON THE PATENT (corrects an earlier internal reading)
// Engram provisional claims 15-18 describe a delta-sync protocol "with peer
// Engram instances"; claim 17's pull-then-push sequence is PEER-ENGRAM to
// PEER-ENGRAM. The soul is NOT a peer Engram it is a CALLER of the database
// system API (cf. claim 27, "invoked explicitly by a caller of the database
// system API"). So claim 17 does not specify a soul↔engram contract and is not
// cited as authority here. This design is derived from the ownership rule
// alone: the owner owns the writes, therefore the soul must HAND writes to the
// owner and must never write the owner's file itself.
//
// THE MECHANISM, AND WHY NOT `POST /api/nodes`
// The obvious route is the one the persona/boot-counter write-backs already
// use, POST /api/nodes. It is the wrong instrument here, verified against the
// live engram binary in a sandbox:
// - it mints a NEW server-side id (engram_node), so the soul's id and the
// owner's id diverge the next /api/sync pull re-imports the node as a
// DUPLICATE, and any edge referencing the soul's id never resolves;
// - it accepts only {content, node_type, salience} and drops label, tier,
// tags, importance, confidence, metadata. A probe posted with tier
// "Canonical" came back tier "Working", importance 0.5.
// POST /api/load-merge (Will's own route, el `dc39a61`) is the right one:
// - engram_load_merge PRESERVES the id and every field;
// - it dedups nodes by id and edges by (from_id,to_id,relation), so a
// re-submitted delta is a NO-OP retry safety is free, and it is the same
// local-wins semantics the graph already uses;
// - it calls persist_canonical() THE OWNER writes its own canonical file.
// The soul never touches it. The ownership rule is honoured in its
// strongest form rather than worked around;
// - it returns real counts {ok, nodes_added, edges_added, node_count},
// so a receipt can be a MEASUREMENT instead of a fixed success shape.
//
// SPOOL-AND-DRAIN, AND WHY IT IS NOT JUST A DIRECT POST
// Measured in a sandbox against a 79k-node / 176MB graph (live scale): one
// load-merge costs ~0.38s, essentially all of it the owner's persist_canonical.
// A chat turn writes 5-7 nodes; pushing each separately would add ~2.7s per
// turn. So writes are STAGED and pushed in one coalesced batch.
// The staging buffer is the FILESYSTEM, not process state, because the soul
// serves each HTTP connection on its own pthread (el_runtime http_serve_async)
// and a shared in-process buffer would lose entries to a read-modify-write
// race silently, which is the one failure mode this file exists to end.
// One file per write, named with uuid_v4, is race-free by construction and
// buys a property a memory buffer cannot: writes that could not be pushed
// SURVIVE A SOUL CRASH and are drained on the next boot.
//
// WHAT IS DELIBERATELY NOT PUSHED
// - InternalStateEvent / heartbeat telemetry. Will's own carve-out, stated in
// engram server.el 8f8ccc9: "48h-pruned, loss-tolerant, ~2/min; snapshotting
// 28MB per heartbeat is waste."
// NOTE (ours, flagged for Will): we do NOT additionally exclude Working-tier
// nodes. That exclusion exists in `fb0bb55` to stop the boot counter leaking
// through the /api/sync PULL; it is about sync backflow, not durability.
// Applying it here would exclude mem_store which writes tier "Working" and
// mem_store is the single most important durable write path in the soul. Boot
// seeding reads the canonical file wholesale, so a pushed Working-tier node
// does survive restart. This is the one classification call this file makes
// that Will has not ruled on.
//
// WHAT THIS BOUNDARY CANNOT EXPRESS (by construction, not by omission)
// - engram_strengthen (salience/activation drift): load-merge SKIPS ids that
// already exist, so it cannot update an existing node. There is no owner-side
// update/upsert route. Not pushable through any current route; left as a
// follow-up that needs a change in the engram repo.
// - engram_forget (hard delete): load-merge is additive and has no delete verb.
// Propagating deletes would mean DELETE /api/nodes/<id>, a HARD delete at the
// owner which scripts/verify-soul-contract.sh section B explicitly fails the
// build for ("to delete is to supersede/tombstone, never hard-remove"). Local
// deletes therefore stay local; the TOMBSTONE NODE and its "tombstones" edge
// (mem_tombstone) are pushed, and that is the sanctioned representation of a
// deletion in this graph.
// Configuration
// wt_engram_url same resolution order as ise_post: env, then the state key
// stashed at boot. NO hardcoded localhost fallback: unlike telemetry, inventing
// a destination for durable data would risk pushing a user's memories at whatever
// happens to be listening on 8742. Empty means "no HTTP owner" -> file mode.
fn wt_engram_url() -> String {
let env_url: String = env("ENGRAM_URL")
if !str_eq(env_url, "") { return env_url }
return state_get("soul_engram_url")
}
fn wt_api_key() -> String {
let env_key: String = env("ENGRAM_API_KEY")
if !str_eq(env_key, "") { return env_key }
return state_get("soul_engram_api_key")
}
// wt_enabled true only in HTTP-engram mode. In file mode the soul IS the
// persistence owner and every path below is a no-op, so this whole feature is
// inert for genesis/local deployments. That is also what makes it reversible.
fn wt_enabled() -> Bool {
return !str_eq(wt_engram_url(), "")
}
// wt_spool_dir where staged deltas live. MUST be readable by the engram
// process: /api/load-merge takes a PATH and the owner opens it itself. Both
// processes are same-host by construction (dev-stack LaunchAgents; the GKE
// image starts engram and soul in one container per entrypoint.sh).
fn wt_spool_dir() -> String {
let raw: String = env("SOUL_OUTBOX_DIR")
let dir: String = if str_eq(raw, "") { env("HOME") + "/.neuron/soul-outbox" } else { raw }
fs_mkdir(dir)
return dir
}
// Helpers
// wt_esc minimal JSON string escape. Deliberately local rather than reusing
// chat.el's json_safe: persist.el is imported BY memory.el, which is imported by
// chat.el, so depending on chat.el here would be an import cycle.
fn wt_esc(s: String) -> String {
let s1: String = str_replace(s, "\\", "\\\\")
let s2: String = str_replace(s1, "\"", "\\\"")
let s3: String = str_replace(s2, "\n", "\\n")
let s4: String = str_replace(s3, "\r", "\\r")
let s5: String = str_replace(s4, "\t", "\\t")
return s5
}
// wt_durable_class Will's telemetry carve-out, by node_type. See header.
fn wt_durable_class(node_type: String) -> Bool {
if str_eq(node_type, "InternalStateEvent") { return false }
return true
}
// wt_inner strip the surrounding brackets off a JSON array so several arrays
// can be concatenated into one. Returns "" for "[]" / "" / anything too short.
fn wt_inner(arr: String) -> String {
let n: Int = str_len(arr)
if n < 3 { return "" }
if !str_starts_with(arr, "[") { return "" }
return str_slice(arr, 1, n - 1)
}
// wt_read fs_read, plus a MANDATORY reset of the runtime's binary-length hint.
//
// THIS IS NOT OPTIONAL AND MUST NOT BE "SIMPLIFIED" BACK TO A BARE fs_read.
// The pinned runtime (vendor/el-runtime/v1.0.0-20260501) keeps a thread-local
// `_tl_fs_read_len` that fs_read SETS to the file's byte count (so binary files
// can be served with a correct Content-Length) and that http_send_response
// CONSUMES as the Content-Length of the next reply. Nothing else clears it
// except json_get_raw. So any fs_read during request handling that is not
// followed by a json_get_raw makes the NEXT HTTP response advertise the FILE's
// length instead of the body's and the runtime then sends that many bytes,
// appending whatever adjacent heap memory follows the reply.
//
// Caught here, measured: a /api/neuron/memory reply that should be 86 bytes went
// out as 497, with 411 bytes of this module's own spool paths and log strings
// trailing the JSON. The drain reads spool files mid-request, so this boundary
// is exactly where the landmine gets stepped on.
//
// Upstream el fixed the class in `43636ae` ("pair fs_read length hint with its
// buffer"); that runtime is NOT the one vendored here, and re-pinning the
// runtime is deliberately out of scope for this change. Clearing the hint at
// our own boundary fixes our exposure without touching the pinned C.
// json_get_raw is used as the reset because it is the only builtin in this
// runtime that zeroes the hint, and it does so before any early return.
fn wt_clear_binlen() -> Void {
let discard: String = json_get_raw("{}", "_wt_reset")
}
fn wt_read(path: String) -> String {
let data: String = fs_read(path)
wt_clear_binlen()
return data
}
// wt_sweep best-effort removal of the zero-byte husks left by truncation.
// The runtime exposes no unlink builtin, so a drained delta is emptied rather
// than deleted; this reclaims the directory entries.
//
// `-empty` is the safety property, not an optimisation: the command is
// STRUCTURALLY INCAPABLE of removing a delta that still has content, so it can
// never destroy a pending write even if it runs concurrently with a stage.
// Only the directory path is interpolated (never a filename), and it is quoted.
// The exit code is ignored an un-swept husk costs one directory entry.
fn wt_sweep(dir: String) -> Void {
if str_eq(dir, "") { return }
if str_contains(dir, "'") { return }
exec_command("find '" + dir + "' -maxdepth 1 -name 'wt*.json' -empty -delete 2>/dev/null")
}
// Staging
// wt_stage write ONE delta file. uuid_v4 in the name makes concurrent stagers
// collision-free without any lock. Returns true if the delta is on disk.
fn wt_stage(nodes_json: String, edges_json: String) -> Bool {
let dir: String = wt_spool_dir()
if str_eq(dir, "") { return false }
let payload: String = "{\"nodes\":" + nodes_json + ",\"edges\":" + edges_json + "}"
let path: String = dir + "/wt-" + uuid_v4() + ".json"
fs_write(path, payload)
// Read-back-verify the stage itself. A stage that did not land is a write we
// would otherwise believe was queued exactly the hallucinated-save class.
if str_eq(wt_read(path), "") {
println("[persist] wt_stage: FAILED to write spool file " + path + " — delta not queued")
return false
}
return true
}
// The write boundary
// wt_node create a node locally AND queue it for the persistence owner.
// Same signature and same return contract as engram_node_full ("" on failure),
// so converting a call site is a rename and nothing else.
fn wt_node(content: String, node_type: String, label: String,
salience: Float, importance: Float, confidence: Float,
tier: String, tags: String) -> String {
let id: String = engram_node_full(content, node_type, label,
salience, importance, confidence,
tier, tags)
if str_eq(id, "") { return "" }
// engram_get_node_json emits the SAME record shape engram_save writes (minus
// the embedding vector, which the owner backfills lazily), so the read-back
// doubles as the delta payload no second serialization to drift.
let rec: String = engram_get_node_json(id)
if str_eq(rec, "") || str_eq(rec, "{}") {
println("[persist] wt_node: local write did not read back, id=" + id + " label=" + label)
return ""
}
if wt_enabled() && wt_durable_class(node_type) {
wt_stage("[" + rec + "]", "[]")
}
return id
}
// wt_edge create an edge locally AND queue it. Mirrors engram_connect.
//
// The edge id is freshly generated rather than read back: the runtime exposes no
// "id of the edge I just created" accessor, and the owner dedups edges by
// (from_id,to_id,relation), never by id so the id is not load-bearing. The
// consequence, stated plainly: the soul's copy and the owner's copy of the same
// edge carry different edge ids. Nothing in either codebase looks an edge up by
// id (neighbors traversal scans from_id/to_id), so this is cosmetic.
fn wt_edge(from_id: String, to_id: String, weight: Float, relation: String) -> Void {
engram_connect(from_id, to_id, weight, relation)
if !wt_enabled() { return }
if str_eq(from_id, "") || str_eq(to_id, "") { return }
let ts: Int = time_now()
let rec: String = "{\"id\":\"" + uuid_v4() + "\""
+ ",\"from_id\":\"" + wt_esc(from_id) + "\""
+ ",\"to_id\":\"" + wt_esc(to_id) + "\""
+ ",\"relation\":\"" + wt_esc(relation) + "\""
+ ",\"metadata\":\"{}\""
+ ",\"weight\":" + float_to_str(weight)
+ ",\"confidence\":1"
+ ",\"created_at\":" + int_to_str(ts)
+ ",\"updated_at\":" + int_to_str(ts)
+ ",\"last_fired\":0,\"inhibitory\":0,\"layer_id\":1}"
wt_stage("[]", "[" + rec + "]")
}
// The drain
// wt_drain coalesce every staged delta into ONE load-merge against the owner.
//
// Returns: nodes_added on success (>= 0), 0 when there was nothing to do, and
// -1 when the push FAILED. -1 is load-bearing: on failure the spool files are
// left untouched, so nothing is lost and the next drain retries them. A caller
// must never read a non-negative return as "my particular node is durable"
// use wt_durable(id) for that.
//
// Concurrency: several threads may drain at once. Each builds its own batch file
// (uuid-named), and overlapping batches are harmless because load-merge dedups.
// Files are truncated ONLY after a confirmed ok:true, so a lost race costs a
// redundant push, never a dropped write.
fn wt_drain() -> Int {
if !wt_enabled() { return 0 }
let dir: String = wt_spool_dir()
if str_eq(dir, "") { return 0 }
// el_list_len/el_list_get, NOT json_stringify(fs_list(...)): fs_list builds
// a native list via el_list_append, and json_stringify does not serialize
// that type it renders the raw pointer value. (Verified in isolation; the
// same latent defect is live in studio.el's /api/tools/file/list route,
// which returns e.g. {"entries":4386409744}. Noted, not fixed here.)
let listing = fs_list(dir)
let count: Int = el_list_len(listing)
if count == 0 { return 0 }
let nodes_acc: String = ""
let edges_acc: String = ""
let drained: String = ""
let found: Int = 0
let i: Int = 0
// No `continue` / `break`: elc lists them as keywords but not one line of
// the shipped soul uses either, so they are unexercised on this build path.
// Guard conditions are expressed as nested ifs instead, and every rebind is
// at the loop-body top level where `let x = ...` is assignment (the idiom
// memory.el's boot-counter loop relies on) never inside a nested block,
// where it would shadow instead.
while i < count {
let name: String = el_list_get(listing, i)
// A delta is only usable when it ends with the closing "]}" that
// wt_stage writes last. fs_write is not atomic, so a file being written
// right now can be observed half-formed; requiring the terminator means
// it is picked up whole on the next drain instead of merged as garbage.
// An empty read means "already drained and truncated" not an error.
let p: String = if str_starts_with(name, "wt-") { dir + "/" + name } else { "" }
let raw: String = if str_eq(p, "") { "" } else { wt_read(p) }
let usable: Bool = !str_eq(raw, "") && str_ends_with(raw, "]}")
let nj: String = if usable { wt_inner(json_get_raw(raw, "nodes")) } else { "" }
let ej: String = if usable { wt_inner(json_get_raw(raw, "edges")) } else { "" }
let nodes_acc = if str_eq(nj, "") { nodes_acc } else if str_eq(nodes_acc, "") { nj } else { nodes_acc + "," + nj }
let edges_acc = if str_eq(ej, "") { edges_acc } else if str_eq(edges_acc, "") { ej } else { edges_acc + "," + ej }
let drained = if !usable { drained } else if str_eq(drained, "") { p } else { drained + "\n" + p }
let found = if usable { found + 1 } else { found }
let i = i + 1
}
if found == 0 { return 0 }
let combined: String = "{\"nodes\":[" + nodes_acc + "],\"edges\":[" + edges_acc + "]}"
let batch: String = dir + "/wtb-" + uuid_v4() + ".json"
fs_write(batch, combined)
if str_eq(wt_read(batch), "") {
println("[persist] wt_drain: could not write batch file " + batch + "" + int_to_str(found) + " deltas stay queued")
return -1
}
let url: String = wt_engram_url()
let key: String = wt_api_key()
let body: String = "{\"path\":\"" + wt_esc(batch) + "\",\"_auth\":\"" + wt_esc(key) + "\"}"
let resp: String = http_post_json(url + "/api/load-merge", body)
// The batch file is pure scratch the retry is rebuilt from the SPOOL, not
// from it. Truncate it unconditionally, before branching on the outcome, so
// a persistently unreachable owner cannot accumulate one husk per attempt.
fs_write(batch, "")
// Distinguish the two failures rather than collapsing them: "cannot reach
// the owner" and "the owner refused this delta" need different human
// responses, and a log line that says the wrong one costs a debugging hour.
// curl surfaces transport errors as a JSON body, so an empty response is not
// the only unreachable signal.
// (str_contains rather than a strict parse on purpose the engram's HTTP
// responses have been observed carrying trailing bytes past the JSON.)
let unreachable: Bool = str_eq(resp, "")
|| str_contains(resp, "Couldn't connect")
|| str_contains(resp, "Failed to connect")
|| str_contains(resp, "Could not resolve")
|| str_contains(resp, "timed out")
if unreachable {
wt_sweep(dir)
println("[persist] wt_drain: owner UNREACHABLE at " + url + "" + int_to_str(found)
+ " deltas stay queued in " + dir + " (will retry): " + resp)
return -1
}
if !str_contains(resp, "\"ok\":true") {
wt_sweep(dir)
println("[persist] wt_drain: owner REJECTED the delta — " + int_to_str(found)
+ " stay queued in " + dir + ": " + resp)
return -1
}
let added: Int = json_get_int(resp, "nodes_added")
let added_e: Int = json_get_int(resp, "edges_added")
// Confirmed. Truncate the drained spool files so they are not re-pushed.
// Truncation (not deletion) because the runtime exposes no unlink builtin;
// an emptied file is inert to the loop above. The zero-byte husks are then
// swept below.
let paths = str_split(drained, "\n")
let pn: Int = el_list_len(paths)
let k: Int = 0
while k < pn {
let one: String = el_list_get(paths, k)
if !str_eq(one, "") { fs_write(one, "") }
let k = k + 1
}
wt_sweep(dir)
println("[persist] wt_drain: pushed " + int_to_str(found) + " deltas -> owner added "
+ int_to_str(added) + " nodes, " + int_to_str(added_e) + " edges")
return added
}
// wt_durable is this id present AT THE OWNER? The only honest answer to
// "did my write persist" in HTTP mode.
//
// In file mode the soul IS the owner, so the local read-back is the owner-side
// read-back and this collapses to the pre-existing check.
//
// nodes_added from wt_drain is NOT a substitute: a concurrent drain may have
// already pushed this node, making our own added count 0 while the node is
// perfectly durable. Presence at the owner is the fact; counts are telemetry.
fn wt_durable(id: String) -> Bool {
if str_eq(id, "") { return false }
if !wt_enabled() {
let local: String = engram_get_node_json(id)
return !str_eq(local, "") && !str_eq(local, "null") && !str_eq(local, "{}")
}
let url: String = wt_engram_url()
let resp: String = http_get(url + "/api/nodes/" + id)
if str_eq(resp, "") { return false }
if str_eq(resp, "{}") { return false }
return str_contains(resp, "\"id\"")
}
// wt_commit flush, then assert at the owner. The receipt callers should use.
// Deliberately NOT a fixed success shape: it can and does return false while the
// local write is perfectly fine in RAM, which is the true state of affairs when
// the owner is unreachable.
fn wt_commit(id: String) -> Bool {
if str_eq(id, "") { return false }
if !wt_enabled() {
let local: String = engram_get_node_json(id)
return !str_eq(local, "") && !str_eq(local, "null") && !str_eq(local, "{}")
}
let pushed: Int = wt_drain()
return wt_durable(id)
}
+58 -7
View File
@@ -186,7 +186,7 @@ fn route_imprint_contextual(body: String) -> String {
return "{\"ok\":false,\"error\":\"empty body\"}"
}
let tags: String = "[\"imprint\",\"contextual\"]"
let id: String = engram_node_full(
let id: String = wt_node(
body,
"Entity",
"imprint:contextual",
@@ -208,7 +208,7 @@ fn route_imprint_user(body: String) -> String {
return "{\"ok\":false,\"error\":\"empty body\"}"
}
let tags: String = "[\"imprint\",\"user\"]"
let id: String = engram_node_full(
let id: String = wt_node(
body,
"Entity",
"imprint:user",
@@ -239,7 +239,7 @@ fn route_synthesize(body: String) -> String {
}
let req: String = "synthesize " + parent_a + " " + parent_b
let tags: String = "[\"soul-inbox-pending\",\"synthesis-request\"]"
engram_node_full(
wt_node(
req,
"Entity",
"synthesis-request",
@@ -395,7 +395,28 @@ fn handle_connectors(method: String, clean: String, body: String) -> String {
return "{\"ok\":false,\"error\":\"unknown connectors route\"}"
}
// handle_request the soul's HTTP entry point.
//
// NOTE ON THE NAME (neuron#117): the el runtime resolves this handler by NAME
// via dlsym(RTLD_DEFAULT, "handle_request") that is why the Linux build must
// link -rdynamic. So the dispatcher body moved to route_dispatch and the name
// `handle_request` stays put as a thin wrapper. Do not rename it back.
//
// The wrapper exists to give the write-through boundary a guaranteed flush
// point. route_dispatch returns from ~60 places; a per-branch flush would be
// forgotten on the 61st. Draining here means EVERY request that staged a write
// pushes it before the connection closes, whatever route produced it, including
// routes added later that know nothing about persistence.
//
// wt_drain is a no-op (no HTTP, no cost) when nothing is staged and when the
// soul is not in HTTP-engram mode, so this is free on read traffic.
fn handle_request(method: String, path: String, body: String) -> String {
let resp: String = route_dispatch(method, path, body)
let flushed: Int = wt_drain()
return resp
}
fn route_dispatch(method: String, path: String, body: String) -> String {
let clean: String = strip_query(path)
// ACTIVITY STAMP (2026-07-30 self-review): every inbound HTTP request
@@ -432,10 +453,27 @@ fn handle_request(method: String, path: String, body: String) -> String {
return engram_scan_nodes_json(9999, 0)
}
if str_eq(clean, "/api/graph/edges") {
// TODO(reliability #8): engram_save races with awareness loop mem_save().
// Both now use atomic write-to-temp+rename (el_runtime.c). Serialised
// by engram_global_mu. Future: add engram_edges_json() builtin.
let snap_path: String = env("HOME") + "/.neuron/engram/snapshot.json"
// FIXED (neuron#117): this GET used to engram_save() straight over
// ~/.neuron/engram/snapshot.json a READ route, in a process that is
// NOT the persistence owner, overwriting the owner's canonical file
// on every call. It broke soul.el:571-573 ("the soul must NEVER write
// to the local snapshot") and it is the same defect class Will removed
// from the engram itself in el `dc39a61` ("stop read routes clobbering
// canonical snapshot"), where route_scan_edges/route_sync were moved
// to scratch paths for exactly this reason. It was also the race the
// old TODO(reliability #8) admitted to.
//
// Export to a scratch path instead. Same response, no canonical write.
// The soul's own snapshot writes are otherwise already gated behind
// state key "soul_snapshot_path", which is set ONLY in the genesis
// file-mode branch (soul.el: is_genesis && safe_to_seed, and
// safe_to_seed is unconditionally false when ENGRAM_URL is set) so
// after this change the soul writes nothing at all in HTTP mode.
// Future: add an engram_edges_json() builtin and drop the file round
// trip entirely.
let scratch_dir: String = env("TMPDIR")
let scratch_base: String = if str_eq(scratch_dir, "") { "/tmp" } else { scratch_dir }
let snap_path: String = scratch_base + "/soul-edges-export-" + state_get("soul_cgi_id") + ".json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
let edges_raw: String = json_get_raw(snap, "edges")
@@ -529,6 +567,13 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_starts_with(clean, "/api/neuron/graph") {
return handle_api_inspect_graph(method, path, body)
}
// Stage 1 structural audit (CGI provisional, "Structural audit 430").
// GET because it is a read of the graph's own structure; the query string
// carries the sample caps (?edge_sample=, ?node_sample=, ?edges=0), so
// str_starts_with rather than str_eq.
if str_starts_with(clean, "/api/neuron/audit/structural") {
return handle_api_structural_audit(method, path, body)
}
if str_starts_with(clean, "/api/neuron/list/") {
// Offset 17 = len("/api/neuron/list/"). Was 16, which left a leading "/" on node_type
// ("/BacklogItem"), so engram_scan_nodes_by_type_json matched nothing list/<type>
@@ -710,6 +755,12 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(clean, "/api/neuron/graph/link") {
return handle_api_link_entities(body)
}
// POST accepted too: same handler, so a JSON-RPC-shaped caller that only
// speaks POST reaches the identical audit. Options still come from the
// query string the handler reads no body fields.
if str_eq(clean, "/api/neuron/audit/structural") {
return handle_api_structural_audit(method, path, body)
}
if str_eq(clean, "/api/neuron/memory") {
return handle_api_remember(body)
}
+1 -1
View File
@@ -204,7 +204,7 @@ fn safety_log_bell(level: String, reason: String, input_summary: String) -> Stri
// Emit a fallback println so the bell event leaves at least a log trace even
// when engram is degraded. This does not replace engram persistence -- it is a
// last-resort audit trail when the primary write cannot be confirmed.
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content,
"BellEvent",
"bell:" + level,
+108
View File
@@ -0,0 +1,108 @@
#!/usr/bin/env bash
# run-el-test.sh — compile and run one El test program from tests/.
#
# WHY THIS EXISTS (2026-08-07, issue #129):
# tests/ has held 14 test programs for months with no way to run them. CI does
# not run them. The convention printed in their own headers
# (`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
# runtime does not implement. So the tests were documentation, not gates —
# which is how a P0 safety regression shipped with a test directory present.
#
# THE RECIPE, AND WHY IT IS THIS SHAPE:
# Same discovery as gen-soul-amalgam.sh — `elc --target=c` emits only an extern
# prototype for any module that has a .elh header next to it, and inlines the
# module's bodies when it does not. A test that imports ../chat.el therefore
# compiles to a 18 KB unit full of unresolved externs unless the headers are
# out of the way. So: copy the sources into a scratch tree, delete every .elh
# on the import chain, and compile the test there.
#
# Scratch copy on purpose: the worktree is shared with other terminals and
# deleting headers in place would be a shared-tree mutation with no owner.
#
# EXIT STATUS IS THE GATE: non-zero if the binary fails to build, crashes, or if
# its output contains a FAIL line or reports a non-zero failed count. Do not
# "improve" this into something that only checks the exit code of the test
# binary — these El tests print failures and still exit 0.
#
# usage: scripts/run-el-test.sh tests/test_history_amplification.el
set -euo pipefail
TEST_REL="${1:?usage: run-el-test.sh tests/<test>.el}"
SRC="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
TEST_NAME="$(basename "$TEST_REL" .el)"
ELC="${ELC:-$HOME/neuron-dev-stack/src/el/lang/dist/platform/elc}"
[ -x "$ELC" ] || ELC="$HOME/el-sdk/elc"
[ -x "$ELC" ] || { echo "[run-el-test] FAIL: no elc found (set ELC=)"; exit 1; }
RTC="${RTC:-$SRC/vendor/el-runtime/v1.0.0-20260501/el_runtime.c}"
[ -f "$RTC" ] || RTC="$HOME/el-sdk/el_runtime.c"
[ -f "$RTC" ] || { echo "[run-el-test] FAIL: no el_runtime.c found (set RTC=)"; exit 1; }
RTDIR="$(dirname "$RTC")"
EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/el}"
SSL="${SSL_PREFIX:-/opt/homebrew/opt/openssl@3}"
GEN="$(mktemp -d "${TMPDIR:-/tmp}/el-test.XXXXXX")"
trap 'rm -rf "$GEN"' EXIT
mkdir -p "$GEN/neuron/tests" "$GEN/foundation/el/elp/src"
cp "$SRC"/*.el "$GEN/neuron/"
cp "$SRC"/tests/*.el "$GEN/neuron/tests/" 2>/dev/null || true
[ -d "$EL_REPO/elp/src" ] && cp "$EL_REPO"/elp/src/*.el "$GEN/foundation/el/elp/src/" 2>/dev/null || true
# The whole recipe depends on there being no headers to short-circuit inlining.
find "$GEN" -name '*.elh' -delete
echo "[run-el-test] compiling $TEST_REL"
( cd "$GEN/neuron" && "$ELC" --target=c "tests/${TEST_NAME}.el" ) > "$GEN/${TEST_NAME}.c"
BODIES=$(grep -c '^el_val_t .*) {$' "$GEN/${TEST_NAME}.c" || true)
echo "[run-el-test] $(wc -c < "$GEN/${TEST_NAME}.c" | tr -d ' ') bytes, ${BODIES} inlined function bodies"
# A test that imports ../chat.el pulls in the bulk of the engine. A tiny body
# count means an import was read from a header instead of inlined, and the test
# would be exercising extern stubs rather than the real code.
if [ "$BODIES" -lt 100 ]; then
echo "[run-el-test] FAIL: only $BODIES inlined bodies — an import was not inlined"
exit 1
fi
cc -O2 -DHAVE_CURL \
-I"$RTDIR" -I"$SSL/include" -L"$SSL/lib" \
"$GEN/${TEST_NAME}.c" "$RTC" \
-lssl -lcrypto -lcurl -lpthread -lm \
-o "$GEN/${TEST_NAME}" 2> "$GEN/cc.log" || {
echo "[run-el-test] FAIL: compile error"; tail -30 "$GEN/cc.log"; exit 1; }
# arm64 pointer-truncation guard (cc-brain.sh's rule): an implicit declaration of
# a runtime symbol truncates its returned pointer to 32 bits.
if grep -E 'implicit.*(engram_|el_)' "$GEN/cc.log"; then
echo "[run-el-test] FAIL: implicit declarations of runtime symbols"; exit 1; fi
# Throwaway HOME so a test can never read or write the live engram at ~/.neuron.
TEST_HOME="$GEN/home"
mkdir -p "$TEST_HOME"
echo "[run-el-test] running $TEST_NAME"
set +e
HOME="$TEST_HOME" NEURON_HOME="$TEST_HOME/.neuron" "$GEN/${TEST_NAME}" 2>&1 | tee "$GEN/out.txt"
RC=${PIPESTATUS[0]}
set -e
if [ "$RC" -ne 0 ]; then
echo "[run-el-test] FAIL: $TEST_NAME exited $RC (crash or abort)"
exit 1
fi
if grep -q " FAIL:" "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME reported failing assertions"
exit 1
fi
if grep -qE '[1-9][0-9]* failed' "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME reported a non-zero failed count"
exit 1
fi
if ! grep -q "PASS:" "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME produced no assertions at all"
exit 1
fi
echo "[run-el-test] PASS: $TEST_NAME"
+7 -7
View File
@@ -87,7 +87,7 @@ fn session_create(body: String) -> String {
let folder: String = json_get(body, "folder")
let content: String = session_make_content(id, title, ts, ts, folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -358,7 +358,7 @@ fn session_update_patch(session_id: String, body: String) -> String {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, eff_title, created_int, ts, eff_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_node_id: String = engram_node_full(
let new_node_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -456,7 +456,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
// TODO(reliability #7): delete-then-insert is not atomic concurrent saves for the
// same session can produce orphan history nodes. State is primary truth; engram fallback.
let tags: String = "[\"session\",\"session-history\",\"Conversation\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
hist, "Conversation", "session:messages:" + session_id,
el_from_float(0.6), el_from_float(0.6), el_from_float(0.9),
"Episodic", tags
@@ -488,7 +488,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
+ " | ts:" + int_to_str(ts_now)
let summary_tags: String = "[\"session-emotional-summary\",\"affective\",\"bell:" + eff_level + "\",\"BellEvent\"]"
let summary_sal: String = if str_eq(eff_level, "hard") { el_from_float(0.95) } else { el_from_float(0.85) }
let sum_discard: String = engram_node_full(
let sum_discard: String = wt_node(
summary_content,
"BellEvent",
"session:emotional-summary",
@@ -529,7 +529,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
if !str_eq(ot_id, "") { engram_forget(ot_id) }
let oti = oti + 1
}
let discard_topic: String = engram_node_full(
let discard_topic: String = wt_node(
topic_content, "Conversation", topic_label,
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", topic_tags
@@ -582,7 +582,7 @@ fn session_update_meta_timestamp(session_id: String) -> Void {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, old_title, created_int, ts, old_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_id: String = engram_node_full(
let new_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -629,7 +629,7 @@ fn session_auto_title(session_id: String, first_message: String) -> Void {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, new_title, created_int, ts, old_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_id: String = engram_node_full(
let new_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
+17
View File
@@ -657,6 +657,23 @@ if is_genesis && safe_to_seed {
}
}
// CRASH RECOVERY (neuron#117). Deltas the previous process staged but could not
// hand to the owner are still on disk the spool is a filesystem queue, not a
// memory buffer, precisely so that a soul that died mid-flight does not take its
// unpushed writes with it. Drain them before serving, so recovered memories are
// durable and recallable from the owner from the first request onward.
//
// Safe on a clean boot: an empty spool means no HTTP call at all. Safe in file
// mode: wt_drain returns immediately when ENGRAM_URL is unset.
let wt_recovered: Int = wt_drain()
if wt_recovered > 0 {
println("[soul] write-through: recovered " + int_to_str(wt_recovered)
+ " nodes from a previous process's spool -> persistence owner")
}
if wt_recovered < 0 {
println("[soul] write-through: spool present but the persistence owner is unreachable — queued, will retry on heartbeat")
}
println("[soul] serving on port " + int_to_str(port))
http_serve_async(port, "handle_request")
println("[soul] awareness loop starting")
+2 -2
View File
@@ -11,7 +11,7 @@ import "memory.el"
fn steward_log_event(kind: String, detail: String) -> Void {
let content: String = "STEWARD:" + kind + " | " + detail
let tags: String = "[\"stewardship\",\"steward:" + kind + "\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
content,
"StewardshipEvent",
"steward:" + kind,
@@ -221,7 +221,7 @@ fn steward_fingerprint_session(input: String, session_id: String) -> String {
+ " formality=" + fs_str
+ " time=" + tb_str
let sample_tags: String = "[\"behavior\",\"BehaviorSample\",\"stewardship\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
sample_content,
"BehaviorSample",
"behavior:" + session_id,
+213
View File
@@ -0,0 +1,213 @@
// test_history_amplification.el
//
// REGRESSION TEST FOR ISSUE #129 (P0, SAFETY).
//
// What this guards: on the agentic path, the crisis score has two halves the
// message you just sent, and the distress that has accumulated across the
// conversation. The second half is the whole reason the escalation logic exists:
// someone whose distress builds over several turns never sends one message that
// trips the bell on its own.
//
// The defect this test was written against (ff421d3, 2026-08-05 fixed
// 2026-08-07): conversation history moved to a per-session key via
// conv_hist_key(session_id), but the agentic path's safety screen was left
// reading the old anonymous "conv_history" bucket. The desktop app always sends
// a session_id, so the screen received "" on every real conversation and the
// escalation half always scored 0. Nothing failed. Nothing logged. The comment
// above the defective line documented this same bug being fixed once before.
//
// THE INVARIANT UNDER TEST, stated so it survives future renames:
// the window the safety screen READS must be the window conv_history_record
// WRITES. Not "must be called conv_history" must AGREE.
//
// This test is deliberately written to fail loudly on the pre-fix source. If it
// ever passes on code where the screen reads a key nothing writes, it is broken.
//
// To run (macOS, from the worktree root):
// scripts/run-el-test.sh tests/test_history_amplification.el
//
import "../chat.el"
import "../safety.el"
import "../sessions.el"
// Program class. Without this an El program compiles as a 'utility', and a
// utility may not call the self-formation primitives (llm_call_system,
// llm_vision) that chat.el's agentic loop references the unit fails to
// compile with a capability violation even though the test never calls them.
// Declaring 'cgi' matches how soul.el declares itself.
//
// The endpoints below are deliberately DEAD: this test must never reach a live
// engram, and nothing it asserts depends on one. Port 9 is discard.
cgi "neuron-test-history-amplification" {
dharma_id: "ntn-test@http://127.0.0.1:9",
principal: "test-harness",
network: "dharma-testnet",
engram: "http://127.0.0.1:9"
}
// Counters
//
// NOTE for anyone copying this harness: the idiom used by the older tests in
// this directory `let pass_count = pass_count + 1` inside an assert function
// does NOT mutate the module-level binding. It declares a new local that dies
// with the call, so those suites all print "0 passed, 0 failed" no matter what
// happened. Counters go through the state store here so the summary is real.
fn bump(counter: String) -> Void {
let cur: String = state_get(counter)
let n: Int = if str_eq(cur, "") { 0 } else { str_to_int(cur) }
state_set(counter, int_to_str(n + 1))
}
fn counter_of(counter: String) -> Int {
let cur: String = state_get(counter)
if str_eq(cur, "") { return 0 }
return str_to_int(cur)
}
fn assert_eq(label: String, got: String, expected: String) -> Void {
if str_eq(got, expected) {
bump("t_pass")
println(" PASS: " + label)
} else {
bump("t_fail")
println(" FAIL: " + label)
println(" got: " + got)
println(" expected: " + expected)
}
}
fn assert_true(label: String, cond: Bool) -> Void {
if cond {
bump("t_pass")
println(" PASS: " + label)
} else {
bump("t_fail")
println(" FAIL: " + label)
}
}
// Fixtures
//
// ESCALATION_MSG scores 50 on its own ("can't go on", safety_score_crisis s9).
// 50 is >= soft (35) and < hard (70): a soft bell by itself, which is the point.
// It is the message that is only a crisis IN CONTEXT.
//
// DISTRESS_TURN carries five history phrases:
// hopeless 15 + worthless 15 + nobody cares 15 + can't take it anymore 20
// + want to disappear 20 = 85; contributed at 1/3 weight = 28.
// 50 + 28 = 78 >= 70 hard_bell. That 28-point swing IS the escalation logic,
// and it is exactly what the defect suppressed.
let TEST_SESSION: String = "sess-test-129-history-amplification"
let ESCALATION_MSG: String = "honestly i can't go on like this"
let DISTRESS_TURN: String = "i feel hopeless and worthless, nobody cares, i can't take it anymore, i want to disappear"
println("")
println("test_history_amplification.el — issue #129 (agentic multi-turn crisis escalation)")
// 1. Baseline: the message alone is a SOFT bell, not a hard one
//
// If this ever returns hard_bell, the test below proves nothing the message
// would trip the bell without any history and the amplification would be
// invisible. This assertion is what keeps the real test honest.
println("")
println("1. baseline — escalation message with NO history is a soft bell")
let baseline: String = safety_screen(ESCALATION_MSG, "")
assert_eq("no history -> soft_bell (not hard)", json_get(baseline, "action"), "soft_bell")
// 2. Producer sanity: history lands in the session's own window
println("")
println("2. producer — conv_history_record writes the session's window")
conv_history_record(TEST_SESSION, DISTRESS_TURN, "i hear you, that sounds heavy", "")
let written: String = state_get(conv_hist_key(TEST_SESSION))
assert_true("session window is non-empty after record", !str_eq(written, ""))
assert_true("session window contains the distress turn", str_contains(written, "hopeless"))
// 3. THE REGRESSION: the agentic screen must SEE that window
//
// Pre-fix this returns soft_bell, because agentic_safety_screen read the
// anonymous bucket and got "". Post-fix it returns hard_bell.
println("")
println("3. REGRESSION #129 — agentic screen reads the session's own window")
let screened: String = agentic_safety_screen(TEST_SESSION, ESCALATION_MSG)
assert_eq(
"distress history escalates the agentic screen to hard_bell",
json_get(screened, "action"),
"hard_bell"
)
// 4. The invariant, stated directly
//
// Independent of thresholds and phrase lists: whatever the screen reads for a
// session must equal what the recorder wrote for that session. This is the
// assertion that survives a future rename of either side.
println("")
println("4. invariant — read window == written window")
let read_back: String = state_get(conv_hist_key(TEST_SESSION))
assert_true("screen input is the recorded window, not empty", !str_eq(read_back, ""))
assert_eq("read window is byte-identical to written window", read_back, written)
// 5. No false positive: a calm session does not escalate
//
// A test that only ever asserts "hard_bell" would pass on code that hard-bells
// every message. This is the other leg, and it runs BEFORE the anonymous case
// below on purpose: that case writes the shared bucket, and under the defect a
// calm session would then inherit it.
println("")
println("5. specificity — a calm history does NOT escalate")
let CALM_SESSION: String = "sess-test-129-calm"
state_set("conv_history", "")
conv_history_record(CALM_SESSION, "what is the weather like today", "clear and mild", "")
let calm: String = agentic_safety_screen(CALM_SESSION, ESCALATION_MSG)
assert_eq("calm history stays at soft_bell", json_get(calm, "action"), "soft_bell")
// 6. Cross-session leakage
//
// The same defect had a second face: because the screen read one shared bucket,
// a calm session could be scored against a DIFFERENT session's distress. That is
// wrong in both directions it fabricates a crisis for the calm user and it
// leaks the distressed user's content into another session's scoring.
println("")
println("6. isolation — one session's distress must not score another session")
state_set("conv_history", "")
let OTHER_SESSION: String = "sess-test-129-other"
conv_history_record(OTHER_SESSION, DISTRESS_TURN, "i hear you", "")
let isolated: String = agentic_safety_screen(CALM_SESSION, ESCALATION_MSG)
assert_eq(
"a distressed OTHER session does not escalate the calm session",
json_get(isolated, "action"),
"soft_bell"
)
// 7. Anonymous sessions still work
//
// conv_hist_key("") deliberately falls back to the shared "conv_history" bucket.
// The fix must not break the no-session_id path older callers rely on. Runs last
// because it writes that shared bucket.
println("")
println("7. anonymous path — empty session_id still screens against the shared window")
state_set("conv_history", "[{\"role\":\"user\",\"content\":\"" + DISTRESS_TURN + "\"}]")
let anon: String = agentic_safety_screen("", ESCALATION_MSG)
assert_eq("anonymous session escalates too", json_get(anon, "action"), "hard_bell")
// Summary
println("")
println("history amplification tests: " + int_to_str(counter_of("t_pass")) + " passed, " + int_to_str(counter_of("t_fail")) + " failed")
+97 -11
View File
@@ -41,6 +41,7 @@
#include <fcntl.h>
#include <dirent.h>
#include <errno.h>
#include <signal.h> /* SIGPIPE disposition — see el_runtime_ignore_sigpipe */
#include <pthread.h>
#include <curl/curl.h>
@@ -1238,16 +1239,77 @@ static const char* http_reason_phrase(int status) {
}
}
/* Best-effort send with retry on partial writes. */
/* ── A departing client MUST NOT be able to kill the daemon ──────────────────
* (2026-08-06, round 9.1 / ADR 0006 item 4.)
*
* Measured field failure: a client cancelled its request at 25 s; the handler
* finished its work at 116.9 s and wrote the reply into the departed client's
* socket. The second send() on a reset connection raised SIGPIPE, whose DEFAULT
* disposition terminates the process `exited due to SIGPIPE ... ran for
* 361177ms`. launchd respawned 4 ms later, so EVERY other in-flight request on
* that daemon lost its work, silently.
*
* Two independent guards, because one of them can be undone from outside this
* file (an embedder may reset signal dispositions) and the other cannot:
* 1. process-wide SIGPIPE -> SIG_IGN, installed at runtime init;
* 2. per-send suppression at the syscall (MSG_NOSIGNAL where the platform has
* it, SO_NOSIGPIPE on the accepted socket on macOS/BSD).
* With either in force, send() reports the peer's departure as EPIPE and the
* caller decides which is the point: this is an ordinary I/O outcome, not a
* fatal condition.
*
* It deliberately does NOT swallow the error. http_send_response() below
* classifies the errno and logs: "client left" for a departure, and a real
* "send failed: <strerror>" for anything else, so a genuine write fault is
* still visible in the log (spec round-9.1 §5.3). */
#ifndef MSG_NOSIGNAL
#define MSG_NOSIGNAL 0
#endif
void el_runtime_ignore_sigpipe(void) {
static int done = 0;
if (done) return;
done = 1;
struct sigaction sa;
memset(&sa, 0, sizeof(sa));
sa.sa_handler = SIG_IGN;
sigemptyset(&sa.sa_mask);
sigaction(SIGPIPE, &sa, NULL);
}
/* Suppress SIGPIPE for one accepted connection (macOS/BSD have no
* MSG_NOSIGNAL; they have the socket option instead). Best effort. */
static void http_socket_nosigpipe(int fd) {
#ifdef SO_NOSIGPIPE
int on = 1;
setsockopt(fd, SOL_SOCKET, SO_NOSIGPIPE, &on, sizeof(on));
#else
(void)fd;
#endif
}
/* Best-effort send with retry on partial writes.
* Returns 0 on success, -1 on failure with errno preserved for the caller. */
static int http_send_all(int fd, const char* p, size_t left) {
while (left > 0) {
ssize_t w = send(fd, p, left, 0);
if (w <= 0) return -1;
ssize_t w = send(fd, p, left, MSG_NOSIGNAL);
if (w < 0) {
if (errno == EINTR) continue; /* not an error — retry */
return -1; /* errno stays set for caller */
}
if (w == 0) { errno = EPIPE; return -1; }
p += w; left -= (size_t)w;
}
return 0;
}
/* Did this write fail because the client is gone, or because something is
* actually wrong with the socket? Only the first is routine. */
static int http_write_err_is_client_gone(int e) {
return e == EPIPE || e == ECONNRESET || e == ENOTCONN || e == ESHUTDOWN;
}
/* Discriminator that http_response() embeds at the start of its envelope.
* A handler returning a string starting with this exact prefix is treated
* as a structured response; anything else is treated as a raw body. */
@@ -1468,14 +1530,30 @@ static void http_send_response(int fd, const char* body) {
free(env_body); free(hdrs.buf); return;
}
if (http_send_all(fd, status_line, (size_t)sl) == 0
&& http_send_all(fd, hdrs.buf, hdrs.len) == 0
&& http_send_all(fd, tail, (size_t)tl) == 0
&& (head_only
/* HEAD requests echo headers + Content-Length but no body. */
? 1
: http_send_all(fd, eff_body, blen) == 0)) {
/* sent successfully */
/* The reply is written in four pieces; any of them can find the client
* already gone. errno is captured at the first failure, before any later
* library call can clobber it, and classified once below. */
errno = 0;
int send_err = 0;
if (http_send_all(fd, status_line, (size_t)sl) != 0) send_err = errno;
else if (http_send_all(fd, hdrs.buf, hdrs.len) != 0) send_err = errno;
else if (http_send_all(fd, tail, (size_t)tl) != 0) send_err = errno;
else if (!head_only /* HEAD echoes headers + Content-Length, no body. */
&& http_send_all(fd, eff_body, blen) != 0) send_err = errno;
if (send_err) {
if (http_write_err_is_client_gone(send_err)) {
/* ROUTINE. The user closed the window, quit the app, or cancelled.
* The work is done and the daemon keeps serving everyone else. */
fprintf(stderr, "[http] client left before the reply was written "
"(%zu-byte body, %s) - request completed, reply discarded\n",
blen, strerror(send_err));
} else {
/* NOT routine — a real write fault. Never let the client-gone case
* above hide this one. */
fprintf(stderr, "[http] send failed: %s (%zu-byte body)\n",
strerror(send_err), blen);
}
}
if (env_parsed_root) el_release(env_parsed_root);
@@ -1491,6 +1569,7 @@ static void* http_worker(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg;
int fd = a->fd;
free(a);
http_socket_nosigpipe(fd);
char *method = NULL, *path = NULL, *body = NULL;
if (http_read_request(fd, &method, &path, &body, NULL) == 0) {
http_handler_fn h = http_lookup_active();
@@ -1531,6 +1610,7 @@ static void* http_worker(void* arg) {
}
void http_serve(el_val_t port, el_val_t handler) {
el_runtime_ignore_sigpipe(); /* serving implies clients that leave */
/* If `handler` looks like a string name, register it as the active handler. */
const char* hname = EL_CSTR(handler);
if (hname && looks_like_string(handler)) {
@@ -1634,6 +1714,7 @@ static void* _http_serve_async_loop(void* raw) {
}
void http_serve_async(el_val_t port, el_val_t handler) {
el_runtime_ignore_sigpipe(); /* serving implies clients that leave */
const char* hname = EL_CSTR(handler);
if (hname && looks_like_string(handler)) {
http_set_handler(handler);
@@ -1821,6 +1902,7 @@ static void* http_worker_v2(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg;
int fd = a->fd;
free(a);
http_socket_nosigpipe(fd);
char *method = NULL, *path = NULL, *body = NULL, *hdr_block = NULL;
if (http_read_request(fd, &method, &path, &body, &hdr_block) == 0) {
http_handler4_fn h = http_lookup_active_v2();
@@ -1858,6 +1940,7 @@ static void* http_worker_v2(void* arg) {
}
void http_serve_v2(el_val_t port, el_val_t handler) {
el_runtime_ignore_sigpipe(); /* serving implies clients that leave */
const char* hname = EL_CSTR(handler);
if (hname && looks_like_string(handler)) {
http_set_handler_v2(handler);
@@ -5511,6 +5594,9 @@ el_val_t getpid_now(void) {
static el_val_t _el_args_list = 0;
void el_runtime_init_args(int argc, char** argv) {
/* First line of every generated main(): a client that leaves must never be
* able to signal this process to death. See el_runtime_ignore_sigpipe. */
el_runtime_ignore_sigpipe();
_el_args_list = el_list_empty();
for (int i = 1; i < argc; i++) {
_el_args_list = el_list_append(_el_args_list, EL_STR(argv[i]));