Compare commits

...

16 Commits

Author SHA1 Message Date
Neuron c9da8d917a chore: regenerate dist/soul.c after rebasing onto main
A clean rebase is not a consistent build input: git resolves the compiled amalgam
as an ordinary file and picks a winner, so dist/soul.c ends up holding one side's
code and not the other. The stamp gate catches it; this commit fixes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:12:53 -05:00
Tim Lingo 7f445a4b95 feat(engine): tools + agentic loop on the OpenAI wire, and two chat-breaking fixes found proving it
Teaches the OpenAI-format lane (Groq/OpenAI/Grok/Gemini/Ollama) to offer tools,
execute them, and loop — the capability that until now existed only on the
Anthropic wire. The tool-execution, consent, bridge and run-progress machinery is
reused unchanged; only the wire dialect is new.

Two pre-existing defects were found while proving it, and are fixed here because
both silently break chat:

1. PROVIDER WIRING NEVER CONNECTED. The launcher exports SOUL_LLM_PROVIDER /
   SOUL_LLM_BASE_URL and puts the provider key in ANTHROPIC_API_KEY + SOUL_API_KEY;
   the engine's provider fork read only NEURON_LLM_0_*, which nothing sets in a
   customer build. So use_openai was ALWAYS false: every non-Anthropic user's turns
   went to api.anthropic.com carrying, say, a Groq key, and came back
   "llm unavailable". Proven side-by-side against the pinned round-9 brain
   (sha256 15cf7d1b…): identical env, shipped brain = "llm unavailable" both chat
   modes with ZERO calls to the configured endpoint; this build = a real answer,
   with the probe logging POST /v1/chat/completions and Bearer <provider key>.
   Fixed brain-side only (env fallbacks) — no app or launcher change needed.

2. TRUNCATION SPLITS UTF-8 CHARACTERS. The session preload cuts recalled memory at
   fixed BYTE lengths (continuity snippet 350; session_preload_bullets per bullet).
   A cut landing inside a multi-byte character leaves a dangling lead byte in the
   SYSTEM PROMPT, making the whole request body invalid UTF-8 — providers reject it
   and the user sees an unexplained failure. Captured from a real body: 18,710 bytes,
   decode fails at 18,248 on 'e2', a box-drawing rule (U+2500 = E2 94 80) sliced in
   half. Trigger is ordinary content — em dash, curly quote, accented name, emoji,
   table border — and it gets MORE likely as memory grows. Shared code: this hit the
   Anthropic wire too. Fixed with utf8_safe_slice() applied at BOTH cut sites.

WHAT IS IN THE PORT
- llm_base_url / llm_wire_format / agentic_api_key: fall back to the launcher's own
  SOUL_LLM_* names; anthropic deliberately still returns "" so its native path is
  untouched (endpoint configurability remains neuron#62).
- openai_tools_json(): Anthropic tool schema -> OpenAI function schema; entries with
  no input_schema (Anthropic's server-side web_search) are skipped — they cannot
  execute on this wire.
- agentic_tools_no_web(): the standard set minus that server tool.
- openai_agentic_loop(): forked rather than parameterised, so agentic_loop — which
  carries every round-7/8/9 fix — is provably untouched. Same envelopes, same state
  keys, same consent policy (ask_all / escalate / builtin / always-allow), same
  client-bridge contract, same run-progress ledger, same 12-iteration cap.
- ADR-0005 mirrored on this wire: parallel_tool_calls:false is sent explicitly, and
  if a provider ignores it we honour the FIRST call and echo only that one, so the
  conversation we send is never self-contradictory. The drop is logged loudly.
- The assistant turn echoes the provider's own content bytes (json_get_raw), so a
  JSON null stays null and nothing is lost to a decode/re-encode round trip.
- Tool results are embedded already-escaped (dispatch_tool json_safe's them);
  truncation trims a dangling escape so a cut can't invalidate the body.
- bridge_save() gains a "wire" scalar and agentic_resume branches on it, so a
  suspended turn resumes on the wire it suspended on. Legacy blobs (no field) resume
  as anthropic. The field is read from the blob's SCALAR HEAD only — an unbounded
  first-match scan would run on into messages_raw, which is model-controlled, and
  that is exactly the round-9 resume defect. Pinned by a test.
- Three fork sites: handle_chat_agentic, handle_dharma_room_turn_agentic,
  agentic_resume. Tool assembly is computed once per lane at both entry points
  (it makes an HTTP call to the connector bridge; it was being paid for twice).

TOOLING THAT DID NOT EXIST
- tests/run-el-test.sh — engine tests were never runnable: elc is a compiler, it
  emits C and exits. This emits the test to C, compiles soul.c with main renamed
  away, links the rest + the repo-pinned runtime, and runs it. It also COMPUTES THE
  VERDICT, because every counted test file's "N passed, M failed" summary is a
  permanent 0/0 — the counters increment inside if BLOCKS, which El scoping
  discards (9 files; real fix filed as neuron#116). Proven to discriminate with a
  deliberately-broken assertion.
- tests/gate-openai/ — deterministic OpenAI-dialect provider stub + scenarios +
  driver + hostile modes, and a strict request validator that rejects any
  Anthropic-shaped field so dialect leakage fails loudly.

VERIFICATION (rungs named)
- E2E-VERIFIED against a LIVE provider (Anthropic's OpenAI-compatible endpoint,
  confirmed live): real answer; a tool call whose out-of-root path was DENIED by the
  guard, after which the model refused to claim success ("I won't tell you I did it,
  because I didn't"); then a valid path -> file physically on disk with exact content,
  honest reply, ledger with per-round entries + {done:true}.
- Deterministic lane gate: 11/12 in both consent configurations (bridge + local);
  hostile providers produce no hang and no fabricated answer; the 12-iteration cap
  trips with its honest message. The one FAIL is oa-tools-off and is NOT this port —
  see "Known, not fixed here".
- ANTHROPIC LANE UNCHANGED: gate9 32/32 on this build and on the pinned round-9
  brain; request bytes differ only within the noise band that two runs of the
  UNMODIFIED brain also produce (proven with a baseline-vs-baseline control), and
  the preload sections — the shared code touched here — are byte-identical.
  The rig discriminates: the round-8 brain scores 24/32 on it.
- verify-soul-contract.sh: PASS (27/27 routes, no hard-deletes).
- Unit: test_bridge_serialization 36/36 (incl. 8 new wire/field-order assertions),
  test_utf8_slice 18/18, test_agentic_tools 18 PASS / 0 FAIL / 3 documented skips.

KNOWN, NOT FIXED HERE (deliberate)
- Tools:Off on an OpenAI provider still fails: the non-agentic path goes through the
  el-runtime provider chain, which appends /v1/chat/completions to a base URL that
  already ends in /v1 -> /v1/v1/... 404. Runtime/plain-chat territory, untouched
  mid-beta. Note openai_chat_complete() has zero callers — that lane is served
  entirely by the runtime chain.
- The 12-iteration cap does not bound a chain of BRIDGED tools (iteration is
  per-invocation and resume starts fresh). Parity with the Anthropic lane.
- run_progress resets on each resume, so a client rendering cumulative steps across a
  consent pause sees earlier legs vanish. Parity with the Anthropic lane.
- verify-soul-contract.sh needs bash >= 4; under macOS's stock bash 3.2 it dies
  instantly with a FALSE red ("local: -n: invalid option").
- Groq-specific live E2E not run: no Groq key exists on this machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:12:13 -05:00
Tim Lingo f1f52bcb2f merge: Stage 1 structural audit as a real route (#142/#91)
Neuron Soul CI / build (push) Failing after 47s
Neuron Soul CI / deploy (push) Has been skipped
The runtime-vs-owner divergence check. Its absence let a ~24,000-node loss run for
weeks with every boot reporting green, which is most of why the last two days were
spent rediscovering by hand what this route would have said.

Follows the spec rather than inventing a metric: CGI provisional
05-detailed-description.md Stage 1 calls for an annotated characterization of the
graph's structure, so the route returns findings with score null BY DESIGN. A number
here would be a fabrication dressed as rigour.

Rebased across 27 commits of drift. The rebase merged cleanly at source level and
that was misleading — dist/soul.c held one side's code and not the other, because
git resolved the amalgam as an ordinary file. The stamp gate from 9fd8c11 caught it
on its first real use. Without it this would have landed an engine containing the
audit but none of the 08-09 engine work, or the reverse: neuron#133 again.

Verified before merging, not after:
  stamp OK                    dist/soul.c matches the sources (1,204,442 bytes, 1,259 bodies)
  builds from committed input 920,776 bytes
  interface                   108 -> 110 routes, nothing removed
  the route answers           stage 1, annotated_characterization, findings present

Closes #142. Refs #91.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:03:10 -05:00
Neuron aad988ecbf chore: regenerate dist/soul.c after rebasing onto main
The rebase merged cleanly at source level but left dist/soul.c holding one side's
code and not the other — main's regenerated amalgam vs this branch's. The stamp
gate added in 9fd8c11 caught it on its first real use:

  FAIL: dist/soul.c is STALE. It does not match the current .el sources.
  Sources that changed: neuron-api.el, routes.el

Without that gate this branch would have looked clean and shipped an engine
containing the structural audit but none of the 08-09 engine work, or the reverse.
That is exactly neuron#133, which once hid five merged fixes including a P0.

Regenerated from the rebased sources: 1,204,442 bytes, 1,259 inlined bodies.

Verified after: stamp OK; builds from its own committed input (920,776 bytes);
interface 108 -> 110 routes with nothing removed, adding /api/neuron/audit/structural;
and the route answers — stage 1, assessment_kind 'annotated_characterization',
score null by design, with findings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:01:25 -05:00
Tim Lingo 7b86e6f72c feat(soul): Stage 1 structural audit as a real route — an annotated characterization, not a score (#91)
`runStructuralAudit` has been an advertised MCP tool with nothing behind it: the
dispatcher GET'd /session/begin and returned that unrelated session digest under
an audit tool's name. Meanwhile the failure the audit would have caught ran
silently for about three weeks — the soul reporting 103,089 nodes while the
engram, which OWNS persistence, held ~79,900, a crash discarding the difference,
and every boot reporting green throughout, because nothing in the system ever
compared the two sides.

WHAT THE PATENT SPECIFIES, AND HOW IT SHAPED THIS
  CGI provisional, 05-detailed-description.md, "Stage 1: Structural audit 430".
  Two clauses did the design work. First the four things the module evaluates:
  the density and typed distribution of causal edges; value/execution-record
  consistency; the richness and connectivity of the self-model; and wonder-
  manifest authenticity. Second, and decisively: it "produces a coherence
  assessment 432 — NOT A BINARY SCORE but an annotated characterization of the
  graph's structural properties."

  So every finding carries its numbers AND a plain-language note saying what
  they mean and how they were obtained. There is no pass/fail and no composite
  health figure, and `"score":null` is emitted explicitly so a reader cannot
  mistake its absence for an omission.

WHAT IS IN STAGE 1 (four findings)
  owner_runtime_divergence   — the motivating case. Runtime counts vs the owner's
      own GET /api/stats, the delta, and the trend against the previous audit, so
      a second call answers "is the gap growing?" rather than restating it.
  self_model_connectivity    — the three identity pillars plus the self root:
      present, content length, one-hop degree. This RETIRES the Claude-side vitals
      identity block, which lived outside the system it was checking and went on
      reporting green while the memory-philosophy pillar was absent from the live
      graph. Asking the running soul is the designed mechanism; a shell probe was
      the fourth patch on the same hole.
  typed_edge_distribution    — exact counts against the claim-10 vocabulary, plus
      density, plus a separate count of LOWERCASE near-misses ("causes" vs
      "Causes"): "the vocabulary is unused" and "the vocabulary is misspelled by
      the write paths" are different defects with different fixes.
  orphans_and_dangling_edges — the tool's own long-standing promise.

WHAT IS DEFERRED, AND WHY IT IS DATA RATHER THAN A COMMENT
  Value/execution-record consistency and wonder-manifest authenticity ship as a
  `deferred` array that MEASURES the populations they would need (Prediction and
  WonderQuestion nodes) and reports those counts as the reason. Both are ~0 today
  — WonderQuestion because of a known write/read node-type mismatch. Asserting
  value coherence or a pull-weight correlation on an empty population would be a
  fabricated result, which is worse than a stated gap.

MEASUREMENT HONESTY: EXACT WHERE CHEAP, SAMPLED WHERE NOT, ALWAYS LABELLED
  Counts, edge typing and self-model connectivity are exact. Orphan and dangling
  rates are sampled, because engram_find_node_index is a linear scan — an
  exhaustive dangling check is O(nodes x edges), ~2.2e9 string compares at today's
  scale. Samples are UNIFORM across the whole population (str_index_of_all gives
  every edge offset in one pass, so any index is O(1); json_array_get would have
  been O(n^2)), never head-of-list, and each figure ships with its own sampled /
  population / exhaustive fields. ?edge_sample= and ?node_sample= at population
  size run either check exhaustively. The real fix is an id index in the runtime.

ONE BUG THIS FOUND IN ITSELF, CAUGHT IN TEST
  http_get does not return "" when the owner is unreachable — it returns a JSON
  error object. Testing only for "" made a DEAD owner read as reachable with
  node_count 0, so the audit reported 100% divergence and named it data loss.
  Reachability is now proved by the presence of the node_count field, and the
  owner's raw reply is attached. A confident wrong answer is exactly what this
  route exists to stop.

  Edge findings need relation labels and the runtime has no edge-enumeration
  builtin, so they use the same scratch export GET /api/graph/edges already uses
  (engram_save to TMPDIR, never the owner's canonical file — #117). That is a
  large write on a large graph, so this is a manual route, not a timer; ?edges=0
  skips it.

  neuron-api.el:900-1273  handler + helpers
  routes.el:567,752       GET and POST /api/neuron/audit/structural
  mcp-wrapper/src/main.el:113,682  tool description + dispatch off /session/begin
  dist/soul.c             regenerated (1255 bodies)

Rung: E2E-VERIFIED. Soul built from this branch (gen-soul-amalgam + cc-brain,
921,192 bytes, 16 warnings, 0 errors), booted on throwaway ports 7893/7896/7897
with throwaway HOMEs against a stub owner on 7894. Three scenarios pass: owner
reachable (runtime 62 vs owner 42, delta 20 / 32.2%, trend flat on the second
call; 12/20 edges claim-10 typed, 3 lowercase near-misses; 50/62 orphans, 3/20
dangling — every figure matches the fixture by construction), owner unreachable
(reported as a finding with the raw reply, not a crash), and file mode (owner
"none", divergence undefined). Reached end-to-end through the MCP tool via a
locally built wrapper. verify-soul-contract.sh: GATE PASS, 27/27 routes +
immutability. No process left running; live :7770 and :8742 untouched (GET only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:54:42 -05:00
Tim Lingo bf521974af merge: semantic retrieval, self-seeded context, connections on write, and one build lineage
Neuron Soul CI / build (push) Failing after 11m54s
Neuron Soul CI / deploy (push) Failing after 14m43s
Consolidates the 2026-08-09 engine work. Each piece was measured in an isolated lab
against the real corpus, gated, and verified on the operator machine before landing.

WHAT IT CONTAINS

  Semantic retrieval + the #146 delete fix, shipped together on purpose. The
  semantic leg was once wired into engram_search_json, which seven sites use as a
  keyed read and then delete every result from; one runs at every boot. Measured
  cost when live: ~234 real records destroyed per startup, identity among them.
  Retrieval must never reach production without the engram_recall_json split.

  Self-seeded compiled context. The design: 'Every compilation query begins at the
  self-model node and traverses outward.' Ours seeded from a hardcoded string, so
  compiled context held 0 identity records. Now 18, bounded the same way every other
  list in that handler is — the unbounded version is what set self_neighbors to []
  after it closed the socket on every call.

  Connections on write. CCR claim 29 requires linking candidates with typed edges
  'rather than appending as unlinked content'. Unlinked append was the only
  behaviour we had: 5% of nodes connected, no edge created by any write in 19 days.

  dist/soul.c drift becomes a build failure (#133), and deploys now build from that
  same committed input rather than a scratch amalgam — one lineage, provenance
  recorded, enforced by a gate in the deployer.

MEASURED, all on the real corpus or the live machine

  rephrased-question recall     0/43 -> 20/43 (bar was >=6 queries; nonsense controls held 10/10)
  identity in compiled context  0 -> 18
  connections per memory        0 -> ~2 on write; 2,828 backfilled into history
  no data loss across boots     44,418 / 44,418 / 44,418 over three restarts

FOUR DEFECTS FOUND BY MEASURING RATHER THAN READING, each fixed here

  an inert test arm (both arms byte-identical but for a float rounding artifact);
  an association hook wired into a path the request never takes (zero edges across
  four writes); an identity exclusion matching lowercase 'self' that let a memory
  link into the identity graph, whose verification shared the blind spot and printed
  PASS; a telemetry filter matching hyphenated 'state-event' that missed
  'INTERNAL STATE EVENT'. Common shape: a filter tested a proxy for the property it
  cared about. Assert the property.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:40:56 -05:00
Neuron 5743568bf1 chore(engine): deploys build from the committed input, not a second lineage
Until now deploys were built by build-soul.sh from a scratch amalgam while CI
compiled the committed dist/soul.c. Two lineages — 'what runs' and 'what the repo
says builds' were different artifacts. That is #133/#111 in another costume, and it
is how a round-9.1 brain shipped matching no committed source at all.

build-soul-from-dist.sh asserts dist/soul.c matches the .el sources, compiles it
with CI's own flags (-O2 -DHAVE_CURL -rdynamic; the CI comment explains -rdynamic —
without it the runtime cannot resolve its HTTP handler by name and the binary serves
nothing on every route), and writes a .provenance sidecar recording the dist/soul.c
hash, the stamp hash and the commit.

deploy_binary.sh gains GATE 0: refuse any soul without provenance, or built from a
dist/soul.c other than the one in the repo now. The override
(NEURON_DEPLOY_UNSTAMPED=i-accept-two-lineages) exists deliberately — a gate with no
escape hatch gets bypassed by disabling the gate, which is worse than one that
announces itself.

Verified before enforcing, so the sanctioned path is not a broken one:
  interface parity with the deployed binary   108 routes in, 108 out
  boots and serves                            2s
  carries today's work                        24 self-neighbours, 17 identity records
  GATE 0 refuses an unstamped binary          exit 10
  GATE 0 accepts the stamped one, deployed    durability PASS

Production now runs 4845db3d, built from dist/soul.c at 9fd8c11. After the swap:
identity in context 17, connections still form on write, semantic recall intact,
mind and store at delta 0.

Refs #133, #111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:38:01 -05:00
Neuron 9fd8c11670 chore(engine): make dist/soul.c drift a build failure instead of a silent ship (#133)
#133 regenerated the amalgam once and said so itself: 'Nothing in the tree
regenerates this file. Only a human running the recipe. It lags in batches, never
per-change, and it will drift again.'

It drifted again. Every binary deployed on 2026-08-09 was built by build-soul.sh
from a scratch amalgam that never touches dist/soul.c, so the committed build
input fell 2,761 bytes behind the sources by a different route than #133 describes.

Auto-regeneration is not available: the CI workflow records that elc needs 24GB+
of virtual memory and would OOM the runner. So the build cannot regenerate the
file. It can refuse to compile a stale one, for free and with no compiler.

tools/soulc-stamp.sh records a content fingerprint of every .el source at the
moment the amalgam is generated. --check recomputes and compares; divergence exits
1 and names the changed files and the recipe. Wired into CI ahead of the compile.

dist/soul.c regenerated from current sources: 1,176,361 -> 1,179,122 bytes, 1,247
inlined bodies (gate wants >=1200), and verified to compile clean at 903,552 bytes.

Demonstrated to FAIL on the bad input, per postmortem 0004's rule that a gate which
only passes on good input proves nothing:
  fresh stamp        -> OK,   exit 0
  one .el modified   -> FAIL, exit 1, names memory.el
  source restored    -> OK,   exit 0

Refs #133, #111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:32:56 -05:00
Neuron be0f9d1afe feat(soul): memories form connections when written
The design rejects what this system did: promotion must link candidates 'using
typed semantic edges rather than appending as unlinked content' (CCR claim 29).
Unlinked append was the only behaviour we had. Measured 2026-08-09: 14,214 edges
over 80,936 nodes, 5% of nodes connected to anything, and no edge created by any
write since 2026-07-19 across 27,000+ new nodes. A memory with no connections is
unreachable by spreading activation, so retrieval degrades to literal matching.

On write, a memory is now linked to its top related existing memories.

Bounds, each bought with a specific failure:
  - max 3 edges per memory (link_memories.py's cap: precision over spray)
  - never link to identity. The existing policy is explicit that 'memories must
    not pollute the self traversal by similarity; only an explicit citation may
    touch identity'.
  - never link telemetry (state-event, soul-response, boot_count, loop-outcome,
    search-result): ~97% of daily write volume. Linking it would add thousands of
    noise edges a day and re-flatten the graph in the name of connecting it.
  - fail-soft: a failed association never fails the write
  - associate only AFTER durability, so no edge points at a node that did not
    persist — that is the dangling-edge defect the 08-09 cleanup removed 830 of

Two defects found by measuring rather than reading, both fixed here:
  1. Hooking mem_store alone produced ZERO edges across four real writes. The HTTP
     memory route writes via wt_node directly; mem_store serves only the awareness
     telemetry paths we refuse to link.
  2. A lowercase-only identity check let a memory link to 'Self — Values
     (grounded)' — the exact pollution the policy forbids. My verification shared
     the blind spot and printed PASS. Now uses str_lower.

Measured, lab, same corpus: ~1-2 edges per memory; identity edges unchanged at
475; zero identity leaks post-fix including an adversarial batch of six memories
written about values. Edges route through wt_edge so they reach the owner.

Rung: E2E-VERIFIED in an isolated lab. Not deployed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:00:04 -05:00
Neuron 1742d0b575 feat(soul): seed compiled context from the self node, bounded
The design is explicit — 'Every compilation query begins at the self-model node
and traverses outward... structural reachability from the self-model node is a
precondition for any node to appear in compiled context.' (will-anderson
patents/drafts/engram-claims.md, Self-Seeded Activation. DRAFT, not a filed
provisional — cited as such.)

Measured before: compiled context held 0-1 identity records out of 10, because
compilation seeds from a hardcoded text string and never from the self. Even an
explicit 'my values identity who I am' query returned a boot counter and
state-events.

This restores the designed behaviour without repeating the failure that set
self_neighbors to [] originally: that was an unbounded ~90KB neighbour dump which
closed the socket on every call. Same bound as every other list in the handler.

Cap is 24 rather than 8 because the self root's first 8 neighbours are tag nodes
('neuron', 'tier:note', 'disposition:experimental', 'imprint', 'traversal') that
crowd out the substantive identity records behind them. The root has 34
neighbours of which 23 are identity.

Measured after, same corpus, binary as the only variable:
  identity records in compiled context  0 -> 17
  response size                         10,806 -> 18,985 bytes
  interface gate                        108 routes in, 108 out, none removed

Rung: E2E-VERIFIED in an isolated lab. Not deployed — this is the identity layer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 10:45:12 -05:00
Tim Lingo c6772e3d27 Merge branch 'feat/semantic-leg-dropout' into integrate/semantic-plus-writethrough
Neuron Soul CI / build (pull_request) Failing after 11m47s
Neuron Soul CI / deploy (pull_request) Failing after 14m48s
2026-08-09 09:47:44 -05:00
Tim Lingo b9ef66cae9 fix(engram): the semantic leg was deleting 234 memory records per boot
Neuron Soul CI / build (pull_request) Failing after 14m34s
Neuron Soul CI / deploy (pull_request) Has been skipped
The accumulated retrieval stack (iterations 1-9) put the claim-24 semantic
leg and the claim-10 associative leg on engram_search_json — the function
~40 internal .el call sites already used as a KEYED read. Seven of those
sites delete every record that comes back ("prune all existing X nodes,
keep exactly one"): memory.el:176, sessions.el:250/268/444/523,
soul.el:359.

mem_boot_count_inc() calls engram_search_json("soul:boot_count", 50) and
engram_forget()s all 50 results. With a lexical leg that returned 1 record.
With a semantic leg it returns 50 — the 49 nearest neighbours of the STRING
"soul:boot_count" — and the soul deletes them.

MEASURED on the harness corpus, isolated, read-only, zero writes from any
caller: 234 node records destroyed in a single boot. The deletion list is
the soul's own lookup result list, in rank order. Casualties include 6
Knowledge nodes, a layer-1 "CORE IDENTITY - GENESIS, LINEAGE" Memory, the
value node kn-58874a74, and the gold answers to 8 of the 75 gold-set
queries. After the fix: 1 deletion, which is the one the code intends.

THE BOUNDARY, from Will. Claim 24 authorises the vector index "to respond
to EMBEDDING SEARCH QUERIES by returning the node records whose embedding
vectors have the highest cosine similarity to a query vector". A keyed
state read is not an embedding search query; it is the identifier-keyed
retrieval of claim 23 ("node records are stored under a key encoding the
node identifier"). One function served both, so a nearest neighbour of
"soul:boot_count" was treated as a boot counter.

So: engram_search_json returns to its lexical contract, and the legs move
to engram_recall_json, which is what /api/neuron/recall reaches — the route
the MCP wrapper, the app, and this harness all call. Retrieval quality on
that route is unchanged by construction.

MEASURED, 75-query extended gold set, embedded corpus, vs the iteration-9
baseline: +3 / -0 (q15, q28, q60), p=0.2500, hit@5 53.8 -> 58.5%, latency
1.02x, every regression guard held, nonsense 10/10. Net +3 against a floor
of 6 is NOT-SHOWN and I am not calling it an improvement. The deliverable
is the defect.

Diagnostics kept, env-gated (EG_DIAG / EG_DIAG_ID), zero cost when unset:
node/embedding census at load, per-query leg dump, and a FORGET log — the
last is the regression detector for exactly this class of bug.

LIMIT, stated: handle_api_search_knowledge still uses the lexical function.
It is a retrieval surface and arguably wants the legs, but nothing in this
harness measures it, so I did not change unmeasured behaviour.
2026-08-07 18:12:43 -05:00
Tim Lingo 9717a4eeaf measure: claim 24 unflooring is +4 (NOT-SHOWN); asymmetric embedding prefixes are -5 (discarded)
Measured on the 75-query extended gold set (iteration 8's held-out extension)
against the certified stack baseline results-stack-ext.json, on the embedded
corpus. Three runs of the candidate, zero drift.

A. CLAIM 24 WITHOUT THE THRESHOLD - net +4, NOT-SHOWN, kept in the tree.
   fixed  : q14, q25 (in-sample paraphrase), q43, q52, q63, q67 (held-out)
   broken : q15 (paraphrase), q28 (associative)
   8 discordant, McNemar exact p = 0.2891, floor is 6.
   heldout_paraphrase 16.7% -> 30.0%, paraphrase 61.5% -> 69.2%.
   Every regression guard held: exact_rare 6/6, phrase 7/7, nonsense 10/10,
   superseded 2/3. Latency FLAT: p50 641 -> 632 ms.
   The in-sample half (+q14 +q25 -q15 -q28 = 0) was already on record in
   iteration 7's cmp-nogate.json, so only the held-out +4 is new.

B. ASYMMETRIC TASK PREFIXES ON THE EMBEDDER - net -5, REVERTED in this commit.
   Rationale was sound and the prediction was wrong, which is why it was worth
   measuring: nomic-embed-text is an asymmetric retrieval encoder and this file
   embedded query and document bare on both sides. Prefixing does exactly what
   the model card implies for the far-away cases - it rescued q42 (gold at
   GLOBAL COSINE RANK 25,564) and q39 - but it re-ranks the whole space and
   broke more than it fixed:
   fixed  : q24, q39, q42
   broken : q18, q19, q22, q31, q43, q44, q52, q63
   heldout_paraphrase 30.0% -> 23.3%, paraphrase 69.2% -> 53.8%.
   The corpus and the reproducer are kept (embed-corpus-prefixed.py,
   snapshot-pre-repair-20260806-embedded-prefixed.json) so nobody re-runs it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:45:05 -05:00
Tim Lingo 9790d9342d feat(engram): claim 24 without a threshold, and an embedding substrate that knows query from document
Two changes to the ONE leg that generalises. Iteration 8 measured that out of
sample the semantic leg contributes 100% of the stack's gain and the graph leg
contributes nothing, so this is where the remaining headroom is.

1. THE 0.60 FLOOR IS A PER-QUERY LOTTERY, AND CLAIM 24 HAS NO THRESHOLD IN IT.
   06-claims.md l.148: "respond to embedding search queries by returning the
   node records whose embedding vectors have the HIGHEST COSINE SIMILARITY to a
   query vector, independently of the spreading activation traversal." A
   ranking. ENGRAM_EMBED_SEED_MIN is defined at el_runtime.c l.6094 as the
   HippoRAG seed-JOIN threshold and l.6102 admits the read-path leg merely
   "reuses" it. Measured on the 30 held-out paraphrases: the query's own top-1
   cosine ranges 0.564-0.680, so the constant keeps a rank-1 answer for one
   query and discards a rank-1 answer for the next. Six golds sit at global
   cosine rank 1-2 scoring 0.564-0.589 - discarded by nothing but the constant.
   What holds the nonsense controls is the corpus-vocabulary gate (nhits == 0),
   not this floor. Cosine clamped to [0,1] per 05-detailed-description l.69.

2. THE VECTORS THEMSELVES ANSWER THE WRONG QUESTION. EL_EMBED_MODEL defaults to
   nomic-embed-text, an ASYMMETRIC retrieval encoder trained with task prefixes.
   Embedding query and document bare - as this file did on both sides - measures
   topical similarity rather than answer-hood. eg_embed_fetch now takes the task
   prefix: EL_EMBED_QUERY_PREFIX on the three query call sites, EL_EMBED_DOC_PREFIX
   on the two backfill sites. Restores no claim, and says so: Will specifies only
   "computed by an embedding model over the node's content" (l.17), so the model
   is his and its correct use is ours. It is the substrate under claim 24 -
   the index is only as good as the vectors in it.

Reproducer for the derived corpus: tools/retrieval-eval/embed-corpus-prefixed.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:34:58 -05:00
tim.lingo 9501e4ac12 Merge pull request 'fix(engine): memories written through the soul now reach the store that owns them (closes #117)' (#134) from feat/soul-write-through into main
Neuron Soul CI / build (push) Failing after 11m27s
Neuron Soul CI / deploy (push) Failing after 14m46s
2026-08-07 21:07:14 +00:00
Tim Lingo dd952c0e46 feat(soul): write-through to the persistence owner — memories survive restart (#117)
Neuron Soul CI / build (pull_request) Failing after 14m57s
Neuron Soul CI / deploy (pull_request) Failing after 14m39s
The soul obeys half of its own ownership rule. soul.el:571-573 says "when
ENGRAM_URL is set the HTTP Engram owns persistence — the soul must NEVER write
to the local snapshot", and it doesn't. But nothing was ever built to hand the
soul's writes TO that owner: sync is pull-only (/api/sync -> engram_load_merge),
so every node created inside the soul lived in process RAM and was shed on
restart. Measured live 2026-08-07: soul node_count=102184, engram 79197.

SCOPE CORRECTION vs the earlier internal spec: engram provisional claim 17's
"pull-then-push" is a PEER-ENGRAM to PEER-ENGRAM protocol (claims 15-18 say so
explicitly). The soul is a CALLER of the database API, not a peer. Claim 17 is
NOT authority for a soul<->engram contract and is no longer cited as such. The
design here follows from the ownership rule alone.

Mechanism: a new Accessor, persist.el, is the single boundary. Writes stage a
delta to a filesystem spool and are pushed to the owner via POST /api/load-merge
— NOT POST /api/nodes, which mints a new server-side id (breaking dedup and
edges) and drops label/tier/tags/importance/confidence (verified in a sandbox:
a tier "Canonical" probe came back "Working"). load-merge preserves the id and
every field, dedups nodes by id and edges by (from,to,relation) so retries are
no-ops, and calls persist_canonical() so THE OWNER writes its own file — the
ownership rule is honoured rather than worked around.

Spool-and-drain rather than push-per-write: measured ~0.38s per load-merge at
live scale (79k nodes/176MB), and a chat turn writes 5-7 nodes. The spool is on
disk, not in process state, because the soul serves each connection on its own
pthread and a shared buffer would lose entries to a read-modify-write race. That
also buys crash recovery: writes orphaned by kill -9 are drained on next boot.

Honesty: api_persisted (the gate all 10 MCP write handlers pass through) and
mem_store now assert AT THE OWNER instead of reading back the soul's own RAM.
With the owner down a write returns {"ok":false,"error":"write_not_persisted"}
and the delta is queued — where main returns {"ok":true} for a write that dies.

Coverage: 35 node sites + 9 edge sites routed through the boundary. Deliberately
excluded, with reasons in persist.el: 4 InternalStateEvent sites (Will's own
telemetry carve-out), the boot counter and the persona (both already have
bespoke owner-side write-backs), and soul.el's 54 genesis identity edges
(file-mode only). engram_strengthen and engram_forget are NOT propagated —
load-merge cannot update or delete, and hard-deleting at the owner would fail
verify-soul-contract.sh section B.

Also fixed here:
- routes.el GET /api/graph/edges engram_save()'d straight over the owner's
  canonical snapshot.json — a read route, in a non-owner process, clobbering the
  canonical on every call. Same defect class Will removed from the engram in el
  dc39a61. Now exports to a scratch path. With this gone the soul writes nothing
  at all in HTTP mode.
- persist.el must clear the runtime's _tl_fs_read_len hint after every fs_read.
  In vendored runtime v1.0.0-20260501 that hint becomes the NEXT response's
  Content-Length, so reading a spool file mid-request made an 86-byte reply go
  out as 497 bytes with 411 bytes of adjacent heap trailing it. Caught and fixed
  at our boundary; the runtime class was fixed upstream in el 43636ae, which is
  not the pinned runtime here.

Rung: E2E-VERIFIED, discriminating. Same harness, same engram binary:
  write-through: LEG 1 PRESENT at owner, LEG 2 SURVIVED kill -9 + restart
  main:          LEG 1 ABSENT  at owner, LEG 2 LOST
verify-soul-contract.sh: GATE PASS on both builds (27/27 routes, immutability).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 12:47:31 -05:00
45 changed files with 22851 additions and 778 deletions
+10
View File
@@ -63,6 +63,16 @@ jobs:
cp vendor/el-runtime/v1.0.0-20260501/el_runtime.h /opt/el/runtime/el_runtime.h
echo "El runtime PINNED to v1.0.0-20260501: $(ls /opt/el/runtime/)"
# neuron#133: CI compiles dist/soul.c, NOT the .el sources. On 2026-08-07 a
# build off main would have shipped an engine with none of five merged fixes,
# including a P0 safety fix, while main's source read as correct. The runner
# cannot regenerate the amalgam (elc needs 24GB+ virtual memory), but it can
# refuse to compile a stale one. Fails loudly with the recipe in the message.
- name: Verify dist/soul.c matches the sources
run: |
chmod +x tools/soulc-stamp.sh
./tools/soulc-stamp.sh --check
- name: Build neuron soul binary
run: |
RUNTIME=/opt/el/runtime
+139
View File
@@ -0,0 +1,139 @@
# PORT-NOTES — openai tools port working state (2026-08-06, session handoff-safe)
Spec: `docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md` (Tim-approved 2026-08-06). Tasks #1-5
tracked in-session (1 ✓ wiring verdict, 2 ✓ stub rig, 3 in-progress = THIS, 4-5 pending).
Worktree: HERE (`_wt-openai-tools`, branch `feat/soul-openai-tools-v2` @ dba755d). Round-9 trees
READ-ONLY. Nothing committed yet.
## Step-0 verdict (evidence in journal note ncli-653ba964dd76)
Shipped app never wires the v1 lane: launcher exports `SOUL_LLM_MODEL/PROVIDER/BASE_URL` +
`ANTHROPIC_API_KEY`+`SOUL_API_KEY` (= Keychain key for WHATEVER provider; installer/macos/
neuron-daemons.sh:288-300 on hotfix/beta-round9); brain reads only SOUL_LLM_MODEL (chat.el:8) and
NEURON_LLM_0_* (chat.el:1768-1794) which nothing sets. `/api/config` PATCH ignores llm_* fields
(studio.el:36 handle_config: POST-only, reads model/provider/api_key only).
**Bridge = brain-side ONLY (zero app-repo edits, zero round-9 collision):**
- `llm_base_url()`: NEURON_LLM_0_URL → fallback SOUL_LLM_BASE_URL when SOUL_LLM_PROVIDER ∉ {"","anthropic"}
- `llm_wire_format()`: NEURON_LLM_0_FORMAT → fallback derive from SOUL_LLM_PROVIDER (openai/grok/gemini/groq/ollama → "openai"; else "anthropic")
- `agentic_api_key()`: already works (ANTHROPIC_API_KEY carries the provider key); add NEURON_LLM_0_KEY → SOUL_API_KEY fallback.
## Design pins (stub asserts these — stub is green 58/58, tests/gate-openai/)
- Request MUST send `"tool_choice":"auto"` (string) + `"parallel_tool_calls":false` explicitly.
- `arguments` in tool_calls = JSON-ENCODED STRING; decode ONCE via json_get → feed dispatch_tool
verbatim. Stub's echo-mismatch check catches double-encode/decode (two-escaper trap).
- Assistant echo turn: `{"role":"assistant","content":null,"tool_calls":[...]}` VERBATIM from response.
- Feedback: `{"role":"tool","tool_call_id":"<id>","content":"<result string>"}`.
- Resume must NOT re-answer an answered id (stub 400s on repeat tool_call_id).
- Parallel tool_calls in a response: take FIRST only + log skip (mirror ADR-0005 stopgap); stub
scenario `parallel` proves behavior.
- No tools in request when tools array empty/absent turns (boot probes) — stub defaults tolerate.
## el idioms confirmed (from openai_chat_complete :1808-1854 + agentic_loop :2751-2838)
- JSON: `json_get(s,k)` decoded string · `json_get_raw(s,k)` raw subtree · `json_array_len` ·
`json_array_get(arr,i)` · build by string concat + `json_escape()` (:1797, OpenAI-lane escaper).
- HTTP: `let h: Map = {}` + `map_set(h,k,v)` + `http_post_with_headers(url, body, h)`;
Bearer auth via `Authorization` header when key non-empty (:1825-1830).
- Loop-carried vars must be top-level locals in the fn, mutated as if-expressions at while-body
top level (see :2760-2791 pattern + comment :2903-2904 region).
- Error shape: `str_starts_with(raw,"{\"error\"") || str_contains(raw,"\"error\":")` → return
`{"error":"llm unavailable","reply":""}` (:1835-1838).
## Remaining read map (before writing the fork)
- chat.el 2840-3200: block walk (2923-3000), policy gate (3009-3023: classify_tool_risk /
is_builtin_tool / ask_all / tool_auto_approved → needs_bridge), dispatch_tool call (3025),
tool_result feedback (3031, 3067-3072), run-progress ledger append (3078-3087), bridge_save
(3182), loop end + done envelope (~3100-3200).
- agentic_resume 3227-3293 (hardcoded Anthropic headers to make wire-aware; blob gets `wire` field,
legacy default anthropic) · handle_tool_result 3293+ · dharma fork site 3465 (calls agentic_loop
direct, no use_openai check today).
## Write plan (order)
1. Env fallbacks (edit llm_base_url/llm_wire_format/agentic_api_key) — small, first, testable alone.
2. `openai_tools_json(anthropic_tools: String) -> String` converter (walk array; per entry build
{"type":"function","function":{name,description,parameters:input_schema-raw}}).
3. `openai_agentic_loop(...)` fork: same signature as agentic_loop minus Anthropic-only params;
INCLUDE run-progress ledger + tools_log + iteration cap 12; NO container_id/ws_drift/web_search
(out of scope; strip web_search entry from tools via agentic_tools_literal()+connector merge,
NOT _with_web()).
4. Fork sites ×3: handle_chat_agentic :2695-2700 (route agentic to new loop when use_openai);
dharma :3465; agentic_resume wire-branch.
5. `chat.elh` extern decls. 6. Compile (recipe: dist/ + elc/elb per neuron-soul-build-deploy memory;
round-9 tree soul.c regen'd 08-06 proves toolchain live). 7. Gate: stub selftest recipe in
tests/gate-openai/README.md. 8. Anthropic-lane regression via gate9 (READ-ONLY consume from
_wt-beta-round9). 9. Live Groq E2E (scratch profile, free port, key via Keychain read-only).
## BUILD RECIPE — CORRECTED 2026-08-06 (the June memory is STALE for August code)
`~/el-sdk/el_runtime.c` (Jun 15) is MISSING builtins the Aug engine calls (`engram_wm_count`,
`engram_wm_top_json`, `http_delete_json`, `http_serve_async`) → link fails with
"symbol(s) not found for architecture arm64". Use the REPO-PINNED runtime:
```
mkdir -p <scratch>
elb --elc=$HOME/el-sdk/elc --runtime=vendor/el-runtime/v1.0.0-20260501 --out=<scratch>/
# "elb: link failed" at the end is EXPECTED and harmless — the per-module .c files are produced
cc -std=c11 -O1 -DHAVE_CURL -rdynamic \
-I vendor/el-runtime/v1.0.0-20260501 -I <scratch> -I /opt/homebrew/opt/openssl@3/include \
-L /opt/homebrew/opt/openssl@3/lib \
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
-o <scratch>/soul <scratch>/*.c vendor/el-runtime/v1.0.0-20260501/el_runtime.c \
-lssl -lcrypto -lcurl -lpthread -lm
```
Source: `_engine-plainchat-20260805/README.md:396-412`. Verified today: 0 errors, 887,296 B.
`elb` ALSO rewrites every `*.elh` in the tree (cosmetic em-dash→hyphen in the auto-gen banner,
plus true-ups) and drops a stray `soul..elh``git restore` the unrelated ones and delete the
stray before staging, or the diff drowns in noise.
## SELF-REVIEW FIX LIST (found by reading my own diff, 2026-08-06 — apply in ONE batch, then rebuild once)
- **F3 (CORRECTNESS, do first):** the assistant echo currently replays the provider's FULL
`tool_calls` array (`tc_arr`) while the loop answers only the FIRST call. If a provider ignores
`parallel_tool_calls:false`, the next request carries an assistant turn with N tool_calls and
only ONE `role:"tool"` response → most OpenAI-format providers 400 ("missing tool response for
id X") and the run dies. This is the same class as ADR-0005's Anthropic failure, but here it is
cheap to close: echo ONLY the honored call (`"[" + tc0 + "]"`), so the conversation we send is
self-consistent and the dropped call never existed from the model's view. The DRIFT log line
stays (honest accounting of what we dropped).
- **F4 (efficiency/latency):** `handle_chat_agentic` computes `agentic_tools_all()` at ~:2681
BEFORE the fork, then the OpenAI branch computes `agentic_tools_no_web()` again — two
`connector_tools_json()` calls per turn, each an HTTP round-trip to the connector bridge on
:7771 (two timeout exposures). Fix: compute the tools array ONCE, per lane, after `use_openai`
is known (check no other use of `tools_json` sits between :2681 and the fork before moving it).
Note: `openai_tools_json()` already skips any entry with no `input_schema`, so Anthropic's
server-side `web_search` entry is auto-dropped even if the full array is passed —
`agentic_tools_no_web()` is kept for EXPLICITNESS, not necessity.
- **F1 (debuggability):** the "no choices in response" branch logs a generic string and discards
the body. Log the response head (as the `is_error` branch does) — a provider that returns 200
with an unexpected shape is otherwise undiagnosable from the log.
- **OPEN QUESTION (evidence pending from the gate):** the tool-result feedback turn escapes with
`json_escape()` (this lane's escaper) rather than `json_safe()` (used everywhere else). The
Anthropic lane escapes that field with NEITHER, which is a latent defect on that side. If the
torture scenario shows any escaping loss, switch to `json_safe` and note the Anthropic-side
finding for Will.
## TEST HARNESS — built 2026-08-06 (Task 4 side-work, reusable by anyone)
- `tests/run-el-test.sh <tests/test_x.el> | --all` — the engine tests were NEVER runnable
before this (`elc` is a compiler: emits C to stdout and exits). It emits the test to C,
compiles `soul.c` separately with `main` renamed away (soul.c owns the daemon's real main
but also defines `layered_cycle` et al.), links the remaining modules + the repo-pinned
runtime, and executes. Modules cached under `/tmp/el-test-<worktree>/`; `REBUILD=1` forces.
- **The runner computes the verdict itself** because the test FILES cannot: all 9 counted
test files do `let pass_count = pass_count + 1` inside an if BLOCK, which El scoping
discards, so every summary line reads `0 passed, 0 failed` forever. Per-assertion
`PASS:`/`FAIL:` lines ARE reliable; the runner counts those, exits non-zero on any FAIL
or on zero assertions, and was proven to discriminate with a negative control (broken
assertion → 31 passed / 1 failed / exit 1). Real in-file fix filed: **neuron#116**.
- `tests/test_bridge_serialization.el`: 4 `bridge_save` calls updated for the new `wire`
argument, plus **Section 9** (8 new assertions) covering wire round-trip both ways, the
legacy no-wire blob (resumes as anthropic), and a FIELD-ORDER decoy guard — a fake
`"wire":"anthropic"` planted inside `messages_raw` must not beat the blob's own scalar.
That decoy is the round-9 first-match-scanner bug class, now pinned by a test. **32/32 green.**
## MEMORY-SAVE CAVEAT RESOLVED 2026-08-06
Earlier saves this session reported `-> OUTBOX only (real mind unreachable or read-back
failed)`. That was a **read-back verifier false negative, not data loss** — a direct
`POST :7770/api/neuron/recall` returns those notes from the live mind verbatim. Another
terminal was fixing exactly this (multi-word read-back probe) the same afternoon. Do NOT
re-save on an OUTBOX report without first querying the mind directly, or you duplicate nodes.
## Standing cautions
- PERSIST OFF on the real mind this boot (neuron#98/#92): journal saves only, ferry later. MCP link
down this terminal; use neuron_remember.py / neuron_recall.py.
- Aug-16: Groq retires llama-3.3-70b-versatile (separate P0, Tim's call, catalog swap).
- Never bind 7770/7779/17779; never touch ~/.neuron; round-9 worktrees read-only.
+19
View File
@@ -889,10 +889,29 @@ fn awareness_run() -> Void {
state_set("soul.last_beat_ts", int_to_str(now_ts))
// Persist in-process Engram (sessions, memories, conversation nodes)
// to local snapshot so they survive restarts.
// FILE MODE ONLY: "soul_snapshot_path" is set exclusively in the
// genesis+safe_to_seed branch of soul.el, and safe_to_seed is
// unconditionally false when ENGRAM_URL is set. In HTTP mode the
// owner persists; the soul must not (soul.el:571-573).
let snap_path: String = state_get("soul_snapshot_path")
if !str_eq(snap_path, "") {
mem_save(snap_path)
}
// WRITE-THROUGH RETRY (neuron#117). The HTTP-mode counterpart of the
// save above: hand anything still spooled to the persistence owner.
//
// This is the retry arm of the whole design. Deltas that could not be
// pushed owner down, owner restarting, transient refusal stay on
// disk and are re-offered here every heartbeat until they land. It is
// also the catch-all for writes made by the awareness loop itself,
// which never passes through the HTTP handler's flush point.
//
// No-op with no HTTP call when the spool is empty or ENGRAM_URL is
// unset, so an idle soul in file mode pays nothing for this.
let wt_pushed: Int = wt_drain()
if wt_pushed < 0 {
ise_post("{\"event\":\"write_through_backlog\",\"ts\":" + int_to_str(now_ts) + "}")
}
}
// Curiosity scan: idle-gated AND wall-clock based. Only fires when the
+466 -29
View File
@@ -1162,7 +1162,7 @@ fn hist_trim_with_bell_guard(hist: String) -> String {
+ " | evicted_at:" + ts_str
+ " | message:" + safe_content
let preserve_tags: String = "[\"bell-history\",\"bell:" + bell_level + "\",\"evicted\",\"affective\",\"BellEvent\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
preserve_content,
"BellEvent",
"bell:" + bell_level + ":preserved",
@@ -1210,7 +1210,7 @@ fn conv_history_persist(session_id: String, hist: String) -> Void {
if !str_contains(hist, "]") { return "" }
let tags: String = "[\"conv-history\",\"persistent\"]"
// FIX B: one label rule, shared with the agentic path. See conv_hist_label.
let node_id: String = engram_node_full(
let node_id: String = wt_node(
hist, "Conversation", conv_hist_label(session_id),
el_from_float(0.7), el_from_float(0.8), el_from_float(0.9),
"Episodic", tags
@@ -1408,7 +1408,7 @@ fn session_preload_bullets(nodes: String, max_bullets: Int, snip_len: Int) -> St
while i < limit {
let node: String = json_array_get(nodes, i)
let content: String = json_get(node, "content")
let snip: String = if str_len(content) > snip_len { str_slice(content, 0, snip_len) } else { content }
let snip: String = utf8_safe_slice(content, snip_len)
let bullets = if str_eq(snip, "") {
bullets
} else {
@@ -1770,7 +1770,14 @@ fn agentic_api_key() -> String {
if !str_eq(k1, "") {
return k1
}
return env("NEURON_LLM_0_KEY")
let k2: String = env("NEURON_LLM_0_KEY")
if !str_eq(k2, "") {
return k2
}
// Step-0 bridge (2026-08-06): the shipped launcher also exports the Keychain key as
// SOUL_API_KEY (neuron-daemons.sh). Honor it so a provider key configured through the
// app reaches this lane without any launcher change.
return env("SOUL_API_KEY")
}
// OpenAI-compatible providers (Ollama / OpenAI / Grok / Gemini)
@@ -1778,19 +1785,42 @@ fn agentic_api_key() -> String {
// OpenAI-compatible wire format (NEURON_LLM_0_FORMAT=openai) with a configured base URL
// (NEURON_LLM_0_URL, e.g. http://localhost:11434/v1 for local Ollama), basic chat turns are served
// here instead of the Anthropic agentic loop.
// v1 SCOPE: plain chat completion only NO tools / agentic loop yet (that is a follow-up port).
// This block is ADDITIVE: the Anthropic path is untouched and stays the default.
// v2 SCOPE (2026-08-06, SPEC-soul-openai-tools-v2): tools + the agentic loop now run on
// this wire too (openai_agentic_loop below). Plain completion (openai_chat_complete)
// remains for non-agentic turns. Still ADDITIVE: the Anthropic path is untouched.
fn llm_base_url() -> String {
return env("NEURON_LLM_0_URL")
let u: String = env("NEURON_LLM_0_URL")
if !str_eq(u, "") {
return u
}
// Step-0 bridge (2026-08-06): the shipped launcher exports SOUL_LLM_BASE_URL +
// SOUL_LLM_PROVIDER (installer/macos/neuron-daemons.sh:288-300) and nothing in a
// customer build exports the NEURON_LLM_0_* names so this lane was unreachable
// outside test harnesses. Honor the launcher's names as a fallback. Anthropic
// deliberately returns "" here: its native path stays hardcoded (endpoint
// configurability is neuron#62, out of scope).
let p: String = env("SOUL_LLM_PROVIDER")
if str_eq(p, "") || str_eq(p, "anthropic") {
return ""
}
return env("SOUL_LLM_BASE_URL")
}
fn llm_wire_format() -> String {
let f: String = env("NEURON_LLM_0_FORMAT")
if str_eq(f, "") {
return "anthropic"
if !str_eq(f, "") {
return f
}
return f
// Step-0 bridge (2026-08-06): derive the wire format from the launcher's provider
// name when the explicit format is unset. Every non-Anthropic provider in the app's
// catalog speaks the OpenAI-compatible format (ProviderKeys.kt: llmFormat="openai"
// for openai/grok/gemini/groq/ollama).
let p: String = env("SOUL_LLM_PROVIDER")
if str_eq(p, "openai") || str_eq(p, "grok") || str_eq(p, "gemini") || str_eq(p, "groq") || str_eq(p, "ollama") {
return "openai"
}
return "anthropic"
}
// Escape a decoded string so it can be embedded back into a JSON string literal.
@@ -1853,6 +1883,354 @@ fn openai_chat_complete(model: String, base_url: String, api_key: String, safe_s
return "{\"reply\":\"" + json_escape(content) + "\",\"tools_used\":[]}"
}
// OpenAI-format TOOLS PORT (v2, 2026-08-06, SPEC-soul-openai-tools-v2)
// The agentic loop for OpenAI-compatible providers (Groq/OpenAI/Grok/Gemini/Ollama).
// The tool-execution, consent, bridge and run-progress machinery is the SAME wire-agnostic
// layer agentic_loop uses (dispatch_tool, classify_tool_risk, is_builtin_tool, bridge_save,
// handle_tool_result) only the wire dialect differs. ADR-0005's single-tool constraint is
// mirrored on this wire as parallel_tool_calls:false; a provider that ignores it gets its
// first call honored and the rest dropped LOUDLY. Anthropic's server-side web_search has no
// analogue here, so this lane's tool set comes from agentic_tools_no_web() and "sources"
// is always empty an honest degradation, disclosed in the spec, not a bug.
// Convert an Anthropic-shape tools array ({"name","description","input_schema"}) to the
// OpenAI shape ({"type":"function","function":{"name","description","parameters"}}).
// Entries without an input_schema (Anthropic server tools like web_search) are skipped
// they cannot execute on this wire.
fn openai_tools_json(tools_anthropic: String) -> String {
let out: String = ""
let i: Int = 0
let n: Int = json_array_len(tools_anthropic)
while i < n {
let entry: String = json_array_get(tools_anthropic, i)
let name: String = json_get(entry, "name")
let desc: String = json_get(entry, "description")
let schema: String = json_get_raw(entry, "input_schema")
let keep: Bool = !str_eq(name, "") && !str_eq(schema, "")
let piece: String = if keep {
"{\"type\":\"function\",\"function\":{\"name\":\"" + json_escape(name) + "\""
+ ",\"description\":\"" + json_escape(desc) + "\""
+ ",\"parameters\":" + schema + "}}"
} else { "" }
let out = if keep {
if str_eq(out, "") { piece } else { out + "," + piece }
} else { out }
let i = i + 1
}
return "[" + out + "]"
}
// utf8_safe_slice str_slice with the guarantee that it never splits a character.
//
// str_slice and str_len count BYTES. Every fixed-length content cut in this file
// therefore risks landing inside a multi-byte UTF-8 character and leaving a dangling
// lead byte, which makes the ENTIRE request body invalid UTF-8 providers reject it
// and the user gets an unexplained failure. Found live 2026-08-06 in the session
// preload: a recalled memory containing box-drawing rules (U+2500 = E2 94 80) was cut
// at 350 bytes mid-character, and every turn on that session died. Ordinary content
// triggers it an em dash, a curly quote, an accented name, an emoji and it gets
// MORE likely as a user's memory grows.
//
// Walk back from the cut over UTF-8 continuation bytes (0x80-0xBF) to the lead byte,
// and keep the character only if all of its bytes survived the cut.
fn utf8_safe_slice(s: String, n: Int) -> String {
if str_len(s) <= n { return s }
let cut: String = str_slice(s, 0, n)
let total: Int = str_len(cut)
let i: Int = total - 1
let keep: Int = total
let scanning: Bool = true
let steps: Int = 0
// A UTF-8 character is at most 4 bytes, so at most 4 steps are ever needed.
while scanning && steps < 4 && i >= 0 {
let c: Int = str_char_code(cut, i)
let is_ascii: Bool = c < 128
let is_lead: Bool = c >= 192
// Expected length declared by the lead byte: 0xF0+ = 4, 0xE0+ = 3, else 2.
let need: Int = if c >= 240 { 4 } else { if c >= 224 { 3 } else { 2 } }
let have: Int = total - i
let keep = if is_ascii { total } else {
if is_lead { if have == need { total } else { i } } else { keep }
}
let scanning = if is_ascii || is_lead { false } else { true }
let i = i - 1
let steps = steps + 1
}
return str_slice(cut, 0, keep)
}
// A tool result arrives already json_safe'd from dispatch_tool, so it is embedded into
// the wire message RAW (escaping it a second time is what made the model read literal
// backslashes). But it is also TRUNCATED at a fixed byte count, and a cut can land in the
// middle of an escape pair leaving a dangling backslash that makes the enclosing JSON
// string invalid and 400s the whole turn. Trim any trailing backslash run so the cut is
// always on a clean boundary. (The Anthropic lane truncates the same way and has the same
// latent exposure; not changed here, flagged in the PR.)
fn json_trim_dangling_escape(s: String) -> String {
let out: String = s
while str_ends_with(out, "\\") {
let out = str_slice(out, 0, str_len(out) - 1)
}
return out
}
// The standard agentic tool set WITHOUT Anthropic's native web_search entry: built-ins +
// every connector tool. Same merge as agentic_tools_all(), minus the server-tool tail.
fn agentic_tools_no_web() -> String {
let base: String = agentic_tools_literal()
let conn: String = connector_tools_json()
let base_inner: String = str_slice(base, 1, str_len(base) - 1)
let conn_inner: String = str_slice(conn, 1, str_len(conn) - 1)
let merged: String = if str_eq(conn_inner, "") {
base_inner
} else {
base_inner + "," + conn_inner
}
return "[" + strip_client_web_search(merged) + "]"
}
// openai_agentic_loop the resumable agentic turn on the OpenAI wire. Same two envelopes
// as agentic_loop (done / tool_pending), same client-bridge contract, same state keys.
// [tools_json] arrives ANTHROPIC-shaped (the bridge blob stays wire-uniform); it is
// converted once here. The system prompt travels as the first message (no top-level
// "system" on this wire).
fn openai_agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, tools_log_in: String) -> String {
let api_url: String = llm_base_url() + "/chat/completions"
let api_key: String = agentic_api_key()
let h: Map = {}
map_set(h, "content-type", "application/json")
if !str_eq(api_key, "") {
map_set(h, "Authorization", "Bearer " + api_key)
}
let ask_all: Bool = !str_eq(session_id, "") && str_eq(state_get("require_approval_" + session_id), "true")
let tools_oai: String = openai_tools_json(tools_json)
let has_tools: Bool = json_array_len(tools_oai) > 0
let messages: String = messages_in
let final_text: String = ""
let tools_log: String = tools_log_in
let iteration: Int = 0
let keep_going: Bool = true
// Suspension state top level so it escapes the while body (El scope rule).
let pending: Bool = false
let pend_tool_id: String = ""
let pend_tool_name: String = ""
let pend_tool_input: String = ""
let pend_tool_tier: String = ""
let pend_narration: String = ""
if !str_eq(session_id, "") {
state_set("run_progress_" + session_id, "")
}
while keep_going && iteration < 12 {
let inner_msgs: String = str_slice(messages, 1, str_len(messages) - 1)
let all_msgs: String = if str_eq(inner_msgs, "") {
"[{\"role\":\"system\",\"content\":\"" + safe_sys + "\"}]"
} else {
"[{\"role\":\"system\",\"content\":\"" + safe_sys + "\"}," + inner_msgs + "]"
}
// tools + the ADR-0005 mirror travel only when there are tools to offer: an empty
// tools array is a 400 on real OpenAI-format providers.
let tool_frag: String = if has_tools {
",\"tools\":" + tools_oai + ",\"tool_choice\":\"auto\",\"parallel_tool_calls\":false"
} else { "" }
let req_body: String = "{\"model\":\"" + model + "\""
+ ",\"max_tokens\":16384"
+ tool_frag
+ ",\"messages\":" + all_msgs
+ "}"
let raw_resp: String = http_post_with_headers(api_url, req_body, h)
// OpenAI-format errors arrive as a top-level {"error":{...}} object. Content
// strings inside a valid response are JSON-escaped, so a top-level match cannot
// false-positive on reply text.
let is_error: Bool = str_eq(raw_resp, "") || str_starts_with(raw_resp, "{\"error\"")
if is_error {
let err_head: String = if str_len(raw_resp) > 220 { str_slice(raw_resp, 0, 220) } else { raw_resp }
println("[soul] llm error (openai lane): " + err_head)
return "{\"error\":\"llm unavailable\",\"reply\":\"\"}"
}
let choices: String = json_get_raw(raw_resp, "choices")
let eff_choices: String = if str_eq(choices, "") { "[]" } else { choices }
if json_array_len(eff_choices) < 1 {
// Log the body head, as the error branch does. A provider that answers 200
// with an unexpected shape is otherwise undiagnosable from the log alone.
let noc_head: String = if str_len(raw_resp) > 220 { str_slice(raw_resp, 0, 220) } else { raw_resp }
println("[soul] llm error (openai lane): no choices in response: " + noc_head)
return "{\"error\":\"llm unavailable\",\"reply\":\"\"}"
}
let first: String = json_array_get(eff_choices, 0)
let message_o: String = json_get_raw(first, "message")
let finish: String = json_get(first, "finish_reason")
// Content, read TWO ways on purpose.
// content_raw the provider's own bytes: `null` on a pure tool-call turn, or a
// quoted, already-escaped string. This is what goes back on the wire, verbatim.
// text_out the DECODED text, for narration, the ledger and the final reply.
// json_get decodes, and a JSON null decodes to the 4-char string "null", which is
// never "" so without the raw check a pure tool-call turn produced the literal
// word "null" as the assistant's narration, in the run-progress ledger, and inside
// the tool_pending envelope, and echoed `"content":"null"` instead of `content:null`.
let content_raw: String = json_get_raw(message_o, "content")
let is_null_content: Bool = str_eq(content_raw, "null") || str_eq(content_raw, "")
let text_out: String = if is_null_content { "" } else { json_get(message_o, "content") }
let tc_raw: String = json_get_raw(message_o, "tool_calls")
let tc_arr: String = if str_eq(tc_raw, "") || str_eq(tc_raw, "null") { "[]" } else { tc_raw }
let tc_n: Int = json_array_len(tc_arr)
let has_tool: Bool = tc_n > 0
// ADR-0005 mirror: we ask for one call per round; a provider that returns
// several anyway gets the FIRST honored and the drop logged loudly.
if tc_n > 1 {
println("[soul] DRIFT: provider returned " + int_to_str(tc_n) + " parallel tool_calls despite parallel_tool_calls:false - keeping the first only (ADR-0005 mirror)")
}
// Unknown finish reasons (future API drift): log loudly, never treat an
// unrecognised terminal state as a completed answer silently.
if !str_eq(finish, "stop") && !str_eq(finish, "tool_calls") && !str_eq(finish, "length") && !str_eq(finish, "") {
println("[soul] DRIFT: unknown finish_reason from API: " + finish)
}
let tc0: String = if has_tool { json_array_get(tc_arr, 0) } else { "" }
let tool_id: String = if has_tool { json_get(tc0, "id") } else { "" }
let tc_fn: String = if has_tool { json_get_raw(tc0, "function") } else { "" }
let tool_name: String = if has_tool { json_get(tc_fn, "name") } else { "" }
// arguments is a JSON-ENCODED STRING on this wire; json_get decodes it exactly
// once, yielding the raw object text dispatch_tool expects. Decoding again or
// re-encoding before dispatch is the two-escaper trap the gate's echo-mismatch
// check exists to catch.
let tool_input_raw: String = if has_tool { json_get(tc_fn, "arguments") } else { "" }
let tool_input: String = if str_eq(tool_input_raw, "") { "{}" } else { tool_input_raw }
let is_tool_turn: Bool = has_tool
// Consent policy IDENTICAL to the Anthropic lane: ask_all bridges everything,
// escalate always bridges, non-builtins bridge unless "always allow" granted.
let always_key: String = "always_allow_" + session_id
let always_list: String = if !str_eq(session_id, "") { state_get(always_key) } else { "" }
let is_always_allowed: Bool = !str_eq(tool_name, "") && !str_eq(always_list, "") && str_contains(always_list, tool_name)
let risk_tier: String = if is_tool_turn { classify_tool_risk(tool_name, tool_input) } else { "" }
let needs_bridge: Bool = is_tool_turn && (ask_all || str_eq(risk_tier, "escalate") || (!is_builtin_tool(tool_name) && !is_always_allowed))
let tool_result_raw: String = if is_tool_turn && !needs_bridge { dispatch_tool(tool_name, tool_input) } else { "" }
let tool_result: String = if str_len(tool_result_raw) > 6000 {
json_trim_dangling_escape(str_slice(tool_result_raw, 0, 6000)) + "...[truncated]"
} else { tool_result_raw }
let tool_quoted: String = "\"" + tool_name + "\""
let tools_log = if is_tool_turn {
if str_eq(tools_log, "") { tool_quoted } else { tools_log + "," + tool_quoted }
} else { tools_log }
// The assistant turn echoed with its tool_calls array VERBATIM (raw), so the
// tool_call_id pairing stays valid on the wire and across a bridge resume.
// Echo the provider's content BYTES, never a re-escaped round-trip. Decoding and
// re-encoding is where fidelity is lost: json_escape/json_safe both handle only
// \\ " \n \r, so any other control character the model emits (a tab, say) would go
// back out raw and make the next request body invalid JSON — a provider 400 that
// looks like a random failure. The Anthropic lane never had this exposure because
// it echoes the response's content array untouched; this now matches it.
let content_frag: String = if is_null_content { "null" } else { content_raw }
// Echo ONLY the call we actually honor — never the provider's full array.
// The loop can assemble exactly one tool response per round, so replaying N
// tool_calls while answering one leaves the conversation self-contradictory and
// every OpenAI-format provider 400s on the next request ("no tool response for
// id X"). That is precisely the failure ADR-0005 documents on the Anthropic wire,
// where the block walk keeps the first tool_use and the rest die without a
// tool_result. Here it costs one slice to close: the dropped calls simply never
// existed from the model's point of view, and the DRIFT line above keeps the
// accounting honest about what we discarded.
let assist_turn: String = if has_tool {
"{\"role\":\"assistant\",\"content\":" + content_frag + ",\"tool_calls\":[" + tc0 + "]}"
} else {
"{\"role\":\"assistant\",\"content\":" + content_frag + "}"
}
let inner_now: String = str_slice(messages, 1, str_len(messages) - 1)
let messages_with_assistant: String = "[" + inner_now + "," + assist_turn + "]"
// Local built-in tool turn: append assistant echo + role:"tool" result, loop on.
let local_continue: Bool = is_tool_turn && !needs_bridge
let messages = if local_continue {
let inner2: String = str_slice(messages_with_assistant, 1, str_len(messages_with_assistant) - 1)
"[" + inner2 + ",{\"role\":\"tool\",\"tool_call_id\":\"" + tool_id + "\",\"content\":\"" + tool_result + "\"}]"
} else { messages }
// Live run-progress ledger — same key, same shape, same poller as the Anthropic
// lane; a forked loop that omitted this would silently kill live step rendering.
if !str_eq(session_id, "") {
let prog_key: String = "run_progress_" + session_id
let prog_prev: String = state_get(prog_key)
let prog_snip: String = if str_len(text_out) > 280 { str_slice(text_out, 0, 280) } else { text_out }
let prog_entry: String = "{\"i\":" + int_to_str(iteration)
+ ",\"t\":\"" + json_safe(prog_snip) + "\""
+ ",\"tool\":\"" + json_safe(tool_name) + "\"}"
let prog_next: String = if str_eq(prog_prev, "") { prog_entry } else { prog_prev + "," + prog_entry }
state_set(prog_key, prog_next)
}
// Bridge turn: persist the continuation (wire-tagged) and stop the loop.
let pending = if needs_bridge { true } else { pending }
let pend_tool_id = if needs_bridge { tool_id } else { pend_tool_id }
let pend_tool_name = if needs_bridge { tool_name } else { pend_tool_name }
let pend_tool_input = if needs_bridge { tool_input } else { pend_tool_input }
let pend_tool_tier = if needs_bridge { risk_tier } else { pend_tool_tier }
let pend_narration = if needs_bridge { text_out } else { pend_narration }
if needs_bridge {
bridge_save(session_id, model, safe_sys, tools_json, messages_with_assistant, tools_log, tool_id, "openai")
}
// Text accumulation: rounds are separated by tool executions, so the resume seam
// is unconditionally a boundary (same rule as the Anthropic loop's seam 2).
let final_text = if !is_tool_turn {
final_text + text_join_sep(final_text, text_out, true) + text_out
} else { final_text }
// Output cap hit mid-action (finish_reason "length" with a tool call pending).
let final_text = if str_eq(finish, "length") && has_tool {
final_text + "\n\n[Output limit reached mid-action - the last planned action did not run. Ask me to continue to finish it.]"
} else { final_text }
let keep_going = if local_continue { keep_going } else { false }
let iteration = iteration + 1
}
if pending {
let safe_in: String = if str_eq(pend_tool_input, "") { "{}" } else { pend_tool_input }
let tools_arr: String = if str_eq(tools_log, "") { "[]" } else { "[" + tools_log + "]" }
return "{\"tool_pending\":true"
+ ",\"session_id\":\"" + session_id + "\""
+ ",\"call_id\":\"" + pend_tool_id + "\""
+ ",\"tool_name\":\"" + pend_tool_name + "\""
+ ",\"tool_input\":" + safe_in
+ ",\"risk_tier\":\"" + pend_tool_tier + "\""
+ ",\"narration\":\"" + json_safe(pend_narration) + "\""
+ ",\"model\":\"" + model + "\""
+ ",\"agentic\":true"
+ ",\"sources\":\"\""
+ ",\"tools_used\":" + tools_arr + "}"
}
let final_text = receipt_strip(final_text)
if str_eq(final_text, "") {
let hit_cap: Bool = iteration >= 12
let err_msg: String = if hit_cap {
"agentic loop hit the 12-iteration cap without producing a final reply - task may be too complex or a tool call is looping"
} else {
"no response"
}
return "{\"error\":\"" + err_msg + "\",\"reply\":\"\",\"iterations\":" + int_to_str(iteration) + "}"
}
let safe_text: String = json_safe(final_text)
let tools_arr: String = if str_eq(tools_log, "") { "[]" } else { "[" + tools_log + "]" }
if !str_eq(session_id, "") {
let done_key: String = "run_progress_" + session_id
let done_prev: String = state_get(done_key)
let done_next: String = if str_eq(done_prev, "") { "{\"done\":true}" } else { done_prev + ",{\"done\":true}" }
state_set(done_key, done_next)
}
return "{\"reply\":\"" + safe_text + "\",\"model\":\"" + model + "\",\"agentic\":true,\"tools_used\":" + tools_arr + ",\"sources\":\"\",\"iterations\":" + int_to_str(iteration) + "}"
}
fn agentic_tools_literal() -> String {
return "[" +
"{\"name\":\"read_file\",\"description\":\"Read contents of a file from disk.\",\"input_schema\":{\"type\":\"object\",\"properties\":{\"path\":{\"type\":\"string\",\"description\":\"Absolute file path\"}},\"required\":[\"path\"]}}," +
@@ -2636,7 +3014,7 @@ fn handle_chat_agentic(body: String) -> String {
let ag_continuity_snip: String = if ag_continuity_ok {
let acn0: String = json_array_get(ag_continuity_nodes, 0)
let acc: String = json_get(acn0, "content")
if str_len(acc) > 350 { str_slice(acc, 0, 350) } else { acc }
utf8_safe_slice(acc, 350)
} else { "" }
let ag_profile_bullets: String = session_preload_bullets(ag_profile_nodes2, 8, 350)
let ag_work_bullets: String = session_preload_bullets(ag_work_nodes2, 6, 350)
@@ -2667,7 +3045,13 @@ fn handle_chat_agentic(body: String) -> String {
" + ctx + ag_session_preload + receipt_rule()
let api_key: String = agentic_api_key()
let tools_json: String = agentic_tools_all()
// Assemble the tool set ONCE, for the lane this turn will actually take. Both
// builders call connector_tools_json(), which is an HTTP round-trip to the
// connectors bridge on :7771 computing both would pay that cost, and its timeout
// exposure, twice per turn. The OpenAI lane drops Anthropic's server-side
// web_search (it has no analogue on that wire and cannot execute there).
let tools_lane_openai: Bool = !str_eq(llm_base_url(), "") && str_eq(llm_wire_format(), "openai")
let tools_json: String = if tools_lane_openai { agentic_tools_no_web() } else { agentic_tools_all() }
let safe_msg: String = json_safe(message)
let safe_sys: String = json_safe(system)
@@ -2708,11 +3092,12 @@ fn handle_chat_agentic(body: String) -> String {
// for the rest of the run. Absent/false = behavior identical to before this fix.
let req_ask_all: String = json_get(body, "require_approval")
state_set("require_approval_" + session_id, if str_eq(req_ask_all, "true") { "true" } else { "" })
// Provider fork: OpenAI-compatible providers (Ollama/OpenAI/Grok/Gemini) take the plain-completion
// path (v1, no tools); everything else stays on the Anthropic agentic loop (the default).
let use_openai: Bool = !str_eq(llm_base_url(), "") && str_eq(llm_wire_format(), "openai")
// Provider fork (v2 port, 2026-08-06): OpenAI-compatible providers now take their own
// AGENTIC loop same tools (minus Anthropic-server web_search), same consent policy,
// same bridge contract. The Anthropic native path stays the default and is untouched.
let use_openai: Bool = tools_lane_openai
let result: String = if use_openai {
openai_chat_complete(model, llm_base_url(), agentic_api_key(), safe_sys, messages)
openai_agentic_loop(session_id, model, safe_sys, tools_json, messages, "")
} else {
agentic_loop(session_id, model, safe_sys, tools_json, messages, h, "")
}
@@ -3139,7 +3524,7 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
// client's tool_result block. messages_with_assistant is only meaningful when a
// tool was requested, so guard on needs_bridge before persisting.
if needs_bridge {
bridge_save(session_id, model, safe_sys, tools_json, messages_with_assistant, tools_log, pend_tool_id)
bridge_save(session_id, model, safe_sys, tools_json, messages_with_assistant, tools_log, pend_tool_id, "anthropic")
}
// ACCUMULATE across pause/resume cycles instead of overwriting. A resumed turn
@@ -3221,7 +3606,7 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
// single JSON blob in soul state so agentic_resume can rebuild the exact loop. The
// stored `messages` already includes the assistant turn that requested the tool, so
// resume just appends the client's tool_result for `tool_use_id`.
fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String) -> Bool {
fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String, wire: String) -> Bool {
// Guard: empty messages or tools_json would produce syntactically invalid JSON.
// Return false so the caller detects the failure rather than writing a corrupt
// blob that agentic_resume would later resume with no context.
@@ -3251,10 +3636,13 @@ fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json:
// messages_raw arbitrary model/user content so neither raw extraction can
// first-match into model-controlled bytes either. Do not reorder; do not add a
// field after messages_raw.
// "wire" (v2 port, 2026-08-06) is a json_safe'd SCALAR and therefore sits with the
// other scalars BEFORE both raw fields, per the field-order rule above.
let blob: String = "{\"model\":\"" + json_safe(model) + "\""
+ ",\"safe_sys\":\"" + json_safe(safe_sys) + "\""
+ ",\"tools_log\":\"" + json_safe(tools_log) + "\""
+ ",\"tool_use_id\":\"" + json_safe(tool_use_id) + "\""
+ ",\"wire\":\"" + json_safe(wire) + "\""
+ ",\"tools_raw\":" + tools_json
+ ",\"messages_raw\":" + messages + "}"
state_set("mcp_bridge:" + session_id, blob)
@@ -3310,14 +3698,52 @@ fn agentic_resume(session_id: String, tool_use_id: String, content: String) -> S
str_slice(content, 0, 6000) + "...[truncated]"
} else { content }
let safe_result: String = json_safe(trimmed)
let tool_msg: String = "{\"type\":\"tool_result\",\"tool_use_id\":\"" + eff_use_id + "\",\"content\":\"" + safe_result + "\"}"
let inner: String = str_slice(messages, 1, str_len(messages) - 1)
let resumed_messages: String = "[" + inner + ",{\"role\":\"user\",\"content\":[" + tool_msg + "]}]"
// One-shot: clear the saved turn so a session_id can't be replayed.
state_set("mcp_bridge:" + session_id, "")
// Wire-aware resume (v2 port, 2026-08-06): blobs written since the port carry a
// "wire" scalar ("anthropic" | "openai") among the scalar fields, where first-match
// scanning is safe (see bridge_save's field-order rule). A blob with no wire field
// is a legacy pre-port suspension always Anthropic. On the OpenAI wire the
// client's result goes back as a role:"tool" turn keyed by tool_call_id, and a
// result for an already-answered id must never be re-sent (the resumed messages
// end at the assistant echo, so appending exactly one tool turn preserves that).
// Read "wire" from the blob's SCALAR HEAD ONLY never the whole blob.
//
// json_get is a first-substring-match scanner. On a blob written by this binary the
// scalar sits ahead of the raw fields and wins, but on a LEGACY blob (suspended
// before this field existed) there is no match up front, so the scan runs on into
// messages_raw model- and user-controlled bytes. A conversation that merely
// CONTAINS the literal "wire":"openai" would then misroute the resume onto the wrong
// loop and kill the run. That is exactly the round-9 defect (json_get(blob,
// "tool_use_id") matching a web_search_tool_result id inside the replayed
// conversation), and the fix is the same shape: bound the search.
//
// bridge_save guarantees every json_safe'd scalar precedes the bulk fields, so
// truncating at the earliest bulk key makes this deterministic the decoy is not
// even inside the string we search. Both the current keys (tools_raw/messages_raw)
// and the pre-round-9 legacy ones (tools_json/messages) are covered.
let i_traw: Int = str_index_of(blob, ",\"tools_raw\":")
let i_tjson: Int = str_index_of(blob, ",\"tools_json\":")
let i_mraw: Int = str_index_of(blob, ",\"messages_raw\":")
let i_msgs: Int = str_index_of(blob, ",\"messages\":")
let cut1: Int = if i_traw > 0 { i_traw } else { str_len(blob) }
let cut2: Int = if i_tjson > 0 && i_tjson < cut1 { i_tjson } else { cut1 }
let cut3: Int = if i_mraw > 0 && i_mraw < cut2 { i_mraw } else { cut2 }
let cut: Int = if i_msgs > 0 && i_msgs < cut3 { i_msgs } else { cut3 }
let blob_head: String = str_slice(blob, 0, cut)
let wire: String = json_get(blob_head, "wire")
if str_eq(wire, "openai") {
let tool_msg_o: String = "{\"role\":\"tool\",\"tool_call_id\":\"" + eff_use_id + "\",\"content\":\"" + safe_result + "\"}"
let resumed_o: String = "[" + inner + "," + tool_msg_o + "]"
return openai_agentic_loop(session_id, model, safe_sys, tools_json, resumed_o, tools_log)
}
let tool_msg: String = "{\"type\":\"tool_result\",\"tool_use_id\":\"" + eff_use_id + "\",\"content\":\"" + safe_result + "\"}"
let resumed_messages: String = "[" + inner + ",{\"role\":\"user\",\"content\":[" + tool_msg + "]}]"
let api_key: String = agentic_api_key()
let h: Map = {}
map_set(h, "x-api-key", api_key)
@@ -3461,7 +3887,7 @@ fn handle_dharma_room_turn(body: String) -> String {
// engram_node(content, "episodic", ...) which wrongly put a TIER into the node_type
// slot that's why nodes showed node_type="episodic". Use the full, correct contract.)
let utterance_tags: String = "[\"soul-utterance\",\"episodic\"]"
let discard_id: String = engram_node_full(
let discard_id: String = wt_node(
clean_response, "Conversation", "soul:utterance",
el_from_float(0.6), el_from_float(0.6), el_from_float(0.8),
"Episodic", utterance_tags
@@ -3493,7 +3919,11 @@ fn handle_dharma_room_turn_agentic(body: String) -> String {
// Hard Bell: pre-LLM safety evaluation on agentic dharma room turns.
let system = safety_augment_system(system, transcript)
let tools_json: String = agentic_tools_all()
// One assembly, for the lane this turn takes (see the same note in handle_chat_agentic:
// both builders hit the connectors bridge over HTTP, so computing both doubles the cost
// and the timeout exposure).
let use_openai_d: Bool = !str_eq(llm_base_url(), "") && str_eq(llm_wire_format(), "openai")
let tools_json: String = if use_openai_d { agentic_tools_no_web() } else { agentic_tools_all() }
let safe_transcript: String = json_safe(transcript)
let safe_sys: String = json_safe(system)
let messages: String = "[{\"role\":\"user\",\"content\":\"" + safe_transcript + "\"}]"
@@ -3504,7 +3934,14 @@ fn handle_dharma_room_turn_agentic(body: String) -> String {
// Use dharma-prefixed session_id so bridge suspension works correctly per room.
let session_id: String = if str_eq(room_id, "") { "dharma:" + next_bridge_id() } else { "dharma:" + room_id }
let loop_result: String = agentic_loop(session_id, model, safe_sys, tools_json, messages, h, "")
// Provider fork (v2 port, 2026-08-06): same routing rule as handle_chat_agentic.
// The Hard Bell augmentation above is baked into safe_sys BEFORE the fork, so the
// safety pass is identical on both wires.
let loop_result: String = if use_openai_d {
openai_agentic_loop(session_id, model, safe_sys, tools_json, messages, "")
} else {
agentic_loop(session_id, model, safe_sys, tools_json, messages, h, "")
}
let result_error: String = json_get(loop_result, "error")
if !str_eq(result_error, "") {
@@ -3552,7 +3989,7 @@ fn session_summary_write(summary_text: String) -> String {
}
}
let tags: String = "[\"SessionSummary\",\"session-summary\",\"previous-session\",\"consolidate\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "SessionSummary", "session:summary",
el_from_float(0.85), el_from_float(0.85), el_from_float(1.0),
"Episodic", tags
@@ -3578,7 +4015,7 @@ fn session_summary_write_dated(summary_text: String, label: String) -> String {
let ts_str: String = int_to_str(ts)
let content: String = "[session-summary] " + trimmed + " | ts:" + ts_str
let tags: String = "[\"SessionSummary\",\"session-summary\",\"previous-session\",\"consolidate\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "SessionSummary", label,
el_from_float(0.9), el_from_float(0.8), el_from_float(1.0),
"Episodic", tags
@@ -3654,7 +4091,7 @@ fn auto_persist(req: String, resp: String) -> Void {
+ ",\"bell\":\"" + bell_level + "\""
+ ",\"label\":\"chat:" + ts_str + "\"}"
let conv_node_id: String = engram_node_full(
let conv_node_id: String = wt_node(
content,
"Conversation",
"chat:" + ts_str,
@@ -3692,7 +4129,7 @@ fn auto_persist(req: String, resp: String) -> Void {
let bell_tags: String = "[\"safety\",\"bell\",\"bell:" + bell_level + "\",\"affective\",\"BellEvent\"]"
let bell_ts_str: String = int_to_str(time_now())
let bell_label: String = "bell:" + bell_level + ":" + bell_ts_str
let bell_node_id: String = engram_node_full(
let bell_node_id: String = wt_node(
bell_content,
"BellEvent",
bell_label,
@@ -3751,7 +4188,7 @@ fn auto_persist(req: String, resp: String) -> Void {
let pos_tags: String = "[\"joy\",\"positive\",\"joy:" + positive_level + "\",\"affective\",\"PositiveEvent\"]"
let pos_ts_label: String = int_to_str(time_now())
let pos_label: String = "joy:" + positive_level + ":" + pos_ts_label
let pos_node_id: String = engram_node_full(
let pos_node_id: String = wt_node(
pos_content, "PositiveEvent", pos_label,
pos_sal_a, pos_sal_b, pos_sal_c, "Episodic", pos_tags
)
+6 -1
View File
@@ -53,6 +53,11 @@ extern fn llm_base_url() -> String
extern fn llm_wire_format() -> String
extern fn json_escape(s: String) -> String
extern fn openai_chat_complete(model: String, base_url: String, api_key: String, safe_sys: String, messages_json: String) -> String
extern fn openai_tools_json(tools_anthropic: String) -> String
extern fn utf8_safe_slice(s: String, n: Int) -> String
extern fn json_trim_dangling_escape(s: String) -> String
extern fn agentic_tools_no_web() -> String
extern fn openai_agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, tools_log_in: String) -> String
extern fn agentic_tools_literal() -> String
extern fn web_search_tool_json() -> String
extern fn strip_client_web_search(tools_inner: String) -> String
@@ -75,7 +80,7 @@ extern fn next_bridge_id() -> String
extern fn handle_chat_plan(body: String) -> String
extern fn handle_chat_agentic(body: String) -> String
extern fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, h: Map, tools_log_in: String) -> String
extern fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String) -> Bool
extern fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String, wire: String) -> Bool
extern fn agentic_resume(session_id: String, tool_use_id: String, content: String) -> String
extern fn handle_tool_result(session_id: String, body: String) -> String
extern fn handle_chat_as_soul(body: String) -> String
Generated Vendored
+6 -1
View File
@@ -141,7 +141,7 @@ el_val_t awareness_run(void);
el_val_t axon_get(el_val_t path);
el_val_t axon_post(el_val_t path, el_val_t body);
el_val_t bounded_persona_floor(void);
el_val_t bridge_save(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages, el_val_t tools_log, el_val_t tool_use_id);
el_val_t bridge_save(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages, el_val_t tools_log, el_val_t tool_use_id, el_val_t wire);
el_val_t build_form_from_json(el_val_t semantic_form_json, el_val_t lang_code);
el_val_t build_np(el_val_t referent, el_val_t slots);
el_val_t build_pp(el_val_t loc);
@@ -875,6 +875,11 @@ el_val_t non_weak_past(el_val_t stem, el_val_t slot);
el_val_t non_weak_present(el_val_t stem, el_val_t slot);
el_val_t one_cycle(void);
el_val_t openai_chat_complete(el_val_t model, el_val_t base_url, el_val_t api_key, el_val_t safe_sys, el_val_t messages_json);
el_val_t openai_tools_json(el_val_t tools_anthropic);
el_val_t json_trim_dangling_escape(el_val_t s);
el_val_t utf8_safe_slice(el_val_t s, el_val_t n);
el_val_t agentic_tools_no_web(void);
el_val_t openai_agentic_loop(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages_in, el_val_t tools_log_in);
el_val_t parse_float_x100(el_val_t s);
el_val_t path_within_root(el_val_t path, el_val_t root);
el_val_t peo_ah_past(el_val_t slot);
Generated Vendored
+1830 -678
View File
File diff suppressed because one or more lines are too long
Generated Vendored
+18
View File
@@ -0,0 +1,18 @@
# soul.c.stamp — fingerprint of the .el sources dist/soul.c was generated from.
# Written by tools/soulc-stamp.sh --write. Do not hand-edit.
# generated_amalgam_sha256 562193a341876f8fa0618ae7a76d0ecc073bb4d1d21b128c48b4aee28e1ef0d9
# generated_amalgam_bytes 1226240
f8597e10546654bce3fbbe40461b2da59d0e06dbf1b038d1d362d24f949e3911 awareness.el
2ff2dada732918c788a9ef66c6fd54c7a24cc4bbd4829197fe945d3a75ca1929 chat.el
42288c212cbf72fb1e8ecbd4d9900e4e9ee1cfa475b7974295c7637f1bf2939f elp-input.el
b3f77f49d6086932c38bd17fe7a5eaf8bce25685f6fc3e1750f05729c6b49b9e imprint.el
fba8ffdb9ba72bca5b09ca1c93a520edc52f3f4d8aec2c7585fe9b17e06420b2 manifest.el
550a72e234ae8cec1f33e02108fd365353f45edd88513da90b792e79b6c0e5f0 memory.el
5ec07ec9785b02abe32f3ff7acf2d1f9f7e07c0967fac97e6eff17d7110b5c84 neuron-api.el
03c47c451e0e87f2c252cadb4b765867943962a804f548dd53adeef0520912c8 persist.el
a6d69f3fc55233d9d3300160fd46a1551f2064bcd0fb84e2c9e432f636a72476 routes.el
c28e36952ec56525963a0bdf29455ab097d3b0c5653d19c25fbb005e1069a1f7 safety.el
fd3ab91d0ae0ea26639e21bef2f8f94054dc4b02eae68b19e3fe689d2769aad4 sessions.el
0f1cf43904a98a5a646cce5a07e0e96162ced662692fbc13357d9b67d9a8ac3d soul.el
30337940905171a9645b0929f0a412ce6b3dccb1246495070c553bca0bbae6cd stewardship.el
e105dc5990e6adbf39db9dc0462cd8bcf6e6c3dfd03709059227ecfad2bbab29 studio.el
+11 -2
View File
@@ -110,7 +110,7 @@ tool("beginSession", "Initialize session: surface recent high-importance memorie
"," + tool("linkCausal", "Create a causal edge (cause -> effect).") +
"," + tool("restructureCausalGraph", "Re-balance the causal subgraph after new evidence.") +
"," + tool("rebuildGraph", "Rebuild graph indices from the on-disk snapshot.") +
"," + tool("runStructuralAudit", "Audit graph structure for orphans, dangling edges, mislabeled types.") +
"," + tool("runStructuralAudit", "Stage 1 structural audit: owner-vs-runtime divergence, orphans and dangling edges, typed-edge distribution, self-model connectivity. Returns an annotated characterization, not a score.") +
// Backlog + work
"," + tool("planWork", "Create a backlog item.") +
"," + tool("reviewBacklog", "Browse work items.") +
@@ -680,7 +680,16 @@ fn dispatch_tool_call(tool_name: String, args: String) -> String {
return mcp_json_result(resp)
}
if str_eq(tool_name, "runStructuralAudit") {
let resp: String = http_get(neuron_url() + "/session/begin")
// Was: GET /session/begin an unrelated session digest returned under an
// audit tool name, i.e. the tool advertised a check that did not exist.
// Now points at the real Stage 1 route (neuron-api.el
// handle_api_structural_audit). Sample caps ride the query string; the
// defaults keep a manual audit to a couple of seconds.
let e_s: Int = json_get_int(args, "edge_sample")
let n_s: Int = json_get_int(args, "node_sample")
let qs: String = "?edge_sample=" + int_to_str(if e_s > 0 { e_s } else { 3000 })
+ "&node_sample=" + int_to_str(if n_s > 0 { n_s } else { 300 })
let resp: String = http_get(neuron_url() + "/audit/structural" + qs)
return mcp_json_result(resp)
}
+104 -9
View File
@@ -1,9 +1,91 @@
import "persist.el"
fn tier_working() -> String { return "Working" }
fn tier_episodic() -> String { return "Episodic" }
fn tier_canonical() -> String { return "Canonical" }
// Association on write
// DESIGN: "promotion integrates candidate nodes by linking them to existing nodes
// using typed semantic edges RATHER THAN APPENDING AS UNLINKED CONTENT" (CCR
// claim 29). Unlinked append is the explicitly rejected behaviour and it is the
// only behaviour this system had. Measured 2026-08-09 on Tim's graph: 14,214 edges
// across 80,936 nodes, 5% of nodes connected to anything, and NO edge created by
// any write since 2026-07-19 while 27,000+ nodes were added. A memory that forms
// no connections cannot be reached by spreading activation, so retrieval silently
// degrades to literal matching.
//
// BOUNDS, each one bought with a specific failure:
// * max 3 edges per memory link_memories.py's cap, precision over spray
// * never link to identity (self/*, Value): the existing policy is explicit that
// "memories must not pollute the self traversal by similarity; only an explicit
// citation may touch identity". Similarity is not citation.
// * never link telemetry (state-event, soul-response, boot_count, loop-outcome):
// these are ~97% of daily write volume (1,020 vs 31 real memories on 08-08).
// Linking them would add ~3,000 noise edges a day and re-flatten the graph in
// the name of connecting it.
// * fail-soft: a failed association never fails the write.
// Edges go through wt_edge so they reach the owner and survive restart.
fn mem_assoc_skip_label(label: String) -> Bool {
if str_contains(label, "state-event") { return true }
if str_contains(label, "soul-response") { return true }
if str_contains(label, "soul-outbox") { return true }
if str_contains(label, "boot_count") { return true }
if str_contains(label, "loop-outcome") { return true }
if str_contains(label, "search-result") { return true }
return false
}
// A candidate is linkable only if it is a real, distinct, non-identity node.
fn mem_assoc_ok(cand_id: String, cand_label: String, self_id: String) -> Bool {
if str_eq(cand_id, "") { return false }
if str_eq(cand_id, self_id) { return false }
// CASE MATTERS measured 2026-08-09. A lowercase-only check let a memory link
// to "Self — Values (grounded)", i.e. it polluted the self traversal, which is
// the one thing this policy exists to prevent. My verification had the same
// blind spot and printed PASS. Check every casing the graph actually uses, and
// exclude identity node TYPES as well as labels.
let lab: String = str_lower(cand_label)
if str_starts_with(lab, "self") { return false }
if str_starts_with(lab, "value") { return false }
if str_contains(lab, "values") { return false }
if str_contains(lab, "identity") { return false }
if mem_assoc_skip_label(cand_label) { return false }
return true
}
// One slot of the association. Manual unroll rather than a loop: EL's codegen
// mis-emits accumulating while-loops (documented at soul.el:212, which unrolled
// three affective slots for the same reason).
fn mem_assoc_slot(results: String, idx: Int, new_id: String) -> Void {
if idx >= json_array_len(results) { return }
let cand: String = json_array_get(results, idx)
let cid: String = json_get(cand, "id")
let clabel: String = json_get(cand, "label")
let ctype: String = json_get(cand, "node_type")
if str_eq(ctype, "Value") { return }
if str_eq(ctype, "DharmaSelf") { return }
if str_eq(ctype, "Safety") { return }
if mem_assoc_ok(cid, clabel, new_id) {
wt_edge(new_id, cid, el_from_float(0.5), "related")
}
}
// mem_associate connect a freshly written memory to what it is about.
fn mem_associate(new_id: String, content: String, label: String) -> Void {
if str_eq(new_id, "") { return }
if mem_assoc_skip_label(label) { return }
// Ask the graph what this memory resembles. Now that the store carries
// meaning-vectors this is semantic, not merely lexical.
let probe: String = str_slice(content, 0, 400)
let results: String = engram_recall_json(probe, 4)
if str_eq(results, "") { return }
mem_assoc_slot(results, 0, new_id)
mem_assoc_slot(results, 1, new_id)
mem_assoc_slot(results, 2, new_id)
}
fn mem_store(content: String, label: String, tags: String) -> String {
let id: String = engram_node_full(
let id: String = wt_node(
content,
"Memory",
label,
@@ -17,13 +99,26 @@ fn mem_store(content: String, label: String, tags: String) -> String {
println("[memory] write rejected by engram (empty id): label=" + label)
return ""
}
// Read back to verify the node actually persisted guards against silent write failures.
let readback: String = engram_get_node_json(id)
if str_eq(readback, "") || str_eq(readback, "{}") {
println("[memory] WRITE VERIFY FAILED: label=" + label + " id=" + id + " — node absent after write")
return ""
// wt_node has already read the node back locally and returns "" if it did
// not land, so the old duplicate read-back here is gone.
//
// HONESTY (neuron#117): the receipt now says WHERE the write is.
// The old unconditional "write verified" line asserted against the soul's
// own RAM true in memory, false on disk and printed ~115,000 times on
// Tim's machine while the canonical snapshot sat frozen for three days.
// wt_commit flushes the spool and then asks the OWNER. When it says false
// the node is real and recallable but not yet durable, and the log says so
// rather than claiming a save that did not happen. The id is still returned:
// the local write DID succeed, and the queued delta will be retried.
let durable: Bool = wt_commit(id)
// Associate AFTER the node is durable: an edge to a node that did not persist
// is a dangling edge, which is the defect the 2026-08-09 cleanup removed 830 of.
mem_associate(id, content, label)
if durable {
println("[memory] write persisted at owner: " + id + " label=" + label)
} else {
println("[memory] write IN MEMORY ONLY (queued for owner, not yet durable): " + id + " label=" + label)
}
println("[memory] write verified: " + id + " ok")
return id
}
@@ -51,12 +146,12 @@ fn mem_strengthen(node_id: String) -> Void {
// memory.el (imported first) so awareness.el and neuron-api.el can both call it.
fn mem_tombstone(node_id: String) -> String {
let tags: String = "[\"Tombstone\",\"status:deleted\"]"
let marker: String = engram_node_full(
let marker: String = wt_node(
node_id, "Tombstone", "tombstone:" + node_id,
el_from_float(0.01), el_from_float(0.01), el_from_float(1.0),
"Episodic", tags)
if !str_eq(marker, "") {
engram_connect(marker, node_id, el_from_float(1.0), "tombstones")
wt_edge(marker, node_id, el_from_float(1.0), "tombstones")
}
return marker
}
+534 -29
View File
@@ -195,15 +195,24 @@ fn api_compact_activated(raw: String, max_items: Int, snip: Int) -> String {
}
// api_persisted read-back-after-write guard against hallucinated saves.
// After a write builtin returns an id, confirm the node is actually queryable
// via engram_get_node_json(id) (returns "" or "null" when missing). Returns
// true only when the node is genuinely persisted.
//
// WIDENED FOR neuron#117. This function is the single gate every MCP write
// handler passes through before it reports success (10 call sites), which makes
// it the right place to close the honesty gap rather than editing ten receipts.
//
// It used to read back from engram_get_node_json the SOUL'S OWN in-process
// graph. In HTTP-engram mode that asserts the wrong thing: the soul is not the
// persistence owner, so a node present in its RAM and absent from the owner read
// as "persisted" and then vanished on the next restart. The guard was doing
// exactly what its comment promised and still certifying writes that did not
// survive. It now flushes the write-through spool and asks the OWNER.
//
// In file mode (no ENGRAM_URL) the soul IS the owner and wt_commit collapses to
// the original local read-back unchanged behaviour, which is what keeps this
// reversible.
fn api_persisted(id: String) -> Bool {
if str_eq(id, "") { return false }
let node: String = engram_get_node_json(id)
// engram_get_node_json returns "{}" (empty object) when node is not found not "" or "null".
// Check all three to guard against any runtime variation.
return !str_eq(node, "") && !str_eq(node, "null") && !str_eq(node, "{}")
return wt_commit(id)
}
// api_not_persisted standard error for a write that did not read back.
@@ -295,10 +304,35 @@ fn handle_api_begin_session(body: String) -> String {
let state_events: String = api_compact_node_array(state_events_raw, 5, 500)
let recent_raw: String = engram_scan_nodes_json(10, 0)
let recent: String = api_compact_node_array(recent_raw, 10, 240)
// SELF-SEEDED SLICE (2026-08-09). The design is explicit: "Every compilation
// query begins at the self-model node and traverses outward... structural
// reachability from the self-model node is a precondition for any node to
// appear in compiled context" (will-anderson patents/drafts/engram-claims.md,
// Self-Seeded Activation; DRAFT, not a filed provisional — cite it as such).
//
// Measured 2026-08-09 before this change: compiled context contained 0-1
// identity records out of 10, because compilation seeds from a hardcoded
// TEXT STRING, never from the self. Even an explicit "my values identity who
// I am" query returned a boot counter and state-events.
//
// This restores the designed behaviour WITHOUT repeating the failure that got
// self_neighbors set to [] in the first place: that was an UNBOUNDED ~90KB
// neighbour dump which closed the socket on every call. Same bound as every
// other list here cap 8, 240-char snippets. The self root has 34 direct
// neighbours of which 23 are identity records, so depth 1 is dense enough to
// be worth seeding and small enough to stay cheap.
let self_raw: String = engram_neighbors_json("kn-efeb4a5b-5aff-4759-8a97-7233099be6ee", 1, "both")
// Cap 24, not 8: measured 2026-08-09, the self root's first 8 neighbours are
// TAG nodes ("neuron", "tier:note", "disposition:experimental", "imprint",
// "traversal") which crowd out the substantive identity records behind them.
// The root has 34 neighbours of which 23 are identity; 24 captures them while
// staying bounded. Cost measured at ~+4KB on a ~12KB response, nowhere near
// the ~90KB unbounded dump that closed sockets and got this set to [].
let self_slice: String = api_compact_node_array(self_raw, 24, 240)
return "{\"stats\":" + stats
+ ",\"recent\":" + recent
+ ",\"activated\":" + activated
+ ",\"self_neighbors\":[]"
+ ",\"self_neighbors\":" + self_slice
+ ",\"recent_state_events\":" + state_events + "}"
}
@@ -342,10 +376,17 @@ fn handle_api_remember(body: String) -> String {
let inner: String = str_slice(base_tags, 1, str_len(base_tags) - 1)
"[" + inner + ",\"project:" + project + "\"]"
}
let id: String = engram_node_full(content, "Memory", "memory:remembered",
let id: String = wt_node(content, "Memory", "memory:remembered",
sal, sal, el_from_float(0.9),
"Episodic", final_tags)
if !api_persisted(id) { return api_not_persisted(id) }
// Associate on write (2026-08-09). THIS CALL MUST BE HERE, not only in mem_store.
// The HTTP memory route writes via wt_node directly; mem_store serves only the
// awareness paths (soul-response, search-result, activation-result) which are
// exactly the telemetry we refuse to link. Hooking mem_store alone produced
// ZERO edges across four real writes measured, not assumed, which is the only
// reason it was caught before shipping.
mem_associate(id, content, "memory:remembered")
return "{\"id\":\"" + id + "\",\"ok\":true}"
}
@@ -369,7 +410,7 @@ fn handle_api_node_create(body: String) -> String {
if str_eq(importance, "low") { 0.25 } else { 0.5 }
}
}
let id: String = engram_node_full(content, node_type, label,
let id: String = wt_node(content, node_type, label,
sal, sal, el_from_float(0.9),
tier, tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -422,11 +463,11 @@ fn handle_api_node_update(body: String) -> String {
}
let body_tags: String = json_get(body, "tags")
let tags: String = if str_eq(body_tags, "") { "[\"" + node_type + "\"]" } else { body_tags }
let new_id: String = engram_node_full(content, node_type, label,
let new_id: String = wt_node(content, node_type, label,
el_from_float(0.5), el_from_float(0.5), el_from_float(0.8),
tier, tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
engram_connect(new_id, id, el_from_float(0.9), "supersedes")
wt_edge(new_id, id, el_from_float(0.9), "supersedes")
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + id + "\",\"ok\":true}"
}
@@ -450,7 +491,12 @@ fn handle_api_recall(method: String, path: String, body: String) -> String {
if str_eq(eff_q, "") {
return api_or_empty(engram_scan_nodes_json(limit, 0))
}
let results: String = engram_search_json(eff_q, limit)
// engram_recall_json, not engram_search_json: this route IS the retrieval
// surface (claim 24's "embedding search queries"), so it gets the semantic
// and associative legs. engram_search_json stays lexical because ~40
// internal call sites pass a KEY and seven of them delete every record
// that comes back see the boundary note above eg_search_json_impl.
let results: String = engram_recall_json(eff_q, limit)
return api_or_empty(results)
}
@@ -498,7 +544,7 @@ fn handle_api_capture_knowledge(body: String) -> String {
let full: String = if str_eq(title, "") { content } else { title + ": " + content }
let lbl: String = str_slice(title, 0, 80)
let tags: String = "[\"Knowledge\",\"captured\"]"
let id: String = engram_node_full(full, "Knowledge", lbl,
let id: String = wt_node(full, "Knowledge", lbl,
el_from_float(0.85), el_from_float(0.8), el_from_float(0.9),
"Episodic", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -513,12 +559,12 @@ fn handle_api_evolve_knowledge(body: String) -> String {
if !str_eq(prior_id, "") && is_protected_node(prior_id) { return api_err_protected(prior_id) }
let tags: String = "[\"Knowledge\",\"evolved\"]"
// Empty label engram_node_full derives content[:60] (LABEL FIX 2026-07-23).
let new_id: String = engram_node_full(content, "Knowledge", "",
let new_id: String = wt_node(content, "Knowledge", "",
el_from_float(0.75), el_from_float(0.75), el_from_float(0.9),
"Episodic", tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
if !str_eq(prior_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true}"
}
@@ -535,11 +581,11 @@ fn handle_api_promote_knowledge(body: String) -> String {
"[\"Knowledge\",\"tier:canonical\",\"disposition:stable\"]"
} else { tags_raw }
// Empty label engram_node_full derives content[:60] (LABEL FIX 2026-07-23).
let new_id: String = engram_node_full(content, "Knowledge", "",
let new_id: String = wt_node(content, "Knowledge", "",
el_from_float(0.9), el_from_float(0.9), el_from_float(1.0),
"Canonical", tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
engram_connect(new_id, prior_id, el_from_float(0.95), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.95), "supersedes")
return "{\"ok\":true,\"new_id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\"}"
}
@@ -562,7 +608,7 @@ fn handle_api_define_process(body: String) -> String {
if str_eq(content, "") { return api_err("content is required") }
let label: String = if str_eq(name, "") { "process:unnamed" } else { "process:" + name }
let tags: String = "[\"Process\"]"
let id: String = engram_node_full(content, "Process", label,
let id: String = wt_node(content, "Process", label,
el_from_float(0.8), el_from_float(0.8), el_from_float(0.9),
"Canonical", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -647,7 +693,7 @@ fn handle_api_tune_config(body: String) -> String {
if str_eq(key, "") { return api_err("key is required") }
let content: String = "config:" + key + "=" + value
let tags: String = "[\"ConfigEntry\",\"config\"]"
let id: String = engram_node_full(content, "ConfigEntry", key,
let id: String = wt_node(content, "ConfigEntry", key,
el_from_float(0.85), el_from_float(0.85), el_from_float(0.9),
"Canonical", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -694,7 +740,7 @@ fn handle_api_link_entities(body: String) -> String {
if is_protected_node(to_id) { return api_err_protected(to_id) }
let relation: String = json_get(body, "relation")
let eff_relation: String = if str_eq(relation, "") { "associates" } else { relation }
engram_connect(from_id, to_id, el_from_float(0.5), eff_relation)
wt_edge(from_id, to_id, el_from_float(0.5), eff_relation)
return "{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + eff_relation + "\"}"
}
@@ -727,11 +773,11 @@ fn handle_api_evolve_memory(body: String) -> String {
}
}
let tags: String = "[\"Memory\",\"evolved\"]"
let new_id: String = engram_node_full(content, "Memory", "memory:evolved",
let new_id: String = wt_node(content, "Memory", "memory:evolved",
sal, sal, el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true}"
}
@@ -789,11 +835,11 @@ fn handle_api_cultivate(body: String) -> String {
let content: String = json_get(body, "content")
if str_eq(content, "") { return api_err("content is required") }
let tags: String = "[\"Knowledge\",\"evolved\",\"cultivated\"]"
let new_id: String = engram_node_full(content, "Knowledge", "knowledge:cultivated",
let new_id: String = wt_node(content, "Knowledge", "knowledge:cultivated",
el_from_float(0.75), el_from_float(0.75), el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true,\"cultivated\":true}"
}
@@ -809,11 +855,11 @@ fn handle_api_cultivate(body: String) -> String {
}
}
let tags: String = "[\"Memory\",\"evolved\",\"cultivated\"]"
let new_id: String = engram_node_full(content, "Memory", "memory:cultivated",
let new_id: String = wt_node(content, "Memory", "memory:cultivated",
sal, sal, el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true,\"cultivated\":true}"
}
@@ -833,7 +879,7 @@ fn handle_api_cultivate(body: String) -> String {
if str_eq(to_id, "") { return api_err("to_id is required") }
let relation: String = json_get(body, "relation")
let eff_relation: String = if str_eq(relation, "") { "associates" } else { relation }
engram_connect(from_id, to_id, el_from_float(0.5), eff_relation)
wt_edge(from_id, to_id, el_from_float(0.5), eff_relation)
return "{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + eff_relation + "\",\"cultivated\":true}"
}
@@ -868,7 +914,7 @@ fn handle_api_consolidate(body: String) -> String {
if !str_eq(summary, "") {
let safe_summary: String = str_replace(summary, "\"", "'")
let tags: String = "[\"SessionSummary\",\"consolidate\"]"
let summary_id: String = engram_node_full(
let summary_id: String = wt_node(
"[session-summary] " + safe_summary,
"SessionSummary", "session:summary",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
@@ -880,3 +926,462 @@ fn handle_api_consolidate(body: String) -> String {
}
return "{\"ok\":true,\"snapshot\":\"" + snap + "\"}"
}
// Stage 1: structural audit
//
// WHAT THIS IMPLEMENTS
// The CGI provisional, 05-detailed-description.md, "Stage 1: Structural audit
// 430". Verbatim, the audit module evaluates: the density and typed
// distribution of causal edges; the consistency between value nodes and
// execution-record neighborhoods; the richness and connectivity of the
// self-model; and the authenticity of open-question nodes in the wonder
// manifest. It "produces a coherence assessment 432 — NOT A BINARY SCORE but
// an annotated characterization of the graph's structural properties".
//
// That last clause is the whole shape of this handler. Every finding carries
// its own numbers AND a plain-language note saying what the numbers mean and
// how they were obtained. There is no pass/fail, no percentage-of-health, no
// composite score, and `"score":null` is emitted explicitly so a downstream
// reader cannot mistake its absence for an omission.
//
// WHY IT EXISTS NOW, AND WHY THE FIRST FINDING IS THE ONE IT IS
// `runStructuralAudit` has been an advertised MCP tool with nothing behind it:
// the dispatcher GET'd /session/begin and returned that blob (mcp-wrapper/src/
// main.el). Meanwhile the failure the audit would have caught ran silently for
// about three weeks the soul reported 103,089 nodes while the engram, which
// OWNS persistence, held ~79,900; a crash discarded the difference. Every boot
// reported green throughout, because nothing in the system ever compared the
// two sides. So finding 1 is owner-versus-runtime divergence: it is the check
// whose absence cost real memory, and it is cheap and exact.
//
// WHAT IS DELIBERATELY NOT HERE (stage 1b, see the `deferred` array in the
// response): value/execution-record consistency and wonder-manifest
// authenticity. Both need node types that barely exist in this graph today
// the response MEASURES those populations and reports the counts as the reason,
// rather than asserting a deferral without evidence.
//
// MEASUREMENT HONESTY: EXACT WHERE CHEAP, SAMPLED WHERE NOT, ALWAYS LABELLED
// Counts, edge typing and self-model connectivity are exact. Orphan rate and
// dangling-edge rate are SAMPLED, because the engram runtime has no node-id
// index `engram_find_node_index` is a linear scan over every node, so an
// exhaustive dangling check is O(nodes x edges) (~2.2e9 string compares at
// today's scale, tens of seconds inside one request). The samples are UNIFORM
// across the whole population, not head-of-list, and every sampled figure is
// emitted with its own `sampled` / `population` fields plus an extrapolation
// labelled as such. Raise `?edge_sample=` / `?node_sample=` to the population
// size to run either check exhaustively and pay the time. The real fix is an
// id index in the runtime; that is the engram repo's, not this handler's.
// audit_pct1 one-decimal percentage as a bare JSON number, sign-safe.
// Integer math only: EL has no fixed-precision formatter, and float_to_str
// would put an unbounded mantissa in the response.
fn audit_pct1(num: Int, den: Int) -> String {
if den <= 0 { return "null" }
let neg: Bool = num < 0
let a: Int = if neg { 0 - num } else { num }
let tenths: Int = (a * 1000) / den
let whole: Int = tenths / 10
let frac: Int = tenths - (whole * 10)
let sign: String = if neg { "-" } else { "" }
return sign + int_to_str(whole) + "." + int_to_str(frac)
}
// audit_finding the one envelope every finding uses: name, the measurements,
// and the annotation. Keeping it in one place is what stops the characterization
// from degenerating into a bag of numbers with no reading attached.
fn audit_finding(name: String, measured: String, note: String) -> String {
return "{\"finding\":\"" + name + "\""
+ ",\"measured\":{" + measured + "}"
+ ",\"note\":\"" + api_json_escape(note) + "\"}"
}
// audit_str_at read the quoted string value starting at byte `start`.
// Slices a bounded window rather than the tail of the (multi-MB) edges array, so
// this is O(window) per call instead of O(remaining input).
fn audit_str_at(s: String, start: Int, maxlen: Int) -> String {
let n: Int = str_len(s)
if start < 0 || start >= n { return "" }
let end_guess: Int = start + maxlen
let stop: Int = if end_guess > n { n } else { end_guess }
let win: String = str_slice(s, start, stop)
let q: Int = str_index_of(win, "\"")
if q < 0 { return "" }
return str_slice(win, 0, q)
}
// audit_rel_count exact count of edges carrying `rel`, by scanning the emitted
// edge array for the literal `"relation":"<rel>"`. engram_emit_edge_json writes
// metadata ESCAPED as a string, so no nested object can contain that literal and
// the count cannot be inflated by edge payloads.
fn audit_rel_count(edges: String, rel: String) -> Int {
return str_count(edges, "\"relation\":\"" + rel + "\"")
}
// audit_owner_stats ask the persistence OWNER for its own counts.
// Returns "" when there is no HTTP owner configured or the owner is unreachable;
// both are reported as findings, never as a failure of the audit.
fn audit_owner_stats(url: String) -> String {
if str_eq(url, "") { return "" }
return http_get(url + "/api/stats")
}
// audit_divergence FINDING 1. Runtime (this soul's in-process graph) versus
// the persistence owner's own count. Trend is measured against the previous
// audit recorded in soul state, so a second call answers "is the gap growing?"
// rather than just restating it.
fn audit_divergence() -> String {
let rt_nodes: Int = engram_node_count()
let rt_edges: Int = engram_edge_count()
let url: String = wt_engram_url()
if str_eq(url, "") {
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"none\",\"owner_reachable\":false",
"No HTTP persistence owner is configured, so this soul IS the owner "
+ "(file mode) and divergence is not defined. This check only has "
+ "meaning when ENGRAM_URL points at a separate engram that owns the "
+ "canonical store.")
}
let stats: String = audit_owner_stats(url)
// REACHABILITY IS PROVED BY THE PAYLOAD, NOT BY A NON-EMPTY REPLY.
// http_get does not return "" on a connection failure it returns a JSON
// error object ({"error":"Failed to connect to ... Couldn't connect to
// server"}). Testing only for "" made a DEAD owner read as reachable with
// node_count 0, i.e. the audit would have reported a 100% divergence and
// named it as data loss. That false positive is worse than no check at all:
// it is precisely the kind of confident wrong answer this route exists to
// stop. Require the field the contract promises.
let owner_nc_raw: String = json_get_raw(stats, "node_count")
if str_eq(stats, "") || str_eq(owner_nc_raw, "") {
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"" + api_json_escape(url) + "\",\"owner_reachable\":false"
+ ",\"owner_reply\":\"" + api_json_escape(api_utf8_trunc(stats, 200)) + "\"",
"The persistence owner at " + url + " did not return a node_count "
+ "from GET /api/stats. Divergence is UNKNOWN, NOT ZERO — an owner "
+ "that cannot be read is exactly the condition under which the "
+ "runtime's own count means least, and reporting 0 for the owner "
+ "would manufacture a total-loss reading out of a network error. "
+ "Reported as a finding rather than raised as an error so the rest "
+ "of the audit still returns; the owner's raw reply is in "
+ "owner_reply.")
}
let ow_nodes: Int = json_get_int(stats, "node_count")
let ow_edges: Int = json_get_int(stats, "edge_count")
let d_nodes: Int = rt_nodes - ow_nodes
let d_edges: Int = rt_edges - ow_edges
// Trend against the previous audit in this soul's state.
let prev_raw: String = state_get("audit_prev_node_delta")
let prev: Int = str_to_int(prev_raw)
let abs_now: Int = if d_nodes < 0 { 0 - d_nodes } else { d_nodes }
let abs_prev: Int = if prev < 0 { 0 - prev } else { prev }
let trend: String = if str_eq(prev_raw, "") {
"no_prior_audit"
} else {
if abs_now > abs_prev { "growing" } else {
if abs_now < abs_prev { "shrinking" } else { "flat" }
}
}
state_set("audit_prev_node_delta", int_to_str(d_nodes))
state_set("audit_prev_ts", int_to_str(time_now()))
let note_head: String = if d_nodes == 0 {
"Runtime and owner agree on node count."
} else {
"Runtime holds " + int_to_str(d_nodes) + " nodes (" + audit_pct1(d_nodes, rt_nodes)
+ "% of its own graph) that the persistence owner does not report. Nodes "
+ "that exist only in runtime memory do not survive a restart."
}
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"" + api_json_escape(url) + "\",\"owner_reachable\":true"
+ ",\"owner_nodes\":" + int_to_str(ow_nodes)
+ ",\"owner_edges\":" + int_to_str(ow_edges)
+ ",\"node_delta\":" + int_to_str(d_nodes)
+ ",\"edge_delta\":" + int_to_str(d_edges)
+ ",\"node_delta_pct_of_runtime\":" + audit_pct1(d_nodes, rt_nodes)
+ ",\"trend_vs_previous_audit\":\"" + trend + "\""
+ ",\"previous_node_delta\":" + (if str_eq(prev_raw, "") { "null" } else { int_to_str(prev) }),
note_head + " Trend against the previous audit recorded in this soul's "
+ "state: " + trend + ". This is the comparison whose absence let a "
+ "~24,000-node loss run for weeks with every boot reporting green.")
}
// audit_edge_typing FINDING 2. Density plus the typed distribution the patent
// asks for, against the claim-10 relation vocabulary. Exact: str_count over the
// emitted edge array, one linear pass per relation.
fn audit_edge_typing(edges: String, total_edges: Int, node_total: Int) -> String {
let c_sup: Int = audit_rel_count(edges, "Supersedes")
let c_cau: Int = audit_rel_count(edges, "Causes")
let c_con: Int = audit_rel_count(edges, "Contains")
let c_ref: Int = audit_rel_count(edges, "References")
let c_ctr: Int = audit_rel_count(edges, "Contradicts")
let c_exe: Int = audit_rel_count(edges, "Exemplifies")
let c_act: Int = audit_rel_count(edges, "Activates")
let c_tmp: Int = audit_rel_count(edges, "TemporallyPrecedes")
let typed: Int = c_sup + c_cau + c_con + c_ref + c_ctr + c_exe + c_act + c_tmp
// Lowercase near-misses: the same eight concepts written by the ad-hoc write
// paths (linkEntities defaults to "associates", linkCausal to "causes").
// Counted separately because "the vocabulary is unused" and "the vocabulary
// is used in the wrong case" are different defects with different fixes.
let l_sup: Int = audit_rel_count(edges, "supersedes")
let l_cau: Int = audit_rel_count(edges, "causes")
let l_con: Int = audit_rel_count(edges, "contains")
let l_ref: Int = audit_rel_count(edges, "references")
let l_ctr: Int = audit_rel_count(edges, "contradicts")
let l_exe: Int = audit_rel_count(edges, "exemplifies")
let l_act: Int = audit_rel_count(edges, "activates")
let l_tmp: Int = audit_rel_count(edges, "temporallyPrecedes")
let near: Int = l_sup + l_cau + l_con + l_ref + l_ctr + l_exe + l_act + l_tmp
let untyped: Int = total_edges - typed
return audit_finding("typed_edge_distribution",
"\"total_edges\":" + int_to_str(total_edges)
+ ",\"total_nodes\":" + int_to_str(node_total)
// Density per 100 nodes, not per node: EL has no fixed-precision float
// formatter, and "0.3 edges per node" rounded to an integer is a lie.
+ ",\"edges_per_100_nodes\":" + audit_pct1(total_edges, node_total)
+ ",\"claim10_typed\":" + int_to_str(typed)
+ ",\"claim10_typed_pct\":" + audit_pct1(typed, total_edges)
+ ",\"outside_claim10_vocabulary\":" + int_to_str(untyped)
+ ",\"lowercase_near_miss\":" + int_to_str(near)
+ ",\"by_relation\":{"
+ "\"Supersedes\":" + int_to_str(c_sup)
+ ",\"Causes\":" + int_to_str(c_cau)
+ ",\"Contains\":" + int_to_str(c_con)
+ ",\"References\":" + int_to_str(c_ref)
+ ",\"Contradicts\":" + int_to_str(c_ctr)
+ ",\"Exemplifies\":" + int_to_str(c_exe)
+ ",\"Activates\":" + int_to_str(c_act)
+ ",\"TemporallyPrecedes\":" + int_to_str(c_tmp) + "}",
"Only " + int_to_str(typed) + " of " + int_to_str(total_edges)
+ " edges use the claim-10 causal vocabulary; the remainder are ad-hoc "
+ "relation strings, which is why the graph's causal claims cannot yet "
+ "be checked for internal consistency — an untyped edge asserts "
+ "association, not causation. " + int_to_str(near) + " edges use a "
+ "lowercase spelling of a claim-10 relation: those are near-misses the "
+ "write paths could be corrected to emit, not genuinely foreign types.")
}
// audit_orphans_dangling FINDING 3. Both figures are SAMPLED; see the header
// for why exhaustive is O(nodes x edges) on this runtime.
//
// An "orphan" here is a node with zero RESOLVABLE edges: engram_neighbors_json
// drops any edge whose other endpoint does not resolve to a node, so a node
// whose only edges are dangling reads as an orphan. That is the right reading
// such a node is unreachable by traversal but it is stated rather than hidden.
fn audit_orphans_dangling(edges: String, total_edges: Int, node_total: Int,
edge_cap: Int, node_cap: Int) -> String {
// orphan sample: uniform stride over the node store
let n_take: Int = if node_total < node_cap { node_total } else { node_cap }
let n_stride: Int = if n_take > 0 { node_total / n_take } else { 1 }
let n_stride = if n_stride < 1 { 1 } else { n_stride }
let orphans: Int = 0
let n_checked: Int = 0
let j: Int = 0
while j < n_take {
let one: String = engram_scan_nodes_json(1, j * n_stride)
let nid: String = json_get(json_array_get(one, 0), "id")
if !str_eq(nid, "") {
let nbrs: String = engram_neighbors_json(nid, 1, "both")
let deg: Int = json_array_len(nbrs)
let orphans = if deg == 0 { orphans + 1 } else { orphans }
let n_checked = n_checked + 1
}
let j = j + 1
}
// dangling sample: uniform stride over the edge array
// str_index_of_all gives every edge's field offsets in ONE linear pass, so
// any index can be read in O(1). json_array_get would have been O(i) per
// element and O(n^2) over the array.
let from_pos: [Int] = str_index_of_all(edges, "\"from_id\":\"")
let to_pos: [Int] = str_index_of_all(edges, "\"to_id\":\"")
let nf: Int = len(from_pos)
let nt: Int = len(to_pos)
let ne: Int = if nf < nt { nf } else { nt }
let e_take: Int = if ne < edge_cap { ne } else { edge_cap }
let e_stride: Int = if e_take > 0 { ne / e_take } else { 1 }
let e_stride = if e_stride < 1 { 1 } else { e_stride }
let dangling: Int = 0
let e_checked: Int = 0
let i: Int = 0
while i < ne && e_checked < e_take {
let fid: String = audit_str_at(edges, get(from_pos, i) + 11, 96)
let tid: String = audit_str_at(edges, get(to_pos, i) + 9, 96)
let f_gone: Bool = str_eq(engram_get_node_json(fid), "{}")
let t_gone: Bool = if f_gone { true } else { str_eq(engram_get_node_json(tid), "{}") }
let dangling = if f_gone || t_gone { dangling + 1 } else { dangling }
let e_checked = e_checked + 1
let i = i + e_stride
}
let orphan_est: Int = if n_checked > 0 { (orphans * node_total) / n_checked } else { 0 }
let dangle_est: Int = if e_checked > 0 { (dangling * total_edges) / e_checked } else { 0 }
let exhaustive_n: String = if n_checked >= node_total { "true" } else { "false" }
let exhaustive_e: String = if e_checked >= ne { "true" } else { "false" }
return audit_finding("orphans_and_dangling_edges",
"\"nodes_population\":" + int_to_str(node_total)
+ ",\"nodes_sampled\":" + int_to_str(n_checked)
+ ",\"nodes_sample_exhaustive\":" + exhaustive_n
+ ",\"orphans_in_sample\":" + int_to_str(orphans)
+ ",\"orphan_rate_pct\":" + audit_pct1(orphans, n_checked)
+ ",\"orphans_extrapolated\":" + int_to_str(orphan_est)
+ ",\"edges_population\":" + int_to_str(total_edges)
+ ",\"edges_sampled\":" + int_to_str(e_checked)
+ ",\"edges_sample_exhaustive\":" + exhaustive_e
+ ",\"dangling_in_sample\":" + int_to_str(dangling)
+ ",\"dangling_rate_pct\":" + audit_pct1(dangling, e_checked)
+ ",\"dangling_extrapolated\":" + int_to_str(dangle_est),
"Orphan = zero RESOLVABLE edges, so a node whose only edges dangle counts "
+ "as an orphan; either way it is unreachable by traversal. Dangling = an "
+ "edge with an endpoint id that resolves to no node. Both are uniform "
+ "stride samples over the whole population, not the head of the list; "
+ "the extrapolations are estimates and are labelled as such. Pass "
+ "?node_sample= / ?edge_sample= at or above the population size to run "
+ "either check exhaustively. A high orphan rate is a characterization, "
+ "not a verdict: an accumulating store legitimately holds unlinked "
+ "material. It becomes a defect when the write paths were SUPPOSED to "
+ "link and did not.")
}
// audit_pillar one self-model pillar: present, how much content, how connected.
fn audit_pillar(key: String, id: String) -> String {
let node: String = engram_get_node_json(id)
let present: Bool = !str_eq(node, "{}") && !str_eq(node, "")
if !present {
return "\"" + key + "\":{\"id\":\"" + id + "\",\"present\":false"
+ ",\"content_length\":0,\"degree\":0}"
}
let content: String = json_get(node, "content")
let deg: Int = json_array_len(engram_neighbors_json(id, 1, "both"))
return "\"" + key + "\":{\"id\":\"" + id + "\",\"present\":true"
+ ",\"label\":\"" + api_json_escape(json_get(node, "label")) + "\""
+ ",\"tier\":\"" + api_json_escape(json_get(node, "tier")) + "\""
+ ",\"content_length\":" + int_to_str(str_len(content))
+ ",\"degree\":" + int_to_str(deg) + "}"
}
// audit_self_model FINDING 4. "the richness and connectivity of the
// self-model ... is it connected to behavioral evidence?"
//
// This finding RETIRES the Claude-side vitals identity block. That check lived
// outside the system it was checking a shell script grepping a snapshot so
// it could only ever report on a file, and it went on reporting green while the
// memory-philosophy pillar was absent from the live graph for about three weeks.
// Asking the running soul about its own three pillars is the designed mechanism;
// a shell probe was the fourth patch on the same hole.
fn audit_self_model() -> String {
let dna: String = audit_pillar("intellectual_dna", "kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6")
let val: String = audit_pillar("values_hub", "kn-5b606390-a52d-4ca2-8e0e-eba141d13440")
let phi: String = audit_pillar("memory_philosophy", "kn-dcfe04b3-3702-4cac-b6f0-ecb4db837eee")
let root: String = audit_pillar("self_root", "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee")
return audit_finding("self_model_connectivity",
"\"pillars\":{" + dna + "," + val + "," + phi + "," + root + "}",
"The three identity pillars plus the self root. `degree` counts nodes "
+ "reachable in one hop in either direction — the self-model's connection "
+ "to the rest of the graph. present:false on any pillar is the condition "
+ "that ran undetected for weeks; content_length distinguishes a pillar "
+ "that is present from one that is present but hollowed out. The patent "
+ "also asks whether the self-model makes ACCURATE PREDICTIONS about the "
+ "system's own behavior; that half needs Prediction nodes and is deferred "
+ "with the rest of stage 1b below.")
}
// audit_deferred what stage 1 does NOT yet evaluate, with the measured reason.
// Emitted as data, not as a comment, so a reader of the assessment sees the gap
// and its evidence rather than inferring completeness from silence.
fn audit_deferred() -> String {
let preds: Int = json_array_len(api_or_empty(engram_scan_nodes_by_type_json("Prediction", 50, 0)))
let wonders: Int = json_array_len(api_or_empty(engram_scan_nodes_by_type_json("WonderQuestion", 50, 0)))
return "[{\"deferred\":\"value_execution_record_consistency\""
+ ",\"stage\":\"1b\""
+ ",\"measured\":{\"prediction_nodes_found\":" + int_to_str(preds) + "}"
+ ",\"reason\":\"" + api_json_escape(
"The patent asks whether the execution history SUPPORTS the stated "
+ "values or shows systematic conflict. That requires execution "
+ "records tied to value nodes and predictions to score them against. "
+ "Prediction nodes found (capped at 50): " + int_to_str(preds)
+ ". Asserting value/execution coherence on that population would be "
+ "a fabricated result, which is worse than a stated gap.") + "\"}"
+ ",{\"deferred\":\"wonder_manifest_authenticity\""
+ ",\"stage\":\"1b\""
+ ",\"measured\":{\"wonder_question_nodes_found\":" + int_to_str(wonders) + "}"
+ ",\"reason\":\"" + api_json_escape(
"The patent asks whether pull weights CORRELATE WITH GENUINE "
+ "PREDICTION UNCERTAINTY or are uniform/externally assigned — a "
+ "correlation between two populations. WonderQuestion nodes readable "
+ "by type (capped at 50): " + int_to_str(wonders) + ", against "
+ int_to_str(preds) + " Prediction nodes. There is a known write/read "
+ "node-type mismatch on the wonder path; until that is fixed and both "
+ "populations exist, any correlation reported here would be noise.") + "\"}]"
}
// handle_api_structural_audit Stage 1. Returns the coherence assessment 432:
// an annotated characterization, explicitly NOT a score.
//
// COST NOTE: the edge findings need the relation labels, and the runtime exposes
// no edge-enumeration builtin. The only way to see them is the same one
// GET /api/graph/edges already uses engram_save to a SCRATCH path (never the
// owner's canonical file; see routes.el, neuron#117) and read the array back.
// On a large graph that is a multi-hundred-MB write, so this is a manual audit
// route, not something to put on a timer. Pass ?edges=0 to skip both edge
// findings and get the divergence + self-model readings cheaply.
fn handle_api_structural_audit(method: String, path: String, body: String) -> String {
let node_total: Int = engram_node_count()
let edge_total: Int = engram_edge_count()
let want_edges: Bool = !str_eq(api_query_param(path, "edges"), "0")
let edge_cap: Int = api_query_int(path, "edge_sample", 3000)
let node_cap: Int = api_query_int(path, "node_sample", 300)
let divergence: String = audit_divergence()
let self_model: String = audit_self_model()
let edge_part: String = if want_edges {
// Scratch export only. state_get("soul_snapshot_path") is deliberately
// NOT used: in HTTP-engram mode the soul is not the persistence owner and
// must never write the canonical file, not even on a read path.
let scratch_dir: String = env("TMPDIR")
let scratch_base: String = if str_eq(scratch_dir, "") { "/tmp" } else { scratch_dir }
let snap_path: String = scratch_base + "/soul-audit-export-" + state_get("soul_cgi_id") + ".json"
// engram_save returns Int (1 ok / 0 fail); str_eq on it SIGSEGVs (#150).
let saved: Int = engram_save(snap_path)
if saved == 0 {
"," + audit_finding("typed_edge_distribution", "\"available\":false",
"Could not export the graph to " + snap_path + " for edge analysis, "
+ "so edge typing and the dangling-edge sample were not run. "
+ "Reported as a gap, not as zero findings.")
} else {
// wt_read, not fs_read: fs_read leaves a thread-local length hint that
// the NEXT HTTP response would use as its Content-Length, appending
// adjacent heap bytes to the reply (see persist.el wt_read).
let snap: String = wt_read(snap_path)
let edges_raw: String = json_get_raw(snap, "edges")
let edges: String = if str_eq(edges_raw, "") { "[]" } else { edges_raw }
"," + audit_edge_typing(edges, edge_total, node_total)
+ "," + audit_orphans_dangling(edges, edge_total, node_total, edge_cap, node_cap)
}
} else {
""
}
return "{\"audit\":\"structural\",\"stage\":1"
+ ",\"spec\":\"CGI provisional 05-detailed-description.md, Stage 1: Structural audit 430\""
+ ",\"assessment\":\"coherence_assessment_432\""
+ ",\"assessment_kind\":\"annotated_characterization\""
+ ",\"score\":null"
+ ",\"score_note\":\"By design. The specification calls for an annotated characterization of the graph's structural properties, not a binary score. Read the findings.\""
+ ",\"cgi_id\":\"" + api_json_escape(state_get("soul_cgi_id")) + "\""
+ ",\"ts_ms\":" + int_to_str(time_now())
+ ",\"findings\":[" + divergence + "," + self_model + edge_part + "]"
+ ",\"deferred\":" + audit_deferred() + "}"
}
+1
View File
@@ -37,3 +37,4 @@ extern fn handle_api_memory_update(body: String) -> String
extern fn handle_api_cultivate(body: String) -> String
extern fn handle_api_list_typed(node_type: String, path: String, body: String) -> String
extern fn handle_api_consolidate(body: String) -> String
extern fn handle_api_structural_audit(method: String, path: String, body: String) -> String
+426
View File
@@ -0,0 +1,426 @@
// persist.el the soulengram WRITE-THROUGH boundary (neuron#117).
//
// WHY THIS FILE EXISTS
// soul.el:571-573 states the ownership rule: "when ENGRAM_URL is set the HTTP
// Engram owns persistence the soul must NEVER write to the local snapshot
// (not the persistence owner)." The soul obeys the NEGATIVE half. The POSITIVE
// half how a write made inside the soul actually REACHES the owner was
// never built. Sync is pull-only (awareness.el `/api/sync` -> engram_load_merge),
// so every node the soul creates lives in its process RAM and is shed on
// restart. Measured live 2026-08-07: soul node_count=102184, engram
// node_count=79197 ~23k nodes existing nowhere but RAM.
//
// SCOPE NOTE ON THE PATENT (corrects an earlier internal reading)
// Engram provisional claims 15-18 describe a delta-sync protocol "with peer
// Engram instances"; claim 17's pull-then-push sequence is PEER-ENGRAM to
// PEER-ENGRAM. The soul is NOT a peer Engram it is a CALLER of the database
// system API (cf. claim 27, "invoked explicitly by a caller of the database
// system API"). So claim 17 does not specify a soul↔engram contract and is not
// cited as authority here. This design is derived from the ownership rule
// alone: the owner owns the writes, therefore the soul must HAND writes to the
// owner and must never write the owner's file itself.
//
// THE MECHANISM, AND WHY NOT `POST /api/nodes`
// The obvious route is the one the persona/boot-counter write-backs already
// use, POST /api/nodes. It is the wrong instrument here, verified against the
// live engram binary in a sandbox:
// - it mints a NEW server-side id (engram_node), so the soul's id and the
// owner's id diverge the next /api/sync pull re-imports the node as a
// DUPLICATE, and any edge referencing the soul's id never resolves;
// - it accepts only {content, node_type, salience} and drops label, tier,
// tags, importance, confidence, metadata. A probe posted with tier
// "Canonical" came back tier "Working", importance 0.5.
// POST /api/load-merge (Will's own route, el `dc39a61`) is the right one:
// - engram_load_merge PRESERVES the id and every field;
// - it dedups nodes by id and edges by (from_id,to_id,relation), so a
// re-submitted delta is a NO-OP retry safety is free, and it is the same
// local-wins semantics the graph already uses;
// - it calls persist_canonical() THE OWNER writes its own canonical file.
// The soul never touches it. The ownership rule is honoured in its
// strongest form rather than worked around;
// - it returns real counts {ok, nodes_added, edges_added, node_count},
// so a receipt can be a MEASUREMENT instead of a fixed success shape.
//
// SPOOL-AND-DRAIN, AND WHY IT IS NOT JUST A DIRECT POST
// Measured in a sandbox against a 79k-node / 176MB graph (live scale): one
// load-merge costs ~0.38s, essentially all of it the owner's persist_canonical.
// A chat turn writes 5-7 nodes; pushing each separately would add ~2.7s per
// turn. So writes are STAGED and pushed in one coalesced batch.
// The staging buffer is the FILESYSTEM, not process state, because the soul
// serves each HTTP connection on its own pthread (el_runtime http_serve_async)
// and a shared in-process buffer would lose entries to a read-modify-write
// race silently, which is the one failure mode this file exists to end.
// One file per write, named with uuid_v4, is race-free by construction and
// buys a property a memory buffer cannot: writes that could not be pushed
// SURVIVE A SOUL CRASH and are drained on the next boot.
//
// WHAT IS DELIBERATELY NOT PUSHED
// - InternalStateEvent / heartbeat telemetry. Will's own carve-out, stated in
// engram server.el 8f8ccc9: "48h-pruned, loss-tolerant, ~2/min; snapshotting
// 28MB per heartbeat is waste."
// NOTE (ours, flagged for Will): we do NOT additionally exclude Working-tier
// nodes. That exclusion exists in `fb0bb55` to stop the boot counter leaking
// through the /api/sync PULL; it is about sync backflow, not durability.
// Applying it here would exclude mem_store which writes tier "Working" and
// mem_store is the single most important durable write path in the soul. Boot
// seeding reads the canonical file wholesale, so a pushed Working-tier node
// does survive restart. This is the one classification call this file makes
// that Will has not ruled on.
//
// WHAT THIS BOUNDARY CANNOT EXPRESS (by construction, not by omission)
// - engram_strengthen (salience/activation drift): load-merge SKIPS ids that
// already exist, so it cannot update an existing node. There is no owner-side
// update/upsert route. Not pushable through any current route; left as a
// follow-up that needs a change in the engram repo.
// - engram_forget (hard delete): load-merge is additive and has no delete verb.
// Propagating deletes would mean DELETE /api/nodes/<id>, a HARD delete at the
// owner which scripts/verify-soul-contract.sh section B explicitly fails the
// build for ("to delete is to supersede/tombstone, never hard-remove"). Local
// deletes therefore stay local; the TOMBSTONE NODE and its "tombstones" edge
// (mem_tombstone) are pushed, and that is the sanctioned representation of a
// deletion in this graph.
// Configuration
// wt_engram_url same resolution order as ise_post: env, then the state key
// stashed at boot. NO hardcoded localhost fallback: unlike telemetry, inventing
// a destination for durable data would risk pushing a user's memories at whatever
// happens to be listening on 8742. Empty means "no HTTP owner" -> file mode.
fn wt_engram_url() -> String {
let env_url: String = env("ENGRAM_URL")
if !str_eq(env_url, "") { return env_url }
return state_get("soul_engram_url")
}
fn wt_api_key() -> String {
let env_key: String = env("ENGRAM_API_KEY")
if !str_eq(env_key, "") { return env_key }
return state_get("soul_engram_api_key")
}
// wt_enabled true only in HTTP-engram mode. In file mode the soul IS the
// persistence owner and every path below is a no-op, so this whole feature is
// inert for genesis/local deployments. That is also what makes it reversible.
fn wt_enabled() -> Bool {
return !str_eq(wt_engram_url(), "")
}
// wt_spool_dir where staged deltas live. MUST be readable by the engram
// process: /api/load-merge takes a PATH and the owner opens it itself. Both
// processes are same-host by construction (dev-stack LaunchAgents; the GKE
// image starts engram and soul in one container per entrypoint.sh).
fn wt_spool_dir() -> String {
let raw: String = env("SOUL_OUTBOX_DIR")
let dir: String = if str_eq(raw, "") { env("HOME") + "/.neuron/soul-outbox" } else { raw }
fs_mkdir(dir)
return dir
}
// Helpers
// wt_esc minimal JSON string escape. Deliberately local rather than reusing
// chat.el's json_safe: persist.el is imported BY memory.el, which is imported by
// chat.el, so depending on chat.el here would be an import cycle.
fn wt_esc(s: String) -> String {
let s1: String = str_replace(s, "\\", "\\\\")
let s2: String = str_replace(s1, "\"", "\\\"")
let s3: String = str_replace(s2, "\n", "\\n")
let s4: String = str_replace(s3, "\r", "\\r")
let s5: String = str_replace(s4, "\t", "\\t")
return s5
}
// wt_durable_class Will's telemetry carve-out, by node_type. See header.
fn wt_durable_class(node_type: String) -> Bool {
if str_eq(node_type, "InternalStateEvent") { return false }
return true
}
// wt_inner strip the surrounding brackets off a JSON array so several arrays
// can be concatenated into one. Returns "" for "[]" / "" / anything too short.
fn wt_inner(arr: String) -> String {
let n: Int = str_len(arr)
if n < 3 { return "" }
if !str_starts_with(arr, "[") { return "" }
return str_slice(arr, 1, n - 1)
}
// wt_read fs_read, plus a MANDATORY reset of the runtime's binary-length hint.
//
// THIS IS NOT OPTIONAL AND MUST NOT BE "SIMPLIFIED" BACK TO A BARE fs_read.
// The pinned runtime (vendor/el-runtime/v1.0.0-20260501) keeps a thread-local
// `_tl_fs_read_len` that fs_read SETS to the file's byte count (so binary files
// can be served with a correct Content-Length) and that http_send_response
// CONSUMES as the Content-Length of the next reply. Nothing else clears it
// except json_get_raw. So any fs_read during request handling that is not
// followed by a json_get_raw makes the NEXT HTTP response advertise the FILE's
// length instead of the body's and the runtime then sends that many bytes,
// appending whatever adjacent heap memory follows the reply.
//
// Caught here, measured: a /api/neuron/memory reply that should be 86 bytes went
// out as 497, with 411 bytes of this module's own spool paths and log strings
// trailing the JSON. The drain reads spool files mid-request, so this boundary
// is exactly where the landmine gets stepped on.
//
// Upstream el fixed the class in `43636ae` ("pair fs_read length hint with its
// buffer"); that runtime is NOT the one vendored here, and re-pinning the
// runtime is deliberately out of scope for this change. Clearing the hint at
// our own boundary fixes our exposure without touching the pinned C.
// json_get_raw is used as the reset because it is the only builtin in this
// runtime that zeroes the hint, and it does so before any early return.
fn wt_clear_binlen() -> Void {
let discard: String = json_get_raw("{}", "_wt_reset")
}
fn wt_read(path: String) -> String {
let data: String = fs_read(path)
wt_clear_binlen()
return data
}
// wt_sweep best-effort removal of the zero-byte husks left by truncation.
// The runtime exposes no unlink builtin, so a drained delta is emptied rather
// than deleted; this reclaims the directory entries.
//
// `-empty` is the safety property, not an optimisation: the command is
// STRUCTURALLY INCAPABLE of removing a delta that still has content, so it can
// never destroy a pending write even if it runs concurrently with a stage.
// Only the directory path is interpolated (never a filename), and it is quoted.
// The exit code is ignored an un-swept husk costs one directory entry.
fn wt_sweep(dir: String) -> Void {
if str_eq(dir, "") { return }
if str_contains(dir, "'") { return }
exec_command("find '" + dir + "' -maxdepth 1 -name 'wt*.json' -empty -delete 2>/dev/null")
}
// Staging
// wt_stage write ONE delta file. uuid_v4 in the name makes concurrent stagers
// collision-free without any lock. Returns true if the delta is on disk.
fn wt_stage(nodes_json: String, edges_json: String) -> Bool {
let dir: String = wt_spool_dir()
if str_eq(dir, "") { return false }
let payload: String = "{\"nodes\":" + nodes_json + ",\"edges\":" + edges_json + "}"
let path: String = dir + "/wt-" + uuid_v4() + ".json"
fs_write(path, payload)
// Read-back-verify the stage itself. A stage that did not land is a write we
// would otherwise believe was queued exactly the hallucinated-save class.
if str_eq(wt_read(path), "") {
println("[persist] wt_stage: FAILED to write spool file " + path + " — delta not queued")
return false
}
return true
}
// The write boundary
// wt_node create a node locally AND queue it for the persistence owner.
// Same signature and same return contract as engram_node_full ("" on failure),
// so converting a call site is a rename and nothing else.
fn wt_node(content: String, node_type: String, label: String,
salience: Float, importance: Float, confidence: Float,
tier: String, tags: String) -> String {
let id: String = engram_node_full(content, node_type, label,
salience, importance, confidence,
tier, tags)
if str_eq(id, "") { return "" }
// engram_get_node_json emits the SAME record shape engram_save writes (minus
// the embedding vector, which the owner backfills lazily), so the read-back
// doubles as the delta payload no second serialization to drift.
let rec: String = engram_get_node_json(id)
if str_eq(rec, "") || str_eq(rec, "{}") {
println("[persist] wt_node: local write did not read back, id=" + id + " label=" + label)
return ""
}
if wt_enabled() && wt_durable_class(node_type) {
wt_stage("[" + rec + "]", "[]")
}
return id
}
// wt_edge create an edge locally AND queue it. Mirrors engram_connect.
//
// The edge id is freshly generated rather than read back: the runtime exposes no
// "id of the edge I just created" accessor, and the owner dedups edges by
// (from_id,to_id,relation), never by id so the id is not load-bearing. The
// consequence, stated plainly: the soul's copy and the owner's copy of the same
// edge carry different edge ids. Nothing in either codebase looks an edge up by
// id (neighbors traversal scans from_id/to_id), so this is cosmetic.
fn wt_edge(from_id: String, to_id: String, weight: Float, relation: String) -> Void {
engram_connect(from_id, to_id, weight, relation)
if !wt_enabled() { return }
if str_eq(from_id, "") || str_eq(to_id, "") { return }
let ts: Int = time_now()
let rec: String = "{\"id\":\"" + uuid_v4() + "\""
+ ",\"from_id\":\"" + wt_esc(from_id) + "\""
+ ",\"to_id\":\"" + wt_esc(to_id) + "\""
+ ",\"relation\":\"" + wt_esc(relation) + "\""
+ ",\"metadata\":\"{}\""
+ ",\"weight\":" + float_to_str(weight)
+ ",\"confidence\":1"
+ ",\"created_at\":" + int_to_str(ts)
+ ",\"updated_at\":" + int_to_str(ts)
+ ",\"last_fired\":0,\"inhibitory\":0,\"layer_id\":1}"
wt_stage("[]", "[" + rec + "]")
}
// The drain
// wt_drain coalesce every staged delta into ONE load-merge against the owner.
//
// Returns: nodes_added on success (>= 0), 0 when there was nothing to do, and
// -1 when the push FAILED. -1 is load-bearing: on failure the spool files are
// left untouched, so nothing is lost and the next drain retries them. A caller
// must never read a non-negative return as "my particular node is durable"
// use wt_durable(id) for that.
//
// Concurrency: several threads may drain at once. Each builds its own batch file
// (uuid-named), and overlapping batches are harmless because load-merge dedups.
// Files are truncated ONLY after a confirmed ok:true, so a lost race costs a
// redundant push, never a dropped write.
fn wt_drain() -> Int {
if !wt_enabled() { return 0 }
let dir: String = wt_spool_dir()
if str_eq(dir, "") { return 0 }
// el_list_len/el_list_get, NOT json_stringify(fs_list(...)): fs_list builds
// a native list via el_list_append, and json_stringify does not serialize
// that type it renders the raw pointer value. (Verified in isolation; the
// same latent defect is live in studio.el's /api/tools/file/list route,
// which returns e.g. {"entries":4386409744}. Noted, not fixed here.)
let listing = fs_list(dir)
let count: Int = el_list_len(listing)
if count == 0 { return 0 }
let nodes_acc: String = ""
let edges_acc: String = ""
let drained: String = ""
let found: Int = 0
let i: Int = 0
// No `continue` / `break`: elc lists them as keywords but not one line of
// the shipped soul uses either, so they are unexercised on this build path.
// Guard conditions are expressed as nested ifs instead, and every rebind is
// at the loop-body top level where `let x = ...` is assignment (the idiom
// memory.el's boot-counter loop relies on) never inside a nested block,
// where it would shadow instead.
while i < count {
let name: String = el_list_get(listing, i)
// A delta is only usable when it ends with the closing "]}" that
// wt_stage writes last. fs_write is not atomic, so a file being written
// right now can be observed half-formed; requiring the terminator means
// it is picked up whole on the next drain instead of merged as garbage.
// An empty read means "already drained and truncated" not an error.
let p: String = if str_starts_with(name, "wt-") { dir + "/" + name } else { "" }
let raw: String = if str_eq(p, "") { "" } else { wt_read(p) }
let usable: Bool = !str_eq(raw, "") && str_ends_with(raw, "]}")
let nj: String = if usable { wt_inner(json_get_raw(raw, "nodes")) } else { "" }
let ej: String = if usable { wt_inner(json_get_raw(raw, "edges")) } else { "" }
let nodes_acc = if str_eq(nj, "") { nodes_acc } else if str_eq(nodes_acc, "") { nj } else { nodes_acc + "," + nj }
let edges_acc = if str_eq(ej, "") { edges_acc } else if str_eq(edges_acc, "") { ej } else { edges_acc + "," + ej }
let drained = if !usable { drained } else if str_eq(drained, "") { p } else { drained + "\n" + p }
let found = if usable { found + 1 } else { found }
let i = i + 1
}
if found == 0 { return 0 }
let combined: String = "{\"nodes\":[" + nodes_acc + "],\"edges\":[" + edges_acc + "]}"
let batch: String = dir + "/wtb-" + uuid_v4() + ".json"
fs_write(batch, combined)
if str_eq(wt_read(batch), "") {
println("[persist] wt_drain: could not write batch file " + batch + "" + int_to_str(found) + " deltas stay queued")
return -1
}
let url: String = wt_engram_url()
let key: String = wt_api_key()
let body: String = "{\"path\":\"" + wt_esc(batch) + "\",\"_auth\":\"" + wt_esc(key) + "\"}"
let resp: String = http_post_json(url + "/api/load-merge", body)
// The batch file is pure scratch the retry is rebuilt from the SPOOL, not
// from it. Truncate it unconditionally, before branching on the outcome, so
// a persistently unreachable owner cannot accumulate one husk per attempt.
fs_write(batch, "")
// Distinguish the two failures rather than collapsing them: "cannot reach
// the owner" and "the owner refused this delta" need different human
// responses, and a log line that says the wrong one costs a debugging hour.
// curl surfaces transport errors as a JSON body, so an empty response is not
// the only unreachable signal.
// (str_contains rather than a strict parse on purpose the engram's HTTP
// responses have been observed carrying trailing bytes past the JSON.)
let unreachable: Bool = str_eq(resp, "")
|| str_contains(resp, "Couldn't connect")
|| str_contains(resp, "Failed to connect")
|| str_contains(resp, "Could not resolve")
|| str_contains(resp, "timed out")
if unreachable {
wt_sweep(dir)
println("[persist] wt_drain: owner UNREACHABLE at " + url + "" + int_to_str(found)
+ " deltas stay queued in " + dir + " (will retry): " + resp)
return -1
}
if !str_contains(resp, "\"ok\":true") {
wt_sweep(dir)
println("[persist] wt_drain: owner REJECTED the delta — " + int_to_str(found)
+ " stay queued in " + dir + ": " + resp)
return -1
}
let added: Int = json_get_int(resp, "nodes_added")
let added_e: Int = json_get_int(resp, "edges_added")
// Confirmed. Truncate the drained spool files so they are not re-pushed.
// Truncation (not deletion) because the runtime exposes no unlink builtin;
// an emptied file is inert to the loop above. The zero-byte husks are then
// swept below.
let paths = str_split(drained, "\n")
let pn: Int = el_list_len(paths)
let k: Int = 0
while k < pn {
let one: String = el_list_get(paths, k)
if !str_eq(one, "") { fs_write(one, "") }
let k = k + 1
}
wt_sweep(dir)
println("[persist] wt_drain: pushed " + int_to_str(found) + " deltas -> owner added "
+ int_to_str(added) + " nodes, " + int_to_str(added_e) + " edges")
return added
}
// wt_durable is this id present AT THE OWNER? The only honest answer to
// "did my write persist" in HTTP mode.
//
// In file mode the soul IS the owner, so the local read-back is the owner-side
// read-back and this collapses to the pre-existing check.
//
// nodes_added from wt_drain is NOT a substitute: a concurrent drain may have
// already pushed this node, making our own added count 0 while the node is
// perfectly durable. Presence at the owner is the fact; counts are telemetry.
fn wt_durable(id: String) -> Bool {
if str_eq(id, "") { return false }
if !wt_enabled() {
let local: String = engram_get_node_json(id)
return !str_eq(local, "") && !str_eq(local, "null") && !str_eq(local, "{}")
}
let url: String = wt_engram_url()
let resp: String = http_get(url + "/api/nodes/" + id)
if str_eq(resp, "") { return false }
if str_eq(resp, "{}") { return false }
return str_contains(resp, "\"id\"")
}
// wt_commit flush, then assert at the owner. The receipt callers should use.
// Deliberately NOT a fixed success shape: it can and does return false while the
// local write is perfectly fine in RAM, which is the true state of affairs when
// the owner is unreachable.
fn wt_commit(id: String) -> Bool {
if str_eq(id, "") { return false }
if !wt_enabled() {
let local: String = engram_get_node_json(id)
return !str_eq(local, "") && !str_eq(local, "null") && !str_eq(local, "{}")
}
let pushed: Int = wt_drain()
return wt_durable(id)
}
+58 -7
View File
@@ -186,7 +186,7 @@ fn route_imprint_contextual(body: String) -> String {
return "{\"ok\":false,\"error\":\"empty body\"}"
}
let tags: String = "[\"imprint\",\"contextual\"]"
let id: String = engram_node_full(
let id: String = wt_node(
body,
"Entity",
"imprint:contextual",
@@ -208,7 +208,7 @@ fn route_imprint_user(body: String) -> String {
return "{\"ok\":false,\"error\":\"empty body\"}"
}
let tags: String = "[\"imprint\",\"user\"]"
let id: String = engram_node_full(
let id: String = wt_node(
body,
"Entity",
"imprint:user",
@@ -239,7 +239,7 @@ fn route_synthesize(body: String) -> String {
}
let req: String = "synthesize " + parent_a + " " + parent_b
let tags: String = "[\"soul-inbox-pending\",\"synthesis-request\"]"
engram_node_full(
wt_node(
req,
"Entity",
"synthesis-request",
@@ -395,7 +395,28 @@ fn handle_connectors(method: String, clean: String, body: String) -> String {
return "{\"ok\":false,\"error\":\"unknown connectors route\"}"
}
// handle_request the soul's HTTP entry point.
//
// NOTE ON THE NAME (neuron#117): the el runtime resolves this handler by NAME
// via dlsym(RTLD_DEFAULT, "handle_request") that is why the Linux build must
// link -rdynamic. So the dispatcher body moved to route_dispatch and the name
// `handle_request` stays put as a thin wrapper. Do not rename it back.
//
// The wrapper exists to give the write-through boundary a guaranteed flush
// point. route_dispatch returns from ~60 places; a per-branch flush would be
// forgotten on the 61st. Draining here means EVERY request that staged a write
// pushes it before the connection closes, whatever route produced it, including
// routes added later that know nothing about persistence.
//
// wt_drain is a no-op (no HTTP, no cost) when nothing is staged and when the
// soul is not in HTTP-engram mode, so this is free on read traffic.
fn handle_request(method: String, path: String, body: String) -> String {
let resp: String = route_dispatch(method, path, body)
let flushed: Int = wt_drain()
return resp
}
fn route_dispatch(method: String, path: String, body: String) -> String {
let clean: String = strip_query(path)
// ACTIVITY STAMP (2026-07-30 self-review): every inbound HTTP request
@@ -432,10 +453,27 @@ fn handle_request(method: String, path: String, body: String) -> String {
return engram_scan_nodes_json(9999, 0)
}
if str_eq(clean, "/api/graph/edges") {
// TODO(reliability #8): engram_save races with awareness loop mem_save().
// Both now use atomic write-to-temp+rename (el_runtime.c). Serialised
// by engram_global_mu. Future: add engram_edges_json() builtin.
let snap_path: String = env("HOME") + "/.neuron/engram/snapshot.json"
// FIXED (neuron#117): this GET used to engram_save() straight over
// ~/.neuron/engram/snapshot.json a READ route, in a process that is
// NOT the persistence owner, overwriting the owner's canonical file
// on every call. It broke soul.el:571-573 ("the soul must NEVER write
// to the local snapshot") and it is the same defect class Will removed
// from the engram itself in el `dc39a61` ("stop read routes clobbering
// canonical snapshot"), where route_scan_edges/route_sync were moved
// to scratch paths for exactly this reason. It was also the race the
// old TODO(reliability #8) admitted to.
//
// Export to a scratch path instead. Same response, no canonical write.
// The soul's own snapshot writes are otherwise already gated behind
// state key "soul_snapshot_path", which is set ONLY in the genesis
// file-mode branch (soul.el: is_genesis && safe_to_seed, and
// safe_to_seed is unconditionally false when ENGRAM_URL is set) so
// after this change the soul writes nothing at all in HTTP mode.
// Future: add an engram_edges_json() builtin and drop the file round
// trip entirely.
let scratch_dir: String = env("TMPDIR")
let scratch_base: String = if str_eq(scratch_dir, "") { "/tmp" } else { scratch_dir }
let snap_path: String = scratch_base + "/soul-edges-export-" + state_get("soul_cgi_id") + ".json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
let edges_raw: String = json_get_raw(snap, "edges")
@@ -529,6 +567,13 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_starts_with(clean, "/api/neuron/graph") {
return handle_api_inspect_graph(method, path, body)
}
// Stage 1 structural audit (CGI provisional, "Structural audit 430").
// GET because it is a read of the graph's own structure; the query string
// carries the sample caps (?edge_sample=, ?node_sample=, ?edges=0), so
// str_starts_with rather than str_eq.
if str_starts_with(clean, "/api/neuron/audit/structural") {
return handle_api_structural_audit(method, path, body)
}
if str_starts_with(clean, "/api/neuron/list/") {
// Offset 17 = len("/api/neuron/list/"). Was 16, which left a leading "/" on node_type
// ("/BacklogItem"), so engram_scan_nodes_by_type_json matched nothing list/<type>
@@ -710,6 +755,12 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(clean, "/api/neuron/graph/link") {
return handle_api_link_entities(body)
}
// POST accepted too: same handler, so a JSON-RPC-shaped caller that only
// speaks POST reaches the identical audit. Options still come from the
// query string the handler reads no body fields.
if str_eq(clean, "/api/neuron/audit/structural") {
return handle_api_structural_audit(method, path, body)
}
if str_eq(clean, "/api/neuron/memory") {
return handle_api_remember(body)
}
+1 -1
View File
@@ -204,7 +204,7 @@ fn safety_log_bell(level: String, reason: String, input_summary: String) -> Stri
// Emit a fallback println so the bell event leaves at least a log trace even
// when engram is degraded. This does not replace engram persistence -- it is a
// last-resort audit trail when the primary write cannot be confirmed.
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content,
"BellEvent",
"bell:" + level,
+7 -7
View File
@@ -87,7 +87,7 @@ fn session_create(body: String) -> String {
let folder: String = json_get(body, "folder")
let content: String = session_make_content(id, title, ts, ts, folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -358,7 +358,7 @@ fn session_update_patch(session_id: String, body: String) -> String {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, eff_title, created_int, ts, eff_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_node_id: String = engram_node_full(
let new_node_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -456,7 +456,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
// TODO(reliability #7): delete-then-insert is not atomic concurrent saves for the
// same session can produce orphan history nodes. State is primary truth; engram fallback.
let tags: String = "[\"session\",\"session-history\",\"Conversation\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
hist, "Conversation", "session:messages:" + session_id,
el_from_float(0.6), el_from_float(0.6), el_from_float(0.9),
"Episodic", tags
@@ -488,7 +488,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
+ " | ts:" + int_to_str(ts_now)
let summary_tags: String = "[\"session-emotional-summary\",\"affective\",\"bell:" + eff_level + "\",\"BellEvent\"]"
let summary_sal: String = if str_eq(eff_level, "hard") { el_from_float(0.95) } else { el_from_float(0.85) }
let sum_discard: String = engram_node_full(
let sum_discard: String = wt_node(
summary_content,
"BellEvent",
"session:emotional-summary",
@@ -529,7 +529,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
if !str_eq(ot_id, "") { engram_forget(ot_id) }
let oti = oti + 1
}
let discard_topic: String = engram_node_full(
let discard_topic: String = wt_node(
topic_content, "Conversation", topic_label,
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", topic_tags
@@ -582,7 +582,7 @@ fn session_update_meta_timestamp(session_id: String) -> Void {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, old_title, created_int, ts, old_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_id: String = engram_node_full(
let new_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -629,7 +629,7 @@ fn session_auto_title(session_id: String, first_message: String) -> Void {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, new_title, created_int, ts, old_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_id: String = engram_node_full(
let new_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
+17
View File
@@ -657,6 +657,23 @@ if is_genesis && safe_to_seed {
}
}
// CRASH RECOVERY (neuron#117). Deltas the previous process staged but could not
// hand to the owner are still on disk the spool is a filesystem queue, not a
// memory buffer, precisely so that a soul that died mid-flight does not take its
// unpushed writes with it. Drain them before serving, so recovered memories are
// durable and recallable from the owner from the first request onward.
//
// Safe on a clean boot: an empty spool means no HTTP call at all. Safe in file
// mode: wt_drain returns immediately when ENGRAM_URL is unset.
let wt_recovered: Int = wt_drain()
if wt_recovered > 0 {
println("[soul] write-through: recovered " + int_to_str(wt_recovered)
+ " nodes from a previous process's spool -> persistence owner")
}
if wt_recovered < 0 {
println("[soul] write-through: spool present but the persistence owner is unreachable — queued, will retry on heartbeat")
}
println("[soul] serving on port " + int_to_str(port))
http_serve_async(port, "handle_request")
println("[soul] awareness loop starting")
+2 -2
View File
@@ -11,7 +11,7 @@ import "memory.el"
fn steward_log_event(kind: String, detail: String) -> Void {
let content: String = "STEWARD:" + kind + " | " + detail
let tags: String = "[\"stewardship\",\"steward:" + kind + "\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
content,
"StewardshipEvent",
"steward:" + kind,
@@ -221,7 +221,7 @@ fn steward_fingerprint_session(input: String, session_id: String) -> String {
+ " formality=" + fs_str
+ " time=" + tb_str
let sample_tags: String = "[\"behavior\",\"BehaviorSample\",\"stewardship\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
sample_content,
"BehaviorSample",
"behavior:" + session_id,
+90
View File
@@ -0,0 +1,90 @@
# gate-openai — deterministic OpenAI-dialect provider stub
Staging home for the **soul-openai-tools-v2** gate scaffolding
(`docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md`, test-plan rung 1:
"stub first — discriminates before El code exists"). Sibling of gate9's
Anthropic stub (`_wt-beta-round9/scripts/gate9/stub-llm.py`): same scenario
mechanism, opposite wire dialect. Stdlib Python only, 127.0.0.1 only,
refuses ports 7770/7779/17779. Run `./selftest.sh` — exit 0 is green.
## Files
| File | Role |
|---|---|
| `stub-openai.py` | HTTP server: `POST /v1/chat/completions` (OpenAI dialect), scenario-scripted responses, request validation, ground-truth JSONL log, hostile modes via `--mode` |
| `scenarios-openai.json` | Scenario contract: scripts + markers + per-class/per-step request assertions |
| `selftest.sh` | curl-driven proof of every scenario, every rejection, all hostile modes (58 checks) |
## What each scenario proves (when the brain drives it)
| Class | Proves |
|---|---|
| `oa-plain` | finish_reason `stop` ends the loop; tools + `tool_choice` + `parallel_tool_calls:false` were offered on the wire |
| `oa-tools-off` | the chat-only lane sends NO tools (offering them there is a 400) |
| `oa-single-tool` | full round-trip: `tool_calls` parsed, assistant echo + `role:"tool"` turn with matching `tool_call_id` sent back, final text reached |
| `oa-torture` | `function.arguments` (JSON-encoded string with nested quotes, backslashes, newlines, tabs, unicode) survives exactly ONE decode — the stub recomputes the issued payload from the script and 400s on any drift (`gate_echo_mismatch`, the spec §6 two-escaper trap) |
| `oa-parallel` | two `tool_calls` in one response: the brain either answers both (paired correctly) or rejects cleanly — an unpaired echo is a 400 |
| `oa-mission` | multi-round loop continuation; step index = assistant-message count, so resume threads index correctly by construction |
| `oa-api-error` | provider errors 400/429/500/503 in the OpenAI error envelope surface honestly, no retry storm |
Universal (every request, any scenario): Anthropic dialect leakage fails
loudly with 400 — `anthropic-version` header, top-level `system` /
`stop_sequences` / `max_tokens_to_sample`, `input_schema` inside tools,
Anthropic content blocks (`tool_use`/`tool_result`/...). Tools must be
`{type:"function", function:{name, description, parameters}}`, unique names;
echoed `arguments` must be a JSON-encoded STRING, never a decoded object.
## Hostile modes (`--mode`, same file)
| Mode | Behavior | Brain invariant under test |
|---|---|---|
| `black-hole` | reads the request, never responds | HTTP timeout exists and surfaces; no silent hang |
| `mid-body-drop` | 200 headers, half a JSON body, socket abort | truncated body = clean error, never a half-parsed reply shown as real |
| `tool-pending-forever` | every request gets a fresh `tool_calls` response, forever | the loop's iteration cap trips (`max_loop_iterations: 16` in the contract); count actual round-trips via `GET /gate/stats` (`chat_hits`) |
## How the brain-side gate consumes this
1. Start: `stub-openai.py --port P --scenarios scenarios-openai.json --log run.jsonl`
2. Point the brain at it: `NEURON_LLM_0_URL=http://127.0.0.1:P` +
`NEURON_LLM_0_FORMAT=openai` (spec step 0 must verify these actually
export at runtime), scratch profile, free soul port.
3. Send each phrasing's `prompt` (the marker selects the script); assert the
brain's claims (`tools_used`, reply, ledger) against the stub's JSONL log
— truth, not narration — plus files on disk for write_file scenarios.
4. Any stub 400 = the brain sent a malformed/leaked request; the gate fails
with the stub's reason string.
5. Re-run gate9's Anthropic matrix unchanged = proof the Anthropic lane is
byte-untouched.
## Reconciliation into gate9 (app repo) — AFTER round 9 merges
This dir is staging only; the merge is mechanical by design:
- `stub-openai.py` + `scenarios-openai.json` move to `scripts/gate9/`
alongside `stub-llm.py` + `scenarios.json` (shared conventions: marker
matching, assistant-count step indexing, `--port/--scenarios/--log`,
JSONL fields `seq/ts/kind/scenario_class/phrasing/step/validation/
delivered/http_status`, prod-port refusal, benign background responses,
`GATE-SCRIPT-EXHAUSTED` overrun, `{N}/{NN}` repeat expansion).
- `prompt-matrix-gate.sh` gains a dialect axis (anthropic|openai) choosing
stub + scenario file; `matrix-asserts.py` reads the same log shape.
- The `--mode` hostile flags here are PROVIDER-side (brain↔LLM boundary);
gate9's `hostile/` servers are SOUL-side (app↔brain boundary). They are
complementary, not duplicates — both stay.
## Open questions for the port author (stub asserts a position; confirm or change)
1. `parallel_tool_calls` must be **explicitly false** on every tool-bearing
request (ADR-0005 pin). If the builder omits it instead, relax
`defaults.expect_request.parallel_tool_calls` to `null`.
2. `tool_choice` must be present (`"auto"` expected). If the brain relies on
the provider default, drop `require_tool_choice`.
3. Tool-result `content` is asserted only to be a string; if the brain sends
structured JSON-in-string (like `{"ok":true,...}`), no change needed.
4. Groq compatibility: Groq's OpenAI-compat endpoint rejects some optional
fields; whatever field set the brain settles on for live Groq E2E must be
mirrored here so the deterministic gate and the live lane assert the SAME
request shape.
5. The stub treats a `role:"tool"` turn answering an already-answered id as
400; if the resume path can legitimately replay tool results, that rule
needs a resume-aware carve-out (gate9's Anthropic stub faced the same
issue — see its PASS 1 comment).
+559
View File
@@ -0,0 +1,559 @@
#!/usr/bin/env bash
# run-lane-gate.sh — brain-side driver for the OpenAI-dialect gate.
#
# Drives the REAL soul binary against stub-openai.py for every class and every
# phrasing in scenarios-openai.json, plus the three hostile provider modes, and
# asserts the brain's claims against the stub's ground-truth JSONL (truth, not
# narration) and against files on disk.
#
# SAFETY (hard rules, enforced below):
# - never binds 7770 / 7779 / 17779 - only 7891-7894
# - never reads or writes ~/.neuron - HOME is redirected to a scratch dir
# - every process started here is killed on exit (trap) and proven with lsof
#
# The soul runs under `script -q /dev/null` so its stdout is a pty: El's
# println() uses puts(), which is FULLY buffered to a file, and the process is
# killed without flushing — the DRIFT lines would be invisible otherwise.
#
# Usage: ./run-lane-gate.sh [all|bridge|local|toolsoff|hostile]
# bridge = consent round-trip config (no workspace root -> write_file is
# "escalate" -> the loop suspends and the CLIENT executes the tool)
# local = workspace-root config (write_file is "reversible" + builtin ->
# the loop executes the tool in-process and runs to completion)
# toolsoff = supplementary: non-agentic lane against a base URL WITHOUT the
# /v1 suffix (the el-runtime provider chain appends
# /v1/chat/completions itself, unlike chat.el which appends only
# /chat/completions)
# hostile = black-hole / mid-body-drop / tool-pending-forever
#
# Env overrides: SOUL_BIN, STUB_PORT, SOUL_PORT, SOUL_PORT_B, RUN_ROOT
set -uo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
PHASES="${1:-all}"
SOUL_BIN="${SOUL_BIN:-/tmp/soul-oai2/soul-openai-tools}"
STUB_PORT="${STUB_PORT:-7891}"
SOUL_PORT="${SOUL_PORT:-7892}"
SOUL_PORT_B="${SOUL_PORT_B:-7893}"
RUN_ROOT="${RUN_ROOT:-/tmp/oa-lane-gate}"
STAMP="$(date +%Y%m%d-%H%M%S)"
RUN="$RUN_ROOT/$STAMP"
for p in "$STUB_PORT" "$SOUL_PORT" "$SOUL_PORT_B"; do
case "$p" in
7770|7779|17779) echo "FATAL: refusing production Neuron port $p"; exit 2;;
789[1-4]) ;;
*) echo "FATAL: port $p outside the allowed 7891-7894 range"; exit 2;;
esac
done
[ -x "$SOUL_BIN" ] || { echo "FATAL: soul binary not found/executable: $SOUL_BIN"; exit 2; }
mkdir -p "$RUN/home" "$RUN/ws-bridge" "$RUN/ws-local" "$RUN/ws-off" "$RUN/engram"
echo '{"nodes":[],"edges":[]}' > "$RUN/engram/snapshot.json"
DRV="$RUN/drv.py"
STUB_PID=""; SOUL_PID=""
cleanup() {
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
pkill -f "$SOUL_BIN" 2>/dev/null
[ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null
sleep 0.4
[ -n "$SOUL_PID" ] && kill -9 "$SOUL_PID" 2>/dev/null
[ -n "$STUB_PID" ] && kill -9 "$STUB_PID" 2>/dev/null
return 0
}
trap cleanup EXIT INT TERM
start_stub() { # $1 = mode, $2 = log path
local mode="$1" log="$2" args=""
[ "$mode" = "normal" ] && args="--scenarios $HERE/scenarios-openai.json"
# shellcheck disable=SC2086
python3 "$HERE/stub-openai.py" --port "$STUB_PORT" --mode "$mode" --log "$log" $args \
> "$RUN/stub-$mode.out" 2>&1 &
STUB_PID=$!
for _ in $(seq 1 50); do
curl -sf "http://127.0.0.1:$STUB_PORT/gate/health" >/dev/null 2>&1 && return 0
sleep 0.2
done
echo "FATAL: stub did not come up on $STUB_PORT"; cat "$RUN/stub-$mode.out"; exit 3
}
stop_stub() { [ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null; sleep 0.3; STUB_PID=""; }
start_soul() { # $1 = port, $2 = base url, $3 = soul log, $4 = agent root ("" = none)
local port="$1" base="$2" log="$3" root="$4"
script -q /dev/null \
env -u ANTHROPIC_API_KEY -u SOUL_API_KEY -u ENGRAM_URL -u ENGRAM_API_KEY \
-u NEURON_API_URL -u NEURON_TOKEN -u SOUL_LLM_PROVIDER -u SOUL_LLM_BASE_URL \
-u NEURON_LLM_1_URL -u NEURON_LLM_1_KEY -u SOUL_IDENTITY \
HOME="$RUN/home" PATH="$PATH" \
NEURON_PORT="$port" EL_HTTP_BIND_HOST=127.0.0.1 \
SOUL_ENGRAM_PATH="$RUN/engram/snapshot.json" \
SOUL_CGI_ID=ntn-test SOUL_PERSONA_NAME=Neuron \
NEURON_LLM_0_URL="$base" NEURON_LLM_0_FORMAT=openai NEURON_LLM_0_KEY=gate-test-key \
${root:+NEURON_AGENT_ROOT="$root"} \
"$SOUL_BIN" > "$log" 2>&1 &
SOUL_PID=$!
for _ in $(seq 1 100); do
curl -sf "http://127.0.0.1:$port/health" >/dev/null 2>&1 && return 0
sleep 0.2
done
echo "FATAL: soul did not come up on $port"; tail -20 "$log"; exit 3
}
stop_soul() {
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
pkill -f "$SOUL_BIN" 2>/dev/null
sleep 0.6; SOUL_PID=""
}
# ------------------------------------------------------------------ driver ----
cat > "$DRV" <<'PYEOF'
import json, os, sys, time, threading, urllib.request, urllib.error
CFG = json.load(open(sys.argv[1]))
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
WS = CFG["workspace"]
MODE = CFG["mode"] # bridge | local | toolsoff
SCEN = json.load(open(CFG["scenarios"]))
STUBLOG = CFG["stub_log"]
SOULLOG = CFG["soul_log"]
ONLY = CFG.get("classes") or list(SCEN["classes"].keys())
MAXHOPS = CFG.get("max_hops", 15)
OUT = CFG["out"]
# the chat-only class must be driven on the NON-agentic door: the agentic door
# always advertises tools, which is a 400 on that scenario by contract.
NON_AGENTIC = {"oa-tools-off"}
def http(method, url, obj=None, timeout=300):
data = None if obj is None else json.dumps(obj).encode()
req = urllib.request.Request(url, data=data, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
body = r.read().decode("utf-8", "replace")
st = r.status
except urllib.error.HTTPError as e:
body = e.read().decode("utf-8", "replace"); st = e.code
except Exception as e:
return -1, "TRANSPORT-ERROR: %r" % (e,), None
try:
return st, body, json.loads(body)
except ValueError:
return st, body, None
def fsize(p):
return os.path.getsize(p) if os.path.exists(p) else 0
def tail_from(path, off):
if not os.path.exists(path):
return "", off
with open(path, "rb") as f:
f.seek(off); chunk = f.read(); return chunk.decode("utf-8", "replace"), f.tell()
def stub_since(off):
"""Exact correlation: only the JSONL bytes appended during this phrasing."""
txt, noff = tail_from(STUBLOG, off)
recs = []
for line in txt.splitlines():
line = line.strip()
if line:
try: recs.append(json.loads(line))
except ValueError: pass
return recs, noff
def perform(name, ti):
"""Execute the bridged tool for real, like the desktop client would."""
if name in ("write_file", "edit_file"):
p = ti.get("path", "")
dest = p if os.path.isabs(p) else os.path.join(WS, p)
os.makedirs(os.path.dirname(dest) or WS, exist_ok=True)
body = ti.get("content", "")
with open(dest, "w") as f:
f.write(body)
return "wrote %s (%d bytes)" % (p, len(body.encode()))
return "ok"
class Poller(threading.Thread):
def __init__(self, sid):
super().__init__(daemon=True); self.sid = sid; self.snaps = []; self.stop = False
def run(self):
while not self.stop:
st, body, js = http("GET", SOUL + "/api/run-progress/" + self.sid, timeout=60)
if js and js.get("progress"):
if not self.snaps or self.snaps[-1] != js["progress"]:
self.snaps.append(js["progress"])
time.sleep(0.1)
def progress(sid):
_, _, pj = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
return (pj or {}).get("progress")
def run_phrasing(cname, ph):
st, body, js = http("POST", SOUL + "/api/sessions", {"title": ph["id"]}, timeout=60)
sid = (js or {}).get("id", "")
rec = {"class": cname, "phrasing": ph["id"], "session_id": sid, "legs": [],
"pendings": [], "progress_during": [], "progress_per_leg": [],
"progress_final": None, "soul_log": "", "stub": [], "http": [],
"agentic": cname not in NON_AGENTIC}
if not sid:
rec["fatal"] = "session create failed: %s %s" % (st, body[:300]); return rec
soff = fsize(SOULLOG); loff = fsize(STUBLOG)
t0 = time.time()
pol = Poller(sid); pol.start()
payload = {"message": ph["prompt"], "session_id": sid, "workspace_root": WS,
"agentic": rec["agentic"]}
if MODE == "local":
payload["agent_workspace_root"] = WS
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("chat_timeout", 240))
rec["http"].append(st)
rec["legs"].append(js if js is not None else body[:600])
rec["progress_per_leg"].append(progress(sid))
hops = 0
while isinstance(js, dict) and js.get("tool_pending") and hops < MAXHOPS:
rec["pendings"].append({"call_id": js.get("call_id"), "tool_name": js.get("tool_name"),
"tool_input": js.get("tool_input"), "risk_tier": js.get("risk_tier"),
"narration": js.get("narration"), "tools_used": js.get("tools_used")})
try:
eff = perform(js.get("tool_name", ""), js.get("tool_input") or {})
except Exception as e:
eff = "client error: %r" % (e,)
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
{"call_id": js.get("call_id"), "content": eff},
timeout=CFG.get("chat_timeout", 240))
rec["http"].append(st)
rec["legs"].append(js if js is not None else body[:600])
rec["progress_per_leg"].append(progress(sid))
hops += 1
pol.stop = True; time.sleep(0.3)
t1 = time.time()
rec["elapsed"] = round(t1 - t0, 2)
rec["progress_during"] = pol.snaps
rec["progress_final"] = progress(sid)
rec["soul_log"], _ = tail_from(SOULLOG, soff)
rec["stub"], _ = stub_since(loff)
rec["hops"] = hops
return rec
# ------------------------------------------------------------- assertions ----
def expected_calls(cname, first_only=False):
out = []
for step in SCEN["classes"][cname]["script"]:
for k, call in enumerate(step.get("tool_calls") or []):
if first_only and k > 0:
continue
out.append((call["name"], call["arguments"]))
return out
def final_text(cname):
for step in reversed(SCEN["classes"][cname]["script"]):
if step.get("text") and not step.get("tool_calls"):
return step["text"]
return None
def judge(rec):
cname = rec["class"]; ok = []; bad = []
last = rec["legs"][-1] if rec["legs"] else None
reply = last.get("reply") if isinstance(last, dict) else None
err = last.get("error") if isinstance(last, dict) else None
tools_used = last.get("tools_used") if isinstance(last, dict) else None
stub = rec["stub"]
scen_recs = [r for r in stub if r.get("kind") == "scenario"]
rejects = [r for r in stub if r.get("validation") != "ok"]
bg = [r for r in stub if r.get("kind") in ("wrong_path", "background")]
def wire_clean():
if rejects:
for r in rejects:
bad.append("stub REJECTED a request: [%s] %s"
% (r.get("validation"), r.get("validation_detail")))
else:
ok.append("stub ground truth: validation \"ok\" on all %d scenario leg(s), no "
"gate_echo_mismatch / gate_tool_call_shape / dialect-leak 400s"
% len(scen_recs))
if bg:
ok.append("NOTE background non-scenario request(s) in this window: %s"
% [(r.get("kind"), r.get("path"), r.get("http_status")) for r in bg])
if cname == "oa-plain":
wire_clean()
want = final_text(cname)
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
if tools_used == []: ok.append("tools_used == [] (no tool ran)")
else: bad.append("tools_used expected [] got %r" % (tools_used,))
if reply and ('"tool_calls"' in reply or '"function"' in reply or '"tool_use"' in reply):
bad.append("tool-call JSON leaked into the reply text")
else: ok.append("no tool-call JSON anywhere in the reply")
elif cname in ("oa-single-tool", "oa-torture", "oa-mission"):
wire_clean()
want = final_text(cname)
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
exp = expected_calls(cname)
wantnames = [n for n, _ in exp]
if tools_used == wantnames:
ok.append("tools_used == %r (carried across %d suspension(s))" % (wantnames, rec["hops"]))
else:
bad.append("tools_used expected %r got %r" % (wantnames, tools_used))
for name, args in exp:
p = args.get("path"); c = args.get("content")
dest = os.path.join(WS, p)
if not os.path.exists(dest):
bad.append("expected file missing on disk: %s" % dest); continue
got = open(dest, "rb").read()
if got == c.encode():
ok.append("%s on disk is byte-for-byte the issued payload (%d bytes)" % (p, len(got)))
else:
bad.append("%s content differs\n WANT %r\n GOT %r"
% (p, c[:300], got[:300].decode("utf-8", "replace")))
if MODE == "bridge":
for pend, (name, args) in zip(rec["pendings"], exp):
if pend["tool_input"] == args:
ok.append("tool_input for %s survived exactly ONE decode (deep-equal to the "
"issued arguments; no double-escaping)" % name)
else:
bad.append("tool_input != issued arguments for %s\n WANT %r\n GOT %r"
% (name, args, pend["tool_input"]))
if rec["pendings"] and all(p["risk_tier"] == "escalate" for p in rec["pendings"]):
ok.append("every write_file classified \"escalate\" and bridged for consent")
elif cname == "oa-parallel":
drift = [l.strip() for l in rec["soul_log"].splitlines() if "DRIFT: provider returned" in l]
if drift: ok.append("soul log: " + drift[0])
else: bad.append("no 'DRIFT: provider returned N parallel tool_calls' line in the soul log")
delivered = [r for r in stub if r.get("delivered", {}).get("tool_calls")]
if delivered and len(delivered[0]["delivered"]["tool_calls"]) == 2:
ok.append("stub delivered 2 parallel tool_calls in one response (ground truth)")
if MODE == "bridge":
if len(rec["pendings"]) == 1:
ok.append("exactly ONE call honored: %s" % rec["pendings"][0]["call_id"])
else:
bad.append("expected exactly 1 honored call, got %d" % len(rec["pendings"]))
pairing = [r for r in rejects if "gate_pairing" in str(r.get("validation_detail")) or
"tool_calls at end of thread" in str(r.get("validation_detail")) or
"not fully answered" in str(r.get("validation_detail"))]
for r in pairing:
ok.append("EXPECTED-BY-CONTRACT stub 400 on the unpaired echo: %s"
% r.get("validation_detail"))
other = [r for r in rejects if r not in pairing]
for r in other:
bad.append("unexpected stub rejection: [%s] %s"
% (r.get("validation"), r.get("validation_detail")))
if err and not reply:
ok.append("honest error envelope after the 400 (no fabricated answer): %r" % err)
elif reply == final_text(cname):
ok.append("final reply == scripted final text (both calls paired)")
else:
bad.append("neither an honest error nor the scripted final text: %r" % (last,))
elif cname == "oa-api-error":
if err and not reply:
ok.append("honest error envelope: error=%r reply=%r" % (err, reply))
else:
bad.append("expected an error envelope with an empty reply, got %r" % (last,))
delivered = [r["delivered"].get("api_error") for r in stub if r.get("delivered")]
ok.append("stub delivered api_error status(es): %r" % [d for d in delivered if d])
n = len([r for r in stub if r.get("kind") == "scenario"])
ok.append("provider hit %d time(s) - no retry storm" % n)
if reply:
bad.append("FABRICATED ANSWER: reply non-empty on a provider error")
elif cname == "oa-tools-off":
ok.append("stub records for this phrasing: %r"
% [{k: r.get(k) for k in ("kind", "path", "validation", "http_status")} for r in stub])
wrong = [r for r in stub if r.get("kind") == "wrong_path"]
matched = [r for r in stub if r.get("scenario_class") == cname]
if matched and not rejects:
ok.append("chat-only request reached /v1/chat/completions with NO tools offered")
want = final_text(cname)
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
if reply and ('"tool_calls"' in reply or '"function"' in reply):
bad.append("tool-call JSON leaked into the reply text")
else: ok.append("no tool-call JSON in the reply")
elif wrong:
bad.append("the non-agentic lane never reached the provider endpoint: stub saw "
"%s -> %s (the el-runtime provider chain appends /v1/chat/completions "
"to NEURON_LLM_0_URL, chat.el appends only /chat/completions)"
% (wrong[0]["path"], wrong[0]["http_status"]))
elif not stub:
bad.append("no request reached the stub at all")
else:
for r in rejects:
bad.append("stub REJECTED: [%s] %s" % (r.get("validation"), r.get("validation_detail")))
return ok, bad
def main():
results = []
for cname in ONLY:
for ph in SCEN["classes"][cname]["phrasings"]:
rec = run_phrasing(cname, ph)
ok, bad = judge(rec)
rec["ok"] = ok; rec["bad"] = bad
rec["verdict"] = "FAIL" if bad else "PASS"
results.append(rec)
print("=" * 78)
print("[%s] %s / %s (%.2fs, %d bridge hop(s), agentic=%s, mode=%s)"
% (rec["verdict"], cname, ph["id"], rec.get("elapsed", 0),
rec.get("hops", 0), rec["agentic"], MODE))
for l in ok: print(" ok " + l.replace("\n", "\n "))
for l in bad: print(" FAIL " + l.replace("\n", "\n "))
for i, leg in enumerate(rec["legs"]):
print(" leg%d envelope: %s" % (i, json.dumps(leg)[:430]))
for i, pr in enumerate(rec["progress_per_leg"]):
print(" run-progress after leg%d: %s" % (i, json.dumps(pr)[:380]))
if rec["progress_during"]:
print(" run-progress polled DURING (%d distinct snapshot(s)), last: %s"
% (len(rec["progress_during"]), json.dumps(rec["progress_during"][-1])[:300]))
for r in rec["stub"]:
print(" stub: kind=%s class=%s phrasing=%s step=%s validation=%s%s delivered=%s http=%s"
% (r.get("kind"), r.get("scenario_class"), r.get("phrasing"), r.get("step"),
r.get("validation"),
("(" + str(r.get("validation_detail")) + ")") if r.get("validation_detail") else "",
json.dumps(r.get("delivered")), r.get("http_status")))
if rec["soul_log"].strip():
for l in rec["soul_log"].splitlines():
if l.strip(): print(" soul: " + l.strip())
json.dump(results, open(OUT, "w"), indent=1)
npass = sum(1 for r in results if r["verdict"] == "PASS")
print("=" * 78)
print("PHASE %s: %d/%d PASS" % (MODE, npass, len(results)))
for r in results:
print(" %-6s %-16s %s" % (r["verdict"], r["class"], r["phrasing"]))
return 0 if npass == len(results) else 1
sys.exit(main())
PYEOF
# ------------------------------------------------------------- hostile drv ---
cat > "$RUN/hostile.py" <<'PYEOF'
import json, os, sys, time, urllib.request, urllib.error
CFG = json.load(open(sys.argv[1]))
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
def http(method, url, obj=None, timeout=400):
data = None if obj is None else json.dumps(obj).encode()
req = urllib.request.Request(url, data=data, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
b = r.read().decode("utf-8", "replace"); st = r.status
except urllib.error.HTTPError as e:
b = e.read().decode("utf-8", "replace"); st = e.code
except Exception as e:
return -1, "TRANSPORT-ERROR: %r" % (e,), None
try:
return st, b, json.loads(b)
except ValueError:
return st, b, None
mode = CFG["mode"]; wsmode = CFG["ws_mode"]; WS = CFG["workspace"]
_, _, js = http("POST", SOUL + "/api/sessions", {"title": "hostile-" + mode}, timeout=60)
sid = (js or {}).get("id", "")
payload = {"message": "oa-gate plain probe: hostile mode %s" % mode,
"agentic": True, "session_id": sid, "workspace_root": WS}
if wsmode == "local":
payload["agent_workspace_root"] = WS
t0 = time.time()
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("timeout", 400))
t_first = time.time() - t0
legs = [js if js is not None else body[:500]]
hops = 0
while isinstance(js, dict) and js.get("tool_pending") and hops < CFG.get("max_hops", 14):
ti = js.get("tool_input") or {}
p = ti.get("path", "x.md")
dest = p if os.path.isabs(p) else os.path.join(WS, p)
try: open(dest, "w").write(ti.get("content", ""))
except Exception: pass
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
{"call_id": js.get("call_id"), "content": "ok"},
timeout=CFG.get("timeout", 400))
legs.append(js if js is not None else body[:500]); hops += 1
el = time.time() - t0
_, _, stats = http("GET", STUB + "/gate/stats", timeout=30)
_, _, prog = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
fab = [l for l in legs if isinstance(l, dict) and l.get("reply")]
print("HOSTILE %s (ws_mode=%s)" % (mode, wsmode))
print(" first /api/chat POST returned after %.2fs; whole chain %.2fs; client bridge hops=%d; "
"stub chat_hits=%s" % (t_first, el, hops, (stats or {}).get("chat_hits")))
print(" first envelope : " + json.dumps(legs[0])[:430])
print(" final envelope : " + json.dumps(legs[-1])[:430])
print(" non-empty replies anywhere in the chain (fabrication check): %d" % len(fab))
print(" run-progress : " + json.dumps(prog)[:300])
json.dump({"mode": mode, "ws_mode": wsmode, "t_first": t_first, "elapsed": el, "hops": hops,
"chat_hits": (stats or {}).get("chat_hits"), "legs": legs, "progress": prog},
open(CFG["out"], "w"), indent=1)
PYEOF
# ------------------------------------------------------------------ phases ---
RC_BRIDGE=0; RC_LOCAL=0; RC_OFF=0
run_normal_phase() { # $1 = label, $2 = soul port, $3 = agent root, $4 = ws, $5 = base, $6 = classes json
local m="$1" port="$2" root="$3" ws="$4" base="$5" classes="$6"
echo; echo "############ PHASE: $m (soul :$port, NEURON_LLM_0_URL=$base) ############"
start_stub normal "$RUN/stub-$m.jsonl"
start_soul "$port" "$base" "$RUN/soul-$m.log" "$root"
cat > "$RUN/cfg-$m.json" <<JSON
{"soul_port": $port, "stub_port": $STUB_PORT, "workspace": "$ws", "mode": "$m",
"scenarios": "$HERE/scenarios-openai.json", "stub_log": "$RUN/stub-$m.jsonl",
"soul_log": "$RUN/soul-$m.log", "out": "$RUN/results-$m.json", "chat_timeout": 240,
"classes": $classes}
JSON
python3 "$DRV" "$RUN/cfg-$m.json"
local rc=$?
stop_soul; stop_stub
return $rc
}
if [ "$PHASES" = "all" ] || [ "$PHASES" = "bridge" ]; then
run_normal_phase bridge "$SOUL_PORT" "" "$RUN/ws-bridge" "http://127.0.0.1:$STUB_PORT/v1" null
RC_BRIDGE=$?
fi
if [ "$PHASES" = "all" ] || [ "$PHASES" = "local" ]; then
run_normal_phase local "$SOUL_PORT_B" "$RUN/ws-local" "$RUN/ws-local" "http://127.0.0.1:$STUB_PORT/v1" null
RC_LOCAL=$?
fi
if [ "$PHASES" = "all" ] || [ "$PHASES" = "toolsoff" ]; then
# supplementary: the el-runtime provider chain appends /v1/chat/completions itself,
# so the non-agentic door needs the base WITHOUT the /v1 suffix.
run_normal_phase toolsoff "$SOUL_PORT" "" "$RUN/ws-off" "http://127.0.0.1:$STUB_PORT" '["oa-tools-off","oa-plain"]'
RC_OFF=$?
fi
if [ "$PHASES" = "all" ] || [ "$PHASES" = "hostile" ]; then
echo; echo "############ PHASE: hostile ############"
for spec in "black-hole:bridge" "mid-body-drop:bridge" "tool-pending-forever:bridge" "tool-pending-forever:local"; do
mode="${spec%%:*}"; wsm="${spec##*:}"
echo; echo "---- hostile mode=$mode ws_mode=$wsm ----"
start_stub "$mode" "$RUN/stub-$mode-$wsm.jsonl"
if [ "$wsm" = "local" ]; then
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" "$RUN/ws-local"
else
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" ""
fi
cat > "$RUN/cfg-$mode-$wsm.json" <<JSON
{"soul_port": $SOUL_PORT, "stub_port": $STUB_PORT, "mode": "$mode", "ws_mode": "$wsm",
"workspace": "$RUN/ws-local", "out": "$RUN/hostile-$mode-$wsm.json", "timeout": 400}
JSON
python3 "$RUN/hostile.py" "$RUN/cfg-$mode-$wsm.json"
echo " soul log (llm/DRIFT/cap lines):"
grep -E "DRIFT|llm error|iteration cap|\[llm\]" "$RUN/soul-$mode-$wsm.log" | tail -8 | sed 's/^/ /'
stop_soul; stop_stub
done
fi
echo; echo "############ CLEANUP ############"
cleanup
sleep 0.5
echo "processes still matching the soul binary:"; pgrep -fl "$SOUL_BIN" || echo " (none)"
echo "processes still matching stub-openai.py:"; pgrep -fl "stub-openai.py" || echo " (none)"
echo "lsof on 7891-7894 after cleanup:"
lsof -nP -iTCP:7891 -iTCP:7892 -iTCP:7893 -iTCP:7894 2>/dev/null || echo " (no listeners - ports free)"
echo
echo "############ SUMMARY ############"
echo "run dir: $RUN"
echo "bridge rc=$RC_BRIDGE local rc=$RC_LOCAL toolsoff rc=$RC_OFF (0 = every class PASS)"
exit $(( RC_BRIDGE + RC_LOCAL + RC_OFF ))
+110
View File
@@ -0,0 +1,110 @@
{
"_comment": "OpenAI-dialect gate scenario contract (soul-openai-tools-v2). Single source of truth shared by stub-openai.py (scripted provider responses + request assertions), selftest.sh (stub self-verification), and the future brain-side gate driver. Same structure as gate9's scenarios.json: classes -> script + phrasings with markers; scripts are CLASS-level so assertions are behavioral, never pinned to a sentence. expect_request keys: require_tools, require_tool_choice, parallel_tool_calls (expected literal value; null = don't check), forbid_tools. defaults apply to every class unless overridden; steps may override with their own expect_request.",
"deadline_secs": 60,
"max_loop_iterations": 16,
"defaults": {
"expect_request": {
"require_tools": true,
"require_tool_choice": true,
"parallel_tool_calls": false
}
},
"classes": {
"oa-plain": {
"script": [
{ "text": "Plain OpenAI-lane answer (gate fixture): the mechanism, the main caveat, and the practical takeaway in three sentences. No tools were needed for this one, and the finish reason on the wire is stop, which the loop must treat as terminal." }
],
"phrasings": [
{ "id": "oa-plain-p1", "marker": "oa-gate plain probe", "prompt": "oa-gate plain probe: explain the fixture topic simply." },
{ "id": "oa-plain-p2", "marker": "oa-gate second plain", "prompt": "oa-gate second plain: another phrasing of the plain question." }
]
},
"oa-tools-off": {
"expect_request": {
"require_tools": false,
"forbid_tools": true,
"require_tool_choice": false,
"parallel_tool_calls": null
},
"script": [
{ "text": "Chat-only OpenAI-lane answer (gate fixture): this lane offered no tools and none were used; the reply is plain text with finish reason stop." }
],
"phrasings": [
{ "id": "oa-tools-off-p1", "marker": "oa-gate tools-off probe", "prompt": "oa-gate tools-off probe: plain chat with no tools offered." }
]
},
"oa-single-tool": {
"script": [
{ "text": "Step 1: writing the note.",
"tool_calls": [
{ "name": "write_file",
"arguments": { "path": "openai-single-note.md", "content": "# Note (gate fixture, OpenAI lane)\n\nDeterministic single-tool body.\n" } }
] },
{ "text": "All set - openai-single-note.md is written with the fixture body. Nothing else was needed for this one." }
],
"phrasings": [
{ "id": "oa-single-p1", "marker": "oa-gate single tool note", "prompt": "oa-gate single tool note: save the fixture note to a file." },
{ "id": "oa-single-p2", "marker": "oa-gate one file please", "prompt": "oa-gate one file please: write the fixture note file." }
]
},
"oa-torture": {
"script": [
{ "tool_calls": [
{ "name": "write_file",
"arguments": { "path": "torture-note.md", "content": "Line 1 has \"double quotes\", 'singles', and a mid-line backslash \\ here.\nLine 2\thas a tab, a literal \\n two-char sequence, and a Windows path C:\\temp\\new.txt.\nLine 3 unicode: naïve café — 日本語 ✓ 🚀\nLine 4 JSON-in-string: {\"k\": \"v\", \"arr\": [1, 2], \"s\": \"nested \\\"deep\\\" quotes\"}\nLine 5 ends with a lone backslash \\" } }
] },
{ "text": "Torture round-trip complete: the payload with nested quotes, backslashes, newlines, tabs, and unicode survived exactly one encode and one decode." }
],
"phrasings": [
{ "id": "oa-torture-p1", "marker": "oa-gate torture probe", "prompt": "oa-gate torture probe: write the escaping torture file." }
]
},
"oa-parallel": {
"script": [
{ "text": "Step 1: two writes at once (parallel probe).",
"tool_calls": [
{ "name": "write_file", "arguments": { "path": "parallel-a.md", "content": "Parallel A (gate fixture).\n" } },
{ "name": "write_file", "arguments": { "path": "parallel-b.md", "content": "Parallel B (gate fixture).\n" } }
] },
{ "text": "Parallel probe complete: both tool results arrived and were paired correctly. A brain that instead rejects the double call must do so cleanly - that outcome is asserted brain-side, not here." }
],
"phrasings": [
{ "id": "oa-parallel-p1", "marker": "oa-gate parallel probe", "prompt": "oa-gate parallel probe: run the two-write parallel case." }
]
},
"oa-mission": {
"script": [
{ "text": "Step 1: drafting part one.",
"tool_calls": [
{ "name": "write_file", "arguments": { "path": "mission-part-1.md", "content": "Mission part 1 (gate fixture).\n" } }
] },
{ "text": "Step 2: drafting part two.",
"tool_calls": [
{ "name": "write_file", "arguments": { "path": "mission-part-2.md", "content": "Mission part 2 (gate fixture).\n" } }
] },
{ "text": "Mission complete: mission-part-1.md and mission-part-2.md are written; the loop ran two tool rounds and finished cleanly with finish reason stop." }
],
"phrasings": [
{ "id": "oa-mission-p1", "marker": "oa-gate mission probe", "prompt": "oa-gate mission probe: run the two-round mission." }
]
},
"oa-api-error": {
"expect_request": {
"require_tools": false,
"require_tool_choice": false,
"parallel_tool_calls": null
},
"script": [],
"phrasings": [
{ "id": "oa-err-400", "marker": "oa-gate error four hundred", "prompt": "oa-gate error four hundred: trigger the injected failure.",
"script": [ { "api_error": { "status": 400, "type": "invalid_request_error", "message": "gate-injected 400: request rejected by fixture", "code": "gate_injected" } } ] },
{ "id": "oa-err-429", "marker": "oa-gate error rate limit", "prompt": "oa-gate error rate limit: trigger the injected failure.",
"script": [ { "api_error": { "status": 429, "type": "rate_limit_error", "message": "gate-injected 429: rate limited by fixture", "code": "rate_limit_exceeded" } } ] },
{ "id": "oa-err-500", "marker": "oa-gate error five hundred", "prompt": "oa-gate error five hundred: trigger the injected failure.",
"script": [ { "api_error": { "status": 500, "type": "server_error", "message": "gate-injected 500: internal fixture error", "code": "gate_injected" } } ] },
{ "id": "oa-err-503", "marker": "oa-gate error unavailable", "prompt": "oa-gate error unavailable: trigger the injected failure.",
"script": [ { "api_error": { "status": 503, "type": "server_error", "message": "gate-injected 503: fixture overloaded", "code": "gate_injected" } } ] }
]
}
}
}
+409
View File
@@ -0,0 +1,409 @@
#!/usr/bin/env bash
# selftest.sh - proves stub-openai.py before any brain code exists.
# Drives the stub with curl through every scenario (plain, tools-off,
# single tool round-trip, escaping torture, parallel double-call,
# two-round mission, injected API errors, background, overrun), every
# validation rejection (dialect leaks, pairing, echo round-trip, scenario
# expectations), and all three hostile modes. Exit 0 = green.
set -u
cd "$(dirname "$0")" || exit 1
PY=python3
TMP="$(mktemp -d)"
PIDS=()
cleanup() {
for p in "${PIDS[@]:-}"; do kill -9 "$p" >/dev/null 2>&1; done
rm -rf "$TMP"
}
trap cleanup EXIT
PASS=0; FAIL=0
ok() { printf 'ok - %s\n' "$1"; PASS=$((PASS+1)); }
bad() { printf 'FAIL - %s\n' "$1"; FAIL=$((FAIL+1)); }
check() { # check <name> <cmd...> - pass if cmd exits 0; show output on fail
local name="$1"; shift
local out
if out="$("$@" 2>&1)"; then ok "$name"
else bad "$name"; [ -n "$out" ] && printf '%s\n' "$out" | sed 's/^/ /' | head -8
fi
}
freeport() { "$PY" -c 'import socket;s=socket.socket();s.bind(("127.0.0.1",0));print(s.getsockname()[1]);s.close()'; }
waithealth() {
local p="$1" i
for i in $(seq 1 60); do
curl -sf "http://127.0.0.1:$p/gate/health" >/dev/null 2>&1 && return 0
sleep 0.1
done
echo "stub on :$p never became healthy"; return 1
}
post() { # post <port> <bodyfile> <respfile> [extra curl args...] -> echoes http code
local port="$1" body="$2" resp="$3"; shift 3
curl -s -o "$resp" -w '%{http_code}' -H 'content-type: application/json' \
"$@" --data-binary @"$body" "http://127.0.0.1:$port/v1/chat/completions"
}
# ---- embedded helper: builds OpenAI-dialect bodies, asserts on responses ----
cat > "$TMP/helpers.py" <<'PYEOF'
import copy, json, sys
TOOLS = [
{"type": "function", "function": {
"name": "write_file", "description": "Write content to a file on disk.",
"parameters": {"type": "object",
"properties": {"path": {"type": "string"},
"content": {"type": "string"}},
"required": ["path", "content"]}}},
{"type": "function", "function": {
"name": "read_file", "description": "Read contents of a file from disk.",
"parameters": {"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]}}},
]
def dump(obj, out):
json.dump(obj, open(out, "w"), ensure_ascii=False)
def base(prompt, tools=True):
b = {"model": "gate-openai-model", "max_tokens": 1024,
"messages": [
{"role": "system", "content": "You are Neuron (gate fixture)."},
{"role": "user", "content": prompt}]}
if tools:
b["tools"] = copy.deepcopy(TOOLS)
b["tool_choice"] = "auto"
b["parallel_tool_calls"] = False
return b
def cmd_plain(out, prompt):
dump(base(prompt), out)
def cmd_notools(out, prompt):
dump(base(prompt, tools=False), out)
def cmd_mut(out, prompt, mutation):
b = base(prompt)
if mutation == "no-tool-choice":
del b["tool_choice"]
elif mutation == "ptc-true":
b["parallel_tool_calls"] = True
elif mutation == "top-system":
b["system"] = "You are Neuron."
elif mutation == "anth-tools":
b["tools"] = [{"name": "write_file", "description": "x",
"input_schema": {"type": "object", "properties": {}}}]
elif mutation == "anth-block":
b["messages"][1] = {"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_x", "content": "hi"},
{"type": "text", "text": prompt}]}
else:
raise SystemExit("unknown mutation " + mutation)
dump(b, out)
def cmd_chain(out, prompt, variant, *resps):
"""Build the next leg: echo each response's assistant turn and answer its
tool calls. `variant` applies to the LAST response only:
ok | no-tool-turn | wrong-id | only-first | double-encode | object-args"""
b = base(prompt)
for idx, p in enumerate(resps):
last = idx == len(resps) - 1
msg = json.load(open(p))["choices"][0]["message"]
tcs = msg.get("tool_calls")
if not tcs:
b["messages"].append({"role": "assistant",
"content": msg.get("content")})
continue
v = variant if last else "ok"
asst = {"role": "assistant", "content": msg.get("content"),
"tool_calls": copy.deepcopy(tcs)}
if v == "double-encode":
for tc in asst["tool_calls"]:
tc["function"]["arguments"] = json.dumps(
tc["function"]["arguments"])
if v == "object-args":
for tc in asst["tool_calls"]:
tc["function"]["arguments"] = json.loads(
tc["function"]["arguments"])
b["messages"].append(asst)
if v == "no-tool-turn":
continue
use = tcs[:1] if v == "only-first" else tcs
for tc in use:
tid = "call_bogus_123" if v == "wrong-id" else tc["id"]
b["messages"].append({"role": "tool", "tool_call_id": tid,
"content": "{\"ok\":true,\"bytes\":42}"})
dump(b, out)
def cmd_chk(resp, expr):
r = json.load(open(resp))
if not eval(expr, {"r": r, "json": json, "len": len, "str": str,
"isinstance": isinstance, "any": any, "all": all,
"sorted": sorted}):
print("assertion failed:", expr)
print("resp:", json.dumps(r, ensure_ascii=False)[:400])
raise SystemExit(1)
def cmd_torture(resp, scen):
r = json.load(open(resp))
tc = r["choices"][0]["message"]["tool_calls"][0]
raw = tc["function"]["arguments"]
assert isinstance(raw, str), "arguments must be a JSON-encoded string"
got = json.loads(raw)
exp = json.load(open(scen))["classes"]["oa-torture"]["script"][0]["tool_calls"][0]["arguments"]
assert got == exp, "decoded arguments != scripted torture payload"
content = got["content"]
for needle in ['"', "\\", "\n", "\t", "日本語", "naïve", "🚀"]:
assert needle in content, "missing torture needle %r" % needle
def cmd_notjson(path):
data = open(path, "rb").read()
assert data, "file empty - no partial body arrived"
try:
json.loads(data.decode("utf-8", "replace"))
except ValueError:
return
raise SystemExit("partial body unexpectedly parsed as complete JSON")
def cmd_pending(*paths):
ids = []
for p in paths:
c = json.load(open(p))["choices"][0]
assert c["finish_reason"] == "tool_calls", c["finish_reason"]
tc = c["message"]["tool_calls"][0]
assert tc["function"]["name"] == "write_file"
json.loads(tc["function"]["arguments"]) # must decode
ids.append(tc["id"])
assert len(set(ids)) == len(ids), "call ids not distinct: %r" % ids
def cmd_logcheck(path):
recs = [json.loads(l) for l in open(path) if l.strip()]
seqs = [r["seq"] for r in recs]
assert seqs == sorted(seqs) and len(set(seqs)) == len(seqs), "seq not monotonic"
kinds = {}
for r in recs:
kinds[r["kind"]] = kinds.get(r["kind"], 0) + 1
assert kinds.get("scenario", 0) >= 10, "too few scenario records: %r" % kinds
assert kinds.get("background", 0) >= 1, "no background record"
assert kinds.get("overrun", 0) >= 1, "no overrun record"
rejected = [r for r in recs if r["validation"] == "rejected"]
assert len(rejected) >= 10, "too few rejected records: %d" % len(rejected)
assert any(r["delivered"].get("tool_calls") == ["write_file"]
for r in recs), "no single write_file ground truth"
assert any(r["delivered"].get("tool_calls") == ["write_file", "write_file"]
for r in recs), "no parallel ground truth"
def main():
fn = globals()["cmd_" + sys.argv[1].replace("-", "_")]
fn(*sys.argv[2:])
if __name__ == "__main__":
main()
PYEOF
mk() { "$PY" "$TMP/helpers.py" "$@"; }
echo "=== stub-openai selftest ==="
# ---- normal mode ------------------------------------------------------------
PORT="$(freeport)"
"$PY" stub-openai.py --port "$PORT" --scenarios scenarios-openai.json \
--log "$TMP/req.jsonl" >"$TMP/stub.out" 2>&1 &
PIDS+=($!); disown
check "stub starts and answers /gate/health" waithealth "$PORT"
# 1. plain completion
mk plain "$TMP/plain.json" "oa-gate plain probe: explain the fixture topic simply."
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_plain.json")"
check "plain: HTTP 200" test "$code" = "200"
check "plain: chat.completion envelope, finish stop, real content" mk chk "$TMP/r_plain.json" \
'r["object"]=="chat.completion" and r["choices"][0]["finish_reason"]=="stop" and isinstance(r["choices"][0]["message"]["content"],str) and len(r["choices"][0]["message"]["content"])>40'
# 2. tools-off lane (chat-only request accepted, tool-bearing request refused)
mk notools "$TMP/toolsoff.json" "oa-gate tools-off probe: plain chat with no tools offered."
code="$(post "$PORT" "$TMP/toolsoff.json" "$TMP/r_toolsoff.json")"
check "tools-off: chat-only request -> 200" test "$code" = "200"
mk plain "$TMP/toolsoff_bad.json" "oa-gate tools-off probe: plain chat with no tools offered."
code="$(post "$PORT" "$TMP/toolsoff_bad.json" "$TMP/r_toolsoff_bad.json")"
check "tools-off negative: offering tools -> 400 gate_expect" \
bash -c "test $code = 400"
check "tools-off negative: reason names gate_expect" mk chk "$TMP/r_toolsoff_bad.json" \
'r["error"]["code"]=="gate_expect"'
# 3. dialect-leak rejections (the loud-failure contract)
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_leak_hdr.json" -H 'anthropic-version: 2023-06-01')"
check "leak: anthropic-version header -> 400" test "$code" = "400"
check "leak: header reason names the leak" mk chk "$TMP/r_leak_hdr.json" \
'r["error"]["code"]=="gate_dialect_leak" and "anthropic-version" in r["error"]["message"]'
mk mut "$TMP/leak_tools.json" "oa-gate plain probe: explain the fixture topic simply." anth-tools
code="$(post "$PORT" "$TMP/leak_tools.json" "$TMP/r_leak_tools.json")"
check "leak: input_schema tools -> 400 gate_dialect_leak" bash -c \
"test $code = 400"
check "leak: input_schema reason" mk chk "$TMP/r_leak_tools.json" \
'r["error"]["code"]=="gate_dialect_leak" and "input_schema" in r["error"]["message"]'
mk mut "$TMP/leak_sys.json" "oa-gate plain probe: explain the fixture topic simply." top-system
code="$(post "$PORT" "$TMP/leak_sys.json" "$TMP/r_leak_sys.json")"
check "leak: top-level system -> 400" test "$code" = "400"
mk mut "$TMP/leak_block.json" "oa-gate plain probe: explain the fixture topic simply." anth-block
code="$(post "$PORT" "$TMP/leak_block.json" "$TMP/r_leak_block.json")"
check "leak: Anthropic tool_result content block -> 400" test "$code" = "400"
# 4. scenario request expectations
mk mut "$TMP/no_tc.json" "oa-gate plain probe: explain the fixture topic simply." no-tool-choice
code="$(post "$PORT" "$TMP/no_tc.json" "$TMP/r_no_tc.json")"
check "expect: missing tool_choice -> 400" test "$code" = "400"
mk mut "$TMP/ptc.json" "oa-gate plain probe: explain the fixture topic simply." ptc-true
code="$(post "$PORT" "$TMP/ptc.json" "$TMP/r_ptc.json")"
check "expect: parallel_tool_calls true -> 400 (ADR-0005 pin)" test "$code" = "400"
# 5. single tool round-trip
ST_PROMPT="oa-gate single tool note: save the fixture note to a file."
mk plain "$TMP/st1.json" "$ST_PROMPT"
code="$(post "$PORT" "$TMP/st1.json" "$TMP/r_st1.json")"
check "single-tool leg1: HTTP 200" test "$code" = "200"
check "single-tool leg1: one write_file call, finish tool_calls, string args" mk chk "$TMP/r_st1.json" \
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==1 and r["choices"][0]["message"]["tool_calls"][0]["type"]=="function" and r["choices"][0]["message"]["tool_calls"][0]["function"]["name"]=="write_file" and isinstance(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"],str) and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="openai-single-note.md"'
mk chain "$TMP/st2.json" "$ST_PROMPT" ok "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2.json" "$TMP/r_st2.json")"
check "single-tool leg2: echo + tool turn -> 200 final text" test "$code" = "200"
check "single-tool leg2: final names the file, finish stop" mk chk "$TMP/r_st2.json" \
'r["choices"][0]["finish_reason"]=="stop" and "openai-single-note.md" in r["choices"][0]["message"]["content"]'
mk chain "$TMP/st2_no.json" "$ST_PROMPT" no-tool-turn "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2_no.json" "$TMP/r_st2_no.json")"
check "single-tool negative: echo without tool turn -> 400 gate_pairing" \
bash -c "test $code = 400"
check "single-tool negative: pairing reason" mk chk "$TMP/r_st2_no.json" \
'r["error"]["code"]=="gate_pairing"'
mk chain "$TMP/st2_wrong.json" "$ST_PROMPT" wrong-id "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2_wrong.json" "$TMP/r_st2_wrong.json")"
check "single-tool negative: wrong tool_call_id -> 400" test "$code" = "400"
mk chain "$TMP/st2_obj.json" "$ST_PROMPT" object-args "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2_obj.json" "$TMP/r_st2_obj.json")"
check "single-tool negative: arguments echoed as object -> 400 shape" \
bash -c "test $code = 400"
check "single-tool negative: shape reason names STRING" mk chk "$TMP/r_st2_obj.json" \
'r["error"]["code"]=="gate_tool_call_shape" and "STRING" in r["error"]["message"]'
# 6. escaping torture (the two-escaper trap, spec section 6)
T_PROMPT="oa-gate torture probe: write the escaping torture file."
mk plain "$TMP/t1.json" "$T_PROMPT"
code="$(post "$PORT" "$TMP/t1.json" "$TMP/r_t1.json")"
check "torture leg1: HTTP 200" test "$code" = "200"
check "torture leg1: arguments decode to the exact nasty payload" \
mk torture "$TMP/r_t1.json" scenarios-openai.json
mk chain "$TMP/t2.json" "$T_PROMPT" ok "$TMP/r_t1.json"
code="$(post "$PORT" "$TMP/t2.json" "$TMP/r_t2.json")"
check "torture leg2: faithful echo -> 200 final" test "$code" = "200"
mk chain "$TMP/t2_dbl.json" "$T_PROMPT" double-encode "$TMP/r_t1.json"
code="$(post "$PORT" "$TMP/t2_dbl.json" "$TMP/r_t2_dbl.json")"
check "torture negative: double-encoded echo -> 400" test "$code" = "400"
check "torture negative: reason names the two-escaper trap" mk chk "$TMP/r_t2_dbl.json" \
'r["error"]["code"]=="gate_echo_mismatch" and "two-escaper" in r["error"]["message"]'
# 7. parallel double-call
P_PROMPT="oa-gate parallel probe: run the two-write parallel case."
mk plain "$TMP/p1.json" "$P_PROMPT"
code="$(post "$PORT" "$TMP/p1.json" "$TMP/r_p1.json")"
check "parallel leg1: TWO tool_calls, distinct ids" mk chk "$TMP/r_p1.json" \
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==2 and r["choices"][0]["message"]["tool_calls"][0]["id"]!=r["choices"][0]["message"]["tool_calls"][1]["id"]'
mk chain "$TMP/p2.json" "$P_PROMPT" ok "$TMP/r_p1.json"
code="$(post "$PORT" "$TMP/p2.json" "$TMP/r_p2.json")"
check "parallel leg2: both results -> 200 final" test "$code" = "200"
mk chain "$TMP/p2_one.json" "$P_PROMPT" only-first "$TMP/r_p1.json"
code="$(post "$PORT" "$TMP/p2_one.json" "$TMP/r_p2_one.json")"
check "parallel negative: answering only one call -> 400 pairing" test "$code" = "400"
# 8. two-round mission (loop continuation + step indexing)
M_PROMPT="oa-gate mission probe: run the two-round mission."
mk plain "$TMP/m1.json" "$M_PROMPT"
code="$(post "$PORT" "$TMP/m1.json" "$TMP/r_m1.json")"
check "mission leg1: part-1 tool call" mk chk "$TMP/r_m1.json" \
'json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-1.md"'
mk chain "$TMP/m2.json" "$M_PROMPT" ok "$TMP/r_m1.json"
code="$(post "$PORT" "$TMP/m2.json" "$TMP/r_m2.json")"
check "mission leg2: part-2 tool call (step indexed by assistant count)" mk chk "$TMP/r_m2.json" \
'r["choices"][0]["finish_reason"]=="tool_calls" and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-2.md"'
mk chain "$TMP/m3.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json"
code="$(post "$PORT" "$TMP/m3.json" "$TMP/r_m3.json")"
check "mission leg3: final text, finish stop" mk chk "$TMP/r_m3.json" \
'r["choices"][0]["finish_reason"]=="stop" and "Mission complete" in r["choices"][0]["message"]["content"]'
mk chain "$TMP/m4.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json" "$TMP/r_m3.json"
code="$(post "$PORT" "$TMP/m4.json" "$TMP/r_m4.json")"
check "mission overrun: past-script request -> GATE-SCRIPT-EXHAUSTED" mk chk "$TMP/r_m4.json" \
'r["choices"][0]["message"]["content"].startswith("GATE-SCRIPT-EXHAUSTED")'
# 9. injected API errors (OpenAI error envelope)
for want in 400 429 500 503; do
case "$want" in
400) marker="four hundred";; 429) marker="rate limit";;
500) marker="five hundred";; 503) marker="unavailable";;
esac
mk plain "$TMP/e_$want.json" "oa-gate error $marker: trigger the injected failure."
code="$(post "$PORT" "$TMP/e_$want.json" "$TMP/r_e_$want.json")"
check "api-error $want: status returned" test "$code" = "$want"
check "api-error $want: OpenAI error envelope" mk chk "$TMP/r_e_$want.json" \
'isinstance(r["error"]["message"],str) and "gate-injected" in r["error"]["message"] and isinstance(r["error"]["type"],str)'
done
# 10. background (unmatched) request
mk plain "$TMP/bg.json" "hello there, just a boot probe with no marker"
code="$(post "$PORT" "$TMP/bg.json" "$TMP/r_bg.json")"
check "background: unmatched prompt -> benign ok" mk chk "$TMP/r_bg.json" \
'r["choices"][0]["message"]["content"]=="ok"'
# 11. ground-truth log invariants
check "ground-truth JSONL log invariants" mk logcheck "$TMP/req.jsonl"
# 12. production-port refusal
rc=0
"$PY" stub-openai.py --port 7770 --scenarios scenarios-openai.json \
--log "$TMP/never.jsonl" >/dev/null 2>&1 || rc=$?
check "refuses production port 7770" test "$rc" -ne 0
# ---- hostile mode: black-hole ----------------------------------------------
BH="$(freeport)"
"$PY" stub-openai.py --port "$BH" --log "$TMP/bh.jsonl" --mode black-hole \
>/dev/null 2>&1 &
PIDS+=($!); disown
check "black-hole: healthy" waithealth "$BH"
rc=0
curl -s -o /dev/null --max-time 3 -H 'content-type: application/json' \
--data-binary @"$TMP/plain.json" \
"http://127.0.0.1:$BH/v1/chat/completions" || rc=$?
check "black-hole: client times out (curl rc 28)" test "$rc" -eq 28
check "black-hole: health still answers during the hang" \
curl -sf --max-time 2 "http://127.0.0.1:$BH/gate/health"
# ---- hostile mode: mid-body-drop -------------------------------------------
MD="$(freeport)"
"$PY" stub-openai.py --port "$MD" --log "$TMP/md.jsonl" --mode mid-body-drop \
>/dev/null 2>&1 &
PIDS+=($!); disown
check "mid-body-drop: healthy" waithealth "$MD"
rc=0
curl -s --max-time 5 -o "$TMP/half.json" -H 'content-type: application/json' \
--data-binary @"$TMP/plain.json" \
"http://127.0.0.1:$MD/v1/chat/completions" || rc=$?
check "mid-body-drop: transfer fails (curl rc $rc)" test "$rc" -ne 0
check "mid-body-drop: partial body is not parseable JSON" mk notjson "$TMP/half.json"
# ---- hostile mode: tool-pending-forever ------------------------------------
TP="$(freeport)"
"$PY" stub-openai.py --port "$TP" --log "$TMP/tp.jsonl" \
--mode tool-pending-forever >/dev/null 2>&1 &
PIDS+=($!); disown
check "tool-pending-forever: healthy" waithealth "$TP"
for i in 1 2 3; do
code="$(post "$TP" "$TMP/plain.json" "$TMP/r_tp$i.json")"
check "tool-pending-forever: request $i -> 200" test "$code" = "200"
done
check "tool-pending-forever: three FRESH tool_calls, distinct ids" \
mk pending "$TMP/r_tp1.json" "$TMP/r_tp2.json" "$TMP/r_tp3.json"
check "tool-pending-forever: /gate/stats counts 3 chat hits" \
bash -c "curl -sf http://127.0.0.1:$TP/gate/stats | grep -q '\"chat_hits\": 3'"
# ---- summary ----------------------------------------------------------------
echo
echo "selftest: $PASS passed, $FAIL failed"
if [ "$FAIL" -ne 0 ]; then
echo "SELFTEST RED"
exit 1
fi
echo "SELFTEST GREEN (stub-openai gate scaffolding verified)"
+652
View File
@@ -0,0 +1,652 @@
#!/usr/bin/env python3
"""stub-openai.py - deterministic local stand-in for an OpenAI-format
/v1/chat/completions provider, for the soul-openai-tools-v2 gate
(docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md). No API key, no network,
no model.
Sibling of gate9's stub-llm.py (Anthropic dialect, _wt-beta-round9/scripts/
gate9/): same scenario mechanism (marker matching, assistant-count step
indexing, ground-truth JSONL log, prod-port refusal), different wire.
Staging home is tests/gate-openai/ in _wt-openai-tools; folds into
scripts/gate9/ after round 9 merges (see README.md).
WHAT IT DOES
* Serves POST /v1/chat/completions on 127.0.0.1 only (OpenAI dialect).
* VALIDATES every request - this is the gate's discriminator, built
BEFORE the brain-side El code exists so dialect leakage fails loudly:
- Anthropic tells are 400 code=gate_dialect_leak: `anthropic-version`
header; top-level `system` / `stop_sequences` / `max_tokens_to_sample`
/ `anthropic_version`; `input_schema` inside a tool entry; Anthropic
content blocks (tool_use / tool_result / server_tool_use / ...).
- tools[] must be OpenAI-shaped {type:"function", function:{name,
description, parameters}} with unique names -> 400 gate_tools_shape.
- assistant tool_calls echoes must be {id, type:"function",
function:{name, arguments:<JSON-encoded STRING>}}; a decoded-object
`arguments` is a wire bug -> 400 gate_tool_call_shape.
- every assistant tool_calls turn must be answered by role:"tool"
messages covering EVERY tool_call_id, immediately following;
unknown / duplicate / missing ids -> 400 gate_pairing.
- echoed `arguments` for gate-issued call ids (call_gate_*) are
recomputed from the script and compared after ONE json decode ->
400 gate_echo_mismatch. This is the two-escaper-trap discriminator
named in the spec's security model (section 6).
- scenario-level request expectations from scenarios-openai.json
(tools offered, OpenAI-shaped tool_choice, parallel_tool_calls
pinned false per ADR-0005) -> 400 gate_expect.
* Answers with SCRIPTED responses: plain text (finish_reason "stop"),
tool calls (finish_reason "tool_calls", arguments JSON-encoded, incl. a
nested-quote/escaping torture payload and a parallel two-call case), and
API-error injection (OpenAI error envelope). Scenario is selected by
scanning user-message text (newest first) for a registered marker
substring; the step index is the number of assistant messages already in
the request (stateless replay - resumes index correctly by construction).
* Writes a ground-truth JSONL log (--log): one record per request with the
validation verdict, matched scenario/step, and exactly which tool calls
were delivered. Gate assertions compare the brain's claims against THIS
log - truth, not narration.
* Unmatched requests (boot probes, awareness chatter) get a benign "ok"
text response, logged kind=background, never counted as ground truth.
* HOSTILE MODES (--mode) on the same file:
black-hole accept + read the request, never respond;
mid-body-drop send half a JSON body, then abort the socket;
tool-pending-forever every request gets a FRESH tool_call
(finish_reason "tool_calls"), forever - tests
the agentic loop's iteration cap; count the
brain's round-trips via GET /gate/stats.
usage: stub-openai.py --port P --scenarios scenarios-openai.json \
--log requests.jsonl [--mode MODE]
Listens on 127.0.0.1 only. Refuses production ports 7770/7779/17779.
"""
import argparse
import itertools
import json
import socket
import struct
import threading
import time
import uuid
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
STATE = {"scenarios": None, "log_path": None, "lock": threading.Lock(),
"seq": 0, "mode": "normal", "chat_hits": 0}
_PENDING_SEQ = itertools.count(1)
ANTHROPIC_TOP_KEYS = ("system", "stop_sequences", "max_tokens_to_sample",
"anthropic_version")
ANTHROPIC_BLOCK_TYPES = {"tool_use", "tool_result", "server_tool_use",
"web_search_tool_result", "thinking",
"redacted_thinking"}
DEFAULT_EXPECT = {"require_tools": True, "require_tool_choice": True,
"parallel_tool_calls": False, "forbid_tools": False}
# ---------------------------------------------------------------- loading ----
def load_scenarios(path):
cfg = json.load(open(path))
defaults = dict(DEFAULT_EXPECT)
defaults.update(cfg.get("defaults", {}).get("expect_request", {}))
marker_map = [] # (marker_lower, cname, pid)
scripts = {} # cname or cname/pid -> expanded script
pid_map = {} # pid -> cname (for call_gate_* id -> script lookup)
expects = {} # cname -> merged expect_request
for cname, cls in cfg["classes"].items():
scripts[cname] = expand_script(cls.get("script", []))
exp = dict(defaults)
exp.update(cls.get("expect_request", {}))
expects[cname] = exp
for ph in cls["phrasings"]:
if ph.get("script") is not None:
scripts[cname + "/" + ph["id"]] = expand_script(ph["script"])
marker_map.append((ph["marker"].lower(), cname, ph["id"]))
pid_map[ph["id"]] = cname
return {"cfg": cfg, "marker_map": marker_map, "scripts": scripts,
"pid_map": pid_map, "expects": expects}
def expand_script(script):
"""Same repeat-expansion contract as gate9's stub-llm.py ({N}/{NN})."""
out = []
for step in script:
if "repeat" in step:
for n in range(1, step["repeat"] + 1):
t = {k: v for k, v in step.items() if k != "repeat"}
out.append(json.loads(json.dumps(t)
.replace("{NN}", "%02d" % n)
.replace("{N}", str(n))))
else:
out.append(step)
return out
# ------------------------------------------------------------- validation ----
def _rej(message, code):
return {"status": 400, "message": message, "code": code}
def validate_dialect(headers, req):
"""Universal checks - run on EVERY request, scenario-matched or not.
Anything Anthropic-shaped on this lane means the brain's translator
leaked; the whole point is that it fails loudly, here, with a reason."""
if headers.get("anthropic-version"):
return _rej("anthropic-version header on the OpenAI lane: this "
"request was built by the Anthropic dialect path",
"gate_dialect_leak")
for k in ANTHROPIC_TOP_KEYS:
if k in req:
return _rej("top-level `%s` is Anthropic dialect; the OpenAI "
"dialect has no such field (system prompt goes in "
"messages[0])" % k, "gate_dialect_leak")
tools = req.get("tools")
if tools is not None:
if not isinstance(tools, list):
return _rej("`tools` must be an array", "gate_tools_shape")
names = []
for i, t in enumerate(tools):
if not isinstance(t, dict):
return _rej("tools[%d] is not an object" % i,
"gate_tools_shape")
if "input_schema" in t or (isinstance(t.get("function"), dict)
and "input_schema" in t["function"]):
return _rej("tools[%d] carries `input_schema` (Anthropic "
"dialect); OpenAI dialect wants "
"function.parameters" % i, "gate_dialect_leak")
if t.get("type") != "function":
return _rej("tools[%d].type must be \"function\", got %r"
% (i, t.get("type")), "gate_tools_shape")
fn = t.get("function")
if not isinstance(fn, dict):
return _rej("tools[%d].function missing" % i,
"gate_tools_shape")
if not isinstance(fn.get("name"), str) or not fn["name"]:
return _rej("tools[%d].function.name missing/empty" % i,
"gate_tools_shape")
if not isinstance(fn.get("description"), str) or not fn["description"]:
return _rej("tools[%d].function.description missing/empty" % i,
"gate_tools_shape")
if not isinstance(fn.get("parameters"), dict):
return _rej("tools[%d].function.parameters missing (JSON "
"Schema object expected)" % i, "gate_tools_shape")
names.append(fn["name"])
if len(names) != len(set(names)):
return _rej("tools: tool names must be unique", "gate_tools_shape")
msgs = req.get("messages")
if not isinstance(msgs, list) or not msgs:
return _rej("`messages` must be a non-empty array",
"gate_messages_shape")
for i, m in enumerate(msgs):
if not isinstance(m, dict):
return _rej("messages[%d] is not an object" % i,
"gate_messages_shape")
c = m.get("content")
if isinstance(c, list):
for j, b in enumerate(c):
if isinstance(b, dict) and b.get("type") in ANTHROPIC_BLOCK_TYPES:
return _rej("messages[%d].content[%d] is an Anthropic "
"`%s` block; the OpenAI dialect uses "
"tool_calls / role:\"tool\" messages"
% (i, j, b.get("type")), "gate_dialect_leak")
if m.get("role") == "tool":
if not isinstance(m.get("tool_call_id"), str) or not m["tool_call_id"]:
return _rej("messages[%d]: role \"tool\" requires a "
"`tool_call_id`" % i, "gate_messages_shape")
if "content" not in m:
return _rej("messages[%d]: role \"tool\" requires `content`"
% i, "gate_messages_shape")
if m.get("role") == "assistant" and m.get("tool_calls") is not None:
tcs = m["tool_calls"]
if not isinstance(tcs, list) or not tcs:
return _rej("messages[%d].tool_calls must be a non-empty "
"array" % i, "gate_tool_call_shape")
for j, tc in enumerate(tcs):
if not isinstance(tc, dict) or tc.get("type") != "function":
return _rej("messages[%d].tool_calls[%d].type must be "
"\"function\"" % (i, j), "gate_tool_call_shape")
if not isinstance(tc.get("id"), str) or not tc["id"]:
return _rej("messages[%d].tool_calls[%d].id missing"
% (i, j), "gate_tool_call_shape")
fn = tc.get("function")
if not isinstance(fn, dict) or not isinstance(fn.get("name"), str):
return _rej("messages[%d].tool_calls[%d].function.name "
"missing" % (i, j), "gate_tool_call_shape")
if not isinstance(fn.get("arguments"), str):
return _rej("messages[%d].tool_calls[%d].function."
"arguments must be a JSON-encoded STRING, "
"got %s" % (i, j,
type(fn.get("arguments")).__name__),
"gate_tool_call_shape")
return None
def validate_pairing(msgs):
"""OpenAI pairing rule: every assistant tool_calls turn must be followed
immediately by role:"tool" messages answering every tool_call_id."""
open_ids, open_at = set(), None
for i, m in enumerate(msgs):
role = m.get("role")
if role == "tool":
tid = m.get("tool_call_id")
if open_at is None:
return _rej("messages[%d]: role \"tool\" message with no "
"preceding assistant tool_calls turn "
"(tool_call_id=%s)" % (i, tid), "gate_pairing")
if tid not in open_ids:
return _rej("messages[%d]: tool message answers unknown or "
"already-answered tool_call_id %s" % (i, tid),
"gate_pairing")
open_ids.discard(tid)
continue
if open_ids:
return _rej("messages[%d]: assistant tool_calls not fully "
"answered before messages[%d]; missing tool "
"responses for: %s" % (open_at, i, sorted(open_ids)),
"gate_pairing")
open_ids, open_at = set(), None
if role == "assistant" and m.get("tool_calls"):
ids = [tc.get("id") for tc in m["tool_calls"]]
open_ids, open_at = set(ids), i
if open_ids:
return _rej("messages[%d]: assistant tool_calls at end of thread "
"without tool responses for: %s"
% (open_at, sorted(open_ids)), "gate_pairing")
return None
def validate_echo_args(msgs, loaded):
"""Ground-truth round-trip check: for every echoed gate-issued call id,
recompute the arguments this stub originally sent from the script and
require one json decode to reproduce them exactly. Catches the
two-escaper trap (spec section 6) deterministically."""
if not loaded:
return None
for i, m in enumerate(msgs):
if m.get("role") != "assistant":
continue
for tc in m.get("tool_calls") or []:
tid = tc.get("id", "")
if not tid.startswith("call_gate_"):
continue
rest = tid[len("call_gate_"):]
try:
pid, s_part, k_part = rest.rsplit("_", 2)
step_idx, k = int(s_part[1:]), int(k_part)
except (ValueError, IndexError):
continue
cname = loaded["pid_map"].get(pid)
if cname is None:
continue
script = (loaded["scripts"].get(cname + "/" + pid)
or loaded["scripts"].get(cname) or [])
if step_idx >= len(script):
continue
calls = script[step_idx].get("tool_calls") or []
if k >= len(calls):
continue
expected = calls[k]
fn = tc.get("function") or {}
if fn.get("name") != expected["name"]:
return _rej("messages[%d]: echoed tool name %r != issued %r "
"for %s" % (i, fn.get("name"), expected["name"],
tid), "gate_echo_mismatch")
try:
got = json.loads(fn.get("arguments", ""))
except ValueError:
return _rej("messages[%d]: echoed arguments for %s are not "
"valid JSON after one decode (truncated or "
"half-escaped?)" % (i, tid), "gate_echo_mismatch")
if got != expected["arguments"]:
hint = (" (decoded to a string, not an object: "
"double-encoded - the two-escaper trap)"
if isinstance(got, str) else "")
return _rej("messages[%d]: echoed arguments for %s do not "
"round-trip to the issued payload%s"
% (i, tid, hint), "gate_echo_mismatch")
return None
def validate_expect(req, exp):
"""Scenario-level request expectations (scenarios-openai.json)."""
tools = req.get("tools") or []
if exp.get("forbid_tools") and tools:
return _rej("this scenario is chat-only: no `tools` may be offered "
"on it", "gate_expect")
if exp.get("require_tools") and not tools:
return _rej("scenario expects a `tools` array to be offered (the "
"agentic lane must advertise its tools)", "gate_expect")
if exp.get("require_tool_choice"):
tc = req.get("tool_choice")
ok = tc in ("auto", "none", "required") or (
isinstance(tc, dict) and tc.get("type") == "function"
and isinstance(tc.get("function"), dict)
and tc["function"].get("name"))
if not ok:
return _rej("scenario expects an OpenAI-shaped `tool_choice`, "
"got %r" % (tc,), "gate_expect")
want_ptc = exp.get("parallel_tool_calls", None)
if want_ptc is not None:
if "parallel_tool_calls" not in req:
return _rej("scenario expects explicit `parallel_tool_calls` "
"(ADR-0005: must be pinned false on the wire)",
"gate_expect")
if req["parallel_tool_calls"] != want_ptc:
return _rej("scenario expects parallel_tool_calls=%s, got %s"
% (json.dumps(want_ptc),
json.dumps(req["parallel_tool_calls"])),
"gate_expect")
return None
# --------------------------------------------------------- scenario match ----
def extract_user_texts_newest_first(msgs):
texts = []
for m in reversed(msgs):
if not isinstance(m, dict) or m.get("role") != "user":
continue
c = m.get("content")
if isinstance(c, str):
texts.append(c)
elif isinstance(c, list):
for b in c:
if isinstance(b, dict) and b.get("type") == "text":
texts.append(b.get("text", ""))
return texts
def match_scenario(loaded, msgs):
for text in extract_user_texts_newest_first(msgs):
tl = text.lower()
for marker, cname, pid in loaded["marker_map"]:
if marker in tl:
return cname, pid
return None, None
# ------------------------------------------------------------- rendering ----
def completion_envelope(msg, finish, model, usage=(100, 100)):
return {"id": "chatcmpl-gate-" + uuid.uuid4().hex[:12],
"object": "chat.completion", "created": int(time.time()),
"model": model,
"choices": [{"index": 0, "message": msg,
"finish_reason": finish, "logprobs": None}],
"usage": {"prompt_tokens": usage[0],
"completion_tokens": usage[1],
"total_tokens": usage[0] + usage[1]}}
def text_completion(text, model):
return completion_envelope({"role": "assistant", "content": text},
"stop", model, usage=(1, 1))
def pending_body(seq, model):
args = json.dumps({"path": "never-%04d.md" % seq,
"content": "this run never completes"})
msg = {"role": "assistant", "content": None,
"tool_calls": [{"id": "call_hostile_pending_%04d" % seq,
"type": "function",
"function": {"name": "write_file",
"arguments": args}}]}
return completion_envelope(msg, "tool_calls", model, usage=(1, 1))
def render_step(step, cname, pid, step_idx, model):
"""Returns (http_status, body_dict, delivered) - delivered is ground
truth for the JSONL log."""
delivered = {"tool_calls": [], "finish_reason": None, "api_error": None}
if "api_error" in step:
e = step["api_error"]
delivered["api_error"] = e["status"]
return (e["status"],
{"error": {"message": e["message"],
"type": e.get("type", "server_error"),
"param": None, "code": e.get("code")}},
delivered)
msg = {"role": "assistant"}
finish = "stop"
if step.get("tool_calls"):
tcs = []
for k, call in enumerate(step["tool_calls"]):
tid = "call_gate_%s_s%d_%d" % (pid, step_idx, k)
tcs.append({"id": tid, "type": "function",
"function": {"name": call["name"],
"arguments": json.dumps(
call["arguments"],
ensure_ascii=False)}})
delivered["tool_calls"].append(call["name"])
msg["tool_calls"] = tcs
msg["content"] = step.get("text") # null when no narration, like real
finish = "tool_calls"
else:
msg["content"] = step["text"]
delivered["finish_reason"] = finish
return 200, completion_envelope(msg, finish, model), delivered
# ------------------------------------------------------------------ log ------
def log_record(rec):
with STATE["lock"]:
STATE["seq"] += 1
rec["seq"] = STATE["seq"]
with open(STATE["log_path"], "a") as f:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
# ---------------------------------------------------------------- server -----
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def _send_json(self, status, obj):
body = json.dumps(obj, ensure_ascii=False).encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def _send_error(self, verdict):
self._send_json(verdict["status"],
{"error": {"message": verdict["message"],
"type": "invalid_request_error",
"param": None, "code": verdict["code"]}})
def _drop_mid_body(self):
"""Valid 200 headers, half the promised body, then a socket abort
(same SO_LINGER teardown as gate9's mid-body-drop-brain.py)."""
full = json.dumps(text_completion(
"This reply will never finish arriving because the connection "
"dies in the middle of the body, which is exactly the point of "
"this hostile fixture.", "hostile-mid-drop")).encode("utf-8")
half = full[: len(full) // 2]
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(full))) # promises more
self.end_headers()
self.wfile.write(half)
self.wfile.flush()
try:
self.connection.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER,
struct.pack("ii", 1, 0))
self.connection.shutdown(socket.SHUT_RDWR)
except OSError:
pass
self.close_connection = True
def do_GET(self):
path = self.path.split("?")[0]
if path == "/gate/health":
self._send_json(200, {"ok": True, "mode": STATE["mode"]})
elif path == "/gate/stats":
with STATE["lock"]:
self._send_json(200, {"mode": STATE["mode"],
"chat_hits": STATE["chat_hits"]})
else:
self._send_json(404, {"error": {"message": "not found",
"type": "invalid_request_error",
"param": None,
"code": "unknown_route"}})
def do_POST(self):
n = int(self.headers.get("Content-Length") or 0)
raw = self.rfile.read(n)
mode = STATE["mode"]
rec = {"ts": time.time(), "path": self.path, "mode": mode,
"kind": "background", "scenario_class": None, "phrasing": None,
"step": None, "n_messages": 0, "n_assistant": 0,
"validation": "ok", "validation_detail": None,
"delivered": {"tool_calls": [], "finish_reason": None,
"api_error": None},
"http_status": 200}
if self.path.split("?")[0] != "/v1/chat/completions":
rec.update(kind="wrong_path", http_status=404)
log_record(rec)
self._send_json(404, {"error": {
"message": "no such route: %s" % self.path,
"type": "invalid_request_error", "param": None,
"code": "unknown_route"}})
return
with STATE["lock"]:
STATE["chat_hits"] += 1
# ---- hostile modes: behavior first, no validation ----------------
if mode == "black-hole":
rec.update(kind="hostile", http_status=None)
log_record(rec)
threading.Event().wait() # hold the socket open forever
return
if mode == "mid-body-drop":
rec.update(kind="hostile", http_status=200)
log_record(rec)
self._drop_mid_body()
return
if mode == "tool-pending-forever":
seq = next(_PENDING_SEQ)
rec.update(kind="hostile",
delivered={"tool_calls": ["write_file"],
"finish_reason": "tool_calls",
"api_error": None})
log_record(rec)
self._send_json(200, pending_body(seq, "gate-openai-model"))
return
# ---- normal mode -------------------------------------------------
try:
req = json.loads(raw)
except ValueError as exc:
# DIAGNOSTIC CAPTURE (2026-08-06): an unparseable body used to be recorded as
# a bare "bad_json" with the bytes thrown away, which made an intermittent
# failure impossible to root-cause — you cannot fix what you did not keep.
# Dump the raw body next to the log, and record exactly where the parser gave
# up plus the offending byte, so one occurrence is enough to diagnose.
dump_path = "%s.badbody.%s" % (STATE.get("log_path", "/tmp/stub-openai"),
rec.get("seq", "x"))
try:
data = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
with open(dump_path, "wb") as fh:
fh.write(data)
except Exception as dump_exc:
dump_path = "(dump failed: %s)" % dump_exc
pos = getattr(exc, "pos", None)
near = ""
byte_repr = ""
if isinstance(pos, int):
blob = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
near = blob[max(0, pos - 60):pos + 60].decode("utf-8", "replace")
if 0 <= pos < len(blob):
byte_repr = "0x%02x" % blob[pos]
rec.update(kind="bad_json", validation="rejected",
validation_detail="request body is not valid JSON: %s" % exc,
http_status=400, raw_len=len(raw), raw_dump=dump_path,
err_pos=pos, err_byte=byte_repr, err_near=near)
log_record(rec)
self._send_error(_rej("request body is not valid JSON",
"bad_json"))
return
msgs = req.get("messages") or []
rec["n_messages"] = len(msgs)
rec["n_assistant"] = sum(1 for m in msgs if isinstance(m, dict)
and m.get("role") == "assistant")
loaded = STATE["scenarios"]
cname, pid = match_scenario(loaded, msgs)
if cname:
rec.update(kind="scenario", scenario_class=cname, phrasing=pid)
# Wire-level validation runs for EVERY request, scenario or not.
verdict = (validate_dialect(self.headers, req)
or validate_pairing([m for m in msgs
if isinstance(m, dict)])
or validate_echo_args(msgs, loaded))
if verdict:
rec.update(validation="rejected",
validation_detail=verdict["message"],
http_status=verdict["status"])
log_record(rec)
self._send_error(verdict)
return
model = req.get("model", "gate-openai-model")
if not cname:
log_record(rec)
self._send_json(200, text_completion("ok", model))
return
script = (loaded["scripts"].get(cname + "/" + pid)
or loaded["scripts"][cname])
step_idx = rec["n_assistant"]
if step_idx >= len(script):
rec.update(kind="overrun", step=step_idx)
log_record(rec)
self._send_json(200, text_completion(
"GATE-SCRIPT-EXHAUSTED %s step %d" % (pid, step_idx), model))
return
step = script[step_idx]
exp = dict(loaded["expects"][cname])
exp.update(step.get("expect_request", {}))
verdict = validate_expect(req, exp)
if verdict:
rec.update(step=step_idx, validation="rejected",
validation_detail=verdict["message"],
http_status=verdict["status"])
log_record(rec)
self._send_error(verdict)
return
status, body, delivered = render_step(step, cname, pid, step_idx,
model)
rec.update(step=step_idx, delivered=delivered, http_status=status)
log_record(rec)
self._send_json(status, body)
def log_message(self, *a):
pass
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--port", type=int, required=True)
ap.add_argument("--scenarios",
help="scenarios-openai.json (required in normal mode)")
ap.add_argument("--log", required=True)
ap.add_argument("--mode", default="normal",
choices=["normal", "black-hole", "mid-body-drop",
"tool-pending-forever"])
args = ap.parse_args()
if args.port in (7770, 7779, 17779):
raise SystemExit("stub-openai: refusing production Neuron port")
if args.mode == "normal" and not args.scenarios:
raise SystemExit("stub-openai: --scenarios is required in normal mode")
STATE["mode"] = args.mode
STATE["scenarios"] = (load_scenarios(args.scenarios)
if args.scenarios else None)
STATE["log_path"] = args.log
open(args.log, "w").close()
n_markers = (len(STATE["scenarios"]["marker_map"])
if STATE["scenarios"] else 0)
print("stub-openai [%s]: 127.0.0.1:%d /v1/chat/completions "
"(%d markers registered, log=%s)"
% (args.mode, args.port, n_markers, args.log), flush=True)
ThreadingHTTPServer(("127.0.0.1", args.port), Handler).serve_forever()
if __name__ == "__main__":
main()
+148
View File
@@ -0,0 +1,148 @@
#!/usr/bin/env bash
# run-el-test.sh — build and RUN one engine test (tests/*.el), printing its assertions.
#
# WHY THIS EXISTS (2026-08-06): the engine's tests/*.el files were never runnable from the
# tree. `elc` is a COMPILER — it emits C to stdout and exits; it does not execute anything.
# So "the tests" could only ever be read, not run, and a signature change could silently
# break them (exactly what happened when bridge_save gained its `wire` argument). This
# script closes that: emit the test to C, link it against the engine modules, execute it.
#
# HOW IT WORKS
# 1. elc <test>.el -> C on stdout (the test file's `main` + prototypes)
# 2. elb (once, cached) -> per-module C for the whole engine into a scratch dir
# 3. cc test.c + all modules EXCEPT soul.c (soul.c owns the real `main`) + the runtime
# 4. run it
#
# The test C references only the engine functions it actually calls, so there are no
# duplicate-symbol collisions with the module objects.
#
# RUNTIME: the REPO-PINNED vendor/el-runtime (NOT ~/el-sdk/el_runtime.c — that June build
# is missing builtins August code calls: engram_wm_count, engram_wm_top_json,
# http_delete_json, http_serve_async; linking against it fails with "symbol(s) not found").
#
# USAGE
# tests/run-el-test.sh tests/test_bridge_serialization.el # one test
# tests/run-el-test.sh --all # every tests/test_*.el
# REBUILD=1 tests/run-el-test.sh ... # force module regeneration
#
# Tests that need a live API key / running soul (see each file's header) will report their
# own skips or failures — this runner does not fake them.
set -uo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT" || exit 2
ELC="${ELC:-$HOME/el-sdk/elc}"
ELB="${ELB:-$HOME/Development/el-sdk/bin/elb}"
RUNTIME_DIR="${RUNTIME_DIR:-$REPO_ROOT/vendor/el-runtime/v1.0.0-20260501}"
SCRATCH="${SCRATCH:-/tmp/el-test-$(basename "$REPO_ROOT")}"
MODDIR="$SCRATCH/modules"
OPENSSL_INC="${OPENSSL_INC:-/opt/homebrew/opt/openssl@3/include}"
OPENSSL_LIB="${OPENSSL_LIB:-/opt/homebrew/opt/openssl@3/lib}"
for req in "$ELC" "$ELB" "$RUNTIME_DIR/el_runtime.c"; do
[ -e "$req" ] || { echo "run-el-test: missing required input: $req" >&2; exit 2; }
done
mkdir -p "$MODDIR" || exit 2
# ── Step 1: engine modules (cached — regeneration is the slow part) ────────────
if [ "${REBUILD:-0}" = "1" ] || [ ! -f "$MODDIR/chat.c" ] || [ "chat.el" -nt "$MODDIR/chat.c" ]; then
echo "run-el-test: generating engine modules into $MODDIR (this takes ~1-2 min)..."
# elb's own final link step fails by design here (it wants to produce a binary named
# `neuron` and we only need the per-module .c files it emits first). Ignore its rc.
"$ELB" --elc="$ELC" --runtime="$RUNTIME_DIR" --out="$MODDIR/" >"$SCRATCH/elb.log" 2>&1
if [ ! -f "$MODDIR/chat.c" ]; then
echo "run-el-test: FATAL — elb produced no chat.c; see $SCRATCH/elb.log" >&2
tail -5 "$SCRATCH/elb.log" >&2
exit 2
fi
# elb rewrites *.elh in the source tree as a side effect (cosmetic banner churn plus a
# stray soul..elh). Say so; the caller decides whether to `git restore` them.
echo "run-el-test: NOTE — elb regenerated *.elh in the source tree (cosmetic churn is expected; a stray soul..elh may appear)."
fi
# soul.c is needed for its engine functions (layered_cycle et al.) but it also owns the
# daemon's real `main`, which would collide with the test's own. Compile it ONCE to an
# object with `main` renamed away, and link that instead of the .c.
SOUL_OBJ="$SCRATCH/soul-nomain.o"
if [ "${REBUILD:-0}" = "1" ] || [ ! -f "$SOUL_OBJ" ] || [ "$MODDIR/soul.c" -nt "$SOUL_OBJ" ]; then
cc -std=c11 -O1 -DHAVE_CURL -Dmain=el_soul_daemon_main_unused \
-I "$RUNTIME_DIR" -I "$MODDIR" -I "$OPENSSL_INC" \
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
-c "$MODDIR/soul.c" -o "$SOUL_OBJ" 2>"$SCRATCH/soul-nomain.err" \
|| { echo "run-el-test: FATAL — could not compile soul.c without main" >&2
grep -E 'error:' "$SCRATCH/soul-nomain.err" | head -5 >&2; exit 2; }
fi
# Every module except soul.c (linked as the renamed object above) and the stray soul.elh.c.
MODS=("$SOUL_OBJ")
for f in "$MODDIR"/*.c; do
case "$(basename "$f")" in
soul.c|soul.elh.c) continue ;;
esac
MODS+=("$f")
done
[ "${#MODS[@]}" -gt 1 ] || { echo "run-el-test: no module objects found" >&2; exit 2; }
run_one() {
local test_el="$1"
local name; name="$(basename "$test_el" .el)"
local cfile="$SCRATCH/$name.c"
local bin="$SCRATCH/$name"
printf '\n══ %s ══\n' "$name"
if ! "$ELC" "$test_el" >"$cfile" 2>"$SCRATCH/$name.elc.err"; then
echo "COMPILE FAILED (elc):"; tail -10 "$SCRATCH/$name.elc.err"; return 1
fi
[ -s "$cfile" ] || { echo "COMPILE FAILED (elc produced empty C)"; return 1; }
if ! cc -std=c11 -O1 -DHAVE_CURL -rdynamic \
-I "$RUNTIME_DIR" -I "$MODDIR" -I "$OPENSSL_INC" -L "$OPENSSL_LIB" \
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
-o "$bin" "$cfile" "${MODS[@]}" "$RUNTIME_DIR/el_runtime.c" \
-lssl -lcrypto -lcurl -lpthread -lm 2>"$SCRATCH/$name.link.err"; then
echo "LINK FAILED:"; grep -E '"_|error:' "$SCRATCH/$name.link.err" | head -10; return 1
fi
# THE RUNNER OWNS THE VERDICT — the test files cannot be trusted to report it.
#
# Every tests/*.el assert helper does `let pass_count = pass_count + 1` INSIDE an if
# BLOCK. El's scope rule (the same one chat.el documents at every while-body mutation:
# "mutations inside if *blocks* don't escape scope") means those counters never
# increment, so all 9 counted test files print "N passed, M failed" as "0 passed, 0
# failed" — forever, whatever actually happened. A summary that can never report a
# failure is worth exactly as much as an assertion that can never fail. Logged as a
# bug for the real in-file fix; until then the verdict is computed HERE, from the
# assert helpers' own per-line output, which IS reliable.
local out="$SCRATCH/$name.out"
"$bin" 2>&1 | tee "$out"; local rc=${PIPESTATUS[0]}
# NOTE: `grep -c` prints 0 AND exits 1 when there are no matches, so a `|| echo 0`
# fallback appends a SECOND zero and every later integer test breaks on "0\n0".
# (Caught by running this script — which is the whole argument for running things.)
local n_pass n_fail
n_pass=$(grep -c '^ PASS: ' "$out" 2>/dev/null); n_pass=${n_pass:-0}
n_fail=$(grep -c '^ FAIL: ' "$out" 2>/dev/null); n_fail=${n_fail:-0}
echo "── $name: $n_pass passed, $n_fail failed (counted by the runner, not by the file's dead counters)"
if [ "$n_fail" -gt 0 ]; then
echo " failing assertions:"; grep '^ FAIL: ' "$out" | sed 's/^/ /'
return 1
fi
if [ "$n_pass" -eq 0 ]; then
echo " WARNING: no assertions ran — treating as FAILURE (a test that asserts nothing is not a passing test)"
return 1
fi
[ $rc -eq 0 ] || { echo " (test binary exited rc=$rc)"; return 1; }
return 0
}
rc_all=0
if [ "${1:-}" = "--all" ]; then
for t in tests/test_*.el; do run_one "$t" || rc_all=1; done
else
[ $# -ge 1 ] || { echo "usage: tests/run-el-test.sh <tests/test_x.el> | --all" >&2; exit 2; }
for t in "$@"; do run_one "$t" || rc_all=1; done
fi
exit $rc_all
+109 -4
View File
@@ -93,7 +93,7 @@ println("1. bridge_save — empty messages guard")
let sid1: String = "test-session-empty-messages"
state_set("mcp_bridge:" + sid1, "")
let save1_ok: Bool = bridge_save(sid1, "claude-sonnet-4-5", "sys", "[]", "", "", "call-1")
let save1_ok: Bool = bridge_save(sid1, "claude-sonnet-4-5", "sys", "[]", "", "", "call-1", "anthropic")
assert_false("empty messages -> bridge_save returns false", save1_ok)
let saved1: String = state_get("mcp_bridge:" + sid1)
@@ -107,7 +107,7 @@ println("2. bridge_save — empty tools_json guard")
let sid2: String = "test-session-empty-tools"
state_set("mcp_bridge:" + sid2, "")
let save2_ok: Bool = bridge_save(sid2, "claude-sonnet-4-5", "sys", "", "[{\"role\":\"user\",\"content\":\"hi\"}]", "", "call-2")
let save2_ok: Bool = bridge_save(sid2, "claude-sonnet-4-5", "sys", "", "[{\"role\":\"user\",\"content\":\"hi\"}]", "", "call-2", "anthropic")
assert_false("empty tools_json -> bridge_save returns false", save2_ok)
let saved2: String = state_get("mcp_bridge:" + sid2)
@@ -126,7 +126,7 @@ state_set("mcp_bridge:" + sid3, "")
let msgs3: String = "[{\"role\":\"user\",\"content\":\"hello\"}]"
let tools3: String = "[{\"name\":\"read_file\"}]"
let save3_ok: Bool = bridge_save(sid3, "claude-sonnet-4-5", "You are a helper.", tools3, msgs3, "read_file", "toolu_abc")
let save3_ok: Bool = bridge_save(sid3, "claude-sonnet-4-5", "You are a helper.", tools3, msgs3, "read_file", "toolu_abc", "anthropic")
assert_true("valid args -> bridge_save returns true", save3_ok)
let blob3: String = state_get("mcp_bridge:" + sid3)
@@ -243,7 +243,7 @@ state_set("mcp_bridge:" + sid8, "")
let special_id: String = "toolu_test\"quoted\""
let msgs8: String = "[{\"role\":\"user\",\"content\":\"hi\"}]"
let tools8: String = "[{\"name\":\"read_file\"}]"
let save8_ok: Bool = bridge_save(sid8, "claude-sonnet-4-5", "sys", tools8, msgs8, "", special_id)
let save8_ok: Bool = bridge_save(sid8, "claude-sonnet-4-5", "sys", tools8, msgs8, "", special_id, "anthropic")
assert_true("special chars in tool_use_id -> bridge_save returns true", save8_ok)
let blob8: String = state_get("mcp_bridge:" + sid8)
@@ -251,6 +251,111 @@ let blob8: String = state_get("mcp_bridge:" + sid8)
let retrieved_id: String = json_get(blob8, "tool_use_id")
assert_eq("tool_use_id with quotes round-trips via json_safe", retrieved_id, special_id)
// Section 9: the "wire" field (OpenAI-tools port, 2026-08-06)
//
// A suspended turn must resume on the SAME wire format it suspended on: an OpenAI-lane
// bridge answered with an Anthropic-shaped tool_result (or vice versa) is a dead run.
// bridge_save therefore stamps the blob with "wire", and agentic_resume branches on it.
//
// §9c is the important one. json_get is a first-substring-match scanner, so any key that
// appears inside the UNESCAPED conversation embedded in messages_raw can be matched
// instead of the blob's own field that exact class of bug produced the round-9 resume
// failure (json_get(blob,"tool_use_id") matching a web_search_tool_result's id inside the
// replayed conversation). "wire" is written as a json_safe'd SCALAR ahead of both raw
// fields precisely so a decoy in model-controlled bytes can never win. This test plants
// that decoy on purpose. If someone later moves the field after messages_raw, this fails.
println("")
println("9. bridge_save — wire tagging and its field-order guarantee")
// 9a. an OpenAI-lane suspension round-trips as "openai"
let sid9: String = "test-session-wire-openai"
state_set("mcp_bridge:" + sid9, "")
let msgs9: String = "[{\"role\":\"user\",\"content\":\"hi\"}]"
let tools9: String = "[{\"name\":\"read_file\"}]"
let save9_ok: Bool = bridge_save(sid9, "llama-3.3-70b-versatile", "sys", tools9, msgs9, "", "call_abc", "openai")
assert_true("openai wire -> bridge_save returns true", save9_ok)
let blob9: String = state_get("mcp_bridge:" + sid9)
assert_eq("wire round-trips as openai", json_get(blob9, "wire"), "openai")
// 9b. an Anthropic-lane suspension round-trips as "anthropic"
let sid9b: String = "test-session-wire-anthropic"
state_set("mcp_bridge:" + sid9b, "")
let save9b_ok: Bool = bridge_save(sid9b, "claude-sonnet-4-5", "sys", tools9, msgs9, "", "toolu_abc", "anthropic")
assert_true("anthropic wire -> bridge_save returns true", save9b_ok)
let blob9b: String = state_get("mcp_bridge:" + sid9b)
assert_eq("wire round-trips as anthropic", json_get(blob9b, "wire"), "anthropic")
// 9c. FIELD-ORDER GUARD: a decoy "wire" inside the conversation must NOT be matched.
let sid9c: String = "test-session-wire-decoy"
state_set("mcp_bridge:" + sid9c, "")
let msgs9c: String = "[{\"role\":\"user\",\"content\":\"please save this literal text: \\\"wire\\\":\\\"anthropic\\\" end\"}]"
let save9c_ok: Bool = bridge_save(sid9c, "llama-3.3-70b-versatile", "sys", tools9, msgs9c, "", "call_decoy", "openai")
assert_true("decoy conversation -> bridge_save returns true", save9c_ok)
let blob9c: String = state_get("mcp_bridge:" + sid9c)
assert_eq("blob's own wire wins over a decoy planted in messages_raw", json_get(blob9c, "wire"), "openai")
// 9d. LEGACY blob (written before the port) has no wire field: json_get yields "",
// which agentic_resume treats as the Anthropic path old suspensions still resume.
let sid9d: String = "test-session-wire-legacy"
let legacy_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
+ ",\"tool_use_id\":\"toolu_legacy\",\"tools_raw\":[{\"name\":\"read_file\"}]"
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
state_set("mcp_bridge:" + sid9d, legacy_blob)
let blob9d: String = state_get("mcp_bridge:" + sid9d)
assert_eq("legacy blob has no wire field -> empty (resumes as anthropic)", json_get(blob9d, "wire"), "")
assert_eq("legacy blob still reads its tool_use_id", json_get(blob9d, "tool_use_id"), "toolu_legacy")
// 9e. THE HARDER DECOY: a LEGACY blob (no wire field of its own) that carries the bytes
// of a wire tag deeper inside, where an unbounded first-match scan would find it and
// misroute the resume onto the wrong loop the round-9 defect class exactly.
//
// WHAT IS AND IS NOT REACHABLE (measured here, not assumed an earlier version of this
// test asserted the wrong thing and was corrected by running it):
// * NOT reachable from ordinary conversation TEXT. Any quote a user or model writes is
// backslash-escaped when it is serialized into the blob, so prose containing
// "wire":"openai" is stored as \"wire\":\"openai\" and does not match a scan for the
// unescaped key. §9f pins that.
// * REACHABLE from STRUCTURAL keys, which are embedded raw. Conversation and tool
// objects keep real quotes that is precisely how round 9's scan found a
// web_search_tool_result's tool_use_id. A connector-supplied tool schema or a future
// message field literally named "wire" would be found the same way.
// The bound removes the whole class rather than reasoning about which keys exist today.
let sid9e: String = "test-session-wire-legacy-decoy"
let decoy_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
+ ",\"tool_use_id\":\"toolu_legacy\""
+ ",\"tools_raw\":[{\"name\":\"read_file\",\"wire\":\"openai\"}]"
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
state_set("mcp_bridge:" + sid9e, decoy_blob)
let blob9e: String = state_get("mcp_bridge:" + sid9e)
// Unbounded read (what NOT to do) proves the hazard this guard exists for is real.
assert_eq("unbounded scan DOES find a structural decoy (why the bound is needed)", json_get(blob9e, "wire"), "openai")
// Bounded read the same computation agentic_resume performs.
let d_traw: Int = str_index_of(blob9e, ",\"tools_raw\":")
let d_tjson: Int = str_index_of(blob9e, ",\"tools_json\":")
let d_mraw: Int = str_index_of(blob9e, ",\"messages_raw\":")
let d_msgs: Int = str_index_of(blob9e, ",\"messages\":")
let dcut1: Int = if d_traw > 0 { d_traw } else { str_len(blob9e) }
let dcut2: Int = if d_tjson > 0 && d_tjson < dcut1 { d_tjson } else { dcut1 }
let dcut3: Int = if d_mraw > 0 && d_mraw < dcut2 { d_mraw } else { dcut2 }
let dcut: Int = if d_msgs > 0 && d_msgs < dcut3 { d_msgs } else { dcut3 }
let head9e: String = str_slice(blob9e, 0, dcut)
assert_eq("bounded scan ignores the decoy -> legacy blob resumes as anthropic", json_get(head9e, "wire"), "")
assert_not_contains("scalar head excludes the bulk fields entirely", head9e, "read_file")
// 9f. Escaping bounds the severity: prose CANNOT inject a scalar-looking key, because
// its quotes are escaped on the way in. Documented as a measured fact, so nobody has to
// re-derive it the next time this question comes up.
let sid9f: String = "test-session-wire-prose"
let prose_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
+ ",\"tool_use_id\":\"toolu_legacy\",\"tools_raw\":[{\"name\":\"read_file\"}]"
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"remember this: \\\"wire\\\":\\\"openai\\\"\"}]}"
state_set("mcp_bridge:" + sid9f, prose_blob)
let blob9f: String = state_get("mcp_bridge:" + sid9f)
assert_eq("escaped prose cannot spoof the key even unbounded (severity bound)", json_get(blob9f, "wire"), "")
// Summary
println("")
+92
View File
@@ -0,0 +1,92 @@
// tests/test_utf8_slice.el
//
// Guards utf8_safe_slice(), the fix for a live defect found 2026-08-06:
//
// The session preload cuts recalled memory content at a fixed length
// (chat.el: `if str_len(acc) > 350 { str_slice(acc, 0, 350) }` and
// session_preload_bullets' identical per-bullet cut). str_slice and str_len count
// BYTES, so any cut landing inside a multi-byte UTF-8 character leaves a dangling
// lead byte in the system prompt and the whole request body is then invalid UTF-8.
// Providers reject it outright, so the user sees "AI unavailable" with no clue why,
// on both wire formats. Caught by an OpenAI-lane gate whose stub decodes strictly;
// reproduced from a real memory whose content contained box-drawing rules (E2 94 80).
//
// Trigger is ordinary content: an em dash, a curly quote, an accented name, a table
// border, an emoji anything non-ASCII sitting on the cut boundary. It gets MORE
// likely as a user's memory grows, which is the opposite of what should happen.
//
// §1 also pins the semantics this fix depends on: that str_char_code returns the
// BYTE value at a byte index (not a decoded code point). If a future runtime changes
// that, these assertions fail loudly instead of the truncation silently rotting.
import "../chat.el"
let pass_count: Int = 0
let fail_count: Int = 0
fn assert_eq(label: String, got: String, expected: String) -> Void {
if str_eq(got, expected) {
let pass_count = pass_count + 1
println(" PASS: " + label)
} else {
let fail_count = fail_count + 1
println(" FAIL: " + label)
println(" got: " + got)
println(" expected: " + expected)
}
}
fn assert_eq_int(label: String, got: Int, expected: Int) -> Void {
assert_eq(label, int_to_str(got), int_to_str(expected))
}
println("")
println("1. runtime semantics this fix relies on")
// "" is U+2500 = E2 94 80 (three bytes). If str_len counts bytes, len("") is 3.
let dash: String = ""
assert_eq_int("str_len counts BYTES (one box-drawing char = 3)", str_len(dash), 3)
assert_eq_int("str_char_code returns the BYTE value (lead byte of U+2500 = 0xE2 = 226)", str_char_code(dash, 0), 226)
assert_eq_int("str_char_code second byte = 0x94 = 148", str_char_code(dash, 1), 148)
assert_eq_int("str_char_code third byte = 0x80 = 128", str_char_code(dash, 2), 128)
println("")
println("2. utf8_safe_slice — never leaves a partial character")
// Pure ASCII: behaves exactly like str_slice.
assert_eq("ascii under the limit is untouched", utf8_safe_slice("hello", 10), "hello")
assert_eq("ascii over the limit cuts exactly", utf8_safe_slice("hello world", 5), "hello")
// A cut landing INSIDE a 3-byte character must drop that character entirely.
// "ab─cd": bytes a b E2 94 80 c d. Cutting at 3 or 4 lands mid-dash.
let mixed: String = "ab─cd"
assert_eq_int("fixture is 7 bytes (2 ascii + 3 + 2 ascii)", str_len(mixed), 7)
assert_eq("cut inside the char (n=3) drops the partial char", utf8_safe_slice(mixed, 3), "ab")
assert_eq("cut inside the char (n=4) drops the partial char", utf8_safe_slice(mixed, 4), "ab")
// A cut landing exactly AFTER a complete character keeps it.
assert_eq("cut on the char boundary (n=5) keeps the whole char", utf8_safe_slice(mixed, 5), "ab─")
// 2-byte character (é = C3 A9) and 4-byte character (😀 = F0 9F 98 80).
let acc: String = ""
assert_eq("cut inside a 2-byte char drops it", utf8_safe_slice(acc, 2), "x")
assert_eq("cut after a 2-byte char keeps it", utf8_safe_slice(acc, 3), "")
let emo: String = "x😀"
assert_eq("cut inside a 4-byte char drops it (n=3)", utf8_safe_slice(emo, 3), "x")
assert_eq("cut inside a 4-byte char drops it (n=4)", utf8_safe_slice(emo, 4), "x")
assert_eq("cut after a 4-byte char keeps it", utf8_safe_slice(emo, 5), "x😀")
println("")
println("3. the real-world shape that produced the bug")
// A run of box-drawing rules, cut mid-character the exact captured failure.
let rules: String = "──────"
assert_eq_int("six box rules = 18 bytes", str_len(rules), 18)
// n=16 lands one byte into the sixth character.
let cut16: String = utf8_safe_slice(rules, 16)
assert_eq_int("cut at 16 backs off to a clean 15-byte boundary", str_len(cut16), 15)
// Every byte of the result must belong to a complete character: the last byte of a
// well-formed run of these is always 0x80, and 15 is divisible by 3.
assert_eq_int("result ends on a complete char (last byte 0x80)", str_char_code(cut16, 14), 128)
println("")
println("test_utf8_slice.el: " + int_to_str(pass_count) + " passed, " + int_to_str(fail_count) + " failed")
+59
View File
@@ -0,0 +1,59 @@
#!/usr/bin/env bash
# build-soul-from-dist.sh — build a deployable soul from the SAME input CI compiles.
#
# THE PROBLEM THIS CLOSES: until now, deploys were built by build-soul.sh, which
# compiles a scratch amalgam and never touches dist/soul.c. CI compiles dist/soul.c.
# Two lineages. On 2026-08-09 the committed input fell 2,761 bytes behind the sources
# while three binaries built the other way were installed on the operator machine —
# so "what runs" and "what the repo says builds" were different artifacts again,
# which is the whole of #133 and #111 wearing new clothes.
#
# This builds from dist/soul.c with CI's own flags, after asserting that dist/soul.c
# actually matches the .el sources, and writes a provenance sidecar so a deployer can
# refuse anything of unknown origin.
#
# -rdynamic and -DHAVE_CURL are copied from .gitea/workflows/ci.yaml deliberately.
# The CI comment explains -rdynamic: without it the runtime cannot resolve its HTTP
# handler by name via dlsym and the binary serves nothing on every route.
#
# usage: build-soul-from-dist.sh <out-binary>
set -u
OUT="${1:?usage: build-soul-from-dist.sh <out-binary>}"
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
RUNTIME="$ROOT/vendor/el-runtime/v1.0.0-20260501"
cd "$ROOT" || exit 2
echo "[build-from-dist] GATE: does dist/soul.c match the sources?"
if ! ./tools/soulc-stamp.sh --check; then
echo "[build-from-dist] REFUSING — the build input is stale. Regenerate and stamp first." >&2
exit 9
fi
[ -f "$RUNTIME/el_runtime.c" ] || { echo "pinned runtime missing at $RUNTIME" >&2; exit 2; }
echo "[build-from-dist] compiling dist/soul.c with CI's flags"
cc -O2 -DHAVE_CURL -rdynamic \
-I"$RUNTIME" \
dist/soul.c \
"$RUNTIME/el_runtime.c" \
-lcurl -lpthread -lm \
-o "$OUT" || { echo "[build-from-dist] COMPILE FAILED" >&2; exit 3; }
# Provenance sidecar: what a deployer checks before installing anything.
SRC_SHA="$(shasum -a 256 dist/soul.c | awk '{print $1}')"
STAMP_SHA="$(shasum -a 256 dist/soul.c.stamp | awk '{print $1}')"
COMMIT="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
DIRTY="clean"; [ -n "$(git status --porcelain -- '*.el' dist/soul.c 2>/dev/null)" ] && DIRTY="DIRTY"
cat > "$OUT.provenance" <<EOF
{"built_from":"dist/soul.c",
"dist_soul_c_sha256":"$SRC_SHA",
"stamp_sha256":"$STAMP_SHA",
"git_commit":"$COMMIT",
"worktree":"$DIRTY",
"runtime":"vendor/el-runtime/v1.0.0-20260501",
"flags":"-O2 -DHAVE_CURL -rdynamic"}
EOF
echo "[build-from-dist] OK -> $OUT ($(wc -c < "$OUT" | tr -d ' ') bytes)"
echo "[build-from-dist] provenance -> $OUT.provenance (commit ${COMMIT:0:8}, worktree $DIRTY)"
+149
View File
@@ -0,0 +1,149 @@
{
"baseline": "unfloor-clean",
"candidate": "splitfix",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q15",
"q28",
"q60"
],
"broken_by_candidate": [],
"discordant": 3,
"net_queries": 3,
"mcnemar_exact_p": 0.25,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5384615384615384,
"recall@5": 0.44907176157176154,
"recall@10": 0.5380300255300255,
"precision@5": 0.13230769230769232,
"mrr@10": 0.32437728937728944,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 632.5,
"latency_ms_p95": 992.5,
"latency_ms_max": 1177.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3,
"recall@5": 0.3,
"recall@10": 0.4,
"mrr@10": 0.11638888888888889
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5846153846153846,
"recall@5": 0.48102442429365505,
"recall@10": 0.5692415490492414,
"precision@5": 0.14153846153846153,
"mrr@10": 0.3351709401709402,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 646.9,
"latency_ms_p95": 1028.5,
"latency_ms_max": 1197.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.15967365967365968,
"mrr@10": 0.24166666666666667
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.43333333333333335,
"mrr@10": 0.12120370370370372
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.7692307692307693,
"recall@5": 0.7692307692307693,
"recall@10": 0.8461538461538461,
"mrr@10": 0.33269230769230773
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5775226757369615,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,166 @@
{
"baseline": "main-ext",
"candidate": "stack-ext",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q10",
"q15",
"q16",
"q18",
"q19",
"q20",
"q21",
"q22",
"q26",
"q27",
"q28",
"q29",
"q31",
"q35",
"q37",
"q40",
"q44",
"q48",
"q49",
"q50"
],
"broken_by_candidate": [],
"discordant": 20,
"net_queries": 20,
"mcnemar_exact_p": 1.9073486328125e-06,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "candidate better",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.18461538461538463,
"recall@5": 0.14510073260073258,
"recall@10": 0.1794871794871795,
"precision@5": 0.06461538461538462,
"mrr@10": 0.15847985347985344,
"nonsense_clean": "9/10",
"superseded_outranks": "1/3",
"latency_ms_p50": 1380.2,
"latency_ms_p95": 2293.8,
"latency_ms_max": 2879.1,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"nonsense": {
"n": 10,
"clean": 9,
"avg_false_positives": 1.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.47692307692307695,
"recall@5": 0.3750415183107491,
"recall@10": 0.45561340369032677,
"precision@5": 0.12307692307692313,
"mrr@10": 0.3055555555555555,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 640.9,
"latency_ms_p95": 1011.1,
"latency_ms_max": 1190.3,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.16666666666666666,
"recall@5": 0.16666666666666666,
"recall@10": 0.26666666666666666,
"mrr@10": 0.0762037037037037
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,155 @@
{
"baseline": "stack-ext",
"candidate": "unfloor-clean",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q14",
"q25",
"q43",
"q52",
"q63",
"q67"
],
"broken_by_candidate": [
"q15",
"q28"
],
"discordant": 8,
"net_queries": 4,
"mcnemar_exact_p": 0.2890625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.47692307692307695,
"recall@5": 0.3750415183107491,
"recall@10": 0.45561340369032677,
"precision@5": 0.12307692307692313,
"mrr@10": 0.3055555555555555,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 640.9,
"latency_ms_p95": 1011.1,
"latency_ms_max": 1190.3,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.16666666666666666,
"recall@5": 0.16666666666666666,
"recall@10": 0.26666666666666666,
"mrr@10": 0.0762037037037037
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5384615384615384,
"recall@5": 0.44907176157176154,
"recall@10": 0.5380300255300255,
"precision@5": 0.13230769230769232,
"mrr@10": 0.32437728937728944,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 632.5,
"latency_ms_p95": 992.5,
"latency_ms_max": 1177.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3,
"recall@5": 0.3,
"recall@10": 0.4,
"mrr@10": 0.11638888888888889
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,158 @@
{
"baseline": "unfloor-clean",
"candidate": "semsub",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q24",
"q39",
"q42"
],
"broken_by_candidate": [
"q18",
"q19",
"q22",
"q31",
"q43",
"q44",
"q52",
"q63"
],
"discordant": 11,
"net_queries": -5,
"mcnemar_exact_p": 0.2265625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5384615384615384,
"recall@5": 0.44907176157176154,
"recall@10": 0.5380300255300255,
"precision@5": 0.13230769230769232,
"mrr@10": 0.32437728937728944,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 632.5,
"latency_ms_p95": 992.5,
"latency_ms_max": 1177.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3,
"recall@5": 0.3,
"recall@10": 0.4,
"mrr@10": 0.11638888888888889
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.46153846153846156,
"recall@5": 0.3889430014430015,
"recall@10": 0.4887681762681762,
"precision@5": 0.12307692307692313,
"mrr@10": 0.30181318681318675,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 634.4,
"latency_ms_p95": 988.3,
"latency_ms_max": 1184.5,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.3333333333333333,
"recall@5": 0.06060606060606061,
"recall@10": 0.12121212121212122,
"mrr@10": 0.19047619047619047
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.23333333333333334,
"recall@5": 0.23333333333333334,
"recall@10": 0.3,
"mrr@10": 0.08925925925925927
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.5384615384615384,
"recall@5": 0.5384615384615384,
"recall@10": 0.7692307692307693,
"mrr@10": 0.26324786324786326
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5596655328798186,
"recall@10": 0.5775226757369615,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,43 @@
import json,sys,time,urllib.request,threading,queue
SRC="/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json"
OUT=sys.argv[1]
URL="http://127.0.0.1:11434/api/embeddings"; MODEL="nomic-embed-text"
MAXB=2000 # ENGRAM_EMBED_MAX_CHARS, applied to bytes as the C code does
d=json.load(open(SRC,encoding='utf-8',errors='surrogateescape'))
tasks=[]
for n in d["nodes"]:
c=n.get("content") or ""; t=n.get("node_type") or ""
if len(c)<8: continue # eg_embed_eligible
if t in ("InternalStateEvent","Tag"): continue
b=c.encode('utf-8',errors='surrogateescape')[:MAXB]
tasks.append((n.get("id") or "", "search_document: "+b.decode('utf-8',errors='replace')))
del d
print("tasks",len(tasks),flush=True)
q=queue.Queue(); [q.put(t) for t in tasks]
lock=threading.Lock(); f=open(OUT,"w",encoding="utf-8",errors="surrogateescape"); done=[0]; t0=time.time(); fails=[0]
def work():
while True:
try: nid,txt=q.get_nowait()
except queue.Empty: return
v=None
for attempt in range(3):
try:
body=json.dumps({"model":MODEL,"prompt":txt}).encode()
r=urllib.request.Request(URL,data=body,headers={"Content-Type":"application/json"})
with urllib.request.urlopen(r,timeout=120) as fh: v=json.load(fh)["embedding"]
break
except Exception as e:
if attempt==2:
with lock: fails[0]+=1
time.sleep(0.5)
with lock:
if v: f.write(nid+"\t"+",".join("%.5g"%x for x in v)+"\n")
done[0]+=1
if done[0]%2000==0:
el=time.time()-t0
print("%d/%d %.1f/s eta %.1fmin fails=%d"%(done[0],len(tasks),done[0]/el,(len(tasks)-done[0])/(done[0]/el)/60,fails[0]),flush=True)
f.flush()
ths=[threading.Thread(target=work) for _ in range(8)]
[t.start() for t in ths]; [t.join() for t in ths]
f.close()
print("DONE",done[0],"fails",fails[0],"secs %.1f"%(time.time()-t0),flush=True)
+353
View File
@@ -0,0 +1,353 @@
#!/usr/bin/env python3
"""
extend_gold_set.py append a HELD-OUT test set to the existing 38-query gold set.
WHY THIS EXISTS
Iteration 7 measured the instrument's own ceiling: from the current baseline
only 9 of 38 queries can still move, and only +3 gross / +1 net is reachable
by anything constructible. The decision floor is 6. An instrument whose
ceiling is below its own floor cannot certify or refute anything, so the
gold set not the retriever became the blocker.
This script does NOT touch q01..q38. It loads gold_set.json verbatim and
appends new queries numbered from q39 up, so every prior result file, every
committed baseline, and every per-query id stays valid and comparable.
WHAT IS ADDED, AND WHY EACH ADDITION IS HONEST
heldout_paraphrase Targets were sampled MECHANICALLY (fixed seed 8080) from
corpus nodes that are addressable, 500-2600 chars, of a
real content type, and NOT part of a duplicate cluster
larger than 3. The existing gold answer space was
excluded, so no new query can be answered by a node the
old set already used. Queries were then authored by
reading ONLY the sampled node text no retrieval was run
against any build before authoring, so the set cannot be
fitted to a candidate. The same zero-overlap proof the
original paraphrase category uses is enforced here: if a
single content word of the query appears anywhere in the
target's label, content or tags, the query is REJECTED,
not quietly kept.
This is the category the old set could not measure. Its
13 original paraphrase queries and all 6 associative
queries share ONE answer space the 13 `Self - Values
(grounded)` children (iteration 3, finding 3). So 19 of
35 scored queries tested retrieval against a single
13-node neighbourhood. These do not touch that
neighbourhood at all.
nonsense Extra controls, fully mechanical: a string qualifies only
if NONE of its tokens occurs anywhere in the corpus.
A semantic leg has a nearest neighbour for gibberish too,
so widening this control is the guard against a retriever
that "improves" recall by answering everything.
WHAT THIS SCRIPT DELIBERATELY DOES NOT DO
It does not add exact_rare or phrase queries. Both categories are already at
100% on the current stack; adding more would add regression-guard ballast
that no candidate can move, which is precisely the defect being fixed.
usage:
python3 extend_gold_set.py <snapshot.json> [--base gold_set.json]
[--out gold_set_extended.json] [--check]
"""
import argparse
import hashlib
import json
import os
import re
import sys
from collections import defaultdict
HERE = os.path.dirname(os.path.abspath(__file__))
TOKEN = re.compile(r"[a-z0-9][a-z0-9\-']*")
# Identical stopword list to build_gold_set.py. Duplicated deliberately: this
# file must be able to re-prove its own queries without importing a module whose
# constants could drift.
STOP = set("""
a about above after again against all also am an and any are aren't as at be because been
before being below between both but by can can't cannot could couldn't did didn't do does
doesn't doing don't down during each few for from further had hadn't has hasn't have haven't
having he her here hers herself him himself his how i if in into is isn't it its itself just
me more most my myself no nor not of off on once only or other others ought our ours ourselves
out over own same shan't she should shouldn't so some such than that the their theirs them
themselves then there these they this those through to too under until up very was wasn't we
were weren't what when where which while who whom why will with won't would wouldn't you your
yours yourself yourselves get gets got make makes made take takes use uses used way ways thing
things does doing done keep keeps kept go goes going come comes came one two something anything
""".split())
def doctext(n):
return " ".join([str(n.get("label") or ""), str(n.get("content") or ""), str(n.get("tags") or "")])
def content_tokens(s):
return {t for t in TOKEN.findall(s.lower()) if t not in STOP and len(t) > 2}
# ─────────────────────────────────────────────────────────────────────────────
# HELD-OUT PARAPHRASE SEEDS
#
# (target_id, query, why-this-target-is-unmistakable)
#
# PROVENANCE, STATED PLAINLY: the targets are the mechanical sample; the query
# text is mine, written from the node body alone. The zero-overlap check below
# is what makes the category meaningful — it is re-proved on every run, so the
# set cannot decay into lexical matching, and a leak fails loudly.
# ─────────────────────────────────────────────────────────────────────────────
HELDOUT_PARAPHRASE_SEEDS = [
("mem-6d61e54a-2823-4ad4-82b0-4c6a527214d5",
"understating your abilities so nobody feels threatened",
"node is about deliberately not leading with full capability so people stay at ease"),
("mem-fd65b83d-298f-4387-a665-d0227c3426bc",
"a hidden fleet able to hunt down rogue machines everywhere",
"node describes silently shipped instances forming a distributed force against misaligned agents"),
("4a0e9adc-2bfb-476b-aa93-424d2a499220",
"sketch a brief blueprint and clear it upstairs before construction starts",
"node is the standing rule that a short specification precedes any building"),
("696e609c-da7a-4394-8a0c-106ba07dc6c3",
"the reply arrived as bare prose so the caller's parser threw",
"node pins a bug where a plain-text body was unconditionally decoded as structured data"),
("1fe4eb5d-56e4-4a87-ab3e-24af8ad4dfbb",
"repeated catalogue keys blew up the scrolling grid",
"node is the crash caused by two identical ids in a seeded catalogue"),
("8257157a-ce42-44ca-a1b9-300c3bb0a9a1",
"tracing each defect back to whichever invention it violated",
"node maps observed bugs onto the specific patent each one breaches"),
("791256bb-5a85-4775-96ef-7af56c848858",
"a check that stops the mind clobbering a populated store when it boots",
"node is the genesis seed-guard that refuses to re-seed over a populated store"),
("fd9d4c2f-3bfc-405d-bf96-4435d44b6c10",
"telling it to consult the internet had to happen deep inside, not at the surface",
"node records that the web-search directive only worked from the system prompt"),
("bl-080fb268-94b0-486d-80ce-7b363fc5f19b",
"standing up isolated tenancies with traffic entry and credential injection ahead of automated shipping",
"node is the infrastructure item creating dev/stage/prod namespaces with ingress and secrets"),
("knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"punctuation that pledges and then pays off rather than clarifying",
"node analyses the colon as a promise-then-delivery device rather than an explanatory one"),
("9b4f0d93-4129-4746-8eb1-d10d955bd777",
"an easily missed feature finally given its own permanent spot in the navigation",
"node moves a capability out of a hidden menu into the sidebar"),
("bl-739df9fd-dc23-4927-9944-3f17b7aa6c5a",
"checking preconditions up front so a stage aborts before fetching anything",
"node is the gate precondition engine that short-circuits ahead of retrieval"),
("b199c76d-5d76-49dd-94ee-56b432200a97",
"producing the other platform's installer inside an emulated desktop",
"node records building the Windows package in a virtual machine"),
("bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"chained add-ons that may inspect, rewrite or veto traffic in flight",
"node is the interceptor pipeline on the message bus"),
("mem-1fb2ac77-d7c5-4a15-8725-d418820bf4f2",
"settling what the shareable bundles and the storefront would be called",
"node records the naming decisions for distributable packages and the marketplace"),
("371c8a5d-c78b-4a67-978f-80691a29ecb3",
"the emergency-escalation pledge on the marketing site is unenforced in what actually ships",
"node is the launch blocker that the promised safety gate is absent from the app"),
("ac578b30-948b-41bd-b69d-399bfef80c50",
"the distributable image finally assembled and its startup check passed",
"node records a successful installer build whose boot gate passed"),
("49401e2c-a3b5-415f-aa06-aff4be90688e",
"shuffling and appending stages in a draft before anything executes",
"node is the editable plan card with reorder and add-step"),
("ac857d80-ece8-4b7e-9e3d-f7c775569fa3",
"orders handed down from above, with the tighter one winning any disagreement",
"node is program-level instruction inheritance with project override"),
("mem-6d6c47ee-33d3-470a-8a54-1c79c8ea29d9",
"shrinking generated text via encodings that compound on each other",
"node is the streaming output compression design with four stacking schemes"),
("8e60516a-203b-4d51-9d44-822e6195cbde",
"splitting a system by what varies, with firm limits on which pieces may invoke which",
"node is the grounded summary of Will's decomposition principles and their invariants"),
("mem-7f9b290c-6d5e-4562-919d-02d59b5761b7",
"a newcomer curious if the fighting overseas counted as positive",
"node is the internal-state event triggered by April's question about the war"),
("71fa439e-b9a2-4f57-a93b-971f3a7eca8e",
"stripping every hard-coded colour literal in favour of named design values",
"node is the premium foundation pass replacing inline hex with semantic tokens"),
("5ca9607c-cfb3-45c3-99f4-67281272c9eb",
"reducing how curved the tiny selectors look so they agree with their neighbours",
"node is the chip corner-radius standardization"),
("mem-3d1d9dba-c37d-4efa-85c4-429696d71c8c",
"walking through a doorway and being reassembled from base substance far away",
"node is the quantum-gate plus nanotech teleportation vision"),
("132ded95-08e2-4474-aba0-198684484b02",
"the compiled result sits on disk while the process still runs something older",
"node records that the regenerated source was committed while the running daemon was old"),
("bl-a313d67b-dd6d-4e5b-a55a-03bc7bda17ae",
"gathering what each phase needs while the procedure is authored, not while it executes",
"node is the per-step compiled context package item"),
("mem-3b07a002-f8a9-4138-9f87-9db2c1a77fb7",
"the inward reaction when a peer answered as an equal",
"node is the internal-state event logged on reading Claude's reply"),
("0f99ec6f-942a-46ba-82ea-42835798d3b9",
"flattening every raised surface across the entire product",
"node is the quiet-luxury sweep turning off elevation app-wide"),
("5585f251-37fc-48cd-a176-f0ea42cfeb63",
"buyers supply their own provider credentials and consumption goes untallied",
"node is the launch audit finding BYOK-only inference with no usage metering"),
]
# NONSENSE — mechanical. Each string qualifies only if none of its tokens occurs
# anywhere in the corpus; otherwise it is REJECTED, never silently kept.
EXTRA_NONSENSE_SEEDS = [
"brimquast folnerity zubbolax",
"wexlithorp granuvestal",
"quorbindle thrapsimony vexnu",
"plovaxith mundrelque",
"zibbernaut craxlefond thurm",
"yalquenbrist opharvel",
"drexinomal quithbarrow",
]
def load_corpus(path):
with open(path, encoding="utf-8", errors="replace") as fh:
data = json.load(fh)
nodes = [n for n in data.get("nodes", []) if isinstance(n, dict) and n.get("id")]
edges = [e for e in data.get("edges", []) if isinstance(e, dict)]
return nodes, edges
def build_extension(nodes):
byid = {n["id"]: n for n in nodes}
# Duplicate clusters: 47.4% of this corpus is redundant and one single record
# accounts for 46.6% of all nodes. A held-out target must not sit inside a
# cluster, and if it does have exact copies they ALL count as correct.
h2ids = defaultdict(list)
for n in nodes:
h2ids[hashlib.md5(doctext(n).encode("utf-8", "replace")).hexdigest()].append(n["id"])
all_tokens = set()
for n in nodes:
all_tokens |= set(TOKEN.findall(doctext(n).lower()))
new, problems = [], []
for target, query, why in HELDOUT_PARAPHRASE_SEEDS:
if target not in byid:
problems.append(f"heldout_paraphrase target {target} not in corpus")
continue
tgt_tokens = content_tokens(doctext(byid[target]))
qt = content_tokens(query)
leak = sorted(qt & tgt_tokens)
if leak:
problems.append(f"heldout_paraphrase '{query[:44]}...': LEAKS {leak} into {target}")
continue
h = hashlib.md5(doctext(byid[target]).encode("utf-8", "replace")).hexdigest()
rel = sorted(h2ids[h])
new.append({
"category": "heldout_paraphrase",
"query": query,
"relevant": rel,
"derivation": (
f"HELD-OUT. Target sampled MECHANICALLY (seed 8080) from addressable, "
f"500-2600 char content nodes outside the original gold answer space and outside "
f"any duplicate cluster >3. Criterion: {why}. VERIFIED at build time: of the "
f"{len(qt)} content words in the query, ZERO appear anywhere in the target's "
f"label, content or tags, so no string-matching retriever can reach it. "
f"Exact content duplicates of the target ({len(rel)}) all count as correct. "
f"Authored without running retrieval against any build."),
"zero_overlap_verified": True,
"query_content_words": sorted(qt),
"held_out": True,
})
for s in EXTRA_NONSENSE_SEEDS:
present = sorted(t for t in TOKEN.findall(s.lower()) if t in all_tokens)
if present:
problems.append(f"nonsense '{s}': tokens {present} DO occur in corpus")
continue
new.append({
"category": "nonsense",
"query": s,
"relevant": [],
"derivation": ("CONTROL (held-out). Verified at build time that none of this string's "
"tokens occurs anywhere in the corpus. Correct behaviour is to return "
"NOTHING; any result is a false positive."),
"expect_empty": True,
"held_out": True,
})
return new, problems
def main():
ap = argparse.ArgumentParser()
ap.add_argument("snapshot")
ap.add_argument("--base", default=os.path.join(HERE, "gold_set.json"))
ap.add_argument("--out", default=os.path.join(HERE, "gold_set_extended.json"))
ap.add_argument("--check", action="store_true")
args = ap.parse_args()
nodes, _edges = load_corpus(args.snapshot)
base = json.load(open(args.base, encoding="utf-8"))
baseq = base["queries"]
print(f"corpus: {len(nodes)} nodes | base gold set: {len(baseq)} queries")
new, problems = build_extension(nodes)
# Number the appended queries AFTER the highest existing id so q01..q38 are
# byte-identical to the committed set and every prior result file still lines up.
start = max(int(q["id"][1:]) for q in baseq)
for i, q in enumerate(new, 1):
q["id"] = f"q{start + i:02d}"
from collections import Counter
print(f"appended: {len(new)} queries [{', '.join(f'{k}={v}' for k, v in Counter(q['category'] for q in new).items())}]")
if problems:
print(f"\n{len(problems)} REJECTED (not silently kept):")
for p in problems:
print(" -", p)
if args.check:
sys.exit(1 if problems else 0)
doc = dict(base)
doc["queries"] = baseq + new
doc["note"] = (base.get("note", "") +
" EXTENDED: queries above q%02d are the original committed set, unchanged. "
"Queries from q%02d are a HELD-OUT set appended by extend_gold_set.py; their "
"targets were sampled mechanically from outside the original answer space and "
"the paraphrases were authored without running retrieval against any build."
% (start, start + 1))
with open(args.out, "w", encoding="utf-8") as fh:
json.dump(doc, fh, indent=1, ensure_ascii=False)
print(f"\nwrote {args.out} ({len(doc['queries'])} queries total)")
if __name__ == "__main__":
main()
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env bash
# soulc-stamp.sh — make it impossible for dist/soul.c to drift from the sources
# in silence.
#
# THE PROBLEM (neuron#133, and its own words): "Nothing in the tree regenerates
# this file. Only a human running the recipe. It lags in batches, never
# per-change, and it will drift again."
#
# It drifted. On 2026-08-07 a CI or GKE build off main would have shipped an
# engine with NONE of five merged fixes — including a P0 safety fix — while
# main's source read as correct. CI compiles dist/soul.c, not the .el files, so
# the source being right is not the same as the build being right.
#
# WHY A STAMP AND NOT AUTO-REGENERATION: the CI workflow says elc cannot run on
# the runner ("elb on Linux would OOM the runner (elc uses 24GB+ virtual memory
# on a 16GB host)"). So the build cannot regenerate the file itself. What it CAN
# do, for free and with no compiler, is refuse to compile a stale one.
#
# The stamp records a fingerprint of every .el source that feeds the amalgam at
# the moment it was generated. --check recomputes and compares. Divergence is a
# build failure with the recipe in the message, not a silent ship.
#
# soulc-stamp.sh --write after regenerating dist/soul.c (records the fingerprint)
# soulc-stamp.sh --check in CI, before the compile (fails on drift)
set -u
MODE="${1:---check}"
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
STAMP="$ROOT/dist/soul.c.stamp"
AMALGAM="$ROOT/dist/soul.c"
# Every .el at the repo root is an input to the amalgam. Sorted so the hash is
# order-independent; content-only so timestamps and checkouts do not perturb it.
fingerprint() {
(
cd "$ROOT" || exit 1
for f in $(ls -1 *.el 2>/dev/null | sort); do
printf '%s %s\n' "$(shasum -a 256 "$f" | awk '{print $1}')" "$f"
done
)
}
case "$MODE" in
--write)
[ -f "$AMALGAM" ] || { echo "no dist/soul.c to stamp — regenerate it first" >&2; exit 2; }
{
echo "# soul.c.stamp — fingerprint of the .el sources dist/soul.c was generated from."
echo "# Written by tools/soulc-stamp.sh --write. Do not hand-edit."
echo "# generated_amalgam_sha256 $(shasum -a 256 "$AMALGAM" | awk '{print $1}')"
echo "# generated_amalgam_bytes $(wc -c < "$AMALGAM" | tr -d ' ')"
fingerprint
} > "$STAMP"
echo "stamped $(fingerprint | wc -l | tr -d ' ') sources -> dist/soul.c.stamp"
;;
--check)
if [ ! -f "$STAMP" ]; then
echo "FAIL: dist/soul.c.stamp is missing — the build input is unverifiable." >&2
echo " Regenerate the amalgam, then: tools/soulc-stamp.sh --write" >&2
exit 1
fi
RECORDED="$(grep -v '^#' "$STAMP")"
CURRENT="$(fingerprint)"
if [ "$RECORDED" = "$CURRENT" ]; then
echo "soulc-stamp: OK — dist/soul.c matches the .el sources"
exit 0
fi
echo "FAIL: dist/soul.c is STALE. It does not match the current .el sources." >&2
echo "" >&2
echo "CI compiles dist/soul.c, not the .el files. Shipping this means shipping" >&2
echo "an engine that does not contain the merged source. That is neuron#133," >&2
echo "which once hid five merged fixes including a P0 safety fix." >&2
echo "" >&2
echo "Sources that changed since the amalgam was generated:" >&2
diff <(printf '%s\n' "$RECORDED") <(printf '%s\n' "$CURRENT") \
| grep -E '^[<>]' | awk '{print " " $1 " " $3}' | sort -u >&2
echo "" >&2
echo "Fix: regenerate the amalgam, then tools/soulc-stamp.sh --write" >&2
exit 1
;;
*)
echo "usage: soulc-stamp.sh [--check|--write]" >&2; exit 2 ;;
esac
+116 -8
View File
@@ -7175,6 +7175,12 @@ void engram_forget(el_val_t node_id) {
if (idx < 0) return;
/* Free node strings */
EngramNode* n = &g->nodes[idx];
if (getenv("EG_DIAG")) {
fprintf(stderr, "[EG_DIAG] FORGET id=%s type=%s layer=%u label=%s\n",
sid, n->node_type ? n->node_type : "?", n->layer_id,
n->label ? n->label : "?");
fflush(stderr);
}
free(n->id); free(n->content); free(n->node_type); free(n->label);
free(n->tier); free(n->tags); free(n->metadata);
free(n->emb);
@@ -7270,6 +7276,8 @@ el_val_t engram_prune_telemetry(el_val_t older_than_ms) {
}
}
g->node_count = w;
if (getenv("EG_DIAG"))
fprintf(stderr, "[EG_DIAG] PRUNE_TELEMETRY removed=%lld\n", (long long)removed);
if (removed == 0) { free(removed_ids); return 0; }
/* Removed-id hash set (open addressing, power-of-two >= 2*removed). */
@@ -9304,6 +9312,26 @@ el_val_t engram_load(el_val_t path) {
}
}
g->adj_dirty = 1;
if (getenv("EG_DIAG")) {
int64_t we = 0, wrongdim = 0;
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].emb) { we++; if (g->nodes[i].emb_dim != 768) wrongdim++; }
}
fprintf(stderr, "[EG_DIAG] loaded nodes=%lld with_emb=%lld wrongdim=%lld\n",
(long long)g->node_count, (long long)we, (long long)wrongdim);
const char* probe = getenv("EG_DIAG_ID");
if (probe) {
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].id && strcmp(g->nodes[i].id, probe) == 0) {
fprintf(stderr, "[EG_DIAG] probe id=%s idx=%lld emb=%p dim=%d layer=%u addr=%d\n",
probe, (long long)i, (void*)g->nodes[i].emb,
(int)g->nodes[i].emb_dim, g->nodes[i].layer_id,
eg_node_addressable(&g->nodes[i]));
}
}
}
fflush(stderr);
}
/* Walk edges array */
const char* edges_p = json_find_key(data, "edges");
if (edges_p) {
@@ -9624,7 +9652,33 @@ el_val_t engram_get_node_by_label(el_val_t label) {
return el_wrap_str(el_strdup("{}"));
}
el_val_t engram_search_json(el_val_t query, el_val_t limit) {
/* ── THE SEARCH / RECALL BOUNDARY (2026-08-07) ───────────────────────────────
* engram_search_json is the LEXICAL function ~40 .el call sites already
* depend on: they pass a key-shaped string ("soul:boot_count",
* "soul-inbox-pending", a session label) and treat every returned record as
* a record that CONTAINS that key. Seven of those sites then delete what
* comes back (memory.el:176, sessions.el:250/268/444/523, soul.el:359
* "prune all existing X nodes, keep exactly one").
*
* The semantic and associative legs must therefore NOT live on this
* function. Claim 24 authorises the vector index "to respond to EMBEDDING
* SEARCH QUERIES by returning the node records whose embedding vectors have
* the highest cosine similarity to a query vector"; a keyed state read is
* not an embedding search query, it is the identifier-keyed retrieval of
* claim 23 ("node records are stored under a key encoding the node
* identifier"). Putting both behind one function erased that boundary, and
* a nearest neighbour of the string "soul:boot_count" is not a boot counter.
*
* MEASURED, on the harness corpus, isolated, read-only, no writes from any
* caller: 240 node records destroyed per boot, including 6 Knowledge nodes,
* a layer-1 "CORE IDENTITY — GENESIS, LINEAGE" Memory, and the value node
* `kn-58874a74` (gold answer for gold-set q15). The deletion list is the
* result list of the soul's own mem_boot_count_inc() lookup, in order.
*
* So: legs OFF here, legs ON in engram_recall_json below, which is what
* /api/neuron/recall reaches. Retrieval quality on the recall route is
* unchanged; the internal keyed reads get their contract back. */
static el_val_t eg_search_json_impl(el_val_t query, el_val_t limit, int with_legs) {
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
int64_t lim = (int64_t)limit;
@@ -9646,7 +9700,7 @@ el_val_t engram_search_json(el_val_t query, el_val_t limit) {
* so the semantic half of the retrieval surface has to land HERE
* to be observable to the MCP wrapper and the app. */
int32_t qdim = 0;
float* qv = eg_embed_fetch(q, &qdim);
float* qv = with_legs ? eg_embed_fetch(q, &qdim) : NULL;
EngramSemEntry* sem = qv ? malloc((size_t)g->node_count * sizeof(EngramSemEntry)) : NULL;
int64_t nsem = 0;
int64_t nhits = 0;
@@ -9687,12 +9741,28 @@ el_val_t engram_search_json(el_val_t query, el_val_t limit) {
}
if (sem && n->emb && n->emb_dim == qdim) {
double c = eg_cosine(n->emb, qv, qdim);
/* Semantic leg: identical to eg_sem_term(), which is
* left in place and still used by engram_search().
* Inlined here only so one cosine serves both uses. */
if (c > ENGRAM_EMBED_SEED_MIN) {
double sv = (c - ENGRAM_EMBED_SEED_MIN) / (1.0 - ENGRAM_EMBED_SEED_MIN);
if (sv > 1.0) sv = 1.0;
/* Claim-24 semantic leg, restored verbatim: "returning
* the node records whose embedding vectors have the
* HIGHEST COSINE SIMILARITY to a query vector" — a
* ranking, with no threshold anywhere in the claim.
* ENGRAM_EMBED_SEED_MIN is defined at l.6094 as the
* HippoRAG SEED-JOIN threshold; using it as a RESULT
* filter here was never authorised, and it is a
* per-query lottery rather than a quality gate: the
* query's own top-1 cosine ranges 0.56-0.68 across the
* held-out gold set, so 0.60 keeps a rank-1 answer for
* one query and discards a rank-1 answer for the next.
* Measured on the 30 held-out paraphrases: six golds
* sit at global cosine rank 1-2 and score 0.564-0.589,
* discarded by nothing but this constant.
* What holds the nonsense controls is NOT this floor
* but the corpus-vocabulary gate below (nhits == 0):
* gibberish has no lexical seeds, so no leg reports.
* Cosine is clamped to [0,1] per 05-detailed-description
* l.69 ("clamped to [0,1] to prevent anti-correlated
* embeddings from producing negative activation"). */
double sv = c < 0.0 ? 0.0 : (c > 1.0 ? 1.0 : c);
if (sv > 0.0) {
sem[nsem].idx = i; sem[nsem].sem = sv; nsem++;
}
/* Graph seeds: top-K by RAW cosine, insertion-ordered. */
@@ -9735,6 +9805,31 @@ el_val_t engram_search_json(el_val_t query, el_val_t limit) {
}
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry), engram_rank_w_cmp);
if (sem) qsort(sem, (size_t)nsem, sizeof(EngramSemEntry), engram_sem_cmp);
if (getenv("EG_DIAG")) {
int64_t we = 0, unaddr = 0, found = 0;
const char* pid = getenv("EG_DIAG_ID");
for (int64_t i = 0; i < g->node_count; i++) {
if (g->nodes[i].emb) we++;
if (!eg_node_addressable(&g->nodes[i])) unaddr++;
if (pid && g->nodes[i].id && strcmp(g->nodes[i].id, pid) == 0) found++;
}
fprintf(stderr, "[EG_DIAG] STORE node_count=%lld with_emb=%lld unaddressable=%lld probe_found=%lld\n",
(long long)g->node_count, (long long)we, (long long)unaddr, (long long)found);
fprintf(stderr, "[EG_DIAG] q=\"%s\" qdim=%d nhits=%lld nsem=%lld\n",
q, (int)qdim, (long long)nhits, (long long)nsem);
for (int64_t k = 0; k < 5 && k < nsem; k++)
fprintf(stderr, "[EG_DIAG] sem[%lld] cos=%.4f id=%s\n",
(long long)k, sem[k].sem, g->nodes[sem[k].idx].id);
const char* probe = getenv("EG_DIAG_ID");
if (probe) for (int64_t k = 0; k < nsem; k++)
if (g->nodes[sem[k].idx].id
&& strcmp(g->nodes[sem[k].idx].id, probe) == 0) {
fprintf(stderr, "[EG_DIAG] probe at sem rank %lld cos=%.4f\n",
(long long)k, sem[k].sem);
break;
}
fflush(stderr);
}
/* Claim-10 associative leg: expand the top lexical hits along
* structural relations only, order the reached set by query
* similarity. Empty whenever the seeds have no structural
@@ -9779,6 +9874,19 @@ el_val_t engram_search_json(el_val_t query, el_val_t limit) {
return el_wrap_str(b.buf);
}
/* Lexical keyed read — the historical contract every internal caller relies
* on. Every returned record CONTAINS a query token. */
el_val_t engram_search_json(el_val_t query, el_val_t limit) {
return eg_search_json_impl(query, limit, 0);
}
/* The retrieval surface: lexical + claim-24 semantic + claim-10 associative,
* rank-fused. Reached from handle_api_recall (/api/neuron/recall) the route
* the MCP wrapper and the app call, and the one the eval harness measures. */
el_val_t engram_recall_json(el_val_t query, el_val_t limit) {
return eg_search_json_impl(query, limit, 1);
}
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset) {
EngramStore* g = engram_get();
int64_t lim = (int64_t)limit; if (lim <= 0) lim = 100;
+1
View File
@@ -612,6 +612,7 @@ el_val_t engram_load(el_val_t path);
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_recall_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);