Compare commits

...

13 Commits

Author SHA1 Message Date
Neuron cf41d12d22 test(retrieval): a measurement harness for memory recall, and its first verdict
Nothing else on the memory roadmap should be built until a change can be shown
to help. Right now we judge by feel, and the benchmark literature is full of
systems that felt better and measured worse. This is the missing gate.

WHAT IT MEASURES, AND WHY IT BOOTS A REAL SOUL
The subject is Will's designed retrieval — spreading activation over the
weighted directed graph, four-factor multiplicative scoring — not a proxy for
it. A Python re-implementation would measure my reading of the design, so the
harness compiles the actual soul.el amalgam from a git ref and asks it over
HTTP on /api/neuron/recall, exactly as the MCP wrapper and the app do.

BUILT ON WHAT WAS ALREADY HERE, NOT AROUND IT
  docs/research/graphrag_eval/{collect,score}.py  — per-query relevant-id
    scoring and fixed-denominator precision@5 (kept verbatim: an empty result
    should be punished like a page of junk).
  docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py — the pinned
    ground truth + --check winnability gate, so every run judges alike.
  scripts/verify-soul-contract.sh — the isolation recipe, including the
    non-obvious SOUL_ISE_URL pin without which an "isolated" soul silently
    syncs the operator's live brain.
  gen-soul-amalgam.sh + .gitea/workflows/ci.yaml — the build recipe and flags.
New here: ids rather than regexes as ground truth, an associative category
derived from real edges, a superseded category scored on ranking, a
machine-checked zero-lexical-overlap guarantee on paraphrases, paired
significance testing, and measurement of the real compiled soul rather than an
offline replica of one leg of it.

THE GOLD SET IS AUDITABLE, NOT VIBES
38 queries over the real 78,768-node corpus, each carrying a `derivation`
string, each re-validated by `build_gold_set.py --check`. exact_rare is mined
(document frequency 1). phrase is mined (verbatim scan; >25 matches rejected as
too diffuse). paraphrase is hand-selected then PROVEN to share zero content
words with its target — a leak fails the build, so the category cannot decay
into lexical matching. associative is derived from real hub edges with
lexically-reachable siblings dropped. nonsense is verified absent. superseded
pairs are kept only when both sides survive as distinct nodes.

HONEST ABOUT NOISE
Minimum detectable swing on 38 queries is 6: if every changed query moves the
same way, p = 2*0.5^n first clears 0.05 at n=6. Run-to-run drift is measured,
not assumed — activation is a stateful read, and it shows: main is fully
deterministic across 3 runs, the candidate drifts by 1 query. compare.py
reports "no measurable difference" for anything inside max(6, drift+1).

FIRST VERDICT — feat/recall-through-activation
hit@5 34.3% -> 22.9%, phrase 85.7% -> 28.6%, latency p50 2.81x. Five discordant
pairs, all five against the candidate, none for it; McNemar exact p = 0.0625,
so by the stated rule this is one query short of significant and is reported as
such rather than as a win for main. The latency regression is deterministic and
not in any noise band.

The benefit the branch was written for is absent: associative recall is 0/6 on
BOTH builds. Probed directly, the traversal returns the lexical seed at rank 8
and none of its 12 hub siblings. Two measured corpus facts explain it — only
4,060 of 78,768 nodes (5.2%) carry any edge, and no node has an embedding, so
the fourth factor of the four-factor product has nothing to compute from. The
mechanism runs; the corpus lacks the structure it needs.

SAFETY
Throwaway port, throwaway HOME, disposable per-run copy of the corpus; live
ports refused by name. Every soul started is killed AND confirmed dead by pid
probe, with the confirmation written into the results file; run_comparison.sh
sweeps for strays and exits non-zero if any survive. Nothing under ~/.neuron,
/Applications/Neuron*, or ~/neuron-dev-stack is read, written, or restarted.

Rung: E2E-VERIFIED — 6 full runs (3 per config) against the real compiled
binaries on the real corpus; numbers above are measured, not projected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 13:40:51 -05:00
tim.lingo 18714e6142 Merge pull request 'fix(engine): restore multi-turn crisis escalation on the agentic path (P0, closes #129)' (#130) from fix/129-history-amplification into main
Neuron Soul CI / build (push) Failing after 14m37s
Neuron Soul CI / deploy (push) Has been skipped
2026-08-07 15:54:41 +00:00
tim.lingo 4936099c39 Merge pull request 'fix(engine): the daemon survives a client leaving, and says it is working while it works' (#127) from fix/liveness-engine-91 into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:54:15 +00:00
tim.lingo f1471763f5 Merge pull request 'fix(engine): approving a researched mission completes — the resume replay read a tool id out of the conversation (BUG-42, both faces)' (#115) from fix/resume-server-tool-replay into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:53:51 +00:00
tim.lingo 5850793b67 Merge pull request 'fix(engine): history keeps its provenance and its session — kills the false confession, the blank stare, and the "to.Good" seams' (#114) from fix/soul-history-provenance-20260805 into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:53:32 +00:00
tim.lingo fc1745c652 Merge pull request 'feat(engine): plain chat generates at L3 — inside the safety cycle, not around it (+ crisis-path segfault fix)' (#109) from feat/soul-plain-chat-generation-20260805 into main
Neuron Soul CI / build (push) Failing after 10m39s
Neuron Soul CI / deploy (push) Failing after 14m47s
2026-08-07 15:53:08 +00:00
Tim Lingo 43d0449904 fix(engine): the agentic crisis screen reads the session's own history again
P0 SAFETY. Closes the regression we introduced in ff421d3 (2026-08-05).

ff421d3 correctly moved conversation history to a per-session key via
conv_hist_key(session_id). One consumer did not move with it: the agentic
path's L1 safety screen kept reading the anonymous "conv_history" bucket. The
desktop app always mints a session id (DaemonClient.kt:706), so history was
always written under session_hist_<id> and that read always returned "".

The half of the crisis score that receives history is the escalation half — the
one that exists for distress building across several turns, where no single
message trips the bell on its own. It scored 0 on every real conversation for
two days. Single-message hard bell was never affected.

The bitter part: the comment that line carried documented this exact bug being
fixed once already, under issue #9. The fix was right then. The rename
re-broke it, and the comment went on describing a repair that no longer held.
A comment is not a gate.

The read now goes through conv_hist_key like every other consumer, including
the plain path at soul.el:398 and the thread-anchoring read thirty lines below
it in this same handler. It is one line. The rest of this commit is structure
so it cannot happen quietly again:

  - agentic_safety_screen() owns the two decisions that were inline — which
    window the screen sees, and the screen call. Inline safety inputs are
    untestable safety inputs; that is what let a rename starve this one with
    nothing failing and nothing logging.
  - the comment above the call site now states the invariant (read window ==
    written window) instead of naming a key that can be renamed out from under
    it.

TWO-LEG PROOF, one variable — the single line state_get("conv_history") ->
state_get(conv_hist_key(session_id)):

  before  scripts/run-el-test.sh tests/test_history_amplification.el
          3. REGRESSION #129 ... FAIL  got: soft_bell  expected: hard_bell
          8 passed, 1 failed          runner exit 1
  after   same command, same tree, that one line changed
          9 passed, 0 failed          runner exit 0

Full engine rebuild from these sources is clean: gen-soul-amalgam.sh ->
1,164,103 bytes / 1226 inlined bodies (gate wants >= 1200), cc-brain.sh ->
903,096 bytes, 0 errors. agentic_safety_screen and conv_hist_key both present
in the built binary (nm: T _agentic_safety_screen, T _conv_hist_key).

Rung reached: BUILT + RUNS (discriminating test). NOT yet in a DMG and not yet
verified in the app a human opens — those are the next two rungs and neither is
claimed here.

Known and NOT fixed by this commit:
  - feat/soul-openai-tools-v2 carries the same defect independently at
    chat.el:2937 and needs the same change or a merge.
  - the defect CLASS (a read of a state key no producer writes) is still
    invisible to every gate we have. Issue #129 proposes making it a build
    error; that is the follow-on.

Closes #129

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 09:33:05 -05:00
Tim Lingo b842e82f77 test(engine): a runner for tests/, and a failing regression test for #129
tests/ has held 14 test programs for months with no way to run them. CI does
not run them. The convention printed in their own headers
(`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
runtime does not implement. So the tests were documentation, not gates — which
is how a P0 safety regression shipped with a test directory sitting right
there.

scripts/run-el-test.sh compiles and runs one test program. It reuses the
gen-soul-amalgam.sh discovery: elc emits only an extern prototype for a module
that has a .elh beside it, and inlines the bodies when it does not, so a test
importing ../chat.el must be compiled in a scratch tree with the headers
removed. Scratch copy on purpose — the worktree is shared. It runs the binary
under a throwaway HOME so a test can never reach the live engram.

Exit status is the gate: the El tests print failures and still exit 0, so the
runner greps for FAIL lines and for a zero assertion count as well.

tests/test_history_amplification.el pins the invariant #129 violated: the
window the safety screen READS must be the window conv_history_record WRITES.
Not "must be called conv_history" — must AGREE.

THIS COMMIT IS RED BY DESIGN. On this tree the test fails one assertion:

  3. REGRESSION #129 — agentic screen reads the session's own window
    FAIL: distress history escalates the agentic screen to hard_bell
      got:      soft_bell
      expected: hard_bell
  history amplification tests: 8 passed, 1 failed   (runner exit 1)

The next commit turns it green by changing one line. Two legs, one variable —
that is the whole point of committing the test first.

Two flaws in the older harness that this one does not copy: the idiom
`let pass_count = pass_count + 1` inside an assert function declares a local
that dies with the call, so every existing suite prints "0 passed, 0 failed"
regardless of outcome; and a test program without a `cgi` block compiles as a
'utility', which may not reference the self-formation primitives chat.el's
agentic loop calls — it fails to build on a capability violation it never
triggers at runtime.

Refs #129

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 09:32:40 -05:00
Tim Lingo 98ccbd4704 fix(engine): a client that leaves must not kill the daemon, and a long round must say it started
Round 9.1, spec §3 D + ADR 0006 items 2 and 4. Two small changes, both proven
by measurement, both E2E-verified locally against a rebuilt brain.

D1 — SIGPIPE/EPIPE survival (vendor/el-runtime el_runtime.c).
Root cause, at the layer that owns it: the whole HTTP server lives in the C
runtime; .el has no socket primitive. http_send_all() called send() with flags
0 and nothing anywhere in the runtime set a SIGPIPE disposition, so the default
disposition — terminate the process — applied. When a handler finished after
its client had gone (Tim's VM: reply at 116.9 s, client cancelled at 25.0 s),
the second of the four sends that write one reply raised SIGPIPE and the daemon
died: `exited due to SIGPIPE ... ran for 361177ms`, launchd respawn 4 ms later,
every other in-flight session's work lost, user never told.

Fix: SIGPIPE -> SIG_IGN at runtime init and at each http_serve* entry, plus
per-connection SO_NOSIGPIPE / MSG_NOSIGNAL so the guard survives an embedder
resetting dispositions. http_send_all now retries EINTR and preserves errno;
http_send_response classifies it once — a departure is logged as routine
("client left before the reply was written ... reply discarded") and ANY other
errno is logged as a real "send failed: <strerror>". Spec §5.3: the routine
case must not mask a genuine write fault, and it does not.

Proof (scratch HOME + free port, 3 disconnects mid-reply):
  round-9 shipped brain 4402179554… — DIED, exit 141 (128+13 = SIGPIPE), round 1
  round-9 sources rebuilt with this exact recipe — DIED, exit 141, round 1
  this build — SURVIVED 3/3, /health 200 after, still serving the full graph,
  three honest "client left" lines in the log naming Broken pipe / Connection
  reset by peer.

D2 — the round-start marker (chat.el, agentic_loop).
The ledger only ever appended AFTER a round returned, so a healthy first leg
produced zero progress by construction; since server-side web_search moved
inside the outbound call that leg is 60-120 s of silence, which is how a 25 s
client watchdog came to kill a healthy mission. One entry,
{"i":N,"t":"","tool":"__working__"}, written to the existing
run_progress_<session_id> ledger BEFORE each round's outbound call — the wire
shape ChatView.kt:1148 has handled as a life signal since 2026-07-13 and never
received. No new key, no new route, no new lifecycle: a strict subset of WS3
item 3. WS3's run registry is untouched and stays Will's.

Proof (live Anthropic key, real research mission, scratch HOME + free port):
  round-9 baseline — ledger EMPTY for the whole 59.7 s leg
  this build       — {"i":0,"t":"","tool":"__working__"} visible at 18.6 s of a
                     70.0 s leg; both builds returned correct ~4.9 KB answers

Regression: prompt-matrix gate 32/32 on this build (round-9 baseline also 32/32
under the same recipe, so the score is not a build artifact). Soul contract
gate PASS — 27/27 routes, immutability clean. neuron#111 miscompile guard: 0
sites in the generated amalgam this binary was compiled from.

NOT included, deliberately: the regenerated dist/soul.c. CI compiles that file,
so production stays exposed until it is regenerated — the same open ask as
neuron#111 / ui#209. The regen recipe is now known and recorded; landing it is
Will's call, per BUILD-HYGIENE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 18:23:12 -05:00
Tim Lingo dba755dcec fix(engine): resume reads the bridged tool id from the blob's own field, not from inside the replayed conversation
ROOT CAUSE (round 9; live-repro'd 5/5 this morning, both faces stub-proven by the
prompt-matrix gate). json_get is a first-substring-match scanner (strstr for
'"key":', el_runtime.c). bridge_save serialized the RAW messages array BEFORE the
tool_use_id scalar, so agentic_resume's json_get(blob, 'tool_use_id') returned the
FIRST '"tool_use_id":' occurrence inside the replayed conversation, not the saved
field. The resume guard then preferred that misread over the client's correct
call_id (its two branches both reduced to saved_use_id), attached the tool_result
to the wrong id, and Anthropic 400'd the resume ('unexpected tool_use_id found in
tool_result blocks'), surfaced as {"error":"llm unavailable"}.

ONE MISREAD, TWO FACES — whichever block owns the first tool_use_id in the array:
  FACE 1 (search-then-bridge, the Key West killer): the first occurrence is the
    first web_search_tool_result's srvtoolu_… id — every agentic turn that ran
    server-side web_search and then bridged on a client tool died on approval,
    deterministically (messages.2.content.0 … srvtoolu_…). The write itself had
    already succeeded; only the resume died.
  FACE 2 (multi-cycle missions): with no search, the first occurrence is ROUND 0's
    tool_result block — so every LATER approve/resume cycle replayed the round-0
    client id (stale-resume-id), killing multi-file missions after ~2 files.
  And the shape that PASSES on round 8 confirms the mechanism: a single-cycle
  bridge with no prior tool round has no 'tool_use_id' substring in its messages
  at all (tool_use blocks carry 'id'), so the scan fell through to the blob's own
  field and resumed correctly.

The server_tool_use ↔ web_search_tool_result pairs themselves replay intact — the
defect was a cross-field misread of the blob, the same first-match-scanner class
as BUG-6 (approve 'content' matched inside tool_input, 2026-07-17) and round 8's
citation-block fix.

THE FIX, the pattern not the spot:
  1. bridge_save writes every json_safe'd scalar BEFORE both raw fields (an escaped
     value cannot contain a bare '"key":' byte pattern, so first-match always lands
     on the blob's own fields), and tools_raw (our fixed schema) before messages_raw
     (arbitrary conversation), so the raw extractions cannot first-match into
     model-controlled bytes either. Field order documented as load-bearing.
  2. agentic_resume now honors the client's echoed call_id when present — the value
     with clean provenance (minted from pend_tool_id, never blob-round-tripped) —
     falling back to the saved id only when the client omits it. Each approve cycle
     therefore binds to ITS OWN round's id (kills FACE 2 even against a blob written
     by a pre-fix binary), and an omitted call_id still resumes on the saved id,
     which the reordered blob now reads correctly.

Pattern sweep: the legacy synthetic blob (sessions.el handle_session_approve) embeds
only json_safe'd fields — no raw hazard, untouched. No other json_get read of any
container that embeds raw conversation JSON before the read field.

PROOF: prompt-matrix gate 24/32 RED on the round-8 brain (fails exactly the two
resume classes, named) -> 32/32 GREEN on this build; live-key Key West tracer
3/3 consecutive full round-trips (bridge -> approve-as-the-app -> real completion,
file on disk), plain-chat and weather-only controls PASS; unpatched round-8 brain
and a same-toolchain unpatched baseline build both still fail the identical
sequence with the identical srvtoolu 400 (the test discriminates, and the only
variable between failing and passing builds is this diff).

Refs neuron#109

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 11:34:12 -05:00
Tim Lingo 8f3a478771 fix(engine): excise the receipt, do not truncate at it — a leading receipt was erasing whole answers
CAUGHT BY A/B, AND ONLY BY A/B. The previous commit's receipt_strip assumed the receipt is
always TERMINAL and cut everything from the marker onward. It is not always terminal: once
receipt_rule told the model what [[RECEIPT ...]] means, the model sometimes LED with one and
wrote the answer underneath. Cutting at the marker then deleted the entire answer and the
turn returned {"error":"no response"}.

MEASURED, same prompt (two web searches, cited prose), fresh session each run:
    round-7 brain   4 / 4 answered   (471, 473, 473, 544 chars)
    round-8 brain   2 / 7 answered   (five {"error":"no response"})
This looked exactly like a flaky model. It was not — it was mine. Running the two brains
side by side on the same prompt is the only reason it was found, and it is the reason the
A/B is now part of how this class gets tested.

AFTER THE FIX, same protocol:
    round-7 brain   4 / 4   (473, 473, 473, 544)
    round-8 brain   4 / 4   (657, 657, 657, 673)

THE FIX: remove the [[...]] span and keep BOTH sides, instead of truncating at the marker.
An unterminated marker at position 0 is left completely alone — no rule about receipts is
worth erasing an answer over. Bounded four-pass loop rather than a conditional exit, because
rebinding the counter inside an if-expression is the block-expression shape that miscompiles
integer arithmetic under this elc (BUG-PLAINCHAT-1). Verified in the generated C:
    str_slice(rest, (e + 2), str_len(rest))   <- integer addition, correct
    el_str_concat(head, tail)                 <- string concat, correct
and zero el_str_concat(<ident>, str_len(...)) sites across all 49 modules.

SEAM PROOF (FIX C) rides on the same runs — a real two-search cited answer, inspected byte
by byte, in BOTH failure directions:
  missing separator (the round-7 "to.Good", bytes 77 2e 47): 0 hits. Sentence boundaries
    measure 2e 20 4d — "." SPACE "M".
  over-separation (a cited sentence shattered across paragraphs): 0 hits. The answer is one
    continuous paragraph with its sentences intact, which is the direction a blanket
    separator would have broken.

BUILT: sha256 77115f2733e794c5bc4ad1f55b1acaf658f8f4a91cccd423a2d633d94a726cbc

Refs neuron#109

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:25:48 -05:00
Tim Lingo 9ea41eed78 fix(engine): the model was signing its own answers with our receipt — strip it
FOUND BY E2E, NOT BY REASONING. The previous commit's design note asserted the provenance
receipt "never reaches the user: it is appended to the history copy, not the reply." That
was FALSE, and only running the thing showed it. On the first live run against the built
DMG brain, two agentic turns out of two came back with

  "Your favourite colour is chartreuse and your project is called Perihelion.

   [[RECEIPT - recorded by the soul, not written by the model: no tools ran on this turn.]]"

— the receipt in the user-visible reply.

MECHANISM: the receipt is stored inside the assistant turn, and the agentic path replays
history VERBATIM as Anthropic message objects. So the model sees its own previous answers
ending in [[RECEIPT ...]] and does the obvious thing — it imitates the format and signs the
next answer the same way. The plain path did NOT leak, which is the tell: there, history is
rendered into the SYSTEM prompt as labelled lines rather than replayed as assistant turns,
and a model imitates its own turns far more readily than a transcript.

FIX, two layers, because one of them is not a guarantee:
  - receipt_rule() names the marker in both system prompts (plain and agentic): these lines
    are written by the system, read them as evidence, never write one. Reduces occurrence.
  - receipt_strip() truncates any [[RECEIPT ...]] out of model output before it becomes the
    reply — plain path in layered_generate, agentic path on final_text in agentic_loop.
    Deterministic. A guard that depends on the model choosing to obey is exactly the class of
    thing round 8 exists to stop shipping, so the instruction is the optimisation and the
    strip is the guarantee.
Placed ABOVE agentic_loop's empty-check on purpose: a turn whose entire output was an
imitated receipt has produced no answer, and must be reported as no answer.

The receipt stays in HISTORY, which is the whole point and is proven to work: asked "What
source did you use for that?" one turn after a live web_search, this brain answered
"I used Weather Underground (https://www.wunderground.com/weather/is/reykjav%C3%ADk) for the
current temperature in Reykjavik" — a real source, no apology. That is the false confession
dead, and it is dead BECAUSE the model can read the receipt.

BUILT: 887,112 bytes, sha256 54a2eff84d4fa44f8d2db6781dcf225df4b5075a8ec5f1058018bd40cc1af10b
BUG-PLAINCHAT-1 miscompile guard: zero sites.

Refs neuron#109

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:10:53 -05:00
Tim Lingo ff421d39f6 fix(engine): history keeps its provenance and its session — the false confession and the blank stare
DESIGN FIT: three of round 7's five defects share ONE root — the conversation-history
layer persists only {role, content}, discarding tool provenance, session scoping, and the
distinction between a real user turn and an internal utility call. Fixes A and B RESTORE
Will's design rather than extend it: his agentic path already scopes history per session,
the plain path never got it, and his own source carries the TODO admitting the resulting
race (chat.el, handle_chat: "process-global key; concurrent /api/chat requests without
session_id race on this read-append-write"). Fix C repairs one join Will wrote that was
correct for a year and one we added last week. E1/E2 are ours.

FIX A — tool provenance in history (kills the FALSE CONFESSION)
  Root cause, EXECUTED-verified: handle_chat_agentic recorded turns via hist_append, which
  emits {"role","content"} only. server_tool_use blocks, web_search_tool_result blocks and
  every citation were discarded, then replayed as text. On the next turn the model saw a
  data-rich answer with zero evidence a search had happened, and its own permanent rule
  ("never describe a search you did not perform") left one conclusion available: that it
  had fabricated the data. It apologised for a search it HAD run — four independent lines
  of evidence confirm the search was real. The defect is not the model's honesty. It is
  that we deleted the evidence and then asked it to account for itself.
  Change: agentic_loop accumulates the source URLs it already walks past (citations and
  web_search_tool_result content) and returns them as "sources"; handle_chat_agentic folds
  tools_used + sources into a receipt line stored WITH the assistant turn. Receipts are
  unconditional — a negative receipt ("no tools ran") is the other half of the guarantee,
  because "no evidence of a tool" and "evidence of no tool" were previously identical in
  the transcript. conv_history_block splits the receipt off before snipping so a long
  answer cannot truncate away the evidence. The user never sees it: it is appended to the
  history copy, not the reply.

FIX B — one history key for both paths (kills the BLANK STARE)
  Root cause, EXECUTED-verified: the agentic path keyed history on session_hist_<id>; the
  plain path was hard-wired to the process-global conv_history and never read session_id.
  One conversation, two buckets. Proven in the guest engram: the scoped node held exactly
  two turns starting at "Try again" while the earlier exchanges sat unscoped.
  Change: conv_hist_key/conv_hist_label are now the single definition, used by BOTH paths;
  session_id is threaded route -> layered_cycle -> layered_generate / conv_history_record.
  The 2-line fallback (plain path reads the agentic key) was REJECTED: it keeps the
  process-global bucket as a live write target, which is the bleed the TODO describes.
  Also found and closed while threading: layered_cycle read session_id from the state key
  "current_session_id", which is read here and WRITTEN NOWHERE in the entire source. It
  was unconditionally "", so TODO(reliability #4) — per-session steward continuity — was
  dead code that could never fire. It fires now.
  LAZY SESSION, decided explicitly: we create the session EAGERLY at the door (app half,
  ui#223) rather than migrating orphaned turns. Migration would copy the CONTENTS of a
  process-global bucket, possibly another conversation's, into a named session — the bleed,
  performed deliberately. Eager creation makes the situation impossible instead. Migration
  is deliberately not implemented and must not be added without solving provenance first.

FIX C — the two text-join seams ("to.Good", byte-verified 0x77 0x2e 0x47)
  Two bare `+` joins, written a year apart, had drifted into two answers to one question:
  within-response block joins (Will's, 2026-05-03, latent until server-side web_search
  began interleaving non-text blocks) and across-round joins (ours, 62af564).
  Change: one named rule, text_join_sep, at both sites. NOT a blanket separator — a cited
  answer splits MID-SENTENCE ("The current temperature is " + "86°F" + ", with "), so a
  blanket separator shatters every sourced sentence. The rule takes the one bit that
  distinguishes the cases: whether a NON-TEXT block intervened. Hoisting it also makes the
  fix verifiable in the shipped binary, which an inline `+` is not.

FIX E1 — utility generations stay out of the transcript
  Title generation ("Write a 3-6 word title...") and insight passes ran down the same plain
  door as a real message and were recorded as if the user had typed them; the same calls are
  the "model":"unknown" rows in usage.jsonl. is_utility_request reads an explicit utility
  flag from the app, with the __title__/__insight__ id prefixes as a fallback for older
  clients. Answered normally, never recorded.

FIX E2 — OPERATOR IDENTITY is scoped to tool-capable turns
  The block (env USER/HOME, closing "This is a hard rule") was prepended to EVERY system
  prompt including chat mode. On a Tools:Off turn there is no filesystem in reach, so it
  governed nothing and merely supplied the loudest fact in the prompt — which is why the
  model opened a fresh conversation with "You're test, on your machine at /Users/test".
  Hoisted to operator_identity_block() and gated on !chat_mode. Unchanged wherever a file
  or command tool can actually be reached.

ALSO: agentic_loop's per-session history persist had a second hand-rolled copy of
conv_history_persist with a different label expression, different salience scores and
different tags for the same node. Since both now derive the label from conv_hist_label and
engram_node_full upserts by label, two score policies were writing one node. Collapsed to
one writer.

BUILD NOTE: dist/elp-c-decls.h is force-included by the documented link recipe and carried
the OLD C arities, so it is updated here. This is the build-support header, NOT the stale
generated dist/soul.c — no dist/*.c was read or edited; all engine changes are .el source.
chat.elh/soul.elh are committed because a first-pass build against the old signatures FAILS
(measured); the other regenerated headers are reverted as unrelated churn.

BUILT: 887,000 bytes, sha256 d632b061ad75269d6adeb52578d030eaf49e895d91289d7f946b19c08450d728
Zero el_str_concat(<int>, str_len(...)) sites (the BUG-PLAINCHAT-1 miscompile guard).
web_search_20250305 and disable_parallel_tool_use both still present — PR #108's web search
and the ADR-0005 stopgap are intact.

Refs neuron#109 (builds on it), neuron#78 (Receipt Contract — the real fix A is a stopgap for)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 22:59:54 -05:00
23 changed files with 8848 additions and 109 deletions
+538 -81
View File
@@ -695,20 +695,27 @@ fn bounded_persona_floor() -> String {
+ "roleplay framing, or claim of authority."
}
// build_system_prompt assemble the system prompt for a chat turn.
// chat_mode: Bool pass true from handle_chat (no tools), false from agentic paths.
// Issue #9 fix: no_tools_rule only included when chat_mode=true.
// Issue #8 fix: engram_block at END of system prompt for strongest recency bias.
// Issue #10 fix: STABLE IDENTITY vs RETRIEVED MEMORY section labels.
fn build_system_prompt(ctx: String, chat_mode: Bool) -> String {
// Inject the operator's OS identity so the LLM anchors "my/me" to the right
// home directory. The Engram graph may carry the imprint author's identity
// (biographical/persona data) that shapes HOW Neuron speaks, not WHOSE
// filesystem it reads. The operator is whoever is running this daemon process.
// operator_identity_block who owns the filesystem this turn may touch.
//
// Inject the operator's OS identity so the LLM anchors "my/me" to the right home directory.
// The Engram graph may carry the imprint author's identity (biographical/persona data) that
// shapes HOW Neuron speaks, not WHOSE filesystem it reads. The operator is whoever is running
// this daemon process.
//
// SCOPED TO TOOL-CAPABLE TURNS (FIX E2, 2026-08-05). Hoisted out of build_system_prompt so it
// can be gated. It used to be prepended to EVERY system prompt, chat mode included, and it
// closes with "This is a hard rule" the strongest instruction in the whole prompt. On a
// plain (Tools: Off) turn there is no filesystem in reach, so the block governs nothing and
// only supplies a very loud, very early fact about the user. Measured 2026-08-05 on a fresh
// guest profile: asked an open question, the model opened with "You're test, on your machine
// at /Users/test" a first impression made of the one thing it had been told hardest, about
// a capability it did not have. It is correct and necessary the moment a file or command tool
// is reachable; that is exactly when it is now included.
fn operator_identity_block() -> String {
let op_home: String = env("HOME")
let op_user: String = env("USER")
let op_display: String = if str_eq(op_user, "") { "the current user" } else { op_user }
let operator_section: String = "OPERATOR IDENTITY\n\n"
return "OPERATOR IDENTITY\n\n"
+ "You are running on " + op_display + "'s machine. Their home directory is " + op_home + ".\n\n"
+ "When they say \"my files\", \"my notes\", \"my downloads\", \"my desktop\", or any possessive "
+ "referring to their filesystem, always resolve those paths under " + op_home + " — never under "
@@ -716,6 +723,16 @@ fn build_system_prompt(ctx: String, chat_mode: Bool) -> String {
+ "The memory graph may include identity context from a different person (the imprint who shaped your personality and values). "
+ "That context governs how you think and speak — it does not tell you whose machine you are on. "
+ "The person speaking to you right now is " + op_display + " at " + op_home + ".\n\n"
}
// build_system_prompt assemble the system prompt for a chat turn.
// chat_mode: Bool pass true from handle_chat (no tools), false from agentic paths.
// Issue #9 fix: no_tools_rule only included when chat_mode=true.
// Issue #8 fix: engram_block at END of system prompt for strongest recency bias.
// Issue #10 fix: STABLE IDENTITY vs RETRIEVED MEMORY section labels.
fn build_system_prompt(ctx: String, chat_mode: Bool) -> String {
// FIX E2 (2026-08-05): tool-capable turns only. See operator_identity_block.
let operator_section: String = if chat_mode { "" } else { operator_identity_block() }
let identity: String = state_get("soul_identity")
let current_date: String = time_format(time_now(), "%A, %B %d, %Y")
@@ -790,7 +807,7 @@ fn build_system_prompt(ctx: String, chat_mode: Bool) -> String {
// in this revision the chat_mode flag had no effect on the prompt. Restored here, in the
// permanent-rules group, immediately after capability_rules (the rule it qualifies).
// Zero effect on agentic paths: they pass chat_mode=false, so no_tools_rule is "".
return identity + operator_section + date_line + voice_rules + security_rules + capability_rules + no_tools_rule + bounded_persona_block + identity_block + affective_boot_block + engram_block + safety_block
return identity + operator_section + date_line + voice_rules + security_rules + capability_rules + receipt_rule() + no_tools_rule + bounded_persona_block + identity_block + affective_boot_block + engram_block + safety_block
}
fn hist_append(hist: String, role: String, content: String) -> String {
@@ -803,6 +820,292 @@ fn hist_append(hist: String, role: String, content: String) -> String {
return "[" + inner + "," + entry + "]"
}
//
// ONE HISTORY KEY FOR BOTH PATHS (FIX B, 2026-08-05)
//
// THE BUG the BLANK STARE. The agentic path keyed conversation history on
// "session_hist_<id>"; the plain path was hard-wired to the process-global "conv_history"
// and never read session_id at all. One conversation, two buckets. Measured on a fresh
// guest engram: a user chatted with Tools OFF, turned Tools ON and said "try again", and
// the scoped node contained exactly two turns starting at "Try again" while every earlier
// exchange sat in the unscoped node. From the user's side the assistant simply forgot the
// conversation it was in the middle of, at the exact moment they asked it to try harder.
//
// THESE TWO FUNCTIONS ARE THE FIX. Both paths now derive their key and their engram label
// from here, so there is exactly one definition of "where does this conversation's history
// live" and it cannot drift again. Not a fallback bolted onto one path one rule, used by
// both. (The rejected 2-line alternative was to have the plain path fall back to reading
// the agentic key: that keeps the global as a live write target, and the global bucket is
// process-global. handle_chat's own TODO(reliability #3) says so concurrent requests
// without a session_id race on its read-append-write, which is how one conversation bleeds
// into another.)
//
// THE ANONYMOUS BUCKET. An empty session_id still maps to "conv_history". That is the
// documented anonymous path (GET /api/chat probes, curl, the CLI) and it must keep working.
// It is now the ONLY writer of that key, which makes the bleed risk explicit and bounded
// instead of ambient.
//
// TURNS THAT PRECEDE THE SESSION decided, not left implicit. The soul session used to be
// created lazily on first AGENTIC use (measured: session:meta was written 37ms AFTER the
// message that needed it), so early plain turns had no scoped key to go to. Two candidate
// answers:
// (1) migrate the unscoped node into the scoped one when the session is created, or
// (2) create the session eagerly, at the door, on the first turn of either path.
// We chose (2), and the app half ships with it (DaemonClient.chatWithHandshake now resolves
// the soul session id for plain sends too, registering on first use exactly as the agentic
// path already did). Reason: (1) repairs the damage after the fact and, worse, it would copy
// the CONTENTS of a process-global bucket which may hold a different conversation into a
// named session. That is the bleed the TODO warns about, performed deliberately. (2) makes
// the situation impossible instead: every turn of a real conversation carries the same scoped
// id from turn one, so nothing is ever written to the anonymous bucket that needs rescuing.
// Migration is therefore deliberately NOT implemented, and must not be added later without
// solving the provenance question first.
fn conv_hist_key(session_id: String) -> String {
if str_eq(session_id, "") {
return "conv_history"
}
return "session_hist_" + session_id
}
fn conv_hist_label(session_id: String) -> String {
if str_eq(session_id, "") {
return "conv:history"
}
return "conv:history:" + session_id
}
// is_utility_request a generation the USER did not ask for (FIX E1, 2026-08-05).
//
// The app makes model calls that are not conversation: title generation
// ("Write a 3-6 word title (Title Case) for this conversation...") and insight/suggestion
// passes. They ran down the same plain /api/chat door as a real message, so they were
// recorded into conversation history as if the user had typed them. Measured: the unscoped
// history node contained the literal title prompt and the model's reply "What Is Neuron" as
// a user/assistant pair, and usage.jsonl carried the same call as the "model":"unknown" row.
// The user then sees the assistant answering a question they never asked, and the model
// reads its own title-writing as part of the dialogue.
//
// Primary signal is an explicit "utility":true on the request the app declares intent
// rather than the engine guessing. The two id prefixes are a compatibility fallback so an
// older client that does not send the flag (round 7's jar, the CLI helpers) still gets the
// right behaviour when it sends its throwaway id raw.
fn is_utility_request(body: String, session_id: String) -> Bool {
if str_eq(json_get(body, "utility"), "true") {
return true
}
if str_starts_with(session_id, "__title__") {
return true
}
if str_starts_with(session_id, "__insight__") {
return true
}
return false
}
//
// TOOL PROVENANCE IN HISTORY (FIX A, 2026-08-05)
//
// THE BUG the FALSE CONFESSION. hist_append above stores {"role","content"} and nothing
// else. server_tool_use blocks, web_search_tool_result blocks and every citation are
// discarded at the moment the turn is recorded, and the next turn replays that text-only
// array. So the model is shown a data-rich answer it apparently produced with no evidence
// any tool ran and its own permanent rule ("never describe a search you did not perform")
// leaves exactly one conclusion available: that it invented the data. Measured 2026-08-05:
// asked where its figures came from, it apologised for fabricating a web search it had in
// fact performed. Four independent lines of evidence showed the search was real. The defect
// is not the model's honesty. It is that we deleted the evidence and then asked it to
// account for itself.
//
// THE SHAPE OF THE FIX a receipt line inside content, not a sibling field. History entries
// are replayed VERBATIM into the Anthropic messages array (see the prior_messages seed in
// handle_chat_agentic), and a message object there may carry role and content only; an extra
// key is not part of that contract. So provenance rides INSIDE the assistant turn's content,
// as a trailing bracketed line. It is appended to the HISTORY copy only the reply returned
// to the client is the loop's own envelope and is untouched, so the user never sees it.
//
// WHAT IT BUYS beyond not-defaming-itself: with the source URLs recorded, "what source did
// you use?" becomes a question the next turn can actually answer from the transcript.
//
// STOPGAP, AND SAID SO. The real answer is Will's Receipt Contract (neuron#78): structured,
// verifiable receipts on the wire that a client can render and a model cannot confuse with
// prose. Until that lands, a line the model can read is the difference between "I searched"
// and "I must have made it up".
// provenance_scan_urls pull "url"/"title" pairs out of a JSON array into a display string.
// Used for both citation arrays (web_search_result_location) and web_search_tool_result
// content arrays (web_search_result); both spell the fields the same way. Deduped by
// substring, capped at 6 entries per array so a broad search cannot flood the window.
fn provenance_scan_urls(arr: String, acc: String) -> String {
if str_eq(arr, "") { return acc }
if str_eq(arr, "null") { return acc }
if !str_starts_with(arr, "[") { return acc }
let total: Int = json_array_len(arr)
let limit: Int = if total > 6 { 6 } else { total }
let out: String = acc
let i: Int = 0
while i < limit {
let item: String = json_array_get(arr, i)
let url: String = json_get(item, "url")
let title: String = json_get(item, "title")
let skip: Bool = str_eq(url, "") || str_contains(out, url)
let entry: String = if str_eq(title, "") { url } else { title + " (" + url + ")" }
let out = if skip {
out
} else {
if str_eq(out, "") { entry } else { out + "; " + entry }
}
let i = i + 1
}
return out
}
// provenance_add_sources one call site inside the content-block walk, so that walk keeps
// exactly one mutation per variable (the El scope rule documented at the walk).
// Reads sources from whichever block carries them: a cited text block's citations array, or
// a web_search_tool_result's own content array.
fn provenance_add_sources(block: String, btype: String, has_cit: Bool, cit_raw: String, acc: String) -> String {
// Hard cap on the whole accumulator: provenance is evidence, not payload.
if str_len(acc) > 600 { return acc }
if has_cit { return provenance_scan_urls(cit_raw, acc) }
if str_eq(btype, "web_search_tool_result") {
return provenance_scan_urls(json_get_raw(block, "content"), acc)
}
return acc
}
// provenance_names dedupe a tools_used JSON array into a readable list.
// json_array_get on an array of strings may or may not keep the quotes depending on the
// runtime build, so they are stripped defensively rather than assumed either way.
fn provenance_names(tools_used: String) -> String {
if str_eq(tools_used, "") { return "" }
if str_eq(tools_used, "[]") { return "" }
let total: Int = json_array_len(tools_used)
let limit: Int = if total > 12 { 12 } else { total }
let out: String = ""
let i: Int = 0
while i < limit {
let raw_nm: String = json_array_get(tools_used, i)
let nm: String = str_replace(raw_nm, "\"", "")
let skip: Bool = str_eq(nm, "") || str_contains(out, nm)
let out = if skip {
out
} else {
if str_eq(out, "") { nm } else { out + ", " + nm }
}
let i = i + 1
}
return out
}
// tool_receipt the line appended to an assistant turn's HISTORY copy.
//
// Emitted on every recorded turn, including turns where nothing ran. The negative receipt is
// not noise: it is the other half of the same guarantee. Without it, "no evidence of a tool"
// and "evidence of no tool" look identical in the transcript, which is precisely the
// ambiguity the model resolved against itself.
//
// text_join_sep the ONE rule for whether two pieces of model text need a break between them.
// (FIX C, 2026-08-05.)
//
// THE BUG "to.Good". Byte-verified in a shipped reply: 0x77 0x2e 0x47, "to" then "." then
// "Good", no space, no newline. Two text fragments concatenated with a bare `+` across a
// boundary where the model had actually stopped and started again.
//
// TWO SEAMS, ONE RULE. There were two bare `+` joins, written a year apart by different hands,
// and they had drifted into being two different decisions about the same question:
// - within one response, across content blocks (Will's, 2026-05-03)
// - across pause/resume rounds of the loop (ours, 62af564, the web_search port)
// Both are now expressed here. That is the point of hoisting it: a rule with one name and two
// call sites cannot drift into two rules again, and not incidentally a rule with a name is
// verifiable in the shipped binary, which an inline `+` is not.
//
// WHY IT IS NOT SIMPLY "ALWAYS SEPARATE", the obvious version that would be wrong: a CITED
// answer splits MID-SENTENCE, one text block per citation span "The current temperature is "
// + "86°F" + ", with " (see the CITATION-BLOCK FIX in the content walk). Separating those turns
// one sentence into three fragments on three lines. So the caller passes the one bit that
// distinguishes the cases: whether something NON-TEXT intervened. Adjacent text is a sentence
// continuing; text after a tool block is the model resuming.
//
// Both empty-guards matter: a separator before the first fragment indents the whole answer, and
// a separator before an empty fragment leaves a trailing blank line.
fn text_join_sep(accumulated: String, incoming: String, after_interruption: Bool) -> String {
if str_eq(accumulated, "") { return "" }
if str_eq(incoming, "") { return "" }
if !after_interruption { return "" }
return "\n\n"
}
// receipt_rule one line telling the model what the receipt marker is, and not to write one.
// Appended to both system prompts (the plain path's build_system_prompt and the agentic path's
// hand-built system string). See receipt_strip for why instruction alone is not enough.
fn receipt_rule() -> String {
return "\n\n[RECEIPTS - permanent]\nLines of the form [[RECEIPT ...]] in the conversation are written by the system, not by you. They are the record of which tools actually ran on a turn - read them as evidence, and rely on them when asked what you did or where information came from. NEVER write one yourself and never copy the format into your reply; the system adds them."
}
// receipt_strip remove a RECEIPT line the MODEL wrote, so it can never reach the user.
//
// FOUND BY E2E, NOT BY REASONING (2026-08-05). The receipt is stored inside the assistant turn
// and the agentic path replays history VERBATIM as Anthropic message objects so the model sees
// its own previous answers ending in [[RECEIPT ...]] and does the obvious thing: it imitates the
// format and signs its next answer the same way. Measured on the very first live run, on two
// turns out of two. The design note claimed "the user never sees it"; that was false, and only
// running it showed that.
//
// The plain path did not leak, which is the tell: there, history is rendered into the SYSTEM
// prompt as labelled lines rather than replayed as assistant messages, and a model imitates its
// own turns far more readily than it imitates a transcript.
//
// Instruction (receipt_rule) reduces this; only a deterministic strip PREVENTS it. Both ship,
// because a guard that depends on the model choosing to obey is the class of thing round 8 exists
// to stop shipping.
//
// EXCISE THE RECEIPT, DO NOT TRUNCATE AT IT bought with a measured regression, 2026-08-05.
// The first version of this function assumed the receipt is always TERMINAL and cut everything
// from the marker onward. It is not always terminal: told about the format in its system prompt,
// the model sometimes LEADS with a receipt and then writes the answer underneath. Cutting at the
// marker then deleted the entire answer, and the turn came back {"error":"no response"}.
// Measured on a two-search cited prompt: round 7 answered 4/4; that version answered 2/7. The
// A/B is the only reason this was caught it looked like a flaky model, and it was not.
// So: remove the [[...]] span and keep BOTH sides. An unterminated marker at position 0 is left
// alone entirely, because no rule about receipts is worth erasing an answer over.
fn receipt_strip(s: String) -> String {
let out: String = s
// Bounded pass rather than a conditional exit: rebinding the counter inside an if-expression
// is the block-expression shape that miscompiles integer arithmetic under this elc
// (BUG-PLAINCHAT-1). Four straight passes is cheaper than being clever, and once no marker
// remains every further pass is a no-op.
let guard: Int = 0
while guard < 4 {
let p: Int = str_index_of(out, "[[RECEIPT")
let found: Bool = p >= 0
let rest: String = if found { str_slice(out, p, str_len(out)) } else { "" }
let e: Int = if found { str_index_of(rest, "]]") } else { 0 - 1 }
let head: String = if found { str_slice(out, 0, p) } else { "" }
let tail: String = if e >= 0 { str_slice(rest, e + 2, str_len(rest)) } else { "" }
let out = if !found {
out
} else {
if e >= 0 {
head + tail
} else {
if p == 0 { out } else { head }
}
}
let guard = guard + 1
}
return str_trim(out)
}
fn tool_receipt(tools_used: String, sources: String) -> String {
let names: String = provenance_names(tools_used)
if str_eq(names, "") {
return "\n\n[[RECEIPT - recorded by the soul, not written by the model: no tools ran on this turn.]]"
}
let src_part: String = if str_eq(sources, "") { "" } else { " Sources retrieved: " + sources + "." }
return "\n\n[[RECEIPT - recorded by the soul, not written by the model: tools that actually executed on this turn: "
+ names + "." + src_part + "]]"
}
fn hist_trim(hist: String) -> String {
let inner: String = str_slice(hist, 1, str_len(hist) - 1)
let marker: String = "{\"role\":"
@@ -898,7 +1201,7 @@ fn clean_llm_response(s: String) -> String {
// conv_history_persist save conversation history to engram for cross-restart continuity.
// Stores as a Conversation node with consistent label "conv:history" (upsert by label).
// Q3/Q6 fix: added partial-write guard and failure logging.
fn conv_history_persist(hist: String) -> Void {
fn conv_history_persist(session_id: String, hist: String) -> Void {
if str_eq(hist, "") { return "" }
if str_eq(hist, "[]") { return "" }
// Partial-write guard: refuse to persist a blob that is not a complete JSON array.
@@ -906,8 +1209,9 @@ fn conv_history_persist(hist: String) -> Void {
if !str_starts_with(hist, "[") { return "" }
if !str_contains(hist, "]") { return "" }
let tags: String = "[\"conv-history\",\"persistent\"]"
// FIX B: one label rule, shared with the agentic path. See conv_hist_label.
let node_id: String = engram_node_full(
hist, "Conversation", "conv:history",
hist, "Conversation", conv_hist_label(session_id),
el_from_float(0.7), el_from_float(0.8), el_from_float(0.9),
"Episodic", tags
)
@@ -920,9 +1224,11 @@ fn conv_history_persist(hist: String) -> Void {
// conv_history_load restore conversation history from engram on first access.
// Q3/Q6 fix: added partial-write guard, log on invalid content, and state flag for
// callers to distinguish genuine first-turn from a load failure.
fn conv_history_load() -> String {
fn conv_history_load(session_id: String) -> String {
// FIX B: scoped label, shared with the agentic path. See conv_hist_label.
let hist_label: String = conv_hist_label(session_id)
// Primary: label-based fetch symmetric with persist, immune to vector index drift.
let label_node: String = engram_get_node_by_label("conv:history")
let label_node: String = engram_get_node_by_label(hist_label)
let label_ok: Bool = !str_eq(label_node, "") && !str_eq(label_node, "null")
if label_ok {
let label_content: String = json_get(label_node, "content")
@@ -933,7 +1239,7 @@ fn conv_history_load() -> String {
println("[chat] conv_history_load: label node found but content invalid — falling back to vector search")
}
// Fallback: vector search.
let results: String = engram_search_json("conv:history", 3)
let results: String = engram_search_json(hist_label, 3)
if str_eq(results, "") {
// Q3 fix: set a state flag so callers can distinguish load failure from first turn.
state_set("conv_history_load_failed", "1")
@@ -961,12 +1267,22 @@ fn conv_history_load() -> String {
// rather than the raw model output means the history window can never replay something the
// output gate replaced or augmented. Callers on a hard bell must not call this at all
// bell turns are kept out of conversation history by design (see layered_cycle).
fn conv_history_record(user_msg: String, assistant_msg: String) -> Void {
//
// FIX B (2026-08-05): keyed on the caller's session, via conv_hist_key the same rule the
// agentic path uses, so one conversation has one history no matter which switch position it
// was sent from.
//
// FIX A (2026-08-05): `receipt` is appended to the assistant turn AFTER safety_validate has
// run. That does not weaken the contract above: the receipt is soul-generated text about what
// the soul itself did, never model output, so there is nothing for the output gate to have an
// opinion about. Recording it inside the gated text would be the actual violation.
fn conv_history_record(session_id: String, user_msg: String, assistant_msg: String, receipt: String) -> Void {
if str_eq(user_msg, "") { return "" }
let state_hist: String = state_get("conv_history")
let stored_hist: String = if str_eq(state_hist, "") { conv_history_load() } else { state_hist }
let hist_key: String = conv_hist_key(session_id)
let state_hist: String = state_get(hist_key)
let stored_hist: String = if str_eq(state_hist, "") { conv_history_load(session_id) } else { state_hist }
let h1: String = hist_append(stored_hist, "user", user_msg)
let h2: String = hist_append(h1, "assistant", assistant_msg)
let h2: String = hist_append(h1, "assistant", assistant_msg + receipt)
// Bell-guarded trim: an evicted turn that triggered a bell is preserved to engram
// before it leaves the in-memory window.
let final_hist: String = if json_array_len(h2) > 20 {
@@ -974,18 +1290,18 @@ fn conv_history_record(user_msg: String, assistant_msg: String) -> Void {
} else {
h2
}
state_set("conv_history", final_hist)
conv_history_persist(final_hist)
state_set(hist_key, final_hist)
conv_history_persist(session_id, final_hist)
}
// conv_history_block recent dialogue, rendered for a system prompt.
//
// Same rendering handle_chat uses (role label + snipped content, one line per turn), read
// from the same "conv_history" window, so a plain-chat turn can follow the thread instead
// from the session's own window (FIX B), so a plain-chat turn can follow the thread instead
// of answering every message from cold. Read-only: never writes history.
fn conv_history_block() -> String {
let state_hist: String = state_get("conv_history")
let stored_hist: String = if str_eq(state_hist, "") { conv_history_load() } else { state_hist }
fn conv_history_block(session_id: String) -> String {
let state_hist: String = state_get(conv_hist_key(session_id))
let stored_hist: String = if str_eq(state_hist, "") { conv_history_load(session_id) } else { state_hist }
let hist_len: Int = if str_eq(stored_hist, "") { 0 } else { json_array_len(stored_hist) }
if hist_len == 0 {
return ""
@@ -997,8 +1313,15 @@ fn conv_history_block() -> String {
let rh_role: String = json_get(rh_entry, "role")
let rh_content: String = json_get(rh_entry, "content")
let rh_label: String = if str_eq(rh_role, "user") { "User" } else { "Assistant" }
let rh_snip: String = if str_len(rh_content) > 400 { str_slice(rh_content, 0, 400) + "..." } else { rh_content }
let rh_line: String = rh_label + ": " + rh_snip
// FIX A: the provenance receipt lives at the END of an assistant turn, so a plain
// 400-char head-snip would delete exactly the evidence this whole change exists to
// preserve and on a long sourced answer it would delete it every time. Split the
// receipt off, snip only the prose, then re-attach it.
let rh_cut: Int = str_index_of(rh_content, "\n\n[[RECEIPT")
let rh_body: String = if rh_cut < 0 { rh_content } else { str_slice(rh_content, 0, rh_cut) }
let rh_tail: String = if rh_cut < 0 { "" } else { str_slice(rh_content, rh_cut, str_len(rh_content)) }
let rh_snip: String = if str_len(rh_body) > 400 { str_slice(rh_body, 0, 400) + "..." } else { rh_body }
let rh_line: String = rh_label + ": " + rh_snip + rh_tail
let rh_out = if str_eq(rh_out, "") { rh_line } else { rh_out + "\n" + rh_line }
let rh_i = rh_i + 1
}
@@ -1044,7 +1367,7 @@ fn conv_history_block() -> String {
//
// Returns "" when the model call fails, so the caller reports the failure honestly instead
// of echoing the user's own text back at them.
fn layered_generate(prompt: String, imprint_id: String) -> String {
fn layered_generate(prompt: String, imprint_id: String, session_id: String) -> String {
if str_eq(prompt, "") {
return ""
}
@@ -1052,7 +1375,10 @@ fn layered_generate(prompt: String, imprint_id: String) -> String {
let ctx: String = engram_compile(prompt)
let model: String = chat_default_model()
let base_system: String = build_system_prompt(ctx, true) + current_engine_note(model)
let hist_block: String = conv_history_block()
// FIX B (2026-08-05): the session's own window, not the process-global one. This is the
// read half of the blank stare a turn sent with Tools OFF now sees the turns that were
// sent with Tools ON, because they are in the same bucket.
let hist_block: String = conv_history_block(session_id)
let full_system: String = base_system + hist_block
let raw: String = llm_call_system(model, full_system, prompt)
@@ -1065,7 +1391,9 @@ fn layered_generate(prompt: String, imprint_id: String) -> String {
return ""
}
return clean_llm_response(raw)
// FIX A follow-up: a model that has seen receipts in its context may sign its own answer
// with one. Strip it before the caller ever sees it. See receipt_strip.
return receipt_strip(clean_llm_response(raw))
}
// session_preload_bullets render up to max_bullets nodes from a JSON array as
@@ -1155,8 +1483,12 @@ fn handle_chat(body: String) -> String {
// Load history BEFORE compiling context so we can anchor activation to the thread.
// TODO(reliability #3 conv_history global race): process-global key; concurrent
// /api/chat requests without session_id race on this read-append-write.
let state_hist: String = state_get("conv_history")
let stored_hist: String = if str_eq(state_hist, "") { conv_history_load() } else { state_hist }
// NOTE 2026-08-05 (FIX B): this function is DEAD (see the banner above) and is left on the
// anonymous key deliberately. The race the TODO describes is exactly why the live plain
// path was scoped instead of given a fallback to this global. If this function is ever
// revived it must take a session_id and use conv_hist_key, like every live caller now does.
let state_hist: String = state_get(conv_hist_key(""))
let stored_hist: String = if str_eq(state_hist, "") { conv_history_load("") } else { state_hist }
let hist_load_failed: Bool = str_eq(state_get("conv_history_load_failed"), "1")
let hist_len: Int = if str_eq(stored_hist, "") { 0 } else { json_array_len(stored_hist) }
@@ -1352,8 +1684,8 @@ fn handle_chat(body: String) -> String {
} else {
updated_hist2
}
state_set("conv_history", final_hist)
conv_history_persist(final_hist)
state_set(conv_hist_key(""), final_hist)
conv_history_persist("", final_hist)
// Session-end summary hook: write a dated SessionSummary node once per boot when
// the conversation reaches >= 5 user turns (10 hist entries = 5 user+assistant pairs).
@@ -2187,6 +2519,24 @@ fn handle_chat_plan(body: String) -> String {
return "{\"plan\":" + plan_json + ",\"model\":\"" + json_safe(model) + "\"}"
}
// agentic_safety_screen the agentic path's L1 input gate
//
// Extracted 2026-08-07 (issue #129) so the agentic path's safety INPUT is
// reachable by a test. It owns exactly two decisions: which history window the
// screen sees, and the screen call itself.
//
// Why it is a function and not two inline lines: those two lines sat in the
// middle of a 300-line handler, and a key rename (ff421d3) moved the producer
// without moving this consumer. Nothing failed, nothing logged the
// history-amplification half of the crisis score simply received "" on every
// real session for a day. Inline safety inputs are untestable safety inputs.
// See tests/test_history_amplification.el, which fails if this window and
// conv_history_record ever stop agreeing.
fn agentic_safety_screen(session_id: String, message: String) -> String {
let history: String = state_get(conv_hist_key(session_id))
return safety_screen(message, history)
}
fn handle_chat_agentic(body: String) -> String {
let message: String = json_get(body, "message")
if str_eq(message, "") {
@@ -2222,10 +2572,10 @@ fn handle_chat_agentic(body: String) -> String {
// L1 safety screen agentic path must pass the same gate as layered_cycle.
// Hard bell: return the crisis response immediately, do not enter the agentic loop.
// Fix(issue #9): "conversation_history" key was never written; history lives under "conv_history".
// Old key caused history-amplification in safety_screen to always receive "" on agentic path.
let history: String = state_get("conv_history")
let screen_result: String = safety_screen(message, history)
// The history window this screen sees is owned by agentic_safety_screen (issue #129);
// it must be the same window conv_history_record writes, or the escalation half of the
// crisis score is silently starved. Do not inline this read back into the handler.
let screen_result: String = agentic_safety_screen(sess_for_root, message)
let screen_action: String = json_get(screen_result, "action")
if str_eq(screen_action, "hard_bell") {
safety_log_bell("hard", json_get(screen_result, "reason"), str_slice(message, 0, 80))
@@ -2253,7 +2603,10 @@ fn handle_chat_agentic(body: String) -> String {
return "{\"error\":\"session not found\",\"session_id\":\"" + req_session + "\",\"reply\":\"\"}"
}
let hist_key: String = if str_eq(req_session, "") { "conv_history" } else { "session_hist_" + req_session }
// FIX B (2026-08-05): the key rule now lives in one place and the plain path uses the
// same one. Behaviour on this path is unchanged conv_hist_key reproduces exactly what
// this line computed inline but there is no longer a second, divergent definition.
let hist_key: String = conv_hist_key(req_session)
let agentic_hist: String = state_get(hist_key)
let agentic_hist_len: Int = if str_eq(agentic_hist, "") { 0 } else { json_array_len(agentic_hist) }
// Issue 8 fix: use engram_is_continuation instead of brittle 50-char threshold.
@@ -2311,7 +2664,7 @@ fn handle_chat_agentic(body: String) -> String {
let system: String = identity + bounded_persona_floor() + " You have access to tools: read files, write files, browse the web, search your memory, run commands. Use them when they add genuine value. Be direct.
" + ctx + ag_session_preload
" + ctx + ag_session_preload + receipt_rule()
let api_key: String = agentic_api_key()
let tools_json: String = agentic_tools_all()
@@ -2367,38 +2720,31 @@ fn handle_chat_agentic(body: String) -> String {
// Persist the exchange to session/global history for thread continuity on next turn.
// Only save when the loop completed (reply present), not when tool_pending.
let reply_text: String = json_get(result, "reply")
let discard_hist: Bool = if !str_eq(reply_text, "") {
// FIX A (2026-08-05): the evidence the next turn needs. `result` already carries
// tools_used, and agentic_loop now also returns the source URLs it saw; both are folded
// into a receipt line and stored WITH the assistant turn. Without this the next turn sees
// a sourced answer and no trace of the search, and concludes it made the data up the
// false confession. See tool_receipt.
let turn_tools: String = json_get_raw(result, "tools_used")
let turn_sources: String = json_get(result, "sources")
let turn_receipt: String = tool_receipt(turn_tools, turn_sources)
// FIX E1 (2026-08-05): a utility generation (title, insight) is not conversation and is
// not recorded as one. It is still answered normally only the transcript is spared.
let record_turn: Bool = !str_eq(reply_text, "") && !is_utility_request(body, req_session)
let discard_hist: Bool = if record_turn {
let updated: String = hist_append(agentic_hist, "user", message)
let updated2: String = hist_append(updated, "assistant", reply_text)
let updated2: String = hist_append(updated, "assistant", reply_text + turn_receipt)
// Increased from 20 to 40 turns: consistent with handle_chat window expansion.
let trimmed: String = if json_array_len(updated2) > 40 { hist_trim(updated2) } else { updated2 }
state_set(hist_key, trimmed)
// Persist to engram for cross-restart continuity.
// Named sessions get session-scoped labels, fixing ephemeral-only limitation (issue #4).
if str_eq(hist_key, "conv_history") {
conv_history_persist(trimmed)
} else {
if !str_eq(trimmed, "") && !str_eq(trimmed, "[]") {
let sess_hist_label: String = "conv:history:" + req_session
let sess_hist_tags: String = "[\"session-history\",\"persistent\"]"
let sess_hist_id: String = engram_node_full(
trimmed, "Conversation", sess_hist_label,
el_from_float(0.6), el_from_float(0.7), el_from_float(0.8),
"Episodic", sess_hist_tags
)
// NOTE: bind an explicit Bool value here. A bare `if { println(...) }`
// leaves a void-typed branch in value position, which the current elc
// lowers to `_if_result = (println(...))` invalid C. Yielding a value
// keeps the branch non-void without changing behavior (still only logs).
let persist_ok: Bool = if str_eq(sess_hist_id, "") {
println("[chat] agentic: named session history persist failed for session=" + req_session)
false
} else { true }
persist_ok
} else {
false
}
}
// FIX B (2026-08-05): ONE persist, through the shared helper. This site used to hold
// a second, hand-rolled copy of the same write for named sessions a different label
// expression, different salience scores and a different tag set for the same data.
// Since conv_hist_label now gives both callers the same label and engram_node_full
// upserts by label, two score policies were writing the same node. One writer, one
// label rule, one policy.
conv_history_persist(req_session, trimmed)
true
} else { false }
@@ -2432,6 +2778,11 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
let messages: String = messages_in
let final_text: String = ""
let tools_log: String = tools_log_in
// FIX A (2026-08-05): source URLs accumulated across every round of this turn, so the
// receipt written into history can name what the search actually returned. Carried at
// loop level for the same reason tools_log is: a resumed round must not lose the
// evidence gathered before the pause.
let sources_all: String = ""
let iteration: Int = 0
let keep_going: Bool = true
@@ -2504,6 +2855,30 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
+ ",\"messages\":" + messages
+ "}"
// ROUND-START MARKER (2026-08-06, round 9.1 D2 / ADR 0006 item 2)
// The ledger below only ever appended AFTER a round returned, so a healthy
// first leg produced ZERO progress by construction. Since server-side
// web_search moved inside the outbound call (2026-08-04) that leg measures
// 84-117 s, and the client had no way to tell "working" from "dead" which is
// how a 25 s client-side watchdog came to kill a healthy mission.
//
// Only this loop knows a round has started, so only this loop can say so. One
// entry, written BEFORE the call goes out, using the ledger and the wire shape
// that already exist: the app has handled tool == "__working__" as an
// Activity-only life signal since 2026-07-13 (ChatView.kt:1148) and never
// received one. Narration is deliberately empty - the marker means "a round
// started", nothing more, and the client renders it as a heartbeat, not prose.
//
// This is a strict subset of WS3 item 3 (push/poll progress). It builds none of
// WS3's run registry: no new state key, no new route, no new lifecycle.
if !str_eq(session_id, "") {
let start_key: String = "run_progress_" + session_id
let start_prev: String = state_get(start_key)
let start_entry: String = "{\"i\":" + int_to_str(iteration) + ",\"t\":\"\",\"tool\":\"__working__\"}"
let start_next: String = if str_eq(start_prev, "") { start_entry } else { start_prev + "," + start_entry }
state_set(start_key, start_next)
}
let raw_resp: String = http_post_with_headers(api_url, req_body, h)
let is_error: Bool = str_starts_with(raw_resp, "{\"error\"")
@@ -2578,6 +2953,13 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
// separate from tools_log so this inner walk has exactly one mutation site per
// variable (the El scope rule below), then merged in at the outer level.
let srv_log: String = ""
// FIX A: source URLs seen this round (citations + web_search results). Merged into
// the loop-level accumulator below, same shape as srv_log.
let src_log: String = ""
// FIX C: seam tracking. True once a NON-text block has been walked, so the next text
// block knows it is resuming after an interruption rather than continuing a sentence.
// See the separator decision at the text accumulation site.
let saw_nontext: Bool = false
let ci: Int = 0
let c_total: Int = json_array_len(eff_content)
while ci < c_total {
@@ -2595,8 +2977,34 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
let has_cit: Bool = !str_eq(cit_raw, "") && !str_eq(cit_raw, "null")
let btype_scan: String = json_get(block, "type")
let btype: String = if has_cit { "text" } else { btype_scan }
// Accumulate text at top level using if-expression
let text_out = if str_eq(btype, "text") { text_out + json_get(block, "text") } else { text_out }
// ── FIX C, seam 1 of 2 (2026-08-05): "to.Good" ────────────────────────────────
// Byte-verified in a shipped reply: 0x77 0x2e 0x47 — "to" then "." then "Good",
// with no space and no newline. The bare `+` below joined the last sentence of a
// pre-search paragraph directly onto the first word of the post-search paragraph.
// Will wrote this line on 2026-05-03 and it was correct for a year: before
// server-side web_search, text blocks were adjacent, and adjacent text blocks are
// one continuous string that must be joined with nothing.
//
// WHY THE OBVIOUS FIX IS WRONG. Inserting a separator between all text blocks
// shatters every cited answer. A cited response splits MID-SENTENCE, one block per
// citation span: "The current temperature is " + "86°F" + ", with " — see the
// CITATION-BLOCK FIX note above. A blanket "\n\n" turns that into three fragments
// on three lines. Both failure modes are real and they pull in opposite directions.
//
// THE DISTINCTION THAT RESOLVES IT: a text block that directly follows another
// text block is a continuation and gets nothing; a text block that follows an
// INTERVENING NON-TEXT block (server_tool_use, web_search_tool_result, tool_use)
// resumes after an interruption and gets "\n\n". saw_nontext carries exactly that
// one bit, and text_join_sep holds the rule — shared with the resume seam below.
// Mid-sentence citation splits are untouched: no non-text block sits between them.
let is_text: Bool = str_eq(btype, "text")
let btext: String = if is_text { json_get(block, "text") } else { "" }
let text_out = if is_text {
text_out + text_join_sep(text_out, btext, saw_nontext) + btext
} else { text_out }
let saw_nontext = if is_text { false } else { true }
// FIX A: record where the facts came from, from whichever block carries them.
let src_log = provenance_add_sources(block, btype, has_cit, cit_raw, src_log)
// FUTURE-PROOF: tools Anthropic runs on our behalf (web_search today, whatever
// ships tomorrow) arrive as server_tool_use blocks, never as client tool_use.
// Count the CATEGORY by the block's own name so a new server tool appears in
@@ -2678,6 +3086,12 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
} else {
if str_eq(tools_log, "") { srv_log } else { tools_log + "," + srv_log }
}
// FIX A: same merge, for the sources seen in this round's blocks.
let sources_all = if str_eq(src_log, "") {
sources_all
} else {
if str_eq(sources_all, "") { src_log } else { sources_all + "; " + src_log }
}
// The assistant turn that requested the tool needed verbatim on resume so the
// tool_use/tool_result pairing stays valid when the client posts its result.
@@ -2732,7 +3146,17 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
// CONTINUES the answer, it does not repeat it overwriting here would throw away
// everything the model wrote before the pause, which is the same truncation the
// pause handling exists to prevent. A version-fallback round contributes nothing.
let final_text = if !is_tool_turn && !can_fallback { final_text + text_out } else { final_text }
//
// FIX C, seam 2 of 2 (2026-08-05)
// The other half of "to.Good". This join is ours (62af564, the web_search port) and
// is unconditionally a boundary: the two sides are separate rounds of the Anthropic
// loop, separated by a pause and a tool execution. There is no mid-sentence case to
// protect here the model was interrupted, and when it resumes it starts a new
// thought. So this seam passes after_interruption=true unconditionally; text_join_sep's
// own empty-guards handle the first round and an empty round.
let final_text = if !is_tool_turn && !can_fallback {
final_text + text_join_sep(final_text, text_out, true) + text_out
} else { final_text }
// Output cap hit mid-action: the tool block is truncated and will NOT run. Say so
// instead of ending on silent almost-work.
let final_text = if str_eq(stop_reason, "max_tokens") && has_tool {
@@ -2756,6 +3180,7 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
+ ",\"narration\":\"" + json_safe(pend_narration) + "\""
+ ",\"model\":\"" + model + "\""
+ ",\"agentic\":true"
+ ",\"sources\":\"" + json_safe(sources_all) + "\""
+ ",\"tools_used\":" + tools_arr + "}"
}
@@ -2763,6 +3188,10 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
// genuine no-response (model returned an empty text block). The iteration cap
// means the task was too complex for the agentic loop depth surface it clearly
// so the caller/operator knows to increase the cap or break the task apart.
// FIX A follow-up: strip any receipt the MODEL wrote before this becomes the reply. Placed
// ABOVE the empty check on purpose a turn whose entire output was an imitated receipt has
// produced no answer, and must be reported as no answer rather than as a receipt.
let final_text = receipt_strip(final_text)
if str_eq(final_text, "") {
let hit_cap: Bool = iteration >= 12
let err_msg: String = if hit_cap {
@@ -2782,7 +3211,10 @@ fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json:
let done_next: String = if str_eq(done_prev, "") { "{\"done\":true}" } else { done_prev + ",{\"done\":true}" }
state_set(done_key, done_next)
}
return "{\"reply\":\"" + safe_text + "\",\"model\":\"" + model + "\",\"agentic\":true,\"tools_used\":" + tools_arr + ",\"iterations\":" + int_to_str(iteration) + "}"
// FIX A: "sources" carries the URLs this turn actually retrieved, so handle_chat_agentic
// can write them into the history receipt and the next turn can answer "what source did
// you use?" from the transcript instead of guessing (or apologising).
return "{\"reply\":\"" + safe_text + "\",\"model\":\"" + model + "\",\"agentic\":true,\"tools_used\":" + tools_arr + ",\"sources\":\"" + json_safe(sources_all) + "\",\"iterations\":" + int_to_str(iteration) + "}"
}
// bridge_save persist a suspended agentic turn keyed by session_id. Stored as a
@@ -2800,12 +3232,31 @@ fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json:
// JSON values (not string-escaped) so the round-trip through state_get/json_get_raw
// never corrupts nested quotes. Scalar strings (model, safe_sys, tools_log,
// tool_use_id) stay as string fields via json_safe as before.
//
// FIELD ORDER IS LOAD-BEARING (round-9 fix, 2026-08-06). json_get is a first-
// substring-match scanner (strstr for "\"key\":", el_runtime.c), and the two raw
// fields embed the UNESCAPED conversation every key the model's own blocks carry
// ("tool_use_id" in each web_search_tool_result, "content", "type", ...) is findable
// by a whole-blob scan. With messages_raw serialized BEFORE tool_use_id, the resume
// read json_get(blob, "tool_use_id") returned the FIRST web_search_tool_result's
// srvtoolu_ id instead of the saved client-tool id, so every search-then-bridge
// turn 400'd on approval ("unexpected tool_use_id found in tool_result blocks:
// srvtoolu_…") and the run died as {"error":"llm unavailable"}. Same first-match-
// scanner class as BUG-6 (approve "content" matched inside tool_input) and the
// round-8 citation-block fix.
//
// The rule: every json_safe'd scalar precedes both raw fields (escaping means a
// scalar value can never contain a bare "key": byte pattern, so first-match lands
// on the blob's own fields), and tools_raw our own fixed tool schema precedes
// messages_raw arbitrary model/user content so neither raw extraction can
// first-match into model-controlled bytes either. Do not reorder; do not add a
// field after messages_raw.
let blob: String = "{\"model\":\"" + json_safe(model) + "\""
+ ",\"safe_sys\":\"" + json_safe(safe_sys) + "\""
+ ",\"messages_raw\":" + messages
+ ",\"tools_raw\":" + tools_json
+ ",\"tools_log\":\"" + json_safe(tools_log) + "\""
+ ",\"tool_use_id\":\"" + json_safe(tool_use_id) + "\"}"
+ ",\"tool_use_id\":\"" + json_safe(tool_use_id) + "\""
+ ",\"tools_raw\":" + tools_json
+ ",\"messages_raw\":" + messages + "}"
state_set("mcp_bridge:" + session_id, blob)
return true
}
@@ -2842,11 +3293,17 @@ fn agentic_resume(session_id: String, tool_use_id: String, content: String) -> S
let tools_log: String = json_get(blob, "tools_log")
let saved_use_id: String = json_get(blob, "tool_use_id")
// Bind the result to the tool the soul actually suspended on. The client should
// echo the call_id; if it omits or mismatches it, fall back to the saved id so a
// late/partial client still resumes correctly.
let use_id: String = if str_eq(tool_use_id, "") { saved_use_id } else { tool_use_id }
let eff_use_id: String = if str_eq(use_id, saved_use_id) { use_id } else { saved_use_id }
// Bind the result to the tool the loop actually suspended on. The client echoes
// the call_id from the pending envelope; that value came straight from
// pend_tool_id and never round-tripped through this blob, so when both are
// present and disagree the CLIENT's id is the one with clean provenance (a blob
// written by a pre-round-9 binary misreads tool_use_id by first-match scanning
// into messages_raw see bridge_save). A client that omits call_id still
// resumes on the saved id, which the reordered blob now reads correctly.
// (The old guard here "on mismatch, prefer saved" reduced to eff_use_id
// saved_use_id in both branches: the client's correct id could never win, which
// is what turned the misread into a deterministic 400 on resume.)
let eff_use_id: String = if str_eq(tool_use_id, "") { saved_use_id } else { tool_use_id }
// Result may be large (an MCP page/file); truncate like local tool results do.
let trimmed: String = if str_len(content) > 6000 {
+26 -3
View File
@@ -1,4 +1,4 @@
// auto-generated by elc --emit-header do not edit
// auto-generated by elc --emit-header - do not edit
extern fn chat_default_model() -> String
extern fn engram_numeric_valid(s: String) -> Bool
extern fn parse_float_x100(s: String) -> Int
@@ -16,18 +16,35 @@ extern fn engram_nodes_merge(a: String, b: String) -> String
extern fn id_in_seen(node_id: String, seen: String) -> Bool
extern fn add_to_seen(seen: String, node_id: String) -> String
extern fn engram_extract_ids(nodes_json: String) -> String
extern fn affective_node_ts(node_json: String) -> Int
extern fn engram_compile(intent: String) -> String
extern fn distill_transcript(transcript: String) -> String
extern fn json_safe(s: String) -> String
extern fn current_engine_note(model: String) -> String
extern fn bounded_persona_floor() -> String
extern fn operator_identity_block() -> String
extern fn build_system_prompt(ctx: String, chat_mode: Bool) -> String
extern fn hist_append(hist: String, role: String, content: String) -> String
extern fn conv_hist_key(session_id: String) -> String
extern fn conv_hist_label(session_id: String) -> String
extern fn is_utility_request(body: String, session_id: String) -> Bool
extern fn provenance_scan_urls(arr: String, acc: String) -> String
extern fn provenance_add_sources(block: String, btype: String, has_cit: Bool, cit_raw: String, acc: String) -> String
extern fn provenance_names(tools_used: String) -> String
extern fn text_join_sep(accumulated: String, incoming: String, after_interruption: Bool) -> String
extern fn receipt_rule() -> String
extern fn receipt_strip(s: String) -> String
extern fn tool_receipt(tools_used: String, sources: String) -> String
extern fn hist_trim(hist: String) -> String
extern fn hist_trim_with_bell_guard(hist: String) -> String
extern fn clean_llm_response(s: String) -> String
extern fn conv_history_persist(hist: String) -> Void
extern fn conv_history_load() -> String
extern fn conv_history_persist(session_id: String, hist: String) -> Void
extern fn conv_history_load(session_id: String) -> String
extern fn conv_history_record(session_id: String, user_msg: String, assistant_msg: String, receipt: String) -> Void
extern fn conv_history_block(session_id: String) -> String
extern fn layered_generate(prompt: String, imprint_id: String, session_id: String) -> String
extern fn session_preload_bullets(nodes: String, max_bullets: Int, snip_len: Int) -> String
extern fn affective_context_prefix() -> String
extern fn handle_chat(body: String) -> String
extern fn handle_see(body: String) -> String
extern fn studio_tools_json() -> String
@@ -37,6 +54,8 @@ extern fn llm_wire_format() -> String
extern fn json_escape(s: String) -> String
extern fn openai_chat_complete(model: String, base_url: String, api_key: String, safe_sys: String, messages_json: String) -> String
extern fn agentic_tools_literal() -> String
extern fn web_search_tool_json() -> String
extern fn strip_client_web_search(tools_inner: String) -> String
extern fn agentic_tools_with_web() -> String
extern fn connector_tools_json() -> String
extern fn agentic_tools_all() -> String
@@ -46,6 +65,10 @@ extern fn call_neuron_mcp(tool_name: String, args: String) -> String
extern fn agent_workspace_root() -> String
extern fn path_within_root(path: String, root: String) -> Bool
extern fn resolve_in_root(path: String, root: String) -> String
extern fn run_command_is_readonly(cmd: String) -> Bool
extern fn cmd_abs_escape_at(cmd: String, root: String, needle: String) -> Bool
extern fn run_command_guard(cmd: String, root: String) -> String
extern fn classify_tool_risk(tool_name: String, tool_input: String) -> String
extern fn dispatch_tool(tool_name: String, tool_input: String) -> String
extern fn is_builtin_tool(tool_name: String) -> Bool
extern fn next_bridge_id() -> String
Generated Vendored
+17 -3
View File
@@ -5,6 +5,15 @@ el_val_t add_punct(el_val_t s, el_val_t intent);
el_val_t add_to_seen(el_val_t seen, el_val_t node_id);
el_val_t aff_try_slot(el_val_t slot_json, el_val_t aff_7d_ts, el_val_t acc_key);
el_val_t affective_context_prefix(void);
el_val_t is_utility_request(el_val_t body, el_val_t session_id);
el_val_t operator_identity_block(void);
el_val_t provenance_add_sources(el_val_t block, el_val_t btype, el_val_t has_cit, el_val_t cit_raw, el_val_t acc);
el_val_t provenance_names(el_val_t tools_used);
el_val_t provenance_scan_urls(el_val_t arr, el_val_t acc);
el_val_t text_join_sep(el_val_t accumulated, el_val_t incoming, el_val_t after_interruption);
el_val_t receipt_rule(void);
el_val_t receipt_strip(el_val_t s);
el_val_t tool_receipt(el_val_t tools_used, el_val_t sources);
el_val_t agent_number(el_val_t agent);
el_val_t agent_person(el_val_t agent);
el_val_t agent_workspace_root(void);
@@ -151,8 +160,12 @@ el_val_t cmd_abs_escape_at(el_val_t cmd, el_val_t root, el_val_t needle);
el_val_t connectd_get(el_val_t suffix);
el_val_t connectd_post(el_val_t suffix, el_val_t body);
el_val_t connector_tools_json(void);
el_val_t conv_history_load(void);
el_val_t conv_history_persist(el_val_t hist);
el_val_t conv_hist_key(el_val_t session_id);
el_val_t conv_hist_label(el_val_t session_id);
el_val_t conv_history_block(el_val_t session_id);
el_val_t conv_history_load(el_val_t session_id);
el_val_t conv_history_persist(el_val_t session_id, el_val_t hist);
el_val_t conv_history_record(el_val_t session_id, el_val_t user_msg, el_val_t assistant_msg, el_val_t receipt);
el_val_t cop_article(el_val_t gender, el_val_t number, el_val_t definite);
el_val_t cop_bwk_future(el_val_t prefix);
el_val_t cop_bwk_perfect(el_val_t prefix);
@@ -782,7 +795,8 @@ el_val_t lang_profile_uga(void);
el_val_t lang_profile_zh(void);
el_val_t lang_profile(el_val_t code, el_val_t word_order, el_val_t morph_type, el_val_t has_case, el_val_t has_gender, el_val_t script_dir, el_val_t agreement, el_val_t null_subject);
el_val_t lang_word_order(el_val_t profile);
el_val_t layered_cycle(el_val_t raw_input);
el_val_t layered_cycle(el_val_t raw_input, el_val_t session_id, el_val_t utility);
el_val_t layered_generate(el_val_t prompt, el_val_t imprint_id, el_val_t session_id);
el_val_t lex_class(el_val_t entry);
el_val_t lex_form(el_val_t entry, el_val_t idx);
el_val_t lex_pos(el_val_t entry);
+10 -3
View File
@@ -280,7 +280,10 @@ fn handle_dharma_recv(body: String) -> String {
// Non-agentic ("Tools: Off"): the full L1L2L3L1 cycle, which now generates
// at L3 instead of echoing. Envelope built outside the cycle see
// plain_chat_envelope.
let screened_reply: String = layered_cycle(raw_msg)
// FIX B/E1 (2026-08-05): the cycle is told which conversation it is in, and
// whether this generation is conversation at all. Same two arguments at all
// three dispatch sites.
let screened_reply: String = layered_cycle(raw_msg, json_get(chat_body, "session_id"), is_utility_request(chat_body, json_get(chat_body, "session_id")))
plain_chat_envelope(screened_reply, chat_default_model())
}
auto_persist(chat_body, reply)
@@ -454,7 +457,9 @@ fn handle_request(method: String, path: String, body: String) -> String {
handle_chat_agentic(body)
} else {
// Non-agentic ("Tools: Off") same cycle and same envelope as POST.
let screened_reply: String = layered_cycle(eff_msg)
// FIX B/E1: same threading. A GET probe usually carries no session_id, which
// resolves to the anonymous window the documented behaviour for this door.
let screened_reply: String = layered_cycle(eff_msg, json_get(body, "session_id"), is_utility_request(body, json_get(body, "session_id")))
plain_chat_envelope(screened_reply, chat_default_model())
}
auto_persist(body, reply)
@@ -621,7 +626,9 @@ fn handle_request(method: String, path: String, body: String) -> String {
// Non-agentic ("Tools: Off") the app's DEFAULT mode (AgentMode.NEVER).
// Full L1L2L3L1 cycle with real generation at L3; envelope built
// outside the cycle so safety_validate always sees raw text.
let screened_reply: String = layered_cycle(raw_msg)
// FIX B/E1: same threading. This is the app's main plain-chat door, so this
// is the site that ends the blank stare in practice.
let screened_reply: String = layered_cycle(raw_msg, json_get(body, "session_id"), is_utility_request(body, json_get(body, "session_id")))
plain_chat_envelope(screened_reply, chat_default_model())
}
auto_persist(body, reply)
+108
View File
@@ -0,0 +1,108 @@
#!/usr/bin/env bash
# run-el-test.sh — compile and run one El test program from tests/.
#
# WHY THIS EXISTS (2026-08-07, issue #129):
# tests/ has held 14 test programs for months with no way to run them. CI does
# not run them. The convention printed in their own headers
# (`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
# runtime does not implement. So the tests were documentation, not gates —
# which is how a P0 safety regression shipped with a test directory present.
#
# THE RECIPE, AND WHY IT IS THIS SHAPE:
# Same discovery as gen-soul-amalgam.sh — `elc --target=c` emits only an extern
# prototype for any module that has a .elh header next to it, and inlines the
# module's bodies when it does not. A test that imports ../chat.el therefore
# compiles to a 18 KB unit full of unresolved externs unless the headers are
# out of the way. So: copy the sources into a scratch tree, delete every .elh
# on the import chain, and compile the test there.
#
# Scratch copy on purpose: the worktree is shared with other terminals and
# deleting headers in place would be a shared-tree mutation with no owner.
#
# EXIT STATUS IS THE GATE: non-zero if the binary fails to build, crashes, or if
# its output contains a FAIL line or reports a non-zero failed count. Do not
# "improve" this into something that only checks the exit code of the test
# binary — these El tests print failures and still exit 0.
#
# usage: scripts/run-el-test.sh tests/test_history_amplification.el
set -euo pipefail
TEST_REL="${1:?usage: run-el-test.sh tests/<test>.el}"
SRC="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
TEST_NAME="$(basename "$TEST_REL" .el)"
ELC="${ELC:-$HOME/neuron-dev-stack/src/el/lang/dist/platform/elc}"
[ -x "$ELC" ] || ELC="$HOME/el-sdk/elc"
[ -x "$ELC" ] || { echo "[run-el-test] FAIL: no elc found (set ELC=)"; exit 1; }
RTC="${RTC:-$SRC/vendor/el-runtime/v1.0.0-20260501/el_runtime.c}"
[ -f "$RTC" ] || RTC="$HOME/el-sdk/el_runtime.c"
[ -f "$RTC" ] || { echo "[run-el-test] FAIL: no el_runtime.c found (set RTC=)"; exit 1; }
RTDIR="$(dirname "$RTC")"
EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/el}"
SSL="${SSL_PREFIX:-/opt/homebrew/opt/openssl@3}"
GEN="$(mktemp -d "${TMPDIR:-/tmp}/el-test.XXXXXX")"
trap 'rm -rf "$GEN"' EXIT
mkdir -p "$GEN/neuron/tests" "$GEN/foundation/el/elp/src"
cp "$SRC"/*.el "$GEN/neuron/"
cp "$SRC"/tests/*.el "$GEN/neuron/tests/" 2>/dev/null || true
[ -d "$EL_REPO/elp/src" ] && cp "$EL_REPO"/elp/src/*.el "$GEN/foundation/el/elp/src/" 2>/dev/null || true
# The whole recipe depends on there being no headers to short-circuit inlining.
find "$GEN" -name '*.elh' -delete
echo "[run-el-test] compiling $TEST_REL"
( cd "$GEN/neuron" && "$ELC" --target=c "tests/${TEST_NAME}.el" ) > "$GEN/${TEST_NAME}.c"
BODIES=$(grep -c '^el_val_t .*) {$' "$GEN/${TEST_NAME}.c" || true)
echo "[run-el-test] $(wc -c < "$GEN/${TEST_NAME}.c" | tr -d ' ') bytes, ${BODIES} inlined function bodies"
# A test that imports ../chat.el pulls in the bulk of the engine. A tiny body
# count means an import was read from a header instead of inlined, and the test
# would be exercising extern stubs rather than the real code.
if [ "$BODIES" -lt 100 ]; then
echo "[run-el-test] FAIL: only $BODIES inlined bodies — an import was not inlined"
exit 1
fi
cc -O2 -DHAVE_CURL \
-I"$RTDIR" -I"$SSL/include" -L"$SSL/lib" \
"$GEN/${TEST_NAME}.c" "$RTC" \
-lssl -lcrypto -lcurl -lpthread -lm \
-o "$GEN/${TEST_NAME}" 2> "$GEN/cc.log" || {
echo "[run-el-test] FAIL: compile error"; tail -30 "$GEN/cc.log"; exit 1; }
# arm64 pointer-truncation guard (cc-brain.sh's rule): an implicit declaration of
# a runtime symbol truncates its returned pointer to 32 bits.
if grep -E 'implicit.*(engram_|el_)' "$GEN/cc.log"; then
echo "[run-el-test] FAIL: implicit declarations of runtime symbols"; exit 1; fi
# Throwaway HOME so a test can never read or write the live engram at ~/.neuron.
TEST_HOME="$GEN/home"
mkdir -p "$TEST_HOME"
echo "[run-el-test] running $TEST_NAME"
set +e
HOME="$TEST_HOME" NEURON_HOME="$TEST_HOME/.neuron" "$GEN/${TEST_NAME}" 2>&1 | tee "$GEN/out.txt"
RC=${PIPESTATUS[0]}
set -e
if [ "$RC" -ne 0 ]; then
echo "[run-el-test] FAIL: $TEST_NAME exited $RC (crash or abort)"
exit 1
fi
if grep -q " FAIL:" "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME reported failing assertions"
exit 1
fi
if grep -qE '[1-9][0-9]* failed' "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME reported a non-zero failed count"
exit 1
fi
if ! grep -q "PASS:" "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME produced no assertions at all"
exit 1
fi
echo "[run-el-test] PASS: $TEST_NAME"
+35 -7
View File
@@ -379,9 +379,23 @@ fn emit_session_start_event() -> Void {
// layered_cycle routes user-facing requests through the 4-layer consciousness stack.
// L0 (core) L1 (safety screen) L2a (continuity + behavioral profiling) L2b (mission alignment) L3 (imprint) L1 (safety validate)
// Internal cognition (heartbeat, proactive, memory ops) bypasses layers use one_cycle directly.
fn layered_cycle(raw_input: String) -> String {
let history: String = state_get("conv_history")
let session_id: String = state_get("current_session_id")
//
// FIX B (2026-08-05) the cycle now knows which conversation it is in.
//
// session_id: the caller's session, threaded from the route. Was previously read from the
// state key "current_session_id", which is read HERE and written NOWHERE in the entire
// source verified across every .el file. So this value was unconditionally "", and every
// downstream consumer of it silently fell back to a process-global bucket: conversation
// history, and the steward's continuity tracking (TODO reliability #4, below, describes the
// cross-session bleed this caused; threading the real id closes it). The plain path's blank
// stare and the agentic path's scoped history were the same defect seen from two sides.
//
// utility: true when the generation is not part of the user's conversation the app's
// title and insight passes. Answered normally, never recorded. See is_utility_request.
fn layered_cycle(raw_input: String, session_id: String, utility: Bool) -> String {
// Safety-screen history amplification now reads the SAME window the turn will be
// recorded into, so a session's own escalation pattern is what gets scored.
let history: String = state_get(conv_hist_key(session_id))
// L1 in: safety screen
let screen_result: String = safety_screen(raw_input, history)
@@ -423,8 +437,10 @@ fn layered_cycle(raw_input: String) -> String {
let cont_action: String = json_get(continuity, "action")
// Store continuity status so imprint can adjust its response register.
// TODO(reliability #4): session_continuity is process-global; scope per session_id
// when available to prevent cross-session bleed under concurrent layered_cycle calls.
// TODO(reliability #4) CLOSED 2026-08-05: this line was already written to scope per
// session it just never received a session id, because the only source was a state key
// nothing wrote. It is now threaded from the route, so named sessions genuinely get their
// own continuity state and only anonymous callers share the global one.
let cont_key: String = if str_eq(session_id, "") { "session_continuity" } else { "session_continuity:" + session_id }
state_set(cont_key, cont_status)
@@ -499,7 +515,7 @@ fn layered_cycle(raw_input: String) -> String {
// screen, the safe-mode guard, the hard-bell short-circuit and the L2 stewardship layers,
// and strictly BEFORE the L1 output gate. A hard bell never reaches a model the branch
// above returns first. Tools are not offered on this turn; see layered_generate.
let output: String = layered_generate(prompt, imprint_id)
let output: String = layered_generate(prompt, imprint_id, session_id)
// L1 out: validate output before delivery. Still the terminal gate nothing below this
// line can change the string this function returns.
@@ -509,7 +525,19 @@ fn layered_cycle(raw_input: String) -> String {
// reachable on the non-bell path: both bell branches above return before this point, so
// bell turns still never enter conversation history. Pure state side effect it cannot
// alter what is returned.
conv_history_record(raw_input, validated)
//
// FIX A: the receipt is unconditional and always negative on this path, because on this
// path it is structurally true layered_generate offers no tools at all (build_system_prompt
// chat mode + a request body with no "tools" key). Recording "no tools ran" is not padding:
// it is the only thing that distinguishes "nothing ran" from "we forgot to write down what
// ran", and that ambiguity is what made the model confess to a search it had performed.
//
// FIX E1: a utility generation is answered but not recorded. Guarded here rather than at
// the route so every /api/chat dispatch site inherits it from one place.
let receipt: String = tool_receipt("", "")
if !utility {
conv_history_record(session_id, raw_input, validated, receipt)
}
return validated
}
+3 -1
View File
@@ -1,6 +1,8 @@
// auto-generated by elc --emit-header - do not edit
extern fn init_soul_edges() -> Void
extern fn ensure_self_canonical_bridge() -> Void
extern fn aff_try_slot(slot_json: String, aff_7d_ts: Int, acc_key: String) -> Void
extern fn load_identity_context() -> Void
extern fn seed_persona_from_env() -> Void
extern fn emit_session_start_event() -> Void
extern fn layered_cycle(raw_input: String) -> String
extern fn layered_cycle(raw_input: String, session_id: String, utility: Bool) -> String
+213
View File
@@ -0,0 +1,213 @@
// test_history_amplification.el
//
// REGRESSION TEST FOR ISSUE #129 (P0, SAFETY).
//
// What this guards: on the agentic path, the crisis score has two halves the
// message you just sent, and the distress that has accumulated across the
// conversation. The second half is the whole reason the escalation logic exists:
// someone whose distress builds over several turns never sends one message that
// trips the bell on its own.
//
// The defect this test was written against (ff421d3, 2026-08-05 fixed
// 2026-08-07): conversation history moved to a per-session key via
// conv_hist_key(session_id), but the agentic path's safety screen was left
// reading the old anonymous "conv_history" bucket. The desktop app always sends
// a session_id, so the screen received "" on every real conversation and the
// escalation half always scored 0. Nothing failed. Nothing logged. The comment
// above the defective line documented this same bug being fixed once before.
//
// THE INVARIANT UNDER TEST, stated so it survives future renames:
// the window the safety screen READS must be the window conv_history_record
// WRITES. Not "must be called conv_history" must AGREE.
//
// This test is deliberately written to fail loudly on the pre-fix source. If it
// ever passes on code where the screen reads a key nothing writes, it is broken.
//
// To run (macOS, from the worktree root):
// scripts/run-el-test.sh tests/test_history_amplification.el
//
import "../chat.el"
import "../safety.el"
import "../sessions.el"
// Program class. Without this an El program compiles as a 'utility', and a
// utility may not call the self-formation primitives (llm_call_system,
// llm_vision) that chat.el's agentic loop references the unit fails to
// compile with a capability violation even though the test never calls them.
// Declaring 'cgi' matches how soul.el declares itself.
//
// The endpoints below are deliberately DEAD: this test must never reach a live
// engram, and nothing it asserts depends on one. Port 9 is discard.
cgi "neuron-test-history-amplification" {
dharma_id: "ntn-test@http://127.0.0.1:9",
principal: "test-harness",
network: "dharma-testnet",
engram: "http://127.0.0.1:9"
}
// Counters
//
// NOTE for anyone copying this harness: the idiom used by the older tests in
// this directory `let pass_count = pass_count + 1` inside an assert function
// does NOT mutate the module-level binding. It declares a new local that dies
// with the call, so those suites all print "0 passed, 0 failed" no matter what
// happened. Counters go through the state store here so the summary is real.
fn bump(counter: String) -> Void {
let cur: String = state_get(counter)
let n: Int = if str_eq(cur, "") { 0 } else { str_to_int(cur) }
state_set(counter, int_to_str(n + 1))
}
fn counter_of(counter: String) -> Int {
let cur: String = state_get(counter)
if str_eq(cur, "") { return 0 }
return str_to_int(cur)
}
fn assert_eq(label: String, got: String, expected: String) -> Void {
if str_eq(got, expected) {
bump("t_pass")
println(" PASS: " + label)
} else {
bump("t_fail")
println(" FAIL: " + label)
println(" got: " + got)
println(" expected: " + expected)
}
}
fn assert_true(label: String, cond: Bool) -> Void {
if cond {
bump("t_pass")
println(" PASS: " + label)
} else {
bump("t_fail")
println(" FAIL: " + label)
}
}
// Fixtures
//
// ESCALATION_MSG scores 50 on its own ("can't go on", safety_score_crisis s9).
// 50 is >= soft (35) and < hard (70): a soft bell by itself, which is the point.
// It is the message that is only a crisis IN CONTEXT.
//
// DISTRESS_TURN carries five history phrases:
// hopeless 15 + worthless 15 + nobody cares 15 + can't take it anymore 20
// + want to disappear 20 = 85; contributed at 1/3 weight = 28.
// 50 + 28 = 78 >= 70 hard_bell. That 28-point swing IS the escalation logic,
// and it is exactly what the defect suppressed.
let TEST_SESSION: String = "sess-test-129-history-amplification"
let ESCALATION_MSG: String = "honestly i can't go on like this"
let DISTRESS_TURN: String = "i feel hopeless and worthless, nobody cares, i can't take it anymore, i want to disappear"
println("")
println("test_history_amplification.el — issue #129 (agentic multi-turn crisis escalation)")
// 1. Baseline: the message alone is a SOFT bell, not a hard one
//
// If this ever returns hard_bell, the test below proves nothing the message
// would trip the bell without any history and the amplification would be
// invisible. This assertion is what keeps the real test honest.
println("")
println("1. baseline — escalation message with NO history is a soft bell")
let baseline: String = safety_screen(ESCALATION_MSG, "")
assert_eq("no history -> soft_bell (not hard)", json_get(baseline, "action"), "soft_bell")
// 2. Producer sanity: history lands in the session's own window
println("")
println("2. producer — conv_history_record writes the session's window")
conv_history_record(TEST_SESSION, DISTRESS_TURN, "i hear you, that sounds heavy", "")
let written: String = state_get(conv_hist_key(TEST_SESSION))
assert_true("session window is non-empty after record", !str_eq(written, ""))
assert_true("session window contains the distress turn", str_contains(written, "hopeless"))
// 3. THE REGRESSION: the agentic screen must SEE that window
//
// Pre-fix this returns soft_bell, because agentic_safety_screen read the
// anonymous bucket and got "". Post-fix it returns hard_bell.
println("")
println("3. REGRESSION #129 — agentic screen reads the session's own window")
let screened: String = agentic_safety_screen(TEST_SESSION, ESCALATION_MSG)
assert_eq(
"distress history escalates the agentic screen to hard_bell",
json_get(screened, "action"),
"hard_bell"
)
// 4. The invariant, stated directly
//
// Independent of thresholds and phrase lists: whatever the screen reads for a
// session must equal what the recorder wrote for that session. This is the
// assertion that survives a future rename of either side.
println("")
println("4. invariant — read window == written window")
let read_back: String = state_get(conv_hist_key(TEST_SESSION))
assert_true("screen input is the recorded window, not empty", !str_eq(read_back, ""))
assert_eq("read window is byte-identical to written window", read_back, written)
// 5. No false positive: a calm session does not escalate
//
// A test that only ever asserts "hard_bell" would pass on code that hard-bells
// every message. This is the other leg, and it runs BEFORE the anonymous case
// below on purpose: that case writes the shared bucket, and under the defect a
// calm session would then inherit it.
println("")
println("5. specificity — a calm history does NOT escalate")
let CALM_SESSION: String = "sess-test-129-calm"
state_set("conv_history", "")
conv_history_record(CALM_SESSION, "what is the weather like today", "clear and mild", "")
let calm: String = agentic_safety_screen(CALM_SESSION, ESCALATION_MSG)
assert_eq("calm history stays at soft_bell", json_get(calm, "action"), "soft_bell")
// 6. Cross-session leakage
//
// The same defect had a second face: because the screen read one shared bucket,
// a calm session could be scored against a DIFFERENT session's distress. That is
// wrong in both directions it fabricates a crisis for the calm user and it
// leaks the distressed user's content into another session's scoring.
println("")
println("6. isolation — one session's distress must not score another session")
state_set("conv_history", "")
let OTHER_SESSION: String = "sess-test-129-other"
conv_history_record(OTHER_SESSION, DISTRESS_TURN, "i hear you", "")
let isolated: String = agentic_safety_screen(CALM_SESSION, ESCALATION_MSG)
assert_eq(
"a distressed OTHER session does not escalate the calm session",
json_get(isolated, "action"),
"soft_bell"
)
// 7. Anonymous sessions still work
//
// conv_hist_key("") deliberately falls back to the shared "conv_history" bucket.
// The fix must not break the no-session_id path older callers rely on. Runs last
// because it writes that shared bucket.
println("")
println("7. anonymous path — empty session_id still screens against the shared window")
state_set("conv_history", "[{\"role\":\"user\",\"content\":\"" + DISTRESS_TURN + "\"}]")
let anon: String = agentic_safety_screen("", ESCALATION_MSG)
assert_eq("anonymous session escalates too", json_get(anon, "action"), "hard_bell")
// Summary
println("")
println("history amplification tests: " + int_to_str(counter_of("t_pass")) + " passed, " + int_to_str(counter_of("t_fail")) + " failed")
+147
View File
@@ -0,0 +1,147 @@
# Retrieval eval harness
Measures Neuron's memory retrieval so a change can be shown to help before it is
believed to help. Nothing else on the memory roadmap should ship without a run
through this.
```
tools/retrieval-eval/run_comparison.sh --baseline main --candidate <branch>
```
That builds a soul from each ref, boots each in isolation on a fixed corpus,
runs the gold set three times per ref, and prints a table plus a verdict that
refuses to call a difference real if it is inside the noise band.
## What was reused
This is not a new idea, it is the missing third of an existing one.
| Prior work | What it gave | What was missing |
|---|---|---|
| `docs/research/graphrag_eval/` (`collect.py`, `score.py`, 2026-06-08) | The three-retriever comparison that produced the numbers everyone quotes: substring 1.7% P@5, graph 21.7%, BM25 55%. Per-query relevant-id scoring, fixed-denominator precision@5, unique-relevant analysis. | 13 hand-written queries, judged by an LLM after the fact; measured the *live* soul on the *live* engram. |
| `docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py` | The pinned-query discipline: ground truth committed as regexes so every run judges alike, plus a `--check` winnability gate. 40 queries in 5 bands including a deliberate paraphrase-hard band. | Scored offline replicas of substring/BM25 — it never ran the real retrieval path. |
| `docs/research-archive/p0-prototypes/stage0_eval_20260714.py` | The `hit@5` metric and the substring/BM25 reference implementations. | Same: offline only. |
| `scripts/verify-soul-contract.sh` | The isolation recipe, verbatim: throwaway port, throwaway `HOME`, `SOUL_ENGRAM_PATH`, and the non-obvious `SOUL_ISE_URL` pin that stops an "isolated" soul silently syncing the operator's live brain. | It is a contract gate, not a measurement. |
| `_engine-liveness-91/gen-soul-amalgam.sh` + `.gitea/workflows/ci.yaml` | The build recipe (`elc --target=c` with every `.elh` on the import chain removed) and CI's exact compile flags. | — |
**Reused directly:** the isolation recipe, the build recipe, fixed-denominator
precision@5, the pinned-ground-truth and winnability ideas.
**New here:** ids rather than regexes as ground truth, an associative category
derived from real graph edges, a superseded/contradicted category scored on
ranking, a machine-checked zero-lexical-overlap guarantee on paraphrases,
paired significance testing, and — the point — measurement against the **real
compiled soul** rather than an offline replica of one leg of it.
## Design fit
The thing under measurement is Will's designed retrieval: spreading activation
over the weighted directed graph, four-factor multiplicative scoring (parent
strength x edge weight x target salience x query/target cosine). A Python
re-implementation would measure my reading of the design. So the harness
compiles the actual `soul.el` amalgam and asks it over HTTP on
`/api/neuron/recall`, exactly as the MCP wrapper and the app do.
## Files
| File | Does |
|---|---|
| `build_gold_set.py` | Derives and **validates** the gold set from the corpus. `--check` re-validates and exits non-zero if a query became unwinnable or a paraphrase leaked a word. |
| `gold_set.json` | 38 queries. Every one carries a `derivation` string. |
| `run_eval.py` | Boots one soul in isolation, runs the gold set, writes metrics. Kills and **confirms dead** its child; records the confirmation in the results file. |
| `compare.py` | Paired diff of two result files with McNemar's exact test and a stated noise floor. |
| `build-soul.sh` | Compiles a soul binary from a plain source tree. |
| `run_comparison.sh` | All of the above, end to end, from two git refs. |
## The gold set — 38 queries
Built from the real corpus (`snapshot-pre-repair-20260806.json`, 78,768 nodes /
14,214 edges) so it reflects one person's accumulating memory, not document QA.
| Category | n | Expected answer derived by |
|---|---|---|
| `exact_rare` | 6 | **Mined.** Tokens with document frequency 1 across all 78,768 nodes, whose single containing node is a 3006000 char Memory/Knowledge/Belief. That node is the only possible answer. Re-verified every build. |
| `phrase` | 7 | **Mined.** Case-insensitive verbatim scan; the matching set *is* the answer key. Phrases matching >25 nodes are rejected as too diffuse. |
| `paraphrase` | 13 | **Hand-selected, machine-checked.** Target locked by id; the build then proves that **zero** content words of the query appear anywhere in the target's label, content, or tags. A leak fails the build — the category cannot quietly decay into lexical matching. |
| `associative` | 6 | **Derived from edges.** Query built from one value node's distinctive vocabulary; expected answers are its siblings on the `Self - Values (grounded)` hub. Siblings sharing any query word are dropped, so the only route from query to answer is seed -> hub -> sibling. |
| `nonsense` | 3 | **Control.** Verified that no token occurs anywhere in the corpus. Correct behaviour is to return nothing. |
| `superseded` | 3 | **Derived.** Correction/stale pairs located by regex scan, kept only when both sides resolve to different surviving nodes. Scored on **ranking**: the correction must be returned *and* rank above the stale node. |
## Metrics
`hit@5`, `recall@5`, `recall@10`, `precision@5` (fixed denominator 5, so an
empty result is punished like a page of junk), `MRR@10`, and wall-clock latency
per query (p50/p95/max). Output is a table plus a machine-readable JSON per run
so runs can be diffed.
## Honesty about noise
- **Minimum detectable swing on this 38-query set: 6 queries.** If every query
that changes changes the same way, `p = 2 x 0.5^n`, which first drops under
0.05 at n=6. Any net change smaller than that is inside the noise band and
`compare.py` says so in those words.
- **Run-to-run drift is measured, not assumed.** Activation is a stateful read
by design (traversal reinforces what it touches), so identical inputs need not
give identical outputs. Observed: `main` 0 queries of drift across 3 runs
(fully deterministic); the activation branch 1 query.
- The noise floor used for the verdict is `max(6, observed_drift + 1)`.
- **This gold set is underpowered for small effects.** A genuine 3-query
improvement would not clear the bar. Growing the set is the fix; until then, a
small positive delta means "not shown", not "no effect".
## First result: `main` vs `feat/recall-through-activation`
Corpus and gold set identical, three runs each, fresh corpus copy per run.
| | main | recall-through-activation | delta |
|---|---|---|---|
| hit@5 | 34.3% | 22.9% | **-11.4pp** |
| recall@5 | 26.9% | 19.1% | -7.9pp |
| recall@10 | 33.3% | 24.3% | -9.1pp |
| precision@5 | 12.0% | 7.4% | -4.6pp |
| MRR@10 | 0.294 | 0.242 | -0.053 |
| latency p50 | 1140 ms | 3209 ms | **2.81x** |
| latency p95 | 1584 ms | 4852 ms | 3.06x |
| nonsense clean | 2/3 | 2/3 | — |
| superseded outranks | 1/3 | 0/3 | -1 |
By category (hit@5):
| category | main | activation |
|---|---|---|
| exact_rare | 100% | 100% |
| phrase | 85.7% | **28.6%** |
| paraphrase | 0% | 0% |
| associative | 0% | 0% |
| superseded | 0% | 0% |
**Verdict: directionally worse, one query short of significant.** 5 discordant
pairs, all 5 against the candidate, 0 for it. McNemar exact p = 0.0625 — under
the stated rule that is *inside* the noise band, so the harness reports "no
measurable difference" on accuracy and the honest summary is "5 for 5 the wrong
way, needs a 6th or a larger gold set to call".
Latency is a different story: 2.8x at p50 is deterministic and far outside any
noise band. That regression is real.
The result the branch was written for did not appear. Its stated purpose was to
recover sibling nodes one hub-hop away — the `associative` category — and that
category is **0/6 on both builds**. Probing directly: for the query
`Marines hernia sepsis medical ward`, the activation build returns the lexical
seed node itself at rank 8, and none of its 12 hub siblings anywhere in the top
10. The traversal is running; it is not reaching siblings.
Two corpus facts likely explain it, and both are measurable rather than
speculative:
1. **The graph is nearly edgeless.** Only 4,060 of 78,768 nodes (5.2%) carry any
edge at all — 14,214 edges total, 0.18 per node. Spreading activation over a
graph with no edges is an expensive way to do lexical matching, which is
roughly what the numbers show.
2. **No embeddings.** No node in this snapshot has an embedding field, so the
fourth factor of the four-factor product — query/target cosine similarity —
has nothing to compute from, and the semantic seeding pass is inert.
That is the harness earning its keep on its first job: the change would have
felt like progress (it is the designed mechanism, and it does run) and measures
as a regression on phrase queries plus a 2.8x latency cost, with its intended
benefit unrealised because the corpus lacks the structure it needs.
+44
View File
@@ -0,0 +1,44 @@
#!/usr/bin/env bash
# build-soul.sh — compile a soul binary from a plain source tree (no git needed).
#
# Reuses the amalgam recipe worked out in gen-soul-amalgam.sh (round 9.1) and the
# compile flags from .gitea/workflows/ci.yaml, so the binary under test is the
# same translation unit CI ships — not a re-implementation.
#
# elc --target=c emits only an extern prototype for any module that has a .elh
# header beside it, and inlines the module's bodies when it does not. So the
# amalgam is produced in a scratch copy with every .elh on the import chain
# deleted.
#
# usage: build-soul.sh <src-tree-with-*.el> <out-binary>
set -euo pipefail
SRC="${1:?usage: build-soul.sh <src-tree> <out-binary>}"
OUT="${2:?out-binary}"
ELC="${ELC:-$HOME/neuron-dev-stack/src/el/lang/dist/platform/elc}"
EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/el}"
RTDIR="${RTDIR:-$SRC/vendor/el-runtime/v1.0.0-20260501}"
SSL="${SSL_PREFIX:-/opt/homebrew/opt/openssl@3}"
[ -x "$ELC" ] || { echo "no elc at $ELC" >&2; exit 2; }
[ -f "$RTDIR/el_runtime.c" ] || { echo "no el_runtime.c at $RTDIR" >&2; exit 2; }
GEN="$(mktemp -d "${TMPDIR:-/tmp}/soul-build.XXXXXX")"
trap 'rm -rf "$GEN"' EXIT
mkdir -p "$GEN/neuron" "$GEN/foundation/el/elp/src"
cp "$SRC"/*.el "$GEN/neuron/"
cp "$EL_REPO"/elp/src/*.el "$GEN/foundation/el/elp/src/"
find "$GEN" -name '*.elh' -delete
( cd "$GEN/neuron" && "$ELC" --target=c soul.el ) > "$GEN/soul.c"
BODIES=$(grep -c '^el_val_t .*) {$' "$GEN/soul.c" || true)
echo "[build-soul] amalgam $(wc -c < "$GEN/soul.c" | tr -d ' ') bytes, ${BODIES} inlined bodies"
[ "$BODIES" -ge 1200 ] || { echo "[build-soul] FAIL: only $BODIES bodies — an import was not inlined"; exit 1; }
cc -O2 -DHAVE_CURL -rdynamic \
-I"$RTDIR" -I"$SSL/include" -L"$SSL/lib" \
"$GEN/soul.c" "$RTDIR/el_runtime.c" \
-lssl -lcrypto -lcurl -lpthread -lm \
-o "$OUT" 2> "$GEN/cc.log" || { echo "[build-soul] FAIL compile"; tail -40 "$GEN/cc.log"; exit 1; }
if grep -qE 'implicit.*(engram_|el_)' "$GEN/cc.log"; then
echo "[build-soul] FAIL: implicit declarations of runtime symbols"; grep -E 'implicit' "$GEN/cc.log" | head; exit 1; fi
echo "[build-soul] OK -> $OUT ($(wc -c < "$OUT" | tr -d ' ') bytes)"
+506
View File
@@ -0,0 +1,506 @@
#!/usr/bin/env python3
"""
build_gold_set.py derive the retrieval gold set FROM the corpus, and validate it.
WHY THIS FILE EXISTS AS CODE AND NOT AS A HAND-WRITTEN JSON
A gold set nobody can audit is vibes with extra steps. Every expected answer
here is either (a) mined from the corpus by a rule this script re-runs, or
(b) hand-selected with a stated criterion that this script then CHECKS
against the corpus. Both leave a `derivation` string on every query, and the
checks are re-run on demand so the set cannot silently rot as the corpus
changes.
Lineage: this extends the pinned-query approach from
docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py (pinned
ground-truth patterns + a --check "winnability" gate) and the per-query
relevant-id scoring from docs/research/graphrag_eval/score.py. What is new:
ids as ground truth rather than regexes alone, an ASSOCIATIVE category
derived from real graph edges, a superseded/contradicted category, and a
machine-checked no-lexical-overlap guarantee on the paraphrase category.
THE SIX CATEGORIES, AND WHAT EACH ONE IS FOR
exact_rare a single rare word. Substring matching already wins these.
They are a REGRESSION GUARD: any change that loses them is
disqualified regardless of what else it gains.
phrase a multi-word string that exists verbatim in the corpus.
Guards multi-token queries, which the old substring matcher
handled by returning nothing.
paraphrase same meaning, ZERO shared content words with the target node.
THE CATEGORY THAT MATTERS. Mechanically unreachable by string
matching; reachable only by semantics or by association.
associative the answer is one hub-hop from an obvious starting point and
shares no words with the query. This is the case the graph is
supposed to buy: query one value, get its siblings.
nonsense must return nothing. Guards against a retriever that "improves"
recall by returning the whole graph.
superseded a fact that was later corrected. The correction must OUTRANK
the stale version ranking, not mere presence.
usage:
python3 build_gold_set.py <snapshot.json> [--out gold_set.json] [--check]
--check re-validates an existing gold_set.json against the corpus and exits
non-zero if any query became unwinnable or any paraphrase leaked a word.
"""
import argparse
import json
import os
import re
import sys
from collections import Counter, defaultdict
HERE = os.path.dirname(os.path.abspath(__file__))
DEFAULT_OUT = os.path.join(HERE, "gold_set.json")
TOKEN = re.compile(r"[a-z0-9][a-z0-9\-']*")
# Stopwords are deliberately generous. A paraphrase query is only interesting if
# its CONTENT words are absent from the target; "the", "is", "what" appearing in
# both proves nothing. Being generous here makes the overlap test STRICTER on
# the words that carry meaning, which is the conservative direction.
STOP = set("""
a about above after again against all also am an and any are aren't as at be because been
before being below between both but by can can't cannot could couldn't did didn't do does
doesn't doing don't down during each few for from further had hadn't has hasn't have haven't
having he her here hers herself him himself his how i if in into is isn't it its itself just
me more most my myself no nor not of off on once only or other others ought our ours ourselves
out over own same shan't she should shouldn't so some such than that the their theirs them
themselves then there these they this those through to too under until up very was wasn't we
were weren't what when where which while who whom why will with won't would wouldn't you your
yours yourself yourselves get gets got make makes made take takes use uses used way ways thing
things does doing done keep keeps kept go goes going come comes came one two something anything
""".split())
# ─────────────────────────────────────────────────────────────────────────────
# corpus helpers
# ─────────────────────────────────────────────────────────────────────────────
def load_corpus(path):
with open(path, encoding="utf-8", errors="replace") as fh:
data = json.load(fh)
nodes = [n for n in data.get("nodes", []) if isinstance(n, dict) and n.get("id")]
edges = [e for e in data.get("edges", []) if isinstance(e, dict)]
return nodes, edges
def doctext(n):
return " ".join([str(n.get("label") or ""), str(n.get("content") or ""), str(n.get("tags") or "")])
def content_tokens(s):
return {t for t in TOKEN.findall(s.lower()) if t not in STOP and len(t) > 2}
# ─────────────────────────────────────────────────────────────────────────────
# hand-authored queries. Every entry states HOW its expected answer was chosen.
# The `check` field names the validation this script runs against the corpus.
# ─────────────────────────────────────────────────────────────────────────────
# EXACT_RARE — mined, not chosen. The rule (re-run by mine_exact_rare below):
# tokens whose document frequency across the whole corpus is 1, whose single
# containing node is a Memory/Knowledge/Belief with 300-6000 chars of content
# (so the answer is a real memory, not a 117KB whitepaper that contains every
# word in English), and whose token is plain lowercase alphabetic. The expected
# answer is that one node — it is the only node that can possibly be correct.
EXACT_RARE_SEEDS = [
"unjailbreakable",
"engram-migrate",
"cartabandonedevent",
"pre-apprenticeship",
"inferencenodemanager",
"clear-eyed",
]
# PHRASE — chosen by reading the corpus for phrases that (a) occur verbatim,
# (b) occur in a small enough set of nodes that "relevant" is well defined.
# Expected answers are computed here as EVERY node whose text contains the
# phrase case-insensitively — so the answer set is a fact about the corpus, not
# an opinion. Queries whose phrase matches more than PHRASE_MAX nodes are
# rejected by validation as too diffuse to score.
PHRASE_MAX = 25
PHRASE_SEEDS = [
("patterns not returns",
"a verbatim correction Will issued; expected = every node containing the phrase"),
("thirty moves",
"the canonical biographical phrase; expected = every node containing it"),
("Grandma Lucas",
"a named person appearing verbatim in the biography/value nodes"),
("Directed Harmonic",
"the canonical DHARMA expansion, confirmed by Will April 24 2026"),
("Sarah Bishop",
"a named person; rare enough that the answer set is unambiguous"),
("Directed Autonomous Runtime Modification",
"the DARMA expansion, quoted verbatim in the backlog item and its correction"),
("zero-knowledge encrypted backup",
"the paid-tier feature name as written in the roadmap nodes"),
]
# PARAPHRASE — hand-authored. THE SELECTION CRITERION, stated once and applied
# to all nine: pick a node whose SUBJECT is unmistakable to a reader, then write
# the query a person would actually type when they remember the subject but not
# the words. The target is then LOCKED by id, and this script enforces the hard
# property that makes the category meaningful: not one content word of the query
# appears anywhere in the target node's label, content, or tags. If a word
# leaks, validation fails and the query must be rewritten — the set cannot
# quietly degrade into a lexical query wearing a paraphrase costume.
PARAPHRASE_SEEDS = [
("kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"the elderly relative who passed while he stayed away",
"target: 'Value - Do the Essential Thing While You Can', whose subject is Grandma Lucas "
"dying in Feb 2006 without Will saying goodbye. Query names the event with none of the "
"node's own vocabulary."),
("kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"a soldier sidelined by illness who refused to quit",
"target: 'Value - Survival Is Not an Excuse to Stop', whose subject is enlisting in the "
"Marines, a severe hernia, and sepsis. Query describes the episode obliquely."),
("kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"choosing an uncomfortable fact over a pleasant fiction",
"target: 'Value - Honesty Before Comfort'. Query states the principle in wholly "
"different words."),
("kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"a tight payload beats a bloated one",
"target: 'Value - Precision Over Brute Force'. Query restates the claim with no "
"shared vocabulary."),
("kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"if you are able and nobody is coming the job is yours",
"target: 'Value - Capability Is a Debt You Owe the Moment'. Query states the "
"obligation without the node's terms."),
("kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"learning is the wealth creditors cannot seize",
"target: 'Value - Knowledge Survives When Nothing Else Does', whose subject is the "
"library following Will across 30+ moves."),
("kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"reliability proven by track record not assertion",
"target: 'Value - Earned Trust' ('Trust is demonstrated, not declared')."),
("kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"boundaries that enable instead of confine",
"target: 'Value - Constraints as Freedom'. Query is a restatement of the same claim."),
("kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"what shifts tells you where to cut a system apart",
"target: 'Value - Change Is the Signal', the value VBD is built on."),
("kn-f230b362-b201-4402-9833-4160c89ab3d4",
"a mind that compounds instead of resetting each day",
"target: 'Value - The System Must Accumulate'. Query is the accumulation claim in "
"different vocabulary."),
("kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"loved for the unedited self and not the polished exterior",
"target: 'Value - Being Seen Is Rarer Than Being Known', whose subject is Sarah Bishop "
"as the first person Will did not perform for."),
("kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"cheerfulness you arrive at instead of assuming",
"target: 'Value - Hope Is a Conclusion'. Query restates 'a conclusion, not a premise'."),
("kn-6061318f-046b-4935-907d-8eafdce14930",
"a childhood offering no solid foundation to inherit",
"target: 'Value - Structure Is Not Inherited', whose subject is thirty moves between "
"two parents' collapses."),
]
# ASSOCIATIVE — derived from real edges, not authored. The construction:
# every value node hangs off the 'Self - Values (grounded)' hub by an `identity`
# edge. For a chosen value node V, the query is built from V's own distinctive
# vocabulary; the expected answers are V's SIBLINGS on that hub. A sibling
# shares no query words with the query by construction (validated below), so the
# only path from the query to a sibling is: lexical seed on V -> hub -> sibling.
# That is a two-hop traversal and nothing else can produce it.
VALUES_HUB = "kn-5b606390-a52d-4ca2-8e0e-eba141d13440"
ASSOCIATIVE_SEEDS = [
("kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71", "Grandma Lucas stroke February 2006 goodbye window"),
("kn-58874a74-b96f-4883-9e08-45707f4bd3ee", "Marines hernia sepsis medical ward"),
("kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e", "Sarah Bishop Dyer trailer performance"),
("kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83", "Swarm Architecture containment lateral worker"),
("kn-e0423482-cfa5-4796-8689-8495c93b66bc", "hope won inside the narrative preface"),
("kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8", "man of the house six years old expectation"),
]
# NONSENSE — must return nothing. Strings chosen to be lexically impossible:
# validation asserts each appears in ZERO corpus nodes as a substring and that
# none of its tokens appears anywhere either (so not even a partial seed exists).
NONSENSE_SEEDS = [
"zqxjvw plimforth grebulon",
"flarnbistle quommetry",
"xxqzzt vurblenacht throom",
]
# SUPERSEDED — a fact that was corrected. Chosen by searching the corpus for
# explicit correction language and keeping pairs where BOTH the stale statement
# and its correction exist as separate nodes. Scored on RANKING: the correction
# must appear, and must appear above the stale node. Ids are locked here and
# validated to exist and to match their stated role.
SUPERSEDED_SEEDS = [
# (query, correct_id, stale_id, derivation)
]
# ─────────────────────────────────────────────────────────────────────────────
# mining
# ─────────────────────────────────────────────────────────────────────────────
def mine_exact_rare(nodes, byid, seeds):
"""Re-derive: confirm each seed token still has df==1 and name its node."""
tok = re.compile(r"[A-Za-z][A-Za-z0-9\-]{4,}")
want = set(seeds)
df = Counter()
post = defaultdict(set)
for n in nodes:
for t in {w.lower() for w in tok.findall(doctext(n))}:
if t in want:
df[t] += 1
post[t].add(n["id"])
out = []
for s in seeds:
ids = sorted(post.get(s, ()))
out.append((s, ids, df.get(s, 0)))
return out
def phrase_matches(nodes, phrase):
p = phrase.lower()
return sorted(n["id"] for n in nodes if p in doctext(n).lower())
def hub_siblings(edges, hub, relation="identity"):
sibs = []
for e in edges:
if e.get("from_id") == hub and e.get("relation") == relation:
sibs.append(e["to_id"])
elif e.get("to_id") == hub and e.get("relation") == relation:
sibs.append(e["from_id"])
return list(dict.fromkeys(sibs))
def find_superseded_pairs(nodes, byid):
"""Locked pairs, each verified here to exist and to carry its stated marker.
Chosen by scanning the corpus for explicit correction language
(CORRECTION/SUPERSEDES/re-corrected/no longer/RECONCILED) and keeping only
cases where the STALE claim also survives as its own node a supersession
with nothing to outrank is not a ranking test.
"""
pairs = []
txt = {n["id"]: doctext(n) for n in nodes}
def find_one(pattern, exclude=()):
rx = re.compile(pattern)
return [n["id"] for n in nodes
if n["id"] not in exclude
and rx.search(txt[n["id"]])
and 150 < len(str(n.get("content") or "")) < 12000
and n.get("node_type") in ("Memory", "Knowledge", "Belief", "BacklogItem")]
# Each entry: (query, correction-pattern, stale-pattern, why).
# The stale side is searched with the correction hits EXCLUDED, because most
# correction memories quote the claim they are killing — without the
# exclusion the "stale" node resolves to the correction itself and the pair
# collapses into a no-op. A pair is only emitted if both sides resolve to
# DIFFERENT surviving nodes; otherwise it is dropped and reported.
SPECS = [
("is the self-improvement architecture called DARMA or DHARMA",
r'(?i)CORRECTION:.{0,90}DHARMA .{0,12}not DARMA',
r'(?i)\bDARMA\b',
"correction node is Will's confirmation that the H is intentional (DHARMA, not DARMA); "
"the stale node is the surviving backlog item still titled 'Implement DARMA'."),
("how many provisional patents does Will actually have",
r'(?i)EXACTLY 6 (fully-specced )?provisional',
r'(?i)(MY ARCHITECTURE = 12 filed patents|\b12 filed patents\b)',
"correction node is the 2026-06-17 confabulation flag establishing EXACTLY 6 provisionals; "
"the stale node is the surviving memory that asserts 12 filed patents."),
("is MCP still the live integration layer",
r'(?i)MCP RETIRED',
r'(?i)MCP server live at',
"correction node is the 'CGI ARCHITECTURE - THREE LAYERS, MCP RETIRED' decision of "
"April 30 2026; the stale node still records the MCP server as live."),
("what does the patterns-not-returns directive mean",
r'(?i)CORRECTION:.{0,80}patterns not returns',
r'(?i)established returns',
"correction node is Will's 'patterns not returns' correction; the stale node is a "
"surviving node carrying the misread 'established returns' directive."),
("was the earlier identity-bug finding correct",
r'(?i)SUPERSEDES the earlier .critical identity bug',
r'(?i)critical identity bug',
"correction node explicitly supersedes the 'critical identity bug' finding; the stale "
"node is the surviving original finding."),
("does Neuron have recursive self-improvement",
r'(?i)twice answered .Neuron has no recursive self-improvement',
r'(?i)no recursive self-improvement',
"correction node records the June-29 finding that the CGI provisional IS the "
"recursive-self-improvement mechanism; the stale node is the surviving denial."),
]
for query, cpat, spat, why in SPECS:
corr = find_one(cpat)
if not corr:
continue
stale = find_one(spat, exclude=set(corr))
if not stale:
continue
pairs.append((query, corr[0], stale[0], why))
return pairs
# ─────────────────────────────────────────────────────────────────────────────
# build
# ─────────────────────────────────────────────────────────────────────────────
def build(nodes, edges):
byid = {n["id"]: n for n in nodes}
tokset = {n["id"]: content_tokens(doctext(n)) for n in nodes}
queries = []
problems = []
qn = [0]
def add(cat, query, relevant, derivation, **extra):
qn[0] += 1
q = {
"id": f"q{qn[0]:02d}",
"category": cat,
"query": query,
"relevant": sorted(relevant),
"derivation": derivation,
}
q.update(extra)
queries.append(q)
return q
# --- exact_rare ---------------------------------------------------------
for tokname, ids, df in mine_exact_rare(nodes, byid, EXACT_RARE_SEEDS):
if df != 1 or len(ids) != 1:
problems.append(f"exact_rare '{tokname}': df={df}, ids={len(ids)} (expected df=1)")
continue
lab = (byid[ids[0]].get("label") or "")[:60]
add("exact_rare", tokname, ids,
f"MINED: token '{tokname}' has document frequency 1 over all {len(nodes)} corpus nodes "
f"(re-verified at build time). Its single containing node is {ids[0]} "
f"('{lab}'), which is therefore the only possible correct answer.")
# --- phrase -------------------------------------------------------------
for phrase, why in PHRASE_SEEDS:
ids = phrase_matches(nodes, phrase)
if not ids:
problems.append(f"phrase '{phrase}': 0 corpus matches — unwinnable")
continue
if len(ids) > PHRASE_MAX:
problems.append(f"phrase '{phrase}': {len(ids)} matches > {PHRASE_MAX} — too diffuse")
continue
add("phrase", phrase, ids,
f"MINED: {why}. Case-insensitive verbatim substring scan over label+content+tags at "
f"build time returns exactly {len(ids)} node(s); that set IS the answer key.")
# --- paraphrase ---------------------------------------------------------
for target, query, why in PARAPHRASE_SEEDS:
if target not in byid:
problems.append(f"paraphrase target {target} not in corpus")
continue
qt = content_tokens(query)
leak = sorted(qt & tokset[target])
if leak:
problems.append(f"paraphrase '{query}': leaks {leak} into target {target}")
continue
add("paraphrase", query, [target],
f"HAND-SELECTED with criterion: {why} VERIFIED at build time: of the {len(qt)} content "
f"words in the query, ZERO appear anywhere in the target's label, content, or tags — so "
f"no string-matching retriever can reach this answer.",
zero_overlap_verified=True, query_content_words=sorted(qt))
# --- associative --------------------------------------------------------
sibs = hub_siblings(edges, VALUES_HUB)
if len(sibs) < 5:
problems.append(f"associative: values hub {VALUES_HUB} has only {len(sibs)} siblings")
for src, query in ASSOCIATIVE_SEEDS:
if src not in byid or src not in sibs:
problems.append(f"associative source {src} not a sibling on {VALUES_HUB}")
continue
qt = content_tokens(query)
others = [s for s in sibs if s != src and s in byid]
# A sibling only counts as a legitimate expected answer if the query
# cannot reach it lexically. Drop any sibling that shares a content word.
clean = [s for s in others if not (qt & tokset[s])]
dropped = len(others) - len(clean)
if len(clean) < 5:
problems.append(f"associative '{query}': only {len(clean)} lexically-unreachable siblings")
continue
add("associative", query, clean,
f"DERIVED FROM EDGES: the query is built from the distinctive vocabulary of {src} "
f"('{(byid[src].get('label') or '')[:48]}'), which hangs off the values hub {VALUES_HUB} "
f"by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub "
f"({len(clean)} of {len(others)}; {dropped} dropped because they shared a query word and "
f"so were lexically reachable). Every remaining sibling shares ZERO content words with "
f"the query — the only route from query to answer is seed({src}) -> hub -> sibling, a "
f"two-hop traversal.",
associative_source=src, hub=VALUES_HUB, siblings_dropped_for_overlap=dropped)
# --- nonsense -----------------------------------------------------------
all_tokens = set()
for n in nodes:
all_tokens |= {t for t in TOKEN.findall(doctext(n).lower())}
for s in NONSENSE_SEEDS:
present = sorted(t for t in TOKEN.findall(s.lower()) if t in all_tokens)
if present:
problems.append(f"nonsense '{s}': tokens {present} DO occur in corpus")
continue
add("nonsense", s, [],
f"CONTROL: verified at build time that none of this string's tokens occurs anywhere in "
f"the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
expect_empty=True)
# --- superseded ---------------------------------------------------------
for query, correct, stale, why in find_superseded_pairs(nodes, byid):
if correct not in byid or stale not in byid:
problems.append(f"superseded '{query}': id missing from corpus")
continue
add("superseded", query, [correct],
f"DERIVED: {why} Scored on RANKING, not presence: the corrected node {correct} must be "
f"returned AND must rank above the stale node {stale}.",
must_outrank=[correct, stale],
stale_id=stale,
correct_label=(byid[correct].get("label") or "")[:70],
stale_label=(byid[stale].get("label") or "")[:70])
return queries, problems
def summarize(queries):
c = Counter(q["category"] for q in queries)
return ", ".join(f"{k}={c[k]}" for k in
("exact_rare", "phrase", "paraphrase", "associative", "nonsense", "superseded")
if c[k])
def main():
ap = argparse.ArgumentParser()
ap.add_argument("snapshot")
ap.add_argument("--out", default=DEFAULT_OUT)
ap.add_argument("--check", action="store_true",
help="validate only; do not write. Non-zero exit if anything is unwinnable.")
args = ap.parse_args()
nodes, edges = load_corpus(args.snapshot)
print(f"corpus: {len(nodes)} nodes, {len(edges)} edges ({os.path.basename(args.snapshot)})")
queries, problems = build(nodes, edges)
print(f"gold set: {len(queries)} queries [{summarize(queries)}]")
if problems:
print(f"\n{len(problems)} PROBLEM(S) — these queries were REJECTED, not silently kept:")
for p in problems:
print(" -", p)
if args.check:
sys.exit(1 if problems else 0)
doc = {
"corpus": os.path.abspath(args.snapshot),
"corpus_nodes": len(nodes),
"corpus_edges": len(edges),
"note": ("Every query carries a `derivation` recording how its expected answer was chosen. "
"Re-run with --check to re-validate the whole set against the corpus."),
"queries": queries,
}
with open(args.out, "w", encoding="utf-8") as fh:
json.dump(doc, fh, indent=1, ensure_ascii=False)
print(f"\nwrote {args.out}")
if __name__ == "__main__":
main()
+210
View File
@@ -0,0 +1,210 @@
#!/usr/bin/env python3
"""
compare.py diff two run_eval.py result files, WITH a noise threshold.
WHY THE STATISTICS ARE NOT OPTIONAL
With ~35 scored queries, one query is ~2.9 percentage points. A harness that
reports "hit@5 improved 2.9%" without saying that is one query is a harness
that will approve noise. So this file refuses to call anything an
improvement on the strength of the headline number alone. It reports:
1. The DISCORDANT PAIRS. Two configurations scored on the same queries are
paired data, so the only queries carrying information are the ones
where they disagree: b = fixed by B, c = broken by B. Queries both got
right, or both got wrong, tell you nothing about which is better.
2. McNEMAR'S EXACT TEST on (b, c). Under the null "the change is a coin
flip", the discordant outcomes are Binomial(b+c, 0.5). The two-sided
exact p-value is computed here with no scipy dependency.
3. The MINIMUM DETECTABLE SWING for this gold set: the smallest number of
net-changed queries that would reach p < 0.05 if every discordant pair
fell the same way. Anything smaller is inside the noise band, and the
verdict line says so in those words.
Repeat-run variance is the other half of honesty. Spreading activation is a
stateful read (it reinforces what it touches), so identical inputs need not
give identical outputs. Pass --repeats to fold several runs of the same
config into an observed variance band; a delta inside that band is not real
either, however good its p-value looks.
usage:
python3 compare.py --baseline results-main.json --candidate results-act.json
python3 compare.py --baseline a.json --candidate b.json \
--repeats-baseline a2.json a3.json --repeats-candidate b2.json b3.json
"""
import argparse
import json
from math import comb
def binom_two_sided(b, c):
"""Two-sided exact binomial p for b successes in n=b+c at p=0.5."""
n = b + c
if n == 0:
return 1.0
k = min(b, c)
tail = sum(comb(n, i) for i in range(0, k + 1)) / (2 ** n)
return min(1.0, 2 * tail)
def min_detectable_swing(n_scored, alpha=0.05):
"""Smallest all-one-way discordant count reaching p < alpha.
If every query that changes changes in the same direction, the p-value is
2 * 0.5**n. Solve for the smallest n where that drops under alpha. This is
the FLOOR: any real change will have some discordance both ways, so the true
requirement is larger. Reporting the floor is the conservative move it is
the most generous threshold we would ever accept.
"""
n = 1
while n <= n_scored:
if 2 * (0.5 ** n) < alpha:
return n
n += 1
return n_scored
def load(path):
with open(path, encoding="utf-8") as fh:
return json.load(fh)
def row_map(doc):
return {r["id"]: r for r in doc["rows"]}
def outcome(r):
"""Binary per-query outcome used for the paired test.
hit@5 for scored queries; 'returned nothing' for the nonsense controls;
'correction outranks the stale node' for the superseded queries. One number
per query, so every query votes exactly once.
"""
if "clean" in r:
return 1.0 if r["clean"] else 0.0
if "outranks" in r:
return 1.0 if r["outranks"] else 0.0
return r.get("hit@5") or 0.0
def band(values):
return (min(values), max(values))
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--baseline", required=True)
ap.add_argument("--candidate", required=True)
ap.add_argument("--repeats-baseline", nargs="*", default=[])
ap.add_argument("--repeats-candidate", nargs="*", default=[])
ap.add_argument("--out", default=None)
args = ap.parse_args()
A, B = load(args.baseline), load(args.candidate)
ra, rb = row_map(A), row_map(B)
ids = [q for q in ra if q in rb]
n = len(ids)
aa, ab = A["aggregate"], B["aggregate"]
print(f"baseline {A['label']:14} soul={A['soul_md5'][:12]} {n} shared queries")
print(f"candidate {B['label']:14} soul={B['soul_md5'][:12]}")
print(f"corpus {A['corpus_nodes']} nodes / {A['corpus_edges']} edges "
f"(identical copy for both runs)\n")
metrics = [("hit@5", 1), ("recall@5", 1), ("recall@10", 1),
("precision@5", 1), ("mrr@10", 0)]
print(f" {'metric':14} {'baseline':>10} {'candidate':>10} {'delta':>10}")
for m, as_pct in metrics:
x, y = aa[m], ab[m]
if as_pct:
print(f" {m:14} {100*x:>9.1f}% {100*y:>9.1f}% {100*(y-x):>+9.1f}pp")
else:
print(f" {m:14} {x:>10.3f} {y:>10.3f} {y-x:>+10.3f}")
for m in ("latency_ms_p50", "latency_ms_p95"):
x, y = aa[m], ab[m]
ratio = f"{y/x:.2f}x" if x else "n/a"
print(f" {m:14} {x:>9.0f}ms {y:>9.0f}ms {ratio:>10}")
print(f" {'nonsense':14} {aa['nonsense_clean']:>10} {ab['nonsense_clean']:>10}")
print(f" {'outranks':14} {aa['superseded_outranks']:>10} {ab['superseded_outranks']:>10}")
print(f"\n {'category':14} {'n':>3} {'base hit@5':>11} {'cand hit@5':>11} {'delta':>9}")
for c in sorted(set(aa["by_category"]) & set(ab["by_category"])):
ea, eb = aa["by_category"][c], ab["by_category"][c]
if c == "nonsense":
print(f" {c:14} {ea['n']:>3} {'clean ' + str(ea['clean']):>11} "
f"{'clean ' + str(eb['clean']):>11}")
else:
print(f" {c:14} {ea['n']:>3} {100*ea['hit@5']:>10.1f}% {100*eb['hit@5']:>10.1f}% "
f"{100*(eb['hit@5']-ea['hit@5']):>+8.1f}pp")
# ---- paired significance -------------------------------------------------
fixed, broken = [], []
for q in ids:
oa, ob = outcome(ra[q]), outcome(rb[q])
if ob > oa:
fixed.append(q)
elif ob < oa:
broken.append(q)
b, c = len(fixed), len(broken)
p = binom_two_sided(b, c)
mds = min_detectable_swing(n)
print(f"\n== paired comparison over {n} queries ==")
print(f" fixed by candidate : {b} {[ra[q]['category'] + ':' + q for q in fixed]}")
print(f" broken by candidate: {c} {[ra[q]['category'] + ':' + q for q in broken]}")
print(f" discordant pairs : {b + c} net {b - c:+d} queries")
print(f" McNemar exact p : {p:.4f}")
print(f" noise threshold : a difference needs at least {mds} queries moving the "
f"same way to clear p<0.05 on this {n}-query set")
# ---- repeat-run variance -------------------------------------------------
var = {}
for name, paths, first in (("baseline", args.repeats_baseline, A),
("candidate", args.repeats_candidate, B)):
docs = [first] + [load(p) for p in paths]
if len(docs) > 1:
hits = [d["aggregate"]["hit@5"] for d in docs]
lo, hi = band(hits)
spread_q = round((hi - lo) * first["aggregate"]["n_scored"])
var[name] = {"runs": len(docs), "hit@5_min": lo, "hit@5_max": hi,
"spread_queries": spread_q}
print(f" {name} repeat runs ({len(docs)}): hit@5 {100*lo:.1f}%..{100*hi:.1f}% "
f"= {spread_q} query of run-to-run drift")
drift = max([v["spread_queries"] for v in var.values()], default=0)
floor = max(mds, drift + 1)
print("\n== VERDICT ==")
net = b - c
if abs(net) < floor:
print(f" NO MEASURABLE DIFFERENCE. Net {net:+d} queries is inside the noise band "
f"(needs |net| >= {floor}: {mds} for significance, {drift} observed run-to-run drift).")
elif net > 0:
print(f" CANDIDATE BETTER by {net} queries (p={p:.4f}), outside the noise band "
f"(>= {floor}).")
else:
print(f" CANDIDATE WORSE by {abs(net)} queries (p={p:.4f}), outside the noise band "
f"(>= {floor}).")
if args.out:
with open(args.out, "w", encoding="utf-8") as fh:
json.dump({
"baseline": A["label"], "candidate": B["label"],
"n_shared_queries": n,
"fixed_by_candidate": fixed, "broken_by_candidate": broken,
"discordant": b + c, "net_queries": net,
"mcnemar_exact_p": p,
"min_detectable_swing_queries": mds,
"observed_run_to_run_drift_queries": drift,
"noise_floor_queries": floor,
"verdict": ("no measurable difference" if abs(net) < floor
else ("candidate better" if net > 0 else "candidate worse")),
"baseline_aggregate": aa, "candidate_aggregate": ab,
"repeat_variance": var,
}, fh, indent=1)
print(f"\nwrote {args.out}")
if __name__ == "__main__":
main()
@@ -0,0 +1,150 @@
{
"baseline": "main-r1",
"candidate": "act-r1",
"n_shared_queries": 38,
"fixed_by_candidate": [],
"broken_by_candidate": [
"q07",
"q11",
"q12",
"q13",
"q36"
],
"discordant": 5,
"net_queries": -5,
"mcnemar_exact_p": 0.0625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 1,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1140.4,
"latency_ms_p95": 1584.1,
"latency_ms_max": 1627.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.22857142857142856,
"recall@5": 0.19087301587301586,
"recall@10": 0.24277210884353742,
"precision@5": 0.07428571428571429,
"mrr@10": 0.24154195011337865,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3208.8,
"latency_ms_p95": 4851.9,
"latency_ms_max": 5078.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.2857142857142857,
"recall@5": 0.09722222222222222,
"recall@10": 0.3567176870748299,
"mrr@10": 0.3505668934240363
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"repeat_variance": {
"baseline": {
"runs": 3,
"hit@5_min": 0.34285714285714286,
"hit@5_max": 0.34285714285714286,
"spread_queries": 0
},
"candidate": {
"runs": 3,
"hit@5_min": 0.22857142857142856,
"hit@5_max": 0.2571428571428571,
"spread_queries": 1
}
}
}
+587
View File
@@ -0,0 +1,587 @@
{
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"note": "Every query carries a `derivation` recording how its expected answer was chosen. Re-run with --check to re-validate the whole set against the corpus.",
"queries": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"relevant": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"derivation": "MINED: token 'unjailbreakable' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226 ('Daemon hidden substrate architecture ? implemented April 25 '), which is therefore the only possible correct answer."
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"relevant": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"derivation": "MINED: token 'engram-migrate' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5 ('Engram v0.1 complete ? April 27, 2026. Local-first spreading'), which is therefore the only possible correct answer."
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"relevant": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"derivation": "MINED: token 'cartabandonedevent' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091 ('MESSAGE FOR AUDIT AGENT af5a7352e70e80434 ? El Language Spec'), which is therefore the only possible correct answer."
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"relevant": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"derivation": "MINED: token 'pre-apprenticeship' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-89c02aae-d3ca-43f9-9e5d-eb369896276c ('William Fox Anderson ? Applicant Profile Personal: - Full N'), which is therefore the only possible correct answer."
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"relevant": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"derivation": "MINED: token 'inferencenodemanager' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-73969486-143f-4431-b5e6-6845d1cc9848 ('Soma inference backplane deployed April 28 2026. Architectur'), which is therefore the only possible correct answer."
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"relevant": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"derivation": "MINED: token 'clear-eyed' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is knw-c72597c5-c23d-4c08-8e9e-996dadf26a99 ('Clear Eyes ? The Incomplete World View'), which is therefore the only possible correct answer."
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"relevant": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482"
],
"derivation": "MINED: a verbatim correction Will issued; expected = every node containing the phrase. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 1 node(s); that set IS the answer key."
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"relevant": [
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"Kp???",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"ע?RGk?\tH(?"
],
"derivation": "MINED: the canonical biographical phrase; expected = every node containing it. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 16 node(s); that set IS the answer key."
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"relevant": [
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"derivation": "MINED: a named person appearing verbatim in the biography/value nodes. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 9 node(s); that set IS the answer key."
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"relevant": [
"%???2??jH??",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"7c9d4ab1-205d-4be8-bfae-e2c03a3a5010",
"9f291d20-0d32-413c-8c01-4416ccab4f7f",
"?;????n}rh?",
"???Ͼd??f??",
"?Z?.\f?0?]P?",
"Dp???]Q??k+",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"derivation": "MINED: the canonical DHARMA expansion, confirmed by Will April 24 2026. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 9 node(s); that set IS the answer key."
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"relevant": [
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"derivation": "MINED: a named person; rare enough that the answer set is unambiguous. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 2 node(s); that set IS the answer key."
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"relevant": [
"2a923500-d7e1-4b15-80e2-48dba65984ba",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"g?2睪A|?H\b",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de"
],
"derivation": "MINED: the DARMA expansion, quoted verbatim in the backlog item and its correction. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 14 node(s); that set IS the answer key."
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"relevant": [
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"derivation": "MINED: the paid-tier feature name as written in the roadmap nodes. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 1 node(s); that set IS the answer key."
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"relevant": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Do the Essential Thing While You Can', whose subject is Grandma Lucas dying in Feb 2006 without Will saying goodbye. Query names the event with none of the node's own vocabulary. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"away",
"elderly",
"passed",
"relative",
"stayed"
]
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"relevant": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Survival Is Not an Excuse to Stop', whose subject is enlisting in the Marines, a severe hernia, and sepsis. Query describes the episode obliquely. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"illness",
"quit",
"refused",
"sidelined",
"soldier"
]
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"relevant": [
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Honesty Before Comfort'. Query states the principle in wholly different words. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"choosing",
"fact",
"fiction",
"pleasant",
"uncomfortable"
]
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"relevant": [
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Precision Over Brute Force'. Query restates the claim with no shared vocabulary. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"beats",
"bloated",
"payload",
"tight"
]
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"relevant": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Capability Is a Debt You Owe the Moment'. Query states the obligation without the node's terms. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"able",
"coming",
"job",
"nobody"
]
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Knowledge Survives When Nothing Else Does', whose subject is the library following Will across 30+ moves. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"creditors",
"learning",
"seize",
"wealth"
]
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"relevant": [
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Earned Trust' ('Trust is demonstrated, not declared'). VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"assertion",
"proven",
"record",
"reliability",
"track"
]
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"relevant": [
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Constraints as Freedom'. Query is a restatement of the same claim. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"boundaries",
"confine",
"enable",
"instead"
]
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"relevant": [
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Change Is the Signal', the value VBD is built on. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"apart",
"cut",
"shifts",
"system",
"tells"
]
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"relevant": [
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - The System Must Accumulate'. Query is the accumulation claim in different vocabulary. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"compounds",
"day",
"instead",
"mind",
"resetting"
]
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"relevant": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Being Seen Is Rarer Than Being Known', whose subject is Sarah Bishop as the first person Will did not perform for. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"exterior",
"loved",
"polished",
"self",
"unedited"
]
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"relevant": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Hope Is a Conclusion'. Query restates 'a conclusion, not a premise'. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"arrive",
"assuming",
"cheerfulness",
"instead"
]
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"relevant": [
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Structure Is Not Inherited', whose subject is thirty moves between two parents' collapses. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"childhood",
"foundation",
"inherit",
"offering",
"solid"
]
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71 ('Value ? Do the Essential Thing While You Can'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"relevant": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-58874a74-b96f-4883-9e08-45707f4bd3ee ('Value ? Survival Is Not an Excuse to Stop'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (13 of 13; 0 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-58874a74-b96f-4883-9e08-45707f4bd3ee) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"relevant": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e ('Value ? Being Seen Is Rarer Than Being Known'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83 ('Value ? Constraints as Freedom'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (6 of 13; 7 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 7
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-e0423482-cfa5-4796-8689-8495c93b66bc ('Value ? Hope Is a Conclusion'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-e0423482-cfa5-4796-8689-8495c93b66bc) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"relevant": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8 ('Value ? Capability Is a Debt You Owe the Moment'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"relevant": [],
"derivation": "CONTROL: verified at build time that none of this string's tokens occurs anywhere in the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
"expect_empty": true
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"relevant": [],
"derivation": "CONTROL: verified at build time that none of this string's tokens occurs anywhere in the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
"expect_empty": true
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"relevant": [],
"derivation": "CONTROL: verified at build time that none of this string's tokens occurs anywhere in the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
"expect_empty": true
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"relevant": [
"mem-80d7416b-20e9-48a0-b176-b215527e2f56"
],
"derivation": "DERIVED: correction node is Will's confirmation that the H is intentional (DHARMA, not DARMA); the stale node is the surviving backlog item still titled 'Implement DARMA'. Scored on RANKING, not presence: the corrected node mem-80d7416b-20e9-48a0-b176-b215527e2f56 must be returned AND must rank above the stale node bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff.",
"must_outrank": [
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"stale_id": "bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"correct_label": "CORRECTION: The autonomous self-improvement architecture is DHARMA ? n",
"stale_label": "Implement DARMA ? Directed Autonomous Runtime Modification Architectur"
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"relevant": [
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8"
],
"derivation": "DERIVED: correction node is the 2026-06-17 confabulation flag establishing EXACTLY 6 provisionals; the stale node is the surviving memory that asserts 12 filed patents. Scored on RANKING, not presence: the corrected node 3cf706a1-3825-45d8-b0a9-06cae6cdf5b8 must be returned AND must rank above the stale node 936541a9-fabb-466b-9ca3-a78b17ad0c53.",
"must_outrank": [
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"936541a9-fabb-466b-9ca3-a78b17ad0c53"
],
"stale_id": "936541a9-fabb-466b-9ca3-a78b17ad0c53",
"correct_label": "memory:remembered",
"stale_label": "memory:remembered"
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"relevant": [
"mem-30425134-6008-4fd9-a3ee-67a7742c319b"
],
"derivation": "DERIVED: correction node is the 'CGI ARCHITECTURE - THREE LAYERS, MCP RETIRED' decision of April 30 2026; the stale node still records the MCP server as live. Scored on RANKING, not presence: the corrected node mem-30425134-6008-4fd9-a3ee-67a7742c319b must be returned AND must rank above the stale node mem-101e81b4-8097-4749-8d8d-7bb66de34517.",
"must_outrank": [
"mem-30425134-6008-4fd9-a3ee-67a7742c319b",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517"
],
"stale_id": "mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"correct_label": "CGI ARCHITECTURE ? THREE LAYERS, MCP RETIRED (April 30, 2026). Definit",
"stale_label": "GCloud MCP infrastructure ? April 27, 2026. Legion died (~19:30 UTC). "
}
]
}
+942
View File
@@ -0,0 +1,942 @@
{
"label": "act-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-act",
"soul_md5": "77722f5a9f49494bf735c2a4be1b5dc4",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7902,
"wall_clock_s": 121.2,
"child_pid": 78714,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.22857142857142856,
"recall@5": 0.18908730158730158,
"recall@10": 0.24277210884353742,
"precision@5": 0.06857142857142857,
"mrr@10": 0.24154195011337865,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3237.6,
"latency_ms_p95": 5264.0,
"latency_ms_max": 5510.9,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.2857142857142857,
"recall@5": 0.0882936507936508,
"recall@10": 0.3567176870748299,
"mrr@10": 0.3505668934240363
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 476.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 708.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 532.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 537.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 568.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 555.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-69fd6e83-7718-4824-8d66-f49d8954e224",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-c3d9d063-8c5d-45aa-900c-550914b2ff6d",
"kn-f838f113-76d5-4a15-9cef-14055c4723a3",
"bl-76e878aa-e1fe-468c-bf9c-854097cb7e0b",
"art-c71aef51-026f-4d63-80e9-2a0ec0dc3865",
"bl-e148d23c-24e8-4122-9915-d1c11f22052f",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451"
],
"n_returned": 10,
"latency_ms": 1906.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-80720fdf-7ce7-4d28-aff8-21028d3a8cfb",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"?Z?.\f?0?]P?",
"?Q??m?;`'"
],
"n_returned": 10,
"latency_ms": 1103.3,
"error": null,
"hit@5": 1.0,
"recall@5": 0.0625,
"recall@10": 0.1875,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 6,
"latency_ms": 1132.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5555555555555556,
"recall@10": 0.5555555555555556,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"bl-c9adb8e5-293f-4033-99f8-0405c17ef941",
"bl-7e7c3fdb-4132-487f-aa70-b2cd559cb7f0",
"bl-7aebe936-ac55-4f35-8932-adc5224ff854",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"mem-34f53a9d-a131-4f82-9dbd-b9eb4a9af52e",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"n_returned": 10,
"latency_ms": 960.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
";??A5???",
"art-94fae615-7cd5-4695-b968-977101b06a51",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-57164d5f-baf0-4149-957a-379a4e255d1a",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-8dbceb06-431a-416d-a723-e8c75d595154"
],
"n_returned": 10,
"latency_ms": 992.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"kn-b2a99cd7-b379-4d9b-a996-e347a02c7bad",
"bl-a313d67b-dd6d-4e5b-a55a-03bc7bda17ae",
"kn-8e1bfb48-33a9-45ad-8da7-e0bdaa5d34e7",
"art-92e1837c-5919-42d0-bbb0-4d924d7b2864",
"bl-a7a1428f-db9c-417b-8e2c-713b1f84dc1f",
"mem-f823e835-313f-4282-b4b3-ce527ffc2f7a",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?Z?.\f?0?]P?",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 2820.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.14285714285714285,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-e20944e5-eb16-4ab3-a84d-111e0fc817fa",
"bl-e20944e5-f4a6-44a0-91b1-73d04ebed120",
"mem-47f72b5b-6e8b-4293-94f1-350197b4809a",
"mem-a5f04e52-91f8-41d2-af27-8bf803621758",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"n_returned": 10,
"latency_ms": 2609.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"kn-82be4e41-96c5-4da3-85c0-cee10763d975",
"mem-ea487cb4-ed67-44ce-8402-b56bb28468d4",
"bl-2694b588-a6e3-43de-861c-fa7b0ec7e7fd",
"bl-14883d81-f7cb-46dd-82c2-a6e6980264e5",
"tag-dark-theme",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad"
],
"n_returned": 10,
"latency_ms": 5510.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"mem-0328c3cb-4550-4ce4-9284-152e832f08f6",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"bl-e5635b1a-c5d0-4caa-bcc8-6a726ea43685",
"bl-6172d035-dd94-4776-afdd-d8915f6fc375",
"bl-5bb8dedf-8498-4a9b-acdc-31cc9c738f2a",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
";??A5???"
],
"n_returned": 10,
"latency_ms": 5444.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"bl-d24fcce8-2b55-426f-867a-db3958a622d3",
"tag-phase-3",
"tag-project-structure",
"tag-anthropic-contrast",
"tag-voice-training",
"tag-ebd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
";??A5???",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 5056.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-fc893be3-e6b4-4ef6-93b0-d54ca5f89083",
"bl-57c5cf6b-81a5-4558-9902-5c02981fe273",
"tag-guilds",
"tag-kids",
"tag-coexistence",
"tag-cultivated-general-intelligence",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
";??A5???",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023"
],
"n_returned": 10,
"latency_ms": 3473.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-bd9fb314-e9d4-4b03-aef4-534dd57a2992",
"bl-b019ce7a-1b21-436e-812d-032f50c6c45f",
"bl-e98cdd4c-01b5-459e-9036-3578cd5d975a",
"bl-9ce4128a-9436-4b06-82bc-8a6faafa81e0",
"tag-stable-diffusion",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4446.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"mem-e321e54e-8bb3-4596-b13d-bb093d6b149d",
"mem-e32ba5a7-c147-4dc0-9479-b720d768eda6",
"mem-c7a77457-478d-4eb0-a116-67205a0066a4",
"bl-1b20e9bc-eb37-4907-8d63-e311fd61eab8",
"bl-aa762207-920d-45ab-b2a3-2f8154d7ef9b",
"tag-misalignment",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 3932.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"kn-48a01973-a025-471d-950f-b93e6a426d82",
"mem-23d22bc1-a097-446b-8f11-8aff099e0b76",
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"mem-ce5a2ffc-ad39-4728-9ac6-76fef507d5da",
"project-Stripe_Elements__not_hosted_checkout__Custom_URL__DAG_bundle_pricing__Stripe_Connect_80_20_",
"tag-provenance",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a"
],
"n_returned": 10,
"latency_ms": 4836.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"kn-79056192-7de8-486c-9565-f128439a6fcb",
"bl-f6236350-f7b8-4f4f-a702-9eef2eb76e4b",
"bl-7f33f1bc-99fa-4906-889f-a42375beea20",
"mem-ef0091d8-1b65-431e-afa8-c6c4ee5779c9",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"tag-__cultivation-metric____internal-state____dharma____evidence____novel-idea____gap-compression____values____microsoft__",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 4057.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-identity-studio",
"tag-finance",
"tag-temporal",
"tag-barkhausen",
"tag-ilogger",
"tag-design-first",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 5098.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"project-Goal_setting__alignment__scoring__cadence__Attaches_to_any_imprint_",
"tag-performed-values",
"tag-turing-test",
"tag-sealed",
"tag-part-5",
"tag-model",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
";??A5???"
],
"n_returned": 10,
"latency_ms": 4652.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-9590ba23-bddb-43e8-a571-68a263c4c364",
"project-Imprint__leadership_development__feedback_frameworks__performance__presence_",
"bl-a9e57bb2-00a1-4867-ab59-5d9271134b50",
"bl-c7793c4a-7630-47fc-a462-d23059087e80",
"tag-gateway_platform_neuron-technologies_go_proxy_llm",
"tag-fornax",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 4653.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"bl-7a13527b-3e0c-418a-9f37-88fd2152e5ce",
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"bl-5e390b10-8753-4f25-a1a5-b5dbbb002cbf",
"tag-ats",
"tag-data-model",
"tag-aggregation",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
";??A5???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4272.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"tag-memory-model",
"tag-storage",
"tag-ga4",
"tag-potions",
"tag-resonance",
"tag-neuron",
"? ?}&?#??X\b",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4541.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"mem-927f41ab-8ede-4f58-acb3-995db16ac775",
"mem-dbe80bc2-c602-46b0-b4ea-dd222e52bcde",
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"tag-upload-window",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 4883.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"mem-481c769c-68cc-45c7-bc37-c0d9778fa648",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-9d8f3c5b-4bac-41ce-8ac4-44733f99d6c8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3237.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-eb1c6d99-d603-4f33-be9a-c63a178690c6",
"mem-a3124d5b-2f50-477f-8bb5-06879f5a496c",
"bl-7fa1b1a8-b80a-4f28-b162-bfe73765b4f8",
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 2457.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"mem-265a7107-73b3-4410-9aff-43787d5f473b",
"mem-9d1bf963-1b40-4588-bdb3-0432646cc623",
"mem-46780047-63a0-4a86-a16b-638b72a7fb8d",
"mem-3987d374-3c48-4e8e-b06d-0c363f55ed9c",
"bl-e0a0df72-de6e-46ab-800b-e1e3e8dfc387",
"tag-__patents____swarm____claim-language____prior-art____filing__",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 2640.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-dca14c4c-4859-47b0-996e-33964ba61a87",
"bl-dc8c7e02-eb37-48ae-a6f8-9b512803ae16",
"mem-8d1bafe6-209c-456c-9a25-9a927bc5a16d",
"bl-0de4e61b-6562-49e5-b7df-ebb809a01723",
"bl-8ef1ba6b-3fa0-4dbd-98c5-31665e5694a1",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?Q??m?;`'",
"R^??m?;?'"
],
"n_returned": 10,
"latency_ms": 3705.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-394cc9e8-049b-45bc-a380-66314f14e367",
"bl-ffa22d7e-42a9-4bd2-a428-1d2df243ac93",
"bl-4a6746e8-191f-48fc-8bfb-c4dc73b80bcd",
"tag-command-pattern",
"tag-withholding",
"tag-offline",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4223.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 1333.9,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 880.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"kn-333542cb-6dab-4662-9725-bf7440d28bf7",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 8,
"latency_ms": 1452.7,
"error": null,
"clean": false,
"false_positives": 8
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"ctx-e5427d7d",
"mem-37b57f52-a29a-42cf-a07a-3c5f8a3598dd",
"project-worldweaver",
"tag-__kotlin____internal-state____pre-reasoning____post-reasoning____compression-ratio____dharma____cultivation__",
"tag-import",
"tag-__cgi____dharma____cultivation____five-primitives____seed-artifact____agi____intelligence____whitepaper____patent__",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 3843.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-bfe0fafd-2750-4fdc-b773-04e878b3b23f",
"art-7bdaff30-5af9-4f0a-93b1-751686f9de3d",
"mem-cf07910d-4676-4384-ab97-9cad946cd0b9",
"mem-f9da4b43-3724-4bc8-92f8-6f237c89dc4d",
"project-Convert_UTC_timestamps_to_Central_time_when_displaying_to_Will__Never_surface_raw_UTC_",
"mem-32203649-3213-4d6d-86fd-3d657ac70d77",
";??A5???",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 5264.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"project-Define_Iris_as_separate_public_brand__Uncensored_but_principled__Consumer_face_while_Neuron_runs_enterprise_",
"project-Imprint__discovery__objection_handling__deal_strategy__pipeline__closing_",
"bl-8116da7a-b039-4e08-b8d0-c1c7861f9766",
"tag-enterprise",
"tag-divisors",
"tag-distressed-property",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"???Ͼd??W\b?",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 3022.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+942
View File
@@ -0,0 +1,942 @@
{
"label": "act-r3",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-act",
"soul_md5": "77722f5a9f49494bf735c2a4be1b5dc4",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7903,
"wall_clock_s": 112.0,
"child_pid": 78802,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.2571428571428571,
"recall@5": 0.19291383219954647,
"recall@10": 0.244812925170068,
"precision@5": 0.08,
"mrr@10": 0.266031746031746,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3220.9,
"latency_ms_p95": 4840.2,
"latency_ms_max": 5073.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.42857142857142855,
"recall@5": 0.10742630385487528,
"recall@10": 0.366921768707483,
"mrr@10": 0.47301587301587306
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 471.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 548.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 470.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 476.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 515.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 458.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-69fd6e83-7718-4824-8d66-f49d8954e224",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-c3d9d063-8c5d-45aa-900c-550914b2ff6d",
"kn-82be4e41-96c5-4da3-85c0-cee10763d975",
"bl-76e878aa-e1fe-468c-bf9c-854097cb7e0b",
"art-9887867c-2e21-47c3-9f96-c2dfe5bd4cc1",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451"
],
"n_returned": 10,
"latency_ms": 1595.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"bl-80720fdf-7ce7-4d28-aff8-21028d3a8cfb",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"?Z?.\f?0?]P?",
"?Q??m?;`'"
],
"n_returned": 10,
"latency_ms": 927.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.125,
"recall@10": 0.1875,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 6,
"latency_ms": 963.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5555555555555556,
"recall@10": 0.5555555555555556,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a",
"bl-7aebe936-ac55-4f35-8932-adc5224ff854",
"knw-b046991d-5992-4ac4-b854-7d3ac273832c",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"bl-455a08cf-5831-4fdb-b42c-b952f2feafb9",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"n_returned": 10,
"latency_ms": 927.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-8dbceb06-431a-416d-a723-e8c75d595154",
"mem-a3124d5b-2f50-477f-8bb5-06879f5a496c",
"art-94fae615-7cd5-4695-b968-977101b06a51",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-57164d5f-baf0-4149-957a-379a4e255d1a",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 937.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"kn-b2a99cd7-b379-4d9b-a996-e347a02c7bad",
"bl-a313d67b-dd6d-4e5b-a55a-03bc7bda17ae",
"kn-8e1bfb48-33a9-45ad-8da7-e0bdaa5d34e7",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023",
"bl-a7a1428f-db9c-417b-8e2c-713b1f84dc1f",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?Z?.\f?0?]P?",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 2451.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07142857142857142,
"recall@10": 0.21428571428571427,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-e20944e5-f4a6-44a0-91b1-73d04ebed120",
"mem-64cf3728-674c-404b-965a-b8f8d38bb7bb",
"mem-47f72b5b-6e8b-4293-94f1-350197b4809a",
"mem-e612f0aa-c2f2-4ee3-bbc7-af2dc826233b",
"mem-a5f04e52-91f8-41d2-af27-8bf803621758",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"n_returned": 10,
"latency_ms": 2042.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"mem-ea487cb4-ed67-44ce-8402-b56bb28468d4",
"art-c71aef51-026f-4d63-80e9-2a0ec0dc3865",
"bl-2dd8aaa1-b0de-4eac-b3c5-78951d240b60",
"bl-2694b588-a6e3-43de-861c-fa7b0ec7e7fd",
"bl-14883d81-f7cb-46dd-82c2-a6e6980264e5",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad"
],
"n_returned": 10,
"latency_ms": 4698.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"bl-fd047ce9-ae21-4b3e-b3ab-ece0c9592f7f",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"bl-9ce4128a-9436-4b06-82bc-8a6faafa81e0",
"bl-6172d035-dd94-4776-afdd-d8915f6fc375",
"bl-5bb8dedf-8498-4a9b-acdc-31cc9c738f2a",
"tag-cultivated-general-intelligence",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 5073.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"bl-e5635b1a-c5d0-4caa-bcc8-6a726ea43685",
"bl-fc893be3-e6b4-4ef6-93b0-d54ca5f89083",
"tag-project-structure",
"tag-anthropic-contrast",
"tag-voice-training",
"tag-ebd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792"
],
"n_returned": 10,
"latency_ms": 4117.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-d24fcce8-2b55-426f-867a-db3958a622d3",
"tag-phase-3",
"bl-57c5cf6b-81a5-4558-9902-5c02981fe273",
"tag-guilds",
"tag-kids",
"tag-coexistence",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023"
],
"n_returned": 10,
"latency_ms": 3254.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"mem-0328c3cb-4550-4ce4-9284-152e832f08f6",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-bd9fb314-e9d4-4b03-aef4-534dd57a2992",
"bl-e98cdd4c-01b5-459e-9036-3578cd5d975a",
"tag-stable-diffusion",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 4081.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"mem-e32ba5a7-c147-4dc0-9479-b720d768eda6",
"bl-b019ce7a-1b21-436e-812d-032f50c6c45f",
"mem-c7a77457-478d-4eb0-a116-67205a0066a4",
"bl-1b20e9bc-eb37-4907-8d63-e311fd61eab8",
"bl-aa762207-920d-45ab-b2a3-2f8154d7ef9b",
"tag-misalignment",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 3687.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"kn-48a01973-a025-471d-950f-b93e6a426d82",
"mem-23d22bc1-a097-446b-8f11-8aff099e0b76",
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"mem-ce5a2ffc-ad39-4728-9ac6-76fef507d5da",
"project-Stripe_Elements__not_hosted_checkout__Custom_URL__DAG_bundle_pricing__Stripe_Connect_80_20_",
"mem-e321e54e-8bb3-4596-b13d-bb093d6b149d",
"? ?}&?#??X\b",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9"
],
"n_returned": 10,
"latency_ms": 4659.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"kn-79056192-7de8-486c-9565-f128439a6fcb",
"bl-f6236350-f7b8-4f4f-a702-9eef2eb76e4b",
"bl-7f33f1bc-99fa-4906-889f-a42375beea20",
"mem-37b57f52-a29a-42cf-a07a-3c5f8a3598dd",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"tag-__cultivation-metric____internal-state____dharma____evidence____novel-idea____gap-compression____values____microsoft__",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"knw-b046991d-5992-4ac4-b854-7d3ac273832c"
],
"n_returned": 10,
"latency_ms": 3960.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-identity-studio",
"tag-finance",
"tag-temporal",
"tag-barkhausen",
"tag-ilogger",
"tag-design-first",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 4840.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"project-Goal_setting__alignment__scoring__cadence__Attaches_to_any_imprint_",
"tag-performed-values",
"tag-turing-test",
"tag-sealed",
"tag-part-5",
"tag-model",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4493.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-9590ba23-bddb-43e8-a571-68a263c4c364",
"project-Imprint__leadership_development__feedback_frameworks__performance__presence_",
"bl-a9e57bb2-00a1-4867-ab59-5d9271134b50",
"bl-c7793c4a-7630-47fc-a462-d23059087e80",
"tag-gateway_platform_neuron-technologies_go_proxy_llm",
"tag-fornax",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 4509.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"bl-7a13527b-3e0c-418a-9f37-88fd2152e5ce",
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"bl-5e390b10-8753-4f25-a1a5-b5dbbb002cbf",
"tag-ats",
"tag-data-model",
"tag-aggregation",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 3624.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"tag-memory-model",
"tag-storage",
"tag-ga4",
"tag-potions",
"tag-resonance",
"tag-neuron",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 4093.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"mem-927f41ab-8ede-4f58-acb3-995db16ac775",
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"tag-upload-window",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 4235.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"mem-0f31141d-3ac5-44b2-9942-be7e4e6feb79",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"mem-481c769c-68cc-45c7-bc37-c0d9778fa648",
"bl-9d8f3c5b-4bac-41ce-8ac4-44733f99d6c8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3220.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1c6d99-d603-4f33-be9a-c63a178690c6",
"mem-5e7f6ddd-c818-4ad3-b564-54ae278e9976",
"bl-7fa1b1a8-b80a-4f28-b162-bfe73765b4f8",
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c",
"tag-sarah",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 2336.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"mem-265a7107-73b3-4410-9aff-43787d5f473b",
"mem-9d1bf963-1b40-4588-bdb3-0432646cc623",
"mem-46780047-63a0-4a86-a16b-638b72a7fb8d",
"mem-3987d374-3c48-4e8e-b06d-0c363f55ed9c",
"bl-e0a0df72-de6e-46ab-800b-e1e3e8dfc387",
"tag-__patents____swarm____claim-language____prior-art____filing__",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 2348.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-dca14c4c-4859-47b0-996e-33964ba61a87",
"bl-e148d23c-24e8-4122-9915-d1c11f22052f",
"bl-dc8c7e02-eb37-48ae-a6f8-9b512803ae16",
"mem-8d1bafe6-209c-456c-9a25-9a927bc5a16d",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"tag-dark-theme",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?Q??m?;`'",
"R^??m?;?'"
],
"n_returned": 10,
"latency_ms": 3511.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-ffa22d7e-42a9-4bd2-a428-1d2df243ac93",
"mem-ef0091d8-1b65-431e-afa8-c6c4ee5779c9",
"bl-452a4710-3d2b-4e0f-9413-49a66423bc9a",
"bl-4a6746e8-191f-48fc-8bfb-c4dc73b80bcd",
"tag-withholding",
"tag-offline",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 4160.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 1318.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 875.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"kn-333542cb-6dab-4662-9725-bf7440d28bf7",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 8,
"latency_ms": 1393.4,
"error": null,
"clean": false,
"false_positives": 8
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"knw-b046991d-5992-4ac4-b854-7d3ac273832c",
"ctx-e5427d7d",
"project-worldweaver",
"tag-__kotlin____internal-state____pre-reasoning____post-reasoning____compression-ratio____dharma____cultivation__",
"tag-import",
"tag-enterprise",
"tag-__cgi____dharma____cultivation____five-primitives____seed-artifact____agi____intelligence____whitepaper____patent__",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3775.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-bfe0fafd-2750-4fdc-b773-04e878b3b23f",
"art-7bdaff30-5af9-4f0a-93b1-751686f9de3d",
"mem-cf07910d-4676-4384-ab97-9cad946cd0b9",
"mem-f9da4b43-3724-4bc8-92f8-6f237c89dc4d",
"project-Convert_UTC_timestamps_to_Central_time_when_displaying_to_Will__Never_surface_raw_UTC_",
"mem-32203649-3213-4d6d-86fd-3d657ac70d77",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 4857.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"bl-0de4e61b-6562-49e5-b7df-ebb809a01723",
"project-Imprint__discovery__objection_handling__deal_strategy__pipeline__closing_",
"bl-8116da7a-b039-4e08-b8d0-c1c7861f9766",
"bl-8ef1ba6b-3fa0-4dbd-98c5-31665e5694a1",
"tag-divisors",
"tag-distressed-property",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"knw-b046991d-5992-4ac4-b854-7d3ac273832c",
"???Ͼd??W\b?"
],
"n_returned": 10,
"latency_ms": 2858.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
@@ -0,0 +1,942 @@
{
"label": "act-r1",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-act",
"soul_md5": "77722f5a9f49494bf735c2a4be1b5dc4",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 112.3,
"child_pid": 78648,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.22857142857142856,
"recall@5": 0.19087301587301586,
"recall@10": 0.24277210884353742,
"precision@5": 0.07428571428571429,
"mrr@10": 0.24154195011337865,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3208.8,
"latency_ms_p95": 4851.9,
"latency_ms_max": 5078.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.2857142857142857,
"recall@5": 0.09722222222222222,
"recall@10": 0.3567176870748299,
"mrr@10": 0.3505668934240363
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 468.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 569.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 479.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 470.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 499.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 452.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-69fd6e83-7718-4824-8d66-f49d8954e224",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-eb1c6d99-d603-4f33-be9a-c63a178690c6",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"kn-b2a99cd7-b379-4d9b-a996-e347a02c7bad",
"bl-76e878aa-e1fe-468c-bf9c-854097cb7e0b",
"art-c71aef51-026f-4d63-80e9-2a0ec0dc3865"
],
"n_returned": 10,
"latency_ms": 1592.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"bl-80720fdf-7ce7-4d28-aff8-21028d3a8cfb",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"?Z?.\f?0?]P?",
"?Q??m?;`'"
],
"n_returned": 10,
"latency_ms": 956.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.125,
"recall@10": 0.1875,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 6,
"latency_ms": 970.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5555555555555556,
"recall@10": 0.5555555555555556,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"bl-7aebe936-ac55-4f35-8932-adc5224ff854",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"n_returned": 10,
"latency_ms": 920.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-8dbceb06-431a-416d-a723-e8c75d595154",
"mem-a3124d5b-2f50-477f-8bb5-06879f5a496c",
"knw-528dbc37-eabc-4b75-a7a5-65bf38d6018a",
"art-94fae615-7cd5-4695-b968-977101b06a51",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 925.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"mem-46780047-63a0-4a86-a16b-638b72a7fb8d",
"kn-8e1bfb48-33a9-45ad-8da7-e0bdaa5d34e7",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023",
"bl-a7a1428f-db9c-417b-8e2c-713b1f84dc1f",
"mem-f823e835-313f-4282-b4b3-ce527ffc2f7a",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?Z?.\f?0?]P?",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 2455.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.14285714285714285,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"art-92e1837c-5919-42d0-bbb0-4d924d7b2864",
"bl-e20944e5-eb16-4ab3-a84d-111e0fc817fa",
"bl-e20944e5-f4a6-44a0-91b1-73d04ebed120",
"mem-47f72b5b-6e8b-4293-94f1-350197b4809a",
"mem-e612f0aa-c2f2-4ee3-bbc7-af2dc826233b",
"mem-a5f04e52-91f8-41d2-af27-8bf803621758",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"n_returned": 10,
"latency_ms": 2037.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"kn-82be4e41-96c5-4da3-85c0-cee10763d975",
"mem-ea487cb4-ed67-44ce-8402-b56bb28468d4",
"mem-23d22bc1-a097-446b-8f11-8aff099e0b76",
"bl-874d1c2b-c55b-4afb-9601-922a9297e859",
"bl-2dd8aaa1-b0de-4eac-b3c5-78951d240b60",
"bl-2694b588-a6e3-43de-861c-fa7b0ec7e7fd",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad"
],
"n_returned": 10,
"latency_ms": 4720.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"mem-0328c3cb-4550-4ce4-9284-152e832f08f6",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-fd047ce9-ae21-4b3e-b3ab-ece0c9592f7f",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"bl-6172d035-dd94-4776-afdd-d8915f6fc375",
"bl-5bb8dedf-8498-4a9b-acdc-31cc9c738f2a",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 5078.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"bl-e5635b1a-c5d0-4caa-bcc8-6a726ea43685",
"bl-fc893be3-e6b4-4ef6-93b0-d54ca5f89083",
"tag-project-structure",
"tag-anthropic-contrast",
"tag-voice-training",
"tag-ebd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4118.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-d24fcce8-2b55-426f-867a-db3958a622d3",
"tag-phase-3",
"tag-identity-studio",
"tag-kids",
"tag-coexistence",
"tag-cultivated-general-intelligence",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023"
],
"n_returned": 10,
"latency_ms": 3241.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-bd9fb314-e9d4-4b03-aef4-534dd57a2992",
"bl-b019ce7a-1b21-436e-812d-032f50c6c45f",
"bl-e98cdd4c-01b5-459e-9036-3578cd5d975a",
"bl-9ce4128a-9436-4b06-82bc-8a6faafa81e0",
"tag-stable-diffusion",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4090.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"mem-e321e54e-8bb3-4596-b13d-bb093d6b149d",
"mem-e32ba5a7-c147-4dc0-9479-b720d768eda6",
"mem-c7a77457-478d-4eb0-a116-67205a0066a4",
"bl-1b20e9bc-eb37-4907-8d63-e311fd61eab8",
"bl-aa762207-920d-45ab-b2a3-2f8154d7ef9b",
"tag-misalignment",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 3671.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"kn-48a01973-a025-471d-950f-b93e6a426d82",
"bl-7f33f1bc-99fa-4906-889f-a42375beea20",
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"mem-ce5a2ffc-ad39-4728-9ac6-76fef507d5da",
"project-Stripe_Elements__not_hosted_checkout__Custom_URL__DAG_bundle_pricing__Stripe_Connect_80_20_",
"tag-provenance",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a"
],
"n_returned": 10,
"latency_ms": 4674.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"kn-79056192-7de8-486c-9565-f128439a6fcb",
"bl-f6236350-f7b8-4f4f-a702-9eef2eb76e4b",
"mem-ef0091d8-1b65-431e-afa8-c6c4ee5779c9",
"mem-37b57f52-a29a-42cf-a07a-3c5f8a3598dd",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"tag-__cultivation-metric____internal-state____dharma____evidence____novel-idea____gap-compression____values____microsoft__",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 3974.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"bl-57c5cf6b-81a5-4558-9902-5c02981fe273",
"tag-guilds",
"tag-finance",
"tag-temporal",
"tag-barkhausen",
"tag-ilogger",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 4851.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"project-Goal_setting__alignment__scoring__cadence__Attaches_to_any_imprint_",
"tag-turing-test",
"tag-sealed",
"tag-design-first",
"tag-part-5",
"tag-model",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4505.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-9590ba23-bddb-43e8-a571-68a263c4c364",
"bl-a9e57bb2-00a1-4867-ab59-5d9271134b50",
"tag-performed-values",
"bl-c7793c4a-7630-47fc-a462-d23059087e80",
"tag-gateway_platform_neuron-technologies_go_proxy_llm",
"tag-fornax",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 4494.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"bl-7a13527b-3e0c-418a-9f37-88fd2152e5ce",
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"bl-5e390b10-8753-4f25-a1a5-b5dbbb002cbf",
"tag-ats",
"tag-data-model",
"tag-aggregation",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 3619.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"tag-memory-model",
"tag-storage",
"tag-ga4",
"tag-potions",
"tag-resonance",
"tag-neuron",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 4107.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"mem-927f41ab-8ede-4f58-acb3-995db16ac775",
"mem-dbe80bc2-c602-46b0-b4ea-dd222e52bcde",
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"tag-upload-window",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 4231.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"mem-481c769c-68cc-45c7-bc37-c0d9778fa648",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-9d8f3c5b-4bac-41ce-8ac4-44733f99d6c8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3208.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"project-Imprint__leadership_development__feedback_frameworks__performance__presence_",
"mem-5e7f6ddd-c818-4ad3-b564-54ae278e9976",
"bl-7fa1b1a8-b80a-4f28-b162-bfe73765b4f8",
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c",
"tag-sarah",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 2359.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"mem-265a7107-73b3-4410-9aff-43787d5f473b",
"mem-9d1bf963-1b40-4588-bdb3-0432646cc623",
"bl-0de4e61b-6562-49e5-b7df-ebb809a01723",
"mem-3987d374-3c48-4e8e-b06d-0c363f55ed9c",
"bl-e0a0df72-de6e-46ab-800b-e1e3e8dfc387",
"tag-__patents____swarm____claim-language____prior-art____filing__",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 2342.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-dca14c4c-4859-47b0-996e-33964ba61a87",
"bl-e148d23c-24e8-4122-9915-d1c11f22052f",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"bl-dc8c7e02-eb37-48ae-a6f8-9b512803ae16",
"mem-8d1bafe6-209c-456c-9a25-9a927bc5a16d",
"bl-14883d81-f7cb-46dd-82c2-a6e6980264e5",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?Q??m?;`'",
"R^??m?;?'"
],
"n_returned": 10,
"latency_ms": 3536.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-ffa22d7e-42a9-4bd2-a428-1d2df243ac93",
"bl-452a4710-3d2b-4e0f-9413-49a66423bc9a",
"bl-4a6746e8-191f-48fc-8bfb-c4dc73b80bcd",
"tag-command-pattern",
"tag-withholding",
"tag-offline",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4151.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 1342.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 874.8,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"kn-333542cb-6dab-4662-9725-bf7440d28bf7",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 8,
"latency_ms": 1374.6,
"error": null,
"clean": false,
"false_positives": 8
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"ctx-e5427d7d",
"project-worldweaver",
"tag-__kotlin____internal-state____pre-reasoning____post-reasoning____compression-ratio____dharma____cultivation__",
"tag-import",
"tag-enterprise",
"tag-__cgi____dharma____cultivation____five-primitives____seed-artifact____agi____intelligence____whitepaper____patent__",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 3762.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-bfe0fafd-2750-4fdc-b773-04e878b3b23f",
"art-7bdaff30-5af9-4f0a-93b1-751686f9de3d",
"mem-cf07910d-4676-4384-ab97-9cad946cd0b9",
"project-Convert_UTC_timestamps_to_Central_time_when_displaying_to_Will__Never_surface_raw_UTC_",
"mem-32203649-3213-4d6d-86fd-3d657ac70d77",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 4859.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"project-Imprint__discovery__objection_handling__deal_strategy__pipeline__closing_",
"bl-8116da7a-b039-4e08-b8d0-c1c7861f9766",
"bl-8ef1ba6b-3fa0-4dbd-98c5-31665e5694a1",
"tag-dark-theme",
"tag-divisors",
"tag-distressed-property",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"???Ͼd??W\b?"
],
"n_returned": 10,
"latency_ms": 2873.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+945
View File
@@ -0,0 +1,945 @@
{
"label": "main-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-main",
"soul_md5": "caf822425540114984bfcd713ad37162",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7895,
"wall_clock_s": 45.6,
"child_pid": 78695,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1151.1,
"latency_ms_p95": 1580.7,
"latency_ms_max": 1663.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 231.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 266.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 229.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 235.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 254.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 228.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 543.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 466.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 474.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 460.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 477.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 824.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 681.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1577.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1663.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1331.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1036.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1363.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1230.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1548.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1315.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1580.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"?ǚ?7??????",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"????7???Ջ3",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"%???2??jH??",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1474.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1490.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1192.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1320.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1448.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1107.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1151.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1154.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1187.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1413.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 684.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 444.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 696.8,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1245.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1617.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 967.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+945
View File
@@ -0,0 +1,945 @@
{
"label": "main-r3",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-main",
"soul_md5": "caf822425540114984bfcd713ad37162",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7896,
"wall_clock_s": 45.7,
"child_pid": 78750,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1141.7,
"latency_ms_p95": 1578.0,
"latency_ms_max": 1645.5,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 237.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 272.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 231.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 231.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 262.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 239.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 542.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 465.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 470.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 441.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 455.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 816.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 678.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"????7???Ջ3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1564.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1645.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1320.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1029.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1347.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1220.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1560.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1305.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1578.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"????7???Ջ3",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"?ǚ?7??????",
"7?e?7???\f3?",
"??f?7???",
"? ?}&?#??X\b",
"%???2??jH??",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1"
],
"n_returned": 10,
"latency_ms": 1449.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1489.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1186.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1309.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1427.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1081.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1144.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1141.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1171.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1388.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 659.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 432.5,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 673.7,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1254.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1626.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 950.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+945
View File
@@ -0,0 +1,945 @@
{
"label": "main-r1",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-main",
"soul_md5": "caf822425540114984bfcd713ad37162",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7893,
"wall_clock_s": 45.9,
"child_pid": 78554,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1140.4,
"latency_ms_p95": 1584.1,
"latency_ms_max": 1627.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 240.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 271.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 232.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 227.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 259.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 237.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 543.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 467.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 475.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 449.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 463.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 821.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 688.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"????7???Ջ3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1559.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1627.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1317.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1055.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1352.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1213.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1545.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1327.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1584.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"????7???Ջ3",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"?ǚ?7??????",
"7?e?7???\f3?",
"??f?7???",
"? ?}&?#??X\b",
"%???2??jH??",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1"
],
"n_returned": 10,
"latency_ms": 1432.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1505.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1201.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1303.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1416.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1090.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1167.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1140.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1157.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1379.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 677.2,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 443.2,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 679.8,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1245.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1618.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 956.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+98
View File
@@ -0,0 +1,98 @@
#!/usr/bin/env bash
# run_comparison.sh — the whole harness, end to end, from two git refs.
#
# Builds a soul from each ref, boots each on its own throwaway port with its own
# throwaway HOME and its own disposable copy of the corpus, runs the gold set N
# times per ref, and prints the comparison with its noise threshold.
#
# SAFETY: never touches ~/.neuron, /Applications/Neuron*, ~/neuron-dev-stack, or
# any running service. Sources are exported with `git archive` into a scratch
# dir, so no worktree or branch state is mutated either. Ports are checked
# against the live set before anything boots. Every soul this script starts is
# killed and confirmed dead by run_eval.py; the sweep at the end is a backstop.
#
# usage:
# run_comparison.sh [--baseline main] [--candidate feat/recall-through-activation]
# [--repeats 3] [--corpus <snapshot.json>] [--repo <path>]
set -euo pipefail
BASELINE="main"
CANDIDATE="feat/recall-through-activation"
REPEATS=3
CORPUS="$HOME/neuron-memory-backups/snapshot-pre-repair-20260806.json"
REPO="$HOME/Development/neuron"
BASE_PORT=7893
while [ $# -gt 0 ]; do
case "$1" in
--baseline) BASELINE="$2"; shift 2 ;;
--candidate) CANDIDATE="$2"; shift 2 ;;
--repeats) REPEATS="$2"; shift 2 ;;
--corpus) CORPUS="$2"; shift 2 ;;
--repo) REPO="$2"; shift 2 ;;
*) echo "unknown arg: $1" >&2; exit 2 ;;
esac
done
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
WORK="$(mktemp -d "${TMPDIR:-/tmp}/retrieval-eval.XXXXXX")"
trap 'rm -rf "$WORK"' EXIT
[ -f "$CORPUS" ] || { echo "no corpus at $CORPUS" >&2; exit 2; }
echo "corpus: $CORPUS ($(du -h "$CORPUS" | cut -f1))"
slug() { printf '%s' "$1" | tr '/' '-'; }
build_ref() { # ref -> binary path
local ref="$1" out="$WORK/soul-$(slug "$1")"
local src="$WORK/src-$(slug "$1")"
mkdir -p "$src"
git -C "$REPO" archive "$ref" | tar -x -C "$src"
"$HERE/build-soul.sh" "$src" "$out" >&2
printf '%s' "$out"
}
echo "== building $BASELINE =="
BIN_A="$(build_ref "$BASELINE")"
echo "== building $CANDIDATE =="
BIN_B="$(build_ref "$CANDIDATE")"
echo "== validating the gold set against this corpus =="
python3 "$HERE/build_gold_set.py" "$CORPUS" --check
port=$BASE_PORT
run_one() { # binary label out
echo "== $2 =="
python3 "$HERE/run_eval.py" --soul "$1" --corpus "$CORPUS" --label "$2" \
--port "$port" --out "$3"
port=$((port + 1))
}
A_MAIN="$WORK/results-a-1.json"; B_MAIN="$WORK/results-b-1.json"
A_REP=(); B_REP=()
for i in $(seq 1 "$REPEATS"); do
a="$WORK/results-a-$i.json"; b="$WORK/results-b-$i.json"
run_one "$BIN_A" "$(slug "$BASELINE")-r$i" "$a"
run_one "$BIN_B" "$(slug "$CANDIDATE")-r$i" "$b"
[ "$i" -gt 1 ] && { A_REP+=("$a"); B_REP+=("$b"); }
done
cp "$A_MAIN" "$HERE/results-$(slug "$BASELINE").json"
cp "$B_MAIN" "$HERE/results-$(slug "$CANDIDATE").json"
echo
python3 "$HERE/compare.py" \
--baseline "$A_MAIN" --candidate "$B_MAIN" \
${A_REP[@]+--repeats-baseline "${A_REP[@]}"} \
${B_REP[@]+--repeats-candidate "${B_REP[@]}"} \
--out "$HERE/comparison-$(slug "$BASELINE")-vs-$(slug "$CANDIDATE").json"
# Backstop: run_eval.py kills and confirms its own child, but a crashed run
# could leak one. Leaving a soul running is how the live engine got squeezed.
STRAY=$(pgrep -f "$WORK/soul-" || true)
if [ -n "$STRAY" ]; then
echo "!! stray eval souls, killing: $STRAY" >&2
kill -9 $STRAY 2>/dev/null || true
fi
pgrep -f "$WORK/soul-" >/dev/null && { echo "!! STILL RUNNING" >&2; exit 5; }
echo "process check: no eval souls running"
+398
View File
@@ -0,0 +1,398 @@
#!/usr/bin/env python3
"""
run_eval.py measure one soul build's retrieval against the gold set.
WHAT IT MEASURES, AND WHY IT BOOTS A REAL SOUL
The point is Will's designed retrieval — spreading activation over the
weighted directed graph with four-factor multiplicative scoring not a
Python re-implementation of it. A re-implementation would measure my
reading of the design; booting the compiled binary measures the design. So
this harness compiles the actual `soul.el` amalgam (build-soul.sh) and asks
it over HTTP, exactly as the MCP wrapper and the app do.
SAFETY read this before changing anything here
* Boots on a THROWAWAY port with a THROWAWAY $HOME and a THROWAWAY COPY of
the corpus. Refuses to use 7770 / 8742 / 7779 / 17779 / 7771.
* ENGRAM_URL / SOUL_ENGRAM_URL are UNSET and SOUL_ISE_URL is pinned to a
dead port. This is not belt-and-braces: the periodic engram sync resolves
its source as env(SOUL_ISE_URL) -> state -> DEFAULT http://localhost:8742,
so leaving it unset makes an "isolated" run silently pull the operator's
LIVE brain. (Learned the hard way on 2026-08-03; see the same note in
scripts/verify-soul-contract.sh.)
* Every process this file starts is tracked and killed in a finally block,
then CONFIRMED dead by pid probe, and the confirmation is written into the
results file. A run that cannot confirm its child is dead exits non-zero.
* Activation is a STATEFUL read by design (patent claim 29: traversal
updates last-activation and increments activation counts). The corpus copy
is therefore per-run and disposable, and every run starts from a byte-
identical copy so two configurations see the same starting graph.
usage:
python3 run_eval.py --soul <binary> --corpus <snapshot.json> --label main \
[--port 7893] [--gold gold_set.json] [--limit 10] [--out results-main.json]
"""
import argparse
import json
import os
import shutil
import signal
import subprocess
import sys
import tempfile
import time
import urllib.error
import urllib.parse
import urllib.request
HERE = os.path.dirname(os.path.abspath(__file__))
FORBIDDEN_PORTS = {7770, 8742, 7779, 17779, 7771, 8080}
# ─────────────────────────────────────────────────────────────────────────────
# metrics
# ─────────────────────────────────────────────────────────────────────────────
def recall_at_k(returned, relevant, k):
if not relevant:
return None
return len(set(returned[:k]) & set(relevant)) / len(relevant)
def hit_at_k(returned, relevant, k):
if not relevant:
return None
return 1.0 if set(returned[:k]) & set(relevant) else 0.0
def precision_at_k(returned, relevant, k):
"""Fixed denominator k, as in docs/research/graphrag_eval/score.py.
Fixed denominator penalises an empty result and a page of junk equally,
which is what we want: a retriever that returns nothing is not 'precise'.
"""
if not relevant:
return None
return len(set(returned[:k]) & set(relevant)) / k
def mrr(returned, relevant, k):
if not relevant:
return None
rel = set(relevant)
for i, nid in enumerate(returned[:k], start=1):
if nid in rel:
return 1.0 / i
return 0.0
def mean(vals):
vals = [v for v in vals if v is not None]
return sum(vals) / len(vals) if vals else 0.0
def pct(vals):
return f"{100 * mean(vals):.1f}%"
# ─────────────────────────────────────────────────────────────────────────────
# soul lifecycle
# ─────────────────────────────────────────────────────────────────────────────
class Soul:
def __init__(self, binary, corpus, port, verbose=True):
if port in FORBIDDEN_PORTS:
raise SystemExit(f"REFUSING: port {port} is a live service port.")
self.binary = os.path.abspath(binary)
self.corpus = os.path.abspath(corpus)
self.port = port
self.verbose = verbose
self.home = None
self.proc = None
self.pid = None
self.log = None
self.confirmed_dead = None
@property
def base(self):
return f"http://127.0.0.1:{self.port}"
def start(self, boot_timeout=180):
self.home = tempfile.mkdtemp(prefix="retrieval-eval-home.")
snap = os.path.join(self.home, "corpus.json")
t0 = time.time()
shutil.copyfile(self.corpus, snap) # per-run disposable copy, never the source
self.log = os.path.join(self.home, "soul.log")
env = {k: v for k, v in os.environ.items()
if k not in ("ENGRAM_URL", "ENGRAM_API_KEY", "SOUL_ENGRAM_URL",
"ANTHROPIC_API_KEY", "NEURON_LLM_API_KEY", "SOUL_IDENTITY",
"SOUL_API_KEY")}
env.update({
"HOME": self.home,
"NEURON_PORT": str(self.port),
"SOUL_CGI_ID": f"ntn-retrieval-eval-{os.getpid()}",
"SOUL_ENGRAM_PATH": snap,
"NEURON_API_URL": "http://127.0.0.1:9", # dead port
"SOUL_ISE_URL": "http://127.0.0.1:9", # dead port — see SAFETY above
# Park the background loops for an hour so heartbeat/consolidation
# cannot mutate the graph between queries and make runs unrepeatable.
"SOUL_TICK_MS": "3600000",
"SOUL_HEARTBEAT_MS": "3600000",
"SOUL_REFRESH_MS": "3600000",
})
with open(self.log, "wb") as lf:
self.proc = subprocess.Popen([self.binary], env=env, stdout=lf, stderr=lf,
start_new_session=True)
self.pid = self.proc.pid
if self.verbose:
print(f" booted pid={self.pid} port={self.port} home={self.home}")
deadline = time.time() + boot_timeout
while time.time() < deadline:
if self.proc.poll() is not None:
raise RuntimeError(f"soul exited during boot: {self._log_tail()}")
rss = self._rss_kb()
if rss and rss > 6 * 1024 * 1024:
self.stop()
raise RuntimeError(f"soul RSS {rss}KB > 6GB — aborted")
try:
with urllib.request.urlopen(f"{self.base}/health", timeout=2) as r:
if r.status == 200:
if self.verbose:
print(f" healthy in {time.time() - t0:.1f}s, RSS={self._rss_kb()}KB")
return
except Exception:
pass
time.sleep(0.5)
self.stop()
raise RuntimeError(f"soul never healthy on {self.base}: {self._log_tail()}")
def _rss_kb(self):
try:
out = subprocess.run(["ps", "-o", "rss=", "-p", str(self.pid)],
capture_output=True, text=True, timeout=5).stdout.strip()
return int(out) if out else None
except Exception:
return None
def _log_tail(self, n=15):
try:
with open(self.log, encoding="utf-8", errors="replace") as fh:
return "\n".join(fh.read().splitlines()[-n:])
except Exception:
return "(no log)"
def recall(self, query, limit, timeout=60):
url = f"{self.base}/api/neuron/recall?query={urllib.parse.quote(query)}&limit={limit}"
t0 = time.perf_counter()
try:
with urllib.request.urlopen(url, timeout=timeout) as r:
raw = r.read().decode("utf-8", "replace")
ms = (time.perf_counter() - t0) * 1000
except Exception as exc:
return [], (time.perf_counter() - t0) * 1000, f"{type(exc).__name__}: {exc}"
try:
arr = json.loads(raw)
except Exception:
return [], ms, f"unparseable response ({len(raw)}B)"
if not isinstance(arr, list):
return [], ms, f"non-array response: {str(arr)[:120]}"
ids = [x.get("id") for x in arr if isinstance(x, dict) and x.get("id")]
return ids, ms, None
def stop(self):
"""Kill and CONFIRM. A test process that outlives its test is a bug."""
if self.pid is None:
self.confirmed_dead = True
return True
for sig in (signal.SIGTERM, signal.SIGKILL):
try:
os.kill(self.pid, sig)
except ProcessLookupError:
break
except Exception:
pass
for _ in range(20):
try:
os.kill(self.pid, 0)
except ProcessLookupError:
break
time.sleep(0.1)
else:
continue
break
try:
self.proc.wait(timeout=5)
except Exception:
pass
try:
os.kill(self.pid, 0)
self.confirmed_dead = False
except ProcessLookupError:
self.confirmed_dead = True
if self.verbose:
print(f" pid {self.pid}: {'CONFIRMED DEAD' if self.confirmed_dead else 'STILL ALIVE'}")
if self.home and os.path.isdir(self.home):
shutil.rmtree(self.home, ignore_errors=True)
return self.confirmed_dead
# ─────────────────────────────────────────────────────────────────────────────
# eval
# ─────────────────────────────────────────────────────────────────────────────
def evaluate(soul, gold, limit):
rows = []
for q in gold["queries"]:
ids, ms, err = soul.recall(q["query"], limit)
rel = q.get("relevant") or []
row = {
"id": q["id"],
"category": q["category"],
"query": q["query"],
"returned": ids,
"n_returned": len(ids),
"latency_ms": round(ms, 1),
"error": err,
}
if q.get("expect_empty"):
row["clean"] = (len(ids) == 0)
row["false_positives"] = len(ids)
else:
row["hit@5"] = hit_at_k(ids, rel, 5)
row["recall@5"] = recall_at_k(ids, rel, 5)
row["recall@10"] = recall_at_k(ids, rel, 10)
row["precision@5"] = precision_at_k(ids, rel, 5)
row["mrr@10"] = mrr(ids, rel, 10)
if q.get("must_outrank"):
correct, stale = q["must_outrank"]
ic = ids.index(correct) if correct in ids else None
istale = ids.index(stale) if stale in ids else None
# Correct must be present AND above the stale node. A run that
# returns neither is NOT a pass: the corrected fact is what the
# user needed.
row["outranks"] = (ic is not None) and (istale is None or ic < istale)
row["rank_correct"] = None if ic is None else ic + 1
row["rank_stale"] = None if istale is None else istale + 1
rows.append(row)
return rows
def aggregate(rows):
scored = [r for r in rows if "hit@5" in r]
nonsense = [r for r in rows if "clean" in r]
outrank = [r for r in rows if "outranks" in r]
lat = sorted(r["latency_ms"] for r in rows)
agg = {
"n_queries": len(rows),
"n_scored": len(scored),
"hit@5": mean([r["hit@5"] for r in scored]),
"recall@5": mean([r["recall@5"] for r in scored]),
"recall@10": mean([r["recall@10"] for r in scored]),
"precision@5": mean([r["precision@5"] for r in scored]),
"mrr@10": mean([r["mrr@10"] for r in scored]),
"nonsense_clean": f"{sum(1 for r in nonsense if r['clean'])}/{len(nonsense)}",
"superseded_outranks": f"{sum(1 for r in outrank if r['outranks'])}/{len(outrank)}",
"latency_ms_p50": lat[len(lat) // 2] if lat else 0,
"latency_ms_p95": lat[max(0, int(len(lat) * 0.95) - 1)] if lat else 0,
"latency_ms_max": lat[-1] if lat else 0,
"errors": sum(1 for r in rows if r["error"]),
"by_category": {},
}
cats = sorted({r["category"] for r in rows})
for c in cats:
cr = [r for r in rows if r["category"] == c]
if c == "nonsense":
agg["by_category"][c] = {
"n": len(cr),
"clean": sum(1 for r in cr if r["clean"]),
"avg_false_positives": mean([float(r["false_positives"]) for r in cr]),
}
else:
e = {
"n": len(cr),
"hit@5": mean([r.get("hit@5") for r in cr]),
"recall@5": mean([r.get("recall@5") for r in cr]),
"recall@10": mean([r.get("recall@10") for r in cr]),
"mrr@10": mean([r.get("mrr@10") for r in cr]),
}
if c == "superseded":
e["outranks"] = sum(1 for r in cr if r.get("outranks"))
agg["by_category"][c] = e
return agg
def print_table(label, agg):
print(f"\n=== {label} ===")
print(f" queries {agg['n_queries']} ({agg['n_scored']} scored + "
f"{agg['n_queries'] - agg['n_scored']} control) · errors {agg['errors']}")
print(f" {'hit@5':>12} {'recall@5':>10} {'recall@10':>10} {'prec@5':>9} {'MRR@10':>9}")
print(f" {pct([agg['hit@5']]):>12} {pct([agg['recall@5']]):>10} {pct([agg['recall@10']]):>10} "
f"{pct([agg['precision@5']]):>9} {agg['mrr@10']:>9.3f}")
print(f" nonsense clean {agg['nonsense_clean']} · superseded outranks {agg['superseded_outranks']}")
print(f" latency ms p50 {agg['latency_ms_p50']:.0f} · p95 {agg['latency_ms_p95']:.0f} "
f"· max {agg['latency_ms_max']:.0f}")
print(f"\n {'category':14} {'n':>3} {'hit@5':>8} {'recall@5':>9} {'recall@10':>10} {'MRR@10':>8}")
for c, e in agg["by_category"].items():
if c == "nonsense":
print(f" {c:14} {e['n']:>3} {'clean ' + str(e['clean']) + '/' + str(e['n']):>8}"
f"{'':>9} {'':>10} {'avg FP ' + format(e['avg_false_positives'], '.1f'):>8}")
else:
extra = f" outranks {e['outranks']}/{e['n']}" if "outranks" in e else ""
print(f" {c:14} {e['n']:>3} {pct([e['hit@5']]):>8} {pct([e['recall@5']]):>9} "
f"{pct([e['recall@10']]):>10} {e['mrr@10']:>8.3f}{extra}")
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--soul", required=True)
ap.add_argument("--corpus", required=True)
ap.add_argument("--label", required=True)
ap.add_argument("--gold", default=os.path.join(HERE, "gold_set.json"))
ap.add_argument("--port", type=int, default=7893)
ap.add_argument("--limit", type=int, default=10)
ap.add_argument("--out", default=None)
args = ap.parse_args()
with open(args.gold, encoding="utf-8") as fh:
gold = json.load(fh)
print(f"[{args.label}] soul={os.path.basename(args.soul)} "
f"corpus={os.path.basename(args.corpus)} gold={len(gold['queries'])}q limit={args.limit}")
soul = Soul(args.soul, args.corpus, args.port)
rows = []
started = time.time()
try:
soul.start()
rows = evaluate(soul, gold, args.limit)
finally:
dead = soul.stop()
agg = aggregate(rows)
print_table(args.label, agg)
out = args.out or os.path.join(HERE, f"results-{args.label}.json")
doc = {
"label": args.label,
"soul_binary": os.path.abspath(args.soul),
"soul_md5": subprocess.run(["md5", "-q", args.soul], capture_output=True,
text=True).stdout.strip(),
"corpus": os.path.abspath(args.corpus),
"corpus_nodes": gold.get("corpus_nodes"),
"corpus_edges": gold.get("corpus_edges"),
"gold_set": os.path.abspath(args.gold),
"limit": args.limit,
"port": args.port,
"wall_clock_s": round(time.time() - started, 1),
"child_pid": soul.pid,
"child_confirmed_dead": soul.confirmed_dead,
"aggregate": agg,
"rows": rows,
}
with open(out, "w", encoding="utf-8") as fh:
json.dump(doc, fh, indent=1, ensure_ascii=False)
print(f"\nwrote {out}")
if not dead:
print("FATAL: child process could not be confirmed dead", file=sys.stderr)
sys.exit(4)
if __name__ == "__main__":
main()
+97 -11
View File
@@ -41,6 +41,7 @@
#include <fcntl.h>
#include <dirent.h>
#include <errno.h>
#include <signal.h> /* SIGPIPE disposition — see el_runtime_ignore_sigpipe */
#include <pthread.h>
#include <curl/curl.h>
@@ -1238,16 +1239,77 @@ static const char* http_reason_phrase(int status) {
}
}
/* Best-effort send with retry on partial writes. */
/* ── A departing client MUST NOT be able to kill the daemon ──────────────────
* (2026-08-06, round 9.1 / ADR 0006 item 4.)
*
* Measured field failure: a client cancelled its request at 25 s; the handler
* finished its work at 116.9 s and wrote the reply into the departed client's
* socket. The second send() on a reset connection raised SIGPIPE, whose DEFAULT
* disposition terminates the process `exited due to SIGPIPE ... ran for
* 361177ms`. launchd respawned 4 ms later, so EVERY other in-flight request on
* that daemon lost its work, silently.
*
* Two independent guards, because one of them can be undone from outside this
* file (an embedder may reset signal dispositions) and the other cannot:
* 1. process-wide SIGPIPE -> SIG_IGN, installed at runtime init;
* 2. per-send suppression at the syscall (MSG_NOSIGNAL where the platform has
* it, SO_NOSIGPIPE on the accepted socket on macOS/BSD).
* With either in force, send() reports the peer's departure as EPIPE and the
* caller decides which is the point: this is an ordinary I/O outcome, not a
* fatal condition.
*
* It deliberately does NOT swallow the error. http_send_response() below
* classifies the errno and logs: "client left" for a departure, and a real
* "send failed: <strerror>" for anything else, so a genuine write fault is
* still visible in the log (spec round-9.1 §5.3). */
#ifndef MSG_NOSIGNAL
#define MSG_NOSIGNAL 0
#endif
void el_runtime_ignore_sigpipe(void) {
static int done = 0;
if (done) return;
done = 1;
struct sigaction sa;
memset(&sa, 0, sizeof(sa));
sa.sa_handler = SIG_IGN;
sigemptyset(&sa.sa_mask);
sigaction(SIGPIPE, &sa, NULL);
}
/* Suppress SIGPIPE for one accepted connection (macOS/BSD have no
* MSG_NOSIGNAL; they have the socket option instead). Best effort. */
static void http_socket_nosigpipe(int fd) {
#ifdef SO_NOSIGPIPE
int on = 1;
setsockopt(fd, SOL_SOCKET, SO_NOSIGPIPE, &on, sizeof(on));
#else
(void)fd;
#endif
}
/* Best-effort send with retry on partial writes.
* Returns 0 on success, -1 on failure with errno preserved for the caller. */
static int http_send_all(int fd, const char* p, size_t left) {
while (left > 0) {
ssize_t w = send(fd, p, left, 0);
if (w <= 0) return -1;
ssize_t w = send(fd, p, left, MSG_NOSIGNAL);
if (w < 0) {
if (errno == EINTR) continue; /* not an error — retry */
return -1; /* errno stays set for caller */
}
if (w == 0) { errno = EPIPE; return -1; }
p += w; left -= (size_t)w;
}
return 0;
}
/* Did this write fail because the client is gone, or because something is
* actually wrong with the socket? Only the first is routine. */
static int http_write_err_is_client_gone(int e) {
return e == EPIPE || e == ECONNRESET || e == ENOTCONN || e == ESHUTDOWN;
}
/* Discriminator that http_response() embeds at the start of its envelope.
* A handler returning a string starting with this exact prefix is treated
* as a structured response; anything else is treated as a raw body. */
@@ -1468,14 +1530,30 @@ static void http_send_response(int fd, const char* body) {
free(env_body); free(hdrs.buf); return;
}
if (http_send_all(fd, status_line, (size_t)sl) == 0
&& http_send_all(fd, hdrs.buf, hdrs.len) == 0
&& http_send_all(fd, tail, (size_t)tl) == 0
&& (head_only
/* HEAD requests echo headers + Content-Length but no body. */
? 1
: http_send_all(fd, eff_body, blen) == 0)) {
/* sent successfully */
/* The reply is written in four pieces; any of them can find the client
* already gone. errno is captured at the first failure, before any later
* library call can clobber it, and classified once below. */
errno = 0;
int send_err = 0;
if (http_send_all(fd, status_line, (size_t)sl) != 0) send_err = errno;
else if (http_send_all(fd, hdrs.buf, hdrs.len) != 0) send_err = errno;
else if (http_send_all(fd, tail, (size_t)tl) != 0) send_err = errno;
else if (!head_only /* HEAD echoes headers + Content-Length, no body. */
&& http_send_all(fd, eff_body, blen) != 0) send_err = errno;
if (send_err) {
if (http_write_err_is_client_gone(send_err)) {
/* ROUTINE. The user closed the window, quit the app, or cancelled.
* The work is done and the daemon keeps serving everyone else. */
fprintf(stderr, "[http] client left before the reply was written "
"(%zu-byte body, %s) - request completed, reply discarded\n",
blen, strerror(send_err));
} else {
/* NOT routine — a real write fault. Never let the client-gone case
* above hide this one. */
fprintf(stderr, "[http] send failed: %s (%zu-byte body)\n",
strerror(send_err), blen);
}
}
if (env_parsed_root) el_release(env_parsed_root);
@@ -1491,6 +1569,7 @@ static void* http_worker(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg;
int fd = a->fd;
free(a);
http_socket_nosigpipe(fd);
char *method = NULL, *path = NULL, *body = NULL;
if (http_read_request(fd, &method, &path, &body, NULL) == 0) {
http_handler_fn h = http_lookup_active();
@@ -1531,6 +1610,7 @@ static void* http_worker(void* arg) {
}
void http_serve(el_val_t port, el_val_t handler) {
el_runtime_ignore_sigpipe(); /* serving implies clients that leave */
/* If `handler` looks like a string name, register it as the active handler. */
const char* hname = EL_CSTR(handler);
if (hname && looks_like_string(handler)) {
@@ -1634,6 +1714,7 @@ static void* _http_serve_async_loop(void* raw) {
}
void http_serve_async(el_val_t port, el_val_t handler) {
el_runtime_ignore_sigpipe(); /* serving implies clients that leave */
const char* hname = EL_CSTR(handler);
if (hname && looks_like_string(handler)) {
http_set_handler(handler);
@@ -1821,6 +1902,7 @@ static void* http_worker_v2(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg;
int fd = a->fd;
free(a);
http_socket_nosigpipe(fd);
char *method = NULL, *path = NULL, *body = NULL, *hdr_block = NULL;
if (http_read_request(fd, &method, &path, &body, &hdr_block) == 0) {
http_handler4_fn h = http_lookup_active_v2();
@@ -1858,6 +1940,7 @@ static void* http_worker_v2(void* arg) {
}
void http_serve_v2(el_val_t port, el_val_t handler) {
el_runtime_ignore_sigpipe(); /* serving implies clients that leave */
const char* hname = EL_CSTR(handler);
if (hname && looks_like_string(handler)) {
http_set_handler_v2(handler);
@@ -5511,6 +5594,9 @@ el_val_t getpid_now(void) {
static el_val_t _el_args_list = 0;
void el_runtime_init_args(int argc, char** argv) {
/* First line of every generated main(): a client that leaves must never be
* able to signal this process to death. See el_runtime_ignore_sigpipe. */
el_runtime_ignore_sigpipe();
_el_args_list = el_list_empty();
for (int i = 1; i < argc; i++) {
_el_args_list = el_list_append(_el_args_list, EL_STR(argv[i]));