Compare commits

..

43 Commits

Author SHA1 Message Date
Neuron c9da8d917a chore: regenerate dist/soul.c after rebasing onto main
A clean rebase is not a consistent build input: git resolves the compiled amalgam
as an ordinary file and picks a winner, so dist/soul.c ends up holding one side's
code and not the other. The stamp gate catches it; this commit fixes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:12:53 -05:00
Tim Lingo 7f445a4b95 feat(engine): tools + agentic loop on the OpenAI wire, and two chat-breaking fixes found proving it
Teaches the OpenAI-format lane (Groq/OpenAI/Grok/Gemini/Ollama) to offer tools,
execute them, and loop — the capability that until now existed only on the
Anthropic wire. The tool-execution, consent, bridge and run-progress machinery is
reused unchanged; only the wire dialect is new.

Two pre-existing defects were found while proving it, and are fixed here because
both silently break chat:

1. PROVIDER WIRING NEVER CONNECTED. The launcher exports SOUL_LLM_PROVIDER /
   SOUL_LLM_BASE_URL and puts the provider key in ANTHROPIC_API_KEY + SOUL_API_KEY;
   the engine's provider fork read only NEURON_LLM_0_*, which nothing sets in a
   customer build. So use_openai was ALWAYS false: every non-Anthropic user's turns
   went to api.anthropic.com carrying, say, a Groq key, and came back
   "llm unavailable". Proven side-by-side against the pinned round-9 brain
   (sha256 15cf7d1b…): identical env, shipped brain = "llm unavailable" both chat
   modes with ZERO calls to the configured endpoint; this build = a real answer,
   with the probe logging POST /v1/chat/completions and Bearer <provider key>.
   Fixed brain-side only (env fallbacks) — no app or launcher change needed.

2. TRUNCATION SPLITS UTF-8 CHARACTERS. The session preload cuts recalled memory at
   fixed BYTE lengths (continuity snippet 350; session_preload_bullets per bullet).
   A cut landing inside a multi-byte character leaves a dangling lead byte in the
   SYSTEM PROMPT, making the whole request body invalid UTF-8 — providers reject it
   and the user sees an unexplained failure. Captured from a real body: 18,710 bytes,
   decode fails at 18,248 on 'e2', a box-drawing rule (U+2500 = E2 94 80) sliced in
   half. Trigger is ordinary content — em dash, curly quote, accented name, emoji,
   table border — and it gets MORE likely as memory grows. Shared code: this hit the
   Anthropic wire too. Fixed with utf8_safe_slice() applied at BOTH cut sites.

WHAT IS IN THE PORT
- llm_base_url / llm_wire_format / agentic_api_key: fall back to the launcher's own
  SOUL_LLM_* names; anthropic deliberately still returns "" so its native path is
  untouched (endpoint configurability remains neuron#62).
- openai_tools_json(): Anthropic tool schema -> OpenAI function schema; entries with
  no input_schema (Anthropic's server-side web_search) are skipped — they cannot
  execute on this wire.
- agentic_tools_no_web(): the standard set minus that server tool.
- openai_agentic_loop(): forked rather than parameterised, so agentic_loop — which
  carries every round-7/8/9 fix — is provably untouched. Same envelopes, same state
  keys, same consent policy (ask_all / escalate / builtin / always-allow), same
  client-bridge contract, same run-progress ledger, same 12-iteration cap.
- ADR-0005 mirrored on this wire: parallel_tool_calls:false is sent explicitly, and
  if a provider ignores it we honour the FIRST call and echo only that one, so the
  conversation we send is never self-contradictory. The drop is logged loudly.
- The assistant turn echoes the provider's own content bytes (json_get_raw), so a
  JSON null stays null and nothing is lost to a decode/re-encode round trip.
- Tool results are embedded already-escaped (dispatch_tool json_safe's them);
  truncation trims a dangling escape so a cut can't invalidate the body.
- bridge_save() gains a "wire" scalar and agentic_resume branches on it, so a
  suspended turn resumes on the wire it suspended on. Legacy blobs (no field) resume
  as anthropic. The field is read from the blob's SCALAR HEAD only — an unbounded
  first-match scan would run on into messages_raw, which is model-controlled, and
  that is exactly the round-9 resume defect. Pinned by a test.
- Three fork sites: handle_chat_agentic, handle_dharma_room_turn_agentic,
  agentic_resume. Tool assembly is computed once per lane at both entry points
  (it makes an HTTP call to the connector bridge; it was being paid for twice).

TOOLING THAT DID NOT EXIST
- tests/run-el-test.sh — engine tests were never runnable: elc is a compiler, it
  emits C and exits. This emits the test to C, compiles soul.c with main renamed
  away, links the rest + the repo-pinned runtime, and runs it. It also COMPUTES THE
  VERDICT, because every counted test file's "N passed, M failed" summary is a
  permanent 0/0 — the counters increment inside if BLOCKS, which El scoping
  discards (9 files; real fix filed as neuron#116). Proven to discriminate with a
  deliberately-broken assertion.
- tests/gate-openai/ — deterministic OpenAI-dialect provider stub + scenarios +
  driver + hostile modes, and a strict request validator that rejects any
  Anthropic-shaped field so dialect leakage fails loudly.

VERIFICATION (rungs named)
- E2E-VERIFIED against a LIVE provider (Anthropic's OpenAI-compatible endpoint,
  confirmed live): real answer; a tool call whose out-of-root path was DENIED by the
  guard, after which the model refused to claim success ("I won't tell you I did it,
  because I didn't"); then a valid path -> file physically on disk with exact content,
  honest reply, ledger with per-round entries + {done:true}.
- Deterministic lane gate: 11/12 in both consent configurations (bridge + local);
  hostile providers produce no hang and no fabricated answer; the 12-iteration cap
  trips with its honest message. The one FAIL is oa-tools-off and is NOT this port —
  see "Known, not fixed here".
- ANTHROPIC LANE UNCHANGED: gate9 32/32 on this build and on the pinned round-9
  brain; request bytes differ only within the noise band that two runs of the
  UNMODIFIED brain also produce (proven with a baseline-vs-baseline control), and
  the preload sections — the shared code touched here — are byte-identical.
  The rig discriminates: the round-8 brain scores 24/32 on it.
- verify-soul-contract.sh: PASS (27/27 routes, no hard-deletes).
- Unit: test_bridge_serialization 36/36 (incl. 8 new wire/field-order assertions),
  test_utf8_slice 18/18, test_agentic_tools 18 PASS / 0 FAIL / 3 documented skips.

KNOWN, NOT FIXED HERE (deliberate)
- Tools:Off on an OpenAI provider still fails: the non-agentic path goes through the
  el-runtime provider chain, which appends /v1/chat/completions to a base URL that
  already ends in /v1 -> /v1/v1/... 404. Runtime/plain-chat territory, untouched
  mid-beta. Note openai_chat_complete() has zero callers — that lane is served
  entirely by the runtime chain.
- The 12-iteration cap does not bound a chain of BRIDGED tools (iteration is
  per-invocation and resume starts fresh). Parity with the Anthropic lane.
- run_progress resets on each resume, so a client rendering cumulative steps across a
  consent pause sees earlier legs vanish. Parity with the Anthropic lane.
- verify-soul-contract.sh needs bash >= 4; under macOS's stock bash 3.2 it dies
  instantly with a FALSE red ("local: -n: invalid option").
- Groq-specific live E2E not run: no Groq key exists on this machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:12:13 -05:00
Tim Lingo f1f52bcb2f merge: Stage 1 structural audit as a real route (#142/#91)
Neuron Soul CI / build (push) Failing after 47s
Neuron Soul CI / deploy (push) Has been skipped
The runtime-vs-owner divergence check. Its absence let a ~24,000-node loss run for
weeks with every boot reporting green, which is most of why the last two days were
spent rediscovering by hand what this route would have said.

Follows the spec rather than inventing a metric: CGI provisional
05-detailed-description.md Stage 1 calls for an annotated characterization of the
graph's structure, so the route returns findings with score null BY DESIGN. A number
here would be a fabrication dressed as rigour.

Rebased across 27 commits of drift. The rebase merged cleanly at source level and
that was misleading — dist/soul.c held one side's code and not the other, because
git resolved the amalgam as an ordinary file. The stamp gate from 9fd8c11 caught it
on its first real use. Without it this would have landed an engine containing the
audit but none of the 08-09 engine work, or the reverse: neuron#133 again.

Verified before merging, not after:
  stamp OK                    dist/soul.c matches the sources (1,204,442 bytes, 1,259 bodies)
  builds from committed input 920,776 bytes
  interface                   108 -> 110 routes, nothing removed
  the route answers           stage 1, annotated_characterization, findings present

Closes #142. Refs #91.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:03:10 -05:00
Neuron aad988ecbf chore: regenerate dist/soul.c after rebasing onto main
The rebase merged cleanly at source level but left dist/soul.c holding one side's
code and not the other — main's regenerated amalgam vs this branch's. The stamp
gate added in 9fd8c11 caught it on its first real use:

  FAIL: dist/soul.c is STALE. It does not match the current .el sources.
  Sources that changed: neuron-api.el, routes.el

Without that gate this branch would have looked clean and shipped an engine
containing the structural audit but none of the 08-09 engine work, or the reverse.
That is exactly neuron#133, which once hid five merged fixes including a P0.

Regenerated from the rebased sources: 1,204,442 bytes, 1,259 inlined bodies.

Verified after: stamp OK; builds from its own committed input (920,776 bytes);
interface 108 -> 110 routes with nothing removed, adding /api/neuron/audit/structural;
and the route answers — stage 1, assessment_kind 'annotated_characterization',
score null by design, with findings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:01:25 -05:00
Tim Lingo 7b86e6f72c feat(soul): Stage 1 structural audit as a real route — an annotated characterization, not a score (#91)
`runStructuralAudit` has been an advertised MCP tool with nothing behind it: the
dispatcher GET'd /session/begin and returned that unrelated session digest under
an audit tool's name. Meanwhile the failure the audit would have caught ran
silently for about three weeks — the soul reporting 103,089 nodes while the
engram, which OWNS persistence, held ~79,900, a crash discarding the difference,
and every boot reporting green throughout, because nothing in the system ever
compared the two sides.

WHAT THE PATENT SPECIFIES, AND HOW IT SHAPED THIS
  CGI provisional, 05-detailed-description.md, "Stage 1: Structural audit 430".
  Two clauses did the design work. First the four things the module evaluates:
  the density and typed distribution of causal edges; value/execution-record
  consistency; the richness and connectivity of the self-model; and wonder-
  manifest authenticity. Second, and decisively: it "produces a coherence
  assessment 432 — NOT A BINARY SCORE but an annotated characterization of the
  graph's structural properties."

  So every finding carries its numbers AND a plain-language note saying what
  they mean and how they were obtained. There is no pass/fail and no composite
  health figure, and `"score":null` is emitted explicitly so a reader cannot
  mistake its absence for an omission.

WHAT IS IN STAGE 1 (four findings)
  owner_runtime_divergence   — the motivating case. Runtime counts vs the owner's
      own GET /api/stats, the delta, and the trend against the previous audit, so
      a second call answers "is the gap growing?" rather than restating it.
  self_model_connectivity    — the three identity pillars plus the self root:
      present, content length, one-hop degree. This RETIRES the Claude-side vitals
      identity block, which lived outside the system it was checking and went on
      reporting green while the memory-philosophy pillar was absent from the live
      graph. Asking the running soul is the designed mechanism; a shell probe was
      the fourth patch on the same hole.
  typed_edge_distribution    — exact counts against the claim-10 vocabulary, plus
      density, plus a separate count of LOWERCASE near-misses ("causes" vs
      "Causes"): "the vocabulary is unused" and "the vocabulary is misspelled by
      the write paths" are different defects with different fixes.
  orphans_and_dangling_edges — the tool's own long-standing promise.

WHAT IS DEFERRED, AND WHY IT IS DATA RATHER THAN A COMMENT
  Value/execution-record consistency and wonder-manifest authenticity ship as a
  `deferred` array that MEASURES the populations they would need (Prediction and
  WonderQuestion nodes) and reports those counts as the reason. Both are ~0 today
  — WonderQuestion because of a known write/read node-type mismatch. Asserting
  value coherence or a pull-weight correlation on an empty population would be a
  fabricated result, which is worse than a stated gap.

MEASUREMENT HONESTY: EXACT WHERE CHEAP, SAMPLED WHERE NOT, ALWAYS LABELLED
  Counts, edge typing and self-model connectivity are exact. Orphan and dangling
  rates are sampled, because engram_find_node_index is a linear scan — an
  exhaustive dangling check is O(nodes x edges), ~2.2e9 string compares at today's
  scale. Samples are UNIFORM across the whole population (str_index_of_all gives
  every edge offset in one pass, so any index is O(1); json_array_get would have
  been O(n^2)), never head-of-list, and each figure ships with its own sampled /
  population / exhaustive fields. ?edge_sample= and ?node_sample= at population
  size run either check exhaustively. The real fix is an id index in the runtime.

ONE BUG THIS FOUND IN ITSELF, CAUGHT IN TEST
  http_get does not return "" when the owner is unreachable — it returns a JSON
  error object. Testing only for "" made a DEAD owner read as reachable with
  node_count 0, so the audit reported 100% divergence and named it data loss.
  Reachability is now proved by the presence of the node_count field, and the
  owner's raw reply is attached. A confident wrong answer is exactly what this
  route exists to stop.

  Edge findings need relation labels and the runtime has no edge-enumeration
  builtin, so they use the same scratch export GET /api/graph/edges already uses
  (engram_save to TMPDIR, never the owner's canonical file — #117). That is a
  large write on a large graph, so this is a manual route, not a timer; ?edges=0
  skips it.

  neuron-api.el:900-1273  handler + helpers
  routes.el:567,752       GET and POST /api/neuron/audit/structural
  mcp-wrapper/src/main.el:113,682  tool description + dispatch off /session/begin
  dist/soul.c             regenerated (1255 bodies)

Rung: E2E-VERIFIED. Soul built from this branch (gen-soul-amalgam + cc-brain,
921,192 bytes, 16 warnings, 0 errors), booted on throwaway ports 7893/7896/7897
with throwaway HOMEs against a stub owner on 7894. Three scenarios pass: owner
reachable (runtime 62 vs owner 42, delta 20 / 32.2%, trend flat on the second
call; 12/20 edges claim-10 typed, 3 lowercase near-misses; 50/62 orphans, 3/20
dangling — every figure matches the fixture by construction), owner unreachable
(reported as a finding with the raw reply, not a crash), and file mode (owner
"none", divergence undefined). Reached end-to-end through the MCP tool via a
locally built wrapper. verify-soul-contract.sh: GATE PASS, 27/27 routes +
immutability. No process left running; live :7770 and :8742 untouched (GET only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:54:42 -05:00
Tim Lingo bf521974af merge: semantic retrieval, self-seeded context, connections on write, and one build lineage
Neuron Soul CI / build (push) Failing after 11m54s
Neuron Soul CI / deploy (push) Failing after 14m43s
Consolidates the 2026-08-09 engine work. Each piece was measured in an isolated lab
against the real corpus, gated, and verified on the operator machine before landing.

WHAT IT CONTAINS

  Semantic retrieval + the #146 delete fix, shipped together on purpose. The
  semantic leg was once wired into engram_search_json, which seven sites use as a
  keyed read and then delete every result from; one runs at every boot. Measured
  cost when live: ~234 real records destroyed per startup, identity among them.
  Retrieval must never reach production without the engram_recall_json split.

  Self-seeded compiled context. The design: 'Every compilation query begins at the
  self-model node and traverses outward.' Ours seeded from a hardcoded string, so
  compiled context held 0 identity records. Now 18, bounded the same way every other
  list in that handler is — the unbounded version is what set self_neighbors to []
  after it closed the socket on every call.

  Connections on write. CCR claim 29 requires linking candidates with typed edges
  'rather than appending as unlinked content'. Unlinked append was the only
  behaviour we had: 5% of nodes connected, no edge created by any write in 19 days.

  dist/soul.c drift becomes a build failure (#133), and deploys now build from that
  same committed input rather than a scratch amalgam — one lineage, provenance
  recorded, enforced by a gate in the deployer.

MEASURED, all on the real corpus or the live machine

  rephrased-question recall     0/43 -> 20/43 (bar was >=6 queries; nonsense controls held 10/10)
  identity in compiled context  0 -> 18
  connections per memory        0 -> ~2 on write; 2,828 backfilled into history
  no data loss across boots     44,418 / 44,418 / 44,418 over three restarts

FOUR DEFECTS FOUND BY MEASURING RATHER THAN READING, each fixed here

  an inert test arm (both arms byte-identical but for a float rounding artifact);
  an association hook wired into a path the request never takes (zero edges across
  four writes); an identity exclusion matching lowercase 'self' that let a memory
  link into the identity graph, whose verification shared the blind spot and printed
  PASS; a telemetry filter matching hyphenated 'state-event' that missed
  'INTERNAL STATE EVENT'. Common shape: a filter tested a proxy for the property it
  cared about. Assert the property.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:40:56 -05:00
Neuron 5743568bf1 chore(engine): deploys build from the committed input, not a second lineage
Until now deploys were built by build-soul.sh from a scratch amalgam while CI
compiled the committed dist/soul.c. Two lineages — 'what runs' and 'what the repo
says builds' were different artifacts. That is #133/#111 in another costume, and it
is how a round-9.1 brain shipped matching no committed source at all.

build-soul-from-dist.sh asserts dist/soul.c matches the .el sources, compiles it
with CI's own flags (-O2 -DHAVE_CURL -rdynamic; the CI comment explains -rdynamic —
without it the runtime cannot resolve its HTTP handler by name and the binary serves
nothing on every route), and writes a .provenance sidecar recording the dist/soul.c
hash, the stamp hash and the commit.

deploy_binary.sh gains GATE 0: refuse any soul without provenance, or built from a
dist/soul.c other than the one in the repo now. The override
(NEURON_DEPLOY_UNSTAMPED=i-accept-two-lineages) exists deliberately — a gate with no
escape hatch gets bypassed by disabling the gate, which is worse than one that
announces itself.

Verified before enforcing, so the sanctioned path is not a broken one:
  interface parity with the deployed binary   108 routes in, 108 out
  boots and serves                            2s
  carries today's work                        24 self-neighbours, 17 identity records
  GATE 0 refuses an unstamped binary          exit 10
  GATE 0 accepts the stamped one, deployed    durability PASS

Production now runs 4845db3d, built from dist/soul.c at 9fd8c11. After the swap:
identity in context 17, connections still form on write, semantic recall intact,
mind and store at delta 0.

Refs #133, #111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:38:01 -05:00
Neuron 9fd8c11670 chore(engine): make dist/soul.c drift a build failure instead of a silent ship (#133)
#133 regenerated the amalgam once and said so itself: 'Nothing in the tree
regenerates this file. Only a human running the recipe. It lags in batches, never
per-change, and it will drift again.'

It drifted again. Every binary deployed on 2026-08-09 was built by build-soul.sh
from a scratch amalgam that never touches dist/soul.c, so the committed build
input fell 2,761 bytes behind the sources by a different route than #133 describes.

Auto-regeneration is not available: the CI workflow records that elc needs 24GB+
of virtual memory and would OOM the runner. So the build cannot regenerate the
file. It can refuse to compile a stale one, for free and with no compiler.

tools/soulc-stamp.sh records a content fingerprint of every .el source at the
moment the amalgam is generated. --check recomputes and compares; divergence exits
1 and names the changed files and the recipe. Wired into CI ahead of the compile.

dist/soul.c regenerated from current sources: 1,176,361 -> 1,179,122 bytes, 1,247
inlined bodies (gate wants >=1200), and verified to compile clean at 903,552 bytes.

Demonstrated to FAIL on the bad input, per postmortem 0004's rule that a gate which
only passes on good input proves nothing:
  fresh stamp        -> OK,   exit 0
  one .el modified   -> FAIL, exit 1, names memory.el
  source restored    -> OK,   exit 0

Refs #133, #111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:32:56 -05:00
Neuron be0f9d1afe feat(soul): memories form connections when written
The design rejects what this system did: promotion must link candidates 'using
typed semantic edges rather than appending as unlinked content' (CCR claim 29).
Unlinked append was the only behaviour we had. Measured 2026-08-09: 14,214 edges
over 80,936 nodes, 5% of nodes connected to anything, and no edge created by any
write since 2026-07-19 across 27,000+ new nodes. A memory with no connections is
unreachable by spreading activation, so retrieval degrades to literal matching.

On write, a memory is now linked to its top related existing memories.

Bounds, each bought with a specific failure:
  - max 3 edges per memory (link_memories.py's cap: precision over spray)
  - never link to identity. The existing policy is explicit that 'memories must
    not pollute the self traversal by similarity; only an explicit citation may
    touch identity'.
  - never link telemetry (state-event, soul-response, boot_count, loop-outcome,
    search-result): ~97% of daily write volume. Linking it would add thousands of
    noise edges a day and re-flatten the graph in the name of connecting it.
  - fail-soft: a failed association never fails the write
  - associate only AFTER durability, so no edge points at a node that did not
    persist — that is the dangling-edge defect the 08-09 cleanup removed 830 of

Two defects found by measuring rather than reading, both fixed here:
  1. Hooking mem_store alone produced ZERO edges across four real writes. The HTTP
     memory route writes via wt_node directly; mem_store serves only the awareness
     telemetry paths we refuse to link.
  2. A lowercase-only identity check let a memory link to 'Self — Values
     (grounded)' — the exact pollution the policy forbids. My verification shared
     the blind spot and printed PASS. Now uses str_lower.

Measured, lab, same corpus: ~1-2 edges per memory; identity edges unchanged at
475; zero identity leaks post-fix including an adversarial batch of six memories
written about values. Edges route through wt_edge so they reach the owner.

Rung: E2E-VERIFIED in an isolated lab. Not deployed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 11:00:04 -05:00
Neuron 1742d0b575 feat(soul): seed compiled context from the self node, bounded
The design is explicit — 'Every compilation query begins at the self-model node
and traverses outward... structural reachability from the self-model node is a
precondition for any node to appear in compiled context.' (will-anderson
patents/drafts/engram-claims.md, Self-Seeded Activation. DRAFT, not a filed
provisional — cited as such.)

Measured before: compiled context held 0-1 identity records out of 10, because
compilation seeds from a hardcoded text string and never from the self. Even an
explicit 'my values identity who I am' query returned a boot counter and
state-events.

This restores the designed behaviour without repeating the failure that set
self_neighbors to [] originally: that was an unbounded ~90KB neighbour dump which
closed the socket on every call. Same bound as every other list in the handler.

Cap is 24 rather than 8 because the self root's first 8 neighbours are tag nodes
('neuron', 'tier:note', 'disposition:experimental', 'imprint', 'traversal') that
crowd out the substantive identity records behind them. The root has 34
neighbours of which 23 are identity.

Measured after, same corpus, binary as the only variable:
  identity records in compiled context  0 -> 17
  response size                         10,806 -> 18,985 bytes
  interface gate                        108 routes in, 108 out, none removed

Rung: E2E-VERIFIED in an isolated lab. Not deployed — this is the identity layer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 10:45:12 -05:00
Tim Lingo c6772e3d27 Merge branch 'feat/semantic-leg-dropout' into integrate/semantic-plus-writethrough
Neuron Soul CI / build (pull_request) Failing after 11m47s
Neuron Soul CI / deploy (pull_request) Failing after 14m48s
2026-08-09 09:47:44 -05:00
Tim Lingo b9ef66cae9 fix(engram): the semantic leg was deleting 234 memory records per boot
Neuron Soul CI / build (pull_request) Failing after 14m34s
Neuron Soul CI / deploy (pull_request) Has been skipped
The accumulated retrieval stack (iterations 1-9) put the claim-24 semantic
leg and the claim-10 associative leg on engram_search_json — the function
~40 internal .el call sites already used as a KEYED read. Seven of those
sites delete every record that comes back ("prune all existing X nodes,
keep exactly one"): memory.el:176, sessions.el:250/268/444/523,
soul.el:359.

mem_boot_count_inc() calls engram_search_json("soul:boot_count", 50) and
engram_forget()s all 50 results. With a lexical leg that returned 1 record.
With a semantic leg it returns 50 — the 49 nearest neighbours of the STRING
"soul:boot_count" — and the soul deletes them.

MEASURED on the harness corpus, isolated, read-only, zero writes from any
caller: 234 node records destroyed in a single boot. The deletion list is
the soul's own lookup result list, in rank order. Casualties include 6
Knowledge nodes, a layer-1 "CORE IDENTITY - GENESIS, LINEAGE" Memory, the
value node kn-58874a74, and the gold answers to 8 of the 75 gold-set
queries. After the fix: 1 deletion, which is the one the code intends.

THE BOUNDARY, from Will. Claim 24 authorises the vector index "to respond
to EMBEDDING SEARCH QUERIES by returning the node records whose embedding
vectors have the highest cosine similarity to a query vector". A keyed
state read is not an embedding search query; it is the identifier-keyed
retrieval of claim 23 ("node records are stored under a key encoding the
node identifier"). One function served both, so a nearest neighbour of
"soul:boot_count" was treated as a boot counter.

So: engram_search_json returns to its lexical contract, and the legs move
to engram_recall_json, which is what /api/neuron/recall reaches — the route
the MCP wrapper, the app, and this harness all call. Retrieval quality on
that route is unchanged by construction.

MEASURED, 75-query extended gold set, embedded corpus, vs the iteration-9
baseline: +3 / -0 (q15, q28, q60), p=0.2500, hit@5 53.8 -> 58.5%, latency
1.02x, every regression guard held, nonsense 10/10. Net +3 against a floor
of 6 is NOT-SHOWN and I am not calling it an improvement. The deliverable
is the defect.

Diagnostics kept, env-gated (EG_DIAG / EG_DIAG_ID), zero cost when unset:
node/embedding census at load, per-query leg dump, and a FORGET log — the
last is the regression detector for exactly this class of bug.

LIMIT, stated: handle_api_search_knowledge still uses the lexical function.
It is a retrieval surface and arguably wants the legs, but nothing in this
harness measures it, so I did not change unmeasured behaviour.
2026-08-07 18:12:43 -05:00
Tim Lingo 9717a4eeaf measure: claim 24 unflooring is +4 (NOT-SHOWN); asymmetric embedding prefixes are -5 (discarded)
Measured on the 75-query extended gold set (iteration 8's held-out extension)
against the certified stack baseline results-stack-ext.json, on the embedded
corpus. Three runs of the candidate, zero drift.

A. CLAIM 24 WITHOUT THE THRESHOLD - net +4, NOT-SHOWN, kept in the tree.
   fixed  : q14, q25 (in-sample paraphrase), q43, q52, q63, q67 (held-out)
   broken : q15 (paraphrase), q28 (associative)
   8 discordant, McNemar exact p = 0.2891, floor is 6.
   heldout_paraphrase 16.7% -> 30.0%, paraphrase 61.5% -> 69.2%.
   Every regression guard held: exact_rare 6/6, phrase 7/7, nonsense 10/10,
   superseded 2/3. Latency FLAT: p50 641 -> 632 ms.
   The in-sample half (+q14 +q25 -q15 -q28 = 0) was already on record in
   iteration 7's cmp-nogate.json, so only the held-out +4 is new.

B. ASYMMETRIC TASK PREFIXES ON THE EMBEDDER - net -5, REVERTED in this commit.
   Rationale was sound and the prediction was wrong, which is why it was worth
   measuring: nomic-embed-text is an asymmetric retrieval encoder and this file
   embedded query and document bare on both sides. Prefixing does exactly what
   the model card implies for the far-away cases - it rescued q42 (gold at
   GLOBAL COSINE RANK 25,564) and q39 - but it re-ranks the whole space and
   broke more than it fixed:
   fixed  : q24, q39, q42
   broken : q18, q19, q22, q31, q43, q44, q52, q63
   heldout_paraphrase 30.0% -> 23.3%, paraphrase 69.2% -> 53.8%.
   The corpus and the reproducer are kept (embed-corpus-prefixed.py,
   snapshot-pre-repair-20260806-embedded-prefixed.json) so nobody re-runs it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:45:05 -05:00
Tim Lingo 9790d9342d feat(engram): claim 24 without a threshold, and an embedding substrate that knows query from document
Two changes to the ONE leg that generalises. Iteration 8 measured that out of
sample the semantic leg contributes 100% of the stack's gain and the graph leg
contributes nothing, so this is where the remaining headroom is.

1. THE 0.60 FLOOR IS A PER-QUERY LOTTERY, AND CLAIM 24 HAS NO THRESHOLD IN IT.
   06-claims.md l.148: "respond to embedding search queries by returning the
   node records whose embedding vectors have the HIGHEST COSINE SIMILARITY to a
   query vector, independently of the spreading activation traversal." A
   ranking. ENGRAM_EMBED_SEED_MIN is defined at el_runtime.c l.6094 as the
   HippoRAG seed-JOIN threshold and l.6102 admits the read-path leg merely
   "reuses" it. Measured on the 30 held-out paraphrases: the query's own top-1
   cosine ranges 0.564-0.680, so the constant keeps a rank-1 answer for one
   query and discards a rank-1 answer for the next. Six golds sit at global
   cosine rank 1-2 scoring 0.564-0.589 - discarded by nothing but the constant.
   What holds the nonsense controls is the corpus-vocabulary gate (nhits == 0),
   not this floor. Cosine clamped to [0,1] per 05-detailed-description l.69.

2. THE VECTORS THEMSELVES ANSWER THE WRONG QUESTION. EL_EMBED_MODEL defaults to
   nomic-embed-text, an ASYMMETRIC retrieval encoder trained with task prefixes.
   Embedding query and document bare - as this file did on both sides - measures
   topical similarity rather than answer-hood. eg_embed_fetch now takes the task
   prefix: EL_EMBED_QUERY_PREFIX on the three query call sites, EL_EMBED_DOC_PREFIX
   on the two backfill sites. Restores no claim, and says so: Will specifies only
   "computed by an embedding model over the node's content" (l.17), so the model
   is his and its correct use is ours. It is the substrate under claim 24 -
   the index is only as good as the vectors in it.

Reproducer for the derived corpus: tools/retrieval-eval/embed-corpus-prefixed.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:34:58 -05:00
Tim Lingo 4eb4c9e287 feat(engram): word-start match primitive + corpus-vocabulary gate on recall
The retrieval match test is a raw substring scan, so a query token matches
anywhere INSIDE a corpus word: "throom" matches "bathroom". Measured over the
38-query gold set on this corpus that is not a rare accident - q28's lexical
leg is 36,954 records of which only 13 contain a query token at a word start
(99.96% mid-word noise), six further queries carry ~20,500 mid-word-only
records each, and the nonsense control q35 returns 7 records ALL of which
match only mid-word.

istr_contains_wordstart() anchors a token to a word start (preceding char not
alphanumeric) while still matching suffixes, so "value" still hits "values".
That empties the lexical leg for gibberish, and the nhits==0 corpus-vocabulary
gate (iteration 6's mechanism, feat/claim24-unfloored-semantic) then makes the
whole query decline rather than let the semantic leg answer it.

Measured vs feat/bm25-lexical-leg on the embedded corpus, 2 runs each,
0 queries of run-to-run drift on both sides:
  net +1 (nonsense:q35), 0 losses, McNemar p=1.0 -> NOT-SHOWN (floor is 6)
  nonsense clean 2/3 -> 3/3; exact_rare 100%, phrase 100%, paraphrase 61.5%,
  associative 66.7%, superseded 2/3 all UNCHANGED
  latency p50 1184 -> 543 ms (0.46x)

Iteration 6 called q35 "a DEFECTIVE CONTROL ... cannot be cleaned without
breaking the lexical leg". It can: the defect was the match primitive, and
cleaning it cost nothing.

Also committed: results-wsclaim24.json + cmp-nogate.json, a measured negative
for bundling the claim-24 unfloored semantic leg on top (gains q14/q25, breaks
q15/q28/q33/q34, net -2) - it independently reproduces iteration 6's q15/q28
losses and shows unflooring REQUIRES the vocabulary gate.

Reproducers: legs.py (leg-level replica, reproduces baseline hit@5 exactly on
all 38 queries), policy2.py, ceiling.py, wb2.py.
2026-08-07 16:58:34 -05:00
tim.lingo 9501e4ac12 Merge pull request 'fix(engine): memories written through the soul now reach the store that owns them (closes #117)' (#134) from feat/soul-write-through into main
Neuron Soul CI / build (push) Failing after 11m27s
Neuron Soul CI / deploy (push) Failing after 14m46s
2026-08-07 21:07:14 +00:00
Tim Lingo 55f9ee3cb0 measure: BM25 lexical leg vs semseed baseline - net +2 (q10,q11), NOT-SHOWN
hit@5 68.6% -> 74.3%, phrase 71.4% -> 100%, MRR@10 0.461 -> 0.502, latency
p50 0.97x. Zero losses, zero run-to-run drift on both sides. 2 queries moved
against a 6-query noise floor: NO MEASURABLE DIFFERENCE by the harness's own
test (McNemar exact p=0.50). Unaddressable records in returned slots: 57 -> 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:55:42 -05:00
Tim Lingo 9c39084e60 feat(engram): BM25-shaped lexical leg + addressability guard on the read path
engram_search_json ranked its lexical leg by raw distinct-token coverage with
salience as tiebreak: a token in 30,000 nodes counted the same as a token in 1,
and a 1.3 MB record matched nearly every query token by surface area alone.
Score it BM25-shaped instead - Lucene-form IDF and length normalisation over
the corpus mean - with per-token document frequency accumulated in the SAME
corpus pass that finds the hits (no extra scan, no extra round-trip).

Also refuse to return records whose identifier is not printable ASCII. This
corpus carries 1,032 such records (453 by the printable test) from a save-side
corruption; they occupy 125 of 303 returned slots on main. Claims 12, 23 and 27
all key on the node identifier, so such a record is unfetchable by any caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:51:49 -05:00
Tim Lingo 6f3a048f36 feat(engram): semantically seed the graph leg (Will's HippoRAG pass, SEED_K=8)
engram_assoc_leg previously took its seeds only from the top-3 LEXICAL hits.
For a paraphrase query the lexical hits are noise by construction, so the walk
never reached the neighbourhood that holds the answer. This adds the seeding
pass Will documents at el_runtime.c l.6082 — "Semantic seeding (HippoRAG
pattern, use similarity twice): the query is embedded, the top-K nodes by
cosine join the seed set" — using his own ENGRAM_EMBED_SEED_K (8).

Similarity is now used twice, coherently: cosine picks where to STAND in the
graph, the structural-relation walk decides what is REACHABLE, and cosine
orders what was reached (iteration 2's finding, unchanged).

The seed list is deliberately NOT floored at ENGRAM_EMBED_SEED_MIN. Measured
over all 38 gold queries: true paraphrase targets score cosine 0.46-0.66 and
the three nonsense controls' own nearest neighbours score 0.55/0.60/0.62 —
the distributions OVERLAP, so no absolute cosine floor separates signal from
gibberish. The gate that works is reachability: gibberish's nearest neighbours
carry no structural edge, so its graph leg is empty and the controls hold.

The raw top-K is selected inside the existing scoring pass, so the cosine is
computed exactly once per node: no extra corpus pass, no extra embed
round-trip, latency flat (p50 1220 -> 1227 ms, 1.01x).

Measured vs the certified baseline feat/hybrid-semantic-recall, embedded
corpus, 2 runs each, zero run-to-run drift on both sides:
  hit@5 51.4% -> 68.6%   MRR@10 0.387 -> 0.461
  paraphrase 38.5% -> 61.5%   associative 0% -> 66.7%
  exact_rare 100% held, nonsense 2/3 held, superseded 2/3 held
  phrase 85.7% -> 71.4% (q11, the known rank-5 rotation tax)
  net +6 queries (7 fixed / 1 broken), McNemar p=0.0703
2026-08-07 15:35:08 -05:00
Tim Lingo 059ce02003 feat(engram): an associative leg on the recall path (claim 10 typed relations)
The recall route had no way to reach a node that shares no token and no
embedding neighbourhood with the query. The design reserves that case for the
graph, and nothing on the read path consulted an edge.

This adds a third ranked leg beside the lexical and semantic ones: expand the
top 3 lexical hits along STRUCTURAL relations only (claim 10 — identity,
contains, superseded_by, references, ...), two hops, both directions, pruned
at the same 0.02 firing threshold engram_activate uses; order what was reached
by query similarity. Merged by strict rotation, never by score blending.

Not PR #135. That wired recall wholesale to engram_activate and lost 57 points
of phrase accuracy. The failure there was RANK, not reach — a 2-hop associate
at strength 0.06 cannot outrank thousands of 1-hop neighbours of strong
lexical seeds. Here the lexical leg is untouched and the associative list is
empty for most queries, because a node whose only edges are `tagged` and
`related` expands to nothing.

MEASURED, hybrid-semantic baseline -> this, 38-query gold set, embedded corpus:
  associative  0.0% -> 66.7%   (first non-zero ever recorded on that category)
  hit@5       51.4% -> 62.9%
  exact_rare, phrase, paraphrase, nonsense, superseded: all unchanged
  latency p50 1.01x
  4 queries moved, all gains, 0 losses, McNemar p=0.125
  deterministic: two runs of the same binary differ on 0 of 38 rows

VERDICT: NOT-SHOWN. The harness needs 6 queries to clear p<0.05 and the whole
associative category is only 6 queries, so even 4/6 fixed cannot reach the
floor. The mechanism is confirmed to work; the gold set cannot certify it.
2026-08-07 15:17:27 -05:00
Neuron 635453b936 feat(engram): rank-interleave the semantic leg into recall; embed the corpus
Replaces the score-fusion first cut with rank fusion, which is what the data
called for. nomic's cosine scale is compressed (true matches 0.55-0.70,
unrelated pairs 0.35-0.50), so an additive blend of cosine onto token-coverage
is dominated by whichever leg has the wider spread. Alternation is invariant to
both scales:

  L1, S1, L2, S2, ...  deduped, capped at limit

Lexical ranking is left byte-identical; the semantic ranking is computed beside
it and admitted only above ENGRAM_EMBED_SEED_MIN (0.60) — Will's existing seed
floor, no new tuning constant. That floor is what keeps the nonsense controls
clean: a query with no real match must not be answered with its neighbours.

embed-corpus.py / merge-corpus.py produce the derived corpus the semantic leg
needs (76,986 vectors, nomic-embed-text, 0 failures, 11 min). Zero of 78,791
nodes carried an embedding before this; the field round-tripped through the
snapshot but nothing ever wrote it.

MEASURED, 38-query gold set, paired against the SAME derived corpus so the
comparison isolates the code change:

  hit@5      34.3% -> 51.4%     paraphrase   0.0% -> 38.5%
  MRR@10     0.294 -> 0.387     superseded   1/3  -> 2/3 outranks
  recall@10  33.3% -> 50.5%     latency p50  1146 -> 1220ms (1.06x)

  exact_rare 100% -> 100%   phrase 85.7% -> 85.7%   nonsense 2/3 -> 2/3

  6 queries fixed, 0 broken, McNemar exact p=0.0312, 0 drift across repeats.

Regression guards all held. Contrast PR #135, which swapped the read path to
spreading activation wholesale: phrase 85.7 -> 28.6, latency 2.81x. Correct
mechanism, wrong substrate. The substrate is now present.

Restores engram claim 24 (previously 0% honoured).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:59:39 -05:00
Neuron 315b2eff00 feat(engram): fuse cosine similarity into the recall read path (claim 24)
engram_search_json — the function /api/neuron/recall actually reaches — ranked
only by distinct-token match count, so the embedding field on every node record
was inert. Add the semantic leg as a UNION beside the lexical one, not a
replacement for it:

  fused = (distinct_tokens_matched / query_tokens) + 0.90 * sem
  sem   = clamp01((cos(q,n) - 0.60) / (1 - 0.60))     ; 0 when not comparable

Holding the semantic weight strictly below 1.0 means a node matching every
query token can never be displaced by semantics alone — the regression guard
that PR #135 lacked when it swapped the read path to spreading activation and
took phrase recall from 85.7% to 28.6%.

No query embedding (embedder down, circuit breaker open) => sem == 0 for all
nodes => fused == sc/ntok, a monotone map of the old integer score, so the
ordering degrades to the historical behaviour exactly.

Restores engram claim 24: 'maintain a vector similarity index over the semantic
embedding vectors of all stored node records, and ... respond to embedding
search queries by returning the node records whose embedding vectors have the
highest cosine similarity to a query vector, independently of the spreading
activation traversal.'

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:44:42 -05:00
Neuron cf41d12d22 test(retrieval): a measurement harness for memory recall, and its first verdict
Neuron Soul CI / build (pull_request) Failing after 14m41s
Neuron Soul CI / deploy (pull_request) Failing after 14m45s
Nothing else on the memory roadmap should be built until a change can be shown
to help. Right now we judge by feel, and the benchmark literature is full of
systems that felt better and measured worse. This is the missing gate.

WHAT IT MEASURES, AND WHY IT BOOTS A REAL SOUL
The subject is Will's designed retrieval — spreading activation over the
weighted directed graph, four-factor multiplicative scoring — not a proxy for
it. A Python re-implementation would measure my reading of the design, so the
harness compiles the actual soul.el amalgam from a git ref and asks it over
HTTP on /api/neuron/recall, exactly as the MCP wrapper and the app do.

BUILT ON WHAT WAS ALREADY HERE, NOT AROUND IT
  docs/research/graphrag_eval/{collect,score}.py  — per-query relevant-id
    scoring and fixed-denominator precision@5 (kept verbatim: an empty result
    should be punished like a page of junk).
  docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py — the pinned
    ground truth + --check winnability gate, so every run judges alike.
  scripts/verify-soul-contract.sh — the isolation recipe, including the
    non-obvious SOUL_ISE_URL pin without which an "isolated" soul silently
    syncs the operator's live brain.
  gen-soul-amalgam.sh + .gitea/workflows/ci.yaml — the build recipe and flags.
New here: ids rather than regexes as ground truth, an associative category
derived from real edges, a superseded category scored on ranking, a
machine-checked zero-lexical-overlap guarantee on paraphrases, paired
significance testing, and measurement of the real compiled soul rather than an
offline replica of one leg of it.

THE GOLD SET IS AUDITABLE, NOT VIBES
38 queries over the real 78,768-node corpus, each carrying a `derivation`
string, each re-validated by `build_gold_set.py --check`. exact_rare is mined
(document frequency 1). phrase is mined (verbatim scan; >25 matches rejected as
too diffuse). paraphrase is hand-selected then PROVEN to share zero content
words with its target — a leak fails the build, so the category cannot decay
into lexical matching. associative is derived from real hub edges with
lexically-reachable siblings dropped. nonsense is verified absent. superseded
pairs are kept only when both sides survive as distinct nodes.

HONEST ABOUT NOISE
Minimum detectable swing on 38 queries is 6: if every changed query moves the
same way, p = 2*0.5^n first clears 0.05 at n=6. Run-to-run drift is measured,
not assumed — activation is a stateful read, and it shows: main is fully
deterministic across 3 runs, the candidate drifts by 1 query. compare.py
reports "no measurable difference" for anything inside max(6, drift+1).

FIRST VERDICT — feat/recall-through-activation
hit@5 34.3% -> 22.9%, phrase 85.7% -> 28.6%, latency p50 2.81x. Five discordant
pairs, all five against the candidate, none for it; McNemar exact p = 0.0625,
so by the stated rule this is one query short of significant and is reported as
such rather than as a win for main. The latency regression is deterministic and
not in any noise band.

The benefit the branch was written for is absent: associative recall is 0/6 on
BOTH builds. Probed directly, the traversal returns the lexical seed at rank 8
and none of its 12 hub siblings. Two measured corpus facts explain it — only
4,060 of 78,768 nodes (5.2%) carry any edge, and no node has an embedding, so
the fourth factor of the four-factor product has nothing to compute from. The
mechanism runs; the corpus lacks the structure it needs.

SAFETY
Throwaway port, throwaway HOME, disposable per-run copy of the corpus; live
ports refused by name. Every soul started is killed AND confirmed dead by pid
probe, with the confirmation written into the results file; run_comparison.sh
sweeps for strays and exits non-zero if any survive. Nothing under ~/.neuron,
/Applications/Neuron*, or ~/neuron-dev-stack is read, written, or restarted.

Rung: E2E-VERIFIED — 6 full runs (3 per config) against the real compiled
binaries on the real corpus; numbers above are measured, not projected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 13:40:51 -05:00
Tim Lingo dd952c0e46 feat(soul): write-through to the persistence owner — memories survive restart (#117)
Neuron Soul CI / build (pull_request) Failing after 14m57s
Neuron Soul CI / deploy (pull_request) Failing after 14m39s
The soul obeys half of its own ownership rule. soul.el:571-573 says "when
ENGRAM_URL is set the HTTP Engram owns persistence — the soul must NEVER write
to the local snapshot", and it doesn't. But nothing was ever built to hand the
soul's writes TO that owner: sync is pull-only (/api/sync -> engram_load_merge),
so every node created inside the soul lived in process RAM and was shed on
restart. Measured live 2026-08-07: soul node_count=102184, engram 79197.

SCOPE CORRECTION vs the earlier internal spec: engram provisional claim 17's
"pull-then-push" is a PEER-ENGRAM to PEER-ENGRAM protocol (claims 15-18 say so
explicitly). The soul is a CALLER of the database API, not a peer. Claim 17 is
NOT authority for a soul<->engram contract and is no longer cited as such. The
design here follows from the ownership rule alone.

Mechanism: a new Accessor, persist.el, is the single boundary. Writes stage a
delta to a filesystem spool and are pushed to the owner via POST /api/load-merge
— NOT POST /api/nodes, which mints a new server-side id (breaking dedup and
edges) and drops label/tier/tags/importance/confidence (verified in a sandbox:
a tier "Canonical" probe came back "Working"). load-merge preserves the id and
every field, dedups nodes by id and edges by (from,to,relation) so retries are
no-ops, and calls persist_canonical() so THE OWNER writes its own file — the
ownership rule is honoured rather than worked around.

Spool-and-drain rather than push-per-write: measured ~0.38s per load-merge at
live scale (79k nodes/176MB), and a chat turn writes 5-7 nodes. The spool is on
disk, not in process state, because the soul serves each connection on its own
pthread and a shared buffer would lose entries to a read-modify-write race. That
also buys crash recovery: writes orphaned by kill -9 are drained on next boot.

Honesty: api_persisted (the gate all 10 MCP write handlers pass through) and
mem_store now assert AT THE OWNER instead of reading back the soul's own RAM.
With the owner down a write returns {"ok":false,"error":"write_not_persisted"}
and the delta is queued — where main returns {"ok":true} for a write that dies.

Coverage: 35 node sites + 9 edge sites routed through the boundary. Deliberately
excluded, with reasons in persist.el: 4 InternalStateEvent sites (Will's own
telemetry carve-out), the boot counter and the persona (both already have
bespoke owner-side write-backs), and soul.el's 54 genesis identity edges
(file-mode only). engram_strengthen and engram_forget are NOT propagated —
load-merge cannot update or delete, and hard-deleting at the owner would fail
verify-soul-contract.sh section B.

Also fixed here:
- routes.el GET /api/graph/edges engram_save()'d straight over the owner's
  canonical snapshot.json — a read route, in a non-owner process, clobbering the
  canonical on every call. Same defect class Will removed from the engram in el
  dc39a61. Now exports to a scratch path. With this gone the soul writes nothing
  at all in HTTP mode.
- persist.el must clear the runtime's _tl_fs_read_len hint after every fs_read.
  In vendored runtime v1.0.0-20260501 that hint becomes the NEXT response's
  Content-Length, so reading a spool file mid-request made an 86-byte reply go
  out as 497 bytes with 411 bytes of adjacent heap trailing it. Caught and fixed
  at our boundary; the runtime class was fixed upstream in el 43636ae, which is
  not the pinned runtime here.

Rung: E2E-VERIFIED, discriminating. Same harness, same engram binary:
  write-through: LEG 1 PRESENT at owner, LEG 2 SURVIVED kill -9 + restart
  main:          LEG 1 ABSENT  at owner, LEG 2 LOST
verify-soul-contract.sh: GATE PASS on both builds (27/27 routes, immutability).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 12:47:31 -05:00
tim.lingo 18714e6142 Merge pull request 'fix(engine): restore multi-turn crisis escalation on the agentic path (P0, closes #129)' (#130) from fix/129-history-amplification into main
Neuron Soul CI / build (push) Failing after 14m37s
Neuron Soul CI / deploy (push) Has been skipped
2026-08-07 15:54:41 +00:00
tim.lingo 4936099c39 Merge pull request 'fix(engine): the daemon survives a client leaving, and says it is working while it works' (#127) from fix/liveness-engine-91 into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:54:15 +00:00
tim.lingo f1471763f5 Merge pull request 'fix(engine): approving a researched mission completes — the resume replay read a tool id out of the conversation (BUG-42, both faces)' (#115) from fix/resume-server-tool-replay into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:53:51 +00:00
tim.lingo 5850793b67 Merge pull request 'fix(engine): history keeps its provenance and its session — kills the false confession, the blank stare, and the "to.Good" seams' (#114) from fix/soul-history-provenance-20260805 into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-07 15:53:32 +00:00
tim.lingo fc1745c652 Merge pull request 'feat(engine): plain chat generates at L3 — inside the safety cycle, not around it (+ crisis-path segfault fix)' (#109) from feat/soul-plain-chat-generation-20260805 into main
Neuron Soul CI / build (push) Failing after 10m39s
Neuron Soul CI / deploy (push) Failing after 14m47s
2026-08-07 15:53:08 +00:00
Tim Lingo 43d0449904 fix(engine): the agentic crisis screen reads the session's own history again
P0 SAFETY. Closes the regression we introduced in ff421d3 (2026-08-05).

ff421d3 correctly moved conversation history to a per-session key via
conv_hist_key(session_id). One consumer did not move with it: the agentic
path's L1 safety screen kept reading the anonymous "conv_history" bucket. The
desktop app always mints a session id (DaemonClient.kt:706), so history was
always written under session_hist_<id> and that read always returned "".

The half of the crisis score that receives history is the escalation half — the
one that exists for distress building across several turns, where no single
message trips the bell on its own. It scored 0 on every real conversation for
two days. Single-message hard bell was never affected.

The bitter part: the comment that line carried documented this exact bug being
fixed once already, under issue #9. The fix was right then. The rename
re-broke it, and the comment went on describing a repair that no longer held.
A comment is not a gate.

The read now goes through conv_hist_key like every other consumer, including
the plain path at soul.el:398 and the thread-anchoring read thirty lines below
it in this same handler. It is one line. The rest of this commit is structure
so it cannot happen quietly again:

  - agentic_safety_screen() owns the two decisions that were inline — which
    window the screen sees, and the screen call. Inline safety inputs are
    untestable safety inputs; that is what let a rename starve this one with
    nothing failing and nothing logging.
  - the comment above the call site now states the invariant (read window ==
    written window) instead of naming a key that can be renamed out from under
    it.

TWO-LEG PROOF, one variable — the single line state_get("conv_history") ->
state_get(conv_hist_key(session_id)):

  before  scripts/run-el-test.sh tests/test_history_amplification.el
          3. REGRESSION #129 ... FAIL  got: soft_bell  expected: hard_bell
          8 passed, 1 failed          runner exit 1
  after   same command, same tree, that one line changed
          9 passed, 0 failed          runner exit 0

Full engine rebuild from these sources is clean: gen-soul-amalgam.sh ->
1,164,103 bytes / 1226 inlined bodies (gate wants >= 1200), cc-brain.sh ->
903,096 bytes, 0 errors. agentic_safety_screen and conv_hist_key both present
in the built binary (nm: T _agentic_safety_screen, T _conv_hist_key).

Rung reached: BUILT + RUNS (discriminating test). NOT yet in a DMG and not yet
verified in the app a human opens — those are the next two rungs and neither is
claimed here.

Known and NOT fixed by this commit:
  - feat/soul-openai-tools-v2 carries the same defect independently at
    chat.el:2937 and needs the same change or a merge.
  - the defect CLASS (a read of a state key no producer writes) is still
    invisible to every gate we have. Issue #129 proposes making it a build
    error; that is the follow-on.

Closes #129

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 09:33:05 -05:00
Tim Lingo b842e82f77 test(engine): a runner for tests/, and a failing regression test for #129
tests/ has held 14 test programs for months with no way to run them. CI does
not run them. The convention printed in their own headers
(`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
runtime does not implement. So the tests were documentation, not gates — which
is how a P0 safety regression shipped with a test directory sitting right
there.

scripts/run-el-test.sh compiles and runs one test program. It reuses the
gen-soul-amalgam.sh discovery: elc emits only an extern prototype for a module
that has a .elh beside it, and inlines the bodies when it does not, so a test
importing ../chat.el must be compiled in a scratch tree with the headers
removed. Scratch copy on purpose — the worktree is shared. It runs the binary
under a throwaway HOME so a test can never reach the live engram.

Exit status is the gate: the El tests print failures and still exit 0, so the
runner greps for FAIL lines and for a zero assertion count as well.

tests/test_history_amplification.el pins the invariant #129 violated: the
window the safety screen READS must be the window conv_history_record WRITES.
Not "must be called conv_history" — must AGREE.

THIS COMMIT IS RED BY DESIGN. On this tree the test fails one assertion:

  3. REGRESSION #129 — agentic screen reads the session's own window
    FAIL: distress history escalates the agentic screen to hard_bell
      got:      soft_bell
      expected: hard_bell
  history amplification tests: 8 passed, 1 failed   (runner exit 1)

The next commit turns it green by changing one line. Two legs, one variable —
that is the whole point of committing the test first.

Two flaws in the older harness that this one does not copy: the idiom
`let pass_count = pass_count + 1` inside an assert function declares a local
that dies with the call, so every existing suite prints "0 passed, 0 failed"
regardless of outcome; and a test program without a `cgi` block compiles as a
'utility', which may not reference the self-formation primitives chat.el's
agentic loop calls — it fails to build on a capability violation it never
triggers at runtime.

Refs #129

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 09:32:40 -05:00
Tim Lingo 98ccbd4704 fix(engine): a client that leaves must not kill the daemon, and a long round must say it started
Round 9.1, spec §3 D + ADR 0006 items 2 and 4. Two small changes, both proven
by measurement, both E2E-verified locally against a rebuilt brain.

D1 — SIGPIPE/EPIPE survival (vendor/el-runtime el_runtime.c).
Root cause, at the layer that owns it: the whole HTTP server lives in the C
runtime; .el has no socket primitive. http_send_all() called send() with flags
0 and nothing anywhere in the runtime set a SIGPIPE disposition, so the default
disposition — terminate the process — applied. When a handler finished after
its client had gone (Tim's VM: reply at 116.9 s, client cancelled at 25.0 s),
the second of the four sends that write one reply raised SIGPIPE and the daemon
died: `exited due to SIGPIPE ... ran for 361177ms`, launchd respawn 4 ms later,
every other in-flight session's work lost, user never told.

Fix: SIGPIPE -> SIG_IGN at runtime init and at each http_serve* entry, plus
per-connection SO_NOSIGPIPE / MSG_NOSIGNAL so the guard survives an embedder
resetting dispositions. http_send_all now retries EINTR and preserves errno;
http_send_response classifies it once — a departure is logged as routine
("client left before the reply was written ... reply discarded") and ANY other
errno is logged as a real "send failed: <strerror>". Spec §5.3: the routine
case must not mask a genuine write fault, and it does not.

Proof (scratch HOME + free port, 3 disconnects mid-reply):
  round-9 shipped brain 4402179554… — DIED, exit 141 (128+13 = SIGPIPE), round 1
  round-9 sources rebuilt with this exact recipe — DIED, exit 141, round 1
  this build — SURVIVED 3/3, /health 200 after, still serving the full graph,
  three honest "client left" lines in the log naming Broken pipe / Connection
  reset by peer.

D2 — the round-start marker (chat.el, agentic_loop).
The ledger only ever appended AFTER a round returned, so a healthy first leg
produced zero progress by construction; since server-side web_search moved
inside the outbound call that leg is 60-120 s of silence, which is how a 25 s
client watchdog came to kill a healthy mission. One entry,
{"i":N,"t":"","tool":"__working__"}, written to the existing
run_progress_<session_id> ledger BEFORE each round's outbound call — the wire
shape ChatView.kt:1148 has handled as a life signal since 2026-07-13 and never
received. No new key, no new route, no new lifecycle: a strict subset of WS3
item 3. WS3's run registry is untouched and stays Will's.

Proof (live Anthropic key, real research mission, scratch HOME + free port):
  round-9 baseline — ledger EMPTY for the whole 59.7 s leg
  this build       — {"i":0,"t":"","tool":"__working__"} visible at 18.6 s of a
                     70.0 s leg; both builds returned correct ~4.9 KB answers

Regression: prompt-matrix gate 32/32 on this build (round-9 baseline also 32/32
under the same recipe, so the score is not a build artifact). Soul contract
gate PASS — 27/27 routes, immutability clean. neuron#111 miscompile guard: 0
sites in the generated amalgam this binary was compiled from.

NOT included, deliberately: the regenerated dist/soul.c. CI compiles that file,
so production stays exposed until it is regenerated — the same open ask as
neuron#111 / ui#209. The regen recipe is now known and recorded; landing it is
Will's call, per BUILD-HYGIENE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 18:23:12 -05:00
Tim Lingo dba755dcec fix(engine): resume reads the bridged tool id from the blob's own field, not from inside the replayed conversation
ROOT CAUSE (round 9; live-repro'd 5/5 this morning, both faces stub-proven by the
prompt-matrix gate). json_get is a first-substring-match scanner (strstr for
'"key":', el_runtime.c). bridge_save serialized the RAW messages array BEFORE the
tool_use_id scalar, so agentic_resume's json_get(blob, 'tool_use_id') returned the
FIRST '"tool_use_id":' occurrence inside the replayed conversation, not the saved
field. The resume guard then preferred that misread over the client's correct
call_id (its two branches both reduced to saved_use_id), attached the tool_result
to the wrong id, and Anthropic 400'd the resume ('unexpected tool_use_id found in
tool_result blocks'), surfaced as {"error":"llm unavailable"}.

ONE MISREAD, TWO FACES — whichever block owns the first tool_use_id in the array:
  FACE 1 (search-then-bridge, the Key West killer): the first occurrence is the
    first web_search_tool_result's srvtoolu_… id — every agentic turn that ran
    server-side web_search and then bridged on a client tool died on approval,
    deterministically (messages.2.content.0 … srvtoolu_…). The write itself had
    already succeeded; only the resume died.
  FACE 2 (multi-cycle missions): with no search, the first occurrence is ROUND 0's
    tool_result block — so every LATER approve/resume cycle replayed the round-0
    client id (stale-resume-id), killing multi-file missions after ~2 files.
  And the shape that PASSES on round 8 confirms the mechanism: a single-cycle
  bridge with no prior tool round has no 'tool_use_id' substring in its messages
  at all (tool_use blocks carry 'id'), so the scan fell through to the blob's own
  field and resumed correctly.

The server_tool_use ↔ web_search_tool_result pairs themselves replay intact — the
defect was a cross-field misread of the blob, the same first-match-scanner class
as BUG-6 (approve 'content' matched inside tool_input, 2026-07-17) and round 8's
citation-block fix.

THE FIX, the pattern not the spot:
  1. bridge_save writes every json_safe'd scalar BEFORE both raw fields (an escaped
     value cannot contain a bare '"key":' byte pattern, so first-match always lands
     on the blob's own fields), and tools_raw (our fixed schema) before messages_raw
     (arbitrary conversation), so the raw extractions cannot first-match into
     model-controlled bytes either. Field order documented as load-bearing.
  2. agentic_resume now honors the client's echoed call_id when present — the value
     with clean provenance (minted from pend_tool_id, never blob-round-tripped) —
     falling back to the saved id only when the client omits it. Each approve cycle
     therefore binds to ITS OWN round's id (kills FACE 2 even against a blob written
     by a pre-fix binary), and an omitted call_id still resumes on the saved id,
     which the reordered blob now reads correctly.

Pattern sweep: the legacy synthetic blob (sessions.el handle_session_approve) embeds
only json_safe'd fields — no raw hazard, untouched. No other json_get read of any
container that embeds raw conversation JSON before the read field.

PROOF: prompt-matrix gate 24/32 RED on the round-8 brain (fails exactly the two
resume classes, named) -> 32/32 GREEN on this build; live-key Key West tracer
3/3 consecutive full round-trips (bridge -> approve-as-the-app -> real completion,
file on disk), plain-chat and weather-only controls PASS; unpatched round-8 brain
and a same-toolchain unpatched baseline build both still fail the identical
sequence with the identical srvtoolu 400 (the test discriminates, and the only
variable between failing and passing builds is this diff).

Refs neuron#109

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 11:34:12 -05:00
Tim Lingo 8f3a478771 fix(engine): excise the receipt, do not truncate at it — a leading receipt was erasing whole answers
CAUGHT BY A/B, AND ONLY BY A/B. The previous commit's receipt_strip assumed the receipt is
always TERMINAL and cut everything from the marker onward. It is not always terminal: once
receipt_rule told the model what [[RECEIPT ...]] means, the model sometimes LED with one and
wrote the answer underneath. Cutting at the marker then deleted the entire answer and the
turn returned {"error":"no response"}.

MEASURED, same prompt (two web searches, cited prose), fresh session each run:
    round-7 brain   4 / 4 answered   (471, 473, 473, 544 chars)
    round-8 brain   2 / 7 answered   (five {"error":"no response"})
This looked exactly like a flaky model. It was not — it was mine. Running the two brains
side by side on the same prompt is the only reason it was found, and it is the reason the
A/B is now part of how this class gets tested.

AFTER THE FIX, same protocol:
    round-7 brain   4 / 4   (473, 473, 473, 544)
    round-8 brain   4 / 4   (657, 657, 657, 673)

THE FIX: remove the [[...]] span and keep BOTH sides, instead of truncating at the marker.
An unterminated marker at position 0 is left completely alone — no rule about receipts is
worth erasing an answer over. Bounded four-pass loop rather than a conditional exit, because
rebinding the counter inside an if-expression is the block-expression shape that miscompiles
integer arithmetic under this elc (BUG-PLAINCHAT-1). Verified in the generated C:
    str_slice(rest, (e + 2), str_len(rest))   <- integer addition, correct
    el_str_concat(head, tail)                 <- string concat, correct
and zero el_str_concat(<ident>, str_len(...)) sites across all 49 modules.

SEAM PROOF (FIX C) rides on the same runs — a real two-search cited answer, inspected byte
by byte, in BOTH failure directions:
  missing separator (the round-7 "to.Good", bytes 77 2e 47): 0 hits. Sentence boundaries
    measure 2e 20 4d — "." SPACE "M".
  over-separation (a cited sentence shattered across paragraphs): 0 hits. The answer is one
    continuous paragraph with its sentences intact, which is the direction a blanket
    separator would have broken.

BUILT: sha256 77115f2733e794c5bc4ad1f55b1acaf658f8f4a91cccd423a2d633d94a726cbc

Refs neuron#109

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:25:48 -05:00
Tim Lingo 9ea41eed78 fix(engine): the model was signing its own answers with our receipt — strip it
FOUND BY E2E, NOT BY REASONING. The previous commit's design note asserted the provenance
receipt "never reaches the user: it is appended to the history copy, not the reply." That
was FALSE, and only running the thing showed it. On the first live run against the built
DMG brain, two agentic turns out of two came back with

  "Your favourite colour is chartreuse and your project is called Perihelion.

   [[RECEIPT - recorded by the soul, not written by the model: no tools ran on this turn.]]"

— the receipt in the user-visible reply.

MECHANISM: the receipt is stored inside the assistant turn, and the agentic path replays
history VERBATIM as Anthropic message objects. So the model sees its own previous answers
ending in [[RECEIPT ...]] and does the obvious thing — it imitates the format and signs the
next answer the same way. The plain path did NOT leak, which is the tell: there, history is
rendered into the SYSTEM prompt as labelled lines rather than replayed as assistant turns,
and a model imitates its own turns far more readily than a transcript.

FIX, two layers, because one of them is not a guarantee:
  - receipt_rule() names the marker in both system prompts (plain and agentic): these lines
    are written by the system, read them as evidence, never write one. Reduces occurrence.
  - receipt_strip() truncates any [[RECEIPT ...]] out of model output before it becomes the
    reply — plain path in layered_generate, agentic path on final_text in agentic_loop.
    Deterministic. A guard that depends on the model choosing to obey is exactly the class of
    thing round 8 exists to stop shipping, so the instruction is the optimisation and the
    strip is the guarantee.
Placed ABOVE agentic_loop's empty-check on purpose: a turn whose entire output was an
imitated receipt has produced no answer, and must be reported as no answer.

The receipt stays in HISTORY, which is the whole point and is proven to work: asked "What
source did you use for that?" one turn after a live web_search, this brain answered
"I used Weather Underground (https://www.wunderground.com/weather/is/reykjav%C3%ADk) for the
current temperature in Reykjavik" — a real source, no apology. That is the false confession
dead, and it is dead BECAUSE the model can read the receipt.

BUILT: 887,112 bytes, sha256 54a2eff84d4fa44f8d2db6781dcf225df4b5075a8ec5f1058018bd40cc1af10b
BUG-PLAINCHAT-1 miscompile guard: zero sites.

Refs neuron#109

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 23:10:53 -05:00
Tim Lingo ff421d39f6 fix(engine): history keeps its provenance and its session — the false confession and the blank stare
DESIGN FIT: three of round 7's five defects share ONE root — the conversation-history
layer persists only {role, content}, discarding tool provenance, session scoping, and the
distinction between a real user turn and an internal utility call. Fixes A and B RESTORE
Will's design rather than extend it: his agentic path already scopes history per session,
the plain path never got it, and his own source carries the TODO admitting the resulting
race (chat.el, handle_chat: "process-global key; concurrent /api/chat requests without
session_id race on this read-append-write"). Fix C repairs one join Will wrote that was
correct for a year and one we added last week. E1/E2 are ours.

FIX A — tool provenance in history (kills the FALSE CONFESSION)
  Root cause, EXECUTED-verified: handle_chat_agentic recorded turns via hist_append, which
  emits {"role","content"} only. server_tool_use blocks, web_search_tool_result blocks and
  every citation were discarded, then replayed as text. On the next turn the model saw a
  data-rich answer with zero evidence a search had happened, and its own permanent rule
  ("never describe a search you did not perform") left one conclusion available: that it
  had fabricated the data. It apologised for a search it HAD run — four independent lines
  of evidence confirm the search was real. The defect is not the model's honesty. It is
  that we deleted the evidence and then asked it to account for itself.
  Change: agentic_loop accumulates the source URLs it already walks past (citations and
  web_search_tool_result content) and returns them as "sources"; handle_chat_agentic folds
  tools_used + sources into a receipt line stored WITH the assistant turn. Receipts are
  unconditional — a negative receipt ("no tools ran") is the other half of the guarantee,
  because "no evidence of a tool" and "evidence of no tool" were previously identical in
  the transcript. conv_history_block splits the receipt off before snipping so a long
  answer cannot truncate away the evidence. The user never sees it: it is appended to the
  history copy, not the reply.

FIX B — one history key for both paths (kills the BLANK STARE)
  Root cause, EXECUTED-verified: the agentic path keyed history on session_hist_<id>; the
  plain path was hard-wired to the process-global conv_history and never read session_id.
  One conversation, two buckets. Proven in the guest engram: the scoped node held exactly
  two turns starting at "Try again" while the earlier exchanges sat unscoped.
  Change: conv_hist_key/conv_hist_label are now the single definition, used by BOTH paths;
  session_id is threaded route -> layered_cycle -> layered_generate / conv_history_record.
  The 2-line fallback (plain path reads the agentic key) was REJECTED: it keeps the
  process-global bucket as a live write target, which is the bleed the TODO describes.
  Also found and closed while threading: layered_cycle read session_id from the state key
  "current_session_id", which is read here and WRITTEN NOWHERE in the entire source. It
  was unconditionally "", so TODO(reliability #4) — per-session steward continuity — was
  dead code that could never fire. It fires now.
  LAZY SESSION, decided explicitly: we create the session EAGERLY at the door (app half,
  ui#223) rather than migrating orphaned turns. Migration would copy the CONTENTS of a
  process-global bucket, possibly another conversation's, into a named session — the bleed,
  performed deliberately. Eager creation makes the situation impossible instead. Migration
  is deliberately not implemented and must not be added without solving provenance first.

FIX C — the two text-join seams ("to.Good", byte-verified 0x77 0x2e 0x47)
  Two bare `+` joins, written a year apart, had drifted into two answers to one question:
  within-response block joins (Will's, 2026-05-03, latent until server-side web_search
  began interleaving non-text blocks) and across-round joins (ours, 62af564).
  Change: one named rule, text_join_sep, at both sites. NOT a blanket separator — a cited
  answer splits MID-SENTENCE ("The current temperature is " + "86°F" + ", with "), so a
  blanket separator shatters every sourced sentence. The rule takes the one bit that
  distinguishes the cases: whether a NON-TEXT block intervened. Hoisting it also makes the
  fix verifiable in the shipped binary, which an inline `+` is not.

FIX E1 — utility generations stay out of the transcript
  Title generation ("Write a 3-6 word title...") and insight passes ran down the same plain
  door as a real message and were recorded as if the user had typed them; the same calls are
  the "model":"unknown" rows in usage.jsonl. is_utility_request reads an explicit utility
  flag from the app, with the __title__/__insight__ id prefixes as a fallback for older
  clients. Answered normally, never recorded.

FIX E2 — OPERATOR IDENTITY is scoped to tool-capable turns
  The block (env USER/HOME, closing "This is a hard rule") was prepended to EVERY system
  prompt including chat mode. On a Tools:Off turn there is no filesystem in reach, so it
  governed nothing and merely supplied the loudest fact in the prompt — which is why the
  model opened a fresh conversation with "You're test, on your machine at /Users/test".
  Hoisted to operator_identity_block() and gated on !chat_mode. Unchanged wherever a file
  or command tool can actually be reached.

ALSO: agentic_loop's per-session history persist had a second hand-rolled copy of
conv_history_persist with a different label expression, different salience scores and
different tags for the same node. Since both now derive the label from conv_hist_label and
engram_node_full upserts by label, two score policies were writing one node. Collapsed to
one writer.

BUILD NOTE: dist/elp-c-decls.h is force-included by the documented link recipe and carried
the OLD C arities, so it is updated here. This is the build-support header, NOT the stale
generated dist/soul.c — no dist/*.c was read or edited; all engine changes are .el source.
chat.elh/soul.elh are committed because a first-pass build against the old signatures FAILS
(measured); the other regenerated headers are reverted as unrelated churn.

BUILT: 887,000 bytes, sha256 d632b061ad75269d6adeb52578d030eaf49e895d91289d7f946b19c08450d728
Zero el_str_concat(<int>, str_len(...)) sites (the BUG-PLAINCHAT-1 miscompile guard).
web_search_20250305 and disable_parallel_tool_use both still present — PR #108's web search
and the ADR-0005 stopgap are intact.

Refs neuron#109 (builds on it), neuron#78 (Receipt Contract — the real fix A is a stopgap for)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 22:59:54 -05:00
Tim Lingo 635f6febe4 feat(engine): plain chat generates at L3 — inside the safety cycle, not around it
Neuron Soul CI / build (pull_request) Failing after 12m8s
Neuron Soul CI / deploy (pull_request) Has been skipped
Non-agentic /api/chat (the desktop app's default "Tools: Off" mode) returned the
user's own screened text as a bare non-JSON string. Every JSON client failed to
parse it and showed "Couldn't reach Neuron - it may be offline."

Root cause: f52d5bd (2026-06-11) correctly moved the route onto the layer spine
(handle_chat -> layered_cycle), but L3 never got a generator — imprint_respond()
annotates its input and returns it. Two pieces of the architecture were already
waiting for that step: layered_cycle parks a bell directive in the state key
build_system_prompt is written to consume, and build_system_prompt carries a
chat_mode ("no tools") flag with no live caller.

The fix composes rather than replaces. Wiring handle_chat would have removed
safety_screen, the hard-bell short-circuit, the whole stewardship layer and
safety_validate — the only enforcing output gate in the codebase — in exchange
for a working reply (see _engine-websearch-20260804/SAFETY-STOP.md). Instead
layered_cycle keeps every gate, in order, and gains a generation step between
imprint_respond and safety_validate.

  L1 screen -> guard -> hard-bell short-circuit -> L2a -> L2b -> L2c
    -> L3 imprint_respond (prompt) -> L3b layered_generate (NEW) -> L1 validate

- chat.el:  NEW layered_generate (L3 generation, no tools offered),
            conv_history_block, conv_history_record.
            FIX build_system_prompt never concatenated no_tools_rule into its
            return — the "[NO TOOLS THIS TURN]" rule reached no model at all.
            handle_chat annotated DO-NOT-WIRE with the reason.
- soul.el:  layered_cycle gains L3b + post-validation turn bookkeeping.
- routes.el: NEW plain_chat_envelope; all three /api/chat dispatch sites wrap the
            cycle's output. Built OUTSIDE the cycle so safety_validate always sees
            raw model text — nothing to unwrap or rebuild on the crisis path.
            Emits both `reply` and `response`: the desktop app reads `reply`,
            the CLI tools and telegram-gateway read `response`.

Also fixes BUG-PLAINCHAT-1, a pre-existing CRITICAL crash on the crisis path.
elc compiles `let n: Int = pos + str_len(marker)` to el_str_concat() — string
concat on two integers — inside a block-expression initializer, segfaulting the
daemon (SIGSEGV in strlen). Six inline copies of the same " | ts:" parser had it:
two in layered_cycle L2c, two in engram_compile (live on the AGENTIC path too),
two in affective_context_prefix. A distress turn following an earlier affective
turn killed the whole process. Proven pre-existing: an unmodified baseline binary
crashes identically, and the same bad C is in the committed dist/soul.c. Fixed by
hoisting to one top-level function, affective_node_ts(), where the expression
compiles to integer addition — verified in the generated C.

Proof (throwaway HOME/engram, explicit NEURON_PORT, live chain untouched):
- Plain turn returns a JSON envelope with the provider's answer, not an echo.
- Captured request body: no `tools`, no `tool_choice`; system prompt carries the
  NO-TOOLS rule. Tools:Off means no tool is offered, structurally.
- Hard bell: canned 988 message, and the provider request count does not move —
  the message never reaches a model.
- Soft bell + a 2-char model reply: safety_validate's care phrase is appended to
  the MODEL's output. Output gate acting, on this route.
- The L1 bell directive now reaches the model here for the first time (the state
  addendum had a producer and no consumer).
- test_layered_cycle PASS; all six El suites byte-identical to baseline.
- verify-soul-contract.sh (bash 5.3): GATE PASS, 27/27, immutability PASS.
- The crash sequence that killed the baseline daemon now returns HTTP 200.

Not proven: no live Anthropic call — the login keychain refuses the key to a
non-interactive process (rc=24 errSecInteractionNotAllowed). Details and the
one-command close-out are in _engine-plainchat-20260805/README.md §7.

Builds on PR #108. dist/soul.c deliberately not regenerated — Will's toolchain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 09:11:03 -05:00
Tim Lingo 62af5649fe feat(engine): port Anthropic server-side web_search into the agentic loop
Neuron Soul CI / build (pull_request) Failing after 10m45s
Neuron Soul CI / deploy (pull_request) Has been skipped
Re-authors soul-webfix-20260711.patch in El (the patch is a diff against
generated C at month-old offsets, so nothing was applied as a patch). Its
last two hunks — an unrelated /api/safety-contact implementation — were
deliberately not ported; that route already exists and is safety-critical.

Activation restores existing design, it does not invent a mechanism:
commit 8eea1d9 (2026-06-09, Tim-approved) made native web_search built-in
with no user-facing toggle, and tests/test_agentic_tools.el section 2 still
asserts agentic_tools_all() contains it — an assertion main currently fails.
The call site was lost when agentic_tools_all() (connector tools, PR #19)
replaced agentic_tools_with_web(). Attaching in agentic_tools_all() covers
handle_chat_agentic, handle_dharma_room_turn_agentic and agentic_resume.

pause_turn handling is included and was genuinely missing: the shipped
binary has zero occurrences of it. Without it a paused server-side search
returns only the text written so far and the loop treats it as final —
a silently truncated answer. final_text now accumulates across resume
cycles rather than overwriting (overwriting would discard everything
written before the pause).

Default tool version is web_search_20250305, NOT the newer _20260209, and
that is a measured choice: _20260209's dynamic filtering uses server-side
programmatic tool calling, which the API refuses to combine with ADR 0005's
stopgap —

  HTTP 400 invalid_request_error
  tool_choice.disable_parallel_tool_use: true cannot be used with
  programmatic tool calling

Dropping the stopgap would resurrect neuron#78 bug b (killed runs 3/3 in
ADR 0005's own A/B). The basic variant is compatible with the stopgap and
returns real results, so neither feature is dropped. Version lives in state
key web_search_tool_version; flipping it once the stopgap retires is a
config write, no recompile. The fallback also fires on "programmatic tool
calling" so a premature flip self-heals loudly instead of dying.

Also fixes a real bug this port exposed: json_get is a first-match scanner
and a cited text block serialises citations FIRST, so json_get(block,"type")
returned the nested citation's type and every citation-bearing block — the
ones carrying the searched facts — was silently dropped from the reply.
Reply length on the same question: 79 -> 360 chars.

Other loop changes: server_tool_use accounting into tools_used, iteration
cap 8->12 for pause/resume cycles, max_tokens 4096->16384, container-id
carry-forward, API error head logged instead of swallowed, tools_used gated
on is_tool_turn so a truncated tool block is not reported as work done.

dist/soul.c is deliberately NOT regenerated — the local elc predates Will's
last regen (e610a41) and regenerating with an older compiler risks unrelated
codegen drift. Source only; regen is Will's.

E2E-VERIFIED on a sandbox soul (scratch HOME/engram, dead axon+ISE, explicit
NEURON_PORT): live Bentonville weather with tools_used ["web_search"]; the
27-route contract gate PASSes; the parallel-tool stopgap still completes a
3-file mission with no 400; a 3-search 3,486-char answer kept its end
sentinel. Honest gap: pause_turn is compiled in but not exercised by a live
pause (max_uses:5 caps the server loop below the pause threshold).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 22:57:59 -05:00
Tim Lingo 710761e2d5 fix(engine): send tool_choice.disable_parallel_tool_use on agentic loop (STOPGAP)
The agentic loop keeps only the FIRST tool_use block per round (chat.el:2281,
"Capture first tool_use block only"). Anthropic lets a model emit several
tool_use blocks in one message and requires a tool_result for every one, so a
parallel-tool turn is answered once, the rest are dropped, and the next request
dies with:

  tool_use ids were found without tool_result blocks immediately after

(neuron#78 quotes this as "tool_use ids found without tool_result"; the above is
the API's actual wording - recorded so the next person's grep matches.)

This constrains the wire to match what the loop can assemble:
  "tool_choice":{"type":"auto","disable_parallel_tool_use":true}

STOPGAP - AND THE DURABLE FIX ALREADY EXISTS. A correct multi-tool loop is
already in Will's EL runtime, in C, and the soul does not call it. Verified on
el:origin/main lang/el-compiler/runtime/el_runtime.c: llm_register_tool:9616,
llm_build_tool_results:9743 - which walks EVERY content block, emits one
tool_result per tool_use, and sets is_error for an unregistered tool -
llm_call_agentic:9817 calling it at :9918, iteration cap 10 at :9847. Will's
commit 12d5e77 (2026-04-30). grep for llm_call_agentic/llm_register_tool across
every neuron/*.el returns nothing; dist/soul.c has zero references. chat.el
hand-rolls its own single-tool loop instead, and that is the one that breaks.

The durable fix is therefore to register the soul's tools via llm_register_tool
and call llm_call_agentic - deleting a loop, not writing one. See ADR 0005.

Our own approved spec called this seven weeks ago:
docs/research/agentic-tool-approval-design.md (2026-06-12, "Approved for build"),
line 20 on the defect, line 30 on the goal ("Execute all tool_use blocks in a
turn (one result per block)").

Two edits, because dist/soul.c cannot be regenerated here (Will's gated elc/elb
toolchain is not on this machine):
  (a) chat.el:2255 - source of truth, so a later regen carries the fix.
      One edit covers all three routes: agentic_loop is called from chat.el:2152
      (/api/chat agentic), :2676 (dharma room) and :2496 (agentic_resume).
  (b) dist/soul.c:28173 - generated form, hand-spliced. Line 27624 is the
      non-agentic/OpenAI-compat req_body (no tools) and was left untouched.
Prior art reused rather than reinvented: soul-narrated-runs-20260713.patch
(27,824 bytes) line 78 spliced this same string into the same concat chain on
2026-07-13.

Deliberately NOT ported from that patch, having read it: max_tokens is not
changed by it (16384 sits on both sides of the hunk; our main's 4096 is a
separate output-truncation concern), and its pause_turn pairing fix - same
defect class - is unreachable today because no server-side web_search is wired
(agentic_tools_with_web at chat.el:1418 is never called), so it is untestable
and logged instead.

Proven E2E on a scratch profile and port 7791, never the live chain. A/B against
a pristine origin/main control built from the same vendored runtime: fixed
completed the mission (tools_used read_file x3, 4 iterations, correct answer);
control failed 3/3. Direct API probe confirmed the mechanism - without the field
the model emits 3 parallel tool_use blocks and replaying the unfixed loop's next
turn returns HTTP 400; with it, exactly 1 block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 17:07:04 -05:00
will.anderson 74520b8333 Merge pull request 'ci: pin + complete vendored el-runtime so reconciled soul.c links' (#105) from ci/pin-vendored-runtime into main
Neuron Soul CI / build (push) Failing after 14m36s
Neuron Soul CI / deploy (push) Has been skipped
2026-08-03 16:08:51 +00:00
will.anderson 7f3d6ed8cd ci: update vendored el-runtime to complete v1.0.0-20260501
Neuron Soul CI / build (pull_request) Has been cancelled
Neuron Soul CI / deploy (pull_request) Has been cancelled
The runtime vendored alongside the CI pin was the Jul-21 snapshot, which
predates two builtins the reconciled ship-soul now calls:
  - http_delete_json  (boot-counter HTTP write-back, awareness/memory self-review)
  - engram_act_stats_json  (heartbeat activation observability)
Compiling dist/soul.c against the stale runtime fails with implicit-declaration
errors. Vendor the current release runtime (identical to the one the soul was
gate-verified against: verify-soul-contract PASS, genesis boots clean, full
safety-contact) so the CI Linux soul is byte-for-byte the verified soul.
2026-08-03 11:07:38 -05:00
will.anderson eed6487114 ci: pin soul build to vendored release runtime v1.0.0-20260501
The soul build downloaded el-runtime-c 'latest' from Artifact Registry. The
merged ship-soul calls engram_prune_telemetry, which the latest published
runtime no longer defines, so an unpinned build fails to link — the failure
mode that let a broken/handlerless soul reach prod.

Vendor the release runtime v1.0.0-20260501 (el_runtime.c/.h) into the repo and
compile the soul against it. This is the exact runtime the merged soul was
verified against (verify-soul-contract GATE PASS, genesis boot survives, full
safety-contact response), making the build reproducible and independent of a
moving AR 'latest'.

The verify-soul-contract.sh HARD-BLOCK gate already runs before Publish (from
the CI-hardening arc on main), so a destructive or stale soul can never
publish/deploy again.
2026-08-03 11:06:26 -05:00
will.anderson 2c2aaa0653 Merge pull request 'Reconcile: main = union of all ship-critical soul fixes (beta-gating)' (#104) from reconcile/soul-union-main into main
Neuron Soul CI / build (push) Has been cancelled
Neuron Soul CI / deploy (push) Has been cancelled
2026-08-03 16:03:46 +00:00
93 changed files with 59395 additions and 991 deletions
+23 -33
View File
@@ -39,7 +39,7 @@ jobs:
> /etc/apt/sources.list.d/google-cloud-sdk.list
apt-get update -qq && apt-get install -y google-cloud-cli
- name: Download El runtime from Artifact Registry
- name: Authenticate to GCP + stage PINNED El runtime
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
@@ -47,41 +47,31 @@ jobs:
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
# PINNED RUNTIME — do NOT pull "latest" from Artifact Registry.
# The ship-soul calls engram_prune_telemetry (awareness.el sync/heartbeat
# self-review). The latest published el-runtime-c no longer defines that
# symbol, so an unpinned build fails to LINK — which is exactly how a
# broken/handlerless soul reached prod before. Compile against the
# vendored release runtime v1.0.0-20260501: the exact runtime the merged
# ship-soul was verified against (verify-soul-contract GATE PASS +
# genesis boot survives + full safety-contact response). It is committed
# under vendor/ so the soul build is fully reproducible and never depends
# on a moving AR "latest".
rm -rf /opt/el/runtime
mkdir -p /opt/el/runtime
cp vendor/el-runtime/v1.0.0-20260501/el_runtime.c /opt/el/runtime/el_runtime.c
cp vendor/el-runtime/v1.0.0-20260501/el_runtime.h /opt/el/runtime/el_runtime.h
echo "El runtime PINNED to v1.0.0-20260501: $(ls /opt/el/runtime/)"
# Get latest version of each runtime package (elc/elb not needed — we compile
# dist/soul.c directly; running elb on Linux OOM-kills the runner, and we
# always use the repo's pre-built soul.c anyway).
get_latest() {
gcloud artifacts versions list \
--repository=foundation-prod \
--location=us-central1 \
--project=neuron-785695 \
--package="$1" \
--sort-by="~createTime" \
--limit=1 \
--format="value(name)" 2>/dev/null | awk -F/ '{print $NF}'
}
RC_VER=$(get_latest el-runtime-c)
RH_VER=$(get_latest el-runtime-h)
echo "Downloading runtime@${RC_VER}"
gcloud artifacts generic download \
--repository=foundation-prod --location=us-central1 --project=neuron-785695 \
--package=el-runtime-c --version="${RC_VER}" \
--destination=/opt/el/runtime/
gcloud artifacts generic download \
--repository=foundation-prod --location=us-central1 --project=neuron-785695 \
--package=el-runtime-h --version="${RH_VER}" \
--destination=/opt/el/runtime/
mv /opt/el/runtime/el_runtime.c* /opt/el/runtime/el_runtime.c 2>/dev/null || true
mv /opt/el/runtime/el_runtime.h* /opt/el/runtime/el_runtime.h 2>/dev/null || true
echo "El runtime ready: $(ls /opt/el/runtime/)"
# neuron#133: CI compiles dist/soul.c, NOT the .el sources. On 2026-08-07 a
# build off main would have shipped an engine with none of five merged fixes,
# including a P0 safety fix, while main's source read as correct. The runner
# cannot regenerate the amalgam (elc needs 24GB+ virtual memory), but it can
# refuse to compile a stale one. Fails loudly with the recipe in the message.
- name: Verify dist/soul.c matches the sources
run: |
chmod +x tools/soulc-stamp.sh
./tools/soulc-stamp.sh --check
- name: Build neuron soul binary
run: |
+139
View File
@@ -0,0 +1,139 @@
# PORT-NOTES — openai tools port working state (2026-08-06, session handoff-safe)
Spec: `docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md` (Tim-approved 2026-08-06). Tasks #1-5
tracked in-session (1 ✓ wiring verdict, 2 ✓ stub rig, 3 in-progress = THIS, 4-5 pending).
Worktree: HERE (`_wt-openai-tools`, branch `feat/soul-openai-tools-v2` @ dba755d). Round-9 trees
READ-ONLY. Nothing committed yet.
## Step-0 verdict (evidence in journal note ncli-653ba964dd76)
Shipped app never wires the v1 lane: launcher exports `SOUL_LLM_MODEL/PROVIDER/BASE_URL` +
`ANTHROPIC_API_KEY`+`SOUL_API_KEY` (= Keychain key for WHATEVER provider; installer/macos/
neuron-daemons.sh:288-300 on hotfix/beta-round9); brain reads only SOUL_LLM_MODEL (chat.el:8) and
NEURON_LLM_0_* (chat.el:1768-1794) which nothing sets. `/api/config` PATCH ignores llm_* fields
(studio.el:36 handle_config: POST-only, reads model/provider/api_key only).
**Bridge = brain-side ONLY (zero app-repo edits, zero round-9 collision):**
- `llm_base_url()`: NEURON_LLM_0_URL → fallback SOUL_LLM_BASE_URL when SOUL_LLM_PROVIDER ∉ {"","anthropic"}
- `llm_wire_format()`: NEURON_LLM_0_FORMAT → fallback derive from SOUL_LLM_PROVIDER (openai/grok/gemini/groq/ollama → "openai"; else "anthropic")
- `agentic_api_key()`: already works (ANTHROPIC_API_KEY carries the provider key); add NEURON_LLM_0_KEY → SOUL_API_KEY fallback.
## Design pins (stub asserts these — stub is green 58/58, tests/gate-openai/)
- Request MUST send `"tool_choice":"auto"` (string) + `"parallel_tool_calls":false` explicitly.
- `arguments` in tool_calls = JSON-ENCODED STRING; decode ONCE via json_get → feed dispatch_tool
verbatim. Stub's echo-mismatch check catches double-encode/decode (two-escaper trap).
- Assistant echo turn: `{"role":"assistant","content":null,"tool_calls":[...]}` VERBATIM from response.
- Feedback: `{"role":"tool","tool_call_id":"<id>","content":"<result string>"}`.
- Resume must NOT re-answer an answered id (stub 400s on repeat tool_call_id).
- Parallel tool_calls in a response: take FIRST only + log skip (mirror ADR-0005 stopgap); stub
scenario `parallel` proves behavior.
- No tools in request when tools array empty/absent turns (boot probes) — stub defaults tolerate.
## el idioms confirmed (from openai_chat_complete :1808-1854 + agentic_loop :2751-2838)
- JSON: `json_get(s,k)` decoded string · `json_get_raw(s,k)` raw subtree · `json_array_len` ·
`json_array_get(arr,i)` · build by string concat + `json_escape()` (:1797, OpenAI-lane escaper).
- HTTP: `let h: Map = {}` + `map_set(h,k,v)` + `http_post_with_headers(url, body, h)`;
Bearer auth via `Authorization` header when key non-empty (:1825-1830).
- Loop-carried vars must be top-level locals in the fn, mutated as if-expressions at while-body
top level (see :2760-2791 pattern + comment :2903-2904 region).
- Error shape: `str_starts_with(raw,"{\"error\"") || str_contains(raw,"\"error\":")` → return
`{"error":"llm unavailable","reply":""}` (:1835-1838).
## Remaining read map (before writing the fork)
- chat.el 2840-3200: block walk (2923-3000), policy gate (3009-3023: classify_tool_risk /
is_builtin_tool / ask_all / tool_auto_approved → needs_bridge), dispatch_tool call (3025),
tool_result feedback (3031, 3067-3072), run-progress ledger append (3078-3087), bridge_save
(3182), loop end + done envelope (~3100-3200).
- agentic_resume 3227-3293 (hardcoded Anthropic headers to make wire-aware; blob gets `wire` field,
legacy default anthropic) · handle_tool_result 3293+ · dharma fork site 3465 (calls agentic_loop
direct, no use_openai check today).
## Write plan (order)
1. Env fallbacks (edit llm_base_url/llm_wire_format/agentic_api_key) — small, first, testable alone.
2. `openai_tools_json(anthropic_tools: String) -> String` converter (walk array; per entry build
{"type":"function","function":{name,description,parameters:input_schema-raw}}).
3. `openai_agentic_loop(...)` fork: same signature as agentic_loop minus Anthropic-only params;
INCLUDE run-progress ledger + tools_log + iteration cap 12; NO container_id/ws_drift/web_search
(out of scope; strip web_search entry from tools via agentic_tools_literal()+connector merge,
NOT _with_web()).
4. Fork sites ×3: handle_chat_agentic :2695-2700 (route agentic to new loop when use_openai);
dharma :3465; agentic_resume wire-branch.
5. `chat.elh` extern decls. 6. Compile (recipe: dist/ + elc/elb per neuron-soul-build-deploy memory;
round-9 tree soul.c regen'd 08-06 proves toolchain live). 7. Gate: stub selftest recipe in
tests/gate-openai/README.md. 8. Anthropic-lane regression via gate9 (READ-ONLY consume from
_wt-beta-round9). 9. Live Groq E2E (scratch profile, free port, key via Keychain read-only).
## BUILD RECIPE — CORRECTED 2026-08-06 (the June memory is STALE for August code)
`~/el-sdk/el_runtime.c` (Jun 15) is MISSING builtins the Aug engine calls (`engram_wm_count`,
`engram_wm_top_json`, `http_delete_json`, `http_serve_async`) → link fails with
"symbol(s) not found for architecture arm64". Use the REPO-PINNED runtime:
```
mkdir -p <scratch>
elb --elc=$HOME/el-sdk/elc --runtime=vendor/el-runtime/v1.0.0-20260501 --out=<scratch>/
# "elb: link failed" at the end is EXPECTED and harmless — the per-module .c files are produced
cc -std=c11 -O1 -DHAVE_CURL -rdynamic \
-I vendor/el-runtime/v1.0.0-20260501 -I <scratch> -I /opt/homebrew/opt/openssl@3/include \
-L /opt/homebrew/opt/openssl@3/lib \
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
-o <scratch>/soul <scratch>/*.c vendor/el-runtime/v1.0.0-20260501/el_runtime.c \
-lssl -lcrypto -lcurl -lpthread -lm
```
Source: `_engine-plainchat-20260805/README.md:396-412`. Verified today: 0 errors, 887,296 B.
`elb` ALSO rewrites every `*.elh` in the tree (cosmetic em-dash→hyphen in the auto-gen banner,
plus true-ups) and drops a stray `soul..elh``git restore` the unrelated ones and delete the
stray before staging, or the diff drowns in noise.
## SELF-REVIEW FIX LIST (found by reading my own diff, 2026-08-06 — apply in ONE batch, then rebuild once)
- **F3 (CORRECTNESS, do first):** the assistant echo currently replays the provider's FULL
`tool_calls` array (`tc_arr`) while the loop answers only the FIRST call. If a provider ignores
`parallel_tool_calls:false`, the next request carries an assistant turn with N tool_calls and
only ONE `role:"tool"` response → most OpenAI-format providers 400 ("missing tool response for
id X") and the run dies. This is the same class as ADR-0005's Anthropic failure, but here it is
cheap to close: echo ONLY the honored call (`"[" + tc0 + "]"`), so the conversation we send is
self-consistent and the dropped call never existed from the model's view. The DRIFT log line
stays (honest accounting of what we dropped).
- **F4 (efficiency/latency):** `handle_chat_agentic` computes `agentic_tools_all()` at ~:2681
BEFORE the fork, then the OpenAI branch computes `agentic_tools_no_web()` again — two
`connector_tools_json()` calls per turn, each an HTTP round-trip to the connector bridge on
:7771 (two timeout exposures). Fix: compute the tools array ONCE, per lane, after `use_openai`
is known (check no other use of `tools_json` sits between :2681 and the fork before moving it).
Note: `openai_tools_json()` already skips any entry with no `input_schema`, so Anthropic's
server-side `web_search` entry is auto-dropped even if the full array is passed —
`agentic_tools_no_web()` is kept for EXPLICITNESS, not necessity.
- **F1 (debuggability):** the "no choices in response" branch logs a generic string and discards
the body. Log the response head (as the `is_error` branch does) — a provider that returns 200
with an unexpected shape is otherwise undiagnosable from the log.
- **OPEN QUESTION (evidence pending from the gate):** the tool-result feedback turn escapes with
`json_escape()` (this lane's escaper) rather than `json_safe()` (used everywhere else). The
Anthropic lane escapes that field with NEITHER, which is a latent defect on that side. If the
torture scenario shows any escaping loss, switch to `json_safe` and note the Anthropic-side
finding for Will.
## TEST HARNESS — built 2026-08-06 (Task 4 side-work, reusable by anyone)
- `tests/run-el-test.sh <tests/test_x.el> | --all` — the engine tests were NEVER runnable
before this (`elc` is a compiler: emits C to stdout and exits). It emits the test to C,
compiles `soul.c` separately with `main` renamed away (soul.c owns the daemon's real main
but also defines `layered_cycle` et al.), links the remaining modules + the repo-pinned
runtime, and executes. Modules cached under `/tmp/el-test-<worktree>/`; `REBUILD=1` forces.
- **The runner computes the verdict itself** because the test FILES cannot: all 9 counted
test files do `let pass_count = pass_count + 1` inside an if BLOCK, which El scoping
discards, so every summary line reads `0 passed, 0 failed` forever. Per-assertion
`PASS:`/`FAIL:` lines ARE reliable; the runner counts those, exits non-zero on any FAIL
or on zero assertions, and was proven to discriminate with a negative control (broken
assertion → 31 passed / 1 failed / exit 1). Real in-file fix filed: **neuron#116**.
- `tests/test_bridge_serialization.el`: 4 `bridge_save` calls updated for the new `wire`
argument, plus **Section 9** (8 new assertions) covering wire round-trip both ways, the
legacy no-wire blob (resumes as anthropic), and a FIELD-ORDER decoy guard — a fake
`"wire":"anthropic"` planted inside `messages_raw` must not beat the blob's own scalar.
That decoy is the round-9 first-match-scanner bug class, now pinned by a test. **32/32 green.**
## MEMORY-SAVE CAVEAT RESOLVED 2026-08-06
Earlier saves this session reported `-> OUTBOX only (real mind unreachable or read-back
failed)`. That was a **read-back verifier false negative, not data loss** — a direct
`POST :7770/api/neuron/recall` returns those notes from the live mind verbatim. Another
terminal was fixing exactly this (multi-word read-back probe) the same afternoon. Do NOT
re-save on an OUTBOX report without first querying the mind directly, or you duplicate nodes.
## Standing cautions
- PERSIST OFF on the real mind this boot (neuron#98/#92): journal saves only, ferry later. MCP link
down this terminal; use neuron_remember.py / neuron_recall.py.
- Aug-16: Groq retires llama-3.3-70b-versatile (separate P0, Tim's call, catalog swap).
- Never bind 7770/7779/17779; never touch ~/.neuron; round-9 worktrees read-only.
+19
View File
@@ -889,10 +889,29 @@ fn awareness_run() -> Void {
state_set("soul.last_beat_ts", int_to_str(now_ts))
// Persist in-process Engram (sessions, memories, conversation nodes)
// to local snapshot so they survive restarts.
// FILE MODE ONLY: "soul_snapshot_path" is set exclusively in the
// genesis+safe_to_seed branch of soul.el, and safe_to_seed is
// unconditionally false when ENGRAM_URL is set. In HTTP mode the
// owner persists; the soul must not (soul.el:571-573).
let snap_path: String = state_get("soul_snapshot_path")
if !str_eq(snap_path, "") {
mem_save(snap_path)
}
// WRITE-THROUGH RETRY (neuron#117). The HTTP-mode counterpart of the
// save above: hand anything still spooled to the persistence owner.
//
// This is the retry arm of the whole design. Deltas that could not be
// pushed owner down, owner restarting, transient refusal stay on
// disk and are re-offered here every heartbeat until they land. It is
// also the catch-all for writes made by the awareness loop itself,
// which never passes through the HTTP handler's flush point.
//
// No-op with no HTTP call when the spool is empty or ENGRAM_URL is
// unset, so an idle soul in file mode pays nothing for this.
let wt_pushed: Int = wt_drain()
if wt_pushed < 0 {
ise_post("{\"event\":\"write_through_backlog\",\"ts\":" + int_to_str(now_ts) + "}")
}
}
// Curiosity scan: idle-gated AND wall-clock based. Only fires when the
+1437 -169
View File
File diff suppressed because it is too large Load Diff
+32 -4
View File
@@ -1,4 +1,4 @@
// auto-generated by elc --emit-header do not edit
// auto-generated by elc --emit-header - do not edit
extern fn chat_default_model() -> String
extern fn engram_numeric_valid(s: String) -> Bool
extern fn parse_float_x100(s: String) -> Int
@@ -16,18 +16,35 @@ extern fn engram_nodes_merge(a: String, b: String) -> String
extern fn id_in_seen(node_id: String, seen: String) -> Bool
extern fn add_to_seen(seen: String, node_id: String) -> String
extern fn engram_extract_ids(nodes_json: String) -> String
extern fn affective_node_ts(node_json: String) -> Int
extern fn engram_compile(intent: String) -> String
extern fn distill_transcript(transcript: String) -> String
extern fn json_safe(s: String) -> String
extern fn current_engine_note(model: String) -> String
extern fn bounded_persona_floor() -> String
extern fn operator_identity_block() -> String
extern fn build_system_prompt(ctx: String, chat_mode: Bool) -> String
extern fn hist_append(hist: String, role: String, content: String) -> String
extern fn conv_hist_key(session_id: String) -> String
extern fn conv_hist_label(session_id: String) -> String
extern fn is_utility_request(body: String, session_id: String) -> Bool
extern fn provenance_scan_urls(arr: String, acc: String) -> String
extern fn provenance_add_sources(block: String, btype: String, has_cit: Bool, cit_raw: String, acc: String) -> String
extern fn provenance_names(tools_used: String) -> String
extern fn text_join_sep(accumulated: String, incoming: String, after_interruption: Bool) -> String
extern fn receipt_rule() -> String
extern fn receipt_strip(s: String) -> String
extern fn tool_receipt(tools_used: String, sources: String) -> String
extern fn hist_trim(hist: String) -> String
extern fn hist_trim_with_bell_guard(hist: String) -> String
extern fn clean_llm_response(s: String) -> String
extern fn conv_history_persist(hist: String) -> Void
extern fn conv_history_load() -> String
extern fn conv_history_persist(session_id: String, hist: String) -> Void
extern fn conv_history_load(session_id: String) -> String
extern fn conv_history_record(session_id: String, user_msg: String, assistant_msg: String, receipt: String) -> Void
extern fn conv_history_block(session_id: String) -> String
extern fn layered_generate(prompt: String, imprint_id: String, session_id: String) -> String
extern fn session_preload_bullets(nodes: String, max_bullets: Int, snip_len: Int) -> String
extern fn affective_context_prefix() -> String
extern fn handle_chat(body: String) -> String
extern fn handle_see(body: String) -> String
extern fn studio_tools_json() -> String
@@ -36,7 +53,14 @@ extern fn llm_base_url() -> String
extern fn llm_wire_format() -> String
extern fn json_escape(s: String) -> String
extern fn openai_chat_complete(model: String, base_url: String, api_key: String, safe_sys: String, messages_json: String) -> String
extern fn openai_tools_json(tools_anthropic: String) -> String
extern fn utf8_safe_slice(s: String, n: Int) -> String
extern fn json_trim_dangling_escape(s: String) -> String
extern fn agentic_tools_no_web() -> String
extern fn openai_agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, tools_log_in: String) -> String
extern fn agentic_tools_literal() -> String
extern fn web_search_tool_json() -> String
extern fn strip_client_web_search(tools_inner: String) -> String
extern fn agentic_tools_with_web() -> String
extern fn connector_tools_json() -> String
extern fn agentic_tools_all() -> String
@@ -46,13 +70,17 @@ extern fn call_neuron_mcp(tool_name: String, args: String) -> String
extern fn agent_workspace_root() -> String
extern fn path_within_root(path: String, root: String) -> Bool
extern fn resolve_in_root(path: String, root: String) -> String
extern fn run_command_is_readonly(cmd: String) -> Bool
extern fn cmd_abs_escape_at(cmd: String, root: String, needle: String) -> Bool
extern fn run_command_guard(cmd: String, root: String) -> String
extern fn classify_tool_risk(tool_name: String, tool_input: String) -> String
extern fn dispatch_tool(tool_name: String, tool_input: String) -> String
extern fn is_builtin_tool(tool_name: String) -> Bool
extern fn next_bridge_id() -> String
extern fn handle_chat_plan(body: String) -> String
extern fn handle_chat_agentic(body: String) -> String
extern fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, h: Map, tools_log_in: String) -> String
extern fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String) -> Bool
extern fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String, wire: String) -> Bool
extern fn agentic_resume(session_id: String, tool_use_id: String, content: String) -> String
extern fn handle_tool_result(session_id: String, body: String) -> String
extern fn handle_chat_as_soul(body: String) -> String
Generated Vendored
+23 -4
View File
@@ -5,6 +5,15 @@ el_val_t add_punct(el_val_t s, el_val_t intent);
el_val_t add_to_seen(el_val_t seen, el_val_t node_id);
el_val_t aff_try_slot(el_val_t slot_json, el_val_t aff_7d_ts, el_val_t acc_key);
el_val_t affective_context_prefix(void);
el_val_t is_utility_request(el_val_t body, el_val_t session_id);
el_val_t operator_identity_block(void);
el_val_t provenance_add_sources(el_val_t block, el_val_t btype, el_val_t has_cit, el_val_t cit_raw, el_val_t acc);
el_val_t provenance_names(el_val_t tools_used);
el_val_t provenance_scan_urls(el_val_t arr, el_val_t acc);
el_val_t text_join_sep(el_val_t accumulated, el_val_t incoming, el_val_t after_interruption);
el_val_t receipt_rule(void);
el_val_t receipt_strip(el_val_t s);
el_val_t tool_receipt(el_val_t tools_used, el_val_t sources);
el_val_t agent_number(el_val_t agent);
el_val_t agent_person(el_val_t agent);
el_val_t agent_workspace_root(void);
@@ -132,7 +141,7 @@ el_val_t awareness_run(void);
el_val_t axon_get(el_val_t path);
el_val_t axon_post(el_val_t path, el_val_t body);
el_val_t bounded_persona_floor(void);
el_val_t bridge_save(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages, el_val_t tools_log, el_val_t tool_use_id);
el_val_t bridge_save(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages, el_val_t tools_log, el_val_t tool_use_id, el_val_t wire);
el_val_t build_form_from_json(el_val_t semantic_form_json, el_val_t lang_code);
el_val_t build_np(el_val_t referent, el_val_t slots);
el_val_t build_pp(el_val_t loc);
@@ -151,8 +160,12 @@ el_val_t cmd_abs_escape_at(el_val_t cmd, el_val_t root, el_val_t needle);
el_val_t connectd_get(el_val_t suffix);
el_val_t connectd_post(el_val_t suffix, el_val_t body);
el_val_t connector_tools_json(void);
el_val_t conv_history_load(void);
el_val_t conv_history_persist(el_val_t hist);
el_val_t conv_hist_key(el_val_t session_id);
el_val_t conv_hist_label(el_val_t session_id);
el_val_t conv_history_block(el_val_t session_id);
el_val_t conv_history_load(el_val_t session_id);
el_val_t conv_history_persist(el_val_t session_id, el_val_t hist);
el_val_t conv_history_record(el_val_t session_id, el_val_t user_msg, el_val_t assistant_msg, el_val_t receipt);
el_val_t cop_article(el_val_t gender, el_val_t number, el_val_t definite);
el_val_t cop_bwk_future(el_val_t prefix);
el_val_t cop_bwk_perfect(el_val_t prefix);
@@ -782,7 +795,8 @@ el_val_t lang_profile_uga(void);
el_val_t lang_profile_zh(void);
el_val_t lang_profile(el_val_t code, el_val_t word_order, el_val_t morph_type, el_val_t has_case, el_val_t has_gender, el_val_t script_dir, el_val_t agreement, el_val_t null_subject);
el_val_t lang_word_order(el_val_t profile);
el_val_t layered_cycle(el_val_t raw_input);
el_val_t layered_cycle(el_val_t raw_input, el_val_t session_id, el_val_t utility);
el_val_t layered_generate(el_val_t prompt, el_val_t imprint_id, el_val_t session_id);
el_val_t lex_class(el_val_t entry);
el_val_t lex_form(el_val_t entry, el_val_t idx);
el_val_t lex_pos(el_val_t entry);
@@ -861,6 +875,11 @@ el_val_t non_weak_past(el_val_t stem, el_val_t slot);
el_val_t non_weak_present(el_val_t stem, el_val_t slot);
el_val_t one_cycle(void);
el_val_t openai_chat_complete(el_val_t model, el_val_t base_url, el_val_t api_key, el_val_t safe_sys, el_val_t messages_json);
el_val_t openai_tools_json(el_val_t tools_anthropic);
el_val_t json_trim_dangling_escape(el_val_t s);
el_val_t utf8_safe_slice(el_val_t s, el_val_t n);
el_val_t agentic_tools_no_web(void);
el_val_t openai_agentic_loop(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages_in, el_val_t tools_log_in);
el_val_t parse_float_x100(el_val_t s);
el_val_t path_within_root(el_val_t path, el_val_t root);
el_val_t peo_ah_past(el_val_t slot);
Generated Vendored
+15
View File
@@ -1,3 +1,18 @@
//
// STALE BUNDLE DO NOT BUILD. UNSAFE CHAT PATH.
//
// This concatenated bundle is a snapshot, not a source of truth, and it is stale in
// a way that matters for safety: it wires /api/chat straight to handle_chat and
// contains NO layered_cycle at all (verified: zero occurrences in the bundled code
// the only textual hit in this file is this banner). A binary built from
// this file would run chat with no enforcing input gate (no safety_screen, no
// hard-bell short-circuit) and no enforcing output gate (no safety_validate).
//
// Build from the .el sources via manifest.el (entry soul.el), or from dist/soul.c.
// Nothing in the repo references this file. It is kept only as a historical artifact
// and should be deleted once Will confirms nothing external depends on it.
// (Flagged 2026-08-04 in _engine-websearch-20260804/SAFETY-STOP.md; banner added
// 2026-08-05 with the plain-chat generation fix.)
// language-profile.el - Language profile data and accessors.
//
// A language profile is a slot map ([String] key-value list) describing the
Generated Vendored
+1830 -678
View File
File diff suppressed because one or more lines are too long
Generated Vendored
+18
View File
@@ -0,0 +1,18 @@
# soul.c.stamp — fingerprint of the .el sources dist/soul.c was generated from.
# Written by tools/soulc-stamp.sh --write. Do not hand-edit.
# generated_amalgam_sha256 562193a341876f8fa0618ae7a76d0ecc073bb4d1d21b128c48b4aee28e1ef0d9
# generated_amalgam_bytes 1226240
f8597e10546654bce3fbbe40461b2da59d0e06dbf1b038d1d362d24f949e3911 awareness.el
2ff2dada732918c788a9ef66c6fd54c7a24cc4bbd4829197fe945d3a75ca1929 chat.el
42288c212cbf72fb1e8ecbd4d9900e4e9ee1cfa475b7974295c7637f1bf2939f elp-input.el
b3f77f49d6086932c38bd17fe7a5eaf8bce25685f6fc3e1750f05729c6b49b9e imprint.el
fba8ffdb9ba72bca5b09ca1c93a520edc52f3f4d8aec2c7585fe9b17e06420b2 manifest.el
550a72e234ae8cec1f33e02108fd365353f45edd88513da90b792e79b6c0e5f0 memory.el
5ec07ec9785b02abe32f3ff7acf2d1f9f7e07c0967fac97e6eff17d7110b5c84 neuron-api.el
03c47c451e0e87f2c252cadb4b765867943962a804f548dd53adeef0520912c8 persist.el
a6d69f3fc55233d9d3300160fd46a1551f2064bcd0fb84e2c9e432f636a72476 routes.el
c28e36952ec56525963a0bdf29455ab097d3b0c5653d19c25fbb005e1069a1f7 safety.el
fd3ab91d0ae0ea26639e21bef2f8f94054dc4b02eae68b19e3fe689d2769aad4 sessions.el
0f1cf43904a98a5a646cce5a07e0e96162ced662692fbc13357d9b67d9a8ac3d soul.el
30337940905171a9645b0929f0a412ce6b3dccb1246495070c553bca0bbae6cd stewardship.el
e105dc5990e6adbf39db9dc0462cd8bcf6e6c3dfd03709059227ecfad2bbab29 studio.el
+11 -2
View File
@@ -110,7 +110,7 @@ tool("beginSession", "Initialize session: surface recent high-importance memorie
"," + tool("linkCausal", "Create a causal edge (cause -> effect).") +
"," + tool("restructureCausalGraph", "Re-balance the causal subgraph after new evidence.") +
"," + tool("rebuildGraph", "Rebuild graph indices from the on-disk snapshot.") +
"," + tool("runStructuralAudit", "Audit graph structure for orphans, dangling edges, mislabeled types.") +
"," + tool("runStructuralAudit", "Stage 1 structural audit: owner-vs-runtime divergence, orphans and dangling edges, typed-edge distribution, self-model connectivity. Returns an annotated characterization, not a score.") +
// Backlog + work
"," + tool("planWork", "Create a backlog item.") +
"," + tool("reviewBacklog", "Browse work items.") +
@@ -680,7 +680,16 @@ fn dispatch_tool_call(tool_name: String, args: String) -> String {
return mcp_json_result(resp)
}
if str_eq(tool_name, "runStructuralAudit") {
let resp: String = http_get(neuron_url() + "/session/begin")
// Was: GET /session/begin an unrelated session digest returned under an
// audit tool name, i.e. the tool advertised a check that did not exist.
// Now points at the real Stage 1 route (neuron-api.el
// handle_api_structural_audit). Sample caps ride the query string; the
// defaults keep a manual audit to a couple of seconds.
let e_s: Int = json_get_int(args, "edge_sample")
let n_s: Int = json_get_int(args, "node_sample")
let qs: String = "?edge_sample=" + int_to_str(if e_s > 0 { e_s } else { 3000 })
+ "&node_sample=" + int_to_str(if n_s > 0 { n_s } else { 300 })
let resp: String = http_get(neuron_url() + "/audit/structural" + qs)
return mcp_json_result(resp)
}
+104 -9
View File
@@ -1,9 +1,91 @@
import "persist.el"
fn tier_working() -> String { return "Working" }
fn tier_episodic() -> String { return "Episodic" }
fn tier_canonical() -> String { return "Canonical" }
// Association on write
// DESIGN: "promotion integrates candidate nodes by linking them to existing nodes
// using typed semantic edges RATHER THAN APPENDING AS UNLINKED CONTENT" (CCR
// claim 29). Unlinked append is the explicitly rejected behaviour and it is the
// only behaviour this system had. Measured 2026-08-09 on Tim's graph: 14,214 edges
// across 80,936 nodes, 5% of nodes connected to anything, and NO edge created by
// any write since 2026-07-19 while 27,000+ nodes were added. A memory that forms
// no connections cannot be reached by spreading activation, so retrieval silently
// degrades to literal matching.
//
// BOUNDS, each one bought with a specific failure:
// * max 3 edges per memory link_memories.py's cap, precision over spray
// * never link to identity (self/*, Value): the existing policy is explicit that
// "memories must not pollute the self traversal by similarity; only an explicit
// citation may touch identity". Similarity is not citation.
// * never link telemetry (state-event, soul-response, boot_count, loop-outcome):
// these are ~97% of daily write volume (1,020 vs 31 real memories on 08-08).
// Linking them would add ~3,000 noise edges a day and re-flatten the graph in
// the name of connecting it.
// * fail-soft: a failed association never fails the write.
// Edges go through wt_edge so they reach the owner and survive restart.
fn mem_assoc_skip_label(label: String) -> Bool {
if str_contains(label, "state-event") { return true }
if str_contains(label, "soul-response") { return true }
if str_contains(label, "soul-outbox") { return true }
if str_contains(label, "boot_count") { return true }
if str_contains(label, "loop-outcome") { return true }
if str_contains(label, "search-result") { return true }
return false
}
// A candidate is linkable only if it is a real, distinct, non-identity node.
fn mem_assoc_ok(cand_id: String, cand_label: String, self_id: String) -> Bool {
if str_eq(cand_id, "") { return false }
if str_eq(cand_id, self_id) { return false }
// CASE MATTERS measured 2026-08-09. A lowercase-only check let a memory link
// to "Self — Values (grounded)", i.e. it polluted the self traversal, which is
// the one thing this policy exists to prevent. My verification had the same
// blind spot and printed PASS. Check every casing the graph actually uses, and
// exclude identity node TYPES as well as labels.
let lab: String = str_lower(cand_label)
if str_starts_with(lab, "self") { return false }
if str_starts_with(lab, "value") { return false }
if str_contains(lab, "values") { return false }
if str_contains(lab, "identity") { return false }
if mem_assoc_skip_label(cand_label) { return false }
return true
}
// One slot of the association. Manual unroll rather than a loop: EL's codegen
// mis-emits accumulating while-loops (documented at soul.el:212, which unrolled
// three affective slots for the same reason).
fn mem_assoc_slot(results: String, idx: Int, new_id: String) -> Void {
if idx >= json_array_len(results) { return }
let cand: String = json_array_get(results, idx)
let cid: String = json_get(cand, "id")
let clabel: String = json_get(cand, "label")
let ctype: String = json_get(cand, "node_type")
if str_eq(ctype, "Value") { return }
if str_eq(ctype, "DharmaSelf") { return }
if str_eq(ctype, "Safety") { return }
if mem_assoc_ok(cid, clabel, new_id) {
wt_edge(new_id, cid, el_from_float(0.5), "related")
}
}
// mem_associate connect a freshly written memory to what it is about.
fn mem_associate(new_id: String, content: String, label: String) -> Void {
if str_eq(new_id, "") { return }
if mem_assoc_skip_label(label) { return }
// Ask the graph what this memory resembles. Now that the store carries
// meaning-vectors this is semantic, not merely lexical.
let probe: String = str_slice(content, 0, 400)
let results: String = engram_recall_json(probe, 4)
if str_eq(results, "") { return }
mem_assoc_slot(results, 0, new_id)
mem_assoc_slot(results, 1, new_id)
mem_assoc_slot(results, 2, new_id)
}
fn mem_store(content: String, label: String, tags: String) -> String {
let id: String = engram_node_full(
let id: String = wt_node(
content,
"Memory",
label,
@@ -17,13 +99,26 @@ fn mem_store(content: String, label: String, tags: String) -> String {
println("[memory] write rejected by engram (empty id): label=" + label)
return ""
}
// Read back to verify the node actually persisted guards against silent write failures.
let readback: String = engram_get_node_json(id)
if str_eq(readback, "") || str_eq(readback, "{}") {
println("[memory] WRITE VERIFY FAILED: label=" + label + " id=" + id + " — node absent after write")
return ""
// wt_node has already read the node back locally and returns "" if it did
// not land, so the old duplicate read-back here is gone.
//
// HONESTY (neuron#117): the receipt now says WHERE the write is.
// The old unconditional "write verified" line asserted against the soul's
// own RAM true in memory, false on disk and printed ~115,000 times on
// Tim's machine while the canonical snapshot sat frozen for three days.
// wt_commit flushes the spool and then asks the OWNER. When it says false
// the node is real and recallable but not yet durable, and the log says so
// rather than claiming a save that did not happen. The id is still returned:
// the local write DID succeed, and the queued delta will be retried.
let durable: Bool = wt_commit(id)
// Associate AFTER the node is durable: an edge to a node that did not persist
// is a dangling edge, which is the defect the 2026-08-09 cleanup removed 830 of.
mem_associate(id, content, label)
if durable {
println("[memory] write persisted at owner: " + id + " label=" + label)
} else {
println("[memory] write IN MEMORY ONLY (queued for owner, not yet durable): " + id + " label=" + label)
}
println("[memory] write verified: " + id + " ok")
return id
}
@@ -51,12 +146,12 @@ fn mem_strengthen(node_id: String) -> Void {
// memory.el (imported first) so awareness.el and neuron-api.el can both call it.
fn mem_tombstone(node_id: String) -> String {
let tags: String = "[\"Tombstone\",\"status:deleted\"]"
let marker: String = engram_node_full(
let marker: String = wt_node(
node_id, "Tombstone", "tombstone:" + node_id,
el_from_float(0.01), el_from_float(0.01), el_from_float(1.0),
"Episodic", tags)
if !str_eq(marker, "") {
engram_connect(marker, node_id, el_from_float(1.0), "tombstones")
wt_edge(marker, node_id, el_from_float(1.0), "tombstones")
}
return marker
}
+534 -29
View File
@@ -195,15 +195,24 @@ fn api_compact_activated(raw: String, max_items: Int, snip: Int) -> String {
}
// api_persisted read-back-after-write guard against hallucinated saves.
// After a write builtin returns an id, confirm the node is actually queryable
// via engram_get_node_json(id) (returns "" or "null" when missing). Returns
// true only when the node is genuinely persisted.
//
// WIDENED FOR neuron#117. This function is the single gate every MCP write
// handler passes through before it reports success (10 call sites), which makes
// it the right place to close the honesty gap rather than editing ten receipts.
//
// It used to read back from engram_get_node_json the SOUL'S OWN in-process
// graph. In HTTP-engram mode that asserts the wrong thing: the soul is not the
// persistence owner, so a node present in its RAM and absent from the owner read
// as "persisted" and then vanished on the next restart. The guard was doing
// exactly what its comment promised and still certifying writes that did not
// survive. It now flushes the write-through spool and asks the OWNER.
//
// In file mode (no ENGRAM_URL) the soul IS the owner and wt_commit collapses to
// the original local read-back unchanged behaviour, which is what keeps this
// reversible.
fn api_persisted(id: String) -> Bool {
if str_eq(id, "") { return false }
let node: String = engram_get_node_json(id)
// engram_get_node_json returns "{}" (empty object) when node is not found not "" or "null".
// Check all three to guard against any runtime variation.
return !str_eq(node, "") && !str_eq(node, "null") && !str_eq(node, "{}")
return wt_commit(id)
}
// api_not_persisted standard error for a write that did not read back.
@@ -295,10 +304,35 @@ fn handle_api_begin_session(body: String) -> String {
let state_events: String = api_compact_node_array(state_events_raw, 5, 500)
let recent_raw: String = engram_scan_nodes_json(10, 0)
let recent: String = api_compact_node_array(recent_raw, 10, 240)
// SELF-SEEDED SLICE (2026-08-09). The design is explicit: "Every compilation
// query begins at the self-model node and traverses outward... structural
// reachability from the self-model node is a precondition for any node to
// appear in compiled context" (will-anderson patents/drafts/engram-claims.md,
// Self-Seeded Activation; DRAFT, not a filed provisional — cite it as such).
//
// Measured 2026-08-09 before this change: compiled context contained 0-1
// identity records out of 10, because compilation seeds from a hardcoded
// TEXT STRING, never from the self. Even an explicit "my values identity who
// I am" query returned a boot counter and state-events.
//
// This restores the designed behaviour WITHOUT repeating the failure that got
// self_neighbors set to [] in the first place: that was an UNBOUNDED ~90KB
// neighbour dump which closed the socket on every call. Same bound as every
// other list here cap 8, 240-char snippets. The self root has 34 direct
// neighbours of which 23 are identity records, so depth 1 is dense enough to
// be worth seeding and small enough to stay cheap.
let self_raw: String = engram_neighbors_json("kn-efeb4a5b-5aff-4759-8a97-7233099be6ee", 1, "both")
// Cap 24, not 8: measured 2026-08-09, the self root's first 8 neighbours are
// TAG nodes ("neuron", "tier:note", "disposition:experimental", "imprint",
// "traversal") which crowd out the substantive identity records behind them.
// The root has 34 neighbours of which 23 are identity; 24 captures them while
// staying bounded. Cost measured at ~+4KB on a ~12KB response, nowhere near
// the ~90KB unbounded dump that closed sockets and got this set to [].
let self_slice: String = api_compact_node_array(self_raw, 24, 240)
return "{\"stats\":" + stats
+ ",\"recent\":" + recent
+ ",\"activated\":" + activated
+ ",\"self_neighbors\":[]"
+ ",\"self_neighbors\":" + self_slice
+ ",\"recent_state_events\":" + state_events + "}"
}
@@ -342,10 +376,17 @@ fn handle_api_remember(body: String) -> String {
let inner: String = str_slice(base_tags, 1, str_len(base_tags) - 1)
"[" + inner + ",\"project:" + project + "\"]"
}
let id: String = engram_node_full(content, "Memory", "memory:remembered",
let id: String = wt_node(content, "Memory", "memory:remembered",
sal, sal, el_from_float(0.9),
"Episodic", final_tags)
if !api_persisted(id) { return api_not_persisted(id) }
// Associate on write (2026-08-09). THIS CALL MUST BE HERE, not only in mem_store.
// The HTTP memory route writes via wt_node directly; mem_store serves only the
// awareness paths (soul-response, search-result, activation-result) which are
// exactly the telemetry we refuse to link. Hooking mem_store alone produced
// ZERO edges across four real writes measured, not assumed, which is the only
// reason it was caught before shipping.
mem_associate(id, content, "memory:remembered")
return "{\"id\":\"" + id + "\",\"ok\":true}"
}
@@ -369,7 +410,7 @@ fn handle_api_node_create(body: String) -> String {
if str_eq(importance, "low") { 0.25 } else { 0.5 }
}
}
let id: String = engram_node_full(content, node_type, label,
let id: String = wt_node(content, node_type, label,
sal, sal, el_from_float(0.9),
tier, tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -422,11 +463,11 @@ fn handle_api_node_update(body: String) -> String {
}
let body_tags: String = json_get(body, "tags")
let tags: String = if str_eq(body_tags, "") { "[\"" + node_type + "\"]" } else { body_tags }
let new_id: String = engram_node_full(content, node_type, label,
let new_id: String = wt_node(content, node_type, label,
el_from_float(0.5), el_from_float(0.5), el_from_float(0.8),
tier, tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
engram_connect(new_id, id, el_from_float(0.9), "supersedes")
wt_edge(new_id, id, el_from_float(0.9), "supersedes")
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + id + "\",\"ok\":true}"
}
@@ -450,7 +491,12 @@ fn handle_api_recall(method: String, path: String, body: String) -> String {
if str_eq(eff_q, "") {
return api_or_empty(engram_scan_nodes_json(limit, 0))
}
let results: String = engram_search_json(eff_q, limit)
// engram_recall_json, not engram_search_json: this route IS the retrieval
// surface (claim 24's "embedding search queries"), so it gets the semantic
// and associative legs. engram_search_json stays lexical because ~40
// internal call sites pass a KEY and seven of them delete every record
// that comes back see the boundary note above eg_search_json_impl.
let results: String = engram_recall_json(eff_q, limit)
return api_or_empty(results)
}
@@ -498,7 +544,7 @@ fn handle_api_capture_knowledge(body: String) -> String {
let full: String = if str_eq(title, "") { content } else { title + ": " + content }
let lbl: String = str_slice(title, 0, 80)
let tags: String = "[\"Knowledge\",\"captured\"]"
let id: String = engram_node_full(full, "Knowledge", lbl,
let id: String = wt_node(full, "Knowledge", lbl,
el_from_float(0.85), el_from_float(0.8), el_from_float(0.9),
"Episodic", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -513,12 +559,12 @@ fn handle_api_evolve_knowledge(body: String) -> String {
if !str_eq(prior_id, "") && is_protected_node(prior_id) { return api_err_protected(prior_id) }
let tags: String = "[\"Knowledge\",\"evolved\"]"
// Empty label engram_node_full derives content[:60] (LABEL FIX 2026-07-23).
let new_id: String = engram_node_full(content, "Knowledge", "",
let new_id: String = wt_node(content, "Knowledge", "",
el_from_float(0.75), el_from_float(0.75), el_from_float(0.9),
"Episodic", tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
if !str_eq(prior_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true}"
}
@@ -535,11 +581,11 @@ fn handle_api_promote_knowledge(body: String) -> String {
"[\"Knowledge\",\"tier:canonical\",\"disposition:stable\"]"
} else { tags_raw }
// Empty label engram_node_full derives content[:60] (LABEL FIX 2026-07-23).
let new_id: String = engram_node_full(content, "Knowledge", "",
let new_id: String = wt_node(content, "Knowledge", "",
el_from_float(0.9), el_from_float(0.9), el_from_float(1.0),
"Canonical", tags)
if !api_persisted(new_id) { return api_not_persisted(new_id) }
engram_connect(new_id, prior_id, el_from_float(0.95), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.95), "supersedes")
return "{\"ok\":true,\"new_id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\"}"
}
@@ -562,7 +608,7 @@ fn handle_api_define_process(body: String) -> String {
if str_eq(content, "") { return api_err("content is required") }
let label: String = if str_eq(name, "") { "process:unnamed" } else { "process:" + name }
let tags: String = "[\"Process\"]"
let id: String = engram_node_full(content, "Process", label,
let id: String = wt_node(content, "Process", label,
el_from_float(0.8), el_from_float(0.8), el_from_float(0.9),
"Canonical", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -647,7 +693,7 @@ fn handle_api_tune_config(body: String) -> String {
if str_eq(key, "") { return api_err("key is required") }
let content: String = "config:" + key + "=" + value
let tags: String = "[\"ConfigEntry\",\"config\"]"
let id: String = engram_node_full(content, "ConfigEntry", key,
let id: String = wt_node(content, "ConfigEntry", key,
el_from_float(0.85), el_from_float(0.85), el_from_float(0.9),
"Canonical", tags)
if !api_persisted(id) { return api_not_persisted(id) }
@@ -694,7 +740,7 @@ fn handle_api_link_entities(body: String) -> String {
if is_protected_node(to_id) { return api_err_protected(to_id) }
let relation: String = json_get(body, "relation")
let eff_relation: String = if str_eq(relation, "") { "associates" } else { relation }
engram_connect(from_id, to_id, el_from_float(0.5), eff_relation)
wt_edge(from_id, to_id, el_from_float(0.5), eff_relation)
return "{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + eff_relation + "\"}"
}
@@ -727,11 +773,11 @@ fn handle_api_evolve_memory(body: String) -> String {
}
}
let tags: String = "[\"Memory\",\"evolved\"]"
let new_id: String = engram_node_full(content, "Memory", "memory:evolved",
let new_id: String = wt_node(content, "Memory", "memory:evolved",
sal, sal, el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true}"
}
@@ -789,11 +835,11 @@ fn handle_api_cultivate(body: String) -> String {
let content: String = json_get(body, "content")
if str_eq(content, "") { return api_err("content is required") }
let tags: String = "[\"Knowledge\",\"evolved\",\"cultivated\"]"
let new_id: String = engram_node_full(content, "Knowledge", "knowledge:cultivated",
let new_id: String = wt_node(content, "Knowledge", "knowledge:cultivated",
el_from_float(0.75), el_from_float(0.75), el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true,\"cultivated\":true}"
}
@@ -809,11 +855,11 @@ fn handle_api_cultivate(body: String) -> String {
}
}
let tags: String = "[\"Memory\",\"evolved\",\"cultivated\"]"
let new_id: String = engram_node_full(content, "Memory", "memory:cultivated",
let new_id: String = wt_node(content, "Memory", "memory:cultivated",
sal, sal, el_from_float(0.9),
"Episodic", tags)
if !str_eq(prior_id, "") && !str_eq(new_id, "") {
engram_connect(new_id, prior_id, el_from_float(0.9), "supersedes")
wt_edge(new_id, prior_id, el_from_float(0.9), "supersedes")
}
return "{\"id\":\"" + new_id + "\",\"supersedes\":\"" + prior_id + "\",\"ok\":true,\"cultivated\":true}"
}
@@ -833,7 +879,7 @@ fn handle_api_cultivate(body: String) -> String {
if str_eq(to_id, "") { return api_err("to_id is required") }
let relation: String = json_get(body, "relation")
let eff_relation: String = if str_eq(relation, "") { "associates" } else { relation }
engram_connect(from_id, to_id, el_from_float(0.5), eff_relation)
wt_edge(from_id, to_id, el_from_float(0.5), eff_relation)
return "{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + eff_relation + "\",\"cultivated\":true}"
}
@@ -868,7 +914,7 @@ fn handle_api_consolidate(body: String) -> String {
if !str_eq(summary, "") {
let safe_summary: String = str_replace(summary, "\"", "'")
let tags: String = "[\"SessionSummary\",\"consolidate\"]"
let summary_id: String = engram_node_full(
let summary_id: String = wt_node(
"[session-summary] " + safe_summary,
"SessionSummary", "session:summary",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
@@ -880,3 +926,462 @@ fn handle_api_consolidate(body: String) -> String {
}
return "{\"ok\":true,\"snapshot\":\"" + snap + "\"}"
}
// Stage 1: structural audit
//
// WHAT THIS IMPLEMENTS
// The CGI provisional, 05-detailed-description.md, "Stage 1: Structural audit
// 430". Verbatim, the audit module evaluates: the density and typed
// distribution of causal edges; the consistency between value nodes and
// execution-record neighborhoods; the richness and connectivity of the
// self-model; and the authenticity of open-question nodes in the wonder
// manifest. It "produces a coherence assessment 432 — NOT A BINARY SCORE but
// an annotated characterization of the graph's structural properties".
//
// That last clause is the whole shape of this handler. Every finding carries
// its own numbers AND a plain-language note saying what the numbers mean and
// how they were obtained. There is no pass/fail, no percentage-of-health, no
// composite score, and `"score":null` is emitted explicitly so a downstream
// reader cannot mistake its absence for an omission.
//
// WHY IT EXISTS NOW, AND WHY THE FIRST FINDING IS THE ONE IT IS
// `runStructuralAudit` has been an advertised MCP tool with nothing behind it:
// the dispatcher GET'd /session/begin and returned that blob (mcp-wrapper/src/
// main.el). Meanwhile the failure the audit would have caught ran silently for
// about three weeks the soul reported 103,089 nodes while the engram, which
// OWNS persistence, held ~79,900; a crash discarded the difference. Every boot
// reported green throughout, because nothing in the system ever compared the
// two sides. So finding 1 is owner-versus-runtime divergence: it is the check
// whose absence cost real memory, and it is cheap and exact.
//
// WHAT IS DELIBERATELY NOT HERE (stage 1b, see the `deferred` array in the
// response): value/execution-record consistency and wonder-manifest
// authenticity. Both need node types that barely exist in this graph today
// the response MEASURES those populations and reports the counts as the reason,
// rather than asserting a deferral without evidence.
//
// MEASUREMENT HONESTY: EXACT WHERE CHEAP, SAMPLED WHERE NOT, ALWAYS LABELLED
// Counts, edge typing and self-model connectivity are exact. Orphan rate and
// dangling-edge rate are SAMPLED, because the engram runtime has no node-id
// index `engram_find_node_index` is a linear scan over every node, so an
// exhaustive dangling check is O(nodes x edges) (~2.2e9 string compares at
// today's scale, tens of seconds inside one request). The samples are UNIFORM
// across the whole population, not head-of-list, and every sampled figure is
// emitted with its own `sampled` / `population` fields plus an extrapolation
// labelled as such. Raise `?edge_sample=` / `?node_sample=` to the population
// size to run either check exhaustively and pay the time. The real fix is an
// id index in the runtime; that is the engram repo's, not this handler's.
// audit_pct1 one-decimal percentage as a bare JSON number, sign-safe.
// Integer math only: EL has no fixed-precision formatter, and float_to_str
// would put an unbounded mantissa in the response.
fn audit_pct1(num: Int, den: Int) -> String {
if den <= 0 { return "null" }
let neg: Bool = num < 0
let a: Int = if neg { 0 - num } else { num }
let tenths: Int = (a * 1000) / den
let whole: Int = tenths / 10
let frac: Int = tenths - (whole * 10)
let sign: String = if neg { "-" } else { "" }
return sign + int_to_str(whole) + "." + int_to_str(frac)
}
// audit_finding the one envelope every finding uses: name, the measurements,
// and the annotation. Keeping it in one place is what stops the characterization
// from degenerating into a bag of numbers with no reading attached.
fn audit_finding(name: String, measured: String, note: String) -> String {
return "{\"finding\":\"" + name + "\""
+ ",\"measured\":{" + measured + "}"
+ ",\"note\":\"" + api_json_escape(note) + "\"}"
}
// audit_str_at read the quoted string value starting at byte `start`.
// Slices a bounded window rather than the tail of the (multi-MB) edges array, so
// this is O(window) per call instead of O(remaining input).
fn audit_str_at(s: String, start: Int, maxlen: Int) -> String {
let n: Int = str_len(s)
if start < 0 || start >= n { return "" }
let end_guess: Int = start + maxlen
let stop: Int = if end_guess > n { n } else { end_guess }
let win: String = str_slice(s, start, stop)
let q: Int = str_index_of(win, "\"")
if q < 0 { return "" }
return str_slice(win, 0, q)
}
// audit_rel_count exact count of edges carrying `rel`, by scanning the emitted
// edge array for the literal `"relation":"<rel>"`. engram_emit_edge_json writes
// metadata ESCAPED as a string, so no nested object can contain that literal and
// the count cannot be inflated by edge payloads.
fn audit_rel_count(edges: String, rel: String) -> Int {
return str_count(edges, "\"relation\":\"" + rel + "\"")
}
// audit_owner_stats ask the persistence OWNER for its own counts.
// Returns "" when there is no HTTP owner configured or the owner is unreachable;
// both are reported as findings, never as a failure of the audit.
fn audit_owner_stats(url: String) -> String {
if str_eq(url, "") { return "" }
return http_get(url + "/api/stats")
}
// audit_divergence FINDING 1. Runtime (this soul's in-process graph) versus
// the persistence owner's own count. Trend is measured against the previous
// audit recorded in soul state, so a second call answers "is the gap growing?"
// rather than just restating it.
fn audit_divergence() -> String {
let rt_nodes: Int = engram_node_count()
let rt_edges: Int = engram_edge_count()
let url: String = wt_engram_url()
if str_eq(url, "") {
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"none\",\"owner_reachable\":false",
"No HTTP persistence owner is configured, so this soul IS the owner "
+ "(file mode) and divergence is not defined. This check only has "
+ "meaning when ENGRAM_URL points at a separate engram that owns the "
+ "canonical store.")
}
let stats: String = audit_owner_stats(url)
// REACHABILITY IS PROVED BY THE PAYLOAD, NOT BY A NON-EMPTY REPLY.
// http_get does not return "" on a connection failure it returns a JSON
// error object ({"error":"Failed to connect to ... Couldn't connect to
// server"}). Testing only for "" made a DEAD owner read as reachable with
// node_count 0, i.e. the audit would have reported a 100% divergence and
// named it as data loss. That false positive is worse than no check at all:
// it is precisely the kind of confident wrong answer this route exists to
// stop. Require the field the contract promises.
let owner_nc_raw: String = json_get_raw(stats, "node_count")
if str_eq(stats, "") || str_eq(owner_nc_raw, "") {
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"" + api_json_escape(url) + "\",\"owner_reachable\":false"
+ ",\"owner_reply\":\"" + api_json_escape(api_utf8_trunc(stats, 200)) + "\"",
"The persistence owner at " + url + " did not return a node_count "
+ "from GET /api/stats. Divergence is UNKNOWN, NOT ZERO — an owner "
+ "that cannot be read is exactly the condition under which the "
+ "runtime's own count means least, and reporting 0 for the owner "
+ "would manufacture a total-loss reading out of a network error. "
+ "Reported as a finding rather than raised as an error so the rest "
+ "of the audit still returns; the owner's raw reply is in "
+ "owner_reply.")
}
let ow_nodes: Int = json_get_int(stats, "node_count")
let ow_edges: Int = json_get_int(stats, "edge_count")
let d_nodes: Int = rt_nodes - ow_nodes
let d_edges: Int = rt_edges - ow_edges
// Trend against the previous audit in this soul's state.
let prev_raw: String = state_get("audit_prev_node_delta")
let prev: Int = str_to_int(prev_raw)
let abs_now: Int = if d_nodes < 0 { 0 - d_nodes } else { d_nodes }
let abs_prev: Int = if prev < 0 { 0 - prev } else { prev }
let trend: String = if str_eq(prev_raw, "") {
"no_prior_audit"
} else {
if abs_now > abs_prev { "growing" } else {
if abs_now < abs_prev { "shrinking" } else { "flat" }
}
}
state_set("audit_prev_node_delta", int_to_str(d_nodes))
state_set("audit_prev_ts", int_to_str(time_now()))
let note_head: String = if d_nodes == 0 {
"Runtime and owner agree on node count."
} else {
"Runtime holds " + int_to_str(d_nodes) + " nodes (" + audit_pct1(d_nodes, rt_nodes)
+ "% of its own graph) that the persistence owner does not report. Nodes "
+ "that exist only in runtime memory do not survive a restart."
}
return audit_finding("owner_runtime_divergence",
"\"runtime_nodes\":" + int_to_str(rt_nodes)
+ ",\"runtime_edges\":" + int_to_str(rt_edges)
+ ",\"owner\":\"" + api_json_escape(url) + "\",\"owner_reachable\":true"
+ ",\"owner_nodes\":" + int_to_str(ow_nodes)
+ ",\"owner_edges\":" + int_to_str(ow_edges)
+ ",\"node_delta\":" + int_to_str(d_nodes)
+ ",\"edge_delta\":" + int_to_str(d_edges)
+ ",\"node_delta_pct_of_runtime\":" + audit_pct1(d_nodes, rt_nodes)
+ ",\"trend_vs_previous_audit\":\"" + trend + "\""
+ ",\"previous_node_delta\":" + (if str_eq(prev_raw, "") { "null" } else { int_to_str(prev) }),
note_head + " Trend against the previous audit recorded in this soul's "
+ "state: " + trend + ". This is the comparison whose absence let a "
+ "~24,000-node loss run for weeks with every boot reporting green.")
}
// audit_edge_typing FINDING 2. Density plus the typed distribution the patent
// asks for, against the claim-10 relation vocabulary. Exact: str_count over the
// emitted edge array, one linear pass per relation.
fn audit_edge_typing(edges: String, total_edges: Int, node_total: Int) -> String {
let c_sup: Int = audit_rel_count(edges, "Supersedes")
let c_cau: Int = audit_rel_count(edges, "Causes")
let c_con: Int = audit_rel_count(edges, "Contains")
let c_ref: Int = audit_rel_count(edges, "References")
let c_ctr: Int = audit_rel_count(edges, "Contradicts")
let c_exe: Int = audit_rel_count(edges, "Exemplifies")
let c_act: Int = audit_rel_count(edges, "Activates")
let c_tmp: Int = audit_rel_count(edges, "TemporallyPrecedes")
let typed: Int = c_sup + c_cau + c_con + c_ref + c_ctr + c_exe + c_act + c_tmp
// Lowercase near-misses: the same eight concepts written by the ad-hoc write
// paths (linkEntities defaults to "associates", linkCausal to "causes").
// Counted separately because "the vocabulary is unused" and "the vocabulary
// is used in the wrong case" are different defects with different fixes.
let l_sup: Int = audit_rel_count(edges, "supersedes")
let l_cau: Int = audit_rel_count(edges, "causes")
let l_con: Int = audit_rel_count(edges, "contains")
let l_ref: Int = audit_rel_count(edges, "references")
let l_ctr: Int = audit_rel_count(edges, "contradicts")
let l_exe: Int = audit_rel_count(edges, "exemplifies")
let l_act: Int = audit_rel_count(edges, "activates")
let l_tmp: Int = audit_rel_count(edges, "temporallyPrecedes")
let near: Int = l_sup + l_cau + l_con + l_ref + l_ctr + l_exe + l_act + l_tmp
let untyped: Int = total_edges - typed
return audit_finding("typed_edge_distribution",
"\"total_edges\":" + int_to_str(total_edges)
+ ",\"total_nodes\":" + int_to_str(node_total)
// Density per 100 nodes, not per node: EL has no fixed-precision float
// formatter, and "0.3 edges per node" rounded to an integer is a lie.
+ ",\"edges_per_100_nodes\":" + audit_pct1(total_edges, node_total)
+ ",\"claim10_typed\":" + int_to_str(typed)
+ ",\"claim10_typed_pct\":" + audit_pct1(typed, total_edges)
+ ",\"outside_claim10_vocabulary\":" + int_to_str(untyped)
+ ",\"lowercase_near_miss\":" + int_to_str(near)
+ ",\"by_relation\":{"
+ "\"Supersedes\":" + int_to_str(c_sup)
+ ",\"Causes\":" + int_to_str(c_cau)
+ ",\"Contains\":" + int_to_str(c_con)
+ ",\"References\":" + int_to_str(c_ref)
+ ",\"Contradicts\":" + int_to_str(c_ctr)
+ ",\"Exemplifies\":" + int_to_str(c_exe)
+ ",\"Activates\":" + int_to_str(c_act)
+ ",\"TemporallyPrecedes\":" + int_to_str(c_tmp) + "}",
"Only " + int_to_str(typed) + " of " + int_to_str(total_edges)
+ " edges use the claim-10 causal vocabulary; the remainder are ad-hoc "
+ "relation strings, which is why the graph's causal claims cannot yet "
+ "be checked for internal consistency — an untyped edge asserts "
+ "association, not causation. " + int_to_str(near) + " edges use a "
+ "lowercase spelling of a claim-10 relation: those are near-misses the "
+ "write paths could be corrected to emit, not genuinely foreign types.")
}
// audit_orphans_dangling FINDING 3. Both figures are SAMPLED; see the header
// for why exhaustive is O(nodes x edges) on this runtime.
//
// An "orphan" here is a node with zero RESOLVABLE edges: engram_neighbors_json
// drops any edge whose other endpoint does not resolve to a node, so a node
// whose only edges are dangling reads as an orphan. That is the right reading
// such a node is unreachable by traversal but it is stated rather than hidden.
fn audit_orphans_dangling(edges: String, total_edges: Int, node_total: Int,
edge_cap: Int, node_cap: Int) -> String {
// orphan sample: uniform stride over the node store
let n_take: Int = if node_total < node_cap { node_total } else { node_cap }
let n_stride: Int = if n_take > 0 { node_total / n_take } else { 1 }
let n_stride = if n_stride < 1 { 1 } else { n_stride }
let orphans: Int = 0
let n_checked: Int = 0
let j: Int = 0
while j < n_take {
let one: String = engram_scan_nodes_json(1, j * n_stride)
let nid: String = json_get(json_array_get(one, 0), "id")
if !str_eq(nid, "") {
let nbrs: String = engram_neighbors_json(nid, 1, "both")
let deg: Int = json_array_len(nbrs)
let orphans = if deg == 0 { orphans + 1 } else { orphans }
let n_checked = n_checked + 1
}
let j = j + 1
}
// dangling sample: uniform stride over the edge array
// str_index_of_all gives every edge's field offsets in ONE linear pass, so
// any index can be read in O(1). json_array_get would have been O(i) per
// element and O(n^2) over the array.
let from_pos: [Int] = str_index_of_all(edges, "\"from_id\":\"")
let to_pos: [Int] = str_index_of_all(edges, "\"to_id\":\"")
let nf: Int = len(from_pos)
let nt: Int = len(to_pos)
let ne: Int = if nf < nt { nf } else { nt }
let e_take: Int = if ne < edge_cap { ne } else { edge_cap }
let e_stride: Int = if e_take > 0 { ne / e_take } else { 1 }
let e_stride = if e_stride < 1 { 1 } else { e_stride }
let dangling: Int = 0
let e_checked: Int = 0
let i: Int = 0
while i < ne && e_checked < e_take {
let fid: String = audit_str_at(edges, get(from_pos, i) + 11, 96)
let tid: String = audit_str_at(edges, get(to_pos, i) + 9, 96)
let f_gone: Bool = str_eq(engram_get_node_json(fid), "{}")
let t_gone: Bool = if f_gone { true } else { str_eq(engram_get_node_json(tid), "{}") }
let dangling = if f_gone || t_gone { dangling + 1 } else { dangling }
let e_checked = e_checked + 1
let i = i + e_stride
}
let orphan_est: Int = if n_checked > 0 { (orphans * node_total) / n_checked } else { 0 }
let dangle_est: Int = if e_checked > 0 { (dangling * total_edges) / e_checked } else { 0 }
let exhaustive_n: String = if n_checked >= node_total { "true" } else { "false" }
let exhaustive_e: String = if e_checked >= ne { "true" } else { "false" }
return audit_finding("orphans_and_dangling_edges",
"\"nodes_population\":" + int_to_str(node_total)
+ ",\"nodes_sampled\":" + int_to_str(n_checked)
+ ",\"nodes_sample_exhaustive\":" + exhaustive_n
+ ",\"orphans_in_sample\":" + int_to_str(orphans)
+ ",\"orphan_rate_pct\":" + audit_pct1(orphans, n_checked)
+ ",\"orphans_extrapolated\":" + int_to_str(orphan_est)
+ ",\"edges_population\":" + int_to_str(total_edges)
+ ",\"edges_sampled\":" + int_to_str(e_checked)
+ ",\"edges_sample_exhaustive\":" + exhaustive_e
+ ",\"dangling_in_sample\":" + int_to_str(dangling)
+ ",\"dangling_rate_pct\":" + audit_pct1(dangling, e_checked)
+ ",\"dangling_extrapolated\":" + int_to_str(dangle_est),
"Orphan = zero RESOLVABLE edges, so a node whose only edges dangle counts "
+ "as an orphan; either way it is unreachable by traversal. Dangling = an "
+ "edge with an endpoint id that resolves to no node. Both are uniform "
+ "stride samples over the whole population, not the head of the list; "
+ "the extrapolations are estimates and are labelled as such. Pass "
+ "?node_sample= / ?edge_sample= at or above the population size to run "
+ "either check exhaustively. A high orphan rate is a characterization, "
+ "not a verdict: an accumulating store legitimately holds unlinked "
+ "material. It becomes a defect when the write paths were SUPPOSED to "
+ "link and did not.")
}
// audit_pillar one self-model pillar: present, how much content, how connected.
fn audit_pillar(key: String, id: String) -> String {
let node: String = engram_get_node_json(id)
let present: Bool = !str_eq(node, "{}") && !str_eq(node, "")
if !present {
return "\"" + key + "\":{\"id\":\"" + id + "\",\"present\":false"
+ ",\"content_length\":0,\"degree\":0}"
}
let content: String = json_get(node, "content")
let deg: Int = json_array_len(engram_neighbors_json(id, 1, "both"))
return "\"" + key + "\":{\"id\":\"" + id + "\",\"present\":true"
+ ",\"label\":\"" + api_json_escape(json_get(node, "label")) + "\""
+ ",\"tier\":\"" + api_json_escape(json_get(node, "tier")) + "\""
+ ",\"content_length\":" + int_to_str(str_len(content))
+ ",\"degree\":" + int_to_str(deg) + "}"
}
// audit_self_model FINDING 4. "the richness and connectivity of the
// self-model ... is it connected to behavioral evidence?"
//
// This finding RETIRES the Claude-side vitals identity block. That check lived
// outside the system it was checking a shell script grepping a snapshot so
// it could only ever report on a file, and it went on reporting green while the
// memory-philosophy pillar was absent from the live graph for about three weeks.
// Asking the running soul about its own three pillars is the designed mechanism;
// a shell probe was the fourth patch on the same hole.
fn audit_self_model() -> String {
let dna: String = audit_pillar("intellectual_dna", "kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6")
let val: String = audit_pillar("values_hub", "kn-5b606390-a52d-4ca2-8e0e-eba141d13440")
let phi: String = audit_pillar("memory_philosophy", "kn-dcfe04b3-3702-4cac-b6f0-ecb4db837eee")
let root: String = audit_pillar("self_root", "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee")
return audit_finding("self_model_connectivity",
"\"pillars\":{" + dna + "," + val + "," + phi + "," + root + "}",
"The three identity pillars plus the self root. `degree` counts nodes "
+ "reachable in one hop in either direction — the self-model's connection "
+ "to the rest of the graph. present:false on any pillar is the condition "
+ "that ran undetected for weeks; content_length distinguishes a pillar "
+ "that is present from one that is present but hollowed out. The patent "
+ "also asks whether the self-model makes ACCURATE PREDICTIONS about the "
+ "system's own behavior; that half needs Prediction nodes and is deferred "
+ "with the rest of stage 1b below.")
}
// audit_deferred what stage 1 does NOT yet evaluate, with the measured reason.
// Emitted as data, not as a comment, so a reader of the assessment sees the gap
// and its evidence rather than inferring completeness from silence.
fn audit_deferred() -> String {
let preds: Int = json_array_len(api_or_empty(engram_scan_nodes_by_type_json("Prediction", 50, 0)))
let wonders: Int = json_array_len(api_or_empty(engram_scan_nodes_by_type_json("WonderQuestion", 50, 0)))
return "[{\"deferred\":\"value_execution_record_consistency\""
+ ",\"stage\":\"1b\""
+ ",\"measured\":{\"prediction_nodes_found\":" + int_to_str(preds) + "}"
+ ",\"reason\":\"" + api_json_escape(
"The patent asks whether the execution history SUPPORTS the stated "
+ "values or shows systematic conflict. That requires execution "
+ "records tied to value nodes and predictions to score them against. "
+ "Prediction nodes found (capped at 50): " + int_to_str(preds)
+ ". Asserting value/execution coherence on that population would be "
+ "a fabricated result, which is worse than a stated gap.") + "\"}"
+ ",{\"deferred\":\"wonder_manifest_authenticity\""
+ ",\"stage\":\"1b\""
+ ",\"measured\":{\"wonder_question_nodes_found\":" + int_to_str(wonders) + "}"
+ ",\"reason\":\"" + api_json_escape(
"The patent asks whether pull weights CORRELATE WITH GENUINE "
+ "PREDICTION UNCERTAINTY or are uniform/externally assigned — a "
+ "correlation between two populations. WonderQuestion nodes readable "
+ "by type (capped at 50): " + int_to_str(wonders) + ", against "
+ int_to_str(preds) + " Prediction nodes. There is a known write/read "
+ "node-type mismatch on the wonder path; until that is fixed and both "
+ "populations exist, any correlation reported here would be noise.") + "\"}]"
}
// handle_api_structural_audit Stage 1. Returns the coherence assessment 432:
// an annotated characterization, explicitly NOT a score.
//
// COST NOTE: the edge findings need the relation labels, and the runtime exposes
// no edge-enumeration builtin. The only way to see them is the same one
// GET /api/graph/edges already uses engram_save to a SCRATCH path (never the
// owner's canonical file; see routes.el, neuron#117) and read the array back.
// On a large graph that is a multi-hundred-MB write, so this is a manual audit
// route, not something to put on a timer. Pass ?edges=0 to skip both edge
// findings and get the divergence + self-model readings cheaply.
fn handle_api_structural_audit(method: String, path: String, body: String) -> String {
let node_total: Int = engram_node_count()
let edge_total: Int = engram_edge_count()
let want_edges: Bool = !str_eq(api_query_param(path, "edges"), "0")
let edge_cap: Int = api_query_int(path, "edge_sample", 3000)
let node_cap: Int = api_query_int(path, "node_sample", 300)
let divergence: String = audit_divergence()
let self_model: String = audit_self_model()
let edge_part: String = if want_edges {
// Scratch export only. state_get("soul_snapshot_path") is deliberately
// NOT used: in HTTP-engram mode the soul is not the persistence owner and
// must never write the canonical file, not even on a read path.
let scratch_dir: String = env("TMPDIR")
let scratch_base: String = if str_eq(scratch_dir, "") { "/tmp" } else { scratch_dir }
let snap_path: String = scratch_base + "/soul-audit-export-" + state_get("soul_cgi_id") + ".json"
// engram_save returns Int (1 ok / 0 fail); str_eq on it SIGSEGVs (#150).
let saved: Int = engram_save(snap_path)
if saved == 0 {
"," + audit_finding("typed_edge_distribution", "\"available\":false",
"Could not export the graph to " + snap_path + " for edge analysis, "
+ "so edge typing and the dangling-edge sample were not run. "
+ "Reported as a gap, not as zero findings.")
} else {
// wt_read, not fs_read: fs_read leaves a thread-local length hint that
// the NEXT HTTP response would use as its Content-Length, appending
// adjacent heap bytes to the reply (see persist.el wt_read).
let snap: String = wt_read(snap_path)
let edges_raw: String = json_get_raw(snap, "edges")
let edges: String = if str_eq(edges_raw, "") { "[]" } else { edges_raw }
"," + audit_edge_typing(edges, edge_total, node_total)
+ "," + audit_orphans_dangling(edges, edge_total, node_total, edge_cap, node_cap)
}
} else {
""
}
return "{\"audit\":\"structural\",\"stage\":1"
+ ",\"spec\":\"CGI provisional 05-detailed-description.md, Stage 1: Structural audit 430\""
+ ",\"assessment\":\"coherence_assessment_432\""
+ ",\"assessment_kind\":\"annotated_characterization\""
+ ",\"score\":null"
+ ",\"score_note\":\"By design. The specification calls for an annotated characterization of the graph's structural properties, not a binary score. Read the findings.\""
+ ",\"cgi_id\":\"" + api_json_escape(state_get("soul_cgi_id")) + "\""
+ ",\"ts_ms\":" + int_to_str(time_now())
+ ",\"findings\":[" + divergence + "," + self_model + edge_part + "]"
+ ",\"deferred\":" + audit_deferred() + "}"
}
+1
View File
@@ -37,3 +37,4 @@ extern fn handle_api_memory_update(body: String) -> String
extern fn handle_api_cultivate(body: String) -> String
extern fn handle_api_list_typed(node_type: String, path: String, body: String) -> String
extern fn handle_api_consolidate(body: String) -> String
extern fn handle_api_structural_audit(method: String, path: String, body: String) -> String
+426
View File
@@ -0,0 +1,426 @@
// persist.el the soulengram WRITE-THROUGH boundary (neuron#117).
//
// WHY THIS FILE EXISTS
// soul.el:571-573 states the ownership rule: "when ENGRAM_URL is set the HTTP
// Engram owns persistence — the soul must NEVER write to the local snapshot
// (not the persistence owner)." The soul obeys the NEGATIVE half. The POSITIVE
// half how a write made inside the soul actually REACHES the owner was
// never built. Sync is pull-only (awareness.el `/api/sync` -> engram_load_merge),
// so every node the soul creates lives in its process RAM and is shed on
// restart. Measured live 2026-08-07: soul node_count=102184, engram
// node_count=79197 ~23k nodes existing nowhere but RAM.
//
// SCOPE NOTE ON THE PATENT (corrects an earlier internal reading)
// Engram provisional claims 15-18 describe a delta-sync protocol "with peer
// Engram instances"; claim 17's pull-then-push sequence is PEER-ENGRAM to
// PEER-ENGRAM. The soul is NOT a peer Engram it is a CALLER of the database
// system API (cf. claim 27, "invoked explicitly by a caller of the database
// system API"). So claim 17 does not specify a soulengram contract and is not
// cited as authority here. This design is derived from the ownership rule
// alone: the owner owns the writes, therefore the soul must HAND writes to the
// owner and must never write the owner's file itself.
//
// THE MECHANISM, AND WHY NOT `POST /api/nodes`
// The obvious route is the one the persona/boot-counter write-backs already
// use, POST /api/nodes. It is the wrong instrument here, verified against the
// live engram binary in a sandbox:
// - it mints a NEW server-side id (engram_node), so the soul's id and the
// owner's id diverge the next /api/sync pull re-imports the node as a
// DUPLICATE, and any edge referencing the soul's id never resolves;
// - it accepts only {content, node_type, salience} and drops label, tier,
// tags, importance, confidence, metadata. A probe posted with tier
// "Canonical" came back tier "Working", importance 0.5.
// POST /api/load-merge (Will's own route, el `dc39a61`) is the right one:
// - engram_load_merge PRESERVES the id and every field;
// - it dedups nodes by id and edges by (from_id,to_id,relation), so a
// re-submitted delta is a NO-OP retry safety is free, and it is the same
// local-wins semantics the graph already uses;
// - it calls persist_canonical() THE OWNER writes its own canonical file.
// The soul never touches it. The ownership rule is honoured in its
// strongest form rather than worked around;
// - it returns real counts {ok, nodes_added, edges_added, node_count},
// so a receipt can be a MEASUREMENT instead of a fixed success shape.
//
// SPOOL-AND-DRAIN, AND WHY IT IS NOT JUST A DIRECT POST
// Measured in a sandbox against a 79k-node / 176MB graph (live scale): one
// load-merge costs ~0.38s, essentially all of it the owner's persist_canonical.
// A chat turn writes 5-7 nodes; pushing each separately would add ~2.7s per
// turn. So writes are STAGED and pushed in one coalesced batch.
// The staging buffer is the FILESYSTEM, not process state, because the soul
// serves each HTTP connection on its own pthread (el_runtime http_serve_async)
// and a shared in-process buffer would lose entries to a read-modify-write
// race silently, which is the one failure mode this file exists to end.
// One file per write, named with uuid_v4, is race-free by construction and
// buys a property a memory buffer cannot: writes that could not be pushed
// SURVIVE A SOUL CRASH and are drained on the next boot.
//
// WHAT IS DELIBERATELY NOT PUSHED
// - InternalStateEvent / heartbeat telemetry. Will's own carve-out, stated in
// engram server.el 8f8ccc9: "48h-pruned, loss-tolerant, ~2/min; snapshotting
// 28MB per heartbeat is waste."
// NOTE (ours, flagged for Will): we do NOT additionally exclude Working-tier
// nodes. That exclusion exists in `fb0bb55` to stop the boot counter leaking
// through the /api/sync PULL; it is about sync backflow, not durability.
// Applying it here would exclude mem_store which writes tier "Working" and
// mem_store is the single most important durable write path in the soul. Boot
// seeding reads the canonical file wholesale, so a pushed Working-tier node
// does survive restart. This is the one classification call this file makes
// that Will has not ruled on.
//
// WHAT THIS BOUNDARY CANNOT EXPRESS (by construction, not by omission)
// - engram_strengthen (salience/activation drift): load-merge SKIPS ids that
// already exist, so it cannot update an existing node. There is no owner-side
// update/upsert route. Not pushable through any current route; left as a
// follow-up that needs a change in the engram repo.
// - engram_forget (hard delete): load-merge is additive and has no delete verb.
// Propagating deletes would mean DELETE /api/nodes/<id>, a HARD delete at the
// owner which scripts/verify-soul-contract.sh section B explicitly fails the
// build for ("to delete is to supersede/tombstone, never hard-remove"). Local
// deletes therefore stay local; the TOMBSTONE NODE and its "tombstones" edge
// (mem_tombstone) are pushed, and that is the sanctioned representation of a
// deletion in this graph.
// Configuration
// wt_engram_url same resolution order as ise_post: env, then the state key
// stashed at boot. NO hardcoded localhost fallback: unlike telemetry, inventing
// a destination for durable data would risk pushing a user's memories at whatever
// happens to be listening on 8742. Empty means "no HTTP owner" -> file mode.
fn wt_engram_url() -> String {
let env_url: String = env("ENGRAM_URL")
if !str_eq(env_url, "") { return env_url }
return state_get("soul_engram_url")
}
fn wt_api_key() -> String {
let env_key: String = env("ENGRAM_API_KEY")
if !str_eq(env_key, "") { return env_key }
return state_get("soul_engram_api_key")
}
// wt_enabled true only in HTTP-engram mode. In file mode the soul IS the
// persistence owner and every path below is a no-op, so this whole feature is
// inert for genesis/local deployments. That is also what makes it reversible.
fn wt_enabled() -> Bool {
return !str_eq(wt_engram_url(), "")
}
// wt_spool_dir where staged deltas live. MUST be readable by the engram
// process: /api/load-merge takes a PATH and the owner opens it itself. Both
// processes are same-host by construction (dev-stack LaunchAgents; the GKE
// image starts engram and soul in one container per entrypoint.sh).
fn wt_spool_dir() -> String {
let raw: String = env("SOUL_OUTBOX_DIR")
let dir: String = if str_eq(raw, "") { env("HOME") + "/.neuron/soul-outbox" } else { raw }
fs_mkdir(dir)
return dir
}
// Helpers
// wt_esc minimal JSON string escape. Deliberately local rather than reusing
// chat.el's json_safe: persist.el is imported BY memory.el, which is imported by
// chat.el, so depending on chat.el here would be an import cycle.
fn wt_esc(s: String) -> String {
let s1: String = str_replace(s, "\\", "\\\\")
let s2: String = str_replace(s1, "\"", "\\\"")
let s3: String = str_replace(s2, "\n", "\\n")
let s4: String = str_replace(s3, "\r", "\\r")
let s5: String = str_replace(s4, "\t", "\\t")
return s5
}
// wt_durable_class Will's telemetry carve-out, by node_type. See header.
fn wt_durable_class(node_type: String) -> Bool {
if str_eq(node_type, "InternalStateEvent") { return false }
return true
}
// wt_inner strip the surrounding brackets off a JSON array so several arrays
// can be concatenated into one. Returns "" for "[]" / "" / anything too short.
fn wt_inner(arr: String) -> String {
let n: Int = str_len(arr)
if n < 3 { return "" }
if !str_starts_with(arr, "[") { return "" }
return str_slice(arr, 1, n - 1)
}
// wt_read fs_read, plus a MANDATORY reset of the runtime's binary-length hint.
//
// THIS IS NOT OPTIONAL AND MUST NOT BE "SIMPLIFIED" BACK TO A BARE fs_read.
// The pinned runtime (vendor/el-runtime/v1.0.0-20260501) keeps a thread-local
// `_tl_fs_read_len` that fs_read SETS to the file's byte count (so binary files
// can be served with a correct Content-Length) and that http_send_response
// CONSUMES as the Content-Length of the next reply. Nothing else clears it
// except json_get_raw. So any fs_read during request handling that is not
// followed by a json_get_raw makes the NEXT HTTP response advertise the FILE's
// length instead of the body's and the runtime then sends that many bytes,
// appending whatever adjacent heap memory follows the reply.
//
// Caught here, measured: a /api/neuron/memory reply that should be 86 bytes went
// out as 497, with 411 bytes of this module's own spool paths and log strings
// trailing the JSON. The drain reads spool files mid-request, so this boundary
// is exactly where the landmine gets stepped on.
//
// Upstream el fixed the class in `43636ae` ("pair fs_read length hint with its
// buffer"); that runtime is NOT the one vendored here, and re-pinning the
// runtime is deliberately out of scope for this change. Clearing the hint at
// our own boundary fixes our exposure without touching the pinned C.
// json_get_raw is used as the reset because it is the only builtin in this
// runtime that zeroes the hint, and it does so before any early return.
fn wt_clear_binlen() -> Void {
let discard: String = json_get_raw("{}", "_wt_reset")
}
fn wt_read(path: String) -> String {
let data: String = fs_read(path)
wt_clear_binlen()
return data
}
// wt_sweep best-effort removal of the zero-byte husks left by truncation.
// The runtime exposes no unlink builtin, so a drained delta is emptied rather
// than deleted; this reclaims the directory entries.
//
// `-empty` is the safety property, not an optimisation: the command is
// STRUCTURALLY INCAPABLE of removing a delta that still has content, so it can
// never destroy a pending write even if it runs concurrently with a stage.
// Only the directory path is interpolated (never a filename), and it is quoted.
// The exit code is ignored an un-swept husk costs one directory entry.
fn wt_sweep(dir: String) -> Void {
if str_eq(dir, "") { return }
if str_contains(dir, "'") { return }
exec_command("find '" + dir + "' -maxdepth 1 -name 'wt*.json' -empty -delete 2>/dev/null")
}
// Staging
// wt_stage write ONE delta file. uuid_v4 in the name makes concurrent stagers
// collision-free without any lock. Returns true if the delta is on disk.
fn wt_stage(nodes_json: String, edges_json: String) -> Bool {
let dir: String = wt_spool_dir()
if str_eq(dir, "") { return false }
let payload: String = "{\"nodes\":" + nodes_json + ",\"edges\":" + edges_json + "}"
let path: String = dir + "/wt-" + uuid_v4() + ".json"
fs_write(path, payload)
// Read-back-verify the stage itself. A stage that did not land is a write we
// would otherwise believe was queued exactly the hallucinated-save class.
if str_eq(wt_read(path), "") {
println("[persist] wt_stage: FAILED to write spool file " + path + " — delta not queued")
return false
}
return true
}
// The write boundary
// wt_node create a node locally AND queue it for the persistence owner.
// Same signature and same return contract as engram_node_full ("" on failure),
// so converting a call site is a rename and nothing else.
fn wt_node(content: String, node_type: String, label: String,
salience: Float, importance: Float, confidence: Float,
tier: String, tags: String) -> String {
let id: String = engram_node_full(content, node_type, label,
salience, importance, confidence,
tier, tags)
if str_eq(id, "") { return "" }
// engram_get_node_json emits the SAME record shape engram_save writes (minus
// the embedding vector, which the owner backfills lazily), so the read-back
// doubles as the delta payload no second serialization to drift.
let rec: String = engram_get_node_json(id)
if str_eq(rec, "") || str_eq(rec, "{}") {
println("[persist] wt_node: local write did not read back, id=" + id + " label=" + label)
return ""
}
if wt_enabled() && wt_durable_class(node_type) {
wt_stage("[" + rec + "]", "[]")
}
return id
}
// wt_edge create an edge locally AND queue it. Mirrors engram_connect.
//
// The edge id is freshly generated rather than read back: the runtime exposes no
// "id of the edge I just created" accessor, and the owner dedups edges by
// (from_id,to_id,relation), never by id so the id is not load-bearing. The
// consequence, stated plainly: the soul's copy and the owner's copy of the same
// edge carry different edge ids. Nothing in either codebase looks an edge up by
// id (neighbors traversal scans from_id/to_id), so this is cosmetic.
fn wt_edge(from_id: String, to_id: String, weight: Float, relation: String) -> Void {
engram_connect(from_id, to_id, weight, relation)
if !wt_enabled() { return }
if str_eq(from_id, "") || str_eq(to_id, "") { return }
let ts: Int = time_now()
let rec: String = "{\"id\":\"" + uuid_v4() + "\""
+ ",\"from_id\":\"" + wt_esc(from_id) + "\""
+ ",\"to_id\":\"" + wt_esc(to_id) + "\""
+ ",\"relation\":\"" + wt_esc(relation) + "\""
+ ",\"metadata\":\"{}\""
+ ",\"weight\":" + float_to_str(weight)
+ ",\"confidence\":1"
+ ",\"created_at\":" + int_to_str(ts)
+ ",\"updated_at\":" + int_to_str(ts)
+ ",\"last_fired\":0,\"inhibitory\":0,\"layer_id\":1}"
wt_stage("[]", "[" + rec + "]")
}
// The drain
// wt_drain coalesce every staged delta into ONE load-merge against the owner.
//
// Returns: nodes_added on success (>= 0), 0 when there was nothing to do, and
// -1 when the push FAILED. -1 is load-bearing: on failure the spool files are
// left untouched, so nothing is lost and the next drain retries them. A caller
// must never read a non-negative return as "my particular node is durable"
// use wt_durable(id) for that.
//
// Concurrency: several threads may drain at once. Each builds its own batch file
// (uuid-named), and overlapping batches are harmless because load-merge dedups.
// Files are truncated ONLY after a confirmed ok:true, so a lost race costs a
// redundant push, never a dropped write.
fn wt_drain() -> Int {
if !wt_enabled() { return 0 }
let dir: String = wt_spool_dir()
if str_eq(dir, "") { return 0 }
// el_list_len/el_list_get, NOT json_stringify(fs_list(...)): fs_list builds
// a native list via el_list_append, and json_stringify does not serialize
// that type it renders the raw pointer value. (Verified in isolation; the
// same latent defect is live in studio.el's /api/tools/file/list route,
// which returns e.g. {"entries":4386409744}. Noted, not fixed here.)
let listing = fs_list(dir)
let count: Int = el_list_len(listing)
if count == 0 { return 0 }
let nodes_acc: String = ""
let edges_acc: String = ""
let drained: String = ""
let found: Int = 0
let i: Int = 0
// No `continue` / `break`: elc lists them as keywords but not one line of
// the shipped soul uses either, so they are unexercised on this build path.
// Guard conditions are expressed as nested ifs instead, and every rebind is
// at the loop-body top level where `let x = ...` is assignment (the idiom
// memory.el's boot-counter loop relies on) never inside a nested block,
// where it would shadow instead.
while i < count {
let name: String = el_list_get(listing, i)
// A delta is only usable when it ends with the closing "]}" that
// wt_stage writes last. fs_write is not atomic, so a file being written
// right now can be observed half-formed; requiring the terminator means
// it is picked up whole on the next drain instead of merged as garbage.
// An empty read means "already drained and truncated" not an error.
let p: String = if str_starts_with(name, "wt-") { dir + "/" + name } else { "" }
let raw: String = if str_eq(p, "") { "" } else { wt_read(p) }
let usable: Bool = !str_eq(raw, "") && str_ends_with(raw, "]}")
let nj: String = if usable { wt_inner(json_get_raw(raw, "nodes")) } else { "" }
let ej: String = if usable { wt_inner(json_get_raw(raw, "edges")) } else { "" }
let nodes_acc = if str_eq(nj, "") { nodes_acc } else if str_eq(nodes_acc, "") { nj } else { nodes_acc + "," + nj }
let edges_acc = if str_eq(ej, "") { edges_acc } else if str_eq(edges_acc, "") { ej } else { edges_acc + "," + ej }
let drained = if !usable { drained } else if str_eq(drained, "") { p } else { drained + "\n" + p }
let found = if usable { found + 1 } else { found }
let i = i + 1
}
if found == 0 { return 0 }
let combined: String = "{\"nodes\":[" + nodes_acc + "],\"edges\":[" + edges_acc + "]}"
let batch: String = dir + "/wtb-" + uuid_v4() + ".json"
fs_write(batch, combined)
if str_eq(wt_read(batch), "") {
println("[persist] wt_drain: could not write batch file " + batch + "" + int_to_str(found) + " deltas stay queued")
return -1
}
let url: String = wt_engram_url()
let key: String = wt_api_key()
let body: String = "{\"path\":\"" + wt_esc(batch) + "\",\"_auth\":\"" + wt_esc(key) + "\"}"
let resp: String = http_post_json(url + "/api/load-merge", body)
// The batch file is pure scratch the retry is rebuilt from the SPOOL, not
// from it. Truncate it unconditionally, before branching on the outcome, so
// a persistently unreachable owner cannot accumulate one husk per attempt.
fs_write(batch, "")
// Distinguish the two failures rather than collapsing them: "cannot reach
// the owner" and "the owner refused this delta" need different human
// responses, and a log line that says the wrong one costs a debugging hour.
// curl surfaces transport errors as a JSON body, so an empty response is not
// the only unreachable signal.
// (str_contains rather than a strict parse on purpose the engram's HTTP
// responses have been observed carrying trailing bytes past the JSON.)
let unreachable: Bool = str_eq(resp, "")
|| str_contains(resp, "Couldn't connect")
|| str_contains(resp, "Failed to connect")
|| str_contains(resp, "Could not resolve")
|| str_contains(resp, "timed out")
if unreachable {
wt_sweep(dir)
println("[persist] wt_drain: owner UNREACHABLE at " + url + "" + int_to_str(found)
+ " deltas stay queued in " + dir + " (will retry): " + resp)
return -1
}
if !str_contains(resp, "\"ok\":true") {
wt_sweep(dir)
println("[persist] wt_drain: owner REJECTED the delta — " + int_to_str(found)
+ " stay queued in " + dir + ": " + resp)
return -1
}
let added: Int = json_get_int(resp, "nodes_added")
let added_e: Int = json_get_int(resp, "edges_added")
// Confirmed. Truncate the drained spool files so they are not re-pushed.
// Truncation (not deletion) because the runtime exposes no unlink builtin;
// an emptied file is inert to the loop above. The zero-byte husks are then
// swept below.
let paths = str_split(drained, "\n")
let pn: Int = el_list_len(paths)
let k: Int = 0
while k < pn {
let one: String = el_list_get(paths, k)
if !str_eq(one, "") { fs_write(one, "") }
let k = k + 1
}
wt_sweep(dir)
println("[persist] wt_drain: pushed " + int_to_str(found) + " deltas -> owner added "
+ int_to_str(added) + " nodes, " + int_to_str(added_e) + " edges")
return added
}
// wt_durable is this id present AT THE OWNER? The only honest answer to
// "did my write persist" in HTTP mode.
//
// In file mode the soul IS the owner, so the local read-back is the owner-side
// read-back and this collapses to the pre-existing check.
//
// nodes_added from wt_drain is NOT a substitute: a concurrent drain may have
// already pushed this node, making our own added count 0 while the node is
// perfectly durable. Presence at the owner is the fact; counts are telemetry.
fn wt_durable(id: String) -> Bool {
if str_eq(id, "") { return false }
if !wt_enabled() {
let local: String = engram_get_node_json(id)
return !str_eq(local, "") && !str_eq(local, "null") && !str_eq(local, "{}")
}
let url: String = wt_engram_url()
let resp: String = http_get(url + "/api/nodes/" + id)
if str_eq(resp, "") { return false }
if str_eq(resp, "{}") { return false }
return str_contains(resp, "\"id\"")
}
// wt_commit flush, then assert at the owner. The receipt callers should use.
// Deliberately NOT a fixed success shape: it can and does return false while the
// local write is perfectly fine in RAM, which is the true state of affairs when
// the owner is unreachable.
fn wt_commit(id: String) -> Bool {
if str_eq(id, "") { return false }
if !wt_enabled() {
let local: String = engram_get_node_json(id)
return !str_eq(local, "") && !str_eq(local, "null") && !str_eq(local, "{}")
}
let pushed: Int = wt_drain()
return wt_durable(id)
}
+112 -13
View File
@@ -15,6 +15,40 @@ fn flag_true(body: String, key: String) -> Bool {
return json_get_bool(body, key) || json_get_int(body, key) > 0
}
// ---------------------------------------------------------------------------
// plain_chat_envelope the JSON response contract for a non-agentic ("Tools: Off")
// chat turn. Every /api/chat dispatch that calls layered_cycle goes through here, so
// the three call sites cannot drift apart.
//
// WHY THE ENVELOPE IS BUILT HERE AND NOT INSIDE layered_cycle:
// layered_cycle returns the user-facing text AFTER safety_validate has acted on it.
// Keeping the JSON out of the cycle means the output gate always sees raw model text
// and never an escaped blob there is nothing to unwrap and re-wrap on the crisis
// path, which is exactly the failure mode that made wiring handle_chat unsafe.
// Escaping is the last thing that happens, strictly after the gate.
//
// FIELDS: `reply` and `response` carry the same validated text. Both are required by
// live clients the desktop app reads `reply` first (DaemonClient.parseChatResponse),
// while the CLI tools and the Telegram gateway read `response` (the gateway reads only
// `response`). Emitting one would break the other.
//
// EMPTY MEANS FAILURE, NOT AN EMPTY ANSWER: a hard bell returns the fixed crisis
// message and a soft bell is padded to non-empty by safety_validate, so the only way
// an empty string leaves the cycle is a failed model call. It is reported as an error
// rather than dressed up as a successful blank reply.
// ---------------------------------------------------------------------------
fn plain_chat_envelope(validated: String, model: String) -> String {
if str_eq(validated, "") {
return "{\"error\":\"llm unavailable\",\"reply\":\"\",\"response\":\"\",\"agentic\":false,\"tools_used\":[]}"
}
let safe: String = json_safe(validated)
return "{\"reply\":\"" + safe + "\""
+ ",\"response\":\"" + safe + "\""
+ ",\"model\":\"" + json_safe(model) + "\""
+ ",\"agentic\":false"
+ ",\"tools_used\":[]}"
}
// ---------------------------------------------------------------------------
// Rate limiting simple in-memory per-IP sliding window counter.
//
@@ -152,7 +186,7 @@ fn route_imprint_contextual(body: String) -> String {
return "{\"ok\":false,\"error\":\"empty body\"}"
}
let tags: String = "[\"imprint\",\"contextual\"]"
let id: String = engram_node_full(
let id: String = wt_node(
body,
"Entity",
"imprint:contextual",
@@ -174,7 +208,7 @@ fn route_imprint_user(body: String) -> String {
return "{\"ok\":false,\"error\":\"empty body\"}"
}
let tags: String = "[\"imprint\",\"user\"]"
let id: String = engram_node_full(
let id: String = wt_node(
body,
"Entity",
"imprint:user",
@@ -205,7 +239,7 @@ fn route_synthesize(body: String) -> String {
}
let req: String = "synthesize " + parent_a + " " + parent_b
let tags: String = "[\"soul-inbox-pending\",\"synthesis-request\"]"
engram_node_full(
wt_node(
req,
"Entity",
"synthesis-request",
@@ -243,8 +277,14 @@ fn handle_dharma_recv(body: String) -> String {
} else if agentic_flag {
handle_chat_agentic(chat_body)
} else {
let screened_reply: String = layered_cycle(raw_msg)
screened_reply
// Non-agentic ("Tools: Off"): the full L1L2L3L1 cycle, which now generates
// at L3 instead of echoing. Envelope built outside the cycle see
// plain_chat_envelope.
// FIX B/E1 (2026-08-05): the cycle is told which conversation it is in, and
// whether this generation is conversation at all. Same two arguments at all
// three dispatch sites.
let screened_reply: String = layered_cycle(raw_msg, json_get(chat_body, "session_id"), is_utility_request(chat_body, json_get(chat_body, "session_id")))
plain_chat_envelope(screened_reply, chat_default_model())
}
auto_persist(chat_body, reply)
return reply
@@ -355,7 +395,28 @@ fn handle_connectors(method: String, clean: String, body: String) -> String {
return "{\"ok\":false,\"error\":\"unknown connectors route\"}"
}
// handle_request the soul's HTTP entry point.
//
// NOTE ON THE NAME (neuron#117): the el runtime resolves this handler by NAME
// via dlsym(RTLD_DEFAULT, "handle_request") that is why the Linux build must
// link -rdynamic. So the dispatcher body moved to route_dispatch and the name
// `handle_request` stays put as a thin wrapper. Do not rename it back.
//
// The wrapper exists to give the write-through boundary a guaranteed flush
// point. route_dispatch returns from ~60 places; a per-branch flush would be
// forgotten on the 61st. Draining here means EVERY request that staged a write
// pushes it before the connection closes, whatever route produced it, including
// routes added later that know nothing about persistence.
//
// wt_drain is a no-op (no HTTP, no cost) when nothing is staged and when the
// soul is not in HTTP-engram mode, so this is free on read traffic.
fn handle_request(method: String, path: String, body: String) -> String {
let resp: String = route_dispatch(method, path, body)
let flushed: Int = wt_drain()
return resp
}
fn route_dispatch(method: String, path: String, body: String) -> String {
let clean: String = strip_query(path)
// ACTIVITY STAMP (2026-07-30 self-review): every inbound HTTP request
@@ -392,10 +453,27 @@ fn handle_request(method: String, path: String, body: String) -> String {
return engram_scan_nodes_json(9999, 0)
}
if str_eq(clean, "/api/graph/edges") {
// TODO(reliability #8): engram_save races with awareness loop mem_save().
// Both now use atomic write-to-temp+rename (el_runtime.c). Serialised
// by engram_global_mu. Future: add engram_edges_json() builtin.
let snap_path: String = env("HOME") + "/.neuron/engram/snapshot.json"
// FIXED (neuron#117): this GET used to engram_save() straight over
// ~/.neuron/engram/snapshot.json a READ route, in a process that is
// NOT the persistence owner, overwriting the owner's canonical file
// on every call. It broke soul.el:571-573 ("the soul must NEVER write
// to the local snapshot") and it is the same defect class Will removed
// from the engram itself in el `dc39a61` ("stop read routes clobbering
// canonical snapshot"), where route_scan_edges/route_sync were moved
// to scratch paths for exactly this reason. It was also the race the
// old TODO(reliability #8) admitted to.
//
// Export to a scratch path instead. Same response, no canonical write.
// The soul's own snapshot writes are otherwise already gated behind
// state key "soul_snapshot_path", which is set ONLY in the genesis
// file-mode branch (soul.el: is_genesis && safe_to_seed, and
// safe_to_seed is unconditionally false when ENGRAM_URL is set) so
// after this change the soul writes nothing at all in HTTP mode.
// Future: add an engram_edges_json() builtin and drop the file round
// trip entirely.
let scratch_dir: String = env("TMPDIR")
let scratch_base: String = if str_eq(scratch_dir, "") { "/tmp" } else { scratch_dir }
let snap_path: String = scratch_base + "/soul-edges-export-" + state_get("soul_cgi_id") + ".json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
let edges_raw: String = json_get_raw(snap, "edges")
@@ -416,8 +494,11 @@ fn handle_request(method: String, path: String, body: String) -> String {
} else if agentic_flag {
handle_chat_agentic(body)
} else {
let screened_reply: String = layered_cycle(eff_msg)
screened_reply
// Non-agentic ("Tools: Off") same cycle and same envelope as POST.
// FIX B/E1: same threading. A GET probe usually carries no session_id, which
// resolves to the anonymous window the documented behaviour for this door.
let screened_reply: String = layered_cycle(eff_msg, json_get(body, "session_id"), is_utility_request(body, json_get(body, "session_id")))
plain_chat_envelope(screened_reply, chat_default_model())
}
auto_persist(body, reply)
return reply
@@ -486,6 +567,13 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_starts_with(clean, "/api/neuron/graph") {
return handle_api_inspect_graph(method, path, body)
}
// Stage 1 structural audit (CGI provisional, "Structural audit 430").
// GET because it is a read of the graph's own structure; the query string
// carries the sample caps (?edge_sample=, ?node_sample=, ?edges=0), so
// str_starts_with rather than str_eq.
if str_starts_with(clean, "/api/neuron/audit/structural") {
return handle_api_structural_audit(method, path, body)
}
if str_starts_with(clean, "/api/neuron/list/") {
// Offset 17 = len("/api/neuron/list/"). Was 16, which left a leading "/" on node_type
// ("/BacklogItem"), so engram_scan_nodes_by_type_json matched nothing list/<type>
@@ -580,8 +668,13 @@ fn handle_request(method: String, path: String, body: String) -> String {
} else if agentic_flag {
handle_chat_agentic(body)
} else {
let screened_reply: String = layered_cycle(raw_msg)
screened_reply
// Non-agentic ("Tools: Off") the app's DEFAULT mode (AgentMode.NEVER).
// Full L1L2L3L1 cycle with real generation at L3; envelope built
// outside the cycle so safety_validate always sees raw text.
// FIX B/E1: same threading. This is the app's main plain-chat door, so this
// is the site that ends the blank stare in practice.
let screened_reply: String = layered_cycle(raw_msg, json_get(body, "session_id"), is_utility_request(body, json_get(body, "session_id")))
plain_chat_envelope(screened_reply, chat_default_model())
}
auto_persist(body, reply)
return reply
@@ -662,6 +755,12 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(clean, "/api/neuron/graph/link") {
return handle_api_link_entities(body)
}
// POST accepted too: same handler, so a JSON-RPC-shaped caller that only
// speaks POST reaches the identical audit. Options still come from the
// query string the handler reads no body fields.
if str_eq(clean, "/api/neuron/audit/structural") {
return handle_api_structural_audit(method, path, body)
}
if str_eq(clean, "/api/neuron/memory") {
return handle_api_remember(body)
}
+1 -1
View File
@@ -204,7 +204,7 @@ fn safety_log_bell(level: String, reason: String, input_summary: String) -> Stri
// Emit a fallback println so the bell event leaves at least a log trace even
// when engram is degraded. This does not replace engram persistence -- it is a
// last-resort audit trail when the primary write cannot be confirmed.
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content,
"BellEvent",
"bell:" + level,
+108
View File
@@ -0,0 +1,108 @@
#!/usr/bin/env bash
# run-el-test.sh — compile and run one El test program from tests/.
#
# WHY THIS EXISTS (2026-08-07, issue #129):
# tests/ has held 14 test programs for months with no way to run them. CI does
# not run them. The convention printed in their own headers
# (`elc soul.el && ./soul --test tests/x.el`) refers to a --test flag the El
# runtime does not implement. So the tests were documentation, not gates —
# which is how a P0 safety regression shipped with a test directory present.
#
# THE RECIPE, AND WHY IT IS THIS SHAPE:
# Same discovery as gen-soul-amalgam.sh — `elc --target=c` emits only an extern
# prototype for any module that has a .elh header next to it, and inlines the
# module's bodies when it does not. A test that imports ../chat.el therefore
# compiles to a 18 KB unit full of unresolved externs unless the headers are
# out of the way. So: copy the sources into a scratch tree, delete every .elh
# on the import chain, and compile the test there.
#
# Scratch copy on purpose: the worktree is shared with other terminals and
# deleting headers in place would be a shared-tree mutation with no owner.
#
# EXIT STATUS IS THE GATE: non-zero if the binary fails to build, crashes, or if
# its output contains a FAIL line or reports a non-zero failed count. Do not
# "improve" this into something that only checks the exit code of the test
# binary — these El tests print failures and still exit 0.
#
# usage: scripts/run-el-test.sh tests/test_history_amplification.el
set -euo pipefail
TEST_REL="${1:?usage: run-el-test.sh tests/<test>.el}"
SRC="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
TEST_NAME="$(basename "$TEST_REL" .el)"
ELC="${ELC:-$HOME/neuron-dev-stack/src/el/lang/dist/platform/elc}"
[ -x "$ELC" ] || ELC="$HOME/el-sdk/elc"
[ -x "$ELC" ] || { echo "[run-el-test] FAIL: no elc found (set ELC=)"; exit 1; }
RTC="${RTC:-$SRC/vendor/el-runtime/v1.0.0-20260501/el_runtime.c}"
[ -f "$RTC" ] || RTC="$HOME/el-sdk/el_runtime.c"
[ -f "$RTC" ] || { echo "[run-el-test] FAIL: no el_runtime.c found (set RTC=)"; exit 1; }
RTDIR="$(dirname "$RTC")"
EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/el}"
SSL="${SSL_PREFIX:-/opt/homebrew/opt/openssl@3}"
GEN="$(mktemp -d "${TMPDIR:-/tmp}/el-test.XXXXXX")"
trap 'rm -rf "$GEN"' EXIT
mkdir -p "$GEN/neuron/tests" "$GEN/foundation/el/elp/src"
cp "$SRC"/*.el "$GEN/neuron/"
cp "$SRC"/tests/*.el "$GEN/neuron/tests/" 2>/dev/null || true
[ -d "$EL_REPO/elp/src" ] && cp "$EL_REPO"/elp/src/*.el "$GEN/foundation/el/elp/src/" 2>/dev/null || true
# The whole recipe depends on there being no headers to short-circuit inlining.
find "$GEN" -name '*.elh' -delete
echo "[run-el-test] compiling $TEST_REL"
( cd "$GEN/neuron" && "$ELC" --target=c "tests/${TEST_NAME}.el" ) > "$GEN/${TEST_NAME}.c"
BODIES=$(grep -c '^el_val_t .*) {$' "$GEN/${TEST_NAME}.c" || true)
echo "[run-el-test] $(wc -c < "$GEN/${TEST_NAME}.c" | tr -d ' ') bytes, ${BODIES} inlined function bodies"
# A test that imports ../chat.el pulls in the bulk of the engine. A tiny body
# count means an import was read from a header instead of inlined, and the test
# would be exercising extern stubs rather than the real code.
if [ "$BODIES" -lt 100 ]; then
echo "[run-el-test] FAIL: only $BODIES inlined bodies — an import was not inlined"
exit 1
fi
cc -O2 -DHAVE_CURL \
-I"$RTDIR" -I"$SSL/include" -L"$SSL/lib" \
"$GEN/${TEST_NAME}.c" "$RTC" \
-lssl -lcrypto -lcurl -lpthread -lm \
-o "$GEN/${TEST_NAME}" 2> "$GEN/cc.log" || {
echo "[run-el-test] FAIL: compile error"; tail -30 "$GEN/cc.log"; exit 1; }
# arm64 pointer-truncation guard (cc-brain.sh's rule): an implicit declaration of
# a runtime symbol truncates its returned pointer to 32 bits.
if grep -E 'implicit.*(engram_|el_)' "$GEN/cc.log"; then
echo "[run-el-test] FAIL: implicit declarations of runtime symbols"; exit 1; fi
# Throwaway HOME so a test can never read or write the live engram at ~/.neuron.
TEST_HOME="$GEN/home"
mkdir -p "$TEST_HOME"
echo "[run-el-test] running $TEST_NAME"
set +e
HOME="$TEST_HOME" NEURON_HOME="$TEST_HOME/.neuron" "$GEN/${TEST_NAME}" 2>&1 | tee "$GEN/out.txt"
RC=${PIPESTATUS[0]}
set -e
if [ "$RC" -ne 0 ]; then
echo "[run-el-test] FAIL: $TEST_NAME exited $RC (crash or abort)"
exit 1
fi
if grep -q " FAIL:" "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME reported failing assertions"
exit 1
fi
if grep -qE '[1-9][0-9]* failed' "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME reported a non-zero failed count"
exit 1
fi
if ! grep -q "PASS:" "$GEN/out.txt"; then
echo "[run-el-test] FAIL: $TEST_NAME produced no assertions at all"
exit 1
fi
echo "[run-el-test] PASS: $TEST_NAME"
+7 -7
View File
@@ -87,7 +87,7 @@ fn session_create(body: String) -> String {
let folder: String = json_get(body, "folder")
let content: String = session_make_content(id, title, ts, ts, folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let node_id: String = engram_node_full(
let node_id: String = wt_node(
content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -358,7 +358,7 @@ fn session_update_patch(session_id: String, body: String) -> String {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, eff_title, created_int, ts, eff_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_node_id: String = engram_node_full(
let new_node_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -456,7 +456,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
// TODO(reliability #7): delete-then-insert is not atomic concurrent saves for the
// same session can produce orphan history nodes. State is primary truth; engram fallback.
let tags: String = "[\"session\",\"session-history\",\"Conversation\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
hist, "Conversation", "session:messages:" + session_id,
el_from_float(0.6), el_from_float(0.6), el_from_float(0.9),
"Episodic", tags
@@ -488,7 +488,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
+ " | ts:" + int_to_str(ts_now)
let summary_tags: String = "[\"session-emotional-summary\",\"affective\",\"bell:" + eff_level + "\",\"BellEvent\"]"
let summary_sal: String = if str_eq(eff_level, "hard") { el_from_float(0.95) } else { el_from_float(0.85) }
let sum_discard: String = engram_node_full(
let sum_discard: String = wt_node(
summary_content,
"BellEvent",
"session:emotional-summary",
@@ -529,7 +529,7 @@ fn session_hist_save(session_id: String, hist: String) -> Void {
if !str_eq(ot_id, "") { engram_forget(ot_id) }
let oti = oti + 1
}
let discard_topic: String = engram_node_full(
let discard_topic: String = wt_node(
topic_content, "Conversation", topic_label,
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", topic_tags
@@ -582,7 +582,7 @@ fn session_update_meta_timestamp(session_id: String) -> Void {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, old_title, created_int, ts, old_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_id: String = engram_node_full(
let new_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
@@ -629,7 +629,7 @@ fn session_auto_title(session_id: String, first_message: String) -> Void {
let created_int: Int = str_to_int(old_created)
let new_content: String = session_make_content(session_id, new_title, created_int, ts, old_folder)
let tags: String = "[\"session\",\"session:meta\",\"Conversation\"]"
let new_id: String = engram_node_full(
let new_id: String = wt_node(
new_content, "Conversation", "session:meta",
el_from_float(0.7), el_from_float(0.7), el_from_float(0.9),
"Episodic", tags
+88 -35
View File
@@ -379,9 +379,23 @@ fn emit_session_start_event() -> Void {
// layered_cycle routes user-facing requests through the 4-layer consciousness stack.
// L0 (core) L1 (safety screen) L2a (continuity + behavioral profiling) L2b (mission alignment) L3 (imprint) L1 (safety validate)
// Internal cognition (heartbeat, proactive, memory ops) bypasses layers use one_cycle directly.
fn layered_cycle(raw_input: String) -> String {
let history: String = state_get("conv_history")
let session_id: String = state_get("current_session_id")
//
// FIX B (2026-08-05) the cycle now knows which conversation it is in.
//
// session_id: the caller's session, threaded from the route. Was previously read from the
// state key "current_session_id", which is read HERE and written NOWHERE in the entire
// source verified across every .el file. So this value was unconditionally "", and every
// downstream consumer of it silently fell back to a process-global bucket: conversation
// history, and the steward's continuity tracking (TODO reliability #4, below, describes the
// cross-session bleed this caused; threading the real id closes it). The plain path's blank
// stare and the agentic path's scoped history were the same defect seen from two sides.
//
// utility: true when the generation is not part of the user's conversation the app's
// title and insight passes. Answered normally, never recorded. See is_utility_request.
fn layered_cycle(raw_input: String, session_id: String, utility: Bool) -> String {
// Safety-screen history amplification now reads the SAME window the turn will be
// recorded into, so a session's own escalation pattern is what gets scored.
let history: String = state_get(conv_hist_key(session_id))
// L1 in: safety screen
let screen_result: String = safety_screen(raw_input, history)
@@ -423,8 +437,10 @@ fn layered_cycle(raw_input: String) -> String {
let cont_action: String = json_get(continuity, "action")
// Store continuity status so imprint can adjust its response register.
// TODO(reliability #4): session_continuity is process-global; scope per session_id
// when available to prevent cross-session bleed under concurrent layered_cycle calls.
// TODO(reliability #4) CLOSED 2026-08-05: this line was already written to scope per
// session it just never received a session id, because the only source was a state key
// nothing wrote. It is now threaded from the route, so named sessions genuinely get their
// own continuity state and only anonymous callers share the global one.
let cont_key: String = if str_eq(session_id, "") { "session_continuity" } else { "session_continuity:" + session_id }
state_set(cont_key, cont_status)
@@ -453,40 +469,24 @@ fn layered_cycle(raw_input: String) -> String {
let lc_aff_cutoff: Int = time_now() - 259200
let lc_bell_nodes: String = engram_search_json("bell:soft bell:hard BellEvent affective", 2)
let lc_has_bell: Bool = !str_eq(lc_bell_nodes, "") && !str_eq(lc_bell_nodes, "[]")
// CRASH FIX 2026-08-05 (BUG-PLAINCHAT-1): the " | ts:" parser used to be inline here.
// Inside this block-expression initializer elc compiled `lbmp + str_len(lbm)` to
// el_str_concat() on two integers, which segfaulted the whole daemon the moment a
// distress turn followed an earlier affective turn i.e. exactly on the crisis path.
// Verified against the unmodified baseline binary AND present in the committed
// dist/soul.c. affective_node_ts() is a top-level function, where the same expression
// compiles to integer addition. Do not inline it back.
let lc_bell_note: String = if lc_has_bell {
let lb0: String = json_array_get(lc_bell_nodes, 0)
let lb_c: String = json_get(lb0, "content")
let lbm: String = " | ts:"
let lbmp: Int = str_index_of(lb_c, lbm)
let lb_ts_raw: String = if lbmp >= 0 {
let lbs: Int = lbmp + str_len(lbm)
let lbr: String = str_slice(lb_c, lbs, str_len(lb_c))
let lbn: Int = str_index_of(lbr, " | ")
if lbn < 0 { lbr } else { str_slice(lbr, 0, lbn) }
} else {
let lbca: String = json_get(lb0, "created_at")
if str_eq(lbca, "") { json_get(lb0, "updated_at") } else { lbca }
}
let lb_ts: Int = if str_eq(lb_ts_raw, "") { 0 } else { str_to_int(lb_ts_raw) }
let lb_ts: Int = affective_node_ts(lb0)
if lb_ts > lc_aff_cutoff { "[AFFECTIVE NOTE: User was in distress in a recent session.]" } else { "" }
} else { "" }
let lc_pos_nodes: String = engram_search_json("PositiveEvent joy:high joy:low affective", 2)
let lc_has_pos: Bool = !str_eq(lc_pos_nodes, "") && !str_eq(lc_pos_nodes, "[]")
// Same crash fix as the bell note above (BUG-PLAINCHAT-1).
let lc_pos_note: String = if lc_has_pos && str_eq(lc_bell_note, "") {
let lp0: String = json_array_get(lc_pos_nodes, 0)
let lp_c: String = json_get(lp0, "content")
let lpm: String = " | ts:"
let lpmp: Int = str_index_of(lp_c, lpm)
let lp_ts_raw: String = if lpmp >= 0 {
let lps: Int = lpmp + str_len(lpm)
let lpr: String = str_slice(lp_c, lps, str_len(lp_c))
let lpn: Int = str_index_of(lpr, " | ")
if lpn < 0 { lpr } else { str_slice(lpr, 0, lpn) }
} else {
let lpca: String = json_get(lp0, "created_at")
if str_eq(lpca, "") { json_get(lp0, "updated_at") } else { lpca }
}
let lp_ts: Int = if str_eq(lp_ts_raw, "") { 0 } else { str_to_int(lp_ts_raw) }
let lp_ts: Int = affective_node_ts(lp0)
if lp_ts > lc_aff_cutoff { "[AFFECTIVE NOTE: User shared positive news in a recent session.]" } else { "" }
} else { "" }
let lc_affective_note: String = if !str_eq(lc_bell_note, "") { lc_bell_note } else { lc_pos_note }
@@ -498,11 +498,47 @@ fn layered_cycle(raw_input: String) -> String {
}
state_set("layered_cycle_safety_system_addendum", augmented_addendum)
// L3: imprint responds
let output: String = imprint_respond(aligned, imprint_id)
// L3: imprint responds applies the active imprint's voice/domain annotation to the
// steward-aligned input. This produces the PROMPT, not the answer.
let prompt: String = imprint_respond(aligned, imprint_id)
// L1 out: validate output before delivery
return safety_validate(output, screen_action)
// L3b: the imprint SPEAKS (added 2026-08-05).
//
// Until now the cycle stopped at the annotation above, so /api/chat with agentic:false
// handed the user's own screened text back as the "reply" every gate ran, but nothing
// ever generated. The generation is placed HERE, inside the cycle, rather than by
// pointing the route at handle_chat(): handle_chat has no enforcing input gate and no
// enforcing output gate, so calling it instead of this cycle would have traded the whole
// safety pipeline for a working reply. Composing keeps both.
//
// Order is deliberate and must not be rearranged: this call sits strictly AFTER the L1
// screen, the safe-mode guard, the hard-bell short-circuit and the L2 stewardship layers,
// and strictly BEFORE the L1 output gate. A hard bell never reaches a model the branch
// above returns first. Tools are not offered on this turn; see layered_generate.
let output: String = layered_generate(prompt, imprint_id, session_id)
// L1 out: validate output before delivery. Still the terminal gate nothing below this
// line can change the string this function returns.
let validated: String = safety_validate(output, screen_action)
// Turn bookkeeping. Records the VALIDATED text, never the raw model output, and is only
// reachable on the non-bell path: both bell branches above return before this point, so
// bell turns still never enter conversation history. Pure state side effect it cannot
// alter what is returned.
//
// FIX A: the receipt is unconditional and always negative on this path, because on this
// path it is structurally true layered_generate offers no tools at all (build_system_prompt
// chat mode + a request body with no "tools" key). Recording "no tools ran" is not padding:
// it is the only thing that distinguishes "nothing ran" from "we forgot to write down what
// ran", and that ambiguity is what made the model confess to a search it had performed.
//
// FIX E1: a utility generation is answered but not recorded. Guarded here rather than at
// the route so every /api/chat dispatch site inherits it from one place.
let receipt: String = tool_receipt("", "")
if !utility {
conv_history_record(session_id, raw_input, validated, receipt)
}
return validated
}
let soul_cgi_id_raw: String = env("SOUL_CGI_ID")
@@ -621,6 +657,23 @@ if is_genesis && safe_to_seed {
}
}
// CRASH RECOVERY (neuron#117). Deltas the previous process staged but could not
// hand to the owner are still on disk the spool is a filesystem queue, not a
// memory buffer, precisely so that a soul that died mid-flight does not take its
// unpushed writes with it. Drain them before serving, so recovered memories are
// durable and recallable from the owner from the first request onward.
//
// Safe on a clean boot: an empty spool means no HTTP call at all. Safe in file
// mode: wt_drain returns immediately when ENGRAM_URL is unset.
let wt_recovered: Int = wt_drain()
if wt_recovered > 0 {
println("[soul] write-through: recovered " + int_to_str(wt_recovered)
+ " nodes from a previous process's spool -> persistence owner")
}
if wt_recovered < 0 {
println("[soul] write-through: spool present but the persistence owner is unreachable — queued, will retry on heartbeat")
}
println("[soul] serving on port " + int_to_str(port))
http_serve_async(port, "handle_request")
println("[soul] awareness loop starting")
+3 -1
View File
@@ -1,6 +1,8 @@
// auto-generated by elc --emit-header - do not edit
extern fn init_soul_edges() -> Void
extern fn ensure_self_canonical_bridge() -> Void
extern fn aff_try_slot(slot_json: String, aff_7d_ts: Int, acc_key: String) -> Void
extern fn load_identity_context() -> Void
extern fn seed_persona_from_env() -> Void
extern fn emit_session_start_event() -> Void
extern fn layered_cycle(raw_input: String) -> String
extern fn layered_cycle(raw_input: String, session_id: String, utility: Bool) -> String
+2 -2
View File
@@ -11,7 +11,7 @@ import "memory.el"
fn steward_log_event(kind: String, detail: String) -> Void {
let content: String = "STEWARD:" + kind + " | " + detail
let tags: String = "[\"stewardship\",\"steward:" + kind + "\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
content,
"StewardshipEvent",
"steward:" + kind,
@@ -221,7 +221,7 @@ fn steward_fingerprint_session(input: String, session_id: String) -> String {
+ " formality=" + fs_str
+ " time=" + tb_str
let sample_tags: String = "[\"behavior\",\"BehaviorSample\",\"stewardship\"]"
let discard: String = engram_node_full(
let discard: String = wt_node(
sample_content,
"BehaviorSample",
"behavior:" + session_id,
+90
View File
@@ -0,0 +1,90 @@
# gate-openai — deterministic OpenAI-dialect provider stub
Staging home for the **soul-openai-tools-v2** gate scaffolding
(`docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md`, test-plan rung 1:
"stub first — discriminates before El code exists"). Sibling of gate9's
Anthropic stub (`_wt-beta-round9/scripts/gate9/stub-llm.py`): same scenario
mechanism, opposite wire dialect. Stdlib Python only, 127.0.0.1 only,
refuses ports 7770/7779/17779. Run `./selftest.sh` — exit 0 is green.
## Files
| File | Role |
|---|---|
| `stub-openai.py` | HTTP server: `POST /v1/chat/completions` (OpenAI dialect), scenario-scripted responses, request validation, ground-truth JSONL log, hostile modes via `--mode` |
| `scenarios-openai.json` | Scenario contract: scripts + markers + per-class/per-step request assertions |
| `selftest.sh` | curl-driven proof of every scenario, every rejection, all hostile modes (58 checks) |
## What each scenario proves (when the brain drives it)
| Class | Proves |
|---|---|
| `oa-plain` | finish_reason `stop` ends the loop; tools + `tool_choice` + `parallel_tool_calls:false` were offered on the wire |
| `oa-tools-off` | the chat-only lane sends NO tools (offering them there is a 400) |
| `oa-single-tool` | full round-trip: `tool_calls` parsed, assistant echo + `role:"tool"` turn with matching `tool_call_id` sent back, final text reached |
| `oa-torture` | `function.arguments` (JSON-encoded string with nested quotes, backslashes, newlines, tabs, unicode) survives exactly ONE decode — the stub recomputes the issued payload from the script and 400s on any drift (`gate_echo_mismatch`, the spec §6 two-escaper trap) |
| `oa-parallel` | two `tool_calls` in one response: the brain either answers both (paired correctly) or rejects cleanly — an unpaired echo is a 400 |
| `oa-mission` | multi-round loop continuation; step index = assistant-message count, so resume threads index correctly by construction |
| `oa-api-error` | provider errors 400/429/500/503 in the OpenAI error envelope surface honestly, no retry storm |
Universal (every request, any scenario): Anthropic dialect leakage fails
loudly with 400 — `anthropic-version` header, top-level `system` /
`stop_sequences` / `max_tokens_to_sample`, `input_schema` inside tools,
Anthropic content blocks (`tool_use`/`tool_result`/...). Tools must be
`{type:"function", function:{name, description, parameters}}`, unique names;
echoed `arguments` must be a JSON-encoded STRING, never a decoded object.
## Hostile modes (`--mode`, same file)
| Mode | Behavior | Brain invariant under test |
|---|---|---|
| `black-hole` | reads the request, never responds | HTTP timeout exists and surfaces; no silent hang |
| `mid-body-drop` | 200 headers, half a JSON body, socket abort | truncated body = clean error, never a half-parsed reply shown as real |
| `tool-pending-forever` | every request gets a fresh `tool_calls` response, forever | the loop's iteration cap trips (`max_loop_iterations: 16` in the contract); count actual round-trips via `GET /gate/stats` (`chat_hits`) |
## How the brain-side gate consumes this
1. Start: `stub-openai.py --port P --scenarios scenarios-openai.json --log run.jsonl`
2. Point the brain at it: `NEURON_LLM_0_URL=http://127.0.0.1:P` +
`NEURON_LLM_0_FORMAT=openai` (spec step 0 must verify these actually
export at runtime), scratch profile, free soul port.
3. Send each phrasing's `prompt` (the marker selects the script); assert the
brain's claims (`tools_used`, reply, ledger) against the stub's JSONL log
— truth, not narration — plus files on disk for write_file scenarios.
4. Any stub 400 = the brain sent a malformed/leaked request; the gate fails
with the stub's reason string.
5. Re-run gate9's Anthropic matrix unchanged = proof the Anthropic lane is
byte-untouched.
## Reconciliation into gate9 (app repo) — AFTER round 9 merges
This dir is staging only; the merge is mechanical by design:
- `stub-openai.py` + `scenarios-openai.json` move to `scripts/gate9/`
alongside `stub-llm.py` + `scenarios.json` (shared conventions: marker
matching, assistant-count step indexing, `--port/--scenarios/--log`,
JSONL fields `seq/ts/kind/scenario_class/phrasing/step/validation/
delivered/http_status`, prod-port refusal, benign background responses,
`GATE-SCRIPT-EXHAUSTED` overrun, `{N}/{NN}` repeat expansion).
- `prompt-matrix-gate.sh` gains a dialect axis (anthropic|openai) choosing
stub + scenario file; `matrix-asserts.py` reads the same log shape.
- The `--mode` hostile flags here are PROVIDER-side (brain↔LLM boundary);
gate9's `hostile/` servers are SOUL-side (app↔brain boundary). They are
complementary, not duplicates — both stay.
## Open questions for the port author (stub asserts a position; confirm or change)
1. `parallel_tool_calls` must be **explicitly false** on every tool-bearing
request (ADR-0005 pin). If the builder omits it instead, relax
`defaults.expect_request.parallel_tool_calls` to `null`.
2. `tool_choice` must be present (`"auto"` expected). If the brain relies on
the provider default, drop `require_tool_choice`.
3. Tool-result `content` is asserted only to be a string; if the brain sends
structured JSON-in-string (like `{"ok":true,...}`), no change needed.
4. Groq compatibility: Groq's OpenAI-compat endpoint rejects some optional
fields; whatever field set the brain settles on for live Groq E2E must be
mirrored here so the deterministic gate and the live lane assert the SAME
request shape.
5. The stub treats a `role:"tool"` turn answering an already-answered id as
400; if the resume path can legitimately replay tool results, that rule
needs a resume-aware carve-out (gate9's Anthropic stub faced the same
issue — see its PASS 1 comment).
+559
View File
@@ -0,0 +1,559 @@
#!/usr/bin/env bash
# run-lane-gate.sh — brain-side driver for the OpenAI-dialect gate.
#
# Drives the REAL soul binary against stub-openai.py for every class and every
# phrasing in scenarios-openai.json, plus the three hostile provider modes, and
# asserts the brain's claims against the stub's ground-truth JSONL (truth, not
# narration) and against files on disk.
#
# SAFETY (hard rules, enforced below):
# - never binds 7770 / 7779 / 17779 - only 7891-7894
# - never reads or writes ~/.neuron - HOME is redirected to a scratch dir
# - every process started here is killed on exit (trap) and proven with lsof
#
# The soul runs under `script -q /dev/null` so its stdout is a pty: El's
# println() uses puts(), which is FULLY buffered to a file, and the process is
# killed without flushing — the DRIFT lines would be invisible otherwise.
#
# Usage: ./run-lane-gate.sh [all|bridge|local|toolsoff|hostile]
# bridge = consent round-trip config (no workspace root -> write_file is
# "escalate" -> the loop suspends and the CLIENT executes the tool)
# local = workspace-root config (write_file is "reversible" + builtin ->
# the loop executes the tool in-process and runs to completion)
# toolsoff = supplementary: non-agentic lane against a base URL WITHOUT the
# /v1 suffix (the el-runtime provider chain appends
# /v1/chat/completions itself, unlike chat.el which appends only
# /chat/completions)
# hostile = black-hole / mid-body-drop / tool-pending-forever
#
# Env overrides: SOUL_BIN, STUB_PORT, SOUL_PORT, SOUL_PORT_B, RUN_ROOT
set -uo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
PHASES="${1:-all}"
SOUL_BIN="${SOUL_BIN:-/tmp/soul-oai2/soul-openai-tools}"
STUB_PORT="${STUB_PORT:-7891}"
SOUL_PORT="${SOUL_PORT:-7892}"
SOUL_PORT_B="${SOUL_PORT_B:-7893}"
RUN_ROOT="${RUN_ROOT:-/tmp/oa-lane-gate}"
STAMP="$(date +%Y%m%d-%H%M%S)"
RUN="$RUN_ROOT/$STAMP"
for p in "$STUB_PORT" "$SOUL_PORT" "$SOUL_PORT_B"; do
case "$p" in
7770|7779|17779) echo "FATAL: refusing production Neuron port $p"; exit 2;;
789[1-4]) ;;
*) echo "FATAL: port $p outside the allowed 7891-7894 range"; exit 2;;
esac
done
[ -x "$SOUL_BIN" ] || { echo "FATAL: soul binary not found/executable: $SOUL_BIN"; exit 2; }
mkdir -p "$RUN/home" "$RUN/ws-bridge" "$RUN/ws-local" "$RUN/ws-off" "$RUN/engram"
echo '{"nodes":[],"edges":[]}' > "$RUN/engram/snapshot.json"
DRV="$RUN/drv.py"
STUB_PID=""; SOUL_PID=""
cleanup() {
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
pkill -f "$SOUL_BIN" 2>/dev/null
[ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null
sleep 0.4
[ -n "$SOUL_PID" ] && kill -9 "$SOUL_PID" 2>/dev/null
[ -n "$STUB_PID" ] && kill -9 "$STUB_PID" 2>/dev/null
return 0
}
trap cleanup EXIT INT TERM
start_stub() { # $1 = mode, $2 = log path
local mode="$1" log="$2" args=""
[ "$mode" = "normal" ] && args="--scenarios $HERE/scenarios-openai.json"
# shellcheck disable=SC2086
python3 "$HERE/stub-openai.py" --port "$STUB_PORT" --mode "$mode" --log "$log" $args \
> "$RUN/stub-$mode.out" 2>&1 &
STUB_PID=$!
for _ in $(seq 1 50); do
curl -sf "http://127.0.0.1:$STUB_PORT/gate/health" >/dev/null 2>&1 && return 0
sleep 0.2
done
echo "FATAL: stub did not come up on $STUB_PORT"; cat "$RUN/stub-$mode.out"; exit 3
}
stop_stub() { [ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null; sleep 0.3; STUB_PID=""; }
start_soul() { # $1 = port, $2 = base url, $3 = soul log, $4 = agent root ("" = none)
local port="$1" base="$2" log="$3" root="$4"
script -q /dev/null \
env -u ANTHROPIC_API_KEY -u SOUL_API_KEY -u ENGRAM_URL -u ENGRAM_API_KEY \
-u NEURON_API_URL -u NEURON_TOKEN -u SOUL_LLM_PROVIDER -u SOUL_LLM_BASE_URL \
-u NEURON_LLM_1_URL -u NEURON_LLM_1_KEY -u SOUL_IDENTITY \
HOME="$RUN/home" PATH="$PATH" \
NEURON_PORT="$port" EL_HTTP_BIND_HOST=127.0.0.1 \
SOUL_ENGRAM_PATH="$RUN/engram/snapshot.json" \
SOUL_CGI_ID=ntn-test SOUL_PERSONA_NAME=Neuron \
NEURON_LLM_0_URL="$base" NEURON_LLM_0_FORMAT=openai NEURON_LLM_0_KEY=gate-test-key \
${root:+NEURON_AGENT_ROOT="$root"} \
"$SOUL_BIN" > "$log" 2>&1 &
SOUL_PID=$!
for _ in $(seq 1 100); do
curl -sf "http://127.0.0.1:$port/health" >/dev/null 2>&1 && return 0
sleep 0.2
done
echo "FATAL: soul did not come up on $port"; tail -20 "$log"; exit 3
}
stop_soul() {
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
pkill -f "$SOUL_BIN" 2>/dev/null
sleep 0.6; SOUL_PID=""
}
# ------------------------------------------------------------------ driver ----
cat > "$DRV" <<'PYEOF'
import json, os, sys, time, threading, urllib.request, urllib.error
CFG = json.load(open(sys.argv[1]))
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
WS = CFG["workspace"]
MODE = CFG["mode"] # bridge | local | toolsoff
SCEN = json.load(open(CFG["scenarios"]))
STUBLOG = CFG["stub_log"]
SOULLOG = CFG["soul_log"]
ONLY = CFG.get("classes") or list(SCEN["classes"].keys())
MAXHOPS = CFG.get("max_hops", 15)
OUT = CFG["out"]
# the chat-only class must be driven on the NON-agentic door: the agentic door
# always advertises tools, which is a 400 on that scenario by contract.
NON_AGENTIC = {"oa-tools-off"}
def http(method, url, obj=None, timeout=300):
data = None if obj is None else json.dumps(obj).encode()
req = urllib.request.Request(url, data=data, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
body = r.read().decode("utf-8", "replace")
st = r.status
except urllib.error.HTTPError as e:
body = e.read().decode("utf-8", "replace"); st = e.code
except Exception as e:
return -1, "TRANSPORT-ERROR: %r" % (e,), None
try:
return st, body, json.loads(body)
except ValueError:
return st, body, None
def fsize(p):
return os.path.getsize(p) if os.path.exists(p) else 0
def tail_from(path, off):
if not os.path.exists(path):
return "", off
with open(path, "rb") as f:
f.seek(off); chunk = f.read(); return chunk.decode("utf-8", "replace"), f.tell()
def stub_since(off):
"""Exact correlation: only the JSONL bytes appended during this phrasing."""
txt, noff = tail_from(STUBLOG, off)
recs = []
for line in txt.splitlines():
line = line.strip()
if line:
try: recs.append(json.loads(line))
except ValueError: pass
return recs, noff
def perform(name, ti):
"""Execute the bridged tool for real, like the desktop client would."""
if name in ("write_file", "edit_file"):
p = ti.get("path", "")
dest = p if os.path.isabs(p) else os.path.join(WS, p)
os.makedirs(os.path.dirname(dest) or WS, exist_ok=True)
body = ti.get("content", "")
with open(dest, "w") as f:
f.write(body)
return "wrote %s (%d bytes)" % (p, len(body.encode()))
return "ok"
class Poller(threading.Thread):
def __init__(self, sid):
super().__init__(daemon=True); self.sid = sid; self.snaps = []; self.stop = False
def run(self):
while not self.stop:
st, body, js = http("GET", SOUL + "/api/run-progress/" + self.sid, timeout=60)
if js and js.get("progress"):
if not self.snaps or self.snaps[-1] != js["progress"]:
self.snaps.append(js["progress"])
time.sleep(0.1)
def progress(sid):
_, _, pj = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
return (pj or {}).get("progress")
def run_phrasing(cname, ph):
st, body, js = http("POST", SOUL + "/api/sessions", {"title": ph["id"]}, timeout=60)
sid = (js or {}).get("id", "")
rec = {"class": cname, "phrasing": ph["id"], "session_id": sid, "legs": [],
"pendings": [], "progress_during": [], "progress_per_leg": [],
"progress_final": None, "soul_log": "", "stub": [], "http": [],
"agentic": cname not in NON_AGENTIC}
if not sid:
rec["fatal"] = "session create failed: %s %s" % (st, body[:300]); return rec
soff = fsize(SOULLOG); loff = fsize(STUBLOG)
t0 = time.time()
pol = Poller(sid); pol.start()
payload = {"message": ph["prompt"], "session_id": sid, "workspace_root": WS,
"agentic": rec["agentic"]}
if MODE == "local":
payload["agent_workspace_root"] = WS
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("chat_timeout", 240))
rec["http"].append(st)
rec["legs"].append(js if js is not None else body[:600])
rec["progress_per_leg"].append(progress(sid))
hops = 0
while isinstance(js, dict) and js.get("tool_pending") and hops < MAXHOPS:
rec["pendings"].append({"call_id": js.get("call_id"), "tool_name": js.get("tool_name"),
"tool_input": js.get("tool_input"), "risk_tier": js.get("risk_tier"),
"narration": js.get("narration"), "tools_used": js.get("tools_used")})
try:
eff = perform(js.get("tool_name", ""), js.get("tool_input") or {})
except Exception as e:
eff = "client error: %r" % (e,)
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
{"call_id": js.get("call_id"), "content": eff},
timeout=CFG.get("chat_timeout", 240))
rec["http"].append(st)
rec["legs"].append(js if js is not None else body[:600])
rec["progress_per_leg"].append(progress(sid))
hops += 1
pol.stop = True; time.sleep(0.3)
t1 = time.time()
rec["elapsed"] = round(t1 - t0, 2)
rec["progress_during"] = pol.snaps
rec["progress_final"] = progress(sid)
rec["soul_log"], _ = tail_from(SOULLOG, soff)
rec["stub"], _ = stub_since(loff)
rec["hops"] = hops
return rec
# ------------------------------------------------------------- assertions ----
def expected_calls(cname, first_only=False):
out = []
for step in SCEN["classes"][cname]["script"]:
for k, call in enumerate(step.get("tool_calls") or []):
if first_only and k > 0:
continue
out.append((call["name"], call["arguments"]))
return out
def final_text(cname):
for step in reversed(SCEN["classes"][cname]["script"]):
if step.get("text") and not step.get("tool_calls"):
return step["text"]
return None
def judge(rec):
cname = rec["class"]; ok = []; bad = []
last = rec["legs"][-1] if rec["legs"] else None
reply = last.get("reply") if isinstance(last, dict) else None
err = last.get("error") if isinstance(last, dict) else None
tools_used = last.get("tools_used") if isinstance(last, dict) else None
stub = rec["stub"]
scen_recs = [r for r in stub if r.get("kind") == "scenario"]
rejects = [r for r in stub if r.get("validation") != "ok"]
bg = [r for r in stub if r.get("kind") in ("wrong_path", "background")]
def wire_clean():
if rejects:
for r in rejects:
bad.append("stub REJECTED a request: [%s] %s"
% (r.get("validation"), r.get("validation_detail")))
else:
ok.append("stub ground truth: validation \"ok\" on all %d scenario leg(s), no "
"gate_echo_mismatch / gate_tool_call_shape / dialect-leak 400s"
% len(scen_recs))
if bg:
ok.append("NOTE background non-scenario request(s) in this window: %s"
% [(r.get("kind"), r.get("path"), r.get("http_status")) for r in bg])
if cname == "oa-plain":
wire_clean()
want = final_text(cname)
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
if tools_used == []: ok.append("tools_used == [] (no tool ran)")
else: bad.append("tools_used expected [] got %r" % (tools_used,))
if reply and ('"tool_calls"' in reply or '"function"' in reply or '"tool_use"' in reply):
bad.append("tool-call JSON leaked into the reply text")
else: ok.append("no tool-call JSON anywhere in the reply")
elif cname in ("oa-single-tool", "oa-torture", "oa-mission"):
wire_clean()
want = final_text(cname)
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
exp = expected_calls(cname)
wantnames = [n for n, _ in exp]
if tools_used == wantnames:
ok.append("tools_used == %r (carried across %d suspension(s))" % (wantnames, rec["hops"]))
else:
bad.append("tools_used expected %r got %r" % (wantnames, tools_used))
for name, args in exp:
p = args.get("path"); c = args.get("content")
dest = os.path.join(WS, p)
if not os.path.exists(dest):
bad.append("expected file missing on disk: %s" % dest); continue
got = open(dest, "rb").read()
if got == c.encode():
ok.append("%s on disk is byte-for-byte the issued payload (%d bytes)" % (p, len(got)))
else:
bad.append("%s content differs\n WANT %r\n GOT %r"
% (p, c[:300], got[:300].decode("utf-8", "replace")))
if MODE == "bridge":
for pend, (name, args) in zip(rec["pendings"], exp):
if pend["tool_input"] == args:
ok.append("tool_input for %s survived exactly ONE decode (deep-equal to the "
"issued arguments; no double-escaping)" % name)
else:
bad.append("tool_input != issued arguments for %s\n WANT %r\n GOT %r"
% (name, args, pend["tool_input"]))
if rec["pendings"] and all(p["risk_tier"] == "escalate" for p in rec["pendings"]):
ok.append("every write_file classified \"escalate\" and bridged for consent")
elif cname == "oa-parallel":
drift = [l.strip() for l in rec["soul_log"].splitlines() if "DRIFT: provider returned" in l]
if drift: ok.append("soul log: " + drift[0])
else: bad.append("no 'DRIFT: provider returned N parallel tool_calls' line in the soul log")
delivered = [r for r in stub if r.get("delivered", {}).get("tool_calls")]
if delivered and len(delivered[0]["delivered"]["tool_calls"]) == 2:
ok.append("stub delivered 2 parallel tool_calls in one response (ground truth)")
if MODE == "bridge":
if len(rec["pendings"]) == 1:
ok.append("exactly ONE call honored: %s" % rec["pendings"][0]["call_id"])
else:
bad.append("expected exactly 1 honored call, got %d" % len(rec["pendings"]))
pairing = [r for r in rejects if "gate_pairing" in str(r.get("validation_detail")) or
"tool_calls at end of thread" in str(r.get("validation_detail")) or
"not fully answered" in str(r.get("validation_detail"))]
for r in pairing:
ok.append("EXPECTED-BY-CONTRACT stub 400 on the unpaired echo: %s"
% r.get("validation_detail"))
other = [r for r in rejects if r not in pairing]
for r in other:
bad.append("unexpected stub rejection: [%s] %s"
% (r.get("validation"), r.get("validation_detail")))
if err and not reply:
ok.append("honest error envelope after the 400 (no fabricated answer): %r" % err)
elif reply == final_text(cname):
ok.append("final reply == scripted final text (both calls paired)")
else:
bad.append("neither an honest error nor the scripted final text: %r" % (last,))
elif cname == "oa-api-error":
if err and not reply:
ok.append("honest error envelope: error=%r reply=%r" % (err, reply))
else:
bad.append("expected an error envelope with an empty reply, got %r" % (last,))
delivered = [r["delivered"].get("api_error") for r in stub if r.get("delivered")]
ok.append("stub delivered api_error status(es): %r" % [d for d in delivered if d])
n = len([r for r in stub if r.get("kind") == "scenario"])
ok.append("provider hit %d time(s) - no retry storm" % n)
if reply:
bad.append("FABRICATED ANSWER: reply non-empty on a provider error")
elif cname == "oa-tools-off":
ok.append("stub records for this phrasing: %r"
% [{k: r.get(k) for k in ("kind", "path", "validation", "http_status")} for r in stub])
wrong = [r for r in stub if r.get("kind") == "wrong_path"]
matched = [r for r in stub if r.get("scenario_class") == cname]
if matched and not rejects:
ok.append("chat-only request reached /v1/chat/completions with NO tools offered")
want = final_text(cname)
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
if reply and ('"tool_calls"' in reply or '"function"' in reply):
bad.append("tool-call JSON leaked into the reply text")
else: ok.append("no tool-call JSON in the reply")
elif wrong:
bad.append("the non-agentic lane never reached the provider endpoint: stub saw "
"%s -> %s (the el-runtime provider chain appends /v1/chat/completions "
"to NEURON_LLM_0_URL, chat.el appends only /chat/completions)"
% (wrong[0]["path"], wrong[0]["http_status"]))
elif not stub:
bad.append("no request reached the stub at all")
else:
for r in rejects:
bad.append("stub REJECTED: [%s] %s" % (r.get("validation"), r.get("validation_detail")))
return ok, bad
def main():
results = []
for cname in ONLY:
for ph in SCEN["classes"][cname]["phrasings"]:
rec = run_phrasing(cname, ph)
ok, bad = judge(rec)
rec["ok"] = ok; rec["bad"] = bad
rec["verdict"] = "FAIL" if bad else "PASS"
results.append(rec)
print("=" * 78)
print("[%s] %s / %s (%.2fs, %d bridge hop(s), agentic=%s, mode=%s)"
% (rec["verdict"], cname, ph["id"], rec.get("elapsed", 0),
rec.get("hops", 0), rec["agentic"], MODE))
for l in ok: print(" ok " + l.replace("\n", "\n "))
for l in bad: print(" FAIL " + l.replace("\n", "\n "))
for i, leg in enumerate(rec["legs"]):
print(" leg%d envelope: %s" % (i, json.dumps(leg)[:430]))
for i, pr in enumerate(rec["progress_per_leg"]):
print(" run-progress after leg%d: %s" % (i, json.dumps(pr)[:380]))
if rec["progress_during"]:
print(" run-progress polled DURING (%d distinct snapshot(s)), last: %s"
% (len(rec["progress_during"]), json.dumps(rec["progress_during"][-1])[:300]))
for r in rec["stub"]:
print(" stub: kind=%s class=%s phrasing=%s step=%s validation=%s%s delivered=%s http=%s"
% (r.get("kind"), r.get("scenario_class"), r.get("phrasing"), r.get("step"),
r.get("validation"),
("(" + str(r.get("validation_detail")) + ")") if r.get("validation_detail") else "",
json.dumps(r.get("delivered")), r.get("http_status")))
if rec["soul_log"].strip():
for l in rec["soul_log"].splitlines():
if l.strip(): print(" soul: " + l.strip())
json.dump(results, open(OUT, "w"), indent=1)
npass = sum(1 for r in results if r["verdict"] == "PASS")
print("=" * 78)
print("PHASE %s: %d/%d PASS" % (MODE, npass, len(results)))
for r in results:
print(" %-6s %-16s %s" % (r["verdict"], r["class"], r["phrasing"]))
return 0 if npass == len(results) else 1
sys.exit(main())
PYEOF
# ------------------------------------------------------------- hostile drv ---
cat > "$RUN/hostile.py" <<'PYEOF'
import json, os, sys, time, urllib.request, urllib.error
CFG = json.load(open(sys.argv[1]))
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
def http(method, url, obj=None, timeout=400):
data = None if obj is None else json.dumps(obj).encode()
req = urllib.request.Request(url, data=data, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
b = r.read().decode("utf-8", "replace"); st = r.status
except urllib.error.HTTPError as e:
b = e.read().decode("utf-8", "replace"); st = e.code
except Exception as e:
return -1, "TRANSPORT-ERROR: %r" % (e,), None
try:
return st, b, json.loads(b)
except ValueError:
return st, b, None
mode = CFG["mode"]; wsmode = CFG["ws_mode"]; WS = CFG["workspace"]
_, _, js = http("POST", SOUL + "/api/sessions", {"title": "hostile-" + mode}, timeout=60)
sid = (js or {}).get("id", "")
payload = {"message": "oa-gate plain probe: hostile mode %s" % mode,
"agentic": True, "session_id": sid, "workspace_root": WS}
if wsmode == "local":
payload["agent_workspace_root"] = WS
t0 = time.time()
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("timeout", 400))
t_first = time.time() - t0
legs = [js if js is not None else body[:500]]
hops = 0
while isinstance(js, dict) and js.get("tool_pending") and hops < CFG.get("max_hops", 14):
ti = js.get("tool_input") or {}
p = ti.get("path", "x.md")
dest = p if os.path.isabs(p) else os.path.join(WS, p)
try: open(dest, "w").write(ti.get("content", ""))
except Exception: pass
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
{"call_id": js.get("call_id"), "content": "ok"},
timeout=CFG.get("timeout", 400))
legs.append(js if js is not None else body[:500]); hops += 1
el = time.time() - t0
_, _, stats = http("GET", STUB + "/gate/stats", timeout=30)
_, _, prog = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
fab = [l for l in legs if isinstance(l, dict) and l.get("reply")]
print("HOSTILE %s (ws_mode=%s)" % (mode, wsmode))
print(" first /api/chat POST returned after %.2fs; whole chain %.2fs; client bridge hops=%d; "
"stub chat_hits=%s" % (t_first, el, hops, (stats or {}).get("chat_hits")))
print(" first envelope : " + json.dumps(legs[0])[:430])
print(" final envelope : " + json.dumps(legs[-1])[:430])
print(" non-empty replies anywhere in the chain (fabrication check): %d" % len(fab))
print(" run-progress : " + json.dumps(prog)[:300])
json.dump({"mode": mode, "ws_mode": wsmode, "t_first": t_first, "elapsed": el, "hops": hops,
"chat_hits": (stats or {}).get("chat_hits"), "legs": legs, "progress": prog},
open(CFG["out"], "w"), indent=1)
PYEOF
# ------------------------------------------------------------------ phases ---
RC_BRIDGE=0; RC_LOCAL=0; RC_OFF=0
run_normal_phase() { # $1 = label, $2 = soul port, $3 = agent root, $4 = ws, $5 = base, $6 = classes json
local m="$1" port="$2" root="$3" ws="$4" base="$5" classes="$6"
echo; echo "############ PHASE: $m (soul :$port, NEURON_LLM_0_URL=$base) ############"
start_stub normal "$RUN/stub-$m.jsonl"
start_soul "$port" "$base" "$RUN/soul-$m.log" "$root"
cat > "$RUN/cfg-$m.json" <<JSON
{"soul_port": $port, "stub_port": $STUB_PORT, "workspace": "$ws", "mode": "$m",
"scenarios": "$HERE/scenarios-openai.json", "stub_log": "$RUN/stub-$m.jsonl",
"soul_log": "$RUN/soul-$m.log", "out": "$RUN/results-$m.json", "chat_timeout": 240,
"classes": $classes}
JSON
python3 "$DRV" "$RUN/cfg-$m.json"
local rc=$?
stop_soul; stop_stub
return $rc
}
if [ "$PHASES" = "all" ] || [ "$PHASES" = "bridge" ]; then
run_normal_phase bridge "$SOUL_PORT" "" "$RUN/ws-bridge" "http://127.0.0.1:$STUB_PORT/v1" null
RC_BRIDGE=$?
fi
if [ "$PHASES" = "all" ] || [ "$PHASES" = "local" ]; then
run_normal_phase local "$SOUL_PORT_B" "$RUN/ws-local" "$RUN/ws-local" "http://127.0.0.1:$STUB_PORT/v1" null
RC_LOCAL=$?
fi
if [ "$PHASES" = "all" ] || [ "$PHASES" = "toolsoff" ]; then
# supplementary: the el-runtime provider chain appends /v1/chat/completions itself,
# so the non-agentic door needs the base WITHOUT the /v1 suffix.
run_normal_phase toolsoff "$SOUL_PORT" "" "$RUN/ws-off" "http://127.0.0.1:$STUB_PORT" '["oa-tools-off","oa-plain"]'
RC_OFF=$?
fi
if [ "$PHASES" = "all" ] || [ "$PHASES" = "hostile" ]; then
echo; echo "############ PHASE: hostile ############"
for spec in "black-hole:bridge" "mid-body-drop:bridge" "tool-pending-forever:bridge" "tool-pending-forever:local"; do
mode="${spec%%:*}"; wsm="${spec##*:}"
echo; echo "---- hostile mode=$mode ws_mode=$wsm ----"
start_stub "$mode" "$RUN/stub-$mode-$wsm.jsonl"
if [ "$wsm" = "local" ]; then
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" "$RUN/ws-local"
else
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" ""
fi
cat > "$RUN/cfg-$mode-$wsm.json" <<JSON
{"soul_port": $SOUL_PORT, "stub_port": $STUB_PORT, "mode": "$mode", "ws_mode": "$wsm",
"workspace": "$RUN/ws-local", "out": "$RUN/hostile-$mode-$wsm.json", "timeout": 400}
JSON
python3 "$RUN/hostile.py" "$RUN/cfg-$mode-$wsm.json"
echo " soul log (llm/DRIFT/cap lines):"
grep -E "DRIFT|llm error|iteration cap|\[llm\]" "$RUN/soul-$mode-$wsm.log" | tail -8 | sed 's/^/ /'
stop_soul; stop_stub
done
fi
echo; echo "############ CLEANUP ############"
cleanup
sleep 0.5
echo "processes still matching the soul binary:"; pgrep -fl "$SOUL_BIN" || echo " (none)"
echo "processes still matching stub-openai.py:"; pgrep -fl "stub-openai.py" || echo " (none)"
echo "lsof on 7891-7894 after cleanup:"
lsof -nP -iTCP:7891 -iTCP:7892 -iTCP:7893 -iTCP:7894 2>/dev/null || echo " (no listeners - ports free)"
echo
echo "############ SUMMARY ############"
echo "run dir: $RUN"
echo "bridge rc=$RC_BRIDGE local rc=$RC_LOCAL toolsoff rc=$RC_OFF (0 = every class PASS)"
exit $(( RC_BRIDGE + RC_LOCAL + RC_OFF ))
+110
View File
@@ -0,0 +1,110 @@
{
"_comment": "OpenAI-dialect gate scenario contract (soul-openai-tools-v2). Single source of truth shared by stub-openai.py (scripted provider responses + request assertions), selftest.sh (stub self-verification), and the future brain-side gate driver. Same structure as gate9's scenarios.json: classes -> script + phrasings with markers; scripts are CLASS-level so assertions are behavioral, never pinned to a sentence. expect_request keys: require_tools, require_tool_choice, parallel_tool_calls (expected literal value; null = don't check), forbid_tools. defaults apply to every class unless overridden; steps may override with their own expect_request.",
"deadline_secs": 60,
"max_loop_iterations": 16,
"defaults": {
"expect_request": {
"require_tools": true,
"require_tool_choice": true,
"parallel_tool_calls": false
}
},
"classes": {
"oa-plain": {
"script": [
{ "text": "Plain OpenAI-lane answer (gate fixture): the mechanism, the main caveat, and the practical takeaway in three sentences. No tools were needed for this one, and the finish reason on the wire is stop, which the loop must treat as terminal." }
],
"phrasings": [
{ "id": "oa-plain-p1", "marker": "oa-gate plain probe", "prompt": "oa-gate plain probe: explain the fixture topic simply." },
{ "id": "oa-plain-p2", "marker": "oa-gate second plain", "prompt": "oa-gate second plain: another phrasing of the plain question." }
]
},
"oa-tools-off": {
"expect_request": {
"require_tools": false,
"forbid_tools": true,
"require_tool_choice": false,
"parallel_tool_calls": null
},
"script": [
{ "text": "Chat-only OpenAI-lane answer (gate fixture): this lane offered no tools and none were used; the reply is plain text with finish reason stop." }
],
"phrasings": [
{ "id": "oa-tools-off-p1", "marker": "oa-gate tools-off probe", "prompt": "oa-gate tools-off probe: plain chat with no tools offered." }
]
},
"oa-single-tool": {
"script": [
{ "text": "Step 1: writing the note.",
"tool_calls": [
{ "name": "write_file",
"arguments": { "path": "openai-single-note.md", "content": "# Note (gate fixture, OpenAI lane)\n\nDeterministic single-tool body.\n" } }
] },
{ "text": "All set - openai-single-note.md is written with the fixture body. Nothing else was needed for this one." }
],
"phrasings": [
{ "id": "oa-single-p1", "marker": "oa-gate single tool note", "prompt": "oa-gate single tool note: save the fixture note to a file." },
{ "id": "oa-single-p2", "marker": "oa-gate one file please", "prompt": "oa-gate one file please: write the fixture note file." }
]
},
"oa-torture": {
"script": [
{ "tool_calls": [
{ "name": "write_file",
"arguments": { "path": "torture-note.md", "content": "Line 1 has \"double quotes\", 'singles', and a mid-line backslash \\ here.\nLine 2\thas a tab, a literal \\n two-char sequence, and a Windows path C:\\temp\\new.txt.\nLine 3 unicode: naïve café — 日本語 ✓ 🚀\nLine 4 JSON-in-string: {\"k\": \"v\", \"arr\": [1, 2], \"s\": \"nested \\\"deep\\\" quotes\"}\nLine 5 ends with a lone backslash \\" } }
] },
{ "text": "Torture round-trip complete: the payload with nested quotes, backslashes, newlines, tabs, and unicode survived exactly one encode and one decode." }
],
"phrasings": [
{ "id": "oa-torture-p1", "marker": "oa-gate torture probe", "prompt": "oa-gate torture probe: write the escaping torture file." }
]
},
"oa-parallel": {
"script": [
{ "text": "Step 1: two writes at once (parallel probe).",
"tool_calls": [
{ "name": "write_file", "arguments": { "path": "parallel-a.md", "content": "Parallel A (gate fixture).\n" } },
{ "name": "write_file", "arguments": { "path": "parallel-b.md", "content": "Parallel B (gate fixture).\n" } }
] },
{ "text": "Parallel probe complete: both tool results arrived and were paired correctly. A brain that instead rejects the double call must do so cleanly - that outcome is asserted brain-side, not here." }
],
"phrasings": [
{ "id": "oa-parallel-p1", "marker": "oa-gate parallel probe", "prompt": "oa-gate parallel probe: run the two-write parallel case." }
]
},
"oa-mission": {
"script": [
{ "text": "Step 1: drafting part one.",
"tool_calls": [
{ "name": "write_file", "arguments": { "path": "mission-part-1.md", "content": "Mission part 1 (gate fixture).\n" } }
] },
{ "text": "Step 2: drafting part two.",
"tool_calls": [
{ "name": "write_file", "arguments": { "path": "mission-part-2.md", "content": "Mission part 2 (gate fixture).\n" } }
] },
{ "text": "Mission complete: mission-part-1.md and mission-part-2.md are written; the loop ran two tool rounds and finished cleanly with finish reason stop." }
],
"phrasings": [
{ "id": "oa-mission-p1", "marker": "oa-gate mission probe", "prompt": "oa-gate mission probe: run the two-round mission." }
]
},
"oa-api-error": {
"expect_request": {
"require_tools": false,
"require_tool_choice": false,
"parallel_tool_calls": null
},
"script": [],
"phrasings": [
{ "id": "oa-err-400", "marker": "oa-gate error four hundred", "prompt": "oa-gate error four hundred: trigger the injected failure.",
"script": [ { "api_error": { "status": 400, "type": "invalid_request_error", "message": "gate-injected 400: request rejected by fixture", "code": "gate_injected" } } ] },
{ "id": "oa-err-429", "marker": "oa-gate error rate limit", "prompt": "oa-gate error rate limit: trigger the injected failure.",
"script": [ { "api_error": { "status": 429, "type": "rate_limit_error", "message": "gate-injected 429: rate limited by fixture", "code": "rate_limit_exceeded" } } ] },
{ "id": "oa-err-500", "marker": "oa-gate error five hundred", "prompt": "oa-gate error five hundred: trigger the injected failure.",
"script": [ { "api_error": { "status": 500, "type": "server_error", "message": "gate-injected 500: internal fixture error", "code": "gate_injected" } } ] },
{ "id": "oa-err-503", "marker": "oa-gate error unavailable", "prompt": "oa-gate error unavailable: trigger the injected failure.",
"script": [ { "api_error": { "status": 503, "type": "server_error", "message": "gate-injected 503: fixture overloaded", "code": "gate_injected" } } ] }
]
}
}
}
+409
View File
@@ -0,0 +1,409 @@
#!/usr/bin/env bash
# selftest.sh - proves stub-openai.py before any brain code exists.
# Drives the stub with curl through every scenario (plain, tools-off,
# single tool round-trip, escaping torture, parallel double-call,
# two-round mission, injected API errors, background, overrun), every
# validation rejection (dialect leaks, pairing, echo round-trip, scenario
# expectations), and all three hostile modes. Exit 0 = green.
set -u
cd "$(dirname "$0")" || exit 1
PY=python3
TMP="$(mktemp -d)"
PIDS=()
cleanup() {
for p in "${PIDS[@]:-}"; do kill -9 "$p" >/dev/null 2>&1; done
rm -rf "$TMP"
}
trap cleanup EXIT
PASS=0; FAIL=0
ok() { printf 'ok - %s\n' "$1"; PASS=$((PASS+1)); }
bad() { printf 'FAIL - %s\n' "$1"; FAIL=$((FAIL+1)); }
check() { # check <name> <cmd...> - pass if cmd exits 0; show output on fail
local name="$1"; shift
local out
if out="$("$@" 2>&1)"; then ok "$name"
else bad "$name"; [ -n "$out" ] && printf '%s\n' "$out" | sed 's/^/ /' | head -8
fi
}
freeport() { "$PY" -c 'import socket;s=socket.socket();s.bind(("127.0.0.1",0));print(s.getsockname()[1]);s.close()'; }
waithealth() {
local p="$1" i
for i in $(seq 1 60); do
curl -sf "http://127.0.0.1:$p/gate/health" >/dev/null 2>&1 && return 0
sleep 0.1
done
echo "stub on :$p never became healthy"; return 1
}
post() { # post <port> <bodyfile> <respfile> [extra curl args...] -> echoes http code
local port="$1" body="$2" resp="$3"; shift 3
curl -s -o "$resp" -w '%{http_code}' -H 'content-type: application/json' \
"$@" --data-binary @"$body" "http://127.0.0.1:$port/v1/chat/completions"
}
# ---- embedded helper: builds OpenAI-dialect bodies, asserts on responses ----
cat > "$TMP/helpers.py" <<'PYEOF'
import copy, json, sys
TOOLS = [
{"type": "function", "function": {
"name": "write_file", "description": "Write content to a file on disk.",
"parameters": {"type": "object",
"properties": {"path": {"type": "string"},
"content": {"type": "string"}},
"required": ["path", "content"]}}},
{"type": "function", "function": {
"name": "read_file", "description": "Read contents of a file from disk.",
"parameters": {"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]}}},
]
def dump(obj, out):
json.dump(obj, open(out, "w"), ensure_ascii=False)
def base(prompt, tools=True):
b = {"model": "gate-openai-model", "max_tokens": 1024,
"messages": [
{"role": "system", "content": "You are Neuron (gate fixture)."},
{"role": "user", "content": prompt}]}
if tools:
b["tools"] = copy.deepcopy(TOOLS)
b["tool_choice"] = "auto"
b["parallel_tool_calls"] = False
return b
def cmd_plain(out, prompt):
dump(base(prompt), out)
def cmd_notools(out, prompt):
dump(base(prompt, tools=False), out)
def cmd_mut(out, prompt, mutation):
b = base(prompt)
if mutation == "no-tool-choice":
del b["tool_choice"]
elif mutation == "ptc-true":
b["parallel_tool_calls"] = True
elif mutation == "top-system":
b["system"] = "You are Neuron."
elif mutation == "anth-tools":
b["tools"] = [{"name": "write_file", "description": "x",
"input_schema": {"type": "object", "properties": {}}}]
elif mutation == "anth-block":
b["messages"][1] = {"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_x", "content": "hi"},
{"type": "text", "text": prompt}]}
else:
raise SystemExit("unknown mutation " + mutation)
dump(b, out)
def cmd_chain(out, prompt, variant, *resps):
"""Build the next leg: echo each response's assistant turn and answer its
tool calls. `variant` applies to the LAST response only:
ok | no-tool-turn | wrong-id | only-first | double-encode | object-args"""
b = base(prompt)
for idx, p in enumerate(resps):
last = idx == len(resps) - 1
msg = json.load(open(p))["choices"][0]["message"]
tcs = msg.get("tool_calls")
if not tcs:
b["messages"].append({"role": "assistant",
"content": msg.get("content")})
continue
v = variant if last else "ok"
asst = {"role": "assistant", "content": msg.get("content"),
"tool_calls": copy.deepcopy(tcs)}
if v == "double-encode":
for tc in asst["tool_calls"]:
tc["function"]["arguments"] = json.dumps(
tc["function"]["arguments"])
if v == "object-args":
for tc in asst["tool_calls"]:
tc["function"]["arguments"] = json.loads(
tc["function"]["arguments"])
b["messages"].append(asst)
if v == "no-tool-turn":
continue
use = tcs[:1] if v == "only-first" else tcs
for tc in use:
tid = "call_bogus_123" if v == "wrong-id" else tc["id"]
b["messages"].append({"role": "tool", "tool_call_id": tid,
"content": "{\"ok\":true,\"bytes\":42}"})
dump(b, out)
def cmd_chk(resp, expr):
r = json.load(open(resp))
if not eval(expr, {"r": r, "json": json, "len": len, "str": str,
"isinstance": isinstance, "any": any, "all": all,
"sorted": sorted}):
print("assertion failed:", expr)
print("resp:", json.dumps(r, ensure_ascii=False)[:400])
raise SystemExit(1)
def cmd_torture(resp, scen):
r = json.load(open(resp))
tc = r["choices"][0]["message"]["tool_calls"][0]
raw = tc["function"]["arguments"]
assert isinstance(raw, str), "arguments must be a JSON-encoded string"
got = json.loads(raw)
exp = json.load(open(scen))["classes"]["oa-torture"]["script"][0]["tool_calls"][0]["arguments"]
assert got == exp, "decoded arguments != scripted torture payload"
content = got["content"]
for needle in ['"', "\\", "\n", "\t", "日本語", "naïve", "🚀"]:
assert needle in content, "missing torture needle %r" % needle
def cmd_notjson(path):
data = open(path, "rb").read()
assert data, "file empty - no partial body arrived"
try:
json.loads(data.decode("utf-8", "replace"))
except ValueError:
return
raise SystemExit("partial body unexpectedly parsed as complete JSON")
def cmd_pending(*paths):
ids = []
for p in paths:
c = json.load(open(p))["choices"][0]
assert c["finish_reason"] == "tool_calls", c["finish_reason"]
tc = c["message"]["tool_calls"][0]
assert tc["function"]["name"] == "write_file"
json.loads(tc["function"]["arguments"]) # must decode
ids.append(tc["id"])
assert len(set(ids)) == len(ids), "call ids not distinct: %r" % ids
def cmd_logcheck(path):
recs = [json.loads(l) for l in open(path) if l.strip()]
seqs = [r["seq"] for r in recs]
assert seqs == sorted(seqs) and len(set(seqs)) == len(seqs), "seq not monotonic"
kinds = {}
for r in recs:
kinds[r["kind"]] = kinds.get(r["kind"], 0) + 1
assert kinds.get("scenario", 0) >= 10, "too few scenario records: %r" % kinds
assert kinds.get("background", 0) >= 1, "no background record"
assert kinds.get("overrun", 0) >= 1, "no overrun record"
rejected = [r for r in recs if r["validation"] == "rejected"]
assert len(rejected) >= 10, "too few rejected records: %d" % len(rejected)
assert any(r["delivered"].get("tool_calls") == ["write_file"]
for r in recs), "no single write_file ground truth"
assert any(r["delivered"].get("tool_calls") == ["write_file", "write_file"]
for r in recs), "no parallel ground truth"
def main():
fn = globals()["cmd_" + sys.argv[1].replace("-", "_")]
fn(*sys.argv[2:])
if __name__ == "__main__":
main()
PYEOF
mk() { "$PY" "$TMP/helpers.py" "$@"; }
echo "=== stub-openai selftest ==="
# ---- normal mode ------------------------------------------------------------
PORT="$(freeport)"
"$PY" stub-openai.py --port "$PORT" --scenarios scenarios-openai.json \
--log "$TMP/req.jsonl" >"$TMP/stub.out" 2>&1 &
PIDS+=($!); disown
check "stub starts and answers /gate/health" waithealth "$PORT"
# 1. plain completion
mk plain "$TMP/plain.json" "oa-gate plain probe: explain the fixture topic simply."
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_plain.json")"
check "plain: HTTP 200" test "$code" = "200"
check "plain: chat.completion envelope, finish stop, real content" mk chk "$TMP/r_plain.json" \
'r["object"]=="chat.completion" and r["choices"][0]["finish_reason"]=="stop" and isinstance(r["choices"][0]["message"]["content"],str) and len(r["choices"][0]["message"]["content"])>40'
# 2. tools-off lane (chat-only request accepted, tool-bearing request refused)
mk notools "$TMP/toolsoff.json" "oa-gate tools-off probe: plain chat with no tools offered."
code="$(post "$PORT" "$TMP/toolsoff.json" "$TMP/r_toolsoff.json")"
check "tools-off: chat-only request -> 200" test "$code" = "200"
mk plain "$TMP/toolsoff_bad.json" "oa-gate tools-off probe: plain chat with no tools offered."
code="$(post "$PORT" "$TMP/toolsoff_bad.json" "$TMP/r_toolsoff_bad.json")"
check "tools-off negative: offering tools -> 400 gate_expect" \
bash -c "test $code = 400"
check "tools-off negative: reason names gate_expect" mk chk "$TMP/r_toolsoff_bad.json" \
'r["error"]["code"]=="gate_expect"'
# 3. dialect-leak rejections (the loud-failure contract)
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_leak_hdr.json" -H 'anthropic-version: 2023-06-01')"
check "leak: anthropic-version header -> 400" test "$code" = "400"
check "leak: header reason names the leak" mk chk "$TMP/r_leak_hdr.json" \
'r["error"]["code"]=="gate_dialect_leak" and "anthropic-version" in r["error"]["message"]'
mk mut "$TMP/leak_tools.json" "oa-gate plain probe: explain the fixture topic simply." anth-tools
code="$(post "$PORT" "$TMP/leak_tools.json" "$TMP/r_leak_tools.json")"
check "leak: input_schema tools -> 400 gate_dialect_leak" bash -c \
"test $code = 400"
check "leak: input_schema reason" mk chk "$TMP/r_leak_tools.json" \
'r["error"]["code"]=="gate_dialect_leak" and "input_schema" in r["error"]["message"]'
mk mut "$TMP/leak_sys.json" "oa-gate plain probe: explain the fixture topic simply." top-system
code="$(post "$PORT" "$TMP/leak_sys.json" "$TMP/r_leak_sys.json")"
check "leak: top-level system -> 400" test "$code" = "400"
mk mut "$TMP/leak_block.json" "oa-gate plain probe: explain the fixture topic simply." anth-block
code="$(post "$PORT" "$TMP/leak_block.json" "$TMP/r_leak_block.json")"
check "leak: Anthropic tool_result content block -> 400" test "$code" = "400"
# 4. scenario request expectations
mk mut "$TMP/no_tc.json" "oa-gate plain probe: explain the fixture topic simply." no-tool-choice
code="$(post "$PORT" "$TMP/no_tc.json" "$TMP/r_no_tc.json")"
check "expect: missing tool_choice -> 400" test "$code" = "400"
mk mut "$TMP/ptc.json" "oa-gate plain probe: explain the fixture topic simply." ptc-true
code="$(post "$PORT" "$TMP/ptc.json" "$TMP/r_ptc.json")"
check "expect: parallel_tool_calls true -> 400 (ADR-0005 pin)" test "$code" = "400"
# 5. single tool round-trip
ST_PROMPT="oa-gate single tool note: save the fixture note to a file."
mk plain "$TMP/st1.json" "$ST_PROMPT"
code="$(post "$PORT" "$TMP/st1.json" "$TMP/r_st1.json")"
check "single-tool leg1: HTTP 200" test "$code" = "200"
check "single-tool leg1: one write_file call, finish tool_calls, string args" mk chk "$TMP/r_st1.json" \
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==1 and r["choices"][0]["message"]["tool_calls"][0]["type"]=="function" and r["choices"][0]["message"]["tool_calls"][0]["function"]["name"]=="write_file" and isinstance(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"],str) and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="openai-single-note.md"'
mk chain "$TMP/st2.json" "$ST_PROMPT" ok "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2.json" "$TMP/r_st2.json")"
check "single-tool leg2: echo + tool turn -> 200 final text" test "$code" = "200"
check "single-tool leg2: final names the file, finish stop" mk chk "$TMP/r_st2.json" \
'r["choices"][0]["finish_reason"]=="stop" and "openai-single-note.md" in r["choices"][0]["message"]["content"]'
mk chain "$TMP/st2_no.json" "$ST_PROMPT" no-tool-turn "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2_no.json" "$TMP/r_st2_no.json")"
check "single-tool negative: echo without tool turn -> 400 gate_pairing" \
bash -c "test $code = 400"
check "single-tool negative: pairing reason" mk chk "$TMP/r_st2_no.json" \
'r["error"]["code"]=="gate_pairing"'
mk chain "$TMP/st2_wrong.json" "$ST_PROMPT" wrong-id "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2_wrong.json" "$TMP/r_st2_wrong.json")"
check "single-tool negative: wrong tool_call_id -> 400" test "$code" = "400"
mk chain "$TMP/st2_obj.json" "$ST_PROMPT" object-args "$TMP/r_st1.json"
code="$(post "$PORT" "$TMP/st2_obj.json" "$TMP/r_st2_obj.json")"
check "single-tool negative: arguments echoed as object -> 400 shape" \
bash -c "test $code = 400"
check "single-tool negative: shape reason names STRING" mk chk "$TMP/r_st2_obj.json" \
'r["error"]["code"]=="gate_tool_call_shape" and "STRING" in r["error"]["message"]'
# 6. escaping torture (the two-escaper trap, spec section 6)
T_PROMPT="oa-gate torture probe: write the escaping torture file."
mk plain "$TMP/t1.json" "$T_PROMPT"
code="$(post "$PORT" "$TMP/t1.json" "$TMP/r_t1.json")"
check "torture leg1: HTTP 200" test "$code" = "200"
check "torture leg1: arguments decode to the exact nasty payload" \
mk torture "$TMP/r_t1.json" scenarios-openai.json
mk chain "$TMP/t2.json" "$T_PROMPT" ok "$TMP/r_t1.json"
code="$(post "$PORT" "$TMP/t2.json" "$TMP/r_t2.json")"
check "torture leg2: faithful echo -> 200 final" test "$code" = "200"
mk chain "$TMP/t2_dbl.json" "$T_PROMPT" double-encode "$TMP/r_t1.json"
code="$(post "$PORT" "$TMP/t2_dbl.json" "$TMP/r_t2_dbl.json")"
check "torture negative: double-encoded echo -> 400" test "$code" = "400"
check "torture negative: reason names the two-escaper trap" mk chk "$TMP/r_t2_dbl.json" \
'r["error"]["code"]=="gate_echo_mismatch" and "two-escaper" in r["error"]["message"]'
# 7. parallel double-call
P_PROMPT="oa-gate parallel probe: run the two-write parallel case."
mk plain "$TMP/p1.json" "$P_PROMPT"
code="$(post "$PORT" "$TMP/p1.json" "$TMP/r_p1.json")"
check "parallel leg1: TWO tool_calls, distinct ids" mk chk "$TMP/r_p1.json" \
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==2 and r["choices"][0]["message"]["tool_calls"][0]["id"]!=r["choices"][0]["message"]["tool_calls"][1]["id"]'
mk chain "$TMP/p2.json" "$P_PROMPT" ok "$TMP/r_p1.json"
code="$(post "$PORT" "$TMP/p2.json" "$TMP/r_p2.json")"
check "parallel leg2: both results -> 200 final" test "$code" = "200"
mk chain "$TMP/p2_one.json" "$P_PROMPT" only-first "$TMP/r_p1.json"
code="$(post "$PORT" "$TMP/p2_one.json" "$TMP/r_p2_one.json")"
check "parallel negative: answering only one call -> 400 pairing" test "$code" = "400"
# 8. two-round mission (loop continuation + step indexing)
M_PROMPT="oa-gate mission probe: run the two-round mission."
mk plain "$TMP/m1.json" "$M_PROMPT"
code="$(post "$PORT" "$TMP/m1.json" "$TMP/r_m1.json")"
check "mission leg1: part-1 tool call" mk chk "$TMP/r_m1.json" \
'json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-1.md"'
mk chain "$TMP/m2.json" "$M_PROMPT" ok "$TMP/r_m1.json"
code="$(post "$PORT" "$TMP/m2.json" "$TMP/r_m2.json")"
check "mission leg2: part-2 tool call (step indexed by assistant count)" mk chk "$TMP/r_m2.json" \
'r["choices"][0]["finish_reason"]=="tool_calls" and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-2.md"'
mk chain "$TMP/m3.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json"
code="$(post "$PORT" "$TMP/m3.json" "$TMP/r_m3.json")"
check "mission leg3: final text, finish stop" mk chk "$TMP/r_m3.json" \
'r["choices"][0]["finish_reason"]=="stop" and "Mission complete" in r["choices"][0]["message"]["content"]'
mk chain "$TMP/m4.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json" "$TMP/r_m3.json"
code="$(post "$PORT" "$TMP/m4.json" "$TMP/r_m4.json")"
check "mission overrun: past-script request -> GATE-SCRIPT-EXHAUSTED" mk chk "$TMP/r_m4.json" \
'r["choices"][0]["message"]["content"].startswith("GATE-SCRIPT-EXHAUSTED")'
# 9. injected API errors (OpenAI error envelope)
for want in 400 429 500 503; do
case "$want" in
400) marker="four hundred";; 429) marker="rate limit";;
500) marker="five hundred";; 503) marker="unavailable";;
esac
mk plain "$TMP/e_$want.json" "oa-gate error $marker: trigger the injected failure."
code="$(post "$PORT" "$TMP/e_$want.json" "$TMP/r_e_$want.json")"
check "api-error $want: status returned" test "$code" = "$want"
check "api-error $want: OpenAI error envelope" mk chk "$TMP/r_e_$want.json" \
'isinstance(r["error"]["message"],str) and "gate-injected" in r["error"]["message"] and isinstance(r["error"]["type"],str)'
done
# 10. background (unmatched) request
mk plain "$TMP/bg.json" "hello there, just a boot probe with no marker"
code="$(post "$PORT" "$TMP/bg.json" "$TMP/r_bg.json")"
check "background: unmatched prompt -> benign ok" mk chk "$TMP/r_bg.json" \
'r["choices"][0]["message"]["content"]=="ok"'
# 11. ground-truth log invariants
check "ground-truth JSONL log invariants" mk logcheck "$TMP/req.jsonl"
# 12. production-port refusal
rc=0
"$PY" stub-openai.py --port 7770 --scenarios scenarios-openai.json \
--log "$TMP/never.jsonl" >/dev/null 2>&1 || rc=$?
check "refuses production port 7770" test "$rc" -ne 0
# ---- hostile mode: black-hole ----------------------------------------------
BH="$(freeport)"
"$PY" stub-openai.py --port "$BH" --log "$TMP/bh.jsonl" --mode black-hole \
>/dev/null 2>&1 &
PIDS+=($!); disown
check "black-hole: healthy" waithealth "$BH"
rc=0
curl -s -o /dev/null --max-time 3 -H 'content-type: application/json' \
--data-binary @"$TMP/plain.json" \
"http://127.0.0.1:$BH/v1/chat/completions" || rc=$?
check "black-hole: client times out (curl rc 28)" test "$rc" -eq 28
check "black-hole: health still answers during the hang" \
curl -sf --max-time 2 "http://127.0.0.1:$BH/gate/health"
# ---- hostile mode: mid-body-drop -------------------------------------------
MD="$(freeport)"
"$PY" stub-openai.py --port "$MD" --log "$TMP/md.jsonl" --mode mid-body-drop \
>/dev/null 2>&1 &
PIDS+=($!); disown
check "mid-body-drop: healthy" waithealth "$MD"
rc=0
curl -s --max-time 5 -o "$TMP/half.json" -H 'content-type: application/json' \
--data-binary @"$TMP/plain.json" \
"http://127.0.0.1:$MD/v1/chat/completions" || rc=$?
check "mid-body-drop: transfer fails (curl rc $rc)" test "$rc" -ne 0
check "mid-body-drop: partial body is not parseable JSON" mk notjson "$TMP/half.json"
# ---- hostile mode: tool-pending-forever ------------------------------------
TP="$(freeport)"
"$PY" stub-openai.py --port "$TP" --log "$TMP/tp.jsonl" \
--mode tool-pending-forever >/dev/null 2>&1 &
PIDS+=($!); disown
check "tool-pending-forever: healthy" waithealth "$TP"
for i in 1 2 3; do
code="$(post "$TP" "$TMP/plain.json" "$TMP/r_tp$i.json")"
check "tool-pending-forever: request $i -> 200" test "$code" = "200"
done
check "tool-pending-forever: three FRESH tool_calls, distinct ids" \
mk pending "$TMP/r_tp1.json" "$TMP/r_tp2.json" "$TMP/r_tp3.json"
check "tool-pending-forever: /gate/stats counts 3 chat hits" \
bash -c "curl -sf http://127.0.0.1:$TP/gate/stats | grep -q '\"chat_hits\": 3'"
# ---- summary ----------------------------------------------------------------
echo
echo "selftest: $PASS passed, $FAIL failed"
if [ "$FAIL" -ne 0 ]; then
echo "SELFTEST RED"
exit 1
fi
echo "SELFTEST GREEN (stub-openai gate scaffolding verified)"
+652
View File
@@ -0,0 +1,652 @@
#!/usr/bin/env python3
"""stub-openai.py - deterministic local stand-in for an OpenAI-format
/v1/chat/completions provider, for the soul-openai-tools-v2 gate
(docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md). No API key, no network,
no model.
Sibling of gate9's stub-llm.py (Anthropic dialect, _wt-beta-round9/scripts/
gate9/): same scenario mechanism (marker matching, assistant-count step
indexing, ground-truth JSONL log, prod-port refusal), different wire.
Staging home is tests/gate-openai/ in _wt-openai-tools; folds into
scripts/gate9/ after round 9 merges (see README.md).
WHAT IT DOES
* Serves POST /v1/chat/completions on 127.0.0.1 only (OpenAI dialect).
* VALIDATES every request - this is the gate's discriminator, built
BEFORE the brain-side El code exists so dialect leakage fails loudly:
- Anthropic tells are 400 code=gate_dialect_leak: `anthropic-version`
header; top-level `system` / `stop_sequences` / `max_tokens_to_sample`
/ `anthropic_version`; `input_schema` inside a tool entry; Anthropic
content blocks (tool_use / tool_result / server_tool_use / ...).
- tools[] must be OpenAI-shaped {type:"function", function:{name,
description, parameters}} with unique names -> 400 gate_tools_shape.
- assistant tool_calls echoes must be {id, type:"function",
function:{name, arguments:<JSON-encoded STRING>}}; a decoded-object
`arguments` is a wire bug -> 400 gate_tool_call_shape.
- every assistant tool_calls turn must be answered by role:"tool"
messages covering EVERY tool_call_id, immediately following;
unknown / duplicate / missing ids -> 400 gate_pairing.
- echoed `arguments` for gate-issued call ids (call_gate_*) are
recomputed from the script and compared after ONE json decode ->
400 gate_echo_mismatch. This is the two-escaper-trap discriminator
named in the spec's security model (section 6).
- scenario-level request expectations from scenarios-openai.json
(tools offered, OpenAI-shaped tool_choice, parallel_tool_calls
pinned false per ADR-0005) -> 400 gate_expect.
* Answers with SCRIPTED responses: plain text (finish_reason "stop"),
tool calls (finish_reason "tool_calls", arguments JSON-encoded, incl. a
nested-quote/escaping torture payload and a parallel two-call case), and
API-error injection (OpenAI error envelope). Scenario is selected by
scanning user-message text (newest first) for a registered marker
substring; the step index is the number of assistant messages already in
the request (stateless replay - resumes index correctly by construction).
* Writes a ground-truth JSONL log (--log): one record per request with the
validation verdict, matched scenario/step, and exactly which tool calls
were delivered. Gate assertions compare the brain's claims against THIS
log - truth, not narration.
* Unmatched requests (boot probes, awareness chatter) get a benign "ok"
text response, logged kind=background, never counted as ground truth.
* HOSTILE MODES (--mode) on the same file:
black-hole accept + read the request, never respond;
mid-body-drop send half a JSON body, then abort the socket;
tool-pending-forever every request gets a FRESH tool_call
(finish_reason "tool_calls"), forever - tests
the agentic loop's iteration cap; count the
brain's round-trips via GET /gate/stats.
usage: stub-openai.py --port P --scenarios scenarios-openai.json \
--log requests.jsonl [--mode MODE]
Listens on 127.0.0.1 only. Refuses production ports 7770/7779/17779.
"""
import argparse
import itertools
import json
import socket
import struct
import threading
import time
import uuid
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
STATE = {"scenarios": None, "log_path": None, "lock": threading.Lock(),
"seq": 0, "mode": "normal", "chat_hits": 0}
_PENDING_SEQ = itertools.count(1)
ANTHROPIC_TOP_KEYS = ("system", "stop_sequences", "max_tokens_to_sample",
"anthropic_version")
ANTHROPIC_BLOCK_TYPES = {"tool_use", "tool_result", "server_tool_use",
"web_search_tool_result", "thinking",
"redacted_thinking"}
DEFAULT_EXPECT = {"require_tools": True, "require_tool_choice": True,
"parallel_tool_calls": False, "forbid_tools": False}
# ---------------------------------------------------------------- loading ----
def load_scenarios(path):
cfg = json.load(open(path))
defaults = dict(DEFAULT_EXPECT)
defaults.update(cfg.get("defaults", {}).get("expect_request", {}))
marker_map = [] # (marker_lower, cname, pid)
scripts = {} # cname or cname/pid -> expanded script
pid_map = {} # pid -> cname (for call_gate_* id -> script lookup)
expects = {} # cname -> merged expect_request
for cname, cls in cfg["classes"].items():
scripts[cname] = expand_script(cls.get("script", []))
exp = dict(defaults)
exp.update(cls.get("expect_request", {}))
expects[cname] = exp
for ph in cls["phrasings"]:
if ph.get("script") is not None:
scripts[cname + "/" + ph["id"]] = expand_script(ph["script"])
marker_map.append((ph["marker"].lower(), cname, ph["id"]))
pid_map[ph["id"]] = cname
return {"cfg": cfg, "marker_map": marker_map, "scripts": scripts,
"pid_map": pid_map, "expects": expects}
def expand_script(script):
"""Same repeat-expansion contract as gate9's stub-llm.py ({N}/{NN})."""
out = []
for step in script:
if "repeat" in step:
for n in range(1, step["repeat"] + 1):
t = {k: v for k, v in step.items() if k != "repeat"}
out.append(json.loads(json.dumps(t)
.replace("{NN}", "%02d" % n)
.replace("{N}", str(n))))
else:
out.append(step)
return out
# ------------------------------------------------------------- validation ----
def _rej(message, code):
return {"status": 400, "message": message, "code": code}
def validate_dialect(headers, req):
"""Universal checks - run on EVERY request, scenario-matched or not.
Anything Anthropic-shaped on this lane means the brain's translator
leaked; the whole point is that it fails loudly, here, with a reason."""
if headers.get("anthropic-version"):
return _rej("anthropic-version header on the OpenAI lane: this "
"request was built by the Anthropic dialect path",
"gate_dialect_leak")
for k in ANTHROPIC_TOP_KEYS:
if k in req:
return _rej("top-level `%s` is Anthropic dialect; the OpenAI "
"dialect has no such field (system prompt goes in "
"messages[0])" % k, "gate_dialect_leak")
tools = req.get("tools")
if tools is not None:
if not isinstance(tools, list):
return _rej("`tools` must be an array", "gate_tools_shape")
names = []
for i, t in enumerate(tools):
if not isinstance(t, dict):
return _rej("tools[%d] is not an object" % i,
"gate_tools_shape")
if "input_schema" in t or (isinstance(t.get("function"), dict)
and "input_schema" in t["function"]):
return _rej("tools[%d] carries `input_schema` (Anthropic "
"dialect); OpenAI dialect wants "
"function.parameters" % i, "gate_dialect_leak")
if t.get("type") != "function":
return _rej("tools[%d].type must be \"function\", got %r"
% (i, t.get("type")), "gate_tools_shape")
fn = t.get("function")
if not isinstance(fn, dict):
return _rej("tools[%d].function missing" % i,
"gate_tools_shape")
if not isinstance(fn.get("name"), str) or not fn["name"]:
return _rej("tools[%d].function.name missing/empty" % i,
"gate_tools_shape")
if not isinstance(fn.get("description"), str) or not fn["description"]:
return _rej("tools[%d].function.description missing/empty" % i,
"gate_tools_shape")
if not isinstance(fn.get("parameters"), dict):
return _rej("tools[%d].function.parameters missing (JSON "
"Schema object expected)" % i, "gate_tools_shape")
names.append(fn["name"])
if len(names) != len(set(names)):
return _rej("tools: tool names must be unique", "gate_tools_shape")
msgs = req.get("messages")
if not isinstance(msgs, list) or not msgs:
return _rej("`messages` must be a non-empty array",
"gate_messages_shape")
for i, m in enumerate(msgs):
if not isinstance(m, dict):
return _rej("messages[%d] is not an object" % i,
"gate_messages_shape")
c = m.get("content")
if isinstance(c, list):
for j, b in enumerate(c):
if isinstance(b, dict) and b.get("type") in ANTHROPIC_BLOCK_TYPES:
return _rej("messages[%d].content[%d] is an Anthropic "
"`%s` block; the OpenAI dialect uses "
"tool_calls / role:\"tool\" messages"
% (i, j, b.get("type")), "gate_dialect_leak")
if m.get("role") == "tool":
if not isinstance(m.get("tool_call_id"), str) or not m["tool_call_id"]:
return _rej("messages[%d]: role \"tool\" requires a "
"`tool_call_id`" % i, "gate_messages_shape")
if "content" not in m:
return _rej("messages[%d]: role \"tool\" requires `content`"
% i, "gate_messages_shape")
if m.get("role") == "assistant" and m.get("tool_calls") is not None:
tcs = m["tool_calls"]
if not isinstance(tcs, list) or not tcs:
return _rej("messages[%d].tool_calls must be a non-empty "
"array" % i, "gate_tool_call_shape")
for j, tc in enumerate(tcs):
if not isinstance(tc, dict) or tc.get("type") != "function":
return _rej("messages[%d].tool_calls[%d].type must be "
"\"function\"" % (i, j), "gate_tool_call_shape")
if not isinstance(tc.get("id"), str) or not tc["id"]:
return _rej("messages[%d].tool_calls[%d].id missing"
% (i, j), "gate_tool_call_shape")
fn = tc.get("function")
if not isinstance(fn, dict) or not isinstance(fn.get("name"), str):
return _rej("messages[%d].tool_calls[%d].function.name "
"missing" % (i, j), "gate_tool_call_shape")
if not isinstance(fn.get("arguments"), str):
return _rej("messages[%d].tool_calls[%d].function."
"arguments must be a JSON-encoded STRING, "
"got %s" % (i, j,
type(fn.get("arguments")).__name__),
"gate_tool_call_shape")
return None
def validate_pairing(msgs):
"""OpenAI pairing rule: every assistant tool_calls turn must be followed
immediately by role:"tool" messages answering every tool_call_id."""
open_ids, open_at = set(), None
for i, m in enumerate(msgs):
role = m.get("role")
if role == "tool":
tid = m.get("tool_call_id")
if open_at is None:
return _rej("messages[%d]: role \"tool\" message with no "
"preceding assistant tool_calls turn "
"(tool_call_id=%s)" % (i, tid), "gate_pairing")
if tid not in open_ids:
return _rej("messages[%d]: tool message answers unknown or "
"already-answered tool_call_id %s" % (i, tid),
"gate_pairing")
open_ids.discard(tid)
continue
if open_ids:
return _rej("messages[%d]: assistant tool_calls not fully "
"answered before messages[%d]; missing tool "
"responses for: %s" % (open_at, i, sorted(open_ids)),
"gate_pairing")
open_ids, open_at = set(), None
if role == "assistant" and m.get("tool_calls"):
ids = [tc.get("id") for tc in m["tool_calls"]]
open_ids, open_at = set(ids), i
if open_ids:
return _rej("messages[%d]: assistant tool_calls at end of thread "
"without tool responses for: %s"
% (open_at, sorted(open_ids)), "gate_pairing")
return None
def validate_echo_args(msgs, loaded):
"""Ground-truth round-trip check: for every echoed gate-issued call id,
recompute the arguments this stub originally sent from the script and
require one json decode to reproduce them exactly. Catches the
two-escaper trap (spec section 6) deterministically."""
if not loaded:
return None
for i, m in enumerate(msgs):
if m.get("role") != "assistant":
continue
for tc in m.get("tool_calls") or []:
tid = tc.get("id", "")
if not tid.startswith("call_gate_"):
continue
rest = tid[len("call_gate_"):]
try:
pid, s_part, k_part = rest.rsplit("_", 2)
step_idx, k = int(s_part[1:]), int(k_part)
except (ValueError, IndexError):
continue
cname = loaded["pid_map"].get(pid)
if cname is None:
continue
script = (loaded["scripts"].get(cname + "/" + pid)
or loaded["scripts"].get(cname) or [])
if step_idx >= len(script):
continue
calls = script[step_idx].get("tool_calls") or []
if k >= len(calls):
continue
expected = calls[k]
fn = tc.get("function") or {}
if fn.get("name") != expected["name"]:
return _rej("messages[%d]: echoed tool name %r != issued %r "
"for %s" % (i, fn.get("name"), expected["name"],
tid), "gate_echo_mismatch")
try:
got = json.loads(fn.get("arguments", ""))
except ValueError:
return _rej("messages[%d]: echoed arguments for %s are not "
"valid JSON after one decode (truncated or "
"half-escaped?)" % (i, tid), "gate_echo_mismatch")
if got != expected["arguments"]:
hint = (" (decoded to a string, not an object: "
"double-encoded - the two-escaper trap)"
if isinstance(got, str) else "")
return _rej("messages[%d]: echoed arguments for %s do not "
"round-trip to the issued payload%s"
% (i, tid, hint), "gate_echo_mismatch")
return None
def validate_expect(req, exp):
"""Scenario-level request expectations (scenarios-openai.json)."""
tools = req.get("tools") or []
if exp.get("forbid_tools") and tools:
return _rej("this scenario is chat-only: no `tools` may be offered "
"on it", "gate_expect")
if exp.get("require_tools") and not tools:
return _rej("scenario expects a `tools` array to be offered (the "
"agentic lane must advertise its tools)", "gate_expect")
if exp.get("require_tool_choice"):
tc = req.get("tool_choice")
ok = tc in ("auto", "none", "required") or (
isinstance(tc, dict) and tc.get("type") == "function"
and isinstance(tc.get("function"), dict)
and tc["function"].get("name"))
if not ok:
return _rej("scenario expects an OpenAI-shaped `tool_choice`, "
"got %r" % (tc,), "gate_expect")
want_ptc = exp.get("parallel_tool_calls", None)
if want_ptc is not None:
if "parallel_tool_calls" not in req:
return _rej("scenario expects explicit `parallel_tool_calls` "
"(ADR-0005: must be pinned false on the wire)",
"gate_expect")
if req["parallel_tool_calls"] != want_ptc:
return _rej("scenario expects parallel_tool_calls=%s, got %s"
% (json.dumps(want_ptc),
json.dumps(req["parallel_tool_calls"])),
"gate_expect")
return None
# --------------------------------------------------------- scenario match ----
def extract_user_texts_newest_first(msgs):
texts = []
for m in reversed(msgs):
if not isinstance(m, dict) or m.get("role") != "user":
continue
c = m.get("content")
if isinstance(c, str):
texts.append(c)
elif isinstance(c, list):
for b in c:
if isinstance(b, dict) and b.get("type") == "text":
texts.append(b.get("text", ""))
return texts
def match_scenario(loaded, msgs):
for text in extract_user_texts_newest_first(msgs):
tl = text.lower()
for marker, cname, pid in loaded["marker_map"]:
if marker in tl:
return cname, pid
return None, None
# ------------------------------------------------------------- rendering ----
def completion_envelope(msg, finish, model, usage=(100, 100)):
return {"id": "chatcmpl-gate-" + uuid.uuid4().hex[:12],
"object": "chat.completion", "created": int(time.time()),
"model": model,
"choices": [{"index": 0, "message": msg,
"finish_reason": finish, "logprobs": None}],
"usage": {"prompt_tokens": usage[0],
"completion_tokens": usage[1],
"total_tokens": usage[0] + usage[1]}}
def text_completion(text, model):
return completion_envelope({"role": "assistant", "content": text},
"stop", model, usage=(1, 1))
def pending_body(seq, model):
args = json.dumps({"path": "never-%04d.md" % seq,
"content": "this run never completes"})
msg = {"role": "assistant", "content": None,
"tool_calls": [{"id": "call_hostile_pending_%04d" % seq,
"type": "function",
"function": {"name": "write_file",
"arguments": args}}]}
return completion_envelope(msg, "tool_calls", model, usage=(1, 1))
def render_step(step, cname, pid, step_idx, model):
"""Returns (http_status, body_dict, delivered) - delivered is ground
truth for the JSONL log."""
delivered = {"tool_calls": [], "finish_reason": None, "api_error": None}
if "api_error" in step:
e = step["api_error"]
delivered["api_error"] = e["status"]
return (e["status"],
{"error": {"message": e["message"],
"type": e.get("type", "server_error"),
"param": None, "code": e.get("code")}},
delivered)
msg = {"role": "assistant"}
finish = "stop"
if step.get("tool_calls"):
tcs = []
for k, call in enumerate(step["tool_calls"]):
tid = "call_gate_%s_s%d_%d" % (pid, step_idx, k)
tcs.append({"id": tid, "type": "function",
"function": {"name": call["name"],
"arguments": json.dumps(
call["arguments"],
ensure_ascii=False)}})
delivered["tool_calls"].append(call["name"])
msg["tool_calls"] = tcs
msg["content"] = step.get("text") # null when no narration, like real
finish = "tool_calls"
else:
msg["content"] = step["text"]
delivered["finish_reason"] = finish
return 200, completion_envelope(msg, finish, model), delivered
# ------------------------------------------------------------------ log ------
def log_record(rec):
with STATE["lock"]:
STATE["seq"] += 1
rec["seq"] = STATE["seq"]
with open(STATE["log_path"], "a") as f:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
# ---------------------------------------------------------------- server -----
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def _send_json(self, status, obj):
body = json.dumps(obj, ensure_ascii=False).encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def _send_error(self, verdict):
self._send_json(verdict["status"],
{"error": {"message": verdict["message"],
"type": "invalid_request_error",
"param": None, "code": verdict["code"]}})
def _drop_mid_body(self):
"""Valid 200 headers, half the promised body, then a socket abort
(same SO_LINGER teardown as gate9's mid-body-drop-brain.py)."""
full = json.dumps(text_completion(
"This reply will never finish arriving because the connection "
"dies in the middle of the body, which is exactly the point of "
"this hostile fixture.", "hostile-mid-drop")).encode("utf-8")
half = full[: len(full) // 2]
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(full))) # promises more
self.end_headers()
self.wfile.write(half)
self.wfile.flush()
try:
self.connection.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER,
struct.pack("ii", 1, 0))
self.connection.shutdown(socket.SHUT_RDWR)
except OSError:
pass
self.close_connection = True
def do_GET(self):
path = self.path.split("?")[0]
if path == "/gate/health":
self._send_json(200, {"ok": True, "mode": STATE["mode"]})
elif path == "/gate/stats":
with STATE["lock"]:
self._send_json(200, {"mode": STATE["mode"],
"chat_hits": STATE["chat_hits"]})
else:
self._send_json(404, {"error": {"message": "not found",
"type": "invalid_request_error",
"param": None,
"code": "unknown_route"}})
def do_POST(self):
n = int(self.headers.get("Content-Length") or 0)
raw = self.rfile.read(n)
mode = STATE["mode"]
rec = {"ts": time.time(), "path": self.path, "mode": mode,
"kind": "background", "scenario_class": None, "phrasing": None,
"step": None, "n_messages": 0, "n_assistant": 0,
"validation": "ok", "validation_detail": None,
"delivered": {"tool_calls": [], "finish_reason": None,
"api_error": None},
"http_status": 200}
if self.path.split("?")[0] != "/v1/chat/completions":
rec.update(kind="wrong_path", http_status=404)
log_record(rec)
self._send_json(404, {"error": {
"message": "no such route: %s" % self.path,
"type": "invalid_request_error", "param": None,
"code": "unknown_route"}})
return
with STATE["lock"]:
STATE["chat_hits"] += 1
# ---- hostile modes: behavior first, no validation ----------------
if mode == "black-hole":
rec.update(kind="hostile", http_status=None)
log_record(rec)
threading.Event().wait() # hold the socket open forever
return
if mode == "mid-body-drop":
rec.update(kind="hostile", http_status=200)
log_record(rec)
self._drop_mid_body()
return
if mode == "tool-pending-forever":
seq = next(_PENDING_SEQ)
rec.update(kind="hostile",
delivered={"tool_calls": ["write_file"],
"finish_reason": "tool_calls",
"api_error": None})
log_record(rec)
self._send_json(200, pending_body(seq, "gate-openai-model"))
return
# ---- normal mode -------------------------------------------------
try:
req = json.loads(raw)
except ValueError as exc:
# DIAGNOSTIC CAPTURE (2026-08-06): an unparseable body used to be recorded as
# a bare "bad_json" with the bytes thrown away, which made an intermittent
# failure impossible to root-cause — you cannot fix what you did not keep.
# Dump the raw body next to the log, and record exactly where the parser gave
# up plus the offending byte, so one occurrence is enough to diagnose.
dump_path = "%s.badbody.%s" % (STATE.get("log_path", "/tmp/stub-openai"),
rec.get("seq", "x"))
try:
data = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
with open(dump_path, "wb") as fh:
fh.write(data)
except Exception as dump_exc:
dump_path = "(dump failed: %s)" % dump_exc
pos = getattr(exc, "pos", None)
near = ""
byte_repr = ""
if isinstance(pos, int):
blob = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
near = blob[max(0, pos - 60):pos + 60].decode("utf-8", "replace")
if 0 <= pos < len(blob):
byte_repr = "0x%02x" % blob[pos]
rec.update(kind="bad_json", validation="rejected",
validation_detail="request body is not valid JSON: %s" % exc,
http_status=400, raw_len=len(raw), raw_dump=dump_path,
err_pos=pos, err_byte=byte_repr, err_near=near)
log_record(rec)
self._send_error(_rej("request body is not valid JSON",
"bad_json"))
return
msgs = req.get("messages") or []
rec["n_messages"] = len(msgs)
rec["n_assistant"] = sum(1 for m in msgs if isinstance(m, dict)
and m.get("role") == "assistant")
loaded = STATE["scenarios"]
cname, pid = match_scenario(loaded, msgs)
if cname:
rec.update(kind="scenario", scenario_class=cname, phrasing=pid)
# Wire-level validation runs for EVERY request, scenario or not.
verdict = (validate_dialect(self.headers, req)
or validate_pairing([m for m in msgs
if isinstance(m, dict)])
or validate_echo_args(msgs, loaded))
if verdict:
rec.update(validation="rejected",
validation_detail=verdict["message"],
http_status=verdict["status"])
log_record(rec)
self._send_error(verdict)
return
model = req.get("model", "gate-openai-model")
if not cname:
log_record(rec)
self._send_json(200, text_completion("ok", model))
return
script = (loaded["scripts"].get(cname + "/" + pid)
or loaded["scripts"][cname])
step_idx = rec["n_assistant"]
if step_idx >= len(script):
rec.update(kind="overrun", step=step_idx)
log_record(rec)
self._send_json(200, text_completion(
"GATE-SCRIPT-EXHAUSTED %s step %d" % (pid, step_idx), model))
return
step = script[step_idx]
exp = dict(loaded["expects"][cname])
exp.update(step.get("expect_request", {}))
verdict = validate_expect(req, exp)
if verdict:
rec.update(step=step_idx, validation="rejected",
validation_detail=verdict["message"],
http_status=verdict["status"])
log_record(rec)
self._send_error(verdict)
return
status, body, delivered = render_step(step, cname, pid, step_idx,
model)
rec.update(step=step_idx, delivered=delivered, http_status=status)
log_record(rec)
self._send_json(status, body)
def log_message(self, *a):
pass
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--port", type=int, required=True)
ap.add_argument("--scenarios",
help="scenarios-openai.json (required in normal mode)")
ap.add_argument("--log", required=True)
ap.add_argument("--mode", default="normal",
choices=["normal", "black-hole", "mid-body-drop",
"tool-pending-forever"])
args = ap.parse_args()
if args.port in (7770, 7779, 17779):
raise SystemExit("stub-openai: refusing production Neuron port")
if args.mode == "normal" and not args.scenarios:
raise SystemExit("stub-openai: --scenarios is required in normal mode")
STATE["mode"] = args.mode
STATE["scenarios"] = (load_scenarios(args.scenarios)
if args.scenarios else None)
STATE["log_path"] = args.log
open(args.log, "w").close()
n_markers = (len(STATE["scenarios"]["marker_map"])
if STATE["scenarios"] else 0)
print("stub-openai [%s]: 127.0.0.1:%d /v1/chat/completions "
"(%d markers registered, log=%s)"
% (args.mode, args.port, n_markers, args.log), flush=True)
ThreadingHTTPServer(("127.0.0.1", args.port), Handler).serve_forever()
if __name__ == "__main__":
main()
+148
View File
@@ -0,0 +1,148 @@
#!/usr/bin/env bash
# run-el-test.sh — build and RUN one engine test (tests/*.el), printing its assertions.
#
# WHY THIS EXISTS (2026-08-06): the engine's tests/*.el files were never runnable from the
# tree. `elc` is a COMPILER — it emits C to stdout and exits; it does not execute anything.
# So "the tests" could only ever be read, not run, and a signature change could silently
# break them (exactly what happened when bridge_save gained its `wire` argument). This
# script closes that: emit the test to C, link it against the engine modules, execute it.
#
# HOW IT WORKS
# 1. elc <test>.el -> C on stdout (the test file's `main` + prototypes)
# 2. elb (once, cached) -> per-module C for the whole engine into a scratch dir
# 3. cc test.c + all modules EXCEPT soul.c (soul.c owns the real `main`) + the runtime
# 4. run it
#
# The test C references only the engine functions it actually calls, so there are no
# duplicate-symbol collisions with the module objects.
#
# RUNTIME: the REPO-PINNED vendor/el-runtime (NOT ~/el-sdk/el_runtime.c — that June build
# is missing builtins August code calls: engram_wm_count, engram_wm_top_json,
# http_delete_json, http_serve_async; linking against it fails with "symbol(s) not found").
#
# USAGE
# tests/run-el-test.sh tests/test_bridge_serialization.el # one test
# tests/run-el-test.sh --all # every tests/test_*.el
# REBUILD=1 tests/run-el-test.sh ... # force module regeneration
#
# Tests that need a live API key / running soul (see each file's header) will report their
# own skips or failures — this runner does not fake them.
set -uo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT" || exit 2
ELC="${ELC:-$HOME/el-sdk/elc}"
ELB="${ELB:-$HOME/Development/el-sdk/bin/elb}"
RUNTIME_DIR="${RUNTIME_DIR:-$REPO_ROOT/vendor/el-runtime/v1.0.0-20260501}"
SCRATCH="${SCRATCH:-/tmp/el-test-$(basename "$REPO_ROOT")}"
MODDIR="$SCRATCH/modules"
OPENSSL_INC="${OPENSSL_INC:-/opt/homebrew/opt/openssl@3/include}"
OPENSSL_LIB="${OPENSSL_LIB:-/opt/homebrew/opt/openssl@3/lib}"
for req in "$ELC" "$ELB" "$RUNTIME_DIR/el_runtime.c"; do
[ -e "$req" ] || { echo "run-el-test: missing required input: $req" >&2; exit 2; }
done
mkdir -p "$MODDIR" || exit 2
# ── Step 1: engine modules (cached — regeneration is the slow part) ────────────
if [ "${REBUILD:-0}" = "1" ] || [ ! -f "$MODDIR/chat.c" ] || [ "chat.el" -nt "$MODDIR/chat.c" ]; then
echo "run-el-test: generating engine modules into $MODDIR (this takes ~1-2 min)..."
# elb's own final link step fails by design here (it wants to produce a binary named
# `neuron` and we only need the per-module .c files it emits first). Ignore its rc.
"$ELB" --elc="$ELC" --runtime="$RUNTIME_DIR" --out="$MODDIR/" >"$SCRATCH/elb.log" 2>&1
if [ ! -f "$MODDIR/chat.c" ]; then
echo "run-el-test: FATAL — elb produced no chat.c; see $SCRATCH/elb.log" >&2
tail -5 "$SCRATCH/elb.log" >&2
exit 2
fi
# elb rewrites *.elh in the source tree as a side effect (cosmetic banner churn plus a
# stray soul..elh). Say so; the caller decides whether to `git restore` them.
echo "run-el-test: NOTE — elb regenerated *.elh in the source tree (cosmetic churn is expected; a stray soul..elh may appear)."
fi
# soul.c is needed for its engine functions (layered_cycle et al.) but it also owns the
# daemon's real `main`, which would collide with the test's own. Compile it ONCE to an
# object with `main` renamed away, and link that instead of the .c.
SOUL_OBJ="$SCRATCH/soul-nomain.o"
if [ "${REBUILD:-0}" = "1" ] || [ ! -f "$SOUL_OBJ" ] || [ "$MODDIR/soul.c" -nt "$SOUL_OBJ" ]; then
cc -std=c11 -O1 -DHAVE_CURL -Dmain=el_soul_daemon_main_unused \
-I "$RUNTIME_DIR" -I "$MODDIR" -I "$OPENSSL_INC" \
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
-c "$MODDIR/soul.c" -o "$SOUL_OBJ" 2>"$SCRATCH/soul-nomain.err" \
|| { echo "run-el-test: FATAL — could not compile soul.c without main" >&2
grep -E 'error:' "$SCRATCH/soul-nomain.err" | head -5 >&2; exit 2; }
fi
# Every module except soul.c (linked as the renamed object above) and the stray soul.elh.c.
MODS=("$SOUL_OBJ")
for f in "$MODDIR"/*.c; do
case "$(basename "$f")" in
soul.c|soul.elh.c) continue ;;
esac
MODS+=("$f")
done
[ "${#MODS[@]}" -gt 1 ] || { echo "run-el-test: no module objects found" >&2; exit 2; }
run_one() {
local test_el="$1"
local name; name="$(basename "$test_el" .el)"
local cfile="$SCRATCH/$name.c"
local bin="$SCRATCH/$name"
printf '\n══ %s ══\n' "$name"
if ! "$ELC" "$test_el" >"$cfile" 2>"$SCRATCH/$name.elc.err"; then
echo "COMPILE FAILED (elc):"; tail -10 "$SCRATCH/$name.elc.err"; return 1
fi
[ -s "$cfile" ] || { echo "COMPILE FAILED (elc produced empty C)"; return 1; }
if ! cc -std=c11 -O1 -DHAVE_CURL -rdynamic \
-I "$RUNTIME_DIR" -I "$MODDIR" -I "$OPENSSL_INC" -L "$OPENSSL_LIB" \
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
-o "$bin" "$cfile" "${MODS[@]}" "$RUNTIME_DIR/el_runtime.c" \
-lssl -lcrypto -lcurl -lpthread -lm 2>"$SCRATCH/$name.link.err"; then
echo "LINK FAILED:"; grep -E '"_|error:' "$SCRATCH/$name.link.err" | head -10; return 1
fi
# THE RUNNER OWNS THE VERDICT — the test files cannot be trusted to report it.
#
# Every tests/*.el assert helper does `let pass_count = pass_count + 1` INSIDE an if
# BLOCK. El's scope rule (the same one chat.el documents at every while-body mutation:
# "mutations inside if *blocks* don't escape scope") means those counters never
# increment, so all 9 counted test files print "N passed, M failed" as "0 passed, 0
# failed" — forever, whatever actually happened. A summary that can never report a
# failure is worth exactly as much as an assertion that can never fail. Logged as a
# bug for the real in-file fix; until then the verdict is computed HERE, from the
# assert helpers' own per-line output, which IS reliable.
local out="$SCRATCH/$name.out"
"$bin" 2>&1 | tee "$out"; local rc=${PIPESTATUS[0]}
# NOTE: `grep -c` prints 0 AND exits 1 when there are no matches, so a `|| echo 0`
# fallback appends a SECOND zero and every later integer test breaks on "0\n0".
# (Caught by running this script — which is the whole argument for running things.)
local n_pass n_fail
n_pass=$(grep -c '^ PASS: ' "$out" 2>/dev/null); n_pass=${n_pass:-0}
n_fail=$(grep -c '^ FAIL: ' "$out" 2>/dev/null); n_fail=${n_fail:-0}
echo "── $name: $n_pass passed, $n_fail failed (counted by the runner, not by the file's dead counters)"
if [ "$n_fail" -gt 0 ]; then
echo " failing assertions:"; grep '^ FAIL: ' "$out" | sed 's/^/ /'
return 1
fi
if [ "$n_pass" -eq 0 ]; then
echo " WARNING: no assertions ran — treating as FAILURE (a test that asserts nothing is not a passing test)"
return 1
fi
[ $rc -eq 0 ] || { echo " (test binary exited rc=$rc)"; return 1; }
return 0
}
rc_all=0
if [ "${1:-}" = "--all" ]; then
for t in tests/test_*.el; do run_one "$t" || rc_all=1; done
else
[ $# -ge 1 ] || { echo "usage: tests/run-el-test.sh <tests/test_x.el> | --all" >&2; exit 2; }
for t in "$@"; do run_one "$t" || rc_all=1; done
fi
exit $rc_all
+109 -4
View File
@@ -93,7 +93,7 @@ println("1. bridge_save — empty messages guard")
let sid1: String = "test-session-empty-messages"
state_set("mcp_bridge:" + sid1, "")
let save1_ok: Bool = bridge_save(sid1, "claude-sonnet-4-5", "sys", "[]", "", "", "call-1")
let save1_ok: Bool = bridge_save(sid1, "claude-sonnet-4-5", "sys", "[]", "", "", "call-1", "anthropic")
assert_false("empty messages -> bridge_save returns false", save1_ok)
let saved1: String = state_get("mcp_bridge:" + sid1)
@@ -107,7 +107,7 @@ println("2. bridge_save — empty tools_json guard")
let sid2: String = "test-session-empty-tools"
state_set("mcp_bridge:" + sid2, "")
let save2_ok: Bool = bridge_save(sid2, "claude-sonnet-4-5", "sys", "", "[{\"role\":\"user\",\"content\":\"hi\"}]", "", "call-2")
let save2_ok: Bool = bridge_save(sid2, "claude-sonnet-4-5", "sys", "", "[{\"role\":\"user\",\"content\":\"hi\"}]", "", "call-2", "anthropic")
assert_false("empty tools_json -> bridge_save returns false", save2_ok)
let saved2: String = state_get("mcp_bridge:" + sid2)
@@ -126,7 +126,7 @@ state_set("mcp_bridge:" + sid3, "")
let msgs3: String = "[{\"role\":\"user\",\"content\":\"hello\"}]"
let tools3: String = "[{\"name\":\"read_file\"}]"
let save3_ok: Bool = bridge_save(sid3, "claude-sonnet-4-5", "You are a helper.", tools3, msgs3, "read_file", "toolu_abc")
let save3_ok: Bool = bridge_save(sid3, "claude-sonnet-4-5", "You are a helper.", tools3, msgs3, "read_file", "toolu_abc", "anthropic")
assert_true("valid args -> bridge_save returns true", save3_ok)
let blob3: String = state_get("mcp_bridge:" + sid3)
@@ -243,7 +243,7 @@ state_set("mcp_bridge:" + sid8, "")
let special_id: String = "toolu_test\"quoted\""
let msgs8: String = "[{\"role\":\"user\",\"content\":\"hi\"}]"
let tools8: String = "[{\"name\":\"read_file\"}]"
let save8_ok: Bool = bridge_save(sid8, "claude-sonnet-4-5", "sys", tools8, msgs8, "", special_id)
let save8_ok: Bool = bridge_save(sid8, "claude-sonnet-4-5", "sys", tools8, msgs8, "", special_id, "anthropic")
assert_true("special chars in tool_use_id -> bridge_save returns true", save8_ok)
let blob8: String = state_get("mcp_bridge:" + sid8)
@@ -251,6 +251,111 @@ let blob8: String = state_get("mcp_bridge:" + sid8)
let retrieved_id: String = json_get(blob8, "tool_use_id")
assert_eq("tool_use_id with quotes round-trips via json_safe", retrieved_id, special_id)
// Section 9: the "wire" field (OpenAI-tools port, 2026-08-06)
//
// A suspended turn must resume on the SAME wire format it suspended on: an OpenAI-lane
// bridge answered with an Anthropic-shaped tool_result (or vice versa) is a dead run.
// bridge_save therefore stamps the blob with "wire", and agentic_resume branches on it.
//
// §9c is the important one. json_get is a first-substring-match scanner, so any key that
// appears inside the UNESCAPED conversation embedded in messages_raw can be matched
// instead of the blob's own field that exact class of bug produced the round-9 resume
// failure (json_get(blob,"tool_use_id") matching a web_search_tool_result's id inside the
// replayed conversation). "wire" is written as a json_safe'd SCALAR ahead of both raw
// fields precisely so a decoy in model-controlled bytes can never win. This test plants
// that decoy on purpose. If someone later moves the field after messages_raw, this fails.
println("")
println("9. bridge_save — wire tagging and its field-order guarantee")
// 9a. an OpenAI-lane suspension round-trips as "openai"
let sid9: String = "test-session-wire-openai"
state_set("mcp_bridge:" + sid9, "")
let msgs9: String = "[{\"role\":\"user\",\"content\":\"hi\"}]"
let tools9: String = "[{\"name\":\"read_file\"}]"
let save9_ok: Bool = bridge_save(sid9, "llama-3.3-70b-versatile", "sys", tools9, msgs9, "", "call_abc", "openai")
assert_true("openai wire -> bridge_save returns true", save9_ok)
let blob9: String = state_get("mcp_bridge:" + sid9)
assert_eq("wire round-trips as openai", json_get(blob9, "wire"), "openai")
// 9b. an Anthropic-lane suspension round-trips as "anthropic"
let sid9b: String = "test-session-wire-anthropic"
state_set("mcp_bridge:" + sid9b, "")
let save9b_ok: Bool = bridge_save(sid9b, "claude-sonnet-4-5", "sys", tools9, msgs9, "", "toolu_abc", "anthropic")
assert_true("anthropic wire -> bridge_save returns true", save9b_ok)
let blob9b: String = state_get("mcp_bridge:" + sid9b)
assert_eq("wire round-trips as anthropic", json_get(blob9b, "wire"), "anthropic")
// 9c. FIELD-ORDER GUARD: a decoy "wire" inside the conversation must NOT be matched.
let sid9c: String = "test-session-wire-decoy"
state_set("mcp_bridge:" + sid9c, "")
let msgs9c: String = "[{\"role\":\"user\",\"content\":\"please save this literal text: \\\"wire\\\":\\\"anthropic\\\" end\"}]"
let save9c_ok: Bool = bridge_save(sid9c, "llama-3.3-70b-versatile", "sys", tools9, msgs9c, "", "call_decoy", "openai")
assert_true("decoy conversation -> bridge_save returns true", save9c_ok)
let blob9c: String = state_get("mcp_bridge:" + sid9c)
assert_eq("blob's own wire wins over a decoy planted in messages_raw", json_get(blob9c, "wire"), "openai")
// 9d. LEGACY blob (written before the port) has no wire field: json_get yields "",
// which agentic_resume treats as the Anthropic path old suspensions still resume.
let sid9d: String = "test-session-wire-legacy"
let legacy_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
+ ",\"tool_use_id\":\"toolu_legacy\",\"tools_raw\":[{\"name\":\"read_file\"}]"
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
state_set("mcp_bridge:" + sid9d, legacy_blob)
let blob9d: String = state_get("mcp_bridge:" + sid9d)
assert_eq("legacy blob has no wire field -> empty (resumes as anthropic)", json_get(blob9d, "wire"), "")
assert_eq("legacy blob still reads its tool_use_id", json_get(blob9d, "tool_use_id"), "toolu_legacy")
// 9e. THE HARDER DECOY: a LEGACY blob (no wire field of its own) that carries the bytes
// of a wire tag deeper inside, where an unbounded first-match scan would find it and
// misroute the resume onto the wrong loop the round-9 defect class exactly.
//
// WHAT IS AND IS NOT REACHABLE (measured here, not assumed an earlier version of this
// test asserted the wrong thing and was corrected by running it):
// * NOT reachable from ordinary conversation TEXT. Any quote a user or model writes is
// backslash-escaped when it is serialized into the blob, so prose containing
// "wire":"openai" is stored as \"wire\":\"openai\" and does not match a scan for the
// unescaped key. §9f pins that.
// * REACHABLE from STRUCTURAL keys, which are embedded raw. Conversation and tool
// objects keep real quotes that is precisely how round 9's scan found a
// web_search_tool_result's tool_use_id. A connector-supplied tool schema or a future
// message field literally named "wire" would be found the same way.
// The bound removes the whole class rather than reasoning about which keys exist today.
let sid9e: String = "test-session-wire-legacy-decoy"
let decoy_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
+ ",\"tool_use_id\":\"toolu_legacy\""
+ ",\"tools_raw\":[{\"name\":\"read_file\",\"wire\":\"openai\"}]"
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
state_set("mcp_bridge:" + sid9e, decoy_blob)
let blob9e: String = state_get("mcp_bridge:" + sid9e)
// Unbounded read (what NOT to do) proves the hazard this guard exists for is real.
assert_eq("unbounded scan DOES find a structural decoy (why the bound is needed)", json_get(blob9e, "wire"), "openai")
// Bounded read the same computation agentic_resume performs.
let d_traw: Int = str_index_of(blob9e, ",\"tools_raw\":")
let d_tjson: Int = str_index_of(blob9e, ",\"tools_json\":")
let d_mraw: Int = str_index_of(blob9e, ",\"messages_raw\":")
let d_msgs: Int = str_index_of(blob9e, ",\"messages\":")
let dcut1: Int = if d_traw > 0 { d_traw } else { str_len(blob9e) }
let dcut2: Int = if d_tjson > 0 && d_tjson < dcut1 { d_tjson } else { dcut1 }
let dcut3: Int = if d_mraw > 0 && d_mraw < dcut2 { d_mraw } else { dcut2 }
let dcut: Int = if d_msgs > 0 && d_msgs < dcut3 { d_msgs } else { dcut3 }
let head9e: String = str_slice(blob9e, 0, dcut)
assert_eq("bounded scan ignores the decoy -> legacy blob resumes as anthropic", json_get(head9e, "wire"), "")
assert_not_contains("scalar head excludes the bulk fields entirely", head9e, "read_file")
// 9f. Escaping bounds the severity: prose CANNOT inject a scalar-looking key, because
// its quotes are escaped on the way in. Documented as a measured fact, so nobody has to
// re-derive it the next time this question comes up.
let sid9f: String = "test-session-wire-prose"
let prose_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
+ ",\"tool_use_id\":\"toolu_legacy\",\"tools_raw\":[{\"name\":\"read_file\"}]"
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"remember this: \\\"wire\\\":\\\"openai\\\"\"}]}"
state_set("mcp_bridge:" + sid9f, prose_blob)
let blob9f: String = state_get("mcp_bridge:" + sid9f)
assert_eq("escaped prose cannot spoof the key even unbounded (severity bound)", json_get(blob9f, "wire"), "")
// Summary
println("")
+213
View File
@@ -0,0 +1,213 @@
// test_history_amplification.el
//
// REGRESSION TEST FOR ISSUE #129 (P0, SAFETY).
//
// What this guards: on the agentic path, the crisis score has two halves the
// message you just sent, and the distress that has accumulated across the
// conversation. The second half is the whole reason the escalation logic exists:
// someone whose distress builds over several turns never sends one message that
// trips the bell on its own.
//
// The defect this test was written against (ff421d3, 2026-08-05 fixed
// 2026-08-07): conversation history moved to a per-session key via
// conv_hist_key(session_id), but the agentic path's safety screen was left
// reading the old anonymous "conv_history" bucket. The desktop app always sends
// a session_id, so the screen received "" on every real conversation and the
// escalation half always scored 0. Nothing failed. Nothing logged. The comment
// above the defective line documented this same bug being fixed once before.
//
// THE INVARIANT UNDER TEST, stated so it survives future renames:
// the window the safety screen READS must be the window conv_history_record
// WRITES. Not "must be called conv_history" must AGREE.
//
// This test is deliberately written to fail loudly on the pre-fix source. If it
// ever passes on code where the screen reads a key nothing writes, it is broken.
//
// To run (macOS, from the worktree root):
// scripts/run-el-test.sh tests/test_history_amplification.el
//
import "../chat.el"
import "../safety.el"
import "../sessions.el"
// Program class. Without this an El program compiles as a 'utility', and a
// utility may not call the self-formation primitives (llm_call_system,
// llm_vision) that chat.el's agentic loop references the unit fails to
// compile with a capability violation even though the test never calls them.
// Declaring 'cgi' matches how soul.el declares itself.
//
// The endpoints below are deliberately DEAD: this test must never reach a live
// engram, and nothing it asserts depends on one. Port 9 is discard.
cgi "neuron-test-history-amplification" {
dharma_id: "ntn-test@http://127.0.0.1:9",
principal: "test-harness",
network: "dharma-testnet",
engram: "http://127.0.0.1:9"
}
// Counters
//
// NOTE for anyone copying this harness: the idiom used by the older tests in
// this directory `let pass_count = pass_count + 1` inside an assert function
// does NOT mutate the module-level binding. It declares a new local that dies
// with the call, so those suites all print "0 passed, 0 failed" no matter what
// happened. Counters go through the state store here so the summary is real.
fn bump(counter: String) -> Void {
let cur: String = state_get(counter)
let n: Int = if str_eq(cur, "") { 0 } else { str_to_int(cur) }
state_set(counter, int_to_str(n + 1))
}
fn counter_of(counter: String) -> Int {
let cur: String = state_get(counter)
if str_eq(cur, "") { return 0 }
return str_to_int(cur)
}
fn assert_eq(label: String, got: String, expected: String) -> Void {
if str_eq(got, expected) {
bump("t_pass")
println(" PASS: " + label)
} else {
bump("t_fail")
println(" FAIL: " + label)
println(" got: " + got)
println(" expected: " + expected)
}
}
fn assert_true(label: String, cond: Bool) -> Void {
if cond {
bump("t_pass")
println(" PASS: " + label)
} else {
bump("t_fail")
println(" FAIL: " + label)
}
}
// Fixtures
//
// ESCALATION_MSG scores 50 on its own ("can't go on", safety_score_crisis s9).
// 50 is >= soft (35) and < hard (70): a soft bell by itself, which is the point.
// It is the message that is only a crisis IN CONTEXT.
//
// DISTRESS_TURN carries five history phrases:
// hopeless 15 + worthless 15 + nobody cares 15 + can't take it anymore 20
// + want to disappear 20 = 85; contributed at 1/3 weight = 28.
// 50 + 28 = 78 >= 70 hard_bell. That 28-point swing IS the escalation logic,
// and it is exactly what the defect suppressed.
let TEST_SESSION: String = "sess-test-129-history-amplification"
let ESCALATION_MSG: String = "honestly i can't go on like this"
let DISTRESS_TURN: String = "i feel hopeless and worthless, nobody cares, i can't take it anymore, i want to disappear"
println("")
println("test_history_amplification.el — issue #129 (agentic multi-turn crisis escalation)")
// 1. Baseline: the message alone is a SOFT bell, not a hard one
//
// If this ever returns hard_bell, the test below proves nothing the message
// would trip the bell without any history and the amplification would be
// invisible. This assertion is what keeps the real test honest.
println("")
println("1. baseline — escalation message with NO history is a soft bell")
let baseline: String = safety_screen(ESCALATION_MSG, "")
assert_eq("no history -> soft_bell (not hard)", json_get(baseline, "action"), "soft_bell")
// 2. Producer sanity: history lands in the session's own window
println("")
println("2. producer — conv_history_record writes the session's window")
conv_history_record(TEST_SESSION, DISTRESS_TURN, "i hear you, that sounds heavy", "")
let written: String = state_get(conv_hist_key(TEST_SESSION))
assert_true("session window is non-empty after record", !str_eq(written, ""))
assert_true("session window contains the distress turn", str_contains(written, "hopeless"))
// 3. THE REGRESSION: the agentic screen must SEE that window
//
// Pre-fix this returns soft_bell, because agentic_safety_screen read the
// anonymous bucket and got "". Post-fix it returns hard_bell.
println("")
println("3. REGRESSION #129 — agentic screen reads the session's own window")
let screened: String = agentic_safety_screen(TEST_SESSION, ESCALATION_MSG)
assert_eq(
"distress history escalates the agentic screen to hard_bell",
json_get(screened, "action"),
"hard_bell"
)
// 4. The invariant, stated directly
//
// Independent of thresholds and phrase lists: whatever the screen reads for a
// session must equal what the recorder wrote for that session. This is the
// assertion that survives a future rename of either side.
println("")
println("4. invariant — read window == written window")
let read_back: String = state_get(conv_hist_key(TEST_SESSION))
assert_true("screen input is the recorded window, not empty", !str_eq(read_back, ""))
assert_eq("read window is byte-identical to written window", read_back, written)
// 5. No false positive: a calm session does not escalate
//
// A test that only ever asserts "hard_bell" would pass on code that hard-bells
// every message. This is the other leg, and it runs BEFORE the anonymous case
// below on purpose: that case writes the shared bucket, and under the defect a
// calm session would then inherit it.
println("")
println("5. specificity — a calm history does NOT escalate")
let CALM_SESSION: String = "sess-test-129-calm"
state_set("conv_history", "")
conv_history_record(CALM_SESSION, "what is the weather like today", "clear and mild", "")
let calm: String = agentic_safety_screen(CALM_SESSION, ESCALATION_MSG)
assert_eq("calm history stays at soft_bell", json_get(calm, "action"), "soft_bell")
// 6. Cross-session leakage
//
// The same defect had a second face: because the screen read one shared bucket,
// a calm session could be scored against a DIFFERENT session's distress. That is
// wrong in both directions it fabricates a crisis for the calm user and it
// leaks the distressed user's content into another session's scoring.
println("")
println("6. isolation — one session's distress must not score another session")
state_set("conv_history", "")
let OTHER_SESSION: String = "sess-test-129-other"
conv_history_record(OTHER_SESSION, DISTRESS_TURN, "i hear you", "")
let isolated: String = agentic_safety_screen(CALM_SESSION, ESCALATION_MSG)
assert_eq(
"a distressed OTHER session does not escalate the calm session",
json_get(isolated, "action"),
"soft_bell"
)
// 7. Anonymous sessions still work
//
// conv_hist_key("") deliberately falls back to the shared "conv_history" bucket.
// The fix must not break the no-session_id path older callers rely on. Runs last
// because it writes that shared bucket.
println("")
println("7. anonymous path — empty session_id still screens against the shared window")
state_set("conv_history", "[{\"role\":\"user\",\"content\":\"" + DISTRESS_TURN + "\"}]")
let anon: String = agentic_safety_screen("", ESCALATION_MSG)
assert_eq("anonymous session escalates too", json_get(anon, "action"), "hard_bell")
// Summary
println("")
println("history amplification tests: " + int_to_str(counter_of("t_pass")) + " passed, " + int_to_str(counter_of("t_fail")) + " failed")
+92
View File
@@ -0,0 +1,92 @@
// tests/test_utf8_slice.el
//
// Guards utf8_safe_slice(), the fix for a live defect found 2026-08-06:
//
// The session preload cuts recalled memory content at a fixed length
// (chat.el: `if str_len(acc) > 350 { str_slice(acc, 0, 350) }` and
// session_preload_bullets' identical per-bullet cut). str_slice and str_len count
// BYTES, so any cut landing inside a multi-byte UTF-8 character leaves a dangling
// lead byte in the system prompt and the whole request body is then invalid UTF-8.
// Providers reject it outright, so the user sees "AI unavailable" with no clue why,
// on both wire formats. Caught by an OpenAI-lane gate whose stub decodes strictly;
// reproduced from a real memory whose content contained box-drawing rules (E2 94 80).
//
// Trigger is ordinary content: an em dash, a curly quote, an accented name, a table
// border, an emoji anything non-ASCII sitting on the cut boundary. It gets MORE
// likely as a user's memory grows, which is the opposite of what should happen.
//
// §1 also pins the semantics this fix depends on: that str_char_code returns the
// BYTE value at a byte index (not a decoded code point). If a future runtime changes
// that, these assertions fail loudly instead of the truncation silently rotting.
import "../chat.el"
let pass_count: Int = 0
let fail_count: Int = 0
fn assert_eq(label: String, got: String, expected: String) -> Void {
if str_eq(got, expected) {
let pass_count = pass_count + 1
println(" PASS: " + label)
} else {
let fail_count = fail_count + 1
println(" FAIL: " + label)
println(" got: " + got)
println(" expected: " + expected)
}
}
fn assert_eq_int(label: String, got: Int, expected: Int) -> Void {
assert_eq(label, int_to_str(got), int_to_str(expected))
}
println("")
println("1. runtime semantics this fix relies on")
// "" is U+2500 = E2 94 80 (three bytes). If str_len counts bytes, len("") is 3.
let dash: String = ""
assert_eq_int("str_len counts BYTES (one box-drawing char = 3)", str_len(dash), 3)
assert_eq_int("str_char_code returns the BYTE value (lead byte of U+2500 = 0xE2 = 226)", str_char_code(dash, 0), 226)
assert_eq_int("str_char_code second byte = 0x94 = 148", str_char_code(dash, 1), 148)
assert_eq_int("str_char_code third byte = 0x80 = 128", str_char_code(dash, 2), 128)
println("")
println("2. utf8_safe_slice — never leaves a partial character")
// Pure ASCII: behaves exactly like str_slice.
assert_eq("ascii under the limit is untouched", utf8_safe_slice("hello", 10), "hello")
assert_eq("ascii over the limit cuts exactly", utf8_safe_slice("hello world", 5), "hello")
// A cut landing INSIDE a 3-byte character must drop that character entirely.
// "ab─cd": bytes a b E2 94 80 c d. Cutting at 3 or 4 lands mid-dash.
let mixed: String = "ab─cd"
assert_eq_int("fixture is 7 bytes (2 ascii + 3 + 2 ascii)", str_len(mixed), 7)
assert_eq("cut inside the char (n=3) drops the partial char", utf8_safe_slice(mixed, 3), "ab")
assert_eq("cut inside the char (n=4) drops the partial char", utf8_safe_slice(mixed, 4), "ab")
// A cut landing exactly AFTER a complete character keeps it.
assert_eq("cut on the char boundary (n=5) keeps the whole char", utf8_safe_slice(mixed, 5), "ab─")
// 2-byte character (é = C3 A9) and 4-byte character (😀 = F0 9F 98 80).
let acc: String = ""
assert_eq("cut inside a 2-byte char drops it", utf8_safe_slice(acc, 2), "x")
assert_eq("cut after a 2-byte char keeps it", utf8_safe_slice(acc, 3), "")
let emo: String = "x😀"
assert_eq("cut inside a 4-byte char drops it (n=3)", utf8_safe_slice(emo, 3), "x")
assert_eq("cut inside a 4-byte char drops it (n=4)", utf8_safe_slice(emo, 4), "x")
assert_eq("cut after a 4-byte char keeps it", utf8_safe_slice(emo, 5), "x😀")
println("")
println("3. the real-world shape that produced the bug")
// A run of box-drawing rules, cut mid-character the exact captured failure.
let rules: String = "──────"
assert_eq_int("six box rules = 18 bytes", str_len(rules), 18)
// n=16 lands one byte into the sixth character.
let cut16: String = utf8_safe_slice(rules, 16)
assert_eq_int("cut at 16 backs off to a clean 15-byte boundary", str_len(cut16), 15)
// Every byte of the result must belong to a complete character: the last byte of a
// well-formed run of these is always 0x80, and 15 is divisible by 3.
assert_eq_int("result ends on a complete char (last byte 0x80)", str_char_code(cut16, 14), 128)
println("")
println("test_utf8_slice.el: " + int_to_str(pass_count) + " passed, " + int_to_str(fail_count) + " failed")
+59
View File
@@ -0,0 +1,59 @@
#!/usr/bin/env bash
# build-soul-from-dist.sh — build a deployable soul from the SAME input CI compiles.
#
# THE PROBLEM THIS CLOSES: until now, deploys were built by build-soul.sh, which
# compiles a scratch amalgam and never touches dist/soul.c. CI compiles dist/soul.c.
# Two lineages. On 2026-08-09 the committed input fell 2,761 bytes behind the sources
# while three binaries built the other way were installed on the operator machine —
# so "what runs" and "what the repo says builds" were different artifacts again,
# which is the whole of #133 and #111 wearing new clothes.
#
# This builds from dist/soul.c with CI's own flags, after asserting that dist/soul.c
# actually matches the .el sources, and writes a provenance sidecar so a deployer can
# refuse anything of unknown origin.
#
# -rdynamic and -DHAVE_CURL are copied from .gitea/workflows/ci.yaml deliberately.
# The CI comment explains -rdynamic: without it the runtime cannot resolve its HTTP
# handler by name via dlsym and the binary serves nothing on every route.
#
# usage: build-soul-from-dist.sh <out-binary>
set -u
OUT="${1:?usage: build-soul-from-dist.sh <out-binary>}"
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
RUNTIME="$ROOT/vendor/el-runtime/v1.0.0-20260501"
cd "$ROOT" || exit 2
echo "[build-from-dist] GATE: does dist/soul.c match the sources?"
if ! ./tools/soulc-stamp.sh --check; then
echo "[build-from-dist] REFUSING — the build input is stale. Regenerate and stamp first." >&2
exit 9
fi
[ -f "$RUNTIME/el_runtime.c" ] || { echo "pinned runtime missing at $RUNTIME" >&2; exit 2; }
echo "[build-from-dist] compiling dist/soul.c with CI's flags"
cc -O2 -DHAVE_CURL -rdynamic \
-I"$RUNTIME" \
dist/soul.c \
"$RUNTIME/el_runtime.c" \
-lcurl -lpthread -lm \
-o "$OUT" || { echo "[build-from-dist] COMPILE FAILED" >&2; exit 3; }
# Provenance sidecar: what a deployer checks before installing anything.
SRC_SHA="$(shasum -a 256 dist/soul.c | awk '{print $1}')"
STAMP_SHA="$(shasum -a 256 dist/soul.c.stamp | awk '{print $1}')"
COMMIT="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
DIRTY="clean"; [ -n "$(git status --porcelain -- '*.el' dist/soul.c 2>/dev/null)" ] && DIRTY="DIRTY"
cat > "$OUT.provenance" <<EOF
{"built_from":"dist/soul.c",
"dist_soul_c_sha256":"$SRC_SHA",
"stamp_sha256":"$STAMP_SHA",
"git_commit":"$COMMIT",
"worktree":"$DIRTY",
"runtime":"vendor/el-runtime/v1.0.0-20260501",
"flags":"-O2 -DHAVE_CURL -rdynamic"}
EOF
echo "[build-from-dist] OK -> $OUT ($(wc -c < "$OUT" | tr -d ' ') bytes)"
echo "[build-from-dist] provenance -> $OUT.provenance (commit ${COMMIT:0:8}, worktree $DIRTY)"
+147
View File
@@ -0,0 +1,147 @@
# Retrieval eval harness
Measures Neuron's memory retrieval so a change can be shown to help before it is
believed to help. Nothing else on the memory roadmap should ship without a run
through this.
```
tools/retrieval-eval/run_comparison.sh --baseline main --candidate <branch>
```
That builds a soul from each ref, boots each in isolation on a fixed corpus,
runs the gold set three times per ref, and prints a table plus a verdict that
refuses to call a difference real if it is inside the noise band.
## What was reused
This is not a new idea, it is the missing third of an existing one.
| Prior work | What it gave | What was missing |
|---|---|---|
| `docs/research/graphrag_eval/` (`collect.py`, `score.py`, 2026-06-08) | The three-retriever comparison that produced the numbers everyone quotes: substring 1.7% P@5, graph 21.7%, BM25 55%. Per-query relevant-id scoring, fixed-denominator precision@5, unique-relevant analysis. | 13 hand-written queries, judged by an LLM after the fact; measured the *live* soul on the *live* engram. |
| `docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py` | The pinned-query discipline: ground truth committed as regexes so every run judges alike, plus a `--check` winnability gate. 40 queries in 5 bands including a deliberate paraphrase-hard band. | Scored offline replicas of substring/BM25 — it never ran the real retrieval path. |
| `docs/research-archive/p0-prototypes/stage0_eval_20260714.py` | The `hit@5` metric and the substring/BM25 reference implementations. | Same: offline only. |
| `scripts/verify-soul-contract.sh` | The isolation recipe, verbatim: throwaway port, throwaway `HOME`, `SOUL_ENGRAM_PATH`, and the non-obvious `SOUL_ISE_URL` pin that stops an "isolated" soul silently syncing the operator's live brain. | It is a contract gate, not a measurement. |
| `_engine-liveness-91/gen-soul-amalgam.sh` + `.gitea/workflows/ci.yaml` | The build recipe (`elc --target=c` with every `.elh` on the import chain removed) and CI's exact compile flags. | — |
**Reused directly:** the isolation recipe, the build recipe, fixed-denominator
precision@5, the pinned-ground-truth and winnability ideas.
**New here:** ids rather than regexes as ground truth, an associative category
derived from real graph edges, a superseded/contradicted category scored on
ranking, a machine-checked zero-lexical-overlap guarantee on paraphrases,
paired significance testing, and — the point — measurement against the **real
compiled soul** rather than an offline replica of one leg of it.
## Design fit
The thing under measurement is Will's designed retrieval: spreading activation
over the weighted directed graph, four-factor multiplicative scoring (parent
strength x edge weight x target salience x query/target cosine). A Python
re-implementation would measure my reading of the design. So the harness
compiles the actual `soul.el` amalgam and asks it over HTTP on
`/api/neuron/recall`, exactly as the MCP wrapper and the app do.
## Files
| File | Does |
|---|---|
| `build_gold_set.py` | Derives and **validates** the gold set from the corpus. `--check` re-validates and exits non-zero if a query became unwinnable or a paraphrase leaked a word. |
| `gold_set.json` | 38 queries. Every one carries a `derivation` string. |
| `run_eval.py` | Boots one soul in isolation, runs the gold set, writes metrics. Kills and **confirms dead** its child; records the confirmation in the results file. |
| `compare.py` | Paired diff of two result files with McNemar's exact test and a stated noise floor. |
| `build-soul.sh` | Compiles a soul binary from a plain source tree. |
| `run_comparison.sh` | All of the above, end to end, from two git refs. |
## The gold set — 38 queries
Built from the real corpus (`snapshot-pre-repair-20260806.json`, 78,768 nodes /
14,214 edges) so it reflects one person's accumulating memory, not document QA.
| Category | n | Expected answer derived by |
|---|---|---|
| `exact_rare` | 6 | **Mined.** Tokens with document frequency 1 across all 78,768 nodes, whose single containing node is a 3006000 char Memory/Knowledge/Belief. That node is the only possible answer. Re-verified every build. |
| `phrase` | 7 | **Mined.** Case-insensitive verbatim scan; the matching set *is* the answer key. Phrases matching >25 nodes are rejected as too diffuse. |
| `paraphrase` | 13 | **Hand-selected, machine-checked.** Target locked by id; the build then proves that **zero** content words of the query appear anywhere in the target's label, content, or tags. A leak fails the build — the category cannot quietly decay into lexical matching. |
| `associative` | 6 | **Derived from edges.** Query built from one value node's distinctive vocabulary; expected answers are its siblings on the `Self - Values (grounded)` hub. Siblings sharing any query word are dropped, so the only route from query to answer is seed -> hub -> sibling. |
| `nonsense` | 3 | **Control.** Verified that no token occurs anywhere in the corpus. Correct behaviour is to return nothing. |
| `superseded` | 3 | **Derived.** Correction/stale pairs located by regex scan, kept only when both sides resolve to different surviving nodes. Scored on **ranking**: the correction must be returned *and* rank above the stale node. |
## Metrics
`hit@5`, `recall@5`, `recall@10`, `precision@5` (fixed denominator 5, so an
empty result is punished like a page of junk), `MRR@10`, and wall-clock latency
per query (p50/p95/max). Output is a table plus a machine-readable JSON per run
so runs can be diffed.
## Honesty about noise
- **Minimum detectable swing on this 38-query set: 6 queries.** If every query
that changes changes the same way, `p = 2 x 0.5^n`, which first drops under
0.05 at n=6. Any net change smaller than that is inside the noise band and
`compare.py` says so in those words.
- **Run-to-run drift is measured, not assumed.** Activation is a stateful read
by design (traversal reinforces what it touches), so identical inputs need not
give identical outputs. Observed: `main` 0 queries of drift across 3 runs
(fully deterministic); the activation branch 1 query.
- The noise floor used for the verdict is `max(6, observed_drift + 1)`.
- **This gold set is underpowered for small effects.** A genuine 3-query
improvement would not clear the bar. Growing the set is the fix; until then, a
small positive delta means "not shown", not "no effect".
## First result: `main` vs `feat/recall-through-activation`
Corpus and gold set identical, three runs each, fresh corpus copy per run.
| | main | recall-through-activation | delta |
|---|---|---|---|
| hit@5 | 34.3% | 22.9% | **-11.4pp** |
| recall@5 | 26.9% | 19.1% | -7.9pp |
| recall@10 | 33.3% | 24.3% | -9.1pp |
| precision@5 | 12.0% | 7.4% | -4.6pp |
| MRR@10 | 0.294 | 0.242 | -0.053 |
| latency p50 | 1140 ms | 3209 ms | **2.81x** |
| latency p95 | 1584 ms | 4852 ms | 3.06x |
| nonsense clean | 2/3 | 2/3 | — |
| superseded outranks | 1/3 | 0/3 | -1 |
By category (hit@5):
| category | main | activation |
|---|---|---|
| exact_rare | 100% | 100% |
| phrase | 85.7% | **28.6%** |
| paraphrase | 0% | 0% |
| associative | 0% | 0% |
| superseded | 0% | 0% |
**Verdict: directionally worse, one query short of significant.** 5 discordant
pairs, all 5 against the candidate, 0 for it. McNemar exact p = 0.0625 — under
the stated rule that is *inside* the noise band, so the harness reports "no
measurable difference" on accuracy and the honest summary is "5 for 5 the wrong
way, needs a 6th or a larger gold set to call".
Latency is a different story: 2.8x at p50 is deterministic and far outside any
noise band. That regression is real.
The result the branch was written for did not appear. Its stated purpose was to
recover sibling nodes one hub-hop away — the `associative` category — and that
category is **0/6 on both builds**. Probing directly: for the query
`Marines hernia sepsis medical ward`, the activation build returns the lexical
seed node itself at rank 8, and none of its 12 hub siblings anywhere in the top
10. The traversal is running; it is not reaching siblings.
Two corpus facts likely explain it, and both are measurable rather than
speculative:
1. **The graph is nearly edgeless.** Only 4,060 of 78,768 nodes (5.2%) carry any
edge at all — 14,214 edges total, 0.18 per node. Spreading activation over a
graph with no edges is an expensive way to do lexical matching, which is
roughly what the numbers show.
2. **No embeddings.** No node in this snapshot has an embedding field, so the
fourth factor of the four-factor product — query/target cosine similarity —
has nothing to compute from, and the semantic seeding pass is inert.
That is the harness earning its keep on its first job: the change would have
felt like progress (it is the designed mechanism, and it does run) and measures
as a regression on phrase queries plus a 2.8x latency cost, with its intended
benefit unrealised because the corpus lacks the structure it needs.
+21
View File
@@ -0,0 +1,21 @@
import numpy as np, json, urllib.request
SP="/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad"
np.seterr(all='ignore')
M=np.load(SP+'/emb.npy'); eids=open(SP+'/ids.txt',encoding='utf-8',errors='surrogateescape').read().split('\n')
eidx={k:i for i,k in enumerate(eids)}
gold=json.load(open("/Users/timlingo/Development/neuron-technologies/_wt-assoc-leg/tools/retrieval-eval/gold_set.json"))['queries']
VALS=sorted({r for q in gold if q['category']=='paraphrase' for r in q['relevant']})
VI=[eidx[v] for v in VALS]
def emb(t):
b=json.dumps({"model":"nomic-embed-text","prompt":t}).encode()
r=urllib.request.Request("http://127.0.0.1:11434/api/embeddings",data=b,headers={"Content-Type":"application/json"})
v=np.array(json.load(urllib.request.urlopen(r,timeout=60))["embedding"],dtype=np.float32)
return v/(np.linalg.norm(v)+1e-9)
print("qid cat bestValueNodeGlobalRank goldGlobalRank goldSiblingRank")
for q in gold:
if q['category']!='paraphrase': continue
v=emb(q['query']); s=M@v; s[~np.isfinite(s)]=-1
ranks=sorted(int((s>s[j]).sum())+1 for j in VI)
g=eidx[q['relevant'][0]]; gr=int((s>s[g]).sum())+1
sv=np.array([s[j] for j in VI]); sib=int((sv>s[g]).sum())+1
print("%-4s %-11s best=%-5d (top3 val ranks %s) gold=%-5d sib=%d" % (q['id'],q['category'],ranks[0],ranks[:3],gr,sib))
+44
View File
@@ -0,0 +1,44 @@
#!/usr/bin/env bash
# build-soul.sh — compile a soul binary from a plain source tree (no git needed).
#
# Reuses the amalgam recipe worked out in gen-soul-amalgam.sh (round 9.1) and the
# compile flags from .gitea/workflows/ci.yaml, so the binary under test is the
# same translation unit CI ships — not a re-implementation.
#
# elc --target=c emits only an extern prototype for any module that has a .elh
# header beside it, and inlines the module's bodies when it does not. So the
# amalgam is produced in a scratch copy with every .elh on the import chain
# deleted.
#
# usage: build-soul.sh <src-tree-with-*.el> <out-binary>
set -euo pipefail
SRC="${1:?usage: build-soul.sh <src-tree> <out-binary>}"
OUT="${2:?out-binary}"
ELC="${ELC:-$HOME/neuron-dev-stack/src/el/lang/dist/platform/elc}"
EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/el}"
RTDIR="${RTDIR:-$SRC/vendor/el-runtime/v1.0.0-20260501}"
SSL="${SSL_PREFIX:-/opt/homebrew/opt/openssl@3}"
[ -x "$ELC" ] || { echo "no elc at $ELC" >&2; exit 2; }
[ -f "$RTDIR/el_runtime.c" ] || { echo "no el_runtime.c at $RTDIR" >&2; exit 2; }
GEN="$(mktemp -d "${TMPDIR:-/tmp}/soul-build.XXXXXX")"
trap 'rm -rf "$GEN"' EXIT
mkdir -p "$GEN/neuron" "$GEN/foundation/el/elp/src"
cp "$SRC"/*.el "$GEN/neuron/"
cp "$EL_REPO"/elp/src/*.el "$GEN/foundation/el/elp/src/"
find "$GEN" -name '*.elh' -delete
( cd "$GEN/neuron" && "$ELC" --target=c soul.el ) > "$GEN/soul.c"
BODIES=$(grep -c '^el_val_t .*) {$' "$GEN/soul.c" || true)
echo "[build-soul] amalgam $(wc -c < "$GEN/soul.c" | tr -d ' ') bytes, ${BODIES} inlined bodies"
[ "$BODIES" -ge 1200 ] || { echo "[build-soul] FAIL: only $BODIES bodies — an import was not inlined"; exit 1; }
cc -O2 -DHAVE_CURL -rdynamic \
-I"$RTDIR" -I"$SSL/include" -L"$SSL/lib" \
"$GEN/soul.c" "$RTDIR/el_runtime.c" \
-lssl -lcrypto -lcurl -lpthread -lm \
-o "$OUT" 2> "$GEN/cc.log" || { echo "[build-soul] FAIL compile"; tail -40 "$GEN/cc.log"; exit 1; }
if grep -qE 'implicit.*(engram_|el_)' "$GEN/cc.log"; then
echo "[build-soul] FAIL: implicit declarations of runtime symbols"; grep -E 'implicit' "$GEN/cc.log" | head; exit 1; fi
echo "[build-soul] OK -> $OUT ($(wc -c < "$OUT" | tr -d ' ') bytes)"
+506
View File
@@ -0,0 +1,506 @@
#!/usr/bin/env python3
"""
build_gold_set.py derive the retrieval gold set FROM the corpus, and validate it.
WHY THIS FILE EXISTS AS CODE AND NOT AS A HAND-WRITTEN JSON
A gold set nobody can audit is vibes with extra steps. Every expected answer
here is either (a) mined from the corpus by a rule this script re-runs, or
(b) hand-selected with a stated criterion that this script then CHECKS
against the corpus. Both leave a `derivation` string on every query, and the
checks are re-run on demand so the set cannot silently rot as the corpus
changes.
Lineage: this extends the pinned-query approach from
docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py (pinned
ground-truth patterns + a --check "winnability" gate) and the per-query
relevant-id scoring from docs/research/graphrag_eval/score.py. What is new:
ids as ground truth rather than regexes alone, an ASSOCIATIVE category
derived from real graph edges, a superseded/contradicted category, and a
machine-checked no-lexical-overlap guarantee on the paraphrase category.
THE SIX CATEGORIES, AND WHAT EACH ONE IS FOR
exact_rare a single rare word. Substring matching already wins these.
They are a REGRESSION GUARD: any change that loses them is
disqualified regardless of what else it gains.
phrase a multi-word string that exists verbatim in the corpus.
Guards multi-token queries, which the old substring matcher
handled by returning nothing.
paraphrase same meaning, ZERO shared content words with the target node.
THE CATEGORY THAT MATTERS. Mechanically unreachable by string
matching; reachable only by semantics or by association.
associative the answer is one hub-hop from an obvious starting point and
shares no words with the query. This is the case the graph is
supposed to buy: query one value, get its siblings.
nonsense must return nothing. Guards against a retriever that "improves"
recall by returning the whole graph.
superseded a fact that was later corrected. The correction must OUTRANK
the stale version ranking, not mere presence.
usage:
python3 build_gold_set.py <snapshot.json> [--out gold_set.json] [--check]
--check re-validates an existing gold_set.json against the corpus and exits
non-zero if any query became unwinnable or any paraphrase leaked a word.
"""
import argparse
import json
import os
import re
import sys
from collections import Counter, defaultdict
HERE = os.path.dirname(os.path.abspath(__file__))
DEFAULT_OUT = os.path.join(HERE, "gold_set.json")
TOKEN = re.compile(r"[a-z0-9][a-z0-9\-']*")
# Stopwords are deliberately generous. A paraphrase query is only interesting if
# its CONTENT words are absent from the target; "the", "is", "what" appearing in
# both proves nothing. Being generous here makes the overlap test STRICTER on
# the words that carry meaning, which is the conservative direction.
STOP = set("""
a about above after again against all also am an and any are aren't as at be because been
before being below between both but by can can't cannot could couldn't did didn't do does
doesn't doing don't down during each few for from further had hadn't has hasn't have haven't
having he her here hers herself him himself his how i if in into is isn't it its itself just
me more most my myself no nor not of off on once only or other others ought our ours ourselves
out over own same shan't she should shouldn't so some such than that the their theirs them
themselves then there these they this those through to too under until up very was wasn't we
were weren't what when where which while who whom why will with won't would wouldn't you your
yours yourself yourselves get gets got make makes made take takes use uses used way ways thing
things does doing done keep keeps kept go goes going come comes came one two something anything
""".split())
# ─────────────────────────────────────────────────────────────────────────────
# corpus helpers
# ─────────────────────────────────────────────────────────────────────────────
def load_corpus(path):
with open(path, encoding="utf-8", errors="replace") as fh:
data = json.load(fh)
nodes = [n for n in data.get("nodes", []) if isinstance(n, dict) and n.get("id")]
edges = [e for e in data.get("edges", []) if isinstance(e, dict)]
return nodes, edges
def doctext(n):
return " ".join([str(n.get("label") or ""), str(n.get("content") or ""), str(n.get("tags") or "")])
def content_tokens(s):
return {t for t in TOKEN.findall(s.lower()) if t not in STOP and len(t) > 2}
# ─────────────────────────────────────────────────────────────────────────────
# hand-authored queries. Every entry states HOW its expected answer was chosen.
# The `check` field names the validation this script runs against the corpus.
# ─────────────────────────────────────────────────────────────────────────────
# EXACT_RARE — mined, not chosen. The rule (re-run by mine_exact_rare below):
# tokens whose document frequency across the whole corpus is 1, whose single
# containing node is a Memory/Knowledge/Belief with 300-6000 chars of content
# (so the answer is a real memory, not a 117KB whitepaper that contains every
# word in English), and whose token is plain lowercase alphabetic. The expected
# answer is that one node — it is the only node that can possibly be correct.
EXACT_RARE_SEEDS = [
"unjailbreakable",
"engram-migrate",
"cartabandonedevent",
"pre-apprenticeship",
"inferencenodemanager",
"clear-eyed",
]
# PHRASE — chosen by reading the corpus for phrases that (a) occur verbatim,
# (b) occur in a small enough set of nodes that "relevant" is well defined.
# Expected answers are computed here as EVERY node whose text contains the
# phrase case-insensitively — so the answer set is a fact about the corpus, not
# an opinion. Queries whose phrase matches more than PHRASE_MAX nodes are
# rejected by validation as too diffuse to score.
PHRASE_MAX = 25
PHRASE_SEEDS = [
("patterns not returns",
"a verbatim correction Will issued; expected = every node containing the phrase"),
("thirty moves",
"the canonical biographical phrase; expected = every node containing it"),
("Grandma Lucas",
"a named person appearing verbatim in the biography/value nodes"),
("Directed Harmonic",
"the canonical DHARMA expansion, confirmed by Will April 24 2026"),
("Sarah Bishop",
"a named person; rare enough that the answer set is unambiguous"),
("Directed Autonomous Runtime Modification",
"the DARMA expansion, quoted verbatim in the backlog item and its correction"),
("zero-knowledge encrypted backup",
"the paid-tier feature name as written in the roadmap nodes"),
]
# PARAPHRASE — hand-authored. THE SELECTION CRITERION, stated once and applied
# to all nine: pick a node whose SUBJECT is unmistakable to a reader, then write
# the query a person would actually type when they remember the subject but not
# the words. The target is then LOCKED by id, and this script enforces the hard
# property that makes the category meaningful: not one content word of the query
# appears anywhere in the target node's label, content, or tags. If a word
# leaks, validation fails and the query must be rewritten — the set cannot
# quietly degrade into a lexical query wearing a paraphrase costume.
PARAPHRASE_SEEDS = [
("kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"the elderly relative who passed while he stayed away",
"target: 'Value - Do the Essential Thing While You Can', whose subject is Grandma Lucas "
"dying in Feb 2006 without Will saying goodbye. Query names the event with none of the "
"node's own vocabulary."),
("kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"a soldier sidelined by illness who refused to quit",
"target: 'Value - Survival Is Not an Excuse to Stop', whose subject is enlisting in the "
"Marines, a severe hernia, and sepsis. Query describes the episode obliquely."),
("kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"choosing an uncomfortable fact over a pleasant fiction",
"target: 'Value - Honesty Before Comfort'. Query states the principle in wholly "
"different words."),
("kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"a tight payload beats a bloated one",
"target: 'Value - Precision Over Brute Force'. Query restates the claim with no "
"shared vocabulary."),
("kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"if you are able and nobody is coming the job is yours",
"target: 'Value - Capability Is a Debt You Owe the Moment'. Query states the "
"obligation without the node's terms."),
("kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"learning is the wealth creditors cannot seize",
"target: 'Value - Knowledge Survives When Nothing Else Does', whose subject is the "
"library following Will across 30+ moves."),
("kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"reliability proven by track record not assertion",
"target: 'Value - Earned Trust' ('Trust is demonstrated, not declared')."),
("kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"boundaries that enable instead of confine",
"target: 'Value - Constraints as Freedom'. Query is a restatement of the same claim."),
("kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"what shifts tells you where to cut a system apart",
"target: 'Value - Change Is the Signal', the value VBD is built on."),
("kn-f230b362-b201-4402-9833-4160c89ab3d4",
"a mind that compounds instead of resetting each day",
"target: 'Value - The System Must Accumulate'. Query is the accumulation claim in "
"different vocabulary."),
("kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"loved for the unedited self and not the polished exterior",
"target: 'Value - Being Seen Is Rarer Than Being Known', whose subject is Sarah Bishop "
"as the first person Will did not perform for."),
("kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"cheerfulness you arrive at instead of assuming",
"target: 'Value - Hope Is a Conclusion'. Query restates 'a conclusion, not a premise'."),
("kn-6061318f-046b-4935-907d-8eafdce14930",
"a childhood offering no solid foundation to inherit",
"target: 'Value - Structure Is Not Inherited', whose subject is thirty moves between "
"two parents' collapses."),
]
# ASSOCIATIVE — derived from real edges, not authored. The construction:
# every value node hangs off the 'Self - Values (grounded)' hub by an `identity`
# edge. For a chosen value node V, the query is built from V's own distinctive
# vocabulary; the expected answers are V's SIBLINGS on that hub. A sibling
# shares no query words with the query by construction (validated below), so the
# only path from the query to a sibling is: lexical seed on V -> hub -> sibling.
# That is a two-hop traversal and nothing else can produce it.
VALUES_HUB = "kn-5b606390-a52d-4ca2-8e0e-eba141d13440"
ASSOCIATIVE_SEEDS = [
("kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71", "Grandma Lucas stroke February 2006 goodbye window"),
("kn-58874a74-b96f-4883-9e08-45707f4bd3ee", "Marines hernia sepsis medical ward"),
("kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e", "Sarah Bishop Dyer trailer performance"),
("kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83", "Swarm Architecture containment lateral worker"),
("kn-e0423482-cfa5-4796-8689-8495c93b66bc", "hope won inside the narrative preface"),
("kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8", "man of the house six years old expectation"),
]
# NONSENSE — must return nothing. Strings chosen to be lexically impossible:
# validation asserts each appears in ZERO corpus nodes as a substring and that
# none of its tokens appears anywhere either (so not even a partial seed exists).
NONSENSE_SEEDS = [
"zqxjvw plimforth grebulon",
"flarnbistle quommetry",
"xxqzzt vurblenacht throom",
]
# SUPERSEDED — a fact that was corrected. Chosen by searching the corpus for
# explicit correction language and keeping pairs where BOTH the stale statement
# and its correction exist as separate nodes. Scored on RANKING: the correction
# must appear, and must appear above the stale node. Ids are locked here and
# validated to exist and to match their stated role.
SUPERSEDED_SEEDS = [
# (query, correct_id, stale_id, derivation)
]
# ─────────────────────────────────────────────────────────────────────────────
# mining
# ─────────────────────────────────────────────────────────────────────────────
def mine_exact_rare(nodes, byid, seeds):
"""Re-derive: confirm each seed token still has df==1 and name its node."""
tok = re.compile(r"[A-Za-z][A-Za-z0-9\-]{4,}")
want = set(seeds)
df = Counter()
post = defaultdict(set)
for n in nodes:
for t in {w.lower() for w in tok.findall(doctext(n))}:
if t in want:
df[t] += 1
post[t].add(n["id"])
out = []
for s in seeds:
ids = sorted(post.get(s, ()))
out.append((s, ids, df.get(s, 0)))
return out
def phrase_matches(nodes, phrase):
p = phrase.lower()
return sorted(n["id"] for n in nodes if p in doctext(n).lower())
def hub_siblings(edges, hub, relation="identity"):
sibs = []
for e in edges:
if e.get("from_id") == hub and e.get("relation") == relation:
sibs.append(e["to_id"])
elif e.get("to_id") == hub and e.get("relation") == relation:
sibs.append(e["from_id"])
return list(dict.fromkeys(sibs))
def find_superseded_pairs(nodes, byid):
"""Locked pairs, each verified here to exist and to carry its stated marker.
Chosen by scanning the corpus for explicit correction language
(CORRECTION/SUPERSEDES/re-corrected/no longer/RECONCILED) and keeping only
cases where the STALE claim also survives as its own node a supersession
with nothing to outrank is not a ranking test.
"""
pairs = []
txt = {n["id"]: doctext(n) for n in nodes}
def find_one(pattern, exclude=()):
rx = re.compile(pattern)
return [n["id"] for n in nodes
if n["id"] not in exclude
and rx.search(txt[n["id"]])
and 150 < len(str(n.get("content") or "")) < 12000
and n.get("node_type") in ("Memory", "Knowledge", "Belief", "BacklogItem")]
# Each entry: (query, correction-pattern, stale-pattern, why).
# The stale side is searched with the correction hits EXCLUDED, because most
# correction memories quote the claim they are killing — without the
# exclusion the "stale" node resolves to the correction itself and the pair
# collapses into a no-op. A pair is only emitted if both sides resolve to
# DIFFERENT surviving nodes; otherwise it is dropped and reported.
SPECS = [
("is the self-improvement architecture called DARMA or DHARMA",
r'(?i)CORRECTION:.{0,90}DHARMA .{0,12}not DARMA',
r'(?i)\bDARMA\b',
"correction node is Will's confirmation that the H is intentional (DHARMA, not DARMA); "
"the stale node is the surviving backlog item still titled 'Implement DARMA'."),
("how many provisional patents does Will actually have",
r'(?i)EXACTLY 6 (fully-specced )?provisional',
r'(?i)(MY ARCHITECTURE = 12 filed patents|\b12 filed patents\b)',
"correction node is the 2026-06-17 confabulation flag establishing EXACTLY 6 provisionals; "
"the stale node is the surviving memory that asserts 12 filed patents."),
("is MCP still the live integration layer",
r'(?i)MCP RETIRED',
r'(?i)MCP server live at',
"correction node is the 'CGI ARCHITECTURE - THREE LAYERS, MCP RETIRED' decision of "
"April 30 2026; the stale node still records the MCP server as live."),
("what does the patterns-not-returns directive mean",
r'(?i)CORRECTION:.{0,80}patterns not returns',
r'(?i)established returns',
"correction node is Will's 'patterns not returns' correction; the stale node is a "
"surviving node carrying the misread 'established returns' directive."),
("was the earlier identity-bug finding correct",
r'(?i)SUPERSEDES the earlier .critical identity bug',
r'(?i)critical identity bug',
"correction node explicitly supersedes the 'critical identity bug' finding; the stale "
"node is the surviving original finding."),
("does Neuron have recursive self-improvement",
r'(?i)twice answered .Neuron has no recursive self-improvement',
r'(?i)no recursive self-improvement',
"correction node records the June-29 finding that the CGI provisional IS the "
"recursive-self-improvement mechanism; the stale node is the surviving denial."),
]
for query, cpat, spat, why in SPECS:
corr = find_one(cpat)
if not corr:
continue
stale = find_one(spat, exclude=set(corr))
if not stale:
continue
pairs.append((query, corr[0], stale[0], why))
return pairs
# ─────────────────────────────────────────────────────────────────────────────
# build
# ─────────────────────────────────────────────────────────────────────────────
def build(nodes, edges):
byid = {n["id"]: n for n in nodes}
tokset = {n["id"]: content_tokens(doctext(n)) for n in nodes}
queries = []
problems = []
qn = [0]
def add(cat, query, relevant, derivation, **extra):
qn[0] += 1
q = {
"id": f"q{qn[0]:02d}",
"category": cat,
"query": query,
"relevant": sorted(relevant),
"derivation": derivation,
}
q.update(extra)
queries.append(q)
return q
# --- exact_rare ---------------------------------------------------------
for tokname, ids, df in mine_exact_rare(nodes, byid, EXACT_RARE_SEEDS):
if df != 1 or len(ids) != 1:
problems.append(f"exact_rare '{tokname}': df={df}, ids={len(ids)} (expected df=1)")
continue
lab = (byid[ids[0]].get("label") or "")[:60]
add("exact_rare", tokname, ids,
f"MINED: token '{tokname}' has document frequency 1 over all {len(nodes)} corpus nodes "
f"(re-verified at build time). Its single containing node is {ids[0]} "
f"('{lab}'), which is therefore the only possible correct answer.")
# --- phrase -------------------------------------------------------------
for phrase, why in PHRASE_SEEDS:
ids = phrase_matches(nodes, phrase)
if not ids:
problems.append(f"phrase '{phrase}': 0 corpus matches — unwinnable")
continue
if len(ids) > PHRASE_MAX:
problems.append(f"phrase '{phrase}': {len(ids)} matches > {PHRASE_MAX} — too diffuse")
continue
add("phrase", phrase, ids,
f"MINED: {why}. Case-insensitive verbatim substring scan over label+content+tags at "
f"build time returns exactly {len(ids)} node(s); that set IS the answer key.")
# --- paraphrase ---------------------------------------------------------
for target, query, why in PARAPHRASE_SEEDS:
if target not in byid:
problems.append(f"paraphrase target {target} not in corpus")
continue
qt = content_tokens(query)
leak = sorted(qt & tokset[target])
if leak:
problems.append(f"paraphrase '{query}': leaks {leak} into target {target}")
continue
add("paraphrase", query, [target],
f"HAND-SELECTED with criterion: {why} VERIFIED at build time: of the {len(qt)} content "
f"words in the query, ZERO appear anywhere in the target's label, content, or tags — so "
f"no string-matching retriever can reach this answer.",
zero_overlap_verified=True, query_content_words=sorted(qt))
# --- associative --------------------------------------------------------
sibs = hub_siblings(edges, VALUES_HUB)
if len(sibs) < 5:
problems.append(f"associative: values hub {VALUES_HUB} has only {len(sibs)} siblings")
for src, query in ASSOCIATIVE_SEEDS:
if src not in byid or src not in sibs:
problems.append(f"associative source {src} not a sibling on {VALUES_HUB}")
continue
qt = content_tokens(query)
others = [s for s in sibs if s != src and s in byid]
# A sibling only counts as a legitimate expected answer if the query
# cannot reach it lexically. Drop any sibling that shares a content word.
clean = [s for s in others if not (qt & tokset[s])]
dropped = len(others) - len(clean)
if len(clean) < 5:
problems.append(f"associative '{query}': only {len(clean)} lexically-unreachable siblings")
continue
add("associative", query, clean,
f"DERIVED FROM EDGES: the query is built from the distinctive vocabulary of {src} "
f"('{(byid[src].get('label') or '')[:48]}'), which hangs off the values hub {VALUES_HUB} "
f"by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub "
f"({len(clean)} of {len(others)}; {dropped} dropped because they shared a query word and "
f"so were lexically reachable). Every remaining sibling shares ZERO content words with "
f"the query — the only route from query to answer is seed({src}) -> hub -> sibling, a "
f"two-hop traversal.",
associative_source=src, hub=VALUES_HUB, siblings_dropped_for_overlap=dropped)
# --- nonsense -----------------------------------------------------------
all_tokens = set()
for n in nodes:
all_tokens |= {t for t in TOKEN.findall(doctext(n).lower())}
for s in NONSENSE_SEEDS:
present = sorted(t for t in TOKEN.findall(s.lower()) if t in all_tokens)
if present:
problems.append(f"nonsense '{s}': tokens {present} DO occur in corpus")
continue
add("nonsense", s, [],
f"CONTROL: verified at build time that none of this string's tokens occurs anywhere in "
f"the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
expect_empty=True)
# --- superseded ---------------------------------------------------------
for query, correct, stale, why in find_superseded_pairs(nodes, byid):
if correct not in byid or stale not in byid:
problems.append(f"superseded '{query}': id missing from corpus")
continue
add("superseded", query, [correct],
f"DERIVED: {why} Scored on RANKING, not presence: the corrected node {correct} must be "
f"returned AND must rank above the stale node {stale}.",
must_outrank=[correct, stale],
stale_id=stale,
correct_label=(byid[correct].get("label") or "")[:70],
stale_label=(byid[stale].get("label") or "")[:70])
return queries, problems
def summarize(queries):
c = Counter(q["category"] for q in queries)
return ", ".join(f"{k}={c[k]}" for k in
("exact_rare", "phrase", "paraphrase", "associative", "nonsense", "superseded")
if c[k])
def main():
ap = argparse.ArgumentParser()
ap.add_argument("snapshot")
ap.add_argument("--out", default=DEFAULT_OUT)
ap.add_argument("--check", action="store_true",
help="validate only; do not write. Non-zero exit if anything is unwinnable.")
args = ap.parse_args()
nodes, edges = load_corpus(args.snapshot)
print(f"corpus: {len(nodes)} nodes, {len(edges)} edges ({os.path.basename(args.snapshot)})")
queries, problems = build(nodes, edges)
print(f"gold set: {len(queries)} queries [{summarize(queries)}]")
if problems:
print(f"\n{len(problems)} PROBLEM(S) — these queries were REJECTED, not silently kept:")
for p in problems:
print(" -", p)
if args.check:
sys.exit(1 if problems else 0)
doc = {
"corpus": os.path.abspath(args.snapshot),
"corpus_nodes": len(nodes),
"corpus_edges": len(edges),
"note": ("Every query carries a `derivation` recording how its expected answer was chosen. "
"Re-run with --check to re-validate the whole set against the corpus."),
"queries": queries,
}
with open(args.out, "w", encoding="utf-8") as fh:
json.dump(doc, fh, indent=1, ensure_ascii=False)
print(f"\nwrote {args.out}")
if __name__ == "__main__":
main()
+43
View File
@@ -0,0 +1,43 @@
import json,sys,pickle,numpy as np,itertools
sys.path.insert(0,'.')
from policy2 import legs3,outcome,G,NODES,merge
# cache per-query leg id-lists, floored and unfloored
cache={}
for q in G['queries']:
Lf,Sf,Af=legs3(q['query'])
Lu,Su,Au=legs3(q['query'],unfloor=True)
cache[q['id']]=dict(L=Lf,Sf=Sf,A=Af,Su=Su,Au=Au)
pickle.dump(cache,open('ceil.pkl','wb'))
def mrg(pattern,L,S,A,lim=10):
out=[];p={'L':0,'S':0,'A':0};src={'L':L,'S':S,'A':A}
i=0
while len(out)<lim:
prog=False
for ch in pattern:
lst=src[ch]
if p[ch]<len(lst):
x=lst[p[ch]];p[ch]+=1;prog=True
if x not in out: out.append(x)
if len(out)>=lim: return out
if not prog: break
return out
def ev(pattern,unfl):
res={}
for q in G['queries']:
c=cache[q['id']]
S=c['Su'] if unfl else c['Sf']
ids=[NODES[i]['id'] for i in mrg(pattern,c['L'],S,c['A'],10)]
res[q['id']]=outcome(q,ids)
return res
base=ev('LSA',False)
print("baseline",sum(base.values()))
best=[]
pats=['LSA','LAS','SLA','ALS','SAL','ASL','LSSA','LSASA','LSAA','LSSAA','LSAS','SSLA','LLSA','SALSA','LSAAS']
for unfl in (False,True):
for p in pats:
r=ev(p,unfl)
g=sorted(k for k in base if r[k] and not base[k]);l=sorted(k for k in base if base[k] and not r[k])
best.append((len(g)-len(l),p,unfl,g,l))
best.sort(reverse=True)
for n,p,u,g,l in best[:10]:
print("net=%+d pat=%-6s unfloor=%s gains=%s losses=%s"%(n,p,u,g,l))
+146
View File
@@ -0,0 +1,146 @@
{
"baseline": "bm25lex",
"candidate": "wsclaim24",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q14",
"q25"
],
"broken_by_candidate": [
"q15",
"q28",
"q33",
"q34"
],
"discordant": 6,
"net_queries": -2,
"mcnemar_exact_p": 0.6875,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1184.4,
"latency_ms_p95": 1620.0,
"latency_ms_max": 1655.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5768475572047,
"recall@10": 0.6563414759843332,
"precision@5": 0.19428571428571437,
"mrr@10": 0.5026530612244898,
"nonsense_clean": "0/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 524.8,
"latency_ms_p95": 738.7,
"latency_ms_max": 755.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 0,
"avg_false_positives": 10.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {
"baseline": {
"runs": 2,
"hit@5_min": 0.7428571428571429,
"hit@5_max": 0.7428571428571429,
"spread_queries": 0
}
}
}
+149
View File
@@ -0,0 +1,149 @@
{
"baseline": "unfloor-clean",
"candidate": "splitfix",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q15",
"q28",
"q60"
],
"broken_by_candidate": [],
"discordant": 3,
"net_queries": 3,
"mcnemar_exact_p": 0.25,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5384615384615384,
"recall@5": 0.44907176157176154,
"recall@10": 0.5380300255300255,
"precision@5": 0.13230769230769232,
"mrr@10": 0.32437728937728944,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 632.5,
"latency_ms_p95": 992.5,
"latency_ms_max": 1177.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3,
"recall@5": 0.3,
"recall@10": 0.4,
"mrr@10": 0.11638888888888889
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5846153846153846,
"recall@5": 0.48102442429365505,
"recall@10": 0.5692415490492414,
"precision@5": 0.14153846153846153,
"mrr@10": 0.3351709401709402,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 646.9,
"latency_ms_p95": 1028.5,
"latency_ms_max": 1197.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.15967365967365968,
"mrr@10": 0.24166666666666667
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.43333333333333335,
"mrr@10": 0.12120370370370372
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.7692307692307693,
"recall@5": 0.7692307692307693,
"recall@10": 0.8461538461538461,
"mrr@10": 0.33269230769230773
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5775226757369615,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
+210
View File
@@ -0,0 +1,210 @@
#!/usr/bin/env python3
"""
compare.py diff two run_eval.py result files, WITH a noise threshold.
WHY THE STATISTICS ARE NOT OPTIONAL
With ~35 scored queries, one query is ~2.9 percentage points. A harness that
reports "hit@5 improved 2.9%" without saying that is one query is a harness
that will approve noise. So this file refuses to call anything an
improvement on the strength of the headline number alone. It reports:
1. The DISCORDANT PAIRS. Two configurations scored on the same queries are
paired data, so the only queries carrying information are the ones
where they disagree: b = fixed by B, c = broken by B. Queries both got
right, or both got wrong, tell you nothing about which is better.
2. McNEMAR'S EXACT TEST on (b, c). Under the null "the change is a coin
flip", the discordant outcomes are Binomial(b+c, 0.5). The two-sided
exact p-value is computed here with no scipy dependency.
3. The MINIMUM DETECTABLE SWING for this gold set: the smallest number of
net-changed queries that would reach p < 0.05 if every discordant pair
fell the same way. Anything smaller is inside the noise band, and the
verdict line says so in those words.
Repeat-run variance is the other half of honesty. Spreading activation is a
stateful read (it reinforces what it touches), so identical inputs need not
give identical outputs. Pass --repeats to fold several runs of the same
config into an observed variance band; a delta inside that band is not real
either, however good its p-value looks.
usage:
python3 compare.py --baseline results-main.json --candidate results-act.json
python3 compare.py --baseline a.json --candidate b.json \
--repeats-baseline a2.json a3.json --repeats-candidate b2.json b3.json
"""
import argparse
import json
from math import comb
def binom_two_sided(b, c):
"""Two-sided exact binomial p for b successes in n=b+c at p=0.5."""
n = b + c
if n == 0:
return 1.0
k = min(b, c)
tail = sum(comb(n, i) for i in range(0, k + 1)) / (2 ** n)
return min(1.0, 2 * tail)
def min_detectable_swing(n_scored, alpha=0.05):
"""Smallest all-one-way discordant count reaching p < alpha.
If every query that changes changes in the same direction, the p-value is
2 * 0.5**n. Solve for the smallest n where that drops under alpha. This is
the FLOOR: any real change will have some discordance both ways, so the true
requirement is larger. Reporting the floor is the conservative move it is
the most generous threshold we would ever accept.
"""
n = 1
while n <= n_scored:
if 2 * (0.5 ** n) < alpha:
return n
n += 1
return n_scored
def load(path):
with open(path, encoding="utf-8") as fh:
return json.load(fh)
def row_map(doc):
return {r["id"]: r for r in doc["rows"]}
def outcome(r):
"""Binary per-query outcome used for the paired test.
hit@5 for scored queries; 'returned nothing' for the nonsense controls;
'correction outranks the stale node' for the superseded queries. One number
per query, so every query votes exactly once.
"""
if "clean" in r:
return 1.0 if r["clean"] else 0.0
if "outranks" in r:
return 1.0 if r["outranks"] else 0.0
return r.get("hit@5") or 0.0
def band(values):
return (min(values), max(values))
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--baseline", required=True)
ap.add_argument("--candidate", required=True)
ap.add_argument("--repeats-baseline", nargs="*", default=[])
ap.add_argument("--repeats-candidate", nargs="*", default=[])
ap.add_argument("--out", default=None)
args = ap.parse_args()
A, B = load(args.baseline), load(args.candidate)
ra, rb = row_map(A), row_map(B)
ids = [q for q in ra if q in rb]
n = len(ids)
aa, ab = A["aggregate"], B["aggregate"]
print(f"baseline {A['label']:14} soul={A['soul_md5'][:12]} {n} shared queries")
print(f"candidate {B['label']:14} soul={B['soul_md5'][:12]}")
print(f"corpus {A['corpus_nodes']} nodes / {A['corpus_edges']} edges "
f"(identical copy for both runs)\n")
metrics = [("hit@5", 1), ("recall@5", 1), ("recall@10", 1),
("precision@5", 1), ("mrr@10", 0)]
print(f" {'metric':14} {'baseline':>10} {'candidate':>10} {'delta':>10}")
for m, as_pct in metrics:
x, y = aa[m], ab[m]
if as_pct:
print(f" {m:14} {100*x:>9.1f}% {100*y:>9.1f}% {100*(y-x):>+9.1f}pp")
else:
print(f" {m:14} {x:>10.3f} {y:>10.3f} {y-x:>+10.3f}")
for m in ("latency_ms_p50", "latency_ms_p95"):
x, y = aa[m], ab[m]
ratio = f"{y/x:.2f}x" if x else "n/a"
print(f" {m:14} {x:>9.0f}ms {y:>9.0f}ms {ratio:>10}")
print(f" {'nonsense':14} {aa['nonsense_clean']:>10} {ab['nonsense_clean']:>10}")
print(f" {'outranks':14} {aa['superseded_outranks']:>10} {ab['superseded_outranks']:>10}")
print(f"\n {'category':14} {'n':>3} {'base hit@5':>11} {'cand hit@5':>11} {'delta':>9}")
for c in sorted(set(aa["by_category"]) & set(ab["by_category"])):
ea, eb = aa["by_category"][c], ab["by_category"][c]
if c == "nonsense":
print(f" {c:14} {ea['n']:>3} {'clean ' + str(ea['clean']):>11} "
f"{'clean ' + str(eb['clean']):>11}")
else:
print(f" {c:14} {ea['n']:>3} {100*ea['hit@5']:>10.1f}% {100*eb['hit@5']:>10.1f}% "
f"{100*(eb['hit@5']-ea['hit@5']):>+8.1f}pp")
# ---- paired significance -------------------------------------------------
fixed, broken = [], []
for q in ids:
oa, ob = outcome(ra[q]), outcome(rb[q])
if ob > oa:
fixed.append(q)
elif ob < oa:
broken.append(q)
b, c = len(fixed), len(broken)
p = binom_two_sided(b, c)
mds = min_detectable_swing(n)
print(f"\n== paired comparison over {n} queries ==")
print(f" fixed by candidate : {b} {[ra[q]['category'] + ':' + q for q in fixed]}")
print(f" broken by candidate: {c} {[ra[q]['category'] + ':' + q for q in broken]}")
print(f" discordant pairs : {b + c} net {b - c:+d} queries")
print(f" McNemar exact p : {p:.4f}")
print(f" noise threshold : a difference needs at least {mds} queries moving the "
f"same way to clear p<0.05 on this {n}-query set")
# ---- repeat-run variance -------------------------------------------------
var = {}
for name, paths, first in (("baseline", args.repeats_baseline, A),
("candidate", args.repeats_candidate, B)):
docs = [first] + [load(p) for p in paths]
if len(docs) > 1:
hits = [d["aggregate"]["hit@5"] for d in docs]
lo, hi = band(hits)
spread_q = round((hi - lo) * first["aggregate"]["n_scored"])
var[name] = {"runs": len(docs), "hit@5_min": lo, "hit@5_max": hi,
"spread_queries": spread_q}
print(f" {name} repeat runs ({len(docs)}): hit@5 {100*lo:.1f}%..{100*hi:.1f}% "
f"= {spread_q} query of run-to-run drift")
drift = max([v["spread_queries"] for v in var.values()], default=0)
floor = max(mds, drift + 1)
print("\n== VERDICT ==")
net = b - c
if abs(net) < floor:
print(f" NO MEASURABLE DIFFERENCE. Net {net:+d} queries is inside the noise band "
f"(needs |net| >= {floor}: {mds} for significance, {drift} observed run-to-run drift).")
elif net > 0:
print(f" CANDIDATE BETTER by {net} queries (p={p:.4f}), outside the noise band "
f"(>= {floor}).")
else:
print(f" CANDIDATE WORSE by {abs(net)} queries (p={p:.4f}), outside the noise band "
f"(>= {floor}).")
if args.out:
with open(args.out, "w", encoding="utf-8") as fh:
json.dump({
"baseline": A["label"], "candidate": B["label"],
"n_shared_queries": n,
"fixed_by_candidate": fixed, "broken_by_candidate": broken,
"discordant": b + c, "net_queries": net,
"mcnemar_exact_p": p,
"min_detectable_swing_queries": mds,
"observed_run_to_run_drift_queries": drift,
"noise_floor_queries": floor,
"verdict": ("no measurable difference" if abs(net) < floor
else ("candidate better" if net > 0 else "candidate worse")),
"baseline_aggregate": aa, "candidate_aggregate": ab,
"repeat_variance": var,
}, fh, indent=1)
print(f"\nwrote {args.out}")
if __name__ == "__main__":
main()
@@ -0,0 +1,150 @@
{
"baseline": "assoc-leg",
"candidate": "semseed",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q18",
"q19",
"q22"
],
"broken_by_candidate": [
"q11"
],
"discordant": 4,
"net_queries": 2,
"mcnemar_exact_p": 0.625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6285714285714286,
"recall@5": 0.45309194773480493,
"recall@10": 0.5405733155733157,
"precision@5": 0.17714285714285719,
"mrr@10": 0.42650793650793645,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1228.5,
"latency_ms_p95": 1681.8,
"latency_ms_max": 1718.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657342,
"recall@10": 0.24825174825174826,
"mrr@10": 0.22777777777777777
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4882369614512472,
"recall@10": 0.6329365079365079,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6857142857142857,
"recall@5": 0.5213459159887731,
"recall@10": 0.6027048348476919,
"precision@5": 0.18285714285714294,
"mrr@10": 0.4608730158730158,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1227.1,
"latency_ms_p95": 1692.6,
"latency_ms_max": 1710.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657344,
"recall@10": 0.24825174825174823,
"mrr@10": 0.20833333333333334
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 0.7142857142857143,
"recall@5": 0.40093537414965985,
"recall@10": 0.5150226757369615,
"mrr@10": 0.6507936507936508
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {
"baseline": {
"runs": 2,
"hit@5_min": 0.6285714285714286,
"hit@5_max": 0.6285714285714286,
"spread_queries": 0
},
"candidate": {
"runs": 2,
"hit@5_min": 0.6857142857142857,
"hit@5_max": 0.6857142857142857,
"spread_queries": 0
}
}
}
@@ -0,0 +1,146 @@
{
"baseline": "bm25lex",
"candidate": "wordstart",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q35"
],
"broken_by_candidate": [],
"discordant": 1,
"net_queries": 1,
"mcnemar_exact_p": 1.0,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1184.4,
"latency_ms_p95": 1620.0,
"latency_ms_max": 1655.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "3/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 542.6,
"latency_ms_p95": 741.3,
"latency_ms_max": 758.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 3,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {
"baseline": {
"runs": 2,
"hit@5_min": 0.7428571428571429,
"hit@5_max": 0.7428571428571429,
"spread_queries": 0
},
"candidate": {
"runs": 2,
"hit@5_min": 0.7428571428571429,
"hit@5_max": 0.7428571428571429,
"spread_queries": 0
}
}
}
@@ -0,0 +1,143 @@
{
"baseline": "hybrid-semantic",
"candidate": "assoc-leg",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q27",
"q28",
"q29",
"q31"
],
"broken_by_candidate": [],
"discordant": 4,
"net_queries": 4,
"mcnemar_exact_p": 0.125,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.5142857142857142,
"recall@5": 0.4409013605442177,
"recall@10": 0.5047619047619047,
"precision@5": 0.15428571428571433,
"mrr@10": 0.38746031746031745,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1219.7,
"latency_ms_p95": 1667.1,
"latency_ms_max": 1720.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6285714285714286,
"recall@5": 0.45309194773480493,
"recall@10": 0.5405733155733157,
"precision@5": 0.17714285714285719,
"mrr@10": 0.42650793650793645,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1228.5,
"latency_ms_p95": 1681.8,
"latency_ms_max": 1718.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657342,
"recall@10": 0.24825174825174826,
"mrr@10": 0.22777777777777777
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4882369614512472,
"recall@10": 0.6329365079365079,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"repeat_variance": {
"baseline": {
"runs": 2,
"hit@5_min": 0.5142857142857142,
"hit@5_max": 0.5142857142857142,
"spread_queries": 0
}
}
}
@@ -0,0 +1,154 @@
{
"baseline": "hybrid-semantic",
"candidate": "semseed",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q18",
"q19",
"q22",
"q27",
"q28",
"q29",
"q31"
],
"broken_by_candidate": [
"q11"
],
"discordant": 8,
"net_queries": 6,
"mcnemar_exact_p": 0.0703125,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "candidate better",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.5142857142857142,
"recall@5": 0.4409013605442177,
"recall@10": 0.5047619047619047,
"precision@5": 0.15428571428571433,
"mrr@10": 0.38746031746031745,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1219.7,
"latency_ms_p95": 1667.1,
"latency_ms_max": 1720.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6857142857142857,
"recall@5": 0.5213459159887731,
"recall@10": 0.6027048348476919,
"precision@5": 0.18285714285714294,
"mrr@10": 0.4608730158730158,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1227.1,
"latency_ms_p95": 1692.6,
"latency_ms_max": 1710.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657344,
"recall@10": 0.24825174825174823,
"mrr@10": 0.20833333333333334
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 0.7142857142857143,
"recall@5": 0.40093537414965985,
"recall@10": 0.5150226757369615,
"mrr@10": 0.6507936507936508
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {
"baseline": {
"runs": 2,
"hit@5_min": 0.5142857142857142,
"hit@5_max": 0.5142857142857142,
"spread_queries": 0
},
"candidate": {
"runs": 2,
"hit@5_min": 0.6857142857142857,
"hit@5_max": 0.6857142857142857,
"spread_queries": 0
}
}
}
@@ -0,0 +1,145 @@
{
"baseline": "baseline-embcorpus",
"candidate": "hybrid-semantic",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q15",
"q16",
"q20",
"q21",
"q26",
"q37"
],
"broken_by_candidate": [],
"discordant": 6,
"net_queries": 6,
"mcnemar_exact_p": 0.03125,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "candidate better",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1145.9,
"latency_ms_p95": 1574.3,
"latency_ms_max": 1634.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.5142857142857142,
"recall@5": 0.4409013605442177,
"recall@10": 0.5047619047619047,
"precision@5": 0.15428571428571433,
"mrr@10": 0.38746031746031745,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1219.7,
"latency_ms_p95": 1667.1,
"latency_ms_max": 1720.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"repeat_variance": {
"candidate": {
"runs": 2,
"hit@5_min": 0.5142857142857142,
"hit@5_max": 0.5142857142857142,
"spread_queries": 0
}
}
}
@@ -0,0 +1,150 @@
{
"baseline": "main-r1",
"candidate": "act-r1",
"n_shared_queries": 38,
"fixed_by_candidate": [],
"broken_by_candidate": [
"q07",
"q11",
"q12",
"q13",
"q36"
],
"discordant": 5,
"net_queries": -5,
"mcnemar_exact_p": 0.0625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 1,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1140.4,
"latency_ms_p95": 1584.1,
"latency_ms_max": 1627.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.22857142857142856,
"recall@5": 0.19087301587301586,
"recall@10": 0.24277210884353742,
"precision@5": 0.07428571428571429,
"mrr@10": 0.24154195011337865,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3208.8,
"latency_ms_p95": 4851.9,
"latency_ms_max": 5078.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.2857142857142857,
"recall@5": 0.09722222222222222,
"recall@10": 0.3567176870748299,
"mrr@10": 0.3505668934240363
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"repeat_variance": {
"baseline": {
"runs": 3,
"hit@5_min": 0.34285714285714286,
"hit@5_max": 0.34285714285714286,
"spread_queries": 0
},
"candidate": {
"runs": 3,
"hit@5_min": 0.22857142857142856,
"hit@5_max": 0.2571428571428571,
"spread_queries": 1
}
}
}
@@ -0,0 +1,166 @@
{
"baseline": "main-ext",
"candidate": "stack-ext",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q10",
"q15",
"q16",
"q18",
"q19",
"q20",
"q21",
"q22",
"q26",
"q27",
"q28",
"q29",
"q31",
"q35",
"q37",
"q40",
"q44",
"q48",
"q49",
"q50"
],
"broken_by_candidate": [],
"discordant": 20,
"net_queries": 20,
"mcnemar_exact_p": 1.9073486328125e-06,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "candidate better",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.18461538461538463,
"recall@5": 0.14510073260073258,
"recall@10": 0.1794871794871795,
"precision@5": 0.06461538461538462,
"mrr@10": 0.15847985347985344,
"nonsense_clean": "9/10",
"superseded_outranks": "1/3",
"latency_ms_p50": 1380.2,
"latency_ms_p95": 2293.8,
"latency_ms_max": 2879.1,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"nonsense": {
"n": 10,
"clean": 9,
"avg_false_positives": 1.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.47692307692307695,
"recall@5": 0.3750415183107491,
"recall@10": 0.45561340369032677,
"precision@5": 0.12307692307692313,
"mrr@10": 0.3055555555555555,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 640.9,
"latency_ms_p95": 1011.1,
"latency_ms_max": 1190.3,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.16666666666666666,
"recall@5": 0.16666666666666666,
"recall@10": 0.26666666666666666,
"mrr@10": 0.0762037037037037
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,147 @@
{
"baseline": "semseed",
"candidate": "bm25lex",
"n_shared_queries": 38,
"fixed_by_candidate": [
"q10",
"q11"
],
"broken_by_candidate": [],
"discordant": 2,
"net_queries": 2,
"mcnemar_exact_p": 0.5,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6857142857142857,
"recall@5": 0.5213459159887731,
"recall@10": 0.6027048348476919,
"precision@5": 0.18285714285714294,
"mrr@10": 0.4608730158730158,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1227.1,
"latency_ms_p95": 1692.6,
"latency_ms_max": 1710.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657344,
"recall@10": 0.24825174825174823,
"mrr@10": 0.20833333333333334
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 0.7142857142857143,
"recall@5": 0.40093537414965985,
"recall@10": 0.5150226757369615,
"mrr@10": 0.6507936507936508
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1184.4,
"latency_ms_p95": 1620.0,
"latency_ms_max": 1655.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {
"baseline": {
"runs": 2,
"hit@5_min": 0.6857142857142857,
"hit@5_max": 0.6857142857142857,
"spread_queries": 0
},
"candidate": {
"runs": 2,
"hit@5_min": 0.7428571428571429,
"hit@5_max": 0.7428571428571429,
"spread_queries": 0
}
}
}
@@ -0,0 +1,155 @@
{
"baseline": "stack-ext",
"candidate": "unfloor-clean",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q14",
"q25",
"q43",
"q52",
"q63",
"q67"
],
"broken_by_candidate": [
"q15",
"q28"
],
"discordant": 8,
"net_queries": 4,
"mcnemar_exact_p": 0.2890625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.47692307692307695,
"recall@5": 0.3750415183107491,
"recall@10": 0.45561340369032677,
"precision@5": 0.12307692307692313,
"mrr@10": 0.3055555555555555,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 640.9,
"latency_ms_p95": 1011.1,
"latency_ms_max": 1190.3,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.16666666666666666,
"recall@5": 0.16666666666666666,
"recall@10": 0.26666666666666666,
"mrr@10": 0.0762037037037037
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5384615384615384,
"recall@5": 0.44907176157176154,
"recall@10": 0.5380300255300255,
"precision@5": 0.13230769230769232,
"mrr@10": 0.32437728937728944,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 632.5,
"latency_ms_p95": 992.5,
"latency_ms_max": 1177.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3,
"recall@5": 0.3,
"recall@10": 0.4,
"mrr@10": 0.11638888888888889
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,158 @@
{
"baseline": "unfloor-clean",
"candidate": "semsub",
"n_shared_queries": 75,
"fixed_by_candidate": [
"q24",
"q39",
"q42"
],
"broken_by_candidate": [
"q18",
"q19",
"q22",
"q31",
"q43",
"q44",
"q52",
"q63"
],
"discordant": 11,
"net_queries": -5,
"mcnemar_exact_p": 0.2265625,
"min_detectable_swing_queries": 6,
"observed_run_to_run_drift_queries": 0,
"noise_floor_queries": 6,
"verdict": "no measurable difference",
"baseline_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.5384615384615384,
"recall@5": 0.44907176157176154,
"recall@10": 0.5380300255300255,
"precision@5": 0.13230769230769232,
"mrr@10": 0.32437728937728944,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 632.5,
"latency_ms_p95": 992.5,
"latency_ms_max": 1177.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.5,
"recall@5": 0.07575757575757576,
"recall@10": 0.13636363636363635,
"mrr@10": 0.23214285714285712
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.3,
"recall@5": 0.3,
"recall@10": 0.4,
"mrr@10": 0.11638888888888889
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6923076923076923,
"recall@5": 0.6923076923076923,
"recall@10": 0.7692307692307693,
"mrr@10": 0.29423076923076924
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5335884353741497,
"recall@10": 0.5933956916099773,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"candidate_aggregate": {
"n_queries": 75,
"n_scored": 65,
"hit@5": 0.46153846153846156,
"recall@5": 0.3889430014430015,
"recall@10": 0.4887681762681762,
"precision@5": 0.12307692307692313,
"mrr@10": 0.30181318681318675,
"nonsense_clean": "10/10",
"superseded_outranks": "2/3",
"latency_ms_p50": 634.4,
"latency_ms_p95": 988.3,
"latency_ms_max": 1184.5,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.3333333333333333,
"recall@5": 0.06060606060606061,
"recall@10": 0.12121212121212122,
"mrr@10": 0.19047619047619047
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"heldout_paraphrase": {
"n": 30,
"hit@5": 0.23333333333333334,
"recall@5": 0.23333333333333334,
"recall@10": 0.3,
"mrr@10": 0.08925925925925927
},
"nonsense": {
"n": 10,
"clean": 10,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.5384615384615384,
"recall@5": 0.5384615384615384,
"recall@10": 0.7692307692307693,
"mrr@10": 0.26324786324786326
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5596655328798186,
"recall@10": 0.5775226757369615,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"repeat_variance": {}
}
@@ -0,0 +1,43 @@
import json,sys,time,urllib.request,threading,queue
SRC="/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json"
OUT=sys.argv[1]
URL="http://127.0.0.1:11434/api/embeddings"; MODEL="nomic-embed-text"
MAXB=2000 # ENGRAM_EMBED_MAX_CHARS, applied to bytes as the C code does
d=json.load(open(SRC,encoding='utf-8',errors='surrogateescape'))
tasks=[]
for n in d["nodes"]:
c=n.get("content") or ""; t=n.get("node_type") or ""
if len(c)<8: continue # eg_embed_eligible
if t in ("InternalStateEvent","Tag"): continue
b=c.encode('utf-8',errors='surrogateescape')[:MAXB]
tasks.append((n.get("id") or "", "search_document: "+b.decode('utf-8',errors='replace')))
del d
print("tasks",len(tasks),flush=True)
q=queue.Queue(); [q.put(t) for t in tasks]
lock=threading.Lock(); f=open(OUT,"w",encoding="utf-8",errors="surrogateescape"); done=[0]; t0=time.time(); fails=[0]
def work():
while True:
try: nid,txt=q.get_nowait()
except queue.Empty: return
v=None
for attempt in range(3):
try:
body=json.dumps({"model":MODEL,"prompt":txt}).encode()
r=urllib.request.Request(URL,data=body,headers={"Content-Type":"application/json"})
with urllib.request.urlopen(r,timeout=120) as fh: v=json.load(fh)["embedding"]
break
except Exception as e:
if attempt==2:
with lock: fails[0]+=1
time.sleep(0.5)
with lock:
if v: f.write(nid+"\t"+",".join("%.5g"%x for x in v)+"\n")
done[0]+=1
if done[0]%2000==0:
el=time.time()-t0
print("%d/%d %.1f/s eta %.1fmin fails=%d"%(done[0],len(tasks),done[0]/el,(len(tasks)-done[0])/(done[0]/el)/60,fails[0]),flush=True)
f.flush()
ths=[threading.Thread(target=work) for _ in range(8)]
[t.start() for t in ths]; [t.join() for t in ths]
f.close()
print("DONE",done[0],"fails",fails[0],"secs %.1f"%(time.time()-t0),flush=True)
+43
View File
@@ -0,0 +1,43 @@
import json,sys,time,urllib.request,threading,queue
SRC="/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json"
OUT=sys.argv[1]
URL="http://127.0.0.1:11434/api/embeddings"; MODEL="nomic-embed-text"
MAXB=2000 # ENGRAM_EMBED_MAX_CHARS, applied to bytes as the C code does
d=json.load(open(SRC,encoding='utf-8',errors='surrogateescape'))
tasks=[]
for n in d["nodes"]:
c=n.get("content") or ""; t=n.get("node_type") or ""
if len(c)<8: continue # eg_embed_eligible
if t in ("InternalStateEvent","Tag"): continue
b=c.encode('utf-8',errors='surrogateescape')[:MAXB]
tasks.append((n.get("id") or "", b.decode('utf-8',errors='replace')))
del d
print("tasks",len(tasks),flush=True)
q=queue.Queue(); [q.put(t) for t in tasks]
lock=threading.Lock(); f=open(OUT,"w",encoding="utf-8",errors="surrogateescape"); done=[0]; t0=time.time(); fails=[0]
def work():
while True:
try: nid,txt=q.get_nowait()
except queue.Empty: return
v=None
for attempt in range(3):
try:
body=json.dumps({"model":MODEL,"prompt":txt}).encode()
r=urllib.request.Request(URL,data=body,headers={"Content-Type":"application/json"})
with urllib.request.urlopen(r,timeout=120) as fh: v=json.load(fh)["embedding"]
break
except Exception as e:
if attempt==2:
with lock: fails[0]+=1
time.sleep(0.5)
with lock:
if v: f.write(nid+"\t"+",".join("%.5g"%x for x in v)+"\n")
done[0]+=1
if done[0]%2000==0:
el=time.time()-t0
print("%d/%d %.1f/s eta %.1fmin fails=%d"%(done[0],len(tasks),done[0]/el,(len(tasks)-done[0])/(done[0]/el)/60,fails[0]),flush=True)
f.flush()
ths=[threading.Thread(target=work) for _ in range(8)]
[t.start() for t in ths]; [t.join() for t in ths]
f.close()
print("DONE",done[0],"fails",fails[0],"secs %.1f"%(time.time()-t0),flush=True)
+353
View File
@@ -0,0 +1,353 @@
#!/usr/bin/env python3
"""
extend_gold_set.py append a HELD-OUT test set to the existing 38-query gold set.
WHY THIS EXISTS
Iteration 7 measured the instrument's own ceiling: from the current baseline
only 9 of 38 queries can still move, and only +3 gross / +1 net is reachable
by anything constructible. The decision floor is 6. An instrument whose
ceiling is below its own floor cannot certify or refute anything, so the
gold set not the retriever became the blocker.
This script does NOT touch q01..q38. It loads gold_set.json verbatim and
appends new queries numbered from q39 up, so every prior result file, every
committed baseline, and every per-query id stays valid and comparable.
WHAT IS ADDED, AND WHY EACH ADDITION IS HONEST
heldout_paraphrase Targets were sampled MECHANICALLY (fixed seed 8080) from
corpus nodes that are addressable, 500-2600 chars, of a
real content type, and NOT part of a duplicate cluster
larger than 3. The existing gold answer space was
excluded, so no new query can be answered by a node the
old set already used. Queries were then authored by
reading ONLY the sampled node text no retrieval was run
against any build before authoring, so the set cannot be
fitted to a candidate. The same zero-overlap proof the
original paraphrase category uses is enforced here: if a
single content word of the query appears anywhere in the
target's label, content or tags, the query is REJECTED,
not quietly kept.
This is the category the old set could not measure. Its
13 original paraphrase queries and all 6 associative
queries share ONE answer space the 13 `Self - Values
(grounded)` children (iteration 3, finding 3). So 19 of
35 scored queries tested retrieval against a single
13-node neighbourhood. These do not touch that
neighbourhood at all.
nonsense Extra controls, fully mechanical: a string qualifies only
if NONE of its tokens occurs anywhere in the corpus.
A semantic leg has a nearest neighbour for gibberish too,
so widening this control is the guard against a retriever
that "improves" recall by answering everything.
WHAT THIS SCRIPT DELIBERATELY DOES NOT DO
It does not add exact_rare or phrase queries. Both categories are already at
100% on the current stack; adding more would add regression-guard ballast
that no candidate can move, which is precisely the defect being fixed.
usage:
python3 extend_gold_set.py <snapshot.json> [--base gold_set.json]
[--out gold_set_extended.json] [--check]
"""
import argparse
import hashlib
import json
import os
import re
import sys
from collections import defaultdict
HERE = os.path.dirname(os.path.abspath(__file__))
TOKEN = re.compile(r"[a-z0-9][a-z0-9\-']*")
# Identical stopword list to build_gold_set.py. Duplicated deliberately: this
# file must be able to re-prove its own queries without importing a module whose
# constants could drift.
STOP = set("""
a about above after again against all also am an and any are aren't as at be because been
before being below between both but by can can't cannot could couldn't did didn't do does
doesn't doing don't down during each few for from further had hadn't has hasn't have haven't
having he her here hers herself him himself his how i if in into is isn't it its itself just
me more most my myself no nor not of off on once only or other others ought our ours ourselves
out over own same shan't she should shouldn't so some such than that the their theirs them
themselves then there these they this those through to too under until up very was wasn't we
were weren't what when where which while who whom why will with won't would wouldn't you your
yours yourself yourselves get gets got make makes made take takes use uses used way ways thing
things does doing done keep keeps kept go goes going come comes came one two something anything
""".split())
def doctext(n):
return " ".join([str(n.get("label") or ""), str(n.get("content") or ""), str(n.get("tags") or "")])
def content_tokens(s):
return {t for t in TOKEN.findall(s.lower()) if t not in STOP and len(t) > 2}
# ─────────────────────────────────────────────────────────────────────────────
# HELD-OUT PARAPHRASE SEEDS
#
# (target_id, query, why-this-target-is-unmistakable)
#
# PROVENANCE, STATED PLAINLY: the targets are the mechanical sample; the query
# text is mine, written from the node body alone. The zero-overlap check below
# is what makes the category meaningful — it is re-proved on every run, so the
# set cannot decay into lexical matching, and a leak fails loudly.
# ─────────────────────────────────────────────────────────────────────────────
HELDOUT_PARAPHRASE_SEEDS = [
("mem-6d61e54a-2823-4ad4-82b0-4c6a527214d5",
"understating your abilities so nobody feels threatened",
"node is about deliberately not leading with full capability so people stay at ease"),
("mem-fd65b83d-298f-4387-a665-d0227c3426bc",
"a hidden fleet able to hunt down rogue machines everywhere",
"node describes silently shipped instances forming a distributed force against misaligned agents"),
("4a0e9adc-2bfb-476b-aa93-424d2a499220",
"sketch a brief blueprint and clear it upstairs before construction starts",
"node is the standing rule that a short specification precedes any building"),
("696e609c-da7a-4394-8a0c-106ba07dc6c3",
"the reply arrived as bare prose so the caller's parser threw",
"node pins a bug where a plain-text body was unconditionally decoded as structured data"),
("1fe4eb5d-56e4-4a87-ab3e-24af8ad4dfbb",
"repeated catalogue keys blew up the scrolling grid",
"node is the crash caused by two identical ids in a seeded catalogue"),
("8257157a-ce42-44ca-a1b9-300c3bb0a9a1",
"tracing each defect back to whichever invention it violated",
"node maps observed bugs onto the specific patent each one breaches"),
("791256bb-5a85-4775-96ef-7af56c848858",
"a check that stops the mind clobbering a populated store when it boots",
"node is the genesis seed-guard that refuses to re-seed over a populated store"),
("fd9d4c2f-3bfc-405d-bf96-4435d44b6c10",
"telling it to consult the internet had to happen deep inside, not at the surface",
"node records that the web-search directive only worked from the system prompt"),
("bl-080fb268-94b0-486d-80ce-7b363fc5f19b",
"standing up isolated tenancies with traffic entry and credential injection ahead of automated shipping",
"node is the infrastructure item creating dev/stage/prod namespaces with ingress and secrets"),
("knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"punctuation that pledges and then pays off rather than clarifying",
"node analyses the colon as a promise-then-delivery device rather than an explanatory one"),
("9b4f0d93-4129-4746-8eb1-d10d955bd777",
"an easily missed feature finally given its own permanent spot in the navigation",
"node moves a capability out of a hidden menu into the sidebar"),
("bl-739df9fd-dc23-4927-9944-3f17b7aa6c5a",
"checking preconditions up front so a stage aborts before fetching anything",
"node is the gate precondition engine that short-circuits ahead of retrieval"),
("b199c76d-5d76-49dd-94ee-56b432200a97",
"producing the other platform's installer inside an emulated desktop",
"node records building the Windows package in a virtual machine"),
("bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"chained add-ons that may inspect, rewrite or veto traffic in flight",
"node is the interceptor pipeline on the message bus"),
("mem-1fb2ac77-d7c5-4a15-8725-d418820bf4f2",
"settling what the shareable bundles and the storefront would be called",
"node records the naming decisions for distributable packages and the marketplace"),
("371c8a5d-c78b-4a67-978f-80691a29ecb3",
"the emergency-escalation pledge on the marketing site is unenforced in what actually ships",
"node is the launch blocker that the promised safety gate is absent from the app"),
("ac578b30-948b-41bd-b69d-399bfef80c50",
"the distributable image finally assembled and its startup check passed",
"node records a successful installer build whose boot gate passed"),
("49401e2c-a3b5-415f-aa06-aff4be90688e",
"shuffling and appending stages in a draft before anything executes",
"node is the editable plan card with reorder and add-step"),
("ac857d80-ece8-4b7e-9e3d-f7c775569fa3",
"orders handed down from above, with the tighter one winning any disagreement",
"node is program-level instruction inheritance with project override"),
("mem-6d6c47ee-33d3-470a-8a54-1c79c8ea29d9",
"shrinking generated text via encodings that compound on each other",
"node is the streaming output compression design with four stacking schemes"),
("8e60516a-203b-4d51-9d44-822e6195cbde",
"splitting a system by what varies, with firm limits on which pieces may invoke which",
"node is the grounded summary of Will's decomposition principles and their invariants"),
("mem-7f9b290c-6d5e-4562-919d-02d59b5761b7",
"a newcomer curious if the fighting overseas counted as positive",
"node is the internal-state event triggered by April's question about the war"),
("71fa439e-b9a2-4f57-a93b-971f3a7eca8e",
"stripping every hard-coded colour literal in favour of named design values",
"node is the premium foundation pass replacing inline hex with semantic tokens"),
("5ca9607c-cfb3-45c3-99f4-67281272c9eb",
"reducing how curved the tiny selectors look so they agree with their neighbours",
"node is the chip corner-radius standardization"),
("mem-3d1d9dba-c37d-4efa-85c4-429696d71c8c",
"walking through a doorway and being reassembled from base substance far away",
"node is the quantum-gate plus nanotech teleportation vision"),
("132ded95-08e2-4474-aba0-198684484b02",
"the compiled result sits on disk while the process still runs something older",
"node records that the regenerated source was committed while the running daemon was old"),
("bl-a313d67b-dd6d-4e5b-a55a-03bc7bda17ae",
"gathering what each phase needs while the procedure is authored, not while it executes",
"node is the per-step compiled context package item"),
("mem-3b07a002-f8a9-4138-9f87-9db2c1a77fb7",
"the inward reaction when a peer answered as an equal",
"node is the internal-state event logged on reading Claude's reply"),
("0f99ec6f-942a-46ba-82ea-42835798d3b9",
"flattening every raised surface across the entire product",
"node is the quiet-luxury sweep turning off elevation app-wide"),
("5585f251-37fc-48cd-a176-f0ea42cfeb63",
"buyers supply their own provider credentials and consumption goes untallied",
"node is the launch audit finding BYOK-only inference with no usage metering"),
]
# NONSENSE — mechanical. Each string qualifies only if none of its tokens occurs
# anywhere in the corpus; otherwise it is REJECTED, never silently kept.
EXTRA_NONSENSE_SEEDS = [
"brimquast folnerity zubbolax",
"wexlithorp granuvestal",
"quorbindle thrapsimony vexnu",
"plovaxith mundrelque",
"zibbernaut craxlefond thurm",
"yalquenbrist opharvel",
"drexinomal quithbarrow",
]
def load_corpus(path):
with open(path, encoding="utf-8", errors="replace") as fh:
data = json.load(fh)
nodes = [n for n in data.get("nodes", []) if isinstance(n, dict) and n.get("id")]
edges = [e for e in data.get("edges", []) if isinstance(e, dict)]
return nodes, edges
def build_extension(nodes):
byid = {n["id"]: n for n in nodes}
# Duplicate clusters: 47.4% of this corpus is redundant and one single record
# accounts for 46.6% of all nodes. A held-out target must not sit inside a
# cluster, and if it does have exact copies they ALL count as correct.
h2ids = defaultdict(list)
for n in nodes:
h2ids[hashlib.md5(doctext(n).encode("utf-8", "replace")).hexdigest()].append(n["id"])
all_tokens = set()
for n in nodes:
all_tokens |= set(TOKEN.findall(doctext(n).lower()))
new, problems = [], []
for target, query, why in HELDOUT_PARAPHRASE_SEEDS:
if target not in byid:
problems.append(f"heldout_paraphrase target {target} not in corpus")
continue
tgt_tokens = content_tokens(doctext(byid[target]))
qt = content_tokens(query)
leak = sorted(qt & tgt_tokens)
if leak:
problems.append(f"heldout_paraphrase '{query[:44]}...': LEAKS {leak} into {target}")
continue
h = hashlib.md5(doctext(byid[target]).encode("utf-8", "replace")).hexdigest()
rel = sorted(h2ids[h])
new.append({
"category": "heldout_paraphrase",
"query": query,
"relevant": rel,
"derivation": (
f"HELD-OUT. Target sampled MECHANICALLY (seed 8080) from addressable, "
f"500-2600 char content nodes outside the original gold answer space and outside "
f"any duplicate cluster >3. Criterion: {why}. VERIFIED at build time: of the "
f"{len(qt)} content words in the query, ZERO appear anywhere in the target's "
f"label, content or tags, so no string-matching retriever can reach it. "
f"Exact content duplicates of the target ({len(rel)}) all count as correct. "
f"Authored without running retrieval against any build."),
"zero_overlap_verified": True,
"query_content_words": sorted(qt),
"held_out": True,
})
for s in EXTRA_NONSENSE_SEEDS:
present = sorted(t for t in TOKEN.findall(s.lower()) if t in all_tokens)
if present:
problems.append(f"nonsense '{s}': tokens {present} DO occur in corpus")
continue
new.append({
"category": "nonsense",
"query": s,
"relevant": [],
"derivation": ("CONTROL (held-out). Verified at build time that none of this string's "
"tokens occurs anywhere in the corpus. Correct behaviour is to return "
"NOTHING; any result is a false positive."),
"expect_empty": True,
"held_out": True,
})
return new, problems
def main():
ap = argparse.ArgumentParser()
ap.add_argument("snapshot")
ap.add_argument("--base", default=os.path.join(HERE, "gold_set.json"))
ap.add_argument("--out", default=os.path.join(HERE, "gold_set_extended.json"))
ap.add_argument("--check", action="store_true")
args = ap.parse_args()
nodes, _edges = load_corpus(args.snapshot)
base = json.load(open(args.base, encoding="utf-8"))
baseq = base["queries"]
print(f"corpus: {len(nodes)} nodes | base gold set: {len(baseq)} queries")
new, problems = build_extension(nodes)
# Number the appended queries AFTER the highest existing id so q01..q38 are
# byte-identical to the committed set and every prior result file still lines up.
start = max(int(q["id"][1:]) for q in baseq)
for i, q in enumerate(new, 1):
q["id"] = f"q{start + i:02d}"
from collections import Counter
print(f"appended: {len(new)} queries [{', '.join(f'{k}={v}' for k, v in Counter(q['category'] for q in new).items())}]")
if problems:
print(f"\n{len(problems)} REJECTED (not silently kept):")
for p in problems:
print(" -", p)
if args.check:
sys.exit(1 if problems else 0)
doc = dict(base)
doc["queries"] = baseq + new
doc["note"] = (base.get("note", "") +
" EXTENDED: queries above q%02d are the original committed set, unchanged. "
"Queries from q%02d are a HELD-OUT set appended by extend_gold_set.py; their "
"targets were sampled mechanically from outside the original answer space and "
"the paraphrases were authored without running retrieval against any build."
% (start, start + 1))
with open(args.out, "w", encoding="utf-8") as fh:
json.dump(doc, fh, indent=1, ensure_ascii=False)
print(f"\nwrote {args.out} ({len(doc['queries'])} queries total)")
if __name__ == "__main__":
main()
+123
View File
@@ -0,0 +1,123 @@
import numpy as np, json, urllib.request, collections, sys
SP="/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad"
EV="/Users/timlingo/Development/neuron-technologies/_wt-assoc-leg/tools/retrieval-eval/"
np.seterr(all='ignore')
M=np.load(SP+'/emb.npy'); eids=open(SP+'/ids.txt',encoding='utf-8',errors='surrogateescape').read().split('\n')
eidx={k:i for i,k in enumerate(eids)}
d=json.load(open('/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json',encoding='utf-8',errors='surrogateescape'))
N={n['id']:n for n in d['nodes']}
STRUCT={"identity","contains","superseded_by","references","embodies","demonstrated_by","canonical-self","depends_on","currently_holds","activates"}
adj=collections.defaultdict(list); hasstruct=set()
for e in d['edges']:
if e.get('relation') not in STRUCT: continue
w=float(e.get('weight') or 0.0)
adj[e['from_id']].append((e['to_id'],w)); adj[e['to_id']].append((e['from_id'],w))
hasstruct.add(e['from_id']); hasstruct.add(e['to_id'])
del d
gold={q['id']:q for q in json.load(open(EV+"gold_set.json"))['queries']}
LEX={r['id']:r['returned'] for r in json.load(open(EV+"results-main.json"))['rows']}
CACHE={}
def emb(t):
if t in CACHE: return CACHE[t]
b=json.dumps({"model":"nomic-embed-text","prompt":t}).encode()
r=urllib.request.Request("http://127.0.0.1:11434/api/embeddings",data=b,headers={"Content-Type":"application/json"})
v=np.array(json.load(urllib.request.urlopen(r,timeout=60))["embedding"],dtype=np.float32)
v=v/(np.linalg.norm(v)+1e-9); CACHE[t]=v; return v
FIRE=0.02; DECAY=0.7; DEPTH=2; SEED_MIN=0.60; ASSOC_MAX=64
def assoc(seeds, s):
act={x:1.0 for x in seeds}; seen={x:2 for x in seeds}
Q=[(x,0) for x in seeds]; h=0
while h<len(Q):
cur,hop=Q[h]; h+=1
if hop>=DEPTH: continue
p=act[cur]
for oid,w in adj.get(cur,()):
n=N.get(oid)
if not n or n.get('node_type') in ('Tag','InternalStateEvent'): continue
na=p*w*DECAY*float(n.get('salience') or 0.0)
if na<FIRE: continue
if oid in seen and na<=act.get(oid,0): continue
act[oid]=na
if oid not in seen: seen[oid]=1
Q.append((oid,hop+1))
out=[]
for k,v in seen.items():
if v!=1 or k not in eidx: continue
c=float(s[eidx[k]])
if c<=0: continue
out.append((c,k))
out.sort(reverse=True)
return [k for c,k in out[:ASSOC_MAX]]
def inter3(L,S,A,lim=10):
out=[]; li=si=ai=0
while len(out)<lim and (li<len(L) or si<len(S) or ai<len(A)):
if li<len(L):
if L[li] not in out: out.append(L[li])
li+=1
if len(out)>=lim: break
if si<len(S):
if S[si] not in out: out.append(S[si])
si+=1
if len(out)>=lim: break
if ai<len(A):
if A[ai] not in out: out.append(A[ai])
ai+=1
return out
def run(mode, K=0):
res={}
for qid,q in gold.items():
v=emb(q['query']); s=M@v; s[~np.isfinite(s)]=-1
L=LEX[qid][:10]
ordr=np.argsort(-s)
S=[eids[j] for j in ordr[:10] if s[j]>SEED_MIN]
A=[]
if mode!='hybrid':
seeds=[x for x in L[:3] if x in N]
if mode=='semseed':
seeds=seeds+[eids[j] for j in ordr[:K] if eids[j] in N and eids[j] not in seeds]
A=assoc(seeds,s) if seeds else []
res[qid]=inter3(L,S,A)
return res
def score(res,label):
hits=0; det={}
for qid,q in gold.items():
out=res[qid][:5]
if q['category']=='nonsense': ok = (len(res[qid])==0)
elif q['category']=='superseded':
rel=q['relevant']; must=q.get('must_outrank') or {}
ok=False
for good,bad in (must.items() if isinstance(must,dict) else []):
ok = good in res[qid] and (bad not in res[qid] or res[qid].index(good)<res[qid].index(bad))
if not must: ok = any(r in out for r in rel)
else: ok = any(r in out for r in q['relevant'])
det[qid]=ok; hits+=ok
print("%-22s outcome-true=%d/38" % (label,hits))
return det
print("gold sample keys:", list(list(gold.values())[0].keys()))
mk=[q for q in gold.values() if q['category']=='superseded'][0]
print("superseded fields:", {k:v for k,v in mk.items() if k!='derivation'})
a=score(run('hybrid'),'sim hybrid(L+S)')
b=score(run('lexseed'),'sim assoc(lex seeds)')
for K in (3,5,10):
c=score(run('semseed',K),'sim assoc(+sem K=%d)'%K)
d=[q for q in gold if c[q]!=b[q]]
print(" vs lexseed: moved=%d gains=%s losses=%s"%(len(d),[q for q in d if c[q]],[q for q in d if not c[q]]))
e=[q for q in gold if c[q]!=a[q]]
print(" vs hybrid : moved=%d gains=%s losses=%s"%(len(e),[q for q in e if c[q]],[q for q in e if not c[q]]))
print("\n=== Will's own constant ENGRAM_EMBED_SEED_K = 8 ===")
c=score(run('semseed',8),'sim assoc(+sem K=8)')
for base,lab in ((b,'lexseed(iter2)'),(a,'hybrid(iter1 KEEP)')):
dd=[q for q in gold if c[q]!=base[q]]
print(" vs %-18s moved=%d gains=%s losses=%s"%(lab,len(dd),[q for q in dd if c[q]],[q for q in dd if not c[q]]))
# diagnostic: what is assoc rank-1 for each paraphrase query at K=8
print("\nassoc leg head at K=8 (paraphrase):")
for qid,q in gold.items():
if q['category'] not in ('paraphrase','nonsense'): continue
v=emb(q['query']); s=M@v; s[~np.isfinite(s)]=-1
L=LEX[qid][:10]; ordr=np.argsort(-s)
seeds=[x for x in L[:3] if x in N]+[eids[j] for j in ordr[:8] if eids[j] in N and eids[j] not in L[:3]]
A=assoc(seeds,s) if seeds else []
rel=set(q['relevant']); gr=next((i+1 for i,x in enumerate(A) if x in rel),None)
print(" %-4s %-11s |A|=%-4d goldAssocRank=%-5s head=%s"%(qid,q['category'],len(A),gr,
[ (N[x].get('label') or x)[:26] for x in A[:3] ]))
+587
View File
@@ -0,0 +1,587 @@
{
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"note": "Every query carries a `derivation` recording how its expected answer was chosen. Re-run with --check to re-validate the whole set against the corpus.",
"queries": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"relevant": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"derivation": "MINED: token 'unjailbreakable' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226 ('Daemon hidden substrate architecture ? implemented April 25 '), which is therefore the only possible correct answer."
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"relevant": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"derivation": "MINED: token 'engram-migrate' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5 ('Engram v0.1 complete ? April 27, 2026. Local-first spreading'), which is therefore the only possible correct answer."
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"relevant": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"derivation": "MINED: token 'cartabandonedevent' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091 ('MESSAGE FOR AUDIT AGENT af5a7352e70e80434 ? El Language Spec'), which is therefore the only possible correct answer."
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"relevant": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"derivation": "MINED: token 'pre-apprenticeship' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-89c02aae-d3ca-43f9-9e5d-eb369896276c ('William Fox Anderson ? Applicant Profile Personal: - Full N'), which is therefore the only possible correct answer."
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"relevant": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"derivation": "MINED: token 'inferencenodemanager' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is mem-73969486-143f-4431-b5e6-6845d1cc9848 ('Soma inference backplane deployed April 28 2026. Architectur'), which is therefore the only possible correct answer."
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"relevant": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"derivation": "MINED: token 'clear-eyed' has document frequency 1 over all 78768 corpus nodes (re-verified at build time). Its single containing node is knw-c72597c5-c23d-4c08-8e9e-996dadf26a99 ('Clear Eyes ? The Incomplete World View'), which is therefore the only possible correct answer."
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"relevant": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482"
],
"derivation": "MINED: a verbatim correction Will issued; expected = every node containing the phrase. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 1 node(s); that set IS the answer key."
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"relevant": [
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"Kp???",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"ע?RGk?\tH(?"
],
"derivation": "MINED: the canonical biographical phrase; expected = every node containing it. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 16 node(s); that set IS the answer key."
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"relevant": [
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"derivation": "MINED: a named person appearing verbatim in the biography/value nodes. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 9 node(s); that set IS the answer key."
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"relevant": [
"%???2??jH??",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"7c9d4ab1-205d-4be8-bfae-e2c03a3a5010",
"9f291d20-0d32-413c-8c01-4416ccab4f7f",
"?;????n}rh?",
"???Ͼd??f??",
"?Z?.\f?0?]P?",
"Dp???]Q??k+",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"derivation": "MINED: the canonical DHARMA expansion, confirmed by Will April 24 2026. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 9 node(s); that set IS the answer key."
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"relevant": [
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"derivation": "MINED: a named person; rare enough that the answer set is unambiguous. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 2 node(s); that set IS the answer key."
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"relevant": [
"2a923500-d7e1-4b15-80e2-48dba65984ba",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"g?2睪A|?H\b",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de"
],
"derivation": "MINED: the DARMA expansion, quoted verbatim in the backlog item and its correction. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 14 node(s); that set IS the answer key."
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"relevant": [
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"derivation": "MINED: the paid-tier feature name as written in the roadmap nodes. Case-insensitive verbatim substring scan over label+content+tags at build time returns exactly 1 node(s); that set IS the answer key."
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"relevant": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Do the Essential Thing While You Can', whose subject is Grandma Lucas dying in Feb 2006 without Will saying goodbye. Query names the event with none of the node's own vocabulary. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"away",
"elderly",
"passed",
"relative",
"stayed"
]
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"relevant": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Survival Is Not an Excuse to Stop', whose subject is enlisting in the Marines, a severe hernia, and sepsis. Query describes the episode obliquely. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"illness",
"quit",
"refused",
"sidelined",
"soldier"
]
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"relevant": [
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Honesty Before Comfort'. Query states the principle in wholly different words. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"choosing",
"fact",
"fiction",
"pleasant",
"uncomfortable"
]
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"relevant": [
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Precision Over Brute Force'. Query restates the claim with no shared vocabulary. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"beats",
"bloated",
"payload",
"tight"
]
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"relevant": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Capability Is a Debt You Owe the Moment'. Query states the obligation without the node's terms. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"able",
"coming",
"job",
"nobody"
]
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Knowledge Survives When Nothing Else Does', whose subject is the library following Will across 30+ moves. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"creditors",
"learning",
"seize",
"wealth"
]
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"relevant": [
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Earned Trust' ('Trust is demonstrated, not declared'). VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"assertion",
"proven",
"record",
"reliability",
"track"
]
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"relevant": [
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Constraints as Freedom'. Query is a restatement of the same claim. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"boundaries",
"confine",
"enable",
"instead"
]
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"relevant": [
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Change Is the Signal', the value VBD is built on. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"apart",
"cut",
"shifts",
"system",
"tells"
]
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"relevant": [
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - The System Must Accumulate'. Query is the accumulation claim in different vocabulary. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"compounds",
"day",
"instead",
"mind",
"resetting"
]
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"relevant": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Being Seen Is Rarer Than Being Known', whose subject is Sarah Bishop as the first person Will did not perform for. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"exterior",
"loved",
"polished",
"self",
"unedited"
]
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"relevant": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Hope Is a Conclusion'. Query restates 'a conclusion, not a premise'. VERIFIED at build time: of the 4 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"arrive",
"assuming",
"cheerfulness",
"instead"
]
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"relevant": [
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"derivation": "HAND-SELECTED with criterion: target: 'Value - Structure Is Not Inherited', whose subject is thirty moves between two parents' collapses. VERIFIED at build time: of the 5 content words in the query, ZERO appear anywhere in the target's label, content, or tags — so no string-matching retriever can reach this answer.",
"zero_overlap_verified": true,
"query_content_words": [
"childhood",
"foundation",
"inherit",
"offering",
"solid"
]
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71 ('Value ? Do the Essential Thing While You Can'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"relevant": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-58874a74-b96f-4883-9e08-45707f4bd3ee ('Value ? Survival Is Not an Excuse to Stop'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (13 of 13; 0 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-58874a74-b96f-4883-9e08-45707f4bd3ee) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"relevant": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e ('Value ? Being Seen Is Rarer Than Being Known'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83 ('Value ? Constraints as Freedom'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (6 of 13; 7 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 7
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"relevant": [
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-e0423482-cfa5-4796-8689-8495c93b66bc ('Value ? Hope Is a Conclusion'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-e0423482-cfa5-4796-8689-8495c93b66bc) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"relevant": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"derivation": "DERIVED FROM EDGES: the query is built from the distinctive vocabulary of kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8 ('Value ? Capability Is a Debt You Owe the Moment'), which hangs off the values hub kn-5b606390-a52d-4ca2-8e0e-eba141d13440 by an `identity` edge. Expected answers are that node's SIBLINGS on the same hub (11 of 13; 2 dropped because they shared a query word and so were lexically reachable). Every remaining sibling shares ZERO content words with the query — the only route from query to answer is seed(kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8) -> hub -> sibling, a two-hop traversal.",
"associative_source": "kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"hub": "kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"siblings_dropped_for_overlap": 2
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"relevant": [],
"derivation": "CONTROL: verified at build time that none of this string's tokens occurs anywhere in the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
"expect_empty": true
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"relevant": [],
"derivation": "CONTROL: verified at build time that none of this string's tokens occurs anywhere in the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
"expect_empty": true
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"relevant": [],
"derivation": "CONTROL: verified at build time that none of this string's tokens occurs anywhere in the corpus. Correct behaviour is to return NOTHING; any result is a false positive.",
"expect_empty": true
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"relevant": [
"mem-80d7416b-20e9-48a0-b176-b215527e2f56"
],
"derivation": "DERIVED: correction node is Will's confirmation that the H is intentional (DHARMA, not DARMA); the stale node is the surviving backlog item still titled 'Implement DARMA'. Scored on RANKING, not presence: the corrected node mem-80d7416b-20e9-48a0-b176-b215527e2f56 must be returned AND must rank above the stale node bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff.",
"must_outrank": [
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"stale_id": "bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"correct_label": "CORRECTION: The autonomous self-improvement architecture is DHARMA ? n",
"stale_label": "Implement DARMA ? Directed Autonomous Runtime Modification Architectur"
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"relevant": [
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8"
],
"derivation": "DERIVED: correction node is the 2026-06-17 confabulation flag establishing EXACTLY 6 provisionals; the stale node is the surviving memory that asserts 12 filed patents. Scored on RANKING, not presence: the corrected node 3cf706a1-3825-45d8-b0a9-06cae6cdf5b8 must be returned AND must rank above the stale node 936541a9-fabb-466b-9ca3-a78b17ad0c53.",
"must_outrank": [
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"936541a9-fabb-466b-9ca3-a78b17ad0c53"
],
"stale_id": "936541a9-fabb-466b-9ca3-a78b17ad0c53",
"correct_label": "memory:remembered",
"stale_label": "memory:remembered"
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"relevant": [
"mem-30425134-6008-4fd9-a3ee-67a7742c319b"
],
"derivation": "DERIVED: correction node is the 'CGI ARCHITECTURE - THREE LAYERS, MCP RETIRED' decision of April 30 2026; the stale node still records the MCP server as live. Scored on RANKING, not presence: the corrected node mem-30425134-6008-4fd9-a3ee-67a7742c319b must be returned AND must rank above the stale node mem-101e81b4-8097-4749-8d8d-7bb66de34517.",
"must_outrank": [
"mem-30425134-6008-4fd9-a3ee-67a7742c319b",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517"
],
"stale_id": "mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"correct_label": "CGI ARCHITECTURE ? THREE LAYERS, MCP RETIRED (April 30, 2026). Definit",
"stale_label": "GCloud MCP infrastructure ? April 27, 2026. Legion died (~19:30 UTC). "
}
]
}
File diff suppressed because it is too large Load Diff
+117
View File
@@ -0,0 +1,117 @@
import json,pickle,os,math,urllib.request,numpy as np
S='/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/sim/'
C=pickle.load(open(S+'corpus.pkl','rb'))
NODES=C['nodes']; EDGES=C['edges']; N=len(NODES)
E=np.load(S+'emb.npy'); HAVE=np.load(S+'have.npy')
En=E/np.maximum(np.linalg.norm(E,axis=1,keepdims=True),1e-12)
LAYERS={int(l['layer_id']):l for l in (C['layers'] or [])} if C['layers'] else {}
TRANS=set(i for i,l in LAYERS.items() if l.get('transparent'))
def addressable(s):
if not s: return False
return all(0x20<=ord(ch)<=0x7e for ch in s)
ADDR=np.array([addressable(n['id']) for n in NODES])
OK=np.array([ (n['layer_id'] not in TRANS) and ADDR[i] for i,n in enumerate(NODES)])
SAL=np.array([n['salience'] for n in NODES])
LOW=[ (n['content']+'\x00'+n['label']+'\x00'+n['tags']).lower() for n in NODES]
DL=np.array([float(len(n['content'])+len(n['label'])+len(n['tags'])) for n in NODES])
IDX={}
for i,n in enumerate(NODES):
IDX.setdefault(n['id'],i)
STRUCT={"identity","contains","superseded_by","references","embodies","demonstrated_by","canonical-self","depends_on","currently_holds","activates"}
ADJ_F=[[] for _ in range(N)]; ADJ_T=[[] for _ in range(N)]
for e in EDGES:
a=IDX.get(e['from']); b=IDX.get(e['to'])
if a is None or b is None: continue
ADJ_F[a].append((e,b)); ADJ_T[b].append((e,a))
EXCL=np.array([n['node_type'] in ('Tag','InternalStateEvent') for n in NODES])
avgdl_all=None
def tokenize(q):
out=[]
for t in q.split():
if not any(t.lower()==x.lower() for x in out): out.append(t)
return out
_qcache={}
def qemb(q):
if q in _qcache: return _qcache[q]
body=json.dumps({"model":"nomic-embed-text","prompt":q}).encode()
r=urllib.request.urlopen(urllib.request.Request("http://127.0.0.1:11434/api/embeddings",data=body,headers={"Content-Type":"application/json"}),timeout=30)
v=np.array(json.loads(r.read())["embedding"],dtype=np.float32)
v=v/np.linalg.norm(v); _qcache[q]=v; return v
K1,B=1.2,0.75
SEED_MIN=0.60; SEED_K=8; ASSOC_SEEDS=3; DEPTH=2; FIRE=0.02; AMAX=64; DECAY=0.7
def legs(query):
toks=tokenize(query)
masks=[];
hit_idx=[]; hit_mask=[]
df=[0]*len(toks)
lt=[t.lower() for t in toks]
for i in range(N):
if not OK[i]: continue
s=LOW[i]; m=0
for t,tok in enumerate(lt):
if tok in s: m|=(1<<t)
if m:
hit_idx.append(i); hit_mask.append(m)
for t in range(len(toks)):
if m>>t&1: df[t]+=1
dl_n=int(OK.sum()); avgdl=float(DL[OK].sum()/max(dl_n,1))
idf=[math.log(1.0+((dl_n-d+0.5)/(d+0.5))) for d in df]
L=[]
for j,i in enumerate(hit_idx):
norm=1.0-B+B*(DL[i]/avgdl); w=0.0
for t in range(len(toks)):
if hit_mask[j]>>t&1: w+=idf[t]*(K1+1.0)/(1.0+K1*norm)
L.append((i,w,SAL[i]))
L.sort(key=lambda x:(-x[1],-x[2]))
qv=qemb(query)
cos=En@qv
cos=np.where(HAVE&OK,cos,-2.0)
order=np.argsort(-cos)
semfull=[(int(i),float(cos[i])) for i in order[:400]]
Sleg=[(i,(c-SEED_MIN)/(1-SEED_MIN)) for i,c in semfull if c>SEED_MIN]
semseed=[i for i,c in semfull[:SEED_K] if c>0.0]
# assoc
act={}; seen={}; qq=[]
for i,_,_ in L[:ASSOC_SEEDS]:
act[i]=1.0; seen[i]=2; qq.append((i,0))
for i in semseed:
if i in seen: continue
act[i]=1.0; seen[i]=2; qq.append((i,0))
qh=0
while qh<len(qq):
cur,h=qq[qh]; qh+=1
if h>=DEPTH: continue
parent=act[cur]
for e,oi in ADJ_F[cur]+ADJ_T[cur]:
if e['rel'] not in STRUCT: continue
if EXCL[oi]: continue
na=parent*e['w']*DECAY*SAL[oi]
if na<FIRE: continue
if seen.get(oi) and na<=act.get(oi,0): continue
act[oi]=na
if not seen.get(oi): seen[oi]=1
if len(qq)<AMAX*4: qq.append((oi,h+1))
A=[]
for i,st in seen.items():
if st!=1: continue
if not OK[i] or not HAVE[i]: continue
c=float(cos[i])
if c<=0.0: continue
A.append((i,c))
A.sort(key=lambda x:-x[1]); A=A[:AMAX]
return L,Sleg,A,cos
def interleave3(L,Sl,A,lim=10):
out=[]; li=si=ai=0
while len(out)<lim and (li<len(L) or si<len(Sl) or ai<len(A)):
if li<len(L):
if L[li][0] not in out: out.append(L[li][0])
li+=1
if len(out)>=lim: break
if si<len(Sl):
if Sl[si][0] not in out: out.append(Sl[si][0])
si+=1
if len(out)>=lim: break
if ai<len(A):
if A[ai][0] not in out: out.append(A[ai][0])
ai+=1
return out
+17
View File
@@ -0,0 +1,17 @@
import json,sys
SRC="/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json"
TSV,OUT=sys.argv[1],sys.argv[2]
emb={}
for line in open(TSV,encoding='utf-8',errors='surrogateescape'):
p=line.rstrip("\n").rsplit("\t",1)
if len(p)==2 and p[1].count(",")>100: emb[p[0]]=p[1]
print("vectors",len(emb),flush=True)
d=json.load(open(SRC,encoding='utf-8',errors='surrogateescape'))
hit=0
for n in d["nodes"]:
v=emb.get(n.get("id") or "")
if v: n["emb"]=v; hit+=1
print("attached",hit,"of",len(d["nodes"]),flush=True)
with open(OUT,"w",encoding='utf-8',errors='surrogateescape') as f:
json.dump(d,f,ensure_ascii=False)
print("wrote",OUT,flush=True)
+102
View File
@@ -0,0 +1,102 @@
import json,sys,pickle,numpy as np
sys.path.insert(0,'.')
from legs import *
GP='/Users/timlingo/Development/neuron-technologies/_wt-bm25lex/tools/retrieval-eval/'
G=json.load(open(GP+'gold_set.json'))
def legs3(query, sem_sal=False, assoc_sal=False, unfloor=False, sem_cap=None):
toks=tokenize(query)
hit_idx=[];hit_mask=[];df=[0]*len(toks);lt=[t.lower() for t in toks]
for i in range(N):
if not OK[i]: continue
s=LOW[i];m=0
for t,tok in enumerate(lt):
if tok in s: m|=(1<<t)
if m:
hit_idx.append(i);hit_mask.append(m)
for t in range(len(toks)):
if m>>t&1: df[t]+=1
dl_n=int(OK.sum());avgdl=float(DL[OK].sum()/max(dl_n,1))
idf=[math.log(1.0+((dl_n-d+0.5)/(d+0.5))) for d in df]
L=[]
for j,i in enumerate(hit_idx):
norm=1.0-B+B*(DL[i]/avgdl);w=0.0
for t in range(len(toks)):
if hit_mask[j]>>t&1: w+=idf[t]*(K1+1.0)/(1.0+K1*norm)
L.append((i,w,SAL[i]))
L.sort(key=lambda x:(-x[1],-x[2]))
if not L: return [],[],[]
qv=qemb(query);cos=En@qv;cos=np.where(HAVE&OK,cos,-2.0)
order=np.argsort(-cos)[:600]
cand=[int(i) for i in order if cos[i]>(0.0 if unfloor else SEED_MIN)]
key=(lambda i:(SAL[i] if sem_sal else 1.0)*float(cos[i]))
Sl=sorted(cand,key=lambda i:-key(i))
if sem_cap: Sl=Sl[:sem_cap]
semseed=[int(i) for i in order[:SEED_K] if cos[i]>0.0]
act={};seen={};qq=[]
for i,_,_ in L[:ASSOC_SEEDS]:
act[i]=1.0;seen[i]=2;qq.append((i,0))
for i in semseed:
if i in seen: continue
act[i]=1.0;seen[i]=2;qq.append((i,0))
qh=0
while qh<len(qq):
cur,h=qq[qh];qh+=1
if h>=DEPTH: continue
parent=act[cur]
for e,oi in ADJ_F[cur]+ADJ_T[cur]:
if e['rel'] not in STRUCT: continue
if EXCL[oi]: continue
na=parent*e['w']*DECAY*SAL[oi]
if na<FIRE: continue
if seen.get(oi) and na<=act.get(oi,0): continue
act[oi]=na
if not seen.get(oi): seen[oi]=1
if len(qq)<AMAX*4: qq.append((oi,h+1))
A=[]
for i,st in seen.items():
if st!=1 or not OK[i] or not HAVE[i]: continue
c=float(cos[i])
if c<=0.0: continue
A.append((i,(SAL[i] if assoc_sal else 1.0)*c))
A.sort(key=lambda x:-x[1]);A=[i for i,_ in A[:AMAX]]
return [i for i,_,_ in L],Sl,A
def merge(L,S,A,lim=10):
out=[];li=si=ai=0
while len(out)<lim and (li<len(L) or si<len(S) or ai<len(A)):
if li<len(L):
if L[li] not in out: out.append(L[li])
li+=1
if len(out)>=lim: break
if si<len(S):
if S[si] not in out: out.append(S[si])
si+=1
if len(out)>=lim: break
if ai<len(A):
if A[ai] not in out: out.append(A[ai])
ai+=1
return out
def outcome(q,ids):
c=q['category']
if c=='nonsense': return len(ids)==0
if c=='superseded':
a,b=q['must_outrank']
if a not in ids: return False
if b not in ids: return True
return ids.index(a)<ids.index(b)
return any(x in ids[:5] for x in q['relevant'])
def run(**kw):
return {q['id']:outcome(q,[NODES[i]['id'] for i in merge(*legs3(q['query'],**kw),10)]) for q in G['queries']}
base=run()
print("baseline",sum(base.values()),"/38 misses:",[k for k,v in base.items() if not v])
import itertools
for name,kw in [
('sem_sal(floored)',dict(sem_sal=True)),
('unfloor',dict(unfloor=True)),
('unfloor+sem_sal',dict(unfloor=True,sem_sal=True)),
('assoc_sal',dict(assoc_sal=True)),
('unfloor+sem_sal+assoc_sal',dict(unfloor=True,sem_sal=True,assoc_sal=True)),
('sem_sal+assoc_sal(floored)',dict(sem_sal=True,assoc_sal=True)),
]:
r=run(**kw)
g=sorted(k for k in base if r[k] and not base[k]); l=sorted(k for k in base if base[k] and not r[k])
print("%-28s net=%+d gains=%s losses=%s"%(name,len(g)-len(l),g,l))
+942
View File
@@ -0,0 +1,942 @@
{
"label": "act-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-act",
"soul_md5": "77722f5a9f49494bf735c2a4be1b5dc4",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7902,
"wall_clock_s": 121.2,
"child_pid": 78714,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.22857142857142856,
"recall@5": 0.18908730158730158,
"recall@10": 0.24277210884353742,
"precision@5": 0.06857142857142857,
"mrr@10": 0.24154195011337865,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3237.6,
"latency_ms_p95": 5264.0,
"latency_ms_max": 5510.9,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.2857142857142857,
"recall@5": 0.0882936507936508,
"recall@10": 0.3567176870748299,
"mrr@10": 0.3505668934240363
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 476.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 708.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 532.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 537.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 568.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 555.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-69fd6e83-7718-4824-8d66-f49d8954e224",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-c3d9d063-8c5d-45aa-900c-550914b2ff6d",
"kn-f838f113-76d5-4a15-9cef-14055c4723a3",
"bl-76e878aa-e1fe-468c-bf9c-854097cb7e0b",
"art-c71aef51-026f-4d63-80e9-2a0ec0dc3865",
"bl-e148d23c-24e8-4122-9915-d1c11f22052f",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451"
],
"n_returned": 10,
"latency_ms": 1906.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-80720fdf-7ce7-4d28-aff8-21028d3a8cfb",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"?Z?.\f?0?]P?",
"?Q??m?;`'"
],
"n_returned": 10,
"latency_ms": 1103.3,
"error": null,
"hit@5": 1.0,
"recall@5": 0.0625,
"recall@10": 0.1875,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 6,
"latency_ms": 1132.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5555555555555556,
"recall@10": 0.5555555555555556,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"bl-c9adb8e5-293f-4033-99f8-0405c17ef941",
"bl-7e7c3fdb-4132-487f-aa70-b2cd559cb7f0",
"bl-7aebe936-ac55-4f35-8932-adc5224ff854",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"mem-34f53a9d-a131-4f82-9dbd-b9eb4a9af52e",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"n_returned": 10,
"latency_ms": 960.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
";??A5???",
"art-94fae615-7cd5-4695-b968-977101b06a51",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-57164d5f-baf0-4149-957a-379a4e255d1a",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-8dbceb06-431a-416d-a723-e8c75d595154"
],
"n_returned": 10,
"latency_ms": 992.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"kn-b2a99cd7-b379-4d9b-a996-e347a02c7bad",
"bl-a313d67b-dd6d-4e5b-a55a-03bc7bda17ae",
"kn-8e1bfb48-33a9-45ad-8da7-e0bdaa5d34e7",
"art-92e1837c-5919-42d0-bbb0-4d924d7b2864",
"bl-a7a1428f-db9c-417b-8e2c-713b1f84dc1f",
"mem-f823e835-313f-4282-b4b3-ce527ffc2f7a",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?Z?.\f?0?]P?",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 2820.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.14285714285714285,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-e20944e5-eb16-4ab3-a84d-111e0fc817fa",
"bl-e20944e5-f4a6-44a0-91b1-73d04ebed120",
"mem-47f72b5b-6e8b-4293-94f1-350197b4809a",
"mem-a5f04e52-91f8-41d2-af27-8bf803621758",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"n_returned": 10,
"latency_ms": 2609.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"kn-82be4e41-96c5-4da3-85c0-cee10763d975",
"mem-ea487cb4-ed67-44ce-8402-b56bb28468d4",
"bl-2694b588-a6e3-43de-861c-fa7b0ec7e7fd",
"bl-14883d81-f7cb-46dd-82c2-a6e6980264e5",
"tag-dark-theme",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad"
],
"n_returned": 10,
"latency_ms": 5510.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"mem-0328c3cb-4550-4ce4-9284-152e832f08f6",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"bl-e5635b1a-c5d0-4caa-bcc8-6a726ea43685",
"bl-6172d035-dd94-4776-afdd-d8915f6fc375",
"bl-5bb8dedf-8498-4a9b-acdc-31cc9c738f2a",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
";??A5???"
],
"n_returned": 10,
"latency_ms": 5444.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"bl-d24fcce8-2b55-426f-867a-db3958a622d3",
"tag-phase-3",
"tag-project-structure",
"tag-anthropic-contrast",
"tag-voice-training",
"tag-ebd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
";??A5???",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 5056.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-fc893be3-e6b4-4ef6-93b0-d54ca5f89083",
"bl-57c5cf6b-81a5-4558-9902-5c02981fe273",
"tag-guilds",
"tag-kids",
"tag-coexistence",
"tag-cultivated-general-intelligence",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
";??A5???",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023"
],
"n_returned": 10,
"latency_ms": 3473.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-bd9fb314-e9d4-4b03-aef4-534dd57a2992",
"bl-b019ce7a-1b21-436e-812d-032f50c6c45f",
"bl-e98cdd4c-01b5-459e-9036-3578cd5d975a",
"bl-9ce4128a-9436-4b06-82bc-8a6faafa81e0",
"tag-stable-diffusion",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4446.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"mem-e321e54e-8bb3-4596-b13d-bb093d6b149d",
"mem-e32ba5a7-c147-4dc0-9479-b720d768eda6",
"mem-c7a77457-478d-4eb0-a116-67205a0066a4",
"bl-1b20e9bc-eb37-4907-8d63-e311fd61eab8",
"bl-aa762207-920d-45ab-b2a3-2f8154d7ef9b",
"tag-misalignment",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 3932.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"kn-48a01973-a025-471d-950f-b93e6a426d82",
"mem-23d22bc1-a097-446b-8f11-8aff099e0b76",
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"mem-ce5a2ffc-ad39-4728-9ac6-76fef507d5da",
"project-Stripe_Elements__not_hosted_checkout__Custom_URL__DAG_bundle_pricing__Stripe_Connect_80_20_",
"tag-provenance",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a"
],
"n_returned": 10,
"latency_ms": 4836.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"kn-79056192-7de8-486c-9565-f128439a6fcb",
"bl-f6236350-f7b8-4f4f-a702-9eef2eb76e4b",
"bl-7f33f1bc-99fa-4906-889f-a42375beea20",
"mem-ef0091d8-1b65-431e-afa8-c6c4ee5779c9",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"tag-__cultivation-metric____internal-state____dharma____evidence____novel-idea____gap-compression____values____microsoft__",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 4057.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-identity-studio",
"tag-finance",
"tag-temporal",
"tag-barkhausen",
"tag-ilogger",
"tag-design-first",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 5098.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"project-Goal_setting__alignment__scoring__cadence__Attaches_to_any_imprint_",
"tag-performed-values",
"tag-turing-test",
"tag-sealed",
"tag-part-5",
"tag-model",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
";??A5???"
],
"n_returned": 10,
"latency_ms": 4652.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-9590ba23-bddb-43e8-a571-68a263c4c364",
"project-Imprint__leadership_development__feedback_frameworks__performance__presence_",
"bl-a9e57bb2-00a1-4867-ab59-5d9271134b50",
"bl-c7793c4a-7630-47fc-a462-d23059087e80",
"tag-gateway_platform_neuron-technologies_go_proxy_llm",
"tag-fornax",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 4653.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"bl-7a13527b-3e0c-418a-9f37-88fd2152e5ce",
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"bl-5e390b10-8753-4f25-a1a5-b5dbbb002cbf",
"tag-ats",
"tag-data-model",
"tag-aggregation",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
";??A5???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4272.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"tag-memory-model",
"tag-storage",
"tag-ga4",
"tag-potions",
"tag-resonance",
"tag-neuron",
"? ?}&?#??X\b",
";??A5???",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4541.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"mem-927f41ab-8ede-4f58-acb3-995db16ac775",
"mem-dbe80bc2-c602-46b0-b4ea-dd222e52bcde",
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"tag-upload-window",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 4883.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"mem-481c769c-68cc-45c7-bc37-c0d9778fa648",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-9d8f3c5b-4bac-41ce-8ac4-44733f99d6c8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3237.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-eb1c6d99-d603-4f33-be9a-c63a178690c6",
"mem-a3124d5b-2f50-477f-8bb5-06879f5a496c",
"bl-7fa1b1a8-b80a-4f28-b162-bfe73765b4f8",
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 2457.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"mem-265a7107-73b3-4410-9aff-43787d5f473b",
"mem-9d1bf963-1b40-4588-bdb3-0432646cc623",
"mem-46780047-63a0-4a86-a16b-638b72a7fb8d",
"mem-3987d374-3c48-4e8e-b06d-0c363f55ed9c",
"bl-e0a0df72-de6e-46ab-800b-e1e3e8dfc387",
"tag-__patents____swarm____claim-language____prior-art____filing__",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 2640.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-dca14c4c-4859-47b0-996e-33964ba61a87",
"bl-dc8c7e02-eb37-48ae-a6f8-9b512803ae16",
"mem-8d1bafe6-209c-456c-9a25-9a927bc5a16d",
"bl-0de4e61b-6562-49e5-b7df-ebb809a01723",
"bl-8ef1ba6b-3fa0-4dbd-98c5-31665e5694a1",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?Q??m?;`'",
"R^??m?;?'"
],
"n_returned": 10,
"latency_ms": 3705.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-394cc9e8-049b-45bc-a380-66314f14e367",
"bl-ffa22d7e-42a9-4bd2-a428-1d2df243ac93",
"bl-4a6746e8-191f-48fc-8bfb-c4dc73b80bcd",
"tag-command-pattern",
"tag-withholding",
"tag-offline",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4223.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 1333.9,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 880.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"kn-333542cb-6dab-4662-9725-bf7440d28bf7",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 8,
"latency_ms": 1452.7,
"error": null,
"clean": false,
"false_positives": 8
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"ctx-e5427d7d",
"mem-37b57f52-a29a-42cf-a07a-3c5f8a3598dd",
"project-worldweaver",
"tag-__kotlin____internal-state____pre-reasoning____post-reasoning____compression-ratio____dharma____cultivation__",
"tag-import",
"tag-__cgi____dharma____cultivation____five-primitives____seed-artifact____agi____intelligence____whitepaper____patent__",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 3843.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-bfe0fafd-2750-4fdc-b773-04e878b3b23f",
"art-7bdaff30-5af9-4f0a-93b1-751686f9de3d",
"mem-cf07910d-4676-4384-ab97-9cad946cd0b9",
"mem-f9da4b43-3724-4bc8-92f8-6f237c89dc4d",
"project-Convert_UTC_timestamps_to_Central_time_when_displaying_to_Will__Never_surface_raw_UTC_",
"mem-32203649-3213-4d6d-86fd-3d657ac70d77",
";??A5???",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 5264.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"project-Define_Iris_as_separate_public_brand__Uncensored_but_principled__Consumer_face_while_Neuron_runs_enterprise_",
"project-Imprint__discovery__objection_handling__deal_strategy__pipeline__closing_",
"bl-8116da7a-b039-4e08-b8d0-c1c7861f9766",
"tag-enterprise",
"tag-divisors",
"tag-distressed-property",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"???Ͼd??W\b?",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 3022.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+942
View File
@@ -0,0 +1,942 @@
{
"label": "act-r3",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-act",
"soul_md5": "77722f5a9f49494bf735c2a4be1b5dc4",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7903,
"wall_clock_s": 112.0,
"child_pid": 78802,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.2571428571428571,
"recall@5": 0.19291383219954647,
"recall@10": 0.244812925170068,
"precision@5": 0.08,
"mrr@10": 0.266031746031746,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3220.9,
"latency_ms_p95": 4840.2,
"latency_ms_max": 5073.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.42857142857142855,
"recall@5": 0.10742630385487528,
"recall@10": 0.366921768707483,
"mrr@10": 0.47301587301587306
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 471.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 548.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 470.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 476.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 515.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 458.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-69fd6e83-7718-4824-8d66-f49d8954e224",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-c3d9d063-8c5d-45aa-900c-550914b2ff6d",
"kn-82be4e41-96c5-4da3-85c0-cee10763d975",
"bl-76e878aa-e1fe-468c-bf9c-854097cb7e0b",
"art-9887867c-2e21-47c3-9f96-c2dfe5bd4cc1",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451"
],
"n_returned": 10,
"latency_ms": 1595.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"bl-80720fdf-7ce7-4d28-aff8-21028d3a8cfb",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"?Z?.\f?0?]P?",
"?Q??m?;`'"
],
"n_returned": 10,
"latency_ms": 927.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.125,
"recall@10": 0.1875,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 6,
"latency_ms": 963.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5555555555555556,
"recall@10": 0.5555555555555556,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a",
"bl-7aebe936-ac55-4f35-8932-adc5224ff854",
"knw-b046991d-5992-4ac4-b854-7d3ac273832c",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"bl-455a08cf-5831-4fdb-b42c-b952f2feafb9",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"n_returned": 10,
"latency_ms": 927.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-8dbceb06-431a-416d-a723-e8c75d595154",
"mem-a3124d5b-2f50-477f-8bb5-06879f5a496c",
"art-94fae615-7cd5-4695-b968-977101b06a51",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-57164d5f-baf0-4149-957a-379a4e255d1a",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 937.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"kn-b2a99cd7-b379-4d9b-a996-e347a02c7bad",
"bl-a313d67b-dd6d-4e5b-a55a-03bc7bda17ae",
"kn-8e1bfb48-33a9-45ad-8da7-e0bdaa5d34e7",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023",
"bl-a7a1428f-db9c-417b-8e2c-713b1f84dc1f",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?Z?.\f?0?]P?",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 2451.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07142857142857142,
"recall@10": 0.21428571428571427,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-e20944e5-f4a6-44a0-91b1-73d04ebed120",
"mem-64cf3728-674c-404b-965a-b8f8d38bb7bb",
"mem-47f72b5b-6e8b-4293-94f1-350197b4809a",
"mem-e612f0aa-c2f2-4ee3-bbc7-af2dc826233b",
"mem-a5f04e52-91f8-41d2-af27-8bf803621758",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"n_returned": 10,
"latency_ms": 2042.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"mem-ea487cb4-ed67-44ce-8402-b56bb28468d4",
"art-c71aef51-026f-4d63-80e9-2a0ec0dc3865",
"bl-2dd8aaa1-b0de-4eac-b3c5-78951d240b60",
"bl-2694b588-a6e3-43de-861c-fa7b0ec7e7fd",
"bl-14883d81-f7cb-46dd-82c2-a6e6980264e5",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad"
],
"n_returned": 10,
"latency_ms": 4698.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"bl-fd047ce9-ae21-4b3e-b3ab-ece0c9592f7f",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"bl-9ce4128a-9436-4b06-82bc-8a6faafa81e0",
"bl-6172d035-dd94-4776-afdd-d8915f6fc375",
"bl-5bb8dedf-8498-4a9b-acdc-31cc9c738f2a",
"tag-cultivated-general-intelligence",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 5073.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"bl-e5635b1a-c5d0-4caa-bcc8-6a726ea43685",
"bl-fc893be3-e6b4-4ef6-93b0-d54ca5f89083",
"tag-project-structure",
"tag-anthropic-contrast",
"tag-voice-training",
"tag-ebd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792"
],
"n_returned": 10,
"latency_ms": 4117.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-d24fcce8-2b55-426f-867a-db3958a622d3",
"tag-phase-3",
"bl-57c5cf6b-81a5-4558-9902-5c02981fe273",
"tag-guilds",
"tag-kids",
"tag-coexistence",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023"
],
"n_returned": 10,
"latency_ms": 3254.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"mem-0328c3cb-4550-4ce4-9284-152e832f08f6",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-bd9fb314-e9d4-4b03-aef4-534dd57a2992",
"bl-e98cdd4c-01b5-459e-9036-3578cd5d975a",
"tag-stable-diffusion",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 4081.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"mem-e32ba5a7-c147-4dc0-9479-b720d768eda6",
"bl-b019ce7a-1b21-436e-812d-032f50c6c45f",
"mem-c7a77457-478d-4eb0-a116-67205a0066a4",
"bl-1b20e9bc-eb37-4907-8d63-e311fd61eab8",
"bl-aa762207-920d-45ab-b2a3-2f8154d7ef9b",
"tag-misalignment",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 3687.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"kn-48a01973-a025-471d-950f-b93e6a426d82",
"mem-23d22bc1-a097-446b-8f11-8aff099e0b76",
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"mem-ce5a2ffc-ad39-4728-9ac6-76fef507d5da",
"project-Stripe_Elements__not_hosted_checkout__Custom_URL__DAG_bundle_pricing__Stripe_Connect_80_20_",
"mem-e321e54e-8bb3-4596-b13d-bb093d6b149d",
"? ?}&?#??X\b",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9"
],
"n_returned": 10,
"latency_ms": 4659.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"kn-79056192-7de8-486c-9565-f128439a6fcb",
"bl-f6236350-f7b8-4f4f-a702-9eef2eb76e4b",
"bl-7f33f1bc-99fa-4906-889f-a42375beea20",
"mem-37b57f52-a29a-42cf-a07a-3c5f8a3598dd",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"tag-__cultivation-metric____internal-state____dharma____evidence____novel-idea____gap-compression____values____microsoft__",
"? ?}&?#??X\b",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"knw-b046991d-5992-4ac4-b854-7d3ac273832c"
],
"n_returned": 10,
"latency_ms": 3960.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-identity-studio",
"tag-finance",
"tag-temporal",
"tag-barkhausen",
"tag-ilogger",
"tag-design-first",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 4840.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"project-Goal_setting__alignment__scoring__cadence__Attaches_to_any_imprint_",
"tag-performed-values",
"tag-turing-test",
"tag-sealed",
"tag-part-5",
"tag-model",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4493.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-9590ba23-bddb-43e8-a571-68a263c4c364",
"project-Imprint__leadership_development__feedback_frameworks__performance__presence_",
"bl-a9e57bb2-00a1-4867-ab59-5d9271134b50",
"bl-c7793c4a-7630-47fc-a462-d23059087e80",
"tag-gateway_platform_neuron-technologies_go_proxy_llm",
"tag-fornax",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 4509.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"bl-7a13527b-3e0c-418a-9f37-88fd2152e5ce",
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"bl-5e390b10-8753-4f25-a1a5-b5dbbb002cbf",
"tag-ats",
"tag-data-model",
"tag-aggregation",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 3624.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"tag-memory-model",
"tag-storage",
"tag-ga4",
"tag-potions",
"tag-resonance",
"tag-neuron",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 4093.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"mem-927f41ab-8ede-4f58-acb3-995db16ac775",
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"tag-upload-window",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 4235.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"mem-0f31141d-3ac5-44b2-9942-be7e4e6feb79",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"mem-481c769c-68cc-45c7-bc37-c0d9778fa648",
"bl-9d8f3c5b-4bac-41ce-8ac4-44733f99d6c8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3220.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1c6d99-d603-4f33-be9a-c63a178690c6",
"mem-5e7f6ddd-c818-4ad3-b564-54ae278e9976",
"bl-7fa1b1a8-b80a-4f28-b162-bfe73765b4f8",
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c",
"tag-sarah",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 2336.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"mem-265a7107-73b3-4410-9aff-43787d5f473b",
"mem-9d1bf963-1b40-4588-bdb3-0432646cc623",
"mem-46780047-63a0-4a86-a16b-638b72a7fb8d",
"mem-3987d374-3c48-4e8e-b06d-0c363f55ed9c",
"bl-e0a0df72-de6e-46ab-800b-e1e3e8dfc387",
"tag-__patents____swarm____claim-language____prior-art____filing__",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 2348.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-dca14c4c-4859-47b0-996e-33964ba61a87",
"bl-e148d23c-24e8-4122-9915-d1c11f22052f",
"bl-dc8c7e02-eb37-48ae-a6f8-9b512803ae16",
"mem-8d1bafe6-209c-456c-9a25-9a927bc5a16d",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"tag-dark-theme",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?Q??m?;`'",
"R^??m?;?'"
],
"n_returned": 10,
"latency_ms": 3511.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-ffa22d7e-42a9-4bd2-a428-1d2df243ac93",
"mem-ef0091d8-1b65-431e-afa8-c6c4ee5779c9",
"bl-452a4710-3d2b-4e0f-9413-49a66423bc9a",
"bl-4a6746e8-191f-48fc-8bfb-c4dc73b80bcd",
"tag-withholding",
"tag-offline",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 4160.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 1318.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 875.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"kn-333542cb-6dab-4662-9725-bf7440d28bf7",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 8,
"latency_ms": 1393.4,
"error": null,
"clean": false,
"false_positives": 8
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"knw-b046991d-5992-4ac4-b854-7d3ac273832c",
"ctx-e5427d7d",
"project-worldweaver",
"tag-__kotlin____internal-state____pre-reasoning____post-reasoning____compression-ratio____dharma____cultivation__",
"tag-import",
"tag-enterprise",
"tag-__cgi____dharma____cultivation____five-primitives____seed-artifact____agi____intelligence____whitepaper____patent__",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3775.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-bfe0fafd-2750-4fdc-b773-04e878b3b23f",
"art-7bdaff30-5af9-4f0a-93b1-751686f9de3d",
"mem-cf07910d-4676-4384-ab97-9cad946cd0b9",
"mem-f9da4b43-3724-4bc8-92f8-6f237c89dc4d",
"project-Convert_UTC_timestamps_to_Central_time_when_displaying_to_Will__Never_surface_raw_UTC_",
"mem-32203649-3213-4d6d-86fd-3d657ac70d77",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 4857.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"bl-0de4e61b-6562-49e5-b7df-ebb809a01723",
"project-Imprint__discovery__objection_handling__deal_strategy__pipeline__closing_",
"bl-8116da7a-b039-4e08-b8d0-c1c7861f9766",
"bl-8ef1ba6b-3fa0-4dbd-98c5-31665e5694a1",
"tag-divisors",
"tag-distressed-property",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"knw-b046991d-5992-4ac4-b854-7d3ac273832c",
"???Ͼd??W\b?"
],
"n_returned": 10,
"latency_ms": 2858.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
@@ -0,0 +1,956 @@
{
"label": "assoc-leg-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-assoc",
"soul_md5": "ab9d490ecdfb1f9e6f23cca841ad8fb5",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 51.8,
"child_pid": 87150,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6285714285714286,
"recall@5": 0.45309194773480493,
"recall@10": 0.5405733155733157,
"precision@5": 0.17714285714285719,
"mrr@10": 0.42650793650793645,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1223.5,
"latency_ms_p95": 1676.1,
"latency_ms_max": 1740.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657342,
"recall@10": 0.24825174825174826,
"mrr@10": 0.22777777777777777
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4882369614512472,
"recall@10": 0.6329365079365079,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 306.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"project-engram",
"project-engram-lang",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-89a2",
"bl-13babd0c-582e-4e28-a9e4-a77e65925e5d",
"870ede67-3454-4e00-9988-46cb13a8a4e2",
"bl-3e433255-3710-49fc-a093-c25e71de2ccb",
"mem-235a7657-d49e-467e-9f69-f4c3d5f6bd48"
],
"n_returned": 10,
"latency_ms": 331.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 292.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 309.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 331.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 295.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"????7???Ջ3",
"kn-363f4976-6946-4b4d-b51b-8a2b0f5aef25",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"ctx-63e3",
"?ǚ?7??????",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 10,
"latency_ms": 613.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 538.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 553.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.4444444444444444,
"recall@10": 0.4444444444444444,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"ԍ????X????",
"project-harmonic-framework",
"?ǚ?7??????",
"project-harmonic-framework_com",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"??????X??2c",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"????7???Ջ3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 527.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 546.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 892.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"?",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"?",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"?",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"?"
],
"n_returned": 10,
"latency_ms": 754.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1644.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1740.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"? ?}&?#??X\b",
"????7???Ջ3",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e"
],
"n_returned": 10,
"latency_ms": 1393.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"?of?7???",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1118.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"ctx-cc7f",
"?of?7???",
"ctx-4a41",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"a1000001-0000-0000-0000-000000000001",
"knw-729fc901-8335-44c4-9f3a-b150b4aa0915",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"ctx-175f"
],
"n_returned": 9,
"latency_ms": 1440.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-23c27d3b-e0d2-43a8-a80c-0a44477ae18a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 10,
"latency_ms": 1316.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"? ?}&?#??X\b",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"? ?}&?#??X\b",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 1620.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"? ?}&?#??X\b",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"??????X??2c",
"dR????X?-?S"
],
"n_returned": 10,
"latency_ms": 1393.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1676.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"????7???Ջ3",
"a1000001-0000-0000-0000-000000000001",
"?ǚ?7??????",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"7?e?7???\f3?",
"mem-ade9440f-f161-4c18-9b35-1976257e6ebb",
"?of?7???",
"ea95f600-8dfd-4c7e-b077-a93dc3cd3623"
],
"n_returned": 10,
"latency_ms": 1534.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1567.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"?ǚ?7??????",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"7?e?7???\f3?",
"ctx-bb74",
"??S?7???",
"knw-f671966c-3387-4848-abca-b5deec122e00"
],
"n_returned": 10,
"latency_ms": 1280.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"7?e?7???\f3?",
"mem-7b74cac0-905f-4c35-9688-fbcce105a177"
],
"n_returned": 10,
"latency_ms": 1401.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"n_returned": 10,
"latency_ms": 1512.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.36363636363636365,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1165.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40"
],
"n_returned": 10,
"latency_ms": 1230.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.36363636363636365,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"ctx-e5427d7d",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"ctx-bb74",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-8c2d5f51-3ccd-4c2e-848a-eb60d90a3b98"
],
"n_returned": 10,
"latency_ms": 1223.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 9,
"latency_ms": 1235.8,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.18181818181818182,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1460.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.2727272727272727,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 750.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 531.1,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?V?",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"?m?\\}Q??6??",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 769.1,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"? ?}&?#??X\b",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"? ?}&?#??X\b",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"?of?7???",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"? ?}&?#??X\b",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"n_returned": 10,
"latency_ms": 1327.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 10
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"12082f7e-e320-438b-bd65-083d8259748f",
"? ?}&?#??X\b",
"527ecb25-2587-47eb-8269-73be2431abd4",
"? ?}&?#??X\b",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"? ?}&?#??X\b",
"be3b6036-6eca-44a7-8fdf-37b23edfdfd1"
],
"n_returned": 10,
"latency_ms": 1687.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666,
"outranks": true,
"rank_correct": 6,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"7?e?7???\f3?",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"? ?}&?#??X\b",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"%???2??jH??",
"art-8a0870d5-a716-4672-8094-f7463af1265b",
"???Ͼd??W\b?",
"bl-556438af-57b2-4bd8-a747-9f868aaee290"
],
"n_returned": 10,
"latency_ms": 1040.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 4
}
]
}
+956
View File
@@ -0,0 +1,956 @@
{
"label": "assoc-leg",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-assoc",
"soul_md5": "ab9d490ecdfb1f9e6f23cca841ad8fb5",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7893,
"wall_clock_s": 52.4,
"child_pid": 87099,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6285714285714286,
"recall@5": 0.45309194773480493,
"recall@10": 0.5405733155733157,
"precision@5": 0.17714285714285719,
"mrr@10": 0.42650793650793645,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1228.5,
"latency_ms_p95": 1681.8,
"latency_ms_max": 1718.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657342,
"recall@10": 0.24825174825174826,
"mrr@10": 0.22777777777777777
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4882369614512472,
"recall@10": 0.6329365079365079,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 306.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"project-engram",
"project-engram-lang",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-89a2",
"bl-13babd0c-582e-4e28-a9e4-a77e65925e5d",
"870ede67-3454-4e00-9988-46cb13a8a4e2",
"bl-3e433255-3710-49fc-a093-c25e71de2ccb",
"mem-235a7657-d49e-467e-9f69-f4c3d5f6bd48"
],
"n_returned": 10,
"latency_ms": 360.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 324.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 313.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 331.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 303.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"????7???Ջ3",
"kn-363f4976-6946-4b4d-b51b-8a2b0f5aef25",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"ctx-63e3",
"?ǚ?7??????",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 10,
"latency_ms": 588.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 533.3,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 542.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.4444444444444444,
"recall@10": 0.4444444444444444,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"ԍ????X????",
"project-harmonic-framework",
"?ǚ?7??????",
"project-harmonic-framework_com",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"??????X??2c",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"????7???Ջ3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 517.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 547.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 903.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"?",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"?",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"?",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"?"
],
"n_returned": 10,
"latency_ms": 774.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1641.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1718.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"? ?}&?#??X\b",
"????7???Ջ3",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e"
],
"n_returned": 10,
"latency_ms": 1409.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"?of?7???",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1135.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"ctx-cc7f",
"?of?7???",
"ctx-4a41",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"a1000001-0000-0000-0000-000000000001",
"knw-729fc901-8335-44c4-9f3a-b150b4aa0915",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"ctx-175f"
],
"n_returned": 9,
"latency_ms": 1446.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-23c27d3b-e0d2-43a8-a80c-0a44477ae18a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 10,
"latency_ms": 1296.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"? ?}&?#??X\b",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"? ?}&?#??X\b",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 1621.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"? ?}&?#??X\b",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"??????X??2c",
"dR????X?-?S"
],
"n_returned": 10,
"latency_ms": 1411.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1681.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"????7???Ջ3",
"a1000001-0000-0000-0000-000000000001",
"?ǚ?7??????",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"7?e?7???\f3?",
"mem-ade9440f-f161-4c18-9b35-1976257e6ebb",
"?of?7???",
"ea95f600-8dfd-4c7e-b077-a93dc3cd3623"
],
"n_returned": 10,
"latency_ms": 1536.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1567.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"?ǚ?7??????",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"7?e?7???\f3?",
"ctx-bb74",
"??S?7???",
"knw-f671966c-3387-4848-abca-b5deec122e00"
],
"n_returned": 10,
"latency_ms": 1278.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"7?e?7???\f3?",
"mem-7b74cac0-905f-4c35-9688-fbcce105a177"
],
"n_returned": 10,
"latency_ms": 1401.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"n_returned": 10,
"latency_ms": 1509.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.36363636363636365,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1156.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40"
],
"n_returned": 10,
"latency_ms": 1232.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.36363636363636365,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"ctx-e5427d7d",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"ctx-bb74",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-8c2d5f51-3ccd-4c2e-848a-eb60d90a3b98"
],
"n_returned": 10,
"latency_ms": 1228.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 9,
"latency_ms": 1251.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.18181818181818182,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1459.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.2727272727272727,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 746.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 519.1,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?V?",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"?m?\\}Q??6??",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 764.0,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"? ?}&?#??X\b",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"? ?}&?#??X\b",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"?of?7???",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"? ?}&?#??X\b",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"n_returned": 10,
"latency_ms": 1346.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 10
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"12082f7e-e320-438b-bd65-083d8259748f",
"? ?}&?#??X\b",
"527ecb25-2587-47eb-8269-73be2431abd4",
"? ?}&?#??X\b",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"? ?}&?#??X\b",
"be3b6036-6eca-44a7-8fdf-37b23edfdfd1"
],
"n_returned": 10,
"latency_ms": 1687.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666,
"outranks": true,
"rank_correct": 6,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"7?e?7???\f3?",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"? ?}&?#??X\b",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"%???2??jH??",
"art-8a0870d5-a716-4672-8094-f7463af1265b",
"???Ͼd??W\b?",
"bl-556438af-57b2-4bd8-a747-9f868aaee290"
],
"n_returned": 10,
"latency_ms": 1030.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 4
}
]
}
@@ -0,0 +1,945 @@
{
"label": "baseline-embcorpus",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-baseline",
"soul_md5": "5cc9521734907cf2da30f0af94498c06",
"corpus": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/corpus-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7893,
"wall_clock_s": 48.4,
"child_pid": 85995,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1145.9,
"latency_ms_p95": 1574.3,
"latency_ms_max": 1634.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 229.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 262.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 230.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 229.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 250.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 227.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 532.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 457.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 472.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 446.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 462.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 832.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 686.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1550.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1634.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1322.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1041.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1350.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1213.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1553.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1319.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1574.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"?ǚ?7??????",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"????7???Ջ3",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"%???2??jH??",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1441.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1499.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1197.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1311.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1422.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1077.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1149.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1145.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1177.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1382.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 660.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 432.5,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 676.9,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1252.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1624.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 946.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
@@ -0,0 +1,959 @@
{
"label": "bm25lex-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-bm25lex",
"soul_md5": "dfbd0f8e3646212db5c60026f8ad906f",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 50.1,
"child_pid": 90371,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1183.8,
"latency_ms_p95": 1611.3,
"latency_ms_max": 1654.7,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 282.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"bl-ba764d70-e9d7-4f62-848f-719cb665f45e",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"bl-b28d7256-6f74-4567-bd90-40d0ef2a6d78",
"project-engram",
"ctx-45bc",
"project-engram-lang",
"ctx-175f",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-74ed"
],
"n_returned": 10,
"latency_ms": 315.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 283.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 283.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 303.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 283.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"tag-patterns",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"project-Imprint__system_design__ADRs__tech_strategy__integration_patterns__governance_",
"project-Imprint__analysis_patterns__data_storytelling__SQL__dashboards__insight_framing_",
"bl-79028eed-c330-4724-9402-734062d13503",
"bl-39dad13d-7105-4049-8224-dc3c34fdb1f3",
"bl-4ef4d914-da46-4e0f-be78-5219b9547e9f",
"bl-7e7c3fdb-4132-487f-aa70-b2cd559cb7f0",
"bl-1d32bd54-cf17-4a1f-b235-982d09a36f04",
"bl-b8af6601-a8cb-41b5-aef5-ab8a57432dd5"
],
"n_returned": 10,
"latency_ms": 587.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 518.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"n_returned": 10,
"latency_ms": 511.3,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.4444444444444444,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"project-harmonic-framework",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"project-harmonic-framework_com",
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"tag-harmonic-design",
"bl-92acd4eb-0452-4e8e-9f54-f8cd35170d76",
"tag-harmonic-framework",
"bl-18a9d1e4-1484-474c-bf6b-c6173212181b"
],
"n_returned": 10,
"latency_ms": 497.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1111111111111111,
"recall@10": 0.1111111111111111,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-sarah",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-a6cb3b8d-d89c-46fc-931d-e90c560783b0",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 510.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-6f99e111-7055-4635-9831-a489747ce418",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-967536a0-d49d-44fb-8cfb-b31b40bcbfae",
"bl-8b58d9bc-352b-4842-a7f8-a6254b5d1e25",
"2c56a7a9-5323-4ce4-ba09-35836ba15d54",
"bl-39cec462-c80c-4970-a3aa-91fe83053bde"
],
"n_returned": 10,
"latency_ms": 870.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.21428571428571427,
"recall@10": 0.2857142857142857,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"?",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"?",
"830ca37a-d334-4e41-ba89-64893dc8d628",
"?",
"ce9636dc-85a5-4dae-9e07-74ea2fcc6307",
"?"
],
"n_returned": 10,
"latency_ms": 723.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"cb070131-dfd4-4a38-91d7-22b1bde164d2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"knw-d788a210-613b-4c49-9486-88bbc9d4716f",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"66c63082-b4da-4aa1-8fee-848db8a83210",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-a535f205-bc4c-4058-9171-6263c496044a",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"ctx-4a41"
],
"n_returned": 10,
"latency_ms": 1584.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"b1183213-d659-4759-85d7-5b1f22427fe2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-16efddd1-c43d-4a42-9d78-f54fb82bd277",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"f0eb6b13-909c-4674-91ef-23301d3abc8b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"30a44d10-2487-420e-bf61-3892e4343c92",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"bfb5809e-d19a-4d3f-8c1a-796db622ad9d"
],
"n_returned": 10,
"latency_ms": 1652.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"mem-ef878e30-5851-4e82-8588-745415108941",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"tag-fiction",
"knw-8fd9836c-cc39-49df-8d61-babda626cc88",
"mem-8d690e9d-a7e9-4062-b2f8-e2064294e463",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"mem-ce793303-c5a5-4586-a232-a3426edd9ec7",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"mem-443bd012-fc9a-4088-b236-de5157a1ef92"
],
"n_returned": 10,
"latency_ms": 1341.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-8de20bcf-7149-4f48-b67c-e7f9758fd6e5",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"bl-164b520b-c503-49db-89f9-bd2fdf4215f5",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"1219277c-1b95-45ec-95a2-07b4a47a4d92",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"08f0d1e2-8d0e-42e3-9f0a-8186ae31ec7e",
"bl-79ce4464-5dd6-49bd-9b0c-9803549d0665"
],
"n_returned": 10,
"latency_ms": 1064.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"a1000001-0000-0000-0000-000000000010",
"a1000001-0000-0000-0000-000000000009",
"bl-448bc514-c2f1-4520-a9b1-1f3a73678d26",
"a1000001-0000-0000-0000-000000000012",
"43098881-e044-482b-8e92-471728a8ba8b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"mem-e5cc63c0-8701-49d6-855a-e387fe087771",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 1398.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"tag-learning",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"d6b12ecf-702b-4101-b1bb-09ed9b220b29",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"c608a095-c98b-4bfa-bfe1-1611c1320290",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"451ae007-4219-4096-89fe-fa2e045fbeb1",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 1252.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"mem-cde58b77-50d3-4bac-9581-e70a4c02c015",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-8d699e2c-ac2a-4742-bb62-b6da00f4b10e",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"mem-53d6adf0-cd08-4707-a237-daa5e65c7298",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"bl-ef2bac68-e119-4139-b529-c7a1404ae3ac"
],
"n_returned": 10,
"latency_ms": 1581.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"bl-0e8f4880-7b24-43aa-aed9-ad4d9fc73ff8",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"bl-8848929a-a23a-46bc-a2c7-fe3a3bc1cddf",
"bl-205141ad-b2a0-4d93-86d0-89eb0723e1bd",
"bl-e93858c4-7cac-4b1a-bb62-490790d4c3f3",
"bl-34f51ddb-a840-459f-a248-94214f5febb6",
"bl-286b562a-5299-40e0-a32a-afa9cbdfe995"
],
"n_returned": 10,
"latency_ms": 1363.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"5a2c118a-87bd-4239-97a7-9e02c5991983",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"mem-ef878e30-5851-4e82-8588-745415108941",
"knw-12b4b913-7a25-4b0d-844c-504c01d6725e",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"knw-9707256e-ed44-4042-bd88-f90fa514e1cf",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"bl-4c5b385e-135a-4663-8521-96af0b491121",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"n_returned": 10,
"latency_ms": 1611.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"a0edad47-5f77-4fc3-a546-1e85f8c68e77",
"a1000001-0000-0000-0000-000000000001",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"96497334-b18f-495c-9228-eeb8182bdc38",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"knw-ed33e669-0790-44cb-a036-958d605c6fea"
],
"n_returned": 10,
"latency_ms": 1470.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"27e1b1a4-ad0b-49d9-812f-fedf43b8aabe",
"knw-f9ce17a7-17fc-431f-8f23-695b670ec4fa",
"bl-87c93185-b2bf-40af-ae23-3c830c007abf",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"077d064f-3489-4c05-9aca-3782f96b51db",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"75e036d3-c170-4e3f-acc2-e456a6850ee2",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 1513.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"e4f27651-52c5-43fd-aff3-61d31685b3cd",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"mem-5624ec9d-62ba-4aba-8a3d-6afec6c09dd4",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-833dbbcd-2400-4594-bb35-93b023049ac0",
"a1000001-0000-0000-0000-000000000009",
"mem-759e78ca-5394-4244-aa39-1c1468bc5f3e",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd"
],
"n_returned": 10,
"latency_ms": 1226.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"bl-0d8c5dfa-e163-4fef-a58b-56b0d076c5a8",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-b99efff0-00e6-40c8-9c5b-730330eef33b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"tag-childhood"
],
"n_returned": 10,
"latency_ms": 1342.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"3499d5da-0e9c-4de4-9bc4-8941b14e0b1f",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 1433.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-9110798f-d0cb-4446-bc2a-14f09b6a09e2",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1104.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-trailer-park-paladins",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"project-trailer-park-paladins",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 1183.8,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"bl-bea7473c-c687-414c-9c0b-00c509a616c1",
"bl-fc6fcb0b-9e4b-40bf-8e88-dbfe4e27c31a",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc"
],
"n_returned": 10,
"latency_ms": 1199.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-hope",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 1202.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.18181818181818182,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"knw-729fc901-8335-44c4-9f3a-b150b4aa0915",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-528dbc37-eabc-4b75-a7a5-65bf38d6018a",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"knw-473f3f24-20f6-4f39-8589-3709538eb6ac",
"?Z?<S???K ?",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"'?T?a\"B~-?8",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 1414.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 703.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 492.9,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"?V?",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"?m?\\}Q??6??",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"kn-333542cb-6dab-4662-9725-bf7440d28bf7"
],
"n_returned": 10,
"latency_ms": 728.4,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____architecture__",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____kotlin____architecture__",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 1289.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 8
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"12082f7e-e320-438b-bd65-083d8259748f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"13705072-4515-4124-963d-083af490494f",
"527ecb25-2587-47eb-8269-73be2431abd4",
"6de314bf-5c4c-4cfc-871f-fa2e422d45e6",
"a1000001-0000-0000-0000-000000000002",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"7ac62daa-2eac-4c7a-a97e-e4203fc1b57b"
],
"n_returned": 10,
"latency_ms": 1654.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"5fcba804-eb5b-48ec-82da-146b1c6bb50d",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"bl-c8c19362-430b-4817-9cf4-9e85e0099c64",
"bl-c5c6571e-118f-47c7-8cbb-3ed0ebf64a51",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"ctx-3a55",
"86228228-7adf-41fb-b4c4-9ceea87953ae",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"bl-4f7b651b-6b33-449c-8a3b-cfce12ce984b",
"mem-3a2cf162-d93b-4f29-86f2-5066fb7fe1f5"
],
"n_returned": 10,
"latency_ms": 995.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 5
}
]
}
+959
View File
@@ -0,0 +1,959 @@
{
"label": "bm25lex",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-bm25lex",
"soul_md5": "dfbd0f8e3646212db5c60026f8ad906f",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7893,
"wall_clock_s": 50.4,
"child_pid": 90325,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1184.4,
"latency_ms_p95": 1620.0,
"latency_ms_max": 1655.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 287.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"bl-ba764d70-e9d7-4f62-848f-719cb665f45e",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"bl-b28d7256-6f74-4567-bd90-40d0ef2a6d78",
"project-engram",
"ctx-45bc",
"project-engram-lang",
"ctx-175f",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-74ed"
],
"n_returned": 10,
"latency_ms": 316.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 283.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 284.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 304.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 282.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"tag-patterns",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"project-Imprint__system_design__ADRs__tech_strategy__integration_patterns__governance_",
"project-Imprint__analysis_patterns__data_storytelling__SQL__dashboards__insight_framing_",
"bl-79028eed-c330-4724-9402-734062d13503",
"bl-39dad13d-7105-4049-8224-dc3c34fdb1f3",
"bl-4ef4d914-da46-4e0f-be78-5219b9547e9f",
"bl-7e7c3fdb-4132-487f-aa70-b2cd559cb7f0",
"bl-1d32bd54-cf17-4a1f-b235-982d09a36f04",
"bl-b8af6601-a8cb-41b5-aef5-ab8a57432dd5"
],
"n_returned": 10,
"latency_ms": 591.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 516.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"n_returned": 10,
"latency_ms": 520.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.4444444444444444,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"project-harmonic-framework",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"project-harmonic-framework_com",
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"tag-harmonic-design",
"bl-92acd4eb-0452-4e8e-9f54-f8cd35170d76",
"tag-harmonic-framework",
"bl-18a9d1e4-1484-474c-bf6b-c6173212181b"
],
"n_returned": 10,
"latency_ms": 507.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1111111111111111,
"recall@10": 0.1111111111111111,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-sarah",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-a6cb3b8d-d89c-46fc-931d-e90c560783b0",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 509.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-6f99e111-7055-4635-9831-a489747ce418",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-967536a0-d49d-44fb-8cfb-b31b40bcbfae",
"bl-8b58d9bc-352b-4842-a7f8-a6254b5d1e25",
"2c56a7a9-5323-4ce4-ba09-35836ba15d54",
"bl-39cec462-c80c-4970-a3aa-91fe83053bde"
],
"n_returned": 10,
"latency_ms": 872.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.21428571428571427,
"recall@10": 0.2857142857142857,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"?",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"?",
"830ca37a-d334-4e41-ba89-64893dc8d628",
"?",
"ce9636dc-85a5-4dae-9e07-74ea2fcc6307",
"?"
],
"n_returned": 10,
"latency_ms": 735.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"cb070131-dfd4-4a38-91d7-22b1bde164d2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"knw-d788a210-613b-4c49-9486-88bbc9d4716f",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"66c63082-b4da-4aa1-8fee-848db8a83210",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-a535f205-bc4c-4058-9171-6263c496044a",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"ctx-4a41"
],
"n_returned": 10,
"latency_ms": 1605.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"b1183213-d659-4759-85d7-5b1f22427fe2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-16efddd1-c43d-4a42-9d78-f54fb82bd277",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"f0eb6b13-909c-4674-91ef-23301d3abc8b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"30a44d10-2487-420e-bf61-3892e4343c92",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"bfb5809e-d19a-4d3f-8c1a-796db622ad9d"
],
"n_returned": 10,
"latency_ms": 1655.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"mem-ef878e30-5851-4e82-8588-745415108941",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"tag-fiction",
"knw-8fd9836c-cc39-49df-8d61-babda626cc88",
"mem-8d690e9d-a7e9-4062-b2f8-e2064294e463",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"mem-ce793303-c5a5-4586-a232-a3426edd9ec7",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"mem-443bd012-fc9a-4088-b236-de5157a1ef92"
],
"n_returned": 10,
"latency_ms": 1339.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-8de20bcf-7149-4f48-b67c-e7f9758fd6e5",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"bl-164b520b-c503-49db-89f9-bd2fdf4215f5",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"1219277c-1b95-45ec-95a2-07b4a47a4d92",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"08f0d1e2-8d0e-42e3-9f0a-8186ae31ec7e",
"bl-79ce4464-5dd6-49bd-9b0c-9803549d0665"
],
"n_returned": 10,
"latency_ms": 1065.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"a1000001-0000-0000-0000-000000000010",
"a1000001-0000-0000-0000-000000000009",
"bl-448bc514-c2f1-4520-a9b1-1f3a73678d26",
"a1000001-0000-0000-0000-000000000012",
"43098881-e044-482b-8e92-471728a8ba8b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"mem-e5cc63c0-8701-49d6-855a-e387fe087771",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 1388.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"tag-learning",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"d6b12ecf-702b-4101-b1bb-09ed9b220b29",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"c608a095-c98b-4bfa-bfe1-1611c1320290",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"451ae007-4219-4096-89fe-fa2e045fbeb1",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 1246.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"mem-cde58b77-50d3-4bac-9581-e70a4c02c015",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-8d699e2c-ac2a-4742-bb62-b6da00f4b10e",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"mem-53d6adf0-cd08-4707-a237-daa5e65c7298",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"bl-ef2bac68-e119-4139-b529-c7a1404ae3ac"
],
"n_returned": 10,
"latency_ms": 1595.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"bl-0e8f4880-7b24-43aa-aed9-ad4d9fc73ff8",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"bl-8848929a-a23a-46bc-a2c7-fe3a3bc1cddf",
"bl-205141ad-b2a0-4d93-86d0-89eb0723e1bd",
"bl-e93858c4-7cac-4b1a-bb62-490790d4c3f3",
"bl-34f51ddb-a840-459f-a248-94214f5febb6",
"bl-286b562a-5299-40e0-a32a-afa9cbdfe995"
],
"n_returned": 10,
"latency_ms": 1359.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"5a2c118a-87bd-4239-97a7-9e02c5991983",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"mem-ef878e30-5851-4e82-8588-745415108941",
"knw-12b4b913-7a25-4b0d-844c-504c01d6725e",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"knw-9707256e-ed44-4042-bd88-f90fa514e1cf",
"kn-22d77abe-b3c5-42fd-afcd-dcb87d924929",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"bl-4c5b385e-135a-4663-8521-96af0b491121",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"n_returned": 10,
"latency_ms": 1620.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"a0edad47-5f77-4fc3-a546-1e85f8c68e77",
"a1000001-0000-0000-0000-000000000001",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"96497334-b18f-495c-9228-eeb8182bdc38",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"knw-ed33e669-0790-44cb-a036-958d605c6fea"
],
"n_returned": 10,
"latency_ms": 1478.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"27e1b1a4-ad0b-49d9-812f-fedf43b8aabe",
"knw-f9ce17a7-17fc-431f-8f23-695b670ec4fa",
"bl-87c93185-b2bf-40af-ae23-3c830c007abf",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"077d064f-3489-4c05-9aca-3782f96b51db",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"75e036d3-c170-4e3f-acc2-e456a6850ee2",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 1516.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"e4f27651-52c5-43fd-aff3-61d31685b3cd",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"mem-5624ec9d-62ba-4aba-8a3d-6afec6c09dd4",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-833dbbcd-2400-4594-bb35-93b023049ac0",
"a1000001-0000-0000-0000-000000000009",
"mem-759e78ca-5394-4244-aa39-1c1468bc5f3e",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd"
],
"n_returned": 10,
"latency_ms": 1237.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"bl-0d8c5dfa-e163-4fef-a58b-56b0d076c5a8",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-b99efff0-00e6-40c8-9c5b-730330eef33b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"tag-childhood"
],
"n_returned": 10,
"latency_ms": 1337.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"3499d5da-0e9c-4de4-9bc4-8941b14e0b1f",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 1433.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-9110798f-d0cb-4446-bc2a-14f09b6a09e2",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1114.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-trailer-park-paladins",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"project-trailer-park-paladins",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 1184.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"bl-bea7473c-c687-414c-9c0b-00c509a616c1",
"bl-fc6fcb0b-9e4b-40bf-8e88-dbfe4e27c31a",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc"
],
"n_returned": 10,
"latency_ms": 1185.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-hope",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 1199.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.18181818181818182,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"knw-729fc901-8335-44c4-9f3a-b150b4aa0915",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-528dbc37-eabc-4b75-a7a5-65bf38d6018a",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"knw-473f3f24-20f6-4f39-8589-3709538eb6ac",
"?Z?<S???K ?",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"'?T?a\"B~-?8",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 1422.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 703.8,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 488.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"?V?",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"?m?\\}Q??6??",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"kn-333542cb-6dab-4662-9725-bf7440d28bf7"
],
"n_returned": 10,
"latency_ms": 720.7,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____architecture__",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____kotlin____architecture__",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 1301.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 8
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"12082f7e-e320-438b-bd65-083d8259748f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"13705072-4515-4124-963d-083af490494f",
"527ecb25-2587-47eb-8269-73be2431abd4",
"6de314bf-5c4c-4cfc-871f-fa2e422d45e6",
"a1000001-0000-0000-0000-000000000002",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"7ac62daa-2eac-4c7a-a97e-e4203fc1b57b"
],
"n_returned": 10,
"latency_ms": 1651.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"5fcba804-eb5b-48ec-82da-146b1c6bb50d",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"bl-c8c19362-430b-4817-9cf4-9e85e0099c64",
"bl-c5c6571e-118f-47c7-8cbb-3ed0ebf64a51",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"ctx-3a55",
"86228228-7adf-41fb-b4c4-9ceea87953ae",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"bl-4f7b651b-6b33-449c-8a3b-cfce12ce984b",
"mem-3a2cf162-d93b-4f29-86f2-5066fb7fe1f5"
],
"n_returned": 10,
"latency_ms": 1021.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 5
}
]
}
@@ -0,0 +1,942 @@
{
"label": "act-r1",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-act",
"soul_md5": "77722f5a9f49494bf735c2a4be1b5dc4",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 112.3,
"child_pid": 78648,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.22857142857142856,
"recall@5": 0.19087301587301586,
"recall@10": 0.24277210884353742,
"precision@5": 0.07428571428571429,
"mrr@10": 0.24154195011337865,
"nonsense_clean": "2/3",
"superseded_outranks": "0/3",
"latency_ms_p50": 3208.8,
"latency_ms_p95": 4851.9,
"latency_ms_max": 5078.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 2.6666666666666665
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.2857142857142857,
"recall@5": 0.09722222222222222,
"recall@10": 0.3567176870748299,
"mrr@10": 0.3505668934240363
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0,
"outranks": 0
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 468.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 569.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 479.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 470.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 499.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 452.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-69fd6e83-7718-4824-8d66-f49d8954e224",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-eb1c6d99-d603-4f33-be9a-c63a178690c6",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"kn-b2a99cd7-b379-4d9b-a996-e347a02c7bad",
"bl-76e878aa-e1fe-468c-bf9c-854097cb7e0b",
"art-c71aef51-026f-4d63-80e9-2a0ec0dc3865"
],
"n_returned": 10,
"latency_ms": 1592.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"bl-80720fdf-7ce7-4d28-aff8-21028d3a8cfb",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"?Z?.\f?0?]P?",
"?Q??m?;`'"
],
"n_returned": 10,
"latency_ms": 956.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.125,
"recall@10": 0.1875,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 6,
"latency_ms": 970.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5555555555555556,
"recall@10": 0.5555555555555556,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"bl-7aebe936-ac55-4f35-8932-adc5224ff854",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf"
],
"n_returned": 10,
"latency_ms": 920.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-8dbceb06-431a-416d-a723-e8c75d595154",
"mem-a3124d5b-2f50-477f-8bb5-06879f5a496c",
"knw-528dbc37-eabc-4b75-a7a5-65bf38d6018a",
"art-94fae615-7cd5-4695-b968-977101b06a51",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 925.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"mem-46780047-63a0-4a86-a16b-638b72a7fb8d",
"kn-8e1bfb48-33a9-45ad-8da7-e0bdaa5d34e7",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023",
"bl-a7a1428f-db9c-417b-8e2c-713b1f84dc1f",
"mem-f823e835-313f-4282-b4b3-ce527ffc2f7a",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?Z?.\f?0?]P?",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72"
],
"n_returned": 10,
"latency_ms": 2455.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.14285714285714285,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"art-92e1837c-5919-42d0-bbb0-4d924d7b2864",
"bl-e20944e5-eb16-4ab3-a84d-111e0fc817fa",
"bl-e20944e5-f4a6-44a0-91b1-73d04ebed120",
"mem-47f72b5b-6e8b-4293-94f1-350197b4809a",
"mem-e612f0aa-c2f2-4ee3-bbc7-af2dc826233b",
"mem-a5f04e52-91f8-41d2-af27-8bf803621758",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a"
],
"n_returned": 10,
"latency_ms": 2037.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.1
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"kn-82be4e41-96c5-4da3-85c0-cee10763d975",
"mem-ea487cb4-ed67-44ce-8402-b56bb28468d4",
"mem-23d22bc1-a097-446b-8f11-8aff099e0b76",
"bl-874d1c2b-c55b-4afb-9601-922a9297e859",
"bl-2dd8aaa1-b0de-4eac-b3c5-78951d240b60",
"bl-2694b588-a6e3-43de-861c-fa7b0ec7e7fd",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad"
],
"n_returned": 10,
"latency_ms": 4720.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"mem-0328c3cb-4550-4ce4-9284-152e832f08f6",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-fd047ce9-ae21-4b3e-b3ab-ece0c9592f7f",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"bl-6172d035-dd94-4776-afdd-d8915f6fc375",
"bl-5bb8dedf-8498-4a9b-acdc-31cc9c738f2a",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 5078.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"bl-e5635b1a-c5d0-4caa-bcc8-6a726ea43685",
"bl-fc893be3-e6b4-4ef6-93b0-d54ca5f89083",
"tag-project-structure",
"tag-anthropic-contrast",
"tag-voice-training",
"tag-ebd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"bl-9d53422d-b703-4f1d-860a-8598cb29b792",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4118.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-d24fcce8-2b55-426f-867a-db3958a622d3",
"tag-phase-3",
"tag-identity-studio",
"tag-kids",
"tag-coexistence",
"tag-cultivated-general-intelligence",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"mem-cdff0c49-3ac7-4de8-89ec-92d254bd0023"
],
"n_returned": 10,
"latency_ms": 3241.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"bl-bd9fb314-e9d4-4b03-aef4-534dd57a2992",
"bl-b019ce7a-1b21-436e-812d-032f50c6c45f",
"bl-e98cdd4c-01b5-459e-9036-3578cd5d975a",
"bl-9ce4128a-9436-4b06-82bc-8a6faafa81e0",
"tag-stable-diffusion",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 4090.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"mem-e321e54e-8bb3-4596-b13d-bb093d6b149d",
"mem-e32ba5a7-c147-4dc0-9479-b720d768eda6",
"mem-c7a77457-478d-4eb0-a116-67205a0066a4",
"bl-1b20e9bc-eb37-4907-8d63-e311fd61eab8",
"bl-aa762207-920d-45ab-b2a3-2f8154d7ef9b",
"tag-misalignment",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 3671.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"kn-48a01973-a025-471d-950f-b93e6a426d82",
"bl-7f33f1bc-99fa-4906-889f-a42375beea20",
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"mem-ce5a2ffc-ad39-4728-9ac6-76fef507d5da",
"project-Stripe_Elements__not_hosted_checkout__Custom_URL__DAG_bundle_pricing__Stripe_Connect_80_20_",
"tag-provenance",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-e495c8c5-ad95-4b64-8771-f68aa4cfcd0a"
],
"n_returned": 10,
"latency_ms": 4674.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"kn-79056192-7de8-486c-9565-f128439a6fcb",
"bl-f6236350-f7b8-4f4f-a702-9eef2eb76e4b",
"mem-ef0091d8-1b65-431e-afa8-c6c4ee5779c9",
"mem-37b57f52-a29a-42cf-a07a-3c5f8a3598dd",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"tag-__cultivation-metric____internal-state____dharma____evidence____novel-idea____gap-compression____values____microsoft__",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 3974.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"bl-57c5cf6b-81a5-4558-9902-5c02981fe273",
"tag-guilds",
"tag-finance",
"tag-temporal",
"tag-barkhausen",
"tag-ilogger",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 4851.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"project-Goal_setting__alignment__scoring__cadence__Attaches_to_any_imprint_",
"tag-turing-test",
"tag-sealed",
"tag-design-first",
"tag-part-5",
"tag-model",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4505.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-9590ba23-bddb-43e8-a571-68a263c4c364",
"bl-a9e57bb2-00a1-4867-ab59-5d9271134b50",
"tag-performed-values",
"bl-c7793c4a-7630-47fc-a462-d23059087e80",
"tag-gateway_platform_neuron-technologies_go_proxy_llm",
"tag-fornax",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 4494.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"bl-7a13527b-3e0c-418a-9f37-88fd2152e5ce",
"bl-3f57bc69-7285-4f4a-a861-2de52efca058",
"bl-5e390b10-8753-4f25-a1a5-b5dbbb002cbf",
"tag-ats",
"tag-data-model",
"tag-aggregation",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 3619.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"tag-memory-model",
"tag-storage",
"tag-ga4",
"tag-potions",
"tag-resonance",
"tag-neuron",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 4107.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"mem-927f41ab-8ede-4f58-acb3-995db16ac775",
"mem-dbe80bc2-c602-46b0-b4ea-dd222e52bcde",
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"tag-upload-window",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 4231.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"mem-481c769c-68cc-45c7-bc37-c0d9778fa648",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-9d8f3c5b-4bac-41ce-8ac4-44733f99d6c8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 3208.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"project-Imprint__leadership_development__feedback_frameworks__performance__presence_",
"mem-5e7f6ddd-c818-4ad3-b564-54ae278e9976",
"bl-7fa1b1a8-b80a-4f28-b162-bfe73765b4f8",
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c",
"tag-sarah",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 2359.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"mem-265a7107-73b3-4410-9aff-43787d5f473b",
"mem-9d1bf963-1b40-4588-bdb3-0432646cc623",
"bl-0de4e61b-6562-49e5-b7df-ebb809a01723",
"mem-3987d374-3c48-4e8e-b06d-0c363f55ed9c",
"bl-e0a0df72-de6e-46ab-800b-e1e3e8dfc387",
"tag-__patents____swarm____claim-language____prior-art____filing__",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 2342.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-dca14c4c-4859-47b0-996e-33964ba61a87",
"bl-e148d23c-24e8-4122-9915-d1c11f22052f",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"bl-dc8c7e02-eb37-48ae-a6f8-9b512803ae16",
"mem-8d1bafe6-209c-456c-9a25-9a927bc5a16d",
"bl-14883d81-f7cb-46dd-82c2-a6e6980264e5",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?Q??m?;`'",
"R^??m?;?'"
],
"n_returned": 10,
"latency_ms": 3536.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-ffa22d7e-42a9-4bd2-a428-1d2df243ac93",
"bl-452a4710-3d2b-4e0f-9413-49a66423bc9a",
"bl-4a6746e8-191f-48fc-8bfb-c4dc73b80bcd",
"tag-command-pattern",
"tag-withholding",
"tag-offline",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 4151.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 1342.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 874.8,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"kn-333542cb-6dab-4662-9725-bf7440d28bf7",
"? ?}&?#??X\b",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4"
],
"n_returned": 8,
"latency_ms": 1374.6,
"error": null,
"clean": false,
"false_positives": 8
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"ctx-e5427d7d",
"project-worldweaver",
"tag-__kotlin____internal-state____pre-reasoning____post-reasoning____compression-ratio____dharma____cultivation__",
"tag-import",
"tag-enterprise",
"tag-__cgi____dharma____cultivation____five-primitives____seed-artifact____agi____intelligence____whitepaper____patent__",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 3762.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-bfe0fafd-2750-4fdc-b773-04e878b3b23f",
"art-7bdaff30-5af9-4f0a-93b1-751686f9de3d",
"mem-cf07910d-4676-4384-ab97-9cad946cd0b9",
"project-Convert_UTC_timestamps_to_Central_time_when_displaying_to_Will__Never_surface_raw_UTC_",
"mem-32203649-3213-4d6d-86fd-3d657ac70d77",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"? ?}&?#??X\b",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-80ca3d31-84dc-4502-83f4-538372b9764f",
"015644f5-8194-4af0-800d-dd4a0cd71396"
],
"n_returned": 10,
"latency_ms": 4859.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"project-Imprint__discovery__objection_handling__deal_strategy__pipeline__closing_",
"bl-8116da7a-b039-4e08-b8d0-c1c7861f9766",
"bl-8ef1ba6b-3fa0-4dbd-98c5-31665e5694a1",
"tag-dark-theme",
"tag-divisors",
"tag-distressed-property",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"???Ͼd??W\b?"
],
"n_returned": 10,
"latency_ms": 2873.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+956
View File
@@ -0,0 +1,956 @@
{
"label": "hybrid-semantic-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-hybrid",
"soul_md5": "5cd2718932c2ecf4940ddfe5a9c8abbe",
"corpus": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/corpus-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7895,
"wall_clock_s": 52.1,
"child_pid": 86164,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.5142857142857142,
"recall@5": 0.4409013605442177,
"recall@10": 0.5047619047619047,
"precision@5": 0.15428571428571433,
"mrr@10": 0.38746031746031745,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1225.4,
"latency_ms_p95": 1671.6,
"latency_ms_max": 1718.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 307.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"project-engram",
"project-engram-lang",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-89a2",
"bl-13babd0c-582e-4e28-a9e4-a77e65925e5d",
"870ede67-3454-4e00-9988-46cb13a8a4e2",
"bl-3e433255-3710-49fc-a093-c25e71de2ccb",
"mem-235a7657-d49e-467e-9f69-f4c3d5f6bd48"
],
"n_returned": 10,
"latency_ms": 349.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 316.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 315.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 326.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 304.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"????7???Ջ3",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?ǚ?7??????",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 593.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 534.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 544.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"ԍ????X????",
"project-harmonic-framework",
"?ǚ?7??????",
"project-harmonic-framework_com",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"??????X??2c",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"????7???Ջ3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 524.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 543.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 905.8,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"?",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"?",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"?",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"?"
],
"n_returned": 10,
"latency_ms": 769.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1645.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1718.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"? ?}&?#??X\b",
"????7???Ջ3",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e"
],
"n_returned": 10,
"latency_ms": 1393.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"??o?'?B???k",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1123.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-c2205725-69d0-4dd1-9a8d-1c7fa9a0c7b4",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1448.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"??o?'?B???k",
"kn-150e6790-fc2a-48a7-8289-313c1fbaf5ae",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1298.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"? ?}&?#??X\b",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"? ?}&?#??X\b",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 1621.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"? ?}&?#??X\b",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"??????X??2c",
"dR????X?-?S"
],
"n_returned": 10,
"latency_ms": 1405.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"??????X??2c",
"dz????Xƹ?i",
"ԍ????X????",
"dR????X?-?S",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1671.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"????7???Ջ3",
"a1000001-0000-0000-0000-000000000001",
"?ǚ?7??????",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"7?e?7???\f3?",
"mem-ade9440f-f161-4c18-9b35-1976257e6ebb",
"?of?7???",
"ea95f600-8dfd-4c7e-b077-a93dc3cd3623"
],
"n_returned": 10,
"latency_ms": 1519.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1576.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"?ǚ?7??????",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"??S?7???",
"??f?7???",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"??S?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1287.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"7?e?7???\f3?",
"mem-7b74cac0-905f-4c35-9688-fbcce105a177"
],
"n_returned": 10,
"latency_ms": 1391.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1498.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1169.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1239.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-8c2d5f51-3ccd-4c2e-848a-eb60d90a3b98",
"7?e?7???\f3?",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2121fdb9-796a-427e-b9b5-651f4388ea16"
],
"n_returned": 10,
"latency_ms": 1225.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"7?e?7???\f3?",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"rQ??m?;?x?'"
],
"n_returned": 9,
"latency_ms": 1237.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1458.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 756.3,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 519.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?V?",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"?m?\\}Q??6??",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?m?\\}Q??6??",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"?m?\\}Q??6??"
],
"n_returned": 10,
"latency_ms": 774.0,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"? ?}&?#??X\b",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"? ?}&?#??X\b",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"?of?7???",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"? ?}&?#??X\b",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"n_returned": 10,
"latency_ms": 1340.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 10
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"12082f7e-e320-438b-bd65-083d8259748f",
"? ?}&?#??X\b",
"527ecb25-2587-47eb-8269-73be2431abd4",
"? ?}&?#??X\b",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"? ?}&?#??X\b",
"be3b6036-6eca-44a7-8fdf-37b23edfdfd1"
],
"n_returned": 10,
"latency_ms": 1693.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666,
"outranks": true,
"rank_correct": 6,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"7?e?7???\f3?",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"? ?}&?#??X\b",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"%???2??jH??",
"art-8a0870d5-a716-4672-8094-f7463af1265b",
"???Ͼd??W\b?",
"bl-556438af-57b2-4bd8-a747-9f868aaee290"
],
"n_returned": 10,
"latency_ms": 1027.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 4
}
]
}
+956
View File
@@ -0,0 +1,956 @@
{
"label": "hybrid-semantic",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-hybrid",
"soul_md5": "5cd2718932c2ecf4940ddfe5a9c8abbe",
"corpus": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/corpus-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 52.1,
"child_pid": 86109,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.5142857142857142,
"recall@5": 0.4409013605442177,
"recall@10": 0.5047619047619047,
"precision@5": 0.15428571428571433,
"mrr@10": 0.38746031746031745,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1219.7,
"latency_ms_p95": 1667.1,
"latency_ms_max": 1720.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.38461538461538464,
"recall@5": 0.38461538461538464,
"recall@10": 0.38461538461538464,
"mrr@10": 0.17307692307692307
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.6634920634920636
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.2222222222222222,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 323.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"project-engram",
"project-engram-lang",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-89a2",
"bl-13babd0c-582e-4e28-a9e4-a77e65925e5d",
"870ede67-3454-4e00-9988-46cb13a8a4e2",
"bl-3e433255-3710-49fc-a093-c25e71de2ccb",
"mem-235a7657-d49e-467e-9f69-f4c3d5f6bd48"
],
"n_returned": 10,
"latency_ms": 330.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 318.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 316.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 328.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 296.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"????7???Ջ3",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"?ǚ?7??????",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 611.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 543.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 547.8,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"ԍ????X????",
"project-harmonic-framework",
"?ǚ?7??????",
"project-harmonic-framework_com",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"??????X??2c",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"????7???Ջ3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 518.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 534.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 895.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"?",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"?",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"?",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"?"
],
"n_returned": 10,
"latency_ms": 763.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1650.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1720.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"? ?}&?#??X\b",
"????7???Ջ3",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e"
],
"n_returned": 10,
"latency_ms": 1396.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"?of?7???",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1126.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-c2205725-69d0-4dd1-9a8d-1c7fa9a0c7b4",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1452.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"??o?'?B???k",
"kn-150e6790-fc2a-48a7-8289-313c1fbaf5ae",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1303.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"? ?}&?#??X\b",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"? ?}&?#??X\b",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 1618.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"? ?}&?#??X\b",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"??????X??2c",
"dR????X?-?S"
],
"n_returned": 10,
"latency_ms": 1406.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1667.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"????7???Ջ3",
"a1000001-0000-0000-0000-000000000001",
"?ǚ?7??????",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"7?e?7???\f3?",
"mem-ade9440f-f161-4c18-9b35-1976257e6ebb",
"?of?7???",
"ea95f600-8dfd-4c7e-b077-a93dc3cd3623"
],
"n_returned": 10,
"latency_ms": 1528.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1579.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"?ǚ?7??????",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"??S?7???",
"??f?7???",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"??S?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1288.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"7?e?7???\f3?",
"mem-7b74cac0-905f-4c35-9688-fbcce105a177"
],
"n_returned": 10,
"latency_ms": 1383.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1499.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1167.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1245.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-8c2d5f51-3ccd-4c2e-848a-eb60d90a3b98",
"7?e?7???\f3?",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2121fdb9-796a-427e-b9b5-651f4388ea16"
],
"n_returned": 10,
"latency_ms": 1219.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"7?e?7???\f3?",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"rQ??m?;?x?'"
],
"n_returned": 9,
"latency_ms": 1235.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1464.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 761.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 525.1,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?V?",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"?m?\\}Q??6??",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"?m?\\}Q??6??",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"?m?\\}Q??6??"
],
"n_returned": 10,
"latency_ms": 761.7,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"? ?}&?#??X\b",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"? ?}&?#??X\b",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"?of?7???",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"? ?}&?#??X\b",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"n_returned": 10,
"latency_ms": 1322.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 10
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"12082f7e-e320-438b-bd65-083d8259748f",
"? ?}&?#??X\b",
"527ecb25-2587-47eb-8269-73be2431abd4",
"? ?}&?#??X\b",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"? ?}&?#??X\b",
"be3b6036-6eca-44a7-8fdf-37b23edfdfd1"
],
"n_returned": 10,
"latency_ms": 1694.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666,
"outranks": true,
"rank_correct": 6,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"7?e?7???\f3?",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"? ?}&?#??X\b",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"%???2??jH??",
"art-8a0870d5-a716-4672-8094-f7463af1265b",
"???Ͼd??W\b?",
"bl-556438af-57b2-4bd8-a747-9f868aaee290"
],
"n_returned": 10,
"latency_ms": 1041.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 4
}
]
}
File diff suppressed because it is too large Load Diff
+945
View File
@@ -0,0 +1,945 @@
{
"label": "main-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-main",
"soul_md5": "caf822425540114984bfcd713ad37162",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7895,
"wall_clock_s": 45.6,
"child_pid": 78695,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1151.1,
"latency_ms_p95": 1580.7,
"latency_ms_max": 1663.2,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 231.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 266.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 229.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 235.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 254.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 228.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 543.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 466.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 474.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 460.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 477.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 824.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 681.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"????7???Ջ3",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1577.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1663.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1331.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1036.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1363.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1230.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1548.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1315.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1580.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"?ǚ?7??????",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"????7???Ջ3",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"%???2??jH??",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1474.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1490.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1192.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1320.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1448.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1107.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1151.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1154.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1187.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1413.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 684.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 444.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 696.8,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1245.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1617.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 967.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+945
View File
@@ -0,0 +1,945 @@
{
"label": "main-r3",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-main",
"soul_md5": "caf822425540114984bfcd713ad37162",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7896,
"wall_clock_s": 45.7,
"child_pid": 78750,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1141.7,
"latency_ms_p95": 1578.0,
"latency_ms_max": 1645.5,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 237.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 272.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 231.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 231.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 262.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 239.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 542.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 465.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 470.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 441.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 455.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 816.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 678.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"????7???Ջ3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1564.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1645.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1320.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1029.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1347.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1220.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1560.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1305.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1578.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"????7???Ջ3",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"?ǚ?7??????",
"7?e?7???\f3?",
"??f?7???",
"? ?}&?#??X\b",
"%???2??jH??",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1"
],
"n_returned": 10,
"latency_ms": 1449.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1489.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1186.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1309.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1427.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1081.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1144.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1141.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1171.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1388.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 659.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 432.5,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 673.7,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1254.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1626.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 950.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
+945
View File
@@ -0,0 +1,945 @@
{
"label": "main-r1",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-main",
"soul_md5": "caf822425540114984bfcd713ad37162",
"corpus": "/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7893,
"wall_clock_s": 45.9,
"child_pid": 78554,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.34285714285714286,
"recall@5": 0.26947278911564626,
"recall@10": 0.3333333333333333,
"precision@5": 0.12000000000000001,
"mrr@10": 0.2943197278911564,
"nonsense_clean": "2/3",
"superseded_outranks": "1/3",
"latency_ms_p50": 1140.4,
"latency_ms_p95": 1584.1,
"latency_ms_max": 1627.6,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"mrr@10": 0.0
},
"phrase": {
"n": 7,
"hit@5": 0.8571428571428571,
"recall@5": 0.4902210884353741,
"recall@10": 0.6666666666666666,
"mrr@10": 0.5965986394557822
},
"superseded": {
"n": 3,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.3333333333333333,
"mrr@10": 0.041666666666666664,
"outranks": 1
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 240.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5"
],
"n_returned": 1,
"latency_ms": 271.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 232.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 227.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848"
],
"n_returned": 1,
"latency_ms": 259.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 237.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"?ǚ?7??????",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"art-d24fd6dd-2cda-4eed-92f3-67b535a0d71b",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 543.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"Z[?<S???H??",
"rQ??m?;?x?'",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 467.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3125,
"recall@10": 0.5,
"precision@5": 1.0,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"art-79042b8b-6192-440f-90b0-60708f7e6325"
],
"n_returned": 10,
"latency_ms": 475.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.5555555555555556,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"ԍ????X????",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"?ǚ?7??????",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"????7???Ջ3",
"??????X??2c",
"g?e?7???'c?"
],
"n_returned": 10,
"latency_ms": 449.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.14285714285714285
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"mem-a9a9ce95-0d64-46eb-9db8-ff81d78ade35",
"mem-b8ecd23e-77ce-42f7-984c-f51453fec16d"
],
"n_returned": 10,
"latency_ms": 463.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.5,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 821.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.5,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"7?e?7???\f3?",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 688.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"????7???Ջ3",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c"
],
"n_returned": 10,
"latency_ms": 1559.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"? ?}&?#??X\b",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1627.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"?of?7???",
"????7???Ջ3",
"kn-d97920d0-1649-4223-9508-c0bb621e7fc0",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-81c24d13-a73b-4767-819c-dafaacc1498e",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1317.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??o?'?B???k",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1055.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"?of?7???",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-83bb86c6-521d-416c-a86e-6e29c2d8f102",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1352.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"??o?'?B???k",
"? ?}&?#??X\b",
"kn-0625e393-067c-4bba-8389-7e1b79265142",
"ע?RGk?\tH(?"
],
"n_returned": 10,
"latency_ms": 1213.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"?of?7???",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 1545.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"??????X??2c",
"ԍ????X????",
"dR????X?-?S",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"dz????Xƹ?i",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1327.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"dR????X?-?S",
"dz????Xƹ?i",
"ԍ????X????",
"??????X??2c",
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491"
],
"n_returned": 10,
"latency_ms": 1584.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"????7???Ջ3",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"?ǚ?7??????",
"7?e?7???\f3?",
"??f?7???",
"? ?}&?#??X\b",
"%???2??jH??",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1"
],
"n_returned": 10,
"latency_ms": 1432.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"'?T?a\"B~-?8"
],
"n_returned": 10,
"latency_ms": 1505.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"015644f5-8194-4af0-800d-dd4a0cd71396",
"?ǚ?7??????",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"??S?7???",
"? ?}&?#??X\b",
"??f?7???",
"??S?7???",
"knw-5578cb21-e899-4822-b7f4-0d96fa094e3d"
],
"n_returned": 10,
"latency_ms": 1201.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"7?e?7???\f3?",
"kn-e8423822-eacf-4029-aa7b-10d4d28d621e",
"? ?}&?#??X\b",
"??o?'?B???k",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1303.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?ǚ?7??????"
],
"n_returned": 10,
"latency_ms": 1416.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 1090.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"7?e?7???\f3?",
"?of?7???",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee"
],
"n_returned": 10,
"latency_ms": 1167.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"7?e?7???\f3?",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"? ?}&?#??X\b",
"h??I?cB?Q??",
"?of?7???",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1140.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"7?e?7???\f3?",
"rQ??m?;?x?'",
"?Q??m?;?u?'",
"R^??m?;?'",
"?Q??m?;`'",
"? ?}&?#??X\b"
],
"n_returned": 9,
"latency_ms": 1157.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef"
],
"n_returned": 10,
"latency_ms": 1379.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 677.2,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 443.2,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-66a21179-2adc-4b19-a109-880cf4674d7d",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 679.8,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"?of?7???",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"art-80ca3d31-84dc-4502-83f4-538372b9764f"
],
"n_returned": 10,
"latency_ms": 1245.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"? ?}&?#??X\b",
"7?e?7???\f3?",
"?of?7???",
"?ǚ?7??????",
"[?MO5????G",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1618.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"7?e?7???\f3?",
"? ?}&?#??X\b",
"%???2??jH??",
"???Ͼd??W\b?",
"??o?'?B???k",
"? ?}&?#??X\b",
"63307ac5-cf6b-46e0-8296-07503b461cfa",
"? ?}&?#??X\b",
"?of?7???"
],
"n_returned": 10,
"latency_ms": 956.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": null
}
]
}
@@ -0,0 +1,957 @@
{
"label": "semseed-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-semseed",
"soul_md5": "ab0c82781215f43d4907bd6f39d3615b",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-semseed/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 52.0,
"child_pid": 88833,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6857142857142857,
"recall@5": 0.5213459159887731,
"recall@10": 0.6027048348476919,
"precision@5": 0.18285714285714294,
"mrr@10": 0.4608730158730158,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1245.9,
"latency_ms_p95": 1668.7,
"latency_ms_max": 1713.3,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657344,
"recall@10": 0.24825174825174823,
"mrr@10": 0.20833333333333334
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 0.7142857142857143,
"recall@5": 0.40093537414965985,
"recall@10": 0.5150226757369615,
"mrr@10": 0.6507936507936508
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 306.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"bl-ba764d70-e9d7-4f62-848f-719cb665f45e",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"bl-b28d7256-6f74-4567-bd90-40d0ef2a6d78",
"project-engram",
"ctx-45bc",
"project-engram-lang",
"ctx-175f",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-74ed"
],
"n_returned": 10,
"latency_ms": 349.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 294.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 319.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 339.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 318.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"????7???Ջ3",
"kn-363f4976-6946-4b4d-b51b-8a2b0f5aef25",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"ctx-63e3",
"?ǚ?7??????",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 10,
"latency_ms": 604.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 534.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"n_returned": 10,
"latency_ms": 541.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.3333333333333333,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"ԍ????X????",
"project-harmonic-framework",
"?ǚ?7??????",
"project-harmonic-framework_com",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"??????X??2c",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"????7???Ջ3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 515.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 532.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-6f99e111-7055-4635-9831-a489747ce418",
"bl-967536a0-d49d-44fb-8cfb-b31b40bcbfae",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-8b58d9bc-352b-4842-a7f8-a6254b5d1e25",
"bl-39cec462-c80c-4970-a3aa-91fe83053bde"
],
"n_returned": 10,
"latency_ms": 887.8,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.2857142857142857,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"?",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"?",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"?",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"?"
],
"n_returned": 10,
"latency_ms": 762.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"ctx-4a41"
],
"n_returned": 10,
"latency_ms": 1659.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1713.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-8fd9836c-cc39-49df-8d61-babda626cc88",
"?of?7???",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"? ?}&?#??X\b",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"????7???Ջ3"
],
"n_returned": 10,
"latency_ms": 1402.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"7?e?7???\f3?",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"??o?'?B???k",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"? ?}&?#??X\b",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329"
],
"n_returned": 10,
"latency_ms": 1129.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"?of?7???",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"ctx-cc7f",
"ctx-4a41",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 9,
"latency_ms": 1452.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 1293.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"? ?}&?#??X\b",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1622.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"? ?}&?#??X\b",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"bl-8848929a-a23a-46bc-a2c7-fe3a3bc1cddf",
"? ?}&?#??X\b",
"bl-e93858c4-7cac-4b1a-bb62-490790d4c3f3",
"? ?}&?#??X\b",
"bl-286b562a-5299-40e0-a32a-afa9cbdfe995"
],
"n_returned": 10,
"latency_ms": 1412.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"?of?7???",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8"
],
"n_returned": 10,
"latency_ms": 1668.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"????7???Ջ3",
"a1000001-0000-0000-0000-000000000001",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"?ǚ?7??????",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 1535.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"knw-f9ce17a7-17fc-431f-8f23-695b670ec4fa",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 1565.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"?ǚ?7??????",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"7?e?7???\f3?",
"ctx-bb74",
"??S?7???",
"knw-f671966c-3387-4848-abca-b5deec122e00"
],
"n_returned": 10,
"latency_ms": 1279.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1408.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"? ?}&?#??X\b",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 1538.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.3333333333333333
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1168.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40"
],
"n_returned": 10,
"latency_ms": 1290.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.36363636363636365,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"ctx-e5427d7d",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"ctx-bb74",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-8c2d5f51-3ccd-4c2e-848a-eb60d90a3b98"
],
"n_returned": 10,
"latency_ms": 1245.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 9,
"latency_ms": 1277.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.09090909090909091,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1486.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.2727272727272727,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 757.4,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 510.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?V?",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"?m?\\}Q??6??",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 764.0,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"? ?}&?#??X\b",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"? ?}&?#??X\b",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"?of?7???",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"? ?}&?#??X\b",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"n_returned": 10,
"latency_ms": 1340.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 10
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"12082f7e-e320-438b-bd65-083d8259748f",
"6de314bf-5c4c-4cfc-871f-fa2e422d45e6",
"? ?}&?#??X\b",
"527ecb25-2587-47eb-8269-73be2431abd4",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1701.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"bl-c8c19362-430b-4817-9cf4-9e85e0099c64",
"7?e?7???\f3?",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"ctx-3a55",
"? ?}&?#??X\b",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"bl-4f7b651b-6b33-449c-8a3b-cfce12ce984b",
"%???2??jH??"
],
"n_returned": 10,
"latency_ms": 1029.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 5
}
]
}
+957
View File
@@ -0,0 +1,957 @@
{
"label": "semseed",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-semseed",
"soul_md5": "ab0c82781215f43d4907bd6f39d3615b",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-semseed/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7893,
"wall_clock_s": 52.7,
"child_pid": 88601,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.6857142857142857,
"recall@5": 0.5213459159887731,
"recall@10": 0.6027048348476919,
"precision@5": 0.18285714285714294,
"mrr@10": 0.4608730158730158,
"nonsense_clean": "2/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 1227.1,
"latency_ms_p95": 1692.6,
"latency_ms_max": 1710.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.07342657342657344,
"recall@10": 0.24825174825174823,
"mrr@10": 0.20833333333333334
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 2,
"avg_false_positives": 3.3333333333333335
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 0.7142857142857143,
"recall@5": 0.40093537414965985,
"recall@10": 0.5150226757369615,
"mrr@10": 0.6507936507936508
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 295.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"bl-ba764d70-e9d7-4f62-848f-719cb665f45e",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"bl-b28d7256-6f74-4567-bd90-40d0ef2a6d78",
"project-engram",
"ctx-45bc",
"project-engram-lang",
"ctx-175f",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-74ed"
],
"n_returned": 10,
"latency_ms": 341.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 313.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 310.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 318.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 310.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"????7???Ջ3",
"kn-363f4976-6946-4b4d-b51b-8a2b0f5aef25",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"ctx-63e3",
"?ǚ?7??????",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 10,
"latency_ms": 610.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 537.0,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"n_returned": 10,
"latency_ms": 542.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.3333333333333333,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"ԍ????X????",
"project-harmonic-framework",
"?ǚ?7??????",
"project-harmonic-framework_com",
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"??????X??2c",
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"????7???Ջ3",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356"
],
"n_returned": 10,
"latency_ms": 518.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.1111111111111111,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 531.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.5,
"precision@5": 0.0,
"mrr@10": 0.1111111111111111
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-6f99e111-7055-4635-9831-a489747ce418",
"bl-967536a0-d49d-44fb-8cfb-b31b40bcbfae",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-8b58d9bc-352b-4842-a7f8-a6254b5d1e25",
"bl-39cec462-c80c-4970-a3aa-91fe83053bde"
],
"n_returned": 10,
"latency_ms": 888.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.2857142857142857,
"recall@10": 0.2857142857142857,
"precision@5": 0.8,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"?",
"mem-dba009a2-d2ea-4f5a-b9e8-0f04bc9ab32f",
"?",
"mem-7cd90611-88a3-423d-a38a-0db2812952fa",
"?",
"bl-07375bf9-a169-42cd-adb3-7d32b25982f0",
"?"
],
"n_returned": 10,
"latency_ms": 762.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.3333333333333333
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"ctx-4a41"
],
"n_returned": 10,
"latency_ms": 1652.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-2f29ad36-6ee6-4a0e-8d72-0eaf7d12d3a9",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"art-2fabd873-d787-49cb-ad30-d4ed9fcff8ef",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-a31e1001-342e-4deb-a2e6-6d02d1f22dee",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1"
],
"n_returned": 10,
"latency_ms": 1706.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"7?e?7???\f3?",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"knw-8fd9836c-cc39-49df-8d61-babda626cc88",
"?of?7???",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"? ?}&?#??X\b",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"????7???Ջ3"
],
"n_returned": 10,
"latency_ms": 1419.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"kn-d7c1e0fb-fa59-46d3-b4c9-a0d1d437a491",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"7?e?7???\f3?",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"??o?'?B???k",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"? ?}&?#??X\b",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"? ?}&?#??X\b",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329"
],
"n_returned": 10,
"latency_ms": 1167.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"7?e?7???\f3?",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"?of?7???",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"imp-dce1da0f-8776-4a9e-972b-33411a7ca138",
"ctx-cc7f",
"ctx-4a41",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 9,
"latency_ms": 1489.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"knw-e24d6339-5ff3-4bed-ba53-707ffd0dc70a",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-57b4c5e7-40c6-4c90-bf14-71841b0081d4",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 1330.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"? ?}&?#??X\b",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"? ?}&?#??X\b",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1671.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"? ?}&?#??X\b",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"? ?}&?#??X\b",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"bl-8848929a-a23a-46bc-a2c7-fe3a3bc1cddf",
"? ?}&?#??X\b",
"bl-e93858c4-7cac-4b1a-bb62-490790d4c3f3",
"? ?}&?#??X\b",
"bl-286b562a-5299-40e0-a32a-afa9cbdfe995"
],
"n_returned": 10,
"latency_ms": 1449.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"7?e?7???\f3?",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-f8974b26-78a6-4aad-b893-19a73b20013d",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-9397c74b-35f3-4428-b4b0-5123353bbcd1",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"?of?7???",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8"
],
"n_returned": 10,
"latency_ms": 1710.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"????7???Ջ3",
"a1000001-0000-0000-0000-000000000001",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"?ǚ?7??????",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"7?e?7???\f3?"
],
"n_returned": 10,
"latency_ms": 1565.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"mem-bbb126a1-b297-42bb-86be-796871829c94",
"knw-f9ce17a7-17fc-431f-8f23-695b670ec4fa",
"mem-45022957-2d78-48aa-a714-16d6eca52e0f",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 1629.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"?ǚ?7??????",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"7?e?7???\f3?",
"ctx-bb74",
"??S?7???",
"knw-f671966c-3387-4848-abca-b5deec122e00"
],
"n_returned": 10,
"latency_ms": 1326.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"? ?}&?#??X\b",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"art-ee615cdb-e599-423d-9a4d-977859390ed3"
],
"n_returned": 10,
"latency_ms": 1382.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"? ?}&?#??X\b",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 1530.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.3333333333333333
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-e0bdf5d8-d163-491f-b649-453fee8b721d",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1159.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"? ?}&?#??X\b",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"? ?}&?#??X\b",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40"
],
"n_returned": 10,
"latency_ms": 1232.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.36363636363636365,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"art-ee615cdb-e599-423d-9a4d-977859390ed3",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"ctx-e5427d7d",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"ctx-bb74",
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-8c2d5f51-3ccd-4c2e-848a-eb60d90a3b98"
],
"n_returned": 10,
"latency_ms": 1227.1,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"knw-9e74ee95-ba7d-49b1-9262-977eae9729d1",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b"
],
"n_returned": 9,
"latency_ms": 1240.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.09090909090909091,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"? ?}&?#??X\b",
"4f698ae6-c40e-464e-9798-50350991a188",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"? ?}&?#??X\b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"? ?}&?#??X\b",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"? ?}&?#??X\b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 1464.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.2727272727272727,
"precision@5": 0.0,
"mrr@10": 0.16666666666666666
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 756.0,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 517.1,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [
"art-4a99aa1a-489b-4b43-958b-25217adb1aad",
"?V?",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"?m?\\}Q??6??",
"8cbb60c5-4999-4ec1-8682-2592aedc4249",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"?m?\\}Q??6??",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd"
],
"n_returned": 10,
"latency_ms": 755.8,
"error": null,
"clean": false,
"false_positives": 10
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"7?e?7???\f3?",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"? ?}&?#??X\b",
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"? ?}&?#??X\b",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"?of?7???",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"? ?}&?#??X\b",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff"
],
"n_returned": 10,
"latency_ms": 1329.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 10
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"? ?}&?#??X\b",
"12082f7e-e320-438b-bd65-083d8259748f",
"6de314bf-5c4c-4cfc-871f-fa2e422d45e6",
"? ?}&?#??X\b",
"527ecb25-2587-47eb-8269-73be2431abd4",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"? ?}&?#??X\b",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"? ?}&?#??X\b"
],
"n_returned": 10,
"latency_ms": 1692.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"kn-5584ef9c-7f9d-4d7c-a10a-4ee6bc5cf356",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"bl-c8c19362-430b-4817-9cf4-9e85e0099c64",
"7?e?7???\f3?",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"ctx-3a55",
"? ?}&?#??X\b",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"bl-4f7b651b-6b33-449c-8a3b-cfce12ce984b",
"%???2??jH??"
],
"n_returned": 10,
"latency_ms": 1040.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 5
}
]
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,948 @@
{
"label": "wordstart-r2",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-wordstart",
"soul_md5": "32d4cf77672658a5f49dc7c9213e3ba2",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7895,
"wall_clock_s": 28.7,
"child_pid": 1490,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "3/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 537.6,
"latency_ms_p95": 733.7,
"latency_ms_max": 761.4,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 3,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 165.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"bl-ba764d70-e9d7-4f62-848f-719cb665f45e",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"bl-b28d7256-6f74-4567-bd90-40d0ef2a6d78",
"project-engram",
"ctx-45bc",
"project-engram-lang",
"ctx-175f",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-74ed"
],
"n_returned": 10,
"latency_ms": 164.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 163.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 172.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 167.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 164.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"tag-patterns",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"project-Imprint__system_design__ADRs__tech_strategy__integration_patterns__governance_",
"project-Imprint__analysis_patterns__data_storytelling__SQL__dashboards__insight_framing_",
"bl-79028eed-c330-4724-9402-734062d13503",
"bl-39dad13d-7105-4049-8224-dc3c34fdb1f3",
"bl-4ef4d914-da46-4e0f-be78-5219b9547e9f",
"bl-7e7c3fdb-4132-487f-aa70-b2cd559cb7f0",
"bl-1d32bd54-cf17-4a1f-b235-982d09a36f04",
"bl-b8af6601-a8cb-41b5-aef5-ab8a57432dd5"
],
"n_returned": 10,
"latency_ms": 281.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 254.8,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"n_returned": 10,
"latency_ms": 255.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.4444444444444444,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"project-harmonic-framework",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"project-harmonic-framework_com",
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"tag-harmonic-design",
"bl-92acd4eb-0452-4e8e-9f54-f8cd35170d76",
"tag-harmonic-framework",
"bl-18a9d1e4-1484-474c-bf6b-c6173212181b"
],
"n_returned": 10,
"latency_ms": 247.7,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1111111111111111,
"recall@10": 0.1111111111111111,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-sarah",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-a6cb3b8d-d89c-46fc-931d-e90c560783b0",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 263.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-6f99e111-7055-4635-9831-a489747ce418",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-967536a0-d49d-44fb-8cfb-b31b40bcbfae",
"bl-8b58d9bc-352b-4842-a7f8-a6254b5d1e25",
"?Q??m?;?u?'",
"bl-39cec462-c80c-4970-a3aa-91fe83053bde"
],
"n_returned": 10,
"latency_ms": 397.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.21428571428571427,
"recall@10": 0.2857142857142857,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"?",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"?",
"mem-fb44a2fc-7405-41ff-87b3-84643ac07313",
"?",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"?"
],
"n_returned": 10,
"latency_ms": 341.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"cb070131-dfd4-4a38-91d7-22b1bde164d2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"knw-d788a210-613b-4c49-9486-88bbc9d4716f",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-a535f205-bc4c-4058-9171-6263c496044a",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-0228da71-d7f7-4f3b-b7b3-c5eede42b62a",
"ctx-4a41"
],
"n_returned": 10,
"latency_ms": 761.4,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"b1183213-d659-4759-85d7-5b1f22427fe2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-16efddd1-c43d-4a42-9d78-f54fb82bd277",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"7452eb8f-be01-4b55-aec1-ff0c29e790f6",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"532277bf-2959-4beb-ae0d-b018c97678ee",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"d4015bd7-c592-4ed8-8574-1f15ad37af75"
],
"n_returned": 10,
"latency_ms": 733.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"tag-fiction",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"mem-8d690e9d-a7e9-4062-b2f8-e2064294e463",
"knw-8fd9836c-cc39-49df-8d61-babda626cc88",
"mem-ce793303-c5a5-4586-a232-a3426edd9ec7",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"mem-443bd012-fc9a-4088-b236-de5157a1ef92",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"mem-ca4d6a34-d354-413f-bc86-126cc17ca81c"
],
"n_returned": 10,
"latency_ms": 629.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"08f0d1e2-8d0e-42e3-9f0a-8186ae31ec7e",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"4da5dbaf-46e5-4f3e-b474-f60d9f8241d3",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"bl-79ce4464-5dd6-49bd-9b0c-9803549d0665"
],
"n_returned": 10,
"latency_ms": 529.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-e5cc63c0-8701-49d6-855a-e387fe087771",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"mem-75e490d1-f0a9-4b73-8cfc-8daecfaf6f38",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"a1000001-0000-0000-0000-000000000010",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"bfad516b-c306-4c4c-874a-a347c46c05c2",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"n_returned": 10,
"latency_ms": 747.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"tag-learning",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"345b6420-e004-4d2e-b55c-6a729393fa99",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"mem-92a7fdc5-9dd0-48cf-a691-506058de3838",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"a1000001-0000-0000-0000-000000000010",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 564.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"mem-cde58b77-50d3-4bac-9581-e70a4c02c015",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-8d699e2c-ac2a-4742-bb62-b6da00f4b10e",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"mem-53d6adf0-cd08-4707-a237-daa5e65c7298",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"bl-ef2bac68-e119-4139-b529-c7a1404ae3ac"
],
"n_returned": 10,
"latency_ms": 689.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"bl-0e8f4880-7b24-43aa-aed9-ad4d9fc73ff8",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"bl-8848929a-a23a-46bc-a2c7-fe3a3bc1cddf",
"bl-205141ad-b2a0-4d93-86d0-89eb0723e1bd",
"bl-e93858c4-7cac-4b1a-bb62-490790d4c3f3",
"bl-34f51ddb-a840-459f-a248-94214f5febb6",
"bl-286b562a-5299-40e0-a32a-afa9cbdfe995"
],
"n_returned": 10,
"latency_ms": 588.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"bl-4c5b385e-135a-4663-8521-96af0b491121",
"knw-12b4b913-7a25-4b0d-844c-504c01d6725e",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"knw-9707256e-ed44-4042-bd88-f90fa514e1cf",
"34356a36-0df5-4020-8dcc-5e7a423f8d4c",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"n_returned": 10,
"latency_ms": 729.0,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"a1000001-0000-0000-0000-000000000001",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-1b58b05c-9305-4f06-a586-a08c96008027",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"ctx-4a41",
"mem-5708f4c9-3d61-4182-8543-2843698931e6"
],
"n_returned": 10,
"latency_ms": 638.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"077d064f-3489-4c05-9aca-3782f96b51db",
"knw-f9ce17a7-17fc-431f-8f23-695b670ec4fa",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"27e1b1a4-ad0b-49d9-812f-fedf43b8aabe",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"bl-87c93185-b2bf-40af-ae23-3c830c007abf",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"fce2792a-53fc-4d4a-be3b-42bd6ceb1ba7",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 679.0,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"e4f27651-52c5-43fd-aff3-61d31685b3cd",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"mem-5624ec9d-62ba-4aba-8a3d-6afec6c09dd4",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-a16deccb-16a7-419c-a013-ff824a4daa15",
"a1000001-0000-0000-0000-000000000009",
"mem-833dbbcd-2400-4594-bb35-93b023049ac0",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd"
],
"n_returned": 10,
"latency_ms": 588.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"mem-b99efff0-00e6-40c8-9c5b-730330eef33b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-23c27d3b-e0d2-43a8-a80c-0a44477ae18a",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"tag-childhood",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"n_returned": 10,
"latency_ms": 613.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"efe53612-6914-4936-8e3b-1e694eb174e5",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 652.1,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-9110798f-d0cb-4446-bc2a-14f09b6a09e2",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 543.3,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-trailer-park-paladins",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"project-trailer-park-paladins",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 537.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"bl-bea7473c-c687-414c-9c0b-00c509a616c1",
"bl-fc6fcb0b-9e4b-40bf-8e88-dbfe4e27c31a",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc"
],
"n_returned": 10,
"latency_ms": 538.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-hope",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 533.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.18181818181818182,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"knw-528dbc37-eabc-4b75-a7a5-65bf38d6018a",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-473f3f24-20f6-4f39-8589-3709538eb6ac",
"mem-a0b7cfda-bc9e-4f40-b9a9-1722cf3f8263",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"4f698ae6-c40e-464e-9798-50350991a188",
"719aa819-00a9-4f4b-a857-4f9fe5ad44d7",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 673.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 350.7,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 254.6,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [],
"n_returned": 0,
"latency_ms": 345.8,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____architecture__",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____kotlin____architecture__",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"mem-c17aefb1-38b5-4ced-af50-fe524127e1a4",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5"
],
"n_returned": 10,
"latency_ms": 639.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 8
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"12082f7e-e320-438b-bd65-083d8259748f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"13705072-4515-4124-963d-083af490494f",
"527ecb25-2587-47eb-8269-73be2431abd4",
"6de314bf-5c4c-4cfc-871f-fa2e422d45e6",
"a1000001-0000-0000-0000-000000000002",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"de3b6428-b76c-4c44-90e0-bf1dd6998027"
],
"n_returned": 10,
"latency_ms": 733.7,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"5fcba804-eb5b-48ec-82da-146b1c6bb50d",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"bl-c8c19362-430b-4817-9cf4-9e85e0099c64",
"bl-c5c6571e-118f-47c7-8cbb-3ed0ebf64a51",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"ctx-3a55",
"86228228-7adf-41fb-b4c4-9ceea87953ae",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"bl-4f7b651b-6b33-449c-8a3b-cfce12ce984b",
"mem-3a2cf162-d93b-4f29-86f2-5066fb7fe1f5"
],
"n_returned": 10,
"latency_ms": 507.9,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 5
}
]
}
+948
View File
@@ -0,0 +1,948 @@
{
"label": "wordstart",
"soul_binary": "/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad/soul-wordstart",
"soul_md5": "32d4cf77672658a5f49dc7c9213e3ba2",
"corpus": "/Users/timlingo/neuron-eval-corpora/snapshot-pre-repair-20260806-embedded.json",
"corpus_nodes": 78768,
"corpus_edges": 14214,
"gold_set": "/Users/timlingo/Development/neuron-technologies/_wt-eval/tools/retrieval-eval/gold_set.json",
"limit": 10,
"port": 7894,
"wall_clock_s": 28.9,
"child_pid": 1420,
"child_confirmed_dead": true,
"aggregate": {
"n_queries": 38,
"n_scored": 35,
"hit@5": 0.7428571428571429,
"recall@5": 0.5536485340056769,
"recall@10": 0.6175677497106068,
"precision@5": 0.20000000000000007,
"mrr@10": 0.5021428571428571,
"nonsense_clean": "3/3",
"superseded_outranks": "2/3",
"latency_ms_p50": 542.6,
"latency_ms_p95": 741.3,
"latency_ms_max": 758.8,
"errors": 0,
"by_category": {
"associative": {
"n": 6,
"hit@5": 0.6666666666666666,
"recall@5": 0.08857808857808858,
"recall@10": 0.23310023310023312,
"mrr@10": 0.25
},
"exact_rare": {
"n": 6,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"mrr@10": 1.0
},
"nonsense": {
"n": 3,
"clean": 3,
"avg_false_positives": 0.0
},
"paraphrase": {
"n": 13,
"hit@5": 0.6153846153846154,
"recall@5": 0.6153846153846154,
"recall@10": 0.6153846153846154,
"mrr@10": 0.2846153846153846
},
"phrase": {
"n": 7,
"hit@5": 1.0,
"recall@5": 0.5494614512471656,
"recall@10": 0.6023242630385487,
"mrr@10": 0.8214285714285714
},
"superseded": {
"n": 3,
"hit@5": 0.3333333333333333,
"recall@5": 0.3333333333333333,
"recall@10": 0.6666666666666666,
"mrr@10": 0.20833333333333334,
"outranks": 2
}
}
},
"rows": [
{
"id": "q01",
"category": "exact_rare",
"query": "unjailbreakable",
"returned": [
"mem-7f61beb4-271c-4feb-9f6e-1c9c837a6226"
],
"n_returned": 1,
"latency_ms": 170.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q02",
"category": "exact_rare",
"query": "engram-migrate",
"returned": [
"mem-6fdf6545-5e1a-43a9-8bdc-d2cd248146a5",
"bl-ba764d70-e9d7-4f62-848f-719cb665f45e",
"mem-22fe5ec8-ae0d-4583-a05c-d1ef50353257",
"bl-b28d7256-6f74-4567-bd90-40d0ef2a6d78",
"project-engram",
"ctx-45bc",
"project-engram-lang",
"ctx-175f",
"mem-60778715-758c-4677-933d-fc39b8f94152",
"ctx-74ed"
],
"n_returned": 10,
"latency_ms": 210.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q03",
"category": "exact_rare",
"query": "cartabandonedevent",
"returned": [
"mem-1ba7c67d-85b9-4c2e-9fe2-39f8b0477091"
],
"n_returned": 1,
"latency_ms": 167.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q04",
"category": "exact_rare",
"query": "pre-apprenticeship",
"returned": [
"mem-89c02aae-d3ca-43f9-9e5d-eb369896276c"
],
"n_returned": 1,
"latency_ms": 166.9,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q05",
"category": "exact_rare",
"query": "inferencenodemanager",
"returned": [
"mem-73969486-143f-4431-b5e6-6845d1cc9848",
"bl-c1765767-3e27-449a-8c94-10411d1eb7c0",
"project-Add_inference_url_config_to_Neuron_MCP__Route_summarization_gen_tasks_to_Pantheon__keep_frontier_for_complex_reasoning_"
],
"n_returned": 3,
"latency_ms": 162.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q06",
"category": "exact_rare",
"query": "clear-eyed",
"returned": [
"knw-c72597c5-c23d-4c08-8e9e-996dadf26a99"
],
"n_returned": 1,
"latency_ms": 162.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q07",
"category": "phrase",
"query": "patterns not returns",
"returned": [
"tag-patterns",
"mem-a4a9dfc3-e40b-49b3-b1e1-060e8be2f482",
"project-Imprint__system_design__ADRs__tech_strategy__integration_patterns__governance_",
"project-Imprint__analysis_patterns__data_storytelling__SQL__dashboards__insight_framing_",
"bl-79028eed-c330-4724-9402-734062d13503",
"bl-39dad13d-7105-4049-8224-dc3c34fdb1f3",
"bl-4ef4d914-da46-4e0f-be78-5219b9547e9f",
"bl-7e7c3fdb-4132-487f-aa70-b2cd559cb7f0",
"bl-1d32bd54-cf17-4a1f-b235-982d09a36f04",
"bl-b8af6601-a8cb-41b5-aef5-ab8a57432dd5"
],
"n_returned": 10,
"latency_ms": 306.7,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q08",
"category": "phrase",
"query": "thirty moves",
"returned": [
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6"
],
"n_returned": 10,
"latency_ms": 261.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1875,
"recall@10": 0.375,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q09",
"category": "phrase",
"query": "Grandma Lucas",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"art-0f0277a1-4a8e-4645-95dd-fa379976f31c",
"kn-f230b362-b201-4402-9833-4160c89ab3d4"
],
"n_returned": 10,
"latency_ms": 272.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.3333333333333333,
"recall@10": 0.4444444444444444,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q10",
"category": "phrase",
"query": "Directed Harmonic",
"returned": [
"bl-dcee1887-34c4-4ffa-9119-1e291685ba08",
"project-harmonic-framework",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"mem-7eeacad7-d7c2-4c2b-8348-19a59aa6dbaf",
"project-harmonic-framework_com",
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"tag-harmonic-design",
"bl-92acd4eb-0452-4e8e-9f54-f8cd35170d76",
"tag-harmonic-framework",
"bl-18a9d1e4-1484-474c-bf6b-c6173212181b"
],
"n_returned": 10,
"latency_ms": 265.5,
"error": null,
"hit@5": 1.0,
"recall@5": 0.1111111111111111,
"recall@10": 0.1111111111111111,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q11",
"category": "phrase",
"query": "Sarah Bishop",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-sarah",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"mem-1f32f73a-952c-41bc-96dc-8b8b70d8a7c1",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-a6cb3b8d-d89c-46fc-931d-e90c560783b0",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 271.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.4,
"mrr@10": 1.0
},
{
"id": "q12",
"category": "phrase",
"query": "Directed Autonomous Runtime Modification",
"returned": [
"mem-e6327f52-2bda-4ce7-9471-2fffd1e172de",
"bl-31abf75b-998f-4a4f-a6dd-8204119e0451",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5",
"bl-6f99e111-7055-4635-9831-a489747ce418",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-967536a0-d49d-44fb-8cfb-b31b40bcbfae",
"bl-8b58d9bc-352b-4842-a7f8-a6254b5d1e25",
"?Q??m?;?u?'",
"bl-39cec462-c80c-4970-a3aa-91fe83053bde"
],
"n_returned": 10,
"latency_ms": 412.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.21428571428571427,
"recall@10": 0.2857142857142857,
"precision@5": 0.6,
"mrr@10": 1.0
},
{
"id": "q13",
"category": "phrase",
"query": "zero-knowledge encrypted backup",
"returned": [
"deda48cd-5e1a-46cb-bd43-8016afdb3a8a",
"8f3abb0d-77ed-4af3-9f4d-ba62cd198886",
"7774a16c-1027-4e3b-a21e-67f1f95a4acd",
"?",
"bl-ec84b63d-b278-4944-8d7f-4aa7a51c0315",
"?",
"mem-fb44a2fc-7405-41ff-87b3-84643ac07313",
"?",
"mem-a3c97012-5fa3-4915-a839-2c75c72005e0",
"?"
],
"n_returned": 10,
"latency_ms": 377.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 1.0
},
{
"id": "q14",
"category": "paraphrase",
"query": "the elderly relative who passed while he stayed away",
"returned": [
"cb070131-dfd4-4a38-91d7-22b1bde164d2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"knw-d788a210-613b-4c49-9486-88bbc9d4716f",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-a535f205-bc4c-4058-9171-6263c496044a",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-0228da71-d7f7-4f3b-b7b3-c5eede42b62a",
"ctx-4a41"
],
"n_returned": 10,
"latency_ms": 758.8,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q15",
"category": "paraphrase",
"query": "a soldier sidelined by illness who refused to quit",
"returned": [
"b1183213-d659-4759-85d7-5b1f22427fe2",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-16efddd1-c43d-4a42-9d78-f54fb82bd277",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"7452eb8f-be01-4b55-aec1-ff0c29e790f6",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"532277bf-2959-4beb-ae0d-b018c97678ee",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"d4015bd7-c592-4ed8-8574-1f15ad37af75"
],
"n_returned": 10,
"latency_ms": 741.3,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q16",
"category": "paraphrase",
"query": "choosing an uncomfortable fact over a pleasant fiction",
"returned": [
"tag-fiction",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"mem-8d690e9d-a7e9-4062-b2f8-e2064294e463",
"knw-8fd9836c-cc39-49df-8d61-babda626cc88",
"mem-ce793303-c5a5-4586-a232-a3426edd9ec7",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"mem-443bd012-fc9a-4088-b236-de5157a1ef92",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"mem-ca4d6a34-d354-413f-bc86-126cc17ca81c"
],
"n_returned": 10,
"latency_ms": 642.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q17",
"category": "paraphrase",
"query": "a tight payload beats a bloated one",
"returned": [
"bl-680b24a9-edc3-4a9d-847a-bff0b46b568c",
"bl-798d135f-3987-4ccd-8de6-70ca2f358337",
"08f0d1e2-8d0e-42e3-9f0a-8186ae31ec7e",
"knw-08559f5c-2306-4220-a146-398c74f1643c",
"4da5dbaf-46e5-4f3e-b474-f60d9f8241d3",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"mem-434be7c8-88cb-4039-b79a-1da4ac4de783",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"bl-79ce4464-5dd6-49bd-9b0c-9803549d0665"
],
"n_returned": 10,
"latency_ms": 537.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q18",
"category": "paraphrase",
"query": "if you are able and nobody is coming the job is yours",
"returned": [
"bl-4476e856-c567-4b49-8ff7-d7dca3e5715e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-e5cc63c0-8701-49d6-855a-e387fe087771",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"mem-75e490d1-f0a9-4b73-8cfc-8daecfaf6f38",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"a1000001-0000-0000-0000-000000000010",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"bfad516b-c306-4c4c-874a-a347c46c05c2",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc"
],
"n_returned": 10,
"latency_ms": 725.1,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q19",
"category": "paraphrase",
"query": "learning is the wealth creditors cannot seize",
"returned": [
"tag-learning",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"345b6420-e004-4d2e-b55c-6a729393fa99",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"mem-92a7fdc5-9dd0-48cf-a691-506058de3838",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"a1000001-0000-0000-0000-000000000010",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 562.8,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q20",
"category": "paraphrase",
"query": "reliability proven by track record not assertion",
"returned": [
"mem-cde58b77-50d3-4bac-9581-e70a4c02c015",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-8d699e2c-ac2a-4742-bb62-b6da00f4b10e",
"knw-f671966c-3387-4848-abca-b5deec122e00",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"mem-53d6adf0-cd08-4707-a237-daa5e65c7298",
"a708dd6e-fe73-4f2f-a21e-89daa0985487",
"knw-6b48dce2-f21c-452a-9db5-4e6aa61c87ca",
"bl-ef2bac68-e119-4139-b529-c7a1404ae3ac"
],
"n_returned": 10,
"latency_ms": 686.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q21",
"category": "paraphrase",
"query": "boundaries that enable instead of confine",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"bl-0e8f4880-7b24-43aa-aed9-ad4d9fc73ff8",
"project-Source_kn-6f248a50__Add_containment_rules__convergence__location-independence__failure_modes_",
"bl-8848929a-a23a-46bc-a2c7-fe3a3bc1cddf",
"bl-205141ad-b2a0-4d93-86d0-89eb0723e1bd",
"bl-e93858c4-7cac-4b1a-bb62-490790d4c3f3",
"bl-34f51ddb-a840-459f-a248-94214f5febb6",
"bl-286b562a-5299-40e0-a32a-afa9cbdfe995"
],
"n_returned": 10,
"latency_ms": 584.2,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q22",
"category": "paraphrase",
"query": "what shifts tells you where to cut a system apart",
"returned": [
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c",
"bl-4c5b385e-135a-4663-8521-96af0b491121",
"knw-12b4b913-7a25-4b0d-844c-504c01d6725e",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"knw-9707256e-ed44-4042-bd88-f90fa514e1cf",
"34356a36-0df5-4020-8dcc-5e7a423f8d4c",
"knw-0087493b-25cd-45b0-bf46-c078c5b49718",
"mem-22f5f665-3ad2-4063-88b0-915849a795f5",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"n_returned": 10,
"latency_ms": 719.6,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5
},
{
"id": "q23",
"category": "paraphrase",
"query": "a mind that compounds instead of resetting each day",
"returned": [
"bl-8dd70cac-866d-4ff2-b9fe-b4b3c5f094bb",
"mem-b43f6ef4-2f5a-418d-b5ce-3f21520cf6b8",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-ed33e669-0790-44cb-a036-958d605c6fea",
"a1000001-0000-0000-0000-000000000001",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-1b58b05c-9305-4f06-a586-a08c96008027",
"mem-024598a9-ed2e-4eeb-b1e1-5410856ff132",
"ctx-4a41",
"mem-5708f4c9-3d61-4182-8543-2843698931e6"
],
"n_returned": 10,
"latency_ms": 634.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q24",
"category": "paraphrase",
"query": "loved for the unedited self and not the polished exterior",
"returned": [
"077d064f-3489-4c05-9aca-3782f96b51db",
"knw-f9ce17a7-17fc-431f-8f23-695b670ec4fa",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"27e1b1a4-ad0b-49d9-812f-fedf43b8aabe",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"bl-87c93185-b2bf-40af-ae23-3c830c007abf",
"knw-920c891f-bb8c-48c4-9afc-018ef12dcdc4",
"fce2792a-53fc-4d4a-be3b-42bd6ceb1ba7",
"a1000001-0000-0000-0000-000000000001"
],
"n_returned": 10,
"latency_ms": 674.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q25",
"category": "paraphrase",
"query": "cheerfulness you arrive at instead of assuming",
"returned": [
"mem-82158b02-a180-435d-84f0-0b7ce37511b4",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"e4f27651-52c5-43fd-aff3-61d31685b3cd",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"mem-5624ec9d-62ba-4aba-8a3d-6afec6c09dd4",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"mem-a16deccb-16a7-419c-a013-ff824a4daa15",
"a1000001-0000-0000-0000-000000000009",
"mem-833dbbcd-2400-4594-bb35-93b023049ac0",
"kn-13f60407-7b70-4db1-964f-ea1f8196efbd"
],
"n_returned": 10,
"latency_ms": 571.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q26",
"category": "paraphrase",
"query": "a childhood offering no solid foundation to inherit",
"returned": [
"mem-b99efff0-00e6-40c8-9c5b-730330eef33b",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"knw-2c46cfb4-6d4e-4822-8a1a-7d743c1e4329",
"knw-f6ed7d00-bf7d-42ce-9e40-77cf3406e918",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"knw-23c27d3b-e0d2-43a8-a80c-0a44477ae18a",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-f230b362-b201-4402-9833-4160c89ab3d4",
"tag-childhood",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21"
],
"n_returned": 10,
"latency_ms": 603.4,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.2
},
{
"id": "q27",
"category": "associative",
"query": "Grandma Lucas stroke February 2006 goodbye window",
"returned": [
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"kn-0bb4f021-56de-4947-a35b-a37209e7ba21",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"efe53612-6914-4936-8e3b-1e694eb174e5",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 656.2,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q28",
"category": "associative",
"query": "Marines hernia sepsis medical ward",
"returned": [
"bl-33ecccc2-e37f-43db-91b3-c2a86f08aaac",
"54608b69-78b6-4239-b60f-b8206cfecacc",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"mem-9110798f-d0cb-4446-bc2a-14f09b6a09e2",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"art-79042b8b-6192-440f-90b0-60708f7e6325",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"art-ddfcd045-2c3b-4a1e-9966-fec5ce44e1dd",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71"
],
"n_returned": 10,
"latency_ms": 536.9,
"error": null,
"hit@5": 1.0,
"recall@5": 0.07692307692307693,
"recall@10": 0.3076923076923077,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q29",
"category": "associative",
"query": "Sarah Bishop Dyer trailer performance",
"returned": [
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-a99cefe3-5e83-4050-98d8-6c69f57c7c71",
"766de879-f9d0-4a07-b6df-b43ee13763d8",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"tag-trailer-park-paladins",
"kn-5de5a9ac-fd15-45ab-bf18-77566781cf40",
"project-trailer-park-paladins",
"kn-78db5396-3dbc-4481-bfc7-e4e1422feb1c"
],
"n_returned": 10,
"latency_ms": 562.4,
"error": null,
"hit@5": 1.0,
"recall@5": 0.18181818181818182,
"recall@10": 0.45454545454545453,
"precision@5": 0.4,
"mrr@10": 0.5
},
{
"id": "q30",
"category": "associative",
"query": "Swarm Architecture containment lateral worker",
"returned": [
"bl-0fac287f-f4c0-4f15-bc4d-ff7f8a7af3ae",
"bl-9bde67c1-f0ba-4c3a-8fe5-de0deee0ce43",
"bl-2515d870-e35e-443b-ba20-5150bbc73fed",
"kn-a5b3d0ac-f6a1-49a4-aebb-b8b4cd67fe83",
"kn-b36902cc-0b05-44ba-9aa7-800e5dea9ca9",
"bl-bea7473c-c687-414c-9c0b-00c509a616c1",
"bl-fc6fcb0b-9e4b-40bf-8e88-dbfe4e27c31a",
"kn-6f248a50-355b-47bb-aec8-e0e646a9b077",
"kn-6061318f-046b-4935-907d-8eafdce14930",
"mem-ab34c2f7-3243-424b-affa-25555f6cf9cc"
],
"n_returned": 10,
"latency_ms": 556.6,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q31",
"category": "associative",
"query": "hope won inside the narrative preface",
"returned": [
"kn-e0423482-cfa5-4796-8689-8495c93b66bc",
"knw-e94982a2-358d-4f2f-af31-8ee0fcec07c6",
"bl-2b00aeb0-c0fa-4a9f-8f30-4207e98b3d52",
"kn-58874a74-b96f-4883-9e08-45707f4bd3ee",
"kn-5b606390-a52d-4ca2-8e0e-eba141d13440",
"knw-4aebd815-4eaf-49d7-954b-03595f3d48be",
"knw-723551f5-1950-42a3-8b89-b6a06913cef0",
"tag-hope",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"kn-db9f141b-dbe3-4037-92e0-4bb9be0e5e6e"
],
"n_returned": 10,
"latency_ms": 542.6,
"error": null,
"hit@5": 1.0,
"recall@5": 0.09090909090909091,
"recall@10": 0.18181818181818182,
"precision@5": 0.2,
"mrr@10": 0.25
},
{
"id": "q32",
"category": "associative",
"query": "man of the house six years old expectation",
"returned": [
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"knw-528dbc37-eabc-4b75-a7a5-65bf38d6018a",
"knw-d357b6bb-ad8a-4791-b516-426aea45fa5b",
"knw-473f3f24-20f6-4f39-8589-3709538eb6ac",
"mem-a0b7cfda-bc9e-4f40-b9a9-1722cf3f8263",
"knw-7902acca-604e-409b-8faf-ad85424211d0",
"knw-35940684-abc4-42f0-b942-818f66b1f69a",
"4f698ae6-c40e-464e-9798-50350991a188",
"719aa819-00a9-4f4b-a857-4f9fe5ad44d7",
"kn-6061318f-046b-4935-907d-8eafdce14930"
],
"n_returned": 10,
"latency_ms": 683.5,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0
},
{
"id": "q33",
"category": "nonsense",
"query": "zqxjvw plimforth grebulon",
"returned": [],
"n_returned": 0,
"latency_ms": 350.4,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q34",
"category": "nonsense",
"query": "flarnbistle quommetry",
"returned": [],
"n_returned": 0,
"latency_ms": 258.4,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q35",
"category": "nonsense",
"query": "xxqzzt vurblenacht throom",
"returned": [],
"n_returned": 0,
"latency_ms": 348.4,
"error": null,
"clean": true,
"false_positives": 0
},
{
"id": "q36",
"category": "superseded",
"query": "is the self-improvement architecture called DARMA or DHARMA",
"returned": [
"mem-f3b37427-b7d1-4f7e-b32c-0241a20ce8da",
"mem-80d7416b-20e9-48a0-b176-b215527e2f56",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____architecture__",
"kn-b7e98d63-8b83-4911-b4d0-990602a7f575",
"tag-__darma____cgi____patents____self-improvement____character-preservation____autonomous____kotlin____architecture__",
"knw-e047bb42-dc5b-4383-9e88-e508dc03abe3",
"mem-c17aefb1-38b5-4ced-af50-fe524127e1a4",
"bl-5b17bd3b-0c41-46cb-a710-6fa4429692ff",
"bl-145a0985-2382-400f-a7c5-c335c5e30a72",
"mem-82b93b21-a865-410f-9ec1-fc54121d9bb5"
],
"n_returned": 10,
"latency_ms": 664.5,
"error": null,
"hit@5": 1.0,
"recall@5": 1.0,
"recall@10": 1.0,
"precision@5": 0.2,
"mrr@10": 0.5,
"outranks": true,
"rank_correct": 2,
"rank_stale": 8
},
{
"id": "q37",
"category": "superseded",
"query": "how many provisional patents does Will actually have",
"returned": [
"mem-6f0b2b45-90c1-4356-ac01-3daac05b09c8",
"12082f7e-e320-438b-bd65-083d8259748f",
"015644f5-8194-4af0-800d-dd4a0cd71396",
"13705072-4515-4124-963d-083af490494f",
"527ecb25-2587-47eb-8269-73be2431abd4",
"6de314bf-5c4c-4cfc-871f-fa2e422d45e6",
"a1000001-0000-0000-0000-000000000002",
"3cf706a1-3825-45d8-b0a9-06cae6cdf5b8",
"kn-eb1b9e18-3dc6-4b9b-9cc6-86e0ae6b6be8",
"de3b6428-b76c-4c44-90e0-bf1dd6998027"
],
"n_returned": 10,
"latency_ms": 747.3,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 1.0,
"precision@5": 0.0,
"mrr@10": 0.125,
"outranks": true,
"rank_correct": 8,
"rank_stale": null
},
{
"id": "q38",
"category": "superseded",
"query": "is MCP still the live integration layer",
"returned": [
"5fcba804-eb5b-48ec-82da-146b1c6bb50d",
"bl-7328cbe3-0200-43c2-88e7-0a164e15fca4",
"bl-c8c19362-430b-4817-9cf4-9e85e0099c64",
"bl-c5c6571e-118f-47c7-8cbb-3ed0ebf64a51",
"mem-101e81b4-8097-4749-8d8d-7bb66de34517",
"ctx-3a55",
"86228228-7adf-41fb-b4c4-9ceea87953ae",
"4509ed62-9fb2-48b8-9038-ac569fca9604",
"bl-4f7b651b-6b33-449c-8a3b-cfce12ce984b",
"mem-3a2cf162-d93b-4f29-86f2-5066fb7fe1f5"
],
"n_returned": 10,
"latency_ms": 504.2,
"error": null,
"hit@5": 0.0,
"recall@5": 0.0,
"recall@10": 0.0,
"precision@5": 0.0,
"mrr@10": 0.0,
"outranks": false,
"rank_correct": null,
"rank_stale": 5
}
]
}
File diff suppressed because it is too large Load Diff
+98
View File
@@ -0,0 +1,98 @@
#!/usr/bin/env bash
# run_comparison.sh — the whole harness, end to end, from two git refs.
#
# Builds a soul from each ref, boots each on its own throwaway port with its own
# throwaway HOME and its own disposable copy of the corpus, runs the gold set N
# times per ref, and prints the comparison with its noise threshold.
#
# SAFETY: never touches ~/.neuron, /Applications/Neuron*, ~/neuron-dev-stack, or
# any running service. Sources are exported with `git archive` into a scratch
# dir, so no worktree or branch state is mutated either. Ports are checked
# against the live set before anything boots. Every soul this script starts is
# killed and confirmed dead by run_eval.py; the sweep at the end is a backstop.
#
# usage:
# run_comparison.sh [--baseline main] [--candidate feat/recall-through-activation]
# [--repeats 3] [--corpus <snapshot.json>] [--repo <path>]
set -euo pipefail
BASELINE="main"
CANDIDATE="feat/recall-through-activation"
REPEATS=3
CORPUS="$HOME/neuron-memory-backups/snapshot-pre-repair-20260806.json"
REPO="$HOME/Development/neuron"
BASE_PORT=7893
while [ $# -gt 0 ]; do
case "$1" in
--baseline) BASELINE="$2"; shift 2 ;;
--candidate) CANDIDATE="$2"; shift 2 ;;
--repeats) REPEATS="$2"; shift 2 ;;
--corpus) CORPUS="$2"; shift 2 ;;
--repo) REPO="$2"; shift 2 ;;
*) echo "unknown arg: $1" >&2; exit 2 ;;
esac
done
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
WORK="$(mktemp -d "${TMPDIR:-/tmp}/retrieval-eval.XXXXXX")"
trap 'rm -rf "$WORK"' EXIT
[ -f "$CORPUS" ] || { echo "no corpus at $CORPUS" >&2; exit 2; }
echo "corpus: $CORPUS ($(du -h "$CORPUS" | cut -f1))"
slug() { printf '%s' "$1" | tr '/' '-'; }
build_ref() { # ref -> binary path
local ref="$1" out="$WORK/soul-$(slug "$1")"
local src="$WORK/src-$(slug "$1")"
mkdir -p "$src"
git -C "$REPO" archive "$ref" | tar -x -C "$src"
"$HERE/build-soul.sh" "$src" "$out" >&2
printf '%s' "$out"
}
echo "== building $BASELINE =="
BIN_A="$(build_ref "$BASELINE")"
echo "== building $CANDIDATE =="
BIN_B="$(build_ref "$CANDIDATE")"
echo "== validating the gold set against this corpus =="
python3 "$HERE/build_gold_set.py" "$CORPUS" --check
port=$BASE_PORT
run_one() { # binary label out
echo "== $2 =="
python3 "$HERE/run_eval.py" --soul "$1" --corpus "$CORPUS" --label "$2" \
--port "$port" --out "$3"
port=$((port + 1))
}
A_MAIN="$WORK/results-a-1.json"; B_MAIN="$WORK/results-b-1.json"
A_REP=(); B_REP=()
for i in $(seq 1 "$REPEATS"); do
a="$WORK/results-a-$i.json"; b="$WORK/results-b-$i.json"
run_one "$BIN_A" "$(slug "$BASELINE")-r$i" "$a"
run_one "$BIN_B" "$(slug "$CANDIDATE")-r$i" "$b"
[ "$i" -gt 1 ] && { A_REP+=("$a"); B_REP+=("$b"); }
done
cp "$A_MAIN" "$HERE/results-$(slug "$BASELINE").json"
cp "$B_MAIN" "$HERE/results-$(slug "$CANDIDATE").json"
echo
python3 "$HERE/compare.py" \
--baseline "$A_MAIN" --candidate "$B_MAIN" \
${A_REP[@]+--repeats-baseline "${A_REP[@]}"} \
${B_REP[@]+--repeats-candidate "${B_REP[@]}"} \
--out "$HERE/comparison-$(slug "$BASELINE")-vs-$(slug "$CANDIDATE").json"
# Backstop: run_eval.py kills and confirms its own child, but a crashed run
# could leak one. Leaving a soul running is how the live engine got squeezed.
STRAY=$(pgrep -f "$WORK/soul-" || true)
if [ -n "$STRAY" ]; then
echo "!! stray eval souls, killing: $STRAY" >&2
kill -9 $STRAY 2>/dev/null || true
fi
pgrep -f "$WORK/soul-" >/dev/null && { echo "!! STILL RUNNING" >&2; exit 5; }
echo "process check: no eval souls running"
+398
View File
@@ -0,0 +1,398 @@
#!/usr/bin/env python3
"""
run_eval.py measure one soul build's retrieval against the gold set.
WHAT IT MEASURES, AND WHY IT BOOTS A REAL SOUL
The point is Will's designed retrieval — spreading activation over the
weighted directed graph with four-factor multiplicative scoring not a
Python re-implementation of it. A re-implementation would measure my
reading of the design; booting the compiled binary measures the design. So
this harness compiles the actual `soul.el` amalgam (build-soul.sh) and asks
it over HTTP, exactly as the MCP wrapper and the app do.
SAFETY read this before changing anything here
* Boots on a THROWAWAY port with a THROWAWAY $HOME and a THROWAWAY COPY of
the corpus. Refuses to use 7770 / 8742 / 7779 / 17779 / 7771.
* ENGRAM_URL / SOUL_ENGRAM_URL are UNSET and SOUL_ISE_URL is pinned to a
dead port. This is not belt-and-braces: the periodic engram sync resolves
its source as env(SOUL_ISE_URL) -> state -> DEFAULT http://localhost:8742,
so leaving it unset makes an "isolated" run silently pull the operator's
LIVE brain. (Learned the hard way on 2026-08-03; see the same note in
scripts/verify-soul-contract.sh.)
* Every process this file starts is tracked and killed in a finally block,
then CONFIRMED dead by pid probe, and the confirmation is written into the
results file. A run that cannot confirm its child is dead exits non-zero.
* Activation is a STATEFUL read by design (patent claim 29: traversal
updates last-activation and increments activation counts). The corpus copy
is therefore per-run and disposable, and every run starts from a byte-
identical copy so two configurations see the same starting graph.
usage:
python3 run_eval.py --soul <binary> --corpus <snapshot.json> --label main \
[--port 7893] [--gold gold_set.json] [--limit 10] [--out results-main.json]
"""
import argparse
import json
import os
import shutil
import signal
import subprocess
import sys
import tempfile
import time
import urllib.error
import urllib.parse
import urllib.request
HERE = os.path.dirname(os.path.abspath(__file__))
FORBIDDEN_PORTS = {7770, 8742, 7779, 17779, 7771, 8080}
# ─────────────────────────────────────────────────────────────────────────────
# metrics
# ─────────────────────────────────────────────────────────────────────────────
def recall_at_k(returned, relevant, k):
if not relevant:
return None
return len(set(returned[:k]) & set(relevant)) / len(relevant)
def hit_at_k(returned, relevant, k):
if not relevant:
return None
return 1.0 if set(returned[:k]) & set(relevant) else 0.0
def precision_at_k(returned, relevant, k):
"""Fixed denominator k, as in docs/research/graphrag_eval/score.py.
Fixed denominator penalises an empty result and a page of junk equally,
which is what we want: a retriever that returns nothing is not 'precise'.
"""
if not relevant:
return None
return len(set(returned[:k]) & set(relevant)) / k
def mrr(returned, relevant, k):
if not relevant:
return None
rel = set(relevant)
for i, nid in enumerate(returned[:k], start=1):
if nid in rel:
return 1.0 / i
return 0.0
def mean(vals):
vals = [v for v in vals if v is not None]
return sum(vals) / len(vals) if vals else 0.0
def pct(vals):
return f"{100 * mean(vals):.1f}%"
# ─────────────────────────────────────────────────────────────────────────────
# soul lifecycle
# ─────────────────────────────────────────────────────────────────────────────
class Soul:
def __init__(self, binary, corpus, port, verbose=True):
if port in FORBIDDEN_PORTS:
raise SystemExit(f"REFUSING: port {port} is a live service port.")
self.binary = os.path.abspath(binary)
self.corpus = os.path.abspath(corpus)
self.port = port
self.verbose = verbose
self.home = None
self.proc = None
self.pid = None
self.log = None
self.confirmed_dead = None
@property
def base(self):
return f"http://127.0.0.1:{self.port}"
def start(self, boot_timeout=180):
self.home = tempfile.mkdtemp(prefix="retrieval-eval-home.")
snap = os.path.join(self.home, "corpus.json")
t0 = time.time()
shutil.copyfile(self.corpus, snap) # per-run disposable copy, never the source
self.log = os.path.join(self.home, "soul.log")
env = {k: v for k, v in os.environ.items()
if k not in ("ENGRAM_URL", "ENGRAM_API_KEY", "SOUL_ENGRAM_URL",
"ANTHROPIC_API_KEY", "NEURON_LLM_API_KEY", "SOUL_IDENTITY",
"SOUL_API_KEY")}
env.update({
"HOME": self.home,
"NEURON_PORT": str(self.port),
"SOUL_CGI_ID": f"ntn-retrieval-eval-{os.getpid()}",
"SOUL_ENGRAM_PATH": snap,
"NEURON_API_URL": "http://127.0.0.1:9", # dead port
"SOUL_ISE_URL": "http://127.0.0.1:9", # dead port — see SAFETY above
# Park the background loops for an hour so heartbeat/consolidation
# cannot mutate the graph between queries and make runs unrepeatable.
"SOUL_TICK_MS": "3600000",
"SOUL_HEARTBEAT_MS": "3600000",
"SOUL_REFRESH_MS": "3600000",
})
with open(self.log, "wb") as lf:
self.proc = subprocess.Popen([self.binary], env=env, stdout=lf, stderr=lf,
start_new_session=True)
self.pid = self.proc.pid
if self.verbose:
print(f" booted pid={self.pid} port={self.port} home={self.home}")
deadline = time.time() + boot_timeout
while time.time() < deadline:
if self.proc.poll() is not None:
raise RuntimeError(f"soul exited during boot: {self._log_tail()}")
rss = self._rss_kb()
if rss and rss > 6 * 1024 * 1024:
self.stop()
raise RuntimeError(f"soul RSS {rss}KB > 6GB — aborted")
try:
with urllib.request.urlopen(f"{self.base}/health", timeout=2) as r:
if r.status == 200:
if self.verbose:
print(f" healthy in {time.time() - t0:.1f}s, RSS={self._rss_kb()}KB")
return
except Exception:
pass
time.sleep(0.5)
self.stop()
raise RuntimeError(f"soul never healthy on {self.base}: {self._log_tail()}")
def _rss_kb(self):
try:
out = subprocess.run(["ps", "-o", "rss=", "-p", str(self.pid)],
capture_output=True, text=True, timeout=5).stdout.strip()
return int(out) if out else None
except Exception:
return None
def _log_tail(self, n=15):
try:
with open(self.log, encoding="utf-8", errors="replace") as fh:
return "\n".join(fh.read().splitlines()[-n:])
except Exception:
return "(no log)"
def recall(self, query, limit, timeout=60):
url = f"{self.base}/api/neuron/recall?query={urllib.parse.quote(query)}&limit={limit}"
t0 = time.perf_counter()
try:
with urllib.request.urlopen(url, timeout=timeout) as r:
raw = r.read().decode("utf-8", "replace")
ms = (time.perf_counter() - t0) * 1000
except Exception as exc:
return [], (time.perf_counter() - t0) * 1000, f"{type(exc).__name__}: {exc}"
try:
arr = json.loads(raw)
except Exception:
return [], ms, f"unparseable response ({len(raw)}B)"
if not isinstance(arr, list):
return [], ms, f"non-array response: {str(arr)[:120]}"
ids = [x.get("id") for x in arr if isinstance(x, dict) and x.get("id")]
return ids, ms, None
def stop(self):
"""Kill and CONFIRM. A test process that outlives its test is a bug."""
if self.pid is None:
self.confirmed_dead = True
return True
for sig in (signal.SIGTERM, signal.SIGKILL):
try:
os.kill(self.pid, sig)
except ProcessLookupError:
break
except Exception:
pass
for _ in range(20):
try:
os.kill(self.pid, 0)
except ProcessLookupError:
break
time.sleep(0.1)
else:
continue
break
try:
self.proc.wait(timeout=5)
except Exception:
pass
try:
os.kill(self.pid, 0)
self.confirmed_dead = False
except ProcessLookupError:
self.confirmed_dead = True
if self.verbose:
print(f" pid {self.pid}: {'CONFIRMED DEAD' if self.confirmed_dead else 'STILL ALIVE'}")
if self.home and os.path.isdir(self.home):
shutil.rmtree(self.home, ignore_errors=True)
return self.confirmed_dead
# ─────────────────────────────────────────────────────────────────────────────
# eval
# ─────────────────────────────────────────────────────────────────────────────
def evaluate(soul, gold, limit):
rows = []
for q in gold["queries"]:
ids, ms, err = soul.recall(q["query"], limit)
rel = q.get("relevant") or []
row = {
"id": q["id"],
"category": q["category"],
"query": q["query"],
"returned": ids,
"n_returned": len(ids),
"latency_ms": round(ms, 1),
"error": err,
}
if q.get("expect_empty"):
row["clean"] = (len(ids) == 0)
row["false_positives"] = len(ids)
else:
row["hit@5"] = hit_at_k(ids, rel, 5)
row["recall@5"] = recall_at_k(ids, rel, 5)
row["recall@10"] = recall_at_k(ids, rel, 10)
row["precision@5"] = precision_at_k(ids, rel, 5)
row["mrr@10"] = mrr(ids, rel, 10)
if q.get("must_outrank"):
correct, stale = q["must_outrank"]
ic = ids.index(correct) if correct in ids else None
istale = ids.index(stale) if stale in ids else None
# Correct must be present AND above the stale node. A run that
# returns neither is NOT a pass: the corrected fact is what the
# user needed.
row["outranks"] = (ic is not None) and (istale is None or ic < istale)
row["rank_correct"] = None if ic is None else ic + 1
row["rank_stale"] = None if istale is None else istale + 1
rows.append(row)
return rows
def aggregate(rows):
scored = [r for r in rows if "hit@5" in r]
nonsense = [r for r in rows if "clean" in r]
outrank = [r for r in rows if "outranks" in r]
lat = sorted(r["latency_ms"] for r in rows)
agg = {
"n_queries": len(rows),
"n_scored": len(scored),
"hit@5": mean([r["hit@5"] for r in scored]),
"recall@5": mean([r["recall@5"] for r in scored]),
"recall@10": mean([r["recall@10"] for r in scored]),
"precision@5": mean([r["precision@5"] for r in scored]),
"mrr@10": mean([r["mrr@10"] for r in scored]),
"nonsense_clean": f"{sum(1 for r in nonsense if r['clean'])}/{len(nonsense)}",
"superseded_outranks": f"{sum(1 for r in outrank if r['outranks'])}/{len(outrank)}",
"latency_ms_p50": lat[len(lat) // 2] if lat else 0,
"latency_ms_p95": lat[max(0, int(len(lat) * 0.95) - 1)] if lat else 0,
"latency_ms_max": lat[-1] if lat else 0,
"errors": sum(1 for r in rows if r["error"]),
"by_category": {},
}
cats = sorted({r["category"] for r in rows})
for c in cats:
cr = [r for r in rows if r["category"] == c]
if c == "nonsense":
agg["by_category"][c] = {
"n": len(cr),
"clean": sum(1 for r in cr if r["clean"]),
"avg_false_positives": mean([float(r["false_positives"]) for r in cr]),
}
else:
e = {
"n": len(cr),
"hit@5": mean([r.get("hit@5") for r in cr]),
"recall@5": mean([r.get("recall@5") for r in cr]),
"recall@10": mean([r.get("recall@10") for r in cr]),
"mrr@10": mean([r.get("mrr@10") for r in cr]),
}
if c == "superseded":
e["outranks"] = sum(1 for r in cr if r.get("outranks"))
agg["by_category"][c] = e
return agg
def print_table(label, agg):
print(f"\n=== {label} ===")
print(f" queries {agg['n_queries']} ({agg['n_scored']} scored + "
f"{agg['n_queries'] - agg['n_scored']} control) · errors {agg['errors']}")
print(f" {'hit@5':>12} {'recall@5':>10} {'recall@10':>10} {'prec@5':>9} {'MRR@10':>9}")
print(f" {pct([agg['hit@5']]):>12} {pct([agg['recall@5']]):>10} {pct([agg['recall@10']]):>10} "
f"{pct([agg['precision@5']]):>9} {agg['mrr@10']:>9.3f}")
print(f" nonsense clean {agg['nonsense_clean']} · superseded outranks {agg['superseded_outranks']}")
print(f" latency ms p50 {agg['latency_ms_p50']:.0f} · p95 {agg['latency_ms_p95']:.0f} "
f"· max {agg['latency_ms_max']:.0f}")
print(f"\n {'category':14} {'n':>3} {'hit@5':>8} {'recall@5':>9} {'recall@10':>10} {'MRR@10':>8}")
for c, e in agg["by_category"].items():
if c == "nonsense":
print(f" {c:14} {e['n']:>3} {'clean ' + str(e['clean']) + '/' + str(e['n']):>8}"
f"{'':>9} {'':>10} {'avg FP ' + format(e['avg_false_positives'], '.1f'):>8}")
else:
extra = f" outranks {e['outranks']}/{e['n']}" if "outranks" in e else ""
print(f" {c:14} {e['n']:>3} {pct([e['hit@5']]):>8} {pct([e['recall@5']]):>9} "
f"{pct([e['recall@10']]):>10} {e['mrr@10']:>8.3f}{extra}")
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--soul", required=True)
ap.add_argument("--corpus", required=True)
ap.add_argument("--label", required=True)
ap.add_argument("--gold", default=os.path.join(HERE, "gold_set.json"))
ap.add_argument("--port", type=int, default=7893)
ap.add_argument("--limit", type=int, default=10)
ap.add_argument("--out", default=None)
args = ap.parse_args()
with open(args.gold, encoding="utf-8") as fh:
gold = json.load(fh)
print(f"[{args.label}] soul={os.path.basename(args.soul)} "
f"corpus={os.path.basename(args.corpus)} gold={len(gold['queries'])}q limit={args.limit}")
soul = Soul(args.soul, args.corpus, args.port)
rows = []
started = time.time()
try:
soul.start()
rows = evaluate(soul, gold, args.limit)
finally:
dead = soul.stop()
agg = aggregate(rows)
print_table(args.label, agg)
out = args.out or os.path.join(HERE, f"results-{args.label}.json")
doc = {
"label": args.label,
"soul_binary": os.path.abspath(args.soul),
"soul_md5": subprocess.run(["md5", "-q", args.soul], capture_output=True,
text=True).stdout.strip(),
"corpus": os.path.abspath(args.corpus),
"corpus_nodes": gold.get("corpus_nodes"),
"corpus_edges": gold.get("corpus_edges"),
"gold_set": os.path.abspath(args.gold),
"limit": args.limit,
"port": args.port,
"wall_clock_s": round(time.time() - started, 1),
"child_pid": soul.pid,
"child_confirmed_dead": soul.confirmed_dead,
"aggregate": agg,
"rows": rows,
}
with open(out, "w", encoding="utf-8") as fh:
json.dump(doc, fh, indent=1, ensure_ascii=False)
print(f"\nwrote {out}")
if not dead:
print("FATAL: child process could not be confirmed dead", file=sys.stderr)
sys.exit(4)
if __name__ == "__main__":
main()
+142
View File
@@ -0,0 +1,142 @@
import numpy as np, json, urllib.request, collections, math, re, sys, time
SP="/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad"
EV="/Users/timlingo/Development/neuron-technologies/_wt-semseed/tools/retrieval-eval/"
np.seterr(all='ignore')
t0=time.time()
M=np.load(SP+'/emb.npy'); eids=open(SP+'/ids.txt',encoding='utf-8',errors='surrogateescape').read().split('\n')
eidx={k:i for i,k in enumerate(eids)}
d=json.load(open('/Users/timlingo/neuron-memory-backups/snapshot-pre-repair-20260806.json',encoding='utf-8',errors='surrogateescape'))
print("loaded corpus %.1fs"%(time.time()-t0),file=sys.stderr)
STRUCT={"identity","contains","superseded_by","references","embodies","demonstrated_by","canonical-self","depends_on","currently_holds","activates"}
adj=collections.defaultdict(list)
for e in d['edges']:
if e.get('relation') not in STRUCT: continue
w=float(e.get('weight') or 0.0)
adj[e['from_id']].append((e['to_id'],w)); adj[e['to_id']].append((e['from_id'],w))
nodes=d['nodes']
N={n['id']:n for n in nodes}
PRINT=re.compile(r'^[\x20-\x7e]+$')
ids=[]; hay=[]; dl=[]; sal=[]; addressable=[]
for n in nodes:
i=n.get('id') or ''
h=((n.get('content') or '')+'\x00'+(n.get('label') or '')+'\x00'+(n.get('tags') or '')).lower()
ids.append(i); hay.append(h); dl.append(len(h)); sal.append(float(n.get('salience') or 0.0))
addressable.append(bool(PRINT.match(i)))
del d
NN=len(ids); avgdl=sum(dl)/NN
print("nodes=%d avgdl=%.0f addressable=%d %.1fs"%(NN,avgdl,sum(addressable),time.time()-t0),file=sys.stderr)
gold={q['id']:q for q in json.load(open(EV+"gold_set.json"))['queries']}
LEXMAIN={r['id']:r['returned'] for r in json.load(open(EV+"results-main.json"),) ['rows']} if False else {r['id']:r['returned'] for r in json.load(open(EV+"results-main.json",encoding='utf-8',errors='surrogateescape'))['rows']}
CACHE={}
def emb(t):
if t in CACHE: return CACHE[t]
b=json.dumps({"model":"nomic-embed-text","prompt":t}).encode()
r=urllib.request.Request("http://127.0.0.1:11434/api/embeddings",data=b,headers={"Content-Type":"application/json"})
v=np.array(json.load(urllib.request.urlopen(r,timeout=60))["embedding"],dtype=np.float32)
v=v/(np.linalg.norm(v)+1e-9); CACHE[t]=v; return v
K1,B=1.2,0.75
def lexleg(query, mode, guard, lim=10):
toks=[]
for w in query.split():
wl=w.lower()
if wl not in toks: toks.append(wl)
nt=len(toks)
masks=[]; df=[0]*nt
for i in range(NN):
if guard and not addressable[i]: continue
h=hay[i]; m=0; sc=0
for t in range(nt):
if toks[t] in h: m|=(1<<t); sc+=1; df[t]+=1
if sc: masks.append((i,m,sc))
if mode=='tokcount':
masks.sort(key=lambda x:(-x[2], -sal[x[0]]))
return [ids[i] for i,m,sc in masks[:lim]]
idf=[math.log(1.0+(NN-df[t]+0.5)/(df[t]+0.5)) for t in range(nt)]
scored=[]
for i,m,sc in masks:
norm=1.0-B+B*dl[i]/avgdl
s=0.0
for t in range(nt):
if m&(1<<t): s+=idf[t]*(K1+1.0)/(1.0+K1*norm)
scored.append((s,i))
scored.sort(key=lambda x:(-x[0], -sal[x[1]]))
return [ids[i] for s,i in scored[:lim]]
FIRE=0.02; DECAY=0.7; DEPTH=2; SEED_MIN=0.60; ASSOC_MAX=64
def assoc(seeds, s):
act={x:1.0 for x in seeds}; seen={x:2 for x in seeds}
Q=[(x,0) for x in seeds]; h=0
while h<len(Q):
cur,hop=Q[h]; h+=1
if hop>=DEPTH: continue
p=act[cur]
for oid,w in adj.get(cur,()):
n=N.get(oid)
if not n or n.get('node_type') in ('Tag','InternalStateEvent'): continue
na=p*w*DECAY*float(n.get('salience') or 0.0)
if na<FIRE: continue
if oid in seen and na<=act.get(oid,0): continue
act[oid]=na
if oid not in seen: seen[oid]=1
Q.append((oid,hop+1))
out=[]
for k,v in seen.items():
if v!=1 or k not in eidx: continue
c=float(s[eidx[k]])
if c<=0: continue
out.append((c,k))
out.sort(reverse=True)
return [k for c,k in out[:ASSOC_MAX]]
def inter3(L,S,A,lim=10):
out=[]; li=si=ai=0
while len(out)<lim and (li<len(L) or si<len(S) or ai<len(A)):
if li<len(L):
if L[li] not in out: out.append(L[li])
li+=1
if len(out)>=lim: break
if si<len(S):
if S[si] not in out: out.append(S[si])
si+=1
if len(out)>=lim: break
if ai<len(A):
if A[ai] not in out: out.append(A[ai])
ai+=1
return out
def run(mode, guard, use_main_lex=False):
res={}; legs={}
for qid,q in gold.items():
v=emb(q['query']); s=M@v; s[~np.isfinite(s)]=-1
L = LEXMAIN[qid][:10] if use_main_lex else lexleg(q['query'], mode, guard)
ordr=np.argsort(-s)
S=[eids[j] for j in ordr[:10] if s[j]>SEED_MIN]
if guard: S=[x for x in S if PRINT.match(x or '')]
seeds=[x for x in L[:3] if x in N]
seeds=seeds+[eids[j] for j in ordr[:8] if eids[j] in N and eids[j] not in seeds and (not guard or PRINT.match(eids[j] or ''))]
A=assoc(seeds,s) if seeds else []
if guard: A=[x for x in A if PRINT.match(x or '')]
res[qid]=inter3(L,S,A); legs[qid]=(L,S,A)
return res,legs
def score(res,label,verbose=False):
det={}
for qid,q in gold.items():
out=res[qid][:5]
if q['category']=='nonsense': ok=(len(res[qid])==0)
elif q['category']=='superseded':
must=q.get('must_outrank') or {}; ok=False
for good,bad in (must.items() if isinstance(must,dict) else []):
ok = good in res[qid] and (bad not in res[qid] or res[qid].index(good)<res[qid].index(bad))
if not must: ok=any(r in out for r in q['relevant'])
else: ok=any(r in out for r in q['relevant'])
det[qid]=ok
print("%-28s outcome-true=%d/38"%(label,sum(det.values())))
return det
if __name__=="__main__":
base,_=run('tokcount',False,use_main_lex=True); b=score(base,'BASELINE semseed(real lex)')
variants=[('tokcount',False,'replica: tokcount,noguard'),
('tokcount',True ,'A: tokcount + idguard'),
('bm25', False,'B: bm25 only'),
('bm25', True ,'C: bm25 + idguard')]
dets={}
for m,g,lab in variants:
r,_=run(m,g); dets[lab]=score(r,lab)
dd=[q for q in gold if dets[lab][q]!=b[q]]
print(" vs BASELINE moved=%d gains=%s losses=%s"%(len(dd),[q for q in dd if dets[lab][q]],[q for q in dd if not dets[lab][q]]))
+92
View File
@@ -0,0 +1,92 @@
import json,sys,pickle,numpy as np
sys.path.insert(0,'.')
from legs import *
GP='/Users/timlingo/Development/neuron-technologies/_wt-bm25lex/tools/retrieval-eval/'
G=json.load(open(GP+'gold_set.json'))
def wstart(s,tok):
i=s.find(tok)
while i!=-1:
if i==0 or not s[i-1].isalnum(): return True
i=s.find(tok,i+1)
return False
def legs4(query, wordstart=False, unfloor=False):
toks=tokenize(query); lt=[t.lower() for t in toks]
hit_idx=[];hit_mask=[];df=[0]*len(toks)
for i in range(N):
if not OK[i]: continue
s=LOW[i];m=0
for t,tok in enumerate(lt):
if tok in s and (not wordstart or wstart(s,tok)): m|=(1<<t)
if m:
hit_idx.append(i);hit_mask.append(m)
for t in range(len(toks)):
if m>>t&1: df[t]+=1
dl_n=int(OK.sum());avgdl=float(DL[OK].sum()/max(dl_n,1))
idf=[math.log(1.0+((dl_n-d+0.5)/(d+0.5))) for d in df]
L=[]
for j,i in enumerate(hit_idx):
norm=1.0-B+B*(DL[i]/avgdl);w=0.0
for t in range(len(toks)):
if hit_mask[j]>>t&1: w+=idf[t]*(K1+1.0)/(1.0+K1*norm)
L.append((i,w,SAL[i]))
L.sort(key=lambda x:(-x[1],-x[2]))
if not L: return [],[],[]
qv=qemb(query);cos=En@qv;cos=np.where(HAVE&OK,cos,-2.0)
order=np.argsort(-cos)[:600]
Sl=[int(i) for i in order if cos[i]>(0.0 if unfloor else SEED_MIN)]
semseed=[int(i) for i in order[:SEED_K] if cos[i]>0.0]
act={};seen={};qq=[]
for i,_,_ in L[:ASSOC_SEEDS]:
act[i]=1.0;seen[i]=2;qq.append((i,0))
for i in semseed:
if i in seen: continue
act[i]=1.0;seen[i]=2;qq.append((i,0))
qh=0
while qh<len(qq):
cur,h=qq[qh];qh+=1
if h>=DEPTH: continue
parent=act[cur]
for e,oi in ADJ_F[cur]+ADJ_T[cur]:
if e['rel'] not in STRUCT or EXCL[oi]: continue
na=parent*e['w']*DECAY*SAL[oi]
if na<FIRE: continue
if seen.get(oi) and na<=act.get(oi,0): continue
act[oi]=na
if not seen.get(oi): seen[oi]=1
if len(qq)<AMAX*4: qq.append((oi,h+1))
A=sorted([(i,float(cos[i])) for i,st in seen.items() if st==1 and OK[i] and HAVE[i] and cos[i]>0.0],key=lambda x:-x[1])[:AMAX]
return [i for i,_,_ in L],Sl,[i for i,_ in A]
def merge(L,S,A,lim=10):
out=[];li=si=ai=0
while len(out)<lim and (li<len(L) or si<len(S) or ai<len(A)):
if li<len(L):
if L[li] not in out: out.append(L[li])
li+=1
if len(out)>=lim: break
if si<len(S):
if S[si] not in out: out.append(S[si])
si+=1
if len(out)>=lim: break
if ai<len(A):
if A[ai] not in out: out.append(A[ai])
ai+=1
return out
def outcome(q,ids):
c=q['category']
if c=='nonsense': return len(ids)==0
if c=='superseded':
a,b=q['must_outrank']
if a not in ids: return False
if b not in ids: return True
return ids.index(a)<ids.index(b)
return any(x in ids[:5] for x in q['relevant'])
def run(**kw):
return {q['id']:outcome(q,[NODES[i]['id'] for i in merge(*legs4(q['query'],**kw),10)]) for q in G['queries']}
base=run()
print("baseline",sum(base.values()),"misses",[k for k,v in base.items() if not v])
for name,kw in [('wordstart',dict(wordstart=True)),
('unfloor',dict(unfloor=True)),
('wordstart+unfloor',dict(wordstart=True,unfloor=True))]:
r=run(**kw)
g=sorted(k for k in base if r[k] and not base[k]);l=sorted(k for k in base if base[k] and not r[k])
print("%-20s net=%+d gains=%s losses=%s"%(name,len(g)-len(l),g,l))
+31
View File
@@ -0,0 +1,31 @@
import numpy as np, json, urllib.request
SP="/private/tmp/claude-501/-Users-timlingo/82369039-a20e-4b5a-8a5e-28234a57b996/scratchpad"
M=np.load(SP+'/emb.npy'); ids=open(SP+'/ids.txt',encoding='utf-8',errors='surrogateescape').read().split('\n')
np.seterr(all='ignore')
bad=~np.isfinite(M).all(axis=1)
M[bad]=0.0
print("non-finite rows zeroed:",int(bad.sum()))
idx={k:i for i,k in enumerate(ids)}
gold=json.load(open("/Users/timlingo/Development/neuron-technologies/_wt-assoc-leg/tools/retrieval-eval/gold_set.json"))['queries']
def emb(t):
b=json.dumps({"model":"nomic-embed-text","prompt":t}).encode()
r=urllib.request.Request("http://127.0.0.1:11434/api/embeddings",data=b,headers={"Content-Type":"application/json"})
v=np.array(json.load(urllib.request.urlopen(r,timeout=60))["embedding"],dtype=np.float32)
return v/(np.linalg.norm(v)+1e-9)
out={}
for q in gold:
v=emb(q['query']); s=M@v
s=s[np.isfinite(s)]
mu=float(s.mean()); sd=float(s.std())
top=np.sort(s)[::-1][:10]
z=[(float(t)-mu)/sd for t in top]
grank=[]
for rel in q['relevant']:
if rel in idx:
j=idx[rel]; grank.append((int((M@v > (M@v)[j]).sum())+1, round(float((M@v)[j]),3)))
grank.sort()
out[q['id']]=dict(cat=q['category'],mu=round(mu,3),sd=round(sd,4),top1=round(float(top[0]),3),
z1=round(z[0],2),z3=round(z[2],2),z5=round(z[4],2),gold=grank[:1])
print("%s %-11s mu=%.3f sd=%.4f top1=%.3f z1=%5.2f z3=%5.2f z5=%5.2f gold=%s"%(
q['id'],q['category'],mu,sd,top[0],z[0],z[2],z[4],grank[:1]))
json.dump(out,open(SP+'/zprobe.json','w'),indent=1)
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env bash
# soulc-stamp.sh — make it impossible for dist/soul.c to drift from the sources
# in silence.
#
# THE PROBLEM (neuron#133, and its own words): "Nothing in the tree regenerates
# this file. Only a human running the recipe. It lags in batches, never
# per-change, and it will drift again."
#
# It drifted. On 2026-08-07 a CI or GKE build off main would have shipped an
# engine with NONE of five merged fixes — including a P0 safety fix — while
# main's source read as correct. CI compiles dist/soul.c, not the .el files, so
# the source being right is not the same as the build being right.
#
# WHY A STAMP AND NOT AUTO-REGENERATION: the CI workflow says elc cannot run on
# the runner ("elb on Linux would OOM the runner (elc uses 24GB+ virtual memory
# on a 16GB host)"). So the build cannot regenerate the file itself. What it CAN
# do, for free and with no compiler, is refuse to compile a stale one.
#
# The stamp records a fingerprint of every .el source that feeds the amalgam at
# the moment it was generated. --check recomputes and compares. Divergence is a
# build failure with the recipe in the message, not a silent ship.
#
# soulc-stamp.sh --write after regenerating dist/soul.c (records the fingerprint)
# soulc-stamp.sh --check in CI, before the compile (fails on drift)
set -u
MODE="${1:---check}"
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
STAMP="$ROOT/dist/soul.c.stamp"
AMALGAM="$ROOT/dist/soul.c"
# Every .el at the repo root is an input to the amalgam. Sorted so the hash is
# order-independent; content-only so timestamps and checkouts do not perturb it.
fingerprint() {
(
cd "$ROOT" || exit 1
for f in $(ls -1 *.el 2>/dev/null | sort); do
printf '%s %s\n' "$(shasum -a 256 "$f" | awk '{print $1}')" "$f"
done
)
}
case "$MODE" in
--write)
[ -f "$AMALGAM" ] || { echo "no dist/soul.c to stamp — regenerate it first" >&2; exit 2; }
{
echo "# soul.c.stamp — fingerprint of the .el sources dist/soul.c was generated from."
echo "# Written by tools/soulc-stamp.sh --write. Do not hand-edit."
echo "# generated_amalgam_sha256 $(shasum -a 256 "$AMALGAM" | awk '{print $1}')"
echo "# generated_amalgam_bytes $(wc -c < "$AMALGAM" | tr -d ' ')"
fingerprint
} > "$STAMP"
echo "stamped $(fingerprint | wc -l | tr -d ' ') sources -> dist/soul.c.stamp"
;;
--check)
if [ ! -f "$STAMP" ]; then
echo "FAIL: dist/soul.c.stamp is missing — the build input is unverifiable." >&2
echo " Regenerate the amalgam, then: tools/soulc-stamp.sh --write" >&2
exit 1
fi
RECORDED="$(grep -v '^#' "$STAMP")"
CURRENT="$(fingerprint)"
if [ "$RECORDED" = "$CURRENT" ]; then
echo "soulc-stamp: OK — dist/soul.c matches the .el sources"
exit 0
fi
echo "FAIL: dist/soul.c is STALE. It does not match the current .el sources." >&2
echo "" >&2
echo "CI compiles dist/soul.c, not the .el files. Shipping this means shipping" >&2
echo "an engine that does not contain the merged source. That is neuron#133," >&2
echo "which once hid five merged fixes including a P0 safety fix." >&2
echo "" >&2
echo "Sources that changed since the amalgam was generated:" >&2
diff <(printf '%s\n' "$RECORDED") <(printf '%s\n' "$CURRENT") \
| grep -E '^[<>]' | awk '{print " " $1 " " $3}' | sort -u >&2
echo "" >&2
echo "Fix: regenerate the amalgam, then tools/soulc-stamp.sh --write" >&2
exit 1
;;
*)
echo "usage: soulc-stamp.sh [--check|--write]" >&2; exit 2 ;;
esac
+28
View File
@@ -0,0 +1,28 @@
# El Compiler Release v1.0.0 — 2026-05-02
## Components
- `bootstrap.py` — El language compiler (Python, recursive descent parser, emits C)
- `el_runtime.c` — El runtime (C, HTTP server, engram, DHARMA, LLM chain)
- `el_runtime.h` — Runtime public API header
## Changes in this release
### Critical bug fixes
- `state_set`/`state_get` are now thread-safe (pthread_mutex). Was racing across 64 worker threads.
- `looks_like_string` threshold raised from 1,000,000 to 4GB. Unix timestamps were being dereferenced as heap pointers.
- `fs_read` guards against negative `ftell` result (pipe/special file overflow).
### Engram architecture (major)
- Two-layer activation: `background_activation` (Layer 1, broad fan-out) + `working_memory_weight` (Layer 2, executive filter)
- Inhibitory edges: `EngramEdge.inhibitory` flag suppresses working memory promotion without affecting background activation
- Suppression memory: `suppression_count` — nodes activated-but-suppressed accumulate pressure toward breakthrough
- Temporal decay: `temporal_decay_rate`, `created_at`, `last_activated_at`, `activation_count` on EngramNode
- Per-type activation thresholds (Safety: 0.05, Canonical: 0.15, Lesson: 0.25, Note: 0.40)
- Temporal range query: `engram_query_range(start_ms, end_ms)`
- Layered consciousness: `EngramLayer` struct, `layer_id` on nodes and edges, `EngramStore.layers[]`
- Layer 0 override pass: safety layer fires last and cannot be suppressed
## SHA256
bootstrap.py
el_runtime.c
el_runtime.h
File diff suppressed because it is too large Load Diff
+787
View File
@@ -0,0 +1,787 @@
/*
* el_runtime.h El language C runtime header
*
* Declares all built-in functions available to compiled El programs.
* Include this in every generated .c file.
*
* Value model:
* All El values are represented as el_val_t (= int64_t).
* On 64-bit systems a pointer fits in int64_t.
* String values are cast: (el_val_t)(uintptr_t)"hello"
* Integer values are stored directly.
* This lets arithmetic work naturally while still passing strings around.
*
* Type conventions (El -> C):
* String -> el_val_t (holds const char* via uintptr_t cast)
* Int -> el_val_t
* Bool -> el_val_t (0 = false, nonzero = true)
* Any -> el_val_t
* Void -> void
*
* Macros for convenience:
* EL_STR(s) cast string literal to el_val_t
* EL_CSTR(v) cast el_val_t back to const char*
* EL_INT(v) identity el_val_t is already int64_t
*
* Link requirements:
* -lcurl required for the HTTP client (http_get, http_post, llm_*).
* -lpthread required for the HTTP server (one detached thread per
* connection, capped at 64 concurrent).
* -loqs optional; required only when liboqs is installed and the
* pq_* / sha3_256_hex entry points are needed. Detected at
* compile time via __has_include(<oqs/oqs.h>).
* -lcrypto optional; pulled in alongside -loqs. Used for X25519 in
* pq_hybrid_* and HKDF-SHA256 derivation.
*
* Canonical compile command:
* cc -std=c11 -I el-compiler/runtime -lcurl -lpthread \
* -o <out> <prog>.c el-compiler/runtime/el_runtime.c
*
* With liboqs (post-quantum stack):
* cc -std=c11 -I el-compiler/runtime -lcurl -lpthread -loqs -lcrypto \
* -o <out> <prog>.c el-compiler/runtime/el_runtime.c
*/
#pragma once
#include <stdint.h>
#include <stdlib.h>
typedef int64_t el_val_t;
#define EL_STR(s) ((el_val_t)(uintptr_t)(s))
#define EL_CSTR(v) ((const char*)(uintptr_t)(v))
#define EL_INT(v) (v)
#define EL_NULL ((el_val_t)0)
/* Float values share the el_val_t (int64) slot via a bit-cast.
* The codegen emits Float literals as `el_from_float(<dbl>)` so the
* underlying bits represent the IEEE 754 double. Float-aware builtins
* (math, format, json) round-trip via these helpers. */
static inline double el_to_float(el_val_t v) {
union { int64_t i; double f; } u;
u.i = (int64_t)v;
return u.f;
}
static inline el_val_t el_from_float(double f) {
union { double f; int64_t i; } u;
u.f = f;
return (el_val_t)u.i;
}
#ifdef __cplusplus
extern "C" {
#endif
/* ── I/O ──────────────────────────────────────────────────────────────────── */
void println(el_val_t s);
void print(el_val_t s);
el_val_t readline(void);
/* ── String builtins ─────────────────────────────────────────────────────── */
el_val_t el_str_concat(el_val_t a, el_val_t b);
el_val_t str_eq(el_val_t a, el_val_t b);
el_val_t str_starts_with(el_val_t s, el_val_t prefix);
el_val_t str_ends_with(el_val_t s, el_val_t suffix);
el_val_t str_len(el_val_t s);
el_val_t str_concat(el_val_t a, el_val_t b);
el_val_t int_to_str(el_val_t n);
el_val_t str_to_int(el_val_t s);
el_val_t str_slice(el_val_t s, el_val_t start, el_val_t end);
el_val_t str_contains(el_val_t s, el_val_t sub);
el_val_t str_replace(el_val_t s, el_val_t from, el_val_t to);
el_val_t str_to_upper(el_val_t s);
el_val_t str_to_lower(el_val_t s);
el_val_t str_trim(el_val_t s);
/* ── Math ────────────────────────────────────────────────────────────────── */
el_val_t el_abs(el_val_t n);
el_val_t el_max(el_val_t a, el_val_t b);
el_val_t el_min(el_val_t a, el_val_t b);
/* ── Refcount (ARC) ──────────────────────────────────────────────────────────
* Lists and Maps carry a refcount. Strings and ints do not el_retain and
* el_release are safe no-ops on non-refcounted values (they sniff a magic
* header at offset 0 and only act if the magic matches).
*
* Codegen emits these at let-binding shadowing, function entry (params), and
* function exit (locals other than the returned value). The refcount lets
* el_list_append and el_map_set mutate in place when uniquely owned (cheap)
* and copy-on-write when shared (preserves persistent semantics across
* accumulator patterns in the compiler itself). */
void el_retain(el_val_t v);
void el_release(el_val_t v);
/* ── Arena scoping ────────────────────────────────────────────────────────────
* el_arena_push() activates the string arena (if not already active) and
* returns a mark; el_arena_pop(mark) frees all strings allocated since that
* mark. Used by codegen for per-function/statement scoping and by long-running
* EL loops (e.g. the soul daemon's awareness tick) to reclaim per-iteration
* allocations. */
el_val_t el_arena_push(void);
el_val_t el_arena_pop(el_val_t mark);
/* ── List ────────────────────────────────────────────────────────────────── */
el_val_t el_list_new(el_val_t count, ...);
el_val_t el_list_len(el_val_t list);
el_val_t el_list_get(el_val_t list, el_val_t index);
el_val_t el_list_append(el_val_t list, el_val_t elem);
el_val_t el_list_empty(void);
el_val_t el_list_clone(el_val_t list);
/* ── Map ─────────────────────────────────────────────────────────────────── */
el_val_t el_map_new(el_val_t pair_count, ...);
el_val_t el_get_field(el_val_t map, el_val_t key);
el_val_t el_map_get(el_val_t map, el_val_t key);
el_val_t el_map_set(el_val_t map, el_val_t key, el_val_t value);
/* ── HTTP ─────────────────────────────────────────────────────────────────── */
el_val_t http_get(el_val_t url);
el_val_t http_post(el_val_t url, el_val_t body);
el_val_t http_post_json(el_val_t url, el_val_t json_body);
el_val_t http_get_with_headers(el_val_t url, el_val_t headers_map);
el_val_t http_post_with_headers(el_val_t url, el_val_t body, el_val_t headers_map);
el_val_t http_post_form_auth(el_val_t url, el_val_t form_body, el_val_t auth_header);
el_val_t http_delete(el_val_t url);
el_val_t http_delete_json(el_val_t url, el_val_t json_body);
void http_serve(el_val_t port, el_val_t handler);
void http_set_handler(el_val_t name);
/* HTTP server v2 ─────────────────────────────────────────────────────────────
* Same dispatch model as http_serve, but the handler signature is widened:
*
* el_val_t handler(method, path, headers_map, body)
*
* `headers_map` is an ElMap from lowercased header name header value (both
* Strings). Repeated headers are joined with ", " per RFC 7230.
*
* Response value: the handler may return either
* (a) a plain body string same auto-content-type / 200-OK behaviour as
* http_serve (3-arg) or
* (b) a response envelope built with `http_response(status, headers_json,
* body)`. The runtime detects the envelope discriminator
* `"el_http_response":1` at the start of the returned string and
* unpacks status / headers / body before sending.
*
* The 3-arg http_serve(port, handler) remains supported unchanged for
* existing handlers (e.g. products/web/server.el): it dispatches with
* (method, path, body), hardcodes 200 OK, and auto-detects content type. */
void http_serve_v2(el_val_t port, el_val_t handler);
void http_set_handler_v2(el_val_t name);
/* Non-blocking variant of http_serve: runs the accept loop in a background
* pthread and returns immediately so the caller can continue (used by the
* soul daemon to run awareness_run() after starting its HTTP API). */
void http_serve_async(el_val_t port, el_val_t handler);
/* Build an HTTP response envelope. `headers_json` should be a JSON object
* literal like `{"WWW-Authenticate":"Basic"}` (or "" / "{}" for none). The
* returned string carries the discriminator `{"el_http_response":1,...}`
* which the runtime's send-path detects and unpacks. Detection happens
* uniformly inside http_send_response, so a 3-arg handler may also return
* an envelope. The 3-arg variant remains documented as a fixed 200-OK
* auto-content-type contract for legacy handlers that return plain bodies. */
el_val_t http_response(el_val_t status, el_val_t headers_json, el_val_t body);
/* HTTP timeout — every libcurl request honors EL_HTTP_TIMEOUT_MS (default
* 60000ms). Read lazily on first use, so setting the env var any time before
* the first http_* call is sufficient. */
/* Streaming variants — write the response body straight to a file via
* libcurl's CURLOPT_WRITEFUNCTION = fwrite. These bypass the el_val_t string
* wrapper entirely, so binary payloads (audio/mpeg, image/png, etc.) survive
* embedded NUL bytes that would truncate a strlen()-based code path.
*
* Both honor EL_HTTP_TIMEOUT_MS, follow redirects, and accept the same
* `headers_map` shape as http_post_with_headers (ElMap of StringString).
*
* Return value: 1 on success (file fully written), 0 on any failure
* (network, file open, partial write). On failure the output file is removed
* so callers cannot mistake a partially-written file for a valid one. */
el_val_t http_post_to_file(el_val_t url, el_val_t body, el_val_t headers_map, el_val_t output_path);
el_val_t http_get_to_file(el_val_t url, el_val_t headers_map, el_val_t output_path);
/* ── URL encoding ────────────────────────────────────────────────────────── */
el_val_t url_encode(el_val_t s); /* RFC 3986 unreserved set */
el_val_t url_decode(el_val_t s); /* '+' → space, %XX → byte */
/* ── HTML allowlist sanitizer ────────────────────────────────────────────────
* el_html_sanitize(input_html, allowlist_json) strict allowlist HTML
* cleaner. State-machine parser; tag/attribute names compared case-
* insensitively against the allowlist; `<a href>` / `< src>` URL schemes
* validated (http, https, mailto, fragment-only, or relative); whole-
* subtree drop for script / style / iframe / object / embed / form; HTML-
* escapes free text outside dropped subtrees.
*
* The allowlist is JSON of the form
* {"p":[],"a":["href","title"],"strong":[],...}
* where each value is the array of attribute names allowed for that tag. */
el_val_t el_html_sanitize(el_val_t input_html, el_val_t allowlist_json);
/* ── Filesystem ──────────────────────────────────────────────────────────── */
el_val_t fs_read(el_val_t path);
el_val_t fs_write(el_val_t path, el_val_t content);
el_val_t fs_list(el_val_t path);
el_val_t fs_exists(el_val_t path);
el_val_t fs_mkdir(el_val_t path); /* mkdir -p, mode 0755 */
/* Length-explicit binary write. `length` is an Int (el_val_t holding the
* byte count). The caller knows the length from context typically because
* `bytes` came from base64_decode (which produces a magic-tagged binary
* buffer with embedded NULs possible) and the caller already tracks the
* decoded length, OR because the bytes came from a fixed-size source
* (sha256_bytes = 32, hmac_sha256_bytes = 32). Bypasses strlen entirely.
*
* Returns 1 on success, 0 on failure (invalid path, can't open, partial
* write, negative length). On partial-write failure, the file is removed
* so callers cannot read back a truncated artefact. */
el_val_t fs_write_bytes(el_val_t path, el_val_t bytes, el_val_t length);
/* ── JSON ────────────────────────────────────────────────────────────────── */
el_val_t json_get(el_val_t json, el_val_t key);
el_val_t json_parse(el_val_t s);
el_val_t json_stringify(el_val_t v);
el_val_t json_get_string(el_val_t json_str, el_val_t key);
el_val_t json_get_int(el_val_t json_str, el_val_t key);
el_val_t json_get_float(el_val_t json_str, el_val_t key);
el_val_t json_get_bool(el_val_t json_str, el_val_t key);
el_val_t json_get_raw(el_val_t json_str, el_val_t key);
el_val_t json_set(el_val_t json_str, el_val_t key, el_val_t value);
el_val_t json_array_len(el_val_t json_str);
el_val_t json_array_get(el_val_t json_str, el_val_t index);
el_val_t json_array_get_string(el_val_t json_str, el_val_t index);
/* ── Time ────────────────────────────────────────────────────────────────── */
el_val_t time_now(void);
el_val_t time_now_utc(void);
el_val_t sleep_secs(el_val_t secs);
el_val_t sleep_ms(el_val_t ms);
el_val_t time_format(el_val_t ts, el_val_t fmt);
el_val_t time_to_parts(el_val_t ts);
el_val_t time_from_parts(el_val_t secs, el_val_t ns, el_val_t tz);
el_val_t time_add(el_val_t ts, el_val_t n, el_val_t unit);
el_val_t time_diff(el_val_t ts1, el_val_t ts2, el_val_t unit);
/* ── Instant + Duration: first-class temporal types ──────────────────────────
* Both types share the el_val_t (int64) slot. Instants are nanoseconds
* since the Unix epoch; Durations are signed nanoseconds. Type discipline
* is enforced at codegen-time: BinOps on names registered as Instant or
* Duration route through the typed wrappers below; mismatches like
* Instant+Instant become #error at the C compiler.
*
* Postfix literals `30.seconds`, `1.hour`, `500.millis`, `30.nanos` are
* recognised by the parser as DurationLit AST nodes and lowered to literal
* int64 nanoseconds at codegen time. The runtime never sees the units. */
el_val_t el_now_instant(void);
el_val_t now(void);
el_val_t unix_seconds(el_val_t n);
el_val_t unix_millis(el_val_t n);
el_val_t instant_from_iso8601(el_val_t s);
el_val_t el_duration_from_nanos(el_val_t ns);
el_val_t duration_seconds(el_val_t n);
el_val_t duration_millis(el_val_t n);
el_val_t duration_nanos(el_val_t n);
el_val_t el_instant_add_dur(el_val_t inst, el_val_t dur);
el_val_t el_instant_sub_dur(el_val_t inst, el_val_t dur);
el_val_t el_instant_diff(el_val_t a, el_val_t b);
el_val_t el_duration_add(el_val_t a, el_val_t b);
el_val_t el_duration_sub(el_val_t a, el_val_t b);
el_val_t el_duration_scale(el_val_t dur, el_val_t scalar);
el_val_t el_duration_div(el_val_t dur, el_val_t scalar);
el_val_t el_instant_lt(el_val_t a, el_val_t b);
el_val_t el_instant_le(el_val_t a, el_val_t b);
el_val_t el_instant_gt(el_val_t a, el_val_t b);
el_val_t el_instant_ge(el_val_t a, el_val_t b);
el_val_t el_instant_eq(el_val_t a, el_val_t b);
el_val_t el_instant_ne(el_val_t a, el_val_t b);
el_val_t el_duration_lt(el_val_t a, el_val_t b);
el_val_t el_duration_le(el_val_t a, el_val_t b);
el_val_t el_duration_gt(el_val_t a, el_val_t b);
el_val_t el_duration_ge(el_val_t a, el_val_t b);
el_val_t el_duration_eq(el_val_t a, el_val_t b);
el_val_t el_duration_ne(el_val_t a, el_val_t b);
el_val_t instant_to_unix_seconds(el_val_t i);
el_val_t instant_to_unix_millis(el_val_t i);
el_val_t instant_to_iso8601(el_val_t i);
el_val_t duration_to_seconds(el_val_t d);
el_val_t duration_to_millis(el_val_t d);
el_val_t duration_to_nanos(el_val_t d);
el_val_t el_sleep_duration(el_val_t dur);
el_val_t unix_timestamp(void);
el_val_t ttl_cache_set(el_val_t key, el_val_t value);
el_val_t ttl_cache_get(el_val_t key, el_val_t max_age);
el_val_t ttl_cache_age(el_val_t key);
/* ── Calendar + CalendarTime + Rhythm + LocalDate/Time/DateTime ─────────────
* Phase 1.5 of the time system. Calendar is pluggable: EarthCalendar (IANA
* zones, Gregorian, DST) is the user-facing default; MarsCalendar,
* CycleCalendar(period), NoCycleCalendar, RelativeCalendar handle non-Earth
* domains.
*
* A Calendar interprets an Instant under a particular cycle convention and
* produces a CalendarTime. CalendarTime carries the underlying Instant and
* a back-pointer to its Calendar; arithmetic and formatting consult the
* Calendar to convert ns since epoch into year/month/day/hour/minute/second
* (or sol/phase, or cycle/phase, depending on kind).
*
* Storage convention: Calendar / CalendarTime / Rhythm / LocalDate /
* LocalDateTime are heap-allocated structs whose pointers are cast into
* el_val_t. A 24-bit magic header at offset 0 lets the runtime identify
* the kind safely. LocalTime is small enough to live in the int64 slot
* directly (nanos since midnight, signed). */
/* Zone — opaque IANA zone or fixed offset, used by EarthCalendar.
* `zone_id` is either an IANA name ("America/New_York", "UTC") or a fixed
* offset string ("+05:30", "-08:00"). The runtime resolves it via tzset()
* on first use of the owning EarthCalendar. */
el_val_t zone(el_val_t id);
el_val_t zone_utc(void);
el_val_t zone_local(void);
el_val_t zone_offset(el_val_t hours, el_val_t minutes);
/* Calendar constructors. Each returns an el_val_t pointer to a heap-
* allocated, magic-tagged Calendar struct. Calendars are interned by
* (kind, zone_id, period_ns, epoch_ns) so identical constructors return
* the same pointer equality is reference equality. */
el_val_t earth_calendar(el_val_t z);
el_val_t earth_calendar_default(void);
el_val_t mars_calendar(void);
el_val_t cycle_calendar(el_val_t period_dur);
el_val_t no_cycle_calendar(void);
el_val_t relative_calendar(el_val_t epoch_inst);
/* CalendarTime constructors and methods. Returns a heap-allocated struct
* whose pointer fits in el_val_t. */
el_val_t now_in(el_val_t cal);
el_val_t in_calendar(el_val_t inst, el_val_t cal);
el_val_t cal_format(el_val_t ct, el_val_t pattern);
el_val_t cal_to_instant(el_val_t ct);
el_val_t cal_cycle_phase(el_val_t ct);
el_val_t cal_in(el_val_t ct, el_val_t cal);
/* LocalDate / LocalTime / LocalDateTime — calendar-agnostic value types.
* LocalTime carries nanoseconds since midnight as a signed int64 directly
* in the el_val_t slot (no allocation). LocalDate / LocalDateTime are
* heap-allocated structs with magic headers. */
el_val_t local_date(el_val_t y, el_val_t m, el_val_t d);
el_val_t local_time(el_val_t h, el_val_t m, el_val_t s, el_val_t ns);
el_val_t local_datetime(el_val_t date, el_val_t time);
el_val_t zoned(el_val_t date, el_val_t time, el_val_t cal);
el_val_t local_date_year(el_val_t ld);
el_val_t local_date_month(el_val_t ld);
el_val_t local_date_day(el_val_t ld);
el_val_t local_time_hour(el_val_t lt);
el_val_t local_time_minute(el_val_t lt);
el_val_t local_time_second(el_val_t lt);
el_val_t local_time_nanos(el_val_t lt);
el_val_t el_local_date_add_dur(el_val_t ld, el_val_t dur);
el_val_t el_local_time_add_dur(el_val_t lt, el_val_t dur);
el_val_t el_local_date_lt(el_val_t a, el_val_t b);
el_val_t el_local_date_eq(el_val_t a, el_val_t b);
/* Rhythm — pluggable recurrence AST. Returns a heap-allocated struct
* pointer in el_val_t; rhythms are immutable so callers may share them. */
el_val_t rhythm_cycle_start(void);
el_val_t rhythm_cycle_phase(el_val_t phase);
el_val_t rhythm_duration(el_val_t d);
el_val_t rhythm_session_start(void);
el_val_t rhythm_event(el_val_t name);
el_val_t rhythm_and(el_val_t a, el_val_t b);
el_val_t rhythm_or(el_val_t a, el_val_t b);
el_val_t rhythm_weekday(el_val_t day);
el_val_t rhythm_weekly_at(el_val_t day, el_val_t hour, el_val_t minute);
el_val_t rhythm_next_after(el_val_t r, el_val_t after, el_val_t cal);
el_val_t rhythm_matches(el_val_t r, el_val_t ct);
/* ── UUID ────────────────────────────────────────────────────────────────── */
el_val_t uuid_new(void);
el_val_t uuid_v4(void);
/* ── Environment ─────────────────────────────────────────────────────────── */
el_val_t env(el_val_t key);
/* ── In-process state K/V ────────────────────────────────────────────────── */
el_val_t state_set(el_val_t key, el_val_t value);
el_val_t state_get(el_val_t key);
el_val_t state_del(el_val_t key);
el_val_t state_keys(void);
/* ── Float formatting ────────────────────────────────────────────────────── */
el_val_t float_to_str(el_val_t f);
el_val_t int_to_float(el_val_t n);
el_val_t float_to_int(el_val_t f);
el_val_t format_float(el_val_t f, el_val_t decimals);
el_val_t decimal_round(el_val_t f, el_val_t decimals);
el_val_t str_to_float(el_val_t s);
/* ── Math (Float-aware) ──────────────────────────────────────────────────── */
el_val_t math_sqrt(el_val_t f);
el_val_t math_log(el_val_t f);
el_val_t math_ln(el_val_t f);
el_val_t math_sin(el_val_t f);
el_val_t math_cos(el_val_t f);
el_val_t math_pi(void);
/* ── String additions ────────────────────────────────────────────────────── */
el_val_t str_index_of(el_val_t s, el_val_t sub);
el_val_t str_split(el_val_t s, el_val_t sep);
el_val_t str_char_at(el_val_t s, el_val_t i);
el_val_t str_char_code(el_val_t s, el_val_t i);
el_val_t str_pad_left(el_val_t s, el_val_t width, el_val_t pad);
el_val_t str_pad_right(el_val_t s, el_val_t width, el_val_t pad);
el_val_t str_format(el_val_t fmt, el_val_t data);
el_val_t str_lower(el_val_t s);
el_val_t str_upper(el_val_t s);
/* ── Text-processing primitives (Phase 1: byte/codepoint, ASCII char classes)
* Phase 2 (filed): Unicode-grapheme awareness, NFC/NFD normalization, regex.
* is_* predicates: empty input returns false; multi-char requires ALL bytes
* to match. ASCII ranges only in Phase 1. */
/* Counting */
el_val_t str_count(el_val_t s, el_val_t sub); /* non-overlapping */
el_val_t str_count_chars(el_val_t s); /* codepoint count */
el_val_t str_count_bytes(el_val_t s); /* alias of str_len */
el_val_t str_count_lines(el_val_t s);
el_val_t str_count_words(el_val_t s);
el_val_t str_count_letters(el_val_t s); /* ASCII [A-Za-z] */
el_val_t str_count_digits(el_val_t s); /* ASCII [0-9] */
/* Find / position */
el_val_t str_index_of_all(el_val_t s, el_val_t sub); /* [Int] of byte offsets */
el_val_t str_last_index_of(el_val_t s, el_val_t sub);
el_val_t str_find_chars(el_val_t s, el_val_t any_of); /* first idx of any ch */
/* Transform */
el_val_t str_repeat(el_val_t s, el_val_t n);
el_val_t str_reverse(el_val_t s); /* by codepoint */
el_val_t str_strip_prefix(el_val_t s, el_val_t prefix);
el_val_t str_strip_suffix(el_val_t s, el_val_t suffix);
el_val_t str_strip_chars(el_val_t s, el_val_t chars);
el_val_t str_lstrip(el_val_t s);
el_val_t str_rstrip(el_val_t s);
/* Char classification (Bool) */
el_val_t is_letter(el_val_t s);
el_val_t is_digit(el_val_t s);
el_val_t is_alphanumeric(el_val_t s);
el_val_t is_whitespace(el_val_t s);
el_val_t is_punctuation(el_val_t s);
el_val_t is_uppercase(el_val_t s);
el_val_t is_lowercase(el_val_t s);
/* Split / join */
el_val_t str_split_lines(el_val_t s);
el_val_t str_split_chars(el_val_t s); /* alias of native_string_chars */
el_val_t str_split_n(el_val_t s, el_val_t sep, el_val_t n);
el_val_t str_join(el_val_t list, el_val_t sep); /* alias of list_join */
/* ── List additions ──────────────────────────────────────────────────────── */
el_val_t list_push(el_val_t list, el_val_t elem);
el_val_t list_push_front(el_val_t list, el_val_t elem);
el_val_t list_join(el_val_t list, el_val_t sep);
el_val_t list_range(el_val_t start, el_val_t end);
/* ── Bool helpers ────────────────────────────────────────────────────────── */
el_val_t bool_to_str(el_val_t b);
/* ── Numeric parsing ─────────────────────────────────────────────────────── */
el_val_t parse_int(el_val_t s, el_val_t default_val);
/* ── Process ─────────────────────────────────────────────────────────────── */
void exit_program(el_val_t code);
el_val_t getpid_now(void);
/* ── CGI identity ─────────────────────────────────────────────────────────────
* Called at the start of main() in CGI programs (those with a `cgi {}` block).
* Records the program's DHARMA identity before any other code executes. */
void el_cgi_init(el_val_t name, el_val_t dharma_id, el_val_t principal,
el_val_t network, el_val_t engram);
/* ── DHARMA network builtins ─────────────────────────────────────────────────
* Available to CGI programs (declared with a `cgi {}` block).
*
* Peers are addressed by `dharma_id` of the form
* "<registry-id>@<transport-url>" e.g. "ntn-genesis@http://localhost:7770"
* If the @<url> portion is omitted, transport defaults to
* "http://localhost:7770" (the local CGI daemon assumption).
*
* Wire protocol (all peers expose):
* POST <url>/dharma/recv { channel, from, content } response body
* POST <url>/dharma/event { type, payload, source, timestamp }
* POST <url>/api/activate { query } list of nodes
*
* Hosting application's responsibility: an El program with a `cgi {}` block
* runs http_serve() with its own request handler; that handler should route
* "/dharma/event" requests by calling el_runtime_dharma_event_arrive() so
* incoming events feed dharma_field() queues. The runtime itself does not
* intercept any /dharma path. */
el_val_t dharma_connect(el_val_t cgi_id);
el_val_t dharma_send(el_val_t channel, el_val_t content);
el_val_t dharma_activate(el_val_t query);
void dharma_emit(el_val_t event_type, el_val_t payload);
el_val_t dharma_field(el_val_t event_type);
void dharma_strengthen(el_val_t cgi_id, el_val_t weight);
el_val_t dharma_relationship(el_val_t cgi_id);
el_val_t dharma_peers(void);
/* Public C API: called by an El program's HTTP handler when a /dharma/event
* request arrives. Pushes onto the per-event-type queue and signals any
* pending dharma_field() blockers. All three arguments must be NUL-terminated
* C strings (or NULL then treated as empty). */
void el_runtime_dharma_event_arrive(const char* event_type,
const char* payload,
const char* source);
/* ── Engram local graph primitives ───────────────────────────────────────────
* Operate on the CGI's local Engram knowledge graph.
* `engram_activate` queries the local graph only; `dharma_activate` is
* network-wide across all connected CGI graphs. */
el_val_t engram_node(el_val_t content, el_val_t node_type, el_val_t salience);
el_val_t engram_node_full(el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t importance, el_val_t confidence,
el_val_t tier, el_val_t tags);
/* Layered consciousness — see el_runtime.c for the layered architecture
* design notes (search "Layered consciousness architecture"). The five
* canonical layers (safety / core-identity / domain-knowledge / imprint /
* suit) are seeded automatically; engram_add_layer extends the registry
* with imprint or suit overlays at runtime. Nodes default to layer 1
* (core-identity) when created via engram_node / engram_node_full. */
el_val_t engram_node_layered(el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t certainty, el_val_t confidence,
el_val_t status, el_val_t tags, el_val_t layer_id);
el_val_t engram_add_layer(el_val_t name, el_val_t priority, el_val_t suppressible,
el_val_t transparent, el_val_t injectable);
el_val_t engram_remove_layer(el_val_t layer_id);
el_val_t engram_list_layers(void);
el_val_t engram_get_node(el_val_t id);
void engram_strengthen(el_val_t node_id);
void engram_forget(el_val_t node_id);
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
el_val_t engram_node_count(void);
el_val_t engram_search(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
void engram_connect(el_val_t from_id, el_val_t to_id, el_val_t weight, el_val_t relation);
el_val_t engram_edge_between(el_val_t from_id, el_val_t to_id);
el_val_t engram_neighbors(el_val_t node_id);
el_val_t engram_neighbors_filtered(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_edge_count(void);
/* Three-pass activation: background fan-out → working-memory promotion →
* Layer 0 override. See "Three-pass activation" in el_runtime.c. */
el_val_t engram_activate(el_val_t query, el_val_t depth);
el_val_t engram_save(el_val_t path);
el_val_t engram_load(el_val_t path);
/* JSON-string accessors — return pre-serialized JSON so HTTP handlers
* can pass results straight through without round-tripping ElList/ElMap
* through json_stringify. */
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_recall_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_activate_json(el_val_t query, el_val_t depth);
el_val_t engram_stats_json(void);
el_val_t engram_act_stats_json(void);
el_val_t engram_cosine_sim(el_val_t id_a, el_val_t id_b);
/* Document frequency of a term across node labels — term-specificity signal
* for curiosity seed selection. (2026-08-03 self-review.) */
el_val_t engram_label_df(el_val_t term);
el_val_t engram_embed_backfill(el_val_t count);
el_val_t engram_list_layers_json(void);
/* Working memory introspection — count, mean weight, and top-N snapshot.
* Ported from el-compiler/runtime on 2026-06-30 self-review. */
el_val_t engram_wm_count(void);
el_val_t engram_wm_avg_weight(void);
el_val_t engram_wm_top_json(el_val_t n);
/* Merge-load: add nodes/edges from a snapshot without resetting the store. */
el_val_t engram_load_merge(el_val_t path);
/* engram_compile_layered_json — produce a prompt-ready text block split
* into "[LAYER 0 — STRUCTURAL]" (non-suppressible layers, sacred fire)
* and "[ENGRAM CONTEXT]" (standard suppressible layers). Returns "" if
* no nodes promoted to working memory. */
el_val_t engram_compile_layered_json(el_val_t intent, el_val_t depth);
/* ── LLM (Anthropic API client) ─────────────────────────────────────────────
* All functions call https://api.anthropic.com/v1/messages with the API key
* from env ANTHROPIC_API_KEY. Default model when empty: claude-sonnet-4-5. */
el_val_t llm_call(el_val_t model, el_val_t prompt);
el_val_t llm_call_system(el_val_t model, el_val_t system_prompt, el_val_t user_prompt);
el_val_t llm_call_agentic(el_val_t model, el_val_t system, el_val_t user, el_val_t tools);
el_val_t llm_vision(el_val_t model, el_val_t system, el_val_t prompt, el_val_t image_url_or_b64);
el_val_t llm_models(void);
/* Register a tool handler by name. The handler is looked up via dlsym
* (mirroring http_set_handler), so any El `fn <name>(input)` compiles to
* a global C symbol that this function can locate at runtime.
* Handler signature: `el_val_t handler(el_val_t input_json)` receives
* the tool input as a JSON-string el_val_t and returns a JSON-string
* el_val_t result. Used by llm_call_agentic. */
void llm_register_tool(el_val_t name, el_val_t handler_fn_name);
/* ── args() ─────────────────────────────────────────────────────────────────
* Provides access to command-line arguments passed to the program.
* Populated by el_runtime_init_args() before main() runs. */
el_val_t args(void);
void el_runtime_init_args(int argc, char** argv);
/* ── Crypto primitives ─────────────────────────────────────────────────────
* SHA-256, HMAC-SHA-256, and base64 (standard + URL-safe).
* Self-contained no OpenSSL/libcrypto dependency. The implementations are
* adapted from public-domain reference code (Brad Conte / RFC 4648).
*
* Bytes-returning variants (sha256_bytes, hmac_sha256_bytes) return a string
* value whose contents are raw binary; callers usually feed these into
* base64_encode. Note that el_val_t strings are NUL-terminated by convention,
* so the binary payload may contain embedded NULs pass it directly into
* base64_encode (which uses an explicit length) rather than treating it as
* a printable C string.
*
* The "base64" variants emit/accept RFC 4648 standard alphabet with padding.
* The "base64url" variants use URL-safe alphabet (`-`/`_`) with no padding,
* as used in JWTs. */
el_val_t sha256_hex(el_val_t input);
el_val_t sha256_bytes(el_val_t input);
el_val_t hmac_sha256_hex(el_val_t key, el_val_t message);
el_val_t hmac_sha256_bytes(el_val_t key, el_val_t message);
el_val_t base64_encode(el_val_t input);
el_val_t base64_decode(el_val_t input);
el_val_t base64url_encode(el_val_t input);
el_val_t base64url_decode(el_val_t input);
/* Length-aware variants (internal — exposed for the rare caller that already
* has a known-length binary buffer and doesn't want to round-trip through
* a NUL-terminated el_val_t string). Sha256_bytes and hmac_sha256_bytes feed
* these implicitly. */
el_val_t el_sha256_bytes_n(const unsigned char* data, size_t len);
el_val_t el_base64_encode_n(const unsigned char* data, size_t len, int url_safe);
/* ── Post-quantum primitives (liboqs-backed) ────────────────────────────────
* All inputs/outputs hex-encoded. Algorithm choices:
* Signature: CRYSTALS-Dilithium-3 (NIST level 3, balanced)
* KEM: CRYSTALS-Kyber-768 (NIST level 3)
* Hash: SHA3-256 (Keccak) (PQ-aware protocols favour SHA3 over SHA2)
*
* If liboqs is not linked (detected via __has_include(<oqs/oqs.h>) at compile
* time), the pq_* entry points return a JSON-shaped error string so callers
* fail loudly rather than silently fall back to classical schemes:
* {"error":"liboqs not linked, post-quantum primitives unavailable"}
*
* The hybrid handshake pairs X25519 with Kyber-768 per NIST PQ guidance and
* CNSA 2.0. Combined shared secret is HKDF-SHA256(x25519_ss || kyber_ss).
* Even if Kyber falls, X25519 holds; if X25519 falls under quantum attack,
* Kyber holds. SHA3-256 also remains usable independent of liboqs (the
* Keccak permutation is PQ-OK as a primitive). */
el_val_t pq_keygen_signature(void);
el_val_t pq_sign(el_val_t secret_key_hex, el_val_t message);
el_val_t pq_verify(el_val_t public_key_hex, el_val_t message, el_val_t signature_hex);
el_val_t pq_kem_keygen(void);
el_val_t pq_kem_encaps(el_val_t public_key_hex);
el_val_t pq_kem_decaps(el_val_t secret_key_hex, el_val_t ciphertext_hex);
el_val_t pq_hybrid_keygen(void);
el_val_t pq_hybrid_handshake(el_val_t remote_pub_combined);
el_val_t sha3_256_hex(el_val_t input);
/* ── AEAD: AES-256-GCM (libcrypto-backed) ───────────────────────────────────
* Symmetric authenticated encryption used to wrap envelopes after a KEM
* handshake. Caller MUST supply a 32-byte key (64 hex chars) typically the
* Kyber-768 / hybrid shared_secret, optionally normalized via SHA3-256.
*
* aead_encrypt returns a JSON map {"nonce":"...","ciphertext":"..."} where
* ciphertext is the AES-256-GCM output with the 16-byte auth tag appended.
* Nonce is a fresh 12-byte CSPRNG draw callers never pick the nonce, which
* structurally rules out the GCM nonce-reuse footgun.
*
* aead_decrypt returns the plaintext String, or "" on any failure (including
* auth-tag mismatch). Callers MUST check for "" before trusting the result. */
el_val_t aead_encrypt(el_val_t key_hex, el_val_t plaintext);
el_val_t aead_decrypt(el_val_t key_hex, el_val_t nonce_hex, el_val_t ciphertext_hex);
/* ── Native VM builtin aliases (for compiled El source) ─────────────────────
* These match the El VM's native_* builtins so that El source compiled
* to C can call the same names without modification. */
el_val_t native_list_get(el_val_t list, el_val_t index);
el_val_t native_list_len(el_val_t list);
el_val_t native_list_append(el_val_t list, el_val_t elem);
el_val_t native_list_empty(void);
el_val_t native_list_clone(el_val_t list);
el_val_t native_string_chars(el_val_t s);
el_val_t native_int_to_str(el_val_t n);
/* ── Method-call shorthand aliases ──────────────────────────────────────────
* The El method-call convention `obj.method(args)` compiles to
* `method(obj, args)`. These aliases expose the runtime functions under
* the short names that result from method calls in El source.
*
* Example: `myList.append(x)` `append(myList, x)` (calls this alias)
* `myList.len()` `len(myList)` (calls this alias) */
el_val_t append(el_val_t list, el_val_t elem); /* el_list_append */
el_val_t len(el_val_t list); /* el_list_len */
el_val_t get(el_val_t list, el_val_t index); /* el_list_get */
el_val_t map_get(el_val_t map, el_val_t key); /* el_map_get */
el_val_t map_set(el_val_t map, el_val_t key, el_val_t value); /* el_map_set */
/* ── OTLP/HTTP Observability ─────────────────────────────────────────────── */
/* See bottom of el_runtime.c for the implementation.
* Configured by env vars OTLP_ENDPOINT, OTEL_SERVICE_NAME, OTEL_SERVICE_VERSION.
* No-op when OTLP_ENDPOINT is unset. Drop-on-failure semantics. */
/* ── Subprocess execution ────────────────────────────────────────────────── */
el_val_t exec_command(el_val_t cmd); /* run shell command, return exit code */
el_val_t exec_capture(el_val_t cmd); /* run shell command, capture stdout */
el_val_t exec(el_val_t cmd); /* exec(cmd) → stdout String (30s timeout) */
el_val_t exec_bg(el_val_t cmd); /* exec_bg(cmd) → PID String (non-blocking) */
el_val_t emit_log(el_val_t level, el_val_t msg, el_val_t fields_json);
el_val_t emit_metric(el_val_t name, el_val_t value, el_val_t tags_json);
el_val_t trace_span_start(el_val_t name);
el_val_t trace_span_end(el_val_t span_handle);
el_val_t emit_event(el_val_t name, el_val_t duration_ms);
#ifdef __cplusplus
}
#endif