Commit Graph

2 Commits

Author SHA1 Message Date
Tim Lingo 77e3c5fa26 test(retrieval): a held-out gold set, and the first out-of-sample number
Iteration 7 measured this harness's own ceiling at +3 gross / +1 net against a
decision floor of 6. An instrument whose ceiling sits below its floor cannot
certify or refute anything, so the gold set — not the retriever — was the
blocker. This iteration builds no retrieval mechanism; it fixes the instrument
and uses it once.

extend_gold_set.py appends 37 queries (q39-q75) WITHOUT touching q01-q38, so
every committed baseline and per-query id stays comparable. 30 held-out
paraphrases: targets sampled mechanically (seed 8080) from addressable
500-2600 char nodes outside the original answer space and outside any duplicate
cluster; queries authored from the node body alone, before any retrieval was
run, and each one re-proved at build time to share ZERO content words with its
target. 7 extra nonsense controls, fully mechanical.

Why this was needed: the original 13 paraphrase and 6 associative queries share
ONE answer space — the 13 `Self - Values (grounded)` children. 19 of 35 scored
queries tested retrieval against a single 13-node neighbourhood.

FIRST OUT-OF-SAMPLE RESULT (main vs the accumulated stack, embedded corpus):
  held-out only : +5 / -0, p=0.0625 — one query short of the floor, NOT-SHOWN
  original 38   : +15 / -0
  full 75       : +20 / -0, p=0.0000, latency 0.46x
In-sample paraphrase 61.5% vs out-of-sample 16.7%: generalisation is real,
directional and 3.7x weaker than the headline number suggested.

Also recorded: 47.4% of this corpus is redundant and ONE record accounts for
46.6% of all 78,768 nodes (36,737 byte-identical copies under distinct ids).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:16:42 -05:00
Neuron cf41d12d22 test(retrieval): a measurement harness for memory recall, and its first verdict
Neuron Soul CI / build (pull_request) Failing after 14m41s
Neuron Soul CI / deploy (pull_request) Failing after 14m45s
Nothing else on the memory roadmap should be built until a change can be shown
to help. Right now we judge by feel, and the benchmark literature is full of
systems that felt better and measured worse. This is the missing gate.

WHAT IT MEASURES, AND WHY IT BOOTS A REAL SOUL
The subject is Will's designed retrieval — spreading activation over the
weighted directed graph, four-factor multiplicative scoring — not a proxy for
it. A Python re-implementation would measure my reading of the design, so the
harness compiles the actual soul.el amalgam from a git ref and asks it over
HTTP on /api/neuron/recall, exactly as the MCP wrapper and the app do.

BUILT ON WHAT WAS ALREADY HERE, NOT AROUND IT
  docs/research/graphrag_eval/{collect,score}.py  — per-query relevant-id
    scoring and fixed-denominator precision@5 (kept verbatim: an empty result
    should be punished like a page of junk).
  docs/research-archive/p0-prototypes/eval_pinned_40q_20260715.py — the pinned
    ground truth + --check winnability gate, so every run judges alike.
  scripts/verify-soul-contract.sh — the isolation recipe, including the
    non-obvious SOUL_ISE_URL pin without which an "isolated" soul silently
    syncs the operator's live brain.
  gen-soul-amalgam.sh + .gitea/workflows/ci.yaml — the build recipe and flags.
New here: ids rather than regexes as ground truth, an associative category
derived from real edges, a superseded category scored on ranking, a
machine-checked zero-lexical-overlap guarantee on paraphrases, paired
significance testing, and measurement of the real compiled soul rather than an
offline replica of one leg of it.

THE GOLD SET IS AUDITABLE, NOT VIBES
38 queries over the real 78,768-node corpus, each carrying a `derivation`
string, each re-validated by `build_gold_set.py --check`. exact_rare is mined
(document frequency 1). phrase is mined (verbatim scan; >25 matches rejected as
too diffuse). paraphrase is hand-selected then PROVEN to share zero content
words with its target — a leak fails the build, so the category cannot decay
into lexical matching. associative is derived from real hub edges with
lexically-reachable siblings dropped. nonsense is verified absent. superseded
pairs are kept only when both sides survive as distinct nodes.

HONEST ABOUT NOISE
Minimum detectable swing on 38 queries is 6: if every changed query moves the
same way, p = 2*0.5^n first clears 0.05 at n=6. Run-to-run drift is measured,
not assumed — activation is a stateful read, and it shows: main is fully
deterministic across 3 runs, the candidate drifts by 1 query. compare.py
reports "no measurable difference" for anything inside max(6, drift+1).

FIRST VERDICT — feat/recall-through-activation
hit@5 34.3% -> 22.9%, phrase 85.7% -> 28.6%, latency p50 2.81x. Five discordant
pairs, all five against the candidate, none for it; McNemar exact p = 0.0625,
so by the stated rule this is one query short of significant and is reported as
such rather than as a win for main. The latency regression is deterministic and
not in any noise band.

The benefit the branch was written for is absent: associative recall is 0/6 on
BOTH builds. Probed directly, the traversal returns the lexical seed at rank 8
and none of its 12 hub siblings. Two measured corpus facts explain it — only
4,060 of 78,768 nodes (5.2%) carry any edge, and no node has an embedding, so
the fourth factor of the four-factor product has nothing to compute from. The
mechanism runs; the corpus lacks the structure it needs.

SAFETY
Throwaway port, throwaway HOME, disposable per-run copy of the corpus; live
ports refused by name. Every soul started is killed AND confirmed dead by pid
probe, with the confirmation written into the results file; run_comparison.sh
sweeps for strays and exits non-zero if any survive. Nothing under ~/.neuron,
/Applications/Neuron*, or ~/neuron-dev-stack is read, written, or restarted.

Rung: E2E-VERIFIED — 6 full runs (3 per config) against the real compiled
binaries on the real corpus; numbers above are measured, not projected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 13:40:51 -05:00