Compare commits

...

8 Commits

Author SHA1 Message Date
Neuron 88e3008735 runtime: resume the learned stance in think, instead of discarding it
El SDK CI - dev / build-and-test (pull_request) Failing after 4m16s
engram_think_json built a NEUTRAL stance on every call — cog_stance_init
with a NULL id, all axis_gain 1.0, bias_dir NULL, reliability 0.5 — and
never loaded the stance the correspondence-beat had been persisting.

That mattered because the faculty enters engram_think ONLY through the
stance: axis_gain[k] warps the per-axis extents and bias_dir seeds the
steering direction. cog_stance_init stores the faculty NAME and nothing
reads it. So with a neutral stance, reason/abduce/induce/plan/analogize
were byte-identical output under different labels, and confidence was
pinned to 0.5 because GeoGradient.confidence IS stance->reliability.

The machinery already existed and only this call site ignored it.
engram_correspondence_beat_json resumes via cog_stance_from_node and
persists via cog_stance_to_node under "stance-<faculty>-<hub>". Every
beat's calibration was written and then thrown away on the next read.
Same defect as the NULL anchor fixed in #142, one line below: a neutral
argument collapsing a capability to a constant.

Resume the same id the beat writes, so learning compounds across beats and
cold boot. Fall back to neutral only when no stance exists — a genuine
uninformed prior rather than a discarded informed one.

Also emit stance_resumed, so confidence 0.5 from a learned-but-unreliable
stance is distinguishable from confidence 0.5 from "no stance exists".
That reporting gap is what let the neutral stance hide.

Verified against a clone of the production store (13,627 nodes):

  before beat, no stance     stance_resumed=false  confidence=0.5
  beat on a NON-keystone     brier 0.00458568 -> 0.00329654
                             reduction 28.11%, n_trials 6000,
                             reliability 0.930726, stance_written=true
  after beat                 stance_resumed=true   confidence=0.930726

Confidence now equals the learned reliability instead of the uninformed
prior. The keystone self-anchor correctly stays at 0.5 — calibration is
deliberately refused on protected identity regions, and that refusal is
now visible as resumed=true with confidence unchanged, rather than being
indistinguishable from the bug.

STILL OPEN: with no learned bias_dir the faculties remain identical in
direction. What distinguishes abduce from induce geometrically is a
design decision about how Neuron thinks, not a plumbing defect, and is
deliberately left to Will.
2026-08-16 11:43:38 -05:00
will.anderson a6cef4b983 runtime: publish the vector index instead of guarding it (#143)
El SDK CI - dev / build-and-test (push) Failing after 10m29s
2026-08-16 16:33:28 +00:00
Neuron 8e9d88fc01 runtime: publish the vector index instead of guarding it
El SDK CI - dev / build-and-test (pull_request) Failing after 10m59s
The crash (SIGTRAP in engram_activate -> eg_vindex_sync -> vindex_insert ->
_realloc) had three read paths mutating five process-global statics.
engram_activate, eg_knn_for_node (whose own comment says "No writes.") and
engram_geo_reify_run_json all called eg_vindex_sync, which frees the index,
reallocs the seen-map and inserts — on a read.

Three moves, in decreasing order of how much they dissolve:

1. Misfiled scratch is not shared state. visited/visit_epoch/visited_cap
   were never owned by the index; they are one traversal's local, hoisted
   into struct VIndex as an allocation optimisation. They want neither a
   lock nor a capability nor a pool — just to go back in the call frame.
   Two concurrent READS stomped each other purely because of this.

2. const IS the capability. Once the scratch leaves the struct, search
   reads and nothing else, so vindex_search takes a const VIndex*. That is
   exactly what a capability-pointer ABI would have bought — a read path
   physically cannot call vindex_insert, enforced by the compiler on every
   future caller — for one qualifier instead of an ABI swept across
   hundreds of builtins.

3. What survives is publication, not ownership. HNSW insert is NOT an
   append: it rewires the neighbour links of already-existing elements and
   reallocs elems[], so the store's append-only property does not transfer
   to the index derived from it. eg_vindex_sync therefore splits into
   eg_vindex_maintain (exclusive, sole mutator) and eg_vindex_view (shared,
   returns const VIndex*). A read path may demand that a current snapshot
   exist — a request to the owner, not a mutation by the reader.

Write-side owner: eg_vindex_note_embedded hooks the embedding-ASSIGNMENT
sites rather than the append sites, because a node with no embedding cannot
be in a vector index — embedding assignment is the event that owns index
membership. One O(log n) insert, no O(node_count) presence scan. This also
retires the "STALENESS (honest tradeoff)" note where a lazily-embedded
older node stayed invisible to route_nearest/autoconnect until a full
rebuild (the embed-gap #20 shape).

Evidence. The existing harness conflated two hazards, which is why fixing
half of it read as failure. Split into four:

  single (3000 vec, ASan+UBSan)          clean  ->  clean
  readers (4 readers, no writer, TSan)   RACE   ->  clean
  unsynchronized (writer+reader, bare)   race   ->  race, expected forever
  published (owner + 4 readers)          n/a    ->  clean, 3000/3000 landed

RESULT: PASS. recall@10 = 0.9365 at ef_search=128 (gate >= 0.90);
determinism byte-identical across two independent builds.

The unsynchronized half is now permanently expected to race, deliberately:
it is the executable proof that the boundary must live above the data
structure, not inside it.

fb32d15's guard is KEPT, correcting this design's own section 5. Measured,
it guards TWO structures and only one was converted here: g->nodes/g->edges
are realloc'd in place (el_runtime.c:7618,7629) and engram_activate_inner's
embed-backfill writes n->emb through exactly such a borrowed pointer.
Deleting the guard reintroduces a measured 11171->9579 edge loss. Its
comment is narrowed to the RAM graph and the deletion precondition named.

That corrects the ordering claim too: the residual is not one ABI that
dissolves everything at once, it is a PROPERTY applied per structure.
Residues evaporate in the order the property is applied, and a residue
whose structure has not been converted must be left standing.
2026-08-16 11:29:17 -05:00
bigmerge e99a4640e2 test: regression harness for the vindex concurrency crash
Promotes the two throwaway sanitizer harnesses used to diagnose the
2026-08-16 soul crash into engram/test/ so the bug cannot silently regress.

The harness has two halves and the PAIR is the point — it is what localises
the defect to concurrency rather than to HNSW logic:

  single      3000 clustered vectors, one thread, ASan+UBSan. The CONTROL.
              Must always be clean. During diagnosis this cleared all 13,820
              real dim-768 vectors from the live store, which DISPROVED an
              inspection-derived hypothesis about an out-of-bounds
              reverse-link write at engram_vindex.c:340.

  concurrent  writer + reader on one shared index, TSan. Currently reports a
              race at engram_vindex.c:195 (visited_reset) reached from both
              vindex_search and vindex_insert, because VIndex still owns its
              visited[]/visit_epoch scratch — so even two concurrent READS
              corrupt each other's traversal.

Verified: half 1 passes, half 2 reproduces the race.

Gated on EXPECT_RACE, default 1, so the concurrent half documents the known
defect without failing the suite today. When the visited set moves to a
per-query checkout pool (hnswlib VisitedListPool style — NOT thread_local,
since http_worker is a thread per connection and a __thread buffer would leak
~55KB per connection), flip EXPECT_RACE=0 and it becomes a real gate.
2026-08-16 11:29:17 -05:00
bigmerge bdc1f99fb9 runtime: guard engram activation against the unsynchronized awareness thread
The soul daemon had two engram callers and only one of them locked.
soul.el:729 starts the HTTP server via http_serve_async (spawning
http_worker threads); soul.el:731 then runs awareness_run() on the MAIN
thread. awareness.el's perceive() -> engram_activate_json() ->
engram_activate() -> eg_vindex_sync() -> vindex_insert() mutates the same
g->nodes/g->edges and the process-global _eg_vindex HNSW index that the
workers touch. g_engram_req_lock existed to serialize exactly this, but it
was only ever taken inside http_worker: engram_req_lock/engram_req_unlock
appear in ZERO .el sources, so the awareness loop ran lock-free beside the
workers on every tick (SOUL_TICK_MS=1000).

Result was a crash-loop under launchd KeepAlive: five crashes in ~4 minutes
on 2026-08-16 with varying faulting frames -- search_layer<-vindex_insert
<-eg_vindex_sync, engram_activate, abort, and one inside xzm_realloc's own
freelist. Varying sites plus a fault in allocator metadata means heap
corruption. The SIGSEGV address 0x65646f4e6d617267 is little-endian ASCII
"gramNode": string bytes dereferenced as an Elem vector pointer.

Diagnosed by bisection rather than inspection:
  - Replaying all 13,820 real dim-768 vectors harvested from the live store
    through the index single-threaded under ASan is 100% clean, which rules
    out an HNSW logic/bounds bug.
  - Two threads on one index trip ThreadSanitizer immediately at
    engram_vindex.c:195 (visited_reset), reached from both vindex_search and
    vindex_insert. VIndex keeps a SHARED visited-epoch scratch buffer, so
    even two concurrent READS corrupt each other's traversal and walk bogus
    element indices.
So this is purely a concurrency defect, not an HNSW logic error. (An
inspection-derived hypothesis about an out-of-bounds reverse-link write at
engram_vindex.c:340 was disproved by the single-threaded run.)

Fix: a thread-local ownership depth (_eg_req_depth) lets engram entry points
self-guard. engram_activate() becomes a wrapper over engram_activate_inner()
that acquires g_engram_req_lock when called with depth 0 (the awareness
thread) and passes through when depth > 0 (nested inside an http_worker that
already holds it), so the non-recursive mutex cannot self-deadlock. The depth
is a plain counter, never a recursive-mutex count, preserving
engram_self_reify_beat_json's contract of genuinely releasing the lock
mid-beat.
2026-08-16 11:29:17 -05:00
will.anderson 44b621e551 runtime: anchor the think read, so Neuron can think at all (#142)
El SDK CI - dev / build-and-test (push) Failing after 13m0s
2026-08-16 16:26:00 +00:00
Neuron ded6ca546f runtime: anchor the think read, so Neuron can think at all
El SDK CI - dev / build-and-test (pull_request) Failing after 13m24s
engram_think_json passed NULL as the anchor. NULL is not "no opinion":
engram_think re-origins at `anchor ? anchor : region->centroid`, so NULL
means "read from the centroid" — and the centroid is the one point where
the gradient is zero by construction. r = x - centroid = 0, so every axis
projection is 0, grad is 0, and direction takes the "at rest" branch at
engram_cognition.c:137.

Measured consequence: EVERY faculty returned an identical null result,
differing only in its label —
  {"direction":[0,0,0,0,0,0,0,0],"spread":0,"magnitude":1,"confidence":0.5}
magnitude 1 is membership evaluated at the centroid, spread 0 is its
distance to itself, confidence 0.5 is the stance fallback. The geometry was
never at fault: /api/drift computes real values (centroid_sep 0.104,
core_disp 0.045) over the very same 87 members. Neuron could not think
because the read was always taken from the region's own centre.

The seeds choose WHICH region; they must also supply the VANTAGE. Anchor at
the first resolvable embedded seed — the same seed eg_geo_build_desc infers
dim from, so the two can never disagree. One seed still yields a real
gradient because the descriptor expands to that seed's neighbourhood, so
the seed's position is distinct from the neighbourhood centroid.

The vector is COPIED, never borrowed: g->nodes is realloc'd in place on
append, so a borrowed EngramNode* dangles across any concurrent write.

Verified against a clone of the production store (13,616 nodes / 37,865
edges):
  self anchor   n_support 87  magnitude 0.00282  spread 18.79
  values hub    n_support 28  magnitude 0.00318  spread 17.72
with distinct unit direction vectors. Previously both returned the zero
vector with magnitude 1 and spread 0.

STILL OPEN, now isolated by this fix: all five faculties return identical
numbers and confidence stays 0.5, because cog_stance_init is passed NULL
for the stance and the faculty enters the computation only through the
stance's axis_gain[] and bias_dir. The faculty label is inert until a
stance is loaded — which is what learn()'s correspondence-beat calibrates.
Same shape as this bug: a neutral parameter collapsing a capability to a
constant.
2026-08-16 11:25:24 -05:00
will.anderson b5b96c05ed runtime: let signal enter as geometry, not as prose about signal (#141)
El SDK CI - dev / build-and-test (push) Failing after 10m36s
2026-08-16 16:13:28 +00:00
8 changed files with 916 additions and 61 deletions
+107
View File
@@ -0,0 +1,107 @@
#!/usr/bin/env bash
# run_vindex_concurrency_tests.sh — regression harness for the 2026-08-16 soul crash.
#
# Four halves. The SET is the point: it separates two hazards the original two-half
# version conflated, and which have fixes in different files.
#
# 1. single ASan+UBSan, one thread. MUST be clean. Hard failure.
#
# 2. readers TSan, N readers, NO writer. Hazard (a): the visited set used
# to live on the index, so two pure READS stamped each other's
# epoch. Fixed in engram_vindex.c (frame-owned VVisit +
# `const VIndex*` search). MUST be clean. Hard failure.
#
# 3. unsynchronized TSan, writer + reader on a BARE index. Hazard (b): in-place
# HNSW insert rewires existing elements' neighbour lists and
# reallocs elems[]. EXPECTED TO RACE, PERMANENTLY. This is not
# a bug to fix inside engram_vindex.c — it is the executable
# proof that a publication boundary must exist above it.
# Not a failure. If it ever goes CLEAN, the test stopped
# interleaving and half 4 is no longer meaningful either.
#
# 4. published TSan, owner + N readers through a publication boundary
# (rwlock: readers shared, owner exclusive) mirroring
# eg_vindex_view / eg_vindex_maintain in lang/runtime/el_runtime.c.
# MUST be clean, and all inserts must land. Hard failure.
#
# See test_vindex_concurrency.c for the full story (SIGSEGV at ASCII address
# "gramNode", heap corruption in xzm_realloc, etc).
#
# usage: run_vindex_concurrency_tests.sh
set -uo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
RUNTIME="$(cd "$HERE/../../lang/runtime" && pwd)"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
SRC="$HERE/test_vindex_concurrency.c"
VINDEX="$RUNTIME/engram_vindex.c"
fail=0
echo "== [1/4] single-threaded control under AddressSanitizer =="
cc -std=c11 -g -O1 -fsanitize=address,undefined -fno-omit-frame-pointer \
-I"$RUNTIME" -o "$WORK/single" "$SRC" "$VINDEX" -lm || { echo "BUILD FAILED"; exit 2; }
if ASAN_OPTIONS=detect_leaks=0 "$WORK/single" single; then
echo " -> OK"
else
echo " -> FAIL: the single-threaded control must always be clean."
echo " If this fails the bug is NOT (only) concurrency — look for a real"
echo " out-of-bounds or lifetime error in engram_vindex.c."
fail=1
fi
cc -std=c11 -g -O1 -fsanitize=thread -fno-omit-frame-pointer \
-I"$RUNTIME" -o "$WORK/conc" "$SRC" "$VINDEX" -lm || { echo "BUILD FAILED"; exit 2; }
# run_tsan <mode> <logfile>; echoes nothing, sets $tsan_raced
run_tsan() {
TSAN_OPTIONS="halt_on_error=0" "$WORK/conc" "$1" >"$2" 2>&1
tsan_rc=$?
if grep -q "ThreadSanitizer: data race" "$2"; then tsan_raced=1; else tsan_raced=0; fi
}
echo
echo "== [2/4] concurrent READERS, no writer (visited-set gate) =="
run_tsan readers "$WORK/readers.log"
if [ "$tsan_raced" = "1" ]; then
echo " -> REGRESSION: two concurrent reads still race."
grep -m1 -A6 "ThreadSanitizer: data race" "$WORK/readers.log" | sed 's/^/ /'
echo " The visited set was supposed to be owned by the call frame."
fail=1
else
echo " -> clean (concurrent reads are safe)"
fi
echo
echo "== [3/4] writer+reader on a BARE index (expected-race probe) =="
run_tsan unsynchronized "$WORK/unsync.log"
if [ "$tsan_raced" = "1" ]; then
echo " -> RACE DETECTED, as expected:"
grep -m1 -A4 "ThreadSanitizer: data race" "$WORK/unsync.log" | sed 's/^/ /'
echo " In-place HNSW insert mutates existing elements. Not fixable inside"
echo " engram_vindex.c — this is why the publication boundary exists."
else
echo " -> NOTE: no race reported. The probe did not interleave; half 4's"
echo " clean result proves less than it should. Investigate."
fi
echo
echo "== [4/4] owner+readers through the publication boundary (boundary gate) =="
run_tsan published "$WORK/pub.log"
if [ "$tsan_raced" = "1" ]; then
echo " -> REGRESSION: the publication boundary did not serialize the owner."
grep -m1 -A6 "ThreadSanitizer: data race" "$WORK/pub.log" | sed 's/^/ /'
fail=1
elif [ "$tsan_rc" != "0" ]; then
echo " -> FAIL: boundary clean under TSan but the run failed:"
tail -3 "$WORK/pub.log" | sed 's/^/ /'
fail=1
else
echo " -> clean (readers project concurrently; the owner's inserts all landed)"
fi
echo
[ "$fail" -eq 0 ] && echo "RESULT: PASS" || echo "RESULT: FAIL"
exit "$fail"
+251
View File
@@ -0,0 +1,251 @@
/* test_vindex_concurrency.c — regression test for the 2026-08-16 soul crash.
*
* WHAT BROKE: the soul daemon crash-looped (5 crashes in ~100s) with SIGSEGV in
* search_layer <- vindex_insert <- eg_vindex_sync, a SIGABRT, and a fault inside
* xzm_realloc's own freelist — i.e. heap corruption. The SIGSEGV address
* 0x65646f4e6d617267 is little-endian ASCII "gramNode": string bytes being
* dereferenced as an Elem vector pointer.
*
* ROOT CAUSE: VIndex owns its traversal scratch (visited[] + visit_epoch), and
* search_layer mutates it via visited_reset(). So the index is unsafe for ANY
* concurrent use — including two concurrent READS. soul.el starts http_serve_async
* (a thread per connection) and then runs awareness_run() on the main thread, which
* reaches the same global index through engram_activate; nothing serialized them.
*
* Neither hnswlib nor FAISS puts the visited set on the index: hnswlib checks one
* out of a VisitedListPool per query, FAISS uses a thread_local VisitedTable.
*
* THE ORIGINAL `concurrent` HALF CONFLATED TWO DISTINCT HAZARDS (2026-08-16). It ran
* a writer against a reader on one bare index, so it could not tell apart:
*
* (a) READ/READ corruption — two searches stamping each other's visited epoch.
* A defect INSIDE engram_vindex.c, fixable there, and now fixed: the visited
* set moved to the call frame and vindex_search takes a `const VIndex*`.
*
* (b) WRITE/READ corruption — vindex_insert rewires the neighbour lists of
* EXISTING elements and reallocs elems[], so an insert is a mutation of the
* whole structure. This is NOT fixable inside engram_vindex.c at any price:
* it is inherent to in-place HNSW. It requires a publication boundary ABOVE
* the data structure (el_runtime.c: eg_vindex_view / eg_vindex_maintain).
*
* Conflating them made the suite unfailable-then-unpassable: fixing (a) left (b)
* still racing, which reads as "the fix did not work" when in fact a different,
* correctly-located fix is what (b) needs. So the halves are now separate:
*
* single N clustered vectors, ONE thread, ASan. The CONTROL. Must always
* be clean. When this passes and a concurrent half fails, the defect
* is concurrency, not an out-of-bounds/logic error in the graph code.
* (On 2026-08-16 this control cleared all 13,820 real dim-768 store
* vectors under ASan, which DISPROVED an inspection-derived hypothesis
* about an out-of-bounds reverse-link write at engram_vindex.c:340.)
*
* readers N reader threads, NO writer, one shared index, TSan. This is
* hazard (a) in isolation. It RACED before the visited set moved off
* the index struct and must be CLEAN now. Hard gate.
*
* unsynchronized writer + reader on a bare index, TSan. Hazard (b) in isolation.
* EXPECTED TO RACE, permanently — it is the executable proof that
* the index cannot be made safe from the inside, and therefore that
* the publication boundary in el_runtime.c has to exist. If this
* ever goes clean, the test stopped interleaving; do not celebrate.
*
* published writer + readers through a publication boundary that mirrors
* eg_vindex_view / eg_vindex_maintain (rwlock: readers shared,
* the single owner exclusive), TSan. Must be CLEAN. Hard gate.
* This is what proves the shape of the runtime fix, in the same
* process, rather than asserting it.
*
* Absence of a crash does NOT mean absence of a race — always read the sanitizer
* verdict, never just the exit code.
*
* Build/run: engram/test/run_vindex_concurrency_tests.sh
*/
#include "engram_vindex.h"
#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdint.h>
#define DIM 128
#define NVEC 3000
#define SEED_N 50
static VIndex* g_ix;
static float* g_vecs;
/* Deterministic filler. Real embeddings are strongly correlated, not uniform noise;
* clustering keeps many candidates near-equidistant, which exercises the diversity
* heuristic and the visited set far harder than random vectors do. */
static void fill_vectors(void) {
g_vecs = (float*)malloc((size_t)NVEC * DIM * sizeof(float));
if (!g_vecs) { fprintf(stderr, "OOM\n"); exit(1); }
for (int i = 0; i < NVEC; i++) {
int cluster = i % 8;
for (int d = 0; d < DIM; d++)
g_vecs[(size_t)i * DIM + d] =
(float)(((d + cluster * 7) % 13) / 13.0) +
(float)(((i * 2654435761u + (unsigned)d) % 97) / 9700.0);
}
}
static void* writer_fn(void* arg) {
(void)arg;
for (int i = SEED_N; i < NVEC; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
return NULL;
}
static void* reader_fn(void* arg) {
(void)arg;
uint64_t ids[8]; float ds[8];
for (int i = 0; i < 20000; i++)
(void)vindex_search(g_ix, g_vecs + (size_t)(i % NVEC) * DIM, 8, 0, ids, ds);
return NULL;
}
static int run_single(void) {
printf("[single] inserting %d vectors on one thread (ASan control)\n", NVEC);
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[single] vindex_create failed\n"); return 1; }
for (int i = 0; i < NVEC; i++) {
if (vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM) != 0) {
fprintf(stderr, "[single] insert %d failed\n", i); return 1;
}
}
if (vindex_size(g_ix) != (size_t)NVEC) {
fprintf(stderr, "[single] size %zu != %d\n", vindex_size(g_ix), NVEC); return 1;
}
uint64_t ids[16]; float ds[16];
for (int q = 0; q < 200; q++) {
int k = vindex_search(g_ix, g_vecs + (size_t)((q * 7) % NVEC) * DIM, 16, 0, ids, ds);
if (k < 0) { fprintf(stderr, "[single] search failed at q=%d\n", q); return 1; }
}
vindex_free(g_ix); g_ix = NULL;
printf("[single] PASS — no memory error (this must ALWAYS pass)\n");
return 0;
}
/* Hazard (b) in isolation: writer + reader on a BARE index, no boundary. */
static int run_unsynchronized(void) {
printf("[unsynchronized] 1 writer + 1 reader on a BARE index (TSan probe)\n");
printf("[unsynchronized] a race here is EXPECTED and PERMANENT — in-place HNSW\n");
printf("[unsynchronized] insert rewires existing elements. This is the proof that\n");
printf("[unsynchronized] the publication boundary must live ABOVE engram_vindex.c.\n");
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[unsynchronized] vindex_create failed\n"); return 1; }
for (int i = 0; i < SEED_N; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_t w, r;
if (pthread_create(&w, NULL, writer_fn, NULL) ||
pthread_create(&r, NULL, reader_fn, NULL)) {
fprintf(stderr, "[unsynchronized] pthread_create failed\n"); return 1;
}
pthread_join(w, NULL);
pthread_join(r, NULL);
vindex_free(g_ix); g_ix = NULL;
printf("[unsynchronized] completed — CHECK THE SANITIZER VERDICT, not this line.\n");
return 0;
}
/* ── hazard (a) in isolation: concurrent READS only ───────────────────────────
* This is what the frame-owned visited set fixes. Before that change, two
* vindex_search calls on one index wrote each other's epoch stamp; TSan reported
* the race at visited_reset and the traversal then walked bogus element indices. */
#define NREADERS 4
static int run_readers(void) {
printf("[readers] %d concurrent readers, NO writer, one shared index (TSan)\n", NREADERS);
printf("[readers] this is the visited-set regression gate — must be CLEAN.\n");
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[readers] vindex_create failed\n"); return 1; }
for (int i = 0; i < NVEC; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_t t[NREADERS];
for (int i = 0; i < NREADERS; i++)
if (pthread_create(&t[i], NULL, reader_fn, NULL)) {
fprintf(stderr, "[readers] pthread_create failed\n"); return 1;
}
for (int i = 0; i < NREADERS; i++) pthread_join(t[i], NULL);
vindex_free(g_ix); g_ix = NULL;
printf("[readers] completed — CHECK THE SANITIZER VERDICT, not this line.\n");
return 0;
}
/* ── the publication boundary, mirroring el_runtime.c ─────────────────────────
* Readers take the boundary SHARED and hold it across the whole search; the one
* owner takes it EXCLUSIVE to extend. Same shape as eg_vindex_view /
* eg_vindex_maintain. Note the reader's index pointer is `const VIndex*` — the
* compiler, not this comment, is what stops a reader inserting. */
static pthread_rwlock_t g_pub = PTHREAD_RWLOCK_INITIALIZER;
static void* pub_writer_fn(void* arg) {
(void)arg;
for (int i = SEED_N; i < NVEC; i++) {
pthread_rwlock_wrlock(&g_pub);
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_rwlock_unlock(&g_pub);
}
return NULL;
}
static void* pub_reader_fn(void* arg) {
(void)arg;
uint64_t ids[8]; float ds[8];
for (int i = 0; i < 5000; i++) {
pthread_rwlock_rdlock(&g_pub);
const VIndex* view = g_ix; /* immutable view */
(void)vindex_search(view, g_vecs + (size_t)(i % NVEC) * DIM, 8, 0, ids, ds);
pthread_rwlock_unlock(&g_pub);
}
return NULL;
}
static int run_published(void) {
printf("[published] 1 owner + %d readers through a publication boundary (TSan)\n", NREADERS);
printf("[published] this is the eg_vindex_view/eg_vindex_maintain gate — must be CLEAN.\n");
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[published] vindex_create failed\n"); return 1; }
for (int i = 0; i < SEED_N; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_t w, r[NREADERS];
if (pthread_create(&w, NULL, pub_writer_fn, NULL)) {
fprintf(stderr, "[published] pthread_create failed\n"); return 1;
}
for (int i = 0; i < NREADERS; i++)
if (pthread_create(&r[i], NULL, pub_reader_fn, NULL)) {
fprintf(stderr, "[published] pthread_create failed\n"); return 1;
}
pthread_join(w, NULL);
for (int i = 0; i < NREADERS; i++) pthread_join(r[i], NULL);
if (vindex_size(g_ix) != (size_t)NVEC) {
fprintf(stderr, "[published] size %zu != %d — the owner lost inserts\n",
vindex_size(g_ix), NVEC);
vindex_free(g_ix); g_ix = NULL; return 1;
}
vindex_free(g_ix); g_ix = NULL;
printf("[published] all %d inserts landed; CHECK THE SANITIZER VERDICT too.\n", NVEC);
return 0;
}
int main(int argc, char** argv) {
const char* mode = (argc > 1) ? argv[1] : "single";
fill_vectors();
int rc;
if (!strcmp(mode, "single")) rc = run_single();
else if (!strcmp(mode, "readers")) rc = run_readers();
else if (!strcmp(mode, "unsynchronized")) rc = run_unsynchronized();
else if (!strcmp(mode, "published")) rc = run_published();
/* back-compat: the pre-split name meant the bare writer+reader probe. */
else if (!strcmp(mode, "concurrent")) rc = run_unsynchronized();
else {
fprintf(stderr, "usage: %s [single|readers|unsynchronized|published]\n", argv[0]);
rc = 2;
}
free(g_vecs);
return rc;
}
+297 -21
View File
@@ -1600,8 +1600,64 @@ typedef struct {
* no longer blocks ingest/reads (measured: non-health latency during a beat * no longer blocks ingest/reads (measured: non-health latency during a beat
* 13.9s sub-second). */ * 13.9s sub-second). */
static pthread_mutex_t g_engram_req_lock = PTHREAD_MUTEX_INITIALIZER; static pthread_mutex_t g_engram_req_lock = PTHREAD_MUTEX_INITIALIZER;
void engram_req_unlock(void){ pthread_mutex_unlock(&g_engram_req_lock); }
void engram_req_lock(void){ pthread_mutex_lock(&g_engram_req_lock); } /* ── AWARENESS-THREAD GUARD (2026-08-16 self-review) ─────────────────────────
* The request lock above serialized http_worker threads against EACH OTHER, but
* the soul daemon has a SECOND, unsynchronized engram caller: soul.el starts the
* HTTP server with http_serve_async (spawning worker threads) and then runs
* awareness_run() on the MAIN thread, whose perceive() -> engram_activate_json()
* -> engram_activate() -> eg_vindex_sync() path mutates the very same RAM graph
* and the process-global _eg_vindex HNSW index. Nothing in any .el source ever
* called engram_req_lock, so that whole loop ran lock-free beside the workers.
*
* Measured consequence (2026-08-16): five crashes in ~4 minutes, all one bug
* SIGSEGV in search_layer<-vindex_insert<-eg_vindex_sync at address
* 0x65646f4e6d617267 (little-endian ASCII "gramNode": a string being
* dereferenced as an Elem vector pointer), plus a SIGABRT and a fault inside
* xzm_realloc's freelist, i.e. corrupted allocator metadata. Confirmed by
* bisection: replaying ALL 13,820 real dim-768 store vectors through the index
* single-threaded under ASan is 100% clean, while two threads on one index trip
* ThreadSanitizer instantly at engram_vindex.c:195 (visited_reset) VIndex keeps
* a SHARED visited-epoch scratch buffer, so even two concurrent READS stomp each
* other's traversal state and walk bogus element indices. So this is purely a
* concurrency defect, not a logic error in the HNSW code.
*
* Fix: a thread-local ownership depth lets engram entry points self-guard. A call
* arriving on the awareness thread (depth 0) acquires the lock; one arriving from
* inside an http_worker that already holds it (depth > 0) is a no-op, so there is
* no self-deadlock on this NON-recursive mutex. Depth is a plain counter, never a
* recursive-mutex count, which preserves engram_self_reify_beat_json's contract of
* really releasing the lock mid-beat (see engram_req_unlock at the reify beat).
*
* SCOPE NARROWED (2026-08-16, vindex publication boundary): this guard originally
* covered TWO hazards the RAM graph AND the process-global _eg_vindex. The vindex
* half is retired: the index now has its own publication boundary (_eg_vindex_rw),
* search takes a `const VIndex*`, and no read path can mutate the index at all.
*
* What REMAINS load-bearing here is the RAM graph alone, and it is a genuine,
* measured hazard independent of the index: g->nodes / g->edges are realloc'd in
* place (el_runtime.c:7618, 7629), so an awareness-thread reader holding
* `EngramNode* n = &g->nodes[i]` across a concurrent append from an http_worker
* holds a dangling pointer and engram_activate_inner's embed-backfill WRITES
* n->emb through exactly such a pointer. That is a separate residue with its own
* fix (the resident graph wants the same publication treatment the index just got);
* until it lands, this guard stays. Do NOT delete it as "the fb32d15 vindex lock". */
static __thread int _eg_req_depth = 0;
void engram_req_unlock(void){ if(_eg_req_depth > 0) _eg_req_depth--; pthread_mutex_unlock(&g_engram_req_lock); }
void engram_req_lock(void){ pthread_mutex_lock(&g_engram_req_lock); _eg_req_depth++; }
/* Acquire only if this thread does not already hold the request lock.
* Returns 1 if this call took ownership (caller must release), 0 if nested. */
static int eg_guard_enter(void){
if (_eg_req_depth > 0) return 0;
pthread_mutex_lock(&g_engram_req_lock);
_eg_req_depth++;
return 1;
}
static void eg_guard_exit(int owned){
if (!owned) return;
if (_eg_req_depth > 0) _eg_req_depth--;
pthread_mutex_unlock(&g_engram_req_lock);
}
static void* http_worker(void* arg) { static void* http_worker(void* arg) {
HttpWorkerArg* a = (HttpWorkerArg*)arg; HttpWorkerArg* a = (HttpWorkerArg*)arg;
@@ -1642,7 +1698,7 @@ static void* http_worker(void* arg) {
(plen == 1 && path[0] == '/')) (plen == 1 && path[0] == '/'))
health_exempt = 1; health_exempt = 1;
} }
if (!health_exempt) pthread_mutex_lock(&g_engram_req_lock); if (!health_exempt) engram_req_lock(); /* tracks _eg_req_depth for eg_guard_enter */
if (h) { if (h) {
el_val_t r = h(EL_STR(dispatch_method), EL_STR(path), EL_STR(body)); el_val_t r = h(EL_STR(dispatch_method), EL_STR(path), EL_STR(body));
const char* rs = EL_CSTR(r); const char* rs = EL_CSTR(r);
@@ -1669,7 +1725,7 @@ static void* http_worker(void* arg) {
} }
/* end of the engram critical section — the response is now a private malloc'd /* end of the engram critical section — the response is now a private malloc'd
* copy; arena teardown + socket write touch no shared engram state. */ * copy; arena teardown + socket write touch no shared engram state. */
if (!health_exempt) pthread_mutex_unlock(&g_engram_req_lock); if (!health_exempt) engram_req_unlock();
el_request_end(); /* free all intermediate strings */ el_request_end(); /* free all intermediate strings */
_tl_http_head_only = head_only; _tl_http_head_only = head_only;
http_send_response(fd, response); http_send_response(fd, response);
@@ -9564,6 +9620,35 @@ static double engram_goal_bias(const EngramNode* n, const char* query) {
* the exact O(n) argmax scan tops up any seed slot the ANN leaves unfilled. * the exact O(n) argmax scan tops up any seed slot the ANN leaves unfilled.
* Single-threaded, matching the adjacent query-embedding cache (no lock). * Single-threaded, matching the adjacent query-embedding cache (no lock).
* Returns NULL when no index is available caller falls back to the O(n) scan. */ * Returns NULL when no index is available caller falls back to the O(n) scan. */
/* ── VINDEX PUBLICATION BOUNDARY (2026-08-16) ────────────────────────────────
* The index is DERIVED GEOMETRY: a projection of the store's embeddings. The
* store is append-only and superseding, so a reader must be able to project
* against geometry that does not move under it.
*
* The HNSW index is NOT itself append-only: vindex_insert rewires the neighbour
* lists of ALREADY-EXISTING elements and reallocs elems[]. So "extend" is a
* mutation of the whole structure, and a reader holding element pointers across
* one is unsafe no matter how pure search itself is (measured: TSan reports the
* elems[] race even after the visited set moved to the call frame).
*
* Hence a publication boundary rather than an ownership discipline:
*
* - eg_vindex_maintain() is the ONLY mutator of the five statics below. It
* takes _eg_vindex_rw EXCLUSIVELY, so it never runs beside a reader.
* - eg_vindex_view() hands back a `const VIndex*` with the boundary held for
* READ. N readers project concurrently; none can mutate, because search
* takes a const index and the compiler enforces it.
*
* A read path may DEMAND that a current snapshot exist that is a request to
* the owner, not a mutation by the reader. What it may not do is mutate the
* geometry it is projecting against. eg_vindex_view/eg_vindex_maintain is
* exactly that split.
*
* Lock ordering: request-outer -> vindex -> store-inner. The vindex boundary is
* never held across a call that can re-enter eg_vindex_view/maintain (verified:
* the four read regions each acquire, search, release without nesting). */
static pthread_rwlock_t _eg_vindex_rw = PTHREAD_RWLOCK_INITIALIZER;
static VIndex* _eg_vindex = NULL; static VIndex* _eg_vindex = NULL;
static int32_t _eg_vindex_dim = 0; static int32_t _eg_vindex_dim = 0;
static int64_t _eg_vindex_built_nc = 0; /* g->node_count at last (re)build */ static int64_t _eg_vindex_built_nc = 0; /* g->node_count at last (re)build */
@@ -9583,8 +9668,10 @@ static int eg_vindex_seen_ensure(int64_t need) {
return 0; return 0;
} }
static VIndex* eg_vindex_sync(EngramStore* g, int32_t dim) { /* THE OWNER. The only function that mutates _eg_vindex* — must be called with
if (!g || dim <= 0) return _eg_vindex; * _eg_vindex_rw held EXCLUSIVELY (see eg_vindex_maintain, the sole caller). */
static void eg_vindex_publish_locked(EngramStore* g, int32_t dim) {
if (!g || dim <= 0) return;
/* Drop a stale index: embedder dim changed, or the resident array shrank /* Drop a stale index: embedder dim changed, or the resident array shrank
* (indices may have been reused/reordered cached node_ids unsafe). */ * (indices may have been reused/reordered cached node_ids unsafe). */
if (_eg_vindex && (_eg_vindex_dim != dim || g->node_count < _eg_vindex_built_nc)) { if (_eg_vindex && (_eg_vindex_dim != dim || g->node_count < _eg_vindex_built_nc)) {
@@ -9594,8 +9681,8 @@ static VIndex* eg_vindex_sync(EngramStore* g, int32_t dim) {
} }
if (!_eg_vindex) { if (!_eg_vindex) {
VIndex* idx = vindex_create((int)dim, 0, 0); VIndex* idx = vindex_create((int)dim, 0, 0);
if (!idx) return NULL; if (!idx) return;
if (eg_vindex_seen_ensure(g->node_count)) { vindex_free(idx); return NULL; } if (eg_vindex_seen_ensure(g->node_count)) { vindex_free(idx); return; }
for (int64_t i = 0; i < g->node_count; i++) { for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i]; EngramNode* n = &g->nodes[i];
if (n->emb && n->emb_dim == dim && vindex_insert(idx, (uint64_t)i, n->emb) == 0) if (n->emb && n->emb_dim == dim && vindex_insert(idx, (uint64_t)i, n->emb) == 0)
@@ -9619,8 +9706,62 @@ static VIndex* eg_vindex_sync(EngramStore* g, int32_t dim) {
} }
_eg_vindex_built_nc = g->node_count; _eg_vindex_built_nc = g->node_count;
} }
}
/* Owner-mediated publish. Takes the boundary EXCLUSIVELY, so it can never run
* beside a reader. Cheap no-op when the published snapshot is already current. */
static void eg_vindex_maintain(EngramStore* g, int32_t dim) {
if (!g || dim <= 0) return;
pthread_rwlock_wrlock(&_eg_vindex_rw);
eg_vindex_publish_locked(g, dim);
pthread_rwlock_unlock(&_eg_vindex_rw);
}
/* READ SIDE. Returns the published snapshot as an IMMUTABLE view, with the
* boundary held for READ the caller MUST pair every call with exactly one
* eg_vindex_view_release(), on every path including error returns.
*
* The returned pointer is `const`: a read path physically cannot call
* vindex_insert on it. That is the compile-time constraint, and it is why this
* replaces eg_vindex_sync rather than wrapping it. May return NULL (no index
* available -> caller falls back to the exact O(n) scan); the boundary is still
* held and still must be released. */
static const VIndex* eg_vindex_view(EngramStore* g, int32_t dim) {
if (g && dim > 0) {
/* Fast path: snapshot already current, take it read-only and go. */
pthread_rwlock_rdlock(&_eg_vindex_rw);
if (_eg_vindex && _eg_vindex_dim == dim && _eg_vindex_built_nc == g->node_count)
return _eg_vindex;
/* Stale or absent. Drop to no lock, ask the owner to publish, re-acquire.
* NEVER upgrade rdlock->wrlock in place: that self-deadlocks. */
pthread_rwlock_unlock(&_eg_vindex_rw);
eg_vindex_maintain(g, dim);
}
pthread_rwlock_rdlock(&_eg_vindex_rw);
return _eg_vindex; return _eg_vindex;
} }
static void eg_vindex_view_release(void) {
pthread_rwlock_unlock(&_eg_vindex_rw);
}
/* WRITE-SIDE MAINTENANCE HOOK. Call after an embedding becomes present on a
* resident ordinal. A node without an embedding cannot be in a vector index at
* all, so embedding-assignment not node append is the event that owns index
* membership. Cheap: one O(log n) HNSW insert, no O(node_count) presence scan.
* A no-op before the first publish (the cold build picks the node up) and on a
* dim mismatch. */
static void eg_vindex_note_embedded(EngramStore* g, int64_t ordinal) {
if (!g || ordinal < 0 || ordinal >= g->node_count) return;
EngramNode* n = &g->nodes[ordinal];
if (!n->emb || n->emb_dim <= 0) return;
pthread_rwlock_wrlock(&_eg_vindex_rw);
if (_eg_vindex && _eg_vindex_dim == n->emb_dim &&
eg_vindex_seen_ensure(g->node_count) == 0 && !_eg_vindex_seen[ordinal]) {
if (vindex_insert(_eg_vindex, (uint64_t)ordinal, n->emb) == 0)
_eg_vindex_seen[ordinal] = 1;
}
pthread_rwlock_unlock(&_eg_vindex_rw);
}
/* ── M9 GEOMETRY PRIMING (ENGRAM_GEOMETRY_PRIMING, default OFF) ────────────── /* ── M9 GEOMETRY PRIMING (ENGRAM_GEOMETRY_PRIMING, default OFF) ──────────────
* Opt-in wiring of the centered relational-neighborhood geometry (engram_geometry.c) * Opt-in wiring of the centered relational-neighborhood geometry (engram_geometry.c)
@@ -9727,7 +9868,9 @@ static int64_t engram_activate_beam(void) {
v = d; return v; v = d; return v;
} }
el_val_t engram_activate(el_val_t query, el_val_t depth) { /* Core activation. Callers must hold the engram request lock — reached only via
* the engram_activate() wrapper below, which self-guards (see eg_guard_enter). */
static el_val_t engram_activate_inner(el_val_t query, el_val_t depth) {
EngramStore* g = engram_get(); EngramStore* g = engram_get();
const char* q = EL_CSTR(query); const char* q = EL_CSTR(query);
int64_t max_depth = (int64_t)depth; if (max_depth <= 0) max_depth = 2; int64_t max_depth = (int64_t)depth; if (max_depth <= 0) max_depth = 2;
@@ -9768,6 +9911,12 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
float* v = eg_embed_fetch(n->content, &d); float* v = eg_embed_fetch(n->content, &d);
if (!v) break; /* embedder down / breaker open — stop this call */ if (!v) break; /* embedder down / breaker open — stop this call */
n->emb = v; n->emb_dim = d; n->emb = v; n->emb_dim = d;
/* Write-side index maintenance: an embedding just became present on
* ordinal i, so the index's owner publishes it now. This is what
* retires the "STALENESS (honest tradeoff)" note above a lazily
* embedded OLDER node no longer waits for a full rebuild to become
* visible to route_nearest / autoconnect. */
eg_vindex_note_embedded(g, i);
backfilled++; backfilled++;
} }
} }
@@ -9966,7 +10115,9 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
* same budget as the exact scan's retry `guard` so dedup/threshold * same budget as the exact scan's retry `guard` so dedup/threshold
* rejects still leave enough distinct seeds. */ * rejects still leave enough distinct seeds. */
{ {
VIndex* vx = eg_vindex_sync(g, q_dim); /* Immutable view: the boundary is held for READ across the whole
* search + harvest, and released at the end of this block. */
const VIndex* vx = eg_vindex_view(g, q_dim);
if (vx && (int64_t)vindex_size(vx) >= ENGRAM_EMBED_SEED_K) { if (vx && (int64_t)vindex_size(vx) >= ENGRAM_EMBED_SEED_K) {
const float* seed_qv = e_eff ? e_eff : q_emb; const float* seed_qv = e_eff ? e_eff : q_emb;
int kreq = ENGRAM_EMBED_SEED_K * 8; int kreq = ENGRAM_EMBED_SEED_K * 8;
@@ -10010,6 +10161,7 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
} }
free(aid); free(ad); free(aid); free(ad);
} }
eg_vindex_view_release();
} }
/* Exact O(n) argmax fallback / top-up (pre-M8 selection, verbatim). /* Exact O(n) argmax fallback / top-up (pre-M8 selection, verbatim).
@@ -10096,9 +10248,11 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
char** vids = malloc((size_t)g->node_count * sizeof(char*)); char** vids = malloc((size_t)g->node_count * sizeof(char*));
if (gmean && vids) { if (gmean && vids) {
for (int64_t i = 0; i < g->node_count; i++) vids[i] = g->nodes[i].id; for (int64_t i = 0; i < g->node_count; i++) vids[i] = g->nodes[i].id;
const VIndex* gvx = eg_vindex_view(g, q_dim);
geo = engram_geometry_descriptor( geo = engram_geometry_descriptor(
g_engram_store, _eg_vindex, vids, (int)g->node_count, g_engram_store, gvx, vids, (int)g->node_count,
seed_ids, (size_t)nsel, NULL, gmean); seed_ids, (size_t)nsel, NULL, gmean);
eg_vindex_view_release();
} }
free(vids); free(vids);
if (geo && geo->n_members > 0) { if (geo && geo->n_members > 0) {
@@ -13248,13 +13402,16 @@ static int eg_knn_for_node(EngramStore* g, int64_t self, int want, uint64_t* out
if(self < 0 || self >= g->node_count) return 0; if(self < 0 || self >= g->node_count) return 0;
EngramNode* n = &g->nodes[self]; EngramNode* n = &g->nodes[self];
if(!n->emb || n->emb_dim <= 0) return 0; if(!n->emb || n->emb_dim <= 0) return 0;
VIndex* vx = eg_vindex_sync(g, n->emb_dim); /* Immutable view held for READ across the search; the harvest below reads
if(!vx) return 0; * only g->nodes, so the boundary is released as soon as the search returns. */
const VIndex* vx = eg_vindex_view(g, n->emb_dim);
if(!vx){ eg_vindex_view_release(); return 0; }
int K = want + 8; int K = want + 8;
uint64_t* ids = (uint64_t*)malloc(sizeof(uint64_t)*(size_t)K); uint64_t* ids = (uint64_t*)malloc(sizeof(uint64_t)*(size_t)K);
float* dist = (float*)malloc(sizeof(float)*(size_t)K); float* dist = (float*)malloc(sizeof(float)*(size_t)K);
if(!ids || !dist){ free(ids); free(dist); return 0; } if(!ids || !dist){ eg_vindex_view_release(); free(ids); free(dist); return 0; }
int m = vindex_search(vx, n->emb, K, 0, ids, dist); int m = vindex_search(vx, n->emb, K, 0, ids, dist);
eg_vindex_view_release();
int c = 0; int c = 0;
for(int j=0; j<m && c<want; j++){ for(int j=0; j<m && c<want; j++){
int64_t bi = (int64_t)ids[j]; int64_t bi = (int64_t)ids[j];
@@ -13285,7 +13442,8 @@ el_val_t engram_autoconnect_node(el_val_t id_v, el_val_t k_v, el_val_t minsim_v)
EngramNode* n = &g->nodes[self]; EngramNode* n = &g->nodes[self];
if((!n->emb || n->emb_dim <= 0) && n->content && eg_embed_eligible(n)){ if((!n->emb || n->emb_dim <= 0) && n->content && eg_embed_eligible(n)){
int32_t d = 0; float* v = eg_embed_fetch(n->content, &d); int32_t d = 0; float* v = eg_embed_fetch(n->content, &d);
if(v && d > 0){ n->emb = v; n->emb_dim = d; if(engram_store_enabled()) eg_store_put_node(n); } if(v && d > 0){ n->emb = v; n->emb_dim = d; if(engram_store_enabled()) eg_store_put_node(n);
eg_vindex_note_embedded(g, self); }
else free(v); else free(v);
} }
if(!n->emb || n->emb_dim <= 0){ jb_puts(&b, "{\"connected\":0,\"reason\":\"unembedded\"}"); return el_wrap_str(b.buf); } if(!n->emb || n->emb_dim <= 0){ jb_puts(&b, "{\"connected\":0,\"reason\":\"unembedded\"}"); return el_wrap_str(b.buf); }
@@ -13420,8 +13578,10 @@ static GeoDescriptor* eg_geo_build_desc(const char* csv) {
char** vids = malloc((size_t)g->node_count * sizeof(char*)); char** vids = malloc((size_t)g->node_count * sizeof(char*));
if (gmean && vids) { if (gmean && vids) {
for (int64_t i = 0; i < g->node_count; i++) vids[i] = g->nodes[i].id; for (int64_t i = 0; i < g->node_count; i++) vids[i] = g->nodes[i].id;
geo = engram_geometry_descriptor(g_engram_store, _eg_vindex, vids, (int)g->node_count, const VIndex* gvx = eg_vindex_view(g, dim);
geo = engram_geometry_descriptor(g_engram_store, gvx, vids, (int)g->node_count,
(const char* const*)ids, (size_t)ns, NULL, gmean); (const char* const*)ids, (size_t)ns, NULL, gmean);
eg_vindex_view_release();
} }
free(vids); free(vids);
for (int i = 0; i < ns; i++) free(ids[i]); for (int i = 0; i < ns; i++) free(ids[i]);
@@ -13458,12 +13618,16 @@ el_val_t engram_geo_reify_run_json(void){
int32_t dim = 0; int32_t dim = 0;
for(int64_t i = 0; i < g->node_count && dim == 0; i++) for(int64_t i = 0; i < g->node_count && dim == 0; i++)
if(g->nodes[i].emb && g->nodes[i].emb_dim > 0) dim = g->nodes[i].emb_dim; if(g->nodes[i].emb && g->nodes[i].emb_dim > 0) dim = g->nodes[i].emb_dim;
VIndex* vx = (dim > 0) ? eg_vindex_sync(g, dim) : NULL;
char** vids = malloc((size_t)g->node_count * sizeof(char*)); char** vids = malloc((size_t)g->node_count * sizeof(char*));
if(!vids) return eg_geo_err("reify oom"); if(!vids) return eg_geo_err("reify oom");
for(int64_t i = 0; i < g->node_count; i++) vids[i] = g->nodes[i].id; for(int64_t i = 0; i < g->node_count; i++) vids[i] = g->nodes[i].id;
/* Held for READ across the whole reify pass: it only searches the index.
* (The multi-second SELF-reify beat below builds a PRIVATE index instead and
* never touches this boundary at all.) */
const VIndex* vx = eg_vindex_view(g, dim);
int persisted = engram_geo_reify_store(g_engram_store, vx, vids, int persisted = engram_geo_reify_store(g_engram_store, vx, vids,
(int)g->node_count, NULL); (int)g->node_count, NULL);
eg_vindex_view_release();
free(vids); free(vids);
int nested = 0; int nested = 0;
if(persisted >= 0){ if(persisted >= 0){
@@ -13773,12 +13937,108 @@ static int eg_cog_is_keystone_seeds(const char* csv) {
el_val_t engram_think_json(el_val_t seeds, el_val_t faculty) { el_val_t engram_think_json(el_val_t seeds, el_val_t faculty) {
GeoDescriptor* g = eg_geo_build_desc(EL_CSTR(seeds)); GeoDescriptor* g = eg_geo_build_desc(EL_CSTR(seeds));
if (!g) return eg_geo_err("geometry unavailable"); if (!g) return eg_geo_err("geometry unavailable");
CogStance st; cog_stance_init(&st, NULL, EL_CSTR(faculty), g->hub_id, NULL, g); /* RESUME THE LEARNED STANCE (2026-08-16 self-review). This built a NEUTRAL
* stance every call all axis_gain 1.0, bias_dir NULL, reliability 0.5
* and never loaded the one the correspondence-beat had been persisting.
*
* That mattered because the faculty enters engram_think ONLY through the
* stance: `gain = stance->axis_gain[k]` warps the per-axis extents, and
* `stance->bias_dir` seeds the steering direction. cog_stance_init stores
* the faculty NAME but nothing reads it. So with a neutral stance,
* reason / abduce / induce / plan / analogize are the same function with
* different labels measured, byte-identical output across all five
* and `confidence` is pinned to the 0.5 uninformed prior, because
* GeoGradient.confidence is just stance->reliability.
*
* The machinery already existed and only this call site ignored it:
* engram_correspondence_beat_json resumes via cog_stance_from_node and
* persists via cog_stance_to_node under the id "stance-<faculty>-<hub>".
* Every beat's calibration was being written and then thrown away on the
* next read. Same defect as the NULL anchor directly above: a neutral
* argument collapsing a capability to a constant.
*
* Resume the same id the beat writes, so learning compounds across beats
* and cold boot. Fall back to neutral only when no stance exists yet
* which is a genuine uninformed prior, not a discarded informed one. */
char sid[256];
snprintf(sid, sizeof sid, "stance-%s-%s",
EL_CSTR(faculty) ? EL_CSTR(faculty) : "reason",
g->hub_id ? g->hub_id : "region");
CogStance st; StoreNode prev; int resumed = 0;
if (g_engram_store && store_get_node(g_engram_store, sid, &prev) == 1) {
if (cog_stance_from_node(&prev, &st) == 0) resumed = 1;
store_node_free(&prev);
}
if (!resumed) cog_stance_init(&st, sid, EL_CSTR(faculty), g->hub_id, NULL, g);
else { free(st.id); st.id = strdup(sid); }
GeoGradient grad; GeoGradient grad;
if (engram_think(g, NULL, &st, &grad) != 0) { cog_stance_free(&st); engram_geo_free(g); return eg_geo_err("think failed"); }
/* ANCHOR THE READ (2026-08-16 self-review). This passed NULL, and NULL is
* not "no opinion" engram_think re-origins at `anchor ? anchor :
* region->centroid`, so NULL means "read from the centroid", and the
* centroid is the ONE point where the gradient is zero by construction:
* r = x - centroid = 0, so every axis projection is 0, grad is 0, and
* direction takes the "at rest" branch. Measured consequence: EVERY
* faculty reason, abduce, induce, plan, analogize returned an
* identical null result, differing only in its label:
* {"direction":[0,0,...],"spread":0,"magnitude":1,"confidence":0.5}
* magnitude 1 is membership evaluated at the centroid, spread 0 is its
* distance to itself, and confidence 0.5 is the stance fallback. The
* geometry was never the problem /api/drift computes real values
* (centroid_sep 0.104, core_disp 0.045) over the very same 87 members.
* Neuron could not think because the read was always taken from the
* region's own centre.
*
* The seeds choose WHICH region; they must also supply the VANTAGE it is
* read from. Anchor at the first resolvable embedded seed the same seed
* eg_geo_build_desc infers `dim` from, so the two never disagree. A single
* seed still yields a real gradient because the descriptor expands to the
* seed's neighbourhood (87 members for the self anchor), so the seed's own
* position is distinct from the neighbourhood centroid.
*
* COPY the vector, never borrow it: g->nodes is realloc'd in place on
* append, so a borrowed EngramNode* is a dangling pointer across any
* concurrent write. 768 floats is 3 KB. */
float* anchor = NULL;
{
EngramStore* eg = engram_get();
const char* csv = EL_CSTR(seeds);
if (eg && csv) {
const char* p = csv;
while (*p && !anchor) {
while (*p == ' ' || *p == ',') p++;
const char* s = p;
while (*p && *p != ',') p++;
const char* e = p; while (e > s && e[-1] == ' ') e--;
if (e > s) {
char* id = strndup(s, (size_t)(e - s));
if (id) {
int64_t idx = engram_find_node_index(id);
if (idx >= 0 && idx < eg->node_count) {
EngramNode* n = &eg->nodes[idx];
if (n->emb && n->emb_dim == g->dim) {
anchor = malloc(sizeof(float) * (size_t)g->dim);
if (anchor) memcpy(anchor, n->emb,
sizeof(float) * (size_t)g->dim);
}
}
free(id);
}
}
}
}
}
if (engram_think(g, anchor, &st, &grad) != 0) { free(anchor); cog_stance_free(&st); engram_geo_free(g); return eg_geo_err("think failed"); }
free(anchor);
JsonBuf b; jb_init(&b); char t[256]; JsonBuf b; jb_init(&b); char t[256];
snprintf(t, sizeof t, "{\"faculty\":\"%s\",\"n_support\":%d,\"magnitude\":%.6g,\"spread\":%.6g,\"confidence\":%.6g,\"dim\":%d", /* stance_resumed distinguishes an INFORMED read from an uninformed one.
EL_CSTR(faculty), grad.n_support, grad.magnitude, grad.spread, grad.confidence, grad.dim); * Without it, confidence 0.5 from a learned-but-unreliable stance and
* confidence 0.5 from "no stance exists" are indistinguishable the same
* reporting gap that let the NULL anchor and the neutral stance hide. */
snprintf(t, sizeof t, "{\"faculty\":\"%s\",\"n_support\":%d,\"magnitude\":%.6g,\"spread\":%.6g,\"confidence\":%.6g,\"stance_resumed\":%s,\"dim\":%d",
EL_CSTR(faculty), grad.n_support, grad.magnitude, grad.spread, grad.confidence,
resumed ? "true" : "false", grad.dim);
jb_puts(&b, t); jb_puts(&b, t);
int emit = grad.dim < 8 ? grad.dim : 8; int emit = grad.dim < 8 ? grad.dim : 8;
jb_puts(&b, ",\"direction\":"); eg_geo_emit_vec(&b, grad.direction, emit); jb_puts(&b, ",\"direction\":"); eg_geo_emit_vec(&b, grad.direction, emit);
@@ -14128,6 +14388,21 @@ el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t di
return el_wrap_str(b.buf); return el_wrap_str(b.buf);
} }
/* Public activation entry point. Serializes against the http_worker threads that
* share g->nodes/g->edges this is the guard the awareness main thread
* (soul.el: awareness_run) was missing entirely. Nested calls from a worker that
* already holds the lock pass straight through.
*
* It no longer guards _eg_vindex: the index has its own publication boundary
* (eg_vindex_view / eg_vindex_maintain) and search cannot mutate it. This guard is
* now about the RAM graph's realloc-in-place ONLY. See the note at eg_guard_enter. */
el_val_t engram_activate(el_val_t query, el_val_t depth) {
int owned = eg_guard_enter();
el_val_t r = engram_activate_inner(query, depth);
eg_guard_exit(owned);
return r;
}
el_val_t engram_activate_json(el_val_t query, el_val_t depth) { el_val_t engram_activate_json(el_val_t query, el_val_t depth) {
/* Run two-layer engram_activate and serialize the result list to JSON. /* Run two-layer engram_activate and serialize the result list to JSON.
* Each entry includes both activation_strength (layer 1 background) and * Each entry includes both activation_strength (layer 1 background) and
@@ -14904,6 +15179,7 @@ el_val_t engram_embed_backfill(el_val_t count) {
float* v = eg_embed_fetch(n->content, &d); float* v = eg_embed_fetch(n->content, &d);
if (!v) break; /* embedder down / breaker open — stop this call */ if (!v) break; /* embedder down / breaker open — stop this call */
n->emb = v; n->emb_dim = d; n->emb = v; n->emb_dim = d;
eg_vindex_note_embedded(g, i); /* write-side index maintenance */
done++; done++;
} }
int64_t total = 0; int64_t total = 0;
+2 -2
View File
@@ -222,7 +222,7 @@ static double eff_w(double weight, double hebb){
} }
GeoDescriptor* engram_geometry_descriptor( GeoDescriptor* engram_geometry_descriptor(
EngramPagedStore* store, VIndex* vindex, EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids, char** vids, int n_vids,
const char* const* seed_ids, size_t n_seeds, const char* const* seed_ids, size_t n_seeds,
const GeoParams* params, const GeoParams* params,
@@ -1401,7 +1401,7 @@ static double geo_weighted_degree(EngramPagedStore* st, const char* id, double e
return deg; return deg;
} }
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex, int engram_geo_reify_store(EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids, char** vids, int n_vids,
const GeoReifyParams* params){ const GeoReifyParams* params){
if(!store) return -1; if(!store) return -1;
+2 -2
View File
@@ -150,7 +150,7 @@ void engram_geo_mean_free(GeoMeanCache* c);
* Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error * Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error
* (no seeds resolvable, OOM). */ * (no seeds resolvable, OOM). */
GeoDescriptor* engram_geometry_descriptor( GeoDescriptor* engram_geometry_descriptor(
EngramPagedStore* store, VIndex* vindex, EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids, char** vids, int n_vids,
const char* const* seed_ids, size_t n_seeds, const char* const* seed_ids, size_t n_seeds,
const GeoParams* params, const GeoParams* params,
@@ -375,7 +375,7 @@ void engram_geo_reify_default_params(GeoReifyParams* p);
* neighborhood (+ member edges), superseding any prior same-hub record with * neighborhood (+ member edges), superseding any prior same-hub record with
* provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0. * provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0.
* Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */ * Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex, int engram_geo_reify_store(EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids, char** vids, int n_vids,
const GeoReifyParams* params); const GeoReifyParams* params);
+68 -34
View File
@@ -74,11 +74,6 @@ struct VIndex {
int entry; /* entry-point element index, -1 if empty */ int entry; /* entry-point element index, -1 if empty */
int max_level; /* current top layer */ int max_level; /* current top layer */
/* scratch: version-stamped visited set (O(1) reset). */
uint32_t* visited;
uint32_t visit_epoch;
size_t visited_cap;
}; };
/* ── small helpers ────────────────────────────────────────────────────────── */ /* ── small helpers ────────────────────────────────────────────────────────── */
@@ -166,37 +161,63 @@ static Pair heap_pop(Heap* h, int is_max){
return top; return top;
} }
/* ── visited set ──────────────────────────────────────────────────────────── */ /* ── visited set — owned by the CALL FRAME, never by the index ──────────────
static int visited_ensure(VIndex* ix){ * This buffer is per-TRAVERSAL scratch. It used to live in struct VIndex as an
if (ix->visited_cap >= ix->cap && ix->visited) return 0; * allocation optimisation, which made every traversal a write to shared state:
size_t nc = ix->cap ? ix->cap : 16; * two concurrent vindex_search calls stamped each other's epoch and then walked
uint32_t* nv = (uint32_t*)realloc(ix->visited, nc*sizeof(uint32_t)); * each other's marks, so even two pure READS corrupted the traversal (measured
if (!nv) return -1; * 2026-08-16: TSan data race at visited_reset, reached from vindex_search on one
if (nc > ix->visited_cap) memset(nv + ix->visited_cap, 0, (nc-ix->visited_cap)*sizeof(uint32_t)); * thread and vindex_insert on another; downstream SIGSEGV dereferencing a bogus
ix->visited = nv; ix->visited_cap = nc; * element index).
*
* It is not an ownership problem and it does not want a lock or a capability —
* it was simply misfiled. A pure function's scratch belongs to the call. Moving
* it here is what lets vindex_search take a `const VIndex*`, which is in turn
* what makes "search does not mutate the index" a COMPILE-TIME property instead
* of a review comment.
*
* Cost: one calloc/free of cap*4 bytes per traversal (~55 KB at the live store's
* 13,820 elements), against thousands of dim-768 dot products in the same call.
* Deliberately NOT __thread: http_worker is a thread per connection, so a
* thread-local buffer would retain ~55 KB per connection for the process life. */
typedef struct {
uint32_t* mark; /* per-element epoch stamp */
uint32_t epoch; /* current traversal's stamp; 0 == "no traversal yet" */
size_t cap;
} VVisit;
/* calloc leaves every stamp 0 and epoch 0; the first visit_reset moves to
* epoch 1, so no element reads as visited before it is marked. */
static int visit_init(VVisit* v, size_t cap){
size_t nc = cap ? cap : 16;
v->mark = (uint32_t*)calloc(nc, sizeof(uint32_t));
if (!v->mark) return -1;
v->cap = nc; v->epoch = 0;
return 0; return 0;
} }
static inline void visited_reset(VIndex* ix){ static void visit_dispose(VVisit* v){ free(v->mark); v->mark = NULL; v->cap = 0; }
if (++ix->visit_epoch == 0){ /* wrapped: clear all */ static inline void visit_reset(VVisit* v){
memset(ix->visited, 0, ix->visited_cap*sizeof(uint32_t)); if (++v->epoch == 0){ /* wrapped: clear all */
ix->visit_epoch = 1; memset(v->mark, 0, v->cap*sizeof(uint32_t));
v->epoch = 1;
} }
} }
static inline int is_visited(VIndex* ix, int e){ return ix->visited[e]==ix->visit_epoch; } static inline int is_visited(const VVisit* v, int e){ return v->mark[e]==v->epoch; }
static inline void mark_visited(VIndex* ix, int e){ ix->visited[e]=ix->visit_epoch; } static inline void mark_visited(VVisit* v, int e){ v->mark[e]=v->epoch; }
/* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */ /* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */
/* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller /* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller
* owns res->a. `q` is a normalised query. */ * owns res->a. `q` is a normalised query. */
static int search_layer(VIndex* ix, const float* q, const int* eps, int neps, static int search_layer(const VIndex* ix, VVisit* vis, const float* q,
const int* eps, int neps,
int ef, int layer, Heap* res /*out, max-heap*/){ int ef, int layer, Heap* res /*out, max-heap*/){
Heap cand = {0,0,0}; /* min-heap: nearest to expand */ Heap cand = {0,0,0}; /* min-heap: nearest to expand */
res->a=NULL; res->n=0; res->cap=0; res->a=NULL; res->n=0; res->cap=0;
visited_reset(ix); visit_reset(vis);
for (int i=0;i<neps;i++){ for (int i=0;i<neps;i++){
int e = eps[i]; int e = eps[i];
if (is_visited(ix,e)) continue; if (is_visited(vis,e)) continue;
mark_visited(ix,e); mark_visited(vis,e);
float d = vdist(ix, q, ix->elems[e].vec); float d = vdist(ix, q, ix->elems[e].vec);
Pair p = { d, e }; Pair p = { d, e };
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; } if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
@@ -212,8 +233,8 @@ static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
NeighList* nl = &ce->links[layer]; NeighList* nl = &ce->links[layer];
for (int i=0;i<nl->count;i++){ for (int i=0;i<nl->count;i++){
int e = nl->ids[i]; int e = nl->ids[i];
if (is_visited(ix,e)) continue; if (is_visited(vis,e)) continue;
mark_visited(ix,e); mark_visited(vis,e);
float d = vdist(ix, q, ix->elems[e].vec); float d = vdist(ix, q, ix->elems[e].vec);
if (res->n < ef || d < res->a[0].d){ if (res->n < ef || d < res->a[0].d){
Pair p = { d, e }; Pair p = { d, e };
@@ -232,7 +253,7 @@ static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
* Keep c only if it is nearer to q than to every already-chosen neighbour; * Keep c only if it is nearer to q than to every already-chosen neighbour;
* backfill from the pruned set (nearest first) to reach M for connectivity. * backfill from the pruned set (nearest first) to reach M for connectivity.
* Writes chosen element indices into out[], returns the count. */ * Writes chosen element indices into out[], returns the count. */
static int select_neighbors(VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){ static int select_neighbors(const VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
(void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */ (void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */
/* sort W ascending by (dist,elem) — deterministic. */ /* sort W ascending by (dist,elem) — deterministic. */
for (int i=1;i<nW;i++){ /* insertion sort (nW small) */ for (int i=1;i<nW;i++){ /* insertion sort (nW small) */
@@ -281,7 +302,7 @@ static int elems_reserve(VIndex* ix){
Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem)); Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem));
if (!ne) return -1; if (!ne) return -1;
ix->elems = ne; ix->cap = nc; ix->elems = ne; ix->cap = nc;
return visited_ensure(ix); return 0;
} }
int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
@@ -307,13 +328,19 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
return 0; return 0;
} }
/* This call frame owns its traversal scratch for the whole insert. ix->cap
* already covers `cur` (elems_reserve ran above), so every reachable element
* index is in range. */
VVisit vis;
if (visit_init(&vis, ix->cap)) return -1;
int ep = ix->entry; int ep = ix->entry;
int L = ix->max_level; int L = ix->max_level;
/* greedy descent through layers above `level` to refine the entry point. */ /* greedy descent through layers above `level` to refine the entry point. */
for (int lc = L; lc > level; lc--){ for (int lc = L; lc > level; lc--){
Heap r = {0,0,0}; Heap r = {0,0,0};
int eps1[1] = { ep }; int eps1[1] = { ep };
if (search_layer(ix, el->vec, eps1, 1, 1, lc, &r)){ return -1; } if (search_layer(ix, &vis, el->vec, eps1, 1, 1, lc, &r)){ visit_dispose(&vis); return -1; }
if (r.n){ ep = r.a[0].e; float bd=r.a[0].d; if (r.n){ ep = r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} } for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} }
free(r.a); free(r.a);
@@ -329,7 +356,7 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
for (int lc = start; lc >= 0; lc--){ for (int lc = start; lc >= 0; lc--){
int Mmax = (lc==0) ? ix->M0 : ix->M; int Mmax = (lc==0) ? ix->M0 : ix->M;
Heap W = {0,0,0}; Heap W = {0,0,0};
if (search_layer(ix, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; } if (search_layer(ix, &vis, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int)); int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
if (!chosen){ free(W.a); rc=-1; break; } if (!chosen){ free(W.a); rc=-1; break; }
int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen); int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen);
@@ -357,13 +384,17 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
} }
done: done:
free(eps_owned); free(eps_owned);
visit_dispose(&vis);
if (rc) return -1; if (rc) return -1;
if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; } if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; }
return 0; return 0;
} }
/* ── search ───────────────────────────────────────────────────────────────── */ /* ── search ───────────────────────────────────────────────────────────────── */
int vindex_search(VIndex* ix, const float* query, int k, int ef_search, /* `ix` is const: search is pure with respect to the index. That is enforced by
* the compiler, not by convention — it is the whole point of moving the visited
* set into the frame below. */
int vindex_search(const VIndex* ix, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out){ uint64_t* node_id_out, float* dist_out){
if (!ix || !query || k <= 0) return -1; if (!ix || !query || k <= 0) return -1;
if (ix->entry < 0) return 0; if (ix->entry < 0) return 0;
@@ -373,11 +404,15 @@ int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
float* q = vec_normalise_copy(query, ix->dim); float* q = vec_normalise_copy(query, ix->dim);
if (!q) return -1; if (!q) return -1;
/* This call frame owns its traversal scratch. */
VVisit vis;
if (visit_init(&vis, ix->cap)){ free(q); return -1; }
int ep = ix->entry; int ep = ix->entry;
for (int lc = ix->max_level; lc > 0; lc--){ for (int lc = ix->max_level; lc > 0; lc--){
Heap r = {0,0,0}; Heap r = {0,0,0};
int eps[1] = { ep }; int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, 1, lc, &r)){ free(q); return -1; } if (search_layer(ix, &vis, q, eps, 1, 1, lc, &r)){ visit_dispose(&vis); free(q); return -1; }
if (r.n){ int b=r.a[0].e; float bd=r.a[0].d; if (r.n){ int b=r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;} for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;}
ep = b; } ep = b; }
@@ -385,7 +420,8 @@ int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
} }
Heap res = {0,0,0}; Heap res = {0,0,0};
int eps[1] = { ep }; int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, ef_search, 0, &res)){ free(res.a); free(q); return -1; } if (search_layer(ix, &vis, q, eps, 1, ef_search, 0, &res)){ visit_dispose(&vis); free(res.a); free(q); return -1; }
visit_dispose(&vis);
free(q); free(q);
/* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */ /* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */
@@ -419,7 +455,6 @@ VIndex* vindex_create(int dim, int M, int ef_construction){
ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0); ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0);
ix->entry = -1; ix->entry = -1;
ix->max_level = 0; ix->max_level = 0;
ix->visit_epoch = 0;
return ix; return ix;
} }
@@ -432,7 +467,6 @@ void vindex_free(VIndex* ix){
free(e->vec); free(e->vec);
} }
free(ix->elems); free(ix->elems);
free(ix->visited);
free(ix); free(ix);
} }
+9 -2
View File
@@ -53,8 +53,15 @@ int vindex_insert(VIndex* idx, uint64_t node_id, const float* vec);
* first (ascending distance). Either out array may be NULL to skip it. * first (ascending distance). Either out array may be NULL to skip it.
* ef_search search-time candidate width; larger == higher recall, slower. * ef_search search-time candidate width; larger == higher recall, slower.
* Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k. * Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k.
* Returns the number of results written, or <0 on error. */ * Returns the number of results written, or <0 on error.
int vindex_search(VIndex* idx, const float* query, int k, int ef_search, *
* `idx` is const BY CONTRACT AND BY TYPE: search does not mutate the index. The
* traversal's visited set is owned by the call frame, so N threads may search one
* index concurrently. Concurrent search against a vindex_insert on the same index
* is still unsafe insert rewires existing elements' neighbour lists and reallocs
* elems[] so the index's owner must not extend a published index under a live
* reader. See eg_vindex_view / eg_vindex_maintain in el_runtime.c. */
int vindex_search(const VIndex* idx, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out); uint64_t* node_id_out, float* dist_out);
/* Number of vectors currently indexed. */ /* Number of vectors currently indexed. */
+180
View File
@@ -0,0 +1,180 @@
# El Runtime — Ownership and Capability ABI
**Status:** §0–§2 verified. §3 re-derived and **built** for the vector index (2026-08-16); not yet applied to the resident RAM graph.
**Date:** 2026-08-16
**Scope:** `lang/runtime/` — every El program (soul, engram, cgi-studio vessels) inherits this by rebuild. Nothing in this document is a change to any El *program*.
**Note on §1's line numbers:** they were read against a checkout that has since shifted by ~135 lines. Verified positions as of `a67452f` are in §2a.
---
## 0. The residual
> **Builtins own memory and reach process state directly.**
That is the residual — the generator. Everything below labelled a "residue" is a deposit left by it. The distinction matters because we have spent significant effort removing deposits, and deposits regenerate.
A residue is fixed. A residual is eliminated. Fixing residues while the residual stands produces exactly the pattern observed on 2026-08-15/16: a run of individually-correct patches, each verified, followed by a new defect of the same shape in a different file.
---
## 1. The residues, measured
Each of these is a distinct merged or proposed fix. Each addresses one deposit. None addresses the residual.
| residue | location | fix that was applied or proposed |
|---|---|---|
| `state_get` leaked its return value per call — 15 MB over 200k calls | builtin | el #140 (merged) |
| VIndex freed under a concurrent reader | `el_runtime.c:9424` | `fb32d15` guard (merged 08:46:43) |
| `_eg_vindex_seen` realloc'd on a read path | `el_runtime.c:9412` | same guard |
| `vindex_insert` on a read path | `el_runtime.c:9434`, `9450` | same guard |
| shared `visited` / epoch scratch stomped by concurrent searches | `engram_vindex.c:7981`, `169186`, `195` | proposed: move to per-search frame |
| nine append sites, none indexing → lazily-embedded nodes invisible | `el_runtime.c:7806, 7988, 8148, 8224, 11526, 11731, 12050, 15295, 15312` | "embed-gap #20", patched by making the *read* path catch up (`9439` comment) |
**Measured:** all file/line references above, read 2026-08-16. Crash frames `engram_activate → eg_vindex_sync → vindex_insert → _realloc → _xzm_xzone_malloc_freelist_outlined` are accounted for by rows 24.
**Inferred, not yet verified:** that the nine append sites do not share a single commit point. This needs one pass before Change C is sized.
---
## 2. Why these are one defect
`eg_vindex_sync` (`el_runtime.c:9419`) has exactly three callers, and **all three are reads**:
- `engram_activate``9802`
- `eg_knn_for_node``13075` (its own header comment states *"No writes."*)
- `engram_geo_reify_run_json``13285`
It mutates five process-global statics (`94009404`): `_eg_vindex`, `_eg_vindex_dim`, `_eg_vindex_built_nc`, `_eg_vindex_seen`, `_eg_vindex_seen_cap`.
Reads mutate because index maintenance was never given an owner on the write side. It got bolted onto reads, because a builtin *could* reach the globals — nothing prevented it. Likewise `state_get` leaked because a builtin *owned* the value it returned; nothing prevented that either.
The store is architecturally append-only and superseding. A read path that mutates contradicts that directly. The contradiction is expressible only because the ABI permits it.
---
## 2a. Verified positions and the fact §1 missed
Read directly at `a67452f`, 2026-08-16. §1's line numbers predate a ~135-line shift; these are current.
| thing | §1 said | actually |
|---|---|---|
| five process-global statics | 94009404 | **95359539** |
| `eg_vindex_seen_ensure` realloc | 9412 | **9547** |
| `eg_vindex_sync` | 9419 | **9554** |
| `vindex_free` on a read path | 9424 | **9559** |
| `vindex_insert` on a read path | 9434 / 9450 | **9569** (build) / **9585** (incremental) |
| caller: `engram_activate_inner` | 9802 | **9939** |
| caller: `eg_knn_for_node` | 13075 | **13212** |
| caller: `engram_geo_reify_run_json` | 13285 | **13422** |
| `fb32d15` guard | — | lock **1602**, depth **1631**, `eg_guard_enter` **1636**, `http_worker` acquire **1687**, `engram_activate` wrapper **14097** |
| VIndex scratch fields | 7981 | **7981** ✓ |
| `search_layer` race site | 195 | **195** ✓ |
**The structural fact §1 and §3 both missed:** *the index does not inherit the store's append-only property.* `vindex_insert` rewires the `NeighList` links of already-existing elements and reallocs `elems[]` — so extending the index mutates the whole structure, not just its tail. This is why "make reads pure" is necessary but **not sufficient**, and why §3 needed a publication boundary rather than only a capability split. It is reproduced as a standing test (`unsynchronized` half, §5).
---
## 3. The change
*(Re-derived 2026-08-16. The previous §3 — a runtime context struct carrying read/write **capability pointers** to every builtin — was written in mutable-store, C-ownership terms. It asked "who is permitted to mutate the shared thing?", which presupposes a shared mutable thing. The engram is immutable and recall is projection; what does not mutate needs no ownership discipline. So the question is not answered, it is dissolved. The implemented change is below.)*
### 3.1 Three moves, in decreasing order of how much they dissolve
**(1) Misfiled scratch is not shared state.** `visited` / `visit_epoch` were never conceptually owned by the index — they are one traversal's local, hoisted into `struct VIndex` as an allocation optimisation. Nothing about them is derived geometry. They want neither a lock nor a capability nor a checkout pool: a pure function's scratch belongs to its call frame, and the fix is to put it back there. This is not "the capability model applied by hand to one global"; it is the deletion of a false ownership claim.
**(2) `const` is the capability, and immutability hands it over for free.** Once the scratch leaves the struct, `search_layer` reads the index and nothing else — so `vindex_search` can take a `const VIndex*`. That is *precisely* the teeth old-§3 wanted from capability pointers: a read path physically cannot call `vindex_insert`, and it is a **compile error**, not a review comment. It costs one qualifier rather than a new ABI swept across hundreds of builtins. The compiler enforces it on every future caller for the same reason.
> The capability type was already in the language. It is spelled `const`.
**(3) What remains is a publication problem, not an ownership problem.** With scratch in the frame and reads const, one hazard survives, and it is real: **HNSW insert is not an append.** `vindex_insert` rewires the `NeighList` links of *already-existing* elements and reallocs `elems[]`. The store's append-only property does **not** transfer to the index derived from it. So a reader projecting against the index while its owner extends it is unsafe no matter how pure search is.
Immutability answers this too, and the answer is publication:
- **`eg_vindex_maintain`** — the sole mutator. Takes the boundary exclusively; never runs beside a reader.
- **`eg_vindex_view`** — returns a `const VIndex*` with the boundary held for read. N readers project concurrently; none can mutate.
A read path may **demand that a current snapshot exist** — that is a request to the owner, not a mutation by the reader. What it may not do is mutate the geometry it is projecting against. `view` / `maintain` is exactly that split, and it is why this replaces `eg_vindex_sync` rather than wrapping it.
**Write-side owner.** Index membership is owned by the event *"an embedding became present on this ordinal"* — not by node append, since a node without an embedding cannot be in a vector index at all. `eg_vindex_note_embedded` hooks the embedding-assignment sites: one O(log n) insert, no O(node_count) presence scan. This also retires the "STALENESS (honest tradeoff)" note in the old `eg_vindex_sync`, where a lazily-embedded *older* node stayed invisible to `route_nearest` / autoconnect until the next full rebuild.
### 3.2 What this does not claim
The **resident RAM graph** (`g->nodes` / `g->edges`) is a *separate* residue of the same residual and is untouched by this change. It is realloc'd in place (`el_runtime.c:7618`, `7629`), so an awareness-thread reader holding `EngramNode* n = &g->nodes[i]` across a concurrent append holds a dangling pointer — and `engram_activate_inner`'s embed-backfill writes `n->emb` through exactly such a pointer. It wants the same publication treatment the index just received. Until that lands, the `fb32d15` guard stays (see §5).
---
## 4. Why this is not a large change
The old §4 argued that El owning its compiler makes a capability-ABI sweep mechanical, since `elc` generates every builtin call site. That argument was load-bearing only for the ABI, and the ABI is gone.
The constraint now travels with the **type of the thing**, not the shape of every call site — so no sweep is needed at all. Measured extent of the implemented change: two qualifiers (`const VIndex*` on `vindex_search`, propagated to `engram_geometry_descriptor` and `engram_geo_reify_store`), one struct field group relocated to a call frame, one rwlock, and three read call sites converted from `eg_vindex_sync` to `view`/`release`.
The payoff of owning the language is unchanged and is now *cheaper*: introduced once, enforced by the compiler on every future builtin, cannot subsequently be forgotten. Contrast the current state, where the same discipline was maintained by hand across hundreds of builtins and demonstrably failed at least six times.
---
## 5. What this deletes
**Deleted (done, 2026-08-16):**
- `eg_vindex_sync` — the function itself. Not renamed: split into `eg_vindex_maintain` (mutating, exclusive, sole owner) and `eg_vindex_view` (const, shared). A name that meant "read paths repair the index" had to stop existing.
- `VIndex::visited` / `visit_epoch` / `visited_cap` — the struct fields, `visited_ensure`, its call from `elems_reserve`, `ix->visit_epoch = 0` in `vindex_create`, and `free(ix->visited)` in `vindex_free`.
- The **proposed** per-search scratch *struct on the index* (a checkout pool / `VisitedListPool`) — never built. The buffer is a plain frame local; a pool is machinery for an ownership question that no longer exists.
- The **proposed** reader-view / owner-handle split for VIndex specifically — superseded. `const` already is the reader view.
- `EXPECT_RACE` in `run_vindex_concurrency_tests.sh` — a knob that let a known defect ride as "expected". Replaced by four halves with real verdicts.
**NOT deleted — the design doc was wrong about this one:**
- `fb32d15` (`eg_guard_enter` / `engram_req_lock` / `_eg_req_depth`). §5 originally called for its removal as "a lock protecting a mutation that ceases to exist." **Measured, it guards two things, and only one of them ceases to exist.** Its own comment names both: the RAM graph *and* `_eg_vindex`. The vindex justification is retired; the RAM-graph justification is independently load-bearing (§3.2), and removing the guard reintroduces the measured 11171→9579 edge-loss defect from 2026-08-14. Its comment has been narrowed to state the RAM graph only. **Precondition for deleting it:** the resident graph gets the same publication boundary the index just got.
- el #140's hand-patch. Left in place — the leak stops being *expressible* only under the abandoned capability-ABI §3, which is not what was built.
**Ordering consequence (revised):** the original ordering claim — "the residual lands first, the residues evaporate rather than get fixed" — did not survive contact. The residual here is not a single ABI that dissolves everything at once; it is a *property* (derived state is published, never edited) applied per structure. The index now has it. The RAM graph does not yet. Residues evaporate **per structure, in the order the property is applied**, and a residue whose structure has not been converted must be left standing, not deleted on the strength of the plan.
---
## 6. Sequencing
1. **Read** how builtins are declared and dispatched, to confirm the call sites are compiler-generated in one place. *(This determines whether §4 holds. If dispatch is scattered, re-size before proceeding.)*
2. Introduce the context type and capability types.
3. Codegen emits the context at every builtin call site.
4. Mechanical sweep of builtin signatures.
5. Move index maintenance behind the write capability; the three read callers take the read capability.
6. Delete the residue-fixes listed in §5.
7. **One** build of soul from el dev — which resolves the `state_get` leak and the crash together, rather than deploying a leak fix that reintroduces the crash.
---
## 7. Open questions
**Answered 2026-08-16:**
- ~~Do the nine append sites share a commit point?~~ **Moot.** The question was mis-aimed: node append is not the event that owns index membership, because a node without an embedding cannot be in a vector index. The five *embedding-assignment* sites are the real owner points (`el_runtime.c:7091, 9839, 13362, 15002`, plus snapshot-restore at `7951`), and three of them carry the ordinal directly — which is all `eg_vindex_note_embedded` needs. The other two run before the node is resident, where the cold build picks it up.
- ~~Does anything outside `lang/runtime/` construct a second `VIndex`?~~ **No.** Swept: the only constructors outside the runtime are `engram/test/*` and `lang/runtime/vindex_bench.c`, all single-threaded and index-private. Inside the runtime, `engram_self_reify_beat_json` builds a **private** index deliberately and never touches the shared boundary — that was already correct and is unchanged.
- ~~Does the HTTP worker pool contend on the same globals?~~ **Yes, and it was never the whole story.** Workers serialize against each other on `engram_req_lock`, but the awareness main thread does not take it at all — that is the gap `fb32d15` closed. Now verified independent of that guard: the index boundary is its own rwlock, so worker/awareness contention on `_eg_vindex` is handled whether or not the request lock is held.
**Still open:**
- The resident RAM graph wants the same publication boundary (§3.2). Until it has one, `fb32d15` cannot be deleted.
- `eg_vindex_view` holds the boundary for read across `engram_geo_reify_store`, which is a long pass. Correct, but it stalls the owner for that duration. If reify latency becomes a problem the answer is a refcounted snapshot, not a shorter lock.
---
## 7a. Evidence (measured 2026-08-16, `engram/test/run_vindex_concurrency_tests.sh`)
| half | before | after |
|---|---|---|
| `single` — 3000 vectors, 1 thread, ASan+UBSan | clean | clean |
| `readers` — 4 readers, no writer, TSan | **race** at `engram_vindex.c:195` (`visited_reset``vindex_search`) | **clean** |
| `unsynchronized` — writer+reader, bare index, TSan | race | **race, expected and permanent** — now the proof the boundary must exist |
| `published` — owner + 4 readers through the boundary, TSan | *(did not exist)* | **clean**, all 3000 inserts landed |
No recall regression: `recall@10 = 0.9365` at `ef_search=128` (gate ≥ 0.90); the determinism test still yields byte-identical results across two independent builds.
Builds locally: all seven engram runtime translation units compile `-Wall -Wextra` clean, and the full engram binary links (`engram/dist/engram.c` + runtime, arm64). The one pre-existing `-Wcomment` warning in `el_runtime.c` is present at `a67452f` too.
---
## 8. What this document is not
It is not an argument for a memory model in general, a garbage collector, process isolation between soul and engram, or a client/server split of the store. Each of those was considered and each addresses mutation that this change removes. They are answers to a question that stops being asked.