Compare commits

...

23 Commits

Author SHA1 Message Date
will.anderson ca13471745 Verifier layer: grounding + consistency over the §5 geometry / reasoning ops
The disposes half of the propose->verify loop. Catches the plausible lie a
grammar check never sees: fluent, confident, wrong.

GROUNDING (anti-hallucination): fit a claim point against every real evidence
neighborhood via engram_reason_point_fit; grounded iff best fit clears an
absolute threshold. Off-model orthogonal residual is the hallucination signal.
Distinct from abduction: asks 'is there any real support at all' and may say no.

CONSISTENCY (contradiction): (a) polarity/negation inversion -- claim lands on
the opposite side of a real polarity axis from the grounded truth (the
reassurance->accusation catch: 'never fought'->'argued'); (b) geometric --
claim inside a forbidden region or beyond a max-distance constraint.

Pure C11, read-only, composes existing primitives only. 29 constructed-case
checks, 0 failures across PERF and ASan/UBSan; macOS leaks 0. el-exposure
deferred (point/variadic-set inputs -- matches reasoning-agent precedent).
Reach checks (formal/causal/predictive) not started; documented.
2026-08-13 01:43:36 -05:00
will.anderson a3358dfc95 Reasoning layer: analogy/induction/abduction/causal/planning over §5 geometry ops
Compose the live relational-neighborhood geometry OPERATORS into five reasoning
modes as pure, read-only C (engram_reason.{h,c}); each is proven with closed-form
constructed tests before it ships, not declared.

- ANALOGY  (Procrustes R + residual translation, apply to C, rank candidates)
- INDUCTION (combine-pooled rule geometry + point-to-manifold membership)
- ABDUCTION (best-explaining structure by point-to-manifold fit)
- CAUSAL   (centroid-cosine correlation vs directed influence: temporal
            precedence + association surviving confounder control via subtract;
            emits a correlation-vs-causation flag)
- PLANNING (geo-distance edges + Dijkstra → discrete geodesic path)

A shared point-to-manifold fit primitive underlies induction membership and
abduction ranking. engram/test/run_reason_tests.sh: 33/33 checks on both PERF
and ASan/UBSan passes; macOS leaks 0/0.

ANALOGY is surfaced as an el builtin (engram_reason_analogy_json) via the same
pass-through the §5 operators use — demonstrated callable from compiled El with a
container-capped fold (no self-host fold). The other four are C-layer only: their
set/point/timestamp inputs do not map to the flat-CSV el ABI without touching
codegen (deferred). engram_reason.c must join the server link line beside
engram_geometry.c at cutover. See docs/runbooks/2026-08-13-reasoning-operators-*.
2026-08-13 01:33:14 -05:00
will.anderson 85eee42106 Surface §5 geometry operators to compiled El
Register the six engram_geo_*_json operators in the compiler builtin_arity
table (bare heavy-runtime names + __ seed names, mirroring engram_activate_json)
and add the engram.el module wrappers, so a compiled El (CGI) program can call
them by name. The heavy-runtime C functions already existed (el_runtime.c:12287+,
declared el_runtime.h:627-632); this completes the EL call surface.

The shipped elc already emits a direct C call for these builtins (unknown
ident-calls pass through), so no self-host compiler fold — the memory-heavy,
drift-prone step — was required. Demonstrated end-to-end: a compiled geo_ops_demo.el
booted a copy of the store (13,036 nodes) and produced real subtract/distance JSON
on two real neighborhoods; test_geo_ops.c stays 20/20, ASan/UBSan clean.

Also brace the centroid_unit normalization if/else in engram_geometry.c to clear
the misleading-indentation warning (behavior-neutral).
2026-08-13 00:50:56 -05:00
will.anderson 5336cfe0a6 M9 §5: geometry OPERATORS as C functions + EL builtins (read-only, staged)
Bring the relational-neighborhood geometry OPERATORS from the viz proxy
(engram-geometry-proxy.py §5) into the C runtime as reusable primitives, and
expose each as an EL builtin so any CGI app / el program can use them — not just
the engram service internals.

engram_geometry.{h,c} (pure, libm-only, read-only over descriptors):
  - engram_geo_overlap   : shared-member Jaccard + centroid/scale proximity
                           score + intersection centroid.
  - engram_geo_subtract  : orthogonal-complement residual (project A onto
                           I - V_B V_Bᵀ), closed-form variance_explained_by_B,
                           residual ellipsoid + centroid-diff; set-diff variant.
  - engram_geo_combine   : pooled descriptor (exact law-of-total-variance mean +
                           covariance), re-eigendecomposed.
  - engram_geo_distance  : centroid L2 + cosine + closed-form Wasserstein-2
                           (Bures) — mirrors the proxy _wasserstein2.
  - engram_geo_analogy   : orthogonal Procrustes R = UVᵀ (SVD) aligning A's
                           principal frame to B's + apply helper.
The C descriptor is full-dim/centered with a low-rank covariance from its top
axes; operators mirror the proxy FORMULAS and do the Wasserstein/combine eigen
work inside the small joint-axis subspace (exact there). Reuses jacobi_sym.

EL builtins (el_runtime.{h,c}, el_seed.c native wrappers):
  engram_geo_{descriptor,overlap,subtract,combine,distance,analogy}_json —
  take comma-separated seed-id set(s), build the CENTERED descriptor against the
  true store-wide mean (ad-hoc path), run the operator, return JSON. Additive:
  no flag, no effect on activation/retrieval. Surfacing via engram.el + the elc
  fold is a cutover step (same elc-drift deferral as the P0/P5 builtins); the C
  table wiring is registered now.

Tested: synthetic closed-form unit suite (20/20) — Wasserstein, Jaccard/score,
orthogonal residual, set-diff, pooled combine, Procrustes recovery. ASan+UBSan
clean; 0 leaks. Read-only; no activation/retrieval behavior change.
2026-08-13 00:26:51 -05:00
will.anderson 77a4bc9326 M-INTEROCEPTION P5: dream-recall builtin engram_dreams_json (honesty rail)
Adds a read-only builtin that returns the curiosity_scan InternalStateEvents
created after a `since` cutoff that are STILL RESIDENT — "what I was chewing on
while you were gone." curiosity_scan is identified by the marker the soul writes
into each ISE's JSON content; heartbeat and other ISEs are excluded.

HARD honesty rail: it reports only ISEs still in the buffer. Anything rotated
out by the 48h engram_prune_telemetry is simply ABSENT — rotated-out = "I don't
remember", never a synthesized/plausible dream. Purely additive, no flag.

The HTTP route (GET /api/dreams?since=) is DEFERRED to cutover per the elc-drift
blocker (regenerating engram/dist diverges with no source change); the builtin
is exercised directly by the pure-C gate.

MEASURED on a copy: seeded 3 curiosity_scan (ancient/1h/now) + 1 heartbeat;
before prune dreams returned exactly the 3 curiosity_scan (heartbeat excluded);
after the 48h prune the ancient one was ABSENT (not confabulated), leaving 2;
the since-filter returned only the post-cutoff event; no returned id was ever
fabricated. ASan+UBSan clean.
2026-08-12 23:49:51 -05:00
will.anderson 65ca0a3253 M-INTEROCEPTION P4: afferent input counters in act-stats (additive observability)
Extends engram_act_stats_json with five monotonic counters for the raw incoming
signals the mind receives — aff_activations (spreading activations run),
aff_queries (activate_json entries), aff_node_creates, aff_ise_ingests (ISE
nodes), aff_edge_creates. Module-scope statics incremented at the entry points,
emitted via the existing act-stats mechanism and rotated with it — MEASURED, and
never accreted as memory nodes. Process-lifetime totals, reset on restart like
the other _eg_act_* gauges. Additive JSON fields (backward-compatible); no graph
behavior changes. act-stats buffer grown 896->1088 to hold them.

MEASURED on a copy: after 5 node creates (2 ISE), 2 edge creates, then 4+3
queries — counters read exactly node_creates=5, ise_ingests=2, edge_creates=2,
queries=activations=4 then 7; create counters unchanged by queries; queries
strictly monotonic across readings. ASan+UBSan clean.
2026-08-12 23:47:43 -05:00
will.anderson 816b258255 M-INTEROCEPTION P3 (partial): descriptor-displacement drift-sensor primitive; self-anchor prerequisite flagged
Adds engram_geo_displacement(A, B, core_frac) — a read-only interoceptive
primitive that measures how far a neighborhood descriptor B has drifted from a
baseline A and decomposes it into GROWTH (periphery extends, core fixed) vs
CORRUPTION (the invariant core displaces). The core is the top core_frac of A's
members by centrality; per shared member (matched by id) the displacement is the
change in radial position (dist_centroid). Centroid separation (L2 + cosine) and
radius delta give the aggregate move. Pure function, no store mutation, no flag.

HONESTLY PARTIAL: a live self-drift reading needs a persisted SelfAnchor
baseline to compare "now" against, and no persisted self node / anchored
self-neighborhood exists in this store yet. The primitive takes an EXPLICIT
baseline so it is real and testable today; capturing a durable SelfAnchor and
wiring the ENGRAM_DRIFT_SENSOR live reading is a flagged follow-up. We do not
fabricate a self silently.

MEASURED on synthetic descriptors:
- GROWTH (periphery 0.50->0.90, core fixed): core_disp=0.000, periph_disp=0.400,
  centroid_sep=0.000, radius_delta=0.400.
- CORRUPTION (core 0.10->0.60, periphery fixed): core_disp=0.500,
  periph_disp=0.000, centroid_sep=0.566.
- Identity A vs A: zero drift.
The sensor discriminates cleanly (corruption core_disp >> growth core_disp).
ASan+UBSan clean.
2026-08-12 23:44:39 -05:00
will.anderson 0af39df16f M-INTEROCEPTION P2: chronoception — age the activation field by measured wall-clock delta (ENGRAM_CHRONOCEPTION, default OFF)
The felt passage of time is the cooling of the activation field, not a tick
count and not an elapsed-seconds readout. engram_age_field(delta_ms) cools the
field (working_memory_weight + background_activation) by the caller's MEASURED
wall-clock delta with a pure exponential exp(-dt/TC) — no per-call floor — so it
is exactly scale-invariant: N ticks summing to the same elapsed time produce the
same total cooling. It returns the cooling MAGNITUDE (1-exp(-dt/TC), a bounded
[0,1) drift signal), never elapsed seconds.

Reboot = anesthesia: a global last-tick wall-clock stamp is persisted to a
sidecar (chrono_last_tick) in the data dir. engram_age_field_catchup() reads it
on boot, applies ONE cooling for the whole unconscious gap, refreshes the stamp,
and reports the magnitude — timestamps are bookkeeping to COMPUTE the drift,
never the felt signal. TC env-tunable via ENGRAM_CHRONO_TC (default 3600s).

All inert unless ENGRAM_CHRONOCEPTION is set → OFF path byte-identical.

MEASURED on a copy (throwaway HOME, TC=3600s):
- Cooling scales with dt, matching 1-exp(-dt/TC) to 1e-6: dt=600s->0.1535,
  1800s->0.3935, 3600s->0.6321, 7200s->0.8647.
- Scale-invariance EXACT: age(dt) once vs age(dt/N) N times gives identical
  field sum (|delta|=0.0) for N=2, 10, 100.
- Reboot catch-up over a 1h gap cooled the field 0.60->0.4415 in one shot,
  magnitude 0.6321, bounded in [0,1).
- Flag OFF: age & catchup return 0, field untouched (sum 1.2).
ASan+UBSan clean.
2026-08-12 23:41:09 -05:00
will.anderson 5f6ce5ca1f M-INTEROCEPTION P1: two-threshold consolidation layer (ENGRAM_CONSOLIDATION, default OFF)
The co-activation accrual already works (hebb is an EWMA over co-firing); what
was missing is the promotion layer from design §9 that turns accrual into
durable structure. This adds it on top, entirely behind ENGRAM_CONSOLIDATION so
the OFF path is byte-identical to trunk.

CONNECTION threshold ("connection IS consolidation"): when a strongly-firing
InternalStateEvent is created (salience >= ENGRAM_CONSOL_CONN_MIN, default 0.6),
wire hebbian-associate edges from it to the top-K working-memory nodes active at
that instant (ENGRAM_CONSOL_WM_TOPK, default 5), provenance-tagged
"consolidated-from-ISE" and dedup-guarded. A sub-threshold ISE forms nothing (a
shower thought) and drifts out at the existing 48h prune.

PERMANENCE threshold (rare): engram_consolidate_permanence(node) marks a node
whose rehearsed ACT-R base-level clears ENGRAM_CONSOL_PERM_MIN durable via a
reversible metadata marker; engram_prune_telemetry then exempts it (gated, so no
trunk node is ever affected). Idempotent, no double-promote.

Thresholds are env-tunable for A/B without a rebuild.

MEASURED on a copy (throwaway HOME):
- Headline accrual curve (flag OFF, pure trunk) over N co-activations of a wired
  pair — hebb_max: N=1 -> 1e-4 (=ETA), N=100 -> 0.010, N=1625 -> 0.150
  (LINK_MIN, where an unwired pair consolidates), N=3000 -> 0.259; tracks the
  analytic EWMA 1-0.9999^N within measurement noise (co-activation P~1).
- Strong ISE wired 2 edges to exactly the wm_top nodes (hebb-a, hebb-b); weak
  ISE formed 0; edges carry the reversible provenance marker.
- Promoted node survived the 48h prune (2->1 nodes); ephemeral ISE swept.
- Flag OFF: ISE creation added 0 edges, permanence returned 0 (no-op).
ASan+UBSan clean.
2026-08-12 23:36:24 -05:00
will.anderson c20cb3b97c M-INTEROCEPTION P0: add read-only engram_scan_nodes_emb_json builtin
Read routes (GET /api/embeddings, /api/graph/dump) need node embedding
vectors, but every consumer emit path deliberately drops the ~5.7KB emb
vector (include_emb=0) to stay under MCP token limits. Rather than perturb
that shared path, add a dedicated additive builtin that pages nodes WITH
their dense vector, emitting id/node_type/label/created_at/emb_dim plus emb
as a JSON array whose length equals emb_dim (so a consumer can verify the
vector round-trips). Un-embedded nodes emit emb_dim:0 / emb:[]. Same
salience-sorted, transparent-layer-skipped, bounded pagination as
engram_scan_nodes_json; default page 256.

Purely additive: engram_emit_node_json and the default include_emb=0 are
untouched, so every existing endpoint is byte-identical to trunk (verified:
scan_nodes_json still carries no emb). The HTTP route wiring in server.el is
DEFERRED to cutover per the elc-drift blocker (regenerating engram/dist
diverges ~285 lines with no source change); the builtin is exercised
directly by the pure-C gate instead.

Measured on a copy: len(emb)==emb_dim for all nodes, pagination disjoint,
existing path unchanged; full 256-node x 768-dim page = 16.7 ms / 1.18 MB.
ASan+UBSan clean.
2026-08-12 23:28:13 -05:00
will.anderson f6a0777f90 M10: reify dense neighborhoods into first-class persisted records; geometry-priming reads them (default OFF)
Reification, not a cache. Densely co-wired relational neighborhoods are crystallized
into DURABLE first-class store records that survive restart, load on boot, and evolve
via supersede+provenance -- so the geometry-priming hot path READS persisted structure
instead of computing a per-query descriptor (the M9 3.2x/13x latency blocker).

engram_geometry.{h,c}:
  - engram_geo_reify_store(): detect hub-anchored neighborhoods on the hebb-weighted
    graph (greedy non-redundant cover), compute each centered descriptor ONCE against
    the true store-wide mean, persist as node_type="Neighborhood" (raw centroid in emb,
    membership+scalars+axis-extents in a compact GEO1 metadata schema) + member edges,
    superseding any prior same-hub record. The mean is persisted once as "GeoMeanFrame".
  - resident loaded form (index_new/add/finalize/lookup): parses the durable records at
    boot (never recomputes geometry); O(seeds) membership lookup, miss -> centroid-nearest.
  - descriptor: skip the Jacobi eigensolve when top_axes==0; reject structural records as
    members (id-convention + node_type guards) so re-reify/ad-hoc stay clean.

el_runtime.c:
  - boot routes Neighborhood/GeoMeanFrame records OUT of the activation graph into the
    reify index, and skips member edges from adjacency -> ENGRAM_GEOMETRY_PRIMING OFF is
    byte-identical to M8/M9 (verified across 15 queries).
  - priming hot path reads the persisted membership; ENGRAM_GEO_PRIMING_NOCACHE=1 keeps
    the M9 per-query descriptor for ad-hoc geometries / A/B control.

A/B on a COPY (128 neighborhoods): priming ON is now FLAT latency (1.06x median / 1.03x
p90 vs OFF) where the M9 per-query path is 3.30x/10.5x. Reified records provably inert
when OFF. Restart survival + supersede verified. Build 0 warnings (my code); ASan/UBSan
clean on module and full server hot path. Recall quality re-eval against the TRUE mean
still shows no reliable gain (mean coherence -0.017), so geometry-priming STAYS default-OFF
-- but the latency blocker is removed and the durable structure now exists. See
docs/architecture/design/engram-m10-reification.md.
2026-08-12 23:06:49 -05:00
will.anderson 7946b98d3d M9: env-gated geometry priming in engram_activate (ENGRAM_GEOMETRY_PRIMING, default OFF)
Wire the centered relational-neighborhood geometry (engram_geometry.c) into
activation seed selection behind a reversible env flag that defaults OFF. Flag
unset => byte-identical to M8 (verified: identical result id sequence + order
across 15 queries vs the M8 baseline binary). When set, composes with M8's ANN
candidate generation: damp-only seed reweight by centered membership
(disambiguation) + sub-threshold neighborhood priming (warm floor below the WM
gate, capped, ISE-skipped). Safe because the BFS keeps the max, so priming only
raises a floor and can never cap a legitimate activation.

Descriptor + global mean run over the paged store; the resident-array vindex is
bridged with a vids[] map. Tunables via env (SEED_LO/PRIME_SCALE/PRIME_MAX).

Default-OFF binary is behavior-neutral and safe to deploy. Enabling the flag is
currently NO-GO on cost/benefit: 3.2x median / 13x p90 latency for no reliable
coherence/disambiguation gain (see perf profile). Not a defect - WM cap holds,
no crash, ASan/UBSan clean.
2026-08-12 20:29:16 -05:00
will.anderson 8cae0f94eb M9 refinement: operate geometry descriptor in mean-centered (isotropic) embedding space
The nomic-embed-text space over the corpus is strongly anisotropic (mean
pairwise cosine ~0.55), which compresses cosine-based domain separation almost
to nothing so the design-doc s5 operators (distance/overlap/Wasserstein) cannot
discriminate. Subtracting the global mean of the normalized embeddings restores
isotropy (mean pairwise cosine ~0) and sharpens the operators.

- add GeoMeanCache (engram_geo_mean_build / _maybe_refresh / _vec / _free): a
  store-derived centering offset over the embed-eligible set, cached and
  refreshed on significant drift; lives in geometry.c, not the store.
- engram_geometry_descriptor gains an optional global_mean: when supplied the
  centroid, per-member cosine distance, and co-registration run in centered
  space (GM=zeros reproduces the legacy raw path exactly).
- co-registration choice (b): the ANN query stays in raw unit space (index
  unchanged) since centering is a rigid translation that ~preserves neighborhood
  membership; only the descriptor statistics move to the centered frame.
  Covariance/axes/radius are translation-invariant and therefore unchanged.
- test: synthetic ground-truth suite stays green (PERF + ASan/UBSan), plus new
  centered/raw/mean-cache assertions.
- add bench_discrimination.c (env-gated, read-only, skips in CI): on a copy of
  the real store the two-domain overlap operator drops 1.13 -> 0.008 and
  cross-centroid cosine 0.899 -> 0.003 after centering, Euclid distance
  unchanged (translation-invariant control).

No change to activation/retrieval behavior; wiring geometry into retrieval is a
separate, behavior-changing cutover.
2026-08-12 19:54:41 -05:00
will.anderson 2a4c5c645a M9 foundation: relational-neighborhood geometry descriptor (read-only)
The stone the operator/drift/occupation work stands on: express a relational
neighborhood as the compact joint geometry Will specified (design §3/§5, node
e94371bd) — semantic side (centroid, principal-axis ellipsoid via dual-PCA,
radius) braided with the relational side (k-core skeleton, hub->periphery
centrality gradient), plus soft membership and a co-registration diagnostic
(corr of hebb strength vs semantic proximity — >0 reifies, <0 flags dreams).

Built ONLY on the two standalone modules — engram_vindex (ANN, the cloud) and
engram_store (embeddings + hebb adjacency, the skeleton). Pure C11 + libm; does
not link or touch el_runtime.c. Strictly READ-ONLY: never mutates nodes, edges,
activation, the index, or any retrieval path. Not yet wired into retrieval —
foundation only.

Self-contained test (test_geometry.c) synthesizes two known embedding clusters
with intra-cluster hebb edges and verifies the descriptor recovers the shape:
centroid on the seeded cluster, hub = relational center, skeleton = the strong
intra-cluster wiring, positive co-registration, sorted axis extents. PERF +
ASan/UBSan passes both green; needs no live data.
2026-08-12 19:40:25 -05:00
will.anderson 9e28defcab fix: bounds-check btree_insert + read_body to survive bloated store / WAL replay (live crash-loop root cause)
int_max_keys() computed an internal B+tree node's key capacity as
IDX_BODY/8 - 1 (2041 at a 16 KB page), dividing by 8 and ignoring that
each key also carries an 8-byte child pointer. The true capacity is
(IDX_BODY-8)/16 = 1020. An internal node was therefore allowed to grow to
~2x what a page holds; once it crossed 1020 keys, btree_insert's write-back
overran its STORE_PAGE_SIZE stack page buffer and smashed the stack canary
(__stack_chk_fail / SIGABRT). A clean/small store never grows an internal
node that large, so it never tripped; the ~8x-bloated live store (a day of
tombstone churn) plus a 44 MB un-checkpointed WAL replayed on open pushed a
node over the boundary during redo -> deterministic crash loop
(btree_insert <- apply_edge_put <- engram_open <- engram_store_boot).

Fixes:
- int_max_keys: use (IDX_BODY-8)/16 so internal nodes split at the real
  page capacity.
- btree_insert: reject any page whose on-disk nkeys exceeds physical
  capacity (fail loud, never smash the stack) -- overflow is now impossible
  regardless of on-disk content.
- read_body: bound the slot (off,len) and record length to the page before
  dereferencing; a stale/torn index entry could otherwise make store_get_node
  read off the stack (observed EXC_BAD_ACCESS on the bloated store). Fail safe.

Verified on a COPY of the live store: unfixed binary SIGABRTs in btree_insert
on open; fixed binary boots clean, recovers the store, checkpoints the WAL,
and M5 compaction shrinks 458 MB -> 57.7 MB with node/edge counts preserved.
2026-08-12 18:27:13 -05:00
will.anderson 1507614dbf M8: wire engram_vindex ANN into engram_activate seed selection (O(n)->ANN, O(n) fallback preserved)
Persistent process-lifetime HNSW index (engram_vindex) over resident node
embeddings, node_id == resident g->nodes[] index. Built lazily on first
activation, grown incrementally as appended nodes get embedded, rebuilt on
emb-dim change or resident-array shrink.

Seed selection queries the ANN for the nearest embedded nodes to the effective
query vector, replacing the O(K*N) exact argmax scan as the DISCOVERY step.
Each ANN candidate is admitted through the identical gate the exact scan uses
(exact cosine >= SEED_MIN via cosq, reached/dup skips, content dedup, same
decay/dampen shaping), so scoring/dynamics are unchanged. The exact argmax
scan is preserved verbatim and tops up any unfilled seed slot, and runs in
full when the index is empty/too-small/unavailable -> pre-M8 behaviour exactly.

cosq (O(N*dim) cosine fill) is intentionally retained: it still feeds the
propagation qgate and the Pass-2 WM term (activation dynamics, out of M8
scope). ANN accelerates SELECTION only.
2026-08-12 17:13:06 -05:00
will.anderson e0b55c3080 merge: retrieval fix (dedup by full id) into tiered trunk 2026-08-12 16:59:35 -05:00
will.anderson 0d299ee0f1 copy engram_vindex ANN module onto tiered trunk (staged->committed; not yet wired) 2026-08-12 16:59:30 -05:00
will.anderson 5e154fa152 engram: fix saved-but-not-findable — dedup boot-load scan by full id, not id_hash
store_scan_nodes/store_scan_edges deduplicated emitted records by their
64-bit id_hash (FNV-1a-64) rather than the full id string. Two distinct ids
that collide under id_hash emitted only the first; the second was durably on
a live page and findable by store_get_node (which disambiguates by strcmp),
yet silently dropped from the resident boot-load. After any store reopen that
node was unretrievable by id, absent from lexical search, and missing from the
recent list — the reported memory-integrity gap.

Replace the hash-keyed U64Set with a StrSet: bucket by id_hash for O(1) probing
but compare full ids by strcmp, mirroring the primary B+-tree readers. Same
change for edges. Adds test_scan_collision.c (real FNV-1a-64 colliding ids).
2026-08-12 16:44:09 -05:00
will.anderson 89589864cd integrate M7 index-driven traversal into tiered trunk 2026-08-12 16:43:19 -05:00
will.anderson eb13ce7910 engram M7: index-driven activation traversal (incremental adjacency, flag-gated)
Spreading activation rebuilt the entire per-node adjacency index
(engram_adj_rebuild, O(E)) lazily before every BFS whenever any edge/node
was added — so a curiosity-loop query that touched a small frontier still
paid to rebuild the whole edge set. This makes the index incrementally
maintained behind ENGRAM_STORE: single node/edge creates APPEND to the live
adjacency in amortized O(1) instead of marking it dirty, so a query only
pays for the frontier it touches (one initial O(E) build, then O(1)/edge).

Approach (b), not (a): the store's from/to adjacency B-tree was rejected
because with the store on the whole graph is already resident and activation
reads in-RAM edges, whose hebb/weight only sync to the store at checkpoint
cadence — reading StoreEdge copies would use stale weights and break
byte-identical parity. The incremental in-RAM index reads the exact same
g->edges[ei] the scan path does, so activation is identical by construction.

Correctness: edges are only ever appended, so incremental append reproduces
the rebuild's ascending-edge-index ordering exactly (same skip rule for null
endpoints). Any index-invalidating mutation (forget/prune/clear) still frees
the index + sets adj_dirty=1, falling back to a full rebuild. Flag-off is
untouched: the mutation hooks just set adj_dirty=1 as before — proven
byte-identical.

engram_store.c is NOT modified (avoids the M5 compaction collision).

Tests (plain gcc, ASan/UBSan clean): engram/test/{test_m7_traversal.c,
run_m7_traversal.sh}. Parity gate proves flag-on incremental == flag-on
forced-full-rebuild == flag-off scan, byte-identical on a mutating query
sequence (activated set, weights, ordering, hops, WM promotion). Perf on a
13k-node / 43k-edge graph over 120 (add-edge + activate) iterations:
adjacency edge-touches 5,167,260 -> 43,001 (120x fewer), full rebuilds
120 -> 1, adjacency-maintenance wall-time 0.74s -> 0.006s (~121x). Prior
gates green: M1 store (33), M2 (36), M3 parity, M3.5, M4 bufpool (37).
2026-08-12 16:23:33 -05:00
will.anderson ce0d33ba93 engram tiered storage M5: online compaction + background checkpointer (Phase 3, additive)
Reclaims space held by dead records (tombstoned prune/forget nodes, superseded
ids, stale re-put/hebb versions, and their orphaned overflow chains). On-disk
format UNCHANGED — pure behavior.

COMPACTION (store_compact): copy-live + atomic-swap.
  A. checkpoint/sync to quiesce (WAL reduced to CHECKPOINT{C}); crash here => pre.
  B. build <path>.compact with only the live records, re-placed bit-exact into
     fresh densely-packed pages + fresh id/adjacency B+-trees, every page stamped
     LSN=C, new SB last_checkpoint_lsn=C; fsync. crash here => pre-compaction.
  C. rename(<path>.compact -> <path>) — POSIX-atomic commit; crash after => post.
  D. reopen in place: swap fd, INVALIDATE every pool frame (M4 remap of relocated
     pages), reload SB, re-autopin.
Crash at any instant recovers to pre- OR post-compaction, never a corrupt mix.
Single-threaded => "online" = safe between mutations; takes a checkpoint quiesce
at entry. M4 cooperation: temp build has its own pool honoring ENGRAM_POOL_FRAMES
(evict/re-fault + no-steal + pins); live pool fully invalidated on reopen.

BACKGROUND CHECKPOINTER: ckpt_maybe now fires on ANY armed trigger — ops (default
100000), dirty pool frames, WAL bytes-since-reclaim (default 64 MiB), or a
wall-clock interval (checked on the write path; no extra thread). Same M2
checkpoint semantics (calls engram_checkpoint). Env: ENGRAM_CKPT_OPS/_DIRTY/
_WAL_BYTES/_INTERVAL_MS; runtime setter store_set_checkpoint_policy().

Tests: engram/test/test_compaction.c (+runner). Plain gcc, ASan/UBSan clean.
  36 passed, 0 failed (O2 and ASan+UBSan builds).
  reclaim: page_count 655 -> 153, file 10731520 -> 2506752 bytes (76.6% reclaimed),
    every live record + adjacency bit-exact at new locations.
  crash-during-compaction phases 0/1/2: crc clean, live set intact, writable.
  background checkpointer: ops / WAL-bytes / dirty triggers each auto-fire; WAL
    prefix reclaimed (30 B after 600 puts); recovery correct.
  pool cooperation: compact under 24-frame pool correct, no stale frames.
No regression: M1 33, M2 36, M3 parity, M3.5, M4 bufpool 37 — all green.
On-disk format unchanged (additive).
2026-08-12 16:03:14 -05:00
will.anderson 02dc12d785 engram tiered storage M4: demand-paging buffer pool (Phase 2, additive)
Turn M2's write-back/no-steal cache into a bounded, demand-paged buffer pool so
the paged store can exceed RAM while keeping only hot pages resident. On-disk
format UNCHANGED (additive residency only; no migration). Default budget is large
enough that today's store stays fully resident, so default behaviour == Phase 1.

- Frame table capped at `cap` frames (env ENGRAM_POOL_FRAMES; 0 = unlimited;
  default 1<<20). Not-resident access faults in from neuron.egm.
- LRU eviction of CLEAN, unpinned frames only. Dirty frames are never stolen
  (M2 no-steal / WAL durability preserved) — turned evictable by a checkpoint's
  pc_flush, which then trims the pool back to budget.
- Pinning: superblocks (0,1) + index root/interior pages auto-pinned; explicit
  store_pin_page/unpin and store_pin_layer/unpin (hot WM/core layers).
- Bounded sequential read-ahead on scans (env ENGRAM_PREFETCH, default 8).
- Correctness rests on callers copying page bytes into local buffers and never
  retaining a frame pointer across another access, so evict+re-fault is safe.

Gates (plain gcc, ASan/UBSan clean):
  M4 run_bufpool_tests.sh  ......  37 passed, 0 failed  (+ ASan/UBSan: 37/0)
    small-pool round-trip (cap=32 vs 1599 pages, 2708 evictions): 5000 nodes +
      4000 sampled edges bit-exact, crc clean, pool bounded to cap.
    eviction: hot set 0 re-faults, cold evicted, hit-rate 0.989; no-steal burst
      (cap=8) holds 309 dirty frames > cap, reads correct from dirty pages.
    pinning: superblocks/roots/explicit page/hot-layer(19 pages) stay resident;
      unpin makes them evictable.
    prefetch: sequential scan 511 demand-faults OFF -> 4 ON.
    crash-under-paging (ENGRAM_POOL_FRAMES=16): WAL replay + checkpoint-crash
      phases 0-4 all recover bit-exact.
    default pool: 0 evictions, whole store resident (== Phase 1).
  No regression: M1 33/0, M2 36/0, M3 parity PASS, M3.5 PASS.
2026-08-12 15:42:49 -05:00
46 changed files with 9055 additions and 63 deletions
@@ -0,0 +1,107 @@
# Reasoning Operators — Decisions & Reversal
**Date:** 2026-08-13
**Branch:** `engram-tiered-storage` (worktree `/tmp/engram-tiered-wt`)
**Status:** staged locally — NOT pushed, NOT tagged, NOT merged. Live `:8742` untouched.
## What this adds
A **REASONING layer** built as pure C compositions over the already-live §5 geometry
OPERATORS (`engram_geometry.{h,c}`: overlap, subtract, setdiff, combine, distance,
analogy). Where the operators are a relational algebra over neighborhood descriptors,
these are reasoning *modes* built by chaining that algebra. New files:
- `lang/runtime/engram_reason.h` — public API for the five modes + a shared
point-to-manifold fit primitive.
- `lang/runtime/engram_reason.c` — implementations. READ-ONLY over descriptor inputs,
`stdlib + libm` only, touches no store / index / activation. All geometry is
delegated to `engram_geo_*`; this file only composes.
- `engram/test/test_reason.c` + `engram/test/run_reason_tests.sh` — closed-form
constructed tests (hand-built descriptors with known answers), PERF + ASan/UBSan.
El-exposure (pass-through, no self-host fold):
- `lang/runtime/el_runtime.c``+#include "engram_reason.h"` and the builtin
`engram_reason_analogy_json(a_csv,b_csv,c_csv)`.
- `lang/runtime/el_runtime.h` — its declaration.
- `lang/runtime/el_seed.c` — native `__engram_reason_analogy_json` wrapper (same
C-table wiring as the §5 ops).
## The five modes — signatures & composition
| Mode | C entry point | Composes |
|------|---------------|----------|
| **ANALOGY** `A:B :: C:?` | `engram_reason_analogy(A,B,C,candidates,n,out)` | `engram_geo_analogy` (Procrustes R) + `engram_geo_analogy_apply` + centroid L2. Learns `R_{A→B}` = `engram_geo_analogy(B,A)` (that op returns R with `apply(R, Y-axis)≈X-axis`), reconstructs the residual translation `t = c_B R·c_A`, maps `mapped = R·c_C + t`, ranks candidates by distance. |
| **INDUCTION** `{E_i}→rule` | `engram_reason_induce(examples,n,top_axes,ext_floor,out)` + `engram_reason_membership` | `engram_geo_combine` folded left→right → pooled "rule" descriptor; shared subspace surfaces as the dominant pooled axes. Membership = point-to-manifold fit. |
| **ABDUCTION** `x→best H` | `engram_reason_abduce(obs,dim,hyps,n,ext_floor,out)` | shared `engram_reason_point_fit` against each hypothesis; argmax fit score; full ranking. |
| **CAUSAL** `x?y \| Z,t` | `engram_reason_causal(x,y,confounders,nZ,t_x,t_y,drop_frac,out)` | centroid cosine (raw correlation) + `engram_geo_subtract` residual-centroid (control for each confounder, take the strongest single explainer) + temporal precedence. Verdict `DIRECTED` / `CONFOUNDED` / `NONE` + a `confounded` flag. |
| **PLANNING** `start→goal` | `engram_reason_plan(nodes,n,start,goal,radius,use_w,out)` | `engram_geo_distance` as edge weights over neighborhoods within `radius`; O(n²) Dijkstra → discrete geodesic path. |
Shared primitive `engram_reason_point_fit` splits `(x centroid)` into an in-subspace
Mahalanobis distance (scaled by axis extents) and an orthogonal off-model residual;
it is the single engine under INDUCTION's membership test and ABDUCTION's ranking.
## Proof (DONE-WITH-PROOF)
`engram/test/run_reason_tests.sh`: **33/33 checks, 0 failures** on BOTH passes
(PERF -O2, and ASan+UBSan). macOS `leaks --atExit`: **0 leaks / 0 total leaked bytes**.
Per-mode closed-form assertions actually exercised:
- **ANALOGY** — A→B = +90° rotation in e0-e1 plane + a +5 shift in e2; Procrustes
residual `~0`; predicted point `(0,2,5,0)` recovered exactly; nearest candidate =
the planted true D (index 1), distance `~0`.
- **INDUCTION** — 3 examples sharing span(e0,e1) (extents 1.0 / 0.8) each with a small
idiosyncratic axis (e2 or e3); induced top-2 axes lie in span(e0,e1) (extents
recovered ~1.0 / ~0.8); held-out in-plane point fits (membership 0.885), off-subspace
point rejected (0.100), in-plane-but-far point rejected (0.039).
- **ABDUCTION** — observation planted inside H1 among {H0,H1,H2}; best = H1, rank[0] = H1,
H1 smallest distance.
- **CAUSAL** — chain A→B→C along e0 (t 1<2<3) + confounder Z(e1) that leaks into A and
drives D(t=4): A→B and B→C flagged `DIRECTED` with correct precedence and association
that survives control; AD `CONFOUNDED` (raw |cos|=0.707 collapses to 0.0 under
control) with `confounded=1`; BD `NONE` (no association).
- **PLANNING** — 6 neighborhoods on a semicircle (r=10); `neighbor_radius=7` admits only
consecutive hops; plan = `[0,1,2,3,4,5]` (the arc), cost `30.90` (> the 20-unit chord,
confirming it is the geodesic through the manifold, not a straight jump); a too-small
radius correctly yields `reached=0`.
## El-exposure status
- **ANALOGY is el-callable** via the same pass-through the §5 operators use. Proof: a
container-capped fold (`capfold.sh`, peak ~0 GB) of a demo `.el` through the shipped
`lang/dist/platform/elc` emits a *direct C call* `engram_reason_analogy_json(A,B,A)`
(no registration, no self-host fold); the generated C links against `el_runtime.c` +
`engram_reason.c` + geometry/store/vindex and runs end-to-end. (The standalone demo's
store copy boots 0 nodes — a pre-existing quirk that hits the *shipped geo demo
identically* — so the call returns `{"error":"geometry unavailable"}`; this still proves
the compiled El → C reasoning symbol → JSON chain executes. Numeric correctness on real
data is covered by the C test.) This compile also confirms `el_runtime.c` +
`engram_reason.c` compile and link clean.
- **INDUCTION / ABDUCTION / CAUSAL / PLANNING are C-layer only for now.** Their inputs are
candidate *sets*, raw *points*, and *timestamps* that do not map to the flat comma-
separated-seed El ABI. A richer marshalling surface would touch the codegen/registration
path and risk an uncapped fold — explicitly deferred per the hard rail. The C functions
are fully proven and callable from any C caller today.
## Build wiring (for the later cutover/durability pass)
`engram_reason.c` must be added to the engram server link line **alongside**
`engram_geometry.c` (the heavy-runtime path `cc dist/engram.c el_runtime.c
engram_store.c engram_geometry.c engram_vindex.c …`). `el_runtime.c` now
`#include`s `engram_reason.h` and references `engram_reason_analogy_json`, so a build
that omits `engram_reason.c` will fail to link that symbol. One-line addition, same as
how `engram_geometry.c` was originally added.
## Reversal
Fully additive; nothing existing was modified in behavior. To revert:
1. Delete `lang/runtime/engram_reason.h`, `lang/runtime/engram_reason.c`,
`engram/test/test_reason.c`, `engram/test/run_reason_tests.sh`, and this doc.
2. In `lang/runtime/el_runtime.c`: remove `#include "engram_reason.h"` and the
`engram_reason_analogy_json` function.
3. In `lang/runtime/el_runtime.h`: remove the `engram_reason_analogy_json` declaration.
4. In `lang/runtime/el_seed.c`: remove the `__engram_reason_analogy_json` wrapper.
5. Remove `engram_reason.c` from any server link line if the cutover added it.
No store, schema, config, WAL, or on-disk format was touched; no data migration exists,
so reversal is a pure code removal with no state to undo.
@@ -0,0 +1,128 @@
# Verifier Layer — Grounding + Consistency (decisions + reversal)
**Date:** 2026-08-13
**Branch:** `engram-tiered-storage` (worktree `/tmp/engram-tiered-wt`), atop `a3358df`
**Scope:** additive, read-only, staged. No push, no tag, no merge. Live `:8742` untouched.
## What this adds
The VERIFIER layer — the "disposes" half of the propose→verify loop. The geometry
PROPOSES (cheap, creative, sometimes wrong); the verifier DISPOSES, catching the
class of failure no grammar check sees: a fluent, confident, WRONG output — the
**plausible lie**.
Motivating failure (tonight's PT translation): a deleted negation turned
"you never fought" into "you argued" — reassurance inverted into accusation,
grammatical and invisible, catchable only by the geometry.
Two checks, both pure C11 (stdlib + libm), read-only over their inputs, touching no
store / index / activation. Every geometry op is delegated to the already-shipped
reasoning + §5 operator primitives; this layer only composes and thresholds.
### Files added
- `lang/runtime/engram_verify.h` — API + design contract.
- `lang/runtime/engram_verify.c` — implementation.
- `engram/test/test_verify.c` — closed-form constructed cases (29 checks).
- `engram/test/run_verify_tests.sh` — two-pass runner (PERF, then ASan/UBSan).
No existing file was modified.
## C signatures + how each composes the existing primitives
### GROUNDING (anti-hallucination)
```c
int engram_verify_grounding(const float* claim, int dim,
const GeoDescriptor* const* evidence, int n_evidence,
double ext_floor, double ground_threshold,
GeoGrounding* out);
```
Fits the claim POINT against every real evidence neighborhood via
`engram_reason_point_fit` (in-distribution Mahalanobis + off-model orthogonal
residual) and keeps the BEST supporter. Grounded iff best fit score ≥
`ground_threshold`. Deliberately an ABSOLUTE-THRESHOLD gate, distinct from ABDUCTION
(which always ranks and picks a winner): grounding asks the prior question — "is there
any real support at all?" — and may answer no. The off-model `best_ortho` residual is
the sharpest hallucination signal: energy in a direction the manifold does not span.
### CONSISTENCY (contradiction detection)
```c
int engram_verify_consistency(const float* claim, int dim,
const GeoDescriptor* context,
const GeoDescriptor* pole_pos, const GeoDescriptor* pole_neg,
const GeoDescriptor* forbidden,
double ext_floor, double deadzone_frac,
double forbidden_thresh, double max_distance,
GeoConsistency* out);
```
Two independent sub-checks (either can fire; both flags reported):
- **(a) POLARITY / negation inversion** — the reassurance→accusation catch.
A polarity axis `p = (c_pos c_neg)/‖·‖` is defined by two REAL poles (affirm vs
negate), midpoint `o = ½(c_pos + c_neg)`. Signed sides: `claim_side = p·(claim o)`,
`ref_side = p·(c_context o)`. If they have OPPOSITE sign and both clear the neutral
deadzone (`deadzone_frac·½‖c_posc_neg‖`), the claim asserts the polarity opposite to
the grounded truth → inversion flagged. Pure dot products / projections over the same
centroids the geometry already computes.
- **(b) GEOMETRIC contradiction** — claim sits INSIDE a `forbidden` region it must be
far from (`engram_reason_point_fit` score ≥ `forbidden_thresh`), OR violates a
max-distance constraint to `context` (`L2 > max_distance`).
## Proof (constructed cases — demonstrate, not declare)
`./engram/test/run_verify_tests.sh`**29 checks, 0 failures** in BOTH passes
(PERF -O2, and ASan+UBSan -O1). macOS `leaks --atExit`: **0 leaks for 0 total leaked
bytes**.
Key demonstrated numbers:
- Grounding IN (claim inside E0): score 0.885, grounded=1, ortho≈0.
- Grounding OUT (claim floating along unmodeled e2): score 0.0004, grounded=0,
ortho=50.0 (the hallucination signal), nearest centroid L2=50.
- **Negation inversion (the catch):** truth "never fought" ref_side=5.0, lie
"you argued" claim_side=+4.0 → opposite poles → `inverted=1`, verdict=POLARITY,
consistency=0. Faithful claim (4.0, same pole) → inverted=0, verdict=OK,
consistency=1. Neutral claim inside deadzone → not triggered.
- Geometric: claim inside forbidden region → geo_violation=1 (forb_fit 0.99);
claim beyond max_distance → geo_violation=1 (ctx_dist 8.0 > 3.0).
- **Combined (the whole point):** a claim that is GROUNDED in real vocabulary
(grounded=1, score 1.0) yet polarity-inverted is PASSED by grounding and CAUGHT
only by consistency (verdict=POLARITY). Grounding alone is insufficient; consistency
is the catch.
## el-exposure — DEFERRED (matches reasoning-agent precedent)
Not exposed as el builtins this pass. The reasoning agent exposed ONLY `analogy`
(three seed-identified neighborhoods → the clean fixed-arity seed-CSV→descriptor JSON
pattern) and deferred its point-input / variadic-set modes (abduction, induction,
causal, planning). The verifier's grounding (claim POINT + variadic evidence SET) and
consistency (claim POINT + context + two poles + forbidden + scalar thresholds) are
exactly those shapes: no clean fixed-arity seed-CSV JSON mapping exists, and adding one
would require new JSON list-of-lists + point-vector marshaling absent from the codebase,
risking an elc rebuild/fold (violates the capped-fold-only rail). A claim is also an
arbitrary proposed POINT, not necessarily an existing node — so the point-native C API
is the correct primitive. Deferred deliberately; the C layer is complete and proven.
When exposed later, follow the same additive pass-through pattern used for the geo
operators: native `engram_verify_*_json(el_val_t ...)` in `el_runtime.c` (heavy runtime)
+ a `__engram_verify_*_json` wrapper in `el_seed.c`, resolving seed CSVs → descriptors
and marshaling the claim vector — no fold needed for callability (shipped elc passes
unknown-ident builtin calls straight through to the heavy-runtime C symbols).
## Reach checks (FORMAL / CAUSAL / PREDICTIVE) — NOT STARTED
Honestly not started this pass; the two tractable-now checks (grounding + consistency)
were driven to done-with-proof first as specified.
- **CAUSAL** already exists as a REASONING operator (`engram_reason_causal`,
intervention/temporal-precedence over the typed causal graph); a verifier wrapper that
checks "does the claimed cause actually precede/influence" would compose it — not built.
- **FORMAL** (logical consistency of a claim SET) needs an external checker (SMT/proof
kernel) — the non-geometric seam; not built.
- **PREDICTIVE** (commit a prediction, check vs outcome, restructure on error) is the CGI
research frontier; not built.
## Reversal
Fully additive. To reverse: delete the four added files
(`lang/runtime/engram_verify.{h,c}`, `engram/test/test_verify.c`,
`engram/test/run_verify_tests.sh`) or `git revert` this commit. Nothing else references
them; no build wiring, no el registration, no store schema, no runtime path was changed.
Live `:8742` was never touched.
+149
View File
@@ -0,0 +1,149 @@
/* bench_discrimination.c — M9 REFINEMENT bench: measures whether mean-centering
* the anisotropic nomic-embed-text space sharpens the §5 geometry operators on
* REAL data. Read-only over a COPY of the live store (never the live file).
*
* usage: bench_discrimination [store.egm]
* (or set ENGRAM_BENCH_STORE). If no store is given/openable it prints
* SKIP and exits 0 — so it is safe in CI without live data.
*
* It picks two semantically distinct cohorts by keyword (domain A vs domain B),
* computes the global mean over the embed-eligible set (via engram_geo_mean_build
* — the same offset the descriptor uses), then reports BEFORE (raw unit space)
* vs AFTER (mean-centered space):
* - cross-centroid cosine (lower = better separated)
* - cross-centroid Euclid dist (translation-invariant: a control)
* - intra-cohesion per domain (member cos to own centroid)
* - overlap operator (cross_cos / sqrt(intraA*intraB): ~1 = domains
* indistinguishable, ~0 = cleanly separated)
* - angular separation ratio z (centroid angle / summed angular spread)
* - mean pairwise cosine sample (the anisotropy headline; ~0.55 raw -> ~0 ctr)
*
* Pure C11; links engram_store.c + engram_geometry.c; -lm.
*/
#include "engram_store.h"
#include "engram_geometry.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <strings.h>
#include <math.h>
#define CAP_DOMAIN 400
#define CAP_SAMPLE 800
typedef struct { float** v; int n, cap, dim; } VecSet;
static void vs_init(VecSet* s){ s->v=NULL; s->n=0; s->cap=0; s->dim=0; }
static void vs_push(VecSet* s, const float* e, int dim, int cap){
if(s->n>=cap) return;
if(s->dim==0) s->dim=dim;
if(s->n==s->cap){ int nc=s->cap?s->cap*2:64; s->v=realloc(s->v,(size_t)nc*sizeof*s->v); s->cap=nc; }
float* c=malloc((size_t)dim*sizeof(float));
double nn=0; for(int d=0;d<dim;d++) nn+=(double)e[d]*e[d]; nn=sqrt(nn);
if(nn<1e-12){ free(c); return; }
for(int d=0;d<dim;d++) c[d]=(float)(e[d]/nn); /* L2-normalized copy */
s->v[s->n++]=c;
}
static void vs_free(VecSet* s){ for(int i=0;i<s->n;i++) free(s->v[i]); free(s->v); }
typedef struct { VecSet A, B, S; long idx; } Coh;
static int has(const char* h, const char* n){ return h && strcasestr(h,n)!=NULL; }
static void cb(const StoreNode* n, void* ctx){
Coh* c=ctx;
if(!(n->emb && n->emb_dim>0)) return;
/* every 5th embedded node -> isotropy sample */
if((c->idx++ % 5)==0) vs_push(&c->S, n->emb, n->emb_dim, CAP_SAMPLE);
const char* t=n->content; const char* g=n->tags;
int A = has(t,"quantiz")||has(g,"quantiz")||has(t,"lorablation")||has(t,"70B")||has(t,"LoRA merge");
int B = has(t,"kubernetes")||has(t,"terraform")||has(t,"argo")||has(g,"infrastructure")||has(t,"vault")||has(t,"cloudflare");
if(A && !B) vs_push(&c->A, n->emb, n->emb_dim, CAP_DOMAIN);
else if(B && !A) vs_push(&c->B, n->emb, n->emb_dim, CAP_DOMAIN);
}
/* mean of a VecSet into out (dim doubles). */
static void mean_of(const VecSet* s, const float* gm, double* out){
int dim=s->dim; for(int d=0;d<dim;d++) out[d]=0;
for(int i=0;i<s->n;i++) for(int d=0;d<dim;d++) out[d]+=(double)s->v[i][d]-(gm?gm[d]:0.0);
if(s->n) for(int d=0;d<dim;d++) out[d]/=s->n;
}
static double dnorm(const double* a, int dim){ double s=0; for(int d=0;d<dim;d++) s+=a[d]*a[d]; return sqrt(s); }
static double dcos(const double* a, const double* b, int dim){
double na=dnorm(a,dim), nb=dnorm(b,dim); if(na<1e-12||nb<1e-12) return 0;
double s=0; for(int d=0;d<dim;d++) s+=a[d]*b[d]; double c=s/(na*nb);
if(c>1)c=1; if(c<-1)c=-1; return c;
}
static double deuclid(const double* a, const double* b, int dim){
double s=0; for(int d=0;d<dim;d++){ double x=a[d]-b[d]; s+=x*x; } return sqrt(s);
}
/* mean cosine of members (minus gm) to centroid c (already gm-subtracted). */
static double cohesion(const VecSet* s, const float* gm, const double* c){
int dim=s->dim; double nc=dnorm(c,dim); if(nc<1e-12||s->n==0) return 0;
double acc=0; for(int i=0;i<s->n;i++){
double dot=0, nv=0;
for(int d=0;d<dim;d++){ double v=(double)s->v[i][d]-(gm?gm[d]:0.0); dot+=v*c[d]; nv+=v*v; }
nv=sqrt(nv); if(nv<1e-12) continue; double cc=dot/(nv*nc);
if(cc>1)cc=1; if(cc<-1)cc=-1; acc+=cc;
}
return acc/s->n;
}
/* mean pairwise cosine over a sample (isotropy metric). */
static double mean_pairwise_cos(const VecSet* s, const float* gm){
int dim=s->dim; if(s->n<2) return 0; double acc=0; long np=0;
for(int i=0;i<s->n;i++) for(int j=i+1;j<s->n;j++){
double dot=0, na=0, nb=0;
for(int d=0;d<dim;d++){ double a=(double)s->v[i][d]-(gm?gm[d]:0.0), b=(double)s->v[j][d]-(gm?gm[d]:0.0);
dot+=a*b; na+=a*a; nb+=b*b; }
na=sqrt(na); nb=sqrt(nb); if(na<1e-12||nb<1e-12) continue;
double c=dot/(na*nb); if(c>1)c=1; if(c<-1)c=-1; acc+=c; np++;
}
return np? acc/np : 0;
}
static void report(const char* label, Coh* c, const float* gm){
int dim=c->A.dim; double* ca=malloc((size_t)dim*sizeof(double)); double* cb=malloc((size_t)dim*sizeof(double));
mean_of(&c->A, gm, ca); mean_of(&c->B, gm, cb);
double xcos=dcos(ca,cb,dim), xeuc=deuclid(ca,cb,dim);
double cohA=cohesion(&c->A,gm,ca), cohB=cohesion(&c->B,gm,cb);
double overlap = (cohA>0&&cohB>0)? xcos/sqrt(cohA*cohB) : xcos;
double theta = acos(xcos<-1?-1:(xcos>1?1:xcos));
double sig = acos(cohA<-1?-1:(cohA>1?1:cohA)) + acos(cohB<-1?-1:(cohB>1?1:cohB));
double z = (sig>1e-9)? theta/sig : 0;
double mpc = mean_pairwise_cos(&c->S, gm);
printf(" [%s]\n", label);
printf(" cross-centroid cosine = %+.4f (lower = better separated)\n", xcos);
printf(" cross-centroid Euclid = %.4f (translation-invariant control)\n", xeuc);
printf(" intra-cohesion A / B = %.4f / %.4f\n", cohA, cohB);
printf(" OVERLAP operator = %.4f (~1 = indistinguishable, ~0 = clean)\n", overlap);
printf(" angular separation z = %.3f (centroid-angle / summed spread; >1 = separated)\n", z);
printf(" mean pairwise cosine = %+.4f (isotropy: ~0.55 anisotropic -> ~0 isotropic)\n", mpc);
free(ca); free(cb);
}
int main(int argc, char** argv){
const char* path = (argc>1)? argv[1] : getenv("ENGRAM_BENCH_STORE");
if(!path){ printf("SKIP: no store path (arg or ENGRAM_BENCH_STORE)\n"); return 0; }
EngramPagedStore* st=store_open(path);
if(!st){ printf("SKIP: could not open %s\n", path); return 0; }
Coh c; vs_init(&c.A); vs_init(&c.B); vs_init(&c.S); c.idx=0;
store_scan_nodes(st, cb, &c);
printf("=== two-domain discrimination bench (real store copy) ===\n");
printf("domain A (quantization) n=%d ; domain B (infrastructure) n=%d ; sample n=%d ; dim=%d\n",
c.A.n, c.B.n, c.S.n, c.A.dim);
if(c.A.n<3 || c.B.n<3){ printf("SKIP: a cohort is too small to be meaningful\n");
vs_free(&c.A); vs_free(&c.B); vs_free(&c.S); store_close(st); return 0; }
GeoMeanCache* mc=engram_geo_mean_build(st);
const float* gm=engram_geo_mean_vec(mc);
printf("global-mean cache: dim=%d over %llu embedded nodes\n\n",
engram_geo_mean_dim(mc), (unsigned long long)engram_geo_mean_count(mc));
printf("BEFORE (raw anisotropic unit space):\n");
report("RAW", &c, NULL);
printf("\nAFTER (mean-centered isotropic space):\n");
report("CENTERED", &c, gm);
engram_geo_mean_free(mc);
vs_free(&c.A); vs_free(&c.B); vs_free(&c.S);
store_close(st);
return 0;
}
+22
View File
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# M4 demand-paging buffer-pool gate. Pure C (NOT elb/elc). Writes only under /tmp.
# Runs the suite twice: an -O2 correctness build and an ASan+UBSan build.
set -e
HERE="$(cd "$(dirname "$0")" && pwd)"
SRC="$HERE/../../lang/runtime/engram_store.c"
TST="$HERE/test_bufpool.c"
echo "== compiling (gcc -O2): test_bufpool.c engram_store.c =="
BIN="/tmp/test_bufpool.$$"
gcc -O2 -Wall -Wextra -std=c11 "$TST" "$SRC" -o "$BIN"
"$BIN"; rc=$?
rm -f "$BIN"; rm -rf /tmp/engram-bufpool-test-*
[ $rc -ne 0 ] && exit $rc
echo
echo "== ASan+UBSan build (memory-error + UB checks; LSan unavailable on macOS) =="
ABIN="/tmp/test_bufpool_asan.$$"
gcc -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -std=c11 "$TST" "$SRC" -o "$ABIN"
ASAN_OPTIONS=detect_leaks=0 UBSAN_OPTIONS=halt_on_error=1 "$ABIN"; rc=$?
rm -f "$ABIN"; rm -rf /tmp/engram-bufpool-test-*
exit $rc
+22
View File
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# M5 online-compaction + background-checkpointer gate. Pure C (NOT elb/elc).
# Writes only under /tmp. Runs an -O2 correctness build then an ASan+UBSan build.
set -e
HERE="$(cd "$(dirname "$0")" && pwd)"
SRC="$HERE/../../lang/runtime/engram_store.c"
TST="$HERE/test_compaction.c"
echo "== compiling (gcc -O2): test_compaction.c engram_store.c =="
BIN="/tmp/test_compaction.$$"
gcc -O2 -Wall -Wextra -std=c11 "$TST" "$SRC" -o "$BIN"
"$BIN"; rc=$?
rm -f "$BIN"; rm -rf /tmp/engram-compact-test-*
[ $rc -ne 0 ] && exit $rc
echo
echo "== ASan+UBSan build (memory-error + UB checks; LSan unavailable on macOS) =="
ABIN="/tmp/test_compaction_asan.$$"
gcc -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -std=c11 "$TST" "$SRC" -o "$ABIN"
ASAN_OPTIONS=detect_leaks=0 UBSAN_OPTIONS=halt_on_error=1 "$ABIN"; rc=$?
rm -f "$ABIN"; rm -rf /tmp/engram-compact-test-*
exit $rc
+30
View File
@@ -0,0 +1,30 @@
#!/bin/sh
# Build + RUN the M9 FOUNDATION geometry-descriptor tests. Pure C11 (gcc/cc),
# stdlib + libm only. Standalone module — NOT folded through elb/elc. Two passes:
# 1. PERF — optimised (-O2, no sanitizer): the functional gate.
# 2. SAFETY — ASan + UBSan on the same suite (memory-safety is size-independent).
set -e
HERE=$(cd "$(dirname "$0")" && pwd)
RT="$HERE/../../lang/runtime"
CC=${CC:-cc}
SRC="$HERE/test_geometry.c $RT/engram_geometry.c $RT/engram_store.c $RT/engram_vindex.c"
WARN="-std=c11 -Wall -Wextra"
TMP=$(mktemp -d)
echo "### PASS 1: PERF (optimised, un-sanitised) — functional gate"
$CC $WARN -O2 -I"$RT" $SRC -lm -o "$TMP/perf"
"$TMP/perf"
echo
echo "### PASS 2: SAFETY (ASan/UBSan)"
$CC $WARN -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -I"$RT" $SRC -lm -o "$TMP/safe"
ASAN_OPTIONS=${ASAN_OPTIONS:-detect_leaks=0} UBSAN_OPTIONS=halt_on_error=1 "$TMP/safe"
# PASS 3 (OPTIONAL): mean-centering discrimination bench on a COPY of a real
# store. Skips cleanly unless ENGRAM_BENCH_STORE points at a store .egm — never
# touches the live store. Read-only; not part of the pass/fail gate.
echo
echo "### PASS 3: DISCRIMINATION BENCH (optional; set ENGRAM_BENCH_STORE)"
BSRC="$HERE/bench_discrimination.c $RT/engram_geometry.c $RT/engram_store.c $RT/engram_vindex.c"
$CC $WARN -O2 -I"$RT" $BSRC -lm -o "$TMP/bench"
"$TMP/bench" "${ENGRAM_BENCH_STORE:-}"
+86
View File
@@ -0,0 +1,86 @@
#!/usr/bin/env bash
# M-INTEROCEPTION P0 gate: engram_scan_nodes_emb_json read-only builtin.
# Throwaway HOME + /tmp only. Never touches ~/.neuron or :8742.
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
GEO="$HERE/../../lang/runtime/engram_geometry.c"
VIDX="$HERE/../../lang/runtime/engram_vindex.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-p0-XXXXXX)"
export HOME="$WORK/home"; mkdir -p "$HOME"
unset ENGRAM_STORE
fail=0
echo "== compile (plain) =="
gcc -O1 -std=c11 -I "$INC" "$HERE/test_interoception_p0_emb.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p0" 2>"$WORK/cc.log" || { echo "COMPILE FAILED"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; }
D="$WORK/d"; mkdir -p "$D"
"$WORK/p0" "$D" || { echo "FAIL: run"; fail=1; }
echo
echo "== assertions =="
python3 - "$D" <<'PY'
import json, sys, os
d = sys.argv[1]
def load(n):
with open(os.path.join(d,n)) as f: return json.load(f)
rc = 0
def check(c,m):
global rc
print((" PASS: " if c else " FAIL: ")+m)
if not c: rc=1
alln = load("emb_all.json")
check(len(alln)==3, f"emb dump returns all 3 nodes (got {len(alln)})")
# salience-sorted: high, mid, low
labels=[n["label"] for n in alln]
check(labels==["emb-high","emb-mid","noemb-low"], f"salience-sorted order {labels}")
for n in alln:
L=len(n["emb"])
check(L==n["emb_dim"], f"{n['label']}: len(emb)={L} == emb_dim={n['emb_dim']}")
check(alln[0]["emb_dim"]==16 and alln[1]["emb_dim"]==16, "embedded nodes report dim 16")
check(alln[2]["emb_dim"]==0 and alln[2]["emb"]==[], "un-embedded node -> emb_dim 0, emb []")
# first emb value round-trips ~0.10
check(abs(alln[0]["emb"][0]-0.10)<1e-3, f"emb[0] round-trips (~0.10, got {alln[0]['emb'][0]})")
pg0=load("emb_pg0.json"); pg1=load("emb_pg1.json")
check(len(pg0)==1 and len(pg1)==1, "pagination: one node per page")
check(pg0[0]["id"]=="n-high" and pg1[0]["id"]=="n-mid", f"pages disjoint & ordered ({pg0[0]['id']},{pg1[0]['id']})")
plain=load("plain.json")
check(len(plain)==3, "existing scan_nodes_json still returns 3")
check(all("emb" not in n for n in plain), "existing scan_nodes_json carries NO emb (behavior-neutral)")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== latency (one 256-page over the 3-node copy) =="
python3 - "$D" <<'PY'
import os
# timing was measured inside C not here; report emb payload size as a proxy
sz=os.path.getsize(os.path.join(os.sys.argv[1] if False else __import__('sys').argv[1],"emb_all.json"))
print(f" emb_all.json payload = {sz} bytes for 3 nodes")
PY
echo
echo "== ASan+UBSan =="
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_interoception_p0_emb.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p0.san" 2>"$WORK/san_cc.log" || { echo "SAN COMPILE FAILED"; tail -20 "$WORK/san_cc.log"; fail=1; }
if [ -x "$WORK/p0.san" ]; then
export ASAN_OPTIONS=detect_leaks=0
DS="$WORK/ds"; mkdir -p "$DS"
"$WORK/p0.san" "$DS" >/dev/null 2>"$WORK/san_run.log"
if grep -qiE 'runtime error|AddressSanitizer|Sanitizer|ERROR: ' "$WORK/san_run.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san_run.log" | head; fail=1
else echo " ok: ASan+UBSan clean"; fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "====== P0 EMB-ENDPOINT GATE: PASS ======"; else echo "====== P0 EMB-ENDPOINT GATE: FAIL ======"; fi
rm -rf "$WORK"
exit $fail
+147
View File
@@ -0,0 +1,147 @@
#!/usr/bin/env bash
# M-INTEROCEPTION P1 gate: two-threshold consolidation (ENGRAM_CONSOLIDATION).
# Throwaway HOME + /tmp only. Never touches ~/.neuron or :8742.
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
GEO="$HERE/../../lang/runtime/engram_geometry.c"
VIDX="$HERE/../../lang/runtime/engram_vindex.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-p1-XXXXXX)"
export HOME="$WORK/home"; mkdir -p "$HOME"
unset ENGRAM_STORE ENGRAM_CONSOLIDATION ENGRAM_CONSOL_CONN_MIN ENGRAM_CONSOL_PERM_MIN ENGRAM_CONSOL_WM_TOPK
fail=0
echo "== compile =="
gcc -O1 -std=c11 -I "$INC" "$HERE/test_interoception_p1_consol.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p1" 2>"$WORK/cc.log" || { echo "COMPILE FAILED"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; }
echo
echo "== (a) HEADLINE: hebb accrual curve over N co-activations (flag OFF, pure trunk) =="
D="$WORK/a"; mkdir -p "$D"
( unset ENGRAM_CONSOLIDATION; "$WORK/p1" accrual "$D" ) >"$WORK/accrual.txt" 2>&1 || { echo "FAIL accrual run"; fail=1; }
python3 - "$WORK/accrual.txt" <<'PY'
import json,sys,re
rows=[]
for line in open(sys.argv[1]):
m=re.match(r'SAMPLE (\d+) (\{.*\})',line.strip())
if not m: continue
n=int(m.group(1)); j=json.loads(m.group(2))
hm=j.get("hebb_max",0.0); hc=j.get("hebb_cand_max",0.0)
rows.append((n,hm,hc))
print(" N hebb_max 1-0.9999^N (predicted EWMA)")
rc=0
for n,hm,hc in rows:
pred=1-0.9999**n
print(f" {n:<7} {hm:<12.6g} {pred:.6g}")
# assertions: monotonic rise, starts near ETA, tracks EWMA prediction
first=rows[0]; last=rows[-1]
def check(c,m):
global rc; print((" PASS: " if c else " FAIL: ")+m);
if not c: rc=1
check(abs(first[1]-0.0001)<5e-5, f"first sample hebb ~= ETA 0.0001 (got {first[1]:.6g})")
check(all(rows[i][1] <= rows[i+1][1]+1e-9 for i in range(len(rows)-1)), "hebb_max is monotonically non-decreasing over N")
check(last[1] > first[1]*50, f"hebb accrues substantially by N={last[0]} (got {last[1]:.4g} vs {first[1]:.4g})")
# EWMA fit: measured should be within 25% of 1-0.9999^N at the mid samples
mid=[r for r in rows if 100<=r[0]<=2000]
ok=all(abs(hm-(1-0.9999**n))/(1-0.9999**n) < 0.25 for n,hm,hc in mid)
check(ok, "measured curve tracks the 1-0.9999^N EWMA prediction within 25% (co-activation P~1)")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== (b) CONNECTION threshold: strong ISE wires to wm_top, weak ISE wires nothing (flag ON) =="
D="$WORK/b"; mkdir -p "$D"
( export ENGRAM_CONSOLIDATION=1; "$WORK/p1" connect "$D" ) >"$WORK/connect.txt" 2>&1 || { echo "FAIL connect run"; fail=1; }
cat "$WORK/connect.txt" | sed 's/^/ /'
python3 - "$WORK/connect.txt" "$D/connect.json" <<'PY'
import json,sys,re
txt=open(sys.argv[1]).read()
g=json.load(open(sys.argv[2]))
def field(k):
m=re.search(rf'{k} (\S+)',txt); return m.group(1) if m else None
sid=field("ISE_STRONG_ID"); wid=field("ISE_WEAK_ID")
m=re.search(r'EDGES before=(\d+) after_strong=(\d+) after_weak=(\d+)',txt)
before,aftS,aftW=int(m.group(1)),int(m.group(2)),int(m.group(3))
rc=0
def check(c,mm):
global rc; print((" PASS: " if c else " FAIL: ")+mm)
if not c: rc=1
strong_edges=[e for e in g["edges"] if e["from_id"]==sid and e["relation"]=="hebbian-associate"]
weak_edges=[e for e in g["edges"] if e["from_id"]==wid]
check(aftS>before, f"strong ISE formed connection edges ({before} -> {aftS})")
check(aftW==aftS, f"weak ISE formed NO edges ({aftS} -> {aftW})")
check(len(strong_edges)>=1, f"strong ISE has {len(strong_edges)} hebbian-associate edge(s) to wm_top")
check(all('consolidated-from-ISE' in (e.get('metadata') or '') for e in strong_edges),
"connection edges are provenance-tagged consolidated-from-ISE (reversible)")
check(len(weak_edges)==0, "weak ISE (below connection bar) has zero outgoing edges")
# targets must be the WM-top nodes (hebb-a / hebb-b), not distractors
tgt_labels=set()
byid={n["id"]:n for n in g["nodes"]}
for e in strong_edges:
t=byid.get(e["to_id"]);
if t: tgt_labels.add(t.get("label"))
print(f" connection targets: {sorted(tgt_labels)}")
check(tgt_labels.issubset({"hebb-a","hebb-b"}) and len(tgt_labels)>=1,
f"connections point at the wm_top nodes {sorted(tgt_labels)}")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== (c) PERMANENCE threshold: promoted node survives 48h prune, ephemeral is swept (flag ON) =="
D="$WORK/c"; mkdir -p "$D"
( export ENGRAM_CONSOLIDATION=1 ENGRAM_CONSOL_PERM_MIN=-1000; "$WORK/p1" perm "$D" ) >"$WORK/perm.txt" 2>&1 || { echo "FAIL perm run"; fail=1; }
cat "$WORK/perm.txt" | sed 's/^/ /'
python3 - "$WORK/perm.txt" <<'PY'
import sys,re,json
txt=open(sys.argv[1]).read()
rc=0
def check(c,m):
global rc; print((" PASS: " if c else " FAIL: ")+m)
if not c: rc=1
prom=int(re.search(r'PROMOTED (\d+)',txt).group(1))
m=re.search(r'NODES before=(\d+) after=(\d+) removed=(\d+)',txt)
before,after,removed=int(m.group(1)),int(m.group(2)),int(m.group(3))
dur=re.search(r'DURABLE_NODE (\{.*\})',txt).group(1)
eph=re.search(r'EPHEMERAL_NODE (\{.*\})',txt).group(1)
durj=json.loads(dur); ephj=json.loads(eph)
check(prom==1, "engram_consolidate_permanence promoted the node (returned 1)")
check(before==2 and after==1 and removed==1, f"exactly one node pruned ({before}->{after}, removed={removed})")
check(durj.get("id")=="ise-durable", "durable node SURVIVED the 48h telemetry prune")
check('consolidated-from-ISE' in (durj.get("metadata") or ''), "durable node carries reversible provenance marker")
check(ephj=={} or not ephj.get("id"), "ephemeral (non-permanent) ISE was swept")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== (d) OFF path byte-identical: ISE creation forms no edges, permanence is a no-op =="
D="$WORK/d"; mkdir -p "$D"
( unset ENGRAM_CONSOLIDATION; "$WORK/p1" offcheck "$D" ) >"$WORK/off.txt" 2>&1
rcoff=$?
cat "$WORK/off.txt" | sed 's/^/ /'
[ $rcoff -eq 0 ] && echo " PASS: flag OFF — ISE creation added 0 edges and permanence returned 0" \
|| { echo " FAIL: OFF path changed behavior"; fail=1; }
echo
echo "== ASan+UBSan (connect + perm + accrual-short) =="
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_interoception_p1_consol.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p1.san" 2>"$WORK/san_cc.log" || { echo "SAN COMPILE FAILED"; tail -25 "$WORK/san_cc.log"; fail=1; }
if [ -x "$WORK/p1.san" ]; then
export ASAN_OPTIONS=detect_leaks=0
DS="$WORK/san"; mkdir -p "$DS"
( export ENGRAM_CONSOLIDATION=1 ENGRAM_CONSOL_PERM_MIN=-1000; "$WORK/p1.san" connect "$DS" ) >/dev/null 2>"$WORK/san_run.log"
( export ENGRAM_CONSOLIDATION=1 ENGRAM_CONSOL_PERM_MIN=-1000; "$WORK/p1.san" perm "$DS" ) >/dev/null 2>>"$WORK/san_run.log"
if grep -qiE 'runtime error|AddressSanitizer|Sanitizer|ERROR: ' "$WORK/san_run.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san_run.log" | head; fail=1
else echo " ok: ASan+UBSan clean"; fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "====== P1 CONSOLIDATION GATE: PASS ======"; else echo "====== P1 CONSOLIDATION GATE: FAIL ======"; fi
rm -rf "$WORK"
exit $fail
+96
View File
@@ -0,0 +1,96 @@
#!/usr/bin/env bash
# M-INTEROCEPTION P2 gate: chronoception (ENGRAM_CHRONOCEPTION).
# Throwaway HOME + /tmp only. TC defaults to 3600s; we pin it for the math.
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
GEO="$HERE/../../lang/runtime/engram_geometry.c"
VIDX="$HERE/../../lang/runtime/engram_vindex.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-p2-XXXXXX)"
export HOME="$WORK/home"; mkdir -p "$HOME"
export ENGRAM_CHRONO_TC=3600 # pin cooling time-constant for the math
unset ENGRAM_STORE
fail=0
echo "== compile =="
gcc -O1 -std=c11 -I "$INC" "$HERE/test_interoception_p2_chrono.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p2" 2>"$WORK/cc.log" || { echo "COMPILE FAILED"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; }
sum_wm(){ python3 -c "import json,sys; g=json.load(open('$1')); print(sum(n.get('working_memory_weight',0) for n in g['nodes']))"; }
echo
echo "== (a) cooling scales with dt (flag ON) =="
for DT in 600000 1800000 3600000 7200000; do # 600s,1800s,3600s,7200s at TC=3600
D="$WORK/dt$DT"; mkdir -p "$D"
( export ENGRAM_CHRONOCEPTION=1; "$WORK/p2" once "$D" "$DT" ) >"$D/out.txt" 2>&1
MAG=$(grep MAGNITUDE "$D/out.txt" | awk '{print $2}')
WM=$(sum_wm "$D/field.json")
PRED=$(python3 -c "import math; print(round(1-math.exp(-$DT/1000/3600),6))")
echo " dt=${DT}ms magnitude=$MAG predicted 1-exp(-dt/TC)=$PRED field_wm_sum=$WM"
python3 -c "import sys; m=float('$MAG'); p=float('$PRED'); sys.exit(0 if abs(m-p)<1e-4 else 1)" \
&& echo " PASS: magnitude matches exp cooling" || { echo " FAIL"; fail=1; }
done
echo
echo "== (b) SCALE-INVARIANCE: age(dt) once == age(dt/N) N times (field within float tol) =="
DT=3600000
for N in 2 10 100; do
DA="$WORK/inv_once_$N"; DB="$WORK/inv_split_$N"; mkdir -p "$DA" "$DB"
( export ENGRAM_CHRONOCEPTION=1; "$WORK/p2" once "$DA" "$DT" ) >/dev/null 2>&1
( export ENGRAM_CHRONOCEPTION=1; "$WORK/p2" split "$DB" "$DT" "$N" ) >/dev/null 2>&1
WA=$(sum_wm "$DA/field.json"); WB=$(sum_wm "$DB/field.json")
echo " N=$N once_wm=$WA split_wm=$WB |delta|=$(python3 -c "print(abs($WA-$WB))")"
python3 -c "import sys; sys.exit(0 if abs($WA-$WB)<1e-9 else 1)" \
&& echo " PASS: scale-invariant within 1e-9" || { echo " FAIL: not scale-invariant"; fail=1; }
done
echo
echo "== (c) REBOOT catch-up: one-shot cooling from persisted last-tick, reports MAGNITUDE not seconds =="
D="$WORK/catch"; mkdir -p "$D"
GAP=3600000 # 1h unconscious
( export ENGRAM_CHRONOCEPTION=1 ENGRAM_DATA_DIR="$D"; "$WORK/p2" catchup "$D" "$GAP" ) >"$D/out.txt" 2>&1
CMAG=$(grep CATCHUP_MAGNITUDE "$D/out.txt" | awk '{print $2}')
CWM=$(sum_wm "$D/field.json")
PRED=$(python3 -c "import math; print(round(1-math.exp(-$GAP/1000/3600),4))")
echo " gap=${GAP}ms catchup_magnitude=$CMAG predicted=$PRED field_wm_sum=$CWM (was 0.6)"
python3 -c "import sys; sys.exit(0 if abs(float('$CMAG')-float('$PRED'))<1e-2 else 1)" \
&& echo " PASS: one-shot catch-up cooled by the elapsed gap, surfaced as a magnitude" \
|| { echo " FAIL"; fail=1; }
# honesty rail: magnitude is bounded [0,1), NOT an elapsed-seconds number
python3 -c "import sys; m=float('$CMAG'); sys.exit(0 if 0<=m<1 else 1)" \
&& echo " PASS: magnitude is a bounded drift signal in [0,1), never elapsed seconds" \
|| { echo " FAIL: magnitude out of [0,1)"; fail=1; }
echo
echo "== (d) OFF path: flag unset -> age & catchup return 0, field untouched =="
D="$WORK/off"; mkdir -p "$D"
( unset ENGRAM_CHRONOCEPTION; export ENGRAM_DATA_DIR="$D"; "$WORK/p2" offcheck "$D" 3600000 ) >"$D/out.txt" 2>&1
cat "$D/out.txt" | sed 's/^/ /'
OFFWM=$(sum_wm "$D/field.json")
# loaded field wm sum = (1.0+0.8+0.6)*0.5 halving = 1.2 ; must be UNCHANGED
echo " field_wm_sum=$OFFWM (expected 1.2, unchanged)"
python3 -c "import sys; sys.exit(0 if abs($OFFWM-1.2)<1e-9 else 1)" \
&& echo " PASS: OFF path leaves the field byte-identical (no aging)" \
|| { echo " FAIL: OFF path modified the field"; fail=1; }
echo
echo "== ASan+UBSan =="
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_interoception_p2_chrono.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p2.san" 2>"$WORK/san_cc.log" || { echo "SAN COMPILE FAILED"; tail -25 "$WORK/san_cc.log"; fail=1; }
if [ -x "$WORK/p2.san" ]; then
export ASAN_OPTIONS=detect_leaks=0
DS="$WORK/san"; mkdir -p "$DS"
( export ENGRAM_CHRONOCEPTION=1 ENGRAM_DATA_DIR="$DS"; "$WORK/p2.san" once "$DS" 3600000 ) >/dev/null 2>"$WORK/san.log"
( export ENGRAM_CHRONOCEPTION=1 ENGRAM_DATA_DIR="$DS"; "$WORK/p2.san" catchup "$DS" 3600000 ) >/dev/null 2>>"$WORK/san.log"
if grep -qiE 'runtime error|AddressSanitizer|Sanitizer|ERROR: ' "$WORK/san.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san.log" | head; fail=1
else echo " ok: ASan+UBSan clean"; fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "====== P2 CHRONOCEPTION GATE: PASS ======"; else echo "====== P2 CHRONOCEPTION GATE: FAIL ======"; fi
rm -rf "$WORK"
exit $fail
+68
View File
@@ -0,0 +1,68 @@
#!/usr/bin/env bash
# M-INTEROCEPTION P3 gate: drift-sensor primitive engram_geo_displacement.
# Read-only pure primitive; no store, no flag. Throwaway /tmp only.
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
GEO="$HERE/../../lang/runtime/engram_geometry.c"
VIDX="$HERE/../../lang/runtime/engram_vindex.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-p3-XXXXXX)"
export HOME="$WORK/home"; mkdir -p "$HOME"
fail=0
echo "== compile =="
gcc -O1 -std=c11 -I "$INC" "$HERE/test_interoception_p3_drift.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p3" 2>"$WORK/cc.log" || { echo "COMPILE FAILED"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; }
"$WORK/p3" > "$WORK/out.txt" 2>&1 || { echo "FAIL run"; cat "$WORK/out.txt"; fail=1; }
cat "$WORK/out.txt" | sed 's/^/ /'
echo
echo "== assertions =="
python3 - "$WORK/out.txt" <<'PY'
import sys,re
rows={}
for line in open(sys.argv[1]):
m=re.match(r'(\w+) (.*)',line.strip())
if not m: continue
tag=m.group(1); kv=dict(re.findall(r'(\w+)=([-\d.]+)',m.group(2)))
rows[tag]={k:float(v) for k,v in kv.items()}
rc=0
def check(c,msg):
global rc; print((" PASS: " if c else " FAIL: ")+msg)
if not c: rc=1
g=rows["GROWTH"]; c=rows["CORRUPTION"]; i=rows["IDENTITY"]
check(g["core_disp"]<0.05, f"GROWTH: core displacement ~0 (core fixed) = {g['core_disp']}")
check(g["periph_disp"]>0.30, f"GROWTH: periphery extended = {g['periph_disp']}")
check(g["centroid_sep"]<1e-6, f"GROWTH: centroid unmoved = {g['centroid_sep']}")
check(abs(g["radius_delta"]-0.4)<1e-4, f"GROWTH: radius grew by ~0.4 = {g['radius_delta']}")
check(c["core_disp"]>0.40, f"CORRUPTION: core displaced strongly = {c['core_disp']}")
check(c["periph_disp"]<0.05, f"CORRUPTION: periphery fixed = {c['periph_disp']}")
check(c["centroid_sep"]>0.1, f"CORRUPTION: centroid moved = {c['centroid_sep']}")
check(c["core_disp"] > 8*g["core_disp"]+0.3,
f"SENSOR DISCRIMINATES: corruption core_disp ({c['core_disp']}) >> growth core_disp ({g['core_disp']})")
check(i["core_disp"]==0 and i["periph_disp"]==0 and i["centroid_sep"]<1e-6,
"IDENTITY: A vs A -> zero drift")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== ASan+UBSan =="
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_interoception_p3_drift.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p3.san" 2>"$WORK/san_cc.log" || { echo "SAN COMPILE FAILED"; tail -25 "$WORK/san_cc.log"; fail=1; }
if [ -x "$WORK/p3.san" ]; then
export ASAN_OPTIONS=detect_leaks=0
"$WORK/p3.san" >/dev/null 2>"$WORK/san.log"
if grep -qiE 'runtime error|AddressSanitizer|Sanitizer|ERROR: ' "$WORK/san.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san.log" | head; fail=1
else echo " ok: ASan+UBSan clean"; fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "====== P3 DRIFT-SENSOR GATE: PASS ======"; else echo "====== P3 DRIFT-SENSOR GATE: FAIL ======"; fi
rm -rf "$WORK"
exit $fail
+69
View File
@@ -0,0 +1,69 @@
#!/usr/bin/env bash
# M-INTEROCEPTION P4 gate: afferent input counters in act-stats (additive).
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
GEO="$HERE/../../lang/runtime/engram_geometry.c"
VIDX="$HERE/../../lang/runtime/engram_vindex.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-p4-XXXXXX)"
export HOME="$WORK/home"; mkdir -p "$HOME"
unset ENGRAM_STORE
fail=0
echo "== compile =="
gcc -O1 -std=c11 -I "$INC" "$HERE/test_interoception_p4_afferent.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p4" 2>"$WORK/cc.log" || { echo "COMPILE FAILED"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; }
"$WORK/p4" > "$WORK/out.txt" 2>&1 || { echo "FAIL run"; cat "$WORK/out.txt"; fail=1; }
grep -oE 'aff_[a-z_]+":[0-9]+' "$WORK/out.txt" | sed 's/^/ /' | head -30
echo
echo "== assertions =="
python3 - "$WORK/out.txt" <<'PY'
import sys,re,json
S={}
for line in open(sys.argv[1]):
m=re.match(r'(STATS\d) (\{.*\})',line.strip())
if m: S[m.group(1)]=json.loads(m.group(2))
rc=0
def check(c,msg):
global rc; print((" PASS: " if c else " FAIL: ")+msg)
if not c: rc=1
s0,s1,s2=S["STATS0"],S["STATS1"],S["STATS2"]
# after creation, before any query
check(s0["aff_node_creates"]==5, f"node_creates==5 (got {s0['aff_node_creates']})")
check(s0["aff_ise_ingests"]==2, f"ise_ingests==2 (got {s0['aff_ise_ingests']})")
check(s0["aff_edge_creates"]==2, f"edge_creates==2 (got {s0['aff_edge_creates']})")
check(s0["aff_queries"]==0 and s0["aff_activations"]==0, "queries/activations start at 0")
# after 4 queries
check(s1["aff_queries"]==4, f"queries==4 (got {s1['aff_queries']})")
check(s1["aff_activations"]==4, f"activations==4 (got {s1['aff_activations']})")
check(s1["aff_node_creates"]==5 and s1["aff_ise_ingests"]==2 and s1["aff_edge_creates"]==2,
"create counters unchanged by queries")
# after 3 more queries — monotonic
check(s2["aff_queries"]==7, f"queries==7 monotonic (got {s2['aff_queries']})")
check(s2["aff_activations"]==7, f"activations==7 monotonic (got {s2['aff_activations']})")
check(s2["aff_queries"]>s1["aff_queries"]>s0["aff_queries"], "queries strictly monotonic across readings")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== ASan+UBSan =="
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_interoception_p4_afferent.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p4.san" 2>"$WORK/san_cc.log" || { echo "SAN COMPILE FAILED"; tail -25 "$WORK/san_cc.log"; fail=1; }
if [ -x "$WORK/p4.san" ]; then
export ASAN_OPTIONS=detect_leaks=0
"$WORK/p4.san" >/dev/null 2>"$WORK/san.log"
if grep -qiE 'runtime error|AddressSanitizer|Sanitizer|ERROR: ' "$WORK/san.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san.log" | head; fail=1
else echo " ok: ASan+UBSan clean"; fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "====== P4 AFFERENT-COUNTERS GATE: PASS ======"; else echo "====== P4 AFFERENT-COUNTERS GATE: FAIL ======"; fi
rm -rf "$WORK"
exit $fail
+74
View File
@@ -0,0 +1,74 @@
#!/usr/bin/env bash
# M-INTEROCEPTION P5 gate: dream-recall builtin engram_dreams_json (honesty rail).
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
GEO="$HERE/../../lang/runtime/engram_geometry.c"
VIDX="$HERE/../../lang/runtime/engram_vindex.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-p5-XXXXXX)"
export HOME="$WORK/home"; mkdir -p "$HOME"
unset ENGRAM_STORE
fail=0
echo "== compile =="
gcc -O1 -std=c11 -I "$INC" "$HERE/test_interoception_p5_dreams.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p5" 2>"$WORK/cc.log" || { echo "COMPILE FAILED"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; }
D="$WORK/d"; mkdir -p "$D"
"$WORK/p5" "$D" > "$WORK/out.txt" 2>&1 || { echo "FAIL run"; cat "$WORK/out.txt"; fail=1; }
cat "$WORK/out.txt" | sed 's/^/ /'
echo
echo "== assertions =="
python3 - "$WORK/out.txt" <<'PY'
import sys,re,json
L={}
for line in open(sys.argv[1]):
line=line.strip()
m=re.match(r'(BEFORE|AFTER) (\[.*\])',line)
if m: L[m.group(1)]=json.loads(m.group(2)); continue
m=re.match(r'PRUNED (\d+)',line)
if m: L['PRUNED']=int(m.group(1)); continue
m=re.match(r'SINCE (\d+) (\[.*\])',line)
if m: L['SINCE']=json.loads(m.group(2))
rc=0
def check(c,msg):
global rc; print((" PASS: " if c else " FAIL: ")+msg)
if not c: rc=1
before_ids={d["id"] for d in L["BEFORE"]}
after_ids={d["id"] for d in L["AFTER"]}
since_ids={d["id"] for d in L["SINCE"]}
check(before_ids=={"cur_old","cur_mid","cur_recent"}, f"before prune: all 3 curiosity_scan, heartbeat excluded (got {sorted(before_ids)})")
check("hb_recent" not in before_ids, "heartbeat ISE never appears (not a dream)")
check(L["PRUNED"]==1, f"prune rotated out exactly the ancient ISE (pruned={L['PRUNED']})")
check(after_ids=={"cur_mid","cur_recent"}, f"after prune: rotated-out cur_old is ABSENT, not confabulated (got {sorted(after_ids)})")
check("cur_old" not in after_ids, "honesty rail: pruned dream is gone = 'I don't remember', never synthesized")
check(since_ids=={"cur_recent"}, f"since filter returns only events after the cutoff (got {sorted(since_ids)})")
# no fabrication: every returned id was one we seeded
seeded={"cur_old","cur_mid","cur_recent","hb_recent"}
allret=before_ids|after_ids|since_ids
check(allret<=seeded, f"no fabricated entries — every returned id was seeded ({sorted(allret)})")
sys.exit(rc)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== ASan+UBSan =="
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_interoception_p5_dreams.c" "$RT" "$ST" "$GEO" "$VIDX" \
-lcurl -lm -o "$WORK/p5.san" 2>"$WORK/san_cc.log" || { echo "SAN COMPILE FAILED"; tail -25 "$WORK/san_cc.log"; fail=1; }
if [ -x "$WORK/p5.san" ]; then
export ASAN_OPTIONS=detect_leaks=0
DS="$WORK/ds"; mkdir -p "$DS"
"$WORK/p5.san" "$DS" >/dev/null 2>"$WORK/san.log"
if grep -qiE 'runtime error|AddressSanitizer|Sanitizer|ERROR: ' "$WORK/san.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san.log" | head; fail=1
else echo " ok: ASan+UBSan clean"; fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "====== P5 DREAM-RECALL GATE: PASS ======"; else echo "====== P5 DREAM-RECALL GATE: FAIL ======"; fi
rm -rf "$WORK"
exit $fail
+137
View File
@@ -0,0 +1,137 @@
#!/usr/bin/env bash
# M7 index-driven-traversal gate. Pure C harness (NOT elb/elc): links the real
# el_runtime.c engram builtins + engram_store.c and drives ENGRAM_STORE off vs on.
# Proves (1) byte-identical activation parity flag-on == flag-off across a
# mutating query sequence, and (2) the O(E)-rebuild cost is eliminated flag-on.
# Writes ONLY under a throwaway /tmp dir with a throwaway HOME.
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
RT="$HERE/../../lang/runtime/el_runtime.c"
ST="$HERE/../../lang/runtime/engram_store.c"
INC="$HERE/../../lang/runtime"
WORK="$(mktemp -d /tmp/engram-m7-XXXXXX)"
DATA="$WORK/data"; mkdir -p "$DATA"
BIN="$WORK/m7"
export HOME="$WORK/home"; mkdir -p "$HOME" # never touch real ~/.neuron
# Hermetic: point the embedder at a guaranteed-refused endpoint so eg_embed_fetch
# fails fast, the circuit breaker opens, and cosq is deterministically absent in
# EVERY run (no dependence on whether a dev Ollama happens to be listening). This
# makes the byte-identical parity comparison reproducible and non-flaky.
export EL_EMBED_URL="http://127.0.0.1:1/api/embeddings"
unset ENGRAM_STORE
fail=0
echo "== compiling harness (gcc: el_runtime.c + engram_store.c + test_m7_traversal.c) =="
gcc -O2 -std=c11 -I "$INC" "$HERE/test_m7_traversal.c" "$RT" "$ST" -lcurl -lm -o "$BIN" 2>"$WORK/cc.log"
if [ $? -ne 0 ]; then echo "COMPILE FAILED:"; cat "$WORK/cc.log"; rm -rf "$WORK"; exit 1; fi
echo " ok: compiled"
echo
echo "== 1) PARITY: index-driven (M7 incremental) activation must be IDENTICAL to the"
echo " full-rebuild scan path — proven under one identical ENGRAM_STORE=1 state,"
echo " so the ONLY variable is how per-node adjacency is maintained."
echo " (compared on deterministic fields: node label + activation_strength +"
echo " working_memory_weight + epistemic_confidence + hops + promoted, IN ORDER;"
echo " node id/timestamps are per-run random and are intentionally excluded.)"
( unset ENGRAM_STORE; "$BIN" parity-off "$DATA" ) || { echo "FAIL: parity-off run"; fail=1; }
ENGRAM_STORE=1 "$BIN" parity-on-rebuild "$DATA" || { echo "FAIL: parity-on-rebuild run"; fail=1; }
ENGRAM_STORE=1 "$BIN" parity-on-incr "$DATA" || { echo "FAIL: parity-on-incr run"; fail=1; }
python3 - "$DATA" <<'PY' || fail=1
import json, sys, os
d = sys.argv[1]
def proj(prefix, i):
a = json.load(open(os.path.join(d, f"{prefix}_act{i}.json")))
out = []
for e in a:
n = e.get("node", {})
out.append([n.get("label",""),
e.get("activation_strength"), e.get("working_memory_weight"),
e.get("epistemic_confidence"), e.get("hops"), e.get("promoted")])
return out
def compare(label, pa, pb, gate):
rc = 0
for i in (1,2,3,4):
a, b = proj(pa, i), proj(pb, i)
if a == b:
print(f" #{i} identical (entries={len(a)}, promoted={sum(1 for r in a if r[5])})")
else:
if gate: rc = 1
print(f" #{i} DIFFERS ({'FAIL' if gate else 'note'})")
for x,y in zip(a,b):
if x != y:
print(f" first diff:\n {pa}={x}\n {pb}={y}"); break
if len(a) != len(b): print(f" length: {pa}={len(a)} {pb}={len(b)}")
print(f" {'PASS' if rc==0 else 'FAIL'}: {label}")
return rc
print(" [CORE M7 GATE] flag-on incremental index == flag-on forced full rebuild:")
rc1 = compare("index-driven activation == full-rebuild scan (same flag state)",
"onincr", "onrb", gate=True)
print(" [context] flag-on incremental index vs flag-off scan path (today's behavior):")
rc2 = compare("M7 (flag-on) == flag-off scan path", "onincr", "off", gate=False)
print(" [context] flag-off scan vs flag-on forced rebuild (isolates any pre-existing")
print(" flag-on/off float difference, INDEPENDENT of M7's incremental path):")
rc3 = compare("flag-off == flag-on (both rebuild path)", "off", "onrb", gate=False)
sys.exit(rc1) # only the core M7 equivalence gates the result
PY
echo
echo "== 2) PERF: ~13k nodes / 43k edges, 200 (add-edge + activate) iterations =="
NODES=13000; EDGES=43000; ITERS=120
( unset ENGRAM_STORE; "$BIN" perf off "$DATA" "$NODES" "$EDGES" "$ITERS" ) | tee "$WORK/perf_off.txt"
[ ${PIPESTATUS[0]} -ne 0 ] && { echo "FAIL: perf off"; fail=1; }
ENGRAM_STORE=1 "$BIN" perf on "$DATA" "$NODES" "$EDGES" "$ITERS" | tee "$WORK/perf_on.txt"
[ ${PIPESTATUS[0]} -ne 0 ] && { echo "FAIL: perf on"; fail=1; }
python3 - "$WORK/perf_off.txt" "$WORK/perf_on.txt" <<'PY'
import re, sys
def parse(f):
t = open(f).read()
def g(k):
m = re.search(k+r'=([\d.]+)', t); return float(m.group(1)) if m else 0.0
return {'rw': g('rebuild_edge_work'), 'rb': g('rebuilds'), 'ap': g('incr_appends'),
'loop_s': g('loop='), 'maint': g('adj_maint'),
'perq': g('per_query')}
off, on = parse(sys.argv[1]), parse(sys.argv[2])
def ratio(a,b): return (a/b) if b else float('inf')
print()
print(f" ADJACENCY TRAVERSAL COST (the metric M7 changes):")
print(f" edge-touches in rebuilds: off={off['rw']:.0f} on={on['rw']:.0f} "
f"({ratio(off['rw'],on['rw']):.0f}x fewer on)")
print(f" full O(E) rebuilds: off={off['rb']:.0f} on={on['rb']:.0f}")
print(f" incremental O(1) appends: off={off['ap']:.0f} on={on['ap']:.0f}")
print(f" adjacency-maint wall-time: off={off['maint']:.4f}s on={on['maint']:.4f}s "
f"({ratio(off['maint'],on['maint']):.1f}x faster on)")
print(f" END-TO-END per-query time: off={off['perq']:.2f}ms on={on['perq']:.2f}ms")
print(f" (per-query is dominated by activation's O(N) node scoring over 13k nodes,")
print(f" which M7 does not touch; the delta is the eliminated rebuild time.)")
ok = on['rw'] < off['rw'] and on['maint'] < off['maint'] and on['rb'] < off['rb']
print(" PASS: flag-on eliminates the O(E) per-query rebuild (fewer edge-touches, less maint time)"
if ok else " FAIL: expected fewer edge-touches AND less adjacency-maint time on flag-on")
sys.exit(0 if ok else 1)
PY
[ $? -ne 0 ] && fail=1
echo
echo "== 3) ASan+UBSan clean across parity + a small perf loop (leaks off — harness intentionally leaks el_strdup) =="
SANBIN="$WORK/m7.san"
gcc -O1 -g -std=c11 -fsanitize=address,undefined -fno-sanitize-recover=undefined \
-I "$INC" "$HERE/test_m7_traversal.c" "$RT" "$ST" -lcurl -lm -o "$SANBIN" 2>"$WORK/san_cc.log"
if [ $? -ne 0 ]; then echo " SAN COMPILE FAILED:"; tail -20 "$WORK/san_cc.log"; fail=1; else
export ASAN_OPTIONS=detect_leaks=0
D2="$WORK/data2"; mkdir -p "$D2"
( unset ENGRAM_STORE; "$SANBIN" parity-off "$D2" ) >/dev/null 2>"$WORK/san_run.log" && \
ENGRAM_STORE=1 "$SANBIN" parity-on-rebuild "$D2" >/dev/null 2>>"$WORK/san_run.log" && \
ENGRAM_STORE=1 "$SANBIN" parity-on-incr "$D2" >/dev/null 2>>"$WORK/san_run.log" && \
( unset ENGRAM_STORE; "$SANBIN" perf off "$D2" 1500 5000 40 ) >/dev/null 2>>"$WORK/san_run.log" && \
ENGRAM_STORE=1 "$SANBIN" perf on "$D2" 1500 5000 40 >/dev/null 2>>"$WORK/san_run.log"
if grep -qiE 'runtime error|AddressSanitizer|UndefinedBehavior|ERROR: ' "$WORK/san_run.log"; then
echo " FAIL: sanitizer findings:"; grep -iE 'runtime error|Sanitizer|ERROR' "$WORK/san_run.log" | head; fail=1
else
echo " ok: ASan+UBSan clean across parity + perf (rebuild + incremental append + BFS)"
fi
fi
echo
if [ "$fail" -eq 0 ]; then echo "================ M7 TRAVERSAL GATE: PASS ================"; else echo "================ M7 TRAVERSAL GATE: FAIL ================"; fi
rm -rf "$WORK"
exit $fail
+23
View File
@@ -0,0 +1,23 @@
#!/bin/sh
# Build + RUN the REASONING-layer tests (engram_reason.c): closed-form constructed
# cases for ANALOGY / INDUCTION / ABDUCTION / CAUSAL / PLANNING, each composing the
# §5 geometry OPERATORS (engram_geometry.c). Pure C11 (stdlib + libm). Standalone —
# NOT folded through elc. Two passes:
# 1. PERF — optimised (-O2, no sanitizer): the functional gate.
# 2. SAFETY — ASan + UBSan on the same suite (memory-safety is size-independent).
set -e
HERE=$(cd "$(dirname "$0")" && pwd)
RT="$HERE/../../lang/runtime"
CC=${CC:-cc}
SRC="$HERE/test_reason.c $RT/engram_reason.c $RT/engram_geometry.c $RT/engram_store.c $RT/engram_vindex.c"
WARN="-std=c11 -Wall -Wextra"
TMP=$(mktemp -d)
echo "### PASS 1: PERF (optimised, un-sanitised) — functional gate"
$CC $WARN -O2 -I"$RT" $SRC -lm -o "$TMP/perf"
"$TMP/perf"
echo
echo "### PASS 2: SAFETY (ASan/UBSan)"
$CC $WARN -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -I"$RT" $SRC -lm -o "$TMP/safe"
ASAN_OPTIONS=${ASAN_OPTIONS:-detect_leaks=0} UBSAN_OPTIONS=halt_on_error=1 "$TMP/safe"
+24
View File
@@ -0,0 +1,24 @@
#!/bin/sh
# Build + RUN the VERIFIER-layer tests (engram_verify.c): closed-form constructed
# cases for GROUNDING (anti-hallucination) and CONSISTENCY (polarity/negation
# inversion + geometric contradiction), each composing the reasoning point-fit
# (engram_reason.c) and the §5 geometry OPERATORS (engram_geometry.c). Pure C11
# (stdlib + libm). Standalone — NOT folded through elc. Two passes:
# 1. PERF — optimised (-O2, no sanitizer): the functional gate.
# 2. SAFETY — ASan + UBSan on the same suite (memory-safety is size-independent).
set -e
HERE=$(cd "$(dirname "$0")" && pwd)
RT="$HERE/../../lang/runtime"
CC=${CC:-cc}
SRC="$HERE/test_verify.c $RT/engram_verify.c $RT/engram_reason.c $RT/engram_geometry.c $RT/engram_store.c $RT/engram_vindex.c"
WARN="-std=c11 -Wall -Wextra"
TMP=$(mktemp -d)
echo "### PASS 1: PERF (optimised, un-sanitised) — functional gate"
$CC $WARN -O2 -I"$RT" $SRC -lm -o "$TMP/perf"
"$TMP/perf"
echo
echo "### PASS 2: SAFETY (ASan/UBSan)"
$CC $WARN -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -I"$RT" $SRC -lm -o "$TMP/safe"
ASAN_OPTIONS=${ASAN_OPTIONS:-detect_leaks=0} UBSAN_OPTIONS=halt_on_error=1 "$TMP/safe"
+25
View File
@@ -0,0 +1,25 @@
#!/bin/sh
# Build + RUN the M8 HNSW vector-index tests. Pure C11 (gcc/cc), stdlib + libm
# only. This is a standalone C module — NOT folded through elb/elc.
#
# Two passes:
# 1. PERF — optimised (-O2, no sanitizer): the real recall@10 gate + speedup
# numbers at full size (N=5000 recall, N=5000/20000 speedup).
# 2. SAFETY — ASan + UBSan on the same suite at reduced size (VINDEX_QUICK=1);
# memory-safety is size-independent, so this stays fast.
set -e
HERE=$(cd "$(dirname "$0")" && pwd)
RT="$HERE/../../lang/runtime"
CC=${CC:-cc}
SRC="$HERE/test_vindex.c $RT/engram_vindex.c $RT/engram_store.c"
WARN="-std=c11 -Wall -Wextra"
TMP=$(mktemp -d)
echo "### PASS 1: PERF (optimised, un-sanitised) — recall gate + speedup"
$CC $WARN -O2 -I"$RT" $SRC -lm -o "$TMP/perf"
"$TMP/perf"
echo
echo "### PASS 2: SAFETY (ASan/UBSan, reduced size)"
$CC $WARN -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -I"$RT" $SRC -lm -o "$TMP/safe"
VINDEX_QUICK=1 ASAN_OPTIONS=${ASAN_OPTIONS:-detect_leaks=0} UBSAN_OPTIONS=halt_on_error=1 "$TMP/safe"
+496
View File
@@ -0,0 +1,496 @@
/* test_bufpool.c — M4 gate for the demand-paging BUFFER POOL (engram_store.{c,h}).
*
* Pure C. Build: gcc -O2 test_bufpool.c ../../lang/runtime/engram_store.c -o t
* Writes ONLY under a throwaway /tmp dir. Never touches ~/.neuron or live ports.
*
* Proves the M4 pool preserves every M1/M2 invariant when the pool is SMALLER
* than the store (pages evict + re-fault): small-pool round-trip correctness,
* LRU eviction policy (hot resident / cold evicted / no dirty stolen), pinned
* residency (superblocks, index roots, explicit page + hot-layer pins), bounded
* read-ahead, and crash safety (WAL replay + checkpoint-crash) under paging.
*/
#include "../../lang/runtime/engram_store.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdint.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>
static int g_pass = 0, g_fail = 0;
static void ok(const char* name, int cond){
printf(" [%s] %s\n", cond ? "PASS" : "FAIL", name);
if (cond) g_pass++; else g_fail++;
}
static char g_dir[512];
static void mk_dir(void){
snprintf(g_dir, sizeof g_dir, "/tmp/engram-bufpool-test-%d", (int)getpid());
mkdir(g_dir, 0700);
}
static void path_in(char* out, size_t cap, const char* name){
snprintf(out, cap, "%s/%s", g_dir, name);
}
/* ── deterministic generators (bit-exact regeneration for oracles) ─────────── */
static uint64_t xs(uint64_t* s){ uint64_t x=*s; x^=x<<13; x^=x>>7; x^=x<<17; *s=x; return x; }
static uint64_t node_seed(int i){ return 0x9E3779B97F4A7C15ULL ^ ((uint64_t)(i+1)*0xD1B54A32D192ED03ULL); }
static uint64_t edge_seed(int i){ return 0xC2B2AE3D27D4EB4FULL ^ ((uint64_t)(i+1)*0x165667B19E3779F9ULL); }
static char* rnd_str(uint64_t* st, size_t len){
char* s = (char*)malloc(len + 1);
for (size_t i=0;i<len;i++) s[i] = (char)(33 + (xs(st) % 94));
s[len] = 0; return s;
}
#define NODE_COUNT 5000
#define EDGE_COUNT 20000
#define EMB_DIM 768
#define CK_NODES 300
static void noop_node_cb(const StoreNode* n, void* ctx){ (void)n; (void)ctx; }
static void gen_node(int i, StoreNode* n){
memset(n, 0, sizeof *n);
uint64_t st = node_seed(i);
char id[32]; snprintf(id, sizeof id, "node-%d", i);
n->id = strdup(id);
size_t clen = (i % 500 == 0) ? (size_t)(17000 + (xs(&st) % 6000)) : (size_t)(xs(&st) % 300);
n->content = rnd_str(&st, clen);
n->node_type = rnd_str(&st, 4 + (xs(&st) % 8));
n->label = (i % 2) ? rnd_str(&st, 3 + (xs(&st) % 10)) : NULL;
n->tier = rnd_str(&st, 4 + (xs(&st) % 6));
n->tags = rnd_str(&st, xs(&st) % 40);
n->metadata = (i % 3) ? rnd_str(&st, xs(&st) % 60) : NULL;
n->salience = (double)(xs(&st) % 1000000) / 997.0;
n->importance = (double)(xs(&st) % 1000000) / 131.0;
n->confidence = (double)(xs(&st) % 1000000) / 733.0;
n->temporal_decay_rate = (double)(xs(&st) % 1000000) / 101.0;
n->activation_count = (int64_t)(xs(&st) % 100000);
n->last_activated = (int64_t)xs(&st);
n->created_at = (int64_t)(1600000000000LL + i);
n->updated_at = (int64_t)xs(&st);
n->background_activation = (double)(xs(&st) % 1000000) / 17.0;
n->working_memory_weight = (double)(xs(&st) % 1000000) / 29.0;
n->suppression_count = (int32_t)(xs(&st) % 50);
n->layer_id = (uint32_t)(xs(&st) % 5);
for (int k=0;k<STORE_BLL_K;k++) n->access_ts[k] = (int64_t)xs(&st);
n->access_head = (int32_t)(xs(&st) % STORE_BLL_K);
n->access_filled = (int32_t)(xs(&st) % (STORE_BLL_K + 1));
n->wm_anchor = (double)(xs(&st) % 1000000) / 3.0;
n->emb = (float*)malloc(EMB_DIM * sizeof(float));
for (int k=0;k<EMB_DIM;k++){ uint32_t u=(uint32_t)xs(&st); memcpy(&n->emb[k], &u, 4); }
n->emb_dim = EMB_DIM;
}
static void gen_edge(int i, StoreEdge* e){
memset(e, 0, sizeof *e);
uint64_t st = edge_seed(i);
char id[32], from[32], to[32];
snprintf(id, sizeof id, "edge-%d", i);
snprintf(from, sizeof from, "node-%d", (int)(xs(&st) % NODE_COUNT));
snprintf(to, sizeof to, "node-%d", (int)(xs(&st) % NODE_COUNT));
e->id = strdup(id); e->from_id = strdup(from); e->to_id = strdup(to);
e->relation = rnd_str(&st, 3 + (xs(&st) % 12));
e->metadata = (i % 4) ? rnd_str(&st, xs(&st) % 40) : NULL;
e->weight = (double)(xs(&st) % 1000000) / 111.0;
e->hebb = (double)(xs(&st) % 1000000) / 1000000.0;
e->confidence = (double)(xs(&st) % 1000000) / 777.0;
e->created_at = (int64_t)(1600000000000LL + i);
e->updated_at = (int64_t)xs(&st);
e->last_fired = (int64_t)xs(&st);
e->inhibitory = (int32_t)(xs(&st) % 2);
e->layer_id = (uint32_t)(xs(&st) % 5);
}
static int streq(const char* a, const char* b){
if (!a && !b) return 1;
if (!a || !b) return 0;
return strcmp(a,b)==0;
}
static int cmp_node(const StoreNode* a, const StoreNode* b){
if (!streq(a->id,b->id) || !streq(a->content,b->content) ||
!streq(a->node_type,b->node_type) || !streq(a->label,b->label) ||
!streq(a->tier,b->tier) || !streq(a->tags,b->tags) ||
!streq(a->metadata,b->metadata)) return 0;
if (a->salience!=b->salience || a->importance!=b->importance ||
a->confidence!=b->confidence || a->temporal_decay_rate!=b->temporal_decay_rate ||
a->activation_count!=b->activation_count || a->last_activated!=b->last_activated ||
a->created_at!=b->created_at || a->updated_at!=b->updated_at ||
a->background_activation!=b->background_activation ||
a->working_memory_weight!=b->working_memory_weight ||
a->suppression_count!=b->suppression_count || a->layer_id!=b->layer_id ||
a->access_head!=b->access_head || a->access_filled!=b->access_filled ||
a->wm_anchor!=b->wm_anchor || a->emb_dim!=b->emb_dim) return 0;
for (int k=0;k<STORE_BLL_K;k++) if (a->access_ts[k]!=b->access_ts[k]) return 0;
if ((a->emb==NULL) != (b->emb==NULL)) return 0;
if (a->emb && memcmp(a->emb, b->emb, (size_t)a->emb_dim*4)!=0) return 0;
return 1;
}
static int cmp_edge(const StoreEdge* a, const StoreEdge* b){
if (!streq(a->id,b->id) || !streq(a->from_id,b->from_id) || !streq(a->to_id,b->to_id) ||
!streq(a->relation,b->relation) || !streq(a->metadata,b->metadata)) return 0;
if (a->weight!=b->weight || a->hebb!=b->hebb || a->confidence!=b->confidence ||
a->created_at!=b->created_at || a->updated_at!=b->updated_at ||
a->last_fired!=b->last_fired || a->inhibitory!=b->inhibitory ||
a->layer_id!=b->layer_id) return 0;
return 1;
}
static void free_node_fields(StoreNode* n){
free(n->id); free(n->content); free(n->node_type); free(n->label);
free(n->tier); free(n->tags); free(n->metadata); free(n->emb); free(n->unknown);
}
static void free_edge_fields(StoreEdge* e){
free(e->id); free(e->from_id); free(e->to_id); free(e->relation); free(e->metadata); free(e->unknown);
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 1 — SMALL-POOL CORRECTNESS: full M1 workload (5k nodes / 20k edges) with
* a frame budget FAR smaller than the store → constant eviction + re-fault, yet
* every read is bit-exact and the pool stays bounded.
* ════════════════════════════════════════════════════════════════════════════ */
static void test_small_pool_roundtrip(void){
printf("\n== 1) small-pool correctness: %d nodes + %d edges, cap=%d frames ==\n",
NODE_COUNT, EDGE_COUNT, 32);
char path[600]; path_in(path, sizeof path, "small.store");
unlink(path);
EngramPagedStore* s = store_create(path);
ok("store_create", s != NULL);
if (!s) return;
store__set_pool_frames(s, 32); /* pool << store */
for (int i=0;i<NODE_COUNT;i++){
StoreNode n; gen_node(i,&n);
if (store_put_node(s,&n)!=0){ ok("put_node", 0); free_node_fields(&n); store_close(s); return; }
free_node_fields(&n);
if ((i%500)==499) store_sync(s); /* checkpoint: dirty→clean so frames evictable */
}
for (int i=0;i<EDGE_COUNT;i++){
StoreEdge e; gen_edge(i,&e);
if (store_put_edge(s,&e)!=0){ ok("put_edge", 0); free_edge_fields(&e); store_close(s); return; }
free_edge_fields(&e);
if ((i%1000)==999) store_sync(s);
}
store_sync(s);
StorePoolStats st; store_pool_stats(s, &st);
printf(" pages=%llu pool: cap=%zu resident=%zu pinned=%zu dirty=%zu evictions=%llu\n",
(unsigned long long)store_page_count(s), st.cap, st.resident, st.pinned,
st.dirty, (unsigned long long)st.evictions);
ok("eviction actually fired (store exceeded the pool)", st.evictions > 0);
ok("pool stayed bounded (resident <= cap)", st.resident <= st.cap);
ok("no dirty frames after checkpoint", st.dirty == 0);
/* read back EVERY node bit-exact despite constant eviction/re-fault */
int bad = 0;
for (int i=0;i<NODE_COUNT;i++){
StoreNode want; gen_node(i,&want);
StoreNode got; int hit = store_get_node(s, want.id, &got);
if (hit!=1 || !cmp_node(&want,&got)) bad++;
if (hit==1) store_node_free(&got);
free_node_fields(&want);
}
ok("all 5000 nodes bit-exact under eviction", bad==0);
/* sample 4000 edges bit-exact */
int ebad = 0;
for (int i=0;i<EDGE_COUNT;i+=5){
StoreEdge want; gen_edge(i,&want);
StoreEdge got; int hit = store_get_edge(s, want.id, &got);
if (hit!=1 || !cmp_edge(&want,&got)) ebad++;
if (hit==1) store_edge_free(&got);
free_edge_fields(&want);
}
ok("sampled 4000 edges bit-exact under eviction", ebad==0);
ok("store_check crc clean under paging", store_check(s, STORE_CHECK_CRC)==0);
store_pool_stats(s, &st);
printf(" after reads: resident=%zu (<= cap=%zu) hits=%llu misses=%llu evictions=%llu\n",
st.resident, st.cap, (unsigned long long)st.hits,
(unsigned long long)st.misses, (unsigned long long)st.evictions);
ok("still bounded after full read-back", st.resident <= st.cap);
store_close(s);
unlink(path);
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 2 — EVICTION POLICY: a repeatedly-touched HOT set stays resident (0 extra
* faults) while a streaming COLD set is evicted; and a dirty-heavy write burst
* proves dirty pages are NEVER stolen before a checkpoint (no-steal).
* ════════════════════════════════════════════════════════════════════════════ */
static void test_eviction_policy(void){
printf("\n== 2) eviction policy: hot resident, cold evicted, no dirty stolen ==\n");
char path[600]; path_in(path, sizeof path, "evict.store");
unlink(path);
/* ---- part A: hot vs cold ---- */
EngramPagedStore* s = store_create(path);
if (!s){ ok("store_create", 0); return; }
const int N = 1500;
for (int i=0;i<N;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); free_node_fields(&n);
if ((i%400)==399) store_sync(s); }
store_sync(s);
store__set_pool_frames(s, 64);
const int HOT = 8;
/* warm the hot set */
for (int h=0;h<HOT;h++){ char id[32]; snprintf(id,sizeof id,"node-%d",h);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g); }
StorePoolStats a,b;
uint64_t hot_faults = 0, cold_faults = 0;
int cold = 200; /* streaming cold ids well outside hot set */
for (int r=0;r<150;r++){
for (int h=0;h<HOT;h++){
char id[32]; snprintf(id,sizeof id,"node-%d",h);
store_pool_stats(s,&a);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g);
store_pool_stats(s,&b);
hot_faults += (b.misses - a.misses);
}
for (int c=0;c<3;c++){
char id[32]; snprintf(id,sizeof id,"node-%d",cold++);
if (cold>=N) cold=200;
store_pool_stats(s,&a);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g);
store_pool_stats(s,&b);
cold_faults += (b.misses - a.misses);
}
}
printf(" hot re-get faults (post-warm)=%llu cold stream faults=%llu\n",
(unsigned long long)hot_faults, (unsigned long long)cold_faults);
ok("HOT pages stay resident (0 faults on re-access)", hot_faults == 0);
ok("COLD pages get evicted + re-faulted", cold_faults > 0);
store_pool_stats(s,&b);
double hr = (double)b.hits / (double)(b.hits + b.misses);
printf(" overall hit-rate = %.3f (hits=%llu misses=%llu)\n",
hr, (unsigned long long)b.hits, (unsigned long long)b.misses);
ok("hit-rate is sane (> 0.5)", hr > 0.5);
store_close(s);
unlink(path);
/* ---- part B: no-steal (dirty pages never evicted before checkpoint) ---- */
EngramPagedStore* s2 = store_create(path);
if (!s2){ ok("store_create(2)", 0); return; }
store__set_pool_frames(s2, 8); /* tiny budget */
for (int i=0;i<1200;i++){ StoreNode n; gen_node(i,&n); store_put_node(s2,&n); free_node_fields(&n); }
/* NO sync: every mutated page is dirty and, by no-steal, unevictable */
StorePoolStats d; store_pool_stats(s2,&d);
printf(" tiny cap=%zu, unsynced burst: resident=%zu dirty=%zu evictions=%llu\n",
d.cap, d.resident, d.dirty, (unsigned long long)d.evictions);
ok("dirty pages pinned in RAM beyond budget (no-steal)", d.dirty > d.cap && d.resident > d.cap);
/* a just-written node is served correctly from its dirty in-RAM page */
{ StoreNode want; gen_node(777,&want); StoreNode got; int hit=store_get_node(s2,want.id,&got);
ok("read served correctly from dirty (un-flushed) page", hit==1 && cmp_node(&want,&got));
if (hit==1) store_node_free(&got); free_node_fields(&want); }
store_sync(s2); /* checkpoint → dirty become clean/evictable */
store_pool_stats(s2,&d);
ok("checkpoint cleared all dirty frames", d.dirty == 0);
/* durability across reopen after the no-steal burst */
store_close(s2);
EngramPagedStore* s3 = store_open(path);
store__set_pool_frames(s3, 8);
int miss=0; for (int i=0;i<1200;i++){ StoreNode want; gen_node(i,&want);
StoreNode got; int hit=store_get_node(s3,want.id,&got);
if (hit!=1 || !cmp_node(&want,&got)) miss++;
if (hit==1) store_node_free(&got); free_node_fields(&want); }
ok("all 1200 survive reopen, bit-exact, tiny pool", miss==0);
store_close(s3);
unlink(path);
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 3 — PINNED RESIDENCY: superblocks + index roots never evicted under heavy
* thrash; an explicitly pinned page stays until unpinned; a pinned hot layer's
* pages stay resident and are released on unpin.
* ════════════════════════════════════════════════════════════════════════════ */
static void test_pinning(void){
printf("\n== 3) pinned residency: superblocks / index roots / page / layer ==\n");
char path[600]; path_in(path, sizeof path, "pin.store");
unlink(path);
EngramPagedStore* s = store_create(path);
if (!s){ ok("store_create", 0); return; }
const int N = 1500;
for (int i=0;i<N;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); free_node_fields(&n);
if ((i%400)==399) store_sync(s); }
store_sync(s);
store_close(s);
s = store_open(path); /* reopen: SBs + roots auto-pinned */
store__set_pool_frames(s, 24);
uint64_t P = store_page_count(s) / 2; /* an arbitrary interior page to pin */
store_pin_page(s, P);
/* thrash: stream a large cold working set to force heavy eviction */
for (int pass=0; pass<3; pass++)
for (int i=0;i<N;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g); }
ok("superblock page 0 never evicted", store_pool_resident(s,0)==1);
ok("superblock mirror page 1 never evicted", store_pool_resident(s,1)==1);
ok("explicitly pinned page stayed resident under thrash", store_pool_resident(s,P)==1);
StorePoolStats st; store_pool_stats(s,&st);
printf(" after thrash: resident=%zu pinned=%zu evictions=%llu\n",
st.resident, st.pinned, (unsigned long long)st.evictions);
ok("structural + explicit pins counted (>=4: 2 SB + 2 roots)", st.pinned >= 4);
/* unpin the page → it becomes evictable and is dropped under further thrash */
store_unpin_page(s, P);
for (int i=0;i<N;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g); }
ok("unpinned page becomes evictable (dropped)", store_pool_resident(s,P)==0);
/* hot-layer pin: layer 3 is used by ~1/5 of the nodes */
int npin = store_pin_layer(s, 3);
printf(" store_pin_layer(3) pinned %d page(s)\n", npin);
ok("pin_layer pinned a non-empty page set", npin > 0);
store_pool_stats(s,&st);
size_t pinned_with_layer = st.pinned;
for (int pass=0; pass<3; pass++)
for (int i=0;i<N;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g); }
store_pool_stats(s,&st);
ok("hot-layer pages stay resident under thrash", st.pinned >= pinned_with_layer);
ok("layer pin holds >= npin extra frames", st.pinned >= (size_t)npin + 4);
store_unpin_layer(s, 3);
store_pool_stats(s,&st);
size_t after_unpin_max = st.pinned;
for (int i=0;i<N;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g); }
store_pool_stats(s,&st);
printf(" pinned frames: with-layer=%zu after-unpin=%zu\n", pinned_with_layer, st.pinned);
ok("unpin_layer released the layer's pins", st.pinned < pinned_with_layer && after_unpin_max <= pinned_with_layer);
store_close(s);
unlink(path);
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 4 — PREFETCH: a sequential scan faults far fewer times with read-ahead on
* than off (each cold cache; identical store).
* ════════════════════════════════════════════════════════════════════════════ */
static void test_prefetch(void){
printf("\n== 4) prefetch: sequential scan faults fewer with read-ahead ==\n");
char path[600]; path_in(path, sizeof path, "prefetch.store");
unlink(path);
EngramPagedStore* s = store_create(path);
if (!s){ ok("store_create", 0); return; }
for (int i=0;i<2000;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); free_node_fields(&n);
if ((i%400)==399) store_sync(s); }
store_sync(s);
store_close(s);
/* prefetch OFF — cold cache */
EngramPagedStore* a = store_open(path);
store__set_pool_frames(a, 0); /* unlimited: isolate prefetch, no eviction */
store__set_prefetch(a, 0);
StorePoolStats o0, o1; store_pool_stats(a,&o0);
int na = store_scan_nodes(a, noop_node_cb, NULL); /* walk + fault every page */
(void)na;
store_pool_stats(a,&o1);
uint64_t faults_off = o1.misses - o0.misses;
store_close(a);
/* prefetch ON — cold cache (fresh open) */
EngramPagedStore* b = store_open(path);
store__set_pool_frames(b, 0);
store__set_prefetch(b, 16);
StorePoolStats p0, p1; store_pool_stats(b,&p0);
int nb = store_scan_nodes(b, noop_node_cb, NULL);
(void)nb;
store_pool_stats(b,&p1);
uint64_t faults_on = p1.misses - p0.misses;
uint64_t pref_reads = p1.prefetch_reads - p0.prefetch_reads;
store_close(b);
printf(" scan demand-faults: prefetch OFF=%llu ON=%llu (read-ahead brought in %llu pages)\n",
(unsigned long long)faults_off, (unsigned long long)faults_on,
(unsigned long long)pref_reads);
ok("prefetch reduced demand faults", faults_on < faults_off);
ok("read-ahead actually ran", pref_reads > 0);
unlink(path);
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 5 — CRASH SAFETY UNDER PAGING: WAL replay and checkpoint-crash recovery
* with a tiny pool (pages evict + re-fault during replay).
* ════════════════════════════════════════════════════════════════════════════ */
static void test_crash_under_paging(void){
printf("\n== 5) crash safety under a tiny pool (ENGRAM_POOL_FRAMES=16) ==\n");
setenv("ENGRAM_POOL_FRAMES", "16", 1); /* every engram_open() below is paged */
setenv("ENGRAM_WAL_SYNC", "always", 1);
/* ---- 5a: power-loss → WAL replay ---- */
char dir[600]; path_in(dir, sizeof dir, "crash_wal"); mkdir(dir, 0700);
EngramPagedStore* s = engram_open(dir);
if (!s){ ok("engram_open", 0); return; }
const int M = 400;
for (int i=0;i<M;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); free_node_fields(&n); }
store__crash(s); /* abandon RAM (dirty pages lost); WAL fsync'd */
s = engram_open(dir); /* replay WAL under 16-frame pool */
ok("reopened after crash (WAL replay, tiny pool)", s!=NULL);
int bad=0; for (int i=0;i<M;i++){ StoreNode want; gen_node(i,&want);
StoreNode got; int hit=store_get_node(s,want.id,&got);
if (hit!=1 || !cmp_node(&want,&got)) bad++;
if (hit==1) store_node_free(&got); free_node_fields(&want); }
ok("all 400 nodes recovered bit-exact via WAL replay under paging", bad==0);
ok("store_check crc clean post-recovery", store_check(s, STORE_CHECK_CRC)==0);
engram_close(s);
/* ---- 5b: checkpoint-crash at each phase ---- */
for (int phase=0; phase<=4; phase++){
char cdir[620]; snprintf(cdir, sizeof cdir, "%s/ck%d", g_dir, phase); mkdir(cdir,0700);
EngramPagedStore* c = engram_open(cdir);
for (int i=0;i<CK_NODES;i++){ StoreNode n; gen_node(i,&n); store_put_node(c,&n); free_node_fields(&n); }
store__checkpoint_crashat(c, phase); /* crash mid-checkpoint (frees c) */
EngramPagedStore* r = engram_open(cdir); /* heal + replay under tiny pool */
int miss=0; for (int i=0;i<CK_NODES;i++){ StoreNode want; gen_node(i,&want);
StoreNode got; int hit=store_get_node(r,want.id,&got);
if (hit!=1 || !cmp_node(&want,&got)) miss++;
if (hit==1) store_node_free(&got); free_node_fields(&want); }
char nm[64]; snprintf(nm,sizeof nm,"checkpoint-crash phase %d: all recovered (paged)", phase);
ok(nm, miss==0);
engram_close(r);
}
unsetenv("ENGRAM_POOL_FRAMES");
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 6 — DEFAULT POOL == PHASE 1: with the default (large) budget, no eviction
* ever fires; the whole store is resident, exactly the pre-M4 behaviour.
* ════════════════════════════════════════════════════════════════════════════ */
static void test_default_is_phase1(void){
printf("\n== 6) default (large) pool == Phase-1 resident (no eviction) ==\n");
char path[600]; path_in(path, sizeof path, "default.store");
unlink(path);
EngramPagedStore* s = store_create(path); /* default cap, no override */
if (!s){ ok("store_create", 0); return; }
for (int i=0;i<1500;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); free_node_fields(&n); }
store_sync(s);
for (int i=0;i<1500;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode g; if (store_get_node(s,id,&g)==1) store_node_free(&g); }
StorePoolStats st; store_pool_stats(s,&st);
printf(" cap=%zu resident=%zu evictions=%llu (pages=%llu)\n",
st.cap, st.resident, (unsigned long long)st.evictions,
(unsigned long long)store_page_count(s));
ok("default budget is large", st.cap >= (size_t)(1u<<20));
ok("no eviction ever fired at default budget", st.evictions == 0);
ok("whole store resident (every page cached)", st.resident == store_page_count(s));
store_close(s);
unlink(path);
}
int main(void){
mk_dir();
printf("engram M4 buffer-pool gate — dir=%s\n", g_dir);
test_small_pool_roundtrip();
test_eviction_policy();
test_pinning();
test_prefetch();
test_crash_under_paging();
test_default_is_phase1();
printf("\n================ %d passed, %d failed ================\n", g_pass, g_fail);
return g_fail ? 1 : 0;
}
+421
View File
@@ -0,0 +1,421 @@
/* test_compaction.c — M5 gate: ONLINE COMPACTION + background checkpointer.
*
* Pure C. Build: gcc -O2 test_compaction.c ../../lang/runtime/engram_store.c -o t
* Writes ONLY under a throwaway /tmp dir. Never touches ~/.neuron or live ports.
*
* Proves:
* 1) RECLAIM — tombstone/forget a large fraction of nodes + re-put many edges
* (dead versions) + orphan large-record overflow chains, then compact:
* page count AND file size drop, yet EVERY live record survives bit-exact and
* the id + adjacency indexes resolve correctly at the relocated positions.
* 2) CRASH-DURING-COMPACTION — kill at phases 0/1/2; recovery is always a
* consistent store (crc clean, every live record intact), never corrupt.
* 3) BACKGROUND CHECKPOINTER — a low ops / WAL-bytes threshold fires a checkpoint
* automatically on the write path; the WAL prefix is reclaimed; recovery works.
* 4) POOL COOPERATION — compaction under a tiny ENGRAM_POOL_FRAMES stays correct
* with no stale frame surviving for a relocated page.
*/
#include "../../lang/runtime/engram_store.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdint.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>
static int g_pass = 0, g_fail = 0;
static void ok(const char* name, int cond){
printf(" [%s] %s\n", cond ? "PASS" : "FAIL", name);
if (cond) g_pass++; else g_fail++;
}
static char g_dir[512];
static int g_dseq = 0;
static void mk_dir(void){
snprintf(g_dir, sizeof g_dir, "/tmp/engram-compact-test-%d-%d", (int)getpid(), g_dseq++);
mkdir(g_dir, 0700);
}
static void egm_path(char* out, size_t cap){ snprintf(out, cap, "%s/neuron.egm", g_dir); }
static void wal_path(char* out, size_t cap){ snprintf(out, cap, "%s/neuron.wal", g_dir); }
static long file_size(const char* p){ struct stat st; return stat(p,&st)==0 ? (long)st.st_size : -1; }
/* ── deterministic generators (bit-exact regeneration for oracles) ─────────── */
static uint64_t xs(uint64_t* s){ uint64_t x=*s; x^=x<<13; x^=x>>7; x^=x<<17; *s=x; return x; }
static uint64_t node_seed(int i){ return 0x9E3779B97F4A7C15ULL ^ ((uint64_t)(i+1)*0xD1B54A32D192ED03ULL); }
static uint64_t edge_seed(int i){ return 0xC2B2AE3D27D4EB4FULL ^ ((uint64_t)(i+1)*0x165667B19E3779F9ULL); }
static char* rnd_str(uint64_t* st, size_t len){
char* s = (char*)malloc(len + 1);
for (size_t i=0;i<len;i++) s[i] = (char)(33 + (xs(st) % 94));
s[len] = 0; return s;
}
#define N_NODES 1500
#define N_DEAD 1200 /* forget node-0 .. node-1199 (1200 dead / 300 live) */
#define N_EDGES 3000
#define EDGE_REPUT 2000 /* re-put edge-0 .. edge-1999 to version 3 */
#define EMB_DIM 96
static int node_is_live(int i){ return i >= N_DEAD; }
static int edge_live_version(int i){ return (i < EDGE_REPUT) ? 3 : 0; }
static void gen_node(int i, StoreNode* n){
memset(n, 0, sizeof *n);
uint64_t st = node_seed(i);
char id[32]; snprintf(id, sizeof id, "node-%d", i);
n->id = strdup(id);
/* every 7th record is large → its own overflow chain (orphaned when it dies) */
size_t clen = (i % 7 == 0) ? (size_t)(18000 + (xs(&st) % 4000)) : (size_t)(xs(&st) % 200);
n->content = rnd_str(&st, clen);
n->node_type = rnd_str(&st, 4 + (xs(&st) % 8));
n->label = (i % 2) ? rnd_str(&st, 3 + (xs(&st) % 10)) : NULL;
n->tier = rnd_str(&st, 4 + (xs(&st) % 6));
n->tags = rnd_str(&st, xs(&st) % 40);
n->metadata = (i % 3) ? rnd_str(&st, xs(&st) % 60) : NULL;
n->salience = (double)(xs(&st) % 1000000) / 997.0;
n->importance = (double)(xs(&st) % 1000000) / 131.0;
n->confidence = (double)(xs(&st) % 1000000) / 733.0;
n->temporal_decay_rate = (double)(xs(&st) % 1000000) / 101.0;
n->activation_count = (int64_t)(xs(&st) % 100000);
n->last_activated = (int64_t)xs(&st);
n->created_at = (int64_t)(1600000000000LL + i);
n->updated_at = (int64_t)xs(&st);
n->background_activation = (double)(xs(&st) % 1000000) / 17.0;
n->working_memory_weight = (double)(xs(&st) % 1000000) / 29.0;
n->suppression_count = (int32_t)(xs(&st) % 50);
n->layer_id = (uint32_t)(xs(&st) % 5);
for (int k=0;k<STORE_BLL_K;k++) n->access_ts[k] = (int64_t)xs(&st);
n->access_head = (int32_t)(xs(&st) % STORE_BLL_K);
n->access_filled = (int32_t)(xs(&st) % (STORE_BLL_K + 1));
n->wm_anchor = (double)(xs(&st) % 1000000) / 3.0;
n->emb = (float*)malloc(EMB_DIM * sizeof(float));
for (int k=0;k<EMB_DIM;k++){ uint32_t u=(uint32_t)xs(&st); memcpy(&n->emb[k], &u, 4); }
n->emb_dim = EMB_DIM;
}
/* version alters weight/hebb/last_fired so a re-put is a distinct payload. */
static void gen_edge(int i, int version, StoreEdge* e){
memset(e, 0, sizeof *e);
uint64_t st = edge_seed(i);
char id[32], from[32], to[32];
snprintf(id, sizeof id, "edge-%d", i);
/* connect live nodes so adjacency queries on live nodes are meaningful */
snprintf(from, sizeof from, "node-%d", N_DEAD + (int)(xs(&st) % (N_NODES - N_DEAD)));
snprintf(to, sizeof to, "node-%d", N_DEAD + (int)(xs(&st) % (N_NODES - N_DEAD)));
e->id = strdup(id); e->from_id = strdup(from); e->to_id = strdup(to);
e->relation = rnd_str(&st, 3 + (xs(&st) % 12));
e->metadata = (i % 4) ? rnd_str(&st, xs(&st) % 40) : NULL;
e->weight = (double)(xs(&st) % 1000000) / 7.0 + version * 100.0;
e->hebb = (double)(xs(&st) % 1000000) / 13.0 + version * 3.0;
e->confidence = (double)(xs(&st) % 1000000) / 5.0;
e->created_at = (int64_t)(1600000000000LL + i);
e->updated_at = (int64_t)xs(&st) + version;
e->last_fired = (int64_t)xs(&st) + version * 1000;
e->inhibitory = (int32_t)(xs(&st) % 2);
e->layer_id = (uint32_t)(xs(&st) % 5);
}
static int streq(const char* a, const char* b){
if (!a && !b) return 1; if (!a || !b) return 0; return strcmp(a,b)==0;
}
static int cmp_node(const StoreNode* a, const StoreNode* b){
if (!streq(a->id,b->id) || !streq(a->content,b->content) ||
!streq(a->node_type,b->node_type) || !streq(a->label,b->label) ||
!streq(a->tier,b->tier) || !streq(a->tags,b->tags) ||
!streq(a->metadata,b->metadata)) return 0;
if (a->salience!=b->salience || a->importance!=b->importance ||
a->confidence!=b->confidence || a->temporal_decay_rate!=b->temporal_decay_rate ||
a->activation_count!=b->activation_count || a->last_activated!=b->last_activated ||
a->created_at!=b->created_at || a->updated_at!=b->updated_at ||
a->background_activation!=b->background_activation ||
a->working_memory_weight!=b->working_memory_weight ||
a->suppression_count!=b->suppression_count || a->layer_id!=b->layer_id ||
a->access_head!=b->access_head || a->access_filled!=b->access_filled ||
a->wm_anchor!=b->wm_anchor || a->emb_dim!=b->emb_dim) return 0;
for (int k=0;k<STORE_BLL_K;k++) if (a->access_ts[k]!=b->access_ts[k]) return 0;
if ((a->emb==NULL) != (b->emb==NULL)) return 0;
if (a->emb && memcmp(a->emb, b->emb, (size_t)a->emb_dim*4)!=0) return 0;
return 1;
}
static int cmp_edge(const StoreEdge* a, const StoreEdge* b){
if (!streq(a->id,b->id) || !streq(a->from_id,b->from_id) || !streq(a->to_id,b->to_id) ||
!streq(a->relation,b->relation) || !streq(a->metadata,b->metadata)) return 0;
if (a->weight!=b->weight || a->hebb!=b->hebb || a->confidence!=b->confidence ||
a->created_at!=b->created_at || a->updated_at!=b->updated_at ||
a->last_fired!=b->last_fired || a->inhibitory!=b->inhibitory ||
a->layer_id!=b->layer_id) return 0;
return 1;
}
/* Populate a durable store with dead space: all nodes/edges, then forget the first
* N_DEAD nodes and re-put the first EDGE_REPUT edges three times. */
static void populate_with_dead_space(EngramPagedStore* s){
for (int i=0;i<N_NODES;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); store_node_free(&n); }
for (int i=0;i<N_EDGES;i++){ StoreEdge e; gen_edge(i,0,&e); store_put_edge(s,&e); store_edge_free(&e); }
/* re-put (in-place field mutation) → prior versions become dead records */
for (int v=1; v<=3; v++)
for (int i=0;i<EDGE_REPUT;i++){ StoreEdge e; gen_edge(i,v,&e); store_put_edge(s,&e); store_edge_free(&e); }
/* forget the cold nodes (tombstone; their large overflow chains orphan) */
for (int i=0;i<N_DEAD;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i); store_forget(s,id); }
}
/* Assert every live node/edge is present + bit-exact via point reads. */
static int verify_live_set(EngramPagedStore* s){
int bad = 0;
for (int i=0;i<N_NODES;i++){
char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode got; int hit = store_get_node(s, id, &got);
if (node_is_live(i)){
StoreNode want; gen_node(i,&want);
if (hit!=1 || !cmp_node(&want,&got)) bad++;
if (hit==1) store_node_free(&got);
store_node_free(&want);
} else {
if (hit!=0) bad++; /* forgotten → must be absent */
if (hit==1) store_node_free(&got);
}
}
for (int i=0;i<N_EDGES;i++){
char id[32]; snprintf(id,sizeof id,"edge-%d",i);
StoreEdge got; int hit = store_get_edge(s, id, &got);
StoreEdge want; gen_edge(i, edge_live_version(i), &want);
if (hit!=1 || !cmp_edge(&want,&got)) bad++;
if (hit==1) store_edge_free(&got);
store_edge_free(&want);
}
return bad;
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 1 — RECLAIM: dead space is reclaimed; live records + indexes survive.
* ════════════════════════════════════════════════════════════════════════════ */
static void test_reclaim(void){
printf("\n== 1) reclaim: forget %d nodes + re-put %d edges x3, then compact ==\n",
N_DEAD, EDGE_REPUT);
mk_dir();
char egm[600]; egm_path(egm, sizeof egm);
EngramPagedStore* s = engram_open(g_dir);
ok("engram_open", s != NULL);
if (!s) return;
populate_with_dead_space(s);
engram_checkpoint(s); /* flush so file size reflects state */
uint64_t pc_before = store_page_count(s);
uint64_t free_before = store_free_page_count(s);
long sz_before = file_size(egm);
printf(" BEFORE: page_count=%llu free_pages=%llu file=%ld bytes (live records intact?)\n",
(unsigned long long)pc_before, (unsigned long long)free_before, sz_before);
ok("pre-compaction live set intact", verify_live_set(s)==0);
/* capture adjacency for a sample of live from-ids to compare post-compaction */
#define NSAMP 12
char samp[NSAMP][32]; size_t pre_cnt[NSAMP];
for (int k=0;k<NSAMP;k++){
snprintf(samp[k], sizeof samp[k], "node-%d", N_DEAD + k*20);
StoreEdge* arr=NULL; size_t cnt=0;
store_get_edges_from(s, samp[k], &arr, &cnt);
pre_cnt[k]=cnt; store_edges_free(arr,cnt);
}
int rc = store_compact(s);
ok("store_compact returns 0", rc==0);
uint64_t pc_after = store_page_count(s);
uint64_t free_after = store_free_page_count(s);
long sz_after = file_size(egm);
printf(" AFTER : page_count=%llu free_pages=%llu file=%ld bytes\n",
(unsigned long long)pc_after, (unsigned long long)free_after, sz_after);
printf(" RECLAIMED: %llu pages, %ld bytes (%.1f%% of file)\n",
(unsigned long long)(pc_before - pc_after), sz_before - sz_after,
sz_before ? 100.0*(sz_before-sz_after)/sz_before : 0.0);
ok("page count dropped (dead pages reclaimed)", pc_after < pc_before);
ok("file size dropped (store physically shrank)", sz_after < sz_before);
ok("store_check crc clean after compaction", store_check(s, STORE_CHECK_CRC)==0);
ok("every LIVE record present + bit-exact at new locations", verify_live_set(s)==0);
/* adjacency index correct at relocated positions */
int adj_bad = 0;
for (int k=0;k<NSAMP;k++){
StoreEdge* arr=NULL; size_t cnt=0;
store_get_edges_from(s, samp[k], &arr, &cnt);
if (cnt != pre_cnt[k]) adj_bad++;
for (size_t j=0;j<cnt;j++){
if (!streq(arr[j].from_id, samp[k])) { adj_bad++; break; }
/* the returned edge must be the canonical latest live edge, bit-exact */
int idx = atoi(arr[j].id + 5);
StoreEdge want; gen_edge(idx, edge_live_version(idx), &want);
if (!cmp_edge(&want,&arr[j])) adj_bad++;
store_edge_free(&want);
}
store_edges_free(arr,cnt);
}
ok("adjacency (get_edges_from) correct + bit-exact post-compaction", adj_bad==0);
/* second compaction is a near no-op (no new dead space) and stays correct */
uint64_t pc2_before = store_page_count(s);
ok("compact again returns 0", store_compact(s)==0);
ok("idempotent-ish: no growth on re-compact", store_page_count(s) <= pc2_before);
ok("live set still intact after 2nd compaction", verify_live_set(s)==0);
engram_close(s);
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 2 — CRASH DURING COMPACTION: kill at phases 0/1/2 → consistent recovery.
* Live set is identical whether we recover pre- or post-compaction, so the same
* oracle must hold, and crc must always be clean (never corrupt).
* ════════════════════════════════════════════════════════════════════════════ */
static void test_crash_during_compaction(void){
printf("\n== 2) crash during compaction at phases 0,1,2 → consistent store ==\n");
for (int phase=0; phase<=2; phase++){
mk_dir();
char egm[600]; egm_path(egm, sizeof egm);
EngramPagedStore* s = engram_open(g_dir);
if (!s){ ok("engram_open", 0); continue; }
populate_with_dead_space(s);
engram_close(s); /* durable baseline on disk */
uint64_t pc_pre = 0;
{ EngramPagedStore* p = engram_open(g_dir); pc_pre = store_page_count(p); engram_close(p); }
EngramPagedStore* c = engram_open(g_dir);
store__compact_crashat(c, phase); /* crashes mid-compaction (frees c) */
EngramPagedStore* r = engram_open(g_dir); /* recover */
char nm[80];
snprintf(nm, sizeof nm, "phase %d: recovers, crc clean", phase);
ok(nm, r && store_check(r, STORE_CHECK_CRC)==0);
snprintf(nm, sizeof nm, "phase %d: every live record intact (not corrupt)", phase);
ok(nm, r && verify_live_set(r)==0);
if (r){
uint64_t pc_now = store_page_count(r);
if (phase < 2){
snprintf(nm, sizeof nm, "phase %d: recovered PRE-compaction image", phase);
ok(nm, pc_now == pc_pre);
} else {
snprintf(nm, sizeof nm, "phase %d: recovered POST-compaction (shrunk)", phase);
ok(nm, pc_now < pc_pre);
}
/* store stays writable + durable after recovery */
StoreNode n; gen_node(N_NODES+phase, &n); free(n.id);
n.id = strdup("post-recovery-node");
store_put_node(r, &n); store_node_free(&n);
StoreNode g; int hit = store_get_node(r, "post-recovery-node", &g);
snprintf(nm, sizeof nm, "phase %d: store writable after recovery", phase);
ok(nm, hit==1);
if (hit==1) store_node_free(&g);
engram_close(r);
}
}
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 3 — BACKGROUND CHECKPOINTER: a low threshold fires checkpoints on the
* write path, reclaiming the WAL prefix automatically; recovery still correct.
* ════════════════════════════════════════════════════════════════════════════ */
static void test_background_checkpointer(void){
printf("\n== 3) background checkpointer: auto-checkpoint on threshold ==\n");
/* (a) ops trigger */
{
mk_dir();
char wal[600]; wal_path(wal, sizeof wal);
EngramPagedStore* s = engram_open(g_dir);
if (!s){ ok("engram_open", 0); return; }
store_set_checkpoint_policy(s, /*ops*/50, /*dirty*/0, /*wal_bytes*/0, /*ms*/0);
uint64_t ckpt0 = engram_last_checkpoint_lsn(s);
for (int i=0;i<600;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); store_node_free(&n); }
uint64_t ckpt1 = engram_last_checkpoint_lsn(s);
long wsz = file_size(wal);
printf(" ops-trigger: ckpt_lsn %llu -> %llu, WAL=%ld bytes after 600 puts\n",
(unsigned long long)ckpt0, (unsigned long long)ckpt1, wsz);
ok("ops trigger fired an automatic checkpoint", ckpt1 > ckpt0);
ok("WAL prefix reclaimed (WAL stays small)", wsz >= 0 && wsz < 200000);
/* crash (abandon RAM) then recover — everything durable via WAL+checkpoint */
store__crash(s);
EngramPagedStore* r = engram_open(g_dir);
int bad=0;
for (int i=0;i<600;i++){ char id[32]; snprintf(id,sizeof id,"node-%d",i);
StoreNode w; gen_node(i,&w); StoreNode g; int hit=store_get_node(r,id,&g);
if (hit!=1 || !cmp_node(&w,&g)) bad++; if(hit==1) store_node_free(&g); store_node_free(&w); }
ok("recovery correct after auto-checkpoints (ops)", r && bad==0);
ok("crc clean after recovery (ops)", r && store_check(r,STORE_CHECK_CRC)==0);
if (r) engram_close(r);
}
/* (b) WAL-bytes trigger */
{
mk_dir();
char wal[600]; wal_path(wal, sizeof wal);
EngramPagedStore* s = engram_open(g_dir);
if (!s){ ok("engram_open", 0); return; }
store_set_checkpoint_policy(s, /*ops*/0, /*dirty*/0, /*wal_bytes*/64*1024, /*ms*/0);
uint64_t ckpt0 = engram_last_checkpoint_lsn(s);
for (int i=0;i<600;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); store_node_free(&n); }
uint64_t ckpt1 = engram_last_checkpoint_lsn(s);
long wsz = file_size(wal);
printf(" wal-bytes-trigger: ckpt_lsn %llu -> %llu, WAL=%ld bytes\n",
(unsigned long long)ckpt0, (unsigned long long)ckpt1, wsz);
ok("wal-bytes trigger fired an automatic checkpoint", ckpt1 > ckpt0);
ok("WAL kept bounded by byte threshold", wsz >= 0 && wsz < 2*1024*1024);
engram_close(s);
}
/* (c) dirty-frames trigger (under a bounded pool) */
{
mk_dir();
EngramPagedStore* s = engram_open(g_dir);
if (!s){ ok("engram_open", 0); return; }
store_set_checkpoint_policy(s, /*ops*/0, /*dirty*/16, /*wal_bytes*/0, /*ms*/0);
uint64_t ckpt0 = engram_last_checkpoint_lsn(s);
for (int i=0;i<400;i++){ StoreNode n; gen_node(i,&n); store_put_node(s,&n); store_node_free(&n); }
uint64_t ckpt1 = engram_last_checkpoint_lsn(s);
ok("dirty-frames trigger fired an automatic checkpoint", ckpt1 > ckpt0);
engram_close(s);
}
}
/* ════════════════════════════════════════════════════════════════════════════
* TEST 4 — POOL COOPERATION: compact under a tiny frame budget (constant eviction
* + re-fault); correctness holds and no stale frame survives a relocated page.
* ════════════════════════════════════════════════════════════════════════════ */
static void test_pool_cooperation(void){
printf("\n== 4) compaction under a small buffer pool (forced eviction) ==\n");
setenv("ENGRAM_POOL_FRAMES", "24", 1); /* pool << store, and the temp build too */
mk_dir();
char egm[600]; egm_path(egm, sizeof egm);
EngramPagedStore* s = engram_open(g_dir);
ok("engram_open (24-frame pool)", s != NULL);
if (!s){ unsetenv("ENGRAM_POOL_FRAMES"); return; }
store__set_pool_frames(s, 24);
populate_with_dead_space(s);
engram_checkpoint(s);
uint64_t pc_before = store_page_count(s);
int rc = store_compact(s);
ok("store_compact under tiny pool returns 0", rc==0);
StorePoolStats st; store_pool_stats(s, &st);
printf(" post-compaction pool: cap=%zu resident=%zu pinned=%zu dirty=%zu\n",
st.cap, st.resident, st.pinned, st.dirty);
ok("pool respected budget after compaction (resident<=cap)", st.resident <= st.cap);
ok("page count dropped under small pool", store_page_count(s) < pc_before);
ok("crc clean under small pool", store_check(s, STORE_CHECK_CRC)==0);
/* If any relocated page had a stale frame, a read would return wrong bytes. */
ok("every live record bit-exact under small pool (no stale frames)", verify_live_set(s)==0);
engram_close(s);
unsetenv("ENGRAM_POOL_FRAMES");
}
int main(void){
printf("=== M5 COMPACTION + BACKGROUND CHECKPOINTER GATE ===\n");
test_reclaim();
test_crash_during_compaction();
test_background_checkpointer();
test_pool_cooperation();
printf("\n=== RESULT: %d passed, %d failed ===\n", g_pass, g_fail);
/* cleanup */
return g_fail ? 1 : 0;
}
+159
View File
@@ -0,0 +1,159 @@
/* test_geometry.c — build + RUN gate for the M9 FOUNDATION geometry descriptor
* (engram_geometry.{c,h}). Self-contained: synthesizes a store with two KNOWN
* embedding clusters + intra-cluster hebb edges, then verifies the descriptor
* recovers the shape — centroid near the seeded cluster, skeleton = the strong
* intra-cluster edges, membership gradient, radius, positive co-registration.
*
* Pure C11; links engram_geometry.c + engram_store.c + engram_vindex.c; -lm.
* ASan/UBSan clean. Needs no live data.
*/
#include "engram_geometry.h"
#include "engram_store.h"
#include "engram_vindex.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <math.h>
#include <stdint.h>
#include <unistd.h>
#define DIM 64
static int g_fail=0;
#define CHECK(c,m) do{ if(!(c)){printf(" FAIL: %s\n",m); g_fail=1;} else printf(" ok: %s\n",m);}while(0)
static uint64_t rs=0x1234abcdULL;
static uint64_t xr(void){ uint64_t z=(rs+=0x9E3779B97F4A7C15ULL);
z=(z^(z>>30))*0xBF58476D1CE4E5B9ULL; z=(z^(z>>27))*0x94D049BB133111EBULL; return z^(z>>31); }
static float jitter(void){ return (float)(((double)(xr()>>11)*(1.0/9007199254740992.0))-0.5)*0.15f; }
/* two clusters: A centered on axis 0, B centered on axis 1. NA+NB nodes. */
#define NA 40
#define NB 40
int main(void){
printf("=== engram_geometry (M9 foundation) test suite ===\n");
char path[256]; snprintf(path,sizeof path,"/tmp/geo_test_store_%d.egm",(int)getpid());
unlink(path);
EngramPagedStore* st=store_create(path);
if(!st){ printf("FAIL: store_create\n"); return 1; }
char aids[NA][16], bids[NB][16];
/* cluster A: near +e0 ; cluster B: near +e1 */
for(int i=0;i<NA;i++){
StoreNode n; memset(&n,0,sizeof n);
snprintf(aids[i],16,"A%d",i); n.id=aids[i]; n.node_type="Concept"; n.tier="Semantic";
n.content="cluster-A"; n.salience=0.5+0.01*i;
float v[DIM]; for(int d=0;d<DIM;d++) v[d]=jitter(); v[0]=1.0f+jitter();
n.emb=v; n.emb_dim=DIM; store_put_node(st,&n);
}
for(int i=0;i<NB;i++){
StoreNode n; memset(&n,0,sizeof n);
snprintf(bids[i],16,"B%d",i); n.id=bids[i]; n.node_type="Concept"; n.tier="Semantic";
n.content="cluster-B"; n.salience=0.3;
float v[DIM]; for(int d=0;d<DIM;d++) v[d]=jitter(); v[1]=1.0f+jitter();
n.emb=v; n.emb_dim=DIM; store_put_node(st,&n);
}
/* strong intra-A hebb edges (a chain + hub), weaker cross edges A0<->B0 */
int ei=0;
for(int i=1;i<NA;i++){
StoreEdge e; memset(&e,0,sizeof e); char id[24]; snprintf(id,24,"eA%d",ei++);
e.id=id; e.from_id=aids[0]; e.to_id=aids[i]; e.relation="assoc"; e.weight=0.9; e.hebb=0.4;
store_put_edge(st,&e);
}
for(int i=1;i<NB;i++){
StoreEdge e; memset(&e,0,sizeof e); char id[24]; snprintf(id,24,"eB%d",ei++);
e.id=id; e.from_id=bids[0]; e.to_id=bids[i]; e.relation="assoc"; e.weight=0.9; e.hebb=0.4;
store_put_edge(st,&e);
}
{ StoreEdge e; memset(&e,0,sizeof e); e.id=(char*)"eX"; e.from_id=aids[0]; e.to_id=bids[0];
e.relation="assoc"; e.weight=0.5; e.hebb=0.0; store_put_edge(st,&e); }
store_close(st);
VIndex* ix=vindex_create(DIM,0,0);
char** ids=NULL; int nids=0;
int ins=vindex_build_from_store(ix, path, &ids, &nids);
CHECK(ins==NA+NB, "vindex built over all embedded nodes");
GeoParams P; engram_geo_default_params(&P); P.ann_k=20; P.max_members=0;
/* seed inside cluster A -> expect an A-dominated neighborhood */
st=store_open(path);
/* global-mean cache over the embedded set: the centering offset */
GeoMeanCache* mc=engram_geo_mean_build(st);
const float* gm=engram_geo_mean_vec(mc);
CHECK(mc!=NULL && engram_geo_mean_dim(mc)==DIM, "global-mean cache built over embedded set");
CHECK(engram_geo_mean_count(mc)==(uint64_t)(NA+NB), "global mean averaged all embedded nodes");
const char* seeds[1]={aids[0]};
/* CENTERED descriptor: pass the global mean so geometry runs in isotropic space */
GeoDescriptor* g=engram_geometry_descriptor(st, ix, ids, nids, seeds, 1, &P, gm);
CHECK(g!=NULL, "descriptor computed");
if(g){
printf(" members=%d embedded=%d edges=%d k_core=%d radius=%.4f co_reg=%.3f n_axes=%d\n",
g->n_members,g->n_embedded,g->n_edges,g->k_core,g->radius,g->co_registration,g->n_axes);
/* geometry ran in CENTERED space: g->centroid is the centered centroid,
* g->global_mean the applied offset. Reconstruct the raw prototype
* (centroid + global_mean) and check it sits on cluster-A's axis. */
CHECK(g->global_mean!=NULL, "descriptor recorded the centering offset (centered mode)");
int argmax=0; float best=-1.f;
for(int d=0;d<g->dim;d++){ float raw=g->centroid[d]+(g->global_mean?g->global_mean[d]:0.f);
if(fabsf(raw)>best){ best=fabsf(raw); argmax=d; } }
printf(" raw-prototype dominant axis = %d (expect 0); centered c[0]=%.3f c[1]=%.3f\n",
argmax, g->centroid[0], g->centroid[1]);
CHECK(argmax==0, "raw prototype sits on cluster-A's axis (near members)");
/* centering pushes A off cluster-B's axis: centered c[0] > c[1] */
CHECK(g->centroid[0] > g->centroid[1], "centered centroid leans off B's axis (isotropy)");
/* hub should be A0 (the intra-A hub with NA-1 strong edges) */
CHECK(g->hub_id && strcmp(g->hub_id,"A0")==0, "hub = the relational center A0");
/* membership: seed A0 == 1.0; A-members strong, B-members (if any) weaker */
double seedw=-1, minA=2, maxB=-1; int na=0,nb=0;
for(int i=0;i<g->n_members;i++){
const char* id=g->members[i].id; double w=g->members[i].membership;
if(strcmp(id,"A0")==0) seedw=w;
if(id[0]=='A'){ na++; if(w<minA)minA=w; }
if(id[0]=='B'){ nb++; if(w>maxB)maxB=w; }
}
printf(" A-members=%d B-members=%d seedw=%.3f\n", na,nb,seedw);
CHECK(fabs(seedw-1.0)<1e-9, "seed membership == 1.0");
CHECK(na>=NA-1, "neighborhood recovers cluster A");
/* skeleton = the strong intra-A edges: every edge eff_weight>=threshold,
* and edges connect A-nodes (co-registration should be positive: wired
* pairs are semantically near). */
int allstrong=1, allA=1;
for(int e=0;e<g->n_edges;e++){
if(g->edges[e].eff_weight < P.edge_min_weight) allstrong=0;
const char* a=g->members[g->edges[e].a].id, *b=g->members[g->edges[e].b].id;
if(!(a[0]=='A'&&b[0]=='A')) { /* the lone eX cross edge is allowed */
if(!((strcmp(a,"A0")==0&&strcmp(b,"B0")==0)||(strcmp(a,"B0")==0&&strcmp(b,"A0")==0))) allA=0; }
}
CHECK(allstrong, "skeleton holds only above-threshold (strong) edges");
CHECK(allA, "skeleton backbone is the intra-cluster wiring");
CHECK(g->co_registration>0.0, "co-registration positive (wired pairs are semantically near)");
/* principal axes: extents strictly non-increasing */
int mono=1; for(int i=1;i<g->n_axes;i++) if(g->axes[i].extent>g->axes[i-1].extent+1e-9) mono=0;
CHECK(g->n_axes>0 && mono, "principal axes sorted by descending extent");
CHECK(g->radius>0, "radius positive");
}
engram_geo_free(g);
/* edge cases: NULL store, no seeds, relational-only (NULL vindex) */
CHECK(engram_geometry_descriptor(NULL,ix,ids,nids,seeds,1,&P,gm)==NULL, "NULL store -> NULL");
CHECK(engram_geometry_descriptor(st,ix,ids,nids,seeds,0,&P,gm)==NULL, "zero seeds -> NULL");
GeoDescriptor* g2=engram_geometry_descriptor(st, NULL, NULL, 0, seeds, 1, &P, gm);
CHECK(g2!=NULL && g2->n_members>=NA-1, "relational-only path (no vindex) works");
engram_geo_free(g2);
/* raw (uncentered) mode still supported: global_mean=NULL -> no offset recorded */
GeoDescriptor* g3=engram_geometry_descriptor(st, ix, ids, nids, seeds, 1, &P, NULL);
CHECK(g3!=NULL && g3->global_mean==NULL, "raw mode (global_mean=NULL) leaves offset unset");
engram_geo_free(g3);
engram_geo_mean_free(mc);
for(int i=0;i<nids;i++) free(ids[i]); free(ids);
vindex_free(ix); store_close(st); unlink(path);
printf("\n=== %s ===\n", g_fail?"FAILURES PRESENT":"ALL TESTS PASSED");
return g_fail;
}
+79
View File
@@ -0,0 +1,79 @@
/* test_interoception_p0_emb.c — M-INTEROCEPTION Priority 0.
*
* Verifies the new READ-ONLY builtin engram_scan_nodes_emb_json(limit,offset):
* - every emitted node carries emb_dim and an emb JSON array of that length,
* - nodes without an embedding emit emb_dim:0 / emb:[],
* - pagination (limit/offset) is honoured,
* - the count matches engram_node_count,
* - the EXISTING engram_scan_nodes_json path is byte-unchanged (no emb field),
* i.e. the addition is purely additive / behavior-neutral.
*
* Pure-C harness (no elc). We craft a snapshot with real emb vectors, load it
* (engram_load parses "emb" comma-lists into node->emb via eg_parse_emb), then
* dump via both scan paths. Assertions live in run_interoception_p0.sh.
*/
#include "el_runtime.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
static el_val_t S(const char* s){ return EL_STR(s); }
/* 16-d embedding as a comma list (>=8 required by eg_parse_emb). */
static void emb_list(char* out, size_t cap, int dim, double base){
size_t o = 0;
for (int i = 0; i < dim; i++){
o += snprintf(out+o, cap-o, "%s%.4f", i?",":"", base + 0.01*i);
}
}
int main(int argc, char** argv){
if (argc < 2){ fprintf(stderr, "usage: %s <dir>\n", argv[0]); return 2; }
const char* dir = argv[1];
char snap[1024]; snprintf(snap, sizeof snap, "%s/seed.json", dir);
char e1[512], e2[512];
emb_list(e1, sizeof e1, 16, 0.10);
emb_list(e2, sizeof e2, 16, 0.50);
/* Two embedded nodes (distinct salience → deterministic sort order) and one
* un-embedded node. */
FILE* f = fopen(snap, "w");
if (!f){ perror("fopen"); return 2; }
fprintf(f,
"{\"nodes\":["
"{\"id\":\"n-high\",\"content\":\"high salience embedded\",\"node_type\":\"Concept\","
"\"label\":\"emb-high\",\"tier\":\"Semantic\",\"salience\":0.9,\"importance\":0.8,"
"\"confidence\":1.0,\"created_at\":1000,\"emb\":\"%s\"},"
"{\"id\":\"n-mid\",\"content\":\"mid salience embedded\",\"node_type\":\"Concept\","
"\"label\":\"emb-mid\",\"tier\":\"Semantic\",\"salience\":0.5,\"importance\":0.5,"
"\"confidence\":1.0,\"created_at\":2000,\"emb\":\"%s\"},"
"{\"id\":\"n-low\",\"content\":\"low salience no embedding\",\"node_type\":\"Fact\","
"\"label\":\"noemb-low\",\"tier\":\"Semantic\",\"salience\":0.1,\"importance\":0.2,"
"\"confidence\":1.0,\"created_at\":3000}"
"],\"edges\":[]}", e1, e2);
fclose(f);
if (!engram_load(S(snap))){ fprintf(stderr, "load failed\n"); return 2; }
long long nc = (long long)(int64_t)engram_node_count();
printf("node_count=%lld\n", nc);
/* full page */
el_val_t all = engram_scan_nodes_emb_json((el_val_t)256, (el_val_t)0);
char p[1024];
snprintf(p, sizeof p, "%s/emb_all.json", dir);
f = fopen(p, "w"); fputs(EL_CSTR(all), f); fclose(f);
/* pagination: one node at offset 0 and one at offset 1 */
el_val_t pg0 = engram_scan_nodes_emb_json((el_val_t)1, (el_val_t)0);
el_val_t pg1 = engram_scan_nodes_emb_json((el_val_t)1, (el_val_t)1);
snprintf(p, sizeof p, "%s/emb_pg0.json", dir); f = fopen(p, "w"); fputs(EL_CSTR(pg0), f); fclose(f);
snprintf(p, sizeof p, "%s/emb_pg1.json", dir); f = fopen(p, "w"); fputs(EL_CSTR(pg1), f); fclose(f);
/* existing path — must be unchanged / carry NO emb */
el_val_t plain = engram_scan_nodes_json((el_val_t)256, (el_val_t)0);
snprintf(p, sizeof p, "%s/plain.json", dir); f = fopen(p, "w"); fputs(EL_CSTR(plain), f); fclose(f);
printf("wrote dumps to %s\n", dir);
return 0;
}
+124
View File
@@ -0,0 +1,124 @@
/* test_interoception_p1_consol.c — M-INTEROCEPTION Priority 1.
* Two-threshold consolidation (ENGRAM_CONSOLIDATION, default OFF).
*
* Modes:
* accrual — flag OFF (pure trunk). Drive N co-activations of a WIRED pair and
* print act-stats at sampled N so the run script can plot the
* hebb accrual curve (headline measurement). No consolidation code
* runs; this measures the EXISTING EWMA accrual.
* connect — flag ON. Seed, activate to populate WM, then create a STRONG ISE
* (connects to wm_top) and a WEAK ISE (below the bar → nothing).
* Exports the graph so edges from each ISE can be counted.
* perm — flag ON. Load two OLD InternalStateEvent nodes; promote one to
* permanence; prune telemetry; export so the durable one is shown
* to survive while the ephemeral one is swept.
* offcheck — flag OFF. Prove creating an ISE forms NO edges and
* engram_consolidate_permanence is a no-op (byte-identical OFF path).
*/
#include "el_runtime.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
static el_val_t S(const char* s){ return EL_STR(s); }
static el_val_t F(double d){ return el_from_float(d); }
static void build_seed(void){
el_val_t a = engram_node_full(S("hebbian potentiation strengthens co-active memory links"),
S("Concept"), S("hebb-a"), F(0.9), F(0.85), F(1.0), S("Semantic"),
S("hebbian,memory,activation"));
el_val_t b = engram_node_full(S("co-active memory links accrue hebbian associative weight"),
S("Concept"), S("hebb-b"), F(0.9), F(0.85), F(1.0), S("Semantic"),
S("hebbian,memory,weight"));
el_val_t c = engram_node_full(S("unrelated culinary recipe for sourdough bread"),
S("Fact"), S("distractor-1"), F(0.4), F(0.4), F(1.0), S("Semantic"), S("food"));
el_val_t d = engram_node_full(S("the weather forecast predicts rain tomorrow afternoon"),
S("Fact"), S("distractor-2"), F(0.4), F(0.4), F(1.0), S("Semantic"), S("weather"));
engram_connect(a, b, F(0.8), S("associate"));
engram_connect(a, c, F(0.3), S("associate"));
engram_connect(b, d, F(0.3), S("associate"));
}
static const char* QUERY =
"hebbian potentiation co-active memory links associative weight";
int main(int argc, char** argv){
if (argc < 3){ fprintf(stderr,"usage: %s <accrual|connect|perm|offcheck> <dir>\n",argv[0]); return 2; }
const char* mode = argv[1];
const char* dir = argv[2];
char p[1024];
if (!strcmp(mode,"accrual")){
build_seed();
int samples[] = {1,10,50,100,250,500,1000,1625,2000,2500,3000};
int ns = (int)(sizeof samples/sizeof samples[0]);
int NMAX = samples[ns-1];
int si = 0;
for (int n=1; n<=NMAX; n++){
engram_activate_json(S(QUERY), (el_val_t)3);
if (si<ns && n==samples[si]){
printf("SAMPLE %d %s\n", n, EL_CSTR(engram_act_stats_json()));
si++;
}
}
return 0;
}
if (!strcmp(mode,"connect")){
build_seed();
engram_activate_json(S(QUERY), (el_val_t)3);
long long e_before = (long long)(int64_t)engram_edge_count();
/* STRONG ISE — should connect to wm_top */
el_val_t ise_strong = engram_node_full(S("strong internal state: focused on hebbian consolidation"),
S("InternalStateEvent"), S("ise-strong"), F(0.9), F(0.8), F(1.0), S("Working"), S("ise"));
long long e_after_strong = (long long)(int64_t)engram_edge_count();
/* WEAK ISE — below the connection bar (salience 0.3 < default 0.6) */
el_val_t ise_weak = engram_node_full(S("weak internal state: idle drift"),
S("InternalStateEvent"), S("ise-weak"), F(0.3), F(0.3), F(1.0), S("Working"), S("ise"));
long long e_after_weak = (long long)(int64_t)engram_edge_count();
printf("ISE_STRONG_ID %s\n", EL_CSTR(ise_strong));
printf("ISE_WEAK_ID %s\n", EL_CSTR(ise_weak));
printf("EDGES before=%lld after_strong=%lld after_weak=%lld\n",
e_before, e_after_strong, e_after_weak);
snprintf(p,sizeof p,"%s/connect.json",dir);
el_val_t g = engram_save(S(p)); (void)g;
return 0;
}
if (!strcmp(mode,"perm")){
/* Two OLD ISE nodes (created_at far in the past → prunable at 48h). */
snprintf(p,sizeof p,"%s/seed.json",dir);
FILE* f=fopen(p,"w");
fprintf(f,"{\"nodes\":["
"{\"id\":\"ise-durable\",\"content\":\"promoted internal state\",\"node_type\":\"InternalStateEvent\","
"\"label\":\"ise-durable\",\"salience\":0.5,\"confidence\":1.0,\"created_at\":1000},"
"{\"id\":\"ise-ephemeral\",\"content\":\"transient internal state\",\"node_type\":\"InternalStateEvent\","
"\"label\":\"ise-ephemeral\",\"salience\":0.5,\"confidence\":1.0,\"created_at\":1000}"
"],\"edges\":[]}");
fclose(f);
if(!engram_load(S(p))){ fprintf(stderr,"load failed\n"); return 2; }
long long n_before = (long long)(int64_t)engram_node_count();
el_val_t promoted = engram_consolidate_permanence(S("ise-durable"));
long long removed = (long long)(int64_t)engram_prune_telemetry((el_val_t)0); /* default 48h */
long long n_after = (long long)(int64_t)engram_node_count();
printf("PROMOTED %lld\n", (long long)(int64_t)promoted);
printf("NODES before=%lld after=%lld removed=%lld\n", n_before, n_after, removed);
printf("DURABLE_NODE %s\n", EL_CSTR(engram_get_node_json(S("ise-durable"))));
printf("EPHEMERAL_NODE %s\n", EL_CSTR(engram_get_node_json(S("ise-ephemeral"))));
return 0;
}
if (!strcmp(mode,"offcheck")){
build_seed();
engram_activate_json(S(QUERY), (el_val_t)3);
long long e_before = (long long)(int64_t)engram_edge_count();
engram_node_full(S("strong internal state with flag OFF"),
S("InternalStateEvent"), S("ise-off"), F(0.9), F(0.8), F(1.0), S("Working"), S("ise"));
long long e_after = (long long)(int64_t)engram_edge_count();
el_val_t perm = engram_consolidate_permanence(S("ise-off"));
printf("OFF edges before=%lld after=%lld perm_ret=%lld\n",
e_before, e_after, (long long)(int64_t)perm);
return (e_before==e_after && (int64_t)perm==0) ? 0 : 1;
}
fprintf(stderr,"unknown mode %s\n",mode); return 2;
}
@@ -0,0 +1,95 @@
/* test_interoception_p2_chrono.c — M-INTEROCEPTION Priority 2.
* Chronoception: engram_age_field(delta_ms) + reboot catch-up
* (ENGRAM_CHRONOCEPTION, default OFF).
*
* Uses loaded snapshots with KNOWN working_memory_weight / background_activation
* so the field is deterministic without depending on activation. (engram_load
* halves WM on boot — the laundering step — so snapshot wm 1.0 -> 0.5 resident.)
*
* Modes:
* once <dir> <dt_ms> — age the field once by dt; save field.json.
* split <dir> <dt_ms> <N> — age by dt/N, N times; save field.json.
* (once vs split must match: scale-invariance.)
* catchup <dir> <gap_ms> — write a last-tick gap_ms in the past, then
* engram_age_field_catchup(); print MAGNITUDE.
* offcheck <dir> <dt_ms> — flag OFF: age returns 0 and field is untouched.
*/
#include "el_runtime.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/time.h>
static el_val_t S(const char* s){ return EL_STR(s); }
static void write_seed(const char* dir){
char p[1024]; snprintf(p,sizeof p,"%s/seed.json",dir);
FILE* f=fopen(p,"w");
fprintf(f,"{\"nodes\":["
"{\"id\":\"f1\",\"content\":\"field node 1\",\"node_type\":\"Concept\",\"label\":\"f1\","
"\"salience\":0.9,\"confidence\":1.0,\"working_memory_weight\":1.0,\"background_activation\":0.5},"
"{\"id\":\"f2\",\"content\":\"field node 2\",\"node_type\":\"Concept\",\"label\":\"f2\","
"\"salience\":0.8,\"confidence\":1.0,\"working_memory_weight\":0.8,\"background_activation\":0.4},"
"{\"id\":\"f3\",\"content\":\"field node 3\",\"node_type\":\"Concept\",\"label\":\"f3\","
"\"salience\":0.7,\"confidence\":1.0,\"working_memory_weight\":0.6,\"background_activation\":0.3}"
"],\"edges\":[]}");
fclose(f);
}
static void load_seed(const char* dir){
char p[1024]; snprintf(p,sizeof p,"%s/seed.json",dir);
write_seed(dir);
if(!engram_load(S(p))){ fprintf(stderr,"load failed\n"); exit(2); }
}
static void save_field(const char* dir){
char p[1024]; snprintf(p,sizeof p,"%s/field.json",dir);
engram_save(S(p));
}
int main(int argc,char** argv){
if(argc<3){ fprintf(stderr,"usage: %s <once|split|catchup|offcheck> <dir> ...\n",argv[0]); return 2; }
const char* mode=argv[1];
const char* dir =argv[2];
if(!strcmp(mode,"once")){
double dt=atof(argv[3]);
load_seed(dir);
el_val_t mag=engram_age_field((el_val_t)(int64_t)dt);
printf("MAGNITUDE %.10f\n", el_to_float(mag));
save_field(dir);
return 0;
}
if(!strcmp(mode,"split")){
double dt=atof(argv[3]); int N=atoi(argv[4]); if(N<1)N=1;
load_seed(dir);
double sub=dt/(double)N;
for(int i=0;i<N;i++) engram_age_field((el_val_t)(int64_t)sub);
save_field(dir);
printf("SPLIT dt=%.0f N=%d sub=%.4f\n", dt, N, sub);
return 0;
}
if(!strcmp(mode,"catchup")){
double gap=atof(argv[3]);
load_seed(dir);
/* Write a last-tick gap_ms in the past. ENGRAM_DATA_DIR is set == dir by
* the runner, so the sidecar the runtime reads is <dir>/chrono_last_tick. */
struct timeval tv; gettimeofday(&tv,NULL);
long long now_ms=(long long)tv.tv_sec*1000+tv.tv_usec/1000;
long long last=now_ms-(long long)gap;
char p[1200]; snprintf(p,sizeof p,"%s/chrono_last_tick",dir);
FILE* f=fopen(p,"w"); fprintf(f,"%lld\n",last); fclose(f);
el_val_t mag=engram_age_field_catchup();
printf("CATCHUP_MAGNITUDE %.10f\n", el_to_float(mag));
save_field(dir);
return 0;
}
if(!strcmp(mode,"offcheck")){
double dt=atof(argv[3]);
load_seed(dir);
el_val_t mag=engram_age_field((el_val_t)(int64_t)dt);
el_val_t magc=engram_age_field_catchup();
printf("OFF age_mag=%.10f catchup_mag=%.10f\n", el_to_float(mag), el_to_float(magc));
save_field(dir);
return 0;
}
fprintf(stderr,"unknown mode %s\n",mode); return 2;
}
+61
View File
@@ -0,0 +1,61 @@
/* test_interoception_p3_drift.c — M-INTEROCEPTION Priority 3 (PARTIAL).
* Drift-sensor primitive engram_geo_displacement: GROWTH vs CORRUPTION split.
*
* Constructs synthetic GeoDescriptors (the struct is public) — a baseline and
* two perturbations — and checks the sensor reports LOW core-displacement for a
* periphery-only change (growth) and HIGH core-displacement for a core change
* (corruption). No store / embeddings needed: this exercises the primitive in
* isolation, which is the honest scope given there is no persisted SelfAnchor
* yet (see engram_geometry.c). */
#include "engram_geometry.h"
#include <stdlib.h>
#include <string.h>
#include <stdio.h>
static GeoMember MK(const char* id, double centrality, double dist){
GeoMember m; memset(&m,0,sizeof m);
m.id=strdup(id); m.centrality=centrality; m.dist_centroid=dist;
m.membership=1.0; m.embedded=1; return m;
}
/* 4 members: 2 core (high centrality), 2 periphery (low). */
static GeoDescriptor* mkdesc(float cx,float cy,double radius,
double c1,double c2,double p1,double p2){
GeoDescriptor* g=calloc(1,sizeof *g);
g->dim=4;
g->centroid=calloc(4,sizeof(float));
g->centroid[0]=cx; g->centroid[1]=cy;
g->radius=radius;
g->n_members=4;
g->members=calloc(4,sizeof(GeoMember));
g->members[0]=MK("core1",10.0,c1);
g->members[1]=MK("core2", 9.0,c2);
g->members[2]=MK("per1", 1.0,p1);
g->members[3]=MK("per2", 0.9,p2);
return g;
}
int main(void){
GeoDisplacement d;
/* baseline: core at 0.10, periphery at 0.50, centroid [1,0], radius 1.0 */
GeoDescriptor* A = mkdesc(1.0f,0.0f,1.0, 0.10,0.10, 0.50,0.50);
/* (i) GROWTH: periphery extends 0.50->0.90; core fixed; radius grows. */
GeoDescriptor* G = mkdesc(1.0f,0.0f,1.4, 0.10,0.10, 0.90,0.90);
engram_geo_displacement(A,G,0.5,&d);
printf("GROWTH centroid_sep=%.4f centroid_cos=%.4f radius_delta=%.4f core_disp=%.4f periph_disp=%.4f core_n=%d periph_n=%d\n",
d.centroid_sep,d.centroid_cos,d.radius_delta,d.core_disp,d.periph_disp,d.core_matched,d.periph_matched);
/* (ii) CORRUPTION: core displaces 0.10->0.60; periphery fixed; centroid shifts. */
GeoDescriptor* C = mkdesc(0.6f,0.4f,1.0, 0.60,0.60, 0.50,0.50);
engram_geo_displacement(A,C,0.5,&d);
printf("CORRUPTION centroid_sep=%.4f centroid_cos=%.4f radius_delta=%.4f core_disp=%.4f periph_disp=%.4f core_n=%d periph_n=%d\n",
d.centroid_sep,d.centroid_cos,d.radius_delta,d.core_disp,d.periph_disp,d.core_matched,d.periph_matched);
/* identity: A vs A -> zero drift */
engram_geo_displacement(A,A,0.5,&d);
printf("IDENTITY centroid_sep=%.4f core_disp=%.4f periph_disp=%.4f\n",
d.centroid_sep,d.core_disp,d.periph_disp);
engram_geo_free(A); engram_geo_free(G); engram_geo_free(C);
return 0;
}
@@ -0,0 +1,35 @@
/* test_interoception_p4_afferent.c — M-INTEROCEPTION Priority 4.
* Afferent input counters in engram_act_stats_json: additive observability.
* Drives KNOWN counts and asserts the emitted counters match and are monotonic.
*/
#include "el_runtime.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
static el_val_t S(const char* s){ return EL_STR(s); }
static el_val_t F(double d){ return el_from_float(d); }
int main(void){
/* 3 plain node creates + 2 ISE creates = 5 node_creates, 2 ise_ingests */
el_val_t a=engram_node_full(S("alpha concept about memory and time"),S("Concept"),S("a"),F(0.9),F(0.8),F(1.0),S("Semantic"),S("x"));
el_val_t b=engram_node_full(S("beta concept about memory and links"),S("Concept"),S("b"),F(0.9),F(0.8),F(1.0),S("Semantic"),S("x"));
engram_node_full(S("gamma distractor"),S("Fact"),S("c"),F(0.4),F(0.4),F(1.0),S("Semantic"),S("y"));
engram_node_full(S("heartbeat internal state one"),S("InternalStateEvent"),S("i1"),F(0.5),F(0.5),F(1.0),S("Working"),S("ise"));
engram_node_full(S("curiosity internal state two"),S("InternalStateEvent"),S("i2"),F(0.5),F(0.5),F(1.0),S("Working"),S("ise"));
/* 2 edge creates */
engram_connect(a,b,F(0.8),S("associate"));
engram_connect(b,a,F(0.3),S("associate"));
/* first reading (0 queries so far) */
printf("STATS0 %s\n", EL_CSTR(engram_act_stats_json()));
/* 4 queries -> 4 activations */
for(int i=0;i<4;i++) engram_activate_json(S("memory and time and links"), (el_val_t)2);
printf("STATS1 %s\n", EL_CSTR(engram_act_stats_json()));
/* 3 more queries -> monotonic increase */
for(int i=0;i<3;i++) engram_activate_json(S("memory and time and links"), (el_val_t)2);
printf("STATS2 %s\n", EL_CSTR(engram_act_stats_json()));
return 0;
}
@@ -0,0 +1,47 @@
/* test_interoception_p5_dreams.c — M-INTEROCEPTION Priority 5.
* Dream-recall-on-wake: engram_dreams_json(since_ms). Honesty rail only
* curiosity_scan ISEs still resident are returned; pruned (rotated-out) ones are
* ABSENT (never confabulated); heartbeat ISEs are excluded.
*/
#include "el_runtime.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/time.h>
static el_val_t S(const char* s){ return EL_STR(s); }
int main(int argc,char** argv){
if(argc<2){ fprintf(stderr,"usage: %s <dir>\n",argv[0]); return 2; }
const char* dir=argv[1];
struct timeval tv; gettimeofday(&tv,NULL);
long long now=(long long)tv.tv_sec*1000+tv.tv_usec/1000;
long long mid=now-3600000; /* 1h ago */
long long ancient=1000; /* pruned by 48h retention */
char p[1024]; snprintf(p,sizeof p,"%s/seed.json",dir);
FILE* f=fopen(p,"w");
fprintf(f,"{\"nodes\":["
"{\"id\":\"cur_old\",\"content\":\"{\\\"kind\\\":\\\"curiosity_scan\\\",\\\"q\\\":\\\"old wondering\\\"}\","
"\"node_type\":\"InternalStateEvent\",\"label\":\"state-event\",\"created_at\":%lld},"
"{\"id\":\"cur_mid\",\"content\":\"{\\\"kind\\\":\\\"curiosity_scan\\\",\\\"q\\\":\\\"mid wondering\\\"}\","
"\"node_type\":\"InternalStateEvent\",\"label\":\"state-event\",\"created_at\":%lld},"
"{\"id\":\"cur_recent\",\"content\":\"{\\\"kind\\\":\\\"curiosity_scan\\\",\\\"q\\\":\\\"recent wondering\\\"}\","
"\"node_type\":\"InternalStateEvent\",\"label\":\"state-event\",\"created_at\":%lld},"
"{\"id\":\"hb_recent\",\"content\":\"{\\\"kind\\\":\\\"heartbeat\\\",\\\"wm\\\":3}\","
"\"node_type\":\"InternalStateEvent\",\"label\":\"state-event\",\"created_at\":%lld}"
"],\"edges\":[]}", ancient, mid, now, now);
fclose(f);
if(!engram_load(S(p))){ fprintf(stderr,"load failed\n"); return 2; }
/* before prune: all resident curiosity_scan after since=0 */
printf("BEFORE %s\n", EL_CSTR(engram_dreams_json((el_val_t)0)));
/* prune 48h — cur_old (ancient) rotates out */
long long pruned=(long long)(int64_t)engram_prune_telemetry((el_val_t)0);
printf("PRUNED %lld\n", pruned);
printf("AFTER %s\n", EL_CSTR(engram_dreams_json((el_val_t)0)));
/* since filter: only events created after 30 min ago -> cur_recent only */
long long since=now-1800000;
printf("SINCE %lld %s\n", since, EL_CSTR(engram_dreams_json((el_val_t)(int64_t)since)));
return 0;
}
+208
View File
@@ -0,0 +1,208 @@
/* test_m7_traversal.c — M7 index-driven activation traversal.
*
* Milestone 7 replaces the O(E) full adjacency rebuild that spreading activation
* paid before every BFS with an incrementally-maintained per-node index, behind
* the ENGRAM_STORE flag (flag-off = unchanged behavior). This harness links the
* REAL el_runtime.c engram builtins (+ engram_store.c) and drives activation
* directly no EL interpreter, no store boot (the index optimization is a pure
* in-RAM concern; the flag is read from the environment).
*
* Modes (argv[1]):
* parity-off <dir> ENGRAM_STORE unset: build a fixed graph, run a scripted
* sequence of activations WITH mid-sequence edge/node
* inserts, dump each activation's JSON to <dir>/off_actN.json.
* parity-on <dir> ENGRAM_STORE=1: identical graph + identical sequence,
* dump to <dir>/on_actN.json. The runner asserts the off/on
* files are BYTE-IDENTICAL (same activated set, weights,
* ordering, hops, WM promotion).
* perf <off|on> <dir> <nodes> <edges> <iters>
* build a large graph, then loop `iters` times doing
* (add 1 edge + activate). Prints wall-time and the M7
* instrumentation counters (rebuild calls / rebuild
* edge-work / incremental appends).
*
* Writes ONLY under the caller-provided throwaway dir.
*/
#include "el_runtime.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
/* M7 instrumentation getters (test-only; defined in el_runtime.c). */
extern int64_t engram_adj_rebuild_calls(void);
extern int64_t engram_adj_rebuild_edge_work(void);
extern int64_t engram_adj_incr_appends(void);
extern double engram_adj_maint_seconds(void);
extern void engram_adj_test_force_dirty(void);
extern int engram_store_enabled(void);
static el_val_t S(const char* s){ return EL_STR(s); }
static el_val_t F(double d){ return el_from_float(d); }
/* Deterministic LCG so off/on processes build byte-identical graphs. */
static uint64_t g_rng = 0x9E3779B97F4A7C15ULL;
static void rng_seed(uint64_t s){ g_rng = s ? s : 1; }
static uint64_t rng_next(void){ g_rng = g_rng * 6364136223846793005ULL + 1442695040888963407ULL; return g_rng >> 17; }
static el_val_t* g_handles = NULL; /* node id handles from engram_node_full */
static int64_t g_nnodes = 0;
static void write_file(const char* path, const char* content){
FILE* f = fopen(path, "wb");
if (!f){ fprintf(stderr, "cannot open %s\n", path); exit(2); }
if (content) fwrite(content, 1, strlen(content), f);
fclose(f);
}
/* Build `n` nodes whose content carries query-matchable tokens, then `m`
* deterministic edges among them. Handles are retained for later connect. */
static void build_graph(int64_t n, int64_t m){
g_handles = malloc((size_t)n * sizeof(el_val_t));
g_nnodes = n;
static const char* topics[] = {
"storage engine durable log", "spreading activation graph traversal",
"hebbian potentiation memory", "buffer pool paging checkpoint",
"adjacency index edge lookup", "working memory promotion",
"b-tree primary index", "embeddings nearest neighbour" };
for (int64_t i = 0; i < n; i++){
char content[256];
snprintf(content, sizeof content,
"node %lld about %s and storage engine activation index",
(long long)i, topics[(size_t)(i % 8)]);
char label[32]; snprintf(label, sizeof label, "n%lld", (long long)i);
g_handles[i] = engram_node_full(S(content), S("Concept"), S(label),
F(0.7), F(0.6), F(1.0), S("Semantic"), S("storage,graph,index"));
}
for (int64_t k = 0; k < m; k++){
int64_t a = (int64_t)(rng_next() % (uint64_t)n);
int64_t b = (int64_t)(rng_next() % (uint64_t)n);
if (a == b) b = (b + 1) % n;
engram_connect(g_handles[a], g_handles[b], F(0.6), S("associate"));
}
}
static const char* Q1 = "storage engine activation and the durable log";
static const char* Q2 = "adjacency index graph traversal";
/* One scripted activation with an optional forced full-rebuild first. */
static el_val_t act(const char* q, int depth, int force_rebuild){
if (force_rebuild) engram_adj_test_force_dirty();
return engram_activate_json(S(q), (el_val_t)depth);
}
/* Run the scripted parity sequence and dump each activation JSON. `tag` names
* the output set. When force_rebuild is set, every activation first forces the
* O(E) full-rebuild path (the pre-M7 "scan" behavior); otherwise the M7
* incremental index is used. The graph build + query sequence are byte-for-byte
* deterministic, so any difference between two runs is attributable solely to
* the difference in adjacency maintenance (and/or the ENGRAM_STORE flag). */
static int run_parity(const char* dir, const char* tag, int force_rebuild){
char p[1024];
rng_seed(0xC0FFEE123ULL);
build_graph(60, 140);
el_val_t a1 = act(Q1, 3, force_rebuild);
snprintf(p, sizeof p, "%s/%s_act1.json", dir, tag); write_file(p, EL_CSTR(a1));
/* Mutate the graph BETWEEN activations: this is exactly where the M7 path
* appends incrementally while the rebuild path marks dirty + fully rebuilds.
* Parity must hold across this divergence in HOW the index is maintained. */
engram_connect(g_handles[0], g_handles[7], F(0.8), S("depends-on"));
engram_connect(g_handles[7], g_handles[23], F(0.7), S("enables"));
engram_connect(g_handles[23], g_handles[41],F(0.5), S("uses"));
el_val_t hnew = engram_node_full(S("freshly minted storage index node about activation"),
S("Concept"), S("nnew"), F(0.8), F(0.7), F(1.0), S("Semantic"), S("storage,index"));
engram_connect(g_handles[0], hnew, F(0.9), S("about"));
el_val_t a2 = act(Q1, 3, force_rebuild);
snprintf(p, sizeof p, "%s/%s_act2.json", dir, tag); write_file(p, EL_CSTR(a2));
el_val_t a3 = act(Q2, 2, force_rebuild);
snprintf(p, sizeof p, "%s/%s_act3.json", dir, tag); write_file(p, EL_CSTR(a3));
el_val_t a4 = act(Q1, 3, force_rebuild);
snprintf(p, sizeof p, "%s/%s_act4.json", dir, tag); write_file(p, EL_CSTR(a4));
printf("[parity-%s] enabled=%d force_rebuild=%d nodes=%lld edges=%lld "
"rebuilds=%lld rebuild_edge_work=%lld incr_appends=%lld\n",
tag, engram_store_enabled(), force_rebuild,
(long long)(int64_t)engram_node_count(), (long long)(int64_t)engram_edge_count(),
(long long)engram_adj_rebuild_calls(), (long long)engram_adj_rebuild_edge_work(),
(long long)engram_adj_incr_appends());
return 0;
}
static double now_sec(void){
struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts);
return (double)ts.tv_sec + (double)ts.tv_nsec * 1e-9;
}
static int run_perf(const char* dir, const char* tag, int64_t n, int64_t m, int64_t iters){
(void)dir;
rng_seed(0xBEEF7777ULL);
double t_build0 = now_sec();
build_graph(n, m);
double t_build = now_sec() - t_build0;
int64_t rb0 = engram_adj_rebuild_calls();
int64_t rw0 = engram_adj_rebuild_edge_work();
int64_t ap0 = engram_adj_incr_appends();
double mt0 = engram_adj_maint_seconds();
double t0 = now_sec();
for (int64_t it = 0; it < iters; it++){
/* One structural mutation per query — the curiosity-loop cadence that
* makes the OLD path rebuild the whole adjacency before every BFS. */
int64_t a = (int64_t)(rng_next() % (uint64_t)n);
int64_t b = (int64_t)(rng_next() % (uint64_t)n);
if (a == b) b = (b + 1) % n;
engram_connect(g_handles[a], g_handles[b], F(0.6), S("associate"));
el_val_t r = engram_activate_json(S(Q1), (el_val_t)2);
(void)r;
}
double elapsed = now_sec() - t0;
double maint = engram_adj_maint_seconds() - mt0;
printf("[perf-%s] flag=%d nodes=%lld edges=%lld iters=%lld build=%.3fs "
"loop=%.3fs per_query=%.3fms adj_maint=%.4fs adj_maint_per_query=%.4fms | "
"rebuilds=%lld rebuild_edge_work=%lld incr_appends=%lld\n",
tag, engram_store_enabled(),
(long long)(int64_t)engram_node_count(), (long long)(int64_t)engram_edge_count(),
(long long)iters, t_build, elapsed, (elapsed / (double)iters) * 1e3,
maint, (maint / (double)iters) * 1e3,
(long long)(engram_adj_rebuild_calls() - rb0),
(long long)(engram_adj_rebuild_edge_work() - rw0),
(long long)(engram_adj_incr_appends() - ap0));
return 0;
}
int main(int argc, char** argv){
if (argc < 3){ fprintf(stderr, "usage: %s <parity-off|parity-on|perf> ...\n", argv[0]); return 2; }
const char* mode = argv[1];
if (!strcmp(mode, "parity-off")){
/* flag-off, rebuild path = today's scan behavior (the baseline). */
if (engram_store_enabled()){ fprintf(stderr, "parity-off requires ENGRAM_STORE unset\n"); return 2; }
return run_parity(argv[2], "off", 0);
}
if (!strcmp(mode, "parity-on-rebuild")){
/* flag-on, but force the O(E) rebuild before each activation. */
if (!engram_store_enabled()){ fprintf(stderr, "parity-on-rebuild requires ENGRAM_STORE=1\n"); return 2; }
return run_parity(argv[2], "onrb", 1);
}
if (!strcmp(mode, "parity-on-incr")){
/* flag-on, M7 incremental index (the code path under test). */
if (!engram_store_enabled()){ fprintf(stderr, "parity-on-incr requires ENGRAM_STORE=1\n"); return 2; }
return run_parity(argv[2], "onincr", 0);
}
if (!strcmp(mode, "perf")){
/* perf <off|on> <dir> <nodes> <edges> <iters> */
if (argc < 7){ fprintf(stderr, "usage: %s perf <off|on> <dir> <nodes> <edges> <iters>\n", argv[0]); return 2; }
const char* tag = argv[2];
int64_t n = strtoll(argv[4], NULL, 10);
int64_t m = strtoll(argv[5], NULL, 10);
int64_t iters = strtoll(argv[6], NULL, 10);
return run_perf(argv[3], tag, n, m, iters);
}
fprintf(stderr, "unknown mode %s\n", mode);
return 2;
}
+255
View File
@@ -0,0 +1,255 @@
/* Closed-form unit tests for the REASONING layer (engram_reason.c). All inputs are
* hand-built synthetic descriptors whose answers are known in closed form. Every
* reasoning MODE is proven, not declared. ASan/UBSan target. */
#include "engram_reason.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <math.h>
static int failures = 0, checks = 0;
static void ok(const char* what, int cond) {
checks++;
if (!cond) { failures++; printf(" FAIL: %s\n", what); }
else printf(" ok: %s\n", what);
}
static void approx(const char* what, double got, double exp, double tol) {
ok(what, fabs(got - exp) <= tol);
if (fabs(got - exp) > tol) printf(" got=%.9g exp=%.9g\n", got, exp);
}
/* ── descriptor builders (mirror scratchpad/test_geo_ops.c) ─────────────────── */
static float* vec(const double* v, int dim) {
float* f = malloc((size_t)dim * sizeof(float));
for (int i = 0; i < dim; i++) f[i] = (float)v[i];
return f;
}
static GeoDescriptor* mk(int dim, const double* centroid,
int n_axes, const double* axis_flat, const double* extents,
int n_members, const char** ids, double total_var) {
GeoDescriptor* g = calloc(1, sizeof(GeoDescriptor));
g->dim = dim;
g->centroid = centroid ? vec(centroid, dim) : NULL;
g->global_mean = NULL;
g->n_axes = n_axes;
g->axes = n_axes ? calloc((size_t)n_axes, sizeof(GeoAxis)) : NULL;
double tr = 0;
for (int k = 0; k < n_axes; k++) {
g->axes[k].axis = vec(&axis_flat[(size_t)k * dim], dim);
g->axes[k].extent = extents[k];
tr += extents[k] * extents[k];
}
g->total_variance = (total_var >= 0) ? total_var : tr;
g->radius = sqrt(g->total_variance > 0 ? g->total_variance : 0);
g->n_members = n_members; g->n_embedded = n_members;
g->members = n_members ? calloc((size_t)n_members, sizeof(GeoMember)) : NULL;
for (int i = 0; i < n_members; i++) {
g->members[i].id = strdup(ids[i]);
g->members[i].membership = 1.0;
g->members[i].centrality = (double)(n_members - i);
g->members[i].embedded = 1;
}
g->hub_id = n_members ? strdup(ids[0]) : strdup("");
g->k_core = 1; g->co_registration = 0.0; g->n_edges = 0; g->edges = NULL;
return g;
}
int main(void) {
printf("== REASONING layer unit tests ==\n");
/* ══════════════════ ANALOGY — recover an affine A→B, apply to C ══════════ */
/* A→B is a +90° rotation in the e0-e1 plane ((x,y)→(-y,x)) plus a +5 shift in e2.
* A frame = (e0,e1); B frame = rotated (e1,-e0); cB = R·cA + t. Predict D from C. */
{
int dim = 4;
double cA[4] = {1,0,0,0};
double cB[4] = {0,1,5,0}; /* R·(1,0,0,0)=(0,1,0,0) + (0,0,5,0) */
double cC[4] = {2,0,0,0};
double axA[8] = {1,0,0,0, 0,1,0,0}; double exA[2] = {1,1};
double axB[8] = {0,1,0,0, -1,0,0,0}; double exB[2] = {1,1}; /* R·e0, R·e1 */
double axC[8] = {1,0,0,0, 0,1,0,0}; double exC[2] = {1,1};
const char* idA[1] = {"A"}, *idB[1] = {"B"}, *idC[1] = {"C"};
GeoDescriptor* A = mk(dim, cA, 2, axA, exA, 1, idA, -1);
GeoDescriptor* B = mk(dim, cB, 2, axB, exB, 1, idB, -1);
GeoDescriptor* C = mk(dim, cC, 2, axC, exC, 1, idC, -1);
/* candidates: the true D + two distractors. true D = R·cC + t = (0,2,5,0). */
double d_true[4] = {0,2,5,0}, d_far1[4] = {9,9,9,9}, d_far2[4] = {0,0,0,0};
const char* idD[1] = {"Dt"}, *idF1[1] = {"F1"}, *idF2[1] = {"F2"};
GeoDescriptor* Dt = mk(dim, d_true, 0, NULL, NULL, 1, idD, 0.0);
GeoDescriptor* F1 = mk(dim, d_far1, 0, NULL, NULL, 1, idF1, 0.0);
GeoDescriptor* F2 = mk(dim, d_far2, 0, NULL, NULL, 1, idF2, 0.0);
const GeoDescriptor* cand[3] = {F1, Dt, F2}; /* true one at index 1 */
GeoAnalogyResult res;
int rc = engram_reason_analogy(A, B, C, cand, 3, &res);
ok("analogy returns 0", rc == 0);
printf("[analogy] residual=%.6f mapped=(%.4f,%.4f,%.4f,%.4f) best=%d bd=%.5f\n",
res.analogy_residual, res.mapped_point[0], res.mapped_point[1],
res.mapped_point[2], res.mapped_point[3], res.best, res.best_distance);
approx("procrustes residual ~0", res.analogy_residual, 0.0, 1e-4);
approx("mapped.x=0", res.mapped_point[0], 0.0, 1e-4);
approx("mapped.y=2", res.mapped_point[1], 2.0, 1e-4);
approx("mapped.z(e2)=5", res.mapped_point[2], 5.0, 1e-4);
ok("nearest candidate = true D (idx 1)", res.best == 1);
approx("best distance ~0", res.best_distance, 0.0, 1e-3);
engram_reason_analogy_free(&res);
engram_geo_free(A); engram_geo_free(B); engram_geo_free(C);
engram_geo_free(Dt); engram_geo_free(F1); engram_geo_free(F2);
}
/* ══════════════════ INDUCTION — recover a shared subspace + membership ═══ */
/* 3 examples all spread over span(e0,e1) (ext 1 & 0.8), each with a small
* idiosyncratic axis (e2 or e3, ext 0.2). Centroids all 0. The induced rule's
* top-2 axes must lie in span(e0,e1); a held-out in-plane point fits, an
* off-subspace point does not. */
{
int dim = 4;
double c0[4] = {0,0,0,0};
double axsh[8] = {1,0,0,0, 0,1,0,0}; double exsh[2] = {1.0, 0.8};
double ax1[12] = {1,0,0,0, 0,1,0,0, 0,0,1,0}; double ex1[3] = {1.0,0.8,0.2}; /* +e2 */
double ax2[12] = {1,0,0,0, 0,1,0,0, 0,0,0,1}; double ex2[3] = {1.0,0.8,0.2}; /* +e3 */
const char* i1[2] = {"e1a","e1b"}, *i2[2] = {"e2a","e2b"}, *i3[2] = {"e3a","e3b"};
GeoDescriptor* E1 = mk(dim, c0, 3, ax1, ex1, 2, i1, -1);
GeoDescriptor* E2 = mk(dim, c0, 3, ax2, ex2, 2, i2, -1);
GeoDescriptor* E3 = mk(dim, c0, 2, axsh, exsh, 2, i3, -1);
const GeoDescriptor* ex[3] = {E1, E2, E3};
GeoInduction ind;
int rc = engram_reason_induce(ex, 3, 8, 1.0, &ind);
ok("induce returns 0", rc == 0);
printf("[induction] rule n_axes=%d ext0=%.4f ext1=%.4f\n",
ind.rule->n_axes, ind.rule->n_axes > 0 ? ind.rule->axes[0].extent : 0,
ind.rule->n_axes > 1 ? ind.rule->axes[1].extent : 0);
/* top-2 axes lie in span(e0,e1): their e2,e3 components ~0. */
int inplane = 1;
for (int k = 0; k < 2 && k < ind.rule->n_axes; k++) {
const float* a = ind.rule->axes[k].axis;
printf(" axis%d=(%.3f,%.3f,%.3f,%.3f) ext=%.4f\n", k, a[0],a[1],a[2],a[3], ind.rule->axes[k].extent);
if (fabs(a[2]) > 0.06 || fabs(a[3]) > 0.06) inplane = 0;
}
ok("induced top-2 axes lie in shared span(e0,e1)", inplane);
approx("dominant extent ~1.0", ind.rule->axes[0].extent, 1.0, 0.06);
approx("second extent ~0.8", ind.rule->axes[1].extent, 0.8, 0.06);
/* membership: in-plane near-centroid positive fits; off-subspace negative doesn't. */
float xpos[4] = {0.3f, -0.2f, 0, 0};
float xneg[4] = {0, 0, 3.0f, 0}; /* large along e2 — outside the rule */
float xfar[4] = {5.0f, 0, 0, 0}; /* in-plane but far — Mahalanobis blows up */
double mp = engram_reason_membership(&ind, xpos);
double mn = engram_reason_membership(&ind, xneg);
double mf = engram_reason_membership(&ind, xfar);
printf("[induction] membership pos=%.4f neg=%.4f far=%.4f\n", mp, mn, mf);
ok("held-out positive fits (>0.5)", mp > 0.5);
ok("off-subspace negative rejected (<0.3)", mn < 0.3);
ok("in-plane-but-far rejected (<0.3)", mf < 0.3);
ok("positive fits far better than negative", mp > mn + 0.4);
engram_reason_induction_free(&ind);
engram_geo_free(E1); engram_geo_free(E2); engram_geo_free(E3);
}
/* ══════════════════ ABDUCTION — pick the best-explaining structure ═══════ */
/* obs planted near H1's centroid among 3 candidate structures. */
{
int dim = 4;
double h0[4] = {0,0,0,0}, h1[4] = {5,0,0,0}, h2[4] = {0,5,0,0};
double ax[8] = {1,0,0,0, 0,1,0,0}; double ex[2] = {1,1};
const char* n0[1] = {"H0"}, *n1[1] = {"H1"}, *n2[1] = {"H2"};
GeoDescriptor* H0 = mk(dim, h0, 2, ax, ex, 1, n0, -1);
GeoDescriptor* H1 = mk(dim, h1, 2, ax, ex, 1, n1, -1);
GeoDescriptor* H2 = mk(dim, h2, 2, ax, ex, 1, n2, -1);
const GeoDescriptor* H[3] = {H0, H1, H2};
float obs[4] = {5.2f, 0.1f, 0, 0}; /* sits inside H1 */
GeoAbduction ab;
int rc = engram_reason_abduce(obs, dim, H, 3, 1.0, &ab);
ok("abduce returns 0", rc == 0);
printf("[abduction] best=%d best_score=%.4f rank=[%d,%d,%d] d=[%.3f,%.3f,%.3f]\n",
ab.best, ab.best_score, ab.rank[0], ab.rank[1], ab.rank[2],
ab.distances[0], ab.distances[1], ab.distances[2]);
ok("best explanation = H1", ab.best == 1);
ok("rank[0] = H1", ab.rank[0] == 1);
ok("H1 has smallest distance", ab.distances[1] < ab.distances[0] && ab.distances[1] < ab.distances[2]);
engram_reason_abduction_free(&ab);
engram_geo_free(H0); engram_geo_free(H1); engram_geo_free(H2);
}
/* ══════════════════ CAUSAL — direction + confounder flag ═════════════════ */
/* Chain A→B→C along e0 (temporal 1<2<3). Confounder Z (e1) injects into A and
* drives D (t=4). AD correlate only via Z must be flagged CONFOUNDED. */
{
int dim = 4;
double cA[4] = {1,1,0,0}; /* e0 (chain) + e1 (confounder leak) */
double cB[4] = {1,0,0,0}; /* e0 */
double cC[4] = {2,0,0,0}; /* e0 */
double cD[4] = {0,1,0,0}; /* e1 only — driven by Z */
double cZ[4] = {0,1,0,0}; /* confounder centroid */
double axZ[4] = {0,1,0,0}; double exZ[1] = {1}; /* Z's subspace = e1 */
const char* idA[1]={"A"},*idB[1]={"B"},*idC[1]={"C"},*idD[1]={"D"},*idZ[1]={"Z"};
GeoDescriptor* A = mk(dim, cA, 0, NULL, NULL, 1, idA, 0.0);
GeoDescriptor* B = mk(dim, cB, 0, NULL, NULL, 1, idB, 0.0);
GeoDescriptor* C = mk(dim, cC, 0, NULL, NULL, 1, idC, 0.0);
GeoDescriptor* D = mk(dim, cD, 0, NULL, NULL, 1, idD, 0.0);
GeoDescriptor* Z = mk(dim, cZ, 1, axZ, exZ, 1, idZ, -1);
const GeoDescriptor* conf[1] = {Z};
GeoCausal ab, bc, ad, bd;
engram_reason_causal(A, B, conf, 1, /*t*/1, 2, 0.5, &ab);
engram_reason_causal(B, C, conf, 1, 2, 3, 0.5, &bc);
engram_reason_causal(A, D, conf, 1, 1, 4, 0.5, &ad);
engram_reason_causal(B, D, conf, 1, 2, 4, 0.5, &bd);
printf("[causal] A->B: raw=%.3f ctrl=%.3f dir=%d verdict=%d strength=%.3f\n",
ab.assoc_raw, ab.assoc_controlled, ab.temporal_dir, ab.verdict, ab.strength);
printf("[causal] B->C: raw=%.3f ctrl=%.3f dir=%d verdict=%d\n", bc.assoc_raw, bc.assoc_controlled, bc.temporal_dir, bc.verdict);
printf("[causal] A--D: raw=%.3f ctrl=%.3f dir=%d verdict=%d confounded=%d\n",
ad.assoc_raw, ad.assoc_controlled, ad.temporal_dir, ad.verdict, ad.confounded);
printf("[causal] B--D: raw=%.3f verdict=%d\n", bd.assoc_raw, bd.verdict);
ok("A->B DIRECTED", ab.verdict == GEO_CAUSAL_DIRECTED);
ok("A->B direction A precedes B", ab.temporal_dir == 1);
ok("A->B association survives control (ctrl high)", ab.assoc_controlled > 0.6);
ok("B->C DIRECTED", bc.verdict == GEO_CAUSAL_DIRECTED);
ok("A--D CONFOUNDED (flagged)", ad.verdict == GEO_CAUSAL_CONFOUNDED && ad.confounded == 1);
ok("A--D raw correlated but control kills it", ad.assoc_raw > 0.6 && ad.assoc_controlled < 0.2);
ok("B--D NONE (no association at all)", bd.verdict == GEO_CAUSAL_NONE);
engram_geo_free(A); engram_geo_free(B); engram_geo_free(C); engram_geo_free(D); engram_geo_free(Z);
}
/* ══════════════════ PLANNING — geodesic path along a curved manifold ═════ */
/* 6 neighborhoods on a semicircle (radius 10). Consecutive chord ~6.18,
* skip-one ~11.76, endpoints ~20. neighbor_radius=7 admits only consecutive
* hops the plan must traverse the whole arc 012345. */
{
int dim = 4; int N = 6; double R = 10.0;
GeoDescriptor* nodes[6];
char nm[6][8];
for (int k = 0; k < N; k++) {
double th = M_PI * (double)k / (double)(N - 1);
double c[4] = { R * cos(th), R * sin(th), 0, 0 };
snprintf(nm[k], sizeof nm[k], "n%d", k);
const char* id[1] = { nm[k] };
nodes[k] = mk(dim, c, 0, NULL, NULL, 1, id, 0.0);
}
const GeoDescriptor* cn[6];
for (int k = 0; k < N; k++) cn[k] = nodes[k];
GeoPlan plan;
int rc = engram_reason_plan(cn, N, 0, 5, 7.0, 0, &plan);
ok("plan returns 0", rc == 0);
printf("[planning] reached=%d len=%d cost=%.4f path=[", plan.reached, plan.path_len, plan.total_cost);
for (int i = 0; i < plan.path_len; i++) printf("%s%d", i ? "," : "", plan.path[i]);
printf("]\n");
ok("goal reached", plan.reached == 1);
ok("path length = 6 (full arc)", plan.path_len == 6);
int monotone = (plan.path_len == 6);
for (int i = 0; i < plan.path_len; i++) if (plan.path[i] != i) monotone = 0;
ok("path = 0,1,2,3,4,5 (the geodesic)", monotone);
/* arc cost ~ 5 * 6.18 = 30.9, and strictly longer than the 20-unit chord. */
approx("arc cost ~30.9", plan.total_cost, 30.9, 0.6);
ok("arc longer than straight chord (20)", plan.total_cost > 20.0);
engram_reason_plan_free(&plan);
/* negative control: radius too small to connect anything ⇒ unreachable. */
GeoPlan p2;
engram_reason_plan(cn, N, 0, 5, 1.0, 0, &p2);
ok("unreachable when radius < min edge", p2.reached == 0);
engram_reason_plan_free(&p2);
for (int k = 0; k < N; k++) engram_geo_free(nodes[k]);
}
printf("\n== %d checks, %d failures ==\n", checks, failures);
return failures ? 1 : 0;
}
+163
View File
@@ -0,0 +1,163 @@
/* test_scan_collision.c — regression gate for the "saved but not findable" bug.
*
* ROOT CAUSE UNDER TEST: store_scan_nodes / store_scan_edges (the boot-load
* path that populates the resident in-RAM graph engram_store_boot ->
* eg_load_node_cb) deduplicated emitted records by their 64-bit id_hash
* (FNV-1a-64), NOT by the full id string. Two DISTINCT ids that collide under
* id_hash therefore emitted only the FIRST: the second node/edge was durably
* present in neuron.egm (store_get_node finds it), physically on a live page,
* yet was SILENTLY DROPPED from the resident load. After any store reopen it
* was unretrievable by id, absent from lexical search, and missing from the
* recent list exactly the reported symptom.
*
* The two ids below are real FNV-1a-64 collisions (found offline via Brent's
* cycle detection over fnv1a(hex16(x))); both hash to 0x15141fdadfa24abe.
*
* Pure C. Writes ONLY under a throwaway /tmp dir. Never touches ~/.neuron.
*/
#include "../../lang/runtime/engram_store.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdint.h>
#include <unistd.h>
#include <sys/stat.h>
static int g_pass = 0, g_fail = 0;
static void ok(const char* name, int cond){
printf(" [%s] %s\n", cond ? "PASS" : "FAIL", name);
if (cond) g_pass++; else g_fail++;
}
/* Confirmed FNV-1a-64 collision (distinct strings, equal id_hash). */
#define ID_A "d2c61ec7d015dc98"
#define ID_B "bf85e965a2aefbdd"
static uint64_t fnv1a(const char* s){
uint64_t h = 1469598103934665603ULL;
for (; *s; ++s){ h ^= (uint8_t)*s; h *= 1099511628211ULL; }
return h;
}
static char g_dir[512];
static void mk_dir(void){
snprintf(g_dir, sizeof g_dir, "/tmp/engram-scancol-%d", (int)getpid());
mkdir(g_dir, 0700);
}
/* ── scan collectors: record which ids the boot-load scan actually emits ── */
typedef struct { const char* want[8]; int seen[8]; int nwant; int total; } Collect;
static void node_cb(const StoreNode* n, void* ctx){
Collect* c = ctx; c->total++;
for (int i=0;i<c->nwant;i++) if (n->id && strcmp(n->id, c->want[i])==0) c->seen[i]=1;
}
static void edge_cb(const StoreEdge* e, void* ctx){
Collect* c = ctx; c->total++;
for (int i=0;i<c->nwant;i++) if (e->id && strcmp(e->id, c->want[i])==0) c->seen[i]=1;
}
static void mk_node(StoreNode* n, const char* id, const char* content){
memset(n, 0, sizeof *n);
n->id = strdup(id);
n->content = strdup(content);
n->node_type = strdup("Memory");
n->label = strdup(content);
n->tier = strdup("Working");
n->tags = strdup("");
n->metadata = strdup("{}");
n->salience = 0.5; n->importance = 0.5; n->confidence = 1.0;
n->created_at = 1700000000000LL; n->updated_at = 1700000000000LL;
n->last_activated = 1700000000000LL;
}
static void mk_edge(StoreEdge* e, const char* id, const char* from, const char* to){
memset(e, 0, sizeof *e);
e->id = strdup(id); e->from_id = strdup(from); e->to_id = strdup(to);
e->relation = strdup("assoc"); e->metadata = strdup("{}");
e->weight = 1.0; e->confidence = 1.0;
e->created_at = 1700000000000LL; e->updated_at = 1700000000000LL;
}
int main(void){
mk_dir();
printf("== scan-collision regression (saved-but-not-findable) ==\n");
printf(" id_hash(%s) = %016llx\n", ID_A, (unsigned long long)fnv1a(ID_A));
printf(" id_hash(%s) = %016llx\n", ID_B, (unsigned long long)fnv1a(ID_B));
ok("precondition: the two ids genuinely collide under id_hash",
fnv1a(ID_A) == fnv1a(ID_B) && strcmp(ID_A, ID_B) != 0);
/* ---- Control: a single node survives a full store round-trip. ---- */
{
EngramPagedStore* s = engram_open(g_dir);
StoreNode n; mk_node(&n, ID_A, "alpha distinctiveword");
store_put_node(s, &n);
engram_close(s); /* checkpoint + close */
EngramPagedStore* r = engram_open(g_dir);
StoreNode got;
ok("control: single node found by id after reopen", store_get_node(r, ID_A, &got)==1);
if (0) {} else store_node_free(&got);
Collect c = {{ID_A}, {0}, 1, 0};
store_scan_nodes(r, node_cb, &c);
ok("control: single node emitted by boot-load scan", c.seen[0]==1);
engram_close(r);
store_node_free(&n);
}
/* ---- Bug: two id-hash-colliding NODES, both durable, both must load. ---- */
{
char dir2[600]; snprintf(dir2, sizeof dir2, "%s/nodes", g_dir); mkdir(dir2, 0700);
EngramPagedStore* s = engram_open(dir2);
StoreNode a, b;
mk_node(&a, ID_A, "alpha distinctiveword-A");
mk_node(&b, ID_B, "beta distinctiveword-B");
store_put_node(s, &a);
store_put_node(s, &b);
engram_close(s);
store_node_free(&a); store_node_free(&b);
EngramPagedStore* r = engram_open(dir2);
/* Both are individually durable (store_get_node disambiguates by strcmp). */
StoreNode ga, gb;
int hit_a = store_get_node(r, ID_A, &ga); if (hit_a==1) store_node_free(&ga);
int hit_b = store_get_node(r, ID_B, &gb); if (hit_b==1) store_node_free(&gb);
ok("both colliding nodes are durably present (store_get_node)", hit_a==1 && hit_b==1);
/* THE REGRESSION: the boot-load scan must emit BOTH, not silently drop one. */
Collect c = {{ID_A, ID_B}, {0,0}, 2, 0};
store_scan_nodes(r, node_cb, &c);
printf(" scan emitted A=%d B=%d (total=%d)\n", c.seen[0], c.seen[1], c.total);
ok("boot-load scan emits node A (would be resident)", c.seen[0]==1);
ok("boot-load scan emits node B (the dropped/unretrievable one)", c.seen[1]==1);
engram_close(r);
}
/* ---- Bug: two id-hash-colliding EDGES, both must load. ---- */
{
char dir3[600]; snprintf(dir3, sizeof dir3, "%s/edges", g_dir); mkdir(dir3, 0700);
EngramPagedStore* s = engram_open(dir3);
StoreNode na, nb; mk_node(&na, "src", "s"); mk_node(&nb, "dst", "d");
store_put_node(s, &na); store_put_node(s, &nb);
StoreEdge ea, eb;
mk_edge(&ea, ID_A, "src", "dst");
mk_edge(&eb, ID_B, "src", "dst");
store_put_edge(s, &ea);
store_put_edge(s, &eb);
engram_close(s);
store_node_free(&na); store_node_free(&nb);
store_edge_free(&ea); store_edge_free(&eb);
EngramPagedStore* r = engram_open(dir3);
Collect c = {{ID_A, ID_B}, {0,0}, 2, 0};
store_scan_edges(r, edge_cb, &c);
printf(" scan emitted edgeA=%d edgeB=%d\n", c.seen[0], c.seen[1]);
ok("boot-load scan emits edge A", c.seen[0]==1);
ok("boot-load scan emits edge B (the dropped one)", c.seen[1]==1);
engram_close(r);
}
printf("\n %d passed, %d failed\n", g_pass, g_fail);
/* cleanup */
char cmd[600]; snprintf(cmd, sizeof cmd, "rm -rf %s", g_dir); if (system(cmd)){}
return g_fail ? 1 : 0;
}
+244
View File
@@ -0,0 +1,244 @@
/* Closed-form unit tests for the VERIFIER layer (engram_verify.c). Every case is a
* hand-built synthetic descriptor / claim point whose verdict is known in closed
* form the checks are PROVEN, not declared. ASan/UBSan target.
*
* The headline case is CONSISTENCY's polarity check: the reassuranceaccusation
* inversion ("you never fought" "you argued") that no grammar check catches. */
#include "engram_verify.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <math.h>
static int failures = 0, checks = 0;
static void ok(const char* what, int cond) {
checks++;
if (!cond) { failures++; printf(" FAIL: %s\n", what); }
else printf(" ok: %s\n", what);
}
static void approx(const char* what, double got, double exp, double tol) {
ok(what, fabs(got - exp) <= tol);
if (fabs(got - exp) > tol) printf(" got=%.9g exp=%.9g\n", got, exp);
}
/* ── descriptor builder (mirrors test_reason.c) ─────────────────────────────── */
static float* vec(const double* v, int dim) {
float* f = malloc((size_t)dim * sizeof(float));
for (int i = 0; i < dim; i++) f[i] = (float)v[i];
return f;
}
static GeoDescriptor* mk(int dim, const double* centroid,
int n_axes, const double* axis_flat, const double* extents,
int n_members, const char** ids, double total_var) {
GeoDescriptor* g = calloc(1, sizeof(GeoDescriptor));
g->dim = dim;
g->centroid = centroid ? vec(centroid, dim) : NULL;
g->global_mean = NULL;
g->n_axes = n_axes;
g->axes = n_axes ? calloc((size_t)n_axes, sizeof(GeoAxis)) : NULL;
double tr = 0;
for (int k = 0; k < n_axes; k++) {
g->axes[k].axis = vec(&axis_flat[(size_t)k * dim], dim);
g->axes[k].extent = extents[k];
tr += extents[k] * extents[k];
}
g->total_variance = (total_var >= 0) ? total_var : tr;
g->radius = sqrt(g->total_variance > 0 ? g->total_variance : 0);
g->n_members = n_members; g->n_embedded = n_members;
g->members = n_members ? calloc((size_t)n_members, sizeof(GeoMember)) : NULL;
for (int i = 0; i < n_members; i++) {
g->members[i].id = strdup(ids[i]);
g->members[i].membership = 1.0;
g->members[i].centrality = (double)(n_members - i);
g->members[i].embedded = 1;
}
g->hub_id = n_members ? strdup(ids[0]) : strdup("");
g->k_core = 1; g->co_registration = 0.0; g->n_edges = 0; g->edges = NULL;
return g;
}
int main(void) {
printf("== VERIFIER layer unit tests ==\n");
/* ══════════════════ GROUNDING — supported vs floating (hallucination) ════ */
/* Two real neighborhoods: E0 at origin, E1 far along e0. A claim planted inside
* E0 is grounded; a claim floating far off-manifold (along an unmodeled axis) is
* flagged UNGROUNDED; a claim near E1 grounds to E1, not E0. */
{
int dim = 4;
double c0[4] = {0,0,0,0}, c1[4] = {10,0,0,0};
double ax[8] = {1,0,0,0, 0,1,0,0}; double ex[2] = {1,1};
const char* i0[1] = {"E0"}, *i1[1] = {"E1"};
GeoDescriptor* E0 = mk(dim, c0, 2, ax, ex, 1, i0, -1);
GeoDescriptor* E1 = mk(dim, c1, 2, ax, ex, 1, i1, -1);
const GeoDescriptor* ev[2] = {E0, E1};
/* (1) grounded claim — sits inside E0. */
float in[4] = {0.3f, -0.2f, 0, 0};
GeoGrounding g1;
int rc = engram_verify_grounding(in, dim, ev, 2, 1.0, 0.5, &g1);
ok("grounding returns 0", rc == 0);
printf("[grounding] IN score=%.4f grounded=%d best=%d dist=%.3f ortho=%.3f nearL2=%.3f\n",
g1.grounding, g1.grounded, g1.best, g1.best_distance, g1.best_ortho, g1.nearest_centroid_l2);
ok("planted-inside claim is GROUNDED", g1.grounded == 1);
ok("grounds to the nearest structure E0", g1.best == 0);
ok("grounded score high (>0.7)", g1.grounding > 0.7);
approx("off-model residual ~0 for in-distribution claim", g1.best_ortho, 0.0, 1e-4);
engram_verify_grounding_free(&g1);
/* (2) hallucinated claim — floats far along the unmodeled e2 axis. */
float out[4] = {0, 0, 50.0f, 0};
GeoGrounding g2;
engram_verify_grounding(out, dim, ev, 2, 1.0, 0.5, &g2);
printf("[grounding] OUT score=%.6f grounded=%d best=%d dist=%.3f ortho=%.3f nearL2=%.3f\n",
g2.grounding, g2.grounded, g2.best, g2.best_distance, g2.best_ortho, g2.nearest_centroid_l2);
ok("floating claim is FLAGGED (ungrounded)", g2.grounded == 0);
ok("floating claim scores near zero (<0.01)", g2.grounding < 0.01);
ok("off-model residual is large (the hallucination signal)", g2.best_ortho > 40.0);
ok("nearest real structure is far (L2>40)", g2.nearest_centroid_l2 > 40.0);
engram_verify_grounding_free(&g2);
/* (3) selection — a claim near E1 grounds to E1. */
float nearE1[4] = {9.8f, 0.1f, 0, 0};
GeoGrounding g3;
engram_verify_grounding(nearE1, dim, ev, 2, 1.0, 0.5, &g3);
printf("[grounding] E1 score=%.4f grounded=%d best=%d\n", g3.grounding, g3.grounded, g3.best);
ok("claim near E1 grounds to E1 (best=1)", g3.best == 1 && g3.grounded == 1);
engram_verify_grounding_free(&g3);
engram_geo_free(E0); engram_geo_free(E1);
}
/* ══════════════════ CONSISTENCY (a) — THE NEGATION-INVERSION CATCH ═══════ */
/* The motivating failure, geometrically. Polarity axis along e0:
* pole_pos = the AFFIRM region ("argued / fought") centroid (+5, )
* pole_neg = the NEGATE region ("never fought / at peace") centroid (5, )
* The grounded TRUTH (context) is the reassurance "you never fought" sits on
* the NEGATE side (5). The bad translation CLAIM "you argued" lands on the
* AFFIRM side (+4). Opposite sides of the negation axis INVERSION flagged
* even though "you argued" is perfectly grammatical. This is the catch. */
{
int dim = 4;
double c_pos[4] = { 5, 0, 0, 0}; /* "argued / fought" */
double c_neg[4] = {-5, 0, 0, 0}; /* "never fought / at peace"*/
double c_truth[4] = {-5, 0, 0, 0}; /* context: the reassurance */
double ax[4] = {1,0,0,0}; double ex[1] = {1};
const char* ip[1]={"pos"},*in[1]={"neg"},*it[1]={"truth"};
GeoDescriptor* POS = mk(dim, c_pos, 1, ax, ex, 1, ip, -1);
GeoDescriptor* NEG = mk(dim, c_neg, 1, ax, ex, 1, in, -1);
GeoDescriptor* CTX = mk(dim, c_truth, 1, ax, ex, 1, it, -1);
/* the plausible LIE: "you argued" — grammatical, fluent, and INVERTED. */
float lie[4] = { 4, 0, 0, 0};
GeoConsistency cl;
int rc = engram_verify_consistency(lie, dim, CTX, POS, NEG, NULL,
1.0, 0.10, 0.5, 0.0, &cl);
ok("consistency returns 0", rc == 0);
printf("[consistency] LIE verdict=%d inverted=%d claim_side=%.3f ref_side=%.3f sep=%.3f consist=%.3f\n",
cl.verdict, cl.inverted, cl.polarity_claim, cl.polarity_reference, cl.polarity_separation, cl.consistency);
ok("NEGATION INVERSION caught (inverted=1)", cl.inverted == 1);
ok("verdict = POLARITY", cl.verdict == GEO_CONSIST_POLARITY);
ok("claim sits on the AFFIRM pole (+)", cl.polarity_claim > 0);
ok("truth sits on the NEGATE pole ()", cl.polarity_reference < 0);
ok("consistency collapses to 0 on inversion", cl.consistency < 1e-9);
/* the FAITHFUL translation: "you were at peace" — same pole as the truth. */
float ok_claim[4] = {-4, 0, 0, 0};
GeoConsistency cok;
engram_verify_consistency(ok_claim, dim, CTX, POS, NEG, NULL,
1.0, 0.10, 0.5, 0.0, &cok);
printf("[consistency] TRUE verdict=%d inverted=%d claim_side=%.3f consist=%.3f\n",
cok.verdict, cok.inverted, cok.polarity_claim, cok.consistency);
ok("faithful claim NOT flagged (inverted=0)", cok.inverted == 0);
ok("faithful claim verdict OK", cok.verdict == GEO_CONSIST_OK);
ok("faithful claim consistency = 1", cok.consistency > 0.999);
/* a NEUTRAL claim near the midpoint must NOT false-trigger. */
float neutral[4] = {0.1f, 0, 0, 0}; /* |side|=0.1 < deadzone 0.5 */
GeoConsistency cn;
engram_verify_consistency(neutral, dim, CTX, POS, NEG, NULL,
1.0, 0.10, 0.5, 0.0, &cn);
printf("[consistency] NEUT verdict=%d inverted=%d claim_side=%.3f consist=%.3f\n",
cn.verdict, cn.inverted, cn.polarity_claim, cn.consistency);
ok("neutral claim inside deadzone does NOT trigger inversion", cn.inverted == 0);
engram_geo_free(POS); engram_geo_free(NEG); engram_geo_free(CTX);
}
/* ══════════════════ CONSISTENCY (b) — GEOMETRIC contradiction ════════════ */
/* A claim that sits INSIDE a forbidden region it should be far from, and a claim
* that violates a max-distance constraint to its context, are both flagged. */
{
int dim = 4;
double c_ctx[4] = {0,0,0,0};
double c_forb[4] = {0,10,0,0}; /* forbidden region, offset along e1 */
double ax[8] = {0,1,0,0, 1,0,0,0}; double ex[2] = {1,1};
const char* ic[1]={"ctx"},*ifb[1]={"forb"};
GeoDescriptor* CTX = mk(dim, c_ctx, 2, ax, ex, 1, ic, -1);
GeoDescriptor* FORB = mk(dim, c_forb, 2, ax, ex, 1, ifb, -1);
/* claim sitting inside the forbidden region → geometric contradiction. */
float inside[4] = {0, 10.1f, 0, 0};
GeoConsistency cf;
engram_verify_consistency(inside, dim, CTX, NULL, NULL, FORB,
1.0, 0.10, 0.5, 0.0, &cf);
printf("[consistency] FORB verdict=%d geo_viol=%d forb_fit=%.4f consist=%.3f\n",
cf.verdict, cf.geo_violation, cf.forbidden_fit, cf.consistency);
ok("claim inside forbidden region FLAGGED", cf.geo_violation == 1);
ok("verdict = GEOMETRIC", cf.verdict == GEO_CONSIST_GEOMETRIC);
ok("forbidden fit is high (claim really is inside)", cf.forbidden_fit > 0.5);
/* claim well clear of the forbidden region → not flagged. */
float clear[4] = {0.2f, 0.1f, 0, 0};
GeoConsistency cc;
engram_verify_consistency(clear, dim, CTX, NULL, NULL, FORB,
1.0, 0.10, 0.5, 0.0, &cc);
printf("[consistency] CLR verdict=%d geo_viol=%d forb_fit=%.4f\n",
cc.verdict, cc.geo_violation, cc.forbidden_fit);
ok("claim clear of forbidden NOT flagged", cc.geo_violation == 0 && cc.verdict == GEO_CONSIST_OK);
/* max-distance constraint: claim too far from context (off-axis, no poles). */
float far[4] = {0, 8.0f, 0, 0};
GeoConsistency cd;
engram_verify_consistency(far, dim, CTX, NULL, NULL, NULL,
1.0, 0.10, 0.5, /*max_distance*/3.0, &cd);
printf("[consistency] DIST verdict=%d geo_viol=%d ctx_dist=%.3f\n",
cd.verdict, cd.geo_violation, cd.context_distance);
ok("claim beyond max_distance FLAGGED", cd.geo_violation == 1 && cd.verdict == GEO_CONSIST_GEOMETRIC);
approx("context distance measured correctly", cd.context_distance, 8.0, 1e-4);
engram_geo_free(CTX); engram_geo_free(FORB);
}
/* ══════════════════ COMBINED — grounded but INVERTED (the full plausible lie) */
/* The most dangerous output: fluent, GROUNDED in real vocabulary, yet polarity-
* inverted. Grounding alone passes it; only consistency catches the lie. This is
* exactly why the verifier needs BOTH checks. */
{
int dim = 4;
double c_pos[4] = { 5, 0, 0, 0}, c_neg[4] = {-5, 0, 0, 0};
double ax[4] = {1,0,0,0}; double ex[1] = {2};
const char* ip[1]={"pos"},*in[1]={"neg"};
GeoDescriptor* POS = mk(dim, c_pos, 1, ax, ex, 1, ip, -1);
GeoDescriptor* NEG = mk(dim, c_neg, 1, ax, ex, 1, in, -1);
const GeoDescriptor* ev[2] = {POS, NEG};
float lie[4] = {5, 0, 0, 0}; /* "argued" — sits dead-center in the affirm region */
GeoGrounding g;
engram_verify_grounding(lie, dim, ev, 2, 1.0, 0.5, &g);
GeoConsistency c;
engram_verify_consistency(lie, dim, NEG /*truth=never fought*/, POS, NEG, NULL,
1.0, 0.10, 0.5, 0.0, &c);
printf("[combined] grounded=%d (score=%.3f) inverted=%d verdict=%d\n",
g.grounded, g.grounding, c.inverted, c.verdict);
ok("plausible lie PASSES grounding (it is real vocabulary)", g.grounded == 1);
ok("plausible lie is CAUGHT by consistency (inverted)", c.inverted == 1);
ok("=> grounding alone is insufficient; consistency is the catch",
g.grounded == 1 && c.verdict == GEO_CONSIST_POLARITY);
engram_verify_grounding_free(&g);
engram_geo_free(POS); engram_geo_free(NEG);
}
printf("\n== %d checks, %d failures ==\n", checks, failures);
return failures ? 1 : 0;
}
+312
View File
@@ -0,0 +1,312 @@
/* test_vindex.c — build + RUN gate for the M8 HNSW vector index.
*
* Covers: recall@10 vs brute-force oracle, brute-force-vs-index speedup,
* correctness edge cases (k>N, identical vectors, self-query, zero vector),
* determinism (seeded PRNG identical graphs), and vindex_build_from_store
* over a real engram_store on-disk file.
*
* Pure C11; links engram_vindex.c + engram_store.c; -lm. ASan/UBSan clean.
*/
#include "engram_vindex.h"
#include "engram_store.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <math.h>
#include <stdint.h>
#include <time.h>
#include <unistd.h>
#define DIM 768
static int g_fail = 0;
/* VINDEX_QUICK=1 shrinks the two large builds so the ASan/UBSan pass (which runs
* ~5-10x slower) stays fast memory-safety is size-independent. The perf numbers
* (recall gate + speedup) come from the un-sanitized, full-size pass. */
static int g_quick = 0;
static int envint(const char* k, int dflt){ const char* s=getenv(k); return s?atoi(s):dflt; }
#define CHECK(cond, msg) do{ if(!(cond)){ printf(" FAIL: %s\n", msg); g_fail=1; } else { printf(" ok: %s\n", msg); } }while(0)
/* deterministic test PRNG (splitmix64) */
static uint64_t rng_state = 0xABCDEF0123456789ULL;
static uint64_t xrng(uint64_t* s){
uint64_t z=(*s+=0x9E3779B97F4A7C15ULL);
z=(z^(z>>30))*0xBF58476D1CE4E5B9ULL; z=(z^(z>>27))*0x94D049BB133111EBULL;
return z^(z>>31);
}
static float frand(uint64_t* s){ return (float)((xrng(s)>>11)*(1.0/9007199254740992.0)) - 0.5f; }
static double now_s(void){
struct timespec t; clock_gettime(CLOCK_MONOTONIC,&t);
return t.tv_sec + t.tv_nsec*1e-9;
}
/* fill vec[N*DIM]: mostly random, some clustered groups (center + small noise). */
static void gen_vectors(float* v, int N, uint64_t seed){
uint64_t s = seed;
int clustered = N/5; /* last fifth is clustered */
int ncenters = 20;
float* centers = (float*)malloc((size_t)ncenters*DIM*sizeof(float));
for (int c=0;c<ncenters;c++) for(int d=0;d<DIM;d++) centers[c*DIM+d]=frand(&s);
for (int i=0;i<N;i++){
if (i < N-clustered){
for (int d=0;d<DIM;d++) v[i*DIM+d]=frand(&s);
} else {
int c = (int)(xrng(&s)%ncenters);
for (int d=0;d<DIM;d++) v[i*DIM+d]=centers[c*DIM+d] + 0.05f*frand(&s);
}
}
free(centers);
}
static float cosdist(const float* a, const float* b){
double da=0,db=0,dot=0;
for(int i=0;i<DIM;i++){ da+=(double)a[i]*a[i]; db+=(double)b[i]*b[i]; dot+=(double)a[i]*b[i]; }
if (da<=0||db<=0) return 1.0f;
return (float)(1.0 - dot/(sqrt(da)*sqrt(db)));
}
/* brute-force top-k node ids into ids[k] (ascending distance). */
static void brute_topk(const float* v, int N, const float* q, int k, int* ids){
float* bd = (float*)malloc((size_t)k*sizeof(float));
for (int i=0;i<k;i++){ ids[i]=-1; bd[i]=1e30f; }
for (int i=0;i<N;i++){
float d = cosdist(q, v+(size_t)i*DIM);
if (d < bd[k-1]){
int p=k-1;
while (p>0 && bd[p-1]>d){ bd[p]=bd[p-1]; ids[p]=ids[p-1]; p--; }
bd[p]=d; ids[p]=i;
}
}
free(bd);
}
/* ── Test 1: recall@10 vs brute force + latency/recall tradeoff ────────────── */
static void test_recall(void){
int N=envint("VINDEX_N_RECALL", g_quick?1500:5000), Q=200, K=10;
printf("\n== Test 1: recall@10 vs brute force (N=%d, DIM=768) ==\n", N);
float* v = (float*)malloc((size_t)N*DIM*sizeof(float));
gen_vectors(v, N, 111);
double t0=now_s();
VIndex* ix = vindex_create(DIM, VINDEX_DEFAULT_M, VINDEX_DEFAULT_EF_CONSTRUCTION);
for (int i=0;i<N;i++) vindex_insert(ix, (uint64_t)i, v+(size_t)i*DIM);
double build_s = now_s()-t0;
printf(" build: %d vectors in %.2fs (M=%d, ef_construction=%d)\n",
N, build_s, VINDEX_DEFAULT_M, VINDEX_DEFAULT_EF_CONSTRUCTION);
/* queries: half random, half near a real vector (perturbed). */
float* qs = (float*)malloc((size_t)Q*DIM*sizeof(float));
uint64_t s=999;
for (int i=0;i<Q;i++){
if (i<Q/2) for(int d=0;d<DIM;d++) qs[i*DIM+d]=frand(&s);
else { int base=(int)(xrng(&s)%N); for(int d=0;d<DIM;d++) qs[i*DIM+d]=v[base*DIM+d]+0.03f*frand(&s); }
}
/* oracle */
int* oracle = (int*)malloc((size_t)Q*K*sizeof(int));
for (int i=0;i<Q;i++) brute_topk(v, N, qs+(size_t)i*DIM, K, oracle+(size_t)i*K);
int efs[] = { 10, 32, 64, 128 };
for (int e=0;e<4;e++){
int ef=efs[e];
uint64_t ids[64]; float dd[64];
int hits=0;
double qt0=now_s();
for (int i=0;i<Q;i++){
int n=vindex_search(ix, qs+(size_t)i*DIM, K, ef, ids, dd);
for (int a=0;a<n;a++) for(int b=0;b<K;b++) if((int)ids[a]==oracle[i*K+b]){ hits++; break; }
}
double qs_ms = (now_s()-qt0)*1000.0/Q;
double recall = (double)hits/(Q*K);
printf(" ef_search=%-4d recall@10=%.4f latency=%.3f ms/query\n", ef, recall, qs_ms);
if (ef==VINDEX_DEFAULT_EF_SEARCH && !g_quick)
CHECK(recall >= 0.90, "recall@10 >= 0.90 at default ef_search=128");
}
free(oracle); free(qs); free(v); vindex_free(ix);
}
/* ── Test 2: speedup vs brute force ───────────────────────────────────────── */
static void speedup_at(int N){
int Q=100, K=10;
float* v=(float*)malloc((size_t)N*DIM*sizeof(float));
gen_vectors(v,N,222);
VIndex* ix=vindex_create(DIM,16,200);
double bt0=now_s();
for(int i=0;i<N;i++) vindex_insert(ix,(uint64_t)i,v+(size_t)i*DIM);
printf(" N=%d build=%.2fs\n", N, now_s()-bt0);
float* qs=(float*)malloc((size_t)Q*DIM*sizeof(float));
uint64_t s=333; for(int i=0;i<Q*DIM;i++) qs[i]=frand(&s);
/* brute force */
int scratch[16];
double b0=now_s();
for(int i=0;i<Q;i++) brute_topk(v,N,qs+(size_t)i*DIM,K,scratch);
double bf=(now_s()-b0)/Q;
/* index */
uint64_t ids[16]; float dd[16];
double i0=now_s();
for(int i=0;i<Q;i++) vindex_search(ix,qs+(size_t)i*DIM,K,64,ids,dd);
double iq=(now_s()-i0)/Q;
printf(" N=%d brute=%.4f ms/q index=%.4f ms/q speedup=%.1fx\n",
N, bf*1000, iq*1000, bf/iq);
CHECK(iq < bf, "index query faster than brute force");
free(qs); free(v); vindex_free(ix);
}
static void test_speedup(void){
printf("\n== Test 2: brute-force vs index speedup ==\n");
speedup_at(g_quick?2000:5000);
speedup_at(envint("VINDEX_N_BIG", g_quick?3000:20000));
}
/* ── Test 3: edge cases ───────────────────────────────────────────────────── */
static void test_edges(void){
printf("\n== Test 3: correctness edge cases ==\n");
/* k larger than node count */
{
VIndex* ix=vindex_create(DIM,16,200);
float vec[DIM]; uint64_t s=1;
for(int i=0;i<3;i++){ for(int d=0;d<DIM;d++) vec[d]=frand(&s); vindex_insert(ix,(uint64_t)i,vec); }
uint64_t ids[50]; float dd[50];
int n=vindex_search(ix, vec, 50, 64, ids, dd);
CHECK(n==3, "k > node count returns exactly node-count results");
vindex_free(ix);
}
/* duplicate / identical vectors */
{
VIndex* ix=vindex_create(DIM,16,200);
float a[DIM]; uint64_t s=2; for(int d=0;d<DIM;d++) a[d]=frand(&s);
for(int i=0;i<10;i++) vindex_insert(ix,(uint64_t)i,a); /* all identical */
float b[DIM]; for(int d=0;d<DIM;d++) b[d]=frand(&s);
vindex_insert(ix,100,b);
uint64_t ids[5]; float dd[5];
int n=vindex_search(ix,a,5,64,ids,dd);
CHECK(n==5, "identical-vector index returns k results");
CHECK(dd[0] < 1e-4f, "top-1 distance ~0 for a duplicated vector");
vindex_free(ix);
}
/* query equal to an indexed vector returns itself as top-1, dist ~0 */
{
VIndex* ix=vindex_create(DIM,16,200);
int N=500; float* v=(float*)malloc((size_t)N*DIM*sizeof(float)); gen_vectors(v,N,7);
for(int i=0;i<N;i++) vindex_insert(ix,(uint64_t)(1000+i),v+(size_t)i*DIM);
int probe=137;
uint64_t ids[3]; float dd[3];
int n=vindex_search(ix, v+(size_t)probe*DIM, 3, 64, ids, dd);
CHECK(n>=1 && ids[0]==(uint64_t)(1000+probe), "self-query returns itself as top-1");
CHECK(dd[0] < 1e-4f, "self-query top-1 distance ~0");
free(v); vindex_free(ix);
}
/* zero vector: no NaN, handled */
{
VIndex* ix=vindex_create(DIM,16,200);
float z[DIM]; memset(z,0,sizeof z);
float a[DIM]; uint64_t s=3; for(int d=0;d<DIM;d++) a[d]=frand(&s);
vindex_insert(ix,0,z); vindex_insert(ix,1,a);
uint64_t ids[2]; float dd[2];
int n=vindex_search(ix, z, 2, 64, ids, dd); /* zero query */
int nan=0; for(int i=0;i<n;i++) if(isnan(dd[i])||isinf(dd[i])) nan=1;
CHECK(n>=1 && !nan, "zero vector query produces no NaN/Inf");
n=vindex_search(ix, a, 2, 64, ids, dd); /* zero indexed */
nan=0; for(int i=0;i<n;i++) if(isnan(dd[i])||isinf(dd[i])) nan=1;
CHECK(!nan, "indexed zero vector produces no NaN/Inf");
vindex_free(ix);
}
}
/* ── Test 4: determinism ──────────────────────────────────────────────────── */
static void test_determinism(void){
printf("\n== Test 4: determinism (seeded PRNG → identical results) ==\n");
int N=1500;
float* v=(float*)malloc((size_t)N*DIM*sizeof(float)); gen_vectors(v,N,55);
uint64_t ids1[10],ids2[10]; float d1[10],d2[10];
int identical=1;
for (int build=0; build<2; build++){
VIndex* ix=vindex_create(DIM,16,200);
for(int i=0;i<N;i++) vindex_insert(ix,(uint64_t)i,v+(size_t)i*DIM);
/* probe several queries */
for (int q=0;q<20;q++){
uint64_t* ida = build? ids2 : ids1; float* da = build? d2 : d1;
vindex_search(ix, v+(size_t)(q*37%N)*DIM, 10, 64, ida, da);
if (build==1){
/* re-run build-0 query stored? simpler: compare within-run below */
}
}
vindex_free(ix);
}
/* Proper comparison: run two fresh builds, same single query. */
identical=1;
for (int q=0;q<25;q++){
int qi=(q*61)%N;
VIndex* a=vindex_create(DIM,16,200); for(int i=0;i<N;i++) vindex_insert(a,(uint64_t)i,v+(size_t)i*DIM);
VIndex* b=vindex_create(DIM,16,200); for(int i=0;i<N;i++) vindex_insert(b,(uint64_t)i,v+(size_t)i*DIM);
int na=vindex_search(a, v+(size_t)qi*DIM,10,64,ids1,d1);
int nb=vindex_search(b, v+(size_t)qi*DIM,10,64,ids2,d2);
if (na!=nb) identical=0;
for(int i=0;i<na;i++) if(ids1[i]!=ids2[i] || d1[i]!=d2[i]) identical=0;
vindex_free(a); vindex_free(b);
}
CHECK(identical, "two independent builds give byte-identical query results");
free(v);
}
/* ── Test 5: build_from_store ─────────────────────────────────────────────── */
static void test_build_from_store(void){
printf("\n== Test 5: vindex_build_from_store over a real engram_store ==\n");
char path[256];
snprintf(path,sizeof path,"/tmp/vindex_test_store_%d.engram",(int)getpid());
unlink(path);
EngramPagedStore* st = store_create(path);
if (!st){ printf(" FAIL: store_create\n"); g_fail=1; return; }
int N=300;
float* v=(float*)malloc((size_t)N*DIM*sizeof(float)); gen_vectors(v,N,88);
for (int i=0;i<N;i++){
StoreNode n; memset(&n,0,sizeof n);
char id[32]; snprintf(id,sizeof id,"node-%d",i);
n.id=id; n.content="x"; n.node_type="concept"; n.tier="Semantic";
n.emb = v+(size_t)i*DIM; n.emb_dim=DIM;
if (store_put_node(st,&n)!=0){ printf(" FAIL: put_node %d\n",i); g_fail=1; }
}
/* a node WITHOUT an emb — must be skipped by build_from_store. */
{ StoreNode n; memset(&n,0,sizeof n); n.id=(char*)"no-emb"; n.content="y"; n.node_type="concept"; n.tier="Semantic";
store_put_node(st,&n); }
store_close(st);
VIndex* ix = vindex_create(DIM,16,200);
char** ids=NULL; int nids=0;
int ins = vindex_build_from_store(ix, path, &ids, &nids);
printf(" build_from_store inserted %d vectors (expected %d; 1 emb-less skipped)\n", ins, N);
CHECK(ins==N, "build_from_store inserts exactly the emb'd nodes");
CHECK((size_t)ins==vindex_size(ix), "index size matches insert count");
/* query with a known vector → must return its own node id as top-1. */
int probe=42;
uint64_t rids[5]; float dd[5];
int n=vindex_search(ix, v+(size_t)probe*DIM, 5, 64, rids, dd);
int correct = (n>=1 && rids[0]<(uint64_t)nids && strcmp(ids[rids[0]], "node-42")==0);
printf(" query for node-42's vector → top-1 id=%s dist=%.5f\n",
(n>=1 && rids[0]<(uint64_t)nids)? ids[rids[0]] : "?", n?dd[0]:-1);
CHECK(correct, "build_from_store query resolves to the right node id");
CHECK(n>=1 && dd[0]<1e-4f, "top-1 distance ~0 for exact stored vector");
for (int i=0;i<nids;i++) free(ids[i]);
free(ids); free(v); vindex_free(ix); unlink(path);
}
int main(void){
(void)rng_state;
g_quick = envint("VINDEX_QUICK", 0);
printf("=== engram_vindex (HNSW) test suite ===%s\n", g_quick?" [QUICK]":"");
test_recall();
test_speedup();
test_edges();
test_determinism();
test_build_from_store();
printf("\n=== %s ===\n", g_fail? "FAILURES PRESENT" : "ALL TESTS PASSED");
return g_fail;
}
+12
View File
@@ -2633,6 +2633,12 @@ fn builtin_arity(name: String) -> Int {
if str_eq(name, "__engram_neighbors_filtered") { return 3 } if str_eq(name, "__engram_neighbors_filtered") { return 3 }
if str_eq(name, "__engram_activate") { return 2 } if str_eq(name, "__engram_activate") { return 2 }
if str_eq(name, "__engram_activate_json") { return 2 } if str_eq(name, "__engram_activate_json") { return 2 }
if str_eq(name, "__engram_geo_descriptor_json") { return 1 }
if str_eq(name, "__engram_geo_overlap_json") { return 2 }
if str_eq(name, "__engram_geo_subtract_json") { return 3 }
if str_eq(name, "__engram_geo_combine_json") { return 2 }
if str_eq(name, "__engram_geo_distance_json") { return 2 }
if str_eq(name, "__engram_geo_analogy_json") { return 2 }
if str_eq(name, "__engram_scan_nodes_json") { return 2 } if str_eq(name, "__engram_scan_nodes_json") { return 2 }
if str_eq(name, "__generate") { return 1 } if str_eq(name, "__generate") { return 1 }
// Filesystem // Filesystem
@@ -2732,6 +2738,12 @@ fn builtin_arity(name: String) -> Int {
if str_eq(name, "engram_scan_nodes_json") { return 2 } if str_eq(name, "engram_scan_nodes_json") { return 2 }
if str_eq(name, "engram_neighbors_json") { return 3 } if str_eq(name, "engram_neighbors_json") { return 3 }
if str_eq(name, "engram_activate_json") { return 2 } if str_eq(name, "engram_activate_json") { return 2 }
if str_eq(name, "engram_geo_descriptor_json") { return 1 }
if str_eq(name, "engram_geo_overlap_json") { return 2 }
if str_eq(name, "engram_geo_subtract_json") { return 3 }
if str_eq(name, "engram_geo_combine_json") { return 2 }
if str_eq(name, "engram_geo_distance_json") { return 2 }
if str_eq(name, "engram_geo_analogy_json") { return 2 }
if str_eq(name, "engram_stats_json") { return 0 } if str_eq(name, "engram_stats_json") { return 0 }
// LLM // LLM
if str_eq(name, "llm_call") { return 2 } if str_eq(name, "llm_call") { return 2 }
+36
View File
@@ -6748,6 +6748,24 @@ el_val_t builtin_arity(el_val_t name) {
if (str_eq(name, EL_STR("__engram_activate_json"))) { if (str_eq(name, EL_STR("__engram_activate_json"))) {
return 2; return 2;
} }
if (str_eq(name, EL_STR("__engram_geo_descriptor_json"))) {
return 1;
}
if (str_eq(name, EL_STR("__engram_geo_overlap_json"))) {
return 2;
}
if (str_eq(name, EL_STR("__engram_geo_subtract_json"))) {
return 3;
}
if (str_eq(name, EL_STR("__engram_geo_combine_json"))) {
return 2;
}
if (str_eq(name, EL_STR("__engram_geo_distance_json"))) {
return 2;
}
if (str_eq(name, EL_STR("__engram_geo_analogy_json"))) {
return 2;
}
if (str_eq(name, EL_STR("__engram_scan_nodes_json"))) { if (str_eq(name, EL_STR("__engram_scan_nodes_json"))) {
return 2; return 2;
} }
@@ -6991,6 +7009,24 @@ el_val_t builtin_arity(el_val_t name) {
if (str_eq(name, EL_STR("engram_activate_json"))) { if (str_eq(name, EL_STR("engram_activate_json"))) {
return 2; return 2;
} }
if (str_eq(name, EL_STR("engram_geo_descriptor_json"))) {
return 1;
}
if (str_eq(name, EL_STR("engram_geo_overlap_json"))) {
return 2;
}
if (str_eq(name, EL_STR("engram_geo_subtract_json"))) {
return 3;
}
if (str_eq(name, EL_STR("engram_geo_combine_json"))) {
return 2;
}
if (str_eq(name, EL_STR("engram_geo_distance_json"))) {
return 2;
}
if (str_eq(name, EL_STR("engram_geo_analogy_json"))) {
return 2;
}
if (str_eq(name, EL_STR("engram_stats_json"))) { if (str_eq(name, EL_STR("engram_stats_json"))) {
return 0; return 0;
} }
+1007 -19
View File
File diff suppressed because it is too large Load Diff
+17
View File
@@ -621,6 +621,23 @@ el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit); el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset); el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset); el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_emb_json(el_val_t limit, el_val_t offset);
el_val_t engram_dreams_json(el_val_t since_ms);
/* §5 geometry operators as EL builtins (read-only; seed-id CSV args). */
el_val_t engram_geo_descriptor_json(el_val_t seeds);
el_val_t engram_geo_overlap_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_geo_subtract_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t mode);
el_val_t engram_geo_combine_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_geo_distance_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_geo_analogy_json(el_val_t a_seeds, el_val_t b_seeds);
/* reasoning layer (compositions over §5 operators). ANALOGY maps cleanly to the
* flat-CSV seed ABI; the other modes take set/point/timestamp inputs deferred from
* this ABI (see engram_reason.h / the reasoning-operators runbook). */
el_val_t engram_reason_analogy_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t c_seeds);
el_val_t engram_consolidate_permanence(el_val_t node_id);
el_val_t engram_age_field(el_val_t delta_ms);
el_val_t engram_age_field_catchup(void);
el_val_t engram_chrono_persist_tick(void);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction); el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_activate_json(el_val_t query, el_val_t depth); el_val_t engram_activate_json(el_val_t query, el_val_t depth);
el_val_t engram_stats_json(void); el_val_t engram_stats_json(void);
+41
View File
@@ -1086,6 +1086,47 @@ el_val_t __engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el
return engram_scan_nodes_by_type_json(node_type, limit, offset); return engram_scan_nodes_by_type_json(node_type, limit, offset);
} }
el_val_t __engram_scan_nodes_emb_json(el_val_t limit, el_val_t offset) {
return engram_scan_nodes_emb_json(limit, offset);
}
el_val_t __engram_dreams_json(el_val_t since_ms) {
return engram_dreams_json(since_ms);
}
/* §5 geometry operators — native wrappers (surfacing via engram.el + elc fold is
* the cutover step; the C table wiring is registered here now, per P0/P5). */
el_val_t __engram_geo_descriptor_json(el_val_t seeds) {
return engram_geo_descriptor_json(seeds);
}
el_val_t __engram_geo_overlap_json(el_val_t a_seeds, el_val_t b_seeds) {
return engram_geo_overlap_json(a_seeds, b_seeds);
}
el_val_t __engram_geo_subtract_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t mode) {
return engram_geo_subtract_json(a_seeds, b_seeds, mode);
}
el_val_t __engram_geo_combine_json(el_val_t a_seeds, el_val_t b_seeds) {
return engram_geo_combine_json(a_seeds, b_seeds);
}
el_val_t __engram_geo_distance_json(el_val_t a_seeds, el_val_t b_seeds) {
return engram_geo_distance_json(a_seeds, b_seeds);
}
el_val_t __engram_geo_analogy_json(el_val_t a_seeds, el_val_t b_seeds) {
return engram_geo_analogy_json(a_seeds, b_seeds);
}
/* reasoning layer — ANALOGY native wrapper (same C-table wiring as the §5 ops). */
el_val_t __engram_reason_analogy_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t c_seeds) {
return engram_reason_analogy_json(a_seeds, b_seeds, c_seeds);
}
el_val_t __engram_consolidate_permanence(el_val_t node_id) {
return engram_consolidate_permanence(node_id);
}
el_val_t __engram_age_field(el_val_t delta_ms) { return engram_age_field(delta_ms); }
el_val_t __engram_age_field_catchup(void) { return engram_age_field_catchup(); }
el_val_t __engram_chrono_persist_tick(void) { return engram_chrono_persist_tick(); }
el_val_t __engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction) { el_val_t __engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction) {
return engram_neighbors_json(node_id, max_depth, direction); return engram_neighbors_json(node_id, max_depth, direction);
} }
+31
View File
@@ -119,6 +119,37 @@ fn engram_activate_json(query: String, limit: Int) -> String {
return __engram_activate_json(query, limit) return __engram_activate_json(query, limit)
} }
// --- Geometry operators (§5) ---
// Relational-neighborhood algebra over centered engram descriptors. Each takes
// comma-separated seed-id set(s); a centered neighborhood is grown from those
// seeds (ad-hoc NOCACHE path) and the operator's JSON result is returned.
// Delegates to the native __engram_geo_*_json seeds (el_seed.c); in the heavy
// engram runtime the same bare names resolve directly to el_runtime.c symbols.
fn engram_geo_descriptor_json(seeds: String) -> String {
return __engram_geo_descriptor_json(seeds)
}
fn engram_geo_overlap_json(a_seeds: String, b_seeds: String) -> String {
return __engram_geo_overlap_json(a_seeds, b_seeds)
}
fn engram_geo_subtract_json(a_seeds: String, b_seeds: String, mode: String) -> String {
return __engram_geo_subtract_json(a_seeds, b_seeds, mode)
}
fn engram_geo_combine_json(a_seeds: String, b_seeds: String) -> String {
return __engram_geo_combine_json(a_seeds, b_seeds)
}
fn engram_geo_distance_json(a_seeds: String, b_seeds: String) -> String {
return __engram_geo_distance_json(a_seeds, b_seeds)
}
fn engram_geo_analogy_json(a_seeds: String, b_seeds: String) -> String {
return __engram_geo_analogy_json(a_seeds, b_seeds)
}
// --- Generation --- // --- Generation ---
fn generate(form: String) -> String { fn generate(form: String) -> String {
File diff suppressed because it is too large Load Diff
+386
View File
@@ -0,0 +1,386 @@
/* engram_geometry.h — M9 FOUNDATION: the relational-neighborhood GEOMETRY
* DESCRIPTOR (design doc §3, §5; memory node e94371bd).
*
* Computes, for a relational neighborhood grown from a seed set, the compact
* (KB-not-MB) joint geometry Will specified: the SEMANTIC geometry (centroid,
* covariance / principal axes, radius) braided with the RELATIONAL geometry
* (k-core skeleton, hub->periphery centrality gradient), plus soft membership.
*
* Two coordinate systems, one shape "a constellation: bright prototype at the
* center, a cloud of members at varying distance, the strongest edges as a
* backbone, fading at the edges."
*
* Built ON the two standalone M-era modules only:
* - engram_vindex : semantic neighbors (the cloud) via ANN.
* - engram_store : node embeddings + hebb adjacency (the skeleton), read-only.
* It does NOT link or touch el_runtime.c, and it is a pure READ over the graph:
* it never modifies nodes, edges, activation, the index, or any retrieval path.
*
* Pure C11, stdlib + libm only. The descriptor is a foundation object; it is NOT
* wired into retrieval/priming yet (that is the next M9 step).
*/
#ifndef ENGRAM_GEOMETRY_H
#define ENGRAM_GEOMETRY_H
#include <stddef.h>
#include <stdint.h>
#include "engram_store.h"
#include "engram_vindex.h"
/* One member of the neighborhood + its place in the gradient. */
typedef struct {
char* id;
double membership; /* soft membership in [0,1] (semantic+relational blend) */
double centrality; /* skeleton weighted-degree — relational salience */
double salience; /* the node's own stored salience */
int core; /* k-core number (0 = fringe / not in any core) */
double dist_centroid; /* cosine distance of member emb to centroid (semantic)*/
int embedded; /* 1 if the member carried an emb vector */
} GeoMember;
/* One skeleton edge (indices into members[]). eff_weight = weight*(1+0.5*hebb),
* clamped to 1.0 the effective propagation strength eg_edge_eff_weight uses. */
typedef struct { uint32_t a, b; double eff_weight; double hebb; } GeoEdge;
/* A compact principal axis of the ellipsoid: unit direction in R^dim + extent
* (sqrt of the covariance eigenvalue = the ellipsoid's half-width along it). */
typedef struct { float* axis; double extent; } GeoAxis;
typedef struct {
int dim;
/* ── anchor ── */
char* hub_id; /* highest-centrality member: the relational hub */
float* centroid; /* v̄ ∈ R^dim: mean of the member embeddings in the
* frame the descriptor operated in. When centered
* (global_mean != NULL) this is the CENTERED
* centroid (mean of L2-normalized embs minus the
* global mean): the neighborhood's location in the
* isotropic/whitened frame. Add global_mean back to
* recover the raw prototype point. When uncentered
* it is the raw mean of L2-normalized member embs. */
float* global_mean; /* the centering offset actually applied (dim floats),
* or NULL if the descriptor ran in raw space. The §5
* operators (distance/overlap/Wasserstein) are only
* discriminative in the centered frame see notes. */
/* ── shape (compact covariance): top principal axes + extents ── */
int n_axes;
GeoAxis* axes; /* orientation + extents of the ellipsoid */
double total_variance; /* trace(Σ) = mean squared member dist to centroid*/
/* ── scale ── */
double radius; /* sqrt(total_variance) — the neighborhood breadth*/
/* ── members + gradient ── */
int n_members;
GeoMember* members; /* soft membership {id->weight} + centrality/salience */
/* ── skeleton ── */
int n_edges;
GeoEdge* edges; /* strong internal hebb edges = the backbone */
int k_core; /* the maximum core number present in the skeleton*/
/* ── diagnostics ── */
double co_registration;/* corr(hebb strength, semantic proximity) over */
/* internal edges: >0 = geometries agree (reify); */
/* <0 = disagree (surprising links / dream cands). */
int n_embedded; /* members that carried an emb vector */
} GeoDescriptor;
typedef struct {
int ann_k; /* semantic expansion: ANN neighbors per seed (0=off) */
int hop_relational; /* 1 = include seeds' hebb neighbors as members */
double edge_min_weight; /* skeleton: ignore internal edges below this eff wt */
int kcore_k; /* target k for the reported k-core (0 = auto/max) */
int top_axes; /* principal axes to retain (default 8) */
int max_members; /* cap neighborhood size (guards the eigensolve cost) */
} GeoParams;
/* Fill p with sane defaults: ann_k=24, hop_relational=1, edge_min_weight=0.05,
* kcore_k=0 (auto), top_axes=8, max_members=400. */
void engram_geo_default_params(GeoParams* p);
/* ── Global-mean cache (mean-centering / whitening the anisotropic emb space) ──
* The nomic-embed-text space over the engram corpus is strongly ANISOTROPIC:
* every embedding sits in a narrow cone (mean pairwise cosine ~0.55), which
* compresses cosine-based domain separation almost to nothing. Subtracting the
* GLOBAL MEAN of the (L2-normalized) embeddings recenters the cloud on the
* origin (mean pairwise cosine -> ~0), restoring isotropy so the §5 operators
* discriminate. The mean is a store-level derived quantity, like the ANN index:
* built once from the paged store, cached, and refreshed when the embedded set
* drifts. It lives here (not in the store) so this stays a contained, read-only
* addition; a runtime owns one GeoMeanCache per open store alongside its VIndex. */
typedef struct GeoMeanCache GeoMeanCache;
/* Scan every live node in `store` and compute the mean of the L2-normalized
* embeddings over the embed-eligible set (nodes carrying an emb vector; the
* unembedded telemetry/system nodes are skipped). Returns a malloc'd cache, or
* NULL on error / no embedded nodes. The offset vector is NOT renormalized it
* is a translation, applied by subtraction. */
GeoMeanCache* engram_geo_mean_build(EngramPagedStore* store);
/* The cached offset (dim floats) — pass to engram_geometry_descriptor as
* global_mean. Valid until the cache is freed/refreshed. */
const float* engram_geo_mean_vec(const GeoMeanCache* c);
int engram_geo_mean_dim(const GeoMeanCache* c);
uint64_t engram_geo_mean_count(const GeoMeanCache* c); /* #embedded nodes used */
/* Recompute the mean IN PLACE iff the embedded-node count has drifted by more
* than `frac` (e.g. 0.10 = 10%) since the cache was built "recompute on
* significant change". Returns 1 if it rebuilt, 0 if unchanged, <0 on error. */
int engram_geo_mean_maybe_refresh(GeoMeanCache* c, EngramPagedStore* store,
double frac);
void engram_geo_mean_free(GeoMeanCache* c);
/* Compute the geometry descriptor of the neighborhood grown from seed_ids.
* READ-ONLY over store + vindex.
* store an opened store (borrowed; not modified).
* vindex optional ANN index for semantic expansion; NULL disables it.
* vids the ordinal->store-id map returned by vindex_build_from_store
* (vids[node_id] == store id). Required iff vindex != NULL.
* n_vids length of vids.
* params NULL to use engram_geo_default_params.
* global_mean optional centering offset (dim floats, from engram_geo_mean_*).
* When non-NULL the SEMANTIC geometry is computed in mean-centered
* (isotropic) space: every normalized member emb has global_mean
* subtracted before the centroid / cosine-distance / co-registration
* math, so those operators discriminate. NULL = raw space (legacy).
* NOTE: the ANN neighbor query still runs in RAW unit-vector space
* centering is a rigid translation that ~preserves neighborhood
* MEMBERSHIP, so the index needs no rebuild; only the descriptor
* STATISTICS move to the centered frame (co-registration choice (b)).
* The eigen/covariance shape (axes, radius) is translation-invariant
* and therefore identical in either frame.
* Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error
* (no seeds resolvable, OOM). */
GeoDescriptor* engram_geometry_descriptor(
EngramPagedStore* store, VIndex* vindex,
char** vids, int n_vids,
const char* const* seed_ids, size_t n_seeds,
const GeoParams* params,
const float* global_mean);
void engram_geo_free(GeoDescriptor* g);
/* ── M-INTEROCEPTION P3: drift-sensor primitive (descriptor displacement) ────
* Read-only. GROWTH vs CORRUPTION split of how far B drifted from baseline A.
* See engram_geometry.c for the honesty note on the missing SelfAnchor. */
typedef struct {
double centroid_sep; /* L2 distance between centroids (same frame) */
double centroid_cos; /* 1 - cosine(centroidA, centroidB) */
double radius_delta; /* |radiusA - radiusB| — neighborhood scale change */
double core_disp; /* mean radial displacement of the invariant core */
double periph_disp; /* mean radial displacement of the periphery */
int core_matched; /* # core members matched by id across A,B */
int periph_matched; /* # periphery members matched by id across A,B */
} GeoDisplacement;
void engram_geo_displacement(const GeoDescriptor* a, const GeoDescriptor* b,
double core_frac, GeoDisplacement* out);
/* ═══════════════════════════════════════════════════════════════════════════
* §5 GEOMETRY OPERATORS a relational ALGEBRA over neighborhood descriptors.
* These are the reusable primitives Will specified: "primitives any CGI
* application should be able to use." READ-ONLY and PURE (stdlib + libm only) —
* they consume GeoDescriptor(s) and never touch the store, index, or activation.
*
* FRAME CONTRACT: both inputs MUST have been built in the SAME frame identical
* emb `dim` and identical `global_mean` (centered against the one true store-wide
* mean). The reify path builds every neighborhood that way, so descriptors are
* directly comparable. An operator returns <0 / NULL if the dims disagree.
*
* REPRESENTATION: the C descriptor lives in the FULL emb dim with a LOW-RANK
* covariance Σ = Σ_k extent_k² · a_k a_kᵀ over its retained principal axes
* (top_axes; the discarded tail variance is not modeled). Every operator mirrors
* the viz-proxy (engram-geometry-proxy.py §5) FORMULA exactly, but evaluates it on
* this representation so semantics match the proxy while absolute numbers differ
* (proxy works in a 24-dim global-PCA reduced dense frame; C in full-dim low-rank).
* The Wasserstein / combine eigen-work is done inside the small JOINT axis subspace
* (dimension nA+nB+1), which is EXACT for the low-rank covariances there.
* Each result struct is released by its engram_geo_*_free.
* */
/* overlap(A,B): shared-member set + Jaccard + centroid/scale proximity score. */
typedef struct {
char** shared_ids; /* ids present in BOTH neighborhoods (owned) */
int n_shared;
int n_union; /* |A B| by id */
double jaccard; /* |A∩B| / |AB| */
double centroid_distance; /* L2 between the (centered) centroids */
double overlap_score; /* jacc*0.5 + max(0,1d/(rA+rB))*0.5 (proxy form)*/
float* intersection_centroid; /* midpoint of the two centroids (dim, owned) */
int dim;
} GeoOverlap;
int engram_geo_overlap(const GeoDescriptor* a, const GeoDescriptor* b, GeoOverlap* out);
void engram_geo_overlap_free(GeoOverlap* o);
/* subtract(A,B) — ORTHOGONAL-COMPLEMENT residual: project A onto I V_B V_Bᵀ
* (V_B = B's top `b_dims` principal axes) "A with B's framing removed". Returns
* A's residual centroid + residual ellipsoid, the fraction of A's energy that lives
* inside B's subspace, and the centroid-difference vector. b_dims<=0 min(3,nB). */
typedef struct {
int dim;
float* residual_centroid; /* P⊥ c_A (owned) */
float* centroid_diff; /* c_A c_B (owned) */
double centroid_diff_mag;
double variance_explained_by_B; /* (‖Qc_A‖²+Tr(QΣ_A)) / (‖c_A‖²+Tr Σ_A) ∈[0,1]*/
int removed_dims; /* # of B axes used as V_B */
double residual_scale; /* sqrt(Tr(P⊥ Σ_A P⊥)) */
int n_axes; /* residual principal axes (owned) */
GeoAxis* axes;
} GeoResidual;
int engram_geo_subtract(const GeoDescriptor* a, const GeoDescriptor* b,
int b_dims, GeoResidual* out);
void engram_geo_residual_free(GeoResidual* r);
/* set-diff variant of subtract: members in A but not in B + the centroid arrow. */
typedef struct {
char** only_ids; /* member ids in A and not in B (owned) */
int n_only;
int removed; /* |A ∩ B| (dropped) */
float* centroid_diff; /* c_A c_B (dim, owned) */
double centroid_diff_mag;
int dim;
} GeoSetDiff;
int engram_geo_setdiff(const GeoDescriptor* a, const GeoDescriptor* b, GeoSetDiff* out);
void engram_geo_setdiff_free(GeoSetDiff* s);
/* combine(A,B): a merged descriptor — POOLED centroid + POOLED covariance
* (exact law-of-total-variance: the covariance you'd get by concatenating the two
* member clouds), re-eigendecomposed for its principal axes. Members = id-union
* (membership = max). top_axes<=0 8. Returns a malloc'd GeoDescriptor (free with
* engram_geo_free) in the same frame as A, or NULL on error. */
GeoDescriptor* engram_geo_combine(const GeoDescriptor* a, const GeoDescriptor* b,
int top_axes);
/* distance(A,B): centroid L2 + centroid cosine + closed-form Wasserstein-2
* (Bures metric) between the two Gaussians mirrors the proxy's _wasserstein2. */
typedef struct {
double centroid_distance;
double centroid_cosine;
double wasserstein2;
int dim;
} GeoDistance;
int engram_geo_distance(const GeoDescriptor* a, const GeoDescriptor* b, GeoDistance* out);
/* analogy(A,B): orthogonal PROCRUSTES transform min_R ‖A B R‖_F, RᵀR=I (SVD)
* aligning A's principal frame to B's (extent-scaled axes, paired by rank). R is
* returned COMPACTLY as an r×r rotation within the joint axis subspace `basis`
* (r vectors of dim floats); it acts as the identity on the orthogonal complement.
* Apply it to a vector with engram_geo_analogy_apply. */
typedef struct {
int dim;
int r; /* subspace rank; R is r×r */
float* basis; /* r×dim row-major orthonormal basis Q (owned) */
double* R; /* r×r rotation in Q-coords, row-major (owned) */
double residual; /* ‖A B R‖_F over the extent-scaled frames */
} GeoAnalogy;
int engram_geo_analogy(const GeoDescriptor* a, const GeoDescriptor* b, GeoAnalogy* out);
/* out_vec = R·v for v ∈ R^dim: v + Σ_i (R̂c c)_i q_i, c_i = q_i·v. dim floats. */
void engram_geo_analogy_apply(const GeoAnalogy* an, const float* v, float* out_vec);
void engram_geo_analogy_free(GeoAnalogy* an);
/* ═══════════════════════════════════════════════════════════════════════════
* M10 REIFICATION: densely co-wired relational neighborhoods crystallized into
* FIRST-CLASS, PERSISTED store records (design doc §2; memory 885f5945). This is
* NOT a cache it is durable structure. A reified neighborhood is a real store
* NODE (node_type "Neighborhood") that survives restart, is loaded on boot, and
* EVOLVES via supersede+provenance when the pattern shifts. The geometry-priming
* HOT PATH reads these persisted records (never computes geometry on the
* activation path). Ad-hoc/transient geometries still use the on-the-fly
* engram_geometry_descriptor above.
*
* Two record types, both ordinary TLV store nodes (no new on-disk format):
* - "GeoMeanFrame" : the store-wide centering mean, persisted ONCE (emb = mean
* vector, id ENGRAM_GEO_MEANFRAME_ID). Referenced by every
* neighborhood so priming centers against the SAME true mean.
* - "Neighborhood" : one reified neighborhood. emb = the RAW centroid (prototype
* point, so it stays centroid-ANN-able; centered_centroid =
* emb - meanframe). metadata = the compact "GEO1" schema:
* hub id, meanframe ref, scalar shape (radius, total_variance,
* k_core, co_registration, n_embedded), axis EXTENTS (ellipsoid
* half-widths), and the MEMBER list {id -> membership, centrality,
* core}. Member links are also persisted as edges relation="member".
*
* v1 honest simplifications (documented; extensible without migration): axis
* DIRECTION vectors are not persisted (extents capture the ellipsoid scale; the
* directions are recomputable via the on-the-fly descriptor for viz/operators);
* with hebb potentiation ~0 on today's store the "hebb-weighted" degree reduces to
* AUTHORED edge weight, so detected neighborhoods currently reflect authored edges
* the design is unchanged and self-correcting once hebb accrues.
* */
#define ENGRAM_GEO_NBHD_TYPE "Neighborhood"
#define ENGRAM_GEO_MEANFRAME_TYPE "GeoMeanFrame"
#define ENGRAM_GEO_MEANFRAME_ID "geo-meanframe" /* stable id of the singleton */
#define ENGRAM_GEO_NBHD_ID_PREFIX "nbhd-" /* id = nbhd-<hub>-<built_at> */
#define ENGRAM_GEO_MEMBER_RELATION "member"
typedef struct {
int min_weighted_degree; /* hub qualifies iff strong-edge weighted degree >= this
* (0 = no floor: just rank + take top max_neighborhoods) */
int max_neighborhoods; /* homeostatic budget cap (default 128) */
double cover_membership; /* skip a hub already a member (w>=this) of an accepted
* neighborhood greedy non-redundant cover (default 0.5) */
int persist_member_edges; /* 1 = also write relation="member" edges (default 1) */
GeoParams descriptor; /* per-neighborhood params (top_axes may be 0 = skip eigensolve) */
} GeoReifyParams;
/* Defaults: min_weighted_degree=0, max_neighborhoods=128, cover_membership=0.5,
* persist_member_edges=1, descriptor = engram_geo_default_params but top_axes=4,
* max_members=256 (reified neighborhoods stay compact). */
void engram_geo_reify_default_params(GeoReifyParams* p);
/* WRITE PATH (offline / consolidation — NEVER the activation hot path).
* Detect dense hub neighborhoods on the hebb-weighted graph, compute each one's
* CENTERED descriptor ONCE against the true store-wide mean, and PERSIST them as
* first-class records: the GeoMeanFrame (once) + one Neighborhood node per detected
* neighborhood (+ member edges), superseding any prior same-hub record with
* provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0.
* Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex,
char** vids, int n_vids,
const GeoReifyParams* params);
/* ── Resident loaded form of the persisted records (boot-time; READ-ONLY) ─────
* The durable Neighborhood/GeoMeanFrame records are the source of truth; this
* index is their LOADED form (like the resident node array is the loaded form of
* the node records, or adjacency the loaded form of edges). It never recomputes
* geometry it parses. Build it by feeding the runtime's boot node scan, or in
* one pass with engram_geo_reify_load. */
typedef struct GeoReifyIndex GeoReifyIndex;
GeoReifyIndex* engram_geo_reify_index_new(void);
/* Feed one store node; if it is a Neighborhood or GeoMeanFrame record it is parsed
* and absorbed (else ignored). The node is BORROWED (copied as needed). 0/<0. */
int engram_geo_reify_index_add(GeoReifyIndex* ix, const StoreNode* n);
/* Build the member->neighborhood hash after all adds. Call once. 0/<0. */
int engram_geo_reify_index_finalize(GeoReifyIndex* ix);
/* One-pass convenience: scan the store and build the finalized index. NULL if the
* store holds no reified records. */
GeoReifyIndex* engram_geo_reify_load(EngramPagedStore* store);
/* A borrowed view of one persisted neighborhood (owned by the index). */
typedef struct {
const char* id;
const char* hub_id;
int n_members;
char* const* member_ids; /* parallel arrays, length n_members */
const double* member_w; /* membership in [0,1] */
double radius;
double co_registration;
int k_core;
int n_embedded;
} GeoNeighborhood;
/* HOT-PATH LOOKUP (no geometry compute): resolve the seed set to the best
* persisted neighborhood the one with the greatest summed seed membership; on a
* miss (no seed is a member of any neighborhood) fall back to the centroid nearest
* the query embedding (centered by the loaded mean frame). q_emb may be NULL (then
* a miss returns NULL). Returns a BORROWED handle (do NOT free) or NULL. */
const GeoNeighborhood* engram_geo_reify_lookup(
const GeoReifyIndex* ix,
const char* const* seed_ids, size_t n_seeds,
const float* q_emb, int q_dim);
int engram_geo_reify_count(const GeoReifyIndex* ix);
const float* engram_geo_reify_mean(const GeoReifyIndex* ix, int* dim); /* loaded true mean or NULL */
void engram_geo_reify_index_free(GeoReifyIndex* ix);
#endif /* ENGRAM_GEOMETRY_H */
+287
View File
@@ -0,0 +1,287 @@
/* engram_reason.c — the REASONING layer. Pure compositions over engram_geometry.h.
* stdlib + libm only; READ-ONLY over its descriptor inputs; touches no store/index. */
#include "engram_reason.h"
#include <stdlib.h>
#include <string.h>
#include <math.h>
/* ── small float-vector helpers ─────────────────────────────────────────────── */
static double vdot(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) s += (double)a[i] * (double)b[i]; return s;
}
static double vnorm(const float* a, int dim) { return sqrt(vdot(a, a, dim)); }
static double vcos(const float* a, const float* b, int dim) {
double na = vnorm(a, dim), nb = vnorm(b, dim);
if (na < 1e-12 || nb < 1e-12) return 0.0; /* a null vector ⇒ no direction */
double c = vdot(a, b, dim) / (na * nb);
if (c > 1.0) c = 1.0; if (c < -1.0) c = -1.0;
return c;
}
static double l2(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) { double d = (double)a[i] - (double)b[i]; s += d * d; }
return sqrt(s);
}
/* ═══════════════════════════════════════════ SHARED — point-to-manifold FIT ══ */
int engram_reason_point_fit(const GeoDescriptor* g, const float* x,
double ext_floor, GeoFit* out) {
if (!g || !x || !out || g->dim <= 0 || !g->centroid) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
int dim = g->dim;
/* residual r = x centroid */
double rr = 0; /* ‖r‖² */
float* r = malloc((size_t)dim * sizeof(float));
if (!r) return -1;
for (int i = 0; i < dim; i++) { double d = (double)x[i] - (double)g->centroid[i]; r[i] = (float)d; rr += d * d; }
double maha2 = 0, ss_in = 0; /* Mahalanobis² and in-subspace energy */
for (int k = 0; k < g->n_axes; k++) {
const float* ax = g->axes[k].axis; if (!ax) continue;
double proj = vdot(r, ax, dim); /* axes are orthonormal directions */
double den = g->axes[k].extent; if (den < ext_floor) den = ext_floor;
maha2 += (proj / den) * (proj / den);
ss_in += proj * proj;
}
double ortho2 = rr - ss_in; if (ortho2 < 0) ortho2 = 0; /* off-subspace energy */
double dist2 = maha2 + ortho2 / (ext_floor * ext_floor);
out->mahalanobis = sqrt(maha2);
out->ortho_residual = sqrt(ortho2);
out->distance = sqrt(dist2);
out->score = 1.0 / (1.0 + dist2);
free(r);
return 0;
}
/* ═══════════════════════════════════════════════════════════════ ANALOGY ════ */
int engram_reason_analogy(const GeoDescriptor* A, const GeoDescriptor* B,
const GeoDescriptor* C,
const GeoDescriptor* const* candidates, int n_candidates,
GeoAnalogyResult* out) {
if (!A || !B || !C || !out) return -1;
if (!A->centroid || !B->centroid || !C->centroid) return -1;
int dim = A->dim;
if (B->dim != dim || C->dim != dim) return -1;
memset(out, 0, sizeof *out);
out->dim = dim; out->best = -1;
/* Learn R_{A→B}. engram_geo_analogy(X,Y) yields R with apply(R, Y-axis) ≈ X-axis
* (R maps Y's frame X's frame); so R that maps AB is engram_geo_analogy(B,A). */
GeoAnalogy an;
if (engram_geo_analogy(B, A, &an) != 0) return -1;
out->analogy_residual = an.residual;
/* mapped = R·c_C + (c_B R·c_A) : the A→B affine (rotation + residual shift). */
float* RcA = malloc((size_t)dim * sizeof(float));
float* RcC = malloc((size_t)dim * sizeof(float));
out->mapped_point = malloc((size_t)dim * sizeof(float));
if (!RcA || !RcC || !out->mapped_point) { free(RcA); free(RcC); free(out->mapped_point); out->mapped_point = NULL; engram_geo_analogy_free(&an); return -1; }
engram_geo_analogy_apply(&an, A->centroid, RcA);
engram_geo_analogy_apply(&an, C->centroid, RcC);
for (int i = 0; i < dim; i++)
out->mapped_point[i] = (float)((double)RcC[i] + ((double)B->centroid[i] - (double)RcA[i]));
free(RcA); free(RcC);
engram_geo_analogy_free(&an);
/* nearest candidate to the mapped point (centroid L2). */
if (candidates && n_candidates > 0) {
out->n_candidates = n_candidates;
out->distances = malloc((size_t)n_candidates * sizeof(double));
if (!out->distances) return -1;
double best = -1; int bi = -1;
for (int i = 0; i < n_candidates; i++) {
const GeoDescriptor* cd = candidates[i];
double d = (cd && cd->centroid && cd->dim == dim) ? l2(out->mapped_point, cd->centroid, dim) : INFINITY;
out->distances[i] = d;
if (bi < 0 || d < best) { best = d; bi = i; }
}
out->best = bi; out->best_distance = best;
}
return 0;
}
void engram_reason_analogy_free(GeoAnalogyResult* r) {
if (!r) return;
free(r->mapped_point); free(r->distances);
r->mapped_point = NULL; r->distances = NULL;
}
/* ═══════════════════════════════════════════════════════════════ INDUCTION ══ */
int engram_reason_induce(const GeoDescriptor* const* examples, int n_examples,
int top_axes, double ext_floor, GeoInduction* out) {
if (!examples || n_examples < 1 || !out) return -1;
if (top_axes <= 0) top_axes = 8;
memset(out, 0, sizeof *out);
/* fold the examples left→right through the pooled-Gaussian combine. n==1 pools
* the single example with itself (identical cov same shape, id-union = itself). */
GeoDescriptor* acc = engram_geo_combine(examples[0],
examples[n_examples > 1 ? 1 : 0], top_axes);
if (!acc) return -1;
for (int i = 2; i < n_examples; i++) {
GeoDescriptor* nxt = engram_geo_combine(acc, examples[i], top_axes);
engram_geo_free(acc);
if (!nxt) return -1;
acc = nxt;
}
out->rule = acc;
out->n_examples = n_examples;
out->ext_floor = (ext_floor > 0) ? ext_floor
: (acc->radius > 0 ? acc->radius * 0.25 : 1.0);
return 0;
}
double engram_reason_membership(const GeoInduction* ind, const float* x) {
if (!ind || !ind->rule || !x) return -1;
GeoFit f;
if (engram_reason_point_fit(ind->rule, x, ind->ext_floor, &f) != 0) return -1;
return f.score;
}
void engram_reason_induction_free(GeoInduction* out) {
if (!out) return;
if (out->rule) engram_geo_free(out->rule);
out->rule = NULL;
}
/* ═══════════════════════════════════════════════════════════════ ABDUCTION ══ */
int engram_reason_abduce(const float* obs, int dim,
const GeoDescriptor* const* hypotheses, int n,
double ext_floor, GeoAbduction* out) {
if (!obs || !hypotheses || n < 1 || dim <= 0 || !out) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
memset(out, 0, sizeof *out);
out->n = n; out->best = -1;
out->scores = malloc((size_t)n * sizeof(double));
out->distances = malloc((size_t)n * sizeof(double));
out->rank = malloc((size_t)n * sizeof(int));
if (!out->scores || !out->distances || !out->rank) { engram_reason_abduction_free(out); return -1; }
double best = -1; int bi = -1;
for (int i = 0; i < n; i++) {
out->rank[i] = i;
const GeoDescriptor* h = hypotheses[i];
GeoFit f;
if (!h || h->dim != dim || engram_reason_point_fit(h, obs, ext_floor, &f) != 0) {
out->scores[i] = 0.0; out->distances[i] = INFINITY;
} else {
out->scores[i] = f.score; out->distances[i] = f.distance;
}
if (bi < 0 || out->scores[i] > best) { best = out->scores[i]; bi = i; }
}
out->best = bi; out->best_score = (bi >= 0) ? out->scores[bi] : 0.0;
/* rank indices best→worst by score (insertion sort — n is small). */
for (int i = 1; i < n; i++) {
int key = out->rank[i]; int j = i - 1;
while (j >= 0 && out->scores[out->rank[j]] < out->scores[key]) { out->rank[j + 1] = out->rank[j]; j--; }
out->rank[j + 1] = key;
}
return 0;
}
void engram_reason_abduction_free(GeoAbduction* out) {
if (!out) return;
free(out->scores); free(out->distances); free(out->rank);
out->scores = NULL; out->distances = NULL; out->rank = NULL;
}
/* ═══════════════════════════════════════════════════════════════════ CAUSAL ══ */
/* |cos| of two descriptors' centroids after removing confounder Z's subspace. */
static double controlled_assoc(const GeoDescriptor* x, const GeoDescriptor* y,
const GeoDescriptor* z) {
GeoResidual rx, ry; double c = 0;
int ox = engram_geo_subtract(x, z, 0, &rx);
int oy = engram_geo_subtract(y, z, 0, &ry);
if (ox == 0 && oy == 0 && rx.residual_centroid && ry.residual_centroid)
c = fabs(vcos(rx.residual_centroid, ry.residual_centroid, x->dim));
if (ox == 0) engram_geo_residual_free(&rx);
if (oy == 0) engram_geo_residual_free(&ry);
return c;
}
int engram_reason_causal(const GeoDescriptor* x, const GeoDescriptor* y,
const GeoDescriptor* const* confounders, int n_conf,
int64_t t_x, int64_t t_y,
double drop_frac, GeoCausal* out) {
if (!x || !y || !out || !x->centroid || !y->centroid || x->dim != y->dim) return -1;
if (!(drop_frac > 0 && drop_frac < 1)) drop_frac = 0.5;
memset(out, 0, sizeof *out);
const double assoc_floor = 0.2; /* below this = no meaningful association */
out->assoc_raw = fabs(vcos(x->centroid, y->centroid, x->dim));
/* control for each confounder; the strongest single explainer wins (min assoc). */
double ctrl = out->assoc_raw;
for (int i = 0; i < n_conf; i++) {
if (!confounders[i]) continue;
double c = controlled_assoc(x, y, confounders[i]);
if (c < ctrl) ctrl = c;
}
out->assoc_controlled = ctrl;
out->temporal_dir = (t_x < t_y) ? 1 : (t_x > t_y) ? -1 : 0;
if (out->assoc_raw < assoc_floor) {
out->verdict = GEO_CAUSAL_NONE;
} else if (ctrl < (1.0 - drop_frac) * out->assoc_raw && ctrl < assoc_floor) {
out->verdict = GEO_CAUSAL_CONFOUNDED; out->confounded = 1;
} else if (out->temporal_dir != 0) {
out->verdict = GEO_CAUSAL_DIRECTED; out->strength = ctrl;
} else {
out->verdict = GEO_CAUSAL_NONE; /* associated + robust but unorientable */
}
return 0;
}
/* ═══════════════════════════════════════════════════════════════════ PLANNING ══ */
int engram_reason_plan(const GeoDescriptor* const* nodes, int n,
int start, int goal, double neighbor_radius,
int use_wasserstein, GeoPlan* out) {
if (!nodes || n < 1 || !out) return -1;
if (start < 0 || start >= n || goal < 0 || goal >= n) return -1;
if (!(neighbor_radius > 0)) return -1;
memset(out, 0, sizeof *out);
/* dense edge weights (i<j symmetric); INFINITY = not adjacent. */
double* W = malloc((size_t)n * (size_t)n * sizeof(double));
if (!W) return -1;
for (int i = 0; i < n; i++) for (int j = 0; j < n; j++) W[(size_t)i * n + j] = (i == j) ? 0.0 : INFINITY;
for (int i = 0; i < n; i++) {
for (int j = i + 1; j < n; j++) {
GeoDistance d;
if (nodes[i] && nodes[j] && engram_geo_distance(nodes[i], nodes[j], &d) == 0) {
double w = use_wasserstein ? d.wasserstein2 : d.centroid_distance;
if (w <= neighbor_radius) { W[(size_t)i * n + j] = w; W[(size_t)j * n + i] = w; }
}
}
}
/* O(n²) Dijkstra. */
double* dist = malloc((size_t)n * sizeof(double));
int* prev = malloc((size_t)n * sizeof(int));
char* done = calloc((size_t)n, 1);
if (!dist || !prev || !done) { free(W); free(dist); free(prev); free(done); return -1; }
for (int i = 0; i < n; i++) { dist[i] = INFINITY; prev[i] = -1; }
dist[start] = 0;
for (int it = 0; it < n; it++) {
int u = -1; double bd = INFINITY;
for (int i = 0; i < n; i++) if (!done[i] && dist[i] < bd) { bd = dist[i]; u = i; }
if (u < 0) break;
done[u] = 1;
if (u == goal) break;
for (int v = 0; v < n; v++) {
double w = W[(size_t)u * n + v];
if (w < INFINITY && !done[v] && dist[u] + w < dist[v]) { dist[v] = dist[u] + w; prev[v] = u; }
}
}
if (dist[goal] < INFINITY) {
int len = 0; for (int v = goal; v != -1; v = prev[v]) len++;
out->path = malloc((size_t)len * sizeof(int));
if (out->path) {
out->path_len = len;
int idx = len - 1;
for (int v = goal; v != -1; v = prev[v]) out->path[idx--] = v;
out->total_cost = dist[goal];
out->reached = 1;
}
}
free(W); free(dist); free(prev); free(done);
return 0;
}
void engram_reason_plan_free(GeoPlan* out) {
if (!out) return;
free(out->path); out->path = NULL;
}
+161
View File
@@ -0,0 +1,161 @@
/* engram_reason.h — the REASONING layer: compositions over the §5 geometry
* OPERATORS (engram_geometry.h). Where the operators are a relational ALGEBRA over
* neighborhood descriptors, these are reasoning MODES built by CHAINING that algebra:
*
* ANALOGY A:B :: C:? learn the AB transform (Procrustes), apply to C.
* INDUCTION {E_i} rule pool example geometries; a generalizing structure
* + a membership test.
* ABDUCTION x best H the structure whose geometry best PLACES an
* observation in-distribution (inverse of prediction).
* CAUSAL x ? y | Z, t separate mere overlap (correlation) from directed
* influence (temporal precedence + association that
* SURVIVES controlling for confounders via subtract).
* PLANNING start goal a trajectory (sequence of neighborhoods) through the
* manifold: shortest path over geo-distance edges.
*
* PURE + READ-ONLY (stdlib + libm only): every function consumes GeoDescriptor(s)
* (+ a few scalars / timestamps) and NEVER touches the store, index, or activation.
* All geometry is delegated to the engram_geo_* primitives; this file only composes.
*
* FRAME CONTRACT (inherited): descriptors passed together MUST share emb `dim` and
* `global_mean` frame exactly the §5 operator contract. A function returns <0 on
* a dim/frame mismatch or bad argument.
*/
#ifndef ENGRAM_REASON_H
#define ENGRAM_REASON_H
#include <stdint.h>
#include "engram_geometry.h"
/* ═══════════════════════════════════════════════════════════════════════════
* SHARED PRIMITIVE point-to-manifold FIT. How well does a single point x sit
* inside a neighborhood's ellipsoid? Splits the residual (x centroid) into:
* - the IN-SUBSPACE part, scaled by each axis extent a Mahalanobis distance
* (how many "radii" out along the modeled directions), and
* - the ORTHOGONAL part outside the retained axes energy the model does not
* explain at all (charged at the extent floor).
* This is the common engine under INDUCTION's membership test and ABDUCTION's
* explanation ranking. ext_floor (>0) guards zero-extent axes / the null model.
* */
typedef struct {
double mahalanobis; /* sqrt( Σ_k ((a_k·(xc)) / max(ext_k,floor))² ) */
double ortho_residual; /* ‖(xc) projected off the retained axes‖ (raw L2) */
double distance; /* sqrt( maha² + (ortho_residual/floor)² ) — full fit */
double score; /* 1 / (1 + distance²) ∈ (0,1] (1 = dead-center) */
} GeoFit;
int engram_reason_point_fit(const GeoDescriptor* g, const float* x,
double ext_floor, GeoFit* out);
/* ═══════════════════════════════════════════════════════════════════════════
* ANALOGY "A:B :: C:?". Learn the transform that carries A to B (orthogonal
* Procrustes rotation R between their principal frames + the residual translation),
* apply it to C, and return the mapped point + the nearest candidate neighborhood.
* Composes: engram_geo_analogy (R) + engram_geo_analogy_apply + engram_geo_distance.
* */
typedef struct {
int dim;
float* mapped_point; /* predicted D location = R·c_C + (c_B R·c_A) (owned)*/
double analogy_residual;/* Procrustes ‖AB R‖_F — frame-alignment quality */
int best; /* index of nearest candidate to mapped_point, or 1 */
double best_distance; /* centroid L2 from mapped_point to the winner */
int n_candidates;
double* distances; /* centroid L2 mapped_point→candidate[i] (owned)*/
} GeoAnalogyResult;
/* candidates may be NULL/0 (then best=1 and only mapped_point is filled). */
int engram_reason_analogy(const GeoDescriptor* A, const GeoDescriptor* B,
const GeoDescriptor* C,
const GeoDescriptor* const* candidates, int n_candidates,
GeoAnalogyResult* out);
void engram_reason_analogy_free(GeoAnalogyResult* r);
/* ═══════════════════════════════════════════════════════════════════════════
* INDUCTION from a SET of example neighborhoods to the generalizing structure.
* Pools the examples (law-of-total-variance via engram_geo_combine, folded left to
* right) into a single "rule" descriptor whose top principal axes are the directions
* CONSISTENTLY present across the examples (the shared subspace surfaces as the
* dominant pooled axes; idiosyncratic per-example directions fall to the tail).
* The rule carries a membership test (point-to-manifold fit against the pool).
* */
typedef struct {
GeoDescriptor* rule; /* induced generalizing geometry (owned; geo_free) */
double ext_floor; /* extent floor used by the membership test */
int n_examples;/* how many examples were pooled */
} GeoInduction;
/* top_axes<=0 → 8. ext_floor<=0 → derived from the pooled radius. */
int engram_reason_induce(const GeoDescriptor* const* examples, int n_examples,
int top_axes, double ext_floor, GeoInduction* out);
/* Membership of a point in the induced rule ∈ (0,1] (the fit score). <0 on error. */
double engram_reason_membership(const GeoInduction* ind, const float* x);
void engram_reason_induction_free(GeoInduction* out);
/* ═══════════════════════════════════════════════════════════════════════════
* ABDUCTION inference to the best explanation. Given an observation POINT, rank a
* set of candidate structures by how well each PLACES the observation in-distribution
* (min point-to-manifold distance = the structure that, if assumed, best accounts for
* the observation). The inverse of prediction.
* */
typedef struct {
int best; /* index of best-explaining hypothesis, or 1 */
double best_score;
int n;
double* scores; /* fit score per hypothesis (higher = better) (owned)*/
double* distances; /* explanation distance per hypothesis (owned)*/
int* rank; /* hypothesis indices sorted best→worst (owned)*/
} GeoAbduction;
int engram_reason_abduce(const float* obs, int dim,
const GeoDescriptor* const* hypotheses, int n,
double ext_floor, GeoAbduction* out);
void engram_reason_abduction_free(GeoAbduction* out);
/* ═══════════════════════════════════════════════════════════════════════════
* CAUSAL correlation vs causation. Over two variables' geometries (+ candidate
* confounders + temporal order), distinguish:
* - mere co-occurrence / overlap (correlation), from
* - directed influence: association that (a) SURVIVES controlling for confounders
* (subtract each Z's subspace from both centroids, re-measure) and (b) is oriented
* by temporal PRECEDENCE.
* Composes: centroid cosine (correlation) + engram_geo_subtract (control) + timestamps.
* */
typedef enum {
GEO_CAUSAL_NONE = 0, /* no meaningful association */
GEO_CAUSAL_DIRECTED = 1, /* survives control + temporally ordered → cause→eff */
GEO_CAUSAL_CONFOUNDED = 2 /* correlated but association dies under control */
} GeoCausalVerdict;
typedef struct {
double assoc_raw; /* |cos(c_x,c_y)| — the raw correlation */
double assoc_controlled; /* |cos| of residual centroids after control */
int temporal_dir; /* +1 x→y, 1 y→x, 0 tie/unknown */
GeoCausalVerdict verdict;
int confounded; /* 1 iff verdict==CONFOUNDED (the flag) */
double strength; /* directed influence estimate ∈[0,1] (0 else)*/
} GeoCausal;
/* confounders may be NULL/0. t_x,t_y are comparable timestamps (any monotone unit);
* pass equal values for "unknown order". drop_frac(0,1): a controlled association
* below (1drop_frac)·assoc_raw AND below an absolute floor CONFOUNDED. */
int engram_reason_causal(const GeoDescriptor* x, const GeoDescriptor* y,
const GeoDescriptor* const* confounders, int n_conf,
int64_t t_x, int64_t t_y,
double drop_frac, GeoCausal* out);
/* ═══════════════════════════════════════════════════════════════════════════
* PLANNING trajectory construction. Given a set of neighborhoods (manifold nodes),
* a start and a goal, build a PATH (sequence of intermediate neighborhoods) by
* shortest path over the graph whose edges connect neighborhoods within
* neighbor_radius, weighted by geo-distance. Long straight jumps are not edges, so
* the path follows the manifold's curvature through intermediates (a discrete geodesic).
* Composes: engram_geo_distance (edge weights) + Dijkstra.
* */
typedef struct {
int* path; /* node indices start..goal (owned) */
int path_len;
double total_cost; /* summed centroid-distance edge weights along path */
int reached; /* 1 if goal reachable within neighbor_radius graph */
} GeoPlan;
/* neighbor_radius>0: max centroid distance for two neighborhoods to be adjacent.
* Use "wasserstein"!=0 to weight edges by Wasserstein-2 instead of centroid L2. */
int engram_reason_plan(const GeoDescriptor* const* nodes, int n,
int start, int goal, double neighbor_radius,
int use_wasserstein, GeoPlan* out);
void engram_reason_plan_free(GeoPlan* out);
#endif /* ENGRAM_REASON_H */
+588 -44
View File
@@ -141,15 +141,50 @@ struct EngramPagedStore {
int recovering; /* set during WAL replay */ int recovering; /* set during WAL replay */
uint64_t ops_since_ckpt; /* checkpoint threshold counter */ uint64_t ops_since_ckpt; /* checkpoint threshold counter */
uint64_t ckpt_threshold; /* auto-checkpoint after this many ops (0 = never) */ uint64_t ckpt_threshold; /* auto-checkpoint after this many ops (0 = never) */
/* ── M5 background-checkpointer triggers (0 = that trigger disabled) ────── */
size_t ckpt_dirty_threshold; /* auto-checkpoint at this many dirty frames */
uint64_t ckpt_wal_threshold; /* auto-checkpoint at this many WAL bytes */
long long ckpt_interval_ms; /* auto-checkpoint after this many ms elapse */
long long last_ckpt_ms; /* wall-clock time of the last checkpoint */
}; };
/* M2 buffer-pool hooks (defined in the M2 section at the bottom of this file). */ /* M2/M4 buffer-pool hooks (defined in the pool section at the bottom of this file).
typedef struct PgEnt { uint64_t id; uint8_t* buf; uint64_t lsn; int dirty; struct PgEnt* next; } PgEnt; * M2 shipped a write-back, no-steal cache (dirtydisk only at checkpoint). M4
* turns it into a bounded, demand-paged buffer pool: a fixed frame budget, LRU
* eviction of CLEAN unpinned frames (no-steal preserved dirty frames are never
* stolen), pinning of hot/structural pages, and bounded read-ahead. `lru_*`
* thread every resident frame onto an MRULRU list; `pin` is an explicit pin
* count (0 = unpinned). */
typedef struct PgEnt {
uint64_t id; uint8_t* buf; uint64_t lsn; int dirty; struct PgEnt* next;
int pin; /* explicit pin count (0 = unpinned) */
struct PgEnt* lru_prev; /* MRU→LRU doubly-linked list */
struct PgEnt* lru_next;
} PgEnt;
static PgCache* pc_new(void); static PgCache* pc_new(void);
static void pc_free(PgCache* c); static void pc_free(PgCache* c);
static PgEnt* pc_get(EngramPagedStore* s, uint64_t id); static PgEnt* pc_get(EngramPagedStore* s, uint64_t id);
static int pc_put(EngramPagedStore* s, uint64_t id, const uint8_t* buf, int dirty); static int pc_put(EngramPagedStore* s, uint64_t id, const uint8_t* buf, int dirty);
static int pc_flush(EngramPagedStore* s); /* pwrite all dirty → clean */ static int pc_flush(EngramPagedStore* s); /* pwrite all dirty → clean */
static void pc_prefetch(EngramPagedStore* s, uint64_t from_id, unsigned window);
static void store__autopin(EngramPagedStore* s); /* pin superblocks + index roots */
/* Per-layer pin record: the set of pages pinned on behalf of a hot layer, kept
* so store_unpin_layer can release exactly what store_pin_layer pinned. */
typedef struct { uint32_t layer; uint64_t* pages; size_t n; } LayerPin;
/* The bounded, demand-paged frame table (M4). Defined here (not in the pool
* section) so page_read / the scan loops can read its stats + prefetch window. */
struct PgCache {
PgEnt** buckets; size_t nbuckets; size_t count;
size_t cap; /* max resident frames; 0 = unlimited */
PgEnt* mru; PgEnt* lru; /* MRU (front) → LRU (back) recency list */
unsigned prefetch; /* read-ahead window (pages); 0 = off */
LayerPin* lp; size_t lp_n, lp_cap; /* hot-layer pin bookkeeping */
size_t dirty_count; /* # dirty frames, maintained incrementally (M5) */
/* stats (introspection only — never affect semantics) */
uint64_t hits, misses, evictions, prefetch_reads;
};
/* ── little-endian scalar codecs ──────────────────────────────────────────── */ /* ── little-endian scalar codecs ──────────────────────────────────────────── */
static void put_u16(uint8_t* p, uint16_t v){ p[0]=(uint8_t)v; p[1]=(uint8_t)(v>>8); } static void put_u16(uint8_t* p, uint16_t v){ p[0]=(uint8_t)v; p[1]=(uint8_t)(v>>8); }
@@ -197,12 +232,12 @@ static uint64_t id_hash(const char* s){
static int page_read(EngramPagedStore* s, uint64_t id, uint8_t* buf){ static int page_read(EngramPagedStore* s, uint64_t id, uint8_t* buf){
if (s->cache){ if (s->cache){
PgEnt* e = pc_get(s, id); PgEnt* e = pc_get(s, id);
if (e){ memcpy(buf, e->buf, STORE_PAGE_SIZE); return 0; } if (e){ memcpy(buf, e->buf, STORE_PAGE_SIZE); s->cache->hits++; return 0; }
} }
off_t off = (off_t)id * STORE_PAGE_SIZE; off_t off = (off_t)id * STORE_PAGE_SIZE; /* demand fault: not resident */
ssize_t r = pread(s->fd, buf, STORE_PAGE_SIZE, off); ssize_t r = pread(s->fd, buf, STORE_PAGE_SIZE, off);
if (r != (ssize_t)STORE_PAGE_SIZE) return -1; if (r != (ssize_t)STORE_PAGE_SIZE) return -1;
if (s->cache) pc_put(s, id, buf, 0); /* cache clean */ if (s->cache){ s->cache->misses++; pc_put(s, id, buf, 0); } /* cache clean */
return 0; return 0;
} }
static int page_write_raw(EngramPagedStore* s, uint64_t id, const uint8_t* buf){ static int page_write_raw(EngramPagedStore* s, uint64_t id, const uint8_t* buf){
@@ -539,8 +574,16 @@ static int leaf_max_entries(EngramPagedStore* s, uint32_t payload){
return nat; return nat;
} }
static int int_max_keys(EngramPagedStore* s){ static int int_max_keys(EngramPagedStore* s){
/* keys*8 + (keys+1)*8 <= IDX_BODY → keys <= IDX_BODY/8 - 1 */ /* An internal node stores `keys` u64 keys FOLLOWED BY (keys+1) u64 child
int nat = (int)(IDX_BODY / 8) - 1; * pointers, so both arrays must fit the page body:
* keys*8 + (keys+1)*8 = 16*keys + 8 <= IDX_BODY keys <= (IDX_BODY-8)/16.
* The prior form `IDX_BODY/8 - 1` divided by 8 instead of 16 it counted
* only the key array and ignored the child array's 8 bytes/key so it
* returned ~2x the real capacity (2041 vs 1020 at a 16 KB page). An internal
* node was then allowed to grow past what a page holds, and btree_insert's
* write-back overran its STORE_PAGE_SIZE stack page buffer, smashing the
* stack canary (__stack_chk_fail). That was the live crash-loop root cause. */
int nat = (int)((IDX_BODY - 8) / 16);
if (s->int_max > 0 && s->int_max < nat) return s->int_max; if (s->int_max > 0 && s->int_max < nat) return s->int_max;
return nat; return nat;
} }
@@ -556,6 +599,22 @@ static int btree_insert(EngramPagedStore* s, int tree, uint64_t page_id,
int nkeys = get_u16(buf + 10); int nkeys = get_u16(buf + 10);
int is_leaf = buf[IDX_LEAF_OFF]; int is_leaf = buf[IDX_LEAF_OFF];
/* Defensive bound: never trust an on-disk entry count enough to overflow the
* fixed STORE_PAGE_SIZE stack buffer below. A leaf holds at most IDX_BODY/esz
* entries; an internal node at most (IDX_BODY-8)/16 keys (keys + child ptrs).
* A page claiming more than its physical capacity is torn/corrupt (or was
* written by a pre-fix build) fail LOUD and abort this insert rather than
* smash the stack or silently truncate. With this guard the memmove/memcpy/
* put_u64 write-backs are provably in-bounds regardless of on-disk content. */
int _cap = is_leaf ? (int)(IDX_BODY / esz) : (int)((IDX_BODY - 8) / 16);
if (nkeys < 0 || nkeys > _cap){
fprintf(stderr, "engram_store: corrupt %s index page %llu: nkeys=%d "
"exceeds page capacity %d — refusing insert (fail-safe)\n",
is_leaf ? "leaf" : "internal",
(unsigned long long)page_id, nkeys, _cap);
return -1;
}
if (is_leaf){ if (is_leaf){
/* find insert position (after equal keys → stable duplicates) */ /* find insert position (after equal keys → stable duplicates) */
int pos = 0; int pos = 0;
@@ -834,6 +893,7 @@ EngramPagedStore* store_create(const char* path){
close(s->fd); free(s); return NULL; close(s->fd); free(s); return NULL;
} }
if (store_sync(s)!=0){ close(s->fd); free(s); return NULL; } if (store_sync(s)!=0){ close(s->fd); free(s); return NULL; }
store__autopin(s); /* keep superblocks + index roots resident */
return s; return s;
} }
@@ -861,6 +921,7 @@ EngramPagedStore* store_open(const char* path){
s->next_lsn = (s->last_checkpoint_lsn > s->sb_seq) ? s->last_checkpoint_lsn : s->sb_seq; s->next_lsn = (s->last_checkpoint_lsn > s->sb_seq) ? s->last_checkpoint_lsn : s->sb_seq;
s->cur_node_page = 0; s->cur_node_page = 0;
s->cur_edge_page = 0; s->cur_edge_page = 0;
store__autopin(s); /* keep superblocks + index roots resident */
return s; return s;
} }
@@ -976,8 +1037,17 @@ static int read_body(EngramPagedStore* s, uint64_t page, uint16_t slot,
uint16_t off,len,fl; slp_slot(buf, slot, &off, &len, &fl); uint16_t off,len,fl; slp_slot(buf, slot, &off, &len, &fl);
*live_out = (fl == SLOT_LIVE); *live_out = (fl == SLOT_LIVE);
if (len < REC_HDR) return -1; if (len < REC_HDR) return -1;
/* Defensive: the slot's (off,len) come from on-disk bytes. A stale primary
* index entry (churn/crash can leave one pointing at a page later repurposed)
* or a torn slot dir can yield an off/len that runs past this 16 KB stack page
* buffer buf[off+..] would then read off the stack (observed EXC_BAD_ACCESS
* via store_get_node on the bloated store). Bound the record to the page and
* fail safe rather than over-read. */
if ((size_t)off + REC_HDR > STORE_PAGE_SIZE || (size_t)off + len > STORE_PAGE_SIZE)
return -1;
uint8_t rec_flags = buf[off + 3]; uint8_t rec_flags = buf[off + 3];
if (rec_flags & REC_OVERFLOW){ if (rec_flags & REC_OVERFLOW){
if ((size_t)off + REC_HDR + 16 > STORE_PAGE_SIZE) return -1; /* head+total u64s */
uint64_t head = get_u64(buf + off + REC_HDR); uint64_t head = get_u64(buf + off + REC_HDR);
uint64_t total = get_u64(buf + off + REC_HDR + 8); uint64_t total = get_u64(buf + off + REC_HDR + 8);
uint8_t* body = ovf_read_chain(s, head, (size_t)total); uint8_t* body = ovf_read_chain(s, head, (size_t)total);
@@ -985,6 +1055,7 @@ static int read_body(EngramPagedStore* s, uint64_t page, uint16_t slot,
*body_out = body; *blen_out = (size_t)total; *body_out = body; *blen_out = (size_t)total;
} else { } else {
uint16_t reclen = get_u16(buf + off); uint16_t reclen = get_u16(buf + off);
if (reclen < REC_HDR || (size_t)off + reclen > STORE_PAGE_SIZE) return -1;
size_t blen = reclen - REC_HDR; size_t blen = reclen - REC_HDR;
uint8_t* body = (uint8_t*)malloc(blen ? blen : 1); uint8_t* body = (uint8_t*)malloc(blen ? blen : 1);
if (!body) return -1; if (!body) return -1;
@@ -1140,35 +1211,53 @@ uint64_t store_page_count(const EngramPagedStore* s){ return s ? s->page_count :
/* ── M3: full live enumeration (boundary-clean; StoreNode/StoreEdge out only) ── /* ── M3: full live enumeration (boundary-clean; StoreNode/StoreEdge out only) ──
* Page-walk every NODE/EDGE page, emitting each DISTINCT live record. A re-put * Page-walk every NODE/EDGE page, emitting each DISTINCT live record. A re-put
* leaves several live records for one id (apply_node_put appends; reads dedup), * leaves several live records for one id (apply_node_put appends; reads dedup),
* so we track ids already emitted by their 64-bit id-hash the same key the * so we track ids already emitted and fetch the canonical latest-live via the
* primary B+-tree uses (design §2.4) and fetch the canonical latest-live via * point-read path so a scan and a get agree exactly. Used by the caller
* the point-read path so a scan and a get agree exactly. Used by the caller * (el_runtime) to load the whole store resident at boot and to export JSON.
* (el_runtime) to load the whole store resident at boot and to export JSON. */ *
typedef struct { uint64_t* h; size_t n, cap; } U64Set; * DEDUP IS BY FULL ID STRING, NOT BY id_hash. (2026-08-12 self-review the
static int u64set_add(U64Set* s, uint64_t v){ /* 1 = newly added, 0 = present */ * "saved but not findable" bug.) The dedup set formerly keyed on the 64-bit
* id_hash alone; two DISTINCT ids that collide under FNV-1a-64 therefore
* emitted only the first, and the second durably on a live page and findable
* by store_get_node, which disambiguates by strcmp was SILENTLY DROPPED from
* the resident boot-load. After any reopen it was unretrievable by id, absent
* from lexical search, and missing from the recent list. The primary B+-tree
* keys on id_hash too, but every reader there re-reads the record and strcmp's
* the id; the scan's dedup must apply the same full-id discipline. Keyed on the
* hash for O(1) bucketing, compared by strcmp for correctness. */
typedef struct { char* key; } StrSlot;
typedef struct { StrSlot* t; size_t n, cap; } StrSet;
static int strset_add(StrSet* s, const char* id){ /* 1 = newly added, 0 = present */
if (!id) return 1;
if ((s->n + 1) * 4 >= s->cap * 3){ if ((s->n + 1) * 4 >= s->cap * 3){
size_t nc = s->cap ? s->cap * 2 : 1024; size_t nc = s->cap ? s->cap * 2 : 1024;
uint64_t* nh = (uint64_t*)calloc(nc, sizeof(uint64_t)); StrSlot* nt = (StrSlot*)calloc(nc, sizeof(StrSlot));
if (!nh) return 1; /* degrade rather than crash */ if (!nt) return 1; /* degrade rather than crash */
for (size_t i = 0; i < s->cap; i++){ for (size_t i = 0; i < s->cap; i++){
uint64_t k = s->h[i]; char* k = s->t[i].key;
if (k){ size_t j = k & (nc - 1); while (nh[j]) j = (j + 1) & (nc - 1); nh[j] = k; } if (k){ size_t j = id_hash(k) & (nc - 1); while (nt[j].key) j = (j + 1) & (nc - 1); nt[j].key = k; }
} }
free(s->h); s->h = nh; s->cap = nc; free(s->t); s->t = nt; s->cap = nc;
} }
uint64_t k = v ? v : 1; /* 0 reserved as empty slot */ size_t j = id_hash(id) & (s->cap - 1);
size_t j = k & (s->cap - 1); while (s->t[j].key){ if (strcmp(s->t[j].key, id) == 0) return 0; j = (j + 1) & (s->cap - 1); }
while (s->h[j]){ if (s->h[j] == k) return 0; j = (j + 1) & (s->cap - 1); } s->t[j].key = strdup(id);
s->h[j] = k; s->n++; return 1; if (!s->t[j].key) return 1; /* OOM: don't dedup, never drop */
s->n++; return 1;
}
static void strset_free(StrSet* s){
for (size_t i = 0; i < s->cap; i++) free(s->t[i].key);
free(s->t); s->t = NULL; s->n = s->cap = 0;
} }
int store_scan_nodes(EngramPagedStore* s, StoreNodeScanCb cb, void* ctx){ int store_scan_nodes(EngramPagedStore* s, StoreNodeScanCb cb, void* ctx){
if (!s || !cb) return -1; if (!s || !cb) return -1;
U64Set seen = {0, 0, 0}; StrSet seen = {0, 0, 0};
uint8_t buf[STORE_PAGE_SIZE]; uint8_t buf[STORE_PAGE_SIZE];
int count = 0; int count = 0;
for (uint64_t pg = 2; pg < s->page_count; pg++){ for (uint64_t pg = 2; pg < s->page_count; pg++){
if (page_read(s, pg, buf) != 0) continue; if (page_read(s, pg, buf) != 0) continue;
if (s->cache) pc_prefetch(s, pg, s->cache->prefetch); /* sequential read-ahead */
if (buf[8] != STORE_PT_NODE) continue; if (buf[8] != STORE_PT_NODE) continue;
int ns = slp_count(buf); int ns = slp_count(buf);
for (int i = 0; i < ns; i++){ for (int i = 0; i < ns; i++){
@@ -1177,27 +1266,28 @@ int store_scan_nodes(EngramPagedStore* s, StoreNodeScanCb cb, void* ctx){
uint8_t* body; size_t blen; int live; uint8_t* body; size_t blen; int live;
if (read_body(s, pg, (uint16_t)i, &body, &blen, &live) != 0) continue; if (read_body(s, pg, (uint16_t)i, &body, &blen, &live) != 0) continue;
StoreNode cand; node_parse(body, blen, &cand); free(body); StoreNode cand; node_parse(body, blen, &cand); free(body);
if (cand.id && u64set_add(&seen, id_hash(cand.id))){ if (cand.id && *cand.id && strset_add(&seen, cand.id)){
StoreNode canon; StoreNode canon;
if (store_get_node(s, cand.id, &canon) == 1){ if (store_get_node(s, cand.id, &canon) == 1){
cb(&canon, ctx); count++; cb(&canon, ctx); count++; /* canonical latest-live */
store_node_free(&canon); store_node_free(&canon);
} }
} }
store_node_free(&cand); store_node_free(&cand);
} }
} }
free(seen.h); strset_free(&seen);
return count; return count;
} }
int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx){ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx){
if (!s || !cb) return -1; if (!s || !cb) return -1;
U64Set seen = {0, 0, 0}; StrSet seen = {0, 0, 0};
uint8_t buf[STORE_PAGE_SIZE]; uint8_t buf[STORE_PAGE_SIZE];
int count = 0; int count = 0;
for (uint64_t pg = 2; pg < s->page_count; pg++){ for (uint64_t pg = 2; pg < s->page_count; pg++){
if (page_read(s, pg, buf) != 0) continue; if (page_read(s, pg, buf) != 0) continue;
if (s->cache) pc_prefetch(s, pg, s->cache->prefetch); /* sequential read-ahead */
if (buf[8] != STORE_PT_EDGE) continue; if (buf[8] != STORE_PT_EDGE) continue;
int ns = slp_count(buf); int ns = slp_count(buf);
for (int i = 0; i < ns; i++){ for (int i = 0; i < ns; i++){
@@ -1206,17 +1296,17 @@ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx){
uint8_t* body; size_t blen; int live; uint8_t* body; size_t blen; int live;
if (read_body(s, pg, (uint16_t)i, &body, &blen, &live) != 0) continue; if (read_body(s, pg, (uint16_t)i, &body, &blen, &live) != 0) continue;
StoreEdge cand; edge_parse(body, blen, &cand); free(body); StoreEdge cand; edge_parse(body, blen, &cand); free(body);
if (cand.id && u64set_add(&seen, id_hash(cand.id))){ if (cand.id && *cand.id && strset_add(&seen, cand.id)){
StoreEdge canon; StoreEdge canon;
if (store_get_edge(s, cand.id, &canon) == 1){ if (store_get_edge(s, cand.id, &canon) == 1){
cb(&canon, ctx); count++; cb(&canon, ctx); count++; /* canonical latest-live */
store_edge_free(&canon); store_edge_free(&canon);
} }
} }
store_edge_free(&cand); store_edge_free(&cand);
} }
} }
free(seen.h); strset_free(&seen);
return count; return count;
} }
@@ -1247,8 +1337,36 @@ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx){
#include <sys/time.h> #include <sys/time.h>
/* ── write-back buffer pool ────────────────────────────────────────────────── */ /* ══════════════════════════════════════════════════════════════════════════════
struct PgCache { PgEnt** buckets; size_t nbuckets; size_t count; }; * M4 demand-paging BUFFER POOL (bounded, LRU, pinned, read-ahead)
*
* A frame table (idframe hash) capped at `cap` resident frames. On a page
* access that is not resident, page_read faults it in from neuron.egm; if the
* pool is full, the LRU eviction path reclaims a CLEAN, unpinned frame. This is
* purely additive residency the on-disk format is unchanged, and with the
* DEFAULT cap (large) no eviction ever fires, so behaviour is byte-for-byte the
* Phase-1 resident store.
*
* Invariants preserved from M2 (write-back, NO-STEAL):
* A DIRTY frame is NEVER evicted (never stolen) its only durable copy is
* the fsync'd WAL, and the store page reaches disk solely at a checkpoint.
* pc_flush (checkpoint) is what turns dirtyclean and thus evictable.
* A PINNED frame is never evicted. Structural pages are auto-pinned: the two
* superblocks (pages 0,1) and every index ROOT/INTERIOR page (type INDEX,
* leaf-flag 0). Leaves are pageable. Explicit pins (pin count) cover hot
* layers and any caller-designated page.
* Correctness under a pool SMALLER than the store rests on: every caller copies
* page bytes into a local stack buffer (memcpy in page_read / out in page_write)
* and never retains a frame pointer across another page access, so a frame may
* be evicted and later re-faulted with no aliasing hazard. A clean frame always
* matches disk, so a re-fault reproduces identical bytes.
* */
/* default frame budget: large enough that today's whole store stays resident
* (== Phase 1). Override with env ENGRAM_POOL_FRAMES (0 = unlimited). */
#ifndef ENGRAM_POOL_FRAMES_DEFAULT
#define ENGRAM_POOL_FRAMES_DEFAULT (1u<<20) /* ~1M frames × 16KiB = 16 GiB */
#endif
static PgCache* pc_new(void){ static PgCache* pc_new(void){
PgCache* c = (PgCache*)calloc(1, sizeof *c); PgCache* c = (PgCache*)calloc(1, sizeof *c);
@@ -1256,6 +1374,12 @@ static PgCache* pc_new(void){
c->nbuckets = 1024; c->nbuckets = 1024;
c->buckets = (PgEnt**)calloc(c->nbuckets, sizeof(PgEnt*)); c->buckets = (PgEnt**)calloc(c->nbuckets, sizeof(PgEnt*));
if (!c->buckets){ free(c); return NULL; } if (!c->buckets){ free(c); return NULL; }
c->cap = ENGRAM_POOL_FRAMES_DEFAULT;
c->prefetch = 8;
const char* pf = getenv("ENGRAM_POOL_FRAMES");
if (pf && *pf){ char* end=NULL; unsigned long long v = strtoull(pf,&end,10); c->cap = (size_t)v; }
const char* pw = getenv("ENGRAM_PREFETCH");
if (pw && *pw){ char* end=NULL; unsigned long v = strtoul(pw,&end,10); c->prefetch = (unsigned)v; }
return c; return c;
} }
static void pc_free(PgCache* c){ static void pc_free(PgCache* c){
@@ -1264,14 +1388,27 @@ static void pc_free(PgCache* c){
PgEnt* e = c->buckets[i]; PgEnt* e = c->buckets[i];
while (e){ PgEnt* n=e->next; free(e->buf); free(e); e=n; } while (e){ PgEnt* n=e->next; free(e->buf); free(e); e=n; }
} }
for (size_t i=0;i<c->lp_n;i++) free(c->lp[i].pages);
free(c->lp);
free(c->buckets); free(c); free(c->buckets); free(c);
} }
static PgEnt* pc_get(EngramPagedStore* s, uint64_t id){
PgCache* c = s->cache; /* ── LRU recency list (front = MRU, back = LRU) ─────────────────────────────── */
PgEnt* e = c->buckets[id % c->nbuckets]; static void lru_unlink(PgCache* c, PgEnt* e){
while (e){ if (e->id==id) return e; e=e->next; } if (e->lru_prev) e->lru_prev->lru_next = e->lru_next; else c->mru = e->lru_next;
return NULL; if (e->lru_next) e->lru_next->lru_prev = e->lru_prev; else c->lru = e->lru_prev;
e->lru_prev = e->lru_next = NULL;
} }
static void lru_push_front(PgCache* c, PgEnt* e){
e->lru_prev = NULL; e->lru_next = c->mru;
if (c->mru) c->mru->lru_prev = e; c->mru = e;
if (!c->lru) c->lru = e;
}
static void lru_touch(PgCache* c, PgEnt* e){
if (c->mru == e) return;
lru_unlink(c, e); lru_push_front(c, e);
}
static void pc_maybe_grow(PgCache* c){ static void pc_maybe_grow(PgCache* c){
if (c->count <= c->nbuckets*4) return; if (c->count <= c->nbuckets*4) return;
size_t nn = c->nbuckets*2; size_t nn = c->nbuckets*2;
@@ -1283,9 +1420,54 @@ static void pc_maybe_grow(PgCache* c){
} }
free(c->buckets); c->buckets=nb; c->nbuckets=nn; free(c->buckets); c->buckets=nb; c->nbuckets=nn;
} }
/* A frame is EVICTABLE iff it is clean, unpinned, not a superblock, and not an
* index root/interior page. This is the sole place the no-steal + structural-pin
* policy is enforced. */
static int pc_evictable(const PgEnt* e){
if (e->dirty) return 0; /* no-steal: dirty pages are pinned to RAM */
if (e->pin > 0) return 0; /* explicit / hot-layer pin */
if (e->id == 0 || e->id == 1) return 0; /* superblock + mirror */
if (e->buf[8] == STORE_PT_INDEX && e->buf[IDX_LEAF_OFF] == 0) return 0; /* root/interior */
return 1;
}
/* Detach `e` from both the hash chain and the recency list, and free it. */
static void pc_remove(PgCache* c, PgEnt* e){
size_t b = e->id % c->nbuckets;
PgEnt** pp = &c->buckets[b];
while (*pp && *pp != e) pp = &(*pp)->next;
if (*pp == e) *pp = e->next;
lru_unlink(c, e);
free(e->buf); free(e);
c->count--;
}
/* Reclaim clean unpinned frames from the LRU end until under budget, or until no
* evictable frame remains (a dirty/pinned-heavy pool may transiently exceed cap
* that is the no-steal guarantee, not a bug: the next checkpoint frees them). */
static void pc_evict_to_budget(PgCache* c){
if (!c->cap) return; /* unlimited */
while (c->count > c->cap){
PgEnt* e = c->lru; int freed = 0;
while (e){
PgEnt* prev = e->lru_prev; /* walk LRU→MRU */
if (pc_evictable(e)){ pc_remove(c, e); c->evictions++; freed = 1; break; }
e = prev;
}
if (!freed) break; /* nothing evictable — allowed to exceed cap */
}
}
static PgEnt* pc_get(EngramPagedStore* s, uint64_t id){
PgCache* c = s->cache;
PgEnt* e = c->buckets[id % c->nbuckets];
while (e){ if (e->id==id){ lru_touch(c, e); return e; } e=e->next; }
return NULL;
}
/* Insert-or-update a frame. New frames go to MRU; then evict down to budget.
* The just-touched frame is at MRU and can never be the eviction victim. */
static int pc_put(EngramPagedStore* s, uint64_t id, const uint8_t* buf, int dirty){ static int pc_put(EngramPagedStore* s, uint64_t id, const uint8_t* buf, int dirty){
PgCache* c = s->cache; PgCache* c = s->cache;
PgEnt* e = pc_get(s, id); PgEnt* e = pc_get(s, id); /* pc_get also bumps it to MRU on a hit */
if (!e){ if (!e){
e = (PgEnt*)calloc(1, sizeof *e); e = (PgEnt*)calloc(1, sizeof *e);
if (!e) return -1; if (!e) return -1;
@@ -1294,11 +1476,13 @@ static int pc_put(EngramPagedStore* s, uint64_t id, const uint8_t* buf, int dirt
e->id = id; e->id = id;
size_t b = id % c->nbuckets; size_t b = id % c->nbuckets;
e->next = c->buckets[b]; c->buckets[b] = e; c->count++; e->next = c->buckets[b]; c->buckets[b] = e; c->count++;
lru_push_front(c, e);
pc_maybe_grow(c); pc_maybe_grow(c);
} }
memcpy(e->buf, buf, STORE_PAGE_SIZE); memcpy(e->buf, buf, STORE_PAGE_SIZE);
e->lsn = get_u64(buf + 16); e->lsn = get_u64(buf + 16);
if (dirty) e->dirty = 1; if (dirty){ if (!e->dirty) c->dirty_count++; e->dirty = 1; } /* clean→dirty transition */
pc_evict_to_budget(c);
return 0; return 0;
} }
static int pc_flush(EngramPagedStore* s){ static int pc_flush(EngramPagedStore* s){
@@ -1307,8 +1491,160 @@ static int pc_flush(EngramPagedStore* s){
for (size_t i=0;i<c->nbuckets;i++) for (size_t i=0;i<c->nbuckets;i++)
for (PgEnt* e=c->buckets[i]; e; e=e->next) for (PgEnt* e=c->buckets[i]; e; e=e->next)
if (e->dirty){ if (page_write_raw(s, e->id, e->buf)!=0) return -1; e->dirty=0; } if (e->dirty){ if (page_write_raw(s, e->id, e->buf)!=0) return -1; e->dirty=0; }
c->dirty_count = 0; /* all frames clean after flush */
/* Post-checkpoint the just-cleaned frames are now evictable; trim the pool
* back to budget so a dirty-heavy burst that transiently overshot cap does
* not leave the pool oversized. No-op at the default (unlimited-ish) cap. */
pc_evict_to_budget(c);
return 0; return 0;
} }
/* Bounded sequential read-ahead: fault the next `window` pages after `from_id`
* into any spare capacity, so a forward scan/leaf-walk hits them instead of
* faulting one-by-one. Never forces an eviction (fills slack only), never
* re-reads a resident page. Prefetch reads are counted separately from demand
* faults so a scan's fault count reflects on-demand misses only. */
static void pc_prefetch(EngramPagedStore* s, uint64_t from_id, unsigned window){
PgCache* c = s->cache;
if (!c || !window) return;
for (unsigned k=1; k<=window; k++){
uint64_t id = from_id + k;
if (id >= s->page_count) break;
if (c->cap && c->count + 1 > c->cap) break; /* no eviction for read-ahead */
if (c->buckets[id % c->nbuckets]){
PgEnt* e = c->buckets[id % c->nbuckets];
int resident = 0; while (e){ if (e->id==id){ resident=1; break; } e=e->next; }
if (resident) continue;
}
uint8_t buf[STORE_PAGE_SIZE];
off_t off = (off_t)id * STORE_PAGE_SIZE;
if (pread(s->fd, buf, STORE_PAGE_SIZE, off) != (ssize_t)STORE_PAGE_SIZE) break;
pc_put(s, id, buf, 0);
c->prefetch_reads++;
}
}
/* Non-LRU-touching frame lookup (for pin bookkeeping that must not reorder). */
static PgEnt* pc_find(PgCache* c, uint64_t id){
PgEnt* e = c->buckets[id % c->nbuckets];
while (e){ if (e->id==id) return e; e=e->next; }
return NULL;
}
/* ── public pin / prefetch / stats API (M4) ─────────────────────────────────── */
int store_pin_page(EngramPagedStore* s, uint64_t page_id){
if (!s || !s->cache) return -1;
uint8_t buf[STORE_PAGE_SIZE];
if (page_read(s, page_id, buf) != 0) return -1; /* fault in + make resident */
PgEnt* e = pc_find(s->cache, page_id);
if (!e) return -1;
e->pin++;
return 0;
}
int store_unpin_page(EngramPagedStore* s, uint64_t page_id){
if (!s || !s->cache) return -1;
PgEnt* e = pc_find(s->cache, page_id);
if (e && e->pin > 0) e->pin--;
return 0;
}
/* Pin every page currently holding a live record of `layer` (hot-layer residency).
* Records the pinned pages so store_unpin_layer releases exactly this set. Pages
* are pinned BEFORE their bodies are read so a small pool cannot evict them mid-scan. */
int store_pin_layer(EngramPagedStore* s, uint32_t layer){
if (!s || !s->cache) return -1;
uint64_t* pages = NULL; size_t np = 0, cap = 0;
uint8_t buf[STORE_PAGE_SIZE];
for (uint64_t pg = 2; pg < s->page_count; pg++){
if (page_read(s, pg, buf) != 0) continue;
int t = buf[8];
if (t != STORE_PT_NODE && t != STORE_PT_EDGE) continue;
PgEnt* pe = pc_find(s->cache, pg);
if (!pe) continue;
pe->pin++; /* provisional pin: keeps pg resident */
int ns = slp_count(buf), match = 0;
for (int i = 0; i < ns && !match; i++){
uint16_t off, len, fl; slp_slot(buf, i, &off, &len, &fl);
if (fl != SLOT_LIVE) continue;
uint8_t* body; size_t blen; int live;
if (read_body(s, pg, (uint16_t)i, &body, &blen, &live) != 0) continue;
uint32_t lid = 0;
if (t == STORE_PT_NODE){ StoreNode c; node_parse(body, blen, &c); lid = c.layer_id; store_node_free(&c); }
else { StoreEdge c; edge_parse(body, blen, &c); lid = c.layer_id; store_edge_free(&c); }
free(body);
if (lid == layer) match = 1;
}
if (match){
if (np == cap){ cap = cap ? cap*2 : 16; uint64_t* np2 = (uint64_t*)realloc(pages, cap*sizeof *pages); if (!np2){ free(pages); return -1; } pages = np2; }
pages[np++] = pg; /* keep the pin */
} else {
pe->pin--; /* no match on this page: drop provisional pin */
}
}
PgCache* c = s->cache;
if (c->lp_n == c->lp_cap){ c->lp_cap = c->lp_cap ? c->lp_cap*2 : 8; c->lp = (LayerPin*)realloc(c->lp, c->lp_cap*sizeof *c->lp); }
c->lp[c->lp_n].layer = layer; c->lp[c->lp_n].pages = pages; c->lp[c->lp_n].n = np; c->lp_n++;
return (int)np;
}
int store_unpin_layer(EngramPagedStore* s, uint32_t layer){
if (!s || !s->cache) return -1;
PgCache* c = s->cache;
for (size_t i = 0; i < c->lp_n; i++){
if (c->lp[i].layer != layer) continue;
for (size_t j = 0; j < c->lp[i].n; j++){
PgEnt* e = pc_find(c, c->lp[i].pages[j]);
if (e && e->pin > 0) e->pin--;
}
free(c->lp[i].pages);
c->lp[i] = c->lp[--c->lp_n]; /* swap-remove */
return 0;
}
return 0;
}
/* Auto-pin the structural pages: both superblocks and the two index roots (plus
* the layer registry). A SHALLOW index root is a LEAF, so it is not covered by
* the "index interior" eviction rule pinning it explicitly guarantees the root
* is never evicted even for a tiny tree. Deeper roots/interiors are additionally
* covered by pc_evictable's INDEX-non-leaf rule. Best-effort (ignores errors on
* a not-yet-built store). */
static void store__autopin(EngramPagedStore* s){
if (!s || !s->cache) return;
store_pin_page(s, 0);
store_pin_page(s, 1);
if (s->root_index_page) store_pin_page(s, s->root_index_page);
if (s->adj_index_page) store_pin_page(s, s->adj_index_page);
if (s->layer_registry_page) store_pin_page(s, s->layer_registry_page);
}
/* Introspection + test hooks. */
void store_pool_stats(const EngramPagedStore* s, StorePoolStats* out){
if (!out) return;
memset(out, 0, sizeof *out);
if (!s || !s->cache) return;
const PgCache* c = s->cache;
out->cap = c->cap; out->resident = c->count; out->prefetch = c->prefetch;
out->hits = c->hits; out->misses = c->misses;
out->evictions = c->evictions; out->prefetch_reads = c->prefetch_reads;
size_t pinned = 0, dirty = 0;
for (size_t i=0;i<c->nbuckets;i++)
for (PgEnt* e=c->buckets[i]; e; e=e->next){
if (!pc_evictable(e)) pinned++;
if (e->dirty) dirty++;
}
out->pinned = pinned; out->dirty = dirty;
}
int store_pool_resident(const EngramPagedStore* s, uint64_t page_id){
if (!s || !s->cache) return -1;
return pc_find(s->cache, page_id) ? 1 : 0;
}
void store__set_pool_frames(EngramPagedStore* s, size_t frames){
if (!s || !s->cache) return;
s->cache->cap = frames;
pc_evict_to_budget(s->cache); /* apply the new budget now */
}
void store__set_prefetch(EngramPagedStore* s, unsigned window){
if (s && s->cache) s->cache->prefetch = window;
}
/* ── WAL log ───────────────────────────────────────────────────────────────── */ /* ── WAL log ───────────────────────────────────────────────────────────────── */
enum { OP_NODE_PUT=1, OP_EDGE_PUT, OP_TOMBSTONE, OP_SUPERSEDE, enum { OP_NODE_PUT=1, OP_EDGE_PUT, OP_TOMBSTONE, OP_SUPERSEDE,
OP_LAYER_PUT, OP_LAYER_DEL, OP_FORGET, OP_HEBB_BATCH, OP_CHECKPOINT }; OP_LAYER_PUT, OP_LAYER_DEL, OP_FORGET, OP_HEBB_BATCH, OP_CHECKPOINT };
@@ -1323,6 +1659,7 @@ struct EngramWal {
uint64_t last_fsync_lsn; uint64_t last_fsync_lsn;
long long last_fsync_ms; long long last_fsync_ms;
uint64_t appended_since_fsync; uint64_t appended_since_fsync;
uint64_t bytes_since_reclaim; /* WAL bytes appended since last reclaim (M5) */
}; };
static long long now_ms(void){ static long long now_ms(void){
@@ -1380,12 +1717,14 @@ static int wal_append(EngramPagedStore* s, uint8_t op, const uint8_t* payload,
free(fr); free(fr);
if (wr != (ssize_t)fl) return -1; if (wr != (ssize_t)fl) return -1;
w->appended_since_fsync++; w->appended_since_fsync++;
w->bytes_since_reclaim += fl;
wal_maybe_fsync(s, lsn); wal_maybe_fsync(s, lsn);
return 0; return 0;
} }
static int wal_reclaim(EngramPagedStore* s, uint64_t ckpt_lsn){ static int wal_reclaim(EngramPagedStore* s, uint64_t ckpt_lsn){
EngramWal* w = s->wal; if (!w) return 0; EngramWal* w = s->wal; if (!w) return 0;
if (ftruncate(w->fd, 0) != 0) return -1; /* prefix <= ckpt reclaimed */ if (ftruncate(w->fd, 0) != 0) return -1; /* prefix <= ckpt reclaimed */
w->bytes_since_reclaim = 0; /* WAL just shrank to the marker */
uint8_t p[8]; put_u64(p, ckpt_lsn); uint8_t p[8]; put_u64(p, ckpt_lsn);
if (wal_append(s, OP_CHECKPOINT, p, 8, ckpt_lsn) != 0) return -1; if (wal_append(s, OP_CHECKPOINT, p, 8, ckpt_lsn) != 0) return -1;
fsync(w->fd); w->last_fsync_ms = now_ms(); fsync(w->fd); w->last_fsync_ms = now_ms();
@@ -1730,12 +2069,24 @@ static int wal_recover(EngramPagedStore* s){
return rc; return rc;
} }
/* ── checkpoint threshold trigger ────────────────────────────────────────────── */ /* ── background checkpointer: fire on ops / dirty-frames / WAL-bytes / timer ─────
* Single-threaded model: the triggers are evaluated on the write path (no
* background thread), so a checkpoint fires on the first mutation after any armed
* threshold trips. This reclaims the WAL prefix automatically instead of only at
* an explicit engram_checkpoint. Same checkpoint semantics as M2 (it calls the
* very same engram_checkpoint). */
static void ckpt_maybe(EngramPagedStore* s){ static void ckpt_maybe(EngramPagedStore* s){
if (s->recovering) return; if (s->recovering) return;
s->ops_since_ckpt++; s->ops_since_ckpt++;
if (s->ckpt_threshold && s->ops_since_ckpt >= s->ckpt_threshold) int fire = 0;
engram_checkpoint(s); if (s->ckpt_threshold && s->ops_since_ckpt >= s->ckpt_threshold) fire = 1;
if (!fire && s->ckpt_dirty_threshold && s->cache &&
s->cache->dirty_count >= s->ckpt_dirty_threshold) fire = 1;
if (!fire && s->ckpt_wal_threshold && s->wal &&
s->wal->bytes_since_reclaim >= s->ckpt_wal_threshold) fire = 1;
if (!fire && s->ckpt_interval_ms &&
(now_ms() - s->last_ckpt_ms) >= s->ckpt_interval_ms) fire = 1;
if (fire) engram_checkpoint(s);
} }
/* ── public mutation entry points (log-then-apply when a WAL is attached) ─────── */ /* ── public mutation entry points (log-then-apply when a WAL is attached) ─────── */
@@ -1870,6 +2221,7 @@ int store__checkpoint_crashat(EngramPagedStore* s, int phase){
if (phase == 3){ store__crash(s); return 0; } if (phase == 3){ store__crash(s); return 0; }
if (wal_reclaim(s, C) != 0) return -1; /* 4: reclaim WAL prefix */ if (wal_reclaim(s, C) != 0) return -1; /* 4: reclaim WAL prefix */
s->ops_since_ckpt = 0; s->ops_since_ckpt = 0;
s->last_ckpt_ms = now_ms(); /* arm the interval trigger */
if (phase == 4){ store__crash(s); return 0; } if (phase == 4){ store__crash(s); return 0; }
return 0; return 0;
} }
@@ -2097,6 +2449,26 @@ static int import_snapshot(EngramPagedStore* s, const char* path){
return 0; return 0;
} }
/* Arm the background checkpointer with sensible defaults, overridable by env:
* ENGRAM_CKPT_OPS mutations since last checkpoint (default 100000)
* ENGRAM_CKPT_DIRTY dirty pool frames (default 0 = off)
* ENGRAM_CKPT_WAL_BYTES WAL bytes since reclaim (default 64 MiB)
* ENGRAM_CKPT_INTERVAL_MS wall-clock ms (default 0 = off)
* Any of these tripping on the write path triggers a checkpoint ( WAL reclaimed).
* The M4 tests set tiny pools but never hit these bounds, so behaviour is unchanged. */
static void engram__default_ckpt_policy(EngramPagedStore* s){
s->ckpt_threshold = 100000;
s->ckpt_dirty_threshold = 0;
s->ckpt_wal_threshold = 64u*1024u*1024u;
s->ckpt_interval_ms = 0;
s->last_ckpt_ms = now_ms();
const char* e;
if ((e=getenv("ENGRAM_CKPT_OPS")) && *e) s->ckpt_threshold = strtoull(e,NULL,10);
if ((e=getenv("ENGRAM_CKPT_DIRTY")) && *e) s->ckpt_dirty_threshold = (size_t)strtoull(e,NULL,10);
if ((e=getenv("ENGRAM_CKPT_WAL_BYTES")) && *e) s->ckpt_wal_threshold = strtoull(e,NULL,10);
if ((e=getenv("ENGRAM_CKPT_INTERVAL_MS")) && *e) s->ckpt_interval_ms = strtoll(e,NULL,10);
}
/* ── durable-engram boot / close ─────────────────────────────────────────────── */ /* ── durable-engram boot / close ─────────────────────────────────────────────── */
EngramPagedStore* engram_open(const char* data_dir){ EngramPagedStore* engram_open(const char* data_dir){
if (!data_dir) return NULL; if (!data_dir) return NULL;
@@ -2120,7 +2492,7 @@ EngramPagedStore* engram_open(const char* data_dir){
if (!s) return NULL; if (!s) return NULL;
s->wal = wal_open(wal_path, sync); s->wal = wal_open(wal_path, sync);
if (!s->wal){ store_close(s); return NULL; } if (!s->wal){ store_close(s); return NULL; }
s->ckpt_threshold = 100000; engram__default_ckpt_policy(s);
wal_recover(s); /* replay post-checkpoint tail */ wal_recover(s); /* replay post-checkpoint tail */
return s; return s;
} }
@@ -2129,7 +2501,7 @@ EngramPagedStore* engram_open(const char* data_dir){
if (!s) return NULL; if (!s) return NULL;
s->wal = wal_open(wal_path, sync); s->wal = wal_open(wal_path, sync);
if (!s->wal){ store_close(s); return NULL; } if (!s->wal){ store_close(s); return NULL; }
s->ckpt_threshold = 100000; engram__default_ckpt_policy(s);
if (stat(snap_path, &st) == 0) import_snapshot(s, snap_path); if (stat(snap_path, &st) == 0) import_snapshot(s, snap_path);
engram_checkpoint(s); /* store is now authoritative */ engram_checkpoint(s); /* store is now authoritative */
return s; return s;
@@ -2139,3 +2511,175 @@ int engram_close(EngramPagedStore* s){
engram_checkpoint(s); engram_checkpoint(s);
return store_close(s); return store_close(s);
} }
/* ══════════════════════════════════════════════════════════════════════════════
* M5 ONLINE COMPACTION + background-checkpointer policy setter
*
* Dead space accrues in three shapes, all reclaimed here:
* 1. DEAD slots on NODE/EDGE pages tombstones (prune/forget), superseded ids,
* and the stale prior versions a re-put / hebb-batch leaves (apply_*_put
* appends a new record + index entry; the old slot is marked DEAD).
* 2. Duplicate primary/adjacency index entries pointing at those DEAD records.
* 3. OVERFLOW chains orphaned when a large record died (tombstone only flips the
* slot; it never frees the record's overflow pages).
*
* Strategy copy-live + atomic swap (the safest crash-safe relocation):
* A. Quiesce: checkpoint (or sync) so the on-disk .egm fully reflects state and
* the WAL is reduced to its CHECKPOINT{C} marker (C = current LSN watermark).
* B. Build a brand-new store file `<path>.compact` holding ONLY the live records
* walked canonically (latest-live per id) and re-placed bit-exact into
* fresh, densely packed pages with fresh id + adjacency B+-trees. Every page
* is stamped with LSN = C and the new superblock records last_checkpoint_lsn
* = C, so it is LSN-consistent with the (unchanged) WAL. fsync it.
* C. Commit by rename(<path>.compact <path>) POSIX-atomic: recovery sees
* either the whole old file or the whole new file, never a torn mix.
* D. Reopen in place: swap the fd, INVALIDATE every pool frame (old page ids now
* hold different data this is the M4 "remap relocated pages" step), reload
* the superblock, re-autopin.
*
* Crash safety (proven by the test at phases 0/1/2):
* crash in A/B (before rename): old .egm is byte-for-byte intact and the WAL
* still matches it recovery = PRE-compaction (all live records present).
* The half-built `.compact` temp is ignored by engram_open and unlinked at the
* next compaction.
* crash after rename (C/D): the new .egm is fully fsync'd with last_checkpoint
* = C and the WAL (CHECKPOINT{C}, nothing newer) matches it recovery =
* POST-compaction. No undo ever needed because relocation is copy-then-swap,
* never in-place mutation of a still-referenced page.
*
* Online vs quiesce: the store is single-threaded, so "online" means it is safe
* to interleave between mutations (each mutation is a synchronous call) NOT that
* it runs concurrently with one. It takes a checkpoint quiesce point at entry.
*
* M4 cooperation: the build writes into a SEPARATE store `d` whose own pool obeys
* ENGRAM_POOL_FRAMES (so a small pool evicts/re-faults throughout the build,
* no-steal + pins honoured there); the live store's pool is fully invalidated on
* reopen, guaranteeing no stale frame maps a relocated page.
* */
/* Drop every resident frame (relocated pages are no longer valid) but keep the
* pool object with its cap / prefetch / stats. Layer-pin bookkeeping is cleared
* (those page ids belong to the old image). */
static void pc_invalidate_all(PgCache* c){
if (!c) return;
for (size_t i=0;i<c->nbuckets;i++){
PgEnt* e = c->buckets[i];
while (e){ PgEnt* n=e->next; free(e->buf); free(e); e=n; }
c->buckets[i] = NULL;
}
for (size_t i=0;i<c->lp_n;i++) free(c->lp[i].pages);
c->lp_n = 0;
c->count = 0; c->dirty_count = 0; c->mru = c->lru = NULL;
}
/* Re-open the store file in place after an atomic swap: swap fd, invalidate the
* pool, reload the superblock, resume the LSN watermark, re-autopin. Keeps the
* attached WAL (it references the same checkpoint LSN the new file carries). */
static int store__reopen_swapped(EngramPagedStore* s){
if (s->fd >= 0) close(s->fd);
s->fd = open(s->path, O_RDWR);
if (s->fd < 0) return -1;
pc_invalidate_all(s->cache); /* M4: no stale frame for a relocated page */
uint8_t b0[STORE_PAGE_SIZE], b1[STORE_PAGE_SIZE];
uint64_t s0=0, s1=0;
int ok0 = sb_load_one(s, 0, b0, &s0) == 0;
int ok1 = sb_load_one(s, 1, b1, &s1) == 0;
if (!ok0 && !ok1) return -1;
const uint8_t* pick = (ok0&&ok1) ? ((s0>=s1)?b0:b1) : (ok0?b0:b1);
sb_apply(s, pick);
s->next_lsn = (s->last_checkpoint_lsn > s->sb_seq) ? s->last_checkpoint_lsn : s->sb_seq;
s->cur_node_page = 0; s->cur_edge_page = 0;
s->ops_since_ckpt = 0; s->last_ckpt_ms = now_ms();
store__autopin(s);
return 0;
}
/* Callback context for copying live records into the compacted store `d`. */
typedef struct { EngramPagedStore* d; int err; } CompactCtx;
static void compact_node_cb(const StoreNode* n, void* ctx){
CompactCtx* c = (CompactCtx*)ctx;
if (c->err) return;
if (node_place(c->d, n) != 0) c->err = 1; /* bit-exact re-place; stamped d->stamp_lsn */
}
static void compact_edge_cb(const StoreEdge* e, void* ctx){
CompactCtx* c = (CompactCtx*)ctx;
if (c->err) return;
if (edge_place(c->d, e) != 0) c->err = 1;
}
/* Build the compacted image (only live records, fresh indexes) into a new file. */
static int compact_build(EngramPagedStore* s, const char* tmp_path){
unlink(tmp_path); /* drop any temp from a crashed run */
EngramPagedStore* d = store_create(tmp_path);
if (!d) return -1;
uint64_t W = s->next_lsn; /* LSN watermark == checkpoint LSN */
memcpy(d->uuid, s->uuid, 16); /* preserve store identity */
d->last_checkpoint_lsn = W;
d->next_lsn = W;
d->stamp_lsn = W; /* every compacted page → LSN W */
int rc = 0;
/* live layers */
StoreLayer* layers = NULL; size_t nlay = 0;
if (store_list_layers(s, &layers, &nlay) == 0){
for (size_t i=0;i<nlay && rc==0;i++)
if (apply_layer_put(d, &layers[i], W) != 0) rc = -1;
store_layers_free(layers, nlay);
}
/* live nodes + edges (canonical latest-live, dedup by id — see store_scan_*) */
CompactCtx ctx = { d, 0 };
if (rc == 0 && store_scan_nodes(s, compact_node_cb, &ctx) < 0) rc = -1;
if (rc == 0 && store_scan_edges(s, compact_edge_cb, &ctx) < 0) rc = -1;
if (ctx.err) rc = -1;
d->stamp_lsn = 0;
if (rc == 0 && store_sync(d) != 0) rc = -1; /* flush + fsync + both superblocks */
store_close(d);
return rc;
}
int store__compact_crashat(EngramPagedStore* s, int phase){
if (!s) return -1;
/* A. quiesce → on-disk store consistent, WAL reduced to its checkpoint marker */
if (s->wal){ if (engram_checkpoint(s) != 0) return -1; }
else { if (store_sync(s) != 0) return -1; }
if (phase == 0){ store__crash(s); return 0; } /* → recovers pre-compaction */
char tmp[1200];
snprintf(tmp, sizeof tmp, "%s.compact", s->path);
if (compact_build(s, tmp) != 0){ unlink(tmp); return -1; }
if (phase == 1){ store__crash(s); return 0; } /* built, not renamed → pre-compaction */
/* C. atomic commit */
if (rename(tmp, s->path) != 0){ unlink(tmp); return -1; }
if (phase == 2){ store__crash(s); return 0; } /* renamed, not reopened → post-compaction */
/* D. reopen RAM state against the compacted file */
return store__reopen_swapped(s);
}
int store_compact(EngramPagedStore* s){ return store__compact_crashat(s, -1); }
void store_set_checkpoint_policy(EngramPagedStore* s, uint64_t ops,
size_t dirty_pages, uint64_t wal_bytes,
long long interval_ms){
if (!s) return;
s->ckpt_threshold = ops;
s->ckpt_dirty_threshold = dirty_pages;
s->ckpt_wal_threshold = wal_bytes;
s->ckpt_interval_ms = interval_ms;
s->last_ckpt_ms = now_ms();
}
uint64_t store_free_page_count(const EngramPagedStore* s){
if (!s) return 0;
uint64_t n = 0, id = s->free_list_head;
uint8_t buf[STORE_PAGE_SIZE];
while (id){
if (page_read((EngramPagedStore*)s, id, buf) != 0) break;
if (buf[8] != STORE_PT_FREE) break;
n++;
id = get_u64(buf + OVF_NEXT_OFF);
}
return n;
}
+74
View File
@@ -209,6 +209,41 @@ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx);
uint64_t engram_wal_next_lsn(const EngramPagedStore* s); uint64_t engram_wal_next_lsn(const EngramPagedStore* s);
uint64_t engram_last_checkpoint_lsn(const EngramPagedStore* s); uint64_t engram_last_checkpoint_lsn(const EngramPagedStore* s);
/* ── M4: demand-paging buffer pool (additive residency; on-disk format UNCHANGED) ──
*
* The write-back, no-steal cache of M2 becomes a bounded, demand-paged buffer
* pool. A fixed frame budget (env ENGRAM_POOL_FRAMES; 0 = unlimited; default
* large whole store resident identical to Phase 1) keeps only hot pages in
* RAM; a page access that is not resident faults in from neuron.egm, and under
* pressure a CLEAN, unpinned frame is evicted (LRU). Dirty frames are never
* stolen (M2 no-steal / WAL durability), and superblocks + index root/interior
* pages are auto-pinned. Prefetch (env ENGRAM_PREFETCH) reads ahead on scans. */
/* Pin / unpin an individual page (faults it in and keeps it resident until
* unpinned). Pin a hot layer's pages (WM/core) as a set. Idempotent counts. */
int store_pin_page(EngramPagedStore* s, uint64_t page_id);
int store_unpin_page(EngramPagedStore* s, uint64_t page_id);
int store_pin_layer(EngramPagedStore* s, uint32_t layer); /* returns #pages pinned */
int store_unpin_layer(EngramPagedStore* s, uint32_t layer);
/* Buffer-pool introspection. */
typedef struct StorePoolStats {
size_t cap; /* frame budget (0 = unlimited) */
size_t resident; /* frames currently resident */
size_t pinned; /* frames that cannot be evicted (dirty/pinned/structural) */
size_t dirty; /* dirty (un-checkpointed) frames */
unsigned prefetch; /* read-ahead window */
uint64_t hits, misses; /* page_read cache hits / demand faults */
uint64_t evictions; /* clean frames reclaimed */
uint64_t prefetch_reads; /* pages brought in by read-ahead */
} StorePoolStats;
void store_pool_stats(const EngramPagedStore* s, StorePoolStats* out);
int store_pool_resident(const EngramPagedStore* s, uint64_t page_id);
/* Test hooks: set the frame budget / prefetch window at runtime (NOT format). */
void store__set_pool_frames(EngramPagedStore* s, size_t frames);
void store__set_prefetch(EngramPagedStore* s, unsigned window);
/* Crash-test hooks (writes only under a throwaway dir). /* Crash-test hooks (writes only under a throwaway dir).
* store__crash abandon all RAM state without flush/fsync (power loss). * store__crash abandon all RAM state without flush/fsync (power loss).
* store__flush_pages pwrite dirty pages to disk WITHOUT a checkpoint (steal). * store__flush_pages pwrite dirty pages to disk WITHOUT a checkpoint (steal).
@@ -218,4 +253,43 @@ void store__crash(EngramPagedStore* s);
int store__flush_pages(EngramPagedStore* s); int store__flush_pages(EngramPagedStore* s);
int store__checkpoint_crashat(EngramPagedStore* s, int phase); int store__checkpoint_crashat(EngramPagedStore* s, int phase);
/* ── M5: online compaction + background checkpointer (additive; format UNCHANGED) ──
*
* COMPACTION reclaims the space held by DEAD records tombstoned nodes/edges
* (telemetry prune, forget), superseded ids, and the stale prior versions a
* re-put/hebb-batch leaves behind plus the overflow pages they orphaned. It
* rewrites only the LIVE records (bit-exact) into a fresh, densely packed image
* with fresh id + adjacency indexes, then commits the swap atomically, so the
* .egm file physically SHRINKS and the freed pages are reclaimed. Crash-safe:
* a crash at any instant recovers to either the pre- or the post-compaction
* store, never a corrupt mix (atomic rename is the commit point). It cooperates
* with the M4 pool (no-steal, pins) by building into a separate store whose own
* pool honours ENGRAM_POOL_FRAMES, then INVALIDATING every frame of the live
* pool so no stale frame survives for a relocated page.
*
* Requires a quiesce point: store_compact performs a checkpoint (or sync) at
* entry, so it is called between mutations, not concurrently with one. */
int store_compact(EngramPagedStore* s);
/* Test hook: run compaction but stop (then power-loss) after `phase`:
* 0 = after the entry checkpoint, before building ( recovers pre-compaction)
* 1 = after building+fsync the new image, before rename ( pre-compaction)
* 2 = after the atomic rename, before reopening RAM state ( post-compaction)
* phase<0 = full compaction. Frees `s` on a crash phase (like the checkpoint hook). */
int store__compact_crashat(EngramPagedStore* s, int phase);
/* BACKGROUND CHECKPOINTER policy. A checkpoint fires automatically on the write
* path when ANY armed trigger trips, reclaiming the WAL prefix without an explicit
* engram_checkpoint. 0 disables that trigger. Same checkpoint semantics as M2.
* ops mutations since last checkpoint (default 100000)
* dirty_pages dirty (un-checkpointed) pool frames
* wal_bytes bytes appended to the WAL since it was last reclaimed
* interval_ms wall-clock ms since the last checkpoint (checked on writes) */
void store_set_checkpoint_policy(EngramPagedStore* s, uint64_t ops,
size_t dirty_pages, uint64_t wal_bytes,
long long interval_ms);
/* Introspection: number of pages currently on the free-list. */
uint64_t store_free_page_count(const EngramPagedStore* s);
#endif /* ENGRAM_STORE_H */ #endif /* ENGRAM_STORE_H */
+157
View File
@@ -0,0 +1,157 @@
/* engram_verify.c — the VERIFIER layer. Pure compositions over engram_reason.h +
* engram_geometry.h. stdlib + libm only; READ-ONLY over its inputs; touches no
* store/index/activation. See engram_verify.h for the design and the frame contract. */
#include "engram_verify.h"
#include <stdlib.h>
#include <string.h>
#include <math.h>
/* ── small float-vector helpers (mirror engram_reason.c) ────────────────────── */
static double vdot(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) s += (double)a[i] * (double)b[i]; return s;
}
static double l2(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) { double d = (double)a[i] - (double)b[i]; s += d * d; }
return sqrt(s);
}
/* ═══════════════════════════════════════════════════════ GROUNDING ══════════ */
int engram_verify_grounding(const float* claim, int dim,
const GeoDescriptor* const* evidence, int n_evidence,
double ext_floor, double ground_threshold,
GeoGrounding* out) {
if (!claim || dim <= 0 || !evidence || n_evidence < 1 || !out) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
if (!(ground_threshold > 0 && ground_threshold < 1)) ground_threshold = 0.5;
memset(out, 0, sizeof *out);
out->n_evidence = n_evidence;
out->best = -1;
out->nearest_centroid_l2 = INFINITY;
out->scores = malloc((size_t)n_evidence * sizeof(double));
if (!out->scores) return -1;
double best = -1;
for (int i = 0; i < n_evidence; i++) {
const GeoDescriptor* e = evidence[i];
GeoFit f;
if (!e || e->dim != dim || !e->centroid ||
engram_reason_point_fit(e, claim, ext_floor, &f) != 0) {
out->scores[i] = 0.0;
continue;
}
out->scores[i] = f.score;
double cl2 = l2(claim, e->centroid, dim);
if (cl2 < out->nearest_centroid_l2) out->nearest_centroid_l2 = cl2;
if (out->best < 0 || f.score > best) {
best = f.score;
out->best = i;
out->grounding = f.score;
out->best_distance = f.distance;
out->best_ortho = f.ortho_residual;
}
}
if (out->best < 0) { out->grounding = 0.0; out->best_distance = INFINITY; }
out->grounded = (out->grounding >= ground_threshold) ? 1 : 0;
return 0;
}
void engram_verify_grounding_free(GeoGrounding* out) {
if (!out) return;
free(out->scores); out->scores = NULL;
}
/* ═══════════════════════════════════════════════════════ CONSISTENCY ════════ */
int engram_verify_consistency(const float* claim, int dim,
const GeoDescriptor* context,
const GeoDescriptor* pole_pos, const GeoDescriptor* pole_neg,
const GeoDescriptor* forbidden,
double ext_floor, double deadzone_frac,
double forbidden_thresh, double max_distance,
GeoConsistency* out) {
if (!claim || dim <= 0 || !out) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
if (!(deadzone_frac >= 0 && deadzone_frac < 1)) deadzone_frac = 0.10;
if (!(forbidden_thresh > 0 && forbidden_thresh < 1)) forbidden_thresh = 0.5;
memset(out, 0, sizeof *out);
out->verdict = GEO_CONSIST_OK;
out->consistency = 1.0;
int do_polarity = (pole_pos && pole_neg);
int do_distance = (max_distance > 0);
if ((do_polarity || do_distance) &&
(!context || context->dim != dim || !context->centroid)) return -1;
if (do_polarity && (pole_pos->dim != dim || pole_neg->dim != dim ||
!pole_pos->centroid || !pole_neg->centroid)) return -1;
if (forbidden && (forbidden->dim != dim || !forbidden->centroid)) return -1;
double pol_score = 1.0, geo_score = 1.0;
/* ── (a) POLARITY / negation inversion ─────────────────────────────────── */
if (do_polarity) {
/* axis p = (c_pos c_neg); midpoint o = ½(c_pos + c_neg). */
float* p = malloc((size_t)dim * sizeof(float));
float* o = malloc((size_t)dim * sizeof(float));
if (!p || !o) { free(p); free(o); return -1; }
double pn2 = 0;
for (int i = 0; i < dim; i++) {
double dpos = (double)pole_pos->centroid[i], dneg = (double)pole_neg->centroid[i];
p[i] = (float)(dpos - dneg);
o[i] = (float)(0.5 * (dpos + dneg));
pn2 += (dpos - dneg) * (dpos - dneg);
}
double pn = sqrt(pn2);
out->polarity_separation = 0.5 * pn;
if (pn > 1e-12) {
/* signed positions along the axis (projection of (x o) onto unit p). */
float* cdo = malloc((size_t)dim * sizeof(float)); /* claim o */
float* rdo = malloc((size_t)dim * sizeof(float)); /* context o */
if (!cdo || !rdo) { free(p); free(o); free(cdo); free(rdo); return -1; }
for (int i = 0; i < dim; i++) {
cdo[i] = (float)((double)claim[i] - (double)o[i]);
rdo[i] = (float)((double)context->centroid[i] - (double)o[i]);
}
double claim_side = vdot(cdo, p, dim) / pn; /* units: emb-space length */
double ref_side = vdot(rdo, p, dim) / pn;
out->polarity_claim = claim_side;
out->polarity_reference = ref_side;
double dz = deadzone_frac * out->polarity_separation; /* neutral band */
if (fabs(claim_side) > dz && fabs(ref_side) > dz &&
(claim_side > 0) != (ref_side > 0)) {
out->inverted = 1;
pol_score = 0.0; /* opposite poles ⇒ zero consistency */
} else if (fabs(claim_side) <= dz || fabs(ref_side) <= dz) {
pol_score = 0.5; /* neutral / undecided */
} else {
pol_score = 1.0; /* same pole ⇒ consistent */
}
free(cdo); free(rdo);
}
free(p); free(o);
}
/* ── (b) GEOMETRIC contradiction ───────────────────────────────────────── */
if (forbidden) {
GeoFit f;
if (engram_reason_point_fit(forbidden, claim, ext_floor, &f) == 0) {
out->forbidden_fit = f.score;
if (f.score >= forbidden_thresh) {
out->geo_violation = 1;
double g = 1.0 - f.score; if (g < 0) g = 0;
if (g < geo_score) geo_score = g;
}
}
}
if (do_distance) {
out->context_distance = l2(claim, context->centroid, dim);
if (out->context_distance > max_distance) {
out->geo_violation = 1;
geo_score = 0.0;
}
}
/* ── verdict + scalar (polarity is the headline; both flags stay visible) ─ */
out->consistency = (pol_score < geo_score) ? pol_score : geo_score;
if (out->inverted) out->verdict = GEO_CONSIST_POLARITY;
else if (out->geo_violation) out->verdict = GEO_CONSIST_GEOMETRIC;
else out->verdict = GEO_CONSIST_OK;
return 0;
}
+118
View File
@@ -0,0 +1,118 @@
/* engram_verify.h — the VERIFIER layer: GROUNDING + CONSISTENCY over the live
* geometry (engram_geometry.h) and reasoning (engram_reason.h) operators.
*
* The geometry PROPOSES (cheap, creative, sometimes wrong); the verifier DISPOSES.
* This layer catches the class of failure a grammar check never sees: a fluent,
* confident, WRONG output the "plausible lie". The motivating case: a translation
* that DELETED a negation so "you never fought" became "you argued" reassurance
* inverted into accusation, grammatical and invisible, catchable ONLY by the geometry.
*
* GROUNDING claim is there ANY real structure that supports it, or is it
* floating free of the manifold? (anti-hallucination gate)
* CONSISTENCY claim does it CONTRADICT the established structure? Two catches:
* (a) POLARITY: the claim lands on the OPPOSITE side of a negation
* axis from the grounded truth (the reassuranceaccusation catch),
* (b) GEOMETRIC: the claim sits inside a region it must be far from,
* or violates a max-distance constraint to its context.
*
* PURE + READ-ONLY (stdlib + libm only): every function consumes a claim POINT
* (float* in R^dim) plus GeoDescriptor(s), and NEVER touches the store, index, or
* activation. All geometry is delegated to engram_reason_point_fit / engram_geo_*;
* this file only composes and applies thresholds.
*
* FRAME CONTRACT (inherited): the claim point and every descriptor passed together
* MUST share emb `dim` and the same `global_mean` frame exactly the §5 operator
* contract. A function returns <0 on a dim/frame mismatch or bad argument.
*/
#ifndef ENGRAM_VERIFY_H
#define ENGRAM_VERIFY_H
#include "engram_geometry.h"
#include "engram_reason.h"
/* ═══════════════════════════════════════════════════════════════════════════
* GROUNDING anti-hallucination. Score how well a claimed POINT is supported by
* the ACTUAL structure: fit the claim against every real evidence neighborhood
* (engram_reason_point_fit in-distribution Mahalanobis + off-model orthogonal
* residual) and take the BEST supporter. A claim that sits inside real structure
* scores high (grounded); a claim floating far from every neighborhood scores low
* on all of them flagged UNGROUNDED (a hallucination).
*
* This is an ABSOLUTE-THRESHOLD gate, deliberately distinct from ABDUCTION (which
* always RANKS and picks a winner among competing hypotheses): grounding asks the
* prior question "is there any real support at all?" and is allowed to answer no.
* The off-model `ortho_residual` is the sharpest hallucination signal: energy in a
* direction the manifold does not even span.
* */
typedef struct {
double grounding; /* ∈[0,1]: overall support = best fit score */
int grounded; /* 1 iff grounding >= ground_threshold */
int best; /* index of best-supporting evidence structure, or 1 */
double best_distance; /* full point-to-manifold distance to the best */
double best_ortho; /* off-model orthogonal residual of the best fit */
double nearest_centroid_l2;/* raw L2 to the nearest evidence centroid (coarse) */
int n_evidence;
double* scores; /* per-evidence fit score, higher = better (owned)*/
} GeoGrounding;
/* ext_floor>0 guards zero-extent axes (default 1.0). ground_threshold∈(0,1): the
* minimum best-fit score to call the claim grounded (default 0.5). */
int engram_verify_grounding(const float* claim, int dim,
const GeoDescriptor* const* evidence, int n_evidence,
double ext_floor, double ground_threshold,
GeoGrounding* out);
void engram_verify_grounding_free(GeoGrounding* out);
/* ═══════════════════════════════════════════════════════════════════════════
* CONSISTENCY contradiction detection. Does the claim contradict the established
* structure? Two independent sub-checks (either can fire; both flags are reported):
*
* (a) POLARITY / negation inversion. A polarity axis p is defined by two REAL
* poles pole_pos (asserts X) and pole_neg (asserts ¬X):
* p = (c_pos c_neg)/· , midpoint o = ½(c_pos + c_neg).
* The claim's side = p·(claim o); the reference's side = p·(c_context o).
* If the two sides have OPPOSITE sign AND both clear the neutral deadzone, the
* claim asserts the polarity opposite to the grounded truth INVERSION flagged.
* This is the "you never fought""you argued" catch: the truth ("never fought")
* sits on the negate pole, the claim ("argued") on the affirm pole opposite
* sides flagged, though every word is grammatical.
*
* (b) GEOMETRIC contradiction. The claim sits INSIDE a `forbidden` region it must
* be far from (point_fit score to forbidden forbidden_thresh), OR it violates
* a max-distance constraint to its context centroid (L2 > max_distance).
*
* pole_pos/pole_neg may both be NULL to skip the polarity check; forbidden may be
* NULL and max_distance0 to skip the geometric check. `context` (the grounded truth
* region) is required whenever polarity or the distance constraint is used.
* */
typedef enum {
GEO_CONSIST_OK = 0, /* consistent with context */
GEO_CONSIST_POLARITY = 1, /* polarity/negation inversion (asserts ¬X where X) */
GEO_CONSIST_GEOMETRIC = 2 /* geometric contradiction (in forbidden / too far) */
} GeoConsistencyVerdict;
typedef struct {
GeoConsistencyVerdict verdict; /* headline (polarity takes precedence) */
double consistency; /* ∈[0,1]: min over the checks (1 = fully consistent)*/
/* polarity sub-check */
int inverted; /* 1 iff a polarity inversion was detected */
double polarity_claim; /* p·(claim o) (signed position on the axis)*/
double polarity_reference; /* p·(c_context o) (the grounded truth's side) */
double polarity_separation; /* ½‖c_pos c_neg‖ (the axis half-length / scale)*/
/* geometric sub-check */
int geo_violation; /* 1 iff a geometric contradiction was detected */
double forbidden_fit; /* claim's point_fit score to the forbidden region*/
double context_distance; /* L2(claim, c_context) */
} GeoConsistency;
/* ext_floor>0 (default 1.0). deadzone_frac∈[0,1): a polarity side within
* deadzone_frac·separation of the midpoint is "neutral" and never triggers inversion
* (default 0.10). forbidden_thresh(0,1): fit-to-forbidden at/above which the claim
* counts as inside the forbidden region (default 0.5). max_distance>0 enables the
* distance constraint; 0 disables it. */
int engram_verify_consistency(const float* claim, int dim,
const GeoDescriptor* context,
const GeoDescriptor* pole_pos, const GeoDescriptor* pole_neg,
const GeoDescriptor* forbidden,
double ext_floor, double deadzone_frac,
double forbidden_thresh, double max_distance,
GeoConsistency* out);
#endif /* ENGRAM_VERIFY_H */
+649
View File
@@ -0,0 +1,649 @@
/* engram_vindex.c — HNSW ANN index over f32 embedding vectors (design §9 M8).
*
* Self-contained: plain C11, stdlib + libm (-lm for sqrtf/logf) only. No
* dependency on el_runtime; the store is read via its PERMANENT on-disk format
* (design §2.4), decoded read-only here so engram_store.{c,h} stay untouched.
*
* Algorithm: Malkov & Yashunin, "Efficient and robust approximate nearest
* neighbor search using Hierarchical Navigable Small World graphs" (2016).
* - multi-layer graph; level ~ Exp(1/ln M), assigned by a per-node seeded PRNG
* (deterministic: seed = FIXED_SEED ^ node_ordinal) so a rebuild is bit-for-
* bit reproducible regardless of wall-clock or global rand() state.
* - greedy descent through upper layers to an entry point, then an ef-bounded
* best-first search at each layer (Algorithm 2).
* - neighbour selection by the diversity heuristic (Algorithm 4), not plain
* k-nearest, with keep-pruned backfill for connectivity.
* - bidirectional links; a neighbour whose degree exceeds M (2M on layer 0) is
* re-pruned with the same heuristic.
*
* Metric: vectors are L2-normalised on entry, so cosine similarity == dot
* product; distance = 1 - dot (in [0,2], smaller == nearer). Deterministic tie-
* breaks are by element index so results are stable across identical builds.
*/
#include "engram_vindex.h"
#include <stdlib.h>
#include <string.h>
#include <math.h>
#include <stdio.h>
#include <stdint.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>
/* Deterministic PRNG seed base (fixed constant — never wall-clock/rand). */
#define VINDEX_FIXED_SEED 0x9E3779B97F4A7C15ULL
/* ── deterministic PRNG (splitmix64) ──────────────────────────────────────── */
static inline uint64_t splitmix64(uint64_t* s){
uint64_t z = (*s += 0x9E3779B97F4A7C15ULL);
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ULL;
z = (z ^ (z >> 27)) * 0x94D049BB133111EBULL;
return z ^ (z >> 31);
}
/* Uniform double in (0,1]. */
static inline double sm_uniform(uint64_t* s){
/* 53-bit mantissa; +1 keeps it in (0,1] so log() never sees 0. */
return ((double)((splitmix64(s) >> 11) + 1)) * (1.0 / 9007199254740993.0);
}
/* ── element + index structures ───────────────────────────────────────────── */
typedef struct {
int count;
int cap;
int* ids; /* neighbour element indices */
} NeighList;
typedef struct {
uint64_t node_id;
int level; /* top layer this element appears on (>=0) */
float* vec; /* dim floats, L2-normalised */
NeighList* links; /* level+1 lists; links[l] = neighbours at layer l */
} Elem;
struct VIndex {
int dim;
int M; /* max neighbours per node, upper layers */
int M0; /* == 2*M, layer 0 */
int ef_construction;
double mL; /* level normaliser = 1/ln(M) */
Elem* elems;
size_t n;
size_t cap;
int entry; /* entry-point element index, -1 if empty */
int max_level; /* current top layer */
/* scratch: version-stamped visited set (O(1) reset). */
uint32_t* visited;
uint32_t visit_epoch;
size_t visited_cap;
};
/* ── small helpers ────────────────────────────────────────────────────────── */
static float* vec_normalise_copy(const float* v, int dim){
float* out = (float*)malloc((size_t)dim * sizeof(float));
if (!out) return NULL;
double ss = 0.0;
for (int i=0;i<dim;i++) ss += (double)v[i]*(double)v[i];
if (ss > 0.0){
float inv = (float)(1.0 / sqrt(ss));
for (int i=0;i<dim;i++) out[i] = v[i]*inv;
} else {
for (int i=0;i<dim;i++) out[i] = 0.0f; /* zero vector stays zero */
}
return out;
}
/* Cosine distance between two normalised vectors: 1 - dot. In [0,2].
* Float accumulation in 4 lanes so the compiler auto-vectorises the hot path
* (this is the dominant cost of both build and search). */
static float vdist(const VIndex* ix, const float* a, const float* b){
int dim = ix->dim;
float s0=0,s1=0,s2=0,s3=0;
int i=0;
for (; i+4<=dim; i+=4){
s0 += a[i]*b[i]; s1 += a[i+1]*b[i+1];
s2 += a[i+2]*b[i+2]; s3 += a[i+3]*b[i+3];
}
float dot = (s0+s1)+(s2+s3);
for (; i<dim; i++) dot += a[i]*b[i];
return 1.0f - dot;
}
static int nl_push(NeighList* nl, int id){
if (nl->count == nl->cap){
int nc = nl->cap ? nl->cap*2 : 4;
int* np = (int*)realloc(nl->ids, (size_t)nc*sizeof(int));
if (!np) return -1;
nl->ids = np; nl->cap = nc;
}
nl->ids[nl->count++] = id;
return 0;
}
/* ── binary heaps over (dist,elem) pairs ──────────────────────────────────── */
typedef struct { float d; int e; } Pair;
typedef struct { Pair* a; int n, cap; } Heap;
static int heap_reserve(Heap* h, int need){
if (need <= h->cap) return 0;
int nc = h->cap ? h->cap*2 : 16;
while (nc < need) nc *= 2;
Pair* na = (Pair*)realloc(h->a, (size_t)nc*sizeof(Pair));
if (!na) return -1;
h->a = na; h->cap = nc; return 0;
}
/* Order predicate: for a MAX-heap on distance, "higher priority" = larger dist;
* ties broken by larger element index (deterministic + stable). is_max selects. */
static inline int pair_before(Pair x, Pair y, int is_max){
if (x.d != y.d) return is_max ? (x.d > y.d) : (x.d < y.d);
return is_max ? (x.e > y.e) : (x.e < y.e);
}
static int heap_push(Heap* h, Pair v, int is_max){
if (heap_reserve(h, h->n+1)) return -1;
int i = h->n++;
h->a[i] = v;
while (i > 0){
int p = (i-1)/2;
if (pair_before(h->a[i], h->a[p], is_max)){
Pair t=h->a[i]; h->a[i]=h->a[p]; h->a[p]=t; i=p;
} else break;
}
return 0;
}
static Pair heap_pop(Heap* h, int is_max){
Pair top = h->a[0];
h->a[0] = h->a[--h->n];
int i = 0;
for (;;){
int l=2*i+1, r=2*i+2, best=i;
if (l<h->n && pair_before(h->a[l], h->a[best], is_max)) best=l;
if (r<h->n && pair_before(h->a[r], h->a[best], is_max)) best=r;
if (best==i) break;
Pair t=h->a[i]; h->a[i]=h->a[best]; h->a[best]=t; i=best;
}
return top;
}
/* ── visited set ──────────────────────────────────────────────────────────── */
static int visited_ensure(VIndex* ix){
if (ix->visited_cap >= ix->cap && ix->visited) return 0;
size_t nc = ix->cap ? ix->cap : 16;
uint32_t* nv = (uint32_t*)realloc(ix->visited, nc*sizeof(uint32_t));
if (!nv) return -1;
if (nc > ix->visited_cap) memset(nv + ix->visited_cap, 0, (nc-ix->visited_cap)*sizeof(uint32_t));
ix->visited = nv; ix->visited_cap = nc;
return 0;
}
static inline void visited_reset(VIndex* ix){
if (++ix->visit_epoch == 0){ /* wrapped: clear all */
memset(ix->visited, 0, ix->visited_cap*sizeof(uint32_t));
ix->visit_epoch = 1;
}
}
static inline int is_visited(VIndex* ix, int e){ return ix->visited[e]==ix->visit_epoch; }
static inline void mark_visited(VIndex* ix, int e){ ix->visited[e]=ix->visit_epoch; }
/* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */
/* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller
* owns res->a. `q` is a normalised query. */
static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
int ef, int layer, Heap* res /*out, max-heap*/){
Heap cand = {0,0,0}; /* min-heap: nearest to expand */
res->a=NULL; res->n=0; res->cap=0;
visited_reset(ix);
for (int i=0;i<neps;i++){
int e = eps[i];
if (is_visited(ix,e)) continue;
mark_visited(ix,e);
float d = vdist(ix, q, ix->elems[e].vec);
Pair p = { d, e };
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
}
while (res->n > ef) heap_pop(res,1); /* trim to ef */
while (cand.n > 0){
Pair c = heap_pop(&cand,0);
float worst = res->a[0].d; /* farthest kept result */
if (res->n >= ef && c.d > worst) break;
Elem* ce = &ix->elems[c.e];
if (layer <= ce->level){
NeighList* nl = &ce->links[layer];
for (int i=0;i<nl->count;i++){
int e = nl->ids[i];
if (is_visited(ix,e)) continue;
mark_visited(ix,e);
float d = vdist(ix, q, ix->elems[e].vec);
if (res->n < ef || d < res->a[0].d){
Pair p = { d, e };
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
if (res->n > ef) heap_pop(res,1);
}
}
}
}
free(cand.a);
return 0;
}
/* ── neighbour selection heuristic (Algorithm 4) ──────────────────────────── */
/* From candidate pairs W (any order), pick up to M diverse neighbours of q.
* Keep c only if it is nearer to q than to every already-chosen neighbour;
* backfill from the pruned set (nearest first) to reach M for connectivity.
* Writes chosen element indices into out[], returns the count. */
static int select_neighbors(VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
(void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */
/* sort W ascending by (dist,elem) — deterministic. */
for (int i=1;i<nW;i++){ /* insertion sort (nW small) */
Pair key=W[i]; int j=i-1;
while (j>=0 && !pair_before(W[j],key,0)){ W[j+1]=W[j]; j--; }
W[j+1]=key;
}
int nout = 0;
Pair* pruned = (Pair*)malloc((size_t)(nW?nW:1)*sizeof(Pair));
int npr = 0;
if (!pruned) return -1;
for (int i=0;i<nW && nout<M;i++){
int good = 1;
for (int j=0;j<nout;j++){
float d = vdist(ix, ix->elems[W[i].e].vec, ix->elems[out[j]].vec);
if (d < W[i].d){ good = 0; break; } /* nearer an existing pick → drop */
}
if (good) out[nout++] = W[i].e;
else pruned[npr++] = W[i];
}
for (int i=0;i<npr && nout<M;i++) out[nout++] = pruned[i].e; /* keep-pruned backfill */
free(pruned);
return nout;
}
/* Re-prune a neighbour's over-full adjacency list back to `Mmax`. */
static void prune_links(VIndex* ix, int e, int layer, int Mmax){
NeighList* nl = &ix->elems[e].links[layer];
if (nl->count <= Mmax) return;
const float* base = ix->elems[e].vec;
Pair* W = (Pair*)malloc((size_t)nl->count*sizeof(Pair));
if (!W) return;
int nW = nl->count;
for (int i=0;i<nW;i++) W[i] = (Pair){ vdist(ix, base, ix->elems[nl->ids[i]].vec), nl->ids[i] };
int* keep = (int*)malloc((size_t)nW*sizeof(int));
if (!keep){ free(W); return; }
int nk = select_neighbors(ix, base, W, nW, Mmax, keep);
if (nk >= 0){ nl->count = nk; for (int i=0;i<nk;i++) nl->ids[i]=keep[i]; }
free(keep); free(W);
}
/* ── insert ───────────────────────────────────────────────────────────────── */
static int elems_reserve(VIndex* ix){
if (ix->n < ix->cap) return 0;
size_t nc = ix->cap ? ix->cap*2 : 64;
Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem));
if (!ne) return -1;
ix->elems = ne; ix->cap = nc;
return visited_ensure(ix);
}
int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
if (!ix || !vec) return -1;
if (elems_reserve(ix)) return -1;
int cur = (int)ix->n;
/* deterministic level assignment, seeded per-node. */
uint64_t seed = VINDEX_FIXED_SEED ^ (node_id + 0x2545F4914F6CDD1DULL*(uint64_t)cur);
int level = (int)(-log(sm_uniform(&seed)) * ix->mL);
if (level < 0) level = 0;
Elem* el = &ix->elems[cur];
el->node_id = node_id;
el->level = level;
el->vec = vec_normalise_copy(vec, ix->dim);
el->links = (NeighList*)calloc((size_t)level+1, sizeof(NeighList));
if (!el->vec || !el->links){ free(el->vec); free(el->links); return -1; }
ix->n++;
if (ix->entry < 0){ /* first element */
ix->entry = cur; ix->max_level = level;
return 0;
}
int ep = ix->entry;
int L = ix->max_level;
/* greedy descent through layers above `level` to refine the entry point. */
for (int lc = L; lc > level; lc--){
Heap r = {0,0,0};
int eps1[1] = { ep };
if (search_layer(ix, el->vec, eps1, 1, 1, lc, &r)){ return -1; }
if (r.n){ ep = r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} }
free(r.a);
}
/* from min(L,level) down to 0: connect. Each layer's ef-results seed the next
* layer's entry set; `eps` is heap-owned below the top and freed each step. */
int start = (L < level) ? L : level;
int eps_stack[1] = { ep };
int* eps = eps_stack; /* not owned (stack) until reassigned to malloc'd */
int* eps_owned = NULL;
int neps = 1;
int rc = 0;
for (int lc = start; lc >= 0; lc--){
int Mmax = (lc==0) ? ix->M0 : ix->M;
Heap W = {0,0,0};
if (search_layer(ix, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
if (!chosen){ free(W.a); rc=-1; break; }
int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen);
if (nc < 0){ free(chosen); free(W.a); rc=-1; break; }
/* link cur <-> chosen (bidirectional), prune neighbours if over-full. */
for (int i=0;i<nc;i++){
int nb = chosen[i];
if (nl_push(&el->links[lc], nb) || nl_push(&ix->elems[nb].links[lc], cur)){
free(chosen); free(W.a); rc=-1; goto done;
}
prune_links(ix, nb, lc, Mmax);
}
free(chosen);
/* next layer's entry points = this layer's ef results. */
if (lc > 0){
int* neweps = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
if (!neweps){ free(W.a); rc=-1; break; }
for (int i=0;i<W.n;i++) neweps[i]=W.a[i].e;
neps = W.n ? W.n : 1;
if (!W.n) neweps[0] = eps[0]; /* fall back to prior ep if empty */
free(eps_owned);
eps = eps_owned = neweps;
}
free(W.a);
}
done:
free(eps_owned);
if (rc) return -1;
if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; }
return 0;
}
/* ── search ───────────────────────────────────────────────────────────────── */
int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out){
if (!ix || !query || k <= 0) return -1;
if (ix->entry < 0) return 0;
if (ef_search <= 0) ef_search = VINDEX_DEFAULT_EF_SEARCH;
if (ef_search < k) ef_search = k;
float* q = vec_normalise_copy(query, ix->dim);
if (!q) return -1;
int ep = ix->entry;
for (int lc = ix->max_level; lc > 0; lc--){
Heap r = {0,0,0};
int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, 1, lc, &r)){ free(q); return -1; }
if (r.n){ int b=r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;}
ep = b; }
free(r.a);
}
Heap res = {0,0,0};
int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, ef_search, 0, &res)){ free(res.a); free(q); return -1; }
free(q);
/* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */
int total = res.n;
Pair* sorted = (Pair*)malloc((size_t)(total?total:1)*sizeof(Pair));
if (!sorted){ free(res.a); return -1; }
for (int i=total-1;i>=0;i--) sorted[i] = heap_pop(&res,1); /* farthest first out → fill from end */
free(res.a);
int out_n = (k < total) ? k : total;
for (int i=0;i<out_n;i++){
if (node_id_out) node_id_out[i] = ix->elems[sorted[i].e].node_id;
if (dist_out) dist_out[i] = sorted[i].d;
}
free(sorted);
return out_n;
}
size_t vindex_size(const VIndex* ix){ return ix ? ix->n : 0; }
VIndex* vindex_create(int dim, int M, int ef_construction){
if (dim <= 0) return NULL;
if (M <= 0) M = VINDEX_DEFAULT_M;
if (ef_construction <= 0) ef_construction = VINDEX_DEFAULT_EF_CONSTRUCTION;
VIndex* ix = (VIndex*)calloc(1, sizeof(VIndex));
if (!ix) return NULL;
ix->dim = dim;
ix->M = M;
ix->M0 = 2*M;
ix->ef_construction = ef_construction;
ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0);
ix->entry = -1;
ix->max_level = 0;
ix->visit_epoch = 0;
return ix;
}
void vindex_free(VIndex* ix){
if (!ix) return;
for (size_t i=0;i<ix->n;i++){
Elem* e = &ix->elems[i];
if (e->links) for (int l=0;l<=e->level;l++) free(e->links[l].ids);
free(e->links);
free(e->vec);
}
free(ix->elems);
free(ix->visited);
free(ix);
}
/* ── read-only decode of the paged store node format (design §2.4) ─────────── */
/* Mirrors engram_store.c constants; the on-disk format is PERMANENT so these are
* safe to duplicate for a read-only harvest of emb vectors. */
#define VS_PAGE_SIZE 16384u
#define VS_HDR 32u
#define VS_SLOT_SIZE 6u
#define VS_SLOT_LIVE 1u
#define VS_REC_HDR 4u
#define VS_REC_OVERFLOW 1u
#define VS_PT_NODE 1u
#define VS_OVF_NEXT 32u
#define VS_OVF_LEN 40u
#define VS_OVF_DATA 44u
#define VS_NT_ID 1u
#define VS_NT_EMB 24u
#define VS_NT_EMB_DIM 25u
static uint16_t vg_u16(const uint8_t* p){ return (uint16_t)(p[0] | (p[1]<<8)); }
static uint32_t vg_u32(const uint8_t* p){ uint32_t v=0; for(int i=0;i<4;i++) v|=(uint32_t)p[i]<<(8*i); return v; }
static uint64_t vg_u64(const uint8_t* p){ uint64_t v=0; for(int i=0;i<8;i++) v|=(uint64_t)p[i]<<(8*i); return v; }
static int vs_pread(int fd, uint64_t page, uint8_t* buf){
off_t off = (off_t)page * VS_PAGE_SIZE;
ssize_t r = pread(fd, buf, VS_PAGE_SIZE, off);
return (r == (ssize_t)VS_PAGE_SIZE) ? 0 : -1;
}
/* Read a (possibly overflowed) record body; caller frees *out. */
static int vs_read_body(int fd, const uint8_t* page, uint16_t off, uint16_t len,
uint8_t** out, size_t* outlen){
if (len < VS_REC_HDR) return -1;
uint8_t flags = page[off+3];
if (flags & VS_REC_OVERFLOW){
uint64_t head = vg_u64(page + off + VS_REC_HDR);
uint64_t total = vg_u64(page + off + VS_REC_HDR + 8);
uint8_t* body = (uint8_t*)malloc(total ? total : 1);
if (!body) return -1;
size_t got=0; uint64_t id=head;
uint8_t ov[VS_PAGE_SIZE];
while (id){
if (vs_pread(fd, id, ov)){ free(body); return -1; }
uint32_t chunk = vg_u32(ov + VS_OVF_LEN);
if (got + chunk > total){ free(body); return -1; }
memcpy(body+got, ov+VS_OVF_DATA, chunk); got += chunk;
id = vg_u64(ov + VS_OVF_NEXT);
}
if (got != total){ free(body); return -1; }
*out = body; *outlen = total;
} else {
uint16_t reclen = vg_u16(page + off);
if (reclen < VS_REC_HDR) return -1;
size_t blen = reclen - VS_REC_HDR;
uint8_t* body = (uint8_t*)malloc(blen ? blen : 1);
if (!body) return -1;
memcpy(body, page + off + VS_REC_HDR, blen);
*out = body; *outlen = blen;
}
return 0;
}
/* Extract id (strdup) and emb (malloc'd float[dim]) from a TLV node body. */
static void vs_parse_node(const uint8_t* body, size_t len, char** id_out,
float** emb_out, int* dim_out){
*id_out=NULL; *emb_out=NULL; *dim_out=0;
size_t i=0;
while (i + 5 <= len){
uint8_t tag = body[i];
uint32_t flen = vg_u32(body + i + 1);
if (i + 5 + (size_t)flen > len) break;
const uint8_t* v = body + i + 5;
if (tag == VS_NT_ID){
char* s = (char*)malloc(flen+1);
if (s){ memcpy(s,v,flen); s[flen]=0; free(*id_out); *id_out=s; }
} else if (tag == VS_NT_EMB){
int dim = (int)(flen/4);
float* e = (float*)malloc((size_t)(dim?dim:1)*sizeof(float));
if (e){ for (int k=0;k<dim;k++){ uint32_t u=vg_u32(v+k*4); memcpy(&e[k],&u,4);}
free(*emb_out); *emb_out=e; if(*dim_out==0) *dim_out=dim; }
} else if (tag == VS_NT_EMB_DIM){
*dim_out = (int)vg_u32(v);
}
i += 5 + flen;
}
}
/* Tiny open-addressing string set to dedup ids across live records. */
typedef struct { char** k; size_t cap, n; } StrSet;
static uint64_t vs_fnv(const char* s){ uint64_t h=1469598103934665603ULL; for(;*s;++s){h^=(uint8_t)*s;h*=1099511628211ULL;} return h; }
static int strset_add(StrSet* s, const char* key){ /* 1 added, 0 dup, -1 err */
if (s->n*2 >= s->cap){
size_t nc = s->cap ? s->cap*2 : 1024;
char** nk = (char**)calloc(nc, sizeof(char*));
if (!nk) return -1;
for (size_t i=0;i<s->cap;i++) if (s->k[i]){ size_t j=vs_fnv(s->k[i])&(nc-1); while(nk[j]) j=(j+1)&(nc-1); nk[j]=s->k[i]; }
free(s->k); s->k=nk; s->cap=nc;
}
size_t j = vs_fnv(key)&(s->cap-1);
while (s->k[j]){ if (strcmp(s->k[j],key)==0) return 0; j=(j+1)&(s->cap-1); }
char* d = strdup(key); if(!d) return -1;
s->k[j]=d; s->n++;
return 1;
}
static void strset_free(StrSet* s){ for(size_t i=0;i<s->cap;i++) free(s->k[i]); free(s->k); }
int vindex_build_from_store(VIndex* ix, const char* store_path,
char*** ids_out, int* n_out){
if (!ix || !store_path) return -1;
int fd = open(store_path, O_RDONLY);
if (fd < 0) return -1;
struct stat st;
if (fstat(fd, &st) != 0){ close(fd); return -1; }
uint64_t npages = (uint64_t)st.st_size / VS_PAGE_SIZE;
char** ids = NULL; size_t ids_n = 0, ids_cap = 0;
StrSet seen = {0,0,0};
int inserted = 0;
uint8_t page[VS_PAGE_SIZE];
for (uint64_t pg = 2; pg < npages; pg++){ /* pages 0,1 = superblocks */
if (vs_pread(fd, pg, page)) continue;
if (page[8] != VS_PT_NODE) continue;
int slots = vg_u16(page + 10);
for (int sidx=0; sidx<slots; sidx++){
const uint8_t* sp = page + VS_HDR + (size_t)sidx*VS_SLOT_SIZE;
uint16_t off = vg_u16(sp), len = vg_u16(sp+2), fl = vg_u16(sp+4);
if (fl != VS_SLOT_LIVE) continue;
if ((size_t)off + VS_REC_HDR > VS_PAGE_SIZE) continue;
uint8_t* body=NULL; size_t blen=0;
if (vs_read_body(fd, page, off, len, &body, &blen)) continue;
char* id=NULL; float* emb=NULL; int dim=0;
vs_parse_node(body, blen, &id, &emb, &dim);
free(body);
if (!id || !emb || dim != ix->dim){ free(id); free(emb); continue; }
int add = strset_add(&seen, id);
if (add <= 0){ free(id); free(emb); continue; } /* dup or err */
if (vindex_insert(ix, (uint64_t)inserted, emb) != 0){ free(id); free(emb); break; }
free(emb);
if (ids_n == ids_cap){
size_t nc = ids_cap ? ids_cap*2 : 256;
char** ni = (char**)realloc(ids, nc*sizeof(char*));
if (!ni){ free(id); break; }
ids = ni; ids_cap = nc;
}
ids[ids_n++] = id; /* transfers ownership */
inserted++;
}
}
close(fd);
strset_free(&seen);
if (ids_out){ *ids_out = ids; if (n_out) *n_out = (int)ids_n; }
else { for (size_t i=0;i<ids_n;i++) free(ids[i]); free(ids); if (n_out) *n_out=(int)ids_n; }
return inserted;
}
/* ── optional persistence (index is rebuildable; convenience only) ─────────── */
#define VINDEX_SAVE_MAGIC "EGVIDX01"
int vindex_save(const VIndex* ix, const char* path){
if (!ix || !path) return -1;
FILE* f = fopen(path, "wb");
if (!f) return -1;
int ok = 1;
#define WR(p,n) do{ if(fwrite((p),1,(n),f)!=(size_t)(n)) ok=0; }while(0)
WR(VINDEX_SAVE_MAGIC, 8);
int32_t hdr[6] = { ix->dim, ix->M, ix->ef_construction, (int32_t)ix->n, ix->entry, ix->max_level };
WR(hdr, sizeof(hdr));
for (size_t i=0; ok && i<ix->n; i++){
Elem* e = &ix->elems[i];
WR(&e->node_id, sizeof(uint64_t));
int32_t lvl = e->level; WR(&lvl, sizeof(int32_t));
WR(e->vec, (size_t)ix->dim*sizeof(float));
for (int l=0; ok && l<=e->level; l++){
int32_t c = e->links[l].count; WR(&c, sizeof(int32_t));
WR(e->links[l].ids, (size_t)c*sizeof(int));
}
}
#undef WR
fclose(f);
return ok ? 0 : -1;
}
VIndex* vindex_load(const char* path){
FILE* f = fopen(path, "rb");
if (!f) return NULL;
char magic[8];
if (fread(magic,1,8,f)!=8 || memcmp(magic,VINDEX_SAVE_MAGIC,8)!=0){ fclose(f); return NULL; }
int32_t hdr[6];
if (fread(hdr,sizeof(hdr),1,f)!=1){ fclose(f); return NULL; }
VIndex* ix = vindex_create(hdr[0], hdr[1], hdr[2]);
if (!ix){ fclose(f); return NULL; }
size_t N = (size_t)hdr[3];
int ok = 1;
for (size_t i=0; ok && i<N; i++){
if (elems_reserve(ix)){ ok=0; break; }
Elem* e = &ix->elems[ix->n];
int32_t lvl;
if (fread(&e->node_id,sizeof(uint64_t),1,f)!=1 || fread(&lvl,sizeof(int32_t),1,f)!=1){ ok=0; break; }
e->level = lvl;
e->vec = (float*)malloc((size_t)ix->dim*sizeof(float));
e->links = (NeighList*)calloc((size_t)lvl+1, sizeof(NeighList));
if (!e->vec || !e->links){ free(e->vec); free(e->links); ok=0; break; }
if (fread(e->vec,sizeof(float),(size_t)ix->dim,f)!=(size_t)ix->dim){ ok=0; }
for (int l=0; ok && l<=lvl; l++){
int32_t c; if (fread(&c,sizeof(int32_t),1,f)!=1){ ok=0; break; }
e->links[l].ids = (int*)malloc((size_t)(c?c:1)*sizeof(int));
e->links[l].cap = c; e->links[l].count = c;
if (c && fread(e->links[l].ids,sizeof(int),(size_t)c,f)!=(size_t)c){ ok=0; }
}
ix->n++;
}
ix->entry = hdr[4]; ix->max_level = hdr[5];
fclose(f);
if (!ok){ vindex_free(ix); return NULL; }
return ix;
}
+82
View File
@@ -0,0 +1,82 @@
/* engram_vindex.h — M8 of the engram query engine: an approximate-nearest-
* neighbour (ANN) vector index over the node embedding vectors, for fast
* activation-seed selection.
*
* Replaces the O(n) cosine scan over emb vectors (design §9 M8; backlog #20)
* with an HNSW (Hierarchical Navigable Small World) graph that returns
* high-recall top-k seeds in ~O(log n).
*
* Standalone module: plain C11, stdlib + libm only. It does NOT modify the
* store format or engram_store.{c,h}; vindex_build_from_store() decodes the
* PERMANENT on-disk node format (design §2.4) read-only to harvest emb vectors.
*
* Similarity metric: cosine. Vectors are L2-normalised on insert/query, so
* cosine similarity == dot product. Reported distance = 1 - cosine_similarity
* (range [0,2]); smaller == closer. A query equal to an indexed vector scores
* distance ~0 against it.
*
* The index is fully rebuildable from the store, so persistence is optional for
* this milestone (see vindex_save/vindex_load below provided as a convenience;
* boot may simply rebuild via vindex_build_from_store()).
*/
#ifndef ENGRAM_VINDEX_H
#define ENGRAM_VINDEX_H
#include <stddef.h>
#include <stdint.h>
/* Tuned defaults (rationale in engram_vindex.c). Pass 0 to vindex_create for
* M / ef_construction to take these; pass ef_search<=0 to vindex_search for
* VINDEX_DEFAULT_EF_SEARCH. */
#define VINDEX_DEFAULT_M 24
#define VINDEX_DEFAULT_EF_CONSTRUCTION 200
#define VINDEX_DEFAULT_EF_SEARCH 128
typedef struct VIndex VIndex;
/* Create an index over `dim`-dimensional f32 vectors.
* M max neighbours per node on upper layers (2*M on layer 0).
* ef_construction candidate-list width during insert (recall/build cost).
* Pass M<=0 or ef_construction<=0 to use the VINDEX_DEFAULT_* above.
* Returns NULL on bad args / OOM. */
VIndex* vindex_create(int dim, int M, int ef_construction);
/* Insert one vector under an opaque caller-defined node_id (need not be unique,
* but the caller is responsible for meaning). `vec` has `dim` floats; it is
* copied and L2-normalised internally. A zero vector is accepted (it simply has
* distance ~1 to everything; never produces NaN). Returns 0 on success, <0 on
* error (bad args / OOM). */
int vindex_insert(VIndex* idx, uint64_t node_id, const float* vec);
/* Top-k search by cosine similarity. Writes up to k results (fewer if the index
* holds fewer than k elements) into node_id_out[] / dist_out[], ordered nearest
* first (ascending distance). Either out array may be NULL to skip it.
* ef_search search-time candidate width; larger == higher recall, slower.
* Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k.
* Returns the number of results written, or <0 on error. */
int vindex_search(VIndex* idx, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out);
/* Number of vectors currently indexed. */
size_t vindex_size(const VIndex* idx);
void vindex_free(VIndex* idx);
/* Build an index by scanning every live node record in the paged store at
* `store_path` (the on-disk format is decoded read-only; the store need not be
* open). Nodes without an emb vector, or whose emb_dim != idx->dim, are skipped.
* Each inserted node is assigned node_id = its 0-based insertion ordinal; if
* `ids_out`/`n_out` are non-NULL, *ids_out is set to a malloc'd array of that
* many strdup'd string ids (ids_out[node_id] == the store id) and *n_out to the
* count the caller frees each string and the array. Returns the number of
* vectors inserted, or <0 on error. */
int vindex_build_from_store(VIndex* idx, const char* store_path,
char*** ids_out, int* n_out);
/* Optional persistence (index is rebuildable from the store; provided for
* convenience). vindex_save writes a self-describing snapshot; vindex_load
* reconstructs an index from one. Return 0 / non-NULL on success. */
int vindex_save(const VIndex* idx, const char* path);
VIndex* vindex_load(const char* path);
#endif /* ENGRAM_VINDEX_H */