engram: reconcile M8 HNSW vindex (#109) onto current dev, restore 3 fixes the branch predated
El SDK CI - dev / build-and-test (pull_request) Failing after 4m49s

Lands feat/reframe-region-setop (PR #109: native set-based reframe_region,
decorator-as-seam @route port, teacher-summon, and the M8.1 activate-latency
work — lazy-memoized cosq via eg_cosq_at + engram_vindex HNSW-accelerated
seed discovery + vindex_harvest_from_store/vindex_bench oracle) onto dev's
actual current HEAD, plus engram-tiered-storage's still-unique test suite.

RECONCILING #109 WITH engram-tiered-storage (M4-M10 HNSW/geometry/reason/
verify work): not a two-way merge. engram_vindex.c's HNSW core (search_layer/
select_neighbors/prune_links/insert) is BYTE-IDENTICAL between the two
branches; #109's copy is a strict superset (adds vindex_harvest_from_store,
used by vindex_bench.c's brute-force-vs-HNSW oracle). engram_reason.c and
engram_verify.c are also byte-identical. #109's own branch point already
carried engram-tiered-storage's M4-M10 lineage forward, so there was nothing
left to merge into #109 for those files. The one thing engram-tiered-storage
had that #109's tree dropped: its full test suite (test_vindex.c,
test_geometry.c, test_reason.c, test_verify.c, test_m7_traversal.c, the
interoception P0-P5 tests, bufpool/compaction tests, and their run_*.sh
harnesses) — ported over here unchanged.

WHY THIS NEEDED HAND RECONCILIATION, NOT A MECHANICAL MERGE: #109's branch
forked from dev on 2026-08-14 15:40 (before restructure-adjacent history
diverged the file's merge-base for `git merge` — it presented as an add/add
conflict). A straight two-dot diff (dev tip -> PR tip) applied cleanly, but
it silently reverted THREE dev fixes landed on 2026-08-14/15, after the
branch point, that the PR's diff had no way to know about:

  1. qgate rescale (2026-08-14 self-review): PR's lazy eg_cosq_at rewrite of
     the query-aware propagation gate dropped the shift-and-floor rescale
     about ENGRAM_EMBED_S0 (measured: unrelated-pair median 0.562->raw gate
     0.67, i.e. "a small tax, not a gate"). Restored the rescale, wrapped
     around the lazy accessor -- the PR's actual improvement (WHEN cosq[oi]
     is computed) is orthogonal to WHAT it gates on and both are kept.
  2. Eviction cause decomposition (2026-08-14 self-review): dev decomposes
     wm_evicted into evict_floor/evict_cap/evict_bll so WM churn is
     diagnosable (identity: evicted == floor+cap+bll+dup_wm+dup_wm_global).
     PR's tree predates this and dropped all three counters + their JSON
     stats fields. Restored declarations, all 4 direct increment sites, the
     eg_wm_carry_over bll increment, and the act-stats JSON fields --
     alongside (not instead of) the PR's own P4 afferent / API-reshape
     counters already in that same struct/JSON.
  3. Hebbian link-formation selection (2026-08-15 self-review, TODAY): dev
     selects the STRONGEST qualifying candidate for consolidation each call;
     PR's tree predates this and reverted to hash-slot order (arbitrary wrt
     association strength) for edge formation -- the one path that writes
     PERMANENT structure. Restored the strongest-candidate while-loop,
     keeping the PR's own genuine improvement at that site
     (engram_adj_on_edge_added incremental-index append instead of a bare
     adj_dirty=1 full-rebuild flag).

engram/src/server.el's 3-way conflicts (autoconnect_on/ise_offgraph_on env
flags, /api/nodes connected-count in responses) were pure additive: dev's
side was empty, PR's side added the feature. Took PR's side whole.

VERIFIED (nsbx sandbox only, live :8742/:7770 never touched):
  - cc -std=c11 -O2, clean link against the real engram/src/server.el via
    elc, zero errors.
  - vindex_bench (built standalone, read-only harvest) against the real
    production store clone (13,671 embedded nodes, 768-dim nomic-embed-text):
    recall@10 = 1.0000 at ef 64/128/200; HNSW search 0.28-0.79ms/query vs
    2.03ms/query brute-force oracle (2.6x-7.2x). HNSW build itself: 46.5s
    for the full 13,671-node set -- see the flagged risk below.
  - Booted the reconciled binary in an isolated nsbx sandbox (:8905, cloned
    snapshot of the live store, 13,424 nodes / 37,656 edges) and called
    /api/activate for real: first call after boot 41.5s (pays the one-time
    HNSW build inline -- matches the standalone bench), second/third calls
    356ms/605ms, no crash, correct results, act-stats JSON (including the
    restored evict_floor/cap/bll fields) reads correctly.

KNOWN RISK TO FLAG BEFORE ANY LIVE CUTOVER (not fixed here; out of scope for
this dev-only land per instructions not to touch :8742/:7770): eg_vindex_sync
builds the HNSW index synchronously, inline, on the first engram_activate()
call after every process start (or index invalidation). On the real node
count that is a ~46s blocking stall on a single-threaded server -- the first
request after every restart (or its concurrent siblings) waits the full
build. Recommend a background/incremental build (or a bounded per-call build
budget) before this ever reaches the live daemon. See PR description / final
report for the fuller writeup.
This commit is contained in:
bigmerge
2026-08-15 16:46:44 -05:00
53 changed files with 13804 additions and 119 deletions
+31
View File
@@ -4,6 +4,37 @@ El is a self-hosting, statically-typed language that compiles to C. This file or
---
## Current work in this worktree — the API reshape / decorated seam (IN PROGRESS, 2026-08-14)
This is the `api-reshape` worktree. The build here reshapes Neuron's external
surface and how it is *declared* — proven on isolated dev-port clones only; **live
prod engram `:8742` is untouched and nothing is promoted.** Full framing lives in
`neuron/docs/architecture/06-cognitive-architecture.md` (Update — 2026-08-14 deep
night) and `02-components.md §5`.
- **Surface collapse.** The ~90 noun-organized CRUD MCP tools collapse to a few
**geometry ops**`read` (the *vantage-read*: re-origin + salience/recency +
an **aperture** → a bounded slice, curing the whole-self dump), `write`,
`relate`, `supersede` (evolve/tombstone/promote, never a hard delete) — plus the
agentic primitives `think`/`attend`/`learn`/`ground`/`assert`. The old noun is a
`type` parameter. Implemented in `tools/api-reshape/surface.el` with a parity
harness (`parity.sh`); aperture proven to bound output. **Not yet:** compiled
into the MCP server, hot-swap, all-alias dispatch.
- **Decorated seam.** `@route(path,method,…)` makes codegen synthesize
`el_route_dispatch` (replacing the hand-written `handle_request` if-else) —
proven decorate→serve on `:8951`. `@manager`/`@engine`/`@accessor` are **parsed
but structurally inert** in the shipped compiler today; the `@route` codegen
lives on the **unmerged branch `feat/el-route-decorators`**. Telemetry-emit and
dharma-bus auto-wiring at the boundary are **staged, not shipped**. In-process,
an `@accessor` reaches the engram via **`engram_*` builtins**, not `http_get`.
**Do not edit** the protected build sources while this is in flight:
`el-compiler/src/codegen.el`, `el-compiler/runtime/el_seed.c` (and the archived
`legacy/el_runtime.c`), the `runtime/engram_*.c` boot files, and `surface.el`
(when present in the reshape tree) — these are owned by the build agents.
---
## What El Is
El compiles `.el` source → C → native binary. Every El value is `el_val_t` (int64_t). Strings are heap pointers cast through int64_t. The compiler is written in El (self-hosting).
+303 -3
View File
@@ -3114,6 +3114,24 @@ fn build_int_names_for_params(params: [Map<String, Any>]) -> Bool {
return true
}
// fn_has_decorator does this FnDef carry a decorator named `name`?
// Reads the `decorators` list [{name, args}] attached by the parser. Absent
// key -> native_list_len returns 0 -> false. This is the multi-decorator-aware
// replacement for the old single `decorator` string check, so a fn may stack
// roles with other decorators (e.g. `@route(...) @manager fn ...`).
fn fn_has_decorator(stmt: Map<String, Any>, name: String) -> Bool {
let dl = stmt["decorators"]
let n: Int = native_list_len(dl)
let i = 0
while i < n {
let d = native_list_get(dl, i)
let dn: String = d["name"]
if str_eq(dn, name) { return true }
let i = i + 1
}
false
}
fn cg_fn(stmt: Map<String, Any>) -> Void {
let fn_name: String = stmt["name"]
// Skip El's `fn main()` - C provides its own main() for top-level stmts
@@ -3125,10 +3143,10 @@ fn cg_fn(stmt: Map<String, Any>) -> Void {
let params_c: String = params_to_c(params)
// VBD role enforcement: dharma_emit / dharma_field may only be called
// from @manager-decorated functions. Surface violations to the C compiler
// via #error directives emitted before the function definition.
let decorator: String = stmt["decorator"]
// via #error directives emitted before the function definition. Read the
// decorator LIST so the role may be stacked with other decorators.
if vbd_has_restricted_call(body) {
if !str_eq(decorator, "manager") {
if !fn_has_decorator(stmt, "manager") {
emit_line("#error \"VBD violation: dharma_emit/dharma_field called from non-@manager fn '" + fn_name + "'\"")
}
}
@@ -3136,6 +3154,15 @@ fn cg_fn(stmt: Map<String, Any>) -> Void {
// arithmetic vs concat on type-annotated identifiers.
build_int_names_for_params(params)
emit_line("el_val_t " + fn_name + "(" + params_c + ") {")
// API-reshape decorator-seam: auto-emit at the decorated-fn boundary
// Every @manager/@accessor fn gets ONE injected call to engram_boundary_beat
// at entry interoception (chrono tick) + telemetry (afferent counter) +
// strengthen (self-activity) + a dharma bus event so a decorated op
// self-reports with ZERO hand-written instrumentation in its body. (VBD role
// = the topmost decorator; write it topmost when stacking with @route.)
if fn_has_decorator(stmt, "manager") || fn_has_decorator(stmt, "accessor") {
emit_line(" engram_boundary_beat(EL_STR(" + c_str_lit(fn_name) + "));")
}
// Seed declared with parameter names so reassignment works
let decl = native_list_empty()
let np: Int = native_list_len(params)
@@ -3677,6 +3704,259 @@ fn cg_decl_streaming(stmt: Map<String, Any>) -> Void {
}
}
// @route dispatcher generation
//
// Scan the token stream for @route-decorated fns and synthesize a generic HTTP
// dispatcher `el_route_dispatch(method, clean, path, body)`. A decorated handler
// must have the uniform signature (method, path, body) -> String. The dispatcher
// matches `clean` (the query-stripped path, supplied by the caller) against each
// route and calls the handler with the ORIGINAL `path` so query strings survive.
// Returns the sentinel "__EL_NO_ROUTE__" when nothing matches, so the caller may
// fall through to any remaining hand-written branches (mixed mode).
//
// Decorator grammar: @route(path, method, kind, suffix)
// path the match string (or the prefix, for compound)
// method "GET" | "POST" | ... ; a '|'-list like "GET|POST"; "ANY"/"" = no guard
// kind "exact" (default) | "prefix" | "suffix" | "compound"
// suffix for "compound": the required str_ends_with suffix
//
// The dispatch table is emitted SPECIFICITY-SORTED (most-specific first), NOT in
// source order, so overlapping prefixes (e.g. /api/x/search vs /api/x) never
// shadow each other regardless of how the handlers are written.
// split_pipe split "GET|POST" on '|' into ["GET","POST"]. Self-contained
// (no dependency on str_split runtime semantics).
fn split_pipe(s: String) -> [String] {
let out: [String] = native_list_empty()
let cur: String = ""
let n: Int = str_len(s)
let i: Int = 0
while i < n {
let ch: String = str_slice(s, i, i + 1)
if str_eq(ch, "|") {
let out = native_list_append(out, cur)
let cur = ""
} else {
let cur = cur + ch
}
let i = i + 1
}
let out = native_list_append(out, cur)
out
}
// route_make_record build a route record map from the @route decorator args.
fn route_make_record(fn_name: String, args: [String]) -> Map<String, Any> {
let na: Int = native_list_len(args)
let rpath: String = ""
if na >= 1 { let rpath = native_list_get(args, 0) }
let rmethod: String = "GET"
if na >= 2 { let rmethod = native_list_get(args, 1) }
let rkind: String = "exact"
if na >= 3 { let rkind = native_list_get(args, 2) }
let rsuffix: String = ""
if na >= 4 { let rsuffix = native_list_get(args, 3) }
{ "name": fn_name, "path": rpath, "method": rmethod, "kind": rkind, "suffix": rsuffix }
}
// route_spec_score higher = more specific = emitted earlier. Ordering:
// exact > compound > suffix > prefix; within a class, a longer path/suffix
// wins (so /api/x/search sorts before /api/x). Guarantees correct dispatch
// independent of source order.
fn route_spec_score(rec: Map<String, Any>) -> Int {
let kind: String = rec["kind"]
let path: String = rec["path"]
let suffix: String = rec["suffix"]
let plen: Int = str_len(path)
let slen: Int = str_len(suffix)
if str_eq(kind, "exact") { return 4000000 + plen }
if str_eq(kind, "compound") { return 3000000 + plen * 100 + slen }
if str_eq(kind, "suffix") { return 2000000 + slen }
return 1000000 + plen
}
// route_sort_desc selection sort of route records by descending specificity.
// N is small (routes per module), so O(n^2) is fine and keeps codegen simple.
fn route_sort_desc(recs: [Map<String, Any>]) -> [Map<String, Any>] {
let n: Int = native_list_len(recs)
let out: [Map<String, Any>] = native_list_empty()
let used: [Bool] = native_list_empty()
let u: Int = 0
while u < n {
let used = native_list_append(used, false)
let u = u + 1
}
let picked: Int = 0
while picked < n {
let best_i: Int = 0 - 1
let best_score: Int = 0 - 1
let i: Int = 0
while i < n {
let is_used: Bool = native_list_get(used, i)
if !is_used {
let sc: Int = route_spec_score(native_list_get(recs, i))
if sc > best_score {
let best_score = sc
let best_i = i
}
}
let i = i + 1
}
let out = native_list_append(out, native_list_get(recs, best_i))
// Rebuild `used` with best_i marked (runtime has no native_list_set).
let new_used: [Bool] = native_list_empty()
let j: Int = 0
while j < n {
if j == best_i {
let new_used = native_list_append(new_used, true)
} else {
let new_used = native_list_append(new_used, native_list_get(used, j))
}
let j = j + 1
}
let used = new_used
let picked = picked + 1
}
out
}
// scan_routes token-level scan collecting every @route-decorated fn as a
// route record. Runs once per module (like scan_fn_sigs) so the dispatcher can
// be synthesized in the streaming backend, which discards per-fn ASTs. Handles
// decorator STACKING: `@route(...) @manager fn` still records the route.
fn scan_routes(tokens: [Any]) -> [Map<String, Any>] {
let total: Int = native_list_len(tokens) / 2
let recs: [Map<String, Any>] = native_list_empty()
let has_pending: Bool = false
let pending_args: [String] = native_list_empty()
let pos: Int = 0
let going: Bool = true
while going {
if pos >= total {
let going = false
} else {
let k: String = tok_kind(tokens, pos)
if str_eq(k, "Eof") {
let going = false
} else {
if str_eq(k, "At") {
let dname: String = tok_value(tokens, pos + 1)
let p: Int = pos + 2
let args: [String] = native_list_empty()
let ka: String = tok_kind(tokens, p)
if str_eq(ka, "LParen") {
let p = p + 1
let running: Bool = true
while running {
let kd: String = tok_kind(tokens, p)
if str_eq(kd, "RParen") {
let running = false
} else {
if str_eq(kd, "Eof") {
let running = false
} else {
if str_eq(kd, "Str") {
let args = native_list_append(args, tok_value(tokens, p))
}
let p = p + 1
}
}
}
if str_eq(tok_kind(tokens, p), "RParen") { let p = p + 1 }
}
if str_eq(dname, "route") {
let has_pending = true
let pending_args = args
}
let pos = p
} else {
if str_eq(k, "Fn") {
let fname: String = tok_value(tokens, pos + 1)
if has_pending {
let recs = native_list_append(recs, route_make_record(fname, pending_args))
let has_pending = false
}
let pos = pos + 2
} else {
let pos = pos + 1
}
}
}
}
}
recs
}
// program_has_routes did scan_routes find any @route fn?
fn program_has_routes(recs: [Map<String, Any>]) -> Bool {
native_list_len(recs) > 0
}
// route_method_guard C boolean prefix guarding on HTTP method, or "" for none.
fn route_method_guard(method: String) -> String {
if str_eq(method, "") { return "" }
if str_eq(method, "ANY") { return "" }
if str_contains(method, "|") {
let parts: [String] = split_pipe(method)
let np: Int = native_list_len(parts)
let expr: String = ""
let i: Int = 0
while i < np {
let m: String = native_list_get(parts, i)
if str_eq(m, "") {
let i = i + 1
} else {
let piece: String = "str_eq(method, EL_STR(" + c_str_lit(m) + "))"
if str_eq(expr, "") {
let expr = piece
} else {
let expr = expr + " || " + piece
}
let i = i + 1
}
}
if str_eq(expr, "") { return "" }
return "(" + expr + ") && "
}
"str_eq(method, EL_STR(" + c_str_lit(method) + ")) && "
}
// route_match_expr C boolean matching `clean` against the route path/kind.
fn route_match_expr(kind: String, path: String, suffix: String) -> String {
if str_eq(kind, "prefix") {
return "str_starts_with(clean, EL_STR(" + c_str_lit(path) + "))"
}
if str_eq(kind, "suffix") {
return "str_ends_with(clean, EL_STR(" + c_str_lit(path) + "))"
}
if str_eq(kind, "compound") {
return "str_starts_with(clean, EL_STR(" + c_str_lit(path) + ")) && str_ends_with(clean, EL_STR(" + c_str_lit(suffix) + "))"
}
"str_eq(clean, EL_STR(" + c_str_lit(path) + "))"
}
// emit_route_dispatch emit the generated el_route_dispatch definition from the
// specificity-sorted route records. No-op if there are no routes.
fn emit_route_dispatch(recs: [Map<String, Any>]) -> Void {
if !program_has_routes(recs) { return }
let sorted: [Map<String, Any>] = route_sort_desc(recs)
emit_line("// ── generated @route dispatcher (specificity-sorted) ──")
emit_line("el_val_t el_route_dispatch(el_val_t method, el_val_t clean, el_val_t path, el_val_t body) {")
let n: Int = native_list_len(sorted)
let i: Int = 0
while i < n {
let rec = native_list_get(sorted, i)
let guard: String = route_method_guard(rec["method"])
let match_e: String = route_match_expr(rec["kind"], rec["path"], rec["suffix"])
let fn_name: String = rec["name"]
emit_line(" if (" + guard + match_e + ") { return " + fn_name + "(method, path, body); }")
let i = i + 1
}
emit_line(" return EL_STR(\"__EL_NO_ROUTE__\");")
emit_line("}")
emit_blank()
}
// emit_streaming_preamble emit #includes, forward decls, and file-scope lets
// using the pre-scanned signature data (no full AST).
fn emit_streaming_preamble(sigs: [Map<String, Any>], source: String) -> Void {
@@ -3769,6 +4049,17 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
emit_streaming_preamble(sigs, source)
el_arena_pop(preamble_mark)
// @route: scan the token stream once for @route-decorated fns. Kept in
// codegen_streaming scope (survives the per-fn arena pops and el_release of
// tokens below via refcount, like `sigs`). If any exist, forward-declare the
// generated dispatcher NOW so hand-written fns (e.g. handle_request) may call
// it before its definition is emitted after the fn-emit loop.
let route_records: [Map<String, Any>] = scan_routes(tokens)
if program_has_routes(route_records) {
emit_line("el_val_t el_route_dispatch(el_val_t method, el_val_t clean, el_val_t path, el_val_t body);")
emit_blank()
}
// Detect whether there is a fn main() and whether there are top-level
// executable stmts (for library detection) from sigs.
let has_el_main: Bool = false
@@ -3988,6 +4279,15 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
}
}
// @route: emit the generated dispatcher definition now after every handler
// fn has been emitted, but before `tokens` is released (route_records holds
// its own refs to the extracted strings). No-op unless the module declared
// at least one @route fn. Emitted before the test/library early-returns so it
// is present in library modules (e.g. neuron's routes.el) too.
let route_arena_mark: Any = el_arena_push()
emit_route_dispatch(route_records)
el_arena_pop(route_arena_mark)
// Tokens fully consumed by the streaming loop release now to free peak heap.
el_release(tokens)
+47 -2
View File
@@ -1782,23 +1782,68 @@ fn parse_stmt(tokens: [Any], pos: Int) -> Map<String, Any> {
return make_result({ "stmt": "TryCatch", "try_body": try_body, "catch_name": catch_name, "catch_body": native_list_empty() }, p)
}
// @decorator - capture decorator name and attach to following stmt
// @decorator - capture decorator name (and optional string args) and
// attach to the following stmt. Backward-compatible: bare @manager /
// @engine / @accessor still parse (no parens -> empty args). Decorators
// STACK: `@route("/p","GET") @manager fn f()` attaches BOTH to f via a
// `decorators` list [{name, args}]. The legacy `decorator` string is kept
// populated (topmost decorator) so the JS backend keeps working unchanged.
if k == "At" {
let p = pos + 1
let dec_name = tok_value(tokens, p)
let p = p + 1
// Optional decorator argument list: @name("a", "b", ...)
let dec_args = native_list_empty()
let ka = tok_kind(tokens, p)
if str_eq(ka, "LParen") {
let p = p + 1
let running_da = true
while running_da {
let kd = tok_kind(tokens, p)
if str_eq(kd, "RParen") {
let running_da = false
} else {
if str_eq(kd, "Eof") {
let running_da = false
} else {
if str_eq(kd, "Str") {
let dec_args = native_list_append(dec_args, tok_value(tokens, p))
}
let p = p + 1
let kc = tok_kind(tokens, p)
if str_eq(kc, "Comma") {
let p = p + 1
}
}
}
}
let p = expect(tokens, p, "RParen")
}
let r = parse_stmt(tokens, p)
let inner = r["node"]
let p2 = r["pos"]
let inner_kind: String = inner["stmt"]
if str_eq(inner_kind, "FnDef") {
// Stack this decorator (topmost-first) onto any decorators the inner
// FnDef already carries from decorators written below this one.
let this_dec = { "name": dec_name, "args": dec_args }
let existing = inner["decorators"]
let dlist = native_list_empty()
let dlist = native_list_append(dlist, this_dec)
let ne: Int = native_list_len(existing)
let ei = 0
while ei < ne {
let dlist = native_list_append(dlist, native_list_get(existing, ei))
let ei = ei + 1
}
let with_dec = {
"stmt": "FnDef",
"name": inner["name"],
"params": inner["params"],
"body": inner["body"],
"ret_type": inner["ret_type"],
"decorator": dec_name
"decorator": dec_name,
"decorators": dlist
}
// r result map fully consumed release to free peak heap.
el_release(r)
+2163 -48
View File
File diff suppressed because it is too large Load Diff
+62 -1
View File
@@ -46,7 +46,6 @@
#include <stdint.h>
#include <stdlib.h>
#include <math.h> /* fmod, sin, sqrt, ... — used by codegen'd float arithmetic */
typedef int64_t el_val_t;
@@ -620,14 +619,69 @@ el_val_t engram_store_close(void);
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_retrieve_geometric_json(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_emb_json(el_val_t limit, el_val_t offset);
el_val_t engram_dreams_json(el_val_t since_ms);
/* §5 geometry operators as EL builtins (read-only; seed-id CSV args). */
el_val_t engram_geo_descriptor_json(el_val_t seeds);
el_val_t engram_geo_overlap_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_geo_subtract_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t mode);
el_val_t engram_geo_combine_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_geo_distance_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_geo_analogy_json(el_val_t a_seeds, el_val_t b_seeds);
/* reasoning layer (compositions over §5 operators). ANALOGY maps cleanly to the
* flat-CSV seed ABI; the other modes take set/point/timestamp inputs deferred from
* this ABI (see engram_reason.h / the reasoning-operators runbook). */
el_val_t engram_reason_analogy_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t c_seeds);
/* COGNITION (2026-08-14): THE ONE OPERATION + grounding, surfaced live. */
el_val_t engram_think_json(el_val_t seeds, el_val_t faculty);
el_val_t engram_ground_json(el_val_t claim, el_val_t evidence, el_val_t for_whom);
el_val_t engram_assert_json(el_val_t claim_id, el_val_t for_whom, el_val_t floor);
el_val_t engram_attend_json(el_val_t node_id, el_val_t observer, el_val_t salience);
el_val_t engram_correspondence_beat_json(el_val_t seeds, el_val_t faculty, el_val_t keystone);
el_val_t engram_consolidate_permanence(el_val_t node_id);
el_val_t engram_age_field(el_val_t delta_ms);
el_val_t engram_age_field_catchup(void);
el_val_t engram_chrono_persist_tick(void);
el_val_t engram_chrono_tick(void);
el_val_t engram_boundary_beat(el_val_t op_name); /* API-reshape decorator-seam auto-emit */
el_val_t engram_self_anchor_capture(void);
el_val_t engram_self_drift_json(void);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_activate_json(el_val_t query, el_val_t depth);
el_val_t engram_stats_json(void);
el_val_t engram_act_stats_json(void);
el_val_t engram_text_health_json(void);
el_val_t engram_cosine_sim(el_val_t id_a, el_val_t id_b);
/* M10 reified-neighborhood read-only HTTP surface (2026-08-13). List/detail of
* the resident reify index; [] until the offline reify writer has run. */
el_val_t engram_geo_reify_list_json(void);
el_val_t engram_geo_reify_get_json(el_val_t id);
/* WRITE: run reification — persist Neighborhood nodes + member edges, rebuild the
* resident index. Wires the previously-dormant engram_geo_reify_store. */
el_val_t engram_geo_reify_run_json(void);
/* SELF-REIFICATION (2026-08-14). ON-BEAT autonomous neighborhood formation:
* gated by ENGRAM_SELF_REIFY (default off → returns {"enabled":false}, writes
* nothing — the live binary is byte-inert until the flag is set). When on, runs
* ONE bounded, incremental, idempotent reification pass (change-detection skips
* unchanged hubs → no re-append; grounded names; residue-preserving supersession;
* soft/overlapping membership; nests only when something changed). Meant to be
* pumped every heartbeat next to Hebbian consolidation. Returns
* {"enabled":true,"reified":N,"skipped":S,"superseded":P,"nested":X,"resident":M,"wrote":b}. */
el_val_t engram_self_reify_beat_json(void);
/* ASYNC EXPLICIT OVERRIDE (degenerate manual case). Rename a live neighborhood
* by id: writes a superseding record with the new name, prepends a residue entry
* (cause="explicit-override", prior name), keeps the same geometry/members. Never
* blocks the beat. Returns {"ok":true,"renamed":<old>,"new_id":<new>,"name":..}. */
el_val_t engram_neighborhood_rename_json(el_val_t id, el_val_t name);
/* Orphan prevention: form up to k semantic-similar edges to a node's nearest
* embedded neighbors (kNN). engram_nearest_json is the read-only probe. */
el_val_t engram_autoconnect_node(el_val_t id, el_val_t k, el_val_t min_sim_pct);
el_val_t engram_nearest_json(el_val_t id, el_val_t k);
/* Telemetry off-graph: append one ISE JSON line to the state-event log tier. */
el_val_t engram_ise_log_append(el_val_t content);
/* Destructively pop up to `max` newly-formed Hebbian associations as a JSON
* array of {from_id,to_id,weight,hebb}. The learning process (soul daemon) is
* not the process that owns persistence (engram HTTP server); this is how a
@@ -680,6 +734,13 @@ el_val_t engram_compile_layered_json(el_val_t intent, el_val_t depth);
el_val_t llm_call(el_val_t model, el_val_t prompt);
el_val_t llm_call_system(el_val_t model, el_val_t system_prompt, el_val_t user_prompt);
el_val_t llm_call_agentic(el_val_t model, el_val_t system, el_val_t user, el_val_t tools);
/* LLM token telemetry (CCR §4.4): usage.{input,output}_tokens parsed from every
* response. Returns Map{input_tokens,output_tokens,total_input_tokens,
* total_output_tokens,calls}. */
el_val_t llm_last_usage(void);
/* Fold usage.{input,output}_tokens from a raw messages response into the counters.
* Called internally on every LLM response; exported for offline testing. */
void llm_record_usage(const char* resp);
el_val_t llm_vision(el_val_t model, el_val_t system, el_val_t prompt, el_val_t image_url_or_b64);
el_val_t llm_models(void);
+339
View File
@@ -0,0 +1,339 @@
/* engram_cognition.c — THE ONE OPERATION. See engram_cognition.h.
* Pure over its inputs (think/warp/express); persistence is additive/supersede
* only. stdlib + libm + engram_store/reason/geometry. Touches no live daemon. */
#include "engram_cognition.h"
#include <stdlib.h>
#include <string.h>
#include <stdio.h>
#include <math.h>
/* ── small helpers ──────────────────────────────────────────────────────────── */
static double vdot(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) s += (double)a[i] * (double)b[i]; return s;
}
static double vnorm(const float* a, int dim) { return sqrt(vdot(a, a, dim)); }
static char* dupstr(const char* s) {
if (!s) return NULL; size_t n = strlen(s) + 1; char* p = malloc(n);
if (p) memcpy(p, s, n); return p;
}
static double clampd(double x, double lo, double hi){ return x<lo?lo:(x>hi?hi:x); }
/* ═══════════════════════════════════════════════ Stance lifecycle ════════════ */
int cog_stance_init(CogStance* s, const char* id, const char* faculty,
const char* anchor_region, const char* for_whom,
const GeoDescriptor* region) {
if (!s || !region) return -1;
memset(s, 0, sizeof *s);
s->id = dupstr(id); s->faculty = dupstr(faculty);
s->anchor_region = dupstr(anchor_region); s->for_whom = dupstr(for_whom);
s->dim = region->dim;
s->n_axes = region->n_axes > COG_MAX_AXES ? COG_MAX_AXES : region->n_axes;
for (int k = 0; k < COG_MAX_AXES; k++) s->axis_gain[k] = 1.0;
s->ext_floor = 1.0; s->drop_frac = 0.5; s->assoc_floor = 0.2;
s->bias_dir = NULL;
s->reliability = 0.5; /* uninformed prior on our own track record */
return 0;
}
void cog_stance_set_frozen_defaults(CogStance* s) {
if (!s) return;
for (int k = 0; k < COG_MAX_AXES; k++) s->axis_gain[k] = 1.0;
s->ext_floor = 1.0; s->drop_frac = 0.5; s->assoc_floor = 0.2;
free(s->bias_dir); s->bias_dir = NULL;
}
void cog_stance_free(CogStance* s) {
if (!s) return;
free(s->id); free(s->faculty); free(s->anchor_region); free(s->for_whom);
free(s->bias_dir);
s->id = s->faculty = s->anchor_region = s->for_whom = NULL; s->bias_dir = NULL;
}
int cog_is_keystone(const CogKeystoneSet* ks, const CogStance* s) {
if (!s) return 0;
if (s->keystone) return 1;
if (!ks || !s->id) return 0;
for (int i = 0; i < ks->n; i++)
if (ks->ids[i] && (strcmp(ks->ids[i], s->id) == 0 ||
(s->anchor_region && strcmp(ks->ids[i], s->anchor_region) == 0))) return 1;
return 0;
}
/* ═══════════════════════════════════════════════ warped fit (think step 2) ══ */
int cog_warped_fit(const GeoDescriptor* g, const float* x,
const CogStance* st, GeoFit* out) {
if (!g || !x || !out || g->dim <= 0 || !g->centroid) return -1;
double ext_floor = (st && st->ext_floor > 0) ? st->ext_floor : 1.0;
int dim = g->dim;
double rr = 0;
float* r = malloc((size_t)dim * sizeof(float));
if (!r) return -1;
for (int i = 0; i < dim; i++) { double d = (double)x[i] - (double)g->centroid[i]; r[i] = (float)d; rr += d * d; }
double maha2 = 0, ss_in = 0;
for (int k = 0; k < g->n_axes; k++) {
const float* ax = g->axes[k].axis; if (!ax) continue;
double proj = vdot(r, ax, dim);
double gain = (st && k < st->n_axes && st->axis_gain[k] > 0) ? st->axis_gain[k] : 1.0;
double den = g->axes[k].extent * gain; if (den < ext_floor) den = ext_floor;
maha2 += (proj / den) * (proj / den);
ss_in += proj * proj;
}
double ortho2 = rr - ss_in; if (ortho2 < 0) ortho2 = 0;
double dist2 = maha2 + ortho2 / (ext_floor * ext_floor);
out->mahalanobis = sqrt(maha2); out->ortho_residual = sqrt(ortho2);
out->distance = sqrt(dist2); out->score = 1.0 / (1.0 + dist2);
free(r);
return 0;
}
/* ═══════════════════════════════════════════════ think (the ONE operation) ══ */
void engram_gradient_free(GeoGradient* g) {
if (!g) return; free(g->direction); g->direction = NULL;
}
int engram_think(const GeoDescriptor* region, const float* anchor,
const CogStance* stance, GeoGradient* out) {
if (!region || !out || region->dim <= 0 || !region->centroid) return -1;
int dim = region->dim;
memset(out, 0, sizeof *out);
out->dim = dim;
const float* x = anchor ? anchor : region->centroid; /* re-origin (step 1) */
GeoFit f;
if (cog_warped_fit(region, x, stance, &f) != 0) return -1; /* fit (step 2) */
/* step 3 — emit a GRADIENT: warped steepest DESCENT of the fit distance². */
double ext_floor = (stance && stance->ext_floor > 0) ? stance->ext_floor : 1.0;
float* grad = calloc((size_t)dim, sizeof(float)); /* ∇ dist² wrt x */
float* r = malloc((size_t)dim * sizeof(float));
out->direction = malloc((size_t)dim * sizeof(float));
if (!grad || !r || !out->direction) { free(grad); free(r); free(out->direction); out->direction = NULL; return -1; }
for (int i = 0; i < dim; i++) r[i] = (float)((double)x[i] - (double)region->centroid[i]);
/* in-subspace: Σ_k 2 (proj/den²) a_k ; also accumulate Σ proj a_k for ortho part */
float* proj_sum = calloc((size_t)dim, sizeof(float));
if (!proj_sum) { free(grad); free(r); free(out->direction); out->direction = NULL; return -1; }
for (int k = 0; k < region->n_axes; k++) {
const float* ax = region->axes[k].axis; if (!ax) continue;
double proj = vdot(r, ax, dim);
double gain = (stance && k < stance->n_axes && stance->axis_gain[k] > 0) ? stance->axis_gain[k] : 1.0;
double den = region->axes[k].extent * gain; if (den < ext_floor) den = ext_floor;
double coef = 2.0 * proj / (den * den);
for (int i = 0; i < dim; i++) { grad[i] += (float)(coef * ax[i]); proj_sum[i] += (float)(proj * ax[i]); }
}
/* orthogonal: (2 r 2 Σ proj a_k) / ext_floor² */
double inv_f2 = 1.0 / (ext_floor * ext_floor);
for (int i = 0; i < dim; i++)
grad[i] += (float)((2.0 * (double)r[i] - 2.0 * (double)proj_sum[i]) * inv_f2);
free(proj_sum);
/* steering = grad (descent), seeded by the stance's bias_dir. */
for (int i = 0; i < dim; i++) out->direction[i] = -grad[i];
if (stance && stance->bias_dir) {
double gn = vnorm(grad, dim), bn = vnorm(stance->bias_dir, dim);
if (bn > 1e-12) {
double scale = (gn > 1e-12 ? gn : 1.0); /* seed at the gradient's scale */
for (int i = 0; i < dim; i++)
out->direction[i] += (float)(scale * (double)stance->bias_dir[i] / bn);
}
}
double dn = vnorm(out->direction, dim);
if (dn > 1e-12) for (int i = 0; i < dim; i++) out->direction[i] /= (float)dn;
else for (int i = 0; i < dim; i++) out->direction[i] = 0.0f; /* at rest */
out->spread = f.distance; /* spiked (0) .. diffuse */
out->confidence = stance ? stance->reliability : 0.5;
out->magnitude = f.score; /* the read's membership */
out->anchor_id = region->hub_id; /* borrowed vantage id */
out->n_support = region->n_members;
out->stance_id = stance ? stance->id : NULL;
free(grad); free(r);
return 0;
}
/* EXPRESSION — the ONLY collapse to a point (a separate faculty from think). */
int engram_express(const GeoGradient* g, const float* anchor, float* out_point) {
if (!g || !anchor || !out_point || !g->direction) return -1;
double commit = clampd(g->confidence, 0.0, 1.0); /* confident => commit far */
for (int i = 0; i < g->dim; i++)
out_point[i] = anchor[i] + g->direction[i] * (float)commit;
return 0;
}
/* ═══════════════════════════════════════════════ Stance serialization ════════ */
/* Compact line schema "STNC1" (mirrors the reify "GEO1" precedent). */
char* cog_stance_to_metadata(const CogStance* s) {
if (!s) return NULL;
size_t cap = 256 + (size_t)s->n_axes * 24 + (size_t)(s->bias_dir ? s->dim * 16 : 0);
char* buf = malloc(cap); if (!buf) return NULL;
size_t o = 0;
o += (size_t)snprintf(buf + o, cap - o, "%s\n", COG_STANCE_META_MAGIC);
o += (size_t)snprintf(buf + o, cap - o, "f %s\n", s->faculty ? s->faculty : "-");
o += (size_t)snprintf(buf + o, cap - o, "r %s\n", s->anchor_region ? s->anchor_region : "-");
o += (size_t)snprintf(buf + o, cap - o, "w %s\n", s->for_whom ? s->for_whom : "-");
o += (size_t)snprintf(buf + o, cap - o, "k %d\n", s->keystone);
o += (size_t)snprintf(buf + o, cap - o, "d %d %d\n", s->dim, s->n_axes);
o += (size_t)snprintf(buf + o, cap - o, "s %.9g %.9g %.9g\n", s->ext_floor, s->drop_frac, s->assoc_floor);
o += (size_t)snprintf(buf + o, cap - o, "g");
for (int k = 0; k < s->n_axes; k++) o += (size_t)snprintf(buf + o, cap - o, " %.9g", s->axis_gain[k]);
o += (size_t)snprintf(buf + o, cap - o, "\n");
o += (size_t)snprintf(buf + o, cap - o, "c %lld %.9g %.9g %.9g %.9g\n",
(long long)s->n_trials, s->brier_sum, s->reliability, s->ema_error, s->last_error);
if (s->bias_dir) {
o += (size_t)snprintf(buf + o, cap - o, "b");
for (int i = 0; i < s->dim; i++) o += (size_t)snprintf(buf + o, cap - o, " %.9g", (double)s->bias_dir[i]);
o += (size_t)snprintf(buf + o, cap - o, "\n");
}
(void)o;
return buf;
}
int cog_stance_to_node(const CogStance* s, StoreNode* out) {
if (!s || !out) return -1;
memset(out, 0, sizeof *out);
out->id = dupstr(s->id);
out->node_type = dupstr(COG_STANCE_NODE_TYPE);
out->content = dupstr(s->faculty ? s->faculty : "stance");
out->label = dupstr(s->faculty ? s->faculty : "stance");
out->metadata = cog_stance_to_metadata(s);
out->importance = s->reliability; /* cached denormalized readout (§2.1) */
out->confidence = s->reliability;
out->temporal_decay_rate = 0.0;
return (out->id && out->node_type && out->metadata) ? 0 : -1;
}
static int parse_floats(const char* line, double* out, int max) {
int n = 0; const char* p = line;
while (*p && n < max) {
while (*p == ' ') p++;
if (!*p) break;
char* end; double v = strtod(p, &end);
if (end == p) break;
out[n++] = v; p = end;
}
return n;
}
int cog_stance_from_node(const StoreNode* n, CogStance* out) {
if (!n || !out || !n->metadata) return -1;
memset(out, 0, sizeof *out);
for (int k = 0; k < COG_MAX_AXES; k++) out->axis_gain[k] = 1.0;
out->ext_floor = 1.0; out->drop_frac = 0.5; out->assoc_floor = 0.2; out->reliability = 0.5;
out->id = dupstr(n->id);
/* verify magic on first line */
const char* m = n->metadata;
if (strncmp(m, COG_STANCE_META_MAGIC, strlen(COG_STANCE_META_MAGIC)) != 0) return -1;
char* copy = dupstr(m); if (!copy) return -1;
for (char* line = strtok(copy, "\n"); line; line = strtok(NULL, "\n")) {
if (line[0] == '\0' || line[1] != ' ') {
if (line[0] == 'g' || line[0] == 'b') { /* vector lines: tag then values */ }
else continue;
}
char tag = line[0];
const char* rest = line + 1; while (*rest == ' ') rest++;
if (tag == 'f') { free(out->faculty); out->faculty = (strcmp(rest, "-") ? dupstr(rest) : NULL); }
else if (tag == 'r') { free(out->anchor_region); out->anchor_region = (strcmp(rest, "-") ? dupstr(rest) : NULL); }
else if (tag == 'w') { free(out->for_whom); out->for_whom = (strcmp(rest, "-") ? dupstr(rest) : NULL); }
else if (tag == 'k') { out->keystone = atoi(rest); }
else if (tag == 'd') { int a=0,b=0; sscanf(rest, "%d %d", &a, &b); out->dim = a; out->n_axes = b > COG_MAX_AXES ? COG_MAX_AXES : b; }
else if (tag == 's') { double v[3]={1,0.5,0.2}; parse_floats(rest, v, 3); out->ext_floor=v[0]; out->drop_frac=v[1]; out->assoc_floor=v[2]; }
else if (tag == 'g') { double v[COG_MAX_AXES]; int c=parse_floats(rest, v, COG_MAX_AXES); for(int k=0;k<c;k++) out->axis_gain[k]=v[k]; }
else if (tag == 'c') { double v[5]={0,0,0.5,0,0}; parse_floats(rest, v, 5); out->n_trials=(int64_t)v[0]; out->brier_sum=v[1]; out->reliability=v[2]; out->ema_error=v[3]; out->last_error=v[4]; }
else if (tag == 'b') { if (out->dim>0){ out->bias_dir=calloc((size_t)out->dim,sizeof(float)); double v[4096]; int c=parse_floats(rest,v,out->dim<4096?out->dim:4096); for(int i=0;i<c;i++) out->bias_dir[i]=(float)v[i]; } }
}
free(copy);
return 0;
}
/* ═══════════════════════════════════════════════ grounding as a RELATION ═════ */
static int put_edge(EngramPagedStore* s, const char* id, const char* from, const char* to,
const char* relation, double weight, const char* meta) {
StoreEdge e; memset(&e, 0, sizeof e);
e.id = (char*)id; e.from_id = (char*)from; e.to_id = (char*)to;
e.relation = (char*)relation; e.weight = weight; e.confidence = weight;
e.metadata = (char*)meta;
return store_put_edge(s, &e);
}
int cog_ground_edge(EngramPagedStore* s, const char* claim_id,
const char* evidence_id, double grounding, const char* for_whom) {
if (!s || !claim_id || !evidence_id) return -1;
char id[512], meta[256];
snprintf(id, sizeof id, "gb-%s-%s-%s", claim_id, evidence_id, for_whom ? for_whom : "global");
snprintf(meta, sizeof meta, "for_whom=%s", for_whom ? for_whom : "-");
return put_edge(s, id, claim_id, evidence_id, COG_GROUNDED_BY_RELATION, grounding, meta);
}
int cog_salient_edge(EngramPagedStore* s, const char* node_id,
const char* observer_id, double salience) {
if (!s || !node_id || !observer_id) return -1;
char id[512];
snprintf(id, sizeof id, "st-%s-%s", node_id, observer_id);
return put_edge(s, id, node_id, observer_id, COG_SALIENT_TO_RELATION, salience, NULL);
}
int cog_assert_gate(EngramPagedStore* s, const char* claim_id,
const char* for_whom, double floor) {
if (!s || !claim_id) return -1;
if (!(floor > 0)) floor = 0.5;
StoreEdge* edges = NULL; size_t n = 0;
if (store_get_edges_from(s, claim_id, &edges, &n) < 0) return -1;
double best = 0.0; int found = 0;
for (size_t i = 0; i < n; i++) {
if (!edges[i].relation || strcmp(edges[i].relation, COG_GROUNDED_BY_RELATION) != 0) continue;
/* grounded-for-whom: match observer if requested; global (for_whom=-) always counts */
int match = 1;
if (for_whom && edges[i].metadata) {
const char* fw = strstr(edges[i].metadata, "for_whom=");
if (fw) { fw += 9; if (strcmp(fw, for_whom) != 0 && strcmp(fw, "-") != 0) match = 0; }
}
if (match) { found = 1; if (edges[i].weight > best) best = edges[i].weight; }
}
store_edges_free(edges, n);
if (!found) return 0; /* ungrounded => refuse assertion (still held) */
return (best >= floor) ? 1 : 0;
}
/* ═══════════════════════════════════════════════ THE CORRESPONDENCE-LOOP ═════ */
int engram_correspondence_beat(const GeoDescriptor* region, const float* anchor,
double outcome_y, CogStance* stance,
int learn, double max_step, CogBeatResult* out) {
if (!region || !stance || !out) return -1;
memset(out, 0, sizeof *out);
if (stance->keystone) { learn = 0; out->wrote_keystone = 1; } /* §6: never write a keystone */
GeoGradient g;
if (engram_think(region, anchor, stance, &g) != 0) return -1; /* PREDICTION */
double p = g.magnitude;
double y = clampd(outcome_y, 0.0, 1.0);
double err = fabs(p - y);
out->correspondence = 1.0 - err;
out->error = err;
out->brier = (p - y) * (p - y);
if (learn) {
/* refine warp: gradient descent of (py)² wrt each axis_gain.
* p = 1/(1+D²); ∂p/∂gain_k = 2 p² proj_k² / (ext_k² gain_k³) (>=0)
* ∂(err²)/∂gain_k = 2 (py) ∂p/∂gain_k
* step = lr · ∂(err²)/∂gain_k, bounded to ±max_step (metastability). */
int dim = region->dim;
const float* x = anchor ? anchor : region->centroid;
float* r = malloc((size_t)dim * sizeof(float));
if (r) {
for (int i = 0; i < dim; i++) r[i] = (float)((double)x[i] - (double)region->centroid[i]);
double lr = 0.5;
double bound = (max_step > 0) ? max_step : 0.05; /* bounded update rate */
for (int k = 0; k < region->n_axes && k < stance->n_axes; k++) {
const float* ax = region->axes[k].axis; if (!ax) continue;
double proj = vdot(r, ax, dim);
double ext = region->axes[k].extent; if (ext < 1e-9) ext = 1e-9;
double gain = stance->axis_gain[k]; if (gain < 1e-6) gain = 1e-6;
double dp_dgain = 2.0 * p * p * (proj * proj) / (ext * ext * gain * gain * gain);
double dErr_dgain = 2.0 * (p - y) * dp_dgain;
double step = -lr * dErr_dgain;
step = clampd(step, -bound, bound);
stance->axis_gain[k] = clampd(gain + step, 0.1, 50.0);
}
free(r);
}
/* calibration */
stance->n_trials += 1;
stance->brier_sum += out->brier;
stance->last_error = err;
stance->ema_error = (stance->n_trials == 1) ? err : 0.9 * stance->ema_error + 0.1 * err;
double mean_brier = stance->brier_sum / (double)stance->n_trials;
stance->reliability = clampd(1.0 - sqrt(mean_brier), 0.0, 1.0);
}
out->reliability = stance->reliability;
engram_gradient_free(&g);
return 0;
}
+216
View File
@@ -0,0 +1,216 @@
/* engram_cognition.h — THE ONE OPERATION.
*
* The buildable form of the "cognition is one operation" theory (design doc
* engram/spec/cognitive-architecture.design.md; memory bdc8a488 / d582a766).
*
* Cognition is ONE operation — think — a directed traversal-READ of the geometry
* from an anchor, steered by a learned STANCE, whose output is a GRADIENT (a
* direction + spread over the geometry), never a point. The named faculties
* (reason / induce / abduce / analogy / relate / plan / ground) are human LABELS
* on regions of think's steering space: each faculty == { think + a named stance }.
* Collapse-to-a-point happens only at EXPRESSION (a separate faculty), never in think.
*
* NAMING (Will's directive): the surface verbs name the cognitive ACT being
* performed (think / reason / induce / ground / verify), not the internal function
* shape. The single frozen primitive underneath every faculty is engram_think,
* which composes over engram_reason_point_fit + the §5 geo-algebra. Those never
* learn. Only the STANCE learns.
*
* "Stance" is the theory's steering PRIOR, deliberately named distinctly: in this
* codebase the token "prior" already means previous-VERSION (supersession). A
* Stance is a learnable bias/disposition over the geometry — which axes matter,
* which way pays off, plus a calibrated track record — attached to a faculty-label
* and a region, and grounded-for-whom.
*
* PURE + (mostly) READ-ONLY, stdlib + libm only. think() and the warp are pure
* over their inputs. Persistence (Stance <-> StoreNode, grounded-by edges) is the
* only part that touches the store, and it is additive / supersede / tombstone —
* never mutate-in-place, never delete. It NEVER touches the live daemon: all
* offline against a scratch store, per the design's rails.
*/
#ifndef ENGRAM_COGNITION_H
#define ENGRAM_COGNITION_H
#include <stdint.h>
#include <stddef.h>
#include "engram_geometry.h"
#include "engram_reason.h"
#include "engram_store.h"
/* Max principal axes a stance warps (matches GeoParams.top_axes default budget). */
#define COG_MAX_AXES 32
/* ═══════════════════════════════════════════════════════════════════════════
* §1 GeoGradient — the OUTPUT of think. A direction + spread over the geometry,
* plus the calibrated confidence and the read it was computed against. NOT a point.
* A spiked gradient (spread→0) = "exact" (deduction); a spread gradient = "fuzzy"
* (prediction). The gradient is ALSO the next steering direction (closed-loop flow).
* ═══════════════════════════════════════════════════════════════════════════ */
typedef struct {
int dim;
float* direction; /* unit steering vector in the anchor's frame (owned) */
double spread; /* 0 = spiked/exact ... large = diffuse/fuzzy */
double confidence; /* calibrated, from the stance's track record (reliab.) */
double magnitude; /* THIS read's own membership/fit estimate ∈(0,1].
* The scalar an expression faculty would SAMPLE; kept
* on the gradient but never used AS a decision by think.*/
const char* anchor_id; /* borrowed: the vantage this was read from */
int n_support; /* neighborhood members that shaped the read */
const char* stance_id; /* borrowed: which stance steered this (provenance) */
} GeoGradient;
void engram_gradient_free(GeoGradient* g);
/* ═══════════════════════════════════════════════════════════════════════════
* §2 Stance — the learnable steering prior, as a first-class object. In memory
* here; persisted as a StoreNode (node_type "Stance") via cog_stance_*serialize.
*
* warp: axis_gain[] per-principal-axis multiplier on extents (which axes
* matter — gain>1 WIDENS an axis so it penalizes less);
* bias_dir[] a steering-direction seed in the region's frame;
* scalars faculty constants this stance overrides (ext_floor, etc).
* calibration: the track record — the ONLY thing the loop (§4) updates
* besides warp: n_trials, a Brier accumulator, reliability
* (→ GeoGradient.confidence), and an EMA error.
* keystone: if set, the correspondence-loop MUST NEVER write warp or
* calibration — read-mostly (self / values). §6 metastability.
* ═══════════════════════════════════════════════════════════════════════════ */
typedef struct {
char* id; /* stance node id (owned) */
char* faculty; /* the act this stance serves: "induce"|"relate"|... */
char* anchor_region; /* node/neighborhood id this stance is attached to */
char* for_whom; /* observer id — grounding is relational (NULL=global) */
int keystone; /* 1 = read-mostly, loop never writes it (§6) */
int dim; /* embedding dim of the region */
int n_axes; /* how many axis_gain entries are live (<= COG_MAX_AXES)*/
double axis_gain[COG_MAX_AXES]; /* per-axis extent multipliers (init 1.0) */
float* bias_dir; /* dim floats, steering seed (owned; NULL = none) */
double ext_floor; /* faculty scalar: the extent floor (init 1.0) */
double drop_frac; /* faculty scalar (causal): confound drop (init 0.5) */
double assoc_floor; /* faculty scalar (causal): assoc floor (init 0.2) */
/* calibration / track record */
int64_t n_trials;
double brier_sum; /* Σ (p y)² */
double reliability; /* calibrated ∈[0,1] → GeoGradient.confidence */
double ema_error; /* EMA of per-trial error */
double last_error;
} CogStance;
/* Initialize a neutral stance (all gains 1.0, default scalars, reliability 0.5).
* dim/n_axes taken from the region descriptor. faculty/id/for_whom are copied. */
int cog_stance_init(CogStance* s, const char* id, const char* faculty,
const char* anchor_region, const char* for_whom,
const GeoDescriptor* region);
void cog_stance_free(CogStance* s);
/* A stance set to today's hard-coded constants == behavioral parity with the
* pre-stance operators (axis_gain all 1.0, ext_floor default, drop_frac 0.5,
* assoc_floor 0.2). This is the FROZEN CONTROL used by the validation. */
void cog_stance_set_frozen_defaults(CogStance* s);
/* ── Serialization: Stance <-> StoreNode (compact line schema "STNC1", mirroring
* the reify "GEO1" precedent). Additive; the node's importance field caches the
* reliability readout. Round-trips exactly (reboot-prove). ──────────────────── */
char* cog_stance_to_metadata(const CogStance* s); /* owned string */
int cog_stance_to_node(const CogStance* s, StoreNode* out);/* fills a StoreNode */
int cog_stance_from_node(const StoreNode* n, CogStance* out);/* parse STNC1 */
#define COG_STANCE_NODE_TYPE "Stance"
#define COG_STANCE_META_MAGIC "STNC1"
/* ═══════════════════════════════════════════════════════════════════════════
* §1.2 think — the ONE operation. Frozen procedure over three steps:
* 1. re-origin on the anchor point (the vantage; the manifold is the read
* neighborhood, passed as `region`);
* 2. fit the anchor under the stance's WARP (engram_reason_point_fit with the
* axis extents multiplied by axis_gain and ext_floor substituted);
* 3. emit a GRADIENT: direction = the warped steepest-descent that reduces the
* fit distance (the "which way pays off" seed + bias_dir), spread from the
* fit distance, confidence from the stance's reliability, magnitude = the
* read's membership estimate. NO point-collapse — that is expression.
*
* `region` — the read neighborhood (built by vantage_read / geometry descriptor).
* `anchor` — the point to read FROM (dim floats). NULL = region centroid (self).
* `stance` — the steering prior. NULL = neutral (frozen defaults) => parity.
* Returns 0 and fills `out` (engram_gradient_free), <0 on error.
* ═══════════════════════════════════════════════════════════════════════════ */
int engram_think(const GeoDescriptor* region, const float* anchor,
const CogStance* stance, GeoGradient* out);
/* The warped fit itself (step 2), exposed for the loop + verifier reuse. Identical
* to engram_reason_point_fit when stance==NULL or all gains==1 && ext_floor default. */
int cog_warped_fit(const GeoDescriptor* region, const float* x,
const CogStance* stance, GeoFit* out);
/* EXPRESSION — the ONLY place a gradient collapses to a point. Samples the gradient
* off the anchor along its steering direction, scaled by (1 spread) so a spiked
* (confident) gradient lands a definite point and a diffuse one barely moves.
* This is deliberately a SEPARATE faculty from think (§1.2, M5). */
int engram_express(const GeoGradient* g, const float* anchor, float* out_point);
/* ═══════════════════════════════════════════════════════════════════════════
* §5 HOLD vs GROUND vs ASSERT. Holding is unconditional (the store gates nothing).
* Grounding is a RELATION — a "grounded-by" edge, probabilistic, grounded-for-whom.
* The honesty floor is checked only at ASSERTION.
* ═══════════════════════════════════════════════════════════════════════════ */
#define COG_GROUNDED_BY_RELATION "grounded-by"
#define COG_SALIENT_TO_RELATION "salient-to"
/* Write a grounded-by edge (additive). weight = grounding ∈(0,1] from the verifier;
* for_whom recorded in edge metadata (grounding is relational). Never a node flag. */
int cog_ground_edge(EngramPagedStore* s, const char* claim_id,
const char* evidence_id, double grounding, const char* for_whom);
/* Write/refresh a salient-to edge: salience is RELATIONAL (grounded-for-whom),
* carried on the edge to the observer — not baked into the node scalar (§2.1). */
int cog_salient_edge(EngramPagedStore* s, const char* node_id,
const char* observer_id, double salience);
/* The honesty floor — a QUERY at assertion time, NOT a schema constraint. Reads the
* claim's stored grounded-by edges (for the given observer) and returns:
* 1 = may assert (best grounding >= floor),
* 0 = REFUSE assertion (holds unconditionally; only asserting is gated),
* <0 = error. The content remains held either way. */
int cog_assert_gate(EngramPagedStore* s, const char* claim_id,
const char* for_whom, double floor);
/* ═══════════════════════════════════════════════════════════════════════════
* §4 THE REFLEXIVE CORRESPONDENCE-LOOP — the learning engine. think scores its
* OWN gradient against outcome, refines the stance on the error, and (optionally)
* writes the (gradient, outcome, error) back as self-describing geometry. This is
* the dormant verifier turned INWARD.
*
* grade(1) SELF-CONSISTENCY (no external world-labels): the outcome is what the
* geometry itself says — the membership determined by the region's SIGNAL subspace
* (the axes reality actually weights). The stance's cheap warped read is graded
* against that geometric truth; error refines the warp so the read corresponds.
* ═══════════════════════════════════════════════════════════════════════════ */
typedef struct {
double correspondence; /* ∈[0,1]: 1 |p y| for this trial */
double error; /* 1 correspondence */
double brier; /* running mean (p y)² across the stance's trials */
double reliability; /* the stance's current calibrated reliability */
int wrote_keystone; /* 1 iff a keystone update was BLOCKED (safety audit) */
} CogBeatResult;
/* One correspondence beat for ONE trial:
* think(region, anchor, stance) -> gradient (a PREDICTION, ungrounded)
* outcome y := grade-1 self-consistency target (in [0,1])
* error = |magnitude y|; refine stance.warp + calibration on the error
* (bounded step; NEVER writes a keystone stance)
* `learn`==0 grades WITHOUT updating (the frozen-control path). `max_step` bounds
* the per-beat warp change (metastability; §6). Returns 0 / <0. */
int engram_correspondence_beat(const GeoDescriptor* region, const float* anchor,
double outcome_y, CogStance* stance,
int learn, double max_step, CogBeatResult* out);
/* ═══════════════════════════════════════════════════════════════════════════
* §6 METASTABILITY. Keystones (self/values) are read-mostly: the loop reads but
* never writes them. Mark by stance flag or by a keystone-id set the loop consults.
* ═══════════════════════════════════════════════════════════════════════════ */
typedef struct { const char** ids; int n; } CogKeystoneSet;
int cog_is_keystone(const CogKeystoneSet* ks, const CogStance* s);
#endif /* ENGRAM_COGNITION_H */
File diff suppressed because it is too large Load Diff
+449
View File
@@ -0,0 +1,449 @@
/* engram_geometry.h — M9 FOUNDATION: the relational-neighborhood GEOMETRY
* DESCRIPTOR (design doc §3, §5; memory node e94371bd).
*
* Computes, for a relational neighborhood grown from a seed set, the compact
* (KB-not-MB) joint geometry Will specified: the SEMANTIC geometry (centroid,
* covariance / principal axes, radius) braided with the RELATIONAL geometry
* (k-core skeleton, hub->periphery centrality gradient), plus soft membership.
*
* Two coordinate systems, one shape "a constellation: bright prototype at the
* center, a cloud of members at varying distance, the strongest edges as a
* backbone, fading at the edges."
*
* Built ON the two standalone M-era modules only:
* - engram_vindex : semantic neighbors (the cloud) via ANN.
* - engram_store : node embeddings + hebb adjacency (the skeleton), read-only.
* It does NOT link or touch el_runtime.c, and it is a pure READ over the graph:
* it never modifies nodes, edges, activation, the index, or any retrieval path.
*
* Pure C11, stdlib + libm only. The descriptor is a foundation object; it is NOT
* wired into retrieval/priming yet (that is the next M9 step).
*/
#ifndef ENGRAM_GEOMETRY_H
#define ENGRAM_GEOMETRY_H
#include <stddef.h>
#include <stdint.h>
#include "engram_store.h"
#include "engram_vindex.h"
/* One member of the neighborhood + its place in the gradient. */
typedef struct {
char* id;
double membership; /* soft membership in [0,1] (semantic+relational blend) */
double centrality; /* skeleton weighted-degree — relational salience */
double salience; /* the node's own stored salience */
int core; /* k-core number (0 = fringe / not in any core) */
double dist_centroid; /* cosine distance of member emb to centroid (semantic)*/
int embedded; /* 1 if the member carried an emb vector */
} GeoMember;
/* One skeleton edge (indices into members[]). eff_weight = weight*(1+0.5*hebb),
* clamped to 1.0 the effective propagation strength eg_edge_eff_weight uses. */
typedef struct { uint32_t a, b; double eff_weight; double hebb; } GeoEdge;
/* A compact principal axis of the ellipsoid: unit direction in R^dim + extent
* (sqrt of the covariance eigenvalue = the ellipsoid's half-width along it). */
typedef struct { float* axis; double extent; } GeoAxis;
typedef struct {
int dim;
/* ── anchor ── */
char* hub_id; /* highest-centrality member: the relational hub */
float* centroid; /* v̄ ∈ R^dim: mean of the member embeddings in the
* frame the descriptor operated in. When centered
* (global_mean != NULL) this is the CENTERED
* centroid (mean of L2-normalized embs minus the
* global mean): the neighborhood's location in the
* isotropic/whitened frame. Add global_mean back to
* recover the raw prototype point. When uncentered
* it is the raw mean of L2-normalized member embs. */
float* global_mean; /* the centering offset actually applied (dim floats),
* or NULL if the descriptor ran in raw space. The §5
* operators (distance/overlap/Wasserstein) are only
* discriminative in the centered frame see notes. */
/* ── shape (compact covariance): top principal axes + extents ── */
int n_axes;
GeoAxis* axes; /* orientation + extents of the ellipsoid */
double total_variance; /* trace(Σ) = mean squared member dist to centroid*/
/* ── scale ── */
double radius; /* sqrt(total_variance) — the neighborhood breadth*/
/* ── members + gradient ── */
int n_members;
GeoMember* members; /* soft membership {id->weight} + centrality/salience */
/* ── skeleton ── */
int n_edges;
GeoEdge* edges; /* strong internal hebb edges = the backbone */
int k_core; /* the maximum core number present in the skeleton*/
/* ── diagnostics ── */
double co_registration;/* corr(hebb strength, semantic proximity) over */
/* internal edges: >0 = geometries agree (reify); */
/* <0 = disagree (surprising links / dream cands). */
int n_embedded; /* members that carried an emb vector */
} GeoDescriptor;
typedef struct {
int ann_k; /* semantic expansion: ANN neighbors per seed (0=off) */
int hop_relational; /* 1 = include seeds' hebb neighbors as members */
double edge_min_weight; /* skeleton: ignore internal edges below this eff wt */
int kcore_k; /* target k for the reported k-core (0 = auto/max) */
int top_axes; /* principal axes to retain (default 8) */
int max_members; /* cap neighborhood size (guards the eigensolve cost) */
} GeoParams;
/* Fill p with sane defaults: ann_k=24, hop_relational=1, edge_min_weight=0.05,
* kcore_k=0 (auto), top_axes=8, max_members=400. */
void engram_geo_default_params(GeoParams* p);
/* ── Global-mean cache (mean-centering / whitening the anisotropic emb space) ──
* The nomic-embed-text space over the engram corpus is strongly ANISOTROPIC:
* every embedding sits in a narrow cone (mean pairwise cosine ~0.55), which
* compresses cosine-based domain separation almost to nothing. Subtracting the
* GLOBAL MEAN of the (L2-normalized) embeddings recenters the cloud on the
* origin (mean pairwise cosine -> ~0), restoring isotropy so the §5 operators
* discriminate. The mean is a store-level derived quantity, like the ANN index:
* built once from the paged store, cached, and refreshed when the embedded set
* drifts. It lives here (not in the store) so this stays a contained, read-only
* addition; a runtime owns one GeoMeanCache per open store alongside its VIndex. */
typedef struct GeoMeanCache GeoMeanCache;
/* Scan every live node in `store` and compute the mean of the L2-normalized
* embeddings over the embed-eligible set (nodes carrying an emb vector; the
* unembedded telemetry/system nodes are skipped). Returns a malloc'd cache, or
* NULL on error / no embedded nodes. The offset vector is NOT renormalized it
* is a translation, applied by subtraction. */
GeoMeanCache* engram_geo_mean_build(EngramPagedStore* store);
/* The cached offset (dim floats) — pass to engram_geometry_descriptor as
* global_mean. Valid until the cache is freed/refreshed. */
const float* engram_geo_mean_vec(const GeoMeanCache* c);
int engram_geo_mean_dim(const GeoMeanCache* c);
uint64_t engram_geo_mean_count(const GeoMeanCache* c); /* #embedded nodes used */
/* Recompute the mean IN PLACE iff the embedded-node count has drifted by more
* than `frac` (e.g. 0.10 = 10%) since the cache was built "recompute on
* significant change". Returns 1 if it rebuilt, 0 if unchanged, <0 on error. */
int engram_geo_mean_maybe_refresh(GeoMeanCache* c, EngramPagedStore* store,
double frac);
void engram_geo_mean_free(GeoMeanCache* c);
/* Compute the geometry descriptor of the neighborhood grown from seed_ids.
* READ-ONLY over store + vindex.
* store an opened store (borrowed; not modified).
* vindex optional ANN index for semantic expansion; NULL disables it.
* vids the ordinal->store-id map returned by vindex_build_from_store
* (vids[node_id] == store id). Required iff vindex != NULL.
* n_vids length of vids.
* params NULL to use engram_geo_default_params.
* global_mean optional centering offset (dim floats, from engram_geo_mean_*).
* When non-NULL the SEMANTIC geometry is computed in mean-centered
* (isotropic) space: every normalized member emb has global_mean
* subtracted before the centroid / cosine-distance / co-registration
* math, so those operators discriminate. NULL = raw space (legacy).
* NOTE: the ANN neighbor query still runs in RAW unit-vector space
* centering is a rigid translation that ~preserves neighborhood
* MEMBERSHIP, so the index needs no rebuild; only the descriptor
* STATISTICS move to the centered frame (co-registration choice (b)).
* The eigen/covariance shape (axes, radius) is translation-invariant
* and therefore identical in either frame.
* Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error
* (no seeds resolvable, OOM). */
GeoDescriptor* engram_geometry_descriptor(
EngramPagedStore* store, VIndex* vindex,
char** vids, int n_vids,
const char* const* seed_ids, size_t n_seeds,
const GeoParams* params,
const float* global_mean);
void engram_geo_free(GeoDescriptor* g);
/* ── M-INTEROCEPTION P3: drift-sensor primitive (descriptor displacement) ────
* Read-only. GROWTH vs CORRUPTION split of how far B drifted from baseline A.
* See engram_geometry.c for the honesty note on the missing SelfAnchor. */
typedef struct {
double centroid_sep; /* L2 distance between centroids (same frame) */
double centroid_cos; /* 1 - cosine(centroidA, centroidB) */
double radius_delta; /* |radiusA - radiusB| — neighborhood scale change */
double core_disp; /* mean radial displacement of the invariant core */
double periph_disp; /* mean radial displacement of the periphery */
int core_matched; /* # core members matched by id across A,B */
int periph_matched; /* # periphery members matched by id across A,B */
} GeoDisplacement;
void engram_geo_displacement(const GeoDescriptor* a, const GeoDescriptor* b,
double core_frac, GeoDisplacement* out);
/* ═══════════════════════════════════════════════════════════════════════════
* §5 GEOMETRY OPERATORS a relational ALGEBRA over neighborhood descriptors.
* These are the reusable primitives Will specified: "primitives any CGI
* application should be able to use." READ-ONLY and PURE (stdlib + libm only) —
* they consume GeoDescriptor(s) and never touch the store, index, or activation.
*
* FRAME CONTRACT: both inputs MUST have been built in the SAME frame identical
* emb `dim` and identical `global_mean` (centered against the one true store-wide
* mean). The reify path builds every neighborhood that way, so descriptors are
* directly comparable. An operator returns <0 / NULL if the dims disagree.
*
* REPRESENTATION: the C descriptor lives in the FULL emb dim with a LOW-RANK
* covariance Σ = Σ_k extent_k² · a_k a_kᵀ over its retained principal axes
* (top_axes; the discarded tail variance is not modeled). Every operator mirrors
* the viz-proxy (engram-geometry-proxy.py §5) FORMULA exactly, but evaluates it on
* this representation so semantics match the proxy while absolute numbers differ
* (proxy works in a 24-dim global-PCA reduced dense frame; C in full-dim low-rank).
* The Wasserstein / combine eigen-work is done inside the small JOINT axis subspace
* (dimension nA+nB+1), which is EXACT for the low-rank covariances there.
* Each result struct is released by its engram_geo_*_free.
* */
/* overlap(A,B): shared-member set + Jaccard + centroid/scale proximity score. */
typedef struct {
char** shared_ids; /* ids present in BOTH neighborhoods (owned) */
int n_shared;
int n_union; /* |A B| by id */
double jaccard; /* |A∩B| / |AB| */
double centroid_distance; /* L2 between the (centered) centroids */
double overlap_score; /* jacc*0.5 + max(0,1d/(rA+rB))*0.5 (proxy form)*/
float* intersection_centroid; /* midpoint of the two centroids (dim, owned) */
int dim;
} GeoOverlap;
int engram_geo_overlap(const GeoDescriptor* a, const GeoDescriptor* b, GeoOverlap* out);
void engram_geo_overlap_free(GeoOverlap* o);
/* subtract(A,B) — ORTHOGONAL-COMPLEMENT residual: project A onto I V_B V_Bᵀ
* (V_B = B's top `b_dims` principal axes) "A with B's framing removed". Returns
* A's residual centroid + residual ellipsoid, the fraction of A's energy that lives
* inside B's subspace, and the centroid-difference vector. b_dims<=0 min(3,nB). */
typedef struct {
int dim;
float* residual_centroid; /* P⊥ c_A (owned) */
float* centroid_diff; /* c_A c_B (owned) */
double centroid_diff_mag;
double variance_explained_by_B; /* (‖Qc_A‖²+Tr(QΣ_A)) / (‖c_A‖²+Tr Σ_A) ∈[0,1]*/
int removed_dims; /* # of B axes used as V_B */
double residual_scale; /* sqrt(Tr(P⊥ Σ_A P⊥)) */
int n_axes; /* residual principal axes (owned) */
GeoAxis* axes;
} GeoResidual;
int engram_geo_subtract(const GeoDescriptor* a, const GeoDescriptor* b,
int b_dims, GeoResidual* out);
void engram_geo_residual_free(GeoResidual* r);
/* set-diff variant of subtract: members in A but not in B + the centroid arrow. */
typedef struct {
char** only_ids; /* member ids in A and not in B (owned) */
int n_only;
int removed; /* |A ∩ B| (dropped) */
float* centroid_diff; /* c_A c_B (dim, owned) */
double centroid_diff_mag;
int dim;
} GeoSetDiff;
int engram_geo_setdiff(const GeoDescriptor* a, const GeoDescriptor* b, GeoSetDiff* out);
void engram_geo_setdiff_free(GeoSetDiff* s);
/* combine(A,B): a merged descriptor — POOLED centroid + POOLED covariance
* (exact law-of-total-variance: the covariance you'd get by concatenating the two
* member clouds), re-eigendecomposed for its principal axes. Members = id-union
* (membership = max). top_axes<=0 8. Returns a malloc'd GeoDescriptor (free with
* engram_geo_free) in the same frame as A, or NULL on error. */
GeoDescriptor* engram_geo_combine(const GeoDescriptor* a, const GeoDescriptor* b,
int top_axes);
/* distance(A,B): centroid L2 + centroid cosine + closed-form Wasserstein-2
* (Bures metric) between the two Gaussians mirrors the proxy's _wasserstein2. */
typedef struct {
double centroid_distance;
double centroid_cosine;
double wasserstein2;
int dim;
} GeoDistance;
int engram_geo_distance(const GeoDescriptor* a, const GeoDescriptor* b, GeoDistance* out);
/* analogy(A,B): orthogonal PROCRUSTES transform min_R ‖A B R‖_F, RᵀR=I (SVD)
* aligning A's principal frame to B's (extent-scaled axes, paired by rank). R is
* returned COMPACTLY as an r×r rotation within the joint axis subspace `basis`
* (r vectors of dim floats); it acts as the identity on the orthogonal complement.
* Apply it to a vector with engram_geo_analogy_apply. */
typedef struct {
int dim;
int r; /* subspace rank; R is r×r */
float* basis; /* r×dim row-major orthonormal basis Q (owned) */
double* R; /* r×r rotation in Q-coords, row-major (owned) */
double residual; /* ‖A B R‖_F over the extent-scaled frames */
} GeoAnalogy;
int engram_geo_analogy(const GeoDescriptor* a, const GeoDescriptor* b, GeoAnalogy* out);
/* out_vec = R·v for v ∈ R^dim: v + Σ_i (R̂c c)_i q_i, c_i = q_i·v. dim floats. */
void engram_geo_analogy_apply(const GeoAnalogy* an, const float* v, float* out_vec);
void engram_geo_analogy_free(GeoAnalogy* an);
/* ═══════════════════════════════════════════════════════════════════════════
* M10 REIFICATION: densely co-wired relational neighborhoods crystallized into
* FIRST-CLASS, PERSISTED store records (design doc §2; memory 885f5945). This is
* NOT a cache it is durable structure. A reified neighborhood is a real store
* NODE (node_type "Neighborhood") that survives restart, is loaded on boot, and
* EVOLVES via supersede+provenance when the pattern shifts. The geometry-priming
* HOT PATH reads these persisted records (never computes geometry on the
* activation path). Ad-hoc/transient geometries still use the on-the-fly
* engram_geometry_descriptor above.
*
* Two record types, both ordinary TLV store nodes (no new on-disk format):
* - "GeoMeanFrame" : the store-wide centering mean, persisted ONCE (emb = mean
* vector, id ENGRAM_GEO_MEANFRAME_ID). Referenced by every
* neighborhood so priming centers against the SAME true mean.
* - "Neighborhood" : one reified neighborhood. emb = the RAW centroid (prototype
* point, so it stays centroid-ANN-able; centered_centroid =
* emb - meanframe). metadata = the compact "GEO1" schema:
* hub id, meanframe ref, scalar shape (radius, total_variance,
* k_core, co_registration, n_embedded), axis EXTENTS (ellipsoid
* half-widths), and the MEMBER list {id -> membership, centrality,
* core}. Member links are also persisted as edges relation="member".
*
* v1 honest simplifications (documented; extensible without migration): axis
* DIRECTION vectors are not persisted (extents capture the ellipsoid scale; the
* directions are recomputable via the on-the-fly descriptor for viz/operators);
* with hebb potentiation ~0 on today's store the "hebb-weighted" degree reduces to
* AUTHORED edge weight, so detected neighborhoods currently reflect authored edges
* the design is unchanged and self-correcting once hebb accrues.
* */
#define ENGRAM_GEO_NBHD_TYPE "Neighborhood"
#define ENGRAM_GEO_MEANFRAME_TYPE "GeoMeanFrame"
#define ENGRAM_GEO_MEANFRAME_ID "geo-meanframe" /* stable id of the singleton */
#define ENGRAM_GEO_NBHD_ID_PREFIX "nbhd-" /* id = nbhd-<hub>-<built_at> */
#define ENGRAM_GEO_MEMBER_RELATION "member"
/* ── One-level nesting (containment DAG). A "super" neighborhood is itself a
* Neighborhood node whose GEO1 metadata carries `level 1` + `c <child_id>` lines
* and which is joined to each child by a "contains" edge (childparent
* "nested-in"). Its id also begins with the "nbhd-" prefix, so the boot path
* routes it into the resident reify index and skips its edges from activation
* adjacency, exactly like a flat neighborhood. */
#define ENGRAM_GEO_SUPER_ID_PREFIX "nbhd-super-"
#define ENGRAM_GEO_SUPER_CONTENT "reified-super-neighborhood"
#define ENGRAM_GEO_CONTAINS_RELATION "contains"
#define ENGRAM_GEO_NESTED_RELATION "nested-in"
/* Per-run counters for the on-beat self-reification operation. All fields are
* out-params filled by engram_geo_reify_store when GeoReifyParams.stats != NULL.
* reified neighborhoods WRITTEN this run (new or materially changed hubs)
* skipped hubs whose signature was UNCHANGED vs their live neighborhood
* (the convergence signal: on a settled store this trends to the
* hub count and `reified` trends to 0 zero appends per beat)
* superseded prior neighborhood records tombstoned into the residue chain
* member_edges relation="member" edges written this run */
typedef struct {
int reified;
int skipped;
int superseded;
int member_edges;
} GeoReifyStats;
typedef struct {
int min_weighted_degree; /* hub qualifies iff strong-edge weighted degree >= this
* (0 = no floor: just rank + take top max_neighborhoods) */
int max_neighborhoods; /* homeostatic budget cap (default 128) */
double cover_membership; /* skip a hub already a member (w>=this) of an accepted
* neighborhood greedy non-redundant cover (default 0.5) */
int persist_member_edges; /* 1 = also write relation="member" edges (default 1) */
GeoParams descriptor; /* per-neighborhood params (top_axes may be 0 = skip eigensolve) */
/* ── SELF-REIFICATION extensions (default 0/NULL = legacy behavior) ──────────
* When these are off, engram_geo_reify_store is byte-for-byte its pre-2026-08-14
* behavior the ENGRAM_SELF_REIFY gate keeps the live binary inert until set. */
int incremental; /* 1 = CHANGE-DETECTION: skip a hub whose neighborhood
* signature (member-set + memberships + coarse geometry)
* is unchanged vs its current live record no re-append,
* no supersede. This is what makes on-beat reification
* idempotent/convergent under the write-barrier. */
int grounded_name; /* 1 = NAME the neighborhood from its most-central member
* labels (grounded, provenance-stamped) instead of the
* fixed content "reified-neighborhood". */
const char* cause; /* supersession CAUSE tag written into the residue chain
* ("autonomous-drift" on the beat, "explicit-override" /
* "rename" for the async manual override). NULL = "reify". */
GeoReifyStats* stats; /* nullable: per-run counters (see above). */
} GeoReifyParams;
/* Defaults: min_weighted_degree=0, max_neighborhoods=128, cover_membership=0.5,
* persist_member_edges=1, descriptor = engram_geo_default_params but top_axes=4,
* max_members=256 (reified neighborhoods stay compact). */
void engram_geo_reify_default_params(GeoReifyParams* p);
/* WRITE PATH (offline / consolidation — NEVER the activation hot path).
* Detect dense hub neighborhoods on the hebb-weighted graph, compute each one's
* CENTERED descriptor ONCE against the true store-wide mean, and PERSIST them as
* first-class records: the GeoMeanFrame (once) + one Neighborhood node per detected
* neighborhood (+ member edges), superseding any prior same-hub record with
* provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0.
* Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex,
char** vids, int n_vids,
const GeoReifyParams* params);
/* NESTING (one level). Reads the already-persisted flat Neighborhood records,
* agglomerates them by centroid cosine >= `min_cos` into groups, and persists one
* PARENT "super" Neighborhood node per group of >= 2 (geometry = mean of child
* centroids; `contains`/`nested-in` edges to children). Tombstones prior super
* records first (idempotent). Returns #parents persisted, or <0. Run AFTER
* engram_geo_reify_store. `min_cos` <= 0 uses the default (0.30). */
int engram_geo_reify_nest(EngramPagedStore* store, double min_cos);
/* ASYNC EXPLICIT OVERRIDE (degenerate manual case). Rename the live neighborhood
* `nbhd_id` to `new_name`: writes a fresh superseding Neighborhood record that
* carries the SAME geometry + members but the new name, tombstones the prior
* record, and PREPENDS a residue entry (cause="explicit-override", the prior
* name) so the maturation trail is preserved. Never blocks the autonomous beat;
* it simply supersedes whatever the beat last wrote. Returns the new record id
* (caller frees) or NULL on failure (id not a live neighborhood). */
char* engram_geo_neighborhood_rename(EngramPagedStore* store,
const char* nbhd_id, const char* new_name);
/* ── Resident loaded form of the persisted records (boot-time; READ-ONLY) ─────
* The durable Neighborhood/GeoMeanFrame records are the source of truth; this
* index is their LOADED form (like the resident node array is the loaded form of
* the node records, or adjacency the loaded form of edges). It never recomputes
* geometry it parses. Build it by feeding the runtime's boot node scan, or in
* one pass with engram_geo_reify_load. */
typedef struct GeoReifyIndex GeoReifyIndex;
GeoReifyIndex* engram_geo_reify_index_new(void);
/* Feed one store node; if it is a Neighborhood or GeoMeanFrame record it is parsed
* and absorbed (else ignored). The node is BORROWED (copied as needed). 0/<0. */
int engram_geo_reify_index_add(GeoReifyIndex* ix, const StoreNode* n);
/* Build the member->neighborhood hash after all adds. Call once. 0/<0. */
int engram_geo_reify_index_finalize(GeoReifyIndex* ix);
/* One-pass convenience: scan the store and build the finalized index. NULL if the
* store holds no reified records. */
GeoReifyIndex* engram_geo_reify_load(EngramPagedStore* store);
/* A borrowed view of one persisted neighborhood (owned by the index). */
typedef struct {
const char* id;
const char* hub_id;
int n_members;
char* const* member_ids; /* parallel arrays, length n_members */
const double* member_w; /* membership in [0,1] */
double radius;
double co_registration;
int k_core;
int n_embedded;
} GeoNeighborhood;
/* HOT-PATH LOOKUP (no geometry compute): resolve the seed set to the best
* persisted neighborhood the one with the greatest summed seed membership; on a
* miss (no seed is a member of any neighborhood) fall back to the centroid nearest
* the query embedding (centered by the loaded mean frame). q_emb may be NULL (then
* a miss returns NULL). Returns a BORROWED handle (do NOT free) or NULL. */
const GeoNeighborhood* engram_geo_reify_lookup(
const GeoReifyIndex* ix,
const char* const* seed_ids, size_t n_seeds,
const float* q_emb, int q_dim);
/* M10 read-only JSON serializers of the resident reify index (caller owns the
* returned malloc'd string; get_cstr returns NULL when id is not found). */
char* engram_geo_reify_list_cstr(const GeoReifyIndex* ix);
char* engram_geo_reify_get_cstr(const GeoReifyIndex* ix, const char* id);
int engram_geo_reify_count(const GeoReifyIndex* ix);
const float* engram_geo_reify_mean(const GeoReifyIndex* ix, int* dim); /* loaded true mean or NULL */
void engram_geo_reify_index_free(GeoReifyIndex* ix);
#endif /* ENGRAM_GEOMETRY_H */
+287
View File
@@ -0,0 +1,287 @@
/* engram_reason.c — the REASONING layer. Pure compositions over engram_geometry.h.
* stdlib + libm only; READ-ONLY over its descriptor inputs; touches no store/index. */
#include "engram_reason.h"
#include <stdlib.h>
#include <string.h>
#include <math.h>
/* ── small float-vector helpers ─────────────────────────────────────────────── */
static double vdot(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) s += (double)a[i] * (double)b[i]; return s;
}
static double vnorm(const float* a, int dim) { return sqrt(vdot(a, a, dim)); }
static double vcos(const float* a, const float* b, int dim) {
double na = vnorm(a, dim), nb = vnorm(b, dim);
if (na < 1e-12 || nb < 1e-12) return 0.0; /* a null vector ⇒ no direction */
double c = vdot(a, b, dim) / (na * nb);
if (c > 1.0) c = 1.0; if (c < -1.0) c = -1.0;
return c;
}
static double l2(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) { double d = (double)a[i] - (double)b[i]; s += d * d; }
return sqrt(s);
}
/* ═══════════════════════════════════════════ SHARED — point-to-manifold FIT ══ */
int engram_reason_point_fit(const GeoDescriptor* g, const float* x,
double ext_floor, GeoFit* out) {
if (!g || !x || !out || g->dim <= 0 || !g->centroid) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
int dim = g->dim;
/* residual r = x centroid */
double rr = 0; /* ‖r‖² */
float* r = malloc((size_t)dim * sizeof(float));
if (!r) return -1;
for (int i = 0; i < dim; i++) { double d = (double)x[i] - (double)g->centroid[i]; r[i] = (float)d; rr += d * d; }
double maha2 = 0, ss_in = 0; /* Mahalanobis² and in-subspace energy */
for (int k = 0; k < g->n_axes; k++) {
const float* ax = g->axes[k].axis; if (!ax) continue;
double proj = vdot(r, ax, dim); /* axes are orthonormal directions */
double den = g->axes[k].extent; if (den < ext_floor) den = ext_floor;
maha2 += (proj / den) * (proj / den);
ss_in += proj * proj;
}
double ortho2 = rr - ss_in; if (ortho2 < 0) ortho2 = 0; /* off-subspace energy */
double dist2 = maha2 + ortho2 / (ext_floor * ext_floor);
out->mahalanobis = sqrt(maha2);
out->ortho_residual = sqrt(ortho2);
out->distance = sqrt(dist2);
out->score = 1.0 / (1.0 + dist2);
free(r);
return 0;
}
/* ═══════════════════════════════════════════════════════════════ ANALOGY ════ */
int engram_reason_analogy(const GeoDescriptor* A, const GeoDescriptor* B,
const GeoDescriptor* C,
const GeoDescriptor* const* candidates, int n_candidates,
GeoAnalogyResult* out) {
if (!A || !B || !C || !out) return -1;
if (!A->centroid || !B->centroid || !C->centroid) return -1;
int dim = A->dim;
if (B->dim != dim || C->dim != dim) return -1;
memset(out, 0, sizeof *out);
out->dim = dim; out->best = -1;
/* Learn R_{A→B}. engram_geo_analogy(X,Y) yields R with apply(R, Y-axis) ≈ X-axis
* (R maps Y's frame X's frame); so R that maps AB is engram_geo_analogy(B,A). */
GeoAnalogy an;
if (engram_geo_analogy(B, A, &an) != 0) return -1;
out->analogy_residual = an.residual;
/* mapped = R·c_C + (c_B R·c_A) : the A→B affine (rotation + residual shift). */
float* RcA = malloc((size_t)dim * sizeof(float));
float* RcC = malloc((size_t)dim * sizeof(float));
out->mapped_point = malloc((size_t)dim * sizeof(float));
if (!RcA || !RcC || !out->mapped_point) { free(RcA); free(RcC); free(out->mapped_point); out->mapped_point = NULL; engram_geo_analogy_free(&an); return -1; }
engram_geo_analogy_apply(&an, A->centroid, RcA);
engram_geo_analogy_apply(&an, C->centroid, RcC);
for (int i = 0; i < dim; i++)
out->mapped_point[i] = (float)((double)RcC[i] + ((double)B->centroid[i] - (double)RcA[i]));
free(RcA); free(RcC);
engram_geo_analogy_free(&an);
/* nearest candidate to the mapped point (centroid L2). */
if (candidates && n_candidates > 0) {
out->n_candidates = n_candidates;
out->distances = malloc((size_t)n_candidates * sizeof(double));
if (!out->distances) return -1;
double best = -1; int bi = -1;
for (int i = 0; i < n_candidates; i++) {
const GeoDescriptor* cd = candidates[i];
double d = (cd && cd->centroid && cd->dim == dim) ? l2(out->mapped_point, cd->centroid, dim) : INFINITY;
out->distances[i] = d;
if (bi < 0 || d < best) { best = d; bi = i; }
}
out->best = bi; out->best_distance = best;
}
return 0;
}
void engram_reason_analogy_free(GeoAnalogyResult* r) {
if (!r) return;
free(r->mapped_point); free(r->distances);
r->mapped_point = NULL; r->distances = NULL;
}
/* ═══════════════════════════════════════════════════════════════ INDUCTION ══ */
int engram_reason_induce(const GeoDescriptor* const* examples, int n_examples,
int top_axes, double ext_floor, GeoInduction* out) {
if (!examples || n_examples < 1 || !out) return -1;
if (top_axes <= 0) top_axes = 8;
memset(out, 0, sizeof *out);
/* fold the examples left→right through the pooled-Gaussian combine. n==1 pools
* the single example with itself (identical cov same shape, id-union = itself). */
GeoDescriptor* acc = engram_geo_combine(examples[0],
examples[n_examples > 1 ? 1 : 0], top_axes);
if (!acc) return -1;
for (int i = 2; i < n_examples; i++) {
GeoDescriptor* nxt = engram_geo_combine(acc, examples[i], top_axes);
engram_geo_free(acc);
if (!nxt) return -1;
acc = nxt;
}
out->rule = acc;
out->n_examples = n_examples;
out->ext_floor = (ext_floor > 0) ? ext_floor
: (acc->radius > 0 ? acc->radius * 0.25 : 1.0);
return 0;
}
double engram_reason_membership(const GeoInduction* ind, const float* x) {
if (!ind || !ind->rule || !x) return -1;
GeoFit f;
if (engram_reason_point_fit(ind->rule, x, ind->ext_floor, &f) != 0) return -1;
return f.score;
}
void engram_reason_induction_free(GeoInduction* out) {
if (!out) return;
if (out->rule) engram_geo_free(out->rule);
out->rule = NULL;
}
/* ═══════════════════════════════════════════════════════════════ ABDUCTION ══ */
int engram_reason_abduce(const float* obs, int dim,
const GeoDescriptor* const* hypotheses, int n,
double ext_floor, GeoAbduction* out) {
if (!obs || !hypotheses || n < 1 || dim <= 0 || !out) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
memset(out, 0, sizeof *out);
out->n = n; out->best = -1;
out->scores = malloc((size_t)n * sizeof(double));
out->distances = malloc((size_t)n * sizeof(double));
out->rank = malloc((size_t)n * sizeof(int));
if (!out->scores || !out->distances || !out->rank) { engram_reason_abduction_free(out); return -1; }
double best = -1; int bi = -1;
for (int i = 0; i < n; i++) {
out->rank[i] = i;
const GeoDescriptor* h = hypotheses[i];
GeoFit f;
if (!h || h->dim != dim || engram_reason_point_fit(h, obs, ext_floor, &f) != 0) {
out->scores[i] = 0.0; out->distances[i] = INFINITY;
} else {
out->scores[i] = f.score; out->distances[i] = f.distance;
}
if (bi < 0 || out->scores[i] > best) { best = out->scores[i]; bi = i; }
}
out->best = bi; out->best_score = (bi >= 0) ? out->scores[bi] : 0.0;
/* rank indices best→worst by score (insertion sort — n is small). */
for (int i = 1; i < n; i++) {
int key = out->rank[i]; int j = i - 1;
while (j >= 0 && out->scores[out->rank[j]] < out->scores[key]) { out->rank[j + 1] = out->rank[j]; j--; }
out->rank[j + 1] = key;
}
return 0;
}
void engram_reason_abduction_free(GeoAbduction* out) {
if (!out) return;
free(out->scores); free(out->distances); free(out->rank);
out->scores = NULL; out->distances = NULL; out->rank = NULL;
}
/* ═══════════════════════════════════════════════════════════════════ CAUSAL ══ */
/* |cos| of two descriptors' centroids after removing confounder Z's subspace. */
static double controlled_assoc(const GeoDescriptor* x, const GeoDescriptor* y,
const GeoDescriptor* z) {
GeoResidual rx, ry; double c = 0;
int ox = engram_geo_subtract(x, z, 0, &rx);
int oy = engram_geo_subtract(y, z, 0, &ry);
if (ox == 0 && oy == 0 && rx.residual_centroid && ry.residual_centroid)
c = fabs(vcos(rx.residual_centroid, ry.residual_centroid, x->dim));
if (ox == 0) engram_geo_residual_free(&rx);
if (oy == 0) engram_geo_residual_free(&ry);
return c;
}
int engram_reason_causal(const GeoDescriptor* x, const GeoDescriptor* y,
const GeoDescriptor* const* confounders, int n_conf,
int64_t t_x, int64_t t_y,
double drop_frac, GeoCausal* out) {
if (!x || !y || !out || !x->centroid || !y->centroid || x->dim != y->dim) return -1;
if (!(drop_frac > 0 && drop_frac < 1)) drop_frac = 0.5;
memset(out, 0, sizeof *out);
const double assoc_floor = 0.2; /* below this = no meaningful association */
out->assoc_raw = fabs(vcos(x->centroid, y->centroid, x->dim));
/* control for each confounder; the strongest single explainer wins (min assoc). */
double ctrl = out->assoc_raw;
for (int i = 0; i < n_conf; i++) {
if (!confounders[i]) continue;
double c = controlled_assoc(x, y, confounders[i]);
if (c < ctrl) ctrl = c;
}
out->assoc_controlled = ctrl;
out->temporal_dir = (t_x < t_y) ? 1 : (t_x > t_y) ? -1 : 0;
if (out->assoc_raw < assoc_floor) {
out->verdict = GEO_CAUSAL_NONE;
} else if (ctrl < (1.0 - drop_frac) * out->assoc_raw && ctrl < assoc_floor) {
out->verdict = GEO_CAUSAL_CONFOUNDED; out->confounded = 1;
} else if (out->temporal_dir != 0) {
out->verdict = GEO_CAUSAL_DIRECTED; out->strength = ctrl;
} else {
out->verdict = GEO_CAUSAL_NONE; /* associated + robust but unorientable */
}
return 0;
}
/* ═══════════════════════════════════════════════════════════════════ PLANNING ══ */
int engram_reason_plan(const GeoDescriptor* const* nodes, int n,
int start, int goal, double neighbor_radius,
int use_wasserstein, GeoPlan* out) {
if (!nodes || n < 1 || !out) return -1;
if (start < 0 || start >= n || goal < 0 || goal >= n) return -1;
if (!(neighbor_radius > 0)) return -1;
memset(out, 0, sizeof *out);
/* dense edge weights (i<j symmetric); INFINITY = not adjacent. */
double* W = malloc((size_t)n * (size_t)n * sizeof(double));
if (!W) return -1;
for (int i = 0; i < n; i++) for (int j = 0; j < n; j++) W[(size_t)i * n + j] = (i == j) ? 0.0 : INFINITY;
for (int i = 0; i < n; i++) {
for (int j = i + 1; j < n; j++) {
GeoDistance d;
if (nodes[i] && nodes[j] && engram_geo_distance(nodes[i], nodes[j], &d) == 0) {
double w = use_wasserstein ? d.wasserstein2 : d.centroid_distance;
if (w <= neighbor_radius) { W[(size_t)i * n + j] = w; W[(size_t)j * n + i] = w; }
}
}
}
/* O(n²) Dijkstra. */
double* dist = malloc((size_t)n * sizeof(double));
int* prev = malloc((size_t)n * sizeof(int));
char* done = calloc((size_t)n, 1);
if (!dist || !prev || !done) { free(W); free(dist); free(prev); free(done); return -1; }
for (int i = 0; i < n; i++) { dist[i] = INFINITY; prev[i] = -1; }
dist[start] = 0;
for (int it = 0; it < n; it++) {
int u = -1; double bd = INFINITY;
for (int i = 0; i < n; i++) if (!done[i] && dist[i] < bd) { bd = dist[i]; u = i; }
if (u < 0) break;
done[u] = 1;
if (u == goal) break;
for (int v = 0; v < n; v++) {
double w = W[(size_t)u * n + v];
if (w < INFINITY && !done[v] && dist[u] + w < dist[v]) { dist[v] = dist[u] + w; prev[v] = u; }
}
}
if (dist[goal] < INFINITY) {
int len = 0; for (int v = goal; v != -1; v = prev[v]) len++;
out->path = malloc((size_t)len * sizeof(int));
if (out->path) {
out->path_len = len;
int idx = len - 1;
for (int v = goal; v != -1; v = prev[v]) out->path[idx--] = v;
out->total_cost = dist[goal];
out->reached = 1;
}
}
free(W); free(dist); free(prev); free(done);
return 0;
}
void engram_reason_plan_free(GeoPlan* out) {
if (!out) return;
free(out->path); out->path = NULL;
}
+161
View File
@@ -0,0 +1,161 @@
/* engram_reason.h — the REASONING layer: compositions over the §5 geometry
* OPERATORS (engram_geometry.h). Where the operators are a relational ALGEBRA over
* neighborhood descriptors, these are reasoning MODES built by CHAINING that algebra:
*
* ANALOGY A:B :: C:? learn the AB transform (Procrustes), apply to C.
* INDUCTION {E_i} rule pool example geometries; a generalizing structure
* + a membership test.
* ABDUCTION x best H the structure whose geometry best PLACES an
* observation in-distribution (inverse of prediction).
* CAUSAL x ? y | Z, t separate mere overlap (correlation) from directed
* influence (temporal precedence + association that
* SURVIVES controlling for confounders via subtract).
* PLANNING start goal a trajectory (sequence of neighborhoods) through the
* manifold: shortest path over geo-distance edges.
*
* PURE + READ-ONLY (stdlib + libm only): every function consumes GeoDescriptor(s)
* (+ a few scalars / timestamps) and NEVER touches the store, index, or activation.
* All geometry is delegated to the engram_geo_* primitives; this file only composes.
*
* FRAME CONTRACT (inherited): descriptors passed together MUST share emb `dim` and
* `global_mean` frame exactly the §5 operator contract. A function returns <0 on
* a dim/frame mismatch or bad argument.
*/
#ifndef ENGRAM_REASON_H
#define ENGRAM_REASON_H
#include <stdint.h>
#include "engram_geometry.h"
/* ═══════════════════════════════════════════════════════════════════════════
* SHARED PRIMITIVE point-to-manifold FIT. How well does a single point x sit
* inside a neighborhood's ellipsoid? Splits the residual (x centroid) into:
* - the IN-SUBSPACE part, scaled by each axis extent a Mahalanobis distance
* (how many "radii" out along the modeled directions), and
* - the ORTHOGONAL part outside the retained axes energy the model does not
* explain at all (charged at the extent floor).
* This is the common engine under INDUCTION's membership test and ABDUCTION's
* explanation ranking. ext_floor (>0) guards zero-extent axes / the null model.
* */
typedef struct {
double mahalanobis; /* sqrt( Σ_k ((a_k·(xc)) / max(ext_k,floor))² ) */
double ortho_residual; /* ‖(xc) projected off the retained axes‖ (raw L2) */
double distance; /* sqrt( maha² + (ortho_residual/floor)² ) — full fit */
double score; /* 1 / (1 + distance²) ∈ (0,1] (1 = dead-center) */
} GeoFit;
int engram_reason_point_fit(const GeoDescriptor* g, const float* x,
double ext_floor, GeoFit* out);
/* ═══════════════════════════════════════════════════════════════════════════
* ANALOGY "A:B :: C:?". Learn the transform that carries A to B (orthogonal
* Procrustes rotation R between their principal frames + the residual translation),
* apply it to C, and return the mapped point + the nearest candidate neighborhood.
* Composes: engram_geo_analogy (R) + engram_geo_analogy_apply + engram_geo_distance.
* */
typedef struct {
int dim;
float* mapped_point; /* predicted D location = R·c_C + (c_B R·c_A) (owned)*/
double analogy_residual;/* Procrustes ‖AB R‖_F — frame-alignment quality */
int best; /* index of nearest candidate to mapped_point, or 1 */
double best_distance; /* centroid L2 from mapped_point to the winner */
int n_candidates;
double* distances; /* centroid L2 mapped_point→candidate[i] (owned)*/
} GeoAnalogyResult;
/* candidates may be NULL/0 (then best=1 and only mapped_point is filled). */
int engram_reason_analogy(const GeoDescriptor* A, const GeoDescriptor* B,
const GeoDescriptor* C,
const GeoDescriptor* const* candidates, int n_candidates,
GeoAnalogyResult* out);
void engram_reason_analogy_free(GeoAnalogyResult* r);
/* ═══════════════════════════════════════════════════════════════════════════
* INDUCTION from a SET of example neighborhoods to the generalizing structure.
* Pools the examples (law-of-total-variance via engram_geo_combine, folded left to
* right) into a single "rule" descriptor whose top principal axes are the directions
* CONSISTENTLY present across the examples (the shared subspace surfaces as the
* dominant pooled axes; idiosyncratic per-example directions fall to the tail).
* The rule carries a membership test (point-to-manifold fit against the pool).
* */
typedef struct {
GeoDescriptor* rule; /* induced generalizing geometry (owned; geo_free) */
double ext_floor; /* extent floor used by the membership test */
int n_examples;/* how many examples were pooled */
} GeoInduction;
/* top_axes<=0 → 8. ext_floor<=0 → derived from the pooled radius. */
int engram_reason_induce(const GeoDescriptor* const* examples, int n_examples,
int top_axes, double ext_floor, GeoInduction* out);
/* Membership of a point in the induced rule ∈ (0,1] (the fit score). <0 on error. */
double engram_reason_membership(const GeoInduction* ind, const float* x);
void engram_reason_induction_free(GeoInduction* out);
/* ═══════════════════════════════════════════════════════════════════════════
* ABDUCTION inference to the best explanation. Given an observation POINT, rank a
* set of candidate structures by how well each PLACES the observation in-distribution
* (min point-to-manifold distance = the structure that, if assumed, best accounts for
* the observation). The inverse of prediction.
* */
typedef struct {
int best; /* index of best-explaining hypothesis, or 1 */
double best_score;
int n;
double* scores; /* fit score per hypothesis (higher = better) (owned)*/
double* distances; /* explanation distance per hypothesis (owned)*/
int* rank; /* hypothesis indices sorted best→worst (owned)*/
} GeoAbduction;
int engram_reason_abduce(const float* obs, int dim,
const GeoDescriptor* const* hypotheses, int n,
double ext_floor, GeoAbduction* out);
void engram_reason_abduction_free(GeoAbduction* out);
/* ═══════════════════════════════════════════════════════════════════════════
* CAUSAL correlation vs causation. Over two variables' geometries (+ candidate
* confounders + temporal order), distinguish:
* - mere co-occurrence / overlap (correlation), from
* - directed influence: association that (a) SURVIVES controlling for confounders
* (subtract each Z's subspace from both centroids, re-measure) and (b) is oriented
* by temporal PRECEDENCE.
* Composes: centroid cosine (correlation) + engram_geo_subtract (control) + timestamps.
* */
typedef enum {
GEO_CAUSAL_NONE = 0, /* no meaningful association */
GEO_CAUSAL_DIRECTED = 1, /* survives control + temporally ordered → cause→eff */
GEO_CAUSAL_CONFOUNDED = 2 /* correlated but association dies under control */
} GeoCausalVerdict;
typedef struct {
double assoc_raw; /* |cos(c_x,c_y)| — the raw correlation */
double assoc_controlled; /* |cos| of residual centroids after control */
int temporal_dir; /* +1 x→y, 1 y→x, 0 tie/unknown */
GeoCausalVerdict verdict;
int confounded; /* 1 iff verdict==CONFOUNDED (the flag) */
double strength; /* directed influence estimate ∈[0,1] (0 else)*/
} GeoCausal;
/* confounders may be NULL/0. t_x,t_y are comparable timestamps (any monotone unit);
* pass equal values for "unknown order". drop_frac(0,1): a controlled association
* below (1drop_frac)·assoc_raw AND below an absolute floor CONFOUNDED. */
int engram_reason_causal(const GeoDescriptor* x, const GeoDescriptor* y,
const GeoDescriptor* const* confounders, int n_conf,
int64_t t_x, int64_t t_y,
double drop_frac, GeoCausal* out);
/* ═══════════════════════════════════════════════════════════════════════════
* PLANNING trajectory construction. Given a set of neighborhoods (manifold nodes),
* a start and a goal, build a PATH (sequence of intermediate neighborhoods) by
* shortest path over the graph whose edges connect neighborhoods within
* neighbor_radius, weighted by geo-distance. Long straight jumps are not edges, so
* the path follows the manifold's curvature through intermediates (a discrete geodesic).
* Composes: engram_geo_distance (edge weights) + Dijkstra.
* */
typedef struct {
int* path; /* node indices start..goal (owned) */
int path_len;
double total_cost; /* summed centroid-distance edge weights along path */
int reached; /* 1 if goal reachable within neighbor_radius graph */
} GeoPlan;
/* neighbor_radius>0: max centroid distance for two neighborhoods to be adjacent.
* Use "wasserstein"!=0 to weight edges by Wasserstein-2 instead of centroid L2. */
int engram_reason_plan(const GeoDescriptor* const* nodes, int n,
int start, int goal, double neighbor_radius,
int use_wasserstein, GeoPlan* out);
void engram_reason_plan_free(GeoPlan* out);
#endif /* ENGRAM_REASON_H */
File diff suppressed because it is too large Load Diff
+109
View File
@@ -209,6 +209,41 @@ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx);
uint64_t engram_wal_next_lsn(const EngramPagedStore* s);
uint64_t engram_last_checkpoint_lsn(const EngramPagedStore* s);
/* ── M4: demand-paging buffer pool (additive residency; on-disk format UNCHANGED) ──
*
* The write-back, no-steal cache of M2 becomes a bounded, demand-paged buffer
* pool. A fixed frame budget (env ENGRAM_POOL_FRAMES; 0 = unlimited; default
* large whole store resident identical to Phase 1) keeps only hot pages in
* RAM; a page access that is not resident faults in from neuron.egm, and under
* pressure a CLEAN, unpinned frame is evicted (LRU). Dirty frames are never
* stolen (M2 no-steal / WAL durability), and superblocks + index root/interior
* pages are auto-pinned. Prefetch (env ENGRAM_PREFETCH) reads ahead on scans. */
/* Pin / unpin an individual page (faults it in and keeps it resident until
* unpinned). Pin a hot layer's pages (WM/core) as a set. Idempotent counts. */
int store_pin_page(EngramPagedStore* s, uint64_t page_id);
int store_unpin_page(EngramPagedStore* s, uint64_t page_id);
int store_pin_layer(EngramPagedStore* s, uint32_t layer); /* returns #pages pinned */
int store_unpin_layer(EngramPagedStore* s, uint32_t layer);
/* Buffer-pool introspection. */
typedef struct StorePoolStats {
size_t cap; /* frame budget (0 = unlimited) */
size_t resident; /* frames currently resident */
size_t pinned; /* frames that cannot be evicted (dirty/pinned/structural) */
size_t dirty; /* dirty (un-checkpointed) frames */
unsigned prefetch; /* read-ahead window */
uint64_t hits, misses; /* page_read cache hits / demand faults */
uint64_t evictions; /* clean frames reclaimed */
uint64_t prefetch_reads; /* pages brought in by read-ahead */
} StorePoolStats;
void store_pool_stats(const EngramPagedStore* s, StorePoolStats* out);
int store_pool_resident(const EngramPagedStore* s, uint64_t page_id);
/* Test hooks: set the frame budget / prefetch window at runtime (NOT format). */
void store__set_pool_frames(EngramPagedStore* s, size_t frames);
void store__set_prefetch(EngramPagedStore* s, unsigned window);
/* Crash-test hooks (writes only under a throwaway dir).
* store__crash abandon all RAM state without flush/fsync (power loss).
* store__flush_pages pwrite dirty pages to disk WITHOUT a checkpoint (steal).
@@ -218,4 +253,78 @@ void store__crash(EngramPagedStore* s);
int store__flush_pages(EngramPagedStore* s);
int store__checkpoint_crashat(EngramPagedStore* s, int phase);
/* ── M5: online compaction + background checkpointer (additive; format UNCHANGED) ──
*
* COMPACTION reclaims the space held by DEAD records tombstoned nodes/edges
* (telemetry prune, forget), superseded ids, and the stale prior versions a
* re-put/hebb-batch leaves behind plus the overflow pages they orphaned. It
* rewrites only the LIVE records (bit-exact) into a fresh, densely packed image
* with fresh id + adjacency indexes, then commits the swap atomically, so the
* .egm file physically SHRINKS and the freed pages are reclaimed. Crash-safe:
* a crash at any instant recovers to either the pre- or the post-compaction
* store, never a corrupt mix (atomic rename is the commit point). It cooperates
* with the M4 pool (no-steal, pins) by building into a separate store whose own
* pool honours ENGRAM_POOL_FRAMES, then INVALIDATING every frame of the live
* pool so no stale frame survives for a relocated page.
*
* Requires a quiesce point: store_compact performs a checkpoint (or sync) at
* entry, so it is called between mutations, not concurrently with one. */
int store_compact(EngramPagedStore* s);
/* Test hook: run compaction but stop (then power-loss) after `phase`:
* 0 = after the entry checkpoint, before building ( recovers pre-compaction)
* 1 = after building+fsync the new image, before rename ( pre-compaction)
* 2 = after the atomic rename, before reopening RAM state ( post-compaction)
* phase<0 = full compaction. Frees `s` on a crash phase (like the checkpoint hook). */
int store__compact_crashat(EngramPagedStore* s, int phase);
/* BACKGROUND CHECKPOINTER policy. A checkpoint fires automatically on the write
* path when ANY armed trigger trips, reclaiming the WAL prefix without an explicit
* engram_checkpoint. 0 disables that trigger. Same checkpoint semantics as M2.
* ops mutations since last checkpoint (default 100000)
* dirty_pages dirty (un-checkpointed) pool frames
* wal_bytes bytes appended to the WAL since it was last reclaimed
* interval_ms wall-clock ms since the last checkpoint (checked on writes) */
void store_set_checkpoint_policy(EngramPagedStore* s, uint64_t ops,
size_t dirty_pages, uint64_t wal_bytes,
long long interval_ms);
/* Introspection: number of pages currently on the free-list. */
uint64_t store_free_page_count(const EngramPagedStore* s);
/* ── CCR §4 managed-memory layer (write-barrier + minor GC + observability) ─────
*
* All flag-gated at store open (default OFF byte-for-byte legacy behaviour):
* ENGRAM_WRITE_BARRIER=1 arm the durable-content write-barrier: a store_put_node
* whose DURABLE fields (content/type/label/tier/tags/metadata/importance/
* confidence/decay/layer/emb) are byte-identical to the last persisted copy
* is SKIPPED entirely no LSN, no WAL, no record, no tombstone. This kills
* the ~99.78% checkpoint full-walk garbage at the source (ephemeral
* activation/WM state is intentionally not re-persisted on think-only cycles).
* ENGRAM_GC=1 (a) node re-puts supersede prior copies (mark DEAD, as edges do)
* so stale versions become reclaimable, and (b) a MINOR GC runs at the head
* of every checkpoint, returning whole dead NODE/EDGE pages to the free list.
* The MAJOR GC is the existing merge-safe store_compact (schedule on a dead-ratio
* threshold from the soul). */
/* Run one minor-GC sweep now: reclaim whole dead NODE/EDGE pages to the free list.
* Returns the number of pages reclaimed (>=0). Safe to call between mutations;
* automatically invoked at each checkpoint when ENGRAM_GC is armed. */
int store_minor_gc(EngramPagedStore* s);
/* GC / cache observability census (CCR §4.4). Page tallies are point-in-time;
* the *_writes / *_skips / *_runs / *_reclaimed counters are cumulative since open. */
typedef struct StoreGcStats {
uint64_t node_pages, edge_pages, index_pages, overflow_pages, free_pages;
uint64_t live_nodes, live_edges; /* live slots on NODE / EDGE pages */
uint64_t dead_slots; /* superseded/tombstoned slots awaiting reclaim */
uint64_t live_bytes, dead_bytes; /* on-page record bytes, live vs dead */
uint64_t durable_writes; /* node puts that actually appended a record */
uint64_t barrier_skips; /* node puts skipped by the write-barrier */
uint64_t minor_gc_runs; /* minor-GC invocations */
uint64_t pages_reclaimed; /* whole pages returned to the free list by minor GC */
int barrier_on, gc_on; /* which gates are armed */
} StoreGcStats;
void store_gc_stats(EngramPagedStore* s, StoreGcStats* out);
#endif /* ENGRAM_STORE_H */
+157
View File
@@ -0,0 +1,157 @@
/* engram_verify.c — the VERIFIER layer. Pure compositions over engram_reason.h +
* engram_geometry.h. stdlib + libm only; READ-ONLY over its inputs; touches no
* store/index/activation. See engram_verify.h for the design and the frame contract. */
#include "engram_verify.h"
#include <stdlib.h>
#include <string.h>
#include <math.h>
/* ── small float-vector helpers (mirror engram_reason.c) ────────────────────── */
static double vdot(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) s += (double)a[i] * (double)b[i]; return s;
}
static double l2(const float* a, const float* b, int dim) {
double s = 0; for (int i = 0; i < dim; i++) { double d = (double)a[i] - (double)b[i]; s += d * d; }
return sqrt(s);
}
/* ═══════════════════════════════════════════════════════ GROUNDING ══════════ */
int engram_verify_grounding(const float* claim, int dim,
const GeoDescriptor* const* evidence, int n_evidence,
double ext_floor, double ground_threshold,
GeoGrounding* out) {
if (!claim || dim <= 0 || !evidence || n_evidence < 1 || !out) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
if (!(ground_threshold > 0 && ground_threshold < 1)) ground_threshold = 0.5;
memset(out, 0, sizeof *out);
out->n_evidence = n_evidence;
out->best = -1;
out->nearest_centroid_l2 = INFINITY;
out->scores = malloc((size_t)n_evidence * sizeof(double));
if (!out->scores) return -1;
double best = -1;
for (int i = 0; i < n_evidence; i++) {
const GeoDescriptor* e = evidence[i];
GeoFit f;
if (!e || e->dim != dim || !e->centroid ||
engram_reason_point_fit(e, claim, ext_floor, &f) != 0) {
out->scores[i] = 0.0;
continue;
}
out->scores[i] = f.score;
double cl2 = l2(claim, e->centroid, dim);
if (cl2 < out->nearest_centroid_l2) out->nearest_centroid_l2 = cl2;
if (out->best < 0 || f.score > best) {
best = f.score;
out->best = i;
out->grounding = f.score;
out->best_distance = f.distance;
out->best_ortho = f.ortho_residual;
}
}
if (out->best < 0) { out->grounding = 0.0; out->best_distance = INFINITY; }
out->grounded = (out->grounding >= ground_threshold) ? 1 : 0;
return 0;
}
void engram_verify_grounding_free(GeoGrounding* out) {
if (!out) return;
free(out->scores); out->scores = NULL;
}
/* ═══════════════════════════════════════════════════════ CONSISTENCY ════════ */
int engram_verify_consistency(const float* claim, int dim,
const GeoDescriptor* context,
const GeoDescriptor* pole_pos, const GeoDescriptor* pole_neg,
const GeoDescriptor* forbidden,
double ext_floor, double deadzone_frac,
double forbidden_thresh, double max_distance,
GeoConsistency* out) {
if (!claim || dim <= 0 || !out) return -1;
if (!(ext_floor > 0)) ext_floor = 1.0;
if (!(deadzone_frac >= 0 && deadzone_frac < 1)) deadzone_frac = 0.10;
if (!(forbidden_thresh > 0 && forbidden_thresh < 1)) forbidden_thresh = 0.5;
memset(out, 0, sizeof *out);
out->verdict = GEO_CONSIST_OK;
out->consistency = 1.0;
int do_polarity = (pole_pos && pole_neg);
int do_distance = (max_distance > 0);
if ((do_polarity || do_distance) &&
(!context || context->dim != dim || !context->centroid)) return -1;
if (do_polarity && (pole_pos->dim != dim || pole_neg->dim != dim ||
!pole_pos->centroid || !pole_neg->centroid)) return -1;
if (forbidden && (forbidden->dim != dim || !forbidden->centroid)) return -1;
double pol_score = 1.0, geo_score = 1.0;
/* ── (a) POLARITY / negation inversion ─────────────────────────────────── */
if (do_polarity) {
/* axis p = (c_pos c_neg); midpoint o = ½(c_pos + c_neg). */
float* p = malloc((size_t)dim * sizeof(float));
float* o = malloc((size_t)dim * sizeof(float));
if (!p || !o) { free(p); free(o); return -1; }
double pn2 = 0;
for (int i = 0; i < dim; i++) {
double dpos = (double)pole_pos->centroid[i], dneg = (double)pole_neg->centroid[i];
p[i] = (float)(dpos - dneg);
o[i] = (float)(0.5 * (dpos + dneg));
pn2 += (dpos - dneg) * (dpos - dneg);
}
double pn = sqrt(pn2);
out->polarity_separation = 0.5 * pn;
if (pn > 1e-12) {
/* signed positions along the axis (projection of (x o) onto unit p). */
float* cdo = malloc((size_t)dim * sizeof(float)); /* claim o */
float* rdo = malloc((size_t)dim * sizeof(float)); /* context o */
if (!cdo || !rdo) { free(p); free(o); free(cdo); free(rdo); return -1; }
for (int i = 0; i < dim; i++) {
cdo[i] = (float)((double)claim[i] - (double)o[i]);
rdo[i] = (float)((double)context->centroid[i] - (double)o[i]);
}
double claim_side = vdot(cdo, p, dim) / pn; /* units: emb-space length */
double ref_side = vdot(rdo, p, dim) / pn;
out->polarity_claim = claim_side;
out->polarity_reference = ref_side;
double dz = deadzone_frac * out->polarity_separation; /* neutral band */
if (fabs(claim_side) > dz && fabs(ref_side) > dz &&
(claim_side > 0) != (ref_side > 0)) {
out->inverted = 1;
pol_score = 0.0; /* opposite poles ⇒ zero consistency */
} else if (fabs(claim_side) <= dz || fabs(ref_side) <= dz) {
pol_score = 0.5; /* neutral / undecided */
} else {
pol_score = 1.0; /* same pole ⇒ consistent */
}
free(cdo); free(rdo);
}
free(p); free(o);
}
/* ── (b) GEOMETRIC contradiction ───────────────────────────────────────── */
if (forbidden) {
GeoFit f;
if (engram_reason_point_fit(forbidden, claim, ext_floor, &f) == 0) {
out->forbidden_fit = f.score;
if (f.score >= forbidden_thresh) {
out->geo_violation = 1;
double g = 1.0 - f.score; if (g < 0) g = 0;
if (g < geo_score) geo_score = g;
}
}
}
if (do_distance) {
out->context_distance = l2(claim, context->centroid, dim);
if (out->context_distance > max_distance) {
out->geo_violation = 1;
geo_score = 0.0;
}
}
/* ── verdict + scalar (polarity is the headline; both flags stay visible) ─ */
out->consistency = (pol_score < geo_score) ? pol_score : geo_score;
if (out->inverted) out->verdict = GEO_CONSIST_POLARITY;
else if (out->geo_violation) out->verdict = GEO_CONSIST_GEOMETRIC;
else out->verdict = GEO_CONSIST_OK;
return 0;
}
+118
View File
@@ -0,0 +1,118 @@
/* engram_verify.h — the VERIFIER layer: GROUNDING + CONSISTENCY over the live
* geometry (engram_geometry.h) and reasoning (engram_reason.h) operators.
*
* The geometry PROPOSES (cheap, creative, sometimes wrong); the verifier DISPOSES.
* This layer catches the class of failure a grammar check never sees: a fluent,
* confident, WRONG output the "plausible lie". The motivating case: a translation
* that DELETED a negation so "you never fought" became "you argued" reassurance
* inverted into accusation, grammatical and invisible, catchable ONLY by the geometry.
*
* GROUNDING claim is there ANY real structure that supports it, or is it
* floating free of the manifold? (anti-hallucination gate)
* CONSISTENCY claim does it CONTRADICT the established structure? Two catches:
* (a) POLARITY: the claim lands on the OPPOSITE side of a negation
* axis from the grounded truth (the reassuranceaccusation catch),
* (b) GEOMETRIC: the claim sits inside a region it must be far from,
* or violates a max-distance constraint to its context.
*
* PURE + READ-ONLY (stdlib + libm only): every function consumes a claim POINT
* (float* in R^dim) plus GeoDescriptor(s), and NEVER touches the store, index, or
* activation. All geometry is delegated to engram_reason_point_fit / engram_geo_*;
* this file only composes and applies thresholds.
*
* FRAME CONTRACT (inherited): the claim point and every descriptor passed together
* MUST share emb `dim` and the same `global_mean` frame exactly the §5 operator
* contract. A function returns <0 on a dim/frame mismatch or bad argument.
*/
#ifndef ENGRAM_VERIFY_H
#define ENGRAM_VERIFY_H
#include "engram_geometry.h"
#include "engram_reason.h"
/* ═══════════════════════════════════════════════════════════════════════════
* GROUNDING anti-hallucination. Score how well a claimed POINT is supported by
* the ACTUAL structure: fit the claim against every real evidence neighborhood
* (engram_reason_point_fit in-distribution Mahalanobis + off-model orthogonal
* residual) and take the BEST supporter. A claim that sits inside real structure
* scores high (grounded); a claim floating far from every neighborhood scores low
* on all of them flagged UNGROUNDED (a hallucination).
*
* This is an ABSOLUTE-THRESHOLD gate, deliberately distinct from ABDUCTION (which
* always RANKS and picks a winner among competing hypotheses): grounding asks the
* prior question "is there any real support at all?" and is allowed to answer no.
* The off-model `ortho_residual` is the sharpest hallucination signal: energy in a
* direction the manifold does not even span.
* */
typedef struct {
double grounding; /* ∈[0,1]: overall support = best fit score */
int grounded; /* 1 iff grounding >= ground_threshold */
int best; /* index of best-supporting evidence structure, or 1 */
double best_distance; /* full point-to-manifold distance to the best */
double best_ortho; /* off-model orthogonal residual of the best fit */
double nearest_centroid_l2;/* raw L2 to the nearest evidence centroid (coarse) */
int n_evidence;
double* scores; /* per-evidence fit score, higher = better (owned)*/
} GeoGrounding;
/* ext_floor>0 guards zero-extent axes (default 1.0). ground_threshold∈(0,1): the
* minimum best-fit score to call the claim grounded (default 0.5). */
int engram_verify_grounding(const float* claim, int dim,
const GeoDescriptor* const* evidence, int n_evidence,
double ext_floor, double ground_threshold,
GeoGrounding* out);
void engram_verify_grounding_free(GeoGrounding* out);
/* ═══════════════════════════════════════════════════════════════════════════
* CONSISTENCY contradiction detection. Does the claim contradict the established
* structure? Two independent sub-checks (either can fire; both flags are reported):
*
* (a) POLARITY / negation inversion. A polarity axis p is defined by two REAL
* poles pole_pos (asserts X) and pole_neg (asserts ¬X):
* p = (c_pos c_neg)/· , midpoint o = ½(c_pos + c_neg).
* The claim's side = p·(claim o); the reference's side = p·(c_context o).
* If the two sides have OPPOSITE sign AND both clear the neutral deadzone, the
* claim asserts the polarity opposite to the grounded truth INVERSION flagged.
* This is the "you never fought""you argued" catch: the truth ("never fought")
* sits on the negate pole, the claim ("argued") on the affirm pole opposite
* sides flagged, though every word is grammatical.
*
* (b) GEOMETRIC contradiction. The claim sits INSIDE a `forbidden` region it must
* be far from (point_fit score to forbidden forbidden_thresh), OR it violates
* a max-distance constraint to its context centroid (L2 > max_distance).
*
* pole_pos/pole_neg may both be NULL to skip the polarity check; forbidden may be
* NULL and max_distance0 to skip the geometric check. `context` (the grounded truth
* region) is required whenever polarity or the distance constraint is used.
* */
typedef enum {
GEO_CONSIST_OK = 0, /* consistent with context */
GEO_CONSIST_POLARITY = 1, /* polarity/negation inversion (asserts ¬X where X) */
GEO_CONSIST_GEOMETRIC = 2 /* geometric contradiction (in forbidden / too far) */
} GeoConsistencyVerdict;
typedef struct {
GeoConsistencyVerdict verdict; /* headline (polarity takes precedence) */
double consistency; /* ∈[0,1]: min over the checks (1 = fully consistent)*/
/* polarity sub-check */
int inverted; /* 1 iff a polarity inversion was detected */
double polarity_claim; /* p·(claim o) (signed position on the axis)*/
double polarity_reference; /* p·(c_context o) (the grounded truth's side) */
double polarity_separation; /* ½‖c_pos c_neg‖ (the axis half-length / scale)*/
/* geometric sub-check */
int geo_violation; /* 1 iff a geometric contradiction was detected */
double forbidden_fit; /* claim's point_fit score to the forbidden region*/
double context_distance; /* L2(claim, c_context) */
} GeoConsistency;
/* ext_floor>0 (default 1.0). deadzone_frac∈[0,1): a polarity side within
* deadzone_frac·separation of the midpoint is "neutral" and never triggers inversion
* (default 0.10). forbidden_thresh(0,1): fit-to-forbidden at/above which the claim
* counts as inside the forbidden region (default 0.5). max_distance>0 enables the
* distance constraint; 0 disables it. */
int engram_verify_consistency(const float* claim, int dim,
const GeoDescriptor* context,
const GeoDescriptor* pole_pos, const GeoDescriptor* pole_neg,
const GeoDescriptor* forbidden,
double ext_floor, double deadzone_frac,
double forbidden_thresh, double max_distance,
GeoConsistency* out);
#endif /* ENGRAM_VERIFY_H */
+689
View File
@@ -0,0 +1,689 @@
/* engram_vindex.c — HNSW ANN index over f32 embedding vectors (design §9 M8).
*
* Self-contained: plain C11, stdlib + libm (-lm for sqrtf/logf) only. No
* dependency on el_runtime; the store is read via its PERMANENT on-disk format
* (design §2.4), decoded read-only here so engram_store.{c,h} stay untouched.
*
* Algorithm: Malkov & Yashunin, "Efficient and robust approximate nearest
* neighbor search using Hierarchical Navigable Small World graphs" (2016).
* - multi-layer graph; level ~ Exp(1/ln M), assigned by a per-node seeded PRNG
* (deterministic: seed = FIXED_SEED ^ node_ordinal) so a rebuild is bit-for-
* bit reproducible regardless of wall-clock or global rand() state.
* - greedy descent through upper layers to an entry point, then an ef-bounded
* best-first search at each layer (Algorithm 2).
* - neighbour selection by the diversity heuristic (Algorithm 4), not plain
* k-nearest, with keep-pruned backfill for connectivity.
* - bidirectional links; a neighbour whose degree exceeds M (2M on layer 0) is
* re-pruned with the same heuristic.
*
* Metric: vectors are L2-normalised on entry, so cosine similarity == dot
* product; distance = 1 - dot (in [0,2], smaller == nearer). Deterministic tie-
* breaks are by element index so results are stable across identical builds.
*/
#include "engram_vindex.h"
#include <stdlib.h>
#include <string.h>
#include <math.h>
#include <stdio.h>
#include <stdint.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>
/* Deterministic PRNG seed base (fixed constant — never wall-clock/rand). */
#define VINDEX_FIXED_SEED 0x9E3779B97F4A7C15ULL
/* ── deterministic PRNG (splitmix64) ──────────────────────────────────────── */
static inline uint64_t splitmix64(uint64_t* s){
uint64_t z = (*s += 0x9E3779B97F4A7C15ULL);
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ULL;
z = (z ^ (z >> 27)) * 0x94D049BB133111EBULL;
return z ^ (z >> 31);
}
/* Uniform double in (0,1]. */
static inline double sm_uniform(uint64_t* s){
/* 53-bit mantissa; +1 keeps it in (0,1] so log() never sees 0. */
return ((double)((splitmix64(s) >> 11) + 1)) * (1.0 / 9007199254740993.0);
}
/* ── element + index structures ───────────────────────────────────────────── */
typedef struct {
int count;
int cap;
int* ids; /* neighbour element indices */
} NeighList;
typedef struct {
uint64_t node_id;
int level; /* top layer this element appears on (>=0) */
float* vec; /* dim floats, L2-normalised */
NeighList* links; /* level+1 lists; links[l] = neighbours at layer l */
} Elem;
struct VIndex {
int dim;
int M; /* max neighbours per node, upper layers */
int M0; /* == 2*M, layer 0 */
int ef_construction;
double mL; /* level normaliser = 1/ln(M) */
Elem* elems;
size_t n;
size_t cap;
int entry; /* entry-point element index, -1 if empty */
int max_level; /* current top layer */
/* scratch: version-stamped visited set (O(1) reset). */
uint32_t* visited;
uint32_t visit_epoch;
size_t visited_cap;
};
/* ── small helpers ────────────────────────────────────────────────────────── */
static float* vec_normalise_copy(const float* v, int dim){
float* out = (float*)malloc((size_t)dim * sizeof(float));
if (!out) return NULL;
double ss = 0.0;
for (int i=0;i<dim;i++) ss += (double)v[i]*(double)v[i];
if (ss > 0.0){
float inv = (float)(1.0 / sqrt(ss));
for (int i=0;i<dim;i++) out[i] = v[i]*inv;
} else {
for (int i=0;i<dim;i++) out[i] = 0.0f; /* zero vector stays zero */
}
return out;
}
/* Cosine distance between two normalised vectors: 1 - dot. In [0,2].
* Float accumulation in 4 lanes so the compiler auto-vectorises the hot path
* (this is the dominant cost of both build and search). */
static float vdist(const VIndex* ix, const float* a, const float* b){
int dim = ix->dim;
float s0=0,s1=0,s2=0,s3=0;
int i=0;
for (; i+4<=dim; i+=4){
s0 += a[i]*b[i]; s1 += a[i+1]*b[i+1];
s2 += a[i+2]*b[i+2]; s3 += a[i+3]*b[i+3];
}
float dot = (s0+s1)+(s2+s3);
for (; i<dim; i++) dot += a[i]*b[i];
return 1.0f - dot;
}
static int nl_push(NeighList* nl, int id){
if (nl->count == nl->cap){
int nc = nl->cap ? nl->cap*2 : 4;
int* np = (int*)realloc(nl->ids, (size_t)nc*sizeof(int));
if (!np) return -1;
nl->ids = np; nl->cap = nc;
}
nl->ids[nl->count++] = id;
return 0;
}
/* ── binary heaps over (dist,elem) pairs ──────────────────────────────────── */
typedef struct { float d; int e; } Pair;
typedef struct { Pair* a; int n, cap; } Heap;
static int heap_reserve(Heap* h, int need){
if (need <= h->cap) return 0;
int nc = h->cap ? h->cap*2 : 16;
while (nc < need) nc *= 2;
Pair* na = (Pair*)realloc(h->a, (size_t)nc*sizeof(Pair));
if (!na) return -1;
h->a = na; h->cap = nc; return 0;
}
/* Order predicate: for a MAX-heap on distance, "higher priority" = larger dist;
* ties broken by larger element index (deterministic + stable). is_max selects. */
static inline int pair_before(Pair x, Pair y, int is_max){
if (x.d != y.d) return is_max ? (x.d > y.d) : (x.d < y.d);
return is_max ? (x.e > y.e) : (x.e < y.e);
}
static int heap_push(Heap* h, Pair v, int is_max){
if (heap_reserve(h, h->n+1)) return -1;
int i = h->n++;
h->a[i] = v;
while (i > 0){
int p = (i-1)/2;
if (pair_before(h->a[i], h->a[p], is_max)){
Pair t=h->a[i]; h->a[i]=h->a[p]; h->a[p]=t; i=p;
} else break;
}
return 0;
}
static Pair heap_pop(Heap* h, int is_max){
Pair top = h->a[0];
h->a[0] = h->a[--h->n];
int i = 0;
for (;;){
int l=2*i+1, r=2*i+2, best=i;
if (l<h->n && pair_before(h->a[l], h->a[best], is_max)) best=l;
if (r<h->n && pair_before(h->a[r], h->a[best], is_max)) best=r;
if (best==i) break;
Pair t=h->a[i]; h->a[i]=h->a[best]; h->a[best]=t; i=best;
}
return top;
}
/* ── visited set ──────────────────────────────────────────────────────────── */
static int visited_ensure(VIndex* ix){
if (ix->visited_cap >= ix->cap && ix->visited) return 0;
size_t nc = ix->cap ? ix->cap : 16;
uint32_t* nv = (uint32_t*)realloc(ix->visited, nc*sizeof(uint32_t));
if (!nv) return -1;
if (nc > ix->visited_cap) memset(nv + ix->visited_cap, 0, (nc-ix->visited_cap)*sizeof(uint32_t));
ix->visited = nv; ix->visited_cap = nc;
return 0;
}
static inline void visited_reset(VIndex* ix){
if (++ix->visit_epoch == 0){ /* wrapped: clear all */
memset(ix->visited, 0, ix->visited_cap*sizeof(uint32_t));
ix->visit_epoch = 1;
}
}
static inline int is_visited(VIndex* ix, int e){ return ix->visited[e]==ix->visit_epoch; }
static inline void mark_visited(VIndex* ix, int e){ ix->visited[e]=ix->visit_epoch; }
/* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */
/* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller
* owns res->a. `q` is a normalised query. */
static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
int ef, int layer, Heap* res /*out, max-heap*/){
Heap cand = {0,0,0}; /* min-heap: nearest to expand */
res->a=NULL; res->n=0; res->cap=0;
visited_reset(ix);
for (int i=0;i<neps;i++){
int e = eps[i];
if (is_visited(ix,e)) continue;
mark_visited(ix,e);
float d = vdist(ix, q, ix->elems[e].vec);
Pair p = { d, e };
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
}
while (res->n > ef) heap_pop(res,1); /* trim to ef */
while (cand.n > 0){
Pair c = heap_pop(&cand,0);
float worst = res->a[0].d; /* farthest kept result */
if (res->n >= ef && c.d > worst) break;
Elem* ce = &ix->elems[c.e];
if (layer <= ce->level){
NeighList* nl = &ce->links[layer];
for (int i=0;i<nl->count;i++){
int e = nl->ids[i];
if (is_visited(ix,e)) continue;
mark_visited(ix,e);
float d = vdist(ix, q, ix->elems[e].vec);
if (res->n < ef || d < res->a[0].d){
Pair p = { d, e };
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
if (res->n > ef) heap_pop(res,1);
}
}
}
}
free(cand.a);
return 0;
}
/* ── neighbour selection heuristic (Algorithm 4) ──────────────────────────── */
/* From candidate pairs W (any order), pick up to M diverse neighbours of q.
* Keep c only if it is nearer to q than to every already-chosen neighbour;
* backfill from the pruned set (nearest first) to reach M for connectivity.
* Writes chosen element indices into out[], returns the count. */
static int select_neighbors(VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
(void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */
/* sort W ascending by (dist,elem) — deterministic. */
for (int i=1;i<nW;i++){ /* insertion sort (nW small) */
Pair key=W[i]; int j=i-1;
while (j>=0 && !pair_before(W[j],key,0)){ W[j+1]=W[j]; j--; }
W[j+1]=key;
}
int nout = 0;
Pair* pruned = (Pair*)malloc((size_t)(nW?nW:1)*sizeof(Pair));
int npr = 0;
if (!pruned) return -1;
for (int i=0;i<nW && nout<M;i++){
int good = 1;
for (int j=0;j<nout;j++){
float d = vdist(ix, ix->elems[W[i].e].vec, ix->elems[out[j]].vec);
if (d < W[i].d){ good = 0; break; } /* nearer an existing pick → drop */
}
if (good) out[nout++] = W[i].e;
else pruned[npr++] = W[i];
}
for (int i=0;i<npr && nout<M;i++) out[nout++] = pruned[i].e; /* keep-pruned backfill */
free(pruned);
return nout;
}
/* Re-prune a neighbour's over-full adjacency list back to `Mmax`. */
static void prune_links(VIndex* ix, int e, int layer, int Mmax){
NeighList* nl = &ix->elems[e].links[layer];
if (nl->count <= Mmax) return;
const float* base = ix->elems[e].vec;
Pair* W = (Pair*)malloc((size_t)nl->count*sizeof(Pair));
if (!W) return;
int nW = nl->count;
for (int i=0;i<nW;i++) W[i] = (Pair){ vdist(ix, base, ix->elems[nl->ids[i]].vec), nl->ids[i] };
int* keep = (int*)malloc((size_t)nW*sizeof(int));
if (!keep){ free(W); return; }
int nk = select_neighbors(ix, base, W, nW, Mmax, keep);
if (nk >= 0){ nl->count = nk; for (int i=0;i<nk;i++) nl->ids[i]=keep[i]; }
free(keep); free(W);
}
/* ── insert ───────────────────────────────────────────────────────────────── */
static int elems_reserve(VIndex* ix){
if (ix->n < ix->cap) return 0;
size_t nc = ix->cap ? ix->cap*2 : 64;
Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem));
if (!ne) return -1;
ix->elems = ne; ix->cap = nc;
return visited_ensure(ix);
}
int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
if (!ix || !vec) return -1;
if (elems_reserve(ix)) return -1;
int cur = (int)ix->n;
/* deterministic level assignment, seeded per-node. */
uint64_t seed = VINDEX_FIXED_SEED ^ (node_id + 0x2545F4914F6CDD1DULL*(uint64_t)cur);
int level = (int)(-log(sm_uniform(&seed)) * ix->mL);
if (level < 0) level = 0;
Elem* el = &ix->elems[cur];
el->node_id = node_id;
el->level = level;
el->vec = vec_normalise_copy(vec, ix->dim);
el->links = (NeighList*)calloc((size_t)level+1, sizeof(NeighList));
if (!el->vec || !el->links){ free(el->vec); free(el->links); return -1; }
ix->n++;
if (ix->entry < 0){ /* first element */
ix->entry = cur; ix->max_level = level;
return 0;
}
int ep = ix->entry;
int L = ix->max_level;
/* greedy descent through layers above `level` to refine the entry point. */
for (int lc = L; lc > level; lc--){
Heap r = {0,0,0};
int eps1[1] = { ep };
if (search_layer(ix, el->vec, eps1, 1, 1, lc, &r)){ return -1; }
if (r.n){ ep = r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} }
free(r.a);
}
/* from min(L,level) down to 0: connect. Each layer's ef-results seed the next
* layer's entry set; `eps` is heap-owned below the top and freed each step. */
int start = (L < level) ? L : level;
int eps_stack[1] = { ep };
int* eps = eps_stack; /* not owned (stack) until reassigned to malloc'd */
int* eps_owned = NULL;
int neps = 1;
int rc = 0;
for (int lc = start; lc >= 0; lc--){
int Mmax = (lc==0) ? ix->M0 : ix->M;
Heap W = {0,0,0};
if (search_layer(ix, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
if (!chosen){ free(W.a); rc=-1; break; }
int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen);
if (nc < 0){ free(chosen); free(W.a); rc=-1; break; }
/* link cur <-> chosen (bidirectional), prune neighbours if over-full. */
for (int i=0;i<nc;i++){
int nb = chosen[i];
if (nl_push(&el->links[lc], nb) || nl_push(&ix->elems[nb].links[lc], cur)){
free(chosen); free(W.a); rc=-1; goto done;
}
prune_links(ix, nb, lc, Mmax);
}
free(chosen);
/* next layer's entry points = this layer's ef results. */
if (lc > 0){
int* neweps = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
if (!neweps){ free(W.a); rc=-1; break; }
for (int i=0;i<W.n;i++) neweps[i]=W.a[i].e;
neps = W.n ? W.n : 1;
if (!W.n) neweps[0] = eps[0]; /* fall back to prior ep if empty */
free(eps_owned);
eps = eps_owned = neweps;
}
free(W.a);
}
done:
free(eps_owned);
if (rc) return -1;
if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; }
return 0;
}
/* ── search ───────────────────────────────────────────────────────────────── */
int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out){
if (!ix || !query || k <= 0) return -1;
if (ix->entry < 0) return 0;
if (ef_search <= 0) ef_search = VINDEX_DEFAULT_EF_SEARCH;
if (ef_search < k) ef_search = k;
float* q = vec_normalise_copy(query, ix->dim);
if (!q) return -1;
int ep = ix->entry;
for (int lc = ix->max_level; lc > 0; lc--){
Heap r = {0,0,0};
int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, 1, lc, &r)){ free(q); return -1; }
if (r.n){ int b=r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;}
ep = b; }
free(r.a);
}
Heap res = {0,0,0};
int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, ef_search, 0, &res)){ free(res.a); free(q); return -1; }
free(q);
/* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */
int total = res.n;
Pair* sorted = (Pair*)malloc((size_t)(total?total:1)*sizeof(Pair));
if (!sorted){ free(res.a); return -1; }
for (int i=total-1;i>=0;i--) sorted[i] = heap_pop(&res,1); /* farthest first out → fill from end */
free(res.a);
int out_n = (k < total) ? k : total;
for (int i=0;i<out_n;i++){
if (node_id_out) node_id_out[i] = ix->elems[sorted[i].e].node_id;
if (dist_out) dist_out[i] = sorted[i].d;
}
free(sorted);
return out_n;
}
size_t vindex_size(const VIndex* ix){ return ix ? ix->n : 0; }
VIndex* vindex_create(int dim, int M, int ef_construction){
if (dim <= 0) return NULL;
if (M <= 0) M = VINDEX_DEFAULT_M;
if (ef_construction <= 0) ef_construction = VINDEX_DEFAULT_EF_CONSTRUCTION;
VIndex* ix = (VIndex*)calloc(1, sizeof(VIndex));
if (!ix) return NULL;
ix->dim = dim;
ix->M = M;
ix->M0 = 2*M;
ix->ef_construction = ef_construction;
ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0);
ix->entry = -1;
ix->max_level = 0;
ix->visit_epoch = 0;
return ix;
}
void vindex_free(VIndex* ix){
if (!ix) return;
for (size_t i=0;i<ix->n;i++){
Elem* e = &ix->elems[i];
if (e->links) for (int l=0;l<=e->level;l++) free(e->links[l].ids);
free(e->links);
free(e->vec);
}
free(ix->elems);
free(ix->visited);
free(ix);
}
/* ── read-only decode of the paged store node format (design §2.4) ─────────── */
/* Mirrors engram_store.c constants; the on-disk format is PERMANENT so these are
* safe to duplicate for a read-only harvest of emb vectors. */
#define VS_PAGE_SIZE 16384u
#define VS_HDR 32u
#define VS_SLOT_SIZE 6u
#define VS_SLOT_LIVE 1u
#define VS_REC_HDR 4u
#define VS_REC_OVERFLOW 1u
#define VS_PT_NODE 1u
#define VS_OVF_NEXT 32u
#define VS_OVF_LEN 40u
#define VS_OVF_DATA 44u
#define VS_NT_ID 1u
#define VS_NT_EMB 24u
#define VS_NT_EMB_DIM 25u
static uint16_t vg_u16(const uint8_t* p){ return (uint16_t)(p[0] | (p[1]<<8)); }
static uint32_t vg_u32(const uint8_t* p){ uint32_t v=0; for(int i=0;i<4;i++) v|=(uint32_t)p[i]<<(8*i); return v; }
static uint64_t vg_u64(const uint8_t* p){ uint64_t v=0; for(int i=0;i<8;i++) v|=(uint64_t)p[i]<<(8*i); return v; }
static int vs_pread(int fd, uint64_t page, uint8_t* buf){
off_t off = (off_t)page * VS_PAGE_SIZE;
ssize_t r = pread(fd, buf, VS_PAGE_SIZE, off);
return (r == (ssize_t)VS_PAGE_SIZE) ? 0 : -1;
}
/* Read a (possibly overflowed) record body; caller frees *out. */
static int vs_read_body(int fd, const uint8_t* page, uint16_t off, uint16_t len,
uint8_t** out, size_t* outlen){
if (len < VS_REC_HDR) return -1;
uint8_t flags = page[off+3];
if (flags & VS_REC_OVERFLOW){
uint64_t head = vg_u64(page + off + VS_REC_HDR);
uint64_t total = vg_u64(page + off + VS_REC_HDR + 8);
uint8_t* body = (uint8_t*)malloc(total ? total : 1);
if (!body) return -1;
size_t got=0; uint64_t id=head;
uint8_t ov[VS_PAGE_SIZE];
while (id){
if (vs_pread(fd, id, ov)){ free(body); return -1; }
uint32_t chunk = vg_u32(ov + VS_OVF_LEN);
if (got + chunk > total){ free(body); return -1; }
memcpy(body+got, ov+VS_OVF_DATA, chunk); got += chunk;
id = vg_u64(ov + VS_OVF_NEXT);
}
if (got != total){ free(body); return -1; }
*out = body; *outlen = total;
} else {
uint16_t reclen = vg_u16(page + off);
if (reclen < VS_REC_HDR) return -1;
size_t blen = reclen - VS_REC_HDR;
uint8_t* body = (uint8_t*)malloc(blen ? blen : 1);
if (!body) return -1;
memcpy(body, page + off + VS_REC_HDR, blen);
*out = body; *outlen = blen;
}
return 0;
}
/* Extract id (strdup) and emb (malloc'd float[dim]) from a TLV node body. */
static void vs_parse_node(const uint8_t* body, size_t len, char** id_out,
float** emb_out, int* dim_out){
*id_out=NULL; *emb_out=NULL; *dim_out=0;
size_t i=0;
while (i + 5 <= len){
uint8_t tag = body[i];
uint32_t flen = vg_u32(body + i + 1);
if (i + 5 + (size_t)flen > len) break;
const uint8_t* v = body + i + 5;
if (tag == VS_NT_ID){
char* s = (char*)malloc(flen+1);
if (s){ memcpy(s,v,flen); s[flen]=0; free(*id_out); *id_out=s; }
} else if (tag == VS_NT_EMB){
int dim = (int)(flen/4);
float* e = (float*)malloc((size_t)(dim?dim:1)*sizeof(float));
if (e){ for (int k=0;k<dim;k++){ uint32_t u=vg_u32(v+k*4); memcpy(&e[k],&u,4);}
free(*emb_out); *emb_out=e; if(*dim_out==0) *dim_out=dim; }
} else if (tag == VS_NT_EMB_DIM){
*dim_out = (int)vg_u32(v);
}
i += 5 + flen;
}
}
/* Tiny open-addressing string set to dedup ids across live records. */
typedef struct { char** k; size_t cap, n; } StrSet;
static uint64_t vs_fnv(const char* s){ uint64_t h=1469598103934665603ULL; for(;*s;++s){h^=(uint8_t)*s;h*=1099511628211ULL;} return h; }
static int strset_add(StrSet* s, const char* key){ /* 1 added, 0 dup, -1 err */
if (s->n*2 >= s->cap){
size_t nc = s->cap ? s->cap*2 : 1024;
char** nk = (char**)calloc(nc, sizeof(char*));
if (!nk) return -1;
for (size_t i=0;i<s->cap;i++) if (s->k[i]){ size_t j=vs_fnv(s->k[i])&(nc-1); while(nk[j]) j=(j+1)&(nc-1); nk[j]=s->k[i]; }
free(s->k); s->k=nk; s->cap=nc;
}
size_t j = vs_fnv(key)&(s->cap-1);
while (s->k[j]){ if (strcmp(s->k[j],key)==0) return 0; j=(j+1)&(s->cap-1); }
char* d = strdup(key); if(!d) return -1;
s->k[j]=d; s->n++;
return 1;
}
static void strset_free(StrSet* s){ for(size_t i=0;i<s->cap;i++) free(s->k[i]); free(s->k); }
int vindex_harvest_from_store(const char* store_path, int dim,
float** vecs_out, char*** ids_out, int* n_out){
if (!store_path || dim <= 0 || !vecs_out) return -1;
int fd = open(store_path, O_RDONLY);
if (fd < 0) return -1;
struct stat st;
if (fstat(fd, &st) != 0){ close(fd); return -1; }
uint64_t npages = (uint64_t)st.st_size / VS_PAGE_SIZE;
float* vecs = NULL; size_t vn = 0, vcap = 0; /* row-major float[vn*dim] */
char** ids = NULL; size_t ids_n = 0, ids_cap = 0;
StrSet seen = {0,0,0};
uint8_t page[VS_PAGE_SIZE];
int failed = 0;
for (uint64_t pg = 2; pg < npages; pg++){ /* pages 0,1 = superblocks */
if (vs_pread(fd, pg, page)) continue;
if (page[8] != VS_PT_NODE) continue;
int slots = vg_u16(page + 10);
for (int sidx=0; sidx<slots; sidx++){
const uint8_t* sp = page + VS_HDR + (size_t)sidx*VS_SLOT_SIZE;
uint16_t off = vg_u16(sp), len = vg_u16(sp+2), fl = vg_u16(sp+4);
if (fl != VS_SLOT_LIVE) continue;
if ((size_t)off + VS_REC_HDR > VS_PAGE_SIZE) continue;
uint8_t* body=NULL; size_t blen=0;
if (vs_read_body(fd, page, off, len, &body, &blen)) continue;
char* id=NULL; float* emb=NULL; int edim=0;
vs_parse_node(body, blen, &id, &emb, &edim);
free(body);
if (!id || !emb || edim != dim){ free(id); free(emb); continue; }
int add = strset_add(&seen, id);
if (add <= 0){ free(id); free(emb); continue; } /* dup or err */
if (vn == vcap){
size_t nc = vcap ? vcap*2 : 1024;
float* nv = (float*)realloc(vecs, nc*(size_t)dim*sizeof(float));
if (!nv){ free(id); free(emb); failed = 1; goto out; }
vecs = nv; vcap = nc;
}
memcpy(vecs + vn*(size_t)dim, emb, (size_t)dim*sizeof(float));
free(emb);
if (ids_n == ids_cap){
size_t nc = ids_cap ? ids_cap*2 : 1024;
char** ni = (char**)realloc(ids, nc*sizeof(char*));
if (!ni){ free(id); failed = 1; goto out; }
ids = ni; ids_cap = nc;
}
ids[ids_n++] = id; /* transfers ownership */
vn++;
}
}
out:
close(fd);
strset_free(&seen);
if (failed){
free(vecs);
for (size_t i=0;i<ids_n;i++) free(ids[i]);
free(ids);
return -1;
}
*vecs_out = vecs;
if (n_out) *n_out = (int)vn;
if (ids_out){ *ids_out = ids; }
else { for (size_t i=0;i<ids_n;i++) free(ids[i]); free(ids); }
return (int)vn;
}
int vindex_build_from_store(VIndex* ix, const char* store_path,
char*** ids_out, int* n_out){
if (!ix || !store_path) return -1;
float* vecs = NULL; char** ids = NULL; int n = 0;
int h = vindex_harvest_from_store(store_path, ix->dim, &vecs, &ids, &n);
if (h < 0) return -1;
int inserted = 0;
for (int i = 0; i < n; i++){
if (vindex_insert(ix, (uint64_t)inserted, vecs + (size_t)i*ix->dim) != 0) break;
inserted++;
}
free(vecs);
if (ids_out){
*ids_out = ids; if (n_out) *n_out = inserted;
/* free any ids beyond what we inserted (insert failure tail) */
for (int i = inserted; i < n; i++) free(ids[i]);
} else {
for (int i = 0; i < n; i++) free(ids[i]);
free(ids);
if (n_out) *n_out = inserted;
}
return inserted;
}
/* ── optional persistence (index is rebuildable; convenience only) ─────────── */
#define VINDEX_SAVE_MAGIC "EGVIDX01"
int vindex_save(const VIndex* ix, const char* path){
if (!ix || !path) return -1;
FILE* f = fopen(path, "wb");
if (!f) return -1;
int ok = 1;
#define WR(p,n) do{ if(fwrite((p),1,(n),f)!=(size_t)(n)) ok=0; }while(0)
WR(VINDEX_SAVE_MAGIC, 8);
int32_t hdr[6] = { ix->dim, ix->M, ix->ef_construction, (int32_t)ix->n, ix->entry, ix->max_level };
WR(hdr, sizeof(hdr));
for (size_t i=0; ok && i<ix->n; i++){
Elem* e = &ix->elems[i];
WR(&e->node_id, sizeof(uint64_t));
int32_t lvl = e->level; WR(&lvl, sizeof(int32_t));
WR(e->vec, (size_t)ix->dim*sizeof(float));
for (int l=0; ok && l<=e->level; l++){
int32_t c = e->links[l].count; WR(&c, sizeof(int32_t));
WR(e->links[l].ids, (size_t)c*sizeof(int));
}
}
#undef WR
fclose(f);
return ok ? 0 : -1;
}
VIndex* vindex_load(const char* path){
FILE* f = fopen(path, "rb");
if (!f) return NULL;
char magic[8];
if (fread(magic,1,8,f)!=8 || memcmp(magic,VINDEX_SAVE_MAGIC,8)!=0){ fclose(f); return NULL; }
int32_t hdr[6];
if (fread(hdr,sizeof(hdr),1,f)!=1){ fclose(f); return NULL; }
VIndex* ix = vindex_create(hdr[0], hdr[1], hdr[2]);
if (!ix){ fclose(f); return NULL; }
size_t N = (size_t)hdr[3];
int ok = 1;
for (size_t i=0; ok && i<N; i++){
if (elems_reserve(ix)){ ok=0; break; }
Elem* e = &ix->elems[ix->n];
int32_t lvl;
if (fread(&e->node_id,sizeof(uint64_t),1,f)!=1 || fread(&lvl,sizeof(int32_t),1,f)!=1){ ok=0; break; }
e->level = lvl;
e->vec = (float*)malloc((size_t)ix->dim*sizeof(float));
e->links = (NeighList*)calloc((size_t)lvl+1, sizeof(NeighList));
if (!e->vec || !e->links){ free(e->vec); free(e->links); ok=0; break; }
if (fread(e->vec,sizeof(float),(size_t)ix->dim,f)!=(size_t)ix->dim){ ok=0; }
for (int l=0; ok && l<=lvl; l++){
int32_t c; if (fread(&c,sizeof(int32_t),1,f)!=1){ ok=0; break; }
e->links[l].ids = (int*)malloc((size_t)(c?c:1)*sizeof(int));
e->links[l].cap = c; e->links[l].count = c;
if (c && fread(e->links[l].ids,sizeof(int),(size_t)c,f)!=(size_t)c){ ok=0; }
}
ix->n++;
}
ix->entry = hdr[4]; ix->max_level = hdr[5];
fclose(f);
if (!ok){ vindex_free(ix); return NULL; }
return ix;
}
+94
View File
@@ -0,0 +1,94 @@
/* engram_vindex.h — M8 of the engram query engine: an approximate-nearest-
* neighbour (ANN) vector index over the node embedding vectors, for fast
* activation-seed selection.
*
* Replaces the O(n) cosine scan over emb vectors (design §9 M8; backlog #20)
* with an HNSW (Hierarchical Navigable Small World) graph that returns
* high-recall top-k seeds in ~O(log n).
*
* Standalone module: plain C11, stdlib + libm only. It does NOT modify the
* store format or engram_store.{c,h}; vindex_build_from_store() decodes the
* PERMANENT on-disk node format (design §2.4) read-only to harvest emb vectors.
*
* Similarity metric: cosine. Vectors are L2-normalised on insert/query, so
* cosine similarity == dot product. Reported distance = 1 - cosine_similarity
* (range [0,2]); smaller == closer. A query equal to an indexed vector scores
* distance ~0 against it.
*
* The index is fully rebuildable from the store, so persistence is optional for
* this milestone (see vindex_save/vindex_load below provided as a convenience;
* boot may simply rebuild via vindex_build_from_store()).
*/
#ifndef ENGRAM_VINDEX_H
#define ENGRAM_VINDEX_H
#include <stddef.h>
#include <stdint.h>
/* Tuned defaults (rationale in engram_vindex.c). Pass 0 to vindex_create for
* M / ef_construction to take these; pass ef_search<=0 to vindex_search for
* VINDEX_DEFAULT_EF_SEARCH. */
#define VINDEX_DEFAULT_M 24
#define VINDEX_DEFAULT_EF_CONSTRUCTION 200
#define VINDEX_DEFAULT_EF_SEARCH 128
typedef struct VIndex VIndex;
/* Create an index over `dim`-dimensional f32 vectors.
* M max neighbours per node on upper layers (2*M on layer 0).
* ef_construction candidate-list width during insert (recall/build cost).
* Pass M<=0 or ef_construction<=0 to use the VINDEX_DEFAULT_* above.
* Returns NULL on bad args / OOM. */
VIndex* vindex_create(int dim, int M, int ef_construction);
/* Insert one vector under an opaque caller-defined node_id (need not be unique,
* but the caller is responsible for meaning). `vec` has `dim` floats; it is
* copied and L2-normalised internally. A zero vector is accepted (it simply has
* distance ~1 to everything; never produces NaN). Returns 0 on success, <0 on
* error (bad args / OOM). */
int vindex_insert(VIndex* idx, uint64_t node_id, const float* vec);
/* Top-k search by cosine similarity. Writes up to k results (fewer if the index
* holds fewer than k elements) into node_id_out[] / dist_out[], ordered nearest
* first (ascending distance). Either out array may be NULL to skip it.
* ef_search search-time candidate width; larger == higher recall, slower.
* Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k.
* Returns the number of results written, or <0 on error. */
int vindex_search(VIndex* idx, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out);
/* Number of vectors currently indexed. */
size_t vindex_size(const VIndex* idx);
void vindex_free(VIndex* idx);
/* Build an index by scanning every live node record in the paged store at
* `store_path` (the on-disk format is decoded read-only; the store need not be
* open). Nodes without an emb vector, or whose emb_dim != idx->dim, are skipped.
* Each inserted node is assigned node_id = its 0-based insertion ordinal; if
* `ids_out`/`n_out` are non-NULL, *ids_out is set to a malloc'd array of that
* many strdup'd string ids (ids_out[node_id] == the store id) and *n_out to the
* count the caller frees each string and the array. Returns the number of
* vectors inserted, or <0 on error. */
int vindex_build_from_store(VIndex* idx, const char* store_path,
char*** ids_out, int* n_out);
/* Read-only harvest of the raw (un-normalised) emb vectors from a paged store,
* applying the SAME filtering vindex_build_from_store does (live records only,
* deduped by store id, emb present with emb_dim == `dim`), in insertion order.
* On success sets *vecs_out to a malloc'd float[n*dim] (row i == the i-th kept
* vector) and *n_out to n; if `ids_out` is non-NULL, sets it to a malloc'd array
* of n strdup'd store ids (ids_out[i] == the id of row i). Caller frees *vecs_out,
* each id string, and the id array. Returns n, or <0 on error. Used both by
* vindex_build_from_store (which then inserts each row) and by benchmarks/oracles
* that need the same vector set the index holds. */
int vindex_harvest_from_store(const char* store_path, int dim,
float** vecs_out, char*** ids_out, int* n_out);
/* Optional persistence (index is rebuildable from the store; provided for
* convenience). vindex_save writes a self-describing snapshot; vindex_load
* reconstructs an index from one. Return 0 / non-NULL on success. */
int vindex_save(const VIndex* idx, const char* path);
VIndex* vindex_load(const char* path);
#endif /* ENGRAM_VINDEX_H */
+238
View File
@@ -0,0 +1,238 @@
/* vindex_bench.c — standalone proof harness for the engram HNSW ANN index.
*
* Measures brute-force cosine top-k (the correctness ORACLE) vs vindex_search
* (HNSW) on: (a) the REAL paged store harvested read-only, and (b) synthetic
* clustered data at several sizes to trace the scaling curve. Reports build time,
* per-query latency (brute vs HNSW), and recall@k (HNSW top-k vs brute top-k).
*
* Read-only: never opens a socket, never writes the store. Safe on an nsbx clone.
*
* Build: cc -O2 -std=c11 vindex_bench.c engram_vindex.c -lm -o vindex_bench
* Usage: vindex_bench store <neuron.egm> <dim> [nqueries] [k] [ef_csv]
* vindex_bench synth <N> [dim] [clusters] [nqueries] [k] [ef_csv]
*/
#include "engram_vindex.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <math.h>
#include <stdint.h>
#include <time.h>
/* ── deterministic PRNG (splitmix64) so runs are reproducible ─────────────── */
static uint64_t g_seed = 0xD1B54A32D192ED03ULL;
static uint64_t sm(void){
uint64_t z = (g_seed += 0x9E3779B97F4A7C15ULL);
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ULL;
z = (z ^ (z >> 27)) * 0x94D049BB133111EBULL;
return z ^ (z >> 31);
}
static double urand(void){ return (double)((sm() >> 11) + 1) * (1.0/9007199254740993.0); }
static double grand(void){ /* Box-Muller */
double u1 = urand(), u2 = urand();
return sqrt(-2.0*log(u1)) * cos(2.0*M_PI*u2);
}
static double now_s(void){
struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts);
return (double)ts.tv_sec + (double)ts.tv_nsec*1e-9;
}
/* L2-normalise a row in place. */
static void l2norm(float* v, int dim){
double ss = 0; for (int i=0;i<dim;i++) ss += (double)v[i]*v[i];
if (ss > 0){ float inv = (float)(1.0/sqrt(ss)); for (int i=0;i<dim;i++) v[i]*=inv; }
}
/* Brute-force top-k by cosine distance (1 - dot on normalised vecs).
* data is n*dim, already L2-normalised. Writes k node ids (row indices) into
* out_ids ascending by distance. Returns nothing; assumes k<=n. */
static void brute_topk(const float* data, int n, int dim, const float* q,
int k, int* out_ids, float* out_d){
/* maintain a small sorted array of the k best (ascending distance). */
for (int i=0;i<k;i++){ out_ids[i]=-1; out_d[i]=2.0f+1.0f; }
for (int i=0;i<n;i++){
const float* r = data + (size_t)i*dim;
float s0=0,s1=0,s2=0,s3=0; int j=0;
for (; j+4<=dim; j+=4){ s0+=q[j]*r[j]; s1+=q[j+1]*r[j+1]; s2+=q[j+2]*r[j+2]; s3+=q[j+3]*r[j+3]; }
float dot=(s0+s1)+(s2+s3); for (; j<dim; j++) dot+=q[j]*r[j];
float d = 1.0f - dot;
if (d >= out_d[k-1]) continue;
int p = k-1;
while (p>0 && out_d[p-1] > d){ out_d[p]=out_d[p-1]; out_ids[p]=out_ids[p-1]; p--; }
out_d[p]=d; out_ids[p]=i;
}
}
/* recall@k: |brute_topk ∩ hnsw_topk| / k. Both are id arrays of length k. */
static double recall_at_k(const int* gt, const uint64_t* ann, int nann, int k){
int hit = 0;
for (int i=0;i<k;i++){
if (gt[i] < 0) continue;
for (int j=0;j<nann;j++){ if ((int)ann[j] == gt[i]){ hit++; break; } }
}
return (double)hit / (double)k;
}
/* Parse "64,128,256" into an int array; returns count. */
static int parse_csv(const char* s, int* out, int maxo){
int n=0; if(!s||!*s) return 0;
const char* p=s;
while(*p && n<maxo){ out[n++]=atoi(p); while(*p && *p!=',') p++; if(*p==',') p++; }
return n;
}
/* Generate n unit vectors on a LOW-DIMENSIONAL MANIFOLD, the property that makes
* real text embeddings tractable for ANN: each vector is a fixed random linear map
* A (dim × LATENT) applied to a latent gaussian z R^LATENT, plus small ambient
* noise, then L2-normalised. Points therefore lie near a `latent`-dim subspace, so
* every point has a well-defined tight neighbourhood (high recall) and the HNSW
* graph is cheap to build unlike near-isotropic 768-d gaussians, where the curse
* of dimensionality makes all points near-equidistant (no structure slow build,
* low recall) and unlike tight clusters (near-duplicates artificial top-k ties).
* `sigma` is the ambient-noise scale. This reproduces the intrinsic-dimensionality
* regime of nomic embeddings, so the scaling curve reflects real-corpus behaviour. */
#define SYNTH_LATENT 48
static void gen_synth(float* data, int n, int dim, int clusters, double sigma){
(void)clusters;
float* A = malloc((size_t)dim*SYNTH_LATENT*sizeof(float)); /* fixed random basis */
for (size_t i=0;i<(size_t)dim*SYNTH_LATENT;i++) A[i]=(float)grand();
float z[SYNTH_LATENT];
for (int i=0;i<n;i++){
for (int l=0;l<SYNTH_LATENT;l++) z[l]=(float)grand();
float* v = data+(size_t)i*dim;
for (int j=0;j<dim;j++){
float acc = (float)(sigma*grand());
const float* row = A + (size_t)j*SYNTH_LATENT;
for (int l=0;l<SYNTH_LATENT;l++) acc += row[l]*z[l];
v[j]=acc;
}
l2norm(v, dim);
}
free(A);
}
/* Build M / ef_construction come from env (VIDX_M / VIDX_EFC) so the scaling
* sweep can trade build cost against graph quality without a recompile. 0 = default. */
static int env_int(const char* k, int dflt){ const char* s=getenv(k); return (s&&*s)?atoi(s):dflt; }
/* Run the full brute-vs-HNSW comparison over an already-normalised dataset. */
static void run_bench(const char* label, float* data, int n, int dim,
int nq, int k, int* efs, int nef, double build_s){
(void)build_s;
int bM = env_int("VIDX_M", 0), bEFC = env_int("VIDX_EFC", 0);
printf("\n=== %s : N=%d dim=%d k=%d queries=%d ===\n", label, n, dim, k, nq);
/* build the index once (shared across ef settings). */
double t0 = now_s();
VIndex* ix = vindex_create(dim, bM, bEFC);
for (int i=0;i<n;i++) vindex_insert(ix, (uint64_t)i, data + (size_t)i*dim);
double bt = now_s()-t0;
printf("HNSW build: M=%d ef_construction=%d -> %.3f s (%.1f k nodes/s)\n",
bM?bM:VINDEX_DEFAULT_M, bEFC?bEFC:VINDEX_DEFAULT_EF_CONSTRUCTION, bt, n/1000.0/bt);
/* choose query vectors: perturb random dataset rows (near-but-not-identical). */
int* qidx = malloc((size_t)nq*sizeof(int));
float* qv = malloc((size_t)nq*dim*sizeof(float));
for (int i=0;i<nq;i++){
int r = (int)(sm() % (uint64_t)n);
qidx[i]=r;
float* dst = qv+(size_t)i*dim; const float* src = data+(size_t)r*dim;
for (int j=0;j<dim;j++) dst[j] = src[j] + (float)(0.01*grand());
l2norm(dst, dim);
}
/* ground truth: brute-force top-k for every query (also the oracle latency). */
int* gt = malloc((size_t)nq*k*sizeof(int));
float* gd = malloc((size_t)k*sizeof(float));
double tb0 = now_s();
for (int i=0;i<nq;i++) brute_topk(data, n, dim, qv+(size_t)i*dim, k, gt+(size_t)i*k, gd);
double brute_ms = (now_s()-tb0)*1000.0/nq;
printf("BRUTE-FORCE : %8.3f ms/query (oracle; O(N*D))\n", brute_ms);
/* HNSW at each ef. */
uint64_t* aid = malloc((size_t)k*sizeof(uint64_t));
float* ad = malloc((size_t)k*sizeof(float));
printf("%-6s %14s %12s %10s\n", "ef", "HNSW ms/query", "speedup", "recall@k");
for (int e=0;e<nef;e++){
int ef = efs[e];
double th0 = now_s();
double rec_sum = 0;
for (int i=0;i<nq;i++){
int m = vindex_search(ix, qv+(size_t)i*dim, k, ef, aid, ad);
rec_sum += recall_at_k(gt+(size_t)i*k, aid, m, k);
}
double hnsw_ms = (now_s()-th0)*1000.0/nq;
printf("%-6d %14.4f %11.1fx %10.4f\n", ef, hnsw_ms, brute_ms/hnsw_ms, rec_sum/nq);
}
free(qidx); free(qv); free(gt); free(gd); free(aid); free(ad);
vindex_free(ix);
}
int main(int argc, char** argv){
setvbuf(stdout, NULL, _IOLBF, 0); /* line-buffered so progress streams to a log */
if (argc < 2){ fprintf(stderr,"usage: %s store <path> <dim> [nq] [k] [ef_csv] | synth <N> [dim] [clusters] [nq] [k] [ef_csv] | sweep <dim> <N_csv> [nq] [k] [ef_csv]\n", argv[0]); return 2; }
int defef[8]; int ndef;
if (strcmp(argv[1],"sweep")==0){
if (argc < 4){ fprintf(stderr,"sweep needs <dim> <N_csv>\n"); return 2; }
int dim = atoi(argv[2]);
int Ns[16]; int nN = parse_csv(argv[3], Ns, 16);
int nq = (argc>4)?atoi(argv[4]):200;
int k = (argc>5)?atoi(argv[5]):10;
ndef = (argc>6)?parse_csv(argv[6],defef,8):parse_csv("64,128,200",defef,8);
for (int s=0;s<nN;s++){
int N = Ns[s];
float* data = malloc((size_t)N*dim*sizeof(float));
if (!data){ fprintf(stderr,"OOM at N=%d\n",N); continue; }
int clusters = N/100; if (clusters < 64) clusters = 64;
gen_synth(data, N, dim, clusters, 1.0);
char lbl[64]; snprintf(lbl,sizeof lbl,"SYNTH N=%d", N);
run_bench(lbl, data, N, dim, nq, k, defef, ndef, 0.0);
free(data);
}
return 0;
}
if (strcmp(argv[1],"store")==0){
if (argc < 4){ fprintf(stderr,"store needs <path> <dim>\n"); return 2; }
const char* path = argv[2]; int dim = atoi(argv[3]);
int nq = (argc>4)?atoi(argv[4]):500;
int k = (argc>5)?atoi(argv[5]):10;
ndef = (argc>6)?parse_csv(argv[6],defef,8):parse_csv("32,64,128,200,400",defef,8);
printf("Harvesting emb vectors from %s (dim=%d) ...\n", path, dim);
float* data=NULL; int n=0;
double t0=now_s();
int h = vindex_harvest_from_store(path, dim, &data, NULL, &n);
double harvest_s = now_s()-t0;
if (h < 0 || n == 0){ fprintf(stderr,"harvest failed (h=%d n=%d) — wrong dim or path?\n", h, n); return 1; }
printf("Harvested %d live embedded nodes in %.2f s\n", n, harvest_s);
for (int i=0;i<n;i++) l2norm(data+(size_t)i*dim, dim); /* oracle needs normalised */
if (nq > n) nq = n;
run_bench("REAL STORE", data, n, dim, nq, k, defef, ndef, 0.0);
free(data);
return 0;
}
if (strcmp(argv[1],"synth")==0){
if (argc < 3){ fprintf(stderr,"synth needs <N>\n"); return 2; }
int N = atoi(argv[2]);
int dim = (argc>3)?atoi(argv[3]):768;
int clusters = (argc>4)?atoi(argv[4]):200;
int nq = (argc>5)?atoi(argv[5]):500;
int k = (argc>6)?atoi(argv[6]):10;
ndef = (argc>7)?parse_csv(argv[7],defef,8):parse_csv("64,128,200",defef,8);
printf("Generating %d synthetic clustered vectors (dim=%d clusters=%d) ...\n", N, dim, clusters);
float* data = malloc((size_t)N*dim*sizeof(float));
if (!data){ fprintf(stderr,"OOM allocating %zu bytes\n", (size_t)N*dim*sizeof(float)); return 1; }
gen_synth(data, N, dim, clusters, 0.35);
char lbl[64]; snprintf(lbl,sizeof lbl,"SYNTH");
run_bench(lbl, data, N, dim, nq, k, defef, ndef, 0.0);
free(data);
return 0;
}
fprintf(stderr,"unknown mode '%s'\n", argv[1]);
return 2;
}