728207aabfe3b20f9c0572a35ae44a403f0e1e77
The original brief targeted engram_activate's O(N*D) cosq prescan, but #109 (this branch) already retires that loop algorithmically (HNSW seed selection + lazy memoized cosine) — GPU-accelerating a loop being deleted isn't real work, so that target was dropped rather than forced. Re-investigated for a genuine remaining GPU-shaped call site (not a manufactured one): HNSW insert's candidate-list distance work is bounded- degree (M=24-48) and sequential/adaptive — too fine-grained for a GPU dispatch to pay off. No O(N^2) pairwise cosine pass exists (dedup only checks the K=8 already-selected seed slots). No concurrent multi-query traffic exists (server.el: the soul's curiosity loop is a single in-process caller). vindex_bench.c's brute_topk — the correctness oracle this same PR adds to validate HNSW recall — is the one real, unforced fit: genuine 1-query-vs-N-vectors, embarrassingly parallel, no adaptivity. Adds: - eg_cosine_batch.metal: batched cosine kernel, single- and multi-query variants, same -2.0 dim-mismatch/zero-norm sentinel as eg_cosine(). - eg_metal_cosine.h/.m: C-callable Objective-C bridge. Lazy one-time device/pipeline init, MTLResourceStorageModeShared buffers, returns false on ANY failure so callers fall back to the scalar CPU loop unconditionally — never partial, never throws. - eg_metal_cosine_stub.c: zero-dependency CPU-only implementation for non-Darwin builds (Linux CI) — same symbols, always returns false, no #ifdef needed at any call site. - build_vindex_bench.sh: one-command build, real bridge + Metal frameworks on Darwin, stub everywhere else. vindex_bench.c: brute_topk_metal / brute_topk_metal_batch call the bridge, falling back to the existing CPU brute_topk on any failure or EL_METAL_COSINE=0. The multi-query batched path exists because the first version (one GPU call per query) measured SLOWER than CPU at N~13.7k — it re-uploaded the full N*D matrix every query. Fixed by uploading the matrix once per query batch. Measured against a real nsbx-sandboxed clone of the live store (never :8742/:7770), 13,671 real embedded nodes, dim=768, 300 real queries: BRUTE-FORCE (CPU): 2.013 ms/query BRUTE-METAL (GPU): 0.117 ms/query (17.2x) id-recall vs CPU oracle: 0.9990 over 300 queries same-rank |Δdist|: max 2.98e-07, mean 7.53e-08 (float32 rounding, not a bug) Synthetic scaling sweep (13k -> 50k nodes, same dim/queries) shows the GPU speedup holding (~11x) as N grows toward the mathematical-foundations doc's 1.3M-node target, with CPU brute-force cost growing linearly as expected. Not wired into engram_activate or the daemon build (nsbx's _build_binary) — vindex_bench is a standalone offline tool, not part of the request-serving binary, so no engram_activate/server-latency claim is made here. The bridge is a reusable primitive (single eg_cosine_batch_metal + batched eg_cosine_batch_metal_multi) other call sites can adopt later without re-deriving any of this. Based on feat/reframe-region-setop (PR #109), not dev directly: the only genuine batch-cosine call site (vindex_bench.c) exists solely on this branch. Flagged explicitly in the PR description as a deliberate deviation from the original "base off dev" instruction.
Description
The Engram programming language — types as knowledge nodes, quantum-sealed prod target
199 MiB
Releases
5
El SDK (latest)
Latest
Languages
Emacs Lisp
95.1%
C
3.9%
HTML
0.3%
Python
0.2%
Shell
0.2%
Other
0.1%