Files
el/lang/runtime/eg_metal_cosine.h
T
bigmerge 728207aabf engram: real Metal batch-cosine kernel, wired into vindex_bench's brute-force oracle
The original brief targeted engram_activate's O(N*D) cosq prescan, but #109
(this branch) already retires that loop algorithmically (HNSW seed selection
+ lazy memoized cosine) — GPU-accelerating a loop being deleted isn't real
work, so that target was dropped rather than forced.

Re-investigated for a genuine remaining GPU-shaped call site (not a
manufactured one): HNSW insert's candidate-list distance work is bounded-
degree (M=24-48) and sequential/adaptive — too fine-grained for a GPU
dispatch to pay off. No O(N^2) pairwise cosine pass exists (dedup only
checks the K=8 already-selected seed slots). No concurrent multi-query
traffic exists (server.el: the soul's curiosity loop is a single in-process
caller). vindex_bench.c's brute_topk — the correctness oracle this same PR
adds to validate HNSW recall — is the one real, unforced fit: genuine
1-query-vs-N-vectors, embarrassingly parallel, no adaptivity.

Adds:
  - eg_cosine_batch.metal: batched cosine kernel, single- and multi-query
    variants, same -2.0 dim-mismatch/zero-norm sentinel as eg_cosine().
  - eg_metal_cosine.h/.m: C-callable Objective-C bridge. Lazy one-time
    device/pipeline init, MTLResourceStorageModeShared buffers, returns
    false on ANY failure so callers fall back to the scalar CPU loop
    unconditionally — never partial, never throws.
  - eg_metal_cosine_stub.c: zero-dependency CPU-only implementation for
    non-Darwin builds (Linux CI) — same symbols, always returns false, no
    #ifdef needed at any call site.
  - build_vindex_bench.sh: one-command build, real bridge + Metal frameworks
    on Darwin, stub everywhere else.

vindex_bench.c: brute_topk_metal / brute_topk_metal_batch call the bridge,
falling back to the existing CPU brute_topk on any failure or
EL_METAL_COSINE=0. The multi-query batched path exists because the first
version (one GPU call per query) measured SLOWER than CPU at N~13.7k — it
re-uploaded the full N*D matrix every query. Fixed by uploading the matrix
once per query batch.

Measured against a real nsbx-sandboxed clone of the live store (never
:8742/:7770), 13,671 real embedded nodes, dim=768, 300 real queries:
  BRUTE-FORCE (CPU):    2.013 ms/query
  BRUTE-METAL (GPU):    0.117 ms/query   (17.2x)
  id-recall vs CPU oracle: 0.9990 over 300 queries
  same-rank |Δdist|: max 2.98e-07, mean 7.53e-08 (float32 rounding, not a bug)

Synthetic scaling sweep (13k -> 50k nodes, same dim/queries) shows the GPU
speedup holding (~11x) as N grows toward the mathematical-foundations doc's
1.3M-node target, with CPU brute-force cost growing linearly as expected.

Not wired into engram_activate or the daemon build (nsbx's _build_binary) —
vindex_bench is a standalone offline tool, not part of the request-serving
binary, so no engram_activate/server-latency claim is made here. The bridge
is a reusable primitive (single eg_cosine_batch_metal + batched
eg_cosine_batch_metal_multi) other call sites can adopt later without
re-deriving any of this.

Based on feat/reframe-region-setop (PR #109), not dev directly: the only
genuine batch-cosine call site (vindex_bench.c) exists solely on this
branch. Flagged explicitly in the PR description as a deliberate deviation
from the original "base off dev" instruction.
2026-08-15 16:48:01 -05:00

86 lines
4.3 KiB
C

/* eg_metal_cosine.h — C-callable bridge to the Metal batched-cosine kernel.
*
* Plain C11 header, safe to #include from el_runtime.c / vindex_bench.c on
* every platform. The implementation (eg_metal_cosine.m) only exists on
* Apple builds; on any other platform (or if Metal init fails for any
* reason at all — no supported GPU, shader compile error, OOM, sandboxing,
* whatever) eg_cosine_batch_metal() returns false and writes nothing, and
* the caller MUST fall back to its existing scalar per-node loop
* unconditionally. This function must never be allowed to crash or hang
* the engram.
*/
#ifndef EG_METAL_COSINE_H
#define EG_METAL_COSINE_H
#include <stdint.h>
#include <stdbool.h>
#ifdef __cplusplus
extern "C" {
#endif
/* Batched cosine similarity: one query vector against `n` node vectors.
*
* query — qdim floats, the query embedding. Raw/unnormalized.
* qdim — query dimensionality (e.g. 768 for nomic-embed-text).
* node_ptrs — array of n pointers, node_ptrs[i] pointing at a (possibly
* differently-owned, possibly NULL) float vector for node i.
* NOT required to be contiguous — this function performs the
* gather into a packed row-major matrix internally, exactly
* mirroring how EngramNode.emb is one malloc per node.
* node_dims — array of n ints, node_dims[i] = that node's real emb_dim
* (0 or mismatched vs qdim ⇒ that node scores -2.0, matching
* eg_cosine's null/dim-mismatch/zero-norm sentinel exactly).
* n — number of nodes.
* out_scores — caller-owned array of n doubles; out_scores[i] is filled
* with the cosine similarity of node i against query, or
* -2.0 for a null/dim-mismatched/zero-norm node — bit-for-bit
* the same contract as eg_cosine(node_ptrs[i], query, qdim).
*
* Returns true iff the GPU path ran and out_scores was fully populated.
* Returns false (out_scores left untouched) on ANY failure or unavailability
* — no Metal-capable device, shader compile failure, allocation failure,
* n<=0, qdim<=0, null query/node_ptrs/node_dims/out_scores. Never partial:
* either every element of out_scores was written, or none were.
*/
bool eg_cosine_batch_metal(const float* query, int32_t qdim,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores);
/* True iff a Metal device + compiled pipeline is available right now (cheap
* after the first call — cached). Purely informational (e.g. for a startup
* log line or /api/stats field); callers should still treat a false return
* from eg_cosine_batch_metal itself as the authoritative fallback signal. */
bool eg_cosine_batch_metal_available(void);
/* Multi-query batched cosine: nq query vectors against the SAME n node
* vectors, in one call. Uploads node_matrix once and reuses it for every
* query, instead of nq separate eg_cosine_batch_metal() calls each paying
* the full gather+upload cost — measured necessary: at N≈13.7k/dim=768,
* repeating the single-query call per query was slower than the CPU
* baseline; batching queries together is what makes the GPU path a real win
* at this shape. Use this whenever multiple queries will run against an
* unchanged (or rarely-changing) node population; use the single-query
* function above for a genuinely one-off comparison.
*
* queries — nq*qdim floats, row-major (query i at queries+i*qdim).
* out_scores — caller-owned nq*n doubles, row-major
* (out_scores[i*n+j] = cosine(queries[i], node j)), same
* -2.0 sentinel semantics as eg_cosine_batch_metal.
*
* Returns true iff the GPU path ran and out_scores was fully populated
* (all nq*n entries); false (untouched) on any failure/unavailability. */
bool eg_cosine_batch_metal_multi(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores);
#ifdef __cplusplus
}
#endif
#endif /* EG_METAL_COSINE_H */