Stop hand-rolling GPU kernels for batch cosine similarity — use ggml (the
MIT-licensed compute library underneath llama.cpp, installed standalone via
Homebrew) as the preferred backend, without ripping out PR #114's
carefully-verified hand-rolled Metal shader.
Structure: one stable public adapter (eg_cosine_batch.h, zero #ifdef at call
sites) backed by three selectable concrete Strategies behind an internal
vtable (eg_cosine_batch_strategy.h) chosen by a Factory (eg_cosine_batch.c):
- eg_cosine_batch_strategy_ggml.c — NEW. ggml + dynamically-loaded Metal
backend plugin (ggml_backend_load_all_from_path
+ ggml_mul_mat for the batched dot
product), gather/scatter around the
-2.0 sentinel contract.
- eg_cosine_batch_strategy_metal_hand.m — PR #114's original hand-rolled
Metal shader bridge, preserved
almost verbatim, now one strategy
among several rather than the only
option. eg_cosine_batch.metal kept
byte-identical to the original.
- eg_cosine_batch_strategy_cpu.c — universal always-false fallback
(direct descendant of PR #114's
eg_metal_cosine_stub.c).
Selection: EL_COSINE_BATCH_STRATEGY=ggml|metal|cpu|auto (default: ggml first,
then hand-rolled Metal, then CPU — first available wins), plus back-compat
EL_METAL_COSINE=0 to disable every GPU-backed strategy. build_vindex_bench.sh
compiles all three strategies on Darwin, CPU-fallback-only elsewhere.
vindex_bench.c now reports BRUTE-GGML and BRUTE-METAL side by side against
the same CPU oracle, on the same dataset, in one run (real numbers vs. real
store snapshot in the PR body).