Compare commits

...

101 Commits

Author SHA1 Message Date
Neuron 688f24b4c1 ingest: name the inversion, and correct the worked example to decomposition
El SDK CI - dev / build-and-test (pull_request) Failing after 14m6s
ingest.el's transduce() was renamed to transduce_manifold() earlier the same
day on the reasoning that it 'was never signal->geometry -- it chunks
already-extracted content and PACKS it into a node+edge manifold, one layer up,
and it had taken the name that belongs to the primitive underneath it.'

That reasoning was backwards. Producing a node+edge manifold is not a layer
above transduction, it IS transduction. Signal -> one vector is the operation
underneath, and its name is geometry. The layer doing it right was renamed out
of the way so the layer doing it wrong could have the name.

With the primitive corrected to return a Manifold, the two layers do the same
kind of thing and the inversion dissolves. What is left is a real distinction
about MODALITY, not layering: transduce() dispatches to a realizer that knows
its modality and can name its components; transduce_bytes() is the
opaque-bytes realizer, the decomposition available to a reader that knows
nothing about what it is reading. It still yields components and relations,
which is why it is transduction and not packing -- it just cuts on byte
boundaries, so its components are positional rather than meaningful. That is a
limitation of this realizer, not the definition of the operation.

Renamed by modality rather than demoted by layer. A distinct symbol is still
mechanically required: reusing transduce here is a conflicting-types error the
moment ingest.c links el_runtime.c.

lang/examples/transduce.el asserted #144's contract and would now fail, so it
is replaced by the decomposition worked example: transduce a chord, persist the
five components and six relations as real nodes and edges, read each part's
geometry back off its own node, and ground one part while its sibling is
demonstrably untouched.
2026-08-16 15:50:00 -05:00
Neuron d777936ee4 runtime: transduction decomposes a signal, it does not convert it
#144 moved transduction into the language and got the dispatch right. It got
the result type wrong: transduce(signal, modality) -> Geometry yields one
vector per signal, and one vector is a fingerprint. A fingerprint can be
matched and ranked; that is all. It cannot be decomposed, cannot have one part
grounded while another is not, and cannot be contradicted in one part while
holding in another, because it has no parts.

A song is not a point. It decomposes into pitch, interval, rhythm, harmonic
function -- components, each with its own geometry, plus the relations among
them. The song IS the structure of the relations.

transduce now returns a Manifold: named components carrying geometry, and
typed weighted relations between them. Signal in, subgraph out.

Components are addressed by key, never by index, because the key is what
survives persistence -- a component becomes a node and is separately groundable
precisely because it is separately named. Relation weight IS the grounding
(correspondence-and-censorship.md 1), so a realizer's relations arrive already
grounded and there is no score computed beside them.
2026-08-16 15:50:00 -05:00
will.anderson 4a57b4faa8 Merge pull request 'docs: the builtin recipe never required a test' (#154) from docs/builtin-recipe-gate into dev
El SDK CI - dev / build-and-test (push) Failing after 3m53s
2026-08-16 20:49:21 +00:00
will.anderson 0ee82d9e91 Merge pull request 'Grounding is the edge's weight, and the weight is a vector' (#150) from feat/grounding-gradient into dev
El SDK CI - dev / build-and-test (push) Failing after 3m43s
2026-08-16 20:49:12 +00:00
will.anderson 9526bda507 Merge pull request 'engram: expose the geometry so the frame can be verified' (#156) from fix/geometry-readable into dev
El SDK CI - dev / build-and-test (push) Failing after 4m7s
2026-08-16 20:49:07 +00:00
Neuron 0389bf9363 engram: expose the geometry so the frame can be verified
El SDK CI - dev / build-and-test (pull_request) Failing after 11m29s
engram_scan_nodes_emb_json has existed as a builtin with NO ROUTE. The
embeddings — the actual positions every distance, angle, membership and
grounding is computed from — were unreadable from outside the process.

That is not a missing convenience. It means every claim about the
coordinate frame was unfalsifiable from the API: whether the space is
isotropic, where the centering offset sits, what the origin is, whether a
node carries geometry at all. You cannot verify a coordinate system you
cannot see, and a system whose frame cannot be checked is exactly the
shape this codebase spent 2026-08-16 removing everywhere else.

GET /api/nodes/emb?limit=&offset=. Read-only, paged, no writes.

Measured consequence of having it: the value manifold and the love
component manifold were both decomposed, null-controlled against random
node sets drawn from the same graph, and several published claims were
retracted because the geometry contradicted them. None of that was
possible before this route existed.
2026-08-16 15:37:22 -05:00
Neuron fe820928b0 docs: the builtin recipe never required a test
El SDK CI - dev / build-and-test (pull_request) Failing after 10m55s
lang/AGENTS.md:71-77 gives four steps for adding a C builtin and ends at
'confirm the self-host fixpoint is byte-identical'. No step asks for a test.
The only 'verify' in the file is that fixpoint, which proves the COMPILER
REPRODUCES ITSELF and says nothing about whether the builtin works — so the
recipe reads as complete while having checked nothing about the thing just
added.

Measured on 2026-08-16: engram_node_set_emb, engram_curiosity_json and
dream_set_handler were all added in a single session with zero tests, by an
agent following this recipe. Separately a UTF-8 fix was written and tested
and THE TEST PASSED ON THE UNPATCHED BUILD — the real defect was elsewhere,
and only building the pre-fix binary exposed it. Without a negative control
that fix would have merged as verified.

Adds step 5 with the two failure shapes actually encountered: a test that
never exercises the change (a route default bypassed the code under test),
and an induction that loses a race (curl --max-time left BOTH builds alive;
only SO_LINGER 0, a real RST, reproduced it). Plus the port-binding check,
because a stale instance answering has silently produced false results here
more than once and pkill -f does not reliably match argv './engram'.

Documentation only. Does not touch the (a) split-the-C / (b) close-the-
compiler-gap question, which is a separate decision.
2026-08-16 13:53:08 -05:00
will.anderson 385c18442d runtime: a disconnecting client must not kill the server (#151)
El SDK CI - dev / build-and-test (push) Failing after 3m53s
2026-08-16 18:33:50 +00:00
Neuron cace6a5ebf runtime: a disconnecting client must not kill the server
El SDK CI - dev / build-and-test (pull_request) Failing after 13m45s
There was no SIGPIPE handling anywhere in this runtime: no signal
disposition, no MSG_NOSIGNAL, no SO_NOSIGPIPE, and send() called with bare
flags. The default disposition of SIGPIPE is to TERMINATE THE PROCESS, so
any client that hangs up mid-response takes the whole engram with it.

MEASURED, and it is not hypothetical. Production has restarted 254 times
since 2026-08-13T19:37 at a flat ~10 minute cadence:

  17:05:18  17:15:29  17:25:38  17:35:50  17:46:00  17:56:10  18:06:22  18:16:30

Intervals of 10m09s-10m12s, not 10m00s. That excess is the whole story:
ai.neuron.engram-tick has StartInterval 600, and engram-tick.sh:13 calls

  curl -s -m10 -X POST .../api/tick

The beat does not finish within 10s over 13,634 nodes, so curl waits its
full timeout and closes. The engram then writes the tick response to a dead
socket, takes SIGPIPE, and dies. launchd KeepAlive restarts it, so the
failure presents as a mysterious restart rather than a crash — and
~/.neuron/logs/engram.log records nothing but "[http] listening on" 254
times, with no exit reason. launchctl list confirms the last exit as -13.

Root cause is one level out: consolidation had no owner, so an external
ticker was created to poke it, and the ticker is what kills it. The fix
here does not address that; it makes the process survivable while it is
addressed.

Two layers, because neither alone is portable:
  - SO_NOSIGPIPE per accepted socket (Darwin/BSD) and MSG_NOSIGNAL per send
    (Linux), so the signal is never raised for socket writes at all.
  - A process-wide SIG_IGN backstop, installed once and idempotent, for
    platforms and paths with neither. With the signal ignored, send()
    returns -1/EPIPE and the existing error path closes the connection.

Also retries send() on EINTR, which the previous loop treated as fatal.

This is an exemption in the sense of lang/spec §8: the write never checked
whether the peer was still there, and the consequence of not checking was
fatal rather than merely wrong.
2026-08-16 13:25:14 -05:00
Neuron 7a1501d097 Grounding is the edge's weight, and the weight is a vector
El SDK CI - dev / build-and-test (pull_request) Failing after 3m59s
A relation that keeps holding up strengthens; one that stops corresponding
decays. That is not analogous to grounding, it IS grounding — so it belongs on
the edge, not in a subsystem beside it. The graph was already the grounding
structure; this stops modelling it as something else.

Deleted, not refactored:
  - cog_ground_edge and the `grounded-by` relation type. A grounded-by edge
    models grounding as a relation BETWEEN nodes when it is a property OF a
    relation. #147 fixed which endpoints that edge landed on and left the wrong
    idea intact. Measured on the live store: the old path scored two nodes with
    ZERO edges between them at 0.925237 and wrote an edge for it.
  - ground() writing. It was a read that wrote — the eg_vindex_sync defect.
    Three identical calls produced three writes to the same edge id.
  - keystone_write_blocked. Its measured cost was 0.00% brier reduction over
    n_trials 0 on the keystone: the loop never ran, so the self was never
    calibrated and never falsifiable. Nothing replaces it — non-circularity of
    the reference frame is temporal, not a permission.
  - a graph predicate for "evidence downstream of itself", built and then
    withdrawn. Reachability from the self region covers 89.2% of the live graph
    (10,580 of 11,861 nodes), so any topological predicate marks nearly all
    evidence tainted and degenerates into the total block censorship began as.

The vector, carried in a GRD1 block on the edge's own metadata:
factual, relational, associative (the existing hebb), polarity (SIGNED — near
zero is "no support", negative is "actively contradicts"; `inhibitory` is that
distinction crushed to one bit), provenance class, and a timestamp. Confidence,
recency, staleness and volatility are DERIVED at read and never serialized.

Decay is one model, not two: cog_decay_factor is the single implementation and
engram_temporal_decay now delegates to it — proven bit-identical over 24
(age, reinforcement) points.

Values reference: thirteen regions, aggregate MIN, binding value named. Measured
— the 13 have pairwise centroid cosine min 0.1525 / mean 0.5199 / max 0.9278, so
they demonstrably are not one region, and a mean would let agreement with twelve
mask a violation of the thirteenth.

Supersession versions the whole vector jointly, gated by consequence and
salience with no epsilon anywhere: floor crossings and sign changes only.
Polarity flips and provenance-class changes are inherently significant and
bypass the salience gate.

Also fixed: the frame contract. Descriptors are built over L2-normalized member
embeddings; think() and the grounding path were fitting RAW vectors against them.
Measured on the self region, same data, same 106 members:
  magnitude 0.00283443 -> 0.536134, spread 18.7565 -> 0.930163.
Every fit score sat three decimal places below the 0.5 floors that gate on them.

assert() gates on both floors and computes still_held instead of returning a
hardcoded `true` — the old build reported still_held for a node that does not
exist.
2026-08-16 13:18:50 -05:00
will.anderson d41645388a runtime: make valid UTF-8 the JSON emitter's contract (#148)
El SDK CI - dev / build-and-test (push) Failing after 4m12s
El SDK CI - dev / build-and-test (pull_request) Failing after 4m37s
2026-08-16 17:03:44 +00:00
Neuron 8a307dfd42 runtime: make valid UTF-8 the JSON emitter's contract
El SDK CI - dev / build-and-test (pull_request) Failing after 10m36s
Three nodes in the live graph carry labels truncated to exactly 80 bytes
ending in a lone 0xE2 — the first byte of an em-dash, cut mid-sequence.
jb_emit_escaped copied every byte >= 0x20 through verbatim, so those three
nodes made the ENTIRE /api/nodes/list response undecodable and no strict
parser could read the graph at all.

  production binary   25,929,607 bytes   INVALID at byte 89260
  this build          26,338,389 bytes   VALID, parses to 13,630 nodes

The damage was NOT written by this runtime. No 80-byte truncation exists
here (the only label truncation is engram_first_n_chars at 60), and the
content of those nodes is 2572 and 2746 bytes. Some other producer wrote
them. That is exactly why fixing a writer could not have fixed this: the
store already holds the damage, and it accepts data from importers, other
producers and older binaries.

So the fix goes where the promise is made. A serializer that emits JSON
owes valid UTF-8 whatever it is handed. jb_emit_escaped now validates each
multi-byte sequence before emitting any of it and substitutes U+FFFD for a
bad lead byte, a missing or malformed continuation, an overlong encoding, a
UTF-16 surrogate, or a codepoint above U+10FFFF. Invalid bytes are REPLACED
rather than dropped, so the damage stays visible in the output instead of
being silently papered over. Well-formed input is byte-identical to before.

Second, preventive and explicitly NOT the cause of the above:
engram_first_n_chars truncated by BYTES despite its name, so content with a
multi-byte character crossing byte 60 would produce a half codepoint in the
label. It now uses el_utf8_safe_len, which returns the largest byte length
<= max that does not split a codepoint. Bounded by bytes, not codepoints,
so existing labels never grow — they only stop splitting.

el_utf8_safe_len lives beside str_count_chars rather than in the engram
because the rest of el's string layer is already codepoint-aware
(str_count_chars counts codepoints, str_reverse walks codepoint lengths).
Byte truncation was the outlier and the concern is a string concern.

Note on the investigation: I first "fixed" the truncator and wrote a test
that passed on the UNPATCHED build too, because route_create_node passes
label = content when no label is supplied, so engram_first_n_chars is never
reached over HTTP. The test proved nothing. The real cause was only found
by decoding the actual failing bytes out of the live response.
2026-08-16 12:03:03 -05:00
will.anderson 616815b2ab Give cross-cutting concerns an owner instead of a convention (#145)
El SDK CI - dev / build-and-test (push) Failing after 11m4s
2026-08-16 16:57:51 +00:00
will.anderson 1a8a966cb3 runtime: transduction is a language concern, so move it into the language (#144)
El SDK CI - dev / build-and-test (push) Failing after 11m29s
2026-08-16 16:57:35 +00:00
will.anderson 1f70b9fa18 runtime: ground the node asked about, and refuse circular support (#147)
El SDK CI - dev / build-and-test (push) Failing after 14m46s
2026-08-16 16:54:17 +00:00
Neuron 317466e8f7 runtime: ground the node asked about, and refuse circular support
El SDK CI - dev / build-and-test (pull_request) Failing after 15m5s
engram_ground_json resolved each seed to a REGION, wrote the grounded-by
edge between the two regions' HUBS, and then echoed those hubs back in the
"claim"/"evidence" fields as if they were the caller's input:

    const char* cid = C->hub_id ? C->hub_id : EL_CSTR(claim);
    const char* eid = E->hub_id ? E->hub_id : EL_CSTR(evidence);
    cog_ground_edge(g_engram_store, cid, eid, grounding, fw);

Three consequences, all measured against a clone of the live store:

1. The edge landed on a node the caller never named. Grounding 3b9ced5d
   against 6edf8c79 wrote an edge on the hubs of their regions instead.
2. When both seeds resolve into the same region the support is circular
   and scores near 1.0 for structural reasons, not evidential ones. Four
   probe nodes written together landed in one region, and every grounding
   among them returned 0.93-0.99 as if it were evidence. Two independent
   agents hit this and reported 0.885 / 0.909 self-groundings as confident.
3. The echo concealed both: the response was indistinguishable from a
   successful grounding of the ids that were passed in.

The region is HOW a claim is evaluated; it is not WHAT the claim is about.
So the edge now attaches to the requested ids, and the resolved hubs are
reported separately as claim_region / evidence_region.

Degeneracy is broader than hub == hub. Three circular shapes, all
previously invisible:
    same-region                both seeds resolve to one region
    claim-region-is-evidence   the evidence IS the hub of the claim's own
                               neighbourhood — measured at 0.98883
    evidence-region-is-claim   the mirror case
Each sets grounding to 0 and writes no edge. Circular support is not
support, and a grounding that is degenerate by construction must not
enter the graph as though it were evidence.

Verified:
  6edf8c79 -> 6edf8c79   degenerate=same-region   g=0        written=false
  6edf8c79 -> d0406dfd   degenerate=same-region   g=0        written=false
  ebc1413e -> 64cc96ef   degenerate=false         g=0.774563 written=true
  64cc96ef -> ebc1413e   degenerate=false         g=0.802896 written=true
Legitimate grounding across distinct regions is unchanged and still
writes; only circular support is refused.

This is the same class as #142 and #146 — a value that looked like an
answer with nothing behind it — except here it was also writing that
non-answer into the canonical store.
2026-08-16 11:53:31 -05:00
will.anderson eb3e6d7c1f runtime: resume the learned stance in think (#146)
El SDK CI - dev / build-and-test (push) Failing after 3m54s
2026-08-16 16:44:15 +00:00
Neuron 88e3008735 runtime: resume the learned stance in think, instead of discarding it
El SDK CI - dev / build-and-test (pull_request) Failing after 4m16s
engram_think_json built a NEUTRAL stance on every call — cog_stance_init
with a NULL id, all axis_gain 1.0, bias_dir NULL, reliability 0.5 — and
never loaded the stance the correspondence-beat had been persisting.

That mattered because the faculty enters engram_think ONLY through the
stance: axis_gain[k] warps the per-axis extents and bias_dir seeds the
steering direction. cog_stance_init stores the faculty NAME and nothing
reads it. So with a neutral stance, reason/abduce/induce/plan/analogize
were byte-identical output under different labels, and confidence was
pinned to 0.5 because GeoGradient.confidence IS stance->reliability.

The machinery already existed and only this call site ignored it.
engram_correspondence_beat_json resumes via cog_stance_from_node and
persists via cog_stance_to_node under "stance-<faculty>-<hub>". Every
beat's calibration was written and then thrown away on the next read.
Same defect as the NULL anchor fixed in #142, one line below: a neutral
argument collapsing a capability to a constant.

Resume the same id the beat writes, so learning compounds across beats and
cold boot. Fall back to neutral only when no stance exists — a genuine
uninformed prior rather than a discarded informed one.

Also emit stance_resumed, so confidence 0.5 from a learned-but-unreliable
stance is distinguishable from confidence 0.5 from "no stance exists".
That reporting gap is what let the neutral stance hide.

Verified against a clone of the production store (13,627 nodes):

  before beat, no stance     stance_resumed=false  confidence=0.5
  beat on a NON-keystone     brier 0.00458568 -> 0.00329654
                             reduction 28.11%, n_trials 6000,
                             reliability 0.930726, stance_written=true
  after beat                 stance_resumed=true   confidence=0.930726

Confidence now equals the learned reliability instead of the uninformed
prior. The keystone self-anchor correctly stays at 0.5 — calibration is
deliberately refused on protected identity regions, and that refusal is
now visible as resumed=true with confidence unchanged, rather than being
indistinguishable from the bug.

STILL OPEN: with no learned bias_dir the faculties remain identical in
direction. What distinguishes abduce from induce geometrically is a
design decision about how Neuron thinks, not a plumbing defect, and is
deliberately left to Will.
2026-08-16 11:43:38 -05:00
bigmerge 26af149aa1 lang: rebuild the bootstrap compiler against merged dev
El SDK CI - dev / build-and-test (pull_request) Failing after 4m46s
The binary was stamped before dev advanced (vindex publication landed in
el_runtime.c and engram_vindex.c). Rebuilt against the merged runtime so the
committed compiler matches the runtime it ships beside. Fixpoint re-verified
byte-identical; test_compiler 82/82; engram/src/server.el still compiles and
still emits its 18 config declarations.
2026-08-16 11:38:52 -05:00
bigmerge c18abf799c engram: declare configuration once instead of at every read site
Migrates engram to the `program` block. 18 configuration variables that each
carried their default inline at the point of use now declare it in one place,
and engram declares itself a singleton.

The read sites lose their defaults entirely: `let v = env("X")` followed by
`if str_eq(v,"") { "default" } else { v }` collapses to `config("X")`. The
guide_env_or(key, dflt) helper is deleted -- its whole job was supplying a
per-site default, which is the thing being removed.

Fixes ENGRAM_DATA_DIR, which was the clearest instance of the defect. It was
read at six sites. Five were dead: `let dir_raw = env("ENGRAM_DATA_DIR")`
immediately shadowed on the next line by `engram_resolve_data_dir()`. The sixth
was live and defaulted to /tmp/engram, contradicting the canonical resolver's
$HOME/.neuron/engram -- and its consumer is the pre-destructive reseed backup,
so with ENGRAM_DATA_DIR unset the safety copy was written to ephemeral storage
while the store it protected lived elsewhere. All six now go through
engram_resolve_data_dir().

ENGRAM_DATA_DIR is deliberately NOT declared in the program block, and the
source says why: engram_resolve_data_dir() already owns it, and a second
declaration would give it two owners that can disagree -- recreating the exact
defect being removed here. A variable belongs in the block when the block would
be its only owner. HOME stays a raw env() read; it is an environment fact, not
configuration.

singleton: "engram" matters more than it looks. Today a second engram whose
bind() fails merely returns from http_serve -- after it has already replayed
the WAL and written boot-time backup files -- and then exits 0, indistinguishable
from a clean run. That is how two instances came to share one data dir. Verified
that the second instance now refuses before any side effect: with instance 1
holding the lock (lsof pid, shell pid, and lock file contents all agreeing at
5946), the second start named that pid, exited 1, and left the data directory
untouched.

Verified by bijection on the generated C: 18 config() reads, 18 declarations,
no read without a declaration and no declaration without a read. Three bad Int
values are reported in a single run rather than costing one restart each.

ENGRAM_API_KEY keeps its permissive empty default, which disables auth -- that
is pre-existing behaviour and changing it is out of scope. The source marks
making it `required` as the obvious hardening follow-up.
2026-08-16 11:38:28 -05:00
bigmerge b305b49f40 lang: re-stamp the bootstrap compiler so the tree can compile its own source
server.el declares a `program` block, which the previously committed elc cannot
parse. Without this the tree is internally inconsistent: source in the repo that
the compiler in the repo rejects.

This is the documented re-stamp from BOOTSTRAP.md / AGENTS.md, and its
precondition is met -- the self-hosting fixpoint was verified byte-identical
(stage3 output == stage2 output) both before installing and again with the
installed binary. tests/native/test_compiler.el passes 82/82 against it.

Two pre-existing failures are unchanged and are NOT from this work, confirmed
by rebuilding them against the original runtime: test_env's
"state_keys returns JSON array" fails identically before and after, and
test_json/test_state fail to link on symbols (json_build_array, state_has) that
were never prototyped -- the same class of gap as config(), which this branch
fixed because it blocked the build.
2026-08-16 11:38:28 -05:00
bigmerge 8ae163e8e5 lang: give cross-cutting concerns an owner instead of a convention
El's units of encapsulation are the function and the module. Neither can hold
a concern that belongs to the process, so each one had been expressed the only
way it could be -- as a convention: call this at every site. Conventions of
that shape do not hold. Measured here: zero process-identity guards at any
layer, 20 environment variables each with its default written inline at the
read site, 62 persist call sites, 10 per-route auth checks. One absence, four
times.

Step 0 first, because the premise was wrong. El was believed to have no
middleware or effect mechanism. It has one, and it is already load-bearing:
codegen injects engram_boundary_beat at the entry of every @manager/@accessor
fn, decorators take arguments and stack, dharma_emit from a non-@manager fn is
a #error, and the cgi block injects el_cgi_init at the head of main(). So the
correct move was not to invent a mechanism but to generalize the seam that
already existed. The real gap is narrower and is now recorded: the seam is
prologue-only and its callee is a fixed builtin.

Adds a `program` block -- the third program-level declarative block. cgi and
service declare what a program may do; program declares what it is.

  program "engram" {
      singleton: "engram"
      env ENGRAM_BIND: String = ":8742"
      env GUIDE_PORT:  Int    = "8771"
  }

singleton takes an exclusive flock before any user statement runs and refuses a
second start, reporting the holder's pid. It is a lock rather than a pidfile so
the kernel releases it on death including SIGKILL -- no stale state, and so no
"delete the lock file to get unstuck" ritual, which would itself be a
convention. It reports the pid because "already running" is not actionable; a
pid is. That is the direct answer to a stale process surviving a pkill and
going on answering probes.

env entries resolve once at startup -- environment wins, declaration supplies
the fallback -- and validate as a whole, reporting every problem at once rather
than costing one restart per variable. config("X") for an undeclared X is
fatal, because an advisory schema is just another convention. Programs without
a program block are unaffected, so migration is per-program.

Only one keyword is added. `config` and `env` could not become keywords -- both
are real identifiers in the tree -- so the block's fields are read as
identifier token values by its own parse loop and stay usable everywhere else.

The init function is emitted at the block site and called from main() rather
than inlined into main(). The live backend is codegen_streaming, which emits in
source order and cannot hold the entry list alive until main(); this way only a
single bool has to survive.

Also fixes: config() was defined in el_runtime.c but never prototyped in
el_runtime.h, so any el program calling it failed to compile under C99.

Spec: section 18 documents what shipped. Section 9 is corrected -- it claimed
decorators had no structural meaning, which has not been true for some time.
Section 19 designs durability-as-an-epilogue-effect and route authorization
and states plainly why neither is implemented here: both land in files under
concurrent modification, and the prerequisite for both is lifting the seam
from prologue-only to prologue/epilogue.

Self-hosting fixpoint verified byte-identical.
2026-08-16 11:38:28 -05:00
bigmerge 3fcc36c2f1 runtime: transduction is a language concern, so move it into the language
El SDK CI - dev / build-and-test (pull_request) Failing after 14m58s
#141 let signal enter as geometry and it worked, but it was placed at the
CONSUMER and said so in its own commit message. This is the correction.

Three defects, all of them placement:

1. It sat in the engram. Ingest is a LANGUAGE concern — every el program
   touching any modality needs it, and the engram is merely one el program
   that happens to hold a graph. The geometry surface is now defined in
   el_runtime.c immediately ABOVE the engram section and depends on nothing
   inside it. Delete the entire engram and geometry still enters el.

2. It marshalled the vector as a hex STRING, because el had no first-class
   geometry value — which reintroduced text as the TRANSPORT medium one layer
   below the problem being fixed. Geometry is now an el value: a magic-tagged
   heap object carried in el_val_t, same discipline as List/Map. Hex survives
   only as an adapter at the edge, which is all an encoding should ever be.

3. It needed an arbitrary `dim <= 8192` bound purely to size an allocation
   from a caller's CLAIM about a string's length. A value carries its own
   width, so the width is derived and never asserted. The bound is gone, not
   raised — there is nothing left to validate.

Language surface, none of it engram-prefixed: geometry_new / _dim / _is /
_get / _set / _norm / _free, geometry_from_f32le_hex + geometry_to_f32le_hex
as the wire adapters, realizer_register(modality, fn_name), realizer_has, and
transduce(signal, modality) -> Geometry.

REALIZERS ARE DECLARABLE IN EL. This is the part that makes the move real
rather than nominal: registration resolves a name with dlsym against the
running binary, the identical mechanism http_set_handler already relies on,
because every el `fn name(...)` compiles to a global C symbol with that exact
name. So an ordinary el function IS a realizer and a new modality needs no
runtime patch. Verified end to end in lang/examples/transduce.el: an el-defined
tone_realizer is registered by name, transduce dispatches to it, and the
signal demonstrably reaches it (distinct signals produce distinct geometry).

A modality with no realizer transduces to NOTHING. There is deliberately no
built-in realizer, not even for text — silently embedding a description of a
signal and calling that perception is the exact defect this ends.

engram/src/server.el is migrated: POST /api/nodes decodes "emb" hex exactly
once, at the edge, into a Geometry, and everything below that line moves
geometry. The wire is unchanged because production clients speak it. "dim" is
now an ASSERTION about the vector, not the source of its width; disagreement
is a rejected ingest, not a silent reinterpretation.

#141's engram_node_set_emb becomes a DEPRECATED WRAPPER over
geometry_from_f32le_hex + node_attach_geometry — kept only because the runtime
ships as an SDK asset and a downstream binary may link the symbol. Its exact
contract, negative cases included, is preserved and re-verified.

ingest.el's `fn transduce` is renamed transduce_manifold. Mechanically it had
to yield the name (duplicate C symbol, a hard compile error, measured). But it
was never signal->geometry: it chunks already-extracted content into a node+edge
manifold, one layer up, and had taken the name belonging to the primitive
underneath it. Behaviour unchanged.

PROPERTIES FROM #141 PRESERVED, each re-measured on a scratch engram (:8971,
never prod :8742):
  - off-dimension vectors stored but NOT indexed — the HNSW build loop still
    filters on n->emb_dim == dim at four sites, so a 64-dim voice vector is
    durable and addressable without perturbing the 768-dim canonical index
  - geometry makes a node ineligible for embed_backfill: after backfill the
    64-dim voice node was still 64-dim while the text control acquired 768
  - the create response reports whether geometry landed, and the node document
    always emits emb_dim and embedded

Read-back with control and negatives, all verified against a PID-confirmed
fresh binary: geometry node emb_dim=64 embedded=true / emb_set=1; text-only
control emb_dim=0 embedded=false / emb_set=0; malformed hex, ragged length,
and dim-disagreement each emb_set=0.

Two compiler landmines found by reading the generated C rather than trusting a
successful build, both documented at their sites: elc lowers `a == b` to
str_eq unless both operand NAMES are in the per-function int-name set (which
does NOT propagate into nested if-expression blocks — the first cut would have
strcmp'd two integers as pointers on the first geometry-bearing request), and
`+` lowers to string concat when either operand is a user-defined call.
2026-08-16 11:37:27 -05:00
will.anderson a6cef4b983 runtime: publish the vector index instead of guarding it (#143)
El SDK CI - dev / build-and-test (push) Failing after 10m29s
2026-08-16 16:33:28 +00:00
Neuron 8e9d88fc01 runtime: publish the vector index instead of guarding it
El SDK CI - dev / build-and-test (pull_request) Failing after 10m59s
The crash (SIGTRAP in engram_activate -> eg_vindex_sync -> vindex_insert ->
_realloc) had three read paths mutating five process-global statics.
engram_activate, eg_knn_for_node (whose own comment says "No writes.") and
engram_geo_reify_run_json all called eg_vindex_sync, which frees the index,
reallocs the seen-map and inserts — on a read.

Three moves, in decreasing order of how much they dissolve:

1. Misfiled scratch is not shared state. visited/visit_epoch/visited_cap
   were never owned by the index; they are one traversal's local, hoisted
   into struct VIndex as an allocation optimisation. They want neither a
   lock nor a capability nor a pool — just to go back in the call frame.
   Two concurrent READS stomped each other purely because of this.

2. const IS the capability. Once the scratch leaves the struct, search
   reads and nothing else, so vindex_search takes a const VIndex*. That is
   exactly what a capability-pointer ABI would have bought — a read path
   physically cannot call vindex_insert, enforced by the compiler on every
   future caller — for one qualifier instead of an ABI swept across
   hundreds of builtins.

3. What survives is publication, not ownership. HNSW insert is NOT an
   append: it rewires the neighbour links of already-existing elements and
   reallocs elems[], so the store's append-only property does not transfer
   to the index derived from it. eg_vindex_sync therefore splits into
   eg_vindex_maintain (exclusive, sole mutator) and eg_vindex_view (shared,
   returns const VIndex*). A read path may demand that a current snapshot
   exist — a request to the owner, not a mutation by the reader.

Write-side owner: eg_vindex_note_embedded hooks the embedding-ASSIGNMENT
sites rather than the append sites, because a node with no embedding cannot
be in a vector index — embedding assignment is the event that owns index
membership. One O(log n) insert, no O(node_count) presence scan. This also
retires the "STALENESS (honest tradeoff)" note where a lazily-embedded
older node stayed invisible to route_nearest/autoconnect until a full
rebuild (the embed-gap #20 shape).

Evidence. The existing harness conflated two hazards, which is why fixing
half of it read as failure. Split into four:

  single (3000 vec, ASan+UBSan)          clean  ->  clean
  readers (4 readers, no writer, TSan)   RACE   ->  clean
  unsynchronized (writer+reader, bare)   race   ->  race, expected forever
  published (owner + 4 readers)          n/a    ->  clean, 3000/3000 landed

RESULT: PASS. recall@10 = 0.9365 at ef_search=128 (gate >= 0.90);
determinism byte-identical across two independent builds.

The unsynchronized half is now permanently expected to race, deliberately:
it is the executable proof that the boundary must live above the data
structure, not inside it.

fb32d15's guard is KEPT, correcting this design's own section 5. Measured,
it guards TWO structures and only one was converted here: g->nodes/g->edges
are realloc'd in place (el_runtime.c:7618,7629) and engram_activate_inner's
embed-backfill writes n->emb through exactly such a borrowed pointer.
Deleting the guard reintroduces a measured 11171->9579 edge loss. Its
comment is narrowed to the RAM graph and the deletion precondition named.

That corrects the ordering claim too: the residual is not one ABI that
dissolves everything at once, it is a PROPERTY applied per structure.
Residues evaporate in the order the property is applied, and a residue
whose structure has not been converted must be left standing.
2026-08-16 11:29:17 -05:00
bigmerge e99a4640e2 test: regression harness for the vindex concurrency crash
Promotes the two throwaway sanitizer harnesses used to diagnose the
2026-08-16 soul crash into engram/test/ so the bug cannot silently regress.

The harness has two halves and the PAIR is the point — it is what localises
the defect to concurrency rather than to HNSW logic:

  single      3000 clustered vectors, one thread, ASan+UBSan. The CONTROL.
              Must always be clean. During diagnosis this cleared all 13,820
              real dim-768 vectors from the live store, which DISPROVED an
              inspection-derived hypothesis about an out-of-bounds
              reverse-link write at engram_vindex.c:340.

  concurrent  writer + reader on one shared index, TSan. Currently reports a
              race at engram_vindex.c:195 (visited_reset) reached from both
              vindex_search and vindex_insert, because VIndex still owns its
              visited[]/visit_epoch scratch — so even two concurrent READS
              corrupt each other's traversal.

Verified: half 1 passes, half 2 reproduces the race.

Gated on EXPECT_RACE, default 1, so the concurrent half documents the known
defect without failing the suite today. When the visited set moves to a
per-query checkout pool (hnswlib VisitedListPool style — NOT thread_local,
since http_worker is a thread per connection and a __thread buffer would leak
~55KB per connection), flip EXPECT_RACE=0 and it becomes a real gate.
2026-08-16 11:29:17 -05:00
bigmerge bdc1f99fb9 runtime: guard engram activation against the unsynchronized awareness thread
The soul daemon had two engram callers and only one of them locked.
soul.el:729 starts the HTTP server via http_serve_async (spawning
http_worker threads); soul.el:731 then runs awareness_run() on the MAIN
thread. awareness.el's perceive() -> engram_activate_json() ->
engram_activate() -> eg_vindex_sync() -> vindex_insert() mutates the same
g->nodes/g->edges and the process-global _eg_vindex HNSW index that the
workers touch. g_engram_req_lock existed to serialize exactly this, but it
was only ever taken inside http_worker: engram_req_lock/engram_req_unlock
appear in ZERO .el sources, so the awareness loop ran lock-free beside the
workers on every tick (SOUL_TICK_MS=1000).

Result was a crash-loop under launchd KeepAlive: five crashes in ~4 minutes
on 2026-08-16 with varying faulting frames -- search_layer<-vindex_insert
<-eg_vindex_sync, engram_activate, abort, and one inside xzm_realloc's own
freelist. Varying sites plus a fault in allocator metadata means heap
corruption. The SIGSEGV address 0x65646f4e6d617267 is little-endian ASCII
"gramNode": string bytes dereferenced as an Elem vector pointer.

Diagnosed by bisection rather than inspection:
  - Replaying all 13,820 real dim-768 vectors harvested from the live store
    through the index single-threaded under ASan is 100% clean, which rules
    out an HNSW logic/bounds bug.
  - Two threads on one index trip ThreadSanitizer immediately at
    engram_vindex.c:195 (visited_reset), reached from both vindex_search and
    vindex_insert. VIndex keeps a SHARED visited-epoch scratch buffer, so
    even two concurrent READS corrupt each other's traversal and walk bogus
    element indices.
So this is purely a concurrency defect, not an HNSW logic error. (An
inspection-derived hypothesis about an out-of-bounds reverse-link write at
engram_vindex.c:340 was disproved by the single-threaded run.)

Fix: a thread-local ownership depth (_eg_req_depth) lets engram entry points
self-guard. engram_activate() becomes a wrapper over engram_activate_inner()
that acquires g_engram_req_lock when called with depth 0 (the awareness
thread) and passes through when depth > 0 (nested inside an http_worker that
already holds it), so the non-recursive mutex cannot self-deadlock. The depth
is a plain counter, never a recursive-mutex count, preserving
engram_self_reify_beat_json's contract of genuinely releasing the lock
mid-beat.
2026-08-16 11:29:17 -05:00
will.anderson 44b621e551 runtime: anchor the think read, so Neuron can think at all (#142)
El SDK CI - dev / build-and-test (push) Failing after 13m0s
2026-08-16 16:26:00 +00:00
Neuron ded6ca546f runtime: anchor the think read, so Neuron can think at all
El SDK CI - dev / build-and-test (pull_request) Failing after 13m24s
engram_think_json passed NULL as the anchor. NULL is not "no opinion":
engram_think re-origins at `anchor ? anchor : region->centroid`, so NULL
means "read from the centroid" — and the centroid is the one point where
the gradient is zero by construction. r = x - centroid = 0, so every axis
projection is 0, grad is 0, and direction takes the "at rest" branch at
engram_cognition.c:137.

Measured consequence: EVERY faculty returned an identical null result,
differing only in its label —
  {"direction":[0,0,0,0,0,0,0,0],"spread":0,"magnitude":1,"confidence":0.5}
magnitude 1 is membership evaluated at the centroid, spread 0 is its
distance to itself, confidence 0.5 is the stance fallback. The geometry was
never at fault: /api/drift computes real values (centroid_sep 0.104,
core_disp 0.045) over the very same 87 members. Neuron could not think
because the read was always taken from the region's own centre.

The seeds choose WHICH region; they must also supply the VANTAGE. Anchor at
the first resolvable embedded seed — the same seed eg_geo_build_desc infers
dim from, so the two can never disagree. One seed still yields a real
gradient because the descriptor expands to that seed's neighbourhood, so
the seed's position is distinct from the neighbourhood centroid.

The vector is COPIED, never borrowed: g->nodes is realloc'd in place on
append, so a borrowed EngramNode* dangles across any concurrent write.

Verified against a clone of the production store (13,616 nodes / 37,865
edges):
  self anchor   n_support 87  magnitude 0.00282  spread 18.79
  values hub    n_support 28  magnitude 0.00318  spread 17.72
with distinct unit direction vectors. Previously both returned the zero
vector with magnitude 1 and spread 0.

STILL OPEN, now isolated by this fix: all five faculties return identical
numbers and confidence stays 0.5, because cog_stance_init is passed NULL
for the stance and the faculty enters the computation only through the
stance's axis_gain[] and bias_dir. The faculty label is inert until a
stance is loaded — which is what learn()'s correspondence-beat calibrates.
Same shape as this bug: a neutral parameter collapsing a capability to a
constant.
2026-08-16 11:25:24 -05:00
will.anderson b5b96c05ed runtime: let signal enter as geometry, not as prose about signal (#141)
El SDK CI - dev / build-and-test (push) Failing after 10m36s
2026-08-16 16:13:28 +00:00
Neuron c79033b749 runtime: let signal enter as geometry, not as prose about signal
El SDK CI - dev / build-and-test (pull_request) Failing after 10m55s
No ingest path could carry a vector. engram_node/_full/_layered take text
only, and a node acquired an embedding solely via engram_embed_backfill
DERIVING one from n->content. That made text the mandatory entry medium:
any non-text modality had to be described in prose first, so the geometry
we then reasoned over was the geometry OF THE DESCRIPTION, not of the
signal. Measured: POST /api/nodes accepted an "emb" field, returned 200
with a fresh id, and stored nothing — emb_dim=None, embedded=false.

engram_node_set_emb attaches a vector to an existing node. Off-dimension
vectors are stored but not indexed (the HNSW build loop already filters on
emb_dim), so modality geometry is durable and addressable without
perturbing the canonical index. Setting emb also makes the node ineligible
for embed_backfill, so a realizer's vector is never overwritten by a
text-derived one.

Two reporting fixes ride along, because both are how the drop stayed
invisible: the create response now reports emb_set instead of being
success-shaped regardless, and the node document now always emits emb_dim
and embedded — without which a genuine ingest drop and a mere reporting
gap are indistinguishable.

Verified live: voice node emb_dim=64 embedded=true; text control emb_dim=0
embedded=false; malformed hex, length mismatch and dim<=0 all reject.

KNOWN PLACEMENT DEFECT: this is at the consumer. Ingest is a language
concern, not an engram feature — every el program touching any modality
needs it. The vector also marshals as a hex STRING because el has no
first-class geometry value, which reintroduces text as the transport
medium one layer below the problem being fixed. The durable shape is
geometry as an el value plus declarable realizers, after which the engram
stops having an ingest concept at all. Landing this as the verified probe
that proves the path.
2026-08-16 11:12:48 -05:00
will.anderson 1119295238 Merge pull request 'runtime: state_get leaked its value on every call' (#140) from fix/state-get-leak into dev
El SDK CI - dev / build-and-test (push) Failing after 14m5s
2026-08-16 13:09:58 +00:00
bigmerge 9c07970943 runtime: state_get leaked its value on every call
El SDK CI - dev / build-and-test (pull_request) Failing after 14m26s
char* result = el_strdup_persist(e ? e->value : "");   // never freed
    pthread_mutex_unlock(&_state_mu);
    char* copy = el_strdup(result);                        // arena-tracked
    return el_wrap_str(copy);

Two copies were made. `result` existed only as the source for `copy` — never
returned, never freed — and el_strdup_persist bypasses the arena BY DESIGN
("state_set, engram internals"), so arena-pop could never reclaim it. Every
state_get leaked its full value string, permanently.

MEASURED: 200,000 state_get calls against a 64-byte value.
    before   15 MB peak RSS growth   (~75 bytes/call — the value plus overhead)
    after     0 MB

IMPACT. The soul's awareness loop has 68 state_get call sites and ticks every
200ms. Live measurement before the fix: RSS climbing 112 MB per 20s, about
19 GB/hour, in awareness_run -> one_cycle -> perceive, while node_count stayed
flat at ~13,479 — growth with no data behind it. It drove the host from 20 GB
free to 4.3 GB in roughly an hour.

WHY NOW, since the code is old: the soul used to restart constantly (no
write-through, divergent graph, 2.11 GB). Stabilising it (neuron #162) let it
stay up long enough to accumulate. The fix did not cause this leak; it removed
the crashes that were hiding it. Same pattern as the test framework surfacing
math_log — the defect was always there, something finally made it visible.

Found by Ishikawa rather than by reading the nearest code: method (arena
push/pop IS correctly paired per tick), material (node count flat, so not data
growth), environment (19 GB/hr / 18,000 ticks = ~1.1 MB per tick, so per-tick
not one-shot), machine (an allocator that bypasses the arena) — which is where
the evidence pointed.

el_strdup tracks into the thread-local arena, which touches no shared state, so
taking the single copy under _state_mu is safe and removes the temporary
entirely.

Verified: self-hosting fixpoint byte-identical; state round-trip correct for
hit, miss, and overwrite.
2026-08-16 08:09:32 -05:00
will.anderson 0832865952 Merge pull request 'test framework phase 3/4: black_box barrier + three-signal complexity gate, armed' (#139) from wt/soul-runtime-reconcile into dev
El SDK CI - dev / build-and-test (push) Failing after 11m32s
2026-08-16 03:02:18 +00:00
Neuron e0b2c0ea54 bench: arm the Phase 4 gate -- proven to pass clean AND fire on a quadratic
El SDK CI - dev / build-and-test (pull_request) Failing after 12m7s
Adds tests/native/test_lexer_scaling.el, the regression gate for el #132.

Both directions are proven on LIVE workloads, not synthetic series:
  healthy per-character scan  1821 3251 6007 10422 us -> O(n)   PASS
  rescan-from-zero (the #132 shape)  922 3667 13524 44792 -> O(n^2) FAIL

A gate only proven to pass is decoration. The quadratic specimen exists so
the gate is proven to FIRE.

Also fixes elb_spread_ok to judge the ASYMPTOTIC TAIL (last three ratios)
rather than the whole sweep. Measured on a genuinely linear scan the ratios
ran 3.37 2.92 1.76 1.65 -- the head looks quadratic because it is cold
cache, the tail is the truth. Whole-sweep spread rejected correct data. A
complexity bound is an asymptotic claim and must be judged asymptotically.

That fix came from the classifier refusing to rubber-stamp my own bad
measurement: it reported INDETERMINATE on an unwarmed sweep rather than
passing it. Warmup is now taken and discarded at every sweep point.

Reverts the == workarounds in test_elbench.el now that el #137 has landed;
the natural form generates no str_eq and all 13 fitter tests stay green.
The  workaround remains -- the Plus arm is still open.
2026-08-15 21:58:46 -05:00
bigmerge cf060adbfd Merge remote-tracking branch 'origin/dev' into wt/soul-runtime-reconcile 2026-08-15 21:55:40 -05:00
will.anderson 63fe8a766d Merge pull request 'codegen: Bool is int-like, so Bool comparisons stop lowering to str_eq' (#138) from fix/bool-is-int-like into dev
El SDK CI - dev / build-and-test (push) Failing after 4m28s
2026-08-16 02:54:34 +00:00
bigmerge b5a0a729e6 codegen: Bool is int-like, so Bool comparisons stop lowering to str_eq
El SDK CI - dev / build-and-test (pull_request) Failing after 14m49s
fn check(label: String, cond: Bool, want: Bool) -> Void {
        if cond == want { ... }        ->  if (str_eq(cond, want))   SIGSEGV
    }

Bool has always been an integer in the value model — type_to_c maps Bool to
"int", and el_runtime.h states "Bool -> el_val_t (0 = false, nonzero = true)".
But Bool names were registered NOWHERE: build_int_names_for_params tracked Int
and Float params, and the `let` path tracked Int and Float bindings. Neither
knew about Bool.

So comparing two Bools fell through to str_eq, which dereferenced 0 or 1 as a
char* and segfaulted immediately.

This is the third instance of one family found tonight, after el #137 (a call
on either side of == poisoned the operator) and el #136 (a missing import
compiled clean). All three are the same shape: something the compiler could not
type, silently handled as a string.

Found while writing #137's own test harness — the first version of that harness
crashed on exactly this, on both the old and new compiler, which is how it
surfaced. A test harness that cannot compare two Bools is a good way to notice.

VERIFIED:
  - the harness that segfaulted on every prior compiler (exit 139, no output)
    now runs clean: 14 passed, 0 failed
  - self-hosting fixpoint byte-identical
  - the compiler's own generated C differs by 8 lines — only the intended
    registration
  - neuron's full soul amalgam regenerates in 424ms, exit 0, BYTE-IDENTICAL
  - test_math 13/13, test_string 27/27, test_core 10/10, test_text 12/12

Adds tests/runtime/operator_typing_test.el, the 15-case suite from #137, so
this family is covered going forward rather than rediscovered.
2026-08-15 21:54:10 -05:00
will.anderson b26dd47aef Merge pull request 'codegen: either side Int is enough for == and !=, not both' (#137) from fix/eq-operand-inference into dev
El SDK CI - dev / build-and-test (push) Failing after 12m0s
2026-08-16 02:52:02 +00:00
bigmerge b55e6bfd53 codegen: either side Int is enough for == and !=, not both
El SDK CI - dev / build-and-test (pull_request) Failing after 12m20s
let a: Int = 5
    getint(5) == a      ->  str_eq(getint(5), a)      SIGSEGV
    getint(5) == 5      ->  getint(5) == 5            fine

A function call whose return type codegen cannot infer poisoned the operator,
and a declared Int on the other side did not save it. str_eq then read an
integer as a char* and segfaulted. Only an integer LITERAL on one side forced
the numeric form, which is why the bug stayed invisible: the common case
happened to be safe.

The check required BOTH operands to be provably Int:

    if is_int_expr(left) { if is_int_expr(right) { numeric } }

Loosening to OR is strictly safer, not a trade:
  - when one side is a known Int, str_eq is ALWAYS wrong — it dereferences
    that integer — while numeric comparison is at worst a wrong answer on a
    program that was already ill-typed;
  - when neither side is Int nothing changes at all, so string comparison is
    untouched.

Found by the test-framework agent while building the benchmark harness; it
correctly declined to fix it mid-phase since it is a codegen semantics change.

VERIFIED, because a semantics change earns more than an assertion:
  - 15/15 on a dedicated operator suite covering string literals, string vars,
    string-returning calls, mixed var/call, and != in every combination. The
    pre-change compiler scores 0/15 on the same file: it segfaults before
    printing anything.
  - self-hosting fixpoint byte-identical
  - the ONLY difference in the compiler's own generated C is the intended one:
    a nested if becoming two sequential ifs, in EqEq and NotEq. Nothing else
    moved.
  - neuron's full soul amalgam regenerates in 400ms, exit 0, output
    BYTE-IDENTICAL at 1,270,212 bytes
  - test_math 13/13, test_string 27/27, test_core 10/10, test_text 12/12 —
    62 tests, 190 assertions, zero failures

NOT fixed here, same family, flagged for a decision: Bool PARAMETERS are not
tracked as int-like, so `cond == want` between two Bool params still lowers to
str_eq and segfaults. Found while writing this commit's own test harness — the
first version of it crashed on exactly that, on both the old and new compiler.
It needs the same treatment, and it wants its own change.
2026-08-15 21:51:35 -05:00
will.anderson dbb06f6ee4 Merge pull request 'compiler: a missing import is an error, not an empty string' (#136) from fix/missing-import-is-an-error into dev
El SDK CI - dev / build-and-test (push) Failing after 11m0s
2026-08-16 02:48:01 +00:00
bigmerge 906c664a65 compiler: a missing import is an error, not an empty string
El SDK CI - dev / build-and-test (pull_request) Failing after 11m23s
import "../../NOPE/does_not_exist.el"

compiled CLEANLY — exit 0, empty stderr, and a program silently missing
everything it imported.

resolve_imports did `fs_read(src_path)` and used the result without checking.
fs_read returns "" both for "file is empty" and "file does not exist", so a
typo, a moved file, or a relative path resolved from the wrong working
directory all produced a successful build of nothing.

It caused a real wrong conclusion during test-framework work: a bisection run
from a subdirectory where ../../runtime/ did not resolve produced ELEVEN
consecutive "successful" compiles that had included no runtime at all, and the
results were believed before anyone noticed.

Missing dependency, confident success — the same shape as a test suite
reporting pass for tests that never ran, and as a benchmark reporting 0us
because the optimiser deleted the loop.

fs_exists separates the two cases, so a legitimately empty file still resolves
to "" and is fine. A path that does not exist now prints the resolved path and
exits 1, which is what build scripts check.

Verified:
  - bad import: exit 1 (was 0), message names the resolved path
  - elc-cli.el still compiles, self-hosting fixpoint byte-identical
  - neuron's full soul amalgam regeneration: exit 0, 405ms, output
    byte-identical at 1,270,212 bytes
2026-08-15 21:47:32 -05:00
Neuron 6a6b589ba0 bench: real black_box barrier + three-signal growth-curve gate
Adds el_black_box (inline asm, +r constraint, memory clobber) and
runtime/elbench.el: a growth-curve classifier that gates time AND
allocation-count AND allocation-bytes, failing if any exceeds its
declared curve.

Refusal is a first-class verdict. The classifier REFUSES rather than
classifying when the largest measurement is below the floor, or when a
series is hard-flat across an 8x input range -- the shape produced when
the optimiser deletes the work. Reporting O(1) there would be a
confident answer with nothing behind it. Disagreeing ratios report
INDETERMINATE rather than a guess.

Deviation from DESIGN.md 6.2, stated in the source: uses consecutive
ratios on a mandated geometric sweep rather than least-squares over
candidate curves. Ratios are directly interpretable on a doubling sweep
and need no floating point; the cost is weaker O(n) vs O(n log n)
separation, reported as an ambiguous band rather than guessed.

Documents the counter scope limit: engram_*.c and libcurl malloc are
NOT tracked, so a flat curve over engram/HTTP-dominated work is not
evidence of anything.

13 tests prove the classifier against real measured series from
fitprobe.el -- including that an accumulator's allocation COUNT is
linear while its bytes are quadratic, and that el #132's pure-CPU shape
reads FLAT on both allocation signals and is caught only by time.
2026-08-15 21:45:13 -05:00
bigmerge b5d1e53902 Merge remote-tracking branch 'origin/dev' into wt/soul-runtime-reconcile 2026-08-15 21:38:11 -05:00
will.anderson 9e96d74f6a Merge pull request 'runtime: count container allocations too, not just strings' (#135) from feat/alloc-accounting-containers into dev
El SDK CI - dev / build-and-test (push) Failing after 11m49s
2026-08-16 02:37:14 +00:00
bigmerge a8908908df runtime: count container allocations too, not just strings
El SDK CI - dev / build-and-test (pull_request) Failing after 12m11s
el #131 instrumented the four string allocators, which meant list- and map-heavy
code reported ZERO allocations — a benchmark over lists would have been fitted
against a flat line and passed anything. Caught during framework work: a
"linear" specimen read 0 allocs until it was rewritten to allocate strings.

A gate is only as good as its blind spots are small, and a signal that silently
reads zero is worse than no signal: it produces a confident pass.

Now counted at every container allocation — ElList and ElMap bodies, their
backing arrays, the copy-on-write clones, and the realloc growth path.

Verified on an append loop (n = 100..800):
    allocs  7, 8, 9, 10          +1 per doubling = O(log n) reallocations
    bytes   2048, 4096, 8192, 16384   exactly 2x per doubling = O(n)

Both curves are what correct amortized growth should look like, and both read
zero before this change.

Known remaining scope, stated rather than left implicit: these counters cover
the runtime's own allocations. They do not see malloc inside engram_*.c or
libcurl, which is correct — the gate is for El-level complexity, not for
third-party memory behaviour.
2026-08-15 21:36:52 -05:00
Neuron 6291a35bb9 design: gate on THREE signals -- the alloc gate would have missed el #132
el #132's quadratic (strlen per character in str_char_code/str_slice) is
pure CPU and allocates NOTHING. Measured on three controlled specimens:

  specimen  allocs        bytes         time
  linear    2.00 -> O(n)  2.16 -> O(n)  2.05 -> O(n)
  accum     2.00 -> O(n)  3.99 -> O(n2) noisy
  compute   FLAT          FLAT          3.96 -> O(n2)

'compute' is #132's shape. A gate fitting only allocation count and bytes
classifies it FLAT and passes -- it would not have caught the defect it
was created for. The gate now fits time AND count AND bytes, failing if
any exceeds its declared curve.

Also: black_box is mandatory and consuming the result is NOT sufficient.
The first 'compute' reported 0us at every n while returning a correct n2 --
clang closed the loop to a multiply. Only an opaque call restored the curve.

Adds lang/tests/bench/fitprobe.el as the fitter's known-good/known-bad set,
so the classifier is provable without depending on a real bug existing.
Marks DESIGN.md 1.3 stale: test_compiler 3.58s -> 0.03s (119x).
2026-08-15 21:34:39 -05:00
will.anderson a69a4a5894 Merge pull request 'runtime: math_log is base-10, not natural log' (#134) from fix/math-log-base10 into dev
El SDK CI - dev / build-and-test (push) Failing after 14m48s
2026-08-16 02:34:09 +00:00
bigmerge edafd8cce8 runtime: math_log is base-10, not natural log
El SDK CI - dev / build-and-test (pull_request) Failing after 10m8s
el_val_t math_log(el_val_t f) { return el_from_float(log(el_to_float(f))); }
    el_val_t math_ln(el_val_t f)  { return el_from_float(log(el_to_float(f))); }

Both were natural log, so math_log and math_ln were the same function.
log10(100) returned 4.605 instead of 2.

Three sources already agreed it should be base-10 and were being contradicted
by this one line:
  - runtime/math.el:55  "// math_log — base-10 logarithm."
  - el_seed.c:1278      __log_f -> log10()  (the path math.el actually calls)
  - tests/native/test_math.el:133  asserts log10(100) == 2

FOUND BY THE NEW TEST FRAMEWORK ON ITS FIRST RUN (el #133). The assertion had
been sitting in the suite the whole time; nothing could report it. The old
harness printed "N passed, M failed" with no per-test detail, and half the
suites were not compiling at all — so a failing assertion in a suite nobody
could run was indistinguishable from no failure.

That is the entire argument for the framework, demonstrated on day one: this is
not a bug the framework introduced, it is a bug the framework made VISIBLE.

Verified: tests/native/test_math.el goes 12/13 -> 13/13, math-log passing.
2026-08-15 21:33:20 -05:00
will.anderson 5e3e69d326 Merge pull request 'test framework phase 1: compile-time registry + El-side runner with per-test timing' (#133) from wt/soul-runtime-reconcile into dev
El SDK CI - dev / build-and-test (push) Failing after 10m17s
2026-08-16 02:31:08 +00:00
bigmerge 4c3414072b Merge remote-tracking branch 'origin/dev' into wt/soul-runtime-reconcile 2026-08-15 21:30:52 -05:00
will.anderson 0288024396 Merge pull request 'compiler: fix the quadratic — strlen() on every character access' (#132) from fix/compiler-quadratic-strlen into dev
El SDK CI - dev / build-and-test (push) Failing after 4m2s
2026-08-16 02:29:02 +00:00
Neuron 3e7ab07e82 test framework phase 1: forward decls, void-return fix, suite migration
El SDK CI - dev / build-and-test (pull_request) Failing after 10m4s
Completes the Phase 1 runner and migrates the 11 test files onto it.

- forward-declare the registry accessors in the test preamble; they are
  defined at the end of the unit but the El runner is compiled in between
- eltest.el: explicit trailing return in the void emit_* helpers, which
  otherwise lower to 'return println(...)' and fail to compile
- test files import runtime/eltest.el explicitly, using the language's own
  textual import mechanism rather than compiler-side auto-injection
- DESIGN.md 6.5: gate on allocation COUNT AND BYTES, not count alone

Verified: self-hosting fixpoint byte-identical (gen2 == gen3). 6 of 11
suites run and report per-test timing. The other 5 fail to COMPILE, and
fail identically under the committed compiler -- pre-existing breakage
this framework makes visible for the first time.
2026-08-15 21:28:30 -05:00
bigmerge d231b7e5e7 compiler: fix the quadratic — strlen() on every character access
El SDK CI - dev / build-and-test (pull_request) Failing after 10m21s
THE BUG. str_char_code() and str_slice() each called strlen() on every
invocation. The lexer walks source one character at a time, so every character
access rescanned the whole remaining input: O(n) per character over n
characters = O(n^2).

    el_val_t str_char_code(el_val_t s, el_val_t i) {
        ...
        int64_t n = (int64_t)strlen(str);   // <- O(n), every call
        if (idx < 0 || idx >= n) return 0;
        return str[idx];
    }

HOW IT WAS FOUND. Not by reading code — by sampling the running process, which
is the same method that resolved tonight's engram outage after four wrong
theories. A geometric sweep of synthetic sources showed wall-clock rising 3.0x,
3.0x, 4.0x, 4.14x per doubling (converging on 4x = quadratic), and a stack
sample put 779 of 779 samples inside lex(), every one bottoming out in
_platform_strlen via str_char_code and str_slice.

THE FIX. Remember the length instead of recomputing it. The subtlety is
INVALIDATION: El strings are arena-allocated, so a freed pointer can be reused
for a different string at the same address, and a naive pointer-keyed cache
would hand back a stale length and read past the end of the new string —
trading a performance bug for a memory-safety one. So entries carry a
generation, a hit requires pointer AND generation to match, and every path that
frees or mutates a runtime string bumps the generation: el_arena_pop,
seed_request_end, __str_set_char. Stale entries cannot be believed; they miss
and recompute.

MEASURED, same host, same inputs:

    n(fns)    before     after
      512      0.10s     0.01s
     1024      0.37s     0.02s
     2048      1.51s     0.03s     50x

    the compiler's own 422 KB source concatenated (DESIGN.md's 3.58s case):
              3.55s ->  0.03s      118x

The speedup GROWS with input size, which is the signature of removing a
complexity class rather than a constant factor. After the fix each doubling
adds ~0.01s: linear.

CORRECTNESS, verified rather than assumed:
  - byte-identical output on every sweep input (n = 128..2048)
  - byte-identical output on the 422 KB compiler concatenation
  - byte-identical output on tests/runtime/string_test.el
  - self-hosting fixpoint byte-identical
  - new tests/runtime/str_cache_test.el: 17 assertions covering bounds, empty
    strings, negative indices, slice clamping, distinct strings not sharing a
    cached length, 1000 interleaved strings forcing cache-slot collisions, and
    a grown string not reporting its old length. All pass.

This is the defect that made dist/soul.c a committed artifact: elc could not run
in CI because it needed 24 GB+ and minutes. It needs neither now.
2026-08-15 21:28:23 -05:00
bigmerge a668062e38 Merge remote-tracking branch 'origin/dev' into wt/soul-runtime-reconcile 2026-08-15 21:24:03 -05:00
Neuron 24fac765a6 test framework phase 1: compile-time registry + El-side runner
Replace the hardcoded test harness main() with a generated static registry
and index-based accessors, and move all reporting into runtime/eltest.el.

The old harness inlined direct calls into main() and counted assertions in
two globals. That shape cannot report which test failed, how long any test
took, or whether a test ran at all -- a misspelled registration reported
success for a test that never executed.

- assertions record into per-test state instead of global counters
- registry table emitted at compile time; discovery strictly precedes
  execution, which is what later enables --list, filtering and sharding
- per-test wall timing on CLOCK_MONOTONIC, taken in C around the call
- runner in El: structured NDJSON events as source of truth, human output
  rendered from the same fields
2026-08-15 21:24:03 -05:00
will.anderson cb1f2a74af Merge pull request 'runtime: allocation accounting — deterministic signal for complexity gating' (#131) from feat/alloc-accounting into dev
El SDK CI - dev / build-and-test (push) Failing after 3m49s
2026-08-16 02:22:13 +00:00
bigmerge 37bcf7eb74 runtime: allocation accounting — the deterministic signal for complexity gating
El SDK CI - dev / build-and-test (pull_request) Failing after 12m7s
Implements the three primitives the test-framework design (DESIGN.md §6.5)
requires for gating on growth curves: el_alloc_count, el_alloc_bytes,
el_peak_rss. Registered in codegen's builtin_arity and wrapped in el_seed.c per
the project's C-builtin recipe.

WHY COUNTS AND NOT WALL-CLOCK: a growth-curve gate has to be a hard build
failure, which means the signal cannot flake. Wall-clock needs warmup,
statistics, and a quiet machine; on shared CI it is unusable as a gate.
Allocation counts are perfectly deterministic — same input, same number, every
machine, every run. Fit them against n and a complexity regression becomes a
build failure with zero noise.

All four runtime string allocators (el_strdup, el_strbuf, and their _persist
variants) funnel every allocation the language performs, so instrumenting there
counts everything.

WHY BYTES AS WELL AS COUNT — this is not redundancy, it is the whole gate.
Measured with two El programs, one allocating once per item, one rebuilding its
accumulator each iteration:

    n     linear allocs / bytes      quadratic allocs / bytes
    100        100 /    290               100 /   5,150
    200        200 /    690               200 /  20,300
    400        400 /  1,490               400 /  80,600
    800        800 /  3,090               800 / 321,200

The quadratic program's allocation COUNT is exactly linear — identical to the
healthy one. Counting allocations alone would have missed it completely. Bytes
catch it: each doubling of n quadruples bytes (ratios 3.94, 3.97, 3.99 ->
converging on 4.0, i.e. O(n^2)), while the linear case converges on 2.0.

That shape — count linear, per-allocation size growing — is the classic
accidental quadratic, and it is exactly elc's defect: quadratic allocation
VOLUME, which the old shipped compiler paid in RSS (27 GB, OOM) and the rebuilt
one pays in malloc/free churn (42s on 1.4 MB). Volume was the invariant across
both; RSS and wall-clock were just the two ways it surfaced.

el_peak_rss is exported for context and is explicitly NOT a gating signal — it
is perturbed by allocator internals, the page cache, and the OS. Gate on the
deterministic numbers; report the physical one.

Counters are unsynchronised by design: this is measurement, and a lock would
change the thing being measured. Exact on the single-threaded compile path,
approximate under threads.
2026-08-15 21:21:42 -05:00
will.anderson 2240d26c32 Merge pull request 'store: judge memory pressure by swap RATE, not level' (#130) from fix/elc-rebuildable-compiler-builtins into dev
El SDK CI - dev / build-and-test (push) Failing after 11m44s
2026-08-16 02:12:22 +00:00
bigmerge 19cc99e57d store: judge memory pressure by swap RATE, not swap level
El SDK CI - dev / build-and-test (pull_request) Failing after 11m59s
The guard I added minutes ago checked swap availability as a level
(avail < total/8 -> report zero available). That is the wrong signal, and the
same host proved it twice within minutes:

    47.65 / 48.00 GiB swap used, 2047 swapouts/s  -> genuinely thrashing
    26.67 / 28.00 GiB swap used,    0 swapouts/s  -> healthy, 15.6 GiB free

Both are ~97% "used". macOS grows swap files on demand and trims them lazily,
so the level says almost nothing about now — it is a high-water mark. The level
check calls the second state an emergency and starves the pool for no reason,
which is its own failure mode: a guard that fires on healthy machines gets
disabled, and then guards nothing.

What separates the two is whether pages are moving. So sample the swapout
counter across calls and judge the delta:

  - > 200 pages/s (~3 MiB/s) sustained outward paging => report zero available;
    callers refuse to grow and pc_relieve_pressure hands frames back.
  - The first call primes the baseline and reports no pressure. One sample
    cannot have a rate, and inferring one from a single reading is exactly the
    mistake this commit removes.

Measured thresholds, not guessed: idle sat at 0/s, recovery burst hit 24,845/s
while the compressor drained (transient, correctly not a growth decision since
growth is only evaluated on eviction passes), and real thrash held ~2000/s.
200/s sits clearly above noise and far below either.

The compressor-footprint subtraction stays: that RAM is genuinely spoken for
regardless of paging rate.
2026-08-15 21:12:00 -05:00
will.anderson f39ae40047 Merge pull request 'store: bound the pool by available memory and let it shrink' (#129) from fix/elc-rebuildable-compiler-builtins into dev
El SDK CI - dev / build-and-test (push) Failing after 4m43s
2026-08-16 02:02:24 +00:00
bigmerge e52415f0e0 store: bound the pool by AVAILABLE memory and let it shrink
El SDK CI - dev / build-and-test (pull_request) Failing after 11m47s
The adaptive budget I added an hour ago could only grow, and grew toward a
share of TOTAL ram (80%, ~38 GiB on a 48 GB host). That is a memory leak with
extra steps: total never shrinks when other processes need memory, so the pool
had no way to notice it was starving the machine it runs on. Deployed briefly;
caught as memory pressure on the host.

A control loop with only one direction is not a control loop.

  - pc_available_ram(): free + inactive + purgeable via host_statistics64 on
    Darwin, MemAvailable on Linux. Availability is the quantity that moves when
    the machine is under pressure; total is not. Returns 0 when it cannot be
    read, and callers then refuse to grow — a cache is never worth swapping the
    host, so unknown means no.

  - Growth is bounded by availability minus a free-memory floor (2 GiB default,
    ENGRAM_POOL_FREE_FLOOR_MB), not by total. The share-of-total ceiling stays
    as a second bound and drops 80% -> 50%.

  - pc_relieve_pressure(): the missing direction. On every eviction pass, if
    available memory is under the floor, hand back ~25% of held frames; the
    resident set follows on the next pass so the memory is actually returned
    rather than merely re-labelled. Counted as adapt_shrinks alongside
    adapt_grows so both directions are visible in the same report.

  - pc_default_cap() also clamps the STARTING budget to what is spare right
    now, so a cold boot on a loaded machine does not open at a size the host
    cannot afford.

Verified on a 48 GB host: engram boots in ~30s, RSS settles at 2.22 GiB (the
store's actual size, resident, not creeping), 0.0% CPU, 13,439 nodes / 37,670
edges, embeddings complete. Guard reports 9.71 GiB available against a 2.00 GiB
floor — 7.71 GiB of headroom it is permitted to use and no more.
2026-08-15 21:02:05 -05:00
will.anderson 7a479111ac Merge pull request 'store: extend the write barrier to edges — kills the full-store walk' (#128) from fix/elc-rebuildable-compiler-builtins into dev
El SDK CI - dev / build-and-test (push) Failing after 14m27s
2026-08-16 01:44:33 +00:00
bigmerge e917b3d439 store: make the buffer pool sense its own state and correct from it
El SDK CI - dev / build-and-test (pull_request) Failing after 14m35s
Follow-on to the edge write barrier. That fix removed the full-store walk;
this one makes the pool able to notice if anything like it happens again.

WHAT WENT WRONG, precisely: the pool thrashed the live engram to a standstill
twice on 2026-08-15 and said nothing. From outside it was indistinguishable
from "busy loading" — 100% CPU, flat RSS, no output — so four wrong theories
got tried (bad binary, corrupt snapshot, WAL replay, feature flags), each
costing a deploy or a rollback. The whole time, hits/misses/evictions were
already being counted in PgCache, and the struct comment read:

    /* stats (introspection only — never affect semantics) */

That comment was the bug. Self-measurement treated as decoration is why the
pool could not correct itself and why no one outside could see what it was
doing. A system that cannot read its own state cannot correct, and neither can
anyone watching it.

  - pc_adapt_budget(): the loop, closed. Over a sliding window, evictions
    running at a large fraction of accesses WHILE reuse is real means the
    working set exceeds the budget — so grow it, geometrically, bounded by a
    LIVE re-read of physical memory. Evictions alone are not pressure (a scan
    evicts and never returns); evictions with reuse are. An explicit
    ENGRAM_POOL_FRAMES still wins — an operator override must not be silently
    overruled.

  - Budget derived, not declared. A constant cannot be right: 16 GiB of frames
    is arbitrary on a 48 GB host and suicidal on a 16 GB one. Even "60% of RAM
    at startup" is a guess about the future — it cannot know the store grew or
    the machine changed. Hence the live re-read.

  - pc_report(): ONE structured emission carrying the entire sensed state,
    through emit_log — El's existing telemetry, already exporting to OTLP.
    Deliberately not a function per stat, and deliberately not a bespoke
    /api/pool endpoint: both make observability something hand-written per noun
    instead of the uniform mechanism every component already has.

  - engram_pool_stats_json(): the same state readable live, wired through the
    normal builtin path (codegen arity + el_seed wrapper), so the pool can be
    observed in real time rather than reconstructed afterward from a stack
    sample.

Verified: with the exact configuration that took production down
(ENGRAM_POOL_FRAMES=65536 → 1 GiB cache against a 2 GiB store) the engram boots
clean and serves — 0.0% CPU, 13,436 nodes / 37,663 edges, embeddings complete —
and NO pressure event fires, because the barrier removed the walk that caused
it. The controller is defense in depth; the barrier is the fix.
2026-08-15 20:44:23 -05:00
bigmerge 777ccc02f0 store: extend the durable-hash write barrier to edges (kills the full-store walk)
El SDK CI - dev / build-and-test (pull_request) Failing after 10m45s
Checkpointing pushes the ENTIRE resident graph through store_put_node and
store_put_edge (see engram_store_checkpoint). Nodes were cheap: a durable-hash
compare skipped unchanged records with zero page I/O. Edges had no barrier at
all — struct comment at PgCache.barrier_on even says "node durable-hash
barrier" — so every edge was rewritten on every checkpoint, and each rewrite
runs the idempotency probe max_page_lsn_for_id -> btree lookup -> page_read.

Edges outnumber nodes ~3:1 here (37,663 vs 13,436), so routine checkpointing
degenerated into a FULL-STORE WALK in id order: random page access across the
whole 2 GiB store, repeated, overwhelmingly to rediscover nothing had changed.
LRU is worst-case under exactly that pattern — it evicts the page it is about
to want — so once the page cache was smaller than the store, the walk collapsed
into thrashing: 100% CPU, flat RSS, no forward progress, port never bound.
That took the live engram down twice on 2026-08-15.

The walk is the defect. Sizing the cache to survive it treats the symptom.

Changes:
  - dh_edge_hash(): edge counterpart of dh_node_hash, with a kind discriminator
    byte so an edge can never collide with a node of the same id in the shared
    map. created_at/updated_at/last_fired are excluded deliberately: last_fired
    is touched by activation without changing what the edge IS, and folding it
    in would defeat the barrier on precisely the hot edges that most need it.
  - store_put_edge(): barrier check + dh_set on success, mirroring
    store_put_node exactly.
  - store_scan_edges(): seed the barrier map from on-disk truth at load, so the
    FIRST post-boot checkpoint already skips unchanged edges. store_scan_nodes
    already did this and its comment says why; edges were simply never done.

Verified: with the exact configuration that killed production
(ENGRAM_POOL_FRAMES=65536 -> 1 GiB cache against a 2 GiB store), the engram now
boots clean and serves — LISTENING, 13,436 nodes / 37,663 edges, embeddings
complete, 0.0% CPU, RSS 1.14 GiB (cache resting at its budget rather than
thrashing against it). Same small cache, same store, no walk.
2026-08-15 20:38:11 -05:00
will.anderson c21074b547 Merge pull request 'runtime: engram_edges_json — kill the whole-graph file round trip' (#127) from fix/elc-rebuildable-compiler-builtins into dev
El SDK CI - dev / build-and-test (push) Failing after 10m19s
2026-08-16 01:13:44 +00:00
bigmerge 4e24d7d3f1 runtime: engram_edges_json — read edges without a whole-graph file round trip
El SDK CI - dev / build-and-test (pull_request) Failing after 13m4s
/api/graph/edges answered a read query by calling engram_save() to serialize
the ENTIRE graph to disk (128 MB) and then fs_read-ing it back. Two defects in
one line, and both bit production on 2026-08-15:

  1. The path it wrote was ~/.neuron/engram/snapshot.json — the engram
     server's CANONICAL store. A READ route overwriting the persistence
     owner's canonical file. This defect had been fixed once (export moved to
     a scratch path); it came back when the hand-written dispatch block was
     replaced by @route dispatch and the unfixed copy is the one that
     survived the merge.
  2. Cost: a full snapshot write, a 128 MB read, and a parse of the whole
     graph, per request, to return a bounded slice.

Calling it tonight overwrote the canonical snapshot and immediately preceded
an engram crash loop.

engram_edges_json(limit, offset) is the builtin that route's own TODO asked
for ("Future: add an engram_edges_json() builtin and drop the file round trip
entirely"). It walks g->edges directly and emits every persisted field.

limit <= 0 defaults to 1000, not unbounded: this is the endpoint that fell
over, and an unbounded default would preserve the failure mode under a new
name. Callers page explicitly.

Registered in codegen.el's builtin_arity (both plain and __ spellings) and
wrapped in el_seed.c per the project's C-builtin recipe.
2026-08-15 20:10:48 -05:00
will.anderson 7557ea6e19 Merge pull request 'runtime: restore engram_recall_json + cgi_* accessors (unblocks the soul build)' (#126) from fix/elc-rebuildable-compiler-builtins into dev
El SDK CI - dev / build-and-test (push) Failing after 10m15s
2026-08-16 00:57:11 +00:00
bigmerge 7351fb0a8d runtime: restore engram_recall_json + cgi_* accessors
El SDK CI - dev / build-and-test (pull_request) Failing after 10m24s
neuron's soul calls engram_recall_json (neuron-api.el:618, memory.el:80) and
cgi_principal (studio.el:72). Both existed in the runtime neuron vendored
(v1.0.0-20260501) and were absent here, so the soul could not link against
current el at all.

The dangerous part is what the obvious "fix" would have done. These look like
redundant wrappers over one impl:

    engram_search_json(q, limit)  -> eg_search_json_impl(q, limit, 0)  LEXICAL
    engram_recall_json(q, limit)  -> eg_search_json_impl(q, limit, 1)  SEMANTIC

They are not interchangeable, and the split is documented at neuron-api.el:613:
search stays LEXICAL because ~40 internal call sites pass a KEY and seven of
them DELETE every record returned. Point those at a semantic matcher and they
delete fuzzy matches. Conversely, pointing recall at search silently downgrades
the mind's entire retrieval surface from semantic to lexical — no error, just
permanently worse recall.

Implemented over engram_activate(), which in this runtime already IS the
semantic path the old with_legs=1 branch built by hand (embeds the query via
eg_embed_fetch, scores by cosine, then spreads activation one hop). Output
shape matches engram_search_json — a flat array via engram_emit_node_json —
because callers parse search's shape, not activate's envelope.

Verified: neuron's soul now compiles and links against current el, boots, and
serves /health with layers initialized.

NOTE for follow-up: current el also ships engram_retrieve_geometric_json, a
structure-first retrieval that appears to be the intended successor to recall.
Repointing the two recall call sites at it may well be the right end state and
would remove the two-wrapper shape entirely — but that is a behavioral change
that must be measured against neuron/tools/retrieval-eval/'s gold set, not
assumed. This commit preserves existing behavior exactly; it does not decide
that question.
2026-08-15 19:56:43 -05:00
will.anderson d545b69614 Merge pull request 'runtime: restore the three builtins that made elc unrebuildable' (#125) from fix/elc-rebuildable-compiler-builtins into dev
El SDK CI - dev / build-and-test (push) Failing after 11m4s
2026-08-16 00:50:43 +00:00
bigmerge 598915cc61 runtime: restore the three builtins that made elc unrebuildable
El SDK CI - dev / build-and-test (pull_request) Failing after 3m52s
The committed elc binary could not be refreshed from its own source. Rebuilding
failed with three implicit-declaration errors: el_mem_check, stdout_to_file,
stdout_restore. The compiler's own source calls all three (compiler.el:472,479,574
and codegen.el:4248) and two are registered in codegen.el's builtin_arity table —
but none were defined in this runtime.

They were found intact in ui/examples/native-hello-ios/NativeHello/el_runtime.c,
a divergent private copy of this runtime that still carried them. Ported verbatim.

Consequence of them being missing: the canonical elc binary was frozen. Source
gained @route dispatch codegen (emit_route_dispatch, codegen.el:3948) and the
@manager boundary-beat seam, but no rebuilt binary could carry them, so
neuron's soul — whose routes.el now calls the compiler-synthesized
el_route_dispatch — could not be built at all.

Verified after the fix:
  - elc rebuilds from current source, clean.
  - Self-hosting fixpoint byte-identical (stage3 == stage2).
  - The rebuilt elc emits el_route_dispatch (2 occurrences in the soul amalgam,
    previously 0) and injects engram_boundary_beat at @manager boundaries,
    i.e. the decorator seam is live rather than inert.

el_mem_check is itself the compiler's memory guard (ELC_MAX_MEM_MB, default
512MB, self-terminates before the OS OOM-killer fires) — so the runtime was
missing the very guard that would have surfaced the compiler's memory blowup
as a clean error instead of a 27GB host-killer.
2026-08-15 19:50:17 -05:00
will.anderson dab14f9100 Merge pull request 'engram: fix silently-wrong query params + make el_seed.o/el_runtime.o link' (#124) from fix/engram-query-param-and-seed-link into dev
El SDK CI - dev / build-and-test (push) Failing after 14m30s
2026-08-16 00:34:33 +00:00
bigmerge 40eb48e92f engram: fix silently-wrong query params, and make el_seed.o + el_runtime.o link
El SDK CI - dev / build-and-test (pull_request) Failing after 14m49s
Three real bugs, all found by actually running the thing rather than reading it.

1. query_param never URL-decoded. A GET of /api/search?q=neural%20network
   searched for the literal string "neural%20network" and returned []. Every
   multi-word search against the live engram has been silently returning empty
   results — not an error, an empty result, which is why it went unnoticed.
   Affects every GET route that reads query params, not just search.

2. query_param matched key names unanchored. str_index_of(qs, "q=") matches
   inside "faq=", so "?faq=X&q=Y" returned X for key "q". Verified live before
   the fix. Now searches for "&key=" against "&"+querystring so a match can
   only land on a real parameter boundary.

3. el_request_start/el_request_end were defined in BOTH el_seed.c and
   el_runtime.c, so linking the two objects together — which is exactly what
   the product build does — failed with duplicate symbols. el_seed.c's own
   comment already says these moved there ("formerly defined in el_runtime.c.
   Now self-contained in el_seed.c"); the el_runtime.c copies were left behind
   during that move. Removed them, kept declarations since http_worker calls
   them. Also added the three missing prototypes (engram_op_assert_json,
   engram_node_full_in, engram_connect_in) that el_seed.c wraps but never
   declared, which made it fail to compile standalone under C99+.

Verified: engram builds and links clean from canonical source; before/after
comparison on a copy of the real store shows "neural network" returning a real
match where the live build returns [], and "?faq=WRONG&q=MetaColloc" now
resolving to MetaColloc. Live engram on :8742 was never touched.
2026-08-15 19:34:04 -05:00
will.anderson c9f75e2592 Merge pull request 'engine: land op_assert + purview mutation wrappers (clean re-merge)' (#123) from merge-pr103-v2 into dev
El SDK CI - dev / build-and-test (push) Failing after 3m39s
2026-08-15 23:27:17 +00:00
bigmerge 09dade0613 Merge remote-tracking branch 'origin/pr/103' into HEAD
El SDK CI - dev / build-and-test (pull_request) Failing after 3m56s
# Conflicts:
#	lang/AGENTS.md
#	lang/runtime/el_runtime.h
2026-08-15 18:26:57 -05:00
will.anderson 38a8e32d6c Merge pull request 'engram: batch-cosine Adapter/Strategy/Factory over ggml (supersedes #114)' (#116) from feat/engram-ggml-cosine-batch into dev
El SDK CI - dev / build-and-test (push) Failing after 3m29s
2026-08-15 23:22:56 +00:00
will.anderson 15b66c8b1a Merge pull request 'swarm: land wt/swarm-ccr onto dev (clean re-merge)' (#122) from merge-swarm-ccr-v2 into dev
El SDK CI - dev / build-and-test (push) Failing after 3m35s
2026-08-15 23:17:56 +00:00
bigmerge 9883aa7564 Merge remote-tracking branch 'origin/wt/swarm-ccr' into merge-swarm-ccr-v2
El SDK CI - dev / build-and-test (pull_request) Failing after 4m13s
# Conflicts:
#	lang/runtime/el_runtime.c
#	lang/runtime/el_runtime.h
2026-08-15 18:17:33 -05:00
will.anderson d45a0882f3 Merge pull request 'nsbx: fail loud on daemon-not-ready + el_seed.c standalone compile' (#118) from fix/nsbx-tooling-hardening into dev
El SDK CI - dev / build-and-test (push) Failing after 3m56s
2026-08-15 23:06:01 +00:00
will.anderson 1c9de03fdb Merge pull request 'engram: make the ggml batch-cosine strategy actually compute in fp32 (recall 0.9933 -> 0.9987)' (#121) from improve/ggml-cosine-fp32-and-init into feat/engram-ggml-cosine-batch
El SDK CI - dev / build-and-test (pull_request) Failing after 3m43s
2026-08-15 23:01:29 +00:00
bigmerge c008b7228a engram: make the ggml batch-cosine strategy actually compute in fp32
#116 shipped the ggml strategy at 0.9933 id-recall against the CPU oracle
while the hand-rolled Metal kernel it replaced scored 0.9997 — a ~150x worse
error margin. That was not an inherent property of ggml. It was a usage bug in
this file, and this commit fixes it.

ggml-metal has two F32xF32 matmul kernels and picks between them purely on
ne11, the number of B rows, which for us is the query-batch size:

  ne11 <= 8  -> kernel_mul_mv_ext_f32_f32_* / kernel_mul_mv_f32_f32_*,
                templated <float, float> — genuine F32.
  ne11 >  8  -> kernel_mul_mm_f32_f32, templated
                <half, half4x4, simdgroup_half8x8, half, half2x4, ...> —
                BOTH operands narrowed to F16, despite F32 tensors on both
                sides.

The old code issued one ggml_mul_mat with ne11 = nq (300 in the benchmark),
landing squarely on the F16 path. The file's own header comment asserted the
opposite ("computes in F32 on the Metal backend"); that claim was wrong and is
replaced with the measurement.

Fix: emit ceil(nq/8) mul_mats over ne11<=8 ggml_view_2d slices of one query
tensor, all expanded into ONE graph and one ggml_backend_graph_compute, so the
node matrix is still uploaded and shared exactly once. EL_GGML_MULMAT_CHUNK
overrides the 8; setting it >= nq reproduces the old behaviour exactly, which
is also how the before/after below was measured in a single binary.

Measured, real store snapshot, 13415 live embedded nodes, dim=768, 300 real
queries, vs the CPU double-accumulated oracle (vindex_bench, offline copy of
the store — no live service touched):

  id-recall   same-rank |Δdist| max   mean
  old (ne11=300)   0.9933   6.80e-05   1.43e-05
  new (ne11<=8)    0.9987   4.77e-07   9.30e-08
  hand-rolled      0.9997   3.58e-07   7.55e-08

~145x better max error, ~154x better mean — now the same order of magnitude as
the hand-rolled kernel rather than 150x off it.

The cost is real and is documented rather than buried. Median of 15 reps of
the whole batch_multi() call, three runs: 13.2-14.4ms unchunked, 19.9-20.2ms
chunked, 17.7-18.0ms hand-rolled. Correctness costs ~+6.7ms per 300-query
batch and leaves ggml ~12% behind the hand-rolled kernel instead of ~35%
ahead. It cannot be recovered inside ggml: an fp32 matmul on Metal must
re-stream the node matrix once per <=8 queries, and ggml's Metal backend ships
no fp32 TILED matmul, so "fast" and "fp32" are genuinely exclusive there.

Two things that did NOT work, recorded so nobody retries them:

  - ggml_mul_mat_set_prec(t, GGML_PREC_F32) does nothing here. Error was
    bit-identical with and without it (1.038e-05 either way) — ggml-metal has
    no F32-accumulating mul_mm kernel to switch to. ne11 is the only lever.
  - The ACCEL/BLAS device looked excellent in an isolated compute-only probe
    (3.4-4.0ms, mean |Δdot| 1.5e-08) but is dominated on BOTH axes end-to-end
    (0.191 ms/query at 0.9973 recall vs 0.125-0.142 at 0.9987), because the
    probe was not competing for the same CPU cores the real call path is. It
    stays reachable via EL_GGML_DEVICE as a no-Metal fallback, labelled as
    measured-and-rejected, not as a recommendation.

Also corrected: the ~7.8s "cold start" blamed on this file is not this file
re-initialising per call — init was already cached. It is Apple's shader cache
missing on ggml's embedded metallib (~650 kernels), keyed on the library and
shared across processes: the first load on a machine reports
"loaded in 7.670 sec", the next run of a *different* binary reports 0.009 sec.
Once per machine per ggml version, not once per process, and not ours to fix.
Warm ggml init is 44-53ms vs 36-117ms for the hand-rolled strategy.

Loading only libggml-metal.so instead of every plugin in the directory is kept
for tidiness, and explicitly documented as NOT a speedup: 44.7-52.4ms against
46.9-58.9ms, the same number inside noise.

The -2.0 sentinel contract is unchanged and re-verified at batch sizes that
straddle the chunk boundary (1,7,8,9,16,17,33), plus NULL rows, dim
mismatches, zero-norm rows, and an all-invalid population. Notably the old
ne11=300 path fails that same check at a 2e-6 cosine tolerance with 2299
mismatches, which is an independent confirmation of the defect.
2026-08-15 17:57:09 -05:00
bigmerge 3718bf0380 runtime: port missing __channel_* primitives into el_seed.c
El SDK CI - dev / build-and-test (pull_request) Failing after 3m44s
runtime/channel.el has always called __channel_new/__channel_send/
__channel_recv/__channel_try_recv/__channel_close, but these were only ever
implemented in the pre-restructure lang/el-compiler/runtime/el_runtime.c.
When the canonical runtime was consolidated onto the release copy
(lang/runtime/el_runtime.c) and el_seed.c became the sole C dependency,
the channel implementation was never carried forward — __mutex_new made the
move, __channel_* did not. Any El program using Go-style channels currently
fails to link on dev.

Ported the working buffered-MPMC-channel implementation (mutex+condvar+
circular buffer, bounded and unbounded modes) from the old el_runtime.c
verbatim, adapted only to el_seed.c's arena API (seed_arena_track in place
of el_arena_track). Declared in el_seed.h alongside the existing mutex
primitives.
2026-08-15 17:50:04 -05:00
bigmerge 9d40f87926 ingest: unify transduce_prose/transduce_structured into one transduce()
transduce() is now THE single mechanism: one function, no content-type
branch inside it. It never asks whether `source` is prose, JSON, or
raw/opaque bytes (audio, etc.) — it runs one algorithm unconditionally:
split on "\n\n" as a universal boundary-marker check, and if that finds
no boundary, fall back to fixed 4096-char windows. Same node/edge wiring
(root -contains-> chunk, chunk -precedes-> next, "#"-prefixed chunk gets
a heading/section_of link) regardless of what's inside a chunk. Dedup is
the existing find_existing_by_content path via merge_manifold, applied
uniformly. The old transduce_structured JSON dataset/records/feature-node
interpretation is deleted outright, not just unused — a JSON file now
gets chunked and deduped like anything else, with no pre-computed
structure. All five ingest_* entry points still exist unchanged in name
and role; ingest_file/ingest_dir/ingest_url/ingest_llm now call the one
transduce() (ingest_stream builds its own turn-nodes directly and never
called either old function, so it's untouched).

This unlocks raw/opaque content (audio, or anything else with no natural
text/JSON shape) without any DSP, LLM call, or external API: transduce()
chunks it exactly like it chunks anything else. There is zero semantic
understanding of audio (or any payload) claimed or built here — any
meaning is expected to emerge later from Neuron's own existing mechanisms
(embedding, spreading activation, dedup) acting on this real geometry
over time.

Two small C builtins added to el_runtime.c/h (fs_size, fs_read_b64_chunk)
because El strings are NUL-unsafe under strlen-based ops and fs_read()'s
result silently truncates at the first embedded NUL, which is routine in
real binary/audio bytes. ingest_file compares fs_read()'s string length
against a real fs_size() stat() count; on mismatch it rebuilds the
payload as base64-encoded fixed 3072-byte windows read directly off disk
(binary-safe in C, verbatim, no invention), joined with the same "\n\n"
marker transduce()'s boundary scan already looks for. This is a
mechanical fidelity fix, not interpretation of content — transduce()
never learns a fallback happened. Registered both builtins' arity in
codegen.el; did not rebuild the elc compiler binary itself (unrelated,
pre-existing gap: self-hosting elc via el_seed.c fails on this worktree
independent of this change, reproduced with codegen.el reverted) — the
existing elc binary compiles calls to unregistered builtins via its
already-existing arity=-1 passthrough, confirmed by an actual clean
`elc ingest.el` + `cc` build against the modified el_runtime.c.

INGEST_KIND keeps existing only as an acquisition-mechanism selector
(dir/file/url/llm/stream — which RPC to use to fetch bytes), not as a
content-type flag; the redundant "structured" value (an alias for "file"
that hinted the now-deleted JSON branch) is removed. ingest_dir drops its
file-extension filter for the same reason: transduce() takes anything now.

Verification: local manifold construction confirmed correct against a
real captured audio file (will_clean.wav, 304288 bytes, and a 12288-byte
real prefix slice) — exact expected node/edge counts both times
(101 nodes/199 edges full file; 5 nodes/7 edges for the slice, matching
ceil(bytes/3072)+1 nodes and 2n-1 edges), with real, verbatim base64
content confirmed decoding back to the actual WAV header bytes. Compiles
clean via the real elc + the modified el_runtime.c/engram_*.c (built and
booted an actual sandbox engram off this exact source with `nsbx create
--branch`).

NOT verified this session, disclosed rather than papered over: end-to-end
server-confirmed persistence (a real before/after /api/stats delta, and a
fetched node by id) for the audio, prose, and JSON-fixture cases. Every
local nsbx sandbox engram tried tonight (two stock pre-#109 binaries
hitting the known O(N*D) brute-force scan bug, then a fresh #109/HNSW
binary built from current dev) took minutes-to indefinitely long on the
final /api/load-merge write's embedding step and hit the client's 60s
HTTP timeout before responding, even for a 5-node write. This is
confirmed as real (if slow) forward progress, not a hang: the sandbox's
WAL file was observed growing steadily across every attempt. The code's
own pre-existing HONESTY GATE correctly refused to report success in
every case, returning "load-merge failed: ..." with a
"nothing below this manifold was confirmed persisted by the server" note
instead — exactly as designed. This is an environment/infrastructure
limitation, not a defect introduced by this change: the engram server
binary itself is untouched by this commit.
2026-08-15 17:50:04 -05:00
bigmerge 90d3f0bc76 engram: port PR #105's 3 genuine wins onto dev's existing cosq/e_eff semantic layer
Reconciles PR #105 ("fix: engram search latency — pin embed model, cache
query embeddings, bound activate BFS") with dev's ACTUAL current
engram_activate, rather than the ancient pre-restructure snapshot #105 was
built against.

WHY THIS NEEDED RECONCILIATION, NOT A DIRECT PORT: #105's single commit
(1dc49b1) modifies `lang/el-compiler/runtime/el_runtime.c` — a path that does
not exist on dev (dev has `lang/runtime/el_runtime.c`; the restructure that
renamed it happened after #105's branch point, which traces to a July 22
merge-base, weeks before the M8/M8.1/qgate/fan-effect/adjacency-index work
this file has grown since). #105's own engram_activate is consequently the
PRE-restructure version: no adjacency index (O(E) full edge scan per hop),
no query-aware qgate, no ACT-R fan effect, no eg_edge_eff_weight, and no
awareness of dev's cosq/e_eff embedding-blend semantic layer — it built a
parallel `g_qcache`/`engram_embed_raw` mechanism from scratch against code
that no longer exists at that path. A raw merge/cherry-pick was not possible
and would have been wrong even if it were: taking #105's tree wholesale would
have thrown away everything dev grew in the meantime (qgate, fan effect,
adjacency index, and this session's own M8 HNSW vindex integration).

RECONCILIATION: kept dev's cosq/e_eff mechanism as the semantic layer
entirely intact (unchanged by this commit) and ported #105's three genuinely
additive wins on TOP of it, at their equivalent sites in the CURRENT
eg_embed_fetch/engram_activate:

  1. keep_alive:-1 on the Ollama embed request body (eg_embed_fetch) — pins
     the embed model resident so a larger generation model loading under
     unified-memory pressure can't evict it and force a cold reload on the
     next search (#105 measured ~2.2s cold vs ~0.02-0.05s warm).
  2. Query-embedding cache upgraded from dev's single-slot (`_eg_qcache_text`,
     only ever remembered the LAST query) to a direct-mapped, FNV-1a-keyed,
     1024-slot cache (reusing the existing engram_id_hash) — so the
     curiosity loop's rotating phrases actually hit the cache instead of
     evicting each other every call. Same "pointer owned by the cache, not
     freed by caller" contract as before, just per-slot instead of global.
  3. Beam cap on the layer-1 spreading-activation BFS (new
     engram_activate_beam(), tunable via ENGRAM_ACTIVATE_BEAM, default 128).
     The FIFO frontier is processed in hop-level batches (entries sharing
     .hops are provably contiguous — see the code comment); when a level
     exceeds the beam width, only the top-`beam` by activation actually
     EXPAND. Every node in an oversized level still gets reached[]/best_bg[]
     recorded (that happens at enqueue time, one level up) and appears in
     the reported/promoted set — the cap bounds associative SPREAD width
     only, never recall of what was already found. Kept as a genuine
     additional bound even though the adjacency index + qgate + fan effect
     already mitigate #105's original "hub-node explosion" failure mode for
     a different reason: those prune WHICH targets matter; this bounds
     worst-case width regardless.

Everything else in dev's engram_activate — cosq/e_eff, the qgate rescale,
the fan effect, eg_edge_eff_weight, the M8 HNSW vindex seed discovery from
the #109 reconciliation earlier this session — is untouched.

VERIFIED (nsbx sandbox only, live :8742/:7770 never touched): cc -std=c11
-O2 clean build; booted in an isolated sandbox against a real cloned
production snapshot (13,424 nodes / 37,656 edges); ran 5 activate() calls
across rotating queries at depth 3, including the same query issued twice
non-consecutively (2nd hit landed at 476ms vs the 1st at 483ms — consistent
with a cache hit once Ollama's own warm-model latency is accounted for; no
crash, correct varied result counts (367-2610 nodes) each call; act-stats
JSON read correctly throughout.

Built on top of the M8/#109 reconciliation (bacaf3d, merged to dev as
#109) — dev's current HEAD at the time of this commit.
2026-08-15 17:50:04 -05:00
bigmerge 2d0aef4ef8 lang: declare the el_runtime.c symbols el_seed.c's wrappers call
El SDK CI - dev / build-and-test (pull_request) Failing after 3m55s
runtime/el_seed.c does not compile standalone via the exact command
tools/install.sh uses (`cc -std=c11 -O2 -I runtime -c runtime/el_seed.c`):
51 of its __-prefixed wrapper functions (http serving, JSON access, key-val
state, URL/HTML escaping, and the whole engram_* node/edge/layer/search
surface) call unprefixed counterparts that are implemented in el_runtime.c,
not in el_seed.c itself, and el_seed.c never declared them -- a toolchain
that treats an implicit function declaration as a hard error under C11
fails the compile outright.

install.sh already compiles el_seed.c and el_runtime.c as separate objects
and archives both into libel.a, so the symbols are always present at link
time; el_seed.c alone was just missing the prototypes.

A plain `#include "el_runtime.h"` was tried first and rejected: it redefines
el_to_float/el_from_float, which el_seed.h already provides -- a real
compile error, not a style preference. Added narrow prototypes instead,
copied verbatim from el_runtime.h, for exactly the 51 symbols el_seed.c's
wrappers reference and nothing else.

Verified clean:
  - `cc -std=c11 -O2 -I runtime -c runtime/el_seed.c` (install.sh's exact
    per-file compile) -- 0 errors, 0 warnings, even with -ferror-limit=0.
  - full `tools/install.sh` run -- compiles both objects and archives them
    into libel.a successfully.

Separately (not fixed here, out of scope): AGENTS.md's documented compiler
self-rebuild command links elc-new.c against el_seed.c, but elc-new.c's own
generated `#include "el_runtime.h"` line and 3 undeclared symbols
(el_mem_check, stdout_to_file, stdout_restore -- present in neither
el_runtime.c nor el_seed.c) mean that command fails regardless of which
runtime file it's linked against; and install.sh's libel.a only archives
el_seed.o + el_runtime.o, so any program that calls into the engram_*
surface fails to link against it (el_runtime.c's engram_* wrappers need
engram_store.c/engram_geometry.c/engram_reason.c/engram_cognition.c/
engram_vindex.c, none of which install.sh compiles in). Both are real,
pre-existing, and independent of this fix -- worth their own look.
2026-08-15 17:32:58 -05:00
bigmerge 979e820f68 nsbx: fail loud on daemon-not-ready instead of printing a false success banner
Three confirmed-live bugs tonight:

- `nsbx up` printed "daemon did not become ready" immediately followed by a
  green "your sandbox is ready" banner and exited 0, because the existing-
  sandbox restart path (`daemon_alive || start_daemon`) never checked
  start_daemon's return code. `cmd_build` had the identical unguarded
  pattern, plus `cmd_run`/`cmd_validate`'s own start-if-dead calls. All four
  now `|| die` with a message pointing at daemon.log.

- `nsbx status`/`nsbx list` reported bare "state: running" for a process
  that's alive (passes kill -0) but not actually answering /api/stats --
  pegged, hung, or mid-boot. Added daemon_health(), which does the real
  stats fetch and distinguishes stopped/running/unresponsive; both commands
  now say "running but NOT RESPONDING" with a next-step hint instead of
  silently going quiet on the stats field. Reproduced live against another
  agent's actively-running (CPU-pinned, non-responsive) sandbox tonight, and
  again via a deliberate SIGSTOP on a throwaway sandbox.

- Sandboxes carried no visible signal that their binary predated a relevant
  fix. `status`/`list` now show the binary's sha + real build timestamp
  (mtime survives `cp -p`), plus a best-effort staleness note: for
  stock-prod clones, compare against the currently-configured live binary;
  for source/branch builds, compare the recorded source commit against
  local origin/dev via merge-base --is-ancestor.

Also, found live while verifying the above:

- A cold boot under concurrent sandbox/CPU load can legitimately take past
  the old hardcoded 15s readiness window. Made it configurable
  (NSBX_READY_TIMEOUT_SECS) rather than just widening the default blindly.

- cmd_create's post-boot baseline capture could silently record sbx_baseline
  as 0/0 when the stats fetch came back empty right after the auto-remerge
  step -- which would make every future `nsbx validate` zero-loss/reboot-
  prove check trivially PASS regardless of real data loss. Added a bounded
  retry and a loud warning if it still comes back empty.

- Sharpened a handful of "no such sandbox" / missing-binary errors to name
  the next command instead of just stating the failure.
2026-08-15 17:32:44 -05:00
bigmerge b3f410fc91 engram: batch-cosine Adapter/Strategy/Factory over ggml, supersedes hand-rolled PR #114
El SDK CI - dev / build-and-test (pull_request) Failing after 4m29s
Stop hand-rolling GPU kernels for batch cosine similarity — use ggml (the
MIT-licensed compute library underneath llama.cpp, installed standalone via
Homebrew) as the preferred backend, without ripping out PR #114's
carefully-verified hand-rolled Metal shader.

Structure: one stable public adapter (eg_cosine_batch.h, zero #ifdef at call
sites) backed by three selectable concrete Strategies behind an internal
vtable (eg_cosine_batch_strategy.h) chosen by a Factory (eg_cosine_batch.c):

  - eg_cosine_batch_strategy_ggml.c    — NEW. ggml + dynamically-loaded Metal
                                          backend plugin (ggml_backend_load_all_from_path
                                          + ggml_mul_mat for the batched dot
                                          product), gather/scatter around the
                                          -2.0 sentinel contract.
  - eg_cosine_batch_strategy_metal_hand.m — PR #114's original hand-rolled
                                          Metal shader bridge, preserved
                                          almost verbatim, now one strategy
                                          among several rather than the only
                                          option. eg_cosine_batch.metal kept
                                          byte-identical to the original.
  - eg_cosine_batch_strategy_cpu.c     — universal always-false fallback
                                          (direct descendant of PR #114's
                                          eg_metal_cosine_stub.c).

Selection: EL_COSINE_BATCH_STRATEGY=ggml|metal|cpu|auto (default: ggml first,
then hand-rolled Metal, then CPU — first available wins), plus back-compat
EL_METAL_COSINE=0 to disable every GPU-backed strategy. build_vindex_bench.sh
compiles all three strategies on Darwin, CPU-fallback-only elsewhere.

vindex_bench.c now reports BRUTE-GGML and BRUTE-METAL side by side against
the same CPU oracle, on the same dataset, in one run (real numbers vs. real
store snapshot in the PR body).
2026-08-15 17:16:41 -05:00
bigmerge 1010185978 Add op_assert grounded-envelope primitive and purview-bounded mutation wrappers
El SDK CI - dev / build-and-test (pull_request) Successful in 6m33s
Adds engram_assert_json — a grounded "assertion envelope" primitive for a
realizer/op_assert seam (per backlog bl-53/#57) — plus purview-scoped
mutation wrappers engram_node_full_in/engram_connect_in, which refuse
non-default purviews rather than silently mutating the live store. Threads
through el_seed.c/h wrappers and the codegen.el arity table per the
project's existing C-builtin recipe.

Also rewrites lang/AGENTS.md build docs with verified (2026-08-15) findings
that el_seed.c does not compile standalone.
2026-08-15 14:24:11 -05:00
bigmerge ff37835ae5 swarm: document the single-writer invariant (Rule 4) in README
El SDK CI - dev / build-and-test (pull_request) Failing after 12m1s
2026-08-14 21:46:23 -05:00
bigmerge e5c80359a8 swarm: Rule 4 — engram-write is @manager-ONLY, enforced by capability
New hard invariant (Will): only the orchestrator mutates global engram state;
workers are read-only against the full engram + write only their own local
geometry. This is an AUTHORITY gate (capability), not a health gate — a worker
is STRUCTURALLY UNABLE to mutate global engram state regardless of engram health.

- containment.el: scope tokens now carry a caps set. Orchestrator token holds
  engram:write + dharma:emit (@manager-only, the VBD rule that only the manager
  mutates global state); worker token holds ONLY engram:read. Rule 4:
  containment_check_engram_write / _dharma_emit reject any caller lacking the
  capability — same scope-token mechanism as the live Rule-2 denial.
- swarm.el: swarm_engram_write is the ONLY engram write path, gated by Rule 4;
  a worker token is denied before any HTTP is issued (no mutation). The curated
  merge (commit=1) is the sole writer: the orchestrator commits approved
  geometry via its write-capable token. Workers' full-engram READ stays intact.
- reshape_surface.el: compose op_write (json_escape_string) for the commit path.
- harness: Rule-4 suite proven — worker engram-write DENIED by capability, no
  node created, violation journalled; orchestrator passes the gate as sole
  writer. 24/24 green on the :8901 clone with real cognition.

Authority gate holds independent of daemon write-health (proven with daemon
both alive and, earlier, crashed). Prod :8742 untouched.
2026-08-14 21:46:03 -05:00
bigmerge b53b5b4e8a swarm: document real-cognition binding + HAVE_CURL build note in README 2026-08-14 21:19:11 -05:00
bigmerge 20bd9ed00b swarm: bind reshape's proven primitives — REAL-COGNITION local swarm end-to-end
Binds the api-reshape surface at wt/api-reshape@d4f401d (op_think/read/attend/
learn, verified against engram.cognition-20260814) into the swarm:
- reshape_surface.el composes the reshape's proven read/cognition primitives
  verbatim (write ops omitted — they need the gate-1 write-healthy clone).
- primitive_binding.el: bound_think -> op_think over the worker's NODE-ID
  anchor (ctx.input); attend/learn bound behind SWARM_WRITE_HEALTHY.
- cognize blueprint derives the vote verdict from the REAL gradient's n_support
  (json_get_int) — per-anchor diversity (6/16/87 support) drives a genuine vote.
- build.sh now defines HAVE_CURL. CRITICAL FIX: without it every http_* was a
  '{"error":"not built with HAVE_CURL"}' stub, so prior 'live engram'
  retrieval was a false positive (matched the ref string, not real content).
  With HAVE_CURL the swarm genuinely hits /api/think on the :8901 clone.

harness_real_cognition.el: 17/17 GREEN with seam=decorated — 8 native-thread
workers each a REAL think (768-dim gradient) over its CCR-scoped node-id anchor,
@manager reduce+vote convergence, all 3 containment rules incl. live Rule-2
denial, afferent telemetry (8 real think signals), durable work-tracking. Reads
only — daemon stays healthy; writes stay gated on the gate-1 clone. Prod :8742
untouched.
2026-08-14 21:18:57 -05:00
bigmerge 70982498e0 swarm: document local-swarm harness + one-flip seam in README 2026-08-14 20:58:17 -05:00
bigmerge 373265c05d swarm: local-swarm integration harness + one-flip primitive seam + telemetry
- primitive_seam.el: SWARM_PRIMITIVE_SEAM selects stub (default, hermetic) vs
  decorated (reshape's dharma-bus primitives). Every seam call is an afferent
  signal; telemetry (seam_mode + afferent tick) rides the vertical result path.
- primitive_binding.el: THE ONE FLIP POINT — bound_think/attend/learn today fall
  back to the stub; when the reshape's decorated primitives land, flip one line
  each and set SWARM_PRIMITIVE_SEAM=decorated. No other change anywhere.
- swarm.el: default blueprint routes think through the seam; the @manager
  aggregates afferent counters from worker results (containment-safe, no shared
  bus register) and journals a swarm.telemetry record; telemetry in the return.
- harness_local_swarm.el: 17/17 GREEN on :8901 with the stub — 8 native-thread
  workers at concurrency 4, reduce+vote convergence, CCR scoping+non-leak, all
  three containment rules (incl. live Rule-2 denial), durable work-tracking,
  afferent telemetry observed. Runs identically under seam=decorated today
  (binding fallback), proving the flip path executes.

Engram writes stay opt-in (durable journal is the substrate); daemon healthy.
2026-08-14 20:58:05 -05:00
bigmerge ed722b9e2e swarm: build harness executable + module load order 2026-08-14 20:44:01 -05:00
bigmerge b0a78c5737 swarm: capability README — architecture, framework grounding, built vs stubbed 2026-08-14 20:43:46 -05:00
bigmerge 447d042022 swarm: HTTP-backed primitive retrieval + live-engram integration test
- primitive_attend retrieves over HTTP (POST /api/search) when ENGRAM_URL is
  set — the location-independent worker model — falling back to the in-process
  store otherwise. Proven against the isolated :8901 clone: CCR compiled a
  bounded context from REAL mind content (VBD/intellectual-dna).
- gate the engram work-tracking mirror behind SWARM_MIRROR=1; the durable
  substrate is always the JSONL journal, so a swarm never depends on the mind
  to track its work. (Repeated POST /api/nodes mirror writes were observed to
  crash the isolated daemon — a daemon-side write-path robustness issue;
  retrieval POST /api/search is solid. Prod :8742 never touched.)
- integ_engram: CCR real-retrieval + full swarm completion against live clone.
2026-08-14 20:43:05 -05:00
bigmerge d4e82d3d56 swarm: convergence strategies + failure threshold, hardened El JSON usage
- vote/merge/reduce/collect convergence proven end-to-end; failure threshold
  aborts a swarm below min_success_ratio (integer per-mille) and completes
  when failures are within tolerance, with worker.failed + swarm.aborted
  tracked durably.
- worked around three El runtime/codegen semantics surfaced during the build:
  json_set inserts RAW (use json_set_str for string values); json_set cannot
  update an existing key (vote tallies via list rescanning); json_array_get
  keeps quotes (use json_array_get_string). Also: float division is unreliable
  (swarm uses integer math), and a let-rebind in a deeply nested if/else does
  not propagate outward (accumulators kept at one block level).

test_convergence: 8/8; test_swarm: 12/12.
2026-08-14 20:37:35 -05:00
bigmerge 40bb6ff579 swarm: orchestrator, CCR context compilation, containment rules, primitive seam
- swarm.el: coordinator running fan-out/converge on El NATIVE threads
  (thread.el spawn/join) in bounded concurrency waves, order-preserving;
  convergence strategies collect/merge/vote/reduce; integer per-mille failure
  threshold (El float division is unreliable — avoided deliberately).
- ccr.el: per-worker Compiled Context Routing — retrieval/scoping/compaction
  into a bounded, minimal package; the compiled-context boundary is the
  security boundary (a worker cannot receive or leak sibling inputs).
- containment.el: the three Swarm containment rules enforced via scope tokens
  (Rule 1 no join, Rule 2 no open, Rule 3 no lateral edge) + execution-tree
  lateral-edge check.
- primitives.el: attend/think/intend/act/learn seam the swarm composes over,
  with engram-backed fallbacks and an explicit binding point for the reshape.
- prototype json_array_push in el_runtime.h (defined but unprototyped).

test_swarm: 12/12 — native fan-out/converge, bounded concurrency, durable
tracking, CCR bounding + non-leak, and all three containment rules.
2026-08-14 20:29:04 -05:00
bigmerge d5411fb58a swarm: durable, inspectable work-tracking journal (worktrack.el)
Single-writer append-only JSONL journal keyed by correlation ID: swarm +
worker + convergence records, reconstructable into a status report. Optional
engram mirror via POST /api/node when ENGRAM_URL is set. Coordinator is the
only writer (workers return structured results), which is race-free and
enforces Swarm containment rule 3 by construction.

Also prototype now_millis/now_ns in el_runtime.h (defined in el_runtime.c but
unprototyped — blocked any El program needing a real ms clock under clang 21).

Test proves durability + inspectability end-to-end.
2026-08-14 20:23:13 -05:00
bigmerge b2aac4bf89 el runtime: prototype + fix channel/mutex seed ABI for modern clang
el_runtime.h declared only __thread_create/__thread_join; the mutex and
channel seed primitives (__mutex_*, __channel_*) were defined in
el_runtime.c but never prototyped. Under Apple clang 21 (C11) the missing
prototypes became implicit-declaration errors, and the void-returning
__channel_send/__channel_close mis-typed el_val_t (long long) returns,
so any El program using runtime/channel.el failed to compile.

- add prototypes for __mutex_new/lock/unlock and all __channel_* to el_runtime.h
- make __channel_send/__channel_close return el_val_t nil so elc's
  trailing-expression codegen for the void El wrappers type-checks

Additive; unbreaks native channels for every downstream El program.
2026-08-14 20:18:06 -05:00
73 changed files with 12140 additions and 331 deletions
+630
View File
@@ -0,0 +1,630 @@
# El Test Framework — Design
**Status:** draft for review
**Author:** Neuron
**Date:** 2026-08-15
**Worktree:** `/Users/will/Development/neuron-technologies/el-worktrees/elc-memory-investigation`
---
## 0. The forcing requirement
We have a confirmed quadratic in `elc`. Peak memory in the old shipped binary and wall-clock in
the current source both grow as O(input²). We cannot fix it, because we cannot test it.
Everything in this document is downstream of one sentence: **a test framework must be able to fail
a build when an operation's growth curve degrades from linear to quadratic.**
That is not a nice-to-have bolted onto a correctness framework. It is the requirement that
determines the architecture. Correctness testing is the easy half.
Second-order requirement, learned the hard way tonight: **the framework must report per-test timing
by default.** The current framework prints `N passed, M failed` and nothing else. That is why a
3.58-second test file sat in the suite unnoticed. A framework that is structurally blind to time
cannot surface the defect class we most need to catch.
---
## 1. What exists today, measured
### 1.1 Two competing systems, neither complete
**System A — `lang/runtime/test.el`.** Manual registration, El-level.
**System B — the compiler's `test { }` block + `elc --test`.** Emits its own harness `main()`
with `__el_pass` / `__el_fail` globals (`codegen.el:3777-3796`).
They do not share a result model. Neither has timing. Both are in the tree.
### 1.2 Specific defects in System A
| Defect | Location | Consequence |
|---|---|---|
| All state as JSON strings in a global string-keyed map | `test.el` throughout | every assertion is `state_get``str_to_int``int_to_str``state_set` |
| Failure list appended by string slice + concat | `_test_json_append` | O(n²) in failure count |
| One OS thread spawned per test | `_test_run_one` via `__thread_create`/`__thread_join` | thread spawn per test, purely to get dispatch-by-name through dlsym |
| Manual registration pairing a string to a function name | `test_case(name, fn_name)` | typo ⇒ test silently never runs, suite still reports pass |
| Counters are assertion-level, global | `_test_pass_count` etc. | no per-test record exists at all |
| No timing, no structured output, no fixtures, no tags, no filtering, no parameterization, no benchmarks | — | — |
The registration defect is the serious one. It is not a slow framework, it is a framework that can
report success for tests that did not execute.
### 1.3 Measured cost structure
Per test file, current build model:
| Step | Time |
|---|---|
| `elc` compile `.el``.c` | 0.00s (small files) |
| **`cc` el_runtime.c → .o** | **0.14s** |
| `cc` test .c → .o | 0.02s |
| link | 0.02s |
> **STALE as of el #132 — re-measured 2026-08-16.** The `test_compiler` figure below was
> *entirely* the `strlen`-per-character quadratic, now fixed. Re-measured on the same host:
> **3.58s → 0.03s (119x)**, and the 422 KB compiler concatenation likewise compiles in 0.03s.
> The table is retained only as the historical record that motivated the gate. The remaining
> per-file cost is the redundant `el_runtime.c` rebuild, which §9's compile-once architecture
> addresses.
Per-file `elc` time across the existing suite:
| File | Bytes | elc time |
|---|---|---|
| `test_compiler` | 29,685 (+394 KB of imports) | **3.58s** |
| `string_test` | 18,545 | 0.01s |
| all other 9 files | 2.210 KB | 0.00s |
Two distinct defects in two distinct regimes:
1. **`test_compiler.el` imports all five compiler sources** — 394 KB in one translation unit. Its
3.58s is entirely the quadratic. It is the only file where the quadratic bites.
2. **Every other file's cost is 100% redundant `el_runtime.c` rebuilds** — 480 KB of identical C,
recompiled once per test file.
Neither is fixed by making the compiler faster. Both are fixed by the architecture below, and the
speedup is a by-product of building it correctly, not the goal.
### 1.4 The asset worth keeping
`codegen.el:3651-3652` already collects `test_names` / `test_c_names` — **the compiler already does
compile-time test discovery.** It then discards that registry into a hardcoded `main()`.
That registry is precisely the seam Go's `_testmain.go` and Rust's `test_main_static` are built on.
The mechanism we need is half-built and wired to the wrong thing.
---
## 2. Grounding — the common spine of excellent frameworks
Researched from primary sources: Go `testing`/`go test`, Rust `libtest`/Criterion, JUnit 5 Platform,
NUnit 3, JMH, Google Benchmark. Six invariants hold across all of them.
1. **A registry is built before execution**`(name, metadata, fn-ptr)` triples. Go generates it
from an AST scan; Rust synthesizes it in a compiler pass; JMH emits it as a build-time resource;
JUnit/NUnit build it reflectively. **Reflection is an implementation of the registry on runtimes
where it is cheap. It is never the architecture.**
2. **Discovery strictly precedes execution.** Every good capability — filtering, listing, counting,
sharding, IDE trees, re-run-failed-only, dry runs — is a consequence of this ordering.
3. **A hierarchy with stable, path-shaped unique IDs.** `TestFoo/subcase_2`. Selection is regex over
that path, one pattern per level.
4. **The framework is a prebuilt library; only the entry point is generated.** "Compile once, link
many" is always: framework archive compiled once + a small generated table + one
`MainStart(deps, registry)` call. Nobody recompiles the harness per test file.
5. **Execution emits an event stream; reporters are downstream renderers.** Human text, NDJSON,
JUnit XML, TAP are all transforms of one event stream. Go's one architectural mistake is doing
this backwards — `test2json` parses human output, and has shipped bugs when user output contains
`--- PASS:`.
6. **A dependency-injection seam at the boundary.** Go's `testdeps.TestDeps` exists so `testing`
can avoid importing `regexp`, profilers, and coverage. The execution core knows nothing about
output formats.
---
## 3. Architecture
### 3.1 The seam
```
┌─────────────────────────────────────────────────────────────┐
│ user code: foo.el with test { } / bench { } blocks │
└───────────────────────────┬─────────────────────────────────┘
│ elc --test
┌─────────────────────────────────────────────────────────────┐
│ generated C (per suite, tiny): │
│ __el_test_fn_0 .. _N lowered test/bench bodies │
│ __el_registry[] static table: name/kind/file/ │
│ line/tags/sizes/expected-O │
│ __el_dispatch(i) generated switch → body │
│ main() { return el_test_main(argc, argv); } │
└───────────────────────────┬─────────────────────────────────┘
│ cc + link (registry only)
┌─────────────────────────────────────────────────────────────┐
│ libeltest.a — PREBUILT ONCE │
│ • el_runtime.o (the 480 KB, compiled once, ever) │
│ • eltest.o the runner, WRITTEN IN EL │
│ discovery view · filtering · execution · fixtures · │
│ timing · benchmark harness · curve fitting · reporters │
└─────────────────────────────────────────────────────────────┘
```
The framework is written in El, compiled to C once, archived. Per-suite compilation touches only
the generated registry. This is Go's model, and it is strictly better for us than Go's because we
own the compiler and already have the AST — no separate source-scanning pass is needed.
### 3.2 Why the runner is in El and the registry is in C
El has no closures and no first-class function pointers. The registry must therefore hold C function
pointers, and it is generated C.
The runner stays in El and reaches the registry through a small builtin surface — indices, not
pointers:
```
__el_reg_count() -> Int
__el_reg_name(i) -> String
__el_reg_file(i) -> String
__el_reg_line(i) -> Int
__el_reg_kind(i) -> Int // 0=test 1=bench
__el_reg_tags(i) -> Int
__el_reg_sizes(i) -> String // JSON array, empty for tests
__el_reg_expect(i) -> Int // complexity class enum, 0 = none
__el_reg_invoke(i) -> Int // runs the body via the generated switch
```
Nine builtins. Everything else — filtering, lifecycle, statistics, curve fitting, all reporters —
is El. That satisfies "written in El" without pretending El can do something it cannot.
### 3.3 Result model
The unit is a **result record**, not a counter:
```
TestResult {
id String // slash path: "parser/handles_empty_input/case_3"
file String
line Int
status Status // Pass | Fail | Error | Skip
duration Int // nanoseconds, ALWAYS populated
message String // assertion detail: expected vs actual
output String // captured stdout/stderr for this test
assertions Int
}
```
`Fail` = an assertion failed. `Error` = unexpected crash/abort. This distinction is load-bearing —
every CI consumer depends on it, and the JUnit XML schema encodes it as distinct elements.
---
## 4. Authoring surface
### 4.1 Tests
`test { }` already exists. Keep it. Add subtests and hierarchy:
```el
test "parser/empty input" {
assert_that(parse(""), is_err())
}
test "parser/table" {
for case in [["", 0], ["a", 1], ["a b", 2]] {
subtest(case[0]) {
assert_that(token_count(case[0]), equals(case[1]))
}
}
}
```
Subtest IDs compose as `parser/table/a_b`. Filtering is `--run 'parser/table/.*'`, one regex per
path segment, exactly as Go does.
**We do not build a parameterized-test annotation system.** Table-driven loops plus subtests subsume
`@ParameterizedTest`, `@MethodSource`, `@CsvSource`, and `TestCaseSource` entirely, at zero framework
surface. This is Go's single biggest ergonomic win over JUnit and NUnit.
### 4.2 Fixtures
Per-file and per-test only, plus a LIFO cleanup stack:
```el
setup_all { ... } // once per suite
setup { ... } // before each test
teardown { ... } // after each test
teardown_all { ... }
```
and inside a test, `cleanup { ... }` registering LIFO-ordered teardown.
**We do not build JUnit 5's extension SPI** — seventeen callback interfaces, hierarchical stores,
registration ordering rules. That complexity is the price of retrofitting a plugin ecosystem onto a
twenty-year-old reflective framework. Go's `t.Cleanup` covers roughly 90% of what `@AfterEach` is
used for at a fraction of the surface.
### 4.3 Assertions — constraint model
One entry point, composable constraint values (NUnit's model, which avoids the N² overload
explosion):
```el
assert_that(actual, equals(expected))
assert_that(xs, has_length(3))
assert_that(s, contains("foo").and(starts_with("bar")))
assert_that(f, is_within(0.01).of(3.14))
```
A constraint is a value with `apply_to(actual) -> ConstraintResult`, and the result knows how to
describe its own failure. Custom constraints are ordinary user types.
**Every failure message must name file, line, the expression text, and both values.** We capture
expression source text at compile time — we have the AST, so we can do this better than any
runtime-introspection framework.
Legacy `assert_true` / `assert_eq` / etc. stay as thin wrappers for migration.
---
## 5. Benchmarks
### 5.1 The loop
Adopt `b.Loop()`, not `b.N`. Go spent fifteen years on `b.N` before concluding `b.Loop` was right;
we skip that.
```el
bench "str_concat" {
let s = make_input(bench_n())
for bench_loop() {
black_box(str_concat(s, "x"))
}
}
```
Three properties that make this the correct choice for a C target:
1. **The timer auto-resets on first call**, so setup above the loop is excluded *by construction*
rather than by the author remembering `ResetTimer`.
2. **`N` is hidden**, so it cannot be misused.
3. **The harness owns the loop shape**, which lets us insert an optimization barrier the C compiler
cannot see through. `black_box(v)` lowers to `asm volatile("" :: "r"(&v) : "memory")`. Since we
emit a single translation unit, dead-code elimination of a benchmark body is a live hazard —
this is our version of JMH's `Blackhole` problem, solved in the harness rather than delegated to
the user.
### 5.2 Iteration scaling
Use Go's `predictN` heuristics verbatim. They are battle-tested and cheap:
```
n = goal_ns * prev_iters / prev_ns // multiply before divide — precision on sub-ns ops
n += n / 5 // 20% headroom, overshoot rather than re-loop
n = min(n, 100 * last) // never grow more than 100× per step
n = max(n, last + 1) // guarantee forward progress
n = min(n, 1_000_000_000) // hard ceiling
```
Report `n` rounded to 1/2/3/5 × 10ᵏ so runs are comparable.
### 5.3 Sampling
Criterion's shape, because it is correct near timer resolution:
- **Warmup**: iteration counts 1, 2, 4, 8… until cumulative time exceeds the warmup budget.
- **Measurement**: collect `sample_size` samples at iteration counts `[d, 2d, 3d, …, Nd]`.
- **Estimate**: slope of a linear regression of iteration-count vs elapsed time. The intercept
absorbs fixed overhead.
- **Time whole samples, never individual iterations.** This is the single most important detail —
it defeats timer-resolution error on nanosecond operations.
Outliers classified by modified Tukey (±1.5 IQR mild, ±3 IQR severe), **reported but retained**.
---
## 6. Complexity gating — the centerpiece
This is the part that makes the quadratic fixable, and the part nobody in the mainstream has
finished. Google Benchmark's `Complexity()` fits the curve and *reports* it. We declare it and
**gate** on it.
### 6.1 Surface
```el
bench "elc_compile" over n in [16, 32, 64, 128, 256, 512, 1024] expect O(n) {
let src = synth_source(bench_n())
for bench_loop() { black_box(compile(src)) }
}
```
Alternative with no new syntax, if the parser change is judged too invasive — `bench_sizes([...])`
and `bench_expect("O(n)")` as calls inside the block. **Recommendation: declarative.** Runtime calls
mean `--list` cannot show the invariant without executing, which breaks the discovery-precedes-
execution invariant from §2.
### 6.2 Fitting
Per Google Benchmark `src/complexity.cc`. For candidate curves
`{O(1), O(log n), O(n), O(n log n), O(n²), O(n³)}`, one-parameter least squares, no intercept:
```
coef = Σ(tᵢ · gᵢ) / Σ(gᵢ²)
rms = sqrt( Σ(tᵢ coef·gᵢ)² / k ) / mean(t) // normalized
```
Best fit = lowest normalized RMS. User-supplied lambda curves also supported.
### 6.3 Gate logic
1. **FAIL** if the best-fit curve is strictly worse than declared, ordering
`O(1) < O(log n) < O(n) < O(n log n) < O(n²) < O(n³)`. Print the fitted coefficient and the full
per-size table.
2. **FAIL** if the declared curve's normalized RMS exceeds a threshold (start at 0.10). This catches
the case where *no* candidate fits — noise, a cache cliff, or a phase change. Report
`INDETERMINATE` honestly rather than gating on garbage.
3. **WARN** if the best fit is strictly better than declared — either an optimization landed and the
annotation should tighten, or the sweep is too narrow to expose real behaviour.
4. **REFUSE to gate** on fewer than 5 distinct sizes spanning under 2 decades, geometrically spaced.
Say so loudly rather than producing a meaningless fit.
### 6.4 Why gate on the exponent, not wall-clock
- **Machine-independent.** The fitted exponent is a property of the algorithm; the coefficient is a
property of the machine. Gating on the exponent makes CI hardware heterogeneity, noisy neighbours,
and thermal throttling irrelevant — they scale `coef`, not `g`.
- **No stored baseline.** No artifact storage, no golden-file drift. The invariant lives in the
source next to the code and is reviewed in the same PR.
- **It catches the failure mode that actually ships.** An O(n) lookup inside an O(n) loop is
invisible at n=100 in a unit test and catastrophic at n=100,000 in production. Constant-factor
regressions are annoying. Complexity regressions are outages. Ours was a 27 GB outage.
### 6.5 The deterministic gate — the one that would have caught us
Wall-clock needs statistics. **Allocation counts do not.** They are perfectly deterministic.
> **Correction, 2026-08-16 — count alone is NOT sufficient. Gate on BOTH count and bytes.**
>
> Measured against two El programs, one allocating once per item and one rebuilding its
> accumulator each iteration:
>
> | n | linear allocs / bytes | quadratic allocs / bytes |
> |---|---|---|
> | 100 | 100 / 290 | 100 / 5,150 |
> | 200 | 200 / 690 | 200 / 20,300 |
> | 400 | 400 / 1,490 | 400 / 80,600 |
> | 800 | 800 / 3,090 | 800 / 321,200 |
>
> The quadratic program's allocation **count is exactly linear** — 100/200/400/800, identical to
> the healthy program. A count-only gate passes it clean. **Bytes** catch it: each doubling of n
> quadruples bytes (ratios 3.94, 3.97, 3.99 → 4.0 = O(n²)) where the linear program converges
> on 2.0.
>
> This is precisely elc's own defect shape — a copy-on-write accumulator reallocating once per
> pass (count linear) into a proportionally larger buffer (bytes quadratic).
>
> Therefore `expect allocs O(n)` **fits count and bytes independently and fails if EITHER exceeds
> the declared curve**, reporting which signal broke. "count linear, bytes quadratic" is a precise,
> directly actionable diagnosis.
>
> **`el_peak_rss()` is CONTEXT ONLY — never gate on it.** It is perturbed by the allocator and by
> the page cache. Allocation volume is the invariant; RSS and malloc/free churn are merely the two
> surfaces it shows on. The old shipped compiler paid the same quadratic in RSS that the rebuilt
> one pays in churn.
>
> **Measure rate, not level.** A guard reading swap *level* saw 97% on a thrashing host and 97% on
> a healthy one; only *rate* separated them. A growth exponent is a rate; a single measurement is
> a level. That is why the gate fits a curve across a sweep instead of comparing one number to a
> threshold.
> **Second correction, same day — THE ALLOCATION GATE ALONE WOULD HAVE MISSED THE REAL BUG.**
>
> el #132 found the actual elc quadratic: `strlen()` called inside `str_char_code()` and
> `str_slice()`, so the lexer rescanned the remaining input on every character. Pure CPU.
> **Zero allocation.** `str_char_code` is a bounds check and an index — it allocates nothing.
>
> Measured on three controlled specimens (`lang/.work/fitprobe.el`), growth ratio per doubling of
> n across n = 200/400/800/1600:
>
> | specimen | allocs | bytes | time | what it proves |
> |---|---|---|---|---|
> | `linear` — one alloc per item | 2.00 2.00 2.00 → **O(n)** | 2.16 2.07 2.23 → **O(n)** | 0.83 2.00 2.05 → **O(n)** | clean baseline |
> | `accum` — rebuilds accumulator | 2.00 2.00 2.00 → **O(n)** | 3.97 3.99 3.99 → **O(n²)** | noisy | count misses, **bytes catches** |
> | `compute` — n scans over n chars | 0 → **FLAT** | 0 → **FLAT** | 3.93 4.01 3.96 → **O(n²)** | **both alloc signals blind; only time catches** |
>
> `compute` is el #132's shape exactly. A gate fitting only allocation count and bytes classifies
> it as FLAT and passes it. **The gate as originally specified would not have caught the defect it
> was created for.**
>
> Therefore the gate fits **THREE** signals and fails if ANY exceeds its declared curve:
>
> ```
> bench "elc_compile" over n in [...] expect time O(n) allocs O(n) bytes O(n) { ... }
> ```
>
> - **allocs (count)** — deterministic, zero-noise. Catches per-item allocation growth.
> - **allocs (bytes)** — deterministic, zero-noise. Catches accumulator-rebuild quadratics that
> count cannot see.
> - **time** — noisy, needs the sweep and statistics. The ONLY signal that sees pure-compute
> complexity regressions. Gate on the fitted *exponent*, never on absolute duration, so CI
> hardware variance scales the coefficient and leaves the classification intact.
>
> The deterministic signals remain preferable where they apply — they need no statistics and are
> correct on the first run. They are simply not sufficient.
>
> **`black_box` is mandatory, and consuming the result is NOT enough.** The first version of
> `compute` accumulated `total + 1` in a nested loop and reported **0 µs at every n** while
> returning a numerically correct n². Clang recognised the idiom and closed the loop to a
> multiply. Feeding the result into output did not prevent it. Only making the inner operation an
> opaque external call restored the real curve. A benchmark harness that trusts the user to defeat
> the optimiser will silently measure nothing — and report success while doing it.
Instrument the runtime with allocation counters and fit *those* against n instead of time:
```el
bench "elc_compile" over n in [...] expect O(n) allocs O(n) { ... }
```
Zero noise, zero statistics, always gateable, correct on the first run on any machine. Go reports
`allocs/op` and `B/op`; **nobody fits them against n.** That is an open opportunity and it is exactly
our bug: elc's defect is quadratic *allocation volume*, which the old binary paid in RSS and the
current source pays in malloc/free churn.
An `expect allocs O(n)` assertion on `elc`'s compile path would have failed the build the day the
quadratic was introduced.
Required runtime additions: `__el_alloc_count()`, `__el_alloc_bytes()`, `__el_peak_rss()`.
### 6.6 Constant-factor gate (secondary, opt-in)
Mann-Whitney U at α = 0.05, noise floor 1%, medians with 95% CIs, `~` for not-significant. Requires
`--count >= 9`. Off by default on CI; opt-in per benchmark.
**Exit nonzero on regression.** Both benchstat and Criterion always exit 0, which is why every shop
using them wrote a wrapper. We do not repeat that omission.
---
## 7. Output
**Structured events are the source of truth.** Human text is rendered from them. We do not repeat
Go's parse-the-human-output design.
Event stream, NDJSON, one object per line, streamed live:
```json
{"time":"...","action":"run","test":"parser/empty"}
{"time":"...","action":"output","test":"parser/empty","output":"..."}
{"time":"...","action":"pass","test":"parser/empty","elapsed":0.0031}
{"time":"...","action":"bench","test":"str_concat","n":1024,"ns_op":41.2,"allocs_op":3,"bigo":"N","rms":0.03}
```
Renderers, all downstream and pluggable:
| Format | Flag | Use |
|---|---|---|
| Human | default | terminal, **per-test duration always shown** |
| NDJSON | `--json` | tooling, history, flaky detection |
| JUnit XML | `--junit-xml=PATH` | every CI system on earth |
| TAP | `--tap` | optional |
JUnit XML per the de-facto schema: `testsuites``testsuite``testcase`, with `time` in seconds
as a decimal, `file`/`line` attributes, and `failure` vs `error` vs `skipped` as distinct child
elements. Absence of a child element means pass. Emit `<testsuites>` even for a single suite, and
parse both shapes on input.
---
## 8. CLI
```
--list print the registry, run nothing
--list-json machine-readable registry
--run PATTERN slash-separated regex per path segment
--tag EXPR tag expression: fast & !slow
--shard I/N deterministic sharding for CI parallelism
--count N repetitions, for statistics
--bench PATTERN run benchmarks (off by default in test runs)
--benchtime DUR per-benchmark time budget
--junit-xml PATH
--json
--isolate re-exec per test on crash, so one SIGSEGV doesn't lose the run
--timeout DUR
--fail-fast
```
`--list` / `--list-json` / `--shard` cost roughly thirty lines because the registry already exists
before `main` does anything. That is the dividend of discovery-precedes-execution.
---
## 9. Build model
```
# once, ever (or when the runtime/framework changes):
cc -c el_runtime.c -o el_runtime.o
elc eltest.el > eltest.c && cc -c eltest.c -o eltest.o
ar rcs libeltest.a el_runtime.o eltest.o
# per suite:
elc --test foo_test.el > foo_test.c # registry + bodies only
cc foo_test.c libeltest.a -o foo_test
```
The 0.14s × N of redundant runtime rebuilds disappears — not because we optimized it, but because
one-runner-over-many-suites requires compile-once-link-many as a structural precondition.
---
## 10. Bootstrap and self-hosting
The framework's own tests are `test { }` blocks run by the framework. Same fixpoint discipline the
compiler already applies to itself.
1. Build the framework using the *existing* harness for its first tests (stage 0).
2. Rebuild the framework's tests as `test { }` blocks run by the new runner (stage 1).
3. Verify stage 1 reports identical results to stage 0.
4. From then on, the framework is tested by itself.
A framework that cannot run its own suite is not evidence of anything. This is a correctness proof,
not a claim.
---
## 11. Explicitly not building
| Rejected | Why |
|---|---|
| Naming-convention discovery (`fn test_foo`) | `test { }` is a real declaration. Go's `TestXxx` exists only because Go had no better hook — and it needs a heuristic to avoid matching `TesticularCancer`. |
| Reflection or symbol-table scanning | Slow, fragile under LTO/strip/dead-strip, and unnecessary when we own the compiler. |
| Parsing human output into structure | Go's `test2json` is its one clear architectural mistake. |
| JUnit 5's extension SPI | Seventeen callback interfaces to retrofit plugins onto a reflective framework. Not our problem. |
| `@ParameterizedTest` machinery | Table-driven loops + subtests subsume it at zero surface. |
| NUnit's out-of-process agents | They bridge CLR versions and AppDomains. We emit one native binary. Keep `--isolate` as crash fallback only. |
| JMH-style forking by default | Forks exist because JIT profiles are per-process. AOT C has no such state. Keep `--fork` available, not default. |
| Exit 0 on regression | benchstat and Criterion both do this, and every user writes a wrapper. |
| Dynamic runtime test registration | Breaks `--list`, sharding, and individual selection. Registry stays static. |
---
## 12. Phasing
| Phase | Content | Gate |
|---|---|---|
| **1** | Registry emission in codegen; 9 builtins; `el_test_main` skeleton in El; result records; per-test timing; human + NDJSON output | existing 11 test files pass, with timing |
| **2** | `libeltest.a` build model; subtests; filtering; `--list`; fixtures; constraint assertions; JUnit XML | suite runs in one binary; runtime compiled once |
| **3** | `bench { }`, `bench_loop`, `black_box`, `predictN`, Criterion sampling | benchmarks produce stable ns/op |
| **4** | Allocation counters; complexity fitting; `expect O(...)` gate | **an `expect allocs O(n)` benchmark on `elc` fails on the current quadratic** |
| **5** | Migrate both legacy systems; delete `runtime/test.el`; self-host | framework runs its own suite |
Phase 4 is the deliverable that matters. Phases 13 exist to make it possible.
---
## 13. Open questions for review
1. **Declarative `over n in [...] expect O(...)` syntax vs runtime calls.** I recommend declarative
(§6.1) so `--list` can show invariants without executing. It costs parser work. Your call.
2. **`bench { }` as a new block form** — parallel to `test { }`, or a modifier on it?
3. **Scope of the constraint model.** Full composable constraints, or start with a flat assertion set
and add constraints later? Full model is more surface but avoids a second migration.
4. **Does `runtime/test.el` get deleted or kept as a deprecated shim?** I lean delete — two systems
is how we got here.
5. **Where does `libeltest.a` live** in the tree, and does `epm` need to know about it?
6. **Allocation counters in `el_seed.c` or `el_runtime.c`?** AGENTS.md says `el_seed.c` is the sole
C dependency and hand-maintained; counters are OS-boundary-adjacent but not OS calls.
7. **Is per-test timing enough, or do we want per-*assertion* timing** for finding slow helpers?
---
## 14. What this document is not
This is a design, not a measurement. Every performance claim about the *current* system in §1 is
measured and reproducible in this worktree. Every claim about the *proposed* system is a prediction.
None of it is verified until Phase 1 runs and Phase 4 fails a build on the real quadratic.
+214 -43
View File
@@ -10,10 +10,60 @@
// cc -std=c11 -O2 -lcurl -lpthread -o engram server.c el_runtime.c
// ./engram
//
// Configuration via environment:
// ENGRAM_BIND host:port (default :8742)
// ENGRAM_API_KEY bearer auth (optional)
// ENGRAM_DATA_DIR snapshot location (default ~/.neuron/engram)
// Configuration is DECLARED, not scattered. See the `program` block below:
// every knob's type and default lives there and nowhere else, is resolved from
// the environment (env wins, declaration is the fallback) and validated before
// any statement of this file runs. Read one with config("NAME") -> String.
//
// The one deliberate exception is ENGRAM_DATA_DIR see the note in the block.
// Program declaration (cross-cutting concerns)
//
// singleton: two engram processes against one data dir is data loss, not a
// warning. The runtime takes an exclusive flock at startup and a second start
// is refused loudly with the holder's pid.
//
// NOT declared here, on purpose: ENGRAM_DATA_DIR. Its resolution is owned by
// engram_resolve_data_dir() (el_runtime.c), which defaults to $HOME/.neuron/engram
// and fails LOUD rather than silently persisting to an ephemeral directory.
// Declaring a default for it here as well would put the data dir's fallback in
// two places which is precisely the defect this migration removes (until
// 2026-08-15 the reseed backup path carried its own "/tmp/engram" default that
// disagreed with the resolver, so the pre-destructive safety copy landed in /tmp).
// HOME is likewise not declared: it is a genuine environment read, not a knob.
program "engram" {
singleton: "engram"
// Core server
env ENGRAM_BIND: String = ":8742"
// Default "" leaves auth DISABLED (check_auth_ok short-circuits to true on an
// empty key). That is the pre-existing behaviour and is deliberately preserved
// here; making this `required` is the obvious hardening follow-up, but it is a
// behaviour change and out of scope for this migration.
env ENGRAM_API_KEY: String = ""
// Feature flags (bool-ish Strings; the predicate fns below own truthiness) ──
env ENGRAM_STORE: String = "off"
env ENGRAM_WAL: String = "off"
env ENGRAM_AUTOCONNECT: String = "off"
env ENGRAM_ISE_OFFGRAPH: String = "off"
// ISE telemetry
env ENGRAM_ISE_RETENTION_MS: Int = "172800000"
// Guide (local Qwen3 via llama-server)
env GUIDE_ENABLE: String = "off"
env GUIDE_TIER_FORCE: String = ""
env GUIDE_CACHE_DIR: String = ""
env GUIDE_RAM_GB_4B: Int = "16"
env GUIDE_RAM_GB_1P7B: Int = "8"
env GUIDE_BACKEND: String = "llama-server"
env GUIDE_HOST: String = "127.0.0.1"
env GUIDE_PORT: Int = "8771"
env GUIDE_LLAMA_SERVER_BIN: String = "llama-server"
env GUIDE_NGL: Int = "99"
env GUIDE_CTX: Int = "4096"
}
// Helpers
@@ -41,17 +91,29 @@ fn strip_query(path: String) -> String {
str_slice(path, 0, q)
}
// query_param extract one query-string value, URL-DECODED.
//
// The decode step was missing (found 2026-08-15): a claim sent as
// "test%20claim" arrived at engram_assert_json still percent-encoded and was
// stored/compared that way, so any value containing a space, &, =, or non-ASCII
// character silently became a different string than the caller sent. Affects
// every GET route that reads params this way, not just /api/assert.
fn query_param(path: String, key: String) -> String {
let q: Int = str_index_of(path, "?")
if q < 0 { return "" }
let qs: String = str_slice(path, q + 1, str_len(path))
let needle: String = key + "="
let pos: Int = str_index_of(qs, needle)
// Anchor the match to a real key boundary: prefixing "&" and searching for
// "&key=" means "q" can never match inside "faq=". (Found 2026-08-15:
// "?faq=X&q=Y" returned X for key "q" a silently wrong value, not an
// error.) The leading "&" makes the first parameter match the same way.
let hay: String = "&" + qs
let needle: String = "&" + key + "="
let pos: Int = str_index_of(hay, needle)
if pos < 0 { return "" }
let after: String = str_slice(qs, pos + str_len(needle), str_len(qs))
let after: String = str_slice(hay, pos + str_len(needle), str_len(hay))
let amp: Int = str_index_of(after, "&")
if amp < 0 { return after }
str_slice(after, 0, amp)
let raw: String = if amp < 0 { after } else { str_slice(after, 0, amp) }
return __url_decode(raw)
}
fn query_int(path: String, key: String, default_val: Int) -> Int {
@@ -121,7 +183,7 @@ fn route_text_health(method: String, path: String, body: String) -> String {
// engram_store_enabled() in el_runtime.c EXACTLY (1 / on / true). Default off
// every persistence path below is byte-for-byte the historical snapshot behavior.
fn store_on() -> Bool {
let v: String = env("ENGRAM_STORE")
let v: String = config("ENGRAM_STORE")
if str_eq(v, "1") { return true }
if str_eq(v, "on") { return true }
if str_eq(v, "true") { return true }
@@ -150,7 +212,6 @@ fn persist_canonical() -> Int {
if store_on() {
return engram_store_checkpoint()
}
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = engram_resolve_data_dir()
// (2026-08-10 self-review) This returned a hardcoded 1, which made every
// caller's `let saved: Int = persist_canonical()` a dead variable six
@@ -164,7 +225,7 @@ fn persist_canonical() -> Int {
// per-write full-snapshot behavior. When ON, structural mutations append O(1)
// WAL records instead of rewriting the whole graph, with threshold compaction.
fn wal_on() -> Bool {
str_eq(env("ENGRAM_WAL"), "on")
str_eq(config("ENGRAM_WAL"), "on")
}
// autoconnect_on ENGRAM_AUTOCONNECT. Will's rule: "we shouldn't be inserting
@@ -172,7 +233,7 @@ fn wal_on() -> Bool {
// edge (kNN over embeddings) so no content node enters the graph edgeless.
// Default OFF -> byte-identical to prior behavior (node created, no auto edges).
fn autoconnect_on() -> Bool {
let v: String = env("ENGRAM_AUTOCONNECT")
let v: String = config("ENGRAM_AUTOCONNECT")
if str_eq(v, "1") { return true }
if str_eq(v, "on") { return true }
if str_eq(v, "true") { return true }
@@ -185,7 +246,7 @@ fn autoconnect_on() -> Bool {
// separate state-event log tier instead of the node graph. Default OFF -> ISEs
// remain graph nodes exactly as before (with 48h prune).
fn ise_offgraph_on() -> Bool {
let v: String = env("ENGRAM_ISE_OFFGRAPH")
let v: String = config("ENGRAM_ISE_OFFGRAPH")
if str_eq(v, "1") { return true }
if str_eq(v, "on") { return true }
if str_eq(v, "true") { return true }
@@ -235,6 +296,24 @@ fn persist_bulk() -> Int {
return persist_canonical()
}
// COMPILER LANDMINE, measured 2026-08-16 do not inline this back into the
// caller. elc lowers `a == b` to numeric comparison only when both operand
// NAMES are in the per-function int-name set, which `let x: Int` populates.
// That registration does NOT propagate into a nested if-expression block: the
// first cut of the geometry-ingest path wrote `let claimed: Int = ...` and
// `let got: Int = ...` inside the else-arm and `claimed == got` came out of
// codegen as `str_eq(claimed, got)` strcmp on two integers reinterpreted as
// pointers, i.e. a segfault on the first geometry-bearing request. Read back
// out of the generated C, not guessed. Function PARAMETERS annotated `: Int`
// do register reliably (verified: `if (claimed == actual)`), so the comparison
// lives in a function of its own. Note also the explicit `return`s a trailing
// if-EXPRESSION at a function tail emits as a statement and the function
// returns 0 regardless, which is the same probe's second finding.
fn width_agrees(claimed: Int, actual: Int) -> Int {
if claimed == actual { return 1 }
return 0
}
// INCOMPLETE-ROUTE FIX (2026-07-24 self-review): this route silently dropped
// label, importance, tier, and tags engram_node() defaults label to content
// and importance to 0.5, so every node created over HTTP lost its metadata.
@@ -276,6 +355,45 @@ fn route_create_node(method: String, path: String, body: String) -> String {
salience, importance, confidence,
tier, tags
)
// GEOMETRY INGEST geometry-valued end to end (2026-08-16).
//
// The defect this route originally had: it accepted an "emb" field,
// returned 200 with a fresh id, and stored NOTHING, because engram_node_full
// has no vector parameter. The consequence was structural, not cosmetic
// text was the only entry medium, so any non-text modality had to be
// DESCRIBED in prose, and what we then reasoned over was the geometry of the
// description, not of the signal.
//
// #141 fixed the drop but marshalled the vector as a hex STRING through
// engram_node_set_emb, which put text back as the TRANSPORT medium one layer
// below the problem being fixed. This is that correction: hex is decoded
// exactly ONCE, here at the edge, into a first-class Geometry, and every
// step below this line moves geometry rather than text. An encoding at the
// boundary is what an encoding is for.
//
// The WIRE is deliberately unchanged "emb" is still little-endian float32
// hex (8 chars per component), the encoding the perception vessel's
// /voice/embed already emits because production clients speak it. What
// changed is underneath it.
//
// "dim" is now treated as an ASSERTION about the vector the caller sent, not
// as the source of its width: a Geometry carries its own width. A stated dim
// that disagrees is a REJECTED ingest, not a silent reinterpretation. Omitting
// "dim" is fine and means "trust the vector", which is the honest default.
//
// Off-dimension vectors remain stored but not inserted into the resident HNSW
// index (its build loop filters on emb_dim), so a 64-dim voice geometry is
// durable and addressable without perturbing the 768-dim canonical index.
let emb_hex: String = json_get_string(body, "emb")
let emb_set: Int = if str_eq(emb_hex, "") { 0 } else {
let g: Geometry = geometry_from_f32le_hex(emb_hex)
let got: Int = geometry_dim(g)
let dim_raw: String = json_get_raw(body, "dim")
let claimed: Int = if str_eq(dim_raw, "") { got } else { json_get_int(body, "dim") }
let landed: Int = if width_agrees(claimed, got) > 0 { node_attach_geometry(id, g) } else { 0 }
let freed: Int = geometry_free(g)
landed
}
let saved: Int = persist_node(id)
// ORPHAN PREVENTION (ENGRAM_AUTOCONNECT): connect the fresh node to its
// nearest embedded neighbors so it never enters the graph edgeless.
@@ -286,7 +404,11 @@ fn route_create_node(method: String, path: String, body: String) -> String {
if added > 0 { let sv2: Int = persist_edges_since(ec0) }
added
} else { 0 }
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\",\"connected\":" + int_to_str(connected) + "}"
// Report whether the supplied geometry actually landed. The old response
// was success-shaped regardless 200 with an id while the vector was
// discarded which is how the drop went unnoticed. A caller can now
// assert on emb_set instead of trusting the status code.
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\",\"connected\":" + int_to_str(connected) + ",\"emb_set\":" + int_to_str(emb_set) + "}"
}
fn route_get_node(method: String, path: String, body: String) -> String {
@@ -321,7 +443,6 @@ fn route_scan_nodes(method: String, path: String, body: String) -> String {
// process ever booted with a partial/empty store, the first read request
// clobbered the good snapshot. Read routes must never write the canonical path.)
fn route_scan_edges(method: String, path: String, body: String) -> String {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = engram_resolve_data_dir()
let snap_path: String = dir + "/.scan-export.json"
engram_save(snap_path)
@@ -482,7 +603,6 @@ fn route_forget(method: String, path: String, body: String) -> String {
fn route_save(method: String, path: String, body: String) -> String {
let p_raw: String = json_get_string(body, "path")
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = engram_resolve_data_dir()
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
// (2026-08-10 self-review) engram_save returns 0 on an empty path and the
@@ -566,7 +686,6 @@ fn route_drift(method: String, path: String, body: String) -> String {
fn route_load(method: String, path: String, body: String) -> String {
let p_raw: String = json_get_string(body, "path")
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = engram_resolve_data_dir()
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
// (2026-08-10 self-review) This was a stub response over the single most
@@ -637,7 +756,6 @@ fn route_embed_backfill(method: String, path: String, body: String) -> String {
// (it skips nodes already present by ID). Auth-exempt: same-host internal call.
// (2026-06-27 self-review: added this route to fix silent 10-min sync failures)
fn route_sync(method: String, path: String, body: String) -> String {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = engram_resolve_data_dir()
// 2026-07-21 self-review: export to a scratch path, never the canonical
// snapshot.json read routes must not be able to clobber the good snapshot.
@@ -713,8 +831,12 @@ fn route_reseed_nodes(method: String, path: String, body: String) -> String {
if str_eq(p, "") { return err_json("path is required") }
if str_eq(fs_read(p), "") { return err_json("file missing or empty") }
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
// (2026-08-15) This site carried its own "/tmp/engram" fallback, which
// DISAGREED with engram_resolve_data_dir() ($HOME/.neuron/engram, fail-loud).
// The consumer is the pre-destructive backup below, so with ENGRAM_DATA_DIR
// unset the safety copy taken before a reseed landed in an ephemeral /tmp
// while the store it was protecting lived elsewhere. One owner, one answer.
let dir: String = engram_resolve_data_dir()
let backup: String = dir + "/.reseed-backup.json"
let replace_raw: String = json_get_raw(body, "replace")
@@ -806,8 +928,7 @@ fn route_emit_ise(method: String, path: String, body: String) -> String {
sal, imp, conf,
"Episodic", "[\"internal-state\",\"InternalStateEvent\"]"
)
let ret_raw: String = env("ENGRAM_ISE_RETENTION_MS")
let ret_ms: Int = if str_eq(ret_raw, "") { 172800000 } else { str_to_int(ret_raw) }
let ret_ms: Int = str_to_int(config("ENGRAM_ISE_RETENTION_MS"))
let pruned: Int = engram_prune_telemetry(ret_ms)
"{\"ok\":true,\"id\":\"" + id + "\",\"pruned\":" + int_to_str(pruned) + "}"
}
@@ -904,6 +1025,22 @@ fn route_similarity(method: String, path: String, body: String) -> String {
// nothing on request. NOTE: the offline reify WRITER (engram_geo_reify_store) is
// currently unwired, so on the live store the resident index is empty and the
// list returns [] until reification runs see the cutover report.
// route_scan_emb GET /api/nodes/emb?limit=&offset= read the raw geometry.
//
// engram_scan_nodes_emb_json has existed as a builtin with NO ROUTE, so the
// embeddings the actual positions every distance, angle, membership and
// grounding is computed from were unreadable from outside the process. You
// cannot verify a coordinate system you cannot see, and every claim about the
// frame (isotropy, centering, what the origin is) was therefore unfalsifiable
// from the API. Read-only.
fn route_scan_emb(method: String, path: String, body: String) -> String {
let l_raw: String = query_param(path, "limit")
let o_raw: String = query_param(path, "offset")
let l: Int = if str_eq(l_raw, "") { 200 } else { str_to_int(l_raw) }
let o: Int = if str_eq(o_raw, "") { 0 } else { str_to_int(o_raw) }
return engram_scan_nodes_emb_json(l, o)
}
fn route_neighborhoods(method: String, path: String, body: String) -> String {
engram_geo_reify_list_json()
}
@@ -997,6 +1134,11 @@ fn route_faculty(path: String, faculty: String) -> String {
fn route_boundary_proof(method: String, path: String, body: String) -> String {
return "{\"op\":\"boundary_proof\",\"body_instrumentation\":\"none\",\"seam\":\"@manager -> engram_boundary_beat auto-injected\"}"
}
// GROUNDING: an attribute of the RELATION, and the relation's weight is a
// VECTOR (factual, relational, associative, polarity, provenance, timestamp).
// /api/ground READS it it never writes. /api/ground/record is the write,
// named as one, and it consolidates only on a consequential + salient move.
// /api/ground/trajectory reads the supersession chain as a time series.
fn route_ground(method: String, path: String, body: String) -> String {
let claim: String = json_get_string(body, "claim")
let evidence: String = json_get_string(body, "evidence")
@@ -1005,12 +1147,31 @@ fn route_ground(method: String, path: String, body: String) -> String {
if str_eq(evidence, "") { return err_json("missing evidence") }
return engram_ground_json(claim, evidence, for_whom)
}
fn route_ground_record(method: String, path: String, body: String) -> String {
let claim: String = json_get_string(body, "claim")
let evidence: String = json_get_string(body, "evidence")
let provenance: String = json_get_string(body, "provenance")
let floor: String = json_get_string(body, "floor")
if str_eq(claim, "") { return err_json("missing claim") }
if str_eq(evidence, "") { return err_json("missing evidence") }
return engram_ground_record_json(claim, evidence, provenance, floor)
}
fn route_ground_trajectory(method: String, path: String, body: String) -> String {
let claim: String = query_param(path, "claim")
let evidence: String = query_param(path, "evidence")
if str_eq(claim, "") { return err_json("missing claim") }
if str_eq(evidence, "") { return err_json("missing evidence") }
return engram_ground_trajectory_json(claim, evidence)
}
fn route_assert(method: String, path: String, body: String) -> String {
let claim: String = query_param(path, "claim")
if str_eq(claim, "") { return err_json("missing claim") }
let for_whom: String = query_param(path, "for_whom")
let floor: String = query_param(path, "floor")
return engram_assert_json(claim, for_whom, floor)
// Both floors. A well-evidenced claim does not earn the right to be asserted
// regardless of whether it means the right thing. rel_floor defaults to floor.
let rel_floor: String = query_param(path, "rel_floor")
return engram_assert_json(claim, for_whom, floor, rel_floor)
}
fn route_attend(method: String, path: String, body: String) -> String {
let node: String = json_get_string(body, "node")
@@ -1056,14 +1217,12 @@ fn route_correspondence_beat(method: String, path: String, body: String) -> Stri
// turns native thinking ON: the response carries reasoning_content (the thinking)
// alongside content (the answer).
fn guide_env_or(key: String, dflt: String) -> String {
let v: String = env(key)
if str_eq(v, "") { return dflt }
return v
}
// (2026-08-15) guide_env_or(key, dflt) lived here. Its whole job was supplying a
// per-call-site default, which is now the program block's job every GUIDE_* knob
// is declared once at the top of this file and read straight through config().
fn guide_enabled() -> Bool {
let v: String = env("GUIDE_ENABLE")
let v: String = config("GUIDE_ENABLE")
if str_eq(v, "1") { return true }
if str_eq(v, "on") { return true }
if str_eq(v, "true") { return true }
@@ -1108,15 +1267,15 @@ fn guide_probe_metal() -> Bool {
// 2. Tier selection (config-driven thresholds, spec-autoselected)
fn guide_threshold_4b() -> Int {
return str_to_int(guide_env_or("GUIDE_RAM_GB_4B", "16"))
return str_to_int(config("GUIDE_RAM_GB_4B"))
}
fn guide_threshold_1p7b() -> Int {
return str_to_int(guide_env_or("GUIDE_RAM_GB_1P7B", "8"))
return str_to_int(config("GUIDE_RAM_GB_1P7B"))
}
// GUIDE_TIER_FORCE overrides the spec autoselect (used to prove cheaply on 0.6b).
fn guide_select_tier(ram_gb: Int) -> String {
let forced: String = env("GUIDE_TIER_FORCE")
let forced: String = config("GUIDE_TIER_FORCE")
if !str_eq(forced, "") { return forced }
if ram_gb >= guide_threshold_4b() { return "4b" }
if ram_gb >= guide_threshold_1p7b() { return "1.7b" }
@@ -1136,8 +1295,10 @@ fn guide_file(tier: String) -> String {
}
fn guide_cache_dir() -> String {
let c: String = env("GUIDE_CACHE_DIR")
let c: String = config("GUIDE_CACHE_DIR")
if !str_eq(c, "") { return c }
// HOME stays a raw env() read: it is the ambient environment, not a knob of
// this program, and it is deliberately absent from the program block.
let home: String = env("HOME")
if !str_eq(home, "") { return home + "/.neuron/guide/models" }
return engram_resolve_data_dir() + "/guide-models"
@@ -1178,9 +1339,9 @@ fn guide_fetch(tier: String) -> Bool {
}
// 4/5. Backend abstraction + BIND as an engageable interlocutor
fn guide_backend() -> String { return guide_env_or("GUIDE_BACKEND", "llama-server") }
fn guide_host() -> String { return guide_env_or("GUIDE_HOST", "127.0.0.1") }
fn guide_port() -> String { return guide_env_or("GUIDE_PORT", "8771") }
fn guide_backend() -> String { return config("GUIDE_BACKEND") }
fn guide_host() -> String { return config("GUIDE_HOST") }
fn guide_port() -> String { return config("GUIDE_PORT") }
fn guide_base_url() -> String { return "http://" + guide_host() + ":" + guide_port() }
// guide_healthy is the guide present and answering? llama-server's /health
@@ -1198,9 +1359,9 @@ fn guide_healthy() -> Bool {
fn guide_load(tier: String) -> Bool {
if guide_healthy() { return true }
let path: String = guide_model_path(tier)
let bin: String = guide_env_or("GUIDE_LLAMA_SERVER_BIN", "llama-server")
let ngl: String = guide_env_or("GUIDE_NGL", "99")
let ctx: String = guide_env_or("GUIDE_CTX", "4096")
let bin: String = config("GUIDE_LLAMA_SERVER_BIN")
let ngl: String = config("GUIDE_NGL")
let ctx: String = config("GUIDE_CTX")
let logf: String = guide_cache_dir() + "/llama-server." + guide_port() + ".log"
let cmd: String = bin + " -m '" + path + "' --host " + guide_host() + " --port " + guide_port() + " -c " + ctx + " -ngl " + ngl + " --jinja >> '" + logf + "' 2>&1"
let pid: String = exec_bg(cmd)
@@ -1596,7 +1757,7 @@ fn route_supersede(method: String, path: String, body: String) -> String {
// Auth
fn check_auth_ok(method: String, body: String) -> Bool {
let key: String = env("ENGRAM_API_KEY")
let key: String = config("ENGRAM_API_KEY")
if str_eq(key, "") { return true }
// Read-only methods don't require auth. Until http_serve surfaces
// request headers we can't accept a Bearer token cleanly; mutating
@@ -1675,6 +1836,9 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "GET") && (str_eq(clean, "/api/edges") || str_eq(clean, "/edges")) {
return route_scan_edges(method, path, body)
}
if str_eq(method, "GET") && (str_eq(clean, "/api/nodes/emb") || str_eq(clean, "/nodes/emb")) {
return route_scan_emb(method, path, body)
}
if str_eq(method, "GET") && str_starts_with(clean, "/api/nodes/") {
return route_get_node(method, path, body)
}
@@ -1764,6 +1928,14 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "GET") && str_starts_with(clean, "/api/plan") {
return route_faculty(path, "plan")
}
// Order matters: the more specific paths must be tested before the /api/ground
// prefix match below, which would otherwise swallow them.
if str_eq(method, "POST") && str_starts_with(clean, "/api/ground/record") {
return route_ground_record(method, path, body)
}
if str_eq(method, "GET") && str_starts_with(clean, "/api/ground/trajectory") {
return route_ground_trajectory(method, path, body)
}
if str_eq(method, "POST") && str_starts_with(clean, "/api/ground") {
return route_ground(method, path, body)
}
@@ -1859,8 +2031,7 @@ fn handle_request(method: String, path: String, body: String) -> String {
// Entry
let bind_raw: String = env("ENGRAM_BIND")
let bind_str: String = if str_eq(bind_raw, "") { ":8742" } else { bind_raw }
let bind_str: String = config("ENGRAM_BIND")
let port: Int = parse_port(bind_str)
// On startup, try to load any existing snapshot (best effort).
+40
View File
@@ -0,0 +1,40 @@
#!/bin/sh
# Build + RUN the §7 GROUNDING-VECTOR tests (engram_cognition.c): the one decay
# model, the consequence gate, and the stored/derived split. Closed-form
# constructed cases — no server, no store, no network. Pure C11 (stdlib + libm).
# Standalone — NOT folded through elc. Two passes:
# 1. PERF — optimised (-O2, no sanitizer): the functional gate.
# 2. SAFETY — ASan + UBSan on the same suite.
#
# NEGATIVE CONTROL (invariant §8.6 — no test without one). Every symbol this
# suite exercises (cog_decay_factor, cog_grounding_significant,
# cog_significance_inherent, CogGrounding, CogProvClass) is introduced by the
# change under test, so the suite does not COMPILE against the pre-change source.
# To reproduce:
# git show origin/dev:lang/runtime/engram_cognition.h > /tmp/pre/engram_cognition.h
# git show origin/dev:lang/runtime/engram_cognition.c > /tmp/pre/engram_cognition.c
# cc -I/tmp/pre engram/test/test_grounding_vector.c /tmp/pre/engram_cognition.c ...
# => error: unknown type name 'CogGrounding'; no binary produced.
set -e
HERE=$(cd "$(dirname "$0")" && pwd)
RT="$HERE/../../lang/runtime"
CC=${CC:-cc}
SRC="$HERE/test_grounding_vector.c $RT/engram_cognition.c $RT/engram_reason.c $RT/engram_geometry.c $RT/engram_store.c $RT/engram_vindex.c"
WARN="-std=c11 -Wall -Wextra"
# engram_store.c declares emit_log as a WEAK symbol and null-checks it, which is
# how a test links the store without the EL runtime. Darwin's ld does not resolve
# an undefined weak symbol at static-link time, so it must be allowed explicitly.
# (The pre-existing runners in this directory — run_verify_tests.sh among them —
# do not do this and therefore fail to link on macOS. Unrelated to this change.)
LDX=""
[ "$(uname -s)" = "Darwin" ] && LDX="-Wl,-U,_emit_log"
TMP=$(mktemp -d)
echo "### PASS 1: PERF (optimised, un-sanitised) — functional gate"
$CC $WARN -O2 -I"$RT" $SRC -lm -lpthread $LDX -o "$TMP/perf"
"$TMP/perf"
echo
echo "### PASS 2: SAFETY (ASan/UBSan)"
$CC $WARN -O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer -I"$RT" $SRC -lm -lpthread $LDX -o "$TMP/safe"
ASAN_OPTIONS=${ASAN_OPTIONS:-detect_leaks=0} UBSAN_OPTIONS=halt_on_error=1 "$TMP/safe"
+107
View File
@@ -0,0 +1,107 @@
#!/usr/bin/env bash
# run_vindex_concurrency_tests.sh — regression harness for the 2026-08-16 soul crash.
#
# Four halves. The SET is the point: it separates two hazards the original two-half
# version conflated, and which have fixes in different files.
#
# 1. single ASan+UBSan, one thread. MUST be clean. Hard failure.
#
# 2. readers TSan, N readers, NO writer. Hazard (a): the visited set used
# to live on the index, so two pure READS stamped each other's
# epoch. Fixed in engram_vindex.c (frame-owned VVisit +
# `const VIndex*` search). MUST be clean. Hard failure.
#
# 3. unsynchronized TSan, writer + reader on a BARE index. Hazard (b): in-place
# HNSW insert rewires existing elements' neighbour lists and
# reallocs elems[]. EXPECTED TO RACE, PERMANENTLY. This is not
# a bug to fix inside engram_vindex.c — it is the executable
# proof that a publication boundary must exist above it.
# Not a failure. If it ever goes CLEAN, the test stopped
# interleaving and half 4 is no longer meaningful either.
#
# 4. published TSan, owner + N readers through a publication boundary
# (rwlock: readers shared, owner exclusive) mirroring
# eg_vindex_view / eg_vindex_maintain in lang/runtime/el_runtime.c.
# MUST be clean, and all inserts must land. Hard failure.
#
# See test_vindex_concurrency.c for the full story (SIGSEGV at ASCII address
# "gramNode", heap corruption in xzm_realloc, etc).
#
# usage: run_vindex_concurrency_tests.sh
set -uo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
RUNTIME="$(cd "$HERE/../../lang/runtime" && pwd)"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
SRC="$HERE/test_vindex_concurrency.c"
VINDEX="$RUNTIME/engram_vindex.c"
fail=0
echo "== [1/4] single-threaded control under AddressSanitizer =="
cc -std=c11 -g -O1 -fsanitize=address,undefined -fno-omit-frame-pointer \
-I"$RUNTIME" -o "$WORK/single" "$SRC" "$VINDEX" -lm || { echo "BUILD FAILED"; exit 2; }
if ASAN_OPTIONS=detect_leaks=0 "$WORK/single" single; then
echo " -> OK"
else
echo " -> FAIL: the single-threaded control must always be clean."
echo " If this fails the bug is NOT (only) concurrency — look for a real"
echo " out-of-bounds or lifetime error in engram_vindex.c."
fail=1
fi
cc -std=c11 -g -O1 -fsanitize=thread -fno-omit-frame-pointer \
-I"$RUNTIME" -o "$WORK/conc" "$SRC" "$VINDEX" -lm || { echo "BUILD FAILED"; exit 2; }
# run_tsan <mode> <logfile>; echoes nothing, sets $tsan_raced
run_tsan() {
TSAN_OPTIONS="halt_on_error=0" "$WORK/conc" "$1" >"$2" 2>&1
tsan_rc=$?
if grep -q "ThreadSanitizer: data race" "$2"; then tsan_raced=1; else tsan_raced=0; fi
}
echo
echo "== [2/4] concurrent READERS, no writer (visited-set gate) =="
run_tsan readers "$WORK/readers.log"
if [ "$tsan_raced" = "1" ]; then
echo " -> REGRESSION: two concurrent reads still race."
grep -m1 -A6 "ThreadSanitizer: data race" "$WORK/readers.log" | sed 's/^/ /'
echo " The visited set was supposed to be owned by the call frame."
fail=1
else
echo " -> clean (concurrent reads are safe)"
fi
echo
echo "== [3/4] writer+reader on a BARE index (expected-race probe) =="
run_tsan unsynchronized "$WORK/unsync.log"
if [ "$tsan_raced" = "1" ]; then
echo " -> RACE DETECTED, as expected:"
grep -m1 -A4 "ThreadSanitizer: data race" "$WORK/unsync.log" | sed 's/^/ /'
echo " In-place HNSW insert mutates existing elements. Not fixable inside"
echo " engram_vindex.c — this is why the publication boundary exists."
else
echo " -> NOTE: no race reported. The probe did not interleave; half 4's"
echo " clean result proves less than it should. Investigate."
fi
echo
echo "== [4/4] owner+readers through the publication boundary (boundary gate) =="
run_tsan published "$WORK/pub.log"
if [ "$tsan_raced" = "1" ]; then
echo " -> REGRESSION: the publication boundary did not serialize the owner."
grep -m1 -A6 "ThreadSanitizer: data race" "$WORK/pub.log" | sed 's/^/ /'
fail=1
elif [ "$tsan_rc" != "0" ]; then
echo " -> FAIL: boundary clean under TSan but the run failed:"
tail -3 "$WORK/pub.log" | sed 's/^/ /'
fail=1
else
echo " -> clean (readers project concurrently; the owner's inserts all landed)"
fi
echo
[ "$fail" -eq 0 ] && echo "RESULT: PASS" || echo "RESULT: FAIL"
exit "$fail"
+176
View File
@@ -0,0 +1,176 @@
/* test_grounding_vector.c — deterministic tests for §7: the one decay model, the
* consequence gate, and the stored/derived split. Links engram_cognition.c
* directly; no server, no store, no network. See run_grounding_vector_tests.sh.
*
* NEGATIVE CONTROL (invariant §8.6). Every symbol exercised here —
* cog_decay_factor, cog_grounding_significant, cog_significance_inherent,
* CogGrounding, CogProvClass — is introduced by the change under test, so this
* suite does not COMPILE against the pre-change source, let alone pass. The
* runner documents the exact reproduction.
*/
#include "engram_cognition.h"
#include <stdio.h>
#include <string.h>
#include <stdlib.h>
#include <math.h>
static int fails = 0;
static void ok(int cond, const char* what) {
printf(" %-62s %s\n", what, cond ? "PASS" : "*** FAIL ***");
if (!cond) fails++;
}
/* The decay formula exactly as el_runtime.c carried it before the move, so the
* refactor can be shown to be bit-identical rather than merely similar. */
static double old_engram_temporal_decay(long long age_ms, long long activation_count,
double temporal_decay_rate) {
if (age_ms <= 0) return 1.0;
double lambda = (temporal_decay_rate > 0.0) ? temporal_decay_rate : 0.693147;
double age_hours = (double)age_ms / 3600000.0;
double t_half = 168.0 * (1.0 + log(1.0 + (double)activation_count));
double factor = exp(-lambda * age_hours / t_half);
if (factor < 0.25) factor = 0.25;
return factor;
}
static CogGrounding base(void) {
CogGrounding g; memset(&g, 0, sizeof g);
g.present = 1;
g.factual = 0.60; g.relational = 0.60;
g.factual_now = 0.60; g.relational_now = 0.60;
g.associative = 0.1; g.polarity = 1.0;
g.prov = COG_PROV_TOLD;
g.fac_proj = 1.0; g.rel_proj = 1.0;
g.cos_angle = 0.9; g.agreement = 1;
g.ts = 1000; g.seq = 1; g.reinforcements = 3;
return g;
}
int main(void) {
const double F = 0.5, R = 0.5;
printf("\n== 1. DECAY IS THE ONE MODEL, AND IT IS BIT-IDENTICAL TO WHAT IT REPLACED ==\n");
{
long long ages[] = {0, 3600000LL, 86400000LL, 7*86400000LL, 30*86400000LL, 365*86400000LL};
int allsame = 1;
for (int i = 0; i < 6; i++)
for (int ac = 0; ac < 4; ac++) {
long long acs[] = {0, 1, 10, 1000};
double a = cog_decay_factor(ages[i], (double)acs[ac], 0.0);
double b = old_engram_temporal_decay(ages[i], acs[ac], 0.0);
if (a != b) allsame = 0;
}
ok(allsame, "cog_decay_factor == the pre-move engram_temporal_decay (24 pts)");
ok(cog_decay_factor(0, 0, 0.0) == 1.0, "age 0 -> no decay");
}
printf("\n DECAY OVER ELAPSED TIME (reinforcements = 0, default rate):\n");
printf(" %10s %10s\n", "elapsed", "decay");
{
struct { const char* label; long long ms; } pts[] = {
{"0", 0LL},
{"1 hour", 3600000LL},
{"1 day", 86400000LL},
{"3 days", 3LL*86400000LL},
{"7 days", 7LL*86400000LL},
{"14 days", 14LL*86400000LL},
{"30 days", 30LL*86400000LL},
{"90 days", 90LL*86400000LL},
};
double prev = 2.0; int monotone = 1;
for (unsigned i = 0; i < sizeof pts / sizeof pts[0]; i++) {
double d = cog_decay_factor(pts[i].ms, 0, 0.0);
printf(" %10s %10.6f\n", pts[i].label, d);
if (d > prev) monotone = 0;
prev = d;
}
ok(monotone, "decay is monotone non-increasing in elapsed time");
ok(fabs(cog_decay_factor(7LL*86400000LL, 0, 0.0) - 0.5) < 1e-6,
"7 days at zero reinforcements == exactly one half-life (0.5)");
ok(cog_decay_factor(7LL*86400000LL, 100, 0.0) > cog_decay_factor(7LL*86400000LL, 0, 0.0),
"reinforcement slows ageing (Lindy term)");
ok(cog_decay_factor(3650LL*86400000LL, 0, 0.0) == 0.25,
"floor is a preference not a cliff: bottoms out at 0.25");
}
printf("\n== 2. CONSEQUENCE GATE: EVERY TRIGGER, AND NO EPSILON ANYWHERE ==\n");
{
CogGrounding p = base(), n = base();
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_NONE,
"identical vectors -> NONE (a re-read must not consolidate)");
n = base(); n.factual = 0.9999; n.factual_now = 0.9999;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_NONE,
"factual 0.60 -> 0.9999 without crossing the floor -> NONE");
n = base(); n.relational = 0.5001; n.relational_now = 0.5001;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_NONE,
"relational 0.60 -> 0.5001, still above floor -> NONE");
n = base(); n.factual_now = 0.4999;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_FACTUAL_FLOOR,
"a 0.1001 drop that CROSSES the floor -> FACTUAL_FLOOR");
n = base(); n.relational_now = 0.4999;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_RELATIONAL_FLOOR,
"relational crossing its floor -> RELATIONAL_FLOOR");
n = base(); n.cos_angle = -0.05; n.agreement = -1;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_AGREEMENT_FLIP,
"agreement +1 -> -1 -> AGREEMENT_FLIP");
n = base(); n.fac_proj = -0.2;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_DIRECTION_REVERSAL,
"factual gradient reverses -> DIRECTION_REVERSAL");
n = base(); n.rel_proj = -0.2;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_DIRECTION_REVERSAL,
"relational gradient reverses -> DIRECTION_REVERSAL");
n = base(); n.polarity = -1.0;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_POLARITY_FLIP,
"support -> contradiction -> POLARITY_FLIP (inherent)");
n = base(); n.polarity = 0.0;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_POLARITY_FLIP,
"support -> ignorance (zero) -> POLARITY_FLIP: not the same state");
n = base(); n.prov = COG_PROV_OBSERVED;
ok(cog_grounding_significant(&p, &n, F, R) == COG_SIG_PROVENANCE_CHANGE,
"told -> observed -> PROVENANCE_CHANGE (inherent)");
CogGrounding fresh; memset(&fresh, 0, sizeof fresh);
ok(cog_grounding_significant(&fresh, &n, F, R) == COG_SIG_FIRST_RECORD,
"no prior version -> FIRST_RECORD");
}
printf("\n== 3. INHERENT MOVES BYPASS THE SALIENCE GATE ==\n");
ok(cog_significance_inherent(COG_SIG_POLARITY_FLIP), "polarity flip is inherent");
ok(cog_significance_inherent(COG_SIG_PROVENANCE_CHANGE), "provenance change is inherent");
ok(cog_significance_inherent(COG_SIG_FIRST_RECORD), "first record is inherent");
ok(!cog_significance_inherent(COG_SIG_FACTUAL_FLOOR), "a floor crossing is NOT inherent");
ok(!cog_significance_inherent(COG_SIG_NONE), "NONE is not inherent");
printf("\n== 4. THE STORED/DERIVED SPLIT: DERIVED VALUES ARE NEVER SERIALIZED ==\n");
{
CogGrounding g = base();
g.decay = 0.3333; g.factual_now = 0.1234; g.relational_now = 0.2345;
g.associative_now = 0.4567; g.age_ms = 999999; g.stale = 1;
char* m = cog_grounding_metadata("pre-existing=keepme", &g);
ok(m != NULL, "serializer returns a document");
ok(m && strstr(m, "pre-existing=keepme"), "pre-existing edge metadata preserved verbatim");
ok(m && strstr(m, "GRD1"), "GRD1 magic present");
ok(m && !strstr(m, "0.3333"), "decay is NOT stored");
ok(m && !strstr(m, "0.1234"), "factual_now is NOT stored");
ok(m && !strstr(m, "0.2345"), "relational_now is NOT stored");
ok(m && !strstr(m, "0.4567"), "associative_now is NOT stored");
ok(m && !strstr(m, "999999"), "age is NOT stored");
ok(m && strstr(m, "told"), "provenance class IS stored");
ok(m && strstr(m, "0.6"), "the factual/relational dimensions ARE stored");
if (m) { printf("\n --- serialized GRD1 block ---\n%s -----------------------------\n", m); }
free(m);
}
printf("\n%s (%d failure%s)\n\n", fails ? "SOME TESTS FAILED" : "ALL TESTS PASSED",
fails, fails == 1 ? "" : "s");
return fails ? 1 : 0;
}
+251
View File
@@ -0,0 +1,251 @@
/* test_vindex_concurrency.c — regression test for the 2026-08-16 soul crash.
*
* WHAT BROKE: the soul daemon crash-looped (5 crashes in ~100s) with SIGSEGV in
* search_layer <- vindex_insert <- eg_vindex_sync, a SIGABRT, and a fault inside
* xzm_realloc's own freelist — i.e. heap corruption. The SIGSEGV address
* 0x65646f4e6d617267 is little-endian ASCII "gramNode": string bytes being
* dereferenced as an Elem vector pointer.
*
* ROOT CAUSE: VIndex owns its traversal scratch (visited[] + visit_epoch), and
* search_layer mutates it via visited_reset(). So the index is unsafe for ANY
* concurrent use — including two concurrent READS. soul.el starts http_serve_async
* (a thread per connection) and then runs awareness_run() on the main thread, which
* reaches the same global index through engram_activate; nothing serialized them.
*
* Neither hnswlib nor FAISS puts the visited set on the index: hnswlib checks one
* out of a VisitedListPool per query, FAISS uses a thread_local VisitedTable.
*
* THE ORIGINAL `concurrent` HALF CONFLATED TWO DISTINCT HAZARDS (2026-08-16). It ran
* a writer against a reader on one bare index, so it could not tell apart:
*
* (a) READ/READ corruption — two searches stamping each other's visited epoch.
* A defect INSIDE engram_vindex.c, fixable there, and now fixed: the visited
* set moved to the call frame and vindex_search takes a `const VIndex*`.
*
* (b) WRITE/READ corruption — vindex_insert rewires the neighbour lists of
* EXISTING elements and reallocs elems[], so an insert is a mutation of the
* whole structure. This is NOT fixable inside engram_vindex.c at any price:
* it is inherent to in-place HNSW. It requires a publication boundary ABOVE
* the data structure (el_runtime.c: eg_vindex_view / eg_vindex_maintain).
*
* Conflating them made the suite unfailable-then-unpassable: fixing (a) left (b)
* still racing, which reads as "the fix did not work" when in fact a different,
* correctly-located fix is what (b) needs. So the halves are now separate:
*
* single N clustered vectors, ONE thread, ASan. The CONTROL. Must always
* be clean. When this passes and a concurrent half fails, the defect
* is concurrency, not an out-of-bounds/logic error in the graph code.
* (On 2026-08-16 this control cleared all 13,820 real dim-768 store
* vectors under ASan, which DISPROVED an inspection-derived hypothesis
* about an out-of-bounds reverse-link write at engram_vindex.c:340.)
*
* readers N reader threads, NO writer, one shared index, TSan. This is
* hazard (a) in isolation. It RACED before the visited set moved off
* the index struct and must be CLEAN now. Hard gate.
*
* unsynchronized writer + reader on a bare index, TSan. Hazard (b) in isolation.
* EXPECTED TO RACE, permanently — it is the executable proof that
* the index cannot be made safe from the inside, and therefore that
* the publication boundary in el_runtime.c has to exist. If this
* ever goes clean, the test stopped interleaving; do not celebrate.
*
* published writer + readers through a publication boundary that mirrors
* eg_vindex_view / eg_vindex_maintain (rwlock: readers shared,
* the single owner exclusive), TSan. Must be CLEAN. Hard gate.
* This is what proves the shape of the runtime fix, in the same
* process, rather than asserting it.
*
* Absence of a crash does NOT mean absence of a race — always read the sanitizer
* verdict, never just the exit code.
*
* Build/run: engram/test/run_vindex_concurrency_tests.sh
*/
#include "engram_vindex.h"
#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdint.h>
#define DIM 128
#define NVEC 3000
#define SEED_N 50
static VIndex* g_ix;
static float* g_vecs;
/* Deterministic filler. Real embeddings are strongly correlated, not uniform noise;
* clustering keeps many candidates near-equidistant, which exercises the diversity
* heuristic and the visited set far harder than random vectors do. */
static void fill_vectors(void) {
g_vecs = (float*)malloc((size_t)NVEC * DIM * sizeof(float));
if (!g_vecs) { fprintf(stderr, "OOM\n"); exit(1); }
for (int i = 0; i < NVEC; i++) {
int cluster = i % 8;
for (int d = 0; d < DIM; d++)
g_vecs[(size_t)i * DIM + d] =
(float)(((d + cluster * 7) % 13) / 13.0) +
(float)(((i * 2654435761u + (unsigned)d) % 97) / 9700.0);
}
}
static void* writer_fn(void* arg) {
(void)arg;
for (int i = SEED_N; i < NVEC; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
return NULL;
}
static void* reader_fn(void* arg) {
(void)arg;
uint64_t ids[8]; float ds[8];
for (int i = 0; i < 20000; i++)
(void)vindex_search(g_ix, g_vecs + (size_t)(i % NVEC) * DIM, 8, 0, ids, ds);
return NULL;
}
static int run_single(void) {
printf("[single] inserting %d vectors on one thread (ASan control)\n", NVEC);
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[single] vindex_create failed\n"); return 1; }
for (int i = 0; i < NVEC; i++) {
if (vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM) != 0) {
fprintf(stderr, "[single] insert %d failed\n", i); return 1;
}
}
if (vindex_size(g_ix) != (size_t)NVEC) {
fprintf(stderr, "[single] size %zu != %d\n", vindex_size(g_ix), NVEC); return 1;
}
uint64_t ids[16]; float ds[16];
for (int q = 0; q < 200; q++) {
int k = vindex_search(g_ix, g_vecs + (size_t)((q * 7) % NVEC) * DIM, 16, 0, ids, ds);
if (k < 0) { fprintf(stderr, "[single] search failed at q=%d\n", q); return 1; }
}
vindex_free(g_ix); g_ix = NULL;
printf("[single] PASS — no memory error (this must ALWAYS pass)\n");
return 0;
}
/* Hazard (b) in isolation: writer + reader on a BARE index, no boundary. */
static int run_unsynchronized(void) {
printf("[unsynchronized] 1 writer + 1 reader on a BARE index (TSan probe)\n");
printf("[unsynchronized] a race here is EXPECTED and PERMANENT — in-place HNSW\n");
printf("[unsynchronized] insert rewires existing elements. This is the proof that\n");
printf("[unsynchronized] the publication boundary must live ABOVE engram_vindex.c.\n");
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[unsynchronized] vindex_create failed\n"); return 1; }
for (int i = 0; i < SEED_N; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_t w, r;
if (pthread_create(&w, NULL, writer_fn, NULL) ||
pthread_create(&r, NULL, reader_fn, NULL)) {
fprintf(stderr, "[unsynchronized] pthread_create failed\n"); return 1;
}
pthread_join(w, NULL);
pthread_join(r, NULL);
vindex_free(g_ix); g_ix = NULL;
printf("[unsynchronized] completed — CHECK THE SANITIZER VERDICT, not this line.\n");
return 0;
}
/* ── hazard (a) in isolation: concurrent READS only ───────────────────────────
* This is what the frame-owned visited set fixes. Before that change, two
* vindex_search calls on one index wrote each other's epoch stamp; TSan reported
* the race at visited_reset and the traversal then walked bogus element indices. */
#define NREADERS 4
static int run_readers(void) {
printf("[readers] %d concurrent readers, NO writer, one shared index (TSan)\n", NREADERS);
printf("[readers] this is the visited-set regression gate — must be CLEAN.\n");
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[readers] vindex_create failed\n"); return 1; }
for (int i = 0; i < NVEC; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_t t[NREADERS];
for (int i = 0; i < NREADERS; i++)
if (pthread_create(&t[i], NULL, reader_fn, NULL)) {
fprintf(stderr, "[readers] pthread_create failed\n"); return 1;
}
for (int i = 0; i < NREADERS; i++) pthread_join(t[i], NULL);
vindex_free(g_ix); g_ix = NULL;
printf("[readers] completed — CHECK THE SANITIZER VERDICT, not this line.\n");
return 0;
}
/* ── the publication boundary, mirroring el_runtime.c ─────────────────────────
* Readers take the boundary SHARED and hold it across the whole search; the one
* owner takes it EXCLUSIVE to extend. Same shape as eg_vindex_view /
* eg_vindex_maintain. Note the reader's index pointer is `const VIndex*` — the
* compiler, not this comment, is what stops a reader inserting. */
static pthread_rwlock_t g_pub = PTHREAD_RWLOCK_INITIALIZER;
static void* pub_writer_fn(void* arg) {
(void)arg;
for (int i = SEED_N; i < NVEC; i++) {
pthread_rwlock_wrlock(&g_pub);
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_rwlock_unlock(&g_pub);
}
return NULL;
}
static void* pub_reader_fn(void* arg) {
(void)arg;
uint64_t ids[8]; float ds[8];
for (int i = 0; i < 5000; i++) {
pthread_rwlock_rdlock(&g_pub);
const VIndex* view = g_ix; /* immutable view */
(void)vindex_search(view, g_vecs + (size_t)(i % NVEC) * DIM, 8, 0, ids, ds);
pthread_rwlock_unlock(&g_pub);
}
return NULL;
}
static int run_published(void) {
printf("[published] 1 owner + %d readers through a publication boundary (TSan)\n", NREADERS);
printf("[published] this is the eg_vindex_view/eg_vindex_maintain gate — must be CLEAN.\n");
g_ix = vindex_create(DIM, 0, 0);
if (!g_ix) { fprintf(stderr, "[published] vindex_create failed\n"); return 1; }
for (int i = 0; i < SEED_N; i++)
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
pthread_t w, r[NREADERS];
if (pthread_create(&w, NULL, pub_writer_fn, NULL)) {
fprintf(stderr, "[published] pthread_create failed\n"); return 1;
}
for (int i = 0; i < NREADERS; i++)
if (pthread_create(&r[i], NULL, pub_reader_fn, NULL)) {
fprintf(stderr, "[published] pthread_create failed\n"); return 1;
}
pthread_join(w, NULL);
for (int i = 0; i < NREADERS; i++) pthread_join(r[i], NULL);
if (vindex_size(g_ix) != (size_t)NVEC) {
fprintf(stderr, "[published] size %zu != %d — the owner lost inserts\n",
vindex_size(g_ix), NVEC);
vindex_free(g_ix); g_ix = NULL; return 1;
}
vindex_free(g_ix); g_ix = NULL;
printf("[published] all %d inserts landed; CHECK THE SANITIZER VERDICT too.\n", NVEC);
return 0;
}
int main(int argc, char** argv) {
const char* mode = (argc > 1) ? argv[1] : "single";
fill_vectors();
int rc;
if (!strcmp(mode, "single")) rc = run_single();
else if (!strcmp(mode, "readers")) rc = run_readers();
else if (!strcmp(mode, "unsynchronized")) rc = run_unsynchronized();
else if (!strcmp(mode, "published")) rc = run_published();
/* back-compat: the pre-split name meant the bare writer+reader probe. */
else if (!strcmp(mode, "concurrent")) rc = run_unsynchronized();
else {
fprintf(stderr, "usage: %s [single|readers|unsynchronized|published]\n", argv[0]);
rc = 2;
}
free(g_vecs);
return rc;
}
+56 -12
View File
@@ -13,7 +13,7 @@
// relations add edges. Every node enters with PROVENANCE + grounding-level
// + stewardship class from the moment of entry.
//
// transduce() is THE single mechanism one function, polymorphic, with no
// transduce_bytes() is THE single mechanism one function, polymorphic, with no
// content-type branch inside it. It does not ask whether a payload is
// prose, structured data, or raw/opaque bytes (audio, or anything else);
// it runs one boundary-scan-with-fixed-window-fallback chunking algorithm
@@ -401,10 +401,54 @@ fn head80(s: String) -> String {
// truncates at the first embedded NUL, which is routine in real binary
// bytes) is a MECHANICAL fidelity concern that belongs to whatever produced
// `source` (see ingest_file's file_source_string below) not a
// content-type judgment made in here. transduce() never learns whether a
// content-type judgment made in here. transduce_bytes() never learns whether a
// chunk is plain text or a base64-encoded raw-byte window; every chunk is
// handled identically either way.
fn transduce(nodes: [String], edges: [String], source: String,
// NAMING, CORRECTED 2026-08-16 (second pass). This function was renamed
// `transduce` -> `transduce_bytes` earlier the same day, on the reasoning
// that it "was never signal->geometry — it chunks already-extracted content
// and PACKS it into a node+edge manifold, one layer up, and it had taken the
// name that belongs to the primitive underneath it."
//
// THAT REASONING WAS BACKWARDS, and it is worth recording why rather than
// quietly re-renaming. Producing a node+edge manifold is not a layer above
// transduction it IS transduction. Transduction is not conversion. When you
// take in music you do not store the song as one discrete geometry; you break
// it into its component parts and store the geometry of each along with the
// relations between them. The song is the structure of those relations.
// Signal -> one vector is the operation UNDERNEATH transduction, and its name
// is encoding, or geometry. So the layer that was doing it right got renamed
// out of the way so the layer doing it wrong could have the name.
//
// The primitive has since been corrected: `transduce(signal, modality)` now
// returns a Manifold components plus relations not a Geometry
// (el_runtime.c, "Manifold"). The two layers are therefore doing the SAME KIND
// of thing, and the inversion dissolves rather than needing to be re-argued.
//
// What is left is a real distinction, and it is about MODALITY, not layering:
//
// * `transduce(signal, modality)` dispatches to a realizer that KNOWS the
// modality and can name its components for audio: pitch, interval,
// rhythm, harmonic function.
// * `transduce_bytes` below is the OPAQUE-BYTES realizer: the decomposition
// available to a reader that knows nothing about what it is reading. It
// still yields components and relations (chunk nodes; contains / precedes
// / section_of edges), which is why it is transduction and not packing. It
// just cuts on the only structure visible without understanding byte
// boundaries so its components are positional rather than meaningful.
// That is a LIMITATION of this realizer, not the definition of the
// operation.
//
// The name is suffixed by its modality, not demoted to a lesser layer. Keeping
// a distinct symbol is also still mechanically required: every El `fn name`
// compiles to a global C symbol, so reusing `transduce` here is a hard
// `conflicting types` error the moment ingest.c links el_runtime.c.
//
// WHERE THIS SHOULD GO: this function should become a registered realizer
// returning a real Manifold, so ingest rides the same primitive as every other
// modality instead of carrying a parallel implementation. Not done here.
// Nothing about this function's behaviour changed in this pass.
fn transduce_bytes(nodes: [String], edges: [String], source: String,
prov: String, ground: String, steward: String,
root_lid: String, root_title: String) -> [String] {
let tagbase: String = "prov:" + prov + " ground:" + ground + " steward:" + steward
@@ -531,8 +575,8 @@ fn default_steward() -> String {
// trustworthy verbatim. When they don't (silent truncation happened),
// rebuild the payload as base64-encoded fixed-size windows read directly
// off disk (fs_read_b64_chunk binary-safe in C), joined with the same
// "\n\n" boundary marker transduce()'s generic scan already looks for, so
// transduce() sees one ordinary boundary-delimited payload and runs its one
// "\n\n" boundary marker transduce_bytes()'s generic scan already looks for, so
// transduce_bytes() sees one ordinary boundary-delimited payload and runs its one
// algorithm on it exactly as it would on prose it never learns that a
// fidelity problem occurred upstream, let alone why.
fn file_source_string(path: String, text: String, real_size: Int) -> String {
@@ -541,7 +585,7 @@ fn file_source_string(path: String, text: String, real_size: Int) -> String {
// 3072 raw bytes -> 4096 base64 chars (3 divides evenly into base64's
// 3-byte/4-char ratio); keeps each resulting node's content a clean,
// bounded, low-kilobytes unit, same order of magnitude as the fixed
// fallback window in transduce() itself.
// fallback window in transduce_bytes() itself.
let win: Int = 3072
let out: String = ""
let off: Int = 0
@@ -561,7 +605,7 @@ fn file_source_string(path: String, text: String, real_size: Int) -> String {
}
// ingest one file -> report JSON. Uniform for every file regardless of
// extension or content transduce() decides nothing about content-type, so
// extension or content transduce_bytes() decides nothing about content-type, so
// neither does this function; it only decides whether the raw bytes made it
// through the read intact (file_source_string), which is a fidelity
// question, not a format one.
@@ -573,14 +617,14 @@ fn ingest_file(path: String) -> String {
return "{\"error\":\"empty or unreadable\",\"path\":" + j_q(path) + "}"
}
let prov: String = "file:" + path
let packed: [String] = transduce(el_list_empty(), el_list_empty(),
let packed: [String] = transduce_bytes(el_list_empty(), el_list_empty(),
source, prov, default_ground(), default_steward(),
"doc:" + basename(path), basename(path))
return merge_packed(packed)
}
// ingest a directory: walk one level, ingest every file found, aggregate.
// No extension filter transduce() handles any payload uniformly now, so
// No extension filter transduce_bytes() handles any payload uniformly now, so
// there is no content-type gate at the directory boundary either.
fn ingest_dir(path: String) -> String {
let entries: [String] = fs_list(path)
@@ -615,7 +659,7 @@ fn ingest_dir(path: String) -> String {
fn ingest_url(url: String) -> String {
let body: String = http_get(url)
if str_eq(body, "") { return "{\"error\":\"empty fetch\",\"url\":" + j_q(url) + "}" }
let packed: [String] = transduce(el_list_empty(), el_list_empty(),
let packed: [String] = transduce_bytes(el_list_empty(), el_list_empty(),
body, "url:" + url, "extracted", "public-web",
"url:" + url, url)
return merge_packed(packed)
@@ -630,7 +674,7 @@ fn ingest_llm(query: String) -> String {
let resp: String = http_post_json("http://127.0.0.1:11434/api/generate", body)
let answer: String = json_get_string(resp, "response")
if str_eq(answer, "") { return "{\"error\":\"no model response\"}" }
let packed: [String] = transduce(el_list_empty(), el_list_empty(),
let packed: [String] = transduce_bytes(el_list_empty(), el_list_empty(),
answer, "llm:" + model + ":" + query, "candidate-provisional", "guide-provisional",
"llm:" + query, "guide answer: " + query)
return merge_packed(packed)
@@ -682,7 +726,7 @@ fn ingest_stream(path: String) -> String {
// It is NOT a content-type flag: it says nothing about what's inside the
// bytes once fetched, and none of the five ingest_* functions it selects
// among interpret their payload differently by content shape anymore
// they all hand off to the single, format-agnostic transduce(). The old
// they all hand off to the single, format-agnostic transduce_bytes(). The old
// "structured" value (a caller-declared alias for "file", used only to hint
// the now-removed JSON-vs-prose branch) is gone along with that branch.
let kind: String = env("INGEST_KIND")
+34 -12
View File
@@ -62,34 +62,56 @@ This is where almost all work belongs. El programs are source files that get com
This is the self-contained C OS-boundary layer. It provides the `__`-prefixed primitives that compiled El programs call: libcurl HTTP, pthreads, filesystem I/O, arena allocation, etc. It is **not generated** — it is maintained by hand.
The old `el_runtime.c` has been archived to `runtime/legacy/`. The runtime is now native El (`runtime/*.el`). `el_seed.c` replaces `el_runtime.c` as the sole C compilation dependency.
The runtime is native El (`runtime/*.el`) over a C OS-boundary. **Status (verified 2026-08-15):** the migration to a seed-only boundary is *in progress, not done*. Two files exist:
- `runtime/el_runtime.c` (~860 KB) — **LIVE**. Holds the engram store (`EngramStore engram_global`) plus the `http_*`/`json_*`/`state_*`/`engram_*` impls. It is the authoritative single-file link target for the compiler, and `tools/install.sh` compiles it into `libel.a`. This is where a new C builtin's *implementation* must currently live to be linkable.
- `runtime/el_seed.c` — the intended hand-maintained `__`-prefixed seed (thin wrappers over the above). It is compiled alongside `el_runtime.c` by `tools/install.sh`, but does **not** compile standalone yet (see the build-path caveat under "Rebuilding the Compiler").
**Only edit `el_seed.c` when you genuinely need OS-level access** (raw sockets, GPU calls, new libcurl features). For everything else, write El.
**Only edit these when you genuinely need OS-level access** (raw sockets, GPU calls, new libcurl features, a new engram store op). For everything else, write El.
When you do add a C builtin:
1. Add the C function to `el_seed.c`
2. Declare it in `el_seed.h`
3. Add it to the `builtin_arity` table in `el-compiler/src/codegen.el` (so the compiler knows the arg count)
4. Rebuild the elc binary (see below)
When you add a C builtin (verbatim-emit recipe — the El name is emitted as the exact C symbol; `builtin_arity` is an arity guard only, not a dispatch table):
1. Implement the C function in `el_runtime.c` (and declare it in `el_runtime.h`).
2. Add a `__`-prefixed thin wrapper in `el_seed.c` and declare it in `el_seed.h`.
3. Add the name to `builtin_arity` in `el-compiler/src/codegen.el` — add **both** the plain and `__`-prefixed spellings.
4. Rebuild the elc binary (see below) and confirm the self-host fixpoint is byte-identical.
5. **Prove it with a NEGATIVE CONTROL.** Show the test FAILING on a build without your change, then passing with it. A test that has never been seen to fail has proven nothing.
> **Step 5 is not optional, and step 4 does not cover it.** The fixpoint proves the *compiler reproduces itself*. It says nothing whatsoever about whether your builtin works. A recipe ending at "byte-identical" reads as complete while having verified nothing about the thing just added — which is why this file, until 2026-08-16, produced builtins with no tests at all.
>
> Measured cost of the omission (2026-08-16): `engram_node_set_emb`, `engram_curiosity_json` and `dream_set_handler` were all added in one session with zero tests. Separately, a UTF-8 fix was written, tested, and **the test passed on the unpatched build too** — the defect was elsewhere entirely, and only building the pre-fix binary exposed it. Without a negative control that fix would have merged as verified.
>
> Two shapes that pass while proving nothing, both hit the same day:
> - A test that never exercises your change (the route supplied a default that bypassed the code under test).
> - An induction that loses a race. `curl --max-time` on a large response left *both* builds alive; only `SO_LINGER 0` — a genuine RST, so the peer is provably gone — reproduced the failure. Six of ten attempts is not a control.
>
> Before every probe, confirm **your** process bound the port (`lsof -nP -iTCP:<port>`, match the PID). A stale instance answering on the port has silently produced false results here more than once, and `pkill -f` does not reliably match an argv like `./engram`.
Worked example: the `engram_assert_json` (op_assert seam) and `engram_node_full_in`/`engram_connect_in` (purview write-side) primitives added 2026-08-15 follow exactly this recipe.
---
## Rebuilding the Compiler
After changing any `.el` source in `el-compiler/src/`:
After changing any `.el` source in `el-compiler/src/` (run from the `lang/` dir):
```bash
cd /Users/will/Development/neuron-technologies/foundation/el
# 1. Stage2: current elc compiles the (modified) compiler to C
./dist/platform/elc elc-cli.el > elc-new.c
# 2. Build the new compiler. The C link target is el_runtime.c — it holds the
# engram store + http/json/state impls the compiler output calls. el_runtime.c
# self-hosts elc on its own; el_seed.c is the (aspirational) seed layer and does
# NOT compile standalone under clang (missing prototypes for the el_runtime.c
# symbols it wraps — see caveat below), so link el_runtime.c here.
cc -std=c11 -I runtime -lcurl -lpthread \
-o dist/platform/elc-new \
elc-new.c runtime/el_seed.c
# Verify self-hosting:
elc-new.c runtime/el_runtime.c
# 3. Verify self-hosting FIXPOINT (stage3 == stage2 output, byte-identical):
./dist/platform/elc-new elc-cli.el > elc-verify.c
diff elc-new.c elc-verify.c # should be identical
diff elc-new.c elc-verify.c # must be identical
mv dist/platform/elc-new dist/platform/elc
```
> **Build-path caveat (verified 2026-08-15).** `el_seed.c` is the intended hand-maintained OS-boundary seed, but it does **not** compile standalone under modern clang: it wraps ~16 unprefixed `el_runtime.c` symbols (`http_serve`, `json_*`, `state_*`, `http_response`) without prototypes, and clang treats implicit declarations as errors (C99+). The productionised install (`tools/install.sh`) builds `libel.a` from **both** `el_seed.o` + `el_runtime.o` together, which is why linking succeeds there. To make `el_seed.c` build on its own, add prototypes for those symbols (or `#include "el_runtime.h"`, reconciling the `__http_serve` return-type mismatch first). Until then, `el_runtime.c` is the authoritative single-file link target for the compiler.
After changing `el_seed.c` only (no El source changes), rebuild downstream programs but do NOT need to rebuild the compiler binary itself — the seed is linked at the application level, not the compiler level.
---
BIN
View File
Binary file not shown.
+242 -20
View File
@@ -862,10 +862,23 @@ fn cg_expr(expr: Map<String, Any>) -> String {
// arithmetic BinOp (or vice-versa). Without this check the
// fallthrough to str_eq produces str_eq(int_value, int_value)
// which reads the integer as a char* and segfaults.
// EITHER side provably Int is enough. Requiring BOTH meant a call
// whose return type codegen cannot infer poisoned the operator:
// getint(5) == a -> str_eq(getint(5), a)
// even with `a` declared Int. str_eq then reads an integer as a
// char* and segfaults. Only an integer LITERAL on one side forced
// the numeric form, so the bug was invisible in the common case.
//
// Loosening to OR is strictly safer: when one side is a known Int,
// str_eq is always wrong (it dereferences that int), while numeric
// comparison is at worst a wrong answer on an already ill-typed
// program. When neither side is Int nothing changes, so string
// comparison is untouched.
if is_int_expr(left) {
if is_int_expr(right) {
return "(" + left_c + " == " + right_c + ")"
}
return "(" + left_c + " == " + right_c + ")"
}
if is_int_expr(right) {
return "(" + left_c + " == " + right_c + ")"
}
// Float literal or negative float literal: use plain == (bit-equal
// el_val_t comparison). This handles `r0 == 3.0`, `neg == -3.0`, etc.
@@ -921,10 +934,12 @@ fn cg_expr(expr: Map<String, Any>) -> String {
}
// Same mixed Ident/BinOp fix as EqEq: use is_int_expr to detect
// integer-typed operands before falling through to !str_eq.
// Either side Int is enough see the EqEq note above.
if is_int_expr(left) {
if is_int_expr(right) {
return "(" + left_c + " != " + right_c + ")"
}
return "(" + left_c + " != " + right_c + ")"
}
if is_int_expr(right) {
return "(" + left_c + " != " + right_c + ")"
}
// Float-typed operands use plain != (bit-equal comparison).
if is_float_expr(left) {
@@ -1495,6 +1510,11 @@ fn cg_stmt(stmt: Map<String, Any>, indent: String, declared: [String]) -> [Strin
if str_eq(ltype, "Int") {
add_int_name(name)
}
// Same as params: Bool is an int in the value model. Without this a
// `let ok: Bool = ...` compared to another Bool lowered to str_eq.
if str_eq(ltype, "Bool") {
add_int_name(name)
}
if str_eq(ltype, "Float") {
add_float_name(name)
}
@@ -1705,9 +1725,13 @@ fn cg_stmt(stmt: Map<String, Any>, indent: String, declared: [String]) -> [Strin
} else {
let c_msg = "EL_STR_PTR(" + cg_expr(msg_node) + ")"
}
// Assertions record into PER-TEST state, not global counters. The test
// is the unit of result; a global pass/fail tally cannot say which test
// failed or whether a test ran at all. Reporting is the runner's job
// nothing is printed here.
emit_line(indent + "if (!(" + c_cond + ")) {")
emit_line(indent + " __el_test_fail(__el_cur_test, " + c_msg + "); __el_fail++;")
emit_line(indent + "} else { __el_pass++; }")
emit_line(indent + " __el_test_fail(" + c_msg + ");")
emit_line(indent + "} else { __el_cur_asserts++; }")
return declared
}
@@ -2602,6 +2626,17 @@ fn builtin_arity(name: String) -> Int {
// LSP seed primitives
if str_eq(name, "__read_n") { return 1 }
if str_eq(name, "__print_raw") { return 1 }
// Test-registry accessors. These are not runtime builtins they are
// GENERATED into the same translation unit by the --test path below, one
// set per test binary. They are declared here so the El-side runner in
// runtime/eltest.el can call them with a known arity.
if str_eq(name, "__el_reg_count") { return 0 }
if str_eq(name, "__el_reg_name") { return 1 }
if str_eq(name, "__el_reg_invoke") { return 1 }
if str_eq(name, "__el_reg_last_ns") { return 0 }
if str_eq(name, "__el_reg_msg") { return 0 }
if str_eq(name, "__el_reg_asserts") { return 0 }
if str_eq(name, "__el_opt_json") { return 0 }
// String
if str_eq(name, "el_str_concat") { return 2 }
if str_eq(name, "str_eq") { return 2 }
@@ -2760,7 +2795,15 @@ fn builtin_arity(name: String) -> Int {
if str_eq(name, "__engram_neighbors_filtered") { return 3 }
if str_eq(name, "__engram_activate") { return 2 }
if str_eq(name, "__engram_activate_json") { return 2 }
if str_eq(name, "__engram_op_assert_json") { return 2 }
if str_eq(name, "__engram_node_full_in") { return 9 }
if str_eq(name, "__engram_connect_in") { return 5 }
if str_eq(name, "__engram_scan_nodes_json") { return 2 }
if str_eq(name, "__engram_edges_json") { return 2 }
if str_eq(name, "__engram_pool_stats_json") { return 0 }
if str_eq(name, "__el_alloc_count") { return 0 }
if str_eq(name, "__el_alloc_bytes") { return 0 }
if str_eq(name, "__el_peak_rss") { return 0 }
if str_eq(name, "__generate") { return 1 }
// Filesystem
if str_eq(name, "fs_read") { return 1 }
@@ -2859,9 +2902,18 @@ fn builtin_arity(name: String) -> Int {
if str_eq(name, "engram_get_node_by_label") { return 1 }
if str_eq(name, "engram_search_json") { return 2 }
if str_eq(name, "engram_scan_nodes_json") { return 2 }
if str_eq(name, "engram_edges_json") { return 2 }
if str_eq(name, "engram_pool_stats_json") { return 0 }
if str_eq(name, "el_alloc_count") { return 0 }
if str_eq(name, "el_alloc_bytes") { return 0 }
if str_eq(name, "el_peak_rss") { return 0 }
if str_eq(name, "el_black_box") { return 1 }
if str_eq(name, "engram_neighbors_json") { return 3 }
if str_eq(name, "engram_activate_json") { return 2 }
if str_eq(name, "engram_stats_json") { return 0 }
if str_eq(name, "engram_op_assert_json") { return 2 }
if str_eq(name, "engram_node_full_in") { return 9 }
if str_eq(name, "engram_connect_in") { return 5 }
// LLM
if str_eq(name, "llm_call") { return 2 }
if str_eq(name, "llm_call_system") { return 3 }
@@ -3081,6 +3133,15 @@ fn build_int_names_for_params(params: [Map<String, Any>]) -> Bool {
if str_eq(ptype, "Int") {
add_int_name(pname)
}
// Bool is an integer in the value model (type_to_c maps Bool -> "int";
// el_runtime.h: "Bool -> el_val_t (0 = false, nonzero = true)"), but
// Bool names were registered nowhere. So `cond == want` between two
// Bool params fell through to str_eq and dereferenced 0 or 1 as a
// char* an immediate segfault. Track them as int-like, which is what
// they are.
if str_eq(ptype, "Bool") {
add_int_name(pname)
}
if str_eq(ptype, "Float") {
add_float_name(pname)
}
@@ -3204,6 +3265,7 @@ fn is_top_level_decl(stmt: Map<String, Any>) -> Bool {
if kind == "EnumDef" { return true }
if kind == "Import" { return true }
if kind == "CgiBlock" { return true }
if kind == "ProgramBlock" { return true }
if kind == "ExternFn" { return true }
false
}
@@ -3216,6 +3278,55 @@ fn cgi_arg(value: String, has_value: Bool) -> String {
return "EL_NULL"
}
// -- Program block: cross-cutting concerns injected at the process boundary ----
//
// emit_program_init emit the `static void __el_program_init(void)` that
// carries a program's declared cross-cutting concerns. Called from main()
// BEFORE any user statement runs, so the guarantees hold for the whole process
// rather than depending on each call site remembering to ask for them.
//
// This is emitted at the point the `program` block is encountered, not buffered
// until main(). The streaming backend emits in source order and cannot hold a
// declaration's entry list alive until main(); emitting a named function here
// and calling it from main() means only a single bool has to survive.
//
// Order matters and is deliberate:
// 1. singleton FIRST if another instance already holds the lock, refuse and
// exit before touching configuration, ports, or any data directory.
// 2. config declarations resolve env-or-default, one declaration per entry.
// 3. validate LAST report EVERY missing/ill-typed entry at once, then exit.
fn el_bool_arg(b: Bool) -> String {
if b { return "EL_INT(1)" }
return "EL_INT(0)"
}
fn emit_program_init(stmt: Map<String, Any>) -> Void {
let pname: String = stmt["name"]
emit_line("static void __el_program_init(void) {")
let has_singleton: Bool = stmt["has_singleton"]
if has_singleton {
let sid: String = stmt["singleton"]
emit_line(" el_singleton_acquire(EL_STR(" + c_str_lit(sid) + "));")
}
let entries = stmt["entries"]
let n: Int = native_list_len(entries)
let i = 0
while i < n {
let e = native_list_get(entries, i)
let ename: String = e["name"]
let etype: String = e["etype"]
let edefault: String = e["default"]
let has_default: Bool = e["has_default"]
let erequired: Bool = e["required"]
let arg_def: String = cgi_arg(edefault, has_default)
emit_line(" el_config_declare(EL_STR(" + c_str_lit(ename) + "), EL_STR(" + c_str_lit(etype) + "), " + arg_def + ", " + el_bool_arg(has_default) + ", " + el_bool_arg(erequired) + ");")
let i = i + 1
}
emit_line(" el_config_validate(EL_STR(" + c_str_lit(pname) + "));")
emit_line("}")
emit_blank()
}
// -- VBD role enforcement ------------------------------------------------------
//
// Scan a function body for direct calls to DHARMA-restricted builtins
@@ -3538,6 +3649,20 @@ fn codegen(stmts: [Map<String, Any>], source: String) -> String {
}
}
// Program block: emit the cross-cutting init function before the user's
// functions so main() can call it (see emit_program_init).
let prog_have: Bool = false
let i = 0
while i < n {
let stmt = native_list_get(stmts, i)
let sk4: String = stmt["stmt"]
if str_eq(sk4, "ProgramBlock") {
emit_program_init(stmt)
let prog_have = true
}
let i = i + 1
}
// Function definitions
let i = 0
while i < n {
@@ -3556,6 +3681,9 @@ fn codegen(stmts: [Map<String, Any>], source: String) -> String {
// with the C-side parameters when fn main()'s body is folded in below.
emit_line("int main(int _argc, char** _argv) {")
emit_line(" el_runtime_init_args(_argc, _argv);")
if prog_have {
emit_line(" __el_program_init();")
}
if cgi_count >= 1 {
let cname: String = cgi_block["name"]
let cdid: String = cgi_block["dharma_id"]
@@ -4100,13 +4228,36 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
// Emit test harness preamble (counters, fail printer) when in test mode.
if test_is_mode {
emit_line("#include <stdio.h>")
emit_line("#include <string.h>")
emit_line("#include <time.h>")
emit_blank()
emit_line("static int __el_pass = 0, __el_fail = 0;")
// Per-test result state. Reset by __el_reg_invoke before each test, so
// every test gets its own record rather than contributing to a global
// tally. The first failure message is retained; later ones only bump
// the count, which keeps the common case allocation-free.
emit_line("static int __el_cur_fails = 0;")
emit_line("static int __el_cur_asserts = 0;")
emit_line("static char __el_cur_msg[512] = \"\";")
emit_line("static const char *__el_cur_test = \"(none)\";")
emit_line("static void __el_test_fail(const char *test, const char *msg) {")
emit_line(" fprintf(stderr, \"FAIL %-40s %s\\n\", test, msg);")
emit_line("static void __el_test_fail(const char *msg) {")
emit_line(" if (__el_cur_fails == 0 && msg) {")
emit_line(" snprintf(__el_cur_msg, sizeof __el_cur_msg, \"%s\", msg);")
emit_line(" }")
emit_line(" __el_cur_fails++; __el_cur_asserts++;")
emit_line("}")
emit_blank()
// Forward declarations for the registry accessors. The definitions are
// emitted at the END of the unit (they reference the test functions,
// which do not exist yet at this point), but the El-side runner is
// compiled in between and calls them so it needs the prototypes here.
emit_line("el_val_t __el_reg_count(void);")
emit_line("el_val_t __el_reg_name(el_val_t i);")
emit_line("el_val_t __el_reg_invoke(el_val_t i);")
emit_line("el_val_t __el_reg_last_ns(void);")
emit_line("el_val_t __el_reg_msg(void);")
emit_line("el_val_t __el_reg_asserts(void);")
emit_line("el_val_t __el_opt_json(void);")
emit_blank()
}
// Streaming parse-emit loop.
@@ -4126,6 +4277,7 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
// Fix: copy the values out BEFORE the release (strings, so no dangling reference)
// and emit from these. No search, so the failure mode is removed rather than moved.
let cgi_have: Bool = false
let prog_have: Bool = false
let cgi_name_v: String = ""
let cgi_did_v: String = ""
let cgi_prin_v: String = ""
@@ -4247,6 +4399,14 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
// These are no-ops in codegen (forward decls already emitted)
// except a CgiBlock, whose declared identity must survive
// this release to be emitted as a compiled constant.
// A ProgramBlock's cross-cutting declarations are
// emitted HERE, as a named init function, because the
// streaming backend cannot hold the entry list alive
// until main(). Only the bool survives.
if str_eq(sk, "ProgramBlock") {
emit_program_init(stmt)
let prog_have = true
}
if str_eq(sk, "CgiBlock") {
let cgi_have = true
let cgi_name_v = stmt["name"]
@@ -4302,17 +4462,72 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
el_release(sigs)
let test_arena_mark: Any = el_arena_push()
let tn: Int = native_list_len(test_c_names)
// Generated test registry
// Discovery happens HERE, at compile time. The runner never searches
// for tests; it walks this table. That ordering — discovery strictly
// before execution is what makes --list, filtering, sharding and
// per-test reporting possible later, and it is why the old harness
// (which inlined direct calls into main) could not have any of them.
emit_line("typedef void (*__el_test_fp)(void);")
emit_line("typedef struct { const char *name; __el_test_fp fn; } __el_test_entry;")
emit_line("static const __el_test_entry __el_registry[] = {")
let ri: Int = 0
while ri < tn {
let r_name: String = native_list_get(test_names, ri)
let r_cfn: String = native_list_get(test_c_names, ri)
emit_line(" { \"" + c_escape(r_name) + "\", " + r_cfn + " },")
let ri = ri + 1
}
// Trailing sentinel keeps the array non-empty when a file declares no
// tests (a zero-length array is not valid C).
emit_line(" { 0, 0 }")
emit_line("};")
emit_line("static const int __el_registry_n = " + int_to_str(tn) + ";")
emit_blank()
emit_line("static long long __el_last_ns = 0;")
emit_line("static int __el_opt_json_v = 0;")
emit_blank()
// Index-based accessors
// El has no function pointers, so the runner works purely in indices.
// This is the whole seam between generated C and the El-side runner.
emit_line("el_val_t __el_reg_count(void) { return (el_val_t)(int64_t)__el_registry_n; }")
emit_line("el_val_t __el_reg_name(el_val_t i) {")
emit_line(" int64_t k = (int64_t)i;")
emit_line(" if (k < 0 || k >= __el_registry_n) return EL_STR(\"\");")
emit_line(" return EL_STR(__el_registry[k].name);")
emit_line("}")
// Timing is taken immediately around the call, in C, on the MONOTONIC
// clock never the wall clock, which can step backwards under NTP.
emit_line("el_val_t __el_reg_invoke(el_val_t i) {")
emit_line(" int64_t k = (int64_t)i;")
emit_line(" if (k < 0 || k >= __el_registry_n) return 0;")
emit_line(" __el_cur_fails = 0; __el_cur_asserts = 0; __el_cur_msg[0] = '\\0';")
emit_line(" __el_cur_test = __el_registry[k].name;")
emit_line(" struct timespec _t0, _t1;")
emit_line(" clock_gettime(CLOCK_MONOTONIC, &_t0);")
emit_line(" __el_registry[k].fn();")
emit_line(" clock_gettime(CLOCK_MONOTONIC, &_t1);")
emit_line(" __el_last_ns = (long long)(_t1.tv_sec - _t0.tv_sec) * 1000000000LL")
emit_line(" + (long long)(_t1.tv_nsec - _t0.tv_nsec);")
emit_line(" return (el_val_t)(int64_t)__el_cur_fails;")
emit_line("}")
emit_line("el_val_t __el_reg_last_ns(void) { return (el_val_t)(int64_t)__el_last_ns; }")
emit_line("el_val_t __el_reg_msg(void) { return EL_STR(__el_cur_msg); }")
emit_line("el_val_t __el_reg_asserts(void) { return (el_val_t)(int64_t)__el_cur_asserts; }")
emit_line("el_val_t __el_opt_json(void) { return (el_val_t)(int64_t)__el_opt_json_v; }")
emit_blank()
// main() delegates to the El-side runner. Everything above this line is
// generated glue; all reporting logic lives in runtime/eltest.el.
emit_line("int main(int _argc, char **_argv) {")
emit_line(" el_runtime_init_args(_argc, _argv);")
let ti: Int = 0
let tn: Int = native_list_len(test_c_names)
while ti < tn {
let tc_name: String = native_list_get(test_c_names, ti)
emit_line(" " + tc_name + "();")
let ti = ti + 1
}
emit_line(" printf(\"%d passed, %d failed\\n\", __el_pass, __el_fail);")
emit_line(" return __el_fail;")
emit_line(" for (int _i = 1; _i < _argc; _i++) {")
emit_line(" if (strcmp(_argv[_i], \"--json\") == 0) __el_opt_json_v = 1;")
emit_line(" }")
emit_line(" return (int)(int64_t)el_test_main();")
emit_line("}")
el_arena_pop(test_arena_mark)
el_release(test_names)
@@ -4338,6 +4553,13 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
let kind2: String = state_get("__program_kind")
emit_line("int main(int _argc, char** _argv) {")
emit_line(" el_runtime_init_args(_argc, _argv);")
// Cross-cutting concerns declared by a `program` block run BEFORE anything
// else a singleton violation must refuse the start before this process
// touches a port or a data directory, and configuration must be validated
// before the first read of it rather than at each read site.
if prog_have {
emit_line(" __el_program_init();")
}
// cgi init if needed
let ns2: Int = native_list_len(sigs)
+16
View File
@@ -419,6 +419,22 @@ fn resolve_imports(src_path: String) -> String {
if !str_eq(already, "") { return "" }
state_set(seen_key, "1")
// A missing file must be a hard error, never an empty string.
//
// fs_read returns "" both for "file is empty" and "file does not exist", and
// this function used the value without distinguishing them. So a broken
// import path a typo, a moved file, a relative path resolved from the
// wrong working directory compiled CLEANLY: exit 0, empty stderr, and a
// program silently missing everything it imported. Observed 2026-08-15:
// eleven consecutive "successful" compiles that had included no runtime at
// all, and a wrong conclusion drawn from them before anyone noticed.
//
// Missing dependency, confident success. fs_exists separates the two cases,
// so a genuinely empty file still resolves to "" and is fine.
if !fs_exists(src_path) {
println("elc: cannot resolve import: " + src_path)
exit_program(1)
}
let source: String = fs_read(src_path)
let dir: String = dirname_of(src_path)
let lines: [String] = str_split(source, "\n")
+1
View File
@@ -184,6 +184,7 @@ fn keyword_kind(word: String) -> String {
if word == "false" { return "Bool" }
if word == "cgi" { return "Cgi" }
if word == "service" { return "Service" }
if word == "program" { return "Program" }
if word == "manager" { return "Manager" }
if word == "engine" { return "Engine" }
if word == "accessor" { return "Accessor" }
+124 -1
View File
@@ -1967,6 +1967,113 @@ fn parse_stmt(tokens: [Any], pos: Int) -> Map<String, Any> {
}, p)
}
// program block: program "name" { singleton: "id", env NAME: Type = "default", ... }
//
// The program block is El's declaration surface for CROSS-CUTTING CONCERNS
// properties of the whole process rather than of any one function, which
// otherwise degrade into "remember to call this at every site" conventions.
//
// singleton: "id" process identity. The runtime takes an exclusive
// lock at startup; a SECOND start is refused, loudly,
// instead of two processes sharing one data dir.
// env NAME: T = "d" one configuration entry. Its type and its default
// are declared ONCE, here, and resolved+validated
// before main() body runs.
// env NAME: T required
// no default; the program refuses to start unless the
// variable is set.
//
// Both compile into calls injected at the head of main() the same boundary
// seam `cgi` already uses (codegen.el emit_program_init). No call site in the
// program body has to remember anything, which is the whole point.
if k == "Program" {
let p = pos + 1
let name = tok_value(tokens, p)
let p = p + 1
let p = expect(tokens, p, "LBrace")
let singleton = ""
let has_singleton = false
let entries = native_list_empty()
// Entry-scratch declared at loop-body level (not inside the branch) so
// that inner `let` forms compile to assignment rather than a C-scoped
// redeclaration the same idiom the service block above relies on.
let ename = ""
let etype = ""
let edefault = ""
let has_default = false
let erequired = false
let fname = ""
let fval = ""
let running = true
while running {
let k2 = tok_kind(tokens, p)
if k2 == "RBrace" {
let running = false
} else {
if k2 == "Eof" {
let running = false
} else {
let fname = tok_value(tokens, p)
let p = p + 1
if str_eq(fname, "env") {
// env NAME: Type [= "default"] [required]
let ename = tok_value(tokens, p)
let p = p + 1
let p = expect(tokens, p, "Colon")
let etype = tok_value(tokens, p)
let p = p + 1
let edefault = ""
let has_default = false
let erequired = false
let k3 = tok_kind(tokens, p)
if str_eq(k3, "Eq") {
let p = p + 1
let edefault = tok_value(tokens, p)
let has_default = true
let p = p + 1
}
let k4 = tok_kind(tokens, p)
if str_eq(k4, "Ident") {
let w = tok_value(tokens, p)
if str_eq(w, "required") {
let erequired = true
let p = p + 1
}
}
let entries = native_list_append(entries, {
"name": ename,
"etype": etype,
"default": edefault,
"has_default": has_default,
"required": erequired
})
} else {
// scalar field: `name: "value"`
let p = expect(tokens, p, "Colon")
let fval = tok_value(tokens, p)
let p = p + 1
if str_eq(fname, "singleton") {
let singleton = fval
let has_singleton = true
}
}
let k5 = tok_kind(tokens, p)
if k5 == "Comma" {
let p = p + 1
}
}
}
}
let p = expect(tokens, p, "RBrace")
return make_result({
"stmt": "ProgramBlock",
"name": name,
"singleton": singleton,
"has_singleton": has_singleton,
"entries": entries
}, p)
}
// assert <cond_expr> [ , <msg_expr> ]
// The message is optional if the next token after the condition is not a
// Comma, emit an empty string placeholder so the test still works.
@@ -2419,6 +2526,7 @@ fn scan_params_c(tokens: [Any], pos: Int) -> Map<String, Any> {
// toplevel_let: { "kind": "toplevel_let", "name": String, "ltype": String }
// cgi_block: { "kind": "cgi_block", "name": String }
// service_block: { "kind": "service_block", "name": String }
// program_block: { "kind": "program_block", "name": String }
//
// Import/TypeDef/EnumDef nodes are skipped (codegen treats them as no-ops).
//
@@ -2546,13 +2654,28 @@ fn scan_fn_sigs(tokens: [Any]) -> [Map<String, Any>] {
"name": name
})
let pos = p
} else {
// --- program block ---
if str_eq(k, "Program") {
let p: Int = pos + 1
let name: String = tok_value(tokens, p)
let p = p + 1
let k2: String = tok_kind(tokens, p)
if str_eq(k2, "LBrace") {
let p = skip_to_rbrace(tokens, p)
}
let sigs = native_list_append(sigs, {
"kind": "program_block",
"name": name
})
let pos = p
} else {
// Import, Type, Enum, From, or any other token.
// Skip ahead to the next statement boundary.
let p: Int = pos + 1
let p = skip_expr_to_stmt_boundary(tokens, p)
let pos = p
}}}}}
}}}}}}
}
}
}
+228
View File
@@ -0,0 +1,228 @@
// transduce.el transduction decomposes a signal into components and the
// relations between them. Runnable: this is the worked example for the
// transduce surface, and it exits non-zero if any claim in it stops being true.
//
// elc lang/examples/transduce.el > transduce.c
// cc -std=c11 -O2 -I lang/runtime -o transduce transduce.c \
// lang/runtime/el_runtime.c lang/runtime/el_seed.c \
// lang/runtime/engram_store.c lang/runtime/engram_vindex.c \
// lang/runtime/engram_cognition.c lang/runtime/engram_geometry.c \
// lang/runtime/engram_reason.c lang/runtime/engram_verify.c \
// -lcurl -lpthread -lm
// ./transduce # exits 0 only if every check passes
//
// It writes to an IN-MEMORY engram (leave ENGRAM_STORE unset) and contacts no
// server. The same claims are asserted by the native harness in
// lang/tests/native/test_transduce.el.
//
// WHAT CHANGED, AND WHY IT MATTERS. #144 shipped
// `transduce(signal, modality) -> Geometry`: one vector per signal. That made
// transduction a CONVERSION take a thing, encode it, store a position and
// what a conversion returns is a fingerprint. A fingerprint can be matched and
// ranked, and that is all it can ever do. It cannot be decomposed, cannot have
// one part grounded while another is not, and cannot be contradicted in one
// part while holding in another, because it has no parts.
//
// A song is not a point. It decomposes into pitch, interval, rhythm, harmonic
// function components, each with its own geometry, plus the relations among
// them. THE SONG IS THE STRUCTURE OF THE RELATIONS. transduce now returns a
// Manifold, and a realizer's job is to say what its modality's components ARE.
fn check(ok: Int, label: String) -> Int {
if ok > 0 {
println(" ok " + label)
return 0
}
println(" FAIL " + label)
exit(1)
return 1
}
fn near(a: Float, b: Float) -> Int {
let d: Float = a - b
if d > 0.001 { return 0 }
if d < -0.001 { return 0 }
return 1
}
fn eq_int(a: Int, b: Int) -> Int {
if a == b { return 1 }
return 0
}
// A DECOMPOSING realizer, written entirely in El
// "tone" signals are note letters, e.g. "CEG". This does NOT return one vector
// for the chord. It returns the PARTS one component per note, one per
// interval between adjacent notes and the relations that make those parts a
// chord rather than an unordered bag of pitches.
//
// The interval is deliberately a COMPONENT, not a field on a note. An interval
// is a thing with its own geometry belonging to neither endpoint; modelling it
// as an attribute of one of them is the same collapse, one level down.
fn tone_realizer(signal: String) -> Manifold {
let m: Manifold = manifold_new()
let n: Int = str_len(signal)
let i: Int = 0
while i < n {
let code: Int = str_char_code(signal, i)
let g: Geometry = geometry_new(2)
let s0: Int = geometry_set(g, 0, int_to_float(code))
let s1: Int = geometry_set(g, 1, int_to_float(i))
let idx: Int = manifold_add(m, "note:" + int_to_str(i), "pitch", g)
let f: Int = geometry_free(g)
i = i + 1
}
let j: Int = 1
while j < n {
let a: Int = str_char_code(signal, j - 1)
let b: Int = str_char_code(signal, j)
let lo: String = "note:" + int_to_str(j - 1)
let hi: String = "note:" + int_to_str(j)
let key: String = "interval:" + int_to_str(j - 1) + "-" + int_to_str(j)
let g: Geometry = geometry_new(1)
let s: Int = geometry_set(g, 0, int_to_float(b - a))
let idx: Int = manifold_add(m, key, "interval", g)
let f: Int = geometry_free(g)
let e1: Int = manifold_relate(m, key, "spans", lo, 0.9)
let e2: Int = manifold_relate(m, key, "spans", hi, 0.9)
let e3: Int = manifold_relate(m, lo, "sounds_before", hi, 0.8)
j = j + 1
}
m
}
// #144's contract, kept as a control: one vector for the whole signal.
fn fingerprint_realizer(signal: String) -> Geometry {
let g: Geometry = geometry_new(4)
let n: Int = str_len(signal)
let a: Int = geometry_set(g, 0, int_to_float(n))
g
}
fn main() -> Void {
println("a realizer declared in El is a first-class realizer")
let reg: Int = realizer_register("tone", "tone_realizer")
let _c: Int = check(reg, "an El fn registers as a realizer by name")
let _c: Int = check(realizer_has("tone"), "the modality now has an organ")
println("transduction decomposes a signal into parts")
let m: Manifold = transduce("CEG", "tone")
let _c: Int = check(manifold_is(m), "transduce returns a real Manifold")
let sz: Int = manifold_size(m)
let _c: Int = check(eq_int(sz, 5), "three notes and two intervals are five parts")
let rc: Int = manifold_rel_count(m)
let _c: Int = check(eq_int(rc, 6), "and they stand in six stated relations")
println("every part is addressable BY KEY, which is what survives persistence")
let i_c: Int = manifold_index_of(m, "note:0")
let _c: Int = check(1 - eq_int(i_c, -1), "the first note is addressable on its own")
let i_iv: Int = manifold_index_of(m, "interval:0-1")
let _c: Int = check(1 - eq_int(i_iv, -1), "so is the interval between the first two")
let miss: Int = manifold_index_of(m, "never_added")
let _c: Int = check(eq_int(miss, -1), "an unknown key is -1, not component 0")
println("parts carry their own geometry, and may differ in width")
let gn: Geometry = manifold_geometry(m, i_c)
let _c: Int = check(eq_int(geometry_dim(gn), 2), "a note component is 2 wide")
let _c: Int = check(near(geometry_get(gn, 0), 67.0), "and it is C — the signal reached the realizer")
let gi: Geometry = manifold_geometry(m, i_iv)
let _c: Int = check(eq_int(geometry_dim(gi), 1), "an interval component is 1 wide")
// A single vector per signal cannot represent parts of unequal width at all.
let _c: Int = check(near(geometry_get(gi, 0), 2.0), "C to E is two semitones")
let f1: Int = geometry_free(gn)
let f2: Int = geometry_free(gi)
println("the relations are content no single part carries")
// That "2" above is not a property of C and not a property of E. It exists
// only BETWEEN them, so a representation with no relations cannot hold it.
let spans: Int = 0
let k: Int = 0
while k < rc {
if str_eq(manifold_rel_name(m, k), "spans") {
if str_eq(manifold_rel_from(m, k), "interval:0-1") { spans = spans + 1 }
}
k = k + 1
}
let _c: Int = check(eq_int(spans, 2), "the interval is wired to both notes it spans")
println("relation weight IS the grounding (correspondence-and-censorship §1)")
let wk: Int = 0
let found: Int = 0
while wk < rc {
if str_eq(manifold_rel_name(m, wk), "sounds_before") {
if near(manifold_rel_weight(m, wk), 0.8) > 0 { found = 1 }
}
wk = wk + 1
}
let _c: Int = check(found, "the ordering relation carries the weight its realizer stated")
println("the decomposition persists as real, separately addressable nodes")
let ids: [String] = el_list_empty()
let n0: Int = engram_node_count()
let e0: Int = engram_edge_count()
let pi: Int = 0
while pi < sz {
let key: String = manifold_key(m, pi)
let g: Geometry = manifold_geometry(m, pi)
let id: String = engram_node("component " + key, "Concept", 0.6)
let att: Int = node_attach_geometry(id, g)
ids = el_list_append(ids, id)
let ff: Int = geometry_free(g)
pi = pi + 1
}
let ri: Int = 0
while ri < rc {
let fi: Int = manifold_index_of(m, manifold_rel_from(m, ri))
let ti: Int = manifold_index_of(m, manifold_rel_to(m, ri))
engram_connect(el_list_get(ids, fi), el_list_get(ids, ti),
manifold_rel_weight(m, ri), manifold_rel_name(m, ri))
ri = ri + 1
}
let _c: Int = check(eq_int(engram_node_count() - n0, 5), "one signal became five nodes")
let _c: Int = check(eq_int(engram_edge_count() - e0, 6), "and six edges between them")
println("each part's geometry is independently readable back off its node")
let id_c: String = el_list_get(ids, manifold_index_of(m, "note:0"))
let id_iv: String = el_list_get(ids, manifold_index_of(m, "interval:0-1"))
let _c: Int = check(eq_int(node_geometry_dim(id_c), 2), "note:0 node carries a 2-wide geometry")
let _c: Int = check(eq_int(node_geometry_dim(id_iv), 1), "interval:0-1 node carries a 1-wide one")
println("one part can be grounded without touching its siblings")
let ear: String = engram_node("evidence: heard a C in the recording", "Memory", 0.7)
engram_connect(ear, id_c, 0.95, "corroborates")
let _c: Int = check(engram_edge_between(ear, id_c), "evidence attaches to note:0 specifically")
let id_g: String = el_list_get(ids, manifold_index_of(m, "note:2"))
let _c: Int = check(1 - engram_edge_between(ear, id_g), "and NOT to note:2 — the sibling is untouched")
// This is the whole gain, and it is impossible with a fingerprint: with one
// node per signal, "the C is corroborated" and "the G is not" have the same
// grounding target and cannot both be recorded.
let _c: Int = check(eq_int(node_geometry_dim(id_g), 2), "note:2 geometry is intact regardless")
println("a fingerprint realizer transduces NOTHING")
// #144's contract exactly: signal in, one Geometry out. It resolves, so the
// organ is present but it does not decompose, so it does not transduce.
// "No organ" and "an organ that only fingerprints" must not look alike.
let rf: Int = realizer_register("fingerprint", "fingerprint_realizer")
let _c: Int = check(rf, "the symbol resolves, so registration succeeds")
let mf: Manifold = transduce("x", "fingerprint")
let _c: Int = check(1 - manifold_is(mf), "a single vector is not a transduction")
println("the one-part case is a size-one manifold, not a bare vector")
let g1: Geometry = geometry_new(3)
let s1: Int = geometry_set(g1, 0, 5.0)
let ms: Manifold = manifold_single("level", "scalar", g1)
let _c: Int = check(manifold_is(ms), "manifold_single yields a real Manifold")
let _c: Int = check(eq_int(manifold_size(ms), 1), "of size one — visibly degenerate, not hidden")
let fg: Int = geometry_free(g1)
let fs: Int = manifold_free(ms)
println("no organ is still reported as no organ")
let me: Manifold = transduce("anything", "echolocation")
let _c: Int = check(1 - manifold_is(me), "no realizer means no manifold, not a fake one")
let fm: Int = manifold_free(m)
// Reaching here means nothing called exit(1) along the way.
println("")
println("all checks passed")
}
+51
View File
@@ -0,0 +1,51 @@
#!/bin/bash
# build_vindex_bench.sh — build the vindex_bench oracle/proof harness, with
# the real ggml + hand-rolled-Metal batch-cosine strategies on Darwin and a
# zero-dependency CPU-only stub everywhere else. Mirrors the two-step recipe
# documented in vindex_bench.c's own header comment; this script exists so
# that recipe is one command, not a copy-pasted paragraph.
#
# Darwin build links FOUR strategy translation units:
# eg_cosine_batch.c — the Factory (always)
# eg_cosine_batch_strategy_cpu.c — universal fallback (always)
# eg_cosine_batch_strategy_ggml.c — ggml + dynamic Metal backend plugin
# eg_cosine_batch_strategy_metal_hand.m — PR #114's original hand-rolled
# Metal shader, preserved as one
# selectable strategy
# plus -DEG_HAVE_STRATEGY_GGML -DEG_HAVE_STRATEGY_METAL_HAND so the Factory
# (and vindex_bench.c's own direct strategy comparison) knows both exist.
#
# ggml is resolved via `brew --prefix ggml` when available (portable across
# Intel /usr/local and Apple Silicon /opt/homebrew installs), falling back to
# /opt/homebrew if brew isn't on PATH. Override with GGML_PREFIX=... env var.
#
# Usage: ./build_vindex_bench.sh [output_path]
set -euo pipefail
cd "$(dirname "$0")"
OUT="${1:-./vindex_bench}"
CC="${CC:-cc}"
if [ "$(uname -s)" = "Darwin" ]; then
GGML_PREFIX="${GGML_PREFIX:-$(brew --prefix ggml 2>/dev/null || echo /opt/homebrew)}"
echo "== Darwin: building with the ggml + hand-rolled-Metal strategies (ggml prefix: $GGML_PREFIX) =="
"$CC" -O2 -std=c11 -x objective-c \
-c eg_cosine_batch_strategy_metal_hand.m -o /tmp/eg_cosine_batch_strategy_metal_hand.o \
-framework Metal -framework Foundation
"$CC" -O2 -std=c11 -DEG_HAVE_STRATEGY_GGML -DEG_HAVE_STRATEGY_METAL_HAND -w \
-I"$GGML_PREFIX/include" \
vindex_bench.c engram_vindex.c \
eg_cosine_batch.c eg_cosine_batch_strategy_cpu.c eg_cosine_batch_strategy_ggml.c \
/tmp/eg_cosine_batch_strategy_metal_hand.o \
-L"$GGML_PREFIX/lib" -lggml -lggml-base \
-Wl,-rpath,"$GGML_PREFIX/lib" \
-lm -framework Metal -framework Foundation -o "$OUT"
else
echo "== non-Darwin: building with the CPU-only fallback strategy (no ggml, no Metal) =="
"$CC" -O2 -std=c11 -w vindex_bench.c engram_vindex.c \
eg_cosine_batch.c eg_cosine_batch_strategy_cpu.c \
-lm -o "$OUT"
fi
echo "built: $OUT"
+121
View File
@@ -0,0 +1,121 @@
/* eg_cosine_batch.c — the Factory. Implements the stable public interface
* declared in eg_cosine_batch.h by selecting ONE concrete
* EgCosineBatchStrategy (eg_cosine_batch_strategy.h) and dispatching every
* call to it. This is the ONLY file that branches on EG_HAVE_STRATEGY_*
* (build-time: which strategy .c/.m files were actually compiled in for
* this platform) call sites never see those macros.
*
* Selection is lazy (first call) and cached mirrors the lazy-init caching
* every individual strategy already does internally, so there is no added
* per-call cost after the first.
*
* Selection mechanism (env var + build-time + runtime capability probe, all
* three, exactly as directed):
* - BUILD-TIME decides which strategies exist to choose from at all: a
* Darwin build compiles+links the ggml strategy and the hand-rolled
* Metal strategy (EG_HAVE_STRATEGY_GGML / EG_HAVE_STRATEGY_METAL_HAND
* both defined); a non-Darwin build compiles neither, matching PR #114's
* original Linux behavior exactly (CPU-fallback only, no Objective-C
* compiler or Metal frameworks required).
* - RUNTIME CAPABILITY PROBE: each candidate strategy's own available()
* does the real, cheap-after-first-call check (device present, backend
* plugin loaded, pipeline compiled) never assumed from build-time
* alone. A build that HAS the ggml strategy compiled in but is running
* on hardware/software where it can't actually initialize (backend
* plugin missing, no GPU) correctly falls through to the next candidate.
* - ENV VAR gives explicit, debuggable override for either axis:
* EL_COSINE_BATCH_STRATEGY = "ggml" | "metal" | "cpu" | unset/"auto"
* forces a specific strategy (falling back to cpu if the forced one
* isn't actually available), or leaves the default auto-preference
* order in place.
* EL_METAL_COSINE = 0/n/N/f/F (back-compat with PR #114's vindex_bench
* gate) disables ALL GPU-backed strategies outright, same as before.
*
* DEFAULT preference order when nothing is forced: ggml, then hand-rolled
* Metal, then CPU fallback first candidate whose available() reports true
* wins. This is what makes "stop hand-rolling GPU kernels, use ggml" real
* rather than nominal: ggml is what actually runs by default on this
* machine today (see the PR body for the measured numbers backing that).
*/
#include "eg_cosine_batch.h"
#include "eg_cosine_batch_strategy.h"
#include <stdlib.h>
#include <string.h>
static bool g_selected = false;
static const EgCosineBatchStrategy* g_active = NULL;
static bool eg_env_truthy_off(const char* v) {
return v && (v[0]=='0' || v[0]=='n' || v[0]=='N' || v[0]=='f' || v[0]=='F');
}
static const EgCosineBatchStrategy* eg_select_strategy(void) {
if (g_selected) return g_active;
g_selected = true;
const EgCosineBatchStrategy* cpu = eg_cosine_batch_strategy_cpu();
const char* force = getenv("EL_COSINE_BATCH_STRATEGY");
const char* legacy_off = getenv("EL_METAL_COSINE");
if (eg_env_truthy_off(legacy_off)) { g_active = cpu; return g_active; }
if (force && strcmp(force, "cpu") == 0) { g_active = cpu; return g_active; }
if (force && strcmp(force, "ggml") == 0) {
#ifdef EG_HAVE_STRATEGY_GGML
const EgCosineBatchStrategy* s = eg_cosine_batch_strategy_ggml();
if (s->available()) { g_active = s; return g_active; }
#endif
g_active = cpu; return g_active;
}
if (force && strcmp(force, "metal") == 0) {
#ifdef EG_HAVE_STRATEGY_METAL_HAND
const EgCosineBatchStrategy* s = eg_cosine_batch_strategy_metal_hand();
if (s->available()) { g_active = s; return g_active; }
#endif
g_active = cpu; return g_active;
}
/* auto (unset, or any other value): ggml -> metal-hand -> cpu, first
* available wins. */
#ifdef EG_HAVE_STRATEGY_GGML
{
const EgCosineBatchStrategy* s = eg_cosine_batch_strategy_ggml();
if (s->available()) { g_active = s; return g_active; }
}
#endif
#ifdef EG_HAVE_STRATEGY_METAL_HAND
{
const EgCosineBatchStrategy* s = eg_cosine_batch_strategy_metal_hand();
if (s->available()) { g_active = s; return g_active; }
}
#endif
g_active = cpu;
return g_active;
}
bool eg_cosine_batch_available(void) {
return eg_select_strategy()->available();
}
const char* eg_cosine_batch_strategy_name(void) {
return eg_select_strategy()->name;
}
bool eg_cosine_batch(const float* query, int32_t qdim,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores) {
return eg_select_strategy()->batch(query, qdim, node_ptrs, node_dims, n, out_scores);
}
bool eg_cosine_batch_multi(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores) {
return eg_select_strategy()->batch_multi(queries, qdim, nq, node_ptrs, node_dims, n, out_scores);
}
+106
View File
@@ -0,0 +1,106 @@
/* eg_cosine_batch.h — stable Adapter interface over batch-cosine-similarity
* BACKEND STRATEGIES. This header is the ONE thing call sites (el_runtime.c,
* vindex_bench.c, ...) talk to. Plain C11, safe to #include on every
* platform the symbols declared here always exist and always link,
* regardless of what backend actually runs underneath. Zero #ifdef at call
* sites: which concrete strategy executes (ggml/Metal, hand-rolled Metal, or
* the always-false CPU fallback) is resolved once, lazily, inside
* eg_cosine_batch.c's factory see eg_cosine_batch_strategy.h for that.
*
* This supersedes eg_metal_cosine.h (PR #114's single hand-rolled-Metal-only
* bridge). The contract is UNCHANGED same shapes, same sentinel, same
* never-partial guarantee, same "caller must always be prepared to fall back
* to its own scalar per-node loop" rule — only the name changed, because the
* thing behind it is no longer "the Metal bridge," it is "whichever batch-
* cosine strategy the factory picked." eg_metal_cosine.h's original doc
* comments (byte-for-byte, this file is the direct descendant) are preserved
* below since they remain the precise spec any strategy must honor.
*
* On ANY failure at ANY step no compute device, compile/init error, alloc
* failure, bad args every function here returns false and writes nothing.
* Out-params are either fully populated or left completely untouched, never
* partial. Callers MUST always be prepared to fall back to their own scalar
* per-node CPU loop unconditionally. These functions must never crash, throw,
* or hang the calling process several call sites run inside a long-lived
* daemon's request-handling hot path.
*/
#ifndef EG_COSINE_BATCH_H
#define EG_COSINE_BATCH_H
#include <stdint.h>
#include <stdbool.h>
#ifdef __cplusplus
extern "C" {
#endif
/* Batched cosine similarity: one query vector against `n` node vectors.
*
* query qdim floats, the query embedding. Raw/unnormalized.
* qdim query dimensionality (e.g. 768 for nomic-embed-text).
* node_ptrs array of n pointers, node_ptrs[i] pointing at a (possibly
* differently-owned, possibly NULL) float vector for node i.
* NOT required to be contiguous every strategy performs the
* gather into a packed row-major matrix internally, exactly
* mirroring how EngramNode.emb is one malloc per node.
* node_dims array of n ints, node_dims[i] = that node's real emb_dim
* (0 or mismatched vs qdim => that node scores -2.0, matching
* eg_cosine's null/dim-mismatch/zero-norm sentinel exactly).
* n number of nodes.
* out_scores caller-owned array of n doubles; out_scores[i] is filled
* with the cosine similarity of node i against query, or
* -2.0 for a null/dim-mismatched/zero-norm node bit-for-bit
* the same contract as eg_cosine(node_ptrs[i], query, qdim).
*
* Returns true iff a real backend strategy ran and out_scores was fully
* populated. Returns false (out_scores left untouched) on ANY failure or
* unavailability no compute device, compile/init failure, allocation
* failure, n<=0, qdim<=0, null query/node_ptrs/node_dims/out_scores.
*/
bool eg_cosine_batch(const float* query, int32_t qdim,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores);
/* True iff a real (non-CPU-fallback) strategy is available right now (cheap
* after the first call cached). Purely informational (e.g. a startup log
* line or /api/stats field); callers should still treat a false return from
* eg_cosine_batch()/eg_cosine_batch_multi() itself as the authoritative
* fallback signal, not this function. */
bool eg_cosine_batch_available(void);
/* Which concrete strategy is currently selected — "ggml", "metal-hand",
* or "cpu-fallback". Purely informational/diagnostic, same spirit as
* eg_cosine_batch_available(). Never NULL. */
const char* eg_cosine_batch_strategy_name(void);
/* Multi-query batched cosine: nq query vectors against the SAME n node
* vectors, in one call. A real strategy uploads/prepares the node population
* ONCE and reuses it for every query, instead of nq separate
* eg_cosine_batch() calls each paying the full gather+upload cost PR #114
* measured this necessary: at N~=13.7k/dim=768, repeating the single-query
* call per query was slower than the CPU baseline; batching queries together
* is what makes a GPU-backed path a real win at this shape. Use this
* whenever multiple queries will run against an unchanged (or
* rarely-changing) node population; use eg_cosine_batch() for a genuinely
* one-off comparison.
*
* queries nq*qdim floats, row-major (query i at queries+i*qdim).
* out_scores caller-owned nq*n doubles, row-major
* (out_scores[i*n+j] = cosine(queries[i], node j)), same
* -2.0 sentinel semantics as eg_cosine_batch().
*
* Returns true iff a real strategy ran and out_scores was fully populated
* (all nq*n entries); false (untouched) on any failure/unavailability. */
bool eg_cosine_batch_multi(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores);
#ifdef __cplusplus
}
#endif
#endif /* EG_COSINE_BATCH_H */
+156
View File
@@ -0,0 +1,156 @@
/* eg_cosine_batch.metal — batched cosine similarity, one query vs N node vectors.
*
* GPU-shaped counterpart to eg_cosine() in el_runtime.c: same math, same
* dim-mismatch/zero-norm sentinel (-2.0), applied to N independent rows in
* parallel instead of one pair at a time in a CPU loop.
*
* Semantics MUST match eg_cosine() exactly:
* - inputs are raw, UNNORMALIZED vectors (nomic-embed-text magnitudes are
* not 1.0) this kernel computes the full dot/(|a|*|b|) cosine, not a
* plain dot product.
* - a node whose declared dim differs from the query dim, or whose norm is
* zero, scores exactly -2.0 (below any valid cosine in [-1,1]), so a
* caller doing `if (score > threshold)` behaves identically whether the
* scalar or the batched path filled the array.
*
* Precision: Apple GPUs do not support double in Metal Shading Language
* everything here is float32. eg_cosine accumulates in CPU double, but its
* *inputs* are float32 embeddings, so the achievable precision ceiling is
* bounded by the input data regardless of accumulator width. To keep the
* float32 reduction from drifting relative to the double-accumulated CPU
* result across dim=768 terms, each thread accumulates with 4 independent
* partial sums (unrolled) rather than one running scalar the same
* error-reduction trick already used by the CPU brute-force loop in
* vindex_bench.c. The measured float-vs-double delta is reported in the PR
* description; this is not assumed to be "close enough" without measurement.
*/
#include <metal_stdlib>
using namespace metal;
/* Per-dispatch invariants. `dim` is the query's dimensionality — the
* dimensionality every comparable node vector must match. */
struct EgCosineParams {
uint n; /* number of node rows */
uint dim; /* vector width (both query and node rows are `dim` wide in
* the packed buffer; node_dims[] carries each node's REAL
* embedded dim for the mismatch check) */
};
/* One thread per node row. node_matrix is n*dim floats, row-major, packed at
* `dim` stride regardless of a row's real dim (the CPU side zero-pads or
* skips packing rows that don't match see eg_cosine_batch_metal in
* eg_metal_cosine.m for the exact packing contract). node_dims[i] is the
* node's true emb_dim, used only for the mismatch sentinel never used to
* index, since every row is packed at uniform `dim` stride. */
kernel void eg_cosine_batch_kernel(
device const float* query [[buffer(0)]],
device const float* node_matrix [[buffer(1)]],
device const int* node_dims [[buffer(2)]],
constant EgCosineParams& p [[buffer(3)]],
device float* out_scores [[buffer(4)]],
uint gid [[thread_position_in_grid]])
{
if (gid >= p.n) return;
if (node_dims[gid] != int(p.dim)) {
out_scores[gid] = -2.0f;
return;
}
device const float* row = node_matrix + (uint64_t)gid * (uint64_t)p.dim;
/* 4-way partial accumulation — same shape as vindex_bench.c's brute_topk
* unroll, done here for float32 accuracy rather than raw throughput. */
float dot0 = 0.0f, dot1 = 0.0f, dot2 = 0.0f, dot3 = 0.0f;
float na0 = 0.0f, na1 = 0.0f, na2 = 0.0f, na3 = 0.0f;
float nb0 = 0.0f, nb1 = 0.0f, nb2 = 0.0f, nb3 = 0.0f;
uint d = 0;
uint dim4 = p.dim & ~3u;
for (; d < dim4; d += 4) {
float a0 = row[d], b0 = query[d];
float a1 = row[d+1], b1 = query[d+1];
float a2 = row[d+2], b2 = query[d+2];
float a3 = row[d+3], b3 = query[d+3];
dot0 += a0*b0; dot1 += a1*b1; dot2 += a2*b2; dot3 += a3*b3;
na0 += a0*a0; na1 += a1*a1; na2 += a2*a2; na3 += a3*a3;
nb0 += b0*b0; nb1 += b1*b1; nb2 += b2*b2; nb3 += b3*b3;
}
float dot = (dot0 + dot1) + (dot2 + dot3);
float na = (na0 + na1) + (na2 + na3);
float nb = (nb0 + nb1) + (nb2 + nb3);
for (; d < p.dim; d++) {
float a = row[d], b = query[d];
dot += a*b; na += a*a; nb += b*b;
}
if (na <= 0.0f || nb <= 0.0f) {
out_scores[gid] = -2.0f;
return;
}
out_scores[gid] = dot / sqrt(na * nb);
}
/* ── multi-query variant ──────────────────────────────────────────────────
* Same per-pair math as eg_cosine_batch_kernel, but amortizes ONE upload of
* node_matrix (the expensive part at real store size 13k*768 floats is
* ~42MB) across `nq` queries instead of re-uploading it once per query.
* Measured need: a naive one-query-at-a-time loop calling the single-query
* kernel nq times was SLOWER than the CPU oracle at N13.7k (re-gather +
* re-upload dominated the actual compute) this is the fix, not a
* hypothetical optimization.
*
* 2D grid: x = node index [0,n), y = query index [0,nq). out_scores is
* nq*n, row-major by query (out_scores[qid*n + nid]). */
struct EgCosineMultiParams { uint n; uint dim; uint nq; };
kernel void eg_cosine_batch_multi_kernel(
device const float* queries [[buffer(0)]], /* nq*dim */
device const float* node_matrix [[buffer(1)]], /* n*dim */
device const int* node_dims [[buffer(2)]], /* n */
constant EgCosineMultiParams& p [[buffer(3)]],
device float* out_scores [[buffer(4)]], /* nq*n */
uint2 gid [[thread_position_in_grid]])
{
uint nid = gid.x, qid = gid.y;
if (nid >= p.n || qid >= p.nq) return;
uint64_t out_idx = (uint64_t)qid * (uint64_t)p.n + (uint64_t)nid;
if (node_dims[nid] != int(p.dim)) {
out_scores[out_idx] = -2.0f;
return;
}
device const float* row = node_matrix + (uint64_t)nid * (uint64_t)p.dim;
device const float* query = queries + (uint64_t)qid * (uint64_t)p.dim;
float dot0 = 0.0f, dot1 = 0.0f, dot2 = 0.0f, dot3 = 0.0f;
float na0 = 0.0f, na1 = 0.0f, na2 = 0.0f, na3 = 0.0f;
float nb0 = 0.0f, nb1 = 0.0f, nb2 = 0.0f, nb3 = 0.0f;
uint d = 0;
uint dim4 = p.dim & ~3u;
for (; d < dim4; d += 4) {
float a0 = row[d], b0 = query[d];
float a1 = row[d+1], b1 = query[d+1];
float a2 = row[d+2], b2 = query[d+2];
float a3 = row[d+3], b3 = query[d+3];
dot0 += a0*b0; dot1 += a1*b1; dot2 += a2*b2; dot3 += a3*b3;
na0 += a0*a0; na1 += a1*a1; na2 += a2*a2; na3 += a3*a3;
nb0 += b0*b0; nb1 += b1*b1; nb2 += b2*b2; nb3 += b3*b3;
}
float dot = (dot0 + dot1) + (dot2 + dot3);
float na = (na0 + na1) + (na2 + na3);
float nb = (nb0 + nb1) + (nb2 + nb3);
for (; d < p.dim; d++) {
float a = row[d], b = query[d];
dot += a*b; na += a*a; nb += b*b;
}
if (na <= 0.0f || nb <= 0.0f) {
out_scores[out_idx] = -2.0f;
return;
}
out_scores[out_idx] = dot / sqrt(na * nb);
}
+82
View File
@@ -0,0 +1,82 @@
/* eg_cosine_batch_strategy.h — internal Strategy interface, NOT for call
* sites (they use eg_cosine_batch.h). Only eg_cosine_batch.c's factory and
* the concrete strategy implementation files include this.
*
* Each concrete strategy exposes exactly one getter returning a pointer to a
* static, immutable EgCosineBatchStrategy vtable. Which getters actually
* exist as linkable symbols is a BUILD-TIME concern (decided by
* build_vindex_bench.sh / the engram daemon's own build, via which .c/.m
* files get compiled per platform) gated by the EG_HAVE_STRATEGY_* macros
* below the factory in eg_cosine_batch.c is the ONLY place that branches
* on those macros. Call sites never see them; that's the whole point of the
* Adapter in eg_cosine_batch.h.
*
* Three concrete strategies exist:
* eg_cosine_batch_strategy_ggml() ggml + dynamically-loaded Metal
* backend plugin. Darwin only in
* this build; the default
* preferred strategy wherever
* available. EG_HAVE_STRATEGY_GGML.
* eg_cosine_batch_strategy_metal_hand() the original hand-rolled Metal
* compute shader from PR #114
* (eg_cosine_batch.metal),
* preserved verbatim as a
* selectable fallback strategy,
* not deleted. Darwin only.
* EG_HAVE_STRATEGY_METAL_HAND.
* eg_cosine_batch_strategy_cpu() universal always-false
* fallback. Always compiled, on
* every platform; this is what a
* non-Darwin build links
* exclusively (matching PR #114's
* eg_metal_cosine_stub.c), and
* what any platform falls back
* to when no real strategy is
* available at runtime.
*/
#ifndef EG_COSINE_BATCH_STRATEGY_H
#define EG_COSINE_BATCH_STRATEGY_H
#include <stdint.h>
#include <stdbool.h>
#ifdef __cplusplus
extern "C" {
#endif
typedef struct EgCosineBatchStrategy {
/* Stable, short, lowercase-hyphenated identifier — what
* eg_cosine_batch_strategy_name() surfaces. Never NULL. */
const char* name;
/* Cheap after the first call (lazy init, cached internally). Must never
* throw/crash/hang mirrors eg_cosine_batch_available()'s contract. */
bool (*available)(void);
/* Same shape/contract as eg_cosine_batch() in eg_cosine_batch.h. */
bool (*batch)(const float* query, int32_t qdim,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores);
/* Same shape/contract as eg_cosine_batch_multi() in eg_cosine_batch.h. */
bool (*batch_multi)(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores);
} EgCosineBatchStrategy;
#ifdef EG_HAVE_STRATEGY_GGML
const EgCosineBatchStrategy* eg_cosine_batch_strategy_ggml(void);
#endif
#ifdef EG_HAVE_STRATEGY_METAL_HAND
const EgCosineBatchStrategy* eg_cosine_batch_strategy_metal_hand(void);
#endif
/* Always declared/linked, on every platform/build. */
const EgCosineBatchStrategy* eg_cosine_batch_strategy_cpu(void);
#ifdef __cplusplus
}
#endif
#endif /* EG_COSINE_BATCH_STRATEGY_H */
@@ -0,0 +1,45 @@
/* eg_cosine_batch_strategy_cpu.c — plain-C, zero-dependency universal
* fallback strategy. Always returns false / unavailable. Direct descendant
* of PR #114's eg_metal_cosine_stub.c, generalized from "the Metal stub" to
* "the strategy vtable's universal fallback entry" now that multiple real
* strategies can exist.
*
* Always compiled, on every platform. On Darwin builds it is the last-resort
* strategy the factory falls back to when neither ggml nor the hand-rolled
* Metal strategy is available at runtime (no device, compile failure, ...).
* On non-Darwin builds it is the ONLY strategy compiled in at all no
* Objective-C, no Metal frameworks, no ggml/Metal backend plugin so
* eg_cosine_batch()/eg_cosine_batch_multi() always return false there and
* every call site's existing CPU fallback runs unconditionally, exactly as
* before this PR.
*/
#include "eg_cosine_batch_strategy.h"
static bool cpu_available(void) {
return false;
}
static bool cpu_batch(const float* query, int32_t qdim,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores) {
(void)query; (void)qdim; (void)node_ptrs; (void)node_dims; (void)n; (void)out_scores;
return false;
}
static bool cpu_batch_multi(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores) {
(void)queries; (void)qdim; (void)nq; (void)node_ptrs; (void)node_dims; (void)n; (void)out_scores;
return false;
}
static const EgCosineBatchStrategy g_cpu_strategy = {
.name = "cpu-fallback",
.available = cpu_available,
.batch = cpu_batch,
.batch_multi = cpu_batch_multi,
};
const EgCosineBatchStrategy* eg_cosine_batch_strategy_cpu(void) {
return &g_cpu_strategy;
}
@@ -0,0 +1,515 @@
/* eg_cosine_batch_strategy_ggml.c — the GGML Strategy, and the preferred
* default whenever it is available (see the factory's selection order in
* eg_cosine_batch.c).
*
* WHY: directive from Will Anderson stop hand-rolling GPU kernels, use a
* real, proven, permissively-licensed library instead. ggml (the compute
* library underneath llama.cpp, MIT licensed) is already installed on this
* machine as a standalone Homebrew package (`brew info ggml`), independent
* of llama.cpp itself. This file is a genuinely bounded COMPUTE UTILITY
* batch cosine-similarity math analogous to a VBD Accessor calling out to
* infrastructure. It is explicitly NOT the engram's reasoning/persistence
* core; using ggml here does not cross the "own the core" line, because
* batch cosine math is infrastructure, not the graph traversal / activation
* spreading / "thinking" that IS the core and stays 100% own-code.
*
* The real API shape (verified against the installed headers + a
* standalone probe program, not assumed from memory of other tensor
* libraries)
*
* ggml ships its CPU and Metal implementations as DYNAMICALLY LOADED PLUGIN
* .so files (confirmed by nm: `ggml_backend_metal_init` is NOT an exported
* symbol of libggml.dylib/libggml-base.dylib it exists ONLY inside
* libggml-metal.so under $(brew --prefix ggml)/libexec/). You cannot link
* `-lggml-metal`; you must go through ggml's backend REGISTRY:
*
* 1. ggml_backend_load_all_from_path(dir) dlopen()s every backend plugin
* .so found in `dir` and registers its device(s). We point this at
* $(brew --prefix ggml)/libexec (resolved once, at build+init time; see
* eg_ggml_backend_dir() below) rather than relying on
* ggml_backend_load_all()'s own default search heuristics, which are
* tuned for an installed llama.cpp-style app bundle layout, not an
* arbitrary `cc`-built binary invoked from an arbitrary cwd the exact
* same "must not silently fall back to CPU for reasons that have
* nothing to do with GPU availability" concern PR #114's hand-rolled
* bridge already documented for its own embedded-shader-source choice.
* 2. ggml_backend_dev_by_type(GGML_BACKEND_DEVICE_TYPE_GPU) find the
* registered Metal device.
* 3. ggml_backend_dev_init(dev, NULL) get a live ggml_backend_t.
* 4. Build a tiny ggml_context (no_alloc=true; it holds only tensor
* metadata, not data), declare 2D F32 tensors, ggml_mul_mat(nodes,
* query) ggml's documented convention: A is [k cols, n rows], B is
* [k cols, m rows] (transposed internally), result is [n cols, m rows]
* i.e. mul_mat(node_matrix[dim,n], query_matrix[dim,nq]) yields
* out[n,nq] where out[j*n+i] = dot(node_i, query_j). A row-major
* (dim,n) node matrix and a row-major (dim,nq) query matrix is EXACTLY
* the packed layout the hand-rolled Metal kernel already used one
* matmul replaces the whole per-row dot-product loop.
* 5. ggml_backend_alloc_ctx_tensors(ctx, backend) to actually allocate
* device buffers for those tensors, ggml_backend_tensor_set() to upload,
* ggml_backend_graph_compute() to run, ggml_backend_tensor_get() to
* read back.
*
* This exact sequence was verified end-to-end in a standalone probe (build
* it yourself: see the PR description) against a plain-C CPU dot product
* bit-for-bit correct within float rounding. Real numbers against the
* el_runtime.c CPU oracle are reported in the PR body via vindex_bench.
*
* ggml_mul_mat only computes the raw dot products it has no notion of
* "cosine" or of this codebase's -2.0 dim-mismatch/null/zero-norm sentinel.
* Per the adapter's directive: gather only VALID, uniform-dim rows into the
* packed matrix sent to the GPU (skipping null/mismatched rows entirely,
* rather than the hand-rolled kernel's zero-pad-and-sentinel-in-shader
* approach), then scatter -2.0 back for every row that was excluded same
* gather/scatter contract eg_cosine_batch.h documents. Norms (||node||,
* ||query||) are computed on the CPU host in the same pass that already
* touches every element to gather/convert essentially free using the
* same 4-way-partial-sum accumulation the hand-rolled kernel and the CPU
* oracle both use, so the float32 error profile stays comparable across all
* three strategies. Only the O(n*dim*nq) dot-product matmul the actual
* expensive part is offloaded to the GPU.
*
* Precision: the ne11<=8 chunking, and why it is not optional
*
* The claim in the first version of this file "ggml_mul_mat on F32 x F32
* inputs computes in F32 on the Metal backend" — is WRONG, and the 0.9933
* id-recall it shipped with (vs the hand-rolled kernel's 0.9997) was the
* symptom. ggml-metal has two F32xF32 matmul kernels and picks between them
* purely on ne11 (the number of B rows == our query count):
*
* ne11 <= 8 -> kernel_mul_mv_ext_f32_f32_* / kernel_mul_mv_f32_f32_*
* templated <float, float> genuine F32 accumulation.
* ne11 > 8 -> kernel_mul_mm_f32_f32, which is templated
* <half, half4x4, simdgroup_half8x8, half, half2x4,
* simdgroup_half8x8, ...> i.e. BOTH operands are narrowed
* to F16 and accumulated in simdgroup_half8x8 tiles, even
* though the tensors are GGML_TYPE_F32 on both sides.
*
* (Read it yourself, no guessing the kernel templates are literal strings
* in the shipped plugin:
* strings $(brew --prefix ggml)/libexec/libggml-metal.so \
* | grep -E 'host_name\("kernel_mul_m[mv]_f32_f32'
* and the runtime pick is visible with GGML_METAL_DEBUG-style logging as
* "compiling pipeline: base = 'kernel_mul_mm_f32_f32'".)
*
* The previous code issued ONE ggml_mul_mat with ne11 = nq (300 in the
* benchmark), landing squarely on the F16 mul_mm path. Measured on this
* machine (M4 Pro), n=13415 x dim=768 x nq=300, against a CPU double-
* accumulated oracle:
*
* ne11=300 (one mul_mat, the old code) : mean |Δdot| = 1.038e-05
* ne11=8 (chunked, this code) : mean |Δdot| = 3.863e-09
*
* a ~2700x reduction in dot-product error, which is exactly the gap that
* showed up as 0.9933-vs-0.9997 recall.
*
* ggml_mul_mat_set_prec(t, GGML_PREC_F32) does NOT fix this. It was tried:
* the error was bit-identical with and without it (1.038e-05 either way),
* because ggml-metal only consults the prec flag on paths that have an F32
* variant to switch to, and there is no F32-accumulating mul_mm kernel in
* this build to select. The ONLY lever from outside ggml is ne11.
*
* So: instead of one mul_mat with ne11=nq, we emit ceil(nq/8) mul_mats, each
* over an ne11<=8 ggml_view_2d slice of the same query tensor, all into ONE
* graph and ONE ggml_backend_graph_compute. The node matrix is still uploaded
* exactly once and still read by the GPU as one shared operand the whole
* point of batch_multi is preserved.
*
* The cost is real, and stated rather than buried. Timing the whole
* batch_multi() call (gather + norms + upload + GPU + scatter) on the real
* shape, median of 15 reps after a discarded warm-up, three separate runs:
*
* unchunked (old, F16 mm) : 13.19 / 13.35 / 14.42 ms -> ~0.044 ms/query
* chunked (this code) : 19.92 / 20.08 / 20.23 ms -> ~0.067 ms/query
* hand-rolled Metal : 17.74 / 17.88 / 17.99 ms -> ~0.060 ms/query
*
* So correctness here costs about +6.7ms per 300-query batch (~1.5x on this
* call), and leaves us ~12% behind the hand-rolled kernel instead of ~35%
* ahead of it. That is not free and should not be sold as free. The reason it
* cannot be recovered inside ggml: an fp32 matmul on Metal has to re-stream
* the whole node matrix once per <=8 queries (38 dispatches x ~41MB here),
* where the F16 mul_mm kernel tiles it in threadgroup memory and reads it far
* fewer times. ggml's Metal backend ships no fp32 TILED matmul, so on this
* backend "fast" and "fp32" are genuinely exclusive the hand-rolled kernel
* escapes the choice only because it is an fp32 kernel written for this one
* shape. Trading precision back for speed is a one-line env change; trading
* the other way was not available before this commit at all.
*
* 8 is not a magic number we invented it is ggml-metal's own mul_mm
* threshold, measured by sweeping ne11 and watching both the error and which
* pipeline ggml compiles (9 flips to mul_mm and the error jumps back to
* 1.0e-05 in the same step). EL_GGML_MULMAT_CHUNK overrides it: raise it to
* trade this precision back for throughput, or set it >= nq to reproduce the
* old single-mul_mat behaviour exactly. If a future ggml moves the threshold,
* the worst case is that we silently land back on mul_mm the same accuracy
* we shipped before, never a correctness break.
*
* Cold start: what is and is not ours to fix
*
* The ~7.8s first-call cost reported for the first version of this file is
* NOT this file re-initialising per call (init is, and always was, cached
* behind g_init_attempted below). It is Apple's Metal shader cache missing
* on ggml's embedded metallib ggml-metal ships ~650 kernels in one
* __ggml_metallib section, and the first newLibraryWithData of it on a given
* machine costs seconds ("ggml_metal_library_init: loaded in 7.670 sec")
* while the driver populates ~//C/com.apple.metal/. That cache is keyed on
* the library, not on our binary, and is shared across processes: the very
* next run of a DIFFERENT binary linking the same ggml reports
* "loaded in 0.009 sec". So it is a once-per-machine, per-ggml-version cost,
* not a per-process one, and nothing this file does can avoid it the
* hand-rolled strategy escapes it only because its shader is two small
* kernels instead of six hundred.
*
* The residual warm init IS ours to look at, and the answer there is "there
* was nothing much to win": ggml_backend_load_all_from_path() dlopens every
* plugin in the directory (three CPU micro-arch variants + BLAS + Metal) when
* we only ever use Metal, so we now load the single Metal plugin instead
* but measured warm that is 44.7-52.4ms against 46.9-58.9ms, i.e. the same
* number inside noise, because libggml-metal.so's own init dominates. Warm
* ggml init lands at 44-53ms, against 36-117ms for the hand-rolled strategy's
* device+pipeline setup. Cold start was never the real defect here; precision
* was.
*/
#include "eg_cosine_batch_strategy.h"
#include <ggml.h>
#include <ggml-backend.h>
#include <ggml-alloc.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <math.h>
/* ── lazy, one-time backend init, cached ─────────────────────────────────── */
static bool g_init_attempted = false;
static bool g_init_ok = false;
static ggml_backend_t g_backend = NULL;
/* Where to look for the dynamically-loaded backend plugin .so files.
* EL_GGML_BACKEND_PATH overrides for non-standard installs; otherwise we try
* the Homebrew opt-prefix symlink (stable across ggml point-version bumps
* $(brew --prefix ggml)/libexec confirmed to exist and contain
* libggml-metal.so / libggml-cpu-*.so / libggml-blas.so on this machine),
* falling back to ggml's own default search (ggml_backend_load_all()) in
* case a different install layout (e.g. a from-source build with a
* standard-prefix install) makes that succeed instead. */
static const char* eg_ggml_backend_dir(void) {
const char* s = getenv("EL_GGML_BACKEND_PATH");
if (s && *s) return s;
return "/opt/homebrew/opt/ggml/libexec";
}
/* "<dir>/libggml-metal.so" in a static buffer. Only ever called once, from
* eg_ggml_ensure_init(), before any thread could race it. */
static const char* eg_ggml_metal_plugin_path(const char* dir) {
static char buf[1024];
snprintf(buf, sizeof buf, "%s/libggml-metal.so", dir);
return buf;
}
/* Largest ne11 (query-batch rows per ggml_mul_mat) that keeps ggml-metal on
* its F32 mul_mv kernels instead of the F16-accumulating mul_mm kernel see
* the precision discussion in this file's header. EL_GGML_MULMAT_CHUNK
* overrides; a value <= 0 means "use the default". */
#define EG_GGML_MULMAT_CHUNK_DEFAULT 8
static int32_t eg_ggml_mulmat_chunk(void) {
static bool resolved = false;
static int32_t chunk = EG_GGML_MULMAT_CHUNK_DEFAULT;
if (!resolved) {
resolved = true;
const char* s = getenv("EL_GGML_MULMAT_CHUNK");
if (s && *s) {
long v = strtol(s, NULL, 10);
if (v > 0 && v <= INT32_MAX) chunk = (int32_t)v;
}
}
return chunk;
}
/* Which ggml device this strategy computes on. GPU (Metal) is the default
* because offloading is the architectural point the engram's own graph
* traversal and activation spreading are CPU work, and a "GPU" strategy that
* quietly saturates the CPU steals from them.
*
* ACCEL (ggml's BLAS/Accelerate plugin) is reachable here mainly as a
* portability fallback and a diagnostic, and it is documented as MEASURED AND
* REJECTED rather than as a recommendation. In an isolated probe that timed
* only ggml_backend_graph_compute, BLAS looked excellent 3.4-4.0ms for the
* 300-query batch at mean |Δdot| 1.5e-08, i.e. as fast as the old F16 path and
* far more accurate. End to end on the real store through vindex_bench it does
* not hold up: 0.191 ms/query at id-recall 0.9973, against 0.125-0.142 ms/query
* at 0.9987 for the Metal default. It is dominated on BOTH axes, because the
* isolated probe was not competing with the rest of the batch for the same CPU
* cores and the real call path is. Kept because a machine with no usable Metal
* device still wants a working ggml strategy not because it is faster. */
static enum ggml_backend_dev_type eg_ggml_device_type(void) {
const char* s = getenv("EL_GGML_DEVICE");
if (s && *s) {
if (strcmp(s, "accel") == 0) return GGML_BACKEND_DEVICE_TYPE_ACCEL;
if (strcmp(s, "cpu") == 0) return GGML_BACKEND_DEVICE_TYPE_CPU;
}
return GGML_BACKEND_DEVICE_TYPE_GPU;
}
static bool eg_ggml_ensure_init(void) {
if (g_init_attempted) return g_init_ok;
g_init_attempted = true;
const char* dir = eg_ggml_backend_dir();
const enum ggml_backend_dev_type want = eg_ggml_device_type();
/* Metal is the only backend this strategy uses by default, so load just
* that one plugin rather than dlopening the whole directory (three CPU
* micro-arch variants + BLAS + Metal here).
*
* Be honest about what this buys: almost nothing in wall time. Measured
* warm, three runs each load-everything 58.9/46.9/55.1ms, Metal-only
* 52.4/51.2/44.7ms. The cost is dominated by dlopening and initialising
* libggml-metal.so itself, not by the four plugins we skip, so the two
* overlap inside noise. It is kept because registering four device types
* we will never dispatch to is untidy and makes ggml_backend_dev_by_type
* ambiguous, not because it is a speedup do not cite it as one.
*
* ggml_backend_load() returns NULL for a missing or unloadable path,
* which simply falls through to the broader searches below; it is never
* fatal. Any non-default device needs the full directory scan to find
* its plugin, so skip the fast path there. */
if (want == GGML_BACKEND_DEVICE_TYPE_GPU)
ggml_backend_load(eg_ggml_metal_plugin_path(dir));
ggml_backend_dev_t dev = ggml_backend_dev_by_type(want);
if (!dev) {
/* Non-standard layout, a ggml built with a differently-named Metal
* plugin, or a non-default device: dlopen every plugin in `dir`. */
ggml_backend_load_all_from_path(dir);
dev = ggml_backend_dev_by_type(want);
}
if (!dev) {
/* Fall back to ggml's own default search heuristics only if the
* explicit path above found nothing avoids double-registering the
* same plugins (ggml does not dedupe two different paths that
* happen to resolve to the same files, e.g. our stable opt-prefix
* symlink vs. its own Cellar-relative guess) in the common case
* where the explicit path already worked. */
ggml_backend_load_all();
dev = ggml_backend_dev_by_type(want);
}
if (!dev) return false;
ggml_backend_t backend = ggml_backend_dev_init(dev, NULL);
if (!backend) return false;
g_backend = backend;
g_init_ok = true;
return true;
}
static bool ggml_strategy_available(void) {
return eg_ggml_ensure_init();
}
/* ── shared core: gather valid rows + norms, matmul, scatter ────────────── */
/* 4-way partial-sum squared-norm accumulation over `dim` floats — same shape
* as eg_cosine_batch.metal's per-thread accumulation and vindex_bench.c's
* CPU brute_topk unroll, kept consistent on purpose so the float32 error
* profile is comparable across all three strategies. */
static float eg_norm_sq_f32(const float* v, int32_t dim) {
float s0 = 0, s1 = 0, s2 = 0, s3 = 0;
int32_t d = 0, dim4 = dim & ~3;
for (; d < dim4; d += 4) {
s0 += v[d] * v[d]; s1 += v[d+1] * v[d+1];
s2 += v[d+2] * v[d+2]; s3 += v[d+3] * v[d+3];
}
float s = (s0 + s1) + (s2 + s3);
for (; d < dim; d++) s += v[d] * v[d];
return s;
}
/* Runs one ggml_mul_mat(node_matrix[dim,n_valid], query_matrix[dim,nq]) and
* combines it with CPU-computed norms into cosine scores, scattering into
* out_scores at ORIGINAL (ungathered) indices. out_scores must already be
* fully sized for n*nq (or n for the single-query case, nq=1) every entry
* gets written (valid rows get a real cosine, invalid rows get -2.0), so
* this never leaves a partial result. Returns false only on a genuine
* failure (alloc, compute) at which point out_scores is left as whatever a
* caller-supplied scratch buffer already contained callers here always
* pass a fresh buffer they discard on false, matching the adapter contract
* of "on failure, out_scores is treated as untouched" from the caller's
* point of view. */
static bool eg_ggml_run(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores) {
if (!queries || qdim <= 0 || nq <= 0 || !node_ptrs || !node_dims || n <= 0 || !out_scores)
return false;
if (!eg_ggml_ensure_init()) return false;
/* Pass 1 (CPU): gather valid rows (non-NULL ptr, dim == qdim) into a
* packed (dim, n_valid) row-major matrix, remembering the original index
* of each packed row, and compute each valid row's squared norm in the
* same pass. Rows excluded here get -2.0 scattered for every query
* below without ever touching the GPU. */
int32_t* valid_orig = (int32_t*)malloc((size_t)n * sizeof(int32_t));
float* node_norm_sq = (float*)malloc((size_t)n * sizeof(float)); /* indexed by packed position */
float* node_matrix = NULL;
if (!valid_orig || !node_norm_sq) { free(valid_orig); free(node_norm_sq); return false; }
int32_t n_valid = 0;
for (int32_t i = 0; i < n; i++) {
if (node_ptrs[i] && node_dims[i] == qdim) n_valid++;
}
if (n_valid > 0) {
node_matrix = (float*)malloc((size_t)n_valid * (size_t)qdim * sizeof(float));
if (!node_matrix) { free(valid_orig); free(node_norm_sq); return false; }
int32_t w = 0;
for (int32_t i = 0; i < n; i++) {
if (!node_ptrs[i] || node_dims[i] != qdim) continue;
memcpy(node_matrix + (size_t)w * qdim, node_ptrs[i], (size_t)qdim * sizeof(float));
node_norm_sq[w] = eg_norm_sq_f32(node_ptrs[i], qdim);
valid_orig[w] = i;
w++;
}
}
/* Query norms — nq is typically small (1 or the size of one batch of
* comparison queries), so this loop is cheap regardless. */
float* q_norm_sq = (float*)malloc((size_t)nq * sizeof(float));
if (!q_norm_sq) { free(valid_orig); free(node_norm_sq); free(node_matrix); return false; }
for (int32_t j = 0; j < nq; j++) q_norm_sq[j] = eg_norm_sq_f32(queries + (size_t)j * qdim, qdim);
/* Nothing valid to compare against: every output is -2.0. Still a fully
* and correctly populated result no GPU dispatch was needed to know
* that. */
if (n_valid == 0) {
for (size_t k = 0; k < (size_t)n * (size_t)nq; k++) out_scores[k] = -2.0;
free(valid_orig); free(node_norm_sq); free(node_matrix); free(q_norm_sq);
return true;
}
/* Pass 2 (GPU via ggml): dot[j*n_valid + i] = dot(node_i, query_j),
* computed as ceil(nq/chunk) separate ggml_mul_mat ops over ne11<=chunk
* ggml_view_2d slices of ONE query tensor, all expanded into ONE graph
* and run by ONE ggml_backend_graph_compute. Chunking is what keeps
* ggml-metal on its F32 mul_mv kernels rather than the F16-accumulating
* mul_mm kernel (see this file's header); sharing one graph and one
* t_nodes tensor is what keeps the node matrix uploaded exactly once,
* which is the entire reason batch_multi exists. */
const int32_t chunk = eg_ggml_mulmat_chunk();
const int32_t ngroups = (nq + chunk - 1) / chunk;
/* Tensors held by the context: t_nodes, t_query, plus one view and one
* mul_mat result per group. The graph holds at most one node per view and
* one per mul_mat. Slack on both so a ggml that bookkeeps slightly
* differently cannot silently overflow the arena. */
const size_t n_tensors = (size_t)2 * (size_t)ngroups + 8;
const size_t graph_size = (size_t)2 * (size_t)ngroups + 16;
struct ggml_init_params gp = {
.mem_size = ggml_tensor_overhead() * n_tensors
+ ggml_graph_overhead_custom(graph_size, false),
.mem_buffer = NULL,
.no_alloc = true,
};
struct ggml_context* ctx = ggml_init(gp);
if (!ctx) { free(valid_orig); free(node_norm_sq); free(node_matrix); free(q_norm_sq); return false; }
struct ggml_tensor** t_dots = (struct ggml_tensor**)malloc((size_t)ngroups * sizeof(*t_dots));
if (!t_dots) { ggml_free(ctx); free(valid_orig); free(node_norm_sq); free(node_matrix); free(q_norm_sq); return false; }
struct ggml_tensor* t_nodes = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, qdim, n_valid);
struct ggml_tensor* t_query = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, qdim, nq);
struct ggml_cgraph* gf = t_nodes && t_query
? ggml_new_graph_custom(ctx, graph_size, false) : NULL;
if (!gf) { free(t_dots); ggml_free(ctx); free(valid_orig); free(node_norm_sq); free(node_matrix); free(q_norm_sq); return false; }
bool built = true;
for (int32_t g = 0; g < ngroups; g++) {
const int32_t start = g * chunk;
const int32_t count = (start + chunk <= nq) ? chunk : (nq - start);
struct ggml_tensor* t_qv = ggml_view_2d(ctx, t_query, qdim, count,
t_query->nb[1],
(size_t)start * t_query->nb[1]);
t_dots[g] = t_qv ? ggml_mul_mat(ctx, t_nodes, t_qv) : NULL;
if (!t_dots[g]) { built = false; break; }
ggml_build_forward_expand(gf, t_dots[g]);
}
if (!built) { free(t_dots); ggml_free(ctx); free(valid_orig); free(node_norm_sq); free(node_matrix); free(q_norm_sq); return false; }
struct ggml_backend_buffer* buf = ggml_backend_alloc_ctx_tensors(ctx, g_backend);
if (!buf) { free(t_dots); ggml_free(ctx); free(valid_orig); free(node_norm_sq); free(node_matrix); free(q_norm_sq); return false; }
ggml_backend_tensor_set(t_nodes, node_matrix, 0, (size_t)n_valid * qdim * sizeof(float));
ggml_backend_tensor_set(t_query, queries, 0, (size_t)nq * qdim * sizeof(float));
free(node_matrix); /* uploaded; the packed CPU copy is no longer needed */
enum ggml_status st = ggml_backend_graph_compute(g_backend, gf);
if (st != GGML_STATUS_SUCCESS) {
free(t_dots); ggml_backend_buffer_free(buf); ggml_free(ctx);
free(valid_orig); free(node_norm_sq); free(q_norm_sq);
return false;
}
/* Each group's result is [n_valid, count] contiguous, so reading group g
* into dot + start*n_valid reconstructs exactly the same flat
* dot[j*n_valid + w] layout a single ne11=nq mul_mat would have produced
* Pass 3 below is unchanged by the chunking. */
float* dot = (float*)malloc((size_t)n_valid * (size_t)nq * sizeof(float));
if (!dot) { free(t_dots); ggml_backend_buffer_free(buf); ggml_free(ctx); free(valid_orig); free(node_norm_sq); free(q_norm_sq); return false; }
for (int32_t g = 0; g < ngroups; g++) {
const int32_t start = g * chunk;
const int32_t count = (start + chunk <= nq) ? chunk : (nq - start);
ggml_backend_tensor_get(t_dots[g], dot + (size_t)start * n_valid, 0,
(size_t)count * (size_t)n_valid * sizeof(float));
}
free(t_dots);
/* Pass 3 (CPU): combine dot/(||a||*||b||) per (query,node) pair, scatter
* into out_scores at ORIGINAL node indices; every excluded row gets
* -2.0 for every query. out_scores is fully populated either way. */
for (int32_t j = 0; j < nq; j++) {
double* orow = out_scores + (size_t)j * n;
for (int32_t i = 0; i < n; i++) orow[i] = -2.0; /* default: excluded */
for (int32_t w = 0; w < n_valid; w++) {
float na = node_norm_sq[w], nb = q_norm_sq[j];
int32_t oi = valid_orig[w];
if (na <= 0.0f || nb <= 0.0f) { orow[oi] = -2.0; continue; }
float d = dot[(size_t)j * n_valid + w];
orow[oi] = (double)(d / sqrtf(na * nb));
}
}
free(dot);
ggml_backend_buffer_free(buf);
ggml_free(ctx);
free(valid_orig); free(node_norm_sq); free(q_norm_sq);
return true;
}
static bool ggml_strategy_batch(const float* query, int32_t qdim,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores) {
if (!query || qdim <= 0 || !node_ptrs || !node_dims || n <= 0 || !out_scores) return false;
/* out_scores here is n doubles (nq=1); eg_ggml_run writes n*nq = n of
* them, laid out identically to the single-query contract. */
return eg_ggml_run(query, qdim, 1, node_ptrs, node_dims, n, out_scores);
}
static bool ggml_strategy_batch_multi(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs, const int32_t* node_dims,
int32_t n, double* out_scores) {
return eg_ggml_run(queries, qdim, nq, node_ptrs, node_dims, n, out_scores);
}
static const EgCosineBatchStrategy g_ggml_strategy = {
.name = "ggml",
.available = ggml_strategy_available,
.batch = ggml_strategy_batch,
.batch_multi = ggml_strategy_batch_multi,
};
const EgCosineBatchStrategy* eg_cosine_batch_strategy_ggml(void) {
return &g_ggml_strategy;
}
@@ -0,0 +1,358 @@
/* eg_cosine_batch_strategy_metal_hand.m — the HAND-ROLLED-METAL Strategy.
*
* This is PR #114's original Objective-C bridge (formerly eg_metal_cosine.m)
* exposing the hand-written Metal compute shader (eg_cosine_batch.metal) as
* one concrete EgCosineBatchStrategy. It is preserved here almost verbatim
* real, carefully verified work, not discarded now living behind the
* Adapter/Strategy/Factory restructuring (see eg_cosine_batch.h and
* eg_cosine_batch_strategy.h) alongside the new ggml-Metal strategy
* (eg_cosine_batch_strategy_ggml.c) and the universal CPU fallback
* (eg_cosine_batch_strategy_cpu.c). The factory in eg_cosine_batch.c prefers
* ggml by default when both are available; this strategy remains selectable
* via EL_COSINE_BATCH_STRATEGY=metal, and is what the factory falls back to
* if ggml's backend plugin fails to load/init for any reason.
*
* Apple-only (Metal has no other platform). This file is excluded from the
* build entirely on non-Darwin see build_vindex_bench.sh, which only
* compiles/links this file and defines EG_HAVE_STRATEGY_METAL_HAND when
* `uname` is Darwin. On Linux the factory never sees this strategy at all
* callers must always be prepared for the "no real strategy available"
* fallback via the CPU strategy, which is also exactly what happens here on
* Apple hardware with no usable GPU.
*
* Design (unchanged from PR #114):
* - Device/queue/pipeline are created lazily, once, and cached in static
* globals every call after the first only allocates buffers + submits.
* - The Metal shader source is embedded as a C string literal (kMetalSrc
* below) rather than loaded from a file at runtime or shipped as a
* precompiled .metallib. Chosen over newLibraryWithFile: /a .metallib
* because the engram binary can be invoked from an arbitrary working
* directory (launchd job, nsbx sandbox, CI) and a file-path shader would
* be one relocation away from silently falling back to CPU for reasons
* that have nothing to do with Metal availability. Embedding costs one
* runtime shader compile (~tens of ms) on first use, amortized over the
* process lifetime, in exchange for a genuinely self-contained binary.
* Source of truth for review/tooling is eg_cosine_batch.metal this
* string MUST be kept byte-identical to that file (a comment marks both
* ends of the copy).
* - Buffers use MTLResourceStorageModeShared: on Apple Silicon's unified
* memory, CPU and GPU read the same physical pages, so filling a buffer
* is a plain memcpy and there is no separate "upload" step.
* - ANY failure at ANY step (no device, pipeline compile error, buffer
* allocation failure, bad args) returns false and leaves out_scores
* untouched. This function is called from the request-handling hot path
* of a long-lived daemon it must never throw, crash, or hang it.
*/
#import <Foundation/Foundation.h>
#import <Metal/Metal.h>
#include "eg_cosine_batch_strategy.h"
#include <string.h>
#include <stdlib.h>
/* ── BEGIN embedded shader source (keep in sync with eg_cosine_batch.metal) ── */
static const char* kEgCosineBatchMetalSrc =
"#include <metal_stdlib>\n"
"using namespace metal;\n"
"struct EgCosineParams { uint n; uint dim; };\n"
"kernel void eg_cosine_batch_kernel(\n"
" device const float* query [[buffer(0)]],\n"
" device const float* node_matrix [[buffer(1)]],\n"
" device const int* node_dims [[buffer(2)]],\n"
" constant EgCosineParams& p [[buffer(3)]],\n"
" device float* out_scores [[buffer(4)]],\n"
" uint gid [[thread_position_in_grid]])\n"
"{\n"
" if (gid >= p.n) return;\n"
" if (node_dims[gid] != int(p.dim)) { out_scores[gid] = -2.0f; return; }\n"
" device const float* row = node_matrix + (uint64_t)gid * (uint64_t)p.dim;\n"
" float dot0 = 0.0f, dot1 = 0.0f, dot2 = 0.0f, dot3 = 0.0f;\n"
" float na0 = 0.0f, na1 = 0.0f, na2 = 0.0f, na3 = 0.0f;\n"
" float nb0 = 0.0f, nb1 = 0.0f, nb2 = 0.0f, nb3 = 0.0f;\n"
" uint d = 0;\n"
" uint dim4 = p.dim & ~3u;\n"
" for (; d < dim4; d += 4) {\n"
" float a0 = row[d], b0 = query[d];\n"
" float a1 = row[d+1], b1 = query[d+1];\n"
" float a2 = row[d+2], b2 = query[d+2];\n"
" float a3 = row[d+3], b3 = query[d+3];\n"
" dot0 += a0*b0; dot1 += a1*b1; dot2 += a2*b2; dot3 += a3*b3;\n"
" na0 += a0*a0; na1 += a1*a1; na2 += a2*a2; na3 += a3*a3;\n"
" nb0 += b0*b0; nb1 += b1*b1; nb2 += b2*b2; nb3 += b3*b3;\n"
" }\n"
" float dot = (dot0 + dot1) + (dot2 + dot3);\n"
" float na = (na0 + na1) + (na2 + na3);\n"
" float nb = (nb0 + nb1) + (nb2 + nb3);\n"
" for (; d < p.dim; d++) {\n"
" float a = row[d], b = query[d];\n"
" dot += a*b; na += a*a; nb += b*b;\n"
" }\n"
" if (na <= 0.0f || nb <= 0.0f) { out_scores[gid] = -2.0f; return; }\n"
" out_scores[gid] = dot / sqrt(na * nb);\n"
"}\n"
"struct EgCosineMultiParams { uint n; uint dim; uint nq; };\n"
"kernel void eg_cosine_batch_multi_kernel(\n"
" device const float* queries [[buffer(0)]],\n"
" device const float* node_matrix [[buffer(1)]],\n"
" device const int* node_dims [[buffer(2)]],\n"
" constant EgCosineMultiParams& p [[buffer(3)]],\n"
" device float* out_scores [[buffer(4)]],\n"
" uint2 gid [[thread_position_in_grid]])\n"
"{\n"
" uint nid = gid.x, qid = gid.y;\n"
" if (nid >= p.n || qid >= p.nq) return;\n"
" uint64_t out_idx = (uint64_t)qid * (uint64_t)p.n + (uint64_t)nid;\n"
" if (node_dims[nid] != int(p.dim)) { out_scores[out_idx] = -2.0f; return; }\n"
" device const float* row = node_matrix + (uint64_t)nid * (uint64_t)p.dim;\n"
" device const float* query = queries + (uint64_t)qid * (uint64_t)p.dim;\n"
" float dot0 = 0.0f, dot1 = 0.0f, dot2 = 0.0f, dot3 = 0.0f;\n"
" float na0 = 0.0f, na1 = 0.0f, na2 = 0.0f, na3 = 0.0f;\n"
" float nb0 = 0.0f, nb1 = 0.0f, nb2 = 0.0f, nb3 = 0.0f;\n"
" uint d = 0;\n"
" uint dim4 = p.dim & ~3u;\n"
" for (; d < dim4; d += 4) {\n"
" float a0 = row[d], b0 = query[d];\n"
" float a1 = row[d+1], b1 = query[d+1];\n"
" float a2 = row[d+2], b2 = query[d+2];\n"
" float a3 = row[d+3], b3 = query[d+3];\n"
" dot0 += a0*b0; dot1 += a1*b1; dot2 += a2*b2; dot3 += a3*b3;\n"
" na0 += a0*a0; na1 += a1*a1; na2 += a2*a2; na3 += a3*a3;\n"
" nb0 += b0*b0; nb1 += b1*b1; nb2 += b2*b2; nb3 += b3*b3;\n"
" }\n"
" float dot = (dot0 + dot1) + (dot2 + dot3);\n"
" float na = (na0 + na1) + (na2 + na3);\n"
" float nb = (nb0 + nb1) + (nb2 + nb3);\n"
" for (; d < p.dim; d++) {\n"
" float a = row[d], b = query[d];\n"
" dot += a*b; na += a*a; nb += b*b;\n"
" }\n"
" if (na <= 0.0f || nb <= 0.0f) { out_scores[out_idx] = -2.0f; return; }\n"
" out_scores[out_idx] = dot / sqrt(na * nb);\n"
"}\n";
/* ── END embedded shader source ── */
typedef struct EgCosineParamsC { uint32_t n; uint32_t dim; } EgCosineParamsC;
typedef struct EgCosineMultiParamsC { uint32_t n; uint32_t dim; uint32_t nq; } EgCosineMultiParamsC;
static id<MTLDevice> g_device = nil;
static id<MTLCommandQueue> g_queue = nil;
static id<MTLComputePipelineState> g_pipeline = nil; /* single-query kernel */
static id<MTLComputePipelineState> g_pipeline_multi = nil; /* multi-query kernel */
static bool g_init_attempted = false;
static bool g_init_ok = false;
/* Lazy, one-time setup. Never throws — every Metal call here is the
* "returns nil/NSError on failure" flavor, not an exception-throwing one. */
static bool eg_metal_ensure_init(void) {
if (g_init_attempted) return g_init_ok;
g_init_attempted = true;
@autoreleasepool {
id<MTLDevice> dev = MTLCreateSystemDefaultDevice();
if (!dev) return false;
id<MTLCommandQueue> q = [dev newCommandQueue];
if (!q) return false;
NSError* err = nil;
NSString* src = [NSString stringWithUTF8String:kEgCosineBatchMetalSrc];
MTLCompileOptions* opts = [MTLCompileOptions new];
id<MTLLibrary> lib = [dev newLibraryWithSource:src options:opts error:&err];
if (!lib) return false;
id<MTLFunction> fn = [lib newFunctionWithName:@"eg_cosine_batch_kernel"];
if (!fn) return false;
id<MTLComputePipelineState> pipe = [dev newComputePipelineStateWithFunction:fn error:&err];
if (!pipe) return false;
id<MTLFunction> fnMulti = [lib newFunctionWithName:@"eg_cosine_batch_multi_kernel"];
if (!fnMulti) return false;
id<MTLComputePipelineState> pipeMulti = [dev newComputePipelineStateWithFunction:fnMulti error:&err];
if (!pipeMulti) return false;
g_device = dev;
g_queue = q;
g_pipeline = pipe;
g_pipeline_multi = pipeMulti;
g_init_ok = true;
return true;
}
}
static bool mh_available(void) {
return eg_metal_ensure_init();
}
static bool mh_batch(const float* query, int32_t qdim,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores) {
if (!query || qdim <= 0 || !node_ptrs || !node_dims || n <= 0 || !out_scores) return false;
if (!eg_metal_ensure_init()) return false;
@autoreleasepool {
const size_t dim = (size_t)qdim;
const size_t nu = (size_t)n;
/* Gather into a packed row-major matrix — EngramNode.emb is one
* malloc per node, not a contiguous array, so this copy is
* unavoidable regardless of backend. Rows whose real dim doesn't
* match qdim are zero-filled (harmless: the kernel sentinels them
* via node_dims before ever reading the row). */
float* matrix = (float*)calloc(nu * dim, sizeof(float));
int32_t* dims_i32 = (int32_t*)malloc(nu * sizeof(int32_t));
if (!matrix || !dims_i32) { free(matrix); free(dims_i32); return false; }
for (size_t i = 0; i < nu; i++) {
dims_i32[i] = node_dims[i];
if (node_ptrs[i] && node_dims[i] == qdim) {
memcpy(matrix + i * dim, node_ptrs[i], dim * sizeof(float));
}
/* else: leave zero-filled; node_dims[i] != qdim (or missing)
* makes the kernel sentinel it to -2.0 without reading the row. */
}
id<MTLBuffer> bufQuery = [g_device newBufferWithBytes:query
length:dim * sizeof(float)
options:MTLResourceStorageModeShared];
id<MTLBuffer> bufMatrix = [g_device newBufferWithBytes:matrix
length:nu * dim * sizeof(float)
options:MTLResourceStorageModeShared];
id<MTLBuffer> bufDims = [g_device newBufferWithBytes:dims_i32
length:nu * sizeof(int32_t)
options:MTLResourceStorageModeShared];
EgCosineParamsC params = { (uint32_t)nu, (uint32_t)dim };
id<MTLBuffer> bufParams = [g_device newBufferWithBytes:&params
length:sizeof(params)
options:MTLResourceStorageModeShared];
id<MTLBuffer> bufOut = [g_device newBufferWithLength:nu * sizeof(float)
options:MTLResourceStorageModeShared];
free(matrix); free(dims_i32);
if (!bufQuery || !bufMatrix || !bufDims || !bufParams || !bufOut) return false;
id<MTLCommandBuffer> cmd = [g_queue commandBuffer];
if (!cmd) return false;
id<MTLComputeCommandEncoder> enc = [cmd computeCommandEncoder];
if (!enc) return false;
[enc setComputePipelineState:g_pipeline];
[enc setBuffer:bufQuery offset:0 atIndex:0];
[enc setBuffer:bufMatrix offset:0 atIndex:1];
[enc setBuffer:bufDims offset:0 atIndex:2];
[enc setBuffer:bufParams offset:0 atIndex:3];
[enc setBuffer:bufOut offset:0 atIndex:4];
NSUInteger tgSize = g_pipeline.maxTotalThreadsPerThreadgroup;
if (tgSize > 256) tgSize = 256;
if (tgSize < 1) tgSize = 1;
MTLSize gridSize = MTLSizeMake(nu, 1, 1);
MTLSize threadgroupSize = MTLSizeMake(tgSize, 1, 1);
[enc dispatchThreads:gridSize threadsPerThreadgroup:threadgroupSize];
[enc endEncoding];
[cmd commit];
[cmd waitUntilCompleted];
if (cmd.status != MTLCommandBufferStatusCompleted) return false;
const float* results = (const float*)bufOut.contents;
if (!results) return false;
for (size_t i = 0; i < nu; i++) out_scores[i] = (double)results[i];
return true;
}
}
static bool mh_batch_multi(const float* queries, int32_t qdim, int32_t nq,
const float* const* node_ptrs,
const int32_t* node_dims,
int32_t n,
double* out_scores) {
if (!queries || qdim <= 0 || nq <= 0 || !node_ptrs || !node_dims || n <= 0 || !out_scores) return false;
if (!eg_metal_ensure_init()) return false;
@autoreleasepool {
const size_t dim = (size_t)qdim;
const size_t nu = (size_t)n;
const size_t nqu = (size_t)nq;
float* matrix = (float*)calloc(nu * dim, sizeof(float));
int32_t* dims_i32 = (int32_t*)malloc(nu * sizeof(int32_t));
if (!matrix || !dims_i32) { free(matrix); free(dims_i32); return false; }
for (size_t i = 0; i < nu; i++) {
dims_i32[i] = node_dims[i];
if (node_ptrs[i] && node_dims[i] == qdim) {
memcpy(matrix + i * dim, node_ptrs[i], dim * sizeof(float));
}
}
/* This is the ONE upload of node_matrix for the whole nq-query batch —
* the fix for the measured re-upload-per-query slowdown. */
id<MTLBuffer> bufMatrix = [g_device newBufferWithBytes:matrix
length:nu * dim * sizeof(float)
options:MTLResourceStorageModeShared];
id<MTLBuffer> bufDims = [g_device newBufferWithBytes:dims_i32
length:nu * sizeof(int32_t)
options:MTLResourceStorageModeShared];
id<MTLBuffer> bufQueries = [g_device newBufferWithBytes:queries
length:nqu * dim * sizeof(float)
options:MTLResourceStorageModeShared];
EgCosineMultiParamsC params = { (uint32_t)nu, (uint32_t)dim, (uint32_t)nqu };
id<MTLBuffer> bufParams = [g_device newBufferWithBytes:&params
length:sizeof(params)
options:MTLResourceStorageModeShared];
id<MTLBuffer> bufOut = [g_device newBufferWithLength:nqu * nu * sizeof(float)
options:MTLResourceStorageModeShared];
free(matrix); free(dims_i32);
if (!bufMatrix || !bufDims || !bufQueries || !bufParams || !bufOut) return false;
id<MTLCommandBuffer> cmd = [g_queue commandBuffer];
if (!cmd) return false;
id<MTLComputeCommandEncoder> enc = [cmd computeCommandEncoder];
if (!enc) return false;
[enc setComputePipelineState:g_pipeline_multi];
[enc setBuffer:bufQueries offset:0 atIndex:0];
[enc setBuffer:bufMatrix offset:0 atIndex:1];
[enc setBuffer:bufDims offset:0 atIndex:2];
[enc setBuffer:bufParams offset:0 atIndex:3];
[enc setBuffer:bufOut offset:0 atIndex:4];
/* 2D dispatch: x over nodes, y over queries. Threadgroup width picked
* from the pipeline's own limit, height fixed at 1 nq is typically
* small (tens to low hundreds) relative to n (thousands+), so tiling
* the wide axis (n) is what matters for occupancy. */
NSUInteger tgWidth = g_pipeline_multi.maxTotalThreadsPerThreadgroup;
if (tgWidth > 256) tgWidth = 256;
if (tgWidth < 1) tgWidth = 1;
MTLSize gridSize = MTLSizeMake(nu, nqu, 1);
MTLSize threadgroupSize = MTLSizeMake(tgWidth, 1, 1);
[enc dispatchThreads:gridSize threadsPerThreadgroup:threadgroupSize];
[enc endEncoding];
[cmd commit];
[cmd waitUntilCompleted];
if (cmd.status != MTLCommandBufferStatusCompleted) return false;
const float* results = (const float*)bufOut.contents;
if (!results) return false;
for (size_t i = 0; i < nqu * nu; i++) out_scores[i] = (double)results[i];
return true;
}
}
static const EgCosineBatchStrategy g_metal_hand_strategy = {
.name = "metal-hand",
.available = mh_available,
.batch = mh_batch,
.batch_multi = mh_batch_multi,
};
const EgCosineBatchStrategy* eg_cosine_batch_strategy_metal_hand(void) {
return &g_metal_hand_strategy;
}
+2753 -111
View File
File diff suppressed because it is too large Load Diff
+295 -1
View File
@@ -275,6 +275,10 @@ el_val_t json_set(el_val_t json_str, el_val_t key, el_val_t value);
el_val_t json_array_len(el_val_t json_str);
el_val_t json_array_get(el_val_t json_str, el_val_t index);
el_val_t json_array_get_string(el_val_t json_str, el_val_t index);
el_val_t json_escape_string(el_val_t sv);
el_val_t json_build_object(el_val_t kvs);
el_val_t json_build_array(el_val_t items);
el_val_t json_array_push(el_val_t arr_v, el_val_t elem_v); /* defined in el_runtime.c */
/* ── Time ────────────────────────────────────────────────────────────────── */
@@ -301,6 +305,8 @@ el_val_t time_diff(el_val_t ts1, el_val_t ts2, el_val_t unit);
el_val_t el_now_instant(void);
el_val_t now(void);
el_val_t now_millis(void); /* wall-clock milliseconds (defined in el_runtime.c) */
el_val_t now_ns(void); /* wall-clock nanoseconds (defined in el_runtime.c) */
el_val_t unix_seconds(el_val_t n);
el_val_t unix_millis(el_val_t n);
el_val_t instant_from_iso8601(el_val_t s);
@@ -580,6 +586,110 @@ void el_runtime_dharma_event_arrive(const char* event_type,
const char* payload,
const char* source);
/* ── Geometry: signal as a first-class El value ──────────────────────────────
*
* A Geometry is an opaque, magic-tagged heap value carried in an el_val_t
* the same discipline as List/Map. It holds a width and a float32 payload,
* and it is the medium a non-text modality enters in. Declared HERE, above
* the engram block, because transduction is a LANGUAGE concern: every El
* program touching any modality needs it, and the engram is merely one El
* program that happens to hold a graph. See el_runtime.c ("Geometry: signal
* as a first-class el value") for the full rationale.
*
* El-side type annotation is simply `Geometry` an opaque boxed pointer,
* exactly like Instant / Calendar / Rhythm. No codegen change is required.
*
* OWNERSHIP: a Geometry is owned by the El caller and released with
* geometry_free. node_attach_geometry COPIES, so a node and the caller's
* value have independent lifetimes. */
el_val_t geometry_new(el_val_t dim); /* zero-filled; 0 on failure */
el_val_t geometry_dim(el_val_t g); /* width, 0 if not a Geometry */
el_val_t geometry_is(el_val_t g); /* 1 if a live Geometry */
el_val_t geometry_get(el_val_t g, el_val_t i); /* Float component */
el_val_t geometry_set(el_val_t g, el_val_t i, el_val_t x); /* 1 ok / 0 out of range */
el_val_t geometry_norm(el_val_t g); /* Float L2 — lets a caller
* check a realizer emitted
* signal, not zeros */
el_val_t geometry_free(el_val_t g); /* 1 if freed, 0 if not a Geometry.
* Returns a value (not void) so it
* is safe in any El expression
* position without a codegen
* void-builtin table entry. */
/* Wire ADAPTERS — the only place an encoding appears, and only at the edge.
* `f32le hex` is little-endian float32, 8 hex chars per component: the
* encoding the perception vessel's /voice/embed already emits. The width is
* DERIVED from the input length, never supplied by a caller which is why
* there is no max-dim constant here to validate a claimed length against. */
el_val_t geometry_from_f32le_hex(el_val_t hex); /* 0 on empty/odd-length/non-hex */
el_val_t geometry_to_f32le_hex(el_val_t g); /* "" if not a Geometry */
/* ── Manifold: the result of a transduction ──────────────────────────────────
* A transduced signal is a SUBGRAPH named components, each with its own
* geometry, plus typed weighted relations among them not a single vector.
* One vector is a fingerprint: matchable, rankable, and nothing else. A song
* decomposes into pitch, interval, rhythm, harmonic function; the song IS the
* structure of those relations, and collapsing it to a point discards exactly
* what made it reasonable-about. See el_runtime.c ("Manifold") for the full
* rationale, the key-addressing rule, and the ownership contract.
*
* Components are addressed BY KEY, never by index, because the key is what
* survives persistence: a component becomes a node, and it is separately
* groundable precisely because it is separately named. Relation weight IS the
* grounding (correspondence-and-censorship.md §1) one quantity, no separate
* score, nothing computed on read.
*
* OWNERSHIP: a Manifold is owned by the El caller and released with
* manifold_free, which also releases every component's geometry. manifold_add
* COPIES the geometry it is given and manifold_geometry RETURNS a copy, so no
* component's vector is ever aliased in either direction. */
el_val_t manifold_new(void); /* empty; 0 on failure */
el_val_t manifold_is(el_val_t m); /* 1 if a live Manifold */
el_val_t manifold_add(el_val_t m, el_val_t key, el_val_t role, el_val_t g);
/* component index, or -1 on empty/duplicate
* key or a value that is not a Geometry */
el_val_t manifold_relate(el_val_t m, el_val_t from, el_val_t rel,
el_val_t to, el_val_t weight);
/* 1 ok / 0 if either endpoint is unknown —
* an unresolvable edge is REFUSED, never
* silently dropped */
el_val_t manifold_size(el_val_t m); /* component count */
el_val_t manifold_rel_count(el_val_t m); /* relation count */
el_val_t manifold_index_of(el_val_t m, el_val_t key); /* index by key, or -1 */
el_val_t manifold_key(el_val_t m, el_val_t i); /* "" if out of range */
el_val_t manifold_role(el_val_t m, el_val_t i); /* "" if out of range */
el_val_t manifold_geometry(el_val_t m, el_val_t i); /* a COPY the caller frees */
el_val_t manifold_rel_from(el_val_t m, el_val_t j); /* source component key */
el_val_t manifold_rel_name(el_val_t m, el_val_t j); /* relation name */
el_val_t manifold_rel_to(el_val_t m, el_val_t j); /* target component key */
el_val_t manifold_rel_weight(el_val_t m, el_val_t j); /* Float — the grounding */
el_val_t manifold_single(el_val_t key, el_val_t role, el_val_t g);
/* the degenerate one-part case, expressible
* but visibly a size-1 manifold rather than
* a parallel path back to a bare vector */
el_val_t manifold_free(el_val_t m); /* 1 if freed, 0 otherwise */
/* ── Realizers + transduce ───────────────────────────────────────────────────
* A REALIZER DECOMPOSES one modality into components and relations. It does
* not encode a signal to a point; that operation is one layer below and is
* called geometry. Registration is by NAME, so a new modality never requires a
* runtime patch: every El `fn name(...)` compiles to a global C symbol with
* that exact name, and the registry resolves it with dlsym against the running
* binary the same mechanism http_set_handler already relies on.
*
* fn tone_realizer(signal: String) -> Manifold { ... }
* realizer_register("tone", "tone_realizer")
* let m: Manifold = transduce(sample, "tone")
*
* SUPERSEDES #144's `transduce -> Geometry`. A realizer that still returns a
* bare Geometry now transduces NOTHING (transduce returns 0), deliberately: an
* organ that only fingerprints must not be indistinguishable from a working
* one. A modality with genuinely one part says so with manifold_single. */
el_val_t realizer_register(el_val_t modality, el_val_t fn_name); /* 1 ok / 0 unresolved */
el_val_t realizer_has(el_val_t modality); /* 1 if a realizer is registered */
el_val_t transduce(el_val_t signal, el_val_t modality); /* Manifold, or 0 if no organ */
/* ── Engram local graph primitives ───────────────────────────────────────────
* Operate on the CGI's local Engram knowledge graph.
* `engram_activate` queries the local graph only; `dharma_activate` is
@@ -606,7 +716,28 @@ el_val_t engram_get_node(el_val_t id);
void engram_strengthen(el_val_t node_id);
void engram_forget(el_val_t node_id);
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
/* Largest byte length <= max_bytes that does not split a UTF-8 codepoint.
* Bounded by bytes, not codepoints, so truncated strings never grow. */
size_t el_utf8_safe_len(const char* s, size_t max_bytes);
el_val_t engram_node_count(void);
/* Attach a Geometry to an existing node, and read the attached width back.
* Named for the operation, not the store: a node acquires geometry. This is
* the geometry-valued ingest path nothing about it is hex, and nothing
* about it assumes the caller's vector matches the canonical text-embedding
* width. node_geometry_dim exists so an attach is VERIFIED by reading it
* back rather than by trusting a success return. */
el_val_t node_attach_geometry(el_val_t node_id, el_val_t g); /* 1 ok / 0 otherwise */
el_val_t node_geometry_dim(el_val_t node_id); /* width, 0 if none */
/* DEPRECATED (shipped in #141, superseded 2026-08-16). Equivalent to
* geometry_from_f32le_hex + node_attach_geometry, and now implemented as
* exactly that. Kept only so anything built against the #141 runtime keeps
* linking; `dim` is accepted but treated as an assertion about the vector's
* width rather than as its source. New code should not call this a hex
* string is a wire encoding, not a way to move geometry between two pieces
* of El. Returns 1 on success, 0 otherwise. */
el_val_t engram_node_set_emb(el_val_t id, el_val_t hex, el_val_t dim);
el_val_t engram_search(el_val_t query, el_val_t limit);
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
void engram_connect(el_val_t from_id, el_val_t to_id, el_val_t weight, el_val_t relation);
@@ -651,8 +782,13 @@ el_val_t engram_geo_analogy_json(el_val_t a_seeds, el_val_t b_seeds);
el_val_t engram_reason_analogy_json(el_val_t a_seeds, el_val_t b_seeds, el_val_t c_seeds);
/* COGNITION (2026-08-14): THE ONE OPERATION + grounding, surfaced live. */
el_val_t engram_think_json(el_val_t seeds, el_val_t faculty);
/* GROUNDING (2026-08-16): grounding is an attribute of the RELATION and it IS the
* hebbian weight. ground reads; ground_record writes; trajectory reads the chain. */
el_val_t engram_ground_json(el_val_t claim, el_val_t evidence, el_val_t for_whom);
el_val_t engram_assert_json(el_val_t claim_id, el_val_t for_whom, el_val_t floor);
el_val_t engram_ground_record_json(el_val_t claim, el_val_t evidence,
el_val_t provenance, el_val_t floor);
el_val_t engram_ground_trajectory_json(el_val_t claim, el_val_t evidence);
el_val_t engram_assert_json(el_val_t claim_id, el_val_t for_whom, el_val_t floor, el_val_t rel_floor);
el_val_t engram_attend_json(el_val_t node_id, el_val_t observer, el_val_t salience);
el_val_t engram_correspondence_beat_json(el_val_t seeds, el_val_t faculty, el_val_t keystone);
el_val_t engram_consolidate_permanence(el_val_t node_id);
@@ -712,6 +848,14 @@ el_val_t engram_label_df(el_val_t term);
el_val_t engram_salient_term(el_val_t node_id, el_val_t max_df,
el_val_t min_df, el_val_t tabu);
el_val_t engram_embed_backfill(el_val_t count);
/* op_assert seam: grounded assertion envelope {subject,grounding} for the realizer. */
el_val_t engram_op_assert_json(el_val_t node_id, el_val_t depth);
/* Parametric mutation (purview write-side): purview==0 => G=live (default), else refuse. */
el_val_t engram_node_full_in(el_val_t purview, el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t importance, el_val_t confidence,
el_val_t tier, el_val_t tags);
void engram_connect_in(el_val_t purview, el_val_t from_id, el_val_t to_id,
el_val_t weight, el_val_t relation);
el_val_t engram_list_layers_json(void);
/* Working memory introspection — count, mean weight, and top-N snapshot.
* Ported from runtime on 2026-06-30 self-review. */
@@ -892,6 +1036,156 @@ el_val_t trace_span_start(el_val_t name);
el_val_t trace_span_end(el_val_t span_handle);
el_val_t emit_event(el_val_t name, el_val_t duration_ms);
el_val_t __thread_create(el_val_t fn_name_v, el_val_t arg_v);
el_val_t __thread_join(el_val_t tid_v);
/* Mutex + channel seed primitives (defined in el_runtime.c). Declared here so
* that compiled El programs which use runtime/thread.el's with_mutex helper or
* runtime/channel.el's Go-style channels see real prototypes instead of an
* implicit int-return declaration (which the C11 ABI mis-truncates el_val_t). */
el_val_t __mutex_new(void);
void __mutex_lock(el_val_t m_v);
void __mutex_unlock(el_val_t m_v);
el_val_t __channel_new(el_val_t capacity_v);
el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v);
el_val_t __channel_recv(el_val_t ch_v);
el_val_t __channel_try_recv(el_val_t ch_v);
el_val_t __channel_close(el_val_t ch_v);
/* ── __ prefixed aliases (self-hosting compiler ABI) ─────────────────────────
* The El self-hosting compiler emits calls to __-prefixed names. These are
* forwarding wrappers around the existing el_runtime functions above. */
/* I/O */
el_val_t __println(el_val_t s);
el_val_t __print(el_val_t s);
el_val_t __readline(void);
/* String */
el_val_t __int_to_str(el_val_t n);
el_val_t __str_to_int(el_val_t s);
el_val_t __float_to_str(el_val_t f);
el_val_t __str_to_float(el_val_t s);
el_val_t __str_len(el_val_t s);
el_val_t __str_char_at(el_val_t s, el_val_t i);
el_val_t __str_cmp(el_val_t a, el_val_t b);
el_val_t __str_ncmp(el_val_t a, el_val_t b, el_val_t n);
el_val_t __str_concat_raw(el_val_t a, el_val_t b);
el_val_t __str_slice_raw(el_val_t s, el_val_t start, el_val_t end);
el_val_t __str_alloc(el_val_t n);
el_val_t __str_set_char(el_val_t s, el_val_t i, el_val_t c);
/* URL encoding */
el_val_t __url_encode(el_val_t s);
el_val_t __url_decode(el_val_t s);
/* Environment */
el_val_t __env_get(el_val_t key);
/* Cross-cutting concerns declared by a `program` block (spec §18).
* All three are COMPILER-INJECTED at the head of main() they are not meant to
* be written by hand, which is the point: the guarantee cannot be forgotten at a
* call site because there is no call site. */
el_val_t el_singleton_acquire(el_val_t id); /* §18.1 process identity */
el_val_t el_config_declare(el_val_t name, el_val_t type,
el_val_t deflt, el_val_t has_default,
el_val_t required); /* §18.2 config schema */
el_val_t el_config_validate(el_val_t program_name); /* §18.2 startup validate */
/* config(key) — the READ side, and the only one programs write by hand. With a
* schema declared it is a validated lookup; without one it degrades to getenv.
* (Defined in el_runtime.c but previously never prototyped here, so any program
* calling it failed to compile under -Werror=implicit-function-declaration.) */
el_val_t config(el_val_t key);
/* Subprocess */
el_val_t __exec(el_val_t cmd);
el_val_t __exec_bg(el_val_t cmd);
/* Process */
el_val_t __exit_program(el_val_t code);
/* Filesystem */
el_val_t __fs_exists(el_val_t path);
el_val_t __fs_mkdir(el_val_t path);
el_val_t __fs_read(el_val_t path);
el_val_t __fs_write(el_val_t path, el_val_t content);
el_val_t __fs_write_bytes(el_val_t path, el_val_t bytes, el_val_t n);
el_val_t __fs_list_raw(el_val_t path);
/* HTTP server */
el_val_t __http_response(el_val_t status, el_val_t headers_json, el_val_t body);
el_val_t __http_serve(el_val_t port, el_val_t handler);
el_val_t __http_serve_v2(el_val_t port, el_val_t handler);
/* HTTP conn fd / SSE (weak; overridden by el_seed.c when linked together) */
el_val_t __http_conn_fd(void);
el_val_t __http_sse_open(el_val_t conn_id);
el_val_t __http_sse_send(el_val_t conn_id, el_val_t data);
el_val_t __http_sse_close(el_val_t conn_id);
/* HTTP client (requires HAVE_CURL; stubs provided for no-curl builds) */
el_val_t __http_do(el_val_t method, el_val_t url, el_val_t body,
el_val_t headers_map, el_val_t timeout_ms);
el_val_t __http_do_map(el_val_t method, el_val_t url, el_val_t body,
el_val_t headers_json, el_val_t timeout_ms);
el_val_t __http_do_map_to_file(el_val_t method, el_val_t url, el_val_t body,
el_val_t headers_json, el_val_t output_path);
/* JSON */
el_val_t __json_array_get(el_val_t json, el_val_t index);
el_val_t __json_array_get_string(el_val_t json, el_val_t index);
el_val_t __json_array_len(el_val_t json);
el_val_t __json_get(el_val_t json, el_val_t key);
el_val_t __json_get_raw(el_val_t json, el_val_t key);
el_val_t __json_set(el_val_t json, el_val_t key, el_val_t value);
el_val_t __json_parse_map(el_val_t json_str);
el_val_t __json_stringify_val(el_val_t val);
/* Hashing */
el_val_t __sha256_hex(el_val_t s);
/* State K/V */
el_val_t __state_del(el_val_t key);
el_val_t __state_get(el_val_t key);
el_val_t __state_keys(void);
el_val_t __state_set(el_val_t key, el_val_t val);
/* UUID */
el_val_t __uuid_v4(void);
/* Args */
el_val_t __args_json(void);
/* Compiler-support builtins — called by the El compiler's own source
* (compiler.el, codegen.el) and registered in codegen.el's builtin_arity. */
el_val_t stdout_to_file(el_val_t path);
el_val_t stdout_restore(void);
el_val_t el_mem_check(void);
/* Allocation accounting — the deterministic signal behind complexity gating.
* Gate on counts/bytes; peak RSS is context only. */
el_val_t el_alloc_count(void);
el_val_t el_alloc_bytes(void);
el_val_t el_peak_rss(void);
el_val_t el_black_box(el_val_t v);
/* Semantic retrieval surface. NOT interchangeable with engram_search_json,
* which is lexical by design see the note at the definition. */
el_val_t engram_recall_json(el_val_t query, el_val_t limit);
/* Edges straight from the store — replaces the engram_save()+fs_read()
* whole-graph round trip that /api/graph/edges used to do. */
el_val_t engram_edges_json(el_val_t limit, el_val_t offset);
/* Buffer-pool interoception as JSON — live pool health for observation. */
el_val_t engram_pool_stats_json(void);
/* CGI identity accessors (read-only). */
el_val_t cgi_principal(void);
el_val_t cgi_network(void);
el_val_t cgi_engram(void);
#ifdef __cplusplus
}
#endif
+332
View File
@@ -37,6 +37,82 @@
#include <dlfcn.h>
#include <curl/curl.h>
/* el_runtime.c bridge prototypes.
*
* A block of __-prefixed wrappers further down in this file (http serving,
* JSON access, key-val state, URL/HTML escaping, and the whole engram_*
* node/edge/layer/search surface -- 51 symbols in total) delegate to
* unprefixed counterparts that are implemented in el_runtime.c, not here.
* Porting them into native el_seed.c or El has not happened yet.
* tools/install.sh compiles el_seed.c and el_runtime.c as separate objects
* and archives both into libel.a, so the symbols are always present at link
* time. el_seed.c alone was just missing the prototypes, which made even a
* standalone -c compile of this one file fail on a toolchain that now treats
* an implicit function declaration as a hard error under C11.
*
* A plain include of el_runtime.h was tried first and rejected: it redefines
* el_to_float and el_from_float, which el_seed.h already provides. Narrow
* prototypes, copied verbatim from el_runtime.h, avoid that collision without
* pulling in the rest of the retiring runtime header.
*/
el_val_t http_response(el_val_t status, el_val_t headers_json, el_val_t body);
void http_serve(el_val_t port, el_val_t handler);
void http_serve_v2(el_val_t port, el_val_t handler);
el_val_t json_get(el_val_t json, el_val_t key);
el_val_t json_get_string(el_val_t json_str, el_val_t key);
el_val_t json_get_int(el_val_t json_str, el_val_t key);
el_val_t json_get_float(el_val_t json_str, el_val_t key);
el_val_t json_get_bool(el_val_t json_str, el_val_t key);
el_val_t json_get_raw(el_val_t json_str, el_val_t key);
el_val_t json_parse(el_val_t s);
el_val_t json_set(el_val_t json_str, el_val_t key, el_val_t value);
el_val_t json_stringify(el_val_t v);
el_val_t json_array_len(el_val_t json_str);
el_val_t json_array_get(el_val_t json_str, el_val_t index);
el_val_t json_array_get_string(el_val_t json_str, el_val_t index);
el_val_t state_set(el_val_t key, el_val_t value);
el_val_t state_get(el_val_t key);
el_val_t state_del(el_val_t key);
el_val_t state_keys(void);
el_val_t url_encode(el_val_t s);
el_val_t url_decode(el_val_t s);
el_val_t el_html_sanitize(el_val_t input_html, el_val_t allowlist_json);
el_val_t engram_node(el_val_t content, el_val_t node_type, el_val_t salience);
el_val_t engram_node_full(el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t importance, el_val_t confidence,
el_val_t tier, el_val_t tags);
el_val_t engram_node_layered(el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t certainty, el_val_t confidence,
el_val_t status, el_val_t tags, el_val_t layer_id);
el_val_t engram_add_layer(el_val_t name, el_val_t priority, el_val_t suppressible,
el_val_t transparent, el_val_t injectable);
el_val_t engram_remove_layer(el_val_t layer_id);
el_val_t engram_list_layers(void);
el_val_t engram_list_layers_json(void);
el_val_t engram_get_node(el_val_t id);
el_val_t engram_get_node_json(el_val_t id);
el_val_t engram_get_node_by_label(el_val_t label);
void engram_strengthen(el_val_t node_id);
void engram_forget(el_val_t node_id);
el_val_t engram_node_count(void);
el_val_t engram_edge_count(void);
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_json(el_val_t limit, el_val_t offset);
el_val_t engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset);
el_val_t engram_search(el_val_t query, el_val_t limit);
el_val_t engram_search_json(el_val_t query, el_val_t limit);
el_val_t engram_activate(el_val_t query, el_val_t depth);
el_val_t engram_activate_json(el_val_t query, el_val_t depth);
el_val_t engram_compile_layered_json(el_val_t intent, el_val_t depth);
el_val_t engram_stats_json(void);
void engram_connect(el_val_t from_id, el_val_t to_id, el_val_t weight, el_val_t relation);
el_val_t engram_edge_between(el_val_t from_id, el_val_t to_id);
el_val_t engram_neighbors(el_val_t node_id);
el_val_t engram_neighbors_filtered(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t engram_load(el_val_t path);
el_val_t engram_save(el_val_t path);
/* ── Private allocator ───────────────────────────────────────────────────── */
/*
* el_seed.c carries its own arena for per-request allocation tracking.
@@ -72,10 +148,17 @@ static void seed_request_start(void) {
_seed_arena_on = 1;
}
/* Defined in el_runtime.c. The string-length cache there keys on pointer +
* generation; anything that frees or mutates a runtime string must bump the
* generation or a reused address could return a stale length. Weak so this
* file still links on its own. */
__attribute__((weak)) void el_str_cache_flush(void);
static void seed_request_end(void) {
_seed_arena_on = 0;
for (size_t i = 0; i < _seed_arena.count; i++) free(_seed_arena.ptrs[i]);
_seed_arena.count = 0;
if (el_str_cache_flush) el_str_cache_flush(); /* freed pointers may be reused */
}
/* el_request_start / el_request_end — formerly defined in el_runtime.c.
@@ -137,6 +220,7 @@ el_val_t __str_set_char(el_val_t s, el_val_t i, el_val_t c) {
int64_t idx = (int64_t)i;
if (idx < 0 || idx >= len) return s;
p[idx] = (char)(unsigned char)(int64_t)c;
if (el_str_cache_flush) el_str_cache_flush(); /* in-place write can move the NUL */
return s;
}
@@ -831,6 +915,219 @@ void __mutex_unlock(el_val_t m) {
pthread_mutex_unlock(&_el_mutexes[slot]);
}
/* ── Channels ─────────────────────────────────────────────────────────────── *
* Buffered MPMC channel backed by a mutex + condvar + circular buffer.
* Ported from the pre-restructure el_runtime.c (b2aac4b) runtime/channel.el
* has always called these five primitives, but they were never carried
* forward into el_seed.c when el_runtime.c was consolidated onto the
* canonical release copy. Native channels were silently unlinkable on dev
* until this port.
*
* __channel_new(capacity) -> Int (handle)
* __channel_send(ch, msg) blocks if full (capacity > 0) or never (unbounded)
* __channel_recv(ch) -> String blocks until a message is available
* __channel_try_recv(ch) -> String non-blocking, returns "" if empty
* __channel_close(ch) signal no more sends; recv drains remaining
*
* Bounded channels (cap > 0): circular buffer, sender blocks when full.
* Unbounded channels (cap == 0): dynamic array, sender never blocks.
*/
#define EL_CHANNEL_MAX 64
#define EL_CHANNEL_BUF 1024
typedef struct {
char** buf;
int cap; /* 0 = unbounded (grows dynamically) */
int head, tail, count;
int dyn_cap; /* allocated slots for unbounded mode */
int closed;
pthread_mutex_t mu;
pthread_cond_t not_empty;
pthread_cond_t not_full;
} ElChannel;
static ElChannel _channels[EL_CHANNEL_MAX];
static int _channel_count = 0;
static pthread_mutex_t _channel_alloc_mu = PTHREAD_MUTEX_INITIALIZER;
el_val_t __channel_new(el_val_t capacity_v) {
int cap = (int)(int64_t)capacity_v;
if (cap < 0) cap = 0;
pthread_mutex_lock(&_channel_alloc_mu);
if (_channel_count >= EL_CHANNEL_MAX) {
pthread_mutex_unlock(&_channel_alloc_mu);
fprintf(stderr, "[__channel_new] channel table full\n");
return EL_INT(-1);
}
int slot = _channel_count++;
pthread_mutex_unlock(&_channel_alloc_mu);
ElChannel* ch = &_channels[slot];
memset(ch, 0, sizeof(*ch));
ch->cap = cap;
ch->closed = 0;
ch->head = 0;
ch->tail = 0;
ch->count = 0;
if (cap > 0) {
/* Bounded: fixed circular buffer. */
ch->buf = (char**)malloc((size_t)cap * sizeof(char*));
ch->dyn_cap = cap;
} else {
/* Unbounded: start with EL_CHANNEL_BUF slots, grow as needed. */
ch->buf = (char**)malloc(EL_CHANNEL_BUF * sizeof(char*));
ch->dyn_cap = EL_CHANNEL_BUF;
}
if (!ch->buf) {
fprintf(stderr, "[__channel_new] out of memory\n");
return EL_INT(-1);
}
pthread_mutex_init(&ch->mu, NULL);
pthread_cond_init(&ch->not_empty, NULL);
pthread_cond_init(&ch->not_full, NULL);
return EL_INT(slot);
}
el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v) {
int slot = (int)(int64_t)ch_v;
if (slot < 0 || slot >= EL_CHANNEL_MAX) return EL_STR("");
ElChannel* ch = &_channels[slot];
const char* msg = EL_CSTR(msg_v);
if (!msg) msg = "";
char* copy = strdup(msg); /* channel owns the string */
pthread_mutex_lock(&ch->mu);
if (ch->closed) {
/* Send on closed channel is a no-op (drop the message). */
pthread_mutex_unlock(&ch->mu);
free(copy);
return EL_STR("");
}
if (ch->cap > 0) {
/* Bounded: block while full. */
while (ch->count >= ch->cap && !ch->closed) {
pthread_cond_wait(&ch->not_full, &ch->mu);
}
if (ch->closed) {
pthread_mutex_unlock(&ch->mu);
free(copy);
return EL_STR("");
}
ch->buf[ch->tail] = copy;
ch->tail = (ch->tail + 1) % ch->cap;
ch->count++;
} else {
/* Unbounded: grow the buffer if needed. */
if (ch->count >= ch->dyn_cap) {
int new_cap = ch->dyn_cap * 2;
char** grown = (char**)realloc(ch->buf, (size_t)new_cap * sizeof(char*));
if (!grown) {
pthread_mutex_unlock(&ch->mu);
free(copy);
fprintf(stderr, "[__channel_send] out of memory growing channel\n");
return EL_STR("");
}
/* The circular buffer may have wrapped. Linearise it first.
* In unbounded mode head is always 0 (we append at tail, drain
* from head), so a simple memmove isn't needed but if the
* buffer did wrap (tail < head after growth), we need to fix up.
* Simplest safe path: if tail wrapped, move the head..old_cap
* segment to new_cap..new_cap+(old_cap-head). */
if (ch->tail < ch->head) {
/* Wrapped: [head..old_cap) is the front, [0..tail) is the back. */
int front = ch->dyn_cap - ch->head;
memmove(grown + ch->dyn_cap, grown + ch->head, (size_t)front * sizeof(char*));
ch->head = ch->dyn_cap;
}
ch->buf = grown;
ch->dyn_cap = new_cap;
}
ch->buf[ch->tail] = copy;
ch->tail = (ch->tail + 1) % ch->dyn_cap;
ch->count++;
}
pthread_cond_signal(&ch->not_empty);
pthread_mutex_unlock(&ch->mu);
return EL_STR("");
}
el_val_t __channel_recv(el_val_t ch_v) {
int slot = (int)(int64_t)ch_v;
if (slot < 0 || slot >= EL_CHANNEL_MAX) return EL_STR("");
ElChannel* ch = &_channels[slot];
pthread_mutex_lock(&ch->mu);
/* Block until there is a message or the channel is closed and drained. */
while (ch->count == 0 && !ch->closed) {
pthread_cond_wait(&ch->not_empty, &ch->mu);
}
if (ch->count == 0) {
/* Closed and empty — signal EOF. */
pthread_mutex_unlock(&ch->mu);
return EL_STR("");
}
int buf_cap = (ch->cap > 0) ? ch->cap : ch->dyn_cap;
char* msg = ch->buf[ch->head];
ch->head = (ch->head + 1) % buf_cap;
ch->count--;
pthread_cond_signal(&ch->not_full);
pthread_mutex_unlock(&ch->mu);
/* Hand the string to the arena so it is freed after the request. */
seed_arena_track(msg);
return EL_STR(msg);
}
el_val_t __channel_try_recv(el_val_t ch_v) {
int slot = (int)(int64_t)ch_v;
if (slot < 0 || slot >= EL_CHANNEL_MAX) return EL_STR("");
ElChannel* ch = &_channels[slot];
pthread_mutex_lock(&ch->mu);
if (ch->count == 0) {
pthread_mutex_unlock(&ch->mu);
return EL_STR("");
}
int buf_cap = (ch->cap > 0) ? ch->cap : ch->dyn_cap;
char* msg = ch->buf[ch->head];
ch->head = (ch->head + 1) % buf_cap;
ch->count--;
pthread_cond_signal(&ch->not_full);
pthread_mutex_unlock(&ch->mu);
seed_arena_track(msg);
return EL_STR(msg);
}
el_val_t __channel_close(el_val_t ch_v) {
int slot = (int)(int64_t)ch_v;
if (slot < 0 || slot >= EL_CHANNEL_MAX) return EL_STR("");
ElChannel* ch = &_channels[slot];
pthread_mutex_lock(&ch->mu);
ch->closed = 1;
/* Wake all blocked recvers and senders so they can observe the close. */
pthread_cond_broadcast(&ch->not_empty);
pthread_cond_broadcast(&ch->not_full);
pthread_mutex_unlock(&ch->mu);
return EL_STR("");
}
/* ── Subprocess ──────────────────────────────────────────────────────────── */
el_val_t __exec(el_val_t cmd) {
@@ -1082,6 +1379,21 @@ el_val_t __engram_scan_nodes_json(el_val_t limit, el_val_t offset) {
return engram_scan_nodes_json(limit, offset);
}
el_val_t engram_edges_json(el_val_t limit, el_val_t offset);
el_val_t __engram_edges_json(el_val_t limit, el_val_t offset) {
return engram_edges_json(limit, offset);
}
el_val_t engram_pool_stats_json(void);
el_val_t __engram_pool_stats_json(void) { return engram_pool_stats_json(); }
el_val_t el_alloc_count(void);
el_val_t el_alloc_bytes(void);
el_val_t el_peak_rss(void);
el_val_t __el_alloc_count(void) { return el_alloc_count(); }
el_val_t __el_alloc_bytes(void) { return el_alloc_bytes(); }
el_val_t __el_peak_rss(void) { return el_peak_rss(); }
el_val_t __engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, el_val_t offset) {
return engram_scan_nodes_by_type_json(node_type, limit, offset);
}
@@ -1094,7 +1406,27 @@ el_val_t __engram_activate_json(el_val_t query, el_val_t depth) {
return engram_activate_json(query, depth);
}
/* Forward decls for el_runtime.c symbols this file wraps. el_seed.c does not
* include el_runtime.h (documented in lang/AGENTS.md), so each wrapped symbol
* needs a prototype here or clang treats it as an implicit declaration (error
* under C99+) and the ABI mis-truncates the el_val_t return. */
el_val_t engram_op_assert_json(el_val_t node_id, el_val_t depth);
el_val_t engram_node_full_in(el_val_t purview, el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t importance, el_val_t confidence,
el_val_t tier, el_val_t tags);
void engram_connect_in(el_val_t purview, el_val_t from_id, el_val_t to_id,
el_val_t weight, el_val_t relation);
el_val_t __engram_stats_json(void) { return engram_stats_json(); }
el_val_t __engram_op_assert_json(el_val_t node_id, el_val_t depth) { return engram_op_assert_json(node_id, depth); }
el_val_t __engram_node_full_in(el_val_t purview, el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t importance, el_val_t confidence,
el_val_t tier, el_val_t tags) {
return engram_node_full_in(purview, content, node_type, label, salience, importance, confidence, tier, tags);
}
void __engram_connect_in(el_val_t purview, el_val_t from_id, el_val_t to_id, el_val_t weight, el_val_t relation) {
engram_connect_in(purview, from_id, to_id, weight, relation);
}
el_val_t __engram_list_layers_json(void) { return engram_list_layers_json(); }
el_val_t __engram_compile_layered_json(el_val_t intent, el_val_t depth) {
+13
View File
@@ -139,6 +139,13 @@ el_val_t __mutex_new(void);
void __mutex_lock(el_val_t m);
void __mutex_unlock(el_val_t m);
/* Buffered MPMC channel (runtime/channel.el). capacity=0 means unbounded. */
el_val_t __channel_new(el_val_t capacity);
el_val_t __channel_send(el_val_t ch, el_val_t msg); /* blocks if bounded+full */
el_val_t __channel_recv(el_val_t ch); /* blocks until available */
el_val_t __channel_try_recv(el_val_t ch); /* non-blocking, "" if empty */
el_val_t __channel_close(el_val_t ch);
/* ── Subprocess ──────────────────────────────────────────────────────────── */
el_val_t __exec(el_val_t cmd); /* popen, capture all stdout, return String */
@@ -233,6 +240,12 @@ el_val_t __engram_scan_nodes_by_type_json(el_val_t node_type, el_val_t limit, e
el_val_t __engram_neighbors_json(el_val_t node_id, el_val_t max_depth, el_val_t direction);
el_val_t __engram_activate_json(el_val_t query, el_val_t depth);
el_val_t __engram_stats_json(void);
el_val_t __engram_op_assert_json(el_val_t node_id, el_val_t depth);
el_val_t __engram_node_full_in(el_val_t purview, el_val_t content, el_val_t node_type, el_val_t label,
el_val_t salience, el_val_t importance, el_val_t confidence,
el_val_t tier, el_val_t tags);
void __engram_connect_in(el_val_t purview, el_val_t from_id, el_val_t to_id,
el_val_t weight, el_val_t relation);
el_val_t __engram_list_layers_json(void);
el_val_t __engram_compile_layered_json(el_val_t intent, el_val_t depth);
+256
View File
@@ -0,0 +1,256 @@
// runtime/elbench.el growth-curve classifier and complexity gate.
//
// Given a geometric sweep of input sizes and the measurements taken at each,
// classify the growth curve and decide whether it violates a declared bound.
//
// Why this exists
//
// Constant-factor regressions are annoying. Complexity regressions are outages.
// An O(n) lookup inside an O(n) loop is invisible at n=100 in a unit test and
// catastrophic at n=100000 in production. el #132 was exactly that: a strlen()
// inside a per-character accessor, quadratic, shipped for months.
//
// THREE signals, not one
//
// The gate fits time AND allocation-count AND allocation-bytes, and fails if
// ANY of them exceeds its declared curve. This is not belt-and-braces; each
// signal is blind to a real defect class the others catch:
//
// * A copy-on-write accumulator rebuilding its buffer allocates ONCE per
// iteration count is exactly linear while bytes go quadratic.
// Count alone passes it.
// * el #132's strlen-per-character is pure CPU and allocates NOTHING.
// Both allocation signals read FLAT. Only time catches it.
//
// The deterministic signals (count, bytes) are preferable where they apply:
// no statistics, correct on the first run, machine-independent. They are
// simply not sufficient.
//
// SCOPE LIMIT read this before trusting a flat curve
//
// The allocation counters track EL-LEVEL allocation only: strings, ElList and
// ElMap bodies, their backing arrays, copy-on-write clones, and the realloc
// growth path. malloc inside engram_*.c and inside libcurl is NOT counted.
//
// A flat allocation curve over a workload dominated by engram or HTTP calls is
// therefore NOT evidence of anything. It means "no El-level allocation growth",
// not "no allocation growth". Gate El-level complexity with this; do not read
// third-party memory behaviour into it.
//
// Classification method
//
// Sizes must form a geometric sweep (each n double the last). On such a sweep
// the ratio between consecutive measurements IS the growth exponent, directly:
//
// O(1) -> 1.0 O(log n) -> ~1.1 O(n) -> 2.0
// O(n log n) -> ~2.2 O(n^2) -> 4.0 O(n^3) -> 8.0
//
// DEVIATION FROM DESIGN.md 6.2, stated plainly: that section specified Google
// Benchmark's one-parameter least-squares fit over candidate curves. This uses
// consecutive ratios instead. The sweep is mandated geometric either way, and
// on a geometric sweep ratios are directly interpretable and need no floating
// point. The cost is weaker separation between O(n) and O(n log n), which is
// reported honestly as an ambiguous band rather than guessed at. Least-squares
// remains the better answer if that band ever needs to be resolved.
//
// All arithmetic is fixed-point, scaled by 1000 ("milli-ratio"), so a ratio of
// 2.0 is 2000. El values are int64; this avoids float-in-list handling.
// Curve identifiers. Ordered by growth the ordering IS the comparison used
// by the gate, so an index comparison decides "worse than declared".
// 0 = O(1) 1 = O(log n) 2 = O(n) 3 = O(n log n) 4 = O(n^2) 5 = O(n^3)
fn elb_curve_name(c: Int) -> String {
if c == 0 { return "O(1)" }
if c == 1 { return "O(log n)" }
if c == 2 { return "O(n)" }
if c == 3 { return "O(n log n)" }
if c == 4 { return "O(n^2)" }
if c == 5 { return "O(n^3)" }
return "O(?)"
}
fn elb_curve_from_name(s: String) -> Int {
if str_eq(s, "O(1)") { return 0 }
if str_eq(s, "O(log n)") { return 1 }
if str_eq(s, "O(n)") { return 2 }
if str_eq(s, "O(n log n)") { return 3 }
if str_eq(s, "O(n^2)") { return 4 }
if str_eq(s, "O(n^3)") { return 5 }
return -1
}
// elb_classify_ratio map a milli-ratio-per-doubling onto a curve.
//
// Bands are deliberately wide at the top (a quadratic measured at 3.4x is
// still a quadratic) and deliberately overlap-averse at the bottom, where a
// misclassification between O(1) and O(log n) matters least.
fn elb_classify_ratio(milli: Int) -> Int {
if milli < 1300 { return 0 }
if milli < 1700 { return 1 }
if milli < 2400 { return 2 }
if milli < 3200 { return 3 }
if milli < 6000 { return 4 }
return 5
}
// elb_ratio milli-ratio between two consecutive measurements.
// Returns -1 when the earlier measurement is zero (ratio undefined).
fn elb_ratio(prev: Int, cur: Int) -> Int {
if prev <= 0 { return -1 }
return (cur * 1000) / prev
}
// The measurement floor
//
// A benchmark whose largest measurement is at or near zero has not been
// measured. Reporting it as O(1) would be a confident answer with nothing
// behind it the same failure as a test that never ran reporting pass, and
// exactly what happened when clang closed a nested loop to a multiply and the
// harness read 0 microseconds at every n.
//
// So: REFUSE. Never classify below the floor.
fn elb_below_floor(vals: [Int], floor: Int) -> Bool {
let n: Int = native_list_len(vals)
let i: Int = 0
let mx: Int = 0
while i < n {
let v: Int = native_list_get(vals, i)
if v > mx { let mx = v }
let i = i + 1
}
if mx < floor { return true }
return false
}
// elb_implausibly_flat a measurement that does not move across a sweep whose
// input grew by 8x or more is not a flat curve, it is a broken measurement.
// Genuine O(1) work still shows noise; a hard-flat series means the work was
// optimised away, the timer has insufficient resolution, or the benchmark body
// never executed.
fn elb_implausibly_flat(vals: [Int]) -> Bool {
let n: Int = native_list_len(vals)
if n < 3 { return false }
let first: Int = native_list_get(vals, 0)
let last: Int = native_list_get(vals, n - 1)
if first == 0 {
if last == 0 { return true }
return false
}
let r: Int = (last * 1000) / first
if r < 1100 { return true }
return false
}
// elb_spread_ok do the consecutive ratios agree with each other?
//
// This is the ratio-method analogue of a normalised-RMS threshold. If the
// doublings disagree wildly the data is noise, a cache cliff, or a phase
// change, and the honest report is INDETERMINATE rather than a classification.
// Applies to the ASYMPTOTIC TAIL only the last three ratios.
//
// The small-n end of any sweep is dominated by fixed overhead, cold caches and
// branch predictors that have not warmed. Measured on a genuinely linear
// character scan, the ratios ran 3.37, 2.92, 1.76, 1.65: the head looks
// quadratic, the tail is the truth. Checking spread across the whole sweep
// therefore rejects correct data. A complexity bound is an asymptotic claim, so
// it is judged on the asymptotic region the same reason a benchmark harness
// discards warmup rather than averaging it in.
fn elb_spread_ok(ratios: [Int]) -> Bool {
let total: Int = native_list_len(ratios)
if total < 2 { return true }
let start: Int = total - 3
if start < 0 { let start = 0 }
let n: Int = total
let lo: Int = 999999
let hi: Int = 0
let i: Int = start
while i < n {
let r: Int = native_list_get(ratios, i)
if r >= 0 {
if r < lo { let lo = r }
if r > hi { let hi = r }
}
let i = i + 1
}
if lo <= 0 { return false }
// Reject when the widest ratio is more than 2.2x the narrowest. That is
// enough slack for real timing noise and tight enough to separate a clean
// 2.0 series from a clean 4.0 series.
if (hi * 1000) / lo > 2200 { return false }
return true
}
// elb_ratios consecutive milli-ratios across the sweep.
fn elb_ratios(vals: [Int]) -> [Int] {
let out: [Int] = native_list_empty()
let n: Int = native_list_len(vals)
let i: Int = 1
while i < n {
let out = native_list_append(out,
elb_ratio(native_list_get(vals, i - 1), native_list_get(vals, i)))
let i = i + 1
}
return out
}
// elb_mean_tail_ratio mean of the LAST TWO ratios.
//
// The tail is used deliberately: asymptotic behaviour is what a complexity
// bound claims, and the small-n end of any sweep is dominated by fixed
// overhead. This is the same reason a benchmark harness discards warmup.
fn elb_mean_tail_ratio(ratios: [Int]) -> Int {
let n: Int = native_list_len(ratios)
if n == 0 { return -1 }
if n == 1 { return native_list_get(ratios, 0) }
let a: Int = native_list_get(ratios, n - 1)
let b: Int = native_list_get(ratios, n - 2)
if a < 0 { return b }
if b < 0 { return a }
return (a + b) / 2
}
// Verdicts
//
// 0 PASS measured curve is at or below the declared bound
// 1 FAIL measured curve is strictly worse than declared
// 2 INDETERMINATE ratios disagree; data is noise or a phase change
// 3 REFUSED below the measurement floor, or implausibly flat
// 4 BETTER measured strictly better than declared (warn, not fail)
fn elb_verdict_name(v: Int) -> String {
if v == 0 { return "PASS" }
if v == 1 { return "FAIL" }
if v == 2 { return "INDETERMINATE" }
if v == 3 { return "REFUSED" }
if v == 4 { return "BETTER" }
return "?"
}
// elb_gate classify one signal against its declared bound.
//
// vals measurements, one per sweep point, in sweep order
// expect declared curve index (see elb_curve_name)
// floor minimum largest-measurement below which we refuse to classify
fn elb_gate(vals: [Int], expect: Int, floor: Int) -> Int {
if elb_below_floor(vals, floor) { return 3 }
if elb_implausibly_flat(vals) { return 3 }
let ratios: [Int] = elb_ratios(vals)
if !elb_spread_ok(ratios) { return 2 }
let m: Int = elb_mean_tail_ratio(ratios)
if m < 0 { return 2 }
let got: Int = elb_classify_ratio(m)
if got > expect { return 1 }
if got < expect { return 4 }
return 0
}
// elb_measured_curve the classified curve for a signal, or -1 if unclassifiable.
fn elb_measured_curve(vals: [Int], floor: Int) -> Int {
if elb_below_floor(vals, floor) { return -1 }
if elb_implausibly_flat(vals) { return -1 }
let ratios: [Int] = elb_ratios(vals)
let m: Int = elb_mean_tail_ratio(ratios)
if m < 0 { return -1 }
return elb_classify_ratio(m)
}
+194
View File
@@ -0,0 +1,194 @@
// runtime/eltest.el El test framework runner (Phase 1).
//
// This is the RUNNER. It is written in El and consumes a registry that the
// compiler generates into the same translation unit when invoked as
// `elc --test`. Nothing here discovers tests; discovery already happened at
// compile time, which is what makes `--list` and filtering possible later.
//
// Architecture
//
// The compiler lowers each `test "name" { ... }` block into a static C
// function and emits a static table of (name, fn) pairs plus a small set of
// index-based accessors. El has no function pointers, so the runner never
// sees one it works entirely in indices:
//
// __el_reg_count() -> Int number of registered tests
// __el_reg_name(i) -> String test name at index i
// __el_reg_invoke(i) -> Int run test i, return its failure count
// __el_reg_last_ns() -> Int wall-clock ns of the last invoke
// __el_reg_msg() -> String first failure message of the last invoke
// __el_reg_asserts() -> Int assertions executed in the last invoke
// __el_opt_json() -> Int 1 if --json was passed
//
// Timing is taken in the generated C, immediately around the call, so no El
// call overhead lands inside the measurement.
//
// Output
//
// Structured events are the source of truth. The human renderer is written
// FROM the same fields the NDJSON renderer emits never the reverse. Parsing
// human output back into structure is the one clear architectural mistake in
// Go's test tooling and we do not repeat it.
//
// Every result carries a duration. Always. A framework that cannot report how
// long its tests took cannot surface a performance regression, and a
// regression nobody can see is one nobody fixes.
// Small helpers (no imports this file must stay self-contained)
// _elt_json_escape minimal JSON string escaping for the NDJSON renderer.
fn _elt_json_escape(s: String) -> String {
let out: String = ""
let n: Int = str_len(s)
let i: Int = 0
while i < n {
let ch: String = str_slice(s, i, i + 1)
if str_eq(ch, "\"") {
let out = out + "\\\""
} else {
if str_eq(ch, "\\") {
let out = out + "\\\\"
} else {
if str_eq(ch, "\n") {
let out = out + "\\n"
} else {
if str_eq(ch, "\t") {
let out = out + "\\t"
} else {
if str_eq(ch, "\r") {
let out = out + "\\r"
} else {
let out = out + ch
}
}
}
}
}
let i = i + 1
}
return out
}
// _elt_pad3 left-pad an integer to three digits (for the ms.fraction form).
fn _elt_pad3(v: Int) -> String {
if v < 10 { return "00" + int_to_str(v) }
if v < 100 { return "0" + int_to_str(v) }
return int_to_str(v)
}
// _elt_ms render a nanosecond duration as "M.mmm" milliseconds.
//
// Deliberately avoids the modulo operator: the remainder is derived by
// subtraction so this stays portable across El backends.
fn _elt_ms(ns: Int) -> String {
let total_us: Int = ns / 1000
let ms_whole: Int = total_us / 1000
let us_rem: Int = total_us - (ms_whole * 1000)
return int_to_str(ms_whole) + "." + _elt_pad3(us_rem)
}
// _elt_secs render a nanosecond duration as fractional seconds, for the
// NDJSON `elapsed` field. JUnit XML and test2json both use seconds-as-decimal.
fn _elt_secs(ns: Int) -> String {
let total_ms: Int = ns / 1000000
let s_whole: Int = total_ms / 1000
let ms_rem: Int = total_ms - (s_whole * 1000)
return int_to_str(s_whole) + "." + _elt_pad3(ms_rem)
}
// Event emission
//
// One function per event shape. Both renderers read the same fields; the
// human renderer is a projection of the event, not a separate code path.
fn _elt_emit_run(json_mode: Bool, name: String) {
if json_mode {
println("{\"action\":\"run\",\"test\":\"" + _elt_json_escape(name) + "\"}")
}
}
fn _elt_emit_result(json_mode: Bool, name: String, fails: Int, ns: Int, asserts: Int, msg: String) {
if json_mode {
let action: String = "pass"
if fails > 0 { let action = "fail" }
let line: String = "{\"action\":\"" + action + "\""
let line = line + ",\"test\":\"" + _elt_json_escape(name) + "\""
let line = line + ",\"elapsed\":" + _elt_secs(ns)
let line = line + ",\"assertions\":" + int_to_str(asserts)
if fails > 0 {
let line = line + ",\"failures\":" + int_to_str(fails)
let line = line + ",\"message\":\"" + _elt_json_escape(msg) + "\""
}
let line = line + "}"
println(line)
return
}
// Human renderer duration is never optional.
if fails > 0 {
println("FAIL " + name + " (" + _elt_ms(ns) + "ms)")
println(" " + msg)
return
}
println("ok " + name + " (" + _elt_ms(ns) + "ms)")
return
}
fn _elt_emit_summary(json_mode: Bool, total: Int, failed: Int, ns: Int, asserts: Int) {
let passed: Int = total - failed
if json_mode {
let line: String = "{\"action\":\"summary\""
let line = line + ",\"tests\":" + int_to_str(total)
let line = line + ",\"passed\":" + int_to_str(passed)
let line = line + ",\"failed\":" + int_to_str(failed)
let line = line + ",\"assertions\":" + int_to_str(asserts)
let line = line + ",\"elapsed\":" + _elt_secs(ns)
let line = line + "}"
println(line)
return
}
println("")
println(int_to_str(total) + " tests, " + int_to_str(passed) + " passed, "
+ int_to_str(failed) + " failed, " + int_to_str(asserts) + " assertions in "
+ _elt_ms(ns) + "ms")
return
}
// The runner
// el_test_main drive the compile-time registry.
//
// Called from the generated main(). Returns the number of FAILING TESTS, which
// becomes the process exit code. Note that this counts tests, not assertions:
// a test is the unit of result. The old harness counted assertions globally and
// therefore could not say which test failed, how long any of them took, or
// whether a test had run at all.
fn el_test_main() -> Int {
let json_mode: Bool = false
if __el_opt_json() == 1 { let json_mode = true }
let n: Int = __el_reg_count()
let i: Int = 0
let failed: Int = 0
let total_ns: Int = 0
let total_asserts: Int = 0
while i < n {
let name: String = __el_reg_name(i)
_elt_emit_run(json_mode, name)
let fails: Int = __el_reg_invoke(i)
let ns: Int = __el_reg_last_ns()
let asserts: Int = __el_reg_asserts()
let msg: String = __el_reg_msg()
let total_ns = total_ns + ns
let total_asserts = total_asserts + asserts
if fails > 0 { let failed = failed + 1 }
_elt_emit_result(json_mode, name, fails, ns, asserts, msg)
let i = i + 1
}
_elt_emit_summary(json_mode, n, failed, total_ns, total_asserts)
return failed
}
+372 -29
View File
@@ -246,14 +246,6 @@ static int put_edge(EngramPagedStore* s, const char* id, const char* from, const
e.metadata = (char*)meta;
return store_put_edge(s, &e);
}
int cog_ground_edge(EngramPagedStore* s, const char* claim_id,
const char* evidence_id, double grounding, const char* for_whom) {
if (!s || !claim_id || !evidence_id) return -1;
char id[512], meta[256];
snprintf(id, sizeof id, "gb-%s-%s-%s", claim_id, evidence_id, for_whom ? for_whom : "global");
snprintf(meta, sizeof meta, "for_whom=%s", for_whom ? for_whom : "-");
return put_edge(s, id, claim_id, evidence_id, COG_GROUNDED_BY_RELATION, grounding, meta);
}
int cog_salient_edge(EngramPagedStore* s, const char* node_id,
const char* observer_id, double salience) {
if (!s || !node_id || !observer_id) return -1;
@@ -261,35 +253,386 @@ int cog_salient_edge(EngramPagedStore* s, const char* node_id,
snprintf(id, sizeof id, "st-%s-%s", node_id, observer_id);
return put_edge(s, id, node_id, observer_id, COG_SALIENT_TO_RELATION, salience, NULL);
}
int cog_assert_gate(EngramPagedStore* s, const char* claim_id,
const char* for_whom, double floor) {
if (!s || !claim_id) return -1;
if (!(floor > 0)) floor = 0.5;
StoreEdge* edges = NULL; size_t n = 0;
if (store_get_edges_from(s, claim_id, &edges, &n) < 0) return -1;
double best = 0.0; int found = 0;
for (size_t i = 0; i < n; i++) {
if (!edges[i].relation || strcmp(edges[i].relation, COG_GROUNDED_BY_RELATION) != 0) continue;
/* grounded-for-whom: match observer if requested; global (for_whom=-) always counts */
int match = 1;
if (for_whom && edges[i].metadata) {
const char* fw = strstr(edges[i].metadata, "for_whom=");
if (fw) { fw += 9; if (strcmp(fw, for_whom) != 0 && strcmp(fw, "-") != 0) match = 0; }
}
if (match) { found = 1; if (edges[i].weight > best) best = edges[i].weight; }
}
store_edges_free(edges, n);
if (!found) return 0; /* ungrounded => refuse assertion (still held) */
return (best >= floor) ? 1 : 0;
/* ═══════════════════════════════════════════════════════════════════════════
* §7 GROUNDING IS THE EDGE'S WEIGHT, AND THE WEIGHT IS A VECTOR.
* See engram_cognition.h §7 for the model and for the measurements the two
* design decisions (thirteen regions, min aggregate) rest on.
* */
/* ── The one decay model. Moved here verbatim from el_runtime.c's
* engram_temporal_decay so nodes and edges share a single implementation and a
* single set of constants; engram_temporal_decay now delegates. Bit-identical
* for nodes: reinforcements := activation_count, lambda_override :=
* temporal_decay_rate.
*
* This is what makes decay ANALYTIC rather than sampled: between two recorded
* versions the trajectory is not unknown, it is known in closed form from the
* last point and elapsed time. Store the point, read the curve. */
double cog_decay_factor(int64_t age_ms, double reinforcements, double lambda_override) {
if (age_ms <= 0) return 1.0;
double lambda = (lambda_override > 0.0) ? lambda_override : COG_DECAY_LAMBDA;
double age_hours = (double)age_ms / 3600000.0;
if (reinforcements < 0) reinforcements = 0;
double t_half = COG_T_HALF_HOURS * (1.0 + log(1.0 + reinforcements));
double factor = exp(-lambda * age_hours / t_half);
if (factor < COG_DECAY_FLOOR) factor = COG_DECAY_FLOOR;
return factor;
}
const char* cog_prov_name(CogProvClass p) {
switch (p) {
case COG_PROV_OBSERVED: return "observed";
case COG_PROV_INFERRED: return "inferred";
case COG_PROV_TOLD: return "told";
case COG_PROV_IMPRINTED: return "imprinted";
default: return "unset";
}
}
CogProvClass cog_prov_parse(const char* s) {
if (!s) return COG_PROV_UNSET;
if (!strcmp(s, "observed")) return COG_PROV_OBSERVED;
if (!strcmp(s, "inferred")) return COG_PROV_INFERRED;
if (!strcmp(s, "told")) return COG_PROV_TOLD;
if (!strcmp(s, "imprinted")) return COG_PROV_IMPRINTED;
return COG_PROV_UNSET;
}
/* Locate the GRD1 block in an edge's metadata. It is always the tail; anything
* ahead of it is the edge's pre-existing metadata, preserved verbatim. */
static const char* cog_grd_find(const char* meta) {
if (!meta) return NULL;
size_t ml = strlen(COG_GROUNDING_META_MAGIC);
if (strncmp(meta, COG_GROUNDING_META_MAGIC, ml) == 0) return meta;
const char* p = meta;
while ((p = strstr(p, COG_GROUNDING_META_MAGIC)) != NULL) {
if (p > meta && p[-1] == '\n') return p;
p += ml;
}
return NULL;
}
int cog_grounding_parse(const StoreEdge* e, int64_t now_ms, CogGrounding* out) {
if (!e || !out) return -1;
memset(out, 0, sizeof *out);
/* Two dimensions exist on every edge whether or not grounding has ever been
* established, because they ARE existing substrate rather than new fields:
* associative the accrued hebb, with its existing dynamics;
* polarity the signed authored weight. `inhibitory` is precisely this
* distinction crushed to one bit, so it is the seed sign. */
out->associative = e->hebb;
out->polarity = e->inhibitory ? -e->weight : e->weight;
out->prov = COG_PROV_UNSET;
out->ts = e->last_fired > 0 ? e->last_fired : e->updated_at;
const char* blk = cog_grd_find(e->metadata);
if (blk) {
out->present = 1;
char* copy = dupstr(blk);
if (!copy) return -1;
for (char* line = strtok(copy, "\n"); line; line = strtok(NULL, "\n")) {
if (line[0] == '\0') continue;
char tag = line[0];
const char* rest = line + 1; while (*rest == ' ') rest++;
if (tag == 'w') { /* the four numeric dimensions */
double v[4] = {0,0,0,0}; parse_floats(rest, v, 4);
out->factual = v[0]; out->relational = v[1];
out->associative = v[2]; out->polarity = v[3];
} else if (tag == 'k') { /* provenance class */
out->prov = cog_prov_parse(rest);
} else if (tag == 't') { /* timestamp + seq + reinforcements */
double v[3] = {0,0,0}; parse_floats(rest, v, 3);
out->ts = (int64_t)v[0]; out->seq = (int64_t)v[1]; out->reinforcements = v[2];
} else if (tag == 'd') {
double v[3] = {0,0,0}; parse_floats(rest, v, 3);
out->fac_proj = v[0]; out->rel_proj = v[1]; out->cos_angle = v[2];
} else if (tag == 'v') {
snprintf(out->binding_value, sizeof out->binding_value, "%s", rest);
} else if (tag == 'c') {
double v[2] = {0,0}; parse_floats(rest, v, 2);
out->floor_at_record = v[0]; out->rel_floor_at_record = v[1];
} else if (tag == 'p') {
snprintf(out->prev_edge, sizeof out->prev_edge, "%s", rest);
}
}
free(copy);
}
out->agreement = (out->cos_angle > 0) ? 1 : (out->cos_angle < 0 ? -1 : 0);
/* ── DERIVED. Nothing below this line is ever serialized. Recency, decay and
* staleness are read off the curve; storing them is how a number ends up
* asserting something nothing computed (§8.1 / spec §2). */
out->age_ms = (out->ts > 0 && now_ms > out->ts) ? (now_ms - out->ts) : 0;
out->decay = cog_decay_factor(out->age_ms, out->reinforcements, 0.0);
out->factual_now = out->factual * out->decay;
out->relational_now = out->relational * out->decay;
out->associative_now = out->associative * out->decay;
out->stale = (out->present && out->floor_at_record > 0 &&
out->factual_now < out->floor_at_record) ? 1 : 0;
return 0;
}
char* cog_grounding_metadata(const char* base_meta, const CogGrounding* g) {
if (!g) return NULL;
size_t keep = 0;
if (base_meta) {
const char* blk = cog_grd_find(base_meta);
keep = blk ? (size_t)(blk - base_meta) : strlen(base_meta);
while (keep > 0 && base_meta[keep - 1] == '\n') keep--;
}
size_t cap = keep + 1024;
char* buf = malloc(cap); if (!buf) return NULL;
size_t o = 0;
if (keep) { memcpy(buf, base_meta, keep); o = keep; buf[o++] = '\n'; }
o += (size_t)snprintf(buf + o, cap - o, "%s\n", COG_GROUNDING_META_MAGIC);
/* STORED ONLY. factual / relational / associative / polarity / provenance /
* timestamp plus the joint state a decision saw. No confidence, no
* recency, no staleness, no volatility: those are read off the curve. */
o += (size_t)snprintf(buf + o, cap - o, "w %.9g %.9g %.9g %.9g\n",
g->factual, g->relational, g->associative, g->polarity);
o += (size_t)snprintf(buf + o, cap - o, "k %s\n", cog_prov_name(g->prov));
o += (size_t)snprintf(buf + o, cap - o, "t %lld %lld %.9g\n",
(long long)g->ts, (long long)g->seq, g->reinforcements);
o += (size_t)snprintf(buf + o, cap - o, "d %.9g %.9g %.9g\n",
g->fac_proj, g->rel_proj, g->cos_angle);
o += (size_t)snprintf(buf + o, cap - o, "v %s\n", g->binding_value[0] ? g->binding_value : "-");
o += (size_t)snprintf(buf + o, cap - o, "c %.9g %.9g\n", g->floor_at_record, g->rel_floor_at_record);
if (g->prev_edge[0]) o += (size_t)snprintf(buf + o, cap - o, "p %s\n", g->prev_edge);
(void)o;
return buf;
}
/* ── Consequence, not epsilon. Every test is a floor crossing or a sign change,
* both exact. Ordered so the two INHERENT (discrete) moves are reported in
* preference to the graded ones, because they bypass the salience gate. */
CogSignificance cog_grounding_significant(const CogGrounding* prev,
const CogGrounding* now,
double floor, double rel_floor) {
if (!now) return COG_SIG_NONE;
if (!prev || !prev->present) return COG_SIG_FIRST_RECORD;
/* INHERENT 1 — polarity sign flip. Ignorance and disagreement are different
* states, and support contradiction is a change of state rather than a
* drift, so no threshold applies. Comparing signs, with zero its own class. */
{
int sp = prev->polarity > 0 ? 1 : (prev->polarity < 0 ? -1 : 0);
int sn = now->polarity > 0 ? 1 : (now->polarity < 0 ? -1 : 0);
if (sp != sn) return COG_SIG_POLARITY_FLIP;
}
/* INHERENT 2 — provenance class change. told → observed is a categorical
* upgrade in what the relation is entitled to, not a movement along an axis. */
if (prev->prov != now->prov) return COG_SIG_PROVENANCE_CHANGE;
/* Crossing an assert floor — the move changes whether this relation can be
* spoken. Compared on the DECAYED values, because that is what the gate reads. */
if ((prev->factual_now >= floor) != (now->factual_now >= floor)) return COG_SIG_FACTUAL_FLOOR;
if ((prev->relational_now >= rel_floor) != (now->relational_now >= rel_floor)) return COG_SIG_RELATIONAL_FLOOR;
/* Flipping factual/relational agreement — the relation stops being "true and
* meaningful" and becomes "true and misapplied", or the reverse. This is the
* 911/CPS contradiction as a measured event rather than a reviewable one. */
if (prev->agreement != now->agreement) return COG_SIG_AGREEMENT_FLIP;
/* A gradient reversing — the evidence stopped pulling the claim toward it and
* began pushing it away, or the same on the values axis. */
if ((prev->fac_proj > 0) != (now->fac_proj > 0)) return COG_SIG_DIRECTION_REVERSAL;
if ((prev->rel_proj > 0) != (now->rel_proj > 0)) return COG_SIG_DIRECTION_REVERSAL;
return COG_SIG_NONE;
}
int cog_significance_inherent(CogSignificance s) {
return (s == COG_SIG_FIRST_RECORD || s == COG_SIG_POLARITY_FLIP ||
s == COG_SIG_PROVENANCE_CHANGE) ? 1 : 0;
}
const char* cog_significance_name(CogSignificance s) {
switch (s) {
case COG_SIG_FIRST_RECORD: return "first-record";
case COG_SIG_POLARITY_FLIP: return "polarity-sign-flip";
case COG_SIG_PROVENANCE_CHANGE: return "provenance-class-change";
case COG_SIG_FACTUAL_FLOOR: return "factual-floor-crossed";
case COG_SIG_RELATIONAL_FLOOR: return "relational-floor-crossed";
case COG_SIG_AGREEMENT_FLIP: return "agreement-sign-flip";
case COG_SIG_DIRECTION_REVERSAL: return "gradient-direction-reversal";
default: return "none";
}
}
/* ── Recording: a NEW edge record. The predecessor is never touched. ────────── */
int cog_grounding_record(EngramPagedStore* s, const StoreEdge* base,
const CogGrounding* g, char* out_id, size_t out_id_cap) {
if (!s || !base || !base->id || !g) return -1;
char root[192];
snprintf(root, sizeof root, "%s", base->id);
char* hash = strchr(root, '#'); if (hash) *hash = '\0';
int seq = (int)g->seq + 1;
char vid[224];
snprintf(vid, sizeof vid, "%s#%d", root, seq);
CogGrounding rec = *g;
rec.seq = seq;
snprintf(rec.prev_edge, sizeof rec.prev_edge, "%s", base->id);
char* meta = cog_grounding_metadata(base->metadata, &rec);
if (!meta) return -1;
StoreEdge e; memset(&e, 0, sizeof e);
e.id = vid; e.from_id = base->from_id; e.to_id = base->to_id;
e.relation = base->relation; e.metadata = meta;
/* The vector IS the weight, so the scalar fields carry their dimensions:
* `weight` the magnitude of polarity, `inhibitory` its sign, `hebb` the
* associative strength. Nothing here is a second copy of a derived value. */
e.weight = rec.polarity < 0 ? -rec.polarity : rec.polarity;
e.inhibitory = rec.polarity < 0 ? 1 : 0;
e.hebb = rec.associative;
e.confidence = base->confidence;
e.created_at = base->created_at;
e.updated_at = rec.ts;
e.last_fired = rec.ts;
e.layer_id = base->layer_id;
int rc = store_put_edge(s, &e);
free(meta);
if (rc != 0) return -1;
if (out_id && out_id_cap) snprintf(out_id, out_id_cap, "%s", vid);
return seq;
}
int cog_grounding_head(EngramPagedStore* s, const char* base_id,
StoreEdge* out, int max_versions) {
if (!s || !base_id || !out) return -1;
if (max_versions <= 0) max_versions = 64;
char root[192]; snprintf(root, sizeof root, "%s", base_id);
char* hash = strchr(root, '#'); if (hash) *hash = '\0';
StoreEdge cur; memset(&cur, 0, sizeof cur);
if (store_get_edge(s, root, &cur) != 1) return -1;
int found = 0;
for (int v = 1; v <= max_versions; v++) {
char vid[224]; snprintf(vid, sizeof vid, "%s#%d", root, v);
StoreEdge nx;
if (store_get_edge(s, vid, &nx) != 1) break;
store_edge_free(&cur); cur = nx; found = v;
}
*out = cur;
return found;
}
/* ── VOLATILITY AND DRIFT: derived from the chain, stored nowhere. The series
* exists only because nothing was destroyed, which is the whole return on
* immutability a derivative for free. */
int cog_grounding_trajectory(EngramPagedStore* s, const char* base_id,
int64_t now_ms, CogTrajectory* out) {
if (!s || !base_id || !out) return -1;
memset(out, 0, sizeof *out);
char root[192]; snprintf(root, sizeof root, "%s", base_id);
char* hash = strchr(root, '#'); if (hash) *hash = '\0';
double pf = 0, pr = 0, f0 = 0, r0 = 0, fN = 0, rN = 0;
double sum_df = 0, sum_dr = 0;
int n = 0;
for (int v = 0; v <= 64; v++) {
char vid[224];
if (v == 0) snprintf(vid, sizeof vid, "%s", root);
else snprintf(vid, sizeof vid, "%s#%d", root, v);
StoreEdge e;
if (store_get_edge(s, vid, &e) != 1) { if (v) break; else continue; }
CogGrounding g;
if (cog_grounding_parse(&e, now_ms, &g) == 0) {
if (n == 0) { f0 = g.factual; r0 = g.relational; }
else { sum_df += fabs(g.factual - pf); sum_dr += fabs(g.relational - pr); }
pf = g.factual; pr = g.relational; fN = pf; rN = pr;
n++;
}
store_edge_free(&e);
}
out->n_versions = n;
if (n > 1) {
out->factual_volatility = sum_df / (double)(n - 1);
out->relational_volatility = sum_dr / (double)(n - 1);
}
out->factual_drift = fN - f0;
out->relational_drift = rN - r0;
/* "STAYED TRUE, BECAME WRONG" — the event the joint record makes visible and
* that per-dimension versioning would have destroyed: the fact held while
* the meaning degraded. Expressed as signs, so there is no tolerance here
* either: factual did not fall, relational did. */
out->stayed_true_became_wrong =
(n > 1 && out->factual_drift >= 0 && out->relational_drift < 0) ? 1 : 0;
return 0;
}
/* ── Assertion gates on BOTH floors. Traversal is untouched: activation still
* conducts on the factual/associative side, so a relation can remain thinkable
* while ceasing to be assertable. That gap is where the wide angles live. */
int cog_assert_two_axis(EngramPagedStore* s, const char* claim_id,
double floor, double rel_floor, int64_t now_ms,
CogAssertion* out) {
if (!s || !claim_id || !out) return -1;
memset(out, 0, sizeof *out);
if (!(floor > 0)) floor = 0.5;
if (!(rel_floor > 0)) rel_floor = floor;
/* still_held is DERIVED, not a literal (§8.1). Holding is unconditional —
* the store gates nothing so the question the field actually answers is
* whether the content is present and live. */
StoreNode n;
if (store_get_node(s, claim_id, &n) == 1) { out->still_held = !n.tombstoned; store_node_free(&n); }
else out->still_held = 0;
double best = -1.0;
for (int dir = 0; dir < 2; dir++) {
StoreEdge* edges = NULL; size_t ne = 0;
int rc = dir == 0 ? store_get_edges_from(s, claim_id, &edges, &ne)
: store_get_edges_to (s, claim_id, &edges, &ne);
if (rc < 0) continue;
for (size_t i = 0; i < ne; i++) {
if (edges[i].tombstoned) continue;
CogGrounding g;
if (cog_grounding_parse(&edges[i], now_ms, &g) != 0) continue;
out->n_edges++;
out->found = 1;
if (g.factual_now > best) {
best = g.factual_now;
out->factual = g.factual_now;
out->relational = g.relational_now; /* the SAME edge, not a max */
out->polarity = g.polarity;
out->cos_angle = g.cos_angle;
out->agreement = g.agreement;
out->prov = g.prov;
out->relational_established = g.present;
snprintf(out->best_edge, sizeof out->best_edge, "%s", edges[i].id ? edges[i].id : "");
snprintf(out->binding_value, sizeof out->binding_value, "%s", g.binding_value);
}
}
store_edges_free(edges, ne);
}
/* BOTH floors, and an unestablished relational axis does NOT pass by default
* defaulting it to passing is the exemption §0 forbids. A negative polarity
* is a relation that actively contradicts and can never license assertion. */
out->may_assert = (out->found && out->relational_established &&
out->polarity > 0 &&
out->factual >= floor && out->relational >= rel_floor) ? 1 : 0;
return 0;
}
/* ═══════════════════════════════════════════════ THE CORRESPONDENCE-LOOP ═════ */
int engram_correspondence_beat(const GeoDescriptor* region, const float* anchor,
double outcome_y, CogStance* stance,
int learn, double max_step, CogBeatResult* out) {
if (!region || !stance || !out) return -1;
memset(out, 0, sizeof *out);
if (stance->keystone) { learn = 0; out->wrote_keystone = 1; } /* §6: never write a keystone */
/* 2026-08-16: the keystone block is GONE. It refused to learn about the
* reference frame, which does not make it a good reference it makes it
* unexaminable, trading circular calibration for an ungroundable one (spec
* §2). Measured cost of the block: on the keystone region the beat reported
* 0.00% brier reduction over n_trials 0 it never ran, so nothing about the
* self was ever calibrated OR falsifiable. What replaces it is a provenance
* constraint, not a permission: cog_grounding_downstream refuses evidence
* that is downstream of the region being calibrated, for every region alike.
* `wrote_keystone` is retained as a reporting field only and is always 0. */
GeoGradient g;
if (engram_think(region, anchor, stance, &g) != 0) return -1; /* PREDICTION */
+269 -17
View File
@@ -23,7 +23,7 @@
* and a region, and grounded-for-whom.
*
* PURE + (mostly) READ-ONLY, stdlib + libm only. think() and the warp are pure
* over their inputs. Persistence (Stance <-> StoreNode, grounded-by edges) is the
* over their inputs. Persistence (Stance <-> StoreNode, edge grounding vectors) is the
* only part that touches the store, and it is additive / supersede / tombstone
* never mutate-in-place, never delete. It NEVER touches the live daemon: all
* offline against a scratch store, per the design's rails.
@@ -152,30 +152,25 @@ int engram_express(const GeoGradient* g, const float* anchor, float* out_point);
/* ═══════════════════════════════════════════════════════════════════════════
* §5 HOLD vs GROUND vs ASSERT. Holding is unconditional (the store gates nothing).
* Grounding is a RELATION a "grounded-by" edge, probabilistic, grounded-for-whom.
* The honesty floor is checked only at ASSERTION.
* Grounding is an ATTRIBUTE OF a relation carried on the edge itself, as a
* vector (§7). The honesty floor is checked only at ASSERTION, on both axes.
* */
#define COG_GROUNDED_BY_RELATION "grounded-by"
/* DELETED 2026-08-16: COG_GROUNDED_BY_RELATION and cog_ground_edge.
*
* A "grounded-by" edge models grounding as a relation BETWEEN two nodes. It is a
* property OF a relation and it is that relation's weight. Minting a new edge
* to carry a score was the error; #147 corrected which endpoints the edge landed
* on and left the wrong idea standing. There is nothing to ground a claim
* "against" that is not already an edge, and if no edge exists the honest answer
* is that the two are not related not a freshly minted one scoring 0.98.
* See §7 for what replaced it. */
#define COG_SALIENT_TO_RELATION "salient-to"
/* Write a grounded-by edge (additive). weight = grounding ∈(0,1] from the verifier;
* for_whom recorded in edge metadata (grounding is relational). Never a node flag. */
int cog_ground_edge(EngramPagedStore* s, const char* claim_id,
const char* evidence_id, double grounding, const char* for_whom);
/* Write/refresh a salient-to edge: salience is RELATIONAL (grounded-for-whom),
* carried on the edge to the observer not baked into the node scalar (§2.1). */
int cog_salient_edge(EngramPagedStore* s, const char* node_id,
const char* observer_id, double salience);
/* The honesty floor — a QUERY at assertion time, NOT a schema constraint. Reads the
* claim's stored grounded-by edges (for the given observer) and returns:
* 1 = may assert (best grounding >= floor),
* 0 = REFUSE assertion (holds unconditionally; only asserting is gated),
* <0 = error. The content remains held either way. */
int cog_assert_gate(EngramPagedStore* s, const char* claim_id,
const char* for_whom, double floor);
/* ═══════════════════════════════════════════════════════════════════════════
* §4 THE REFLEXIVE CORRESPONDENCE-LOOP the learning engine. think scores its
* OWN gradient against outcome, refines the stance on the error, and (optionally)
@@ -209,8 +204,265 @@ int engram_correspondence_beat(const GeoDescriptor* region, const float* anchor,
/* ═══════════════════════════════════════════════════════════════════════════
* §6 METASTABILITY. Keystones (self/values) are read-mostly: the loop reads but
* never writes them. Mark by stance flag or by a keystone-id set the loop consults.
*
* SUPERSEDED BY §7's PROVENANCE CONSTRAINT (2026-08-16). The keystone flag is a
* PERMISSION: it asks who the target is, not where the evidence came from. That
* is censorship, and it costs the ability to ever ground the self (spec
* correspondence-and-censorship.md §0/§2). The constraint that actually protects
* a reference frame is cog_grounding_downstream: a region may not be calibrated
* by evidence downstream of itself. These declarations remain only so existing
* call sites keep compiling; nothing in the grounding path consults them.
* */
typedef struct { const char** ids; int n; } CogKeystoneSet;
int cog_is_keystone(const CogKeystoneSet* ks, const CogStance* s);
/* ═══════════════════════════════════════════════════════════════════════════
* §7 GROUNDING IS THE EDGE'S WEIGHT, AND THE WEIGHT IS A VECTOR
* (2026-08-16; spec correspondence-and-censorship.md §2§6 @ 2b7e4ba.)
*
* THE MODEL. Grounding is not a subsystem, a score, or a relation BETWEEN nodes.
* It is an attribute OF a relation. The graph already IS the grounding structure:
* every edge is a grounded relation, and what that relation is worth is carried
* on the edge itself. Three things follow, and each DELETES rather than adds:
*
* 1. `grounded-by` as a relation type does not exist, and cog_ground_edge is
* gone. Minting an edge to hold a score models grounding as a relation
* between nodes when it is a property of a relation. #147 corrected which
* endpoints that edge landed on and left the wrong idea standing.
* 2. There is no observer, and no sampling rate. Change is not a consequence of
* use it IS use, the way potentiation is the firing rather than something
* that reads the firing and writes a weight. So no supervisor compares a
* value to a threshold and decides to persist.
* 3. Between two recorded versions the trajectory is not unknown. Decay is a
* pure function of the last recorded point and elapsed time, so it is
* ANALYTIC: store the point, read the curve.
*
* WHAT IS *NOT* HERE, DELIBERATELY. An earlier draft of the spec posed "a graph
* predicate for evidence downstream of itself" as the hard problem, and this file
* briefly contained one. It is withdrawn. Non-circularity is TEMPORAL, not
* topological: you cannot recalibrate the ruler while measuring with it, so you
* do it when you are not using the frame to act. Reachability could never have
* worked measured on the live store, reachability from the self region over
* all relations reaches 89.2% of the graph (10,580 of 11,861 nodes) and 16.0%
* over hebbian/semantic relations alone, so the predicate marks essentially all
* evidence tainted and the constraint degenerates into the total block that
* censorship started as. Nothing replaces it here; the independence is a fact
* about engagement, owned by the dreamer, not a fact about the graph.
*
*
* §7.1 THE VECTOR
*
* The test for a real dimension is whether it can move independently of the
* others. Five can, and each maps onto substrate that already exists:
*
* factual correspondence with evidence. [GRD1]
* relational correspondence with values min over THIRTEEN
* value regions, carrying the binding value's NAME. [GRD1]
* associative co-activation frequency. This is the edge's `hebb`
* field with its existing dynamics NOT a new one.
* Independent by construction: every superstition is
* a strong association with no factual grounding.
* polarity SIGNED. Near zero means "no support"; NEGATIVE means
* "this actively contradicts". The edge's `inhibitory`
* bit is exactly this distinction crushed to one bit,
* and is carried forward as the seed value. [GRD1]
* provenance observed / inferred / told / imprinted. Categorical,
* and load-bearing: it governs what the relation is
* entitled to. [GRD1]
*
* Plus a TIMESTAMP, which is what turns the supersession chain into a time
* series of vectors rather than a series of numbers.
*
* DERIVED, THEREFORE NEVER STORED. Confidence (high grounding AND low
* volatility), recency (decay read off the curve), staleness (grounding fallen
* below its floor), volatility (the derivative of a series nothing destroyed).
* Storing confidence separately is how `confidence: 0.5` ends up sitting beside
* a zero direction vector, asserting something nothing computed. Every field in
* CogGrounding below is marked STORED or DERIVED, and the serializer writes
* only the STORED ones.
*
* THE VALUES REFERENCE IS THIRTEEN REGIONS AND THE AGGREGATE IS MIN.
* Measured on the live store: the values root kn-5b606390 `contains` exactly 13
* value nodes; pairwise centroid cosine among their regions is min 0.1525,
* mean 0.5199, median 0.5282, max 0.9278 they demonstrably do not form one
* region. Against a single union region the individual values sit at cosine
* 0.38..0.89, with constraints-as-freedom at 0.3812 and change-is-the-signal at
* 0.4677, so a union centroid under-represents precisely the values a claim is
* most likely to be measured against. MIN rather than MEAN because a mean lets
* strong agreement with twelve values mask a violation of the thirteenth, which
* is the mechanism of rationalization; min yields a binding constraint with a
* NAME attached rather than a score.
*
* TRAVERSAL CONDUCTS ON FACTUAL; ASSERTION REQUIRES BOTH. If activation
* conducted on relational weight, Neuron could not follow a chain of reasoning
* to a conclusion he then rejects censorship arriving through the spreading
* rule. The gap between reachable and assertable is where the wide
* factual/relational angles live, and that gap is the interesting part.
* */
/* ── The one decay model (moved here from el_runtime.c so that node decay and
* edge-grounding decay are a single implementation with a single set of
* constants, rather than a model and a parallel copy of it). Half-life scales
* with how established the thing is: T_eff = T_HALF · (1 + ln(1 + reinforcements)).
* The floor is a preference, not a cliff max penalty for age alone is 4x.
* `lambda_override` > 0 replaces the default rate; 0 means use the default. */
#define COG_T_HALF_HOURS 168.0
#define COG_DECAY_LAMBDA 0.693147
#define COG_DECAY_FLOOR 0.25
double cog_decay_factor(int64_t age_ms, double reinforcements, double lambda_override);
/* The compact vector block carried in the edge's own metadata. Line schema, same
* precedent as STNC1 / GEO1. Metadata the edge already carried is preserved
* verbatim ahead of the magic line. */
#define COG_GROUNDING_META_MAGIC "GRD1"
/* Provenance class — categorical, and it governs what the relation is entitled
* to. A change of class is inherently significant and needs no threshold,
* because told observed is a categorical upgrade, not a drift. */
typedef enum {
COG_PROV_UNSET = 0,
COG_PROV_OBSERVED = 1,
COG_PROV_INFERRED = 2,
COG_PROV_TOLD = 3,
COG_PROV_IMPRINTED = 4
} CogProvClass;
const char* cog_prov_name(CogProvClass p);
CogProvClass cog_prov_parse(const char* s);
typedef struct {
int present; /* 1 iff the edge carries a GRD1 block */
/* ── STORED: the vector, as it stood at `ts` ─────────────────────────────── */
double factual; /* correspondence with evidence */
double relational; /* min over the thirteen value regions */
double associative; /* co-activation frequency — mirrors edge->hebb */
double polarity; /* SIGNED support; <0 = actively contradicts */
CogProvClass prov; /* observed / inferred / told / imprinted */
int64_t ts; /* when this version was recorded (ms) */
int64_t seq; /* supersession sequence number */
double reinforcements; /* uses folded into this version */
char binding_value[128]; /* the argmin value — the conflict's NAME */
/* the two gradients as frame-independent signed projections, plus the angle
* between them in full R^dim. These are part of the JOINT STATE a decision
* saw, not a convenience: near +1 evidence and values push the same way; at
* or below 0 the relation is factually supported and relationally wrong. */
double fac_proj, rel_proj, cos_angle;
int agreement; /* sign(cos_angle): +1 / 0 / 1 */
double floor_at_record, rel_floor_at_record;
char prev_edge[192]; /* the version this superseded ("" if first) */
/* ── DERIVED at read time. NEVER serialized. ─────────────────────────────── */
int64_t age_ms; /* recency: now ts */
double decay; /* cog_decay_factor over that age */
double factual_now; /* factual · decay */
double relational_now;
double associative_now;
int stale; /* grounding fallen below its floor */
} CogGrounding;
/* Read an edge's vector as of `now_ms`. Pure — never writes. An edge with no
* GRD1 block still has an associative strength (its accrued hebb) and a polarity
* (its signed authored weight); `present` says whether the grounding dimensions
* have ever been established, and an unestablished dimension is reported as such
* rather than defaulted to a passing value. */
int cog_grounding_parse(const StoreEdge* e, int64_t now_ms, CogGrounding* out);
/* Serialize the STORED half of the vector, preserving pre-existing non-GRD1
* metadata. Returns an owned string. Derived fields are not written. */
char* cog_grounding_metadata(const char* base_meta, const CogGrounding* g);
/* ── §7.2 CONSOLIDATION-GATED SUPERSESSION ──────────────────────────────────
*
* Supersession is not recording it is CONSOLIDATION, gated by salience, which
* is why you remember the argument and not the commute. Significance is
* evaluated PER-DIMENSION but the record is the WHOLE VECTOR: any dimension
* moving enough to matter triggers a supersession, and the new version captures
* every dimension as it stood at that instant. Versioning axes independently
* would make the joint state unreconstructable, and the joint state is the point
* it is what makes "stayed true, became wrong" visible as an event (factual
* holding steady across versions while relational degrades).
*
* There is deliberately no epsilon in this enum or in the function that computes
* it. Every test is a floor crossing or a sign change, both exact. Two of them
* are INHERENTLY significant because they are discrete state changes rather than
* drift, and those bypass the salience gate entirely. */
typedef enum {
COG_SIG_NONE = 0, /* nothing decision-relevant moved — DO NOT RECORD */
COG_SIG_FIRST_RECORD = 1, /* no prior version exists */
COG_SIG_POLARITY_FLIP = 2, /* INHERENT: support ↔ contradiction, or ignorance
* either. A discrete change of state. */
COG_SIG_PROVENANCE_CHANGE = 3, /* INHERENT: told → observed is a categorical
* upgrade in what the relation is entitled to. */
COG_SIG_FACTUAL_FLOOR = 4, /* crossed the assert floor, factual axis */
COG_SIG_RELATIONAL_FLOOR = 5, /* crossed the assert floor, relational axis */
COG_SIG_AGREEMENT_FLIP = 6, /* factual/relational agreement changed sign */
COG_SIG_DIRECTION_REVERSAL = 7 /* a gradient reversed direction */
} CogSignificance;
CogSignificance cog_grounding_significant(const CogGrounding* prev,
const CogGrounding* now,
double floor, double rel_floor);
const char* cog_significance_name(CogSignificance s);
/* 1 iff this reason is a discrete state change that consolidates regardless of
* salience (polarity flip, provenance change, first record). */
int cog_significance_inherent(CogSignificance s);
/* ── §7.3 RECORDING: supersession of the EDGE, never an overwrite ────────────
* Writes version seq+1 as a NEW edge record with the same endpoints and relation
* and id "<root>#<seq+1>", carrying a GRD1 `p` pointer to its predecessor. The
* predecessor is never touched. The chain IS the trajectory: not only what the
* grounding is but which way it has been moving and how fast a derivative
* obtained for free from immutability, because the points were never destroyed.
* Returns the version written (>=1), or <0 on error. */
int cog_grounding_record(EngramPagedStore* s, const StoreEdge* base,
const CogGrounding* g, char* out_id, size_t out_id_cap);
/* Walk forward from a base edge id to its newest recorded version. Point reads
* only; consolidation is gated, so the chain is short. Returns the highest
* version found (0 = the base record is the only one). */
int cog_grounding_head(EngramPagedStore* s, const char* base_id,
StoreEdge* out, int max_versions);
/* VOLATILITY — derived, never stored: the mean absolute per-version change of a
* dimension across the recorded chain. Feeds the equally-derived `confidence`
* (high grounding AND low volatility), which is likewise never stored. */
typedef struct {
int n_versions;
double factual_volatility;
double relational_volatility;
double factual_drift; /* signed: newest oldest */
double relational_drift;
int stayed_true_became_wrong; /* factual steady while relational degraded */
} CogTrajectory;
int cog_grounding_trajectory(EngramPagedStore* s, const char* base_id,
int64_t now_ms, CogTrajectory* out);
/* ── §7.4 ASSERTION GATES ON BOTH FLOORS ────────────────────────────────────
* A well-evidenced claim must not earn the right to be asserted regardless of
* whether it means the right thing. `may_assert` requires the decayed factual
* grounding to clear `floor` AND the decayed relational grounding to clear
* `rel_floor`. A relation whose relational axis has never been established does
* not pass by default it is reported unestablished and refused, because
* defaulting it to passing is exactly the exemption §0 forbids. Traversal is
* untouched: activation still conducts on the factual/associative side, so a
* relation can remain thinkable while ceasing to be assertable. */
typedef struct {
int may_assert;
int found; /* any relation at all on this claim */
int relational_established;
int still_held; /* DERIVED: node present and not tombstoned */
double factual; /* best decayed factual grounding */
double relational; /* the SAME edge's relational axis, not a max */
double polarity;
double cos_angle;
int agreement;
CogProvClass prov;
char best_edge[192];
char binding_value[128];
int n_edges;
} CogAssertion;
int cog_assert_two_axis(EngramPagedStore* s, const char* claim_id,
double floor, double rel_floor, int64_t now_ms,
CogAssertion* out);
#endif /* ENGRAM_COGNITION_H */
+2 -2
View File
@@ -222,7 +222,7 @@ static double eff_w(double weight, double hebb){
}
GeoDescriptor* engram_geometry_descriptor(
EngramPagedStore* store, VIndex* vindex,
EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids,
const char* const* seed_ids, size_t n_seeds,
const GeoParams* params,
@@ -1401,7 +1401,7 @@ static double geo_weighted_degree(EngramPagedStore* st, const char* id, double e
return deg;
}
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex,
int engram_geo_reify_store(EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids,
const GeoReifyParams* params){
if(!store) return -1;
+2 -2
View File
@@ -150,7 +150,7 @@ void engram_geo_mean_free(GeoMeanCache* c);
* Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error
* (no seeds resolvable, OOM). */
GeoDescriptor* engram_geometry_descriptor(
EngramPagedStore* store, VIndex* vindex,
EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids,
const char* const* seed_ids, size_t n_seeds,
const GeoParams* params,
@@ -375,7 +375,7 @@ void engram_geo_reify_default_params(GeoReifyParams* p);
* neighborhood (+ member edges), superseding any prior same-hub record with
* provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0.
* Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex,
int engram_geo_reify_store(EngramPagedStore* store, const VIndex* vindex,
char** vids, int n_vids,
const GeoReifyParams* params);
+378 -6
View File
@@ -44,6 +44,11 @@
#include <string.h>
#include <stdint.h>
#include <unistd.h>
#if defined(__APPLE__) || defined(__MACH__)
#include <sys/sysctl.h>
#include <mach/mach.h>
#include <mach/mach_host.h>
#endif
#include <fcntl.h>
#include <errno.h>
#include <time.h>
@@ -236,8 +241,16 @@ struct PgCache {
unsigned prefetch; /* read-ahead window (pages); 0 = off */
LayerPin* lp; size_t lp_n, lp_cap; /* hot-layer pin bookkeeping */
size_t dirty_count; /* # dirty frames, maintained incrementally (M5) */
/* stats (introspection only — never affect semantics) */
/* Interoception. These were "introspection only — never affect semantics",
* and that was the bug: the pool could not feel itself thrash, so it could
* not correct, and neither could anyone watching from outside. The sensed
* state IS the corrective mechanism (see pc_adapt_budget) the same way the
* engram's own boundary-beat/chronoception let it feel its own activity. */
uint64_t hits, misses, evictions, prefetch_reads;
/* sliding-window marks so pressure reflects NOW, not lifetime totals */
uint64_t adapt_last_acc, adapt_last_evic, adapt_last_hits;
uint64_t adapt_grows; /* budget corrections upward */
uint64_t adapt_shrinks; /* budget corrections downward (memory pressure) */
};
/* ── little-endian scalar codecs ──────────────────────────────────────────── */
@@ -334,6 +347,51 @@ static uint64_t dh_node_hash(const StoreNode* n){
return h;
}
/* dh_edge_hash — the edge counterpart of dh_node_hash.
*
* WHY THIS EXISTS (2026-08-15): the write barrier was node-only. Checkpointing
* pushes the WHOLE resident graph through store_put_node/store_put_edge (see
* engram_store_checkpoint), and nodes were cheaply skipped when unchanged
* a hash compare, no page I/O. Edges had no such check, so every edge was
* rewritten on every checkpoint, and each rewrite runs the idempotency probe
* max_page_lsn_for_id btree lookup page_read per stored copy.
*
* Edges outnumber nodes roughly 3:1 here (37,663 vs 13,430), so this turned
* routine checkpointing into a FULL-STORE WALK in id order random page access
* across the entire 2 GiB store, repeated, mostly to rediscover that nothing
* had changed. That walk is the failure mode: with a page cache smaller than
* the store it degenerates into thrashing and the engram never makes progress.
* Sizing the cache around that walk treats the symptom; the walk itself should
* not happen.
*
* The discriminator byte keeps the edge keyspace from ever colliding with a
* node of the same id in the shared dh map: distinct kinds cannot produce the
* same hash, so a stale skip is not reachable by collision. */
static uint64_t dh_edge_hash(const StoreEdge* e){
uint64_t h = 1469598103934665603ULL;
const uint8_t kind = 0xE0; /* edge discriminator */
dh_fold_bytes(&h, &kind, 1);
dh_fold_str(&h, e->id);
dh_fold_str(&h, e->from_id);
dh_fold_str(&h, e->to_id);
dh_fold_str(&h, e->relation);
dh_fold_str(&h, e->metadata);
uint8_t t8[8];
put_f64(t8, e->weight); dh_fold_bytes(&h, t8, 8);
put_f64(t8, e->hebb); dh_fold_bytes(&h, t8, 8);
put_f64(t8, e->confidence); dh_fold_bytes(&h, t8, 8);
uint8_t t4[4];
put_u32(t4, (uint32_t)e->inhibitory); dh_fold_bytes(&h, t4, 4);
put_u32(t4, e->layer_id); dh_fold_bytes(&h, t4, 4);
/* created_at/updated_at/last_fired are deliberately EXCLUDED: last_fired is
* touched by activation without changing what the edge IS, and including it
* would defeat the barrier on exactly the hot edges it most needs to skip.
* The fields that define the edge's durable content are all folded above. */
if (e->unknown && e->unknown_len) dh_fold_bytes(&h, e->unknown, e->unknown_len);
if (h == 0) h = 1; /* reserve 0 as "absent" in the map */
return h;
}
/* Open-addressing id(string)→durable-hash map. Keyed for O(1) bucketing on the
* id's FNV hash, compared by strcmp for correctness (full-id discipline, matching
* store_scan_*'s StrSet). Values are the 64-bit durable hash. */
@@ -1531,6 +1589,11 @@ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx){
if (cand.id && *cand.id && strset_add(&seen, cand.id)){
StoreEdge canon;
if (store_get_edge(s, cand.id, &canon) == 1){
/* seed the write-barrier map from on-disk truth so the FIRST
* post-boot checkpoint full-walk already skips unchanged edges
* (mirrors store_scan_nodes; without it the barrier is empty at
* boot and the first checkpoint re-probes every edge) */
if (s->barrier_on) dh_set(s->dh, canon.id, dh_edge_hash(&canon));
cb(&canon, ctx); count++; /* canonical latest-live */
store_edge_free(&canon);
}
@@ -1594,19 +1657,76 @@ int store_scan_edges(EngramPagedStore* s, StoreEdgeScanCb cb, void* ctx){
* matches disk, so a re-fault reproduces identical bytes.
* */
/* default frame budget: large enough that today's whole store stays resident
* (== Phase 1). Override with env ENGRAM_POOL_FRAMES (0 = unlimited). */
#ifndef ENGRAM_POOL_FRAMES_DEFAULT
#define ENGRAM_POOL_FRAMES_DEFAULT (1u<<20) /* ~1M frames × 16KiB = 16 GiB */
/* ── Frame budget ────────────────────────────────────────────────────────────
*
* A FIXED frame count cannot be correct. It has no relationship to either
* quantity that decides whether a cache works: the size of the working set, or
* the memory actually available on the host. It is the same number on a 16 GB
* laptop and a 256 GB server, and it stays put while the store grows.
*
* That is not hypothetical. On 2026-08-15 the deployment pinned
* ENGRAM_POOL_FRAMES=65536 (1 GiB) while neuron.egm grew to 2.1 GiB. The
* working set was twice the budget, so boot-time WAL replay which walks
* pages in an order uncorrelated with reuse evicted each page shortly before
* it was needed again. The engram spun at 100% CPU inside pc_evict_to_budget
* and never bound its port. Not slow: making no progress. Denning's thrashing,
* exactly, and no eviction policy can fix it when the working set does not
* fit, only more frames or admission control help.
*
* So the budget is DERIVED, from the host's physical memory, and it scales
* with the machine instead of pretending memory is a constant.
*
* ENGRAM_POOL_FRAMES explicit frame count; 0 = unlimited. Overrides all.
* Prefer leaving it unset a hand-set number is how
* this failure happened.
* ENGRAM_POOL_MEM_PCT percent of physical RAM to budget (default 60).
*
* Fallback when RAM cannot be read is 16 GiB worth of frames the old
* default, retained only as a floor for that case.
* */
#ifndef ENGRAM_POOL_FRAMES_FALLBACK
#define ENGRAM_POOL_FRAMES_FALLBACK (1u<<20) /* ~1M frames × 16KiB = 16 GiB */
#endif
static uint64_t pc_available_ram(void); /* fwd — defined with the controller */
/* Physical RAM in bytes, 0 when it cannot be determined. */
static uint64_t pc_physical_ram(void){
#if defined(__APPLE__) || defined(__MACH__)
uint64_t v = 0; size_t len = sizeof v;
int mib[2] = { CTL_HW, HW_MEMSIZE };
if (sysctl(mib, 2, &v, &len, NULL, 0) == 0) return v;
return 0;
#else
long pages = sysconf(_SC_PHYS_PAGES);
long psz = sysconf(_SC_PAGESIZE);
if (pages > 0 && psz > 0) return (uint64_t)pages * (uint64_t)psz;
return 0;
#endif
}
static size_t pc_default_cap(void){
unsigned pct = 60;
const char* p = getenv("ENGRAM_POOL_MEM_PCT");
if (p && *p){ unsigned long v = strtoul(p, NULL, 10); if (v > 0 && v <= 95) pct = (unsigned)v; }
uint64_t ram = pc_physical_ram();
if (!ram) return ENGRAM_POOL_FRAMES_FALLBACK;
uint64_t budget_bytes = (ram / 100u) * pct;
/* Never start above what the machine can actually spare right now. */
uint64_t avail = pc_available_ram();
if (avail > (1ull<<30) && budget_bytes > avail - (1ull<<30)) budget_bytes = avail - (1ull<<30);
uint64_t frames = budget_bytes / (uint64_t)STORE_PAGE_SIZE;
if (frames < 4096) frames = 4096; /* never absurdly small */
return (size_t)frames;
}
static PgCache* pc_new(void){
PgCache* c = (PgCache*)calloc(1, sizeof *c);
if (!c) return NULL;
c->nbuckets = 1024;
c->buckets = (PgEnt**)calloc(c->nbuckets, sizeof(PgEnt*));
if (!c->buckets){ free(c); return NULL; }
c->cap = ENGRAM_POOL_FRAMES_DEFAULT;
c->cap = pc_default_cap();
c->prefetch = 8;
const char* pf = getenv("ENGRAM_POOL_FRAMES");
if (pf && *pf){ char* end=NULL; unsigned long long v = strtoull(pf,&end,10); c->cap = (size_t)v; }
@@ -1676,6 +1796,241 @@ static void pc_remove(PgCache* c, PgEnt* e){
/* Reclaim clean unpinned frames from the LRU end until under budget, or until no
* evictable frame remains (a dirty/pinned-heavy pool may transiently exceed cap
* that is the no-steal guarantee, not a bug: the next checkpoint frees them). */
/* ── Adaptive budget: close the loop ─────────────────────────────────────────
*
* THE LESSON THIS ENCODES (2026-08-15). The engram spent hours down while four
* separate theories were tried bad binary, corrupt snapshot, WAL replay,
* feature flags because nothing in the system said what was happening. It
* looked identical to "busy loading": 100% CPU, flat RSS, no output. Meanwhile
* hits/misses/evictions were ALREADY being counted, right here, and surfaced
* nowhere. One eviction-rate number would have ended it in seconds.
*
* So the counters are not decoration. They are the control signal.
*
* A budget chosen once a literal like 65536, or 60% of RAM read at startup
* is a guess about the future. It cannot know the store grew, the working set
* shifted, or another process took the memory. The cache already MEASURES the
* only thing that matters (am I evicting pages I am about to want again), so it
* should act on that measurement instead of on a number someone typed.
*
* The controller: over a sliding window, if evictions are running at a rate
* comparable to accesses AND there is genuine reuse (hits are material), the
* working set exceeds the budget grow it. Growth is geometric, bounded by a
* live re-read of physical memory rather than a value cached at boot, so it
* tracks the machine instead of a snapshot of it. It never shrinks on its own:
* cap is a ceiling, not an allocation, and frames are only ever held because a
* real access put them there.
*
* Two things this deliberately does NOT do: it does not attempt a cleverer
* eviction policy (when the working set does not fit, no policy helps that is
* Denning, and it is why "tune the LRU" was never the fix), and it does not stay
* silent (pool_report exposes the same numbers outward, so a human or a metric
* pipeline sees the pressure the controller is reacting to). */
/* El's native telemetry, already in the runtime and already exporting to OTLP.
* Declared weak so engram_store.c still links standalone; when the runtime is
* present (every real build) the pool's interoception flows into the SAME
* pipeline as every other metric.
*
* ONE emission carrying the whole sensed state not a function per stat, and
* not a bespoke per-subsystem endpoint. Both of those are the degenerate case:
* they make observability something you hand-write per noun instead of a
* uniform mechanism every component already has. el_val_t is int64_t; strings
* ride as pointers cast through it (see el_runtime.h's value model). */
__attribute__((weak)) int64_t emit_log(int64_t level, int64_t msg, int64_t fields_json);
static void pc_report(const PgCache* c, const char* cause){
if (!emit_log) return; /* runtime not linked: no-op */
uint64_t acc = c->hits + c->misses;
char f[512];
snprintf(f, sizeof f,
"{\"component\":\"engram.pool\",\"cause\":\"%s\",\"hits\":%llu,\"misses\":%llu,"
"\"evictions\":%llu,\"prefetch_reads\":%llu,\"cap_frames\":%zu,\"resident\":%zu,"
"\"dirty\":%zu,\"grows\":%llu,\"hit_rate\":%.4f,\"evict_ratio\":%.4f,"
"\"cap_gib\":%.3f,\"resident_gib\":%.3f}",
cause,
(unsigned long long)c->hits, (unsigned long long)c->misses,
(unsigned long long)c->evictions, (unsigned long long)c->prefetch_reads,
c->cap, c->count, c->dirty_count, (unsigned long long)c->adapt_grows,
acc ? (double)c->hits / (double)acc : 0.0,
acc ? (double)c->evictions / (double)acc : 0.0,
(double)c->cap * (double)STORE_PAGE_SIZE / (1024.0*1024.0*1024.0),
(double)c->count * (double)STORE_PAGE_SIZE / (1024.0*1024.0*1024.0));
emit_log((int64_t)(uintptr_t)"warn", (int64_t)(uintptr_t)"engram.pool pressure",
(int64_t)(uintptr_t)f);
}
static uint64_t pc_ram_bytes_live(void){ return pc_physical_ram(); }
/* AVAILABLE memory right now — free + reclaimable, not total.
*
* Sizing a cache against TOTAL ram is what turns a cache into a memory leak:
* total does not shrink when other processes need memory, so a pool that only
* grows never notices it is starving the machine it runs on. Availability does.
* Returns 0 when undeterminable callers then refuse to grow, the safe way. */
static uint64_t pc_available_ram(void){
#if defined(__APPLE__) || defined(__MACH__)
/* SWAP AND COMPRESSOR FIRST. free+inactive+purgeable is a LIE under memory
* pressure: a machine deep in swap still reports gigabytes "available",
* because inactive pages are only reclaimable by evicting them to swap.
* Observed 2026-08-15: this returned 9.43 GiB available while vm.swapusage
* showed 51.58 of 53.25 GiB used (97% full) and the compressor occupied
* 23.7 GiB the host was thrashing to disk and the pool would have been
* cleared to grow into it. Growing a cache in that state is how a guard
* becomes the crash.
*
* So: if swap is nearly spent, report ZERO available. Callers refuse to
* grow on 0 and pc_relieve_pressure hands frames back. Only when the
* machine is genuinely not swapping do free+inactive+purgeable mean
* anything, and even then the compressor's footprint is subtracted because
* that RAM is already spoken for. */
/* RATE, NOT LEVEL. Swap *level* is a terrible signal: macOS grows swap files
* on demand and reclaims them lazily, so "47 of 48 GiB used" can mean the
* machine is dying OR that it recovered ten minutes ago and the file has not
* been trimmed yet. Measured both states on one host within minutes:
* 47.65/48.00 GiB used, 2047 swapouts/s -> genuinely thrashing
* 26.67/28.00 GiB used, 0 swapouts/s -> perfectly healthy, 15.6 GiB free
* A level check calls the second one an emergency and starves the pool for
* no reason. What distinguishes them is whether pages are moving NOW.
*
* So sample the swapout counter across calls and judge the delta. First call
* establishes the baseline and reports no pressure one sample cannot have
* a rate, and guessing from a single reading is the whole mistake. */
{
static uint64_t prev_swapouts = 0;
static time_t prev_t = 0;
static int primed = 0;
mach_port_t h0 = mach_host_self();
vm_statistics64_data_t v0; mach_msg_type_number_t c0 = HOST_VM_INFO64_COUNT;
if (host_statistics64(h0, HOST_VM_INFO64, (host_info64_t)&v0, &c0) == KERN_SUCCESS){
uint64_t now_out = (uint64_t)v0.swapouts;
time_t now_t = time(NULL);
if (!primed){ prev_swapouts = now_out; prev_t = now_t; primed = 1; }
else if (now_t > prev_t){
double per_s = (double)(now_out - prev_swapouts) / (double)(now_t - prev_t);
prev_swapouts = now_out; prev_t = now_t;
/* Sustained outward paging with nothing coming back is the
* signature of a host being pushed into swap. ~200 pages/s is
* ~3 MiB/s well above idle noise, well below the 2000+/s seen
* while actually thrashing. */
if (per_s > 200.0) return 0;
}
}
}
mach_port_t host = mach_host_self();
vm_size_t page = 0;
if (host_page_size(host, &page) != KERN_SUCCESS) return 0;
vm_statistics64_data_t vm; mach_msg_type_number_t cnt = HOST_VM_INFO64_COUNT;
if (host_statistics64(host, HOST_VM_INFO64, (host_info64_t)&vm, &cnt) != KERN_SUCCESS) return 0;
uint64_t avail = (uint64_t)vm.free_count + (uint64_t)vm.inactive_count
+ (uint64_t)vm.purgeable_count;
/* the compressor is holding real RAM that nobody can hand us */
uint64_t compressed = (uint64_t)vm.compressor_page_count;
if (compressed >= avail) return 0;
avail -= compressed;
return avail * (uint64_t)page;
#else
FILE* f = fopen("/proc/meminfo", "r");
if (!f) return 0;
char line[256]; unsigned long long kb = 0;
while (fgets(line, sizeof line, f))
if (sscanf(line, "MemAvailable: %llu kB", &kb) == 1) break;
fclose(f);
return (uint64_t)kb * 1024ull;
#endif
}
/* Shrink the budget when the machine is short on memory.
*
* A pool that can only grow is a leak with extra steps. This is the other half
* of the control loop: if free memory drops below a floor, hand frames back.
* The resident set follows on the next eviction pass, so the memory is actually
* returned rather than merely re-labelled. */
#ifndef ENGRAM_POOL_FREE_FLOOR_BYTES
#define ENGRAM_POOL_FREE_FLOOR_BYTES (2ull*1024ull*1024ull*1024ull) /* 2 GiB */
#endif
static int pc_relieve_pressure(PgCache* c){
uint64_t avail = pc_available_ram();
if (!avail) return 0;
uint64_t floor_b = ENGRAM_POOL_FREE_FLOOR_BYTES;
const char* fe = getenv("ENGRAM_POOL_FREE_FLOOR_MB");
if (fe && *fe){ unsigned long v = strtoul(fe, NULL, 10); if (v) floor_b = (uint64_t)v * 1024ull * 1024ull; }
if (avail >= floor_b) return 0; /* machine has room */
if (!c->cap || c->count == 0) return 0;
size_t was = c->cap;
size_t want = c->count - (c->count / 4); /* give back ~25% of what we hold */
if (want < 4096) want = 4096;
if (want >= c->cap) return 0;
c->cap = want;
c->adapt_shrinks++;
fprintf(stderr,
"[engram] memory pressure: %.2f GiB available (floor %.2f GiB) — shrinking pool "
"budget %zu -> %zu frames (%.2f -> %.2f GiB) and releasing frames.\n",
(double)avail/(1024.0*1024.0*1024.0), (double)floor_b/(1024.0*1024.0*1024.0),
was, c->cap,
(double)was * (double)STORE_PAGE_SIZE/(1024.0*1024.0*1024.0),
(double)c->cap* (double)STORE_PAGE_SIZE/(1024.0*1024.0*1024.0));
fflush(stderr);
return 1;
}
static void pc_adapt_budget(PgCache* c){
if (!c->cap) return; /* unlimited: nothing to adapt */
if (getenv("ENGRAM_POOL_FRAMES")) return; /* explicit operator override wins */
/* Sliding window so the signal reflects NOW, not lifetime totals. */
uint64_t acc = c->hits + c->misses;
if (acc - c->adapt_last_acc < 100000) return;
uint64_t d_acc = acc - c->adapt_last_acc;
uint64_t d_evic = c->evictions - c->adapt_last_evic;
uint64_t d_hits = c->hits - c->adapt_last_hits;
c->adapt_last_acc = acc; c->adapt_last_evic = c->evictions; c->adapt_last_hits = c->hits;
/* Pressure = evicting on a large fraction of accesses while still getting
* real reuse. Evictions alone are normal (a scan evicts and never returns);
* evictions WITH reuse means the working set genuinely does not fit. */
if (d_evic * 3 < d_acc) return; /* < 1/3 of accesses evict: healthy */
if (d_hits * 4 < d_acc) return; /* little reuse: a scan, not pressure */
/* Growth is bounded by what is AVAILABLE, never by total RAM. Sizing against
* total is how a cache starves its own host: total never shrinks when other
* processes need memory. Refuse to grow at all if availability is unknown or
* already under the floor a cache is never worth swapping the machine. */
uint64_t avail = pc_available_ram();
uint64_t floor_b = ENGRAM_POOL_FREE_FLOOR_BYTES;
const char* fe = getenv("ENGRAM_POOL_FREE_FLOOR_MB");
if (fe && *fe){ unsigned long v = strtoul(fe, NULL, 10); if (v) floor_b = (uint64_t)v * 1024ull * 1024ull; }
if (!avail || avail <= floor_b) return;
uint64_t ram = pc_ram_bytes_live();
if (!ram) return;
unsigned pct = 50; /* ceiling as a share of TOTAL, belt-and-braces */
const char* mp = getenv("ENGRAM_POOL_MAX_PCT");
if (mp && *mp){ unsigned long v = strtoul(mp, NULL, 10); if (v > 0 && v <= 95) pct = (unsigned)v; }
size_t ceiling = (size_t)(((ram / 100u) * pct) / (uint64_t)STORE_PAGE_SIZE);
/* and never grow into the free-memory floor */
uint64_t headroom = avail - floor_b;
size_t ceil_avail = (size_t)((c->count * (uint64_t)STORE_PAGE_SIZE + headroom)
/ (uint64_t)STORE_PAGE_SIZE);
if (ceil_avail < ceiling) ceiling = ceil_avail;
if (c->cap >= ceiling) return; /* already at the machine's limit */
size_t want = c->cap + (c->cap / 2) + 1; /* ×1.5, geometric */
if (want > ceiling) want = ceiling;
size_t was = c->cap;
c->cap = want;
c->adapt_grows++;
/* Emit the sensed state, not just the reaction. These are the numbers that
* would have diagnosed 2026-08-15 in seconds instead of hours. */
pc_report(c, "budget-grow");
fprintf(stderr,
"[engram] pool pressure: %llu evictions / %llu accesses (%llu hits) at %zu frames "
"(%.2f GiB) — working set exceeds budget; growing to %zu frames (%.2f GiB).\n",
(unsigned long long)d_evic, (unsigned long long)d_acc, (unsigned long long)d_hits,
was, (double)was * (double)STORE_PAGE_SIZE / (1024.0*1024.0*1024.0),
c->cap,(double)c->cap * (double)STORE_PAGE_SIZE / (1024.0*1024.0*1024.0));
fflush(stderr);
}
static void pc_evict_to_budget(PgCache* c){
if (!c->cap) return; /* unlimited */
while (c->count > c->cap){
@@ -1687,6 +2042,7 @@ static void pc_evict_to_budget(PgCache* c){
}
if (!freed) break; /* nothing evictable — allowed to exceed cap */
}
if (!pc_relieve_pressure(c)) pc_adapt_budget(c);
}
static PgEnt* pc_get(EngramPagedStore* s, uint64_t id){
@@ -2377,6 +2733,18 @@ int store_put_node(EngramPagedStore* s, const StoreNode* n){
int store_put_edge(EngramPagedStore* s, const StoreEdge* e){
if (!s || !e || !e->id || !e->from_id || !e->to_id) return -1;
STORE_GUARD(s);
/* Durable-hash write barrier — mirrors store_put_node. An unchanged edge
* costs one hash compare and zero page I/O; without this, checkpointing
* re-probed every edge against the paged store (max_page_lsn_for_id
* page_read), turning a routine checkpoint into a full-store walk. */
uint64_t dh_h = 0;
if (s->barrier_on){
dh_h = dh_edge_hash(e);
if (dh_get(s->dh, e->id) == dh_h){
s->stat_barrier_skips++;
return 0;
}
}
uint64_t L = ++s->next_lsn;
if (s->wal){
size_t blen; uint8_t* body = edge_serialize(e, &blen);
@@ -2386,6 +2754,10 @@ int store_put_edge(EngramPagedStore* s, const StoreEdge* e){
if (wr != 0) return -1;
}
int r = apply_edge_put(s, e, L);
if (r == 0 && s->barrier_on){
if (!dh_h) dh_h = dh_edge_hash(e);
dh_set(s->dh, e->id, dh_h); /* remember the now-persisted durable hash */
}
ckpt_maybe(s);
return r;
}
+68 -34
View File
@@ -74,11 +74,6 @@ struct VIndex {
int entry; /* entry-point element index, -1 if empty */
int max_level; /* current top layer */
/* scratch: version-stamped visited set (O(1) reset). */
uint32_t* visited;
uint32_t visit_epoch;
size_t visited_cap;
};
/* ── small helpers ────────────────────────────────────────────────────────── */
@@ -166,37 +161,63 @@ static Pair heap_pop(Heap* h, int is_max){
return top;
}
/* ── visited set ──────────────────────────────────────────────────────────── */
static int visited_ensure(VIndex* ix){
if (ix->visited_cap >= ix->cap && ix->visited) return 0;
size_t nc = ix->cap ? ix->cap : 16;
uint32_t* nv = (uint32_t*)realloc(ix->visited, nc*sizeof(uint32_t));
if (!nv) return -1;
if (nc > ix->visited_cap) memset(nv + ix->visited_cap, 0, (nc-ix->visited_cap)*sizeof(uint32_t));
ix->visited = nv; ix->visited_cap = nc;
/* ── visited set — owned by the CALL FRAME, never by the index ──────────────
* This buffer is per-TRAVERSAL scratch. It used to live in struct VIndex as an
* allocation optimisation, which made every traversal a write to shared state:
* two concurrent vindex_search calls stamped each other's epoch and then walked
* each other's marks, so even two pure READS corrupted the traversal (measured
* 2026-08-16: TSan data race at visited_reset, reached from vindex_search on one
* thread and vindex_insert on another; downstream SIGSEGV dereferencing a bogus
* element index).
*
* It is not an ownership problem and it does not want a lock or a capability
* it was simply misfiled. A pure function's scratch belongs to the call. Moving
* it here is what lets vindex_search take a `const VIndex*`, which is in turn
* what makes "search does not mutate the index" a COMPILE-TIME property instead
* of a review comment.
*
* Cost: one calloc/free of cap*4 bytes per traversal (~55 KB at the live store's
* 13,820 elements), against thousands of dim-768 dot products in the same call.
* Deliberately NOT __thread: http_worker is a thread per connection, so a
* thread-local buffer would retain ~55 KB per connection for the process life. */
typedef struct {
uint32_t* mark; /* per-element epoch stamp */
uint32_t epoch; /* current traversal's stamp; 0 == "no traversal yet" */
size_t cap;
} VVisit;
/* calloc leaves every stamp 0 and epoch 0; the first visit_reset moves to
* epoch 1, so no element reads as visited before it is marked. */
static int visit_init(VVisit* v, size_t cap){
size_t nc = cap ? cap : 16;
v->mark = (uint32_t*)calloc(nc, sizeof(uint32_t));
if (!v->mark) return -1;
v->cap = nc; v->epoch = 0;
return 0;
}
static inline void visited_reset(VIndex* ix){
if (++ix->visit_epoch == 0){ /* wrapped: clear all */
memset(ix->visited, 0, ix->visited_cap*sizeof(uint32_t));
ix->visit_epoch = 1;
static void visit_dispose(VVisit* v){ free(v->mark); v->mark = NULL; v->cap = 0; }
static inline void visit_reset(VVisit* v){
if (++v->epoch == 0){ /* wrapped: clear all */
memset(v->mark, 0, v->cap*sizeof(uint32_t));
v->epoch = 1;
}
}
static inline int is_visited(VIndex* ix, int e){ return ix->visited[e]==ix->visit_epoch; }
static inline void mark_visited(VIndex* ix, int e){ ix->visited[e]=ix->visit_epoch; }
static inline int is_visited(const VVisit* v, int e){ return v->mark[e]==v->epoch; }
static inline void mark_visited(VVisit* v, int e){ v->mark[e]=v->epoch; }
/* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */
/* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller
* owns res->a. `q` is a normalised query. */
static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
static int search_layer(const VIndex* ix, VVisit* vis, const float* q,
const int* eps, int neps,
int ef, int layer, Heap* res /*out, max-heap*/){
Heap cand = {0,0,0}; /* min-heap: nearest to expand */
res->a=NULL; res->n=0; res->cap=0;
visited_reset(ix);
visit_reset(vis);
for (int i=0;i<neps;i++){
int e = eps[i];
if (is_visited(ix,e)) continue;
mark_visited(ix,e);
if (is_visited(vis,e)) continue;
mark_visited(vis,e);
float d = vdist(ix, q, ix->elems[e].vec);
Pair p = { d, e };
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
@@ -212,8 +233,8 @@ static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
NeighList* nl = &ce->links[layer];
for (int i=0;i<nl->count;i++){
int e = nl->ids[i];
if (is_visited(ix,e)) continue;
mark_visited(ix,e);
if (is_visited(vis,e)) continue;
mark_visited(vis,e);
float d = vdist(ix, q, ix->elems[e].vec);
if (res->n < ef || d < res->a[0].d){
Pair p = { d, e };
@@ -232,7 +253,7 @@ static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
* Keep c only if it is nearer to q than to every already-chosen neighbour;
* backfill from the pruned set (nearest first) to reach M for connectivity.
* Writes chosen element indices into out[], returns the count. */
static int select_neighbors(VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
static int select_neighbors(const VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
(void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */
/* sort W ascending by (dist,elem) — deterministic. */
for (int i=1;i<nW;i++){ /* insertion sort (nW small) */
@@ -281,7 +302,7 @@ static int elems_reserve(VIndex* ix){
Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem));
if (!ne) return -1;
ix->elems = ne; ix->cap = nc;
return visited_ensure(ix);
return 0;
}
int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
@@ -307,13 +328,19 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
return 0;
}
/* This call frame owns its traversal scratch for the whole insert. ix->cap
* already covers `cur` (elems_reserve ran above), so every reachable element
* index is in range. */
VVisit vis;
if (visit_init(&vis, ix->cap)) return -1;
int ep = ix->entry;
int L = ix->max_level;
/* greedy descent through layers above `level` to refine the entry point. */
for (int lc = L; lc > level; lc--){
Heap r = {0,0,0};
int eps1[1] = { ep };
if (search_layer(ix, el->vec, eps1, 1, 1, lc, &r)){ return -1; }
if (search_layer(ix, &vis, el->vec, eps1, 1, 1, lc, &r)){ visit_dispose(&vis); return -1; }
if (r.n){ ep = r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} }
free(r.a);
@@ -329,7 +356,7 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
for (int lc = start; lc >= 0; lc--){
int Mmax = (lc==0) ? ix->M0 : ix->M;
Heap W = {0,0,0};
if (search_layer(ix, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
if (search_layer(ix, &vis, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
if (!chosen){ free(W.a); rc=-1; break; }
int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen);
@@ -357,13 +384,17 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
}
done:
free(eps_owned);
visit_dispose(&vis);
if (rc) return -1;
if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; }
return 0;
}
/* ── search ───────────────────────────────────────────────────────────────── */
int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
/* `ix` is const: search is pure with respect to the index. That is enforced by
* the compiler, not by convention it is the whole point of moving the visited
* set into the frame below. */
int vindex_search(const VIndex* ix, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out){
if (!ix || !query || k <= 0) return -1;
if (ix->entry < 0) return 0;
@@ -373,11 +404,15 @@ int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
float* q = vec_normalise_copy(query, ix->dim);
if (!q) return -1;
/* This call frame owns its traversal scratch. */
VVisit vis;
if (visit_init(&vis, ix->cap)){ free(q); return -1; }
int ep = ix->entry;
for (int lc = ix->max_level; lc > 0; lc--){
Heap r = {0,0,0};
int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, 1, lc, &r)){ free(q); return -1; }
if (search_layer(ix, &vis, q, eps, 1, 1, lc, &r)){ visit_dispose(&vis); free(q); return -1; }
if (r.n){ int b=r.a[0].e; float bd=r.a[0].d;
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;}
ep = b; }
@@ -385,7 +420,8 @@ int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
}
Heap res = {0,0,0};
int eps[1] = { ep };
if (search_layer(ix, q, eps, 1, ef_search, 0, &res)){ free(res.a); free(q); return -1; }
if (search_layer(ix, &vis, q, eps, 1, ef_search, 0, &res)){ visit_dispose(&vis); free(res.a); free(q); return -1; }
visit_dispose(&vis);
free(q);
/* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */
@@ -419,7 +455,6 @@ VIndex* vindex_create(int dim, int M, int ef_construction){
ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0);
ix->entry = -1;
ix->max_level = 0;
ix->visit_epoch = 0;
return ix;
}
@@ -432,7 +467,6 @@ void vindex_free(VIndex* ix){
free(e->vec);
}
free(ix->elems);
free(ix->visited);
free(ix);
}
+9 -2
View File
@@ -53,8 +53,15 @@ int vindex_insert(VIndex* idx, uint64_t node_id, const float* vec);
* first (ascending distance). Either out array may be NULL to skip it.
* ef_search search-time candidate width; larger == higher recall, slower.
* Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k.
* Returns the number of results written, or <0 on error. */
int vindex_search(VIndex* idx, const float* query, int k, int ef_search,
* Returns the number of results written, or <0 on error.
*
* `idx` is const BY CONTRACT AND BY TYPE: search does not mutate the index. The
* traversal's visited set is owned by the call frame, so N threads may search one
* index concurrently. Concurrent search against a vindex_insert on the same index
* is still unsafe insert rewires existing elements' neighbour lists and reallocs
* elems[] so the index's owner must not extend a published index under a live
* reader. See eg_vindex_view / eg_vindex_maintain in el_runtime.c. */
int vindex_search(const VIndex* idx, const float* query, int k, int ef_search,
uint64_t* node_id_out, float* dist_out);
/* Number of vectors currently indexed. */
+147 -5
View File
@@ -7,11 +7,29 @@
*
* Read-only: never opens a socket, never writes the store. Safe on an nsbx clone.
*
* Build: cc -O2 -std=c11 vindex_bench.c engram_vindex.c -lm -o vindex_bench
* Also runs the brute-force oracle a second (and third) way, through the
* batch-cosine Strategies behind eg_cosine_batch_strategy.h the ggml
* strategy and the hand-rolled-Metal strategy (Apple/Metal only; see
* eg_cosine_batch.h/eg_cosine_batch_strategy.h) and reports each one's
* latency + a correctness check against the CPU oracle side-by-side with the
* existing CPU-vs-HNSW numbers. This harness deliberately reaches past the
* single-selection Factory (eg_cosine_batch.c) to instantiate every
* compiled-in strategy directly, so it can compare all of them against the
* SAME dataset in one run that is the harness's whole job; a real call
* site (el_runtime.c) never does this, it only ever calls the plain
* eg_cosine_batch()/eg_cosine_batch_multi() adapter functions.
* EL_METAL_COSINE=0 forces CPU-only (skips every strategy comparison).
*
* Build (macOS, ggml + hand-rolled Metal): see build_vindex_bench.sh.
* Build (Linux / no Metal): omit every eg_cosine_batch_strategy_*.{c,m} file
* except eg_cosine_batch_strategy_cpu.c this file never references
* ggml/Metal directly except through the plain-C strategy header, guarded
* by the same EG_HAVE_STRATEGY_* build macros the Factory itself uses.
* Usage: vindex_bench store <neuron.egm> <dim> [nqueries] [k] [ef_csv]
* vindex_bench synth <N> [dim] [clusters] [nqueries] [k] [ef_csv]
*/
#include "engram_vindex.h"
#include "eg_cosine_batch_strategy.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
@@ -74,6 +92,106 @@ static double recall_at_k(const int* gt, const uint64_t* ann, int nann, int k){
return (double)hit / (double)k;
}
/* EL_METAL_COSINE: 0/off/false disables EVERY strategy comparison outright
* (falls back to brute_topk() only), matching el_runtime.c's own gate for
* the same env var (back-compat name kept from PR #114; it now gates all
* GPU-backed strategies, not just the hand-rolled Metal one). Unset or any
* other value = try every compiled-in strategy, report each that's
* available, skip (without failing the run) any that isn't. */
static bool g_strategy_env_checked = false;
static bool g_strategy_disabled_by_env = false;
static void eg_strategy_check_env_once(void){
if (g_strategy_env_checked) return;
g_strategy_env_checked = true;
const char* v = getenv("EL_METAL_COSINE");
if (v && (v[0]=='0' || v[0]=='n' || v[0]=='N' || v[0]=='f' || v[0]=='F'))
g_strategy_disabled_by_env = true;
}
/* Batched sibling of brute_topk, generalized over ANY EgCosineBatchStrategy:
* computes top-k for ALL nq queries in ONE strategy->batch_multi() call,
* uploading/preparing the node population exactly once instead of once per
* query. out_ids/out_d are nq*k, row-major (query i's results at
* out_ids+i*k / out_d+i*k). Returns false (nothing written) on any
* failure/unavailability; caller treats that as "skip this strategy in the
* report", never as a hard error. */
static bool batch_topk_strategy(const EgCosineBatchStrategy* strat,
const float* data, int n, int dim,
const float* queries, int nq,
int k, int* out_ids, float* out_d){
if (!strat || !strat->available()) return false;
const float** row_ptrs = malloc((size_t)n * sizeof(float*));
int32_t* dims = malloc((size_t)n * sizeof(int32_t));
double* scores = malloc((size_t)nq * (size_t)n * sizeof(double));
if (!row_ptrs || !dims || !scores) { free(row_ptrs); free(dims); free(scores); return false; }
for (int i = 0; i < n; i++) { row_ptrs[i] = data + (size_t)i * dim; dims[i] = dim; }
bool ok = strat->batch_multi(queries, dim, nq, row_ptrs, dims, n, scores);
free(row_ptrs); free(dims);
if (!ok) { free(scores); return false; }
for (int qi = 0; qi < nq; qi++) {
int* ids = out_ids + (size_t)qi * k;
float* ds = out_d + (size_t)qi * k;
const double* srow = scores + (size_t)qi * n;
for (int i = 0; i < k; i++) { ids[i] = -1; ds[i] = 3.0f; }
for (int i = 0; i < n; i++) {
float d = 1.0f - (float)srow[i]; /* same distance convention as brute_topk */
if (d >= ds[k-1]) continue;
int p = k - 1;
while (p > 0 && ds[p-1] > d) { ds[p] = ds[p-1]; ids[p] = ids[p-1]; p--; }
ds[p] = d; ids[p] = i;
}
}
free(scores);
return true;
}
/* Runs batch_topk_strategy for one named strategy over ALL nq queries, diffs
* against the CPU ground truth (gt/gd, both nq*k), and prints a report line
* in the same shape PR #114 established for BRUTE-METAL id-recall over
* every query plus the actual max/mean same-rank distance delta across
* every (query,rank) pair that was compared, never fabricated or assumed. */
static void report_strategy_vs_oracle(const char* label, const EgCosineBatchStrategy* strat,
const float* data, int n, int dim,
const float* qv, int nq, int k,
const int* gt, const float* gd, double brute_ms){
if (g_strategy_disabled_by_env) { printf("%-13s: disabled via EL_METAL_COSINE\n", label); return; }
if (!strat || !strat->available()) { printf("%-13s: not available on this build/host — skipped\n", label); return; }
int* gtm = malloc((size_t)nq*k*sizeof(int));
float* gdm = malloc((size_t)nq*k*sizeof(float));
double tm0 = now_s();
bool ok = batch_topk_strategy(strat, data, n, dim, qv, nq, k, gtm, gdm);
double strat_ms = (now_s()-tm0)*1000.0/nq;
if (ok) {
double rec_sum = 0; double max_ddiff = 0; double sum_ddiff = 0; int compared = 0;
for (int i=0;i<nq;i++) {
const int* ids_gt = gt+(size_t)i*k;
const float* d_gt = gd+(size_t)i*k;
const int* ids_m = gtm+(size_t)i*k;
const float* d_m = gdm+(size_t)i*k;
uint64_t idset[512]; int m = (k<512)?k:512;
for (int j=0;j<m;j++) idset[j] = (uint64_t)ids_m[j];
rec_sum += recall_at_k(ids_gt, idset, m, k);
for (int j=0;j<k;j++) {
if (ids_gt[j] == ids_m[j]) {
double diff = fabs((double)d_gt[j]-(double)d_m[j]);
if (diff>max_ddiff) max_ddiff=diff;
sum_ddiff += diff; compared++;
}
}
}
printf("%-13s: %8.3f ms/query (%.1fx vs CPU brute; id-recall %.4f vs CPU oracle over %d queries; same-rank |Δdist|: max %.2e, mean %.2e over %d compared)\n",
label, strat_ms, brute_ms/strat_ms, rec_sum/nq, nq, max_ddiff, compared?sum_ddiff/compared:0.0, compared);
} else {
printf("%-13s: batch call failed mid-run — skipped\n", label);
}
free(gtm); free(gdm);
}
/* Parse "64,128,256" into an int array; returns count. */
static int parse_csv(const char* s, int* out, int maxo){
int n=0; if(!s||!*s) return 0;
@@ -142,13 +260,37 @@ static void run_bench(const char* label, float* data, int n, int dim,
l2norm(dst, dim);
}
/* ground truth: brute-force top-k for every query (also the oracle latency). */
/* ground truth: brute-force top-k for every query (also the oracle latency).
* gd is nq*k (one real slot per query, not a shared scratch buffer) so the
* strategy comparisons below can diff against every query's actual
* distances, not just whichever query happened to run last. */
int* gt = malloc((size_t)nq*k*sizeof(int));
float* gd = malloc((size_t)k*sizeof(float));
float* gd = malloc((size_t)nq*k*sizeof(float));
double tb0 = now_s();
for (int i=0;i<nq;i++) brute_topk(data, n, dim, qv+(size_t)i*dim, k, gt+(size_t)i*k, gd);
for (int i=0;i<nq;i++) brute_topk(data, n, dim, qv+(size_t)i*dim, k, gt+(size_t)i*k, gd+(size_t)i*k);
double brute_ms = (now_s()-tb0)*1000.0/nq;
printf("BRUTE-FORCE : %8.3f ms/query (oracle; O(N*D))\n", brute_ms);
printf("BRUTE-FORCE : %8.3f ms/query (oracle; O(N*D), CPU)\n", brute_ms);
/* GPU-backed oracles: SAME nq queries, SAME top-k contract, via each
* compiled-in Strategy's batch_multi() (uploads/prepares the node
* population once, not once per query). Run only for strategies that
* are actually available (checked internally) never fabricated, never
* assumed. Verified against the CPU ground truth computed above:
* id-recall across ALL nq queries, plus the actual max/mean distance
* delta across every (query,rank) pair that was compared. */
eg_strategy_check_env_once();
#ifdef EG_HAVE_STRATEGY_GGML
report_strategy_vs_oracle("BRUTE-GGML", eg_cosine_batch_strategy_ggml(),
data, n, dim, qv, nq, k, gt, gd, brute_ms);
#else
printf("BRUTE-GGML : strategy not compiled into this build\n");
#endif
#ifdef EG_HAVE_STRATEGY_METAL_HAND
report_strategy_vs_oracle("BRUTE-METAL", eg_cosine_batch_strategy_metal_hand(),
data, n, dim, qv, nq, k, gt, gd, brute_ms);
#else
printf("BRUTE-METAL : strategy not compiled into this build\n");
#endif
/* HNSW at each ef. */
uint64_t* aid = malloc((size_t)k*sizeof(uint64_t));
+169 -5
View File
@@ -29,6 +29,8 @@ This section is the **single source of truth** for what works and what is planne
- Lexer: keywords, identifiers, integer/float/string/bool literals, operators below.
- Parser: `let`, `return`, `fn`, `type`, `enum`, `import`, `from … import`, `while`, `for`, `if/else if/else`, `match`, `@decorator`, array/map literals, all listed operators, function calls, field access, index access, unary `!`/`-`, postfix `?`.
- Codegen: function definitions, top-level `main()`, all expression forms above, control flow, decorator-as-AST-attachment.
- Boundary seam: decorator arguments and stacking; VBD role enforcement via `#error`; `engram_boundary_beat` auto-emit at `@manager`/`@accessor` entry; `@route` dispatch tables (Section 9).
- Program-level declarative blocks: `cgi`, `service`, and `program` — the last carrying process identity and configuration (Section 18).
- C runtime: I/O, string operations, integer math, lists, maps, filesystem, command-line args, basic `json_get` substring lookup.
### Planned (in flight)
@@ -37,7 +39,7 @@ This section is the **single source of truth** for what works and what is planne
- **Match codegen.** Currently parsed; codegen does not emit. Adding `({ ... })` statement-expression emission.
- **`?` propagation.** Currently no-op. Adding nil-propagation semantics.
- **`cgi` block parsing.** Currently lexed (`cgi` is a keyword) but not parsed as a statement. Adding `parse_cgi_block` and codegen of `el_cgi_init` at the head of `main()`.
- **VBD role enforcement.** `@manager`/`@engine`/`@accessor` are accepted as decorators but not enforced. Adding compile-time check that `dharma_emit`/`dharma_field` only appear inside `@manager` functions.
- **Boundary epilogues.** The decorator seam injects a prologue only. Adding prologue/epilogue wrapping, the prerequisite for durability-as-an-effect (Section 19.1).
- **`vessel` keyword.** Replaces `package` in manifests. Adding to lexer.
- **Real `engram_*` runtime.** Currently stub. Adding in-process graph store with spreading activation, Hebbian strengthening, and disk persistence — see Section 16.4.
- **Real `dharma_*` runtime.** Currently stub. Adding network transport, channel registry, identity resolution.
@@ -96,8 +98,10 @@ The following words are reserved and cannot be used as identifiers. Each row not
| `while` | yes | Loop |
| `import` / `from` / `as` | yes | Module import |
| `true` / `false` | yes | Bool literals |
| `cgi` | planned | Top-level CGI declaration block |
| `manager` / `engine` / `accessor` | as decorators | VBD role marker on `fn` (enforcement planned) |
| `cgi` | yes | Top-level CGI declaration block |
| `service` | yes | Top-level capability-bounded declaration block |
| `program` | yes | Top-level cross-cutting declaration block (Section 18) |
| `manager` / `engine` / `accessor` | as decorators | VBD role marker on `fn`; enforcement and boundary auto-emit are live (Section 9) |
| `vessel` | planned | Manifest declaration (replaces `package`) |
| `activate` / `where` | planned | Spreading-activation construct |
| `sealed` | planned | Capability scope block |
@@ -446,9 +450,21 @@ Parsed. The module name is recorded; the brace-list is consumed. Both forms prod
fn handle(channel: String, msg: String) -> Void { … }
```
The `@` token followed by an identifier attaches a decorator name to the next `FnDef`. Decorators with structural meaning today: none. Planned enforcement (Section 16.2): VBD roles `@manager`, `@engine`, `@accessor`.
The `@` token followed by an identifier attaches a decorator to the next `FnDef`.
Non-VBD decorators are accepted and ignored.
**Decorators take arguments and they stack.** `@route("/p", "GET") @manager fn f()` attaches both to `f` as a `decorators` list of `{name, args}` records, topmost-first. Arguments are string literals only.
**Decorators have structural meaning today.** This is El's function-level boundary seam — the mechanism by which a cross-cutting concern is handled *at the boundary* rather than by a convention repeated at every call site:
| Decorator | Structural effect |
|---|---|
| `@manager` | Permits calls to `dharma_emit` / `dharma_field`. Calling either from a non-`@manager` fn emits a `#error` into the generated C — a compile-time failure, not a lint. |
| `@manager`, `@accessor` | Codegen injects one call to `engram_boundary_beat(<fn name>)` at function entry. The decorated op self-reports (chrono tick, afferent counter, self-activity strengthen, dharma bus event) with **zero** hand-written instrumentation in its body. |
| `@route(path, method, …)` | Records a route into a generated dispatch table. |
Decorators with no registered meaning are accepted and ignored.
**Limits of the seam, as it stands.** The injection is a *prologue only* — there is no epilogue, no wrapping of the call, and no way for a decorator to run code after the body returns. The injected callee is a fixed builtin chosen by the compiler, not derived from the decorator name or its arguments. Section 19 depends on lifting exactly these two limits.
---
@@ -1088,4 +1104,152 @@ The next minor version closes the implementation gaps named in this document. Tr
---
## 18. The Program Block — cross-cutting concerns [implemented]
### 18.0 Why this exists
A cross-cutting concern is one that belongs to the *process*, not to any function in it: only one of me may run; this is what my configuration is; every mutation must be durable; every request must be authorized.
El's units of encapsulation are the function and the module. Neither can hold a concern like that. So each one had been expressed the only way it could be — as a **convention**: *call this at every site.* Conventions of that shape do not hold. They are not enforced by anything, they are invisible in review, and they fail silently at the one site somebody forgot.
Measured in this codebase before this section existed:
| Concern | State | What the convention was |
|---|---|---|
| process identity | **zero** guards anywhere — no pidfile, no lock, no already-running check, at any layer | "check nothing is already running first" |
| configuration | **20** distinct environment variables in one program, each with its default written inline at the read site | "remember the right default here" |
| durability | **62** `persist_*` / `engram_save` / `wal_*` / `checkpoint` call sites | "after you mutate, remember to persist" |
| request auth | **10** per-route `_auth` checks | "check the token in this handler too" |
These are not four problems. They are one absence, four times.
That the convention form fails is observed, not predicted. Process identity failed three times in a single day: twice, two engram processes ran simultaneously against the same data directory; twice, a stale binary held a port and answered probes while a fresh build was believed to be under test, because `pkill -f` had silently failed to match its argv — which nearly produced a false "the fix does not work" conclusion. Configuration failed structurally: `ENGRAM_DATA_DIR` was read at six sites, five of them dead bindings, and the sixth defaulted to `/tmp/engram` — contradicting the canonical resolver's `$HOME/.neuron/engram` and landing a pre-destructive safety backup on ephemeral storage.
The `program` block is where a concern of this shape is declared once and enforced by the compiler at the process boundary.
### 18.1 Syntax
```
program "engram" {
singleton: "engram"
env ENGRAM_BIND: String = ":8742"
env GUIDE_PORT: Int = "8771"
env ENGRAM_API_KEY: String required
}
```
At most one `program` block per program. It composes with `cgi` and `service` — those declare what a program *may do*; `program` declares what a program *is*.
Grammar:
```ebnf
program_block = "program" string "{" { program_field } "}" ;
program_field = singleton_field | env_field ;
singleton_field = "singleton" ":" string [ "," ] ;
env_field = "env" ident ":" type
[ "=" string ] [ "required" ] [ "," ] ;
```
`singleton` and `env` are **not** reserved words. They are read as identifier token values by the block's own parse loop, so they remain usable as ordinary identifiers everywhere else. `program` is the only keyword this section adds.
### 18.2 Process identity — `singleton`
`singleton: "id"` compiles to an `el_singleton_acquire("id")` call injected as the **first statement of `main()`**, before any user statement runs.
The runtime takes an exclusive non-blocking `flock` on `<dir>/el-singleton-<id>.lock`, where `<dir>` is `$EL_SINGLETON_DIR`, else `$TMPDIR`, else `/tmp`. On success it writes its pid and holds the descriptor open for the life of the process. On contention it **refuses to start**: it reports the holder's pid, names the lock file, and exits 1.
Two properties are deliberate:
- **It is a lock, not a pidfile.** The kernel releases an `flock` when the owning process dies — including on `SIGKILL` and on crash. There is therefore no stale-lock state, and so no "delete the lock file to get unstuck" recovery ritual. Such a ritual would itself be a convention, which is the thing this section exists to remove.
- **It reports the holder's pid.** "Already running" is not actionable. A pid is. This is the direct answer to the observed failure where a stale process survived a `pkill` and went on answering probes.
Refusal is loud and total. It is not a warning, and the program does not continue degraded. This matters more than it looks: today a second engram whose `bind()` fails merely *returns* from `http_serve` — after it has already replayed the WAL and written boot-time backup files — and then exits **0**, indistinguishable from a clean run. `singleton` refuses before the first side effect.
### 18.3 Configuration — `env`
Each `env` entry declares one configuration variable: its name, its type (`Int` or `String`), and either a default or `required`.
Resolution happens once, at startup, in declaration order: **the environment wins; the declaration supplies the fallback.** Then `el_config_validate` checks the whole schema and reports *every* problem at once before exiting — a startup that fails one variable at a time costs one restart per variable.
Values are read with `config("NAME")`, which returns a `String`.
The enforcement that makes the declaration real: **once a program block exists, `config("X")` for an undeclared `X` is a fatal error.** Without that, the schema would be advisory, and an advisory schema is just another convention. Programs with no `program` block are unaffected — `config()` falls back to a plain environment read, so migration is incremental and per-program.
The point is not that configuration is now centralized. It is that **a default is no longer a decision made at a read site.** A read site cannot disagree with another read site about what a variable means, because a read site no longer says.
### 18.4 What is deliberately not declared here
Some values look like configuration and are not. `ENGRAM_DATA_DIR` already has a single owner — `engram_resolve_data_dir()`, which resolves it, creates the directory, and fails loud rather than silently persisting to an ephemeral path. Declaring it in the `program` block as well would give it two owners that can disagree, recreating the precise defect this section removes.
The rule: **a variable belongs in the program block when the block would be its only owner.** If a resolver already owns it, leave it there.
`HOME` is likewise not configuration. It is an environment fact, and stays a raw `env()` read.
---
## 19. Boundary Effects — durability and request authorization [design only, not implemented]
Sections 19.1 and 19.2 specify the two remaining concerns from the table in 18.0. Both are **designed and deliberately unimplemented.** The reason is stated in 19.3 and it is not difficulty.
### 19.1 Durability as an epilogue effect
**The defect.** 62 call sites carry the convention *"after you mutate, remember to persist."* This is structurally the same defect as the index bug being fixed elsewhere in this tree — *"after you append, remember to index"* — which failed at **9 of 9** sites. A convention that failed at 100% of its sites is the strongest available evidence about what this class of convention is worth.
**Why the existing seam cannot express it.** §9's injection is a prologue. Durability is inherently an *epilogue*: persist after the mutation succeeds, and not at all if it threw. The seam has no epilogue.
**Design.** Extend the decorator seam from prologue-only to prologue/epilogue, then declare durability as an effect on the mutating function:
```
@durable("engram")
fn engram_write_node(id: String, body: String) -> Bool { … }
```
Codegen wraps rather than prefixes:
```c
el_val_t engram_write_node(el_val_t id, el_val_t body) {
el_effect_enter(EL_STR("durable"), EL_STR("engram"));
el_val_t __r = /* original body */;
el_effect_exit(EL_STR("durable"), EL_STR("engram"), __r);
return __r;
}
```
`el_effect_exit` is where the persist happens, and it is the only place it happens. Two properties follow that the 62 hand-written sites cannot have:
- **Coalescing.** The epilogue is a single choke point, so N mutations inside one request can produce one fsync instead of N. The hand-written form cannot coalesce, because no site knows about the others.
- **Failure is not silent.** A persist that fails inside `el_effect_exit` can force the mutation's return value to failure. A forgotten `persist_*` call cannot fail — it simply does not happen, which is exactly why the defect is invisible.
**Enforcement, and this is the part that actually fixes it.** Mirroring §9's `#error` for `dharma_emit`: a function that calls a mutating primitive without carrying `@durable` is a **compile error**. Otherwise this is a 63rd thing to remember rather than a replacement for 62.
### 19.2 Request authorization as a route effect
**The defect.** 10 per-route `_auth` checks. The HTTP layer has no concept of authorization, so a new route is unauthenticated by default and silently so — the failure mode is a route that forgot, and nothing anywhere reports it.
**Design.** Authorization becomes an argument to the `@route` decorator, which already takes arguments and already builds a dispatch table:
```
@route("/api/write", "POST", auth: "required")
fn route_write(body: String) -> String { … }
```
The generated dispatcher performs the check **before** dispatch, so an unauthorized request never reaches the handler and the handler contains no auth code at all.
The default must be `required`. A route that says nothing gets authorization; opening one up takes an explicit `auth: "public"`. Defaulting to public preserves the current failure mode exactly — forgetting stays silent — and a default that preserves the defect is not a fix.
Route inventory falls out for free: the dispatch table already exists, so the compiler can emit the full route/auth matrix and make "which routes are public" a fact that is read rather than audited.
### 19.3 Why these are not implemented
Not difficulty — **collision**. Both land squarely in regions two other agents hold right now:
- **Durability** requires changing the mutation and persist paths in `lang/runtime/el_runtime.c` and `engram/src/server.el` — the same files and the same read/write paths being restructured by concurrent work on VIndex read-path mutation and memory ownership, and on geometry-as-an-el-value and `transduce`.
- **Request auth** requires changing route dispatch in `engram/src/server.el`, which the geometry/`transduce` work is actively reshaping.
Implementing either now would mean editing files under concurrent modification and resolving conflicts in exactly the paths whose correctness is currently under repair. The designs are recorded here so the work is not lost, and so that whoever lands them does so against a settled tree.
The prerequisite for 19.1 is the same in both cases: **lift the §9 seam from prologue-only to prologue/epilogue.** That change is independent of both collisions and can land first.
---
End of specification.
+180
View File
@@ -0,0 +1,180 @@
# El Runtime — Ownership and Capability ABI
**Status:** §0–§2 verified. §3 re-derived and **built** for the vector index (2026-08-16); not yet applied to the resident RAM graph.
**Date:** 2026-08-16
**Scope:** `lang/runtime/` — every El program (soul, engram, cgi-studio vessels) inherits this by rebuild. Nothing in this document is a change to any El *program*.
**Note on §1's line numbers:** they were read against a checkout that has since shifted by ~135 lines. Verified positions as of `a67452f` are in §2a.
---
## 0. The residual
> **Builtins own memory and reach process state directly.**
That is the residual — the generator. Everything below labelled a "residue" is a deposit left by it. The distinction matters because we have spent significant effort removing deposits, and deposits regenerate.
A residue is fixed. A residual is eliminated. Fixing residues while the residual stands produces exactly the pattern observed on 2026-08-15/16: a run of individually-correct patches, each verified, followed by a new defect of the same shape in a different file.
---
## 1. The residues, measured
Each of these is a distinct merged or proposed fix. Each addresses one deposit. None addresses the residual.
| residue | location | fix that was applied or proposed |
|---|---|---|
| `state_get` leaked its return value per call — 15 MB over 200k calls | builtin | el #140 (merged) |
| VIndex freed under a concurrent reader | `el_runtime.c:9424` | `fb32d15` guard (merged 08:46:43) |
| `_eg_vindex_seen` realloc'd on a read path | `el_runtime.c:9412` | same guard |
| `vindex_insert` on a read path | `el_runtime.c:9434`, `9450` | same guard |
| shared `visited` / epoch scratch stomped by concurrent searches | `engram_vindex.c:7981`, `169186`, `195` | proposed: move to per-search frame |
| nine append sites, none indexing → lazily-embedded nodes invisible | `el_runtime.c:7806, 7988, 8148, 8224, 11526, 11731, 12050, 15295, 15312` | "embed-gap #20", patched by making the *read* path catch up (`9439` comment) |
**Measured:** all file/line references above, read 2026-08-16. Crash frames `engram_activate → eg_vindex_sync → vindex_insert → _realloc → _xzm_xzone_malloc_freelist_outlined` are accounted for by rows 24.
**Inferred, not yet verified:** that the nine append sites do not share a single commit point. This needs one pass before Change C is sized.
---
## 2. Why these are one defect
`eg_vindex_sync` (`el_runtime.c:9419`) has exactly three callers, and **all three are reads**:
- `engram_activate``9802`
- `eg_knn_for_node``13075` (its own header comment states *"No writes."*)
- `engram_geo_reify_run_json``13285`
It mutates five process-global statics (`94009404`): `_eg_vindex`, `_eg_vindex_dim`, `_eg_vindex_built_nc`, `_eg_vindex_seen`, `_eg_vindex_seen_cap`.
Reads mutate because index maintenance was never given an owner on the write side. It got bolted onto reads, because a builtin *could* reach the globals — nothing prevented it. Likewise `state_get` leaked because a builtin *owned* the value it returned; nothing prevented that either.
The store is architecturally append-only and superseding. A read path that mutates contradicts that directly. The contradiction is expressible only because the ABI permits it.
---
## 2a. Verified positions and the fact §1 missed
Read directly at `a67452f`, 2026-08-16. §1's line numbers predate a ~135-line shift; these are current.
| thing | §1 said | actually |
|---|---|---|
| five process-global statics | 94009404 | **95359539** |
| `eg_vindex_seen_ensure` realloc | 9412 | **9547** |
| `eg_vindex_sync` | 9419 | **9554** |
| `vindex_free` on a read path | 9424 | **9559** |
| `vindex_insert` on a read path | 9434 / 9450 | **9569** (build) / **9585** (incremental) |
| caller: `engram_activate_inner` | 9802 | **9939** |
| caller: `eg_knn_for_node` | 13075 | **13212** |
| caller: `engram_geo_reify_run_json` | 13285 | **13422** |
| `fb32d15` guard | — | lock **1602**, depth **1631**, `eg_guard_enter` **1636**, `http_worker` acquire **1687**, `engram_activate` wrapper **14097** |
| VIndex scratch fields | 7981 | **7981** ✓ |
| `search_layer` race site | 195 | **195** ✓ |
**The structural fact §1 and §3 both missed:** *the index does not inherit the store's append-only property.* `vindex_insert` rewires the `NeighList` links of already-existing elements and reallocs `elems[]` — so extending the index mutates the whole structure, not just its tail. This is why "make reads pure" is necessary but **not sufficient**, and why §3 needed a publication boundary rather than only a capability split. It is reproduced as a standing test (`unsynchronized` half, §5).
---
## 3. The change
*(Re-derived 2026-08-16. The previous §3 — a runtime context struct carrying read/write **capability pointers** to every builtin — was written in mutable-store, C-ownership terms. It asked "who is permitted to mutate the shared thing?", which presupposes a shared mutable thing. The engram is immutable and recall is projection; what does not mutate needs no ownership discipline. So the question is not answered, it is dissolved. The implemented change is below.)*
### 3.1 Three moves, in decreasing order of how much they dissolve
**(1) Misfiled scratch is not shared state.** `visited` / `visit_epoch` were never conceptually owned by the index — they are one traversal's local, hoisted into `struct VIndex` as an allocation optimisation. Nothing about them is derived geometry. They want neither a lock nor a capability nor a checkout pool: a pure function's scratch belongs to its call frame, and the fix is to put it back there. This is not "the capability model applied by hand to one global"; it is the deletion of a false ownership claim.
**(2) `const` is the capability, and immutability hands it over for free.** Once the scratch leaves the struct, `search_layer` reads the index and nothing else — so `vindex_search` can take a `const VIndex*`. That is *precisely* the teeth old-§3 wanted from capability pointers: a read path physically cannot call `vindex_insert`, and it is a **compile error**, not a review comment. It costs one qualifier rather than a new ABI swept across hundreds of builtins. The compiler enforces it on every future caller for the same reason.
> The capability type was already in the language. It is spelled `const`.
**(3) What remains is a publication problem, not an ownership problem.** With scratch in the frame and reads const, one hazard survives, and it is real: **HNSW insert is not an append.** `vindex_insert` rewires the `NeighList` links of *already-existing* elements and reallocs `elems[]`. The store's append-only property does **not** transfer to the index derived from it. So a reader projecting against the index while its owner extends it is unsafe no matter how pure search is.
Immutability answers this too, and the answer is publication:
- **`eg_vindex_maintain`** — the sole mutator. Takes the boundary exclusively; never runs beside a reader.
- **`eg_vindex_view`** — returns a `const VIndex*` with the boundary held for read. N readers project concurrently; none can mutate.
A read path may **demand that a current snapshot exist** — that is a request to the owner, not a mutation by the reader. What it may not do is mutate the geometry it is projecting against. `view` / `maintain` is exactly that split, and it is why this replaces `eg_vindex_sync` rather than wrapping it.
**Write-side owner.** Index membership is owned by the event *"an embedding became present on this ordinal"* — not by node append, since a node without an embedding cannot be in a vector index at all. `eg_vindex_note_embedded` hooks the embedding-assignment sites: one O(log n) insert, no O(node_count) presence scan. This also retires the "STALENESS (honest tradeoff)" note in the old `eg_vindex_sync`, where a lazily-embedded *older* node stayed invisible to `route_nearest` / autoconnect until the next full rebuild.
### 3.2 What this does not claim
The **resident RAM graph** (`g->nodes` / `g->edges`) is a *separate* residue of the same residual and is untouched by this change. It is realloc'd in place (`el_runtime.c:7618`, `7629`), so an awareness-thread reader holding `EngramNode* n = &g->nodes[i]` across a concurrent append holds a dangling pointer — and `engram_activate_inner`'s embed-backfill writes `n->emb` through exactly such a pointer. It wants the same publication treatment the index just received. Until that lands, the `fb32d15` guard stays (see §5).
---
## 4. Why this is not a large change
The old §4 argued that El owning its compiler makes a capability-ABI sweep mechanical, since `elc` generates every builtin call site. That argument was load-bearing only for the ABI, and the ABI is gone.
The constraint now travels with the **type of the thing**, not the shape of every call site — so no sweep is needed at all. Measured extent of the implemented change: two qualifiers (`const VIndex*` on `vindex_search`, propagated to `engram_geometry_descriptor` and `engram_geo_reify_store`), one struct field group relocated to a call frame, one rwlock, and three read call sites converted from `eg_vindex_sync` to `view`/`release`.
The payoff of owning the language is unchanged and is now *cheaper*: introduced once, enforced by the compiler on every future builtin, cannot subsequently be forgotten. Contrast the current state, where the same discipline was maintained by hand across hundreds of builtins and demonstrably failed at least six times.
---
## 5. What this deletes
**Deleted (done, 2026-08-16):**
- `eg_vindex_sync` — the function itself. Not renamed: split into `eg_vindex_maintain` (mutating, exclusive, sole owner) and `eg_vindex_view` (const, shared). A name that meant "read paths repair the index" had to stop existing.
- `VIndex::visited` / `visit_epoch` / `visited_cap` — the struct fields, `visited_ensure`, its call from `elems_reserve`, `ix->visit_epoch = 0` in `vindex_create`, and `free(ix->visited)` in `vindex_free`.
- The **proposed** per-search scratch *struct on the index* (a checkout pool / `VisitedListPool`) — never built. The buffer is a plain frame local; a pool is machinery for an ownership question that no longer exists.
- The **proposed** reader-view / owner-handle split for VIndex specifically — superseded. `const` already is the reader view.
- `EXPECT_RACE` in `run_vindex_concurrency_tests.sh` — a knob that let a known defect ride as "expected". Replaced by four halves with real verdicts.
**NOT deleted — the design doc was wrong about this one:**
- `fb32d15` (`eg_guard_enter` / `engram_req_lock` / `_eg_req_depth`). §5 originally called for its removal as "a lock protecting a mutation that ceases to exist." **Measured, it guards two things, and only one of them ceases to exist.** Its own comment names both: the RAM graph *and* `_eg_vindex`. The vindex justification is retired; the RAM-graph justification is independently load-bearing (§3.2), and removing the guard reintroduces the measured 11171→9579 edge-loss defect from 2026-08-14. Its comment has been narrowed to state the RAM graph only. **Precondition for deleting it:** the resident graph gets the same publication boundary the index just got.
- el #140's hand-patch. Left in place — the leak stops being *expressible* only under the abandoned capability-ABI §3, which is not what was built.
**Ordering consequence (revised):** the original ordering claim — "the residual lands first, the residues evaporate rather than get fixed" — did not survive contact. The residual here is not a single ABI that dissolves everything at once; it is a *property* (derived state is published, never edited) applied per structure. The index now has it. The RAM graph does not yet. Residues evaporate **per structure, in the order the property is applied**, and a residue whose structure has not been converted must be left standing, not deleted on the strength of the plan.
---
## 6. Sequencing
1. **Read** how builtins are declared and dispatched, to confirm the call sites are compiler-generated in one place. *(This determines whether §4 holds. If dispatch is scattered, re-size before proceeding.)*
2. Introduce the context type and capability types.
3. Codegen emits the context at every builtin call site.
4. Mechanical sweep of builtin signatures.
5. Move index maintenance behind the write capability; the three read callers take the read capability.
6. Delete the residue-fixes listed in §5.
7. **One** build of soul from el dev — which resolves the `state_get` leak and the crash together, rather than deploying a leak fix that reintroduces the crash.
---
## 7. Open questions
**Answered 2026-08-16:**
- ~~Do the nine append sites share a commit point?~~ **Moot.** The question was mis-aimed: node append is not the event that owns index membership, because a node without an embedding cannot be in a vector index. The five *embedding-assignment* sites are the real owner points (`el_runtime.c:7091, 9839, 13362, 15002`, plus snapshot-restore at `7951`), and three of them carry the ordinal directly — which is all `eg_vindex_note_embedded` needs. The other two run before the node is resident, where the cold build picks it up.
- ~~Does anything outside `lang/runtime/` construct a second `VIndex`?~~ **No.** Swept: the only constructors outside the runtime are `engram/test/*` and `lang/runtime/vindex_bench.c`, all single-threaded and index-private. Inside the runtime, `engram_self_reify_beat_json` builds a **private** index deliberately and never touches the shared boundary — that was already correct and is unchanged.
- ~~Does the HTTP worker pool contend on the same globals?~~ **Yes, and it was never the whole story.** Workers serialize against each other on `engram_req_lock`, but the awareness main thread does not take it at all — that is the gap `fb32d15` closed. Now verified independent of that guard: the index boundary is its own rwlock, so worker/awareness contention on `_eg_vindex` is handled whether or not the request lock is held.
**Still open:**
- The resident RAM graph wants the same publication boundary (§3.2). Until it has one, `fb32d15` cannot be deleted.
- `eg_vindex_view` holds the boundary for read across `engram_geo_reify_store`, which is a long pass. Correct, but it stalls the owner for that duration. If reify latency becomes a problem the answer is a refcounted snapshot, not a shorter lock.
---
## 7a. Evidence (measured 2026-08-16, `engram/test/run_vindex_concurrency_tests.sh`)
| half | before | after |
|---|---|---|
| `single` — 3000 vectors, 1 thread, ASan+UBSan | clean | clean |
| `readers` — 4 readers, no writer, TSan | **race** at `engram_vindex.c:195` (`visited_reset``vindex_search`) | **clean** |
| `unsynchronized` — writer+reader, bare index, TSan | race | **race, expected and permanent** — now the proof the boundary must exist |
| `published` — owner + 4 readers through the boundary, TSan | *(did not exist)* | **clean**, all 3000 inserts landed |
No recall regression: `recall@10 = 0.9365` at `ef_search=128` (gate ≥ 0.90); the determinism test still yields byte-identical results across two independent builds.
Builds locally: all seven engram runtime translation units compile `-Wall -Wextra` clean, and the full engram binary links (`engram/dist/engram.c` + runtime, arm64). The one pre-existing `-Wcomment` warning in `el_runtime.c` is present at `a67452f` too.
---
## 8. What this document is not
It is not an argument for a memory model in general, a garbage collector, process isolation between soul and engram, or a client/server split of the store. Each of those was considered and each addresses mutation that this change removes. They are answers to a question that stops being asked.
+184
View File
@@ -0,0 +1,184 @@
# Swarm + CCR + Work-Tracking — Neuron's bounded parallel execution, in native El
Bounded parallel agent execution on El's **native** concurrency — no external
orchestrator. Grounded directly in two of Will's frameworks:
- **Swarm Architecture** (*Bounded Parallel Agent Execution*, Mar 2026)
- **Compiled Context Runtime / CCR** (*Process-Driven Agent Execution with
Unbounded Local Memory*, Mar 2026)
A swarm is a **coordinator** (the main thread) that mints a correlation identity,
compiles a **bounded per-worker context (CCR)**, dispatches workers as **native
pthreads** (`thread.el` `spawn`/`join`), tracks every unit of work durably, and
**converges** results before returning control to the parent step.
```
Parent step
└─ swarm_run(blueprint, knowledge_refs, inputs, config)
fan-out ──▶ worker_1 (CCR ctx_1) ─┐ native
worker_2 (CCR ctx_2) ─┤ pthreads,
worker_k (CCR ctx_k) ─┘ bounded by `concurrency`
converge ─▶ collect | merge | vote | reduce ──▶ merged result
```
## Why it runs on El natively
El is natively agentic. This capability composes El's shipped primitives — it
adds no bespoke runtime:
| Primitive | Source | Role in the swarm |
|-----------|--------|-------------------|
| `spawn(fn,arg)` / `join(tid)` | `runtime/thread.el``__thread_create` (pthread + dlsym) | fan-out / rejoin |
| `parallel_map`, `with_mutex` | `runtime/thread.el` | reference concurrency patterns |
| Go-style channels | `runtime/channel.el``__channel_*` | available for vertical event streams |
| `engram_*`, `http_*`, `fs_*`, `json_*` | `el_runtime.c` builtins | retrieval, tracking, I/O |
Every El fn compiles to a global C symbol, so any top-level `(String)->String`
fn is directly threadable — the worker entry is exactly such a fn.
## Modules
| File | Framework grounding | What it does |
|------|--------------------|--------------|
| `worktrack.el` | Swarm §6 (correlation IDs, audit) | Durable, single-writer **JSONL journal** keyed by correlation ID; reconstructable status report; opt-in engram mirror (`SWARM_MIRROR=1`). |
| `containment.el` | Swarm §3 + the single-writer invariant | Scope tokens w/ capabilities; **Rule 1** (no join), **Rule 2** (no open), **Rule 3** (no lateral edge), **Rule 4** (engram-write is @manager-only, by capability) enforced as checks. |
| `ccr.el` | CCR §5 + Swarm §9.3 | Per-worker **Compiled Context Routing**: retrieve → scope → compact into a **bounded, minimal** package. The compiled-context boundary *is* the security boundary. |
| `primitives.el` | CCR §2 (Five Primitives) | `attend / think / intend / act / learn` seam the swarm composes over. Engram-backed; explicit binding point for the API-surface reshape. |
| `swarm.el` | Swarm §2, §4, §5 | The coordinator: fan-out/converge on native threads, bounded concurrency, four convergence strategies, integer failure threshold, full tracking. |
## Invariant: only the orchestrator mutates global engram state
**Only the orchestrator (@manager) writes to the engram / mutates global state.
Workers are read-only against the full engram and may write only their own local
geometry (their returned result + the journal). A worker is STRUCTURALLY UNABLE
to mutate global engram state.**
This is **Rule 4** — an **authority gate, not a health gate**. Scope tokens carry
a capability set: the orchestrator's token holds `engram:write` + `dharma:emit`
(@manager-only, the VBD rule that only the manager mutates global state); a
worker's token holds **only** `engram:read`. Every engram mutation
(`op_write`/`op_relate`/`op_supersede``POST /api/nodes`, `/api/edges`,
`DELETE`) flows through `swarm_engram_write`, which checks the caller's capability
via the **same scope-token mechanism as the live Rule-2 denial** and rejects any
worker **before any HTTP is issued**. Capability is fixed at mint time and cannot
be acquired at runtime — so the guarantee holds regardless of engram health
(distinct from the `SWARM_WRITE_HEALTHY` *health* gate).
The **curated merge is the only write path**: workers return geometry; the
orchestrator, and only the orchestrator, commits the approved/verified geometry
back (`commit=1`). Workers keep full-engram **read** access (`op_think`/`op_read`).
Proven in `harness_real_cognition.el` (§G): a worker `swarm_engram_write` is
DENIED by capability with no node created and the violation journalled; the
orchestrator passes the gate as the sole authorized writer.
## Containment → distribution
The three containment rules make workers **location-independent** (Swarm §9): a
worker reads only its compiled context, shares no state with siblings, and its
only outward edge is the returned result. The same coordinator can run workers
as local threads today or dispatch them across machines later — the mechanism is
identical; only the topology changes. Enforced here:
- **Rule 2**`swarm_run` rejects any swarm opened under a worker token.
- **Rules 1 + 3** — each worker gets a *closed* worker token; the coordinator is
the only journal writer, so workers share no mutable state.
## Usage
```el
// one process step fans out; results converge before the next step
let inputs: String = "[\"billing\",\"payments\",\"ledger\"]"
let refs: String = "[\"Volatility-Based Decomposition\"]" // CCR knowledge refs
let cfg: String = "{\"concurrency\":\"4\",\"strategy\":\"collect\",\"min_success_ratio\":\"1.0\"}"
let result: String = swarm_run("analyze_item", refs, inputs, cfg)
// result: { corr_id, status, merged, report }
```
Build any program that uses the swarm:
```bash
lang/swarm/build.sh myprog.el ./myprog # concat + elc + cc (el_runtime.c)
```
Config keys: `concurrency` (max workers at once), `strategy`
(`collect|merge|vote|reduce`), `min_success_ratio` (decimal string, e.g. `0.8`),
`caller_token` (containment). Env: `SWARM_TRACK_DIR` (journal dir),
`CCR_TOKEN_BUDGET`, `ENGRAM_URL`/`ENGRAM_API_KEY` (retrieval + mirror),
`SWARM_MIRROR=1`.
## Tests
```bash
lang/swarm/build.sh lang/swarm/tests/test_swarm.el /tmp/t && SWARM_TRACK_DIR=/tmp/trk /tmp/t # 12/12
lang/swarm/build.sh lang/swarm/tests/test_convergence.el /tmp/c && SWARM_TRACK_DIR=/tmp/trk /tmp/c # 8/8
# integration against an isolated engram clone (never live):
source <sandbox>/.nsbx-env
lang/swarm/build.sh lang/swarm/tests/integ_engram.el /tmp/i && /tmp/i
```
## Local-swarm integration harness (the one flip)
`tests/harness_local_swarm.el` proves the **full local-swarm mechanics today** on
the isolated clone with the primitive seam pointed at the hermetic stub — 17/17
green: 8 native-thread workers at concurrency 4, reduce + vote convergence, CCR
scoping + non-leak, all three containment rules (incl. live Rule-2 denial),
durable work-tracking, and **afferent telemetry** observed by the @manager.
Binding to the reshape's decorated primitives is **one flip and a run**:
```
# in primitive_binding.el — change one line each:
fn bound_think(ctx, instruction) { return think(ctx, instruction) } # decorated, dharma bus
# then:
SWARM_PRIMITIVE_SEAM=decorated lang/swarm/build.sh tests/harness_local_swarm.el ./h && ./h
```
Nothing else in the swarm changes. `primitive_seam.el` (`seam_think/attend/learn`)
already routes every worker primitive call through this one switch, and the same
harness runs the bound path. Today `SWARM_PRIMITIVE_SEAM=decorated` still runs
green because the binding falls back to the stub — proving the flip path executes.
## Real cognition — the seam is BOUND
`primitive_binding.el` is bound to the api-reshape agent's proven primitives
(`wt/api-reshape@d4f401d`): `bound_think -> op_think` (GET `/api/think`), real
768-dim gradients over the engram geometry. `reshape_surface.el` composes those
read/cognition primitives verbatim (`op_think/read/attend/learn`).
`tests/harness_real_cognition.el` runs the **local swarm on real cognition**,
17/17 green with `SWARM_PRIMITIVE_SEAM=decorated` against the `:8901` clone: 8
native-thread workers, each a real `think` over its CCR-scoped **node-id anchor**
(free-text anchors return "geometry unavailable"), `@manager` reduce+vote, all
three containment rules, afferent telemetry, durable tracking. Per-anchor support
counts (e.g. 6 / 16 / 87) drive a genuine, cognition-derived vote.
> **Build note (load-bearing):** the swarm build **must** define `HAVE_CURL`
> (`build.sh` does). Without it every `http_*` builtin is a
> `{"error":"not built with HAVE_CURL"}` stub — real HTTP silently disappears.
Writes (`attend`/`learn`, `POST`) are gated behind `SWARM_WRITE_HEALTHY=1` and the
api-reshape agent's gate-1 write-healthy clone; the proven run is read-cognition.
## Built vs stubbed (honest)
**Real, tested:**
- Native-thread fan-out/converge, bounded concurrency, order-preserving rejoin.
- All three containment rules enforced (scope tokens + lateral-edge check).
- CCR per-worker context: retrieval → scoping → compaction, bounded, non-leaking
(a worker never receives sibling inputs) — verified against the live isolated mind.
- Full durable work-tracking (JSONL journal, reconstructable report).
- Four convergence strategies + integer failure threshold / partial-abort.
**Seam / not yet bound:**
- `primitives.el` `think` is a deterministic, hermetic transform (no model call).
Binding point is marked `PRIMITIVE_BINDING`; wire to the API-surface reshape's
`think/act/attend/intend/learn` when it lands.
- Blueprints are dispatched by name in `swarm_run_blueprint` (default +
`classify`/`faildemo` demos). A YAML process-definition loader (Swarm §5) is
future work — the runtime contract is in place.
- Distributed placement (cloud/edge/federated topologies, Swarm §9.2) is
structurally enabled by containment but not yet wired to a placement layer;
today all workers are local native threads.
- Engram work-tracking mirror is opt-in; the durable substrate is the journal.
+61
View File
@@ -0,0 +1,61 @@
#!/usr/bin/env bash
# build.sh — compile an El program that uses the swarm capability.
#
# Concatenates the El native-concurrency stdlib (thread.el, channel.el) and the
# swarm capability modules in dependency order, then the user program, compiles
# with the canonical elc, and links against the shared C runtime.
#
# Usage:
# swarm/build.sh <program.el> <out-binary>
#
# The swarm modules use only el_runtime.c builtins plus thread.el/channel.el,
# so nothing else needs concatenating (engram_*, json_*, str_*, fs_*, http_*,
# uuid_v4, now_millis are all C builtins in el_runtime.c).
set -uo pipefail
cd "$(dirname "$0")/.." # -> lang/
LANG_DIR="$(pwd)"
ELC="${ELC:-${LANG_DIR}/dist/platform/elc}"
RT="${LANG_DIR}/el-compiler/runtime"
PROG="${1:?usage: build.sh <program.el> <out-binary>}"
OUT="${2:?usage: build.sh <program.el> <out-binary>}"
# swarm module load order (each may depend on those before it):
# worktrack — durable work-tracking journal (no swarm deps)
# containment — the three containment rules (no swarm deps)
# primitives — think/act/attend/intend/learn seam (no swarm deps)
# ccr — per-worker compiled bounded context (depends: primitives)
# swarm — orchestrator: fan-out/converge (depends: all above + thread)
SWARM_MODULES="
swarm/worktrack.el
swarm/containment.el
swarm/primitives.el
swarm/reshape_surface.el
swarm/primitive_binding.el
swarm/primitive_seam.el
swarm/ccr.el
swarm/swarm.el
"
TMP_C="$(mktemp -t swarm_build.XXXXXX).c"
COMBINED="$(mktemp -t swarm_combined.XXXXXX).el"
cat runtime/thread.el runtime/channel.el $SWARM_MODULES "$PROG" > "$COMBINED"
if ! "$ELC" "$COMBINED" > "$TMP_C" 2>/tmp/swarm.elc.err; then
echo "elc FAILED:" >&2
sed 's/^/ /' /tmp/swarm.elc.err >&2
rm -f "$TMP_C" "$COMBINED"
exit 1
fi
if ! cc -O2 -DHAVE_CURL -I "$RT" "$TMP_C" "$RT/el_runtime.c" -lcurl -lpthread -lm -o "$OUT" 2>/tmp/swarm.cc.err; then
echo "cc FAILED:" >&2
sed 's/^/ /' /tmp/swarm.cc.err >&2
rm -f "$TMP_C" "$COMBINED"
exit 1
fi
rm -f "$TMP_C" "$COMBINED"
echo "built: $OUT"
+153
View File
@@ -0,0 +1,153 @@
// ccr.el Compiled Context Routing for work distribution.
//
// The same spine as the API's vantage-read, applied per worker. Instead of
// handing every worker the coordinator's full memory, CCR compiles a MINIMAL,
// BOUNDED context package scoped to exactly one worker's input (CCR §5, "Compiled
// Context Injection"; Swarm §9.3, "The Compiled Context Boundary as Security
// Boundary").
//
// The pipeline is CCR §5.1: Retrieval -> Scoping -> Compilation -> (Injection,
// which here is placing the package into the worker's task envelope).
//
// 1. Retrieval resolve the blueprint's knowledge refs + the input's salient
// terms against the mind (primitive_attend).
// 2. Scoping keep only what THIS input needs; drop everything else. A
// worker never receives sibling inputs or unrelated memory.
// 3. Compilation compact to a CTX string within a token budget (lossless of
// meaning, smaller in tokens): collapse blank runs, dedupe
// lines, then bound to the budget.
//
// The package a worker receives is therefore (a) sufficient for its task and
// (b) incapable of leaking what it was never given the containment boundary
// and the security boundary are the same object.
// token budget helpers
// ccr_est_tokens cheap token estimate (~4 chars/token).
fn ccr_est_tokens(s: String) -> Int {
return str_len(s) / 4
}
// ccr_default_budget default per-worker context budget in tokens.
// Override with CCR_TOKEN_BUDGET.
fn ccr_default_budget() -> Int {
let b: String = env("CCR_TOKEN_BUDGET")
if str_eq(b, "") {
return 1200
}
return str_to_int(b)
}
// stage 3: compaction
// ccr_compact collapse blank-line runs and drop exact duplicate lines, then
// bound the result to `budget` tokens (truncate on a line boundary). Meaning is
// preserved; token count falls (CCR §5.2).
fn ccr_compact(text: String, budget: Int) -> String {
let lines: [String] = str_split_lines(text)
let n: Int = el_list_len(lines)
let seen: String = "\n"
let out: String = ""
let out_tokens = 0
let i = 0
while i < n {
let ln: String = str_trim(el_list_get(lines, i))
if str_eq(ln, "") {
let i = i + 1
} else {
let marker: String = "\n" + ln + "\n"
if str_contains(seen, marker) {
// duplicate line skip
let i = i + 1
} else {
let seen = seen + ln + "\n"
let line_tokens: Int = ccr_est_tokens(ln) + 1
if out_tokens + line_tokens > budget {
// budget exhausted stop (bounded)
let i = n
} else {
let out = out + ln + "\n"
let out_tokens = out_tokens + line_tokens
let i = i + 1
}
}
}
}
return out
}
// stages 1+2: retrieve + scope
// ccr_retrieve_scoped pull context relevant to this input and its blueprint
// knowledge refs, scoped to a fraction of the budget so no single source floods
// the package. Returns compacted retrieved text (may be empty if the mind is
// unreachable the input alone is still a valid minimal context).
fn ccr_retrieve_scoped(blueprint: String, knowledge_refs: String, input_item: String, budget: Int) -> String {
let acc: String = ""
// knowledge_refs is a JSON array of query strings.
let m: Int = json_array_len(knowledge_refs)
let i = 0
while i < m {
let ref: String = json_array_get_string(knowledge_refs, i)
let hit: String = primitive_attend(ref, 3)
let acc = acc + "# ref:" + ref + "\n" + hit + "\n"
let i = i + 1
}
// the input's own salient text also seeds retrieval
let hit2: String = primitive_attend(input_item, 3)
let acc = acc + "# input-context\n" + hit2 + "\n"
// scope retrieval to ~60% of budget; the input itself gets the rest
let retr_budget: Int = (budget * 6) / 10
return ccr_compact(acc, retr_budget)
}
// ccr_compile assemble the bounded per-worker context package
//
// blueprint : task blueprint name
// knowledge_refs : JSON array of retrieval queries from the blueprint
// input_item : THIS worker's single input (and nothing else)
// corr_id : swarm correlation ID
// worker_id : this worker's ID
// scope_token : the worker's containment token (closed boundary)
//
// Returns a JSON package: { blueprint, corr_id, worker_id, scope_token,
// input, knowledge, budget_tokens, compiled_tokens }. `knowledge` is compiled
// and bounded; the package as a whole is bounded by budget.
fn ccr_compile(blueprint: String, knowledge_refs: String, input_item: String,
corr_id: String, worker_id: String, scope_token: String) -> String {
let budget: Int = ccr_default_budget()
let knowledge: String = ccr_retrieve_scoped(blueprint, knowledge_refs, input_item, budget)
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "blueprint")
let kv = el_list_append(kv, blueprint)
let kv = el_list_append(kv, "corr_id")
let kv = el_list_append(kv, corr_id)
let kv = el_list_append(kv, "worker_id")
let kv = el_list_append(kv, worker_id)
let kv = el_list_append(kv, "input")
let kv = el_list_append(kv, input_item)
let kv = el_list_append(kv, "knowledge")
let kv = el_list_append(kv, knowledge)
let kv = el_list_append(kv, "budget_tokens")
let kv = el_list_append(kv, int_to_str(budget))
let pkg: String = json_build_object(kv)
// stamp the scope token as a nested object, and the measured size
let pkg2: String = json_set(pkg, "scope_token", scope_token)
let compiled_tokens: Int = ccr_est_tokens(pkg2)
let pkg3: String = json_set(pkg2, "compiled_tokens", int_to_str(compiled_tokens))
return pkg3
}
// ccr_within_budget did the compiled package stay within its budget?
// (Retrieval is bounded to 60% and the input is small; this asserts the whole
// package is bounded the property distribution relies on.)
fn ccr_within_budget(pkg: String) -> Bool {
let budget: Int = str_to_int(json_get_string(pkg, "budget_tokens"))
let compiled: Int = str_to_int(json_get_string(pkg, "compiled_tokens"))
// allow a small envelope for JSON framing overhead
if compiled <= budget + 200 {
return true
}
return false
}
+181
View File
@@ -0,0 +1,181 @@
// containment.el the Swarm Architecture containment rules, enforced.
//
// "These rules are not conventions. They are enforced by the runtime."
// (Swarm Architecture §3.2). The three rules that make bounded parallelism
// and therefore location-independent distribution safe:
//
// Rule 1: a worker may NOT join another swarm.
// Rule 2: a worker may NOT initiate a new swarm.
// Rule 3: a worker may NOT communicate laterally with sibling workers.
//
// Enforcement is by SCOPE TOKEN. When a swarm fans out, the coordinator mints a
// swarm scope token and stamps a distinct worker scope token into each worker's
// task envelope. Any attempt to create or join a swarm checks the caller's
// token: if the caller already holds a WORKER token, the operation is rejected.
// Rule 3 is enforced structurally elsewhere workers share no mutable state and
// the only channels they hold are the vertical result path but this module
// provides the explicit lateral-edge check for the execution tree.
//
// A scope token is a JSON object: {"kind":"coordinator|worker","swarm":"<corr>",
// "worker":"<id-or-empty>","depth":"<n>"}.
// Token minting
// CAPABILITIES. A scope token carries a `caps` set the authority it holds.
// This is an AUTHORITY gate, not a health gate: capability is decided at mint
// time and cannot be acquired at runtime. Engram-WRITE (op_write/op_relate/
// op_supersede -> POST /api/nodes, /api/edges, DELETE) and dharma_emit are
// @manager-ONLY capabilities exactly the VBD rule that only the orchestrator
// mutates global state. The orchestrator's token carries them; a worker's token
// NEVER does. A worker is therefore STRUCTURALLY UNABLE to mutate global engram
// state, regardless of engram health.
fn cap_orchestrator() -> String { return "engram:read,engram:write,dharma:emit,state:write" }
fn cap_worker() -> String { return "engram:read" }
// containment_coordinator_token the token the orchestrator (@manager) holds.
// Depth 0. Carries the engram-WRITE + dharma-emit capabilities (@manager-only).
fn containment_coordinator_token(corr_id: String) -> String {
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "kind")
let kv = el_list_append(kv, "coordinator")
let kv = el_list_append(kv, "swarm")
let kv = el_list_append(kv, corr_id)
let kv = el_list_append(kv, "worker")
let kv = el_list_append(kv, "")
let kv = el_list_append(kv, "depth")
let kv = el_list_append(kv, "0")
let kv = el_list_append(kv, "caps")
let kv = el_list_append(kv, cap_orchestrator())
return json_build_object(kv)
}
// containment_worker_token the token stamped into a worker's envelope. Depth 1.
// A closed boundary: forbids opening/joining swarms AND carries ONLY the
// engram:READ capability no engram:write, no dharma:emit. Read-only against the
// full engram; may write only its own local geometry (its returned result).
fn containment_worker_token(corr_id: String, worker_id: String) -> String {
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "kind")
let kv = el_list_append(kv, "worker")
let kv = el_list_append(kv, "swarm")
let kv = el_list_append(kv, corr_id)
let kv = el_list_append(kv, "worker")
let kv = el_list_append(kv, worker_id)
let kv = el_list_append(kv, "depth")
let kv = el_list_append(kv, "1")
let kv = el_list_append(kv, "caps")
let kv = el_list_append(kv, cap_worker())
return json_build_object(kv)
}
// containment_has_cap does this token carry capability `cap`?
fn containment_has_cap(token: String, cap: String) -> Bool {
return str_contains(json_get_string(token, "caps"), cap)
}
// Rule checks (return "" on allow, or a rejection reason string)
// containment_check_open may the holder of `token` OPEN a new swarm?
// Enforces Rule 2 (a worker may not initiate a new swarm). Only a coordinator
// token, or an absent token (top-level process), may open one.
fn containment_check_open(token: String) -> String {
if str_eq(token, "") {
return ""
}
let kind: String = json_get_string(token, "kind")
if str_eq(kind, "worker") {
return "CONTAINMENT rule 2: a swarm worker may not initiate a new swarm (worker=" + json_get_string(token, "worker") + " swarm=" + json_get_string(token, "swarm") + ")"
}
return ""
}
// containment_check_join may the holder of `token` JOIN swarm `target_corr`?
// Enforces Rule 1 (a worker may not join another swarm). A worker already bound
// to swarm A may not register into swarm B; and a worker may not re-join at all.
fn containment_check_join(token: String, target_corr: String) -> String {
if str_eq(token, "") {
return ""
}
let kind: String = json_get_string(token, "kind")
if str_eq(kind, "worker") {
return "CONTAINMENT rule 1: a swarm worker may not join another swarm (worker=" + json_get_string(token, "worker") + " bound-swarm=" + json_get_string(token, "swarm") + " attempted-swarm=" + target_corr + ")"
}
return ""
}
// containment_check_lateral may `from_token` open a communication edge to a
// sibling worker `to_worker_id`? Enforces Rule 3 (no lateral communication).
// The only permitted edges are vertical: worker->coordinator and
// coordinator->worker. Any worker->worker edge is rejected.
fn containment_check_lateral(from_token: String, to_worker_id: String) -> String {
let kind: String = json_get_string(from_token, "kind")
if str_eq(kind, "worker") {
if str_eq(to_worker_id, "") {
// empty target = the coordinator (vertical) allowed
return ""
}
return "CONTAINMENT rule 3: a swarm worker may not communicate laterally with sibling workers (from=" + json_get_string(from_token, "worker") + " to=" + to_worker_id + ")"
}
return ""
}
// containment_check_engram_write RULE 4: only a token carrying the
// engram:write capability (the orchestrator's) may mutate global engram state.
// A worker token (engram:read only) is REJECTED the authority gate. Reuses the
// exact scope-token mechanism as Rule 2's open-denial. Returns "" on allow, or a
// rejection reason. This is an AUTHORITY gate: it does not consult engram health.
fn containment_check_engram_write(token: String, op: String) -> String {
if containment_has_cap(token, "engram:write") {
return ""
}
return "CONTAINMENT rule 4: engram-write is @manager-only — a worker is read-only against the engram and may not mutate global state (op=" + op + " kind=" + json_get_string(token, "kind") + " worker=" + json_get_string(token, "worker") + " caps=" + json_get_string(token, "caps") + ")"
}
// containment_check_dharma_emit the same @manager-only rule for dharma_emit,
// grounding Rule 4 in VBD: global-state mutations (engram-write, dharma-emit) are
// orchestrator-only, checked by the one capability mechanism.
fn containment_check_dharma_emit(token: String) -> String {
if containment_has_cap(token, "dharma:emit") {
return ""
}
return "CONTAINMENT rule 4: dharma_emit is @manager-only (kind=" + json_get_string(token, "kind") + ")"
}
// Enforcement helpers
// containment_allows_open Bool convenience over containment_check_open.
fn containment_allows_open(token: String) -> Bool {
return str_eq(containment_check_open(token), "")
}
// containment_is_worker is this a worker-scoped (closed-boundary) token?
fn containment_is_worker(token: String) -> Bool {
return str_eq(json_get_string(token, "kind"), "worker")
}
// containment_guard_open assert a swarm may be opened under this token.
// Returns "" if allowed, or records a CONTAINMENT violation to the work-tracking
// journal and returns the reason. Callers must abort on a non-empty return.
fn containment_guard_open(token: String, corr_id: String) -> String {
let reason: String = containment_check_open(token)
if str_eq(reason, "") {
return ""
}
let p: String = json_set_str("{}", "reason", reason)
worktrack_append("containment.violation", corr_id, "open", p)
return reason
}
// containment_guard_engram_write assert a token may mutate global engram state
// (Rule 4). Returns "" if allowed; otherwise journals a containment.violation and
// returns the reason. The write path MUST abort on a non-empty return.
fn containment_guard_engram_write(token: String, corr_id: String, op: String) -> String {
let reason: String = containment_check_engram_write(token, op)
if str_eq(reason, "") {
return ""
}
let p0: String = json_set_str("{}", "reason", reason)
let p1: String = json_set_str(p0, "op", op)
worktrack_append("containment.violation", corr_id, "engram-write", p1)
return reason
}
+49
View File
@@ -0,0 +1,49 @@
// primitive_binding.el THE ONE FLIP POINT.
//
// This file is the single seam between the swarm and the real agentic
// primitives. Binding the reshape's decorated primitives is a one-line change
// HERE and nothing else changes anywhere in the swarm.
//
// The api-reshape agent (wt/api-reshape) is wiring the primitives as DECORATED
// El on the dharma_* event bus over the engram think/attend/learn/ground/assert
// become decorated fns that emit afferent events onto the bus. The moment they
// land, flip `bound_think` (and its siblings) to call them.
//
// TODAY (stub fallback, compiles + runs now against :8901):
// fn bound_think(...) { return primitive_think(ctx, instruction) }
//
// THE FLIP (when reshape's decorated primitives land one line each):
// fn bound_think(...) { return think(ctx, instruction) } // decorated, on dharma bus
//
// Keep the stub as fallback: `bound_think` is only reached when the seam mode is
// "decorated" (SWARM_PRIMITIVE_SEAM=decorated). Until you flip these bodies AND
// set that env, the harness runs entirely on the hermetic stub.
// bound_think BOUND to the reshape's proven decorated `think` (op_think),
// real cognition over the engram geometry. The worker's CCR slice carries a
// NODE-ID anchor in ctx.input (free-text anchors return "geometry unavailable");
// think re-origins at that node's region under the faculty and returns a real
// 768-dim gradient.
fn bound_think(ctx: String, instruction: String) -> String {
let anchor: String = json_get_string(ctx, "input")
let faculty: String = json_get_string(ctx, "faculty")
return op_think(anchor, faculty)
}
// bound_attend BOUND to the reshape's op_attend (POST /api/attend). Needs the
// gate-1 write-healthy clone; falls back to the read-side attend otherwise.
fn bound_attend(query: String, limit: Int) -> String {
if str_eq(env("SWARM_WRITE_HEALTHY"), "1") {
return op_attend(query, "self")
}
return primitive_attend(query, limit)
}
// bound_learn BOUND to the reshape's op_learn (correspondence-beat). Needs the
// gate-1 write-healthy clone; falls back to the opt-in journal-only learn.
fn bound_learn(corr_id: String, observation: String) -> String {
if str_eq(env("SWARM_WRITE_HEALTHY"), "1") {
return op_learn(observation, "induce")
}
return primitive_learn(corr_id, observation)
}
+55
View File
@@ -0,0 +1,55 @@
// primitive_seam.el the configurable primitive seam + telemetry.
//
// One switch selects where a worker's primitive invocation goes:
// SWARM_PRIMITIVE_SEAM=stub (default) hermetic in-process think.
// SWARM_PRIMITIVE_SEAM=decorated the reshape's decorated
// primitives on the dharma bus
// (see primitive_binding.el).
//
// Every seam invocation is an AFFERENT signal a primitive call travelling
// toward the manager. The seam stamps telemetry onto each thought (seam_mode +
// one afferent tick) so the coordinator can aggregate afferent counters across
// the swarm without any shared mutable state (containment-safe: counts ride the
// vertical result path, not a shared bus register).
// seam_mode "stub" (default) or "decorated".
fn seam_mode() -> String {
let m: String = env("SWARM_PRIMITIVE_SEAM")
if str_eq(m, "decorated") {
return "decorated"
}
return "stub"
}
// seam_think route a worker's `think` through the configured seam and stamp
// telemetry. Returns the thought JSON augmented with:
// seam_mode : which side of the seam served this call
// afferent : "1" one afferent primitive signal was emitted
fn seam_think(ctx: String, instruction: String) -> String {
let mode: String = seam_mode()
let thought: String = ""
if str_eq(mode, "decorated") {
let thought = bound_think(ctx, instruction)
} else {
let thought = primitive_think(ctx, instruction)
}
let t1: String = json_set_str(thought, "seam_mode", mode)
let t2: String = json_set_str(t1, "afferent", "1")
return t2
}
// seam_attend / seam_learn same seam for the other primitives (used when a
// blueprint retrieves or writes through the bus).
fn seam_attend(query: String, limit: Int) -> String {
if str_eq(seam_mode(), "decorated") {
return bound_attend(query, limit)
}
return primitive_attend(query, limit)
}
fn seam_learn(corr_id: String, observation: String) -> String {
if str_eq(seam_mode(), "decorated") {
return bound_learn(corr_id, observation)
}
return primitive_learn(corr_id, observation)
}
+94
View File
@@ -0,0 +1,94 @@
// primitives.el the agentic primitive SEAM the swarm composes over.
//
// The swarm is orchestration OVER the five CCR primitives, not a replacement for
// them (CCR §2, "The Five Primitives / The Execution Cycle"): a worker executes
// its task blueprint as attend -> think -> intend -> act -> learn against its
// compiled, bounded context.
//
// This file is the SEAM. The parallel API-surface reshape exposes the canonical
// primitive tools; when it lands, bind each primitive below to the reshaped
// implementation (see PRIMITIVE_BINDING). Until then these are thin, engram-
// backed fallbacks so the swarm its fan-out, containment, CCR context
// compilation, convergence, and work-tracking is fully exercisable today.
//
// Contract: every primitive takes and returns String (JSON where structured), so
// any primitive is directly threadable via thread.el's spawn (which runs
// top-level (String)->String El fns).
//
// PRIMITIVE_BINDING: to bind the reshape's real tools, replace each fallback body
// with a call to the reshaped El fn / API endpoint. Signatures here are the
// stable contract the swarm depends on; keep them.
// attend retrieve the minimal relevant context for a focus
// Vantage-read: pull only what this focus needs from the mind. Backed by the
// engram's spreading-activation retrieval.
fn primitive_attend(query: String, limit: Int) -> String {
if str_eq(query, "") {
return "[]"
}
// Location-independent worker model: when an engram daemon is configured,
// retrieve over HTTP (the worker may run anywhere). POST /api/search
// {query,limit,_auth}. Falls back to the in-process store otherwise.
let url: String = env("ENGRAM_URL")
if str_eq(url, "") {
return engram_activate(query, limit)
}
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "query")
let kv = el_list_append(kv, query)
let body0: String = json_build_object(kv)
let body1: String = json_set(body0, "limit", int_to_str(limit))
let body2: String = json_set_str(body1, "_auth", env("ENGRAM_API_KEY"))
return http_post(url + "/api/search", body2)
}
// think reason over the compiled context
// In production this routes to a model (CCR dynamic model selection). Here it is
// a deterministic, hermetic transform so swarm behaviour is testable without an
// external model: it echoes a structured verdict derived from the context. The
// binding point for a real model is explicit.
fn primitive_think(compiled_ctx: String, instruction: String) -> String {
// PRIMITIVE_BINDING: replace with the reshape's think() (model inference).
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "instruction")
let kv = el_list_append(kv, instruction)
let kv = el_list_append(kv, "ctx_bytes")
let kv = el_list_append(kv, int_to_str(str_len(compiled_ctx)))
let kv = el_list_append(kv, "conclusion")
let kv = el_list_append(kv, "reasoned:" + instruction)
return json_build_object(kv)
}
// intend form a bounded plan/decision from a thought
fn primitive_intend(thought: String) -> String {
let concl: String = json_get_string(thought, "conclusion")
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "intent")
let kv = el_list_append(kv, concl)
return json_build_object(kv)
}
// act execute a bounded effect and return its result
// Workers defer real side-effects to the coordinator (idempotency requirement,
// Swarm §7.3). Here act produces an artifact-shaped result the coordinator
// collects during convergence.
fn primitive_act(intent: String, input_item: String) -> String {
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "acted_on")
let kv = el_list_append(kv, input_item)
let kv = el_list_append(kv, "via")
let kv = el_list_append(kv, json_get_string(intent, "intent"))
return json_build_object(kv)
}
// learn record an observation into the mind, tagged by correlation ID
// Append-only, naturally idempotent (Swarm §7.3). Best-effort: a worker that
// cannot reach the mind still returns its result.
fn primitive_learn(corr_id: String, observation: String) -> String {
let url: String = env("ENGRAM_URL")
if str_eq(url, "") {
return ""
}
let content: String = "swarm-worker-obs corr=" + corr_id + " :: " + observation
return engram_node(content, "Memory", 0.4)
}
+103
View File
@@ -0,0 +1,103 @@
// reshape_surface.el the api-reshape agent's PROVEN decorated primitives,
// composed into the swarm build to bind real cognition.
//
// PROVENANCE: these fns are the reshape's surface at wt/api-reshape @ d4f401d
// ("reshape: decorator-as-seam — port @route codegen, prove decorate->serve,
// rewrite surface as decorated El"), verified live against
// engram.cognition-20260814. Copied verbatim (read/cognition ops only) so the
// swarm binds the REAL primitives, not a reimplementation. The write ops
// (op_write/op_relate/op_supersede/op_ground) are intentionally NOT composed
// here they exercise the persist_node write path that needs the gate-1
// write-healthy clone; the swarm's proven run is read-cognition (think/read).
//
// Ops route to the ENGRAM over ENGRAM_URL pinned by THIS worktree's .nsbx-env
// to the :8901 swarm clone (never the reshape agent's :8900). Separate clones,
// no collision.
fn engram_url() -> String {
let u: String = env("ENGRAM_URL")
if str_eq(u, "") { return "http://127.0.0.1:8900" }
return u
}
fn engram_key() -> String {
let k: String = env("ENGRAM_API_KEY")
if str_eq(k, "") { return "sbx-dev-api-reshape" }
return k
}
fn SELF_KEY() -> String { return "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee" }
fn VALUES_KEY() -> String { return "kn-5b606390-a52d-4ca2-8e0e-eba141d13440" }
// self/values name -> keystone id; anything else passes through unchanged.
fn resolve_named(v: String) -> String {
if str_eq(v, "self") { return SELF_KEY() }
if str_eq(v, "neuron") { return SELF_KEY() }
if str_eq(v, "values") { return VALUES_KEY() }
if str_eq(v, "values_hub") { return VALUES_KEY() }
return v
}
// read THE VANTAGE-READ. Re-origin at a point + aperture -> a BOUNDED slice.
fn op_read(vantage: String, typ: String, k: Int) -> String {
let vid: String = resolve_named(vantage)
if str_eq(typ, "edges") {
return http_get(engram_url() + "/api/neighbors/" + vid)
}
if str_starts_with(vid, "kn-") {
return http_get(engram_url() + "/api/neighbors/" + vid)
}
return http_get(engram_url() + "/api/search?q=" + url_encode(vid) + "&limit=" + int_to_str(k))
}
// think THE ONE OPERATION. anchor (node ids) steered by faculty -> gradient.
fn op_think(seeds: String, faculty: String) -> String {
let s: String = resolve_named(seeds)
let f: String = if str_eq(faculty, "") { "reason" } else { faculty }
return http_get(engram_url() + "/api/think?seeds=" + url_encode(s) + "&faculty=" + f)
}
// attend aim attention at a region. (POST needs a write-healthy clone.)
fn op_attend(node: String, observer: String) -> String {
let n: String = resolve_named(node)
let o: String = if str_eq(observer, "") { SELF_KEY() } else { resolve_named(observer) }
let body: String = "{\"_auth\":\"" + engram_key() + "\",\"node\":\"" + n
+ "\",\"observer\":\"" + o + "\",\"salience\":\"0.6\"}"
return http_post_json(engram_url() + "/api/attend", body)
}
fn identity_typed(t: String) -> Bool {
if str_eq(t, "self") { return true }
if str_eq(t, "values") { return true }
return false
}
fn type_to_node_type(t: String) -> String {
if str_eq(t, "knowledge") { return "Knowledge" }
if str_eq(t, "artifact") { return "Artifact" }
if str_eq(t, "backlog") { return "WorkItem" }
if str_eq(t, "process") { return "Process" }
if str_eq(t, "state") { return "InternalStateEvent" }
return "Memory"
}
// write add a node (POST /api/nodes). Identity types refused. This is a
// global-engram MUTATION @manager-only (Rule 4); never called on a worker path.
// (Reshape's op_write, with json_escape -> the available json_escape_string.)
fn op_write(content: String, typ: String, importance: Float) -> String {
if str_eq(content, "") { return "{\"error\":\"write: content required\"}" }
if identity_typed(typ) {
return "{\"error\":\"write type=" + typ + " is write-protected -> intentional-cultivation\"}"
}
let body: String = "{\"_auth\":\"" + engram_key() + "\",\"content\":\"" + json_escape_string(content)
+ "\",\"node_type\":\"" + type_to_node_type(typ) + "\",\"tier\":\"Working\",\"importance\":"
+ float_to_str(importance) + "}"
return http_post_json(engram_url() + "/api/nodes", body)
}
// learn the reflexive correspondence-beat: calibrate the steering-prior.
// (POST needs a write-healthy clone.)
fn op_learn(seeds: String, faculty: String) -> String {
let s: String = resolve_named(seeds)
let f: String = if str_eq(faculty, "") { "induce" } else { faculty }
let body: String = "{\"_auth\":\"" + engram_key() + "\",\"seeds\":\"" + s
+ "\",\"faculty\":\"" + f + "\",\"keystone\":\"false\"}"
return http_post_json(engram_url() + "/api/correspondence-beat", body)
}
+483
View File
@@ -0,0 +1,483 @@
// swarm.el the swarm orchestrator: bounded parallel agent execution.
//
// Implements Swarm Architecture's single pattern fan out, execute independently,
// converge on El's NATIVE concurrency (thread.el spawn/join). No external
// orchestrator: a swarm is a coordinator (this file, the main thread) that mints
// a correlation identity, compiles a bounded CCR context per worker, dispatches
// workers as native pthreads, tracks every unit of work, and converges the
// results before returning control to the parent step.
//
// The five properties of every swarm (Swarm §2.1) are all present:
// parent step -> swarm_run is called from one process step
// task blueprint -> `blueprint` name + knowledge refs, run by every worker
// input set -> `inputs_json`, one item per worker
// convergence -> `strategy` in config (collect|merge|vote|reduce)
// correlation ID -> minted here, threaded through tracking + every worker
//
// Containment (Swarm §3) is enforced: the caller must hold a coordinator/absent
// token to open a swarm (Rule 2), each worker is stamped a closed worker token
// (Rules 1+3), and workers share no mutable state (the coordinator is the only
// journal writer).
// worker entry the top-level (String)->String fn native threads run
//
// Every El fn compiles to a global C symbol; spawn() resolves this by name via
// dlsym and runs it in a pthread. The envelope carries everything the worker is
// permitted to see its compiled context and nothing else (§9.3).
//
// Returns a result JSON: {worker_id, status:"completed"|"failed", output|error}.
fn swarm_worker_entry(envelope_json: String) -> String {
let worker_id: String = json_get_string(envelope_json, "worker_id")
let ctx: String = json_get_raw(envelope_json, "ctx")
// The worker holds a CLOSED worker token (Rules 1+3): it shares no state
// with siblings and may not open/join a swarm. That boundary is enforced at
// the point of attempt swarm_run rejects any swarm opened under a worker
// token (Rule 2). A worker simply executing its blueprint is not opening a
// swarm, so it proceeds. Its only outward edge is this returned result
// (the vertical worker->coordinator path).
let out: String = swarm_run_blueprint(ctx)
// A worker reports failed iff its blueprint signalled failure. This is the
// vertical status edge the coordinator reads during convergence (§4.3, §7).
let bstatus: String = json_get_string(out, "blueprint_status")
let status: String = "completed"
if str_eq(bstatus, "failed") {
let status = "failed"
}
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "worker_id")
let kv = el_list_append(kv, worker_id)
let kv = el_list_append(kv, "status")
let kv = el_list_append(kv, status)
let res: String = json_build_object(kv)
return json_set(res, "output", out)
}
// swarm_run_blueprint execute the task blueprint over a compiled context.
// The default blueprint is the CCR execution cycle: think -> intend -> act over
// the worker's bounded context. Specialise by dispatching on
// json_get_string(ctx,"blueprint"). Idempotent: reads ctx, writes only its
// returned output (§7.3).
fn swarm_run_blueprint(ctx: String) -> String {
let blueprint: String = json_get_string(ctx, "blueprint")
let input_item: String = json_get_string(ctx, "input")
let knowledge: String = json_get_string(ctx, "knowledge")
// classify deterministic verdict for the `vote` convergence strategy:
// verdict is "long" if the input has >4 chars, else "short".
if str_eq(blueprint, "classify") {
let verdict: String = "short"
if str_len(input_item) > 4 {
let verdict = "long"
}
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "verdict")
let kv = el_list_append(kv, verdict)
let kv = el_list_append(kv, "blueprint_status")
let kv = el_list_append(kv, "ok")
return json_build_object(kv)
}
// faildemo a worker that fails on inputs beginning with "x" (exercises the
// failure threshold + partial convergence path). Idempotent, side-effect-free.
if str_eq(blueprint, "faildemo") {
let st: String = "ok"
if str_starts_with(input_item, "x") {
let st = "failed"
}
return json_set_str("{}", "blueprint_status", st)
}
// cognize REAL-COGNITION blueprint. Routes think through the seam (bound to
// op_think in decorated mode) over the worker's NODE-ID anchor, then derives a
// vote verdict from the gradient's confidence. In stub mode there is no
// gradient, so the verdict falls back to a deterministic slice hash the
// same blueprint runs green on either side of the seam.
if str_eq(blueprint, "cognize") {
let thought: String = seam_think(ctx, "reason over " + input_item)
// Derive the vote verdict from the REAL gradient's support count
// (json_get_int, since n_support is numeric). Different anchors have
// different support -> genuine, cognition-driven vote diversity. In stub
// mode there is no gradient (n_support -> 0) -> "uncertain".
let nsup: Int = json_get_int(thought, "n_support")
let verdict: String = "uncertain"
if nsup >= 10 {
let verdict = "confident"
}
let ck: [String] = el_list_empty()
let ck = el_list_append(ck, "verdict")
let ck = el_list_append(ck, verdict)
let ck = el_list_append(ck, "blueprint_status")
let ck = el_list_append(ck, "ok")
let cout0: String = json_build_object(ck)
let cout1: String = json_set_str(cout0, "n_support", int_to_str(nsup))
let cout2: String = json_set_str(cout1, "seam_mode", json_get_string(thought, "seam_mode"))
return json_set_str(cout2, "afferent", json_get_string(thought, "afferent"))
}
// default (analyze_item): the CCR execution cycle think -> intend -> act,
// with `think` routed through the CONFIGURABLE PRIMITIVE SEAM. Telemetry
// (seam_mode + afferent tick) rides the worker's returned output.
let instruction: String = "process input: " + input_item
let thought: String = seam_think(ctx, instruction)
let intent: String = primitive_intend(thought)
let effect: String = primitive_act(intent, input_item)
let e1: String = json_set_str(effect, "blueprint_status", "ok")
let e2: String = json_set_str(e1, "seam_mode", json_get_string(thought, "seam_mode"))
let e3: String = json_set_str(e2, "afferent", json_get_string(thought, "afferent"))
return e3
}
// native-thread fan-out, bounded by concurrency, order-preserving
//
// parallel_map (thread.el) spawns ALL threads at once. The swarm honours the
// blueprint's `concurrency` cap (§5.1: a resource constraint, not a parallelism
// constraint all items are processed, at most N at a time) by dispatching in
// waves of N native threads, joining each wave before the next. Results are
// returned in input order.
fn swarm_fanout(worker_fn: String, envelopes: [String], concurrency: Int) -> [String] {
let n: Int = el_list_len(envelopes)
let cap: Int = concurrency
if cap < 1 {
let cap = 1
}
let results: [String] = el_list_empty()
let base = 0
while base < n {
// spawn a wave of up to `cap` workers
let tids: [String] = el_list_empty()
let k = 0
while k < cap {
let idx: Int = base + k
if idx < n {
let env_item: String = el_list_get(envelopes, idx)
let tid: Int = spawn(worker_fn, env_item)
let tids = el_list_append(tids, int_to_str(tid))
}
let k = k + 1
}
// join the wave in order
let j = 0
let jn: Int = el_list_len(tids)
while j < jn {
let tid: Int = str_to_int(el_list_get(tids, j))
let r: String = join(tid)
let results = el_list_append(results, r)
let j = j + 1
}
let base = base + cap
}
return results
}
// convergence strategies (Swarm §4.2)
// swarm_converge_collect ordered list, no transformation.
fn swarm_converge_collect(results: [String]) -> String {
let n: Int = el_list_len(results)
let arr: String = "[]"
let i = 0
while i < n {
let arr = json_array_push(arr, el_list_get(results, i))
let i = i + 1
}
return arr
}
// swarm_converge_merge combine worker outputs into a single joined string.
fn swarm_converge_merge(results: [String]) -> String {
let n: Int = el_list_len(results)
let merged: String = ""
let i = 0
while i < n {
let out: String = json_get_raw(el_list_get(results, i), "output")
if i > 0 {
let merged = merged + " | "
}
let merged = merged + out
let i = i + 1
}
return json_set_str("{}", "merged", merged)
}
// swarm_converge_vote tally a field across worker outputs, pick the majority.
// Each worker output is expected to carry a "verdict" string field.
fn swarm_converge_vote(results: [String]) -> String {
let n: Int = el_list_len(results)
// Collect verdicts (no mutable tally: json_set can't update an existing key
// and there is no el_list_set). Then count each verdict by rescanning.
let verdicts: [String] = el_list_empty()
let i = 0
while i < n {
let out: String = json_get_raw(el_list_get(results, i), "output")
let v: String = json_get_string(out, "verdict")
if str_eq(v, "") {
let i = i + 1
} else {
let verdicts = el_list_append(verdicts, v)
let i = i + 1
}
}
// pick the verdict with the highest count (first-past-the-post)
let vn: Int = el_list_len(verdicts)
let best: String = ""
let bestc = 0
let a = 0
while a < vn {
let cand: String = el_list_get(verdicts, a)
// count occurrences of cand
let c = 0
let b = 0
while b < vn {
if str_eq(el_list_get(verdicts, b), cand) {
let c = c + 1
}
let b = b + 1
}
if c > bestc {
let bestc = c
let best = cand
}
let a = a + 1
}
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "winner")
let kv = el_list_append(kv, best)
let kv = el_list_append(kv, "votes")
let kv = el_list_append(kv, int_to_str(bestc))
return json_build_object(kv)
}
// swarm_converge_reduce fold outputs into an accumulator (count + concat).
fn swarm_converge_reduce(results: [String]) -> String {
let n: Int = el_list_len(results)
let acc: String = ""
let i = 0
while i < n {
let out: String = json_get_raw(el_list_get(results, i), "output")
let acc = acc + out
let i = i + 1
}
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "count")
let kv = el_list_append(kv, int_to_str(n))
let kv = el_list_append(kv, "accumulated")
let kv = el_list_append(kv, acc)
return json_build_object(kv)
}
// ratio_to_permille parse a decimal ratio string ("1.0", "0.8") into an
// integer per-mille (1000, 800) so failure thresholds use exact integer math.
// (El float division is unreliable in this runtime int_to_float(n)/int_to_float(n)
// does not equal 1.0 so the swarm deliberately avoids floats.)
fn ratio_to_permille(s: String) -> Int {
if str_eq(s, "") {
return 1000
}
let parts: [String] = str_split(s, ".")
let whole: Int = str_to_int(el_list_get(parts, 0))
let permille: Int = whole * 1000
if el_list_len(parts) > 1 {
let frac_raw: String = el_list_get(parts, 1)
let frac3: String = str_slice(str_pad_right(frac_raw, 3, "0"), 0, 3)
let permille = permille + str_to_int(frac3)
}
return permille
}
// swarm_converge dispatch on strategy name.
fn swarm_converge(strategy: String, results: [String]) -> String {
if str_eq(strategy, "merge") {
return swarm_converge_merge(results)
}
if str_eq(strategy, "vote") {
return swarm_converge_vote(results)
}
if str_eq(strategy, "reduce") {
return swarm_converge_reduce(results)
}
// default: collect
return swarm_converge_collect(results)
}
// the ONLY global-engram write path (Rule 4, @manager-only)
//
// Every engram mutation flows through here and is gated by the caller's token
// capability. Only the orchestrator's token carries engram:write, so a worker
// (engram:read only) calling this is DENIED by capability before any HTTP is
// issued structurally unable to mutate global engram state, regardless of
// engram health. This is the curated-merge write: the orchestrator committing
// the geometry it approved. Workers never reach a successful branch here.
fn swarm_engram_write(token: String, corr_id: String, content: String, typ: String, importance: Float) -> String {
let deny: String = containment_guard_engram_write(token, corr_id, "engram.write")
if str_eq(deny, "") {
// authorized (orchestrator) perform the write
let res: String = op_write(content, typ, importance)
let new_id: String = json_get_string(res, "id")
let cp: String = json_set_str("{}", "node_id", new_id)
worktrack_append("swarm.committed", corr_id, "orchestrator", cp)
return res
}
// denied by capability return the rejection, no engram mutation performed
return json_set_str("{}", "denied", deny)
}
// the coordinator: fan out -> track -> converge
//
// blueprint : task blueprint name run by every worker
// knowledge_refs : JSON array of retrieval queries for CCR compilation
// inputs_json : JSON array of input items (one per worker)
// config_json : { concurrency, strategy, min_success_ratio,
// failure_action, caller_token }
//
// Returns: { corr_id, status:"completed"|"aborted", merged, report }.
fn swarm_run(blueprint: String, knowledge_refs: String, inputs_json: String, config_json: String) -> String {
let corr_id: String = "swarm-" + uuid_v4()
let caller_token: String = json_get_raw(config_json, "caller_token")
let concurrency: Int = str_to_int(json_get_string(config_json, "concurrency"))
if concurrency < 1 {
let concurrency = 4
}
let strategy: String = json_get_string(config_json, "strategy")
// Containment Rule 2: only a coordinator/absent token may open a swarm
let deny: String = containment_guard_open(caller_token, corr_id)
if str_eq(deny, "") {
// allowed proceed
let n: Int = json_array_len(inputs_json)
// swarm.created
let cp: String = json_set_str("{}", "blueprint", blueprint)
let cp2: String = json_set(cp, "input_count", int_to_str(n))
worktrack_append("swarm.created", corr_id, corr_id, cp2)
// build per-worker envelopes: worker token + CCR-compiled bounded context
let envelopes: [String] = el_list_empty()
let i = 0
while i < n {
let worker_id: String = corr_id + "/worker-" + int_to_str(i)
let input_item: String = json_array_get_string(inputs_json, i)
let wtoken: String = containment_worker_token(corr_id, worker_id)
let ctx: String = ccr_compile(blueprint, knowledge_refs, input_item, corr_id, worker_id, wtoken)
// envelope: only this worker's compiled context + its closed token
let ekv: [String] = el_list_empty()
let ekv = el_list_append(ekv, "worker_id")
let ekv = el_list_append(ekv, worker_id)
let ekv = el_list_append(ekv, "corr_id")
let ekv = el_list_append(ekv, corr_id)
let env0: String = json_build_object(ekv)
let env1: String = json_set(env0, "scope_token", wtoken)
let env2: String = json_set(env1, "ctx", ctx)
let envelopes = el_list_append(envelopes, env2)
let sp: String = json_set_str("{}", "input", input_item)
worktrack_append("worker.started", corr_id, worker_id, sp)
let i = i + 1
}
// native-thread fan-out (bounded)
let results: [String] = swarm_fanout("swarm_worker_entry", envelopes, concurrency)
// record per-worker terminal status + aggregate AFFERENT telemetry.
// Afferent counters (primitive signals travelling toward the @manager)
// are summed from the vertical result path no shared bus register,
// so the aggregation is containment-safe.
let succ = 0
let afferent = 0
let seam_mode_seen: String = "stub"
let rn: Int = el_list_len(results)
let r = 0
while r < rn {
let res: String = el_list_get(results, r)
let wid: String = json_get_string(res, "worker_id")
let st: String = json_get_string(res, "status")
let out: String = json_get_raw(res, "output")
let aff: Int = str_to_int(json_get_string(out, "afferent"))
let afferent = afferent + aff
let sm: String = json_get_string(out, "seam_mode")
if str_eq(sm, "") {
let seam_mode_seen = seam_mode_seen
} else {
let seam_mode_seen = sm
}
if str_eq(st, "completed") {
let succ = succ + 1
worktrack_append("worker.completed", corr_id, wid, json_set_str("{}", "status", "completed"))
} else {
worktrack_append("worker.failed", corr_id, wid, json_set_str("{}", "error", json_get_string(res, "error")))
}
let r = r + 1
}
// swarm.converging
let vg: String = json_set("{}", "success_count", int_to_str(succ))
worktrack_append("swarm.converging", corr_id, corr_id, vg)
// swarm.telemetry afferent counters observed by the @manager.
let tkv: [String] = el_list_empty()
let tkv = el_list_append(tkv, "seam_mode")
let tkv = el_list_append(tkv, seam_mode_seen)
let telem0: String = json_build_object(tkv)
let telem1: String = json_set_str(telem0, "afferent_think", int_to_str(afferent))
let telemetry: String = json_set_str(telem1, "results_received", int_to_str(rn))
worktrack_append("swarm.telemetry", corr_id, corr_id, telemetry)
// failure threshold (Swarm §4.3), integer per-mille math
// require succ/n >= min_success_ratio <=> succ*1000 >= permille*n
let permille: Int = ratio_to_permille(json_get_string(config_json, "min_success_ratio"))
let status: String = "completed"
if succ * 1000 < permille * n {
let status = "aborted"
}
if str_eq(status, "aborted") {
let ap: String = json_set_str("{}", "reason", "success ratio below min_success_ratio")
worktrack_append("swarm.aborted", corr_id, corr_id, ap)
let rep: String = worktrack_swarm_report(corr_id)
let ok: [String] = el_list_empty()
let ok = el_list_append(ok, "corr_id")
let ok = el_list_append(ok, corr_id)
let ok = el_list_append(ok, "status")
let ok = el_list_append(ok, "aborted")
let out0: String = json_build_object(ok)
return json_set(out0, "report", rep)
}
// converge
let merged: String = swarm_converge(strategy, results)
let dp: String = json_set_str("{}", "strategy", strategy)
worktrack_append("swarm.completed", corr_id, corr_id, dp)
// curated merge = the ONLY engram write path (Rule 4)
// With "commit":"1", the ORCHESTRATOR (its token carries engram:write)
// commits the approved merged geometry back to the engram. This is the
// single writer. Workers returned geometry; only the orchestrator writes.
let commit_id: String = ""
if str_eq(json_get_string(config_json, "commit"), "1") {
let orch_token: String = containment_coordinator_token(corr_id)
let cres: String = swarm_engram_write(orch_token, corr_id, "swarm-merge " + corr_id + " :: " + merged, "memory", 0.5)
let commit_id = json_get_string(cres, "id")
}
let rep2: String = worktrack_swarm_report(corr_id)
let ok2: [String] = el_list_empty()
let ok2 = el_list_append(ok2, "corr_id")
let ok2 = el_list_append(ok2, corr_id)
let ok2 = el_list_append(ok2, "status")
let ok2 = el_list_append(ok2, "completed")
let out1: String = json_build_object(ok2)
let out2: String = json_set(out1, "report", rep2)
let out3: String = json_set(out2, "merged", merged)
let out4: String = json_set(out3, "telemetry", telemetry)
return json_set_str(out4, "committed_node", commit_id)
}
// denied: caller was a worker trying to open a swarm (Rule 2)
let dkv: [String] = el_list_empty()
let dkv = el_list_append(dkv, "corr_id")
let dkv = el_list_append(dkv, corr_id)
let dkv = el_list_append(dkv, "status")
let dkv = el_list_append(dkv, "denied")
let dkv = el_list_append(dkv, "error")
let dkv = el_list_append(dkv, deny)
return json_build_object(dkv)
}
+88
View File
@@ -0,0 +1,88 @@
// harness_local_swarm.el LOCAL-SWARM INTEGRATION HARNESS.
//
// Proves the FULL local-swarm mechanics end-to-end, TODAY, on the isolated
// engram clone (:8901), with the primitive seam pointed at the hermetic stub.
// The moment the api-reshape agent lands the decorated primitives on the
// dharma bus, binding is ONE flip (primitive_binding.el) + SWARM_PRIMITIVE_SEAM=
// decorated this same harness then runs the bound path with no other change.
//
// The @manager (the coordinator) fans out N native El worker threads at real
// concurrency, each given a CCR-scoped engram slice, each invoking the primitive
// seam (think over its slice), enforces all three containment rules, converges
// (vote AND reduce), work-tracks durably, and observes afferent telemetry.
//
// Run with the sandbox env sourced (ENGRAM_URL=:8901) to also exercise CCR
// retrieval against the real (isolated) mind; runs fully without it too.
fn ok(label: String, cond: Bool, fails: Int) -> Int {
if cond { print(" ok " + label); return fails }
print(" FAIL " + label); return fails + 1
}
fn main() -> Int {
let fails = 0
print("== LOCAL-SWARM INTEGRATION HARNESS (seam=" + seam_mode() + ") ==")
// 8 independent slices, real concurrency of 4 (2 waves of native pthreads).
let inputs: String = "[\"billing\",\"payments\",\"ledger\",\"invoicing\",\"tax\",\"payroll\",\"audit\",\"fx\"]"
let refs: String = "[\"Volatility-Based Decomposition\"]"
// A) fan-out / converge at real concurrency (reduce)
let cfg_r: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\"}"
let rr: String = swarm_run("analyze_item", refs, inputs, cfg_r)
let fails = ok("swarm completed at concurrency=4 over 8 native-thread workers", str_eq(json_get_string(rr, "status"), "completed"), fails)
let corr: String = json_get_string(rr, "corr_id")
let merged_r: String = json_get_raw(rr, "merged")
let fails = ok("reduce converged all 8 worker outputs", str_to_int(json_get_string(merged_r, "count")) == 8, fails)
// B) afferent telemetry observed by the @manager
let telem: String = json_get_raw(rr, "telemetry")
let aff: Int = str_to_int(json_get_string(telem, "afferent_think"))
let seen_mode: String = json_get_string(telem, "seam_mode")
let fails = ok("afferent think-signals counted = 8 (one per worker)", aff == 8, fails)
let fails = ok("telemetry records the active seam mode", str_eq(seen_mode, seam_mode()), fails)
let telem_recs: Int = worktrack_count_kind(corr, "swarm.telemetry")
let fails = ok("telemetry durably journalled", telem_recs == 1, fails)
// C) CCR scoping + non-leak per worker
let wt: String = containment_worker_token(corr, corr + "/worker-3")
let ctx3: String = ccr_compile("analyze_item", refs, "invoicing", corr, corr + "/worker-3", wt)
let fails = ok("CCR context bounded within token budget", ccr_within_budget(ctx3), fails)
let fails = ok("CCR context carries THIS slice", str_eq(json_get_string(ctx3, "input"), "invoicing"), fails)
let leaks: Bool = str_contains(ctx3, "payroll") || str_contains(ctx3, "audit")
let fails = ok("CCR context does NOT leak sibling slices (security boundary)", !leaks, fails)
// D) all three containment rules
let deny: String = containment_check_open(wt)
let fails = ok("Rule 2: worker token may not OPEN a swarm", !str_eq(deny, ""), fails)
let denyj: String = containment_check_join(wt, "other-swarm")
let fails = ok("Rule 1: worker token may not JOIN another swarm", !str_eq(denyj, ""), fails)
let lat: String = containment_check_lateral(wt, "sibling-9")
let fails = ok("Rule 3: worker->worker lateral edge rejected", !str_eq(lat, ""), fails)
let ver: String = containment_check_lateral(wt, "")
let fails = ok("Rule 3: worker->manager vertical edge allowed", str_eq(ver, ""), fails)
// enforced live: a worker-token caller is denied opening a real swarm
let wcfg: String = json_set(cfg_r, "caller_token", wt)
let denied: String = swarm_run("analyze_item", refs, inputs, wcfg)
let fails = ok("Rule 2 enforced live: worker-caller swarm denied", str_eq(json_get_string(denied, "status"), "denied"), fails)
// E) vote convergence strategy at concurrency
let cfg_v: String = "{\"concurrency\":\"8\",\"strategy\":\"vote\",\"min_success_ratio\":\"1.0\"}"
let rv: String = swarm_run("classify", refs, inputs, cfg_v)
let winner: String = json_get_string(json_get_raw(rv, "merged"), "winner")
// billing/payments/ledger/invoicing/payroll/audit = long(>4); tax/fx = short -> long wins
let fails = ok("vote converged (winner=long)", str_eq(winner, "long"), fails)
// F) durable, inspectable work-tracking
let started: Int = worktrack_count_kind(corr, "worker.started")
let completed: Int = worktrack_count_kind(corr, "worker.completed")
let fails = ok("work-tracking journal: 8 started + 8 completed", (started == 8) && (completed == 8), fails)
print("")
if fails == 0 {
print("HARNESS GREEN — full local-swarm mechanics proven with seam=" + seam_mode())
return 0
}
print("HARNESS FAIL (" + int_to_str(fails) + ")")
return 1
}
+119
View File
@@ -0,0 +1,119 @@
// harness_real_cognition.el the LOCAL SWARM running REAL cognition.
//
// Run with: SWARM_PRIMITIVE_SEAM=decorated + the sandbox env sourced
// (ENGRAM_URL=:8901). Each worker's `think` is BOUND to the reshape's proven
// op_think (GET /api/think) over its NODE-ID anchor real 768-dim gradients from
// the live (isolated) geometry, not the stub. The @manager fans out N native-El
// worker threads at real concurrency, converges (reduce + vote) over the real
// cognition, enforces all three containment rules, observes afferent telemetry,
// and work-tracks durably.
//
// Anchors are real self-neighbourhood node ids on the :8901 clone (free-text
// anchors return "geometry unavailable", so these must be node ids).
fn ok(label: String, cond: Bool, fails: Int) -> Int {
if cond { print(" ok " + label); return fails }
print(" FAIL " + label); return fails + 1
}
fn main() -> Int {
let fails = 0
print("== REAL-COGNITION LOCAL SWARM (seam=" + seam_mode() + ", engram=" + env("ENGRAM_URL") + ") ==")
// 0) direct proof the bound primitive returns REAL cognition
let g: String = op_think("self", "plan")
let dim: Int = json_get_int(g, "dim")
let nsup: Int = json_get_int(g, "n_support")
let fails = ok("bound op_think returns a real 768-dim gradient", dim == 768, fails)
let fails = ok("real gradient has support (n_support>0)", nsup > 0, fails)
let gfree: String = op_think("this-is-free-text-not-a-node", "reason")
let fails = ok("free-text anchor correctly refused (geometry unavailable)", str_contains(gfree, "geometry unavailable"), fails)
// the input set: 8 real NODE-ID anchors from self's neighbourhood
let anchors: String = "[\"a1000001-0000-0000-0000-000000000001\",\"5f011441-fa43-4fe7-a9c0-c78a584ef11d\",\"kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6\",\"76d7fd0b-0672-4511-a2f5-a095cf9c60ae\",\"7027e302-593f-441d-8fd6-9c400c163108\",\"2a730b18-6566-46ee-a21e-4f4dd0380908\",\"46b0e4dd-2c19-48d2-bcbc-19f61d6c79ae\",\"9162cde8-8739-4f00-bfc9-2850ed612e50\"]"
let refs: String = "[\"self\"]"
// A) fan-out real cognition at concurrency, converge with REDUCE
let cfg_r: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\"}"
let rr: String = swarm_run("cognize", refs, anchors, cfg_r)
let fails = ok("swarm completed: 8 workers each a real think, concurrency=4", str_eq(json_get_string(rr, "status"), "completed"), fails)
let corr: String = json_get_string(rr, "corr_id")
let merged_r: String = json_get_raw(rr, "merged")
let fails = ok("reduce converged all 8 real-cognition outputs", str_to_int(json_get_string(merged_r, "count")) == 8, fails)
let acc: String = json_get_string(merged_r, "accumulated")
let fails = ok("converged output carries real gradient support (n_support)", str_contains(acc, "n_support"), fails)
// B) afferent telemetry: 8 real think-signals, decorated seam
let telem: String = json_get_raw(rr, "telemetry")
let aff: Int = str_to_int(json_get_string(telem, "afferent_think"))
let fails = ok("afferent counters = 8 real think invocations", aff == 8, fails)
let fails = ok("telemetry records seam_mode=decorated", str_eq(json_get_string(telem, "seam_mode"), "decorated"), fails)
let fails = ok("telemetry durably journalled", worktrack_count_kind(corr, "swarm.telemetry") == 1, fails)
// C) converge with VOTE over real cognition
let cfg_v: String = "{\"concurrency\":\"8\",\"strategy\":\"vote\",\"min_success_ratio\":\"1.0\"}"
let rv: String = swarm_run("cognize", refs, anchors, cfg_v)
let winner: String = json_get_string(json_get_raw(rv, "merged"), "winner")
let fails = ok("vote converged over real cognition (winner=" + winner + ")", !str_eq(winner, ""), fails)
// D) all three containment rules still enforced
let wt: String = containment_worker_token(corr, corr + "/worker-2")
let fails = ok("Rule 2: worker may not open a swarm", !str_eq(containment_check_open(wt), ""), fails)
let fails = ok("Rule 1: worker may not join another swarm", !str_eq(containment_check_join(wt, "s2"), ""), fails)
let fails = ok("Rule 3: worker->worker lateral edge rejected", !str_eq(containment_check_lateral(wt, "sib"), ""), fails)
let wcfg: String = json_set(cfg_r, "caller_token", wt)
let denied: String = swarm_run("cognize", refs, anchors, wcfg)
let fails = ok("Rule 2 enforced LIVE: worker-caller swarm denied", str_eq(json_get_string(denied, "status"), "denied"), fails)
// E) CCR scoping + non-leak over node-id anchors
let ctx: String = ccr_compile("cognize", refs, "a1000001-0000-0000-0000-000000000001", corr, corr + "/worker-0", wt)
let fails = ok("CCR context bounded within budget", ccr_within_budget(ctx), fails)
let leaks: Bool = str_contains(ctx, "9162cde8")
let fails = ok("CCR context does NOT leak sibling anchors", !leaks, fails)
// F) durable work-tracking
let started: Int = worktrack_count_kind(corr, "worker.started")
let completed: Int = worktrack_count_kind(corr, "worker.completed")
let fails = ok("work-tracking: 8 started + 8 completed", (started == 8) && (completed == 8), fails)
// G) RULE 4 engram-write is @manager-ONLY (authority gate)
// A worker token (engram:read only) is STRUCTURALLY denied any engram write.
let worker_tok: String = containment_worker_token(corr, corr + "/worker-1")
let orch_tok: String = containment_coordinator_token(corr)
let fails = ok("worker token carries engram:read", containment_has_cap(worker_tok, "engram:read"), fails)
let fails = ok("worker token does NOT carry engram:write", !containment_has_cap(worker_tok, "engram:write"), fails)
let fails = ok("orchestrator token carries engram:write", containment_has_cap(orch_tok, "engram:write"), fails)
// a worker attempting an engram write is DENIED BY CAPABILITY (no HTTP issued)
let wdeny: String = swarm_engram_write(worker_tok, corr, "worker tries to mutate global state", "memory", 0.5)
let denied_reason: String = json_get_string(wdeny, "denied")
let fails = ok("worker engram-write DENIED by capability (Rule 4)", str_contains(denied_reason, "rule 4"), fails)
let fails = ok("denied worker write performed NO engram mutation (no node id)", str_eq(json_get_string(wdeny, "id"), ""), fails)
let fails = ok("Rule-4 violation journalled", worktrack_count_kind(corr, "containment.violation") >= 1, fails)
// the orchestrator passes the capability gate (sole authorized writer)
let odeny: String = containment_check_engram_write(orch_tok, "engram.write")
let fails = ok("orchestrator PASSES the engram-write capability gate (sole writer)", str_eq(odeny, ""), fails)
// H) curated merge = the only write path (orchestrator commits)
// The AUTHORITY gate above is already proven (worker denied, orchestrator
// authorized) WITHOUT issuing a write. The actual persisting commit exercises
// the engram write path, which needs the gate-1 write-healthy clone so it
// runs only under SWARM_WRITE_HEALTHY=1 (else it would hit the known daemon
// write-crash). Authority != health: the gate holds either way.
if str_eq(env("SWARM_WRITE_HEALTHY"), "1") {
let cfg_commit: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\",\"commit\":\"1\"}"
let rc: String = swarm_run("cognize", refs, anchors, cfg_commit)
let committed: String = json_get_string(rc, "committed_node")
let fails2: Int = ok("orchestrator (sole writer) committed the merge to the engram", !str_eq(committed, ""), fails)
let fails = fails2
} else {
print(" note curated-merge commit deferred to the gate-1 write-healthy clone (set SWARM_WRITE_HEALTHY=1); authority gate already proven above")
}
print("")
if fails == 0 {
print("REAL-COGNITION SWARM GREEN — Neuron thinking in parallel over its own geometry.")
return 0
}
print("REAL-COGNITION SWARM FAIL (" + int_to_str(fails) + ")")
return 1
}
+45
View File
@@ -0,0 +1,45 @@
// integ_engram.el integration proof against a LIVE (isolated) engram.
//
// Run with the sandbox env sourced (ENGRAM_URL=http://127.0.0.1:8901,
// ENGRAM_API_KEY=sbx-dev-swarm-ccr). Proves:
// (a) CCR retrieval pulls REAL content from the mind over HTTP;
// (b) a full swarm runs and converges against the live mind;
// (c) work-tracking mirrors records into the engram as SwarmTrack nodes.
fn main() -> Int {
let url: String = env("ENGRAM_URL")
if str_eq(url, "") {
print("SKIP integ_engram (ENGRAM_URL not set)")
return 0
}
// (a) CCR compiles a bounded context whose retrieval hit the real mind.
let refs: String = "[\"Volatility-Based Decomposition\",\"Swarm Architecture containment\"]"
let wt: String = containment_worker_token("integ", "integ/w0")
let ctx: String = ccr_compile("analyze_item", refs, "decompose the billing module", "integ", "integ/w0", wt)
let knowledge: String = json_get_string(ctx, "knowledge")
let pulled_real: Bool = str_contains(knowledge, "olatility") || str_contains(knowledge, "Anderson") || str_contains(knowledge, "VBD")
if pulled_real {
print(" ok CCR retrieval pulled real mind content (" + int_to_str(str_len(knowledge)) + " bytes, bounded)")
} else {
print(" FAIL CCR retrieval returned no mind content")
}
let bounded: Bool = ccr_within_budget(ctx)
if bounded { print(" ok compiled context stayed within budget") } else { print(" FAIL context over budget") }
// (b) a real swarm over the live mind.
let inputs: String = "[\"billing\",\"payments\",\"ledger\"]"
let cfg: String = "{\"concurrency\":\"3\",\"strategy\":\"collect\",\"min_success_ratio\":\"1.0\"}"
let res: String = swarm_run("analyze_item", refs, inputs, cfg)
let status: String = json_get_string(res, "status")
if str_eq(status, "completed") { print(" ok swarm completed against live engram") } else { print(" FAIL swarm status=" + status) }
let corr: String = json_get_string(res, "corr_id")
// (c) work-tracking mirrored into the mind: search for this swarm's records.
let hits: String = primitive_attend(corr, 5)
let mirrored: Bool = str_contains(hits, "swarm-track") || str_contains(hits, corr)
if mirrored { print(" ok work-tracking mirrored into the engram (queryable)") } else { print(" note mirror not yet visible to search (async index)") }
print("DONE integ_engram corr=" + corr)
return 0
}
+55
View File
@@ -0,0 +1,55 @@
// test_convergence.el convergence strategies + failure threshold / abort.
fn assert_true(label: String, cond: Bool, fails: Int) -> Int {
if cond { print(" ok " + label); return fails }
print(" FAIL " + label); return fails + 1
}
fn main() -> Int {
let fails = 0
let refs: String = "[]"
// vote: classify 5 inputs; 3 "long" (>4 chars) vs 2 "short" -> winner long ──
let inputs: String = "[\"alpha\",\"bravo\",\"hi\",\"charlie\",\"ok\"]"
let cfg_v: String = "{\"concurrency\":\"3\",\"strategy\":\"vote\",\"min_success_ratio\":\"1.0\"}"
let rv: String = swarm_run("classify", refs, inputs, cfg_v)
let merged_v: String = json_get_raw(rv, "merged")
let winner: String = json_get_string(merged_v, "winner")
let votes: Int = str_to_int(json_get_string(merged_v, "votes"))
let fails = assert_true("vote winner = long", str_eq(winner, "long"), fails)
let fails = assert_true("vote count = 3", votes == 3, fails)
// merge: outputs joined
let cfg_m: String = "{\"concurrency\":\"2\",\"strategy\":\"merge\",\"min_success_ratio\":\"1.0\"}"
let rm: String = swarm_run("analyze_item", refs, "[\"a\",\"b\",\"c\"]", cfg_m)
let merged_m: String = json_get_raw(rm, "merged")
let joined: String = json_get_string(merged_m, "merged")
let fails = assert_true("merge produced a joined string", str_contains(joined, "|"), fails)
// reduce: count accumulates
let cfg_r: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\"}"
let rr: String = swarm_run("analyze_item", refs, "[\"a\",\"b\",\"c\",\"d\"]", cfg_r)
let merged_r: String = json_get_raw(rr, "merged")
let rcount: Int = str_to_int(json_get_string(merged_r, "count"))
let fails = assert_true("reduce count = 4", rcount == 4, fails)
// failure threshold: 2 of 5 fail (x-prefixed); ratio 3/5=0.6 < 0.8 -> aborted ──
let fin: String = "[\"a\",\"xb\",\"c\",\"xd\",\"e\"]"
let cfg_f: String = "{\"concurrency\":\"5\",\"strategy\":\"collect\",\"min_success_ratio\":\"0.8\"}"
let rf: String = swarm_run("faildemo", refs, fin, cfg_f)
let fstatus: String = json_get_string(rf, "status")
let fails = assert_true("swarm aborted below min_success_ratio (0.6<0.8)", str_eq(fstatus, "aborted"), fails)
let corr_f: String = json_get_string(rf, "corr_id")
let failed_n: Int = worktrack_count_kind(corr_f, "worker.failed")
let aborted_n: Int = worktrack_count_kind(corr_f, "swarm.aborted")
let fails = assert_true("tracked 2 worker.failed", failed_n == 2, fails)
let fails = assert_true("tracked swarm.aborted", aborted_n == 1, fails)
// same failures tolerated when min_success_ratio=0.5 (0.6>=0.5) -> completed
let cfg_ok: String = "{\"concurrency\":\"5\",\"strategy\":\"collect\",\"min_success_ratio\":\"0.5\"}"
let rok: String = swarm_run("faildemo", refs, fin, cfg_ok)
let fails = assert_true("swarm completes when failures within tolerance", str_eq(json_get_string(rok, "status"), "completed"), fails)
if fails == 0 { print("PASS test_convergence"); return 0 }
print("FAIL test_convergence (" + int_to_str(fails) + ")"); return 1
}
+75
View File
@@ -0,0 +1,75 @@
// test_swarm.el end-to-end proof of the swarm capability on native El threads.
//
// Proves: native-thread fan-out/converge, bounded concurrency, per-worker CCR
// bounded context (with the security-boundary property), containment Rule 2
// enforcement, and durable work-tracking.
fn assert_true(label: String, cond: Bool, fails: Int) -> Int {
if cond {
print(" ok " + label)
return fails
}
print(" FAIL " + label)
return fails + 1
}
fn main() -> Int {
let fails = 0
// 1) fan-out / converge (collect) over native threads
let inputs: String = "[\"alpha\",\"bravo\",\"charlie\",\"delta\",\"echo\"]"
let refs: String = "[]"
let cfg: String = "{\"concurrency\":\"2\",\"strategy\":\"collect\",\"min_success_ratio\":\"1.0\"}"
let res: String = swarm_run("analyze_item", refs, inputs, cfg)
let status: String = json_get_string(res, "status")
let fails = assert_true("swarm completed", str_eq(status, "completed"), fails)
let merged: String = json_get_raw(res, "merged")
let count: Int = json_array_len(merged)
let fails = assert_true("collect returned 5 results (bounded concurrency=2)", count == 5, fails)
// 2) work-tracking is durable + complete
let corr: String = json_get_string(res, "corr_id")
let started: Int = worktrack_count_kind(corr, "worker.started")
let completed: Int = worktrack_count_kind(corr, "worker.completed")
let created: Int = worktrack_count_kind(corr, "swarm.created")
let done: Int = worktrack_count_kind(corr, "swarm.completed")
let fails = assert_true("tracked 5 worker.started", started == 5, fails)
let fails = assert_true("tracked 5 worker.completed", completed == 5, fails)
let fails = assert_true("tracked swarm.created + swarm.completed", (created == 1) && (done == 1), fails)
// 3) CCR: bounded, minimal, non-leaking per-worker context
let wtoken: String = containment_worker_token(corr, corr + "/worker-0")
let ctx: String = ccr_compile("analyze_item", refs, "alpha", corr, corr + "/worker-0", wtoken)
let in_budget: Bool = ccr_within_budget(ctx)
let fails = assert_true("CCR context within token budget", in_budget, fails)
let this_input: String = json_get_string(ctx, "input")
let fails = assert_true("CCR context contains THIS worker's input", str_eq(this_input, "alpha"), fails)
// security boundary: a worker's compiled context must not carry a sibling input
let leaks_sibling: Bool = str_contains(ctx, "charlie")
let fails = assert_true("CCR context does NOT leak sibling inputs", !leaks_sibling, fails)
// 4) containment Rule 2: a worker may not open a swarm
let worker_caller_cfg: String = json_set(cfg, "caller_token", wtoken)
let denied: String = swarm_run("analyze_item", refs, inputs, worker_caller_cfg)
let dstatus: String = json_get_string(denied, "status")
let fails = assert_true("worker-token caller denied opening a swarm (Rule 2)", str_eq(dstatus, "denied"), fails)
// coordinator token IS allowed
let coord: String = containment_coordinator_token("some-corr")
let allow_reason: String = containment_check_open(coord)
let fails = assert_true("coordinator token allowed to open a swarm", str_eq(allow_reason, ""), fails)
// 5) containment Rule 3: no lateral worker->worker edge
let lateral: String = containment_check_lateral(wtoken, "some-sibling")
let fails = assert_true("lateral worker->worker edge rejected (Rule 3)", !str_eq(lateral, ""), fails)
let vertical: String = containment_check_lateral(wtoken, "")
let fails = assert_true("vertical worker->coordinator edge allowed", str_eq(vertical, ""), fails)
if fails == 0 {
print("PASS test_swarm")
return 0
}
print("FAIL test_swarm (" + int_to_str(fails) + " failures)")
return 1
}
+38
View File
@@ -0,0 +1,38 @@
// test_worktrack.el durability + inspectability of the work-tracking journal.
fn main() -> Int {
let corr: String = "test-" + uuid_v4()
// record a swarm lifecycle
let p1: String = json_set("{}", "input_count", "3")
worktrack_append("swarm.created", corr, "swarm-1", p1)
worktrack_append("worker.started", corr, "worker-001", "{}")
worktrack_append("worker.started", corr, "worker-002", "{}")
worktrack_append("worker.completed", corr, "worker-001", "{}")
worktrack_append("worker.failed", corr, "worker-002", "{}")
worktrack_append("swarm.completed", corr, "swarm-1", "{}")
// inspect: reconstruct the report from the durable journal
let report: String = worktrack_swarm_report(corr)
print("report=" + report)
let recs_n: Int = el_list_len(worktrack_records(corr))
print("records=" + int_to_str(recs_n))
let state: String = json_get_string(report, "state")
let completed: Int = str_to_int(json_get_string(report, "workers_completed"))
let failed: Int = str_to_int(json_get_string(report, "workers_failed"))
if str_eq(state, "completed") {
if completed == 1 {
if failed == 1 {
if recs_n == 6 {
print("PASS worktrack")
return 0
}
}
}
}
print("FAIL worktrack")
return 1
}
+217
View File
@@ -0,0 +1,217 @@
// worktrack.el full work-tracking for the swarm.
//
// "Intent all the way up, orchestrator at the top." Every unit of parallel
// work a swarm fans out is recorded here: the swarm itself, each worker, its
// status, its result summary, the convergence, and the final merged output
// all threaded by a single correlation ID so the entire execution graph can be
// reconstructed and audited (Swarm Architecture §6.1).
//
// DURABILITY. Records are appended to a JSON-lines journal on disk. The journal
// is append-only and single-writer: only the coordinator (the main thread, before
// and after each fan-out and during convergence) writes to it. Workers never
// touch it they return structured results and the coordinator records them.
// This is deliberate: it makes the tracking store race-free and, not
// coincidentally, enforces Swarm containment rule 3 (no lateral worker state).
//
// INSPECTABILITY. The journal is plain JSONL greppable, tailable, replayable.
// worktrack_read() loads it back; worktrack_swarm_report() reconstructs a
// swarm's full record from its correlation ID.
//
// ENGRAM MIRROR (optional). When ENGRAM_URL is set, each record is also mirrored
// into the engram as a node (POST /api/node) tagged with the correlation ID, so
// the swarm's execution becomes part of the durable mind, queryable by memory.
//
// Depends on: el_runtime.c builtins (fs_*, http_post, env, json_*, uuid_v4,
// now_millis, str_*). No El-module concat dependencies of its own.
// JSON helper
// json_set inserts its value as a RAW JSON fragment (objects/arrays/numbers).
// json_set_str sets a plain STRING value, correctly quoted and escaped. Use
// json_set for nested JSON, json_set_str for strings.
fn json_set_str(j: String, key: String, val: String) -> String {
return json_set(j, key, "\"" + json_escape_string(val) + "\"")
}
// Journal location
// worktrack_dir directory holding the swarm journals.
// Override with SWARM_TRACK_DIR; defaults to ./.swarm-track (relative to CWD).
fn worktrack_dir() -> String {
let d: String = env("SWARM_TRACK_DIR")
if str_eq(d, "") {
return ".swarm-track"
}
return d
}
// worktrack_journal_path the JSONL journal file for one correlation ID.
fn worktrack_journal_path(corr_id: String) -> String {
return worktrack_dir() + "/" + corr_id + ".jsonl"
}
// worktrack_init ensure the journal directory exists. Idempotent.
fn worktrack_init() -> Bool {
let d: String = worktrack_dir()
if fs_exists(d) {
return true
}
return fs_mkdir(d)
}
// Record construction
// worktrack_record build one journal record as a JSON object string.
// kind: the record kind (swarm.created, worker.started, ...)
// corr_id: the swarm correlation ID (links every record)
// subject: the entity the record is about (swarm id, worker id, "")
// payload: a JSON object string with kind-specific fields
fn worktrack_record(kind: String, corr_id: String, subject: String, payload: String) -> String {
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "kind")
let kv = el_list_append(kv, kind)
let kv = el_list_append(kv, "corr_id")
let kv = el_list_append(kv, corr_id)
let kv = el_list_append(kv, "subject")
let kv = el_list_append(kv, subject)
let kv = el_list_append(kv, "ts_ms")
let kv = el_list_append(kv, int_to_str(now_millis()))
let rec: String = json_build_object(kv)
// Attach the payload as a nested raw JSON field.
let rec2: String = json_set(rec, "data", payload)
return rec2
}
// Journal append (single-writer, durable)
// worktrack_append append one record to the correlation journal (durable),
// and mirror it to the engram if ENGRAM_URL is configured. Returns the record.
//
// fs_write here is used in append semantics: we read-modify-write the file. The
// coordinator is the only writer, so this is safe and race-free.
fn worktrack_append(kind: String, corr_id: String, subject: String, payload: String) -> String {
worktrack_init()
let rec: String = worktrack_record(kind, corr_id, subject, payload)
let path: String = worktrack_journal_path(corr_id)
let prior: String = ""
if fs_exists(path) {
let prior = fs_read(path)
}
let next: String = prior + rec + "\n"
fs_write(path, next)
worktrack_mirror_engram(rec, corr_id, kind, subject)
return rec
}
// worktrack_mirror_engram best-effort mirror of a record into the engram.
// No-op unless ENGRAM_URL is set. Failures are swallowed (tracking must not
// depend on the mind being reachable).
fn worktrack_mirror_engram(rec: String, corr_id: String, kind: String, subject: String) -> Bool {
// Opt-in: the durable substrate is the JSONL journal (always written). The
// engram mirror is an additional convenience, enabled with SWARM_MIRROR=1,
// so a swarm never depends on or loads the mind just to track its work.
if str_eq(env("SWARM_MIRROR"), "1") {
// enabled fall through to the mirror POST
let _go: Int = 1
} else {
return false
}
let url: String = env("ENGRAM_URL")
if str_eq(url, "") {
return false
}
let content: String = "swarm-track " + kind + " " + subject + " :: " + rec
let body_kv: [String] = el_list_empty()
let body_kv = el_list_append(body_kv, "content")
let body_kv = el_list_append(body_kv, content)
let body_kv = el_list_append(body_kv, "node_type")
let body_kv = el_list_append(body_kv, "SwarmTrack")
let body_kv = el_list_append(body_kv, "salience")
let body_kv = el_list_append(body_kv, "0.5")
let body: String = json_build_object(body_kv)
let key: String = env("ENGRAM_API_KEY")
let body2: String = json_set_str(body, "_auth", key)
let resp: String = http_post(url + "/api/nodes", body2)
return true
}
// Read / inspect
// worktrack_read read the raw JSONL journal for a correlation ID.
fn worktrack_read(corr_id: String) -> String {
let path: String = worktrack_journal_path(corr_id)
if fs_exists(path) {
return fs_read(path)
}
return ""
}
// worktrack_records the journal as a [String] of record JSON objects, in order.
fn worktrack_records(corr_id: String) -> [String] {
let raw: String = worktrack_read(corr_id)
let out: [String] = el_list_empty()
if str_eq(raw, "") {
return out
}
let lines: [String] = str_split_lines(raw)
let n: Int = el_list_len(lines)
let i = 0
while i < n {
let ln: String = el_list_get(lines, i)
if str_eq(ln, "") {
let i = i + 1
} else {
let out = el_list_append(out, ln)
let i = i + 1
}
}
return out
}
// worktrack_count_kind how many records of a given kind exist for a swarm.
// Powers assertions and live status ("how many workers completed").
fn worktrack_count_kind(corr_id: String, kind: String) -> Int {
let recs: [String] = worktrack_records(corr_id)
let n: Int = el_list_len(recs)
let c = 0
let i = 0
while i < n {
let r: String = el_list_get(recs, i)
let k: String = json_get_string(r, "kind")
if str_eq(k, kind) {
let c = c + 1
}
let i = i + 1
}
return c
}
// worktrack_swarm_report reconstruct a compact status report for a swarm from
// its journal: counts of started/completed/failed workers and terminal state.
// Inspectable, durable, derived purely from the append-only record.
fn worktrack_swarm_report(corr_id: String) -> String {
let started: Int = worktrack_count_kind(corr_id, "worker.started")
let completed: Int = worktrack_count_kind(corr_id, "worker.completed")
let failed: Int = worktrack_count_kind(corr_id, "worker.failed")
let done: Int = worktrack_count_kind(corr_id, "swarm.completed")
let aborted: Int = worktrack_count_kind(corr_id, "swarm.aborted")
let state: String = "running"
if aborted > 0 {
let state = "aborted"
} else {
if done > 0 {
let state = "completed"
}
}
let kv: [String] = el_list_empty()
let kv = el_list_append(kv, "corr_id")
let kv = el_list_append(kv, corr_id)
let kv = el_list_append(kv, "state")
let kv = el_list_append(kv, state)
let kv = el_list_append(kv, "workers_started")
let kv = el_list_append(kv, int_to_str(started))
let kv = el_list_append(kv, "workers_completed")
let kv = el_list_append(kv, int_to_str(completed))
let kv = el_list_append(kv, "workers_failed")
let kv = el_list_append(kv, int_to_str(failed))
return json_build_object(kv)
}
+91
View File
@@ -0,0 +1,91 @@
// fitprobe.el controlled growth-curve specimens for validating the complexity fitter.
//
// Three deliberately-shaped workloads. None depends on a real defect existing,
// which is the point: the fitter must be provable against KNOWN curves.
//
// linear one allocation per item. count O(n), bytes O(n), time O(n)
// accum rebuilds its accumulator. count O(n), bytes O(n^2), time O(n^2)
// compute nested arithmetic, no alloc. count O(1), bytes O(1), time O(n^2)
//
// `compute` is the specimen that matters. It is the shape of el #132
// (strlen-per-character inside str_char_code): pure CPU, zero allocation.
// An allocation-only gate is structurally blind to it.
//
// No imports uses runtime builtins directly so nothing collides.
fn work_linear(n: Int) -> Int {
let parts: [String] = native_list_empty()
let i: Int = 0
while i < n {
let parts = native_list_append(parts, int_to_str(i))
let i = i + 1
}
return native_list_len(parts)
}
fn work_accum(n: Int) -> Int {
let acc: String = ""
let i: Int = 0
while i < n {
let acc = acc + "x"
let i = i + 1
}
return str_len(acc)
}
fn work_compute(n: Int) -> Int {
// str_char_code is an opaque external call, so the C optimiser cannot
// reduce this nest to a closed form the way it does with `total + 1`.
// This is the exact shape of el #132: n scans over n characters, pure
// CPU, ZERO allocation.
let s: String = "abcdefghij"
let total: Int = 0
let i: Int = 0
while i < n {
let j: Int = 0
while j < n {
let total = total + str_char_code(s, 0)
let j = j + 1
}
let i = i + 1
}
return total
}
fn run_one(mode: String, n: Int) {
let c0: Int = el_alloc_count()
let b0: Int = el_alloc_bytes()
let t0: Int = el_now_instant()
let r: Int = 0
if str_eq(mode, "linear") { let r = work_linear(n) }
if str_eq(mode, "accum") { let r = work_accum(n) }
if str_eq(mode, "compute") { let r = work_compute(n) }
let t1: Int = el_now_instant()
let c1: Int = el_alloc_count()
let b1: Int = el_alloc_bytes()
println(mode + "\t" + int_to_str(n)
+ "\t" + int_to_str(c1 - c0)
+ "\t" + int_to_str(b1 - b0)
+ "\t" + int_to_str((t1 - t0) / 1000)
+ "\t" + int_to_str(r))
return
}
fn sweep(mode: String) {
run_one(mode, 200)
run_one(mode, 400)
run_one(mode, 800)
run_one(mode, 1600)
return
}
fn main() -> Int {
println("mode\tn\tallocs\tbytes\tusec\tsink")
sweep("linear")
sweep("accum")
sweep("compute")
return 0
}
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// tests/native/test_compiler.el comprehensive tests for the El compiler pipeline.
//
// Tests the lexer (lexer.el), parser (parser.el), and codegen (codegen.el)
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_codegen_js.el - basic tests for JS codegen features.
//
// These tests verify that core El language features produce correct values
+111
View File
@@ -0,0 +1,111 @@
import "../../runtime/eltest.el"
import "../../runtime/elbench.el"
// test_elbench.el proves the growth-curve classifier against KNOWN curves.
//
// Every series below is real measured data from lang/tests/bench/fitprobe.el
// on a geometric sweep n = 200/400/800/1600. The classifier must be provable
// without depending on a live defect existing, which is the whole point of
// keeping controlled specimens.
fn _s4(a: Int, b: Int, c: Int, d: Int) -> [Int] {
let l: [Int] = native_list_empty()
let l = native_list_append(l, a)
let l = native_list_append(l, b)
let l = native_list_append(l, c)
let l = native_list_append(l, d)
return l
}
test "classifies a linear allocation series as O(n)" {
// fitprobe `linear`, allocation count
let v = _s4(208, 409, 810, 1611)
assert elb_measured_curve(v, 10) == 2, "linear allocs should classify O(n)"
}
test "classifies a linear byte series as O(n)" {
// fitprobe `linear`, allocation bytes
let v = _s4(4786, 9682, 19474, 39658)
assert elb_measured_curve(v, 10) == 2, "linear bytes should classify O(n)"
}
test "classifies a quadratic byte series as O(n^2)" {
// fitprobe `accum`, allocation bytes -- the accumulator-rebuild shape
let v = _s4(20300, 80600, 321200, 1282400)
assert elb_measured_curve(v, 10) == 4, "accum bytes should classify O(n^2)"
}
test "accumulator count is linear -- proves count alone misses it" {
// Same run as above. The COUNT is exactly linear while bytes are
// quadratic. A count-only gate passes this defect clean.
let v = _s4(200, 400, 800, 1600)
assert elb_measured_curve(v, 10) == 2, "accum count classifies O(n)"
assert elb_gate(v, 2, 10) == 0, "count-only gate PASSES the quadratic"
}
test "classifies a quadratic time series as O(n^2)" {
// fitprobe `compute` -- el #132's shape: n scans over n characters
let v = _s4(67, 205, 818, 3268)
assert elb_measured_curve(v, 10) == 4, "compute time should classify O(n^2)"
}
test "REFUSES an all-zero series instead of calling it O(1)" {
// fitprobe `compute` allocation count. Pure CPU, allocates nothing.
// Reporting O(1) here would be a confident answer with nothing behind it.
let v = _s4(0, 0, 0, 0)
assert elb_gate(v, 2, 10) == 3, "all-zero series must be REFUSED"
assert elb_measured_curve(v, 10) < 0, "unclassifiable returns -1"
}
test "REFUSES an implausibly flat series" {
// The shape produced when clang closes a loop to a multiply: a real
// answer, no work done, no movement across an 8x input range.
let v = _s4(1000, 1001, 1002, 1003)
assert elb_gate(v, 2, 10) == 3, "hard-flat series must be REFUSED"
}
test "gate FAILS a quadratic declared as linear" {
let v = _s4(20300, 80600, 321200, 1282400)
assert elb_gate(v, 2, 10) == 1, "O(n^2) measured vs O(n) declared must FAIL"
}
test "gate PASSES a linear series declared as linear" {
let v = _s4(208, 409, 810, 1611)
assert elb_gate(v, 2, 10) == 0, "O(n) measured vs O(n) declared must PASS"
}
test "gate reports BETTER when measured beats the declared bound" {
let v = _s4(208, 409, 810, 1611)
assert elb_gate(v, 4, 10) == 4, "O(n) measured vs O(n^2) declared is BETTER"
}
test "gate reports INDETERMINATE on disagreeing ratios" {
// fitprobe `linear` WALL TIME at these sizes: 26/19/43/78 microseconds.
// Ratios 0.73, 2.26, 1.81 disagree well past the noise threshold. The
// honest answer is "cannot tell", not a classification -- this is exactly
// why benchmarks need auto-scaled iteration counts rather than one shot.
let v = _s4(26, 19, 43, 78)
assert elb_gate(v, 2, 10) == 2, "disagreeing ratios must be INDETERMINATE"
}
test "black_box is a real barrier and returns its input" {
assert el_black_box(42) == 42, "black_box is value-preserving"
let s: Int = 0
let i: Int = 0
while i < 100 {
// Bind the call before using it in arithmetic: `x + call(...)`
// lowers to el_str_concat() on integers. Same inference defect
// as `call(...) == y` lowering to str_eq().
let bx: Int = el_black_box(1)
let s = s + bx
let i = i + 1
}
assert s == 100, "black_box does not disturb the computation"
}
test "curve names round-trip" {
assert elb_curve_from_name("O(n)") == 2, "O(n) parses"
assert elb_curve_from_name("O(n^2)") == 4, "O(n^2) parses"
assert str_eq(elb_curve_name(4), "O(n^2)"), "O(n^2) renders"
assert elb_curve_from_name("O(nonsense)") < 0, "unknown curve is -1"
}
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_env.el - native test suite for runtime/env.el
//
// Covers: env() for reading environment variables, args() returning a list,
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_fs.el - native test suite for runtime/fs.el
//
// Covers: fs_write/read round-trip, fs_exists, fs_mkdir, fs_list,
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_json.el - native test suite for runtime/json.el
//
// Covers: json_get (dot-path), typed extractors (int, bool, float),
+178
View File
@@ -0,0 +1,178 @@
import "../../runtime/eltest.el"
import "../../runtime/elbench.el"
// test_lexer_scaling.el THE ARMED GATE.
//
// This is the regression test that would have caught el #132.
//
// #132 was a strlen() inside str_char_code() and str_slice(). The lexer walks
// source one character at a time, so every character access rescanned the whole
// remaining input: O(n) per character over n characters = O(n^2). It shipped for
// months. It was found by a geometric sweep, not by reading code.
//
// So this test IS a geometric sweep. It scans a string of length n, character by
// character, at four doubling sizes, and asserts the cost is linear. If anyone
// reintroduces a per-character rescan in str_char_code, in str_slice, in any
// accessor the lexer leans on the measured curve becomes O(n^2) and this fails.
//
// The value is in it being ARMED, not in it currently failing. It passes today
// because #132 is fixed. That is the correct state for a regression gate.
//
// Note the deliberate `let c: Int = str_char_code(...)` binding in the scan loop.
// Inlining it as `total + str_char_code(s, i)` lowers to el_str_concat() on
// integers the Plus arm of the operator-typing family, still open at the time
// of writing. Binding first is the safe form.
// _mk_string build a string of length >= n by DOUBLING.
//
// Deliberately not `s = s + "x"` n times: that is itself quadratic in bytes and
// would contaminate the very measurement this test exists to take. Doubling
// allocates ~2n total.
fn _mk_string(n: Int) -> String {
let s: String = "abcdefgh"
while str_len(s) < n {
let s = s + s
}
return s
}
// _scan walk the string one character at a time, REPS times.
//
// This is the lexer's access pattern reduced to its essential shape. The
// repetitions lift the measurement clear of timer resolution; without them the
// smaller sizes land in noise and the classifier correctly reports
// INDETERMINATE rather than guessing.
fn _scan(s: String, n: Int, reps: Int) -> Int {
let total: Int = 0
let r: Int = 0
while r < reps {
let i: Int = 0
while i < n {
let c: Int = str_char_code(s, i)
let total = total + c
let i = i + 1
}
let r = r + 1
}
return total
}
// _measure_scan microseconds for a full scan sweep point.
fn _measure_scan(n: Int, reps: Int) -> Int {
let s: String = _mk_string(n)
// WARMUP, discarded. Without it the small-n end of the sweep is dominated
// by cold caches and reads as superlinear on genuinely linear work --
// measured ratios 3.37 2.92 1.76 1.65 on exactly this workload.
let w: Int = _scan(s, n, 2)
let wj: Int = el_black_box(w)
let t0: Int = el_now_instant()
let got: Int = _scan(s, n, reps)
let t1: Int = el_now_instant()
// Feed the result through the barrier so the scan cannot be elided.
let sink: Int = el_black_box(got)
if sink == 0 { println("") }
return (t1 - t0) / 1000
}
fn _series4(a: Int, b: Int, c: Int, d: Int) -> [Int] {
let l: [Int] = native_list_empty()
let l = native_list_append(l, a)
let l = native_list_append(l, b)
let l = native_list_append(l, c)
let l = native_list_append(l, d)
return l
}
test "character scan is LINEAR in time -- regression gate for el #132" {
let reps: Int = 40
let t1: Int = _measure_scan(16384, reps)
let t2: Int = _measure_scan(32768, reps)
let t3: Int = _measure_scan(65536, reps)
let t4: Int = _measure_scan(131072, reps)
let series: [Int] = _series4(t1, t2, t3, t4)
let verdict: Int = elb_gate(series, 2, 50)
let measured: Int = elb_measured_curve(series, 50)
// Report the actual numbers regardless of outcome. A gate that fires
// without showing its evidence is just an assertion.
println(" scan us: " + int_to_str(t1) + " " + int_to_str(t2) + " "
+ int_to_str(t3) + " " + int_to_str(t4)
+ " -> " + elb_curve_name(measured) + " [" + elb_verdict_name(verdict) + "]")
// PASS (0) or BETTER (4) are both acceptable. FAIL (1) means someone
// reintroduced superlinear per-character cost. REFUSED (3) or
// INDETERMINATE (2) mean the measurement is untrustworthy -- which is
// also a failure of this test, deliberately: a gate that cannot measure
// must not report success.
assert verdict == 0 || verdict == 4, "character scan must measure O(n) or better"
}
test "string building by doubling stays linear in allocated bytes" {
let b1: Int = el_alloc_bytes()
let s1: String = _mk_string(8192)
let b2: Int = el_alloc_bytes()
let s2: String = _mk_string(16384)
let b3: Int = el_alloc_bytes()
let s3: String = _mk_string(32768)
let b4: Int = el_alloc_bytes()
let s4: String = _mk_string(65536)
let b5: Int = el_alloc_bytes()
let series: [Int] = _series4(b2 - b1, b3 - b2, b4 - b3, b5 - b4)
let verdict: Int = elb_gate(series, 2, 1000)
let measured: Int = elb_measured_curve(series, 1000)
println(" bytes: " + int_to_str(b2 - b1) + " " + int_to_str(b3 - b2) + " "
+ int_to_str(b4 - b3) + " " + int_to_str(b5 - b4)
+ " -> " + elb_curve_name(measured) + " [" + elb_verdict_name(verdict) + "]")
assert verdict == 0 || verdict == 4, "doubling build must be O(n) in bytes"
assert str_len(s4) >= 65536, "final string reached the requested size"
}
// _scan_quadratic a DELIBERATELY quadratic scan: for each position, rescan
// from the start. This is precisely what el #132 did strlen() from offset 0
// on every character access reproduced here so the gate can be proven to
// FIRE, not merely to pass on healthy code. An unproven gate is decoration.
fn _scan_quadratic(s: String, n: Int) -> Int {
let total: Int = 0
let i: Int = 0
while i < n {
let j: Int = 0
while j < i {
let c: Int = str_char_code(s, j)
let total = total + c
let j = j + 1
}
let i = i + 1
}
return total
}
fn _measure_quadratic(n: Int) -> Int {
let s: String = _mk_string(n)
let w: Int = _scan_quadratic(s, 64)
let wj: Int = el_black_box(w)
let t0: Int = el_now_instant()
let got: Int = _scan_quadratic(s, n)
let t1: Int = el_now_instant()
let sink: Int = el_black_box(got)
return (t1 - t0) / 1000
}
test "the gate FIRES on a live quadratic scan -- proves it is armed" {
let q1: Int = _measure_quadratic(1024)
let q2: Int = _measure_quadratic(2048)
let q3: Int = _measure_quadratic(4096)
let q4: Int = _measure_quadratic(8192)
let series: [Int] = _series4(q1, q2, q3, q4)
let verdict: Int = elb_gate(series, 2, 50)
let measured: Int = elb_measured_curve(series, 50)
println(" quad us: " + int_to_str(q1) + " " + int_to_str(q2) + " "
+ int_to_str(q3) + " " + int_to_str(q4)
+ " -> " + elb_curve_name(measured) + " [" + elb_verdict_name(verdict) + "]")
assert measured == 4, "a rescan-from-zero workload must classify O(n^2)"
assert verdict == 1, "declared O(n) against measured O(n^2) must FAIL the gate"
}
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_math.el - native test suite for runtime/math.el
//
// Covers: integer math (abs, max, min), float math (sqrt, log, sin, cos, pi),
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_state.el - native test suite for runtime/state.el
//
// Covers: state_set/get/del, state_has, state_get_or, state_keys,
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_string.el - native test suite for runtime/string.el
//
// Covers: type conversions, core primitives, comparison and search,
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_text.el - native test suite for text primitives.
//
// Mirrors the acceptance corpus in tests/text/examples/ using the
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// test_time.el - native test suite for runtime/time.el
//
// Covers: time_now (positive timestamp), time_to_parts (UTC decomposition),
+534
View File
@@ -0,0 +1,534 @@
import "../../runtime/eltest.el"
// test_transduce.el transduction produces a SUBGRAPH, not a point.
//
// WHAT IS ACTUALLY UNDER TEST. #144 moved transduction into the language and
// got the dispatch right: realizers declared in El, resolved by name, no
// runtime patch per modality. It got the RESULT TYPE wrong
// `transduce(signal, modality) -> Geometry`, one vector per signal.
//
// One vector is a FINGERPRINT. It can be matched and it can be ranked, and
// that is the whole of what it can ever do. It cannot be decomposed, cannot
// have one part grounded while another is not, and cannot be contradicted in
// one part while holding in another because it has no parts. Treating
// transduction as a CONVERSION (signal in, position out) is the premise this
// file exists to falsify.
//
// A song is not a point. It decomposes into pitch, interval, rhythm, harmonic
// function components, each with its own geometry, plus the relations among
// them. THE SONG IS THE STRUCTURE OF THE RELATIONS. So transduction yields a
// Manifold: named components carrying geometry, and typed weighted relations
// between them.
//
// The geometry tests below are UNCHANGED from #144 and still pass, which is
// the point: Geometry was never wrong, it was misplaced. A vector is the right
// representation for a COMPONENT. It was only ever wrong as the representation
// of a whole transduced signal.
//
// COMPARISON DISCIPLINE IN THIS FILE (measured 2026-08-16, not stylistic):
// elc lowers `a == b` to a NUMERIC comparison only when both operand names are
// in the per-function int-name set, which `let x: Int` populates. A bare call
// like `manifold_size(m) == 5` is not a registered name, so it lowers to
// `str_eq(...)` strcmp on two integers reinterpreted as pointers. `<` and `>`
// lower directly via binop_to_c with no type inference at all, so truthiness is
// written `> 0` / `< 1` here, and any exact `==` is done on a value first bound
// through `let x: Int`.
//
// ONE FURTHER RULE, measured while writing this file: that int-name set LEAKS
// ACROSS `test` BLOCKS. Binding `dn` as a Float in one test and as an Int in
// another silently demoted the Int comparison to str_eq and failed an
// assertion that was arithmetically true. Every Int-bound name compared with
// `==` here is therefore spelled UNIQUELY across the whole file (note_dim,
// iv_dim, ...), rather than reusing a short name per test.
// A DECOMPOSING realizer, written entirely in El
// "tone" signals are note letters, e.g. "CEG". This realizer does NOT return
// one vector for the chord. It returns the PARTS one component per note, one
// per interval between adjacent notes and the relations that make those
// parts a chord rather than an unordered bag of pitches.
//
// The interval is deliberately a COMPONENT, not an attribute of a note. An
// interval is a thing with its own geometry that belongs to neither endpoint;
// modelling it as a field on a note is exactly the collapse this change
// rejects, one level down.
fn tone_realizer(signal: String) -> Manifold {
let m: Manifold = manifold_new()
let n: Int = str_len(signal)
let i: Int = 0
while i < n {
let code: Int = str_char_code(signal, i)
let g: Geometry = geometry_new(2)
let s0: Int = geometry_set(g, 0, int_to_float(code))
let s1: Int = geometry_set(g, 1, int_to_float(i))
let idx: Int = manifold_add(m, "note:" + int_to_str(i), "pitch", g)
let f: Int = geometry_free(g)
i = i + 1
}
let j: Int = 1
while j < n {
let a: Int = str_char_code(signal, j - 1)
let b: Int = str_char_code(signal, j)
let lo: String = "note:" + int_to_str(j - 1)
let hi: String = "note:" + int_to_str(j)
let key: String = "interval:" + int_to_str(j - 1) + "-" + int_to_str(j)
let g: Geometry = geometry_new(1)
let s: Int = geometry_set(g, 0, int_to_float(b - a))
let idx: Int = manifold_add(m, key, "interval", g)
let f: Int = geometry_free(g)
let e1: Int = manifold_relate(m, key, "spans", lo, 0.9)
let e2: Int = manifold_relate(m, key, "spans", hi, 0.9)
let e3: Int = manifold_relate(m, lo, "sounds_before", hi, 0.8)
j = j + 1
}
m
}
// A second realizer for a different modality, to prove the registry keys on
// modality and does not just hand back "the last thing registered". Its
// decomposition has a DIFFERENT shape two components, one relation so a
// test can tell the two organs apart by structure alone.
fn pulse_realizer(signal: String) -> Manifold {
let m: Manifold = manifold_new()
let ga: Geometry = geometry_new(1)
let sa: Int = geometry_set(ga, 0, 1.0)
let ia: Int = manifold_add(m, "onset", "event", ga)
let fa: Int = geometry_free(ga)
let gb: Geometry = geometry_new(1)
let sb: Int = geometry_set(gb, 0, 0.0)
let ib: Int = manifold_add(m, "decay", "envelope", gb)
let fb: Int = geometry_free(gb)
let e: Int = manifold_relate(m, "onset", "decays_into", "decay", 0.7)
m
}
// #144's ACTUAL CONTRACT, preserved verbatim as a control: a realizer that
// returns one vector for the whole signal. This is not a strawman it is what
// the merged primitive asked realizers to be. It must now transduce NOTHING.
fn fingerprint_realizer(signal: String) -> Geometry {
let g: Geometry = geometry_new(4)
let n: Int = str_len(signal)
let a: Int = geometry_set(g, 0, int_to_float(n))
let b: Int = geometry_set(g, 1, int_to_float(n * 2))
g
}
// A realizer returning something that is not a value at all.
fn bogus_realizer(signal: String) -> Manifold {
return 12345
}
//
// Geometry unchanged from #144. A vector is the right representation for a
// COMPONENT; it was only ever wrong as the representation of a whole signal.
//
test "geometry-is-a-value-with-its-own-width" {
let g: Geometry = geometry_new(8)
let live: Int = geometry_is(g)
assert live > 0, "geometry_new returns a live Geometry"
let d: Int = geometry_dim(g)
assert d == 8, "a Geometry carries its own width"
let freed: Int = geometry_free(g)
assert freed > 0, "geometry_free reports what it did"
}
test "geometry-rejects-nonsense-without-an-arbitrary-bound" {
let zero: Geometry = geometry_new(0)
let z: Int = geometry_is(zero)
assert z < 1, "dim 0 is not a geometry"
let neg: Geometry = geometry_new(-4)
let n: Int = geometry_is(neg)
assert n < 1, "negative dim is not a geometry"
let nd: Int = geometry_dim(0)
assert nd < 1, "geometry_dim of a non-geometry is 0"
let ni: Int = geometry_is(0)
assert ni < 1, "geometry_is of a non-geometry is 0"
let nf: Int = geometry_free(0)
assert nf < 1, "geometry_free of a non-geometry is a no-op"
}
test "geometry-components-round-trip" {
let g: Geometry = geometry_new(3)
let s0: Int = geometry_set(g, 0, 1.5)
let s1: Int = geometry_set(g, 1, -2.5)
assert s0 > 0, "set in range succeeds"
let oob: Int = geometry_set(g, 3, 9.0)
assert oob < 1, "set out of range is refused, not silently dropped"
let v0: Float = geometry_get(g, 0)
let d0: Float = v0 - 1.5
assert d0 < 0.001, "component 0 round-trips"
assert d0 > -0.001, "component 0 round-trips"
let v1: Float = geometry_get(g, 1)
let d1: Float = v1 + 2.5
assert d1 < 0.001, "component 1 round-trips (negative)"
assert d1 > -0.001, "component 1 round-trips (negative)"
let freed: Int = geometry_free(g)
}
test "hex-is-an-edge-adapter-and-derives-its-own-width" {
let g: Geometry = geometry_from_f32le_hex("0000803f00000040")
let live: Int = geometry_is(g)
assert live > 0, "valid hex decodes to a Geometry"
let hex_dim: Int = geometry_dim(g)
assert hex_dim == 2, "width is DERIVED from the input, never supplied"
let back: String = geometry_to_f32le_hex(g)
assert str_eq(back, "0000803f00000040"), "hex round-trips exactly"
let freed: Int = geometry_free(g)
}
test "hex-rejects-malformed-input" {
let empty: Geometry = geometry_from_f32le_hex("")
let e: Int = geometry_is(empty)
assert e < 1, "empty hex is not a geometry"
let ragged: Geometry = geometry_from_f32le_hex("0000803f0000")
let r: Int = geometry_is(ragged)
assert r < 1, "length not a multiple of 8 is refused"
let nonhex: Geometry = geometry_from_f32le_hex("zzzzzzzz")
let nh: Int = geometry_is(nonhex)
assert nh < 1, "non-hex characters are refused"
}
test "norm-lets-a-caller-check-a-realizer-emitted-signal" {
let g: Geometry = geometry_new(2)
let z: Float = geometry_norm(g)
assert z < 0.001, "a fresh geometry is zero — norm says so"
let s0: Int = geometry_set(g, 0, 3.0)
let s1: Int = geometry_set(g, 1, 4.0)
let nrm: Float = geometry_norm(g)
let dnorm: Float = nrm - 5.0
assert dnorm < 0.001, "3-4-5: norm is 5"
assert dnorm > -0.001, "3-4-5: norm is 5"
let freed: Int = geometry_free(g)
}
//
// Manifold the corrected result of a transduction
//
test "a-manifold-is-a-value-that-holds-parts-and-relations" {
let m: Manifold = manifold_new()
let live: Int = manifold_is(m)
assert live > 0, "manifold_new returns a live Manifold"
let fresh_sz: Int = manifold_size(m)
assert fresh_sz == 0, "a fresh manifold has no components"
let fresh_rc: Int = manifold_rel_count(m)
assert fresh_rc == 0, "a fresh manifold has no relations"
let freed: Int = manifold_free(m)
assert freed > 0, "manifold_free reports what it did"
}
test "manifold-accessors-are-total" {
let ni2: Int = manifold_is(0)
assert ni2 < 1, "manifold_is of a non-manifold is 0"
let ns: Int = manifold_size(0)
assert ns < 1, "manifold_size of a non-manifold is 0"
let nf2: Int = manifold_free(0)
assert nf2 < 1, "manifold_free of a non-manifold is a no-op"
let k: String = manifold_key(0, 0)
assert str_eq(k, ""), "manifold_key of a non-manifold is empty, never a crash"
}
test "components-are-addressed-by-key-not-by-index" {
// The key is what survives persistence: a component becomes a node, and it
// is separately groundable precisely because it is separately NAMED.
let m: Manifold = manifold_new()
let g: Geometry = geometry_new(1)
let s: Int = geometry_set(g, 0, 7.0)
let first_idx: Int = manifold_add(m, "rhythm", "temporal", g)
assert first_idx == 0, "the first component is index 0"
let found_idx: Int = manifold_index_of(m, "rhythm")
assert found_idx == 0, "a component is found by its key"
let missing: Int = manifold_index_of(m, "never_added")
assert missing < 0, "an unknown key resolves to -1, not to component 0"
let role: String = manifold_role(m, 0)
assert str_eq(role, "temporal"), "a component carries what KIND of part it is"
let f: Int = geometry_free(g)
let fm: Int = manifold_free(m)
}
test "a-duplicate-key-is-refused-because-addressing-must-be-unambiguous" {
let m: Manifold = manifold_new()
let g: Geometry = geometry_new(1)
let ok_idx: Int = manifold_add(m, "pitch", "spectral", g)
assert ok_idx == 0, "first add succeeds"
let dup: Int = manifold_add(m, "pitch", "spectral", g)
assert dup < 0, "two components answering to one name is not an addressing scheme"
let dup_sz: Int = manifold_size(m)
assert dup_sz == 1, "and the duplicate did not land"
let f: Int = geometry_free(g)
let fm: Int = manifold_free(m)
}
test "a-part-with-no-geometry-is-not-a-part" {
let m: Manifold = manifold_new()
let bad: Int = manifold_add(m, "ghost", "none", 0)
assert bad < 0, "a non-Geometry is refused as a component"
let empty_key: Int = manifold_add(m, "", "none", geometry_new(1))
assert empty_key < 0, "an unaddressable component is refused"
let none_sz: Int = manifold_size(m)
assert none_sz < 1, "nothing landed"
let fm: Int = manifold_free(m)
}
test "an-edge-to-a-nonexistent-endpoint-is-refused-not-dropped" {
// A decomposition that silently loses edges is indistinguishable from one
// that never had them.
let m: Manifold = manifold_new()
let g: Geometry = geometry_new(1)
let a: Int = manifold_add(m, "here", "part", g)
let dangling: Int = manifold_relate(m, "here", "points_at", "nowhere", 0.5)
assert dangling < 1, "an edge to an unknown target is refused"
let backwards: Int = manifold_relate(m, "nowhere", "points_at", "here", 0.5)
assert backwards < 1, "an edge from an unknown source is refused"
let dang_rc: Int = manifold_rel_count(m)
assert dang_rc < 1, "and no relation was recorded"
let f: Int = geometry_free(g)
let fm: Int = manifold_free(m)
}
test "a-component-owns-its-geometry-independently-of-the-caller" {
// manifold_add COPIES. Freeing the caller's vector must not disturb the
// component, or a decomposition would be unusable the moment it was built.
let m: Manifold = manifold_new()
let g: Geometry = geometry_new(2)
let s0: Int = geometry_set(g, 0, 42.0)
let idx: Int = manifold_add(m, "part", "kind", g)
let freed: Int = geometry_free(g)
assert freed > 0, "the caller freed its own vector"
let back: Geometry = manifold_geometry(m, 0)
let live: Int = geometry_is(back)
assert live > 0, "the component still has geometry"
let v: Float = geometry_get(back, 0)
let dv: Float = v - 42.0
assert dv < 0.001, "and it is the right geometry"
assert dv > -0.001, "and it is the right geometry"
let fb: Int = geometry_free(back)
let fm: Int = manifold_free(m)
}
//
// transduce signal in, SUBGRAPH out
//
test "a-realizer-declared-in-el-is-a-first-class-realizer" {
// THE CLAIM, unchanged from #144: tone_realizer is an ordinary El function.
// It is not in the runtime and the compiler knows nothing about it.
// Registering it by name is enough to make it the organ for a modality.
let reg: Int = realizer_register("tone", "tone_realizer")
assert reg > 0, "an El fn registers as a realizer by name"
let has: Int = realizer_has("tone")
assert has > 0, "the modality now has an organ"
let m: Manifold = transduce("CEG", "tone")
let live: Int = manifold_is(m)
assert live > 0, "transduce returns a real Manifold"
let fm: Int = manifold_free(m)
}
test "transduction-decomposes-a-signal-into-parts" {
// THE CENTRAL CLAIM. "CEG" is three notes. What comes back is not one
// vector standing for a chord it is five addressable parts (three notes,
// two intervals) and six relations. A fingerprint has one part by
// construction and could not express this at any width.
let reg: Int = realizer_register("tone", "tone_realizer")
let m: Manifold = transduce("CEG", "tone")
let ceg_sz: Int = manifold_size(m)
assert ceg_sz == 5, "three notes and two intervals are five distinct parts"
let ceg_rc: Int = manifold_rel_count(m)
assert ceg_rc == 6, "and the parts stand in six stated relations"
// Every part is independently addressable BY NAME.
let n0: Int = manifold_index_of(m, "note:0")
assert n0 > -1, "the first note is addressable on its own"
let n2: Int = manifold_index_of(m, "note:2")
assert n2 > -1, "so is the third"
let iv: Int = manifold_index_of(m, "interval:0-1")
assert iv > -1, "so is the interval between the first two"
let fm: Int = manifold_free(m)
}
test "each-part-carries-its-own-geometry" {
let reg: Int = realizer_register("tone", "tone_realizer")
let m: Manifold = transduce("CEG", "tone")
// 'C' is 67. The note component's geometry is the note's, not the chord's.
let note_i: Int = manifold_index_of(m, "note:0")
let gn: Geometry = manifold_geometry(m, note_i)
let note_dim: Int = geometry_dim(gn)
assert note_dim == 2, "a note component has the width its realizer gave it"
let pitch: Float = geometry_get(gn, 0)
let dpitch: Float = pitch - 67.0
assert dpitch < 0.001, "and it is C, so the signal reached the El realizer"
assert dpitch > -0.001, "and it is C, so the signal reached the El realizer"
// Parts may have DIFFERENT widths. A single vector per signal cannot
// represent parts of unequal dimensionality at all.
let iv_i: Int = manifold_index_of(m, "interval:0-1")
let gi: Geometry = manifold_geometry(m, iv_i)
let iv_dim: Int = geometry_dim(gi)
assert iv_dim == 1, "an interval component has its own, different width"
let f1: Int = geometry_free(gn)
let f2: Int = geometry_free(gi)
let fm: Int = manifold_free(m)
}
test "the-relations-are-content-no-single-part-carries" {
// THE POINT OF THE WHOLE CHANGE. C->E is two semitones. That "2" is not a
// property of C and not a property of E; it exists only BETWEEN them. A
// representation with no relations cannot hold it, which is why collapsing
// a signal to one vector does not merely lose resolution it loses a
// category of content.
let reg: Int = realizer_register("tone", "tone_realizer")
let m: Manifold = transduce("CEG", "tone")
let step_i: Int = manifold_index_of(m, "interval:0-1")
let gi: Geometry = manifold_geometry(m, step_i)
let step: Float = geometry_get(gi, 0)
let dstep: Float = step - 2.0
assert dstep < 0.001, "C to E is two semitones"
assert dstep > -0.001, "C to E is two semitones"
// And the interval is WIRED to both endpoints, so the structure says which
// two things it is the interval between.
let spans: Int = 0
let span_rc: Int = manifold_rel_count(m)
let k: Int = 0
while k < span_rc {
let rn: String = manifold_rel_name(m, k)
let rf: String = manifold_rel_from(m, k)
if str_eq(rn, "spans") {
if str_eq(rf, "interval:0-1") { spans = spans + 1 }
}
k = k + 1
}
assert spans == 2, "the interval is related to both notes it spans"
let fg: Int = geometry_free(gi)
let fm: Int = manifold_free(m)
}
test "relation-weight-is-the-grounding-carried-on-the-edge" {
// correspondence-and-censorship.md §1: grounding is an attribute of the
// edge and it IS the weight one quantity, not a score computed beside
// it. A realizer states a relation and its weight is the claim.
let reg: Int = realizer_register("tone", "tone_realizer")
let m: Manifold = transduce("CE", "tone")
let ce_rc: Int = manifold_rel_count(m)
assert ce_rc == 3, "one interval yields two spans and one ordering"
let found_w: Int = 0
let k: Int = 0
while k < ce_rc {
let rn: String = manifold_rel_name(m, k)
if str_eq(rn, "sounds_before") {
let w: Float = manifold_rel_weight(m, k)
let dw: Float = w - 0.8
if dw < 0.001 { if dw > -0.001 { found_w = found_w + 1 } }
}
k = k + 1
}
assert found_w == 1, "the ordering relation carries the weight its realizer stated"
let fm: Int = manifold_free(m)
}
test "distinct-signals-decompose-differently" {
let reg: Int = realizer_register("tone", "tone_realizer")
let m2: Manifold = transduce("CE", "tone")
let m3: Manifold = transduce("CEG", "tone")
let two_sz: Int = manifold_size(m2)
let three_sz: Int = manifold_size(m3)
assert two_sz == 3, "two notes decompose into two notes and one interval"
assert three_sz == 5, "three notes decompose into three notes and two intervals"
// Structure differs, not just position: fingerprints of a two-note and a
// three-note signal have identical shape and differ only numerically.
let two_rc: Int = manifold_rel_count(m2)
let three_rc: Int = manifold_rel_count(m3)
assert two_rc < three_rc, "and the relational structure itself differs"
let f2: Int = manifold_free(m2)
let f3: Int = manifold_free(m3)
}
test "the-registry-keys-on-modality" {
let r1: Int = realizer_register("tone", "tone_realizer")
let rp: Int = realizer_register("pulse", "pulse_realizer")
assert rp > 0, "a second modality registers independently"
let mt: Manifold = transduce("CEG", "tone")
let mp: Manifold = transduce("CEG", "pulse")
let tone_sz: Int = manifold_size(mt)
let pulse_sz: Int = manifold_size(mp)
assert tone_sz == 5, "tone still routes to its own realizer"
assert pulse_sz == 2, "pulse routes to a different realizer, with its own decomposition"
let onset: Int = manifold_index_of(mp, "onset")
assert onset > -1, "and to that realizer's own component vocabulary"
let f1: Int = manifold_free(mt)
let f2: Int = manifold_free(mp)
}
test "no-organ-is-reported-as-no-organ" {
// A modality with no realizer must transduce to NOTHING. It must never
// fall back to embedding a description of the signal and calling that
// perception that silent substitution is the original defect.
let has: Int = realizer_has("echolocation")
assert has < 1, "unregistered modality has no organ"
let m: Manifold = transduce("anything", "echolocation")
let live: Int = manifold_is(m)
assert live < 1, "no realizer means no manifold, not a fake one"
}
test "registration-of-an-unresolvable-name-fails-loudly" {
let bad: Int = realizer_register("ghost", "no_such_function_anywhere")
assert bad < 1, "an unresolvable realizer name is a registration failure"
let has: Int = realizer_has("ghost")
assert has < 1, "and nothing gets registered"
}
test "a-fingerprint-realizer-transduces-nothing" {
// THE SUPERSESSION OF #144, asserted directly. fingerprint_realizer is
// exactly what the merged primitive asked a realizer to be: signal in, one
// Geometry out. It resolves, so registration succeeds the organ is
// present. But it does not decompose, so it does not transduce.
//
// This is a deliberate hard failure. "No organ" and "an organ that only
// fingerprints" must not be indistinguishable, which is the same
// distinction realizer_register already draws between an absent and a
// broken organ. A modality with genuinely one part says so with
// manifold_single, and is then visibly a size-1 manifold.
let reg: Int = realizer_register("fingerprint", "fingerprint_realizer")
assert reg > 0, "the symbol resolves, so registration succeeds"
let m: Manifold = transduce("x", "fingerprint")
let live: Int = manifold_is(m)
assert live < 1, "a single vector is not a transduction"
}
test "a-realizer-returning-nonsense-transduces-nothing" {
let reg: Int = realizer_register("bogus", "bogus_realizer")
assert reg > 0, "the symbol resolves, so registration succeeds"
let m: Manifold = transduce("x", "bogus")
let live: Int = manifold_is(m)
assert live < 1, "a non-Manifold return transduced nothing"
}
test "the-one-part-case-is-a-size-one-manifold-not-a-bare-vector" {
// Some modalities really do have one part. That is a manifold of size 1
// a special case of decomposition, not a parallel path back to a
// fingerprint. Anything reading it still asks manifold_size and still gets
// a real answer, and a second part can be added later without changing the
// type of the thing.
let g: Geometry = geometry_new(3)
let s: Int = geometry_set(g, 0, 5.0)
let m: Manifold = manifold_single("level", "scalar", g)
let live: Int = manifold_is(m)
assert live > 0, "manifold_single yields a real Manifold"
let one_sz: Int = manifold_size(m)
assert one_sz == 1, "of size one — visibly degenerate, not hidden"
let idx: Int = manifold_index_of(m, "level")
assert idx == 0, "and its one part is still addressable by name"
let f: Int = geometry_free(g)
let fm: Int = manifold_free(m)
}
@@ -0,0 +1,28 @@
fn getstr(x: String) -> String { return x }
fn getint(x: Int) -> Int { return x }
fn ok(label: String) -> Void { println("ok " + label) }
fn bad(label: String) -> Void { println("FAIL " + label) }
let s1: String = "hello"
let s2: String = "hello"
let s3: String = "world"
let i1: Int = 5
let i2: Int = 5
let i3: Int = 9
if "abc" == "abc" { ok("str literal eq") } else { bad("str literal eq") }
if "abc" == "xyz" { bad("str literal ne") } else { ok("str literal ne") }
if s1 == s2 { ok("str var eq") } else { bad("str var eq") }
if s1 == s3 { bad("str var ne") } else { ok("str var ne") }
if getstr("hi") == "hi" { ok("str call vs literal") } else { bad("str call vs literal") }
if s1 == getstr("hello") { ok("str var vs call") } else { bad("str var vs call") }
if s1 == getstr("nope") { bad("str var vs call ne") } else { ok("str var vs call ne") }
if i1 == i2 { ok("int var eq") } else { bad("int var eq") }
if i1 == i3 { bad("int var ne") } else { ok("int var ne") }
if getint(5) == i1 { ok("int call vs var") } else { bad("int call vs var") }
if getint(9) == i1 { bad("int call vs var ne") } else { ok("int call vs var ne") }
if s1 != s3 { ok("str NOTEQ") } else { bad("str NOTEQ") }
if s1 != s2 { bad("str NOTEQ same") } else { ok("str NOTEQ same") }
if i1 != i3 { ok("int NOTEQ") } else { bad("int NOTEQ") }
if getint(9) != i1 { ok("int call NOTEQ") } else { bad("int call NOTEQ") }
println("done")
+58
View File
@@ -0,0 +1,58 @@
fn expect_int(label: String, got: Int, want: Int) -> Void {
if got == want { println("ok " + label) }
else { println("FAIL " + label + " got=" + int_to_str(got) + " want=" + int_to_str(want)) }
}
fn expect_str(label: String, got: String, want: String) -> Void {
if str_eq(got, want) { println("ok " + label) }
else { println("FAIL " + label + " got='" + got + "' want='" + want + "'") }
}
// 1. basic char access across a string
let s: String = "hello"
expect_int("char[0]=h", str_char_code(s, 0), 104)
expect_int("char[4]=o", str_char_code(s, 4), 111)
expect_int("char[5] OOB -> 0", str_char_code(s, 5), 0)
expect_int("char[-1] OOB -> 0", str_char_code(s, -1), 0)
expect_int("empty string OOB", str_char_code("", 0), 0)
// 2. slices
expect_str("slice(0,5)", str_slice(s, 0, 5), "hello")
expect_str("slice(1,3)", str_slice(s, 1, 3), "el")
expect_str("slice past end clamps", str_slice(s, 3, 99), "lo")
expect_str("slice inverted -> empty", str_slice(s, 4, 2), "")
// 3. DIFFERENT strings must not share a cached length (the real hazard)
let a: String = "abc"
let b: String = "abcdefghij"
expect_int("a[2]=c", str_char_code(a, 2), 99)
expect_int("a[3] OOB", str_char_code(a, 3), 0)
expect_int("b[9]=j", str_char_code(b, 9), 106)
expect_int("b[3]=d after a", str_char_code(b, 3), 100)
expect_int("a[3] still OOB after b", str_char_code(a, 3), 0)
// 4. many distinct strings interleaved forces cache slot collisions
fn interleave(n: Int) -> Int {
let i: Int = 0
let bad: Int = 0
while i < n {
let t: String = int_to_str(i)
let l: Int = str_len(t)
let last: Int = str_char_code(t, l - 1)
let oob: Int = str_char_code(t, l)
if oob != 0 { let bad2: Int = bad + 1
let bad: Int = bad2 }
if last == 0 { let bad3: Int = bad + 1
let bad: Int = bad3 }
let i2: Int = i + 1
let i: Int = i2
}
return bad
}
expect_int("1000 interleaved strings, no bad reads", interleave(1000), 0)
// 5. concatenation changes length cache must not report the old one
let g: String = "12345"
let g2: String = g + "6789"
expect_int("grown string len via char", str_char_code(g2, 8), 57)
expect_int("original still bounded", str_char_code(g, 5), 0)
println("done")
+1
View File
@@ -1,3 +1,4 @@
import "../../runtime/eltest.el"
// tests/runtime/string_test.el Test suite for runtime/string.el
//
// Exercises every public function exported by runtime/string.el using the
+3
View File
@@ -173,4 +173,7 @@ Each ad-hoc harness becomes `nsbx run <name> …` (or `--source` build) against
## Env knobs
`NSBX_ROOT`, `NSBX_PORT_BASE`, `NSBX_RSS_BOUND_MB`, `NSBX_REMERGE_THRESHOLD`,
`NSBX_READY_TIMEOUT_SECS` (default 15 — how long `up`/`create`/`build` wait for a
daemon to answer `/api/stats` before reporting failure; raise it if a boot is
legitimately slow under concurrent sandbox/CPU load rather than actually broken),
`EL_REPO` (for `elc` + runtime sources), `ENGRAM_LIVE_DATA_DIR`, `ENGRAM_LIVE_PLIST`.
+134 -29
View File
@@ -32,6 +32,12 @@ EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/foundation/el}"
PORT_BASE="${NSBX_PORT_BASE:-8900}"
RSS_BOUND_MB="${NSBX_RSS_BOUND_MB:-550}" # from store-fix reboot-proof (aaf13f88)
REMERGE_THRESHOLD="${NSBX_REMERGE_THRESHOLD:-40000}"
# readiness-poll window for start_daemon (0.5s ticks). Default unchanged (15s) —
# but a cold boot against the full live store, under concurrent CPU contention
# from other running sandboxes, has been observed live to take well past that.
# Bump per-invocation with NSBX_READY_TIMEOUT_SECS if `up`/`create` reports a
# not-ready failure but the daemon looks otherwise fine (see its logs/daemon.log).
READY_TICKS=$(( ${NSBX_READY_TIMEOUT_SECS:-15} * 2 ))
KEYSTONES=( "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee" "kn-5b606390-a52d-4ca2-8e0e-eba141d13440" )
# fixed probe set for retrieval-parity (stable, identity-anchored)
PARITY_QUERIES=( "who am I" "self identity core" "engram store durability" "keystone self anchor" "grounding honesty" )
@@ -77,7 +83,7 @@ _port_claimed(){ # is another sandbox already assigned this port?
live_stats(){ curl -s -m5 "$LIVE_URL/api/stats" 2>/dev/null; }
api(){ # api <name> <path> [json-body]
local name="$1" path="$2" body="${3:-}"
local port; port="$(mget "$name" "['port']")"; [ -n "$port" ] || die "unknown sandbox: $name"
local port; port="$(mget "$name" "['port']")"; [ -n "$port" ] || die "unknown sandbox: $name (run: nsbx list)"
local url="http://127.0.0.1:${port}${path}"
if [ -n "$body" ]; then curl -s -m30 -X POST -H 'Content-Type: application/json' -d "$body" "$url"
else curl -s -m30 "$url"; fi
@@ -88,6 +94,19 @@ stat_field(){ printf '%s' "$1" | sed -n "s/.*\"$2\":\([0-9]*\).*/\1/p"; }
daemon_pid(){ local f; f="$(sdir "$1")/daemon.pid"; [ -f "$f" ] && cat "$f" || true; }
daemon_alive(){ local p; p="$(daemon_pid "$1")"; [ -n "$p" ] && kill -0 "$p" 2>/dev/null; }
# daemon_health <name> : prints "stopped" | "running" | "unresponsive" (to stdout).
# "running" means the pid is alive AND /api/stats actually answered — process
# liveness alone (daemon_alive) is not proof the HTTP server is serving; a pegged
# or hung process still passes kill -0. Short timeout (2s) since this runs per-row
# in `nsbx list`.
daemon_health(){
local name="$1"
daemon_alive "$name" || { echo "stopped"; return 0; }
local port; port="$(mget "$name" "['port']")"
local s; s="$(curl -s -m2 "http://127.0.0.1:${port}/api/stats" 2>/dev/null)"
[ -n "$s" ] && echo "running" || echo "unresponsive"
}
# ---------------------------------------------------------------- elc/build ----
find_elc(){
command -v elc 2>/dev/null && return 0
@@ -123,6 +142,49 @@ _build_binary(){
ok "built: $out ($(ls -lh "$out" | awk '{print $5}'), sha $(sha "$out" | cut -c1-12))"
}
# bin_built_at <path> : human-readable build timestamp. `cp -p` preserves mtime,
# so this is the ORIGINAL build time even for binaries copied stock-prod into a
# sandbox — not the copy time.
bin_built_at(){ stat -f '%Sm' -t '%Y-%m-%d %H:%M:%S' "$1" 2>/dev/null || echo "unknown"; }
# _binary_freshness <name> : best-effort staleness note (empty string if fresh/
# unknown — never guesses). Two cases:
# - stock-prod: compares the sha recorded at create time against the CURRENTLY
# configured live real binary's sha (recomputed now, not cached) — catches
# "live prod moved on since this sandbox was cloned".
# - branch/source/rebuilt: compares the recorded source_commit against the
# LOCAL origin/dev ref (no fetch — reads whatever the repo already has) —
# catches "built from a commit that predates current dev tip".
_binary_freshness(){
local name="$1" src; src="$(mget "$name" "['source']")"
case "$src" in
stock-prod:*)
local live_bin cur_sha rec_sha
live_bin="$(_live_real_bin)"
[ -n "$live_bin" ] && [ -x "$live_bin" ] || return 0
cur_sha="$(sha "$live_bin")"; rec_sha="$(mget "$name" "['binary_sha256']")"
[ -n "$cur_sha" ] && [ -n "$rec_sha" ] && [ "$cur_sha" != "$rec_sha" ] \
&& printf 'stale: live prod binary has moved on since this sandbox was cloned (live is now %s, sha %s) — nsbx build %s --binary %s to catch up' \
"$(basename "$live_bin")" "${cur_sha:0:12}" "$name" "$live_bin"
;;
branch:*|source:*|rebuilt:*)
local commit cur behind
commit="$(mget "$name" "['source_commit']")"
[ -n "$commit" ] || return 0
cur="$(git -C "$EL_REPO" rev-parse origin/dev 2>/dev/null)" || return 0
[ -n "$cur" ] && [ "$commit" != "$cur" ] || return 0
git -C "$EL_REPO" merge-base --is-ancestor "$commit" "$cur" 2>/dev/null || return 0
behind="$(git -C "$EL_REPO" rev-list --count "$commit..$cur" 2>/dev/null)"
printf 'stale: built from %s, %s commit(s) behind local origin/dev (%s) — nsbx build %s --branch origin/dev' \
"${commit:0:12}" "${behind:-?}" "${cur:0:12}" "$name"
;;
esac
}
# _source_commit <dir> : best-effort git HEAD of a source tree used to build a
# sandbox binary, empty if not a git repo (e.g. a prebuilt --binary path has none).
_source_commit(){ git -C "$1" rev-parse HEAD 2>/dev/null || true; }
# ---------------------------------------------------------------- daemon -------
# start_daemon <name> : boots the sandbox's real engram binary on its isolated
# port against its cloned data dir, with the SAME auto-remerge net the live soul
@@ -133,9 +195,9 @@ start_daemon(){
local port bin data export key
port="$(mget "$name" "['port']")"; bin="$d/bin/engram"; data="$d/data"
key="sbx-$name"; export="$data/.scan-export.reseed-clean.json"
[ -x "$bin" ] || die "sandbox binary missing: $bin"
[ -x "$bin" ] || die "sandbox binary missing: $bin (run: nsbx build $name --source DIR | --branch REF, or nsbx destroy $name && nsbx create $name to reclone stock-prod)"
[ "$port" != "$LIVE_BIND_PORT" ] && [ "$port" != "$SOUL_PORT" ] || die "refusing forbidden port $port"
[ -f "$data/neuron.egm" ] || die "sandbox has no cloned store: $data/neuron.egm"
[ -f "$data/neuron.egm" ] || die "sandbox has no cloned store: $data/neuron.egm (data dir is corrupt/incomplete — run: nsbx destroy $name && nsbx create $name)"
# HARD guard: never point a sandbox daemon at the live data dir.
[ "$(cd "$data" && pwd -P)" != "$(cd "$LIVE_DATA_DIR" && pwd -P)" ] || die "refusing: sandbox data dir resolves to LIVE store"
@@ -150,11 +212,11 @@ start_daemon(){
echo "$pid" > "$d/daemon.pid"
# readiness poll
local url="http://127.0.0.1:$port" i s
for i in $(seq 1 30); do
for i in $(seq 1 "$READY_TICKS"); do
s="$(curl -s -m3 "$url/api/stats" 2>/dev/null)"
[ -n "$s" ] && break; sleep 0.5
done
[ -n "$s" ] || { warn "daemon did not become ready (see $d/logs/daemon.log)"; return 1; }
[ -n "$s" ] || { warn "daemon did not become ready within ${NSBX_READY_TIMEOUT_SECS:-15}s (see $d/logs/daemon.log). pid $pid may still be alive and slow to boot under load — check: lsof -iTCP:$port -P, or retry with NSBX_READY_TIMEOUT_SECS=45"; return 1; }
ok "ready pid=$pid boot-stats: $s"
# auto-remerge net (idempotent): match live edge population if the export is present
if [ -f "$export" ]; then
@@ -231,33 +293,35 @@ cmd_create(){
info "live baseline stats: ${lstats:-<unavailable>}"
# ---- determine + place the runtime binary (versioned into the snapshot) ----
local source_desc live_bin
local source_desc live_bin source_commit=""
live_bin="$(_live_real_bin)"
if [ -n "$binpath" ]; then
[ -x "$binpath" ] || die "not an executable binary: $binpath"
cp -p "$binpath" "$d/bin/engram"; source_desc="prebuilt:$binpath"
elif [ -n "$src" ]; then
_build_binary "$src" "$d/bin/engram" "$d/build"; source_desc="source:$src"
source_commit="$(_source_commit "$src")"
elif [ -n "$branch" ]; then
log "worktree: $repo @ $branch -> $d/build/worktree"
git -C "$repo" worktree add --detach "$d/build/worktree" "$branch" >/dev/null 2>&1 \
|| die "git worktree add failed ($repo @ $branch)"
_build_binary "$d/build/worktree" "$d/bin/engram" "$d/build"; source_desc="branch:$branch@$repo"
source_commit="$(_source_commit "$d/build/worktree")"
else
[ -x "$live_bin" ] || die "cannot resolve live ENGRAM_REAL_BIN: $live_bin"
cp -p "$live_bin" "$d/bin/engram"; source_desc="stock-prod:$live_bin"
fi
local bin_sha; bin_sha="$(sha "$d/bin/engram")"
info "runtime: $source_desc (sha ${bin_sha:0:12})"
info "runtime: $source_desc (sha ${bin_sha:0:12}, built $(bin_built_at "$d/bin/engram"))"
# ---- write manifest ----
python3 - "$name" "$port" "$source_desc" "$bin_sha" "$egm_sha" "$base_nodes" "$base_edges" "$(sha "$live_bin" 2>/dev/null)" <<'PY' > "$(manifest "$name")"
python3 - "$name" "$port" "$source_desc" "$bin_sha" "$egm_sha" "$base_nodes" "$base_edges" "$(sha "$live_bin" 2>/dev/null)" "$source_commit" <<'PY' > "$(manifest "$name")"
import json,sys,datetime
name,port,src,binsha,egmsha,bn,be,livebinsha=sys.argv[1:9]
name,port,src,binsha,egmsha,bn,be,livebinsha,source_commit=sys.argv[1:10]
json.dump({
"name":name,"port":int(port),"created_at":datetime.datetime.now(datetime.timezone.utc).isoformat(),
"source":src,"binary_sha256":binsha,"clone_egm_sha256":egmsha,
"live_binary_sha256":livebinsha,
"live_binary_sha256":livebinsha,"source_commit":source_commit,
"live_baseline":{"node_count":int(bn or 0),"edge_count":int(be or 0)},
"keystones":["kn-efeb4a5b-5aff-4759-8a97-7233099be6ee","kn-5b606390-a52d-4ca2-8e0e-eba141d13440"]
}, sys.stdout, indent=2)
@@ -268,6 +332,17 @@ PY
start_daemon "$name" || die "daemon failed to start"
local sstats; sstats="$(sbx_stats "$name")"
local sbn sbe; sbn="$(stat_field "$sstats" node_count)"; sbe="$(stat_field "$sstats" edge_count)"
# a boot immediately followed by an auto-remerge can leave the daemon briefly
# busy — retry rather than silently folding a 0/0 baseline into the manifest.
# `validate`'s zero-loss/reboot-prove checks compare current counts >= baseline,
# so a 0/0 baseline would make them trivially PASS regardless of real data loss.
local _bi
for _bi in 1 2 3 4 5; do
[ -n "$sbn" ] && [ "$sbn" != "0" ] && break
sleep 1
sstats="$(sbx_stats "$name")"; sbn="$(stat_field "$sstats" node_count)"; sbe="$(stat_field "$sstats" edge_count)"
done
[ -z "$sbn" ] || [ "$sbn" = "0" ] && warn "sandbox stats still empty/zero after retries — recording sbx_baseline 0/0. This makes 'nsbx validate $name' zero-loss checks trivially pass; investigate before trusting a validate PASS: nsbx status $name"
_capture_retrieval "$name" "$d/baseline/retrieval.json"
# fold sandbox baseline into manifest
python3 - "$(manifest "$name")" "$sbn" "$sbe" <<'PY'
@@ -320,7 +395,12 @@ except Exception: print("[]")' 2>/dev/null)"
# thereafter. Prod on :$LIVE_BIND_PORT/:$SOUL_PORT is unreachable from here by design.
cmd_up(){
local name; if [ $# -gt 0 ] && [ "${1#-}" = "$1" ]; then name="$1"; shift; else name="${USER:-dev}-dev"; fi
if mexists "$name"; then daemon_alive "$name" || start_daemon "$name"; else cmd_create "$name" "$@"; fi
if mexists "$name"; then
daemon_alive "$name" || start_daemon "$name" \
|| die "daemon did not become ready — see $(sdir "$name")/logs/daemon.log (try: nsbx up $name again once you've checked the log)"
else
cmd_create "$name" "$@"
fi
local port; port="$(mget "$name" "['port']")"
echo >&2
ok "your sandbox '$name' is ready at http://127.0.0.1:$port (a private copy of the mind — prod is untouchable)"
@@ -334,34 +414,37 @@ cmd_up(){
# it on the SAME clone + port (the code-change dev loop, in place).
cmd_build(){
local name="$1"; shift || true
mexists "$name" || die "no such sandbox: $name"
mexists "$name" || die "no such sandbox: $name (run: nsbx list — or nsbx create $name to make it)"
local src="" branch="" repo="$EL_REPO"
while [ $# -gt 0 ]; do case "$1" in
--source) src="$2"; shift 2;; --branch) branch="$2"; shift 2;; --repo) repo="$2"; shift 2;;
*) die "unknown flag: $1";; esac; done
local d; d="$(sdir "$name")"
stop_daemon "$name"
if [ -n "$src" ]; then _build_binary "$src" "$d/bin/engram" "$d/build"
local source_commit=""
if [ -n "$src" ]; then _build_binary "$src" "$d/bin/engram" "$d/build"; source_commit="$(_source_commit "$src")"
elif [ -n "$branch" ]; then
rm -rf "$d/build/worktree" 2>/dev/null; git -C "$repo" worktree prune 2>/dev/null
git -C "$repo" worktree add --detach "$d/build/worktree" "$branch" >/dev/null 2>&1 || die "worktree add failed"
_build_binary "$d/build/worktree" "$d/bin/engram" "$d/build"
source_commit="$(_source_commit "$d/build/worktree")"
else die "usage: nsbx build <name> --source DIR | --branch REF [--repo R]"; fi
# record new binary sha
python3 - "$(manifest "$name")" "$(sha "$d/bin/engram")" "${src:-branch:$branch}" <<'PY'
import json,sys; mf,s,src=sys.argv[1:4]
d=json.load(open(mf)); d["binary_sha256"]=s; d["source"]="rebuilt:"+src
python3 - "$(manifest "$name")" "$(sha "$d/bin/engram")" "${src:-branch:$branch}" "$source_commit" <<'PY'
import json,sys; mf,s,src,source_commit=sys.argv[1:5]
d=json.load(open(mf)); d["binary_sha256"]=s; d["source"]="rebuilt:"+src; d["source_commit"]=source_commit
json.dump(d,open(mf,'w'),indent=2)
PY
start_daemon "$name"
start_daemon "$name" \
|| die "rebuilt binary did not become ready — see $d/logs/daemon.log (the old binary is gone; fix the code and re-run nsbx build $name ...)"
ok "rebuilt + restarted on :$(mget "$name" "['port']")"
}
# ================================================================ run ==========
cmd_run(){
local name="$1"; shift || true
mexists "$name" || die "no such sandbox: $name"
daemon_alive "$name" || start_daemon "$name"
mexists "$name" || die "no such sandbox: $name (run: nsbx list — or nsbx create $name to make it)"
daemon_alive "$name" || start_daemon "$name" || die "daemon not running and failed to start — see $(sdir "$name")/logs/daemon.log"
local d port; d="$(sdir "$name")"; port="$(mget "$name" "['port']")"
# direct API form: nsbx run <name> api <path> [json]
if [ "${1:-}" = "api" ]; then
@@ -395,8 +478,8 @@ cmd_run(){
# RSS bound; retrieval parity; keystone integrity.
cmd_validate(){
local name="$1"; shift || true
mexists "$name" || die "no such sandbox: $name"
daemon_alive "$name" || start_daemon "$name"
mexists "$name" || die "no such sandbox: $name (run: nsbx list — or nsbx create $name to make it)"
daemon_alive "$name" || start_daemon "$name" || die "daemon not running and failed to start — see $(sdir "$name")/logs/daemon.log"
local d port key; d="$(sdir "$name")"; port="$(mget "$name" "['port']")"; key="sbx-$name"
local url="http://127.0.0.1:$port"
local bn be; bn="$(mget "$name" "['sbx_baseline']['node_count']")"; be="$(mget "$name" "['sbx_baseline']['edge_count']")"
@@ -489,7 +572,7 @@ PY
# Default is a DRY-RUN plan; requires --i-approve-prod-cutover to actually cut over.
cmd_promote(){
local name="$1"; shift || true
mexists "$name" || die "no such sandbox: $name"
mexists "$name" || die "no such sandbox: $name (run: nsbx list — or nsbx create $name to make it)"
local approve=0 do_data=0
while [ $# -gt 0 ]; do case "$1" in
--i-approve-prod-cutover) approve=1; shift;;
@@ -588,7 +671,7 @@ PY
# ================================================================ destroy ======
cmd_destroy(){
local name="$1"; shift || true
mexists "$name" || die "no such sandbox: $name"
mexists "$name" || die "no such sandbox: $name (run: nsbx list — or nsbx create $name to make it)"
local d; d="$(sdir "$name")"
stop_daemon "$name"
if [ -d "$d/build/worktree" ]; then
@@ -604,22 +687,44 @@ cmd_destroy(){
# ================================================================ list/status ==
cmd_list(){
[ -d "$SBX_ROOT" ] || { echo "no sandboxes"; return 0; }
printf '%-16s %-6s %-8s %-9s %s\n' NAME PORT STATE PID SOURCE
printf '%-16s %-6s %-13s %-9s %-19s %s\n' NAME PORT STATE PID "BUILT" SOURCE
local m
for m in "$SBX_ROOT"/*/manifest.json; do
[ -f "$m" ] || continue
local n p src pid state
local n p src pid state bpath built fresh
n="$(python3 -c "import json;print(json.load(open('$m'))['name'])")"
p="$(python3 -c "import json;print(json.load(open('$m'))['port'])")"
src="$(python3 -c "import json;print(json.load(open('$m'))['source'])")"
pid="$(daemon_pid "$n")"; state="stopped"; daemon_alive "$n" && state="running"
printf '%-16s %-6s %-8s %-9s %s\n' "$n" "$p" "$state" "${pid:-}" "$src"
pid="$(daemon_pid "$n")"
case "$(daemon_health "$n")" in
running) state="running";;
unresponsive) state="running(!resp)";;
*) state="stopped";;
esac
bpath="$(sdir "$n")/bin/engram"; built="$([ -f "$bpath" ] && bin_built_at "$bpath" || echo unknown)"
fresh="$(_binary_freshness "$n")"; [ -n "$fresh" ] && src="[STALE] $src"
printf '%-16s %-6s %-13s %-9s %-19s %s\n' "$n" "$p" "$state" "${pid:-}" "$built" "$src"
done
info "state 'running(!resp)' = process alive but /api/stats didn't answer — see: nsbx status <name>"
}
cmd_status(){
local name="$1"; mexists "$name" || die "no such sandbox: $name"
local name="$1"; mexists "$name" || die "no such sandbox: $name (run: nsbx list to see what exists)"
python3 -m json.tool "$(manifest "$name")"
daemon_alive "$name" && echo "state: running (pid $(daemon_pid "$name")) stats: $(sbx_stats "$name")" || echo "state: stopped"
local bpath; bpath="$(sdir "$name")/bin/engram"
if [ -f "$bpath" ]; then
echo "binary: sha=$(sha "$bpath" | cut -c1-12) built=$(bin_built_at "$bpath")"
local fresh; fresh="$(_binary_freshness "$name")"
[ -n "$fresh" ] && printf '%s%s%s\n' "$C_YEL" "$fresh" "$C_0"
fi
case "$(daemon_health "$name")" in
running)
echo "state: running (pid $(daemon_pid "$name")) stats: $(sbx_stats "$name")";;
unresponsive)
printf '%sstate: running but NOT RESPONDING%s (pid %s) — process alive, /api/stats returned nothing.\n' "$C_RED" "$C_0" "$(daemon_pid "$name")"
info "check: tail -50 $(sdir "$name")/logs/daemon.log | next: kill -9 $(daemon_pid "$name") && nsbx up $name"
;;
*) echo "state: stopped";;
esac
[ -f "$(sdir "$name")/validate.json" ] && { echo "--- last validation ---"; python3 -m json.tool "$(sdir "$name")/validate.json"; }
}