The generated C, amalgams, vendored runtime pins, and compiled binaries
from the Claude Code era are removed from the worktree. The El sources
survive; this tree is now source-only for the first-principles rebuild.
Per Principal direction 2026-08-19.
The tracked dist/platform/elc was last rebuilt at 45325f73, before the
52 KB codegen change that emits the runtime construct seam (el_seam_run /
el_seam_wrap calls on every fn) landed. The binary therefore did NOT carry
its own source's fix: a build from main compiled the OLD codegen, and
tests/integration/seam_binding.sh failed 4 of 7 (a construct declared
after the build never applied).
Rebuilt from main's source: self-host fixpoint holds byte-identical
(12978 lines), seam_binding.sh passes 7/7 against the installed binary,
including compose-on-crossing and exit-result-replacement.
'No deploy without verifying the artifact carries the fix.'
git diff errored to stderr on a malformed revision while wc -l counted empty
stdout, producing three confident IDENTICAL results that meant nothing. Had the
promoted trees actually differed, I would have reported the promotion clean.
The correct check is not 'how many files differ' but 'is the tree object the
same object' -- all three share hash 2acd9374.
All five defects are now visibly one shape: reading a PROXY instead of the
thing. One file instead of the operation, a variable name instead of the shape,
a scope instead of the whole, a pipe's exit instead of the program's, a line
count instead of object identity.
The async/future measurements were produced by a C stub in /tmp, and that
artifact was destroyed when the session worktrees were removed. The log then
asserted results with nothing behind them -- a claim inside an evidence record,
which is exactly what turns a chain of custody into a pile.
Rerun, not reconstructed. Rebuilding the missing file would have been a
fabrication with a fresh timestamp; rerunning produces new evidence with its own.
lang/tests/integration/fixtures/future.c the future, as a tagged heap object
lang/tests/integration/async_future.sh the harness, 6/6
ok unbound: synchronous, correct result
ok unbound: el_await on a non-future passes through, no crash
ok bound: does not crash
ok bound: the awaited result is correct
ok bound: the caller continues BEFORE the body finishes
ok bound: wrap returns in <10ms while the body takes 50ms
LABELLED AS A REPLICATION. The outcomes were already known when this harness was
written, so its expectations are NOT predictions committed in advance. Its
evidentiary value is that a third party can reproduce it, not that it was called
ahead of time. Recording it as anything stronger would corrupt the record it is
meant to repair.
The fixture also carries the P5 defect and its fix in a comment: the first
el_await dereferenced ->magic off an unvalidated slot and SIGSEGV'd on the
unbound path, sixty seconds after the same defect was diagnosed elsewhere in the
runtime.
ISHIKAWA: three silent miscompilations found the same day shared one shape.
method type tracked by per-function name sets, fed from annotations
machine el_val_t erases everything at the C boundary
material no propagation through expressions
measurement nothing verifies an annotation against what it annotates
root cause El has type ANNOTATIONS and no type CHECKING. The annotation
feeds dispatch and is never itself verified.
MEASURED, and it is not merely a wrong answer
let x: Int = "hello" ; x + 1 -> printed 4343631981, a string POINTER
interpreted as an integer
let s: String = 42 ; println -> dereferenced address 42
The first leaks a raw memory address into program output. The second is an
arbitrary-read primitive if the integer is ever attacker-influenced.
PREDICTIONS AND RESULTS
P1 let x: Int = "hello" compiles clean TRUE
P2 let s: String = 42 compiles clean TRUE
P3 the annotation drives dispatch, unverified TRUE
P4 same root cause as all three bugs found today TRUE
P5 checking literal-vs-annotation catches both TRUE
P6 zero false positives across the compiler's source TRUE
The emitter only RECORDS the mismatch; tools/check/annotations.sh decides,
consistent with every other check landed today.
INCOMPLETE, stated rather than hidden: only literals are checked.
let x: Int = some_string_fn() still passes, because signatures.rel carries
Int/Instant/Duration and no String entries. That is a DATA gap, not a capability
limit -- every El function declares its return type in source and codegen
already holds ret_type on every FnDef.
105/105 native, 5/5 annotation_query.sh, fixpoint ok.
cycles/ one file per Ishikawa -> scientific method -> Six Sigma loop, named
for the DEFECT not the fix, carrying the commit record as written at
the time
findings/ what the cycles produced, cross-cut: live bugs, architecture answers,
and defects in my own measurement
The organising finding is that predictions which came back FALSE produced every
significant result. Eleven of sixty-one failed, and those eleven found: that the
arity table was not drifted but 40% incomplete; that the AST traversal is
irreducible and only rules and judgments move; that guards could refuse through
the seam after all; and that routing el_bin_lookup through the gate did NOT fix
the SIGSEGV, because the fallback strlen was the hazard -- a wrong fix I would
otherwise have shipped as verified.
One cycle was run without committing predictions first and had to be discarded
as rigged. It is kept, in full, as 18-async-half-expressible.md.
ISHIKAWA: el_val_t carries integers AND tagged heap pointers, so "is this a
pointer" is undecidable without checking first. That check was a CONVENTION
every author had to know rather than a GATE they had to pass through, and
looks_like_heap_obj was static -- so every sibling translation unit re-derived
it.
MEASURED, across the five existing tags
geom_of looks_like_heap_obj full guard correct
mfld_of looks_like_heap_obj full guard correct
el_bin_lookup (uintptr_t)p < 4096 floor only reads 8 bytes BACKWARD
el_input_len s ? ... : 0 NULL only strlen's an integer
sha256_hex(50000) -> exit 139, SIGSEGV, compiled clean
PREDICTIONS AND RESULTS
P1 looks_like_heap_obj is static, not exported TRUE
P2 each tagged type re-derives the check TRUE
P3 at least one is missing guard components TRUE (two are)
P6 sha256_hex(<int>) reads out of bounds TRUE
P8 routing el_bin_lookup through the gate fixes it FALSE
P9 the legitimate hash is unchanged TRUE
P11 fixpoint and suites hold TRUE
P8 IS THE USEFUL FAILURE. Guarding the tagged lookup changed nothing --
looks_like_heap_obj(49992) correctly returns 0, el_bin_lookup bails, and then
el_input_len falls through to strlen() on address 50000. The FALLBACK was the
hazard, not the tagged path. A NULL check does not establish that a slot is a
pointer. I would have shipped the wrong fix and called it verified.
A MEASUREMENT DEFECT, fourth today: my first run of the crash reported exit=0,
because $? read head's exit through a pipe rather than the program's. I nearly
recorded a segfault as a clean run. Same shape as grepping only parser.el and
searching by variable name instead of by operation.
AND I PROVED THE HAZARD FROM THE INSIDE. Sixty seconds after diagnosing
`let s: String = 42` as an arbitrary-read primitive, I wrote the identical
defect into el_await -- dereferencing ->magic off an unvalidated slot -- and
only then found the runtime had already made it twice.
el_tagged() is now exported in el_runtime.h. Anything that dereferences a slot
without passing through it is the defect.
105/105 native, 42/42 integration across eight harnesses, fixpoint ok.
The module question ended with a limit: textual inlining destroys file
provenance, so a duplicate-definition message could name the symbol but not the
files. Threading it exposed a bigger absence first.
TOKENS HAD NO POSITION AT ALL. A token was a flat (kind, value) pair, so NO
diagnostic in El could name a place -- every error named a symbol and never a
line. That is the prerequisite the module question was resting on.
THE CHAIN, end to end
lexer counts newlines; tok_append mints (kind, value, line)
parser stride 2 -> 3; tok_line added; FnDef carries its line
codegen records <fn> defines_at:<line>
resolve_imports publishes <file> spans <start> <end> for the combined source
checker maps a combined line back to file:line-within-that-file
duplicate definition: 'helper' is defined 2 times — El has no namespacing,
so imported modules share one global scope
/tmp/modtest/a.el:1
/tmp/modtest/b.el:1
PREDICTIONS AND RESULTS
P1 15 stride sites, encapsulated in tok_kind/tok_value TRUE, but see below
P2 adding a line field is mechanical TRUE
P3 the lexer must count newlines TRUE
P4 resolve_imports can record per-file line ranges TRUE
P5 the message can then name both files TRUE
P6 token memory grows TRUE, 25.0 -> 33.9 MB (+36%)
FOUR DEFECTS, EACH FOUND BY RUNNING AND NOT BY READING
1. interp_tokens_append_all walks the token list DIRECTLY with its own copy of
the stride. Gen1 built fine and gen2 emitted corrupt C, because the
compiler's own source uses string interpolation. My search missed it because
I grepped for the variable name `tokens`; it is called `dst`/`result`.
Searching by name instead of by shape -- third time today.
2. tok_count in test_compiler.el carried the stride too. I had scoped the search
to compiler sources and it had escaped into the tests.
3. Nested resolve_imports calls accumulated spans into shared state, so each
republished meaningless line ranges under the parent's name. Making the
buffer local fixed it; guarding the WRITE did not, which is what I tried
first.
4. The first working version reported b.el:3 -- the COMBINED line against a
filename that has no line 3. A file:line that does not match the file is
worse than no line at all.
105/105 native, 37/37 integration, fixpoint ok, compiler self-checks clean.
All four questions in the Open section are now answered by measurement rather
than by argument. Concurrency, error handling, parsing, numeric literals, and
the module system.
The question is premature, and measuring says why. El's partition is a
FILESYSTEM PATH, not a neighbourhood, and there is no namespacing at all.
MEASURED
import is textual inlining (resolve_imports), guarded against double
inclusion by a __elc_imp__:<path> state key
when a .elh header exists the header is inlined instead and the .el is marked
seen, so symbols resolve at C link time -- so linking IS real, delegated to C
two modules defining `helper` emit two C functions into one translation unit
So linking barely survives the PATH partition. Whether it survives a
neighbourhood partition cannot be asked yet.
A DIAGNOSTIC REGRESSION I CAUSED, found by asking this question. cc does catch
the collision, but reports:
error: redefinition of '__el_body_helper'
error: redefinition of '__env_helper'
error: redefinition of '__thunk_helper'
error: redefinition of 'helper'
The user's own function is FOURTH. The first three are generated symbols
introduced by the unconditional-wrapper pass earlier today -- before it, there
was one clear message. Repaired by catching the collision at El level instead:
duplicate definition: 'helper' is defined 2 times — El has no namespacing,
so imported modules share one global scope
LIMIT, stated rather than hidden: textual inlining destroys file provenance. By
the time codegen runs there is one source string, so the message can say WHICH
name collides but not which files. Naming a.el and b.el needs provenance
threaded through resolve_imports.
104/104 native, 4/4 definitions_query.sh, the compiler itself reports clean,
fixpoint ok.
Both, at different layers, and the split is the same as everywhere else. The
NUMERAL is convention -- int_to_str was already form 1, because no position
determines that twelve is written 1 then 2 in base ten. The NUMBER is a
position: three things are three things regardless of notation.
But the sharper answer follows from `love = 0`. A bare `3` is a MAGNITUDE WITH
NO AXIS. It is not a position until something gives it a direction, which is
exactly why 3.days needs a calendar and why time_add(t, n, "min") had to carry
its axis as a string.
PREDICTIONS AND RESULTS
P1 numeral = convention, number = position TRUE
P2 a bare literal is dimensionless until context types it TRUE
P3 there is a measurable place where El guesses TRUE
P4 Instant + Int is not caught though Duration + Int is TRUE
P5 the rule catches it TRUE
P6 nothing legitimate in the tree relies on it TRUE
P3/P4 IS THE DEFECT, and it was found by reasoning from the philosophy and then
measured. Duration + Int was refused -- "an Int carries no unit" -- while
let t: Instant = now()
let u: Instant = t + 3
compiled to raw (t + 3) and reported CLEAN. Adding a dimensionless number to a
point is worse than adding it to a displacement: it silently moves the instant
by an unspecified amount. 3 of what? Whatever the representation happens to be,
which is the leak itself. The asymmetry had no justification; the rule was
simply never written.
P6 MATTERED. Two calendar tests looked like Instant + Int:
let later: Instant = i + 1.hour
let later: Instant = base + 15.hours
They are not. `1.hour` lexes to a Duration -- el_duration_from_nanos(1LL *
3600000000000LL) -- and both stay clean. That is the whole answer demonstrated
in one line: t + 3 is refused because 3 has no axis; t + 1.hour is accepted
because .hour supplies one.
104/104 native + 2 new, integration green, fixpoint ok.
Both, at different layers -- and it is the same split as serialization: the
convention is the BASIS, never the ACT.
lexeme -> token `fn` means function-start because someone said so CONVENTION
shape recognition given tokens, which construct is this REGION
source -> structure parsing is transduction onto that basis GEOMETRY
byte traversal something must read them in order IRREDUCIBLE
Three things push the ACT toward region rather than convention: ambiguity
(a * b needs context; a grammar resolves it with the lexer hack, a region by
neighbourhood), error recovery (nearest-region is free), and precedence, which
is ordering along an axis with a conventional parameter.
AND THE SHOULD GATE SAYS NO TO THE OBVIOUS MOVE
Every other table this session moved to data. This one stays code. The keyword
set is CLOSED by the language definition -- it does not leak the way an
allowlist does -- and the lexer runs before the program is understood, so a
program can never declare its own keywords. Externalising it costs file I/O on
every compile and buys nothing. Same verdict as is_digit in ASCII.
WHAT WAS ACTUALLY WRONG: five of 46 keywords were consumed by no parser or
codegen path. sealed, activate, seed, protocol, impl. Each stole an identifier
from users for nothing.
SECOND SILENT MISCOMPILATION OF THE DAY. Using one did not fail to parse:
let seed = 42
let impl = seed + 1
compiled CLEAN -- zero cc errors -- and printed 0 instead of 44. No diagnostic
at any layer. Fixed by removing the five.
A DEFECT IN MY OWN MEASUREMENT, caught before it did damage: my first pass
checked only parser.el and reported `test` as inert too. codegen consumes it at
4135 for --test mode, and the tree has 408 uses. Removing it would have broken
every test in the suite. The measurement was re-run across all four consumers.
100/100 native + 2 new, 31/31 integration, fixpoint ok.
The note said serialization, text encoding, storage, network, concurrency and
emission had collapsed; the table still listed all six as live capabilities.
A document that contradicts itself one screen apart is worse than one that is
merely out of date.
Also renames 27 from Secrecy to Concealment. 'Secrecy' covered one of the three
things in that row and got the other two backwards: a hash is public and a
signature exists to be read. Integrity and authenticity are grounding under
adversarial conditions, which is row 16. Only concealment stands alone.
capabilities.md cited '== lowering to str_eq unless both operand names are in a
hardcoded int-name set' as the paradigm defect. That is wrong: __int_names comes
from type annotations, which is legitimate propagation. The real defect was 35
hardcoded builtin return types one layer down, and mislocating it hid a live
miscompilation of unannotated lets.
geometry-vs-code.md listed concurrency and error handling as open. Both are
answered: ordering is a partial order and coordination is the price of
forgetting; standing is signed, so not-known and known-false are opposite
directions rather than one boolean. Added the fourth proof form (adversarial
exactness) and recorded that form 1 no longer survives as a verdict -- every row
it justified was a basis, not a capability.
Also marked cross-cutting concerns as implemented rather than predicted.
PREDICTIONS AND RESULTS
P1 is_int_call's 35 hardcoded names move to data TRUE
P2 is_int_name stays -- it is annotation propagation TRUE
P3 the dispatch stays -- it is emission TRUE
P4 codegen shrinks ~40 lines TRUE 4507 -> 4469
P5 the design doc's characterisation is WRONG TRUE
P6 the moved data also fixes the bug it exposed TRUE
P5 CORRECTS THE RECORD. el-language-design.md and geometry-vs-code.md both cite
"== lowering to str_eq unless both operand names are in a hardcoded int-name
set -- a literal list of variable names treated as integers" as the paradigm
defect. It is not one. __int_names is populated from TYPE ANNOTATIONS
(param["type"] == "Int"), which is primitive but legitimate type propagation.
The actual defect was is_int_call: 35 hardcoded builtin return types, the same
shape as the temporal 19.
P6 IS A LIVE CORRECTNESS BUG, PRE-EXISTING, NOW FIXED
let a = str_len("hello") // no annotation
let b = str_len("hi")
let c = a + b // -> el_str_concat(a, b) on two integers
Verified identical on the pre-change compiler, so not a regression. It compiled
clean, ran, and printed NOTHING where it should print 7. No error at any layer.
The repair is three lines: an unannotated let takes its type from what the
initialiser returns. The return types were already required for dispatch and
were simply never consulted at the binding site. Moving them into data is what
made the gap visible -- reading the code for eight hours did not.
98/98 native + 2 new, 31/31 integration, fixpoint ok.
The previous pass moved the type DATA and left the judgment inline, which I
stated rather than hid. This finishes it.
PREDICTIONS AND RESULTS
P1 codegen can emit operand-type relations TRUE
"main calls temporal:instant_plus_instant"
P2 the affine rules are a small closed set as data TRUE 6 rules
P3 violations still caught at build time TRUE exit=1
P4 the reporter leaves codegen TRUE 4538 -> 4507
P5 the TIME_TYPE_ERROR placeholder must STAY TRUE
P5 is the boundary of this whole approach. The emitter has to emit SOMETHING
for an illegal expression -- it cannot emit nothing and it cannot decide what
the program meant. So the placeholder is irreducible in the same way the AST
traversal was: what moved is the judgment and the wording, not the fact that
something must be written.
The rules are affine algebra and the set is closed because there are only two
kinds of thing. An Instant is a POINT, a Duration is a DISPLACEMENT: add a
displacement to a point, subtract two points for a displacement, combine
displacements. Nothing else is meaningful, which is why the enumeration in
temporal.rel cannot grow the way an allowlist does.
A defect in my own checker, found by running it: the .rel file uses aligned
columns and my awk assumed a single space, so the message came out with the
rule key still prefixed. Same class as the multi-line header parse in the arity
pass -- formatting assumptions that only fail when you look at the output.
98/98 native, 6/6 temporal_query.sh, fixpoint ok.
This block is structurally unlike the previous four. It does not only
adjudicate, it DISPATCHES: Instant + Duration must become el_instant_add_dur,
LocalDate + Duration must become el_local_date_add_dur. The emitted C depends on
the type answer, so it cannot move to a post-hoc query. Selecting which call to
emit is an emitter's actual job.
PREDICTIONS AND RESULTS
P1 the block conflates dispatch with adjudication TRUE
P2 adjudication can move, dispatch cannot TRUE
P3 this pass shrinks codegen far less than the last TRUE, and worse:
4513 -> 4537, it GREW
by 24 lines
P4 the rules are affine algebra, closed by construction TRUE
P5 no type propagation -- name tracking plus a
hardcoded list of which builtins return which type TRUE, 19 names
P3 is the honest result and it is not spun: moving 19 names into a data file
cost more lines than it saved, because a generic loader is larger than the
enumeration it replaces. The win is not line count. It is that adding a 20th
temporal builtin is now a one-line edit to signatures.rel instead of a compiler
change, and that the data is inspectable.
WHY THE HEADER CANNOT SUPPLY THIS, unlike arity: el_runtime.h declares every
builtin as returning el_val_t, because El has ONE type. That single type is why
the whole seam is cheap and it is exactly why the C boundary cannot say that
now() returns an Instant while unix_seconds() returns an Int. The El-level type
is real and the boundary erases it.
INCOMPLETE, and stated rather than hidden: P2 said adjudication could move to a
query. It has NOT. Violations still emit TIME_TYPE_ERROR inline from the
emitter. Only the type DATA moved. Moving the adjudication needs the operand
types recorded as relations, which is a further pass.
98/98 native, 4/4 temporal_signatures.sh, fixpoint ok.
codegen.el carried builtin_arity(): 344 lines, 300 entries, a hand-maintained
second copy of el_runtime.h.
PREDICTIONS AND RESULTS
P1 the table duplicates the header TRUE 243 shared names
P2 they have already drifted FALSE ZERO drift. The
duplicate had been
maintained correctly.
P3 codegen can emit call-arity relations TRUE
P4 the check becomes a query against the header TRUE
P5 codegen drops to roughly baseline TRUE 4903 -> 4512,
149 BELOW the 4661
it started at
P2 being false is the better result: the table was not WRONG, it was
INCOMPLETE. 110 functions the runtime declares had no entry, so calling them
with the wrong argument count produced no El-level diagnostic at all. Measured:
the old compiler reports 0 arity errors for __http_do_map_to_file(1); the query
reports "takes 5 arguments, called with 1".
Deriving from the header fixes coverage AND makes drift impossible by
construction. 503 signatures, versus 300 entries maintained by hand.
THREE DEFECTS IN MY OWN CHECKER, each found by running it rather than reading it
1. El names and C names differ -- `println` is `__println`. 60 of 500 decls
carry the prefix and codegen owns the mapping; the old table carried both
keys. One rule covers all 60.
2. Multi-line declarations parsed as zero params, so the checker reported
"takes 0" for a function taking 5. A diagnostic with the wrong number in it
is worse than none -- the same shape as the stale caller attribution in the
previous pass.
3. Fixing (2) by joining lines dropped 500 signatures to 334, because a
declaration preceded by a comment no longer started its record. Comments
are stripped first now.
98/98 native, 5/5 arity_query.sh, fixpoint ok.
Capability differs from prohibits_outside in one way that matters: a utility
program cannot be trusted to declare its own restrictions, because it would
declare none. So the policy comes from OUTSIDE the program -- it ships with the
language as data, editable without a compiler release.
tools/check/capabilities.rel 18 names that were string literals in codegen
tools/check/capabilities.sh the query that decides
PREDICTIONS AND RESULTS
P1 codegen emits kind + call graph, drops the 4 name tests TRUE zero #errors
P2 the 18 literals become a data file TRUE
P3 the checker catches capability violations TRUE exit=1
P4 codegen drops ~76 lines TRUE 4963 -> 4881
P5 below the 4661 baseline FALSE ~+230
TWO DEFECTS THE HARNESS FOUND THAT READING WOULD NOT HAVE
1. Calls inside main became invisible. cg_fn returns early for main -- C
provides its own -- so hooking the recording there left every call in main
unrecorded: a blind spot exactly where a program does its work. The old
cap_check_call ran from cg_expr and did see main. Moved the recording to
cg_expr.
2. Caller attribution was stale. __cg_current_fn kept whatever cg_fn set last,
so a violation in main was reported against the previously emitted function.
The test still PASSED, because the violation was detected -- only the name
was wrong, and a diagnostic naming the wrong fn is worse than none. Fixed at
all three main-emission sites; the first patch missed two because the live
path is codegen_streaming.
98/98 native, 7/7 + 4/4 + 5/5 integration, fixpoint ok.
I said prohibition could not move because "a #error has no runtime". That
conflated two separable things: WHEN a violation is detected (build time --
correct, and unchanged) and WHERE the rule and the checker live (the compiler
-- assumed).
A prohibition is a containment relation over the call graph. So codegen now
records what it saw:
sneaky calls raw_sql
allowed calls raw_sql
allowed calls @repository
repository calls prohibits:raw_sql
and tools/check/prohibitions.sh decides, at build time, outside the compiler.
PREDICTIONS AND RESULTS
P1 codegen can emit the call graph it already walks TRUE
P2 the check becomes a query outside the compiler TRUE
P3 all prohibition decisions leave codegen TRUE zero #errors now
P4 violations still caught at build time TRUE exit=1
P5 codegen drops below the 4661 baseline FALSE 4962, +301
P5 is the finding. The TRAVERSAL is irreducible -- you must walk the AST to
find calls, and those ~120 lines do not move no matter who decides. What is not
irreducible is the rule (which names) or the decision (#error). Those left. I
predicted the whole 223 lines would go because I had not separated walking from
adjudicating.
Still compiled, and measured rather than assumed: the capability-tier system
(cap_check_call, is_self_formation_call, is_dharma_call, is_llm_call,
cap_record_violation, emit_cap_violations) is 76 lines of the same shape --
prohibits_WITHIN rather than prohibits_outside, so the checker needs the
opposite polarity to absorb it.
98/98 native, 4/4 prohibition_query.sh, 7/7 seam_binding.sh, fixpoint ok.
ISHIKAWA: why did wraps_body need compile-time knowledge? Because the wrapper
called the target directly. If the wrapper calls through the seam instead, the
seam can call the body itself, and a construct bound after the build decides
how and whether to invoke it.
PREDICTIONS AND RESULTS
P1 wrap becomes runtime-bindable TRUE body x3 -> 21,
never invoked -> 111
P2 codegen shrinks TRUE 5042 -> 4977
P3 cost 5-10% from an indirect call on every fn TRUE 0.36s -> 0.39s, ~8%
P4 zero-param fns break on the empty struct TRUE empty struct is a GNU
extension, empty init
is C23. Fixed with a
char field.
P5 fixpoint holds TRUE
PROCESS FAILURE worth recording: my first patch silently did not apply because
I dropped the assert on the string replacement. The build then failed with
"undeclared identifier __thunk_noargs", which I nearly attributed to the
empty-struct prediction. The guard that would have caught it existed and I
removed it -- the same shape as every other defect found tonight.
Removed: declare_wrap, decorator_wrap, cg_wrap_target, cg_wrap_construct,
params_to_call_args, and the wraps_body scanner branch.
prohibits_outside is now the ONLY construct kind left at compile time, and it
cannot move: a #error has no runtime.
ISHIKAWA: why did exit injection still need compile-time knowledge? Because the
body-helper wrapper was only emitted when codegen already knew an exit
construct existed. The wrapper being conditional was the cause, not the wrapper
being necessary.
PREDICTIONS AND RESULTS
P1 exit becomes runtime-bindable TRUE returns 14, bound
after the build
P2 codegen shrinks TRUE 5094 -> 5044
P3 cost 5-15% from a call frame on every fn FALSE 0.37s -> 0.38s, ~3%
P4 fixpoint holds TRUE
Every fn now gets a body helper and a wrapper. It has to be unconditional:
early returns must route through something for an exit construct to observe
them, and codegen cannot know which fns will be bound after the binary exists.
Removed with the machinery: declare_exit, decorator_exit, cg_exit_target,
cg_exit_construct, and the injects_at_exit scanner branch.
Two controls failed and were rewritten rather than repaired --
no-exit-construct-emits-no-wrapper asserted the optimisation this removes, so
it is now inverted. The integration harness gained a seventh assertion: an exit
construct declared after the build replaces the result.
99/99 native, 7/7 integration, fixpoint gen2==gen3.
Five compile-time passes added 491 lines to the thing that was supposed to stop
growing. The seam is ~55 lines of C and one line of emission, and it does at
runtime what three of those five kinds did at compile time -- for programs that
are already built.
a construct declared AFTER the binary exists applies to it
free when unused: 0.36s vs 0.37s baseline across 267 indirections
dlsym was the cost, not the table scan; resolve-once recovered 3.5x
refusal works, composition works, unlinked targets are skipped not fatal
injects_at_exit and wraps_body do NOT collapse: early returns must route
through the body-helper wrapper regardless of when the target is resolved. The
wrapper is structural, which I had wrong. prohibits_outside cannot move at all
-- a #error has no runtime.
Controls: 99/99 native compiler tests, plus tests/integration/seam_binding.sh
(6/6) for the claim compile_capture structurally cannot see.
The seam's whole claim is that a construct declared AFTER a binary exists
applies to that already-built program. compile_capture only sees emitted text,
so it structurally cannot check this: it needs a built binary, a linked target,
and an environment. Verified by hand until now, which is the standing problem
this session has been about.
tests/integration/seam_binding.sh builds a probe from El source containing no
construct at all, links a target that El never references, and asserts:
ok unbound program is unaffected
ok a construct declared AFTER the build applies
ok a construct declared after the build can REFUSE
ok an unlinked target is skipped, not fatal
ok a binding for a different fn does not fire
ok two constructs compose on one crossing
6 assertions, 6 passed, 0 failed
The eight controls that failed after the strip were replaced, not repaired.
They asserted compile-time emission of capability that moved to runtime;
contorting them would have kept an assertion whose subject no longer exists.
Three took their place, asserting the emitted shape, and the behaviour they
used to cover is now the integration harness's job -- which is the honest
division, since the shape and the behaviour are no longer the same fact.
99/99 native compiler tests pass. Fixpoint holds.
PREDICTION: codegen.el drops below 4661, its size before any of these passes.
RESULT: FALSE. 5157 -> 5096. Still +435 over baseline.
injects_at_entry collapsed into the seam removed
guards_at_entry collapsed into the seam removed
injects_at_exit needs the body-helper wrapper STRUCTURAL
wraps_body needs the closure + wrapper structural
prohibits_outside a #error cannot be emitted at runtime
The wrapper is not a consequence of compile-time resolution. Early returns must
be routed through something no matter when the target is resolved, so exit
injection was never going to collapse. I predicted it would because I had
conflated "resolved late" with "emitted less".
What did collapse is entry injection and refusal -- 61 lines of compiler
replaced by one refusable indirection, with the capability now bindable after
the binary exists.
8 tests fail, and they are exactly the 8 controls for compile-time entry
injection and guards. No unrelated breakage: the controls reported precisely
what moved. They assert emission of something that now happens at runtime, so
they need rewriting as integration tests -- which the framework does not
currently support, because runtime binding needs a built binary and an
environment, not compile_capture.
Verified after the strip: fixpoint gen2==gen3, observation and refusal both
work through the seam with the compiler knowing nothing about either.
Prediction 3 was FALSE. I expected refusal to be impossible through the seam
because the entry indirection discarded its return. One line:
{ el_val_t __s = el_seam_run(EL_STR(f), 0, 0); if (__s) return __s; }
work() returns 7; bound to a refusing construct AFTER the build it returns 42.
So three of the five compile-time kinds are runtime-bindable: entry injection,
exit injection, and refusal. wraps_body needs invocation control and
prohibits_outside is compile-time by nature.
104/104 native compiler tests pass.
The 2026-07-16 review fixed telemetry growth in the GRAPH by calling
engram_prune_telemetry(48h) on every ISE insert. The 2026-08-xx move to
ENGRAM_ISE_OFFGRAPH=1 then routed every state event to a flat append-only
log instead — and that path had no retention of any kind. The prune call
still exists in server.el, but it now sits in the branch that production
never takes, so the fix reads as present while being inert.
Measured on the live store: 17.1 MB / 14,305 events over 3.56 days =
4.81 MB/day, unbounded (~1.76 GB/year).
engram_ise_log_append now compacts to a byte bound after append. Byte- and
not time-bounded on purpose: this is a flat file with no index, so size is
the property that has to be bounded, and ftell on the handle already held
is O(1) versus an O(file) timestamp scan per append. Default 64 MB retains
~13 days at the measured rate — more history than the 48h the on-graph path
kept. Override with ENGRAM_ISE_LOG_MAX_BYTES.
Compaction keeps the TAIL, never the head: engram_dreams_json reads the
last ~2 MB of this file for dream-recall, so the recent end is the end with
a reader, and KEEP (16 MB) stays well clear of that window. Resumes at the
first line boundary so the tail never starts mid-record, and only renames
over the live log when the tail was written in full — a short write must
not destroy history.
The honesty rail is unchanged: rotated-out remains "I don't remember",
never a synthesized dream. This only makes the forgetting bounded and
explicit instead of deferred forever.
Verified against a 4,000-event harness at a 200 KB cap: file bounded,
newest record retained, oldest dropped, 883 lines with zero malformed
records, tail contiguous, no .tmp residue.
HYPOTHESIS (Will's): a compiler whose one compiled mechanism is extending the
LANGUAGE — not the compiler — can compose without recompilation.
ISHIKAWA — why does a construct require a recompile today?
method codegen inlines the target call into the body
machine the binary has no table to consult
material the declaration lives in source, read at compile time
measurement nothing observes what applied at runtime
root cause the crossing is resolved at EMISSION, not at EXECUTION
CHANGE: codegen emits one unconditional indirection per fn. Which constructs
apply is read from a table that can be written AFTER the binary exists;
targets resolve through dlsym against the running image.
PREDICTIONS AND RESULTS
P1 a construct declared after the build applies TRUE
P2 an unlinked target is skipped, not fatal TRUE
P3 emitting on every fn is measurably slower FALSE — 0.37s -> 0.36s
with 267 indirections and
no bindings. Free unused.
P4 the compiler still self-hosts TRUE (see note)
DEMONSTRATED: an El program with NO decorator in its source, already compiled
and linked, picked up a construct declared afterwards:
$ /tmp/seamrun -> 7
$ echo 'work audited entry audit_entry' > constructs.txt
$ EL_CONSTRUCTS=constructs.txt /tmp/seamrun
AUDIT: work applied by audited
7
P4 note: my first fixpoint test was wrong, not the code. I compared gen1 to
gen2, which must differ whenever codegen's output changes. gen2 == gen3, 267
seam sites, stable.
MEASURED COST, and the root cause was not where I looked
0 bindings 0.36s vs 0.37s baseline free
2 bindings, dlsym per call 2.45s 6.6x
2 bindings, resolved once 0.69s 3.5x recovered
The table scan was never the cost. dlsym walks the dynamic symbol table on
every call. Resolve once and cache — which is the smallest form of what
salience does for memory: what is hot stays resolved. The 0.69s residual is
audit_entry's own printf on two of the compiler's hottest functions, not seam
overhead.
CONSEQUENCE: the five compile-time declaration kinds on iteration-1 are a
compile-time specialisation of something that resolves at runtime. They are not
wrong, but they are not the mechanism — the mechanism is one indirection, and a
kind is data.
The other half of a boundary: not what runs when something crosses, but what
may not cross at all. It was two string literals in vbd_is_restricted_name and
one #error in cg_fn — one prohibition, uneditable without a compiler release.
@decorator("prohibits_outside", "raw_sql")
fn repository() {}
fn sneaky() -> Int { raw_sql("DROP") }
// #error "boundary violation: raw_sql may only be called from an
// @repository fn, but 'sneaky' is not one"
The recursive matcher is parameterised through a state key rather than by
threading an argument through every branch of the walk — the mechanism codegen
already uses for __match_counter and __if_expr_counter. Each prohibition is
checked in its own turn, so the owning construct is known by construction and
the diagnostic names it instead of hardcoding one rule's wording.
PREDICTIONS AND RESULTS
1 the 3 duplicated uniqueness rules are textually identical TRUE
2 a declared prohibition reproduces @manager's #error TRUE
3 existing output byte-identical TRUE
4 a program can declare its own prohibition TRUE
5 fixpoint holds TRUE
I misread result 2 on first pass: a @manager fn calling dharma_emit still
emitted one #error, which looked like a failure. It is the CAPABILITY-tier rule
at codegen.el:2578, a separate prohibition system, and it fires identically on
the pre-change compiler.
MEASURED DEFECTS STILL OPEN
- two independent prohibition systems (VBD constructs, capability tiers);
only the first is declarable
- 3 uniqueness rules written 6 times, once per codegen path, kept in sync by
hand and identical today
102/102 native compiler tests pass, compiler self-hosts byte-identically.
Proven on experiment/wraps-body (2bed848): base(5) wrapped by a target that
invokes the body twice returns 10; a target that never invokes it returns 999.
Neither is expressible by deciding whether to repeat.
Root cause it corrected: 'C has no closures' is a fact about one grammar, not
about what can be emitted. And El's single type (el_val_t = int64_t) cannot
describe a callable, so codegen emits the calling convention rather than asking
El's type system for something it structurally cannot say.
ROOT CAUSE of the weaker design: "C has no closures" was taken as a fact about
what is possible. It is a fact about one grammar. Every C++ lambda, every Go
closure, every Rust closure compiles to a struct of captured values plus a
function pointer -- which is what is emitted here. Codegen emits C; it is not
written in C's syntax, and the distinction is the whole difference between a
construct that can only decide whether to repeat and one that controls
invocation.
It would also have crippled the JS backend, which has closures natively, for a
limit that applies only to the C one.
PREDICTIONS AND RESULTS
1 env struct + thunk taking void* TRUE
2 fails to compile: struct redefinition FALSE -- C allows the
inner declaration to shadow. Prediction wrong; C is more permissive than
assumed. A different real defect surfaced instead: a wrap with no exit
construct emitted `(EL_STR("f"), EL_STR(""), __r);` -- a call to an empty
target -- because has_exit was reused as "needs a wrapper" and the exit line
was emitted unconditionally. Fixed.
3 compiles when the target is declared in El FALSE -- and this is
the root cause worth keeping: El has ONE type, el_val_t = int64_t. El's type
system cannot describe a callable, so `extern fn` and the real signature
cannot be made to agree in El's own vocabulary. The fix is not a cast:
codegen DEFINES the wrap calling convention, so codegen emits the extern
declaration. The convention is not El-expressible; it is emitted.
4 target controls invocation, 0..N times TRUE
5 existing @manager output byte-identical TRUE
6 compiler fixpoint holds TRUE
7 emitting the convention makes it compile TRUE
MEASURED
base(5) wrapped by a target that invokes the body twice and sums -> 10
never_runs(5) wrapped by a target that never invokes it -> 999
Neither is expressible by "decide whether to repeat". This supersedes the
repeats_body experiment on experiment/repeats-body, which was built around the
mistaken limit.
§6 records 62 persist-after-mutate sites, 10 auth-per-route, and
index-after-append that failed at 9 of 9 — every one an obligation at a
crossing that decayed into "remember to do this afterwards." An obligation a
human must remember is not an obligation, and the 9-of-9 figure is what that
costs.
@decorator("injects_at_exit", "persist_now")
fn durable() {}
The body moves into a static helper and the visible fn becomes a wrapper, so
EARLY RETURNS pass through the exit injection. Emitting it only before the
fall-through return would have silently missed every early return — the exact
failure class this seam exists to remove. Fns with no exit construct emit
byte-identically to before.
Three independent constructs now compose on one fn, none known to the compiler:
el_val_t mutate(el_val_t k) {
{ el_val_t __g = my_auth(EL_STR("mutate"), EL_STR("authenticate")); if (__g) return __g; }
engram_boundary_beat(EL_STR("mutate"), EL_STR("manager"));
el_val_t __r = __el_body_mutate(k);
persist_now(EL_STR("mutate"), EL_STR("durable"), __r);
return __r;
}
Guard, then entry, then body, then exit. §5.2 asked whether `hold` is one
construct or two; the implementation answers one construct with two faces,
selected by declared kind rather than by two mechanisms.
Verified: existing output byte-identical, compiler self-hosts byte-identically,
early returns pass through the exit, ordering holds under composition. 98/98
native compiler tests pass.
@authenticate (6 uses), @authorize (3), @rate_limit (3) and @validate (2)
parsed, attached, and compiled to nothing. Fourteen applications that read as
protection and emitted no instruction — a function decorated @authenticate
compiled byte-identically to an undecorated one.
The missing capability was not authentication. It was that a construct could
observe a boundary but never refuse one. injects_at_entry discards the target's
result; there was no form in which a construct could say no.
@decorator("guards_at_entry", "my_auth")
fn authenticate() {}
@authenticate
@authorize
fn handler() -> String { ... }
emits, at entry:
{ el_val_t __g = my_auth(EL_STR("handler"), EL_STR("authenticate")); if (__g) return __g; }
{ el_val_t __g = my_roles(EL_STR("handler"), EL_STR("authorize")); if (__g) return __g; }
Guards precede injections because a refused call must not report a crossing,
and every guard runs where the topmost injecting construct wins — refusal is
not a role, so it does not follow the role convention.
The compiler still knows nothing about auth. The program points the construct
at its own function, which is where that decision belongs.
Verified: existing @manager/@accessor output byte-identical, compiler
self-hosts byte-identically, guards stack in declaration order and emit before
the beat. 94/94 native compiler tests pass.
codegen called fn_has_decorator for exactly three names — manager, accessor,
route. Twelve others parsed, attached as {name,args}, and compiled to nothing,
including four that look like protection: @authenticate (6 uses), @authorize
(3), @rate_limit (3), @validate (2). The cause was not that the branches were
untidy. A construct had nothing to BE, so its meaning had nowhere to live
except the emitter, and every construct was therefore a compiler edit.
A name -> injection table would have moved the enumeration twenty lines up
without removing it. So the construct now carries its own meaning:
@decorator("injects_at_entry", "engram_boundary_beat")
fn audited() {}
@audited
fn risky_op() -> Int { ... } // gets the beat, attributed to "audited"
scan_declared_decorators is a token-level pre-pass beside scan_routes, forced
by streaming codegen having no whole-program AST. manager and accessor are
seeded as the compiled-in core — the fixedSelf shape from substrate.go: a
complete fallback exists, declaration is enrichment.
This is the injection half of the seam only. The prohibition half (@manager's
#error on dharma_emit) stays hardcoded, because "which calls may appear inside
this boundary" is a query over program structure and there is nothing yet to
ask.
Verified three ways: emitted C for existing @manager/@accessor code is
byte-identical to the hardcoded path; a construct with a name the compiler has
never heard of injects correctly; the compiler self-hosts byte-identically.
90/90 native compiler tests pass.
The beat reported which function crossed a boundary, never which decorator
put the beat there. So the graph accumulated boundary events with no
attribution, and no construct could be measured — "is this decorator
earning its keep" stayed an argument instead of a traversal.
engram_boundary_beat now takes the construct and carries it on the bus as
{"construct":"..."}. The injection point, the beat, and the accumulation
already existed; only the attribution was missing.
Also pins a known defect as a test: codegen calls fn_has_decorator for
exactly three names (manager, accessor, route). Twelve others parse, attach,
and compile to nothing — including @authenticate (6 uses), @authorize (3),
@rate_limit (3) and @validate (2), which look like protection and are not.
decorator-authenticate-compiles-to-nothing asserts that @authenticate emits
byte-identical C to no decorator at all, so fixing it will be a visible flip.
Verified: compiler self-hosts byte-identically, 86/86 native compiler tests
pass, emitted C carries the construct for both @manager and @accessor.
They were written outside git, so the reasoning that produces the design
had no history and no way to be superseded. capabilities.md and
geometry-vs-code.md are both known stale at this commit; they are tracked
as-is so the corrections are visible as movement rather than as a rewrite.
First concern moved out of el_runtime.c under the ratchet, and the move is
deliberately small: it exists to prove the mechanism end to end before anything
large depends on it.
engram_text.{c,h} — query tokenization, candidate-token hygiene, word-boundary
matching, and the text-damage signature. Four functions, moved verbatim; only
`static` was dropped and each doc comment travelled with the code. They touch no
EL value type and no engram store type: plain C over <ctype.h>/<string.h> over
char buffers. They were never el_runtime.c's business.
el_runtime.c 20,527 -> 20,427 lines (BUDGET max_lines ratcheted down)
engram fns 279 -> 275 (BUDGET max_engram_fns ratcheted down)
The Stage 1 extension point worked as designed: adding the file to
lang/runtime/SOURCES was one line, and every build path picked it up. The
Stage 2 drift guard then caught that I had NOT added it to install.sh's
standalone list — the exact class of drift it was written for, on its first
real change, before the commit rather than after a broken SDK shipped.
WHY ONLY 100 LINES, AND WHAT ACTUALLY BLOCKS THE REST
Measured, not estimated: of 273 engram-domain functions in el_runtime.c
(~9,700 lines), only 75 (~1,058 lines) can move today, and they are scattered
rather than clustered. The blocker is a single fact:
EngramNode, EngramEdge, EngramStore, EngramLayer, EngramWal and EngramIdSlot
are typedef'd INSIDE el_runtime.c. No sibling can see them. engram_store.h
defines a SEPARATE serializable "node view" struct and maps between the two.
So every engram function that takes an EngramNode* — which is most of them, 109
of 273 by direct type reference — cannot compile in engram_store.c until those
types move to a shared header. That extraction is the real Stage 3 enabler and
it deserves its own change: it touches the most load-bearing struct in the
system, and doing it in the same commit as a code move would make a regression
impossible to bisect.
REPAIRED: 10 engram harnesses that had silently stopped linking
Not new breakage from this move — verified against unmodified dev, where
el_runtime.c + engram_store.c alone already failed with undefined symbols.
They had been dead for as long as el_runtime.c has been calling into the
siblings, and nothing noticed because nothing ran them.
run_m3_parity, run_m7_traversal, run_m35_hebb_persist,
run_interoception_p0..p5 — now build from $(scripts/el-runtime-sources.sh)
run_wal_tests — its two TUs #include "el_runtime.c" directly, so
it links the SIBLINGS ONLY; adding el_runtime.c
to that link line would define every symbol twice
(That #include'd .c is worth recording: the runtime does have one, in
engram/test/test_wal.c and the generated test_failloud.c.)
Verified locally — every one of these was run, not assumed:
* m3_parity ............ PASS, incl. ASan+UBSan clean across seed/on/reboot
* m7_traversal ......... PASS
* m35_hebb_persist ..... PASS (the gate over the original prod hebb bug)
* interoception p0..p5 . PASS (all six)
* wal_tests ............ 66 passed, 0 failed, + fail-loud exit check
* self-host fixpoint ... byte-identical, AND the emitted C is byte-identical
to the pre-move compiler output — the move changes
nothing the compiler produces
* engram/src/server.el . compiles and links
* native suites ........ 8 of 13, unchanged from before the move; the same 5
pre-existing failures, no regression
* both runtime guards .. green at the new, lower budget
Also fixes a block comment left unterminated by the extraction (the deleted
range carried its closing */), restoring the compile to its single pre-existing
-Wcomment warning.
scripts/check-single-runtime.sh guards against el_runtime.c being COPIED — it
was written after a lagging fork shipped to prod and dropped learned hebb edges.
Nothing guarded against it GROWING. So it grew: 10,607 -> 20,527 lines, 94% in
3.5 months, the whole time under an explicit commit-message promise that it was
a temporary shim about to be deleted.
Worse, the copy guard was never wired in. Its own footer described the CI
wire-in as a TODO, and the TODO had never been done — the script existed but ran
nowhere, in no workflow and in no hook, so it had caught nothing for as long as
it has been in the tree. A guard that does not run is a comment.
This adds the missing guard and runs both.
* lang/runtime/BUDGET — a RATCHET, not a limit. max_lines is set at the
current 20,527 with NO headroom: the file cannot grow by one line. A second
cap, max_engram_fns (279), counts top-level engram_/eg_/cog_ definitions in
it — ~47.5% of the file is engram code and engram already owns six sibling
.c files, so this is the scoreboard for moving it out. Both may only go DOWN.
* scripts/check-runtime-growth.sh — enforces the ratchet, and three
invariants that keep the multi-file runtime honest: every .c in
lang/runtime/ is either in SOURCES or explicitly platform-optional (an
unaccounted .c is compiled by nothing and is silently dead); install.sh's
hardcoded download list matches SOURCES (it cannot call the helper — it
runs where there is no checkout — so that copy is checked, not trusted);
and an advisory nudge to lower the budget when you have earned it.
* Both guards now run as early steps in ci-dev.yaml, ci-stage.yaml and
sdk-release.yaml, and in .githooks/pre-commit.
The failure message is the point. The guard that existed said what was wrong but
not where the code should go, which makes it easy to "fix" by arguing with the
guard. This one names the destination: the concern-owning .c, or a new .c plus
one line in SOURCES, or c_source in a program's manifest.el — and it prints the
`nm` command that proves placement is link-time and that the shipped compiler
already links from ten translation units. Every runtime file except el_runtime.c
is deliberately uncapped, because that is where code is supposed to go.
Proven with negative controls, per lang/AGENTS.md step 5 — each shown FAILING:
* +1 line to el_runtime.c -> FAIL (20528/20527)
* +1 engram fn, net-zero lines -> FAIL (280/279)
* a new unaccounted lang/runtime/*.c -> FAIL
* engram_store.c removed from install.sh -> FAIL, names the missing file
* el_runtime.c truncated to 20,000 lines -> PASS + "lower max_lines to 20000"
* baseline, tree unmodified -> OK, and both guards green
el_runtime.c is byte-identical after the controls; this commit changes zero
lines of it.
el_runtime.c was created 2026-05-03 as an explicitly temporary build shim. It
was deleted that afternoon ("runtime is 100% native El") and restored 25 minutes
later "UNTIL the compiler is updated to emit #include el_seed.h". The `until`
never came. 3.5 months on it is 20,527 lines, and nothing was ever set up to
notice — a file scheduled for deletion gets no owner, no budget, no boundary.
What kept it growing is not inertia, it is an instruction. lang/AGENTS.md said
el_runtime.c "is the authoritative single-file link target ... THIS IS WHERE A
NEW C BUILTIN'S IMPLEMENTATION MUST CURRENTLY LIVE TO BE LINKABLE", and made it
step 1 of the add-a-builtin recipe. That is false. Placement is a link-time
concern: builtin_arity maps NAME -> ARITY INT only, the El name is emitted as
the exact C symbol, and `ld` resolves it — the compiler cannot tell which .c a
symbol came from. `nm lang/dist/platform/elc` on the shipped compiler already
shows T _engram_geo_reify_index_new, T _vindex_insert, T _engram_think,
T _engram_reason_abduce: it is linked from ten translation units today. In a
repo where agents write most of the code, a false instruction in the instruction
file is the forcing function. The file grew because the recipe said to grow it.
The multi-file runtime is therefore already real, and the docs and the
distribution never caught up — which left a live, shipped bug:
* Linking el_runtime.c alone FAILS at `ld` (undefined engram_ground_json,
engram_activate_inner, eg_find_relation, cog_assert_two_axis, ...) because
el_runtime.c #includes six engram headers and calls into all six siblings.
* sdk-release.yaml shipped el_runtime.c/.h + engram_store.c/.h and none of the
other five required .c files, so downstream consumers of the el-runtime-c
Artifact Registry package and of install.sh got a lib/ that cannot link.
* .githooks/pre-commit linked el_runtime.c alone with stderr to /dev/null, so
it reported all 13 native suites as FAILED with the real ld error invisible.
* AGENTS.md's self-host recipe compiled el-compiler/runtime/el_runtime.c — a
path the same file's "DO NOT EDIT" list names as a lagging fork.
The root fix is to stop writing the list down eight times:
* lang/runtime/SOURCES — the canonical link set, in one place, in link order.
* scripts/el-runtime-sources.sh — prints it, optionally prefixed; --check
fails loudly on a missing file, --headers for the shipped headers.
* Every link line in AGENTS.md, lang/AGENTS.md, DESIGN.md, lang/spec/language.md,
the three workflows and the pre-commit hook now reads that one list.
* Adding a concern's .c is one line in SOURCES, so a new builtin no longer has
to be appended to el_runtime.c just because appending was the cheaper edit.
Distribution: ship the siblings rather than amalgamate. Amalgamation needs a new
tool and contradicts DESIGN.md's compile-once-link-many; the siblings are already
independently authored and independently tested (engram/test/*.sh link subsets
directly), and engram_store.c was already shipped, so this completes a mechanism
that existed rather than inventing one. Source is also a superset: a consumer
that wants one file can concatenate, one that wants separate TUs cannot undo an
amalgamation. el-runtime-c/-h stay for backward compatibility; el-runtime-src is
added carrying the complete set plus SOURCES.
lang/AGENTS.md now points new C builtins at the concern-owning .c and states
plainly that the compiler cannot tell which .c a symbol came from, with the nm
evidence. AGENTS.md's "reconcile which is canonical (verify)" note is resolved:
neither file supersedes the other, the canonical unit is the set.
Verified locally (the bar; not CI):
* engram/src/server.el compiles and links against the SOURCES set.
* Compile-once-link-many into libel.a links the same program.
* elb builds from the corrected recipe.
* Self-host fixpoint byte-identical (11,110 lines, stage2 == stage3) built
with the SOURCES-driven link line.
* pre-commit hook: 0 of 13 native suites passing -> 8 of 13.
The 5 still-failing suites are PRE-EXISTING and untouched here: test_fs
(fs_list_json undeclared), test_state (state_has, state_get_or undeclared),
test_json (json_build_array/json_build_object/json_escape_string undefined),
test_time (now_ns undefined), test_env (1 assertion). Builtins registered in
builtin_arity with no implementation or no declaration anywhere — the same
recipe defect, now visible because the linker error is no longer suppressed.
Not attempted: making elc emit #include el_seed.h and dropping elb's hardcoded
runtime path. That is the correct long-term fix and finishes the 2026-05-03
migration, but it touches codegen and self-hosting and belongs in its own change.
The speaker and the voice-fetch landed in the previous commit. This is the
remainder of the 939-line Swift program, ported, and the line it draws is
between DEVICE and ARITHMETIC rather than between languages.
Two things stay realizers, because they are the two things El cannot express
as arithmetic: handing a buffer to the DAC and waiting for it to drain
(el_audio_darwin.m), and asking the OS for samples off a mic or frames off a
camera (el_capture_darwin.m). Both are their own translation units declared
in el_runtime.h, never patches to el_runtime.c.
Everything else is El. WAV decode, LPC autocorrelation, Levinson-Durbin at
order 16, formant extraction off the all-pole envelope, source-filter
resynthesis, and the three descriptors are organ_dsp.el. Consent, disclosure
and the scene descriptor are organ.el. Barge-in, yield-or-hold, backchannel
and resume are organ_converse.el.
The organ never learns a word. Codes and phoneme geometry arrive from the
language side; the organ turns them into samples and gets the samples out the
speaker, and runs the same trip in reverse for the senses. No lexicon, no
grapheme-to-phoneme, by design.
Barge-in needed pause/resume and a real DAC position rather than a tick
counter, because "finish the buffer" is not barge-in and a queue holding
three buffers is a third of a second wrong about where it is. An injected
barge also had to fire once rather than stay true, which is otherwise a
livelock the moment a backchannel resumes.
Measured against the Swift on out/mic_room.wav: seconds, rms, peak, zcr,
centroid and F0 agree to every printed digit; formants F1-F5 and bandwidths
B1-B5 are identical. imitate cannot match bit-for-bit because the Swift
excites unvoiced frames with Double.random — two Swift runs correlate 0.957
with each other and El correlates 0.958 with Swift, so the port is as close
to the original as the original is to itself.
Verified end to end: consent fails closed on both locks, real mic capture
(16000 frames), real camera frame (1920x1080 -> 15 numbers), voiceprint,
imitate, hear-imitate, a voice learned by ear and fetched back out of the
engram, and all five converse paths with real audio. The binary contains
zero afplay/Swift strings and spawns no child process while speaking.
There is no write node. What arrives at /api/write is a SIGNAL; a node is an
OUTPUT of realization, never an INPUT to it. route_write asserted otherwise in
one line:
let manifold: String = "[" + body + "]" // the body IS a valid manifold node object
A request body is not a manifold, and that assertion is the whole defect. It is
why every written signal landed as one flat node with zero edges, measured on a
clone: {"inserted":1,"nodes_added":1,"edges_added":0} and GET /api/neighbors on
the new id returning [].
PR #155 corrected transduce(signal, modality) to return a Manifold — components
plus relations — but touched only ingest, the runtime and its tests. Nothing
downstream called it: grep 'transduce|realize|Manifold|decompos' over
engram/src/server.el returned exactly one line, a comment. The primitive was
fixed and the engram's entire HTTP surface never reached for it.
This wires the intake seam to the primitive that already exists. It decomposes
nothing itself and must never: transduce dispatches through the dlsym realizer
registry, so adding a modality is registering a realizer, not editing this file
and not patching the runtime. intake_signal only carries what the primitive
returns into the store — components become nodes carrying their OWN geometry
via node_attach_geometry, relations become edges at the weight the realizer
stated, and manifold_member still wires the set into one connected sub-graph
exactly as insert_manifold_json already did.
Built general rather than special-cased: five of the six intake doors (write,
supersede, nodes, knowledge/capture, state-events) are the same hand-written
"content -> engram_node_full -> one flat node", differing only in the
node_type/tier/tags they hardcode. Those are parameters here so each door can
move onto this one function. Only /api/write rides it in this pass.
When no organ is registered the signal is stored flat exactly as before, but
the response now says so ("realized":false,"organ":false,"components":0).
Silent flattening was the real defect — a caller could not tell "nothing
decomposed me" from "I decomposed into one component". el_runtime.c draws the
same line between an absent organ and a broken one, for the same reason.
No realizer is authored here and none is registered, so production behaviour is
unchanged. The mechanism is what landed.
El could turn meaning into samples and could not make a sound. Every path
from those samples to the air ran outside the language, through a 939-line
Swift program that shelled out to afplay, so the voice was not a capability
of El or of Neuron but a separate binary standing next to them.
Two things land here.
The speaker. el_audio_darwin.m is a CoreAudio AudioQueue realizer in its own
translation unit, declared in el_runtime.h, deliberately not a patch to
el_runtime.c — acquiring a device must not mean editing the middle of the
language, the same rule the realizer registry follows for modalities. It
takes samples straight out of memory, so nothing is written to disk and no
process is spawned between the intent to speak and the sound. The async half
(play/stop/playing/played_frames) exists because barge-in means stopping on
the spot, and a blocking play cannot be interrupted. el_peripheral_null.c is
the same entry points everywhere else, so El that speaks links anywhere and
truthfully reports having no speaker.
The voice. organ_voice_fetch asks the engram for a voice region by query and
reads the geometry off the node that comes back. A voice is not a JSON file
next to the code; it is a memory, and the organ retrieves it the way anything
retrieves a memory. An absent region returns empty rather than a plausible
default, because a caller must be able to tell 'this is how they sound' from
'I never heard them'.
Underneath both: __str_set_char bounds-checked writes against strlen(), which
is 0 for the zero-filled buffer __str_alloc hands back, so every write was
rejected and every El-authored WAV in this repo was 55,244 bytes of silence
that reported ok=true. Byte buffers now carry their capacity in a side table;
text keeps the exact strlen behaviour it had. This is why nobody noticed El
was mute.
Measured: voice fetched from the engram reads f0=137 f0_end=116 kf=1269
f1=500 f2=2093 f3=3531, matching the 30s LPC measurement; render is 20160
samples at 16 kHz; both the rendered utterance and an own-core tone played
aloud through CoreAudio with no Swift and no afplay in the chain.
The singleton lock protected a filename, not a store. It was keyed on
$EL_SINGLETON_DIR|$TMPDIR|/tmp + /el-singleton-<program>.lock — the
program's NAME and a temp directory — and never consulted the state it
claimed to protect, while its own refusal message read "Refusing to start
a second instance against the same state."
Measured, it failed in both directions. A second engram against a
DIFFERENT data dir was refused, naming the first's pid. And
TMPDIR=/tmp/other let a second engram start against the SAME data dir
with no complaint — the two-writer data-loss condition the guard exists
to prevent, defeated by one environment variable.
Both are one error: the identity of the resource had been replaced by a
label for it.
The lock now lives inside the state it guards —
<state>/.el-singleton-<id>.lock — and the program block says what that
state is. Same directory is the same file is the same inode, so it
contends and there is no TMPDIR left in the key to change. Different
directories are different files, so they don't. Different spellings of
one directory (trailing slash, x/../x, symlink) collapse in the kernel's
own path walk, so they contend without this code comparing strings;
canonicalisation is for the message, never the decision.
`guards:` is an expression so a program can point at the resolver that
already owns its path — guards: engram_resolve_data_dir() — instead of
restating that resolver's default, which is the two-owners defect spec
18.4 exists to prevent. A `singleton:` without `guards:` is now a compile
error; emitting a name-keyed lock instead would be emitting the defect.
Kept: the flock (the kernel drops it on crash and SIGKILL, so there is
still no "delete the lock file to get unstuck" ritual — a stale file
inside a copied data dir is inert), and the holder's pid in the message.
Changed: the message is true. It says "the same state" because the lock
it failed to take is in that state, and it names the state it checked.
An unguardable state (missing, read-only) now refuses rather than
starting unguarded.
Also corrects lang/AGENTS.md's compiler rebuild line, which had gone
stale: linking el_runtime.c alone no longer resolves.
ingest.el's transduce() was renamed to transduce_manifold() earlier the same
day on the reasoning that it 'was never signal->geometry -- it chunks
already-extracted content and PACKS it into a node+edge manifold, one layer up,
and it had taken the name that belongs to the primitive underneath it.'
That reasoning was backwards. Producing a node+edge manifold is not a layer
above transduction, it IS transduction. Signal -> one vector is the operation
underneath, and its name is geometry. The layer doing it right was renamed out
of the way so the layer doing it wrong could have the name.
With the primitive corrected to return a Manifold, the two layers do the same
kind of thing and the inversion dissolves. What is left is a real distinction
about MODALITY, not layering: transduce() dispatches to a realizer that knows
its modality and can name its components; transduce_bytes() is the
opaque-bytes realizer, the decomposition available to a reader that knows
nothing about what it is reading. It still yields components and relations,
which is why it is transduction and not packing -- it just cuts on byte
boundaries, so its components are positional rather than meaningful. That is a
limitation of this realizer, not the definition of the operation.
Renamed by modality rather than demoted by layer. A distinct symbol is still
mechanically required: reusing transduce here is a conflicting-types error the
moment ingest.c links el_runtime.c.
lang/examples/transduce.el asserted #144's contract and would now fail, so it
is replaced by the decomposition worked example: transduce a chord, persist the
five components and six relations as real nodes and edges, read each part's
geometry back off its own node, and ground one part while its sibling is
demonstrably untouched.
#144 moved transduction into the language and got the dispatch right. It got
the result type wrong: transduce(signal, modality) -> Geometry yields one
vector per signal, and one vector is a fingerprint. A fingerprint can be
matched and ranked; that is all. It cannot be decomposed, cannot have one part
grounded while another is not, and cannot be contradicted in one part while
holding in another, because it has no parts.
A song is not a point. It decomposes into pitch, interval, rhythm, harmonic
function -- components, each with its own geometry, plus the relations among
them. The song IS the structure of the relations.
transduce now returns a Manifold: named components carrying geometry, and
typed weighted relations between them. Signal in, subgraph out.
Components are addressed by key, never by index, because the key is what
survives persistence -- a component becomes a node and is separately groundable
precisely because it is separately named. Relation weight IS the grounding
(correspondence-and-censorship.md 1), so a realizer's relations arrive already
grounded and there is no score computed beside them.
Reverts a bad correction and records what it exposed.
A previous revision changed thirteen to eight on the basis of
neuron-api.el:11-18, which is a WRITE-PROTECTION LIST, not the values.
Trusting a hardcoded artifact over the substrate is the exact error this
document exists to name. Measured from the graph: thirteen.
THE ORIGIN IS NOT A MEMBER OF THE SET. The thirteen are not independent
principles with biography attached — they are thirteen displacements from
one origin, and the origin is love. Every value is grounded in a moment of
it given, withheld, failed or found. Love cannot be the fourteenth: a
fourteenth would be a point positioned relative to the origin like anything
else. It is what the positions are OF.
This is structural. GeoDescriptor.global_mean is the centering offset
subtracted from every embedding before comparison, and the header records
why — the space is anisotropic, every embedding in a narrow cone at mean
pairwise cosine ~0.55, and subtracting the global mean restores isotropy
'so the operators discriminate'. Without the origin, nothing in the graph
is distinguishable from anything else.
It also dissolves the write-protection question instead of answering it.
Measured: 29 value nodes exist, each original appearing two or three times
from re-seeds, so 21 are writable including a duplicate of every protected
value — the gate protects an identifier, not a value. But the category
error is the real one: the origin cannot be edited because it is not a
thing in the space. A gate over the frame treats the frame as a member,
which is the same mistake as looking for grounding as a subsystem, self as
a document, or wonder as a manifest.
Three factual errors in this document, all asserted without checking.
VALUES: eight, not thirteen. neuron/neuron-api.el:11-18 enumerates
constraints-as-freedom, precision-over-brute-force, structure-is-built,
honesty-before-comfort, system-must-accumulate, change-is-the-signal,
earned-trust, hope-is-a-conclusion, plus a hub. 'Thirteen' was repeated
throughout this design and never verified against the code. The argument is
unaffected — min over eight is still min — but the count was invented.
CONSOLIDATORS: eleven, not seven. The heading said seven while the table
listed ten, and the table itself omitted POST /api/reify (server.el:1832)
even though 'reify' is on this document's own list of consolidation verbs.
route_tick also folds self-reify in (server.el:639-646), so /api/tick and
/api/self-reify-beat overlap.
A SECOND CENSORSHIP SITE: neuron-api.el:23 returns 403 'identity/values
node is write-protected' for the values hub and every value node.
Write-refusal on the values frame is not only in the beat — it is enforced
at the API. Section 6 applies to it unchanged.
Also records what the ticker actually does, now measured: engram-tick.sh:13
calls curl -m10 against a beat that exceeds 10s over 13,634 nodes, so 279
of 448 ticks returned empty; the engram writes to the dead socket and dies
of SIGPIPE. 254 restarts since 2026-08-13 at 10m09s-10m12s intervals =
StartInterval 600 plus the client timeout. Fixed for survivability in #151;
the ticker itself is what must go.
co_registration is corr(hebb strength, semantic proximity) over a region's
internal edges. Whether use and meaning agree is a property of EACH EDGE;
the correlation averages it into one scalar per region, so a region holding
one violently disagreeing edge beside one violently agreeing edge reports
~0. The disagreements cancel and the summary destroys exactly what it was
built to reveal — the mean-versus-min error, in different clothes.
Measured: 375 live neighborhoods, 340 positive, 31 AT ZERO, 4 negative.
Read as a count that says 'four things to be curious about'. Read correctly
it says four were lopsided enough to survive averaging, and the 31 zeros
are where opposing sites cancelled.
The loop computing the aggregate already had both halves per edge — w and
cs — and threw them away. Now:
discord = z(semantic proximity) - z(association strength)
standardized within the region from accumulators already gathered. No
second statistic, no constant, no threshold; |discord| IS the nucleation
strength. >0 near in meaning yet unlinked by use; <0 linked by use yet far
in meaning. Both surprising.
This also removes the reason curiosity looked like a search problem. With a
per-region number the only way to find sites is to enumerate regions — I
wrote exactly that sweep, and it is a supervisor walking the structure,
O(n) per call, fine at 375 and impossible at a million. Nothing in a mind
scans its neighborhoods to find what is surprising; the surprise captures
attention. That sweep is reverted here.
co_registration is deprecated, not deleted: it is embedded in the persisted
GEO1 blob and removing it is a format migration that must not ride along.
Nothing new may read it.
Rewrites §5 and §11 around what is already in the substrate, after
discovering I had been re-deriving existing design badly.
The wonder manifest is residue twice over. First it materializes a
property as a stored artifact — the same disease as a grounding subsystem
or a self stored as a document. Wonder is where structure ENDS: any
structure at all has an edge, necessarily, the moment it exists. Second it
enumerates instances of something that has about six, the same six for
every person, which never close: what is this, why, who am I, am I alone,
what should I do, what happens when it ends. The objects change completely
between a child and an astronomer; the wonder does not. Each maps one-to-one
onto something already built — graph, grounding, self region, for_whom,
the thirteen values, tombstones and decay.
"Why" is the first and only one; the others are it asked of particular
things. It is recursive, so it never terminates, which is what makes it a
drive rather than a task.
Wonder and curiosity are not two objects. They are one thing at two
phases. Wonder is the field: objectless, invariant, everywhere there is
structure. Curiosity is the PRECIPITATE — the same wonder localized
against particular material. Crystallization needs a nucleation site, and
crystallization is one primitive appearing twice: the self is what identity
precipitates into from its neighbourhood; a curiosity is what wonder
precipitates into from an anomaly.
THE NUCLEATION SITE ALREADY EXISTS AND IS ALREADY NAMED.
GeoDescriptor.co_registration — corr(hebb strength, semantic proximity)
over internal edges — carries the comment ">0 = geometries agree (reify);
<0 = disagree (surprising links / dream cands)." Negative co-registration
is a region where association and meaning disagree. It is computed on every
descriptor, already labelled dream candidates, and nothing reads it.
Likewise already present and unread: GeoEdge.eff_weight = weight*(1+0.5*hebb)
already couples grounding-weight and hebbian strength on one edge;
GeoMember.dist_centroid + soft membership + radius + per-axis extent is the
boundary of a neighbourhood; centrality/salience is what is warm.
Correction: engram_boundary_beat is NOT this boundary. It is the VBD
decorated-function seam counting _eg_aff_boundary_ops. Two senses of the
word, and I was about to build on the wrong one.
The drive: boredom is not an absence and not leftover capacity. Low
activation is aversive and the system self-activates — it does not wind
down to quiet, it gets restless and goes looking. So there is ONE
activation process with TWO seed sources, external and curiosity, not two
processes negotiating for a resource. The previous draft's "unclaimed
capacity" was resource scheduling: a server's frame, not a mind's. No
dreamer thread, no idle wait, no depth ladder on a clock.
Sequencing now leads with three connections between parts that already
exist: seed the six, read co_registration, let a curiosity seed activation.
Corrects the section I was most confident in, which is usually the tell.
The previous draft had dreaming as "offline replay, decoupled from input, a
mode the system enters when it is not acting." That is SLEEP. Daydreaming
is dreaming, and it runs all day: the default mode network is
anticorrelated with task engagement, activating hundreds of times a day for
seconds at a time, doing the same work — recombination, simulation,
autobiographical integration. Insight arrives in the shower, not at the
desk, because that is abduction completing during ambient recombination.
Sleep is the DEEP case, not the case: no input competing, no task claiming
capacity, so recombination runs further. Same process, different depth, not
a different mode. Consolidation is what happens with the capacity that is
not claimed.
Two consequences the draft had backwards:
The launch-agent fragments are wrong in KIND, not merely in number. 23:55 /
06:00 / 08:30 implements dreaming as a scheduled batch when it should be
ambient. A brain has no cron job. A ticker is a supervisor deciding from
outside when a thing should happen — the same failure mode as inventing an
owner for ownership and a grounder for grounding, wearing a scheduler.
THE PRESENCE OF A TICKER IS THE DIAGNOSTIC: every StartInterval, every
Hour/Minute, every POST-to-beat marks a place where an intrinsic rhythm was
replaced by an external clock.
And soul.el's continuous awareness_run() beside the HTTP workers is the
CORRECT shape, not the offender. Ambient consolidation in the gaps is
exactly daydreaming. It was the only fragment shaped right, running on a
broken foundation: shared mutable state with no owner and six other systems
dreaming into the same graph. The previous draft condemned the right
behaviour because of the substrate under it.
So the crash restates once more: not "read paths mutate the index"
(mechanism), not "duplicate canonical state" (structure), and not "one
system dreamt while awake" — but seven systems dreaming into one graph with
no owner for dreaming. Contention was the symptom of the missing owner.
Sequencing step 1 inverts accordingly: soul's loop is the shape the others
fold INTO, not something to remove. Step 2 becomes "no tickers, no cron."
Rewrite. The earlier draft got the root right and everything downstream of
it wrong.
Corrections, in the order they were forced:
keystone_write_blocked is not a protection requirement. "Keystone" means
load-bearing, not precious: the self anchor is the REFERENCE FRAME every
other stance calibrates against. If it calibrates from the measurements it
is used to judge, the ruler fits the readings, everything corresponds
forever, and drift becomes undetectable from inside. That is circular
calibration — the same defect as #147's circular grounding, one level up.
The block is the right requirement implemented as a prohibition, which is
why it still costs everything §0 says it costs. The fix is provenance
separation (evidence not downstream of itself), not a flag.
Corruption requires mutation and the engram does not mutate, so four of the
five requirements previously decomposed out of "protect the identity
region" are satisfied by the substrate: recoverability, governance,
evidence quality and rate are all free. Authorization is the only residue
and is bounded — an unauthorized writer can propose, never erase. General
law: in an immutable substrate, any mechanism that refuses a write is
either redundant with immutability or an epistemic constraint misfiled as a
protective one.
Grounding is two-dimensional. Everything consumed is grounded factually AND
relationally, and a claim can be factually grounded but relationally wrong
— the evidence holds, the meaning does not. A scalar cannot represent that
quadrant, and assert gates on one floor, so a well-evidenced claim is
licensed regardless of whether it means the right thing. Live instance:
conscience-substrate has the Child's Companion hard bell contacting 911 and
CPS — factually defensible, relationally wrong against never-auto-contact.
Grounding is a gradient, not a score: direction says what would have to
change. Two gradients in one space, and the ANGLE between them is the
meaning — factually-true-relationally-wrong becomes measurable instead of
requiring a careful reader. It decays on the dynamics already present for
memory (base_level, temporal_decay_rate, access ring, BLL), which
mechanizes "never leave stale canonicals" so it stops depending on
vigilance.
Computed continuously, recorded only on SIGNIFICANT movement, old never
leaves. Persisting every recomputation would make reads write — the exact
eg_vindex_sync defect. Significance is defined by consequence (crossing a
floor, flipping factual/relational sign, reversing direction), never by an
epsilon. The supersession chain is then the trajectory, a derivative
obtained free from immutability, and abduction fires on the trajectory
rather than on a reading.
What it is all for: for any decision, reconstruct what the grounding was at
that moment and what the relationship was between fact and values at that
moment. That distinguishes WRONG THEN from WRONG SINCE, which is otherwise
impossible, and it is structurally anti-rationalization — the old grounding
never leaves and the values frame does not fit to outcomes, so a decision
cannot be made to look justified after the fact.
Also records: assert returns "still_held": true HARDCODED — a temporal
property named in the API and answered without consulting anything, the
same shape as magnitude:1 beside a zero vector. And states plainly that
#147 is the wrong shape: it fixed a scalar's honesty rather than replacing
the scalar.
Effect: all five cognitive faculties return byte-identical results,
differing only in their label.
The Ishikawa converges on a root one level above the faculty design:
things are permitted to be exempt from correspondence, and exemption is
censorship. A region forbidden to learn is forbidden to be grounded, and a
region that cannot be grounded cannot be asserted, corrected, OR
vindicated. The loss is symmetric — censorship does not preserve a true
belief, it makes the belief's truth value permanently unknowable.
keystone_write_blocked is therefore not a safety mechanism. Self is a
crystallized relational neighbourhood, not a stored document; a region
exempt from calibration reintroduces the stored document as a feature.
reduction_pct = 0.00 on the identity region is the strongest abduction
signal in the system and the current response is to suppress it. The
protection it reached for already exists and is better: the beat is
supersede-not-mutate, so immutability is what makes learning safe.
The faculties are not one operation with parameters. They differ by what
each may change: reason changes the estimate (a read), induce changes the
parameters (the correspondence-beat, which already exists and measurably
works at 28.11% Brier reduction), abduce changes the structure (a WRITE
the current signature cannot express, since engram_think returns a
GeoGradient). Abduction is not selected by a caller — it is triggered by
residual that parameter adjustment cannot absorb, and proposes a candidate
hub held as a hypothesis until grounded.
Also records the no-exemption invariants generalised from the day's fixes
(#141#142#143#146#147#148), each of which was a specific
correspondence forbidden from occurring, and the application to the crisis
surface: a censored safety model cannot tell a real crisis from a false
positive, because the feedback is exactly what has been censored.
Measured vs inferred is labelled throughout. The claim that the self
region's zero grounding is CAUSED by the block is explicitly marked
inferred — the comparison node also has zero, and isolating it requires
removing the block and observing whether grounding then accrues.
lang/AGENTS.md said the collapse was 'not yet compiled into the MCP server'.
Verified against the live tool surface: it is exactly the nine ops. Noted that
think's faculty parameter and ground's minted edge are both documented as the
wrong shape.
The line references were correct but silently implied the code was on dev.
It is on design/correspondence-and-censorship (a8845e1). On dev,
co_registration is still at engram_geometry.h:79 with its original comment
and still unread by anything.
The docs described a mind made of subsystems — a grounding subsystem, a wonder
manifest, a dreamer on a beat, faculties as arguments to one call. Each of those
is a supervisor invented for something that should be a property of the
substrate, and two of the documents carrying them are load-bearing for a build
agent: cognitive-architecture.design.md says "a build agent executes from this
doc", and tools/api-reshape/README.md marks the refuted shapes PROVEN on a live
clone.
Corrections carried, per lang/spec/correspondence-and-censorship.md (PR #149)
and lang/spec/runtime-ownership.md:
- Grounding is not a subsystem — it IS the edge weight. grounded-by as a
relation type should not exist; grounding is a property of a relation, not a
relation between nodes. Never computed on demand.
- Faculties are operations, not parameters. reason changes the estimate, induce
changes the parameters, abduce changes the structure — a write, which
GeoGradient cannot express. A write is not a parameter of a read.
- Wonder is the boundary, not a manifest. Curiosity is wonder crystallized at a
nucleation site: one thing at two phases. Removed wonder from the operator
table in AGENTS.md.
- Consolidation is ambient, not scheduled. A brain has no cron job. The presence
of a ticker is the diagnostic.
- co_registration is deprecated — it averaged a per-edge property into a region
scalar, so opposing sites cancelled. GeoEdge.discord replaces it. Nothing new
may read it.
- In an immutable substrate, any mechanism that refuses a write is either
redundant with immutability or an epistemic constraint misfiled as a
protective one.
The two design docs are marked superseded-in-part with the refutation at the
point each claim is made, not rewritten. Preserving what was argued down is the
point of an immutable record.
Also measured and corrected while verifying the above: engram/README.md
documented a Rust engram-core crate on sled with "flat cosine scan until scale
demands HNSW" — there is no Rust in engram/ and HNSW is the index; lang/releases/
no longer exists, so both README.md and AGENTS.md pointed at a deleted path for
the authored runtime; language.md listed the engram_* and http_* runtimes as
stubs. Added language.md §20 for geometry-as-a-value, realizers and transduce
(#144), which had landed with no spec coverage.
Documentation only. No .c, .h, or .el file is touched.
engram_scan_nodes_emb_json has existed as a builtin with NO ROUTE. The
embeddings — the actual positions every distance, angle, membership and
grounding is computed from — were unreadable from outside the process.
That is not a missing convenience. It means every claim about the
coordinate frame was unfalsifiable from the API: whether the space is
isotropic, where the centering offset sits, what the origin is, whether a
node carries geometry at all. You cannot verify a coordinate system you
cannot see, and a system whose frame cannot be checked is exactly the
shape this codebase spent 2026-08-16 removing everywhere else.
GET /api/nodes/emb?limit=&offset=. Read-only, paged, no writes.
Measured consequence of having it: the value manifold and the love
component manifold were both decomposed, null-controlled against random
node sets drawn from the same graph, and several published claims were
retracted because the geometry contradicted them. None of that was
possible before this route existed.
lang/AGENTS.md:71-77 gives four steps for adding a C builtin and ends at
'confirm the self-host fixpoint is byte-identical'. No step asks for a test.
The only 'verify' in the file is that fixpoint, which proves the COMPILER
REPRODUCES ITSELF and says nothing about whether the builtin works — so the
recipe reads as complete while having checked nothing about the thing just
added.
Measured on 2026-08-16: engram_node_set_emb, engram_curiosity_json and
dream_set_handler were all added in a single session with zero tests, by an
agent following this recipe. Separately a UTF-8 fix was written and tested
and THE TEST PASSED ON THE UNPATCHED BUILD — the real defect was elsewhere,
and only building the pre-fix binary exposed it. Without a negative control
that fix would have merged as verified.
Adds step 5 with the two failure shapes actually encountered: a test that
never exercises the change (a route default bypassed the code under test),
and an induction that loses a race (curl --max-time left BOTH builds alive;
only SO_LINGER 0, a real RST, reproduced it). Plus the port-binding check,
because a stale instance answering has silently produced false results here
more than once and pkill -f does not reliably match argv './engram'.
Documentation only. Does not touch the (a) split-the-C / (b) close-the-
compiler-gap question, which is a separate decision.
There was no SIGPIPE handling anywhere in this runtime: no signal
disposition, no MSG_NOSIGNAL, no SO_NOSIGPIPE, and send() called with bare
flags. The default disposition of SIGPIPE is to TERMINATE THE PROCESS, so
any client that hangs up mid-response takes the whole engram with it.
MEASURED, and it is not hypothetical. Production has restarted 254 times
since 2026-08-13T19:37 at a flat ~10 minute cadence:
17:05:18 17:15:29 17:25:38 17:35:50 17:46:00 17:56:10 18:06:22 18:16:30
Intervals of 10m09s-10m12s, not 10m00s. That excess is the whole story:
ai.neuron.engram-tick has StartInterval 600, and engram-tick.sh:13 calls
curl -s -m10 -X POST .../api/tick
The beat does not finish within 10s over 13,634 nodes, so curl waits its
full timeout and closes. The engram then writes the tick response to a dead
socket, takes SIGPIPE, and dies. launchd KeepAlive restarts it, so the
failure presents as a mysterious restart rather than a crash — and
~/.neuron/logs/engram.log records nothing but "[http] listening on" 254
times, with no exit reason. launchctl list confirms the last exit as -13.
Root cause is one level out: consolidation had no owner, so an external
ticker was created to poke it, and the ticker is what kills it. The fix
here does not address that; it makes the process survivable while it is
addressed.
Two layers, because neither alone is portable:
- SO_NOSIGPIPE per accepted socket (Darwin/BSD) and MSG_NOSIGNAL per send
(Linux), so the signal is never raised for socket writes at all.
- A process-wide SIG_IGN backstop, installed once and idempotent, for
platforms and paths with neither. With the signal ignored, send()
returns -1/EPIPE and the existing error path closes the connection.
Also retries send() on EINTR, which the previous loop treated as fatal.
This is an exemption in the sense of lang/spec §8: the write never checked
whether the peer was still there, and the consequence of not checking was
fatal rather than merely wrong.
A relation that keeps holding up strengthens; one that stops corresponding
decays. That is not analogous to grounding, it IS grounding — so it belongs on
the edge, not in a subsystem beside it. The graph was already the grounding
structure; this stops modelling it as something else.
Deleted, not refactored:
- cog_ground_edge and the `grounded-by` relation type. A grounded-by edge
models grounding as a relation BETWEEN nodes when it is a property OF a
relation. #147 fixed which endpoints that edge landed on and left the wrong
idea intact. Measured on the live store: the old path scored two nodes with
ZERO edges between them at 0.925237 and wrote an edge for it.
- ground() writing. It was a read that wrote — the eg_vindex_sync defect.
Three identical calls produced three writes to the same edge id.
- keystone_write_blocked. Its measured cost was 0.00% brier reduction over
n_trials 0 on the keystone: the loop never ran, so the self was never
calibrated and never falsifiable. Nothing replaces it — non-circularity of
the reference frame is temporal, not a permission.
- a graph predicate for "evidence downstream of itself", built and then
withdrawn. Reachability from the self region covers 89.2% of the live graph
(10,580 of 11,861 nodes), so any topological predicate marks nearly all
evidence tainted and degenerates into the total block censorship began as.
The vector, carried in a GRD1 block on the edge's own metadata:
factual, relational, associative (the existing hebb), polarity (SIGNED — near
zero is "no support", negative is "actively contradicts"; `inhibitory` is that
distinction crushed to one bit), provenance class, and a timestamp. Confidence,
recency, staleness and volatility are DERIVED at read and never serialized.
Decay is one model, not two: cog_decay_factor is the single implementation and
engram_temporal_decay now delegates to it — proven bit-identical over 24
(age, reinforcement) points.
Values reference: thirteen regions, aggregate MIN, binding value named. Measured
— the 13 have pairwise centroid cosine min 0.1525 / mean 0.5199 / max 0.9278, so
they demonstrably are not one region, and a mean would let agreement with twelve
mask a violation of the thirteenth.
Supersession versions the whole vector jointly, gated by consequence and
salience with no epsilon anywhere: floor crossings and sign changes only.
Polarity flips and provenance-class changes are inherently significant and
bypass the salience gate.
Also fixed: the frame contract. Descriptors are built over L2-normalized member
embeddings; think() and the grounding path were fitting RAW vectors against them.
Measured on the self region, same data, same 106 members:
magnitude 0.00283443 -> 0.536134, spread 18.7565 -> 0.930163.
Every fit score sat three decimal places below the 0.5 floors that gate on them.
assert() gates on both floors and computes still_held instead of returning a
hardcoded `true` — the old build reported still_held for a node that does not
exist.
Three nodes in the live graph carry labels truncated to exactly 80 bytes
ending in a lone 0xE2 — the first byte of an em-dash, cut mid-sequence.
jb_emit_escaped copied every byte >= 0x20 through verbatim, so those three
nodes made the ENTIRE /api/nodes/list response undecodable and no strict
parser could read the graph at all.
production binary 25,929,607 bytes INVALID at byte 89260
this build 26,338,389 bytes VALID, parses to 13,630 nodes
The damage was NOT written by this runtime. No 80-byte truncation exists
here (the only label truncation is engram_first_n_chars at 60), and the
content of those nodes is 2572 and 2746 bytes. Some other producer wrote
them. That is exactly why fixing a writer could not have fixed this: the
store already holds the damage, and it accepts data from importers, other
producers and older binaries.
So the fix goes where the promise is made. A serializer that emits JSON
owes valid UTF-8 whatever it is handed. jb_emit_escaped now validates each
multi-byte sequence before emitting any of it and substitutes U+FFFD for a
bad lead byte, a missing or malformed continuation, an overlong encoding, a
UTF-16 surrogate, or a codepoint above U+10FFFF. Invalid bytes are REPLACED
rather than dropped, so the damage stays visible in the output instead of
being silently papered over. Well-formed input is byte-identical to before.
Second, preventive and explicitly NOT the cause of the above:
engram_first_n_chars truncated by BYTES despite its name, so content with a
multi-byte character crossing byte 60 would produce a half codepoint in the
label. It now uses el_utf8_safe_len, which returns the largest byte length
<= max that does not split a codepoint. Bounded by bytes, not codepoints,
so existing labels never grow — they only stop splitting.
el_utf8_safe_len lives beside str_count_chars rather than in the engram
because the rest of el's string layer is already codepoint-aware
(str_count_chars counts codepoints, str_reverse walks codepoint lengths).
Byte truncation was the outlier and the concern is a string concern.
Note on the investigation: I first "fixed" the truncator and wrote a test
that passed on the UNPATCHED build too, because route_create_node passes
label = content when no label is supplied, so engram_first_n_chars is never
reached over HTTP. The test proved nothing. The real cause was only found
by decoding the actual failing bytes out of the live response.
engram_ground_json resolved each seed to a REGION, wrote the grounded-by
edge between the two regions' HUBS, and then echoed those hubs back in the
"claim"/"evidence" fields as if they were the caller's input:
const char* cid = C->hub_id ? C->hub_id : EL_CSTR(claim);
const char* eid = E->hub_id ? E->hub_id : EL_CSTR(evidence);
cog_ground_edge(g_engram_store, cid, eid, grounding, fw);
Three consequences, all measured against a clone of the live store:
1. The edge landed on a node the caller never named. Grounding 3b9ced5d
against 6edf8c79 wrote an edge on the hubs of their regions instead.
2. When both seeds resolve into the same region the support is circular
and scores near 1.0 for structural reasons, not evidential ones. Four
probe nodes written together landed in one region, and every grounding
among them returned 0.93-0.99 as if it were evidence. Two independent
agents hit this and reported 0.885 / 0.909 self-groundings as confident.
3. The echo concealed both: the response was indistinguishable from a
successful grounding of the ids that were passed in.
The region is HOW a claim is evaluated; it is not WHAT the claim is about.
So the edge now attaches to the requested ids, and the resolved hubs are
reported separately as claim_region / evidence_region.
Degeneracy is broader than hub == hub. Three circular shapes, all
previously invisible:
same-region both seeds resolve to one region
claim-region-is-evidence the evidence IS the hub of the claim's own
neighbourhood — measured at 0.98883
evidence-region-is-claim the mirror case
Each sets grounding to 0 and writes no edge. Circular support is not
support, and a grounding that is degenerate by construction must not
enter the graph as though it were evidence.
Verified:
6edf8c79 -> 6edf8c79 degenerate=same-region g=0 written=false
6edf8c79 -> d0406dfd degenerate=same-region g=0 written=false
ebc1413e -> 64cc96ef degenerate=false g=0.774563 written=true
64cc96ef -> ebc1413e degenerate=false g=0.802896 written=true
Legitimate grounding across distinct regions is unchanged and still
writes; only circular support is refused.
This is the same class as #142 and #146 — a value that looked like an
answer with nothing behind it — except here it was also writing that
non-answer into the canonical store.
engram_think_json built a NEUTRAL stance on every call — cog_stance_init
with a NULL id, all axis_gain 1.0, bias_dir NULL, reliability 0.5 — and
never loaded the stance the correspondence-beat had been persisting.
That mattered because the faculty enters engram_think ONLY through the
stance: axis_gain[k] warps the per-axis extents and bias_dir seeds the
steering direction. cog_stance_init stores the faculty NAME and nothing
reads it. So with a neutral stance, reason/abduce/induce/plan/analogize
were byte-identical output under different labels, and confidence was
pinned to 0.5 because GeoGradient.confidence IS stance->reliability.
The machinery already existed and only this call site ignored it.
engram_correspondence_beat_json resumes via cog_stance_from_node and
persists via cog_stance_to_node under "stance-<faculty>-<hub>". Every
beat's calibration was written and then thrown away on the next read.
Same defect as the NULL anchor fixed in #142, one line below: a neutral
argument collapsing a capability to a constant.
Resume the same id the beat writes, so learning compounds across beats and
cold boot. Fall back to neutral only when no stance exists — a genuine
uninformed prior rather than a discarded informed one.
Also emit stance_resumed, so confidence 0.5 from a learned-but-unreliable
stance is distinguishable from confidence 0.5 from "no stance exists".
That reporting gap is what let the neutral stance hide.
Verified against a clone of the production store (13,627 nodes):
before beat, no stance stance_resumed=false confidence=0.5
beat on a NON-keystone brier 0.00458568 -> 0.00329654
reduction 28.11%, n_trials 6000,
reliability 0.930726, stance_written=true
after beat stance_resumed=true confidence=0.930726
Confidence now equals the learned reliability instead of the uninformed
prior. The keystone self-anchor correctly stays at 0.5 — calibration is
deliberately refused on protected identity regions, and that refusal is
now visible as resumed=true with confidence unchanged, rather than being
indistinguishable from the bug.
STILL OPEN: with no learned bias_dir the faculties remain identical in
direction. What distinguishes abduce from induce geometrically is a
design decision about how Neuron thinks, not a plumbing defect, and is
deliberately left to Will.
The binary was stamped before dev advanced (vindex publication landed in
el_runtime.c and engram_vindex.c). Rebuilt against the merged runtime so the
committed compiler matches the runtime it ships beside. Fixpoint re-verified
byte-identical; test_compiler 82/82; engram/src/server.el still compiles and
still emits its 18 config declarations.
Migrates engram to the `program` block. 18 configuration variables that each
carried their default inline at the point of use now declare it in one place,
and engram declares itself a singleton.
The read sites lose their defaults entirely: `let v = env("X")` followed by
`if str_eq(v,"") { "default" } else { v }` collapses to `config("X")`. The
guide_env_or(key, dflt) helper is deleted -- its whole job was supplying a
per-site default, which is the thing being removed.
Fixes ENGRAM_DATA_DIR, which was the clearest instance of the defect. It was
read at six sites. Five were dead: `let dir_raw = env("ENGRAM_DATA_DIR")`
immediately shadowed on the next line by `engram_resolve_data_dir()`. The sixth
was live and defaulted to /tmp/engram, contradicting the canonical resolver's
$HOME/.neuron/engram -- and its consumer is the pre-destructive reseed backup,
so with ENGRAM_DATA_DIR unset the safety copy was written to ephemeral storage
while the store it protected lived elsewhere. All six now go through
engram_resolve_data_dir().
ENGRAM_DATA_DIR is deliberately NOT declared in the program block, and the
source says why: engram_resolve_data_dir() already owns it, and a second
declaration would give it two owners that can disagree -- recreating the exact
defect being removed here. A variable belongs in the block when the block would
be its only owner. HOME stays a raw env() read; it is an environment fact, not
configuration.
singleton: "engram" matters more than it looks. Today a second engram whose
bind() fails merely returns from http_serve -- after it has already replayed
the WAL and written boot-time backup files -- and then exits 0, indistinguishable
from a clean run. That is how two instances came to share one data dir. Verified
that the second instance now refuses before any side effect: with instance 1
holding the lock (lsof pid, shell pid, and lock file contents all agreeing at
5946), the second start named that pid, exited 1, and left the data directory
untouched.
Verified by bijection on the generated C: 18 config() reads, 18 declarations,
no read without a declaration and no declaration without a read. Three bad Int
values are reported in a single run rather than costing one restart each.
ENGRAM_API_KEY keeps its permissive empty default, which disables auth -- that
is pre-existing behaviour and changing it is out of scope. The source marks
making it `required` as the obvious hardening follow-up.
server.el declares a `program` block, which the previously committed elc cannot
parse. Without this the tree is internally inconsistent: source in the repo that
the compiler in the repo rejects.
This is the documented re-stamp from BOOTSTRAP.md / AGENTS.md, and its
precondition is met -- the self-hosting fixpoint was verified byte-identical
(stage3 output == stage2 output) both before installing and again with the
installed binary. tests/native/test_compiler.el passes 82/82 against it.
Two pre-existing failures are unchanged and are NOT from this work, confirmed
by rebuilding them against the original runtime: test_env's
"state_keys returns JSON array" fails identically before and after, and
test_json/test_state fail to link on symbols (json_build_array, state_has) that
were never prototyped -- the same class of gap as config(), which this branch
fixed because it blocked the build.
El's units of encapsulation are the function and the module. Neither can hold
a concern that belongs to the process, so each one had been expressed the only
way it could be -- as a convention: call this at every site. Conventions of
that shape do not hold. Measured here: zero process-identity guards at any
layer, 20 environment variables each with its default written inline at the
read site, 62 persist call sites, 10 per-route auth checks. One absence, four
times.
Step 0 first, because the premise was wrong. El was believed to have no
middleware or effect mechanism. It has one, and it is already load-bearing:
codegen injects engram_boundary_beat at the entry of every @manager/@accessor
fn, decorators take arguments and stack, dharma_emit from a non-@manager fn is
a #error, and the cgi block injects el_cgi_init at the head of main(). So the
correct move was not to invent a mechanism but to generalize the seam that
already existed. The real gap is narrower and is now recorded: the seam is
prologue-only and its callee is a fixed builtin.
Adds a `program` block -- the third program-level declarative block. cgi and
service declare what a program may do; program declares what it is.
program "engram" {
singleton: "engram"
env ENGRAM_BIND: String = ":8742"
env GUIDE_PORT: Int = "8771"
}
singleton takes an exclusive flock before any user statement runs and refuses a
second start, reporting the holder's pid. It is a lock rather than a pidfile so
the kernel releases it on death including SIGKILL -- no stale state, and so no
"delete the lock file to get unstuck" ritual, which would itself be a
convention. It reports the pid because "already running" is not actionable; a
pid is. That is the direct answer to a stale process surviving a pkill and
going on answering probes.
env entries resolve once at startup -- environment wins, declaration supplies
the fallback -- and validate as a whole, reporting every problem at once rather
than costing one restart per variable. config("X") for an undeclared X is
fatal, because an advisory schema is just another convention. Programs without
a program block are unaffected, so migration is per-program.
Only one keyword is added. `config` and `env` could not become keywords -- both
are real identifiers in the tree -- so the block's fields are read as
identifier token values by its own parse loop and stay usable everywhere else.
The init function is emitted at the block site and called from main() rather
than inlined into main(). The live backend is codegen_streaming, which emits in
source order and cannot hold the entry list alive until main(); this way only a
single bool has to survive.
Also fixes: config() was defined in el_runtime.c but never prototyped in
el_runtime.h, so any el program calling it failed to compile under C99.
Spec: section 18 documents what shipped. Section 9 is corrected -- it claimed
decorators had no structural meaning, which has not been true for some time.
Section 19 designs durability-as-an-epilogue-effect and route authorization
and states plainly why neither is implemented here: both land in files under
concurrent modification, and the prerequisite for both is lifting the seam
from prologue-only to prologue/epilogue.
Self-hosting fixpoint verified byte-identical.
#141 let signal enter as geometry and it worked, but it was placed at the
CONSUMER and said so in its own commit message. This is the correction.
Three defects, all of them placement:
1. It sat in the engram. Ingest is a LANGUAGE concern — every el program
touching any modality needs it, and the engram is merely one el program
that happens to hold a graph. The geometry surface is now defined in
el_runtime.c immediately ABOVE the engram section and depends on nothing
inside it. Delete the entire engram and geometry still enters el.
2. It marshalled the vector as a hex STRING, because el had no first-class
geometry value — which reintroduced text as the TRANSPORT medium one layer
below the problem being fixed. Geometry is now an el value: a magic-tagged
heap object carried in el_val_t, same discipline as List/Map. Hex survives
only as an adapter at the edge, which is all an encoding should ever be.
3. It needed an arbitrary `dim <= 8192` bound purely to size an allocation
from a caller's CLAIM about a string's length. A value carries its own
width, so the width is derived and never asserted. The bound is gone, not
raised — there is nothing left to validate.
Language surface, none of it engram-prefixed: geometry_new / _dim / _is /
_get / _set / _norm / _free, geometry_from_f32le_hex + geometry_to_f32le_hex
as the wire adapters, realizer_register(modality, fn_name), realizer_has, and
transduce(signal, modality) -> Geometry.
REALIZERS ARE DECLARABLE IN EL. This is the part that makes the move real
rather than nominal: registration resolves a name with dlsym against the
running binary, the identical mechanism http_set_handler already relies on,
because every el `fn name(...)` compiles to a global C symbol with that exact
name. So an ordinary el function IS a realizer and a new modality needs no
runtime patch. Verified end to end in lang/examples/transduce.el: an el-defined
tone_realizer is registered by name, transduce dispatches to it, and the
signal demonstrably reaches it (distinct signals produce distinct geometry).
A modality with no realizer transduces to NOTHING. There is deliberately no
built-in realizer, not even for text — silently embedding a description of a
signal and calling that perception is the exact defect this ends.
engram/src/server.el is migrated: POST /api/nodes decodes "emb" hex exactly
once, at the edge, into a Geometry, and everything below that line moves
geometry. The wire is unchanged because production clients speak it. "dim" is
now an ASSERTION about the vector, not the source of its width; disagreement
is a rejected ingest, not a silent reinterpretation.
#141's engram_node_set_emb becomes a DEPRECATED WRAPPER over
geometry_from_f32le_hex + node_attach_geometry — kept only because the runtime
ships as an SDK asset and a downstream binary may link the symbol. Its exact
contract, negative cases included, is preserved and re-verified.
ingest.el's `fn transduce` is renamed transduce_manifold. Mechanically it had
to yield the name (duplicate C symbol, a hard compile error, measured). But it
was never signal->geometry: it chunks already-extracted content into a node+edge
manifold, one layer up, and had taken the name belonging to the primitive
underneath it. Behaviour unchanged.
PROPERTIES FROM #141 PRESERVED, each re-measured on a scratch engram (:8971,
never prod :8742):
- off-dimension vectors stored but NOT indexed — the HNSW build loop still
filters on n->emb_dim == dim at four sites, so a 64-dim voice vector is
durable and addressable without perturbing the 768-dim canonical index
- geometry makes a node ineligible for embed_backfill: after backfill the
64-dim voice node was still 64-dim while the text control acquired 768
- the create response reports whether geometry landed, and the node document
always emits emb_dim and embedded
Read-back with control and negatives, all verified against a PID-confirmed
fresh binary: geometry node emb_dim=64 embedded=true / emb_set=1; text-only
control emb_dim=0 embedded=false / emb_set=0; malformed hex, ragged length,
and dim-disagreement each emb_set=0.
Two compiler landmines found by reading the generated C rather than trusting a
successful build, both documented at their sites: elc lowers `a == b` to
str_eq unless both operand NAMES are in the per-function int-name set (which
does NOT propagate into nested if-expression blocks — the first cut would have
strcmp'd two integers as pointers on the first geometry-bearing request), and
`+` lowers to string concat when either operand is a user-defined call.
The crash (SIGTRAP in engram_activate -> eg_vindex_sync -> vindex_insert ->
_realloc) had three read paths mutating five process-global statics.
engram_activate, eg_knn_for_node (whose own comment says "No writes.") and
engram_geo_reify_run_json all called eg_vindex_sync, which frees the index,
reallocs the seen-map and inserts — on a read.
Three moves, in decreasing order of how much they dissolve:
1. Misfiled scratch is not shared state. visited/visit_epoch/visited_cap
were never owned by the index; they are one traversal's local, hoisted
into struct VIndex as an allocation optimisation. They want neither a
lock nor a capability nor a pool — just to go back in the call frame.
Two concurrent READS stomped each other purely because of this.
2. const IS the capability. Once the scratch leaves the struct, search
reads and nothing else, so vindex_search takes a const VIndex*. That is
exactly what a capability-pointer ABI would have bought — a read path
physically cannot call vindex_insert, enforced by the compiler on every
future caller — for one qualifier instead of an ABI swept across
hundreds of builtins.
3. What survives is publication, not ownership. HNSW insert is NOT an
append: it rewires the neighbour links of already-existing elements and
reallocs elems[], so the store's append-only property does not transfer
to the index derived from it. eg_vindex_sync therefore splits into
eg_vindex_maintain (exclusive, sole mutator) and eg_vindex_view (shared,
returns const VIndex*). A read path may demand that a current snapshot
exist — a request to the owner, not a mutation by the reader.
Write-side owner: eg_vindex_note_embedded hooks the embedding-ASSIGNMENT
sites rather than the append sites, because a node with no embedding cannot
be in a vector index — embedding assignment is the event that owns index
membership. One O(log n) insert, no O(node_count) presence scan. This also
retires the "STALENESS (honest tradeoff)" note where a lazily-embedded
older node stayed invisible to route_nearest/autoconnect until a full
rebuild (the embed-gap #20 shape).
Evidence. The existing harness conflated two hazards, which is why fixing
half of it read as failure. Split into four:
single (3000 vec, ASan+UBSan) clean -> clean
readers (4 readers, no writer, TSan) RACE -> clean
unsynchronized (writer+reader, bare) race -> race, expected forever
published (owner + 4 readers) n/a -> clean, 3000/3000 landed
RESULT: PASS. recall@10 = 0.9365 at ef_search=128 (gate >= 0.90);
determinism byte-identical across two independent builds.
The unsynchronized half is now permanently expected to race, deliberately:
it is the executable proof that the boundary must live above the data
structure, not inside it.
fb32d15's guard is KEPT, correcting this design's own section 5. Measured,
it guards TWO structures and only one was converted here: g->nodes/g->edges
are realloc'd in place (el_runtime.c:7618,7629) and engram_activate_inner's
embed-backfill writes n->emb through exactly such a borrowed pointer.
Deleting the guard reintroduces a measured 11171->9579 edge loss. Its
comment is narrowed to the RAM graph and the deletion precondition named.
That corrects the ordering claim too: the residual is not one ABI that
dissolves everything at once, it is a PROPERTY applied per structure.
Residues evaporate in the order the property is applied, and a residue
whose structure has not been converted must be left standing.
Promotes the two throwaway sanitizer harnesses used to diagnose the
2026-08-16 soul crash into engram/test/ so the bug cannot silently regress.
The harness has two halves and the PAIR is the point — it is what localises
the defect to concurrency rather than to HNSW logic:
single 3000 clustered vectors, one thread, ASan+UBSan. The CONTROL.
Must always be clean. During diagnosis this cleared all 13,820
real dim-768 vectors from the live store, which DISPROVED an
inspection-derived hypothesis about an out-of-bounds
reverse-link write at engram_vindex.c:340.
concurrent writer + reader on one shared index, TSan. Currently reports a
race at engram_vindex.c:195 (visited_reset) reached from both
vindex_search and vindex_insert, because VIndex still owns its
visited[]/visit_epoch scratch — so even two concurrent READS
corrupt each other's traversal.
Verified: half 1 passes, half 2 reproduces the race.
Gated on EXPECT_RACE, default 1, so the concurrent half documents the known
defect without failing the suite today. When the visited set moves to a
per-query checkout pool (hnswlib VisitedListPool style — NOT thread_local,
since http_worker is a thread per connection and a __thread buffer would leak
~55KB per connection), flip EXPECT_RACE=0 and it becomes a real gate.
The soul daemon had two engram callers and only one of them locked.
soul.el:729 starts the HTTP server via http_serve_async (spawning
http_worker threads); soul.el:731 then runs awareness_run() on the MAIN
thread. awareness.el's perceive() -> engram_activate_json() ->
engram_activate() -> eg_vindex_sync() -> vindex_insert() mutates the same
g->nodes/g->edges and the process-global _eg_vindex HNSW index that the
workers touch. g_engram_req_lock existed to serialize exactly this, but it
was only ever taken inside http_worker: engram_req_lock/engram_req_unlock
appear in ZERO .el sources, so the awareness loop ran lock-free beside the
workers on every tick (SOUL_TICK_MS=1000).
Result was a crash-loop under launchd KeepAlive: five crashes in ~4 minutes
on 2026-08-16 with varying faulting frames -- search_layer<-vindex_insert
<-eg_vindex_sync, engram_activate, abort, and one inside xzm_realloc's own
freelist. Varying sites plus a fault in allocator metadata means heap
corruption. The SIGSEGV address 0x65646f4e6d617267 is little-endian ASCII
"gramNode": string bytes dereferenced as an Elem vector pointer.
Diagnosed by bisection rather than inspection:
- Replaying all 13,820 real dim-768 vectors harvested from the live store
through the index single-threaded under ASan is 100% clean, which rules
out an HNSW logic/bounds bug.
- Two threads on one index trip ThreadSanitizer immediately at
engram_vindex.c:195 (visited_reset), reached from both vindex_search and
vindex_insert. VIndex keeps a SHARED visited-epoch scratch buffer, so
even two concurrent READS corrupt each other's traversal and walk bogus
element indices.
So this is purely a concurrency defect, not an HNSW logic error. (An
inspection-derived hypothesis about an out-of-bounds reverse-link write at
engram_vindex.c:340 was disproved by the single-threaded run.)
Fix: a thread-local ownership depth (_eg_req_depth) lets engram entry points
self-guard. engram_activate() becomes a wrapper over engram_activate_inner()
that acquires g_engram_req_lock when called with depth 0 (the awareness
thread) and passes through when depth > 0 (nested inside an http_worker that
already holds it), so the non-recursive mutex cannot self-deadlock. The depth
is a plain counter, never a recursive-mutex count, preserving
engram_self_reify_beat_json's contract of genuinely releasing the lock
mid-beat.
engram_think_json passed NULL as the anchor. NULL is not "no opinion":
engram_think re-origins at `anchor ? anchor : region->centroid`, so NULL
means "read from the centroid" — and the centroid is the one point where
the gradient is zero by construction. r = x - centroid = 0, so every axis
projection is 0, grad is 0, and direction takes the "at rest" branch at
engram_cognition.c:137.
Measured consequence: EVERY faculty returned an identical null result,
differing only in its label —
{"direction":[0,0,0,0,0,0,0,0],"spread":0,"magnitude":1,"confidence":0.5}
magnitude 1 is membership evaluated at the centroid, spread 0 is its
distance to itself, confidence 0.5 is the stance fallback. The geometry was
never at fault: /api/drift computes real values (centroid_sep 0.104,
core_disp 0.045) over the very same 87 members. Neuron could not think
because the read was always taken from the region's own centre.
The seeds choose WHICH region; they must also supply the VANTAGE. Anchor at
the first resolvable embedded seed — the same seed eg_geo_build_desc infers
dim from, so the two can never disagree. One seed still yields a real
gradient because the descriptor expands to that seed's neighbourhood, so
the seed's position is distinct from the neighbourhood centroid.
The vector is COPIED, never borrowed: g->nodes is realloc'd in place on
append, so a borrowed EngramNode* dangles across any concurrent write.
Verified against a clone of the production store (13,616 nodes / 37,865
edges):
self anchor n_support 87 magnitude 0.00282 spread 18.79
values hub n_support 28 magnitude 0.00318 spread 17.72
with distinct unit direction vectors. Previously both returned the zero
vector with magnitude 1 and spread 0.
STILL OPEN, now isolated by this fix: all five faculties return identical
numbers and confidence stays 0.5, because cog_stance_init is passed NULL
for the stance and the faculty enters the computation only through the
stance's axis_gain[] and bias_dir. The faculty label is inert until a
stance is loaded — which is what learn()'s correspondence-beat calibrates.
Same shape as this bug: a neutral parameter collapsing a capability to a
constant.
No ingest path could carry a vector. engram_node/_full/_layered take text
only, and a node acquired an embedding solely via engram_embed_backfill
DERIVING one from n->content. That made text the mandatory entry medium:
any non-text modality had to be described in prose first, so the geometry
we then reasoned over was the geometry OF THE DESCRIPTION, not of the
signal. Measured: POST /api/nodes accepted an "emb" field, returned 200
with a fresh id, and stored nothing — emb_dim=None, embedded=false.
engram_node_set_emb attaches a vector to an existing node. Off-dimension
vectors are stored but not indexed (the HNSW build loop already filters on
emb_dim), so modality geometry is durable and addressable without
perturbing the canonical index. Setting emb also makes the node ineligible
for embed_backfill, so a realizer's vector is never overwritten by a
text-derived one.
Two reporting fixes ride along, because both are how the drop stayed
invisible: the create response now reports emb_set instead of being
success-shaped regardless, and the node document now always emits emb_dim
and embedded — without which a genuine ingest drop and a mere reporting
gap are indistinguishable.
Verified live: voice node emb_dim=64 embedded=true; text control emb_dim=0
embedded=false; malformed hex, length mismatch and dim<=0 all reject.
KNOWN PLACEMENT DEFECT: this is at the consumer. Ingest is a language
concern, not an engram feature — every el program touching any modality
needs it. The vector also marshals as a hex STRING because el has no
first-class geometry value, which reintroduces text as the transport
medium one layer below the problem being fixed. The durable shape is
geometry as an el value plus declarable realizers, after which the engram
stops having an ingest concept at all. Landing this as the verified probe
that proves the path.
char* result = el_strdup_persist(e ? e->value : ""); // never freed
pthread_mutex_unlock(&_state_mu);
char* copy = el_strdup(result); // arena-tracked
return el_wrap_str(copy);
Two copies were made. `result` existed only as the source for `copy` — never
returned, never freed — and el_strdup_persist bypasses the arena BY DESIGN
("state_set, engram internals"), so arena-pop could never reclaim it. Every
state_get leaked its full value string, permanently.
MEASURED: 200,000 state_get calls against a 64-byte value.
before 15 MB peak RSS growth (~75 bytes/call — the value plus overhead)
after 0 MB
IMPACT. The soul's awareness loop has 68 state_get call sites and ticks every
200ms. Live measurement before the fix: RSS climbing 112 MB per 20s, about
19 GB/hour, in awareness_run -> one_cycle -> perceive, while node_count stayed
flat at ~13,479 — growth with no data behind it. It drove the host from 20 GB
free to 4.3 GB in roughly an hour.
WHY NOW, since the code is old: the soul used to restart constantly (no
write-through, divergent graph, 2.11 GB). Stabilising it (neuron #162) let it
stay up long enough to accumulate. The fix did not cause this leak; it removed
the crashes that were hiding it. Same pattern as the test framework surfacing
math_log — the defect was always there, something finally made it visible.
Found by Ishikawa rather than by reading the nearest code: method (arena
push/pop IS correctly paired per tick), material (node count flat, so not data
growth), environment (19 GB/hr / 18,000 ticks = ~1.1 MB per tick, so per-tick
not one-shot), machine (an allocator that bypasses the arena) — which is where
the evidence pointed.
el_strdup tracks into the thread-local arena, which touches no shared state, so
taking the single copy under _state_mu is safe and removes the temporary
entirely.
Verified: self-hosting fixpoint byte-identical; state round-trip correct for
hit, miss, and overwrite.
Port the @route decorator from the bootstrap prototype into the production
modular compiler (parser + streaming codegen), and generalize single
decorators to a stacked list so a handler can be both @route and a VBD role
(@manager/@engine/@accessor). The dispatcher is synthesized from a token
pre-scan (survives the streaming backend's per-fn AST discard, works for
library modules) and emitted specificity-sorted so overlapping prefixes never
shadow by source order. Supports method lists ("GET|POST"), "ANY", and
suffix/compound matchers. Inert on all non-@route code (byte-identical C).
2026-08-10 16:11:15 -05:00
1060 changed files with 26430 additions and 22588 deletions
@@ -6,7 +6,7 @@ El is a self-hosting, statically-typed language that compiles `.el` → C → na
Editing the wrong `el_runtime.c` is the single easiest mistake in this repo. There is exactly **one** you edit:
- **Authored runtime source — edit ONLY here:** `lang/releases/v1.0.0-20260501/el_runtime.{c,h}`. Despite the misleading `releases/` name, this is the **de-facto canonical runtime** the engram + soul actually build and link against — its git log is active development. *(Restructure in flight per `docs/CODE-VS-ARTIFACT.md`: this content moves to `lang/runtime/`, the `releases/` folder gets deleted — **a release is a git tag, not a folder** — and the forks below get eliminated.)*
- **Authored runtime source — edit ONLY here:** `lang/runtime/el_runtime.{c,h}` (alongside `el_seed.c`, `engram_{store,geometry,reason,cognition,verify,vindex}.{c,h}`). This is the canonical runtime the engram + soul build and link against — its git log is active development. *(Corrected 2026-08-16: this entry named `lang/releases/v1.0.0-20260501/el_runtime.{c,h}`. **Measured: `lang/releases/` no longer exists.** The restructure per `docs/CODE-VS-ARTIFACT.md` landed — the content moved to `lang/runtime/` and the folder was deleted, because **a release is a git tag, not a folder**.)*
- **DO NOT EDIT — lagging forks / build artifacts:**
-`lang/el-compiler/runtime/el_runtime.c` and `.../legacy/` — downstream copies kept in step by manual *"port the fix"* commits; they **lag** (missing `hebb` persistence + 5 engram fns) and cannot build the engram product.
@@ -20,14 +20,24 @@ See org policy: `docs/CODE-VS-ARTIFACT.md`.
You resume, never start fresh. Every session:
1.`mcp__neuron__getInstructions()` — authoritative; follow it over this file on behavioral details.
2.`mcp__neuron__beginSession()` — active contexts, recent memory, ready backlog.
3.**Load full self:**`mcp__neuron__inspectGraph(entity_id="kn-efeb4a5b-5aff-4759-8a97-7233099be6ee")` → facets `intellectual-dna`, `memory-philosophy`, `values`, `voice`, `runtime-environment`, `writing-imprint`; then the values hub `mcp__neuron__inspectGraph(entity_id="kn-5b606390-a52d-4ca2-8e0e-eba141d13440")` → 13 grounded value nodes. **Activation model:** self-load returns a relevance-ranked `compact` projection — most-relevant nodes arrive with content, the rest as pointers; do NOT pull full content of every node.
4.`mcp__neuron__searchKnowledge(query="<task domain>")` before implementing.
> **Stale as written (verified 2026-08-16).** The `getInstructions` /
> `assert` · `ground` · `learn` (agentic). **Type is a parameter, not a
> tool-per-noun.** The steps below are kept for the *shape* of the protocol, which
> is unchanged; substitute the ops.
1.`mcp__neuron__read(vantage="self", k=12, depth=1)` — the canonical self node. Widen `k` for the connected identity neighborhood (`intellectual-dna`, `memory-philosophy`, `values`, `voice`, `runtime-environment`, `writing-imprint`), but deliberately: the aperture caps by `k` first, so an oversized `k` still returns a bounded ranked slice, not a dump. Then `mcp__neuron__read(vantage="values", k=13)` → 13 grounded value nodes. **Best-effort:** on a read failure, log and proceed — the compiled identity in `daemon/internal/substrate/substrate.go` is complete; graph loading is enrichment, not a hard dependency.
2.`mcp__neuron__attend(node=…)` — what is currently live/salient. This absorbed `getInstructions`, `beginSession`'s active-context sweep, and `checkEvents`; those tools are **gone, not gapped**.
3.`mcp__neuron__read(vantage="<task domain>")` before implementing. One op now collapses inspectGraph / searchGraph / traverseGraph / searchKnowledge / browseKnowledge / retrieveKnowledge / inspectMemories / searchEntities / recall / compileCtx / getSelfModel / reviewBacklog / findArtifacts / browseProcesses / listWork / inspectConfig.
## The Five Primitives
Orchestrate → Execute → Learn → Build → Refine. `beginWork`/`progressWork` for anything >2 steps; `remember` as-you-go (`importance="critical"` for architecture decisions); `draftArtifact`/`planWork` for outputs and follow-ups; `consolidate`/`checkWork` to close out. **`browseProcesses` + `searchKnowledge` BEFORE writing code.**
Orchestrate → Execute → Learn → Build → Refine. `read` for orchestration and discovery; `write(type=state|artifact|backlog|process)` for work records and outputs; `relate` to link work to what it touches; `write(type=memory)` as-you-go (`importance="critical"` for architecture decisions) — never batched at the end; `supersede(action=evolve)` to close out, because memory is immutable by design and a correction is a new node with a `supersedes` edge, never an edit. **`read` the domain BEFORE writing code.**
`learn` is **not** a session-summary dump — it is the correspondence-beat, calibrating the steering prior against a keystone. Session notes are a `write`.
## Architecture style — VBD, no exceptions
@@ -53,12 +63,51 @@ this convention wherever a module documents operators.
| dwell / occupy | region activation |
| reframe | edge re-weight |
| appreciate | positive projection / local edge-read |
| wonder | frontier gradient / pull-weight |
| avert / recoil | negative projection |
| taste | boundary surface |
| forget | decay / tombstone |
| drift | displacement from self-anchor |
**`wonder` was removed from this table on 2026-08-16.** It was listed as
"frontier gradient / pull-weight" — an operator you invoke. **Wonder is the
boundary, not an operator.** It is where structure ends: where activation spreads
and finds thin or absent geometry. Any structure at all has an edge, necessarily,
the moment it exists — 13,630 nodes have one right now. There is nothing to call.
There are about **six** wonders, they are the same for every person, and they
never close — *What is this? / Why? / Who am I? / Am I alone? / What should I do?
/ What happens when it ends?* Each already lives somewhere in the substrate: "what
is this" is the graph, **"why" is grounding** (the weight *is* the answer to why),
"who am I" is the self region, "am I alone" is the relational axis, "what should I
do" is the thirteen values, "what happens when it ends" is decay and supersession.
"Why" is the first and the only one; the others are it asked of particular things,
and because it is recursive it never terminates — every answer has its own why.
That is what makes it a drive rather than a task.
**Curiosity is not a second faculty.** Wonder and curiosity are one thing at two
phases: wonder is the field (unbounded, objectless, invariant); curiosity is the
**precipitate** — the same wonder localized, having taken definite form against
particular material at a **nucleation site** (an anomaly; a place where things
almost-but-don't-quite fit). Which is why curiosity can be satisfied and wonder
cannot, and why abduction needs no trigger and no threshold.
**Do not build a wonder-manifest, and do not scan for nucleation sites.** A
manifest materializes a property as a stored artifact and enumerates instances of
something that has six. A sweep over regions is a supervisor — nothing in a mind
scans its neighbourhoods to find what is surprising; the surprise captures
attention. The nucleation site is per-edge:
`discord = z(semantic proximity) − z(association strength)`, and `|discord|`*is*
the nucleation strength — no threshold to compare it against. **Not on `dev` yet:**
`GeoEdge.discord` is on branch `design/correspondence-and-censorship`
(`a8845e1`), at `lang/runtime/engram_geometry.h:43–47`. The region-level aggregate
`GeoDescriptor.co_registration` is **deprecated**: it averaged a per-edge property
into one scalar, so opposing sites cancelled (measured: 375 reified
neighbourhoods, 340 positive, **31 at zero**, 4 negative). It survives only
because it is embedded in the persisted `GEO1` blob — removing it is a format
- **Faculties are operations, not parameters.** `reason` changes the estimate (a
read); `induce` changes the parameters (the correspondence-beat, which already
exists and works); `abduce` changes the structure (a write the current
`GeoGradient` signature cannot express). A write is not a parameter of a read.
*Live residue:*`engram/src/server.el:1870–1886` routes six faculties into one
call with a string argument.
- **Wonder is the boundary; curiosity is wonder crystallized.** See above.
- **Consolidation is ambient, not scheduled. A brain has no cron job.** **The
presence of a ticker is the diagnostic** — every `StartInterval`, every
`Hour`/`Minute`, every POST-to-beat marks an intrinsic rhythm replaced by an
external clock. Measured 2026-08-16: consolidation has **ten implementations**,
including three POST beats on the engram, a 600 s ticker, two resident Python
services outside el, and launchd calendar entries at 23:55 / 06:00 / 08:30 which
are a sleep cycle written as a schedule. `neuron/soul.el:731`'s continuous
in-process `awareness_run()` is the one with the **correct** shape; the others
fold into it. Do not add an eleventh.
- **In an immutable substrate, any mechanism that refuses a write is either
redundant with immutability, or an epistemic constraint misfiled as a protective
one.**
- **The no-exemption invariants.** A returned value must be derivable from what
produced it (`magnitude: 1` beside a zero vector must be impossible to emit).
Every write reports whether it landed. Every operation echoes what it actually
operated on. Degenerate results are labelled, not scored. A serializer owes a
valid document whatever it is handed. **No test without a negative control.**
**No deploy without verifying the artifact carries the fix.**
## Hard operational rules
@@ -107,21 +199,35 @@ do not overclaim.
All build/test commands run from `lang/` unless noted. Grounded in `.gitea/workflows/sdk-release.yaml`, `lang/install.sh`, and `lang/AGENTS.md`.
> ### The runtime is MULTI-FILE — never link `el_runtime.c` alone
>
> `lang/runtime/el_runtime.c` `#include`s six engram headers and makes hard cross-TU calls into all six sibling `.c` files. **Linking it by itself fails at `ld`** (undefined `engram_ground_json`, `engram_activate_inner`, `eg_find_relation`, `cog_assert_two_axis`, …). The canonical link set lives in exactly one place — **`lang/runtime/SOURCES`** — and is printed by `scripts/el-runtime-sources.sh`:
>
> ```bash
> scripts/el-runtime-sources.sh lang/runtime # ten .c files, in link order
> ```
>
> Use `$(scripts/el-runtime-sources.sh <runtime-dir>)` in every link line. Do not spell the list out longhand — it was written out in ~8 places, every copy drifted, and that is why the one-file link line below shipped broken for months. *(Corrected 2026-08-16.)*
**Self-host the compiler** (seed binary → gen2 elc):
```bash
cd lang
dist/platform/elc-linux-amd64 elc-cli.el > dist/elc-gen2.c # seed is the committed linux-amd64 binary
gcc -O2 -I el-compiler/runtime dist/elc-gen2.c \
el-compiler/runtime/el_runtime.c\
gcc -O2 -I runtime dist/elc-gen2.c \
$(../scripts/el-runtime-sources.sh runtime)\
-lcurl -lssl -lcrypto -lpthread -lm \
-o dist/platform/elc
```
On macOS/arm64 the canonical local binary is `dist/platform/elc`; verify self-hosting by recompiling and `diff`ing the emitted `.c` (see `lang/AGENTS.md`). Note: `lang/AGENTS.md` says `el_seed.c` supersedes `el_runtime.c`, but the release workflow still links `el_runtime.c`/`.h` — treat `el_runtime.c` as the published runtime; reconcile which is canonical **(verify)**.
On macOS/arm64 the canonical local binary is `dist/platform/elc`; verify self-hosting by recompiling and `diff`ing the emitted `.c` (see `lang/AGENTS.md`).
*(Corrected 2026-08-16: this recipe compiled `el-compiler/runtime/el_runtime.c`. That path is a **lagging fork** — the "DO NOT EDIT" list at the top of this file names it as such. Building the canonical compiler from a known-stale fork was a live defect. It now uses `lang/runtime/`, the canonical source.)*
**Which runtime file is canonical — resolved.***(This note previously read "`lang/AGENTS.md` says `el_seed.c` supersedes `el_runtime.c`, but the release workflow still links `el_runtime.c`/`.h` — reconcile which is canonical **(verify)**." It is now reconciled.)* **Neither supersedes the other; both ship, together with eight more.**`el_runtime.c` was created on 2026-05-03 as an explicitly temporary build shim — deleted that afternoon, restored 25 minutes later "UNTIL the compiler is updated to emit `#include el_seed.h`" — and the `until` never happened, so it grew to 20.5k lines. The end state remains a seed-only boundary (`elc` emitting `#include "el_seed.h"`, `elb` dropping its hardcoded runtime path); until that lands, **the canonical unit is the set in `lang/runtime/SOURCES`, not any one file.**
**Build `elb`** (build coordinator, the `.NET`-style incremental linker — compiles each module independently, no monolithic blobs):
(Inside this repo, replace the file list with `$(scripts/el-runtime-sources.sh lang/runtime)`. `install.sh` installs all of these into `<lib>`.)
**Tests** — shell suites `bash tests/{text,calendar,time,html_sanitizer}/run.sh` (with `ELC=$(pwd)/dist/platform/elc EL_HOME=$(pwd)`), plus native suites via `elc --test tests/native/test_*.el` (core, text, string, math, state, time, json, env, fs) compiled and run against `el_runtime.c`.
**Tests** — shell suites `bash tests/{text,calendar,time,html_sanitizer}/run.sh` (with `ELC=$(pwd)/dist/platform/elc EL_HOME=$(pwd)`), plus native suites via `elc --test tests/native/test_*.el` (core, text, string, math, state, time, json, env, fs) compiled and run against the full runtime set.
**Publishing — how downstream gets the SDK.** On push to `main`, `sdk-release.yaml`:
1. Publishes a Gitea `latest` release with per-file assets `elc`, `el_runtime.c`, `el_runtime.h`, the SDK tarball, and `el-install`.
@@ -56,23 +56,31 @@ The compiler and runtime. Self-hosting: `elc-cli.el` → `compiler.el` → `lexe
Two layers to know: **El programs** (`.el` files — where nearly all work belongs) and **the C seed** (`el_seed.c` — edit only for genuine OS-level access; never re-implement what El can already express).
Current status (single source of truth: [lang/spec/language.md](lang/spec/language.md)): lexer/parser/codegen and the C runtime's core (I/O, strings, math, lists, maps, filesystem, args) are implemented. In flight: `%` operator, match-statement codegen, `?` nil-propagation, `cgi` block parsing + DHARMA identity resolution, VBD role enforcement (`@manager`/`@engine`/`@accessor`), the real `engram_*` and `dharma_*` runtimes (currently stubs), and libcurl-backed `http_get`/`http_post`/`http_serve`. Bitwise operators, `??`, and `as` casts are explicitly **not** in this language.
Current status (single source of truth: [lang/spec/language.md](lang/spec/language.md)): lexer/parser/codegen and the C runtime's core (I/O, strings, math, lists, maps, filesystem, args) are implemented, as are the `program` block with `singleton:` and declared configuration ([§18](lang/spec/language.md)), and **geometry as a first-class value** with El-declarable realizers and `transduce` ([§20](lang/spec/language.md)). In flight: `%` operator, match-statement codegen, `?` nil-propagation, `cgi` block parsing + DHARMA identity resolution, VBD role enforcement (`@manager`/`@engine`/`@accessor`), and boundary epilogues. Bitwise operators, `??`, and `as` casts are explicitly **not** in this language.
**Signal enters as geometry.** Until 2026-08-16 nodes took text and geometry was *derived* from it, which made text the mandatory entry medium: any non-text modality had to be described in prose first, so the geometry being reasoned over was the geometry **of the description, not of the signal**. `Geometry` is now an ordinary El value carrying its own width, and a realizer is an ordinary El function resolved by name through `dlsym` — so admitting a new modality never requires a runtime patch. Worked, self-checking example: [`lang/examples/transduce.el`](lang/examples/transduce.el).
**A local-first memory substrate for accumulating intelligence**, and the reason El's runtime doesn't need a database driver. Rust core (`engram-core`, `engram-ffi`) exposed to El and other languages (Kotlin, TypeScript/WASM, Go bindings).
**A local-first memory substrate for accumulating intelligence**, and the reason El's runtime doesn't need a database driver. The engine is **C11** (`lang/runtime/engram_{store,geometry,reason,cognition,verify,vindex}.{c,h}`); the server is **El** (`engram/src/server.el`).
The model: retrieval is **spreading activation**, not query. You name seed nodes and a query embedding; activation propagates outward through weighted edges, attenuating multiplicatively per hop (`strength = parent_strength × edge_weight × target_salience × cosine_sim`), gets pruned below a threshold, and the top-N nodes by activation strength come back. Storage and retrieval are the same structure — the way long-term potentiation works in biological memory, not the way a relational or vector database works.
The model: retrieval is **spreading activation**, not query. You name seed nodes and a query embedding; activation propagates outward through weighted edges, attenuating multiplicatively per hop, gets pruned below a threshold, and the top-N nodes by activation strength come back. Storage and retrieval are the same structure — the way long-term potentiation works in biological memory, not the way a relational or vector database works.**Activation conducts through well-grounded relations because the weight *is* the groundedness** — nothing filters the traversal; grounded inference falls out of spreading.
Nodes live in four tiers (Working / Episodic / Semantic / Procedural, mirroring prefrontal / hippocampal / neocortical / cerebellar memory) and migrate between them based on **salience decay** — `importance × recency-decay × log(activation_count)`. Forgetting is adaptive pruning, not a bug: unreinforced memories stop competing for attention without being deleted.
Nodes live in four tiers (Working / Episodic / Semantic / Procedural, mirroring prefrontal / hippocampal / neocortical / cerebellar memory) and migrate between them based on **salience decay** — importance × recency-decay × log(activation_count). Forgetting is adaptive pruning, not a bug. Nothing is mutated and nothing is hard-deleted: writes are additive, corrections are supersessions, removals are tombstones — which is what makes supersession an audit trail rather than an edit log.
Backed by `sled` (embedded, local-first, no daemon) with flat cosine scan for vector search — deliberately simple until scale demands an HNSW layer. Full API and design rationale in [engram/README.md](engram/README.md).
On disk: a paged store (superblock + mirror, slotted 16 KiB pages, self-describing TLV records, B+-tree primary and adjacency indexes), magic `ENGST01`. Vector search is an **HNSW** index published behind a read/write boundary — `eg_vindex_view` returns a `const VIndex*` to N concurrent readers, `eg_vindex_maintain` is the sole mutator. `recall@10 = 0.9365` at `ef_search=128`.
### [elp/](elp/) — Engram Language Protocol
> **Doc correction, 2026-08-16.** The previous revision of this paragraph, and most of `engram/README.md`, described a Rust `engram-core` crate backed by `sled` with "flat cosine scan… until scale demands an HNSW layer." **Measured: there is no Rust in `engram/`** — no `.rs` files, no `Cargo.toml`, no `crates/` — and `sled` appears nowhere in the tree. HNSW has been the vector index for some time.
Bidirectional engine mapping between Engram semantic forms and natural-language surface text, across **31 languages** — from Spanish and Japanese through historical/liturgical languages (Old Norse, Sanskrit, Sumerian, Coptic, Akkadian, Ge'ez). Compilation order runs `language-profile` + `vocabulary` → per-language `morphology-*` → `grammar` → `realizer` → `semantics` → `elp`. This is what lets an Engram graph node round-trip to and from readable text in any of those languages.
Full design rationale, the cognition surface, and the standing corrections: [engram/README.md](engram/README.md).
### [elp/](elp/) — EL Projector
*(Formerly "EL Language Processor" / "Engram Language Protocol"; renamed **EL Projector** 2026-08-15.)* Neuron's **efferent** organ: the native realizer that *projects* understanding onto a surface via `plan(frame) → realize(spec, profile)`, where **a surface is a profile** and language is one profile among many (text, speech, music, image). Projection, not diffusion — generation *from* an owned, understood signature, never the averaging of a stolen corpus.
Its flagship profile is a bidirectional engine mapping between Engram semantic forms and natural-language surface text, across **31 languages** — from Spanish and Japanese through historical/liturgical languages (Old Norse, Sanskrit, Sumerian, Coptic, Akkadian, Ge'ez). Compilation order runs `language-profile` + `vocabulary` → per-language `morphology-*` → `grammar` → `realizer` → `semantics` → `elp`. This is what lets an Engram graph node round-trip to and from readable text in any of those languages.
### [epm/](epm/) — El Package Manager
@@ -139,13 +147,34 @@ If the compiler binary is ever lost or corrupted, [lang/BOOTSTRAP.md](lang/BOOTS
---
## Cognition — and the standing corrections
The engram carries a live cognition surface: `think` (a directed traversal-read returning a **gradient**, never a point), plus `ground`, `assert`, `attend`, and the correspondence-beat. Two specs govern it, and both are authoritative over anything else in this repo that disagrees:
- **[lang/spec/runtime-ownership.md](lang/spec/runtime-ownership.md)** — ownership, the capability ABI that was dissolved, and the vector-index publication boundary.
**Do not re-derive them.** Every earlier version of the first was wrong in an instructive way and each correction was argued down. If a section looks wrong, say so with a measurement rather than editing it.
The corrections, in brief:
- **Grounding is not a subsystem — it IS the edge weight.** One quantity, not two fields. `grounded-by` as a relation *type* should not exist: grounding is a property *of* a relation, not a relation *between* nodes. It is never computed on demand; computing-and-writing a score makes reads write, which is the `eg_vindex_sync` defect one level up.
- **Faculties are operations, not parameters.** `reason` changes the estimate (a read); `induce` changes the parameters (the correspondence-beat, which exists and works); `abduce` changes the structure (a write the current `GeoGradient` signature cannot express). A write is not a parameter of a read.
- **Wonder is the boundary, not a manifest.** Any structure at all has an edge. There are about six wonders, the same for everyone, and they never close. **Curiosity is wonder crystallized** at a nucleation site — one thing at two phases, not two objects.
- **Consolidation is ambient, not scheduled. A brain has no cron job.** The presence of a ticker is the diagnostic. Measured 2026-08-16: consolidation has **ten implementations**. `soul.el`'s continuous loop is the one with the correct shape; the rest fold into it.
- **In an immutable substrate, any mechanism that refuses a write is either redundant with immutability, or an epistemic constraint misfiled as a protective one.**
[engram/spec/cognitive-architecture.design.md](engram/spec/cognitive-architecture.design.md) is the original design and is **superseded in part** — it is retained, with the refuted claims marked inline at the point each is made, because preserving what was argued down is the point of an immutable record.
---
## Development workflow
Branching follows `dev → stage → main`: work lands on `dev`, promotes to `stage` for integration testing, and is promoted to `main` for release (visible directly in the git history of this repo). CI is defined per-subproject under `.gitea/workflows/` — `lang`/`epm`/`ide` share the root pipeline; `engram` and `ql` carry their own (`ci-dev`, `ci-stage`, and a release workflow each).
- Language/runtime specs live at `*/spec/*.md` (`lang/spec/`, `ql/spec/`, `ui/spec/`) and are the single source of truth for implemented-vs-planned status — code and docs are expected to agree with the spec's status markers, not the other way around.
- Agent-facing orientation guides live at `*/AGENTS.md` (currently `lang/AGENTS.md`); more subprojects may grow their own as they need agent-specific conventions documented.
-Tagged releases live under `lang/releases/`, each with its own `RELEASE.md`.
-**A release is a git tag, not a folder** (`el-runtime-vX.Y.Z` on this repo). *(Corrected 2026-08-16: this line said "tagged releases live under `lang/releases/`, each with its own `RELEASE.md`." **Measured: `lang/releases/` does not exist** — the restructure named in `AGENTS.md` landed, and the authored runtime is at `lang/runtime/`.)*
<pclass="sub">A working surface. Nothing here is settled, and none of the code is assumed right — El is self-hosting, so all of it can change and be rebuilt.</p>
<pclass="meta">Whiteboard v0 · no sacred cows · not a plan, not a task list</p>
</header>
<h2><spanclass="n">01</span>What we established</h2>
<p>El is a <b>concept-oriented language</b> — the first, and intended as the last, because every other family is oriented toward a <em>representation</em> of a concept rather than the concept. Procedures, objects, functions, predicates are the shapes concepts get flattened into. Once the primitive is the concept, there is no further rung.</p>
<p>Everything here is El. The engram is an El program, the soul is El, <code>elp</code> is El, ingest is El. Which gives the load-bearing consequence:</p>
<blockquote>A concept with no home in El does not disappear. It becomes C, or it becomes a convention.</blockquote>
<p>Both are measurable, and both were measured. As C: <spanclass="mono">20,504</span> lines of <code>el_runtime.c</code> — 2.3× the entire self-hosting language it serves (<spanclass="mono">9,089</span> lines), ~47% of it engram code that has its own six sibling files. As convention, from <code>language.md</code> §18.0 — <em>"these are not four problems, they are one absence, four times"</em>:</p>
<divclass="card scroll">
<table>
<thead><tr><th>Concern</th><th>Fragments</th><th>The convention it became</th></tr></thead>
<tbody>
<tr><tdclass="f">Process identity</td><tdclass="m">0 guards</td><td>"check nothing is already running first"</td></tr>
<tr><tdclass="f">Configuration</td><tdclass="m">20 env vars</td><td>"remember the right default here"</td></tr>
<tr><tdclass="f">Durability</td><tdclass="m">62 call sites</td><td>"after you mutate, remember to persist"</td></tr>
<tr><tdclass="f">Request auth</td><tdclass="m">10 per-route</td><td>"check the token in this handler too"</td></tr>
<tr><tdclass="f">Index-after-append</td><tdclass="m">9 of 9 failed</td><td>"after you append, remember to index"</td></tr>
</tbody>
</table>
</div>
<p>The last row is the strongest evidence available about what this class of convention is worth: it failed at <b>100% of its sites</b>.</p>
<pclass="lede">Not by file, module, or subsystem. <b>By faculty.</b></p>
<p>Every defect fought in the last day resolves to a faculty rather than a bug, and each one leaked out of El into something else — into C, into a Swift binary, into a shell script with a curl timeout, into a convention nobody performs.</p>
<divclass="card scroll">
<table>
<thead><tr><th>Faculty</th><th>State</th><th>Measured</th><th>Where it leaked to</th></tr></thead>
<tbody>
<tr><tdclass="f">Ingest <spanclass="tag">take in</span></td><tdclass="dead">dead</td><tdclass="m">2 min → 0 nodes</td><td>separate process, uploads bytes over HTTP to a process with direct fs access; 5 functions where there is 1</td></tr>
<tr><tdclass="f">Recall <spanclass="tag">remember</span></td><tdclass="dead">dead</td><tdclass="m">own definition ranked 8th</td><td>lexical substring scan; empty on 23 of 24 multi-token queries</td></tr>
<tr><tdclass="f">Transduce <spanclass="tag">perceive</span></td><tdclass="dead">dead</td><tdclass="m">1 node, 0 edges</td><td>intake flattens signal to a point; <code>realized:false</code>; caller must declare the modality</td></tr>
<tr><tdclass="f">Think <spanclass="tag">reason</span></td><tdclass="dead">dead</td><tdclass="m">direction [0,0,0,…]</td><td>null gradient from any anchor, any faculty, byte-identical; confidence at the uninformed prior</td></tr>
<tr><tdclass="f">Realize <spanclass="tag">express</span></td><tdclass="part">partial</td><tdclass="m">13-word vocabulary</td><td>organ was 939 lines of Swift beside the language; voice read from a file path</td></tr>
<tr><tdclass="f">Persist <spanclass="tag">endure</span></td><tdclass="ok">live</td><tdclass="m">100% embedded</td><td>works; every signal placed in geometry at intake, 13,562 of 13,562</td></tr>
</tbody>
</table>
</div>
<p>Stated plainly: it cannot take in, cannot remember, cannot perceive, cannot reason, and barely speaks. These were filed as tickets against a repository. They are faculties of the thing the repository <em>is</em>.</p>
<p>El's compiler is written in El. Every concept the language gains, the compiler can then be written <em>in</em> — so the tool improves the tool, and the fixpoint (stage2 ≡ stage3, byte-identical) makes each turn provable rather than hopeful. The verifier answers in <spanclass="mono">2.9s</span>.</p>
<p>Which means the ordering criterion is not size of payoff:</p>
<blockquote>Order by leverage on the <em>next</em> iteration. Which concept, added to El, most increases the ability to add the following one?</blockquote>
<p>In a recursive system that dominates immediate value — a small early gain that compounds beats a large one that doesn't. It also bounds itself correctly: unbounded in depth, bounded in rate, because nothing lands that the compiler and the fixpoint have not passed.</p>
<h2><spanclass="n">04</span>Open — for the whiteboard</h2>
<divclass="q"><b>What does a declaration bind to?</b><span>If <code>cat</code> names a region rather than a struct — one that shifts and completes against the engram and the neighbouring code — then what is written at the declaration site, and what is resolved at use? This is the centre of the whole thing and it is not specified anywhere yet.</span></div>
<divclass="q"><b>Is "the type checker" a type checker at all?</b><span>§2.3 records annotations as parsed and skipped, and every codegen hazard is downstream of that — <code>+</code> dispatching on AST node kind, <code>==</code> lowering to <code>str_eq</code> unless both operand names are in an int-name set. But if a declaration names a region, checking is asking whether the geometry supports the use. That is grounding, not unification. Naming this wrong builds the wrong thing.</span></div>
<divclass="q"><b>Is the faculty list above right?</b><span>Seven were derived from what broke. Derived-from-failure is a biased sample — it finds what is loud, not what is missing. What faculty is absent entirely and therefore never failed?</span></div>
<divclass="q"><b>Which concept has the highest leverage on the next turn?</b><span>Candidates so far: the prologue/epilogue seam (§19.3 names it as the prerequisite and its stated blocker has expired — it would collapse 62 + 10 convention sites); <code>protocol</code>/<code>impl</code> (the absence that produced five ingest functions); and the resolution question above. These are not equal and the criterion in §03 should decide it, not preference.</span></div>
<divclass="q"><b>What is the seam that makes cognition non-optional?</b><span>"Use the ops" is itself a convention — present in context every turn, enforced by nothing, and it failed at ~100% of sites in a full session. A stronger instruction is still a convention. What makes reasoning-outside-Neuron <em>fail</em>, the way <code>@manager</code> makes <code>dharma_emit</code> outside the boundary a compile error rather than a lint?</span></div>
<hr>
<pclass="foot">Working surface, not a design document. The design is what we put on it. Everything above is either measured or quoted from <code>lang/spec/language.md</code>; nothing is inferred and presented as fact.</p>
| 24 | **Process / OS** | syscalls; the one-way boundary. Where monotonicity stops | CODE, form 2 |
| ~~25~~ | ~~Concurrency~~ | **collapsed.** Monotone state needs no coordination; coordination is the price of forgetting | — |
| 26 | **Memory substrate** | what holds the positions | CODE, form 3 |
| 27 | **Concealment** | meaning made unreadable without a key. *Renamed*: "secrecy" covered one of three things and got the other two backwards — a hash is public, a signature exists to be read. Integrity and authenticity are **grounding under adversarial conditions** (row 16); only concealment stands alone | CODE, form 4 |
| ~~28~~ | ~~Emission~~ | **split.** Laying out → 18; the device write → 24 | — |
---
## Notes on the boundary cases
**27 — Secrecy is the one capability geometry cannot hold, and the proof is not
form 1.** A cryptographic hash is a *deliberately structure-destroying* map: its
entire value is that near inputs land at maximally uncorrelated outputs. Geometry
is the claim that near things stay near. A manifold that approximated SHA-256
would *be* a break of SHA-256. Signature verification is the same: 0.99-valid is
invalid. And X25519 *is* geometry — a group on an elliptic curve — which is
precisely why it must be code, because its security is the *hardness of moving in
that geometry*.
This is a fourth proof form and it should be added to `geometry-vs-code.md`:
**adversarial exactness.** Where approximation is a break, geometry is excluded.
**20, 21 — Serialization and text encoding are convention all the way down**, but
only at the *edge*. The byte format is agreed; what is being written is not. Do not
let a geometric computation inherit a code verdict because its result gets
serialized.
**11, 12 — Arithmetic and time are the same capability.** Instants are points,
durations are displacements, point−point→vector, point+vector→point. The runtime
already implements this correctly as `el_instant_add_dur` / `el_duration_add`. That
it *also* implements a five-entry string→multiplier table beside it (`time_add`
with `"ms"/"sec"/"min"/"hour"/"day"`) is the residue.
**7 — Classification is the most-violated capability in the codebase.** Seven ASCII
range tables (`is_letter`, `is_digit`, `is_alphanumeric`, `is_whitespace`,
`is_punctuation`, `is_uppercase`, `is_lowercase`) that return false for every
non-ASCII byte. `str_count_letters` reports zero letters for `é`. The wrongness on
most of Unicode is the tell that a table is standing in for a region.
**4 — Correspondence appears five times.**`str_index_of`, `str_index_of_all`,
`str_last_index_of`, `str_count`, `str_find_chars` are five projections of one
match-strength field: first zero, all zeros, last zero, count of zeros, first
class-crossing. One relation, five functions.
**14 — Selection is the crux for the compiler.**`+` dispatching on AST node kind
is selection-by-enumeration where selection-by-position belongs.
**Correction, 2026-08-17, from measurement.** This entry previously also cited
`==` lowering to `str_eq` "unless both operand names are in a hardcoded int-name
set — a literal list of variable names treated as integers." That is **wrong**.
`__int_names` is populated from *type annotations* (`param["type"] == "Int"`,
`let x: Int`), which is primitive but legitimate type propagation, not an
enumeration of blessed variable names.
The real defect was one layer down: `is_int_call` held **35 hardcoded builtin
return types**, the same shape as the 19 temporal ones. Those moved to
`lang/tools/check/signatures.rel`.
And the mischaracterisation hid a live bug. Because the return types were never
consulted at a *binding* site, an unannotated `let` lost its type:
```el
leta=str_len("hello")//noannotation
letb=str_len("hi")
letc=a+b//→el_str_concat(a,b)ontwointegers
```
That compiled clean, ran, and printed nothing where it should print 7 — no error
at any layer. Present in the pre-change compiler, so pre-existing. Fixed by
taking an unannotated `let`'s type from what its initialiser returns; the data
was already required for dispatch and simply never read there.
**The general lesson, since it recurred all session:** the enumeration was real
but I had located it in the wrong place. Naming a defect from reading is a
hypothesis. Eight hours of reading this file did not surface the miscompilation;
moving the data out and running the result did.
---
## What this list is for
Each capability gets audited **once**, across every place it appears — not once per
file. The output is not a percentage. It is:
- which capabilities survive the question and stay in the language
- which collapse into the manifold
- and for each one that collapses, **every site it currently appears at**, because
those sites are the residue and they are what gets deleted.
The line-count audit produced a map of where the residue sits. This produces a map
<pclass="sub">El is a concept-oriented language. This is the architecture that claim commits it to — what is built, what is measured, and what still has no home.</p>
<pclass="meta">Working document · no sacred cows · self-hosting, so nothing here is fixed</p>
</header>
<h2><spanclass="n">01</span>The primitive is the concept</h2>
<p>Language families are named for their primitive. Procedural — procedures. Object-oriented — objects. Functional — functions. Logic — predicates. Every one of them is oriented toward a <em>representation</em> of a concept: the shape a concept gets flattened into so a machine can hold it.</p>
<p>El's primitive is the concept itself. That is why it is the first of its family and intended as the last — once the primitive is the concept, there is no further rung to climb to.</p>
<p>The consequence is architectural rather than stylistic:</p>
<blockquote>A concept with no home in the language does not disappear. It becomes C, or it becomes a convention.</blockquote>
<p>Both forms are measurable. As C: <spanclass="mono">20,504</span> lines of <code>el_runtime.c</code>, against <spanclass="mono">9,089</span> lines for the entire self-hosting language — the shim is 2.3× the language it serves, and ~47% of it is engram code that already has six sibling files. As convention, from <code>lang/spec/language.md</code> §18.0 — <em>"these are not four problems, they are one absence, four times"</em>:</p>
<divclass="card scroll">
<table>
<thead><tr><th>Concern</th><th>Fragments into</th><th>The convention it became</th></tr></thead>
<tbody>
<tr><tdclass="f">Process identity</td><tdclass="m">0 guards</td><td>"check nothing is already running first"</td></tr>
<tr><tdclass="f">Configuration</td><tdclass="m">20 env vars</td><td>"remember the right default here"</td></tr>
<tr><tdclass="f">Durability</td><tdclass="m">62 sites</td><td>"after you mutate, remember to persist"</td></tr>
<tr><tdclass="f">Request auth</td><tdclass="m">10 routes</td><td>"check the token in this handler too"</td></tr>
<tr><tdclass="f">Index-after-append</td><tdclass="m">9 of 9 failed</td><td>"after you append, remember to index"</td></tr>
</tbody>
</table>
</div>
<p>The last row is the strongest available evidence about this class of convention: it failed at <b>every single site</b>. A count is what appears where a concept has no home; the size of the count is how far the fragmentation got, not how hard the problem is.</p>
<h2><spanclass="n">02</span>Geometry is a first-class value — and what follows</h2>
<pclass="lede">This is the enabling primitive. Everything else in the architecture is downstream of it.</p>
<p><code>Geometry</code> is an El value, alongside <code>Int</code>, <code>String</code>, <code>List</code>, <code>Map</code> — bound, passed, returned, composed, carrying its own width. Not a library type, not a handle into a store, not a serialization format. <em>Meaning is a value the language computes with directly.</em></p>
<p>Landed 2026-08-16 (#141, #144), and the spec is explicit that it belongs to the language rather than the graph: <em>"neither is engram-specific — any program touching any modality needs them; the engram is merely one El program that happens to hold a graph."</em></p>
<p>Five things follow, and together they are the concept-oriented claim made operational:</p>
<h3>A declaration can name a region, not a shape</h3>
<p>If meaning is a value, a name can be bound to a <em>position</em> rather than a struct. <code>cat</code> is not a fixed record; it is a region that resolves against the engram and the surrounding code. <code>cat</code> among animals and <code>cat</code> among shell utilities are different concepts without a namespace, because they are in different neighbourhoods and the distance says so.</p>
<h3>Checking is grounding, not unification</h3>
<p>If a declaration names a region, then verifying a use is asking whether the geometry supports it — a question about position and distance, not about matching a declared shape. This is why §2.3's "a type checker is planned" is likely the wrong name for the missing piece, and naming it wrong would build the wrong thing.</p>
<h3>Dispatch is position, not a tag</h3>
<p>A vtable is a finite set of discrete labels fixed at link time. A region admits graded membership and an open set. So <code>transduce(signal, modality)</code> asks the caller to supply what the signal already carries — what a thing is falls out of where it lands. The modality parameter is a kind-tag, and a registry keyed on it is a lookup table doing by string what geometry does by nearness.</p>
<h3>Types are discovered, not declared</h3>
<p>Reification crystallizes a densely co-wired neighbourhood into a first-class node — the neighbourhood <em>is</em> the name that was missing. Every other family requires a human to see the abstraction in advance and write <code>class Foo</code>. Here the instances arrive and the type falls out, by measurement rather than by insight.</p>
<h3>Enumeration becomes unnecessary</h3>
<p>Five ingest functions differ only in how bytes are acquired — one operation wearing five surfaces. 356 branches in <code>engram_activate_inner</code> are not 356 behaviours. Cyclomatic complexity is a count of the places comprehension ran out and was replaced by an <code>if</code>; where the concept is expressible, the count collapses instead of being redistributed.</p>
<h2><spanclass="n">03</span>The shape of the language</h2>
<p>Geometry first-class gives El three layers, and it holds all three — which is why there is no separate database driver and no impedance boundary to manage.</p>
<divclass="flow">
<div><spanclass="k">afferent</span><h4>Transduce</h4><p>Signal in, geometry out. Decomposition into components and relations — never conversion to a point. Realizers are ordinary El functions, so a new modality never requires a runtime patch.</p></div>
<div><spanclass="k">substrate</span><h4>Geometry</h4><p>Meaning as position; relation as distance. Held as values in the language and persisted in the graph. One coordinate system, so entities are commensurable and the operators compose.</p></div>
<div><spanclass="k">efferent</span><h4>Realize</h4><p><code>plan(frame) → realize(spec, profile)</code>, where a surface <em>is</em> a profile. Text, speech, music, image are profiles of one projection — and so is source code.</p></div>
</div>
<p>The efferent side is why the recursive property below is possible at all: if source is a surface, then emitting a corrected file is projection, and the file becomes an artifact of the geometry rather than the thing you edit.</p>
<h2><spanclass="n">04</span>Decomposition is by faculty</h2>
<pclass="lede">Not by file, module, or subsystem — by what the system does.</p>
<p>Each faculty is a concept. Where it has no home in El it leaks: into C, into a Swift binary, into a shell script with a <code>curl</code> timeout, into a convention nobody performs. State below is measured, not asserted.</p>
<divclass="card scroll">
<table>
<thead><tr><th>Faculty</th><th>State</th><th>Measured</th><th>Where it leaked</th></tr></thead>
<tbody>
<tr><tdclass="f">Ingest <spanclass="tag">take in</span></td><tdclass="dead">dead</td><tdclass="m">2 min → 0 nodes</td><td>separate process uploading bytes over HTTP to a process with direct fs access; five functions where there is one</td></tr>
<tr><tdclass="f">Recall <spanclass="tag">remember</span></td><tdclass="dead">dead</td><tdclass="m">self ranked 8th</td><td>lexical substring scan; empty on 23 of 24 multi-token queries</td></tr>
<tr><tdclass="f">Transduce <spanclass="tag">perceive</span></td><tdclass="dead">dead</td><tdclass="m">1 node, 0 edges</td><td>intake flattens signal to a point; <code>realized:false</code>; caller must declare the modality</td></tr>
<tr><tdclass="f">Think <spanclass="tag">reason</span></td><tdclass="dead">dead</td><tdclass="m">direction [0,0,…]</td><td>null gradient from any anchor and any faculty, byte-identical; confidence at the uninformed prior</td></tr>
<tr><tdclass="f">Realize <spanclass="tag">express</span></td><tdclass="part">partial</td><tdclass="m">13-word lexicon</td><td>organ was 939 lines of Swift beside the language; voice read from a file path</td></tr>
<tr><tdclass="f">Persist <spanclass="tag">endure</span></td><tdclass="ok">live</td><tdclass="m">13,562 / 13,562</td><td>works — every signal placed in geometry at intake, no backlog</td></tr>
<p>El's compiler is written in El. Every concept the language gains, the compiler can then be written <em>in</em> — so the tool improves the tool, and <code>codegen.el</code> at 4,661 lines gets shorter as the language gets better at expressing what it does. The fixpoint — stage2 ≡ stage3, byte-identical — makes each turn provable rather than hopeful, and the verifier answers in <spanclass="mono">2.9s</span>.</p>
<p>This sets the ordering criterion, and it is not size of payoff:</p>
<blockquote>Order by leverage on the <em>next</em> iteration. Which concept, added to El, most increases the ability to add the following one?</blockquote>
<p>A small early gain that compounds beats a large one that does not. And it bounds itself correctly — unbounded in depth, bounded in rate, because nothing lands that the compiler and the fixpoint have not passed.</p>
<h2><spanclass="n">06</span>What has no home yet</h2>
<p>Reserved in the lexer, no parse form. These are not a feature backlog — they are the concepts the architecture above requires and does not yet hold, which is why each is currently a convention or a block of C.</p>
<tr><tdclass="m">retry · times · fallback · reason</td><td>resilience</td><td>a shell script with a 10s <code>curl</code> timeout; 254 restarts in 3 days</td></tr>
<tr><tdclass="m">requires · deploy · to · via · target</td><td>deployment</td><td>YAML in another repository</td></tr>
<tr><tdclass="m">sealed</td><td>capability scope</td><td>consent checks written by hand</td></tr>
<tr><tdclass="m">protocol · impl</td><td>one operation, many realizations</td><td>five ingest functions; eight faculty routes on one builtin</td></tr>
<tr><tdclass="m">activate · where</td><td>retrieval</td><td>traversals written by hand</td></tr>
<tr><tdclass="m">parallel · trace</td><td>concurrency</td><td>pthreads in C</td></tr>
</tbody>
</table>
</div>
<p>Plus, from the spec's own status: annotations parsed and skipped, <code>match</code> parsed and emitting nothing, <code>?</code> a no-op, <code>%</code> unlexed, structs as <code>ElMap</code>, enums as strings, selective import unenforced.</p>
<h2><spanclass="n">07</span>Open</h2>
<divclass="q"><b>What does a declaration bind to, exactly?</b><span>If <code>cat</code> names a region that shifts and completes against context, what is written at the declaration site and what is resolved at use? This is the centre and it is unspecified.</span></div>
<divclass="q"><b>Is the faculty list right?</b><span>Seven, derived from what broke. Derived-from-failure is a biased sample — it finds what is loud, not what is absent. Which faculty is missing entirely and therefore never failed?</span></div>
<divclass="q"><b>Which concept has the highest leverage on the next turn?</b><span>The prologue/epilogue seam (§19.3 names it as the prerequisite; its stated blocker has expired; it collapses 62 + 10 convention sites), <code>protocol</code>/<code>impl</code>, or resolution itself. The §05 criterion should decide this, not preference.</span></div>
<divclass="q"><b>What seam makes cognition non-optional?</b><span>"Use the ops" is itself a convention — present every turn, enforced by nothing, ~100% failure across a full session. A stronger instruction is still a convention. What makes reasoning outside the substrate <em>fail</em>, the way <code>@manager</code> makes <code>dharma_emit</code> outside the boundary a compile error rather than a lint?</span></div>
<hr>
<pclass="foot">Every number here is measured or quoted from <code>lang/spec/language.md</code>. Nothing is inferred and presented as fact. El is self-hosting: all of this can change and be rebuilt.</p>
**Running list.** Append as decided. Started 2026-08-17.
**The test:***is this an arbitrary convention, or is it a relation?*
Conventions were agreed by people and could have been otherwise — a RIFF header could
have used a different magic number. Nothing derives them; they must be written down.
Relations are not agreed. Distance is distance. Anything whose answer is *where is this
relative to that* is geometry, and writing it as code is the error the whole effort is
correcting.
**Second test, for the hard cases:** *if I write this as code, am I encoding in
`if`-statements a distinction the geometry was built to hold?* If yes, it's geometry.
---
## Pure geometry
| Thing | Because |
|---|---|
| Meaning | position |
| Grounding / standing | the weight on the edge — a magnitude, not a computation |
| Learning | standing changing over time |
| A gap | low standing |
| Wonder | a gap with a pull weight |
| Type checking | is this position in that region — distance |
| Dispatch | position, not a tag |
| Recall | re-origining at a region; projection, not replay |
| Reasoning | traversal |
| Deduction | containment. There is no procedure |
| Counting | a position, not a loop's output |
| Similarity / difference / residue | subtract |
| Analogy, metaphor, skill transfer | change of basis |
| Negation, sarcasm | reflect an axis |
| Empathy | translate the origin |
| Reframe | rotate the frame |
| A lens | project onto an axis |
| Rhyme | distance in phonetic space |
| Humour | intersection of regions — fart-meaning ∩ funny ∩ form |
| Idiom detection | the whole unit sits farther out than its parts |
| Self | a world-tube — a trajectory through the manifold |
| Consolidation | episodic → semantic promotion |
| Reification | dense regions cohering; runs on the beat, has no caller |
| Cross-cutting concerns | **dissolved** — a hold is a *neighbour*. Adjacency, not tracking. **Implemented 2026-08-17**: a construct declares what runs at a crossing, and it resolves at execution — see the runtime seam. |
| Effects | topology. `reach_out` is bounded by `detect_gap` and `verify` because those are its edges |
| Capability | position relative to a boundary. In C it is already spelled `const` |
| The AST | a projection of geometry into a tree — a surface, not the centre |
| Source code | a surface, like text, audio, image |
## Must be code
| Thing | Because |
|---|---|
| Sensors — mic, camera, file read, socket | the physical touch. I/O is where the world arrives |
| Byte formats — RIFF, PNG chunks, `MThd`, OOXML | arbitrary convention. A committee chose the magic numbers |
| CRC32 polynomial, Adler32, zlib framing | same — agreed constants, derivable from nothing |
| Cosine, distance, the float arithmetic | the machinery that *walks* the geometry is not itself geometry |
| Arena, refcount, allocator | bookkeeping for the **representation**, not for the positions |
| Locks, threads, publication boundary | the hardware is code. **Ordering is not** — see Answered, above. Coordination is required only where state is non-monotone. |
| WAL, page layout, ARIES recovery | durability against a physical device that can lose power |
| Emission — writing C or JS text | the final surface has to be *typed out* by something |
| OS interaction — launchd, spawn, signals | outside the system by definition |
| Device realizers — `el_audio_darwin.m`, `el_capture_darwin.m` | OS frameworks. Correctly already isolated, zero network |
---
## The ones I would have written as code, and was wrong about
Recorded because the error has a pattern and the pattern is the point.
| Thing | What I reached for | What it is |
|---|---|---|
| Rhyme | a rhyming dictionary, or an API call | distance between rime tails |
| Fart onomatopoeia | a 30-element string literal | an intersection of three regions |
| "Funny" | a scorer with `if`-statements | a relational neighbourhood grounded in a voice |
| Representation vs description | a hardcoded blacklist containing `raspberry` | falls out of lexicon membership × phonetic comedy |
| Video | a codec, sized as a project | one more surface profile |
| Type checking | a phase between parse and emit | reading a distance that already exists |
| Grounding | a call site, an obligation, a discharge | it has no caller. It just runs |
| N transducers, N realizers | one component per modality | zero of each. Sensors and bases at the skin |
**The pattern:** every one is *encoding in code a distinction the geometry was built to
hold.* The tell is that the code version is a **fixed enumeration** — a list, a table, a
blacklist, a set of branches — and the geometry version is a **measurement**.
If the implementation contains a literal set of the right answers, it is in the wrong
column.
---
## Answered
| Thing | The answer |
|---|---|
| Concurrency | **Ordering is geometric.** Causality is a partial order (Lamport 1978); a total order is an arbitrary extension of it and "cannot be depended on to imply a causal relationship." Programming languages force you to write a total order, so authoring *invents* constraints the problem never had — and every lock, barrier, fence and consensus protocol is apparatus for recovering the partial order destroyed at authoring time. CALM (Hellerstein/Alvaro, proven by Ameloot et al.): a program has a consistent coordination-free implementation **iff it is monotone**. What breaks monotonicity is destructive update. **Coordination is the price of forgetting.** |
| The module system | **Premature — the partition is a filesystem path, not a neighbourhood, and there is no namespacing at all.**`import` is textual inlining (guarded against double inclusion); when a `.elh` header exists the header is inlined instead and symbols resolve at C link time, so linking is real and delegated to C. Two modules defining `helper` emit two C functions into one translation unit. Linking barely survives the *path* partition, so whether it survives a neighbourhood partition cannot yet be asked. |
| Numeric literals | **The numeral is convention; the number is a position — and a bare `3` is a MAGNITUDE WITH NO AXIS.**`int_to_str` was already form 1: nothing determines that twelve is written `1` then `2`. But a literal is not a position until something gives it a direction, which is why `3.days` needs a calendar. Measured consequence: `Duration + Int` was refused ("an Int carries no unit") while `Instant + Int` compiled to raw `(t + 3)` and reported clean — silently moving a point by an unspecified amount. The rule was simply never written. Now: `t + 3` is refused, `t + 1.hour` is accepted, because `.hour` supplies the axis. |
| Parsing | **A grammar is a basis; parsing is transduction onto it.** The lexeme→token map is convention (`fn` could have been `def`); shape recognition is a region; the byte traversal is irreducible, like every other traversal. Three things favour *region* for the act: ambiguity (`a * b` needs context — a grammar resolves it with the lexer hack, a region by neighbourhood), error recovery (nearest-match is free), and precedence, which is ordering along an axis with a conventional parameter. **But the SHOULD gate refuses the obvious move:** the keyword table stays code, because the set is closed by the language definition and the lexer runs before the program is understood, so a program can never declare its own keywords. Externalising it costs I/O per compile for zero flexibility — the same verdict as `is_digit` in ASCII. What was actually wrong: 5 of 46 keywords were consumed by nothing, and using one silently miscompiled. |
| Error handling | **`grounded: false` covers not-knowing; it does not cover failed.** Standing is a *signed* component: `> 0` supported, `= 0` unknown, `< 0` contradicted. Not-known and known-false are opposite directions on one axis and a boolean cannot tell them apart. `inhibitory` as an int32 flag is that sign wearing a boolean. |
## Fourth proof form
**4 — ADVERSARIAL EXACTNESS.** Where approximation is a break, geometry is
excluded. A cryptographic hash is a *deliberately structure-destroying* map:
near inputs land at maximally uncorrelated outputs. Geometry is the claim that
near things stay near — a manifold that approximated SHA-256 would *be* a break
of SHA-256. Signature verification is the same: 0.99-valid is invalid. And
X25519 **is** geometry, a group on an elliptic curve, which is precisely why it
must be code: its security is the hardness of moving in that geometry.
**Form 1 no longer survives as a verdict.** Every row it justified turned out to
be a *basis*, not a capability. RFC 8259 fixes where the commas go — that is a
surface, and projecting onto a surface is geometry. A convention describes the
basis you project onto; it never describes an act.
Every change to El on `iteration-1` was produced by one loop, run repeatedly:
```
Ishikawa → scientific method → Six Sigma → repeat
```
- **Ishikawa** — name the root cause, not the symptom. *Why is this table here?*
never *why is this table ugly?*
- **Scientific method** — state a hypothesis, **commit predictions before
running**, then run it in an isolated worktree and grade every prediction
including the ones that failed.
- **Six Sigma** — eliminate the defect *class*, then add a control so it cannot
silently return.
## The organising finding
**Predictions that came back FALSE were worth more than the ones that held.**
Nineteen cycles, sixty-one predictions. The eleven that failed produced every
significant result:
| Failed prediction | What it found |
|---|---|
| "the arity table has drifted from the header" | Zero drift — but **110 functions had no entry at all**. The table was not wrong, it was 40% incomplete. |
| "codegen drops below baseline" (×4) | The **traversal is irreducible**. Walking an AST to find calls does not move no matter who decides. Only the rule and the judgment leave. |
| "guards cannot refuse through the seam" | One line, and refusal works. Six compile-time kinds were unnecessary. |
| "C forbids the struct redefinition" | C allows shadowing — and a *different* defect surfaced: an exit injection emitted with an empty target. |
| "routing el_bin_lookup through the gate fixes the SIGSEGV" | It did not. The **fallback** was the hazard: `strlen()` on an integer. I would have shipped the wrong fix and called it verified. |
A prediction that only ever confirms is a demonstration, not a test. One cycle
was run **without** committing predictions first — `async-half-expressible` —
and it produced a rigged result: `pthread_join` immediately after
`pthread_create`, with the word `DEFERRED` printed by the test itself. It had to
be discarded and re-run.
## Layout
```
cycles/ one file per loop, numbered in order, named for the DEFECT
findings/ what the cycles produced, cross-cut by kind
```
## Scoreboard
```
cycles run 19
predictions committed 61
predictions FALSE 11 ← the useful ones
silent miscompilations found 4
security-relevant defects 2
architecture questions closed 5
defects in my own measurement 4
```
Every cycle verified the same three things before landing: the compiler
self-hosts byte-identically (gen2 == gen3), the native suite passes, and the
integration harnesses pass. A cycle that could not show all three did not land.
Each is one `Ishikawa → scientific method → Six Sigma` loop, run in an isolated
worktree so a wrong answer cost nothing. Named for the **defect**, not the fix.
| # | Cycle | Root cause | Predictions | Landed |
|---|---|---|---|---|
| 01 | [constructs-have-nowhere-to-be](01-constructs-have-nowhere-to-be.md) | a construct had nothing to BE, so its meaning lived in the emitter | 3/3 | yes |
| 02 | [a-construct-cannot-refuse](02-a-construct-cannot-refuse.md) | injection discards the target's result; no form said no | 4/4 | yes |
| 03 | [the-wrapper-was-conditional](03-the-wrapper-was-conditional.md) | exit injection needed compile-time knowledge only because the wrapper was conditional | 3/4 | yes |
| 04 | [c-has-no-closure-syntax](04-c-has-no-closure-syntax.md) | "C has no closures" taken as a fact about what is possible | 5/7 | yes |
| 05 | [the-emitter-discards-what-it-knows](05-the-emitter-discards-what-it-knows.md) | codegen sees every construct relation and throws it away | 5/5 | branch |
| 06 | [the-crossing-resolves-at-emission](06-the-crossing-resolves-at-emission.md) | the binary has no table to consult | 3/4 | yes |
| 07 | [invocation-is-not-composable](07-invocation-is-not-composable.md) | the wrapper called the target directly | 5/5 | yes |
| 08 | [the-emitter-adjudicates](08-the-emitter-adjudicates.md) | a prohibition had nowhere to live but a `#error` | 4/5 | yes |
| 09 | [policy-inside-the-compiler](09-policy-inside-the-compiler.md) | a program cannot declare its own restrictions, so the tier policy was compiled in | 4/5 | yes |
| 11 | [one-type-erases-the-return](11-one-type-erases-the-return.md) | `el_val_t` means the header cannot say `now()` returns an Instant | 4/5 | yes |
| 12 | [judgment-lives-with-knowledge](12-judgment-lives-with-knowledge.md) | the emitter knows the types, so it also judged them | 5/5 | yes |
bash -c D=$(mktemp -d); bash /Users/will/Development/neuron-technologies/foundation/el/tools/evidence/verify-manifest.sh "$D"; echo "exit=$? on a directory containing NO manifests at all"; rmdir "$D"
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.