b5d1e53902dd38ca88769930de2ea1524cbb3cc4
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
19cc99e57d |
store: judge memory pressure by swap RATE, not swap level
El SDK CI - dev / build-and-test (pull_request) Failing after 11m59s
The guard I added minutes ago checked swap availability as a level
(avail < total/8 -> report zero available). That is the wrong signal, and the
same host proved it twice within minutes:
47.65 / 48.00 GiB swap used, 2047 swapouts/s -> genuinely thrashing
26.67 / 28.00 GiB swap used, 0 swapouts/s -> healthy, 15.6 GiB free
Both are ~97% "used". macOS grows swap files on demand and trims them lazily,
so the level says almost nothing about now — it is a high-water mark. The level
check calls the second state an emergency and starves the pool for no reason,
which is its own failure mode: a guard that fires on healthy machines gets
disabled, and then guards nothing.
What separates the two is whether pages are moving. So sample the swapout
counter across calls and judge the delta:
- > 200 pages/s (~3 MiB/s) sustained outward paging => report zero available;
callers refuse to grow and pc_relieve_pressure hands frames back.
- The first call primes the baseline and reports no pressure. One sample
cannot have a rate, and inferring one from a single reading is exactly the
mistake this commit removes.
Measured thresholds, not guessed: idle sat at 0/s, recovery burst hit 24,845/s
while the compressor drained (transient, correctly not a growth decision since
growth is only evaluated on eviction passes), and real thrash held ~2000/s.
200/s sits clearly above noise and far below either.
The compressor-footprint subtraction stays: that RAM is genuinely spoken for
regardless of paging rate.
|
||
|
|
e52415f0e0 |
store: bound the pool by AVAILABLE memory and let it shrink
El SDK CI - dev / build-and-test (pull_request) Failing after 11m47s
The adaptive budget I added an hour ago could only grow, and grew toward a
share of TOTAL ram (80%, ~38 GiB on a 48 GB host). That is a memory leak with
extra steps: total never shrinks when other processes need memory, so the pool
had no way to notice it was starving the machine it runs on. Deployed briefly;
caught as memory pressure on the host.
A control loop with only one direction is not a control loop.
- pc_available_ram(): free + inactive + purgeable via host_statistics64 on
Darwin, MemAvailable on Linux. Availability is the quantity that moves when
the machine is under pressure; total is not. Returns 0 when it cannot be
read, and callers then refuse to grow — a cache is never worth swapping the
host, so unknown means no.
- Growth is bounded by availability minus a free-memory floor (2 GiB default,
ENGRAM_POOL_FREE_FLOOR_MB), not by total. The share-of-total ceiling stays
as a second bound and drops 80% -> 50%.
- pc_relieve_pressure(): the missing direction. On every eviction pass, if
available memory is under the floor, hand back ~25% of held frames; the
resident set follows on the next pass so the memory is actually returned
rather than merely re-labelled. Counted as adapt_shrinks alongside
adapt_grows so both directions are visible in the same report.
- pc_default_cap() also clamps the STARTING budget to what is spare right
now, so a cold boot on a loaded machine does not open at a size the host
cannot afford.
Verified on a 48 GB host: engram boots in ~30s, RSS settles at 2.22 GiB (the
store's actual size, resident, not creeping), 0.0% CPU, 13,439 nodes / 37,670
edges, embeddings complete. Guard reports 9.71 GiB available against a 2.00 GiB
floor — 7.71 GiB of headroom it is permitted to use and no more.
|
||
|
|
e917b3d439 |
store: make the buffer pool sense its own state and correct from it
El SDK CI - dev / build-and-test (pull_request) Failing after 14m35s
Follow-on to the edge write barrier. That fix removed the full-store walk;
this one makes the pool able to notice if anything like it happens again.
WHAT WENT WRONG, precisely: the pool thrashed the live engram to a standstill
twice on 2026-08-15 and said nothing. From outside it was indistinguishable
from "busy loading" — 100% CPU, flat RSS, no output — so four wrong theories
got tried (bad binary, corrupt snapshot, WAL replay, feature flags), each
costing a deploy or a rollback. The whole time, hits/misses/evictions were
already being counted in PgCache, and the struct comment read:
/* stats (introspection only — never affect semantics) */
That comment was the bug. Self-measurement treated as decoration is why the
pool could not correct itself and why no one outside could see what it was
doing. A system that cannot read its own state cannot correct, and neither can
anyone watching it.
- pc_adapt_budget(): the loop, closed. Over a sliding window, evictions
running at a large fraction of accesses WHILE reuse is real means the
working set exceeds the budget — so grow it, geometrically, bounded by a
LIVE re-read of physical memory. Evictions alone are not pressure (a scan
evicts and never returns); evictions with reuse are. An explicit
ENGRAM_POOL_FRAMES still wins — an operator override must not be silently
overruled.
- Budget derived, not declared. A constant cannot be right: 16 GiB of frames
is arbitrary on a 48 GB host and suicidal on a 16 GB one. Even "60% of RAM
at startup" is a guess about the future — it cannot know the store grew or
the machine changed. Hence the live re-read.
- pc_report(): ONE structured emission carrying the entire sensed state,
through emit_log — El's existing telemetry, already exporting to OTLP.
Deliberately not a function per stat, and deliberately not a bespoke
/api/pool endpoint: both make observability something hand-written per noun
instead of the uniform mechanism every component already has.
- engram_pool_stats_json(): the same state readable live, wired through the
normal builtin path (codegen arity + el_seed wrapper), so the pool can be
observed in real time rather than reconstructed afterward from a stack
sample.
Verified: with the exact configuration that took production down
(ENGRAM_POOL_FRAMES=65536 → 1 GiB cache against a 2 GiB store) the engram boots
clean and serves — 0.0% CPU, 13,436 nodes / 37,663 edges, embeddings complete —
and NO pressure event fires, because the barrier removed the walk that caused
it. The controller is defense in depth; the barrier is the fix.
|
||
|
|
777ccc02f0 |
store: extend the durable-hash write barrier to edges (kills the full-store walk)
El SDK CI - dev / build-and-test (pull_request) Failing after 10m45s
Checkpointing pushes the ENTIRE resident graph through store_put_node and
store_put_edge (see engram_store_checkpoint). Nodes were cheap: a durable-hash
compare skipped unchanged records with zero page I/O. Edges had no barrier at
all — struct comment at PgCache.barrier_on even says "node durable-hash
barrier" — so every edge was rewritten on every checkpoint, and each rewrite
runs the idempotency probe max_page_lsn_for_id -> btree lookup -> page_read.
Edges outnumber nodes ~3:1 here (37,663 vs 13,436), so routine checkpointing
degenerated into a FULL-STORE WALK in id order: random page access across the
whole 2 GiB store, repeated, overwhelmingly to rediscover nothing had changed.
LRU is worst-case under exactly that pattern — it evicts the page it is about
to want — so once the page cache was smaller than the store, the walk collapsed
into thrashing: 100% CPU, flat RSS, no forward progress, port never bound.
That took the live engram down twice on 2026-08-15.
The walk is the defect. Sizing the cache to survive it treats the symptom.
Changes:
- dh_edge_hash(): edge counterpart of dh_node_hash, with a kind discriminator
byte so an edge can never collide with a node of the same id in the shared
map. created_at/updated_at/last_fired are excluded deliberately: last_fired
is touched by activation without changing what the edge IS, and folding it
in would defeat the barrier on precisely the hot edges that most need it.
- store_put_edge(): barrier check + dh_set on success, mirroring
store_put_node exactly.
- store_scan_edges(): seed the barrier map from on-disk truth at load, so the
FIRST post-boot checkpoint already skips unchanged edges. store_scan_nodes
already did this and its comment says why; edges were simply never done.
Verified: with the exact configuration that killed production
(ENGRAM_POOL_FRAMES=65536 -> 1 GiB cache against a 2 GiB store), the engram now
boots clean and serves — LISTENING, 13,436 nodes / 37,663 edges, embeddings
complete, 0.0% CPU, RSS 1.14 GiB (cache resting at its budget rather than
thrashing against it). Same small cache, same store, no walk.
|
||
|
|
01826421c4 |
seam: implement decorated-fn boundary auto-emit; prove on clone
Will waived diff review -> build it for real. Add engram_boundary_beat() to the runtime (afferent counter++ + engram_chrono_tick + engram_strengthen(self-anchor) + dharma_emit) and two act-stats counters (aff_boundary_ops, dharma_emits). codegen cg_fn injects ONE engram_boundary_beat(op) at the entry of every @manager/@accessor fn (fn_has_decorator, so it fires under @route @manager too) — a decorated op self-reports with ZERO hand-written instrumentation. Rebuilt elc self-host + the cognition engram in the worktree; ran it as the clone daemon on :8900. Proof (/api/boundary-proof, @manager, empty body, 5x): aff_boundary_ops 0->5, dharma_emits 0->5, self activation_count 1510->1513, chrono stamp advanced. Brought in feat/cognitive-architecture engram runtime+server for the build. strengthen = activation bump (not content/edge write) -> identity protection intact. Live :8742 untouched; no push, no cutover. |