bigmerge bacaf3d39c
El SDK CI - dev / build-and-test (pull_request) Failing after 4m49s
engram: reconcile M8 HNSW vindex (#109) onto current dev, restore 3 fixes the branch predated
Lands feat/reframe-region-setop (PR #109: native set-based reframe_region,
decorator-as-seam @route port, teacher-summon, and the M8.1 activate-latency
work — lazy-memoized cosq via eg_cosq_at + engram_vindex HNSW-accelerated
seed discovery + vindex_harvest_from_store/vindex_bench oracle) onto dev's
actual current HEAD, plus engram-tiered-storage's still-unique test suite.

RECONCILING #109 WITH engram-tiered-storage (M4-M10 HNSW/geometry/reason/
verify work): not a two-way merge. engram_vindex.c's HNSW core (search_layer/
select_neighbors/prune_links/insert) is BYTE-IDENTICAL between the two
branches; #109's copy is a strict superset (adds vindex_harvest_from_store,
used by vindex_bench.c's brute-force-vs-HNSW oracle). engram_reason.c and
engram_verify.c are also byte-identical. #109's own branch point already
carried engram-tiered-storage's M4-M10 lineage forward, so there was nothing
left to merge into #109 for those files. The one thing engram-tiered-storage
had that #109's tree dropped: its full test suite (test_vindex.c,
test_geometry.c, test_reason.c, test_verify.c, test_m7_traversal.c, the
interoception P0-P5 tests, bufpool/compaction tests, and their run_*.sh
harnesses) — ported over here unchanged.

WHY THIS NEEDED HAND RECONCILIATION, NOT A MECHANICAL MERGE: #109's branch
forked from dev on 2026-08-14 15:40 (before restructure-adjacent history
diverged the file's merge-base for `git merge` — it presented as an add/add
conflict). A straight two-dot diff (dev tip -> PR tip) applied cleanly, but
it silently reverted THREE dev fixes landed on 2026-08-14/15, after the
branch point, that the PR's diff had no way to know about:

  1. qgate rescale (2026-08-14 self-review): PR's lazy eg_cosq_at rewrite of
     the query-aware propagation gate dropped the shift-and-floor rescale
     about ENGRAM_EMBED_S0 (measured: unrelated-pair median 0.562->raw gate
     0.67, i.e. "a small tax, not a gate"). Restored the rescale, wrapped
     around the lazy accessor -- the PR's actual improvement (WHEN cosq[oi]
     is computed) is orthogonal to WHAT it gates on and both are kept.
  2. Eviction cause decomposition (2026-08-14 self-review): dev decomposes
     wm_evicted into evict_floor/evict_cap/evict_bll so WM churn is
     diagnosable (identity: evicted == floor+cap+bll+dup_wm+dup_wm_global).
     PR's tree predates this and dropped all three counters + their JSON
     stats fields. Restored declarations, all 4 direct increment sites, the
     eg_wm_carry_over bll increment, and the act-stats JSON fields --
     alongside (not instead of) the PR's own P4 afferent / API-reshape
     counters already in that same struct/JSON.
  3. Hebbian link-formation selection (2026-08-15 self-review, TODAY): dev
     selects the STRONGEST qualifying candidate for consolidation each call;
     PR's tree predates this and reverted to hash-slot order (arbitrary wrt
     association strength) for edge formation -- the one path that writes
     PERMANENT structure. Restored the strongest-candidate while-loop,
     keeping the PR's own genuine improvement at that site
     (engram_adj_on_edge_added incremental-index append instead of a bare
     adj_dirty=1 full-rebuild flag).

engram/src/server.el's 3-way conflicts (autoconnect_on/ise_offgraph_on env
flags, /api/nodes connected-count in responses) were pure additive: dev's
side was empty, PR's side added the feature. Took PR's side whole.

VERIFIED (nsbx sandbox only, live :8742/:7770 never touched):
  - cc -std=c11 -O2, clean link against the real engram/src/server.el via
    elc, zero errors.
  - vindex_bench (built standalone, read-only harvest) against the real
    production store clone (13,671 embedded nodes, 768-dim nomic-embed-text):
    recall@10 = 1.0000 at ef 64/128/200; HNSW search 0.28-0.79ms/query vs
    2.03ms/query brute-force oracle (2.6x-7.2x). HNSW build itself: 46.5s
    for the full 13,671-node set -- see the flagged risk below.
  - Booted the reconciled binary in an isolated nsbx sandbox (:8905, cloned
    snapshot of the live store, 13,424 nodes / 37,656 edges) and called
    /api/activate for real: first call after boot 41.5s (pays the one-time
    HNSW build inline -- matches the standalone bench), second/third calls
    356ms/605ms, no crash, correct results, act-stats JSON (including the
    restored evict_floor/cap/bll fields) reads correctly.

KNOWN RISK TO FLAG BEFORE ANY LIVE CUTOVER (not fixed here; out of scope for
this dev-only land per instructions not to touch :8742/:7770): eg_vindex_sync
builds the HNSW index synchronously, inline, on the first engram_activate()
call after every process start (or index invalidation). On the real node
count that is a ~46s blocking stall on a single-threaded server -- the first
request after every restart (or its concurrent siblings) waits the full
build. Recommend a background/incremental build (or a bounded per-call build
budget) before this ever reaches the live daemon. See PR description / final
report for the fuller writeup.
2026-08-15 16:46:44 -05:00

El

A self-hosting, statically-typed language that compiles to C — built around a graph-native runtime instead of a database driver.

El is the execution substrate for the Neuron agent runtime, the DHARMA network, and the Engram knowledge graph. This repository is the monorepo for the whole stack: the language itself, the graph memory engine it's built to talk to natively, and the tools (package manager, IDE, UI framework, diagramming) built on top of it.


Why El exists

Every other language treats persistent, associative state as something you reach for through a driver — a SQL client, an ORM, a Redis library bolted on from outside. El inverts that: graph operations (engram_*) are runtime primitives, on the same footing as string or list operations. There is no separate database driver because the database is not separate.

El has four defining properties:

  1. Self-hosting compiler. The compiler (lexer.el, parser.el, codegen.el, compiler.el) is written in El. It compiles El source to C, which cc compiles against a fixed runtime into a native binary. A Rust genesis compiler bootstrapped the first iteration; the self-hosted binary at lang/dist/platform/elc has been the canonical compiler ever since — every binary in dist/platform/ was produced by an earlier version of itself compiling el-compiler/src/. The chain is auditable: source is the ground truth, not the binary. See lang/BOOTSTRAP.md for the full recovery path if that binary is ever lost.
  2. C compilation target. Every compiled program is plain C11. Every El value is el_val_t (int64_t); strings are heap pointers cast through it. Functions become C functions; top-level statements become main().
  3. Graph-native runtime. The runtime provides first-class graph operations over an in-process Engram store — no separate DB driver, no ORM.
  4. DHARMA-aware identity. A cgi block declares a program's DHARMA identity at compile time. The runtime resolves identity before user code runs, so dharma_* calls have a stable principal and channel surface throughout.

Architecture map

                     ┌─────────────┐
                     │    lang     │   El compiler + C runtime
                     │ (El itself) │   everything below is written in it,
                     └──────┬──────┘   or compiles down through it
                            │
              ┌─────────────┼─────────────┐
              │             │             │
       ┌──────▼─────┐ ┌─────▼─────┐ ┌─────▼─────┐
       │   engram   │ │    epm    │ │    ide    │
       │ graph/mem  │ │  package  │ │  editor + │
       │  substrate │ │  manager  │ │    LSP    │
       └──────┬─────┘ └───────────┘ └───────────┘
              │
      ┌───────┼────────────────┬─────────────────────┐
      │       │                │                     │
┌─────▼───┐ ┌─▼──────────┐  ┌──▼──────────┐    ┌─────▼──────┐
│   elp   │ │ ql         │  │  ui         │    │   arbor    │
│  NLG /  │ │engram-el.  │  |spreading-   │    |arbor       │
│ 31 langs│ │studio+tests│  |activation UI│    |diagram lang│
└─────────┘ └────────────┘  └─────────────┘    └────────────┘

lang is the foundation — the compiler and C runtime everything else builds on. engram is the graph-native memory/state engine that gives El its identity (property 3 above). Everything else is either a tool for working with El (epm, ide) or a system built on top of Engram's graph model (elp, ql, ui, arbor).


Repository layout

lang/ — the El language

The compiler and runtime. Self-hosting: elc-cli.elcompiler.ellexer.el / parser.el / codegen.el / codegen-js.el, textually inlined and compiled in one pass. Compiles to C11 and links against el-compiler/runtime/el_seed.c, a hand-maintained OS-boundary layer (libcurl HTTP, pthreads, filesystem, arena allocation) — everything else in the runtime is native El (runtime/*.el).

Two layers to know: El programs (.el files — where nearly all work belongs) and the C seed (el_seed.c — edit only for genuine OS-level access; never re-implement what El can already express).

Current status (single source of truth: lang/spec/language.md): lexer/parser/codegen and the C runtime's core (I/O, strings, math, lists, maps, filesystem, args) are implemented. In flight: % operator, match-statement codegen, ? nil-propagation, cgi block parsing + DHARMA identity resolution, VBD role enforcement (@manager/@engine/@accessor), the real engram_* and dharma_* runtimes (currently stubs), and libcurl-backed http_get/http_post/http_serve. Bitwise operators, ??, and as casts are explicitly not in this language.

Key docs: AGENTS.md (agent-facing orientation), BOOTSTRAP.md (compiler recovery from scratch), spec/language.md, spec/codegen-js.md.

engram/ — graph intelligence substrate

A local-first memory substrate for accumulating intelligence, and the reason El's runtime doesn't need a database driver. Rust core (engram-core, engram-ffi) exposed to El and other languages (Kotlin, TypeScript/WASM, Go bindings).

The model: retrieval is spreading activation, not query. You name seed nodes and a query embedding; activation propagates outward through weighted edges, attenuating multiplicatively per hop (strength = parent_strength × edge_weight × target_salience × cosine_sim), gets pruned below a threshold, and the top-N nodes by activation strength come back. Storage and retrieval are the same structure — the way long-term potentiation works in biological memory, not the way a relational or vector database works.

Nodes live in four tiers (Working / Episodic / Semantic / Procedural, mirroring prefrontal / hippocampal / neocortical / cerebellar memory) and migrate between them based on salience decayimportance × recency-decay × log(activation_count). Forgetting is adaptive pruning, not a bug: unreinforced memories stop competing for attention without being deleted.

Backed by sled (embedded, local-first, no daemon) with flat cosine scan for vector search — deliberately simple until scale demands an HNSW layer. Full API and design rationale in engram/README.md.

elp/ — Engram Language Protocol

Bidirectional engine mapping between Engram semantic forms and natural-language surface text, across 31 languages — from Spanish and Japanese through historical/liturgical languages (Old Norse, Sanskrit, Sumerian, Coptic, Akkadian, Ge'ez). Compilation order runs language-profile + vocabulary → per-language morphology-*grammarrealizersemanticselp. This is what lets an Engram graph node round-trip to and from readable text in any of those languages.

epm/ — El Package Manager

Manages vessels (El's package unit): publish, install, resolve dependencies. Vessels are stored in Engram as graph nodes, not files in a registry index — epm reads the local manifest.el, talks to Engram over HTTP, and writes resolved vessels to .epm/vessels/. Source: registry.el, install.el, update.el, manifest.el.

ide/ — El IDE

Three vessels: el-ide-server (HTTP backend — file ops, build/run, LSP bridge, plugin host, settings), el-lsp (the language server — completion, hover, diagnostics, outline, format, type graph), and el-plugin-host (first-party plugin lifecycle: install/remove/enable/disable). ide/projects/ and ide/examples/ hold sample projects, including the canonical hello-friends first-program walkthrough.

ql/ — engram-el

The El-native integration layer for a live Engram server — not a library (no importable modules, no build artifact), a set of standalone .el programs run directly via el run-file. Three components: Studio (studio/studio.el, a full terminal graph explorer), a Hebbian field-model proof of concept, and El builtin / LLM-builtin smoke test suites. This is the reference for correct patterns when an El program uses Engram as its substrate. Spec: ql/spec/elql.md.

ui/ — el-ui

A frontend framework where component state is an Engram graph and reactivity is spreading activation — not virtual-DOM diffing (React), Proxy-based dependency tracking (Vue), or compile-time analysis (Svelte). Re-renders are activated and propagated the same way associative memory retrieval works in engram/.

~15 vessels covering the full frontend surface: el-platform (env/fs/network/clock abstraction), el-config, el-html (SSR emit primitives), el-layout, el-style (design tokens/themes), el-i18n, el-auth / el-identity (JWT, sessions, OAuth PKCE — Engram-native), el-services (REST/gRPC/WebSocket bindings), el-aop (@authenticate/@authorize/@cache/@rate_limit decorators), el-secrets, el-graph (graph rendering/editor), el-publish (App Store / Play Store automation), and el-ui-compiler (El→JS component compiler; currently a stub pending a JS backend in elc). Spec: ui/spec/framework.md.

arbor/ — diagram language

A .arbor diagram language and toolchain: arbor-core (NodeId/shape/edge-kind types), arbor-parse (recursive-descent parser), arbor-diagram (IR + Mermaid serializer + architecture-diagram builders), arbor-layout (hierarchical layout — rank assignment, positioning, group bounds), arbor-render (SVG renderer), arbor-cli. (The architecture map above is the kind of diagram this is for.)


Getting started

Install the El SDK from the latest release:

bash lang/install.sh
# EL_VERSION=v1.0.0   bash lang/install.sh   # pin a specific release tag
# EL_PREFIX=/opt/el   bash lang/install.sh   # custom install prefix

Or build the compiler from source and verify the self-hosting chain:

cd lang
./dist/platform/elc elc-cli.el > elc-new.c
cc -std=c11 -I el-compiler/runtime -lcurl -lpthread \
   -o dist/platform/elc-new \
   elc-new.c el-compiler/runtime/el_seed.c

# Confirm the new binary reproduces itself exactly
./dist/platform/elc-new elc-cli.el > elc-verify.c
diff elc-new.c elc-verify.c   # should be identical

mv dist/platform/elc-new dist/platform/elc

Run your first program:

./lang/dist/platform/elc lang/examples/hello.el > hello.c
cc -std=c11 -I lang/el-compiler/runtime -lcurl -lpthread \
   -o hello hello.c lang/el-compiler/runtime/el_seed.c
./hello

More examples in lang/examples/, including a full starter project at lang/examples/hello-project/.

If the compiler binary is ever lost or corrupted, lang/BOOTSTRAP.md is the authoritative recovery path.


Development workflow

Branching follows dev → stage → main: work lands on dev, promotes to stage for integration testing, and is promoted to main for release (visible directly in the git history of this repo). CI is defined per-subproject under .gitea/workflows/lang/epm/ide share the root pipeline; engram and ql carry their own (ci-dev, ci-stage, and a release workflow each).

  • Language/runtime specs live at */spec/*.md (lang/spec/, ql/spec/, ui/spec/) and are the single source of truth for implemented-vs-planned status — code and docs are expected to agree with the spec's status markers, not the other way around.
  • Agent-facing orientation guides live at */AGENTS.md (currently lang/AGENTS.md); more subprojects may grow their own as they need agent-specific conventions documented.
  • Tagged releases live under lang/releases/, each with its own RELEASE.md.

Status

This is an actively developed, internal monorepo — not yet published under an open license. Treat everything here as proprietary to Neuron Technologies unless told otherwise.

S
Description
The Engram programming language — types as knowledge nodes, quantum-sealed prod target
Readme
199 MiB
2026-07-22 22:02:53 +00:00
Languages
Emacs Lisp 95.1%
C 3.9%
HTML 0.3%
Python 0.2%
Shell 0.2%
Other 0.1%