bigmerge 90d3f0bc76 engram: port PR #105's 3 genuine wins onto dev's existing cosq/e_eff semantic layer
Reconciles PR #105 ("fix: engram search latency — pin embed model, cache
query embeddings, bound activate BFS") with dev's ACTUAL current
engram_activate, rather than the ancient pre-restructure snapshot #105 was
built against.

WHY THIS NEEDED RECONCILIATION, NOT A DIRECT PORT: #105's single commit
(1dc49b1) modifies `lang/el-compiler/runtime/el_runtime.c` — a path that does
not exist on dev (dev has `lang/runtime/el_runtime.c`; the restructure that
renamed it happened after #105's branch point, which traces to a July 22
merge-base, weeks before the M8/M8.1/qgate/fan-effect/adjacency-index work
this file has grown since). #105's own engram_activate is consequently the
PRE-restructure version: no adjacency index (O(E) full edge scan per hop),
no query-aware qgate, no ACT-R fan effect, no eg_edge_eff_weight, and no
awareness of dev's cosq/e_eff embedding-blend semantic layer — it built a
parallel `g_qcache`/`engram_embed_raw` mechanism from scratch against code
that no longer exists at that path. A raw merge/cherry-pick was not possible
and would have been wrong even if it were: taking #105's tree wholesale would
have thrown away everything dev grew in the meantime (qgate, fan effect,
adjacency index, and this session's own M8 HNSW vindex integration).

RECONCILIATION: kept dev's cosq/e_eff mechanism as the semantic layer
entirely intact (unchanged by this commit) and ported #105's three genuinely
additive wins on TOP of it, at their equivalent sites in the CURRENT
eg_embed_fetch/engram_activate:

  1. keep_alive:-1 on the Ollama embed request body (eg_embed_fetch) — pins
     the embed model resident so a larger generation model loading under
     unified-memory pressure can't evict it and force a cold reload on the
     next search (#105 measured ~2.2s cold vs ~0.02-0.05s warm).
  2. Query-embedding cache upgraded from dev's single-slot (`_eg_qcache_text`,
     only ever remembered the LAST query) to a direct-mapped, FNV-1a-keyed,
     1024-slot cache (reusing the existing engram_id_hash) — so the
     curiosity loop's rotating phrases actually hit the cache instead of
     evicting each other every call. Same "pointer owned by the cache, not
     freed by caller" contract as before, just per-slot instead of global.
  3. Beam cap on the layer-1 spreading-activation BFS (new
     engram_activate_beam(), tunable via ENGRAM_ACTIVATE_BEAM, default 128).
     The FIFO frontier is processed in hop-level batches (entries sharing
     .hops are provably contiguous — see the code comment); when a level
     exceeds the beam width, only the top-`beam` by activation actually
     EXPAND. Every node in an oversized level still gets reached[]/best_bg[]
     recorded (that happens at enqueue time, one level up) and appears in
     the reported/promoted set — the cap bounds associative SPREAD width
     only, never recall of what was already found. Kept as a genuine
     additional bound even though the adjacency index + qgate + fan effect
     already mitigate #105's original "hub-node explosion" failure mode for
     a different reason: those prune WHICH targets matter; this bounds
     worst-case width regardless.

Everything else in dev's engram_activate — cosq/e_eff, the qgate rescale,
the fan effect, eg_edge_eff_weight, the M8 HNSW vindex seed discovery from
the #109 reconciliation earlier this session — is untouched.

VERIFIED (nsbx sandbox only, live :8742/:7770 never touched): cc -std=c11
-O2 clean build; booted in an isolated sandbox against a real cloned
production snapshot (13,424 nodes / 37,656 edges); ran 5 activate() calls
across rotating queries at depth 3, including the same query issued twice
non-consecutively (2nd hit landed at 476ms vs the 1st at 483ms — consistent
with a cache hit once Ollama's own warm-model latency is accounted for; no
crash, correct varied result counts (367-2610 nodes) each call; act-stats
JSON read correctly throughout.

Built on top of the M8/#109 reconciliation (bacaf3d, merged to dev as
#109) — dev's current HEAD at the time of this commit.
2026-08-15 17:50:04 -05:00

El

A self-hosting, statically-typed language that compiles to C — built around a graph-native runtime instead of a database driver.

El is the execution substrate for the Neuron agent runtime, the DHARMA network, and the Engram knowledge graph. This repository is the monorepo for the whole stack: the language itself, the graph memory engine it's built to talk to natively, and the tools (package manager, IDE, UI framework, diagramming) built on top of it.


Why El exists

Every other language treats persistent, associative state as something you reach for through a driver — a SQL client, an ORM, a Redis library bolted on from outside. El inverts that: graph operations (engram_*) are runtime primitives, on the same footing as string or list operations. There is no separate database driver because the database is not separate.

El has four defining properties:

  1. Self-hosting compiler. The compiler (lexer.el, parser.el, codegen.el, compiler.el) is written in El. It compiles El source to C, which cc compiles against a fixed runtime into a native binary. A Rust genesis compiler bootstrapped the first iteration; the self-hosted binary at lang/dist/platform/elc has been the canonical compiler ever since — every binary in dist/platform/ was produced by an earlier version of itself compiling el-compiler/src/. The chain is auditable: source is the ground truth, not the binary. See lang/BOOTSTRAP.md for the full recovery path if that binary is ever lost.
  2. C compilation target. Every compiled program is plain C11. Every El value is el_val_t (int64_t); strings are heap pointers cast through it. Functions become C functions; top-level statements become main().
  3. Graph-native runtime. The runtime provides first-class graph operations over an in-process Engram store — no separate DB driver, no ORM.
  4. DHARMA-aware identity. A cgi block declares a program's DHARMA identity at compile time. The runtime resolves identity before user code runs, so dharma_* calls have a stable principal and channel surface throughout.

Architecture map

                     ┌─────────────┐
                     │    lang     │   El compiler + C runtime
                     │ (El itself) │   everything below is written in it,
                     └──────┬──────┘   or compiles down through it
                            │
              ┌─────────────┼─────────────┐
              │             │             │
       ┌──────▼─────┐ ┌─────▼─────┐ ┌─────▼─────┐
       │   engram   │ │    epm    │ │    ide    │
       │ graph/mem  │ │  package  │ │  editor + │
       │  substrate │ │  manager  │ │    LSP    │
       └──────┬─────┘ └───────────┘ └───────────┘
              │
      ┌───────┼────────────────┬─────────────────────┐
      │       │                │                     │
┌─────▼───┐ ┌─▼──────────┐  ┌──▼──────────┐    ┌─────▼──────┐
│   elp   │ │ ql         │  │  ui         │    │   arbor    │
│  NLG /  │ │engram-el.  │  |spreading-   │    |arbor       │
│ 31 langs│ │studio+tests│  |activation UI│    |diagram lang│
└─────────┘ └────────────┘  └─────────────┘    └────────────┘

lang is the foundation — the compiler and C runtime everything else builds on. engram is the graph-native memory/state engine that gives El its identity (property 3 above). Everything else is either a tool for working with El (epm, ide) or a system built on top of Engram's graph model (elp, ql, ui, arbor).


Repository layout

lang/ — the El language

The compiler and runtime. Self-hosting: elc-cli.elcompiler.ellexer.el / parser.el / codegen.el / codegen-js.el, textually inlined and compiled in one pass. Compiles to C11 and links against el-compiler/runtime/el_seed.c, a hand-maintained OS-boundary layer (libcurl HTTP, pthreads, filesystem, arena allocation) — everything else in the runtime is native El (runtime/*.el).

Two layers to know: El programs (.el files — where nearly all work belongs) and the C seed (el_seed.c — edit only for genuine OS-level access; never re-implement what El can already express).

Current status (single source of truth: lang/spec/language.md): lexer/parser/codegen and the C runtime's core (I/O, strings, math, lists, maps, filesystem, args) are implemented. In flight: % operator, match-statement codegen, ? nil-propagation, cgi block parsing + DHARMA identity resolution, VBD role enforcement (@manager/@engine/@accessor), the real engram_* and dharma_* runtimes (currently stubs), and libcurl-backed http_get/http_post/http_serve. Bitwise operators, ??, and as casts are explicitly not in this language.

Key docs: AGENTS.md (agent-facing orientation), BOOTSTRAP.md (compiler recovery from scratch), spec/language.md, spec/codegen-js.md.

engram/ — graph intelligence substrate

A local-first memory substrate for accumulating intelligence, and the reason El's runtime doesn't need a database driver. Rust core (engram-core, engram-ffi) exposed to El and other languages (Kotlin, TypeScript/WASM, Go bindings).

The model: retrieval is spreading activation, not query. You name seed nodes and a query embedding; activation propagates outward through weighted edges, attenuating multiplicatively per hop (strength = parent_strength × edge_weight × target_salience × cosine_sim), gets pruned below a threshold, and the top-N nodes by activation strength come back. Storage and retrieval are the same structure — the way long-term potentiation works in biological memory, not the way a relational or vector database works.

Nodes live in four tiers (Working / Episodic / Semantic / Procedural, mirroring prefrontal / hippocampal / neocortical / cerebellar memory) and migrate between them based on salience decayimportance × recency-decay × log(activation_count). Forgetting is adaptive pruning, not a bug: unreinforced memories stop competing for attention without being deleted.

Backed by sled (embedded, local-first, no daemon) with flat cosine scan for vector search — deliberately simple until scale demands an HNSW layer. Full API and design rationale in engram/README.md.

elp/ — Engram Language Protocol

Bidirectional engine mapping between Engram semantic forms and natural-language surface text, across 31 languages — from Spanish and Japanese through historical/liturgical languages (Old Norse, Sanskrit, Sumerian, Coptic, Akkadian, Ge'ez). Compilation order runs language-profile + vocabulary → per-language morphology-*grammarrealizersemanticselp. This is what lets an Engram graph node round-trip to and from readable text in any of those languages.

epm/ — El Package Manager

Manages vessels (El's package unit): publish, install, resolve dependencies. Vessels are stored in Engram as graph nodes, not files in a registry index — epm reads the local manifest.el, talks to Engram over HTTP, and writes resolved vessels to .epm/vessels/. Source: registry.el, install.el, update.el, manifest.el.

ide/ — El IDE

Three vessels: el-ide-server (HTTP backend — file ops, build/run, LSP bridge, plugin host, settings), el-lsp (the language server — completion, hover, diagnostics, outline, format, type graph), and el-plugin-host (first-party plugin lifecycle: install/remove/enable/disable). ide/projects/ and ide/examples/ hold sample projects, including the canonical hello-friends first-program walkthrough.

ql/ — engram-el

The El-native integration layer for a live Engram server — not a library (no importable modules, no build artifact), a set of standalone .el programs run directly via el run-file. Three components: Studio (studio/studio.el, a full terminal graph explorer), a Hebbian field-model proof of concept, and El builtin / LLM-builtin smoke test suites. This is the reference for correct patterns when an El program uses Engram as its substrate. Spec: ql/spec/elql.md.

ui/ — el-ui

A frontend framework where component state is an Engram graph and reactivity is spreading activation — not virtual-DOM diffing (React), Proxy-based dependency tracking (Vue), or compile-time analysis (Svelte). Re-renders are activated and propagated the same way associative memory retrieval works in engram/.

~15 vessels covering the full frontend surface: el-platform (env/fs/network/clock abstraction), el-config, el-html (SSR emit primitives), el-layout, el-style (design tokens/themes), el-i18n, el-auth / el-identity (JWT, sessions, OAuth PKCE — Engram-native), el-services (REST/gRPC/WebSocket bindings), el-aop (@authenticate/@authorize/@cache/@rate_limit decorators), el-secrets, el-graph (graph rendering/editor), el-publish (App Store / Play Store automation), and el-ui-compiler (El→JS component compiler; currently a stub pending a JS backend in elc). Spec: ui/spec/framework.md.

arbor/ — diagram language

A .arbor diagram language and toolchain: arbor-core (NodeId/shape/edge-kind types), arbor-parse (recursive-descent parser), arbor-diagram (IR + Mermaid serializer + architecture-diagram builders), arbor-layout (hierarchical layout — rank assignment, positioning, group bounds), arbor-render (SVG renderer), arbor-cli. (The architecture map above is the kind of diagram this is for.)


Getting started

Install the El SDK from the latest release:

bash lang/install.sh
# EL_VERSION=v1.0.0   bash lang/install.sh   # pin a specific release tag
# EL_PREFIX=/opt/el   bash lang/install.sh   # custom install prefix

Or build the compiler from source and verify the self-hosting chain:

cd lang
./dist/platform/elc elc-cli.el > elc-new.c
cc -std=c11 -I el-compiler/runtime -lcurl -lpthread \
   -o dist/platform/elc-new \
   elc-new.c el-compiler/runtime/el_seed.c

# Confirm the new binary reproduces itself exactly
./dist/platform/elc-new elc-cli.el > elc-verify.c
diff elc-new.c elc-verify.c   # should be identical

mv dist/platform/elc-new dist/platform/elc

Run your first program:

./lang/dist/platform/elc lang/examples/hello.el > hello.c
cc -std=c11 -I lang/el-compiler/runtime -lcurl -lpthread \
   -o hello hello.c lang/el-compiler/runtime/el_seed.c
./hello

More examples in lang/examples/, including a full starter project at lang/examples/hello-project/.

If the compiler binary is ever lost or corrupted, lang/BOOTSTRAP.md is the authoritative recovery path.


Development workflow

Branching follows dev → stage → main: work lands on dev, promotes to stage for integration testing, and is promoted to main for release (visible directly in the git history of this repo). CI is defined per-subproject under .gitea/workflows/lang/epm/ide share the root pipeline; engram and ql carry their own (ci-dev, ci-stage, and a release workflow each).

  • Language/runtime specs live at */spec/*.md (lang/spec/, ql/spec/, ui/spec/) and are the single source of truth for implemented-vs-planned status — code and docs are expected to agree with the spec's status markers, not the other way around.
  • Agent-facing orientation guides live at */AGENTS.md (currently lang/AGENTS.md); more subprojects may grow their own as they need agent-specific conventions documented.
  • Tagged releases live under lang/releases/, each with its own RELEASE.md.

Status

This is an actively developed, internal monorepo — not yet published under an open license. Treat everything here as proprietary to Neuron Technologies unless told otherwise.

S
Description
The Engram programming language — types as knowledge nodes, quantum-sealed prod target
Readme
203 MiB
2026-07-22 22:02:53 +00:00
Languages
Emacs Lisp 95.1%
C 3.9%
HTML 0.3%
Python 0.2%
Shell 0.2%
Other 0.1%