Compare commits

..

34 Commits

Author SHA1 Message Date
bigmerge 827257d3a4 Remove __pycache__ .pyc files accidentally included in the projector commit
El SDK CI - dev / build-and-test (pull_request) Successful in 6m28s
2026-08-15 14:27:23 -05:00
bigmerge 4bbfdcceff Add native audio/image efferent surfaces + projector proof-of-shape
audio-surface.el / image-surface.el: own-core additive-synthesis WAV and
raster-PNG renderers (integer-only DSP, since EL has no floats), rendered
from learned engram signatures via a pluggable surface-profile
abstraction (surface-profile.el). audio-demo.el / image-demo.el are
drivers. NOTE: demo files hardcode absolute paths to this worktree's own
directory — will need a path fixup before landing.

elp/projector/ is a Python package the author's own README marks as
"STAGING/PROOF-OF-SHAPE — not the deliverable", superseded by the native
.el surface-profile work above; kept as a validated architecture proof.
Generated output (elp/faculty/{out,sig}, elp/projector/out,
__pycache__) intentionally excluded.
2026-08-15 14:26:59 -05:00
will.anderson d71fc4c1c0 Merge pull request 'promote stage -> main: reconciled el runtime (engram search + natives + durable truncation fix + Windows port)' (#82) from stage into main
El SDK Release / build-and-release (push) Successful in 8m31s
El SDK CI - dev / build-and-test (pull_request) Successful in 8m41s
2026-07-22 21:44:01 +00:00
will.anderson a118d19393 Merge pull request 'promote dev -> stage: el cluster (#66 engram + #79 truncation fix + release-runtime Windows port)' (#81) from dev into stage
El SDK CI - stage / build-and-test (push) Successful in 7m58s
El SDK Release / build-and-release (pull_request) Successful in 4m16s
2026-07-22 21:20:17 +00:00
will.anderson c6aa1e5c53 Merge pull request 'Land el cluster: #66 engram search + natives, #79 truncation fix, + release-runtime Windows port (reconciled)' (#80) from reconcile/el-cluster-windows-runtime into dev
El SDK CI - dev / build-and-test (push) Successful in 8m3s
El SDK CI - stage / build-and-test (pull_request) Successful in 4m25s
Land el cluster (#66 + #79 + release-runtime Windows-port reconciliation) into dev
2026-07-22 21:06:36 +00:00
will.anderson ff577391f2 reconcile(release-runtime): Windows-port + complete v1.0.0 release runtime so the desktop soul cross-compiles
El SDK CI - dev / build-and-test (pull_request) Successful in 7m7s
The desktop soul (neuron/dist) compiles against the v1.0.0-20260501 release
runtime. After #66 landed the engram natives (tokenized/ranked search,
engram_prune_telemetry) and #79 the durable truncation fix into this runtime,
two gaps remained before it could cross-compile the Windows brain:

1. Windows OS boundary: the release runtime had no Win32 path. Ported the same
   _WIN32-guarded shim the mainline runtime carries (#69): #ifdef _WIN32 ->
   el_platform_win.h (winsock/dlsym/popen + WSAStartup ctor), SOCKET fd guards
   and el_closesocket() at every socket site, CreateProcessA for exec_bg, the
   tm_zone/mingw guard, an el_setsockopt optval wrapper (GCC14), and curl-less
   libcurl stubs. Every change is _WIN32/HAVE_CURL-gated — the POSIX build is
   byte-identical (gcc -fsyntax-only clean; native behaviour unchanged).

2. Header exports: the release el_runtime.h omitted symbols the soul dist calls
   that are defined in this runtime's .c — the http_handler_fn/http_handler4_fn
   typedefs and el_arena_push/pop, engram_prune_telemetry, engram_get_node_by_label.
   Declaration-only, POSIX-neutral; fixes implicit-declaration/unknown-type
   errors under the C11 mingw build.

Result: x86_64-w64-mingw32-gcc compiles el_runtime.c + all 48 soul modules
clean; POSIX gcc -fsyntax-only clean. This is the Windows-port PR the runtime
needed on main (the release-runtime counterpart to #69), landed via stage.
2026-07-22 15:56:08 -05:00
will.anderson ee0d5f9b97 Merge #79: durable HTTP response-truncation fix, both runtimes (via stage) 2026-07-22 15:46:02 -05:00
will.anderson 391bd818ea Merge #66: tokenized+ranked engram lexical search + engram natives (via stage) 2026-07-22 15:45:54 -05:00
will.anderson 43636aed99 runtime: pair fs_read length hint with its buffer in BOTH runtimes — kill response truncation for good
El SDK Release / build-and-release (pull_request) Failing after 7s
The binary-safe fs_read length (_tl_fs_read_len) was consumed by the HTTP
response path for ANY body, even when a handler wrapped a smaller file into a
larger reply. Content-Length then lied AND the send stopped short: the
safety-contact (988) routes returned 178 of 208/218 bytes, cut mid-'set_at' —
unparseable JSON. The desktop app read that as failure. On Windows the shipped
brain is an OLD build without even the per-handler workaround, so EVERY reply
truncated: the app can't read confirmations and refuses the new user.

Durable fix: pair the length hint with the exact buffer pointer it describes
(_tl_fs_read_buf). Apply the raw byte count ONLY when the response IS that
buffer (binary file serving stays correct); every wrapped/enveloped/derived
body is measured with strlen. Reset both at request start and in fs_read /
json_get_raw. This also closes the stale-hint heap over-read (a length larger
than a later body would read past it out the socket) that a plain max() leaves
open — so this class of bug dies on every platform, not just where a handler
happened to be patched.

Applied identically to the mainline runtime (lang/el-compiler/runtime) AND the
frozen release runtime (lang/releases/v1.0.0-20260501) the desktop souls
compile against — the release copy still carried the raw leak, which is why the
Windows brain kept truncating. Same proven approach as PR #78 (Tim Lingo),
extended to cover the release runtime and rebased onto current main.

Both runtimes: gcc -fsyntax-only clean.
2026-07-22 15:04:25 -05:00
will.anderson 8f8ccc945e self-review 2026-07-22: persist canonical snapshot on write routes; newest-first tie-break in node listings
El SDK Release / build-and-release (pull_request) Failing after 14m24s
Durability: the 2026-07-21 fix stopped read routes writing the canonical
snapshot but left no save on ANY write path — every mutation lived in RAM
until a manual POST /api/save. Observed live: two restarts reverted the
store to a 17h-old snapshot, destroying same-day writes. persist_canonical()
now runs after node/edge create, knowledge capture, forget, strengthen, and
load-merge. ISE telemetry excluded deliberately (48h-pruned, loss-tolerant,
~2/min; snapshotting 28MB per heartbeat is waste).

Listing order: scan routes sort by salience with store-order ties, so
equal-salience telemetry (all ISEs are 0.3) returned OLDEST first — a
limited /api/nodes query silently returned a stale window, and a 41h-old
heartbeat series read as a live outage during this review. Ties now break
newest-first by created_at.
2026-07-22 08:51:33 -05:00
will.anderson 409ec99397 self-review 2026-07-22: ACT-R/Petrov base-level WM decay replaces per-call multiplicative carry-over
The old carry-over (weight *= 0.7 per engram_activate call) was call-rate-
dependent — carried context died in seconds under rapid curiosity scans and
lingered for hours under quiet loops — and a decayed scalar cannot represent
access frequency at all.

Now: k=10 access-timestamp ring + Petrov (2006) closed-form tail, d=0.5.
WM promotion and engram_strengthen record presentations; carry-over evicts
at base-level tau=-3.0 (Soar forgetting, ~403s single-touch) and shapes the
weight held at promotion (wm_anchor) with the ACT-R retrieval logistic
(s=0.4) — a pure function of wall-clock time, idempotent per call.
Persisted as access_ts/wm_anchor in snapshots; legacy nodes fall back to
the optimized form ln(n/(1-d)) - d*ln(L). base_level exposed in both node
serializers for observability.

Backing spec: 2026-07-21 integration brief (bl-b17facdd). Verified live:
carried weight ~anchor seconds after two disjoint activations (old code:
0.49x); frequency-hot nodes hold B=1.9 vs -0.14 single-touch.
2026-07-22 08:44:39 -05:00
will.anderson dc39a61e2c self-review 2026-07-21: stop read routes clobbering canonical snapshot; add /api/load-merge
Root cause of the 2026-05→07 identity-node loss: route_scan_edges and
route_sync serialized state by engram_save()ing over the canonical
snapshot.json on every GET, so one bad boot load meant the first read
request overwrote the good snapshot. Read routes now export to scratch
paths. Boot guard preserves evidence on non-empty-file/zero-node loads
and keeps a boot-time backup on good loads. New POST /api/load-merge
(explicit path required) used to restore 385 identity nodes + 1115
edges from the 2026-05-13 backup.
2026-07-21 08:50:38 -05:00
will.anderson eba9eac8a8 self-review 2026-07-19: port stranded fixes to the release runtime (production copy)
Three fixes that existed elsewhere but never reached the runtime the engram
binary actually builds against:

- tokenized + ranked query matching (search/search_json/activate seeds/
  goal_bias) ported from the el-compiler copy (e3dabe3, 2026-07-14) — the
  production engram kept whole-query Ctrl-F for 5 days after the fix
  'shipped'. Multi-word curiosity seeds went 0 -> 36 activated. Kept the
  ISE seed exclusion the el-compiler copy dropped.
- Knowledge -> 0.20 WM threshold after tier checks (dev-line 4bf7716):
  Semantic/Episodic Knowledge nodes fell to the 0.40 note default and only
  entered WM via breakthrough.
- goal_bias: Knowledge in is_knowledge + curiosity-seed technical terms
  (dev-line d53516b).

Also: seed_epoch was a running pairwise average, not the mean it claimed —
exponentially over-weighted later seeds in the temporal-proximity bonus.
Fixed to a true int64-sum mean. Stale INHIBITION_FACTOR comment corrected.

Root cause captured as knowledge: two runtime copies + branch-per-fix
without merge discipline stranded the entire dev semantic layer (cosine
activation, embeddings) out of production. Reconciliation planned as P1.
2026-07-19 08:46:47 -05:00
will.anderson ab6b52a0b4 self-review 2026-07-18: fix soul SIGABRT double-free + engram route scoping sweep
1. engram_neighbors_json (release runtime): BFS frontier/visited strings were
   el_strdup'd (arena-tracked) but manually freed, so el_request_end()
   double-freed every one — SIGABRT in http_worker under load (2 prod crashes
   today via /api/neuron/session/begin and /api/neuron/graph; reproduced and
   verified fixed with ASAN). Introduced when porting from the dev runtime,
   which correctly uses plain strdup. Third instance of the
   arena-vs-manual-free class (after EngramNode 07-15 and idmap keys 07-16).

2. server.el: let-in-if scoping sweep — defaults assigned inside if-blocks
   never mutated the outer binding, so /api/search and /api/activate always
   ran with q="", created nodes got node_type=""/salience=0.0, edges got
   relation=""/weight=0.0, and save/load with no path hit engram_save("").
   Rewritten to the let-if-else expression form. /api/activate now also
   rejects empty queries instead of wiping carried WM weights.

3. engram_activate: retrieval reinforcement (ACT-R base-level learning) —
   nodes promoted to WM that survive both capacity caps now get
   last_activated/activation_count updated, so frequently retrieved memories
   decay slower than abandoned ones. Scoped to promoted-only to avoid
   flattening dampening across BFS fan-out.
2026-07-18 08:48:04 -05:00
will.anderson 2baa0b9a41 Merge pull request 'release: promote stage -> main (ci publish hardening for sdk-release)' (#77) from stage into main
El SDK Release / build-and-release (push) Successful in 7m55s
2026-07-15 21:21:28 +00:00
will.anderson 6a8b2461cd Merge pull request 'release: promote dev -> stage (ci publish hardening for stage/main)' (#76) from dev into stage
El SDK CI - stage / build-and-test (push) Successful in 8m19s
El SDK Release / build-and-release (pull_request) Failing after 13m1s
2026-07-15 21:16:11 +00:00
will.anderson bcb356fe69 Merge pull request 'ci(stage,main): decouple ci-base rebuild, make SDK publish fail loudly' (#75) from hotfix/ci-stage-main-publish-hardening into dev
El SDK CI - stage / build-and-test (pull_request) Successful in 4m27s
El SDK CI - dev / build-and-test (push) Failing after 14m3s
2026-07-15 21:15:27 +00:00
will.anderson dd7827059a ci(stage,main): decouple ci-base rebuild, make SDK publish fail loudly
El SDK CI - dev / build-and-test (pull_request) Failing after 14m30s
Mirror the PR #72 fix (applied to ci-dev.yaml) onto ci-stage.yaml and
sdk-release.yaml. The stage and prod release jobs reported FAILURE even
when the el-runtime-c/-h publish SUCCEEDED, because the ancillary ci-base
Docker rebuild (a CI-cache optimization on the fragile host-mode GCE
runner) reddened the whole job.

- Rebuild ci-base step: continue-on-error: true — never blocks/reddens
  the job; the SDK publish is the deliverable.
- Publish step: set -euo pipefail + empty-key guard + active-account echo
  so a real publish failure still fails loud and is diagnosable.
2026-07-15 16:14:50 -05:00
will.anderson 208e36c899 Merge pull request 'release: promote stage -> main (tokenized search, get_node_by_label, epm fix, win portability)' (#74) from stage into main
El SDK Release / build-and-release (push) Successful in 8m23s
2026-07-15 18:24:39 +00:00
will.anderson b97ce74d1f Merge pull request 'release: promote dev -> stage (tokenized search, get_node_by_label, epm fix)' (#73) from dev into stage
El SDK CI - stage / build-and-test (push) Failing after 8m45s
El SDK Release / build-and-release (pull_request) Successful in 4m1s
2026-07-15 17:20:11 +00:00
will.anderson 155a449c4e Merge pull request 'ci(dev): make SDK publish fail loudly, decouple ci-base rebuild' (#72) from hotfix/ci-dev-publish-hardening into dev
El SDK CI - dev / build-and-test (push) Successful in 8m56s
El SDK CI - stage / build-and-test (pull_request) Successful in 4m10s
2026-07-15 16:34:14 +00:00
will.anderson 4696fd6833 ci(dev): make SDK publish fail loudly, decouple ci-base rebuild
El SDK CI - dev / build-and-test (pull_request) Successful in 8m51s
The dev push build went green-then-red while nothing published: the
Publish step had no set -e, so an auth/upload failure exited 0 (silent
no-publish), while the ci-base rebuild (set -euo pipefail + Docker on the
host-mode runner) hard-failed the job. Add set -euo pipefail + an empty-key
guard + active-account echo to the Publish step so failures surface with a
retrievable log, and mark the ci-base cache rebuild continue-on-error so
the fragile Docker step can never block the actual SDK artifact publish.
2026-07-15 11:33:37 -05:00
will.anderson 581a351fb1 Merge pull request 'integrate: stack PRs #65–#69 (elc OOM guard, tokenized+semantic engram search, get_node_by_label, win portability) for green CI' (#71) from hotfix/stage-elc-engram-integration into dev
El SDK CI - dev / build-and-test (push) Failing after 14m31s
2026-07-15 15:49:57 +00:00
will.anderson 8ce8656de2 epm: declare cross-module callees as extern fn so strict compilers accept generated C
El SDK CI - dev / build-and-test (pull_request) Successful in 7m33s
epm's sibling modules (registry/install/update) call functions defined in other
modules and in the El runtime (config, read_installed, registry_find,
manifest_deps, manifest_name, registry_latest_version, registry_token,
install_vessel, installed_version) without importing them, so elc emits no C
prototype for those calls. gcc<=13 treated the resulting implicit declarations
as warnings; gcc>=14 and clang reject them as hard errors, which is why the
"Build epm" CI step fails and blocks the whole dev/stage pipeline.

Add `extern fn` forward declarations -- El's own separate-compilation mechanism
-- for each cross-module callee at the top of registry/install/update. This
gives elc the correct C prototype in every generated translation unit, so the
calls compile cleanly and still resolve at link time. Simply suppressing
-Wimplicit-function-declaration would be unsafe: an implicit int return
truncates the 64-bit pointer returns of config/registry_find into a latent
crash, so declaring the true signatures is the correct fix. Localized to epm;
touches neither elc nor the runtime.
2026-07-15 10:14:43 -05:00
will.anderson 1e49560f1f Merge remote-tracking branch 'origin/feat/engram-semantic-search' into hotfix/stage-elc-engram-integration
El SDK CI - dev / build-and-test (pull_request) Failing after 14m39s
# Conflicts:
#	lang/el-compiler/runtime/el_runtime.c
2026-07-15 09:33:05 -05:00
will.anderson e8f0b5a9de Merge remote-tracking branch 'origin/fix/engram-lexical-tokenized-search' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 40287c4cfc Merge remote-tracking branch 'origin/hotfix/win-runtime-portability' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 0481bea44d Merge remote-tracking branch 'origin/hotfix/runtime-engram-get-node-by-label' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 9d565ca080 Merge remote-tracking branch 'origin/hotfix/elc-fixes' into hotfix/stage-elc-engram-integration 2026-07-15 09:28:44 -05:00
will.anderson 4773dd0aa2 runtime: make Windows soul reproducible from a clean el checkout
El SDK Release / build-and-release (pull_request) Failing after 16s
Two el_runtime portability defects only ever lived in staged local copies
used to hand-build neuron-ui PR #136's curl-enabled Windows neuron.exe.
gcc 15 promotes both to hard errors, so a clean el checkout cannot rebuild
that soul. Upstream the minimal fixes so the build is reproducible:

- http_serve_async: cast setsockopt optval to (const char*). Win32/mingw
  setsockopt wants const char*, not int*; the cast is a no-op on POSIX and
  matches the four already-cast sites elsewhere in this file.
- engram_save persist path: map fsync -> _commit in the _WIN32-only
  el_platform_win.h (io.h already included). Windows has no fsync(); the
  POSIX path is untouched.
2026-07-15 04:24:08 -05:00
will.anderson b4967af13e feat(engram): semantic search layer via nomic-embed-text (cosine ∪ lexical)
Lexical istr_contains alone can't surface a node whose words don't appear
in the query. This adds an optional dense-vector layer: node content and the
query are embedded through Ollama (nomic-embed-text), and nodes are ranked by
cosine similarity unioned with lexical hits, so a paraphrase query reaches the
right node.

Wired into all three query entry points in el_runtime.c:
  - engram_search_json (HTTP /api/search): collect lexical ∪ semantic
    candidates, score (lexical base 1.0 + cosine; pure-semantic = cosine),
    rank, emit top-N. Stable sort preserves old order when semantic is off.
  - engram_search (internal el_val twin): lexical ∪ semantic union.
  - engram_activate seed loop (HTTP /api/activate): a node seeds if it
    lexically matches OR clears the cosine threshold; pure-semantic seeds
    enter scaled by cosine so paraphrase spreads without overpowering.

Degradable by design: the whole layer is gated on HAVE_CURL plus a one-shot
runtime probe. If curl is compiled out, Ollama is unreachable, or
ENGRAM_SEMANTIC=0, every entry point yields zero semantic signal and callers
fall back byte-for-byte to the pre-existing lexical search.

Node embeddings are cached in process memory keyed by node id with an FNV-1a
content hash for invalidation; the query is embedded once per call — so the
graph is not re-embedded on every query. nomic task prefixes
(search_query:/search_document:) are applied for retrieval separation.

Build steps gain -DHAVE_CURL so the engram artifact compiles the layer in
(-lcurl was already linked). Env: ENGRAM_SEMANTIC, ENGRAM_EMBED_URL,
ENGRAM_EMBED_MODEL, ENGRAM_SEMANTIC_MIN (cosine threshold, default 0.6).
2026-07-14 18:48:16 -05:00
will.anderson e3dabe3e08 fix(engram): tokenized + ranked lexical search, not whole-query Ctrl-F
El SDK Release / build-and-release (pull_request) Failing after 14m46s
engram search/activate/goal-bias matched the ENTIRE raw query string as a
single case-insensitive substring (istr_contains(field, q)). Multi-word
queries like "windows msi signing" only matched a node containing that exact
contiguous run, so real multi-word queries returned ZERO on a graph saturated
with the answer. This is Ctrl-F, not search — and search is the core of the
engram being useful.

Fix: split the query on whitespace into distinct tokens; a node matches if it
contains ANY token in content/label/tags. Rank by distinct tokens matched
(desc) then salience (desc). istr_contains is kept unchanged as the per-token
primitive. Single-token queries are a strict special case (score 0 or 1) so
the many single-word callers do not regress.

Sites changed (all in el_runtime.c):
- new helpers engram_tokenize_query / engram_node_match_score / engram_rank_cmp
- engram_search           (internal el_val_t path)
- engram_search_json      (HTTP /api/search path)
- engram_activate seed loop (HTTP /api/activate path; seed activation scaled
  by token coverage so full-query matches seed more strongly)
- engram_goal_bias overlap bonus upgraded to graded token coverage

Proof (6591-node snapshot copy, rebuilt binary on :8799, POST JSON path):
  windows msi signing  0 -> 20   Will Anderson  0 -> 20
  windows msi          0 -> 20   tokenized search fix  0 -> 20
Single-word parity preserved (VBD/volatility/elc capped at limit; unkey = all
matching nodes). Top hits are relevant (e.g. "Will Anderson" surfaces the
Project Design and VBD whitepapers).

Note: GET ?q=a%20b still returns 0 because query_param (server.el) does not
URL-decode — a separate EL-layer bug; the soul's POST-JSON path is fixed here.
2026-07-14 18:39:07 -05:00
will.anderson 0a0a2bcb44 parser: bound token reads to Eof so malformed input errors instead of OOMing
El SDK Release / build-and-release (pull_request) Failing after 16s
Out-of-range tok_kind/tok_value reads returned runtime null (el_list_get OOB
-> 0) rather than the Eof sentinel, so the inner parse loops (parse_block,
call-arg, array-literal, match-arm) that terminate only on their close
delimiter or k=="Eof" never saw Eof once the cursor ran past the single
trailing Eof token. On unclosed-delimiter input the parser then appended AST
nodes forever -> unbounded allocation -> ~700GB -> OOM (observed compiling
neuron/sessions.el).

Fix at the choke point: tok_kind returns "Eof" and tok_value returns "" for
out-of-range positions, restoring the parser-wide contract that reads at/after
the end yield Eof. expect() no longer steps past the Eof sentinel on mismatch.
This terminates every overrun loop simultaneously; a malformed program now
surfaces as a normal (best-effort) parse end instead of exhausting memory.

Requires a self-hosted bootstrap rebuild of elc to take effect.
2026-07-14 14:21:39 -05:00
will.anderson 5c41c66a0f Merge pull request 'fix(windows): guard el_mem_check with _WIN32 — rusage is POSIX-only' (#60) from fix/windows-rusage-guard into stage
El SDK CI - stage / build-and-test (push) Failing after 13m21s
fix(windows): guard el_mem_check with _WIN32 — rusage is POSIX-only
2026-06-25 16:48:13 +00:00
38 changed files with 4517 additions and 314 deletions
+15
View File
@@ -214,9 +214,18 @@ jobs:
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
# Fail loudly: previously this step had no `set -e`, so an auth or
# upload failure was swallowed (step exited 0 on the trailing echo)
# and the SDK silently never published. Surface failures now.
set -euo pipefail
if [ -z "${GCP_SA_KEY:-}" ]; then
echo "FATAL: GCP_SA_KEY secret is empty — cannot authenticate to publish" >&2
exit 1
fi
echo "${GCP_SA_KEY}" > /tmp/gcp-key.json
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
echo "Publishing as active account: $(gcloud config get-value account 2>/dev/null)"
VERSION="${GITHUB_SHA:0:8}"
@@ -268,6 +277,12 @@ jobs:
# Patches ci-base:dev in-place: pulls the existing image (which has all
# system deps — Node, Go, gcloud, Docker CLI, etc.) and overlays the freshly
# built El SDK on top. Keeps the full ci-base rebuild fast and incremental.
#
# continue-on-error: this is a CI-cache optimization, NOT the release
# artifact. It runs Docker (pull/build/push ~600MB) on the host-mode GCE
# runner where DinD/Docker availability is fragile. A failure here must
# never block or redden the job — the SDK publish above is the deliverable.
continue-on-error: true
if: github.event_name == 'push'
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+15
View File
@@ -212,12 +212,21 @@ jobs:
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
# Fail loudly: previously this step had no `set -e`, so an auth or
# upload failure was swallowed (step exited 0 on the trailing echo)
# and the SDK silently never published. Surface failures now.
set -euo pipefail
if [ -z "${GCP_SA_KEY:-}" ]; then
echo "FATAL: GCP_SA_KEY secret is empty — cannot authenticate to publish" >&2
exit 1
fi
echo "${GCP_SA_KEY}" > /tmp/gcp-key.json
apt-get install -y -qq apt-transport-https ca-certificates curl
echo "deb [trusted=yes] https://packages.cloud.google.com/apt cloud-sdk main" > /etc/apt/sources.list.d/google-cloud-sdk.list
apt-get update -qq && apt-get install -y google-cloud-cli
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
echo "Publishing as active account: $(gcloud config get-value account 2>/dev/null)"
VERSION="${GITHUB_SHA:0:8}"
@@ -253,6 +262,12 @@ jobs:
# Patches ci-base:stage in-place: pulls the existing image (which has all
# system deps — Node, Go, gcloud, Docker CLI, etc.) and overlays the freshly
# built El SDK on top. Keeps the full ci-base rebuild fast and incremental.
#
# continue-on-error: this is a CI-cache optimization, NOT the release
# artifact. It runs Docker (pull/build/push ~600MB) on the host-mode GCE
# runner where DinD/Docker availability is fragile. A failure here must
# never block or redden the job — the SDK publish above is the deliverable.
continue-on-error: true
if: github.event_name == 'push'
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+15
View File
@@ -288,12 +288,21 @@ jobs:
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
run: |
# Fail loudly: previously this step had no `set -e`, so an auth or
# upload failure was swallowed (step exited 0 on the trailing echo)
# and the SDK silently never published. Surface failures now.
set -euo pipefail
if [ -z "${GCP_SA_KEY:-}" ]; then
echo "FATAL: GCP_SA_KEY secret is empty — cannot authenticate to publish" >&2
exit 1
fi
echo "${GCP_SA_KEY}" > /tmp/gcp-key.json
apt-get install -y -qq apt-transport-https ca-certificates curl
echo "deb [trusted=yes] https://packages.cloud.google.com/apt cloud-sdk main" > /etc/apt/sources.list.d/google-cloud-sdk.list
apt-get update -qq && apt-get install -y google-cloud-cli
gcloud auth activate-service-account --key-file=/tmp/gcp-key.json
gcloud config set project neuron-785695
echo "Publishing as active account: $(gcloud config get-value account 2>/dev/null)"
VERSION="${GITHUB_SHA:0:8}"
@@ -345,6 +354,12 @@ jobs:
# Patches ci-base:latest in-place: pulls the existing image (which has all
# system deps — Node, Go, gcloud, Docker CLI, etc.) and overlays the freshly
# built El SDK on top. Keeps the full ci-base rebuild fast and incremental.
#
# continue-on-error: this is a CI-cache optimization, NOT the release
# artifact. It runs Docker (pull/build/push ~600MB) on the host-mode GCE
# runner where DinD/Docker availability is fragile. A failure here must
# never block or redden the job — the SDK publish above is the deliverable.
continue-on-error: true
if: github.event_name == 'push'
env:
GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+91
View File
@@ -0,0 +1,91 @@
> **STATUS: STAGING / PROOF-OF-SHAPE — not the deliverable.** This Python package
> proved the architecture end-to-end against the proven realizer faculty (faithful
> md/docx/midi from real geometry: 0 ungrounded claims, SACRED polarity). Per Will's
> steer, the DELIVERABLE is NATIVE: the seam lives on the existing EL realizer as
> **surface-as-profile** — see `../src/surface-profile.el` and
> `../tests/examples/surface-profile-demo.el` (compiles + runs through elc → C →
> binary). The concepts below (one geometry-carrying frame; surface = a pluggable
> profile; plan/realize; deterministic-from-meaning) are exactly what the native
> module implements. Keep this package as the validated proof; build native.
# Efferent Multimodal Projector
**geometry → any surface, faithfully.** Neuron's own document-generation faculty:
the efferent twin of the ingest organ. Ingest is afferent (world → geometry);
this is efferent (geometry → an arbitrary-format document / any modality).
Built against the **proven** realizer faculty (neuron-talk sidecar `:8756`,
artifact `art-7affa557`). The live soul (`:8742` / `:7770`) is contacted **only**
through the read-only, GET-only `engram_client` — never mutated.
## The pipeline (surface-agnostic)
```
geometry region + surface/format spec
→ PLAN (manifold → document skeleton/DAG; the geometry IS the outline) plan.py
→ REALIZE (proven realizer, scaled sentence → passage, each section faithful) realize.py
→ COHERE (document-level flow / transitions, not stitched sentences) cohere.py
→ EMIT (pluggable SurfaceProjector → the target surface) projectors/
```
**The surface is a PARAMETER.** `pipeline.build_ir(...)` builds ONE
surface-neutral `DocumentIR` (`document_ir.py`); `pipeline.emit(doc, surface)`
projects it to whichever surface you name. Markdown, docx, and MIDI are the same
IR emitted three ways.
## The pivot: a geometry-carrying IR
`DocumentIR` is **not** a text tree. Every `Block` carries BOTH:
- `.sentences` — realized faithful text (what **text** projectors read),
- `.provenance` — the source geometry: `subj_id / relation / obj / polarity /
confidence / importance / salience / node_id` (what **music / image / video**
projectors read).
That single decision is what makes the projector multimodal: text renders the
words; music/image decode the geometry. A claim with no provenance cannot exist
in the IR — faithfulness is structural.
## The one shared seam
`projectors/base.py` — `SurfaceProjector.project(frame: DocumentIR) -> bytes`
(+ `surface / media_type / ext / modality / profile`). Register with
`register()`. Adding a surface changes nothing upstream.
`TwoStageProjector` blesses the peer plan/realize decomposition:
`spec = plan(frame)`, `bytes = realize(spec)`, `project = realize∘plan`; the
`profile` is the pluggable per-surface knob (text lang-profile, music
instr/mode-profile). `projectors/midi.py` is the reference two-stage impl.
## Surfaces
| surface | modality | status | emitter |
|---|---|---|---|
| `markdown` | text | landed | own (str) |
| `docx` | text | landed | own minimal OOXML (stdlib `zipfile`+XML, no lib) |
| `midi` | audio | landed (symbolic-music proof) | own minimal SMF (stdlib `struct`, no lib) |
| `audio` (WAV) | audio | peer agent (additive synth) | conforms to `TwoStageProjector` |
| `image` | image | documented seam | `projectors/seams.py` |
| `video` | video | documented seam (image×sound×time) | `projectors/seams.py` |
Music maps: relation → scale degree (same relation → same pitch), **polarity →
major/minor third (SACRED negation is audible)**, confidence → duration,
importance → velocity, section → register. Deterministic projection from meaning
— nothing invented.
## Faithfulness
`provenance.py` audits the IR: **zero** ungrounded claims, SACRED polarity
preserved (negations reported, never dropped), COHERE introduces no new geometry
(connectives are marked). `trace_table()` emits the geometry → section → claim
table.
## Run
```bash
PY=~/Desktop/lang-realizers/venv/bin/python
PYTHONPATH=~/Desktop/neuron-talk:~/Desktop/lang-realizers $PY generate.py
# writes ./out/{neuron-self,engram-temporal}.{md,docx,mid} + *.audit.json + *.provenance.md
```
Requires the proven realizer env (spaCy + the neuron-talk/lang-realizers engine)
and the read-only engram at `:8742`.
+79
View File
@@ -0,0 +1,79 @@
"""cohere.py — COHERE stage: document-level flow, not stitched sentences.
Fidelity is REALIZE's job; FLOW is this stage's. The hard part beyond sentence
fidelity is that a document must read as one thing. We add connective tissue at
the passage level:
* an opening abstract that names what the document covers (built ONLY from the
section headings that already exist — it introduces no new claim),
* a short transition lead into each section after the first, drawn from a
fixed set of discourse connectives ("Beyond that,", "Relatedly,", ...) that
carry no propositional content,
* ordering so the highest-grounded section leads.
CRITICAL: every connective is marked ``kind="connective"`` in its provenance, so
the faithfulness audit can prove COHERE introduced ZERO new geometry claims. A
transition is discourse glue, never a fact.
"""
from __future__ import annotations
from document_ir import Block, DocumentIR, Provenance
# discourse connectives — pure flow, no propositional content
_TRANSITIONS = [
"Beyond that,", "Relatedly,", "In the same region,", "From there,",
"Alongside this,", "Further,", "Turning to the next facet,",
]
def _connective_prov() -> Provenance:
return Provenance(subj_id=None, subject=None, relation="", obj=None,
polarity="aff", confidence=1.0, node_id=None,
kind="connective")
def _abstract_block(doc: DocumentIR) -> Block:
"""A grounded opening: names the sections, asserts nothing new."""
headings = [s.heading for s in doc.sections]
if not headings:
return Block(role="lead")
if len(headings) == 1:
body = f"This document, generated from Neuron's geometry, covers {headings[0]}."
else:
listed = ", ".join(headings[:-1]) + f", and {headings[-1]}"
body = ("This document is projected directly from Neuron's meaning-geometry. "
f"It traces {listed}.")
b = Block(role="lead")
b.sentences.append(body)
b.provenance.append(_connective_prov())
return b
def cohere_document(doc: DocumentIR, *, add_abstract: bool = True,
add_transitions: bool = True) -> DocumentIR:
"""Order sections by grounding, add abstract + transitions (flow only)."""
# order: strongest-grounded section (mean confidence x #claims) first,
# but keep an explicitly-first section if the plan pinned one via level 1.
def _score(sec):
provs = [p for p in sec.all_provenance() if p.kind == "fact"]
if not provs:
return 0.0
mean_conf = sum(p.confidence for p in provs) / len(provs)
return mean_conf * len(provs)
doc.sections.sort(key=_score, reverse=True)
if add_transitions:
for i, sec in enumerate(doc.sections):
if i == 0 or not sec.blocks:
continue
lead = _TRANSITIONS[(i - 1) % len(_TRANSITIONS)]
first = sec.blocks[0]
if first.sentences:
# prepend the connective to the first sentence (flow, no new claim)
first.sentences[0] = f"{lead} {first.sentences[0][0].lower()}{first.sentences[0][1:]}"
if add_abstract:
doc.meta["abstract"] = _abstract_block(doc)
return doc
+111
View File
@@ -0,0 +1,111 @@
"""document_ir.py — the surface-neutral, GEOMETRY-CARRYING document intermediate.
This is the pivot of the whole efferent projector. A DocumentIR is NOT a text
tree. It is a projection of a meaning-geometry region that carries, at every
leaf, BOTH:
* the realized surface text (``Block.sentences``) — what a TEXT projector reads,
* the source geometry (``Block.provenance``) — what a MUSIC / IMAGE /
VIDEO projector reads.
Because the IR holds the geometry, not just the words, the SAME
plan -> realize -> cohere pipeline drives every surface. A markdown projector
renders the sentences; a music projector reads the provenance edges (salience,
importance, polarity, relation) and maps them onto a symbolic-music surface;
an image/video projector (documented seam) would read the same geometry.
Nothing in this module invents content. Every :class:`Provenance` points at a
real engram node id and a real relation. That is the faithfulness contract made
structural: a claim with no provenance cannot exist in the IR.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from typing import Any
# --------------------------------------------------------------------------- #
# Provenance — the geometry an emitted claim traces to. FAITHFULNESS is here.
# --------------------------------------------------------------------------- #
@dataclass
class Provenance:
"""One geometry edge behind one realized claim.
``kind`` distinguishes a FACT (a structural edge asserted by the geometry,
spoken as fact) from an INTERPRETATION (something attributed, spoken with
attribution) — the facts-as-facts + interpretations-attributed discipline
(memory 80927e26). ``polarity`` is SACRED: a negated edge stays negated.
"""
subj_id: str | None # source engram node id of the subject
subject: str | None # normalized subject surface
relation: str # predicate lemma (e.g. "use", "contain", "be")
obj: str | None # normalized object / complement surface
polarity: str = "aff" # "aff" | "neg" (SACRED — never silently flipped)
confidence: float = 0.0 # extraction confidence in [0,1]
node_id: str | None = None # engram node the claim was extracted from
kind: str = "fact" # "fact" | "interpretation"
importance: float = 0.0 # source node importance (drives music/emphasis)
salience: float = 0.0 # source node salience
def trace(self) -> str:
arrow = "-->" if self.polarity == "aff" else "--NOT-->"
return (f"[{(self.node_id or '?')[:8]}] {self.subject!r} {arrow}"
f"{self.relation} {self.obj!r} (conf {self.confidence:.2f})")
@dataclass
class Block:
"""A passage: one or more faithful sentences + the geometry they trace to.
``sentences`` and ``provenance`` are index-aligned where possible: sentence
``i`` was realized from ``provenance[i]``. A COHERE transition sentence with
no new geometry carries a provenance whose ``kind == "connective"`` so the
audit can see it introduced no new claim.
"""
sentences: list[str] = field(default_factory=list)
provenance: list[Provenance] = field(default_factory=list)
role: str = "body" # "body" | "lead" | "transition"
def text(self) -> str:
return " ".join(s.rstrip(". ") + "." for s in self.sentences if s.strip())
@dataclass
class Section:
heading: str
level: int = 2 # markdown heading level / outline depth
blocks: list[Block] = field(default_factory=list)
seed_ids: list[str] = field(default_factory=list) # geometry nodes of section
summary: str = "" # one-line grounded gloss (for pptx bullets / TOC)
def all_provenance(self) -> list[Provenance]:
out: list[Provenance] = []
for b in self.blocks:
out.extend(b.provenance)
return out
@dataclass
class DocumentIR:
"""The surface-neutral document. Built ONCE, projected to ANY surface."""
title: str
subtitle: str = ""
sections: list[Section] = field(default_factory=list)
seed_id: str | None = None # the geometry region root
format_spec: dict[str, Any] = field(default_factory=dict) # requested shape
meta: dict[str, Any] = field(default_factory=dict)
# -- geometry facets (what non-text projectors consume) ----------------- #
def all_provenance(self) -> list[Provenance]:
out: list[Provenance] = []
for s in self.sections:
out.extend(s.all_provenance())
return out
def claim_count(self) -> int:
return sum(1 for p in self.all_provenance() if p.kind in ("fact", "interpretation"))
def ungrounded_count(self) -> int:
"""Claims with no traceable node — MUST be zero for a faithful doc."""
return sum(1 for p in self.all_provenance()
if p.kind in ("fact", "interpretation") and not p.node_id)
+81
View File
@@ -0,0 +1,81 @@
"""generate.py — drive the projector: one geometry region -> many surfaces.
Proves the thesis with REAL output: builds ONE surface-neutral DocumentIR from
Neuron's OWN self-geometry (read-only against the live soul via the proven
faculty), then EMITS it to Markdown, docx, and MIDI — the same plan/realize/
cohere, three surfaces. Writes the files + the faithfulness audit to ./out/.
"""
from __future__ import annotations
import json
import os
import sys
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
import pipeline # noqa: E402
import provenance # noqa: E402
from geometry import load_self_region # noqa: E402
OUT = os.path.join(_HERE, "out")
def _emit_all(doc, stem):
"""Emit one IR to every text/audio surface + audit + provenance."""
for surface in ("markdown", "docx", "midi"):
data = pipeline.emit(doc, surface)
proj = pipeline.get_projector(surface)
path = os.path.join(OUT, f"{stem}.{proj.ext}")
with open(path, "wb") as f:
f.write(data)
print(f" emitted {surface:9s} -> {os.path.basename(path)} ({len(data)} bytes)")
a = provenance.audit(doc)
with open(os.path.join(OUT, f"{stem}.audit.json"), "w") as f:
json.dump(a, f, indent=2)
with open(os.path.join(OUT, f"{stem}.provenance.md"), "w") as f:
f.write(provenance.trace_table(doc))
print(" audit:", {k: a[k] for k in ("claims", "ungrounded_claims",
"negations_preserved", "distinct_source_nodes", "faithful")})
return a
def main():
os.makedirs(OUT, exist_ok=True)
print("surfaces registered:", pipeline.available_surfaces())
# ---- Document 1: Neuron's self-description (marquee) ------------------- #
print("\n[1] Neuron self-description")
region = load_self_region(max_nodes=9)
print(" self region:", region)
doc1 = pipeline.build_ir(
None, region=region,
title="Neuron: A Self-Description from Its Own Geometry",
subtitle="Projected efferently from the engram — every claim traces a node.",
format_spec={"genre": "self-description", "register": "expository"},
max_sections=5, conf_floor=0.6)
print(f" IR: {len(doc1.sections)} sections, {doc1.claim_count()} claims, "
f"ungrounded={doc1.ungrounded_count()}")
_emit_all(doc1, "neuron-self")
# ---- Document 2: a coherent, clean whitepaper-style section ------------ #
print("\n[2] Whitepaper-style section (coherent clean region)")
doc2, _ = pipeline.project(
["chronoception", "time", "awareness", "engram", "temporal"],
surface="markdown",
title="Temporal Awareness in the Engram",
subtitle="A section projected from the geometry of chronoception.",
format_spec={"genre": "whitepaper-section", "register": "technical"},
max_sections=4)
print(f" IR: {len(doc2.sections)} sections, {doc2.claim_count()} claims, "
f"ungrounded={doc2.ungrounded_count()}")
_emit_all(doc2, "engram-temporal")
# echo both markdowns so they are visible in the run log
for stem, doc in (("neuron-self", doc1), ("engram-temporal", doc2)):
print(f"\n===== GENERATED MARKDOWN — {stem} =====\n")
print(pipeline.emit(doc, "markdown").decode())
if __name__ == "__main__":
main()
+129
View File
@@ -0,0 +1,129 @@
"""geometry.py — READ-ONLY loader for a meaning-geometry region.
The efferent projector never writes to the soul. This module reaches the
geometry through the PROVEN, read-only neuron-talk faculty (``engram_client``,
GET-only, which physically refuses non-GET methods) against the running sidecar
soul. The live daemon :8742 / :7770 is contacted ONLY through that read-only
client — never mutated.
A "region" is a seed node plus a bounded neighborhood: the manifold that will
become the document's skeleton. We pool a few single-term lexical searches
(the engram search is a single-term matcher) and, when available, walk one hop
of reified neighbors, then rank by self/importance signal.
"""
from __future__ import annotations
import os
import sys
# Wire in the proven faculty (own-the-core: we reuse it, we do not fork it).
_NT = os.path.expanduser("~/Desktop/neuron-talk")
_LR = os.path.expanduser("~/Desktop/lang-realizers")
for _p in (_NT, _LR):
if _p not in sys.path:
sys.path.insert(0, _p)
from engram_client import ReadOnlyEngramClient # noqa: E402
class Region:
"""A geometry region: ranked nodes + the reified edges among them."""
def __init__(self, seed: str, nodes: list[dict], edges: list[dict]):
self.seed = seed
self.nodes = nodes # ranked engram node dicts
self.edges = edges # [{src, dst, edge, ...}]
self.by_id = {n["id"]: n for n in nodes if n.get("id")}
def __repr__(self):
return f"<Region seed={self.seed!r} nodes={len(self.nodes)} edges={len(self.edges)}>"
def _prose_quality(content: str) -> float:
"""Reward clean expository prose; penalize shouty banner-dense nodes.
A high ALLCAPS-word ratio or very short content signals a banner/telegraphic
memory node that extracts into garbage. Clean declarative prose scores high.
"""
if not content or not content.strip():
return 0.0
words = content.split()
if len(words) < 8:
return 0.1
caps = sum(1 for w in words if len(w) > 2 and w.strip(".,:;'\"-").isupper())
caps_ratio = caps / max(1, len(words))
# sentences with lowercase interior words read as prose
lower = sum(1 for w in words if w[:1].islower())
lower_ratio = lower / max(1, len(words))
return max(0.0, 1.2 * lower_ratio - 2.0 * caps_ratio)
def _relevance(content: str, terms: list[str]) -> float:
"""Topical relevance to the seed terms — keeps a region ON-THEME so a clean
but off-topic node cannot hijack the document."""
if not terms:
return 0.0
low = (content or "").lower()
hits = sum(1 for t in terms if t.lower() in low)
return hits / max(1, len(terms))
def _node_rank(n: dict, terms: list[str] | None = None) -> float:
return (float(n.get("importance") or 0.0) * 2.0
+ float(n.get("salience") or 0.0)
+ 1.5 * _prose_quality(n.get("content") or "")
+ 2.0 * _relevance(n.get("content") or "", terms or [])
+ (0.5 if (n.get("content") or "").strip() else 0.0))
def load_region(seed_terms: list[str] | str, *, client: ReadOnlyEngramClient | None = None,
max_nodes: int = 10, per_term: int = 20, hop: bool = True) -> Region:
"""Pull a bounded geometry region around ``seed_terms`` (read-only).
``seed_terms`` may be a single string or several probe terms; results are
pooled and de-duplicated. When ``hop`` and the reified neighbor endpoint is
live, one hop of neighbors is folded in so the region is a real
neighborhood, not just a keyword hit list.
"""
client = client or ReadOnlyEngramClient()
if isinstance(seed_terms, str):
seed_terms = [seed_terms]
pool: dict[str, dict] = {}
for term in seed_terms:
for n in client.search(term, limit=per_term):
if isinstance(n, dict) and n.get("id"):
pool.setdefault(n["id"], n)
ranked = sorted(pool.values(), key=lambda n: _node_rank(n, seed_terms),
reverse=True)
nodes = ranked[:max_nodes]
edges: list[dict] = []
if hop and nodes:
present = {n["id"] for n in nodes}
for n in list(nodes):
try:
for nb in client.neighbors(n["id"]):
node = nb.get("node") if isinstance(nb, dict) else None
edge = nb.get("edge") if isinstance(nb, dict) else None
if node and node.get("id"):
edges.append({"src": n["id"], "dst": node["id"],
"edge": edge})
# fold a strong neighbor into the region (bounded)
if (node["id"] not in present and len(nodes) < max_nodes + 6
and _node_rank(node, seed_terms) > 0.4):
present.add(node["id"])
nodes.append(node)
except Exception: # noqa: BLE001 — read-only best-effort; never fatal
continue
return Region(seed=", ".join(seed_terms), nodes=nodes, edges=edges)
def load_self_region(client: ReadOnlyEngramClient | None = None,
max_nodes: int = 10) -> Region:
"""The self/identity region — Neuron's own geometry, for self-description."""
return load_region(["self", "identity", "Neuron", "values", "memory",
"imprint", "consciousness"],
client=client, max_nodes=max_nodes)
+67
View File
@@ -0,0 +1,67 @@
"""pipeline.py — the Efferent Multimodal Projector, top level.
geometry region + surface/format spec
-> PLAN (manifold -> document skeleton/DAG)
-> REALIZE (proven realizer, sentence -> passage, each section faithful)
-> COHERE (document-level flow / transitions, not stitched sentences)
-> EMIT (pluggable SurfaceProjector -> the target surface)
THE SURFACE IS A PARAMETER. ``project(...)`` builds the geometry-carrying
DocumentIR once, then hands it to whichever surface projector the caller named.
Markdown, docx, and midi (music) are all the SAME IR emitted differently. That
is the efferent multimodal projector: geometry -> any surface.
"""
from __future__ import annotations
import os
import sys
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
sys.path.insert(0, os.path.join(_HERE, "projectors"))
from cohere import cohere_document # noqa: E402
from document_ir import DocumentIR # noqa: E402
from geometry import Region, load_region # noqa: E402
from plan import plan_document # noqa: E402
from realize import realize_document # noqa: E402
# registering the projectors (import for side-effect: each self-registers)
import projectors.markdown # noqa: E402,F401
import projectors.docx # noqa: E402,F401
import projectors.midi # noqa: E402,F401
import projectors.seams # noqa: E402,F401
from projectors.base import available_surfaces, get_projector # noqa: E402
def build_ir(seed_terms, *, title: str, subtitle: str = "",
format_spec: dict | None = None,
region: Region | None = None,
max_sections: int = 8, conf_floor: float = 0.55) -> DocumentIR:
"""geometry -> PLAN -> REALIZE -> COHERE = the surface-neutral DocumentIR."""
region = region or load_region(seed_terms)
doc = plan_document(region, title=title, subtitle=subtitle,
format_spec=format_spec or {},
conf_floor=conf_floor, max_sections=max_sections)
doc = realize_document(doc)
doc = cohere_document(doc)
return doc
def emit(doc: DocumentIR, surface: str) -> bytes:
"""EMIT: project the built IR onto one surface (surface = a parameter)."""
return get_projector(surface).project(doc)
def project(seed_terms, *, surface: str, title: str, subtitle: str = "",
format_spec: dict | None = None, region: Region | None = None,
max_sections: int = 8) -> tuple[DocumentIR, bytes]:
"""The full efferent projection: geometry + surface -> (IR, bytes)."""
doc = build_ir(seed_terms, title=title, subtitle=subtitle,
format_spec=format_spec, region=region,
max_sections=max_sections)
return doc, emit(doc, surface)
__all__ = ["build_ir", "emit", "project", "available_surfaces",
"get_projector", "load_region", "DocumentIR"]
+192
View File
@@ -0,0 +1,192 @@
"""plan.py — PLAN stage: geometry region -> document skeleton (a DAG/outline).
The manifold becomes the skeleton. We extract faithful propositions from the
region's nodes (the proven neuron-talk extractor, SACRED polarity preserved),
apply a quality floor, then GROUP them into sections. Grouping is by source
node — each engram node is one coherent topic, so one salient node becomes one
section. The section ORDER is the node ranking (importance/salience): the
geometry decides the outline, not a template.
Output: a DocumentIR whose sections carry seed node ids and empty blocks. REALIZE
fills the blocks; the plan owns the structure.
"""
from __future__ import annotations
import os
import re
import sys
_NT = os.path.expanduser("~/Desktop/neuron-talk")
_LR = os.path.expanduser("~/Desktop/lang-realizers")
for _p in (_NT, _LR):
if _p not in sys.path:
sys.path.insert(0, _p)
import propositions # noqa: E402 (the proven, faithful extractor)
from document_ir import DocumentIR, Section # noqa: E402
from geometry import Region # noqa: E402
# --------------------------------------------------------------------------- #
# Proposition quality — keep only clean, well-grounded claims.
# --------------------------------------------------------------------------- #
_JUNK_RE = re.compile(r"[.][a-z]{1,3}\b|[^A-Za-z0-9 '\-]") # ".o", stray symbols
def _has_banner_token(s: str) -> bool:
"""True if any word is an ALLCAPS banner token (DHARMA, ENGRAM, MEASURED)."""
for w in (s or "").split():
core = w.strip(".,:;'\"-")
if len(core) > 2 and core.isupper():
return True
return False
def _clean_prop(p, floor: float) -> bool:
if p.confidence < floor:
return False
if not p.subject or not (p.object or (p.obj_np is not None)):
return False
subj = (p.subject or "").strip()
obj = (p.object or "").strip()
if len(subj) < 2:
return False
# banner-derived shouty fragments read as garbage in prose
if _has_banner_token(subj) or _has_banner_token(obj):
return False
if propositions._is_shouty(p.sentence or ""):
return False
# junk tokens: file-extension fragments (".o"), stray non-word symbols
if _JUNK_RE.search(subj) or _JUNK_RE.search(obj):
return False
# a proposition whose object repeats the subject is usually a parse artifact
if obj and subj.lower() == obj.lower():
return False
# a bare copula with no real complement ("X is it") reads as noise
if p.predicate == "be" and obj.lower() in ("it", "no", "nothing", "empty", ""):
return False
return True
def _dedup(props):
"""Drop duplicate claims. Two axes: (a) identical (pred,obj,polarity), and
(b) same (subject,predicate) — which collapses a mis-split compound like
"detection is post-hoc eval" -> "Detection is post/hoc/eval" into one claim
(keep the highest-confidence surface)."""
props = sorted(props, key=lambda p: p.confidence, reverse=True)
seen_po, seen_sp, out = set(), set(), []
for p in props:
subj = (p.subject or "").lower()
po = (p.predicate, (p.object or "").lower(), p.polarity)
sp = (subj, p.predicate, p.polarity)
if po in seen_po or sp in seen_sp:
continue
seen_po.add(po)
seen_sp.add(sp)
out.append(p)
return out
# --------------------------------------------------------------------------- #
# Heading derivation — a clean human heading from a node.
# --------------------------------------------------------------------------- #
_HEADING_RE = re.compile(r"^\s*#{1,4}\s+(.{2,70})\s*$", re.M)
# node-type / system labels that are NOT topical headings
_NONTOPIC_LABEL = re.compile(r"^(memory|node|knowledge|doc|session)[:/]", re.I)
def _titlecase_banner(s: str) -> str:
"""A shouty banner ("CHRONOCEPTION — SCALE-INVARIANCE") makes a fine title
once Title-cased. Keep short acronyms uppercase."""
def fix(w):
core = w.strip("—-:,.")
if len(core) <= 3 and core.isupper():
return w # acronym
return w.capitalize()
return " ".join(fix(w) for w in s.split())
def _clean_heading(text: str) -> str | None:
"""First line only, no markdown, capped, banner Title-cased. None if unusable."""
if not text:
return None
line = text.strip().splitlines()[0]
line = re.sub(r"^#+\s*", "", line).strip().strip("#").strip()
# cut at a natural break so a long banner heading stays a heading, not a para
for sep in ("", " ", ": ", ". "):
if sep in line and len(line) > 48:
line = line.split(sep)[0].strip()
break
if not (3 <= len(line) <= 64):
return None
if propositions._is_shouty(line):
line = _titlecase_banner(line)
return line or None
def _heading_for(node: dict, fallback: str) -> str:
label = (node.get("label") or "").strip()
content = node.get("content") or ""
candidates: list[str] = []
# a node-type label ("memory:remembered") is never a topic — skip it
if label and not _NONTOPIC_LABEL.match(label):
candidates.append(label)
m = _HEADING_RE.search(content)
if m:
candidates.append(m.group(1))
# the leading banner/first sentence of the content is often the real title
first = re.split(r"(?<=[.\n])", content.strip(), maxsplit=1)[0] if content.strip() else ""
candidates.append(first)
for c in candidates:
h = _clean_heading(c)
if h:
return h
return fallback
def plan_document(region: Region, *, title: str, subtitle: str = "",
format_spec: dict | None = None,
conf_floor: float = 0.55,
max_sections: int = 8,
max_claims_per_section: int = 6) -> DocumentIR:
"""Region -> DocumentIR skeleton. The geometry dictates the outline."""
format_spec = format_spec or {}
doc = DocumentIR(title=title, subtitle=subtitle,
seed_id=region.nodes[0]["id"] if region.nodes else None,
format_spec=format_spec)
made = 0
seen_headings: set[str] = set()
for node in region.nodes:
if made >= max_sections:
break
props = propositions.extract(node.get("content") or "",
node_id=node.get("id"),
node_importance=float(node.get("importance") or 0.0),
max_sentences=10)
props = [p for p in props if _clean_prop(p, conf_floor)]
props = _dedup(props)
props.sort(key=lambda p: p.confidence, reverse=True)
props = props[:max_claims_per_section]
if not props:
continue
heading = _heading_for(node, fallback=f"Region {made + 1}")
# cross-section dedup: a topic appears once. Distinguish by top claim
# subject, else drop the collision so the outline stays clean.
if heading.lower() in seen_headings:
subj = (props[0].subject or "").strip().title()
alt = f"{heading}: {subj}" if subj and subj.lower() not in heading.lower() else None
if alt and alt.lower() not in seen_headings and len(alt) <= 64:
heading = alt
else:
continue
seen_headings.add(heading.lower())
sec = Section(heading=heading, level=2, seed_ids=[node["id"]])
# stash the planned propositions on the section for REALIZE
sec.__dict__["_planned_props"] = props
sec.__dict__["_node"] = node
doc.sections.append(sec)
made += 1
return doc
+106
View File
@@ -0,0 +1,106 @@
"""base.py — the SurfaceProjector interface + registry.
THE key abstraction of the efferent projector: a projector is a pure function
from the surface-neutral, geometry-carrying DocumentIR to bytes on a target
SURFACE. The surface is a PARAMETER. Adding a surface = registering one more
projector; nothing upstream (plan/realize/cohere) changes.
DocumentIR --project--> bytes (per surface)
A TEXT projector reads ``block.sentences``. A NON-TEXT projector (music, image,
video) reads ``block.provenance`` — the geometry the IR carries — and decodes it
onto its surface. Both consume the SAME IR. That symmetry is the whole design:
the realizer generalizes into a multimodal projector, geometry -> any surface.
"""
from __future__ import annotations
from typing import Protocol, runtime_checkable
import sys
import os
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from document_ir import DocumentIR # noqa: E402
@runtime_checkable
class SurfaceProjector(Protocol):
"""Geometry-document -> one surface. Implementations MUST be pure & faithful.
THE ONE SHARED SEAM. Every surface — text, music, image, video — conforms to
this single contract:
project(frame: DocumentIR) -> bytes
where ``frame`` is the geometry-carrying meaning-geometry (the SemFrame at
document scale; a single utterance is the degenerate one-section frame).
RECOMMENDED INTERNAL SHAPE (the peer music/text decomposition, blessed here
so all surfaces share it): a projector may split ``project`` into
spec = self.plan(frame) # meaning-geometry -> surface-specific spec
bytes = self.realize(spec) # spec -> surface, via this projector's PROFILE
``project`` is then ``realize(plan(frame))``. The PROFILE (a text lang-profile,
a music instr/mode-profile, an image layout-profile) is a property of the
projector instance — the pluggable knob. See :class:`TwoStageProjector`.
A TEXT projector's plan reads ``frame`` sentences; a MUSIC/IMAGE projector's
plan reads ``frame.all_provenance()`` — the geometry — and derives its spec
(pitch/harmony/rhythm, or layout) FROM the meaning, deterministically. Same
frame, different profile.
"""
surface: str # "markdown" | "docx" | "midi" | "audio" | "image" | "video"
media_type: str # MIME type of the emitted bytes
ext: str # file extension (no dot)
modality: str # "text" | "audio" | "image" | "video"
profile: object # the pluggable per-surface profile (may be None)
def project(self, doc: DocumentIR) -> bytes:
"""Emit the document on this surface. Returns raw bytes."""
...
class TwoStageProjector:
"""Optional base for the peer plan()/realize() decomposition.
Subclasses implement ``plan(frame) -> spec`` and ``realize(spec) -> bytes``;
``project`` is their composition. This is exactly the peer music interface
(spec = plan(frame, profile); surface = realize(spec, profile)) expressed so
that it still satisfies the single ``SurfaceProjector.project`` seam. Text,
music, and image projectors can all subclass this and remain interchangeable.
"""
surface: str = ""
media_type: str = ""
ext: str = ""
modality: str = ""
profile: object = None
def plan(self, doc: DocumentIR): # -> spec
raise NotImplementedError
def realize(self, spec) -> bytes:
raise NotImplementedError
def project(self, doc: DocumentIR) -> bytes:
return self.realize(self.plan(doc))
_REGISTRY: dict[str, SurfaceProjector] = {}
def register(projector: SurfaceProjector) -> SurfaceProjector:
_REGISTRY[projector.surface] = projector
return projector
def get_projector(surface: str) -> SurfaceProjector:
if surface not in _REGISTRY:
raise KeyError(f"no projector registered for surface {surface!r}; "
f"have {sorted(_REGISTRY)}")
return _REGISTRY[surface]
def available_surfaces() -> list[str]:
return sorted(_REGISTRY)
+113
View File
@@ -0,0 +1,113 @@
"""docx.py — the .docx surface projector: an OWN minimal OOXML emitter.
Own-the-core: a .docx is just a ZIP of a few XML parts (WordprocessingML). We
emit it with the standard library only — ``zipfile`` + string XML — no
python-docx, no external dependency. This proves a "richer structured format"
surface without importing anyone else's toolkit.
Parts emitted (the minimal valid set + a styles part for real headings):
[Content_Types].xml
_rels/.rels
word/_rels/document.xml.rels
word/styles.xml (Title / Heading1 / Heading2 / Normal)
word/document.xml (the content)
Like the markdown projector it reads only the IR's realized sentences; it
invents nothing. The surface differs, the faithful content does not.
"""
from __future__ import annotations
import io
import os
import sys
import zipfile
from xml.sax.saxutils import escape
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from document_ir import DocumentIR # noqa: E402
from projectors.base import register # noqa: E402
_CONTENT_TYPES = """<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">
<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>
<Default Extension="xml" ContentType="application/xml"/>
<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>
<Override PartName="/word/styles.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.styles+xml"/>
</Types>"""
_RELS = """<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">
<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/>
</Relationships>"""
_DOC_RELS = """<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">
<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/styles" Target="styles.xml"/>
</Relationships>"""
_W = "http://schemas.openxmlformats.org/wordprocessingml/2006/main"
_STYLES = f"""<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<w:styles xmlns:w="{_W}">
<w:style w:type="paragraph" w:default="1" w:styleId="Normal"><w:name w:val="Normal"/>
<w:rPr><w:sz w:val="22"/></w:rPr></w:style>
<w:style w:type="paragraph" w:styleId="Title"><w:name w:val="Title"/>
<w:pPr><w:spacing w:after="240"/></w:pPr>
<w:rPr><w:b/><w:sz w:val="52"/></w:rPr></w:style>
<w:style w:type="paragraph" w:styleId="Subtitle"><w:name w:val="Subtitle"/>
<w:rPr><w:i/><w:sz w:val="28"/><w:color w:val="555555"/></w:rPr></w:style>
<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/>
<w:pPr><w:spacing w:before="240" w:after="120"/><w:outlineLvl w:val="0"/></w:pPr>
<w:rPr><w:b/><w:sz w:val="34"/></w:rPr></w:style>
<w:style w:type="paragraph" w:styleId="Heading2"><w:name w:val="heading 2"/>
<w:pPr><w:spacing w:before="200" w:after="100"/><w:outlineLvl w:val="1"/></w:pPr>
<w:rPr><w:b/><w:sz w:val="28"/></w:rPr></w:style>
</w:styles>"""
def _para(text: str, style: str | None = None) -> str:
ppr = f"<w:pPr><w:pStyle w:val=\"{style}\"/></w:pPr>" if style else ""
return (f"<w:p>{ppr}<w:r><w:t xml:space=\"preserve\">"
f"{escape(text)}</w:t></w:r></w:p>")
class DocxProjector:
surface = "docx"
media_type = ("application/vnd.openxmlformats-officedocument."
"wordprocessingml.document")
ext = "docx"
modality = "text"
def _document_xml(self, doc: DocumentIR) -> str:
body: list[str] = [_para(doc.title, "Title")]
if doc.subtitle:
body.append(_para(doc.subtitle, "Subtitle"))
abstract = doc.meta.get("abstract")
if abstract is not None and abstract.sentences:
body.append(_para(abstract.text()))
for sec in doc.sections:
style = "Heading1" if sec.level <= 1 else "Heading2"
body.append(_para(sec.heading, style))
for block in sec.blocks:
t = block.text()
if t:
body.append(_para(t))
return (f"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>"
f"<w:document xmlns:w=\"{_W}\"><w:body>"
+ "".join(body)
+ "<w:sectPr><w:pgSz w:w=\"12240\" w:h=\"15840\"/>"
"<w:pgMar w:top=\"1440\" w:right=\"1440\" w:bottom=\"1440\" "
"w:left=\"1440\"/></w:sectPr></w:body></w:document>")
def project(self, doc: DocumentIR) -> bytes:
buf = io.BytesIO()
with zipfile.ZipFile(buf, "w", zipfile.ZIP_DEFLATED) as z:
z.writestr("[Content_Types].xml", _CONTENT_TYPES)
z.writestr("_rels/.rels", _RELS)
z.writestr("word/_rels/document.xml.rels", _DOC_RELS)
z.writestr("word/styles.xml", _STYLES)
z.writestr("word/document.xml", self._document_xml(doc))
return buf.getvalue()
register(DocxProjector())
+45
View File
@@ -0,0 +1,45 @@
"""markdown.py — the Markdown surface projector (text facet).
The most tractable surface, and the reference implementation: reads the IR's
realized sentences and lays them out as Markdown. Introduces no content — it is
pure typography over the faithful text the realizer produced.
"""
from __future__ import annotations
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from document_ir import DocumentIR # noqa: E402
from projectors.base import register # noqa: E402
class MarkdownProjector:
surface = "markdown"
media_type = "text/markdown"
ext = "md"
modality = "text"
def render_str(self, doc: DocumentIR) -> str:
lines: list[str] = [f"# {doc.title}"]
if doc.subtitle:
lines.append(f"\n*{doc.subtitle}*")
abstract = doc.meta.get("abstract")
if abstract is not None and abstract.sentences:
lines.append("")
lines.append(abstract.text())
for sec in doc.sections:
lines.append("")
lines.append(f"{'#' * max(2, sec.level)} {sec.heading}")
for block in sec.blocks:
body = block.text()
if body:
lines.append("")
lines.append(body)
return "\n".join(lines) + "\n"
def project(self, doc: DocumentIR) -> bytes:
return self.render_str(doc).encode("utf-8")
register(MarkdownProjector())
+133
View File
@@ -0,0 +1,133 @@
"""midi.py — the MUSIC surface projector: geometry -> symbolic music (MIDI).
The first NON-TEXT surface, and the proof of the general shape. "Music is
language and it is math" (Will): symbolic music is tractable and geometry-native,
so it is the natural efferent twin to try first after text.
CRUCIALLY this projector does NOT read the realized sentences. It reads the IR's
GEOMETRY facet — ``block.provenance`` — and DECODES each edge onto a musical
surface. That is the whole thesis of the multimodal projector: the same
geometry-carrying IR drives text AND music; a text projector reads the words, a
music projector reads the meaning-geometry. The mapping is deterministic and
faithful to the geometry's structure:
relation lemma -> scale degree (same relation -> same pitch class;
meaning has a consistent sonic form)
polarity -> mode (aff = major third above; neg = minor
third / lowered — SACRED polarity is
audible, a negated edge sounds negated)
confidence -> note duration (stronger grounding rings longer)
importance -> velocity (more important source = louder)
section -> phrase + register shift (structure becomes musical form)
Own-the-core: a Standard MIDI File is a header chunk + a track chunk of
delta-timed events. We emit the raw bytes with ``struct`` — no external MIDI
library. Format 0, one track.
"""
from __future__ import annotations
import io
import os
import struct
import sys
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from document_ir import DocumentIR, Provenance # noqa: E402
from projectors.base import TwoStageProjector, register # noqa: E402
_TICKS = 480 # ticks per quarter note
_C_MAJOR = [0, 2, 4, 5, 7, 9, 11] # semitone offsets of a diatonic scale
def _vlq(n: int) -> bytes:
"""MIDI variable-length quantity encoding of a delta time."""
if n == 0:
return b"\x00"
out = bytearray()
out.append(n & 0x7F)
n >>= 7
while n:
out.insert(0, (n & 0x7F) | 0x80)
n >>= 7
return bytes(out)
def _degree_for(relation: str) -> int:
"""Stable scale degree for a relation lemma (same relation -> same pitch)."""
if not relation:
return 0
return sum(ord(c) for c in relation.lower()) % len(_C_MAJOR)
def _note_for(p: Provenance, base: int) -> tuple[int, int, int]:
"""(pitch, velocity, duration_ticks) for one geometry edge."""
root = base + _C_MAJOR[_degree_for(p.relation)]
# polarity -> mode: affirmed edges take the bright major third, negated edges
# take the darker minor third. The negation is AUDIBLE and never dropped.
third = 4 if p.polarity == "aff" else 3
pitch = max(24, min(96, root + (third if p.confidence >= 0.5 else 0)))
velocity = int(56 + 60 * min(1.0, max(0.0, p.importance)))
velocity = max(40, min(120, velocity))
# confidence -> duration: quarter .. dotted-half
dur = int(_TICKS * (0.5 + 1.5 * min(1.0, max(0.0, p.confidence))))
return pitch, velocity, dur
# a mode-profile: the pluggable musical knob (the peer's mode_profile). Scale +
# tempo. Swapping this profile re-voices the SAME geometry — surface as parameter.
_DEFAULT_PROFILE = {"scale": _C_MAJOR, "tempo_us": 500000,
"registers": [60, 55, 64, 50, 67, 48], "program": 0}
class MidiProjector(TwoStageProjector):
"""geometry -> symbolic music, in the shared two-stage shape.
``plan(frame)`` -> a music_spec: an ordered list of note dicts derived
deterministically from the frame's provenance geometry
(the peer's ``plan(frame, profile) -> spec``).
``realize(spec)`` -> Standard MIDI File bytes (the peer's
``realize(spec, profile) -> surface``; here the surface
is symbolic MIDI, the minimal audio proof — a richer
additive-synth audio projector conforms identically).
"""
surface = "midi"
media_type = "audio/midi"
ext = "mid"
modality = "audio"
def __init__(self, profile: dict | None = None):
self.profile = profile or _DEFAULT_PROFILE
# -- stage 1: meaning-geometry -> music_spec (reads the GEOMETRY facet) -- #
def plan(self, doc: DocumentIR) -> list[dict]:
registers = self.profile["registers"]
spec: list[dict] = []
for si, sec in enumerate(doc.sections):
base = registers[si % len(registers)]
provs = [p for p in sec.all_provenance()
if p.kind in ("fact", "interpretation")]
for i, p in enumerate(provs):
pitch, vel, dur = _note_for(p, base)
spec.append({"pitch": pitch, "velocity": vel, "dur": dur,
"rest_before": (_TICKS // 2) if (si > 0 and i == 0) else 0,
"relation": p.relation, "polarity": p.polarity})
return spec
# -- stage 2: music_spec -> MIDI bytes (own-core, no library) ------------ #
def realize(self, spec: list[dict]) -> bytes:
ev = bytearray()
ev += _vlq(0) + b"\xFF\x51\x03" + struct.pack(">I", self.profile["tempo_us"])[1:]
ev += _vlq(0) + bytes([0xC0, self.profile["program"] & 0x7F])
for note in spec:
ev += _vlq(note["rest_before"]) + bytes([0x90, note["pitch"], note["velocity"]])
ev += _vlq(note["dur"]) + bytes([0x80, note["pitch"], 0])
ev += _vlq(0) + b"\xFF\x2F\x00"
track = bytes(ev)
buf = io.BytesIO()
buf.write(b"MThd" + struct.pack(">IHHH", 6, 0, 1, _TICKS))
buf.write(b"MTrk" + struct.pack(">I", len(track)) + track)
return buf.getvalue()
register(MidiProjector())
+60
View File
@@ -0,0 +1,60 @@
"""seams.py — documented efferent seams for IMAGE and VIDEO surfaces.
These are NOT implemented (per the build rails: architect, do not overbuild).
They are registered as first-class seams so the interface PROVES it accepts
future non-text projectors without any upstream change. Each documents exactly
what its decoder would read from the geometry-carrying IR, making the multimodal
generalization concrete rather than hand-wavy.
The symmetry that guarantees these are possible, not moonshots: they are the
efferent twins of multimodal INGEST. If meaning can HOLD an image (ingest as
first-class geometry), meaning can PROJECT one back. Video = image x sound x
TIME, and the engram already stores time (chronoception). So video falls out of
an image projector + the music projector + the stored temporal ordering.
"""
from __future__ import annotations
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from document_ir import DocumentIR # noqa: E402
from projectors.base import register # noqa: E402
class _Seam:
"""A registered-but-unimplemented projector. Names its decoder contract."""
def project(self, doc: DocumentIR) -> bytes: # pragma: no cover - seam
raise NotImplementedError(
f"{self.surface!r} projector is a documented seam, not yet built. "
f"Decoder contract: {self.decoder_contract}")
class ImageProjector(_Seam):
surface = "image"
media_type = "image/png"
ext = "png"
modality = "image"
decoder_contract = (
"reads block.provenance as a spatial layout — nodes become regions, edges "
"become adjacencies; salience/importance drive size/contrast; polarity "
"drives figure/ground. The efferent twin of image ingest (a geometry->raster "
"decoder, learned or engineered), exactly mirroring the embedder that turned "
"the image INTO geometry.")
class VideoProjector(_Seam):
surface = "video"
media_type = "video/mp4"
ext = "mp4"
modality = "video"
decoder_contract = (
"image x sound x TIME. Composes the image projector (per-keyframe geometry "
"layout) with the midi/music projector (score) along the geometry's stored "
"temporal ordering (chronoception). Needs no new principle once image + music "
"exist — only a muxer.")
register(ImageProjector())
register(VideoProjector())
+63
View File
@@ -0,0 +1,63 @@
"""provenance.py — the faithfulness audit + geometry->section trace.
A document projected from geometry is only worth anything if every claim traces
back. This module walks the DocumentIR and proves the discipline held:
* ZERO ungrounded claims (every fact/interpretation has a real node id),
* every emitted sentence maps to a geometry edge (or is a marked connective),
* SACRED polarity survived (negations are reported, never silently dropped),
* COHERE introduced no new geometry (connectives carry no claim).
It emits both a machine verdict and a human-readable geometry->section table.
"""
from __future__ import annotations
from document_ir import DocumentIR
def audit(doc: DocumentIR) -> dict:
provs = doc.all_provenance()
facts = [p for p in provs if p.kind in ("fact", "interpretation")]
connectives = [p for p in provs if p.kind == "connective"]
ungrounded = [p for p in facts if not p.node_id]
negations = [p for p in facts if p.polarity == "neg"]
node_ids = sorted({p.node_id for p in facts if p.node_id})
return {
"claims": len(facts),
"connectives": len(connectives),
"ungrounded_claims": len(ungrounded),
"negations_preserved": len(negations),
"distinct_source_nodes": len(node_ids),
"faithful": len(ungrounded) == 0,
"source_nodes": node_ids,
}
def trace_table(doc: DocumentIR) -> str:
"""Human-readable geometry -> section -> claim provenance table."""
lines = ["# Provenance — every claim traces geometry", ""]
lines.append(f"**Document:** {doc.title}")
a = audit(doc)
lines.append(f"**Claims:** {a['claims']} · **Ungrounded:** "
f"{a['ungrounded_claims']} · **Negations preserved:** "
f"{a['negations_preserved']} · **Source nodes:** "
f"{a['distinct_source_nodes']} · **Faithful:** "
f"{'YES' if a['faithful'] else 'NO'}")
lines.append("")
for si, sec in enumerate(doc.sections, 1):
lines.append(f"## {si}. {sec.heading}")
lines.append(f"_seed nodes: {', '.join(i[:8] for i in sec.seed_ids)}_")
lines.append("")
lines.append("| # | realized claim | traces geometry edge |")
lines.append("|---|----------------|----------------------|")
n = 0
for block in sec.blocks:
for sent, prov in zip(block.sentences, block.provenance):
if prov.kind == "connective":
continue
n += 1
edge = prov.trace().replace("|", "\\|")
s = sent.replace("|", "\\|")
lines.append(f"| {n} | {s} | {edge} |")
lines.append("")
return "\n".join(lines) + "\n"
+112
View File
@@ -0,0 +1,112 @@
"""realize.py — REALIZE stage: fill each planned section with faithful passages.
Scales the PROVEN realizer from a single assertion to a passage. For each
planned proposition we build a realizer-ready clause (the proven
``_prop_to_clause`` mapping) and run it through the proven engine
(``engine.realize``), which is a deterministic grammar with the SACRED negation
contract — it never invents. Each realized sentence is paired with a
:class:`Provenance` that pins it to the exact geometry edge it came from.
"Passage, not a list of sentences": within a section we lightly vary sentence
openings and group related claims, but we add NO content the geometry did not
assert. The only non-geometry words are function words the grammar already owns
(articles, "and", conjunction of same-subject claims). Document-level flow is
COHERE's job; this stage owns intra-section fluency + fidelity.
"""
from __future__ import annotations
import os
import sys
_NT = os.path.expanduser("~/Desktop/neuron-talk")
_LR = os.path.expanduser("~/Desktop/lang-realizers")
for _p in (_NT, _LR):
if _p not in sys.path:
sys.path.insert(0, _p)
import engine # noqa: E402 (the proven no-LLM realizer)
from dialogue import _prop_to_clause # noqa: E402 (proven prop -> clause)
from document_ir import Block, DocumentIR, Provenance, Section # noqa: E402
def _provenance_from(p, kind: str = "fact") -> Provenance:
return Provenance(
subj_id=p.source_node_id, subject=p.subject, relation=p.predicate,
obj=p.object, polarity=p.polarity, confidence=round(float(p.confidence), 3),
node_id=p.source_node_id, kind=kind,
importance=float(getattr(p, "node_importance", 0.0) or 0.0),
salience=0.0,
)
import re as _re
# a well-formed declarative opens with a determiner, a proper noun, "I", or a
# capitalized head — not a mis-parsed object pronoun or a copula fragment.
_BAD_OPENERS = _re.compile(r"^(Me |It is I|There is|This is it|That is it)\b")
_VACUOUS = _re.compile(r"^\w+ (is|are|was|were) (it|no|nothing|empty|those|this|that)\.?$",
_re.I)
def _good_sentence(text: str) -> bool:
"""Fluency gate — drops degenerate realizations. NEVER loosens faithfulness;
it only refuses to SPEAK a claim whose surface came out malformed."""
words = text.rstrip(".").split()
if len(words) < 3:
return False
if _BAD_OPENERS.search(text):
return False
if _VACUOUS.match(text):
return False
# a sentence that is mostly one-letter/two-letter tokens is a parse artifact
short = sum(1 for w in words if len(w.strip(".,'")) <= 2)
if short > len(words) / 2:
return False
return True
def _realize_prop(p, lang: str = "en") -> tuple[str, Provenance] | None:
"""One proposition -> (faithful sentence, provenance) or None if it drops."""
clause = _prop_to_clause(p)
text = engine.realize(clause, lang)
if not text or not text.strip():
return None
text = text.strip()
if not text.endswith((".", "!", "?")):
text += "."
# capitalize first character (proper nouns / "I" already handled by grammar)
text = text[0].upper() + text[1:]
if not _good_sentence(text):
return None
return text, _provenance_from(p)
def realize_document(doc: DocumentIR, lang: str = "en") -> DocumentIR:
"""Fill every planned section's blocks with faithful, realized passages."""
for sec in doc.sections:
planned = sec.__dict__.get("_planned_props", [])
block = Block(role="body")
summary_bits: list[str] = []
for p in planned:
r = _realize_prop(p, lang)
if r is None:
continue
text, prov = r
block.sentences.append(text)
block.provenance.append(prov)
if len(summary_bits) < 1:
# a short grounded gloss for TOC / pptx bullets
obj = (prov.obj or "").strip().rstrip(".")
if obj:
summary_bits.append(obj)
if block.sentences:
sec.blocks.append(block)
sec.summary = summary_bits[0] if summary_bits else ""
# drop the transient planning payload; the IR is now self-contained
sec.__dict__.pop("_planned_props", None)
sec.__dict__.pop("_node", None)
# prune sections that realized to nothing
doc.sections = [s for s in doc.sections if s.blocks]
return doc
+73
View File
@@ -0,0 +1,73 @@
// audio-demo.el - Drive the native audio surface: render a tone per instrument
// from its LEARNED signature, then render a small meaning-phrase "piece".
// Entry point: top-level statement calls main() (same convention as the
// examples' top-level println(run_test())).
fn micros_to_str(xs: [Int]) -> String {
let n: Int = native_list_len(xs)
let out: String = ""
let i: Int = 0
while i < n {
if i > 0 { let out: String = out + "," }
let out: String = out + int_to_str(native_list_get(xs, i))
let i: Int = i + 1
}
return out
}
// Render a 1.0s A4 (midi 69) tone from a signature file, print the parsed
// partials (proving the numbers came from the engram .sig), write the WAV.
fn render_tone(name: String, sigpath: String, outpath: String, table: [Int]) -> Int {
let lines: [String] = sig_load(sigpath)
let partials: [Int] = parse_micros(sig_field(lines, "partials"))
println("[" + name + "] partials_n=" + sig_field(lines, "partials_n") + " parsed_partials_micro(scale 1e6)=" + micros_to_str(partials))
println("[" + name + "] raw partials line from .sig = " + sig_field(lines, "partials"))
let freq: Int = freq_of_midi(69)
let note: [Int] = synth_from_sig(lines, freq, 1000, 900, 44100, table)
let n: Int = native_list_len(note)
let ok: Int = wav_write(outpath, note, n, 44100)
println("[" + name + "] rendered " + int_to_str(n) + " samples -> " + outpath + " (write_ok=" + int_to_str(ok) + ")")
return n
}
fn run_demo() -> Int {
let table: [Int] = sin_table()
fs_mkdir("/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out")
println("=== TONES: render A4 (midi 69) from each learned signature ===")
render_tone("flute", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/flute.sig", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/tone-flute.wav", table)
render_tone("clarinet", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/clarinet.sig", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/tone-clarinet.wav", table)
render_tone("violin", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/violin.sig", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/tone-violin.wav", table)
render_tone("piano", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/piano.sig", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/tone-piano.wav", table)
render_tone("organ", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/organ.sig", "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/tone-organ.wav", table)
println("")
println("=== PIECE: a 6-frame meaning phrase (incl. a NEG frame) ===")
let frames: [[String]] = native_list_empty()
let frames: [[String]] = native_list_append(frames, audio_frame("agent", "aff", "0.9", "0.8", "0", "s1"))
let frames: [[String]] = native_list_append(frames, audio_frame("theme", "aff", "0.7", "0.6", "0", "s2"))
let frames: [[String]] = native_list_append(frames, audio_frame("cause", "aff", "0.8", "0.9", "1", "s3"))
let frames: [[String]] = native_list_append(frames, audio_frame("negation", "neg", "0.85", "0.7", "0", "s4"))
let frames: [[String]] = native_list_append(frames, audio_frame("goal", "aff", "0.6", "0.5", "1", "s5"))
let frames: [[String]] = native_list_append(frames, audio_frame("result", "aff", "0.95", "1.0", "0", "s6"))
// Print the plan so the NEG frame's minor third (+3) vs major (+4) is visible.
let nf: Int = native_list_len(frames)
let fi: Int = 0
while fi < nf {
let frame: [String] = native_list_get(frames, fi)
let plan: [Int] = plan_note(frame)
let pol: String = surface_get(frame, "polarity")
let third_name: String = "major(+4)"
if str_eq(pol, "neg") { let third_name: String = "MINOR(+3)" }
println("frame " + int_to_str(fi) + " relation=" + surface_get(frame, "relation") + " polarity=" + pol + " -> midi=" + int_to_str(native_list_get(plan, 0)) + " dur_ms=" + int_to_str(native_list_get(plan, 1)) + " amp_pm=" + int_to_str(native_list_get(plan, 2)) + " third=" + third_name)
let fi: Int = fi + 1
}
let piano_lines: [String] = sig_load("/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/piano.sig")
let total: Int = realize_audio(frames, piano_lines, "/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/piece.wav", 44100, table)
println("PIECE rendered " + int_to_str(total) + " samples -> /Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/piece.wav")
return total
}
println("audio-demo main returned samples=" + int_to_str(run_demo()))
+400
View File
@@ -0,0 +1,400 @@
// audio-surface.el - Native own-core additive-synthesis audio surface.
//
// The AUDIO efferent seam, native, no Python and no library. This renders real
// PCM .wav bytes from instrument SIGNATURES read from engram-sourced .sig data
// files (elp/faculty/sig/*.sig) - the partial amplitudes are NEVER literals in
// this source; they are parsed from the learned signature at run time. That is
// the whole proof: render-from-learned-signatures.
//
// EL has no float arithmetic operator (codegen emits raw int64 ops for + - * /
// on the shared 64-bit slot) and no float-arithmetic natives - so ALL synthesis
// math here is own-core INTEGER fixed-point. Angles use a quarter-wave sine
// table (scale 10000) from a fixed-point Taylor series; amplitudes are parsed to
// micro (scale 1e6) straight from the .sig text; frequencies are milliHz ints.
//
// Pipeline mirrors the two-stage projector (midi.py): plan_note(frame) reads a
// frame's meaning-geometry slot-map and derives (pitch, duration, amplitude);
// realize_audio SUPERPOSES the signature's partials (the compose op) and
// serialises RIFF/WAVE. Same frame -> midi OR audio.
// -- integer decimal + string helpers -----------------------------------------
fn str_to_int_el(s: String) -> Int {
let n: Int = str_len(s)
let i: Int = 0
let v: Int = 0
let neg: Bool = false
while i < n {
let c: Int = str_char_code(s, i)
if c == 45 { let neg: Bool = true }
if c >= 48 {
if c < 58 {
let v: Int = v * 10 + (c - 48)
}
}
let i: Int = i + 1
}
if neg { return 0 - v }
return v
}
fn parse_micro(s: String) -> Int {
let dot: Int = str_index_of(s, ".")
if dot < 0 {
return str_to_int_el(s) * 1000000
}
let n: Int = str_len(s)
let ipart: String = str_slice(s, 0, dot)
let fpart: String = str_slice(s, dot + 1, n)
let iv: Int = str_to_int_el(ipart)
let fv: Int = 0
let scale: Int = 100000
let fn2: Int = str_len(fpart)
let i: Int = 0
while i < 6 {
let d: Int = 0
if i < fn2 {
let d: Int = str_char_code(fpart, i) - 48
}
let fv: Int = fv + d * scale
let scale: Int = scale / 10
let i: Int = i + 1
}
return iv * 1000000 + fv
}
// -- signature (engram data file) loader ---------------------------------------
fn sig_load(path: String) -> [String] {
let text: String = fs_read(path)
return str_split(text, "\n")
}
fn sig_field(lines: [String], key: String) -> String {
let pref: String = key + ": "
let n: Int = native_list_len(lines)
let plen: Int = str_len(pref)
let i: Int = 0
while i < n {
let ln: String = native_list_get(lines, i)
if str_starts_with(ln, pref) {
return str_slice(ln, plen, str_len(ln))
}
let i: Int = i + 1
}
return ""
}
fn parse_micros(csv: String) -> [Int] {
let parts: [String] = str_split(csv, ",")
let n: Int = native_list_len(parts)
let out: [Int] = native_list_empty()
let i: Int = 0
while i < n {
let out: [Int] = native_list_append(out, parse_micro(native_list_get(parts, i)))
let i: Int = i + 1
}
return out
}
// -- fixed-point sine (own-core, quarter-wave Taylor table, scale 10000) --------
fn sin_table() -> [Int] {
let HP: Int = 1570796
let t: [Int] = native_list_empty()
let q: Int = 0
while q < 257 {
let x: Int = q * HP / 256
let x2: Int = x * x / 1000000
let x3: Int = x2 * x / 1000000
let x5: Int = x3 * x2 / 1000000
let x7: Int = x5 * x2 / 1000000
let x9: Int = x7 * x2 / 1000000
let s: Int = x - x3 / 6 + x5 / 120 - x7 / 5040 + x9 / 362880
let t: [Int] = native_list_append(t, s / 100)
let q: Int = q + 1
}
return t
}
fn sin_lookup(t: [Int], phase: Int) -> Int {
let p: Int = phase % 1024
if p < 0 { let p: Int = p + 1024 }
let quad: Int = p / 256
let r: Int = p % 256
if quad == 0 { return native_list_get(t, r) }
if quad == 1 { return native_list_get(t, 256 - r) }
if quad == 2 { return 0 - native_list_get(t, r) }
return 0 - native_list_get(t, 256 - r)
}
fn isqrt_int(n: Int) -> Int {
if n <= 0 { return 0 }
let x: Int = n
let y: Int = (x + 1) / 2
while y < x {
let x: Int = y
let y: Int = (x + n / x) / 2
}
return x
}
// freq_of_midi: equal-tempered frequency in milliHz. 440000 mHz at midi 69.
fn freq_of_midi(m: Int) -> Int {
let f: Int = 440000
if m > 69 {
let k: Int = m - 69
let i: Int = 0
while i < k {
let f: Int = f * 1059463 / 1000000
let i: Int = i + 1
}
return f
}
if m < 69 {
let k: Int = 69 - m
let i: Int = 0
while i < k {
let f: Int = f * 1000000 / 1059463
let i: Int = i + 1
}
return f
}
return f
}
// -- envelope (ADSR), scale 1000 -----------------------------------------------
fn adsr_env(i: Int, total: Int, atk_n: Int, dec_n: Int, sus_pm: Int, rel_n: Int) -> Int {
if i < atk_n {
if atk_n == 0 { return 1000 }
return 1000 * i / atk_n
}
if i < atk_n + dec_n {
if dec_n == 0 { return sus_pm }
return 1000 - (1000 - sus_pm) * (i - atk_n) / dec_n
}
let rel_start: Int = total - rel_n
if i < rel_start {
return sus_pm
}
if rel_n == 0 { return 0 }
let left: Int = total - i
return sus_pm * left / rel_n
}
// -- note synthesis: SUPERPOSE the learned partials -> [Int] samples -----------
fn note_samples(freq_mHz: Int, dur_ms: Int, rate: Int, partials: [Int], sumP: Int, b_micro: Int, vib_rate: Int, vib_cents: Int, atk_ms: Int, dec_ms: Int, sus_pm: Int, rel_ms: Int, amp_pm: Int, table: [Int]) -> [Int] {
let total: Int = dur_ms * rate / 1000
let atk_n: Int = atk_ms * rate / 1000
let dec_n: Int = dec_ms * rate / 1000
let rel_n: Int = rel_ms * rate / 1000
let np: Int = native_list_len(partials)
let half_mhz: Int = rate * 1000 / 2
let out: [Int] = native_list_empty()
let i: Int = 0
while i < total {
let acc: Int = 0
let k: Int = 0
while k < np {
let harm: Int = k + 1
let amp_k: Int = native_list_get(partials, k)
let factor: Int = 1000000
if b_micro > 0 {
let val: Int = 1000000 + b_micro * harm * harm
let factor: Int = isqrt_int(val * 1000000)
}
let fn_mhz: Int = freq_mHz * harm
let fn_mhz: Int = fn_mhz * factor / 1000000
if vib_cents > 0 {
if vib_rate > 0 {
let vphase: Int = i * vib_rate * 1024 / rate
let vs: Int = sin_lookup(table, vphase)
let vibf: Int = 1000000 + (vib_cents * vs * 833) / 10000
let fn_mhz: Int = fn_mhz * vibf / 1000000
}
}
if fn_mhz <= half_mhz {
let phase: Int = i * fn_mhz * 1024 / (rate * 1000)
let sv: Int = sin_lookup(table, phase)
let acc: Int = acc + sv * amp_k / 1000000
}
let k: Int = k + 1
}
let env: Int = adsr_env(i, total, atk_n, dec_n, sus_pm, rel_n)
let s16: Int = acc * 2800000 / sumP
let s16: Int = s16 * env / 1000
let s16: Int = s16 * amp_pm / 1000
if s16 > 32767 { let s16: Int = 32767 }
if s16 < 0 - 32767 { let s16: Int = 0 - 32767 }
let out: [Int] = native_list_append(out, s16)
let i: Int = i + 1
}
return out
}
fn synth_from_sig(lines: [String], freq_mHz: Int, dur_ms: Int, amp_pm: Int, rate: Int, table: [Int]) -> [Int] {
let partials: [Int] = parse_micros(sig_field(lines, "partials"))
let np: Int = native_list_len(partials)
let sumP: Int = 0
let j: Int = 0
while j < np {
let pj: Int = native_list_get(partials, j)
let sumP: Int = sumP + pj
let j: Int = j + 1
}
if sumP <= 0 { let sumP: Int = 1000000 }
let adsr: [String] = str_split(sig_field(lines, "adsr"), ",")
let atk_ms: Int = parse_micro(native_list_get(adsr, 0)) / 1000
let dec_ms: Int = parse_micro(native_list_get(adsr, 1)) / 1000
let sus_pm: Int = parse_micro(native_list_get(adsr, 2)) / 1000
let rel_ms: Int = parse_micro(native_list_get(adsr, 3)) / 1000
let b_micro: Int = parse_micro(sig_field(lines, "inharmonicity_B"))
let vib_rate: Int = str_to_int_el(sig_field(lines, "vibrato_rate_hz"))
let vib_cents: Int = str_to_int_el(sig_field(lines, "vibrato_depth_cents"))
return note_samples(freq_mHz, dur_ms, rate, partials, sumP, b_micro, vib_rate, vib_cents, atk_ms, dec_ms, sus_pm, rel_ms, amp_pm, table)
}
// -- byte-buffer helpers (own-core, no library) --------------------------------
fn put_tag(buf: String, pos: Int, s: String) -> String {
let n: Int = str_len(s)
let i: Int = 0
while i < n {
let buf: String = __str_set_char(buf, pos + i, str_char_code(s, i))
let i: Int = i + 1
}
return buf
}
fn put_u32le(buf: String, pos: Int, v: Int) -> String {
let buf: String = __str_set_char(buf, pos, v % 256)
let buf: String = __str_set_char(buf, pos + 1, (v / 256) % 256)
let buf: String = __str_set_char(buf, pos + 2, (v / 65536) % 256)
let buf: String = __str_set_char(buf, pos + 3, (v / 16777216) % 256)
return buf
}
fn put_u16le(buf: String, pos: Int, v: Int) -> String {
let buf: String = __str_set_char(buf, pos, v % 256)
let buf: String = __str_set_char(buf, pos + 1, (v / 256) % 256)
return buf
}
// -- WAV serializer: own-core RIFF/WAVE, PCM mono 16-bit -----------------------
fn wav_write(path: String, samples: [Int], n: Int, rate: Int) -> Int {
let data_len: Int = n * 2
let total: Int = 44 + data_len
let buf: String = __str_alloc(total)
let buf: String = put_tag(buf, 0, "RIFF")
let buf: String = put_u32le(buf, 4, 36 + data_len)
let buf: String = put_tag(buf, 8, "WAVE")
let buf: String = put_tag(buf, 12, "fmt ")
let buf: String = put_u32le(buf, 16, 16)
let buf: String = put_u16le(buf, 20, 1)
let buf: String = put_u16le(buf, 22, 1)
let buf: String = put_u32le(buf, 24, rate)
let buf: String = put_u32le(buf, 28, rate * 2)
let buf: String = put_u16le(buf, 32, 2)
let buf: String = put_u16le(buf, 34, 16)
let buf: String = put_tag(buf, 36, "data")
let buf: String = put_u32le(buf, 40, data_len)
let i: Int = 0
while i < n {
let v: Int = native_list_get(samples, i)
if v < 0 { let v: Int = v + 65536 }
let buf: String = __str_set_char(buf, 44 + i * 2, v % 256)
let buf: String = __str_set_char(buf, 44 + i * 2 + 1, (v / 256) % 256)
let i: Int = i + 1
}
let ok: Int = fs_write_bytes(path, buf, total)
return ok
}
// -- plan: frame slot-map -> note atom (pitch, duration, amplitude) ------------
fn audio_frame(relation: String, polarity: String, confidence: String, importance: String, salience: String, subj_id: String) -> [String] {
let f: [String] = native_list_empty()
let f: [String] = native_list_append(f, "relation")
let f: [String] = native_list_append(f, relation)
let f: [String] = native_list_append(f, "polarity")
let f: [String] = native_list_append(f, polarity)
let f: [String] = native_list_append(f, "confidence")
let f: [String] = native_list_append(f, confidence)
let f: [String] = native_list_append(f, "importance")
let f: [String] = native_list_append(f, importance)
let f: [String] = native_list_append(f, "salience")
let f: [String] = native_list_append(f, salience)
let f: [String] = native_list_append(f, "subj_id")
let f: [String] = native_list_append(f, subj_id)
return f
}
fn degree_offset(deg: Int) -> Int {
if deg == 0 { return 0 }
if deg == 1 { return 2 }
if deg == 2 { return 4 }
if deg == 3 { return 5 }
if deg == 4 { return 7 }
if deg == 5 { return 9 }
return 11
}
// returns [midi, dur_ms, amp_pm]
fn plan_note(frame: [String]) -> [Int] {
let relation: String = surface_get(frame, "relation")
let polarity: String = surface_get(frame, "polarity")
let confidence: String = surface_get(frame, "confidence")
let importance: String = surface_get(frame, "importance")
let salience: String = surface_get(frame, "salience")
let rn: Int = str_len(relation)
let csum: Int = 0
let i: Int = 0
while i < rn {
let cc: Int = str_char_code(relation, i)
let csum: Int = csum + cc
let i: Int = i + 1
}
let deg: Int = csum % 7
let third: Int = 4
if str_eq(polarity, "neg") { let third: Int = 3 }
let sal_oct: Int = str_to_int_el(salience)
let doff: Int = degree_offset(deg)
let midi: Int = 60 + sal_oct * 12 + doff + third
let conf_micro: Int = parse_micro(confidence)
let dur_ms: Int = 200 + conf_micro / 1000
let imp_micro: Int = parse_micro(importance)
let amp_pm: Int = 400 + imp_micro / 2000
let out: [Int] = native_list_empty()
let out: [Int] = native_list_append(out, midi)
let out: [Int] = native_list_append(out, dur_ms)
let out: [Int] = native_list_append(out, amp_pm)
return out
}
fn realize_audio(frames: [[String]], sig_lines: [String], path: String, rate: Int, table: [Int]) -> Int {
let nf: Int = native_list_len(frames)
let all: [Int] = native_list_empty()
let count: Int = 0
let fi: Int = 0
while fi < nf {
let frame: [String] = native_list_get(frames, fi)
let plan: [Int] = plan_note(frame)
let midi: Int = native_list_get(plan, 0)
let dur_ms: Int = native_list_get(plan, 1)
let amp_pm: Int = native_list_get(plan, 2)
let freq: Int = freq_of_midi(midi)
let note: [Int] = synth_from_sig(sig_lines, freq, dur_ms, amp_pm, rate, table)
let nn: Int = native_list_len(note)
let j: Int = 0
while j < nn {
let all: [Int] = native_list_append(all, native_list_get(note, j))
let j: Int = j + 1
}
let count: Int = count + nn
let fi: Int = fi + 1
}
let ok: Int = wav_write(path, all, count, rate)
return count
}
+65
View File
@@ -0,0 +1,65 @@
// image-demo.el - Drive the native PNG surface: plan a scene from a small
// meaning phrase (incl. a NEG frame) and emit a byte-valid 64x64 PNG whose
// palette is read from elp/faculty/sig/scene.basis.
fn img_frame(relation: String, polarity: String, confidence: String, importance: String, salience: String, subj_id: String) -> [String] {
let f: [String] = native_list_empty()
let f: [String] = native_list_append(f, "relation")
let f: [String] = native_list_append(f, relation)
let f: [String] = native_list_append(f, "polarity")
let f: [String] = native_list_append(f, polarity)
let f: [String] = native_list_append(f, "confidence")
let f: [String] = native_list_append(f, confidence)
let f: [String] = native_list_append(f, "importance")
let f: [String] = native_list_append(f, importance)
let f: [String] = native_list_append(f, "salience")
let f: [String] = native_list_append(f, salience)
let f: [String] = native_list_append(f, "subj_id")
let f: [String] = native_list_append(f, subj_id)
return f
}
fn rgb_str(c: [Int]) -> String {
return int_to_str(native_list_get(c, 0)) + "," + int_to_str(native_list_get(c, 1)) + "," + int_to_str(native_list_get(c, 2))
}
fn run_image() -> Int {
fs_mkdir("/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out")
let table: [Int] = crc_table()
println("crc_table[1]=" + int_to_str(native_list_get(table, 1)) + " (expect 1996959894 / 0x77073096)")
let basis: [String] = basis_load("/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/sig/scene.basis")
let warm: [Int] = parse_rgb(basis_field(basis, "warm"))
let cool: [Int] = parse_rgb(basis_field(basis, "cool"))
let bg: [Int] = parse_rgb(basis_field(basis, "bg"))
println("basis warm=" + rgb_str(warm) + " cool=" + rgb_str(cool) + " bg=" + rgb_str(bg) + " (read from scene.basis)")
let frames: [[String]] = native_list_empty()
let frames: [[String]] = native_list_append(frames, img_frame("agent", "aff", "0.9", "0.8", "0", "s1"))
let frames: [[String]] = native_list_append(frames, img_frame("theme", "aff", "0.7", "0.6", "1", "s2"))
let frames: [[String]] = native_list_append(frames, img_frame("cause", "aff", "0.8", "0.9", "0", "s3"))
let frames: [[String]] = native_list_append(frames, img_frame("negation", "neg", "0.85", "0.7", "1", "s4"))
let frames: [[String]] = native_list_append(frames, img_frame("goal", "aff", "0.6", "0.5", "0", "s5"))
let frames: [[String]] = native_list_append(frames, img_frame("result", "aff", "0.95", "1.0", "1", "s6"))
let shapes: [[Int]] = plan_scene(frames, warm, cool)
let ns: Int = native_list_len(shapes)
println("planned " + int_to_str(ns) + " shapes:")
let si: Int = 0
while si < ns {
let sh: [Int] = native_list_get(shapes, si)
let pol: String = surface_get(native_list_get(frames, si), "polarity")
println(" shape " + int_to_str(si) + " type=" + int_to_str(native_list_get(sh, 0)) + " x=" + int_to_str(native_list_get(sh, 1)) + " y=" + int_to_str(native_list_get(sh, 2)) + " size=" + int_to_str(native_list_get(sh, 3)) + " rgb=" + int_to_str(native_list_get(sh, 4)) + "," + int_to_str(native_list_get(sh, 5)) + "," + int_to_str(native_list_get(sh, 6)) + " polarity=" + pol)
let si: Int = si + 1
}
let raw: [Int] = rasterize(64, 64, shapes, bg)
println("rasterized raw (filtered scanlines) bytes=" + int_to_str(native_list_len(raw)) + " (expect 12352)")
let png: [Int] = png_build(64, 64, raw, table)
let plen: Int = native_list_len(png)
let ok: Int = png_write("/Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/scene.png", png)
println("PNG bytes=" + int_to_str(plen) + " -> /Users/will/Development/neuron-technologies/foundation/el/.claude/worktrees/agent-aaf04b0a9714c4070/elp/faculty/out/scene.png (write_ok=" + int_to_str(ok) + ")")
return plen
}
println("image-demo returned png_bytes=" + int_to_str(run_image()))
+412
View File
@@ -0,0 +1,412 @@
// image-surface.el - Native own-core raster PNG surface (the image efferent
// twin of audio). Renders a 64x64 RGB scene deterministically from a frame's
// meaning-geometry, then serialises a byte-valid PNG entirely own-core:
// 8-byte magic, IHDR, IDAT (zlib STORED/uncompressed DEFLATE + Adler32), IEND,
// with a per-chunk CRC32 computed via software xor32 (EL has no bitwise ops).
//
// The RGB palette basis is read from elp/faculty/sig/scene.basis (data, not
// literals) - the same read-from-learned discipline as the audio signatures.
// Integer-only throughout; pixels are composed functionally (painter's order)
// so no list mutation is needed.
// -- small int/parse helpers (self-contained) ----------------------------------
fn i_str_to_int(s: String) -> Int {
let n: Int = str_len(s)
let i: Int = 0
let v: Int = 0
while i < n {
let c: Int = str_char_code(s, i)
if c >= 48 {
if c < 58 {
let v: Int = v * 10 + (c - 48)
}
}
let i: Int = i + 1
}
return v
}
fn basis_load(path: String) -> [String] {
return str_split(fs_read(path), "\n")
}
fn basis_field(lines: [String], key: String) -> String {
let pref: String = key + ": "
let n: Int = native_list_len(lines)
let plen: Int = str_len(pref)
let i: Int = 0
while i < n {
let ln: String = native_list_get(lines, i)
if str_starts_with(ln, pref) {
return str_slice(ln, plen, str_len(ln))
}
let i: Int = i + 1
}
return ""
}
fn parse_rgb(csv: String) -> [Int] {
let parts: [String] = str_split(csv, ",")
let out: [Int] = native_list_empty()
let n: Int = native_list_len(parts)
let i: Int = 0
while i < n {
let v: Int = i_str_to_int(native_list_get(parts, i))
let out: [Int] = native_list_append(out, v)
let i: Int = i + 1
}
return out
}
// -- software 32-bit XOR (no bitwise ops in EL) --------------------------------
fn xor32(a: Int, b: Int) -> Int {
let r: Int = 0
let bit: Int = 1
let i: Int = 0
while i < 32 {
let abit: Int = (a / bit) % 2
let bbit: Int = (b / bit) % 2
if abit != bbit {
let add: Int = bit
let r: Int = r + add
}
let bit: Int = bit * 2
let i: Int = i + 1
}
return r
}
// -- CRC32 (table-driven, table built with xor32) ------------------------------
fn crc_table() -> [Int] {
let t: [Int] = native_list_empty()
let n: Int = 0
while n < 256 {
let c: Int = n
let k: Int = 0
while k < 8 {
if c % 2 == 1 {
let h: Int = c / 2
let c: Int = xor32(h, 3988292384)
} else {
let c: Int = c / 2
}
let k: Int = k + 1
}
let t: [Int] = native_list_append(t, c)
let n: Int = n + 1
}
return t
}
fn crc32_of(bytes: [Int], table: [Int]) -> Int {
let crc: Int = 4294967295
let n: Int = native_list_len(bytes)
let i: Int = 0
while i < n {
let b: Int = native_list_get(bytes, i)
let lo: Int = crc % 256
let idx: Int = xor32(lo, b) % 256
let tv: Int = native_list_get(table, idx)
let hi: Int = crc / 256
let crc: Int = xor32(hi, tv)
let i: Int = i + 1
}
return xor32(crc, 4294967295)
}
// -- Adler32 (for the zlib trailer) --------------------------------------------
fn adler32_of(bytes: [Int]) -> Int {
let a: Int = 1
let b: Int = 0
let n: Int = native_list_len(bytes)
let i: Int = 0
while i < n {
let byte: Int = native_list_get(bytes, i)
let a: Int = (a + byte) % 65521
let b: Int = (b + a) % 65521
let i: Int = i + 1
}
return b * 65536 + a
}
// -- byte-list append helpers --------------------------------------------------
fn app_u32be(dst: [Int], v: Int) -> [Int] {
let dst: [Int] = native_list_append(dst, (v / 16777216) % 256)
let dst: [Int] = native_list_append(dst, (v / 65536) % 256)
let dst: [Int] = native_list_append(dst, (v / 256) % 256)
let dst: [Int] = native_list_append(dst, v % 256)
return dst
}
fn app_tag(dst: [Int], s: String) -> [Int] {
let n: Int = str_len(s)
let i: Int = 0
while i < n {
let dst: [Int] = native_list_append(dst, str_char_code(s, i))
let i: Int = i + 1
}
return dst
}
fn app_all(dst: [Int], src: [Int]) -> [Int] {
let n: Int = native_list_len(src)
let i: Int = 0
while i < n {
let dst: [Int] = native_list_append(dst, native_list_get(src, i))
let i: Int = i + 1
}
return dst
}
// -- plan: frame meaning-geometry -> shape atoms -------------------------------
// shape = [type, x, y, size, r, g, b] (type 0=rect 1=disc 2=triangle)
fn charsum(s: String) -> Int {
let n: Int = str_len(s)
let i: Int = 0
let acc: Int = 0
while i < n {
let c: Int = str_char_code(s, i)
let acc: Int = acc + c
let i: Int = i + 1
}
return acc
}
fn micro_of(s: String) -> Int {
let dot: Int = str_index_of(s, ".")
if dot < 0 { return i_str_to_int(s) * 1000000 }
let n: Int = str_len(s)
let fp: String = str_slice(s, dot + 1, n)
let ip: String = str_slice(s, 0, dot)
let iv: Int = i_str_to_int(ip)
let fv: Int = 0
let scale: Int = 100000
let fl: Int = str_len(fp)
let i: Int = 0
while i < 6 {
let d: Int = 0
if i < fl { let d: Int = str_char_code(fp, i) - 48 }
let fv: Int = fv + d * scale
let scale: Int = scale / 10
let i: Int = i + 1
}
return iv * 1000000 + fv
}
fn plan_scene(frames: [[String]], warm: [Int], cool: [Int]) -> [[Int]] {
let shapes: [[Int]] = native_list_empty()
let nf: Int = native_list_len(frames)
let fi: Int = 0
while fi < nf {
let fr: [String] = native_list_get(frames, fi)
let relation: String = surface_get(fr, "relation")
let polarity: String = surface_get(fr, "polarity")
let confidence: String = surface_get(fr, "confidence")
let importance: String = surface_get(fr, "importance")
let salience: String = surface_get(fr, "salience")
// relation -> shape type
let stype: Int = charsum(relation) % 3
// confidence -> size (8..22)
let cmi: Int = micro_of(confidence)
let size: Int = 8 + cmi / 71428
// salience -> y
let sal: Int = i_str_to_int(salience)
let y: Int = 6 + sal * 26
// subj_id/index -> x
let x: Int = 4 + (fi * 10) % 48
// polarity -> warm/cool base color
let br: Int = native_list_get(warm, 0)
let bg2: Int = native_list_get(warm, 1)
let bb: Int = native_list_get(warm, 2)
if str_eq(polarity, "neg") {
let br: Int = native_list_get(cool, 0)
let bg2: Int = native_list_get(cool, 1)
let bb: Int = native_list_get(cool, 2)
}
// importance -> brightness (500..1000 permille)
let imi: Int = micro_of(importance)
let bpm: Int = 500 + imi / 2000
let r: Int = br * bpm / 1000
let g: Int = bg2 * bpm / 1000
let b: Int = bb * bpm / 1000
let sh: [Int] = native_list_empty()
let sh: [Int] = native_list_append(sh, stype)
let sh: [Int] = native_list_append(sh, x)
let sh: [Int] = native_list_append(sh, y)
let sh: [Int] = native_list_append(sh, size)
let sh: [Int] = native_list_append(sh, r)
let sh: [Int] = native_list_append(sh, g)
let sh: [Int] = native_list_append(sh, b)
let shapes: [[Int]] = native_list_append(shapes, sh)
let fi: Int = fi + 1
}
return shapes
}
// covers: is (px,py) inside this shape?
fn covers(sh: [Int], px: Int, py: Int) -> Bool {
let stype: Int = native_list_get(sh, 0)
let sx: Int = native_list_get(sh, 1)
let sy: Int = native_list_get(sh, 2)
let size: Int = native_list_get(sh, 3)
let cx: Int = sx + size / 2
if stype == 0 {
if px >= sx {
if px < sx + size {
if py >= sy {
if py < sy + size {
return true
}
}
}
}
return false
}
if stype == 1 {
let rad: Int = size / 2
let dx: Int = px - cx
let dy: Int = py - (sy + rad)
if dx * dx + dy * dy <= rad * rad {
return true
}
return false
}
// triangle: apex at top (sy), base at sy+size
if py >= sy {
if py < sy + size {
let dyv: Int = py - sy
let halfw: Int = dyv / 2
let dxv: Int = px - cx
let adx: Int = dxv
if adx < 0 { let adx: Int = 0 - dxv }
if adx <= halfw {
return true
}
}
}
return false
}
// pixel_color: painter's algorithm - last covering shape wins. Returns [r,g,b].
fn pixel_color(px: Int, py: Int, shapes: [[Int]], bg: [Int]) -> [Int] {
let r: Int = native_list_get(bg, 0)
let g: Int = native_list_get(bg, 1)
let b: Int = native_list_get(bg, 2)
let n: Int = native_list_len(shapes)
let i: Int = 0
while i < n {
let sh: [Int] = native_list_get(shapes, i)
if covers(sh, px, py) {
let r: Int = native_list_get(sh, 4)
let g: Int = native_list_get(sh, 5)
let b: Int = native_list_get(sh, 6)
}
let i: Int = i + 1
}
let out: [Int] = native_list_empty()
let out: [Int] = native_list_append(out, r)
let out: [Int] = native_list_append(out, g)
let out: [Int] = native_list_append(out, b)
return out
}
// rasterize: build the raw (filtered) scanline byte stream, filter byte 0 / row.
fn rasterize(w: Int, h: Int, shapes: [[Int]], bg: [Int]) -> [Int] {
let raw: [Int] = native_list_empty()
let y: Int = 0
while y < h {
let raw: [Int] = native_list_append(raw, 0)
let x: Int = 0
while x < w {
let col: [Int] = pixel_color(x, y, shapes, bg)
let raw: [Int] = native_list_append(raw, native_list_get(col, 0))
let raw: [Int] = native_list_append(raw, native_list_get(col, 1))
let raw: [Int] = native_list_append(raw, native_list_get(col, 2))
let x: Int = x + 1
}
let y: Int = y + 1
}
return raw
}
// zlib stream with a single STORED (uncompressed) DEFLATE block + Adler32.
fn zlib_store(raw: [Int]) -> [Int] {
let z: [Int] = native_list_empty()
let z: [Int] = native_list_append(z, 120)
let z: [Int] = native_list_append(z, 1)
let z: [Int] = native_list_append(z, 1)
let len: Int = native_list_len(raw)
let nlen: Int = 65535 - len
let z: [Int] = native_list_append(z, len % 256)
let z: [Int] = native_list_append(z, (len / 256) % 256)
let z: [Int] = native_list_append(z, nlen % 256)
let z: [Int] = native_list_append(z, (nlen / 256) % 256)
let z: [Int] = app_all(z, raw)
let ad: Int = adler32_of(raw)
let z: [Int] = app_u32be(z, ad)
return z
}
// append a full PNG chunk: length + (type+data) + crc32(type+data).
fn app_chunk(png: [Int], type_and_data: [Int], table: [Int]) -> [Int] {
let total: Int = native_list_len(type_and_data)
let dlen: Int = total - 4
let png: [Int] = app_u32be(png, dlen)
let png: [Int] = app_all(png, type_and_data)
let crc: Int = crc32_of(type_and_data, table)
let png: [Int] = app_u32be(png, crc)
return png
}
fn png_build(w: Int, h: Int, raw: [Int], table: [Int]) -> [Int] {
let png: [Int] = native_list_empty()
// 8-byte signature
let png: [Int] = native_list_append(png, 137)
let png: [Int] = native_list_append(png, 80)
let png: [Int] = native_list_append(png, 78)
let png: [Int] = native_list_append(png, 71)
let png: [Int] = native_list_append(png, 13)
let png: [Int] = native_list_append(png, 10)
let png: [Int] = native_list_append(png, 26)
let png: [Int] = native_list_append(png, 10)
// IHDR
let ihdr: [Int] = native_list_empty()
let ihdr: [Int] = app_tag(ihdr, "IHDR")
let ihdr: [Int] = app_u32be(ihdr, w)
let ihdr: [Int] = app_u32be(ihdr, h)
let ihdr: [Int] = native_list_append(ihdr, 8)
let ihdr: [Int] = native_list_append(ihdr, 2)
let ihdr: [Int] = native_list_append(ihdr, 0)
let ihdr: [Int] = native_list_append(ihdr, 0)
let ihdr: [Int] = native_list_append(ihdr, 0)
let png: [Int] = app_chunk(png, ihdr, table)
// IDAT
let z: [Int] = zlib_store(raw)
let idat: [Int] = native_list_empty()
let idat: [Int] = app_tag(idat, "IDAT")
let idat: [Int] = app_all(idat, z)
let png: [Int] = app_chunk(png, idat, table)
// IEND
let iend: [Int] = native_list_empty()
let iend: [Int] = app_tag(iend, "IEND")
let png: [Int] = app_chunk(png, iend, table)
return png
}
fn png_write(path: String, png: [Int]) -> Int {
let n: Int = native_list_len(png)
let buf: String = __str_alloc(n)
let i: Int = 0
while i < n {
let buf: String = __str_set_char(buf, i, native_list_get(png, i))
let i: Int = i + 1
}
let ok: Int = fs_write_bytes(path, buf, n)
return ok
}
+153
View File
@@ -0,0 +1,153 @@
// surface-profile.el - Surface profile data and accessors.
//
// THE NATIVE EFFERENT SEAM: surface = a pluggable PROFILE, using the exact same
// slot-map mechanism as language-profile.el. A language profile tells the
// realizer HOW to shape a natural-language surface (word order, morphology); a
// SURFACE profile tells the realizer WHICH surface to project meaning onto
// (markdown, docx, html, plain, or a non-text medium like symbolic music).
//
// The generalization is exact: realize_lang(form, profile) already renders a
// SemForm parameterized by a [String] profile read via lang_get. Surface is one
// more axis of that same profile vector. One frame (sem_frame), one plan step
// (sem_to_spec), one render (realize) the surface is DATA, not a code path,
// precisely as language is data. Adding a surface means adding a profile, no
// engine change. This is the multimodal projector, native: geometry -> any
// surface, the efferent twin of ingest.
//
// Surface slot keys:
// surface - "markdown" | "docx" | "html" | "plain" | "midi" | "image"
// modality - "text" | "audio" | "image" | "video"
// media_type - MIME type of the emitted surface
// head_open - string prepended to a heading (e.g. "## " for markdown)
// head_close - string appended to a heading (e.g. "" for markdown, "</h2>" for html)
// emph_open - string opening emphasis (e.g. "*")
// emph_close - string closing emphasis (e.g. "*")
// item_mark - list-item marker (e.g. "- ")
// para_sep - paragraph separator (e.g. "\n\n")
//
// For a TEXT modality the render composes these markers around the surface that
// the EXISTING realizer produces (realize_lang / sem_realize). For a non-text
// modality (audio/image) the profile declares modality + media_type and the
// render dispatches to the medium projector, which reads the SAME frame's
// geometry (its intent/affect/structure) and projects it onto sound or pixels
// deterministic-from-meaning, nothing invented. That dispatch point is where a
// music profile or image profile conforms, native, no parallel layer.
// -- Constructor -------------------------------------------------------------
fn surface_profile(surface: String, modality: String, media_type: String, head_open: String, head_close: String, emph_open: String, emph_close: String, item_mark: String, para_sep: String) -> [String] {
let r: [String] = native_list_empty()
let r = native_list_append(r, "surface")
let r = native_list_append(r, surface)
let r = native_list_append(r, "modality")
let r = native_list_append(r, modality)
let r = native_list_append(r, "media_type")
let r = native_list_append(r, media_type)
let r = native_list_append(r, "head_open")
let r = native_list_append(r, head_open)
let r = native_list_append(r, "head_close")
let r = native_list_append(r, head_close)
let r = native_list_append(r, "emph_open")
let r = native_list_append(r, emph_open)
let r = native_list_append(r, "emph_close")
let r = native_list_append(r, emph_close)
let r = native_list_append(r, "item_mark")
let r = native_list_append(r, item_mark)
let r = native_list_append(r, "para_sep")
let r = native_list_append(r, para_sep)
return r
}
// -- Accessor (same convention as lang_get; standalone so this is a leaf) -----
fn surface_get(profile: [String], key: String) -> String {
let n: Int = native_list_len(profile)
let i: Int = 0
while i < n - 1 {
let k: String = native_list_get(profile, i)
if str_eq(k, key) {
return native_list_get(profile, i + 1)
}
let i = i + 2
}
return ""
}
fn surface_is_text(profile: [String]) -> Bool {
return str_eq(surface_get(profile, "modality"), "text")
}
// -- Built-in TEXT surface profiles ------------------------------------------
// Markdown: headings with "## ", emphasis with "*", "- " list items.
fn surface_profile_markdown() -> [String] {
return surface_profile("markdown", "text", "text/markdown", "## ", "", "*", "*", "- ", "\n\n")
}
// Plain text: no markup at all headings become bare uppercase-free lines.
fn surface_profile_plain() -> [String] {
return surface_profile("plain", "text", "text/plain", "", "", "", "", " - ", "\n\n")
}
// HTML: block-level heading/emphasis tags.
fn surface_profile_html() -> [String] {
return surface_profile("html", "text", "text/html", "<h2>", "</h2>", "<em>", "</em>", "<li>", "\n")
}
// docx: WordprocessingML is structural, not inline-markup; the head/emph slots
// carry the run/style intent that the OOXML emitter maps to <w:pStyle>. Declared
// here so docx is a first-class surface on the same seam.
fn surface_profile_docx() -> [String] {
return surface_profile("docx", "text", "application/vnd.openxmlformats-officedocument.wordprocessingml.document", "Heading2:", "", "b:", "", "bullet:", "\n")
}
// -- Built-in NON-TEXT surface profiles (the multimodal seam) ----------------
// Symbolic music (MIDI): modality=audio. The render dispatches to the music
// projector, which reads the SAME frame's intent/affect and projects it to
// pitch/rhythm deterministic-from-meaning. head/emph slots are empty because
// the medium is not textual; media_type names the surface. A music profile
// (scale/mode/instrument) is layered onto this by the audio agent, native.
fn surface_profile_midi() -> [String] {
return surface_profile("midi", "audio", "audio/midi", "", "", "", "", "", "")
}
// Synthesized audio (WAV): modality=audio, peer to midi. The richer audio
// surface the render SUPERPOSES ingested tonal primitives (sine at f0*n per an
// ingested instrument signature) into PCM, own-core, exactly as midi writes an
// SMF via struct. A music profile (scale/mode/instrument/adsr) layers onto this
// as its own [String] slot-map read by the same getter. Same frame -> midi OR
// audio, interchangeable; this is the audio agent's native conforming point.
fn surface_profile_audio() -> [String] {
return surface_profile("audio", "audio", "audio/wav", "", "", "", "", "", "")
}
// Image (raster): modality=image. Documented seam the render dispatches to the
// image projector, the efferent twin of image ingest, reading the same frame.
fn surface_profile_image() -> [String] {
return surface_profile("image", "image", "image/png", "", "", "", "", "", "")
}
// -- Composition helpers: wrap realized TEXT with the surface's markers -------
//
// These take text the EXISTING realizer already produced and shape it for the
// surface. They add NO content pure surface typography over faithful text,
// exactly as the language profile adds no content, only linguistic form.
fn surface_heading(profile: [String], text: String) -> String {
let o: String = surface_get(profile, "head_open")
let c: String = surface_get(profile, "head_close")
return o + text + c
}
fn surface_emph(profile: [String], text: String) -> String {
let o: String = surface_get(profile, "emph_open")
let c: String = surface_get(profile, "emph_close")
return o + text + c
}
// A section: a heading + a paragraph separator + the (already realized) body.
fn surface_section(profile: [String], heading: String, body: String) -> String {
let sep: String = surface_get(profile, "para_sep")
return surface_heading(profile, heading) + sep + body
}
@@ -0,0 +1,26 @@
// surface-profile-demo.el - ONE SemFrame, realized ONCE, projected to THREE
// surfaces via surface profiles. Proves surface-as-profile natively: the frame
// and the realized sentence are identical; only the surface PROFILE differs.
fn demo() -> String {
// 1. The shared frame (meaning-geometry): assert(Neuron, contain, the memory).
let frame: [String] = sem_frame("assert", "Neuron", "the memory", "")
// 2. REALIZE once via the EXISTING native realizer (language = a profile).
let sentence: String = sem_realize(frame)
// 3. PROJECT the same realized sentence onto three surfaces (surface = a
// profile). Same frame, same sentence, different surface one render.
let heading: String = "Memory"
let md: String = surface_section(surface_profile_markdown(), heading, sentence)
let html: String = surface_section(surface_profile_html(), heading, sentence)
let plain: String = surface_section(surface_profile_plain(), heading, sentence)
// 4. Report the non-text seam: a surface profile can declare an audio/image
// medium; the render dispatches to the medium projector on the SAME frame.
let midi_media: String = surface_get(surface_profile_midi(), "media_type")
return "MD=[" + md + "] HTML=[" + html + "] PLAIN=[" + plain + "] MIDI_MEDIA=" + midi_media
}
println(demo())
+1 -1
View File
@@ -81,7 +81,7 @@ jobs:
# Link to produce the engram binary
- name: Link engram binary
run: |
cc -std=c11 -O2 \
cc -std=c11 -O2 -DHAVE_CURL \
-I /usr/local/lib/el \
-o dist/engram \
dist/engram.c \
+1 -1
View File
@@ -88,7 +88,7 @@ jobs:
# Link to produce the engram binary
- name: Link engram binary
run: |
cc -std=c11 -O2 \
cc -std=c11 -O2 -DHAVE_CURL \
-I /usr/local/lib/el \
-o dist/engram \
dist/engram.c \
+1 -1
View File
@@ -62,7 +62,7 @@ jobs:
# Link to produce the engram binary
- name: Link engram binary
run: |
cc -std=c11 -O2 \
cc -std=c11 -O2 -DHAVE_CURL \
-I /usr/local/lib/el \
-o dist/engram \
dist/engram.c \
BIN
View File
Binary file not shown.
+142 -95
View File
@@ -10,6 +10,7 @@ el_val_t query_param(el_val_t path, el_val_t key);
el_val_t query_int(el_val_t path, el_val_t key, el_val_t default_val);
el_val_t extract_id(el_val_t path, el_val_t prefix);
el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body);
el_val_t persist_canonical(void);
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_get_node(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body);
@@ -20,18 +21,23 @@ el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_neighbors(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body);
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body);
el_val_t check_auth_ok(el_val_t method, el_val_t body);
el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body);
el_val_t bind_raw;
el_val_t bind_str;
el_val_t port;
el_val_t data_dir_raw;
el_val_t data_dir;
el_val_t snapshot_path;
el_val_t boot_snap;
el_val_t parse_port(el_val_t bind) {
el_val_t colon = str_index_of(bind, EL_STR(":"));
@@ -110,17 +116,22 @@ el_val_t route_stats(el_val_t method, el_val_t path, el_val_t body) {
return 0;
}
el_val_t persist_canonical(void) {
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_1 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_1 = (EL_STR("/tmp/engram")); } else { _if_result_1 = (dir_raw); } _if_result_1; });
engram_save(el_str_concat(dir, EL_STR("/snapshot.json")));
return 1;
return 0;
}
el_val_t route_create_node(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
el_val_t node_type = json_get_string(body, EL_STR("node_type"));
if (str_eq(node_type, EL_STR(""))) {
node_type = EL_STR("Memory");
}
el_val_t salience = json_get_float(body, EL_STR("salience"));
if (salience == el_from_float(0.0)) {
salience = el_from_float(0.5);
}
el_val_t nt_raw = json_get_string(body, EL_STR("node_type"));
el_val_t node_type = ({ el_val_t _if_result_2 = 0; if (str_eq(nt_raw, EL_STR(""))) { _if_result_2 = (EL_STR("Memory")); } else { _if_result_2 = (nt_raw); } _if_result_2; });
el_val_t sal_raw = json_get_float(body, EL_STR("salience"));
el_val_t salience = ({ el_val_t _if_result_3 = 0; if ((sal_raw == el_from_float(0.0))) { _if_result_3 = (el_from_float(0.5)); } else { _if_result_3 = (sal_raw); } _if_result_3; });
el_val_t id = engram_node(content, node_type, salience);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"id\":\""), id), EL_STR("\",\"content\":\"")), content), EL_STR("\",\"node_type\":\"")), node_type), EL_STR("\"}"));
return 0;
}
@@ -146,11 +157,9 @@ el_val_t route_scan_nodes(el_val_t method, el_val_t path, el_val_t body) {
}
el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
el_val_t snap_path = el_str_concat(dir, EL_STR("/snapshot.json"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_4 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_4 = (EL_STR("/tmp/engram")); } else { _if_result_4 = (dir_raw); } _if_result_4; });
el_val_t snap_path = el_str_concat(dir, EL_STR("/.scan-export.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
if (str_eq(snap, EL_STR(""))) {
@@ -165,36 +174,22 @@ el_val_t route_scan_edges(el_val_t method, el_val_t path, el_val_t body) {
}
el_val_t route_search(el_val_t method, el_val_t path, el_val_t body) {
el_val_t q = EL_STR("");
if (str_eq(method, EL_STR("GET"))) {
q = query_param(path, EL_STR("q"));
} else {
q = json_get_string(body, EL_STR("query"));
}
el_val_t limit = query_int(path, EL_STR("limit"), 20);
if (limit == 0) {
limit = json_get_int(body, EL_STR("limit"));
}
if (limit == 0) {
limit = 20;
}
el_val_t q = ({ el_val_t _if_result_5 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_5 = (query_param(path, EL_STR("q"))); } else { _if_result_5 = (json_get_string(body, EL_STR("query"))); } _if_result_5; });
el_val_t lim_url = query_int(path, EL_STR("limit"), 0);
el_val_t lim_body = json_get_int(body, EL_STR("limit"));
el_val_t lim_either = ({ el_val_t _if_result_6 = 0; if ((lim_url > 0)) { _if_result_6 = (lim_url); } else { _if_result_6 = (lim_body); } _if_result_6; });
el_val_t limit = ({ el_val_t _if_result_7 = 0; if ((lim_either > 0)) { _if_result_7 = (lim_either); } else { _if_result_7 = (20); } _if_result_7; });
return engram_search_json(q, limit);
return 0;
}
el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
el_val_t q = EL_STR("");
el_val_t depth = 3;
if (str_eq(method, EL_STR("GET"))) {
q = query_param(path, EL_STR("q"));
depth = query_int(path, EL_STR("depth"), 3);
} else {
q = json_get_string(body, EL_STR("query"));
el_val_t bd = json_get_int(body, EL_STR("depth"));
if (bd > 0) {
depth = bd;
}
el_val_t q = ({ el_val_t _if_result_8 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_8 = (query_param(path, EL_STR("q"))); } else { _if_result_8 = (json_get_string(body, EL_STR("query"))); } _if_result_8; });
if (str_eq(q, EL_STR(""))) {
return err_json(EL_STR("missing query"));
}
el_val_t d_raw = ({ el_val_t _if_result_9 = 0; if (str_eq(method, EL_STR("GET"))) { _if_result_9 = (query_int(path, EL_STR("depth"), 3)); } else { _if_result_9 = (json_get_int(body, EL_STR("depth"))); } _if_result_9; });
el_val_t depth = ({ el_val_t _if_result_10 = 0; if ((d_raw > 0)) { _if_result_10 = (d_raw); } else { _if_result_10 = (3); } _if_result_10; });
return el_str_concat(el_str_concat(EL_STR("{\"results\":"), engram_activate_json(q, depth)), EL_STR("}"));
return 0;
}
@@ -202,15 +197,12 @@ el_val_t route_activate(el_val_t method, el_val_t path, el_val_t body) {
el_val_t route_create_edge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t from_id = json_get_string(body, EL_STR("from_id"));
el_val_t to_id = json_get_string(body, EL_STR("to_id"));
el_val_t relation = json_get_string(body, EL_STR("relation"));
if (str_eq(relation, EL_STR(""))) {
relation = EL_STR("associates");
}
el_val_t weight = json_get_float(body, EL_STR("weight"));
if (weight == el_from_float(0.0)) {
weight = el_from_float(0.5);
}
el_val_t rel_raw = json_get_string(body, EL_STR("relation"));
el_val_t relation = ({ el_val_t _if_result_11 = 0; if (str_eq(rel_raw, EL_STR(""))) { _if_result_11 = (EL_STR("associates")); } else { _if_result_11 = (rel_raw); } _if_result_11; });
el_val_t w_raw = json_get_float(body, EL_STR("weight"));
el_val_t weight = ({ el_val_t _if_result_12 = 0; if ((w_raw == el_from_float(0.0))) { _if_result_12 = (el_from_float(0.5)); } else { _if_result_12 = (w_raw); } _if_result_12; });
engram_connect(from_id, to_id, weight, relation);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"from_id\":\""), from_id), EL_STR("\",\"to_id\":\"")), to_id), EL_STR("\",\"relation\":\"")), relation), EL_STR("\"}"));
return 0;
}
@@ -231,6 +223,7 @@ el_val_t route_strengthen(el_val_t method, el_val_t path, el_val_t body) {
return err_json(EL_STR("missing node_id"));
}
engram_strengthen(id);
el_val_t saved = persist_canonical();
return ok_json();
return 0;
}
@@ -241,29 +234,40 @@ el_val_t route_forget(el_val_t method, el_val_t path, el_val_t body) {
return err_json(EL_STR("missing id"));
}
engram_forget(id);
el_val_t saved = persist_canonical();
return ok_json();
return 0;
}
el_val_t route_create_ise(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
if (str_eq(content, EL_STR(""))) {
return err_json(EL_STR("missing content"));
}
el_val_t sal = el_from_float(0.3);
el_val_t imp = el_from_float(0.3);
el_val_t conf = el_from_float(0.8);
el_val_t id = engram_node_full(content, EL_STR("InternalStateEvent"), EL_STR("state-event"), sal, imp, conf, EL_STR("Episodic"), EL_STR("[\"internal-state\",\"InternalStateEvent\"]"));
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\"}"));
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p_raw = json_get_string(body, EL_STR("path"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_13 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_13 = (EL_STR("/tmp/engram")); } else { _if_result_13 = (dir_raw); } _if_result_13; });
el_val_t p = ({ el_val_t _if_result_14 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_14 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_14 = (p_raw); } _if_result_14; });
engram_save(p);
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"path\":\""), p), EL_STR("\"}"));
return 0;
}
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p_raw = json_get_string(body, EL_STR("path"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_15 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_15 = (EL_STR("/tmp/engram")); } else { _if_result_15 = (dir_raw); } _if_result_15; });
el_val_t p = ({ el_val_t _if_result_16 = 0; if (str_eq(p_raw, EL_STR(""))) { _if_result_16 = (el_str_concat(dir, EL_STR("/snapshot.json"))); } else { _if_result_16 = (p_raw); } _if_result_16; });
engram_load(p);
return ok_json();
return 0;
}
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
return EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}");
return 0;
}
el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
el_val_t snap_path = el_str_concat(dir, EL_STR("/sync-export.json"));
el_val_t dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
el_val_t dir = ({ el_val_t _if_result_17 = 0; if (str_eq(dir_raw, EL_STR(""))) { _if_result_17 = (EL_STR("/tmp/engram")); } else { _if_result_17 = (dir_raw); } _if_result_17; });
el_val_t snap_path = el_str_concat(dir, EL_STR("/.sync-export.json"));
engram_save(snap_path);
el_val_t snap = fs_read(snap_path);
if (str_eq(snap, EL_STR(""))) {
@@ -273,36 +277,68 @@ el_val_t route_sync(el_val_t method, el_val_t path, el_val_t body) {
return 0;
}
el_val_t route_save(el_val_t method, el_val_t path, el_val_t body) {
el_val_t route_load_merge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p = json_get_string(body, EL_STR("path"));
if (str_eq(p, EL_STR(""))) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
p = el_str_concat(dir, EL_STR("/snapshot.json"));
return err_json(EL_STR("path is required"));
}
engram_save(p);
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"path\":\""), p), EL_STR("\"}"));
if (str_eq(fs_read(p), EL_STR(""))) {
return err_json(EL_STR("file missing or empty"));
}
el_val_t before_n = engram_node_count();
el_val_t before_e = engram_edge_count();
engram_load_merge(p);
el_val_t added_n = (engram_node_count() - before_n);
el_val_t added_e = (engram_edge_count() - before_e);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"nodes_added\":"), int_to_str(added_n)), EL_STR(",\"edges_added\":")), int_to_str(added_e)), EL_STR(",\"node_count\":")), int_to_str(engram_node_count())), EL_STR("}"));
return 0;
}
el_val_t route_load(el_val_t method, el_val_t path, el_val_t body) {
el_val_t p = json_get_string(body, EL_STR("path"));
if (str_eq(p, EL_STR(""))) {
el_val_t dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(dir, EL_STR(""))) {
dir = EL_STR("/tmp/engram");
}
p = el_str_concat(dir, EL_STR("/snapshot.json"));
el_val_t route_emit_ise(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
if (str_eq(content, EL_STR(""))) {
return err_json(EL_STR("missing content"));
}
engram_load(p);
return ok_json();
el_val_t sal = el_from_float(0.3);
el_val_t imp = el_from_float(0.3);
el_val_t conf = el_from_float(0.8);
el_val_t id = engram_node_full(content, EL_STR("InternalStateEvent"), EL_STR("state-event"), sal, imp, conf, EL_STR("Episodic"), EL_STR("[\"internal-state\",\"InternalStateEvent\"]"));
el_val_t ret_raw = env(EL_STR("ENGRAM_ISE_RETENTION_MS"));
el_val_t ret_ms = ({ el_val_t _if_result_18 = 0; if (str_eq(ret_raw, EL_STR(""))) { _if_result_18 = (172800000); } else { _if_result_18 = (str_to_int(ret_raw)); } _if_result_18; });
el_val_t pruned = engram_prune_telemetry(ret_ms);
return el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\",\"pruned\":")), int_to_str(pruned)), EL_STR("}"));
return 0;
}
el_val_t route_health(el_val_t method, el_val_t path, el_val_t body) {
return EL_STR("{\"status\":\"ok\",\"engine\":\"engram-runtime-native\"}");
el_val_t route_capture_knowledge(el_val_t method, el_val_t path, el_val_t body) {
el_val_t content = json_get_string(body, EL_STR("content"));
if (str_eq(content, EL_STR(""))) {
return err_json(EL_STR("missing content"));
}
el_val_t title = json_get_string(body, EL_STR("title"));
el_val_t label = ({ el_val_t _if_result_19 = 0; if (str_eq(title, EL_STR(""))) { _if_result_19 = (str_slice(content, 0, 60)); } else { _if_result_19 = (title); } _if_result_19; });
el_val_t category_raw = json_get_string(body, EL_STR("category"));
el_val_t category = ({ el_val_t _if_result_20 = 0; if (str_eq(category_raw, EL_STR(""))) { _if_result_20 = (EL_STR("other")); } else { _if_result_20 = (category_raw); } _if_result_20; });
el_val_t ktier_raw = json_get_string(body, EL_STR("tier"));
el_val_t ktier = ({ el_val_t _if_result_21 = 0; if (str_eq(ktier_raw, EL_STR(""))) { _if_result_21 = (EL_STR("note")); } else { _if_result_21 = (ktier_raw); } _if_result_21; });
el_val_t project = json_get_string(body, EL_STR("project"));
el_val_t tags_raw = json_get_raw(body, EL_STR("tags"));
el_val_t tags_base = ({ el_val_t _if_result_22 = 0; if (str_eq(tags_raw, EL_STR(""))) { _if_result_22 = (EL_STR("[]")); } else { _if_result_22 = (tags_raw); } _if_result_22; });
el_val_t base_len = str_len(tags_base);
el_val_t head = str_slice(tags_base, 0, (base_len - 1));
el_val_t sep = ({ el_val_t _if_result_23 = 0; if (str_eq(head, EL_STR("["))) { _if_result_23 = (EL_STR("")); } else { _if_result_23 = (EL_STR(",")); } _if_result_23; });
el_val_t safe_cat = str_replace(category, EL_STR("\""), EL_STR("'"));
el_val_t safe_tier = str_replace(ktier, EL_STR("\""), EL_STR("'"));
el_val_t safe_proj = str_replace(project, EL_STR("\""), EL_STR("'"));
el_val_t proj_tag = ({ el_val_t _if_result_24 = 0; if (str_eq(safe_proj, EL_STR(""))) { _if_result_24 = (EL_STR("")); } else { _if_result_24 = (el_str_concat(el_str_concat(EL_STR(",\"project:"), safe_proj), EL_STR("\""))); } _if_result_24; });
el_val_t tags = el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(head, sep), EL_STR("\"category:")), safe_cat), EL_STR("\",\"tier:")), safe_tier), EL_STR("\"")), proj_tag), EL_STR("]"));
el_val_t sal = el_from_float(0.5);
el_val_t imp = el_from_float(0.5);
el_val_t conf = el_from_float(0.9);
el_val_t id = engram_node_full(content, EL_STR("Knowledge"), label, sal, imp, conf, EL_STR("Semantic"), tags);
el_val_t saved = persist_canonical();
return el_str_concat(el_str_concat(EL_STR("{\"ok\":true,\"id\":\""), id), EL_STR("\"}"));
return 0;
}
@@ -329,12 +365,15 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
return route_health(method, path, body);
}
}
if (str_eq(method, EL_STR("POST")) && str_starts_with(clean, EL_STR("/api/neuron/state-events"))) {
return route_create_ise(method, path, body);
if (str_eq(method, EL_STR("POST")) && str_eq(clean, EL_STR("/api/neuron/state-events"))) {
return route_emit_ise(method, path, body);
}
if (!check_auth_ok(method, body)) {
return err_json(EL_STR("unauthorized"));
}
if (str_eq(method, EL_STR("POST")) && str_eq(clean, EL_STR("/api/neuron/knowledge/capture"))) {
return route_capture_knowledge(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/stats")) || str_eq(clean, EL_STR("/stats")))) {
return route_stats(method, path, body);
}
@@ -374,32 +413,40 @@ el_val_t handle_request(el_val_t method, el_val_t path, el_val_t body) {
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/strengthen")) || str_eq(clean, EL_STR("/strengthen")))) {
return route_strengthen(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && (str_eq(clean, EL_STR("/api/sync")) || str_eq(clean, EL_STR("/sync")))) {
return route_sync(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/save")) || str_eq(clean, EL_STR("/save")))) {
return route_save(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/load")) || str_eq(clean, EL_STR("/load")))) {
return route_load(method, path, body);
}
if (str_eq(method, EL_STR("POST")) && (str_eq(clean, EL_STR("/api/load-merge")) || str_eq(clean, EL_STR("/load-merge")))) {
return route_load_merge(method, path, body);
}
if (str_eq(method, EL_STR("GET")) && str_eq(clean, EL_STR("/api/sync"))) {
return route_sync(method, path, body);
}
return el_str_concat(el_str_concat(EL_STR("{\"error\":\"not found\",\"path\":\""), clean), EL_STR("\"}"));
return 0;
}
int main(int _argc, char** _argv) {
el_runtime_init_args(_argc, _argv);
bind_str = env(EL_STR("ENGRAM_BIND"));
if (str_eq(bind_str, EL_STR(""))) {
bind_str = EL_STR(":8742");
}
bind_raw = env(EL_STR("ENGRAM_BIND"));
bind_str = ({ el_val_t _if_result_25 = 0; if (str_eq(bind_raw, EL_STR(""))) { _if_result_25 = (EL_STR(":8742")); } else { _if_result_25 = (bind_raw); } _if_result_25; });
port = parse_port(bind_str);
data_dir = env(EL_STR("ENGRAM_DATA_DIR"));
if (str_eq(data_dir, EL_STR(""))) {
data_dir = EL_STR("/tmp/engram");
}
data_dir_raw = env(EL_STR("ENGRAM_DATA_DIR"));
data_dir = ({ el_val_t _if_result_26 = 0; if (str_eq(data_dir_raw, EL_STR(""))) { _if_result_26 = (EL_STR("/tmp/engram")); } else { _if_result_26 = (data_dir_raw); } _if_result_26; });
snapshot_path = el_str_concat(data_dir, EL_STR("/snapshot.json"));
engram_load(snapshot_path);
boot_snap = fs_read(snapshot_path);
if (!str_eq(boot_snap, EL_STR(""))) {
if (engram_node_count() == 0) {
println(EL_STR("[engram] WARNING: snapshot.json is non-empty but load produced 0 nodes \xe2\x80\x94 preserving copy at snapshot.failed-load.json"));
fs_write(el_str_concat(data_dir, EL_STR("/snapshot.failed-load.json")), boot_snap);
} else {
fs_write(el_str_concat(data_dir, EL_STR("/snapshot.boot-backup.json")), boot_snap);
}
}
println(EL_STR("[engram] runtime-native graph engine"));
println(el_str_concat(EL_STR("[engram] data_dir="), data_dir));
println(el_str_concat(EL_STR("[engram] node_count="), int_to_str(engram_node_count())));
+190 -54
View File
@@ -76,13 +76,43 @@ fn route_stats(method: String, path: String, body: String) -> String {
engram_stats_json()
}
// (2026-07-18 self-review) Scoping sweep: `let` inside an if-block creates an
// inner scope only it does NOT mutate the outer binding (documented with
// evidence in awareness.el, 2026-05-25). Every default/reassignment below used
// that broken pattern, so defaults never applied: nodes were created with
// node_type="" and salience=0.0, /api/search and /api/activate ALWAYS ran with
// q="" regardless of input, edges defaulted to relation=""/weight=0.0, and
// save/load with no "path" hit engram_save(""). Rewritten to the
// `let x = if cond { a } else { b }` expression form (the pattern the newer
// routes route_emit_ise/route_capture_knowledge already use correctly).
// persist_canonical save the canonical snapshot after a durable write.
//
// WHY (2026-07-22 self-review): the 2026-07-21 fix correctly stopped READ
// routes from writing the canonical snapshot.json but nothing was left
// that saved it on WRITE. Every mutation (node create, edge create,
// knowledge capture, forget, merge) lived only in RAM until someone POSTed
// /api/save manually; a process restart silently discarded everything since
// the last manual save. Observed live: two engram restarts during the
// 2026-07-22 review reverted the store to a ~17h-old snapshot, destroying
// same-day writes. Reads must never write the canonical; writes must always
// persist it. ISE telemetry is deliberately excluded (48h-pruned, loss-
// tolerant, ~2/min snapshotting the whole store per heartbeat is waste;
// any durable write that follows persists the pruning too).
fn persist_canonical() -> Int {
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
engram_save(dir + "/snapshot.json")
return 1
}
fn route_create_node(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
let node_type: String = json_get_string(body, "node_type")
if str_eq(node_type, "") { let node_type = "Memory" }
let salience: Float = json_get_float(body, "salience")
if salience == 0.0 { let salience = 0.5 }
let nt_raw: String = json_get_string(body, "node_type")
let node_type: String = if str_eq(nt_raw, "") { "Memory" } else { nt_raw }
let sal_raw: Float = json_get_float(body, "salience")
let salience: Float = if sal_raw == 0.0 { 0.5 } else { sal_raw }
let id: String = engram_node(content, node_type, salience)
let saved: Int = persist_canonical()
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\"}"
}
@@ -103,13 +133,14 @@ fn route_scan_nodes(method: String, path: String, body: String) -> String {
}
// route_scan_edges bulk export of all edges as a JSON array. Implemented
// via engram_save fs_read of the canonical on-disk snapshot, which the
// runtime keeps in lockstep with the in-memory graph. Live against the
// running graph, not a stale export.
// via engram_save fs_read of a SCRATCH export path. (2026-07-21 self-review:
// previously this saved over the canonical snapshot.json on every GET if the
// process ever booted with a partial/empty store, the first read request
// clobbered the good snapshot. Read routes must never write the canonical path.)
fn route_scan_edges(method: String, path: String, body: String) -> String {
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let snap_path: String = dir + "/snapshot.json"
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let snap_path: String = dir + "/.scan-export.json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
if str_eq(snap, "") { return "[]" }
@@ -122,40 +153,34 @@ fn route_scan_edges(method: String, path: String, body: String) -> String {
}
fn route_search(method: String, path: String, body: String) -> String {
let q: String = ""
if str_eq(method, "GET") {
let q = query_param(path, "q")
} else {
let q = json_get_string(body, "query")
}
let limit: Int = query_int(path, "limit", 20)
if limit == 0 { let limit = json_get_int(body, "limit") }
if limit == 0 { let limit = 20 }
let q: String = if str_eq(method, "GET") { query_param(path, "q") } else { json_get_string(body, "query") }
let lim_url: Int = query_int(path, "limit", 0)
let lim_body: Int = json_get_int(body, "limit")
let lim_either: Int = if lim_url > 0 { lim_url } else { lim_body }
let limit: Int = if lim_either > 0 { lim_either } else { 20 }
return engram_search_json(q, limit)
}
fn route_activate(method: String, path: String, body: String) -> String {
let q: String = ""
let depth: Int = 3
if str_eq(method, "GET") {
let q = query_param(path, "q")
let depth = query_int(path, "depth", 3)
} else {
let q = json_get_string(body, "query")
let bd: Int = json_get_int(body, "depth")
if bd > 0 { let depth = bd }
}
let q: String = if str_eq(method, "GET") { query_param(path, "q") } else { json_get_string(body, "query") }
// Guard: engram_activate with an empty query matches zero seeds, which
// zeroes ALL carried working-memory weights (documented in awareness.el
// perceive()). Never let an empty activation through to wipe WM.
if str_eq(q, "") { return err_json("missing query") }
let d_raw: Int = if str_eq(method, "GET") { query_int(path, "depth", 3) } else { json_get_int(body, "depth") }
let depth: Int = if d_raw > 0 { d_raw } else { 3 }
return "{\"results\":" + engram_activate_json(q, depth) + "}"
}
fn route_create_edge(method: String, path: String, body: String) -> String {
let from_id: String = json_get_string(body, "from_id")
let to_id: String = json_get_string(body, "to_id")
let relation: String = json_get_string(body, "relation")
if str_eq(relation, "") { let relation = "associates" }
let weight: Float = json_get_float(body, "weight")
if weight == 0.0 { let weight = 0.5 }
let rel_raw: String = json_get_string(body, "relation")
let relation: String = if str_eq(rel_raw, "") { "associates" } else { rel_raw }
let w_raw: Float = json_get_float(body, "weight")
let weight: Float = if w_raw == 0.0 { 0.5 } else { w_raw }
engram_connect(from_id, to_id, weight, relation)
let saved: Int = persist_canonical()
"{\"ok\":true,\"from_id\":\"" + from_id + "\",\"to_id\":\"" + to_id + "\",\"relation\":\"" + relation + "\"}"
}
@@ -170,6 +195,7 @@ fn route_strengthen(method: String, path: String, body: String) -> String {
let id: String = json_get_string(body, "node_id")
if str_eq(id, "") { return err_json("missing node_id") }
engram_strengthen(id)
let saved: Int = persist_canonical()
ok_json()
}
@@ -177,27 +203,24 @@ fn route_forget(method: String, path: String, body: String) -> String {
let id: String = extract_id(path, "/api/nodes/")
if str_eq(id, "") { return err_json("missing id") }
engram_forget(id)
let saved: Int = persist_canonical()
ok_json()
}
fn route_save(method: String, path: String, body: String) -> String {
let p: String = json_get_string(body, "path")
if str_eq(p, "") {
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let p = dir + "/snapshot.json"
}
let p_raw: String = json_get_string(body, "path")
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
engram_save(p)
"{\"ok\":true,\"path\":\"" + p + "\"}"
}
fn route_load(method: String, path: String, body: String) -> String {
let p: String = json_get_string(body, "path")
if str_eq(p, "") {
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let p = dir + "/snapshot.json"
}
let p_raw: String = json_get_string(body, "path")
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
engram_load(p)
ok_json()
}
@@ -219,15 +242,36 @@ fn route_health(method: String, path: String, body: String) -> String {
// (it skips nodes already present by ID). Auth-exempt: same-host internal call.
// (2026-06-27 self-review: added this route to fix silent 10-min sync failures)
fn route_sync(method: String, path: String, body: String) -> String {
let dir: String = env("ENGRAM_DATA_DIR")
if str_eq(dir, "") { let dir = "/tmp/engram" }
let snap_path: String = dir + "/snapshot.json"
let dir_raw: String = env("ENGRAM_DATA_DIR")
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
// 2026-07-21 self-review: export to a scratch path, never the canonical
// snapshot.json read routes must not be able to clobber the good snapshot.
let snap_path: String = dir + "/.sync-export.json"
engram_save(snap_path)
let snap: String = fs_read(snap_path)
if str_eq(snap, "") { return "{\"nodes\":[],\"edges\":[]}" }
return snap
}
// route_load_merge POST /api/load-merge {"path": "..."} merge a snapshot
// file into the live store WITHOUT resetting it (engram_load_merge skips nodes
// already present by id). Added 2026-07-21 self-review to restore the 244 kn-
// identity Knowledge nodes lost from the snapshot lineage between 05-13 and
// 07-13. Requires an explicit path: refuses to run without one so it can never
// be triggered accidentally against a default.
fn route_load_merge(method: String, path: String, body: String) -> String {
let p: String = json_get_string(body, "path")
if str_eq(p, "") { return err_json("path is required") }
if str_eq(fs_read(p), "") { return err_json("file missing or empty") }
let before_n: Int = engram_node_count()
let before_e: Int = engram_edge_count()
engram_load_merge(p)
let added_n: Int = engram_node_count() - before_n
let added_e: Int = engram_edge_count() - before_e
let saved: Int = persist_canonical()
"{\"ok\":true,\"nodes_added\":" + int_to_str(added_n) + ",\"edges_added\":" + int_to_str(added_e) + ",\"node_count\":" + int_to_str(engram_node_count()) + "}"
}
// route_emit_ise write an InternalStateEvent node from the soul daemon.
//
// Endpoint: POST /api/neuron/state-events
@@ -241,10 +285,20 @@ fn route_sync(method: String, path: String, body: String) -> String {
//
// Salience/importance set to match engram_node_full ISE defaults used by the
// in-process fallback path in awareness.el (salience=0.3, importance=0.3,
// confidence=0.8, tier=Episodic). High temporal_decay_rate (1.617) ISEs
// are inherently transient; they should decay faster than structural knowledge.
// confidence=0.8, tier=Episodic).
// (2026-06-26 self-review: added this route after discovering ise_post was
// silently failing the soul posts here but the endpoint didn't exist.)
//
// Retention (2026-07-16 self-review): an earlier comment here claimed ISEs
// got temporal_decay_rate=1.617 that was never implemented (engram_node_full
// hardcodes 0.0), and per-node decay only dampens activation anyway; it never
// removes nodes. By 2026-07-16 ISEs were 75% of the store (10,175 of 13,522
// nodes, ~4,300/day, unbounded). ISEs are already WM-excluded in
// engram_activate, so the fix is retention, not decay: every insert calls
// engram_prune_telemetry(), a single O(nodes+edges) compaction pass that
// removes ISEs older than ENGRAM_ISE_RETENTION_MS (default 48h), protecting
// "session-start" labels and self_review events as durable history. At
// ~3 ISEs/min this bounds telemetry at ~8.6k nodes instead of growing forever.
fn route_emit_ise(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
if str_eq(content, "") { return err_json("missing content") }
@@ -256,6 +310,64 @@ fn route_emit_ise(method: String, path: String, body: String) -> String {
sal, imp, conf,
"Episodic", "[\"internal-state\",\"InternalStateEvent\"]"
)
let ret_raw: String = env("ENGRAM_ISE_RETENTION_MS")
let ret_ms: Int = if str_eq(ret_raw, "") { 172800000 } else { str_to_int(ret_raw) }
let pruned: Int = engram_prune_telemetry(ret_ms)
"{\"ok\":true,\"id\":\"" + id + "\",\"pruned\":" + int_to_str(pruned) + "}"
}
// Knowledge capture
//
// route_capture_knowledge direct Knowledge-node capture over HTTP.
//
// Endpoint: POST /api/neuron/knowledge/capture (auth required: "_auth" in body)
// Body: {"content": "...", "title": "...", "category": "...",
// "tier": "note|lesson|canonical", "tags": [...], "project": "...",
// "_auth": "<key>"}
//
// WHY (2026-07-15 self-review): the world-ingestor integrator was designed
// against this endpoint (its MCP-unavailable fallback), but the route never
// existed every direct push 404'd, and because the auth gate ran before
// routing, the failure surfaced as {"error":"unauthorized"} and was
// misdiagnosed for two weeks while world knowledge silently dropped.
// POST /api/nodes was no substitute: it discards label/tags/tier, which
// makes captured knowledge invisible to tag-scoped search and curiosity.
//
// The incoming knowledge tier (note/lesson/canonical) is preserved as a
// "tier:<x>" tag rather than mapped onto Engram's cognitive tiers Knowledge
// nodes land in Semantic (stable reference), and the epistemic tier stays
// queryable without inventing a lossy mapping.
fn route_capture_knowledge(method: String, path: String, body: String) -> String {
let content: String = json_get_string(body, "content")
if str_eq(content, "") { return err_json("missing content") }
let title: String = json_get_string(body, "title")
let label: String = if str_eq(title, "") { str_slice(content, 0, 60) } else { title }
let category_raw: String = json_get_string(body, "category")
let category: String = if str_eq(category_raw, "") { "other" } else { category_raw }
let ktier_raw: String = json_get_string(body, "tier")
let ktier: String = if str_eq(ktier_raw, "") { "note" } else { ktier_raw }
let project: String = json_get_string(body, "project")
let tags_raw: String = json_get_raw(body, "tags")
let tags_base: String = if str_eq(tags_raw, "") { "[]" } else { tags_raw }
// Merge category/tier/project markers into the tag array. Search matches
// against the tags string, so these make captures findable by facet.
let base_len: Int = str_len(tags_base)
let head: String = str_slice(tags_base, 0, base_len - 1)
let sep: String = if str_eq(head, "[") { "" } else { "," }
let safe_cat: String = str_replace(category, "\"", "'")
let safe_tier: String = str_replace(ktier, "\"", "'")
let safe_proj: String = str_replace(project, "\"", "'")
let proj_tag: String = if str_eq(safe_proj, "") { "" } else { ",\"project:" + safe_proj + "\"" }
let tags: String = head + sep + "\"category:" + safe_cat + "\",\"tier:" + safe_tier + "\"" + proj_tag + "]"
let sal: Float = 0.5
let imp: Float = 0.5
let conf: Float = 0.9
let id: String = engram_node_full(
content, "Knowledge", label,
sal, imp, conf,
"Semantic", tags
)
let saved: Int = persist_canonical()
"{\"ok\":true,\"id\":\"" + id + "\"}"
}
@@ -295,6 +407,12 @@ fn handle_request(method: String, path: String, body: String) -> String {
return err_json("unauthorized")
}
// Knowledge capture (auth enforced above; the world-ingestor integrator
// and any headless session without MCP push knowledge through this)
if str_eq(method, "POST") && str_eq(clean, "/api/neuron/knowledge/capture") {
return route_capture_knowledge(method, path, body)
}
// Stats
if str_eq(method, "GET") && (str_eq(clean, "/api/stats") || str_eq(clean, "/stats")) {
return route_stats(method, path, body)
@@ -351,6 +469,9 @@ fn handle_request(method: String, path: String, body: String) -> String {
if str_eq(method, "POST") && (str_eq(clean, "/api/load") || str_eq(clean, "/load")) {
return route_load(method, path, body)
}
if str_eq(method, "POST") && (str_eq(clean, "/api/load-merge") || str_eq(clean, "/load-merge")) {
return route_load_merge(method, path, body)
}
// Sync soul daemon periodic pull of non-ISE knowledge into in-process graph
if str_eq(method, "GET") && str_eq(clean, "/api/sync") {
@@ -362,16 +483,31 @@ fn handle_request(method: String, path: String, body: String) -> String {
// Entry
let bind_str: String = env("ENGRAM_BIND")
if str_eq(bind_str, "") { let bind_str = ":8742" }
let bind_raw: String = env("ENGRAM_BIND")
let bind_str: String = if str_eq(bind_raw, "") { ":8742" } else { bind_raw }
let port: Int = parse_port(bind_str)
// On startup, try to load any existing snapshot (best effort).
let data_dir: String = env("ENGRAM_DATA_DIR")
if str_eq(data_dir, "") { let data_dir = "/tmp/engram" }
let data_dir_raw: String = env("ENGRAM_DATA_DIR")
let data_dir: String = if str_eq(data_dir_raw, "") { "/tmp/engram" } else { data_dir_raw }
let snapshot_path: String = data_dir + "/snapshot.json"
engram_load(snapshot_path)
// 2026-07-21 self-review boot guard: if the snapshot file has content but the
// load produced 0 nodes, something is wrong (corrupt file / parse failure).
// Preserve the evidence and warn loudly and since read routes no longer write
// the canonical path, a bad boot can no longer clobber the good snapshot.
let boot_snap: String = fs_read(snapshot_path)
if !str_eq(boot_snap, "") {
if engram_node_count() == 0 {
println("[engram] WARNING: snapshot.json is non-empty but load produced 0 nodes — preserving copy at snapshot.failed-load.json")
fs_write(data_dir + "/snapshot.failed-load.json", boot_snap)
} else {
// Good load: keep a boot-time backup of the snapshot as loaded.
fs_write(data_dir + "/snapshot.boot-backup.json", boot_snap)
}
}
println("[engram] runtime-native graph engine")
println("[engram] data_dir=" + data_dir)
println("[engram] node_count=" + int_to_str(engram_node_count()))
+10
View File
@@ -17,6 +17,16 @@
// 4. Append dep to order after all its transitive deps
// 5. Deduplicate: skip already-ordered vessels
// Cross-module forward declarations
// Defined in sibling epm modules; resolved at link time. The `extern fn` decls
// give elc the C prototypes so generated install.c compiles cleanly under strict
// compilers (gcc>=14 / clang) that reject implicit function declarations.
extern fn manifest_name(src: String) -> String // manifest.el
extern fn manifest_deps(src: String) -> String // manifest.el
extern fn registry_token() -> String // registry.el
extern fn registry_find(name: String, version: String) -> String // registry.el
extern fn registry_latest_version(name: String) -> String // registry.el
// Install paths
// packages_dir returns the root directory for installed vessels.
+9
View File
@@ -14,6 +14,15 @@
// EPM_REGISTRY_ORG org name that hosts vessel repos (default: neuron-technologies)
// EPM_TOKEN Gitea personal access token (required for publish)
// Cross-module forward declarations
// These symbols are defined in sibling epm modules or the El runtime and are
// resolved at link time. The `extern fn` decls give elc the C prototype so the
// generated registry.c compiles cleanly under strict compilers (gcc>=14 / clang)
// that reject implicit function declarations. Signature arity must match the
// definition; return/param types are informational (all lower to el_val_t).
extern fn config(key: String) -> String // El runtime builtin
extern fn read_installed() -> String // install.el
// Config helpers
// registry_api_url returns the Gitea API base URL with no trailing slash.
+9
View File
@@ -6,6 +6,15 @@
// Depends on: registry.el (registry_latest_version, registry_find),
// install.el (read_installed, install_vessel, installed_version)
// Cross-module forward declarations
// Defined in sibling epm modules; resolved at link time. The `extern fn` decls
// give elc the C prototypes so generated update.c compiles cleanly under strict
// compilers (gcc>=14 / clang) that reject implicit function declarations.
extern fn read_installed() -> String // install.el
extern fn installed_version(name: String) -> String // install.el
extern fn install_vessel(name: String, version: String) -> Bool // install.el
extern fn registry_latest_version(name: String) -> String // registry.el
// Semver helpers
// semver_part extracts the Nth dot-separated component from a semver string.
@@ -75,6 +75,7 @@ static inline void* el_win_dlsym(void* handle, const char* name) {
#include <direct.h> /* _mkdir */
#define mkdir(path, mode) _mkdir(path) /* POSIX mkdir(path,mode) → _mkdir(path) */
#define timegm _mkgmtime /* UTC tm → time_t */
#define fsync(fd) _commit(fd) /* no fsync() on Windows; _commit() (<io.h>) is the equiv */
/* setenv/unsetenv: not in the Windows CRT; map to _putenv_s / SetEnvironmentVariable. */
static inline int setenv(const char* name, const char* value, int overwrite) {
+457 -40
View File
@@ -82,8 +82,14 @@ static _Thread_local ElArena _tl_arena = {NULL, 0, 0};
static _Thread_local int _tl_arena_active = 0;
/* Binary-safe fs_read length — set by fs_read, consumed by http_send_response.
* Allows serving PNGs and other binary files without strlen truncation. */
static _Thread_local size_t _tl_fs_read_len = 0;
* Allows serving PNGs and other binary files without strlen truncation.
* PAIRED with the buffer pointer it describes: the length may only be applied
* to the exact buffer fs_read returned. Without the pairing, any handler that
* fs_read a file and then WRAPPED it into a larger response had that response
* truncated to the file's length (Content-Length lied AND the send stopped
* short) the safety-contact onboarding trap, 2026-07-17. */
static _Thread_local size_t _tl_fs_read_len = 0;
static _Thread_local const char* _tl_fs_read_buf = NULL;
static void el_arena_track(char* p) {
if (!_tl_arena_active || !p) return;
@@ -101,6 +107,8 @@ static void el_arena_track(char* p) {
void el_request_start(void) {
_tl_arena.count = 0;
_tl_arena_active = 1;
_tl_fs_read_len = 0; /* never let a previous request's file length */
_tl_fs_read_buf = NULL; /* leak into this response's byte accounting */
}
/* Called by http_worker after the El handler returns and the response is sent.
@@ -1484,11 +1492,14 @@ static void http_send_response(int fd, const char* body) {
}
const char* eff_body = is_envelope ? env_body : body;
/* Use the real byte count from fs_read if available (handles binary files
* with embedded null bytes PNG, WOFF2, etc.). Fall back to strlen for
* normal text/JSON responses where _tl_fs_read_len is 0. */
size_t blen = (_tl_fs_read_len > 0) ? _tl_fs_read_len : strlen(eff_body);
/* Use the real byte count from fs_read ONLY when this body IS the exact
* buffer fs_read returned (binary files with embedded null bytes PNG,
* WOFF2, etc.). Any other body wrapped, enveloped, or derived must be
* measured with strlen, or it is truncated/over-read to the file's size. */
size_t blen = (_tl_fs_read_len > 0 && eff_body == _tl_fs_read_buf)
? _tl_fs_read_len : strlen(eff_body);
_tl_fs_read_len = 0; /* consume — one-shot per response */
_tl_fs_read_buf = NULL;
int head_only = _tl_http_head_only;
JsonBuf hdrs; jb_init(&hdrs);
@@ -1568,11 +1579,22 @@ static void* http_worker(void* arg) {
const char* rs = EL_CSTR(r);
/* Copy response out BEFORE arena teardown.
* For binary files, _tl_fs_read_len holds the real byte count
* use memcpy instead of strdup so null bytes are preserved. */
size_t rlen = _tl_fs_read_len > 0 ? _tl_fs_read_len : (rs ? strlen(rs) : 0);
* use memcpy instead of strdup so null bytes are preserved.
* The stored length applies ONLY when the response IS the exact
* fs_read buffer; a wrapped/derived response must use strlen or
* it gets truncated (or over-read) to the file's length. */
size_t rlen;
if (_tl_fs_read_len > 0 && rs && rs == _tl_fs_read_buf) {
rlen = _tl_fs_read_len; /* raw file bytes — binary-safe */
} else {
rlen = rs ? strlen(rs) : 0;
_tl_fs_read_len = 0; /* hint doesn't describe this body */
_tl_fs_read_buf = NULL;
}
response = malloc(rlen + 1);
if (response && rs) { memcpy(response, rs, rlen); response[rlen] = '\0'; }
else if (response) { response[0] = '\0'; }
if (_tl_fs_read_len > 0) _tl_fs_read_buf = response; /* hint follows the copy */
} else {
response = el_strdup_persist("el-runtime: no http handler registered");
}
@@ -1822,10 +1844,20 @@ static void* http_worker_v2(void* arg) {
el_val_t hmap = http_build_headers_map(hdr_block ? hdr_block : "");
el_val_t r = h(EL_STR(dispatch_method), EL_STR(path), hmap, EL_STR(body));
const char* rs = EL_CSTR(r);
size_t rlen = _tl_fs_read_len > 0 ? _tl_fs_read_len : (rs ? strlen(rs) : 0);
/* Same pairing rule as the v1 worker: the fs_read length is only
* trustworthy for the exact buffer fs_read returned. */
size_t rlen;
if (_tl_fs_read_len > 0 && rs && rs == _tl_fs_read_buf) {
rlen = _tl_fs_read_len; /* raw file bytes — binary-safe */
} else {
rlen = rs ? strlen(rs) : 0;
_tl_fs_read_len = 0; /* hint doesn't describe this body */
_tl_fs_read_buf = NULL;
}
response = malloc(rlen + 1);
if (response && rs) { memcpy(response, rs, rlen); response[rlen] = '\0'; }
else if (response) { response[0] = '\0'; }
if (_tl_fs_read_len > 0) _tl_fs_read_buf = response; /* hint follows the copy */
el_release(hmap);
} else {
response = el_strdup_persist(
@@ -1963,8 +1995,9 @@ void http_serve_async(el_val_t port, el_val_t handler) {
int sock = socket(AF_INET6, SOCK_STREAM, 0);
if (sock < 0) { perror("socket"); return; }
int yes = 1; int no = 0;
setsockopt(sock, SOL_SOCKET, SO_REUSEADDR, &yes, sizeof(yes));
setsockopt(sock, IPPROTO_IPV6, IPV6_V6ONLY, &no, sizeof(no));
/* Win32/mingw setsockopt takes optval as (const char*); the cast is portable on POSIX too. */
setsockopt(sock, SOL_SOCKET, SO_REUSEADDR, (const char*)&yes, sizeof(yes));
setsockopt(sock, IPPROTO_IPV6, IPV6_V6ONLY, (const char*)&no, sizeof(no));
struct sockaddr_in6 addr;
memset(&addr, 0, sizeof(addr));
addr.sin6_family = AF_INET6;
@@ -2023,6 +2056,7 @@ el_val_t http_response(el_val_t status, el_val_t headers_json, el_val_t body) {
el_val_t fs_read(el_val_t pathv) {
const char* path = EL_CSTR(pathv);
_tl_fs_read_len = 0;
_tl_fs_read_buf = NULL;
if (!path) return el_wrap_str(el_strdup(""));
FILE* f = fopen(path, "rb");
if (!f) return el_wrap_str(el_strdup(""));
@@ -2034,6 +2068,7 @@ el_val_t fs_read(el_val_t pathv) {
size_t got = fread(buf, 1, (size_t)sz, f);
buf[got] = '\0';
_tl_fs_read_len = got; /* store real byte count for binary-safe send */
_tl_fs_read_buf = buf; /* ...valid ONLY for this exact buffer */
fclose(f);
return el_wrap_str(buf);
}
@@ -3576,8 +3611,10 @@ el_val_t json_get_raw(el_val_t json_str, el_val_t key) {
const char* k = EL_CSTR(key);
const char* p = json_find_key(json, k);
/* Clear fs_read binary-length hint — result is a fresh null-terminated
* string, not the raw file bytes, so Content-Length must use strlen. */
* string, not the raw file bytes, so Content-Length must use strlen.
* (Kept although the pointer pairing now makes this redundant.) */
_tl_fs_read_len = 0;
_tl_fs_read_buf = NULL;
if (!p) return el_wrap_str(el_strdup(""));
const char* end = json_skip_value(p);
size_t n = (size_t)(end - p);
@@ -6826,6 +6863,312 @@ static int istr_contains(const char* hay, const char* needle) {
return 0;
}
/* ── Tokenized query matching ───────────────────────────────────────────
* The engram query surface (search / activate / goal-bias) historically
* matched the ENTIRE raw query string as a single case-insensitive
* substring via istr_contains(field, q). That is Ctrl-F, not search:
* a multi-word query like "windows msi signing" only matched a node whose
* text contained that exact contiguous run, so real multi-word queries
* returned zero. istr_contains stays as the per-TOKEN primitive; these
* helpers split the query on whitespace and match ANY token, then rank by
* how many DISTINCT tokens a node covers. Single-token queries are a strict
* special case (score is 0 or 1) so single-word callers never regress. */
#define ENGRAM_MAX_QTOKENS 32
#define ENGRAM_QTOK_LEN 256
/* Split q on whitespace into up to ENGRAM_MAX_QTOKENS distinct
* (case-insensitive) tokens. Returns the token count. Over-long tokens are
* truncated to ENGRAM_QTOK_LEN-1; over-count tokens are ignored. */
static int engram_tokenize_query(const char* q,
char toks[][ENGRAM_QTOK_LEN], int maxtok) {
int n = 0;
if (!q) return 0;
const char* p = q;
while (*p && n < maxtok) {
while (*p && isspace((unsigned char)*p)) p++;
if (!*p) break;
char buf[ENGRAM_QTOK_LEN];
size_t tl = 0;
while (*p && !isspace((unsigned char)*p)) {
if (tl < sizeof(buf) - 1) buf[tl++] = *p;
p++;
}
buf[tl] = '\0';
if (tl == 0) continue;
int dup = 0;
for (int s = 0; s < n; s++) {
if (strcasecmp(toks[s], buf) == 0) { dup = 1; break; }
}
if (dup) continue;
memcpy(toks[n], buf, tl + 1);
n++;
}
return n;
}
/* Count how many of the ntok distinct query tokens appear (case-insensitive)
* in the node's content, label, or tags. 0 == no match. */
static int engram_node_match_score(const EngramNode* n,
char toks[][ENGRAM_QTOK_LEN], int ntok) {
int score = 0;
for (int t = 0; t < ntok; t++) {
if (istr_contains(n->content, toks[t]) ||
istr_contains(n->label, toks[t]) ||
istr_contains(n->tags, toks[t]))
score++;
}
return score;
}
/* Rank entry: distinct-token match count (primary, desc) then salience
* (tiebreak, desc). */
typedef struct { int64_t idx; int score; double salience; } EngramRankEntry;
static int engram_rank_cmp(const void* a, const void* b) {
const EngramRankEntry* ea = (const EngramRankEntry*)a;
const EngramRankEntry* eb = (const EngramRankEntry*)b;
if (ea->score != eb->score) return eb->score - ea->score; /* desc */
if (ea->salience < eb->salience) return 1;
if (ea->salience > eb->salience) return -1;
return 0;
}
/* ══════════════════════════════════════════════════════════════════════════
* SEMANTIC SEARCH LAYER nomic-embed-text via Ollama /api/embeddings
*
* Augments the lexical (istr_contains) matcher with dense-vector retrieval.
* Node content and the query are embedded through a local Ollama server;
* nodes are ranked by cosine similarity and UNIONED with lexical hits. This
* lets a paraphrase query surface a node whose words never appear in it.
*
* DEGRADABLE BY DESIGN. The whole layer is gated on HAVE_CURL plus a one-shot
* runtime probe of the embedding endpoint. If curl is not compiled in, or
* Ollama is unreachable, or ENGRAM_SEMANTIC=0, every entry point returns
* "no semantic signal" and callers fall back to pure lexical behaviour
* byte-for-byte the pre-existing search.
*
* CACHE. Node embeddings are computed lazily on first use and cached in
* process memory keyed by node id, with an FNV-1a content hash for
* invalidation (edited content re-embeds). The query is embedded once per
* search call. This is what "avoid re-embedding the whole graph every query"
* buys us: a warm cache serves cosine from RAM. (A cold process still pays
* O(N) embed calls the first time each node is scanned persisting the cache
* to a snapshot sidecar is the documented next step, not done here.)
*
* nomic task prefixes ("search_query:" / "search_document:") are applied
* because nomic-embed-text is trained with them; they materially improve
* retrieval separation (empirically: paraphrase 0.72 vs distractors <0.48).
*
* ENV:
* ENGRAM_SEMANTIC "0" disables; unset/other = auto-probe
* ENGRAM_EMBED_URL default http://localhost:11434/api/embeddings
* ENGRAM_EMBED_MODEL default nomic-embed-text
* ENGRAM_SEMANTIC_MIN cosine threshold for a pure-semantic match (def 0.6)
* */
static double engram_semantic_min(void) {
static double v = -1.0;
if (v >= 0.0) return v;
const char* s = getenv("ENGRAM_SEMANTIC_MIN");
double d = 0.6;
if (s && *s) { char* e = NULL; double t = strtod(s, &e);
if (e != s && t >= 0.0 && t <= 1.0) d = t; }
v = d; return v;
}
#ifdef HAVE_CURL
typedef struct { char* id; uint64_t hash; float* vec; int dim; } EngramEmbEntry;
static EngramEmbEntry* g_emb_items = NULL;
static int64_t g_emb_count = 0, g_emb_cap = 0;
static int g_emb_state = 0; /* 0=unprobed, 1=available, -1=disabled */
static uint64_t engram_fnv1a(const char* s) {
uint64_t h = 1469598103934665603ULL;
if (s) for (const unsigned char* p = (const unsigned char*)s; *p; p++) {
h ^= *p; h *= 1099511628211ULL;
}
return h;
}
/* Parse "embedding":[f,f,...] from an Ollama response. malloc'd vec, or NULL. */
static float* engram_parse_embedding(const char* json, int* out_dim) {
if (!json) return NULL;
const char* p = strstr(json, "\"embedding\"");
if (!p) return NULL;
p = strchr(p, '[');
if (!p) return NULL;
p++;
int cap = 1024, n = 0;
float* v = malloc((size_t)cap * sizeof(float));
if (!v) return NULL;
while (*p && *p != ']') {
while (*p == ' ' || *p == '\t' || *p == '\n' || *p == '\r' || *p == ',') p++;
if (*p == ']' || !*p) break;
char* e = NULL;
double d = strtod(p, &e);
if (e == p) break;
if (n >= cap) { cap *= 2; float* nv = realloc(v, (size_t)cap * sizeof(float));
if (!nv) { free(v); return NULL; } v = nv; }
v[n++] = (float)d;
p = e;
}
if (n == 0) { free(v); return NULL; }
*out_dim = n;
return v;
}
/* JSON-escape src into a malloc'd buffer (no surrounding quotes). */
static char* engram_json_escape(const char* src) {
if (!src) src = "";
size_t n = strlen(src);
char* out = malloc(n * 2 + 1);
if (!out) return NULL;
size_t j = 0;
for (size_t i = 0; i < n; i++) {
unsigned char c = (unsigned char)src[i];
if (c == '"') { out[j++] = '\\'; out[j++] = '"'; }
else if (c == '\\') { out[j++] = '\\'; out[j++] = '\\'; }
else if (c == '\n') { out[j++] = '\\'; out[j++] = 'n'; }
else if (c == '\r') { out[j++] = '\\'; out[j++] = 'r'; }
else if (c == '\t') { out[j++] = '\\'; out[j++] = 't'; }
else if (c < 0x20) { /* drop other control bytes */ }
else { out[j++] = (char)c; }
}
out[j] = '\0';
return out;
}
/* Embed `prefix+text` via Ollama. Returns malloc'd vec (caller frees), or NULL. */
static float* engram_embed_raw(const char* prefix, const char* text, int* out_dim) {
if (!text) return NULL;
const char* url = getenv("ENGRAM_EMBED_URL");
if (!url || !*url) url = "http://localhost:11434/api/embeddings";
const char* model = getenv("ENGRAM_EMBED_MODEL");
if (!model || !*model) model = "nomic-embed-text";
/* Bound content length to keep latency/memory sane on huge nodes. */
char* trunc = NULL;
size_t maxlen = 8192;
if (strlen(text) > maxlen) {
trunc = malloc(maxlen + 1);
if (trunc) { memcpy(trunc, text, maxlen); trunc[maxlen] = '\0'; text = trunc; }
}
char* esc_prefix = engram_json_escape(prefix ? prefix : "");
char* esc = engram_json_escape(text);
free(trunc);
if (!esc || !esc_prefix) { free(esc); free(esc_prefix); return NULL; }
size_t blen = strlen(esc) + strlen(esc_prefix) + strlen(model) + 64;
char* body = malloc(blen);
if (!body) { free(esc); free(esc_prefix); return NULL; }
snprintf(body, blen, "{\"model\":\"%s\",\"prompt\":\"%s%s\"}", model, esc_prefix, esc);
free(esc); free(esc_prefix);
CURL* c = curl_easy_init();
if (!c) { free(body); return NULL; }
HttpBuf rb; httpbuf_init(&rb);
struct curl_slist* h = curl_slist_append(NULL, "Content-Type: application/json");
char errbuf[CURL_ERROR_SIZE]; errbuf[0] = '\0';
curl_easy_setopt(c, CURLOPT_URL, url);
curl_easy_setopt(c, CURLOPT_WRITEFUNCTION, http_write_cb);
curl_easy_setopt(c, CURLOPT_WRITEDATA, &rb);
curl_easy_setopt(c, CURLOPT_POST, 1L);
curl_easy_setopt(c, CURLOPT_POSTFIELDS, body);
curl_easy_setopt(c, CURLOPT_POSTFIELDSIZE, (long)strlen(body));
curl_easy_setopt(c, CURLOPT_HTTPHEADER, h);
curl_easy_setopt(c, CURLOPT_TIMEOUT_MS, el_http_timeout_ms());
curl_easy_setopt(c, CURLOPT_NOSIGNAL, 1L);
curl_easy_setopt(c, CURLOPT_ERRORBUFFER, errbuf);
CURLcode rc = curl_easy_perform(c);
curl_slist_free_all(h);
curl_easy_cleanup(c);
free(body);
if (rc != CURLE_OK) { free(rb.data); return NULL; }
float* v = engram_parse_embedding(rb.data, out_dim);
free(rb.data);
return v;
}
/* One-shot probe: is semantic search available? Caches the verdict. */
static int engram_semantic_enabled(void) {
if (g_emb_state != 0) return g_emb_state == 1;
const char* s = getenv("ENGRAM_SEMANTIC");
if (s && strcmp(s, "0") == 0) { g_emb_state = -1; return 0; }
int dim = 0;
float* v = engram_embed_raw("search_query: ", "probe", &dim);
if (v && dim > 0) { free(v); g_emb_state = 1; return 1; }
free(v);
g_emb_state = -1; return 0;
}
/* Embed the query. Returns malloc'd vec (caller frees), or NULL if semantic off. */
static float* engram_embed_query(const char* q, int* dim) {
if (!engram_semantic_enabled()) return NULL;
if (!q || !*q) return NULL;
return engram_embed_raw("search_query: ", q, dim);
}
/* Cached node embedding. Returns a pointer OWNED BY THE CACHE — do not free. */
static const float* engram_node_vec(EngramNode* n, int* out_dim) {
if (!n || !n->id) return NULL;
uint64_t h = engram_fnv1a(n->content);
for (int64_t i = 0; i < g_emb_count; i++) {
if (g_emb_items[i].id && strcmp(g_emb_items[i].id, n->id) == 0) {
if (g_emb_items[i].hash == h && g_emb_items[i].vec) {
*out_dim = g_emb_items[i].dim; return g_emb_items[i].vec;
}
/* content changed → re-embed in place */
int dim = 0;
float* v = engram_embed_raw("search_document: ", n->content ? n->content : "", &dim);
if (!v) return NULL;
free(g_emb_items[i].vec);
g_emb_items[i].vec = v; g_emb_items[i].dim = dim; g_emb_items[i].hash = h;
*out_dim = dim; return v;
}
}
int dim = 0;
float* v = engram_embed_raw("search_document: ", n->content ? n->content : "", &dim);
if (!v) return NULL;
if (g_emb_count >= g_emb_cap) {
int64_t nc = g_emb_cap ? g_emb_cap * 2 : 256;
EngramEmbEntry* ni = realloc(g_emb_items, (size_t)nc * sizeof(EngramEmbEntry));
if (!ni) { free(v); return NULL; }
g_emb_items = ni; g_emb_cap = nc;
}
g_emb_items[g_emb_count].id = strdup(n->id);
g_emb_items[g_emb_count].hash = h;
g_emb_items[g_emb_count].vec = v;
g_emb_items[g_emb_count].dim = dim;
g_emb_count++;
*out_dim = dim; return v;
}
static double engram_cosine(const float* a, const float* b, int dim) {
double dot = 0, na = 0, nb = 0;
for (int i = 0; i < dim; i++) { dot += (double)a[i] * b[i];
na += (double)a[i] * a[i];
nb += (double)b[i] * b[i]; }
if (na <= 0 || nb <= 0) return 0.0;
return dot / (sqrt(na) * sqrt(nb));
}
/* Cosine of node n against the query vector; 0 if unavailable / dim mismatch. */
static double engram_node_cosine(EngramNode* n, const float* qvec, int qdim) {
if (!qvec || qdim <= 0) return 0.0;
int ndim = 0;
const float* nv = engram_node_vec(n, &ndim);
if (!nv || ndim != qdim) return 0.0;
return engram_cosine(qvec, nv, qdim);
}
#else /* !HAVE_CURL — semantic layer compiled out; callers stay pure-lexical.
* Only the two boundary functions the always-compiled search/activate
* code calls are stubbed; the query embed always yields NULL so every
* cosine is 0 and every caller collapses to lexical-only. */
static float* engram_embed_query(const char* q, int* dim) { (void)q; (void)dim; return NULL; }
static double engram_node_cosine(EngramNode* n, const float* qvec, int qdim) {
(void)n; (void)qvec; (void)qdim; return 0.0;
}
#endif /* HAVE_CURL */
el_val_t engram_search(el_val_t query, el_val_t limit) {
EngramStore* g = engram_get();
const char* q = EL_CSTR(query);
@@ -6833,21 +7176,45 @@ el_val_t engram_search(el_val_t query, el_val_t limit) {
if (lim <= 0) lim = 100;
el_val_t lst = el_list_empty();
if (!q || !*q) return lst;
int64_t found = 0;
for (int64_t i = 0; i < g->node_count && found < lim; i++) {
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
if (ntok == 0) return lst;
/* Semantic augmentation: embed the query once; a node is a hit if it covers
* >=1 query token (tokenized-lexical, #66) OR its cosine clears the
* threshold (#67). qvec is NULL (cosine 0) when semantic is unavailable
* pure tokenized-lexical, byte-identical to the lexical-only behaviour. */
int qdim = 0;
float* qvec = engram_embed_query(q, &qdim);
double sem_min = engram_semantic_min();
EngramRankEntry* hits = malloc((size_t)g->node_count * sizeof(EngramRankEntry));
if (!hits) { free(qvec); return lst; }
int64_t nhits = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
/* Filter transparent layers: nodes whose layer is `transparent=1`
* shape output but are invisible to introspection ("what do you
* know about yourself"). They still surface via engram_activate
* + engram_compile_layered_json that's the legitimate path. */
if (engram_layer_is_transparent(n->layer_id)) continue;
if (istr_contains(n->content, q) ||
istr_contains(n->label, q) ||
istr_contains(n->tags, q)) {
lst = el_list_append(lst, engram_node_to_map(n));
found++;
int sc = engram_node_match_score(n, toks, ntok);
double sem = qvec ? engram_node_cosine(n, qvec, qdim) : 0.0;
if (sc > 0 || sem >= sem_min) {
hits[nhits].idx = i;
hits[nhits].score = sc;
hits[nhits].salience = n->salience;
nhits++;
}
}
/* Rank by distinct tokens matched (desc) then salience (desc), then cap.
* Pure-semantic hits (token score 0) sort after every lexical hit a
* lexical semantic union with lexical precedence. */
qsort(hits, (size_t)nhits, sizeof(EngramRankEntry), engram_rank_cmp);
int64_t end = nhits < lim ? nhits : lim;
for (int64_t k = 0; k < end; k++) {
lst = el_list_append(lst, engram_node_to_map(&g->nodes[hits[k].idx]));
}
free(hits);
free(qvec);
return lst;
}
@@ -7124,10 +7491,14 @@ static double engram_temporal_proximity_bonus(int64_t node_created,
static double engram_goal_bias(const EngramNode* n, const char* query) {
if (!query || !*query) return 1.0;
double bias = 1.0;
/* Direct lexical overlap: node content/label/tags share text with query. */
if (istr_contains(n->content, query) || istr_contains(n->label, query) ||
istr_contains(n->tags, query)) {
bias += 0.5;
/* Direct lexical overlap, graded by token coverage: a node covering all
* query tokens gets the full +0.5; partial coverage gets a proportional
* share. Single-token queries full +0.5 on match, identical to before. */
{
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(query, toks, ENGRAM_MAX_QTOKENS);
int sc = engram_node_match_score(n, toks, ntok);
if (sc > 0 && ntok > 0) bias += 0.5 * ((double)sc / (double)ntok);
}
/* Node-type resonance with query intent. */
int technical_query = istr_contains(query, "code") ||
@@ -7193,14 +7564,31 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
if (!seeds) {
free(best_bg); free(best_hops); free(reached); return out;
}
/* Tokenized + semantic seeding: a node seeds if it covers >=1 query token
* (tokenized-lexical, #66) OR its cosine clears the threshold (#67). A
* lexical seed's activation is scaled by token coverage (fraction of
* distinct query tokens covered) so a node matching all words seeds more
* strongly than one matching a single word; single-word queries coverage
* 1.0. A pure-semantic seed (no token match) is instead down-weighted by
* its cosine so paraphrase matches spread without overpowering exact seeds.
* q_vec is NULL (cosine 0) when semantic is unavailable the seed set is
* exactly the tokenized-lexical one. q_vec is freed right after this loop
* so the many downstream early-returns need no cleanup change. */
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
int q_dim = 0;
float* q_vec = engram_embed_query(q, &q_dim);
double q_sem_min = engram_semantic_min();
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
if (istr_contains(n->content, q) ||
istr_contains(n->label, q) ||
istr_contains(n->tags, q)) {
int sc = engram_node_match_score(n, toks, ntok);
double sem = q_vec ? engram_node_cosine(n, q_vec, q_dim) : 0.0;
if (sc > 0 || sem >= q_sem_min) {
double tdecay = engram_temporal_decay(n, now_ms);
double dampen = engram_activation_dampen(n);
double act = n->salience * tdecay * dampen;
if (sc > 0) act *= (ntok > 0 ? (double)sc / (double)ntok : 1.0);
else act *= sem; /* pure-semantic seed: down-weight by cosine */
seeds[seed_count].idx = i;
seeds[seed_count].act = act;
seeds[seed_count].created_at = n->created_at;
@@ -7210,6 +7598,7 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
reached[i] = 1;
}
}
free(q_vec);
/* Compute mean seed created_at for temporal proximity bonus. */
int64_t seed_epoch = 0;
if (seed_count > 0) {
@@ -7797,22 +8186,50 @@ el_val_t engram_search_json(el_val_t query, el_val_t limit) {
if (lim <= 0) lim = 100;
JsonBuf b; jb_init(&b);
jb_putc(&b, '[');
int first = 1;
int64_t found = 0;
if (q && *q) {
for (int64_t i = 0; i < g->node_count && found < lim; i++) {
EngramNode* n = &g->nodes[i];
/* Filter transparent layers — same as engram_search. */
if (engram_layer_is_transparent(n->layer_id)) continue;
if (istr_contains(n->content, q) ||
istr_contains(n->label, q) ||
istr_contains(n->tags, q)) {
if (!first) jb_putc(&b, ',');
engram_emit_node_json(&b, n);
first = 0;
found++;
if (q && *q && g->node_count > 0) {
/* Collect candidates from the UNION of tokenized-lexical and semantic
* matches, score each, rank by score, emit the top `lim`. A node is a
* candidate if it covers >=1 query token (tokenized-lexical, #66) OR its
* query cosine clears the threshold (#67). Lexical score is the distinct
* token count (>=1), so any lexical hit outranks a pure-semantic hit
* (cosine < 1); pure-semantic hits are scored by cosine alone. When
* semantic is unavailable qvec is NULL, sem is 0, only tokenized-lexical
* hits are collected, and the stable insertion sort preserves order. */
char toks[ENGRAM_MAX_QTOKENS][ENGRAM_QTOK_LEN];
int ntok = engram_tokenize_query(q, toks, ENGRAM_MAX_QTOKENS);
int qdim = 0;
float* qvec = engram_embed_query(q, &qdim);
double sem_min = engram_semantic_min();
typedef struct { int64_t idx; double score; } Cand;
Cand* cand = malloc((size_t)g->node_count * sizeof(Cand));
if (cand) {
int64_t nc = 0;
for (int64_t i = 0; i < g->node_count; i++) {
EngramNode* n = &g->nodes[i];
if (engram_layer_is_transparent(n->layer_id)) continue;
int sc = engram_node_match_score(n, toks, ntok);
double sem = qvec ? engram_node_cosine(n, qvec, qdim) : 0.0;
if (sc > 0 || sem >= sem_min) {
cand[nc].idx = i;
cand[nc].score = (double)sc + sem;
nc++;
}
}
/* Insertion sort by score desc; stable for equal scores. */
for (int64_t i = 1; i < nc; i++) {
Cand k = cand[i]; int64_t j = i - 1;
while (j >= 0 && cand[j].score < k.score) { cand[j + 1] = cand[j]; j--; }
cand[j + 1] = k;
}
int first = 1;
for (int64_t i = 0; i < nc && i < lim; i++) {
if (!first) jb_putc(&b, ',');
engram_emit_node_json(&b, &g->nodes[cand[i].idx]);
first = 0;
}
free(cand);
}
free(qvec);
}
jb_putc(&b, ']');
return el_wrap_str(jb_finish(&b));
+25 -1
View File
@@ -23,10 +23,29 @@ fn tok_at(tokens: [Any], pos: Int) -> Map<String, Any> {
}
fn tok_kind(tokens: [Any], pos: Int) -> String {
// Out-of-range reads must report the Eof sentinel so every `== "Eof"`
// termination guard in the parser fires. Without this, reading past the
// single trailing Eof token returns runtime null (el_list_get OOB -> 0),
// which matches no delimiter, letting inner parse loops append AST nodes
// forever on malformed input -> unbounded allocation -> OOM.
let n: Int = native_list_len(tokens) / 2
if pos < 0 {
return "Eof"
}
if pos >= n {
return "Eof"
}
native_list_get(tokens, pos * 2)
}
fn tok_value(tokens: [Any], pos: Int) -> String {
let n: Int = native_list_len(tokens) / 2
if pos < 0 {
return ""
}
if pos >= n {
return ""
}
native_list_get(tokens, pos * 2 + 1)
}
@@ -35,7 +54,12 @@ fn expect(tokens: [Any], pos: Int, kind: String) -> Int {
if k == kind {
return pos + 1
}
// On mismatch just advance; error recovery is best-effort
// On mismatch, error recovery is best-effort. But never step PAST the Eof
// sentinel: once at Eof a mismatch means the input ended early, and
// advancing would run the cursor off the token list.
if k == "Eof" {
return pos
}
pos + 1
}
@@ -0,0 +1,186 @@
#ifndef EL_PLATFORM_WIN_H
#define EL_PLATFORM_WIN_H
/*
* el_platform_win.h — Windows OS-boundary shim for el_runtime.c.
*
* Branch: feat/windows-el-runtime. Included ONLY when _WIN32 is defined; the POSIX build is
* untouched. Goal: let el_runtime.c (a BSD-sockets / dlfcn / fork host) compile and link with
* mingw-w64 into a native neuron.exe, with no behavioural change to the Linux/macOS build.
*
* What it maps:
* - sockets : winsock2 (same call names: socket/bind/listen/accept/recv/send/setsockopt).
* Sockets close with closesocket() (see el_closesocket), and the stack must be
* started once with WSAStartup — done automatically via a load-time constructor.
* - dlsym : el_runtime.c uses dlsym(RTLD_DEFAULT, name) to resolve callback/tool symbols
* exported by the main module. Windows equivalent: GetProcAddress on the process
* module. Link the soul with -Wl,--export-all-symbols so the symbols are findable.
* - popen : mapped to _popen/_pclose.
* - threads : UNCHANGED. mingw-w64 ships winpthreads, so <pthread.h> + -lpthread just work.
*/
#ifndef WIN32_LEAN_AND_MEAN
#define WIN32_LEAN_AND_MEAN
#endif
#include <winsock2.h>
#include <ws2tcpip.h>
#include <windows.h>
#include <io.h>
#include <process.h>
/* Portable headers mingw-w64 provides (verified present). */
#include <stdarg.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <strings.h> /* strcasecmp */
#include <ctype.h>
#include <math.h>
#include <time.h>
#include <sys/time.h> /* mingw-w64 provides gettimeofday here */
#include <sys/types.h>
#include <sys/stat.h>
#include <fcntl.h>
#include <dirent.h>
#include <errno.h>
#include <pthread.h>
/* ── socket close ─────────────────────────────────────────────────────────── */
/* Winsock closes sockets with closesocket(), not close() (close() is for file fds). The POSIX
build defines the same helper as close() so the call sites are identical across platforms. */
static inline int el_closesocket(SOCKET s) { return closesocket(s); }
/* ── setsockopt optval type ───────────────────────────────────────────────── */
/* Winsock's setsockopt takes optval as (const char*); POSIX takes (const void*), so el_runtime.c
passes &int directly. GCC 14+ makes that an error under -Wincompatible-pointer-types. Wrap it so
the runtime's POSIX-style call sites compile unchanged (defined before the macro so the wrapper
itself resolves to the real winsock setsockopt). */
static inline int el_setsockopt(SOCKET s, int level, int optname, const void* optval, int optlen) {
return setsockopt(s, level, optname, (const char*)optval, optlen);
}
#define setsockopt(s, l, o, v, n) el_setsockopt((s), (l), (o), (v), (int)(n))
/* ── winsock init (once, at load) ─────────────────────────────────────────── */
static void el__win_net_init(void) {
static int inited = 0;
if (!inited) { WSADATA w; WSAStartup(MAKEWORD(2, 2), &w); inited = 1; }
}
__attribute__((constructor)) static void el__win_ctor(void) { el__win_net_init(); }
/* ── dlsym → GetProcAddress ───────────────────────────────────────────────── */
#ifndef RTLD_DEFAULT
#define RTLD_DEFAULT ((void*)0)
#endif
static inline void* el_win_dlsym(void* handle, const char* name) {
(void)handle;
return (void*)(uintptr_t)GetProcAddress(GetModuleHandleA(NULL), name);
}
#define dlsym(h, n) el_win_dlsym((h), (n))
/* ── popen / pclose ───────────────────────────────────────────────────────── */
#define popen _popen
#define pclose _pclose
/* ── misc POSIX → Win32 shims ─────────────────────────────────────────────── */
#include <direct.h> /* _mkdir */
#define mkdir(path, mode) _mkdir(path) /* POSIX mkdir(path,mode) → _mkdir(path) */
#define timegm _mkgmtime /* UTC tm → time_t */
/* setenv/unsetenv: not in the Windows CRT; map to _putenv_s / SetEnvironmentVariable. */
static inline int setenv(const char* name, const char* value, int overwrite) {
(void)overwrite;
return _putenv_s(name, value ? value : "");
}
static inline int unsetenv(const char* name) {
/* _putenv_s(name, "") sets VAR="" rather than removing it.
* SetEnvironmentVariableA(name, NULL) truly deletes it from the Win32
* env block; then we sync the CRT cache with _putenv("NAME="). */
SetEnvironmentVariableA(name, NULL);
size_t len = strlen(name);
char *buf = (char*)malloc(len + 2);
if (!buf) return -1;
memcpy(buf, name, len);
buf[len] = '=';
buf[len + 1] = '\0';
_putenv(buf);
free(buf);
return 0;
}
/* nanosleep — not available in MSVC/UCRT; approximate with Sleep(). */
static inline int el_nanosleep(const struct timespec *req, struct timespec *rem) {
(void)rem;
DWORD ms = (DWORD)((req->tv_sec * 1000ULL) + (req->tv_nsec / 1000000ULL));
Sleep(ms ? ms : 1);
return 0;
}
#define nanosleep(req, rem) el_nanosleep((req), (rem))
/* localtime_r/gmtime_r: Windows offers localtime_s/gmtime_s with reversed arg order. */
static inline struct tm* localtime_r(const time_t* t, struct tm* out) {
return localtime_s(out, t) == 0 ? out : (struct tm*)0;
}
static inline struct tm* gmtime_r(const time_t* t, struct tm* out) {
return gmtime_s(out, t) == 0 ? out : (struct tm*)0;
}
/* ── libcurl: degradable stubs for the curl-less Windows build ─────────────── */
/* The curl-less validation build (WITH_CURL=0) links no libcurl. el_runtime.c uses libcurl
* unconditionally for its HTTP client / LLM layer; these stubs let it compile and link so the
* runtime, HTTP *server*, graph and memory work natively on Windows. Live outbound HTTP/LLM calls
* degrade to a runtime error (curl_easy_perform returns an error) — matching the documented
* curl-less contract. When HAVE_CURL is defined (WITH_CURL=1) the real <curl/curl.h> is used and
* this whole block is compiled out. POSIX never sees this header, so the POSIX build is untouched. */
#ifndef HAVE_CURL
typedef void CURL;
typedef int CURLcode;
#define CURLE_OK 0
#define CURLE_HTTP_RETURNED_ERROR 22
#define CURL_ERROR_SIZE 256
/* Option ids: values are irrelevant to the no-op setopt below; kept distinct for readability. */
#define CURLOPT_URL 10002
#define CURLOPT_WRITEFUNCTION 20011
#define CURLOPT_WRITEDATA 10001
#define CURLOPT_POSTFIELDS 10015
#define CURLOPT_POSTFIELDSIZE 120
#define CURLOPT_POST 47
#define CURLOPT_HTTPHEADER 10023
#define CURLOPT_TIMEOUT_MS 155
#define CURLOPT_NOSIGNAL 99
#define CURLOPT_USERAGENT 10018
#define CURLOPT_FOLLOWLOCATION 52
#define CURLOPT_ERRORBUFFER 10010
#define CURLOPT_CUSTOMREQUEST 10036
#define CURLOPT_FAILONERROR 45
struct curl_slist { char* data; struct curl_slist* next; };
static inline struct curl_slist* curl_slist_append(struct curl_slist* list, const char* s) {
struct curl_slist* node = (struct curl_slist*)malloc(sizeof(struct curl_slist));
if (!node) return list;
node->data = s ? strdup(s) : NULL;
node->next = NULL;
if (!list) return node;
struct curl_slist* p = list;
while (p->next) p = p->next;
p->next = node;
return list;
}
static inline void curl_slist_free_all(struct curl_slist* list) {
while (list) { struct curl_slist* n = list->next; free(list->data); free(list); list = n; }
}
static inline CURL* curl_easy_init(void) { return (CURL*)malloc(1); }
static inline CURLcode curl_easy_setopt(CURL* h, int opt, ...) { (void)h; (void)opt; return CURLE_OK; }
static inline CURLcode curl_easy_perform(CURL* h) { (void)h; return 7 /* CURLE_COULDNT_CONNECT */; }
static inline void curl_easy_cleanup(CURL* h) { free(h); }
static inline const char* curl_easy_strerror(CURLcode c) {
(void)c; return "libcurl not built in (curl-less build)";
}
#endif /* !HAVE_CURL */
#endif /* EL_PLATFORM_WIN_H */
File diff suppressed because it is too large Load Diff
@@ -758,6 +758,18 @@ el_val_t trace_span_start(el_val_t name);
el_val_t trace_span_end(el_val_t span_handle);
el_val_t emit_event(el_val_t name, el_val_t duration_ms);
/* ── Runtime symbols required by the soul modules ──────────────────────────── */
/* All implemented in el_runtime.c but omitted from this release header; the soul dist modules
* reference them directly, so the public header must export them. Declarations only — mirrors the
* mainline el_runtime.h and is platform-independent (no behavioural change to the POSIX build). */
typedef el_val_t (*http_handler_fn)(el_val_t method, el_val_t path, el_val_t body);
typedef el_val_t (*http_handler4_fn)(el_val_t method, el_val_t path, el_val_t body, el_val_t headers);
el_val_t el_arena_push(void);
el_val_t el_arena_pop(el_val_t mark);
void http_serve_async(el_val_t port, el_val_t handler);
el_val_t engram_get_node_by_label(el_val_t label);
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
#ifdef __cplusplus
}
#endif