Compare commits
48 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| fe820928b0 | |||
| 385c18442d | |||
| cace6a5ebf | |||
| d41645388a | |||
| 8a307dfd42 | |||
| 616815b2ab | |||
| 1a8a966cb3 | |||
| 1f70b9fa18 | |||
| 317466e8f7 | |||
| eb3e6d7c1f | |||
| 88e3008735 | |||
| 26af149aa1 | |||
| c18abf799c | |||
| b305b49f40 | |||
| 8ae163e8e5 | |||
| 3fcc36c2f1 | |||
| a6cef4b983 | |||
| 8e9d88fc01 | |||
| e99a4640e2 | |||
| bdc1f99fb9 | |||
| 44b621e551 | |||
| ded6ca546f | |||
| b5b96c05ed | |||
| c79033b749 | |||
| 1119295238 | |||
| 9c07970943 | |||
| 0832865952 | |||
| e0b2c0ea54 | |||
| cf060adbfd | |||
| 63fe8a766d | |||
| b5a0a729e6 | |||
| b26dd47aef | |||
| b55e6bfd53 | |||
| dbb06f6ee4 | |||
| 906c664a65 | |||
| 6a6b589ba0 | |||
| b5d1e53902 | |||
| 9e96d74f6a | |||
| a8908908df | |||
| 6291a35bb9 | |||
| a69a4a5894 | |||
| edafd8cce8 | |||
| 5e3e69d326 | |||
| 4c3414072b | |||
| 0288024396 | |||
| 3e7ab07e82 | |||
| a668062e38 | |||
| 24fac765a6 |
@@ -0,0 +1,630 @@
|
|||||||
|
# El Test Framework — Design
|
||||||
|
|
||||||
|
**Status:** draft for review
|
||||||
|
**Author:** Neuron
|
||||||
|
**Date:** 2026-08-15
|
||||||
|
**Worktree:** `/Users/will/Development/neuron-technologies/el-worktrees/elc-memory-investigation`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0. The forcing requirement
|
||||||
|
|
||||||
|
We have a confirmed quadratic in `elc`. Peak memory in the old shipped binary and wall-clock in
|
||||||
|
the current source both grow as O(input²). We cannot fix it, because we cannot test it.
|
||||||
|
|
||||||
|
Everything in this document is downstream of one sentence: **a test framework must be able to fail
|
||||||
|
a build when an operation's growth curve degrades from linear to quadratic.**
|
||||||
|
|
||||||
|
That is not a nice-to-have bolted onto a correctness framework. It is the requirement that
|
||||||
|
determines the architecture. Correctness testing is the easy half.
|
||||||
|
|
||||||
|
Second-order requirement, learned the hard way tonight: **the framework must report per-test timing
|
||||||
|
by default.** The current framework prints `N passed, M failed` and nothing else. That is why a
|
||||||
|
3.58-second test file sat in the suite unnoticed. A framework that is structurally blind to time
|
||||||
|
cannot surface the defect class we most need to catch.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. What exists today, measured
|
||||||
|
|
||||||
|
### 1.1 Two competing systems, neither complete
|
||||||
|
|
||||||
|
**System A — `lang/runtime/test.el`.** Manual registration, El-level.
|
||||||
|
|
||||||
|
**System B — the compiler's `test { }` block + `elc --test`.** Emits its own harness `main()`
|
||||||
|
with `__el_pass` / `__el_fail` globals (`codegen.el:3777-3796`).
|
||||||
|
|
||||||
|
They do not share a result model. Neither has timing. Both are in the tree.
|
||||||
|
|
||||||
|
### 1.2 Specific defects in System A
|
||||||
|
|
||||||
|
| Defect | Location | Consequence |
|
||||||
|
|---|---|---|
|
||||||
|
| All state as JSON strings in a global string-keyed map | `test.el` throughout | every assertion is `state_get` → `str_to_int` → `int_to_str` → `state_set` |
|
||||||
|
| Failure list appended by string slice + concat | `_test_json_append` | O(n²) in failure count |
|
||||||
|
| One OS thread spawned per test | `_test_run_one` via `__thread_create`/`__thread_join` | thread spawn per test, purely to get dispatch-by-name through dlsym |
|
||||||
|
| Manual registration pairing a string to a function name | `test_case(name, fn_name)` | typo ⇒ test silently never runs, suite still reports pass |
|
||||||
|
| Counters are assertion-level, global | `_test_pass_count` etc. | no per-test record exists at all |
|
||||||
|
| No timing, no structured output, no fixtures, no tags, no filtering, no parameterization, no benchmarks | — | — |
|
||||||
|
|
||||||
|
The registration defect is the serious one. It is not a slow framework, it is a framework that can
|
||||||
|
report success for tests that did not execute.
|
||||||
|
|
||||||
|
### 1.3 Measured cost structure
|
||||||
|
|
||||||
|
Per test file, current build model:
|
||||||
|
|
||||||
|
| Step | Time |
|
||||||
|
|---|---|
|
||||||
|
| `elc` compile `.el` → `.c` | 0.00s (small files) |
|
||||||
|
| **`cc` el_runtime.c → .o** | **0.14s** |
|
||||||
|
| `cc` test .c → .o | 0.02s |
|
||||||
|
| link | 0.02s |
|
||||||
|
|
||||||
|
> **STALE as of el #132 — re-measured 2026-08-16.** The `test_compiler` figure below was
|
||||||
|
> *entirely* the `strlen`-per-character quadratic, now fixed. Re-measured on the same host:
|
||||||
|
> **3.58s → 0.03s (119x)**, and the 422 KB compiler concatenation likewise compiles in 0.03s.
|
||||||
|
> The table is retained only as the historical record that motivated the gate. The remaining
|
||||||
|
> per-file cost is the redundant `el_runtime.c` rebuild, which §9's compile-once architecture
|
||||||
|
> addresses.
|
||||||
|
|
||||||
|
Per-file `elc` time across the existing suite:
|
||||||
|
|
||||||
|
| File | Bytes | elc time |
|
||||||
|
|---|---|---|
|
||||||
|
| `test_compiler` | 29,685 (+394 KB of imports) | **3.58s** |
|
||||||
|
| `string_test` | 18,545 | 0.01s |
|
||||||
|
| all other 9 files | 2.2–10 KB | 0.00s |
|
||||||
|
|
||||||
|
Two distinct defects in two distinct regimes:
|
||||||
|
|
||||||
|
1. **`test_compiler.el` imports all five compiler sources** — 394 KB in one translation unit. Its
|
||||||
|
3.58s is entirely the quadratic. It is the only file where the quadratic bites.
|
||||||
|
2. **Every other file's cost is 100% redundant `el_runtime.c` rebuilds** — 480 KB of identical C,
|
||||||
|
recompiled once per test file.
|
||||||
|
|
||||||
|
Neither is fixed by making the compiler faster. Both are fixed by the architecture below, and the
|
||||||
|
speedup is a by-product of building it correctly, not the goal.
|
||||||
|
|
||||||
|
### 1.4 The asset worth keeping
|
||||||
|
|
||||||
|
`codegen.el:3651-3652` already collects `test_names` / `test_c_names` — **the compiler already does
|
||||||
|
compile-time test discovery.** It then discards that registry into a hardcoded `main()`.
|
||||||
|
|
||||||
|
That registry is precisely the seam Go's `_testmain.go` and Rust's `test_main_static` are built on.
|
||||||
|
The mechanism we need is half-built and wired to the wrong thing.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Grounding — the common spine of excellent frameworks
|
||||||
|
|
||||||
|
Researched from primary sources: Go `testing`/`go test`, Rust `libtest`/Criterion, JUnit 5 Platform,
|
||||||
|
NUnit 3, JMH, Google Benchmark. Six invariants hold across all of them.
|
||||||
|
|
||||||
|
1. **A registry is built before execution** — `(name, metadata, fn-ptr)` triples. Go generates it
|
||||||
|
from an AST scan; Rust synthesizes it in a compiler pass; JMH emits it as a build-time resource;
|
||||||
|
JUnit/NUnit build it reflectively. **Reflection is an implementation of the registry on runtimes
|
||||||
|
where it is cheap. It is never the architecture.**
|
||||||
|
|
||||||
|
2. **Discovery strictly precedes execution.** Every good capability — filtering, listing, counting,
|
||||||
|
sharding, IDE trees, re-run-failed-only, dry runs — is a consequence of this ordering.
|
||||||
|
|
||||||
|
3. **A hierarchy with stable, path-shaped unique IDs.** `TestFoo/subcase_2`. Selection is regex over
|
||||||
|
that path, one pattern per level.
|
||||||
|
|
||||||
|
4. **The framework is a prebuilt library; only the entry point is generated.** "Compile once, link
|
||||||
|
many" is always: framework archive compiled once + a small generated table + one
|
||||||
|
`MainStart(deps, registry)` call. Nobody recompiles the harness per test file.
|
||||||
|
|
||||||
|
5. **Execution emits an event stream; reporters are downstream renderers.** Human text, NDJSON,
|
||||||
|
JUnit XML, TAP are all transforms of one event stream. Go's one architectural mistake is doing
|
||||||
|
this backwards — `test2json` parses human output, and has shipped bugs when user output contains
|
||||||
|
`--- PASS:`.
|
||||||
|
|
||||||
|
6. **A dependency-injection seam at the boundary.** Go's `testdeps.TestDeps` exists so `testing`
|
||||||
|
can avoid importing `regexp`, profilers, and coverage. The execution core knows nothing about
|
||||||
|
output formats.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Architecture
|
||||||
|
|
||||||
|
### 3.1 The seam
|
||||||
|
|
||||||
|
```
|
||||||
|
┌─────────────────────────────────────────────────────────────┐
|
||||||
|
│ user code: foo.el with test { } / bench { } blocks │
|
||||||
|
└───────────────────────────┬─────────────────────────────────┘
|
||||||
|
│ elc --test
|
||||||
|
▼
|
||||||
|
┌─────────────────────────────────────────────────────────────┐
|
||||||
|
│ generated C (per suite, tiny): │
|
||||||
|
│ __el_test_fn_0 .. _N lowered test/bench bodies │
|
||||||
|
│ __el_registry[] static table: name/kind/file/ │
|
||||||
|
│ line/tags/sizes/expected-O │
|
||||||
|
│ __el_dispatch(i) generated switch → body │
|
||||||
|
│ main() { return el_test_main(argc, argv); } │
|
||||||
|
└───────────────────────────┬─────────────────────────────────┘
|
||||||
|
│ cc + link (registry only)
|
||||||
|
▼
|
||||||
|
┌─────────────────────────────────────────────────────────────┐
|
||||||
|
│ libeltest.a — PREBUILT ONCE │
|
||||||
|
│ • el_runtime.o (the 480 KB, compiled once, ever) │
|
||||||
|
│ • eltest.o the runner, WRITTEN IN EL │
|
||||||
|
│ discovery view · filtering · execution · fixtures · │
|
||||||
|
│ timing · benchmark harness · curve fitting · reporters │
|
||||||
|
└─────────────────────────────────────────────────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
The framework is written in El, compiled to C once, archived. Per-suite compilation touches only
|
||||||
|
the generated registry. This is Go's model, and it is strictly better for us than Go's because we
|
||||||
|
own the compiler and already have the AST — no separate source-scanning pass is needed.
|
||||||
|
|
||||||
|
### 3.2 Why the runner is in El and the registry is in C
|
||||||
|
|
||||||
|
El has no closures and no first-class function pointers. The registry must therefore hold C function
|
||||||
|
pointers, and it is generated C.
|
||||||
|
|
||||||
|
The runner stays in El and reaches the registry through a small builtin surface — indices, not
|
||||||
|
pointers:
|
||||||
|
|
||||||
|
```
|
||||||
|
__el_reg_count() -> Int
|
||||||
|
__el_reg_name(i) -> String
|
||||||
|
__el_reg_file(i) -> String
|
||||||
|
__el_reg_line(i) -> Int
|
||||||
|
__el_reg_kind(i) -> Int // 0=test 1=bench
|
||||||
|
__el_reg_tags(i) -> Int
|
||||||
|
__el_reg_sizes(i) -> String // JSON array, empty for tests
|
||||||
|
__el_reg_expect(i) -> Int // complexity class enum, 0 = none
|
||||||
|
__el_reg_invoke(i) -> Int // runs the body via the generated switch
|
||||||
|
```
|
||||||
|
|
||||||
|
Nine builtins. Everything else — filtering, lifecycle, statistics, curve fitting, all reporters —
|
||||||
|
is El. That satisfies "written in El" without pretending El can do something it cannot.
|
||||||
|
|
||||||
|
### 3.3 Result model
|
||||||
|
|
||||||
|
The unit is a **result record**, not a counter:
|
||||||
|
|
||||||
|
```
|
||||||
|
TestResult {
|
||||||
|
id String // slash path: "parser/handles_empty_input/case_3"
|
||||||
|
file String
|
||||||
|
line Int
|
||||||
|
status Status // Pass | Fail | Error | Skip
|
||||||
|
duration Int // nanoseconds, ALWAYS populated
|
||||||
|
message String // assertion detail: expected vs actual
|
||||||
|
output String // captured stdout/stderr for this test
|
||||||
|
assertions Int
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`Fail` = an assertion failed. `Error` = unexpected crash/abort. This distinction is load-bearing —
|
||||||
|
every CI consumer depends on it, and the JUnit XML schema encodes it as distinct elements.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Authoring surface
|
||||||
|
|
||||||
|
### 4.1 Tests
|
||||||
|
|
||||||
|
`test { }` already exists. Keep it. Add subtests and hierarchy:
|
||||||
|
|
||||||
|
```el
|
||||||
|
test "parser/empty input" {
|
||||||
|
assert_that(parse(""), is_err())
|
||||||
|
}
|
||||||
|
|
||||||
|
test "parser/table" {
|
||||||
|
for case in [["", 0], ["a", 1], ["a b", 2]] {
|
||||||
|
subtest(case[0]) {
|
||||||
|
assert_that(token_count(case[0]), equals(case[1]))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Subtest IDs compose as `parser/table/a_b`. Filtering is `--run 'parser/table/.*'`, one regex per
|
||||||
|
path segment, exactly as Go does.
|
||||||
|
|
||||||
|
**We do not build a parameterized-test annotation system.** Table-driven loops plus subtests subsume
|
||||||
|
`@ParameterizedTest`, `@MethodSource`, `@CsvSource`, and `TestCaseSource` entirely, at zero framework
|
||||||
|
surface. This is Go's single biggest ergonomic win over JUnit and NUnit.
|
||||||
|
|
||||||
|
### 4.2 Fixtures
|
||||||
|
|
||||||
|
Per-file and per-test only, plus a LIFO cleanup stack:
|
||||||
|
|
||||||
|
```el
|
||||||
|
setup_all { ... } // once per suite
|
||||||
|
setup { ... } // before each test
|
||||||
|
teardown { ... } // after each test
|
||||||
|
teardown_all { ... }
|
||||||
|
```
|
||||||
|
|
||||||
|
and inside a test, `cleanup { ... }` registering LIFO-ordered teardown.
|
||||||
|
|
||||||
|
**We do not build JUnit 5's extension SPI** — seventeen callback interfaces, hierarchical stores,
|
||||||
|
registration ordering rules. That complexity is the price of retrofitting a plugin ecosystem onto a
|
||||||
|
twenty-year-old reflective framework. Go's `t.Cleanup` covers roughly 90% of what `@AfterEach` is
|
||||||
|
used for at a fraction of the surface.
|
||||||
|
|
||||||
|
### 4.3 Assertions — constraint model
|
||||||
|
|
||||||
|
One entry point, composable constraint values (NUnit's model, which avoids the N² overload
|
||||||
|
explosion):
|
||||||
|
|
||||||
|
```el
|
||||||
|
assert_that(actual, equals(expected))
|
||||||
|
assert_that(xs, has_length(3))
|
||||||
|
assert_that(s, contains("foo").and(starts_with("bar")))
|
||||||
|
assert_that(f, is_within(0.01).of(3.14))
|
||||||
|
```
|
||||||
|
|
||||||
|
A constraint is a value with `apply_to(actual) -> ConstraintResult`, and the result knows how to
|
||||||
|
describe its own failure. Custom constraints are ordinary user types.
|
||||||
|
|
||||||
|
**Every failure message must name file, line, the expression text, and both values.** We capture
|
||||||
|
expression source text at compile time — we have the AST, so we can do this better than any
|
||||||
|
runtime-introspection framework.
|
||||||
|
|
||||||
|
Legacy `assert_true` / `assert_eq` / etc. stay as thin wrappers for migration.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Benchmarks
|
||||||
|
|
||||||
|
### 5.1 The loop
|
||||||
|
|
||||||
|
Adopt `b.Loop()`, not `b.N`. Go spent fifteen years on `b.N` before concluding `b.Loop` was right;
|
||||||
|
we skip that.
|
||||||
|
|
||||||
|
```el
|
||||||
|
bench "str_concat" {
|
||||||
|
let s = make_input(bench_n())
|
||||||
|
for bench_loop() {
|
||||||
|
black_box(str_concat(s, "x"))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Three properties that make this the correct choice for a C target:
|
||||||
|
|
||||||
|
1. **The timer auto-resets on first call**, so setup above the loop is excluded *by construction*
|
||||||
|
rather than by the author remembering `ResetTimer`.
|
||||||
|
2. **`N` is hidden**, so it cannot be misused.
|
||||||
|
3. **The harness owns the loop shape**, which lets us insert an optimization barrier the C compiler
|
||||||
|
cannot see through. `black_box(v)` lowers to `asm volatile("" :: "r"(&v) : "memory")`. Since we
|
||||||
|
emit a single translation unit, dead-code elimination of a benchmark body is a live hazard —
|
||||||
|
this is our version of JMH's `Blackhole` problem, solved in the harness rather than delegated to
|
||||||
|
the user.
|
||||||
|
|
||||||
|
### 5.2 Iteration scaling
|
||||||
|
|
||||||
|
Use Go's `predictN` heuristics verbatim. They are battle-tested and cheap:
|
||||||
|
|
||||||
|
```
|
||||||
|
n = goal_ns * prev_iters / prev_ns // multiply before divide — precision on sub-ns ops
|
||||||
|
n += n / 5 // 20% headroom, overshoot rather than re-loop
|
||||||
|
n = min(n, 100 * last) // never grow more than 100× per step
|
||||||
|
n = max(n, last + 1) // guarantee forward progress
|
||||||
|
n = min(n, 1_000_000_000) // hard ceiling
|
||||||
|
```
|
||||||
|
|
||||||
|
Report `n` rounded to 1/2/3/5 × 10ᵏ so runs are comparable.
|
||||||
|
|
||||||
|
### 5.3 Sampling
|
||||||
|
|
||||||
|
Criterion's shape, because it is correct near timer resolution:
|
||||||
|
|
||||||
|
- **Warmup**: iteration counts 1, 2, 4, 8… until cumulative time exceeds the warmup budget.
|
||||||
|
- **Measurement**: collect `sample_size` samples at iteration counts `[d, 2d, 3d, …, Nd]`.
|
||||||
|
- **Estimate**: slope of a linear regression of iteration-count vs elapsed time. The intercept
|
||||||
|
absorbs fixed overhead.
|
||||||
|
- **Time whole samples, never individual iterations.** This is the single most important detail —
|
||||||
|
it defeats timer-resolution error on nanosecond operations.
|
||||||
|
|
||||||
|
Outliers classified by modified Tukey (±1.5 IQR mild, ±3 IQR severe), **reported but retained**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Complexity gating — the centerpiece
|
||||||
|
|
||||||
|
This is the part that makes the quadratic fixable, and the part nobody in the mainstream has
|
||||||
|
finished. Google Benchmark's `Complexity()` fits the curve and *reports* it. We declare it and
|
||||||
|
**gate** on it.
|
||||||
|
|
||||||
|
### 6.1 Surface
|
||||||
|
|
||||||
|
```el
|
||||||
|
bench "elc_compile" over n in [16, 32, 64, 128, 256, 512, 1024] expect O(n) {
|
||||||
|
let src = synth_source(bench_n())
|
||||||
|
for bench_loop() { black_box(compile(src)) }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Alternative with no new syntax, if the parser change is judged too invasive — `bench_sizes([...])`
|
||||||
|
and `bench_expect("O(n)")` as calls inside the block. **Recommendation: declarative.** Runtime calls
|
||||||
|
mean `--list` cannot show the invariant without executing, which breaks the discovery-precedes-
|
||||||
|
execution invariant from §2.
|
||||||
|
|
||||||
|
### 6.2 Fitting
|
||||||
|
|
||||||
|
Per Google Benchmark `src/complexity.cc`. For candidate curves
|
||||||
|
`{O(1), O(log n), O(n), O(n log n), O(n²), O(n³)}`, one-parameter least squares, no intercept:
|
||||||
|
|
||||||
|
```
|
||||||
|
coef = Σ(tᵢ · gᵢ) / Σ(gᵢ²)
|
||||||
|
rms = sqrt( Σ(tᵢ − coef·gᵢ)² / k ) / mean(t) // normalized
|
||||||
|
```
|
||||||
|
|
||||||
|
Best fit = lowest normalized RMS. User-supplied lambda curves also supported.
|
||||||
|
|
||||||
|
### 6.3 Gate logic
|
||||||
|
|
||||||
|
1. **FAIL** if the best-fit curve is strictly worse than declared, ordering
|
||||||
|
`O(1) < O(log n) < O(n) < O(n log n) < O(n²) < O(n³)`. Print the fitted coefficient and the full
|
||||||
|
per-size table.
|
||||||
|
2. **FAIL** if the declared curve's normalized RMS exceeds a threshold (start at 0.10). This catches
|
||||||
|
the case where *no* candidate fits — noise, a cache cliff, or a phase change. Report
|
||||||
|
`INDETERMINATE` honestly rather than gating on garbage.
|
||||||
|
3. **WARN** if the best fit is strictly better than declared — either an optimization landed and the
|
||||||
|
annotation should tighten, or the sweep is too narrow to expose real behaviour.
|
||||||
|
4. **REFUSE to gate** on fewer than 5 distinct sizes spanning under 2 decades, geometrically spaced.
|
||||||
|
Say so loudly rather than producing a meaningless fit.
|
||||||
|
|
||||||
|
### 6.4 Why gate on the exponent, not wall-clock
|
||||||
|
|
||||||
|
- **Machine-independent.** The fitted exponent is a property of the algorithm; the coefficient is a
|
||||||
|
property of the machine. Gating on the exponent makes CI hardware heterogeneity, noisy neighbours,
|
||||||
|
and thermal throttling irrelevant — they scale `coef`, not `g`.
|
||||||
|
- **No stored baseline.** No artifact storage, no golden-file drift. The invariant lives in the
|
||||||
|
source next to the code and is reviewed in the same PR.
|
||||||
|
- **It catches the failure mode that actually ships.** An O(n) lookup inside an O(n) loop is
|
||||||
|
invisible at n=100 in a unit test and catastrophic at n=100,000 in production. Constant-factor
|
||||||
|
regressions are annoying. Complexity regressions are outages. Ours was a 27 GB outage.
|
||||||
|
|
||||||
|
### 6.5 The deterministic gate — the one that would have caught us
|
||||||
|
|
||||||
|
Wall-clock needs statistics. **Allocation counts do not.** They are perfectly deterministic.
|
||||||
|
|
||||||
|
> **Correction, 2026-08-16 — count alone is NOT sufficient. Gate on BOTH count and bytes.**
|
||||||
|
>
|
||||||
|
> Measured against two El programs, one allocating once per item and one rebuilding its
|
||||||
|
> accumulator each iteration:
|
||||||
|
>
|
||||||
|
> | n | linear allocs / bytes | quadratic allocs / bytes |
|
||||||
|
> |---|---|---|
|
||||||
|
> | 100 | 100 / 290 | 100 / 5,150 |
|
||||||
|
> | 200 | 200 / 690 | 200 / 20,300 |
|
||||||
|
> | 400 | 400 / 1,490 | 400 / 80,600 |
|
||||||
|
> | 800 | 800 / 3,090 | 800 / 321,200 |
|
||||||
|
>
|
||||||
|
> The quadratic program's allocation **count is exactly linear** — 100/200/400/800, identical to
|
||||||
|
> the healthy program. A count-only gate passes it clean. **Bytes** catch it: each doubling of n
|
||||||
|
> quadruples bytes (ratios 3.94, 3.97, 3.99 → 4.0 = O(n²)) where the linear program converges
|
||||||
|
> on 2.0.
|
||||||
|
>
|
||||||
|
> This is precisely elc's own defect shape — a copy-on-write accumulator reallocating once per
|
||||||
|
> pass (count linear) into a proportionally larger buffer (bytes quadratic).
|
||||||
|
>
|
||||||
|
> Therefore `expect allocs O(n)` **fits count and bytes independently and fails if EITHER exceeds
|
||||||
|
> the declared curve**, reporting which signal broke. "count linear, bytes quadratic" is a precise,
|
||||||
|
> directly actionable diagnosis.
|
||||||
|
>
|
||||||
|
> **`el_peak_rss()` is CONTEXT ONLY — never gate on it.** It is perturbed by the allocator and by
|
||||||
|
> the page cache. Allocation volume is the invariant; RSS and malloc/free churn are merely the two
|
||||||
|
> surfaces it shows on. The old shipped compiler paid the same quadratic in RSS that the rebuilt
|
||||||
|
> one pays in churn.
|
||||||
|
>
|
||||||
|
> **Measure rate, not level.** A guard reading swap *level* saw 97% on a thrashing host and 97% on
|
||||||
|
> a healthy one; only *rate* separated them. A growth exponent is a rate; a single measurement is
|
||||||
|
> a level. That is why the gate fits a curve across a sweep instead of comparing one number to a
|
||||||
|
> threshold.
|
||||||
|
|
||||||
|
> **Second correction, same day — THE ALLOCATION GATE ALONE WOULD HAVE MISSED THE REAL BUG.**
|
||||||
|
>
|
||||||
|
> el #132 found the actual elc quadratic: `strlen()` called inside `str_char_code()` and
|
||||||
|
> `str_slice()`, so the lexer rescanned the remaining input on every character. Pure CPU.
|
||||||
|
> **Zero allocation.** `str_char_code` is a bounds check and an index — it allocates nothing.
|
||||||
|
>
|
||||||
|
> Measured on three controlled specimens (`lang/.work/fitprobe.el`), growth ratio per doubling of
|
||||||
|
> n across n = 200/400/800/1600:
|
||||||
|
>
|
||||||
|
> | specimen | allocs | bytes | time | what it proves |
|
||||||
|
> |---|---|---|---|---|
|
||||||
|
> | `linear` — one alloc per item | 2.00 2.00 2.00 → **O(n)** | 2.16 2.07 2.23 → **O(n)** | 0.83 2.00 2.05 → **O(n)** | clean baseline |
|
||||||
|
> | `accum` — rebuilds accumulator | 2.00 2.00 2.00 → **O(n)** | 3.97 3.99 3.99 → **O(n²)** | noisy | count misses, **bytes catches** |
|
||||||
|
> | `compute` — n scans over n chars | 0 → **FLAT** | 0 → **FLAT** | 3.93 4.01 3.96 → **O(n²)** | **both alloc signals blind; only time catches** |
|
||||||
|
>
|
||||||
|
> `compute` is el #132's shape exactly. A gate fitting only allocation count and bytes classifies
|
||||||
|
> it as FLAT and passes it. **The gate as originally specified would not have caught the defect it
|
||||||
|
> was created for.**
|
||||||
|
>
|
||||||
|
> Therefore the gate fits **THREE** signals and fails if ANY exceeds its declared curve:
|
||||||
|
>
|
||||||
|
> ```
|
||||||
|
> bench "elc_compile" over n in [...] expect time O(n) allocs O(n) bytes O(n) { ... }
|
||||||
|
> ```
|
||||||
|
>
|
||||||
|
> - **allocs (count)** — deterministic, zero-noise. Catches per-item allocation growth.
|
||||||
|
> - **allocs (bytes)** — deterministic, zero-noise. Catches accumulator-rebuild quadratics that
|
||||||
|
> count cannot see.
|
||||||
|
> - **time** — noisy, needs the sweep and statistics. The ONLY signal that sees pure-compute
|
||||||
|
> complexity regressions. Gate on the fitted *exponent*, never on absolute duration, so CI
|
||||||
|
> hardware variance scales the coefficient and leaves the classification intact.
|
||||||
|
>
|
||||||
|
> The deterministic signals remain preferable where they apply — they need no statistics and are
|
||||||
|
> correct on the first run. They are simply not sufficient.
|
||||||
|
>
|
||||||
|
> **`black_box` is mandatory, and consuming the result is NOT enough.** The first version of
|
||||||
|
> `compute` accumulated `total + 1` in a nested loop and reported **0 µs at every n** while
|
||||||
|
> returning a numerically correct n². Clang recognised the idiom and closed the loop to a
|
||||||
|
> multiply. Feeding the result into output did not prevent it. Only making the inner operation an
|
||||||
|
> opaque external call restored the real curve. A benchmark harness that trusts the user to defeat
|
||||||
|
> the optimiser will silently measure nothing — and report success while doing it.
|
||||||
|
|
||||||
|
Instrument the runtime with allocation counters and fit *those* against n instead of time:
|
||||||
|
|
||||||
|
```el
|
||||||
|
bench "elc_compile" over n in [...] expect O(n) allocs O(n) { ... }
|
||||||
|
```
|
||||||
|
|
||||||
|
Zero noise, zero statistics, always gateable, correct on the first run on any machine. Go reports
|
||||||
|
`allocs/op` and `B/op`; **nobody fits them against n.** That is an open opportunity and it is exactly
|
||||||
|
our bug: elc's defect is quadratic *allocation volume*, which the old binary paid in RSS and the
|
||||||
|
current source pays in malloc/free churn.
|
||||||
|
|
||||||
|
An `expect allocs O(n)` assertion on `elc`'s compile path would have failed the build the day the
|
||||||
|
quadratic was introduced.
|
||||||
|
|
||||||
|
Required runtime additions: `__el_alloc_count()`, `__el_alloc_bytes()`, `__el_peak_rss()`.
|
||||||
|
|
||||||
|
### 6.6 Constant-factor gate (secondary, opt-in)
|
||||||
|
|
||||||
|
Mann-Whitney U at α = 0.05, noise floor 1%, medians with 95% CIs, `~` for not-significant. Requires
|
||||||
|
`--count >= 9`. Off by default on CI; opt-in per benchmark.
|
||||||
|
|
||||||
|
**Exit nonzero on regression.** Both benchstat and Criterion always exit 0, which is why every shop
|
||||||
|
using them wrote a wrapper. We do not repeat that omission.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Output
|
||||||
|
|
||||||
|
**Structured events are the source of truth.** Human text is rendered from them. We do not repeat
|
||||||
|
Go's parse-the-human-output design.
|
||||||
|
|
||||||
|
Event stream, NDJSON, one object per line, streamed live:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"time":"...","action":"run","test":"parser/empty"}
|
||||||
|
{"time":"...","action":"output","test":"parser/empty","output":"..."}
|
||||||
|
{"time":"...","action":"pass","test":"parser/empty","elapsed":0.0031}
|
||||||
|
{"time":"...","action":"bench","test":"str_concat","n":1024,"ns_op":41.2,"allocs_op":3,"bigo":"N","rms":0.03}
|
||||||
|
```
|
||||||
|
|
||||||
|
Renderers, all downstream and pluggable:
|
||||||
|
|
||||||
|
| Format | Flag | Use |
|
||||||
|
|---|---|---|
|
||||||
|
| Human | default | terminal, **per-test duration always shown** |
|
||||||
|
| NDJSON | `--json` | tooling, history, flaky detection |
|
||||||
|
| JUnit XML | `--junit-xml=PATH` | every CI system on earth |
|
||||||
|
| TAP | `--tap` | optional |
|
||||||
|
|
||||||
|
JUnit XML per the de-facto schema: `testsuites` → `testsuite` → `testcase`, with `time` in seconds
|
||||||
|
as a decimal, `file`/`line` attributes, and `failure` vs `error` vs `skipped` as distinct child
|
||||||
|
elements. Absence of a child element means pass. Emit `<testsuites>` even for a single suite, and
|
||||||
|
parse both shapes on input.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. CLI
|
||||||
|
|
||||||
|
```
|
||||||
|
--list print the registry, run nothing
|
||||||
|
--list-json machine-readable registry
|
||||||
|
--run PATTERN slash-separated regex per path segment
|
||||||
|
--tag EXPR tag expression: fast & !slow
|
||||||
|
--shard I/N deterministic sharding for CI parallelism
|
||||||
|
--count N repetitions, for statistics
|
||||||
|
--bench PATTERN run benchmarks (off by default in test runs)
|
||||||
|
--benchtime DUR per-benchmark time budget
|
||||||
|
--junit-xml PATH
|
||||||
|
--json
|
||||||
|
--isolate re-exec per test on crash, so one SIGSEGV doesn't lose the run
|
||||||
|
--timeout DUR
|
||||||
|
--fail-fast
|
||||||
|
```
|
||||||
|
|
||||||
|
`--list` / `--list-json` / `--shard` cost roughly thirty lines because the registry already exists
|
||||||
|
before `main` does anything. That is the dividend of discovery-precedes-execution.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Build model
|
||||||
|
|
||||||
|
```
|
||||||
|
# once, ever (or when the runtime/framework changes):
|
||||||
|
cc -c el_runtime.c -o el_runtime.o
|
||||||
|
elc eltest.el > eltest.c && cc -c eltest.c -o eltest.o
|
||||||
|
ar rcs libeltest.a el_runtime.o eltest.o
|
||||||
|
|
||||||
|
# per suite:
|
||||||
|
elc --test foo_test.el > foo_test.c # registry + bodies only
|
||||||
|
cc foo_test.c libeltest.a -o foo_test
|
||||||
|
```
|
||||||
|
|
||||||
|
The 0.14s × N of redundant runtime rebuilds disappears — not because we optimized it, but because
|
||||||
|
one-runner-over-many-suites requires compile-once-link-many as a structural precondition.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Bootstrap and self-hosting
|
||||||
|
|
||||||
|
The framework's own tests are `test { }` blocks run by the framework. Same fixpoint discipline the
|
||||||
|
compiler already applies to itself.
|
||||||
|
|
||||||
|
1. Build the framework using the *existing* harness for its first tests (stage 0).
|
||||||
|
2. Rebuild the framework's tests as `test { }` blocks run by the new runner (stage 1).
|
||||||
|
3. Verify stage 1 reports identical results to stage 0.
|
||||||
|
4. From then on, the framework is tested by itself.
|
||||||
|
|
||||||
|
A framework that cannot run its own suite is not evidence of anything. This is a correctness proof,
|
||||||
|
not a claim.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Explicitly not building
|
||||||
|
|
||||||
|
| Rejected | Why |
|
||||||
|
|---|---|
|
||||||
|
| Naming-convention discovery (`fn test_foo`) | `test { }` is a real declaration. Go's `TestXxx` exists only because Go had no better hook — and it needs a heuristic to avoid matching `TesticularCancer`. |
|
||||||
|
| Reflection or symbol-table scanning | Slow, fragile under LTO/strip/dead-strip, and unnecessary when we own the compiler. |
|
||||||
|
| Parsing human output into structure | Go's `test2json` is its one clear architectural mistake. |
|
||||||
|
| JUnit 5's extension SPI | Seventeen callback interfaces to retrofit plugins onto a reflective framework. Not our problem. |
|
||||||
|
| `@ParameterizedTest` machinery | Table-driven loops + subtests subsume it at zero surface. |
|
||||||
|
| NUnit's out-of-process agents | They bridge CLR versions and AppDomains. We emit one native binary. Keep `--isolate` as crash fallback only. |
|
||||||
|
| JMH-style forking by default | Forks exist because JIT profiles are per-process. AOT C has no such state. Keep `--fork` available, not default. |
|
||||||
|
| Exit 0 on regression | benchstat and Criterion both do this, and every user writes a wrapper. |
|
||||||
|
| Dynamic runtime test registration | Breaks `--list`, sharding, and individual selection. Registry stays static. |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Phasing
|
||||||
|
|
||||||
|
| Phase | Content | Gate |
|
||||||
|
|---|---|---|
|
||||||
|
| **1** | Registry emission in codegen; 9 builtins; `el_test_main` skeleton in El; result records; per-test timing; human + NDJSON output | existing 11 test files pass, with timing |
|
||||||
|
| **2** | `libeltest.a` build model; subtests; filtering; `--list`; fixtures; constraint assertions; JUnit XML | suite runs in one binary; runtime compiled once |
|
||||||
|
| **3** | `bench { }`, `bench_loop`, `black_box`, `predictN`, Criterion sampling | benchmarks produce stable ns/op |
|
||||||
|
| **4** | Allocation counters; complexity fitting; `expect O(...)` gate | **an `expect allocs O(n)` benchmark on `elc` fails on the current quadratic** |
|
||||||
|
| **5** | Migrate both legacy systems; delete `runtime/test.el`; self-host | framework runs its own suite |
|
||||||
|
|
||||||
|
Phase 4 is the deliverable that matters. Phases 1–3 exist to make it possible.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Open questions for review
|
||||||
|
|
||||||
|
1. **Declarative `over n in [...] expect O(...)` syntax vs runtime calls.** I recommend declarative
|
||||||
|
(§6.1) so `--list` can show invariants without executing. It costs parser work. Your call.
|
||||||
|
2. **`bench { }` as a new block form** — parallel to `test { }`, or a modifier on it?
|
||||||
|
3. **Scope of the constraint model.** Full composable constraints, or start with a flat assertion set
|
||||||
|
and add constraints later? Full model is more surface but avoids a second migration.
|
||||||
|
4. **Does `runtime/test.el` get deleted or kept as a deprecated shim?** I lean delete — two systems
|
||||||
|
is how we got here.
|
||||||
|
5. **Where does `libeltest.a` live** in the tree, and does `epm` need to know about it?
|
||||||
|
6. **Allocation counters in `el_seed.c` or `el_runtime.c`?** AGENTS.md says `el_seed.c` is the sole
|
||||||
|
C dependency and hand-maintained; counters are OS-boundary-adjacent but not OS calls.
|
||||||
|
7. **Is per-test timing enough, or do we want per-*assertion* timing** for finding slow helpers?
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. What this document is not
|
||||||
|
|
||||||
|
This is a design, not a measurement. Every performance claim about the *current* system in §1 is
|
||||||
|
measured and reproducible in this worktree. Every claim about the *proposed* system is a prediction.
|
||||||
|
None of it is verified until Phase 1 runs and Phase 4 fails a build on the real quadratic.
|
||||||
+145
-37
@@ -10,10 +10,60 @@
|
|||||||
// cc -std=c11 -O2 -lcurl -lpthread -o engram server.c el_runtime.c
|
// cc -std=c11 -O2 -lcurl -lpthread -o engram server.c el_runtime.c
|
||||||
// ./engram
|
// ./engram
|
||||||
//
|
//
|
||||||
// Configuration via environment:
|
// Configuration is DECLARED, not scattered. See the `program` block below:
|
||||||
// ENGRAM_BIND — host:port (default :8742)
|
// every knob's type and default lives there and nowhere else, is resolved from
|
||||||
// ENGRAM_API_KEY — bearer auth (optional)
|
// the environment (env wins, declaration is the fallback) and validated before
|
||||||
// ENGRAM_DATA_DIR — snapshot location (default ~/.neuron/engram)
|
// any statement of this file runs. Read one with config("NAME") -> String.
|
||||||
|
//
|
||||||
|
// The one deliberate exception is ENGRAM_DATA_DIR — see the note in the block.
|
||||||
|
|
||||||
|
// ── Program declaration (cross-cutting concerns) ──────────────────────────────
|
||||||
|
//
|
||||||
|
// singleton: two engram processes against one data dir is data loss, not a
|
||||||
|
// warning. The runtime takes an exclusive flock at startup and a second start
|
||||||
|
// is refused loudly with the holder's pid.
|
||||||
|
//
|
||||||
|
// NOT declared here, on purpose: ENGRAM_DATA_DIR. Its resolution is owned by
|
||||||
|
// engram_resolve_data_dir() (el_runtime.c), which defaults to $HOME/.neuron/engram
|
||||||
|
// and fails LOUD rather than silently persisting to an ephemeral directory.
|
||||||
|
// Declaring a default for it here as well would put the data dir's fallback in
|
||||||
|
// two places — which is precisely the defect this migration removes (until
|
||||||
|
// 2026-08-15 the reseed backup path carried its own "/tmp/engram" default that
|
||||||
|
// disagreed with the resolver, so the pre-destructive safety copy landed in /tmp).
|
||||||
|
// HOME is likewise not declared: it is a genuine environment read, not a knob.
|
||||||
|
program "engram" {
|
||||||
|
singleton: "engram"
|
||||||
|
|
||||||
|
// ── Core server ──
|
||||||
|
env ENGRAM_BIND: String = ":8742"
|
||||||
|
// Default "" leaves auth DISABLED (check_auth_ok short-circuits to true on an
|
||||||
|
// empty key). That is the pre-existing behaviour and is deliberately preserved
|
||||||
|
// here; making this `required` is the obvious hardening follow-up, but it is a
|
||||||
|
// behaviour change and out of scope for this migration.
|
||||||
|
env ENGRAM_API_KEY: String = ""
|
||||||
|
|
||||||
|
// ── Feature flags (bool-ish Strings; the predicate fns below own truthiness) ──
|
||||||
|
env ENGRAM_STORE: String = "off"
|
||||||
|
env ENGRAM_WAL: String = "off"
|
||||||
|
env ENGRAM_AUTOCONNECT: String = "off"
|
||||||
|
env ENGRAM_ISE_OFFGRAPH: String = "off"
|
||||||
|
|
||||||
|
// ── ISE telemetry ──
|
||||||
|
env ENGRAM_ISE_RETENTION_MS: Int = "172800000"
|
||||||
|
|
||||||
|
// ── Guide (local Qwen3 via llama-server) ──
|
||||||
|
env GUIDE_ENABLE: String = "off"
|
||||||
|
env GUIDE_TIER_FORCE: String = ""
|
||||||
|
env GUIDE_CACHE_DIR: String = ""
|
||||||
|
env GUIDE_RAM_GB_4B: Int = "16"
|
||||||
|
env GUIDE_RAM_GB_1P7B: Int = "8"
|
||||||
|
env GUIDE_BACKEND: String = "llama-server"
|
||||||
|
env GUIDE_HOST: String = "127.0.0.1"
|
||||||
|
env GUIDE_PORT: Int = "8771"
|
||||||
|
env GUIDE_LLAMA_SERVER_BIN: String = "llama-server"
|
||||||
|
env GUIDE_NGL: Int = "99"
|
||||||
|
env GUIDE_CTX: Int = "4096"
|
||||||
|
}
|
||||||
|
|
||||||
// ── Helpers ───────────────────────────────────────────────────────────────────
|
// ── Helpers ───────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
@@ -133,7 +183,7 @@ fn route_text_health(method: String, path: String, body: String) -> String {
|
|||||||
// engram_store_enabled() in el_runtime.c EXACTLY (1 / on / true). Default off →
|
// engram_store_enabled() in el_runtime.c EXACTLY (1 / on / true). Default off →
|
||||||
// every persistence path below is byte-for-byte the historical snapshot behavior.
|
// every persistence path below is byte-for-byte the historical snapshot behavior.
|
||||||
fn store_on() -> Bool {
|
fn store_on() -> Bool {
|
||||||
let v: String = env("ENGRAM_STORE")
|
let v: String = config("ENGRAM_STORE")
|
||||||
if str_eq(v, "1") { return true }
|
if str_eq(v, "1") { return true }
|
||||||
if str_eq(v, "on") { return true }
|
if str_eq(v, "on") { return true }
|
||||||
if str_eq(v, "true") { return true }
|
if str_eq(v, "true") { return true }
|
||||||
@@ -162,7 +212,6 @@ fn persist_canonical() -> Int {
|
|||||||
if store_on() {
|
if store_on() {
|
||||||
return engram_store_checkpoint()
|
return engram_store_checkpoint()
|
||||||
}
|
}
|
||||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
|
||||||
let dir: String = engram_resolve_data_dir()
|
let dir: String = engram_resolve_data_dir()
|
||||||
// (2026-08-10 self-review) This returned a hardcoded 1, which made every
|
// (2026-08-10 self-review) This returned a hardcoded 1, which made every
|
||||||
// caller's `let saved: Int = persist_canonical()` a dead variable — six
|
// caller's `let saved: Int = persist_canonical()` a dead variable — six
|
||||||
@@ -176,7 +225,7 @@ fn persist_canonical() -> Int {
|
|||||||
// per-write full-snapshot behavior. When ON, structural mutations append O(1)
|
// per-write full-snapshot behavior. When ON, structural mutations append O(1)
|
||||||
// WAL records instead of rewriting the whole graph, with threshold compaction.
|
// WAL records instead of rewriting the whole graph, with threshold compaction.
|
||||||
fn wal_on() -> Bool {
|
fn wal_on() -> Bool {
|
||||||
str_eq(env("ENGRAM_WAL"), "on")
|
str_eq(config("ENGRAM_WAL"), "on")
|
||||||
}
|
}
|
||||||
|
|
||||||
// autoconnect_on — ENGRAM_AUTOCONNECT. Will's rule: "we shouldn't be inserting
|
// autoconnect_on — ENGRAM_AUTOCONNECT. Will's rule: "we shouldn't be inserting
|
||||||
@@ -184,7 +233,7 @@ fn wal_on() -> Bool {
|
|||||||
// edge (kNN over embeddings) so no content node enters the graph edgeless.
|
// edge (kNN over embeddings) so no content node enters the graph edgeless.
|
||||||
// Default OFF -> byte-identical to prior behavior (node created, no auto edges).
|
// Default OFF -> byte-identical to prior behavior (node created, no auto edges).
|
||||||
fn autoconnect_on() -> Bool {
|
fn autoconnect_on() -> Bool {
|
||||||
let v: String = env("ENGRAM_AUTOCONNECT")
|
let v: String = config("ENGRAM_AUTOCONNECT")
|
||||||
if str_eq(v, "1") { return true }
|
if str_eq(v, "1") { return true }
|
||||||
if str_eq(v, "on") { return true }
|
if str_eq(v, "on") { return true }
|
||||||
if str_eq(v, "true") { return true }
|
if str_eq(v, "true") { return true }
|
||||||
@@ -197,7 +246,7 @@ fn autoconnect_on() -> Bool {
|
|||||||
// separate state-event log tier instead of the node graph. Default OFF -> ISEs
|
// separate state-event log tier instead of the node graph. Default OFF -> ISEs
|
||||||
// remain graph nodes exactly as before (with 48h prune).
|
// remain graph nodes exactly as before (with 48h prune).
|
||||||
fn ise_offgraph_on() -> Bool {
|
fn ise_offgraph_on() -> Bool {
|
||||||
let v: String = env("ENGRAM_ISE_OFFGRAPH")
|
let v: String = config("ENGRAM_ISE_OFFGRAPH")
|
||||||
if str_eq(v, "1") { return true }
|
if str_eq(v, "1") { return true }
|
||||||
if str_eq(v, "on") { return true }
|
if str_eq(v, "on") { return true }
|
||||||
if str_eq(v, "true") { return true }
|
if str_eq(v, "true") { return true }
|
||||||
@@ -247,6 +296,24 @@ fn persist_bulk() -> Int {
|
|||||||
return persist_canonical()
|
return persist_canonical()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// COMPILER LANDMINE, measured 2026-08-16 — do not inline this back into the
|
||||||
|
// caller. elc lowers `a == b` to numeric comparison only when both operand
|
||||||
|
// NAMES are in the per-function int-name set, which `let x: Int` populates.
|
||||||
|
// That registration does NOT propagate into a nested if-expression block: the
|
||||||
|
// first cut of the geometry-ingest path wrote `let claimed: Int = ...` and
|
||||||
|
// `let got: Int = ...` inside the else-arm and `claimed == got` came out of
|
||||||
|
// codegen as `str_eq(claimed, got)` — strcmp on two integers reinterpreted as
|
||||||
|
// pointers, i.e. a segfault on the first geometry-bearing request. Read back
|
||||||
|
// out of the generated C, not guessed. Function PARAMETERS annotated `: Int`
|
||||||
|
// do register reliably (verified: `if (claimed == actual)`), so the comparison
|
||||||
|
// lives in a function of its own. Note also the explicit `return`s — a trailing
|
||||||
|
// if-EXPRESSION at a function tail emits as a statement and the function
|
||||||
|
// returns 0 regardless, which is the same probe's second finding.
|
||||||
|
fn width_agrees(claimed: Int, actual: Int) -> Int {
|
||||||
|
if claimed == actual { return 1 }
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
// INCOMPLETE-ROUTE FIX (2026-07-24 self-review): this route silently dropped
|
// INCOMPLETE-ROUTE FIX (2026-07-24 self-review): this route silently dropped
|
||||||
// label, importance, tier, and tags — engram_node() defaults label to content
|
// label, importance, tier, and tags — engram_node() defaults label to content
|
||||||
// and importance to 0.5, so every node created over HTTP lost its metadata.
|
// and importance to 0.5, so every node created over HTTP lost its metadata.
|
||||||
@@ -288,6 +355,45 @@ fn route_create_node(method: String, path: String, body: String) -> String {
|
|||||||
salience, importance, confidence,
|
salience, importance, confidence,
|
||||||
tier, tags
|
tier, tags
|
||||||
)
|
)
|
||||||
|
// GEOMETRY INGEST — geometry-valued end to end (2026-08-16).
|
||||||
|
//
|
||||||
|
// The defect this route originally had: it accepted an "emb" field,
|
||||||
|
// returned 200 with a fresh id, and stored NOTHING, because engram_node_full
|
||||||
|
// has no vector parameter. The consequence was structural, not cosmetic —
|
||||||
|
// text was the only entry medium, so any non-text modality had to be
|
||||||
|
// DESCRIBED in prose, and what we then reasoned over was the geometry of the
|
||||||
|
// description, not of the signal.
|
||||||
|
//
|
||||||
|
// #141 fixed the drop but marshalled the vector as a hex STRING through
|
||||||
|
// engram_node_set_emb, which put text back as the TRANSPORT medium one layer
|
||||||
|
// below the problem being fixed. This is that correction: hex is decoded
|
||||||
|
// exactly ONCE, here at the edge, into a first-class Geometry, and every
|
||||||
|
// step below this line moves geometry rather than text. An encoding at the
|
||||||
|
// boundary is what an encoding is for.
|
||||||
|
//
|
||||||
|
// The WIRE is deliberately unchanged — "emb" is still little-endian float32
|
||||||
|
// hex (8 chars per component), the encoding the perception vessel's
|
||||||
|
// /voice/embed already emits — because production clients speak it. What
|
||||||
|
// changed is underneath it.
|
||||||
|
//
|
||||||
|
// "dim" is now treated as an ASSERTION about the vector the caller sent, not
|
||||||
|
// as the source of its width: a Geometry carries its own width. A stated dim
|
||||||
|
// that disagrees is a REJECTED ingest, not a silent reinterpretation. Omitting
|
||||||
|
// "dim" is fine and means "trust the vector", which is the honest default.
|
||||||
|
//
|
||||||
|
// Off-dimension vectors remain stored but not inserted into the resident HNSW
|
||||||
|
// index (its build loop filters on emb_dim), so a 64-dim voice geometry is
|
||||||
|
// durable and addressable without perturbing the 768-dim canonical index.
|
||||||
|
let emb_hex: String = json_get_string(body, "emb")
|
||||||
|
let emb_set: Int = if str_eq(emb_hex, "") { 0 } else {
|
||||||
|
let g: Geometry = geometry_from_f32le_hex(emb_hex)
|
||||||
|
let got: Int = geometry_dim(g)
|
||||||
|
let dim_raw: String = json_get_raw(body, "dim")
|
||||||
|
let claimed: Int = if str_eq(dim_raw, "") { got } else { json_get_int(body, "dim") }
|
||||||
|
let landed: Int = if width_agrees(claimed, got) > 0 { node_attach_geometry(id, g) } else { 0 }
|
||||||
|
let freed: Int = geometry_free(g)
|
||||||
|
landed
|
||||||
|
}
|
||||||
let saved: Int = persist_node(id)
|
let saved: Int = persist_node(id)
|
||||||
// ORPHAN PREVENTION (ENGRAM_AUTOCONNECT): connect the fresh node to its
|
// ORPHAN PREVENTION (ENGRAM_AUTOCONNECT): connect the fresh node to its
|
||||||
// nearest embedded neighbors so it never enters the graph edgeless.
|
// nearest embedded neighbors so it never enters the graph edgeless.
|
||||||
@@ -298,7 +404,11 @@ fn route_create_node(method: String, path: String, body: String) -> String {
|
|||||||
if added > 0 { let sv2: Int = persist_edges_since(ec0) }
|
if added > 0 { let sv2: Int = persist_edges_since(ec0) }
|
||||||
added
|
added
|
||||||
} else { 0 }
|
} else { 0 }
|
||||||
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\",\"connected\":" + int_to_str(connected) + "}"
|
// Report whether the supplied geometry actually landed. The old response
|
||||||
|
// was success-shaped regardless — 200 with an id while the vector was
|
||||||
|
// discarded — which is how the drop went unnoticed. A caller can now
|
||||||
|
// assert on emb_set instead of trusting the status code.
|
||||||
|
"{\"id\":\"" + id + "\",\"content\":\"" + content + "\",\"node_type\":\"" + node_type + "\",\"connected\":" + int_to_str(connected) + ",\"emb_set\":" + int_to_str(emb_set) + "}"
|
||||||
}
|
}
|
||||||
|
|
||||||
fn route_get_node(method: String, path: String, body: String) -> String {
|
fn route_get_node(method: String, path: String, body: String) -> String {
|
||||||
@@ -333,7 +443,6 @@ fn route_scan_nodes(method: String, path: String, body: String) -> String {
|
|||||||
// process ever booted with a partial/empty store, the first read request
|
// process ever booted with a partial/empty store, the first read request
|
||||||
// clobbered the good snapshot. Read routes must never write the canonical path.)
|
// clobbered the good snapshot. Read routes must never write the canonical path.)
|
||||||
fn route_scan_edges(method: String, path: String, body: String) -> String {
|
fn route_scan_edges(method: String, path: String, body: String) -> String {
|
||||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
|
||||||
let dir: String = engram_resolve_data_dir()
|
let dir: String = engram_resolve_data_dir()
|
||||||
let snap_path: String = dir + "/.scan-export.json"
|
let snap_path: String = dir + "/.scan-export.json"
|
||||||
engram_save(snap_path)
|
engram_save(snap_path)
|
||||||
@@ -494,7 +603,6 @@ fn route_forget(method: String, path: String, body: String) -> String {
|
|||||||
|
|
||||||
fn route_save(method: String, path: String, body: String) -> String {
|
fn route_save(method: String, path: String, body: String) -> String {
|
||||||
let p_raw: String = json_get_string(body, "path")
|
let p_raw: String = json_get_string(body, "path")
|
||||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
|
||||||
let dir: String = engram_resolve_data_dir()
|
let dir: String = engram_resolve_data_dir()
|
||||||
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
|
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
|
||||||
// (2026-08-10 self-review) engram_save returns 0 on an empty path and the
|
// (2026-08-10 self-review) engram_save returns 0 on an empty path and the
|
||||||
@@ -578,7 +686,6 @@ fn route_drift(method: String, path: String, body: String) -> String {
|
|||||||
|
|
||||||
fn route_load(method: String, path: String, body: String) -> String {
|
fn route_load(method: String, path: String, body: String) -> String {
|
||||||
let p_raw: String = json_get_string(body, "path")
|
let p_raw: String = json_get_string(body, "path")
|
||||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
|
||||||
let dir: String = engram_resolve_data_dir()
|
let dir: String = engram_resolve_data_dir()
|
||||||
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
|
let p: String = if str_eq(p_raw, "") { dir + "/snapshot.json" } else { p_raw }
|
||||||
// (2026-08-10 self-review) This was a stub response over the single most
|
// (2026-08-10 self-review) This was a stub response over the single most
|
||||||
@@ -649,7 +756,6 @@ fn route_embed_backfill(method: String, path: String, body: String) -> String {
|
|||||||
// (it skips nodes already present by ID). Auth-exempt: same-host internal call.
|
// (it skips nodes already present by ID). Auth-exempt: same-host internal call.
|
||||||
// (2026-06-27 self-review: added this route to fix silent 10-min sync failures)
|
// (2026-06-27 self-review: added this route to fix silent 10-min sync failures)
|
||||||
fn route_sync(method: String, path: String, body: String) -> String {
|
fn route_sync(method: String, path: String, body: String) -> String {
|
||||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
|
||||||
let dir: String = engram_resolve_data_dir()
|
let dir: String = engram_resolve_data_dir()
|
||||||
// 2026-07-21 self-review: export to a scratch path, never the canonical
|
// 2026-07-21 self-review: export to a scratch path, never the canonical
|
||||||
// snapshot.json — read routes must not be able to clobber the good snapshot.
|
// snapshot.json — read routes must not be able to clobber the good snapshot.
|
||||||
@@ -725,8 +831,12 @@ fn route_reseed_nodes(method: String, path: String, body: String) -> String {
|
|||||||
if str_eq(p, "") { return err_json("path is required") }
|
if str_eq(p, "") { return err_json("path is required") }
|
||||||
if str_eq(fs_read(p), "") { return err_json("file missing or empty") }
|
if str_eq(fs_read(p), "") { return err_json("file missing or empty") }
|
||||||
|
|
||||||
let dir_raw: String = env("ENGRAM_DATA_DIR")
|
// (2026-08-15) This site carried its own "/tmp/engram" fallback, which
|
||||||
let dir: String = if str_eq(dir_raw, "") { "/tmp/engram" } else { dir_raw }
|
// DISAGREED with engram_resolve_data_dir() ($HOME/.neuron/engram, fail-loud).
|
||||||
|
// The consumer is the pre-destructive backup below, so with ENGRAM_DATA_DIR
|
||||||
|
// unset the safety copy taken before a reseed landed in an ephemeral /tmp
|
||||||
|
// while the store it was protecting lived elsewhere. One owner, one answer.
|
||||||
|
let dir: String = engram_resolve_data_dir()
|
||||||
let backup: String = dir + "/.reseed-backup.json"
|
let backup: String = dir + "/.reseed-backup.json"
|
||||||
|
|
||||||
let replace_raw: String = json_get_raw(body, "replace")
|
let replace_raw: String = json_get_raw(body, "replace")
|
||||||
@@ -818,8 +928,7 @@ fn route_emit_ise(method: String, path: String, body: String) -> String {
|
|||||||
sal, imp, conf,
|
sal, imp, conf,
|
||||||
"Episodic", "[\"internal-state\",\"InternalStateEvent\"]"
|
"Episodic", "[\"internal-state\",\"InternalStateEvent\"]"
|
||||||
)
|
)
|
||||||
let ret_raw: String = env("ENGRAM_ISE_RETENTION_MS")
|
let ret_ms: Int = str_to_int(config("ENGRAM_ISE_RETENTION_MS"))
|
||||||
let ret_ms: Int = if str_eq(ret_raw, "") { 172800000 } else { str_to_int(ret_raw) }
|
|
||||||
let pruned: Int = engram_prune_telemetry(ret_ms)
|
let pruned: Int = engram_prune_telemetry(ret_ms)
|
||||||
"{\"ok\":true,\"id\":\"" + id + "\",\"pruned\":" + int_to_str(pruned) + "}"
|
"{\"ok\":true,\"id\":\"" + id + "\",\"pruned\":" + int_to_str(pruned) + "}"
|
||||||
}
|
}
|
||||||
@@ -1068,14 +1177,12 @@ fn route_correspondence_beat(method: String, path: String, body: String) -> Stri
|
|||||||
// turns native thinking ON: the response carries reasoning_content (the thinking)
|
// turns native thinking ON: the response carries reasoning_content (the thinking)
|
||||||
// alongside content (the answer).
|
// alongside content (the answer).
|
||||||
|
|
||||||
fn guide_env_or(key: String, dflt: String) -> String {
|
// (2026-08-15) guide_env_or(key, dflt) lived here. Its whole job was supplying a
|
||||||
let v: String = env(key)
|
// per-call-site default, which is now the program block's job — every GUIDE_* knob
|
||||||
if str_eq(v, "") { return dflt }
|
// is declared once at the top of this file and read straight through config().
|
||||||
return v
|
|
||||||
}
|
|
||||||
|
|
||||||
fn guide_enabled() -> Bool {
|
fn guide_enabled() -> Bool {
|
||||||
let v: String = env("GUIDE_ENABLE")
|
let v: String = config("GUIDE_ENABLE")
|
||||||
if str_eq(v, "1") { return true }
|
if str_eq(v, "1") { return true }
|
||||||
if str_eq(v, "on") { return true }
|
if str_eq(v, "on") { return true }
|
||||||
if str_eq(v, "true") { return true }
|
if str_eq(v, "true") { return true }
|
||||||
@@ -1120,15 +1227,15 @@ fn guide_probe_metal() -> Bool {
|
|||||||
|
|
||||||
// ── 2. Tier selection (config-driven thresholds, spec-autoselected) ────────────
|
// ── 2. Tier selection (config-driven thresholds, spec-autoselected) ────────────
|
||||||
fn guide_threshold_4b() -> Int {
|
fn guide_threshold_4b() -> Int {
|
||||||
return str_to_int(guide_env_or("GUIDE_RAM_GB_4B", "16"))
|
return str_to_int(config("GUIDE_RAM_GB_4B"))
|
||||||
}
|
}
|
||||||
fn guide_threshold_1p7b() -> Int {
|
fn guide_threshold_1p7b() -> Int {
|
||||||
return str_to_int(guide_env_or("GUIDE_RAM_GB_1P7B", "8"))
|
return str_to_int(config("GUIDE_RAM_GB_1P7B"))
|
||||||
}
|
}
|
||||||
|
|
||||||
// GUIDE_TIER_FORCE overrides the spec autoselect (used to prove cheaply on 0.6b).
|
// GUIDE_TIER_FORCE overrides the spec autoselect (used to prove cheaply on 0.6b).
|
||||||
fn guide_select_tier(ram_gb: Int) -> String {
|
fn guide_select_tier(ram_gb: Int) -> String {
|
||||||
let forced: String = env("GUIDE_TIER_FORCE")
|
let forced: String = config("GUIDE_TIER_FORCE")
|
||||||
if !str_eq(forced, "") { return forced }
|
if !str_eq(forced, "") { return forced }
|
||||||
if ram_gb >= guide_threshold_4b() { return "4b" }
|
if ram_gb >= guide_threshold_4b() { return "4b" }
|
||||||
if ram_gb >= guide_threshold_1p7b() { return "1.7b" }
|
if ram_gb >= guide_threshold_1p7b() { return "1.7b" }
|
||||||
@@ -1148,8 +1255,10 @@ fn guide_file(tier: String) -> String {
|
|||||||
}
|
}
|
||||||
|
|
||||||
fn guide_cache_dir() -> String {
|
fn guide_cache_dir() -> String {
|
||||||
let c: String = env("GUIDE_CACHE_DIR")
|
let c: String = config("GUIDE_CACHE_DIR")
|
||||||
if !str_eq(c, "") { return c }
|
if !str_eq(c, "") { return c }
|
||||||
|
// HOME stays a raw env() read: it is the ambient environment, not a knob of
|
||||||
|
// this program, and it is deliberately absent from the program block.
|
||||||
let home: String = env("HOME")
|
let home: String = env("HOME")
|
||||||
if !str_eq(home, "") { return home + "/.neuron/guide/models" }
|
if !str_eq(home, "") { return home + "/.neuron/guide/models" }
|
||||||
return engram_resolve_data_dir() + "/guide-models"
|
return engram_resolve_data_dir() + "/guide-models"
|
||||||
@@ -1190,9 +1299,9 @@ fn guide_fetch(tier: String) -> Bool {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// ── 4/5. Backend abstraction + BIND as an engageable interlocutor ──────────────
|
// ── 4/5. Backend abstraction + BIND as an engageable interlocutor ──────────────
|
||||||
fn guide_backend() -> String { return guide_env_or("GUIDE_BACKEND", "llama-server") }
|
fn guide_backend() -> String { return config("GUIDE_BACKEND") }
|
||||||
fn guide_host() -> String { return guide_env_or("GUIDE_HOST", "127.0.0.1") }
|
fn guide_host() -> String { return config("GUIDE_HOST") }
|
||||||
fn guide_port() -> String { return guide_env_or("GUIDE_PORT", "8771") }
|
fn guide_port() -> String { return config("GUIDE_PORT") }
|
||||||
fn guide_base_url() -> String { return "http://" + guide_host() + ":" + guide_port() }
|
fn guide_base_url() -> String { return "http://" + guide_host() + ":" + guide_port() }
|
||||||
|
|
||||||
// guide_healthy — is the guide present and answering? llama-server's /health
|
// guide_healthy — is the guide present and answering? llama-server's /health
|
||||||
@@ -1210,9 +1319,9 @@ fn guide_healthy() -> Bool {
|
|||||||
fn guide_load(tier: String) -> Bool {
|
fn guide_load(tier: String) -> Bool {
|
||||||
if guide_healthy() { return true }
|
if guide_healthy() { return true }
|
||||||
let path: String = guide_model_path(tier)
|
let path: String = guide_model_path(tier)
|
||||||
let bin: String = guide_env_or("GUIDE_LLAMA_SERVER_BIN", "llama-server")
|
let bin: String = config("GUIDE_LLAMA_SERVER_BIN")
|
||||||
let ngl: String = guide_env_or("GUIDE_NGL", "99")
|
let ngl: String = config("GUIDE_NGL")
|
||||||
let ctx: String = guide_env_or("GUIDE_CTX", "4096")
|
let ctx: String = config("GUIDE_CTX")
|
||||||
let logf: String = guide_cache_dir() + "/llama-server." + guide_port() + ".log"
|
let logf: String = guide_cache_dir() + "/llama-server." + guide_port() + ".log"
|
||||||
let cmd: String = bin + " -m '" + path + "' --host " + guide_host() + " --port " + guide_port() + " -c " + ctx + " -ngl " + ngl + " --jinja >> '" + logf + "' 2>&1"
|
let cmd: String = bin + " -m '" + path + "' --host " + guide_host() + " --port " + guide_port() + " -c " + ctx + " -ngl " + ngl + " --jinja >> '" + logf + "' 2>&1"
|
||||||
let pid: String = exec_bg(cmd)
|
let pid: String = exec_bg(cmd)
|
||||||
@@ -1608,7 +1717,7 @@ fn route_supersede(method: String, path: String, body: String) -> String {
|
|||||||
// ── Auth ──────────────────────────────────────────────────────────────────────
|
// ── Auth ──────────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
fn check_auth_ok(method: String, body: String) -> Bool {
|
fn check_auth_ok(method: String, body: String) -> Bool {
|
||||||
let key: String = env("ENGRAM_API_KEY")
|
let key: String = config("ENGRAM_API_KEY")
|
||||||
if str_eq(key, "") { return true }
|
if str_eq(key, "") { return true }
|
||||||
// Read-only methods don't require auth. Until http_serve surfaces
|
// Read-only methods don't require auth. Until http_serve surfaces
|
||||||
// request headers we can't accept a Bearer token cleanly; mutating
|
// request headers we can't accept a Bearer token cleanly; mutating
|
||||||
@@ -1871,8 +1980,7 @@ fn handle_request(method: String, path: String, body: String) -> String {
|
|||||||
|
|
||||||
// ── Entry ─────────────────────────────────────────────────────────────────────
|
// ── Entry ─────────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
let bind_raw: String = env("ENGRAM_BIND")
|
let bind_str: String = config("ENGRAM_BIND")
|
||||||
let bind_str: String = if str_eq(bind_raw, "") { ":8742" } else { bind_raw }
|
|
||||||
let port: Int = parse_port(bind_str)
|
let port: Int = parse_port(bind_str)
|
||||||
|
|
||||||
// On startup, try to load any existing snapshot (best effort).
|
// On startup, try to load any existing snapshot (best effort).
|
||||||
|
|||||||
Executable
+107
@@ -0,0 +1,107 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# run_vindex_concurrency_tests.sh — regression harness for the 2026-08-16 soul crash.
|
||||||
|
#
|
||||||
|
# Four halves. The SET is the point: it separates two hazards the original two-half
|
||||||
|
# version conflated, and which have fixes in different files.
|
||||||
|
#
|
||||||
|
# 1. single ASan+UBSan, one thread. MUST be clean. Hard failure.
|
||||||
|
#
|
||||||
|
# 2. readers TSan, N readers, NO writer. Hazard (a): the visited set used
|
||||||
|
# to live on the index, so two pure READS stamped each other's
|
||||||
|
# epoch. Fixed in engram_vindex.c (frame-owned VVisit +
|
||||||
|
# `const VIndex*` search). MUST be clean. Hard failure.
|
||||||
|
#
|
||||||
|
# 3. unsynchronized TSan, writer + reader on a BARE index. Hazard (b): in-place
|
||||||
|
# HNSW insert rewires existing elements' neighbour lists and
|
||||||
|
# reallocs elems[]. EXPECTED TO RACE, PERMANENTLY. This is not
|
||||||
|
# a bug to fix inside engram_vindex.c — it is the executable
|
||||||
|
# proof that a publication boundary must exist above it.
|
||||||
|
# Not a failure. If it ever goes CLEAN, the test stopped
|
||||||
|
# interleaving and half 4 is no longer meaningful either.
|
||||||
|
#
|
||||||
|
# 4. published TSan, owner + N readers through a publication boundary
|
||||||
|
# (rwlock: readers shared, owner exclusive) mirroring
|
||||||
|
# eg_vindex_view / eg_vindex_maintain in lang/runtime/el_runtime.c.
|
||||||
|
# MUST be clean, and all inserts must land. Hard failure.
|
||||||
|
#
|
||||||
|
# See test_vindex_concurrency.c for the full story (SIGSEGV at ASCII address
|
||||||
|
# "gramNode", heap corruption in xzm_realloc, etc).
|
||||||
|
#
|
||||||
|
# usage: run_vindex_concurrency_tests.sh
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
RUNTIME="$(cd "$HERE/../../lang/runtime" && pwd)"
|
||||||
|
WORK="$(mktemp -d)"
|
||||||
|
trap 'rm -rf "$WORK"' EXIT
|
||||||
|
|
||||||
|
SRC="$HERE/test_vindex_concurrency.c"
|
||||||
|
VINDEX="$RUNTIME/engram_vindex.c"
|
||||||
|
|
||||||
|
fail=0
|
||||||
|
|
||||||
|
echo "== [1/4] single-threaded control under AddressSanitizer =="
|
||||||
|
cc -std=c11 -g -O1 -fsanitize=address,undefined -fno-omit-frame-pointer \
|
||||||
|
-I"$RUNTIME" -o "$WORK/single" "$SRC" "$VINDEX" -lm || { echo "BUILD FAILED"; exit 2; }
|
||||||
|
if ASAN_OPTIONS=detect_leaks=0 "$WORK/single" single; then
|
||||||
|
echo " -> OK"
|
||||||
|
else
|
||||||
|
echo " -> FAIL: the single-threaded control must always be clean."
|
||||||
|
echo " If this fails the bug is NOT (only) concurrency — look for a real"
|
||||||
|
echo " out-of-bounds or lifetime error in engram_vindex.c."
|
||||||
|
fail=1
|
||||||
|
fi
|
||||||
|
|
||||||
|
cc -std=c11 -g -O1 -fsanitize=thread -fno-omit-frame-pointer \
|
||||||
|
-I"$RUNTIME" -o "$WORK/conc" "$SRC" "$VINDEX" -lm || { echo "BUILD FAILED"; exit 2; }
|
||||||
|
|
||||||
|
# run_tsan <mode> <logfile>; echoes nothing, sets $tsan_raced
|
||||||
|
run_tsan() {
|
||||||
|
TSAN_OPTIONS="halt_on_error=0" "$WORK/conc" "$1" >"$2" 2>&1
|
||||||
|
tsan_rc=$?
|
||||||
|
if grep -q "ThreadSanitizer: data race" "$2"; then tsan_raced=1; else tsan_raced=0; fi
|
||||||
|
}
|
||||||
|
|
||||||
|
echo
|
||||||
|
echo "== [2/4] concurrent READERS, no writer (visited-set gate) =="
|
||||||
|
run_tsan readers "$WORK/readers.log"
|
||||||
|
if [ "$tsan_raced" = "1" ]; then
|
||||||
|
echo " -> REGRESSION: two concurrent reads still race."
|
||||||
|
grep -m1 -A6 "ThreadSanitizer: data race" "$WORK/readers.log" | sed 's/^/ /'
|
||||||
|
echo " The visited set was supposed to be owned by the call frame."
|
||||||
|
fail=1
|
||||||
|
else
|
||||||
|
echo " -> clean (concurrent reads are safe)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo
|
||||||
|
echo "== [3/4] writer+reader on a BARE index (expected-race probe) =="
|
||||||
|
run_tsan unsynchronized "$WORK/unsync.log"
|
||||||
|
if [ "$tsan_raced" = "1" ]; then
|
||||||
|
echo " -> RACE DETECTED, as expected:"
|
||||||
|
grep -m1 -A4 "ThreadSanitizer: data race" "$WORK/unsync.log" | sed 's/^/ /'
|
||||||
|
echo " In-place HNSW insert mutates existing elements. Not fixable inside"
|
||||||
|
echo " engram_vindex.c — this is why the publication boundary exists."
|
||||||
|
else
|
||||||
|
echo " -> NOTE: no race reported. The probe did not interleave; half 4's"
|
||||||
|
echo " clean result proves less than it should. Investigate."
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo
|
||||||
|
echo "== [4/4] owner+readers through the publication boundary (boundary gate) =="
|
||||||
|
run_tsan published "$WORK/pub.log"
|
||||||
|
if [ "$tsan_raced" = "1" ]; then
|
||||||
|
echo " -> REGRESSION: the publication boundary did not serialize the owner."
|
||||||
|
grep -m1 -A6 "ThreadSanitizer: data race" "$WORK/pub.log" | sed 's/^/ /'
|
||||||
|
fail=1
|
||||||
|
elif [ "$tsan_rc" != "0" ]; then
|
||||||
|
echo " -> FAIL: boundary clean under TSan but the run failed:"
|
||||||
|
tail -3 "$WORK/pub.log" | sed 's/^/ /'
|
||||||
|
fail=1
|
||||||
|
else
|
||||||
|
echo " -> clean (readers project concurrently; the owner's inserts all landed)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo
|
||||||
|
[ "$fail" -eq 0 ] && echo "RESULT: PASS" || echo "RESULT: FAIL"
|
||||||
|
exit "$fail"
|
||||||
@@ -0,0 +1,251 @@
|
|||||||
|
/* test_vindex_concurrency.c — regression test for the 2026-08-16 soul crash.
|
||||||
|
*
|
||||||
|
* WHAT BROKE: the soul daemon crash-looped (5 crashes in ~100s) with SIGSEGV in
|
||||||
|
* search_layer <- vindex_insert <- eg_vindex_sync, a SIGABRT, and a fault inside
|
||||||
|
* xzm_realloc's own freelist — i.e. heap corruption. The SIGSEGV address
|
||||||
|
* 0x65646f4e6d617267 is little-endian ASCII "gramNode": string bytes being
|
||||||
|
* dereferenced as an Elem vector pointer.
|
||||||
|
*
|
||||||
|
* ROOT CAUSE: VIndex owns its traversal scratch (visited[] + visit_epoch), and
|
||||||
|
* search_layer mutates it via visited_reset(). So the index is unsafe for ANY
|
||||||
|
* concurrent use — including two concurrent READS. soul.el starts http_serve_async
|
||||||
|
* (a thread per connection) and then runs awareness_run() on the main thread, which
|
||||||
|
* reaches the same global index through engram_activate; nothing serialized them.
|
||||||
|
*
|
||||||
|
* Neither hnswlib nor FAISS puts the visited set on the index: hnswlib checks one
|
||||||
|
* out of a VisitedListPool per query, FAISS uses a thread_local VisitedTable.
|
||||||
|
*
|
||||||
|
* THE ORIGINAL `concurrent` HALF CONFLATED TWO DISTINCT HAZARDS (2026-08-16). It ran
|
||||||
|
* a writer against a reader on one bare index, so it could not tell apart:
|
||||||
|
*
|
||||||
|
* (a) READ/READ corruption — two searches stamping each other's visited epoch.
|
||||||
|
* A defect INSIDE engram_vindex.c, fixable there, and now fixed: the visited
|
||||||
|
* set moved to the call frame and vindex_search takes a `const VIndex*`.
|
||||||
|
*
|
||||||
|
* (b) WRITE/READ corruption — vindex_insert rewires the neighbour lists of
|
||||||
|
* EXISTING elements and reallocs elems[], so an insert is a mutation of the
|
||||||
|
* whole structure. This is NOT fixable inside engram_vindex.c at any price:
|
||||||
|
* it is inherent to in-place HNSW. It requires a publication boundary ABOVE
|
||||||
|
* the data structure (el_runtime.c: eg_vindex_view / eg_vindex_maintain).
|
||||||
|
*
|
||||||
|
* Conflating them made the suite unfailable-then-unpassable: fixing (a) left (b)
|
||||||
|
* still racing, which reads as "the fix did not work" when in fact a different,
|
||||||
|
* correctly-located fix is what (b) needs. So the halves are now separate:
|
||||||
|
*
|
||||||
|
* single N clustered vectors, ONE thread, ASan. The CONTROL. Must always
|
||||||
|
* be clean. When this passes and a concurrent half fails, the defect
|
||||||
|
* is concurrency, not an out-of-bounds/logic error in the graph code.
|
||||||
|
* (On 2026-08-16 this control cleared all 13,820 real dim-768 store
|
||||||
|
* vectors under ASan, which DISPROVED an inspection-derived hypothesis
|
||||||
|
* about an out-of-bounds reverse-link write at engram_vindex.c:340.)
|
||||||
|
*
|
||||||
|
* readers N reader threads, NO writer, one shared index, TSan. This is
|
||||||
|
* hazard (a) in isolation. It RACED before the visited set moved off
|
||||||
|
* the index struct and must be CLEAN now. Hard gate.
|
||||||
|
*
|
||||||
|
* unsynchronized writer + reader on a bare index, TSan. Hazard (b) in isolation.
|
||||||
|
* EXPECTED TO RACE, permanently — it is the executable proof that
|
||||||
|
* the index cannot be made safe from the inside, and therefore that
|
||||||
|
* the publication boundary in el_runtime.c has to exist. If this
|
||||||
|
* ever goes clean, the test stopped interleaving; do not celebrate.
|
||||||
|
*
|
||||||
|
* published writer + readers through a publication boundary that mirrors
|
||||||
|
* eg_vindex_view / eg_vindex_maintain (rwlock: readers shared,
|
||||||
|
* the single owner exclusive), TSan. Must be CLEAN. Hard gate.
|
||||||
|
* This is what proves the shape of the runtime fix, in the same
|
||||||
|
* process, rather than asserting it.
|
||||||
|
*
|
||||||
|
* Absence of a crash does NOT mean absence of a race — always read the sanitizer
|
||||||
|
* verdict, never just the exit code.
|
||||||
|
*
|
||||||
|
* Build/run: engram/test/run_vindex_concurrency_tests.sh
|
||||||
|
*/
|
||||||
|
#include "engram_vindex.h"
|
||||||
|
|
||||||
|
#include <pthread.h>
|
||||||
|
#include <stdio.h>
|
||||||
|
#include <stdlib.h>
|
||||||
|
#include <string.h>
|
||||||
|
#include <stdint.h>
|
||||||
|
|
||||||
|
#define DIM 128
|
||||||
|
#define NVEC 3000
|
||||||
|
#define SEED_N 50
|
||||||
|
|
||||||
|
static VIndex* g_ix;
|
||||||
|
static float* g_vecs;
|
||||||
|
|
||||||
|
/* Deterministic filler. Real embeddings are strongly correlated, not uniform noise;
|
||||||
|
* clustering keeps many candidates near-equidistant, which exercises the diversity
|
||||||
|
* heuristic and the visited set far harder than random vectors do. */
|
||||||
|
static void fill_vectors(void) {
|
||||||
|
g_vecs = (float*)malloc((size_t)NVEC * DIM * sizeof(float));
|
||||||
|
if (!g_vecs) { fprintf(stderr, "OOM\n"); exit(1); }
|
||||||
|
for (int i = 0; i < NVEC; i++) {
|
||||||
|
int cluster = i % 8;
|
||||||
|
for (int d = 0; d < DIM; d++)
|
||||||
|
g_vecs[(size_t)i * DIM + d] =
|
||||||
|
(float)(((d + cluster * 7) % 13) / 13.0) +
|
||||||
|
(float)(((i * 2654435761u + (unsigned)d) % 97) / 9700.0);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
static void* writer_fn(void* arg) {
|
||||||
|
(void)arg;
|
||||||
|
for (int i = SEED_N; i < NVEC; i++)
|
||||||
|
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
|
||||||
|
return NULL;
|
||||||
|
}
|
||||||
|
|
||||||
|
static void* reader_fn(void* arg) {
|
||||||
|
(void)arg;
|
||||||
|
uint64_t ids[8]; float ds[8];
|
||||||
|
for (int i = 0; i < 20000; i++)
|
||||||
|
(void)vindex_search(g_ix, g_vecs + (size_t)(i % NVEC) * DIM, 8, 0, ids, ds);
|
||||||
|
return NULL;
|
||||||
|
}
|
||||||
|
|
||||||
|
static int run_single(void) {
|
||||||
|
printf("[single] inserting %d vectors on one thread (ASan control)\n", NVEC);
|
||||||
|
g_ix = vindex_create(DIM, 0, 0);
|
||||||
|
if (!g_ix) { fprintf(stderr, "[single] vindex_create failed\n"); return 1; }
|
||||||
|
for (int i = 0; i < NVEC; i++) {
|
||||||
|
if (vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM) != 0) {
|
||||||
|
fprintf(stderr, "[single] insert %d failed\n", i); return 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (vindex_size(g_ix) != (size_t)NVEC) {
|
||||||
|
fprintf(stderr, "[single] size %zu != %d\n", vindex_size(g_ix), NVEC); return 1;
|
||||||
|
}
|
||||||
|
uint64_t ids[16]; float ds[16];
|
||||||
|
for (int q = 0; q < 200; q++) {
|
||||||
|
int k = vindex_search(g_ix, g_vecs + (size_t)((q * 7) % NVEC) * DIM, 16, 0, ids, ds);
|
||||||
|
if (k < 0) { fprintf(stderr, "[single] search failed at q=%d\n", q); return 1; }
|
||||||
|
}
|
||||||
|
vindex_free(g_ix); g_ix = NULL;
|
||||||
|
printf("[single] PASS — no memory error (this must ALWAYS pass)\n");
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* Hazard (b) in isolation: writer + reader on a BARE index, no boundary. */
|
||||||
|
static int run_unsynchronized(void) {
|
||||||
|
printf("[unsynchronized] 1 writer + 1 reader on a BARE index (TSan probe)\n");
|
||||||
|
printf("[unsynchronized] a race here is EXPECTED and PERMANENT — in-place HNSW\n");
|
||||||
|
printf("[unsynchronized] insert rewires existing elements. This is the proof that\n");
|
||||||
|
printf("[unsynchronized] the publication boundary must live ABOVE engram_vindex.c.\n");
|
||||||
|
g_ix = vindex_create(DIM, 0, 0);
|
||||||
|
if (!g_ix) { fprintf(stderr, "[unsynchronized] vindex_create failed\n"); return 1; }
|
||||||
|
for (int i = 0; i < SEED_N; i++)
|
||||||
|
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
|
||||||
|
|
||||||
|
pthread_t w, r;
|
||||||
|
if (pthread_create(&w, NULL, writer_fn, NULL) ||
|
||||||
|
pthread_create(&r, NULL, reader_fn, NULL)) {
|
||||||
|
fprintf(stderr, "[unsynchronized] pthread_create failed\n"); return 1;
|
||||||
|
}
|
||||||
|
pthread_join(w, NULL);
|
||||||
|
pthread_join(r, NULL);
|
||||||
|
vindex_free(g_ix); g_ix = NULL;
|
||||||
|
printf("[unsynchronized] completed — CHECK THE SANITIZER VERDICT, not this line.\n");
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── hazard (a) in isolation: concurrent READS only ───────────────────────────
|
||||||
|
* This is what the frame-owned visited set fixes. Before that change, two
|
||||||
|
* vindex_search calls on one index wrote each other's epoch stamp; TSan reported
|
||||||
|
* the race at visited_reset and the traversal then walked bogus element indices. */
|
||||||
|
#define NREADERS 4
|
||||||
|
|
||||||
|
static int run_readers(void) {
|
||||||
|
printf("[readers] %d concurrent readers, NO writer, one shared index (TSan)\n", NREADERS);
|
||||||
|
printf("[readers] this is the visited-set regression gate — must be CLEAN.\n");
|
||||||
|
g_ix = vindex_create(DIM, 0, 0);
|
||||||
|
if (!g_ix) { fprintf(stderr, "[readers] vindex_create failed\n"); return 1; }
|
||||||
|
for (int i = 0; i < NVEC; i++)
|
||||||
|
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
|
||||||
|
|
||||||
|
pthread_t t[NREADERS];
|
||||||
|
for (int i = 0; i < NREADERS; i++)
|
||||||
|
if (pthread_create(&t[i], NULL, reader_fn, NULL)) {
|
||||||
|
fprintf(stderr, "[readers] pthread_create failed\n"); return 1;
|
||||||
|
}
|
||||||
|
for (int i = 0; i < NREADERS; i++) pthread_join(t[i], NULL);
|
||||||
|
vindex_free(g_ix); g_ix = NULL;
|
||||||
|
printf("[readers] completed — CHECK THE SANITIZER VERDICT, not this line.\n");
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── the publication boundary, mirroring el_runtime.c ─────────────────────────
|
||||||
|
* Readers take the boundary SHARED and hold it across the whole search; the one
|
||||||
|
* owner takes it EXCLUSIVE to extend. Same shape as eg_vindex_view /
|
||||||
|
* eg_vindex_maintain. Note the reader's index pointer is `const VIndex*` — the
|
||||||
|
* compiler, not this comment, is what stops a reader inserting. */
|
||||||
|
static pthread_rwlock_t g_pub = PTHREAD_RWLOCK_INITIALIZER;
|
||||||
|
|
||||||
|
static void* pub_writer_fn(void* arg) {
|
||||||
|
(void)arg;
|
||||||
|
for (int i = SEED_N; i < NVEC; i++) {
|
||||||
|
pthread_rwlock_wrlock(&g_pub);
|
||||||
|
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
|
||||||
|
pthread_rwlock_unlock(&g_pub);
|
||||||
|
}
|
||||||
|
return NULL;
|
||||||
|
}
|
||||||
|
|
||||||
|
static void* pub_reader_fn(void* arg) {
|
||||||
|
(void)arg;
|
||||||
|
uint64_t ids[8]; float ds[8];
|
||||||
|
for (int i = 0; i < 5000; i++) {
|
||||||
|
pthread_rwlock_rdlock(&g_pub);
|
||||||
|
const VIndex* view = g_ix; /* immutable view */
|
||||||
|
(void)vindex_search(view, g_vecs + (size_t)(i % NVEC) * DIM, 8, 0, ids, ds);
|
||||||
|
pthread_rwlock_unlock(&g_pub);
|
||||||
|
}
|
||||||
|
return NULL;
|
||||||
|
}
|
||||||
|
|
||||||
|
static int run_published(void) {
|
||||||
|
printf("[published] 1 owner + %d readers through a publication boundary (TSan)\n", NREADERS);
|
||||||
|
printf("[published] this is the eg_vindex_view/eg_vindex_maintain gate — must be CLEAN.\n");
|
||||||
|
g_ix = vindex_create(DIM, 0, 0);
|
||||||
|
if (!g_ix) { fprintf(stderr, "[published] vindex_create failed\n"); return 1; }
|
||||||
|
for (int i = 0; i < SEED_N; i++)
|
||||||
|
(void)vindex_insert(g_ix, (uint64_t)i, g_vecs + (size_t)i * DIM);
|
||||||
|
|
||||||
|
pthread_t w, r[NREADERS];
|
||||||
|
if (pthread_create(&w, NULL, pub_writer_fn, NULL)) {
|
||||||
|
fprintf(stderr, "[published] pthread_create failed\n"); return 1;
|
||||||
|
}
|
||||||
|
for (int i = 0; i < NREADERS; i++)
|
||||||
|
if (pthread_create(&r[i], NULL, pub_reader_fn, NULL)) {
|
||||||
|
fprintf(stderr, "[published] pthread_create failed\n"); return 1;
|
||||||
|
}
|
||||||
|
pthread_join(w, NULL);
|
||||||
|
for (int i = 0; i < NREADERS; i++) pthread_join(r[i], NULL);
|
||||||
|
if (vindex_size(g_ix) != (size_t)NVEC) {
|
||||||
|
fprintf(stderr, "[published] size %zu != %d — the owner lost inserts\n",
|
||||||
|
vindex_size(g_ix), NVEC);
|
||||||
|
vindex_free(g_ix); g_ix = NULL; return 1;
|
||||||
|
}
|
||||||
|
vindex_free(g_ix); g_ix = NULL;
|
||||||
|
printf("[published] all %d inserts landed; CHECK THE SANITIZER VERDICT too.\n", NVEC);
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
int main(int argc, char** argv) {
|
||||||
|
const char* mode = (argc > 1) ? argv[1] : "single";
|
||||||
|
fill_vectors();
|
||||||
|
int rc;
|
||||||
|
if (!strcmp(mode, "single")) rc = run_single();
|
||||||
|
else if (!strcmp(mode, "readers")) rc = run_readers();
|
||||||
|
else if (!strcmp(mode, "unsynchronized")) rc = run_unsynchronized();
|
||||||
|
else if (!strcmp(mode, "published")) rc = run_published();
|
||||||
|
/* back-compat: the pre-split name meant the bare writer+reader probe. */
|
||||||
|
else if (!strcmp(mode, "concurrent")) rc = run_unsynchronized();
|
||||||
|
else {
|
||||||
|
fprintf(stderr, "usage: %s [single|readers|unsynchronized|published]\n", argv[0]);
|
||||||
|
rc = 2;
|
||||||
|
}
|
||||||
|
free(g_vecs);
|
||||||
|
return rc;
|
||||||
|
}
|
||||||
+27
-12
@@ -13,7 +13,7 @@
|
|||||||
// relations add edges. Every node enters with PROVENANCE + grounding-level
|
// relations add edges. Every node enters with PROVENANCE + grounding-level
|
||||||
// + stewardship class from the moment of entry.
|
// + stewardship class from the moment of entry.
|
||||||
//
|
//
|
||||||
// transduce() is THE single mechanism — one function, polymorphic, with no
|
// transduce_manifold() is THE single mechanism — one function, polymorphic, with no
|
||||||
// content-type branch inside it. It does not ask whether a payload is
|
// content-type branch inside it. It does not ask whether a payload is
|
||||||
// prose, structured data, or raw/opaque bytes (audio, or anything else);
|
// prose, structured data, or raw/opaque bytes (audio, or anything else);
|
||||||
// it runs one boundary-scan-with-fixed-window-fallback chunking algorithm
|
// it runs one boundary-scan-with-fixed-window-fallback chunking algorithm
|
||||||
@@ -401,10 +401,25 @@ fn head80(s: String) -> String {
|
|||||||
// truncates at the first embedded NUL, which is routine in real binary
|
// truncates at the first embedded NUL, which is routine in real binary
|
||||||
// bytes) is a MECHANICAL fidelity concern that belongs to whatever produced
|
// bytes) is a MECHANICAL fidelity concern that belongs to whatever produced
|
||||||
// `source` (see ingest_file's file_source_string below) — not a
|
// `source` (see ingest_file's file_source_string below) — not a
|
||||||
// content-type judgment made in here. transduce() never learns whether a
|
// content-type judgment made in here. transduce_manifold() never learns whether a
|
||||||
// chunk is plain text or a base64-encoded raw-byte window; every chunk is
|
// chunk is plain text or a base64-encoded raw-byte window; every chunk is
|
||||||
// handled identically either way.
|
// handled identically either way.
|
||||||
fn transduce(nodes: [String], edges: [String], source: String,
|
// RENAMED transduce -> transduce_manifold (2026-08-16). Two reasons, and the
|
||||||
|
// first is not the interesting one:
|
||||||
|
//
|
||||||
|
// 1. Mechanical: `transduce` is now a LANGUAGE primitive in el_runtime.h
|
||||||
|
// (transduce(signal, modality) -> Geometry). Every El `fn name(...)`
|
||||||
|
// compiles to a global C symbol with that exact name, so keeping this
|
||||||
|
// name here is a hard `conflicting types for 'transduce'` compile error
|
||||||
|
// the moment ingest.c links el_runtime.c. Measured, not anticipated.
|
||||||
|
//
|
||||||
|
// 2. Actual: this function was never signal->geometry. It chunks already-
|
||||||
|
// extracted content and PACKS it into a node+edge manifold — a real
|
||||||
|
// operation, but one layer up, and it had taken the name that belongs to
|
||||||
|
// the primitive underneath it. `transduce` is where a signal becomes
|
||||||
|
// geometry; `transduce_manifold` is where extracted content becomes
|
||||||
|
// structure. Nothing about this function's behaviour changed.
|
||||||
|
fn transduce_manifold(nodes: [String], edges: [String], source: String,
|
||||||
prov: String, ground: String, steward: String,
|
prov: String, ground: String, steward: String,
|
||||||
root_lid: String, root_title: String) -> [String] {
|
root_lid: String, root_title: String) -> [String] {
|
||||||
let tagbase: String = "prov:" + prov + " ground:" + ground + " steward:" + steward
|
let tagbase: String = "prov:" + prov + " ground:" + ground + " steward:" + steward
|
||||||
@@ -531,8 +546,8 @@ fn default_steward() -> String {
|
|||||||
// trustworthy verbatim. When they don't (silent truncation happened),
|
// trustworthy verbatim. When they don't (silent truncation happened),
|
||||||
// rebuild the payload as base64-encoded fixed-size windows read directly
|
// rebuild the payload as base64-encoded fixed-size windows read directly
|
||||||
// off disk (fs_read_b64_chunk — binary-safe in C), joined with the same
|
// off disk (fs_read_b64_chunk — binary-safe in C), joined with the same
|
||||||
// "\n\n" boundary marker transduce()'s generic scan already looks for, so
|
// "\n\n" boundary marker transduce_manifold()'s generic scan already looks for, so
|
||||||
// transduce() sees one ordinary boundary-delimited payload and runs its one
|
// transduce_manifold() sees one ordinary boundary-delimited payload and runs its one
|
||||||
// algorithm on it exactly as it would on prose — it never learns that a
|
// algorithm on it exactly as it would on prose — it never learns that a
|
||||||
// fidelity problem occurred upstream, let alone why.
|
// fidelity problem occurred upstream, let alone why.
|
||||||
fn file_source_string(path: String, text: String, real_size: Int) -> String {
|
fn file_source_string(path: String, text: String, real_size: Int) -> String {
|
||||||
@@ -541,7 +556,7 @@ fn file_source_string(path: String, text: String, real_size: Int) -> String {
|
|||||||
// 3072 raw bytes -> 4096 base64 chars (3 divides evenly into base64's
|
// 3072 raw bytes -> 4096 base64 chars (3 divides evenly into base64's
|
||||||
// 3-byte/4-char ratio); keeps each resulting node's content a clean,
|
// 3-byte/4-char ratio); keeps each resulting node's content a clean,
|
||||||
// bounded, low-kilobytes unit, same order of magnitude as the fixed
|
// bounded, low-kilobytes unit, same order of magnitude as the fixed
|
||||||
// fallback window in transduce() itself.
|
// fallback window in transduce_manifold() itself.
|
||||||
let win: Int = 3072
|
let win: Int = 3072
|
||||||
let out: String = ""
|
let out: String = ""
|
||||||
let off: Int = 0
|
let off: Int = 0
|
||||||
@@ -561,7 +576,7 @@ fn file_source_string(path: String, text: String, real_size: Int) -> String {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// ingest one file -> report JSON. Uniform for every file regardless of
|
// ingest one file -> report JSON. Uniform for every file regardless of
|
||||||
// extension or content — transduce() decides nothing about content-type, so
|
// extension or content — transduce_manifold() decides nothing about content-type, so
|
||||||
// neither does this function; it only decides whether the raw bytes made it
|
// neither does this function; it only decides whether the raw bytes made it
|
||||||
// through the read intact (file_source_string), which is a fidelity
|
// through the read intact (file_source_string), which is a fidelity
|
||||||
// question, not a format one.
|
// question, not a format one.
|
||||||
@@ -573,14 +588,14 @@ fn ingest_file(path: String) -> String {
|
|||||||
return "{\"error\":\"empty or unreadable\",\"path\":" + j_q(path) + "}"
|
return "{\"error\":\"empty or unreadable\",\"path\":" + j_q(path) + "}"
|
||||||
}
|
}
|
||||||
let prov: String = "file:" + path
|
let prov: String = "file:" + path
|
||||||
let packed: [String] = transduce(el_list_empty(), el_list_empty(),
|
let packed: [String] = transduce_manifold(el_list_empty(), el_list_empty(),
|
||||||
source, prov, default_ground(), default_steward(),
|
source, prov, default_ground(), default_steward(),
|
||||||
"doc:" + basename(path), basename(path))
|
"doc:" + basename(path), basename(path))
|
||||||
return merge_packed(packed)
|
return merge_packed(packed)
|
||||||
}
|
}
|
||||||
|
|
||||||
// ingest a directory: walk one level, ingest every file found, aggregate.
|
// ingest a directory: walk one level, ingest every file found, aggregate.
|
||||||
// No extension filter — transduce() handles any payload uniformly now, so
|
// No extension filter — transduce_manifold() handles any payload uniformly now, so
|
||||||
// there is no content-type gate at the directory boundary either.
|
// there is no content-type gate at the directory boundary either.
|
||||||
fn ingest_dir(path: String) -> String {
|
fn ingest_dir(path: String) -> String {
|
||||||
let entries: [String] = fs_list(path)
|
let entries: [String] = fs_list(path)
|
||||||
@@ -615,7 +630,7 @@ fn ingest_dir(path: String) -> String {
|
|||||||
fn ingest_url(url: String) -> String {
|
fn ingest_url(url: String) -> String {
|
||||||
let body: String = http_get(url)
|
let body: String = http_get(url)
|
||||||
if str_eq(body, "") { return "{\"error\":\"empty fetch\",\"url\":" + j_q(url) + "}" }
|
if str_eq(body, "") { return "{\"error\":\"empty fetch\",\"url\":" + j_q(url) + "}" }
|
||||||
let packed: [String] = transduce(el_list_empty(), el_list_empty(),
|
let packed: [String] = transduce_manifold(el_list_empty(), el_list_empty(),
|
||||||
body, "url:" + url, "extracted", "public-web",
|
body, "url:" + url, "extracted", "public-web",
|
||||||
"url:" + url, url)
|
"url:" + url, url)
|
||||||
return merge_packed(packed)
|
return merge_packed(packed)
|
||||||
@@ -630,7 +645,7 @@ fn ingest_llm(query: String) -> String {
|
|||||||
let resp: String = http_post_json("http://127.0.0.1:11434/api/generate", body)
|
let resp: String = http_post_json("http://127.0.0.1:11434/api/generate", body)
|
||||||
let answer: String = json_get_string(resp, "response")
|
let answer: String = json_get_string(resp, "response")
|
||||||
if str_eq(answer, "") { return "{\"error\":\"no model response\"}" }
|
if str_eq(answer, "") { return "{\"error\":\"no model response\"}" }
|
||||||
let packed: [String] = transduce(el_list_empty(), el_list_empty(),
|
let packed: [String] = transduce_manifold(el_list_empty(), el_list_empty(),
|
||||||
answer, "llm:" + model + ":" + query, "candidate-provisional", "guide-provisional",
|
answer, "llm:" + model + ":" + query, "candidate-provisional", "guide-provisional",
|
||||||
"llm:" + query, "guide answer: " + query)
|
"llm:" + query, "guide answer: " + query)
|
||||||
return merge_packed(packed)
|
return merge_packed(packed)
|
||||||
@@ -682,7 +697,7 @@ fn ingest_stream(path: String) -> String {
|
|||||||
// It is NOT a content-type flag: it says nothing about what's inside the
|
// It is NOT a content-type flag: it says nothing about what's inside the
|
||||||
// bytes once fetched, and none of the five ingest_* functions it selects
|
// bytes once fetched, and none of the five ingest_* functions it selects
|
||||||
// among interpret their payload differently by content shape anymore —
|
// among interpret their payload differently by content shape anymore —
|
||||||
// they all hand off to the single, format-agnostic transduce(). The old
|
// they all hand off to the single, format-agnostic transduce_manifold(). The old
|
||||||
// "structured" value (a caller-declared alias for "file", used only to hint
|
// "structured" value (a caller-declared alias for "file", used only to hint
|
||||||
// the now-removed JSON-vs-prose branch) is gone along with that branch.
|
// the now-removed JSON-vs-prose branch) is gone along with that branch.
|
||||||
let kind: String = env("INGEST_KIND")
|
let kind: String = env("INGEST_KIND")
|
||||||
|
|||||||
@@ -73,6 +73,17 @@ When you add a C builtin (verbatim-emit recipe — the El name is emitted as the
|
|||||||
2. Add a `__`-prefixed thin wrapper in `el_seed.c` and declare it in `el_seed.h`.
|
2. Add a `__`-prefixed thin wrapper in `el_seed.c` and declare it in `el_seed.h`.
|
||||||
3. Add the name to `builtin_arity` in `el-compiler/src/codegen.el` — add **both** the plain and `__`-prefixed spellings.
|
3. Add the name to `builtin_arity` in `el-compiler/src/codegen.el` — add **both** the plain and `__`-prefixed spellings.
|
||||||
4. Rebuild the elc binary (see below) and confirm the self-host fixpoint is byte-identical.
|
4. Rebuild the elc binary (see below) and confirm the self-host fixpoint is byte-identical.
|
||||||
|
5. **Prove it with a NEGATIVE CONTROL.** Show the test FAILING on a build without your change, then passing with it. A test that has never been seen to fail has proven nothing.
|
||||||
|
|
||||||
|
> **Step 5 is not optional, and step 4 does not cover it.** The fixpoint proves the *compiler reproduces itself*. It says nothing whatsoever about whether your builtin works. A recipe ending at "byte-identical" reads as complete while having verified nothing about the thing just added — which is why this file, until 2026-08-16, produced builtins with no tests at all.
|
||||||
|
>
|
||||||
|
> Measured cost of the omission (2026-08-16): `engram_node_set_emb`, `engram_curiosity_json` and `dream_set_handler` were all added in one session with zero tests. Separately, a UTF-8 fix was written, tested, and **the test passed on the unpatched build too** — the defect was elsewhere entirely, and only building the pre-fix binary exposed it. Without a negative control that fix would have merged as verified.
|
||||||
|
>
|
||||||
|
> Two shapes that pass while proving nothing, both hit the same day:
|
||||||
|
> - A test that never exercises your change (the route supplied a default that bypassed the code under test).
|
||||||
|
> - An induction that loses a race. `curl --max-time` on a large response left *both* builds alive; only `SO_LINGER 0` — a genuine RST, so the peer is provably gone — reproduced the failure. Six of ten attempts is not a control.
|
||||||
|
>
|
||||||
|
> Before every probe, confirm **your** process bound the port (`lsof -nP -iTCP:<port>`, match the PID). A stale instance answering on the port has silently produced false results here more than once, and `pkill -f` does not reliably match an argv like `./engram`.
|
||||||
|
|
||||||
Worked example: the `engram_assert_json` (op_assert seam) and `engram_node_full_in`/`engram_connect_in` (purview write-side) primitives added 2026-08-15 follow exactly this recipe.
|
Worked example: the `engram_assert_json` (op_assert seam) and `engram_node_full_in`/`engram_connect_in` (purview write-side) primitives added 2026-08-15 follow exactly this recipe.
|
||||||
|
|
||||||
|
|||||||
Vendored
BIN
Binary file not shown.
+226
-20
@@ -862,10 +862,23 @@ fn cg_expr(expr: Map<String, Any>) -> String {
|
|||||||
// arithmetic BinOp (or vice-versa). Without this check the
|
// arithmetic BinOp (or vice-versa). Without this check the
|
||||||
// fallthrough to str_eq produces str_eq(int_value, int_value)
|
// fallthrough to str_eq produces str_eq(int_value, int_value)
|
||||||
// which reads the integer as a char* and segfaults.
|
// which reads the integer as a char* and segfaults.
|
||||||
|
// EITHER side provably Int is enough. Requiring BOTH meant a call
|
||||||
|
// whose return type codegen cannot infer poisoned the operator:
|
||||||
|
// getint(5) == a -> str_eq(getint(5), a)
|
||||||
|
// even with `a` declared Int. str_eq then reads an integer as a
|
||||||
|
// char* and segfaults. Only an integer LITERAL on one side forced
|
||||||
|
// the numeric form, so the bug was invisible in the common case.
|
||||||
|
//
|
||||||
|
// Loosening to OR is strictly safer: when one side is a known Int,
|
||||||
|
// str_eq is always wrong (it dereferences that int), while numeric
|
||||||
|
// comparison is at worst a wrong answer on an already ill-typed
|
||||||
|
// program. When neither side is Int nothing changes, so string
|
||||||
|
// comparison is untouched.
|
||||||
if is_int_expr(left) {
|
if is_int_expr(left) {
|
||||||
if is_int_expr(right) {
|
return "(" + left_c + " == " + right_c + ")"
|
||||||
return "(" + left_c + " == " + right_c + ")"
|
}
|
||||||
}
|
if is_int_expr(right) {
|
||||||
|
return "(" + left_c + " == " + right_c + ")"
|
||||||
}
|
}
|
||||||
// Float literal or negative float literal: use plain == (bit-equal
|
// Float literal or negative float literal: use plain == (bit-equal
|
||||||
// el_val_t comparison). This handles `r0 == 3.0`, `neg == -3.0`, etc.
|
// el_val_t comparison). This handles `r0 == 3.0`, `neg == -3.0`, etc.
|
||||||
@@ -921,10 +934,12 @@ fn cg_expr(expr: Map<String, Any>) -> String {
|
|||||||
}
|
}
|
||||||
// Same mixed Ident/BinOp fix as EqEq: use is_int_expr to detect
|
// Same mixed Ident/BinOp fix as EqEq: use is_int_expr to detect
|
||||||
// integer-typed operands before falling through to !str_eq.
|
// integer-typed operands before falling through to !str_eq.
|
||||||
|
// Either side Int is enough — see the EqEq note above.
|
||||||
if is_int_expr(left) {
|
if is_int_expr(left) {
|
||||||
if is_int_expr(right) {
|
return "(" + left_c + " != " + right_c + ")"
|
||||||
return "(" + left_c + " != " + right_c + ")"
|
}
|
||||||
}
|
if is_int_expr(right) {
|
||||||
|
return "(" + left_c + " != " + right_c + ")"
|
||||||
}
|
}
|
||||||
// Float-typed operands use plain != (bit-equal comparison).
|
// Float-typed operands use plain != (bit-equal comparison).
|
||||||
if is_float_expr(left) {
|
if is_float_expr(left) {
|
||||||
@@ -1495,6 +1510,11 @@ fn cg_stmt(stmt: Map<String, Any>, indent: String, declared: [String]) -> [Strin
|
|||||||
if str_eq(ltype, "Int") {
|
if str_eq(ltype, "Int") {
|
||||||
add_int_name(name)
|
add_int_name(name)
|
||||||
}
|
}
|
||||||
|
// Same as params: Bool is an int in the value model. Without this a
|
||||||
|
// `let ok: Bool = ...` compared to another Bool lowered to str_eq.
|
||||||
|
if str_eq(ltype, "Bool") {
|
||||||
|
add_int_name(name)
|
||||||
|
}
|
||||||
if str_eq(ltype, "Float") {
|
if str_eq(ltype, "Float") {
|
||||||
add_float_name(name)
|
add_float_name(name)
|
||||||
}
|
}
|
||||||
@@ -1705,9 +1725,13 @@ fn cg_stmt(stmt: Map<String, Any>, indent: String, declared: [String]) -> [Strin
|
|||||||
} else {
|
} else {
|
||||||
let c_msg = "EL_STR_PTR(" + cg_expr(msg_node) + ")"
|
let c_msg = "EL_STR_PTR(" + cg_expr(msg_node) + ")"
|
||||||
}
|
}
|
||||||
|
// Assertions record into PER-TEST state, not global counters. The test
|
||||||
|
// is the unit of result; a global pass/fail tally cannot say which test
|
||||||
|
// failed or whether a test ran at all. Reporting is the runner's job —
|
||||||
|
// nothing is printed here.
|
||||||
emit_line(indent + "if (!(" + c_cond + ")) {")
|
emit_line(indent + "if (!(" + c_cond + ")) {")
|
||||||
emit_line(indent + " __el_test_fail(__el_cur_test, " + c_msg + "); __el_fail++;")
|
emit_line(indent + " __el_test_fail(" + c_msg + ");")
|
||||||
emit_line(indent + "} else { __el_pass++; }")
|
emit_line(indent + "} else { __el_cur_asserts++; }")
|
||||||
return declared
|
return declared
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -2602,6 +2626,17 @@ fn builtin_arity(name: String) -> Int {
|
|||||||
// LSP seed primitives
|
// LSP seed primitives
|
||||||
if str_eq(name, "__read_n") { return 1 }
|
if str_eq(name, "__read_n") { return 1 }
|
||||||
if str_eq(name, "__print_raw") { return 1 }
|
if str_eq(name, "__print_raw") { return 1 }
|
||||||
|
// Test-registry accessors. These are not runtime builtins — they are
|
||||||
|
// GENERATED into the same translation unit by the --test path below, one
|
||||||
|
// set per test binary. They are declared here so the El-side runner in
|
||||||
|
// runtime/eltest.el can call them with a known arity.
|
||||||
|
if str_eq(name, "__el_reg_count") { return 0 }
|
||||||
|
if str_eq(name, "__el_reg_name") { return 1 }
|
||||||
|
if str_eq(name, "__el_reg_invoke") { return 1 }
|
||||||
|
if str_eq(name, "__el_reg_last_ns") { return 0 }
|
||||||
|
if str_eq(name, "__el_reg_msg") { return 0 }
|
||||||
|
if str_eq(name, "__el_reg_asserts") { return 0 }
|
||||||
|
if str_eq(name, "__el_opt_json") { return 0 }
|
||||||
// String
|
// String
|
||||||
if str_eq(name, "el_str_concat") { return 2 }
|
if str_eq(name, "el_str_concat") { return 2 }
|
||||||
if str_eq(name, "str_eq") { return 2 }
|
if str_eq(name, "str_eq") { return 2 }
|
||||||
@@ -2872,6 +2907,7 @@ fn builtin_arity(name: String) -> Int {
|
|||||||
if str_eq(name, "el_alloc_count") { return 0 }
|
if str_eq(name, "el_alloc_count") { return 0 }
|
||||||
if str_eq(name, "el_alloc_bytes") { return 0 }
|
if str_eq(name, "el_alloc_bytes") { return 0 }
|
||||||
if str_eq(name, "el_peak_rss") { return 0 }
|
if str_eq(name, "el_peak_rss") { return 0 }
|
||||||
|
if str_eq(name, "el_black_box") { return 1 }
|
||||||
if str_eq(name, "engram_neighbors_json") { return 3 }
|
if str_eq(name, "engram_neighbors_json") { return 3 }
|
||||||
if str_eq(name, "engram_activate_json") { return 2 }
|
if str_eq(name, "engram_activate_json") { return 2 }
|
||||||
if str_eq(name, "engram_stats_json") { return 0 }
|
if str_eq(name, "engram_stats_json") { return 0 }
|
||||||
@@ -3097,6 +3133,15 @@ fn build_int_names_for_params(params: [Map<String, Any>]) -> Bool {
|
|||||||
if str_eq(ptype, "Int") {
|
if str_eq(ptype, "Int") {
|
||||||
add_int_name(pname)
|
add_int_name(pname)
|
||||||
}
|
}
|
||||||
|
// Bool is an integer in the value model (type_to_c maps Bool -> "int";
|
||||||
|
// el_runtime.h: "Bool -> el_val_t (0 = false, nonzero = true)"), but
|
||||||
|
// Bool names were registered nowhere. So `cond == want` between two
|
||||||
|
// Bool params fell through to str_eq and dereferenced 0 or 1 as a
|
||||||
|
// char* — an immediate segfault. Track them as int-like, which is what
|
||||||
|
// they are.
|
||||||
|
if str_eq(ptype, "Bool") {
|
||||||
|
add_int_name(pname)
|
||||||
|
}
|
||||||
if str_eq(ptype, "Float") {
|
if str_eq(ptype, "Float") {
|
||||||
add_float_name(pname)
|
add_float_name(pname)
|
||||||
}
|
}
|
||||||
@@ -3220,6 +3265,7 @@ fn is_top_level_decl(stmt: Map<String, Any>) -> Bool {
|
|||||||
if kind == "EnumDef" { return true }
|
if kind == "EnumDef" { return true }
|
||||||
if kind == "Import" { return true }
|
if kind == "Import" { return true }
|
||||||
if kind == "CgiBlock" { return true }
|
if kind == "CgiBlock" { return true }
|
||||||
|
if kind == "ProgramBlock" { return true }
|
||||||
if kind == "ExternFn" { return true }
|
if kind == "ExternFn" { return true }
|
||||||
false
|
false
|
||||||
}
|
}
|
||||||
@@ -3232,6 +3278,55 @@ fn cgi_arg(value: String, has_value: Bool) -> String {
|
|||||||
return "EL_NULL"
|
return "EL_NULL"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// -- Program block: cross-cutting concerns injected at the process boundary ----
|
||||||
|
//
|
||||||
|
// emit_program_init — emit the `static void __el_program_init(void)` that
|
||||||
|
// carries a program's declared cross-cutting concerns. Called from main()
|
||||||
|
// BEFORE any user statement runs, so the guarantees hold for the whole process
|
||||||
|
// rather than depending on each call site remembering to ask for them.
|
||||||
|
//
|
||||||
|
// This is emitted at the point the `program` block is encountered, not buffered
|
||||||
|
// until main(). The streaming backend emits in source order and cannot hold a
|
||||||
|
// declaration's entry list alive until main(); emitting a named function here
|
||||||
|
// and calling it from main() means only a single bool has to survive.
|
||||||
|
//
|
||||||
|
// Order matters and is deliberate:
|
||||||
|
// 1. singleton FIRST — if another instance already holds the lock, refuse and
|
||||||
|
// exit before touching configuration, ports, or any data directory.
|
||||||
|
// 2. config declarations — resolve env-or-default, one declaration per entry.
|
||||||
|
// 3. validate LAST — report EVERY missing/ill-typed entry at once, then exit.
|
||||||
|
fn el_bool_arg(b: Bool) -> String {
|
||||||
|
if b { return "EL_INT(1)" }
|
||||||
|
return "EL_INT(0)"
|
||||||
|
}
|
||||||
|
|
||||||
|
fn emit_program_init(stmt: Map<String, Any>) -> Void {
|
||||||
|
let pname: String = stmt["name"]
|
||||||
|
emit_line("static void __el_program_init(void) {")
|
||||||
|
let has_singleton: Bool = stmt["has_singleton"]
|
||||||
|
if has_singleton {
|
||||||
|
let sid: String = stmt["singleton"]
|
||||||
|
emit_line(" el_singleton_acquire(EL_STR(" + c_str_lit(sid) + "));")
|
||||||
|
}
|
||||||
|
let entries = stmt["entries"]
|
||||||
|
let n: Int = native_list_len(entries)
|
||||||
|
let i = 0
|
||||||
|
while i < n {
|
||||||
|
let e = native_list_get(entries, i)
|
||||||
|
let ename: String = e["name"]
|
||||||
|
let etype: String = e["etype"]
|
||||||
|
let edefault: String = e["default"]
|
||||||
|
let has_default: Bool = e["has_default"]
|
||||||
|
let erequired: Bool = e["required"]
|
||||||
|
let arg_def: String = cgi_arg(edefault, has_default)
|
||||||
|
emit_line(" el_config_declare(EL_STR(" + c_str_lit(ename) + "), EL_STR(" + c_str_lit(etype) + "), " + arg_def + ", " + el_bool_arg(has_default) + ", " + el_bool_arg(erequired) + ");")
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
emit_line(" el_config_validate(EL_STR(" + c_str_lit(pname) + "));")
|
||||||
|
emit_line("}")
|
||||||
|
emit_blank()
|
||||||
|
}
|
||||||
|
|
||||||
// -- VBD role enforcement ------------------------------------------------------
|
// -- VBD role enforcement ------------------------------------------------------
|
||||||
//
|
//
|
||||||
// Scan a function body for direct calls to DHARMA-restricted builtins
|
// Scan a function body for direct calls to DHARMA-restricted builtins
|
||||||
@@ -3554,6 +3649,20 @@ fn codegen(stmts: [Map<String, Any>], source: String) -> String {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Program block: emit the cross-cutting init function before the user's
|
||||||
|
// functions so main() can call it (see emit_program_init).
|
||||||
|
let prog_have: Bool = false
|
||||||
|
let i = 0
|
||||||
|
while i < n {
|
||||||
|
let stmt = native_list_get(stmts, i)
|
||||||
|
let sk4: String = stmt["stmt"]
|
||||||
|
if str_eq(sk4, "ProgramBlock") {
|
||||||
|
emit_program_init(stmt)
|
||||||
|
let prog_have = true
|
||||||
|
}
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
|
||||||
// Function definitions
|
// Function definitions
|
||||||
let i = 0
|
let i = 0
|
||||||
while i < n {
|
while i < n {
|
||||||
@@ -3572,6 +3681,9 @@ fn codegen(stmts: [Map<String, Any>], source: String) -> String {
|
|||||||
// with the C-side parameters when fn main()'s body is folded in below.
|
// with the C-side parameters when fn main()'s body is folded in below.
|
||||||
emit_line("int main(int _argc, char** _argv) {")
|
emit_line("int main(int _argc, char** _argv) {")
|
||||||
emit_line(" el_runtime_init_args(_argc, _argv);")
|
emit_line(" el_runtime_init_args(_argc, _argv);")
|
||||||
|
if prog_have {
|
||||||
|
emit_line(" __el_program_init();")
|
||||||
|
}
|
||||||
if cgi_count >= 1 {
|
if cgi_count >= 1 {
|
||||||
let cname: String = cgi_block["name"]
|
let cname: String = cgi_block["name"]
|
||||||
let cdid: String = cgi_block["dharma_id"]
|
let cdid: String = cgi_block["dharma_id"]
|
||||||
@@ -4116,13 +4228,36 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
|||||||
// Emit test harness preamble (counters, fail printer) when in test mode.
|
// Emit test harness preamble (counters, fail printer) when in test mode.
|
||||||
if test_is_mode {
|
if test_is_mode {
|
||||||
emit_line("#include <stdio.h>")
|
emit_line("#include <stdio.h>")
|
||||||
|
emit_line("#include <string.h>")
|
||||||
|
emit_line("#include <time.h>")
|
||||||
emit_blank()
|
emit_blank()
|
||||||
emit_line("static int __el_pass = 0, __el_fail = 0;")
|
// Per-test result state. Reset by __el_reg_invoke before each test, so
|
||||||
|
// every test gets its own record rather than contributing to a global
|
||||||
|
// tally. The first failure message is retained; later ones only bump
|
||||||
|
// the count, which keeps the common case allocation-free.
|
||||||
|
emit_line("static int __el_cur_fails = 0;")
|
||||||
|
emit_line("static int __el_cur_asserts = 0;")
|
||||||
|
emit_line("static char __el_cur_msg[512] = \"\";")
|
||||||
emit_line("static const char *__el_cur_test = \"(none)\";")
|
emit_line("static const char *__el_cur_test = \"(none)\";")
|
||||||
emit_line("static void __el_test_fail(const char *test, const char *msg) {")
|
emit_line("static void __el_test_fail(const char *msg) {")
|
||||||
emit_line(" fprintf(stderr, \"FAIL %-40s %s\\n\", test, msg);")
|
emit_line(" if (__el_cur_fails == 0 && msg) {")
|
||||||
|
emit_line(" snprintf(__el_cur_msg, sizeof __el_cur_msg, \"%s\", msg);")
|
||||||
|
emit_line(" }")
|
||||||
|
emit_line(" __el_cur_fails++; __el_cur_asserts++;")
|
||||||
emit_line("}")
|
emit_line("}")
|
||||||
emit_blank()
|
emit_blank()
|
||||||
|
// Forward declarations for the registry accessors. The definitions are
|
||||||
|
// emitted at the END of the unit (they reference the test functions,
|
||||||
|
// which do not exist yet at this point), but the El-side runner is
|
||||||
|
// compiled in between and calls them — so it needs the prototypes here.
|
||||||
|
emit_line("el_val_t __el_reg_count(void);")
|
||||||
|
emit_line("el_val_t __el_reg_name(el_val_t i);")
|
||||||
|
emit_line("el_val_t __el_reg_invoke(el_val_t i);")
|
||||||
|
emit_line("el_val_t __el_reg_last_ns(void);")
|
||||||
|
emit_line("el_val_t __el_reg_msg(void);")
|
||||||
|
emit_line("el_val_t __el_reg_asserts(void);")
|
||||||
|
emit_line("el_val_t __el_opt_json(void);")
|
||||||
|
emit_blank()
|
||||||
}
|
}
|
||||||
|
|
||||||
// Streaming parse-emit loop.
|
// Streaming parse-emit loop.
|
||||||
@@ -4142,6 +4277,7 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
|||||||
// Fix: copy the values out BEFORE the release (strings, so no dangling reference)
|
// Fix: copy the values out BEFORE the release (strings, so no dangling reference)
|
||||||
// and emit from these. No search, so the failure mode is removed rather than moved.
|
// and emit from these. No search, so the failure mode is removed rather than moved.
|
||||||
let cgi_have: Bool = false
|
let cgi_have: Bool = false
|
||||||
|
let prog_have: Bool = false
|
||||||
let cgi_name_v: String = ""
|
let cgi_name_v: String = ""
|
||||||
let cgi_did_v: String = ""
|
let cgi_did_v: String = ""
|
||||||
let cgi_prin_v: String = ""
|
let cgi_prin_v: String = ""
|
||||||
@@ -4263,6 +4399,14 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
|||||||
// These are no-ops in codegen (forward decls already emitted)
|
// These are no-ops in codegen (forward decls already emitted)
|
||||||
// — except a CgiBlock, whose declared identity must survive
|
// — except a CgiBlock, whose declared identity must survive
|
||||||
// this release to be emitted as a compiled constant.
|
// this release to be emitted as a compiled constant.
|
||||||
|
// A ProgramBlock's cross-cutting declarations are
|
||||||
|
// emitted HERE, as a named init function, because the
|
||||||
|
// streaming backend cannot hold the entry list alive
|
||||||
|
// until main(). Only the bool survives.
|
||||||
|
if str_eq(sk, "ProgramBlock") {
|
||||||
|
emit_program_init(stmt)
|
||||||
|
let prog_have = true
|
||||||
|
}
|
||||||
if str_eq(sk, "CgiBlock") {
|
if str_eq(sk, "CgiBlock") {
|
||||||
let cgi_have = true
|
let cgi_have = true
|
||||||
let cgi_name_v = stmt["name"]
|
let cgi_name_v = stmt["name"]
|
||||||
@@ -4318,17 +4462,72 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
|||||||
el_release(sigs)
|
el_release(sigs)
|
||||||
|
|
||||||
let test_arena_mark: Any = el_arena_push()
|
let test_arena_mark: Any = el_arena_push()
|
||||||
|
let tn: Int = native_list_len(test_c_names)
|
||||||
|
|
||||||
|
// ── Generated test registry ──────────────────────────────────────────
|
||||||
|
// Discovery happens HERE, at compile time. The runner never searches
|
||||||
|
// for tests; it walks this table. That ordering — discovery strictly
|
||||||
|
// before execution — is what makes --list, filtering, sharding and
|
||||||
|
// per-test reporting possible later, and it is why the old harness
|
||||||
|
// (which inlined direct calls into main) could not have any of them.
|
||||||
|
emit_line("typedef void (*__el_test_fp)(void);")
|
||||||
|
emit_line("typedef struct { const char *name; __el_test_fp fn; } __el_test_entry;")
|
||||||
|
emit_line("static const __el_test_entry __el_registry[] = {")
|
||||||
|
let ri: Int = 0
|
||||||
|
while ri < tn {
|
||||||
|
let r_name: String = native_list_get(test_names, ri)
|
||||||
|
let r_cfn: String = native_list_get(test_c_names, ri)
|
||||||
|
emit_line(" { \"" + c_escape(r_name) + "\", " + r_cfn + " },")
|
||||||
|
let ri = ri + 1
|
||||||
|
}
|
||||||
|
// Trailing sentinel keeps the array non-empty when a file declares no
|
||||||
|
// tests (a zero-length array is not valid C).
|
||||||
|
emit_line(" { 0, 0 }")
|
||||||
|
emit_line("};")
|
||||||
|
emit_line("static const int __el_registry_n = " + int_to_str(tn) + ";")
|
||||||
|
emit_blank()
|
||||||
|
emit_line("static long long __el_last_ns = 0;")
|
||||||
|
emit_line("static int __el_opt_json_v = 0;")
|
||||||
|
emit_blank()
|
||||||
|
|
||||||
|
// ── Index-based accessors ────────────────────────────────────────────
|
||||||
|
// El has no function pointers, so the runner works purely in indices.
|
||||||
|
// This is the whole seam between generated C and the El-side runner.
|
||||||
|
emit_line("el_val_t __el_reg_count(void) { return (el_val_t)(int64_t)__el_registry_n; }")
|
||||||
|
emit_line("el_val_t __el_reg_name(el_val_t i) {")
|
||||||
|
emit_line(" int64_t k = (int64_t)i;")
|
||||||
|
emit_line(" if (k < 0 || k >= __el_registry_n) return EL_STR(\"\");")
|
||||||
|
emit_line(" return EL_STR(__el_registry[k].name);")
|
||||||
|
emit_line("}")
|
||||||
|
// Timing is taken immediately around the call, in C, on the MONOTONIC
|
||||||
|
// clock — never the wall clock, which can step backwards under NTP.
|
||||||
|
emit_line("el_val_t __el_reg_invoke(el_val_t i) {")
|
||||||
|
emit_line(" int64_t k = (int64_t)i;")
|
||||||
|
emit_line(" if (k < 0 || k >= __el_registry_n) return 0;")
|
||||||
|
emit_line(" __el_cur_fails = 0; __el_cur_asserts = 0; __el_cur_msg[0] = '\\0';")
|
||||||
|
emit_line(" __el_cur_test = __el_registry[k].name;")
|
||||||
|
emit_line(" struct timespec _t0, _t1;")
|
||||||
|
emit_line(" clock_gettime(CLOCK_MONOTONIC, &_t0);")
|
||||||
|
emit_line(" __el_registry[k].fn();")
|
||||||
|
emit_line(" clock_gettime(CLOCK_MONOTONIC, &_t1);")
|
||||||
|
emit_line(" __el_last_ns = (long long)(_t1.tv_sec - _t0.tv_sec) * 1000000000LL")
|
||||||
|
emit_line(" + (long long)(_t1.tv_nsec - _t0.tv_nsec);")
|
||||||
|
emit_line(" return (el_val_t)(int64_t)__el_cur_fails;")
|
||||||
|
emit_line("}")
|
||||||
|
emit_line("el_val_t __el_reg_last_ns(void) { return (el_val_t)(int64_t)__el_last_ns; }")
|
||||||
|
emit_line("el_val_t __el_reg_msg(void) { return EL_STR(__el_cur_msg); }")
|
||||||
|
emit_line("el_val_t __el_reg_asserts(void) { return (el_val_t)(int64_t)__el_cur_asserts; }")
|
||||||
|
emit_line("el_val_t __el_opt_json(void) { return (el_val_t)(int64_t)__el_opt_json_v; }")
|
||||||
|
emit_blank()
|
||||||
|
|
||||||
|
// main() delegates to the El-side runner. Everything above this line is
|
||||||
|
// generated glue; all reporting logic lives in runtime/eltest.el.
|
||||||
emit_line("int main(int _argc, char **_argv) {")
|
emit_line("int main(int _argc, char **_argv) {")
|
||||||
emit_line(" el_runtime_init_args(_argc, _argv);")
|
emit_line(" el_runtime_init_args(_argc, _argv);")
|
||||||
let ti: Int = 0
|
emit_line(" for (int _i = 1; _i < _argc; _i++) {")
|
||||||
let tn: Int = native_list_len(test_c_names)
|
emit_line(" if (strcmp(_argv[_i], \"--json\") == 0) __el_opt_json_v = 1;")
|
||||||
while ti < tn {
|
emit_line(" }")
|
||||||
let tc_name: String = native_list_get(test_c_names, ti)
|
emit_line(" return (int)(int64_t)el_test_main();")
|
||||||
emit_line(" " + tc_name + "();")
|
|
||||||
let ti = ti + 1
|
|
||||||
}
|
|
||||||
emit_line(" printf(\"%d passed, %d failed\\n\", __el_pass, __el_fail);")
|
|
||||||
emit_line(" return __el_fail;")
|
|
||||||
emit_line("}")
|
emit_line("}")
|
||||||
el_arena_pop(test_arena_mark)
|
el_arena_pop(test_arena_mark)
|
||||||
el_release(test_names)
|
el_release(test_names)
|
||||||
@@ -4354,6 +4553,13 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
|||||||
let kind2: String = state_get("__program_kind")
|
let kind2: String = state_get("__program_kind")
|
||||||
emit_line("int main(int _argc, char** _argv) {")
|
emit_line("int main(int _argc, char** _argv) {")
|
||||||
emit_line(" el_runtime_init_args(_argc, _argv);")
|
emit_line(" el_runtime_init_args(_argc, _argv);")
|
||||||
|
// Cross-cutting concerns declared by a `program` block run BEFORE anything
|
||||||
|
// else — a singleton violation must refuse the start before this process
|
||||||
|
// touches a port or a data directory, and configuration must be validated
|
||||||
|
// before the first read of it rather than at each read site.
|
||||||
|
if prog_have {
|
||||||
|
emit_line(" __el_program_init();")
|
||||||
|
}
|
||||||
|
|
||||||
// cgi init if needed
|
// cgi init if needed
|
||||||
let ns2: Int = native_list_len(sigs)
|
let ns2: Int = native_list_len(sigs)
|
||||||
|
|||||||
@@ -419,6 +419,22 @@ fn resolve_imports(src_path: String) -> String {
|
|||||||
if !str_eq(already, "") { return "" }
|
if !str_eq(already, "") { return "" }
|
||||||
state_set(seen_key, "1")
|
state_set(seen_key, "1")
|
||||||
|
|
||||||
|
// A missing file must be a hard error, never an empty string.
|
||||||
|
//
|
||||||
|
// fs_read returns "" both for "file is empty" and "file does not exist", and
|
||||||
|
// this function used the value without distinguishing them. So a broken
|
||||||
|
// import path — a typo, a moved file, a relative path resolved from the
|
||||||
|
// wrong working directory — compiled CLEANLY: exit 0, empty stderr, and a
|
||||||
|
// program silently missing everything it imported. Observed 2026-08-15:
|
||||||
|
// eleven consecutive "successful" compiles that had included no runtime at
|
||||||
|
// all, and a wrong conclusion drawn from them before anyone noticed.
|
||||||
|
//
|
||||||
|
// Missing dependency, confident success. fs_exists separates the two cases,
|
||||||
|
// so a genuinely empty file still resolves to "" and is fine.
|
||||||
|
if !fs_exists(src_path) {
|
||||||
|
println("elc: cannot resolve import: " + src_path)
|
||||||
|
exit_program(1)
|
||||||
|
}
|
||||||
let source: String = fs_read(src_path)
|
let source: String = fs_read(src_path)
|
||||||
let dir: String = dirname_of(src_path)
|
let dir: String = dirname_of(src_path)
|
||||||
let lines: [String] = str_split(source, "\n")
|
let lines: [String] = str_split(source, "\n")
|
||||||
|
|||||||
@@ -184,6 +184,7 @@ fn keyword_kind(word: String) -> String {
|
|||||||
if word == "false" { return "Bool" }
|
if word == "false" { return "Bool" }
|
||||||
if word == "cgi" { return "Cgi" }
|
if word == "cgi" { return "Cgi" }
|
||||||
if word == "service" { return "Service" }
|
if word == "service" { return "Service" }
|
||||||
|
if word == "program" { return "Program" }
|
||||||
if word == "manager" { return "Manager" }
|
if word == "manager" { return "Manager" }
|
||||||
if word == "engine" { return "Engine" }
|
if word == "engine" { return "Engine" }
|
||||||
if word == "accessor" { return "Accessor" }
|
if word == "accessor" { return "Accessor" }
|
||||||
|
|||||||
@@ -1967,6 +1967,113 @@ fn parse_stmt(tokens: [Any], pos: Int) -> Map<String, Any> {
|
|||||||
}, p)
|
}, p)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// program block: program "name" { singleton: "id", env NAME: Type = "default", ... }
|
||||||
|
//
|
||||||
|
// The program block is El's declaration surface for CROSS-CUTTING CONCERNS —
|
||||||
|
// properties of the whole process rather than of any one function, which
|
||||||
|
// otherwise degrade into "remember to call this at every site" conventions.
|
||||||
|
//
|
||||||
|
// singleton: "id" — process identity. The runtime takes an exclusive
|
||||||
|
// lock at startup; a SECOND start is refused, loudly,
|
||||||
|
// instead of two processes sharing one data dir.
|
||||||
|
// env NAME: T = "d" — one configuration entry. Its type and its default
|
||||||
|
// are declared ONCE, here, and resolved+validated
|
||||||
|
// before main() body runs.
|
||||||
|
// env NAME: T required
|
||||||
|
// — no default; the program refuses to start unless the
|
||||||
|
// variable is set.
|
||||||
|
//
|
||||||
|
// Both compile into calls injected at the head of main() — the same boundary
|
||||||
|
// seam `cgi` already uses (codegen.el emit_program_init). No call site in the
|
||||||
|
// program body has to remember anything, which is the whole point.
|
||||||
|
if k == "Program" {
|
||||||
|
let p = pos + 1
|
||||||
|
let name = tok_value(tokens, p)
|
||||||
|
let p = p + 1
|
||||||
|
let p = expect(tokens, p, "LBrace")
|
||||||
|
let singleton = ""
|
||||||
|
let has_singleton = false
|
||||||
|
let entries = native_list_empty()
|
||||||
|
// Entry-scratch declared at loop-body level (not inside the branch) so
|
||||||
|
// that inner `let` forms compile to assignment rather than a C-scoped
|
||||||
|
// redeclaration — the same idiom the service block above relies on.
|
||||||
|
let ename = ""
|
||||||
|
let etype = ""
|
||||||
|
let edefault = ""
|
||||||
|
let has_default = false
|
||||||
|
let erequired = false
|
||||||
|
let fname = ""
|
||||||
|
let fval = ""
|
||||||
|
let running = true
|
||||||
|
while running {
|
||||||
|
let k2 = tok_kind(tokens, p)
|
||||||
|
if k2 == "RBrace" {
|
||||||
|
let running = false
|
||||||
|
} else {
|
||||||
|
if k2 == "Eof" {
|
||||||
|
let running = false
|
||||||
|
} else {
|
||||||
|
let fname = tok_value(tokens, p)
|
||||||
|
let p = p + 1
|
||||||
|
if str_eq(fname, "env") {
|
||||||
|
// env NAME: Type [= "default"] [required]
|
||||||
|
let ename = tok_value(tokens, p)
|
||||||
|
let p = p + 1
|
||||||
|
let p = expect(tokens, p, "Colon")
|
||||||
|
let etype = tok_value(tokens, p)
|
||||||
|
let p = p + 1
|
||||||
|
let edefault = ""
|
||||||
|
let has_default = false
|
||||||
|
let erequired = false
|
||||||
|
let k3 = tok_kind(tokens, p)
|
||||||
|
if str_eq(k3, "Eq") {
|
||||||
|
let p = p + 1
|
||||||
|
let edefault = tok_value(tokens, p)
|
||||||
|
let has_default = true
|
||||||
|
let p = p + 1
|
||||||
|
}
|
||||||
|
let k4 = tok_kind(tokens, p)
|
||||||
|
if str_eq(k4, "Ident") {
|
||||||
|
let w = tok_value(tokens, p)
|
||||||
|
if str_eq(w, "required") {
|
||||||
|
let erequired = true
|
||||||
|
let p = p + 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let entries = native_list_append(entries, {
|
||||||
|
"name": ename,
|
||||||
|
"etype": etype,
|
||||||
|
"default": edefault,
|
||||||
|
"has_default": has_default,
|
||||||
|
"required": erequired
|
||||||
|
})
|
||||||
|
} else {
|
||||||
|
// scalar field: `name: "value"`
|
||||||
|
let p = expect(tokens, p, "Colon")
|
||||||
|
let fval = tok_value(tokens, p)
|
||||||
|
let p = p + 1
|
||||||
|
if str_eq(fname, "singleton") {
|
||||||
|
let singleton = fval
|
||||||
|
let has_singleton = true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let k5 = tok_kind(tokens, p)
|
||||||
|
if k5 == "Comma" {
|
||||||
|
let p = p + 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let p = expect(tokens, p, "RBrace")
|
||||||
|
return make_result({
|
||||||
|
"stmt": "ProgramBlock",
|
||||||
|
"name": name,
|
||||||
|
"singleton": singleton,
|
||||||
|
"has_singleton": has_singleton,
|
||||||
|
"entries": entries
|
||||||
|
}, p)
|
||||||
|
}
|
||||||
|
|
||||||
// assert <cond_expr> [ , <msg_expr> ]
|
// assert <cond_expr> [ , <msg_expr> ]
|
||||||
// The message is optional — if the next token after the condition is not a
|
// The message is optional — if the next token after the condition is not a
|
||||||
// Comma, emit an empty string placeholder so the test still works.
|
// Comma, emit an empty string placeholder so the test still works.
|
||||||
@@ -2419,6 +2526,7 @@ fn scan_params_c(tokens: [Any], pos: Int) -> Map<String, Any> {
|
|||||||
// toplevel_let: { "kind": "toplevel_let", "name": String, "ltype": String }
|
// toplevel_let: { "kind": "toplevel_let", "name": String, "ltype": String }
|
||||||
// cgi_block: { "kind": "cgi_block", "name": String }
|
// cgi_block: { "kind": "cgi_block", "name": String }
|
||||||
// service_block: { "kind": "service_block", "name": String }
|
// service_block: { "kind": "service_block", "name": String }
|
||||||
|
// program_block: { "kind": "program_block", "name": String }
|
||||||
//
|
//
|
||||||
// Import/TypeDef/EnumDef nodes are skipped (codegen treats them as no-ops).
|
// Import/TypeDef/EnumDef nodes are skipped (codegen treats them as no-ops).
|
||||||
//
|
//
|
||||||
@@ -2546,13 +2654,28 @@ fn scan_fn_sigs(tokens: [Any]) -> [Map<String, Any>] {
|
|||||||
"name": name
|
"name": name
|
||||||
})
|
})
|
||||||
let pos = p
|
let pos = p
|
||||||
|
} else {
|
||||||
|
// --- program block ---
|
||||||
|
if str_eq(k, "Program") {
|
||||||
|
let p: Int = pos + 1
|
||||||
|
let name: String = tok_value(tokens, p)
|
||||||
|
let p = p + 1
|
||||||
|
let k2: String = tok_kind(tokens, p)
|
||||||
|
if str_eq(k2, "LBrace") {
|
||||||
|
let p = skip_to_rbrace(tokens, p)
|
||||||
|
}
|
||||||
|
let sigs = native_list_append(sigs, {
|
||||||
|
"kind": "program_block",
|
||||||
|
"name": name
|
||||||
|
})
|
||||||
|
let pos = p
|
||||||
} else {
|
} else {
|
||||||
// Import, Type, Enum, From, or any other token.
|
// Import, Type, Enum, From, or any other token.
|
||||||
// Skip ahead to the next statement boundary.
|
// Skip ahead to the next statement boundary.
|
||||||
let p: Int = pos + 1
|
let p: Int = pos + 1
|
||||||
let p = skip_expr_to_stmt_boundary(tokens, p)
|
let p = skip_expr_to_stmt_boundary(tokens, p)
|
||||||
let pos = p
|
let pos = p
|
||||||
}}}}}
|
}}}}}}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,213 @@
|
|||||||
|
// transduce.el — geometry as a first-class El value, and a realizer written
|
||||||
|
// in El. Runnable: this is the worked example for the transduce surface, and
|
||||||
|
// it doubles as an executable proof because it checks every claim it makes.
|
||||||
|
//
|
||||||
|
// elc lang/examples/transduce.el > transduce.c
|
||||||
|
// cc -std=c11 -O2 -I lang/runtime -o transduce transduce.c \
|
||||||
|
// lang/runtime/el_runtime.c lang/runtime/el_seed.c \
|
||||||
|
// lang/runtime/engram_*.c -lcurl -lpthread -lm
|
||||||
|
// ./transduce # exits 0 only if every check passes
|
||||||
|
//
|
||||||
|
// (A `test "..."` form of the same checks lives in
|
||||||
|
// lang/tests/native/test_transduce.el, for when the native harness is
|
||||||
|
// repaired — the shipped elc currently emits calls to __el_reg_count and
|
||||||
|
// friends without emitting their definitions, which breaks every native test
|
||||||
|
// equally, test_math.el included. Verified 2026-08-16, unrelated to this work.)
|
||||||
|
//
|
||||||
|
// WHY THIS EXISTS. Until 2026-08-16 no El ingest path could carry a vector:
|
||||||
|
// nodes took text, and geometry was DERIVED from that text. Text was the
|
||||||
|
// mandatory entry medium, so any non-text modality had to be DESCRIBED in
|
||||||
|
// prose first and the geometry we reasoned over was the geometry OF THE
|
||||||
|
// DESCRIPTION, not of the signal. Two things fix that, and both are shown
|
||||||
|
// below: geometry is a VALUE that carries its own width, and a REALIZER is an
|
||||||
|
// ordinary El function — so admitting a new modality never requires a runtime
|
||||||
|
// patch.
|
||||||
|
//
|
||||||
|
// COMPARISON DISCIPLINE (measured, not stylistic): elc lowers `a == b`
|
||||||
|
// numerically only when both operand NAMES are in the per-function int-name
|
||||||
|
// set that `let x: Int` populates. A bare `f(x) == 0` is not a registered
|
||||||
|
// name and lowers to str_eq — strcmp on two integers as pointers. `<` and `>`
|
||||||
|
// lower directly with no inference, so truthiness is written `> 0` / `< 1`.
|
||||||
|
|
||||||
|
// ── A realizer, written entirely in El ──────────────────────────────────────
|
||||||
|
// Not in the runtime. Not known to the compiler. Registered by NAME and
|
||||||
|
// dispatched to through transduce(). That is the whole claim.
|
||||||
|
fn tone_realizer(signal: String) -> Geometry {
|
||||||
|
let g: Geometry = geometry_new(4)
|
||||||
|
let n: Int = str_len(signal)
|
||||||
|
let a: Int = geometry_set(g, 0, int_to_float(n))
|
||||||
|
let b: Int = geometry_set(g, 1, int_to_float(n * 2))
|
||||||
|
let c: Int = geometry_set(g, 2, int_to_float(n * 3))
|
||||||
|
let d: Int = geometry_set(g, 3, int_to_float(n * 4))
|
||||||
|
g
|
||||||
|
}
|
||||||
|
|
||||||
|
// A second modality, to show the registry keys on modality rather than just
|
||||||
|
// returning whatever was registered last.
|
||||||
|
fn pulse_realizer(signal: String) -> Geometry {
|
||||||
|
let g: Geometry = geometry_new(2)
|
||||||
|
let a: Int = geometry_set(g, 0, 1.0)
|
||||||
|
let b: Int = geometry_set(g, 1, 0.0)
|
||||||
|
g
|
||||||
|
}
|
||||||
|
|
||||||
|
// A deliberately BROKEN realizer: returns something that is not a Geometry.
|
||||||
|
fn bogus_realizer(signal: String) -> Geometry {
|
||||||
|
return 12345
|
||||||
|
}
|
||||||
|
|
||||||
|
// Fails FAST rather than accumulating a count, for a measured reason: a first
|
||||||
|
// cut wrote `let fails: Int = fails + check(...)` and `+` lowered to STRING
|
||||||
|
// CONCAT, because elc dispatches `+` on whether both operands are known-Int and
|
||||||
|
// a user-defined fn call is not — so the counter printed 4343632752, a pointer.
|
||||||
|
// Nothing was wrong with the checks; the tally was lying. Exiting at the first
|
||||||
|
// failure needs no arithmetic at all, so there is nothing left to get wrong.
|
||||||
|
fn check(ok: Int, label: String) -> Int {
|
||||||
|
if ok > 0 {
|
||||||
|
println(" ok " + label)
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
println(" FAIL " + label)
|
||||||
|
exit(1)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
fn near(a: Float, b: Float) -> Int {
|
||||||
|
let d: Float = a - b
|
||||||
|
if d > 0.001 { return 0 }
|
||||||
|
if d < -0.001 { return 0 }
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
fn eq_int(a: Int, b: Int) -> Int {
|
||||||
|
if a == b { return 1 }
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
fn main() -> Void {
|
||||||
|
println("geometry is a value that carries its own width")
|
||||||
|
let g8: Geometry = geometry_new(8)
|
||||||
|
let _c: Int = check(geometry_is(g8), "geometry_new returns a live Geometry")
|
||||||
|
let d8: Int = geometry_dim(g8)
|
||||||
|
let _c: Int = check(eq_int(d8, 8), "a Geometry carries its own width (8)")
|
||||||
|
let _c: Int = check(geometry_free(g8), "geometry_free reports what it did")
|
||||||
|
|
||||||
|
println("nonsense is refused — with no arbitrary max-dim bound")
|
||||||
|
// #141 needed `dim <= 8192` only to bound an allocation sized from a
|
||||||
|
// caller's CLAIM about a string's length. A value that carries its own
|
||||||
|
// width has nothing left to validate.
|
||||||
|
let z: Geometry = geometry_new(0)
|
||||||
|
let zi: Int = geometry_is(z)
|
||||||
|
let _c: Int = check(1 - zi, "dim 0 is not a geometry")
|
||||||
|
let ng: Geometry = geometry_new(-4)
|
||||||
|
let ngi: Int = geometry_is(ng)
|
||||||
|
let _c: Int = check(1 - ngi, "negative dim is not a geometry")
|
||||||
|
let nd: Int = geometry_dim(0)
|
||||||
|
let _c: Int = check(1 - nd, "geometry_dim of a non-geometry is 0, not a crash")
|
||||||
|
let nf: Int = geometry_free(0)
|
||||||
|
let _c: Int = check(1 - nf, "geometry_free of a non-geometry is a no-op")
|
||||||
|
|
||||||
|
println("components round-trip, and out-of-range is refused")
|
||||||
|
let g3: Geometry = geometry_new(3)
|
||||||
|
let s0: Int = geometry_set(g3, 0, 1.5)
|
||||||
|
let s1: Int = geometry_set(g3, 1, -2.5)
|
||||||
|
let _c: Int = check(s0, "set in range succeeds")
|
||||||
|
let oob: Int = geometry_set(g3, 3, 9.0)
|
||||||
|
let _c: Int = check(1 - oob, "set out of range is refused, not silently dropped")
|
||||||
|
let _c: Int = check(near(geometry_get(g3, 0), 1.5), "component 0 round-trips")
|
||||||
|
let _c: Int = check(near(geometry_get(g3, 1), -2.5), "component 1 round-trips (negative)")
|
||||||
|
let ff3: Int = geometry_free(g3)
|
||||||
|
|
||||||
|
println("hex is an EDGE adapter, and derives its own width")
|
||||||
|
// little-endian float32: 1.0 = 0000803f, 2.0 = 00000040
|
||||||
|
let gh: Geometry = geometry_from_f32le_hex("0000803f00000040")
|
||||||
|
let _c: Int = check(geometry_is(gh), "valid hex decodes to a Geometry")
|
||||||
|
let dh: Int = geometry_dim(gh)
|
||||||
|
let _c: Int = check(eq_int(dh, 2), "width DERIVED from input, never supplied")
|
||||||
|
let _c: Int = check(near(geometry_get(gh, 0), 1.0), "first component decoded")
|
||||||
|
let _c: Int = check(near(geometry_get(gh, 1), 2.0), "second component decoded")
|
||||||
|
let back: String = geometry_to_f32le_hex(gh)
|
||||||
|
let _c: Int = check(str_eq(back, "0000803f00000040"), "hex round-trips exactly")
|
||||||
|
let ffh: Int = geometry_free(gh)
|
||||||
|
|
||||||
|
println("malformed hex is refused")
|
||||||
|
let he: Geometry = geometry_from_f32le_hex("")
|
||||||
|
let hei: Int = geometry_is(he)
|
||||||
|
let _c: Int = check(1 - hei, "empty hex is not a geometry")
|
||||||
|
let hr: Geometry = geometry_from_f32le_hex("0000803f0000")
|
||||||
|
let hri: Int = geometry_is(hr)
|
||||||
|
let _c: Int = check(1 - hri, "length not a multiple of 8 is refused")
|
||||||
|
let hn: Geometry = geometry_from_f32le_hex("zzzzzzzz")
|
||||||
|
let hni: Int = geometry_is(hn)
|
||||||
|
let _c: Int = check(1 - hni, "non-hex characters are refused")
|
||||||
|
|
||||||
|
println("a realizer declared in El is a first-class realizer")
|
||||||
|
let reg: Int = realizer_register("tone", "tone_realizer")
|
||||||
|
let _c: Int = check(reg, "an El fn registers as a realizer BY NAME")
|
||||||
|
let _c: Int = check(realizer_has("tone"), "the modality now has an organ")
|
||||||
|
let gt: Geometry = transduce("aaa", "tone")
|
||||||
|
let _c: Int = check(geometry_is(gt), "transduce returns real geometry")
|
||||||
|
let dt: Int = geometry_dim(gt)
|
||||||
|
let _c: Int = check(eq_int(dt, 4), "the El realizer determined the width, not the runtime")
|
||||||
|
// str_len("aaa") == 3, so component 0 must be 3.0 — proof the signal
|
||||||
|
// actually reached the El function rather than a stub answering for it.
|
||||||
|
let _c: Int = check(near(geometry_get(gt, 0), 3.0), "the signal REACHED the El realizer")
|
||||||
|
let fft: Int = geometry_free(gt)
|
||||||
|
|
||||||
|
println("distinct signals transduce to distinct geometry")
|
||||||
|
let g1: Geometry = transduce("aa", "tone")
|
||||||
|
let g2: Geometry = transduce("aaaaa", "tone")
|
||||||
|
let a1: Float = geometry_get(g1, 0)
|
||||||
|
let a2: Float = geometry_get(g2, 0)
|
||||||
|
// 5 - 2 = 3. If transduction were a stub these would be equal.
|
||||||
|
let _c: Int = check(near(a2 - a1, 3.0), "different signals produce different geometry")
|
||||||
|
let ff1: Int = geometry_free(g1)
|
||||||
|
let ff2: Int = geometry_free(g2)
|
||||||
|
|
||||||
|
println("the registry keys on modality")
|
||||||
|
let r2: Int = realizer_register("pulse", "pulse_realizer")
|
||||||
|
let _c: Int = check(r2, "a second modality registers independently")
|
||||||
|
let mt: Geometry = transduce("aaa", "tone")
|
||||||
|
let mp: Geometry = transduce("aaa", "pulse")
|
||||||
|
let mdt: Int = geometry_dim(mt)
|
||||||
|
let mdp: Int = geometry_dim(mp)
|
||||||
|
let _c: Int = check(eq_int(mdt, 4), "tone still routes to its own realizer")
|
||||||
|
let _c: Int = check(eq_int(mdp, 2), "pulse routes to a different realizer")
|
||||||
|
let ffm1: Int = geometry_free(mt)
|
||||||
|
let ffm2: Int = geometry_free(mp)
|
||||||
|
|
||||||
|
println("no organ is reported as no organ")
|
||||||
|
// A modality with no realizer must transduce to NOTHING. It must never
|
||||||
|
// fall back to embedding a description of the signal and calling that
|
||||||
|
// perception — that silent substitution is the defect this all exists to end.
|
||||||
|
let eh: Int = realizer_has("echolocation")
|
||||||
|
let _c: Int = check(1 - eh, "unregistered modality has no organ")
|
||||||
|
let ge: Geometry = transduce("anything", "echolocation")
|
||||||
|
let gei: Int = geometry_is(ge)
|
||||||
|
let _c: Int = check(1 - gei, "no realizer means NO geometry, not fake geometry")
|
||||||
|
|
||||||
|
println("an unresolvable realizer name fails at WIRING time")
|
||||||
|
let bad: Int = realizer_register("ghost", "no_such_function_anywhere")
|
||||||
|
let _c: Int = check(1 - bad, "unresolvable realizer name is a registration failure")
|
||||||
|
let gh2: Int = realizer_has("ghost")
|
||||||
|
let _c: Int = check(1 - gh2, "and nothing gets registered")
|
||||||
|
|
||||||
|
println("a realizer returning non-geometry transduces nothing")
|
||||||
|
let rb: Int = realizer_register("bogus", "bogus_realizer")
|
||||||
|
let _c: Int = check(rb, "the symbol resolves, so registration succeeds")
|
||||||
|
let gb: Geometry = transduce("x", "bogus")
|
||||||
|
let gbi: Int = geometry_is(gb)
|
||||||
|
let _c: Int = check(1 - gbi, "contract enforced at the boundary: nothing handed back")
|
||||||
|
|
||||||
|
println("norm lets a caller check a realizer emitted signal, not zeros")
|
||||||
|
let gn: Geometry = geometry_new(2)
|
||||||
|
let _c: Int = check(near(geometry_norm(gn), 0.0), "a fresh geometry is zero — norm says so")
|
||||||
|
let n0: Int = geometry_set(gn, 0, 3.0)
|
||||||
|
let n1: Int = geometry_set(gn, 1, 4.0)
|
||||||
|
let _c: Int = check(near(geometry_norm(gn), 5.0), "3-4-5: norm is 5")
|
||||||
|
let ffn: Int = geometry_free(gn)
|
||||||
|
|
||||||
|
// Reaching here means nothing called exit(1) along the way.
|
||||||
|
println("")
|
||||||
|
println("all checks passed")
|
||||||
|
}
|
||||||
+1156
-55
File diff suppressed because it is too large
Load Diff
@@ -586,6 +586,60 @@ void el_runtime_dharma_event_arrive(const char* event_type,
|
|||||||
const char* payload,
|
const char* payload,
|
||||||
const char* source);
|
const char* source);
|
||||||
|
|
||||||
|
/* ── Geometry: signal as a first-class El value ──────────────────────────────
|
||||||
|
*
|
||||||
|
* A Geometry is an opaque, magic-tagged heap value carried in an el_val_t —
|
||||||
|
* the same discipline as List/Map. It holds a width and a float32 payload,
|
||||||
|
* and it is the medium a non-text modality enters in. Declared HERE, above
|
||||||
|
* the engram block, because transduction is a LANGUAGE concern: every El
|
||||||
|
* program touching any modality needs it, and the engram is merely one El
|
||||||
|
* program that happens to hold a graph. See el_runtime.c ("Geometry: signal
|
||||||
|
* as a first-class el value") for the full rationale.
|
||||||
|
*
|
||||||
|
* El-side type annotation is simply `Geometry` — an opaque boxed pointer,
|
||||||
|
* exactly like Instant / Calendar / Rhythm. No codegen change is required.
|
||||||
|
*
|
||||||
|
* OWNERSHIP: a Geometry is owned by the El caller and released with
|
||||||
|
* geometry_free. node_attach_geometry COPIES, so a node and the caller's
|
||||||
|
* value have independent lifetimes. */
|
||||||
|
|
||||||
|
el_val_t geometry_new(el_val_t dim); /* zero-filled; 0 on failure */
|
||||||
|
el_val_t geometry_dim(el_val_t g); /* width, 0 if not a Geometry */
|
||||||
|
el_val_t geometry_is(el_val_t g); /* 1 if a live Geometry */
|
||||||
|
el_val_t geometry_get(el_val_t g, el_val_t i); /* Float component */
|
||||||
|
el_val_t geometry_set(el_val_t g, el_val_t i, el_val_t x); /* 1 ok / 0 out of range */
|
||||||
|
el_val_t geometry_norm(el_val_t g); /* Float L2 — lets a caller
|
||||||
|
* check a realizer emitted
|
||||||
|
* signal, not zeros */
|
||||||
|
el_val_t geometry_free(el_val_t g); /* 1 if freed, 0 if not a Geometry.
|
||||||
|
* Returns a value (not void) so it
|
||||||
|
* is safe in any El expression
|
||||||
|
* position without a codegen
|
||||||
|
* void-builtin table entry. */
|
||||||
|
|
||||||
|
/* Wire ADAPTERS — the only place an encoding appears, and only at the edge.
|
||||||
|
* `f32le hex` is little-endian float32, 8 hex chars per component: the
|
||||||
|
* encoding the perception vessel's /voice/embed already emits. The width is
|
||||||
|
* DERIVED from the input length, never supplied by a caller — which is why
|
||||||
|
* there is no max-dim constant here to validate a claimed length against. */
|
||||||
|
el_val_t geometry_from_f32le_hex(el_val_t hex); /* 0 on empty/odd-length/non-hex */
|
||||||
|
el_val_t geometry_to_f32le_hex(el_val_t g); /* "" if not a Geometry */
|
||||||
|
|
||||||
|
/* ── Realizers + transduce ───────────────────────────────────────────────────
|
||||||
|
* A REALIZER maps one modality into geometry. Registration is by NAME, so a
|
||||||
|
* new modality never requires a runtime patch: every El `fn name(...)`
|
||||||
|
* compiles to a global C symbol with that exact name, and the registry
|
||||||
|
* resolves it with dlsym against the running binary — the same mechanism
|
||||||
|
* http_set_handler already relies on.
|
||||||
|
*
|
||||||
|
* fn tone_realizer(signal: String) -> Geometry { ... }
|
||||||
|
* realizer_register("tone", "tone_realizer")
|
||||||
|
* let g: Geometry = transduce(sample, "tone")
|
||||||
|
*/
|
||||||
|
el_val_t realizer_register(el_val_t modality, el_val_t fn_name); /* 1 ok / 0 unresolved */
|
||||||
|
el_val_t realizer_has(el_val_t modality); /* 1 if a realizer is registered */
|
||||||
|
el_val_t transduce(el_val_t signal, el_val_t modality); /* Geometry, or 0 if no organ */
|
||||||
|
|
||||||
/* ── Engram local graph primitives ───────────────────────────────────────────
|
/* ── Engram local graph primitives ───────────────────────────────────────────
|
||||||
* Operate on the CGI's local Engram knowledge graph.
|
* Operate on the CGI's local Engram knowledge graph.
|
||||||
* `engram_activate` queries the local graph only; `dharma_activate` is
|
* `engram_activate` queries the local graph only; `dharma_activate` is
|
||||||
@@ -612,7 +666,28 @@ el_val_t engram_get_node(el_val_t id);
|
|||||||
void engram_strengthen(el_val_t node_id);
|
void engram_strengthen(el_val_t node_id);
|
||||||
void engram_forget(el_val_t node_id);
|
void engram_forget(el_val_t node_id);
|
||||||
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
|
el_val_t engram_prune_telemetry(el_val_t older_than_ms);
|
||||||
|
/* Largest byte length <= max_bytes that does not split a UTF-8 codepoint.
|
||||||
|
* Bounded by bytes, not codepoints, so truncated strings never grow. */
|
||||||
|
size_t el_utf8_safe_len(const char* s, size_t max_bytes);
|
||||||
|
|
||||||
el_val_t engram_node_count(void);
|
el_val_t engram_node_count(void);
|
||||||
|
/* Attach a Geometry to an existing node, and read the attached width back.
|
||||||
|
* Named for the operation, not the store: a node acquires geometry. This is
|
||||||
|
* the geometry-valued ingest path — nothing about it is hex, and nothing
|
||||||
|
* about it assumes the caller's vector matches the canonical text-embedding
|
||||||
|
* width. node_geometry_dim exists so an attach is VERIFIED by reading it
|
||||||
|
* back rather than by trusting a success return. */
|
||||||
|
el_val_t node_attach_geometry(el_val_t node_id, el_val_t g); /* 1 ok / 0 otherwise */
|
||||||
|
el_val_t node_geometry_dim(el_val_t node_id); /* width, 0 if none */
|
||||||
|
|
||||||
|
/* DEPRECATED (shipped in #141, superseded 2026-08-16). Equivalent to
|
||||||
|
* geometry_from_f32le_hex + node_attach_geometry, and now implemented as
|
||||||
|
* exactly that. Kept only so anything built against the #141 runtime keeps
|
||||||
|
* linking; `dim` is accepted but treated as an assertion about the vector's
|
||||||
|
* width rather than as its source. New code should not call this — a hex
|
||||||
|
* string is a wire encoding, not a way to move geometry between two pieces
|
||||||
|
* of El. Returns 1 on success, 0 otherwise. */
|
||||||
|
el_val_t engram_node_set_emb(el_val_t id, el_val_t hex, el_val_t dim);
|
||||||
el_val_t engram_search(el_val_t query, el_val_t limit);
|
el_val_t engram_search(el_val_t query, el_val_t limit);
|
||||||
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
|
el_val_t engram_scan_nodes(el_val_t limit, el_val_t offset);
|
||||||
void engram_connect(el_val_t from_id, el_val_t to_id, el_val_t weight, el_val_t relation);
|
void engram_connect(el_val_t from_id, el_val_t to_id, el_val_t weight, el_val_t relation);
|
||||||
@@ -952,6 +1027,22 @@ el_val_t __url_decode(el_val_t s);
|
|||||||
/* Environment */
|
/* Environment */
|
||||||
el_val_t __env_get(el_val_t key);
|
el_val_t __env_get(el_val_t key);
|
||||||
|
|
||||||
|
/* Cross-cutting concerns declared by a `program` block (spec §18).
|
||||||
|
* All three are COMPILER-INJECTED at the head of main() — they are not meant to
|
||||||
|
* be written by hand, which is the point: the guarantee cannot be forgotten at a
|
||||||
|
* call site because there is no call site. */
|
||||||
|
el_val_t el_singleton_acquire(el_val_t id); /* §18.1 process identity */
|
||||||
|
el_val_t el_config_declare(el_val_t name, el_val_t type,
|
||||||
|
el_val_t deflt, el_val_t has_default,
|
||||||
|
el_val_t required); /* §18.2 config schema */
|
||||||
|
el_val_t el_config_validate(el_val_t program_name); /* §18.2 startup validate */
|
||||||
|
|
||||||
|
/* config(key) — the READ side, and the only one programs write by hand. With a
|
||||||
|
* schema declared it is a validated lookup; without one it degrades to getenv.
|
||||||
|
* (Defined in el_runtime.c but previously never prototyped here, so any program
|
||||||
|
* calling it failed to compile under -Werror=implicit-function-declaration.) */
|
||||||
|
el_val_t config(el_val_t key);
|
||||||
|
|
||||||
/* Subprocess */
|
/* Subprocess */
|
||||||
el_val_t __exec(el_val_t cmd);
|
el_val_t __exec(el_val_t cmd);
|
||||||
el_val_t __exec_bg(el_val_t cmd);
|
el_val_t __exec_bg(el_val_t cmd);
|
||||||
@@ -1022,6 +1113,7 @@ el_val_t el_mem_check(void);
|
|||||||
el_val_t el_alloc_count(void);
|
el_val_t el_alloc_count(void);
|
||||||
el_val_t el_alloc_bytes(void);
|
el_val_t el_alloc_bytes(void);
|
||||||
el_val_t el_peak_rss(void);
|
el_val_t el_peak_rss(void);
|
||||||
|
el_val_t el_black_box(el_val_t v);
|
||||||
|
|
||||||
/* Semantic retrieval surface. NOT interchangeable with engram_search_json,
|
/* Semantic retrieval surface. NOT interchangeable with engram_search_json,
|
||||||
* which is lexical by design — see the note at the definition. */
|
* which is lexical by design — see the note at the definition. */
|
||||||
|
|||||||
@@ -0,0 +1,256 @@
|
|||||||
|
// runtime/elbench.el — growth-curve classifier and complexity gate.
|
||||||
|
//
|
||||||
|
// Given a geometric sweep of input sizes and the measurements taken at each,
|
||||||
|
// classify the growth curve and decide whether it violates a declared bound.
|
||||||
|
//
|
||||||
|
// ── Why this exists ──────────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// Constant-factor regressions are annoying. Complexity regressions are outages.
|
||||||
|
// An O(n) lookup inside an O(n) loop is invisible at n=100 in a unit test and
|
||||||
|
// catastrophic at n=100000 in production. el #132 was exactly that: a strlen()
|
||||||
|
// inside a per-character accessor, quadratic, shipped for months.
|
||||||
|
//
|
||||||
|
// ── THREE signals, not one ───────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// The gate fits time AND allocation-count AND allocation-bytes, and fails if
|
||||||
|
// ANY of them exceeds its declared curve. This is not belt-and-braces; each
|
||||||
|
// signal is blind to a real defect class the others catch:
|
||||||
|
//
|
||||||
|
// * A copy-on-write accumulator rebuilding its buffer allocates ONCE per
|
||||||
|
// iteration — count is exactly linear — while bytes go quadratic.
|
||||||
|
// Count alone passes it.
|
||||||
|
// * el #132's strlen-per-character is pure CPU and allocates NOTHING.
|
||||||
|
// Both allocation signals read FLAT. Only time catches it.
|
||||||
|
//
|
||||||
|
// The deterministic signals (count, bytes) are preferable where they apply:
|
||||||
|
// no statistics, correct on the first run, machine-independent. They are
|
||||||
|
// simply not sufficient.
|
||||||
|
//
|
||||||
|
// ── SCOPE LIMIT — read this before trusting a flat curve ─────────────────────
|
||||||
|
//
|
||||||
|
// The allocation counters track EL-LEVEL allocation only: strings, ElList and
|
||||||
|
// ElMap bodies, their backing arrays, copy-on-write clones, and the realloc
|
||||||
|
// growth path. malloc inside engram_*.c and inside libcurl is NOT counted.
|
||||||
|
//
|
||||||
|
// A flat allocation curve over a workload dominated by engram or HTTP calls is
|
||||||
|
// therefore NOT evidence of anything. It means "no El-level allocation growth",
|
||||||
|
// not "no allocation growth". Gate El-level complexity with this; do not read
|
||||||
|
// third-party memory behaviour into it.
|
||||||
|
//
|
||||||
|
// ── Classification method ────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// Sizes must form a geometric sweep (each n double the last). On such a sweep
|
||||||
|
// the ratio between consecutive measurements IS the growth exponent, directly:
|
||||||
|
//
|
||||||
|
// O(1) -> 1.0 O(log n) -> ~1.1 O(n) -> 2.0
|
||||||
|
// O(n log n) -> ~2.2 O(n^2) -> 4.0 O(n^3) -> 8.0
|
||||||
|
//
|
||||||
|
// DEVIATION FROM DESIGN.md 6.2, stated plainly: that section specified Google
|
||||||
|
// Benchmark's one-parameter least-squares fit over candidate curves. This uses
|
||||||
|
// consecutive ratios instead. The sweep is mandated geometric either way, and
|
||||||
|
// on a geometric sweep ratios are directly interpretable and need no floating
|
||||||
|
// point. The cost is weaker separation between O(n) and O(n log n), which is
|
||||||
|
// reported honestly as an ambiguous band rather than guessed at. Least-squares
|
||||||
|
// remains the better answer if that band ever needs to be resolved.
|
||||||
|
//
|
||||||
|
// All arithmetic is fixed-point, scaled by 1000 ("milli-ratio"), so a ratio of
|
||||||
|
// 2.0 is 2000. El values are int64; this avoids float-in-list handling.
|
||||||
|
|
||||||
|
// Curve identifiers. Ordered by growth — the ordering IS the comparison used
|
||||||
|
// by the gate, so an index comparison decides "worse than declared".
|
||||||
|
// 0 = O(1) 1 = O(log n) 2 = O(n) 3 = O(n log n) 4 = O(n^2) 5 = O(n^3)
|
||||||
|
|
||||||
|
fn elb_curve_name(c: Int) -> String {
|
||||||
|
if c == 0 { return "O(1)" }
|
||||||
|
if c == 1 { return "O(log n)" }
|
||||||
|
if c == 2 { return "O(n)" }
|
||||||
|
if c == 3 { return "O(n log n)" }
|
||||||
|
if c == 4 { return "O(n^2)" }
|
||||||
|
if c == 5 { return "O(n^3)" }
|
||||||
|
return "O(?)"
|
||||||
|
}
|
||||||
|
|
||||||
|
fn elb_curve_from_name(s: String) -> Int {
|
||||||
|
if str_eq(s, "O(1)") { return 0 }
|
||||||
|
if str_eq(s, "O(log n)") { return 1 }
|
||||||
|
if str_eq(s, "O(n)") { return 2 }
|
||||||
|
if str_eq(s, "O(n log n)") { return 3 }
|
||||||
|
if str_eq(s, "O(n^2)") { return 4 }
|
||||||
|
if str_eq(s, "O(n^3)") { return 5 }
|
||||||
|
return -1
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_classify_ratio — map a milli-ratio-per-doubling onto a curve.
|
||||||
|
//
|
||||||
|
// Bands are deliberately wide at the top (a quadratic measured at 3.4x is
|
||||||
|
// still a quadratic) and deliberately overlap-averse at the bottom, where a
|
||||||
|
// misclassification between O(1) and O(log n) matters least.
|
||||||
|
fn elb_classify_ratio(milli: Int) -> Int {
|
||||||
|
if milli < 1300 { return 0 }
|
||||||
|
if milli < 1700 { return 1 }
|
||||||
|
if milli < 2400 { return 2 }
|
||||||
|
if milli < 3200 { return 3 }
|
||||||
|
if milli < 6000 { return 4 }
|
||||||
|
return 5
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_ratio — milli-ratio between two consecutive measurements.
|
||||||
|
// Returns -1 when the earlier measurement is zero (ratio undefined).
|
||||||
|
fn elb_ratio(prev: Int, cur: Int) -> Int {
|
||||||
|
if prev <= 0 { return -1 }
|
||||||
|
return (cur * 1000) / prev
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── The measurement floor ────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// A benchmark whose largest measurement is at or near zero has not been
|
||||||
|
// measured. Reporting it as O(1) would be a confident answer with nothing
|
||||||
|
// behind it — the same failure as a test that never ran reporting pass, and
|
||||||
|
// exactly what happened when clang closed a nested loop to a multiply and the
|
||||||
|
// harness read 0 microseconds at every n.
|
||||||
|
//
|
||||||
|
// So: REFUSE. Never classify below the floor.
|
||||||
|
fn elb_below_floor(vals: [Int], floor: Int) -> Bool {
|
||||||
|
let n: Int = native_list_len(vals)
|
||||||
|
let i: Int = 0
|
||||||
|
let mx: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let v: Int = native_list_get(vals, i)
|
||||||
|
if v > mx { let mx = v }
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
if mx < floor { return true }
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_implausibly_flat — a measurement that does not move across a sweep whose
|
||||||
|
// input grew by 8x or more is not a flat curve, it is a broken measurement.
|
||||||
|
// Genuine O(1) work still shows noise; a hard-flat series means the work was
|
||||||
|
// optimised away, the timer has insufficient resolution, or the benchmark body
|
||||||
|
// never executed.
|
||||||
|
fn elb_implausibly_flat(vals: [Int]) -> Bool {
|
||||||
|
let n: Int = native_list_len(vals)
|
||||||
|
if n < 3 { return false }
|
||||||
|
let first: Int = native_list_get(vals, 0)
|
||||||
|
let last: Int = native_list_get(vals, n - 1)
|
||||||
|
if first == 0 {
|
||||||
|
if last == 0 { return true }
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
let r: Int = (last * 1000) / first
|
||||||
|
if r < 1100 { return true }
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_spread_ok — do the consecutive ratios agree with each other?
|
||||||
|
//
|
||||||
|
// This is the ratio-method analogue of a normalised-RMS threshold. If the
|
||||||
|
// doublings disagree wildly the data is noise, a cache cliff, or a phase
|
||||||
|
// change, and the honest report is INDETERMINATE rather than a classification.
|
||||||
|
// Applies to the ASYMPTOTIC TAIL only — the last three ratios.
|
||||||
|
//
|
||||||
|
// The small-n end of any sweep is dominated by fixed overhead, cold caches and
|
||||||
|
// branch predictors that have not warmed. Measured on a genuinely linear
|
||||||
|
// character scan, the ratios ran 3.37, 2.92, 1.76, 1.65: the head looks
|
||||||
|
// quadratic, the tail is the truth. Checking spread across the whole sweep
|
||||||
|
// therefore rejects correct data. A complexity bound is an asymptotic claim, so
|
||||||
|
// it is judged on the asymptotic region — the same reason a benchmark harness
|
||||||
|
// discards warmup rather than averaging it in.
|
||||||
|
fn elb_spread_ok(ratios: [Int]) -> Bool {
|
||||||
|
let total: Int = native_list_len(ratios)
|
||||||
|
if total < 2 { return true }
|
||||||
|
let start: Int = total - 3
|
||||||
|
if start < 0 { let start = 0 }
|
||||||
|
let n: Int = total
|
||||||
|
let lo: Int = 999999
|
||||||
|
let hi: Int = 0
|
||||||
|
let i: Int = start
|
||||||
|
while i < n {
|
||||||
|
let r: Int = native_list_get(ratios, i)
|
||||||
|
if r >= 0 {
|
||||||
|
if r < lo { let lo = r }
|
||||||
|
if r > hi { let hi = r }
|
||||||
|
}
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
if lo <= 0 { return false }
|
||||||
|
// Reject when the widest ratio is more than 2.2x the narrowest. That is
|
||||||
|
// enough slack for real timing noise and tight enough to separate a clean
|
||||||
|
// 2.0 series from a clean 4.0 series.
|
||||||
|
if (hi * 1000) / lo > 2200 { return false }
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_ratios — consecutive milli-ratios across the sweep.
|
||||||
|
fn elb_ratios(vals: [Int]) -> [Int] {
|
||||||
|
let out: [Int] = native_list_empty()
|
||||||
|
let n: Int = native_list_len(vals)
|
||||||
|
let i: Int = 1
|
||||||
|
while i < n {
|
||||||
|
let out = native_list_append(out,
|
||||||
|
elb_ratio(native_list_get(vals, i - 1), native_list_get(vals, i)))
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_mean_tail_ratio — mean of the LAST TWO ratios.
|
||||||
|
//
|
||||||
|
// The tail is used deliberately: asymptotic behaviour is what a complexity
|
||||||
|
// bound claims, and the small-n end of any sweep is dominated by fixed
|
||||||
|
// overhead. This is the same reason a benchmark harness discards warmup.
|
||||||
|
fn elb_mean_tail_ratio(ratios: [Int]) -> Int {
|
||||||
|
let n: Int = native_list_len(ratios)
|
||||||
|
if n == 0 { return -1 }
|
||||||
|
if n == 1 { return native_list_get(ratios, 0) }
|
||||||
|
let a: Int = native_list_get(ratios, n - 1)
|
||||||
|
let b: Int = native_list_get(ratios, n - 2)
|
||||||
|
if a < 0 { return b }
|
||||||
|
if b < 0 { return a }
|
||||||
|
return (a + b) / 2
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Verdicts ─────────────────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// 0 PASS measured curve is at or below the declared bound
|
||||||
|
// 1 FAIL measured curve is strictly worse than declared
|
||||||
|
// 2 INDETERMINATE ratios disagree; data is noise or a phase change
|
||||||
|
// 3 REFUSED below the measurement floor, or implausibly flat
|
||||||
|
// 4 BETTER measured strictly better than declared (warn, not fail)
|
||||||
|
|
||||||
|
fn elb_verdict_name(v: Int) -> String {
|
||||||
|
if v == 0 { return "PASS" }
|
||||||
|
if v == 1 { return "FAIL" }
|
||||||
|
if v == 2 { return "INDETERMINATE" }
|
||||||
|
if v == 3 { return "REFUSED" }
|
||||||
|
if v == 4 { return "BETTER" }
|
||||||
|
return "?"
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_gate — classify one signal against its declared bound.
|
||||||
|
//
|
||||||
|
// vals measurements, one per sweep point, in sweep order
|
||||||
|
// expect declared curve index (see elb_curve_name)
|
||||||
|
// floor minimum largest-measurement below which we refuse to classify
|
||||||
|
fn elb_gate(vals: [Int], expect: Int, floor: Int) -> Int {
|
||||||
|
if elb_below_floor(vals, floor) { return 3 }
|
||||||
|
if elb_implausibly_flat(vals) { return 3 }
|
||||||
|
let ratios: [Int] = elb_ratios(vals)
|
||||||
|
if !elb_spread_ok(ratios) { return 2 }
|
||||||
|
let m: Int = elb_mean_tail_ratio(ratios)
|
||||||
|
if m < 0 { return 2 }
|
||||||
|
let got: Int = elb_classify_ratio(m)
|
||||||
|
if got > expect { return 1 }
|
||||||
|
if got < expect { return 4 }
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
// elb_measured_curve — the classified curve for a signal, or -1 if unclassifiable.
|
||||||
|
fn elb_measured_curve(vals: [Int], floor: Int) -> Int {
|
||||||
|
if elb_below_floor(vals, floor) { return -1 }
|
||||||
|
if elb_implausibly_flat(vals) { return -1 }
|
||||||
|
let ratios: [Int] = elb_ratios(vals)
|
||||||
|
let m: Int = elb_mean_tail_ratio(ratios)
|
||||||
|
if m < 0 { return -1 }
|
||||||
|
return elb_classify_ratio(m)
|
||||||
|
}
|
||||||
@@ -0,0 +1,194 @@
|
|||||||
|
// runtime/eltest.el — El test framework runner (Phase 1).
|
||||||
|
//
|
||||||
|
// This is the RUNNER. It is written in El and consumes a registry that the
|
||||||
|
// compiler generates into the same translation unit when invoked as
|
||||||
|
// `elc --test`. Nothing here discovers tests; discovery already happened at
|
||||||
|
// compile time, which is what makes `--list` and filtering possible later.
|
||||||
|
//
|
||||||
|
// ── Architecture ─────────────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// The compiler lowers each `test "name" { ... }` block into a static C
|
||||||
|
// function and emits a static table of (name, fn) pairs plus a small set of
|
||||||
|
// index-based accessors. El has no function pointers, so the runner never
|
||||||
|
// sees one — it works entirely in indices:
|
||||||
|
//
|
||||||
|
// __el_reg_count() -> Int number of registered tests
|
||||||
|
// __el_reg_name(i) -> String test name at index i
|
||||||
|
// __el_reg_invoke(i) -> Int run test i, return its failure count
|
||||||
|
// __el_reg_last_ns() -> Int wall-clock ns of the last invoke
|
||||||
|
// __el_reg_msg() -> String first failure message of the last invoke
|
||||||
|
// __el_reg_asserts() -> Int assertions executed in the last invoke
|
||||||
|
// __el_opt_json() -> Int 1 if --json was passed
|
||||||
|
//
|
||||||
|
// Timing is taken in the generated C, immediately around the call, so no El
|
||||||
|
// call overhead lands inside the measurement.
|
||||||
|
//
|
||||||
|
// ── Output ───────────────────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// Structured events are the source of truth. The human renderer is written
|
||||||
|
// FROM the same fields the NDJSON renderer emits — never the reverse. Parsing
|
||||||
|
// human output back into structure is the one clear architectural mistake in
|
||||||
|
// Go's test tooling and we do not repeat it.
|
||||||
|
//
|
||||||
|
// Every result carries a duration. Always. A framework that cannot report how
|
||||||
|
// long its tests took cannot surface a performance regression, and a
|
||||||
|
// regression nobody can see is one nobody fixes.
|
||||||
|
|
||||||
|
// ── Small helpers (no imports — this file must stay self-contained) ──────────
|
||||||
|
|
||||||
|
// _elt_json_escape — minimal JSON string escaping for the NDJSON renderer.
|
||||||
|
fn _elt_json_escape(s: String) -> String {
|
||||||
|
let out: String = ""
|
||||||
|
let n: Int = str_len(s)
|
||||||
|
let i: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let ch: String = str_slice(s, i, i + 1)
|
||||||
|
if str_eq(ch, "\"") {
|
||||||
|
let out = out + "\\\""
|
||||||
|
} else {
|
||||||
|
if str_eq(ch, "\\") {
|
||||||
|
let out = out + "\\\\"
|
||||||
|
} else {
|
||||||
|
if str_eq(ch, "\n") {
|
||||||
|
let out = out + "\\n"
|
||||||
|
} else {
|
||||||
|
if str_eq(ch, "\t") {
|
||||||
|
let out = out + "\\t"
|
||||||
|
} else {
|
||||||
|
if str_eq(ch, "\r") {
|
||||||
|
let out = out + "\\r"
|
||||||
|
} else {
|
||||||
|
let out = out + ch
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// _elt_pad3 — left-pad an integer to three digits (for the ms.fraction form).
|
||||||
|
fn _elt_pad3(v: Int) -> String {
|
||||||
|
if v < 10 { return "00" + int_to_str(v) }
|
||||||
|
if v < 100 { return "0" + int_to_str(v) }
|
||||||
|
return int_to_str(v)
|
||||||
|
}
|
||||||
|
|
||||||
|
// _elt_ms — render a nanosecond duration as "M.mmm" milliseconds.
|
||||||
|
//
|
||||||
|
// Deliberately avoids the modulo operator: the remainder is derived by
|
||||||
|
// subtraction so this stays portable across El backends.
|
||||||
|
fn _elt_ms(ns: Int) -> String {
|
||||||
|
let total_us: Int = ns / 1000
|
||||||
|
let ms_whole: Int = total_us / 1000
|
||||||
|
let us_rem: Int = total_us - (ms_whole * 1000)
|
||||||
|
return int_to_str(ms_whole) + "." + _elt_pad3(us_rem)
|
||||||
|
}
|
||||||
|
|
||||||
|
// _elt_secs — render a nanosecond duration as fractional seconds, for the
|
||||||
|
// NDJSON `elapsed` field. JUnit XML and test2json both use seconds-as-decimal.
|
||||||
|
fn _elt_secs(ns: Int) -> String {
|
||||||
|
let total_ms: Int = ns / 1000000
|
||||||
|
let s_whole: Int = total_ms / 1000
|
||||||
|
let ms_rem: Int = total_ms - (s_whole * 1000)
|
||||||
|
return int_to_str(s_whole) + "." + _elt_pad3(ms_rem)
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Event emission ───────────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// One function per event shape. Both renderers read the same fields; the
|
||||||
|
// human renderer is a projection of the event, not a separate code path.
|
||||||
|
|
||||||
|
fn _elt_emit_run(json_mode: Bool, name: String) {
|
||||||
|
if json_mode {
|
||||||
|
println("{\"action\":\"run\",\"test\":\"" + _elt_json_escape(name) + "\"}")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn _elt_emit_result(json_mode: Bool, name: String, fails: Int, ns: Int, asserts: Int, msg: String) {
|
||||||
|
if json_mode {
|
||||||
|
let action: String = "pass"
|
||||||
|
if fails > 0 { let action = "fail" }
|
||||||
|
let line: String = "{\"action\":\"" + action + "\""
|
||||||
|
let line = line + ",\"test\":\"" + _elt_json_escape(name) + "\""
|
||||||
|
let line = line + ",\"elapsed\":" + _elt_secs(ns)
|
||||||
|
let line = line + ",\"assertions\":" + int_to_str(asserts)
|
||||||
|
if fails > 0 {
|
||||||
|
let line = line + ",\"failures\":" + int_to_str(fails)
|
||||||
|
let line = line + ",\"message\":\"" + _elt_json_escape(msg) + "\""
|
||||||
|
}
|
||||||
|
let line = line + "}"
|
||||||
|
println(line)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
// Human renderer — duration is never optional.
|
||||||
|
if fails > 0 {
|
||||||
|
println("FAIL " + name + " (" + _elt_ms(ns) + "ms)")
|
||||||
|
println(" " + msg)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
println("ok " + name + " (" + _elt_ms(ns) + "ms)")
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
fn _elt_emit_summary(json_mode: Bool, total: Int, failed: Int, ns: Int, asserts: Int) {
|
||||||
|
let passed: Int = total - failed
|
||||||
|
if json_mode {
|
||||||
|
let line: String = "{\"action\":\"summary\""
|
||||||
|
let line = line + ",\"tests\":" + int_to_str(total)
|
||||||
|
let line = line + ",\"passed\":" + int_to_str(passed)
|
||||||
|
let line = line + ",\"failed\":" + int_to_str(failed)
|
||||||
|
let line = line + ",\"assertions\":" + int_to_str(asserts)
|
||||||
|
let line = line + ",\"elapsed\":" + _elt_secs(ns)
|
||||||
|
let line = line + "}"
|
||||||
|
println(line)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
println("")
|
||||||
|
println(int_to_str(total) + " tests, " + int_to_str(passed) + " passed, "
|
||||||
|
+ int_to_str(failed) + " failed, " + int_to_str(asserts) + " assertions in "
|
||||||
|
+ _elt_ms(ns) + "ms")
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── The runner ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
// el_test_main — drive the compile-time registry.
|
||||||
|
//
|
||||||
|
// Called from the generated main(). Returns the number of FAILING TESTS, which
|
||||||
|
// becomes the process exit code. Note that this counts tests, not assertions:
|
||||||
|
// a test is the unit of result. The old harness counted assertions globally and
|
||||||
|
// therefore could not say which test failed, how long any of them took, or
|
||||||
|
// whether a test had run at all.
|
||||||
|
fn el_test_main() -> Int {
|
||||||
|
let json_mode: Bool = false
|
||||||
|
if __el_opt_json() == 1 { let json_mode = true }
|
||||||
|
|
||||||
|
let n: Int = __el_reg_count()
|
||||||
|
let i: Int = 0
|
||||||
|
let failed: Int = 0
|
||||||
|
let total_ns: Int = 0
|
||||||
|
let total_asserts: Int = 0
|
||||||
|
|
||||||
|
while i < n {
|
||||||
|
let name: String = __el_reg_name(i)
|
||||||
|
_elt_emit_run(json_mode, name)
|
||||||
|
|
||||||
|
let fails: Int = __el_reg_invoke(i)
|
||||||
|
let ns: Int = __el_reg_last_ns()
|
||||||
|
let asserts: Int = __el_reg_asserts()
|
||||||
|
let msg: String = __el_reg_msg()
|
||||||
|
|
||||||
|
let total_ns = total_ns + ns
|
||||||
|
let total_asserts = total_asserts + asserts
|
||||||
|
if fails > 0 { let failed = failed + 1 }
|
||||||
|
|
||||||
|
_elt_emit_result(json_mode, name, fails, ns, asserts, msg)
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
|
||||||
|
_elt_emit_summary(json_mode, n, failed, total_ns, total_asserts)
|
||||||
|
return failed
|
||||||
|
}
|
||||||
@@ -222,7 +222,7 @@ static double eff_w(double weight, double hebb){
|
|||||||
}
|
}
|
||||||
|
|
||||||
GeoDescriptor* engram_geometry_descriptor(
|
GeoDescriptor* engram_geometry_descriptor(
|
||||||
EngramPagedStore* store, VIndex* vindex,
|
EngramPagedStore* store, const VIndex* vindex,
|
||||||
char** vids, int n_vids,
|
char** vids, int n_vids,
|
||||||
const char* const* seed_ids, size_t n_seeds,
|
const char* const* seed_ids, size_t n_seeds,
|
||||||
const GeoParams* params,
|
const GeoParams* params,
|
||||||
@@ -1401,7 +1401,7 @@ static double geo_weighted_degree(EngramPagedStore* st, const char* id, double e
|
|||||||
return deg;
|
return deg;
|
||||||
}
|
}
|
||||||
|
|
||||||
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex,
|
int engram_geo_reify_store(EngramPagedStore* store, const VIndex* vindex,
|
||||||
char** vids, int n_vids,
|
char** vids, int n_vids,
|
||||||
const GeoReifyParams* params){
|
const GeoReifyParams* params){
|
||||||
if(!store) return -1;
|
if(!store) return -1;
|
||||||
|
|||||||
@@ -150,7 +150,7 @@ void engram_geo_mean_free(GeoMeanCache* c);
|
|||||||
* Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error
|
* Returns a malloc'd descriptor (free with engram_geo_free), or NULL on error
|
||||||
* (no seeds resolvable, OOM). */
|
* (no seeds resolvable, OOM). */
|
||||||
GeoDescriptor* engram_geometry_descriptor(
|
GeoDescriptor* engram_geometry_descriptor(
|
||||||
EngramPagedStore* store, VIndex* vindex,
|
EngramPagedStore* store, const VIndex* vindex,
|
||||||
char** vids, int n_vids,
|
char** vids, int n_vids,
|
||||||
const char* const* seed_ids, size_t n_seeds,
|
const char* const* seed_ids, size_t n_seeds,
|
||||||
const GeoParams* params,
|
const GeoParams* params,
|
||||||
@@ -375,7 +375,7 @@ void engram_geo_reify_default_params(GeoReifyParams* p);
|
|||||||
* neighborhood (+ member edges), superseding any prior same-hub record with
|
* neighborhood (+ member edges), superseding any prior same-hub record with
|
||||||
* provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0.
|
* provenance. Read-then-write over `store`. Returns #neighborhoods persisted, or <0.
|
||||||
* Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */
|
* Skips existing Neighborhood/GeoMeanFrame nodes when detecting (idempotent re-reify). */
|
||||||
int engram_geo_reify_store(EngramPagedStore* store, VIndex* vindex,
|
int engram_geo_reify_store(EngramPagedStore* store, const VIndex* vindex,
|
||||||
char** vids, int n_vids,
|
char** vids, int n_vids,
|
||||||
const GeoReifyParams* params);
|
const GeoReifyParams* params);
|
||||||
|
|
||||||
|
|||||||
@@ -74,11 +74,6 @@ struct VIndex {
|
|||||||
|
|
||||||
int entry; /* entry-point element index, -1 if empty */
|
int entry; /* entry-point element index, -1 if empty */
|
||||||
int max_level; /* current top layer */
|
int max_level; /* current top layer */
|
||||||
|
|
||||||
/* scratch: version-stamped visited set (O(1) reset). */
|
|
||||||
uint32_t* visited;
|
|
||||||
uint32_t visit_epoch;
|
|
||||||
size_t visited_cap;
|
|
||||||
};
|
};
|
||||||
|
|
||||||
/* ── small helpers ────────────────────────────────────────────────────────── */
|
/* ── small helpers ────────────────────────────────────────────────────────── */
|
||||||
@@ -166,37 +161,63 @@ static Pair heap_pop(Heap* h, int is_max){
|
|||||||
return top;
|
return top;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* ── visited set ──────────────────────────────────────────────────────────── */
|
/* ── visited set — owned by the CALL FRAME, never by the index ──────────────
|
||||||
static int visited_ensure(VIndex* ix){
|
* This buffer is per-TRAVERSAL scratch. It used to live in struct VIndex as an
|
||||||
if (ix->visited_cap >= ix->cap && ix->visited) return 0;
|
* allocation optimisation, which made every traversal a write to shared state:
|
||||||
size_t nc = ix->cap ? ix->cap : 16;
|
* two concurrent vindex_search calls stamped each other's epoch and then walked
|
||||||
uint32_t* nv = (uint32_t*)realloc(ix->visited, nc*sizeof(uint32_t));
|
* each other's marks, so even two pure READS corrupted the traversal (measured
|
||||||
if (!nv) return -1;
|
* 2026-08-16: TSan data race at visited_reset, reached from vindex_search on one
|
||||||
if (nc > ix->visited_cap) memset(nv + ix->visited_cap, 0, (nc-ix->visited_cap)*sizeof(uint32_t));
|
* thread and vindex_insert on another; downstream SIGSEGV dereferencing a bogus
|
||||||
ix->visited = nv; ix->visited_cap = nc;
|
* element index).
|
||||||
|
*
|
||||||
|
* It is not an ownership problem and it does not want a lock or a capability —
|
||||||
|
* it was simply misfiled. A pure function's scratch belongs to the call. Moving
|
||||||
|
* it here is what lets vindex_search take a `const VIndex*`, which is in turn
|
||||||
|
* what makes "search does not mutate the index" a COMPILE-TIME property instead
|
||||||
|
* of a review comment.
|
||||||
|
*
|
||||||
|
* Cost: one calloc/free of cap*4 bytes per traversal (~55 KB at the live store's
|
||||||
|
* 13,820 elements), against thousands of dim-768 dot products in the same call.
|
||||||
|
* Deliberately NOT __thread: http_worker is a thread per connection, so a
|
||||||
|
* thread-local buffer would retain ~55 KB per connection for the process life. */
|
||||||
|
typedef struct {
|
||||||
|
uint32_t* mark; /* per-element epoch stamp */
|
||||||
|
uint32_t epoch; /* current traversal's stamp; 0 == "no traversal yet" */
|
||||||
|
size_t cap;
|
||||||
|
} VVisit;
|
||||||
|
|
||||||
|
/* calloc leaves every stamp 0 and epoch 0; the first visit_reset moves to
|
||||||
|
* epoch 1, so no element reads as visited before it is marked. */
|
||||||
|
static int visit_init(VVisit* v, size_t cap){
|
||||||
|
size_t nc = cap ? cap : 16;
|
||||||
|
v->mark = (uint32_t*)calloc(nc, sizeof(uint32_t));
|
||||||
|
if (!v->mark) return -1;
|
||||||
|
v->cap = nc; v->epoch = 0;
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
static inline void visited_reset(VIndex* ix){
|
static void visit_dispose(VVisit* v){ free(v->mark); v->mark = NULL; v->cap = 0; }
|
||||||
if (++ix->visit_epoch == 0){ /* wrapped: clear all */
|
static inline void visit_reset(VVisit* v){
|
||||||
memset(ix->visited, 0, ix->visited_cap*sizeof(uint32_t));
|
if (++v->epoch == 0){ /* wrapped: clear all */
|
||||||
ix->visit_epoch = 1;
|
memset(v->mark, 0, v->cap*sizeof(uint32_t));
|
||||||
|
v->epoch = 1;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
static inline int is_visited(VIndex* ix, int e){ return ix->visited[e]==ix->visit_epoch; }
|
static inline int is_visited(const VVisit* v, int e){ return v->mark[e]==v->epoch; }
|
||||||
static inline void mark_visited(VIndex* ix, int e){ ix->visited[e]=ix->visit_epoch; }
|
static inline void mark_visited(VVisit* v, int e){ v->mark[e]=v->epoch; }
|
||||||
|
|
||||||
/* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */
|
/* ── search one layer (Algorithm 2): best-first, ef-bounded ───────────────── */
|
||||||
/* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller
|
/* Returns results as an unsorted Heap (max-heap on distance, size<=ef). Caller
|
||||||
* owns res->a. `q` is a normalised query. */
|
* owns res->a. `q` is a normalised query. */
|
||||||
static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
|
static int search_layer(const VIndex* ix, VVisit* vis, const float* q,
|
||||||
|
const int* eps, int neps,
|
||||||
int ef, int layer, Heap* res /*out, max-heap*/){
|
int ef, int layer, Heap* res /*out, max-heap*/){
|
||||||
Heap cand = {0,0,0}; /* min-heap: nearest to expand */
|
Heap cand = {0,0,0}; /* min-heap: nearest to expand */
|
||||||
res->a=NULL; res->n=0; res->cap=0;
|
res->a=NULL; res->n=0; res->cap=0;
|
||||||
visited_reset(ix);
|
visit_reset(vis);
|
||||||
for (int i=0;i<neps;i++){
|
for (int i=0;i<neps;i++){
|
||||||
int e = eps[i];
|
int e = eps[i];
|
||||||
if (is_visited(ix,e)) continue;
|
if (is_visited(vis,e)) continue;
|
||||||
mark_visited(ix,e);
|
mark_visited(vis,e);
|
||||||
float d = vdist(ix, q, ix->elems[e].vec);
|
float d = vdist(ix, q, ix->elems[e].vec);
|
||||||
Pair p = { d, e };
|
Pair p = { d, e };
|
||||||
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
|
if (heap_push(&cand,p,0) || heap_push(res,p,1)){ free(cand.a); return -1; }
|
||||||
@@ -212,8 +233,8 @@ static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
|
|||||||
NeighList* nl = &ce->links[layer];
|
NeighList* nl = &ce->links[layer];
|
||||||
for (int i=0;i<nl->count;i++){
|
for (int i=0;i<nl->count;i++){
|
||||||
int e = nl->ids[i];
|
int e = nl->ids[i];
|
||||||
if (is_visited(ix,e)) continue;
|
if (is_visited(vis,e)) continue;
|
||||||
mark_visited(ix,e);
|
mark_visited(vis,e);
|
||||||
float d = vdist(ix, q, ix->elems[e].vec);
|
float d = vdist(ix, q, ix->elems[e].vec);
|
||||||
if (res->n < ef || d < res->a[0].d){
|
if (res->n < ef || d < res->a[0].d){
|
||||||
Pair p = { d, e };
|
Pair p = { d, e };
|
||||||
@@ -232,7 +253,7 @@ static int search_layer(VIndex* ix, const float* q, const int* eps, int neps,
|
|||||||
* Keep c only if it is nearer to q than to every already-chosen neighbour;
|
* Keep c only if it is nearer to q than to every already-chosen neighbour;
|
||||||
* backfill from the pruned set (nearest first) to reach M for connectivity.
|
* backfill from the pruned set (nearest first) to reach M for connectivity.
|
||||||
* Writes chosen element indices into out[], returns the count. */
|
* Writes chosen element indices into out[], returns the count. */
|
||||||
static int select_neighbors(VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
|
static int select_neighbors(const VIndex* ix, const float* q, Pair* W, int nW, int M, int* out){
|
||||||
(void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */
|
(void)q; /* q's distances are precomputed in W[].d; kept for call-site clarity */
|
||||||
/* sort W ascending by (dist,elem) — deterministic. */
|
/* sort W ascending by (dist,elem) — deterministic. */
|
||||||
for (int i=1;i<nW;i++){ /* insertion sort (nW small) */
|
for (int i=1;i<nW;i++){ /* insertion sort (nW small) */
|
||||||
@@ -281,7 +302,7 @@ static int elems_reserve(VIndex* ix){
|
|||||||
Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem));
|
Elem* ne = (Elem*)realloc(ix->elems, nc*sizeof(Elem));
|
||||||
if (!ne) return -1;
|
if (!ne) return -1;
|
||||||
ix->elems = ne; ix->cap = nc;
|
ix->elems = ne; ix->cap = nc;
|
||||||
return visited_ensure(ix);
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
|
int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
|
||||||
@@ -307,13 +328,19 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
|
|||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* This call frame owns its traversal scratch for the whole insert. ix->cap
|
||||||
|
* already covers `cur` (elems_reserve ran above), so every reachable element
|
||||||
|
* index is in range. */
|
||||||
|
VVisit vis;
|
||||||
|
if (visit_init(&vis, ix->cap)) return -1;
|
||||||
|
|
||||||
int ep = ix->entry;
|
int ep = ix->entry;
|
||||||
int L = ix->max_level;
|
int L = ix->max_level;
|
||||||
/* greedy descent through layers above `level` to refine the entry point. */
|
/* greedy descent through layers above `level` to refine the entry point. */
|
||||||
for (int lc = L; lc > level; lc--){
|
for (int lc = L; lc > level; lc--){
|
||||||
Heap r = {0,0,0};
|
Heap r = {0,0,0};
|
||||||
int eps1[1] = { ep };
|
int eps1[1] = { ep };
|
||||||
if (search_layer(ix, el->vec, eps1, 1, 1, lc, &r)){ return -1; }
|
if (search_layer(ix, &vis, el->vec, eps1, 1, 1, lc, &r)){ visit_dispose(&vis); return -1; }
|
||||||
if (r.n){ ep = r.a[0].e; float bd=r.a[0].d;
|
if (r.n){ ep = r.a[0].e; float bd=r.a[0].d;
|
||||||
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} }
|
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; ep=r.a[i].e;} }
|
||||||
free(r.a);
|
free(r.a);
|
||||||
@@ -329,7 +356,7 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
|
|||||||
for (int lc = start; lc >= 0; lc--){
|
for (int lc = start; lc >= 0; lc--){
|
||||||
int Mmax = (lc==0) ? ix->M0 : ix->M;
|
int Mmax = (lc==0) ? ix->M0 : ix->M;
|
||||||
Heap W = {0,0,0};
|
Heap W = {0,0,0};
|
||||||
if (search_layer(ix, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
|
if (search_layer(ix, &vis, el->vec, eps, neps, ix->ef_construction, lc, &W)){ rc=-1; break; }
|
||||||
int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
|
int* chosen = (int*)malloc((size_t)(W.n?W.n:1)*sizeof(int));
|
||||||
if (!chosen){ free(W.a); rc=-1; break; }
|
if (!chosen){ free(W.a); rc=-1; break; }
|
||||||
int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen);
|
int nc = select_neighbors(ix, el->vec, W.a, W.n, Mmax, chosen);
|
||||||
@@ -357,13 +384,17 @@ int vindex_insert(VIndex* ix, uint64_t node_id, const float* vec){
|
|||||||
}
|
}
|
||||||
done:
|
done:
|
||||||
free(eps_owned);
|
free(eps_owned);
|
||||||
|
visit_dispose(&vis);
|
||||||
if (rc) return -1;
|
if (rc) return -1;
|
||||||
if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; }
|
if (level > ix->max_level){ ix->max_level = level; ix->entry = cur; }
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* ── search ───────────────────────────────────────────────────────────────── */
|
/* ── search ───────────────────────────────────────────────────────────────── */
|
||||||
int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
|
/* `ix` is const: search is pure with respect to the index. That is enforced by
|
||||||
|
* the compiler, not by convention — it is the whole point of moving the visited
|
||||||
|
* set into the frame below. */
|
||||||
|
int vindex_search(const VIndex* ix, const float* query, int k, int ef_search,
|
||||||
uint64_t* node_id_out, float* dist_out){
|
uint64_t* node_id_out, float* dist_out){
|
||||||
if (!ix || !query || k <= 0) return -1;
|
if (!ix || !query || k <= 0) return -1;
|
||||||
if (ix->entry < 0) return 0;
|
if (ix->entry < 0) return 0;
|
||||||
@@ -373,11 +404,15 @@ int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
|
|||||||
float* q = vec_normalise_copy(query, ix->dim);
|
float* q = vec_normalise_copy(query, ix->dim);
|
||||||
if (!q) return -1;
|
if (!q) return -1;
|
||||||
|
|
||||||
|
/* This call frame owns its traversal scratch. */
|
||||||
|
VVisit vis;
|
||||||
|
if (visit_init(&vis, ix->cap)){ free(q); return -1; }
|
||||||
|
|
||||||
int ep = ix->entry;
|
int ep = ix->entry;
|
||||||
for (int lc = ix->max_level; lc > 0; lc--){
|
for (int lc = ix->max_level; lc > 0; lc--){
|
||||||
Heap r = {0,0,0};
|
Heap r = {0,0,0};
|
||||||
int eps[1] = { ep };
|
int eps[1] = { ep };
|
||||||
if (search_layer(ix, q, eps, 1, 1, lc, &r)){ free(q); return -1; }
|
if (search_layer(ix, &vis, q, eps, 1, 1, lc, &r)){ visit_dispose(&vis); free(q); return -1; }
|
||||||
if (r.n){ int b=r.a[0].e; float bd=r.a[0].d;
|
if (r.n){ int b=r.a[0].e; float bd=r.a[0].d;
|
||||||
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;}
|
for (int i=1;i<r.n;i++) if (r.a[i].d<bd){bd=r.a[i].d; b=r.a[i].e;}
|
||||||
ep = b; }
|
ep = b; }
|
||||||
@@ -385,7 +420,8 @@ int vindex_search(VIndex* ix, const float* query, int k, int ef_search,
|
|||||||
}
|
}
|
||||||
Heap res = {0,0,0};
|
Heap res = {0,0,0};
|
||||||
int eps[1] = { ep };
|
int eps[1] = { ep };
|
||||||
if (search_layer(ix, q, eps, 1, ef_search, 0, &res)){ free(res.a); free(q); return -1; }
|
if (search_layer(ix, &vis, q, eps, 1, ef_search, 0, &res)){ visit_dispose(&vis); free(res.a); free(q); return -1; }
|
||||||
|
visit_dispose(&vis);
|
||||||
free(q);
|
free(q);
|
||||||
|
|
||||||
/* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */
|
/* res is a max-heap of size<=ef; pop into ascending order, keep nearest k. */
|
||||||
@@ -419,7 +455,6 @@ VIndex* vindex_create(int dim, int M, int ef_construction){
|
|||||||
ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0);
|
ix->mL = 1.0 / log((double)M > 1.0 ? (double)M : 2.0);
|
||||||
ix->entry = -1;
|
ix->entry = -1;
|
||||||
ix->max_level = 0;
|
ix->max_level = 0;
|
||||||
ix->visit_epoch = 0;
|
|
||||||
return ix;
|
return ix;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -432,7 +467,6 @@ void vindex_free(VIndex* ix){
|
|||||||
free(e->vec);
|
free(e->vec);
|
||||||
}
|
}
|
||||||
free(ix->elems);
|
free(ix->elems);
|
||||||
free(ix->visited);
|
|
||||||
free(ix);
|
free(ix);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -53,8 +53,15 @@ int vindex_insert(VIndex* idx, uint64_t node_id, const float* vec);
|
|||||||
* first (ascending distance). Either out array may be NULL to skip it.
|
* first (ascending distance). Either out array may be NULL to skip it.
|
||||||
* ef_search — search-time candidate width; larger == higher recall, slower.
|
* ef_search — search-time candidate width; larger == higher recall, slower.
|
||||||
* Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k.
|
* Pass <=0 for VINDEX_DEFAULT_EF_SEARCH. Internally clamped to >=k.
|
||||||
* Returns the number of results written, or <0 on error. */
|
* Returns the number of results written, or <0 on error.
|
||||||
int vindex_search(VIndex* idx, const float* query, int k, int ef_search,
|
*
|
||||||
|
* `idx` is const BY CONTRACT AND BY TYPE: search does not mutate the index. The
|
||||||
|
* traversal's visited set is owned by the call frame, so N threads may search one
|
||||||
|
* index concurrently. Concurrent search against a vindex_insert on the same index
|
||||||
|
* is still unsafe — insert rewires existing elements' neighbour lists and reallocs
|
||||||
|
* elems[] — so the index's owner must not extend a published index under a live
|
||||||
|
* reader. See eg_vindex_view / eg_vindex_maintain in el_runtime.c. */
|
||||||
|
int vindex_search(const VIndex* idx, const float* query, int k, int ef_search,
|
||||||
uint64_t* node_id_out, float* dist_out);
|
uint64_t* node_id_out, float* dist_out);
|
||||||
|
|
||||||
/* Number of vectors currently indexed. */
|
/* Number of vectors currently indexed. */
|
||||||
|
|||||||
+169
-5
@@ -29,6 +29,8 @@ This section is the **single source of truth** for what works and what is planne
|
|||||||
- Lexer: keywords, identifiers, integer/float/string/bool literals, operators below.
|
- Lexer: keywords, identifiers, integer/float/string/bool literals, operators below.
|
||||||
- Parser: `let`, `return`, `fn`, `type`, `enum`, `import`, `from … import`, `while`, `for`, `if/else if/else`, `match`, `@decorator`, array/map literals, all listed operators, function calls, field access, index access, unary `!`/`-`, postfix `?`.
|
- Parser: `let`, `return`, `fn`, `type`, `enum`, `import`, `from … import`, `while`, `for`, `if/else if/else`, `match`, `@decorator`, array/map literals, all listed operators, function calls, field access, index access, unary `!`/`-`, postfix `?`.
|
||||||
- Codegen: function definitions, top-level `main()`, all expression forms above, control flow, decorator-as-AST-attachment.
|
- Codegen: function definitions, top-level `main()`, all expression forms above, control flow, decorator-as-AST-attachment.
|
||||||
|
- Boundary seam: decorator arguments and stacking; VBD role enforcement via `#error`; `engram_boundary_beat` auto-emit at `@manager`/`@accessor` entry; `@route` dispatch tables (Section 9).
|
||||||
|
- Program-level declarative blocks: `cgi`, `service`, and `program` — the last carrying process identity and configuration (Section 18).
|
||||||
- C runtime: I/O, string operations, integer math, lists, maps, filesystem, command-line args, basic `json_get` substring lookup.
|
- C runtime: I/O, string operations, integer math, lists, maps, filesystem, command-line args, basic `json_get` substring lookup.
|
||||||
|
|
||||||
### Planned (in flight)
|
### Planned (in flight)
|
||||||
@@ -37,7 +39,7 @@ This section is the **single source of truth** for what works and what is planne
|
|||||||
- **Match codegen.** Currently parsed; codegen does not emit. Adding `({ ... })` statement-expression emission.
|
- **Match codegen.** Currently parsed; codegen does not emit. Adding `({ ... })` statement-expression emission.
|
||||||
- **`?` propagation.** Currently no-op. Adding nil-propagation semantics.
|
- **`?` propagation.** Currently no-op. Adding nil-propagation semantics.
|
||||||
- **`cgi` block parsing.** Currently lexed (`cgi` is a keyword) but not parsed as a statement. Adding `parse_cgi_block` and codegen of `el_cgi_init` at the head of `main()`.
|
- **`cgi` block parsing.** Currently lexed (`cgi` is a keyword) but not parsed as a statement. Adding `parse_cgi_block` and codegen of `el_cgi_init` at the head of `main()`.
|
||||||
- **VBD role enforcement.** `@manager`/`@engine`/`@accessor` are accepted as decorators but not enforced. Adding compile-time check that `dharma_emit`/`dharma_field` only appear inside `@manager` functions.
|
- **Boundary epilogues.** The decorator seam injects a prologue only. Adding prologue/epilogue wrapping, the prerequisite for durability-as-an-effect (Section 19.1).
|
||||||
- **`vessel` keyword.** Replaces `package` in manifests. Adding to lexer.
|
- **`vessel` keyword.** Replaces `package` in manifests. Adding to lexer.
|
||||||
- **Real `engram_*` runtime.** Currently stub. Adding in-process graph store with spreading activation, Hebbian strengthening, and disk persistence — see Section 16.4.
|
- **Real `engram_*` runtime.** Currently stub. Adding in-process graph store with spreading activation, Hebbian strengthening, and disk persistence — see Section 16.4.
|
||||||
- **Real `dharma_*` runtime.** Currently stub. Adding network transport, channel registry, identity resolution.
|
- **Real `dharma_*` runtime.** Currently stub. Adding network transport, channel registry, identity resolution.
|
||||||
@@ -96,8 +98,10 @@ The following words are reserved and cannot be used as identifiers. Each row not
|
|||||||
| `while` | yes | Loop |
|
| `while` | yes | Loop |
|
||||||
| `import` / `from` / `as` | yes | Module import |
|
| `import` / `from` / `as` | yes | Module import |
|
||||||
| `true` / `false` | yes | Bool literals |
|
| `true` / `false` | yes | Bool literals |
|
||||||
| `cgi` | planned | Top-level CGI declaration block |
|
| `cgi` | yes | Top-level CGI declaration block |
|
||||||
| `manager` / `engine` / `accessor` | as decorators | VBD role marker on `fn` (enforcement planned) |
|
| `service` | yes | Top-level capability-bounded declaration block |
|
||||||
|
| `program` | yes | Top-level cross-cutting declaration block (Section 18) |
|
||||||
|
| `manager` / `engine` / `accessor` | as decorators | VBD role marker on `fn`; enforcement and boundary auto-emit are live (Section 9) |
|
||||||
| `vessel` | planned | Manifest declaration (replaces `package`) |
|
| `vessel` | planned | Manifest declaration (replaces `package`) |
|
||||||
| `activate` / `where` | planned | Spreading-activation construct |
|
| `activate` / `where` | planned | Spreading-activation construct |
|
||||||
| `sealed` | planned | Capability scope block |
|
| `sealed` | planned | Capability scope block |
|
||||||
@@ -446,9 +450,21 @@ Parsed. The module name is recorded; the brace-list is consumed. Both forms prod
|
|||||||
fn handle(channel: String, msg: String) -> Void { … }
|
fn handle(channel: String, msg: String) -> Void { … }
|
||||||
```
|
```
|
||||||
|
|
||||||
The `@` token followed by an identifier attaches a decorator name to the next `FnDef`. Decorators with structural meaning today: none. Planned enforcement (Section 16.2): VBD roles `@manager`, `@engine`, `@accessor`.
|
The `@` token followed by an identifier attaches a decorator to the next `FnDef`.
|
||||||
|
|
||||||
Non-VBD decorators are accepted and ignored.
|
**Decorators take arguments and they stack.** `@route("/p", "GET") @manager fn f()` attaches both to `f` as a `decorators` list of `{name, args}` records, topmost-first. Arguments are string literals only.
|
||||||
|
|
||||||
|
**Decorators have structural meaning today.** This is El's function-level boundary seam — the mechanism by which a cross-cutting concern is handled *at the boundary* rather than by a convention repeated at every call site:
|
||||||
|
|
||||||
|
| Decorator | Structural effect |
|
||||||
|
|---|---|
|
||||||
|
| `@manager` | Permits calls to `dharma_emit` / `dharma_field`. Calling either from a non-`@manager` fn emits a `#error` into the generated C — a compile-time failure, not a lint. |
|
||||||
|
| `@manager`, `@accessor` | Codegen injects one call to `engram_boundary_beat(<fn name>)` at function entry. The decorated op self-reports (chrono tick, afferent counter, self-activity strengthen, dharma bus event) with **zero** hand-written instrumentation in its body. |
|
||||||
|
| `@route(path, method, …)` | Records a route into a generated dispatch table. |
|
||||||
|
|
||||||
|
Decorators with no registered meaning are accepted and ignored.
|
||||||
|
|
||||||
|
**Limits of the seam, as it stands.** The injection is a *prologue only* — there is no epilogue, no wrapping of the call, and no way for a decorator to run code after the body returns. The injected callee is a fixed builtin chosen by the compiler, not derived from the decorator name or its arguments. Section 19 depends on lifting exactly these two limits.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -1088,4 +1104,152 @@ The next minor version closes the implementation gaps named in this document. Tr
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## 18. The Program Block — cross-cutting concerns [implemented]
|
||||||
|
|
||||||
|
### 18.0 Why this exists
|
||||||
|
|
||||||
|
A cross-cutting concern is one that belongs to the *process*, not to any function in it: only one of me may run; this is what my configuration is; every mutation must be durable; every request must be authorized.
|
||||||
|
|
||||||
|
El's units of encapsulation are the function and the module. Neither can hold a concern like that. So each one had been expressed the only way it could be — as a **convention**: *call this at every site.* Conventions of that shape do not hold. They are not enforced by anything, they are invisible in review, and they fail silently at the one site somebody forgot.
|
||||||
|
|
||||||
|
Measured in this codebase before this section existed:
|
||||||
|
|
||||||
|
| Concern | State | What the convention was |
|
||||||
|
|---|---|---|
|
||||||
|
| process identity | **zero** guards anywhere — no pidfile, no lock, no already-running check, at any layer | "check nothing is already running first" |
|
||||||
|
| configuration | **20** distinct environment variables in one program, each with its default written inline at the read site | "remember the right default here" |
|
||||||
|
| durability | **62** `persist_*` / `engram_save` / `wal_*` / `checkpoint` call sites | "after you mutate, remember to persist" |
|
||||||
|
| request auth | **10** per-route `_auth` checks | "check the token in this handler too" |
|
||||||
|
|
||||||
|
These are not four problems. They are one absence, four times.
|
||||||
|
|
||||||
|
That the convention form fails is observed, not predicted. Process identity failed three times in a single day: twice, two engram processes ran simultaneously against the same data directory; twice, a stale binary held a port and answered probes while a fresh build was believed to be under test, because `pkill -f` had silently failed to match its argv — which nearly produced a false "the fix does not work" conclusion. Configuration failed structurally: `ENGRAM_DATA_DIR` was read at six sites, five of them dead bindings, and the sixth defaulted to `/tmp/engram` — contradicting the canonical resolver's `$HOME/.neuron/engram` and landing a pre-destructive safety backup on ephemeral storage.
|
||||||
|
|
||||||
|
The `program` block is where a concern of this shape is declared once and enforced by the compiler at the process boundary.
|
||||||
|
|
||||||
|
### 18.1 Syntax
|
||||||
|
|
||||||
|
```
|
||||||
|
program "engram" {
|
||||||
|
singleton: "engram"
|
||||||
|
env ENGRAM_BIND: String = ":8742"
|
||||||
|
env GUIDE_PORT: Int = "8771"
|
||||||
|
env ENGRAM_API_KEY: String required
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
At most one `program` block per program. It composes with `cgi` and `service` — those declare what a program *may do*; `program` declares what a program *is*.
|
||||||
|
|
||||||
|
Grammar:
|
||||||
|
|
||||||
|
```ebnf
|
||||||
|
program_block = "program" string "{" { program_field } "}" ;
|
||||||
|
program_field = singleton_field | env_field ;
|
||||||
|
singleton_field = "singleton" ":" string [ "," ] ;
|
||||||
|
env_field = "env" ident ":" type
|
||||||
|
[ "=" string ] [ "required" ] [ "," ] ;
|
||||||
|
```
|
||||||
|
|
||||||
|
`singleton` and `env` are **not** reserved words. They are read as identifier token values by the block's own parse loop, so they remain usable as ordinary identifiers everywhere else. `program` is the only keyword this section adds.
|
||||||
|
|
||||||
|
### 18.2 Process identity — `singleton`
|
||||||
|
|
||||||
|
`singleton: "id"` compiles to an `el_singleton_acquire("id")` call injected as the **first statement of `main()`**, before any user statement runs.
|
||||||
|
|
||||||
|
The runtime takes an exclusive non-blocking `flock` on `<dir>/el-singleton-<id>.lock`, where `<dir>` is `$EL_SINGLETON_DIR`, else `$TMPDIR`, else `/tmp`. On success it writes its pid and holds the descriptor open for the life of the process. On contention it **refuses to start**: it reports the holder's pid, names the lock file, and exits 1.
|
||||||
|
|
||||||
|
Two properties are deliberate:
|
||||||
|
|
||||||
|
- **It is a lock, not a pidfile.** The kernel releases an `flock` when the owning process dies — including on `SIGKILL` and on crash. There is therefore no stale-lock state, and so no "delete the lock file to get unstuck" recovery ritual. Such a ritual would itself be a convention, which is the thing this section exists to remove.
|
||||||
|
- **It reports the holder's pid.** "Already running" is not actionable. A pid is. This is the direct answer to the observed failure where a stale process survived a `pkill` and went on answering probes.
|
||||||
|
|
||||||
|
Refusal is loud and total. It is not a warning, and the program does not continue degraded. This matters more than it looks: today a second engram whose `bind()` fails merely *returns* from `http_serve` — after it has already replayed the WAL and written boot-time backup files — and then exits **0**, indistinguishable from a clean run. `singleton` refuses before the first side effect.
|
||||||
|
|
||||||
|
### 18.3 Configuration — `env`
|
||||||
|
|
||||||
|
Each `env` entry declares one configuration variable: its name, its type (`Int` or `String`), and either a default or `required`.
|
||||||
|
|
||||||
|
Resolution happens once, at startup, in declaration order: **the environment wins; the declaration supplies the fallback.** Then `el_config_validate` checks the whole schema and reports *every* problem at once before exiting — a startup that fails one variable at a time costs one restart per variable.
|
||||||
|
|
||||||
|
Values are read with `config("NAME")`, which returns a `String`.
|
||||||
|
|
||||||
|
The enforcement that makes the declaration real: **once a program block exists, `config("X")` for an undeclared `X` is a fatal error.** Without that, the schema would be advisory, and an advisory schema is just another convention. Programs with no `program` block are unaffected — `config()` falls back to a plain environment read, so migration is incremental and per-program.
|
||||||
|
|
||||||
|
The point is not that configuration is now centralized. It is that **a default is no longer a decision made at a read site.** A read site cannot disagree with another read site about what a variable means, because a read site no longer says.
|
||||||
|
|
||||||
|
### 18.4 What is deliberately not declared here
|
||||||
|
|
||||||
|
Some values look like configuration and are not. `ENGRAM_DATA_DIR` already has a single owner — `engram_resolve_data_dir()`, which resolves it, creates the directory, and fails loud rather than silently persisting to an ephemeral path. Declaring it in the `program` block as well would give it two owners that can disagree, recreating the precise defect this section removes.
|
||||||
|
|
||||||
|
The rule: **a variable belongs in the program block when the block would be its only owner.** If a resolver already owns it, leave it there.
|
||||||
|
|
||||||
|
`HOME` is likewise not configuration. It is an environment fact, and stays a raw `env()` read.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 19. Boundary Effects — durability and request authorization [design only, not implemented]
|
||||||
|
|
||||||
|
Sections 19.1 and 19.2 specify the two remaining concerns from the table in 18.0. Both are **designed and deliberately unimplemented.** The reason is stated in 19.3 and it is not difficulty.
|
||||||
|
|
||||||
|
### 19.1 Durability as an epilogue effect
|
||||||
|
|
||||||
|
**The defect.** 62 call sites carry the convention *"after you mutate, remember to persist."* This is structurally the same defect as the index bug being fixed elsewhere in this tree — *"after you append, remember to index"* — which failed at **9 of 9** sites. A convention that failed at 100% of its sites is the strongest available evidence about what this class of convention is worth.
|
||||||
|
|
||||||
|
**Why the existing seam cannot express it.** §9's injection is a prologue. Durability is inherently an *epilogue*: persist after the mutation succeeds, and not at all if it threw. The seam has no epilogue.
|
||||||
|
|
||||||
|
**Design.** Extend the decorator seam from prologue-only to prologue/epilogue, then declare durability as an effect on the mutating function:
|
||||||
|
|
||||||
|
```
|
||||||
|
@durable("engram")
|
||||||
|
fn engram_write_node(id: String, body: String) -> Bool { … }
|
||||||
|
```
|
||||||
|
|
||||||
|
Codegen wraps rather than prefixes:
|
||||||
|
|
||||||
|
```c
|
||||||
|
el_val_t engram_write_node(el_val_t id, el_val_t body) {
|
||||||
|
el_effect_enter(EL_STR("durable"), EL_STR("engram"));
|
||||||
|
el_val_t __r = /* original body */;
|
||||||
|
el_effect_exit(EL_STR("durable"), EL_STR("engram"), __r);
|
||||||
|
return __r;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`el_effect_exit` is where the persist happens, and it is the only place it happens. Two properties follow that the 62 hand-written sites cannot have:
|
||||||
|
|
||||||
|
- **Coalescing.** The epilogue is a single choke point, so N mutations inside one request can produce one fsync instead of N. The hand-written form cannot coalesce, because no site knows about the others.
|
||||||
|
- **Failure is not silent.** A persist that fails inside `el_effect_exit` can force the mutation's return value to failure. A forgotten `persist_*` call cannot fail — it simply does not happen, which is exactly why the defect is invisible.
|
||||||
|
|
||||||
|
**Enforcement, and this is the part that actually fixes it.** Mirroring §9's `#error` for `dharma_emit`: a function that calls a mutating primitive without carrying `@durable` is a **compile error**. Otherwise this is a 63rd thing to remember rather than a replacement for 62.
|
||||||
|
|
||||||
|
### 19.2 Request authorization as a route effect
|
||||||
|
|
||||||
|
**The defect.** 10 per-route `_auth` checks. The HTTP layer has no concept of authorization, so a new route is unauthenticated by default and silently so — the failure mode is a route that forgot, and nothing anywhere reports it.
|
||||||
|
|
||||||
|
**Design.** Authorization becomes an argument to the `@route` decorator, which already takes arguments and already builds a dispatch table:
|
||||||
|
|
||||||
|
```
|
||||||
|
@route("/api/write", "POST", auth: "required")
|
||||||
|
fn route_write(body: String) -> String { … }
|
||||||
|
```
|
||||||
|
|
||||||
|
The generated dispatcher performs the check **before** dispatch, so an unauthorized request never reaches the handler and the handler contains no auth code at all.
|
||||||
|
|
||||||
|
The default must be `required`. A route that says nothing gets authorization; opening one up takes an explicit `auth: "public"`. Defaulting to public preserves the current failure mode exactly — forgetting stays silent — and a default that preserves the defect is not a fix.
|
||||||
|
|
||||||
|
Route inventory falls out for free: the dispatch table already exists, so the compiler can emit the full route/auth matrix and make "which routes are public" a fact that is read rather than audited.
|
||||||
|
|
||||||
|
### 19.3 Why these are not implemented
|
||||||
|
|
||||||
|
Not difficulty — **collision**. Both land squarely in regions two other agents hold right now:
|
||||||
|
|
||||||
|
- **Durability** requires changing the mutation and persist paths in `lang/runtime/el_runtime.c` and `engram/src/server.el` — the same files and the same read/write paths being restructured by concurrent work on VIndex read-path mutation and memory ownership, and on geometry-as-an-el-value and `transduce`.
|
||||||
|
- **Request auth** requires changing route dispatch in `engram/src/server.el`, which the geometry/`transduce` work is actively reshaping.
|
||||||
|
|
||||||
|
Implementing either now would mean editing files under concurrent modification and resolving conflicts in exactly the paths whose correctness is currently under repair. The designs are recorded here so the work is not lost, and so that whoever lands them does so against a settled tree.
|
||||||
|
|
||||||
|
The prerequisite for 19.1 is the same in both cases: **lift the §9 seam from prologue-only to prologue/epilogue.** That change is independent of both collisions and can land first.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
End of specification.
|
End of specification.
|
||||||
|
|||||||
@@ -0,0 +1,180 @@
|
|||||||
|
# El Runtime — Ownership and Capability ABI
|
||||||
|
|
||||||
|
**Status:** §0–§2 verified. §3 re-derived and **built** for the vector index (2026-08-16); not yet applied to the resident RAM graph.
|
||||||
|
**Date:** 2026-08-16
|
||||||
|
**Scope:** `lang/runtime/` — every El program (soul, engram, cgi-studio vessels) inherits this by rebuild. Nothing in this document is a change to any El *program*.
|
||||||
|
|
||||||
|
**Note on §1's line numbers:** they were read against a checkout that has since shifted by ~135 lines. Verified positions as of `a67452f` are in §2a.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0. The residual
|
||||||
|
|
||||||
|
> **Builtins own memory and reach process state directly.**
|
||||||
|
|
||||||
|
That is the residual — the generator. Everything below labelled a "residue" is a deposit left by it. The distinction matters because we have spent significant effort removing deposits, and deposits regenerate.
|
||||||
|
|
||||||
|
A residue is fixed. A residual is eliminated. Fixing residues while the residual stands produces exactly the pattern observed on 2026-08-15/16: a run of individually-correct patches, each verified, followed by a new defect of the same shape in a different file.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. The residues, measured
|
||||||
|
|
||||||
|
Each of these is a distinct merged or proposed fix. Each addresses one deposit. None addresses the residual.
|
||||||
|
|
||||||
|
| residue | location | fix that was applied or proposed |
|
||||||
|
|---|---|---|
|
||||||
|
| `state_get` leaked its return value per call — 15 MB over 200k calls | builtin | el #140 (merged) |
|
||||||
|
| VIndex freed under a concurrent reader | `el_runtime.c:9424` | `fb32d15` guard (merged 08:46:43) |
|
||||||
|
| `_eg_vindex_seen` realloc'd on a read path | `el_runtime.c:9412` | same guard |
|
||||||
|
| `vindex_insert` on a read path | `el_runtime.c:9434`, `9450` | same guard |
|
||||||
|
| shared `visited` / epoch scratch stomped by concurrent searches | `engram_vindex.c:79–81`, `169–186`, `195` | proposed: move to per-search frame |
|
||||||
|
| nine append sites, none indexing → lazily-embedded nodes invisible | `el_runtime.c:7806, 7988, 8148, 8224, 11526, 11731, 12050, 15295, 15312` | "embed-gap #20", patched by making the *read* path catch up (`9439` comment) |
|
||||||
|
|
||||||
|
**Measured:** all file/line references above, read 2026-08-16. Crash frames `engram_activate → eg_vindex_sync → vindex_insert → _realloc → _xzm_xzone_malloc_freelist_outlined` are accounted for by rows 2–4.
|
||||||
|
|
||||||
|
**Inferred, not yet verified:** that the nine append sites do not share a single commit point. This needs one pass before Change C is sized.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Why these are one defect
|
||||||
|
|
||||||
|
`eg_vindex_sync` (`el_runtime.c:9419`) has exactly three callers, and **all three are reads**:
|
||||||
|
|
||||||
|
- `engram_activate` — `9802`
|
||||||
|
- `eg_knn_for_node` — `13075` (its own header comment states *"No writes."*)
|
||||||
|
- `engram_geo_reify_run_json` — `13285`
|
||||||
|
|
||||||
|
It mutates five process-global statics (`9400–9404`): `_eg_vindex`, `_eg_vindex_dim`, `_eg_vindex_built_nc`, `_eg_vindex_seen`, `_eg_vindex_seen_cap`.
|
||||||
|
|
||||||
|
Reads mutate because index maintenance was never given an owner on the write side. It got bolted onto reads, because a builtin *could* reach the globals — nothing prevented it. Likewise `state_get` leaked because a builtin *owned* the value it returned; nothing prevented that either.
|
||||||
|
|
||||||
|
The store is architecturally append-only and superseding. A read path that mutates contradicts that directly. The contradiction is expressible only because the ABI permits it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2a. Verified positions and the fact §1 missed
|
||||||
|
|
||||||
|
Read directly at `a67452f`, 2026-08-16. §1's line numbers predate a ~135-line shift; these are current.
|
||||||
|
|
||||||
|
| thing | §1 said | actually |
|
||||||
|
|---|---|---|
|
||||||
|
| five process-global statics | 9400–9404 | **9535–9539** |
|
||||||
|
| `eg_vindex_seen_ensure` realloc | 9412 | **9547** |
|
||||||
|
| `eg_vindex_sync` | 9419 | **9554** |
|
||||||
|
| `vindex_free` on a read path | 9424 | **9559** |
|
||||||
|
| `vindex_insert` on a read path | 9434 / 9450 | **9569** (build) / **9585** (incremental) |
|
||||||
|
| caller: `engram_activate_inner` | 9802 | **9939** |
|
||||||
|
| caller: `eg_knn_for_node` | 13075 | **13212** |
|
||||||
|
| caller: `engram_geo_reify_run_json` | 13285 | **13422** |
|
||||||
|
| `fb32d15` guard | — | lock **1602**, depth **1631**, `eg_guard_enter` **1636**, `http_worker` acquire **1687**, `engram_activate` wrapper **14097** |
|
||||||
|
| VIndex scratch fields | 79–81 | **79–81** ✓ |
|
||||||
|
| `search_layer` race site | 195 | **195** ✓ |
|
||||||
|
|
||||||
|
**The structural fact §1 and §3 both missed:** *the index does not inherit the store's append-only property.* `vindex_insert` rewires the `NeighList` links of already-existing elements and reallocs `elems[]` — so extending the index mutates the whole structure, not just its tail. This is why "make reads pure" is necessary but **not sufficient**, and why §3 needed a publication boundary rather than only a capability split. It is reproduced as a standing test (`unsynchronized` half, §5).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. The change
|
||||||
|
|
||||||
|
*(Re-derived 2026-08-16. The previous §3 — a runtime context struct carrying read/write **capability pointers** to every builtin — was written in mutable-store, C-ownership terms. It asked "who is permitted to mutate the shared thing?", which presupposes a shared mutable thing. The engram is immutable and recall is projection; what does not mutate needs no ownership discipline. So the question is not answered, it is dissolved. The implemented change is below.)*
|
||||||
|
|
||||||
|
### 3.1 Three moves, in decreasing order of how much they dissolve
|
||||||
|
|
||||||
|
**(1) Misfiled scratch is not shared state.** `visited` / `visit_epoch` were never conceptually owned by the index — they are one traversal's local, hoisted into `struct VIndex` as an allocation optimisation. Nothing about them is derived geometry. They want neither a lock nor a capability nor a checkout pool: a pure function's scratch belongs to its call frame, and the fix is to put it back there. This is not "the capability model applied by hand to one global"; it is the deletion of a false ownership claim.
|
||||||
|
|
||||||
|
**(2) `const` is the capability, and immutability hands it over for free.** Once the scratch leaves the struct, `search_layer` reads the index and nothing else — so `vindex_search` can take a `const VIndex*`. That is *precisely* the teeth old-§3 wanted from capability pointers: a read path physically cannot call `vindex_insert`, and it is a **compile error**, not a review comment. It costs one qualifier rather than a new ABI swept across hundreds of builtins. The compiler enforces it on every future caller for the same reason.
|
||||||
|
|
||||||
|
> The capability type was already in the language. It is spelled `const`.
|
||||||
|
|
||||||
|
**(3) What remains is a publication problem, not an ownership problem.** With scratch in the frame and reads const, one hazard survives, and it is real: **HNSW insert is not an append.** `vindex_insert` rewires the `NeighList` links of *already-existing* elements and reallocs `elems[]`. The store's append-only property does **not** transfer to the index derived from it. So a reader projecting against the index while its owner extends it is unsafe no matter how pure search is.
|
||||||
|
|
||||||
|
Immutability answers this too, and the answer is publication:
|
||||||
|
|
||||||
|
- **`eg_vindex_maintain`** — the sole mutator. Takes the boundary exclusively; never runs beside a reader.
|
||||||
|
- **`eg_vindex_view`** — returns a `const VIndex*` with the boundary held for read. N readers project concurrently; none can mutate.
|
||||||
|
|
||||||
|
A read path may **demand that a current snapshot exist** — that is a request to the owner, not a mutation by the reader. What it may not do is mutate the geometry it is projecting against. `view` / `maintain` is exactly that split, and it is why this replaces `eg_vindex_sync` rather than wrapping it.
|
||||||
|
|
||||||
|
**Write-side owner.** Index membership is owned by the event *"an embedding became present on this ordinal"* — not by node append, since a node without an embedding cannot be in a vector index at all. `eg_vindex_note_embedded` hooks the embedding-assignment sites: one O(log n) insert, no O(node_count) presence scan. This also retires the "STALENESS (honest tradeoff)" note in the old `eg_vindex_sync`, where a lazily-embedded *older* node stayed invisible to `route_nearest` / autoconnect until the next full rebuild.
|
||||||
|
|
||||||
|
### 3.2 What this does not claim
|
||||||
|
|
||||||
|
The **resident RAM graph** (`g->nodes` / `g->edges`) is a *separate* residue of the same residual and is untouched by this change. It is realloc'd in place (`el_runtime.c:7618`, `7629`), so an awareness-thread reader holding `EngramNode* n = &g->nodes[i]` across a concurrent append holds a dangling pointer — and `engram_activate_inner`'s embed-backfill writes `n->emb` through exactly such a pointer. It wants the same publication treatment the index just received. Until that lands, the `fb32d15` guard stays (see §5).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Why this is not a large change
|
||||||
|
|
||||||
|
The old §4 argued that El owning its compiler makes a capability-ABI sweep mechanical, since `elc` generates every builtin call site. That argument was load-bearing only for the ABI, and the ABI is gone.
|
||||||
|
|
||||||
|
The constraint now travels with the **type of the thing**, not the shape of every call site — so no sweep is needed at all. Measured extent of the implemented change: two qualifiers (`const VIndex*` on `vindex_search`, propagated to `engram_geometry_descriptor` and `engram_geo_reify_store`), one struct field group relocated to a call frame, one rwlock, and three read call sites converted from `eg_vindex_sync` to `view`/`release`.
|
||||||
|
|
||||||
|
The payoff of owning the language is unchanged and is now *cheaper*: introduced once, enforced by the compiler on every future builtin, cannot subsequently be forgotten. Contrast the current state, where the same discipline was maintained by hand across hundreds of builtins and demonstrably failed at least six times.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. What this deletes
|
||||||
|
|
||||||
|
**Deleted (done, 2026-08-16):**
|
||||||
|
|
||||||
|
- `eg_vindex_sync` — the function itself. Not renamed: split into `eg_vindex_maintain` (mutating, exclusive, sole owner) and `eg_vindex_view` (const, shared). A name that meant "read paths repair the index" had to stop existing.
|
||||||
|
- `VIndex::visited` / `visit_epoch` / `visited_cap` — the struct fields, `visited_ensure`, its call from `elems_reserve`, `ix->visit_epoch = 0` in `vindex_create`, and `free(ix->visited)` in `vindex_free`.
|
||||||
|
- The **proposed** per-search scratch *struct on the index* (a checkout pool / `VisitedListPool`) — never built. The buffer is a plain frame local; a pool is machinery for an ownership question that no longer exists.
|
||||||
|
- The **proposed** reader-view / owner-handle split for VIndex specifically — superseded. `const` already is the reader view.
|
||||||
|
- `EXPECT_RACE` in `run_vindex_concurrency_tests.sh` — a knob that let a known defect ride as "expected". Replaced by four halves with real verdicts.
|
||||||
|
|
||||||
|
**NOT deleted — the design doc was wrong about this one:**
|
||||||
|
|
||||||
|
- `fb32d15` (`eg_guard_enter` / `engram_req_lock` / `_eg_req_depth`). §5 originally called for its removal as "a lock protecting a mutation that ceases to exist." **Measured, it guards two things, and only one of them ceases to exist.** Its own comment names both: the RAM graph *and* `_eg_vindex`. The vindex justification is retired; the RAM-graph justification is independently load-bearing (§3.2), and removing the guard reintroduces the measured 11171→9579 edge-loss defect from 2026-08-14. Its comment has been narrowed to state the RAM graph only. **Precondition for deleting it:** the resident graph gets the same publication boundary the index just got.
|
||||||
|
- el #140's hand-patch. Left in place — the leak stops being *expressible* only under the abandoned capability-ABI §3, which is not what was built.
|
||||||
|
|
||||||
|
**Ordering consequence (revised):** the original ordering claim — "the residual lands first, the residues evaporate rather than get fixed" — did not survive contact. The residual here is not a single ABI that dissolves everything at once; it is a *property* (derived state is published, never edited) applied per structure. The index now has it. The RAM graph does not yet. Residues evaporate **per structure, in the order the property is applied**, and a residue whose structure has not been converted must be left standing, not deleted on the strength of the plan.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Sequencing
|
||||||
|
|
||||||
|
1. **Read** how builtins are declared and dispatched, to confirm the call sites are compiler-generated in one place. *(This determines whether §4 holds. If dispatch is scattered, re-size before proceeding.)*
|
||||||
|
2. Introduce the context type and capability types.
|
||||||
|
3. Codegen emits the context at every builtin call site.
|
||||||
|
4. Mechanical sweep of builtin signatures.
|
||||||
|
5. Move index maintenance behind the write capability; the three read callers take the read capability.
|
||||||
|
6. Delete the residue-fixes listed in §5.
|
||||||
|
7. **One** build of soul from el dev — which resolves the `state_get` leak and the crash together, rather than deploying a leak fix that reintroduces the crash.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Open questions
|
||||||
|
|
||||||
|
**Answered 2026-08-16:**
|
||||||
|
|
||||||
|
- ~~Do the nine append sites share a commit point?~~ **Moot.** The question was mis-aimed: node append is not the event that owns index membership, because a node without an embedding cannot be in a vector index. The five *embedding-assignment* sites are the real owner points (`el_runtime.c:7091, 9839, 13362, 15002`, plus snapshot-restore at `7951`), and three of them carry the ordinal directly — which is all `eg_vindex_note_embedded` needs. The other two run before the node is resident, where the cold build picks it up.
|
||||||
|
- ~~Does anything outside `lang/runtime/` construct a second `VIndex`?~~ **No.** Swept: the only constructors outside the runtime are `engram/test/*` and `lang/runtime/vindex_bench.c`, all single-threaded and index-private. Inside the runtime, `engram_self_reify_beat_json` builds a **private** index deliberately and never touches the shared boundary — that was already correct and is unchanged.
|
||||||
|
- ~~Does the HTTP worker pool contend on the same globals?~~ **Yes, and it was never the whole story.** Workers serialize against each other on `engram_req_lock`, but the awareness main thread does not take it at all — that is the gap `fb32d15` closed. Now verified independent of that guard: the index boundary is its own rwlock, so worker/awareness contention on `_eg_vindex` is handled whether or not the request lock is held.
|
||||||
|
|
||||||
|
**Still open:**
|
||||||
|
|
||||||
|
- The resident RAM graph wants the same publication boundary (§3.2). Until it has one, `fb32d15` cannot be deleted.
|
||||||
|
- `eg_vindex_view` holds the boundary for read across `engram_geo_reify_store`, which is a long pass. Correct, but it stalls the owner for that duration. If reify latency becomes a problem the answer is a refcounted snapshot, not a shorter lock.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7a. Evidence (measured 2026-08-16, `engram/test/run_vindex_concurrency_tests.sh`)
|
||||||
|
|
||||||
|
| half | before | after |
|
||||||
|
|---|---|---|
|
||||||
|
| `single` — 3000 vectors, 1 thread, ASan+UBSan | clean | clean |
|
||||||
|
| `readers` — 4 readers, no writer, TSan | **race** at `engram_vindex.c:195` (`visited_reset` ← `vindex_search`) | **clean** |
|
||||||
|
| `unsynchronized` — writer+reader, bare index, TSan | race | **race, expected and permanent** — now the proof the boundary must exist |
|
||||||
|
| `published` — owner + 4 readers through the boundary, TSan | *(did not exist)* | **clean**, all 3000 inserts landed |
|
||||||
|
|
||||||
|
No recall regression: `recall@10 = 0.9365` at `ef_search=128` (gate ≥ 0.90); the determinism test still yields byte-identical results across two independent builds.
|
||||||
|
|
||||||
|
Builds locally: all seven engram runtime translation units compile `-Wall -Wextra` clean, and the full engram binary links (`engram/dist/engram.c` + runtime, arm64). The one pre-existing `-Wcomment` warning in `el_runtime.c` is present at `a67452f` too.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. What this document is not
|
||||||
|
|
||||||
|
It is not an argument for a memory model in general, a garbage collector, process isolation between soul and engram, or a client/server split of the store. Each of those was considered and each addresses mutation that this change removes. They are answers to a question that stops being asked.
|
||||||
@@ -0,0 +1,91 @@
|
|||||||
|
// fitprobe.el — controlled growth-curve specimens for validating the complexity fitter.
|
||||||
|
//
|
||||||
|
// Three deliberately-shaped workloads. None depends on a real defect existing,
|
||||||
|
// which is the point: the fitter must be provable against KNOWN curves.
|
||||||
|
//
|
||||||
|
// linear — one allocation per item. count O(n), bytes O(n), time O(n)
|
||||||
|
// accum — rebuilds its accumulator. count O(n), bytes O(n^2), time O(n^2)
|
||||||
|
// compute — nested arithmetic, no alloc. count O(1), bytes O(1), time O(n^2)
|
||||||
|
//
|
||||||
|
// `compute` is the specimen that matters. It is the shape of el #132
|
||||||
|
// (strlen-per-character inside str_char_code): pure CPU, zero allocation.
|
||||||
|
// An allocation-only gate is structurally blind to it.
|
||||||
|
//
|
||||||
|
// No imports — uses runtime builtins directly so nothing collides.
|
||||||
|
|
||||||
|
fn work_linear(n: Int) -> Int {
|
||||||
|
let parts: [String] = native_list_empty()
|
||||||
|
let i: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let parts = native_list_append(parts, int_to_str(i))
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
return native_list_len(parts)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn work_accum(n: Int) -> Int {
|
||||||
|
let acc: String = ""
|
||||||
|
let i: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let acc = acc + "x"
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
return str_len(acc)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn work_compute(n: Int) -> Int {
|
||||||
|
// str_char_code is an opaque external call, so the C optimiser cannot
|
||||||
|
// reduce this nest to a closed form the way it does with `total + 1`.
|
||||||
|
// This is the exact shape of el #132: n scans over n characters, pure
|
||||||
|
// CPU, ZERO allocation.
|
||||||
|
let s: String = "abcdefghij"
|
||||||
|
let total: Int = 0
|
||||||
|
let i: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let j: Int = 0
|
||||||
|
while j < n {
|
||||||
|
let total = total + str_char_code(s, 0)
|
||||||
|
let j = j + 1
|
||||||
|
}
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
return total
|
||||||
|
}
|
||||||
|
|
||||||
|
fn run_one(mode: String, n: Int) {
|
||||||
|
let c0: Int = el_alloc_count()
|
||||||
|
let b0: Int = el_alloc_bytes()
|
||||||
|
let t0: Int = el_now_instant()
|
||||||
|
|
||||||
|
let r: Int = 0
|
||||||
|
if str_eq(mode, "linear") { let r = work_linear(n) }
|
||||||
|
if str_eq(mode, "accum") { let r = work_accum(n) }
|
||||||
|
if str_eq(mode, "compute") { let r = work_compute(n) }
|
||||||
|
|
||||||
|
let t1: Int = el_now_instant()
|
||||||
|
let c1: Int = el_alloc_count()
|
||||||
|
let b1: Int = el_alloc_bytes()
|
||||||
|
|
||||||
|
println(mode + "\t" + int_to_str(n)
|
||||||
|
+ "\t" + int_to_str(c1 - c0)
|
||||||
|
+ "\t" + int_to_str(b1 - b0)
|
||||||
|
+ "\t" + int_to_str((t1 - t0) / 1000)
|
||||||
|
+ "\t" + int_to_str(r))
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
fn sweep(mode: String) {
|
||||||
|
run_one(mode, 200)
|
||||||
|
run_one(mode, 400)
|
||||||
|
run_one(mode, 800)
|
||||||
|
run_one(mode, 1600)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
fn main() -> Int {
|
||||||
|
println("mode\tn\tallocs\tbytes\tusec\tsink")
|
||||||
|
sweep("linear")
|
||||||
|
sweep("accum")
|
||||||
|
sweep("compute")
|
||||||
|
return 0
|
||||||
|
}
|
||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// tests/native/test_compiler.el — comprehensive tests for the El compiler pipeline.
|
// tests/native/test_compiler.el — comprehensive tests for the El compiler pipeline.
|
||||||
//
|
//
|
||||||
// Tests the lexer (lexer.el), parser (parser.el), and codegen (codegen.el)
|
// Tests the lexer (lexer.el), parser (parser.el), and codegen (codegen.el)
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_codegen_js.el - basic tests for JS codegen features.
|
// test_codegen_js.el - basic tests for JS codegen features.
|
||||||
//
|
//
|
||||||
// These tests verify that core El language features produce correct values
|
// These tests verify that core El language features produce correct values
|
||||||
|
|||||||
@@ -0,0 +1,111 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
|
import "../../runtime/elbench.el"
|
||||||
|
|
||||||
|
// test_elbench.el — proves the growth-curve classifier against KNOWN curves.
|
||||||
|
//
|
||||||
|
// Every series below is real measured data from lang/tests/bench/fitprobe.el
|
||||||
|
// on a geometric sweep n = 200/400/800/1600. The classifier must be provable
|
||||||
|
// without depending on a live defect existing, which is the whole point of
|
||||||
|
// keeping controlled specimens.
|
||||||
|
|
||||||
|
fn _s4(a: Int, b: Int, c: Int, d: Int) -> [Int] {
|
||||||
|
let l: [Int] = native_list_empty()
|
||||||
|
let l = native_list_append(l, a)
|
||||||
|
let l = native_list_append(l, b)
|
||||||
|
let l = native_list_append(l, c)
|
||||||
|
let l = native_list_append(l, d)
|
||||||
|
return l
|
||||||
|
}
|
||||||
|
|
||||||
|
test "classifies a linear allocation series as O(n)" {
|
||||||
|
// fitprobe `linear`, allocation count
|
||||||
|
let v = _s4(208, 409, 810, 1611)
|
||||||
|
assert elb_measured_curve(v, 10) == 2, "linear allocs should classify O(n)"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "classifies a linear byte series as O(n)" {
|
||||||
|
// fitprobe `linear`, allocation bytes
|
||||||
|
let v = _s4(4786, 9682, 19474, 39658)
|
||||||
|
assert elb_measured_curve(v, 10) == 2, "linear bytes should classify O(n)"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "classifies a quadratic byte series as O(n^2)" {
|
||||||
|
// fitprobe `accum`, allocation bytes -- the accumulator-rebuild shape
|
||||||
|
let v = _s4(20300, 80600, 321200, 1282400)
|
||||||
|
assert elb_measured_curve(v, 10) == 4, "accum bytes should classify O(n^2)"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "accumulator count is linear -- proves count alone misses it" {
|
||||||
|
// Same run as above. The COUNT is exactly linear while bytes are
|
||||||
|
// quadratic. A count-only gate passes this defect clean.
|
||||||
|
let v = _s4(200, 400, 800, 1600)
|
||||||
|
assert elb_measured_curve(v, 10) == 2, "accum count classifies O(n)"
|
||||||
|
assert elb_gate(v, 2, 10) == 0, "count-only gate PASSES the quadratic"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "classifies a quadratic time series as O(n^2)" {
|
||||||
|
// fitprobe `compute` -- el #132's shape: n scans over n characters
|
||||||
|
let v = _s4(67, 205, 818, 3268)
|
||||||
|
assert elb_measured_curve(v, 10) == 4, "compute time should classify O(n^2)"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "REFUSES an all-zero series instead of calling it O(1)" {
|
||||||
|
// fitprobe `compute` allocation count. Pure CPU, allocates nothing.
|
||||||
|
// Reporting O(1) here would be a confident answer with nothing behind it.
|
||||||
|
let v = _s4(0, 0, 0, 0)
|
||||||
|
assert elb_gate(v, 2, 10) == 3, "all-zero series must be REFUSED"
|
||||||
|
assert elb_measured_curve(v, 10) < 0, "unclassifiable returns -1"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "REFUSES an implausibly flat series" {
|
||||||
|
// The shape produced when clang closes a loop to a multiply: a real
|
||||||
|
// answer, no work done, no movement across an 8x input range.
|
||||||
|
let v = _s4(1000, 1001, 1002, 1003)
|
||||||
|
assert elb_gate(v, 2, 10) == 3, "hard-flat series must be REFUSED"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "gate FAILS a quadratic declared as linear" {
|
||||||
|
let v = _s4(20300, 80600, 321200, 1282400)
|
||||||
|
assert elb_gate(v, 2, 10) == 1, "O(n^2) measured vs O(n) declared must FAIL"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "gate PASSES a linear series declared as linear" {
|
||||||
|
let v = _s4(208, 409, 810, 1611)
|
||||||
|
assert elb_gate(v, 2, 10) == 0, "O(n) measured vs O(n) declared must PASS"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "gate reports BETTER when measured beats the declared bound" {
|
||||||
|
let v = _s4(208, 409, 810, 1611)
|
||||||
|
assert elb_gate(v, 4, 10) == 4, "O(n) measured vs O(n^2) declared is BETTER"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "gate reports INDETERMINATE on disagreeing ratios" {
|
||||||
|
// fitprobe `linear` WALL TIME at these sizes: 26/19/43/78 microseconds.
|
||||||
|
// Ratios 0.73, 2.26, 1.81 disagree well past the noise threshold. The
|
||||||
|
// honest answer is "cannot tell", not a classification -- this is exactly
|
||||||
|
// why benchmarks need auto-scaled iteration counts rather than one shot.
|
||||||
|
let v = _s4(26, 19, 43, 78)
|
||||||
|
assert elb_gate(v, 2, 10) == 2, "disagreeing ratios must be INDETERMINATE"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "black_box is a real barrier and returns its input" {
|
||||||
|
assert el_black_box(42) == 42, "black_box is value-preserving"
|
||||||
|
let s: Int = 0
|
||||||
|
let i: Int = 0
|
||||||
|
while i < 100 {
|
||||||
|
// Bind the call before using it in arithmetic: `x + call(...)`
|
||||||
|
// lowers to el_str_concat() on integers. Same inference defect
|
||||||
|
// as `call(...) == y` lowering to str_eq().
|
||||||
|
let bx: Int = el_black_box(1)
|
||||||
|
let s = s + bx
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
assert s == 100, "black_box does not disturb the computation"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "curve names round-trip" {
|
||||||
|
assert elb_curve_from_name("O(n)") == 2, "O(n) parses"
|
||||||
|
assert elb_curve_from_name("O(n^2)") == 4, "O(n^2) parses"
|
||||||
|
assert str_eq(elb_curve_name(4), "O(n^2)"), "O(n^2) renders"
|
||||||
|
assert elb_curve_from_name("O(nonsense)") < 0, "unknown curve is -1"
|
||||||
|
}
|
||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_env.el - native test suite for runtime/env.el
|
// test_env.el - native test suite for runtime/env.el
|
||||||
//
|
//
|
||||||
// Covers: env() for reading environment variables, args() returning a list,
|
// Covers: env() for reading environment variables, args() returning a list,
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_fs.el - native test suite for runtime/fs.el
|
// test_fs.el - native test suite for runtime/fs.el
|
||||||
//
|
//
|
||||||
// Covers: fs_write/read round-trip, fs_exists, fs_mkdir, fs_list,
|
// Covers: fs_write/read round-trip, fs_exists, fs_mkdir, fs_list,
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_json.el - native test suite for runtime/json.el
|
// test_json.el - native test suite for runtime/json.el
|
||||||
//
|
//
|
||||||
// Covers: json_get (dot-path), typed extractors (int, bool, float),
|
// Covers: json_get (dot-path), typed extractors (int, bool, float),
|
||||||
|
|||||||
@@ -0,0 +1,178 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
|
import "../../runtime/elbench.el"
|
||||||
|
|
||||||
|
// test_lexer_scaling.el — THE ARMED GATE.
|
||||||
|
//
|
||||||
|
// This is the regression test that would have caught el #132.
|
||||||
|
//
|
||||||
|
// #132 was a strlen() inside str_char_code() and str_slice(). The lexer walks
|
||||||
|
// source one character at a time, so every character access rescanned the whole
|
||||||
|
// remaining input: O(n) per character over n characters = O(n^2). It shipped for
|
||||||
|
// months. It was found by a geometric sweep, not by reading code.
|
||||||
|
//
|
||||||
|
// So this test IS a geometric sweep. It scans a string of length n, character by
|
||||||
|
// character, at four doubling sizes, and asserts the cost is linear. If anyone
|
||||||
|
// reintroduces a per-character rescan — in str_char_code, in str_slice, in any
|
||||||
|
// accessor the lexer leans on — the measured curve becomes O(n^2) and this fails.
|
||||||
|
//
|
||||||
|
// The value is in it being ARMED, not in it currently failing. It passes today
|
||||||
|
// because #132 is fixed. That is the correct state for a regression gate.
|
||||||
|
//
|
||||||
|
// Note the deliberate `let c: Int = str_char_code(...)` binding in the scan loop.
|
||||||
|
// Inlining it as `total + str_char_code(s, i)` lowers to el_str_concat() on
|
||||||
|
// integers — the Plus arm of the operator-typing family, still open at the time
|
||||||
|
// of writing. Binding first is the safe form.
|
||||||
|
|
||||||
|
// _mk_string — build a string of length >= n by DOUBLING.
|
||||||
|
//
|
||||||
|
// Deliberately not `s = s + "x"` n times: that is itself quadratic in bytes and
|
||||||
|
// would contaminate the very measurement this test exists to take. Doubling
|
||||||
|
// allocates ~2n total.
|
||||||
|
fn _mk_string(n: Int) -> String {
|
||||||
|
let s: String = "abcdefgh"
|
||||||
|
while str_len(s) < n {
|
||||||
|
let s = s + s
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
// _scan — walk the string one character at a time, REPS times.
|
||||||
|
//
|
||||||
|
// This is the lexer's access pattern reduced to its essential shape. The
|
||||||
|
// repetitions lift the measurement clear of timer resolution; without them the
|
||||||
|
// smaller sizes land in noise and the classifier correctly reports
|
||||||
|
// INDETERMINATE rather than guessing.
|
||||||
|
fn _scan(s: String, n: Int, reps: Int) -> Int {
|
||||||
|
let total: Int = 0
|
||||||
|
let r: Int = 0
|
||||||
|
while r < reps {
|
||||||
|
let i: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let c: Int = str_char_code(s, i)
|
||||||
|
let total = total + c
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
let r = r + 1
|
||||||
|
}
|
||||||
|
return total
|
||||||
|
}
|
||||||
|
|
||||||
|
// _measure_scan — microseconds for a full scan sweep point.
|
||||||
|
fn _measure_scan(n: Int, reps: Int) -> Int {
|
||||||
|
let s: String = _mk_string(n)
|
||||||
|
// WARMUP, discarded. Without it the small-n end of the sweep is dominated
|
||||||
|
// by cold caches and reads as superlinear on genuinely linear work --
|
||||||
|
// measured ratios 3.37 2.92 1.76 1.65 on exactly this workload.
|
||||||
|
let w: Int = _scan(s, n, 2)
|
||||||
|
let wj: Int = el_black_box(w)
|
||||||
|
let t0: Int = el_now_instant()
|
||||||
|
let got: Int = _scan(s, n, reps)
|
||||||
|
let t1: Int = el_now_instant()
|
||||||
|
// Feed the result through the barrier so the scan cannot be elided.
|
||||||
|
let sink: Int = el_black_box(got)
|
||||||
|
if sink == 0 { println("") }
|
||||||
|
return (t1 - t0) / 1000
|
||||||
|
}
|
||||||
|
|
||||||
|
fn _series4(a: Int, b: Int, c: Int, d: Int) -> [Int] {
|
||||||
|
let l: [Int] = native_list_empty()
|
||||||
|
let l = native_list_append(l, a)
|
||||||
|
let l = native_list_append(l, b)
|
||||||
|
let l = native_list_append(l, c)
|
||||||
|
let l = native_list_append(l, d)
|
||||||
|
return l
|
||||||
|
}
|
||||||
|
|
||||||
|
test "character scan is LINEAR in time -- regression gate for el #132" {
|
||||||
|
let reps: Int = 40
|
||||||
|
let t1: Int = _measure_scan(16384, reps)
|
||||||
|
let t2: Int = _measure_scan(32768, reps)
|
||||||
|
let t3: Int = _measure_scan(65536, reps)
|
||||||
|
let t4: Int = _measure_scan(131072, reps)
|
||||||
|
let series: [Int] = _series4(t1, t2, t3, t4)
|
||||||
|
|
||||||
|
let verdict: Int = elb_gate(series, 2, 50)
|
||||||
|
let measured: Int = elb_measured_curve(series, 50)
|
||||||
|
|
||||||
|
// Report the actual numbers regardless of outcome. A gate that fires
|
||||||
|
// without showing its evidence is just an assertion.
|
||||||
|
println(" scan us: " + int_to_str(t1) + " " + int_to_str(t2) + " "
|
||||||
|
+ int_to_str(t3) + " " + int_to_str(t4)
|
||||||
|
+ " -> " + elb_curve_name(measured) + " [" + elb_verdict_name(verdict) + "]")
|
||||||
|
|
||||||
|
// PASS (0) or BETTER (4) are both acceptable. FAIL (1) means someone
|
||||||
|
// reintroduced superlinear per-character cost. REFUSED (3) or
|
||||||
|
// INDETERMINATE (2) mean the measurement is untrustworthy -- which is
|
||||||
|
// also a failure of this test, deliberately: a gate that cannot measure
|
||||||
|
// must not report success.
|
||||||
|
assert verdict == 0 || verdict == 4, "character scan must measure O(n) or better"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "string building by doubling stays linear in allocated bytes" {
|
||||||
|
let b1: Int = el_alloc_bytes()
|
||||||
|
let s1: String = _mk_string(8192)
|
||||||
|
let b2: Int = el_alloc_bytes()
|
||||||
|
let s2: String = _mk_string(16384)
|
||||||
|
let b3: Int = el_alloc_bytes()
|
||||||
|
let s3: String = _mk_string(32768)
|
||||||
|
let b4: Int = el_alloc_bytes()
|
||||||
|
let s4: String = _mk_string(65536)
|
||||||
|
let b5: Int = el_alloc_bytes()
|
||||||
|
|
||||||
|
let series: [Int] = _series4(b2 - b1, b3 - b2, b4 - b3, b5 - b4)
|
||||||
|
let verdict: Int = elb_gate(series, 2, 1000)
|
||||||
|
let measured: Int = elb_measured_curve(series, 1000)
|
||||||
|
println(" bytes: " + int_to_str(b2 - b1) + " " + int_to_str(b3 - b2) + " "
|
||||||
|
+ int_to_str(b4 - b3) + " " + int_to_str(b5 - b4)
|
||||||
|
+ " -> " + elb_curve_name(measured) + " [" + elb_verdict_name(verdict) + "]")
|
||||||
|
|
||||||
|
assert verdict == 0 || verdict == 4, "doubling build must be O(n) in bytes"
|
||||||
|
assert str_len(s4) >= 65536, "final string reached the requested size"
|
||||||
|
}
|
||||||
|
|
||||||
|
// _scan_quadratic — a DELIBERATELY quadratic scan: for each position, rescan
|
||||||
|
// from the start. This is precisely what el #132 did — strlen() from offset 0
|
||||||
|
// on every character access — reproduced here so the gate can be proven to
|
||||||
|
// FIRE, not merely to pass on healthy code. An unproven gate is decoration.
|
||||||
|
fn _scan_quadratic(s: String, n: Int) -> Int {
|
||||||
|
let total: Int = 0
|
||||||
|
let i: Int = 0
|
||||||
|
while i < n {
|
||||||
|
let j: Int = 0
|
||||||
|
while j < i {
|
||||||
|
let c: Int = str_char_code(s, j)
|
||||||
|
let total = total + c
|
||||||
|
let j = j + 1
|
||||||
|
}
|
||||||
|
let i = i + 1
|
||||||
|
}
|
||||||
|
return total
|
||||||
|
}
|
||||||
|
|
||||||
|
fn _measure_quadratic(n: Int) -> Int {
|
||||||
|
let s: String = _mk_string(n)
|
||||||
|
let w: Int = _scan_quadratic(s, 64)
|
||||||
|
let wj: Int = el_black_box(w)
|
||||||
|
let t0: Int = el_now_instant()
|
||||||
|
let got: Int = _scan_quadratic(s, n)
|
||||||
|
let t1: Int = el_now_instant()
|
||||||
|
let sink: Int = el_black_box(got)
|
||||||
|
return (t1 - t0) / 1000
|
||||||
|
}
|
||||||
|
|
||||||
|
test "the gate FIRES on a live quadratic scan -- proves it is armed" {
|
||||||
|
let q1: Int = _measure_quadratic(1024)
|
||||||
|
let q2: Int = _measure_quadratic(2048)
|
||||||
|
let q3: Int = _measure_quadratic(4096)
|
||||||
|
let q4: Int = _measure_quadratic(8192)
|
||||||
|
let series: [Int] = _series4(q1, q2, q3, q4)
|
||||||
|
|
||||||
|
let verdict: Int = elb_gate(series, 2, 50)
|
||||||
|
let measured: Int = elb_measured_curve(series, 50)
|
||||||
|
println(" quad us: " + int_to_str(q1) + " " + int_to_str(q2) + " "
|
||||||
|
+ int_to_str(q3) + " " + int_to_str(q4)
|
||||||
|
+ " -> " + elb_curve_name(measured) + " [" + elb_verdict_name(verdict) + "]")
|
||||||
|
|
||||||
|
assert measured == 4, "a rescan-from-zero workload must classify O(n^2)"
|
||||||
|
assert verdict == 1, "declared O(n) against measured O(n^2) must FAIL the gate"
|
||||||
|
}
|
||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_math.el - native test suite for runtime/math.el
|
// test_math.el - native test suite for runtime/math.el
|
||||||
//
|
//
|
||||||
// Covers: integer math (abs, max, min), float math (sqrt, log, sin, cos, pi),
|
// Covers: integer math (abs, max, min), float math (sqrt, log, sin, cos, pi),
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_state.el - native test suite for runtime/state.el
|
// test_state.el - native test suite for runtime/state.el
|
||||||
//
|
//
|
||||||
// Covers: state_set/get/del, state_has, state_get_or, state_keys,
|
// Covers: state_set/get/del, state_has, state_get_or, state_keys,
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_string.el - native test suite for runtime/string.el
|
// test_string.el - native test suite for runtime/string.el
|
||||||
//
|
//
|
||||||
// Covers: type conversions, core primitives, comparison and search,
|
// Covers: type conversions, core primitives, comparison and search,
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_text.el - native test suite for text primitives.
|
// test_text.el - native test suite for text primitives.
|
||||||
//
|
//
|
||||||
// Mirrors the acceptance corpus in tests/text/examples/ using the
|
// Mirrors the acceptance corpus in tests/text/examples/ using the
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// test_time.el - native test suite for runtime/time.el
|
// test_time.el - native test suite for runtime/time.el
|
||||||
//
|
//
|
||||||
// Covers: time_now (positive timestamp), time_to_parts (UTC decomposition),
|
// Covers: time_now (positive timestamp), time_to_parts (UTC decomposition),
|
||||||
|
|||||||
@@ -0,0 +1,234 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
|
// test_transduce.el — geometry as a first-class El value, and realizers
|
||||||
|
// declared in El rather than patched into the runtime.
|
||||||
|
//
|
||||||
|
// WHAT IS ACTUALLY UNDER TEST. Until 2026-08-16 no El ingest path could carry
|
||||||
|
// a vector: nodes took text, and geometry was DERIVED from that text. Text was
|
||||||
|
// therefore the mandatory entry medium, so any non-text modality had to be
|
||||||
|
// DESCRIBED in prose first and the geometry we reasoned over was the geometry
|
||||||
|
// OF THE DESCRIPTION, not of the signal. The fix has two halves, and this file
|
||||||
|
// exercises both:
|
||||||
|
//
|
||||||
|
// 1. Geometry is a VALUE — it carries its own width, so nothing has to
|
||||||
|
// assert a width against a string's length.
|
||||||
|
// 2. A REALIZER is an ordinary El function. `tone_realizer` below is not in
|
||||||
|
// the runtime, is not known to the compiler, and is not special in any
|
||||||
|
// way; it is registered BY NAME and dispatched to through transduce().
|
||||||
|
// That is the load-bearing claim: adding a modality must not require a
|
||||||
|
// runtime patch, or nothing has actually moved into the language.
|
||||||
|
//
|
||||||
|
// COMPARISON DISCIPLINE IN THIS FILE (measured 2026-08-16, not stylistic):
|
||||||
|
// elc lowers `a == b` to a NUMERIC comparison only when both operand names are
|
||||||
|
// in the per-function int-name set, which `let x: Int` populates. A bare call
|
||||||
|
// like `geometry_is(g) == 0` is not a registered name, so it lowers to
|
||||||
|
// `str_eq(...)` — strcmp on two integers reinterpreted as pointers. `<` and `>`
|
||||||
|
// lower directly via binop_to_c with no type inference at all, so truthiness is
|
||||||
|
// written `> 0` / `< 1` here, and any exact `==` is done on a value first bound
|
||||||
|
// through `let x: Int`.
|
||||||
|
|
||||||
|
// ── A realizer, written entirely in El ──────────────────────────────────────
|
||||||
|
// Maps a "tone" signal into a 4-component geometry. Deliberately trivial —
|
||||||
|
// what is being proven is that an El function can BE a realizer, not that
|
||||||
|
// this is good acoustics. The one real property it has: distinct signals
|
||||||
|
// produce distinct geometry, so the test can tell transduction from a stub.
|
||||||
|
fn tone_realizer(signal: String) -> Geometry {
|
||||||
|
let g: Geometry = geometry_new(4)
|
||||||
|
let n: Int = str_len(signal)
|
||||||
|
let a: Int = geometry_set(g, 0, int_to_float(n))
|
||||||
|
let b: Int = geometry_set(g, 1, int_to_float(n * 2))
|
||||||
|
let c: Int = geometry_set(g, 2, int_to_float(n * 3))
|
||||||
|
let d: Int = geometry_set(g, 3, int_to_float(n * 4))
|
||||||
|
g
|
||||||
|
}
|
||||||
|
|
||||||
|
// A second realizer for a different modality, to prove the registry keys on
|
||||||
|
// modality and does not just hand back "the last thing registered".
|
||||||
|
fn pulse_realizer(signal: String) -> Geometry {
|
||||||
|
let g: Geometry = geometry_new(2)
|
||||||
|
let a: Int = geometry_set(g, 0, 1.0)
|
||||||
|
let b: Int = geometry_set(g, 1, 0.0)
|
||||||
|
g
|
||||||
|
}
|
||||||
|
|
||||||
|
// A deliberately BROKEN realizer: it returns something that is not a Geometry.
|
||||||
|
// transduce() must not hand this back to a caller as if it were one.
|
||||||
|
fn bogus_realizer(signal: String) -> Geometry {
|
||||||
|
return 12345
|
||||||
|
}
|
||||||
|
|
||||||
|
test "geometry-is-a-value-with-its-own-width" {
|
||||||
|
let g: Geometry = geometry_new(8)
|
||||||
|
let live: Int = geometry_is(g)
|
||||||
|
assert live > 0, "geometry_new returns a live Geometry"
|
||||||
|
let d: Int = geometry_dim(g)
|
||||||
|
assert d == 8, "a Geometry carries its own width"
|
||||||
|
let freed: Int = geometry_free(g)
|
||||||
|
assert freed > 0, "geometry_free reports what it did"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "geometry-rejects-nonsense-without-an-arbitrary-bound" {
|
||||||
|
// dim <= 0 is not a width. Note there is deliberately no MAX dim here:
|
||||||
|
// #141 needed `dim <= 8192` only to bound an allocation sized from a
|
||||||
|
// caller's claim about a string. A value that carries its own width has
|
||||||
|
// nothing left to validate, so the only failure left is allocation.
|
||||||
|
let zero: Geometry = geometry_new(0)
|
||||||
|
let z: Int = geometry_is(zero)
|
||||||
|
assert z < 1, "dim 0 is not a geometry"
|
||||||
|
let neg: Geometry = geometry_new(-4)
|
||||||
|
let n: Int = geometry_is(neg)
|
||||||
|
assert n < 1, "negative dim is not a geometry"
|
||||||
|
// Accessors must be total: a non-geometry is 0-width, never a crash.
|
||||||
|
let nd: Int = geometry_dim(0)
|
||||||
|
assert nd < 1, "geometry_dim of a non-geometry is 0"
|
||||||
|
let ni: Int = geometry_is(0)
|
||||||
|
assert ni < 1, "geometry_is of a non-geometry is 0"
|
||||||
|
let nf: Int = geometry_free(0)
|
||||||
|
assert nf < 1, "geometry_free of a non-geometry is a no-op"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "geometry-components-round-trip" {
|
||||||
|
let g: Geometry = geometry_new(3)
|
||||||
|
let s0: Int = geometry_set(g, 0, 1.5)
|
||||||
|
let s1: Int = geometry_set(g, 1, -2.5)
|
||||||
|
assert s0 > 0, "set in range succeeds"
|
||||||
|
let oob: Int = geometry_set(g, 3, 9.0)
|
||||||
|
assert oob < 1, "set out of range is refused, not silently dropped"
|
||||||
|
let v0: Float = geometry_get(g, 0)
|
||||||
|
let d0: Float = v0 - 1.5
|
||||||
|
assert d0 < 0.001, "component 0 round-trips"
|
||||||
|
assert d0 > -0.001, "component 0 round-trips"
|
||||||
|
let v1: Float = geometry_get(g, 1)
|
||||||
|
let d1: Float = v1 + 2.5
|
||||||
|
assert d1 < 0.001, "component 1 round-trips (negative)"
|
||||||
|
assert d1 > -0.001, "component 1 round-trips (negative)"
|
||||||
|
let freed: Int = geometry_free(g)
|
||||||
|
}
|
||||||
|
|
||||||
|
test "hex-is-an-edge-adapter-and-derives-its-own-width" {
|
||||||
|
// 2 components, little-endian float32: 1.0 = 0000803f, 2.0 = 00000040.
|
||||||
|
let g: Geometry = geometry_from_f32le_hex("0000803f00000040")
|
||||||
|
let live: Int = geometry_is(g)
|
||||||
|
assert live > 0, "valid hex decodes to a Geometry"
|
||||||
|
let d: Int = geometry_dim(g)
|
||||||
|
assert d == 2, "width is DERIVED from the input, never supplied"
|
||||||
|
let a: Float = geometry_get(g, 0)
|
||||||
|
let da: Float = a - 1.0
|
||||||
|
assert da < 0.001, "first component decoded"
|
||||||
|
assert da > -0.001, "first component decoded"
|
||||||
|
let b: Float = geometry_get(g, 1)
|
||||||
|
let db: Float = b - 2.0
|
||||||
|
assert db < 0.001, "second component decoded"
|
||||||
|
assert db > -0.001, "second component decoded"
|
||||||
|
// Egress adapter is the exact inverse.
|
||||||
|
let back: String = geometry_to_f32le_hex(g)
|
||||||
|
assert str_eq(back, "0000803f00000040"), "hex round-trips exactly"
|
||||||
|
let freed: Int = geometry_free(g)
|
||||||
|
}
|
||||||
|
|
||||||
|
test "hex-rejects-malformed-input" {
|
||||||
|
let empty: Geometry = geometry_from_f32le_hex("")
|
||||||
|
let e: Int = geometry_is(empty)
|
||||||
|
assert e < 1, "empty hex is not a geometry"
|
||||||
|
let ragged: Geometry = geometry_from_f32le_hex("0000803f0000")
|
||||||
|
let r: Int = geometry_is(ragged)
|
||||||
|
assert r < 1, "length not a multiple of 8 is refused"
|
||||||
|
let nonhex: Geometry = geometry_from_f32le_hex("zzzzzzzz")
|
||||||
|
let nh: Int = geometry_is(nonhex)
|
||||||
|
assert nh < 1, "non-hex characters are refused"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "a-realizer-declared-in-el-is-a-first-class-realizer" {
|
||||||
|
// THE CLAIM: tone_realizer is an ordinary El function. It is not in the
|
||||||
|
// runtime and the compiler knows nothing about it. Registering it by name
|
||||||
|
// is enough to make it the organ for a modality.
|
||||||
|
let reg: Int = realizer_register("tone", "tone_realizer")
|
||||||
|
assert reg > 0, "an El fn registers as a realizer by name"
|
||||||
|
let has: Int = realizer_has("tone")
|
||||||
|
assert has > 0, "the modality now has an organ"
|
||||||
|
|
||||||
|
let g: Geometry = transduce("aaa", "tone")
|
||||||
|
let live: Int = geometry_is(g)
|
||||||
|
assert live > 0, "transduce returns real geometry"
|
||||||
|
let d: Int = geometry_dim(g)
|
||||||
|
assert d == 4, "the El realizer determined the width, not the runtime"
|
||||||
|
// str_len("aaa") == 3, so component 0 must be 3.0 — proof the signal
|
||||||
|
// actually reached the El function rather than a stub answering for it.
|
||||||
|
let c0: Float = geometry_get(g, 0)
|
||||||
|
let dc: Float = c0 - 3.0
|
||||||
|
assert dc < 0.001, "the signal reached the El realizer"
|
||||||
|
assert dc > -0.001, "the signal reached the El realizer"
|
||||||
|
let freed: Int = geometry_free(g)
|
||||||
|
}
|
||||||
|
|
||||||
|
test "distinct-signals-transduce-to-distinct-geometry" {
|
||||||
|
let reg: Int = realizer_register("tone", "tone_realizer")
|
||||||
|
let g1: Geometry = transduce("aa", "tone")
|
||||||
|
let g2: Geometry = transduce("aaaaa", "tone")
|
||||||
|
let a: Float = geometry_get(g1, 0)
|
||||||
|
let b: Float = geometry_get(g2, 0)
|
||||||
|
let diff: Float = b - a
|
||||||
|
// 5 - 2 = 3. If transduction were a stub these would be equal.
|
||||||
|
assert diff > 2.9, "different signals produce different geometry"
|
||||||
|
assert diff < 3.1, "different signals produce different geometry"
|
||||||
|
let f1: Int = geometry_free(g1)
|
||||||
|
let f2: Int = geometry_free(g2)
|
||||||
|
}
|
||||||
|
|
||||||
|
test "the-registry-keys-on-modality" {
|
||||||
|
let r1: Int = realizer_register("tone", "tone_realizer")
|
||||||
|
let r2: Int = realizer_register("pulse", "pulse_realizer")
|
||||||
|
assert r2 > 0, "a second modality registers independently"
|
||||||
|
let gt: Geometry = transduce("aaa", "tone")
|
||||||
|
let gp: Geometry = transduce("aaa", "pulse")
|
||||||
|
let dt: Int = geometry_dim(gt)
|
||||||
|
let dp: Int = geometry_dim(gp)
|
||||||
|
assert dt == 4, "tone still routes to its own realizer"
|
||||||
|
assert dp == 2, "pulse routes to a different realizer"
|
||||||
|
let f1: Int = geometry_free(gt)
|
||||||
|
let f2: Int = geometry_free(gp)
|
||||||
|
}
|
||||||
|
|
||||||
|
test "no-organ-is-reported-as-no-organ" {
|
||||||
|
// A modality with no realizer must transduce to NOTHING. It must never
|
||||||
|
// fall back to embedding a description of the signal and calling that
|
||||||
|
// perception — that silent substitution is the entire defect this change
|
||||||
|
// exists to end.
|
||||||
|
let has: Int = realizer_has("echolocation")
|
||||||
|
assert has < 1, "unregistered modality has no organ"
|
||||||
|
let g: Geometry = transduce("anything", "echolocation")
|
||||||
|
let live: Int = geometry_is(g)
|
||||||
|
assert live < 1, "no realizer means no geometry, not fake geometry"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "registration-of-an-unresolvable-name-fails-loudly" {
|
||||||
|
// Reported at the moment of WIRING, not later as "this modality mysteriously
|
||||||
|
// produces nothing". Distinguishing "no organ" from "broken organ" is the
|
||||||
|
// lesson that made this whole change necessary.
|
||||||
|
let bad: Int = realizer_register("ghost", "no_such_function_anywhere")
|
||||||
|
assert bad < 1, "an unresolvable realizer name is a registration failure"
|
||||||
|
let has: Int = realizer_has("ghost")
|
||||||
|
assert has < 1, "and nothing gets registered"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "a-realizer-returning-non-geometry-transduces-nothing" {
|
||||||
|
let reg: Int = realizer_register("bogus", "bogus_realizer")
|
||||||
|
assert reg > 0, "the symbol resolves, so registration succeeds"
|
||||||
|
// ...but the contract is enforced at the boundary, so the caller never
|
||||||
|
// receives a value that would misbehave far away from here.
|
||||||
|
let g: Geometry = transduce("x", "bogus")
|
||||||
|
let live: Int = geometry_is(g)
|
||||||
|
assert live < 1, "a non-Geometry return transduced nothing"
|
||||||
|
}
|
||||||
|
|
||||||
|
test "norm-lets-a-caller-check-a-realizer-emitted-signal" {
|
||||||
|
let g: Geometry = geometry_new(2)
|
||||||
|
let z: Float = geometry_norm(g)
|
||||||
|
assert z < 0.001, "a fresh geometry is zero — norm says so"
|
||||||
|
let s0: Int = geometry_set(g, 0, 3.0)
|
||||||
|
let s1: Int = geometry_set(g, 1, 4.0)
|
||||||
|
let n: Float = geometry_norm(g)
|
||||||
|
let dn: Float = n - 5.0
|
||||||
|
assert dn < 0.001, "3-4-5: norm is 5"
|
||||||
|
assert dn > -0.001, "3-4-5: norm is 5"
|
||||||
|
let freed: Int = geometry_free(g)
|
||||||
|
}
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
fn getstr(x: String) -> String { return x }
|
||||||
|
fn getint(x: Int) -> Int { return x }
|
||||||
|
fn ok(label: String) -> Void { println("ok " + label) }
|
||||||
|
fn bad(label: String) -> Void { println("FAIL " + label) }
|
||||||
|
|
||||||
|
let s1: String = "hello"
|
||||||
|
let s2: String = "hello"
|
||||||
|
let s3: String = "world"
|
||||||
|
let i1: Int = 5
|
||||||
|
let i2: Int = 5
|
||||||
|
let i3: Int = 9
|
||||||
|
|
||||||
|
if "abc" == "abc" { ok("str literal eq") } else { bad("str literal eq") }
|
||||||
|
if "abc" == "xyz" { bad("str literal ne") } else { ok("str literal ne") }
|
||||||
|
if s1 == s2 { ok("str var eq") } else { bad("str var eq") }
|
||||||
|
if s1 == s3 { bad("str var ne") } else { ok("str var ne") }
|
||||||
|
if getstr("hi") == "hi" { ok("str call vs literal") } else { bad("str call vs literal") }
|
||||||
|
if s1 == getstr("hello") { ok("str var vs call") } else { bad("str var vs call") }
|
||||||
|
if s1 == getstr("nope") { bad("str var vs call ne") } else { ok("str var vs call ne") }
|
||||||
|
if i1 == i2 { ok("int var eq") } else { bad("int var eq") }
|
||||||
|
if i1 == i3 { bad("int var ne") } else { ok("int var ne") }
|
||||||
|
if getint(5) == i1 { ok("int call vs var") } else { bad("int call vs var") }
|
||||||
|
if getint(9) == i1 { bad("int call vs var ne") } else { ok("int call vs var ne") }
|
||||||
|
if s1 != s3 { ok("str NOTEQ") } else { bad("str NOTEQ") }
|
||||||
|
if s1 != s2 { bad("str NOTEQ same") } else { ok("str NOTEQ same") }
|
||||||
|
if i1 != i3 { ok("int NOTEQ") } else { bad("int NOTEQ") }
|
||||||
|
if getint(9) != i1 { ok("int call NOTEQ") } else { bad("int call NOTEQ") }
|
||||||
|
println("done")
|
||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import "../../runtime/eltest.el"
|
||||||
// tests/runtime/string_test.el — Test suite for runtime/string.el
|
// tests/runtime/string_test.el — Test suite for runtime/string.el
|
||||||
//
|
//
|
||||||
// Exercises every public function exported by runtime/string.el using the
|
// Exercises every public function exported by runtime/string.el using the
|
||||||
|
|||||||
Reference in New Issue
Block a user