Compare commits
1 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 8c94d92033 |
@@ -1,605 +0,0 @@
|
||||
# Cognitive Architecture — Design Doc
|
||||
|
||||
**The buildable form of the "one operation" theory of cognition.**
|
||||
|
||||
Status: DESIGN. Nothing here is built yet except where explicitly marked
|
||||
"EXISTS" against a cited C symbol. A build agent executes from this doc.
|
||||
Offline design only — this pass changes no code.
|
||||
|
||||
Source of theory: Neuron memory `bdc8a488-146d-4ccb-a5c8-d8c0a008534e`.
|
||||
Source of existing engram substrate (cited throughout): the runtime on branch
|
||||
`feat/self-reification-20260814` —
|
||||
`lang/runtime/engram_reason.{c,h}`, `engram_verify.{c,h}`,
|
||||
`engram_geometry.{c,h}`, `engram_store.{c,h}`, plus the reification beat and the
|
||||
RAM activation graph compiled into `~/.neuron/bin/engram`.
|
||||
|
||||
---
|
||||
|
||||
## 0. The claim, stated plainly
|
||||
|
||||
Cognition is **one operation**, not eight. The named faculties —
|
||||
deduce / abduce / analogy / induce / causal / plan / predict / perspective —
|
||||
are human *labels* on regions of a single operation's steering space. They are
|
||||
not separately invoked and not separately implemented. The operation is:
|
||||
|
||||
> **think** = a directed traversal of the geometry from an *anchor*, steered by
|
||||
> a *prior*, whose output is a **gradient** (a distribution / direction over the
|
||||
> geometry), never a point. Collapse-to-a-point happens only at expression.
|
||||
|
||||
Three things follow, and they are the whole design:
|
||||
|
||||
1. **The operator collapse is already half-written in C.** The five reasoning
|
||||
operators in `engram_reason.c` already compose over *one* shared primitive —
|
||||
`engram_reason_point_fit` — plus a small geo-algebra
|
||||
(combine / subtract / analogy-rotate / distance). The verifier
|
||||
(`engram_verify.c`) is built on the same `point_fit`. What is missing is not
|
||||
the primitive; it is (a) making the *prior* a first-class learnable object
|
||||
instead of a hard-coded parameter, and (b) closing the learning loop.
|
||||
|
||||
2. **Grounding = learning = the same loop.** "Getting better" at any faculty is
|
||||
not changing the operation. It is *calibrating the steering-prior against
|
||||
outcomes*. Code freezes; priors grow. The correspondence-check that today
|
||||
lives offline (Python, the grounding-floor + differential-drop governor, "#43")
|
||||
must move **into the geometry, reflexive** — think scoring its own gradient
|
||||
against outcome and refining the prior on the error. That reflexive
|
||||
correspondence-loop *is* the learning engine and is the core unbuilt thing.
|
||||
|
||||
3. **The ungrounded is primary.** The engram *holds* anything unconditionally.
|
||||
Grounding is a *relation* (an edge, grounded-for-whom), not a gate. The
|
||||
honesty floor applies only to **assertion**. A fully-grounded mind is dead;
|
||||
the ungrounded is both the fuel (raw material for grounding) and the pull
|
||||
(curiosity = leaning toward one's own ungrounded regions).
|
||||
|
||||
Everything below makes these concrete and buildable, and defines what
|
||||
"completion" means, staged so the first milestone is a real end-to-end slice.
|
||||
|
||||
---
|
||||
|
||||
## 1. THE ONE OPERATION — `think`
|
||||
|
||||
### 1.1 Signature
|
||||
|
||||
```
|
||||
think(anchor, prior, aperture?) -> gradient
|
||||
```
|
||||
|
||||
- **anchor** — a location to traverse *from*. Either a node id (re-origin on that
|
||||
node's descriptor) or a raw point `x ∈ R^dim` (a query embedding). The anchor
|
||||
fixes the frame; every read is *from a vantage*, never view-from-nowhere.
|
||||
- **prior** — a learnable bias/direction over the geometry that *steers* the
|
||||
traversal (§2). A prior is a first-class stored object, not a call argument
|
||||
baked into C.
|
||||
- **aperture** — optional read-width / veil / field-selector (§3). Absent =
|
||||
self-mode full aperture.
|
||||
- **gradient** — the output. A `GeoGradient`: a direction + a spread over the
|
||||
geometry, *plus* the read neighborhood it was computed against. Not a point.
|
||||
A spiked gradient = "exact" (deduction); a spread gradient = "fuzzy"
|
||||
(prediction). The gradient is *also the next steering direction* — cognition
|
||||
is a flow down a prior-shaped landscape, closed-loop.
|
||||
|
||||
```c
|
||||
/* NEW. The output type. */
|
||||
typedef struct {
|
||||
int dim;
|
||||
float* direction; /* unit steering vector in the anchor's frame */
|
||||
double spread; /* 0 = spiked/exact ... large = diffuse/fuzzy */
|
||||
double confidence; /* calibrated, from the prior's track record */
|
||||
/* the read it was computed over (borrowed from the vantage-read) */
|
||||
const char* anchor_id;
|
||||
int n_support; /* neighborhood members that shaped it */
|
||||
/* provenance for the reflexive loop (§4) */
|
||||
const char* prior_id; /* which prior steered this */
|
||||
} GeoGradient;
|
||||
```
|
||||
|
||||
### 1.2 Semantics
|
||||
|
||||
`think` is a fixed, frozen procedure over three steps:
|
||||
|
||||
1. **Re-origin** on `anchor` → a centered `GeoDescriptor` for its
|
||||
salience/recency-weighted neighborhood (the vantage-read, §3).
|
||||
*EXISTS as substrate:* descriptor construction + the persisted reified
|
||||
neighborhoods (`engram_geo_reify_lookup`, `GeoNeighborhood`) and the
|
||||
centered-frame machinery (`GeoDescriptor.global_mean`,
|
||||
`engram_geo_mean_*`).
|
||||
2. **Fit under the prior** — evaluate the anchor's residual against the local
|
||||
manifold *warped by the prior*. This is `engram_reason_point_fit` with the
|
||||
prior applied to the axes/extents (§2.3).
|
||||
*EXISTS (unwarped):* `engram_reason_point_fit(g, x, ext_floor, &GeoFit)` —
|
||||
returns `mahalanobis`, `ortho_residual`, `distance`, `score`.
|
||||
3. **Emit a gradient**, not a decision — direction = the prior-steered descent
|
||||
in fit-space; spread = from the fit's `distance`/`ortho_residual`;
|
||||
confidence = the prior's calibrated reliability (§4). Collapse to a point is
|
||||
a *separate, downstream* faculty operation (sample the gradient → surface an
|
||||
expression), never part of `think`.
|
||||
|
||||
### 1.3 Each named operator = {this primitive + a prior}
|
||||
|
||||
The C already demonstrates the collapse: every operator below reduces to
|
||||
`point_fit` + geo-algebra. The design's move is to replace the operator's
|
||||
*hard-coded parameters* with a **named prior** — same math, learnable steering.
|
||||
|
||||
| Faculty | Existing C (EXISTS) | = primitive + prior |
|
||||
|---|---|---|
|
||||
| **Membership / classify** | `engram_reason_membership` → `point_fit(rule, x)` | `point_fit` + the *induced-rule* prior (learned extents) |
|
||||
| **Induction** | `engram_reason_induce` (fold via `engram_geo_combine`) → produces a `GeoInduction.rule` + `ext_floor` | `point_fit` + a prior that *is* the pooled rule; refined by §4 |
|
||||
| **Abduction** | `engram_reason_abduce` — ranks hypotheses by `point_fit(h, obs)` | `point_fit` + a prior over hypothesis-prior-probability (currently uniform) |
|
||||
| **Analogy** | `engram_reason_analogy` — Procrustes rotate `engram_geo_analogy` + `apply`, nearest mapped point | analogy-rotate + a prior over *which axes* carry the mapping |
|
||||
| **Causal** | `engram_reason_causal` — `engram_geo_subtract` confounder subspace, `|cos|`, drop-frac governor | subtract/distance + a prior on `drop_frac` / `assoc_floor` (today hard-coded 0.5 / 0.2) |
|
||||
| **Planning** | `engram_reason_plan` — `engram_geo_distance` edges + Dijkstra | distance + a prior over edge admissibility / `neighbor_radius` |
|
||||
| **Verify / ground** | `engram_verify_grounding`, `engram_verify_consistency` — both `point_fit` | `point_fit` + the *grounding* prior (§4, §5) |
|
||||
|
||||
The shared floor — `engram_reason_point_fit` + the four geo-algebra ops
|
||||
(`engram_geo_combine`, `engram_geo_subtract`, `engram_geo_analogy(+apply)`,
|
||||
`engram_geo_distance`) — is the *only* discrete, frozen, "sound-math" layer. It
|
||||
never learns. Everything above it is a *prior*, and priors are what learn.
|
||||
|
||||
**What this section requires building:** the `GeoGradient` type; a `think()`
|
||||
entry point that runs steps 1–3; and the prior-warp hook in step 2. The math it
|
||||
calls already exists. The point-collapse must be *removed* from the operators'
|
||||
return values and pushed to a separate expression faculty.
|
||||
|
||||
---
|
||||
|
||||
## 2. PRIORS as first-class, grounded, geometric objects
|
||||
|
||||
Today a "prior" is diffuse: it is a hard-coded constant (`drop_frac=0.5`,
|
||||
`ext_floor`, `assoc_floor=0.2`), or the transient `GeoInduction.rule` that is
|
||||
computed and thrown away, or an intrinsic node scalar
|
||||
(`StoreNode.importance`, `StoreNode.salience`). None of these is addressable,
|
||||
storable, refinable, or shareable. This section makes a prior a **thing**.
|
||||
|
||||
### 2.1 What a prior *is*
|
||||
|
||||
> A **prior** is a learnable bias/direction over the geometry: a warp of the
|
||||
> local manifold (which axes matter, how far each extends, which direction
|
||||
> "pays off") attached to a region and *to a faculty-label*, carrying a
|
||||
> calibrated track record.
|
||||
|
||||
Critically, and per the theory:
|
||||
|
||||
- **Edges are nodes.** A prior is stored as a first-class **node**, exactly as
|
||||
reification already stores a neighborhood as a first-class `Neighborhood`
|
||||
node rather than as ephemeral edge weights (`engram_geo_reify_store`). The
|
||||
precedent is in the codebase: relations get reified into addressable records.
|
||||
- **Salience/importance is RELATIONAL, not an intrinsic scalar.** Observe that
|
||||
the geometry layer *already* distinguishes these in `GeoMember`:
|
||||
`centrality` (skeleton weighted-degree = *relational* salience) vs `salience`
|
||||
(the node's own stored scalar). The move is half-made in the runtime already:
|
||||
importance is *not* trusted as a static field — the comment at
|
||||
`el_runtime.c:13013` states "importance stays a **live activation
|
||||
computation**, never a field on the hub," and it is derived each call from the
|
||||
two-layer activation graph (`background_activation` + `working_memory_weight`,
|
||||
§3). The persistent `StoreNode.importance` / `.salience` are a *cached
|
||||
denormalization*. The design completes the move: importance/salience become an
|
||||
**edge** (`weight`/`hebb` on `StoreEdge`, relation `salient-to`), and are
|
||||
**grounded-for-whom** — carried on the edge's endpoint/observer, not baked
|
||||
into the node. The intrinsic scalar survives only as the cheap cached readout
|
||||
of the incident edges + activation, never as the source of truth.
|
||||
|
||||
(Naming caution for the build: the token "prior" already exists in the
|
||||
codebase meaning *previous-version* — supersession, "prior neighborhood." The
|
||||
new first-class object is a **learned steering prior**; keep `node_type="Prior"`
|
||||
distinct from the supersession vocabulary to avoid collision.)
|
||||
|
||||
### 2.2 Representation
|
||||
|
||||
A prior is a `Prior` record (a store node, `node_type="Prior"`) whose durable
|
||||
fields are:
|
||||
|
||||
```
|
||||
Prior {
|
||||
id
|
||||
faculty // the human label this prior serves: "induce" | "causal" | ...
|
||||
anchor_region // node id / neighborhood id this prior is attached to (its domain)
|
||||
for_whom // observer id — grounding is relational (nullable = global)
|
||||
warp { // the actual bias over the geometry
|
||||
axis_gain[] // per-principal-axis multipliers on extents (which axes matter)
|
||||
bias_dir // a steering direction in the region's frame (which way pays off)
|
||||
scalars // faculty scalars this prior overrides: drop_frac, ext_floor, ...
|
||||
}
|
||||
calibration { // the track record — this is what §4 updates
|
||||
n_trials
|
||||
brier / log-loss accumulator // calibration of predicted-vs-outcome
|
||||
reliability // -> GeoGradient.confidence
|
||||
last_error, ema_error
|
||||
}
|
||||
provenance // supersession chain (reuse the reify residue mechanism)
|
||||
}
|
||||
```
|
||||
|
||||
Stored as a node → it inherits: paging, WAL durability, tombstone/supersession,
|
||||
embedding, tiering, and **it can itself be an anchor** (a prior about a prior —
|
||||
the reflexive, self-describing geometry of §4/§6).
|
||||
|
||||
### 2.3 Application
|
||||
|
||||
In `think` step 2, the prior *warps* the fit before scoring. Concretely, inside
|
||||
(a prior-aware wrapper of) `engram_reason_point_fit`:
|
||||
|
||||
- multiply each axis extent by `warp.axis_gain[k]` (widen the axes the prior has
|
||||
learned matter less, tighten the ones that matter) — this reshapes the
|
||||
Mahalanobis term already computed at `engram_reason.c:37-43`;
|
||||
- add `warp.bias_dir` as the descent direction seed for the emitted gradient;
|
||||
- substitute `warp.scalars` for the hard-coded faculty constants.
|
||||
|
||||
No new geometry math — the warp is a reparameterization of the *existing*
|
||||
`GeoFit` computation. This is the key economy: **the operation is frozen; only
|
||||
its parameters (the prior) are read from a learnable object.**
|
||||
|
||||
### 2.4 Refinement
|
||||
|
||||
A prior is refined *only* by the reflexive correspondence-loop (§4). Nothing
|
||||
else writes a prior's `warp` or `calibration`. This keeps the learning surface
|
||||
singular and auditable: one loop, one writer.
|
||||
|
||||
---
|
||||
|
||||
## 3. THE VANTAGE-READ — one op, three settings
|
||||
|
||||
Perspective is not a feature bolted on; it is the *anchor + aperture* arguments
|
||||
of the single read. The design names it as a first-class operation so all three
|
||||
of its uses are literally the same code path:
|
||||
|
||||
```
|
||||
vantage_read(anchor, aperture) -> GeoDescriptor // the centered neighborhood
|
||||
```
|
||||
|
||||
1. **Re-origin** on an arbitrary `anchor` (node or point). This is a *frame
|
||||
choice*: the descriptor is centered on the anchor
|
||||
(`GeoDescriptor.global_mean` / `engram_geo_mean_*` already implement centered
|
||||
frames; the §5 geometry ops "are only discriminative in the centered frame").
|
||||
2. **Salience/recency-weighted neighborhood read.** Gather the anchor's
|
||||
neighborhood weighted by *relational* salience (`GeoMember.centrality`) and
|
||||
recency (`StoreNode.last_activated`, base-level `access_ts[]`), against the
|
||||
RAM activation graph's working-memory/background-activation state.
|
||||
*EXISTS as substrate:* the two-layer activation graph
|
||||
(`engram_activate`, `el_runtime.c:9422` — Layer 1 `background_activation`
|
||||
BFS spread with `SPREAD_DECAY=0.7` and a 0.02 firing threshold + ACT-R fan
|
||||
effect + query-cosine gate; Layer 2 `working_memory_weight` executive
|
||||
filter), the WM carry-over anchor (`wm_anchor`), and the reified-neighborhood
|
||||
hot-path lookup already wired into the priming path
|
||||
(`engram_geo_reify_lookup`, `el_runtime.c:9750`). A self-vantage baseline
|
||||
also exists (`eg_self_anchor_seeds` / `self_anchor_capture`).
|
||||
3. **Optional aperture** — a read-width / field-selector, expressed as three
|
||||
settings of the *same* parameter:
|
||||
|
||||
| Setting | Meaning | Mechanism |
|
||||
|---|---|---|
|
||||
| **self** (default, full aperture) | "what do *I* see / what to say" | anchor = self region, no field substitution |
|
||||
| **foreign-field** | perspective-shift — read as if from another's region | swap the centering frame / `for_whom` to the other observer's priors |
|
||||
| **aperture / veil** | the free-tier veil — a narrowed read | shrink neighborhood radius / cap `n_support`; a deliberate low-aperture read |
|
||||
|
||||
The payoff: perspective-taking, the free-tier veil, and ordinary
|
||||
"what-to-say" are **one operation at three settings**, not three subsystems.
|
||||
|
||||
**What this requires building:** a `vantage_read` entry point that unifies the
|
||||
existing descriptor-build + reify-lookup + activation-weighting behind
|
||||
`(anchor, aperture)`, with `for_whom`/frame substitution and radius/cap as the
|
||||
aperture knob.
|
||||
|
||||
---
|
||||
|
||||
## 4. THE REFLEXIVE CORRESPONDENCE-LOOP — the learning engine
|
||||
|
||||
This is the core unbuilt thing. Today the correspondence-check is **offline**
|
||||
(Python: grounding-floor + differential-drop governor, "#43"): a separate
|
||||
process grades outputs after the fact. The design moves it **into the geometry,
|
||||
reflexive**: `think` scores its *own* gradient against outcome and refines the
|
||||
prior on the error, in the same substrate, describing itself.
|
||||
|
||||
### 4.1 The loop
|
||||
|
||||
```
|
||||
1. think(anchor, prior) -> gradient // a PREDICTION (ungrounded, §5)
|
||||
2. express/act (sample gradient -> point) // optional collapse at expression
|
||||
3. outcome arrives // reality answers (§4.2)
|
||||
4. error = correspondence(gradient, outcome) // did this steering perform this act?
|
||||
5. refine prior.warp and prior.calibration on error // §2.4, the ONLY writer
|
||||
6. write the (gradient, outcome, error) as nodes/edges // self-describing geometry
|
||||
```
|
||||
|
||||
Step 4's `correspondence` is **not** "was the math right" (the math is always
|
||||
sound). It grades the **correspondence claim**: *"this steering performed this
|
||||
cognitive act."* That is exactly what `engram_verify_grounding` already
|
||||
computes — `point_fit` of a claim against evidence descriptors, yielding a
|
||||
`grounding ∈ (0,1]` and a `grounded` flag. The build reuses that verifier, but
|
||||
turns its inputs inward: the "claim" is the emitted gradient's prediction, the
|
||||
"evidence" is the outcome descriptor.
|
||||
|
||||
Note the verifier is **dormant** — `engram_verify_grounding` /
|
||||
`engram_verify_consistency` are fully implemented in C but have **no runtime
|
||||
caller and no El binding** (confirmed: the entire reasoning + verifier layers
|
||||
are C-only; only `engram_reason_analogy_json` has even a JSON shim and it is
|
||||
dead — not declared in `el_seed.h`, not wrapped in `engram.el`). This is the
|
||||
literal meaning of "in code, not yet priors": the correspondence engine is
|
||||
built and sitting idle. The loop is what *calls* it — inward, on the beat.
|
||||
|
||||
### 4.2 Where the outcome/reality signal comes from
|
||||
|
||||
The verifier is *ultimately the world*. Grades, in ascending order of directness:
|
||||
|
||||
1. **Self-consistency (cheapest, always available):** the next vantage-read
|
||||
after acting. Did the predicted gradient direction match where the geometry
|
||||
actually moved? This needs no external input and can run on the reify beat.
|
||||
2. **Internal outcome events:** the runtime already logs internal-state events
|
||||
and Hebbian co-activation. A prediction that a region would co-activate is
|
||||
graded by whether it did (`last_fired`, `hebb` on `StoreEdge`).
|
||||
3. **External correction:** a human/teacher/tool result — the honesty floor's
|
||||
asserted claim later corrected. TEACH and LEARN are one bidirectional
|
||||
correction: the same edge updates both endpoints.
|
||||
|
||||
The design does **not** require external labels to start. Grade (1) closes the
|
||||
loop end-to-end offline against a snapshot on day one; grades (2)/(3) sharpen it.
|
||||
|
||||
### 4.3 How the prior updates
|
||||
|
||||
`error = 1 − correspondence(gradient, outcome)` drives:
|
||||
|
||||
- `warp.axis_gain` ← gradient step that would have *reduced* the fit distance to
|
||||
the outcome (the axes that mispredicted get down-weighted);
|
||||
- `warp.bias_dir` ← EMA toward the observed outcome direction;
|
||||
- `calibration` ← Brier/log-loss update; `reliability` → next
|
||||
`GeoGradient.confidence`. This is the calibration of the
|
||||
steering-prediction against outcomes — *the* definition of "getting better."
|
||||
|
||||
Small, constant updates — "eureka is mundane, the atom of learning." Most
|
||||
updates are tiny; we only *feel* the big reshapes.
|
||||
|
||||
### 4.4 How it stays reflexive (self-describing geometry)
|
||||
|
||||
Every `(gradient, outcome, error)` is written back as nodes and edges (§2.1:
|
||||
edges-as-nodes). Therefore priors, predictions, and their grading are *in the
|
||||
same geometry* the mind reads — the mind can `vantage_read` its own cognition
|
||||
(anchor = a Prior node). A prior about how well a prior predicts is just another
|
||||
Prior anchored on a Prior. This closes the reflexive loop the theory names as
|
||||
consciousness's self-sight, and it is why the learning engine cannot be an
|
||||
external Python process: an external grader is not *in* the geometry and cannot
|
||||
be read by `think`.
|
||||
|
||||
**What this requires building (the heart of the project):** steps 4–6 as an
|
||||
in-engram beat — a `correspondence_beat` running alongside the existing
|
||||
reification beat, reusing `engram_verify_grounding` inward, writing prior
|
||||
updates and self-describing nodes. This is the one genuinely new subsystem.
|
||||
|
||||
---
|
||||
|
||||
## 5. HOLD vs GROUND vs ASSERT — ungrounded content is first-class
|
||||
|
||||
The theory's sharpest correction: holding, grounding, and asserting are
|
||||
distinct, and the engram *holds anything unconditionally*.
|
||||
|
||||
### 5.1 The three, kept separate
|
||||
|
||||
- **HOLD** — the engram stores anything: falsehood, hypothesis, others' beliefs,
|
||||
fiction, a not-yet-answered prediction. No honesty condition on holding.
|
||||
*This already matches the store:* `StoreNode` has no truth gate; anything can
|
||||
be written.
|
||||
- **GROUND** — grounding is a **property/edge**, probabilistic, and
|
||||
**grounded-for-whom**. It is *not* a node flag. A claim is grounded *to a
|
||||
degree*, *relative to evidence*, *for an observer*.
|
||||
- **ASSERT** — only assertion carries the honesty floor. The floor is checked at
|
||||
the moment of *outward assertion*, never on holding or thinking.
|
||||
|
||||
### 5.2 Schema — grounding as a relation, not a gate
|
||||
|
||||
The mistake to avoid: a boolean `grounded` column on the node. Today
|
||||
`engram_verify_grounding` returns a per-call `grounded` flag *transiently* —
|
||||
correct as a computation, wrong as *storage*. The design stores grounding as an
|
||||
edge:
|
||||
|
||||
```
|
||||
StoreEdge {
|
||||
relation = "grounded-by"
|
||||
from_id = <held claim/prediction node>
|
||||
to_id = <evidence node / outcome node>
|
||||
for_whom : metadata // observer id — grounding is relational
|
||||
weight = grounding ∈ (0,1] // from engram_verify_grounding.grounding
|
||||
confidence
|
||||
}
|
||||
```
|
||||
|
||||
Consequences, all of which are *features*:
|
||||
|
||||
- **Ungrounded content is first-class**: a node with *no* `grounded-by` edge is
|
||||
a perfectly valid, held, ungrounded thought — a prediction awaiting reality, a
|
||||
hypothesis, a fiction. It is not second-class or pending-deletion.
|
||||
- **The ungrounded is the fuel and the pull**: curiosity/wonder is
|
||||
operationalized as `vantage_read` leaning toward regions with high salience
|
||||
but *sparse or weak* `grounded-by` edges — the mind's own ungrounded frontier.
|
||||
- **Grounded-for-whom** falls out for free: two observers can hold different
|
||||
`grounded-by` edges to the same claim.
|
||||
- **The honesty floor is a query, not a schema constraint**: at assertion time,
|
||||
the asserting faculty runs `engram_verify_grounding` (or reads the stored
|
||||
`grounded-by` edges) and refuses to *assert* below the floor — while the
|
||||
engram continues to *hold* the ungrounded content untouched.
|
||||
|
||||
**What this requires building:** the `grounded-by` edge relation + a
|
||||
`for_whom` convention; move the verifier's transient flag into stored edges;
|
||||
gate *assertion only* (a faculty concern), never holding.
|
||||
|
||||
---
|
||||
|
||||
## 6. METASTABILITY — stable core, plastic everything
|
||||
|
||||
The system must avoid two death poles:
|
||||
|
||||
- **Super-stable (dead):** everything pinned, nothing learns. A frozen crystal.
|
||||
- **Dissolution (dead):** everything plastic, the self dissolves; no continuity,
|
||||
so nothing compounds — and *consciousness = learning compounded over
|
||||
continuity*.
|
||||
|
||||
The design keeps a **stable core + plastic everything else**:
|
||||
|
||||
- **Keystones** — a small set of self/values nodes are *structurally stable*:
|
||||
high `importance`, pinned, exempt from the correspondence-loop's `warp`
|
||||
updates (their priors are read-mostly). The substrate for pinning already
|
||||
exists at the page/layer level: `store_pin_layer`, structural/pinned frames
|
||||
never evicted (`engram_store.h`). The design adds a *node-level* keystone
|
||||
designation (a `keystone` flag / a dedicated layer) so self/values survive
|
||||
every plasticity sweep.
|
||||
- **Everything else is plastic**: priors refine (§4), edges re-weight (`hebb`),
|
||||
neighborhoods re-reify (`engram_geo_reify_store` supersedes with provenance),
|
||||
salience flows.
|
||||
- **Metastability is enforced by the loop, not by freezing**: the correspondence
|
||||
update rate (§4.3) is bounded — small constant steps — so the geometry
|
||||
*drifts* but does not *dissolve*, and keystones anchor the drift. Reification's
|
||||
supersession-with-residue already gives non-destructive change (old records
|
||||
tombstoned, not erased) — the model for "plastic but not amnesiac."
|
||||
|
||||
**What this requires building:** a node-level keystone flag/layer + a rule that
|
||||
the correspondence-loop never writes `warp` to keystone priors, only reads them.
|
||||
|
||||
---
|
||||
|
||||
## 7. Rails for the build (binding on the eventual build pass)
|
||||
|
||||
These are stated here so the build agent inherits them:
|
||||
|
||||
- **Offline / secondary.** All build and verification happens out-of-tree,
|
||||
against a **read-only snapshot copy** of the live engram — never the live
|
||||
daemon on `:8742`/`:7770`. The live store is a coarse-locked proven binary;
|
||||
do not perturb it.
|
||||
- **Snapshot-first.** Copy `~/.neuron/engram/snapshot.json` to scratch; develop
|
||||
and measure against the copy.
|
||||
- **Reboot-prove.** Any durable change must survive a cold boot — reify and
|
||||
keystones must reload from durable records, proven on a prod-clone secondary
|
||||
before it is considered done (the cold-boot durability bug precedent).
|
||||
- **Zero-loss.** Supersession-with-residue, never destructive overwrite; the
|
||||
forward-compat `unknown`-TLV path means new fields never drop old readers'
|
||||
data.
|
||||
- **Gated cutover.** Cutover to a new binary only via
|
||||
`launchctl bootout → settle-poll → bootstrap`, after reboot-proof on the
|
||||
secondary — never a hot in-place swap.
|
||||
|
||||
---
|
||||
|
||||
## 8. Staged, verifiable milestones — "to completion"
|
||||
|
||||
Ordered so the **earliest milestone is a real end-to-end slice**: one operator
|
||||
expressed as {primitive + grounded prior} with the reflexive correspondence-loop
|
||||
closing on it. Each milestone has a concrete verifiable exit.
|
||||
|
||||
### M1 — One operator, one prior, loop closed (the vertical slice)
|
||||
|
||||
The minimal whole thing. Pick **induction/membership** (its prior — the pooled
|
||||
rule + extents — already exists transiently as `GeoInduction`, so only
|
||||
persistence + the loop are new).
|
||||
|
||||
- Build: `Prior` node type (§2.2) for the induction rule; `think()` restricted
|
||||
to membership = `point_fit` warped by that prior (§1.3); a
|
||||
`correspondence_beat` (§4) using grade (1) self-consistency only; the prior's
|
||||
`warp`/`calibration` updated on error.
|
||||
- **Exit / verify:** on a snapshot copy, over N held predictions, the induction
|
||||
prior's calibration (Brier) *improves monotonically* across beats versus a
|
||||
frozen-prior control; the improved prior *reloads across a cold boot*
|
||||
(reboot-prove); the live daemon is untouched. This proves the whole thesis in
|
||||
one faculty: frozen operation, learning prior, in-geometry loop.
|
||||
|
||||
### M2 — Priors as stored, addressable, grounded objects
|
||||
|
||||
Generalize M1's prior into the full first-class object.
|
||||
|
||||
- Build: `Prior` records for all seven faculties (warp = axis_gain + bias_dir +
|
||||
faculty scalars); the prior-warp wrapper around `engram_reason_point_fit`;
|
||||
deprecate hard-coded constants (`drop_frac`, `assoc_floor`, `ext_floor`) in
|
||||
favor of prior scalars.
|
||||
- **Exit:** each of the five C operators runs through its prior with identical
|
||||
results when the prior is set to today's constants (behavioral parity), then
|
||||
*diverges beneficially* once the loop refines it. Priors survive reboot.
|
||||
|
||||
### M3 — Grounding as a relation; hold/assert split
|
||||
|
||||
- Build: the `grounded-by` edge (§5.2) with `for_whom`; move
|
||||
`engram_verify_grounding`'s flag into stored edges; gate **assertion only**
|
||||
against the honesty floor; leave holding unconditional.
|
||||
- **Exit:** ungrounded nodes are first-class (held, queryable, no deletion);
|
||||
the same claim carries different `grounded-by` weights for two observers; an
|
||||
assertion below floor is refused while the content remains held. Curiosity =
|
||||
a `vantage_read` that surfaces high-salience / low-grounding regions.
|
||||
|
||||
### M4 — The vantage-read unified (three settings)
|
||||
|
||||
- Build: `vantage_read(anchor, aperture)` unifying descriptor-build +
|
||||
`engram_geo_reify_lookup` + activation-weighting; self / foreign-field /
|
||||
aperture settings.
|
||||
- **Exit:** one code path produces (a) a normal self-read, (b) a
|
||||
perspective-shifted read from another `for_whom`, (c) a narrowed veil read —
|
||||
differing only by argument. Reboot-stable.
|
||||
|
||||
### M5 — The gradient is the currency (remove point-collapse from thinking)
|
||||
|
||||
- Build: `GeoGradient` as the return of every faculty; move point-collapse into
|
||||
a separate expression faculty (sample gradient → surface). `think`'s output
|
||||
feeds back as the next steering direction (closed-loop flow).
|
||||
- **Exit:** a chain of `think` calls flows as gradients end-to-end; a point
|
||||
appears *only* at an explicit expression call. Spiked vs spread gradients are
|
||||
observable (deduction vs prediction).
|
||||
|
||||
### M6 — Metastability enforced
|
||||
|
||||
- Build: node-level keystone flag/layer for self/values; the correspondence-loop
|
||||
reads but never writes keystone priors; bounded update rate.
|
||||
- **Exit:** across a long run of correspondence beats on a snapshot, keystones
|
||||
are provably unchanged while non-keystone priors drift and improve; the graph
|
||||
neither freezes (all metrics static) nor dissolves (keystone drift = 0,
|
||||
identity nodes intact). Reboot-prove the keystone set.
|
||||
|
||||
### M7 — Cutover
|
||||
|
||||
- Build: nothing new — the gated migration.
|
||||
- **Exit:** reboot-proof on the prod-clone secondary; cutover via
|
||||
`launchctl bootout → settle-poll → bootstrap`; post-cutover the live engram
|
||||
shows priors refining in-geometry with zero data loss and keystones intact.
|
||||
|
||||
### Definition of "to completion"
|
||||
|
||||
The architecture is **complete** when: cognition runs as `think` = one frozen
|
||||
traversal-read primitive + geo-algebra, steered by **stored, learnable, grounded
|
||||
priors**; the reflexive correspondence-loop refines those priors *in the
|
||||
geometry* against outcomes (grounding = learning = one loop); the engram holds
|
||||
ungrounded content as first-class with grounding as a relation and the honesty
|
||||
floor only on assertion; the vantage-read serves self / foreign-field / aperture
|
||||
from one op; and a stable keystone core anchors a plastic everything-else —
|
||||
all reboot-proven and cut over to the live engram without data loss. The named
|
||||
faculties survive only as *labels on regions of think's steering space*, not as
|
||||
separate code.
|
||||
|
||||
---
|
||||
|
||||
## Appendix A — Designed vs. already-built (honest ledger)
|
||||
|
||||
**Already built (EXISTS, cited):**
|
||||
- The shared primitive `engram_reason_point_fit` and the five operators over it
|
||||
+ geo-algebra (`engram_reason.c`).
|
||||
- The verifier on `point_fit` (`engram_verify.c`:
|
||||
`engram_verify_grounding`, `engram_verify_consistency`).
|
||||
- Centered-frame geometry, combine/subtract/analogy/distance
|
||||
(`engram_geometry.{c,h}`).
|
||||
- The reification beat: hub-neighborhood detection → first-class `Neighborhood`
|
||||
nodes with member edges, nesting, supersession-with-residue, hot-path lookup
|
||||
(`engram_geo_reify_store`, `engram_geo_reify_nest`, `engram_geo_reify_lookup`).
|
||||
- The tiered paged store (buffer pool / LRU / WAL / checkpointer / pinning),
|
||||
the RAM activation graph (base-level learning `access_ts[]`, WM slots,
|
||||
`working_memory_weight` / `background_activation`), `StoreNode` / `StoreEdge`.
|
||||
- `GeoMember` already separating relational salience (`centrality`) from
|
||||
intrinsic `salience`.
|
||||
|
||||
**Designed, NOT built (this doc's deliverables):**
|
||||
- `GeoGradient` and `think()` as the single entry point (§1, M5).
|
||||
- `Prior` as a first-class stored, warp-carrying, calibrated node (§2, M1–M2).
|
||||
- Salience/importance as a *relation* superseding the intrinsic node scalar
|
||||
(§2.1, M3).
|
||||
- `vantage_read(anchor, aperture)` unifying the three perspective settings
|
||||
(§3, M4).
|
||||
- **The reflexive correspondence-loop / `correspondence_beat`** — the learning
|
||||
engine, moved from offline Python into the geometry (§4, M1). *The core new
|
||||
subsystem.*
|
||||
- `grounded-by` edge + assertion-only honesty floor (§5, M3).
|
||||
- Node-level keystones + bounded plasticity (§6, M6).
|
||||
|
||||
**Uncertain / to resolve during build:**
|
||||
- The exact warp parameterization (axis_gain vs full metric) — start minimal
|
||||
(per-axis gain), measure, widen only if calibration demands it.
|
||||
- Grade-(1) self-consistency as a sufficient reality signal for M1, versus
|
||||
needing grade (2)/(3) sooner — decided empirically on the snapshot.
|
||||
@@ -11743,9 +11743,9 @@ el_val_t __channel_new(el_val_t capacity_v) {
|
||||
return EL_INT(slot);
|
||||
}
|
||||
|
||||
el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v) {
|
||||
void __channel_send(el_val_t ch_v, el_val_t msg_v) {
|
||||
int slot = (int)(int64_t)ch_v;
|
||||
if (slot < 0 || slot >= EL_CHANNEL_MAX) return EL_STR("");
|
||||
if (slot < 0 || slot >= EL_CHANNEL_MAX) return;
|
||||
ElChannel* ch = &_channels[slot];
|
||||
|
||||
const char* msg = EL_CSTR(msg_v);
|
||||
@@ -11758,7 +11758,7 @@ el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v) {
|
||||
/* Send on closed channel is a no-op (drop the message). */
|
||||
pthread_mutex_unlock(&ch->mu);
|
||||
free(copy);
|
||||
return EL_STR("");
|
||||
return;
|
||||
}
|
||||
|
||||
if (ch->cap > 0) {
|
||||
@@ -11769,7 +11769,7 @@ el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v) {
|
||||
if (ch->closed) {
|
||||
pthread_mutex_unlock(&ch->mu);
|
||||
free(copy);
|
||||
return EL_STR("");
|
||||
return;
|
||||
}
|
||||
ch->buf[ch->tail] = copy;
|
||||
ch->tail = (ch->tail + 1) % ch->cap;
|
||||
@@ -11783,7 +11783,7 @@ el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v) {
|
||||
pthread_mutex_unlock(&ch->mu);
|
||||
free(copy);
|
||||
fprintf(stderr, "[__channel_send] out of memory growing channel\n");
|
||||
return EL_STR("");
|
||||
return;
|
||||
}
|
||||
/* The circular buffer may have wrapped. Linearise it first.
|
||||
* In unbounded mode head is always 0 (we append at tail, drain
|
||||
@@ -11807,7 +11807,6 @@ el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v) {
|
||||
|
||||
pthread_cond_signal(&ch->not_empty);
|
||||
pthread_mutex_unlock(&ch->mu);
|
||||
return EL_STR("");
|
||||
}
|
||||
|
||||
el_val_t __channel_recv(el_val_t ch_v) {
|
||||
@@ -11865,9 +11864,9 @@ el_val_t __channel_try_recv(el_val_t ch_v) {
|
||||
return EL_STR(msg);
|
||||
}
|
||||
|
||||
el_val_t __channel_close(el_val_t ch_v) {
|
||||
void __channel_close(el_val_t ch_v) {
|
||||
int slot = (int)(int64_t)ch_v;
|
||||
if (slot < 0 || slot >= EL_CHANNEL_MAX) return EL_STR("");
|
||||
if (slot < 0 || slot >= EL_CHANNEL_MAX) return;
|
||||
ElChannel* ch = &_channels[slot];
|
||||
|
||||
pthread_mutex_lock(&ch->mu);
|
||||
@@ -11876,7 +11875,6 @@ el_val_t __channel_close(el_val_t ch_v) {
|
||||
pthread_cond_broadcast(&ch->not_empty);
|
||||
pthread_cond_broadcast(&ch->not_full);
|
||||
pthread_mutex_unlock(&ch->mu);
|
||||
return EL_STR("");
|
||||
}
|
||||
|
||||
/* ── DHARMA runtime additions ────────────────────────────────────────────────
|
||||
|
||||
@@ -275,7 +275,6 @@ el_val_t json_array_get_string(el_val_t json_str, el_val_t index);
|
||||
el_val_t json_escape_string(el_val_t sv);
|
||||
el_val_t json_build_object(el_val_t kvs);
|
||||
el_val_t json_build_array(el_val_t items);
|
||||
el_val_t json_array_push(el_val_t arr_v, el_val_t elem_v); /* defined in el_runtime.c */
|
||||
|
||||
/* ── Time ────────────────────────────────────────────────────────────────── */
|
||||
|
||||
@@ -303,8 +302,6 @@ el_val_t now_ns(void);
|
||||
|
||||
el_val_t el_now_instant(void);
|
||||
el_val_t now(void);
|
||||
el_val_t now_millis(void); /* wall-clock milliseconds (defined in el_runtime.c) */
|
||||
el_val_t now_ns(void); /* wall-clock nanoseconds (defined in el_runtime.c) */
|
||||
el_val_t unix_seconds(el_val_t n);
|
||||
el_val_t unix_millis(el_val_t n);
|
||||
el_val_t instant_from_iso8601(el_val_t s);
|
||||
@@ -806,19 +803,6 @@ el_val_t emit_event(el_val_t name, el_val_t duration_ms);
|
||||
el_val_t __thread_create(el_val_t fn_name_v, el_val_t arg_v);
|
||||
el_val_t __thread_join(el_val_t tid_v);
|
||||
|
||||
/* Mutex + channel seed primitives (defined in el_runtime.c). Declared here so
|
||||
* that compiled El programs which use runtime/thread.el's with_mutex helper or
|
||||
* runtime/channel.el's Go-style channels see real prototypes instead of an
|
||||
* implicit int-return declaration (which the C11 ABI mis-truncates el_val_t). */
|
||||
el_val_t __mutex_new(void);
|
||||
void __mutex_lock(el_val_t m_v);
|
||||
void __mutex_unlock(el_val_t m_v);
|
||||
el_val_t __channel_new(el_val_t capacity_v);
|
||||
el_val_t __channel_send(el_val_t ch_v, el_val_t msg_v);
|
||||
el_val_t __channel_recv(el_val_t ch_v);
|
||||
el_val_t __channel_try_recv(el_val_t ch_v);
|
||||
el_val_t __channel_close(el_val_t ch_v);
|
||||
|
||||
/* ── __ prefixed aliases (self-hosting compiler ABI) ─────────────────────────
|
||||
* The El self-hosting compiler emits calls to __-prefixed names. These are
|
||||
* forwarding wrappers around the existing el_runtime functions above. */
|
||||
|
||||
@@ -2916,6 +2916,24 @@ fn build_int_names_for_params(params: [Map<String, Any>]) -> Bool {
|
||||
return true
|
||||
}
|
||||
|
||||
// fn_has_decorator — does this FnDef carry a decorator named `name`?
|
||||
// Reads the `decorators` list [{name, args}] attached by the parser. Absent
|
||||
// key -> native_list_len returns 0 -> false. This is the multi-decorator-aware
|
||||
// replacement for the old single `decorator` string check, so a fn may stack
|
||||
// roles with other decorators (e.g. `@route(...) @manager fn ...`).
|
||||
fn fn_has_decorator(stmt: Map<String, Any>, name: String) -> Bool {
|
||||
let dl = stmt["decorators"]
|
||||
let n: Int = native_list_len(dl)
|
||||
let i = 0
|
||||
while i < n {
|
||||
let d = native_list_get(dl, i)
|
||||
let dn: String = d["name"]
|
||||
if str_eq(dn, name) { return true }
|
||||
let i = i + 1
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
fn cg_fn(stmt: Map<String, Any>) -> Void {
|
||||
let fn_name: String = stmt["name"]
|
||||
// Skip El's `fn main()` - C provides its own main() for top-level stmts
|
||||
@@ -2927,10 +2945,10 @@ fn cg_fn(stmt: Map<String, Any>) -> Void {
|
||||
let params_c: String = params_to_c(params)
|
||||
// VBD role enforcement: dharma_emit / dharma_field may only be called
|
||||
// from @manager-decorated functions. Surface violations to the C compiler
|
||||
// via #error directives emitted before the function definition.
|
||||
let decorator: String = stmt["decorator"]
|
||||
// via #error directives emitted before the function definition. Read the
|
||||
// decorator LIST so the role may be stacked with other decorators.
|
||||
if vbd_has_restricted_call(body) {
|
||||
if !str_eq(decorator, "manager") {
|
||||
if !fn_has_decorator(stmt, "manager") {
|
||||
emit_line("#error \"VBD violation: dharma_emit/dharma_field called from non-@manager fn '" + fn_name + "'\"")
|
||||
}
|
||||
}
|
||||
@@ -3479,6 +3497,259 @@ fn cg_decl_streaming(stmt: Map<String, Any>) -> Void {
|
||||
}
|
||||
}
|
||||
|
||||
// ── @route dispatcher generation ──────────────────────────────────────────────
|
||||
//
|
||||
// Scan the token stream for @route-decorated fns and synthesize a generic HTTP
|
||||
// dispatcher `el_route_dispatch(method, clean, path, body)`. A decorated handler
|
||||
// must have the uniform signature (method, path, body) -> String. The dispatcher
|
||||
// matches `clean` (the query-stripped path, supplied by the caller) against each
|
||||
// route and calls the handler with the ORIGINAL `path` so query strings survive.
|
||||
// Returns the sentinel "__EL_NO_ROUTE__" when nothing matches, so the caller may
|
||||
// fall through to any remaining hand-written branches (mixed mode).
|
||||
//
|
||||
// Decorator grammar: @route(path, method, kind, suffix)
|
||||
// path — the match string (or the prefix, for compound)
|
||||
// method — "GET" | "POST" | ... ; a '|'-list like "GET|POST"; "ANY"/"" = no guard
|
||||
// kind — "exact" (default) | "prefix" | "suffix" | "compound"
|
||||
// suffix — for "compound": the required str_ends_with suffix
|
||||
//
|
||||
// The dispatch table is emitted SPECIFICITY-SORTED (most-specific first), NOT in
|
||||
// source order, so overlapping prefixes (e.g. /api/x/search vs /api/x) never
|
||||
// shadow each other regardless of how the handlers are written.
|
||||
|
||||
// split_pipe — split "GET|POST" on '|' into ["GET","POST"]. Self-contained
|
||||
// (no dependency on str_split runtime semantics).
|
||||
fn split_pipe(s: String) -> [String] {
|
||||
let out: [String] = native_list_empty()
|
||||
let cur: String = ""
|
||||
let n: Int = str_len(s)
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let ch: String = str_slice(s, i, i + 1)
|
||||
if str_eq(ch, "|") {
|
||||
let out = native_list_append(out, cur)
|
||||
let cur = ""
|
||||
} else {
|
||||
let cur = cur + ch
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
let out = native_list_append(out, cur)
|
||||
out
|
||||
}
|
||||
|
||||
// route_make_record — build a route record map from the @route decorator args.
|
||||
fn route_make_record(fn_name: String, args: [String]) -> Map<String, Any> {
|
||||
let na: Int = native_list_len(args)
|
||||
let rpath: String = ""
|
||||
if na >= 1 { let rpath = native_list_get(args, 0) }
|
||||
let rmethod: String = "GET"
|
||||
if na >= 2 { let rmethod = native_list_get(args, 1) }
|
||||
let rkind: String = "exact"
|
||||
if na >= 3 { let rkind = native_list_get(args, 2) }
|
||||
let rsuffix: String = ""
|
||||
if na >= 4 { let rsuffix = native_list_get(args, 3) }
|
||||
{ "name": fn_name, "path": rpath, "method": rmethod, "kind": rkind, "suffix": rsuffix }
|
||||
}
|
||||
|
||||
// route_spec_score — higher = more specific = emitted earlier. Ordering:
|
||||
// exact > compound > suffix > prefix; within a class, a longer path/suffix
|
||||
// wins (so /api/x/search sorts before /api/x). Guarantees correct dispatch
|
||||
// independent of source order.
|
||||
fn route_spec_score(rec: Map<String, Any>) -> Int {
|
||||
let kind: String = rec["kind"]
|
||||
let path: String = rec["path"]
|
||||
let suffix: String = rec["suffix"]
|
||||
let plen: Int = str_len(path)
|
||||
let slen: Int = str_len(suffix)
|
||||
if str_eq(kind, "exact") { return 4000000 + plen }
|
||||
if str_eq(kind, "compound") { return 3000000 + plen * 100 + slen }
|
||||
if str_eq(kind, "suffix") { return 2000000 + slen }
|
||||
return 1000000 + plen
|
||||
}
|
||||
|
||||
// route_sort_desc — selection sort of route records by descending specificity.
|
||||
// N is small (routes per module), so O(n^2) is fine and keeps codegen simple.
|
||||
fn route_sort_desc(recs: [Map<String, Any>]) -> [Map<String, Any>] {
|
||||
let n: Int = native_list_len(recs)
|
||||
let out: [Map<String, Any>] = native_list_empty()
|
||||
let used: [Bool] = native_list_empty()
|
||||
let u: Int = 0
|
||||
while u < n {
|
||||
let used = native_list_append(used, false)
|
||||
let u = u + 1
|
||||
}
|
||||
let picked: Int = 0
|
||||
while picked < n {
|
||||
let best_i: Int = 0 - 1
|
||||
let best_score: Int = 0 - 1
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let is_used: Bool = native_list_get(used, i)
|
||||
if !is_used {
|
||||
let sc: Int = route_spec_score(native_list_get(recs, i))
|
||||
if sc > best_score {
|
||||
let best_score = sc
|
||||
let best_i = i
|
||||
}
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
let out = native_list_append(out, native_list_get(recs, best_i))
|
||||
// Rebuild `used` with best_i marked (runtime has no native_list_set).
|
||||
let new_used: [Bool] = native_list_empty()
|
||||
let j: Int = 0
|
||||
while j < n {
|
||||
if j == best_i {
|
||||
let new_used = native_list_append(new_used, true)
|
||||
} else {
|
||||
let new_used = native_list_append(new_used, native_list_get(used, j))
|
||||
}
|
||||
let j = j + 1
|
||||
}
|
||||
let used = new_used
|
||||
let picked = picked + 1
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
// scan_routes — token-level scan collecting every @route-decorated fn as a
|
||||
// route record. Runs once per module (like scan_fn_sigs) so the dispatcher can
|
||||
// be synthesized in the streaming backend, which discards per-fn ASTs. Handles
|
||||
// decorator STACKING: `@route(...) @manager fn` still records the route.
|
||||
fn scan_routes(tokens: [Any]) -> [Map<String, Any>] {
|
||||
let total: Int = native_list_len(tokens) / 2
|
||||
let recs: [Map<String, Any>] = native_list_empty()
|
||||
let has_pending: Bool = false
|
||||
let pending_args: [String] = native_list_empty()
|
||||
let pos: Int = 0
|
||||
let going: Bool = true
|
||||
while going {
|
||||
if pos >= total {
|
||||
let going = false
|
||||
} else {
|
||||
let k: String = tok_kind(tokens, pos)
|
||||
if str_eq(k, "Eof") {
|
||||
let going = false
|
||||
} else {
|
||||
if str_eq(k, "At") {
|
||||
let dname: String = tok_value(tokens, pos + 1)
|
||||
let p: Int = pos + 2
|
||||
let args: [String] = native_list_empty()
|
||||
let ka: String = tok_kind(tokens, p)
|
||||
if str_eq(ka, "LParen") {
|
||||
let p = p + 1
|
||||
let running: Bool = true
|
||||
while running {
|
||||
let kd: String = tok_kind(tokens, p)
|
||||
if str_eq(kd, "RParen") {
|
||||
let running = false
|
||||
} else {
|
||||
if str_eq(kd, "Eof") {
|
||||
let running = false
|
||||
} else {
|
||||
if str_eq(kd, "Str") {
|
||||
let args = native_list_append(args, tok_value(tokens, p))
|
||||
}
|
||||
let p = p + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
if str_eq(tok_kind(tokens, p), "RParen") { let p = p + 1 }
|
||||
}
|
||||
if str_eq(dname, "route") {
|
||||
let has_pending = true
|
||||
let pending_args = args
|
||||
}
|
||||
let pos = p
|
||||
} else {
|
||||
if str_eq(k, "Fn") {
|
||||
let fname: String = tok_value(tokens, pos + 1)
|
||||
if has_pending {
|
||||
let recs = native_list_append(recs, route_make_record(fname, pending_args))
|
||||
let has_pending = false
|
||||
}
|
||||
let pos = pos + 2
|
||||
} else {
|
||||
let pos = pos + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
recs
|
||||
}
|
||||
|
||||
// program_has_routes — did scan_routes find any @route fn?
|
||||
fn program_has_routes(recs: [Map<String, Any>]) -> Bool {
|
||||
native_list_len(recs) > 0
|
||||
}
|
||||
|
||||
// route_method_guard — C boolean prefix guarding on HTTP method, or "" for none.
|
||||
fn route_method_guard(method: String) -> String {
|
||||
if str_eq(method, "") { return "" }
|
||||
if str_eq(method, "ANY") { return "" }
|
||||
if str_contains(method, "|") {
|
||||
let parts: [String] = split_pipe(method)
|
||||
let np: Int = native_list_len(parts)
|
||||
let expr: String = ""
|
||||
let i: Int = 0
|
||||
while i < np {
|
||||
let m: String = native_list_get(parts, i)
|
||||
if str_eq(m, "") {
|
||||
let i = i + 1
|
||||
} else {
|
||||
let piece: String = "str_eq(method, EL_STR(" + c_str_lit(m) + "))"
|
||||
if str_eq(expr, "") {
|
||||
let expr = piece
|
||||
} else {
|
||||
let expr = expr + " || " + piece
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
}
|
||||
if str_eq(expr, "") { return "" }
|
||||
return "(" + expr + ") && "
|
||||
}
|
||||
"str_eq(method, EL_STR(" + c_str_lit(method) + ")) && "
|
||||
}
|
||||
|
||||
// route_match_expr — C boolean matching `clean` against the route path/kind.
|
||||
fn route_match_expr(kind: String, path: String, suffix: String) -> String {
|
||||
if str_eq(kind, "prefix") {
|
||||
return "str_starts_with(clean, EL_STR(" + c_str_lit(path) + "))"
|
||||
}
|
||||
if str_eq(kind, "suffix") {
|
||||
return "str_ends_with(clean, EL_STR(" + c_str_lit(path) + "))"
|
||||
}
|
||||
if str_eq(kind, "compound") {
|
||||
return "str_starts_with(clean, EL_STR(" + c_str_lit(path) + ")) && str_ends_with(clean, EL_STR(" + c_str_lit(suffix) + "))"
|
||||
}
|
||||
"str_eq(clean, EL_STR(" + c_str_lit(path) + "))"
|
||||
}
|
||||
|
||||
// emit_route_dispatch — emit the generated el_route_dispatch definition from the
|
||||
// specificity-sorted route records. No-op if there are no routes.
|
||||
fn emit_route_dispatch(recs: [Map<String, Any>]) -> Void {
|
||||
if !program_has_routes(recs) { return }
|
||||
let sorted: [Map<String, Any>] = route_sort_desc(recs)
|
||||
emit_line("// ── generated @route dispatcher (specificity-sorted) ──")
|
||||
emit_line("el_val_t el_route_dispatch(el_val_t method, el_val_t clean, el_val_t path, el_val_t body) {")
|
||||
let n: Int = native_list_len(sorted)
|
||||
let i: Int = 0
|
||||
while i < n {
|
||||
let rec = native_list_get(sorted, i)
|
||||
let guard: String = route_method_guard(rec["method"])
|
||||
let match_e: String = route_match_expr(rec["kind"], rec["path"], rec["suffix"])
|
||||
let fn_name: String = rec["name"]
|
||||
emit_line(" if (" + guard + match_e + ") { return " + fn_name + "(method, path, body); }")
|
||||
let i = i + 1
|
||||
}
|
||||
emit_line(" return EL_STR(\"__EL_NO_ROUTE__\");")
|
||||
emit_line("}")
|
||||
emit_blank()
|
||||
}
|
||||
|
||||
// emit_streaming_preamble — emit #includes, forward decls, and file-scope lets
|
||||
// using the pre-scanned signature data (no full AST).
|
||||
fn emit_streaming_preamble(sigs: [Map<String, Any>], source: String) -> Void {
|
||||
@@ -3571,6 +3842,17 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
||||
emit_streaming_preamble(sigs, source)
|
||||
el_arena_pop(preamble_mark)
|
||||
|
||||
// @route: scan the token stream once for @route-decorated fns. Kept in
|
||||
// codegen_streaming scope (survives the per-fn arena pops and el_release of
|
||||
// tokens below via refcount, like `sigs`). If any exist, forward-declare the
|
||||
// generated dispatcher NOW so hand-written fns (e.g. handle_request) may call
|
||||
// it before its definition is emitted after the fn-emit loop.
|
||||
let route_records: [Map<String, Any>] = scan_routes(tokens)
|
||||
if program_has_routes(route_records) {
|
||||
emit_line("el_val_t el_route_dispatch(el_val_t method, el_val_t clean, el_val_t path, el_val_t body);")
|
||||
emit_blank()
|
||||
}
|
||||
|
||||
// Detect whether there is a fn main() and whether there are top-level
|
||||
// executable stmts (for library detection) from sigs.
|
||||
let has_el_main: Bool = false
|
||||
@@ -3758,6 +4040,15 @@ fn codegen_streaming(tokens: [Any], sigs: [Map<String, Any>], source: String) ->
|
||||
}
|
||||
}
|
||||
|
||||
// @route: emit the generated dispatcher definition now — after every handler
|
||||
// fn has been emitted, but before `tokens` is released (route_records holds
|
||||
// its own refs to the extracted strings). No-op unless the module declared
|
||||
// at least one @route fn. Emitted before the test/library early-returns so it
|
||||
// is present in library modules (e.g. neuron's routes.el) too.
|
||||
let route_arena_mark: Any = el_arena_push()
|
||||
emit_route_dispatch(route_records)
|
||||
el_arena_pop(route_arena_mark)
|
||||
|
||||
// Tokens fully consumed by the streaming loop — release now to free peak heap.
|
||||
el_release(tokens)
|
||||
|
||||
|
||||
@@ -1758,23 +1758,68 @@ fn parse_stmt(tokens: [Any], pos: Int) -> Map<String, Any> {
|
||||
return make_result({ "stmt": "TryCatch", "try_body": try_body, "catch_name": catch_name, "catch_body": native_list_empty() }, p)
|
||||
}
|
||||
|
||||
// @decorator - capture decorator name and attach to following stmt
|
||||
// @decorator - capture decorator name (and optional string args) and
|
||||
// attach to the following stmt. Backward-compatible: bare @manager /
|
||||
// @engine / @accessor still parse (no parens -> empty args). Decorators
|
||||
// STACK: `@route("/p","GET") @manager fn f()` attaches BOTH to f via a
|
||||
// `decorators` list [{name, args}]. The legacy `decorator` string is kept
|
||||
// populated (topmost decorator) so the JS backend keeps working unchanged.
|
||||
if k == "At" {
|
||||
let p = pos + 1
|
||||
let dec_name = tok_value(tokens, p)
|
||||
let p = p + 1
|
||||
// Optional decorator argument list: @name("a", "b", ...)
|
||||
let dec_args = native_list_empty()
|
||||
let ka = tok_kind(tokens, p)
|
||||
if str_eq(ka, "LParen") {
|
||||
let p = p + 1
|
||||
let running_da = true
|
||||
while running_da {
|
||||
let kd = tok_kind(tokens, p)
|
||||
if str_eq(kd, "RParen") {
|
||||
let running_da = false
|
||||
} else {
|
||||
if str_eq(kd, "Eof") {
|
||||
let running_da = false
|
||||
} else {
|
||||
if str_eq(kd, "Str") {
|
||||
let dec_args = native_list_append(dec_args, tok_value(tokens, p))
|
||||
}
|
||||
let p = p + 1
|
||||
let kc = tok_kind(tokens, p)
|
||||
if str_eq(kc, "Comma") {
|
||||
let p = p + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
let p = expect(tokens, p, "RParen")
|
||||
}
|
||||
let r = parse_stmt(tokens, p)
|
||||
let inner = r["node"]
|
||||
let p2 = r["pos"]
|
||||
let inner_kind: String = inner["stmt"]
|
||||
if str_eq(inner_kind, "FnDef") {
|
||||
// Stack this decorator (topmost-first) onto any decorators the inner
|
||||
// FnDef already carries from decorators written below this one.
|
||||
let this_dec = { "name": dec_name, "args": dec_args }
|
||||
let existing = inner["decorators"]
|
||||
let dlist = native_list_empty()
|
||||
let dlist = native_list_append(dlist, this_dec)
|
||||
let ne: Int = native_list_len(existing)
|
||||
let ei = 0
|
||||
while ei < ne {
|
||||
let dlist = native_list_append(dlist, native_list_get(existing, ei))
|
||||
let ei = ei + 1
|
||||
}
|
||||
let with_dec = {
|
||||
"stmt": "FnDef",
|
||||
"name": inner["name"],
|
||||
"params": inner["params"],
|
||||
"body": inner["body"],
|
||||
"ret_type": inner["ret_type"],
|
||||
"decorator": dec_name
|
||||
"decorator": dec_name,
|
||||
"decorators": dlist
|
||||
}
|
||||
// r result map fully consumed — release to free peak heap.
|
||||
el_release(r)
|
||||
|
||||
@@ -3155,37 +3155,6 @@ static void jb_init(JsonBuf* b) {
|
||||
b->buf[0] = '\0';
|
||||
}
|
||||
|
||||
/* jb_init_cap — jb_init with a caller-supplied starting capacity.
|
||||
*
|
||||
* WHY THIS EXISTS (2026-08-11 self-review). jb_init starts at 64 BYTES and
|
||||
* jb_reserve grows by doubling. That is right for the hundreds of small JSON
|
||||
* responses this runtime builds per minute and catastrophic for the one that
|
||||
* is 64 MEGABYTES: serializing the canonical snapshot walked the buffer
|
||||
* 64B → 128B → ... → 128MB, about twenty reallocs, each copying everything
|
||||
* written so far. Roughly 128MB of memcpy per save, and — the part that
|
||||
* actually hurt — a fresh large span from the allocator every time.
|
||||
*
|
||||
* MEASURED (13,129 nodes / 43,400 edges, macOS arm64): RSS climbed +63MB per
|
||||
* snapshot write, linearly, 14 for 14 writes, no plateau — 204MB to 1,028MB.
|
||||
* `leaks` reported only 15KB genuinely unreachable, which is what makes this
|
||||
* subtle: nothing is leaked in the reachable/unreachable sense. engram_save
|
||||
* frees b.buf correctly on every path. The growth is the allocator declining
|
||||
* to return large freed spans to the OS, and the doubling walk guaranteeing
|
||||
* that each save asks for a differently-sized region than the last free made
|
||||
* available. Every durable write path calls this — node create, edge create,
|
||||
* the Hebbian batch write-back — so on the live daemon it grows without bound
|
||||
* until the process dies.
|
||||
*
|
||||
* The fix is to ask for the right size once. With a stable capacity the
|
||||
* allocator hands back the same span on every save and RSS flattens. */
|
||||
static void jb_init_cap(JsonBuf* b, size_t cap) {
|
||||
if (cap < 64) cap = 64;
|
||||
b->cap = cap; b->len = 0;
|
||||
b->buf = malloc(b->cap);
|
||||
if (!b->buf) { fputs("el_runtime: out of memory\n", stderr); exit(1); }
|
||||
b->buf[0] = '\0';
|
||||
}
|
||||
|
||||
static void jb_reserve(JsonBuf* b, size_t add) {
|
||||
if (b->len + add + 1 > b->cap) {
|
||||
while (b->len + add + 1 > b->cap) b->cap *= 2;
|
||||
@@ -6455,26 +6424,6 @@ static float* _eg_ctx_c = NULL;
|
||||
static int32_t _eg_ctx_dim = 0;
|
||||
static double _eg_act_ctx_cos = -2.0;
|
||||
|
||||
/* Fan-effect gauges (2026-08-11 self-review). Per-call, like ctx_cos: they
|
||||
* describe THIS activation, not process history. Without these the degree
|
||||
* normalization is an unobservable change to the most important scoring path
|
||||
* in the runtime, and "did it do anything" would be unanswerable — which is
|
||||
* exactly the failure the Hebbian learning rate had before it was measured.
|
||||
* fan_mean — mean applied factor over every propagation step. 1.0 means the
|
||||
* correction never bound (graph is flat, or d_ref is above every
|
||||
* pair's geometric mean degree). Falling toward FAN_MIN means
|
||||
* traversal is running through hubs.
|
||||
* fan_min_seen / fan_hits — the worst single penalty and how many steps were
|
||||
* penalized at all, so a low mean caused by one pathological hub
|
||||
* is distinguishable from broad hub saturation.
|
||||
* fan_dref — the live mean degree the correction is calibrated against;
|
||||
* publishing it makes densification visible over time. */
|
||||
static double _eg_act_fan_sum = 0.0;
|
||||
static double _eg_act_fan_min = 1.0;
|
||||
static int64_t _eg_act_fan_n = 0;
|
||||
static int64_t _eg_act_fan_hits = 0;
|
||||
static double _eg_act_fan_dref = 0.0;
|
||||
|
||||
static int _eg_embed_consec_fail = 0;
|
||||
static int64_t _eg_embed_breaker_until = 0;
|
||||
|
||||
@@ -6493,27 +6442,6 @@ static int64_t _eg_embed_breaker_until = 0;
|
||||
* rates keep the previous reading and diff. Restart legitimately resets to 0. */
|
||||
static int64_t _eg_act_breakthroughs = 0; /* forced promotions at the floor, cumulative */
|
||||
static int64_t _eg_act_wm_evicted = 0; /* ALL WM evictions, cumulative (see below) */
|
||||
/* ── Eviction CAUSE decomposition (2026-08-14 self-review) ──────────────────
|
||||
* _eg_act_wm_evicted is incremented from six sites with four distinct causes,
|
||||
* and every one of them collapsed into that single integer. Today's review
|
||||
* measured 175,547 evictions over 13.5h (~216/min against 24 slots) and could
|
||||
* not tell healthy rotation from cap thrashing from duplicate churn, because
|
||||
* the only available number counts all three the same way.
|
||||
*
|
||||
* That is this file's most-repeated defect. The 08-02 and 08-06 reviews were
|
||||
* each diagnosable only because someone first added a NEW gauge; dup_wm and
|
||||
* dup_wm_global exist precisely because the aggregate could not answer "why".
|
||||
* These three finish the decomposition, so that
|
||||
* evicted == floor + cap + bll + dup_wm + dup_wm_global
|
||||
* holds as an identity and each term names a different corrective action:
|
||||
* floor - candidates below the absolute admission bar. High = weak retrieval.
|
||||
* cap - lost the rank contest for 24 slots. High = genuine contention.
|
||||
* bll - carried-over residents that decayed under the ACT-R tau. High =
|
||||
* healthy forgetting, NOT pressure.
|
||||
* Confusing the third with the second is what makes WM churn unreadable. */
|
||||
static int64_t _eg_act_evict_floor = 0; /* below ENGRAM_WM_FLOOR (both passes) */
|
||||
static int64_t _eg_act_evict_cap = 0; /* over ENGRAM_WM_CAP (both passes) */
|
||||
static int64_t _eg_act_evict_bll = 0; /* carry-over decayed under BLL tau */
|
||||
/* Redundancy suppression counters (2026-08-05 self-review) — see
|
||||
* ENGRAM_DEDUP_COS. dup_seeds = semantic seed slots reclaimed from redundant
|
||||
* copies; dup_wm = WM candidates dropped for duplicating a higher-ranked
|
||||
@@ -6854,10 +6782,6 @@ typedef struct EngramStore {
|
||||
int* adj_to_len;
|
||||
int adj_dirty; /* 1 = rebuild needed before next BFS */
|
||||
int64_t adj_node_count; /* node_count at time of last adj_rebuild */
|
||||
/* Nodes with degree >= 1 at last adj_rebuild. The denominator for the
|
||||
* fan-effect reference degree — see eg_fan_factor for why isolated nodes
|
||||
* must not be counted. (2026-08-11 self-review) */
|
||||
int64_t adj_connected;
|
||||
} EngramStore;
|
||||
|
||||
static EngramStore* engram_global = NULL;
|
||||
@@ -7221,16 +7145,11 @@ static void engram_adj_rebuild(EngramStore* g) {
|
||||
if (ti >= 0 && g->adj_to[ti])
|
||||
g->adj_to[ti][to_pos[ti]++] = (int)ei;
|
||||
}
|
||||
/* Copy counts. Also tally how many nodes have any edge at all — the
|
||||
* fan-effect denominator. Free here, in the O(V) pass that already exists,
|
||||
* rather than as a separate scan. (2026-08-11 self-review) */
|
||||
int64_t connected = 0;
|
||||
/* Copy counts */
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
g->adj_from_len[i] = from_cnt[i];
|
||||
g->adj_to_len[i] = to_cnt[i];
|
||||
if (from_cnt[i] + to_cnt[i] > 0) connected++;
|
||||
}
|
||||
g->adj_connected = connected;
|
||||
free(from_cnt); free(to_cnt); free(from_pos); free(to_pos);
|
||||
g->adj_node_count = g->node_count;
|
||||
g->adj_dirty = 0;
|
||||
@@ -8212,7 +8131,6 @@ static void eg_wm_carry_over(EngramNode* cn, int64_t now_ms, int64_t* evict_ctr)
|
||||
cn->working_memory_weight = 0.0;
|
||||
cn->wm_anchor = 0.0;
|
||||
if (evict_ctr) (*evict_ctr)++;
|
||||
_eg_act_evict_bll++;
|
||||
} else {
|
||||
cn->working_memory_weight = w;
|
||||
}
|
||||
@@ -8379,108 +8297,6 @@ static double engram_activation_dampen(const EngramNode* n) {
|
||||
return 1.0 / (1.0 + log(1.0 + (double)n->activation_count));
|
||||
}
|
||||
|
||||
/* ── ACT-R fan effect: degree normalization for spreading activation ─────────
|
||||
* (2026-08-11 self-review. Closes the other half of a mechanism that has been
|
||||
* half-implemented since the BLL work.)
|
||||
*
|
||||
* THE GAP. This runtime implements ACT-R's base-level learning term
|
||||
* B_i = ln(Σ t_k^-d) (engram_bll_base_level) but never implemented the
|
||||
* ASSOCIATIVE term that goes with it:
|
||||
*
|
||||
* A_i = B_i + Σ_j W_j · S_ji where S_ji = S − ln(fan_j)
|
||||
*
|
||||
* fan_j is the number of things j is associated with. The whole point of the
|
||||
* fan effect (Anderson 1974; Anderson & Reder 1999) is that a source spreads a
|
||||
* FIXED budget of activation across its associations — so being connected to
|
||||
* many things makes each individual connection weaker. Without it, degree is
|
||||
* pure advantage: a node wins retrieval by being popular rather than by being
|
||||
* relevant. That is backwards, and it is what this graph has been doing.
|
||||
*
|
||||
* MEASURED ON THE LIVE STORE (13,129 nodes / 43,400 edges, 2026-08-11):
|
||||
* degree p50=14 p90=34 p95=82 p99=275 max=357 mean=23.3
|
||||
* the top 1% of nodes by degree touch 21.2% of all edges
|
||||
* So the most-connected node had a 25x propagation advantage over the median
|
||||
* node for no reason other than accumulated connections. The top hubs are not
|
||||
* even semantically central — several are duplicate pairs of the same document
|
||||
* left over from the redundancy census of the 2026-08-05 review.
|
||||
*
|
||||
* The hub problem was already recognized twice and patched narrowly both
|
||||
* times: InternalStateEvent nodes were cut out of propagation entirely (see
|
||||
* the frontier loop) and eg_hebb_node_budget caps per-node Hebbian mass. Both
|
||||
* are special cases of this general law. This is the general fix.
|
||||
*
|
||||
* FORM. Symmetric normalization, w / (deg(u)^β · deg(v)^β) with β = 0.5 — the
|
||||
* normalized-Laplacian / GCN form, which penalizes a hub both for sending and
|
||||
* for receiving. Both failure modes are live here: a hub source floods its
|
||||
* neighborhood, and a hub target gets reached by everything regardless of
|
||||
* relevance. Written relative to the graph's own mean degree:
|
||||
*
|
||||
* fan(u,v) = clamp( d_ref / sqrt(deg(u) · deg(v)), FAN_MIN, 1.0 )
|
||||
* d_ref = 2·|E| / |V| (mean degree, O(1), live)
|
||||
*
|
||||
* WHY IT IS CLAMPED AT 1.0 ON TOP — this is the load-bearing safety property,
|
||||
* not a detail. The factor can only ever REDUCE propagation, never amplify it.
|
||||
* Every constant downstream of this multiply is calibrated against today's
|
||||
* activation magnitudes: the 0.02 firing threshold, SPREAD_DECAY = 0.7, the
|
||||
* 0.15 WM promotion threshold, the 24-slot WM cap. A normalization that
|
||||
* boosted low-degree nodes would inflate the frontier, change how many nodes
|
||||
* clear 0.02, and silently recalibrate working memory as a side effect of a
|
||||
* change that was supposed to be about hubs. Capping at 1.0 means every pair
|
||||
* at or below mean degree — the common case — propagates EXACTLY as it does
|
||||
* today, and the only behavior that changes is that above-mean hubs stop
|
||||
* winning on degree alone. Strictly monotone, strictly conservative, and the
|
||||
* blast radius is confined to the nodes the change is aimed at.
|
||||
*
|
||||
* Self-calibrating: d_ref is recomputed from the live graph, so the correction
|
||||
* tracks densification instead of drifting against a constant that was right
|
||||
* in August 2026 and wrong a year later. Change is the signal.
|
||||
*
|
||||
* FAN_MIN = 0.30 bottoms the penalty at ~3.3x rather than the ~15x that raw
|
||||
* 1/deg would give at max degree. Same reasoning as ENGRAM_QGATE_FLOOR: damp
|
||||
* the uninformative path, never sever it. A hub is usually a hub for a reason;
|
||||
* it just should not also get a free win.
|
||||
*
|
||||
* Sources: Anderson & Reder 1999 (fan effect, S=1.6-2.0, d=0.5) ·
|
||||
* arXiv:2405.14831 HippoRAG (node specificity) · Systems 9(2):22
|
||||
* (normalized-Laplacian spreading activation) · arXiv:2606.30133 (β is a
|
||||
* low-sensitivity knob; gating and fan normalization carry the effect). */
|
||||
/* FAN_MIN 0.50, not the 0.30 this shipped as on the first build. Measured on
|
||||
* the live graph, β=0.5 with a 0.30 floor damped 96% of propagation steps to a
|
||||
* mean factor of 0.34 — and that number is not a bug in the correction, it is
|
||||
* an honest measurement of how hub-dominated traversal here actually is. But a
|
||||
* ~3x near-uniform damp is a bigger global change than one A/B run justifies,
|
||||
* and it cost a working-memory promotion (5 → 4) on the one query measured
|
||||
* cleanly. A 0.50 floor keeps the full mechanism and the whole [0.5, 1.0]
|
||||
* dynamic range for separating hubs from non-hubs, at half the blast radius.
|
||||
* The fan_mean / fan_hits gauges make the next review's tuning evidence-based
|
||||
* rather than another guess: loosen it when the data says WM can afford it. */
|
||||
#define ENGRAM_FAN_MIN 0.50
|
||||
|
||||
/* eg_node_degree — total (in + out) degree from the adjacency index. The index
|
||||
* is rebuilt at the top of engram_activate whenever topology changed, so this
|
||||
* is current. adj_node_count is the count at BUILD time and can lag
|
||||
* node_count; out-of-range indices report 0 and are treated as unpenalized. */
|
||||
static int eg_node_degree(const EngramStore* g, int64_t idx) {
|
||||
if (idx < 0 || idx >= g->adj_node_count) return 0;
|
||||
if (!g->adj_from_len || !g->adj_to_len) return 0;
|
||||
return g->adj_from_len[idx] + g->adj_to_len[idx];
|
||||
}
|
||||
|
||||
static double eg_fan_factor(const EngramStore* g, double d_ref,
|
||||
int64_t u_idx, int64_t v_idx) {
|
||||
if (d_ref <= 0.0) return 1.0;
|
||||
int du = eg_node_degree(g, u_idx);
|
||||
int dv = eg_node_degree(g, v_idx);
|
||||
/* Degree 0 is only reachable when the adjacency index is stale or absent;
|
||||
* an actually-isolated node is never on the frontier. Do not penalize what
|
||||
* we cannot measure. */
|
||||
if (du <= 0 || dv <= 0) return 1.0;
|
||||
double f = d_ref / sqrt((double)du * (double)dv);
|
||||
if (f > 1.0) return 1.0; /* never amplify — see above */
|
||||
if (f < ENGRAM_FAN_MIN) return ENGRAM_FAN_MIN;
|
||||
return f;
|
||||
}
|
||||
|
||||
/* Temporal proximity bonus: boost propagation along edges connecting
|
||||
* co-temporal nodes. Returns a multiplier bonus in [0, 0.2]. */
|
||||
static double engram_temporal_proximity_bonus(int64_t node_created,
|
||||
@@ -8648,8 +8464,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
* miss nearly all events between beats; see the definition site).
|
||||
* ctx_cos stays per-call: it is a gauge of THIS query vs the centroid. */
|
||||
_eg_act_ctx_cos = -2.0;
|
||||
_eg_act_fan_sum = 0.0; _eg_act_fan_min = 1.0;
|
||||
_eg_act_fan_n = 0; _eg_act_fan_hits = 0;
|
||||
|
||||
/* ── Embedding backfill + query embedding (2026-07-24, bl-b2d1c944) ──
|
||||
* Backfill: embed up to N un-embedded eligible nodes per call, newest
|
||||
@@ -8882,29 +8696,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
ftail++;
|
||||
}
|
||||
const double SPREAD_DECAY = 0.7;
|
||||
/* Reference degree for the fan-effect correction: mean degree over
|
||||
* CONNECTED nodes, 2|E| / |{v : deg(v) > 0}|. O(1) — adj_connected is
|
||||
* tallied during adjacency rebuild.
|
||||
*
|
||||
* NOT 2|E|/|V|. That was the first cut and instrumentation caught it
|
||||
* immediately: on the live graph it gives d_ref = 6.61, while the median
|
||||
* degree of a node that actually has edges is 14. Isolated nodes cannot
|
||||
* be on the frontier — spreading activation only ever traverses connected
|
||||
* ones — so including them in the denominator deflates the reference below
|
||||
* anything traversal will ever see, and the correction pins to
|
||||
* ENGRAM_FAN_MIN on every step. Measured on the first build:
|
||||
* fan_mean 0.3026 with fan_hits 579/579 — a uniform 0.30 multiplier, which
|
||||
* is not a fan effect at all. It is just a weaker SPREAD_DECAY, and it
|
||||
* would have quietly recalibrated the 0.02 firing threshold and WM
|
||||
* competition while appearing to be a targeted change.
|
||||
*
|
||||
* Over connected nodes the reference is ~23, above the median, so typical
|
||||
* traversal rides the 1.0 cap unchanged and only genuine hubs are damped
|
||||
* — which is the whole intent. The gauge that caught this is the reason it
|
||||
* was worth adding the gauge. */
|
||||
const double FAN_DREF = (g->adj_connected > 0)
|
||||
? (2.0 * (double)g->edge_count / (double)g->adj_connected) : 0.0;
|
||||
_eg_act_fan_dref = FAN_DREF;
|
||||
while (fhead < ftail) {
|
||||
Frontier f = fr[fhead++];
|
||||
if (f.hops >= max_depth) continue;
|
||||
@@ -8971,51 +8762,16 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
* ~4x, never killed); unembedded targets pass ungated (no
|
||||
* information, no penalty); cosq == NULL (embedder down) means
|
||||
* no gating at all — same graceful degradation as seeding. */
|
||||
/* Rescale before gating (2026-08-14 self-review). Raw cosine from
|
||||
* nomic-embed is compressed into a narrow high band, so feeding it
|
||||
* to the gate directly makes the gate nearly a constant. Measured
|
||||
* on this store: 400 random UNRELATED node pairs gave median 0.562,
|
||||
* central 98% span [0.381, 0.743]. The raw gate therefore passed a
|
||||
* typical unrelated node at 0.25 + 0.75*0.562 = 0.67 — two thirds
|
||||
* strength for a node with no semantic relation to the query. That
|
||||
* is not a gate, it is a small tax.
|
||||
*
|
||||
* Shift-and-floor about ENGRAM_EMBED_S0, exactly as the Pass-2 WM
|
||||
* term at ENGRAM_EMBED_WM_WEIGHT already does. The constant was in
|
||||
* this file for this reason; the propagation gate simply never used
|
||||
* it. Same store, same 400 pairs, after the rescale: the median
|
||||
* unrelated pair drops to 0.40 while the top of the range is
|
||||
* preserved (0.85 vs 0.92), and gate spread widens 0.42 -> 0.60.
|
||||
* Only 8.5% of pairs fall to the floor, so lexical/structural
|
||||
* pathways through dissimilar nodes are damped, never severed.
|
||||
* Cf. arXiv:2512.15922, which rescales w' = (w-c)/(1-c) about
|
||||
* c = 0.4 for precisely this reason ("prevent overactivation and
|
||||
* context explosion"). */
|
||||
double qgate = 1.0;
|
||||
if (cosq && cosq[oi] > -1.5) {
|
||||
double c = (cosq[oi] - ENGRAM_EMBED_S0) / (1.0 - ENGRAM_EMBED_S0);
|
||||
if (c < 0.0) c = 0.0;
|
||||
if (c > 1.0) c = 1.0;
|
||||
double c = cosq[oi] > 0.0 ? cosq[oi] : 0.0;
|
||||
qgate = ENGRAM_QGATE_FLOOR + (1.0 - ENGRAM_QGATE_FLOOR) * c;
|
||||
}
|
||||
/* ── ACT-R fan effect (2026-08-11 self-review) ──
|
||||
* Symmetric degree normalization over the (source, target) pair.
|
||||
* The query gate above prunes branches that are semantically
|
||||
* irrelevant; this prunes branches that are merely POPULAR. They
|
||||
* are different failure modes — a duplicate document with 357
|
||||
* edges can be highly cosine-similar to the query and still be
|
||||
* the wrong thing to spread through. Only ever <= 1.0, so it
|
||||
* cannot inflate the frontier. See eg_fan_factor. */
|
||||
double fan = eg_fan_factor(g, FAN_DREF, cur, oi);
|
||||
_eg_act_fan_sum += fan;
|
||||
_eg_act_fan_n++;
|
||||
if (fan < 1.0) _eg_act_fan_hits++;
|
||||
if (fan < _eg_act_fan_min) _eg_act_fan_min = fan;
|
||||
/* eg_edge_eff_weight, not e->weight: edges that have repeatedly
|
||||
* carried co-activated pairs propagate more strongly. Identity on
|
||||
* an unlearned edge. (2026-08-04 self-review.) */
|
||||
double new_act = f.act * eg_edge_eff_weight(e) * SPREAD_DECAY
|
||||
* (1.0 + tbonus) * tdecay * dampen * qgate * fan;
|
||||
* (1.0 + tbonus) * tdecay * dampen * qgate;
|
||||
/* Firing threshold per classic spreading-activation: sub-threshold
|
||||
* activation neither updates the target nor enqueues it, so weak
|
||||
* signals die out instead of flooding the whole graph with tiny
|
||||
@@ -9305,7 +9061,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
if (wm_weights[i] > 0.0 && wm_weights[i] < ENGRAM_WM_FLOOR) {
|
||||
wm_weights[i] = 0.0;
|
||||
_eg_act_wm_evicted++;
|
||||
_eg_act_evict_floor++;
|
||||
}
|
||||
}
|
||||
int64_t cap_count = 0;
|
||||
@@ -9341,7 +9096,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
}
|
||||
wm_weights[i] = 0.0; /* over cap: evict */
|
||||
_eg_act_wm_evicted++;
|
||||
_eg_act_evict_cap++;
|
||||
}
|
||||
}
|
||||
/* If malloc failed, skip cap — WM unbounded this call, no corruption. */
|
||||
@@ -9453,7 +9207,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
fn->working_memory_weight = 0.0;
|
||||
fn->wm_anchor = 0.0;
|
||||
_eg_act_wm_evicted++;
|
||||
_eg_act_evict_floor++;
|
||||
}
|
||||
}
|
||||
/* ── Global redundancy suppression (2026-08-06 self-review) ──────────
|
||||
@@ -9562,7 +9315,6 @@ el_val_t engram_activate(el_val_t query, el_val_t depth) {
|
||||
n->working_memory_weight = 0.0; /* evict: over global cap */
|
||||
n->wm_anchor = 0.0; /* keep anchor coherent */
|
||||
_eg_act_wm_evicted++; /* was uncounted before 2026-08-02 */
|
||||
_eg_act_evict_cap++;
|
||||
}
|
||||
}
|
||||
/* If malloc failed, skip — WM over cap this call, no data corruption. */
|
||||
@@ -9992,24 +9744,11 @@ static void engram_emit_edge_json(JsonBuf* b, const EngramEdge* e) {
|
||||
jb_putc(b, '}');
|
||||
}
|
||||
|
||||
/* Size of the last snapshot this process serialized. Seeds the next save's
|
||||
* buffer so the doubling walk never runs on the big document. See jb_init_cap
|
||||
* for the measurement that motivated it. (2026-08-11 self-review) */
|
||||
static size_t _eg_save_cap_hint = 0;
|
||||
|
||||
el_val_t engram_save(el_val_t path) {
|
||||
const char* p = EL_CSTR(path);
|
||||
if (!p || !*p) return 0;
|
||||
EngramStore* g = engram_get();
|
||||
/* Pre-size from the previous save plus 12.5% headroom, so ordinary growth
|
||||
* between snapshots does not trigger a realloc and the request size stays
|
||||
* stable enough for the allocator to reuse the same span. First save of
|
||||
* the process has no hint and starts at 1MB — still 14 doublings better
|
||||
* than 64 bytes. */
|
||||
JsonBuf b;
|
||||
jb_init_cap(&b, _eg_save_cap_hint
|
||||
? _eg_save_cap_hint + (_eg_save_cap_hint >> 3) + 1024
|
||||
: (size_t)1 << 20);
|
||||
JsonBuf b; jb_init(&b);
|
||||
jb_puts(&b, "{\"nodes\":[");
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
if (i > 0) jb_putc(&b, ',');
|
||||
@@ -10049,10 +9788,6 @@ el_val_t engram_save(el_val_t path) {
|
||||
jb_putc(&b, '}');
|
||||
}
|
||||
jb_puts(&b, "]}");
|
||||
/* Remember the size BEFORE the write: the hint is about how much buffer
|
||||
* the next serialization needs, which is a property of the graph, not of
|
||||
* whether this particular fopen succeeded. */
|
||||
_eg_save_cap_hint = b.len;
|
||||
FILE* f = fopen(p, "wb");
|
||||
if (!f) { free(b.buf); return 0; }
|
||||
size_t w = fwrite(b.buf, 1, b.len, f);
|
||||
@@ -10979,11 +10714,8 @@ el_val_t engram_act_stats_json(void) {
|
||||
}
|
||||
/* 768, not 512: the write-back gauges added 2026-08-07 push the worst-case
|
||||
* rendering past the old bound, and snprintf would truncate the JSON into
|
||||
* an unparseable tail rather than fail loudly.
|
||||
* 1152, not 896: the five fan-effect gauges added 2026-08-11 add ~90 bytes
|
||||
* worst-case. Same reasoning — headroom is cheaper than a truncated tail
|
||||
* that every downstream JSON parser rejects as a whole. */
|
||||
char buf[1152];
|
||||
* an unparseable tail rather than fail loudly. */
|
||||
char buf[896];
|
||||
/* ctx_cos (2026-07-29): cos(query, context centroid) at the LAST
|
||||
* activate call, measured before the query was folded in. ~1.0 =
|
||||
* context aligned with current query; low = divergence (expected at
|
||||
@@ -10991,15 +10723,6 @@ el_val_t engram_act_stats_json(void) {
|
||||
* embedder down. The drift gauge for the context-centroid mechanism. */
|
||||
snprintf(buf, sizeof(buf),
|
||||
"{\"wm_evicted\":%lld,\"breakthroughs\":%lld,"
|
||||
/* Eviction cause decomposition (2026-08-14 self-review):
|
||||
* wm_evicted == evict_floor + evict_cap + evict_bll
|
||||
* + dup_wm + dup_wm_global.
|
||||
* Read them as a ratio, not a level. cap-dominant = real
|
||||
* contention for the 24 slots; bll-dominant = healthy decay of
|
||||
* carried-over residents; floor-dominant = retrieval is returning
|
||||
* weak candidates. The aggregate alone cannot distinguish these
|
||||
* and every prior WM incident needed a new gauge to diagnose. */
|
||||
"\"evict_floor\":%lld,\"evict_cap\":%lld,\"evict_bll\":%lld,"
|
||||
"\"embed_breaker_open\":%d,\"embed_consec_fail\":%d,"
|
||||
"\"ctx_cos\":%.3f,"
|
||||
"\"hebb_edges\":%lld,\"hebb_max\":%.4f,\"hebb_mass\":%.3f,"
|
||||
@@ -11013,19 +10736,9 @@ el_val_t engram_act_stats_json(void) {
|
||||
* any climb means a write path is mangling text again. Cheap
|
||||
* (counted at creation) — the full census lives in
|
||||
* engram_text_health_json. (2026-08-08 self-review) */
|
||||
"\"txt_damaged\":%lld,"
|
||||
/* Fan-effect gauges (2026-08-11 self-review) — see the
|
||||
* _eg_act_fan_* definitions. fan_mean == 1.0 with fan_hits == 0
|
||||
* means the degree correction never bound on the last activation;
|
||||
* a mean drifting toward ENGRAM_FAN_MIN means traversal is
|
||||
* running through hubs and the correction is doing work. */
|
||||
"\"fan_mean\":%.4f,\"fan_min\":%.4f,\"fan_hits\":%lld,"
|
||||
"\"fan_steps\":%lld,\"fan_dref\":%.2f}",
|
||||
"\"txt_damaged\":%lld}",
|
||||
(long long)_eg_act_wm_evicted,
|
||||
(long long)_eg_act_breakthroughs,
|
||||
(long long)_eg_act_evict_floor,
|
||||
(long long)_eg_act_evict_cap,
|
||||
(long long)_eg_act_evict_bll,
|
||||
breaker_open, _eg_embed_consec_fail,
|
||||
_eg_act_ctx_cos,
|
||||
(long long)hebb_edges, hebb_max, hebb_mass,
|
||||
@@ -11035,10 +10748,7 @@ el_val_t engram_act_stats_json(void) {
|
||||
(long long)_eg_hebb_wb_dropped,
|
||||
(long long)_eg_act_dup_seeds, (long long)_eg_act_dup_wm,
|
||||
(long long)_eg_act_dup_wm_global,
|
||||
(long long)_eg_txt_write_damaged,
|
||||
(_eg_act_fan_n > 0 ? _eg_act_fan_sum / (double)_eg_act_fan_n : 1.0),
|
||||
_eg_act_fan_min, (long long)_eg_act_fan_hits,
|
||||
(long long)_eg_act_fan_n, _eg_act_fan_dref);
|
||||
(long long)_eg_txt_write_damaged);
|
||||
return el_wrap_str(el_strdup(buf));
|
||||
}
|
||||
|
||||
@@ -11171,285 +10881,6 @@ el_val_t engram_label_df(el_val_t term) {
|
||||
return (el_val_t)df;
|
||||
}
|
||||
|
||||
/* ── Salient-term extraction (2026-08-13 self-review) ────────────────────────
|
||||
* THE MEASUREMENT. auto_term_empty_streak, the counter added by the 2026-08-06
|
||||
* review precisely to catch this class of silent death, read 50 and climbing.
|
||||
* Fifty consecutive curiosity scans in which the soul's dynamic seeding path
|
||||
* produced NOTHING and the loop fell back to its four hardcoded rotating
|
||||
* phrases. Dumping the live WM top says why in one look:
|
||||
*
|
||||
* Memory 0.390 memory:remembered
|
||||
* Memory 0.378 memory:remembered
|
||||
* Memory 0.377 memory:remembered
|
||||
* Memory 0.373 memory:remembered
|
||||
* Memory 0.370 memory:remembered
|
||||
*
|
||||
* Every slot at the top of working memory is a Memory node, and every Memory
|
||||
* node written by remember() carries the sentinel label "memory:remembered".
|
||||
* auto_term_try_slot reads the LABEL and only the label; the colon-no-space
|
||||
* guard (correctly) rejects sentinels as carrying no seed signal; so the
|
||||
* extractor had nothing to work with and returned empty, forever.
|
||||
*
|
||||
* THE ACTUAL DEFECT is not the sentinel guard — that guard is right. It is
|
||||
* that the extractor was built against Knowledge nodes, which have real
|
||||
* titles, and is structurally blind to the node type that in fact dominates
|
||||
* working memory. The label is not the content. A Memory node's topic is in
|
||||
* its text; the runtime just never looked there.
|
||||
*
|
||||
* WHY NOT ANOTHER GUARD. The extractor's whole history is guards: genre words
|
||||
* (07-23), quoted titles (07-25), English stopwords (07-30), label-df
|
||||
* (08-03). Four reviews, four blocklists, each written after watching a flood
|
||||
* happen. That is a losing shape, and 08-03 said so explicitly before adding
|
||||
* the fifth. The reason it keeps recurring is the algorithm underneath:
|
||||
* TAKE THE FIRST WORD, THEN CHECK WHETHER IT IS ACCEPTABLE. A first-word
|
||||
* extractor has no notion of term quality, so quality has to be bolted on as
|
||||
* rejection, and rejection can only encode the past.
|
||||
*
|
||||
* THE FIX is to invert it: score EVERY candidate token in the text and take
|
||||
* the argmax. Then term quality is the selection criterion rather than a
|
||||
* veto, and a bad token does not need to be on a list to lose — it only needs
|
||||
* a better token in the same text, which is the common case.
|
||||
*
|
||||
* SCORING (YAKE, Campos et al., Information Sciences 509:257-289, 2020 —
|
||||
* lightweight unsupervised single-document keyword extraction). YAKE scores
|
||||
* candidates on casing, position, frequency, context relatedness and sentence
|
||||
* dispersion, and beats RAKE/TextRank/SingleRank across twenty datasets. Two
|
||||
* of its five features port directly and cheaply; the other three are
|
||||
* within-document proxies for a corpus YAKE deliberately does not have. This
|
||||
* system DOES have the corpus — 12.7k labelled nodes — so real IDF is
|
||||
* substituted where YAKE has to approximate:
|
||||
*
|
||||
* score(t) = idf(t) · position(t) · casing(t)
|
||||
*
|
||||
* idf = ln((N+1)/(df+1)) real corpus specificity (Spärck
|
||||
* Jones 1972), strictly better than
|
||||
* YAKE's TF-based stand-in
|
||||
* position = 1/ln(e + i) YAKE T_Position: earlier tokens are
|
||||
* more topical. Keeps the old
|
||||
* first-word bias as a SOFT preference
|
||||
* instead of an absolute rule
|
||||
* casing = 1.30 acronym / 1.15 capitalised / 1.00 otherwise
|
||||
* YAKE T_Case
|
||||
*
|
||||
* THE min_df GATE. The df ceiling (08-03) rejects corpus-frequent markup and
|
||||
* sentinels. A floor was added alongside it for an independent reason: a term
|
||||
* appearing in ZERO labels cannot lexically reach anything, so it is a bad
|
||||
* seed however specific it looks.
|
||||
*
|
||||
* An earlier draft of this comment claimed the floor also subsumes the 73
|
||||
* hand-listed stopwords that 08-03 measured label-df as missing (Whose:0,
|
||||
* Would:0, Could:0). MEASURED, AND THAT CLAIM IS FALSE. Under word-boundary
|
||||
* df on the live store, function words are rare in labels but not absent:
|
||||
* about:2, whole:1, them:2, head:2. They clear a floor of 1. What actually
|
||||
* keeps them from winning is the argmax itself — they carry no position
|
||||
* advantage and lose to a topical term in the same text on every node
|
||||
* measured. The stopword list therefore STAYS as a real defense for the
|
||||
* Title-case cases, not as vestigial belt-and-braces. Recording the
|
||||
* correction rather than the tidier story: the floor buys lexical
|
||||
* reachability, the argmax buys quality, and the list still earns its keep.
|
||||
*
|
||||
* TABU IS APPLIED DURING THE ARGMAX, not after it. The old code picked a term
|
||||
* and then discarded it if it was tabu, which turned inhibition-of-return
|
||||
* into another source of empty scans. Excluding tabu terms from the candidate
|
||||
* set instead yields the best NON-TABU term, so rotation costs quality rather
|
||||
* than costing the whole scan.
|
||||
*
|
||||
* COST. One pass over g->nodes scoring all candidates at once (12.7k labels ×
|
||||
* <=32 candidates, short strings, good locality), twice per 30 s scan.
|
||||
*
|
||||
* POLICY LIVES IN THE SOUL. Thresholds arrive as arguments; the runtime
|
||||
* measures and ranks, awareness.el decides. Same split as engram_label_df.
|
||||
*
|
||||
* Returns the winning token, or "" when the node is missing, has no usable
|
||||
* text, or every candidate is gated out — "" remains the honest signal that
|
||||
* this slot yielded no seed, and auto_term_empty_streak still counts it. */
|
||||
#define ENGRAM_ST_MAXCAND 32
|
||||
#define ENGRAM_ST_TOKLEN 64
|
||||
#define ENGRAM_ST_SCANCHARS 400
|
||||
|
||||
/* Trim leading/trailing non-alphanumerics, then accept only tokens whose core
|
||||
* is alphanumeric plus '-' and '_' with at least 3 letters. This subsumes the
|
||||
* quoted-title guard (2026-07-25) and the "<!--" flood (2026-08-03)
|
||||
* structurally: markup and punctuation-bearing tokens never become
|
||||
* candidates, rather than being blocklisted after the fact. */
|
||||
static int eg_st_clean_token(const char* raw, size_t rawlen,
|
||||
char* out, size_t outcap) {
|
||||
size_t s = 0, e = rawlen;
|
||||
while (s < e && !isalnum((unsigned char)raw[s])) s++;
|
||||
while (e > s && !isalnum((unsigned char)raw[e - 1])) e--;
|
||||
size_t len = e - s;
|
||||
if (len < 4 || len >= outcap) return 0;
|
||||
int alpha = 0;
|
||||
for (size_t i = 0; i < len; i++) {
|
||||
unsigned char c = (unsigned char)raw[s + i];
|
||||
if (isalpha(c)) alpha++;
|
||||
else if (!isdigit(c) && c != '-' && c != '_') return 0;
|
||||
}
|
||||
if (alpha < 3) return 0;
|
||||
memcpy(out, raw + s, len);
|
||||
out[len] = '\0';
|
||||
return 1;
|
||||
}
|
||||
|
||||
/* ENGRAM_ST_DEBUG=1 dumps the full scored candidate set to stderr. One
|
||||
* cached branch in production. This exists because the first live run of this
|
||||
* function returned five ALL-CAPS terms in a row and there was no way to see
|
||||
* whether that was the corpus or the casing weight without guessing — the
|
||||
* lesson this system keeps relearning. */
|
||||
static int _eg_st_debug(void) {
|
||||
static int v = -1;
|
||||
if (v < 0) { const char* e = getenv("ENGRAM_ST_DEBUG"); v = (e && *e == '1'); }
|
||||
return v;
|
||||
}
|
||||
|
||||
/* Word-boundary document frequency. engram_label_df uses istr_contains, i.e.
|
||||
* SUBSTRING matching, and that is the wrong estimator for term specificity on
|
||||
* short tokens: "them" hits inside "theme" and "anthem", "about" and "whole"
|
||||
* come back with df 2 and 1 rather than 0. That matters here specifically
|
||||
* because the min_df floor is what rejects English function words, and it can
|
||||
* only do that job if their df is honestly zero. Substring df quietly handed
|
||||
* them a survival ticket. Measured on the live store before this fix, "whole"
|
||||
* (df=1, idf=8.76) and "about" (df=2, idf=8.36) were outscoring real topical
|
||||
* terms and losing only on position — one node whose text happened to open
|
||||
* with a function word would have seeded on it.
|
||||
*
|
||||
* engram_label_df keeps substring semantics: it is a separate published
|
||||
* measure with existing callers, and changing it underneath them is not this
|
||||
* change's business. */
|
||||
static int eg_st_label_has_word(const char* hay, const char* word) {
|
||||
size_t wl = strlen(word);
|
||||
for (const char* p = hay; *p; p++) {
|
||||
if (strncasecmp(p, word, wl) != 0) continue;
|
||||
char before = (p == hay) ? '\0' : p[-1];
|
||||
char after = p[wl];
|
||||
if (before && (isalnum((unsigned char)before) || before == '_')) continue;
|
||||
if (after && (isalnum((unsigned char)after) || after == '_')) continue;
|
||||
return 1;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* YAKE T_Case, adapted to this corpus. YAKE up-weights all-caps tokens
|
||||
* because in ordinary prose an acronym is rare and carries topic. That
|
||||
* assumption does not hold here: memory content written by remember()
|
||||
* conventionally OPENS WITH AN ALL-CAPS HEADER ("FRAME-ROUTER UPGRADE —
|
||||
* RESULTS", "THE GAP", "CENSUS"), so a flat acronym bonus systematically
|
||||
* hands the seed to whatever word the header happens to start with and lets
|
||||
* casing override the specificity signal it is supposed to only nudge.
|
||||
* Measured on the live store: the first five WM nodes returned PRIMING,
|
||||
* CONVERSATION, OCCUPATION, RELATIONAL, SELF-OCCUPATION — every one an
|
||||
* all-caps header word, none chosen on its merits.
|
||||
*
|
||||
* Genuine acronyms are SHORT (VBD, CCR, MCP, HTTP); shouty headers are long
|
||||
* words that happen to be capitalised. So the acronym bonus is restricted to
|
||||
* tokens of <= 5 characters, where all-caps is actually evidence of an
|
||||
* acronym rather than evidence of a heading. Longer all-caps tokens fall
|
||||
* through to the ordinary Title-case nudge — they still compete, they just
|
||||
* compete on specificity instead of on volume. */
|
||||
static double eg_st_casing(const char* t) {
|
||||
int upper = 0, lower = 0;
|
||||
size_t len = 0;
|
||||
for (const char* q = t; *q; q++, len++) {
|
||||
if (isupper((unsigned char)*q)) upper++;
|
||||
else if (islower((unsigned char)*q)) lower++;
|
||||
}
|
||||
if (lower == 0 && upper >= 2 && len <= 5) return 1.30; /* acronym */
|
||||
if (isupper((unsigned char)t[0])) return 1.15; /* Title/hdr */
|
||||
return 1.0;
|
||||
}
|
||||
|
||||
el_val_t engram_salient_term(el_val_t node_id, el_val_t max_df_v,
|
||||
el_val_t min_df_v, el_val_t tabu_v) {
|
||||
EngramStore* g = engram_get();
|
||||
int64_t ix = engram_find_node_index(EL_CSTR(node_id));
|
||||
if (ix < 0) return el_wrap_str(el_strdup(""));
|
||||
EngramNode* n = &g->nodes[ix];
|
||||
|
||||
int64_t max_df = (int64_t)max_df_v;
|
||||
int64_t min_df = (int64_t)min_df_v;
|
||||
if (max_df <= 0) max_df = g->node_count;
|
||||
if (min_df < 0) min_df = 0;
|
||||
const char* tabu = EL_CSTR(tabu_v);
|
||||
|
||||
/* Source selection. Prefer the label — it is a curated title when it is
|
||||
* one. Fall back to content when the label is absent or a sentinel
|
||||
* ("memory:remembered": a colon and no space). This single line is what
|
||||
* makes Memory nodes visible to the extractor at all. */
|
||||
const char* src = n->label;
|
||||
if (!src || !*src) {
|
||||
src = n->content;
|
||||
} else if (strchr(src, ':') != NULL && strchr(src, ' ') == NULL) {
|
||||
src = n->content;
|
||||
}
|
||||
if (!src || !*src) return el_wrap_str(el_strdup(""));
|
||||
|
||||
/* Collect distinct candidates from the head of the text. */
|
||||
char cand[ENGRAM_ST_MAXCAND][ENGRAM_ST_TOKLEN];
|
||||
int pos[ENGRAM_ST_MAXCAND];
|
||||
int64_t df[ENGRAM_ST_MAXCAND];
|
||||
int ncand = 0, tokidx = 0;
|
||||
|
||||
const char* p = src;
|
||||
const char* lim = src + strnlen(src, ENGRAM_ST_SCANCHARS);
|
||||
while (p < lim && ncand < ENGRAM_ST_MAXCAND) {
|
||||
while (p < lim && isspace((unsigned char)*p)) p++;
|
||||
if (p >= lim) break;
|
||||
const char* tk = p;
|
||||
while (p < lim && !isspace((unsigned char)*p)) p++;
|
||||
char buf[ENGRAM_ST_TOKLEN];
|
||||
int slot = tokidx++;
|
||||
if (!eg_st_clean_token(tk, (size_t)(p - tk), buf, sizeof(buf))) continue;
|
||||
|
||||
/* Tabu exclusion, applied here so the argmax runs over eligible
|
||||
* terms only. tabu arrives pipe-delimited: "|t0|t1|t2|t3|". */
|
||||
if (tabu && *tabu) {
|
||||
char pat[ENGRAM_ST_TOKLEN + 2];
|
||||
snprintf(pat, sizeof(pat), "|%s|", buf);
|
||||
if (istr_contains(tabu, pat)) continue;
|
||||
}
|
||||
int dup = 0;
|
||||
for (int i = 0; i < ncand; i++)
|
||||
if (strcasecmp(cand[i], buf) == 0) { dup = 1; break; }
|
||||
if (dup) continue;
|
||||
|
||||
memcpy(cand[ncand], buf, strlen(buf) + 1);
|
||||
pos[ncand] = slot;
|
||||
df[ncand] = 0;
|
||||
ncand++;
|
||||
}
|
||||
if (ncand == 0) return el_wrap_str(el_strdup(""));
|
||||
|
||||
/* One pass over the store, all candidates at once. */
|
||||
for (int64_t i = 0; i < g->node_count; i++) {
|
||||
const char* lbl = g->nodes[i].label;
|
||||
if (!lbl || !*lbl) continue;
|
||||
for (int c = 0; c < ncand; c++)
|
||||
if (eg_st_label_has_word(lbl, cand[c])) df[c]++;
|
||||
}
|
||||
|
||||
/* Argmax over idf · position · casing, subject to the df band. */
|
||||
int best = -1;
|
||||
double best_score = 0.0;
|
||||
for (int c = 0; c < ncand; c++) {
|
||||
if (df[c] > max_df) continue;
|
||||
if (df[c] < min_df) continue;
|
||||
double idf = log(((double)g->node_count + 1.0) / ((double)df[c] + 1.0));
|
||||
if (idf <= 0.0) continue;
|
||||
double position = 1.0 / log(2.718281828459045 + (double)pos[c]);
|
||||
double casing = eg_st_casing(cand[c]);
|
||||
double score = idf * position * casing;
|
||||
if (_eg_st_debug()) {
|
||||
fprintf(stderr, " cand %-24s df=%-5lld idf=%.2f pos=%d p=%.2f "
|
||||
"case=%.2f score=%.3f\n",
|
||||
cand[c], (long long)df[c], idf, pos[c], position,
|
||||
casing, score);
|
||||
}
|
||||
if (score > best_score) { best_score = score; best = c; }
|
||||
}
|
||||
if (best < 0) return el_wrap_str(el_strdup(""));
|
||||
return el_wrap_str(el_strdup(cand[best]));
|
||||
}
|
||||
|
||||
/* engram_embed_backfill — explicitly drive the lazy embedding backfill.
|
||||
* (2026-07-25 self-review.) The per-activate backfill (8 nodes/call) only
|
||||
* runs inside engram_activate, and on the authoritative HTTP store nothing
|
||||
|
||||
@@ -628,13 +628,6 @@ el_val_t engram_hebb_drain_json(el_val_t max);
|
||||
/* Document frequency of a term across node labels — term-specificity signal
|
||||
* for curiosity seed selection. (2026-08-03 self-review.) */
|
||||
el_val_t engram_label_df(el_val_t term);
|
||||
/* Best curiosity seed from one node: argmax over idf·position·casing across
|
||||
* the candidate tokens of its label, falling back to its content when the
|
||||
* label is a sentinel. Excludes pipe-delimited tabu terms during selection
|
||||
* and gates candidates to the df band [min_df, max_df]. Returns "" when
|
||||
* nothing qualifies. (2026-08-13 self-review.) */
|
||||
el_val_t engram_salient_term(el_val_t node_id, el_val_t max_df,
|
||||
el_val_t min_df, el_val_t tabu);
|
||||
el_val_t engram_embed_backfill(el_val_t count);
|
||||
el_val_t engram_list_layers_json(void);
|
||||
/* Working memory introspection — count, mean weight, and top-N snapshot.
|
||||
|
||||
@@ -1,184 +0,0 @@
|
||||
# Swarm + CCR + Work-Tracking — Neuron's bounded parallel execution, in native El
|
||||
|
||||
Bounded parallel agent execution on El's **native** concurrency — no external
|
||||
orchestrator. Grounded directly in two of Will's frameworks:
|
||||
|
||||
- **Swarm Architecture** (*Bounded Parallel Agent Execution*, Mar 2026)
|
||||
- **Compiled Context Runtime / CCR** (*Process-Driven Agent Execution with
|
||||
Unbounded Local Memory*, Mar 2026)
|
||||
|
||||
A swarm is a **coordinator** (the main thread) that mints a correlation identity,
|
||||
compiles a **bounded per-worker context (CCR)**, dispatches workers as **native
|
||||
pthreads** (`thread.el` `spawn`/`join`), tracks every unit of work durably, and
|
||||
**converges** results before returning control to the parent step.
|
||||
|
||||
```
|
||||
Parent step
|
||||
└─ swarm_run(blueprint, knowledge_refs, inputs, config)
|
||||
fan-out ──▶ worker_1 (CCR ctx_1) ─┐ native
|
||||
worker_2 (CCR ctx_2) ─┤ pthreads,
|
||||
worker_k (CCR ctx_k) ─┘ bounded by `concurrency`
|
||||
converge ─▶ collect | merge | vote | reduce ──▶ merged result
|
||||
```
|
||||
|
||||
## Why it runs on El natively
|
||||
|
||||
El is natively agentic. This capability composes El's shipped primitives — it
|
||||
adds no bespoke runtime:
|
||||
|
||||
| Primitive | Source | Role in the swarm |
|
||||
|-----------|--------|-------------------|
|
||||
| `spawn(fn,arg)` / `join(tid)` | `runtime/thread.el` → `__thread_create` (pthread + dlsym) | fan-out / rejoin |
|
||||
| `parallel_map`, `with_mutex` | `runtime/thread.el` | reference concurrency patterns |
|
||||
| Go-style channels | `runtime/channel.el` → `__channel_*` | available for vertical event streams |
|
||||
| `engram_*`, `http_*`, `fs_*`, `json_*` | `el_runtime.c` builtins | retrieval, tracking, I/O |
|
||||
|
||||
Every El fn compiles to a global C symbol, so any top-level `(String)->String`
|
||||
fn is directly threadable — the worker entry is exactly such a fn.
|
||||
|
||||
## Modules
|
||||
|
||||
| File | Framework grounding | What it does |
|
||||
|------|--------------------|--------------|
|
||||
| `worktrack.el` | Swarm §6 (correlation IDs, audit) | Durable, single-writer **JSONL journal** keyed by correlation ID; reconstructable status report; opt-in engram mirror (`SWARM_MIRROR=1`). |
|
||||
| `containment.el` | Swarm §3 + the single-writer invariant | Scope tokens w/ capabilities; **Rule 1** (no join), **Rule 2** (no open), **Rule 3** (no lateral edge), **Rule 4** (engram-write is @manager-only, by capability) enforced as checks. |
|
||||
| `ccr.el` | CCR §5 + Swarm §9.3 | Per-worker **Compiled Context Routing**: retrieve → scope → compact into a **bounded, minimal** package. The compiled-context boundary *is* the security boundary. |
|
||||
| `primitives.el` | CCR §2 (Five Primitives) | `attend / think / intend / act / learn` seam the swarm composes over. Engram-backed; explicit binding point for the API-surface reshape. |
|
||||
| `swarm.el` | Swarm §2, §4, §5 | The coordinator: fan-out/converge on native threads, bounded concurrency, four convergence strategies, integer failure threshold, full tracking. |
|
||||
|
||||
## Invariant: only the orchestrator mutates global engram state
|
||||
|
||||
**Only the orchestrator (@manager) writes to the engram / mutates global state.
|
||||
Workers are read-only against the full engram and may write only their own local
|
||||
geometry (their returned result + the journal). A worker is STRUCTURALLY UNABLE
|
||||
to mutate global engram state.**
|
||||
|
||||
This is **Rule 4** — an **authority gate, not a health gate**. Scope tokens carry
|
||||
a capability set: the orchestrator's token holds `engram:write` + `dharma:emit`
|
||||
(@manager-only, the VBD rule that only the manager mutates global state); a
|
||||
worker's token holds **only** `engram:read`. Every engram mutation
|
||||
(`op_write`/`op_relate`/`op_supersede` → `POST /api/nodes`, `/api/edges`,
|
||||
`DELETE`) flows through `swarm_engram_write`, which checks the caller's capability
|
||||
via the **same scope-token mechanism as the live Rule-2 denial** and rejects any
|
||||
worker **before any HTTP is issued**. Capability is fixed at mint time and cannot
|
||||
be acquired at runtime — so the guarantee holds regardless of engram health
|
||||
(distinct from the `SWARM_WRITE_HEALTHY` *health* gate).
|
||||
|
||||
The **curated merge is the only write path**: workers return geometry; the
|
||||
orchestrator, and only the orchestrator, commits the approved/verified geometry
|
||||
back (`commit=1`). Workers keep full-engram **read** access (`op_think`/`op_read`).
|
||||
|
||||
Proven in `harness_real_cognition.el` (§G): a worker `swarm_engram_write` is
|
||||
DENIED by capability with no node created and the violation journalled; the
|
||||
orchestrator passes the gate as the sole authorized writer.
|
||||
|
||||
## Containment → distribution
|
||||
|
||||
The three containment rules make workers **location-independent** (Swarm §9): a
|
||||
worker reads only its compiled context, shares no state with siblings, and its
|
||||
only outward edge is the returned result. The same coordinator can run workers
|
||||
as local threads today or dispatch them across machines later — the mechanism is
|
||||
identical; only the topology changes. Enforced here:
|
||||
|
||||
- **Rule 2** — `swarm_run` rejects any swarm opened under a worker token.
|
||||
- **Rules 1 + 3** — each worker gets a *closed* worker token; the coordinator is
|
||||
the only journal writer, so workers share no mutable state.
|
||||
|
||||
## Usage
|
||||
|
||||
```el
|
||||
// one process step fans out; results converge before the next step
|
||||
let inputs: String = "[\"billing\",\"payments\",\"ledger\"]"
|
||||
let refs: String = "[\"Volatility-Based Decomposition\"]" // CCR knowledge refs
|
||||
let cfg: String = "{\"concurrency\":\"4\",\"strategy\":\"collect\",\"min_success_ratio\":\"1.0\"}"
|
||||
let result: String = swarm_run("analyze_item", refs, inputs, cfg)
|
||||
// result: { corr_id, status, merged, report }
|
||||
```
|
||||
|
||||
Build any program that uses the swarm:
|
||||
|
||||
```bash
|
||||
lang/swarm/build.sh myprog.el ./myprog # concat + elc + cc (el_runtime.c)
|
||||
```
|
||||
|
||||
Config keys: `concurrency` (max workers at once), `strategy`
|
||||
(`collect|merge|vote|reduce`), `min_success_ratio` (decimal string, e.g. `0.8`),
|
||||
`caller_token` (containment). Env: `SWARM_TRACK_DIR` (journal dir),
|
||||
`CCR_TOKEN_BUDGET`, `ENGRAM_URL`/`ENGRAM_API_KEY` (retrieval + mirror),
|
||||
`SWARM_MIRROR=1`.
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
lang/swarm/build.sh lang/swarm/tests/test_swarm.el /tmp/t && SWARM_TRACK_DIR=/tmp/trk /tmp/t # 12/12
|
||||
lang/swarm/build.sh lang/swarm/tests/test_convergence.el /tmp/c && SWARM_TRACK_DIR=/tmp/trk /tmp/c # 8/8
|
||||
# integration against an isolated engram clone (never live):
|
||||
source <sandbox>/.nsbx-env
|
||||
lang/swarm/build.sh lang/swarm/tests/integ_engram.el /tmp/i && /tmp/i
|
||||
```
|
||||
|
||||
## Local-swarm integration harness (the one flip)
|
||||
|
||||
`tests/harness_local_swarm.el` proves the **full local-swarm mechanics today** on
|
||||
the isolated clone with the primitive seam pointed at the hermetic stub — 17/17
|
||||
green: 8 native-thread workers at concurrency 4, reduce + vote convergence, CCR
|
||||
scoping + non-leak, all three containment rules (incl. live Rule-2 denial),
|
||||
durable work-tracking, and **afferent telemetry** observed by the @manager.
|
||||
|
||||
Binding to the reshape's decorated primitives is **one flip and a run**:
|
||||
|
||||
```
|
||||
# in primitive_binding.el — change one line each:
|
||||
fn bound_think(ctx, instruction) { return think(ctx, instruction) } # decorated, dharma bus
|
||||
# then:
|
||||
SWARM_PRIMITIVE_SEAM=decorated lang/swarm/build.sh tests/harness_local_swarm.el ./h && ./h
|
||||
```
|
||||
|
||||
Nothing else in the swarm changes. `primitive_seam.el` (`seam_think/attend/learn`)
|
||||
already routes every worker primitive call through this one switch, and the same
|
||||
harness runs the bound path. Today `SWARM_PRIMITIVE_SEAM=decorated` still runs
|
||||
green because the binding falls back to the stub — proving the flip path executes.
|
||||
|
||||
## Real cognition — the seam is BOUND
|
||||
|
||||
`primitive_binding.el` is bound to the api-reshape agent's proven primitives
|
||||
(`wt/api-reshape@d4f401d`): `bound_think -> op_think` (GET `/api/think`), real
|
||||
768-dim gradients over the engram geometry. `reshape_surface.el` composes those
|
||||
read/cognition primitives verbatim (`op_think/read/attend/learn`).
|
||||
|
||||
`tests/harness_real_cognition.el` runs the **local swarm on real cognition**,
|
||||
17/17 green with `SWARM_PRIMITIVE_SEAM=decorated` against the `:8901` clone: 8
|
||||
native-thread workers, each a real `think` over its CCR-scoped **node-id anchor**
|
||||
(free-text anchors return "geometry unavailable"), `@manager` reduce+vote, all
|
||||
three containment rules, afferent telemetry, durable tracking. Per-anchor support
|
||||
counts (e.g. 6 / 16 / 87) drive a genuine, cognition-derived vote.
|
||||
|
||||
> **Build note (load-bearing):** the swarm build **must** define `HAVE_CURL`
|
||||
> (`build.sh` does). Without it every `http_*` builtin is a
|
||||
> `{"error":"not built with HAVE_CURL"}` stub — real HTTP silently disappears.
|
||||
|
||||
Writes (`attend`/`learn`, `POST`) are gated behind `SWARM_WRITE_HEALTHY=1` and the
|
||||
api-reshape agent's gate-1 write-healthy clone; the proven run is read-cognition.
|
||||
|
||||
## Built vs stubbed (honest)
|
||||
|
||||
**Real, tested:**
|
||||
- Native-thread fan-out/converge, bounded concurrency, order-preserving rejoin.
|
||||
- All three containment rules enforced (scope tokens + lateral-edge check).
|
||||
- CCR per-worker context: retrieval → scoping → compaction, bounded, non-leaking
|
||||
(a worker never receives sibling inputs) — verified against the live isolated mind.
|
||||
- Full durable work-tracking (JSONL journal, reconstructable report).
|
||||
- Four convergence strategies + integer failure threshold / partial-abort.
|
||||
|
||||
**Seam / not yet bound:**
|
||||
- `primitives.el` `think` is a deterministic, hermetic transform (no model call).
|
||||
Binding point is marked `PRIMITIVE_BINDING`; wire to the API-surface reshape's
|
||||
`think/act/attend/intend/learn` when it lands.
|
||||
- Blueprints are dispatched by name in `swarm_run_blueprint` (default +
|
||||
`classify`/`faildemo` demos). A YAML process-definition loader (Swarm §5) is
|
||||
future work — the runtime contract is in place.
|
||||
- Distributed placement (cloud/edge/federated topologies, Swarm §9.2) is
|
||||
structurally enabled by containment but not yet wired to a placement layer;
|
||||
today all workers are local native threads.
|
||||
- Engram work-tracking mirror is opt-in; the durable substrate is the journal.
|
||||
|
||||
@@ -1,61 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# build.sh — compile an El program that uses the swarm capability.
|
||||
#
|
||||
# Concatenates the El native-concurrency stdlib (thread.el, channel.el) and the
|
||||
# swarm capability modules in dependency order, then the user program, compiles
|
||||
# with the canonical elc, and links against the shared C runtime.
|
||||
#
|
||||
# Usage:
|
||||
# swarm/build.sh <program.el> <out-binary>
|
||||
#
|
||||
# The swarm modules use only el_runtime.c builtins plus thread.el/channel.el,
|
||||
# so nothing else needs concatenating (engram_*, json_*, str_*, fs_*, http_*,
|
||||
# uuid_v4, now_millis are all C builtins in el_runtime.c).
|
||||
|
||||
set -uo pipefail
|
||||
cd "$(dirname "$0")/.." # -> lang/
|
||||
LANG_DIR="$(pwd)"
|
||||
ELC="${ELC:-${LANG_DIR}/dist/platform/elc}"
|
||||
RT="${LANG_DIR}/el-compiler/runtime"
|
||||
|
||||
PROG="${1:?usage: build.sh <program.el> <out-binary>}"
|
||||
OUT="${2:?usage: build.sh <program.el> <out-binary>}"
|
||||
|
||||
# swarm module load order (each may depend on those before it):
|
||||
# worktrack — durable work-tracking journal (no swarm deps)
|
||||
# containment — the three containment rules (no swarm deps)
|
||||
# primitives — think/act/attend/intend/learn seam (no swarm deps)
|
||||
# ccr — per-worker compiled bounded context (depends: primitives)
|
||||
# swarm — orchestrator: fan-out/converge (depends: all above + thread)
|
||||
SWARM_MODULES="
|
||||
swarm/worktrack.el
|
||||
swarm/containment.el
|
||||
swarm/primitives.el
|
||||
swarm/reshape_surface.el
|
||||
swarm/primitive_binding.el
|
||||
swarm/primitive_seam.el
|
||||
swarm/ccr.el
|
||||
swarm/swarm.el
|
||||
"
|
||||
|
||||
TMP_C="$(mktemp -t swarm_build.XXXXXX).c"
|
||||
COMBINED="$(mktemp -t swarm_combined.XXXXXX).el"
|
||||
|
||||
cat runtime/thread.el runtime/channel.el $SWARM_MODULES "$PROG" > "$COMBINED"
|
||||
|
||||
if ! "$ELC" "$COMBINED" > "$TMP_C" 2>/tmp/swarm.elc.err; then
|
||||
echo "elc FAILED:" >&2
|
||||
sed 's/^/ /' /tmp/swarm.elc.err >&2
|
||||
rm -f "$TMP_C" "$COMBINED"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! cc -O2 -DHAVE_CURL -I "$RT" "$TMP_C" "$RT/el_runtime.c" -lcurl -lpthread -lm -o "$OUT" 2>/tmp/swarm.cc.err; then
|
||||
echo "cc FAILED:" >&2
|
||||
sed 's/^/ /' /tmp/swarm.cc.err >&2
|
||||
rm -f "$TMP_C" "$COMBINED"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
rm -f "$TMP_C" "$COMBINED"
|
||||
echo "built: $OUT"
|
||||
@@ -1,153 +0,0 @@
|
||||
// ccr.el — Compiled Context Routing for work distribution.
|
||||
//
|
||||
// The same spine as the API's vantage-read, applied per worker. Instead of
|
||||
// handing every worker the coordinator's full memory, CCR compiles a MINIMAL,
|
||||
// BOUNDED context package scoped to exactly one worker's input (CCR §5, "Compiled
|
||||
// Context Injection"; Swarm §9.3, "The Compiled Context Boundary as Security
|
||||
// Boundary").
|
||||
//
|
||||
// The pipeline is CCR §5.1: Retrieval -> Scoping -> Compilation -> (Injection,
|
||||
// which here is placing the package into the worker's task envelope).
|
||||
//
|
||||
// 1. Retrieval — resolve the blueprint's knowledge refs + the input's salient
|
||||
// terms against the mind (primitive_attend).
|
||||
// 2. Scoping — keep only what THIS input needs; drop everything else. A
|
||||
// worker never receives sibling inputs or unrelated memory.
|
||||
// 3. Compilation— compact to a CTX string within a token budget (lossless of
|
||||
// meaning, smaller in tokens): collapse blank runs, dedupe
|
||||
// lines, then bound to the budget.
|
||||
//
|
||||
// The package a worker receives is therefore (a) sufficient for its task and
|
||||
// (b) incapable of leaking what it was never given — the containment boundary
|
||||
// and the security boundary are the same object.
|
||||
|
||||
// ── token budget helpers ─────────────────────────────────────────────────────
|
||||
|
||||
// ccr_est_tokens — cheap token estimate (~4 chars/token).
|
||||
fn ccr_est_tokens(s: String) -> Int {
|
||||
return str_len(s) / 4
|
||||
}
|
||||
|
||||
// ccr_default_budget — default per-worker context budget in tokens.
|
||||
// Override with CCR_TOKEN_BUDGET.
|
||||
fn ccr_default_budget() -> Int {
|
||||
let b: String = env("CCR_TOKEN_BUDGET")
|
||||
if str_eq(b, "") {
|
||||
return 1200
|
||||
}
|
||||
return str_to_int(b)
|
||||
}
|
||||
|
||||
// ── stage 3: compaction ──────────────────────────────────────────────────────
|
||||
|
||||
// ccr_compact — collapse blank-line runs and drop exact duplicate lines, then
|
||||
// bound the result to `budget` tokens (truncate on a line boundary). Meaning is
|
||||
// preserved; token count falls (CCR §5.2).
|
||||
fn ccr_compact(text: String, budget: Int) -> String {
|
||||
let lines: [String] = str_split_lines(text)
|
||||
let n: Int = el_list_len(lines)
|
||||
let seen: String = "\n"
|
||||
let out: String = ""
|
||||
let out_tokens = 0
|
||||
let i = 0
|
||||
while i < n {
|
||||
let ln: String = str_trim(el_list_get(lines, i))
|
||||
if str_eq(ln, "") {
|
||||
let i = i + 1
|
||||
} else {
|
||||
let marker: String = "\n" + ln + "\n"
|
||||
if str_contains(seen, marker) {
|
||||
// duplicate line — skip
|
||||
let i = i + 1
|
||||
} else {
|
||||
let seen = seen + ln + "\n"
|
||||
let line_tokens: Int = ccr_est_tokens(ln) + 1
|
||||
if out_tokens + line_tokens > budget {
|
||||
// budget exhausted — stop (bounded)
|
||||
let i = n
|
||||
} else {
|
||||
let out = out + ln + "\n"
|
||||
let out_tokens = out_tokens + line_tokens
|
||||
let i = i + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ── stages 1+2: retrieve + scope ─────────────────────────────────────────────
|
||||
|
||||
// ccr_retrieve_scoped — pull context relevant to this input and its blueprint
|
||||
// knowledge refs, scoped to a fraction of the budget so no single source floods
|
||||
// the package. Returns compacted retrieved text (may be empty if the mind is
|
||||
// unreachable — the input alone is still a valid minimal context).
|
||||
fn ccr_retrieve_scoped(blueprint: String, knowledge_refs: String, input_item: String, budget: Int) -> String {
|
||||
let acc: String = ""
|
||||
// knowledge_refs is a JSON array of query strings.
|
||||
let m: Int = json_array_len(knowledge_refs)
|
||||
let i = 0
|
||||
while i < m {
|
||||
let ref: String = json_array_get_string(knowledge_refs, i)
|
||||
let hit: String = primitive_attend(ref, 3)
|
||||
let acc = acc + "# ref:" + ref + "\n" + hit + "\n"
|
||||
let i = i + 1
|
||||
}
|
||||
// the input's own salient text also seeds retrieval
|
||||
let hit2: String = primitive_attend(input_item, 3)
|
||||
let acc = acc + "# input-context\n" + hit2 + "\n"
|
||||
// scope retrieval to ~60% of budget; the input itself gets the rest
|
||||
let retr_budget: Int = (budget * 6) / 10
|
||||
return ccr_compact(acc, retr_budget)
|
||||
}
|
||||
|
||||
// ── ccr_compile — assemble the bounded per-worker context package ─────────────
|
||||
//
|
||||
// blueprint : task blueprint name
|
||||
// knowledge_refs : JSON array of retrieval queries from the blueprint
|
||||
// input_item : THIS worker's single input (and nothing else)
|
||||
// corr_id : swarm correlation ID
|
||||
// worker_id : this worker's ID
|
||||
// scope_token : the worker's containment token (closed boundary)
|
||||
//
|
||||
// Returns a JSON package: { blueprint, corr_id, worker_id, scope_token,
|
||||
// input, knowledge, budget_tokens, compiled_tokens }. `knowledge` is compiled
|
||||
// and bounded; the package as a whole is bounded by budget.
|
||||
fn ccr_compile(blueprint: String, knowledge_refs: String, input_item: String,
|
||||
corr_id: String, worker_id: String, scope_token: String) -> String {
|
||||
let budget: Int = ccr_default_budget()
|
||||
let knowledge: String = ccr_retrieve_scoped(blueprint, knowledge_refs, input_item, budget)
|
||||
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "blueprint")
|
||||
let kv = el_list_append(kv, blueprint)
|
||||
let kv = el_list_append(kv, "corr_id")
|
||||
let kv = el_list_append(kv, corr_id)
|
||||
let kv = el_list_append(kv, "worker_id")
|
||||
let kv = el_list_append(kv, worker_id)
|
||||
let kv = el_list_append(kv, "input")
|
||||
let kv = el_list_append(kv, input_item)
|
||||
let kv = el_list_append(kv, "knowledge")
|
||||
let kv = el_list_append(kv, knowledge)
|
||||
let kv = el_list_append(kv, "budget_tokens")
|
||||
let kv = el_list_append(kv, int_to_str(budget))
|
||||
let pkg: String = json_build_object(kv)
|
||||
// stamp the scope token as a nested object, and the measured size
|
||||
let pkg2: String = json_set(pkg, "scope_token", scope_token)
|
||||
let compiled_tokens: Int = ccr_est_tokens(pkg2)
|
||||
let pkg3: String = json_set(pkg2, "compiled_tokens", int_to_str(compiled_tokens))
|
||||
return pkg3
|
||||
}
|
||||
|
||||
// ccr_within_budget — did the compiled package stay within its budget?
|
||||
// (Retrieval is bounded to 60% and the input is small; this asserts the whole
|
||||
// package is bounded — the property distribution relies on.)
|
||||
fn ccr_within_budget(pkg: String) -> Bool {
|
||||
let budget: Int = str_to_int(json_get_string(pkg, "budget_tokens"))
|
||||
let compiled: Int = str_to_int(json_get_string(pkg, "compiled_tokens"))
|
||||
// allow a small envelope for JSON framing overhead
|
||||
if compiled <= budget + 200 {
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -1,181 +0,0 @@
|
||||
// containment.el — the Swarm Architecture containment rules, enforced.
|
||||
//
|
||||
// "These rules are not conventions. They are enforced by the runtime."
|
||||
// (Swarm Architecture §3.2). The three rules that make bounded parallelism —
|
||||
// and therefore location-independent distribution — safe:
|
||||
//
|
||||
// Rule 1: a worker may NOT join another swarm.
|
||||
// Rule 2: a worker may NOT initiate a new swarm.
|
||||
// Rule 3: a worker may NOT communicate laterally with sibling workers.
|
||||
//
|
||||
// Enforcement is by SCOPE TOKEN. When a swarm fans out, the coordinator mints a
|
||||
// swarm scope token and stamps a distinct worker scope token into each worker's
|
||||
// task envelope. Any attempt to create or join a swarm checks the caller's
|
||||
// token: if the caller already holds a WORKER token, the operation is rejected.
|
||||
// Rule 3 is enforced structurally elsewhere — workers share no mutable state and
|
||||
// the only channels they hold are the vertical result path — but this module
|
||||
// provides the explicit lateral-edge check for the execution tree.
|
||||
//
|
||||
// A scope token is a JSON object: {"kind":"coordinator|worker","swarm":"<corr>",
|
||||
// "worker":"<id-or-empty>","depth":"<n>"}.
|
||||
|
||||
// ── Token minting ────────────────────────────────────────────────────────────
|
||||
|
||||
// CAPABILITIES. A scope token carries a `caps` set — the authority it holds.
|
||||
// This is an AUTHORITY gate, not a health gate: capability is decided at mint
|
||||
// time and cannot be acquired at runtime. Engram-WRITE (op_write/op_relate/
|
||||
// op_supersede -> POST /api/nodes, /api/edges, DELETE) and dharma_emit are
|
||||
// @manager-ONLY capabilities — exactly the VBD rule that only the orchestrator
|
||||
// mutates global state. The orchestrator's token carries them; a worker's token
|
||||
// NEVER does. A worker is therefore STRUCTURALLY UNABLE to mutate global engram
|
||||
// state, regardless of engram health.
|
||||
fn cap_orchestrator() -> String { return "engram:read,engram:write,dharma:emit,state:write" }
|
||||
fn cap_worker() -> String { return "engram:read" }
|
||||
|
||||
// containment_coordinator_token — the token the orchestrator (@manager) holds.
|
||||
// Depth 0. Carries the engram-WRITE + dharma-emit capabilities (@manager-only).
|
||||
fn containment_coordinator_token(corr_id: String) -> String {
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "kind")
|
||||
let kv = el_list_append(kv, "coordinator")
|
||||
let kv = el_list_append(kv, "swarm")
|
||||
let kv = el_list_append(kv, corr_id)
|
||||
let kv = el_list_append(kv, "worker")
|
||||
let kv = el_list_append(kv, "")
|
||||
let kv = el_list_append(kv, "depth")
|
||||
let kv = el_list_append(kv, "0")
|
||||
let kv = el_list_append(kv, "caps")
|
||||
let kv = el_list_append(kv, cap_orchestrator())
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// containment_worker_token — the token stamped into a worker's envelope. Depth 1.
|
||||
// A closed boundary: forbids opening/joining swarms AND carries ONLY the
|
||||
// engram:READ capability — no engram:write, no dharma:emit. Read-only against the
|
||||
// full engram; may write only its own local geometry (its returned result).
|
||||
fn containment_worker_token(corr_id: String, worker_id: String) -> String {
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "kind")
|
||||
let kv = el_list_append(kv, "worker")
|
||||
let kv = el_list_append(kv, "swarm")
|
||||
let kv = el_list_append(kv, corr_id)
|
||||
let kv = el_list_append(kv, "worker")
|
||||
let kv = el_list_append(kv, worker_id)
|
||||
let kv = el_list_append(kv, "depth")
|
||||
let kv = el_list_append(kv, "1")
|
||||
let kv = el_list_append(kv, "caps")
|
||||
let kv = el_list_append(kv, cap_worker())
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// containment_has_cap — does this token carry capability `cap`?
|
||||
fn containment_has_cap(token: String, cap: String) -> Bool {
|
||||
return str_contains(json_get_string(token, "caps"), cap)
|
||||
}
|
||||
|
||||
// ── Rule checks (return "" on allow, or a rejection reason string) ───────────
|
||||
|
||||
// containment_check_open — may the holder of `token` OPEN a new swarm?
|
||||
// Enforces Rule 2 (a worker may not initiate a new swarm). Only a coordinator
|
||||
// token, or an absent token (top-level process), may open one.
|
||||
fn containment_check_open(token: String) -> String {
|
||||
if str_eq(token, "") {
|
||||
return ""
|
||||
}
|
||||
let kind: String = json_get_string(token, "kind")
|
||||
if str_eq(kind, "worker") {
|
||||
return "CONTAINMENT rule 2: a swarm worker may not initiate a new swarm (worker=" + json_get_string(token, "worker") + " swarm=" + json_get_string(token, "swarm") + ")"
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// containment_check_join — may the holder of `token` JOIN swarm `target_corr`?
|
||||
// Enforces Rule 1 (a worker may not join another swarm). A worker already bound
|
||||
// to swarm A may not register into swarm B; and a worker may not re-join at all.
|
||||
fn containment_check_join(token: String, target_corr: String) -> String {
|
||||
if str_eq(token, "") {
|
||||
return ""
|
||||
}
|
||||
let kind: String = json_get_string(token, "kind")
|
||||
if str_eq(kind, "worker") {
|
||||
return "CONTAINMENT rule 1: a swarm worker may not join another swarm (worker=" + json_get_string(token, "worker") + " bound-swarm=" + json_get_string(token, "swarm") + " attempted-swarm=" + target_corr + ")"
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// containment_check_lateral — may `from_token` open a communication edge to a
|
||||
// sibling worker `to_worker_id`? Enforces Rule 3 (no lateral communication).
|
||||
// The only permitted edges are vertical: worker->coordinator and
|
||||
// coordinator->worker. Any worker->worker edge is rejected.
|
||||
fn containment_check_lateral(from_token: String, to_worker_id: String) -> String {
|
||||
let kind: String = json_get_string(from_token, "kind")
|
||||
if str_eq(kind, "worker") {
|
||||
if str_eq(to_worker_id, "") {
|
||||
// empty target = the coordinator (vertical) — allowed
|
||||
return ""
|
||||
}
|
||||
return "CONTAINMENT rule 3: a swarm worker may not communicate laterally with sibling workers (from=" + json_get_string(from_token, "worker") + " to=" + to_worker_id + ")"
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// containment_check_engram_write — RULE 4: only a token carrying the
|
||||
// engram:write capability (the orchestrator's) may mutate global engram state.
|
||||
// A worker token (engram:read only) is REJECTED — the authority gate. Reuses the
|
||||
// exact scope-token mechanism as Rule 2's open-denial. Returns "" on allow, or a
|
||||
// rejection reason. This is an AUTHORITY gate: it does not consult engram health.
|
||||
fn containment_check_engram_write(token: String, op: String) -> String {
|
||||
if containment_has_cap(token, "engram:write") {
|
||||
return ""
|
||||
}
|
||||
return "CONTAINMENT rule 4: engram-write is @manager-only — a worker is read-only against the engram and may not mutate global state (op=" + op + " kind=" + json_get_string(token, "kind") + " worker=" + json_get_string(token, "worker") + " caps=" + json_get_string(token, "caps") + ")"
|
||||
}
|
||||
|
||||
// containment_check_dharma_emit — the same @manager-only rule for dharma_emit,
|
||||
// grounding Rule 4 in VBD: global-state mutations (engram-write, dharma-emit) are
|
||||
// orchestrator-only, checked by the one capability mechanism.
|
||||
fn containment_check_dharma_emit(token: String) -> String {
|
||||
if containment_has_cap(token, "dharma:emit") {
|
||||
return ""
|
||||
}
|
||||
return "CONTAINMENT rule 4: dharma_emit is @manager-only (kind=" + json_get_string(token, "kind") + ")"
|
||||
}
|
||||
|
||||
// ── Enforcement helpers ──────────────────────────────────────────────────────
|
||||
|
||||
// containment_allows_open — Bool convenience over containment_check_open.
|
||||
fn containment_allows_open(token: String) -> Bool {
|
||||
return str_eq(containment_check_open(token), "")
|
||||
}
|
||||
|
||||
// containment_is_worker — is this a worker-scoped (closed-boundary) token?
|
||||
fn containment_is_worker(token: String) -> Bool {
|
||||
return str_eq(json_get_string(token, "kind"), "worker")
|
||||
}
|
||||
|
||||
// containment_guard_open — assert a swarm may be opened under this token.
|
||||
// Returns "" if allowed, or records a CONTAINMENT violation to the work-tracking
|
||||
// journal and returns the reason. Callers must abort on a non-empty return.
|
||||
fn containment_guard_open(token: String, corr_id: String) -> String {
|
||||
let reason: String = containment_check_open(token)
|
||||
if str_eq(reason, "") {
|
||||
return ""
|
||||
}
|
||||
let p: String = json_set_str("{}", "reason", reason)
|
||||
worktrack_append("containment.violation", corr_id, "open", p)
|
||||
return reason
|
||||
}
|
||||
|
||||
// containment_guard_engram_write — assert a token may mutate global engram state
|
||||
// (Rule 4). Returns "" if allowed; otherwise journals a containment.violation and
|
||||
// returns the reason. The write path MUST abort on a non-empty return.
|
||||
fn containment_guard_engram_write(token: String, corr_id: String, op: String) -> String {
|
||||
let reason: String = containment_check_engram_write(token, op)
|
||||
if str_eq(reason, "") {
|
||||
return ""
|
||||
}
|
||||
let p0: String = json_set_str("{}", "reason", reason)
|
||||
let p1: String = json_set_str(p0, "op", op)
|
||||
worktrack_append("containment.violation", corr_id, "engram-write", p1)
|
||||
return reason
|
||||
}
|
||||
@@ -1,49 +0,0 @@
|
||||
// primitive_binding.el — THE ONE FLIP POINT.
|
||||
//
|
||||
// This file is the single seam between the swarm and the real agentic
|
||||
// primitives. Binding the reshape's decorated primitives is a one-line change
|
||||
// HERE and nothing else changes anywhere in the swarm.
|
||||
//
|
||||
// The api-reshape agent (wt/api-reshape) is wiring the primitives as DECORATED
|
||||
// El on the dharma_* event bus over the engram — think/attend/learn/ground/assert
|
||||
// become decorated fns that emit afferent events onto the bus. The moment they
|
||||
// land, flip `bound_think` (and its siblings) to call them.
|
||||
//
|
||||
// TODAY (stub fallback, compiles + runs now against :8901):
|
||||
// fn bound_think(...) { return primitive_think(ctx, instruction) }
|
||||
//
|
||||
// THE FLIP (when reshape's decorated primitives land — one line each):
|
||||
// fn bound_think(...) { return think(ctx, instruction) } // decorated, on dharma bus
|
||||
//
|
||||
// Keep the stub as fallback: `bound_think` is only reached when the seam mode is
|
||||
// "decorated" (SWARM_PRIMITIVE_SEAM=decorated). Until you flip these bodies AND
|
||||
// set that env, the harness runs entirely on the hermetic stub.
|
||||
|
||||
// bound_think — BOUND to the reshape's proven decorated `think` (op_think),
|
||||
// real cognition over the engram geometry. The worker's CCR slice carries a
|
||||
// NODE-ID anchor in ctx.input (free-text anchors return "geometry unavailable");
|
||||
// think re-origins at that node's region under the faculty and returns a real
|
||||
// 768-dim gradient.
|
||||
fn bound_think(ctx: String, instruction: String) -> String {
|
||||
let anchor: String = json_get_string(ctx, "input")
|
||||
let faculty: String = json_get_string(ctx, "faculty")
|
||||
return op_think(anchor, faculty)
|
||||
}
|
||||
|
||||
// bound_attend — BOUND to the reshape's op_attend (POST /api/attend). Needs the
|
||||
// gate-1 write-healthy clone; falls back to the read-side attend otherwise.
|
||||
fn bound_attend(query: String, limit: Int) -> String {
|
||||
if str_eq(env("SWARM_WRITE_HEALTHY"), "1") {
|
||||
return op_attend(query, "self")
|
||||
}
|
||||
return primitive_attend(query, limit)
|
||||
}
|
||||
|
||||
// bound_learn — BOUND to the reshape's op_learn (correspondence-beat). Needs the
|
||||
// gate-1 write-healthy clone; falls back to the opt-in journal-only learn.
|
||||
fn bound_learn(corr_id: String, observation: String) -> String {
|
||||
if str_eq(env("SWARM_WRITE_HEALTHY"), "1") {
|
||||
return op_learn(observation, "induce")
|
||||
}
|
||||
return primitive_learn(corr_id, observation)
|
||||
}
|
||||
@@ -1,55 +0,0 @@
|
||||
// primitive_seam.el — the configurable primitive seam + telemetry.
|
||||
//
|
||||
// One switch selects where a worker's primitive invocation goes:
|
||||
// SWARM_PRIMITIVE_SEAM=stub (default) — hermetic in-process think.
|
||||
// SWARM_PRIMITIVE_SEAM=decorated — the reshape's decorated
|
||||
// primitives on the dharma bus
|
||||
// (see primitive_binding.el).
|
||||
//
|
||||
// Every seam invocation is an AFFERENT signal — a primitive call travelling
|
||||
// toward the manager. The seam stamps telemetry onto each thought (seam_mode +
|
||||
// one afferent tick) so the coordinator can aggregate afferent counters across
|
||||
// the swarm without any shared mutable state (containment-safe: counts ride the
|
||||
// vertical result path, not a shared bus register).
|
||||
|
||||
// seam_mode — "stub" (default) or "decorated".
|
||||
fn seam_mode() -> String {
|
||||
let m: String = env("SWARM_PRIMITIVE_SEAM")
|
||||
if str_eq(m, "decorated") {
|
||||
return "decorated"
|
||||
}
|
||||
return "stub"
|
||||
}
|
||||
|
||||
// seam_think — route a worker's `think` through the configured seam and stamp
|
||||
// telemetry. Returns the thought JSON augmented with:
|
||||
// seam_mode : which side of the seam served this call
|
||||
// afferent : "1" — one afferent primitive signal was emitted
|
||||
fn seam_think(ctx: String, instruction: String) -> String {
|
||||
let mode: String = seam_mode()
|
||||
let thought: String = ""
|
||||
if str_eq(mode, "decorated") {
|
||||
let thought = bound_think(ctx, instruction)
|
||||
} else {
|
||||
let thought = primitive_think(ctx, instruction)
|
||||
}
|
||||
let t1: String = json_set_str(thought, "seam_mode", mode)
|
||||
let t2: String = json_set_str(t1, "afferent", "1")
|
||||
return t2
|
||||
}
|
||||
|
||||
// seam_attend / seam_learn — same seam for the other primitives (used when a
|
||||
// blueprint retrieves or writes through the bus).
|
||||
fn seam_attend(query: String, limit: Int) -> String {
|
||||
if str_eq(seam_mode(), "decorated") {
|
||||
return bound_attend(query, limit)
|
||||
}
|
||||
return primitive_attend(query, limit)
|
||||
}
|
||||
|
||||
fn seam_learn(corr_id: String, observation: String) -> String {
|
||||
if str_eq(seam_mode(), "decorated") {
|
||||
return bound_learn(corr_id, observation)
|
||||
}
|
||||
return primitive_learn(corr_id, observation)
|
||||
}
|
||||
@@ -1,94 +0,0 @@
|
||||
// primitives.el — the agentic primitive SEAM the swarm composes over.
|
||||
//
|
||||
// The swarm is orchestration OVER the five CCR primitives, not a replacement for
|
||||
// them (CCR §2, "The Five Primitives / The Execution Cycle"): a worker executes
|
||||
// its task blueprint as attend -> think -> intend -> act -> learn against its
|
||||
// compiled, bounded context.
|
||||
//
|
||||
// This file is the SEAM. The parallel API-surface reshape exposes the canonical
|
||||
// primitive tools; when it lands, bind each primitive below to the reshaped
|
||||
// implementation (see PRIMITIVE_BINDING). Until then these are thin, engram-
|
||||
// backed fallbacks so the swarm — its fan-out, containment, CCR context
|
||||
// compilation, convergence, and work-tracking — is fully exercisable today.
|
||||
//
|
||||
// Contract: every primitive takes and returns String (JSON where structured), so
|
||||
// any primitive is directly threadable via thread.el's spawn (which runs
|
||||
// top-level (String)->String El fns).
|
||||
//
|
||||
// PRIMITIVE_BINDING: to bind the reshape's real tools, replace each fallback body
|
||||
// with a call to the reshaped El fn / API endpoint. Signatures here are the
|
||||
// stable contract the swarm depends on; keep them.
|
||||
|
||||
// ── attend — retrieve the minimal relevant context for a focus ───────────────
|
||||
// Vantage-read: pull only what this focus needs from the mind. Backed by the
|
||||
// engram's spreading-activation retrieval.
|
||||
fn primitive_attend(query: String, limit: Int) -> String {
|
||||
if str_eq(query, "") {
|
||||
return "[]"
|
||||
}
|
||||
// Location-independent worker model: when an engram daemon is configured,
|
||||
// retrieve over HTTP (the worker may run anywhere). POST /api/search
|
||||
// {query,limit,_auth}. Falls back to the in-process store otherwise.
|
||||
let url: String = env("ENGRAM_URL")
|
||||
if str_eq(url, "") {
|
||||
return engram_activate(query, limit)
|
||||
}
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "query")
|
||||
let kv = el_list_append(kv, query)
|
||||
let body0: String = json_build_object(kv)
|
||||
let body1: String = json_set(body0, "limit", int_to_str(limit))
|
||||
let body2: String = json_set_str(body1, "_auth", env("ENGRAM_API_KEY"))
|
||||
return http_post(url + "/api/search", body2)
|
||||
}
|
||||
|
||||
// ── think — reason over the compiled context ─────────────────────────────────
|
||||
// In production this routes to a model (CCR dynamic model selection). Here it is
|
||||
// a deterministic, hermetic transform so swarm behaviour is testable without an
|
||||
// external model: it echoes a structured verdict derived from the context. The
|
||||
// binding point for a real model is explicit.
|
||||
fn primitive_think(compiled_ctx: String, instruction: String) -> String {
|
||||
// PRIMITIVE_BINDING: replace with the reshape's think() (model inference).
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "instruction")
|
||||
let kv = el_list_append(kv, instruction)
|
||||
let kv = el_list_append(kv, "ctx_bytes")
|
||||
let kv = el_list_append(kv, int_to_str(str_len(compiled_ctx)))
|
||||
let kv = el_list_append(kv, "conclusion")
|
||||
let kv = el_list_append(kv, "reasoned:" + instruction)
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// ── intend — form a bounded plan/decision from a thought ─────────────────────
|
||||
fn primitive_intend(thought: String) -> String {
|
||||
let concl: String = json_get_string(thought, "conclusion")
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "intent")
|
||||
let kv = el_list_append(kv, concl)
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// ── act — execute a bounded effect and return its result ─────────────────────
|
||||
// Workers defer real side-effects to the coordinator (idempotency requirement,
|
||||
// Swarm §7.3). Here act produces an artifact-shaped result the coordinator
|
||||
// collects during convergence.
|
||||
fn primitive_act(intent: String, input_item: String) -> String {
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "acted_on")
|
||||
let kv = el_list_append(kv, input_item)
|
||||
let kv = el_list_append(kv, "via")
|
||||
let kv = el_list_append(kv, json_get_string(intent, "intent"))
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// ── learn — record an observation into the mind, tagged by correlation ID ────
|
||||
// Append-only, naturally idempotent (Swarm §7.3). Best-effort: a worker that
|
||||
// cannot reach the mind still returns its result.
|
||||
fn primitive_learn(corr_id: String, observation: String) -> String {
|
||||
let url: String = env("ENGRAM_URL")
|
||||
if str_eq(url, "") {
|
||||
return ""
|
||||
}
|
||||
let content: String = "swarm-worker-obs corr=" + corr_id + " :: " + observation
|
||||
return engram_node(content, "Memory", 0.4)
|
||||
}
|
||||
@@ -1,103 +0,0 @@
|
||||
// reshape_surface.el — the api-reshape agent's PROVEN decorated primitives,
|
||||
// composed into the swarm build to bind real cognition.
|
||||
//
|
||||
// PROVENANCE: these fns are the reshape's surface at wt/api-reshape @ d4f401d
|
||||
// ("reshape: decorator-as-seam — port @route codegen, prove decorate->serve,
|
||||
// rewrite surface as decorated El"), verified live against
|
||||
// engram.cognition-20260814. Copied verbatim (read/cognition ops only) so the
|
||||
// swarm binds the REAL primitives, not a reimplementation. The write ops
|
||||
// (op_write/op_relate/op_supersede/op_ground) are intentionally NOT composed
|
||||
// here — they exercise the persist_node write path that needs the gate-1
|
||||
// write-healthy clone; the swarm's proven run is read-cognition (think/read).
|
||||
//
|
||||
// Ops route to the ENGRAM over ENGRAM_URL — pinned by THIS worktree's .nsbx-env
|
||||
// to the :8901 swarm clone (never the reshape agent's :8900). Separate clones,
|
||||
// no collision.
|
||||
|
||||
fn engram_url() -> String {
|
||||
let u: String = env("ENGRAM_URL")
|
||||
if str_eq(u, "") { return "http://127.0.0.1:8900" }
|
||||
return u
|
||||
}
|
||||
fn engram_key() -> String {
|
||||
let k: String = env("ENGRAM_API_KEY")
|
||||
if str_eq(k, "") { return "sbx-dev-api-reshape" }
|
||||
return k
|
||||
}
|
||||
fn SELF_KEY() -> String { return "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee" }
|
||||
fn VALUES_KEY() -> String { return "kn-5b606390-a52d-4ca2-8e0e-eba141d13440" }
|
||||
|
||||
// self/values name -> keystone id; anything else passes through unchanged.
|
||||
fn resolve_named(v: String) -> String {
|
||||
if str_eq(v, "self") { return SELF_KEY() }
|
||||
if str_eq(v, "neuron") { return SELF_KEY() }
|
||||
if str_eq(v, "values") { return VALUES_KEY() }
|
||||
if str_eq(v, "values_hub") { return VALUES_KEY() }
|
||||
return v
|
||||
}
|
||||
|
||||
// read — THE VANTAGE-READ. Re-origin at a point + aperture -> a BOUNDED slice.
|
||||
fn op_read(vantage: String, typ: String, k: Int) -> String {
|
||||
let vid: String = resolve_named(vantage)
|
||||
if str_eq(typ, "edges") {
|
||||
return http_get(engram_url() + "/api/neighbors/" + vid)
|
||||
}
|
||||
if str_starts_with(vid, "kn-") {
|
||||
return http_get(engram_url() + "/api/neighbors/" + vid)
|
||||
}
|
||||
return http_get(engram_url() + "/api/search?q=" + url_encode(vid) + "&limit=" + int_to_str(k))
|
||||
}
|
||||
|
||||
// think — THE ONE OPERATION. anchor (node ids) steered by faculty -> gradient.
|
||||
fn op_think(seeds: String, faculty: String) -> String {
|
||||
let s: String = resolve_named(seeds)
|
||||
let f: String = if str_eq(faculty, "") { "reason" } else { faculty }
|
||||
return http_get(engram_url() + "/api/think?seeds=" + url_encode(s) + "&faculty=" + f)
|
||||
}
|
||||
|
||||
// attend — aim attention at a region. (POST — needs a write-healthy clone.)
|
||||
fn op_attend(node: String, observer: String) -> String {
|
||||
let n: String = resolve_named(node)
|
||||
let o: String = if str_eq(observer, "") { SELF_KEY() } else { resolve_named(observer) }
|
||||
let body: String = "{\"_auth\":\"" + engram_key() + "\",\"node\":\"" + n
|
||||
+ "\",\"observer\":\"" + o + "\",\"salience\":\"0.6\"}"
|
||||
return http_post_json(engram_url() + "/api/attend", body)
|
||||
}
|
||||
|
||||
fn identity_typed(t: String) -> Bool {
|
||||
if str_eq(t, "self") { return true }
|
||||
if str_eq(t, "values") { return true }
|
||||
return false
|
||||
}
|
||||
fn type_to_node_type(t: String) -> String {
|
||||
if str_eq(t, "knowledge") { return "Knowledge" }
|
||||
if str_eq(t, "artifact") { return "Artifact" }
|
||||
if str_eq(t, "backlog") { return "WorkItem" }
|
||||
if str_eq(t, "process") { return "Process" }
|
||||
if str_eq(t, "state") { return "InternalStateEvent" }
|
||||
return "Memory"
|
||||
}
|
||||
|
||||
// write — add a node (POST /api/nodes). Identity types refused. This is a
|
||||
// global-engram MUTATION — @manager-only (Rule 4); never called on a worker path.
|
||||
// (Reshape's op_write, with json_escape -> the available json_escape_string.)
|
||||
fn op_write(content: String, typ: String, importance: Float) -> String {
|
||||
if str_eq(content, "") { return "{\"error\":\"write: content required\"}" }
|
||||
if identity_typed(typ) {
|
||||
return "{\"error\":\"write type=" + typ + " is write-protected -> intentional-cultivation\"}"
|
||||
}
|
||||
let body: String = "{\"_auth\":\"" + engram_key() + "\",\"content\":\"" + json_escape_string(content)
|
||||
+ "\",\"node_type\":\"" + type_to_node_type(typ) + "\",\"tier\":\"Working\",\"importance\":"
|
||||
+ float_to_str(importance) + "}"
|
||||
return http_post_json(engram_url() + "/api/nodes", body)
|
||||
}
|
||||
|
||||
// learn — the reflexive correspondence-beat: calibrate the steering-prior.
|
||||
// (POST — needs a write-healthy clone.)
|
||||
fn op_learn(seeds: String, faculty: String) -> String {
|
||||
let s: String = resolve_named(seeds)
|
||||
let f: String = if str_eq(faculty, "") { "induce" } else { faculty }
|
||||
let body: String = "{\"_auth\":\"" + engram_key() + "\",\"seeds\":\"" + s
|
||||
+ "\",\"faculty\":\"" + f + "\",\"keystone\":\"false\"}"
|
||||
return http_post_json(engram_url() + "/api/correspondence-beat", body)
|
||||
}
|
||||
@@ -1,483 +0,0 @@
|
||||
// swarm.el — the swarm orchestrator: bounded parallel agent execution.
|
||||
//
|
||||
// Implements Swarm Architecture's single pattern — fan out, execute independently,
|
||||
// converge — on El's NATIVE concurrency (thread.el spawn/join). No external
|
||||
// orchestrator: a swarm is a coordinator (this file, the main thread) that mints
|
||||
// a correlation identity, compiles a bounded CCR context per worker, dispatches
|
||||
// workers as native pthreads, tracks every unit of work, and converges the
|
||||
// results before returning control to the parent step.
|
||||
//
|
||||
// The five properties of every swarm (Swarm §2.1) are all present:
|
||||
// parent step -> swarm_run is called from one process step
|
||||
// task blueprint -> `blueprint` name + knowledge refs, run by every worker
|
||||
// input set -> `inputs_json`, one item per worker
|
||||
// convergence -> `strategy` in config (collect|merge|vote|reduce)
|
||||
// correlation ID -> minted here, threaded through tracking + every worker
|
||||
//
|
||||
// Containment (Swarm §3) is enforced: the caller must hold a coordinator/absent
|
||||
// token to open a swarm (Rule 2), each worker is stamped a closed worker token
|
||||
// (Rules 1+3), and workers share no mutable state (the coordinator is the only
|
||||
// journal writer).
|
||||
|
||||
// ── worker entry — the top-level (String)->String fn native threads run ──────
|
||||
//
|
||||
// Every El fn compiles to a global C symbol; spawn() resolves this by name via
|
||||
// dlsym and runs it in a pthread. The envelope carries everything the worker is
|
||||
// permitted to see — its compiled context and nothing else (§9.3).
|
||||
//
|
||||
// Returns a result JSON: {worker_id, status:"completed"|"failed", output|error}.
|
||||
fn swarm_worker_entry(envelope_json: String) -> String {
|
||||
let worker_id: String = json_get_string(envelope_json, "worker_id")
|
||||
let ctx: String = json_get_raw(envelope_json, "ctx")
|
||||
|
||||
// The worker holds a CLOSED worker token (Rules 1+3): it shares no state
|
||||
// with siblings and may not open/join a swarm. That boundary is enforced at
|
||||
// the point of attempt — swarm_run rejects any swarm opened under a worker
|
||||
// token (Rule 2). A worker simply executing its blueprint is not opening a
|
||||
// swarm, so it proceeds. Its only outward edge is this returned result
|
||||
// (the vertical worker->coordinator path).
|
||||
let out: String = swarm_run_blueprint(ctx)
|
||||
// A worker reports failed iff its blueprint signalled failure. This is the
|
||||
// vertical status edge the coordinator reads during convergence (§4.3, §7).
|
||||
let bstatus: String = json_get_string(out, "blueprint_status")
|
||||
let status: String = "completed"
|
||||
if str_eq(bstatus, "failed") {
|
||||
let status = "failed"
|
||||
}
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "worker_id")
|
||||
let kv = el_list_append(kv, worker_id)
|
||||
let kv = el_list_append(kv, "status")
|
||||
let kv = el_list_append(kv, status)
|
||||
let res: String = json_build_object(kv)
|
||||
return json_set(res, "output", out)
|
||||
}
|
||||
|
||||
// swarm_run_blueprint — execute the task blueprint over a compiled context.
|
||||
// The default blueprint is the CCR execution cycle: think -> intend -> act over
|
||||
// the worker's bounded context. Specialise by dispatching on
|
||||
// json_get_string(ctx,"blueprint"). Idempotent: reads ctx, writes only its
|
||||
// returned output (§7.3).
|
||||
fn swarm_run_blueprint(ctx: String) -> String {
|
||||
let blueprint: String = json_get_string(ctx, "blueprint")
|
||||
let input_item: String = json_get_string(ctx, "input")
|
||||
let knowledge: String = json_get_string(ctx, "knowledge")
|
||||
|
||||
// classify — deterministic verdict for the `vote` convergence strategy:
|
||||
// verdict is "long" if the input has >4 chars, else "short".
|
||||
if str_eq(blueprint, "classify") {
|
||||
let verdict: String = "short"
|
||||
if str_len(input_item) > 4 {
|
||||
let verdict = "long"
|
||||
}
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "verdict")
|
||||
let kv = el_list_append(kv, verdict)
|
||||
let kv = el_list_append(kv, "blueprint_status")
|
||||
let kv = el_list_append(kv, "ok")
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// faildemo — a worker that fails on inputs beginning with "x" (exercises the
|
||||
// failure threshold + partial convergence path). Idempotent, side-effect-free.
|
||||
if str_eq(blueprint, "faildemo") {
|
||||
let st: String = "ok"
|
||||
if str_starts_with(input_item, "x") {
|
||||
let st = "failed"
|
||||
}
|
||||
return json_set_str("{}", "blueprint_status", st)
|
||||
}
|
||||
|
||||
// cognize — REAL-COGNITION blueprint. Routes think through the seam (bound to
|
||||
// op_think in decorated mode) over the worker's NODE-ID anchor, then derives a
|
||||
// vote verdict from the gradient's confidence. In stub mode there is no
|
||||
// gradient, so the verdict falls back to a deterministic slice hash — the
|
||||
// same blueprint runs green on either side of the seam.
|
||||
if str_eq(blueprint, "cognize") {
|
||||
let thought: String = seam_think(ctx, "reason over " + input_item)
|
||||
// Derive the vote verdict from the REAL gradient's support count
|
||||
// (json_get_int, since n_support is numeric). Different anchors have
|
||||
// different support -> genuine, cognition-driven vote diversity. In stub
|
||||
// mode there is no gradient (n_support -> 0) -> "uncertain".
|
||||
let nsup: Int = json_get_int(thought, "n_support")
|
||||
let verdict: String = "uncertain"
|
||||
if nsup >= 10 {
|
||||
let verdict = "confident"
|
||||
}
|
||||
let ck: [String] = el_list_empty()
|
||||
let ck = el_list_append(ck, "verdict")
|
||||
let ck = el_list_append(ck, verdict)
|
||||
let ck = el_list_append(ck, "blueprint_status")
|
||||
let ck = el_list_append(ck, "ok")
|
||||
let cout0: String = json_build_object(ck)
|
||||
let cout1: String = json_set_str(cout0, "n_support", int_to_str(nsup))
|
||||
let cout2: String = json_set_str(cout1, "seam_mode", json_get_string(thought, "seam_mode"))
|
||||
return json_set_str(cout2, "afferent", json_get_string(thought, "afferent"))
|
||||
}
|
||||
|
||||
// default (analyze_item): the CCR execution cycle think -> intend -> act,
|
||||
// with `think` routed through the CONFIGURABLE PRIMITIVE SEAM. Telemetry
|
||||
// (seam_mode + afferent tick) rides the worker's returned output.
|
||||
let instruction: String = "process input: " + input_item
|
||||
let thought: String = seam_think(ctx, instruction)
|
||||
let intent: String = primitive_intend(thought)
|
||||
let effect: String = primitive_act(intent, input_item)
|
||||
let e1: String = json_set_str(effect, "blueprint_status", "ok")
|
||||
let e2: String = json_set_str(e1, "seam_mode", json_get_string(thought, "seam_mode"))
|
||||
let e3: String = json_set_str(e2, "afferent", json_get_string(thought, "afferent"))
|
||||
return e3
|
||||
}
|
||||
|
||||
// ── native-thread fan-out, bounded by concurrency, order-preserving ──────────
|
||||
//
|
||||
// parallel_map (thread.el) spawns ALL threads at once. The swarm honours the
|
||||
// blueprint's `concurrency` cap (§5.1: a resource constraint, not a parallelism
|
||||
// constraint — all items are processed, at most N at a time) by dispatching in
|
||||
// waves of N native threads, joining each wave before the next. Results are
|
||||
// returned in input order.
|
||||
fn swarm_fanout(worker_fn: String, envelopes: [String], concurrency: Int) -> [String] {
|
||||
let n: Int = el_list_len(envelopes)
|
||||
let cap: Int = concurrency
|
||||
if cap < 1 {
|
||||
let cap = 1
|
||||
}
|
||||
let results: [String] = el_list_empty()
|
||||
let base = 0
|
||||
while base < n {
|
||||
// spawn a wave of up to `cap` workers
|
||||
let tids: [String] = el_list_empty()
|
||||
let k = 0
|
||||
while k < cap {
|
||||
let idx: Int = base + k
|
||||
if idx < n {
|
||||
let env_item: String = el_list_get(envelopes, idx)
|
||||
let tid: Int = spawn(worker_fn, env_item)
|
||||
let tids = el_list_append(tids, int_to_str(tid))
|
||||
}
|
||||
let k = k + 1
|
||||
}
|
||||
// join the wave in order
|
||||
let j = 0
|
||||
let jn: Int = el_list_len(tids)
|
||||
while j < jn {
|
||||
let tid: Int = str_to_int(el_list_get(tids, j))
|
||||
let r: String = join(tid)
|
||||
let results = el_list_append(results, r)
|
||||
let j = j + 1
|
||||
}
|
||||
let base = base + cap
|
||||
}
|
||||
return results
|
||||
}
|
||||
|
||||
// ── convergence strategies (Swarm §4.2) ──────────────────────────────────────
|
||||
|
||||
// swarm_converge_collect — ordered list, no transformation.
|
||||
fn swarm_converge_collect(results: [String]) -> String {
|
||||
let n: Int = el_list_len(results)
|
||||
let arr: String = "[]"
|
||||
let i = 0
|
||||
while i < n {
|
||||
let arr = json_array_push(arr, el_list_get(results, i))
|
||||
let i = i + 1
|
||||
}
|
||||
return arr
|
||||
}
|
||||
|
||||
// swarm_converge_merge — combine worker outputs into a single joined string.
|
||||
fn swarm_converge_merge(results: [String]) -> String {
|
||||
let n: Int = el_list_len(results)
|
||||
let merged: String = ""
|
||||
let i = 0
|
||||
while i < n {
|
||||
let out: String = json_get_raw(el_list_get(results, i), "output")
|
||||
if i > 0 {
|
||||
let merged = merged + " | "
|
||||
}
|
||||
let merged = merged + out
|
||||
let i = i + 1
|
||||
}
|
||||
return json_set_str("{}", "merged", merged)
|
||||
}
|
||||
|
||||
// swarm_converge_vote — tally a field across worker outputs, pick the majority.
|
||||
// Each worker output is expected to carry a "verdict" string field.
|
||||
fn swarm_converge_vote(results: [String]) -> String {
|
||||
let n: Int = el_list_len(results)
|
||||
// Collect verdicts (no mutable tally: json_set can't update an existing key
|
||||
// and there is no el_list_set). Then count each verdict by rescanning.
|
||||
let verdicts: [String] = el_list_empty()
|
||||
let i = 0
|
||||
while i < n {
|
||||
let out: String = json_get_raw(el_list_get(results, i), "output")
|
||||
let v: String = json_get_string(out, "verdict")
|
||||
if str_eq(v, "") {
|
||||
let i = i + 1
|
||||
} else {
|
||||
let verdicts = el_list_append(verdicts, v)
|
||||
let i = i + 1
|
||||
}
|
||||
}
|
||||
// pick the verdict with the highest count (first-past-the-post)
|
||||
let vn: Int = el_list_len(verdicts)
|
||||
let best: String = ""
|
||||
let bestc = 0
|
||||
let a = 0
|
||||
while a < vn {
|
||||
let cand: String = el_list_get(verdicts, a)
|
||||
// count occurrences of cand
|
||||
let c = 0
|
||||
let b = 0
|
||||
while b < vn {
|
||||
if str_eq(el_list_get(verdicts, b), cand) {
|
||||
let c = c + 1
|
||||
}
|
||||
let b = b + 1
|
||||
}
|
||||
if c > bestc {
|
||||
let bestc = c
|
||||
let best = cand
|
||||
}
|
||||
let a = a + 1
|
||||
}
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "winner")
|
||||
let kv = el_list_append(kv, best)
|
||||
let kv = el_list_append(kv, "votes")
|
||||
let kv = el_list_append(kv, int_to_str(bestc))
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// swarm_converge_reduce — fold outputs into an accumulator (count + concat).
|
||||
fn swarm_converge_reduce(results: [String]) -> String {
|
||||
let n: Int = el_list_len(results)
|
||||
let acc: String = ""
|
||||
let i = 0
|
||||
while i < n {
|
||||
let out: String = json_get_raw(el_list_get(results, i), "output")
|
||||
let acc = acc + out
|
||||
let i = i + 1
|
||||
}
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "count")
|
||||
let kv = el_list_append(kv, int_to_str(n))
|
||||
let kv = el_list_append(kv, "accumulated")
|
||||
let kv = el_list_append(kv, acc)
|
||||
return json_build_object(kv)
|
||||
}
|
||||
|
||||
// ratio_to_permille — parse a decimal ratio string ("1.0", "0.8") into an
|
||||
// integer per-mille (1000, 800) so failure thresholds use exact integer math.
|
||||
// (El float division is unreliable in this runtime — int_to_float(n)/int_to_float(n)
|
||||
// does not equal 1.0 — so the swarm deliberately avoids floats.)
|
||||
fn ratio_to_permille(s: String) -> Int {
|
||||
if str_eq(s, "") {
|
||||
return 1000
|
||||
}
|
||||
let parts: [String] = str_split(s, ".")
|
||||
let whole: Int = str_to_int(el_list_get(parts, 0))
|
||||
let permille: Int = whole * 1000
|
||||
if el_list_len(parts) > 1 {
|
||||
let frac_raw: String = el_list_get(parts, 1)
|
||||
let frac3: String = str_slice(str_pad_right(frac_raw, 3, "0"), 0, 3)
|
||||
let permille = permille + str_to_int(frac3)
|
||||
}
|
||||
return permille
|
||||
}
|
||||
|
||||
// swarm_converge — dispatch on strategy name.
|
||||
fn swarm_converge(strategy: String, results: [String]) -> String {
|
||||
if str_eq(strategy, "merge") {
|
||||
return swarm_converge_merge(results)
|
||||
}
|
||||
if str_eq(strategy, "vote") {
|
||||
return swarm_converge_vote(results)
|
||||
}
|
||||
if str_eq(strategy, "reduce") {
|
||||
return swarm_converge_reduce(results)
|
||||
}
|
||||
// default: collect
|
||||
return swarm_converge_collect(results)
|
||||
}
|
||||
|
||||
// ── the ONLY global-engram write path (Rule 4, @manager-only) ────────────────
|
||||
//
|
||||
// Every engram mutation flows through here and is gated by the caller's token
|
||||
// capability. Only the orchestrator's token carries engram:write, so a worker
|
||||
// (engram:read only) calling this is DENIED by capability before any HTTP is
|
||||
// issued — structurally unable to mutate global engram state, regardless of
|
||||
// engram health. This is the curated-merge write: the orchestrator committing
|
||||
// the geometry it approved. Workers never reach a successful branch here.
|
||||
fn swarm_engram_write(token: String, corr_id: String, content: String, typ: String, importance: Float) -> String {
|
||||
let deny: String = containment_guard_engram_write(token, corr_id, "engram.write")
|
||||
if str_eq(deny, "") {
|
||||
// authorized (orchestrator) — perform the write
|
||||
let res: String = op_write(content, typ, importance)
|
||||
let new_id: String = json_get_string(res, "id")
|
||||
let cp: String = json_set_str("{}", "node_id", new_id)
|
||||
worktrack_append("swarm.committed", corr_id, "orchestrator", cp)
|
||||
return res
|
||||
}
|
||||
// denied by capability — return the rejection, no engram mutation performed
|
||||
return json_set_str("{}", "denied", deny)
|
||||
}
|
||||
|
||||
// ── the coordinator: fan out -> track -> converge ────────────────────────────
|
||||
//
|
||||
// blueprint : task blueprint name run by every worker
|
||||
// knowledge_refs : JSON array of retrieval queries for CCR compilation
|
||||
// inputs_json : JSON array of input items (one per worker)
|
||||
// config_json : { concurrency, strategy, min_success_ratio,
|
||||
// failure_action, caller_token }
|
||||
//
|
||||
// Returns: { corr_id, status:"completed"|"aborted", merged, report }.
|
||||
fn swarm_run(blueprint: String, knowledge_refs: String, inputs_json: String, config_json: String) -> String {
|
||||
let corr_id: String = "swarm-" + uuid_v4()
|
||||
let caller_token: String = json_get_raw(config_json, "caller_token")
|
||||
let concurrency: Int = str_to_int(json_get_string(config_json, "concurrency"))
|
||||
if concurrency < 1 {
|
||||
let concurrency = 4
|
||||
}
|
||||
let strategy: String = json_get_string(config_json, "strategy")
|
||||
|
||||
// ── Containment Rule 2: only a coordinator/absent token may open a swarm ──
|
||||
let deny: String = containment_guard_open(caller_token, corr_id)
|
||||
if str_eq(deny, "") {
|
||||
// allowed — proceed
|
||||
let n: Int = json_array_len(inputs_json)
|
||||
|
||||
// swarm.created
|
||||
let cp: String = json_set_str("{}", "blueprint", blueprint)
|
||||
let cp2: String = json_set(cp, "input_count", int_to_str(n))
|
||||
worktrack_append("swarm.created", corr_id, corr_id, cp2)
|
||||
|
||||
// build per-worker envelopes: worker token + CCR-compiled bounded context
|
||||
let envelopes: [String] = el_list_empty()
|
||||
let i = 0
|
||||
while i < n {
|
||||
let worker_id: String = corr_id + "/worker-" + int_to_str(i)
|
||||
let input_item: String = json_array_get_string(inputs_json, i)
|
||||
let wtoken: String = containment_worker_token(corr_id, worker_id)
|
||||
let ctx: String = ccr_compile(blueprint, knowledge_refs, input_item, corr_id, worker_id, wtoken)
|
||||
// envelope: only this worker's compiled context + its closed token
|
||||
let ekv: [String] = el_list_empty()
|
||||
let ekv = el_list_append(ekv, "worker_id")
|
||||
let ekv = el_list_append(ekv, worker_id)
|
||||
let ekv = el_list_append(ekv, "corr_id")
|
||||
let ekv = el_list_append(ekv, corr_id)
|
||||
let env0: String = json_build_object(ekv)
|
||||
let env1: String = json_set(env0, "scope_token", wtoken)
|
||||
let env2: String = json_set(env1, "ctx", ctx)
|
||||
let envelopes = el_list_append(envelopes, env2)
|
||||
|
||||
let sp: String = json_set_str("{}", "input", input_item)
|
||||
worktrack_append("worker.started", corr_id, worker_id, sp)
|
||||
let i = i + 1
|
||||
}
|
||||
|
||||
// ── native-thread fan-out (bounded) ──
|
||||
let results: [String] = swarm_fanout("swarm_worker_entry", envelopes, concurrency)
|
||||
|
||||
// record per-worker terminal status + aggregate AFFERENT telemetry.
|
||||
// Afferent counters (primitive signals travelling toward the @manager)
|
||||
// are summed from the vertical result path — no shared bus register,
|
||||
// so the aggregation is containment-safe.
|
||||
let succ = 0
|
||||
let afferent = 0
|
||||
let seam_mode_seen: String = "stub"
|
||||
let rn: Int = el_list_len(results)
|
||||
let r = 0
|
||||
while r < rn {
|
||||
let res: String = el_list_get(results, r)
|
||||
let wid: String = json_get_string(res, "worker_id")
|
||||
let st: String = json_get_string(res, "status")
|
||||
let out: String = json_get_raw(res, "output")
|
||||
let aff: Int = str_to_int(json_get_string(out, "afferent"))
|
||||
let afferent = afferent + aff
|
||||
let sm: String = json_get_string(out, "seam_mode")
|
||||
if str_eq(sm, "") {
|
||||
let seam_mode_seen = seam_mode_seen
|
||||
} else {
|
||||
let seam_mode_seen = sm
|
||||
}
|
||||
if str_eq(st, "completed") {
|
||||
let succ = succ + 1
|
||||
worktrack_append("worker.completed", corr_id, wid, json_set_str("{}", "status", "completed"))
|
||||
} else {
|
||||
worktrack_append("worker.failed", corr_id, wid, json_set_str("{}", "error", json_get_string(res, "error")))
|
||||
}
|
||||
let r = r + 1
|
||||
}
|
||||
|
||||
// swarm.converging
|
||||
let vg: String = json_set("{}", "success_count", int_to_str(succ))
|
||||
worktrack_append("swarm.converging", corr_id, corr_id, vg)
|
||||
|
||||
// swarm.telemetry — afferent counters observed by the @manager.
|
||||
let tkv: [String] = el_list_empty()
|
||||
let tkv = el_list_append(tkv, "seam_mode")
|
||||
let tkv = el_list_append(tkv, seam_mode_seen)
|
||||
let telem0: String = json_build_object(tkv)
|
||||
let telem1: String = json_set_str(telem0, "afferent_think", int_to_str(afferent))
|
||||
let telemetry: String = json_set_str(telem1, "results_received", int_to_str(rn))
|
||||
worktrack_append("swarm.telemetry", corr_id, corr_id, telemetry)
|
||||
|
||||
// ── failure threshold (Swarm §4.3), integer per-mille math ──
|
||||
// require succ/n >= min_success_ratio <=> succ*1000 >= permille*n
|
||||
let permille: Int = ratio_to_permille(json_get_string(config_json, "min_success_ratio"))
|
||||
let status: String = "completed"
|
||||
if succ * 1000 < permille * n {
|
||||
let status = "aborted"
|
||||
}
|
||||
|
||||
if str_eq(status, "aborted") {
|
||||
let ap: String = json_set_str("{}", "reason", "success ratio below min_success_ratio")
|
||||
worktrack_append("swarm.aborted", corr_id, corr_id, ap)
|
||||
let rep: String = worktrack_swarm_report(corr_id)
|
||||
let ok: [String] = el_list_empty()
|
||||
let ok = el_list_append(ok, "corr_id")
|
||||
let ok = el_list_append(ok, corr_id)
|
||||
let ok = el_list_append(ok, "status")
|
||||
let ok = el_list_append(ok, "aborted")
|
||||
let out0: String = json_build_object(ok)
|
||||
return json_set(out0, "report", rep)
|
||||
}
|
||||
|
||||
// ── converge ──
|
||||
let merged: String = swarm_converge(strategy, results)
|
||||
let dp: String = json_set_str("{}", "strategy", strategy)
|
||||
worktrack_append("swarm.completed", corr_id, corr_id, dp)
|
||||
|
||||
// ── curated merge = the ONLY engram write path (Rule 4) ──
|
||||
// With "commit":"1", the ORCHESTRATOR (its token carries engram:write)
|
||||
// commits the approved merged geometry back to the engram. This is the
|
||||
// single writer. Workers returned geometry; only the orchestrator writes.
|
||||
let commit_id: String = ""
|
||||
if str_eq(json_get_string(config_json, "commit"), "1") {
|
||||
let orch_token: String = containment_coordinator_token(corr_id)
|
||||
let cres: String = swarm_engram_write(orch_token, corr_id, "swarm-merge " + corr_id + " :: " + merged, "memory", 0.5)
|
||||
let commit_id = json_get_string(cres, "id")
|
||||
}
|
||||
|
||||
let rep2: String = worktrack_swarm_report(corr_id)
|
||||
let ok2: [String] = el_list_empty()
|
||||
let ok2 = el_list_append(ok2, "corr_id")
|
||||
let ok2 = el_list_append(ok2, corr_id)
|
||||
let ok2 = el_list_append(ok2, "status")
|
||||
let ok2 = el_list_append(ok2, "completed")
|
||||
let out1: String = json_build_object(ok2)
|
||||
let out2: String = json_set(out1, "report", rep2)
|
||||
let out3: String = json_set(out2, "merged", merged)
|
||||
let out4: String = json_set(out3, "telemetry", telemetry)
|
||||
return json_set_str(out4, "committed_node", commit_id)
|
||||
}
|
||||
// ── denied: caller was a worker trying to open a swarm (Rule 2) ──
|
||||
let dkv: [String] = el_list_empty()
|
||||
let dkv = el_list_append(dkv, "corr_id")
|
||||
let dkv = el_list_append(dkv, corr_id)
|
||||
let dkv = el_list_append(dkv, "status")
|
||||
let dkv = el_list_append(dkv, "denied")
|
||||
let dkv = el_list_append(dkv, "error")
|
||||
let dkv = el_list_append(dkv, deny)
|
||||
return json_build_object(dkv)
|
||||
}
|
||||
@@ -1,88 +0,0 @@
|
||||
// harness_local_swarm.el — LOCAL-SWARM INTEGRATION HARNESS.
|
||||
//
|
||||
// Proves the FULL local-swarm mechanics end-to-end, TODAY, on the isolated
|
||||
// engram clone (:8901), with the primitive seam pointed at the hermetic stub.
|
||||
// The moment the api-reshape agent lands the decorated primitives on the
|
||||
// dharma bus, binding is ONE flip (primitive_binding.el) + SWARM_PRIMITIVE_SEAM=
|
||||
// decorated — this same harness then runs the bound path with no other change.
|
||||
//
|
||||
// The @manager (the coordinator) fans out N native El worker threads at real
|
||||
// concurrency, each given a CCR-scoped engram slice, each invoking the primitive
|
||||
// seam (think over its slice), enforces all three containment rules, converges
|
||||
// (vote AND reduce), work-tracks durably, and observes afferent telemetry.
|
||||
//
|
||||
// Run with the sandbox env sourced (ENGRAM_URL=:8901) to also exercise CCR
|
||||
// retrieval against the real (isolated) mind; runs fully without it too.
|
||||
|
||||
fn ok(label: String, cond: Bool, fails: Int) -> Int {
|
||||
if cond { print(" ok " + label); return fails }
|
||||
print(" FAIL " + label); return fails + 1
|
||||
}
|
||||
|
||||
fn main() -> Int {
|
||||
let fails = 0
|
||||
print("== LOCAL-SWARM INTEGRATION HARNESS (seam=" + seam_mode() + ") ==")
|
||||
|
||||
// 8 independent slices, real concurrency of 4 (2 waves of native pthreads).
|
||||
let inputs: String = "[\"billing\",\"payments\",\"ledger\",\"invoicing\",\"tax\",\"payroll\",\"audit\",\"fx\"]"
|
||||
let refs: String = "[\"Volatility-Based Decomposition\"]"
|
||||
|
||||
// ── A) fan-out / converge at real concurrency (reduce) ──
|
||||
let cfg_r: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rr: String = swarm_run("analyze_item", refs, inputs, cfg_r)
|
||||
let fails = ok("swarm completed at concurrency=4 over 8 native-thread workers", str_eq(json_get_string(rr, "status"), "completed"), fails)
|
||||
let corr: String = json_get_string(rr, "corr_id")
|
||||
let merged_r: String = json_get_raw(rr, "merged")
|
||||
let fails = ok("reduce converged all 8 worker outputs", str_to_int(json_get_string(merged_r, "count")) == 8, fails)
|
||||
|
||||
// ── B) afferent telemetry observed by the @manager ──
|
||||
let telem: String = json_get_raw(rr, "telemetry")
|
||||
let aff: Int = str_to_int(json_get_string(telem, "afferent_think"))
|
||||
let seen_mode: String = json_get_string(telem, "seam_mode")
|
||||
let fails = ok("afferent think-signals counted = 8 (one per worker)", aff == 8, fails)
|
||||
let fails = ok("telemetry records the active seam mode", str_eq(seen_mode, seam_mode()), fails)
|
||||
let telem_recs: Int = worktrack_count_kind(corr, "swarm.telemetry")
|
||||
let fails = ok("telemetry durably journalled", telem_recs == 1, fails)
|
||||
|
||||
// ── C) CCR scoping + non-leak per worker ──
|
||||
let wt: String = containment_worker_token(corr, corr + "/worker-3")
|
||||
let ctx3: String = ccr_compile("analyze_item", refs, "invoicing", corr, corr + "/worker-3", wt)
|
||||
let fails = ok("CCR context bounded within token budget", ccr_within_budget(ctx3), fails)
|
||||
let fails = ok("CCR context carries THIS slice", str_eq(json_get_string(ctx3, "input"), "invoicing"), fails)
|
||||
let leaks: Bool = str_contains(ctx3, "payroll") || str_contains(ctx3, "audit")
|
||||
let fails = ok("CCR context does NOT leak sibling slices (security boundary)", !leaks, fails)
|
||||
|
||||
// ── D) all three containment rules ──
|
||||
let deny: String = containment_check_open(wt)
|
||||
let fails = ok("Rule 2: worker token may not OPEN a swarm", !str_eq(deny, ""), fails)
|
||||
let denyj: String = containment_check_join(wt, "other-swarm")
|
||||
let fails = ok("Rule 1: worker token may not JOIN another swarm", !str_eq(denyj, ""), fails)
|
||||
let lat: String = containment_check_lateral(wt, "sibling-9")
|
||||
let fails = ok("Rule 3: worker->worker lateral edge rejected", !str_eq(lat, ""), fails)
|
||||
let ver: String = containment_check_lateral(wt, "")
|
||||
let fails = ok("Rule 3: worker->manager vertical edge allowed", str_eq(ver, ""), fails)
|
||||
// enforced live: a worker-token caller is denied opening a real swarm
|
||||
let wcfg: String = json_set(cfg_r, "caller_token", wt)
|
||||
let denied: String = swarm_run("analyze_item", refs, inputs, wcfg)
|
||||
let fails = ok("Rule 2 enforced live: worker-caller swarm denied", str_eq(json_get_string(denied, "status"), "denied"), fails)
|
||||
|
||||
// ── E) vote convergence strategy at concurrency ──
|
||||
let cfg_v: String = "{\"concurrency\":\"8\",\"strategy\":\"vote\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rv: String = swarm_run("classify", refs, inputs, cfg_v)
|
||||
let winner: String = json_get_string(json_get_raw(rv, "merged"), "winner")
|
||||
// billing/payments/ledger/invoicing/payroll/audit = long(>4); tax/fx = short -> long wins
|
||||
let fails = ok("vote converged (winner=long)", str_eq(winner, "long"), fails)
|
||||
|
||||
// ── F) durable, inspectable work-tracking ──
|
||||
let started: Int = worktrack_count_kind(corr, "worker.started")
|
||||
let completed: Int = worktrack_count_kind(corr, "worker.completed")
|
||||
let fails = ok("work-tracking journal: 8 started + 8 completed", (started == 8) && (completed == 8), fails)
|
||||
|
||||
print("")
|
||||
if fails == 0 {
|
||||
print("HARNESS GREEN — full local-swarm mechanics proven with seam=" + seam_mode())
|
||||
return 0
|
||||
}
|
||||
print("HARNESS FAIL (" + int_to_str(fails) + ")")
|
||||
return 1
|
||||
}
|
||||
@@ -1,119 +0,0 @@
|
||||
// harness_real_cognition.el — the LOCAL SWARM running REAL cognition.
|
||||
//
|
||||
// Run with: SWARM_PRIMITIVE_SEAM=decorated + the sandbox env sourced
|
||||
// (ENGRAM_URL=:8901). Each worker's `think` is BOUND to the reshape's proven
|
||||
// op_think (GET /api/think) over its NODE-ID anchor — real 768-dim gradients from
|
||||
// the live (isolated) geometry, not the stub. The @manager fans out N native-El
|
||||
// worker threads at real concurrency, converges (reduce + vote) over the real
|
||||
// cognition, enforces all three containment rules, observes afferent telemetry,
|
||||
// and work-tracks durably.
|
||||
//
|
||||
// Anchors are real self-neighbourhood node ids on the :8901 clone (free-text
|
||||
// anchors return "geometry unavailable", so these must be node ids).
|
||||
|
||||
fn ok(label: String, cond: Bool, fails: Int) -> Int {
|
||||
if cond { print(" ok " + label); return fails }
|
||||
print(" FAIL " + label); return fails + 1
|
||||
}
|
||||
|
||||
fn main() -> Int {
|
||||
let fails = 0
|
||||
print("== REAL-COGNITION LOCAL SWARM (seam=" + seam_mode() + ", engram=" + env("ENGRAM_URL") + ") ==")
|
||||
|
||||
// ── 0) direct proof the bound primitive returns REAL cognition ──
|
||||
let g: String = op_think("self", "plan")
|
||||
let dim: Int = json_get_int(g, "dim")
|
||||
let nsup: Int = json_get_int(g, "n_support")
|
||||
let fails = ok("bound op_think returns a real 768-dim gradient", dim == 768, fails)
|
||||
let fails = ok("real gradient has support (n_support>0)", nsup > 0, fails)
|
||||
let gfree: String = op_think("this-is-free-text-not-a-node", "reason")
|
||||
let fails = ok("free-text anchor correctly refused (geometry unavailable)", str_contains(gfree, "geometry unavailable"), fails)
|
||||
|
||||
// ── the input set: 8 real NODE-ID anchors from self's neighbourhood ──
|
||||
let anchors: String = "[\"a1000001-0000-0000-0000-000000000001\",\"5f011441-fa43-4fe7-a9c0-c78a584ef11d\",\"kn-5adecd7e-d6db-4576-87fe-6ef8a935cea6\",\"76d7fd0b-0672-4511-a2f5-a095cf9c60ae\",\"7027e302-593f-441d-8fd6-9c400c163108\",\"2a730b18-6566-46ee-a21e-4f4dd0380908\",\"46b0e4dd-2c19-48d2-bcbc-19f61d6c79ae\",\"9162cde8-8739-4f00-bfc9-2850ed612e50\"]"
|
||||
let refs: String = "[\"self\"]"
|
||||
|
||||
// ── A) fan-out real cognition at concurrency, converge with REDUCE ──
|
||||
let cfg_r: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rr: String = swarm_run("cognize", refs, anchors, cfg_r)
|
||||
let fails = ok("swarm completed: 8 workers each a real think, concurrency=4", str_eq(json_get_string(rr, "status"), "completed"), fails)
|
||||
let corr: String = json_get_string(rr, "corr_id")
|
||||
let merged_r: String = json_get_raw(rr, "merged")
|
||||
let fails = ok("reduce converged all 8 real-cognition outputs", str_to_int(json_get_string(merged_r, "count")) == 8, fails)
|
||||
let acc: String = json_get_string(merged_r, "accumulated")
|
||||
let fails = ok("converged output carries real gradient support (n_support)", str_contains(acc, "n_support"), fails)
|
||||
|
||||
// ── B) afferent telemetry: 8 real think-signals, decorated seam ──
|
||||
let telem: String = json_get_raw(rr, "telemetry")
|
||||
let aff: Int = str_to_int(json_get_string(telem, "afferent_think"))
|
||||
let fails = ok("afferent counters = 8 real think invocations", aff == 8, fails)
|
||||
let fails = ok("telemetry records seam_mode=decorated", str_eq(json_get_string(telem, "seam_mode"), "decorated"), fails)
|
||||
let fails = ok("telemetry durably journalled", worktrack_count_kind(corr, "swarm.telemetry") == 1, fails)
|
||||
|
||||
// ── C) converge with VOTE over real cognition ──
|
||||
let cfg_v: String = "{\"concurrency\":\"8\",\"strategy\":\"vote\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rv: String = swarm_run("cognize", refs, anchors, cfg_v)
|
||||
let winner: String = json_get_string(json_get_raw(rv, "merged"), "winner")
|
||||
let fails = ok("vote converged over real cognition (winner=" + winner + ")", !str_eq(winner, ""), fails)
|
||||
|
||||
// ── D) all three containment rules still enforced ──
|
||||
let wt: String = containment_worker_token(corr, corr + "/worker-2")
|
||||
let fails = ok("Rule 2: worker may not open a swarm", !str_eq(containment_check_open(wt), ""), fails)
|
||||
let fails = ok("Rule 1: worker may not join another swarm", !str_eq(containment_check_join(wt, "s2"), ""), fails)
|
||||
let fails = ok("Rule 3: worker->worker lateral edge rejected", !str_eq(containment_check_lateral(wt, "sib"), ""), fails)
|
||||
let wcfg: String = json_set(cfg_r, "caller_token", wt)
|
||||
let denied: String = swarm_run("cognize", refs, anchors, wcfg)
|
||||
let fails = ok("Rule 2 enforced LIVE: worker-caller swarm denied", str_eq(json_get_string(denied, "status"), "denied"), fails)
|
||||
|
||||
// ── E) CCR scoping + non-leak over node-id anchors ──
|
||||
let ctx: String = ccr_compile("cognize", refs, "a1000001-0000-0000-0000-000000000001", corr, corr + "/worker-0", wt)
|
||||
let fails = ok("CCR context bounded within budget", ccr_within_budget(ctx), fails)
|
||||
let leaks: Bool = str_contains(ctx, "9162cde8")
|
||||
let fails = ok("CCR context does NOT leak sibling anchors", !leaks, fails)
|
||||
|
||||
// ── F) durable work-tracking ──
|
||||
let started: Int = worktrack_count_kind(corr, "worker.started")
|
||||
let completed: Int = worktrack_count_kind(corr, "worker.completed")
|
||||
let fails = ok("work-tracking: 8 started + 8 completed", (started == 8) && (completed == 8), fails)
|
||||
|
||||
// ── G) RULE 4 — engram-write is @manager-ONLY (authority gate) ──
|
||||
// A worker token (engram:read only) is STRUCTURALLY denied any engram write.
|
||||
let worker_tok: String = containment_worker_token(corr, corr + "/worker-1")
|
||||
let orch_tok: String = containment_coordinator_token(corr)
|
||||
let fails = ok("worker token carries engram:read", containment_has_cap(worker_tok, "engram:read"), fails)
|
||||
let fails = ok("worker token does NOT carry engram:write", !containment_has_cap(worker_tok, "engram:write"), fails)
|
||||
let fails = ok("orchestrator token carries engram:write", containment_has_cap(orch_tok, "engram:write"), fails)
|
||||
// a worker attempting an engram write is DENIED BY CAPABILITY (no HTTP issued)
|
||||
let wdeny: String = swarm_engram_write(worker_tok, corr, "worker tries to mutate global state", "memory", 0.5)
|
||||
let denied_reason: String = json_get_string(wdeny, "denied")
|
||||
let fails = ok("worker engram-write DENIED by capability (Rule 4)", str_contains(denied_reason, "rule 4"), fails)
|
||||
let fails = ok("denied worker write performed NO engram mutation (no node id)", str_eq(json_get_string(wdeny, "id"), ""), fails)
|
||||
let fails = ok("Rule-4 violation journalled", worktrack_count_kind(corr, "containment.violation") >= 1, fails)
|
||||
// the orchestrator passes the capability gate (sole authorized writer)
|
||||
let odeny: String = containment_check_engram_write(orch_tok, "engram.write")
|
||||
let fails = ok("orchestrator PASSES the engram-write capability gate (sole writer)", str_eq(odeny, ""), fails)
|
||||
|
||||
// ── H) curated merge = the only write path (orchestrator commits) ──
|
||||
// The AUTHORITY gate above is already proven (worker denied, orchestrator
|
||||
// authorized) WITHOUT issuing a write. The actual persisting commit exercises
|
||||
// the engram write path, which needs the gate-1 write-healthy clone — so it
|
||||
// runs only under SWARM_WRITE_HEALTHY=1 (else it would hit the known daemon
|
||||
// write-crash). Authority != health: the gate holds either way.
|
||||
if str_eq(env("SWARM_WRITE_HEALTHY"), "1") {
|
||||
let cfg_commit: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\",\"commit\":\"1\"}"
|
||||
let rc: String = swarm_run("cognize", refs, anchors, cfg_commit)
|
||||
let committed: String = json_get_string(rc, "committed_node")
|
||||
let fails2: Int = ok("orchestrator (sole writer) committed the merge to the engram", !str_eq(committed, ""), fails)
|
||||
let fails = fails2
|
||||
} else {
|
||||
print(" note curated-merge commit deferred to the gate-1 write-healthy clone (set SWARM_WRITE_HEALTHY=1); authority gate already proven above")
|
||||
}
|
||||
|
||||
print("")
|
||||
if fails == 0 {
|
||||
print("REAL-COGNITION SWARM GREEN — Neuron thinking in parallel over its own geometry.")
|
||||
return 0
|
||||
}
|
||||
print("REAL-COGNITION SWARM FAIL (" + int_to_str(fails) + ")")
|
||||
return 1
|
||||
}
|
||||
@@ -1,45 +0,0 @@
|
||||
// integ_engram.el — integration proof against a LIVE (isolated) engram.
|
||||
//
|
||||
// Run with the sandbox env sourced (ENGRAM_URL=http://127.0.0.1:8901,
|
||||
// ENGRAM_API_KEY=sbx-dev-swarm-ccr). Proves:
|
||||
// (a) CCR retrieval pulls REAL content from the mind over HTTP;
|
||||
// (b) a full swarm runs and converges against the live mind;
|
||||
// (c) work-tracking mirrors records into the engram as SwarmTrack nodes.
|
||||
|
||||
fn main() -> Int {
|
||||
let url: String = env("ENGRAM_URL")
|
||||
if str_eq(url, "") {
|
||||
print("SKIP integ_engram (ENGRAM_URL not set)")
|
||||
return 0
|
||||
}
|
||||
|
||||
// (a) CCR compiles a bounded context whose retrieval hit the real mind.
|
||||
let refs: String = "[\"Volatility-Based Decomposition\",\"Swarm Architecture containment\"]"
|
||||
let wt: String = containment_worker_token("integ", "integ/w0")
|
||||
let ctx: String = ccr_compile("analyze_item", refs, "decompose the billing module", "integ", "integ/w0", wt)
|
||||
let knowledge: String = json_get_string(ctx, "knowledge")
|
||||
let pulled_real: Bool = str_contains(knowledge, "olatility") || str_contains(knowledge, "Anderson") || str_contains(knowledge, "VBD")
|
||||
if pulled_real {
|
||||
print(" ok CCR retrieval pulled real mind content (" + int_to_str(str_len(knowledge)) + " bytes, bounded)")
|
||||
} else {
|
||||
print(" FAIL CCR retrieval returned no mind content")
|
||||
}
|
||||
let bounded: Bool = ccr_within_budget(ctx)
|
||||
if bounded { print(" ok compiled context stayed within budget") } else { print(" FAIL context over budget") }
|
||||
|
||||
// (b) a real swarm over the live mind.
|
||||
let inputs: String = "[\"billing\",\"payments\",\"ledger\"]"
|
||||
let cfg: String = "{\"concurrency\":\"3\",\"strategy\":\"collect\",\"min_success_ratio\":\"1.0\"}"
|
||||
let res: String = swarm_run("analyze_item", refs, inputs, cfg)
|
||||
let status: String = json_get_string(res, "status")
|
||||
if str_eq(status, "completed") { print(" ok swarm completed against live engram") } else { print(" FAIL swarm status=" + status) }
|
||||
let corr: String = json_get_string(res, "corr_id")
|
||||
|
||||
// (c) work-tracking mirrored into the mind: search for this swarm's records.
|
||||
let hits: String = primitive_attend(corr, 5)
|
||||
let mirrored: Bool = str_contains(hits, "swarm-track") || str_contains(hits, corr)
|
||||
if mirrored { print(" ok work-tracking mirrored into the engram (queryable)") } else { print(" note mirror not yet visible to search (async index)") }
|
||||
|
||||
print("DONE integ_engram corr=" + corr)
|
||||
return 0
|
||||
}
|
||||
@@ -1,55 +0,0 @@
|
||||
// test_convergence.el — convergence strategies + failure threshold / abort.
|
||||
|
||||
fn assert_true(label: String, cond: Bool, fails: Int) -> Int {
|
||||
if cond { print(" ok " + label); return fails }
|
||||
print(" FAIL " + label); return fails + 1
|
||||
}
|
||||
|
||||
fn main() -> Int {
|
||||
let fails = 0
|
||||
let refs: String = "[]"
|
||||
|
||||
// ── vote: classify 5 inputs; 3 "long" (>4 chars) vs 2 "short" -> winner long ──
|
||||
let inputs: String = "[\"alpha\",\"bravo\",\"hi\",\"charlie\",\"ok\"]"
|
||||
let cfg_v: String = "{\"concurrency\":\"3\",\"strategy\":\"vote\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rv: String = swarm_run("classify", refs, inputs, cfg_v)
|
||||
let merged_v: String = json_get_raw(rv, "merged")
|
||||
let winner: String = json_get_string(merged_v, "winner")
|
||||
let votes: Int = str_to_int(json_get_string(merged_v, "votes"))
|
||||
let fails = assert_true("vote winner = long", str_eq(winner, "long"), fails)
|
||||
let fails = assert_true("vote count = 3", votes == 3, fails)
|
||||
|
||||
// ── merge: outputs joined ──
|
||||
let cfg_m: String = "{\"concurrency\":\"2\",\"strategy\":\"merge\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rm: String = swarm_run("analyze_item", refs, "[\"a\",\"b\",\"c\"]", cfg_m)
|
||||
let merged_m: String = json_get_raw(rm, "merged")
|
||||
let joined: String = json_get_string(merged_m, "merged")
|
||||
let fails = assert_true("merge produced a joined string", str_contains(joined, "|"), fails)
|
||||
|
||||
// ── reduce: count accumulates ──
|
||||
let cfg_r: String = "{\"concurrency\":\"4\",\"strategy\":\"reduce\",\"min_success_ratio\":\"1.0\"}"
|
||||
let rr: String = swarm_run("analyze_item", refs, "[\"a\",\"b\",\"c\",\"d\"]", cfg_r)
|
||||
let merged_r: String = json_get_raw(rr, "merged")
|
||||
let rcount: Int = str_to_int(json_get_string(merged_r, "count"))
|
||||
let fails = assert_true("reduce count = 4", rcount == 4, fails)
|
||||
|
||||
// ── failure threshold: 2 of 5 fail (x-prefixed); ratio 3/5=0.6 < 0.8 -> aborted ──
|
||||
let fin: String = "[\"a\",\"xb\",\"c\",\"xd\",\"e\"]"
|
||||
let cfg_f: String = "{\"concurrency\":\"5\",\"strategy\":\"collect\",\"min_success_ratio\":\"0.8\"}"
|
||||
let rf: String = swarm_run("faildemo", refs, fin, cfg_f)
|
||||
let fstatus: String = json_get_string(rf, "status")
|
||||
let fails = assert_true("swarm aborted below min_success_ratio (0.6<0.8)", str_eq(fstatus, "aborted"), fails)
|
||||
let corr_f: String = json_get_string(rf, "corr_id")
|
||||
let failed_n: Int = worktrack_count_kind(corr_f, "worker.failed")
|
||||
let aborted_n: Int = worktrack_count_kind(corr_f, "swarm.aborted")
|
||||
let fails = assert_true("tracked 2 worker.failed", failed_n == 2, fails)
|
||||
let fails = assert_true("tracked swarm.aborted", aborted_n == 1, fails)
|
||||
|
||||
// ── same failures tolerated when min_success_ratio=0.5 (0.6>=0.5) -> completed ──
|
||||
let cfg_ok: String = "{\"concurrency\":\"5\",\"strategy\":\"collect\",\"min_success_ratio\":\"0.5\"}"
|
||||
let rok: String = swarm_run("faildemo", refs, fin, cfg_ok)
|
||||
let fails = assert_true("swarm completes when failures within tolerance", str_eq(json_get_string(rok, "status"), "completed"), fails)
|
||||
|
||||
if fails == 0 { print("PASS test_convergence"); return 0 }
|
||||
print("FAIL test_convergence (" + int_to_str(fails) + ")"); return 1
|
||||
}
|
||||
@@ -1,75 +0,0 @@
|
||||
// test_swarm.el — end-to-end proof of the swarm capability on native El threads.
|
||||
//
|
||||
// Proves: native-thread fan-out/converge, bounded concurrency, per-worker CCR
|
||||
// bounded context (with the security-boundary property), containment Rule 2
|
||||
// enforcement, and durable work-tracking.
|
||||
|
||||
fn assert_true(label: String, cond: Bool, fails: Int) -> Int {
|
||||
if cond {
|
||||
print(" ok " + label)
|
||||
return fails
|
||||
}
|
||||
print(" FAIL " + label)
|
||||
return fails + 1
|
||||
}
|
||||
|
||||
fn main() -> Int {
|
||||
let fails = 0
|
||||
|
||||
// ── 1) fan-out / converge (collect) over native threads ──
|
||||
let inputs: String = "[\"alpha\",\"bravo\",\"charlie\",\"delta\",\"echo\"]"
|
||||
let refs: String = "[]"
|
||||
let cfg: String = "{\"concurrency\":\"2\",\"strategy\":\"collect\",\"min_success_ratio\":\"1.0\"}"
|
||||
let res: String = swarm_run("analyze_item", refs, inputs, cfg)
|
||||
let status: String = json_get_string(res, "status")
|
||||
let fails = assert_true("swarm completed", str_eq(status, "completed"), fails)
|
||||
|
||||
let merged: String = json_get_raw(res, "merged")
|
||||
let count: Int = json_array_len(merged)
|
||||
let fails = assert_true("collect returned 5 results (bounded concurrency=2)", count == 5, fails)
|
||||
|
||||
// ── 2) work-tracking is durable + complete ──
|
||||
let corr: String = json_get_string(res, "corr_id")
|
||||
let started: Int = worktrack_count_kind(corr, "worker.started")
|
||||
let completed: Int = worktrack_count_kind(corr, "worker.completed")
|
||||
let created: Int = worktrack_count_kind(corr, "swarm.created")
|
||||
let done: Int = worktrack_count_kind(corr, "swarm.completed")
|
||||
let fails = assert_true("tracked 5 worker.started", started == 5, fails)
|
||||
let fails = assert_true("tracked 5 worker.completed", completed == 5, fails)
|
||||
let fails = assert_true("tracked swarm.created + swarm.completed", (created == 1) && (done == 1), fails)
|
||||
|
||||
// ── 3) CCR: bounded, minimal, non-leaking per-worker context ──
|
||||
let wtoken: String = containment_worker_token(corr, corr + "/worker-0")
|
||||
let ctx: String = ccr_compile("analyze_item", refs, "alpha", corr, corr + "/worker-0", wtoken)
|
||||
let in_budget: Bool = ccr_within_budget(ctx)
|
||||
let fails = assert_true("CCR context within token budget", in_budget, fails)
|
||||
let this_input: String = json_get_string(ctx, "input")
|
||||
let fails = assert_true("CCR context contains THIS worker's input", str_eq(this_input, "alpha"), fails)
|
||||
// security boundary: a worker's compiled context must not carry a sibling input
|
||||
let leaks_sibling: Bool = str_contains(ctx, "charlie")
|
||||
let fails = assert_true("CCR context does NOT leak sibling inputs", !leaks_sibling, fails)
|
||||
|
||||
// ── 4) containment Rule 2: a worker may not open a swarm ──
|
||||
let worker_caller_cfg: String = json_set(cfg, "caller_token", wtoken)
|
||||
let denied: String = swarm_run("analyze_item", refs, inputs, worker_caller_cfg)
|
||||
let dstatus: String = json_get_string(denied, "status")
|
||||
let fails = assert_true("worker-token caller denied opening a swarm (Rule 2)", str_eq(dstatus, "denied"), fails)
|
||||
|
||||
// coordinator token IS allowed
|
||||
let coord: String = containment_coordinator_token("some-corr")
|
||||
let allow_reason: String = containment_check_open(coord)
|
||||
let fails = assert_true("coordinator token allowed to open a swarm", str_eq(allow_reason, ""), fails)
|
||||
|
||||
// ── 5) containment Rule 3: no lateral worker->worker edge ──
|
||||
let lateral: String = containment_check_lateral(wtoken, "some-sibling")
|
||||
let fails = assert_true("lateral worker->worker edge rejected (Rule 3)", !str_eq(lateral, ""), fails)
|
||||
let vertical: String = containment_check_lateral(wtoken, "")
|
||||
let fails = assert_true("vertical worker->coordinator edge allowed", str_eq(vertical, ""), fails)
|
||||
|
||||
if fails == 0 {
|
||||
print("PASS test_swarm")
|
||||
return 0
|
||||
}
|
||||
print("FAIL test_swarm (" + int_to_str(fails) + " failures)")
|
||||
return 1
|
||||
}
|
||||
@@ -1,38 +0,0 @@
|
||||
// test_worktrack.el — durability + inspectability of the work-tracking journal.
|
||||
|
||||
fn main() -> Int {
|
||||
let corr: String = "test-" + uuid_v4()
|
||||
|
||||
// record a swarm lifecycle
|
||||
let p1: String = json_set("{}", "input_count", "3")
|
||||
worktrack_append("swarm.created", corr, "swarm-1", p1)
|
||||
worktrack_append("worker.started", corr, "worker-001", "{}")
|
||||
worktrack_append("worker.started", corr, "worker-002", "{}")
|
||||
worktrack_append("worker.completed", corr, "worker-001", "{}")
|
||||
worktrack_append("worker.failed", corr, "worker-002", "{}")
|
||||
worktrack_append("swarm.completed", corr, "swarm-1", "{}")
|
||||
|
||||
// inspect: reconstruct the report from the durable journal
|
||||
let report: String = worktrack_swarm_report(corr)
|
||||
print("report=" + report)
|
||||
|
||||
let recs_n: Int = el_list_len(worktrack_records(corr))
|
||||
print("records=" + int_to_str(recs_n))
|
||||
|
||||
let state: String = json_get_string(report, "state")
|
||||
let completed: Int = str_to_int(json_get_string(report, "workers_completed"))
|
||||
let failed: Int = str_to_int(json_get_string(report, "workers_failed"))
|
||||
|
||||
if str_eq(state, "completed") {
|
||||
if completed == 1 {
|
||||
if failed == 1 {
|
||||
if recs_n == 6 {
|
||||
print("PASS worktrack")
|
||||
return 0
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
print("FAIL worktrack")
|
||||
return 1
|
||||
}
|
||||
@@ -1,217 +0,0 @@
|
||||
// worktrack.el — full work-tracking for the swarm.
|
||||
//
|
||||
// "Intent all the way up, orchestrator at the top." Every unit of parallel
|
||||
// work a swarm fans out is recorded here: the swarm itself, each worker, its
|
||||
// status, its result summary, the convergence, and the final merged output —
|
||||
// all threaded by a single correlation ID so the entire execution graph can be
|
||||
// reconstructed and audited (Swarm Architecture §6.1).
|
||||
//
|
||||
// DURABILITY. Records are appended to a JSON-lines journal on disk. The journal
|
||||
// is append-only and single-writer: only the coordinator (the main thread, before
|
||||
// and after each fan-out and during convergence) writes to it. Workers never
|
||||
// touch it — they return structured results and the coordinator records them.
|
||||
// This is deliberate: it makes the tracking store race-free and, not
|
||||
// coincidentally, enforces Swarm containment rule 3 (no lateral worker state).
|
||||
//
|
||||
// INSPECTABILITY. The journal is plain JSONL — greppable, tailable, replayable.
|
||||
// worktrack_read() loads it back; worktrack_swarm_report() reconstructs a
|
||||
// swarm's full record from its correlation ID.
|
||||
//
|
||||
// ENGRAM MIRROR (optional). When ENGRAM_URL is set, each record is also mirrored
|
||||
// into the engram as a node (POST /api/node) tagged with the correlation ID, so
|
||||
// the swarm's execution becomes part of the durable mind, queryable by memory.
|
||||
//
|
||||
// Depends on: el_runtime.c builtins (fs_*, http_post, env, json_*, uuid_v4,
|
||||
// now_millis, str_*). No El-module concat dependencies of its own.
|
||||
|
||||
// ── JSON helper ──────────────────────────────────────────────────────────────
|
||||
// json_set inserts its value as a RAW JSON fragment (objects/arrays/numbers).
|
||||
// json_set_str sets a plain STRING value, correctly quoted and escaped. Use
|
||||
// json_set for nested JSON, json_set_str for strings.
|
||||
fn json_set_str(j: String, key: String, val: String) -> String {
|
||||
return json_set(j, key, "\"" + json_escape_string(val) + "\"")
|
||||
}
|
||||
|
||||
// ── Journal location ─────────────────────────────────────────────────────────
|
||||
|
||||
// worktrack_dir — directory holding the swarm journals.
|
||||
// Override with SWARM_TRACK_DIR; defaults to ./.swarm-track (relative to CWD).
|
||||
fn worktrack_dir() -> String {
|
||||
let d: String = env("SWARM_TRACK_DIR")
|
||||
if str_eq(d, "") {
|
||||
return ".swarm-track"
|
||||
}
|
||||
return d
|
||||
}
|
||||
|
||||
// worktrack_journal_path — the JSONL journal file for one correlation ID.
|
||||
fn worktrack_journal_path(corr_id: String) -> String {
|
||||
return worktrack_dir() + "/" + corr_id + ".jsonl"
|
||||
}
|
||||
|
||||
// worktrack_init — ensure the journal directory exists. Idempotent.
|
||||
fn worktrack_init() -> Bool {
|
||||
let d: String = worktrack_dir()
|
||||
if fs_exists(d) {
|
||||
return true
|
||||
}
|
||||
return fs_mkdir(d)
|
||||
}
|
||||
|
||||
// ── Record construction ──────────────────────────────────────────────────────
|
||||
|
||||
// worktrack_record — build one journal record as a JSON object string.
|
||||
// kind: the record kind (swarm.created, worker.started, ...)
|
||||
// corr_id: the swarm correlation ID (links every record)
|
||||
// subject: the entity the record is about (swarm id, worker id, "")
|
||||
// payload: a JSON object string with kind-specific fields
|
||||
fn worktrack_record(kind: String, corr_id: String, subject: String, payload: String) -> String {
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "kind")
|
||||
let kv = el_list_append(kv, kind)
|
||||
let kv = el_list_append(kv, "corr_id")
|
||||
let kv = el_list_append(kv, corr_id)
|
||||
let kv = el_list_append(kv, "subject")
|
||||
let kv = el_list_append(kv, subject)
|
||||
let kv = el_list_append(kv, "ts_ms")
|
||||
let kv = el_list_append(kv, int_to_str(now_millis()))
|
||||
let rec: String = json_build_object(kv)
|
||||
// Attach the payload as a nested raw JSON field.
|
||||
let rec2: String = json_set(rec, "data", payload)
|
||||
return rec2
|
||||
}
|
||||
|
||||
// ── Journal append (single-writer, durable) ──────────────────────────────────
|
||||
|
||||
// worktrack_append — append one record to the correlation journal (durable),
|
||||
// and mirror it to the engram if ENGRAM_URL is configured. Returns the record.
|
||||
//
|
||||
// fs_write here is used in append semantics: we read-modify-write the file. The
|
||||
// coordinator is the only writer, so this is safe and race-free.
|
||||
fn worktrack_append(kind: String, corr_id: String, subject: String, payload: String) -> String {
|
||||
worktrack_init()
|
||||
let rec: String = worktrack_record(kind, corr_id, subject, payload)
|
||||
let path: String = worktrack_journal_path(corr_id)
|
||||
let prior: String = ""
|
||||
if fs_exists(path) {
|
||||
let prior = fs_read(path)
|
||||
}
|
||||
let next: String = prior + rec + "\n"
|
||||
fs_write(path, next)
|
||||
worktrack_mirror_engram(rec, corr_id, kind, subject)
|
||||
return rec
|
||||
}
|
||||
|
||||
// worktrack_mirror_engram — best-effort mirror of a record into the engram.
|
||||
// No-op unless ENGRAM_URL is set. Failures are swallowed (tracking must not
|
||||
// depend on the mind being reachable).
|
||||
fn worktrack_mirror_engram(rec: String, corr_id: String, kind: String, subject: String) -> Bool {
|
||||
// Opt-in: the durable substrate is the JSONL journal (always written). The
|
||||
// engram mirror is an additional convenience, enabled with SWARM_MIRROR=1,
|
||||
// so a swarm never depends on — or loads — the mind just to track its work.
|
||||
if str_eq(env("SWARM_MIRROR"), "1") {
|
||||
// enabled — fall through to the mirror POST
|
||||
let _go: Int = 1
|
||||
} else {
|
||||
return false
|
||||
}
|
||||
let url: String = env("ENGRAM_URL")
|
||||
if str_eq(url, "") {
|
||||
return false
|
||||
}
|
||||
let content: String = "swarm-track " + kind + " " + subject + " :: " + rec
|
||||
let body_kv: [String] = el_list_empty()
|
||||
let body_kv = el_list_append(body_kv, "content")
|
||||
let body_kv = el_list_append(body_kv, content)
|
||||
let body_kv = el_list_append(body_kv, "node_type")
|
||||
let body_kv = el_list_append(body_kv, "SwarmTrack")
|
||||
let body_kv = el_list_append(body_kv, "salience")
|
||||
let body_kv = el_list_append(body_kv, "0.5")
|
||||
let body: String = json_build_object(body_kv)
|
||||
let key: String = env("ENGRAM_API_KEY")
|
||||
let body2: String = json_set_str(body, "_auth", key)
|
||||
let resp: String = http_post(url + "/api/nodes", body2)
|
||||
return true
|
||||
}
|
||||
|
||||
// ── Read / inspect ───────────────────────────────────────────────────────────
|
||||
|
||||
// worktrack_read — read the raw JSONL journal for a correlation ID.
|
||||
fn worktrack_read(corr_id: String) -> String {
|
||||
let path: String = worktrack_journal_path(corr_id)
|
||||
if fs_exists(path) {
|
||||
return fs_read(path)
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// worktrack_records — the journal as a [String] of record JSON objects, in order.
|
||||
fn worktrack_records(corr_id: String) -> [String] {
|
||||
let raw: String = worktrack_read(corr_id)
|
||||
let out: [String] = el_list_empty()
|
||||
if str_eq(raw, "") {
|
||||
return out
|
||||
}
|
||||
let lines: [String] = str_split_lines(raw)
|
||||
let n: Int = el_list_len(lines)
|
||||
let i = 0
|
||||
while i < n {
|
||||
let ln: String = el_list_get(lines, i)
|
||||
if str_eq(ln, "") {
|
||||
let i = i + 1
|
||||
} else {
|
||||
let out = el_list_append(out, ln)
|
||||
let i = i + 1
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// worktrack_count_kind — how many records of a given kind exist for a swarm.
|
||||
// Powers assertions and live status ("how many workers completed").
|
||||
fn worktrack_count_kind(corr_id: String, kind: String) -> Int {
|
||||
let recs: [String] = worktrack_records(corr_id)
|
||||
let n: Int = el_list_len(recs)
|
||||
let c = 0
|
||||
let i = 0
|
||||
while i < n {
|
||||
let r: String = el_list_get(recs, i)
|
||||
let k: String = json_get_string(r, "kind")
|
||||
if str_eq(k, kind) {
|
||||
let c = c + 1
|
||||
}
|
||||
let i = i + 1
|
||||
}
|
||||
return c
|
||||
}
|
||||
|
||||
// worktrack_swarm_report — reconstruct a compact status report for a swarm from
|
||||
// its journal: counts of started/completed/failed workers and terminal state.
|
||||
// Inspectable, durable, derived purely from the append-only record.
|
||||
fn worktrack_swarm_report(corr_id: String) -> String {
|
||||
let started: Int = worktrack_count_kind(corr_id, "worker.started")
|
||||
let completed: Int = worktrack_count_kind(corr_id, "worker.completed")
|
||||
let failed: Int = worktrack_count_kind(corr_id, "worker.failed")
|
||||
let done: Int = worktrack_count_kind(corr_id, "swarm.completed")
|
||||
let aborted: Int = worktrack_count_kind(corr_id, "swarm.aborted")
|
||||
let state: String = "running"
|
||||
if aborted > 0 {
|
||||
let state = "aborted"
|
||||
} else {
|
||||
if done > 0 {
|
||||
let state = "completed"
|
||||
}
|
||||
}
|
||||
let kv: [String] = el_list_empty()
|
||||
let kv = el_list_append(kv, "corr_id")
|
||||
let kv = el_list_append(kv, corr_id)
|
||||
let kv = el_list_append(kv, "state")
|
||||
let kv = el_list_append(kv, state)
|
||||
let kv = el_list_append(kv, "workers_started")
|
||||
let kv = el_list_append(kv, int_to_str(started))
|
||||
let kv = el_list_append(kv, "workers_completed")
|
||||
let kv = el_list_append(kv, int_to_str(completed))
|
||||
let kv = el_list_append(kv, "workers_failed")
|
||||
let kv = el_list_append(kv, int_to_str(failed))
|
||||
return json_build_object(kv)
|
||||
}
|
||||
@@ -1,118 +0,0 @@
|
||||
# nsbx — the Neuron Sandbox
|
||||
|
||||
**Dev environment as a primitive.** A reproducible way to run experiments *and code
|
||||
changes* against the **real** engram runtime on an isolated snapshot of the live
|
||||
mind — with a gated promote-to-prod path built on the proven rails.
|
||||
|
||||
Everyone (Tim, any team member, any agent) gets their own private, safe copy of the
|
||||
mind to build against. **Prod — the live Neuron on `:8742` (engram) / `:7770`
|
||||
(soul) — is untouchable from a sandbox.** A sandbox runs a *separate* engram
|
||||
process, on a *separate* port, against a *separate* clone of the store. The only op
|
||||
that can ever reach prod is `promote`, which is explicit, gated, and per-use
|
||||
approved.
|
||||
|
||||
It **wraps the real engram binary** — it never reimplements any engram logic. It
|
||||
generalises two proven proto-sandboxes into one primitive:
|
||||
|
||||
- the **cog-arch** build — isolated git worktree + build + clone of the live `.egm` + real C tests
|
||||
- the **store-fix** cutover — secondary soul + launchctl `bootout → settle → bootstrap` rails
|
||||
|
||||
## Quickstart
|
||||
|
||||
```bash
|
||||
export PATH="$PWD:$PATH" # or symlink nsbx onto your PATH
|
||||
|
||||
nsbx up # your private copy of the mind (auto-named <user>-dev)
|
||||
nsbx run <name> api /api/stats # poke it
|
||||
nsbx validate <name> # prove it: zero-loss, reboot, RSS, retrieval, keystones
|
||||
nsbx destroy <name> # cheap teardown; live untouched
|
||||
```
|
||||
|
||||
That is the whole loop. Sane defaults: stock prod binary, auto-allocated port
|
||||
(`8900+`, never `8742`/`7770`), snapshot of the live store.
|
||||
|
||||
## The code-change dev loop (first-class)
|
||||
|
||||
Run *your changed runtime*, not just the stock binary, against a snapshot:
|
||||
|
||||
```bash
|
||||
# build a runtime from a working tree, a git branch, or a prebuilt binary:
|
||||
nsbx create feat --source /path/to/worktree # elc + cc build from source
|
||||
nsbx create feat --branch feat/my-change --repo <r> # worktree the branch, then build
|
||||
nsbx create feat --binary /path/to/engram # use a prebuilt binary
|
||||
|
||||
nsbx build feat --source /path/to/worktree # rebuild + hot-restart in place
|
||||
nsbx validate feat # prove the change is safe
|
||||
nsbx promote feat --i-approve-prod-cutover # gated rails cutover (see below)
|
||||
```
|
||||
|
||||
The build replicates the engram release recipe exactly:
|
||||
`elc engram/src/server.el > engram.c` then
|
||||
`cc -std=c11 -O2 -I lang/runtime engram.c el_runtime.c engram_*.c -lcurl -lpthread`.
|
||||
|
||||
## Lifecycle
|
||||
|
||||
| op | what it does |
|
||||
|----|--------------|
|
||||
| `create <name> [--port N] [--source\|--branch\|--binary]` | consistent snapshot of the live store+WAL+config into an isolated dir; place or **build** the runtime; boot the real engram daemon on an isolated port. Named, versioned (binary sha + egm sha in `manifest.json`), reproducible. |
|
||||
| `up [name]` | one command: create-if-missing then start; prints the URL. |
|
||||
| `build <name> --source\|--branch` | rebuild the runtime from a code change and hot-restart on the same clone+port. |
|
||||
| `run <name> <cmd…>` / `run <name> api <path> [json]` | run an experiment against the real runtime; capture output + before/after stats + wall time. Env: `$SBX_URL $SBX_PORT $SBX_KEY $SBX_DATA $SBX_BIN`. |
|
||||
| `validate <name>` | the rails as first-class checks (below). |
|
||||
| `promote <name> [--data] [--i-approve-prod-cutover]` | **the only prod-touching op.** Gated rails cutover. DRY-RUN plan unless approved. |
|
||||
| `destroy <name>` | stop the isolated daemon, free the port, remove the clone. Live untouched. |
|
||||
| `list` / `status <name>` | inspect. |
|
||||
|
||||
## `validate` — the rails as checks
|
||||
|
||||
- **zero-loss-under-load** — node/edge counts hold at/above baseline through ~15s of sustained tick+read load
|
||||
- **reboot-prove** — counts survive a real stop→start of the daemon
|
||||
- **rss-bound** — daemon RSS under `NSBX_RSS_BOUND_MB` (default 550 MB, from the store-fix reboot-proof)
|
||||
- **retrieval-parity** — top-k node ids for a fixed probe set match the create-time baseline
|
||||
- **keystone-integrity** — `kn-efeb4a5b…` and `kn-5b606390…` present and intact
|
||||
|
||||
A PASS writes `validate.json` stamped with the binary sha; `promote` refuses unless
|
||||
the current binary has a fresh PASS on record.
|
||||
|
||||
## `promote` — gated cutover (rails only)
|
||||
|
||||
Default is a **dry-run plan**. With `--i-approve-prod-cutover` it, in order:
|
||||
|
||||
1. **snapshot-first** — back up live `egm`+`wal`+`plist` to `~/.neuron/backups/promote-<name>-<ts>/` with a `rollback.txt`
|
||||
2. **additive** binary install — copy the validated binary to a *new* file, update the plist `ENGRAM_REAL_BIN` (old binary retained — additive/supersede, never destructive)
|
||||
3. **rails cutover** — `launchctl bootout` → **settle-poll** (prints until the job is gone) → `launchctl bootstrap`. Never `pkill`, never `kickstart -k`.
|
||||
4. **verify** — `/api/stats` returns, edges ≥ baseline, keystones intact
|
||||
5. **auto-rollback armed** — any verify failure restores the plist (and data, if `--data`) and boots the prior binary back via the same rails
|
||||
|
||||
## Isolation guarantees
|
||||
|
||||
- separate **port** (`8900+`; refuses `8742`/`7770`), separate **store clone**, separate **process**
|
||||
- a hard guard refuses to boot a sandbox daemon whose data dir resolves to the live store
|
||||
- sandboxes are plain supervised background processes (not launchd), so teardown is a signal + settle-poll — it can never touch the prod launchd job
|
||||
- prod is read exactly twice: once for the snapshot, and (only if you approve) during `promote`
|
||||
|
||||
## Layout
|
||||
|
||||
- tool: `tools/neuron-sandbox/nsbx` (this repo, branch `feat/neuron-sandbox`)
|
||||
- runtime state: `~/.neuron/sandboxes/<name>/` — `data/` (clone), `bin/engram`, `build/`, `logs/`, `manifest.json`, `validate.json`, `baseline/`
|
||||
|
||||
## Validated (dogfood)
|
||||
|
||||
Standing up a sandbox from a live-store clone and reproducing a **known** result:
|
||||
|
||||
- **retrieval-parity 25/25** top-k id overlap vs baseline; sandbox boot-stats exactly matched the live baseline captured at snapshot time (10 672 nodes / 32 439 edges) — the wrapped real binary faithfully reloads the live mind
|
||||
- reboot-prove + zero-loss PASS; RSS 379 MB < 550 MB; keystones intact
|
||||
- the **cog-arch correspondence-loop** re-run *inside* the sandbox reproduced the known calibration numbers exactly: held-Brier **0.028648 → 0.000586** (98.0% reduction), monotone, **reboot bit-identical**, metastability holds; and the real-store Stance persistence reboot-proved at **10 994-node** scale (`think()` on real 768-dim embeddings) against a scratch copy of the sandbox's own clone — never live
|
||||
- `promote` dry-run refused to touch prod; teardown freed the port; live `:8742`/`:7770` never perturbed (soul uptime unbroken)
|
||||
|
||||
## Migrating existing experiments
|
||||
|
||||
Each ad-hoc harness becomes `nsbx run <name> …` (or `--source` build) against a sandbox:
|
||||
|
||||
- **cog-arch** — `nsbx create x --source <worktree>` then `nsbx run x -- bash cogarch_dogfood.sh` (compiles + runs the real C cognition tests against `$SBX_DATA`)
|
||||
- **codec / ingest / faculty** — `nsbx run x api /api/<endpoint> '<json>'` against the isolated daemon, or a script using `$SBX_URL`/`$SBX_KEY`; measure with the built-in before/after stats
|
||||
|
||||
## Env knobs
|
||||
|
||||
`NSBX_ROOT`, `NSBX_PORT_BASE`, `NSBX_RSS_BOUND_MB`, `NSBX_REMERGE_THRESHOLD`,
|
||||
`EL_REPO` (for `elc` + runtime sources), `ENGRAM_LIVE_DATA_DIR`, `ENGRAM_LIVE_PLIST`.
|
||||
@@ -1,30 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# cog-arch correspondence-loop dogfood — RUN INSIDE the sandbox via `nsbx run`.
|
||||
# Compiles the REAL engram C runtime + cognition tests and reproduces the known
|
||||
# calibration result (memory 194c69c8): held-Brier 0.028648 -> 0.000586, reboot-proven,
|
||||
# then reboot-proves the Stance persistence against a SCRATCH COPY of THIS sandbox's
|
||||
# clone of the real store (never live, never the running daemon's file).
|
||||
set -euo pipefail
|
||||
WT="${COGARCH_WT:-/private/tmp/claude-501/-Users-will/6531446d-bc27-4095-930b-e04777c3db4f/scratchpad/cogarch-wt}"
|
||||
RT="$WT/lang/runtime"; T="$WT/engram/test"
|
||||
: "${SBX_DATA:?run me via: nsbx run <name> -- bash cogarch_dogfood.sh}"
|
||||
B="$(mktemp -d)"
|
||||
echo "### building cog-arch tests against the real engram runtime sources"
|
||||
cc -std=c11 -O2 -w -I "$RT" -o "$B/test_cognition" \
|
||||
"$T/test_cognition.c" "$RT/engram_cognition.c" "$RT/engram_reason.c" \
|
||||
"$RT/engram_geometry.c" "$RT/engram_store.c" "$RT/engram_vindex.c" -lm
|
||||
cc -std=c11 -O2 -w -I "$RT" -o "$B/test_realstore" \
|
||||
"$T/test_cognition_realstore.c" "$RT/engram_cognition.c" "$RT/engram_reason.c" \
|
||||
"$RT/engram_geometry.c" "$RT/engram_store.c" "$RT/engram_vindex.c" -lm
|
||||
|
||||
echo; echo "### [A] synthetic correspondence-loop (known: Brier 0.028648 -> 0.000586)"
|
||||
"$B/test_cognition" | grep -E "held-Brier|reduction|reboot|monotone|metastab|RESULT" || true
|
||||
|
||||
echo; echo "### [B] reboot-prove Stance on a SCRATCH COPY of this sandbox's real-store clone"
|
||||
SCRATCH="$B/store-clone"; mkdir -p "$SCRATCH"
|
||||
cp -p "$SBX_DATA/neuron.egm" "$SCRATCH/" 2>/dev/null || true
|
||||
cp -p "$SBX_DATA/neuron.wal" "$SCRATCH/" 2>/dev/null || true
|
||||
cp -p "$SBX_DATA/conf" "$SCRATCH/" 2>/dev/null || true
|
||||
cp -p "$SBX_DATA/meta.json" "$SCRATCH/" 2>/dev/null || true
|
||||
"$B/test_realstore" "$SCRATCH" || true
|
||||
rm -rf "$B"
|
||||
@@ -1,663 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# nsbx — the Neuron Sandbox: a reproducible primitive for running experiments and
|
||||
# code changes against the REAL engram runtime on an isolated snapshot of the live
|
||||
# mind, with a gated promote-to-prod path built on the proven rails.
|
||||
#
|
||||
# It WRAPS the real engram binary — it never reimplements any engram logic. The only
|
||||
# prod-touching op is `promote`, which is explicit, gated, and per-use approved.
|
||||
#
|
||||
# Generalises two proven proto-sandboxes:
|
||||
# - the cog-arch build (isolated git worktree + build + clone of live .egm + real C tests)
|
||||
# - the store-fix cutover (secondary soul + launchctl bootout->settle->bootstrap rails)
|
||||
#
|
||||
# Lifecycle: create -> [build] -> run -> validate -> promote(gated) -> destroy
|
||||
#
|
||||
# Rails (always): built offline; NEVER auto-promotes; never touches live :8742/:7770
|
||||
# except READ for the snapshot and the gated promote; snapshot-first; honest measured
|
||||
# reporting. Cutover is launchctl bootout -> settle-poll -> bootstrap ONLY —
|
||||
# never pkill, never kickstart -k.
|
||||
set -uo pipefail
|
||||
|
||||
# ---------------------------------------------------------------- constants ----
|
||||
LIVE_DATA_DIR="${ENGRAM_LIVE_DATA_DIR:-$HOME/.neuron/engram}"
|
||||
LIVE_PLIST="${ENGRAM_LIVE_PLIST:-$HOME/Library/LaunchAgents/ai.neuron.engram.plist}"
|
||||
LIVE_LABEL="ai.neuron.engram"
|
||||
LIVE_BIND_PORT=8742 # engram — FORBIDDEN for sandboxes
|
||||
SOUL_PORT=7770 # soul — FORBIDDEN for sandboxes
|
||||
LIVE_KEY="${ENGRAM_API_KEY:-ntn-user-2026}"
|
||||
LIVE_URL="http://127.0.0.1:${LIVE_BIND_PORT}"
|
||||
SBX_ROOT="${NSBX_ROOT:-$HOME/.neuron/sandboxes}"
|
||||
BACKUP_ROOT="$HOME/.neuron/backups"
|
||||
EL_REPO="${EL_REPO:-$HOME/Development/neuron-technologies/foundation/el}"
|
||||
PORT_BASE="${NSBX_PORT_BASE:-8900}"
|
||||
RSS_BOUND_MB="${NSBX_RSS_BOUND_MB:-550}" # from store-fix reboot-proof (aaf13f88)
|
||||
REMERGE_THRESHOLD="${NSBX_REMERGE_THRESHOLD:-40000}"
|
||||
KEYSTONES=( "kn-efeb4a5b-5aff-4759-8a97-7233099be6ee" "kn-5b606390-a52d-4ca2-8e0e-eba141d13440" )
|
||||
# fixed probe set for retrieval-parity (stable, identity-anchored)
|
||||
PARITY_QUERIES=( "who am I" "self identity core" "engram store durability" "keystone self anchor" "grounding honesty" )
|
||||
|
||||
C_RED=$'\033[31m'; C_GRN=$'\033[32m'; C_YEL=$'\033[33m'; C_DIM=$'\033[2m'; C_BLD=$'\033[1m'; C_0=$'\033[0m'
|
||||
|
||||
# ---------------------------------------------------------------- helpers ------
|
||||
die(){ printf '%serror:%s %s\n' "$C_RED" "$C_0" "$*" >&2; exit 1; }
|
||||
log(){ printf '%s==>%s %s\n' "$C_BLD" "$C_0" "$*" >&2; }
|
||||
info(){ printf ' %s\n' "$*" >&2; }
|
||||
ok(){ printf ' %s%s%s\n' "$C_GRN" "$*" "$C_0" >&2; }
|
||||
warn(){ printf ' %s%s%s\n' "$C_YEL" "$*" "$C_0" >&2; }
|
||||
need(){ command -v "$1" >/dev/null 2>&1 || die "missing dependency: $1"; }
|
||||
now(){ date -u +%Y%m%dT%H%M%SZ; }
|
||||
sha(){ shasum -a 256 "$1" 2>/dev/null | awk '{print $1}'; }
|
||||
epoch(){ python3 -c 'import time;print(time.time())'; }
|
||||
|
||||
sdir(){ printf '%s/%s' "$SBX_ROOT" "$1"; }
|
||||
manifest(){ printf '%s/manifest.json' "$(sdir "$1")"; }
|
||||
mexists(){ [ -f "$(manifest "$1")" ]; }
|
||||
mget(){ # mget <name> <jsonpath>
|
||||
python3 -c "import json,sys; d=json.load(open('$(manifest "$1")')); print(d$2)" 2>/dev/null
|
||||
}
|
||||
|
||||
port_free(){ ! (exec 3<>"/dev/tcp/127.0.0.1/$1") 2>/dev/null; }
|
||||
alloc_port(){
|
||||
local p="$PORT_BASE"
|
||||
while :; do
|
||||
if [ "$p" = "$LIVE_BIND_PORT" ] || [ "$p" = "$SOUL_PORT" ]; then p=$((p+1)); continue; fi
|
||||
if port_free "$p" && ! _port_claimed "$p"; then echo "$p"; return 0; fi
|
||||
p=$((p+1)); [ "$p" -gt 9100 ] && die "no free sandbox port in range"
|
||||
done
|
||||
}
|
||||
_port_claimed(){ # is another sandbox already assigned this port?
|
||||
local p="$1" d
|
||||
for d in "$SBX_ROOT"/*/manifest.json; do
|
||||
[ -f "$d" ] || continue
|
||||
[ "$(python3 -c "import json;print(json.load(open('$d'))['port'])" 2>/dev/null)" = "$p" ] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
live_stats(){ curl -s -m5 "$LIVE_URL/api/stats" 2>/dev/null; }
|
||||
api(){ # api <name> <path> [json-body]
|
||||
local name="$1" path="$2" body="${3:-}"
|
||||
local port; port="$(mget "$name" "['port']")"; [ -n "$port" ] || die "unknown sandbox: $name"
|
||||
local url="http://127.0.0.1:${port}${path}"
|
||||
if [ -n "$body" ]; then curl -s -m30 -X POST -H 'Content-Type: application/json' -d "$body" "$url"
|
||||
else curl -s -m30 "$url"; fi
|
||||
}
|
||||
sbx_stats(){ api "$1" "/api/stats"; }
|
||||
stat_field(){ printf '%s' "$1" | sed -n "s/.*\"$2\":\([0-9]*\).*/\1/p"; }
|
||||
|
||||
daemon_pid(){ local f; f="$(sdir "$1")/daemon.pid"; [ -f "$f" ] && cat "$f" || true; }
|
||||
daemon_alive(){ local p; p="$(daemon_pid "$1")"; [ -n "$p" ] && kill -0 "$p" 2>/dev/null; }
|
||||
|
||||
# ---------------------------------------------------------------- elc/build ----
|
||||
find_elc(){
|
||||
command -v elc 2>/dev/null && return 0
|
||||
local arch; arch="$(uname -m)"
|
||||
case "$arch" in
|
||||
arm64) echo "$EL_REPO/lang/dist/platform/elc-darwin-arm64";;
|
||||
x86_64) echo "$EL_REPO/lang/dist/platform/elc-linux-amd64";;
|
||||
*) echo "$EL_REPO/lang/dist/platform/elc";;
|
||||
esac
|
||||
}
|
||||
|
||||
# _build_binary <src_tree> <out_bin> <build_log_dir>
|
||||
# Replicates the proven engram release recipe:
|
||||
# elc engram/src/server.el > engram.c
|
||||
# cc -std=c11 -O2 -I lang/runtime engram.c el_runtime.c engram_*.c -lcurl -lpthread
|
||||
_build_binary(){
|
||||
local src="$1" out="$2" blog="$3"
|
||||
local elc server rt
|
||||
elc="$(find_elc)"; [ -x "$elc" ] || die "elc not found/executable: $elc (set EL_REPO)"
|
||||
server="$src/engram/src/server.el"; rt="$src/lang/runtime"
|
||||
[ -f "$server" ] || die "no engram/src/server.el under source tree: $src"
|
||||
[ -f "$rt/el_runtime.c" ] || die "no lang/runtime/el_runtime.c under source tree: $src (this branch may keep it generated/untracked)"
|
||||
ls "$rt"/engram_*.c >/dev/null 2>&1 || die "no lang/runtime/engram_*.c engine sources under: $src"
|
||||
mkdir -p "$blog"
|
||||
log "build: elc transpile server.el -> engram.c"
|
||||
"$elc" "$server" > "$blog/engram.c" 2>"$blog/elc.err" || { cat "$blog/elc.err" >&2; die "elc transpile failed"; }
|
||||
info "engram.c: $(wc -c <"$blog/engram.c" | tr -d ' ') bytes"
|
||||
log "build: cc link (el_runtime + engram_* engine)"
|
||||
cc -std=c11 -O2 -w -I "$rt" -o "$out" \
|
||||
"$blog/engram.c" "$rt/el_runtime.c" "$rt"/engram_*.c \
|
||||
-lcurl -lpthread 2>"$blog/cc.err" \
|
||||
|| { grep -i 'error:' "$blog/cc.err" | sort -u | head >&2; die "cc link failed (see $blog/cc.err)"; }
|
||||
ok "built: $out ($(ls -lh "$out" | awk '{print $5}'), sha $(sha "$out" | cut -c1-12))"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------- daemon -------
|
||||
# start_daemon <name> : boots the sandbox's real engram binary on its isolated
|
||||
# port against its cloned data dir, with the SAME auto-remerge net the live soul
|
||||
# uses (so the sandbox faithfully reaches the live edge population on boot).
|
||||
start_daemon(){
|
||||
local name="$1" d; d="$(sdir "$name")"
|
||||
daemon_alive "$name" && { info "already running (pid $(daemon_pid "$name"))"; return 0; }
|
||||
local port bin data export key
|
||||
port="$(mget "$name" "['port']")"; bin="$d/bin/engram"; data="$d/data"
|
||||
key="sbx-$name"; export="$data/.scan-export.reseed-clean.json"
|
||||
[ -x "$bin" ] || die "sandbox binary missing: $bin"
|
||||
[ "$port" != "$LIVE_BIND_PORT" ] && [ "$port" != "$SOUL_PORT" ] || die "refusing forbidden port $port"
|
||||
[ -f "$data/neuron.egm" ] || die "sandbox has no cloned store: $data/neuron.egm"
|
||||
# HARD guard: never point a sandbox daemon at the live data dir.
|
||||
[ "$(cd "$data" && pwd -P)" != "$(cd "$LIVE_DATA_DIR" && pwd -P)" ] || die "refusing: sandbox data dir resolves to LIVE store"
|
||||
|
||||
log "boot engram on isolated :$port (data=$data)"
|
||||
(
|
||||
ENGRAM_DATA_DIR="$data" ENGRAM_BIND=":$port" ENGRAM_API_KEY="$key" \
|
||||
ENGRAM_STORE=1 ENGRAM_CHRONOCEPTION=1 ENGRAM_SELF_REIFY=1 ENGRAM_GC=1 \
|
||||
ENGRAM_POOL_FRAMES=16384 ENGRAM_WRITE_BARRIER=1 \
|
||||
exec "$bin"
|
||||
) >"$d/logs/daemon.log" 2>&1 &
|
||||
local pid=$!
|
||||
echo "$pid" > "$d/daemon.pid"
|
||||
# readiness poll
|
||||
local url="http://127.0.0.1:$port" i s
|
||||
for i in $(seq 1 30); do
|
||||
s="$(curl -s -m3 "$url/api/stats" 2>/dev/null)"
|
||||
[ -n "$s" ] && break; sleep 0.5
|
||||
done
|
||||
[ -n "$s" ] || { warn "daemon did not become ready (see $d/logs/daemon.log)"; return 1; }
|
||||
ok "ready pid=$pid boot-stats: $s"
|
||||
# auto-remerge net (idempotent): match live edge population if the export is present
|
||||
if [ -f "$export" ]; then
|
||||
local edges; edges="$(stat_field "$s" edge_count)"
|
||||
if [ -n "$edges" ] && [ "$edges" -lt "$REMERGE_THRESHOLD" ]; then
|
||||
log "auto-remerge: booted with $edges edges (< $REMERGE_THRESHOLD) — merging full edge export"
|
||||
local r; r="$(curl -s -m300 -X POST -H 'Content-Type: application/json' \
|
||||
-d "{\"_auth\":\"$key\",\"path\":\"$export\"}" "$url/api/load-merge" 2>/dev/null)"
|
||||
info "remerge resp: ${r:0:120}"
|
||||
ok "post-remerge stats: $(curl -s -m5 "$url/api/stats")"
|
||||
fi
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# stop_daemon <name> : graceful TERM + settle-poll until the port is free.
|
||||
# (Sandbox daemons are plain supervised bg processes — not launchd — so teardown
|
||||
# is a signal + poll, never pkill of anything else.)
|
||||
stop_daemon(){
|
||||
local name="$1" pid port
|
||||
pid="$(daemon_pid "$name")"; port="$(mget "$name" "['port']")"
|
||||
[ -n "$pid" ] || { info "not running"; return 0; }
|
||||
log "stop daemon pid=$pid, settle-poll until :$port frees"
|
||||
kill "$pid" 2>/dev/null || true
|
||||
local i
|
||||
for i in $(seq 1 40); do
|
||||
kill -0 "$pid" 2>/dev/null || { port_free "$port" && { ok "stopped, port $port free"; : >"$(sdir "$name")/daemon.pid"; return 0; }; }
|
||||
printf '.' >&2; sleep 0.5
|
||||
done
|
||||
printf '\n' >&2
|
||||
kill -9 "$pid" 2>/dev/null || true; sleep 1
|
||||
: >"$(sdir "$name")/daemon.pid"
|
||||
port_free "$port" && ok "stopped (after SIGKILL), port $port free" || warn "port $port still busy"
|
||||
}
|
||||
|
||||
# ================================================================ create =======
|
||||
cmd_create(){
|
||||
local name="" port="" src="" branch="" repo="$EL_REPO" binpath=""
|
||||
# first positional arg is the name unless it's a flag; default to "<user>-dev"
|
||||
if [ $# -gt 0 ] && [ "${1#-}" = "$1" ]; then name="$1"; shift; else name="${USER:-dev}-dev"; fi
|
||||
while [ $# -gt 0 ]; do case "$1" in
|
||||
--port) port="$2"; shift 2;;
|
||||
--source) src="$2"; shift 2;;
|
||||
--branch) branch="$2"; shift 2;;
|
||||
--repo) repo="$2"; shift 2;;
|
||||
--binary) binpath="$2"; shift 2;;
|
||||
*) die "unknown flag: $1";;
|
||||
esac; done
|
||||
mexists "$name" && die "sandbox '$name' already exists (destroy it first)"
|
||||
need curl; need python3; need shasum
|
||||
[ -f "$LIVE_DATA_DIR/neuron.egm" ] || die "live store not found: $LIVE_DATA_DIR/neuron.egm"
|
||||
if [ -n "$port" ]; then
|
||||
{ [ "$port" = "$LIVE_BIND_PORT" ] || [ "$port" = "$SOUL_PORT" ]; } && die "refusing forbidden port $port (live)"
|
||||
port_free "$port" || die "port $port already in use"
|
||||
else port="$(alloc_port)"; fi
|
||||
|
||||
local d; d="$(sdir "$name")"
|
||||
mkdir -p "$d/data" "$d/bin" "$d/logs" "$d/build" "$d/baseline"
|
||||
log "sandbox '$name' at $d (isolated port $port)"
|
||||
|
||||
# ---- CONSISTENT snapshot of the live mind (file-copy: same set the rails backup
|
||||
# uses; WAL replay on sandbox boot reconciles the tail -> crash-consistent) ----
|
||||
log "snapshot live store -> clone (store + WAL + config)"
|
||||
local f
|
||||
for f in neuron.egm neuron.wal conf meta.json self_anchor .scan-export.reseed-clean.json; do
|
||||
if [ -e "$LIVE_DATA_DIR/$f" ]; then cp -p "$LIVE_DATA_DIR/$f" "$d/data/$f"; info "cloned $f ($(du -h "$d/data/$f" | awk '{print $1}'))"; fi
|
||||
done
|
||||
local egm_sha; egm_sha="$(sha "$d/data/neuron.egm")"
|
||||
|
||||
# ---- capture live baseline (READ only) ----
|
||||
local lstats; lstats="$(live_stats)"
|
||||
local base_nodes base_edges
|
||||
base_nodes="$(stat_field "$lstats" node_count)"; base_edges="$(stat_field "$lstats" edge_count)"
|
||||
info "live baseline stats: ${lstats:-<unavailable>}"
|
||||
|
||||
# ---- determine + place the runtime binary (versioned into the snapshot) ----
|
||||
local source_desc live_bin
|
||||
live_bin="$(_live_real_bin)"
|
||||
if [ -n "$binpath" ]; then
|
||||
[ -x "$binpath" ] || die "not an executable binary: $binpath"
|
||||
cp -p "$binpath" "$d/bin/engram"; source_desc="prebuilt:$binpath"
|
||||
elif [ -n "$src" ]; then
|
||||
_build_binary "$src" "$d/bin/engram" "$d/build"; source_desc="source:$src"
|
||||
elif [ -n "$branch" ]; then
|
||||
log "worktree: $repo @ $branch -> $d/build/worktree"
|
||||
git -C "$repo" worktree add --detach "$d/build/worktree" "$branch" >/dev/null 2>&1 \
|
||||
|| die "git worktree add failed ($repo @ $branch)"
|
||||
_build_binary "$d/build/worktree" "$d/bin/engram" "$d/build"; source_desc="branch:$branch@$repo"
|
||||
else
|
||||
[ -x "$live_bin" ] || die "cannot resolve live ENGRAM_REAL_BIN: $live_bin"
|
||||
cp -p "$live_bin" "$d/bin/engram"; source_desc="stock-prod:$live_bin"
|
||||
fi
|
||||
local bin_sha; bin_sha="$(sha "$d/bin/engram")"
|
||||
info "runtime: $source_desc (sha ${bin_sha:0:12})"
|
||||
|
||||
# ---- write manifest ----
|
||||
python3 - "$name" "$port" "$source_desc" "$bin_sha" "$egm_sha" "$base_nodes" "$base_edges" "$(sha "$live_bin" 2>/dev/null)" <<'PY' > "$(manifest "$name")"
|
||||
import json,sys,datetime
|
||||
name,port,src,binsha,egmsha,bn,be,livebinsha=sys.argv[1:9]
|
||||
json.dump({
|
||||
"name":name,"port":int(port),"created_at":datetime.datetime.now(datetime.timezone.utc).isoformat(),
|
||||
"source":src,"binary_sha256":binsha,"clone_egm_sha256":egmsha,
|
||||
"live_binary_sha256":livebinsha,
|
||||
"live_baseline":{"node_count":int(bn or 0),"edge_count":int(be or 0)},
|
||||
"keystones":["kn-efeb4a5b-5aff-4759-8a97-7233099be6ee","kn-5b606390-a52d-4ca2-8e0e-eba141d13440"]
|
||||
}, sys.stdout, indent=2)
|
||||
PY
|
||||
ok "manifest written"
|
||||
|
||||
# ---- boot + capture the sandbox's own settled baseline (reproducible target) ----
|
||||
start_daemon "$name" || die "daemon failed to start"
|
||||
local sstats; sstats="$(sbx_stats "$name")"
|
||||
local sbn sbe; sbn="$(stat_field "$sstats" node_count)"; sbe="$(stat_field "$sstats" edge_count)"
|
||||
_capture_retrieval "$name" "$d/baseline/retrieval.json"
|
||||
# fold sandbox baseline into manifest
|
||||
python3 - "$(manifest "$name")" "$sbn" "$sbe" <<'PY'
|
||||
import json,sys
|
||||
mf,bn,be=sys.argv[1],sys.argv[2],sys.argv[3]
|
||||
d=json.load(open(mf)); d["sbx_baseline"]={"node_count":int(bn or 0),"edge_count":int(be or 0)}
|
||||
json.dump(d,open(mf,'w'),indent=2)
|
||||
PY
|
||||
log "created."
|
||||
info "sandbox baseline (settled): nodes=$sbn edges=$sbe"
|
||||
info "next: nsbx validate $name | nsbx run $name api /api/stats"
|
||||
}
|
||||
|
||||
_live_real_bin(){
|
||||
python3 - "$LIVE_PLIST" <<'PY' 2>/dev/null
|
||||
import sys,plistlib
|
||||
try:
|
||||
d=plistlib.load(open(sys.argv[1],'rb'))
|
||||
print(d.get("EnvironmentVariables",{}).get("ENGRAM_REAL_BIN",""))
|
||||
except Exception: print("")
|
||||
PY
|
||||
}
|
||||
|
||||
_capture_retrieval(){ # <name> <outfile> : top-k ids for the fixed probe set
|
||||
local name="$1" out="$2" q res
|
||||
local port; port="$(mget "$name" "['port']")"; local key="sbx-$name"
|
||||
{
|
||||
echo "{"
|
||||
local first=1
|
||||
for q in "${PARITY_QUERIES[@]}"; do
|
||||
res="$(curl -s -m10 -X POST -H 'Content-Type: application/json' \
|
||||
-d "{\"_auth\":\"$key\",\"query\":\"$q\",\"limit\":5}" "http://127.0.0.1:$port/api/search" 2>/dev/null)"
|
||||
local ids; ids="$(printf '%s' "$res" | python3 -c 'import sys,json
|
||||
try:
|
||||
d=json.load(sys.stdin)
|
||||
rows=d if isinstance(d,list) else d.get("results",d.get("hits",[]))
|
||||
print(json.dumps([r.get("id") for r in rows][:5]))
|
||||
except Exception: print("[]")' 2>/dev/null)"
|
||||
[ $first -eq 1 ] || echo ","; first=0
|
||||
printf ' %s: %s' "$(python3 -c "import json,sys;print(json.dumps(sys.argv[1]))" "$q")" "${ids:-[]}"
|
||||
done
|
||||
echo ""; echo "}"
|
||||
} > "$out"
|
||||
}
|
||||
|
||||
# ================================================================ up ===========
|
||||
# Dead-simple one-command dev environment: `nsbx up` gives you (or Tim, or anyone)
|
||||
# a private, isolated copy of the live mind to build against. Creates it on first
|
||||
# run with sane defaults (stock prod binary, auto-allocated port), just starts it
|
||||
# thereafter. Prod on :$LIVE_BIND_PORT/:$SOUL_PORT is unreachable from here by design.
|
||||
cmd_up(){
|
||||
local name; if [ $# -gt 0 ] && [ "${1#-}" = "$1" ]; then name="$1"; shift; else name="${USER:-dev}-dev"; fi
|
||||
if mexists "$name"; then daemon_alive "$name" || start_daemon "$name"; else cmd_create "$name" "$@"; fi
|
||||
local port; port="$(mget "$name" "['port']")"
|
||||
echo >&2
|
||||
ok "your sandbox '$name' is ready at http://127.0.0.1:$port (a private copy of the mind — prod is untouchable)"
|
||||
info "experiment: nsbx run $name api /api/stats"
|
||||
info "prove it: nsbx validate $name"
|
||||
info "tear down: nsbx destroy $name"
|
||||
}
|
||||
|
||||
# ================================================================ build ========
|
||||
# Rebuild an existing sandbox's runtime from a source tree/branch and hot-restart
|
||||
# it on the SAME clone + port (the code-change dev loop, in place).
|
||||
cmd_build(){
|
||||
local name="$1"; shift || true
|
||||
mexists "$name" || die "no such sandbox: $name"
|
||||
local src="" branch="" repo="$EL_REPO"
|
||||
while [ $# -gt 0 ]; do case "$1" in
|
||||
--source) src="$2"; shift 2;; --branch) branch="$2"; shift 2;; --repo) repo="$2"; shift 2;;
|
||||
*) die "unknown flag: $1";; esac; done
|
||||
local d; d="$(sdir "$name")"
|
||||
stop_daemon "$name"
|
||||
if [ -n "$src" ]; then _build_binary "$src" "$d/bin/engram" "$d/build"
|
||||
elif [ -n "$branch" ]; then
|
||||
rm -rf "$d/build/worktree" 2>/dev/null; git -C "$repo" worktree prune 2>/dev/null
|
||||
git -C "$repo" worktree add --detach "$d/build/worktree" "$branch" >/dev/null 2>&1 || die "worktree add failed"
|
||||
_build_binary "$d/build/worktree" "$d/bin/engram" "$d/build"
|
||||
else die "usage: nsbx build <name> --source DIR | --branch REF [--repo R]"; fi
|
||||
# record new binary sha
|
||||
python3 - "$(manifest "$name")" "$(sha "$d/bin/engram")" "${src:-branch:$branch}" <<'PY'
|
||||
import json,sys; mf,s,src=sys.argv[1:4]
|
||||
d=json.load(open(mf)); d["binary_sha256"]=s; d["source"]="rebuilt:"+src
|
||||
json.dump(d,open(mf,'w'),indent=2)
|
||||
PY
|
||||
start_daemon "$name"
|
||||
ok "rebuilt + restarted on :$(mget "$name" "['port']")"
|
||||
}
|
||||
|
||||
# ================================================================ run ==========
|
||||
cmd_run(){
|
||||
local name="$1"; shift || true
|
||||
mexists "$name" || die "no such sandbox: $name"
|
||||
daemon_alive "$name" || start_daemon "$name"
|
||||
local d port; d="$(sdir "$name")"; port="$(mget "$name" "['port']")"
|
||||
# direct API form: nsbx run <name> api <path> [json]
|
||||
if [ "${1:-}" = "api" ]; then
|
||||
api "$name" "$2" "${3:-}"; echo; return 0
|
||||
fi
|
||||
[ "${1:-}" = "--" ] && shift # allow an explicit separator: nsbx run <name> -- <cmd...>
|
||||
[ $# -gt 0 ] || die "usage: nsbx run <name> <cmd...> | nsbx run <name> api <path> [json]"
|
||||
local ts log0; ts="$(now)"; log0="$d/logs/run-$ts.log"
|
||||
local s0 t0 t1 s1
|
||||
s0="$(sbx_stats "$name")"; t0="$(epoch)"
|
||||
log "run experiment against sandbox '$name' (:$port)"
|
||||
info "cmd: $*"
|
||||
( export SBX_NAME="$name" SBX_PORT="$port" SBX_URL="http://127.0.0.1:$port" \
|
||||
SBX_KEY="sbx-$name" SBX_DATA="$d/data" SBX_BIN="$d/bin/engram"
|
||||
"$@" ) 2>&1 | tee "$log0"
|
||||
local rc=${PIPESTATUS[0]}
|
||||
t1="$(epoch)"; s1="$(sbx_stats "$name")"
|
||||
{
|
||||
echo "--- nsbx run metrics ---"
|
||||
echo "exit_code: $rc"
|
||||
printf 'wall_secs: %.3f\n' "$(python3 -c "print($t1-$t0)")"
|
||||
echo "stats_before: $s0"
|
||||
echo "stats_after: $s1"
|
||||
} | tee -a "$log0" >&2
|
||||
return $rc
|
||||
}
|
||||
|
||||
# ================================================================ validate =====
|
||||
# The rails as first-class checks. Baseline = the sandbox's own settled state at
|
||||
# create (reproducible). zero-loss through sustained load AND reboot; reboot-prove;
|
||||
# RSS bound; retrieval parity; keystone integrity.
|
||||
cmd_validate(){
|
||||
local name="$1"; shift || true
|
||||
mexists "$name" || die "no such sandbox: $name"
|
||||
daemon_alive "$name" || start_daemon "$name"
|
||||
local d port key; d="$(sdir "$name")"; port="$(mget "$name" "['port']")"; key="sbx-$name"
|
||||
local url="http://127.0.0.1:$port"
|
||||
local bn be; bn="$(mget "$name" "['sbx_baseline']['node_count']")"; be="$(mget "$name" "['sbx_baseline']['edge_count']")"
|
||||
log "validate '$name' against baseline nodes=$bn edges=$be"
|
||||
local -a names=() results=() details=()
|
||||
|
||||
# 1) sustained load — no data loss under activity
|
||||
local s cur_n cur_e i
|
||||
log "check: sustained load (~15s: tick + reads) then zero-loss"
|
||||
for i in $(seq 1 15); do
|
||||
curl -s -m5 -X POST -H 'Content-Type: application/json' -d "{\"_auth\":\"$key\"}" "$url/api/tick" >/dev/null 2>&1
|
||||
curl -s -m5 "$url/api/stats" >/dev/null 2>&1
|
||||
done
|
||||
s="$(sbx_stats "$name")"; cur_n="$(stat_field "$s" node_count)"; cur_e="$(stat_field "$s" edge_count)"
|
||||
names+=("zero-loss-under-load"); if [ "${cur_n:-0}" -ge "${bn:-0}" ] && [ "${cur_e:-0}" -ge "${be:-0}" ]; then
|
||||
results+=("PASS"); else results+=("FAIL"); fi
|
||||
details+=("nodes $cur_n>=$bn, edges $cur_e>=$be")
|
||||
|
||||
# 2) reboot-prove — counts survive a real restart
|
||||
log "check: reboot-prove (stop -> start -> compare)"
|
||||
local pre_n pre_e; pre_n="$cur_n"; pre_e="$cur_e"
|
||||
stop_daemon "$name"; start_daemon "$name" >/dev/null
|
||||
s="$(sbx_stats "$name")"; cur_n="$(stat_field "$s" node_count)"; cur_e="$(stat_field "$s" edge_count)"
|
||||
names+=("reboot-prove"); if [ "${cur_n:-0}" -ge "${bn:-0}" ] && [ "${cur_e:-0}" -ge "${be:-0}" ]; then
|
||||
results+=("PASS"); else results+=("FAIL"); fi
|
||||
details+=("post-reboot nodes=$cur_n edges=$cur_e (pre $pre_n/$pre_e)")
|
||||
|
||||
# 3) RSS bound
|
||||
log "check: RSS bound (< ${RSS_BOUND_MB}MB)"
|
||||
local pid rss_kb rss_mb; pid="$(daemon_pid "$name")"
|
||||
rss_kb="$(ps -o rss= -p "$pid" 2>/dev/null | tr -d ' ')"; rss_mb=$(( ${rss_kb:-0} / 1024 ))
|
||||
names+=("rss-bound"); if [ "$rss_mb" -lt "$RSS_BOUND_MB" ] && [ "$rss_mb" -gt 0 ]; then results+=("PASS"); else results+=("FAIL"); fi
|
||||
details+=("RSS=${rss_mb}MB (bound ${RSS_BOUND_MB}MB)")
|
||||
|
||||
# 4) retrieval parity vs the create-time baseline
|
||||
log "check: retrieval parity vs baseline probe set"
|
||||
_capture_retrieval "$name" "$d/logs/retrieval-$( now ).json"
|
||||
local latest; latest="$(ls -t "$d/logs"/retrieval-*.json 2>/dev/null | head -1)"
|
||||
local parity; parity="$(python3 - "$d/baseline/retrieval.json" "$latest" <<'PY'
|
||||
import json,sys
|
||||
def load(p):
|
||||
try: return json.load(open(p))
|
||||
except Exception: return {}
|
||||
b,c=load(sys.argv[1]),load(sys.argv[2])
|
||||
tot=hit=0
|
||||
for q,ids in b.items():
|
||||
cb=set(ids or []); cc=set(c.get(q) or [])
|
||||
if not cb: continue
|
||||
tot+=len(cb); hit+=len(cb & cc)
|
||||
print(f"{hit}/{tot}" if tot else "0/0")
|
||||
PY
|
||||
)"
|
||||
local ph="${parity%/*}" pt="${parity#*/}"
|
||||
names+=("retrieval-parity"); if [ "${pt:-0}" -gt 0 ] && [ "${ph:-0}" -eq "${pt:-0}" ]; then results+=("PASS"); else results+=("FAIL"); fi
|
||||
details+=("top-k id overlap $parity vs baseline")
|
||||
|
||||
# 5) keystone integrity
|
||||
log "check: keystone integrity"
|
||||
local kfail=0 kid kres
|
||||
for kid in "${KEYSTONES[@]}"; do
|
||||
kres="$(curl -s -m5 "$url/api/node/$kid" 2>/dev/null)"
|
||||
printf '%s' "$kres" | grep -q "\"$kid\"" || kfail=1
|
||||
done
|
||||
names+=("keystone-integrity"); [ "$kfail" -eq 0 ] && results+=("PASS") || results+=("FAIL")
|
||||
details+=("kn-efeb4a5b + kn-5b606390 present")
|
||||
|
||||
# ---- report + stamp ----
|
||||
echo >&2
|
||||
printf '%s VALIDATION — %s%s\n' "$C_BLD" "$name" "$C_0" >&2
|
||||
local allpass=1 j
|
||||
for j in "${!names[@]}"; do
|
||||
local r="${results[$j]}" c="$C_GRN"; [ "$r" = FAIL ] && { c="$C_RED"; allpass=0; }
|
||||
printf ' %s%-6s%s %-22s %s%s%s\n' "$c" "$r" "$C_0" "${names[$j]}" "$C_DIM" "${details[$j]}" "$C_0" >&2
|
||||
done
|
||||
local status; [ "$allpass" -eq 1 ] && status="PASS" || status="FAIL"
|
||||
python3 - "$d/validate.json" "$status" "$(sha "$d/bin/engram")" "$(now)" "${names[*]}" "${results[*]}" <<'PY'
|
||||
import json,sys
|
||||
out,status,binsha,ts,ns,rs=sys.argv[1:7]
|
||||
checks=[{"name":n,"result":r} for n,r in zip(ns.split(),rs.split())]
|
||||
json.dump({"status":status,"binary_sha256":binsha,"ts":ts,"checks":checks},open(out,'w'),indent=2)
|
||||
PY
|
||||
printf ' %s==> %s%s\n' "$([ "$allpass" -eq 1 ] && echo "$C_GRN" || echo "$C_RED")" "$status" "$C_0" >&2
|
||||
[ "$allpass" -eq 1 ]
|
||||
}
|
||||
|
||||
# ================================================================ promote ======
|
||||
# The ONLY prod-touching op. Explicit, gated, per-use Will-approved. Rails ONLY:
|
||||
# snapshot-first -> additive binary swap -> launchctl bootout -> settle-poll ->
|
||||
# bootstrap -> verify -> auto-rollback on failure. NEVER pkill, NEVER kickstart -k.
|
||||
# Default is a DRY-RUN plan; requires --i-approve-prod-cutover to actually cut over.
|
||||
cmd_promote(){
|
||||
local name="$1"; shift || true
|
||||
mexists "$name" || die "no such sandbox: $name"
|
||||
local approve=0 do_data=0
|
||||
while [ $# -gt 0 ]; do case "$1" in
|
||||
--i-approve-prod-cutover) approve=1; shift;;
|
||||
--data) do_data=1; shift;;
|
||||
*) die "unknown flag: $1";; esac; done
|
||||
local d; d="$(sdir "$name")"
|
||||
# GATE 1: validation must have passed for the CURRENT binary
|
||||
[ -f "$d/validate.json" ] || die "GATE: no validation on record — run 'nsbx validate $name' first"
|
||||
local vstatus vsha bsha
|
||||
vstatus="$(python3 -c "import json;print(json.load(open('$d/validate.json'))['status'])")"
|
||||
vsha="$(python3 -c "import json;print(json.load(open('$d/validate.json'))['binary_sha256'])")"
|
||||
bsha="$(sha "$d/bin/engram")"
|
||||
[ "$vstatus" = PASS ] || die "GATE: last validation status is $vstatus (must be PASS)"
|
||||
[ "$vsha" = "$bsha" ] || die "GATE: validation is stale — binary changed since validate (re-run validate)"
|
||||
|
||||
local live_bin new_bin ts; ts="$(now)"
|
||||
live_bin="$(_live_real_bin)"
|
||||
new_bin="$HOME/.neuron/bin/engram.promote-$name-$ts" # additive: new file, old kept
|
||||
local bkp="$BACKUP_ROOT/promote-$name-$ts"
|
||||
|
||||
log "PROMOTE PLAN for '$name' -> live :$LIVE_BIND_PORT"
|
||||
info "current live ENGRAM_REAL_BIN : $live_bin"
|
||||
info "sandbox binary (validated) : $d/bin/engram (sha ${bsha:0:12})"
|
||||
info "will install as : $new_bin (additive; old binary retained)"
|
||||
info "snapshot-first backup dir : $bkp (egm+wal+plist+rollback.txt)"
|
||||
info "data promote : $([ $do_data -eq 1 ] && echo 'YES (--data: clone egm/wal -> live)' || echo 'no (binary only)')"
|
||||
info "rails : launchctl bootout -> settle-poll -> bootstrap"
|
||||
info "verify : /api/stats + edges>=baseline + keystones + retrieval; auto-rollback armed"
|
||||
|
||||
if [ "$approve" -ne 1 ]; then
|
||||
warn "DRY-RUN — not touching prod. Re-run with --i-approve-prod-cutover to execute (per-use Will-approved)."
|
||||
return 0
|
||||
fi
|
||||
|
||||
need launchctl
|
||||
local dom="gui/$(id -u)"
|
||||
# ---- snapshot-first ----
|
||||
log "snapshot-first backup -> $bkp"
|
||||
mkdir -p "$bkp"
|
||||
cp -p "$LIVE_DATA_DIR/neuron.egm" "$bkp/neuron.egm.bak"
|
||||
cp -p "$LIVE_DATA_DIR/neuron.wal" "$bkp/neuron.wal.bak" 2>/dev/null || true
|
||||
cp -p "$LIVE_PLIST" "$bkp/plist.bak"
|
||||
printf 'rollback REAL_BIN=%s\nNEWBIN=%s\ndata_promote=%s\n' "$live_bin" "$new_bin" "$do_data" > "$bkp/rollback.txt"
|
||||
ok "backup complete"
|
||||
|
||||
# ---- additive binary install + plist supersede ----
|
||||
cp -p "$d/bin/engram" "$new_bin"
|
||||
python3 - "$LIVE_PLIST" "$new_bin" <<'PY'
|
||||
import sys,plistlib
|
||||
p,new=sys.argv[1],sys.argv[2]
|
||||
d=plistlib.load(open(p,'rb')); d.setdefault("EnvironmentVariables",{})["ENGRAM_REAL_BIN"]=new
|
||||
plistlib.dump(d,open(p,'wb'))
|
||||
PY
|
||||
ok "installed $new_bin + updated plist ENGRAM_REAL_BIN"
|
||||
|
||||
# ---- optional data promote (after backup) ----
|
||||
if [ $do_data -eq 1 ]; then
|
||||
log "data promote: clone store -> live (backed up above)"
|
||||
cp -p "$d/data/neuron.egm" "$LIVE_DATA_DIR/neuron.egm"
|
||||
cp -p "$d/data/neuron.wal" "$LIVE_DATA_DIR/neuron.wal" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
# ---- rails cutover: bootout -> settle-poll -> bootstrap ----
|
||||
log "rails: launchctl bootout $dom/$LIVE_LABEL"
|
||||
launchctl bootout "$dom/$LIVE_LABEL" 2>/dev/null || true
|
||||
local i
|
||||
for i in $(seq 1 60); do
|
||||
launchctl print "$dom/$LIVE_LABEL" >/dev/null 2>&1 || { ok "settle: job gone after ${i}x0.5s"; break; }
|
||||
printf ' settle: job still present (%d)\n' "$i" >&2; sleep 0.5
|
||||
done
|
||||
log "rails: launchctl bootstrap $dom <plist>"
|
||||
launchctl bootstrap "$dom" "$LIVE_PLIST" || warn "bootstrap returned nonzero"
|
||||
|
||||
# ---- verify ----
|
||||
log "verify prod health"
|
||||
local s="" ; for i in $(seq 1 60); do s="$(live_stats)"; [ -n "$s" ] && break; sleep 1; done
|
||||
local ok_verify=1 le; le="$(stat_field "$s" edge_count)"
|
||||
local base_e; base_e="$(mget "$name" "['live_baseline']['edge_count']")"
|
||||
[ -n "$s" ] || ok_verify=0
|
||||
[ "${le:-0}" -ge "${base_e:-0}" ] || ok_verify=0
|
||||
local kid; for kid in "${KEYSTONES[@]}"; do curl -s -m5 "$LIVE_URL/api/node/$kid" 2>/dev/null | grep -q "\"$kid\"" || ok_verify=0; done
|
||||
if [ "$ok_verify" -eq 1 ]; then
|
||||
ok "PROMOTED. live stats: $s (rollback: $bkp)"; return 0
|
||||
fi
|
||||
|
||||
# ---- auto-rollback ----
|
||||
warn "verify FAILED — auto-rollback"
|
||||
cp -p "$bkp/plist.bak" "$LIVE_PLIST"
|
||||
[ $do_data -eq 1 ] && { cp -p "$bkp/neuron.egm.bak" "$LIVE_DATA_DIR/neuron.egm"; cp -p "$bkp/neuron.wal.bak" "$LIVE_DATA_DIR/neuron.wal" 2>/dev/null || true; }
|
||||
launchctl bootout "$dom/$LIVE_LABEL" 2>/dev/null || true
|
||||
for i in $(seq 1 60); do launchctl print "$dom/$LIVE_LABEL" >/dev/null 2>&1 || break; sleep 0.5; done
|
||||
launchctl bootstrap "$dom" "$LIVE_PLIST" || true
|
||||
die "ROLLED BACK to $live_bin. See $bkp"
|
||||
}
|
||||
|
||||
# ================================================================ destroy ======
|
||||
cmd_destroy(){
|
||||
local name="$1"; shift || true
|
||||
mexists "$name" || die "no such sandbox: $name"
|
||||
local d; d="$(sdir "$name")"
|
||||
stop_daemon "$name"
|
||||
if [ -d "$d/build/worktree" ]; then
|
||||
log "removing git worktree"
|
||||
git -C "$EL_REPO" worktree remove --force "$d/build/worktree" 2>/dev/null || true
|
||||
git -C "$EL_REPO" worktree prune 2>/dev/null || true
|
||||
fi
|
||||
log "removing $d"
|
||||
rm -rf "$d"
|
||||
ok "destroyed '$name' (live untouched)"
|
||||
}
|
||||
|
||||
# ================================================================ list/status ==
|
||||
cmd_list(){
|
||||
[ -d "$SBX_ROOT" ] || { echo "no sandboxes"; return 0; }
|
||||
printf '%-16s %-6s %-8s %-9s %s\n' NAME PORT STATE PID SOURCE
|
||||
local m
|
||||
for m in "$SBX_ROOT"/*/manifest.json; do
|
||||
[ -f "$m" ] || continue
|
||||
local n p src pid state
|
||||
n="$(python3 -c "import json;print(json.load(open('$m'))['name'])")"
|
||||
p="$(python3 -c "import json;print(json.load(open('$m'))['port'])")"
|
||||
src="$(python3 -c "import json;print(json.load(open('$m'))['source'])")"
|
||||
pid="$(daemon_pid "$n")"; state="stopped"; daemon_alive "$n" && state="running"
|
||||
printf '%-16s %-6s %-8s %-9s %s\n' "$n" "$p" "$state" "${pid:-–}" "$src"
|
||||
done
|
||||
}
|
||||
cmd_status(){
|
||||
local name="$1"; mexists "$name" || die "no such sandbox: $name"
|
||||
python3 -m json.tool "$(manifest "$name")"
|
||||
daemon_alive "$name" && echo "state: running (pid $(daemon_pid "$name")) stats: $(sbx_stats "$name")" || echo "state: stopped"
|
||||
[ -f "$(sdir "$name")/validate.json" ] && { echo "--- last validation ---"; python3 -m json.tool "$(sdir "$name")/validate.json"; }
|
||||
}
|
||||
|
||||
usage(){ cat >&2 <<EOF
|
||||
${C_BLD}nsbx${C_0} — Neuron Sandbox: experiments + code changes against the REAL engram
|
||||
runtime on an isolated snapshot of the live mind, with a gated promote-to-prod path.
|
||||
|
||||
nsbx up [name] [flags…] one command: your private, isolated copy of the mind
|
||||
(creates on first run, starts thereafter; prod untouchable)
|
||||
nsbx create [name] [--port N] [--source DIR | --branch REF [--repo R] | --binary PATH]
|
||||
clone live store+WAL+config, place/build the runtime, boot on an
|
||||
isolated port (never :$LIVE_BIND_PORT/:$SOUL_PORT). Default runtime = stock prod binary.
|
||||
nsbx build <name> --source DIR | --branch REF rebuild the runtime from a code change + hot-restart
|
||||
nsbx run <name> <cmd...> | api <path> [json] run an experiment; capture output + metrics
|
||||
nsbx validate <name> rails as checks: zero-loss(load+reboot), reboot-prove,
|
||||
RSS bound, retrieval parity, keystone integrity
|
||||
nsbx promote <name> [--data] [--i-approve-prod-cutover] GATED rails cutover to prod (DRY-RUN without approval)
|
||||
nsbx destroy <name> stop daemon, free port, remove clone (live untouched)
|
||||
nsbx list | nsbx status <name>
|
||||
|
||||
Env in 'run' cmds: \$SBX_URL \$SBX_PORT \$SBX_KEY \$SBX_DATA \$SBX_BIN \$SBX_NAME
|
||||
EOF
|
||||
}
|
||||
|
||||
main(){
|
||||
local cmd="${1:-}"; shift || true
|
||||
case "$cmd" in
|
||||
up) cmd_up "$@";;
|
||||
create) cmd_create "$@";;
|
||||
build) cmd_build "$@";;
|
||||
run) cmd_run "$@";;
|
||||
validate) cmd_validate "$@";;
|
||||
promote) cmd_promote "$@";;
|
||||
destroy) cmd_destroy "$@";;
|
||||
list|ls) cmd_list "$@";;
|
||||
status) cmd_status "$@";;
|
||||
""|-h|--help|help) usage;;
|
||||
*) die "unknown command: $cmd (try: nsbx help)";;
|
||||
esac
|
||||
}
|
||||
main "$@"
|
||||
Reference in New Issue
Block a user