Merge pull request 'singleton: guard the state, not the program's name' (#157) from fix/singleton-guards-the-state into dev
El SDK CI - dev / build-and-test (push) Failing after 11m15s

This commit was merged in pull request #157.
This commit is contained in:
2026-08-17 00:57:49 +00:00
8 changed files with 213 additions and 49 deletions
+36 -7
View File
@@ -1143,6 +1143,7 @@ The `program` block is where a concern of this shape is declared once and enforc
```
program "engram" {
singleton: "engram"
guards: engram_resolve_data_dir()
env ENGRAM_BIND: String = ":8742"
env GUIDE_PORT: Int = "8771"
env ENGRAM_API_KEY: String required
@@ -1155,24 +1156,50 @@ Grammar:
```ebnf
program_block = "program" string "{" { program_field } "}" ;
program_field = singleton_field | env_field ;
program_field = singleton_field | guards_field | env_field ;
singleton_field = "singleton" ":" string [ "," ] ;
guards_field = "guards" ":" expr [ "," ] ;
env_field = "env" ident ":" type
[ "=" string ] [ "required" ] [ "," ] ;
```
`singleton` and `env` are **not** reserved words. They are read as identifier token values by the block's own parse loop, so they remain usable as ordinary identifiers everywhere else. `program` is the only keyword this section adds.
`singleton`, `guards` and `env` are **not** reserved words. They are read as identifier token values by the block's own parse loop, so they remain usable as ordinary identifiers everywhere else. `program` is the only keyword this section adds.
### 18.2 Process identity — `singleton`
### 18.2 Process identity — `singleton` and `guards`
`singleton: "id"` compiles to an `el_singleton_acquire("id")` call injected as the **first statement of `main()`**, before any user statement runs.
`singleton: "id"` with `guards: <expr>` compiles to `el_singleton_acquire("id", <expr>)`, injected as the **first statement of `main()`**, before any user statement runs. `<expr>` evaluates to the path of the **state** the singleton protects.
The runtime takes an exclusive non-blocking `flock` on `<dir>/el-singleton-<id>.lock`, where `<dir>` is `$EL_SINGLETON_DIR`, else `$TMPDIR`, else `/tmp`. On success it writes its pid and holds the descriptor open for the life of the process. On contention it **refuses to start**: it reports the holder's pid, names the lock file, and exits 1.
**`guards:` is mandatory.** A `singleton:` without one is a compile error. This is not defensive strictness; it is the correction of a defect measured in this tree on 2026-08-16, and the rule the rest of this section exists to state:
Two properties are deliberate:
> **Guard the thing, not the name.** A lock that protects state must be keyed on the state.
- **It is a lock, not a pidfile.** The kernel releases an `flock` when the owning process dies — including on `SIGKILL` and on crash. There is therefore no stale-lock state, and so no "delete the lock file to get unstuck" recovery ritual. Such a ritual would itself be a convention, which is the thing this section exists to remove.
Until that date the lock was `<dir>/el-singleton-<id>.lock` where `<dir>` was `$EL_SINGLETON_DIR`, else `$TMPDIR`, else `/tmp`. It was keyed on the program's **name** and on a temp directory, and it never consulted the state it claimed to protect — while its own refusal message read *"Refusing to start a second instance against the same state."* Measured, it failed in **both** directions:
| Situation | Correct answer | Name-keyed lock gave |
|---|---|---|
| same data dir, same `$TMPDIR` | refuse | refuse ✅ |
| same data dir, different `$TMPDIR` | refuse | **started** ❌ — the two-writer data-loss condition, defeated by one environment variable |
| different data dirs, same `$TMPDIR` | both start | **refused**, naming an unrelated pid ❌ |
| same dir spelled differently, different `$TMPDIR` | refuse | **started** ❌ |
Both failure directions are one error: the identity of a resource had been replaced by a label for it. The false negative is the dangerous one — a guard whose bypass is `TMPDIR=/tmp/other` is not a guard.
**The mechanism.** The lock file lives **inside the guarded directory**: `<state>/.el-singleton-<id>.lock`. The runtime takes an exclusive non-blocking `flock` on it, writes its pid, and holds the descriptor open for the life of the process.
That single placement decision is the whole fix, and it is why there is no hashing, no canonical-path registry, and no environment variable left to subvert:
- **Same directory** ⇒ same file ⇒ same inode ⇒ the `flock` contends. `$TMPDIR` is not in the key, so there is nothing to change to get past it. `$EL_SINGLETON_DIR` no longer exists.
- **Different directories** ⇒ different files ⇒ no contention. Two stores are two stores; they were never in conflict, and are no longer treated as if they were.
- **Different spellings of one directory** — trailing slash, `x/../x`, a symlink — resolve to the same inode during the kernel's own path walk, so they contend without this code comparing strings. Path canonicalisation happens only to make the diagnostic name one directory in one spelling; the *decision* never depends on it.
- **An unguardable state** — the directory is missing, or read-only — is a **refusal**, not a fallback. Starting unguarded against the store the guard exists to protect is the failure being removed.
**Why `guards:` is an expression and not a string.** The runtime cannot know, generically, which environment variable holds an arbitrary program's state; and a program whose state path already has an owner must not restate it. The engram's data dir is resolved by `engram_resolve_data_dir()`, which owns both the `$ENGRAM_DATA_DIR` read and the `$HOME/.neuron/engram` fallback (§18.4). Writing `guards: engram_resolve_data_dir()` points the guard at that owner. A `guards:` that took a string would force the path's default to be written down twice, and a guard that resolved the path its own way could end up locking a directory the program never writes to — the same two-owners defect §18.4 exists to prevent.
Three properties are deliberate:
- **It is a lock, not a pidfile.** The kernel releases an `flock` when the owning process dies — including on `SIGKILL` and on crash. There is therefore no stale-lock state, and so no "delete the lock file to get unstuck" recovery ritual. Such a ritual would itself be a convention, which is the thing this section exists to remove. (A lock file left behind inside a copied data directory — `cp -Rc` and friends — is inert: it carries no lock, only a stale pid string that the next holder overwrites.)
- **It reports the holder's pid.** "Already running" is not actionable. A pid is. This is the direct answer to the observed failure where a stale process survived a `pkill` and went on answering probes.
- **The message is true.** It names the state it checked and the lock it failed to take, and it says "the same state" only because the lock it contended for is *in* that state. A diagnostic that asserts a check that did not happen is worse than no diagnostic: it is what let the name-keyed version read as correct for as long as it did.
Refusal is loud and total. It is not a warning, and the program does not continue degraded. This matters more than it looks: today a second engram whose `bind()` fails merely *returns* from `http_serve` — after it has already replayed the WAL and written boot-time backup files — and then exits **0**, indistinguishable from a clean run. `singleton` refuses before the first side effect.
@@ -1194,6 +1221,8 @@ Some values look like configuration and are not. `ENGRAM_DATA_DIR` already has a
The rule: **a variable belongs in the program block when the block would be its only owner.** If a resolver already owns it, leave it there.
This is also why `guards:` (§18.2) takes an expression: it lets the block *reference* the existing owner — `guards: engram_resolve_data_dir()` — rather than become a second one.
`HOME` is likewise not configuration. It is an environment fact, and stays a raw `env()` read.
---