nsbx: fail loud on daemon-not-ready instead of printing a false success banner

Three confirmed-live bugs tonight:

- `nsbx up` printed "daemon did not become ready" immediately followed by a
  green "your sandbox is ready" banner and exited 0, because the existing-
  sandbox restart path (`daemon_alive || start_daemon`) never checked
  start_daemon's return code. `cmd_build` had the identical unguarded
  pattern, plus `cmd_run`/`cmd_validate`'s own start-if-dead calls. All four
  now `|| die` with a message pointing at daemon.log.

- `nsbx status`/`nsbx list` reported bare "state: running" for a process
  that's alive (passes kill -0) but not actually answering /api/stats --
  pegged, hung, or mid-boot. Added daemon_health(), which does the real
  stats fetch and distinguishes stopped/running/unresponsive; both commands
  now say "running but NOT RESPONDING" with a next-step hint instead of
  silently going quiet on the stats field. Reproduced live against another
  agent's actively-running (CPU-pinned, non-responsive) sandbox tonight, and
  again via a deliberate SIGSTOP on a throwaway sandbox.

- Sandboxes carried no visible signal that their binary predated a relevant
  fix. `status`/`list` now show the binary's sha + real build timestamp
  (mtime survives `cp -p`), plus a best-effort staleness note: for
  stock-prod clones, compare against the currently-configured live binary;
  for source/branch builds, compare the recorded source commit against
  local origin/dev via merge-base --is-ancestor.

Also, found live while verifying the above:

- A cold boot under concurrent sandbox/CPU load can legitimately take past
  the old hardcoded 15s readiness window. Made it configurable
  (NSBX_READY_TIMEOUT_SECS) rather than just widening the default blindly.

- cmd_create's post-boot baseline capture could silently record sbx_baseline
  as 0/0 when the stats fetch came back empty right after the auto-remerge
  step -- which would make every future `nsbx validate` zero-loss/reboot-
  prove check trivially PASS regardless of real data loss. Added a bounded
  retry and a loud warning if it still comes back empty.

- Sharpened a handful of "no such sandbox" / missing-binary errors to name
  the next command instead of just stating the failure.
This commit is contained in:
bigmerge
2026-08-15 17:32:44 -05:00
parent 2555e363a6
commit 979e820f68
2 changed files with 137 additions and 29 deletions
+3
View File
@@ -173,4 +173,7 @@ Each ad-hoc harness becomes `nsbx run <name> …` (or `--source` build) against
## Env knobs
`NSBX_ROOT`, `NSBX_PORT_BASE`, `NSBX_RSS_BOUND_MB`, `NSBX_REMERGE_THRESHOLD`,
`NSBX_READY_TIMEOUT_SECS` (default 15 — how long `up`/`create`/`build` wait for a
daemon to answer `/api/stats` before reporting failure; raise it if a boot is
legitimately slow under concurrent sandbox/CPU load rather than actually broken),
`EL_REPO` (for `elc` + runtime sources), `ENGRAM_LIVE_DATA_DIR`, `ENGRAM_LIVE_PLIST`.