nsbx: fail loud on daemon-not-ready instead of printing a false success banner
Three confirmed-live bugs tonight: - `nsbx up` printed "daemon did not become ready" immediately followed by a green "your sandbox is ready" banner and exited 0, because the existing- sandbox restart path (`daemon_alive || start_daemon`) never checked start_daemon's return code. `cmd_build` had the identical unguarded pattern, plus `cmd_run`/`cmd_validate`'s own start-if-dead calls. All four now `|| die` with a message pointing at daemon.log. - `nsbx status`/`nsbx list` reported bare "state: running" for a process that's alive (passes kill -0) but not actually answering /api/stats -- pegged, hung, or mid-boot. Added daemon_health(), which does the real stats fetch and distinguishes stopped/running/unresponsive; both commands now say "running but NOT RESPONDING" with a next-step hint instead of silently going quiet on the stats field. Reproduced live against another agent's actively-running (CPU-pinned, non-responsive) sandbox tonight, and again via a deliberate SIGSTOP on a throwaway sandbox. - Sandboxes carried no visible signal that their binary predated a relevant fix. `status`/`list` now show the binary's sha + real build timestamp (mtime survives `cp -p`), plus a best-effort staleness note: for stock-prod clones, compare against the currently-configured live binary; for source/branch builds, compare the recorded source commit against local origin/dev via merge-base --is-ancestor. Also, found live while verifying the above: - A cold boot under concurrent sandbox/CPU load can legitimately take past the old hardcoded 15s readiness window. Made it configurable (NSBX_READY_TIMEOUT_SECS) rather than just widening the default blindly. - cmd_create's post-boot baseline capture could silently record sbx_baseline as 0/0 when the stats fetch came back empty right after the auto-remerge step -- which would make every future `nsbx validate` zero-loss/reboot- prove check trivially PASS regardless of real data loss. Added a bounded retry and a loud warning if it still comes back empty. - Sharpened a handful of "no such sandbox" / missing-binary errors to name the next command instead of just stating the failure.
This commit is contained in:
@@ -173,4 +173,7 @@ Each ad-hoc harness becomes `nsbx run <name> …` (or `--source` build) against
|
||||
## Env knobs
|
||||
|
||||
`NSBX_ROOT`, `NSBX_PORT_BASE`, `NSBX_RSS_BOUND_MB`, `NSBX_REMERGE_THRESHOLD`,
|
||||
`NSBX_READY_TIMEOUT_SECS` (default 15 — how long `up`/`create`/`build` wait for a
|
||||
daemon to answer `/api/stats` before reporting failure; raise it if a boot is
|
||||
legitimately slow under concurrent sandbox/CPU load rather than actually broken),
|
||||
`EL_REPO` (for `elc` + runtime sources), `ENGRAM_LIVE_DATA_DIR`, `ENGRAM_LIVE_PLIST`.
|
||||
|
||||
Reference in New Issue
Block a user