Files
el/engram/test/run_http_sigpipe_test.sh
T
Neuron dab081c8f8
El SDK CI - dev / build-and-test (pull_request) Failing after 3m47s
runtime: a client hanging up must not kill the server
send() to a socket whose peer has closed raises SIGPIPE, whose default
disposition is terminate, and nothing in this runtime suppressed it. grep for
SIGPIPE|MSG_NOSIGNAL|SO_NOSIGPIPE|sigaction across lang/runtime/ and engram/src/
returned zero hits. So any abandoned request could kill the process, and one
did, every ten minutes, for three days.

Measured in production, not inferred:
  - 254 restarts between 2026-08-13T19:37Z and 2026-08-16T18:16Z
  - `launchctl list ai.neuron.engram` -> LastExitStatus = 13
    (raw wait status 13 = killed by signal 13 = SIGPIPE)
  - trigger: ai.neuron.engram-tick.plist, StartInterval 600 — matching the
    restart cadence exactly — running `curl -s -m10 -X POST /api/tick`. When
    /api/tick exceeded curl's 10s timeout the client hung up, and the eventual
    response write killed the server.
  - the tick log recorded an EMPTY response in 279 of 448 ticks.
  - launchd KeepAlive:true restarted it each time, so the loop was invisible
    except as a PID that kept changing.

Nothing was ever logged about it and nothing could have been: a signal-killed
process never reaches a line where it could write one. engram.log is 2.7 MB of
nothing but repeated "[http] listening on [::]:8742". That is the signature of
this bug, not a gap in logging.

Note what was already correct: http_send_all's `w <= 0` branch handles a dead
peer properly. It had never once executed, because the signal killed the process
before send() could return -1/EPIPE. The error handling was written and
unreachable.

The suppression is per-socket/per-call (SO_NOSIGPIPE on macOS/BSD, MSG_NOSIGNAL
on Linux) rather than a global signal(SIGPIPE, SIG_IGN). Both fix the crash;
only this one has zero blast radius. A global handler would also change the
disposition of writes to ordinary pipes in the exec/fs paths, which this fix has
no business touching. Every HTTP response funnels through http_send_all, so one
function covers all three accept loops.

Negative control (invariant 8.6). Same data, same endpoint, same probe:
  unpatched -> SERVER DIED 43.3s after the hang-up   (exit 1)
  patched   -> survived 2 rounds, still listening, still serving (exit 0)

The probe is C, not Python, and it is worth saying why the obvious version of it
is useless: closing with SO_LINGER=0 emits RST, the server's FIRST write returns
ECONNRESET rather than raising SIGPIPE, http_send_all stops cleanly, and the test
PASSES ON THE UNPATCHED BUILD. It has to be a normal close (FIN) — first write
succeeds, peer answers RST, second write raises the signal. And liveness must be
checked ~45s later, not immediately: the kill lands when the server reaches its
write, not when the client leaves. My first probe got both wrong and reported a
false pass.

Not fixed here, and worth separate attention: /api/activate takes ~16s and
/api/tick exceeds 10s at all, and the heartbeat's -m10 is shorter than the work
it asks for.
2026-08-16 13:37:49 -05:00

33 lines
1.5 KiB
Bash
Executable File

#!/bin/sh
# Build + RUN the SIGPIPE integration probe: a client that hangs up mid-response
# must not kill the server. Pure C11, no Python, no third-party anything.
#
# This is an INTEGRATION probe — it needs a running engram, and it needs the PID
# so it can tell "still serving" from "restarted by a supervisor underneath me".
# Point it at a SCRATCH instance, never at production:
#
# cp -Rc ~/.neuron/engram /tmp/engram-scratch # APFS clone, instant
# EL_SINGLETON_DIR=/tmp ENGRAM_DATA_DIR=/tmp/engram-scratch \
# ENGRAM_STORE=1 ENGRAM_API_KEY=ntn-user-2026 ENGRAM_BIND=":18753" ./engram &
# lsof -nP -iTCP:18753 -sTCP:LISTEN # confirm YOUR pid owns the port
# ./engram/test/run_http_sigpipe_test.sh 18753 <pid>
#
# NEGATIVE CONTROL (invariant §8.6 — no test without one). This probe was shown
# to FAIL on the pre-change build before the fix was accepted. Measured, same
# data, same endpoint, same probe:
# unpatched -> SERVER DIED 43.3s after the hang-up (exit 1)
# patched -> survived 2 rounds, still listening, still serving (exit 0)
# To reproduce the failing side, build the runtime at the parent commit and run
# this against it.
set -e
HERE=$(cd "$(dirname "$0")" && pwd)
CC=${CC:-cc}
PORT=${1:?usage: run_http_sigpipe_test.sh <port> <pid> [rounds] [settle_seconds]}
PID=${2:?usage: run_http_sigpipe_test.sh <port> <pid> [rounds] [settle_seconds]}
ROUNDS=${3:-2}
SETTLE=${4:-45}
TMP=$(mktemp -d)
$CC -std=c11 -Wall -Wextra -O2 "$HERE/test_http_sigpipe.c" -o "$TMP/probe"
"$TMP/probe" "$PORT" "$PID" "$ROUNDS" "$SETTLE"