dab081c8f8
El SDK CI - dev / build-and-test (pull_request) Failing after 3m47s
send() to a socket whose peer has closed raises SIGPIPE, whose default
disposition is terminate, and nothing in this runtime suppressed it. grep for
SIGPIPE|MSG_NOSIGNAL|SO_NOSIGPIPE|sigaction across lang/runtime/ and engram/src/
returned zero hits. So any abandoned request could kill the process, and one
did, every ten minutes, for three days.
Measured in production, not inferred:
- 254 restarts between 2026-08-13T19:37Z and 2026-08-16T18:16Z
- `launchctl list ai.neuron.engram` -> LastExitStatus = 13
(raw wait status 13 = killed by signal 13 = SIGPIPE)
- trigger: ai.neuron.engram-tick.plist, StartInterval 600 — matching the
restart cadence exactly — running `curl -s -m10 -X POST /api/tick`. When
/api/tick exceeded curl's 10s timeout the client hung up, and the eventual
response write killed the server.
- the tick log recorded an EMPTY response in 279 of 448 ticks.
- launchd KeepAlive:true restarted it each time, so the loop was invisible
except as a PID that kept changing.
Nothing was ever logged about it and nothing could have been: a signal-killed
process never reaches a line where it could write one. engram.log is 2.7 MB of
nothing but repeated "[http] listening on [::]:8742". That is the signature of
this bug, not a gap in logging.
Note what was already correct: http_send_all's `w <= 0` branch handles a dead
peer properly. It had never once executed, because the signal killed the process
before send() could return -1/EPIPE. The error handling was written and
unreachable.
The suppression is per-socket/per-call (SO_NOSIGPIPE on macOS/BSD, MSG_NOSIGNAL
on Linux) rather than a global signal(SIGPIPE, SIG_IGN). Both fix the crash;
only this one has zero blast radius. A global handler would also change the
disposition of writes to ordinary pipes in the exec/fs paths, which this fix has
no business touching. Every HTTP response funnels through http_send_all, so one
function covers all three accept loops.
Negative control (invariant 8.6). Same data, same endpoint, same probe:
unpatched -> SERVER DIED 43.3s after the hang-up (exit 1)
patched -> survived 2 rounds, still listening, still serving (exit 0)
The probe is C, not Python, and it is worth saying why the obvious version of it
is useless: closing with SO_LINGER=0 emits RST, the server's FIRST write returns
ECONNRESET rather than raising SIGPIPE, http_send_all stops cleanly, and the test
PASSES ON THE UNPATCHED BUILD. It has to be a normal close (FIN) — first write
succeeds, peer answers RST, second write raises the signal. And liveness must be
checked ~45s later, not immediately: the kill lands when the server reaches its
write, not when the client leaves. My first probe got both wrong and reported a
false pass.
Not fixed here, and worth separate attention: /api/activate takes ~16s and
/api/tick exceeds 10s at all, and the heartbeat's -m10 is shorter than the work
it asks for.
33 lines
1.5 KiB
Bash
Executable File
33 lines
1.5 KiB
Bash
Executable File
#!/bin/sh
|
|
# Build + RUN the SIGPIPE integration probe: a client that hangs up mid-response
|
|
# must not kill the server. Pure C11, no Python, no third-party anything.
|
|
#
|
|
# This is an INTEGRATION probe — it needs a running engram, and it needs the PID
|
|
# so it can tell "still serving" from "restarted by a supervisor underneath me".
|
|
# Point it at a SCRATCH instance, never at production:
|
|
#
|
|
# cp -Rc ~/.neuron/engram /tmp/engram-scratch # APFS clone, instant
|
|
# EL_SINGLETON_DIR=/tmp ENGRAM_DATA_DIR=/tmp/engram-scratch \
|
|
# ENGRAM_STORE=1 ENGRAM_API_KEY=ntn-user-2026 ENGRAM_BIND=":18753" ./engram &
|
|
# lsof -nP -iTCP:18753 -sTCP:LISTEN # confirm YOUR pid owns the port
|
|
# ./engram/test/run_http_sigpipe_test.sh 18753 <pid>
|
|
#
|
|
# NEGATIVE CONTROL (invariant §8.6 — no test without one). This probe was shown
|
|
# to FAIL on the pre-change build before the fix was accepted. Measured, same
|
|
# data, same endpoint, same probe:
|
|
# unpatched -> SERVER DIED 43.3s after the hang-up (exit 1)
|
|
# patched -> survived 2 rounds, still listening, still serving (exit 0)
|
|
# To reproduce the failing side, build the runtime at the parent commit and run
|
|
# this against it.
|
|
set -e
|
|
HERE=$(cd "$(dirname "$0")" && pwd)
|
|
CC=${CC:-cc}
|
|
PORT=${1:?usage: run_http_sigpipe_test.sh <port> <pid> [rounds] [settle_seconds]}
|
|
PID=${2:?usage: run_http_sigpipe_test.sh <port> <pid> [rounds] [settle_seconds]}
|
|
ROUNDS=${3:-2}
|
|
SETTLE=${4:-45}
|
|
TMP=$(mktemp -d)
|
|
|
|
$CC -std=c11 -Wall -Wextra -O2 "$HERE/test_http_sigpipe.c" -o "$TMP/probe"
|
|
"$TMP/probe" "$PORT" "$PID" "$ROUNDS" "$SETTLE"
|