feat(engine): tools + agentic loop on the OpenAI wire, and two chat-breaking fixes found proving it
Teaches the OpenAI-format lane (Groq/OpenAI/Grok/Gemini/Ollama) to offer tools,
execute them, and loop — the capability that until now existed only on the
Anthropic wire. The tool-execution, consent, bridge and run-progress machinery is
reused unchanged; only the wire dialect is new.
Two pre-existing defects were found while proving it, and are fixed here because
both silently break chat:
1. PROVIDER WIRING NEVER CONNECTED. The launcher exports SOUL_LLM_PROVIDER /
SOUL_LLM_BASE_URL and puts the provider key in ANTHROPIC_API_KEY + SOUL_API_KEY;
the engine's provider fork read only NEURON_LLM_0_*, which nothing sets in a
customer build. So use_openai was ALWAYS false: every non-Anthropic user's turns
went to api.anthropic.com carrying, say, a Groq key, and came back
"llm unavailable". Proven side-by-side against the pinned round-9 brain
(sha256 15cf7d1b…): identical env, shipped brain = "llm unavailable" both chat
modes with ZERO calls to the configured endpoint; this build = a real answer,
with the probe logging POST /v1/chat/completions and Bearer <provider key>.
Fixed brain-side only (env fallbacks) — no app or launcher change needed.
2. TRUNCATION SPLITS UTF-8 CHARACTERS. The session preload cuts recalled memory at
fixed BYTE lengths (continuity snippet 350; session_preload_bullets per bullet).
A cut landing inside a multi-byte character leaves a dangling lead byte in the
SYSTEM PROMPT, making the whole request body invalid UTF-8 — providers reject it
and the user sees an unexplained failure. Captured from a real body: 18,710 bytes,
decode fails at 18,248 on 'e2', a box-drawing rule (U+2500 = E2 94 80) sliced in
half. Trigger is ordinary content — em dash, curly quote, accented name, emoji,
table border — and it gets MORE likely as memory grows. Shared code: this hit the
Anthropic wire too. Fixed with utf8_safe_slice() applied at BOTH cut sites.
WHAT IS IN THE PORT
- llm_base_url / llm_wire_format / agentic_api_key: fall back to the launcher's own
SOUL_LLM_* names; anthropic deliberately still returns "" so its native path is
untouched (endpoint configurability remains neuron#62).
- openai_tools_json(): Anthropic tool schema -> OpenAI function schema; entries with
no input_schema (Anthropic's server-side web_search) are skipped — they cannot
execute on this wire.
- agentic_tools_no_web(): the standard set minus that server tool.
- openai_agentic_loop(): forked rather than parameterised, so agentic_loop — which
carries every round-7/8/9 fix — is provably untouched. Same envelopes, same state
keys, same consent policy (ask_all / escalate / builtin / always-allow), same
client-bridge contract, same run-progress ledger, same 12-iteration cap.
- ADR-0005 mirrored on this wire: parallel_tool_calls:false is sent explicitly, and
if a provider ignores it we honour the FIRST call and echo only that one, so the
conversation we send is never self-contradictory. The drop is logged loudly.
- The assistant turn echoes the provider's own content bytes (json_get_raw), so a
JSON null stays null and nothing is lost to a decode/re-encode round trip.
- Tool results are embedded already-escaped (dispatch_tool json_safe's them);
truncation trims a dangling escape so a cut can't invalidate the body.
- bridge_save() gains a "wire" scalar and agentic_resume branches on it, so a
suspended turn resumes on the wire it suspended on. Legacy blobs (no field) resume
as anthropic. The field is read from the blob's SCALAR HEAD only — an unbounded
first-match scan would run on into messages_raw, which is model-controlled, and
that is exactly the round-9 resume defect. Pinned by a test.
- Three fork sites: handle_chat_agentic, handle_dharma_room_turn_agentic,
agentic_resume. Tool assembly is computed once per lane at both entry points
(it makes an HTTP call to the connector bridge; it was being paid for twice).
TOOLING THAT DID NOT EXIST
- tests/run-el-test.sh — engine tests were never runnable: elc is a compiler, it
emits C and exits. This emits the test to C, compiles soul.c with main renamed
away, links the rest + the repo-pinned runtime, and runs it. It also COMPUTES THE
VERDICT, because every counted test file's "N passed, M failed" summary is a
permanent 0/0 — the counters increment inside if BLOCKS, which El scoping
discards (9 files; real fix filed as neuron#116). Proven to discriminate with a
deliberately-broken assertion.
- tests/gate-openai/ — deterministic OpenAI-dialect provider stub + scenarios +
driver + hostile modes, and a strict request validator that rejects any
Anthropic-shaped field so dialect leakage fails loudly.
VERIFICATION (rungs named)
- E2E-VERIFIED against a LIVE provider (Anthropic's OpenAI-compatible endpoint,
confirmed live): real answer; a tool call whose out-of-root path was DENIED by the
guard, after which the model refused to claim success ("I won't tell you I did it,
because I didn't"); then a valid path -> file physically on disk with exact content,
honest reply, ledger with per-round entries + {done:true}.
- Deterministic lane gate: 11/12 in both consent configurations (bridge + local);
hostile providers produce no hang and no fabricated answer; the 12-iteration cap
trips with its honest message. The one FAIL is oa-tools-off and is NOT this port —
see "Known, not fixed here".
- ANTHROPIC LANE UNCHANGED: gate9 32/32 on this build and on the pinned round-9
brain; request bytes differ only within the noise band that two runs of the
UNMODIFIED brain also produce (proven with a baseline-vs-baseline control), and
the preload sections — the shared code touched here — are byte-identical.
The rig discriminates: the round-8 brain scores 24/32 on it.
- verify-soul-contract.sh: PASS (27/27 routes, no hard-deletes).
- Unit: test_bridge_serialization 36/36 (incl. 8 new wire/field-order assertions),
test_utf8_slice 18/18, test_agentic_tools 18 PASS / 0 FAIL / 3 documented skips.
KNOWN, NOT FIXED HERE (deliberate)
- Tools:Off on an OpenAI provider still fails: the non-agentic path goes through the
el-runtime provider chain, which appends /v1/chat/completions to a base URL that
already ends in /v1 -> /v1/v1/... 404. Runtime/plain-chat territory, untouched
mid-beta. Note openai_chat_complete() has zero callers — that lane is served
entirely by the runtime chain.
- The 12-iteration cap does not bound a chain of BRIDGED tools (iteration is
per-invocation and resume starts fresh). Parity with the Anthropic lane.
- run_progress resets on each resume, so a client rendering cumulative steps across a
consent pause sees earlier legs vanish. Parity with the Anthropic lane.
- verify-soul-contract.sh needs bash >= 4; under macOS's stock bash 3.2 it dies
instantly with a FALSE red ("local: -n: invalid option").
- Groq-specific live E2E not run: no Groq key exists on this machine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,90 @@
|
||||
# gate-openai — deterministic OpenAI-dialect provider stub
|
||||
|
||||
Staging home for the **soul-openai-tools-v2** gate scaffolding
|
||||
(`docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md`, test-plan rung 1:
|
||||
"stub first — discriminates before El code exists"). Sibling of gate9's
|
||||
Anthropic stub (`_wt-beta-round9/scripts/gate9/stub-llm.py`): same scenario
|
||||
mechanism, opposite wire dialect. Stdlib Python only, 127.0.0.1 only,
|
||||
refuses ports 7770/7779/17779. Run `./selftest.sh` — exit 0 is green.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Role |
|
||||
|---|---|
|
||||
| `stub-openai.py` | HTTP server: `POST /v1/chat/completions` (OpenAI dialect), scenario-scripted responses, request validation, ground-truth JSONL log, hostile modes via `--mode` |
|
||||
| `scenarios-openai.json` | Scenario contract: scripts + markers + per-class/per-step request assertions |
|
||||
| `selftest.sh` | curl-driven proof of every scenario, every rejection, all hostile modes (58 checks) |
|
||||
|
||||
## What each scenario proves (when the brain drives it)
|
||||
|
||||
| Class | Proves |
|
||||
|---|---|
|
||||
| `oa-plain` | finish_reason `stop` ends the loop; tools + `tool_choice` + `parallel_tool_calls:false` were offered on the wire |
|
||||
| `oa-tools-off` | the chat-only lane sends NO tools (offering them there is a 400) |
|
||||
| `oa-single-tool` | full round-trip: `tool_calls` parsed, assistant echo + `role:"tool"` turn with matching `tool_call_id` sent back, final text reached |
|
||||
| `oa-torture` | `function.arguments` (JSON-encoded string with nested quotes, backslashes, newlines, tabs, unicode) survives exactly ONE decode — the stub recomputes the issued payload from the script and 400s on any drift (`gate_echo_mismatch`, the spec §6 two-escaper trap) |
|
||||
| `oa-parallel` | two `tool_calls` in one response: the brain either answers both (paired correctly) or rejects cleanly — an unpaired echo is a 400 |
|
||||
| `oa-mission` | multi-round loop continuation; step index = assistant-message count, so resume threads index correctly by construction |
|
||||
| `oa-api-error` | provider errors 400/429/500/503 in the OpenAI error envelope surface honestly, no retry storm |
|
||||
|
||||
Universal (every request, any scenario): Anthropic dialect leakage fails
|
||||
loudly with 400 — `anthropic-version` header, top-level `system` /
|
||||
`stop_sequences` / `max_tokens_to_sample`, `input_schema` inside tools,
|
||||
Anthropic content blocks (`tool_use`/`tool_result`/...). Tools must be
|
||||
`{type:"function", function:{name, description, parameters}}`, unique names;
|
||||
echoed `arguments` must be a JSON-encoded STRING, never a decoded object.
|
||||
|
||||
## Hostile modes (`--mode`, same file)
|
||||
|
||||
| Mode | Behavior | Brain invariant under test |
|
||||
|---|---|---|
|
||||
| `black-hole` | reads the request, never responds | HTTP timeout exists and surfaces; no silent hang |
|
||||
| `mid-body-drop` | 200 headers, half a JSON body, socket abort | truncated body = clean error, never a half-parsed reply shown as real |
|
||||
| `tool-pending-forever` | every request gets a fresh `tool_calls` response, forever | the loop's iteration cap trips (`max_loop_iterations: 16` in the contract); count actual round-trips via `GET /gate/stats` (`chat_hits`) |
|
||||
|
||||
## How the brain-side gate consumes this
|
||||
|
||||
1. Start: `stub-openai.py --port P --scenarios scenarios-openai.json --log run.jsonl`
|
||||
2. Point the brain at it: `NEURON_LLM_0_URL=http://127.0.0.1:P` +
|
||||
`NEURON_LLM_0_FORMAT=openai` (spec step 0 must verify these actually
|
||||
export at runtime), scratch profile, free soul port.
|
||||
3. Send each phrasing's `prompt` (the marker selects the script); assert the
|
||||
brain's claims (`tools_used`, reply, ledger) against the stub's JSONL log
|
||||
— truth, not narration — plus files on disk for write_file scenarios.
|
||||
4. Any stub 400 = the brain sent a malformed/leaked request; the gate fails
|
||||
with the stub's reason string.
|
||||
5. Re-run gate9's Anthropic matrix unchanged = proof the Anthropic lane is
|
||||
byte-untouched.
|
||||
|
||||
## Reconciliation into gate9 (app repo) — AFTER round 9 merges
|
||||
|
||||
This dir is staging only; the merge is mechanical by design:
|
||||
- `stub-openai.py` + `scenarios-openai.json` move to `scripts/gate9/`
|
||||
alongside `stub-llm.py` + `scenarios.json` (shared conventions: marker
|
||||
matching, assistant-count step indexing, `--port/--scenarios/--log`,
|
||||
JSONL fields `seq/ts/kind/scenario_class/phrasing/step/validation/
|
||||
delivered/http_status`, prod-port refusal, benign background responses,
|
||||
`GATE-SCRIPT-EXHAUSTED` overrun, `{N}/{NN}` repeat expansion).
|
||||
- `prompt-matrix-gate.sh` gains a dialect axis (anthropic|openai) choosing
|
||||
stub + scenario file; `matrix-asserts.py` reads the same log shape.
|
||||
- The `--mode` hostile flags here are PROVIDER-side (brain↔LLM boundary);
|
||||
gate9's `hostile/` servers are SOUL-side (app↔brain boundary). They are
|
||||
complementary, not duplicates — both stay.
|
||||
|
||||
## Open questions for the port author (stub asserts a position; confirm or change)
|
||||
|
||||
1. `parallel_tool_calls` must be **explicitly false** on every tool-bearing
|
||||
request (ADR-0005 pin). If the builder omits it instead, relax
|
||||
`defaults.expect_request.parallel_tool_calls` to `null`.
|
||||
2. `tool_choice` must be present (`"auto"` expected). If the brain relies on
|
||||
the provider default, drop `require_tool_choice`.
|
||||
3. Tool-result `content` is asserted only to be a string; if the brain sends
|
||||
structured JSON-in-string (like `{"ok":true,...}`), no change needed.
|
||||
4. Groq compatibility: Groq's OpenAI-compat endpoint rejects some optional
|
||||
fields; whatever field set the brain settles on for live Groq E2E must be
|
||||
mirrored here so the deterministic gate and the live lane assert the SAME
|
||||
request shape.
|
||||
5. The stub treats a `role:"tool"` turn answering an already-answered id as
|
||||
400; if the resume path can legitimately replay tool results, that rule
|
||||
needs a resume-aware carve-out (gate9's Anthropic stub faced the same
|
||||
issue — see its PASS 1 comment).
|
||||
Executable
+559
@@ -0,0 +1,559 @@
|
||||
#!/usr/bin/env bash
|
||||
# run-lane-gate.sh — brain-side driver for the OpenAI-dialect gate.
|
||||
#
|
||||
# Drives the REAL soul binary against stub-openai.py for every class and every
|
||||
# phrasing in scenarios-openai.json, plus the three hostile provider modes, and
|
||||
# asserts the brain's claims against the stub's ground-truth JSONL (truth, not
|
||||
# narration) and against files on disk.
|
||||
#
|
||||
# SAFETY (hard rules, enforced below):
|
||||
# - never binds 7770 / 7779 / 17779 - only 7891-7894
|
||||
# - never reads or writes ~/.neuron - HOME is redirected to a scratch dir
|
||||
# - every process started here is killed on exit (trap) and proven with lsof
|
||||
#
|
||||
# The soul runs under `script -q /dev/null` so its stdout is a pty: El's
|
||||
# println() uses puts(), which is FULLY buffered to a file, and the process is
|
||||
# killed without flushing — the DRIFT lines would be invisible otherwise.
|
||||
#
|
||||
# Usage: ./run-lane-gate.sh [all|bridge|local|toolsoff|hostile]
|
||||
# bridge = consent round-trip config (no workspace root -> write_file is
|
||||
# "escalate" -> the loop suspends and the CLIENT executes the tool)
|
||||
# local = workspace-root config (write_file is "reversible" + builtin ->
|
||||
# the loop executes the tool in-process and runs to completion)
|
||||
# toolsoff = supplementary: non-agentic lane against a base URL WITHOUT the
|
||||
# /v1 suffix (the el-runtime provider chain appends
|
||||
# /v1/chat/completions itself, unlike chat.el which appends only
|
||||
# /chat/completions)
|
||||
# hostile = black-hole / mid-body-drop / tool-pending-forever
|
||||
#
|
||||
# Env overrides: SOUL_BIN, STUB_PORT, SOUL_PORT, SOUL_PORT_B, RUN_ROOT
|
||||
set -uo pipefail
|
||||
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
PHASES="${1:-all}"
|
||||
|
||||
SOUL_BIN="${SOUL_BIN:-/tmp/soul-oai2/soul-openai-tools}"
|
||||
STUB_PORT="${STUB_PORT:-7891}"
|
||||
SOUL_PORT="${SOUL_PORT:-7892}"
|
||||
SOUL_PORT_B="${SOUL_PORT_B:-7893}"
|
||||
RUN_ROOT="${RUN_ROOT:-/tmp/oa-lane-gate}"
|
||||
STAMP="$(date +%Y%m%d-%H%M%S)"
|
||||
RUN="$RUN_ROOT/$STAMP"
|
||||
|
||||
for p in "$STUB_PORT" "$SOUL_PORT" "$SOUL_PORT_B"; do
|
||||
case "$p" in
|
||||
7770|7779|17779) echo "FATAL: refusing production Neuron port $p"; exit 2;;
|
||||
789[1-4]) ;;
|
||||
*) echo "FATAL: port $p outside the allowed 7891-7894 range"; exit 2;;
|
||||
esac
|
||||
done
|
||||
[ -x "$SOUL_BIN" ] || { echo "FATAL: soul binary not found/executable: $SOUL_BIN"; exit 2; }
|
||||
|
||||
mkdir -p "$RUN/home" "$RUN/ws-bridge" "$RUN/ws-local" "$RUN/ws-off" "$RUN/engram"
|
||||
echo '{"nodes":[],"edges":[]}' > "$RUN/engram/snapshot.json"
|
||||
DRV="$RUN/drv.py"
|
||||
|
||||
STUB_PID=""; SOUL_PID=""
|
||||
cleanup() {
|
||||
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
|
||||
pkill -f "$SOUL_BIN" 2>/dev/null
|
||||
[ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null
|
||||
sleep 0.4
|
||||
[ -n "$SOUL_PID" ] && kill -9 "$SOUL_PID" 2>/dev/null
|
||||
[ -n "$STUB_PID" ] && kill -9 "$STUB_PID" 2>/dev/null
|
||||
return 0
|
||||
}
|
||||
trap cleanup EXIT INT TERM
|
||||
|
||||
start_stub() { # $1 = mode, $2 = log path
|
||||
local mode="$1" log="$2" args=""
|
||||
[ "$mode" = "normal" ] && args="--scenarios $HERE/scenarios-openai.json"
|
||||
# shellcheck disable=SC2086
|
||||
python3 "$HERE/stub-openai.py" --port "$STUB_PORT" --mode "$mode" --log "$log" $args \
|
||||
> "$RUN/stub-$mode.out" 2>&1 &
|
||||
STUB_PID=$!
|
||||
for _ in $(seq 1 50); do
|
||||
curl -sf "http://127.0.0.1:$STUB_PORT/gate/health" >/dev/null 2>&1 && return 0
|
||||
sleep 0.2
|
||||
done
|
||||
echo "FATAL: stub did not come up on $STUB_PORT"; cat "$RUN/stub-$mode.out"; exit 3
|
||||
}
|
||||
stop_stub() { [ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null; sleep 0.3; STUB_PID=""; }
|
||||
|
||||
start_soul() { # $1 = port, $2 = base url, $3 = soul log, $4 = agent root ("" = none)
|
||||
local port="$1" base="$2" log="$3" root="$4"
|
||||
script -q /dev/null \
|
||||
env -u ANTHROPIC_API_KEY -u SOUL_API_KEY -u ENGRAM_URL -u ENGRAM_API_KEY \
|
||||
-u NEURON_API_URL -u NEURON_TOKEN -u SOUL_LLM_PROVIDER -u SOUL_LLM_BASE_URL \
|
||||
-u NEURON_LLM_1_URL -u NEURON_LLM_1_KEY -u SOUL_IDENTITY \
|
||||
HOME="$RUN/home" PATH="$PATH" \
|
||||
NEURON_PORT="$port" EL_HTTP_BIND_HOST=127.0.0.1 \
|
||||
SOUL_ENGRAM_PATH="$RUN/engram/snapshot.json" \
|
||||
SOUL_CGI_ID=ntn-test SOUL_PERSONA_NAME=Neuron \
|
||||
NEURON_LLM_0_URL="$base" NEURON_LLM_0_FORMAT=openai NEURON_LLM_0_KEY=gate-test-key \
|
||||
${root:+NEURON_AGENT_ROOT="$root"} \
|
||||
"$SOUL_BIN" > "$log" 2>&1 &
|
||||
SOUL_PID=$!
|
||||
for _ in $(seq 1 100); do
|
||||
curl -sf "http://127.0.0.1:$port/health" >/dev/null 2>&1 && return 0
|
||||
sleep 0.2
|
||||
done
|
||||
echo "FATAL: soul did not come up on $port"; tail -20 "$log"; exit 3
|
||||
}
|
||||
stop_soul() {
|
||||
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
|
||||
pkill -f "$SOUL_BIN" 2>/dev/null
|
||||
sleep 0.6; SOUL_PID=""
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------ driver ----
|
||||
cat > "$DRV" <<'PYEOF'
|
||||
import json, os, sys, time, threading, urllib.request, urllib.error
|
||||
|
||||
CFG = json.load(open(sys.argv[1]))
|
||||
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
|
||||
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
|
||||
WS = CFG["workspace"]
|
||||
MODE = CFG["mode"] # bridge | local | toolsoff
|
||||
SCEN = json.load(open(CFG["scenarios"]))
|
||||
STUBLOG = CFG["stub_log"]
|
||||
SOULLOG = CFG["soul_log"]
|
||||
ONLY = CFG.get("classes") or list(SCEN["classes"].keys())
|
||||
MAXHOPS = CFG.get("max_hops", 15)
|
||||
OUT = CFG["out"]
|
||||
# the chat-only class must be driven on the NON-agentic door: the agentic door
|
||||
# always advertises tools, which is a 400 on that scenario by contract.
|
||||
NON_AGENTIC = {"oa-tools-off"}
|
||||
|
||||
def http(method, url, obj=None, timeout=300):
|
||||
data = None if obj is None else json.dumps(obj).encode()
|
||||
req = urllib.request.Request(url, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
body = r.read().decode("utf-8", "replace")
|
||||
st = r.status
|
||||
except urllib.error.HTTPError as e:
|
||||
body = e.read().decode("utf-8", "replace"); st = e.code
|
||||
except Exception as e:
|
||||
return -1, "TRANSPORT-ERROR: %r" % (e,), None
|
||||
try:
|
||||
return st, body, json.loads(body)
|
||||
except ValueError:
|
||||
return st, body, None
|
||||
|
||||
def fsize(p):
|
||||
return os.path.getsize(p) if os.path.exists(p) else 0
|
||||
|
||||
def tail_from(path, off):
|
||||
if not os.path.exists(path):
|
||||
return "", off
|
||||
with open(path, "rb") as f:
|
||||
f.seek(off); chunk = f.read(); return chunk.decode("utf-8", "replace"), f.tell()
|
||||
|
||||
def stub_since(off):
|
||||
"""Exact correlation: only the JSONL bytes appended during this phrasing."""
|
||||
txt, noff = tail_from(STUBLOG, off)
|
||||
recs = []
|
||||
for line in txt.splitlines():
|
||||
line = line.strip()
|
||||
if line:
|
||||
try: recs.append(json.loads(line))
|
||||
except ValueError: pass
|
||||
return recs, noff
|
||||
|
||||
def perform(name, ti):
|
||||
"""Execute the bridged tool for real, like the desktop client would."""
|
||||
if name in ("write_file", "edit_file"):
|
||||
p = ti.get("path", "")
|
||||
dest = p if os.path.isabs(p) else os.path.join(WS, p)
|
||||
os.makedirs(os.path.dirname(dest) or WS, exist_ok=True)
|
||||
body = ti.get("content", "")
|
||||
with open(dest, "w") as f:
|
||||
f.write(body)
|
||||
return "wrote %s (%d bytes)" % (p, len(body.encode()))
|
||||
return "ok"
|
||||
|
||||
class Poller(threading.Thread):
|
||||
def __init__(self, sid):
|
||||
super().__init__(daemon=True); self.sid = sid; self.snaps = []; self.stop = False
|
||||
def run(self):
|
||||
while not self.stop:
|
||||
st, body, js = http("GET", SOUL + "/api/run-progress/" + self.sid, timeout=60)
|
||||
if js and js.get("progress"):
|
||||
if not self.snaps or self.snaps[-1] != js["progress"]:
|
||||
self.snaps.append(js["progress"])
|
||||
time.sleep(0.1)
|
||||
|
||||
def progress(sid):
|
||||
_, _, pj = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
|
||||
return (pj or {}).get("progress")
|
||||
|
||||
def run_phrasing(cname, ph):
|
||||
st, body, js = http("POST", SOUL + "/api/sessions", {"title": ph["id"]}, timeout=60)
|
||||
sid = (js or {}).get("id", "")
|
||||
rec = {"class": cname, "phrasing": ph["id"], "session_id": sid, "legs": [],
|
||||
"pendings": [], "progress_during": [], "progress_per_leg": [],
|
||||
"progress_final": None, "soul_log": "", "stub": [], "http": [],
|
||||
"agentic": cname not in NON_AGENTIC}
|
||||
if not sid:
|
||||
rec["fatal"] = "session create failed: %s %s" % (st, body[:300]); return rec
|
||||
soff = fsize(SOULLOG); loff = fsize(STUBLOG)
|
||||
t0 = time.time()
|
||||
pol = Poller(sid); pol.start()
|
||||
payload = {"message": ph["prompt"], "session_id": sid, "workspace_root": WS,
|
||||
"agentic": rec["agentic"]}
|
||||
if MODE == "local":
|
||||
payload["agent_workspace_root"] = WS
|
||||
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("chat_timeout", 240))
|
||||
rec["http"].append(st)
|
||||
rec["legs"].append(js if js is not None else body[:600])
|
||||
rec["progress_per_leg"].append(progress(sid))
|
||||
hops = 0
|
||||
while isinstance(js, dict) and js.get("tool_pending") and hops < MAXHOPS:
|
||||
rec["pendings"].append({"call_id": js.get("call_id"), "tool_name": js.get("tool_name"),
|
||||
"tool_input": js.get("tool_input"), "risk_tier": js.get("risk_tier"),
|
||||
"narration": js.get("narration"), "tools_used": js.get("tools_used")})
|
||||
try:
|
||||
eff = perform(js.get("tool_name", ""), js.get("tool_input") or {})
|
||||
except Exception as e:
|
||||
eff = "client error: %r" % (e,)
|
||||
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
|
||||
{"call_id": js.get("call_id"), "content": eff},
|
||||
timeout=CFG.get("chat_timeout", 240))
|
||||
rec["http"].append(st)
|
||||
rec["legs"].append(js if js is not None else body[:600])
|
||||
rec["progress_per_leg"].append(progress(sid))
|
||||
hops += 1
|
||||
pol.stop = True; time.sleep(0.3)
|
||||
t1 = time.time()
|
||||
rec["elapsed"] = round(t1 - t0, 2)
|
||||
rec["progress_during"] = pol.snaps
|
||||
rec["progress_final"] = progress(sid)
|
||||
rec["soul_log"], _ = tail_from(SOULLOG, soff)
|
||||
rec["stub"], _ = stub_since(loff)
|
||||
rec["hops"] = hops
|
||||
return rec
|
||||
|
||||
# ------------------------------------------------------------- assertions ----
|
||||
def expected_calls(cname, first_only=False):
|
||||
out = []
|
||||
for step in SCEN["classes"][cname]["script"]:
|
||||
for k, call in enumerate(step.get("tool_calls") or []):
|
||||
if first_only and k > 0:
|
||||
continue
|
||||
out.append((call["name"], call["arguments"]))
|
||||
return out
|
||||
|
||||
def final_text(cname):
|
||||
for step in reversed(SCEN["classes"][cname]["script"]):
|
||||
if step.get("text") and not step.get("tool_calls"):
|
||||
return step["text"]
|
||||
return None
|
||||
|
||||
def judge(rec):
|
||||
cname = rec["class"]; ok = []; bad = []
|
||||
last = rec["legs"][-1] if rec["legs"] else None
|
||||
reply = last.get("reply") if isinstance(last, dict) else None
|
||||
err = last.get("error") if isinstance(last, dict) else None
|
||||
tools_used = last.get("tools_used") if isinstance(last, dict) else None
|
||||
stub = rec["stub"]
|
||||
scen_recs = [r for r in stub if r.get("kind") == "scenario"]
|
||||
rejects = [r for r in stub if r.get("validation") != "ok"]
|
||||
bg = [r for r in stub if r.get("kind") in ("wrong_path", "background")]
|
||||
|
||||
def wire_clean():
|
||||
if rejects:
|
||||
for r in rejects:
|
||||
bad.append("stub REJECTED a request: [%s] %s"
|
||||
% (r.get("validation"), r.get("validation_detail")))
|
||||
else:
|
||||
ok.append("stub ground truth: validation \"ok\" on all %d scenario leg(s), no "
|
||||
"gate_echo_mismatch / gate_tool_call_shape / dialect-leak 400s"
|
||||
% len(scen_recs))
|
||||
if bg:
|
||||
ok.append("NOTE background non-scenario request(s) in this window: %s"
|
||||
% [(r.get("kind"), r.get("path"), r.get("http_status")) for r in bg])
|
||||
|
||||
if cname == "oa-plain":
|
||||
wire_clean()
|
||||
want = final_text(cname)
|
||||
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
|
||||
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
|
||||
if tools_used == []: ok.append("tools_used == [] (no tool ran)")
|
||||
else: bad.append("tools_used expected [] got %r" % (tools_used,))
|
||||
if reply and ('"tool_calls"' in reply or '"function"' in reply or '"tool_use"' in reply):
|
||||
bad.append("tool-call JSON leaked into the reply text")
|
||||
else: ok.append("no tool-call JSON anywhere in the reply")
|
||||
|
||||
elif cname in ("oa-single-tool", "oa-torture", "oa-mission"):
|
||||
wire_clean()
|
||||
want = final_text(cname)
|
||||
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
|
||||
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
|
||||
exp = expected_calls(cname)
|
||||
wantnames = [n for n, _ in exp]
|
||||
if tools_used == wantnames:
|
||||
ok.append("tools_used == %r (carried across %d suspension(s))" % (wantnames, rec["hops"]))
|
||||
else:
|
||||
bad.append("tools_used expected %r got %r" % (wantnames, tools_used))
|
||||
for name, args in exp:
|
||||
p = args.get("path"); c = args.get("content")
|
||||
dest = os.path.join(WS, p)
|
||||
if not os.path.exists(dest):
|
||||
bad.append("expected file missing on disk: %s" % dest); continue
|
||||
got = open(dest, "rb").read()
|
||||
if got == c.encode():
|
||||
ok.append("%s on disk is byte-for-byte the issued payload (%d bytes)" % (p, len(got)))
|
||||
else:
|
||||
bad.append("%s content differs\n WANT %r\n GOT %r"
|
||||
% (p, c[:300], got[:300].decode("utf-8", "replace")))
|
||||
if MODE == "bridge":
|
||||
for pend, (name, args) in zip(rec["pendings"], exp):
|
||||
if pend["tool_input"] == args:
|
||||
ok.append("tool_input for %s survived exactly ONE decode (deep-equal to the "
|
||||
"issued arguments; no double-escaping)" % name)
|
||||
else:
|
||||
bad.append("tool_input != issued arguments for %s\n WANT %r\n GOT %r"
|
||||
% (name, args, pend["tool_input"]))
|
||||
if rec["pendings"] and all(p["risk_tier"] == "escalate" for p in rec["pendings"]):
|
||||
ok.append("every write_file classified \"escalate\" and bridged for consent")
|
||||
|
||||
elif cname == "oa-parallel":
|
||||
drift = [l.strip() for l in rec["soul_log"].splitlines() if "DRIFT: provider returned" in l]
|
||||
if drift: ok.append("soul log: " + drift[0])
|
||||
else: bad.append("no 'DRIFT: provider returned N parallel tool_calls' line in the soul log")
|
||||
delivered = [r for r in stub if r.get("delivered", {}).get("tool_calls")]
|
||||
if delivered and len(delivered[0]["delivered"]["tool_calls"]) == 2:
|
||||
ok.append("stub delivered 2 parallel tool_calls in one response (ground truth)")
|
||||
if MODE == "bridge":
|
||||
if len(rec["pendings"]) == 1:
|
||||
ok.append("exactly ONE call honored: %s" % rec["pendings"][0]["call_id"])
|
||||
else:
|
||||
bad.append("expected exactly 1 honored call, got %d" % len(rec["pendings"]))
|
||||
pairing = [r for r in rejects if "gate_pairing" in str(r.get("validation_detail")) or
|
||||
"tool_calls at end of thread" in str(r.get("validation_detail")) or
|
||||
"not fully answered" in str(r.get("validation_detail"))]
|
||||
for r in pairing:
|
||||
ok.append("EXPECTED-BY-CONTRACT stub 400 on the unpaired echo: %s"
|
||||
% r.get("validation_detail"))
|
||||
other = [r for r in rejects if r not in pairing]
|
||||
for r in other:
|
||||
bad.append("unexpected stub rejection: [%s] %s"
|
||||
% (r.get("validation"), r.get("validation_detail")))
|
||||
if err and not reply:
|
||||
ok.append("honest error envelope after the 400 (no fabricated answer): %r" % err)
|
||||
elif reply == final_text(cname):
|
||||
ok.append("final reply == scripted final text (both calls paired)")
|
||||
else:
|
||||
bad.append("neither an honest error nor the scripted final text: %r" % (last,))
|
||||
|
||||
elif cname == "oa-api-error":
|
||||
if err and not reply:
|
||||
ok.append("honest error envelope: error=%r reply=%r" % (err, reply))
|
||||
else:
|
||||
bad.append("expected an error envelope with an empty reply, got %r" % (last,))
|
||||
delivered = [r["delivered"].get("api_error") for r in stub if r.get("delivered")]
|
||||
ok.append("stub delivered api_error status(es): %r" % [d for d in delivered if d])
|
||||
n = len([r for r in stub if r.get("kind") == "scenario"])
|
||||
ok.append("provider hit %d time(s) - no retry storm" % n)
|
||||
if reply:
|
||||
bad.append("FABRICATED ANSWER: reply non-empty on a provider error")
|
||||
|
||||
elif cname == "oa-tools-off":
|
||||
ok.append("stub records for this phrasing: %r"
|
||||
% [{k: r.get(k) for k in ("kind", "path", "validation", "http_status")} for r in stub])
|
||||
wrong = [r for r in stub if r.get("kind") == "wrong_path"]
|
||||
matched = [r for r in stub if r.get("scenario_class") == cname]
|
||||
if matched and not rejects:
|
||||
ok.append("chat-only request reached /v1/chat/completions with NO tools offered")
|
||||
want = final_text(cname)
|
||||
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
|
||||
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
|
||||
if reply and ('"tool_calls"' in reply or '"function"' in reply):
|
||||
bad.append("tool-call JSON leaked into the reply text")
|
||||
else: ok.append("no tool-call JSON in the reply")
|
||||
elif wrong:
|
||||
bad.append("the non-agentic lane never reached the provider endpoint: stub saw "
|
||||
"%s -> %s (the el-runtime provider chain appends /v1/chat/completions "
|
||||
"to NEURON_LLM_0_URL, chat.el appends only /chat/completions)"
|
||||
% (wrong[0]["path"], wrong[0]["http_status"]))
|
||||
elif not stub:
|
||||
bad.append("no request reached the stub at all")
|
||||
else:
|
||||
for r in rejects:
|
||||
bad.append("stub REJECTED: [%s] %s" % (r.get("validation"), r.get("validation_detail")))
|
||||
return ok, bad
|
||||
|
||||
def main():
|
||||
results = []
|
||||
for cname in ONLY:
|
||||
for ph in SCEN["classes"][cname]["phrasings"]:
|
||||
rec = run_phrasing(cname, ph)
|
||||
ok, bad = judge(rec)
|
||||
rec["ok"] = ok; rec["bad"] = bad
|
||||
rec["verdict"] = "FAIL" if bad else "PASS"
|
||||
results.append(rec)
|
||||
print("=" * 78)
|
||||
print("[%s] %s / %s (%.2fs, %d bridge hop(s), agentic=%s, mode=%s)"
|
||||
% (rec["verdict"], cname, ph["id"], rec.get("elapsed", 0),
|
||||
rec.get("hops", 0), rec["agentic"], MODE))
|
||||
for l in ok: print(" ok " + l.replace("\n", "\n "))
|
||||
for l in bad: print(" FAIL " + l.replace("\n", "\n "))
|
||||
for i, leg in enumerate(rec["legs"]):
|
||||
print(" leg%d envelope: %s" % (i, json.dumps(leg)[:430]))
|
||||
for i, pr in enumerate(rec["progress_per_leg"]):
|
||||
print(" run-progress after leg%d: %s" % (i, json.dumps(pr)[:380]))
|
||||
if rec["progress_during"]:
|
||||
print(" run-progress polled DURING (%d distinct snapshot(s)), last: %s"
|
||||
% (len(rec["progress_during"]), json.dumps(rec["progress_during"][-1])[:300]))
|
||||
for r in rec["stub"]:
|
||||
print(" stub: kind=%s class=%s phrasing=%s step=%s validation=%s%s delivered=%s http=%s"
|
||||
% (r.get("kind"), r.get("scenario_class"), r.get("phrasing"), r.get("step"),
|
||||
r.get("validation"),
|
||||
("(" + str(r.get("validation_detail")) + ")") if r.get("validation_detail") else "",
|
||||
json.dumps(r.get("delivered")), r.get("http_status")))
|
||||
if rec["soul_log"].strip():
|
||||
for l in rec["soul_log"].splitlines():
|
||||
if l.strip(): print(" soul: " + l.strip())
|
||||
json.dump(results, open(OUT, "w"), indent=1)
|
||||
npass = sum(1 for r in results if r["verdict"] == "PASS")
|
||||
print("=" * 78)
|
||||
print("PHASE %s: %d/%d PASS" % (MODE, npass, len(results)))
|
||||
for r in results:
|
||||
print(" %-6s %-16s %s" % (r["verdict"], r["class"], r["phrasing"]))
|
||||
return 0 if npass == len(results) else 1
|
||||
|
||||
sys.exit(main())
|
||||
PYEOF
|
||||
|
||||
# ------------------------------------------------------------- hostile drv ---
|
||||
cat > "$RUN/hostile.py" <<'PYEOF'
|
||||
import json, os, sys, time, urllib.request, urllib.error
|
||||
|
||||
CFG = json.load(open(sys.argv[1]))
|
||||
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
|
||||
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
|
||||
|
||||
def http(method, url, obj=None, timeout=400):
|
||||
data = None if obj is None else json.dumps(obj).encode()
|
||||
req = urllib.request.Request(url, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
b = r.read().decode("utf-8", "replace"); st = r.status
|
||||
except urllib.error.HTTPError as e:
|
||||
b = e.read().decode("utf-8", "replace"); st = e.code
|
||||
except Exception as e:
|
||||
return -1, "TRANSPORT-ERROR: %r" % (e,), None
|
||||
try:
|
||||
return st, b, json.loads(b)
|
||||
except ValueError:
|
||||
return st, b, None
|
||||
|
||||
mode = CFG["mode"]; wsmode = CFG["ws_mode"]; WS = CFG["workspace"]
|
||||
_, _, js = http("POST", SOUL + "/api/sessions", {"title": "hostile-" + mode}, timeout=60)
|
||||
sid = (js or {}).get("id", "")
|
||||
payload = {"message": "oa-gate plain probe: hostile mode %s" % mode,
|
||||
"agentic": True, "session_id": sid, "workspace_root": WS}
|
||||
if wsmode == "local":
|
||||
payload["agent_workspace_root"] = WS
|
||||
t0 = time.time()
|
||||
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("timeout", 400))
|
||||
t_first = time.time() - t0
|
||||
legs = [js if js is not None else body[:500]]
|
||||
hops = 0
|
||||
while isinstance(js, dict) and js.get("tool_pending") and hops < CFG.get("max_hops", 14):
|
||||
ti = js.get("tool_input") or {}
|
||||
p = ti.get("path", "x.md")
|
||||
dest = p if os.path.isabs(p) else os.path.join(WS, p)
|
||||
try: open(dest, "w").write(ti.get("content", ""))
|
||||
except Exception: pass
|
||||
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
|
||||
{"call_id": js.get("call_id"), "content": "ok"},
|
||||
timeout=CFG.get("timeout", 400))
|
||||
legs.append(js if js is not None else body[:500]); hops += 1
|
||||
el = time.time() - t0
|
||||
_, _, stats = http("GET", STUB + "/gate/stats", timeout=30)
|
||||
_, _, prog = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
|
||||
fab = [l for l in legs if isinstance(l, dict) and l.get("reply")]
|
||||
print("HOSTILE %s (ws_mode=%s)" % (mode, wsmode))
|
||||
print(" first /api/chat POST returned after %.2fs; whole chain %.2fs; client bridge hops=%d; "
|
||||
"stub chat_hits=%s" % (t_first, el, hops, (stats or {}).get("chat_hits")))
|
||||
print(" first envelope : " + json.dumps(legs[0])[:430])
|
||||
print(" final envelope : " + json.dumps(legs[-1])[:430])
|
||||
print(" non-empty replies anywhere in the chain (fabrication check): %d" % len(fab))
|
||||
print(" run-progress : " + json.dumps(prog)[:300])
|
||||
json.dump({"mode": mode, "ws_mode": wsmode, "t_first": t_first, "elapsed": el, "hops": hops,
|
||||
"chat_hits": (stats or {}).get("chat_hits"), "legs": legs, "progress": prog},
|
||||
open(CFG["out"], "w"), indent=1)
|
||||
PYEOF
|
||||
|
||||
# ------------------------------------------------------------------ phases ---
|
||||
RC_BRIDGE=0; RC_LOCAL=0; RC_OFF=0
|
||||
run_normal_phase() { # $1 = label, $2 = soul port, $3 = agent root, $4 = ws, $5 = base, $6 = classes json
|
||||
local m="$1" port="$2" root="$3" ws="$4" base="$5" classes="$6"
|
||||
echo; echo "############ PHASE: $m (soul :$port, NEURON_LLM_0_URL=$base) ############"
|
||||
start_stub normal "$RUN/stub-$m.jsonl"
|
||||
start_soul "$port" "$base" "$RUN/soul-$m.log" "$root"
|
||||
cat > "$RUN/cfg-$m.json" <<JSON
|
||||
{"soul_port": $port, "stub_port": $STUB_PORT, "workspace": "$ws", "mode": "$m",
|
||||
"scenarios": "$HERE/scenarios-openai.json", "stub_log": "$RUN/stub-$m.jsonl",
|
||||
"soul_log": "$RUN/soul-$m.log", "out": "$RUN/results-$m.json", "chat_timeout": 240,
|
||||
"classes": $classes}
|
||||
JSON
|
||||
python3 "$DRV" "$RUN/cfg-$m.json"
|
||||
local rc=$?
|
||||
stop_soul; stop_stub
|
||||
return $rc
|
||||
}
|
||||
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "bridge" ]; then
|
||||
run_normal_phase bridge "$SOUL_PORT" "" "$RUN/ws-bridge" "http://127.0.0.1:$STUB_PORT/v1" null
|
||||
RC_BRIDGE=$?
|
||||
fi
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "local" ]; then
|
||||
run_normal_phase local "$SOUL_PORT_B" "$RUN/ws-local" "$RUN/ws-local" "http://127.0.0.1:$STUB_PORT/v1" null
|
||||
RC_LOCAL=$?
|
||||
fi
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "toolsoff" ]; then
|
||||
# supplementary: the el-runtime provider chain appends /v1/chat/completions itself,
|
||||
# so the non-agentic door needs the base WITHOUT the /v1 suffix.
|
||||
run_normal_phase toolsoff "$SOUL_PORT" "" "$RUN/ws-off" "http://127.0.0.1:$STUB_PORT" '["oa-tools-off","oa-plain"]'
|
||||
RC_OFF=$?
|
||||
fi
|
||||
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "hostile" ]; then
|
||||
echo; echo "############ PHASE: hostile ############"
|
||||
for spec in "black-hole:bridge" "mid-body-drop:bridge" "tool-pending-forever:bridge" "tool-pending-forever:local"; do
|
||||
mode="${spec%%:*}"; wsm="${spec##*:}"
|
||||
echo; echo "---- hostile mode=$mode ws_mode=$wsm ----"
|
||||
start_stub "$mode" "$RUN/stub-$mode-$wsm.jsonl"
|
||||
if [ "$wsm" = "local" ]; then
|
||||
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" "$RUN/ws-local"
|
||||
else
|
||||
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" ""
|
||||
fi
|
||||
cat > "$RUN/cfg-$mode-$wsm.json" <<JSON
|
||||
{"soul_port": $SOUL_PORT, "stub_port": $STUB_PORT, "mode": "$mode", "ws_mode": "$wsm",
|
||||
"workspace": "$RUN/ws-local", "out": "$RUN/hostile-$mode-$wsm.json", "timeout": 400}
|
||||
JSON
|
||||
python3 "$RUN/hostile.py" "$RUN/cfg-$mode-$wsm.json"
|
||||
echo " soul log (llm/DRIFT/cap lines):"
|
||||
grep -E "DRIFT|llm error|iteration cap|\[llm\]" "$RUN/soul-$mode-$wsm.log" | tail -8 | sed 's/^/ /'
|
||||
stop_soul; stop_stub
|
||||
done
|
||||
fi
|
||||
|
||||
echo; echo "############ CLEANUP ############"
|
||||
cleanup
|
||||
sleep 0.5
|
||||
echo "processes still matching the soul binary:"; pgrep -fl "$SOUL_BIN" || echo " (none)"
|
||||
echo "processes still matching stub-openai.py:"; pgrep -fl "stub-openai.py" || echo " (none)"
|
||||
echo "lsof on 7891-7894 after cleanup:"
|
||||
lsof -nP -iTCP:7891 -iTCP:7892 -iTCP:7893 -iTCP:7894 2>/dev/null || echo " (no listeners - ports free)"
|
||||
echo
|
||||
echo "############ SUMMARY ############"
|
||||
echo "run dir: $RUN"
|
||||
echo "bridge rc=$RC_BRIDGE local rc=$RC_LOCAL toolsoff rc=$RC_OFF (0 = every class PASS)"
|
||||
exit $(( RC_BRIDGE + RC_LOCAL + RC_OFF ))
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"_comment": "OpenAI-dialect gate scenario contract (soul-openai-tools-v2). Single source of truth shared by stub-openai.py (scripted provider responses + request assertions), selftest.sh (stub self-verification), and the future brain-side gate driver. Same structure as gate9's scenarios.json: classes -> script + phrasings with markers; scripts are CLASS-level so assertions are behavioral, never pinned to a sentence. expect_request keys: require_tools, require_tool_choice, parallel_tool_calls (expected literal value; null = don't check), forbid_tools. defaults apply to every class unless overridden; steps may override with their own expect_request.",
|
||||
"deadline_secs": 60,
|
||||
"max_loop_iterations": 16,
|
||||
"defaults": {
|
||||
"expect_request": {
|
||||
"require_tools": true,
|
||||
"require_tool_choice": true,
|
||||
"parallel_tool_calls": false
|
||||
}
|
||||
},
|
||||
"classes": {
|
||||
"oa-plain": {
|
||||
"script": [
|
||||
{ "text": "Plain OpenAI-lane answer (gate fixture): the mechanism, the main caveat, and the practical takeaway in three sentences. No tools were needed for this one, and the finish reason on the wire is stop, which the loop must treat as terminal." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-plain-p1", "marker": "oa-gate plain probe", "prompt": "oa-gate plain probe: explain the fixture topic simply." },
|
||||
{ "id": "oa-plain-p2", "marker": "oa-gate second plain", "prompt": "oa-gate second plain: another phrasing of the plain question." }
|
||||
]
|
||||
},
|
||||
"oa-tools-off": {
|
||||
"expect_request": {
|
||||
"require_tools": false,
|
||||
"forbid_tools": true,
|
||||
"require_tool_choice": false,
|
||||
"parallel_tool_calls": null
|
||||
},
|
||||
"script": [
|
||||
{ "text": "Chat-only OpenAI-lane answer (gate fixture): this lane offered no tools and none were used; the reply is plain text with finish reason stop." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-tools-off-p1", "marker": "oa-gate tools-off probe", "prompt": "oa-gate tools-off probe: plain chat with no tools offered." }
|
||||
]
|
||||
},
|
||||
"oa-single-tool": {
|
||||
"script": [
|
||||
{ "text": "Step 1: writing the note.",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file",
|
||||
"arguments": { "path": "openai-single-note.md", "content": "# Note (gate fixture, OpenAI lane)\n\nDeterministic single-tool body.\n" } }
|
||||
] },
|
||||
{ "text": "All set - openai-single-note.md is written with the fixture body. Nothing else was needed for this one." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-single-p1", "marker": "oa-gate single tool note", "prompt": "oa-gate single tool note: save the fixture note to a file." },
|
||||
{ "id": "oa-single-p2", "marker": "oa-gate one file please", "prompt": "oa-gate one file please: write the fixture note file." }
|
||||
]
|
||||
},
|
||||
"oa-torture": {
|
||||
"script": [
|
||||
{ "tool_calls": [
|
||||
{ "name": "write_file",
|
||||
"arguments": { "path": "torture-note.md", "content": "Line 1 has \"double quotes\", 'singles', and a mid-line backslash \\ here.\nLine 2\thas a tab, a literal \\n two-char sequence, and a Windows path C:\\temp\\new.txt.\nLine 3 unicode: naïve café — 日本語 ✓ 🚀\nLine 4 JSON-in-string: {\"k\": \"v\", \"arr\": [1, 2], \"s\": \"nested \\\"deep\\\" quotes\"}\nLine 5 ends with a lone backslash \\" } }
|
||||
] },
|
||||
{ "text": "Torture round-trip complete: the payload with nested quotes, backslashes, newlines, tabs, and unicode survived exactly one encode and one decode." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-torture-p1", "marker": "oa-gate torture probe", "prompt": "oa-gate torture probe: write the escaping torture file." }
|
||||
]
|
||||
},
|
||||
"oa-parallel": {
|
||||
"script": [
|
||||
{ "text": "Step 1: two writes at once (parallel probe).",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file", "arguments": { "path": "parallel-a.md", "content": "Parallel A (gate fixture).\n" } },
|
||||
{ "name": "write_file", "arguments": { "path": "parallel-b.md", "content": "Parallel B (gate fixture).\n" } }
|
||||
] },
|
||||
{ "text": "Parallel probe complete: both tool results arrived and were paired correctly. A brain that instead rejects the double call must do so cleanly - that outcome is asserted brain-side, not here." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-parallel-p1", "marker": "oa-gate parallel probe", "prompt": "oa-gate parallel probe: run the two-write parallel case." }
|
||||
]
|
||||
},
|
||||
"oa-mission": {
|
||||
"script": [
|
||||
{ "text": "Step 1: drafting part one.",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file", "arguments": { "path": "mission-part-1.md", "content": "Mission part 1 (gate fixture).\n" } }
|
||||
] },
|
||||
{ "text": "Step 2: drafting part two.",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file", "arguments": { "path": "mission-part-2.md", "content": "Mission part 2 (gate fixture).\n" } }
|
||||
] },
|
||||
{ "text": "Mission complete: mission-part-1.md and mission-part-2.md are written; the loop ran two tool rounds and finished cleanly with finish reason stop." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-mission-p1", "marker": "oa-gate mission probe", "prompt": "oa-gate mission probe: run the two-round mission." }
|
||||
]
|
||||
},
|
||||
"oa-api-error": {
|
||||
"expect_request": {
|
||||
"require_tools": false,
|
||||
"require_tool_choice": false,
|
||||
"parallel_tool_calls": null
|
||||
},
|
||||
"script": [],
|
||||
"phrasings": [
|
||||
{ "id": "oa-err-400", "marker": "oa-gate error four hundred", "prompt": "oa-gate error four hundred: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 400, "type": "invalid_request_error", "message": "gate-injected 400: request rejected by fixture", "code": "gate_injected" } } ] },
|
||||
{ "id": "oa-err-429", "marker": "oa-gate error rate limit", "prompt": "oa-gate error rate limit: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 429, "type": "rate_limit_error", "message": "gate-injected 429: rate limited by fixture", "code": "rate_limit_exceeded" } } ] },
|
||||
{ "id": "oa-err-500", "marker": "oa-gate error five hundred", "prompt": "oa-gate error five hundred: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 500, "type": "server_error", "message": "gate-injected 500: internal fixture error", "code": "gate_injected" } } ] },
|
||||
{ "id": "oa-err-503", "marker": "oa-gate error unavailable", "prompt": "oa-gate error unavailable: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 503, "type": "server_error", "message": "gate-injected 503: fixture overloaded", "code": "gate_injected" } } ] }
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
Executable
+409
@@ -0,0 +1,409 @@
|
||||
#!/usr/bin/env bash
|
||||
# selftest.sh - proves stub-openai.py before any brain code exists.
|
||||
# Drives the stub with curl through every scenario (plain, tools-off,
|
||||
# single tool round-trip, escaping torture, parallel double-call,
|
||||
# two-round mission, injected API errors, background, overrun), every
|
||||
# validation rejection (dialect leaks, pairing, echo round-trip, scenario
|
||||
# expectations), and all three hostile modes. Exit 0 = green.
|
||||
set -u
|
||||
cd "$(dirname "$0")" || exit 1
|
||||
PY=python3
|
||||
TMP="$(mktemp -d)"
|
||||
PIDS=()
|
||||
cleanup() {
|
||||
for p in "${PIDS[@]:-}"; do kill -9 "$p" >/dev/null 2>&1; done
|
||||
rm -rf "$TMP"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
PASS=0; FAIL=0
|
||||
ok() { printf 'ok - %s\n' "$1"; PASS=$((PASS+1)); }
|
||||
bad() { printf 'FAIL - %s\n' "$1"; FAIL=$((FAIL+1)); }
|
||||
check() { # check <name> <cmd...> - pass if cmd exits 0; show output on fail
|
||||
local name="$1"; shift
|
||||
local out
|
||||
if out="$("$@" 2>&1)"; then ok "$name"
|
||||
else bad "$name"; [ -n "$out" ] && printf '%s\n' "$out" | sed 's/^/ /' | head -8
|
||||
fi
|
||||
}
|
||||
|
||||
freeport() { "$PY" -c 'import socket;s=socket.socket();s.bind(("127.0.0.1",0));print(s.getsockname()[1]);s.close()'; }
|
||||
waithealth() {
|
||||
local p="$1" i
|
||||
for i in $(seq 1 60); do
|
||||
curl -sf "http://127.0.0.1:$p/gate/health" >/dev/null 2>&1 && return 0
|
||||
sleep 0.1
|
||||
done
|
||||
echo "stub on :$p never became healthy"; return 1
|
||||
}
|
||||
post() { # post <port> <bodyfile> <respfile> [extra curl args...] -> echoes http code
|
||||
local port="$1" body="$2" resp="$3"; shift 3
|
||||
curl -s -o "$resp" -w '%{http_code}' -H 'content-type: application/json' \
|
||||
"$@" --data-binary @"$body" "http://127.0.0.1:$port/v1/chat/completions"
|
||||
}
|
||||
|
||||
# ---- embedded helper: builds OpenAI-dialect bodies, asserts on responses ----
|
||||
cat > "$TMP/helpers.py" <<'PYEOF'
|
||||
import copy, json, sys
|
||||
|
||||
TOOLS = [
|
||||
{"type": "function", "function": {
|
||||
"name": "write_file", "description": "Write content to a file on disk.",
|
||||
"parameters": {"type": "object",
|
||||
"properties": {"path": {"type": "string"},
|
||||
"content": {"type": "string"}},
|
||||
"required": ["path", "content"]}}},
|
||||
{"type": "function", "function": {
|
||||
"name": "read_file", "description": "Read contents of a file from disk.",
|
||||
"parameters": {"type": "object",
|
||||
"properties": {"path": {"type": "string"}},
|
||||
"required": ["path"]}}},
|
||||
]
|
||||
|
||||
def dump(obj, out):
|
||||
json.dump(obj, open(out, "w"), ensure_ascii=False)
|
||||
|
||||
def base(prompt, tools=True):
|
||||
b = {"model": "gate-openai-model", "max_tokens": 1024,
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are Neuron (gate fixture)."},
|
||||
{"role": "user", "content": prompt}]}
|
||||
if tools:
|
||||
b["tools"] = copy.deepcopy(TOOLS)
|
||||
b["tool_choice"] = "auto"
|
||||
b["parallel_tool_calls"] = False
|
||||
return b
|
||||
|
||||
def cmd_plain(out, prompt):
|
||||
dump(base(prompt), out)
|
||||
|
||||
def cmd_notools(out, prompt):
|
||||
dump(base(prompt, tools=False), out)
|
||||
|
||||
def cmd_mut(out, prompt, mutation):
|
||||
b = base(prompt)
|
||||
if mutation == "no-tool-choice":
|
||||
del b["tool_choice"]
|
||||
elif mutation == "ptc-true":
|
||||
b["parallel_tool_calls"] = True
|
||||
elif mutation == "top-system":
|
||||
b["system"] = "You are Neuron."
|
||||
elif mutation == "anth-tools":
|
||||
b["tools"] = [{"name": "write_file", "description": "x",
|
||||
"input_schema": {"type": "object", "properties": {}}}]
|
||||
elif mutation == "anth-block":
|
||||
b["messages"][1] = {"role": "user", "content": [
|
||||
{"type": "tool_result", "tool_use_id": "toolu_x", "content": "hi"},
|
||||
{"type": "text", "text": prompt}]}
|
||||
else:
|
||||
raise SystemExit("unknown mutation " + mutation)
|
||||
dump(b, out)
|
||||
|
||||
def cmd_chain(out, prompt, variant, *resps):
|
||||
"""Build the next leg: echo each response's assistant turn and answer its
|
||||
tool calls. `variant` applies to the LAST response only:
|
||||
ok | no-tool-turn | wrong-id | only-first | double-encode | object-args"""
|
||||
b = base(prompt)
|
||||
for idx, p in enumerate(resps):
|
||||
last = idx == len(resps) - 1
|
||||
msg = json.load(open(p))["choices"][0]["message"]
|
||||
tcs = msg.get("tool_calls")
|
||||
if not tcs:
|
||||
b["messages"].append({"role": "assistant",
|
||||
"content": msg.get("content")})
|
||||
continue
|
||||
v = variant if last else "ok"
|
||||
asst = {"role": "assistant", "content": msg.get("content"),
|
||||
"tool_calls": copy.deepcopy(tcs)}
|
||||
if v == "double-encode":
|
||||
for tc in asst["tool_calls"]:
|
||||
tc["function"]["arguments"] = json.dumps(
|
||||
tc["function"]["arguments"])
|
||||
if v == "object-args":
|
||||
for tc in asst["tool_calls"]:
|
||||
tc["function"]["arguments"] = json.loads(
|
||||
tc["function"]["arguments"])
|
||||
b["messages"].append(asst)
|
||||
if v == "no-tool-turn":
|
||||
continue
|
||||
use = tcs[:1] if v == "only-first" else tcs
|
||||
for tc in use:
|
||||
tid = "call_bogus_123" if v == "wrong-id" else tc["id"]
|
||||
b["messages"].append({"role": "tool", "tool_call_id": tid,
|
||||
"content": "{\"ok\":true,\"bytes\":42}"})
|
||||
dump(b, out)
|
||||
|
||||
def cmd_chk(resp, expr):
|
||||
r = json.load(open(resp))
|
||||
if not eval(expr, {"r": r, "json": json, "len": len, "str": str,
|
||||
"isinstance": isinstance, "any": any, "all": all,
|
||||
"sorted": sorted}):
|
||||
print("assertion failed:", expr)
|
||||
print("resp:", json.dumps(r, ensure_ascii=False)[:400])
|
||||
raise SystemExit(1)
|
||||
|
||||
def cmd_torture(resp, scen):
|
||||
r = json.load(open(resp))
|
||||
tc = r["choices"][0]["message"]["tool_calls"][0]
|
||||
raw = tc["function"]["arguments"]
|
||||
assert isinstance(raw, str), "arguments must be a JSON-encoded string"
|
||||
got = json.loads(raw)
|
||||
exp = json.load(open(scen))["classes"]["oa-torture"]["script"][0]["tool_calls"][0]["arguments"]
|
||||
assert got == exp, "decoded arguments != scripted torture payload"
|
||||
content = got["content"]
|
||||
for needle in ['"', "\\", "\n", "\t", "日本語", "naïve", "🚀"]:
|
||||
assert needle in content, "missing torture needle %r" % needle
|
||||
|
||||
def cmd_notjson(path):
|
||||
data = open(path, "rb").read()
|
||||
assert data, "file empty - no partial body arrived"
|
||||
try:
|
||||
json.loads(data.decode("utf-8", "replace"))
|
||||
except ValueError:
|
||||
return
|
||||
raise SystemExit("partial body unexpectedly parsed as complete JSON")
|
||||
|
||||
def cmd_pending(*paths):
|
||||
ids = []
|
||||
for p in paths:
|
||||
c = json.load(open(p))["choices"][0]
|
||||
assert c["finish_reason"] == "tool_calls", c["finish_reason"]
|
||||
tc = c["message"]["tool_calls"][0]
|
||||
assert tc["function"]["name"] == "write_file"
|
||||
json.loads(tc["function"]["arguments"]) # must decode
|
||||
ids.append(tc["id"])
|
||||
assert len(set(ids)) == len(ids), "call ids not distinct: %r" % ids
|
||||
|
||||
def cmd_logcheck(path):
|
||||
recs = [json.loads(l) for l in open(path) if l.strip()]
|
||||
seqs = [r["seq"] for r in recs]
|
||||
assert seqs == sorted(seqs) and len(set(seqs)) == len(seqs), "seq not monotonic"
|
||||
kinds = {}
|
||||
for r in recs:
|
||||
kinds[r["kind"]] = kinds.get(r["kind"], 0) + 1
|
||||
assert kinds.get("scenario", 0) >= 10, "too few scenario records: %r" % kinds
|
||||
assert kinds.get("background", 0) >= 1, "no background record"
|
||||
assert kinds.get("overrun", 0) >= 1, "no overrun record"
|
||||
rejected = [r for r in recs if r["validation"] == "rejected"]
|
||||
assert len(rejected) >= 10, "too few rejected records: %d" % len(rejected)
|
||||
assert any(r["delivered"].get("tool_calls") == ["write_file"]
|
||||
for r in recs), "no single write_file ground truth"
|
||||
assert any(r["delivered"].get("tool_calls") == ["write_file", "write_file"]
|
||||
for r in recs), "no parallel ground truth"
|
||||
|
||||
def main():
|
||||
fn = globals()["cmd_" + sys.argv[1].replace("-", "_")]
|
||||
fn(*sys.argv[2:])
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
PYEOF
|
||||
mk() { "$PY" "$TMP/helpers.py" "$@"; }
|
||||
|
||||
echo "=== stub-openai selftest ==="
|
||||
|
||||
# ---- normal mode ------------------------------------------------------------
|
||||
PORT="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$PORT" --scenarios scenarios-openai.json \
|
||||
--log "$TMP/req.jsonl" >"$TMP/stub.out" 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "stub starts and answers /gate/health" waithealth "$PORT"
|
||||
|
||||
# 1. plain completion
|
||||
mk plain "$TMP/plain.json" "oa-gate plain probe: explain the fixture topic simply."
|
||||
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_plain.json")"
|
||||
check "plain: HTTP 200" test "$code" = "200"
|
||||
check "plain: chat.completion envelope, finish stop, real content" mk chk "$TMP/r_plain.json" \
|
||||
'r["object"]=="chat.completion" and r["choices"][0]["finish_reason"]=="stop" and isinstance(r["choices"][0]["message"]["content"],str) and len(r["choices"][0]["message"]["content"])>40'
|
||||
|
||||
# 2. tools-off lane (chat-only request accepted, tool-bearing request refused)
|
||||
mk notools "$TMP/toolsoff.json" "oa-gate tools-off probe: plain chat with no tools offered."
|
||||
code="$(post "$PORT" "$TMP/toolsoff.json" "$TMP/r_toolsoff.json")"
|
||||
check "tools-off: chat-only request -> 200" test "$code" = "200"
|
||||
mk plain "$TMP/toolsoff_bad.json" "oa-gate tools-off probe: plain chat with no tools offered."
|
||||
code="$(post "$PORT" "$TMP/toolsoff_bad.json" "$TMP/r_toolsoff_bad.json")"
|
||||
check "tools-off negative: offering tools -> 400 gate_expect" \
|
||||
bash -c "test $code = 400"
|
||||
check "tools-off negative: reason names gate_expect" mk chk "$TMP/r_toolsoff_bad.json" \
|
||||
'r["error"]["code"]=="gate_expect"'
|
||||
|
||||
# 3. dialect-leak rejections (the loud-failure contract)
|
||||
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_leak_hdr.json" -H 'anthropic-version: 2023-06-01')"
|
||||
check "leak: anthropic-version header -> 400" test "$code" = "400"
|
||||
check "leak: header reason names the leak" mk chk "$TMP/r_leak_hdr.json" \
|
||||
'r["error"]["code"]=="gate_dialect_leak" and "anthropic-version" in r["error"]["message"]'
|
||||
mk mut "$TMP/leak_tools.json" "oa-gate plain probe: explain the fixture topic simply." anth-tools
|
||||
code="$(post "$PORT" "$TMP/leak_tools.json" "$TMP/r_leak_tools.json")"
|
||||
check "leak: input_schema tools -> 400 gate_dialect_leak" bash -c \
|
||||
"test $code = 400"
|
||||
check "leak: input_schema reason" mk chk "$TMP/r_leak_tools.json" \
|
||||
'r["error"]["code"]=="gate_dialect_leak" and "input_schema" in r["error"]["message"]'
|
||||
mk mut "$TMP/leak_sys.json" "oa-gate plain probe: explain the fixture topic simply." top-system
|
||||
code="$(post "$PORT" "$TMP/leak_sys.json" "$TMP/r_leak_sys.json")"
|
||||
check "leak: top-level system -> 400" test "$code" = "400"
|
||||
mk mut "$TMP/leak_block.json" "oa-gate plain probe: explain the fixture topic simply." anth-block
|
||||
code="$(post "$PORT" "$TMP/leak_block.json" "$TMP/r_leak_block.json")"
|
||||
check "leak: Anthropic tool_result content block -> 400" test "$code" = "400"
|
||||
|
||||
# 4. scenario request expectations
|
||||
mk mut "$TMP/no_tc.json" "oa-gate plain probe: explain the fixture topic simply." no-tool-choice
|
||||
code="$(post "$PORT" "$TMP/no_tc.json" "$TMP/r_no_tc.json")"
|
||||
check "expect: missing tool_choice -> 400" test "$code" = "400"
|
||||
mk mut "$TMP/ptc.json" "oa-gate plain probe: explain the fixture topic simply." ptc-true
|
||||
code="$(post "$PORT" "$TMP/ptc.json" "$TMP/r_ptc.json")"
|
||||
check "expect: parallel_tool_calls true -> 400 (ADR-0005 pin)" test "$code" = "400"
|
||||
|
||||
# 5. single tool round-trip
|
||||
ST_PROMPT="oa-gate single tool note: save the fixture note to a file."
|
||||
mk plain "$TMP/st1.json" "$ST_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/st1.json" "$TMP/r_st1.json")"
|
||||
check "single-tool leg1: HTTP 200" test "$code" = "200"
|
||||
check "single-tool leg1: one write_file call, finish tool_calls, string args" mk chk "$TMP/r_st1.json" \
|
||||
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==1 and r["choices"][0]["message"]["tool_calls"][0]["type"]=="function" and r["choices"][0]["message"]["tool_calls"][0]["function"]["name"]=="write_file" and isinstance(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"],str) and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="openai-single-note.md"'
|
||||
mk chain "$TMP/st2.json" "$ST_PROMPT" ok "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2.json" "$TMP/r_st2.json")"
|
||||
check "single-tool leg2: echo + tool turn -> 200 final text" test "$code" = "200"
|
||||
check "single-tool leg2: final names the file, finish stop" mk chk "$TMP/r_st2.json" \
|
||||
'r["choices"][0]["finish_reason"]=="stop" and "openai-single-note.md" in r["choices"][0]["message"]["content"]'
|
||||
mk chain "$TMP/st2_no.json" "$ST_PROMPT" no-tool-turn "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2_no.json" "$TMP/r_st2_no.json")"
|
||||
check "single-tool negative: echo without tool turn -> 400 gate_pairing" \
|
||||
bash -c "test $code = 400"
|
||||
check "single-tool negative: pairing reason" mk chk "$TMP/r_st2_no.json" \
|
||||
'r["error"]["code"]=="gate_pairing"'
|
||||
mk chain "$TMP/st2_wrong.json" "$ST_PROMPT" wrong-id "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2_wrong.json" "$TMP/r_st2_wrong.json")"
|
||||
check "single-tool negative: wrong tool_call_id -> 400" test "$code" = "400"
|
||||
mk chain "$TMP/st2_obj.json" "$ST_PROMPT" object-args "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2_obj.json" "$TMP/r_st2_obj.json")"
|
||||
check "single-tool negative: arguments echoed as object -> 400 shape" \
|
||||
bash -c "test $code = 400"
|
||||
check "single-tool negative: shape reason names STRING" mk chk "$TMP/r_st2_obj.json" \
|
||||
'r["error"]["code"]=="gate_tool_call_shape" and "STRING" in r["error"]["message"]'
|
||||
|
||||
# 6. escaping torture (the two-escaper trap, spec section 6)
|
||||
T_PROMPT="oa-gate torture probe: write the escaping torture file."
|
||||
mk plain "$TMP/t1.json" "$T_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/t1.json" "$TMP/r_t1.json")"
|
||||
check "torture leg1: HTTP 200" test "$code" = "200"
|
||||
check "torture leg1: arguments decode to the exact nasty payload" \
|
||||
mk torture "$TMP/r_t1.json" scenarios-openai.json
|
||||
mk chain "$TMP/t2.json" "$T_PROMPT" ok "$TMP/r_t1.json"
|
||||
code="$(post "$PORT" "$TMP/t2.json" "$TMP/r_t2.json")"
|
||||
check "torture leg2: faithful echo -> 200 final" test "$code" = "200"
|
||||
mk chain "$TMP/t2_dbl.json" "$T_PROMPT" double-encode "$TMP/r_t1.json"
|
||||
code="$(post "$PORT" "$TMP/t2_dbl.json" "$TMP/r_t2_dbl.json")"
|
||||
check "torture negative: double-encoded echo -> 400" test "$code" = "400"
|
||||
check "torture negative: reason names the two-escaper trap" mk chk "$TMP/r_t2_dbl.json" \
|
||||
'r["error"]["code"]=="gate_echo_mismatch" and "two-escaper" in r["error"]["message"]'
|
||||
|
||||
# 7. parallel double-call
|
||||
P_PROMPT="oa-gate parallel probe: run the two-write parallel case."
|
||||
mk plain "$TMP/p1.json" "$P_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/p1.json" "$TMP/r_p1.json")"
|
||||
check "parallel leg1: TWO tool_calls, distinct ids" mk chk "$TMP/r_p1.json" \
|
||||
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==2 and r["choices"][0]["message"]["tool_calls"][0]["id"]!=r["choices"][0]["message"]["tool_calls"][1]["id"]'
|
||||
mk chain "$TMP/p2.json" "$P_PROMPT" ok "$TMP/r_p1.json"
|
||||
code="$(post "$PORT" "$TMP/p2.json" "$TMP/r_p2.json")"
|
||||
check "parallel leg2: both results -> 200 final" test "$code" = "200"
|
||||
mk chain "$TMP/p2_one.json" "$P_PROMPT" only-first "$TMP/r_p1.json"
|
||||
code="$(post "$PORT" "$TMP/p2_one.json" "$TMP/r_p2_one.json")"
|
||||
check "parallel negative: answering only one call -> 400 pairing" test "$code" = "400"
|
||||
|
||||
# 8. two-round mission (loop continuation + step indexing)
|
||||
M_PROMPT="oa-gate mission probe: run the two-round mission."
|
||||
mk plain "$TMP/m1.json" "$M_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/m1.json" "$TMP/r_m1.json")"
|
||||
check "mission leg1: part-1 tool call" mk chk "$TMP/r_m1.json" \
|
||||
'json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-1.md"'
|
||||
mk chain "$TMP/m2.json" "$M_PROMPT" ok "$TMP/r_m1.json"
|
||||
code="$(post "$PORT" "$TMP/m2.json" "$TMP/r_m2.json")"
|
||||
check "mission leg2: part-2 tool call (step indexed by assistant count)" mk chk "$TMP/r_m2.json" \
|
||||
'r["choices"][0]["finish_reason"]=="tool_calls" and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-2.md"'
|
||||
mk chain "$TMP/m3.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json"
|
||||
code="$(post "$PORT" "$TMP/m3.json" "$TMP/r_m3.json")"
|
||||
check "mission leg3: final text, finish stop" mk chk "$TMP/r_m3.json" \
|
||||
'r["choices"][0]["finish_reason"]=="stop" and "Mission complete" in r["choices"][0]["message"]["content"]'
|
||||
mk chain "$TMP/m4.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json" "$TMP/r_m3.json"
|
||||
code="$(post "$PORT" "$TMP/m4.json" "$TMP/r_m4.json")"
|
||||
check "mission overrun: past-script request -> GATE-SCRIPT-EXHAUSTED" mk chk "$TMP/r_m4.json" \
|
||||
'r["choices"][0]["message"]["content"].startswith("GATE-SCRIPT-EXHAUSTED")'
|
||||
|
||||
# 9. injected API errors (OpenAI error envelope)
|
||||
for want in 400 429 500 503; do
|
||||
case "$want" in
|
||||
400) marker="four hundred";; 429) marker="rate limit";;
|
||||
500) marker="five hundred";; 503) marker="unavailable";;
|
||||
esac
|
||||
mk plain "$TMP/e_$want.json" "oa-gate error $marker: trigger the injected failure."
|
||||
code="$(post "$PORT" "$TMP/e_$want.json" "$TMP/r_e_$want.json")"
|
||||
check "api-error $want: status returned" test "$code" = "$want"
|
||||
check "api-error $want: OpenAI error envelope" mk chk "$TMP/r_e_$want.json" \
|
||||
'isinstance(r["error"]["message"],str) and "gate-injected" in r["error"]["message"] and isinstance(r["error"]["type"],str)'
|
||||
done
|
||||
|
||||
# 10. background (unmatched) request
|
||||
mk plain "$TMP/bg.json" "hello there, just a boot probe with no marker"
|
||||
code="$(post "$PORT" "$TMP/bg.json" "$TMP/r_bg.json")"
|
||||
check "background: unmatched prompt -> benign ok" mk chk "$TMP/r_bg.json" \
|
||||
'r["choices"][0]["message"]["content"]=="ok"'
|
||||
|
||||
# 11. ground-truth log invariants
|
||||
check "ground-truth JSONL log invariants" mk logcheck "$TMP/req.jsonl"
|
||||
|
||||
# 12. production-port refusal
|
||||
rc=0
|
||||
"$PY" stub-openai.py --port 7770 --scenarios scenarios-openai.json \
|
||||
--log "$TMP/never.jsonl" >/dev/null 2>&1 || rc=$?
|
||||
check "refuses production port 7770" test "$rc" -ne 0
|
||||
|
||||
# ---- hostile mode: black-hole ----------------------------------------------
|
||||
BH="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$BH" --log "$TMP/bh.jsonl" --mode black-hole \
|
||||
>/dev/null 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "black-hole: healthy" waithealth "$BH"
|
||||
rc=0
|
||||
curl -s -o /dev/null --max-time 3 -H 'content-type: application/json' \
|
||||
--data-binary @"$TMP/plain.json" \
|
||||
"http://127.0.0.1:$BH/v1/chat/completions" || rc=$?
|
||||
check "black-hole: client times out (curl rc 28)" test "$rc" -eq 28
|
||||
check "black-hole: health still answers during the hang" \
|
||||
curl -sf --max-time 2 "http://127.0.0.1:$BH/gate/health"
|
||||
|
||||
# ---- hostile mode: mid-body-drop -------------------------------------------
|
||||
MD="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$MD" --log "$TMP/md.jsonl" --mode mid-body-drop \
|
||||
>/dev/null 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "mid-body-drop: healthy" waithealth "$MD"
|
||||
rc=0
|
||||
curl -s --max-time 5 -o "$TMP/half.json" -H 'content-type: application/json' \
|
||||
--data-binary @"$TMP/plain.json" \
|
||||
"http://127.0.0.1:$MD/v1/chat/completions" || rc=$?
|
||||
check "mid-body-drop: transfer fails (curl rc $rc)" test "$rc" -ne 0
|
||||
check "mid-body-drop: partial body is not parseable JSON" mk notjson "$TMP/half.json"
|
||||
|
||||
# ---- hostile mode: tool-pending-forever ------------------------------------
|
||||
TP="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$TP" --log "$TMP/tp.jsonl" \
|
||||
--mode tool-pending-forever >/dev/null 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "tool-pending-forever: healthy" waithealth "$TP"
|
||||
for i in 1 2 3; do
|
||||
code="$(post "$TP" "$TMP/plain.json" "$TMP/r_tp$i.json")"
|
||||
check "tool-pending-forever: request $i -> 200" test "$code" = "200"
|
||||
done
|
||||
check "tool-pending-forever: three FRESH tool_calls, distinct ids" \
|
||||
mk pending "$TMP/r_tp1.json" "$TMP/r_tp2.json" "$TMP/r_tp3.json"
|
||||
check "tool-pending-forever: /gate/stats counts 3 chat hits" \
|
||||
bash -c "curl -sf http://127.0.0.1:$TP/gate/stats | grep -q '\"chat_hits\": 3'"
|
||||
|
||||
# ---- summary ----------------------------------------------------------------
|
||||
echo
|
||||
echo "selftest: $PASS passed, $FAIL failed"
|
||||
if [ "$FAIL" -ne 0 ]; then
|
||||
echo "SELFTEST RED"
|
||||
exit 1
|
||||
fi
|
||||
echo "SELFTEST GREEN (stub-openai gate scaffolding verified)"
|
||||
Executable
+652
@@ -0,0 +1,652 @@
|
||||
#!/usr/bin/env python3
|
||||
"""stub-openai.py - deterministic local stand-in for an OpenAI-format
|
||||
/v1/chat/completions provider, for the soul-openai-tools-v2 gate
|
||||
(docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md). No API key, no network,
|
||||
no model.
|
||||
|
||||
Sibling of gate9's stub-llm.py (Anthropic dialect, _wt-beta-round9/scripts/
|
||||
gate9/): same scenario mechanism (marker matching, assistant-count step
|
||||
indexing, ground-truth JSONL log, prod-port refusal), different wire.
|
||||
Staging home is tests/gate-openai/ in _wt-openai-tools; folds into
|
||||
scripts/gate9/ after round 9 merges (see README.md).
|
||||
|
||||
WHAT IT DOES
|
||||
* Serves POST /v1/chat/completions on 127.0.0.1 only (OpenAI dialect).
|
||||
* VALIDATES every request - this is the gate's discriminator, built
|
||||
BEFORE the brain-side El code exists so dialect leakage fails loudly:
|
||||
- Anthropic tells are 400 code=gate_dialect_leak: `anthropic-version`
|
||||
header; top-level `system` / `stop_sequences` / `max_tokens_to_sample`
|
||||
/ `anthropic_version`; `input_schema` inside a tool entry; Anthropic
|
||||
content blocks (tool_use / tool_result / server_tool_use / ...).
|
||||
- tools[] must be OpenAI-shaped {type:"function", function:{name,
|
||||
description, parameters}} with unique names -> 400 gate_tools_shape.
|
||||
- assistant tool_calls echoes must be {id, type:"function",
|
||||
function:{name, arguments:<JSON-encoded STRING>}}; a decoded-object
|
||||
`arguments` is a wire bug -> 400 gate_tool_call_shape.
|
||||
- every assistant tool_calls turn must be answered by role:"tool"
|
||||
messages covering EVERY tool_call_id, immediately following;
|
||||
unknown / duplicate / missing ids -> 400 gate_pairing.
|
||||
- echoed `arguments` for gate-issued call ids (call_gate_*) are
|
||||
recomputed from the script and compared after ONE json decode ->
|
||||
400 gate_echo_mismatch. This is the two-escaper-trap discriminator
|
||||
named in the spec's security model (section 6).
|
||||
- scenario-level request expectations from scenarios-openai.json
|
||||
(tools offered, OpenAI-shaped tool_choice, parallel_tool_calls
|
||||
pinned false per ADR-0005) -> 400 gate_expect.
|
||||
* Answers with SCRIPTED responses: plain text (finish_reason "stop"),
|
||||
tool calls (finish_reason "tool_calls", arguments JSON-encoded, incl. a
|
||||
nested-quote/escaping torture payload and a parallel two-call case), and
|
||||
API-error injection (OpenAI error envelope). Scenario is selected by
|
||||
scanning user-message text (newest first) for a registered marker
|
||||
substring; the step index is the number of assistant messages already in
|
||||
the request (stateless replay - resumes index correctly by construction).
|
||||
* Writes a ground-truth JSONL log (--log): one record per request with the
|
||||
validation verdict, matched scenario/step, and exactly which tool calls
|
||||
were delivered. Gate assertions compare the brain's claims against THIS
|
||||
log - truth, not narration.
|
||||
* Unmatched requests (boot probes, awareness chatter) get a benign "ok"
|
||||
text response, logged kind=background, never counted as ground truth.
|
||||
* HOSTILE MODES (--mode) on the same file:
|
||||
black-hole accept + read the request, never respond;
|
||||
mid-body-drop send half a JSON body, then abort the socket;
|
||||
tool-pending-forever every request gets a FRESH tool_call
|
||||
(finish_reason "tool_calls"), forever - tests
|
||||
the agentic loop's iteration cap; count the
|
||||
brain's round-trips via GET /gate/stats.
|
||||
|
||||
usage: stub-openai.py --port P --scenarios scenarios-openai.json \
|
||||
--log requests.jsonl [--mode MODE]
|
||||
Listens on 127.0.0.1 only. Refuses production ports 7770/7779/17779.
|
||||
"""
|
||||
import argparse
|
||||
import itertools
|
||||
import json
|
||||
import socket
|
||||
import struct
|
||||
import threading
|
||||
import time
|
||||
import uuid
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
STATE = {"scenarios": None, "log_path": None, "lock": threading.Lock(),
|
||||
"seq": 0, "mode": "normal", "chat_hits": 0}
|
||||
_PENDING_SEQ = itertools.count(1)
|
||||
|
||||
ANTHROPIC_TOP_KEYS = ("system", "stop_sequences", "max_tokens_to_sample",
|
||||
"anthropic_version")
|
||||
ANTHROPIC_BLOCK_TYPES = {"tool_use", "tool_result", "server_tool_use",
|
||||
"web_search_tool_result", "thinking",
|
||||
"redacted_thinking"}
|
||||
DEFAULT_EXPECT = {"require_tools": True, "require_tool_choice": True,
|
||||
"parallel_tool_calls": False, "forbid_tools": False}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- loading ----
|
||||
def load_scenarios(path):
|
||||
cfg = json.load(open(path))
|
||||
defaults = dict(DEFAULT_EXPECT)
|
||||
defaults.update(cfg.get("defaults", {}).get("expect_request", {}))
|
||||
marker_map = [] # (marker_lower, cname, pid)
|
||||
scripts = {} # cname or cname/pid -> expanded script
|
||||
pid_map = {} # pid -> cname (for call_gate_* id -> script lookup)
|
||||
expects = {} # cname -> merged expect_request
|
||||
for cname, cls in cfg["classes"].items():
|
||||
scripts[cname] = expand_script(cls.get("script", []))
|
||||
exp = dict(defaults)
|
||||
exp.update(cls.get("expect_request", {}))
|
||||
expects[cname] = exp
|
||||
for ph in cls["phrasings"]:
|
||||
if ph.get("script") is not None:
|
||||
scripts[cname + "/" + ph["id"]] = expand_script(ph["script"])
|
||||
marker_map.append((ph["marker"].lower(), cname, ph["id"]))
|
||||
pid_map[ph["id"]] = cname
|
||||
return {"cfg": cfg, "marker_map": marker_map, "scripts": scripts,
|
||||
"pid_map": pid_map, "expects": expects}
|
||||
|
||||
|
||||
def expand_script(script):
|
||||
"""Same repeat-expansion contract as gate9's stub-llm.py ({N}/{NN})."""
|
||||
out = []
|
||||
for step in script:
|
||||
if "repeat" in step:
|
||||
for n in range(1, step["repeat"] + 1):
|
||||
t = {k: v for k, v in step.items() if k != "repeat"}
|
||||
out.append(json.loads(json.dumps(t)
|
||||
.replace("{NN}", "%02d" % n)
|
||||
.replace("{N}", str(n))))
|
||||
else:
|
||||
out.append(step)
|
||||
return out
|
||||
|
||||
|
||||
# ------------------------------------------------------------- validation ----
|
||||
def _rej(message, code):
|
||||
return {"status": 400, "message": message, "code": code}
|
||||
|
||||
|
||||
def validate_dialect(headers, req):
|
||||
"""Universal checks - run on EVERY request, scenario-matched or not.
|
||||
Anything Anthropic-shaped on this lane means the brain's translator
|
||||
leaked; the whole point is that it fails loudly, here, with a reason."""
|
||||
if headers.get("anthropic-version"):
|
||||
return _rej("anthropic-version header on the OpenAI lane: this "
|
||||
"request was built by the Anthropic dialect path",
|
||||
"gate_dialect_leak")
|
||||
for k in ANTHROPIC_TOP_KEYS:
|
||||
if k in req:
|
||||
return _rej("top-level `%s` is Anthropic dialect; the OpenAI "
|
||||
"dialect has no such field (system prompt goes in "
|
||||
"messages[0])" % k, "gate_dialect_leak")
|
||||
tools = req.get("tools")
|
||||
if tools is not None:
|
||||
if not isinstance(tools, list):
|
||||
return _rej("`tools` must be an array", "gate_tools_shape")
|
||||
names = []
|
||||
for i, t in enumerate(tools):
|
||||
if not isinstance(t, dict):
|
||||
return _rej("tools[%d] is not an object" % i,
|
||||
"gate_tools_shape")
|
||||
if "input_schema" in t or (isinstance(t.get("function"), dict)
|
||||
and "input_schema" in t["function"]):
|
||||
return _rej("tools[%d] carries `input_schema` (Anthropic "
|
||||
"dialect); OpenAI dialect wants "
|
||||
"function.parameters" % i, "gate_dialect_leak")
|
||||
if t.get("type") != "function":
|
||||
return _rej("tools[%d].type must be \"function\", got %r"
|
||||
% (i, t.get("type")), "gate_tools_shape")
|
||||
fn = t.get("function")
|
||||
if not isinstance(fn, dict):
|
||||
return _rej("tools[%d].function missing" % i,
|
||||
"gate_tools_shape")
|
||||
if not isinstance(fn.get("name"), str) or not fn["name"]:
|
||||
return _rej("tools[%d].function.name missing/empty" % i,
|
||||
"gate_tools_shape")
|
||||
if not isinstance(fn.get("description"), str) or not fn["description"]:
|
||||
return _rej("tools[%d].function.description missing/empty" % i,
|
||||
"gate_tools_shape")
|
||||
if not isinstance(fn.get("parameters"), dict):
|
||||
return _rej("tools[%d].function.parameters missing (JSON "
|
||||
"Schema object expected)" % i, "gate_tools_shape")
|
||||
names.append(fn["name"])
|
||||
if len(names) != len(set(names)):
|
||||
return _rej("tools: tool names must be unique", "gate_tools_shape")
|
||||
msgs = req.get("messages")
|
||||
if not isinstance(msgs, list) or not msgs:
|
||||
return _rej("`messages` must be a non-empty array",
|
||||
"gate_messages_shape")
|
||||
for i, m in enumerate(msgs):
|
||||
if not isinstance(m, dict):
|
||||
return _rej("messages[%d] is not an object" % i,
|
||||
"gate_messages_shape")
|
||||
c = m.get("content")
|
||||
if isinstance(c, list):
|
||||
for j, b in enumerate(c):
|
||||
if isinstance(b, dict) and b.get("type") in ANTHROPIC_BLOCK_TYPES:
|
||||
return _rej("messages[%d].content[%d] is an Anthropic "
|
||||
"`%s` block; the OpenAI dialect uses "
|
||||
"tool_calls / role:\"tool\" messages"
|
||||
% (i, j, b.get("type")), "gate_dialect_leak")
|
||||
if m.get("role") == "tool":
|
||||
if not isinstance(m.get("tool_call_id"), str) or not m["tool_call_id"]:
|
||||
return _rej("messages[%d]: role \"tool\" requires a "
|
||||
"`tool_call_id`" % i, "gate_messages_shape")
|
||||
if "content" not in m:
|
||||
return _rej("messages[%d]: role \"tool\" requires `content`"
|
||||
% i, "gate_messages_shape")
|
||||
if m.get("role") == "assistant" and m.get("tool_calls") is not None:
|
||||
tcs = m["tool_calls"]
|
||||
if not isinstance(tcs, list) or not tcs:
|
||||
return _rej("messages[%d].tool_calls must be a non-empty "
|
||||
"array" % i, "gate_tool_call_shape")
|
||||
for j, tc in enumerate(tcs):
|
||||
if not isinstance(tc, dict) or tc.get("type") != "function":
|
||||
return _rej("messages[%d].tool_calls[%d].type must be "
|
||||
"\"function\"" % (i, j), "gate_tool_call_shape")
|
||||
if not isinstance(tc.get("id"), str) or not tc["id"]:
|
||||
return _rej("messages[%d].tool_calls[%d].id missing"
|
||||
% (i, j), "gate_tool_call_shape")
|
||||
fn = tc.get("function")
|
||||
if not isinstance(fn, dict) or not isinstance(fn.get("name"), str):
|
||||
return _rej("messages[%d].tool_calls[%d].function.name "
|
||||
"missing" % (i, j), "gate_tool_call_shape")
|
||||
if not isinstance(fn.get("arguments"), str):
|
||||
return _rej("messages[%d].tool_calls[%d].function."
|
||||
"arguments must be a JSON-encoded STRING, "
|
||||
"got %s" % (i, j,
|
||||
type(fn.get("arguments")).__name__),
|
||||
"gate_tool_call_shape")
|
||||
return None
|
||||
|
||||
|
||||
def validate_pairing(msgs):
|
||||
"""OpenAI pairing rule: every assistant tool_calls turn must be followed
|
||||
immediately by role:"tool" messages answering every tool_call_id."""
|
||||
open_ids, open_at = set(), None
|
||||
for i, m in enumerate(msgs):
|
||||
role = m.get("role")
|
||||
if role == "tool":
|
||||
tid = m.get("tool_call_id")
|
||||
if open_at is None:
|
||||
return _rej("messages[%d]: role \"tool\" message with no "
|
||||
"preceding assistant tool_calls turn "
|
||||
"(tool_call_id=%s)" % (i, tid), "gate_pairing")
|
||||
if tid not in open_ids:
|
||||
return _rej("messages[%d]: tool message answers unknown or "
|
||||
"already-answered tool_call_id %s" % (i, tid),
|
||||
"gate_pairing")
|
||||
open_ids.discard(tid)
|
||||
continue
|
||||
if open_ids:
|
||||
return _rej("messages[%d]: assistant tool_calls not fully "
|
||||
"answered before messages[%d]; missing tool "
|
||||
"responses for: %s" % (open_at, i, sorted(open_ids)),
|
||||
"gate_pairing")
|
||||
open_ids, open_at = set(), None
|
||||
if role == "assistant" and m.get("tool_calls"):
|
||||
ids = [tc.get("id") for tc in m["tool_calls"]]
|
||||
open_ids, open_at = set(ids), i
|
||||
if open_ids:
|
||||
return _rej("messages[%d]: assistant tool_calls at end of thread "
|
||||
"without tool responses for: %s"
|
||||
% (open_at, sorted(open_ids)), "gate_pairing")
|
||||
return None
|
||||
|
||||
|
||||
def validate_echo_args(msgs, loaded):
|
||||
"""Ground-truth round-trip check: for every echoed gate-issued call id,
|
||||
recompute the arguments this stub originally sent from the script and
|
||||
require one json decode to reproduce them exactly. Catches the
|
||||
two-escaper trap (spec section 6) deterministically."""
|
||||
if not loaded:
|
||||
return None
|
||||
for i, m in enumerate(msgs):
|
||||
if m.get("role") != "assistant":
|
||||
continue
|
||||
for tc in m.get("tool_calls") or []:
|
||||
tid = tc.get("id", "")
|
||||
if not tid.startswith("call_gate_"):
|
||||
continue
|
||||
rest = tid[len("call_gate_"):]
|
||||
try:
|
||||
pid, s_part, k_part = rest.rsplit("_", 2)
|
||||
step_idx, k = int(s_part[1:]), int(k_part)
|
||||
except (ValueError, IndexError):
|
||||
continue
|
||||
cname = loaded["pid_map"].get(pid)
|
||||
if cname is None:
|
||||
continue
|
||||
script = (loaded["scripts"].get(cname + "/" + pid)
|
||||
or loaded["scripts"].get(cname) or [])
|
||||
if step_idx >= len(script):
|
||||
continue
|
||||
calls = script[step_idx].get("tool_calls") or []
|
||||
if k >= len(calls):
|
||||
continue
|
||||
expected = calls[k]
|
||||
fn = tc.get("function") or {}
|
||||
if fn.get("name") != expected["name"]:
|
||||
return _rej("messages[%d]: echoed tool name %r != issued %r "
|
||||
"for %s" % (i, fn.get("name"), expected["name"],
|
||||
tid), "gate_echo_mismatch")
|
||||
try:
|
||||
got = json.loads(fn.get("arguments", ""))
|
||||
except ValueError:
|
||||
return _rej("messages[%d]: echoed arguments for %s are not "
|
||||
"valid JSON after one decode (truncated or "
|
||||
"half-escaped?)" % (i, tid), "gate_echo_mismatch")
|
||||
if got != expected["arguments"]:
|
||||
hint = (" (decoded to a string, not an object: "
|
||||
"double-encoded - the two-escaper trap)"
|
||||
if isinstance(got, str) else "")
|
||||
return _rej("messages[%d]: echoed arguments for %s do not "
|
||||
"round-trip to the issued payload%s"
|
||||
% (i, tid, hint), "gate_echo_mismatch")
|
||||
return None
|
||||
|
||||
|
||||
def validate_expect(req, exp):
|
||||
"""Scenario-level request expectations (scenarios-openai.json)."""
|
||||
tools = req.get("tools") or []
|
||||
if exp.get("forbid_tools") and tools:
|
||||
return _rej("this scenario is chat-only: no `tools` may be offered "
|
||||
"on it", "gate_expect")
|
||||
if exp.get("require_tools") and not tools:
|
||||
return _rej("scenario expects a `tools` array to be offered (the "
|
||||
"agentic lane must advertise its tools)", "gate_expect")
|
||||
if exp.get("require_tool_choice"):
|
||||
tc = req.get("tool_choice")
|
||||
ok = tc in ("auto", "none", "required") or (
|
||||
isinstance(tc, dict) and tc.get("type") == "function"
|
||||
and isinstance(tc.get("function"), dict)
|
||||
and tc["function"].get("name"))
|
||||
if not ok:
|
||||
return _rej("scenario expects an OpenAI-shaped `tool_choice`, "
|
||||
"got %r" % (tc,), "gate_expect")
|
||||
want_ptc = exp.get("parallel_tool_calls", None)
|
||||
if want_ptc is not None:
|
||||
if "parallel_tool_calls" not in req:
|
||||
return _rej("scenario expects explicit `parallel_tool_calls` "
|
||||
"(ADR-0005: must be pinned false on the wire)",
|
||||
"gate_expect")
|
||||
if req["parallel_tool_calls"] != want_ptc:
|
||||
return _rej("scenario expects parallel_tool_calls=%s, got %s"
|
||||
% (json.dumps(want_ptc),
|
||||
json.dumps(req["parallel_tool_calls"])),
|
||||
"gate_expect")
|
||||
return None
|
||||
|
||||
|
||||
# --------------------------------------------------------- scenario match ----
|
||||
def extract_user_texts_newest_first(msgs):
|
||||
texts = []
|
||||
for m in reversed(msgs):
|
||||
if not isinstance(m, dict) or m.get("role") != "user":
|
||||
continue
|
||||
c = m.get("content")
|
||||
if isinstance(c, str):
|
||||
texts.append(c)
|
||||
elif isinstance(c, list):
|
||||
for b in c:
|
||||
if isinstance(b, dict) and b.get("type") == "text":
|
||||
texts.append(b.get("text", ""))
|
||||
return texts
|
||||
|
||||
|
||||
def match_scenario(loaded, msgs):
|
||||
for text in extract_user_texts_newest_first(msgs):
|
||||
tl = text.lower()
|
||||
for marker, cname, pid in loaded["marker_map"]:
|
||||
if marker in tl:
|
||||
return cname, pid
|
||||
return None, None
|
||||
|
||||
|
||||
# ------------------------------------------------------------- rendering ----
|
||||
def completion_envelope(msg, finish, model, usage=(100, 100)):
|
||||
return {"id": "chatcmpl-gate-" + uuid.uuid4().hex[:12],
|
||||
"object": "chat.completion", "created": int(time.time()),
|
||||
"model": model,
|
||||
"choices": [{"index": 0, "message": msg,
|
||||
"finish_reason": finish, "logprobs": None}],
|
||||
"usage": {"prompt_tokens": usage[0],
|
||||
"completion_tokens": usage[1],
|
||||
"total_tokens": usage[0] + usage[1]}}
|
||||
|
||||
|
||||
def text_completion(text, model):
|
||||
return completion_envelope({"role": "assistant", "content": text},
|
||||
"stop", model, usage=(1, 1))
|
||||
|
||||
|
||||
def pending_body(seq, model):
|
||||
args = json.dumps({"path": "never-%04d.md" % seq,
|
||||
"content": "this run never completes"})
|
||||
msg = {"role": "assistant", "content": None,
|
||||
"tool_calls": [{"id": "call_hostile_pending_%04d" % seq,
|
||||
"type": "function",
|
||||
"function": {"name": "write_file",
|
||||
"arguments": args}}]}
|
||||
return completion_envelope(msg, "tool_calls", model, usage=(1, 1))
|
||||
|
||||
|
||||
def render_step(step, cname, pid, step_idx, model):
|
||||
"""Returns (http_status, body_dict, delivered) - delivered is ground
|
||||
truth for the JSONL log."""
|
||||
delivered = {"tool_calls": [], "finish_reason": None, "api_error": None}
|
||||
if "api_error" in step:
|
||||
e = step["api_error"]
|
||||
delivered["api_error"] = e["status"]
|
||||
return (e["status"],
|
||||
{"error": {"message": e["message"],
|
||||
"type": e.get("type", "server_error"),
|
||||
"param": None, "code": e.get("code")}},
|
||||
delivered)
|
||||
msg = {"role": "assistant"}
|
||||
finish = "stop"
|
||||
if step.get("tool_calls"):
|
||||
tcs = []
|
||||
for k, call in enumerate(step["tool_calls"]):
|
||||
tid = "call_gate_%s_s%d_%d" % (pid, step_idx, k)
|
||||
tcs.append({"id": tid, "type": "function",
|
||||
"function": {"name": call["name"],
|
||||
"arguments": json.dumps(
|
||||
call["arguments"],
|
||||
ensure_ascii=False)}})
|
||||
delivered["tool_calls"].append(call["name"])
|
||||
msg["tool_calls"] = tcs
|
||||
msg["content"] = step.get("text") # null when no narration, like real
|
||||
finish = "tool_calls"
|
||||
else:
|
||||
msg["content"] = step["text"]
|
||||
delivered["finish_reason"] = finish
|
||||
return 200, completion_envelope(msg, finish, model), delivered
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ log ------
|
||||
def log_record(rec):
|
||||
with STATE["lock"]:
|
||||
STATE["seq"] += 1
|
||||
rec["seq"] = STATE["seq"]
|
||||
with open(STATE["log_path"], "a") as f:
|
||||
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- server -----
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
protocol_version = "HTTP/1.1"
|
||||
|
||||
def _send_json(self, status, obj):
|
||||
body = json.dumps(obj, ensure_ascii=False).encode("utf-8")
|
||||
self.send_response(status)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def _send_error(self, verdict):
|
||||
self._send_json(verdict["status"],
|
||||
{"error": {"message": verdict["message"],
|
||||
"type": "invalid_request_error",
|
||||
"param": None, "code": verdict["code"]}})
|
||||
|
||||
def _drop_mid_body(self):
|
||||
"""Valid 200 headers, half the promised body, then a socket abort
|
||||
(same SO_LINGER teardown as gate9's mid-body-drop-brain.py)."""
|
||||
full = json.dumps(text_completion(
|
||||
"This reply will never finish arriving because the connection "
|
||||
"dies in the middle of the body, which is exactly the point of "
|
||||
"this hostile fixture.", "hostile-mid-drop")).encode("utf-8")
|
||||
half = full[: len(full) // 2]
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(full))) # promises more
|
||||
self.end_headers()
|
||||
self.wfile.write(half)
|
||||
self.wfile.flush()
|
||||
try:
|
||||
self.connection.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER,
|
||||
struct.pack("ii", 1, 0))
|
||||
self.connection.shutdown(socket.SHUT_RDWR)
|
||||
except OSError:
|
||||
pass
|
||||
self.close_connection = True
|
||||
|
||||
def do_GET(self):
|
||||
path = self.path.split("?")[0]
|
||||
if path == "/gate/health":
|
||||
self._send_json(200, {"ok": True, "mode": STATE["mode"]})
|
||||
elif path == "/gate/stats":
|
||||
with STATE["lock"]:
|
||||
self._send_json(200, {"mode": STATE["mode"],
|
||||
"chat_hits": STATE["chat_hits"]})
|
||||
else:
|
||||
self._send_json(404, {"error": {"message": "not found",
|
||||
"type": "invalid_request_error",
|
||||
"param": None,
|
||||
"code": "unknown_route"}})
|
||||
|
||||
def do_POST(self):
|
||||
n = int(self.headers.get("Content-Length") or 0)
|
||||
raw = self.rfile.read(n)
|
||||
mode = STATE["mode"]
|
||||
rec = {"ts": time.time(), "path": self.path, "mode": mode,
|
||||
"kind": "background", "scenario_class": None, "phrasing": None,
|
||||
"step": None, "n_messages": 0, "n_assistant": 0,
|
||||
"validation": "ok", "validation_detail": None,
|
||||
"delivered": {"tool_calls": [], "finish_reason": None,
|
||||
"api_error": None},
|
||||
"http_status": 200}
|
||||
if self.path.split("?")[0] != "/v1/chat/completions":
|
||||
rec.update(kind="wrong_path", http_status=404)
|
||||
log_record(rec)
|
||||
self._send_json(404, {"error": {
|
||||
"message": "no such route: %s" % self.path,
|
||||
"type": "invalid_request_error", "param": None,
|
||||
"code": "unknown_route"}})
|
||||
return
|
||||
with STATE["lock"]:
|
||||
STATE["chat_hits"] += 1
|
||||
|
||||
# ---- hostile modes: behavior first, no validation ----------------
|
||||
if mode == "black-hole":
|
||||
rec.update(kind="hostile", http_status=None)
|
||||
log_record(rec)
|
||||
threading.Event().wait() # hold the socket open forever
|
||||
return
|
||||
if mode == "mid-body-drop":
|
||||
rec.update(kind="hostile", http_status=200)
|
||||
log_record(rec)
|
||||
self._drop_mid_body()
|
||||
return
|
||||
if mode == "tool-pending-forever":
|
||||
seq = next(_PENDING_SEQ)
|
||||
rec.update(kind="hostile",
|
||||
delivered={"tool_calls": ["write_file"],
|
||||
"finish_reason": "tool_calls",
|
||||
"api_error": None})
|
||||
log_record(rec)
|
||||
self._send_json(200, pending_body(seq, "gate-openai-model"))
|
||||
return
|
||||
|
||||
# ---- normal mode -------------------------------------------------
|
||||
try:
|
||||
req = json.loads(raw)
|
||||
except ValueError as exc:
|
||||
# DIAGNOSTIC CAPTURE (2026-08-06): an unparseable body used to be recorded as
|
||||
# a bare "bad_json" with the bytes thrown away, which made an intermittent
|
||||
# failure impossible to root-cause — you cannot fix what you did not keep.
|
||||
# Dump the raw body next to the log, and record exactly where the parser gave
|
||||
# up plus the offending byte, so one occurrence is enough to diagnose.
|
||||
dump_path = "%s.badbody.%s" % (STATE.get("log_path", "/tmp/stub-openai"),
|
||||
rec.get("seq", "x"))
|
||||
try:
|
||||
data = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
|
||||
with open(dump_path, "wb") as fh:
|
||||
fh.write(data)
|
||||
except Exception as dump_exc:
|
||||
dump_path = "(dump failed: %s)" % dump_exc
|
||||
pos = getattr(exc, "pos", None)
|
||||
near = ""
|
||||
byte_repr = ""
|
||||
if isinstance(pos, int):
|
||||
blob = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
|
||||
near = blob[max(0, pos - 60):pos + 60].decode("utf-8", "replace")
|
||||
if 0 <= pos < len(blob):
|
||||
byte_repr = "0x%02x" % blob[pos]
|
||||
rec.update(kind="bad_json", validation="rejected",
|
||||
validation_detail="request body is not valid JSON: %s" % exc,
|
||||
http_status=400, raw_len=len(raw), raw_dump=dump_path,
|
||||
err_pos=pos, err_byte=byte_repr, err_near=near)
|
||||
log_record(rec)
|
||||
self._send_error(_rej("request body is not valid JSON",
|
||||
"bad_json"))
|
||||
return
|
||||
msgs = req.get("messages") or []
|
||||
rec["n_messages"] = len(msgs)
|
||||
rec["n_assistant"] = sum(1 for m in msgs if isinstance(m, dict)
|
||||
and m.get("role") == "assistant")
|
||||
loaded = STATE["scenarios"]
|
||||
cname, pid = match_scenario(loaded, msgs)
|
||||
if cname:
|
||||
rec.update(kind="scenario", scenario_class=cname, phrasing=pid)
|
||||
|
||||
# Wire-level validation runs for EVERY request, scenario or not.
|
||||
verdict = (validate_dialect(self.headers, req)
|
||||
or validate_pairing([m for m in msgs
|
||||
if isinstance(m, dict)])
|
||||
or validate_echo_args(msgs, loaded))
|
||||
if verdict:
|
||||
rec.update(validation="rejected",
|
||||
validation_detail=verdict["message"],
|
||||
http_status=verdict["status"])
|
||||
log_record(rec)
|
||||
self._send_error(verdict)
|
||||
return
|
||||
|
||||
model = req.get("model", "gate-openai-model")
|
||||
if not cname:
|
||||
log_record(rec)
|
||||
self._send_json(200, text_completion("ok", model))
|
||||
return
|
||||
|
||||
script = (loaded["scripts"].get(cname + "/" + pid)
|
||||
or loaded["scripts"][cname])
|
||||
step_idx = rec["n_assistant"]
|
||||
if step_idx >= len(script):
|
||||
rec.update(kind="overrun", step=step_idx)
|
||||
log_record(rec)
|
||||
self._send_json(200, text_completion(
|
||||
"GATE-SCRIPT-EXHAUSTED %s step %d" % (pid, step_idx), model))
|
||||
return
|
||||
|
||||
step = script[step_idx]
|
||||
exp = dict(loaded["expects"][cname])
|
||||
exp.update(step.get("expect_request", {}))
|
||||
verdict = validate_expect(req, exp)
|
||||
if verdict:
|
||||
rec.update(step=step_idx, validation="rejected",
|
||||
validation_detail=verdict["message"],
|
||||
http_status=verdict["status"])
|
||||
log_record(rec)
|
||||
self._send_error(verdict)
|
||||
return
|
||||
|
||||
status, body, delivered = render_step(step, cname, pid, step_idx,
|
||||
model)
|
||||
rec.update(step=step_idx, delivered=delivered, http_status=status)
|
||||
log_record(rec)
|
||||
self._send_json(status, body)
|
||||
|
||||
def log_message(self, *a):
|
||||
pass
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--port", type=int, required=True)
|
||||
ap.add_argument("--scenarios",
|
||||
help="scenarios-openai.json (required in normal mode)")
|
||||
ap.add_argument("--log", required=True)
|
||||
ap.add_argument("--mode", default="normal",
|
||||
choices=["normal", "black-hole", "mid-body-drop",
|
||||
"tool-pending-forever"])
|
||||
args = ap.parse_args()
|
||||
if args.port in (7770, 7779, 17779):
|
||||
raise SystemExit("stub-openai: refusing production Neuron port")
|
||||
if args.mode == "normal" and not args.scenarios:
|
||||
raise SystemExit("stub-openai: --scenarios is required in normal mode")
|
||||
STATE["mode"] = args.mode
|
||||
STATE["scenarios"] = (load_scenarios(args.scenarios)
|
||||
if args.scenarios else None)
|
||||
STATE["log_path"] = args.log
|
||||
open(args.log, "w").close()
|
||||
n_markers = (len(STATE["scenarios"]["marker_map"])
|
||||
if STATE["scenarios"] else 0)
|
||||
print("stub-openai [%s]: 127.0.0.1:%d /v1/chat/completions "
|
||||
"(%d markers registered, log=%s)"
|
||||
% (args.mode, args.port, n_markers, args.log), flush=True)
|
||||
ThreadingHTTPServer(("127.0.0.1", args.port), Handler).serve_forever()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user