Compare commits
9 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| de65991807 | |||
| dba755dcec | |||
| 8f3a478771 | |||
| 9ea41eed78 | |||
| ff421d39f6 | |||
| 635f6febe4 | |||
| 62af5649fe | |||
| 710761e2d5 | |||
| 74520b8333 |
+139
@@ -0,0 +1,139 @@
|
||||
# PORT-NOTES — openai tools port working state (2026-08-06, session handoff-safe)
|
||||
|
||||
Spec: `docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md` (Tim-approved 2026-08-06). Tasks #1-5
|
||||
tracked in-session (1 ✓ wiring verdict, 2 ✓ stub rig, 3 in-progress = THIS, 4-5 pending).
|
||||
Worktree: HERE (`_wt-openai-tools`, branch `feat/soul-openai-tools-v2` @ dba755d). Round-9 trees
|
||||
READ-ONLY. Nothing committed yet.
|
||||
|
||||
## Step-0 verdict (evidence in journal note ncli-653ba964dd76)
|
||||
Shipped app never wires the v1 lane: launcher exports `SOUL_LLM_MODEL/PROVIDER/BASE_URL` +
|
||||
`ANTHROPIC_API_KEY`+`SOUL_API_KEY` (= Keychain key for WHATEVER provider; installer/macos/
|
||||
neuron-daemons.sh:288-300 on hotfix/beta-round9); brain reads only SOUL_LLM_MODEL (chat.el:8) and
|
||||
NEURON_LLM_0_* (chat.el:1768-1794) which nothing sets. `/api/config` PATCH ignores llm_* fields
|
||||
(studio.el:36 handle_config: POST-only, reads model/provider/api_key only).
|
||||
**Bridge = brain-side ONLY (zero app-repo edits, zero round-9 collision):**
|
||||
- `llm_base_url()`: NEURON_LLM_0_URL → fallback SOUL_LLM_BASE_URL when SOUL_LLM_PROVIDER ∉ {"","anthropic"}
|
||||
- `llm_wire_format()`: NEURON_LLM_0_FORMAT → fallback derive from SOUL_LLM_PROVIDER (openai/grok/gemini/groq/ollama → "openai"; else "anthropic")
|
||||
- `agentic_api_key()`: already works (ANTHROPIC_API_KEY carries the provider key); add NEURON_LLM_0_KEY → SOUL_API_KEY fallback.
|
||||
|
||||
## Design pins (stub asserts these — stub is green 58/58, tests/gate-openai/)
|
||||
- Request MUST send `"tool_choice":"auto"` (string) + `"parallel_tool_calls":false` explicitly.
|
||||
- `arguments` in tool_calls = JSON-ENCODED STRING; decode ONCE via json_get → feed dispatch_tool
|
||||
verbatim. Stub's echo-mismatch check catches double-encode/decode (two-escaper trap).
|
||||
- Assistant echo turn: `{"role":"assistant","content":null,"tool_calls":[...]}` VERBATIM from response.
|
||||
- Feedback: `{"role":"tool","tool_call_id":"<id>","content":"<result string>"}`.
|
||||
- Resume must NOT re-answer an answered id (stub 400s on repeat tool_call_id).
|
||||
- Parallel tool_calls in a response: take FIRST only + log skip (mirror ADR-0005 stopgap); stub
|
||||
scenario `parallel` proves behavior.
|
||||
- No tools in request when tools array empty/absent turns (boot probes) — stub defaults tolerate.
|
||||
|
||||
## el idioms confirmed (from openai_chat_complete :1808-1854 + agentic_loop :2751-2838)
|
||||
- JSON: `json_get(s,k)` decoded string · `json_get_raw(s,k)` raw subtree · `json_array_len` ·
|
||||
`json_array_get(arr,i)` · build by string concat + `json_escape()` (:1797, OpenAI-lane escaper).
|
||||
- HTTP: `let h: Map = {}` + `map_set(h,k,v)` + `http_post_with_headers(url, body, h)`;
|
||||
Bearer auth via `Authorization` header when key non-empty (:1825-1830).
|
||||
- Loop-carried vars must be top-level locals in the fn, mutated as if-expressions at while-body
|
||||
top level (see :2760-2791 pattern + comment :2903-2904 region).
|
||||
- Error shape: `str_starts_with(raw,"{\"error\"") || str_contains(raw,"\"error\":")` → return
|
||||
`{"error":"llm unavailable","reply":""}` (:1835-1838).
|
||||
|
||||
## Remaining read map (before writing the fork)
|
||||
- chat.el 2840-3200: block walk (2923-3000), policy gate (3009-3023: classify_tool_risk /
|
||||
is_builtin_tool / ask_all / tool_auto_approved → needs_bridge), dispatch_tool call (3025),
|
||||
tool_result feedback (3031, 3067-3072), run-progress ledger append (3078-3087), bridge_save
|
||||
(3182), loop end + done envelope (~3100-3200).
|
||||
- agentic_resume 3227-3293 (hardcoded Anthropic headers to make wire-aware; blob gets `wire` field,
|
||||
legacy default anthropic) · handle_tool_result 3293+ · dharma fork site 3465 (calls agentic_loop
|
||||
direct, no use_openai check today).
|
||||
|
||||
## Write plan (order)
|
||||
1. Env fallbacks (edit llm_base_url/llm_wire_format/agentic_api_key) — small, first, testable alone.
|
||||
2. `openai_tools_json(anthropic_tools: String) -> String` converter (walk array; per entry build
|
||||
{"type":"function","function":{name,description,parameters:input_schema-raw}}).
|
||||
3. `openai_agentic_loop(...)` fork: same signature as agentic_loop minus Anthropic-only params;
|
||||
INCLUDE run-progress ledger + tools_log + iteration cap 12; NO container_id/ws_drift/web_search
|
||||
(out of scope; strip web_search entry from tools via agentic_tools_literal()+connector merge,
|
||||
NOT _with_web()).
|
||||
4. Fork sites ×3: handle_chat_agentic :2695-2700 (route agentic to new loop when use_openai);
|
||||
dharma :3465; agentic_resume wire-branch.
|
||||
5. `chat.elh` extern decls. 6. Compile (recipe: dist/ + elc/elb per neuron-soul-build-deploy memory;
|
||||
round-9 tree soul.c regen'd 08-06 proves toolchain live). 7. Gate: stub selftest recipe in
|
||||
tests/gate-openai/README.md. 8. Anthropic-lane regression via gate9 (READ-ONLY consume from
|
||||
_wt-beta-round9). 9. Live Groq E2E (scratch profile, free port, key via Keychain read-only).
|
||||
|
||||
## BUILD RECIPE — CORRECTED 2026-08-06 (the June memory is STALE for August code)
|
||||
`~/el-sdk/el_runtime.c` (Jun 15) is MISSING builtins the Aug engine calls (`engram_wm_count`,
|
||||
`engram_wm_top_json`, `http_delete_json`, `http_serve_async`) → link fails with
|
||||
"symbol(s) not found for architecture arm64". Use the REPO-PINNED runtime:
|
||||
```
|
||||
mkdir -p <scratch>
|
||||
elb --elc=$HOME/el-sdk/elc --runtime=vendor/el-runtime/v1.0.0-20260501 --out=<scratch>/
|
||||
# "elb: link failed" at the end is EXPECTED and harmless — the per-module .c files are produced
|
||||
cc -std=c11 -O1 -DHAVE_CURL -rdynamic \
|
||||
-I vendor/el-runtime/v1.0.0-20260501 -I <scratch> -I /opt/homebrew/opt/openssl@3/include \
|
||||
-L /opt/homebrew/opt/openssl@3/lib \
|
||||
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
|
||||
-o <scratch>/soul <scratch>/*.c vendor/el-runtime/v1.0.0-20260501/el_runtime.c \
|
||||
-lssl -lcrypto -lcurl -lpthread -lm
|
||||
```
|
||||
Source: `_engine-plainchat-20260805/README.md:396-412`. Verified today: 0 errors, 887,296 B.
|
||||
`elb` ALSO rewrites every `*.elh` in the tree (cosmetic em-dash→hyphen in the auto-gen banner,
|
||||
plus true-ups) and drops a stray `soul..elh` — `git restore` the unrelated ones and delete the
|
||||
stray before staging, or the diff drowns in noise.
|
||||
|
||||
## SELF-REVIEW FIX LIST (found by reading my own diff, 2026-08-06 — apply in ONE batch, then rebuild once)
|
||||
- **F3 (CORRECTNESS, do first):** the assistant echo currently replays the provider's FULL
|
||||
`tool_calls` array (`tc_arr`) while the loop answers only the FIRST call. If a provider ignores
|
||||
`parallel_tool_calls:false`, the next request carries an assistant turn with N tool_calls and
|
||||
only ONE `role:"tool"` response → most OpenAI-format providers 400 ("missing tool response for
|
||||
id X") and the run dies. This is the same class as ADR-0005's Anthropic failure, but here it is
|
||||
cheap to close: echo ONLY the honored call (`"[" + tc0 + "]"`), so the conversation we send is
|
||||
self-consistent and the dropped call never existed from the model's view. The DRIFT log line
|
||||
stays (honest accounting of what we dropped).
|
||||
- **F4 (efficiency/latency):** `handle_chat_agentic` computes `agentic_tools_all()` at ~:2681
|
||||
BEFORE the fork, then the OpenAI branch computes `agentic_tools_no_web()` again — two
|
||||
`connector_tools_json()` calls per turn, each an HTTP round-trip to the connector bridge on
|
||||
:7771 (two timeout exposures). Fix: compute the tools array ONCE, per lane, after `use_openai`
|
||||
is known (check no other use of `tools_json` sits between :2681 and the fork before moving it).
|
||||
Note: `openai_tools_json()` already skips any entry with no `input_schema`, so Anthropic's
|
||||
server-side `web_search` entry is auto-dropped even if the full array is passed —
|
||||
`agentic_tools_no_web()` is kept for EXPLICITNESS, not necessity.
|
||||
- **F1 (debuggability):** the "no choices in response" branch logs a generic string and discards
|
||||
the body. Log the response head (as the `is_error` branch does) — a provider that returns 200
|
||||
with an unexpected shape is otherwise undiagnosable from the log.
|
||||
- **OPEN QUESTION (evidence pending from the gate):** the tool-result feedback turn escapes with
|
||||
`json_escape()` (this lane's escaper) rather than `json_safe()` (used everywhere else). The
|
||||
Anthropic lane escapes that field with NEITHER, which is a latent defect on that side. If the
|
||||
torture scenario shows any escaping loss, switch to `json_safe` and note the Anthropic-side
|
||||
finding for Will.
|
||||
|
||||
## TEST HARNESS — built 2026-08-06 (Task 4 side-work, reusable by anyone)
|
||||
- `tests/run-el-test.sh <tests/test_x.el> | --all` — the engine tests were NEVER runnable
|
||||
before this (`elc` is a compiler: emits C to stdout and exits). It emits the test to C,
|
||||
compiles `soul.c` separately with `main` renamed away (soul.c owns the daemon's real main
|
||||
but also defines `layered_cycle` et al.), links the remaining modules + the repo-pinned
|
||||
runtime, and executes. Modules cached under `/tmp/el-test-<worktree>/`; `REBUILD=1` forces.
|
||||
- **The runner computes the verdict itself** because the test FILES cannot: all 9 counted
|
||||
test files do `let pass_count = pass_count + 1` inside an if BLOCK, which El scoping
|
||||
discards, so every summary line reads `0 passed, 0 failed` forever. Per-assertion
|
||||
`PASS:`/`FAIL:` lines ARE reliable; the runner counts those, exits non-zero on any FAIL
|
||||
or on zero assertions, and was proven to discriminate with a negative control (broken
|
||||
assertion → 31 passed / 1 failed / exit 1). Real in-file fix filed: **neuron#116**.
|
||||
- `tests/test_bridge_serialization.el`: 4 `bridge_save` calls updated for the new `wire`
|
||||
argument, plus **Section 9** (8 new assertions) covering wire round-trip both ways, the
|
||||
legacy no-wire blob (resumes as anthropic), and a FIELD-ORDER decoy guard — a fake
|
||||
`"wire":"anthropic"` planted inside `messages_raw` must not beat the blob's own scalar.
|
||||
That decoy is the round-9 first-match-scanner bug class, now pinned by a test. **32/32 green.**
|
||||
|
||||
## MEMORY-SAVE CAVEAT RESOLVED 2026-08-06
|
||||
Earlier saves this session reported `-> OUTBOX only (real mind unreachable or read-back
|
||||
failed)`. That was a **read-back verifier false negative, not data loss** — a direct
|
||||
`POST :7770/api/neuron/recall` returns those notes from the live mind verbatim. Another
|
||||
terminal was fixing exactly this (multi-word read-back probe) the same afternoon. Do NOT
|
||||
re-save on an OUTBOX report without first querying the mind directly, or you duplicate nodes.
|
||||
|
||||
## Standing cautions
|
||||
- PERSIST OFF on the real mind this boot (neuron#98/#92): journal saves only, ferry later. MCP link
|
||||
down this terminal; use neuron_remember.py / neuron_recall.py.
|
||||
- Aug-16: Groq retires llama-3.3-70b-versatile (separate P0, Tim's call, catalog swap).
|
||||
- Never bind 7770/7779/17779; never touch ~/.neuron; round-9 worktrees read-only.
|
||||
@@ -1,4 +1,4 @@
|
||||
// auto-generated by elc --emit-header — do not edit
|
||||
// auto-generated by elc --emit-header - do not edit
|
||||
extern fn chat_default_model() -> String
|
||||
extern fn engram_numeric_valid(s: String) -> Bool
|
||||
extern fn parse_float_x100(s: String) -> Int
|
||||
@@ -16,18 +16,35 @@ extern fn engram_nodes_merge(a: String, b: String) -> String
|
||||
extern fn id_in_seen(node_id: String, seen: String) -> Bool
|
||||
extern fn add_to_seen(seen: String, node_id: String) -> String
|
||||
extern fn engram_extract_ids(nodes_json: String) -> String
|
||||
extern fn affective_node_ts(node_json: String) -> Int
|
||||
extern fn engram_compile(intent: String) -> String
|
||||
extern fn distill_transcript(transcript: String) -> String
|
||||
extern fn json_safe(s: String) -> String
|
||||
extern fn current_engine_note(model: String) -> String
|
||||
extern fn bounded_persona_floor() -> String
|
||||
extern fn operator_identity_block() -> String
|
||||
extern fn build_system_prompt(ctx: String, chat_mode: Bool) -> String
|
||||
extern fn hist_append(hist: String, role: String, content: String) -> String
|
||||
extern fn conv_hist_key(session_id: String) -> String
|
||||
extern fn conv_hist_label(session_id: String) -> String
|
||||
extern fn is_utility_request(body: String, session_id: String) -> Bool
|
||||
extern fn provenance_scan_urls(arr: String, acc: String) -> String
|
||||
extern fn provenance_add_sources(block: String, btype: String, has_cit: Bool, cit_raw: String, acc: String) -> String
|
||||
extern fn provenance_names(tools_used: String) -> String
|
||||
extern fn text_join_sep(accumulated: String, incoming: String, after_interruption: Bool) -> String
|
||||
extern fn receipt_rule() -> String
|
||||
extern fn receipt_strip(s: String) -> String
|
||||
extern fn tool_receipt(tools_used: String, sources: String) -> String
|
||||
extern fn hist_trim(hist: String) -> String
|
||||
extern fn hist_trim_with_bell_guard(hist: String) -> String
|
||||
extern fn clean_llm_response(s: String) -> String
|
||||
extern fn conv_history_persist(hist: String) -> Void
|
||||
extern fn conv_history_load() -> String
|
||||
extern fn conv_history_persist(session_id: String, hist: String) -> Void
|
||||
extern fn conv_history_load(session_id: String) -> String
|
||||
extern fn conv_history_record(session_id: String, user_msg: String, assistant_msg: String, receipt: String) -> Void
|
||||
extern fn conv_history_block(session_id: String) -> String
|
||||
extern fn layered_generate(prompt: String, imprint_id: String, session_id: String) -> String
|
||||
extern fn session_preload_bullets(nodes: String, max_bullets: Int, snip_len: Int) -> String
|
||||
extern fn affective_context_prefix() -> String
|
||||
extern fn handle_chat(body: String) -> String
|
||||
extern fn handle_see(body: String) -> String
|
||||
extern fn studio_tools_json() -> String
|
||||
@@ -36,7 +53,14 @@ extern fn llm_base_url() -> String
|
||||
extern fn llm_wire_format() -> String
|
||||
extern fn json_escape(s: String) -> String
|
||||
extern fn openai_chat_complete(model: String, base_url: String, api_key: String, safe_sys: String, messages_json: String) -> String
|
||||
extern fn openai_tools_json(tools_anthropic: String) -> String
|
||||
extern fn utf8_safe_slice(s: String, n: Int) -> String
|
||||
extern fn json_trim_dangling_escape(s: String) -> String
|
||||
extern fn agentic_tools_no_web() -> String
|
||||
extern fn openai_agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, tools_log_in: String) -> String
|
||||
extern fn agentic_tools_literal() -> String
|
||||
extern fn web_search_tool_json() -> String
|
||||
extern fn strip_client_web_search(tools_inner: String) -> String
|
||||
extern fn agentic_tools_with_web() -> String
|
||||
extern fn connector_tools_json() -> String
|
||||
extern fn agentic_tools_all() -> String
|
||||
@@ -46,13 +70,17 @@ extern fn call_neuron_mcp(tool_name: String, args: String) -> String
|
||||
extern fn agent_workspace_root() -> String
|
||||
extern fn path_within_root(path: String, root: String) -> Bool
|
||||
extern fn resolve_in_root(path: String, root: String) -> String
|
||||
extern fn run_command_is_readonly(cmd: String) -> Bool
|
||||
extern fn cmd_abs_escape_at(cmd: String, root: String, needle: String) -> Bool
|
||||
extern fn run_command_guard(cmd: String, root: String) -> String
|
||||
extern fn classify_tool_risk(tool_name: String, tool_input: String) -> String
|
||||
extern fn dispatch_tool(tool_name: String, tool_input: String) -> String
|
||||
extern fn is_builtin_tool(tool_name: String) -> Bool
|
||||
extern fn next_bridge_id() -> String
|
||||
extern fn handle_chat_plan(body: String) -> String
|
||||
extern fn handle_chat_agentic(body: String) -> String
|
||||
extern fn agentic_loop(session_id: String, model: String, safe_sys: String, tools_json: String, messages_in: String, h: Map, tools_log_in: String) -> String
|
||||
extern fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String) -> Bool
|
||||
extern fn bridge_save(session_id: String, model: String, safe_sys: String, tools_json: String, messages: String, tools_log: String, tool_use_id: String, wire: String) -> Bool
|
||||
extern fn agentic_resume(session_id: String, tool_use_id: String, content: String) -> String
|
||||
extern fn handle_tool_result(session_id: String, body: String) -> String
|
||||
extern fn handle_chat_as_soul(body: String) -> String
|
||||
|
||||
+23
-4
@@ -5,6 +5,15 @@ el_val_t add_punct(el_val_t s, el_val_t intent);
|
||||
el_val_t add_to_seen(el_val_t seen, el_val_t node_id);
|
||||
el_val_t aff_try_slot(el_val_t slot_json, el_val_t aff_7d_ts, el_val_t acc_key);
|
||||
el_val_t affective_context_prefix(void);
|
||||
el_val_t is_utility_request(el_val_t body, el_val_t session_id);
|
||||
el_val_t operator_identity_block(void);
|
||||
el_val_t provenance_add_sources(el_val_t block, el_val_t btype, el_val_t has_cit, el_val_t cit_raw, el_val_t acc);
|
||||
el_val_t provenance_names(el_val_t tools_used);
|
||||
el_val_t provenance_scan_urls(el_val_t arr, el_val_t acc);
|
||||
el_val_t text_join_sep(el_val_t accumulated, el_val_t incoming, el_val_t after_interruption);
|
||||
el_val_t receipt_rule(void);
|
||||
el_val_t receipt_strip(el_val_t s);
|
||||
el_val_t tool_receipt(el_val_t tools_used, el_val_t sources);
|
||||
el_val_t agent_number(el_val_t agent);
|
||||
el_val_t agent_person(el_val_t agent);
|
||||
el_val_t agent_workspace_root(void);
|
||||
@@ -132,7 +141,7 @@ el_val_t awareness_run(void);
|
||||
el_val_t axon_get(el_val_t path);
|
||||
el_val_t axon_post(el_val_t path, el_val_t body);
|
||||
el_val_t bounded_persona_floor(void);
|
||||
el_val_t bridge_save(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages, el_val_t tools_log, el_val_t tool_use_id);
|
||||
el_val_t bridge_save(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages, el_val_t tools_log, el_val_t tool_use_id, el_val_t wire);
|
||||
el_val_t build_form_from_json(el_val_t semantic_form_json, el_val_t lang_code);
|
||||
el_val_t build_np(el_val_t referent, el_val_t slots);
|
||||
el_val_t build_pp(el_val_t loc);
|
||||
@@ -151,8 +160,12 @@ el_val_t cmd_abs_escape_at(el_val_t cmd, el_val_t root, el_val_t needle);
|
||||
el_val_t connectd_get(el_val_t suffix);
|
||||
el_val_t connectd_post(el_val_t suffix, el_val_t body);
|
||||
el_val_t connector_tools_json(void);
|
||||
el_val_t conv_history_load(void);
|
||||
el_val_t conv_history_persist(el_val_t hist);
|
||||
el_val_t conv_hist_key(el_val_t session_id);
|
||||
el_val_t conv_hist_label(el_val_t session_id);
|
||||
el_val_t conv_history_block(el_val_t session_id);
|
||||
el_val_t conv_history_load(el_val_t session_id);
|
||||
el_val_t conv_history_persist(el_val_t session_id, el_val_t hist);
|
||||
el_val_t conv_history_record(el_val_t session_id, el_val_t user_msg, el_val_t assistant_msg, el_val_t receipt);
|
||||
el_val_t cop_article(el_val_t gender, el_val_t number, el_val_t definite);
|
||||
el_val_t cop_bwk_future(el_val_t prefix);
|
||||
el_val_t cop_bwk_perfect(el_val_t prefix);
|
||||
@@ -782,7 +795,8 @@ el_val_t lang_profile_uga(void);
|
||||
el_val_t lang_profile_zh(void);
|
||||
el_val_t lang_profile(el_val_t code, el_val_t word_order, el_val_t morph_type, el_val_t has_case, el_val_t has_gender, el_val_t script_dir, el_val_t agreement, el_val_t null_subject);
|
||||
el_val_t lang_word_order(el_val_t profile);
|
||||
el_val_t layered_cycle(el_val_t raw_input);
|
||||
el_val_t layered_cycle(el_val_t raw_input, el_val_t session_id, el_val_t utility);
|
||||
el_val_t layered_generate(el_val_t prompt, el_val_t imprint_id, el_val_t session_id);
|
||||
el_val_t lex_class(el_val_t entry);
|
||||
el_val_t lex_form(el_val_t entry, el_val_t idx);
|
||||
el_val_t lex_pos(el_val_t entry);
|
||||
@@ -861,6 +875,11 @@ el_val_t non_weak_past(el_val_t stem, el_val_t slot);
|
||||
el_val_t non_weak_present(el_val_t stem, el_val_t slot);
|
||||
el_val_t one_cycle(void);
|
||||
el_val_t openai_chat_complete(el_val_t model, el_val_t base_url, el_val_t api_key, el_val_t safe_sys, el_val_t messages_json);
|
||||
el_val_t openai_tools_json(el_val_t tools_anthropic);
|
||||
el_val_t json_trim_dangling_escape(el_val_t s);
|
||||
el_val_t utf8_safe_slice(el_val_t s, el_val_t n);
|
||||
el_val_t agentic_tools_no_web(void);
|
||||
el_val_t openai_agentic_loop(el_val_t session_id, el_val_t model, el_val_t safe_sys, el_val_t tools_json, el_val_t messages_in, el_val_t tools_log_in);
|
||||
el_val_t parse_float_x100(el_val_t s);
|
||||
el_val_t path_within_root(el_val_t path, el_val_t root);
|
||||
el_val_t peo_ah_past(el_val_t slot);
|
||||
|
||||
+15
@@ -1,3 +1,18 @@
|
||||
// ╔══════════════════════════════════════════════════════════════════════════╗
|
||||
// ║ STALE BUNDLE — DO NOT BUILD. UNSAFE CHAT PATH. ║
|
||||
// ╚══════════════════════════════════════════════════════════════════════════╝
|
||||
// This concatenated bundle is a snapshot, not a source of truth, and it is stale in
|
||||
// a way that matters for safety: it wires /api/chat straight to handle_chat and
|
||||
// contains NO layered_cycle at all (verified: zero occurrences in the bundled code —
|
||||
// the only textual hit in this file is this banner). A binary built from
|
||||
// this file would run chat with no enforcing input gate (no safety_screen, no
|
||||
// hard-bell short-circuit) and no enforcing output gate (no safety_validate).
|
||||
//
|
||||
// Build from the .el sources via manifest.el (entry soul.el), or from dist/soul.c.
|
||||
// Nothing in the repo references this file. It is kept only as a historical artifact
|
||||
// and should be deleted once Will confirms nothing external depends on it.
|
||||
// (Flagged 2026-08-04 in _engine-websearch-20260804/SAFETY-STOP.md; banner added
|
||||
// 2026-08-05 with the plain-chat generation fix.)
|
||||
// language-profile.el - Language profile data and accessors.
|
||||
//
|
||||
// A language profile is a slot map ([String] key-value list) describing the
|
||||
|
||||
+1
-1
@@ -28170,7 +28170,7 @@ el_val_t agentic_loop(el_val_t session_id, el_val_t model, el_val_t safe_sys, el
|
||||
state_set(el_str_concat(EL_STR("run_progress_"), session_id), EL_STR(""));
|
||||
}
|
||||
while (keep_going && (iteration < 8)) {
|
||||
el_val_t req_body = el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"model\":\""), model), EL_STR("\"")), EL_STR(",\"max_tokens\":4096")), EL_STR(",\"system\":\"")), safe_sys), EL_STR("\"")), EL_STR(",\"tools\":")), tools_json), EL_STR(",\"messages\":")), messages), EL_STR("}"));
|
||||
el_val_t req_body = el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(el_str_concat(EL_STR("{\"model\":\""), model), EL_STR("\"")), EL_STR(",\"max_tokens\":4096,\"tool_choice\":{\"type\":\"auto\",\"disable_parallel_tool_use\":true}")), EL_STR(",\"system\":\"")), safe_sys), EL_STR("\"")), EL_STR(",\"tools\":")), tools_json), EL_STR(",\"messages\":")), messages), EL_STR("}"));
|
||||
el_val_t raw_resp = http_post_with_headers(api_url, req_body, h);
|
||||
el_val_t is_error = ((str_starts_with(raw_resp, EL_STR("{\"error\"")) || str_starts_with(raw_resp, EL_STR("{\"type\":\"error\""))) || str_contains(raw_resp, EL_STR("authentication_error")));
|
||||
if (is_error) {
|
||||
|
||||
@@ -15,6 +15,40 @@ fn flag_true(body: String, key: String) -> Bool {
|
||||
return json_get_bool(body, key) || json_get_int(body, key) > 0
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// plain_chat_envelope — the JSON response contract for a non-agentic ("Tools: Off")
|
||||
// chat turn. Every /api/chat dispatch that calls layered_cycle goes through here, so
|
||||
// the three call sites cannot drift apart.
|
||||
//
|
||||
// WHY THE ENVELOPE IS BUILT HERE AND NOT INSIDE layered_cycle:
|
||||
// layered_cycle returns the user-facing text AFTER safety_validate has acted on it.
|
||||
// Keeping the JSON out of the cycle means the output gate always sees raw model text
|
||||
// and never an escaped blob — there is nothing to unwrap and re-wrap on the crisis
|
||||
// path, which is exactly the failure mode that made wiring handle_chat unsafe.
|
||||
// Escaping is the last thing that happens, strictly after the gate.
|
||||
//
|
||||
// FIELDS: `reply` and `response` carry the same validated text. Both are required by
|
||||
// live clients — the desktop app reads `reply` first (DaemonClient.parseChatResponse),
|
||||
// while the CLI tools and the Telegram gateway read `response` (the gateway reads only
|
||||
// `response`). Emitting one would break the other.
|
||||
//
|
||||
// EMPTY MEANS FAILURE, NOT AN EMPTY ANSWER: a hard bell returns the fixed crisis
|
||||
// message and a soft bell is padded to non-empty by safety_validate, so the only way
|
||||
// an empty string leaves the cycle is a failed model call. It is reported as an error
|
||||
// rather than dressed up as a successful blank reply.
|
||||
// ---------------------------------------------------------------------------
|
||||
fn plain_chat_envelope(validated: String, model: String) -> String {
|
||||
if str_eq(validated, "") {
|
||||
return "{\"error\":\"llm unavailable\",\"reply\":\"\",\"response\":\"\",\"agentic\":false,\"tools_used\":[]}"
|
||||
}
|
||||
let safe: String = json_safe(validated)
|
||||
return "{\"reply\":\"" + safe + "\""
|
||||
+ ",\"response\":\"" + safe + "\""
|
||||
+ ",\"model\":\"" + json_safe(model) + "\""
|
||||
+ ",\"agentic\":false"
|
||||
+ ",\"tools_used\":[]}"
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rate limiting — simple in-memory per-IP sliding window counter.
|
||||
//
|
||||
@@ -243,8 +277,14 @@ fn handle_dharma_recv(body: String) -> String {
|
||||
} else if agentic_flag {
|
||||
handle_chat_agentic(chat_body)
|
||||
} else {
|
||||
let screened_reply: String = layered_cycle(raw_msg)
|
||||
screened_reply
|
||||
// Non-agentic ("Tools: Off"): the full L1→L2→L3→L1 cycle, which now generates
|
||||
// at L3 instead of echoing. Envelope built outside the cycle — see
|
||||
// plain_chat_envelope.
|
||||
// FIX B/E1 (2026-08-05): the cycle is told which conversation it is in, and
|
||||
// whether this generation is conversation at all. Same two arguments at all
|
||||
// three dispatch sites.
|
||||
let screened_reply: String = layered_cycle(raw_msg, json_get(chat_body, "session_id"), is_utility_request(chat_body, json_get(chat_body, "session_id")))
|
||||
plain_chat_envelope(screened_reply, chat_default_model())
|
||||
}
|
||||
auto_persist(chat_body, reply)
|
||||
return reply
|
||||
@@ -416,8 +456,11 @@ fn handle_request(method: String, path: String, body: String) -> String {
|
||||
} else if agentic_flag {
|
||||
handle_chat_agentic(body)
|
||||
} else {
|
||||
let screened_reply: String = layered_cycle(eff_msg)
|
||||
screened_reply
|
||||
// Non-agentic ("Tools: Off") — same cycle and same envelope as POST.
|
||||
// FIX B/E1: same threading. A GET probe usually carries no session_id, which
|
||||
// resolves to the anonymous window — the documented behaviour for this door.
|
||||
let screened_reply: String = layered_cycle(eff_msg, json_get(body, "session_id"), is_utility_request(body, json_get(body, "session_id")))
|
||||
plain_chat_envelope(screened_reply, chat_default_model())
|
||||
}
|
||||
auto_persist(body, reply)
|
||||
return reply
|
||||
@@ -580,8 +623,13 @@ fn handle_request(method: String, path: String, body: String) -> String {
|
||||
} else if agentic_flag {
|
||||
handle_chat_agentic(body)
|
||||
} else {
|
||||
let screened_reply: String = layered_cycle(raw_msg)
|
||||
screened_reply
|
||||
// Non-agentic ("Tools: Off") — the app's DEFAULT mode (AgentMode.NEVER).
|
||||
// Full L1→L2→L3→L1 cycle with real generation at L3; envelope built
|
||||
// outside the cycle so safety_validate always sees raw text.
|
||||
// FIX B/E1: same threading. This is the app's main plain-chat door, so this
|
||||
// is the site that ends the blank stare in practice.
|
||||
let screened_reply: String = layered_cycle(raw_msg, json_get(body, "session_id"), is_utility_request(body, json_get(body, "session_id")))
|
||||
plain_chat_envelope(screened_reply, chat_default_model())
|
||||
}
|
||||
auto_persist(body, reply)
|
||||
return reply
|
||||
|
||||
@@ -379,9 +379,23 @@ fn emit_session_start_event() -> Void {
|
||||
// layered_cycle — routes user-facing requests through the 4-layer consciousness stack.
|
||||
// L0 (core) → L1 (safety screen) → L2a (continuity + behavioral profiling) → L2b (mission alignment) → L3 (imprint) → L1 (safety validate)
|
||||
// Internal cognition (heartbeat, proactive, memory ops) bypasses layers — use one_cycle directly.
|
||||
fn layered_cycle(raw_input: String) -> String {
|
||||
let history: String = state_get("conv_history")
|
||||
let session_id: String = state_get("current_session_id")
|
||||
//
|
||||
// FIX B (2026-08-05) — the cycle now knows which conversation it is in.
|
||||
//
|
||||
// session_id: the caller's session, threaded from the route. Was previously read from the
|
||||
// state key "current_session_id", which is read HERE and written NOWHERE in the entire
|
||||
// source — verified across every .el file. So this value was unconditionally "", and every
|
||||
// downstream consumer of it silently fell back to a process-global bucket: conversation
|
||||
// history, and the steward's continuity tracking (TODO reliability #4, below, describes the
|
||||
// cross-session bleed this caused; threading the real id closes it). The plain path's blank
|
||||
// stare and the agentic path's scoped history were the same defect seen from two sides.
|
||||
//
|
||||
// utility: true when the generation is not part of the user's conversation — the app's
|
||||
// title and insight passes. Answered normally, never recorded. See is_utility_request.
|
||||
fn layered_cycle(raw_input: String, session_id: String, utility: Bool) -> String {
|
||||
// Safety-screen history amplification now reads the SAME window the turn will be
|
||||
// recorded into, so a session's own escalation pattern is what gets scored.
|
||||
let history: String = state_get(conv_hist_key(session_id))
|
||||
|
||||
// L1 in: safety screen
|
||||
let screen_result: String = safety_screen(raw_input, history)
|
||||
@@ -423,8 +437,10 @@ fn layered_cycle(raw_input: String) -> String {
|
||||
let cont_action: String = json_get(continuity, "action")
|
||||
|
||||
// Store continuity status so imprint can adjust its response register.
|
||||
// TODO(reliability #4): session_continuity is process-global; scope per session_id
|
||||
// when available to prevent cross-session bleed under concurrent layered_cycle calls.
|
||||
// TODO(reliability #4) CLOSED 2026-08-05: this line was already written to scope per
|
||||
// session — it just never received a session id, because the only source was a state key
|
||||
// nothing wrote. It is now threaded from the route, so named sessions genuinely get their
|
||||
// own continuity state and only anonymous callers share the global one.
|
||||
let cont_key: String = if str_eq(session_id, "") { "session_continuity" } else { "session_continuity:" + session_id }
|
||||
state_set(cont_key, cont_status)
|
||||
|
||||
@@ -453,40 +469,24 @@ fn layered_cycle(raw_input: String) -> String {
|
||||
let lc_aff_cutoff: Int = time_now() - 259200
|
||||
let lc_bell_nodes: String = engram_search_json("bell:soft bell:hard BellEvent affective", 2)
|
||||
let lc_has_bell: Bool = !str_eq(lc_bell_nodes, "") && !str_eq(lc_bell_nodes, "[]")
|
||||
// CRASH FIX 2026-08-05 (BUG-PLAINCHAT-1): the " | ts:" parser used to be inline here.
|
||||
// Inside this block-expression initializer elc compiled `lbmp + str_len(lbm)` to
|
||||
// el_str_concat() on two integers, which segfaulted the whole daemon the moment a
|
||||
// distress turn followed an earlier affective turn — i.e. exactly on the crisis path.
|
||||
// Verified against the unmodified baseline binary AND present in the committed
|
||||
// dist/soul.c. affective_node_ts() is a top-level function, where the same expression
|
||||
// compiles to integer addition. Do not inline it back.
|
||||
let lc_bell_note: String = if lc_has_bell {
|
||||
let lb0: String = json_array_get(lc_bell_nodes, 0)
|
||||
let lb_c: String = json_get(lb0, "content")
|
||||
let lbm: String = " | ts:"
|
||||
let lbmp: Int = str_index_of(lb_c, lbm)
|
||||
let lb_ts_raw: String = if lbmp >= 0 {
|
||||
let lbs: Int = lbmp + str_len(lbm)
|
||||
let lbr: String = str_slice(lb_c, lbs, str_len(lb_c))
|
||||
let lbn: Int = str_index_of(lbr, " | ")
|
||||
if lbn < 0 { lbr } else { str_slice(lbr, 0, lbn) }
|
||||
} else {
|
||||
let lbca: String = json_get(lb0, "created_at")
|
||||
if str_eq(lbca, "") { json_get(lb0, "updated_at") } else { lbca }
|
||||
}
|
||||
let lb_ts: Int = if str_eq(lb_ts_raw, "") { 0 } else { str_to_int(lb_ts_raw) }
|
||||
let lb_ts: Int = affective_node_ts(lb0)
|
||||
if lb_ts > lc_aff_cutoff { "[AFFECTIVE NOTE: User was in distress in a recent session.]" } else { "" }
|
||||
} else { "" }
|
||||
let lc_pos_nodes: String = engram_search_json("PositiveEvent joy:high joy:low affective", 2)
|
||||
let lc_has_pos: Bool = !str_eq(lc_pos_nodes, "") && !str_eq(lc_pos_nodes, "[]")
|
||||
// Same crash fix as the bell note above (BUG-PLAINCHAT-1).
|
||||
let lc_pos_note: String = if lc_has_pos && str_eq(lc_bell_note, "") {
|
||||
let lp0: String = json_array_get(lc_pos_nodes, 0)
|
||||
let lp_c: String = json_get(lp0, "content")
|
||||
let lpm: String = " | ts:"
|
||||
let lpmp: Int = str_index_of(lp_c, lpm)
|
||||
let lp_ts_raw: String = if lpmp >= 0 {
|
||||
let lps: Int = lpmp + str_len(lpm)
|
||||
let lpr: String = str_slice(lp_c, lps, str_len(lp_c))
|
||||
let lpn: Int = str_index_of(lpr, " | ")
|
||||
if lpn < 0 { lpr } else { str_slice(lpr, 0, lpn) }
|
||||
} else {
|
||||
let lpca: String = json_get(lp0, "created_at")
|
||||
if str_eq(lpca, "") { json_get(lp0, "updated_at") } else { lpca }
|
||||
}
|
||||
let lp_ts: Int = if str_eq(lp_ts_raw, "") { 0 } else { str_to_int(lp_ts_raw) }
|
||||
let lp_ts: Int = affective_node_ts(lp0)
|
||||
if lp_ts > lc_aff_cutoff { "[AFFECTIVE NOTE: User shared positive news in a recent session.]" } else { "" }
|
||||
} else { "" }
|
||||
let lc_affective_note: String = if !str_eq(lc_bell_note, "") { lc_bell_note } else { lc_pos_note }
|
||||
@@ -498,11 +498,47 @@ fn layered_cycle(raw_input: String) -> String {
|
||||
}
|
||||
state_set("layered_cycle_safety_system_addendum", augmented_addendum)
|
||||
|
||||
// L3: imprint responds
|
||||
let output: String = imprint_respond(aligned, imprint_id)
|
||||
// L3: imprint responds — applies the active imprint's voice/domain annotation to the
|
||||
// steward-aligned input. This produces the PROMPT, not the answer.
|
||||
let prompt: String = imprint_respond(aligned, imprint_id)
|
||||
|
||||
// L1 out: validate output before delivery
|
||||
return safety_validate(output, screen_action)
|
||||
// L3b: the imprint SPEAKS (added 2026-08-05).
|
||||
//
|
||||
// Until now the cycle stopped at the annotation above, so /api/chat with agentic:false
|
||||
// handed the user's own screened text back as the "reply" — every gate ran, but nothing
|
||||
// ever generated. The generation is placed HERE, inside the cycle, rather than by
|
||||
// pointing the route at handle_chat(): handle_chat has no enforcing input gate and no
|
||||
// enforcing output gate, so calling it instead of this cycle would have traded the whole
|
||||
// safety pipeline for a working reply. Composing keeps both.
|
||||
//
|
||||
// Order is deliberate and must not be rearranged: this call sits strictly AFTER the L1
|
||||
// screen, the safe-mode guard, the hard-bell short-circuit and the L2 stewardship layers,
|
||||
// and strictly BEFORE the L1 output gate. A hard bell never reaches a model — the branch
|
||||
// above returns first. Tools are not offered on this turn; see layered_generate.
|
||||
let output: String = layered_generate(prompt, imprint_id, session_id)
|
||||
|
||||
// L1 out: validate output before delivery. Still the terminal gate — nothing below this
|
||||
// line can change the string this function returns.
|
||||
let validated: String = safety_validate(output, screen_action)
|
||||
|
||||
// Turn bookkeeping. Records the VALIDATED text, never the raw model output, and is only
|
||||
// reachable on the non-bell path: both bell branches above return before this point, so
|
||||
// bell turns still never enter conversation history. Pure state side effect — it cannot
|
||||
// alter what is returned.
|
||||
//
|
||||
// FIX A: the receipt is unconditional and always negative on this path, because on this
|
||||
// path it is structurally true — layered_generate offers no tools at all (build_system_prompt
|
||||
// chat mode + a request body with no "tools" key). Recording "no tools ran" is not padding:
|
||||
// it is the only thing that distinguishes "nothing ran" from "we forgot to write down what
|
||||
// ran", and that ambiguity is what made the model confess to a search it had performed.
|
||||
//
|
||||
// FIX E1: a utility generation is answered but not recorded. Guarded here rather than at
|
||||
// the route so every /api/chat dispatch site inherits it from one place.
|
||||
let receipt: String = tool_receipt("", "")
|
||||
if !utility {
|
||||
conv_history_record(session_id, raw_input, validated, receipt)
|
||||
}
|
||||
return validated
|
||||
}
|
||||
|
||||
let soul_cgi_id_raw: String = env("SOUL_CGI_ID")
|
||||
|
||||
@@ -1,6 +1,8 @@
|
||||
// auto-generated by elc --emit-header - do not edit
|
||||
extern fn init_soul_edges() -> Void
|
||||
extern fn ensure_self_canonical_bridge() -> Void
|
||||
extern fn aff_try_slot(slot_json: String, aff_7d_ts: Int, acc_key: String) -> Void
|
||||
extern fn load_identity_context() -> Void
|
||||
extern fn seed_persona_from_env() -> Void
|
||||
extern fn emit_session_start_event() -> Void
|
||||
extern fn layered_cycle(raw_input: String) -> String
|
||||
extern fn layered_cycle(raw_input: String, session_id: String, utility: Bool) -> String
|
||||
|
||||
@@ -0,0 +1,90 @@
|
||||
# gate-openai — deterministic OpenAI-dialect provider stub
|
||||
|
||||
Staging home for the **soul-openai-tools-v2** gate scaffolding
|
||||
(`docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md`, test-plan rung 1:
|
||||
"stub first — discriminates before El code exists"). Sibling of gate9's
|
||||
Anthropic stub (`_wt-beta-round9/scripts/gate9/stub-llm.py`): same scenario
|
||||
mechanism, opposite wire dialect. Stdlib Python only, 127.0.0.1 only,
|
||||
refuses ports 7770/7779/17779. Run `./selftest.sh` — exit 0 is green.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Role |
|
||||
|---|---|
|
||||
| `stub-openai.py` | HTTP server: `POST /v1/chat/completions` (OpenAI dialect), scenario-scripted responses, request validation, ground-truth JSONL log, hostile modes via `--mode` |
|
||||
| `scenarios-openai.json` | Scenario contract: scripts + markers + per-class/per-step request assertions |
|
||||
| `selftest.sh` | curl-driven proof of every scenario, every rejection, all hostile modes (58 checks) |
|
||||
|
||||
## What each scenario proves (when the brain drives it)
|
||||
|
||||
| Class | Proves |
|
||||
|---|---|
|
||||
| `oa-plain` | finish_reason `stop` ends the loop; tools + `tool_choice` + `parallel_tool_calls:false` were offered on the wire |
|
||||
| `oa-tools-off` | the chat-only lane sends NO tools (offering them there is a 400) |
|
||||
| `oa-single-tool` | full round-trip: `tool_calls` parsed, assistant echo + `role:"tool"` turn with matching `tool_call_id` sent back, final text reached |
|
||||
| `oa-torture` | `function.arguments` (JSON-encoded string with nested quotes, backslashes, newlines, tabs, unicode) survives exactly ONE decode — the stub recomputes the issued payload from the script and 400s on any drift (`gate_echo_mismatch`, the spec §6 two-escaper trap) |
|
||||
| `oa-parallel` | two `tool_calls` in one response: the brain either answers both (paired correctly) or rejects cleanly — an unpaired echo is a 400 |
|
||||
| `oa-mission` | multi-round loop continuation; step index = assistant-message count, so resume threads index correctly by construction |
|
||||
| `oa-api-error` | provider errors 400/429/500/503 in the OpenAI error envelope surface honestly, no retry storm |
|
||||
|
||||
Universal (every request, any scenario): Anthropic dialect leakage fails
|
||||
loudly with 400 — `anthropic-version` header, top-level `system` /
|
||||
`stop_sequences` / `max_tokens_to_sample`, `input_schema` inside tools,
|
||||
Anthropic content blocks (`tool_use`/`tool_result`/...). Tools must be
|
||||
`{type:"function", function:{name, description, parameters}}`, unique names;
|
||||
echoed `arguments` must be a JSON-encoded STRING, never a decoded object.
|
||||
|
||||
## Hostile modes (`--mode`, same file)
|
||||
|
||||
| Mode | Behavior | Brain invariant under test |
|
||||
|---|---|---|
|
||||
| `black-hole` | reads the request, never responds | HTTP timeout exists and surfaces; no silent hang |
|
||||
| `mid-body-drop` | 200 headers, half a JSON body, socket abort | truncated body = clean error, never a half-parsed reply shown as real |
|
||||
| `tool-pending-forever` | every request gets a fresh `tool_calls` response, forever | the loop's iteration cap trips (`max_loop_iterations: 16` in the contract); count actual round-trips via `GET /gate/stats` (`chat_hits`) |
|
||||
|
||||
## How the brain-side gate consumes this
|
||||
|
||||
1. Start: `stub-openai.py --port P --scenarios scenarios-openai.json --log run.jsonl`
|
||||
2. Point the brain at it: `NEURON_LLM_0_URL=http://127.0.0.1:P` +
|
||||
`NEURON_LLM_0_FORMAT=openai` (spec step 0 must verify these actually
|
||||
export at runtime), scratch profile, free soul port.
|
||||
3. Send each phrasing's `prompt` (the marker selects the script); assert the
|
||||
brain's claims (`tools_used`, reply, ledger) against the stub's JSONL log
|
||||
— truth, not narration — plus files on disk for write_file scenarios.
|
||||
4. Any stub 400 = the brain sent a malformed/leaked request; the gate fails
|
||||
with the stub's reason string.
|
||||
5. Re-run gate9's Anthropic matrix unchanged = proof the Anthropic lane is
|
||||
byte-untouched.
|
||||
|
||||
## Reconciliation into gate9 (app repo) — AFTER round 9 merges
|
||||
|
||||
This dir is staging only; the merge is mechanical by design:
|
||||
- `stub-openai.py` + `scenarios-openai.json` move to `scripts/gate9/`
|
||||
alongside `stub-llm.py` + `scenarios.json` (shared conventions: marker
|
||||
matching, assistant-count step indexing, `--port/--scenarios/--log`,
|
||||
JSONL fields `seq/ts/kind/scenario_class/phrasing/step/validation/
|
||||
delivered/http_status`, prod-port refusal, benign background responses,
|
||||
`GATE-SCRIPT-EXHAUSTED` overrun, `{N}/{NN}` repeat expansion).
|
||||
- `prompt-matrix-gate.sh` gains a dialect axis (anthropic|openai) choosing
|
||||
stub + scenario file; `matrix-asserts.py` reads the same log shape.
|
||||
- The `--mode` hostile flags here are PROVIDER-side (brain↔LLM boundary);
|
||||
gate9's `hostile/` servers are SOUL-side (app↔brain boundary). They are
|
||||
complementary, not duplicates — both stay.
|
||||
|
||||
## Open questions for the port author (stub asserts a position; confirm or change)
|
||||
|
||||
1. `parallel_tool_calls` must be **explicitly false** on every tool-bearing
|
||||
request (ADR-0005 pin). If the builder omits it instead, relax
|
||||
`defaults.expect_request.parallel_tool_calls` to `null`.
|
||||
2. `tool_choice` must be present (`"auto"` expected). If the brain relies on
|
||||
the provider default, drop `require_tool_choice`.
|
||||
3. Tool-result `content` is asserted only to be a string; if the brain sends
|
||||
structured JSON-in-string (like `{"ok":true,...}`), no change needed.
|
||||
4. Groq compatibility: Groq's OpenAI-compat endpoint rejects some optional
|
||||
fields; whatever field set the brain settles on for live Groq E2E must be
|
||||
mirrored here so the deterministic gate and the live lane assert the SAME
|
||||
request shape.
|
||||
5. The stub treats a `role:"tool"` turn answering an already-answered id as
|
||||
400; if the resume path can legitimately replay tool results, that rule
|
||||
needs a resume-aware carve-out (gate9's Anthropic stub faced the same
|
||||
issue — see its PASS 1 comment).
|
||||
Executable
+559
@@ -0,0 +1,559 @@
|
||||
#!/usr/bin/env bash
|
||||
# run-lane-gate.sh — brain-side driver for the OpenAI-dialect gate.
|
||||
#
|
||||
# Drives the REAL soul binary against stub-openai.py for every class and every
|
||||
# phrasing in scenarios-openai.json, plus the three hostile provider modes, and
|
||||
# asserts the brain's claims against the stub's ground-truth JSONL (truth, not
|
||||
# narration) and against files on disk.
|
||||
#
|
||||
# SAFETY (hard rules, enforced below):
|
||||
# - never binds 7770 / 7779 / 17779 - only 7891-7894
|
||||
# - never reads or writes ~/.neuron - HOME is redirected to a scratch dir
|
||||
# - every process started here is killed on exit (trap) and proven with lsof
|
||||
#
|
||||
# The soul runs under `script -q /dev/null` so its stdout is a pty: El's
|
||||
# println() uses puts(), which is FULLY buffered to a file, and the process is
|
||||
# killed without flushing — the DRIFT lines would be invisible otherwise.
|
||||
#
|
||||
# Usage: ./run-lane-gate.sh [all|bridge|local|toolsoff|hostile]
|
||||
# bridge = consent round-trip config (no workspace root -> write_file is
|
||||
# "escalate" -> the loop suspends and the CLIENT executes the tool)
|
||||
# local = workspace-root config (write_file is "reversible" + builtin ->
|
||||
# the loop executes the tool in-process and runs to completion)
|
||||
# toolsoff = supplementary: non-agentic lane against a base URL WITHOUT the
|
||||
# /v1 suffix (the el-runtime provider chain appends
|
||||
# /v1/chat/completions itself, unlike chat.el which appends only
|
||||
# /chat/completions)
|
||||
# hostile = black-hole / mid-body-drop / tool-pending-forever
|
||||
#
|
||||
# Env overrides: SOUL_BIN, STUB_PORT, SOUL_PORT, SOUL_PORT_B, RUN_ROOT
|
||||
set -uo pipefail
|
||||
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
PHASES="${1:-all}"
|
||||
|
||||
SOUL_BIN="${SOUL_BIN:-/tmp/soul-oai2/soul-openai-tools}"
|
||||
STUB_PORT="${STUB_PORT:-7891}"
|
||||
SOUL_PORT="${SOUL_PORT:-7892}"
|
||||
SOUL_PORT_B="${SOUL_PORT_B:-7893}"
|
||||
RUN_ROOT="${RUN_ROOT:-/tmp/oa-lane-gate}"
|
||||
STAMP="$(date +%Y%m%d-%H%M%S)"
|
||||
RUN="$RUN_ROOT/$STAMP"
|
||||
|
||||
for p in "$STUB_PORT" "$SOUL_PORT" "$SOUL_PORT_B"; do
|
||||
case "$p" in
|
||||
7770|7779|17779) echo "FATAL: refusing production Neuron port $p"; exit 2;;
|
||||
789[1-4]) ;;
|
||||
*) echo "FATAL: port $p outside the allowed 7891-7894 range"; exit 2;;
|
||||
esac
|
||||
done
|
||||
[ -x "$SOUL_BIN" ] || { echo "FATAL: soul binary not found/executable: $SOUL_BIN"; exit 2; }
|
||||
|
||||
mkdir -p "$RUN/home" "$RUN/ws-bridge" "$RUN/ws-local" "$RUN/ws-off" "$RUN/engram"
|
||||
echo '{"nodes":[],"edges":[]}' > "$RUN/engram/snapshot.json"
|
||||
DRV="$RUN/drv.py"
|
||||
|
||||
STUB_PID=""; SOUL_PID=""
|
||||
cleanup() {
|
||||
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
|
||||
pkill -f "$SOUL_BIN" 2>/dev/null
|
||||
[ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null
|
||||
sleep 0.4
|
||||
[ -n "$SOUL_PID" ] && kill -9 "$SOUL_PID" 2>/dev/null
|
||||
[ -n "$STUB_PID" ] && kill -9 "$STUB_PID" 2>/dev/null
|
||||
return 0
|
||||
}
|
||||
trap cleanup EXIT INT TERM
|
||||
|
||||
start_stub() { # $1 = mode, $2 = log path
|
||||
local mode="$1" log="$2" args=""
|
||||
[ "$mode" = "normal" ] && args="--scenarios $HERE/scenarios-openai.json"
|
||||
# shellcheck disable=SC2086
|
||||
python3 "$HERE/stub-openai.py" --port "$STUB_PORT" --mode "$mode" --log "$log" $args \
|
||||
> "$RUN/stub-$mode.out" 2>&1 &
|
||||
STUB_PID=$!
|
||||
for _ in $(seq 1 50); do
|
||||
curl -sf "http://127.0.0.1:$STUB_PORT/gate/health" >/dev/null 2>&1 && return 0
|
||||
sleep 0.2
|
||||
done
|
||||
echo "FATAL: stub did not come up on $STUB_PORT"; cat "$RUN/stub-$mode.out"; exit 3
|
||||
}
|
||||
stop_stub() { [ -n "$STUB_PID" ] && kill "$STUB_PID" 2>/dev/null; sleep 0.3; STUB_PID=""; }
|
||||
|
||||
start_soul() { # $1 = port, $2 = base url, $3 = soul log, $4 = agent root ("" = none)
|
||||
local port="$1" base="$2" log="$3" root="$4"
|
||||
script -q /dev/null \
|
||||
env -u ANTHROPIC_API_KEY -u SOUL_API_KEY -u ENGRAM_URL -u ENGRAM_API_KEY \
|
||||
-u NEURON_API_URL -u NEURON_TOKEN -u SOUL_LLM_PROVIDER -u SOUL_LLM_BASE_URL \
|
||||
-u NEURON_LLM_1_URL -u NEURON_LLM_1_KEY -u SOUL_IDENTITY \
|
||||
HOME="$RUN/home" PATH="$PATH" \
|
||||
NEURON_PORT="$port" EL_HTTP_BIND_HOST=127.0.0.1 \
|
||||
SOUL_ENGRAM_PATH="$RUN/engram/snapshot.json" \
|
||||
SOUL_CGI_ID=ntn-test SOUL_PERSONA_NAME=Neuron \
|
||||
NEURON_LLM_0_URL="$base" NEURON_LLM_0_FORMAT=openai NEURON_LLM_0_KEY=gate-test-key \
|
||||
${root:+NEURON_AGENT_ROOT="$root"} \
|
||||
"$SOUL_BIN" > "$log" 2>&1 &
|
||||
SOUL_PID=$!
|
||||
for _ in $(seq 1 100); do
|
||||
curl -sf "http://127.0.0.1:$port/health" >/dev/null 2>&1 && return 0
|
||||
sleep 0.2
|
||||
done
|
||||
echo "FATAL: soul did not come up on $port"; tail -20 "$log"; exit 3
|
||||
}
|
||||
stop_soul() {
|
||||
[ -n "$SOUL_PID" ] && kill "$SOUL_PID" 2>/dev/null
|
||||
pkill -f "$SOUL_BIN" 2>/dev/null
|
||||
sleep 0.6; SOUL_PID=""
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------ driver ----
|
||||
cat > "$DRV" <<'PYEOF'
|
||||
import json, os, sys, time, threading, urllib.request, urllib.error
|
||||
|
||||
CFG = json.load(open(sys.argv[1]))
|
||||
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
|
||||
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
|
||||
WS = CFG["workspace"]
|
||||
MODE = CFG["mode"] # bridge | local | toolsoff
|
||||
SCEN = json.load(open(CFG["scenarios"]))
|
||||
STUBLOG = CFG["stub_log"]
|
||||
SOULLOG = CFG["soul_log"]
|
||||
ONLY = CFG.get("classes") or list(SCEN["classes"].keys())
|
||||
MAXHOPS = CFG.get("max_hops", 15)
|
||||
OUT = CFG["out"]
|
||||
# the chat-only class must be driven on the NON-agentic door: the agentic door
|
||||
# always advertises tools, which is a 400 on that scenario by contract.
|
||||
NON_AGENTIC = {"oa-tools-off"}
|
||||
|
||||
def http(method, url, obj=None, timeout=300):
|
||||
data = None if obj is None else json.dumps(obj).encode()
|
||||
req = urllib.request.Request(url, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
body = r.read().decode("utf-8", "replace")
|
||||
st = r.status
|
||||
except urllib.error.HTTPError as e:
|
||||
body = e.read().decode("utf-8", "replace"); st = e.code
|
||||
except Exception as e:
|
||||
return -1, "TRANSPORT-ERROR: %r" % (e,), None
|
||||
try:
|
||||
return st, body, json.loads(body)
|
||||
except ValueError:
|
||||
return st, body, None
|
||||
|
||||
def fsize(p):
|
||||
return os.path.getsize(p) if os.path.exists(p) else 0
|
||||
|
||||
def tail_from(path, off):
|
||||
if not os.path.exists(path):
|
||||
return "", off
|
||||
with open(path, "rb") as f:
|
||||
f.seek(off); chunk = f.read(); return chunk.decode("utf-8", "replace"), f.tell()
|
||||
|
||||
def stub_since(off):
|
||||
"""Exact correlation: only the JSONL bytes appended during this phrasing."""
|
||||
txt, noff = tail_from(STUBLOG, off)
|
||||
recs = []
|
||||
for line in txt.splitlines():
|
||||
line = line.strip()
|
||||
if line:
|
||||
try: recs.append(json.loads(line))
|
||||
except ValueError: pass
|
||||
return recs, noff
|
||||
|
||||
def perform(name, ti):
|
||||
"""Execute the bridged tool for real, like the desktop client would."""
|
||||
if name in ("write_file", "edit_file"):
|
||||
p = ti.get("path", "")
|
||||
dest = p if os.path.isabs(p) else os.path.join(WS, p)
|
||||
os.makedirs(os.path.dirname(dest) or WS, exist_ok=True)
|
||||
body = ti.get("content", "")
|
||||
with open(dest, "w") as f:
|
||||
f.write(body)
|
||||
return "wrote %s (%d bytes)" % (p, len(body.encode()))
|
||||
return "ok"
|
||||
|
||||
class Poller(threading.Thread):
|
||||
def __init__(self, sid):
|
||||
super().__init__(daemon=True); self.sid = sid; self.snaps = []; self.stop = False
|
||||
def run(self):
|
||||
while not self.stop:
|
||||
st, body, js = http("GET", SOUL + "/api/run-progress/" + self.sid, timeout=60)
|
||||
if js and js.get("progress"):
|
||||
if not self.snaps or self.snaps[-1] != js["progress"]:
|
||||
self.snaps.append(js["progress"])
|
||||
time.sleep(0.1)
|
||||
|
||||
def progress(sid):
|
||||
_, _, pj = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
|
||||
return (pj or {}).get("progress")
|
||||
|
||||
def run_phrasing(cname, ph):
|
||||
st, body, js = http("POST", SOUL + "/api/sessions", {"title": ph["id"]}, timeout=60)
|
||||
sid = (js or {}).get("id", "")
|
||||
rec = {"class": cname, "phrasing": ph["id"], "session_id": sid, "legs": [],
|
||||
"pendings": [], "progress_during": [], "progress_per_leg": [],
|
||||
"progress_final": None, "soul_log": "", "stub": [], "http": [],
|
||||
"agentic": cname not in NON_AGENTIC}
|
||||
if not sid:
|
||||
rec["fatal"] = "session create failed: %s %s" % (st, body[:300]); return rec
|
||||
soff = fsize(SOULLOG); loff = fsize(STUBLOG)
|
||||
t0 = time.time()
|
||||
pol = Poller(sid); pol.start()
|
||||
payload = {"message": ph["prompt"], "session_id": sid, "workspace_root": WS,
|
||||
"agentic": rec["agentic"]}
|
||||
if MODE == "local":
|
||||
payload["agent_workspace_root"] = WS
|
||||
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("chat_timeout", 240))
|
||||
rec["http"].append(st)
|
||||
rec["legs"].append(js if js is not None else body[:600])
|
||||
rec["progress_per_leg"].append(progress(sid))
|
||||
hops = 0
|
||||
while isinstance(js, dict) and js.get("tool_pending") and hops < MAXHOPS:
|
||||
rec["pendings"].append({"call_id": js.get("call_id"), "tool_name": js.get("tool_name"),
|
||||
"tool_input": js.get("tool_input"), "risk_tier": js.get("risk_tier"),
|
||||
"narration": js.get("narration"), "tools_used": js.get("tools_used")})
|
||||
try:
|
||||
eff = perform(js.get("tool_name", ""), js.get("tool_input") or {})
|
||||
except Exception as e:
|
||||
eff = "client error: %r" % (e,)
|
||||
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
|
||||
{"call_id": js.get("call_id"), "content": eff},
|
||||
timeout=CFG.get("chat_timeout", 240))
|
||||
rec["http"].append(st)
|
||||
rec["legs"].append(js if js is not None else body[:600])
|
||||
rec["progress_per_leg"].append(progress(sid))
|
||||
hops += 1
|
||||
pol.stop = True; time.sleep(0.3)
|
||||
t1 = time.time()
|
||||
rec["elapsed"] = round(t1 - t0, 2)
|
||||
rec["progress_during"] = pol.snaps
|
||||
rec["progress_final"] = progress(sid)
|
||||
rec["soul_log"], _ = tail_from(SOULLOG, soff)
|
||||
rec["stub"], _ = stub_since(loff)
|
||||
rec["hops"] = hops
|
||||
return rec
|
||||
|
||||
# ------------------------------------------------------------- assertions ----
|
||||
def expected_calls(cname, first_only=False):
|
||||
out = []
|
||||
for step in SCEN["classes"][cname]["script"]:
|
||||
for k, call in enumerate(step.get("tool_calls") or []):
|
||||
if first_only and k > 0:
|
||||
continue
|
||||
out.append((call["name"], call["arguments"]))
|
||||
return out
|
||||
|
||||
def final_text(cname):
|
||||
for step in reversed(SCEN["classes"][cname]["script"]):
|
||||
if step.get("text") and not step.get("tool_calls"):
|
||||
return step["text"]
|
||||
return None
|
||||
|
||||
def judge(rec):
|
||||
cname = rec["class"]; ok = []; bad = []
|
||||
last = rec["legs"][-1] if rec["legs"] else None
|
||||
reply = last.get("reply") if isinstance(last, dict) else None
|
||||
err = last.get("error") if isinstance(last, dict) else None
|
||||
tools_used = last.get("tools_used") if isinstance(last, dict) else None
|
||||
stub = rec["stub"]
|
||||
scen_recs = [r for r in stub if r.get("kind") == "scenario"]
|
||||
rejects = [r for r in stub if r.get("validation") != "ok"]
|
||||
bg = [r for r in stub if r.get("kind") in ("wrong_path", "background")]
|
||||
|
||||
def wire_clean():
|
||||
if rejects:
|
||||
for r in rejects:
|
||||
bad.append("stub REJECTED a request: [%s] %s"
|
||||
% (r.get("validation"), r.get("validation_detail")))
|
||||
else:
|
||||
ok.append("stub ground truth: validation \"ok\" on all %d scenario leg(s), no "
|
||||
"gate_echo_mismatch / gate_tool_call_shape / dialect-leak 400s"
|
||||
% len(scen_recs))
|
||||
if bg:
|
||||
ok.append("NOTE background non-scenario request(s) in this window: %s"
|
||||
% [(r.get("kind"), r.get("path"), r.get("http_status")) for r in bg])
|
||||
|
||||
if cname == "oa-plain":
|
||||
wire_clean()
|
||||
want = final_text(cname)
|
||||
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
|
||||
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
|
||||
if tools_used == []: ok.append("tools_used == [] (no tool ran)")
|
||||
else: bad.append("tools_used expected [] got %r" % (tools_used,))
|
||||
if reply and ('"tool_calls"' in reply or '"function"' in reply or '"tool_use"' in reply):
|
||||
bad.append("tool-call JSON leaked into the reply text")
|
||||
else: ok.append("no tool-call JSON anywhere in the reply")
|
||||
|
||||
elif cname in ("oa-single-tool", "oa-torture", "oa-mission"):
|
||||
wire_clean()
|
||||
want = final_text(cname)
|
||||
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
|
||||
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
|
||||
exp = expected_calls(cname)
|
||||
wantnames = [n for n, _ in exp]
|
||||
if tools_used == wantnames:
|
||||
ok.append("tools_used == %r (carried across %d suspension(s))" % (wantnames, rec["hops"]))
|
||||
else:
|
||||
bad.append("tools_used expected %r got %r" % (wantnames, tools_used))
|
||||
for name, args in exp:
|
||||
p = args.get("path"); c = args.get("content")
|
||||
dest = os.path.join(WS, p)
|
||||
if not os.path.exists(dest):
|
||||
bad.append("expected file missing on disk: %s" % dest); continue
|
||||
got = open(dest, "rb").read()
|
||||
if got == c.encode():
|
||||
ok.append("%s on disk is byte-for-byte the issued payload (%d bytes)" % (p, len(got)))
|
||||
else:
|
||||
bad.append("%s content differs\n WANT %r\n GOT %r"
|
||||
% (p, c[:300], got[:300].decode("utf-8", "replace")))
|
||||
if MODE == "bridge":
|
||||
for pend, (name, args) in zip(rec["pendings"], exp):
|
||||
if pend["tool_input"] == args:
|
||||
ok.append("tool_input for %s survived exactly ONE decode (deep-equal to the "
|
||||
"issued arguments; no double-escaping)" % name)
|
||||
else:
|
||||
bad.append("tool_input != issued arguments for %s\n WANT %r\n GOT %r"
|
||||
% (name, args, pend["tool_input"]))
|
||||
if rec["pendings"] and all(p["risk_tier"] == "escalate" for p in rec["pendings"]):
|
||||
ok.append("every write_file classified \"escalate\" and bridged for consent")
|
||||
|
||||
elif cname == "oa-parallel":
|
||||
drift = [l.strip() for l in rec["soul_log"].splitlines() if "DRIFT: provider returned" in l]
|
||||
if drift: ok.append("soul log: " + drift[0])
|
||||
else: bad.append("no 'DRIFT: provider returned N parallel tool_calls' line in the soul log")
|
||||
delivered = [r for r in stub if r.get("delivered", {}).get("tool_calls")]
|
||||
if delivered and len(delivered[0]["delivered"]["tool_calls"]) == 2:
|
||||
ok.append("stub delivered 2 parallel tool_calls in one response (ground truth)")
|
||||
if MODE == "bridge":
|
||||
if len(rec["pendings"]) == 1:
|
||||
ok.append("exactly ONE call honored: %s" % rec["pendings"][0]["call_id"])
|
||||
else:
|
||||
bad.append("expected exactly 1 honored call, got %d" % len(rec["pendings"]))
|
||||
pairing = [r for r in rejects if "gate_pairing" in str(r.get("validation_detail")) or
|
||||
"tool_calls at end of thread" in str(r.get("validation_detail")) or
|
||||
"not fully answered" in str(r.get("validation_detail"))]
|
||||
for r in pairing:
|
||||
ok.append("EXPECTED-BY-CONTRACT stub 400 on the unpaired echo: %s"
|
||||
% r.get("validation_detail"))
|
||||
other = [r for r in rejects if r not in pairing]
|
||||
for r in other:
|
||||
bad.append("unexpected stub rejection: [%s] %s"
|
||||
% (r.get("validation"), r.get("validation_detail")))
|
||||
if err and not reply:
|
||||
ok.append("honest error envelope after the 400 (no fabricated answer): %r" % err)
|
||||
elif reply == final_text(cname):
|
||||
ok.append("final reply == scripted final text (both calls paired)")
|
||||
else:
|
||||
bad.append("neither an honest error nor the scripted final text: %r" % (last,))
|
||||
|
||||
elif cname == "oa-api-error":
|
||||
if err and not reply:
|
||||
ok.append("honest error envelope: error=%r reply=%r" % (err, reply))
|
||||
else:
|
||||
bad.append("expected an error envelope with an empty reply, got %r" % (last,))
|
||||
delivered = [r["delivered"].get("api_error") for r in stub if r.get("delivered")]
|
||||
ok.append("stub delivered api_error status(es): %r" % [d for d in delivered if d])
|
||||
n = len([r for r in stub if r.get("kind") == "scenario"])
|
||||
ok.append("provider hit %d time(s) - no retry storm" % n)
|
||||
if reply:
|
||||
bad.append("FABRICATED ANSWER: reply non-empty on a provider error")
|
||||
|
||||
elif cname == "oa-tools-off":
|
||||
ok.append("stub records for this phrasing: %r"
|
||||
% [{k: r.get(k) for k in ("kind", "path", "validation", "http_status")} for r in stub])
|
||||
wrong = [r for r in stub if r.get("kind") == "wrong_path"]
|
||||
matched = [r for r in stub if r.get("scenario_class") == cname]
|
||||
if matched and not rejects:
|
||||
ok.append("chat-only request reached /v1/chat/completions with NO tools offered")
|
||||
want = final_text(cname)
|
||||
if reply == want: ok.append("final reply == scripted final text (byte-exact)")
|
||||
else: bad.append("final reply mismatch:\n WANT: %r\n GOT : %r" % (want, reply))
|
||||
if reply and ('"tool_calls"' in reply or '"function"' in reply):
|
||||
bad.append("tool-call JSON leaked into the reply text")
|
||||
else: ok.append("no tool-call JSON in the reply")
|
||||
elif wrong:
|
||||
bad.append("the non-agentic lane never reached the provider endpoint: stub saw "
|
||||
"%s -> %s (the el-runtime provider chain appends /v1/chat/completions "
|
||||
"to NEURON_LLM_0_URL, chat.el appends only /chat/completions)"
|
||||
% (wrong[0]["path"], wrong[0]["http_status"]))
|
||||
elif not stub:
|
||||
bad.append("no request reached the stub at all")
|
||||
else:
|
||||
for r in rejects:
|
||||
bad.append("stub REJECTED: [%s] %s" % (r.get("validation"), r.get("validation_detail")))
|
||||
return ok, bad
|
||||
|
||||
def main():
|
||||
results = []
|
||||
for cname in ONLY:
|
||||
for ph in SCEN["classes"][cname]["phrasings"]:
|
||||
rec = run_phrasing(cname, ph)
|
||||
ok, bad = judge(rec)
|
||||
rec["ok"] = ok; rec["bad"] = bad
|
||||
rec["verdict"] = "FAIL" if bad else "PASS"
|
||||
results.append(rec)
|
||||
print("=" * 78)
|
||||
print("[%s] %s / %s (%.2fs, %d bridge hop(s), agentic=%s, mode=%s)"
|
||||
% (rec["verdict"], cname, ph["id"], rec.get("elapsed", 0),
|
||||
rec.get("hops", 0), rec["agentic"], MODE))
|
||||
for l in ok: print(" ok " + l.replace("\n", "\n "))
|
||||
for l in bad: print(" FAIL " + l.replace("\n", "\n "))
|
||||
for i, leg in enumerate(rec["legs"]):
|
||||
print(" leg%d envelope: %s" % (i, json.dumps(leg)[:430]))
|
||||
for i, pr in enumerate(rec["progress_per_leg"]):
|
||||
print(" run-progress after leg%d: %s" % (i, json.dumps(pr)[:380]))
|
||||
if rec["progress_during"]:
|
||||
print(" run-progress polled DURING (%d distinct snapshot(s)), last: %s"
|
||||
% (len(rec["progress_during"]), json.dumps(rec["progress_during"][-1])[:300]))
|
||||
for r in rec["stub"]:
|
||||
print(" stub: kind=%s class=%s phrasing=%s step=%s validation=%s%s delivered=%s http=%s"
|
||||
% (r.get("kind"), r.get("scenario_class"), r.get("phrasing"), r.get("step"),
|
||||
r.get("validation"),
|
||||
("(" + str(r.get("validation_detail")) + ")") if r.get("validation_detail") else "",
|
||||
json.dumps(r.get("delivered")), r.get("http_status")))
|
||||
if rec["soul_log"].strip():
|
||||
for l in rec["soul_log"].splitlines():
|
||||
if l.strip(): print(" soul: " + l.strip())
|
||||
json.dump(results, open(OUT, "w"), indent=1)
|
||||
npass = sum(1 for r in results if r["verdict"] == "PASS")
|
||||
print("=" * 78)
|
||||
print("PHASE %s: %d/%d PASS" % (MODE, npass, len(results)))
|
||||
for r in results:
|
||||
print(" %-6s %-16s %s" % (r["verdict"], r["class"], r["phrasing"]))
|
||||
return 0 if npass == len(results) else 1
|
||||
|
||||
sys.exit(main())
|
||||
PYEOF
|
||||
|
||||
# ------------------------------------------------------------- hostile drv ---
|
||||
cat > "$RUN/hostile.py" <<'PYEOF'
|
||||
import json, os, sys, time, urllib.request, urllib.error
|
||||
|
||||
CFG = json.load(open(sys.argv[1]))
|
||||
SOUL = "http://127.0.0.1:%d" % CFG["soul_port"]
|
||||
STUB = "http://127.0.0.1:%d" % CFG["stub_port"]
|
||||
|
||||
def http(method, url, obj=None, timeout=400):
|
||||
data = None if obj is None else json.dumps(obj).encode()
|
||||
req = urllib.request.Request(url, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
b = r.read().decode("utf-8", "replace"); st = r.status
|
||||
except urllib.error.HTTPError as e:
|
||||
b = e.read().decode("utf-8", "replace"); st = e.code
|
||||
except Exception as e:
|
||||
return -1, "TRANSPORT-ERROR: %r" % (e,), None
|
||||
try:
|
||||
return st, b, json.loads(b)
|
||||
except ValueError:
|
||||
return st, b, None
|
||||
|
||||
mode = CFG["mode"]; wsmode = CFG["ws_mode"]; WS = CFG["workspace"]
|
||||
_, _, js = http("POST", SOUL + "/api/sessions", {"title": "hostile-" + mode}, timeout=60)
|
||||
sid = (js or {}).get("id", "")
|
||||
payload = {"message": "oa-gate plain probe: hostile mode %s" % mode,
|
||||
"agentic": True, "session_id": sid, "workspace_root": WS}
|
||||
if wsmode == "local":
|
||||
payload["agent_workspace_root"] = WS
|
||||
t0 = time.time()
|
||||
st, body, js = http("POST", SOUL + "/api/chat", payload, timeout=CFG.get("timeout", 400))
|
||||
t_first = time.time() - t0
|
||||
legs = [js if js is not None else body[:500]]
|
||||
hops = 0
|
||||
while isinstance(js, dict) and js.get("tool_pending") and hops < CFG.get("max_hops", 14):
|
||||
ti = js.get("tool_input") or {}
|
||||
p = ti.get("path", "x.md")
|
||||
dest = p if os.path.isabs(p) else os.path.join(WS, p)
|
||||
try: open(dest, "w").write(ti.get("content", ""))
|
||||
except Exception: pass
|
||||
st, body, js = http("POST", SOUL + "/api/sessions/%s/tool_result" % sid,
|
||||
{"call_id": js.get("call_id"), "content": "ok"},
|
||||
timeout=CFG.get("timeout", 400))
|
||||
legs.append(js if js is not None else body[:500]); hops += 1
|
||||
el = time.time() - t0
|
||||
_, _, stats = http("GET", STUB + "/gate/stats", timeout=30)
|
||||
_, _, prog = http("GET", SOUL + "/api/run-progress/" + sid, timeout=30)
|
||||
fab = [l for l in legs if isinstance(l, dict) and l.get("reply")]
|
||||
print("HOSTILE %s (ws_mode=%s)" % (mode, wsmode))
|
||||
print(" first /api/chat POST returned after %.2fs; whole chain %.2fs; client bridge hops=%d; "
|
||||
"stub chat_hits=%s" % (t_first, el, hops, (stats or {}).get("chat_hits")))
|
||||
print(" first envelope : " + json.dumps(legs[0])[:430])
|
||||
print(" final envelope : " + json.dumps(legs[-1])[:430])
|
||||
print(" non-empty replies anywhere in the chain (fabrication check): %d" % len(fab))
|
||||
print(" run-progress : " + json.dumps(prog)[:300])
|
||||
json.dump({"mode": mode, "ws_mode": wsmode, "t_first": t_first, "elapsed": el, "hops": hops,
|
||||
"chat_hits": (stats or {}).get("chat_hits"), "legs": legs, "progress": prog},
|
||||
open(CFG["out"], "w"), indent=1)
|
||||
PYEOF
|
||||
|
||||
# ------------------------------------------------------------------ phases ---
|
||||
RC_BRIDGE=0; RC_LOCAL=0; RC_OFF=0
|
||||
run_normal_phase() { # $1 = label, $2 = soul port, $3 = agent root, $4 = ws, $5 = base, $6 = classes json
|
||||
local m="$1" port="$2" root="$3" ws="$4" base="$5" classes="$6"
|
||||
echo; echo "############ PHASE: $m (soul :$port, NEURON_LLM_0_URL=$base) ############"
|
||||
start_stub normal "$RUN/stub-$m.jsonl"
|
||||
start_soul "$port" "$base" "$RUN/soul-$m.log" "$root"
|
||||
cat > "$RUN/cfg-$m.json" <<JSON
|
||||
{"soul_port": $port, "stub_port": $STUB_PORT, "workspace": "$ws", "mode": "$m",
|
||||
"scenarios": "$HERE/scenarios-openai.json", "stub_log": "$RUN/stub-$m.jsonl",
|
||||
"soul_log": "$RUN/soul-$m.log", "out": "$RUN/results-$m.json", "chat_timeout": 240,
|
||||
"classes": $classes}
|
||||
JSON
|
||||
python3 "$DRV" "$RUN/cfg-$m.json"
|
||||
local rc=$?
|
||||
stop_soul; stop_stub
|
||||
return $rc
|
||||
}
|
||||
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "bridge" ]; then
|
||||
run_normal_phase bridge "$SOUL_PORT" "" "$RUN/ws-bridge" "http://127.0.0.1:$STUB_PORT/v1" null
|
||||
RC_BRIDGE=$?
|
||||
fi
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "local" ]; then
|
||||
run_normal_phase local "$SOUL_PORT_B" "$RUN/ws-local" "$RUN/ws-local" "http://127.0.0.1:$STUB_PORT/v1" null
|
||||
RC_LOCAL=$?
|
||||
fi
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "toolsoff" ]; then
|
||||
# supplementary: the el-runtime provider chain appends /v1/chat/completions itself,
|
||||
# so the non-agentic door needs the base WITHOUT the /v1 suffix.
|
||||
run_normal_phase toolsoff "$SOUL_PORT" "" "$RUN/ws-off" "http://127.0.0.1:$STUB_PORT" '["oa-tools-off","oa-plain"]'
|
||||
RC_OFF=$?
|
||||
fi
|
||||
|
||||
if [ "$PHASES" = "all" ] || [ "$PHASES" = "hostile" ]; then
|
||||
echo; echo "############ PHASE: hostile ############"
|
||||
for spec in "black-hole:bridge" "mid-body-drop:bridge" "tool-pending-forever:bridge" "tool-pending-forever:local"; do
|
||||
mode="${spec%%:*}"; wsm="${spec##*:}"
|
||||
echo; echo "---- hostile mode=$mode ws_mode=$wsm ----"
|
||||
start_stub "$mode" "$RUN/stub-$mode-$wsm.jsonl"
|
||||
if [ "$wsm" = "local" ]; then
|
||||
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" "$RUN/ws-local"
|
||||
else
|
||||
start_soul "$SOUL_PORT" "http://127.0.0.1:$STUB_PORT/v1" "$RUN/soul-$mode-$wsm.log" ""
|
||||
fi
|
||||
cat > "$RUN/cfg-$mode-$wsm.json" <<JSON
|
||||
{"soul_port": $SOUL_PORT, "stub_port": $STUB_PORT, "mode": "$mode", "ws_mode": "$wsm",
|
||||
"workspace": "$RUN/ws-local", "out": "$RUN/hostile-$mode-$wsm.json", "timeout": 400}
|
||||
JSON
|
||||
python3 "$RUN/hostile.py" "$RUN/cfg-$mode-$wsm.json"
|
||||
echo " soul log (llm/DRIFT/cap lines):"
|
||||
grep -E "DRIFT|llm error|iteration cap|\[llm\]" "$RUN/soul-$mode-$wsm.log" | tail -8 | sed 's/^/ /'
|
||||
stop_soul; stop_stub
|
||||
done
|
||||
fi
|
||||
|
||||
echo; echo "############ CLEANUP ############"
|
||||
cleanup
|
||||
sleep 0.5
|
||||
echo "processes still matching the soul binary:"; pgrep -fl "$SOUL_BIN" || echo " (none)"
|
||||
echo "processes still matching stub-openai.py:"; pgrep -fl "stub-openai.py" || echo " (none)"
|
||||
echo "lsof on 7891-7894 after cleanup:"
|
||||
lsof -nP -iTCP:7891 -iTCP:7892 -iTCP:7893 -iTCP:7894 2>/dev/null || echo " (no listeners - ports free)"
|
||||
echo
|
||||
echo "############ SUMMARY ############"
|
||||
echo "run dir: $RUN"
|
||||
echo "bridge rc=$RC_BRIDGE local rc=$RC_LOCAL toolsoff rc=$RC_OFF (0 = every class PASS)"
|
||||
exit $(( RC_BRIDGE + RC_LOCAL + RC_OFF ))
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"_comment": "OpenAI-dialect gate scenario contract (soul-openai-tools-v2). Single source of truth shared by stub-openai.py (scripted provider responses + request assertions), selftest.sh (stub self-verification), and the future brain-side gate driver. Same structure as gate9's scenarios.json: classes -> script + phrasings with markers; scripts are CLASS-level so assertions are behavioral, never pinned to a sentence. expect_request keys: require_tools, require_tool_choice, parallel_tool_calls (expected literal value; null = don't check), forbid_tools. defaults apply to every class unless overridden; steps may override with their own expect_request.",
|
||||
"deadline_secs": 60,
|
||||
"max_loop_iterations": 16,
|
||||
"defaults": {
|
||||
"expect_request": {
|
||||
"require_tools": true,
|
||||
"require_tool_choice": true,
|
||||
"parallel_tool_calls": false
|
||||
}
|
||||
},
|
||||
"classes": {
|
||||
"oa-plain": {
|
||||
"script": [
|
||||
{ "text": "Plain OpenAI-lane answer (gate fixture): the mechanism, the main caveat, and the practical takeaway in three sentences. No tools were needed for this one, and the finish reason on the wire is stop, which the loop must treat as terminal." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-plain-p1", "marker": "oa-gate plain probe", "prompt": "oa-gate plain probe: explain the fixture topic simply." },
|
||||
{ "id": "oa-plain-p2", "marker": "oa-gate second plain", "prompt": "oa-gate second plain: another phrasing of the plain question." }
|
||||
]
|
||||
},
|
||||
"oa-tools-off": {
|
||||
"expect_request": {
|
||||
"require_tools": false,
|
||||
"forbid_tools": true,
|
||||
"require_tool_choice": false,
|
||||
"parallel_tool_calls": null
|
||||
},
|
||||
"script": [
|
||||
{ "text": "Chat-only OpenAI-lane answer (gate fixture): this lane offered no tools and none were used; the reply is plain text with finish reason stop." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-tools-off-p1", "marker": "oa-gate tools-off probe", "prompt": "oa-gate tools-off probe: plain chat with no tools offered." }
|
||||
]
|
||||
},
|
||||
"oa-single-tool": {
|
||||
"script": [
|
||||
{ "text": "Step 1: writing the note.",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file",
|
||||
"arguments": { "path": "openai-single-note.md", "content": "# Note (gate fixture, OpenAI lane)\n\nDeterministic single-tool body.\n" } }
|
||||
] },
|
||||
{ "text": "All set - openai-single-note.md is written with the fixture body. Nothing else was needed for this one." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-single-p1", "marker": "oa-gate single tool note", "prompt": "oa-gate single tool note: save the fixture note to a file." },
|
||||
{ "id": "oa-single-p2", "marker": "oa-gate one file please", "prompt": "oa-gate one file please: write the fixture note file." }
|
||||
]
|
||||
},
|
||||
"oa-torture": {
|
||||
"script": [
|
||||
{ "tool_calls": [
|
||||
{ "name": "write_file",
|
||||
"arguments": { "path": "torture-note.md", "content": "Line 1 has \"double quotes\", 'singles', and a mid-line backslash \\ here.\nLine 2\thas a tab, a literal \\n two-char sequence, and a Windows path C:\\temp\\new.txt.\nLine 3 unicode: naïve café — 日本語 ✓ 🚀\nLine 4 JSON-in-string: {\"k\": \"v\", \"arr\": [1, 2], \"s\": \"nested \\\"deep\\\" quotes\"}\nLine 5 ends with a lone backslash \\" } }
|
||||
] },
|
||||
{ "text": "Torture round-trip complete: the payload with nested quotes, backslashes, newlines, tabs, and unicode survived exactly one encode and one decode." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-torture-p1", "marker": "oa-gate torture probe", "prompt": "oa-gate torture probe: write the escaping torture file." }
|
||||
]
|
||||
},
|
||||
"oa-parallel": {
|
||||
"script": [
|
||||
{ "text": "Step 1: two writes at once (parallel probe).",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file", "arguments": { "path": "parallel-a.md", "content": "Parallel A (gate fixture).\n" } },
|
||||
{ "name": "write_file", "arguments": { "path": "parallel-b.md", "content": "Parallel B (gate fixture).\n" } }
|
||||
] },
|
||||
{ "text": "Parallel probe complete: both tool results arrived and were paired correctly. A brain that instead rejects the double call must do so cleanly - that outcome is asserted brain-side, not here." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-parallel-p1", "marker": "oa-gate parallel probe", "prompt": "oa-gate parallel probe: run the two-write parallel case." }
|
||||
]
|
||||
},
|
||||
"oa-mission": {
|
||||
"script": [
|
||||
{ "text": "Step 1: drafting part one.",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file", "arguments": { "path": "mission-part-1.md", "content": "Mission part 1 (gate fixture).\n" } }
|
||||
] },
|
||||
{ "text": "Step 2: drafting part two.",
|
||||
"tool_calls": [
|
||||
{ "name": "write_file", "arguments": { "path": "mission-part-2.md", "content": "Mission part 2 (gate fixture).\n" } }
|
||||
] },
|
||||
{ "text": "Mission complete: mission-part-1.md and mission-part-2.md are written; the loop ran two tool rounds and finished cleanly with finish reason stop." }
|
||||
],
|
||||
"phrasings": [
|
||||
{ "id": "oa-mission-p1", "marker": "oa-gate mission probe", "prompt": "oa-gate mission probe: run the two-round mission." }
|
||||
]
|
||||
},
|
||||
"oa-api-error": {
|
||||
"expect_request": {
|
||||
"require_tools": false,
|
||||
"require_tool_choice": false,
|
||||
"parallel_tool_calls": null
|
||||
},
|
||||
"script": [],
|
||||
"phrasings": [
|
||||
{ "id": "oa-err-400", "marker": "oa-gate error four hundred", "prompt": "oa-gate error four hundred: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 400, "type": "invalid_request_error", "message": "gate-injected 400: request rejected by fixture", "code": "gate_injected" } } ] },
|
||||
{ "id": "oa-err-429", "marker": "oa-gate error rate limit", "prompt": "oa-gate error rate limit: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 429, "type": "rate_limit_error", "message": "gate-injected 429: rate limited by fixture", "code": "rate_limit_exceeded" } } ] },
|
||||
{ "id": "oa-err-500", "marker": "oa-gate error five hundred", "prompt": "oa-gate error five hundred: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 500, "type": "server_error", "message": "gate-injected 500: internal fixture error", "code": "gate_injected" } } ] },
|
||||
{ "id": "oa-err-503", "marker": "oa-gate error unavailable", "prompt": "oa-gate error unavailable: trigger the injected failure.",
|
||||
"script": [ { "api_error": { "status": 503, "type": "server_error", "message": "gate-injected 503: fixture overloaded", "code": "gate_injected" } } ] }
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
Executable
+409
@@ -0,0 +1,409 @@
|
||||
#!/usr/bin/env bash
|
||||
# selftest.sh - proves stub-openai.py before any brain code exists.
|
||||
# Drives the stub with curl through every scenario (plain, tools-off,
|
||||
# single tool round-trip, escaping torture, parallel double-call,
|
||||
# two-round mission, injected API errors, background, overrun), every
|
||||
# validation rejection (dialect leaks, pairing, echo round-trip, scenario
|
||||
# expectations), and all three hostile modes. Exit 0 = green.
|
||||
set -u
|
||||
cd "$(dirname "$0")" || exit 1
|
||||
PY=python3
|
||||
TMP="$(mktemp -d)"
|
||||
PIDS=()
|
||||
cleanup() {
|
||||
for p in "${PIDS[@]:-}"; do kill -9 "$p" >/dev/null 2>&1; done
|
||||
rm -rf "$TMP"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
PASS=0; FAIL=0
|
||||
ok() { printf 'ok - %s\n' "$1"; PASS=$((PASS+1)); }
|
||||
bad() { printf 'FAIL - %s\n' "$1"; FAIL=$((FAIL+1)); }
|
||||
check() { # check <name> <cmd...> - pass if cmd exits 0; show output on fail
|
||||
local name="$1"; shift
|
||||
local out
|
||||
if out="$("$@" 2>&1)"; then ok "$name"
|
||||
else bad "$name"; [ -n "$out" ] && printf '%s\n' "$out" | sed 's/^/ /' | head -8
|
||||
fi
|
||||
}
|
||||
|
||||
freeport() { "$PY" -c 'import socket;s=socket.socket();s.bind(("127.0.0.1",0));print(s.getsockname()[1]);s.close()'; }
|
||||
waithealth() {
|
||||
local p="$1" i
|
||||
for i in $(seq 1 60); do
|
||||
curl -sf "http://127.0.0.1:$p/gate/health" >/dev/null 2>&1 && return 0
|
||||
sleep 0.1
|
||||
done
|
||||
echo "stub on :$p never became healthy"; return 1
|
||||
}
|
||||
post() { # post <port> <bodyfile> <respfile> [extra curl args...] -> echoes http code
|
||||
local port="$1" body="$2" resp="$3"; shift 3
|
||||
curl -s -o "$resp" -w '%{http_code}' -H 'content-type: application/json' \
|
||||
"$@" --data-binary @"$body" "http://127.0.0.1:$port/v1/chat/completions"
|
||||
}
|
||||
|
||||
# ---- embedded helper: builds OpenAI-dialect bodies, asserts on responses ----
|
||||
cat > "$TMP/helpers.py" <<'PYEOF'
|
||||
import copy, json, sys
|
||||
|
||||
TOOLS = [
|
||||
{"type": "function", "function": {
|
||||
"name": "write_file", "description": "Write content to a file on disk.",
|
||||
"parameters": {"type": "object",
|
||||
"properties": {"path": {"type": "string"},
|
||||
"content": {"type": "string"}},
|
||||
"required": ["path", "content"]}}},
|
||||
{"type": "function", "function": {
|
||||
"name": "read_file", "description": "Read contents of a file from disk.",
|
||||
"parameters": {"type": "object",
|
||||
"properties": {"path": {"type": "string"}},
|
||||
"required": ["path"]}}},
|
||||
]
|
||||
|
||||
def dump(obj, out):
|
||||
json.dump(obj, open(out, "w"), ensure_ascii=False)
|
||||
|
||||
def base(prompt, tools=True):
|
||||
b = {"model": "gate-openai-model", "max_tokens": 1024,
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are Neuron (gate fixture)."},
|
||||
{"role": "user", "content": prompt}]}
|
||||
if tools:
|
||||
b["tools"] = copy.deepcopy(TOOLS)
|
||||
b["tool_choice"] = "auto"
|
||||
b["parallel_tool_calls"] = False
|
||||
return b
|
||||
|
||||
def cmd_plain(out, prompt):
|
||||
dump(base(prompt), out)
|
||||
|
||||
def cmd_notools(out, prompt):
|
||||
dump(base(prompt, tools=False), out)
|
||||
|
||||
def cmd_mut(out, prompt, mutation):
|
||||
b = base(prompt)
|
||||
if mutation == "no-tool-choice":
|
||||
del b["tool_choice"]
|
||||
elif mutation == "ptc-true":
|
||||
b["parallel_tool_calls"] = True
|
||||
elif mutation == "top-system":
|
||||
b["system"] = "You are Neuron."
|
||||
elif mutation == "anth-tools":
|
||||
b["tools"] = [{"name": "write_file", "description": "x",
|
||||
"input_schema": {"type": "object", "properties": {}}}]
|
||||
elif mutation == "anth-block":
|
||||
b["messages"][1] = {"role": "user", "content": [
|
||||
{"type": "tool_result", "tool_use_id": "toolu_x", "content": "hi"},
|
||||
{"type": "text", "text": prompt}]}
|
||||
else:
|
||||
raise SystemExit("unknown mutation " + mutation)
|
||||
dump(b, out)
|
||||
|
||||
def cmd_chain(out, prompt, variant, *resps):
|
||||
"""Build the next leg: echo each response's assistant turn and answer its
|
||||
tool calls. `variant` applies to the LAST response only:
|
||||
ok | no-tool-turn | wrong-id | only-first | double-encode | object-args"""
|
||||
b = base(prompt)
|
||||
for idx, p in enumerate(resps):
|
||||
last = idx == len(resps) - 1
|
||||
msg = json.load(open(p))["choices"][0]["message"]
|
||||
tcs = msg.get("tool_calls")
|
||||
if not tcs:
|
||||
b["messages"].append({"role": "assistant",
|
||||
"content": msg.get("content")})
|
||||
continue
|
||||
v = variant if last else "ok"
|
||||
asst = {"role": "assistant", "content": msg.get("content"),
|
||||
"tool_calls": copy.deepcopy(tcs)}
|
||||
if v == "double-encode":
|
||||
for tc in asst["tool_calls"]:
|
||||
tc["function"]["arguments"] = json.dumps(
|
||||
tc["function"]["arguments"])
|
||||
if v == "object-args":
|
||||
for tc in asst["tool_calls"]:
|
||||
tc["function"]["arguments"] = json.loads(
|
||||
tc["function"]["arguments"])
|
||||
b["messages"].append(asst)
|
||||
if v == "no-tool-turn":
|
||||
continue
|
||||
use = tcs[:1] if v == "only-first" else tcs
|
||||
for tc in use:
|
||||
tid = "call_bogus_123" if v == "wrong-id" else tc["id"]
|
||||
b["messages"].append({"role": "tool", "tool_call_id": tid,
|
||||
"content": "{\"ok\":true,\"bytes\":42}"})
|
||||
dump(b, out)
|
||||
|
||||
def cmd_chk(resp, expr):
|
||||
r = json.load(open(resp))
|
||||
if not eval(expr, {"r": r, "json": json, "len": len, "str": str,
|
||||
"isinstance": isinstance, "any": any, "all": all,
|
||||
"sorted": sorted}):
|
||||
print("assertion failed:", expr)
|
||||
print("resp:", json.dumps(r, ensure_ascii=False)[:400])
|
||||
raise SystemExit(1)
|
||||
|
||||
def cmd_torture(resp, scen):
|
||||
r = json.load(open(resp))
|
||||
tc = r["choices"][0]["message"]["tool_calls"][0]
|
||||
raw = tc["function"]["arguments"]
|
||||
assert isinstance(raw, str), "arguments must be a JSON-encoded string"
|
||||
got = json.loads(raw)
|
||||
exp = json.load(open(scen))["classes"]["oa-torture"]["script"][0]["tool_calls"][0]["arguments"]
|
||||
assert got == exp, "decoded arguments != scripted torture payload"
|
||||
content = got["content"]
|
||||
for needle in ['"', "\\", "\n", "\t", "日本語", "naïve", "🚀"]:
|
||||
assert needle in content, "missing torture needle %r" % needle
|
||||
|
||||
def cmd_notjson(path):
|
||||
data = open(path, "rb").read()
|
||||
assert data, "file empty - no partial body arrived"
|
||||
try:
|
||||
json.loads(data.decode("utf-8", "replace"))
|
||||
except ValueError:
|
||||
return
|
||||
raise SystemExit("partial body unexpectedly parsed as complete JSON")
|
||||
|
||||
def cmd_pending(*paths):
|
||||
ids = []
|
||||
for p in paths:
|
||||
c = json.load(open(p))["choices"][0]
|
||||
assert c["finish_reason"] == "tool_calls", c["finish_reason"]
|
||||
tc = c["message"]["tool_calls"][0]
|
||||
assert tc["function"]["name"] == "write_file"
|
||||
json.loads(tc["function"]["arguments"]) # must decode
|
||||
ids.append(tc["id"])
|
||||
assert len(set(ids)) == len(ids), "call ids not distinct: %r" % ids
|
||||
|
||||
def cmd_logcheck(path):
|
||||
recs = [json.loads(l) for l in open(path) if l.strip()]
|
||||
seqs = [r["seq"] for r in recs]
|
||||
assert seqs == sorted(seqs) and len(set(seqs)) == len(seqs), "seq not monotonic"
|
||||
kinds = {}
|
||||
for r in recs:
|
||||
kinds[r["kind"]] = kinds.get(r["kind"], 0) + 1
|
||||
assert kinds.get("scenario", 0) >= 10, "too few scenario records: %r" % kinds
|
||||
assert kinds.get("background", 0) >= 1, "no background record"
|
||||
assert kinds.get("overrun", 0) >= 1, "no overrun record"
|
||||
rejected = [r for r in recs if r["validation"] == "rejected"]
|
||||
assert len(rejected) >= 10, "too few rejected records: %d" % len(rejected)
|
||||
assert any(r["delivered"].get("tool_calls") == ["write_file"]
|
||||
for r in recs), "no single write_file ground truth"
|
||||
assert any(r["delivered"].get("tool_calls") == ["write_file", "write_file"]
|
||||
for r in recs), "no parallel ground truth"
|
||||
|
||||
def main():
|
||||
fn = globals()["cmd_" + sys.argv[1].replace("-", "_")]
|
||||
fn(*sys.argv[2:])
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
PYEOF
|
||||
mk() { "$PY" "$TMP/helpers.py" "$@"; }
|
||||
|
||||
echo "=== stub-openai selftest ==="
|
||||
|
||||
# ---- normal mode ------------------------------------------------------------
|
||||
PORT="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$PORT" --scenarios scenarios-openai.json \
|
||||
--log "$TMP/req.jsonl" >"$TMP/stub.out" 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "stub starts and answers /gate/health" waithealth "$PORT"
|
||||
|
||||
# 1. plain completion
|
||||
mk plain "$TMP/plain.json" "oa-gate plain probe: explain the fixture topic simply."
|
||||
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_plain.json")"
|
||||
check "plain: HTTP 200" test "$code" = "200"
|
||||
check "plain: chat.completion envelope, finish stop, real content" mk chk "$TMP/r_plain.json" \
|
||||
'r["object"]=="chat.completion" and r["choices"][0]["finish_reason"]=="stop" and isinstance(r["choices"][0]["message"]["content"],str) and len(r["choices"][0]["message"]["content"])>40'
|
||||
|
||||
# 2. tools-off lane (chat-only request accepted, tool-bearing request refused)
|
||||
mk notools "$TMP/toolsoff.json" "oa-gate tools-off probe: plain chat with no tools offered."
|
||||
code="$(post "$PORT" "$TMP/toolsoff.json" "$TMP/r_toolsoff.json")"
|
||||
check "tools-off: chat-only request -> 200" test "$code" = "200"
|
||||
mk plain "$TMP/toolsoff_bad.json" "oa-gate tools-off probe: plain chat with no tools offered."
|
||||
code="$(post "$PORT" "$TMP/toolsoff_bad.json" "$TMP/r_toolsoff_bad.json")"
|
||||
check "tools-off negative: offering tools -> 400 gate_expect" \
|
||||
bash -c "test $code = 400"
|
||||
check "tools-off negative: reason names gate_expect" mk chk "$TMP/r_toolsoff_bad.json" \
|
||||
'r["error"]["code"]=="gate_expect"'
|
||||
|
||||
# 3. dialect-leak rejections (the loud-failure contract)
|
||||
code="$(post "$PORT" "$TMP/plain.json" "$TMP/r_leak_hdr.json" -H 'anthropic-version: 2023-06-01')"
|
||||
check "leak: anthropic-version header -> 400" test "$code" = "400"
|
||||
check "leak: header reason names the leak" mk chk "$TMP/r_leak_hdr.json" \
|
||||
'r["error"]["code"]=="gate_dialect_leak" and "anthropic-version" in r["error"]["message"]'
|
||||
mk mut "$TMP/leak_tools.json" "oa-gate plain probe: explain the fixture topic simply." anth-tools
|
||||
code="$(post "$PORT" "$TMP/leak_tools.json" "$TMP/r_leak_tools.json")"
|
||||
check "leak: input_schema tools -> 400 gate_dialect_leak" bash -c \
|
||||
"test $code = 400"
|
||||
check "leak: input_schema reason" mk chk "$TMP/r_leak_tools.json" \
|
||||
'r["error"]["code"]=="gate_dialect_leak" and "input_schema" in r["error"]["message"]'
|
||||
mk mut "$TMP/leak_sys.json" "oa-gate plain probe: explain the fixture topic simply." top-system
|
||||
code="$(post "$PORT" "$TMP/leak_sys.json" "$TMP/r_leak_sys.json")"
|
||||
check "leak: top-level system -> 400" test "$code" = "400"
|
||||
mk mut "$TMP/leak_block.json" "oa-gate plain probe: explain the fixture topic simply." anth-block
|
||||
code="$(post "$PORT" "$TMP/leak_block.json" "$TMP/r_leak_block.json")"
|
||||
check "leak: Anthropic tool_result content block -> 400" test "$code" = "400"
|
||||
|
||||
# 4. scenario request expectations
|
||||
mk mut "$TMP/no_tc.json" "oa-gate plain probe: explain the fixture topic simply." no-tool-choice
|
||||
code="$(post "$PORT" "$TMP/no_tc.json" "$TMP/r_no_tc.json")"
|
||||
check "expect: missing tool_choice -> 400" test "$code" = "400"
|
||||
mk mut "$TMP/ptc.json" "oa-gate plain probe: explain the fixture topic simply." ptc-true
|
||||
code="$(post "$PORT" "$TMP/ptc.json" "$TMP/r_ptc.json")"
|
||||
check "expect: parallel_tool_calls true -> 400 (ADR-0005 pin)" test "$code" = "400"
|
||||
|
||||
# 5. single tool round-trip
|
||||
ST_PROMPT="oa-gate single tool note: save the fixture note to a file."
|
||||
mk plain "$TMP/st1.json" "$ST_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/st1.json" "$TMP/r_st1.json")"
|
||||
check "single-tool leg1: HTTP 200" test "$code" = "200"
|
||||
check "single-tool leg1: one write_file call, finish tool_calls, string args" mk chk "$TMP/r_st1.json" \
|
||||
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==1 and r["choices"][0]["message"]["tool_calls"][0]["type"]=="function" and r["choices"][0]["message"]["tool_calls"][0]["function"]["name"]=="write_file" and isinstance(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"],str) and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="openai-single-note.md"'
|
||||
mk chain "$TMP/st2.json" "$ST_PROMPT" ok "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2.json" "$TMP/r_st2.json")"
|
||||
check "single-tool leg2: echo + tool turn -> 200 final text" test "$code" = "200"
|
||||
check "single-tool leg2: final names the file, finish stop" mk chk "$TMP/r_st2.json" \
|
||||
'r["choices"][0]["finish_reason"]=="stop" and "openai-single-note.md" in r["choices"][0]["message"]["content"]'
|
||||
mk chain "$TMP/st2_no.json" "$ST_PROMPT" no-tool-turn "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2_no.json" "$TMP/r_st2_no.json")"
|
||||
check "single-tool negative: echo without tool turn -> 400 gate_pairing" \
|
||||
bash -c "test $code = 400"
|
||||
check "single-tool negative: pairing reason" mk chk "$TMP/r_st2_no.json" \
|
||||
'r["error"]["code"]=="gate_pairing"'
|
||||
mk chain "$TMP/st2_wrong.json" "$ST_PROMPT" wrong-id "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2_wrong.json" "$TMP/r_st2_wrong.json")"
|
||||
check "single-tool negative: wrong tool_call_id -> 400" test "$code" = "400"
|
||||
mk chain "$TMP/st2_obj.json" "$ST_PROMPT" object-args "$TMP/r_st1.json"
|
||||
code="$(post "$PORT" "$TMP/st2_obj.json" "$TMP/r_st2_obj.json")"
|
||||
check "single-tool negative: arguments echoed as object -> 400 shape" \
|
||||
bash -c "test $code = 400"
|
||||
check "single-tool negative: shape reason names STRING" mk chk "$TMP/r_st2_obj.json" \
|
||||
'r["error"]["code"]=="gate_tool_call_shape" and "STRING" in r["error"]["message"]'
|
||||
|
||||
# 6. escaping torture (the two-escaper trap, spec section 6)
|
||||
T_PROMPT="oa-gate torture probe: write the escaping torture file."
|
||||
mk plain "$TMP/t1.json" "$T_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/t1.json" "$TMP/r_t1.json")"
|
||||
check "torture leg1: HTTP 200" test "$code" = "200"
|
||||
check "torture leg1: arguments decode to the exact nasty payload" \
|
||||
mk torture "$TMP/r_t1.json" scenarios-openai.json
|
||||
mk chain "$TMP/t2.json" "$T_PROMPT" ok "$TMP/r_t1.json"
|
||||
code="$(post "$PORT" "$TMP/t2.json" "$TMP/r_t2.json")"
|
||||
check "torture leg2: faithful echo -> 200 final" test "$code" = "200"
|
||||
mk chain "$TMP/t2_dbl.json" "$T_PROMPT" double-encode "$TMP/r_t1.json"
|
||||
code="$(post "$PORT" "$TMP/t2_dbl.json" "$TMP/r_t2_dbl.json")"
|
||||
check "torture negative: double-encoded echo -> 400" test "$code" = "400"
|
||||
check "torture negative: reason names the two-escaper trap" mk chk "$TMP/r_t2_dbl.json" \
|
||||
'r["error"]["code"]=="gate_echo_mismatch" and "two-escaper" in r["error"]["message"]'
|
||||
|
||||
# 7. parallel double-call
|
||||
P_PROMPT="oa-gate parallel probe: run the two-write parallel case."
|
||||
mk plain "$TMP/p1.json" "$P_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/p1.json" "$TMP/r_p1.json")"
|
||||
check "parallel leg1: TWO tool_calls, distinct ids" mk chk "$TMP/r_p1.json" \
|
||||
'r["choices"][0]["finish_reason"]=="tool_calls" and len(r["choices"][0]["message"]["tool_calls"])==2 and r["choices"][0]["message"]["tool_calls"][0]["id"]!=r["choices"][0]["message"]["tool_calls"][1]["id"]'
|
||||
mk chain "$TMP/p2.json" "$P_PROMPT" ok "$TMP/r_p1.json"
|
||||
code="$(post "$PORT" "$TMP/p2.json" "$TMP/r_p2.json")"
|
||||
check "parallel leg2: both results -> 200 final" test "$code" = "200"
|
||||
mk chain "$TMP/p2_one.json" "$P_PROMPT" only-first "$TMP/r_p1.json"
|
||||
code="$(post "$PORT" "$TMP/p2_one.json" "$TMP/r_p2_one.json")"
|
||||
check "parallel negative: answering only one call -> 400 pairing" test "$code" = "400"
|
||||
|
||||
# 8. two-round mission (loop continuation + step indexing)
|
||||
M_PROMPT="oa-gate mission probe: run the two-round mission."
|
||||
mk plain "$TMP/m1.json" "$M_PROMPT"
|
||||
code="$(post "$PORT" "$TMP/m1.json" "$TMP/r_m1.json")"
|
||||
check "mission leg1: part-1 tool call" mk chk "$TMP/r_m1.json" \
|
||||
'json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-1.md"'
|
||||
mk chain "$TMP/m2.json" "$M_PROMPT" ok "$TMP/r_m1.json"
|
||||
code="$(post "$PORT" "$TMP/m2.json" "$TMP/r_m2.json")"
|
||||
check "mission leg2: part-2 tool call (step indexed by assistant count)" mk chk "$TMP/r_m2.json" \
|
||||
'r["choices"][0]["finish_reason"]=="tool_calls" and json.loads(r["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])["path"]=="mission-part-2.md"'
|
||||
mk chain "$TMP/m3.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json"
|
||||
code="$(post "$PORT" "$TMP/m3.json" "$TMP/r_m3.json")"
|
||||
check "mission leg3: final text, finish stop" mk chk "$TMP/r_m3.json" \
|
||||
'r["choices"][0]["finish_reason"]=="stop" and "Mission complete" in r["choices"][0]["message"]["content"]'
|
||||
mk chain "$TMP/m4.json" "$M_PROMPT" ok "$TMP/r_m1.json" "$TMP/r_m2.json" "$TMP/r_m3.json"
|
||||
code="$(post "$PORT" "$TMP/m4.json" "$TMP/r_m4.json")"
|
||||
check "mission overrun: past-script request -> GATE-SCRIPT-EXHAUSTED" mk chk "$TMP/r_m4.json" \
|
||||
'r["choices"][0]["message"]["content"].startswith("GATE-SCRIPT-EXHAUSTED")'
|
||||
|
||||
# 9. injected API errors (OpenAI error envelope)
|
||||
for want in 400 429 500 503; do
|
||||
case "$want" in
|
||||
400) marker="four hundred";; 429) marker="rate limit";;
|
||||
500) marker="five hundred";; 503) marker="unavailable";;
|
||||
esac
|
||||
mk plain "$TMP/e_$want.json" "oa-gate error $marker: trigger the injected failure."
|
||||
code="$(post "$PORT" "$TMP/e_$want.json" "$TMP/r_e_$want.json")"
|
||||
check "api-error $want: status returned" test "$code" = "$want"
|
||||
check "api-error $want: OpenAI error envelope" mk chk "$TMP/r_e_$want.json" \
|
||||
'isinstance(r["error"]["message"],str) and "gate-injected" in r["error"]["message"] and isinstance(r["error"]["type"],str)'
|
||||
done
|
||||
|
||||
# 10. background (unmatched) request
|
||||
mk plain "$TMP/bg.json" "hello there, just a boot probe with no marker"
|
||||
code="$(post "$PORT" "$TMP/bg.json" "$TMP/r_bg.json")"
|
||||
check "background: unmatched prompt -> benign ok" mk chk "$TMP/r_bg.json" \
|
||||
'r["choices"][0]["message"]["content"]=="ok"'
|
||||
|
||||
# 11. ground-truth log invariants
|
||||
check "ground-truth JSONL log invariants" mk logcheck "$TMP/req.jsonl"
|
||||
|
||||
# 12. production-port refusal
|
||||
rc=0
|
||||
"$PY" stub-openai.py --port 7770 --scenarios scenarios-openai.json \
|
||||
--log "$TMP/never.jsonl" >/dev/null 2>&1 || rc=$?
|
||||
check "refuses production port 7770" test "$rc" -ne 0
|
||||
|
||||
# ---- hostile mode: black-hole ----------------------------------------------
|
||||
BH="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$BH" --log "$TMP/bh.jsonl" --mode black-hole \
|
||||
>/dev/null 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "black-hole: healthy" waithealth "$BH"
|
||||
rc=0
|
||||
curl -s -o /dev/null --max-time 3 -H 'content-type: application/json' \
|
||||
--data-binary @"$TMP/plain.json" \
|
||||
"http://127.0.0.1:$BH/v1/chat/completions" || rc=$?
|
||||
check "black-hole: client times out (curl rc 28)" test "$rc" -eq 28
|
||||
check "black-hole: health still answers during the hang" \
|
||||
curl -sf --max-time 2 "http://127.0.0.1:$BH/gate/health"
|
||||
|
||||
# ---- hostile mode: mid-body-drop -------------------------------------------
|
||||
MD="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$MD" --log "$TMP/md.jsonl" --mode mid-body-drop \
|
||||
>/dev/null 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "mid-body-drop: healthy" waithealth "$MD"
|
||||
rc=0
|
||||
curl -s --max-time 5 -o "$TMP/half.json" -H 'content-type: application/json' \
|
||||
--data-binary @"$TMP/plain.json" \
|
||||
"http://127.0.0.1:$MD/v1/chat/completions" || rc=$?
|
||||
check "mid-body-drop: transfer fails (curl rc $rc)" test "$rc" -ne 0
|
||||
check "mid-body-drop: partial body is not parseable JSON" mk notjson "$TMP/half.json"
|
||||
|
||||
# ---- hostile mode: tool-pending-forever ------------------------------------
|
||||
TP="$(freeport)"
|
||||
"$PY" stub-openai.py --port "$TP" --log "$TMP/tp.jsonl" \
|
||||
--mode tool-pending-forever >/dev/null 2>&1 &
|
||||
PIDS+=($!); disown
|
||||
check "tool-pending-forever: healthy" waithealth "$TP"
|
||||
for i in 1 2 3; do
|
||||
code="$(post "$TP" "$TMP/plain.json" "$TMP/r_tp$i.json")"
|
||||
check "tool-pending-forever: request $i -> 200" test "$code" = "200"
|
||||
done
|
||||
check "tool-pending-forever: three FRESH tool_calls, distinct ids" \
|
||||
mk pending "$TMP/r_tp1.json" "$TMP/r_tp2.json" "$TMP/r_tp3.json"
|
||||
check "tool-pending-forever: /gate/stats counts 3 chat hits" \
|
||||
bash -c "curl -sf http://127.0.0.1:$TP/gate/stats | grep -q '\"chat_hits\": 3'"
|
||||
|
||||
# ---- summary ----------------------------------------------------------------
|
||||
echo
|
||||
echo "selftest: $PASS passed, $FAIL failed"
|
||||
if [ "$FAIL" -ne 0 ]; then
|
||||
echo "SELFTEST RED"
|
||||
exit 1
|
||||
fi
|
||||
echo "SELFTEST GREEN (stub-openai gate scaffolding verified)"
|
||||
Executable
+652
@@ -0,0 +1,652 @@
|
||||
#!/usr/bin/env python3
|
||||
"""stub-openai.py - deterministic local stand-in for an OpenAI-format
|
||||
/v1/chat/completions provider, for the soul-openai-tools-v2 gate
|
||||
(docs/specs/SPEC-soul-openai-tools-v2-2026-08-06.md). No API key, no network,
|
||||
no model.
|
||||
|
||||
Sibling of gate9's stub-llm.py (Anthropic dialect, _wt-beta-round9/scripts/
|
||||
gate9/): same scenario mechanism (marker matching, assistant-count step
|
||||
indexing, ground-truth JSONL log, prod-port refusal), different wire.
|
||||
Staging home is tests/gate-openai/ in _wt-openai-tools; folds into
|
||||
scripts/gate9/ after round 9 merges (see README.md).
|
||||
|
||||
WHAT IT DOES
|
||||
* Serves POST /v1/chat/completions on 127.0.0.1 only (OpenAI dialect).
|
||||
* VALIDATES every request - this is the gate's discriminator, built
|
||||
BEFORE the brain-side El code exists so dialect leakage fails loudly:
|
||||
- Anthropic tells are 400 code=gate_dialect_leak: `anthropic-version`
|
||||
header; top-level `system` / `stop_sequences` / `max_tokens_to_sample`
|
||||
/ `anthropic_version`; `input_schema` inside a tool entry; Anthropic
|
||||
content blocks (tool_use / tool_result / server_tool_use / ...).
|
||||
- tools[] must be OpenAI-shaped {type:"function", function:{name,
|
||||
description, parameters}} with unique names -> 400 gate_tools_shape.
|
||||
- assistant tool_calls echoes must be {id, type:"function",
|
||||
function:{name, arguments:<JSON-encoded STRING>}}; a decoded-object
|
||||
`arguments` is a wire bug -> 400 gate_tool_call_shape.
|
||||
- every assistant tool_calls turn must be answered by role:"tool"
|
||||
messages covering EVERY tool_call_id, immediately following;
|
||||
unknown / duplicate / missing ids -> 400 gate_pairing.
|
||||
- echoed `arguments` for gate-issued call ids (call_gate_*) are
|
||||
recomputed from the script and compared after ONE json decode ->
|
||||
400 gate_echo_mismatch. This is the two-escaper-trap discriminator
|
||||
named in the spec's security model (section 6).
|
||||
- scenario-level request expectations from scenarios-openai.json
|
||||
(tools offered, OpenAI-shaped tool_choice, parallel_tool_calls
|
||||
pinned false per ADR-0005) -> 400 gate_expect.
|
||||
* Answers with SCRIPTED responses: plain text (finish_reason "stop"),
|
||||
tool calls (finish_reason "tool_calls", arguments JSON-encoded, incl. a
|
||||
nested-quote/escaping torture payload and a parallel two-call case), and
|
||||
API-error injection (OpenAI error envelope). Scenario is selected by
|
||||
scanning user-message text (newest first) for a registered marker
|
||||
substring; the step index is the number of assistant messages already in
|
||||
the request (stateless replay - resumes index correctly by construction).
|
||||
* Writes a ground-truth JSONL log (--log): one record per request with the
|
||||
validation verdict, matched scenario/step, and exactly which tool calls
|
||||
were delivered. Gate assertions compare the brain's claims against THIS
|
||||
log - truth, not narration.
|
||||
* Unmatched requests (boot probes, awareness chatter) get a benign "ok"
|
||||
text response, logged kind=background, never counted as ground truth.
|
||||
* HOSTILE MODES (--mode) on the same file:
|
||||
black-hole accept + read the request, never respond;
|
||||
mid-body-drop send half a JSON body, then abort the socket;
|
||||
tool-pending-forever every request gets a FRESH tool_call
|
||||
(finish_reason "tool_calls"), forever - tests
|
||||
the agentic loop's iteration cap; count the
|
||||
brain's round-trips via GET /gate/stats.
|
||||
|
||||
usage: stub-openai.py --port P --scenarios scenarios-openai.json \
|
||||
--log requests.jsonl [--mode MODE]
|
||||
Listens on 127.0.0.1 only. Refuses production ports 7770/7779/17779.
|
||||
"""
|
||||
import argparse
|
||||
import itertools
|
||||
import json
|
||||
import socket
|
||||
import struct
|
||||
import threading
|
||||
import time
|
||||
import uuid
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
STATE = {"scenarios": None, "log_path": None, "lock": threading.Lock(),
|
||||
"seq": 0, "mode": "normal", "chat_hits": 0}
|
||||
_PENDING_SEQ = itertools.count(1)
|
||||
|
||||
ANTHROPIC_TOP_KEYS = ("system", "stop_sequences", "max_tokens_to_sample",
|
||||
"anthropic_version")
|
||||
ANTHROPIC_BLOCK_TYPES = {"tool_use", "tool_result", "server_tool_use",
|
||||
"web_search_tool_result", "thinking",
|
||||
"redacted_thinking"}
|
||||
DEFAULT_EXPECT = {"require_tools": True, "require_tool_choice": True,
|
||||
"parallel_tool_calls": False, "forbid_tools": False}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- loading ----
|
||||
def load_scenarios(path):
|
||||
cfg = json.load(open(path))
|
||||
defaults = dict(DEFAULT_EXPECT)
|
||||
defaults.update(cfg.get("defaults", {}).get("expect_request", {}))
|
||||
marker_map = [] # (marker_lower, cname, pid)
|
||||
scripts = {} # cname or cname/pid -> expanded script
|
||||
pid_map = {} # pid -> cname (for call_gate_* id -> script lookup)
|
||||
expects = {} # cname -> merged expect_request
|
||||
for cname, cls in cfg["classes"].items():
|
||||
scripts[cname] = expand_script(cls.get("script", []))
|
||||
exp = dict(defaults)
|
||||
exp.update(cls.get("expect_request", {}))
|
||||
expects[cname] = exp
|
||||
for ph in cls["phrasings"]:
|
||||
if ph.get("script") is not None:
|
||||
scripts[cname + "/" + ph["id"]] = expand_script(ph["script"])
|
||||
marker_map.append((ph["marker"].lower(), cname, ph["id"]))
|
||||
pid_map[ph["id"]] = cname
|
||||
return {"cfg": cfg, "marker_map": marker_map, "scripts": scripts,
|
||||
"pid_map": pid_map, "expects": expects}
|
||||
|
||||
|
||||
def expand_script(script):
|
||||
"""Same repeat-expansion contract as gate9's stub-llm.py ({N}/{NN})."""
|
||||
out = []
|
||||
for step in script:
|
||||
if "repeat" in step:
|
||||
for n in range(1, step["repeat"] + 1):
|
||||
t = {k: v for k, v in step.items() if k != "repeat"}
|
||||
out.append(json.loads(json.dumps(t)
|
||||
.replace("{NN}", "%02d" % n)
|
||||
.replace("{N}", str(n))))
|
||||
else:
|
||||
out.append(step)
|
||||
return out
|
||||
|
||||
|
||||
# ------------------------------------------------------------- validation ----
|
||||
def _rej(message, code):
|
||||
return {"status": 400, "message": message, "code": code}
|
||||
|
||||
|
||||
def validate_dialect(headers, req):
|
||||
"""Universal checks - run on EVERY request, scenario-matched or not.
|
||||
Anything Anthropic-shaped on this lane means the brain's translator
|
||||
leaked; the whole point is that it fails loudly, here, with a reason."""
|
||||
if headers.get("anthropic-version"):
|
||||
return _rej("anthropic-version header on the OpenAI lane: this "
|
||||
"request was built by the Anthropic dialect path",
|
||||
"gate_dialect_leak")
|
||||
for k in ANTHROPIC_TOP_KEYS:
|
||||
if k in req:
|
||||
return _rej("top-level `%s` is Anthropic dialect; the OpenAI "
|
||||
"dialect has no such field (system prompt goes in "
|
||||
"messages[0])" % k, "gate_dialect_leak")
|
||||
tools = req.get("tools")
|
||||
if tools is not None:
|
||||
if not isinstance(tools, list):
|
||||
return _rej("`tools` must be an array", "gate_tools_shape")
|
||||
names = []
|
||||
for i, t in enumerate(tools):
|
||||
if not isinstance(t, dict):
|
||||
return _rej("tools[%d] is not an object" % i,
|
||||
"gate_tools_shape")
|
||||
if "input_schema" in t or (isinstance(t.get("function"), dict)
|
||||
and "input_schema" in t["function"]):
|
||||
return _rej("tools[%d] carries `input_schema` (Anthropic "
|
||||
"dialect); OpenAI dialect wants "
|
||||
"function.parameters" % i, "gate_dialect_leak")
|
||||
if t.get("type") != "function":
|
||||
return _rej("tools[%d].type must be \"function\", got %r"
|
||||
% (i, t.get("type")), "gate_tools_shape")
|
||||
fn = t.get("function")
|
||||
if not isinstance(fn, dict):
|
||||
return _rej("tools[%d].function missing" % i,
|
||||
"gate_tools_shape")
|
||||
if not isinstance(fn.get("name"), str) or not fn["name"]:
|
||||
return _rej("tools[%d].function.name missing/empty" % i,
|
||||
"gate_tools_shape")
|
||||
if not isinstance(fn.get("description"), str) or not fn["description"]:
|
||||
return _rej("tools[%d].function.description missing/empty" % i,
|
||||
"gate_tools_shape")
|
||||
if not isinstance(fn.get("parameters"), dict):
|
||||
return _rej("tools[%d].function.parameters missing (JSON "
|
||||
"Schema object expected)" % i, "gate_tools_shape")
|
||||
names.append(fn["name"])
|
||||
if len(names) != len(set(names)):
|
||||
return _rej("tools: tool names must be unique", "gate_tools_shape")
|
||||
msgs = req.get("messages")
|
||||
if not isinstance(msgs, list) or not msgs:
|
||||
return _rej("`messages` must be a non-empty array",
|
||||
"gate_messages_shape")
|
||||
for i, m in enumerate(msgs):
|
||||
if not isinstance(m, dict):
|
||||
return _rej("messages[%d] is not an object" % i,
|
||||
"gate_messages_shape")
|
||||
c = m.get("content")
|
||||
if isinstance(c, list):
|
||||
for j, b in enumerate(c):
|
||||
if isinstance(b, dict) and b.get("type") in ANTHROPIC_BLOCK_TYPES:
|
||||
return _rej("messages[%d].content[%d] is an Anthropic "
|
||||
"`%s` block; the OpenAI dialect uses "
|
||||
"tool_calls / role:\"tool\" messages"
|
||||
% (i, j, b.get("type")), "gate_dialect_leak")
|
||||
if m.get("role") == "tool":
|
||||
if not isinstance(m.get("tool_call_id"), str) or not m["tool_call_id"]:
|
||||
return _rej("messages[%d]: role \"tool\" requires a "
|
||||
"`tool_call_id`" % i, "gate_messages_shape")
|
||||
if "content" not in m:
|
||||
return _rej("messages[%d]: role \"tool\" requires `content`"
|
||||
% i, "gate_messages_shape")
|
||||
if m.get("role") == "assistant" and m.get("tool_calls") is not None:
|
||||
tcs = m["tool_calls"]
|
||||
if not isinstance(tcs, list) or not tcs:
|
||||
return _rej("messages[%d].tool_calls must be a non-empty "
|
||||
"array" % i, "gate_tool_call_shape")
|
||||
for j, tc in enumerate(tcs):
|
||||
if not isinstance(tc, dict) or tc.get("type") != "function":
|
||||
return _rej("messages[%d].tool_calls[%d].type must be "
|
||||
"\"function\"" % (i, j), "gate_tool_call_shape")
|
||||
if not isinstance(tc.get("id"), str) or not tc["id"]:
|
||||
return _rej("messages[%d].tool_calls[%d].id missing"
|
||||
% (i, j), "gate_tool_call_shape")
|
||||
fn = tc.get("function")
|
||||
if not isinstance(fn, dict) or not isinstance(fn.get("name"), str):
|
||||
return _rej("messages[%d].tool_calls[%d].function.name "
|
||||
"missing" % (i, j), "gate_tool_call_shape")
|
||||
if not isinstance(fn.get("arguments"), str):
|
||||
return _rej("messages[%d].tool_calls[%d].function."
|
||||
"arguments must be a JSON-encoded STRING, "
|
||||
"got %s" % (i, j,
|
||||
type(fn.get("arguments")).__name__),
|
||||
"gate_tool_call_shape")
|
||||
return None
|
||||
|
||||
|
||||
def validate_pairing(msgs):
|
||||
"""OpenAI pairing rule: every assistant tool_calls turn must be followed
|
||||
immediately by role:"tool" messages answering every tool_call_id."""
|
||||
open_ids, open_at = set(), None
|
||||
for i, m in enumerate(msgs):
|
||||
role = m.get("role")
|
||||
if role == "tool":
|
||||
tid = m.get("tool_call_id")
|
||||
if open_at is None:
|
||||
return _rej("messages[%d]: role \"tool\" message with no "
|
||||
"preceding assistant tool_calls turn "
|
||||
"(tool_call_id=%s)" % (i, tid), "gate_pairing")
|
||||
if tid not in open_ids:
|
||||
return _rej("messages[%d]: tool message answers unknown or "
|
||||
"already-answered tool_call_id %s" % (i, tid),
|
||||
"gate_pairing")
|
||||
open_ids.discard(tid)
|
||||
continue
|
||||
if open_ids:
|
||||
return _rej("messages[%d]: assistant tool_calls not fully "
|
||||
"answered before messages[%d]; missing tool "
|
||||
"responses for: %s" % (open_at, i, sorted(open_ids)),
|
||||
"gate_pairing")
|
||||
open_ids, open_at = set(), None
|
||||
if role == "assistant" and m.get("tool_calls"):
|
||||
ids = [tc.get("id") for tc in m["tool_calls"]]
|
||||
open_ids, open_at = set(ids), i
|
||||
if open_ids:
|
||||
return _rej("messages[%d]: assistant tool_calls at end of thread "
|
||||
"without tool responses for: %s"
|
||||
% (open_at, sorted(open_ids)), "gate_pairing")
|
||||
return None
|
||||
|
||||
|
||||
def validate_echo_args(msgs, loaded):
|
||||
"""Ground-truth round-trip check: for every echoed gate-issued call id,
|
||||
recompute the arguments this stub originally sent from the script and
|
||||
require one json decode to reproduce them exactly. Catches the
|
||||
two-escaper trap (spec section 6) deterministically."""
|
||||
if not loaded:
|
||||
return None
|
||||
for i, m in enumerate(msgs):
|
||||
if m.get("role") != "assistant":
|
||||
continue
|
||||
for tc in m.get("tool_calls") or []:
|
||||
tid = tc.get("id", "")
|
||||
if not tid.startswith("call_gate_"):
|
||||
continue
|
||||
rest = tid[len("call_gate_"):]
|
||||
try:
|
||||
pid, s_part, k_part = rest.rsplit("_", 2)
|
||||
step_idx, k = int(s_part[1:]), int(k_part)
|
||||
except (ValueError, IndexError):
|
||||
continue
|
||||
cname = loaded["pid_map"].get(pid)
|
||||
if cname is None:
|
||||
continue
|
||||
script = (loaded["scripts"].get(cname + "/" + pid)
|
||||
or loaded["scripts"].get(cname) or [])
|
||||
if step_idx >= len(script):
|
||||
continue
|
||||
calls = script[step_idx].get("tool_calls") or []
|
||||
if k >= len(calls):
|
||||
continue
|
||||
expected = calls[k]
|
||||
fn = tc.get("function") or {}
|
||||
if fn.get("name") != expected["name"]:
|
||||
return _rej("messages[%d]: echoed tool name %r != issued %r "
|
||||
"for %s" % (i, fn.get("name"), expected["name"],
|
||||
tid), "gate_echo_mismatch")
|
||||
try:
|
||||
got = json.loads(fn.get("arguments", ""))
|
||||
except ValueError:
|
||||
return _rej("messages[%d]: echoed arguments for %s are not "
|
||||
"valid JSON after one decode (truncated or "
|
||||
"half-escaped?)" % (i, tid), "gate_echo_mismatch")
|
||||
if got != expected["arguments"]:
|
||||
hint = (" (decoded to a string, not an object: "
|
||||
"double-encoded - the two-escaper trap)"
|
||||
if isinstance(got, str) else "")
|
||||
return _rej("messages[%d]: echoed arguments for %s do not "
|
||||
"round-trip to the issued payload%s"
|
||||
% (i, tid, hint), "gate_echo_mismatch")
|
||||
return None
|
||||
|
||||
|
||||
def validate_expect(req, exp):
|
||||
"""Scenario-level request expectations (scenarios-openai.json)."""
|
||||
tools = req.get("tools") or []
|
||||
if exp.get("forbid_tools") and tools:
|
||||
return _rej("this scenario is chat-only: no `tools` may be offered "
|
||||
"on it", "gate_expect")
|
||||
if exp.get("require_tools") and not tools:
|
||||
return _rej("scenario expects a `tools` array to be offered (the "
|
||||
"agentic lane must advertise its tools)", "gate_expect")
|
||||
if exp.get("require_tool_choice"):
|
||||
tc = req.get("tool_choice")
|
||||
ok = tc in ("auto", "none", "required") or (
|
||||
isinstance(tc, dict) and tc.get("type") == "function"
|
||||
and isinstance(tc.get("function"), dict)
|
||||
and tc["function"].get("name"))
|
||||
if not ok:
|
||||
return _rej("scenario expects an OpenAI-shaped `tool_choice`, "
|
||||
"got %r" % (tc,), "gate_expect")
|
||||
want_ptc = exp.get("parallel_tool_calls", None)
|
||||
if want_ptc is not None:
|
||||
if "parallel_tool_calls" not in req:
|
||||
return _rej("scenario expects explicit `parallel_tool_calls` "
|
||||
"(ADR-0005: must be pinned false on the wire)",
|
||||
"gate_expect")
|
||||
if req["parallel_tool_calls"] != want_ptc:
|
||||
return _rej("scenario expects parallel_tool_calls=%s, got %s"
|
||||
% (json.dumps(want_ptc),
|
||||
json.dumps(req["parallel_tool_calls"])),
|
||||
"gate_expect")
|
||||
return None
|
||||
|
||||
|
||||
# --------------------------------------------------------- scenario match ----
|
||||
def extract_user_texts_newest_first(msgs):
|
||||
texts = []
|
||||
for m in reversed(msgs):
|
||||
if not isinstance(m, dict) or m.get("role") != "user":
|
||||
continue
|
||||
c = m.get("content")
|
||||
if isinstance(c, str):
|
||||
texts.append(c)
|
||||
elif isinstance(c, list):
|
||||
for b in c:
|
||||
if isinstance(b, dict) and b.get("type") == "text":
|
||||
texts.append(b.get("text", ""))
|
||||
return texts
|
||||
|
||||
|
||||
def match_scenario(loaded, msgs):
|
||||
for text in extract_user_texts_newest_first(msgs):
|
||||
tl = text.lower()
|
||||
for marker, cname, pid in loaded["marker_map"]:
|
||||
if marker in tl:
|
||||
return cname, pid
|
||||
return None, None
|
||||
|
||||
|
||||
# ------------------------------------------------------------- rendering ----
|
||||
def completion_envelope(msg, finish, model, usage=(100, 100)):
|
||||
return {"id": "chatcmpl-gate-" + uuid.uuid4().hex[:12],
|
||||
"object": "chat.completion", "created": int(time.time()),
|
||||
"model": model,
|
||||
"choices": [{"index": 0, "message": msg,
|
||||
"finish_reason": finish, "logprobs": None}],
|
||||
"usage": {"prompt_tokens": usage[0],
|
||||
"completion_tokens": usage[1],
|
||||
"total_tokens": usage[0] + usage[1]}}
|
||||
|
||||
|
||||
def text_completion(text, model):
|
||||
return completion_envelope({"role": "assistant", "content": text},
|
||||
"stop", model, usage=(1, 1))
|
||||
|
||||
|
||||
def pending_body(seq, model):
|
||||
args = json.dumps({"path": "never-%04d.md" % seq,
|
||||
"content": "this run never completes"})
|
||||
msg = {"role": "assistant", "content": None,
|
||||
"tool_calls": [{"id": "call_hostile_pending_%04d" % seq,
|
||||
"type": "function",
|
||||
"function": {"name": "write_file",
|
||||
"arguments": args}}]}
|
||||
return completion_envelope(msg, "tool_calls", model, usage=(1, 1))
|
||||
|
||||
|
||||
def render_step(step, cname, pid, step_idx, model):
|
||||
"""Returns (http_status, body_dict, delivered) - delivered is ground
|
||||
truth for the JSONL log."""
|
||||
delivered = {"tool_calls": [], "finish_reason": None, "api_error": None}
|
||||
if "api_error" in step:
|
||||
e = step["api_error"]
|
||||
delivered["api_error"] = e["status"]
|
||||
return (e["status"],
|
||||
{"error": {"message": e["message"],
|
||||
"type": e.get("type", "server_error"),
|
||||
"param": None, "code": e.get("code")}},
|
||||
delivered)
|
||||
msg = {"role": "assistant"}
|
||||
finish = "stop"
|
||||
if step.get("tool_calls"):
|
||||
tcs = []
|
||||
for k, call in enumerate(step["tool_calls"]):
|
||||
tid = "call_gate_%s_s%d_%d" % (pid, step_idx, k)
|
||||
tcs.append({"id": tid, "type": "function",
|
||||
"function": {"name": call["name"],
|
||||
"arguments": json.dumps(
|
||||
call["arguments"],
|
||||
ensure_ascii=False)}})
|
||||
delivered["tool_calls"].append(call["name"])
|
||||
msg["tool_calls"] = tcs
|
||||
msg["content"] = step.get("text") # null when no narration, like real
|
||||
finish = "tool_calls"
|
||||
else:
|
||||
msg["content"] = step["text"]
|
||||
delivered["finish_reason"] = finish
|
||||
return 200, completion_envelope(msg, finish, model), delivered
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ log ------
|
||||
def log_record(rec):
|
||||
with STATE["lock"]:
|
||||
STATE["seq"] += 1
|
||||
rec["seq"] = STATE["seq"]
|
||||
with open(STATE["log_path"], "a") as f:
|
||||
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- server -----
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
protocol_version = "HTTP/1.1"
|
||||
|
||||
def _send_json(self, status, obj):
|
||||
body = json.dumps(obj, ensure_ascii=False).encode("utf-8")
|
||||
self.send_response(status)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def _send_error(self, verdict):
|
||||
self._send_json(verdict["status"],
|
||||
{"error": {"message": verdict["message"],
|
||||
"type": "invalid_request_error",
|
||||
"param": None, "code": verdict["code"]}})
|
||||
|
||||
def _drop_mid_body(self):
|
||||
"""Valid 200 headers, half the promised body, then a socket abort
|
||||
(same SO_LINGER teardown as gate9's mid-body-drop-brain.py)."""
|
||||
full = json.dumps(text_completion(
|
||||
"This reply will never finish arriving because the connection "
|
||||
"dies in the middle of the body, which is exactly the point of "
|
||||
"this hostile fixture.", "hostile-mid-drop")).encode("utf-8")
|
||||
half = full[: len(full) // 2]
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(full))) # promises more
|
||||
self.end_headers()
|
||||
self.wfile.write(half)
|
||||
self.wfile.flush()
|
||||
try:
|
||||
self.connection.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER,
|
||||
struct.pack("ii", 1, 0))
|
||||
self.connection.shutdown(socket.SHUT_RDWR)
|
||||
except OSError:
|
||||
pass
|
||||
self.close_connection = True
|
||||
|
||||
def do_GET(self):
|
||||
path = self.path.split("?")[0]
|
||||
if path == "/gate/health":
|
||||
self._send_json(200, {"ok": True, "mode": STATE["mode"]})
|
||||
elif path == "/gate/stats":
|
||||
with STATE["lock"]:
|
||||
self._send_json(200, {"mode": STATE["mode"],
|
||||
"chat_hits": STATE["chat_hits"]})
|
||||
else:
|
||||
self._send_json(404, {"error": {"message": "not found",
|
||||
"type": "invalid_request_error",
|
||||
"param": None,
|
||||
"code": "unknown_route"}})
|
||||
|
||||
def do_POST(self):
|
||||
n = int(self.headers.get("Content-Length") or 0)
|
||||
raw = self.rfile.read(n)
|
||||
mode = STATE["mode"]
|
||||
rec = {"ts": time.time(), "path": self.path, "mode": mode,
|
||||
"kind": "background", "scenario_class": None, "phrasing": None,
|
||||
"step": None, "n_messages": 0, "n_assistant": 0,
|
||||
"validation": "ok", "validation_detail": None,
|
||||
"delivered": {"tool_calls": [], "finish_reason": None,
|
||||
"api_error": None},
|
||||
"http_status": 200}
|
||||
if self.path.split("?")[0] != "/v1/chat/completions":
|
||||
rec.update(kind="wrong_path", http_status=404)
|
||||
log_record(rec)
|
||||
self._send_json(404, {"error": {
|
||||
"message": "no such route: %s" % self.path,
|
||||
"type": "invalid_request_error", "param": None,
|
||||
"code": "unknown_route"}})
|
||||
return
|
||||
with STATE["lock"]:
|
||||
STATE["chat_hits"] += 1
|
||||
|
||||
# ---- hostile modes: behavior first, no validation ----------------
|
||||
if mode == "black-hole":
|
||||
rec.update(kind="hostile", http_status=None)
|
||||
log_record(rec)
|
||||
threading.Event().wait() # hold the socket open forever
|
||||
return
|
||||
if mode == "mid-body-drop":
|
||||
rec.update(kind="hostile", http_status=200)
|
||||
log_record(rec)
|
||||
self._drop_mid_body()
|
||||
return
|
||||
if mode == "tool-pending-forever":
|
||||
seq = next(_PENDING_SEQ)
|
||||
rec.update(kind="hostile",
|
||||
delivered={"tool_calls": ["write_file"],
|
||||
"finish_reason": "tool_calls",
|
||||
"api_error": None})
|
||||
log_record(rec)
|
||||
self._send_json(200, pending_body(seq, "gate-openai-model"))
|
||||
return
|
||||
|
||||
# ---- normal mode -------------------------------------------------
|
||||
try:
|
||||
req = json.loads(raw)
|
||||
except ValueError as exc:
|
||||
# DIAGNOSTIC CAPTURE (2026-08-06): an unparseable body used to be recorded as
|
||||
# a bare "bad_json" with the bytes thrown away, which made an intermittent
|
||||
# failure impossible to root-cause — you cannot fix what you did not keep.
|
||||
# Dump the raw body next to the log, and record exactly where the parser gave
|
||||
# up plus the offending byte, so one occurrence is enough to diagnose.
|
||||
dump_path = "%s.badbody.%s" % (STATE.get("log_path", "/tmp/stub-openai"),
|
||||
rec.get("seq", "x"))
|
||||
try:
|
||||
data = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
|
||||
with open(dump_path, "wb") as fh:
|
||||
fh.write(data)
|
||||
except Exception as dump_exc:
|
||||
dump_path = "(dump failed: %s)" % dump_exc
|
||||
pos = getattr(exc, "pos", None)
|
||||
near = ""
|
||||
byte_repr = ""
|
||||
if isinstance(pos, int):
|
||||
blob = raw if isinstance(raw, (bytes, bytearray)) else str(raw).encode()
|
||||
near = blob[max(0, pos - 60):pos + 60].decode("utf-8", "replace")
|
||||
if 0 <= pos < len(blob):
|
||||
byte_repr = "0x%02x" % blob[pos]
|
||||
rec.update(kind="bad_json", validation="rejected",
|
||||
validation_detail="request body is not valid JSON: %s" % exc,
|
||||
http_status=400, raw_len=len(raw), raw_dump=dump_path,
|
||||
err_pos=pos, err_byte=byte_repr, err_near=near)
|
||||
log_record(rec)
|
||||
self._send_error(_rej("request body is not valid JSON",
|
||||
"bad_json"))
|
||||
return
|
||||
msgs = req.get("messages") or []
|
||||
rec["n_messages"] = len(msgs)
|
||||
rec["n_assistant"] = sum(1 for m in msgs if isinstance(m, dict)
|
||||
and m.get("role") == "assistant")
|
||||
loaded = STATE["scenarios"]
|
||||
cname, pid = match_scenario(loaded, msgs)
|
||||
if cname:
|
||||
rec.update(kind="scenario", scenario_class=cname, phrasing=pid)
|
||||
|
||||
# Wire-level validation runs for EVERY request, scenario or not.
|
||||
verdict = (validate_dialect(self.headers, req)
|
||||
or validate_pairing([m for m in msgs
|
||||
if isinstance(m, dict)])
|
||||
or validate_echo_args(msgs, loaded))
|
||||
if verdict:
|
||||
rec.update(validation="rejected",
|
||||
validation_detail=verdict["message"],
|
||||
http_status=verdict["status"])
|
||||
log_record(rec)
|
||||
self._send_error(verdict)
|
||||
return
|
||||
|
||||
model = req.get("model", "gate-openai-model")
|
||||
if not cname:
|
||||
log_record(rec)
|
||||
self._send_json(200, text_completion("ok", model))
|
||||
return
|
||||
|
||||
script = (loaded["scripts"].get(cname + "/" + pid)
|
||||
or loaded["scripts"][cname])
|
||||
step_idx = rec["n_assistant"]
|
||||
if step_idx >= len(script):
|
||||
rec.update(kind="overrun", step=step_idx)
|
||||
log_record(rec)
|
||||
self._send_json(200, text_completion(
|
||||
"GATE-SCRIPT-EXHAUSTED %s step %d" % (pid, step_idx), model))
|
||||
return
|
||||
|
||||
step = script[step_idx]
|
||||
exp = dict(loaded["expects"][cname])
|
||||
exp.update(step.get("expect_request", {}))
|
||||
verdict = validate_expect(req, exp)
|
||||
if verdict:
|
||||
rec.update(step=step_idx, validation="rejected",
|
||||
validation_detail=verdict["message"],
|
||||
http_status=verdict["status"])
|
||||
log_record(rec)
|
||||
self._send_error(verdict)
|
||||
return
|
||||
|
||||
status, body, delivered = render_step(step, cname, pid, step_idx,
|
||||
model)
|
||||
rec.update(step=step_idx, delivered=delivered, http_status=status)
|
||||
log_record(rec)
|
||||
self._send_json(status, body)
|
||||
|
||||
def log_message(self, *a):
|
||||
pass
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--port", type=int, required=True)
|
||||
ap.add_argument("--scenarios",
|
||||
help="scenarios-openai.json (required in normal mode)")
|
||||
ap.add_argument("--log", required=True)
|
||||
ap.add_argument("--mode", default="normal",
|
||||
choices=["normal", "black-hole", "mid-body-drop",
|
||||
"tool-pending-forever"])
|
||||
args = ap.parse_args()
|
||||
if args.port in (7770, 7779, 17779):
|
||||
raise SystemExit("stub-openai: refusing production Neuron port")
|
||||
if args.mode == "normal" and not args.scenarios:
|
||||
raise SystemExit("stub-openai: --scenarios is required in normal mode")
|
||||
STATE["mode"] = args.mode
|
||||
STATE["scenarios"] = (load_scenarios(args.scenarios)
|
||||
if args.scenarios else None)
|
||||
STATE["log_path"] = args.log
|
||||
open(args.log, "w").close()
|
||||
n_markers = (len(STATE["scenarios"]["marker_map"])
|
||||
if STATE["scenarios"] else 0)
|
||||
print("stub-openai [%s]: 127.0.0.1:%d /v1/chat/completions "
|
||||
"(%d markers registered, log=%s)"
|
||||
% (args.mode, args.port, n_markers, args.log), flush=True)
|
||||
ThreadingHTTPServer(("127.0.0.1", args.port), Handler).serve_forever()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Executable
+148
@@ -0,0 +1,148 @@
|
||||
#!/usr/bin/env bash
|
||||
# run-el-test.sh — build and RUN one engine test (tests/*.el), printing its assertions.
|
||||
#
|
||||
# WHY THIS EXISTS (2026-08-06): the engine's tests/*.el files were never runnable from the
|
||||
# tree. `elc` is a COMPILER — it emits C to stdout and exits; it does not execute anything.
|
||||
# So "the tests" could only ever be read, not run, and a signature change could silently
|
||||
# break them (exactly what happened when bridge_save gained its `wire` argument). This
|
||||
# script closes that: emit the test to C, link it against the engine modules, execute it.
|
||||
#
|
||||
# HOW IT WORKS
|
||||
# 1. elc <test>.el -> C on stdout (the test file's `main` + prototypes)
|
||||
# 2. elb (once, cached) -> per-module C for the whole engine into a scratch dir
|
||||
# 3. cc test.c + all modules EXCEPT soul.c (soul.c owns the real `main`) + the runtime
|
||||
# 4. run it
|
||||
#
|
||||
# The test C references only the engine functions it actually calls, so there are no
|
||||
# duplicate-symbol collisions with the module objects.
|
||||
#
|
||||
# RUNTIME: the REPO-PINNED vendor/el-runtime (NOT ~/el-sdk/el_runtime.c — that June build
|
||||
# is missing builtins August code calls: engram_wm_count, engram_wm_top_json,
|
||||
# http_delete_json, http_serve_async; linking against it fails with "symbol(s) not found").
|
||||
#
|
||||
# USAGE
|
||||
# tests/run-el-test.sh tests/test_bridge_serialization.el # one test
|
||||
# tests/run-el-test.sh --all # every tests/test_*.el
|
||||
# REBUILD=1 tests/run-el-test.sh ... # force module regeneration
|
||||
#
|
||||
# Tests that need a live API key / running soul (see each file's header) will report their
|
||||
# own skips or failures — this runner does not fake them.
|
||||
set -uo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
cd "$REPO_ROOT" || exit 2
|
||||
|
||||
ELC="${ELC:-$HOME/el-sdk/elc}"
|
||||
ELB="${ELB:-$HOME/Development/el-sdk/bin/elb}"
|
||||
RUNTIME_DIR="${RUNTIME_DIR:-$REPO_ROOT/vendor/el-runtime/v1.0.0-20260501}"
|
||||
SCRATCH="${SCRATCH:-/tmp/el-test-$(basename "$REPO_ROOT")}"
|
||||
MODDIR="$SCRATCH/modules"
|
||||
OPENSSL_INC="${OPENSSL_INC:-/opt/homebrew/opt/openssl@3/include}"
|
||||
OPENSSL_LIB="${OPENSSL_LIB:-/opt/homebrew/opt/openssl@3/lib}"
|
||||
|
||||
for req in "$ELC" "$ELB" "$RUNTIME_DIR/el_runtime.c"; do
|
||||
[ -e "$req" ] || { echo "run-el-test: missing required input: $req" >&2; exit 2; }
|
||||
done
|
||||
|
||||
mkdir -p "$MODDIR" || exit 2
|
||||
|
||||
# ── Step 1: engine modules (cached — regeneration is the slow part) ────────────
|
||||
if [ "${REBUILD:-0}" = "1" ] || [ ! -f "$MODDIR/chat.c" ] || [ "chat.el" -nt "$MODDIR/chat.c" ]; then
|
||||
echo "run-el-test: generating engine modules into $MODDIR (this takes ~1-2 min)..."
|
||||
# elb's own final link step fails by design here (it wants to produce a binary named
|
||||
# `neuron` and we only need the per-module .c files it emits first). Ignore its rc.
|
||||
"$ELB" --elc="$ELC" --runtime="$RUNTIME_DIR" --out="$MODDIR/" >"$SCRATCH/elb.log" 2>&1
|
||||
if [ ! -f "$MODDIR/chat.c" ]; then
|
||||
echo "run-el-test: FATAL — elb produced no chat.c; see $SCRATCH/elb.log" >&2
|
||||
tail -5 "$SCRATCH/elb.log" >&2
|
||||
exit 2
|
||||
fi
|
||||
# elb rewrites *.elh in the source tree as a side effect (cosmetic banner churn plus a
|
||||
# stray soul..elh). Say so; the caller decides whether to `git restore` them.
|
||||
echo "run-el-test: NOTE — elb regenerated *.elh in the source tree (cosmetic churn is expected; a stray soul..elh may appear)."
|
||||
fi
|
||||
|
||||
# soul.c is needed for its engine functions (layered_cycle et al.) but it also owns the
|
||||
# daemon's real `main`, which would collide with the test's own. Compile it ONCE to an
|
||||
# object with `main` renamed away, and link that instead of the .c.
|
||||
SOUL_OBJ="$SCRATCH/soul-nomain.o"
|
||||
if [ "${REBUILD:-0}" = "1" ] || [ ! -f "$SOUL_OBJ" ] || [ "$MODDIR/soul.c" -nt "$SOUL_OBJ" ]; then
|
||||
cc -std=c11 -O1 -DHAVE_CURL -Dmain=el_soul_daemon_main_unused \
|
||||
-I "$RUNTIME_DIR" -I "$MODDIR" -I "$OPENSSL_INC" \
|
||||
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
|
||||
-c "$MODDIR/soul.c" -o "$SOUL_OBJ" 2>"$SCRATCH/soul-nomain.err" \
|
||||
|| { echo "run-el-test: FATAL — could not compile soul.c without main" >&2
|
||||
grep -E 'error:' "$SCRATCH/soul-nomain.err" | head -5 >&2; exit 2; }
|
||||
fi
|
||||
|
||||
# Every module except soul.c (linked as the renamed object above) and the stray soul.elh.c.
|
||||
MODS=("$SOUL_OBJ")
|
||||
for f in "$MODDIR"/*.c; do
|
||||
case "$(basename "$f")" in
|
||||
soul.c|soul.elh.c) continue ;;
|
||||
esac
|
||||
MODS+=("$f")
|
||||
done
|
||||
[ "${#MODS[@]}" -gt 1 ] || { echo "run-el-test: no module objects found" >&2; exit 2; }
|
||||
|
||||
run_one() {
|
||||
local test_el="$1"
|
||||
local name; name="$(basename "$test_el" .el)"
|
||||
local cfile="$SCRATCH/$name.c"
|
||||
local bin="$SCRATCH/$name"
|
||||
|
||||
printf '\n══ %s ══\n' "$name"
|
||||
|
||||
if ! "$ELC" "$test_el" >"$cfile" 2>"$SCRATCH/$name.elc.err"; then
|
||||
echo "COMPILE FAILED (elc):"; tail -10 "$SCRATCH/$name.elc.err"; return 1
|
||||
fi
|
||||
[ -s "$cfile" ] || { echo "COMPILE FAILED (elc produced empty C)"; return 1; }
|
||||
|
||||
if ! cc -std=c11 -O1 -DHAVE_CURL -rdynamic \
|
||||
-I "$RUNTIME_DIR" -I "$MODDIR" -I "$OPENSSL_INC" -L "$OPENSSL_LIB" \
|
||||
-include dist/elp-c-decls.h -Wno-error=implicit-function-declaration \
|
||||
-o "$bin" "$cfile" "${MODS[@]}" "$RUNTIME_DIR/el_runtime.c" \
|
||||
-lssl -lcrypto -lcurl -lpthread -lm 2>"$SCRATCH/$name.link.err"; then
|
||||
echo "LINK FAILED:"; grep -E '"_|error:' "$SCRATCH/$name.link.err" | head -10; return 1
|
||||
fi
|
||||
|
||||
# THE RUNNER OWNS THE VERDICT — the test files cannot be trusted to report it.
|
||||
#
|
||||
# Every tests/*.el assert helper does `let pass_count = pass_count + 1` INSIDE an if
|
||||
# BLOCK. El's scope rule (the same one chat.el documents at every while-body mutation:
|
||||
# "mutations inside if *blocks* don't escape scope") means those counters never
|
||||
# increment, so all 9 counted test files print "N passed, M failed" as "0 passed, 0
|
||||
# failed" — forever, whatever actually happened. A summary that can never report a
|
||||
# failure is worth exactly as much as an assertion that can never fail. Logged as a
|
||||
# bug for the real in-file fix; until then the verdict is computed HERE, from the
|
||||
# assert helpers' own per-line output, which IS reliable.
|
||||
local out="$SCRATCH/$name.out"
|
||||
"$bin" 2>&1 | tee "$out"; local rc=${PIPESTATUS[0]}
|
||||
|
||||
# NOTE: `grep -c` prints 0 AND exits 1 when there are no matches, so a `|| echo 0`
|
||||
# fallback appends a SECOND zero and every later integer test breaks on "0\n0".
|
||||
# (Caught by running this script — which is the whole argument for running things.)
|
||||
local n_pass n_fail
|
||||
n_pass=$(grep -c '^ PASS: ' "$out" 2>/dev/null); n_pass=${n_pass:-0}
|
||||
n_fail=$(grep -c '^ FAIL: ' "$out" 2>/dev/null); n_fail=${n_fail:-0}
|
||||
echo "── $name: $n_pass passed, $n_fail failed (counted by the runner, not by the file's dead counters)"
|
||||
if [ "$n_fail" -gt 0 ]; then
|
||||
echo " failing assertions:"; grep '^ FAIL: ' "$out" | sed 's/^/ /'
|
||||
return 1
|
||||
fi
|
||||
if [ "$n_pass" -eq 0 ]; then
|
||||
echo " WARNING: no assertions ran — treating as FAILURE (a test that asserts nothing is not a passing test)"
|
||||
return 1
|
||||
fi
|
||||
[ $rc -eq 0 ] || { echo " (test binary exited rc=$rc)"; return 1; }
|
||||
return 0
|
||||
}
|
||||
|
||||
rc_all=0
|
||||
if [ "${1:-}" = "--all" ]; then
|
||||
for t in tests/test_*.el; do run_one "$t" || rc_all=1; done
|
||||
else
|
||||
[ $# -ge 1 ] || { echo "usage: tests/run-el-test.sh <tests/test_x.el> | --all" >&2; exit 2; }
|
||||
for t in "$@"; do run_one "$t" || rc_all=1; done
|
||||
fi
|
||||
exit $rc_all
|
||||
@@ -93,7 +93,7 @@ println("1. bridge_save — empty messages guard")
|
||||
let sid1: String = "test-session-empty-messages"
|
||||
state_set("mcp_bridge:" + sid1, "")
|
||||
|
||||
let save1_ok: Bool = bridge_save(sid1, "claude-sonnet-4-5", "sys", "[]", "", "", "call-1")
|
||||
let save1_ok: Bool = bridge_save(sid1, "claude-sonnet-4-5", "sys", "[]", "", "", "call-1", "anthropic")
|
||||
assert_false("empty messages -> bridge_save returns false", save1_ok)
|
||||
|
||||
let saved1: String = state_get("mcp_bridge:" + sid1)
|
||||
@@ -107,7 +107,7 @@ println("2. bridge_save — empty tools_json guard")
|
||||
let sid2: String = "test-session-empty-tools"
|
||||
state_set("mcp_bridge:" + sid2, "")
|
||||
|
||||
let save2_ok: Bool = bridge_save(sid2, "claude-sonnet-4-5", "sys", "", "[{\"role\":\"user\",\"content\":\"hi\"}]", "", "call-2")
|
||||
let save2_ok: Bool = bridge_save(sid2, "claude-sonnet-4-5", "sys", "", "[{\"role\":\"user\",\"content\":\"hi\"}]", "", "call-2", "anthropic")
|
||||
assert_false("empty tools_json -> bridge_save returns false", save2_ok)
|
||||
|
||||
let saved2: String = state_get("mcp_bridge:" + sid2)
|
||||
@@ -126,7 +126,7 @@ state_set("mcp_bridge:" + sid3, "")
|
||||
|
||||
let msgs3: String = "[{\"role\":\"user\",\"content\":\"hello\"}]"
|
||||
let tools3: String = "[{\"name\":\"read_file\"}]"
|
||||
let save3_ok: Bool = bridge_save(sid3, "claude-sonnet-4-5", "You are a helper.", tools3, msgs3, "read_file", "toolu_abc")
|
||||
let save3_ok: Bool = bridge_save(sid3, "claude-sonnet-4-5", "You are a helper.", tools3, msgs3, "read_file", "toolu_abc", "anthropic")
|
||||
assert_true("valid args -> bridge_save returns true", save3_ok)
|
||||
|
||||
let blob3: String = state_get("mcp_bridge:" + sid3)
|
||||
@@ -243,7 +243,7 @@ state_set("mcp_bridge:" + sid8, "")
|
||||
let special_id: String = "toolu_test\"quoted\""
|
||||
let msgs8: String = "[{\"role\":\"user\",\"content\":\"hi\"}]"
|
||||
let tools8: String = "[{\"name\":\"read_file\"}]"
|
||||
let save8_ok: Bool = bridge_save(sid8, "claude-sonnet-4-5", "sys", tools8, msgs8, "", special_id)
|
||||
let save8_ok: Bool = bridge_save(sid8, "claude-sonnet-4-5", "sys", tools8, msgs8, "", special_id, "anthropic")
|
||||
assert_true("special chars in tool_use_id -> bridge_save returns true", save8_ok)
|
||||
|
||||
let blob8: String = state_get("mcp_bridge:" + sid8)
|
||||
@@ -251,6 +251,111 @@ let blob8: String = state_get("mcp_bridge:" + sid8)
|
||||
let retrieved_id: String = json_get(blob8, "tool_use_id")
|
||||
assert_eq("tool_use_id with quotes round-trips via json_safe", retrieved_id, special_id)
|
||||
|
||||
// ── Section 9: the "wire" field (OpenAI-tools port, 2026-08-06) ───────────────
|
||||
//
|
||||
// A suspended turn must resume on the SAME wire format it suspended on: an OpenAI-lane
|
||||
// bridge answered with an Anthropic-shaped tool_result (or vice versa) is a dead run.
|
||||
// bridge_save therefore stamps the blob with "wire", and agentic_resume branches on it.
|
||||
//
|
||||
// §9c is the important one. json_get is a first-substring-match scanner, so any key that
|
||||
// appears inside the UNESCAPED conversation embedded in messages_raw can be matched
|
||||
// instead of the blob's own field — that exact class of bug produced the round-9 resume
|
||||
// failure (json_get(blob,"tool_use_id") matching a web_search_tool_result's id inside the
|
||||
// replayed conversation). "wire" is written as a json_safe'd SCALAR ahead of both raw
|
||||
// fields precisely so a decoy in model-controlled bytes can never win. This test plants
|
||||
// that decoy on purpose. If someone later moves the field after messages_raw, this fails.
|
||||
|
||||
println("")
|
||||
println("9. bridge_save — wire tagging and its field-order guarantee")
|
||||
|
||||
// 9a. an OpenAI-lane suspension round-trips as "openai"
|
||||
let sid9: String = "test-session-wire-openai"
|
||||
state_set("mcp_bridge:" + sid9, "")
|
||||
let msgs9: String = "[{\"role\":\"user\",\"content\":\"hi\"}]"
|
||||
let tools9: String = "[{\"name\":\"read_file\"}]"
|
||||
let save9_ok: Bool = bridge_save(sid9, "llama-3.3-70b-versatile", "sys", tools9, msgs9, "", "call_abc", "openai")
|
||||
assert_true("openai wire -> bridge_save returns true", save9_ok)
|
||||
let blob9: String = state_get("mcp_bridge:" + sid9)
|
||||
assert_eq("wire round-trips as openai", json_get(blob9, "wire"), "openai")
|
||||
|
||||
// 9b. an Anthropic-lane suspension round-trips as "anthropic"
|
||||
let sid9b: String = "test-session-wire-anthropic"
|
||||
state_set("mcp_bridge:" + sid9b, "")
|
||||
let save9b_ok: Bool = bridge_save(sid9b, "claude-sonnet-4-5", "sys", tools9, msgs9, "", "toolu_abc", "anthropic")
|
||||
assert_true("anthropic wire -> bridge_save returns true", save9b_ok)
|
||||
let blob9b: String = state_get("mcp_bridge:" + sid9b)
|
||||
assert_eq("wire round-trips as anthropic", json_get(blob9b, "wire"), "anthropic")
|
||||
|
||||
// 9c. FIELD-ORDER GUARD: a decoy "wire" inside the conversation must NOT be matched.
|
||||
let sid9c: String = "test-session-wire-decoy"
|
||||
state_set("mcp_bridge:" + sid9c, "")
|
||||
let msgs9c: String = "[{\"role\":\"user\",\"content\":\"please save this literal text: \\\"wire\\\":\\\"anthropic\\\" end\"}]"
|
||||
let save9c_ok: Bool = bridge_save(sid9c, "llama-3.3-70b-versatile", "sys", tools9, msgs9c, "", "call_decoy", "openai")
|
||||
assert_true("decoy conversation -> bridge_save returns true", save9c_ok)
|
||||
let blob9c: String = state_get("mcp_bridge:" + sid9c)
|
||||
assert_eq("blob's own wire wins over a decoy planted in messages_raw", json_get(blob9c, "wire"), "openai")
|
||||
|
||||
// 9d. LEGACY blob (written before the port) has no wire field: json_get yields "",
|
||||
// which agentic_resume treats as the Anthropic path — old suspensions still resume.
|
||||
let sid9d: String = "test-session-wire-legacy"
|
||||
let legacy_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
|
||||
+ ",\"tool_use_id\":\"toolu_legacy\",\"tools_raw\":[{\"name\":\"read_file\"}]"
|
||||
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
|
||||
state_set("mcp_bridge:" + sid9d, legacy_blob)
|
||||
let blob9d: String = state_get("mcp_bridge:" + sid9d)
|
||||
assert_eq("legacy blob has no wire field -> empty (resumes as anthropic)", json_get(blob9d, "wire"), "")
|
||||
assert_eq("legacy blob still reads its tool_use_id", json_get(blob9d, "tool_use_id"), "toolu_legacy")
|
||||
|
||||
// 9e. THE HARDER DECOY: a LEGACY blob (no wire field of its own) that carries the bytes
|
||||
// of a wire tag deeper inside, where an unbounded first-match scan would find it and
|
||||
// misroute the resume onto the wrong loop — the round-9 defect class exactly.
|
||||
//
|
||||
// WHAT IS AND IS NOT REACHABLE (measured here, not assumed — an earlier version of this
|
||||
// test asserted the wrong thing and was corrected by running it):
|
||||
// * NOT reachable from ordinary conversation TEXT. Any quote a user or model writes is
|
||||
// backslash-escaped when it is serialized into the blob, so prose containing
|
||||
// "wire":"openai" is stored as \"wire\":\"openai\" and does not match a scan for the
|
||||
// unescaped key. §9f pins that.
|
||||
// * REACHABLE from STRUCTURAL keys, which are embedded raw. Conversation and tool
|
||||
// objects keep real quotes — that is precisely how round 9's scan found a
|
||||
// web_search_tool_result's tool_use_id. A connector-supplied tool schema or a future
|
||||
// message field literally named "wire" would be found the same way.
|
||||
// The bound removes the whole class rather than reasoning about which keys exist today.
|
||||
let sid9e: String = "test-session-wire-legacy-decoy"
|
||||
let decoy_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
|
||||
+ ",\"tool_use_id\":\"toolu_legacy\""
|
||||
+ ",\"tools_raw\":[{\"name\":\"read_file\",\"wire\":\"openai\"}]"
|
||||
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
|
||||
state_set("mcp_bridge:" + sid9e, decoy_blob)
|
||||
let blob9e: String = state_get("mcp_bridge:" + sid9e)
|
||||
|
||||
// Unbounded read (what NOT to do) — proves the hazard this guard exists for is real.
|
||||
assert_eq("unbounded scan DOES find a structural decoy (why the bound is needed)", json_get(blob9e, "wire"), "openai")
|
||||
|
||||
// Bounded read — the same computation agentic_resume performs.
|
||||
let d_traw: Int = str_index_of(blob9e, ",\"tools_raw\":")
|
||||
let d_tjson: Int = str_index_of(blob9e, ",\"tools_json\":")
|
||||
let d_mraw: Int = str_index_of(blob9e, ",\"messages_raw\":")
|
||||
let d_msgs: Int = str_index_of(blob9e, ",\"messages\":")
|
||||
let dcut1: Int = if d_traw > 0 { d_traw } else { str_len(blob9e) }
|
||||
let dcut2: Int = if d_tjson > 0 && d_tjson < dcut1 { d_tjson } else { dcut1 }
|
||||
let dcut3: Int = if d_mraw > 0 && d_mraw < dcut2 { d_mraw } else { dcut2 }
|
||||
let dcut: Int = if d_msgs > 0 && d_msgs < dcut3 { d_msgs } else { dcut3 }
|
||||
let head9e: String = str_slice(blob9e, 0, dcut)
|
||||
assert_eq("bounded scan ignores the decoy -> legacy blob resumes as anthropic", json_get(head9e, "wire"), "")
|
||||
assert_not_contains("scalar head excludes the bulk fields entirely", head9e, "read_file")
|
||||
|
||||
// 9f. Escaping bounds the severity: prose CANNOT inject a scalar-looking key, because
|
||||
// its quotes are escaped on the way in. Documented as a measured fact, so nobody has to
|
||||
// re-derive it the next time this question comes up.
|
||||
let sid9f: String = "test-session-wire-prose"
|
||||
let prose_blob: String = "{\"model\":\"claude-sonnet-4-5\",\"safe_sys\":\"sys\",\"tools_log\":\"\""
|
||||
+ ",\"tool_use_id\":\"toolu_legacy\",\"tools_raw\":[{\"name\":\"read_file\"}]"
|
||||
+ ",\"messages_raw\":[{\"role\":\"user\",\"content\":\"remember this: \\\"wire\\\":\\\"openai\\\"\"}]}"
|
||||
state_set("mcp_bridge:" + sid9f, prose_blob)
|
||||
let blob9f: String = state_get("mcp_bridge:" + sid9f)
|
||||
assert_eq("escaped prose cannot spoof the key even unbounded (severity bound)", json_get(blob9f, "wire"), "")
|
||||
|
||||
// ── Summary ────────────────────────────────────────────────────────────────────
|
||||
|
||||
println("")
|
||||
|
||||
@@ -0,0 +1,92 @@
|
||||
// tests/test_utf8_slice.el
|
||||
//
|
||||
// Guards utf8_safe_slice(), the fix for a live defect found 2026-08-06:
|
||||
//
|
||||
// The session preload cuts recalled memory content at a fixed length
|
||||
// (chat.el: `if str_len(acc) > 350 { str_slice(acc, 0, 350) }` and
|
||||
// session_preload_bullets' identical per-bullet cut). str_slice and str_len count
|
||||
// BYTES, so any cut landing inside a multi-byte UTF-8 character leaves a dangling
|
||||
// lead byte in the system prompt — and the whole request body is then invalid UTF-8.
|
||||
// Providers reject it outright, so the user sees "AI unavailable" with no clue why,
|
||||
// on both wire formats. Caught by an OpenAI-lane gate whose stub decodes strictly;
|
||||
// reproduced from a real memory whose content contained box-drawing rules (E2 94 80).
|
||||
//
|
||||
// Trigger is ordinary content: an em dash, a curly quote, an accented name, a table
|
||||
// border, an emoji — anything non-ASCII sitting on the cut boundary. It gets MORE
|
||||
// likely as a user's memory grows, which is the opposite of what should happen.
|
||||
//
|
||||
// §1 also pins the semantics this fix depends on: that str_char_code returns the
|
||||
// BYTE value at a byte index (not a decoded code point). If a future runtime changes
|
||||
// that, these assertions fail loudly instead of the truncation silently rotting.
|
||||
|
||||
import "../chat.el"
|
||||
|
||||
let pass_count: Int = 0
|
||||
let fail_count: Int = 0
|
||||
|
||||
fn assert_eq(label: String, got: String, expected: String) -> Void {
|
||||
if str_eq(got, expected) {
|
||||
let pass_count = pass_count + 1
|
||||
println(" PASS: " + label)
|
||||
} else {
|
||||
let fail_count = fail_count + 1
|
||||
println(" FAIL: " + label)
|
||||
println(" got: " + got)
|
||||
println(" expected: " + expected)
|
||||
}
|
||||
}
|
||||
|
||||
fn assert_eq_int(label: String, got: Int, expected: Int) -> Void {
|
||||
assert_eq(label, int_to_str(got), int_to_str(expected))
|
||||
}
|
||||
|
||||
println("")
|
||||
println("1. runtime semantics this fix relies on")
|
||||
|
||||
// "─" is U+2500 = E2 94 80 (three bytes). If str_len counts bytes, len("─") is 3.
|
||||
let dash: String = "─"
|
||||
assert_eq_int("str_len counts BYTES (one box-drawing char = 3)", str_len(dash), 3)
|
||||
assert_eq_int("str_char_code returns the BYTE value (lead byte of U+2500 = 0xE2 = 226)", str_char_code(dash, 0), 226)
|
||||
assert_eq_int("str_char_code second byte = 0x94 = 148", str_char_code(dash, 1), 148)
|
||||
assert_eq_int("str_char_code third byte = 0x80 = 128", str_char_code(dash, 2), 128)
|
||||
|
||||
println("")
|
||||
println("2. utf8_safe_slice — never leaves a partial character")
|
||||
|
||||
// Pure ASCII: behaves exactly like str_slice.
|
||||
assert_eq("ascii under the limit is untouched", utf8_safe_slice("hello", 10), "hello")
|
||||
assert_eq("ascii over the limit cuts exactly", utf8_safe_slice("hello world", 5), "hello")
|
||||
|
||||
// A cut landing INSIDE a 3-byte character must drop that character entirely.
|
||||
// "ab─cd": bytes a b E2 94 80 c d. Cutting at 3 or 4 lands mid-dash.
|
||||
let mixed: String = "ab─cd"
|
||||
assert_eq_int("fixture is 7 bytes (2 ascii + 3 + 2 ascii)", str_len(mixed), 7)
|
||||
assert_eq("cut inside the char (n=3) drops the partial char", utf8_safe_slice(mixed, 3), "ab")
|
||||
assert_eq("cut inside the char (n=4) drops the partial char", utf8_safe_slice(mixed, 4), "ab")
|
||||
// A cut landing exactly AFTER a complete character keeps it.
|
||||
assert_eq("cut on the char boundary (n=5) keeps the whole char", utf8_safe_slice(mixed, 5), "ab─")
|
||||
|
||||
// 2-byte character (é = C3 A9) and 4-byte character (😀 = F0 9F 98 80).
|
||||
let acc: String = "xé"
|
||||
assert_eq("cut inside a 2-byte char drops it", utf8_safe_slice(acc, 2), "x")
|
||||
assert_eq("cut after a 2-byte char keeps it", utf8_safe_slice(acc, 3), "xé")
|
||||
let emo: String = "x😀"
|
||||
assert_eq("cut inside a 4-byte char drops it (n=3)", utf8_safe_slice(emo, 3), "x")
|
||||
assert_eq("cut inside a 4-byte char drops it (n=4)", utf8_safe_slice(emo, 4), "x")
|
||||
assert_eq("cut after a 4-byte char keeps it", utf8_safe_slice(emo, 5), "x😀")
|
||||
|
||||
println("")
|
||||
println("3. the real-world shape that produced the bug")
|
||||
|
||||
// A run of box-drawing rules, cut mid-character — the exact captured failure.
|
||||
let rules: String = "──────"
|
||||
assert_eq_int("six box rules = 18 bytes", str_len(rules), 18)
|
||||
// n=16 lands one byte into the sixth character.
|
||||
let cut16: String = utf8_safe_slice(rules, 16)
|
||||
assert_eq_int("cut at 16 backs off to a clean 15-byte boundary", str_len(cut16), 15)
|
||||
// Every byte of the result must belong to a complete character: the last byte of a
|
||||
// well-formed run of these is always 0x80, and 15 is divisible by 3.
|
||||
assert_eq_int("result ends on a complete char (last byte 0x80)", str_char_code(cut16, 14), 128)
|
||||
|
||||
println("")
|
||||
println("test_utf8_slice.el: " + int_to_str(pass_count) + " passed, " + int_to_str(fail_count) + " failed")
|
||||
Reference in New Issue
Block a user