runtime: make valid UTF-8 the JSON emitter's contract #148

Merged
will.anderson merged 1 commits from fix/utf8-truncation into dev 2026-08-16 17:03:49 +00:00
Owner

Three nodes in the live graph carry labels truncated to exactly 80 bytes ending in a lone 0xE2 — the first byte of an em-dash, cut mid-sequence. jb_emit_escaped copied every byte ≥ 0x20 through verbatim, so those three nodes made the entire /api/nodes/list response undecodable and no strict parser could read the graph at all.

build size result
production binary 25,929,607 B INVALID at byte 89260
this build 26,338,389 B VALID — parses to 13,630 nodes

The damage was not written by this runtime

No 80-byte truncation exists here (the only label truncation is engram_first_n_chars at 60), and the content of those nodes is 2572 and 2746 bytes. Some other producer wrote them.

That is exactly why fixing a writer could not have fixed this: the store already holds the damage, and it accepts data from importers, other producers and older binaries. So the fix goes where the promise is made. A serializer that emits JSON owes valid UTF-8 whatever it is handed.

jb_emit_escaped now validates each multi-byte sequence before emitting any of it, substituting U+FFFD for a bad lead byte, a missing or malformed continuation, an overlong encoding, a UTF-16 surrogate, or a codepoint above U+10FFFF. Invalid bytes are replaced, not dropped, so damage stays visible rather than silently papered over. Well-formed input is byte-identical to before.

Second change — preventive, and explicitly not the cause of the above

engram_first_n_chars truncated by bytes despite its name, so content with a multi-byte character crossing byte 60 would produce a half codepoint. It now uses el_utf8_safe_len, which returns the largest byte length ≤ max that does not split a codepoint. Bounded by bytes, not codepoints, so existing labels never grow — they only stop splitting.

el_utf8_safe_len lives beside str_count_chars rather than in the engram, because the rest of el’s string layer is already codepoint-aware (str_count_chars counts codepoints, str_reverse walks codepoint lengths). Byte truncation was the outlier, and the concern is a string concern.

Note on the investigation

I first “fixed” the truncator and wrote a test that passed on the unpatched build too — because route_create_node passes label = content when no label is supplied, so engram_first_n_chars is never reached over HTTP. The test proved nothing. The real cause was found only by decoding the actual failing bytes out of the live response.

Three nodes in the live graph carry labels truncated to exactly **80 bytes ending in a lone `0xE2`** — the first byte of an em-dash, cut mid-sequence. `jb_emit_escaped` copied every byte ≥ 0x20 through verbatim, so those three nodes made the **entire** `/api/nodes/list` response undecodable and no strict parser could read the graph at all. | build | size | result | |---|---|---| | production binary | 25,929,607 B | **INVALID at byte 89260** | | this build | 26,338,389 B | **VALID — parses to 13,630 nodes** | ### The damage was not written by this runtime No 80-byte truncation exists here (the only label truncation is `engram_first_n_chars` at 60), and the content of those nodes is 2572 and 2746 bytes. Some other producer wrote them. That is exactly why fixing a *writer* could not have fixed this: the store already holds the damage, and it accepts data from importers, other producers and older binaries. **So the fix goes where the promise is made.** A serializer that emits JSON owes valid UTF-8 whatever it is handed. `jb_emit_escaped` now validates each multi-byte sequence before emitting any of it, substituting U+FFFD for a bad lead byte, a missing or malformed continuation, an overlong encoding, a UTF-16 surrogate, or a codepoint above U+10FFFF. Invalid bytes are **replaced, not dropped**, so damage stays visible rather than silently papered over. Well-formed input is byte-identical to before. ### Second change — preventive, and explicitly *not* the cause of the above `engram_first_n_chars` truncated by **bytes** despite its name, so content with a multi-byte character crossing byte 60 would produce a half codepoint. It now uses `el_utf8_safe_len`, which returns the largest byte length ≤ max that does not split a codepoint. Bounded by bytes, not codepoints, so existing labels never grow — they only stop splitting. `el_utf8_safe_len` lives beside `str_count_chars` rather than in the engram, because the rest of el’s string layer is already codepoint-aware (`str_count_chars` counts codepoints, `str_reverse` walks codepoint lengths). Byte truncation was the outlier, and the concern is a string concern. ### Note on the investigation I first “fixed” the truncator and wrote a test that **passed on the unpatched build too** — because `route_create_node` passes `label = content` when no label is supplied, so `engram_first_n_chars` is never reached over HTTP. The test proved nothing. The real cause was found only by decoding the actual failing bytes out of the live response.
will.anderson added 1 commit 2026-08-16 17:03:31 +00:00
runtime: make valid UTF-8 the JSON emitter's contract
El SDK CI - dev / build-and-test (pull_request) Failing after 10m36s
8a307dfd42
Three nodes in the live graph carry labels truncated to exactly 80 bytes
ending in a lone 0xE2 — the first byte of an em-dash, cut mid-sequence.
jb_emit_escaped copied every byte >= 0x20 through verbatim, so those three
nodes made the ENTIRE /api/nodes/list response undecodable and no strict
parser could read the graph at all.

  production binary   25,929,607 bytes   INVALID at byte 89260
  this build          26,338,389 bytes   VALID, parses to 13,630 nodes

The damage was NOT written by this runtime. No 80-byte truncation exists
here (the only label truncation is engram_first_n_chars at 60), and the
content of those nodes is 2572 and 2746 bytes. Some other producer wrote
them. That is exactly why fixing a writer could not have fixed this: the
store already holds the damage, and it accepts data from importers, other
producers and older binaries.

So the fix goes where the promise is made. A serializer that emits JSON
owes valid UTF-8 whatever it is handed. jb_emit_escaped now validates each
multi-byte sequence before emitting any of it and substitutes U+FFFD for a
bad lead byte, a missing or malformed continuation, an overlong encoding, a
UTF-16 surrogate, or a codepoint above U+10FFFF. Invalid bytes are REPLACED
rather than dropped, so the damage stays visible in the output instead of
being silently papered over. Well-formed input is byte-identical to before.

Second, preventive and explicitly NOT the cause of the above:
engram_first_n_chars truncated by BYTES despite its name, so content with a
multi-byte character crossing byte 60 would produce a half codepoint in the
label. It now uses el_utf8_safe_len, which returns the largest byte length
<= max that does not split a codepoint. Bounded by bytes, not codepoints,
so existing labels never grow — they only stop splitting.

el_utf8_safe_len lives beside str_count_chars rather than in the engram
because the rest of el's string layer is already codepoint-aware
(str_count_chars counts codepoints, str_reverse walks codepoint lengths).
Byte truncation was the outlier and the concern is a string concern.

Note on the investigation: I first "fixed" the truncator and wrote a test
that passed on the UNPATCHED build too, because route_create_node passes
label = content when no label is supplied, so engram_first_n_chars is never
reached over HTTP. The test proved nothing. The real cause was only found
by decoding the actual failing bytes out of the live response.
will.anderson merged commit d41645388a into dev 2026-08-16 17:03:49 +00:00
Sign in to join this conversation.