~/blog

TextDecoder without stream:true mangles split UTF-8 chunks

published

#encoding#streams#javascript

TL;DR

decoder.decode(chunk) defaults to stream: false, which flushes the decoder at the end of every call. A multi-byte character split across two chunks becomes replacement characters — é turns into ��, two of them, not one. Pass { stream: true } for every chunk except the last, or skip the manual loop and use TextDecoderStream.

The problem

You are reading a response body chunk by chunk, because you want progress or because the payload is too big to buffer:

const decoder = new TextDecoder()
let text = ''
for await (const chunk of response.body) {
  text += decoder.decode(chunk)   // <- the bug
}

This passes every test you write. Then a real 4 MB export goes through and a handful of characters in the middle are wrong:

Jos� Mart�nez
S�o Paulo

Not all of them. Roughly one per chunk boundary, at positions that move when the payload size changes. Re-running gives a different set of broken characters, which is how this ends up filed as a flaky network bug.

Why it happens

Two facts collide.

Chunks split on byte boundaries, not character boundaries. A ReadableStream from fetch hands you whatever bytes arrived. Nothing aligns them to UTF-8 sequences. é is C3 A9; a chunk can end after C3 and the next one begins with A9.

decode() flushes unless you tell it not to. MDN documents stream as “a boolean flag indicating whether additional data will follow in subsequent calls to decode()… It defaults to false.” Flushing means the decoder handles end-of-queue: the UTF-8 decoder spec says “if byte is end-of-queue and UTF-8 bytes needed is not 0, then set UTF-8 bytes needed to 0 and return error”, and in the default replacement error mode an error emits U+FFFD.

So the split é produces two replacement characters, from two separate failures:

CallBytes seenDecoder stateOutput
decode(chunkA)… C3needs 1 more byte, then hits end-of-queue… + U+FFFD
decode(chunkB)A9 …continuation byte with 0 bytes neededU+FFFD + …

A three-byte character split after its first byte yields three. This is also why the bug is invisible in development: ASCII is single-byte, so a chunk boundary in hello world can never fall inside a character.

With stream: true the decoder keeps its partial sequence between calls. The spec puts it plainly: “the way streaming works is to not handle end-of-queue here when this’s do not flush is true… That way in a subsequent invocation this’s decoder is not set anew in the first step of the algorithm and its state is preserved.”

What to do

Option 1 — set the flag, and flush at the end. The final call matters: without it, a truncated sequence at the very end of the body is silently dropped instead of becoming U+FFFD.

const decoder = new TextDecoder()
let text = ''
for await (const chunk of response.body) {
  text += decoder.decode(chunk, { stream: true })
}
text += decoder.decode()   // flush: no args means stream:false, empty input

Option 2 — let the platform do it. TextDecoderStream is a transform stream that holds the same state for you, and it composes with anything else in the pipeline:

const stream = response.body.pipeThrough(new TextDecoderStream())
for await (const str of stream) {
  process(str)              // already correctly decoded text
}

Option 3 — do not stream at all. If you do not actually need incremental output, await response.text() reads the whole body and decodes it once. The streaming loop is only worth writing when you want the first bytes before the last ones arrive.

A quick check that reproduces it without a network at all:

const bytes = new TextEncoder().encode('café')   // 63 61 66 C3 A9
const d = new TextDecoder()

// split the é in half
d.decode(bytes.slice(0, 4)) + d.decode(bytes.slice(4))
// 'caf�' + '�'  ->  'caf��'

const s = new TextDecoder()
s.decode(bytes.slice(0, 4), { stream: true }) + s.decode(bytes.slice(4))
// 'caf' + 'é'  ->  'café'

Caveats

References