TextDecoder without stream:true mangles split UTF-8 chunks
published
TL;DR
decoder.decode(chunk) defaults to stream: false, which flushes the decoder at the end of every call. A multi-byte character split across two chunks becomes replacement characters — é turns into ��, two of them, not one. Pass { stream: true } for every chunk except the last, or skip the manual loop and use TextDecoderStream.
The problem
You are reading a response body chunk by chunk, because you want progress or because the payload is too big to buffer:
const decoder = new TextDecoder()
let text = ''
for await (const chunk of response.body) {
text += decoder.decode(chunk) // <- the bug
}
This passes every test you write. Then a real 4 MB export goes through and a handful of characters in the middle are wrong:
Jos� Mart�nez
S�o Paulo
Not all of them. Roughly one per chunk boundary, at positions that move when the payload size changes. Re-running gives a different set of broken characters, which is how this ends up filed as a flaky network bug.
Why it happens
Two facts collide.
Chunks split on byte boundaries, not character boundaries. A ReadableStream from fetch hands you whatever bytes arrived. Nothing aligns them to UTF-8 sequences. é is C3 A9; a chunk can end after C3 and the next one begins with A9.
decode() flushes unless you tell it not to. MDN documents stream as “a boolean flag indicating whether additional data will follow in subsequent calls to decode()… It defaults to false.” Flushing means the decoder handles end-of-queue: the UTF-8 decoder spec says “if byte is end-of-queue and UTF-8 bytes needed is not 0, then set UTF-8 bytes needed to 0 and return error”, and in the default replacement error mode an error emits U+FFFD.
So the split é produces two replacement characters, from two separate failures:
| Call | Bytes seen | Decoder state | Output |
|---|---|---|---|
decode(chunkA) | … C3 | needs 1 more byte, then hits end-of-queue | … + U+FFFD |
decode(chunkB) | A9 … | continuation byte with 0 bytes needed | U+FFFD + … |
A three-byte character split after its first byte yields three. This is also why the bug is invisible in development: ASCII is single-byte, so a chunk boundary in hello world can never fall inside a character.
With stream: true the decoder keeps its partial sequence between calls. The spec puts it plainly: “the way streaming works is to not handle end-of-queue here when this’s do not flush is true… That way in a subsequent invocation this’s decoder is not set anew in the first step of the algorithm and its state is preserved.”
What to do
Option 1 — set the flag, and flush at the end. The final call matters: without it, a truncated sequence at the very end of the body is silently dropped instead of becoming U+FFFD.
const decoder = new TextDecoder()
let text = ''
for await (const chunk of response.body) {
text += decoder.decode(chunk, { stream: true })
}
text += decoder.decode() // flush: no args means stream:false, empty input
Option 2 — let the platform do it. TextDecoderStream is a transform stream that holds the same state for you, and it composes with anything else in the pipeline:
const stream = response.body.pipeThrough(new TextDecoderStream())
for await (const str of stream) {
process(str) // already correctly decoded text
}
Option 3 — do not stream at all. If you do not actually need incremental output, await response.text() reads the whole body and decodes it once. The streaming loop is only worth writing when you want the first bytes before the last ones arrive.
A quick check that reproduces it without a network at all:
const bytes = new TextEncoder().encode('café') // 63 61 66 C3 A9
const d = new TextDecoder()
// split the é in half
d.decode(bytes.slice(0, 4)) + d.decode(bytes.slice(4))
// 'caf�' + '�' -> 'caf��'
const s = new TextDecoder()
s.decode(bytes.slice(0, 4), { stream: true }) + s.decode(bytes.slice(4))
// 'caf' + 'é' -> 'café'
Caveats
- Reusing a decoder across unrelated bodies is its own bug. The state that makes streaming work is per-instance. If you stop mid-stream without a final flush and then decode a different payload with the same
TextDecoder, the leftover partial sequence is prepended to the new one. Either flush or construct a new decoder per body. fatal: truechanges the symptom, not the cause. With that option a split sequence throws aTypeErrorinstead of producing U+FFFD. That is often better — a loud failure beats silent mojibake — but the fix is still thestreamflag.- This is not UTF-8-specific. Any multi-byte encoding a
TextDecodersupports has the same boundary problem; UTF-16 splits on two-byte units, so it breaks on far more characters than UTF-8 does. ignoreBOMand the BOM are unaffected. The BOM is consumed on the first decode of a stream, which is another reason not to recycle decoders between bodies.- Node’s
string_decodermodule exists for the same reason and predatesTextDecoder. If you are in Node and already using it,StringDecoder#writeis the streaming call and#endis the flush — same two-method shape.