Your download filename breaks on the first accented character
published
TL;DR
Content-Disposition: attachment; filename="café.csv" is not a UTF-8 field. Browsers that read it as ISO-8859-1 save the file as café.csv. Send filename*=UTF-8''caf%C3%A9.csv instead — and do not build that value with bare encodeURIComponent, because it leaves ' and * unescaped and those are illegal there.
The problem
You generate a CSV export named after something the user typed. It works all through development, because your test data is report.csv. Then a real user exports café ñandú.csv and the file lands in their Downloads folder as:
café ñandú.csv
Nothing threw. The response was a 200. The bytes of the file are perfect. Only the name is wrong, which is the kind of bug that gets filed as “the export is broken” three weeks later.
Every claim in this post was executed on Node v26.2.0. That exact mojibake is reproducible in one line:
Buffer.from('café ñandú.csv', 'utf8').toString('latin1')
// → 'café ñandú.csv'
That is the whole bug. You wrote UTF-8 bytes into a header field that is being decoded as ISO-8859-1, and each two-byte sequence came out as two Latin-1 characters.
Why it happens
HTTP header field values are not UTF-8 by default, and Content-Disposition inherits that. MDN’s guidance on the plain filename parameter is to stay inside ASCII:
Prefer ASCII characters if possible (the client may percent-encode it, as long as the server implementation decodes it).
Percent-encoding the plain filename is not the escape hatch it looks like either:
Avoid percent escape sequences in
filename, because they are handled inconsistently across browsers. (Firefox and Chrome decode them, while Safari does not.)
So the naive fix — percent-encode and hope — gives you a file called caf%C3%A9.csv on one browser and café.csv on another. That is worse than the mojibake, because now the behaviour depends on the client.
The actual mechanism is a separate parameter, filename*, whose value uses the extended syntax from RFC 5987 (now RFC 8187): a charset, an optional language, and a percent-encoded value, joined by single quotes.
| Header sent | What the browser saves |
|---|---|
filename="café.csv" | café.csv — UTF-8 bytes read as Latin-1 |
filename="caf%C3%A9.csv" | café.csv on Chrome/Firefox, caf%C3%A9.csv on Safari |
filename*=UTF-8''caf%C3%A9.csv | café.csv everywhere that understands the parameter |
When both are present, MDN is explicit about precedence:
When both
filenameandfilename*are present in a single header field value,filename*is preferred overfilenamewhen both are understood. It’s recommended to include both for maximum compatibility.
The part that bites after you know about filename*
The advice you will find is “percent-encode it with encodeURIComponent”. That is wrong at the edges, and it fails silently rather than throwing.
RFC 8187 restricts the unescaped characters in an ext-value to attr-char, which is alphanumerics plus !#$&+-.^_ and a backtick, pipe and tilde. encodeURIComponent does not escape !'()*. Two of those — the apostrophe and the asterisk — are not attr-char, and the apostrophe is the field’s own delimiter:
encodeURIComponent("quote'star*.csv")
// → "quote'star*.csv" ← nothing was escaped
A filename containing an apostrophe therefore injects a third single quote into a field whose grammar is charset'language'value. Executed against the RFC 8187 grammar, that output does not validate.
What to do
Send both parameters, and escape the extended one properly:
function contentDisposition(filename) {
// ASCII fallback for clients that ignore filename*
const ascii = filename.replace(/[^\x20-\x7e]/g, '_').replace(/["\\]/g, '_')
// RFC 8187 ext-value: escape everything encodeURIComponent leaves behind
const encoded = encodeURIComponent(filename).replace(
/['()!*]/g,
(c) => '%' + c.charCodeAt(0).toString(16).toUpperCase()
)
return `attachment; filename="${ascii}"; filename*=UTF-8''${encoded}`
}
contentDisposition('café ñandú.csv')
// attachment; filename="caf_ _and_.csv"; filename*=UTF-8''caf%C3%A9%20%C3%B1and%C3%BA.csv
contentDisposition("quote'star*.csv")
// attachment; filename="quote'star*.csv"; filename*=UTF-8''quote%27star%2A.csv
Both outputs are copied from a run on Node v26.2.0, along with отчёт.csv and a filename
containing a double quote. In all four the ext-value validates against the RFC 8187 grammar
and decodeURIComponent round-trips it back to the original name.
Three things that matter in those eight lines:
- The ASCII fallback is sanitised, not transliterated. Replacing
éwithereads nicer, and MDN suggests it, but it needs a real transliteration table to be correct across scripts — a Cyrillic or CJK filename has no ASCII equivalent at all. Substituting_is the version that cannot be wrong. - Quotes and backslashes are stripped from the fallback because it is a quoted-string. A filename containing
"ends the parameter early and everything after it is parsed as more header. - The
filename*value carries no quotes.filename*="UTF-8''x"is malformed; the ext-value is bare.
If you are on Express or Koa, res.download() and the content-disposition package already do all of this. Reach for the manual version when you are writing the header yourself — a Worker, a Lambda, a raw Response.
Caveats
- This is about the name, never the bytes. A mojibake filename means the header encoding is wrong; the file body is untouched and still valid.
filename*is widely supported now, but the fallback is not decoration: some download managers and older tooling still read onlyfilename, which is exactly why the spec’s advice is to send both.- Sanitising for the header is not sanitising for the filesystem. Path separators, reserved Windows device names (
CON,PRN,NUL), leading dots and trailing spaces are a different problem, and the browser — not your header — decides most of it. - The mojibake above assumes the client decodes as ISO-8859-1. A client that assumes UTF-8 will show the name correctly, which is why this bug reproduces for some users and not others and gets closed as “cannot reproduce”.