~/blog

The u flag makes /[a-z]/i match two non-ASCII characters

published

#javascript#regex#unicode

TL;DR

/^[a-z]+$/i and /^[a-z]+$/iu are not the same validator. The u flag switches case-insensitive matching from Unicode Default Case Conversion to simple case folding, and simple case folding maps some characters from outside Basic Latin into it. Two of them reach [a-z]: K (Kelvin sign) and ſ (long s). So an ASCII-only gate gets looser the moment you add u. Keep u, drop i, and write both ranges: /^[a-zA-Z]+$/u.

The problem

You have a username validator. It is deliberately strict — ASCII letters only, nothing clever:

const USERNAME = /^[a-z]+$/i;

A linter asks you to add the u flag. ESLint’s require-unicode-regexp exists for good reasons: u makes surrogate pairs behave, turns invalid escapes into errors instead of silent literals, and switches off a pile of Annex B legacy behaviour. It reads like a pure correctness upgrade, so it sails through review:

const USERNAME = /^[a-z]+$/iu;

The validator is now more permissive than it was before. On Node v26.2.0:

const KELVIN = 'K'; // KELVIN SIGN
const LONGS = 'ſ'; // LATIN SMALL LETTER LONG S

/^[a-z]+$/i.test(KELVIN); // false
/^[a-z]+$/iu.test(KELVIN); // true   <-- added u, got looser

/^[a-z]+$/i.test(LONGS); // false
/^[a-z]+$/iu.test(LONGS); // true

It is not confined to the lowercase range, and it reaches \w too:

/^[A-Z]+$/iu.test(KELVIN); // true
/^\w+$/iu.test(KELVIN); // true

So these both pass a check that was written to mean “ASCII letters, nothing else”:

/^[a-z]+$/iu.test('admi' + LONGS); // true
/^[a-zA-Z]+$/iu.test('Ban' + KELVIN); // true

The second one is the interesting shape. BanK renders as BanK in almost every font — it is the Kelvin sign, which looks exactly like a capital K.

Why it happens

It is the i flag that changes meaning, not the u flag. MDN’s RegExp.prototype.ignoreCase page states both halves.

Without u, case-insensitive matching uses Unicode Default Case Conversion, the same algorithm behind toUpperCase, and that algorithm is explicitly barred from crossing into Basic Latin:

If the regex is Unicode-unaware, case mapping uses the Unicode Default Case Conversion — the same algorithm used in String.prototype.toUpperCase(). This algorithm prevents code points outside the Basic Latin block to be mapped to code points within it, so ſ and K mentioned previously are not matched by /[a-z]/i.

With u, it uses simple case folding instead, and simple case folding has no such rule:

If the regex is Unicode-aware, the case mapping happens through simple case folding specified in CaseFolding.txt … It may however map code points outside the Basic Latin block to code points within it — for example, ſ (U+017F LATIN SMALL LETTER LONG S) case-folds to s (U+0073) and K (U+212A KELVIN SIGN) case-folds to k (U+006B). Therefore, ſ and K can be matched by /[a-z]/ui.

Two checks confirm it is the interaction and not the u flag alone. Without i, u matches neither character:

/^[a-z]+$/u.test(KELVIN); // false
/^[a-z]+$/u.test(LONGS); // false

And the v flag inherits the same folding:

/^[a-z]+$/iv.test(KELVIN); // true

ESLint documents the side effect on its own rule page, which is worth knowing before you argue with a reviewer about it:

In some cases, adding the u flag to a regular expression using both the i flag and the \w character class can change its behavior due to Unicode case folding.

The part that turns it into a bug

A widened validator on its own is a nuisance. It becomes a real problem because the regex and your normalizer disagree about what the string is, and they disagree in opposite directions for the two characters:

KELVIN.toLowerCase(); // 'k'   — collapses onto ASCII k
LONGS.toLowerCase(); // 'ſ'   — unchanged
LONGS.toUpperCase(); // 'S'   — collapses onto ASCII S

So the classic validate-then-normalize pipeline lets a collision through:

const taken = new Set(['bank']);

const candidate = 'Ban' + KELVIN; // renders as "BanK"

/^[a-zA-Z]+$/iu.test(candidate); // true  — validator says fine
taken.has(candidate.toLowerCase()); // true  — but it lowercases to "bank"

Whether that is a duplicate-account bug, a bypassed uniqueness check or a lookalike-name problem depends on which side runs first. All three come from the same place: the regex approved a string that is not the string anything downstream will handle.

What to do

Keep u. The linter is right that you want it. Move the case handling out of the flag and into the pattern:

const USERNAME = /^[a-zA-Z]+$/u;

Verified against the four strings that break the original, plus four that must still pass:

PatternBanKadmiſabc / ABCASCII-only?
/^[a-z]+$/iuacceptsacceptsacceptsno
/^[a-z]+$/irejectsrejectsacceptsyes, but loses u
/^[a-zA-Z]+$/urejectsrejectsacceptsyes
/^[\x41-\x5A\x61-\x7A]+$/urejectsrejectsacceptsyes

The last row is the same thing spelled in codepoints. Use it when you want the intent to survive a future reader who “simplifies” the ranges back into [a-z] with an i.

If your gate is genuinely about codepoints rather than letters, test them directly and skip the character class:

const isAsciiLetters = (s) =>
  s.length > 0 &&
  [...s].every((c) => {
    const cp = c.codePointAt(0);
    return (cp >= 0x41 && cp <= 0x5a) || (cp >= 0x61 && cp <= 0x7a);
  });

Caveats

References