Cached at:
10/02/26, 08:34 AM
# Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx
Source: [https://rahuljuliato.com/posts/emacs-rx-in-typescript](https://rahuljuliato.com/posts/emacs-rx-in-typescript)
## Intro[https://rahuljuliato.com/posts/emacs-rx-in-typescript#intro](https://rahuljuliato.com/posts/emacs-rx-in-typescript#intro)
Quick, what does this match?
That's the official regexp from[semver\.org](https://semver.org/)\. It validates version numbers like:
Don't get me wrong, I love regexps, but in practice you probably spend a bunch of time writing one, testing it against some cases, and moving on, proud of your achievement\!
Some time passes and lucky future you \(or unlucky someone else\) has to change it\. Dramatic pause here\.
I bet you've been there\. Now your options are probably: decode it again from the start, rewrite the whole thing, or, in the age of AI, ask \(and hopefully not blindly accept\) an LLM for a new recipe\.
Emacs has had a nice answer for more readable regexps for a long time: the`rx`macro\. I started using it all the time in Emacs Lisp, as reviewers always suggested it to me\. Later, I started missing this DSL in JavaScript and TypeScript, so I wrote a small version of it for my projects\.
So, what about reading that SemVer regexp like`semver`in the code below?
The same strings match, and you get named groups as a bonus\. By the end of this post you'll know every piece of it\.
> **TL;DR:**jump straight to the[cheat sheet](https://rahuljuliato.com/posts/emacs-rx-in-typescript#cheat-sheet), the[side\-by\-side examples](https://rahuljuliato.com/posts/emacs-rx-in-typescript#examples-js-ts-regex-vs-rx), the[full source](https://rahuljuliato.com/posts/emacs-rx-in-typescript#full-source), or grab the[gist](https://gist.github.com/LionyxML/b6c078a2d1b13cad42666db1773a9cec)to sneak a peek at the result\.
> **NOTE:**the`RX`here has nothing to do with[RxJS](https://rxjs.dev/), which is an amazing library for reactive programming with observables\.
## A taste of rx in Emacs Lisp[https://rahuljuliato.com/posts/emacs-rx-in-typescript#a-taste-of-rx-in-emacs-lisp](https://rahuljuliato.com/posts/emacs-rx-in-typescript#a-taste-of-rx-in-emacs-lisp)
With`rx`you describe a regexp as a tree of named forms, and Emacs turns it into the regexp string for you:
A few things to notice:
1. **Strings are literals\.**`"\("`means a parenthesis\. You don't need to escape anything by hand\.
2. **Sequence is implicit\.**Every form takes a list of things and matches them one after the other\. You don't need to wrap them in a`seq`, even though`seq`exists\.
3. **Groups appear only when needed\.**`\(\+ digit\)`becomes`\[\[:digit:\]\]\+`, not`\\\(?:\[\[:digit:\]\]\\\)\+`\.
The proposed JavaScript/TypeScript version in this post reads like this:
## Under the hood[https://rahuljuliato.com/posts/emacs-rx-in-typescript#under-the-hood](https://rahuljuliato.com/posts/emacs-rx-in-typescript#under-the-hood)
If you want strings to be literals, you can't represent a regexp piece as a plain`string`, otherwise you can't tell`"\("`\(a literal parenthesis\) apart from`"\(?:\.\.\.\)"`\(a group you built\)\. So every piece is a small object:
`src`is the regexp text\.`kind`records how that text behaves when you glue it to other things:
- `atom`: a single unit, like`a`,`\\d`,`\[a\-z\]`or`\(\.\.\.\)`\. You can put a quantifier right after it\.
- `seq`: safe to concatenate, but a quantifier needs`\(?:\.\.\.\)`around it\.`abc`is a`seq`, and so is`a\+`, since`a\+?`would silently turn into a lazy quantifier\.
- `alt`: has a`\|`at the top level, so it needs`\(?:\.\.\.\)`almost everywhere\.
Plain strings go through`literal`, which escapes them:
With that in place,`seq`joins nodes and only brackets alternations:
\(That backreference check is one of those bugs you only find by writing tests, or when it happens to you in prod\.`backref\(1\)`followed by the literal`"0"`gives you backreference number ten\.\)
Every quantifier is a`seq`of its arguments plus a suffix, bracketed only when the body isn't an atom:
Because each quantifier calls`seq`on its arguments, you get the implicit sequence for free:`optional\("\-", group\(x\)\)`becomes`\(?:\-\(x\)\)?`\.
And finally, the two entry points\. As in Emacs,`rx`returns a string\.`RX`returns a`RegExp`you can use right away:
`RX\.flags`exists because Emacs controls case folding through the`case\-fold\-search`variable, and JavaScript puts it on the regexp itself\.
That's the whole engine\! Now, let's build our vocabulary\.
## Character sets[https://rahuljuliato.com/posts/emacs-rx-in-typescript#character-sets](https://rahuljuliato.com/posts/emacs-rx-in-typescript#character-sets)
In Emacs you write`\(any "a\-z" "\_"\)`\. Inside those strings,`a\-z`is a range, and a`\-`at either end is a plain dash\. I kept the same rule:
The dash comes out escaped because sets can merge\. If you combine`anyOf\("\+\-"\)`with`anyOf\("0\-9"\)`, an unescaped`\-`would end up in the middle and create a range from`\+`to`0`\. Escaping it costs one backslash\.
And merging is the reason why`RxNode`has a`set`field\. It is there to hold the text that goes between`\[`and`\]`, so`anyOf`can take other sets as arguments:
`not`negates a set, and it knows the shorthand classes:
The simple email check, which most of us have written as`/^\[^\\s@\]\+@\[^\\s@\]\+\\\.\[^\\s@\]\+$/`at some point, becomes:
The rest of the Emacs character classes are there too:`digit`,`hexDigit`,`space`,`blank`,`wordChar`,`notWordChar`,`alpha`,`alnum`,`lower`,`upper`,`punct`,`control`,`graphic`,`printing`,`ascii`and`nonascii`\. One difference: in Emacs they understand Unicode, and mine are ASCII only\.`alpha`won't match`é`\.
Two more come from rx's symbol list, and people \(me, many times\) mix them up:
In rx,`anything`really means anything, newlines included\. Here is where the difference shows up:
## Alternatives, and the longest match[https://rahuljuliato.com/posts/emacs-rx-in-typescript#alternatives-and-the-longest-match](https://rahuljuliato.com/posts/emacs-rx-in-typescript#alternatives-and-the-longest-match)
`or`works as you'd expect, and gets bracketed when it lands inside a sequence:
Did you notice the order changed? I copied that behavior from Emacs\. When every branch of an`or`is a plain string, rx hands them to`regexp\-opt`, which builds a pattern that prefers the longest match:
JavaScript alternation takes the first branch that matches, going left to right\. So the naive regexp for a list of keywords has a 'bug':
I don't build a trie like`regexp\-opt`does\. Sorting the strings by length, longest first, is enough to get the same behavior:
As in Emacs,`or\(\)`with no branches returns`unmatchable`, which is`\(?\!\)`here\. It's handy when you build the branch list at runtime and it might come out empty\.
## Repetition, greedy and lazy[https://rahuljuliato.com/posts/emacs-rx-in-typescript#repetition-greedy-and-lazy](https://rahuljuliato.com/posts/emacs-rx-in-typescript#repetition-greedy-and-lazy)
Emacs has`\(= n \.\.\.\)`,`\(\>= n \.\.\.\)`and`\(\*\* n m \.\.\.\)`\. Here they are`repeat`,`atLeast`and`between`:
The lazy versions`\*?`,`\+?`and`??`are`zeroOrMoreLazy`,`oneOrMoreLazy`and`optionalLazy`\. The classic HTML tag example:
## Groups and backreferences[https://rahuljuliato.com/posts/emacs-rx-in-typescript#groups-and-backreferences](https://rahuljuliato.com/posts/emacs-rx-in-typescript#groups-and-backreferences)
`group`is a capturing group, and`backref`points back to it:
Emacs also has`\(group\-n N \.\.\.\)`to pick the group number\. JavaScript can't do that, but it has named groups, which serve the same purpose and read better:
`backref`accepts a name as well:
## Anchors[https://rahuljuliato.com/posts/emacs-rx-in-typescript#anchors](https://rahuljuliato.com/posts/emacs-rx-in-typescript#anchors)
`rx`distinguishes the start of the string \(`bos`\) from the start of a line \(`bol`\)\. In JavaScript both are`^`, and the`m`flag decides which one you get\. I kept both names so the intent shows in the code:
Why not add`start`and`end`automatically? Because you only want them when validating a whole string\. When searching inside a text, as in`split`,`replace`or`matchAll`, a hidden`^`and`$`would break everything\. Emacs agrees:`bos`and`eos`are explicit in`rx`too\.
`wordBoundary`and`notWordBoundary`map straight to`\\b`and`\\B`\. Emacs also has`bow`and`eow`\(`\\<`and`\\\>`\), start and end of a word\. JavaScript lacks those, so I combined`\\b`with a lookaround:
## Literal and raw[https://rahuljuliato.com/posts/emacs-rx-in-typescript#literal-and-raw](https://rahuljuliato.com/posts/emacs-rx-in-typescript#literal-and-raw)
Plain strings are already literals, but`rx`has an explicit`literal`form for strings computed at runtime, and I kept it\. It documents that the value came from somewhere else:
The opposite direction is`rx`'s`\(regexp \.\.\.\)`form, the escape hatch\. Here it's`raw`, and it receives either a string or an existing`RegExp`\. It lets you adopt the DSL in a codebase full of old regexps without rewriting all of them, like:
`raw`can't see inside the text it gets, so it adds brackets whenever it's combined with something else\. It's an extra`\(?:\)`, and the regexp still works\.
## Shall we try SemVer again?[https://rahuljuliato.com/posts/emacs-rx-in-typescript#shall-we-try-semver-again](https://rahuljuliato.com/posts/emacs-rx-in-typescript#shall-we-try-semver-again)
Back to the regexp from the intro\. In Emacs you would give names to the pieces with`rx\-define`or`rx\-let`\. In TypeScript those are just`const`s:
Now you can read the spec in the code\. A numeric identifier is`0`, or a non\-zero digit followed by any number of digits\. A pre\-release is a dotted list of identifiers, and so is build metadata\.`dotted`is a plain function returning a node, which is as far as abstraction needs to go here\.
It matches the same strings as the official regexp, and the named groups give you a result like:
Next time the spec changes, you can understand what the current regex does at a glance, instead of fighting an army of punctuation\.
## What's missing from Emacs rx[https://rahuljuliato.com/posts/emacs-rx-in-typescript#what-s-missing-from-emacs-rx](https://rahuljuliato.com/posts/emacs-rx-in-typescript#what-s-missing-from-emacs-rx)
I tried to map every`rx`form, and a few have no JavaScript equivalent:
- `point`: JavaScript regexps don't know about a cursor\.
- `symbol\-start`,`symbol\-end`,`syntax`,`category`: these depend on Emacs syntax tables\.
- `intersection`: possible with the`v`flag, but I haven't needed it\.
- `minimal\-match`/`maximal\-match`: these flip the greediness of everything inside them\. Doable, but it would need a separate pass, and the`\*Lazy`functions cover my use cases\.
- `eval`: TypeScript already evaluates expressions everywhere, so you get it for free\.
And one addition Emacs doesn't need:`RX\.flags`\.
## Cheat sheet[https://rahuljuliato.com/posts/emacs-rx-in-typescript#cheat-sheet](https://rahuljuliato.com/posts/emacs-rx-in-typescript#cheat-sheet)
Emacs rxTypeScriptJS regexp \(roughly\)`seq`,`:`,`and``seq\(\.\.\.\)`, implicit in every form`ab``or`,`\|``or\(\.\.\.\)``a\|b``any`,`in`,`char``anyOf\("a\-z", "\_", digit\)``\[a\-z\_\\d\]``not\-char``notChar\(\.\.\.\)``\[^\.\.\.\]``not``not\(charset\)``\\D`,`\[^\.\.\.\]``\*`,`\+`,`?``zeroOrMore`,`oneOrMore`,`optional``x\*`,`x\+`,`x?``\*?`,`\+?`,`??``zeroOrMoreLazy`,`oneOrMoreLazy`,`optionalLazy``x\*?`,`x\+?`,`x??``=`,`\>=`,`\*\*``repeat`,`atLeast`,`between``x\{n\}`,`x\{n,\}`,`x\{n,m\}``group``group\(\.\.\.\)``\(\.\.\.\)``group\-n``named\("name", \.\.\.\)``\(?<name\>\.\.\.\)``backref``backref\(1\)`,`backref\("name"\)``\\1`,`\\k<name\>``literal``literal\(s\)``s`, escaped:`1\\\+1``regexp`,`regex``raw\("\.\.\."\)`,`raw\(/\.\.\./\)``\(?:\.\.\.\)`, as\-is`rx\-define`,`rx\-let``const`\(none\)`bos`,`eos``start`,`end``^`,`$``bol`,`eol``lineStart`,`lineEnd`\(with the`m`flag\)`^`,`$``bow`,`eow``wordStart`,`wordEnd``\\b\(?=\\w\)`,`\\b\(?<=\\w\)``word\-boundary``wordBoundary``\\b``not\-word\-boundary``notWordBoundary``\\B``nonl`,`not\-newline``notNewline``\.``anychar`,`anything``anything``\[\\s\\S\]``unmatchable``unmatchable``\(?\!\)``digit``digit``\\d``hex\-digit`,`xdigit``hexDigit``\[0\-9a\-fA\-F\]``space`,`whitespace``space``\\s``blank``blank``\[ \\t\]``word`,`wordchar``wordChar``\\w``not\-wordchar``notWordChar``\\W``alpha`,`letter``alpha``\[a\-zA\-Z\]``alnum``alnum``\[a\-zA\-Z0\-9\]``lower`,`upper``lower`,`upper``\[a\-z\]`,`\[A\-Z\]``punct`,`punctuation``punct``\[\!\-/:\-@\[\-\`\{\-~\]``cntrl`,`control``control``\[\\x00\-\\x1f\\x7f\]``graph`,`graphic``graphic``\[\!\-~\]``print`,`printing``printing``\[ \-~\]``ascii`,`nonascii``ascii`,`nonascii``\[\\x00\-\\x7f\]`,`\[\\u0080\-\\uffff\]`## Examples: JS/TS regex vs RX[https://rahuljuliato.com/posts/emacs-rx-in-typescript#examples-js-ts-regex-vs-rx](https://rahuljuliato.com/posts/emacs-rx-in-typescript#examples-js-ts-regex-vs-rx)
Each example below shows the goal, the Emacs`rx`form in a comment, the regexp you would write by hand, and the`RX`version\. When`RX`produces a different regexp text, the`// =\>`line shows it\. The results at the bottom come from running both against the same strings\.
### Digits only[https://rahuljuliato.com/posts/emacs-rx-in-typescript#digits-only](https://rahuljuliato.com/posts/emacs-rx-in-typescript#digits-only)
The whole string is digits\.
### Letters only[https://rahuljuliato.com/posts/emacs-rx-in-typescript#letters-only](https://rahuljuliato.com/posts/emacs-rx-in-typescript#letters-only)
The whole string is ASCII letters\.
### Optional letter[https://rahuljuliato.com/posts/emacs-rx-in-typescript#optional-letter](https://rahuljuliato.com/posts/emacs-rx-in-typescript#optional-letter)
Both spellings, color and colour\.
### Two words[https://rahuljuliato.com/posts/emacs-rx-in-typescript#two-words](https://rahuljuliato.com/posts/emacs-rx-in-typescript#two-words)
Two words separated by a space\.
### Phone number[https://rahuljuliato.com/posts/emacs-rx-in-typescript#phone-number](https://rahuljuliato.com/posts/emacs-rx-in-typescript#phone-number)
\(123\) 456\-7890, parentheses and all\.
### Hex color[https://rahuljuliato.com/posts/emacs-rx-in-typescript#hex-color](https://rahuljuliato.com/posts/emacs-rx-in-typescript#hex-color)
`\#ff00aa`\-style colors\.
### Signed integer[https://rahuljuliato.com/posts/emacs-rx-in-typescript#signed-integer](https://rahuljuliato.com/posts/emacs-rx-in-typescript#signed-integer)
An optional sign, then digits\.
### Simple email[https://rahuljuliato.com/posts/emacs-rx-in-typescript#simple-email](https://rahuljuliato.com/posts/emacs-rx-in-typescript#simple-email)
[Something@something\.something](mailto:
[email protected]), no spaces\.
### No digits[https://rahuljuliato.com/posts/emacs-rx-in-typescript#no-digits](https://rahuljuliato.com/posts/emacs-rx-in-typescript#no-digits)
A string without any digit\.
### CSV line[https://rahuljuliato.com/posts/emacs-rx-in-typescript#csv-line](https://rahuljuliato.com/posts/emacs-rx-in-typescript#csv-line)
Exactly three comma\-separated fields\.
### One of many[https://rahuljuliato.com/posts/emacs-rx-in-typescript#one-of-many](https://rahuljuliato.com/posts/emacs-rx-in-typescript#one-of-many)
A fixed list of words\.
### Title and name[https://rahuljuliato.com/posts/emacs-rx-in-typescript#title-and-name](https://rahuljuliato.com/posts/emacs-rx-in-typescript#title-and-name)
`mr`or`ms`, then a name, keeping the title\.
### Between[https://rahuljuliato.com/posts/emacs-rx-in-typescript#between](https://rahuljuliato.com/posts/emacs-rx-in-typescript#between)
Two to four digits\.
### At least[https://rahuljuliato.com/posts/emacs-rx-in-typescript#at-least](https://rahuljuliato.com/posts/emacs-rx-in-typescript#at-least)
Three or more digits, anywhere\.
### Repeated word[https://rahuljuliato.com/posts/emacs-rx-in-typescript#repeated-word](https://rahuljuliato.com/posts/emacs-rx-in-typescript#repeated-word)
The same word twice\.
### Matching tags[https://rahuljuliato.com/posts/emacs-rx-in-typescript#matching-tags](https://rahuljuliato.com/posts/emacs-rx-in-typescript#matching-tags)
An open tag and its own closing tag\.
### Whole word[https://rahuljuliato.com/posts/emacs-rx-in-typescript#whole-word](https://rahuljuliato.com/posts/emacs-rx-in-typescript#whole-word)
`cat`as a word, not inside another one\.
### Case\-insensitive[https://rahuljuliato.com/posts/emacs-rx-in-typescript#case-insensitive](https://rahuljuliato.com/posts/emacs-rx-in-typescript#case-insensitive)
`hello`, in any case\.
## Full source[https://rahuljuliato.com/posts/emacs-rx-in-typescript#full-source](https://rahuljuliato.com/posts/emacs-rx-in-typescript#full-source)
It's a single file with no dependencies\. Copy it into your project and start deleting the forms you don't need, or adding the ones you miss\.
You can check the same code, plus all the examples from this post \(and a few more\), in this[gist](https://gist.github.com/LionyxML/b6c078a2d1b13cad42666db1773a9cec)\. If you'd rather not set anything up, paste it into the[TypeScript Playground](https://www.typescriptlang.org/play), hit "Run", and check the "Logs" tab\.
## Wrapping up[https://rahuljuliato.com/posts/emacs-rx-in-typescript#wrapping-up](https://rahuljuliato.com/posts/emacs-rx-in-typescript#wrapping-up)
None of this is new\. On the Emacs side, as I said before,`rx`has shipped for decades, and the Elisp version is more complete than mine\.
The idea of describing patterns with a small DSL instead of raw syntax isn't new either\. Plenty of people have tried it, each in their own way\. One project I like a lot in this space is[Zod](https://zod.dev/), which I wrote about in my[Zod quick tutorial](https://rahuljuliato.com/posts/zod-tutorial)\. It's not a regexp builder: you compose small schema pieces, and Zod gives you back a parser and a TypeScript type from the same construction\. It follows the same spirit, though: build big things out of small named pieces you can read\.
If you write Elisp and have never tried`rx`, open`\*scratch\*`, type`\(rx \(\+ digit\)\)`, and`C\-x C\-e`it\. If you write JavaScript or TypeScript, the file above is yours\. And if you port it to another language, send me a link\.