Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' Rx

Intro

Quick, what does this match?

/^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-((?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\.(?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\+([0-9a-zA-Z-]+(?:\.[0-9a-zA-Z-]+)*))?$/;

// Take
// your
// time...
//
// ...still decoding?
//
// OK, keep reading :)

That's the official regexp from semver.org. It validates version numbers like:

// matches
"1.2.3"
"0.10.0"
"2.0.0-rc.1"
"1.0.0-alpha.1+build.5"
"1.0.0+20260930"

// doesn't match
"01.2.3"   // leading zero
"1.2"      // missing patch
"v1.2.3"   // no "v" prefix allowed
"1.0.0-01" // numeric pre-release with a leading zero
"1.2.3-"   // empty pre-release

Don't get me wrong, I love regexps, but in practice you probably spend a bunch of time writing one, testing it against some cases, and moving on, proud of your achievement!

Some time passes and lucky future you (or unlucky someone else) has to change it. Dramatic pause here.

I bet you've been there. Now your options are probably: decode it again from the start, rewrite the whole thing, or, in the age of AI, ask (and hopefully not blindly accept) an LLM for a new recipe.

Emacs has had a nice answer for more readable regexps for a long time: the rx macro. I started using it all the time in Emacs Lisp, as reviewers always suggested it to me. Later, I started missing this DSL in JavaScript and TypeScript, so I wrote a small version of it for my projects.

So, what about reading that SemVer regexp like semver in the code below?

const num = or("0", seq(anyOf("1-9"), zeroOrMore(digit)));
const idChar = anyOf(alnum, "-");
const preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, "-"), zeroOrMore(idChar)));
const dotted = (x: Item) => seq(x, zeroOrMore(".", x));

const semver = RX(
  start,
  named("major", num), ".",
  named("minor", num), ".",
  named("patch", num),
  optional("-", named("pre", dotted(preId))),
  optional("+", named("build", dotted(oneOrMore(idChar)))),
  end,
);

The same strings match, and you get named groups as a bonus. By the end of this post you'll know every piece of it.

TL;DR: jump straight to the , the , the , or grab the gist to sneak a peek at the result.
NOTE: the RX here has nothing to do with RxJS, which is an amazing library for reactive programming with observables.

A taste of rx in Emacs Lisp

With rx you describe a regexp as a tree of named forms, and Emacs turns it into the regexp string for you:

(rx bos (+ digit) eos)
;; => "\\`[[:digit:]]+\\'"

(rx bol "colo" (? "u") "r" eol)
;; => "^colou?r$"

(rx bos "(" (= 3 digit) ")" space (= 3 digit) "-" (= 4 digit) eos)
;; => "\\`([[:digit:]]\\{3\\})[[:space:]][[:digit:]]\\{3\\}-[[:digit:]]\\{4\\}\\'"

A few things to notice:

  1. Strings are literals. "(" means a parenthesis. You don't need to escape anything by hand.
  2. Sequence is implicit. Every form takes a list of things and matches them one after the other. You don't need to wrap them in a seq, even though seq exists.
  3. Groups appear only when needed. (+ digit) becomes [[:digit:]]+, not \(?:[[:digit:]]\)+.

The proposed JavaScript/TypeScript version in this post reads like this:

const phone = RX(
  start, "(", repeat(3, digit), ")", space,
  repeat(3, digit), "-", repeat(4, digit), end,
);
// => /^\(\d{3}\)\s\d{3}-\d{4}$/

Under the hood

If you want strings to be literals, you can't represent a regexp piece as a plain string, otherwise you can't tell "(" (a literal parenthesis) apart from "(?:...)" (a group you built). So every piece is a small object:

type Kind = "atom" | "seq" | "alt";

interface RxNode {
  readonly src: string;
  readonly kind: Kind;
  readonly set?: string; // char sets only, see below
  readonly neg?: boolean;
}

type Item = string | RxNode;

src is the regexp text. kind records how that text behaves when you glue it to other things:

  • atom: a single unit, like a, \d, [a-z] or (...). You can put a quantifier right after it.
  • seq: safe to concatenate, but a quantifier needs (?:...) around it. abc is a seq, and so is a+, since a+? would silently turn into a lazy quantifier.
  • alt: has a | at the top level, so it needs (?:...) almost everywhere.

Plain strings go through literal, which escapes them:

const esc = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\
amp;"); const literal = (s: string): RxNode => ({ src: esc(s), kind: s.length === 1 ? "atom" : "seq", }); const toNode = (x: Item): RxNode => (typeof x === "string" ? literal(x) : x);

With that in place, seq joins nodes and only brackets alternations:

const seq = (...xs: Item[]): RxNode => {
  const nodes = xs.map(toNode).filter((n) => n.src !== "");
  if (nodes.length === 0) return { src: "", kind: "seq" };
  if (nodes.length === 1) return nodes[0];

  let src = "";
  for (const n of nodes) {
	const part = n.kind === "alt" ? `(?:${n.src})` : n.src;
	// `\1` followed by a literal `0` would read as `\10`
	if (/\\\d+$/.test(src) && /^\d/.test(part)) src += "(?:)";
	src += part;
  }
  return { src, kind: "seq" };
};

(That backreference check is one of those bugs you only find by writing tests, or when it happens to you in prod. backref(1) followed by the literal "0" gives you backreference number ten.)

Every quantifier is a seq of its arguments plus a suffix, bracketed only when the body isn't an atom:

const quantifiable = (n: RxNode) =>
  n.kind === "atom" ? n.src : `(?:${n.src})`;

const quantifier =
  (suffix: string) =>
  (...xs: Item[]): RxNode => ({
	src: quantifiable(seq(...xs)) + suffix,
	kind: "seq",
  });

const zeroOrMore = quantifier("*");
const oneOrMore = quantifier("+");
const optional = quantifier("?");

Because each quantifier calls seq on its arguments, you get the implicit sequence for free: optional("-", group(x)) becomes (?:-(x))?.

And finally, the two entry points. As in Emacs, rx returns a string. RX returns a RegExp you can use right away:

const rx = (...xs: Item[]): string => seq(...xs).src;

function RX(...xs: Item[]): RegExp {
  return new RegExp(rx(...xs));
}

RX.flags = (flags: string, ...xs: Item[]): RegExp =>
  new RegExp(rx(...xs), flags);

RX.flags exists because Emacs controls case folding through the case-fold-search variable, and JavaScript puts it on the regexp itself.

That's the whole engine! Now, let's build our vocabulary.

Character sets

In Emacs you write (any "a-z" "_"). Inside those strings, a-z is a range, and a - at either end is a plain dash. I kept the same rule:

const hexDigit = anyOf("0-9a-fA-F");
RX(start, "#", repeat(6, hexDigit), end);
// => /^#[0-9a-fA-F]{6}$/

RX(start, optional(anyOf("+-")), oneOrMore(digit), end);
// => /^[+\-]?\d+$/

The dash comes out escaped because sets can merge. If you combine anyOf("+-") with anyOf("0-9"), an unescaped - would end up in the middle and create a range from + to 0. Escaping it costs one backslash.

And merging is the reason why RxNode has a set field. It is there to hold the text that goes between [ and ], so anyOf can take other sets as arguments:

const lower = anyOf("a-z");
const upper = anyOf("A-Z");
const alpha = anyOf(lower, upper);   // [a-zA-Z]
const alnum = anyOf(alpha, "0-9");   // [a-zA-Z0-9]

not negates a set, and it knows the shorthand classes:

not(digit);                // \D
not(anyOf(space, "@"));    // [^\s@]
notChar(",");              // [^,]   (rx's not-char)

The simple email check, which most of us have written as /^[^\s@]+@[^\s@]+\.[^\s@]+$/ at some point, becomes:

const part = oneOrMore(not(anyOf(space, "@")));

RX(start, part, "@", part, ".", part, end);
// => /^[^\s@]+@[^\s@]+\.[^\s@]+$/

The rest of the Emacs character classes are there too: digit, hexDigit, space, blank, wordChar, notWordChar, alpha, alnum, lower, upper, punct, control, graphic, printing, ascii and nonascii. One difference: in Emacs they understand Unicode, and mine are ASCII only. alpha won't match é.

Two more come from rx's symbol list, and people (me, many times) mix them up:

const notNewline: RxNode = { src: ".", kind: "atom" }; // rx: nonl
const anything = set("\\s\\S");                        // rx: anything / anychar

In rx, anything really means anything, newlines included. Here is where the difference shows up:

const code = "a = 1; /* first\n   second */ b = 2;";

RX("/*", zeroOrMoreLazy(notNewline), "*/").exec(code);
// => null

RX("/*", zeroOrMoreLazy(anything), "*/").exec(code)?.[0];
// => "/* first\n   second */"

Alternatives, and the longest match

or works as you'd expect, and gets bracketed when it lands inside a sequence:

RX(start, or("cat", "dog", "bird"), end);
// => /^(?:bird|cat|dog)$/

Did you notice the order changed? I copied that behavior from Emacs. When every branch of an or is a plain string, rx hands them to regexp-opt, which builds a pattern that prefers the longest match:

(rx (or "in" "int" "interface"))
;; => "\\(?:in\\(?:t\\(?:erface\\)?\\)?\\)"

JavaScript alternation takes the first branch that matches, going left to right. So the naive regexp for a list of keywords has a 'bug':

/in|int|interface/.exec("interface Foo")?.[0];
// => "in"

RX(or("in", "int", "interface")).exec("interface Foo")?.[0];
// => "interface"

I don't build a trie like regexp-opt does. Sorting the strings by length, longest first, is enough to get the same behavior:

const or = (...xs: Item[]): RxNode => {
  if (xs.length === 0) return unmatchable;
  if (xs.length === 1) return toNode(xs[0]);

  const branches = xs.every((x) => typeof x === "string")
	? [...(xs as string[])].sort((a, b) => b.length - a.length)
	: xs;

  return { src: branches.map((x) => toNode(x).src).join("|"), kind: "alt" };
};

As in Emacs, or() with no branches returns unmatchable, which is (?!) here. It's handy when you build the branch list at runtime and it might come out empty.

Repetition, greedy and lazy

Emacs has (= n ...), (>= n ...) and (** n m ...). Here they are repeat, atLeast and between:

RX(start, between(2, 4, digit), end);  // /^\d{2,4}$/
RX(atLeast(3, digit));                 // /\d{3,}/

The lazy versions *?, +? and ?? are zeroOrMoreLazy, oneOrMoreLazy and optionalLazy. The classic HTML tag example:

const html = "bold and italic";

RX("<", oneOrMore(notNewline), ">").exec(html)?.[0];
// => "bold and italic"

RX("<", oneOrMoreLazy(notNewline), ">").exec(html)?.[0];
// => ""

Groups and backreferences

group is a capturing group, and backref points back to it:

RX(start, group(oneOrMore(wordChar)), space, backref(1), end);
// => /^(\w+)\s\1$/   matches "hello hello", not "hello world"

Emacs also has (group-n N ...) to pick the group number. JavaScript can't do that, but it has named groups, which serve the same purpose and read better:

const date = RX(
  named("y", repeat(4, digit)), "-",
  named("m", repeat(2, digit)), "-",
  named("d", repeat(2, digit)),
);

date.exec("2026-09-30")?.groups;
// => { y: '2026', m: '09', d: '30' }

backref accepts a name as well:

RX(
  "<", named("tag", oneOrMore(wordChar)), ">",
  zeroOrMoreLazy(notNewline),
  "",
);
// => /<(?\w+)>.*?<\/\k>/

Anchors

rx distinguishes the start of the string (bos) from the start of a line (bol). In JavaScript both are ^, and the m flag decides which one you get. I kept both names so the intent shows in the code:

const text = "TODO: write post\nDONE: fix rx\nTODO: publish";
const todo = RX.flags(
  "gm",
  lineStart, "TODO: ", named("task", oneOrMore(notNewline)), lineEnd,
);

[...text.matchAll(todo)].map((m) => m.groups?.task);
// => [ 'write post', 'publish' ]

Why not add start and end automatically? Because you only want them when validating a whole string. When searching inside a text, as in split, replace or matchAll, a hidden ^ and $ would break everything. Emacs agrees: bos and eos are explicit in rx too.

wordBoundary and notWordBoundary map straight to \b and \B. Emacs also has bow and eow (\< and \>), start and end of a word. JavaScript lacks those, so I combined \b with a lookaround:

const wordStart = zeroWidth("\\b(?=\\w)");
const wordEnd = zeroWidth("\\b(?<=\\w)");

Literal and raw

Plain strings are already literals, but rx has an explicit literal form for strings computed at runtime, and I kept it. It documents that the value came from somewhere else:

const userInput = "1+1=2? (maybe)";

new RegExp(userInput).test(userInput);  // false, oops
RX(literal(userInput)).test(userInput); // true

The opposite direction is rx's (regexp ...) form, the escape hatch. Here it's raw, and it receives either a string or an existing RegExp. It lets you adopt the DSL in a codebase full of old regexps without rewriting all of them, like:

const legacyZip = /\d{5}(?:-\d{4})?/;

RX(start, repeat(2, upper), " ", raw(legacyZip), end);
// => /^[A-Z]{2} (?:\d{5}(?:-\d{4})?)$/

raw can't see inside the text it gets, so it adds brackets whenever it's combined with something else. It's an extra (?:), and the regexp still works.

Shall we try SemVer again?

Back to the regexp from the intro. In Emacs you would give names to the pieces with rx-define or rx-let. In TypeScript those are just consts:

const num = or("0", seq(anyOf("1-9"), zeroOrMore(digit)));
const idChar = anyOf(alnum, "-");
const preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, "-"), zeroOrMore(idChar)));
const dotted = (x: Item) => seq(x, zeroOrMore(".", x));

const semver = RX(
  start,
  named("major", num), ".",
  named("minor", num), ".",
  named("patch", num),
  optional("-", named("pre", dotted(preId))),
  optional("+", named("build", dotted(oneOrMore(idChar)))),
  end,
);

Now you can read the spec in the code. A numeric identifier is 0, or a non-zero digit followed by any number of digits. A pre-release is a dotted list of identifiers, and so is build metadata. dotted is a plain function returning a node, which is as far as abstraction needs to go here.

It matches the same strings as the official regexp, and the named groups give you a result like:

semver.exec("1.0.0-alpha.1+build.5")?.groups;
// => {
//   major: '1',
//   minor: '0',
//   patch: '0',
//   pre: 'alpha.1',
//   build: 'build.5'
// }

Next time the spec changes, you can understand what the current regex does at a glance, instead of fighting an army of punctuation.

What's missing from Emacs rx

I tried to map every rx form, and a few have no JavaScript equivalent:

  • point: JavaScript regexps don't know about a cursor.
  • symbol-start, symbol-end, syntax, category: these depend on Emacs syntax tables.
  • intersection: possible with the v flag, but I haven't needed it.
  • minimal-match / maximal-match: these flip the greediness of everything inside them. Doable, but it would need a separate pass, and the *Lazy functions cover my use cases.
  • eval: TypeScript already evaluates expressions everywhere, so you get it for free.

And one addition Emacs doesn't need: RX.flags.

Cheat sheet

Emacs rx TypeScript JS regexp (roughly)
seq, :, and seq(...), implicit in every form ab
or, | or(...) a|b
any, in, char anyOf("a-z", "_", digit) [a-z_\d]
not-char notChar(...) [^...]
not not(charset) \D, [^...]
*, +, ? zeroOrMore, oneOrMore, optional x*, x+, x?
*?, +?, ?? zeroOrMoreLazy, oneOrMoreLazy, optionalLazy x*?, x+?, x??
=, >=, ** repeat, atLeast, between x{n}, x{n,}, x{n,m}
group group(...) (...)
group-n named("name", ...) (?...)
backref backref(1), backref("name") \1, \k
literal literal(s) s, escaped: 1\+1
regexp, regex raw("..."), raw(/.../) (?:...), as-is
rx-define, rx-let const (none)
bos, eos start, end ^, $
bol, eol lineStart, lineEnd (with the m flag) ^, $
bow, eow wordStart, wordEnd \b(?=\w), \b(?<=\w)
word-boundary wordBoundary \b
not-word-boundary notWordBoundary \B
nonl, not-newline notNewline .
anychar, anything anything [\s\S]
unmatchable unmatchable (?!)
digit digit \d
hex-digit, xdigit hexDigit [0-9a-fA-F]
space, whitespace space \s
blank blank [ \t]
word, wordchar wordChar \w
not-wordchar notWordChar \W
alpha, letter alpha [a-zA-Z]
alnum alnum [a-zA-Z0-9]
lower, upper lower, upper [a-z], [A-Z]
punct, punctuation punct [!-/:-@[-`{-~]
cntrl, control control [\x00-\x1f\x7f]
graph, graphic graphic [!-~]
print, printing printing [ -~]
ascii, nonascii ascii, nonascii [\x00-\x7f], [\u0080-\uffff]

Examples: JS/TS regex vs RX

Each example below shows the goal, the Emacs rx form in a comment, the regexp you would write by hand, and the RX version. When RX produces a different regexp text, the // => line shows it. The results at the bottom come from running both against the same strings.

Digits only

The whole string is digits.

// Emacs: (rx bos (+ digit) eos)
const regex = /^\d+$/;
const dsl = RX(start, oneOrMore(digit), end);

// "123" -> true
// "12a" -> false

Letters only

The whole string is ASCII letters.

// Emacs: (rx bos (+ alpha) eos)
const regex = /^[a-zA-Z]+$/;
const dsl = RX(start, oneOrMore(alpha), end);

// "Hello" -> true
// "He11o" -> false

Optional letter

Both spellings, color and colour.

// Emacs: (rx bos "colo" (? "u") "r" eos)
const regex = /^colou?r$/;
const dsl = RX(start, "colo", optional("u"), "r", end);

// "color" -> true
// "colour" -> true
// "colouur" -> false

Two words

Two words separated by a space.

// Emacs: (rx bos (+ wordchar) space (+ wordchar) eos)
const regex = /^\w+\s\w+$/;
const word = oneOrMore(wordChar);
const dsl = RX(start, word, space, word, end);

// "hello world" -> true
// "hello" -> false

Phone number

(123) 456-7890, parentheses and all.

// Emacs: (rx bos "(" (= 3 digit) ")" space (= 3 digit) "-" (= 4 digit) eos)
const regex = /^\(\d{3}\)\s\d{3}-\d{4}$/;
const dsl = RX(
  start, "(", repeat(3, digit), ")", space,
  repeat(3, digit), "-", repeat(4, digit), end,
);

// "(123) 456-7890" -> true
// "123-456-7890" -> false

Hex color

#ff00aa-style colors.

// Emacs: (rx bos "#" (= 6 hex-digit) eos)
const regex = /^#[0-9a-fA-F]{6}$/;
const dsl = RX(start, "#", repeat(6, hexDigit), end);

// "#ff00aa" -> true
// "#ff00ag" -> false

Signed integer

An optional sign, then digits.

// Emacs: (rx bos (? (any "+-")) (+ digit) eos)
const regex = /^[+-]?\d+$/;
const dsl = RX(start, optional(anyOf("+-")), oneOrMore(digit), end);
// => /^[+\-]?\d+$/

// "-42" -> true
// "42" -> true
// "*42" -> false

Simple email

Something@something.something, no spaces.

// Emacs: (rx-let ((part (+ (not (any space "@")))))
//   (rx bos part "@" part "." part eos))
const regex = /^[^\s@]+@[^\s@]+\.[^\s@]+$/;
const part = oneOrMore(not(anyOf(space, "@")));
const dsl = RX(start, part, "@", part, ".", part, end);

// "a@b.com" -> true
// "a b@c.com" -> false
// "a@b" -> false

No digits

A string without any digit.

// Emacs: (rx bos (+ (not digit)) eos)
const regex = /^[^\d]+$/;
const dsl = RX(start, oneOrMore(not(digit)), end);
// => /^\D+$/

// "abc" -> true
// "a1b" -> false

CSV line

Exactly three comma-separated fields.

// Emacs: (rx-let ((field (+ (not-char ","))))
//   (rx bos field "," field "," field eos))
const regex = /^[^,]+,[^,]+,[^,]+$/;
const field = oneOrMore(notChar(","));
const dsl = RX(start, field, ",", field, ",", field, end);

// "a,b,c" -> true
// "a,b," -> false

One of many

A fixed list of words.

// Emacs: (rx bos (or "cat" "dog" "bird") eos)
const regex = /^(?:cat|dog|bird)$/;
const dsl = RX(start, or("cat", "dog", "bird"), end);
// => /^(?:bird|cat|dog)$/

// "cat" -> true
// "bird" -> true
// "cow" -> false

Title and name

mr or ms, then a name, keeping the title.

// Emacs: (rx bos (group (or "mr" "ms")) space (+ wordchar) eos)
const regex = /^(mr|ms)\s\w+$/;
const dsl = RX(start, group(or("mr", "ms")), space, oneOrMore(wordChar), end);

// "mr john" -> true
// "dr john" -> false

Between

Two to four digits.

// Emacs: (rx bos (** 2 4 digit) eos)
const regex = /^\d{2,4}$/;
const dsl = RX(start, between(2, 4, digit), end);

// "12" -> true
// "12345" -> false

At least

Three or more digits, anywhere.

// Emacs: (rx (>= 3 digit))
const regex = /\d{3,}/;
const dsl = RX(atLeast(3, digit));

// "12" -> false
// "a123" -> true

Repeated word

The same word twice.

// Emacs: (rx bos (group (+ wordchar)) space (backref 1) eos)
const regex = /^(\w+)\s\1$/;
const dsl = RX(start, group(oneOrMore(wordChar)), space, backref(1), end);

// "hello hello" -> true
// "hello world" -> false

Matching tags

An open tag and its own closing tag.

// Emacs: (rx "<" (group-n 1 (+ wordchar)) ">" (*? nonl) "")
const regex = /<(?\w+)>.*?<\/\k>/;
const dsl = RX(
  "<", named("tag", oneOrMore(wordChar)), ">",
  zeroOrMoreLazy(notNewline),
  "",
);

// "bold" -> true
// "oops" -> false

Whole word

cat as a word, not inside another one.

// Emacs: (rx word-boundary "cat" word-boundary)
const regex = /\bcat\b/;
const dsl = RX(wordBoundary, "cat", wordBoundary);

// "the cat sat" -> true
// "concatenate" -> false

Case-insensitive

hello, in any case.

// Emacs: (let ((case-fold-search t))
//   (string-match-p (rx bos "hello" eos) "HeLLo"))
const regex = /^hello$/i;
const dsl = RX.flags("i", start, "hello", end);

// "HeLLo" -> true
// "help" -> false

Full source

It's a single file with no dependencies. Copy it into your project and start deleting the forms you don't need, or adding the ones you miss.

You can check the same code, plus all the examples from this post (and a few more), in this gist. If you'd rather not set anything up, paste it into the TypeScript Playground, hit "Run", and check the "Logs" tab.

/* =========================================================
 * CORE
 * ========================================================= */

// How a node behaves when combined with others:
//   atom -> single unit, a quantifier can be glued right after it
//   seq  -> safe to concatenate, needs (?:) to be quantified
//   alt  -> has a top-level `|`, needs (?:) almost everywhere
type Kind = "atom" | "seq" | "alt";

interface RxNode {
  readonly src: string;
  readonly kind: Kind;
  // char sets only: the text that goes inside [ ], so sets can merge
  readonly set?: string;
  readonly neg?: boolean;
}

type Item = string | RxNode;

const esc = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\
amp;"); const escSet = (s: string) => s.replace(/[\]\\^-]/g, "\\
amp;"); // rx: (literal EXPR) — a string computed at runtime, matched as-is const literal = (s: string): RxNode => ({ src: esc(s), kind: s.length === 1 ? "atom" : "seq", }); const toNode = (x: Item): RxNode => (typeof x === "string" ? literal(x) : x); // wrap a node so a quantifier applies to all of it const quantifiable = (n: RxNode) => n.kind === "atom" ? n.src : `(?:${n.src})`; // rx: (regexp EXPR) — escape hatch, trust the regexp as-is const raw = (re: string | RegExp): RxNode => ({ src: typeof re === "string" ? re : re.source, kind: "alt", }); /* ========================================================= * COMPOSITION * ========================================================= */ const seq = (...xs: Item[]): RxNode => { const nodes = xs.map(toNode).filter((n) => n.src !== ""); if (nodes.length === 0) return { src: "", kind: "seq" }; if (nodes.length === 1) return nodes[0]; let src = ""; for (const n of nodes) { const part = n.kind === "alt" ? `(?:${n.src})` : n.src; // `\1` followed by a literal `0` would read as `\10` if (/\\\d+$/.test(src) && /^\d/.test(part)) src += "(?:)"; src += part; } return { src, kind: "seq" }; }; const unmatchable: RxNode = { src: "(?!)", kind: "atom" }; // Like rx: when every branch is a plain string, try the longest first, // so or("in", "int") matches "int" instead of stopping at "in". const or = (...xs: Item[]): RxNode => { if (xs.length === 0) return unmatchable; if (xs.length === 1) return toNode(xs[0]); const branches = xs.every((x) => typeof x === "string") ? [...(xs as string[])].sort((a, b) => b.length - a.length) : xs; return { src: branches.map((x) => toNode(x).src).join("|"), kind: "alt" }; }; /* ========================================================= * CHARACTER SETS * ========================================================= */ const set = (body: string, neg = false): RxNode => ({ src: neg ? `[^${body}]` : `[${body}]`, kind: "atom", set: body, neg, }); // class escapes are sets too, so they can go inside anyOf(...) const classEscape = (e: string): RxNode => ({ src: e, kind: "atom", set: e, neg: false, }); // Same reading as rx: inside a string, "a-z" is a range, while a `-` // at the start or end is just a dash ("+-" is plus or minus). const intervals = (s: string): string => { let body = ""; let i = 0; while (i < s.length) { if (i < s.length - 2 && s[i + 1] === "-") { body += `${escSet(s[i])}-${escSet(s[i + 2])}`; i += 3; } else { body += escSet(s[i]); i += 1; } } return body; }; // rx: (any "a-z" "_" digit) — also known as `in` and `char` const anyOf = (...xs: Item[]): RxNode => { const body = xs .map((x) => { if (typeof x === "string") return intervals(x); if (x.set === undefined || x.neg) throw new Error(`anyOf: not a positive char set: ${x.src}`); return x.set; }) .join(""); return set(body); }; // rx: (not charset) — not(digit) -> \D, not(anyOf(",;")) -> [^,;] const not = (x: Item): RxNode => { const n = typeof x === "string" ? anyOf(x) : x; if (n.set === undefined) throw new Error(`not: not a char set: ${n.src}`); if (n.neg) return set(n.set); if (/^\\[dswDSW]$/.test(n.src)) { const c = n.src[1]; const flipped = c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase(); return classEscape(`\\${flipped}`); } return set(n.set, true); }; // rx: (not-char "a-z" ...) — shorthand for (not (any ...)) const notChar = (...xs: Item[]) => not(anyOf(...xs)); // rx char classes, `[[:name:]]` in Emacs const digit = classEscape("\\d"); const space = classEscape("\\s"); const wordChar = classEscape("\\w"); const notWordChar = not(wordChar); const lower = anyOf("a-z"); const upper = anyOf("A-Z"); const alpha = anyOf(lower, upper); const alnum = anyOf(alpha, "0-9"); const hexDigit = anyOf("0-9a-fA-F"); const blank = set(" \\t"); const control = set("\\x00-\\x1f\\x7f"); const punct = anyOf("!-/:-@[-`{-~"); const graphic = anyOf("!-~"); const printing = anyOf(" -~"); const ascii = set("\\x00-\\x7f"); const nonascii = set("\\u0080-\\uffff"); // rx: `nonl` is any char but newline; `anything` really is anything const notNewline: RxNode = { src: ".", kind: "atom" }; const anything = set("\\s\\S"); /* ========================================================= * ANCHORS (zero-width) * ========================================================= */ const zeroWidth = (src: string): RxNode => ({ src, kind: "seq" }); // rx: bos / eos const start = zeroWidth("^"); const end = zeroWidth("$"); // rx: bol / eol — same symbols, only per line with the "m" flag const lineStart = start; const lineEnd = end; const wordBoundary = zeroWidth("\\b"); const notWordBoundary = zeroWidth("\\B"); // rx: bow / eow — JS has no \< \>, so a boundary plus a lookaround const wordStart = zeroWidth("\\b(?=\\w)"); const wordEnd = zeroWidth("\\b(?<=\\w)"); /* ========================================================= * GROUPS & BACKREFERENCES * ========================================================= */ const group = (...xs: Item[]): RxNode => ({ src: `(${seq(...xs).src})`, kind: "atom", }); // rx has (group-n N ...); JS can't pick group numbers, but it can name them const named = (name: string, ...xs: Item[]): RxNode => ({ src: `(?<${name}>${seq(...xs).src})`, kind: "atom", }); const backref = (ref: number | string): RxNode => ({ src: typeof ref === "number" ? `\\${ref}` : `\\k<${ref}>`, kind: "atom", }); /* ========================================================= * QUANTIFIERS * ========================================================= */ const quantifier = (suffix: string) => (...xs: Item[]): RxNode => ({ src: quantifiable(seq(...xs)) + suffix, kind: "seq", }); // greedy — rx: * + ? const zeroOrMore = quantifier("*"); const oneOrMore = quantifier("+"); const optional = quantifier("?"); // lazy — rx: *? +? ?? const zeroOrMoreLazy = quantifier("*?"); const oneOrMoreLazy = quantifier("+?"); const optionalLazy = quantifier("??"); // rx: (= n ...) (>= n ...) (** n m ...) const repeat = (n: number, ...xs: Item[]) => quantifier(`{${n}}`)(...xs); const atLeast = (n: number, ...xs: Item[]) => quantifier(`{${n},}`)(...xs); const between = (n: number, m: number, ...xs: Item[]) => quantifier(`{${n},${m}}`)(...xs); /* ========================================================= * ENTRY POINTS * ========================================================= */ // rx(...) -> the regexp source string (like Emacs, rx returns a string) // RX(...) -> a ready-to-use RegExp, no more `new RegExp(seq(...))` const rx = (...xs: Item[]): string => seq(...xs).src; function RX(...xs: Item[]): RegExp { return new RegExp(rx(...xs)); } // Emacs uses `case-fold-search` for this; JS puts it on the regexp RX.flags = (flags: string, ...xs: Item[]): RegExp => new RegExp(rx(...xs), flags);

Wrapping up

None of this is new. On the Emacs side, as I said before, rx has shipped for decades, and the Elisp version is more complete than mine.

The idea of describing patterns with a small DSL instead of raw syntax isn't new either. Plenty of people have tried it, each in their own way. One project I like a lot in this space is Zod, which I wrote about in my Zod quick tutorial. It's not a regexp builder: you compose small schema pieces, and Zod gives you back a parser and a TypeScript type from the same construction. It follows the same spirit, though: build big things out of small named pieces you can read.

If you write Elisp and have never tried rx, open *scratch*, type (rx (+ digit)), and C-x C-e it. If you write JavaScript or TypeScript, the file above is yours. And if you port it to another language, send me a link.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论