A slex Tutorial

This tutorial introduces lex-style scanning and the .slex spec files you will use to build the IC lexer in PA 1. A lexer is a spec file: an ordered list of regular-expression rules, each optionally paired with a block of Scala to run when it matches. The slex library reads the spec and runs it directly: there is no generated code to look at, no extra build step to run, and the spec file is the lexer.

Every example below is a real file in the tutorial’s companion repository: a spec file (toy.slex, mini.slex, …) plus a small program that loads and runs it. Clone the repository (which also contains the library source – read it if you’d like!) and run and modify the examples as you read:

git clone https://github.com/williams-cs/cs434-f26-slex-tutorial
cd cs434-f26-slex-tutorial
sbt "runMain toy.demo"

Each demo lexes the tutorial’s sample input, and anything you pass is lexed instead: sbt "runMain toy.demo if (2 <= 3)". The tutorial’s claims about the examples are also an MUnit suite – sbt test runs them all.

Note

Editor support. If you use VSCode, install the course’s “slex & scup” extension before reading on: download the .vsix and run code --install-extension slex-scup.vsix (or, in VSCode, Extensions → “···” → Install from VSIX…). What it does for .slex files is described in Editor Support below.

The Idea: Tokens as Regular Expressions

Lex (1975) and its successor flex are scanner generators: you write a specification (an ordered list of regular expressions, each paired with an action to run when it matches) and the tool writes the scanner. The engine decides where each token begins and ends; your actions decide what to do about it.

The division of labor works because the token boundaries of a programming language are regular, even though much of what a scanner must compute (keywords vs. identifiers, escape sequences, range checks) is not.

A slex spec reads like lex. A rule is a pattern, optionally followed by a { ... } block of code to run when the pattern matches:

# two rules, from some .slex file
[0-9]+       { Token(NUM, text, pos) }      # produce a token
[ \t\r\n]+                                  # no block: match and discard

Inside a block, two names are always available: text is the matched characters (the lexeme), and pos is where they started. The block either builds a Token or calls error(...). A rule with no block is a skip rule: whatever it matches is thrown away, which is all whitespace and comments need. In the spec file itself, # starts a comment and blank lines are ignored.

A complete lexer is three small files. First, the token kinds: a Scala enum, which is nothing more than a list of names, one per kind of token. You write it by hand, in an ordinary source file:

// K.scala -- the token kinds of a toy while-language
package toy

enum K:
  case WHILE, IF, LEQ, LT, LP, RP, NUM

Second, the spec file, which describes the token language. Here is all of it for the toy language:

# toy.slex -- the token language of the toy while-language

%kinds toy.K

%{
  // Helper functions for the blocks below.  This is ordinary
  // Scala; anything defined here can be used in any block.
  def toNumber(s: String): Int = s.toInt
%}

%token
  WHILE "while"   IF    "if"      LEQ   "<="
  LT    "<"       LP    "("       RP    ")"

# whitespace: no block, so matches are discarded
[ \t\r\n]+

# a number: the block runs on the matched digits and builds the
# token, storing the parsed Int in the token's value field
[0-9]+      { Token(NUM, text, pos, toNumber(text)) }

Reading it top to bottom:

Third, the program that uses the lexer. It loads the spec by naming the enum and gets back a working lexer:

// Toy.scala -- load the spec and lex something
package toy

import slex.{LexError, LexSpec}

@main def demo(input: String*) =
  val toy = LexSpec.load[K]("toy.slex")
  val line = if input.isEmpty then "while (2 <= 10)" else input.mkString(" ")
  try
    for t <- toy.tokenize(line) do
      println(t)   // [WHILE,while,1], [LP,(,1], [NUM,2,1], ...
  catch case e: LexError => println(s"lex error at ${e.pos}: ${e.message}")

The tokens are scup.Token[K] values: a kind, the lexeme text, the position, and a value field where a block can stash a computed result, like the parsed Int above (or, in IC, an unescaped string). More on that below.

The Two Laws

At every step, a scanner must answer: which rule matches next, and how much input does it take? Several rules may match, at several different lengths. Lex settled the question with two tie-breaking laws, and slex follows them. Most of the characteristic behavior of a lexer traces back to these two laws.

(The little specs in this section show just the rules each law needs. In the companion repository they are rel.slex and kw.slex, each with its %kinds line and a tiny driver: sbt "runMain laws.rel" and sbt "runMain laws.kw" show the laws on real input.)

Law 1: maximal munch. The rule that matches the longest lexeme wins, regardless of where it appears in the list.

# rel.slex -- two relational operators
%token
  LT  "<"
  LEQ "<="

On the input <=<, the LT row comes first and matches at the start, but LEQ matches two characters there, and length beats listing order: the tokens are LEQ then LT. This is why x<=y can never lex as x, <, =, y, the same rule you applied by hand in HW 2.

Law 2: the first rule wins ties. If several rules match the same longest lexeme, the one listed first wins.

# kw.slex -- one keyword and an identifier rule
%token
  WHILE "while"

[a-z]+       { Token(ID, text, pos) }

On the input while, both the keyword row and the identifier rule match all five characters. The tie goes to the keyword because its row is listed first. To handle keywords in a slex spec, put the %token table above the identifier rule; no keyword map is needed, because first-rule-wins picks the keyword row on an exact match. But note what happens on whilewhilst: the identifier rule matches more characters, and Law 1 outranks Law 2, so it lexes as one ID. A keyword row only ever claims a lexeme that is exactly the keyword.

Order the rules carelessly and Law 2 bites: list the identifier rule above the %token table and it wins every tie, so while becomes an ordinary identifier and the keyword rows are dead code.

(The other classic design puts the keyword lookup in the identifier rule’s action, one rule consulting a Map[String, K], and your block is ordinary Scala, so nothing stops you. With the %token table doing the same job declaratively, IC’s spec never needs to.)

One footnote to Law 1: a zero-length match never wins. A rule like a* matches the empty string at every position; if that counted, the scanner would emit empty tokens forever without advancing. slex ignores matches of length zero, so a rule fires only when it matches at least one character.

Watching It Scan

A scan can print a trace of its decisions. Pass trace = true to tokenize (or scan) and the lexer prints, one munch at a time, exactly what the engine is doing: which rules are still alive after each character, every point where some rule accepts (and who wins the tie), and the final munch:

spec.tokenize("whilewhilst", trace = true)

(A String => Unit sink also works there, so trace = println means the same thing, and scup’s parse takes the same argument; you will meet it again in the scup tutorial.) The laws and mini demos take a -trace flag:

sbt "runMain laws.kw -trace whilewhilst"

The trace opens with the numbered rule list and the machine running it (the DFA those rules compile to; Inspecting the Machine shows it), then reports each munch. Here is the trace for whilewhilst, with both laws visible in a single lexeme (middle characters elided):

Rules (listing order breaks ties):
  #1   "while"
  #2   [a-z]+
Engine: DFA (8 states)

scan at 1:1: "whilewhilst"
  'w'   alive: #1 "while", #2 [a-z]+
        accept: #2 [a-z]+  (best is now 1 char)
  ...
  'e'   alive: #1 "while", #2 [a-z]+
        accept: #1 "while", #2 [a-z]+ -- first listed wins: #1 "while"  (best is now 5 chars)
  'w'   alive: #2 [a-z]+
        accept: #2 [a-z]+  (best is now 6 chars)
  ...
  't'   alive: #2 [a-z]+
        accept: #2 [a-z]+  (best is now 11 chars)
  (end of input)
  longest match: #2 [a-z]+ takes 11 chars
  => [ID,whilewhilst,1]

Both laws appear in it. At five characters both rules accept, and the tie goes to the row listed first; the first listed wins: #1 line is Law 2 claiming the keyword. But the identifier rule is still alive, keeps accepting, and the longest accept takes the munch. The last two lines are Law 1 outranking Law 2, lexing whilewhilst as one ID.

One more thing to watch for: when the machine consumes characters past its last accept while hunting for a longer match and then dies, the longest match line reports the retreat, (backing up 2), and scanning resumes after the lexeme actually taken. That give-back is the cost of maximal munch; an HW 2 problem steps through an example.

The trace is for watching the algorithm now and for debugging your IC lexer later. When a lexeme comes out with the wrong kind or the wrong length (a keyword lexed as an identifier, <= split in two, a comment rule eating half the file), trace = true on one small failing input shows which rules were alive, who accepted where, and why the winner won.

Inspecting the Machine

The trace shows the machine running; the check CLI’s -dump flag shows the machines themselves: the NFA the rules compile to, and the DFA built from it. For rel.slex (the two relational operators from Law 1), the whole thing fits on a screen:

$ sbt "runMain slex.check -dump rel.slex"
rel.slex: OK (2 rules, 0 actions)
rel.slex: NFA 11 states, DFA 3 states

Rules (listing order breaks ties):
  #1   "<"   -> LT        (line 6)
  #2   "<="  -> LEQ       (line 6)

NFA (11 states; state 0 is the start, with an ε-edge to each rule's fragment):
     0 --ε--> 1, 5
     1 --ε--> 2
     2 --ε--> 3
     3 --'<'--> 4
     4 accepts #1 "<"
     5 --ε--> 6
     6 --ε--> 7
     7 --'<'--> 8
     8 --ε--> 9
     9 --'='--> 10
    10 accepts #2 "<="

DFA (3 states; state 0 is the start; each state lists its NFA states):
     0 = {0, 1, 2, 3, 5, 6, 7}
       --'<'--> 1
     1 = {4, 8, 9}  accepts #1 "<"
       --'='--> 2
     2 = {10}  accepts #2 "<="

The rule table at the top summarizes the spec: each rule’s pattern, the token it produces, and its line, in priority order, because the listing order is the tie-breaker. Below it is the combined NFA those rules compile to: state 0 fans out by ε-edges to every rule’s fragment, states 1–4 are the fragment for <, states 5–10 the fragment for <= (follow the chain: ε-glue, match <, ε-glue, match =), and each fragment ends in an accept state tagged with its rule’s number, the tag that makes first-rule-wins a one-line min. The machine has more ε-edges than the one you would draw by hand, because Thompson’s construction glues fragments together instead of merging states; simulate both and they accept the same strings.

Below the NFA is what the subset construction from HW 1 makes of it: a DFA whose every state is a set of NFA states (the braces), with one edge per character instead of ε-glue. For rel.slex it is the three-state machine you would have drawn by hand – state 0 has read nothing, state 1 has read < and accepts rule #1, state 2 has read <= and accepts rule #2 – and that table, not the NFA, is what a scan walks: one array lookup per character. Where two rules accept in the same DFA state, the dump says so and names the winner, which is Law 2 precomputed.

-dump works on any spec. Point it at mini.slex (103 NFA states, 30 DFA states) or, in PA 1, at your ic.slex when you want to see what a real token language compiles to, or when you are debugging: a rule whose accept state you cannot reach from its fragment’s start is a rule that can never fire.

Writing Patterns

A rule’s pattern is a lex-style regular expression, ended by the first whitespace outside a character class (write \ for a literal space):

These are just the five inductive cases of the regular-expression definition from class (empty string, single symbol, concatenation, alternation, star); everything else (+, ?, classes, literal strings) is sugar for those five. A malformed pattern (say, [a-z) is reported, with its line and what was expected, the moment the spec loads. If you use the course VSCode extension, it is also a squiggle under the offending character as you type.

A Complete Example: A Miniature IC Lexer

Here is a small cut of the IC lexer, showing all the moving parts at once:

# mini.slex -- a small cut of the IC lexer

%kinds mini.K

%token
  WHILE  "while"   IF     "if"      ELSE   "else"
  LEQ    "<="      EQ     "=="      ASSIGN "="
  LT     "<"       LP     "("       RP     ")"
  LB     "{"       RB     "}"       SEMI   ";"

[ \t\r\n]+
//[^\n]*

[a-zA-Z][a-zA-Z0-9_]*  { Token(ID, text, pos) }
[0-9]+                 { Token(NUM, text, pos) }
// Mini.scala -- lexing one line with the mini lexer
val mini = LexSpec.load[K]("mini.slex")
mini.tokenize("while (i <= 10) { i = i1; } // done")
// WHILE:"while", LP:"(", ID:"i", LEQ:"<=", NUM:"10", RP:")",
// LB:"{", ID:"i", ASSIGN:"=", ID:"i1", SEMI:";", RB:"}"

Things to notice:

Actions: The Escape Hatch

The patterns decide where tokens begin and end; the blocks decide everything else. A block is arbitrary Scala, and that is where all the non-regular work of a real scanner lives.

Transforming the lexeme. A string literal’s token should carry the string’s contents, not its source spelling. The block is where the quotes come off:

# one rule from strings.slex
"[^"\n]*"  { Token(STR, text.substring(1, text.length - 1), pos) }

Your IC lexer’s version will go further and translate escape sequences (\n, \t, \", \\) into the characters they denote, a few lines of ordinary string processing (a helper in the %{ ... %} prelude keeps the rule readable) that no regular expression could do for you.

Checking and erroring. Blocks report lexical errors with the in-scope error(msg) helper, which is how a scanner rejects lexemes that are regular in shape but illegal in value. IC’s integer literals must fit in 32 bits, and [0-9]+ has no opinion about that:

# one rule from strings.slex
[0-9]+     {
    val n = text.toIntOption.getOrElse(error("Integer out of range."))
    Token(NUM, text, pos, n)
}

(toIntOption, not toInt, because a 40-digit literal should produce your error, not a NumberFormatException.) Note the fourth argument: the parsed Int rides along in the token’s value field, so later phases never re-parse the text.

The catch-all idiom. What about a string literal that never closes? No string rule matches it, so without help the scanner would report only illegal character '"'. The fix is a second, lower-priority rule that matches just the opening delimiter and whose block is the error:

# two rules from strings.slex: the real rule, then the catch-all
"[^"\n]*"  { Token(STR, text.substring(1, text.length - 1), pos) }
"          { error("Beginning of unterminated string literal.") }

This works because whenever a properly terminated string starts here, the first rule matches a longer lexeme (even the empty string "" is two characters), so maximal munch picks it and the catch-all never fires. The lone-quote rule can only win when no terminated string begins at this position (the error case), so the message can describe the real problem and its location. The same idiom handles unterminated /* ... */ comments in PA 1: a real comment rule first, then a rule for a bare /* whose block errors; a final .|\n rule turns any other stray character into a precise message.

The Enum Contract

Token kinds are a real Scala 3 enum in your project, not strings: your parser’s actions and your compiler’s later phases pattern-match on them with full type checking. The spec binds to the enum at load time: every name the spec uses must be an enum member. A typo in a %token row is caught the moment the spec loads, with its line and a did-you-mean; a typo inside a block is an ordinary compile error (the enum’s members are in scope there).

Enum members with no %token row are simply the ones whose tokens a block builds (ID, INTLIT). The check CLI lists them, so a glance shows exactly which kinds your blocks are responsible for:

$ sbt "runMain slex.check ic.slex"
ic.slex: OK (52 rules, 8 actions)
ic.slex: kinds produced by action blocks (no %token row):
  ID, CLS_ID, INTLIT, STRINGLIT, BOOLLIT

A free consequence of the %token table: the parser’s error messages print fixed-text kinds as their source text (expected '{' or 'extends') with no display-name map anywhere.

Errors Reach the Host

A block’s error(msg) throws slex.LexError(msg, pos); input no rule matches at all throws the same. Your compiler’s lexer shim (a provided, ~40-line class) converts it to IC’s own LexicalError. Since blocks are plain Scala, a block that needs full control can also throw LexicalError directly, adjusting the column arithmetic as it pleases.

Editor Support

The course VSCode extension (“slex & scup”, installed in the note at the top of this tutorial; course projects also recommend it automatically when opened) understands .slex files: Scala highlighting inside blocks and preludes, Format Document for the column layout, outline, squiggles under malformed patterns as you type, and, with the check CLI wired in, every load-time validation as a squiggle on save. Navigation understands the spec: go-to-definition on a token kind steps from its %token row to the action that constructs it to the enum member itself, and %kinds jumps to the enum declaration. Inside blocks and the prelude, go-to-definition and hover behave like ordinary Scala once the project has compiled (the build’s generated actions object powers them; until the next compile after an edit, navigation quietly falls back to the spec-level answers).

Under the Hood

The engine does no code generation: the spec is interpreted by the same three algorithms you build or run by hand this semester, plus a small driver. The library’s four core files (Re, Nfa, Dfa, Lexer) total about 650 lines; read them end to end.

  1. Re.scala — the regular-expression AST: the five inductive cases, with combinators and character classes as sugar. Each pattern in your spec is parsed into one of these (by ReSyntax.scala, itself written with a scup grammar over character tokens).

  2. Nfa.scala — Nfa.compile runs the Thompson construction (the same construction you implement in the HW), turning each rule’s Re into an NFA fragment with one start and one accept state. The fragments are joined into a single combined machine: a fresh start state with an ε-edge to every rule’s fragment, and each fragment’s accept state tagged with its rule’s index. One machine runs all the rules at once.

    Nfa.longestMatch is the NFA simulation you built in Lab 1: keep a set of current states, and repeatedly take the ε-closure and step on the next character. Every time the current set contains an accepting state, record (length so far, lowest rule tag present) as the best match so far. When the machine dies or the input ends, the remembered best implements maximal munch, and taking the lowest tag implements first-rule-wins; each is about one line of code. It is no longer what a scan runs by default, but it is one flag away: tokenize(src, engine = Engine.Nfa).

  3. Dfa.scala — Dfa.compile runs the subset construction from HW 1 over the ASCII alphabet: each DFA state is a set of NFA states, the start state is the ε-closure of the NFA’s start, and the transition from a state on a character is the ε-closure of everything its members reach on that character. Per state it precomputes what the scan needs: the lowest rule tag among the accepting members (first-rule-wins, folded into the table) and, for traces, every accepting tag and which rules are still alive. Dfa.longestMatch is then a table walk – one array lookup per character – that remembers the last accepting length exactly as the simulation does. Same answer, about a hundred times faster on IC programs. (The trace argument of Watching It Scan is either engine narrating through one shared voice; Inspecting the Machine prints both machines.)

  4. Lexer.scala — the driver: repeatedly call the chosen engine’s longestMatch, run the winning rule’s action (or discard, for a skip rule), advance the line/column counters over the lexeme, and loop. Input is ASCII: the first character outside it ends the scan with a LexError at its position, whichever engine runs.

Where You Will Use It

In PA 1 you will write ic.slex, the complete IC token language: a %token table of every keyword and operator, rules for whitespace, both comment forms, identifiers, and literals, plus blocks for string-escape processing, integer range checks, and unterminated-string/comment errors. The two laws and the catch-all idiom above cover most of that assignment; the finished spec plus the ~25-line TokenKind enum is the entire lexer.

One note on realism. Production compilers (clang, rustc, javac) hand-write their lexers rather than generate them, chiefly for the quality of their error messages (and, less than folklore suggests, for speed). Writing PA 1’s blocks will show you why: the hard parts of a real scanner are the ones regular expressions cannot express.