
This tutorial introduces lex-style scanning and the
.slex spec files you will use to build the IC lexer in PA
1. A lexer is a spec file: an ordered list of
regular-expression rules, each optionally paired with a block of Scala
to run when it matches. The slex library reads the spec and runs it
directly: there is no generated code to look at, no extra build step to
run, and the spec file is the lexer.
Every example below is a real file in the tutorial’s companion
repository: a spec file (toy.slex, mini.slex,
…) plus a small program that loads and runs it. Clone the repository
(which also contains the library source – read it if you’d like!) and
run and modify the examples as you read:
git clone https://github.com/williams-cs/cs434-f26-slex-tutorial
cd cs434-f26-slex-tutorial
sbt "runMain toy.demo"
Each demo lexes the tutorial’s sample input, and anything you pass is
lexed instead: sbt "runMain toy.demo if (2 <= 3)". The
tutorial’s claims about the examples are also an MUnit suite –
sbt test runs them all.
Editor support. If you use VSCode, install the
course’s “slex & scup” extension before reading on: download the .vsix and run
code --install-extension slex-scup.vsix (or, in VSCode,
Extensions → “···” → Install from VSIX…). What it does for
.slex files is described in Editor Support below.
The Idea: Tokens as Regular Expressions
Lex (1975) and its successor flex are scanner generators: you write a specification (an ordered list of regular expressions, each paired with an action to run when it matches) and the tool writes the scanner. The engine decides where each token begins and ends; your actions decide what to do about it.
The division of labor works because the token boundaries of a programming language are regular, even though much of what a scanner must compute (keywords vs. identifiers, escape sequences, range checks) is not.
A slex spec reads like lex. A rule is a pattern, optionally followed
by a { ... } block of code to run when the pattern
matches:
# two rules, from some .slex file
[0-9]+ { Token(NUM, text, pos) } # produce a token
[ \t\r\n]+ # no block: match and discard
Inside a block, two names are always available: text is
the matched characters (the lexeme), and pos is
where they started. The block either builds a Token or
calls error(...). A rule with no block is a skip
rule: whatever it matches is thrown away, which is all whitespace and
comments need. In the spec file itself, # starts a comment
and blank lines are ignored.
A complete lexer is three small files. First, the token
kinds: a Scala enum, which is nothing more than a list
of names, one per kind of token. You write it by hand, in an ordinary
source file:
// K.scala -- the token kinds of a toy while-language
package toy
enum K:
case WHILE, IF, LEQ, LT, LP, RP, NUMSecond, the spec file, which describes the token language. Here is all of it for the toy language:
# toy.slex -- the token language of the toy while-language
%kinds toy.K
%{
// Helper functions for the blocks below. This is ordinary
// Scala; anything defined here can be used in any block.
def toNumber(s: String): Int = s.toInt
%}
%token
WHILE "while" IF "if" LEQ "<="
LT "<" LP "(" RP ")"
# whitespace: no block, so matches are discarded
[ \t\r\n]+
# a number: the block runs on the matched digits and builds the
# token, storing the parsed Int in the token's value field
[0-9]+ { Token(NUM, text, pos, toNumber(text)) }
Reading it top to bottom:
%kindsnames the enum. Every name the spec uses must be a member of the enum; a typo is caught, with a suggestion, the moment the spec loads. The enum’s members are also automatically usable in blocks, which is why the last rule can sayNUMbare.- The
%{ ... %}prelude holds helper functions (and imports) for the blocks.toNumberis overkill for the toy (the block could just saytext.toInt), but it shows where shared helpers live; your IC lexer’s string-escape processing will want one. - Each
%tokenrow is a fixed-text token: the row is itself a rule matching exactly that text, sitting at that spot in the rule order. Every keyword, operator, and punctuation mark is one row; no block is needed, because there is nothing to compute. - Not every enum member has a row:
NUMhas none, because its tokens are built by the block at the bottom.
Third, the program that uses the lexer. It loads the spec by naming the enum and gets back a working lexer:
// Toy.scala -- load the spec and lex something
package toy
import slex.{LexError, LexSpec}
@main def demo(input: String*) =
val toy = LexSpec.load[K]("toy.slex")
val line = if input.isEmpty then "while (2 <= 10)" else input.mkString(" ")
try
for t <- toy.tokenize(line) do
println(t) // [WHILE,while,1], [LP,(,1], [NUM,2,1], ...
catch case e: LexError => println(s"lex error at ${e.pos}: ${e.message}")The tokens are scup.Token[K] values: a kind, the lexeme
text, the position, and a value field where a block can
stash a computed result, like the parsed Int above (or, in
IC, an unescaped string). More on that below.
The Two Laws
At every step, a scanner must answer: which rule matches next, and how much input does it take? Several rules may match, at several different lengths. Lex settled the question with two tie-breaking laws, and slex follows them. Most of the characteristic behavior of a lexer traces back to these two laws.
(The little specs in this section show just the rules each law needs.
In the companion repository they are rel.slex and
kw.slex, each with its %kinds line and a tiny
driver: sbt "runMain laws.rel" and
sbt "runMain laws.kw" show the laws on real input.)
Law 1: maximal munch. The rule that matches the longest lexeme wins, regardless of where it appears in the list.
# rel.slex -- two relational operators
%token
LT "<"
LEQ "<="
On the input <=<, the LT row comes
first and matches at the start, but LEQ matches
two characters there, and length beats listing order: the
tokens are LEQ then LT. This is why
x<=y can never lex as x, <,
=, y, the same rule you applied by hand in HW
2.
Law 2: the first rule wins ties. If several rules match the same longest lexeme, the one listed first wins.
# kw.slex -- one keyword and an identifier rule
%token
WHILE "while"
[a-z]+ { Token(ID, text, pos) }
On the input while, both the keyword row and the
identifier rule match all five characters. The tie goes to the keyword
because its row is listed first. To handle keywords in a slex spec, put
the %token table above the identifier rule; no keyword map
is needed, because first-rule-wins picks the keyword row on an exact
match. But note what happens on whilewhilst: the identifier
rule matches more characters, and Law 1 outranks Law 2, so it
lexes as one ID. A keyword row only ever claims a lexeme
that is exactly the keyword.
Order the rules carelessly and Law 2 bites: list the identifier rule
above the %token table and it wins every tie, so
while becomes an ordinary identifier and the keyword rows
are dead code.
(The other classic design puts the keyword lookup in the identifier
rule’s action, one rule consulting a
Map[String, K], and your block is ordinary Scala, so
nothing stops you. With the %token table doing the same job
declaratively, IC’s spec never needs to.)
One footnote to Law 1: a zero-length match never wins. A
rule like a* matches the empty string at every position; if
that counted, the scanner would emit empty tokens forever without
advancing. slex ignores matches of length zero, so a rule fires only
when it matches at least one character.
Watching It Scan
A scan can print a trace of its decisions. Pass
trace = true to tokenize (or
scan) and the lexer prints, one munch at a time, exactly
what the engine is doing: which rules are still alive after each
character, every point where some rule accepts (and who wins the tie),
and the final munch:
spec.tokenize("whilewhilst", trace = true)(A String => Unit sink also works there, so
trace = println means the same thing, and scup’s
parse takes the same argument; you will meet it again in
the scup tutorial.) The laws and mini demos take a -trace
flag:
sbt "runMain laws.kw -trace whilewhilst"
The trace opens with the numbered rule list and the machine running
it (the DFA those rules compile to; Inspecting the Machine shows it),
then reports each munch. Here is the trace for whilewhilst,
with both laws visible in a single lexeme (middle characters
elided):
Rules (listing order breaks ties):
#1 "while"
#2 [a-z]+
Engine: DFA (8 states)
scan at 1:1: "whilewhilst"
'w' alive: #1 "while", #2 [a-z]+
accept: #2 [a-z]+ (best is now 1 char)
...
'e' alive: #1 "while", #2 [a-z]+
accept: #1 "while", #2 [a-z]+ -- first listed wins: #1 "while" (best is now 5 chars)
'w' alive: #2 [a-z]+
accept: #2 [a-z]+ (best is now 6 chars)
...
't' alive: #2 [a-z]+
accept: #2 [a-z]+ (best is now 11 chars)
(end of input)
longest match: #2 [a-z]+ takes 11 chars
=> [ID,whilewhilst,1]
Both laws appear in it. At five characters both rules accept, and the
tie goes to the row listed first; the first listed wins: #1
line is Law 2 claiming the keyword. But the identifier rule is still
alive, keeps accepting, and the longest accept takes the munch.
The last two lines are Law 1 outranking Law 2, lexing
whilewhilst as one ID.
One more thing to watch for: when the machine consumes characters
past its last accept while hunting for a longer match and then
dies, the longest match line reports the retreat,
(backing up 2), and scanning resumes after the lexeme
actually taken. That give-back is the cost of maximal munch; an HW 2
problem steps through an example.
The trace is for watching the algorithm now and for debugging your IC
lexer later. When a lexeme comes out with the wrong kind or the wrong
length (a keyword lexed as an identifier, <= split in
two, a comment rule eating half the file), trace = true on
one small failing input shows which rules were alive, who accepted
where, and why the winner won.
Inspecting the Machine
The trace shows the machine running; the check CLI’s
-dump flag shows the machines themselves: the NFA the rules
compile to, and the DFA built from it. For rel.slex (the
two relational operators from Law 1), the whole thing fits on a
screen:
$ sbt "runMain slex.check -dump rel.slex"
rel.slex: OK (2 rules, 0 actions)
rel.slex: NFA 11 states, DFA 3 states
Rules (listing order breaks ties):
#1 "<" -> LT (line 6)
#2 "<=" -> LEQ (line 6)
NFA (11 states; state 0 is the start, with an ε-edge to each rule's fragment):
0 --ε--> 1, 5
1 --ε--> 2
2 --ε--> 3
3 --'<'--> 4
4 accepts #1 "<"
5 --ε--> 6
6 --ε--> 7
7 --'<'--> 8
8 --ε--> 9
9 --'='--> 10
10 accepts #2 "<="
DFA (3 states; state 0 is the start; each state lists its NFA states):
0 = {0, 1, 2, 3, 5, 6, 7}
--'<'--> 1
1 = {4, 8, 9} accepts #1 "<"
--'='--> 2
2 = {10} accepts #2 "<="
The rule table at the top summarizes the spec: each rule’s pattern,
the token it produces, and its line, in priority order, because the
listing order is the tie-breaker. Below it is the combined NFA those
rules compile to: state 0 fans out by ε-edges to every rule’s fragment,
states 1–4 are the fragment for <, states 5–10 the
fragment for <= (follow the chain: ε-glue, match
<, ε-glue, match =), and each fragment ends
in an accept state tagged with its rule’s number, the tag that
makes first-rule-wins a one-line min. The machine has more
ε-edges than the one you would draw by hand, because Thompson’s
construction glues fragments together instead of merging states;
simulate both and they accept the same strings.
Below the NFA is what the subset construction from HW 1
makes of it: a DFA whose every state is a set of NFA states (the
braces), with one edge per character instead of ε-glue. For
rel.slex it is the three-state machine you would have drawn
by hand – state 0 has read nothing, state 1 has read <
and accepts rule #1, state 2 has read <= and accepts
rule #2 – and that table, not the NFA, is what a scan walks: one array
lookup per character. Where two rules accept in the same DFA state, the
dump says so and names the winner, which is Law 2 precomputed.
-dump works on any spec. Point it at
mini.slex (103 NFA states, 30 DFA states) or, in PA 1, at
your ic.slex when you want to see what a real token
language compiles to, or when you are debugging: a rule whose accept
state you cannot reach from its fragment’s start is a rule that can
never fire.
Writing Patterns
A rule’s pattern is a lex-style regular expression, ended by the
first whitespace outside a character class (write \ for a
literal space):
- literal characters:
while; .: any character except newline (as in lex);- escapes
\n\t\r, and\xfor any metacharacterx(so\*\+matches the two characters*+); - postfix
*(zero or more),+(one or more),?(optional); - alternation
|and grouping( ); - character classes
[abc], ranges[a-z0-9], and negation[^...](so//[^\n]*matches a line comment).
These are just the five inductive cases of the regular-expression
definition from class (empty string, single symbol, concatenation,
alternation, star); everything else (+, ?,
classes, literal strings) is sugar for those five. A malformed pattern
(say, [a-z) is reported, with its line and what was
expected, the moment the spec loads. If you use the course VSCode
extension, it is also a squiggle under the offending character as you
type.
A Complete Example: A Miniature IC Lexer
Here is a small cut of the IC lexer, showing all the moving parts at once:
# mini.slex -- a small cut of the IC lexer
%kinds mini.K
%token
WHILE "while" IF "if" ELSE "else"
LEQ "<=" EQ "==" ASSIGN "="
LT "<" LP "(" RP ")"
LB "{" RB "}" SEMI ";"
[ \t\r\n]+
//[^\n]*
[a-zA-Z][a-zA-Z0-9_]* { Token(ID, text, pos) }
[0-9]+ { Token(NUM, text, pos) }
// Mini.scala -- lexing one line with the mini lexer
val mini = LexSpec.load[K]("mini.slex")
mini.tokenize("while (i <= 10) { i = i1; } // done")
// WHILE:"while", LP:"(", ID:"i", LEQ:"<=", NUM:"10", RP:")",
// LB:"{", ID:"i", ASSIGN:"=", ID:"i1", SEMI:";", RB:"}"Things to notice:
Scanning is one token at a time.
spec.scan(src)starts a scan ofsrc; eachnext()call returnsSome(token), orNoneat the end of the input.tokenizeis the scan-everything convenience built on top.Skip rules discard their lexemes: no whitespace or comment tokens appear in the output. But the discarded text still advances the line and column counters, so positions stay consistent: after
12\n 34, the secondNUMis at line 2, column 3.Every token carries the position of its first character, which is what error messages in every later phase of the compiler will point at.
Unmatched input is an error, not a crash. If no rule matches even one character, the scan throws a
LexErrornaming the offender and where it is. Because tokens are scanned one at a time, every token before the bad character is produced before the error is thrown, exactly the “process all the valid tokens, then report the first lexical error” behavior PA 1 requires.
Actions: The Escape Hatch
The patterns decide where tokens begin and end; the blocks decide everything else. A block is arbitrary Scala, and that is where all the non-regular work of a real scanner lives.
Transforming the lexeme. A string literal’s token should carry the string’s contents, not its source spelling. The block is where the quotes come off:
# one rule from strings.slex
"[^"\n]*" { Token(STR, text.substring(1, text.length - 1), pos) }
Your IC lexer’s version will go further and translate escape
sequences (\n, \t, \",
\\) into the characters they denote, a few lines of
ordinary string processing (a helper in the %{ ... %}
prelude keeps the rule readable) that no regular expression could do for
you.
Checking and erroring. Blocks report lexical errors with the
in-scope error(msg) helper, which is how a scanner rejects
lexemes that are regular in shape but illegal in
value. IC’s integer literals must fit in 32 bits, and
[0-9]+ has no opinion about that:
# one rule from strings.slex
[0-9]+ {
val n = text.toIntOption.getOrElse(error("Integer out of range."))
Token(NUM, text, pos, n)
}
(toIntOption, not toInt, because a 40-digit
literal should produce your error, not a
NumberFormatException.) Note the fourth argument: the
parsed Int rides along in the token’s value
field, so later phases never re-parse the text.
The catch-all idiom. What about a string literal that never
closes? No string rule matches it, so without help the scanner would
report only illegal character '"'. The fix is a second,
lower-priority rule that matches just the opening delimiter and whose
block is the error:
# two rules from strings.slex: the real rule, then the catch-all
"[^"\n]*" { Token(STR, text.substring(1, text.length - 1), pos) }
" { error("Beginning of unterminated string literal.") }
This works because whenever a properly terminated string starts here,
the first rule matches a longer lexeme (even the empty string
"" is two characters), so maximal munch picks it and the
catch-all never fires. The lone-quote rule can only win when no
terminated string begins at this position (the error case), so the
message can describe the real problem and its location. The same idiom
handles unterminated /* ... */ comments in PA 1: a real
comment rule first, then a rule for a bare /* whose block
errors; a final .|\n rule turns any other stray character
into a precise message.
The Enum Contract
Token kinds are a real Scala 3 enum in your project, not strings:
your parser’s actions and your compiler’s later phases pattern-match on
them with full type checking. The spec binds to the enum at load time:
every name the spec uses must be an enum member. A typo in a
%token row is caught the moment the spec loads, with its
line and a did-you-mean; a typo inside a block is an ordinary compile
error (the enum’s members are in scope there).
Enum members with no %token row are simply the ones
whose tokens a block builds (ID, INTLIT). The
check CLI lists them, so a glance shows exactly which kinds your blocks
are responsible for:
$ sbt "runMain slex.check ic.slex"
ic.slex: OK (52 rules, 8 actions)
ic.slex: kinds produced by action blocks (no %token row):
ID, CLS_ID, INTLIT, STRINGLIT, BOOLLIT
A free consequence of the %token table: the parser’s
error messages print fixed-text kinds as their source text
(expected '{' or 'extends') with no display-name map
anywhere.
Errors Reach the Host
A block’s error(msg) throws
slex.LexError(msg, pos); input no rule matches at all
throws the same. Your compiler’s lexer shim (a provided, ~40-line class)
converts it to IC’s own LexicalError. Since blocks are
plain Scala, a block that needs full control can also throw
LexicalError directly, adjusting the column arithmetic as
it pleases.
Editor Support
The course VSCode extension (“slex & scup”, installed in the note
at the top of this tutorial; course projects also recommend it
automatically when opened) understands .slex files: Scala
highlighting inside blocks and preludes, Format Document
for the column layout, outline, squiggles under malformed patterns as
you type, and, with the check CLI wired in, every load-time validation
as a squiggle on save. Navigation understands the spec: go-to-definition
on a token kind steps from its %token row to the action
that constructs it to the enum member itself, and %kinds
jumps to the enum declaration. Inside blocks and the prelude,
go-to-definition and hover behave like ordinary Scala once the project
has compiled (the build’s generated actions object powers them; until
the next compile after an edit, navigation quietly falls back to the
spec-level answers).
Under the Hood
The engine does no code generation: the spec is interpreted by the same three algorithms you build or run by hand this semester, plus a small driver. The library’s four core files (Re, Nfa, Dfa, Lexer) total about 650 lines; read them end to end.
Re.scala — the regular-expression AST: the five inductive cases, with combinators and character classes as sugar. Each pattern in your spec is parsed into one of these (by ReSyntax.scala, itself written with a scup grammar over character tokens).
Nfa.scala —
Nfa.compileruns the Thompson construction (the same construction you implement in the HW), turning each rule’sReinto an NFA fragment with one start and one accept state. The fragments are joined into a single combined machine: a fresh start state with an ε-edge to every rule’s fragment, and each fragment’s accept state tagged with its rule’s index. One machine runs all the rules at once.Nfa.longestMatchis the NFA simulation you built in Lab 1: keep a set of current states, and repeatedly take the ε-closure and step on the next character. Every time the current set contains an accepting state, record(length so far, lowest rule tag present)as the best match so far. When the machine dies or the input ends, the remembered best implements maximal munch, and taking the lowest tag implements first-rule-wins; each is about one line of code. It is no longer what a scan runs by default, but it is one flag away:tokenize(src, engine = Engine.Nfa).Dfa.scala —
Dfa.compileruns the subset construction from HW 1 over the ASCII alphabet: each DFA state is a set of NFA states, the start state is the ε-closure of the NFA’s start, and the transition from a state on a character is the ε-closure of everything its members reach on that character. Per state it precomputes what the scan needs: the lowest rule tag among the accepting members (first-rule-wins, folded into the table) and, for traces, every accepting tag and which rules are still alive.Dfa.longestMatchis then a table walk – one array lookup per character – that remembers the last accepting length exactly as the simulation does. Same answer, about a hundred times faster on IC programs. (Thetraceargument of Watching It Scan is either engine narrating through one shared voice; Inspecting the Machine prints both machines.)Lexer.scala — the driver: repeatedly call the chosen engine’s
longestMatch, run the winning rule’s action (or discard, for a skip rule), advance the line/column counters over the lexeme, and loop. Input is ASCII: the first character outside it ends the scan with aLexErrorat its position, whichever engine runs.
Where You Will Use It
In PA 1 you will write ic.slex, the complete IC token
language: a %token table of every keyword and operator,
rules for whitespace, both comment forms, identifiers, and literals,
plus blocks for string-escape processing, integer range checks, and
unterminated-string/comment errors. The two laws and the catch-all idiom
above cover most of that assignment; the finished spec plus the ~25-line
TokenKind enum is the entire lexer.
One note on realism. Production compilers (clang, rustc, javac) hand-write their lexers rather than generate them, chiefly for the quality of their error messages (and, less than folklore suggests, for speed). Writing PA 1’s blocks will show you why: the hard parts of a real scanner are the ones regular expressions cannot express.