
This tutorial introduces LL(1) predictive parsing and the
.scup spec files we use all semester: in the Calculator
(Lab 2), the RE compiler (HW 4), and the IC parser you will build in PA
2. A parser is a spec file: a yacc-flavored grammar
whose productions carry blocks of Scala that build your AST. The scup
library reads the spec and runs it directly: there is no generated
parser to look at, no extra build step, and the spec file is the
grammar.
Every example below is a real file in the tutorial’s companion
repository: a grammar spec (calc0.scup,
calc.scup, conflict.scup, …) plus a small
program that loads and runs it. Clone the repository (which also
contains the library source – read it, it is short) and run and modify
the examples as you read:
git clone https://github.com/williams-cs/cs434-f26-scup-tutorial
cd cs434-f26-scup-tutorial
sbt "runMain calc0.demo"
The demo parses the tutorial’s sample input, and anything you pass is
parsed instead: sbt "runMain calc0.demo (1+2)*3". The
tutorial’s claims about the examples are also an MUnit suite –
sbt test runs them all.
Editor support. If you use VSCode, install the
course’s “slex & scup” extension before reading on: download the .vsix and run
code --install-extension slex-scup.vsix (or, in VSCode,
Extensions → “···” → Install from VSIX…). What it does for
.scup files is described in Editor Support below.
The Idea: Grammars You Can Run
Yacc (1975) and its successors (bison, CUP, ANTLR) are parser generators: you write a grammar, one rule per nonterminal with an action to run at each production, and the tool writes the parser. The engine decides which production applies; your actions decide what to build.
A .scup file reads like yacc. Before parsing anything,
the library analyzes your grammar, computing the same nullable,
FIRST, and FOLLOW sets you compute by hand in HW 2, and either builds an
LL(1) prediction table or rejects the grammar with a message naming the
conflicting productions. Parsing is then predictive: one token
of lookahead selects each production, the engine never backtracks, and
error messages state exactly which tokens would have been legal.
Here is a complete calculator, deliberately plain: one rule per precedence level, written with nothing but rules, alternatives, and action blocks:
# calc0.scup -- a first calculator, deliberately plain: one rule per
# precedence level, written with nothing but rules, alternatives, and
# action blocks. Naive on purpose (no subtraction; the recursion
# leans right) -- the tutorial fixes and shrinks it into calc.scup,
# and it is exactly the starter grammar you will repair in Lab 2.
%start expr
%type expr: Int
%type exprTail: Int
%type term: Int
%type termTail: Int
%type factor: Int
expr ::= term exprTail { $1 + $2 } ;
exprTail ::= "+" expr { $2 }
| %empty { 0 } ;
term ::= factor termTail { $1 * $2 } ;
termTail ::= "*" term { $2 }
| %empty { 1 } ;
factor ::= NUM { $1.text.toInt }
| "(" expr ")" { $2 } ;
Read it as the
expr/term/factor stratification
you built by hand in HW 2: an expr is terms
added together, a term is factors multiplied,
and a factor is a number or a parenthesized expression. A
rule lists its productions after ::=, separated by
|; %empty is the empty production; and a
production’s { ... } block computes the rule’s value from
its symbols’ values ($1, $2, …; the next
sections have much more to say about them). %start names
the start symbol, and each %type line declares a rule’s
value type. One idiom to notice: an absent tail contributes its
operation’s identity (0 for +, 1
for *), so { $1 + $2 } is right whether or not
a tail followed.
And as its header comment admits, the grammar is naive: no
subtraction, and its recursion leans right, which is harmless for
+ and * but wrong the moment -
arrives. Both flaws fall to shorthands introduced below, which will also
shrink the file by half; the repaired, finished calc.scup
appears in the %binary section.
The terminals in quotes resolve through a companion
.slex token spec. The calculator’s is a dozen lines,
straight out of the slex tutorial (its
token-kind enum, calc.K, is the usual one-line-per-name
Scala file next door):
# calc.slex -- the calculator's token language
%kinds calc.K
%token
PLUS "+" MINUS "-" STAR "*"
SLASH "/" LP "(" RP ")"
COMMA ","
[ \t\r\n]+
[0-9]+ { Token(NUM, text, pos) }
The host is the short program that loads the two specs and runs them – lexer errors and parse errors both reach it as ordinary exceptions:
// Calc0.scala -- load the plain calculator and run it
package calc0
import calc.K
import slex.{LexError, LexSpec}
import scup.{GrammarSpec, ParseError}
@main def demo(input: String*) =
val tracing = input.headOption.contains("-trace") // see the tutorial's Watching It Parse
val words = if tracing then input.drop(1) else input
val lex = LexSpec.load[K]("calc.slex")
val g = GrammarSpec.load("calc0.scup", lex.binding)
val line = if words.isEmpty then "1+2*3" else words.mkString(" ")
try println(g.parseAs[Int](lex.tokenize(line), trace = tracing)) // 7
catch
case e: LexError => println(s"lex error at ${e.pos}: ${e.message}")
case e: ParseError => println(e.getMessage)Watching It Parse
A parse can print a trace of its decisions. Pass
trace = true to any parse call, one extra argument:
g.parseAs[Int](lex.tokenize("1+2*3"), trace = true)(A String => Unit sink also works there, so
trace = println means the same thing, and a test can
collect the lines into a buffer.)
— or, in the companion repository, give the calculator the
-trace flag:
sbt "runMain calc0.demo -trace 1+2*3"
The trace opens with the grammar’s productions, numbered (for this grammar, the eight productions you just read), and then one line per step the parser takes:
Productions:
( 1) expr ::= term exprTail
( 2) term ::= factor termTail
( 3) factor ::= NUM
( 4) | '(' expr ')'
( 5) termTail ::= '*' term
( 6) | %empty
( 7) exprTail ::= '+' expr
( 8) | %empty
parse: 1 + 2 * 3
expr NUM '1' @1:1 predict (1)
term NUM '1' @1:1 predict (2)
factor NUM '1' @1:1 predict (3)
match NUM '1' @1:1
factor = 1
termTail '+' @1:2 predict (6)
termTail = 1
term = 1
exprTail '+' @1:2 predict (7)
match '+' @1:2
expr NUM '2' @1:3 predict (1)
term NUM '2' @1:3 predict (2)
factor NUM '2' @1:3 predict (3)
match NUM '2' @1:3
factor = 2
termTail '*' @1:4 predict (5)
match '*' @1:4
term NUM '3' @1:5 predict (2)
factor NUM '3' @1:5 predict (3)
match NUM '3' @1:5
factor = 3
termTail end of input predict (6)
termTail = 1
term = 3
termTail = 3
term = 6
exprTail end of input predict (8)
exprTail = 0
expr = 6
exprTail = 6
expr = 7
7
Read it top to bottom, following the indentation:
- A
predictline is the only kind of decision the algorithm makes: at ruleexpr, with lookaheadNUM '1', the LL(1) table selects production (1). The parser never guesses and never backs up; each production is entered because the lookahead token selected it. - A
matchline consumes one token of input; nothing else does. %emptyis predicted like anything else: attermTailwith lookahead'+', the table picks the empty production (6), because+cannot begin atermTail. That is how the parse climbs back out of a level.- As each rule finishes, the value its action built rides upward:
factor = 2, thenterm = 6after the* 3, thenexpr = 7at the root. These values are the subject of the next section.
The trace is for watching the algorithm now and for debugging
grammars later: in PA 2, when your IC parser commits to a production you
did not expect, trace = println on one small failing input
shows the decision that led it there, and the lookahead token that
forced it.
Tokens In, Values Out
Symbols in a production are terminals and nonterminals. A terminal is
written either as a token-kind name (NUM, ID)
or, for fixed-text tokens, as its literal spelling in quotes:
"while", "<=", "(". The
literals resolve through the companion .slex spec’s
%token table, so the grammar never mentions
LPAREN-style names for punctuation, and error messages
print expected '(' for free. A nonterminal is a rule
name.
In a block, $1, $2, … refer to the
production’s symbols, in order, 1-based over all of
them. A terminal’s value is its token: $2.text,
$2.line, $2.pos, and $2.value
(where your lexer block stashed a parsed Int or unescaped
string). A nonterminal’s value is whatever its rule’s block built. What
a block does not mention is simply dropped; where yacc actions skip
$2, yours never name it:
# one rule from some .scup file
point ::= "(" NUM "," NUM ")" { Point($2.text.toInt, $4.text.toInt) } ;
Declaring %type point: Point types the block: its
$i parameters and its result are checked by the Scala
compiler against the grammar, with errors reported at this
file’s own line numbers: get the symbol index or the type wrong and the
spec does not load. Undeclared rules default to Any;
declare everything, as the IC grammar does.
Two defaults keep specs terse. A one-symbol production with no block
passes its value through (see factor’s first form used
alone: stmt ::= block | ifStmt ... ;). And a production of
more than one symbol requires a block, unless the whole rule is
action-free (more on that mode below).
Choice Is Predictive
The alternatives of a rule are not tried in order; there is no order. When the engine reaches a rule, it looks at the next token and asks: which production can begin with this token? For that question to have a unique answer, the productions’ FIRST sets must not overlap. If they do, the grammar is rejected at load, not at parse time, and not for some unlucky input:
# from conflict.scup -- a rule that is not LL(1)
s ::= "(" b
| "(" c ;
$ sbt "runMain scup.check conflict.scup"
conflict.scup:12: rule 's' is not LL(1): productions 1 and 2 can
both begin with '(' (FIRST/FIRST conflict); left-factor the common
prefix
This is the LL(1) condition from lecture, enforced mechanically. The fix is the one you practice in HW 2: left-factor the common prefix, deferring the decision until one token suffices:
# the same rule, left-factored
s ::= "(" sTail { ... } ;
sTail ::= b | c ;
You will do the same rewrite many times in PA 2: an IC class member
commits to field vs. method only after
Type ID, and a statement beginning with an expression
commits to assignment vs. call at the following
= or ;.
Lists, Options, and Greedy Repetition
LL(1) grammars cannot use the left-recursive list idiom every yacc
grammar contains, so scup builds the standard shorthands in: the EBNF
glyphs x*, x+, and x? as postfix
modifiers on any symbol, written exactly as the grammars in the course
handouts write them, plus %list for the ubiquitous
separated list (which no standard notation covers):
# the shorthands, as they appear in a production
stmt* # zero or more -> List
digit+ # one or more -> List
expr? # zero or one -> Option
%list(formal, ",") # comma-separated -> List, empty allowed
The list forms produce a List of the symbol’s values,
and x? produces an Option (Some
of the value, or None), which blocks default away in one
line: $3.getOrElse("") (the companion repository’s
opt.scup is a two-line demonstration). An
%empty alternative with no block likewise produces
None, so a rule shaped x ::= ... | %empty ;
declares an Option[...] type and its other alternatives
wrap their results in Some. Internally each shorthand
becomes an ordinary (right-recursive) rule that shows up in dumps under
its surface name (stmt*, %list(formal, ',')),
and the engine runs the loops iteratively, so lists of any length cost
constant stack.
The finished calculator uses the list form to parse whole
programs: calc.scup in the companion repository
declares program ::= %list(expr, ",") ;, turning
1+2*3, 10 into List(7, 10). Its trace
(sbt "runMain calc.demo -trace 1+2, 10") is where loop
lines first appear; the engine runs each list as a loop, deciding on one
token whether to continue:
(',' expr)*: ',' @1:4 continues loop
(',' expr)*: end of input stops loop
%list(expr, ',') = List(3, 10)
One subtlety: repetition is greedy. If a token could either continue a repetition or follow it, the repetition continues. This is the standard choice in practical LL tools, and the one that binds trailing operators to the innermost construct. A token that conflicts with the repeated symbol’s own FIRST set is still a hard error.
A parenthesized alternation of terminals, as in
("-" | "!"), is the one inline group allowed; anything more
structured gets its own named rule, yacc-style.
Left Recursion,
%binary, and %chain
The natural grammar for subtraction is left-recursive —
expr ::= expr '-' term | term
— and left recursion is fatal for a parser that must decide using one
token of lookahead: the rule would recurse forever without consuming
anything. scup detects it and tells you (leftrec.scup in
the companion repository is exactly this grammar;
sbt "runMain scup.check leftrec.scup" shows the report).
The rewrite you would do by hand (HW 2 again) parses
term ('-' term)* and folds the results leftward; a
%left/%right block followed by a
%binary rule packages the whole ladder, lowest precedence
first:
# from ic.scup: the whole binary-operator ladder
%left "||"
%left "&&"
%left "+" "-"
%left "*" "/" "%"
expr ::= %binary(unaryExpr) { BinOpNode($1, $3, $2.text, $1.line) } ;
One block is the combine for every level: $1
and $3 are the operands, $2 the operator
token. 9-2-3 still parses to (9-2)-3 — the
tree leans left even though the grammar recursion leans right. Your IC
parser’s six binary levels are exactly the declaration above.
Here is calc.scup, the plain calculator from the top of
this tutorial rebuilt with these shorthands: the two tail rules per
level are gone, subtraction and division are in, the associativity flaw
is fixed (9-2-3 is 4), and whole programs are
one %list line:
# calc.scup -- the calculator's grammar
%start program
%type program: List[Int]
%type expr: Int
%type term: Int
%type factor: Int
program ::= %list(expr, ",") ;
%left "+" "-"
expr ::= %binary(term) { if $2.text == "+" then $1 + $3 else $1 - $3 } ;
%left "*" "/"
term ::= %binary(factor) { if $2.text == "*" then $1 * $3 else $1 / $3 } ;
factor ::= NUM { $1.text.toInt }
| "(" expr ")" { $2 } ;
Run it with sbt "runMain calc.demo", and with
-trace to compare its parse of 1+2*3 against
calc0’s. This repair, a %left block and a
%binary rule replacing a right-leaning tail, is what Lab 2
asks you to do.
The other left-recursion idiom is the postfix loop:
e.f, e[i], e.m(args),
e.length, each wrapping the expression built so far.
%chain(base, suffixRule) parses base suffix*
and folds leftward, with $$ bound to the accumulated value
inside the suffix rule’s blocks:
# the postfix loop from ic.scup (abridged)
postfixExpr ::= %chain(primary, postfixOp) ;
postfixOp ::= "." "length" { LengthNode($$, $$.line) }
| "[" expr "]" { ArrayLocationNode($$, $2, $$.line) } ;
Each alternative of the suffix rule builds the node for one more
postfix step around the value accumulated so far. The full rule in
ic.scup handles .f and .m(args) the same way:
it parses what follows the . into a small enum and matches
on it, with every case wrapping $$.
The Analysis You Do by Hand
When a grammar loads, scup computes nullable, FIRST, and FOLLOW by
the fixpoint algorithms of Dragon 4.4 (the same computations, on the
same kind of grammar, as HW 2), then builds the LL(1) table. All of it
is inspectable: the check CLI loads a grammar, runs the analysis, and
(with -dump) prints the productions (numbered), the
nullable set, FIRST and FOLLOW for every rule, and the prediction table
mapping each (rule, lookahead) pair to a production number: the same
table you construct on paper, in the same shape as the Dragon book’s
Figure 4.17:
Here is the whole thing for bare.scup, an action-free
grammar in the companion repository whose two rules are
s ::= "(" nums ")" ;
nums ::= %list(NUM, ",") ;
$ sbt "runMain scup.check -dump bare.scup"
bare.scup: OK -- LL(1), 2 rules (start: s)
Grammar (start symbol: s)
Productions:
( 1) s ::= '(' nums ')'
( 2) nums ::= %list(NUM, ',')
( 3) %list(NUM, ',') ::= NUM (',' NUM)*
( 4) | %empty
( 5) (',' NUM)* ::= ',' NUM (',' NUM)*
( 6) | %empty
Nullable: nums, %list(NUM, ','), (',' NUM)*
FIRST:
s : '('
nums : NUM
%list(NUM, ',') : NUM
(',' NUM)* : ','
FOLLOW:
s : $
nums : ')'
%list(NUM, ',') : ')'
(',' NUM)* : ')'
LL(1) table (entries are production numbers; $ is end of input):
'(' ')' ',' NUM
s 1 . . .
nums . 2 . 2
%list(NUM, ',') . 4 . 3
(',' NUM)* . 6 5 .
The %list shorthand has become the two ordinary
right-recursive rules described above — productions 3–6
— and that they are nullable with the FIRST and FOLLOW sets you would
compute for them by hand. The table is read exactly like the one you
build on paper: parsing nums and seeing ')'
selects production 2; a . cell is an error, and its row is
the “expected …” set a ParseError reports. Point the
command at any .scup file, including one you write.
Two good uses: to check your homework, type an HW 2 exercise grammar
into a .scup file (no actions, no types, no Scala) and
compare the dump against your hand computation; and to debug a grammar,
since when a conflict report surprises you, the FOLLOW sets in the dump
usually explain it.
Grammars Without Actions
That homework flow works because a bare grammar is a valid spec: a
rule with no blocks and no %type still loads, LL(1)-checks,
and parses, building a generic labeled tree (STree) with
one node per production, single-symbol productions passing through. Zero
Scala required. (In this mode any name no rule defines is treated as a
terminal, so the checker prints a note for each one that looks like a
rule name; a misspelled nonterminal would otherwise silently “check out”
as a terminal.) This is also a fine way to draft a grammar in
PA 2: get the factoring LL(1)-clean first, then add %types
and blocks rule by rule.
The Dangling Else
One conflict is famous enough to have a standard resolution. In
stmt ::= 'if' '(' expr ')' stmt elseTail | ...
elseTail ::= 'else' stmt | %empty
an else can begin elseTail but can
also follow it (when the if was nested inside
another if). That is a FIRST/FOLLOW conflict, and scup
rejects it like any other, unless you say you mean it:
# from ic.scup: the dangling else
%greedy
elseTail ::= "else" stmt { $2 }
| %empty { EmptyStmtNode() } ;
%greedy resolves a FIRST/FOLLOW conflict in favor of the
production that consumes input: each else binds to the
nearest unmatched if. This is yacc’s prefer-shift
resolution of the same conflict, made explicit. It is the only
%greedy in the whole IC grammar; anywhere else, left-factor
instead.
Errors
Because parsing is predictive, error messages are exact: the engine knows which terminals have table entries at the point of failure:
ParseError: expected ')', '*', '+' or '-' at line 1, column 4
(while parsing 'expr')
parseAll/parseAs require the whole input to
be consumed and throw ParseError otherwise; your compiler’s
host catches it and reports a SyntaxError in IC’s format.
Fixed-text terminals print as their %token spelling
(expected '{') with no display-name map anywhere. There is
no “furthest failure” guesswork as in backtracking parsers: the reported
position is the exact token where prediction failed.
Preludes and Hosts
A %{ ... %} prelude at the top of the spec holds the
imports your blocks need (your AST package, the token-kind enum), small
helper defs, and parser-only types: little enums
and case classes that left-factored tail rules return for the owning
rule to match on. Prelude definitions live inside the spec’s compiled
actions; no other phase of your compiler can see them, which is what you
want for a parser-only type. Types the rest of the compiler uses (your
AST) belong in ordinary project source files. The host is the
short Scala object that loads the two specs and owns the entry
point:
// ICParser.scala -- the host object (provided in PA 2)
object ICParser:
private lazy val g = GrammarSpec.load("ic.scup", Lexer.spec.binding)
def parse(source: String): ProgramNode =
val toks = Lexer.specTokens(source)
try ProgramNode(g.parseAs[List[ClassNode]](toks), 1)
catch case e: ParseError => throw SyntaxError(e.expected, e.pos)The binding carries what the grammar needs from the
token spec: the enum, the %token literal table, and the
display names.
Editor Support
The course VSCode extension (“slex & scup”, installed in the note
at the top of this tutorial) understands .scup files: Scala
highlighting inside blocks, Format Document for the
::=/| layout, outline over rules, and, with
the check CLI wired in, LL(1) conflicts as squiggles on the offending
rule. Go-to-definition follows a rule reference to its definition, a
"literal" to its %token row in the companion
.slex, a token kind through that .slex to the
action constructing it (or the enum member), and
$1/$$ in an action to the production symbol it
names. Inside { ... } actions, go-to-definition and hover
behave like ordinary Scala once the project has compiled; hovering
$1 even shows the type the grammar gives it.
Under the Hood
Your spec is interpreted, not compiled to a parser: the loader builds the same rule/production structures the library’s embedded Scala DSL builds, and three short files implement everything; read them.
Syntax.scala —
Grammar,Rule, productions, the EBNF modifiers, precedence ladders. This is the embedded DSL the spec surface maps onto:lazy val factor: Rule[Int] = rule( NUM --> { t => t.text.toInt } | (!LP, expr, !RP) --> { e => e })is
factor ::= NUM { ... } | "(" expr ")" { ... } ;written as Scala values: productions are symbol tuples,-->attaches a typed action,!marks punctuation to drop (the spec’s blocks just ignore unmentioned$iinstead), andleftAssoc/operators(...)are%binary’s underpinnings. The DSL remains fully usable (the pattern parser inside slex is written with it, and the companion repository’sdsl/DslCalc.scalais the whole calculator written this way:sbt "runMain dsl.demo"), but the spec file is the surface the course uses. Your action blocks are compiled behind the scenes into functions over the$ivalues, which is why type errors in them point at your spec file’s own lines.Analysis.scala — nullable, FIRST, FOLLOW, and the table, as fixpoints over the productions; conflict and left-recursion detection; the dump. After HW 2, this file is your homework computations mechanized, Dragon 4.4 almost line for line.
Parse.scala — the predictive engine: look up (rule, lookahead) in the table, run the chosen production’s symbols left to right, apply the action. That is the entire runtime.
Where You Will Use It
In Lab 2 you complete the Calculator: add operators,
repair its deliberately right-associative starter grammar with a
%left block and %binary, and add
%list(expr, ",") programs; in HW 3 you
implement the nullable/FIRST fixpoints yourself
(hw/First.scala), checking your implementation against the
library’s dump.
In HW 4 you write the RE compiler’s parser: a
grammar for regular expressions whose atoms all start with distinct
terminals, plus a let-binding form handled by a small
substitution pass after parsing.
In PA 2 you write ic.scup, the IC
grammar. The IC specification’s grammar appendix gives most of the LL(1)
factorings; the statement rule and primary are deliberately
left for you, and the conflict reports will point out where an attempt
is not yet LL(1). Tokens flow straight in from ic.slex, so
the front end is one line per phase:
source --ic.slex--> tokens --ic.scup--> AST