A scup Tutorial

This tutorial introduces LL(1) predictive parsing and the .scup spec files we use all semester: in the Calculator (Lab 2), the RE compiler (HW 4), and the IC parser you will build in PA 2. A parser is a spec file: a yacc-flavored grammar whose productions carry blocks of Scala that build your AST. The scup library reads the spec and runs it directly: there is no generated parser to look at, no extra build step, and the spec file is the grammar.

Every example below is a real file in the tutorial’s companion repository: a grammar spec (calc0.scup, calc.scup, conflict.scup, …) plus a small program that loads and runs it. Clone the repository (which also contains the library source – read it, it is short) and run and modify the examples as you read:

git clone https://github.com/williams-cs/cs434-f26-scup-tutorial
cd cs434-f26-scup-tutorial
sbt "runMain calc0.demo"

The demo parses the tutorial’s sample input, and anything you pass is parsed instead: sbt "runMain calc0.demo (1+2)*3". The tutorial’s claims about the examples are also an MUnit suite – sbt test runs them all.

Note

Editor support. If you use VSCode, install the course’s “slex & scup” extension before reading on: download the .vsix and run code --install-extension slex-scup.vsix (or, in VSCode, Extensions → “···” → Install from VSIX…). What it does for .scup files is described in Editor Support below.

The Idea: Grammars You Can Run

Yacc (1975) and its successors (bison, CUP, ANTLR) are parser generators: you write a grammar, one rule per nonterminal with an action to run at each production, and the tool writes the parser. The engine decides which production applies; your actions decide what to build.

A .scup file reads like yacc. Before parsing anything, the library analyzes your grammar, computing the same nullable, FIRST, and FOLLOW sets you compute by hand in HW 2, and either builds an LL(1) prediction table or rejects the grammar with a message naming the conflicting productions. Parsing is then predictive: one token of lookahead selects each production, the engine never backtracks, and error messages state exactly which tokens would have been legal.

Here is a complete calculator, deliberately plain: one rule per precedence level, written with nothing but rules, alternatives, and action blocks:

# calc0.scup -- a first calculator, deliberately plain: one rule per
# precedence level, written with nothing but rules, alternatives, and
# action blocks.  Naive on purpose (no subtraction; the recursion
# leans right) -- the tutorial fixes and shrinks it into calc.scup,
# and it is exactly the starter grammar you will repair in Lab 2.

%start  expr
%type   expr:     Int
%type   exprTail: Int
%type   term:     Int
%type   termTail: Int
%type   factor:   Int

expr     ::= term exprTail  { $1 + $2 } ;

exprTail ::= "+" expr  { $2 }
           | %empty    { 0 } ;

term     ::= factor termTail  { $1 * $2 } ;

termTail ::= "*" term  { $2 }
           | %empty    { 1 } ;

factor   ::= NUM           { $1.text.toInt }
           | "(" expr ")"  { $2 } ;

Read it as the expr/term/factor stratification you built by hand in HW 2: an expr is terms added together, a term is factors multiplied, and a factor is a number or a parenthesized expression. A rule lists its productions after ::=, separated by |; %empty is the empty production; and a production’s { ... } block computes the rule’s value from its symbols’ values ($1, $2, …; the next sections have much more to say about them). %start names the start symbol, and each %type line declares a rule’s value type. One idiom to notice: an absent tail contributes its operation’s identity (0 for +, 1 for *), so { $1 + $2 } is right whether or not a tail followed.

And as its header comment admits, the grammar is naive: no subtraction, and its recursion leans right, which is harmless for + and * but wrong the moment - arrives. Both flaws fall to shorthands introduced below, which will also shrink the file by half; the repaired, finished calc.scup appears in the %binary section.

The terminals in quotes resolve through a companion .slex token spec. The calculator’s is a dozen lines, straight out of the slex tutorial (its token-kind enum, calc.K, is the usual one-line-per-name Scala file next door):

# calc.slex -- the calculator's token language

%kinds calc.K

%token
  PLUS  "+"   MINUS "-"   STAR  "*"
  SLASH "/"   LP    "("   RP    ")"
  COMMA ","

[ \t\r\n]+

[0-9]+      { Token(NUM, text, pos) }

The host is the short program that loads the two specs and runs them – lexer errors and parse errors both reach it as ordinary exceptions:

// Calc0.scala -- load the plain calculator and run it
package calc0

import calc.K
import slex.{LexError, LexSpec}
import scup.{GrammarSpec, ParseError}

@main def demo(input: String*) =
  val tracing = input.headOption.contains("-trace")   // see the tutorial's Watching It Parse
  val words   = if tracing then input.drop(1) else input
  val lex  = LexSpec.load[K]("calc.slex")
  val g    = GrammarSpec.load("calc0.scup", lex.binding)
  val line = if words.isEmpty then "1+2*3" else words.mkString(" ")
  try println(g.parseAs[Int](lex.tokenize(line), trace = tracing))   // 7
  catch
    case e: LexError   => println(s"lex error at ${e.pos}: ${e.message}")
    case e: ParseError => println(e.getMessage)

Watching It Parse

A parse can print a trace of its decisions. Pass trace = true to any parse call, one extra argument:

g.parseAs[Int](lex.tokenize("1+2*3"), trace = true)

(A String => Unit sink also works there, so trace = println means the same thing, and a test can collect the lines into a buffer.)

— or, in the companion repository, give the calculator the -trace flag:

sbt "runMain calc0.demo -trace 1+2*3"

The trace opens with the grammar’s productions, numbered (for this grammar, the eight productions you just read), and then one line per step the parser takes:

Productions:
  ( 1)         expr ::= term exprTail
  ( 2)         term ::= factor termTail
  ( 3)       factor ::= NUM
  ( 4)                | '(' expr ')'
  ( 5)     termTail ::= '*' term
  ( 6)                | %empty
  ( 7)     exprTail ::= '+' expr
  ( 8)                | %empty
parse: 1 + 2 * 3
expr           NUM '1' @1:1     predict (1)
  term         NUM '1' @1:1     predict (2)
    factor     NUM '1' @1:1     predict (3)
      match NUM '1' @1:1
    factor = 1
    termTail   '+' @1:2         predict (6)
    termTail = 1
  term = 1
  exprTail     '+' @1:2         predict (7)
    match '+' @1:2
    expr       NUM '2' @1:3     predict (1)
      term     NUM '2' @1:3     predict (2)
        factor NUM '2' @1:3     predict (3)
          match NUM '2' @1:3
        factor = 2
        termTail '*' @1:4         predict (5)
          match '*' @1:4
          term NUM '3' @1:5     predict (2)
            factor NUM '3' @1:5     predict (3)
              match NUM '3' @1:5
            factor = 3
            termTail end of input     predict (6)
            termTail = 1
          term = 3
        termTail = 3
      term = 6
      exprTail end of input     predict (8)
      exprTail = 0
    expr = 6
  exprTail = 6
expr = 7
7

Read it top to bottom, following the indentation:

The trace is for watching the algorithm now and for debugging grammars later: in PA 2, when your IC parser commits to a production you did not expect, trace = println on one small failing input shows the decision that led it there, and the lookahead token that forced it.

Tokens In, Values Out

Symbols in a production are terminals and nonterminals. A terminal is written either as a token-kind name (NUM, ID) or, for fixed-text tokens, as its literal spelling in quotes: "while", "<=", "(". The literals resolve through the companion .slex spec’s %token table, so the grammar never mentions LPAREN-style names for punctuation, and error messages print expected '(' for free. A nonterminal is a rule name.

In a block, $1, $2, … refer to the production’s symbols, in order, 1-based over all of them. A terminal’s value is its token: $2.text, $2.line, $2.pos, and $2.value (where your lexer block stashed a parsed Int or unescaped string). A nonterminal’s value is whatever its rule’s block built. What a block does not mention is simply dropped; where yacc actions skip $2, yours never name it:

# one rule from some .scup file
point ::= "(" NUM "," NUM ")"   { Point($2.text.toInt, $4.text.toInt) } ;

Declaring %type point: Point types the block: its $i parameters and its result are checked by the Scala compiler against the grammar, with errors reported at this file’s own line numbers: get the symbol index or the type wrong and the spec does not load. Undeclared rules default to Any; declare everything, as the IC grammar does.

Two defaults keep specs terse. A one-symbol production with no block passes its value through (see factor’s first form used alone: stmt ::= block | ifStmt ... ;). And a production of more than one symbol requires a block, unless the whole rule is action-free (more on that mode below).

Choice Is Predictive

The alternatives of a rule are not tried in order; there is no order. When the engine reaches a rule, it looks at the next token and asks: which production can begin with this token? For that question to have a unique answer, the productions’ FIRST sets must not overlap. If they do, the grammar is rejected at load, not at parse time, and not for some unlucky input:

# from conflict.scup -- a rule that is not LL(1)
s ::= "(" b
    | "(" c ;
$ sbt "runMain scup.check conflict.scup"
conflict.scup:12: rule 's' is not LL(1): productions 1 and 2 can
both begin with '(' (FIRST/FIRST conflict); left-factor the common
prefix

This is the LL(1) condition from lecture, enforced mechanically. The fix is the one you practice in HW 2: left-factor the common prefix, deferring the decision until one token suffices:

# the same rule, left-factored
s     ::= "(" sTail   { ... } ;
sTail ::= b | c ;

You will do the same rewrite many times in PA 2: an IC class member commits to field vs. method only after Type ID, and a statement beginning with an expression commits to assignment vs. call at the following = or ;.

Lists, Options, and Greedy Repetition

LL(1) grammars cannot use the left-recursive list idiom every yacc grammar contains, so scup builds the standard shorthands in: the EBNF glyphs x*, x+, and x? as postfix modifiers on any symbol, written exactly as the grammars in the course handouts write them, plus %list for the ubiquitous separated list (which no standard notation covers):

# the shorthands, as they appear in a production
stmt*               # zero or more    -> List
digit+              # one or more     -> List
expr?               # zero or one     -> Option
%list(formal, ",")  # comma-separated -> List, empty allowed

The list forms produce a List of the symbol’s values, and x? produces an Option (Some of the value, or None), which blocks default away in one line: $3.getOrElse("") (the companion repository’s opt.scup is a two-line demonstration). An %empty alternative with no block likewise produces None, so a rule shaped x ::= ... | %empty ; declares an Option[...] type and its other alternatives wrap their results in Some. Internally each shorthand becomes an ordinary (right-recursive) rule that shows up in dumps under its surface name (stmt*, %list(formal, ',')), and the engine runs the loops iteratively, so lists of any length cost constant stack.

The finished calculator uses the list form to parse whole programs: calc.scup in the companion repository declares program ::= %list(expr, ",") ;, turning 1+2*3, 10 into List(7, 10). Its trace (sbt "runMain calc.demo -trace 1+2, 10") is where loop lines first appear; the engine runs each list as a loop, deciding on one token whether to continue:

    (',' expr)*: ',' @1:4 continues loop
    (',' expr)*: end of input stops loop
  %list(expr, ',') = List(3, 10)

One subtlety: repetition is greedy. If a token could either continue a repetition or follow it, the repetition continues. This is the standard choice in practical LL tools, and the one that binds trailing operators to the innermost construct. A token that conflicts with the repeated symbol’s own FIRST set is still a hard error.

A parenthesized alternation of terminals, as in ("-" | "!"), is the one inline group allowed; anything more structured gets its own named rule, yacc-style.

Left Recursion, %binary, and %chain

The natural grammar for subtraction is left-recursive —

expr ::= expr '-' term | term

— and left recursion is fatal for a parser that must decide using one token of lookahead: the rule would recurse forever without consuming anything. scup detects it and tells you (leftrec.scup in the companion repository is exactly this grammar; sbt "runMain scup.check leftrec.scup" shows the report). The rewrite you would do by hand (HW 2 again) parses term ('-' term)* and folds the results leftward; a %left/%right block followed by a %binary rule packages the whole ladder, lowest precedence first:

# from ic.scup: the whole binary-operator ladder
%left "||"
%left "&&"
%left "+" "-"
%left "*" "/" "%"
expr ::= %binary(unaryExpr)   { BinOpNode($1, $3, $2.text, $1.line) } ;

One block is the combine for every level: $1 and $3 are the operands, $2 the operator token. 9-2-3 still parses to (9-2)-3 — the tree leans left even though the grammar recursion leans right. Your IC parser’s six binary levels are exactly the declaration above.

Here is calc.scup, the plain calculator from the top of this tutorial rebuilt with these shorthands: the two tail rules per level are gone, subtraction and division are in, the associativity flaw is fixed (9-2-3 is 4), and whole programs are one %list line:

# calc.scup -- the calculator's grammar

%start  program
%type   program: List[Int]
%type   expr:    Int
%type   term:    Int
%type   factor:  Int

program ::= %list(expr, ",") ;

%left "+" "-"
expr    ::= %binary(term)  { if $2.text == "+" then $1 + $3 else $1 - $3 } ;

%left "*" "/"
term    ::= %binary(factor)  { if $2.text == "*" then $1 * $3 else $1 / $3 } ;

factor  ::= NUM           { $1.text.toInt }
          | "(" expr ")"  { $2 } ;

Run it with sbt "runMain calc.demo", and with -trace to compare its parse of 1+2*3 against calc0’s. This repair, a %left block and a %binary rule replacing a right-leaning tail, is what Lab 2 asks you to do.

The other left-recursion idiom is the postfix loop: e.f, e[i], e.m(args), e.length, each wrapping the expression built so far. %chain(base, suffixRule) parses base suffix* and folds leftward, with $$ bound to the accumulated value inside the suffix rule’s blocks:

# the postfix loop from ic.scup (abridged)
postfixExpr ::= %chain(primary, postfixOp) ;

postfixOp   ::= "." "length"   { LengthNode($$, $$.line) }
              | "[" expr "]"   { ArrayLocationNode($$, $2, $$.line) } ;

Each alternative of the suffix rule builds the node for one more postfix step around the value accumulated so far. The full rule in ic.scup handles .f and .m(args) the same way: it parses what follows the . into a small enum and matches on it, with every case wrapping $$.

The Analysis You Do by Hand

When a grammar loads, scup computes nullable, FIRST, and FOLLOW by the fixpoint algorithms of Dragon 4.4 (the same computations, on the same kind of grammar, as HW 2), then builds the LL(1) table. All of it is inspectable: the check CLI loads a grammar, runs the analysis, and (with -dump) prints the productions (numbered), the nullable set, FIRST and FOLLOW for every rule, and the prediction table mapping each (rule, lookahead) pair to a production number: the same table you construct on paper, in the same shape as the Dragon book’s Figure 4.17:

Here is the whole thing for bare.scup, an action-free grammar in the companion repository whose two rules are

s    ::= "(" nums ")" ;

nums ::= %list(NUM, ",") ;
$ sbt "runMain scup.check -dump bare.scup"
bare.scup: OK -- LL(1), 2 rules (start: s)
Grammar (start symbol: s)

Productions:
  ( 1)            s ::= '(' nums ')'
  ( 2)         nums ::= %list(NUM, ',')
  ( 3) %list(NUM, ',') ::= NUM (',' NUM)*
  ( 4)                | %empty
  ( 5)   (',' NUM)* ::= ',' NUM (',' NUM)*
  ( 6)                | %empty

Nullable: nums, %list(NUM, ','), (',' NUM)*

FIRST:
             s : '('
          nums : NUM
  %list(NUM, ',') : NUM
    (',' NUM)* : ','

FOLLOW:
             s : $
          nums : ')'
  %list(NUM, ',') : ')'
    (',' NUM)* : ')'

LL(1) table (entries are production numbers; $ is end of input):

                '('  ')'  ','  NUM
             s    1    .    .    .
          nums    .    2    .    2
  %list(NUM, ',')    .    4    .    3
    (',' NUM)*    .    6    5    .

The %list shorthand has become the two ordinary right-recursive rules described above — productions 3–6 — and that they are nullable with the FIRST and FOLLOW sets you would compute for them by hand. The table is read exactly like the one you build on paper: parsing nums and seeing ')' selects production 2; a . cell is an error, and its row is the “expected …” set a ParseError reports. Point the command at any .scup file, including one you write.

Two good uses: to check your homework, type an HW 2 exercise grammar into a .scup file (no actions, no types, no Scala) and compare the dump against your hand computation; and to debug a grammar, since when a conflict report surprises you, the FOLLOW sets in the dump usually explain it.

Grammars Without Actions

That homework flow works because a bare grammar is a valid spec: a rule with no blocks and no %type still loads, LL(1)-checks, and parses, building a generic labeled tree (STree) with one node per production, single-symbol productions passing through. Zero Scala required. (In this mode any name no rule defines is treated as a terminal, so the checker prints a note for each one that looks like a rule name; a misspelled nonterminal would otherwise silently “check out” as a terminal.) This is also a fine way to draft a grammar in PA 2: get the factoring LL(1)-clean first, then add %types and blocks rule by rule.

The Dangling Else

One conflict is famous enough to have a standard resolution. In

stmt     ::= 'if' '(' expr ')' stmt elseTail | ...
elseTail ::= 'else' stmt | %empty

an else can begin elseTail but can also follow it (when the if was nested inside another if). That is a FIRST/FOLLOW conflict, and scup rejects it like any other, unless you say you mean it:

# from ic.scup: the dangling else
%greedy
elseTail ::= "else" stmt   { $2 }
           | %empty        { EmptyStmtNode() } ;

%greedy resolves a FIRST/FOLLOW conflict in favor of the production that consumes input: each else binds to the nearest unmatched if. This is yacc’s prefer-shift resolution of the same conflict, made explicit. It is the only %greedy in the whole IC grammar; anywhere else, left-factor instead.

Errors

Because parsing is predictive, error messages are exact: the engine knows which terminals have table entries at the point of failure:

ParseError: expected ')', '*', '+' or '-' at line 1, column 4
(while parsing 'expr')

parseAll/parseAs require the whole input to be consumed and throw ParseError otherwise; your compiler’s host catches it and reports a SyntaxError in IC’s format. Fixed-text terminals print as their %token spelling (expected '{') with no display-name map anywhere. There is no “furthest failure” guesswork as in backtracking parsers: the reported position is the exact token where prediction failed.

Preludes and Hosts

A %{ ... %} prelude at the top of the spec holds the imports your blocks need (your AST package, the token-kind enum), small helper defs, and parser-only types: little enums and case classes that left-factored tail rules return for the owning rule to match on. Prelude definitions live inside the spec’s compiled actions; no other phase of your compiler can see them, which is what you want for a parser-only type. Types the rest of the compiler uses (your AST) belong in ordinary project source files. The host is the short Scala object that loads the two specs and owns the entry point:

// ICParser.scala -- the host object (provided in PA 2)
object ICParser:
  private lazy val g = GrammarSpec.load("ic.scup", Lexer.spec.binding)

  def parse(source: String): ProgramNode =
    val toks = Lexer.specTokens(source)
    try ProgramNode(g.parseAs[List[ClassNode]](toks), 1)
    catch case e: ParseError => throw SyntaxError(e.expected, e.pos)

The binding carries what the grammar needs from the token spec: the enum, the %token literal table, and the display names.

Editor Support

The course VSCode extension (“slex & scup”, installed in the note at the top of this tutorial) understands .scup files: Scala highlighting inside blocks, Format Document for the ::=/| layout, outline over rules, and, with the check CLI wired in, LL(1) conflicts as squiggles on the offending rule. Go-to-definition follows a rule reference to its definition, a "literal" to its %token row in the companion .slex, a token kind through that .slex to the action constructing it (or the enum member), and $1/$$ in an action to the production symbol it names. Inside { ... } actions, go-to-definition and hover behave like ordinary Scala once the project has compiled; hovering $1 even shows the type the grammar gives it.

Under the Hood

Your spec is interpreted, not compiled to a parser: the loader builds the same rule/production structures the library’s embedded Scala DSL builds, and three short files implement everything; read them.

  1. Syntax.scala — Grammar, Rule, productions, the EBNF modifiers, precedence ladders. This is the embedded DSL the spec surface maps onto:

    lazy val factor: Rule[Int] = rule(
        NUM --> { t => t.text.toInt }
      | (!LP, expr, !RP) --> { e => e })

    is factor ::= NUM { ... } | "(" expr ")" { ... } ; written as Scala values: productions are symbol tuples, --> attaches a typed action, ! marks punctuation to drop (the spec’s blocks just ignore unmentioned $i instead), and leftAssoc/ operators(...) are %binary’s underpinnings. The DSL remains fully usable (the pattern parser inside slex is written with it, and the companion repository’s dsl/DslCalc.scala is the whole calculator written this way: sbt "runMain dsl.demo"), but the spec file is the surface the course uses. Your action blocks are compiled behind the scenes into functions over the $i values, which is why type errors in them point at your spec file’s own lines.

  2. Analysis.scala — nullable, FIRST, FOLLOW, and the table, as fixpoints over the productions; conflict and left-recursion detection; the dump. After HW 2, this file is your homework computations mechanized, Dragon 4.4 almost line for line.

  3. Parse.scala — the predictive engine: look up (rule, lookahead) in the table, run the chosen production’s symbols left to right, apply the action. That is the entire runtime.

Where You Will Use It

In Lab 2 you complete the Calculator: add operators, repair its deliberately right-associative starter grammar with a %left block and %binary, and add %list(expr, ",") programs; in HW 3 you implement the nullable/FIRST fixpoints yourself (hw/First.scala), checking your implementation against the library’s dump.

In HW 4 you write the RE compiler’s parser: a grammar for regular expressions whose atoms all start with distinct terminals, plus a let-binding form handled by a small substitution pass after parsing.

In PA 2 you write ic.scup, the IC grammar. The IC specification’s grammar appendix gives most of the LL(1) factorings; the statement rule and primary are deliberately left for you, and the conflict reports will point out where an attempt is not yet LL(1). Tokens flow straight in from ic.slex, so the front end is one line per phase:

source  --ic.slex-->  tokens  --ic.scup-->  AST