After this assignment, you will be able to:
- rewrite an ambiguous, left-recursive grammar into LL(1) form, with one nonterminal per precedence level.
- implement the nullable and FIRST fixpoint computations that an LL(1) parser generator runs over your grammar.
- explain how a predictive parser chooses productions, and diagnose and repair FIRST/FIRST and FIRST/FOLLOW conflicts.
Overview
This week we turn grammars into working parsers with
scup, the course’s LL(1) parsing library. It builds on last
week’s homework but focuses more on pragmatics and implementation. As
such, there is not a lot of new theoretical material and the main tasks
are related to implementing a small parser for a calculator. You’ll get
an invite for a repository for that – you’ll want to get started on that
as you work through the problems. On Friday we’ll go over the questions
below, and the lab session will be spent on finishing up the
calculator.
A parser is a .scup spec file: a yacc-flavored grammar,
one rule per nonterminal, each production carrying a block of Scala to
run when it applies. When the spec is loaded, the library computes
nullable, FIRST, and FOLLOW (the same computations as HW 2), rejects the
grammar if it is not LL(1), and parses with one token of lookahead and
no backtracking. The spec file is the grammar; this is how we will build
the parser for IC in PA 2.
The engine behind the specs is three short files
(Syntax.scala, Analysis.scala,
Parse.scala), and you should read all of them this
week; every question below is easier once you understand how rules,
productions, %list, %binary, and the analysis
are implemented. Two design decisions matter most:
Conflicts are rejected up front: a rule’s alternatives are not tried in order, so one token of lookahead must select a unique production. If two productions could begin with the same token, the spec is refused at load, with a message naming the line and the productions, and the fix is the left-factoring you practiced in HW 2.
Left recursion is rejected too: a rule like \(E \rt E + T\) can never decide anything with one token of lookahead. The library detects the cycle and refuses; a
%leftblock plus a%binaryrule packages the standard rewrite (iterate, then fold leftward).
Readings
- The scup tutorial – the LL(1) grammar DSL
and worked examples with the library.
Start here. - Dragon 4.4 (predictive parsing, FIRST and FOLLOW, LL(1) grammars) – the theory the library mechanizes; HW 2’s computations become running code this week.
- The scup source code – read all of it; it is short: Syntax.scala, Analysis.scala, Parse.scala.
Also included in the Calculator starter.
Exercises
The following grammar describes the language of regular expressions:
\[\begin{array}{rcl} R & \rightarrow & R ~\mathit{bar}~ R ~~|~~ R~ R ~~|~~ R ~\mathit{star}~ ~~|~~ ( ~R~ ) ~~|~~ \epsilon ~~|~~ \mathit{letter} \end{array}\]
where bar, star, letter, ‘(’, and ‘)’ are all terminals. This is an ambiguous grammar. The Kleene star operation has higher precedence than concatenation; and, in turn, concatenation has higher precedence than alternation.
Write an unambiguous grammar that accepts the same language, respects the desired operator precedence, and is LL(1): it must contain no left recursion, and every choice must be decidable with one token of lookahead. Use one nonterminal per precedence level, and use the \(X \rt \alpha~ \beta^{*}\) (“an \(\alpha\) followed by any number of \(\beta\)s”) style of rule where iteration is needed; that is what a postfix
*implements in a spec.HintTwo of the three levels are unusual: concatenation has no operator token (one operand simply follows another), and star is postfix rather than infix. Neither is an infix chain like the calculator’s; both are the iteration shape, \(X \rt \alpha~ \beta^{*}\).
Write the parse tree for the expression \(a ~|~ b ~c *~ d ~|~ e\) using your grammar.
Transcribe your grammar into an action-free
.scupspec (plain rules, no blocks) and check your LL(1) claim mechanically. Accept the GitHub invitation email to your Calculator repository and clone it; the rest of this assignment lives there. Put your spec at its root as, say,re.scup, and runsbt "runMain scup.check re.scup". Keep the file, since Question 2 dumps its analysis.
Implementing the analysis. Your Calculator repository’s
src/main/scala/hw/First.scalacomputes nullable and FIRST over a tiny stand-alone BNF representation: nonterminals are strings, terminals are characters, a grammar is aMapfrom nonterminal to productions. The fixpoint drivers are provided: each asks a per-production question over and over until the answer stops changing. Implement the two questions (each is a few lines) and check them with the provided tests, which must all pass:sbt "testOnly hw.FirstTests"seqNullable(syms, nullable): can a production with these symbols derive \(\epsilon\), given the nonterminals known to be nullable so far?seqFirst(syms, nullable, first): FIRST of a production with these symbols. Walk left to right, accumulating each symbol’s FIRST, and stop after the first symbol that is not nullable.scup runs these same fixpoints over your grammars. Run
sbt "runMain scup.check -dump re.scup"on the spec you wrote for Question 1, and check the FIRST sets and the LL(1) table against your hand computation of both.
Prediction is not backtracking. Many parsing tools handle a rule’s alternatives by trying them in order, rewinding the input when one fails, so alternatives sharing a prefix “work” by reparsing. scup takes a different approach.
What happens with the spec
s ::= "a" "b" { ... } | "a" "c" { ... } ;and when does it happen: when the spec is loaded, or only on some inputs?
Rewrite the rule so it is LL(1) and accepts the same language.
A backtracking parser given the ordered alternatives
s ::= "a" | "a" "b"silently fails on the inputab: the first alternative wins, and the leftoverbthen makes the parse fail somewhere else entirely. Explain why this whole class of ordering bug cannot arise in an LL(1) tool.
Left recursion.
Transcribe the left-recursive rule
\[\begin{array}{rcl} E & \rightarrow & E ~-~ \mathit{num} ~~|~~ \mathit{num} \end{array}\]
directly into a
.scuprule (i.e.,e ::= e "-" NUM { ... } | NUM ;) and explain precisely what happens, and when. Why can’t any parser decide this rule with one token of lookahead?Rewrite the rule without left recursion using a postfix
*, and give the fold that builds a left-leaning tree, so that \(9-2-3\) means \((9-2)-3\).The spec layer packages exactly this pattern as a
%leftblock plus a%binaryrule. Using them, write the two-level grammar for+ - * /with the usual precedence and left associativity, over a token type of your choosing.
The dangling else. Recall the classic ambiguous statement grammar:
\[\begin{array}{rcl} S & \rightarrow & \mathbf{if}~ E ~\mathbf{then}~ S \\ & | & \mathbf{if}~ E ~\mathbf{then}~ S ~\mathbf{else}~ S\\ & | & \mathit{other} \end{array}\]
The grammar is ambiguous (an
elsemay attach to eitherif), and a predictive parser surfaces the problem as a conflict.The two
ifproductions share the prefix \(\mathbf{if}~ E ~\mathbf{then}~ S\), so they must be left-factored:\[\begin{array}{rcl} S & \rightarrow & \mathbf{if}~ E ~\mathbf{then}~ S ~ S' ~~|~~ \mathit{other}\\ S' & \rightarrow & \mathbf{else}~ S ~~|~~ \epsilon \end{array}\]
Show that \(S'\) has a FIRST/FOLLOW conflict: compute FIRST of its first production and FOLLOW(\(S'\)), and exhibit the overlap.
scup accepts this rule only marked
%greedy, which resolves the conflict in favor of the consuming production. To whichifdoes theelseof \(\mathbf{if}~ a~ \mathbf{then}~ \mathbf{if}~ b ~\mathbf{then}~ s_1~ \mathbf{else}~ s_2\) attach under that resolution? Answer by tracing the predictive parse: which rule is the engine in when theelsearrives?Why does scup make you opt in with
%greedy, rather than resolving every FIRST/FOLLOW conflict this way by default?
The programming companion to this homework is Lab
2: completing the calculator’s grammar in the same repository. Your
hw/First.scala implementation from Question 2 is submitted
together with the lab.