HW 3: LL(1) Parsing

Learning Objectives

After this assignment, you will be able to:

Overview

This week we turn grammars into working parsers with scup, the course’s LL(1) parsing library. It builds on last week’s homework but focuses more on pragmatics and implementation. As such, there is not a lot of new theoretical material and the main tasks are related to implementing a small parser for a calculator. You’ll get an invite for a repository for that – you’ll want to get started on that as you work through the problems. On Friday we’ll go over the questions below, and the lab session will be spent on finishing up the calculator.

A parser is a .scup spec file: a yacc-flavored grammar, one rule per nonterminal, each production carrying a block of Scala to run when it applies. When the spec is loaded, the library computes nullable, FIRST, and FOLLOW (the same computations as HW 2), rejects the grammar if it is not LL(1), and parses with one token of lookahead and no backtracking. The spec file is the grammar; this is how we will build the parser for IC in PA 2.

The engine behind the specs is three short files (Syntax.scala, Analysis.scala, Parse.scala), and you should read all of them this week; every question below is easier once you understand how rules, productions, %list, %binary, and the analysis are implemented. Two design decisions matter most:

Readings

Exercises

  1. The following grammar describes the language of regular expressions:

    \[\begin{array}{rcl} R & \rightarrow & R ~\mathit{bar}~ R ~~|~~ R~ R ~~|~~ R ~\mathit{star}~ ~~|~~ ( ~R~ ) ~~|~~ \epsilon ~~|~~ \mathit{letter} \end{array}\]

    where bar, star, letter, ‘(’, and ‘)’ are all terminals. This is an ambiguous grammar. The Kleene star operation has higher precedence than concatenation; and, in turn, concatenation has higher precedence than alternation.

    1. Write an unambiguous grammar that accepts the same language, respects the desired operator precedence, and is LL(1): it must contain no left recursion, and every choice must be decidable with one token of lookahead. Use one nonterminal per precedence level, and use the \(X \rt \alpha~ \beta^{*}\) (“an \(\alpha\) followed by any number of \(\beta\)s”) style of rule where iteration is needed; that is what a postfix * implements in a spec.

      Hint

      Two of the three levels are unusual: concatenation has no operator token (one operand simply follows another), and star is postfix rather than infix. Neither is an infix chain like the calculator’s; both are the iteration shape, \(X \rt \alpha~ \beta^{*}\).

    2. Write the parse tree for the expression \(a ~|~ b ~c *~ d ~|~ e\) using your grammar.

    3. Transcribe your grammar into an action-free .scup spec (plain rules, no blocks) and check your LL(1) claim mechanically. Accept the GitHub invitation email to your Calculator repository and clone it; the rest of this assignment lives there. Put your spec at its root as, say, re.scup, and run sbt "runMain scup.check re.scup". Keep the file, since Question 2 dumps its analysis.

  2. Implementing the analysis. Your Calculator repository’s src/main/scala/hw/First.scala computes nullable and FIRST over a tiny stand-alone BNF representation: nonterminals are strings, terminals are characters, a grammar is a Map from nonterminal to productions. The fixpoint drivers are provided: each asks a per-production question over and over until the answer stops changing. Implement the two questions (each is a few lines) and check them with the provided tests, which must all pass:

    sbt "testOnly hw.FirstTests"
    1. seqNullable(syms, nullable): can a production with these symbols derive \(\epsilon\), given the nonterminals known to be nullable so far?

    2. seqFirst(syms, nullable, first): FIRST of a production with these symbols. Walk left to right, accumulating each symbol’s FIRST, and stop after the first symbol that is not nullable.

    3. scup runs these same fixpoints over your grammars. Run sbt "runMain scup.check -dump re.scup" on the spec you wrote for Question 1, and check the FIRST sets and the LL(1) table against your hand computation of both.

  3. Prediction is not backtracking. Many parsing tools handle a rule’s alternatives by trying them in order, rewinding the input when one fails, so alternatives sharing a prefix “work” by reparsing. scup takes a different approach.

    1. What happens with the spec

      s ::= "a" "b"  { ... }
          | "a" "c"  { ... } ;

      and when does it happen: when the spec is loaded, or only on some inputs?

    2. Rewrite the rule so it is LL(1) and accepts the same language.

    3. A backtracking parser given the ordered alternatives s ::= "a" | "a" "b" silently fails on the input ab: the first alternative wins, and the leftover b then makes the parse fail somewhere else entirely. Explain why this whole class of ordering bug cannot arise in an LL(1) tool.

  4. Left recursion.

    1. Transcribe the left-recursive rule

      \[\begin{array}{rcl} E & \rightarrow & E ~-~ \mathit{num} ~~|~~ \mathit{num} \end{array}\]

      directly into a .scup rule (i.e., e ::= e "-" NUM { ... } | NUM ;) and explain precisely what happens, and when. Why can’t any parser decide this rule with one token of lookahead?

    2. Rewrite the rule without left recursion using a postfix *, and give the fold that builds a left-leaning tree, so that \(9-2-3\) means \((9-2)-3\).

    3. The spec layer packages exactly this pattern as a %left block plus a %binary rule. Using them, write the two-level grammar for + - * / with the usual precedence and left associativity, over a token type of your choosing.

  5. The dangling else. Recall the classic ambiguous statement grammar:

    \[\begin{array}{rcl} S & \rightarrow & \mathbf{if}~ E ~\mathbf{then}~ S \\ & | & \mathbf{if}~ E ~\mathbf{then}~ S ~\mathbf{else}~ S\\ & | & \mathit{other} \end{array}\]

    The grammar is ambiguous (an else may attach to either if), and a predictive parser surfaces the problem as a conflict.

    1. The two if productions share the prefix \(\mathbf{if}~ E ~\mathbf{then}~ S\), so they must be left-factored:

      \[\begin{array}{rcl} S & \rightarrow & \mathbf{if}~ E ~\mathbf{then}~ S ~ S' ~~|~~ \mathit{other}\\ S' & \rightarrow & \mathbf{else}~ S ~~|~~ \epsilon \end{array}\]

      Show that \(S'\) has a FIRST/FOLLOW conflict: compute FIRST of its first production and FOLLOW(\(S'\)), and exhibit the overlap.

    2. scup accepts this rule only marked %greedy, which resolves the conflict in favor of the consuming production. To which if does the else of \(\mathbf{if}~ a~ \mathbf{then}~ \mathbf{if}~ b ~\mathbf{then}~ s_1~ \mathbf{else}~ s_2\) attach under that resolution? Answer by tracing the predictive parse: which rule is the engine in when the else arrives?

    3. Why does scup make you opt in with %greedy, rather than resolving every FIRST/FOLLOW conflict this way by default?

The programming companion to this homework is Lab 2: completing the calculator’s grammar in the same repository. Your hw/First.scala implementation from Question 2 is submitted together with the lab.