HW 2: Lexical Analyzers, Grammars, and Top-Down Parsing

Learning Objectives

After this assignment, you will be able to:

Overview

Readings

Exercises

  1. Minimize the states of the following DFA. Label each state in the minimized DFA with the set of states from the original DFA to which it corresponds.

  2. Consider the following lexical analysis specification:

    (aba)+   { return Tok1; } 
    (a(b)+a) { return Tok2; } 
    (a|b)    { return Tok3; } 

    In case of tokens with the same length, the token whose pattern occurs first in the above list is returned.

    1. Build an NFA that accepts strings matching one of the above three patterns using Thompson’s construction.

    2. Transform your NFA into a DFA. Label the DFA states with the set of NFA states to which they correspond. Indicate the final states in the DFA and label each of these states with the (unique) token being returned in that state.

    3. Show the steps in the functioning of the lexer for the input string abaabbaba. Indicate what tokens does the lexer return for successive calls to getToken(). For each of these calls indicate the DFA states being traversed in the automaton.

    4. Check your answers against slex, which does exactly what you just did by hand: Thompson’s construction, then the subset construction (each DFA state labeled with its set of NFA states), then a scan of the DFA with maximal munch and the first-listed rule winning ties. Your clone of the slex tutorial repository contains this problem’s specification as hw2.slex. To compare machines with your parts 1 and 2 (state numbers will differ; shapes should not), run

      sbt "runMain slex.check -dump hw2.slex"

      and to watch part 3’s scan, munch by munch, run

      sbt "runMain hw.lex -trace abaabbaba"

      and compare the output to your hand simulation: the alive rule sets after each character correspond to your DFA states, each accept: line marks a final state (and shows who wins a tie), and each backing up line is an excursion past the last accept that maximal munch gives back. Nothing to hand in for this part, but remember that the trace works on every spec (tokenize(src, trace = true)), including the ic.slex you will write for PA 1.

  3. Dragon 4.2.1

  4. Dragon 4.2.3 (a) — (c) and (d).

  5. Dragon 4.3.1

  6. Consider the following grammar:

    \[\begin{array}{rcl} S & \rightarrow & a ~ S ~ b ~ S ~~|~~ b ~ S ~ a ~ S ~~|~~ \epsilon \end{array}\]

    1. Show that the grammar is ambiguous by constructing two different rightmost derivations for some string.

    2. Construct the corresponding parse trees for this string.

  7. Consider the following grammar:

    \[\begin{array}{rcl} S & \rightarrow & B ~ C ~ z \\ % A & \rightarrow & v ~|~ \epsilon\\ B & \rightarrow & x ~ B ~|~ D \\ % C & \rightarrow & u ~ A ~|~ \epsilon \\ C & \rightarrow & u ~ v ~|~ u \\ D & \rightarrow & y ~ D ~|~ \epsilon \\ \end{array}\]

    1. Is this grammar LL(1)? If it is not, explain why, and then modify the grammar to be LL(1) before proceeding.

    2. Compute the FIRST and FOLLOW sets for the (possibly modified) grammar.

    3. Construct the LL(1) parsing table.

      NOTE: There is a typo in the book in the description of how to construct the parsing table. On page 224, step 1 of Algorithm 4.31 should refer to \(\text{FIRST}(\alpha)\), and not \(\text{FIRST}(A)\).

    4. Show the steps taken to parse xxyuz with your table. (Use Fig. 4.21 as an example of how to show the parser’s progress.)

  8. Consider the following grammar for statements:

    \[\begin{array}{rcl} \Stmt & \rightarrow & \t{if} ~ \t{E} ~\t{then} ~ \Stmt ~ \StmtTail \\ & | & \t{while} ~ \t{E} ~ \Stmt \\ & | & \t{\{} ~ \List ~ \t{\}} \\ & | & \t{S} \\ ~\\ \StmtTail & \rightarrow & \t{else} ~ \Stmt \\ & | & \epsilon \\~\\ \List & \rightarrow & \Stmt ~ \ListTail \\ ~\\ \ListTail & \rightarrow & \t{;} ~ \List \\ & | & \epsilon \end{array}\]

    In this grammar, semicolons separate consecutive statements, similar to how commas separate entries in a list. You can assume E and S are terminals that represent other expression and statement forms that we do not currently care about. If we resolve the typical conflict regarding expansion of the optional else part of an if statement by preferring to consume an else from the input whenever we see one, we can build a predictive parser for this grammar.

    1. Build the LL(1) predictive parser table for this grammar.

    2. Using Figure 4.21 in the Dragon book as a model, show the steps taken by your parser on input

          if E then S else while E { S }
    3. Use the techniques outlined in Dragon 4.4.5 to add error-correcting rules to your table.

    4. Describe the behavior of your parser on the following two inputs:

      • if E then S ; if E then S }

      • while E { S ; if E S ; }

Note

Bottom Up Parsing. There is a whole second family of parsing techniques we largely skip: bottom-up (LR/LALR) parsing, which powered earlier tools like yacc and bison and dominated compiler construction for decades. However, they have given way to other techniques and tools in practice much of the time and we’ll focus on the foundations of those parsers instead.