After this assignment, you will be able to:
- build a lexical analyzer from a list of token patterns via Thompson’s construction, the subset construction, and DFA minimization, and trace its longest-match behavior on an input.
- show that a grammar is ambiguous, and transform grammars by left-factoring and eliminating left recursion.
- compute FIRST and FOLLOW sets, build an LL(1) parsing table, and trace a predictive parser on an input, including panic-mode error recovery.
Overview
We’ll finish our exploration of Lexical Analysis this week by 1) looking more deeply at the automatic construction of a lexical analyzer from a specification of token patterns (problems 1 and 2), and 2) implementing the lexical analyzer for IC. The second part is described in a separate handout on PA 1.
We’ll also begin to look at the next phase of the compiler: Syntax Analysis. The reading on this topic covers grammars and the first of two general parsing techniques. The problems explore properties of grammars that make them suitable for describing programming languages and automatic parsing.
Readings
- Dragon 3.8, 3.9.6
- Dragon 4.1-4.4
- The slex tutorial – lex-style scanning: declare tokens as regular expressions.
- The slex source code – the engine is HW 2’s Thompson construction, HW 1’s subset construction, and a simple DFA simulation.
Exercises
Minimize the states of the following DFA. Label each state in the minimized DFA with the set of states from the original DFA to which it corresponds.
Consider the following lexical analysis specification:
(aba)+ { return Tok1; } (a(b)+a) { return Tok2; } (a|b) { return Tok3; }In case of tokens with the same length, the token whose pattern occurs first in the above list is returned.
Build an NFA that accepts strings matching one of the above three patterns using Thompson’s construction.
Transform your NFA into a DFA. Label the DFA states with the set of NFA states to which they correspond. Indicate the final states in the DFA and label each of these states with the (unique) token being returned in that state.
Show the steps in the functioning of the lexer for the input string
abaabbaba. Indicate what tokens does the lexer return for successive calls togetToken(). For each of these calls indicate the DFA states being traversed in the automaton.Check your answers against slex, which does exactly what you just did by hand: Thompson’s construction, then the subset construction (each DFA state labeled with its set of NFA states), then a scan of the DFA with maximal munch and the first-listed rule winning ties. Your clone of the slex tutorial repository contains this problem’s specification as
hw2.slex. To compare machines with your parts 1 and 2 (state numbers will differ; shapes should not), runsbt "runMain slex.check -dump hw2.slex"and to watch part 3’s scan, munch by munch, run
sbt "runMain hw.lex -trace abaabbaba"and compare the output to your hand simulation: the
aliverule sets after each character correspond to your DFA states, eachaccept:line marks a final state (and shows who wins a tie), and eachbacking upline is an excursion past the last accept that maximal munch gives back. Nothing to hand in for this part, but remember that the trace works on every spec (tokenize(src, trace = true)), including theic.slexyou will write for PA 1.
Dragon 4.2.1
Dragon 4.2.3 (a) — (c) and (d).
Dragon 4.3.1
Consider the following grammar:
\[\begin{array}{rcl} S & \rightarrow & a ~ S ~ b ~ S ~~|~~ b ~ S ~ a ~ S ~~|~~ \epsilon \end{array}\]
Show that the grammar is ambiguous by constructing two different rightmost derivations for some string.
Construct the corresponding parse trees for this string.
Consider the following grammar:
\[\begin{array}{rcl} S & \rightarrow & B ~ C ~ z \\ % A & \rightarrow & v ~|~ \epsilon\\ B & \rightarrow & x ~ B ~|~ D \\ % C & \rightarrow & u ~ A ~|~ \epsilon \\ C & \rightarrow & u ~ v ~|~ u \\ D & \rightarrow & y ~ D ~|~ \epsilon \\ \end{array}\]
Is this grammar LL(1)? If it is not, explain why, and then modify the grammar to be LL(1) before proceeding.
Compute the FIRST and FOLLOW sets for the (possibly modified) grammar.
Construct the LL(1) parsing table.
NOTE: There is a typo in the book in the description of how to construct the parsing table. On page 224, step 1 of Algorithm 4.31 should refer to \(\text{FIRST}(\alpha)\), and not \(\text{FIRST}(A)\).
Show the steps taken to parse
xxyuzwith your table. (Use Fig. 4.21 as an example of how to show the parser’s progress.)
Consider the following grammar for statements:
\[\begin{array}{rcl} \Stmt & \rightarrow & \t{if} ~ \t{E} ~\t{then} ~ \Stmt ~ \StmtTail \\ & | & \t{while} ~ \t{E} ~ \Stmt \\ & | & \t{\{} ~ \List ~ \t{\}} \\ & | & \t{S} \\ ~\\ \StmtTail & \rightarrow & \t{else} ~ \Stmt \\ & | & \epsilon \\~\\ \List & \rightarrow & \Stmt ~ \ListTail \\ ~\\ \ListTail & \rightarrow & \t{;} ~ \List \\ & | & \epsilon \end{array}\]
In this grammar, semicolons separate consecutive statements, similar to how commas separate entries in a list. You can assume
EandSare terminals that represent other expression and statement forms that we do not currently care about. If we resolve the typical conflict regarding expansion of the optionalelsepart of anifstatement by preferring to consume anelsefrom the input whenever we see one, we can build a predictive parser for this grammar.Build the LL(1) predictive parser table for this grammar.
Using Figure 4.21 in the Dragon book as a model, show the steps taken by your parser on input
if E then S else while E { S }Use the techniques outlined in Dragon 4.4.5 to add error-correcting rules to your table.
Describe the behavior of your parser on the following two inputs:
if E then S ; if E then S }while E { S ; if E S ; }
Bottom Up Parsing. There is a whole second family of parsing techniques we largely skip: bottom-up (LR/LALR) parsing, which powered earlier tools like yacc and bison and dominated compiler construction for decades. However, they have given way to other techniques and tools in practice much of the time and we’ll focus on the foundations of those parsers instead.