HW 4: Abstract Syntax Trees and Symbol Tables

Learning Objectives

After this assignment, you will be able to:

Overview

This week’s goal is to explore the internal data structures used by a compiler to organize and manipulate a source program. These data structures are the core of a compiler, and developing a good set of programming abstractions is essential for managing implementation complexity.

There are several specific topics for the week:

  1. Building abstract syntax trees (ASTs) during parsing, and examining other intermediate representations.

  2. Developing a general structure for symbols information.

Readings

The readings are from a variety of sources, listed roughly in the order in which they will be most useful. Some parts are pretty light and fill in background material; others are more directly related to the problems below.

Exercises

  1. This question requires some programming. You may work with a partner if you wish. The problem explores how parsers construct values — and, in particular, abstract syntax trees — as they go. In scup every rule produces a value: %type expr: Int says rule expr matches tokens and its blocks build an Int. Until now our calculator evaluated on the fly:

      %left "+" "-"
      expr ::= %binary(term)  { if $2.text == "+" then $1 + $3 else $1 - $3 } ;

    Modern compilers do essentially no real work during parsing. Instead, the parser builds an abstract syntax tree, and all subsequent analysis is performed by separate passes over the tree. For the calculator, that means AST classes like:

      abstract class Expr
      case class Add(left: Expr, right: Expr) extends Expr
      case class Number(value: Int) extends Expr

    and a grammar of %type expr: Expr whose operator blocks produce tree constructors instead of arithmetic:

      expr ::= %binary(term)  { if $2.text == "+" then Add($1, $3) else Sub($1, $3) } ;

    Starting from your Calculator from Lab 2:

    1. Define AST case classes for the calculator language and change the grammar’s %types to Expr, so that parsing builds a tree and performs no arithmetic. Then add a separate evalExpr(e: Expr): Int function to Calc.scala as a pattern-matching tree traversal (see pattern matching on case classes in the Scala tour), and change eval(source) to parse the source to an Expr and evaluate it with that function.

    2. Your Lab 2 calculator already parses comma-separated programs with program ::= %list(expr, ","). Change program’s %type to List[Expr], so that “4,1+7” parses to a list of trees. (Tip: pass trace = true to parseAs to watch the parser work through an input and check that it matches your intuition.)

    3. Update evalList in Calc.scala to parse a program to a List[Expr] and evaluate each tree with evalExpr, and check that all your Lab 2 tests still pass.

    No need to bring this code to our meetings, but be prepared to talk about the above items.

  2.  [Adapted from Cooper and Torczon]

    • Show how the code fragment

        if (c[i] != 0) {
          a[i] = b[i] / c[i];
        } else {
          a[i] = b[i];
        }
        println(a[i]);

      might be represented in an abstract syntax tree, in a control flow graph, and in quadruples (or three-address code — see the simple TAC definition).

    • Discuss the advantages of each representation.

    • For what applications would one representation be preferable to the others?

  3. This question also involves programming. It is more substantial than most of the homework questions, so please don’t wait until the last minute. You may work with a partner if you like. The following grammar describes the language of regular expressions, with several unusual characteristics described below:

    \[\begin{array}{rcl} R & \rightarrow & R ~ \tm{|} ~ R \\ & | & R ~ \tm{.} ~ R \\ & | & R ~ \tm{*} \\ & | & R ~ \tm{?} \\ & | & R ~ \tm{+} \\ & | & \tm{(} ~ R ~ \tm{)} \\ & | & \nt{letter} \\ & | & \tm{[} ~ L ~ \tm{]} \\ & | & \tm{let} ~ \nt{id} ~ \tm{=} ~ R ~ \tm{in} ~ R \\ & | & \nt{id} \\ & | & \tm{@} \\[1ex] L & \rightarrow & L ~ \nt{letter} \\ & | & \nt{letter} \end{array}\]

    The *, ?, and + operators have higher precedence than concatenation (.); and, in turn, concatenation has higher precedence than alternation. A letter can be any lower-case letter in a–z, and @ stands for \(\epsilon\). The term [\(L\)] indicates an alternation between all of the letters in the letter list \(L\). Here are some examples:

    • a.b+: a followed by one or more b’s.

    • [abc]: any of a, b, or c

    • a.(b|c)*: a followed by any number of b’s and c’s.

    • a|@: either a or \(\epsilon\), which is represented by @.

    In order to describe more interesting patterns succinctly, our language has “let”-bindings for id’s (which are identifiers starting with capital letters), as in the following:

    let C = (c.o.w)* in
      C.m.C.m.C

    which is the same as (c.o.w)*.m.(c.o.w)*.m.(c.o.w)*. Bindings can be nested, and one binding can redefine an already bound name:

    let C = c.o.w in
      C.C.
      let C = m.o.o in
        C*

    which is equivalent to c.o.w.c.o.w.(m.o.o)*.

    The starter code for this problem includes a complete lexer (re.lex.RELexer) and a skeletal scup grammar (re.parser.REParser) for regular expressions. You are to complete the parser, design an AST package for regular expressions, and then write code to translate regular expressions into NFA descriptions that can be executed on the NFA simulator included in the starter (package nfa). The simulator uses the same set-of-states algorithm you wrote in Lab 1, but its .nfa file format is different (edges are labeled with any of a–z) — don’t confuse the two.

    To begin, accept the GitHub invitation email to your team’s HW 4 “RECompiler” repository and clone it. It is set up like the project for the first question and PA 2. You are free to again work with a partner or two on this question.

    1. The main method in re.Main currently reads a regular expression from its first argument. It is executed from the command line as follows:

        sbt "run ex1.re"

      Before you can successfully parse expressions, however, you must complete the grammar in re.scup. Structure it with one rule per precedence level, as you did for a simpler version of this language in HW 3.

    2. Design a hierarchy of case classes to represent regular expression ASTs. The root of your hierarchy should be the re.ast.RENode class that I have provided. Your hierarchy should contain a reasonable, minimal collection of classes. Not every concrete syntactic form needs to have an analog in the abstract syntax (eg, “[L]” can be expressed as an alternation, etc.).

    3. Extend the parser to generate an AST for the parsed regular expression.

      Your AST classes should use Scala collections — List, Option, and friends — to store their data. Options are useful for storing “optional” values; I encourage you to use them where appropriate.

      For let-bindings, remember that grammar rules are static, so you cannot thread an environment through the parse. Instead, have your blocks build a small intermediate tree of your own that still contains lets and names, and then substitute definitions for names in a second “resolve” pass in re.parser.REParser, so bindings never show up in your final AST (see the notes at the top of re.scup). The resolve pass maintains a map from id’s to their definitions and should throw a re.error.REError if it encounters an id that has not been defined.

      You may use either static or dynamic scoping for names. The two differ only when a definition mentions another name that is later rebound. Consider:

      let A = a in
        let B = A in
          (let A = c in
            B
          ).B
      • Static scoping: a name inside a definition means the binding in scope where the definition is written. When B is defined, A means a, so B means a everywhere, and the whole expression is a.a. To implement this, resolve each definition before adding it to the map.

      • Dynamic scoping: a name inside a definition means the binding in scope where the definition is used. The first use of B occurs inside let A = c, so there B means c; the second use of B occurs outside it, where B means a. The whole expression is c.a. To implement this, store each definition in the map unresolved, and resolve it when you substitute it for a name.

      Either choice is fine; the provided tests do not distinguish them. You may assume that definitions are never circular (for example, let X = X.a in X will not appear), so the dynamic approach always terminates.

    4. Complete the re.PrettyPrint class for generating printable, fully parenthesized forms of expressions represented by a RENode (for example, a.b* prints as (a.(b*))), and extend the main method to print the parsed expression.

    5. Write a re.NFABuilder class to build a NFA for a regular expression represented by a RENode. I have provided the re.NFA class to help in this step — your builder simply needs to create a new NFA object and invoke the appropriate methods on it to create states and edges. See the classes documentation for details on the NFA class. This process will be a recursive traversal of the AST. Think carefully about what information would be most useful to propagate down the tree and back up during the traversal.

    6. Use the NFA.toString() to write the resulting NFA to a file. Specifically, when re.Main is run on an input file x.re, it must write exactly one file named x.nfa (that is, the input’s base name with a .nfa extension) into the current directory — the provided tests require this. The starter’s Main.writeNfaFile helper handles that naming; just call it with your NFA. So, running on ex1.re produces ex1.nfa, which can then be run with the NFA simulator:

      sbt "runMain nfa.NFASimulator ex1.nfa ab cow"

      The alphabet for this simulator is a–z, plus @ for \(\epsilon\).

    If ex1.re contains the expression “a.(b|c)*”, your program should generate an NFA similar to the following:

    7
    6
    0 a:(1) ; 
    1 @:(2,3,6) ; 
    2 b:(4) ; 
    3 c:(5) ; 
    4 @:(6) ; 
    5 @:(6) ; 
    6 @:(1) ; 

    Running

      sbt "runMain nfa.NFASimulator ex1.nfa ab cow abccccc a"

    will then result in

    ab: yes
    cow: no
    abccccc: yes
    a: yes

    sbt test runs the full pipeline the same way: for each tests/X.re, re.PipelineTests runs re.Main to produce X.nfa and checks the simulator’s answers against tests/X.re.expected. Be sure to test your program on more sophisticated examples, too — add your own .re/.re.expected pairs to tests/.

    To help you debug the last two steps, the NFA class also contains a printDot() method that generates the file “nfa.dot”. This graphical representation of the NFA can be viewed by issuing the following command from the command line to generate a PDF file:

    dot -Tpdf < nfa.dot > nfa.pdf

    The dot command is part of Graphviz. To install it on your laptop, use brew install graphviz on a Mac or sudo apt install graphviz on Ubuntu/WSL.

    Please come to your meetings with your final code committed to your GitHub repository. Time permitting, I’d like to do mini-code reviews of your solutions.