Lab 1: NFA Simulation

Due: 10pm, Wednesday, Sep. 23

Learning Objectives

After this lab, you will be able to:

Overview

Russ Cox describes how inefficient some regular expression pattern matchers can be. In this lab you will build a pattern matcher that can outperform Java’s regular expression package. More specifically, you will write a program that, given a description of an NFA, simulates its behavior on an input string.

Write your answers to the written questions (1, 2, 4, and 5) in the WRITEUP.md file at the top of your repository. Submit your finished simulator and writeup via Gradescope.

You may work on this lab with a partner if you like. Just use one of your github repositories and add the other as a member, and also in Gradescope when you submit the repository.

Questions

  1. To motivate the design and to see how a little bit of theory can dramatically improve software design, we start by comparing the two simulation algorithms in section “Regular Expression Search Algorithms” section of Cox’s article.

    Here is the NFA corresponding to “a?a?a?aaa”:

    Enumerate all paths taken through the NFA using the backtracking search algorithm on input aaab. How many are there? Given the NFA for \((\texttt{a?})^n\texttt{a}^n\), how many paths may need to be explored to test an input string of length \(n+1\)? (A Big-O bound is fine, and a one or two sentence explanation is sufficient.)

  2. Now, show the steps performed by Algorithm 3.22 in Dragon (page 156). (This is a more precise and succinct description of the second algorithm Cox describes). It is sufficient to show the states in \(S\) each time line (3) is executed. You will find it useful to compute the \(\epsilon\)-closure for each state in the NFA.

  3. Implement Algorithm 3.22. Start with the NFA simulator starter: accept the GitHub invitation email to your Lab 1 repository in the course organization and clone it. The starter parses the input file for you (NFAReader.scala) and provides the main program (NFASimulator.scala); your entire job is the two methods in NFA.scala — epsClose and accepts. The starter’s unit tests fail until both work, so sbt test tells you when you are done. You must write your simulator in Scala, within that sbt project: the Gradescope autograder builds the project with sbt and requires the main class to be nfa.NFASimulator.

    The input to your program will be the transition table for an NFA over the alphabet \(\{a,b\}\). The format of the input is:

    • a number \(n\) indicating the number of states (which will be labeled \(0\), \(1\), ..., \(n-1\));

    • a number \(q \in 0..n-1\) indicating the only accepting state of the NFA. (Assume 0 is the start state).

    • the transitions for each state on a, b, and \(\epsilon\).

    As an example, the table for the following NFA

    would be

    4
    3
    (0,1) (0)   ()
    ()    (2)   ()
    ()    (3)   (1)
    ()    ()    (0)

    Several sample input files (t0.nfa, t1.nfa, t2.nfa) are included in the starter project. Your program should output “yes” or “no” for each string it tests. The name of the file containing the input NFA and the test strings should be passed to your program as command-line arguments, as in:

    sbt "run t0.nfa aabb abab cow"

    NOTE: I am far more interested in clarity and correctness than efficiency. The Dragon book describes a fairly low-level, efficient implementation, but you are not expected to follow that approach. Just use standard scala.collection (or similar) data structures — Stacks, Lists, Vectors, Sets, arrays, etc. — to implement a reasonable, straight-forward solution.

  4. The java.util.regex package contains a Pattern class that uses the first algorithm. You can test a string against a regular expression using this class as follows:

    java.util.regex.Pattern.matches("a?a?a?aaa", "aaaa")

    Compare the performance of your program to a program that uses this method. For convenience, the data files part-4/e\(n\).nfa in the starter project (part-4/e0.nfa ... part-4/e40.nfa) contain the NFAs for \((\texttt{a?})^n\texttt{a}^n\). (The provided example is both legal Java code and legal Scala code.)

    The Scala read-eval-print loop is a convenient way to run these comparisons: sbt console, run from your project directory, starts a REPL with your simulator on the classpath, so both approaches are one line each. Scala’s * repeats a string, which saves you from typing out the regular expressions, and a little helper times anything you give it:

    $ sbt console
    scala> def time[A](op: => A): A = { 
        val t0 = System.nanoTime(); 
        val r = op; 
        println(s"${(System.nanoTime() - t0) / 1000000} ms"); 
        r 
    }
    
    scala> time(java.util.regex.Pattern.matches("a?" * 25 + "a" * 25, "a" * 25))
    904 ms
    val res0: Boolean = true
    
    scala> val e25 = nfa.NFAReader.read(new java.io.File("part-4/e25.nfa"))
    scala> time(e25.accepts("a" * 25))
    1 ms
    val res1: Boolean = true

    (Quit with :quit or Ctrl-D.) Record your measurements in WRITEUP.md — a small table of \(n\) against the two times is ideal — along with a sentence on what you observed.

  5. HW 1’s last exercise explored converting an NFA to a DFA before simulation. Why is this preferable? What are the downsides to DFA conversion? Do you think the potential issues impact lexical analysis for programming languages substantially?