After this assignment, you will be able to:
- specify the complete token set of a real language as an ordered list of regular-expression rules, using maximal munch and rule order to resolve overlapping matches.
- write scanner actions for translating escape sequences, range-checking literals, and reporting malformed tokens with precise positions.
- test a compiler phase systematically against a specification suite of expected-output files.
Overview
In this programming assignment, you will implement the scanner for
your IC compiler. The IC language specification document is available on
the course web page. You will write the scanner as a slex spec
file: ic.slex, at the root of your repository,
declares IC’s tokens as an ordered list of regular-expression rules with
action blocks, and slex’s engine applies maximal munch and
first-rule-wins to do the rest — exactly the ideas we have studied with
regular expressions and NFAs, made concrete in one file. The slex tutorial is required reading before you start;
the rule-priority laws and the action idioms it presents are most of
this assignment.
A note on realism: we use a lex-style library here — one whose engine you can readily read and understand, as it builds on your homework questions. Production compilers such as clang and rustc hand-write their lexers, primarily for the quality of their error messages. The action code in this assignment is where you will see why.
Implementation Details
You will specify the IC token rules, build a driver program for the lexer, and write a test suite. You are required to implement the following:
enum TokenKind, in the starter’sic/lex/Token.scala. Extend the provided enumeration with a case for each kind of token in IC. Every name the spec file uses must be a member of this enum (a typo is caught, with a suggestion, when the spec loads), and the enum’s members are automatically usable in the spec’s action blocks. The lexer returnsTokenobjects (also inToken.scala), each carrying at least:kind, which kind of token was matched;text, the text of the token (for string literals this is the string’s contents, with escape sequences already processed);pos, the position (line and column) where the token occurs in the input file. Positions use 1-based line and column numbers and refer to the token’s first character; for a string literal, that is the opening quote. slex hands each rule’s action exactly this position.
The
Tokenclass extendsscup.Tokenish; that is what will let a parser consume your token stream in PA 2.ic.slex, the spec file at the root of the repository — this is the assignment. The starter spec recognizes only a token or two, but it shows all the idioms: a%tokenrow, a skip rule, and an action block. It also includes the rule that skips C-style/* ... */block comments, since that regular expression is notoriously tricky to get right. Extend the spec to all of IC://line comments, keywords, identifiers and class identifiers, integer and string literals, and every operator and punctuation symbol, following the lexical rules in the IC specification carefully. (ic/lex/Lexer.scala, provided in full, is just the ~40-line host shim that loads the spec; you should not need to modify it.) The regular expressions determine where each token begins and ends (remember maximal munch:<=is one token, not two — slex guarantees this as long as your rules are right); your action blocks must handle everything the regular expressions cannot:Keywords. A
%tokenrow for each keyword, anywhere above the identifier rule — first-rule-wins does the rest, with no keyword table anywhere. (The tutorial also shows the action-side lookup alternative.)String escapes. Only
\t,\n,\", and\\are legal escapes; the string-literal action must translate them and produce the string’s contents, and any other escape is a lexical error.Integer range. Integer literals have no leading zeros and must be in range; the action checks the value and reports out-of-range literals as errors.
Unterminated strings and comments. Use the tutorial’s low-priority catch-all idiom: after the real string-literal rule, a rule matching just
"whose action throwsLexicalError, and likewise a bare-/*rule after the provided comment rule. That is how your lexer reports what is unterminated and where, instead of a generic illegal-character error.
class Compiler. This will be the main class of your compiler at the end of the semester. At this point, this class is just a testbed for your lexer. It takes a filename as an argument, opens that file, and tokenizes its contents with your lexer. The code should print a representation of each token read from the file to the standard output, one token per line. Your output must include the following information: the token kind, the text of the token (if any), and the line number for that token. At the command line, your program must be invoked as follows:sbt "run <file.ic>"I have given you code in the
ic.Compilerclass to handle the optional command line argument “-d”, as in:sbt "run -d <file.ic>"This option will turn on debugging messages generated by calls to
ic.Util.debug(...). (See theic.UtilandCompilersource code for more details.) If “-d” is not provided, calls toic.Util.debug(...)will have no effect. You should use this printing mechanism to aid in testing and debugging your code, so that you can selectively turn on and off whatever logging you would like.The starter also includes
./icc.sh, a tiny wrapper around this command that hides sbt’s log messages:./icc.sh [-d] <file.ic>takes exactly the same arguments as the compiler and prints only your compiler’s own output.class LexicalError. Your lexer should also detect and report any lexical analysis errors it may encounter. You must use the provided exception class for lexical errors, which contains the position where the error occurred and an error message. Your action blocks report errors with the in-scopeerror(...)helper (for out-of-range literals, bad escapes, unterminated strings and comments); it and any stray character no rule matches surface as slex’sLexError, which the providednext()converts into aLexicalErrorfor you. (A block may also throwLexicalErrordirectly when it wants to adjust the reported position.) The main method must catchLexicalErrorand terminate the execution. Your program must always report the first lexical error in the file; slex scans left to right, one token pernext()call, so this comes for free if you print each token as you read it and do not catch errors too early.
Output Format
As mentioned above, you are free to print out the relevant information about tokens in any reasonable way provided that you print them out one per line with no blank lines between them. The exact content of error messages is left to you. In addition, the last line printed by your code should be either
Success.
or
Failed.
depending on whether or not any lexical errors were found. Please match those lines exactly to ensure my test scripts can properly validate your code. For similar reasons, the output should not contain any other text beyond what is specified here.
Utility Functions
I have provided a few utility methods in the ic.Util
class. You are free to (and should!) use these methods in your code. In
particular, make use of Util.debug to print diagnostic
messages for debugging. Also, use Scala’s assert function
to assert that specific conditions (e.g.: preconditions, postconditions,
invariants) are always true at run time.
Code Structure
All of the classes you write should be in or under the package
ic, containing the following:
the class
Compilercontaining the main method;the
ic.lexsub-package, containing theLexer,Token, andTokenKindclasses;the
ic.errorsub-package, containing theLexicalErrorclass.
The slex library sources are included in the starter; you can read
them but should not need to modify them. Your ic.slex sits
at the project root, next to build.sbt; the build compiles
its action blocks automatically, so type errors in a block appear (in
sbt and in VSCode) at the spec file’s own line numbers.
Testing the scanner
You must test your lexer. You should develop a thorough test suite that tests all legal tokens and as many lexical errors as you can think of. There are two types of tests that you should consider writing:
Specification Tests: These test the behavior of the whole system. In this case, that means running your compiler on different input files to ensure that it meets the requirements laid out in this handout. You should consider automating your testing as much as possible to make it easy to re-run them. The starter includes a test runner in
src/test/scala/tests/SpecTests.scalaand the course pa1 test suite intests/pa1. Runsbt "specTests tests/pa1"to compile every
.icfile in that directory and compare the result to the matching.ic.expectedfile. (Add-v, as insbt "specTests -v tests/pa1", to see the source and full output of each failing test.) Use that suite, but do not stop there: you should also write your own tests — add them alongside the provided ones intests/. I can help out with that if you’d like.Unit Tests: These tests exercise the behavior of individual components of your system in a modular way. I strongly suggest writing unit tests as you develop your code. The starter includes several example MUnit tests in
src/test/scala; run them withsbt test. The spec reloads on every test run, so you can grow it rule by rule — test your string-literal rule on every escape and every malformed string long before the full spec exists.
When a test lexes wrongly and staring at the spec does not explain
why, make the scanner narrate: tokenize(src, trace = true)
on the one failing line prints, munch by munch, which rules were alive,
who accepted, and why the winner won (the tutorial’s Watching It Scan section, and HW 2
problem 2, show you how to read it).
We will test your lexer against our own test cases.
Submission
You will develop your code, with version control, in your GitHub repository — push your final version of your code and supporting files to GitHub by the deadline, with a descriptive message such as “pa1 submission”. For grading, you will then submit your project to Gradescope by the deadline.
As in any other large program, much of the value in a compiler is in how easily it can be maintained. For this reason, a high value will be placed here on both clarity and brevity – both in documentation and code. Make sure your code structure is well-documented.
Also include a brief summary of your project in
README.md. At this, simply describe your basic code
structure and testing strategy. Also mention any known bugs and other
information that may be useful when grading your assignment.
Your project directory should be organized as follows:
CONFIG.md- a quick overview of how to build and run the project.README.md- your project write up.build.sbt- the sbt build definition. You should not need to modify this./src- all of your source code (src/main/scala) and unit tests (src/test/scala)./test- your specification test cases.
Once you have pushed what you believe to be your final version, please verify all of your code has been added and committed properly. You can do this by simply inspecting the files through the server’s web interface, but the most reliable way is to clone a new copy of your project into a temporary directory, build it, and run it in a few short examples.