Grades and deadlines
-
Initial deadline: 23:59 on Wednesday 23 September
-
How to submit: Turn in a completed
lab2.1.ymlfile, as well as youraichat.mdif needed, to the CS department submit system underSI413/lab2.1I recommend using the
clubtool to submit, by runningclub -cSI413 -plab2.1 lab2.1.yml aichat.md -
Collaboration and AI: Recall the course policy on collaboration and AI usage. For this lab:
-
You can collaborate with your classmates as long as that is cited and you do your own work. (You are all designing different languages.)
-
You may use the class Gemini Gem as long as you cite where/how you used it and you turn in a complete transcript of all relevant conversations.
For the gem, switch to the “Pro” mode and write “transcript to produce a complete block of the conversation for you to copy/paste.
-
You may not use other AI tools besides the course-specific Gemini Gem. If you think there is something useful that you aren’t able to do with the Gem, talk to your instructor!
-
-
Grading:
In this lab you will complete the following tasks:
- Design a programming language that handles basic string operations as before, plus boolean literals, boolean operations, and (unscoped) variables.
- Write an example program and test cases demonstrating how your language works.
- Write a token specification list for your language.
- Write a context-free grammar using those token types to formally specify your langauge’s syntax.
- Describe the new semantics of your language
- Participate in peer review as both a reviewer and a reviewee
If your submission meets all requirements, you will receive 7 points towards your total lab grade.
-
Resubmissions:
We will follow the same resubmission policy for all labs this semester. For any given deadline you must at leat demonstrate significant progress or you get a zero that cannot be revised. The initial deadline for this lab is the same regardless of any other pending resubmissions on previous labs. If your work is not yet up to the standard, you will receive a new deadline of one week later (up to the end of the semester) and a chance to revise for full credit.
Downloads
YAML file to complete
Download the file lab2.1 and fill it in as you complete the tasks for the lab.
ptree tool
We made a Java program called ptree which you can use to check your
token and grammar specs. Here is the github page
You do not need to compile and build this from source — just download the release jar directly. You can copy-paste this command to just download it into your current folder from the command line:
wget https://github.com/si413usna/ptree/releases/download/v1.0/ptree.jar
To test out your token specs, put those into a file tokenSpec.txt, and
some source code in your language into example.prog (or whatever you
want to call these files), and then run
java -jar ptree.jar tokenSpec.txt example.prog
To test your grammar as well, put that into MyLang.g4 (or whatever you
want to call that file) and run
java -jar ptree.jar tokenSpec.txt MyLang.g4 example.prog
You should definitely do this to test your own tokens and grammar! We will use exactly the same tools in Java, with the same format of the token and grammar spec files, for the rest of the semester’s interpreters and compilers. So it’s worth your effort to get comfortable with this!
Task 1: Language Design
Ground rules
As before, source code should be plaintext using ASCII printable characters, space, and newline.
Your language may build off one or more of the languages from the previous unit, but it doesn’t need to do so. In particular, you can change anything about how those languages work if you want to, or you can start with something completely different. In either case, you need to include the complete specs for your new language in what you turn in for this lab.
Your language must be original and unique. It must have its own name and should not look and/or work quite the same as any other language.
Required capabilities
Your language needs to support all of the things from the previous unit:
- String literals
- String concatenation
- String reversal
- String input
- String printing
- Source code comments
Additionally, your language must also support:
-
Boolean literals: Some way to specify a literal “true” or “false” value in source code
-
Boolean logic operations: Logical OR and NOT operations. (You don’t need a separate operator for AND if you have these two…)
-
String less-than: Testing whether one string comes before another one in lexicographic order, returning a boolean true/false
-
Boolean to string conversion: some kind of operator that turns a boolean variable into a string. You can decide how true and false values should be converted, just as long as they aren’t the same thing!
-
String chop: This is a new operation on strings. Given two strings a and b, the chop operation will return a substring of a, starting at the beginning, and going up to (but not including) the first occurrence of b, if any.
So for example, chopping
battleaxewithaxewould producebattle; choppingone,two,threewith,would produceone, and choppingyingwithyangwould produceying. -
Variables: Let the programmer assign a name to the result of evaluating an expression, then use that name in later expressions. There should be a way to hold either strings or boolean values in variables.
Tips
There is some syntax to figure out with boolean literals and the new operations, but you should be comfortable doing that already from previous weeks.
The really new things are variables and having two different types. Think about how you want to handle that in your language!
-
What does it look like to declare a variable and assign it a value? Maybe these happen in a single statement, maybe not — up to you.
-
Is the programmer allowed to re-assign a variable with a new value after assigning it initially?
-
Variables values can either be strings or booleans. Should the programmer specify the type when declaring the variable (like in C, C++, Java)?
-
Should there be special rules for what variable names are allowed to be for each type (Perl kind of does this with sigils)?
-
Or should we let the programmer assign whatever type they want to whatever variable they want, like in Python?
All of this is up to you, and we expect to see a lot of different choices explored in what you turn in. (Ideas not listed above are also welcome!)
REQUIREMENTS
Fill in the language_name field in your YAML file with a name you choose
for your programming language.
To receive credit for this lab, your language needs to:
-
Be original. The syntax and semantics should not be very similar to any single existing language, or to any of your classmates’ languages.
-
Support the required capabilities
-
Not have unnecessary extra features
Task 2: Example program
Write an example program in your language which works equivalently to the following Python program.
(Of course, your example program will not have these function definitions since your program will have those operations built-in!)
def reverse(s):
return s[::-1]
def chop(a, b):
return a.split(b)[0]
def bool2str(x):
return 'T' if x else 'F'
# this is where the program actually starts!
order = input()
label = reverse(chop(order, " salad"))
rushhour = True
manager = False
weird = not (label + "sandwich" < "mustard") or manager or False
print("eat up: [" + label + "]")
status = "hot: " + bool2str(not rushhour or weird)
print(status)
REQUIREMENTS
Fill in the complete source code (along with comments) under the
example_program field in your YAML file.
Fill in your test cases under example_input_1, example_output_1, etc.
Your example program must:
-
Demonstrate all of your language’s capabilities (including comments)
-
Be easy to read and follow
-
Have the exact same behavior as the python program above
Task 3: Scanner specification
Now let’s start specifying the syntax of your language.
Give a table of token names (like STRLIT or ID or BOOLLIT) and corresponding regexes for all of the possible tokens in your language.
The list should be in order of highest to lowest priority tokens. Regular expressions should use Java syntax.
Use the token name ignore for regexes that should be ignored (like comments
and whitespace).
REQUIREMENTS
Fill in the table in the tokens field of your YAML file.
Your tokens must:
- Cover every character in any valid source code for your language
- Use Java regex syntax
- Work with the
ptreetest program to correctly tokenize your example program
Task 4: Parser specification
Complete the syntax spec for your language by providing a context-free grammar (CFG) for your language.
Use the token names from the token spec that you just wrote above. Each token name should be all caps and will constitute the terminal symbols in your grammar.
The nonterminals in your grammar should be written in lower case. The first nonterminal will be treated as the start symbol, which will be the root node of any correct parse tree for your language covering the complete program.
Follow the ANTLR syntax for writing the grammar. If your grammar is in
MyGrammar.g4, then it must start with a line like:
parser grammar MyGrammar;
What comes next must be a complete listing of your tokens, like:
tokens { YOUR, TOKENS, GO, HERE }
And then we get to the actual grammar rules. Each nonterminal should
appear just once; use the | operator to separate multiple expansions
of the same nonterminal. And there is no epsilon (or lambda) symbol in
ANTLR syntax; you just leave it blank.
So for example what we might write in class examples as
prog -> stmt prog
prog -> ε
expr -> INT
expr -> ADDOP expr
expr -> expr MULOP expr
would be written in ANTLR syntax like
prog
: stmt prog
|
;
expr
: INT
| ADDOP expr
| expr MULOP expr
;
Don’t forget the semicolon at the end of each rule!
REQUIREMENTS
Fill in the grammar field of your YAML file with CFG production rules.
Your grammar must:
- Use the token types from the previous part as terminal symbols
- Allow any valid program in your language to parse as a single tree with the start symbol (first nonterminal) as the root
- Follow the ANTLR syntax and work with the
ptreetest program - Work to correctly parse your example program
Task 5: Semantics specification
Indicate the meaning of any syntactically valid program in your language, answering questions such as:
- What do all the different operators do?
- What do “true” and “false” literals look like?
- How will boolean values get printed?
- How do variables work regarding types and reassignment (see discussion above)?
You can assume we all know what “string concatenation”, “reversal”, “lexicographical comparison”, etc., all mean. The question here is to connect the parts of your syntax with these meanings.
REQUIREMENTS
Fill in the semantics field of your YAML file with a specification of
the language semantics, connecting the required capabilities to syntax
in your particular language.
Your semantics spec must:
- Be clear and easy to understand
- Not be overly long or explain “obvious” aspects
- Make it possible for anyone to know exactly how a program in your language should behave
Task 6: Peer review
Just like with Lab 1.1, I want you to review each other’s language specs carefully. The job of the reviewer is to see if they can understand exactly how the language works without any extra explanations. Think about edge cases and details to help each other get it right, and challenge your self to really dig into the details of where the language might “break” or be underspecified.
By indicating Y that your spec has passed “final review”, you are
saying that the version you are turning in has been successfully
reviewed and the reviewer agrees it covers all the requirements with no
ambiguities or UB.
REQUIREMENTS
Fill in the last four questions in your YAML file, indicating whose spec
you have or will review, who has reviewed your spec, and a clear Y or
N indicating whether the reviewer agrees that the version you are
turning in satisfies all the requirements of the language and is
clearly presented.