2.4 KiB
Arithmetic Grammar
Syntax of Expressions
An expression can be defined as:
exp ::= int | exp + exp | exp - exp | exp * exp
| exp / exp | - exp | ( exp )
7 + (10/3) \times (-2)
Applying this to the above expression:
exp -> exp + exp
-> int + exp
-> 7 + exp
-> 7 + exp * exp
-> 7 + (exp) * exp
-> 7 + (exp / exp) * exp
-> 7 + (int / int) * exp
-> 7 + (10 / 3) * (-exp)
-> 7 + (10 / 3) * (-int)
-> 7 + (10 / 3) * (-2)
This grammar is ambiguous: this means one input expression could be generated in several different ways.
5 - 4 \times 7
exp -> exp - exp
-> int - exp
-> 5 - exp
-> 5 - exp * exp
...
-> 5 - 4 * 7
However, there is another way to derive this expression starting with *
exp -> exp * exp
-> exp - exp * exp
...
-> 5 - 4 * 7
These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
exp ::= mexp | mexp + exp | mexp - exp
mexp ::= term | term * mexp | term / mexp --multiplicative expression
term ::= int | - term | ( exp )
exp -> mexp - exp
-> term - exp
-> int - exp
-> 5 - exp
-> 5 - mexp
-> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp
-> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7
This grammar is unique (non-ambiguous)
Semantics of Expressions
On the left-hand side, the + is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
$[![ exp + exp ]!] = [![exp ]!] + [![exp ]!]$ | [\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!] ... same for all binary operations
[\![ -exp]\!] = - [\![exp ]\!]
[\![ x]\!] = x
[\![(exp) ]\!] = [\![exp ]\!] - This is because parentheses change order of operations, not the operation itself.
Addition can be rewritten:
[\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!])
int ::= digit | int digit
digit ::= 0 | 1 | 2 | 3 ... | 9
[\![d_0 ]\!] = value(d_0)
[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)
Scanners and Parsers
Scanners take the source language as input and output a stream of tokens.
A token is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
The grammar of tokens is always regular: this means it can be generated and recognised by a DFA (deterministic finite automaton).

