Files
notes/docs/lectures/compilers/02_arithmetic.md
T
2026-10-04 15:24:17 +01:00

111 lines
2.4 KiB
Markdown

# Arithmetic Grammar
## Syntax of Expressions
An expression can be defined as:
```haskell
exp ::= int | exp + exp | exp - exp | exp * exp
| exp / exp | - exp | ( exp )
```
$$
7 + (10/3) \times (-2)
$$
Applying this to the above expression:
```haskell
exp -> exp + exp
-> int + exp
-> 7 + exp
-> 7 + exp * exp
-> 7 + (exp) * exp
-> 7 + (exp / exp) * exp
-> 7 + (int / int) * exp
-> 7 + (10 / 3) * (-exp)
-> 7 + (10 / 3) * (-int)
-> 7 + (10 / 3) * (-2)
```
This grammar is **ambiguous**: this means one input expression could be generated in several different ways.
$$
5 - 4 \times 7
$$
```haskell
exp -> exp - exp
-> int - exp
-> 5 - exp
-> 5 - exp * exp
...
-> 5 - 4 * 7
```
However, there is another way to derive this expression starting with `*`
```haskell
exp -> exp * exp
-> exp - exp * exp
...
-> 5 - 4 * 7
```
These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
![](img/e.png)
```haskell
exp ::= mexp | mexp + exp | mexp - exp
mexp ::= term | term * mexp | term / mexp --multiplicative expression
term ::= int | - term | ( exp )
exp -> mexp - exp
-> term - exp
-> int - exp
-> 5 - exp
-> 5 - mexp
-> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp
-> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7
```
This grammar is unique (non-ambiguous)
## Semantics of Expressions
On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
$[\![ -exp]\!] = - [\![exp ]\!]$
$[\![ x]\!] = x$
$[\![(exp) ]\!] = [\![exp ]\!]$ - This is because parentheses change order of operations, not the operation itself.
Addition can be rewritten:
$$
[\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!])
$$
```haskell
int ::= digit | int digit
digit ::= 0 | 1 | 2 | 3 ... | 9
```
$[\![d_0 ]\!] = value(d_0)$
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
## Scanners and Parsers
![img](img/f.png)
Scanners take the source language as input and output a stream of tokens.
A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).