111 lines
2.4 KiB
Markdown
111 lines
2.4 KiB
Markdown
# Arithmetic Grammar
|
|
|
|
## Syntax of Expressions
|
|
|
|
An expression can be defined as:
|
|
|
|
```haskell
|
|
exp ::= int | exp + exp | exp - exp | exp * exp
|
|
| exp / exp | - exp | ( exp )
|
|
```
|
|
|
|
$$
|
|
7 + (10/3) \times (-2)
|
|
$$
|
|
|
|
Applying this to the above expression:
|
|
|
|
```haskell
|
|
exp -> exp + exp
|
|
-> int + exp
|
|
-> 7 + exp
|
|
-> 7 + exp * exp
|
|
-> 7 + (exp) * exp
|
|
-> 7 + (exp / exp) * exp
|
|
-> 7 + (int / int) * exp
|
|
-> 7 + (10 / 3) * (-exp)
|
|
-> 7 + (10 / 3) * (-int)
|
|
-> 7 + (10 / 3) * (-2)
|
|
```
|
|
|
|
This grammar is **ambiguous**: this means one input expression could be generated in several different ways.
|
|
|
|
$$
|
|
5 - 4 \times 7
|
|
$$
|
|
|
|
```haskell
|
|
exp -> exp - exp
|
|
-> int - exp
|
|
-> 5 - exp
|
|
-> 5 - exp * exp
|
|
...
|
|
-> 5 - 4 * 7
|
|
```
|
|
|
|
However, there is another way to derive this expression starting with `*`
|
|
|
|
```haskell
|
|
exp -> exp * exp
|
|
-> exp - exp * exp
|
|
...
|
|
-> 5 - 4 * 7
|
|
```
|
|
|
|
These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
|
|
|
|

|
|
|
|
```haskell
|
|
exp ::= mexp | mexp + exp | mexp - exp
|
|
mexp ::= term | term * mexp | term / mexp --multiplicative expression
|
|
term ::= int | - term | ( exp )
|
|
|
|
exp -> mexp - exp
|
|
-> term - exp
|
|
-> int - exp
|
|
-> 5 - exp
|
|
-> 5 - mexp
|
|
-> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp
|
|
-> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7
|
|
```
|
|
|
|
This grammar is unique (non-ambiguous)
|
|
|
|
## Semantics of Expressions
|
|
|
|
On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
|
|
|
|
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
|
|
|
|
$[\![ -exp]\!] = - [\![exp ]\!]$
|
|
|
|
$[\![ x]\!] = x$
|
|
|
|
$[\![(exp) ]\!] = [\![exp ]\!]$ - This is because parentheses change order of operations, not the operation itself.
|
|
|
|
Addition can be rewritten:
|
|
|
|
$$
|
|
[\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!])
|
|
$$
|
|
|
|
```haskell
|
|
int ::= digit | int digit
|
|
digit ::= 0 | 1 | 2 | 3 ... | 9
|
|
```
|
|
|
|
$[\![d_0 ]\!] = value(d_0)$
|
|
|
|
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
|
|
|
|
## Scanners and Parsers
|
|
|
|

|
|
|
|
Scanners take the source language as input and output a stream of tokens.
|
|
|
|
A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
|
|
|
|
The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).
|