# Arithmetic Grammar ## Syntax of Expressions An expression can be defined as: ```haskell exp ::= int | exp + exp | exp - exp | exp * exp | exp / exp | - exp | ( exp ) ``` $$ 7 + (10/3) \times (-2) $$ Applying this to the above expression: ```haskell exp -> exp + exp -> int + exp -> 7 + exp -> 7 + exp * exp -> 7 + (exp) * exp -> 7 + (exp / exp) * exp -> 7 + (int / int) * exp -> 7 + (10 / 3) * (-exp) -> 7 + (10 / 3) * (-int) -> 7 + (10 / 3) * (-2) ``` This grammar is **ambiguous**: this means one input expression could be generated in several different ways. $$ 5 - 4 \times 7 $$ ```haskell exp -> exp - exp -> int - exp -> 5 - exp -> 5 - exp * exp ... -> 5 - 4 * 7 ``` However, there is another way to derive this expression starting with `*` ```haskell exp -> exp * exp -> exp - exp * exp ... -> 5 - 4 * 7 ``` These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS. ![](img/e.png) ```haskell exp ::= mexp | mexp + exp | mexp - exp mexp ::= term | term * mexp | term / mexp --multiplicative expression term ::= int | - term | ( exp ) exp -> mexp - exp -> term - exp -> int - exp -> 5 - exp -> 5 - mexp -> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp -> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7 ``` This grammar is unique (non-ambiguous) ## Semantics of Expressions On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation. $[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations $[\![ -exp]\!] = - [\![exp ]\!]$ $[\![ x]\!] = x$ $[\![(exp) ]\!] = [\![exp ]\!]$ - This is because parentheses change order of operations, not the operation itself. Addition can be rewritten: $$ [\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!]) $$ ```haskell int ::= digit | int digit digit ::= 0 | 1 | 2 | 3 ... | 9 ``` $[\![d_0 ]\!] = value(d_0)$ $[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$ ## Scanners and Parsers ![img](img/f.png) Scanners take the source language as input and output a stream of tokens. A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses. The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).