Files
notes/docs/lectures/compilers/02_arithmetic.md
T

2.4 KiB

Arithmetic Grammar

Syntax of Expressions

An expression can be defined as:

exp ::= int | exp + exp | exp - exp | exp * exp
		| exp / exp | - exp | ( exp )

7 + (10/3) \times (-2)

Applying this to the above expression:

exp -> exp + exp
-> int + exp
-> 7 + exp
-> 7 + exp * exp
-> 7 + (exp) * exp
-> 7 + (exp / exp) * exp
-> 7 + (int / int) * exp
-> 7 + (10 / 3) * (-exp)
-> 7 + (10 / 3) * (-int)
-> 7 + (10 / 3) * (-2)

This grammar is ambiguous, this means one input expression could be generated in several different ways.


5 - 4 \times 7
exp -> exp - exp 
-> int - exp 
-> 5 - exp
-> 5 - exp * exp
...
-> 5 - 4 * 7

However there is another way to derive this expression starting with *

exp -> exp * exp
-> exp - exp * exp
...
-> 5 - 4 * 7

These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS.

exp  ::= mexp | mexp + exp | mexp - exp
mexp ::= term | term * mexp | term / mexp --multiplicative expression
term ::= int | - term | ( exp )

exp -> mexp - exp
-> term - exp
-> int - exp
-> 5 - exp
-> 5 - mexp
-> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp
-> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7

This grammar is unique (non-ambiguous)

Semantics of Expressions

On the left hand side the + is just a symbol, however on the right hand side it is an arithmetic sum operation.

$[![ exp + exp ]!] = [![exp ]!] + [![exp ]!]$ | [\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!] ... same for all binary operations

[\![ -exp]\!] = - [\![exp ]\!]

[\![ x]\!] = x

[\![(exp) ]\!] = [\![exp ]\!] - This is because parentheses change order of operations, not the operation itself.

Addition can be rewritten:


[\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!])
int ::= digit | int digit
digit ::= 0 | 1 | 2 | 3 ... | 9

[\![d_0 ]\!] = value(d_0)

[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)

Scanners and Parsers

img

Scanners take the source language as input and outputs a stream of tokens.

A token is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis.

The grammar of tokens is always regular, this means it can be generated and recognised by a DFA (deterministic finite automata).