Add the rest of university notes

This commit is contained in:
John Gatward committed 2026-10-04 14:02:35 +01:00
1 parent c1b84c7f7d
commit d0f27f276b
366 files changed
+9844 -110

No files matched your search

+30
View File
@@ -0,0 +1,30 @@
# Compilers - COMP 3012
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to program written in a target programming language. A compiler is written in an **implementation language**.
![img](img/a.png)
Notation of semantics of program $A$: $[\![A]\!]$
An **interpreter** is a program that takes a source program and executes the program.
![img](img/b.png)
> NOTE: Java uses both. A java source program is compiled into byte code (by a compiler) which is then executed by an interpreter (called java virtual machine - JVM). JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, where as converting IR to executable code is called **back end**.
* The front end focuses on understand the source-language program.
* The back end focuses on mapping programs to the target machine
![img](img/c.png)
* The front end, intermediate representation and the back end are all part of the compiler.
IR is stored as an Abstract Syntax Tree **AST**.
The syntactic details needed for parsing the source program are represented in the structure of the tree.
The **IR** could be broken down into many sub-steps i.e. a IR1 could be created which is then ran through an optimiser to create IR2 which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
![img](img/d.png)
+112
View File
@@ -0,0 +1,112 @@
# Arithmetic Grammar
## Syntax of Expressions
An expression can be defined as:
```haskell
exp ::= int | exp + exp | exp - exp | exp * exp
| exp / exp | - exp | ( exp )
```
$$
7 + (10/3) \times (-2)
$$
Applying this to the above expression:
```haskell
exp -> exp + exp
-> int + exp
-> 7 + exp
-> 7 + exp * exp
-> 7 + (exp) * exp
-> 7 + (exp / exp) * exp
-> 7 + (int / int) * exp
-> 7 + (10 / 3) * (-exp)
-> 7 + (10 / 3) * (-int)
-> 7 + (10 / 3) * (-2)
```
This grammar is **ambiguous**, this means one input expression could be generated in several different ways.
$$
5 - 4 \times 7
$$
```haskell
exp -> exp - exp
-> int - exp
-> 5 - exp
-> 5 - exp * exp
...
-> 5 - 4 * 7
```
However there is another way to derive this expression starting with `*`
```haskell
exp -> exp * exp
-> exp - exp * exp
...
-> 5 - 4 * 7
```
These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
![](img/e.png)
```haskell
exp ::= mexp | mexp + exp | mexp - exp
mexp ::= term | term * mexp | term / mexp --multiplicative expression
term ::= int | - term | ( exp )
exp -> mexp - exp
-> term - exp
-> int - exp
-> 5 - exp
-> 5 - mexp
-> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp
-> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7
```
This grammar is unique (non-ambiguous)
## Semantics of Expressions
On the left hand side the $+$ is just a symbol, however on the right hand side it is an arithmetic sum operation.
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
$[\![ -exp]\!] = - [\![exp ]\!]$
$[\![ x]\!] = x$
$[\![(exp) ]\!] = [\![exp ]\!]$ - This is because parentheses change order of operations, not the operation itself.
Addition can be rewritten:
$$
[\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!])
$$
```haskell
int ::= digit | int digit
digit ::= 0 | 1 | 2 | 3 ... | 9
```
$[\![d_0 ]\!] = value(d_0)$
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
## Scanners and Parsers
![img](img/f.png)
Scanners take the source language as input and outputs a stream of tokens.
A **token** is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis.
The **grammar of tokens is always regular**, this means it can be generated and recognised by a DFA (deterministic finite automata).
+129
View File
@@ -0,0 +1,129 @@
# Triangle Abstract Machine
**TAM** instruction set
```assembly
LOADL (int)
NEG
ADD
SUB
MUL
DIV
```
TAM works on a stack of integers.
##### Executing a TAM program
```assembly
LOADL 7
ADD --adds top two numbers on the stack
LOADL 2
SUB -- note its 15-2
LOADL 4
DIV --integer division
```
The stack during this program:
$$
\begin{bmatrix}
{8} \\
{5}
\end{bmatrix}
\implies
\begin{bmatrix}
{7} \\
{8} \\
{5}
\end{bmatrix}
\implies
\begin{bmatrix}
{15} \\
{5}
\end{bmatrix}
\implies
\begin{bmatrix}
{2} \\
{15} \\
{5}
\end{bmatrix}
\implies
\begin{bmatrix}
{13} \\
{5}
\end{bmatrix}
\implies
\begin{bmatrix}
{4} \\
{13} \\
{5}
\end{bmatrix}
\implies
\begin{bmatrix}
{3} \\
{5}
\end{bmatrix}
$$
## Compiler Complete Example
Program in **Arith**
```c
5 * ((8 + 7) - 2) / 4
```
**A**bstract **S**yntax **T**ree
![img](img/h.png)
**TAM** program
```assembly
LOADL 5
LOADL 8
LOADL 7
ADD
LOADL 2
SUB
LOADL 4
DIV
MUL
```
## Implementing TAM in Haskell
```haskell
module TAM where
data TamInstruction = LOADL Int
| ADD | SUB
| MUL | DIV
| NEG
deriving(Eq, Show)
type Stack = [Int]
execute :: [TamInstruction] -> Stack -> Stack
execute [] s = s --if stack empty, then return the stack
execute (LOADL n : tp) s = execute tp (n : s) --push n to top of stack
execute (ADD : tp) (a : b : s) = execute tp ((a+b):s) --push a+b
...
execute (DIV : tp) (a : b : s) = execute tp ((a`div`b):s)
```
Quicker way to write the execute function using `absOpToConcrOp`
```haskell
convOp :: TamInstruction -> Int -> Int -> Int
convOp ADD = (+)
convOp SUB = (-)
convOp MUL = (*)
convOp DIV = (`div`)
execute :: [TamInstruction] -> Stack -> Stack
execute [] s = s
execute (LOADL n : tp) s = execute tp (n : s)
execute (NEG : tp) (a : s) = execute tp (-a : s)
execute (op : tp) (a : b : s) = execute tp ((convOp op a b) : s)
```
@@ -0,0 +1,41 @@
# Functional Parsers
In our parser - there's a lot of repeated code and a lot of cases.
Types of scanner and parser are very similar
```haskell
scanToken :: String -> Maybe (Token, String)
parseTerm :: [Token] -> Maybe (AST, [Token])
parseExp :: [Token] -> Maybe (AST, [Token])
general :: [c] -> Maybe (a, [c])
lessGeneral :: String -> Maybe (a, String)
--remember :t string :: [char]
-- [c] list of characters or tokens
Parser a :: String -> [(a, String)]
-- no maybe needed as failure is now returning an empty list
--The parser of type a, we can now define generic functions that operate on a given type => less repeated code
--This is an instance of a typeclass (monad yikes)
```
Do notation
```haskell
Parser a = String -> [(a, String)]
symbol :: String -> Parser ()
-- no need to define result, as all it does it succeed or fail
exp :: Parser AST
parseParenthesis :: Parser AST
parseParenthesis = do symbol '('
t <- exp
symbol ')'
return t
```
+101
View File
@@ -0,0 +1,101 @@
# Functor
Parsing an expression in parenthesis:
```haskell
parseP :: Parser AST
parseP = do symbol '('
t <- exp
symbol ')'
return t
```
Before we write this sort of code, we need to understand `type classes` (especially `monads`)
## Types vs Typeclasses
| Types | Type classes |
| ------ | ------------ |
| Bool | Eq |
| Char | Show |
| AST | Num |
| String | Functor |
| | Monad |
**Eq**: typeclass equality; A type can only be typeclass equality if two like types can be compared
A type can be a *member* (instance) of a type class, meaning that if has the properties/functions that the class requires
e.g. `Bool` is an instance of `Eq` and `Show`
###### Is there a type that is **not** in `Eq`?
```haskell
(\c -> c :: Int) == (\c -> c :: Int)
```
**ERROR**: No instance for `Eq(Int -> Int)`
Why?
```haskell
f :: Int -> Int
g :: Int -> Int
```
Then `f == g` should be `fn == gn` for every n, the computer cannot do this (halting problem).
## Type Constructors
A type constructor takes a type to construct a new type.
`Maybe` - not a type but a type constructor
`Maybe String` - a type
```haskell
newtype Parser a = P (String -> [a, String])
```
**Parser** is a type constructor
**Parser AST** is a type
Functor is a typeclass of which `parser` is an instance
##### Functor
```haskell
class Functor f where
fmap :: (a -> b) -> fa -> fb
instance Functor Maybe where
fmap g (Just x) = Just (g x)
fmap g Nothing = Nothing -- fmap id = id
-- lists
instance Functor [] where
fmap g [] = []
fmap g (t:ts) = (g t) : fmap g ts
-- goal: write parser as a functor
newtype Parser a = P ( String -> [a, String] )
-- Need: fmap :: (a->b) -> Parser a -> Parser b
instance Functor Parser where
fmap g pa = -- parser pa
P (\str -> map (\(x,s) -> (gx,s))
parse pa str)
```
##### Rules of Functors
```haskell
fmap id = id -- identity
fmap (f . g) = fmap f . fmap g
```
Haskell doesn't enforce these rules however it is convention.
@@ -0,0 +1,95 @@
# Applicative Functors
Types: `Bool`, `Int`, `Char`, `[Char] = String`
Type Constructors: `Maybe`, `[]`
(type) classes: `Eq`, `Show`, `Functor`
```haskell
newtype Parser a = P ( String -> [a, String])
parse :: Parser a -> String -> [(a, String)]
parse (P p) s = p s -- s's can be cancelled from both sides
instance Functor Parser where
-- fmap :: (a -> b) -> Parser a -> Parser b
fmap g pa = P (\s -> [(g x, s1) |
(x,s1) <- parse pa s])
```
Applicative - motivation
```haskell
Functor f
fmap0 :: a -> f a
fmap1 :: (a -> b) -> f a -> f b
-- cannot do this with functors ie cannot deal with multiple parameters
fmap2 :: (a -> b -> c) -> f a -> f b -> f c
fmap3 :: (a -> ... n) -> f a -> ... f n
```
`Functor f` can do `fmap1` however cannot do `fmap0` or `fmap2` etc.
**Remember**: `a -> b -> c == a -> (b -> c)`
For `fmap2` we can use `fmap2 :: (a -> (b -> c)) -> f a -> f (a -> b)`
would need: `f(b -> c) -> f b -> f c`
```haskell
class Functor f => Applicative f where
pure :: a -> f a
(<*>) :: f (a -> b) -> f a -> f b
-- <*> infix operator
-- NOTE its f (a -> b) and not (a -> b) in fmap1
-- fmap1 not part of the applicative class
```
Writing `fmap3` in an applicative functor
```haskell
fmap3 :: g x y z = (pure g) <*> x <*> y <*> z
```
##### Example Maybe
```haskell
instance Applicative Maybe where
-- pure :: a -> Maybe a
pure x = Just x
-- (<*>) :: Maybe (a -> b) -> Maybe a -> Maybe b
Just g <*> (Just x) = Just (g x)
_ <*> _ = Nothing
```
##### Example Lists
```haskell
instance Applicative [] where
-- pure :: a -> [a]
pure x = [x]
-- (<*>) :: [a -> b] -> [a] -> [b]
gs <*> xs = [g x | g <- gs, x <- xs]
```
##### Example Parser
```haskell
instance Applicative Parser where
-- pure :: a -> Parser a
-- newtype Parser a = P ( String -> [(a, String)] )
pure x = P (\s -> [(x,s)])
-- <*> :: Parser (a -> b) -> Parser a -> Parser b
pf <*> pa = P (\s -> [ (f x, s2) |
(f, s1) <- parse pf s,
(x, s2) <- parse pa s1)])
```
All parse does is apply a parser
`parse :: Parser a -> String -> [(a, String)]`
Where `P` is the constructor
`parse ( P p ) = p`
@@ -0,0 +1,466 @@
### Functor Class of Parsers
```haskell
newtype Parser a = P ( String -> [(a, String)] )
parse :: Parser a -> String -> [(a, String)]
parse (P f) src = f src
item :: Parser Char
item = P (\src -> case src of
[] -> []
(c:src') -> [(c,src')] )
symbol :: String -> Parser ()
integer :: Parser Int
binary :: Parser Int
intORbin :: Parser Int
expr :: Parser AST
```
```
λ> parse (symbol "something") "nothing"
[]
λ> parse (symbol "<=") "<= something nothing"
[((), "something nothing")]
NOTE: does nothing because all we have implemented for symbol is ()
λ> integer "123 blah blah"
[(123, "blah blah")]
λ> parse binary "101 blah"
[(5, "blah")]
λ> parse intORbin "101 blah"
[(101, "blah"), (5, "blah")]
λ> parse expr "1+2*3"
[(BinOp Addition (LitInteger 1) BinOp Multiplication (LitInteger 2) (LitInteger 3)), "")]
NOTE: expr defined in ArtihExpr
```
Defining the functor parser
```haskell
instance Functor Parser where
-- must not give type of fmap as it is already given in functor class
-- good practice to comment type
-- fmap :: (a -> b) -> Parser a -> Parser b
--first assume returns one value
-- doesnt fail, doesn't produce more than one result
fmap g pa = P (\src -> let [(x,src1)] = parse pa src
in [(g x, src1)] )
```
```
λ> parse (fmap (+3) integer) "42 blah blah"
[(45, blah blah)]
λ> parse (fmap evaluate expr) "1+2*3"
[(7,"")]
λ> parse (fmap (+3) integer) "42 blah blah"
*** Exception Non-exhaustive patterns
λ> parse (fmap (+3) intORbin) "101 blah"
*** Exception Non-exhaustive patterns
```
fixing `fmap`
```haskell
fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src])
-- using list comprehension
```
```
λ> parse (fmap (+3) intORbin) "101 blah"
[(104, "blah"), (8, "blah")]
```
### Applicative Class of Parsers
```haskell
instance Applicative Parser where
-- pure :: a -> Parser a
-- commenting type for good practice
pure x = P (\src -> [(x, src)])
-- (<*>) :: Parser (a -> b) -> Parser a -> Parser b
simpleFun :: Parser (Int -> Int)
-- parser the function "double" or "square"
```
```
λ> parse (fmap (\f -> f 3) simpleFun) "double blah"
[(6, "blah")]
a parser that returns a function as a result
λ> parse simpleFun "double blah blah"
parse simpleFun "double blah blah" :: [(Int -> Int, String)]
-- the function
```
```haskell
instance Applicative Parser where
-- pure :: a -> Parser a
-- commenting type for good practice
pure x = P (\src -> [(x, src)])
-- (<*>) :: Parser (a -> b) -> Parser a -> Parser b
pf <*> pa = P (\src -> let [(f,src1)] = parse pf src
[(x,src2)] = parse pa src
in [(f x, src2)] )
-- this works if the two parsers both give one, different result
```
```
λ> parse (simpleFun <*> integer) "double 7"
[(14, "")]
λ> parse (simpleFun <*> integer) "square 7"
[(49, "")]
λ> parse (simpleFun <*> integer) "cube 7"
*** Exception non-exhaustive pattern
λ> parse (simpleFun <*> intORbin) "square 101"
*** Exception non-exhaustive pattern
-- fails bc intORbin gives two results
```
Using list comprehension
```haskell
pf <*> pa = P (\src -> [ (f x, src2) | (f,src1) <- parse pf src,
(x,src2) <- parse pa src1 ] )
```
```
λ> parse (simpleFun <*> integer) "cube 7"
[]
λ> parse (simpleFun <*> intORbin) "square 101"
[(10201, ""), (25, "")]
```
### Monad Class of Parser
Monad class will facilitate the use of `do` notation.
```haskell
instance Monad Parser where
-- return :: a -> Parser a
-- we dont have to define return as its automatically defined as
-- return = pure
--only method we need to define for the monad class is bind >>=
-- (>>=) :: Parser a -> (a -> Parser b) -> Parser b
pa >>= fpb = P (\src -> let [(x, src1)] = parse pa src
[(y, src2)] = parse (fpb x) src1
in [(y,src2)] )
checkNum :: Int -> Parser Bool
checkNum n = fmap (==n) integer
```
```
λ> parse (checkNum 7) " 7 blah blah"
[(True, "blah blah")]
λ> parse (checkNum 6) " 7 blah blah"
[(False, "blah blah")]
λ> parse (checkNum 7) " no blah blah"
[]
λ> parse (binary >>= checkNum) "101 5"
[(True, "")]
λ> parse (binary >>= checkNum) "101 6"
[(False, "")]
λ> parse (binary >>= checkNum) "no 101 6"
*** Exception non-exhaustive pattern
λ> parse (intORbin >>= checkNum) "101 6"
*** Exception non-exhaustive pattern
--cant cope with multiple values
```
Using list comprehension
```haskell
pa >>= fpb = P (\src -> [ (y,src2) | (x,src1) <- parse pa src,
(y,src2) <- parse (fpb x) src1 ] )
```
```
λ> parse (binary >>= checkNum) "no 101 6"
[]
λ> parse (intORbin >>= checkNum) "110 6"
[(False,""), (True, "")]
-- false is 110 (base 10) != 6
-- true is 110 (base 2) == 6
```
Improving the definition further
As we unpack and repack `(y,src2)`, we can just call it `r` (result)
```haskell
pa >>= fpb = P (\src -> [ r | (x,src1) <- parse pa src,
r <- parse (fpb x) src1 ] )
```
```
λ> parse (intORbin >>= checkNum) "113 113"
[(True,""), (True, "113 ")]
-- the integer part recognises 113 == 113
-- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true
```
What is the do notation and how is it connected to the bind function, we will show this by writing a simple parser
```haskell
pairSum :: Parser Int
-- read (parse) an integer, bind it to a function, map it to another parser
pairSum = integer >>= \n -> integer >>= \m -> return (n+m)
```
```
λ> parse pairSum "3 8"
[(11, "")]
```
Rewriting `pairSum` with `do`
```haskell
pairSum :: Parser Int
-- apply integer and then put it into variable n
-- apply integer and bind to variable m
pairSum = do n <- integer
m <- integer
return (n+m)
--much cleaner & easier to understand
```
```
parse (symbol "number" >>= \u -> integer) "number 9"
[(9, "")]
parse (symbol "number" >> integer) "number 9"
[(9, "")]
NOTE: >> is a non-dependant bind
```
```haskell
the grammer
--funApp ::= ( simpleFun integer )
-- will be a parser that returns an integer
funApp :: Parser Int
funApp = symbol '(' >> (simpleFun <*> integer) >>= \y -> symbol ')' >> return y
```
```
λ> parse funApp "(double 5)"
[(10, "")]
```
Rewrite with `do`
```haskell
funApp = do symbol '('
f <- simpleFun
x <-integer
symbol ')'
return (f x)
```
### Alternative Class of Parser
```haskell
instance Alternative Parser where
-- empty :: Parser a
empty = P (\src -> [])
-- (<|>) :: Parser a -> Parser a -> Parser a
p1 <|> p2 = P (\src -> case parse p1 src of
[] -> parse p2 src
rs -> rs)
-- if p1 fails, then parse with p2, else return result rs
```
```
λ> parse (symbol "abc" <|> symbol "acb") "abc"
[("abc", "")]
λ> parse (symbol "abc" <|> symbol "acb") "xyz"
[]
λ> parse (integer <|> binary) "1101"
[(1101,"")]
λ> parse (binary <|> integer) "1101"
[(13,"")]
-- will only apply p2 if p1 fails
λ> parse (binary <|> integer) "1201"
[(1,"201")]
-- binary successfully parses "1" and leaves "201"
```
Using parallel choice notation `<||>`
```haskell
(<||>) :: Parser a -> Parser a -> Parser a
p1 <||> p2 = P (\src -> parse p1 src ++ parse p2 src)
```
```
λ> parse (binary <||> integer) "1101"
[(13, ""), (1101, "")]
```
### Explaining the `FunParser.hs` library
```haskell
satisfy :: Parser a -> (a -> Bool) -> Parser a
satisfy p cond = do x <- p
if (cond x) then return x
else empty
-- the way to denote failure is empty (from alternitve class)
```
```
λ> parse (satisfy integer (>10)) "42"
[(42, "")]
λ> parse (satisfy integer (>10)) "9"
[]
```
Writing a satisfy function just for characters
```haskell
sat :: (Char -> Bool) -> Parser Char
-- item parses 1 character
sat cond = satisfy item cond
```
```
λ> parse (sat isUpper) "a"
[]
λ> parse (sat isUpper) "A"
['A',""]
```
```haskell
lower :: Parser Char
lower = sat isLower
upper :: Parser Char
upper = sat isUpper
digit :: Parser Char
digit = sat isDigit
--and so on for others like letter & alphaNumeric
char :: Char -> Parser Char
char c = sat (==c)
```
```
λ> parse (char 'A') "not a captial a"
[]
λ> parse (char 'A') "A not a captial a"
['A'," not a capital a"]
```
```haskell
string :: String -> Parser String
string [] = return [] --list as string is list of chars
string (c:cs) = do char c
string cs
return (c:cs)
```
```
λ> parse (string "hello") "hello everybody"
[("hello", "everybody")]
λ> parse (string "hello") " hello everybody"
[]
λ> parse (sat isSpace) " hello"
[(' ',"hello")]
λ> parse (many (sat isSpace)) " hello"
[(' ',"hello")]
```
We have to fix leading white space causing failure
```haskell
space :: Parser ()
-- a parser that succeeds or fails and does not return anything
space = do many (sat isSpace)
return ()
-- writing a parser to ignore white space
token :: Parser a -> Parser a
token p = do space
x <- p
space
return x
```
```
λ> parse (token (string "hello")) " hello everybody"
[("hello","everybody")]
```
```haskell
symbol :: String -> Parser String
symbol = token (string s)
```
```
λ> parse (symbol "hello") " hello everybody"
[("hello","everybody")]
```
#### Defining parsers for arithmetic expressions
```haskell
-- expr ::= mexpr + exp | mexpr - exp | mexpr
expr :: Parser AST
expr = do t1 <- mexpr
symbol '+'
t2 <- expr
return (BinOp Addition t1 t2)
<|>
do t1 <- mexpr
symbol '-'
t2 <- expr
return (BinOp Subtraction t1 t2)
<|>
mexpr
--we can optimise this grammer as all symbols start with mexpr
-- expr ::= mexpr ( + expr | - expr | empty)
expr :: Parser AST
expr = do t1 <- mexpr
(do symbol '+'
t2 <- expr
return (BinOp Addition t1 t2)
<|>
do symbol '-'
t2 <- expr
return (BinOp Subtraction t1 t2)
<|>
return t1)
```
+149
View File
@@ -0,0 +1,149 @@
# Compiling Variables
A variable is identified by a alphanumeric string. We can store this as a list of pairs, with the variables identifier and its value.
Variable Environment or VarEnv - `[(Identifier, Stack Address)]`
A stack address is an integer value that specifies where in the stack that variable is contained. The bottom of the stack is reserved for variable values.
The bottom of the stack is indexed `0`.
Lets say our environment consists of 3 variables named x,y,z. It would look like:
`[("z",2), ("y",1), ("x",0)]`
| Variables | Stack (Values) | Index |
| :-------: | :------------: | :---: |
| x | 7 | 0 |
| y | 2 | 1 |
| z | 9 | 2 |
To get the value of a variable from the stackk, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
`LOAD a` - copy address a to top of stack
`STORE a` - pop top of stack to address a
For example if `LOADL 2` is called, it will effect the stack in the following way:
| Variables | Stack (Values) | Index |
| :-------: | :------------: | :---: |
| x | 7 | 0 |
| y | 2 | 1 |
| z | 9 | 2 |
| | … | |
| | 9 | |
```haskell
expCode :: VarEnv -> Expr -> [TAMInst]
```
Before we just called the abstract syntax tree `AST` however with the extended grammar now we will have multiple ASTs, one for programs, one for commands, expressions. The AST for expressions we call `Expr`.
Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list.
## Declaration of Variables
```js
let var x; //no value given means initialised to 0
var y := 5 //note no semicolon
var z;
in ...
```
For the code above, we need to generate a VarEnv. The compiler needs to generate a variable environment and TAM code.
VarEnv: `[("z",2), ("y",1), ("x",0)]`
TAM code stack: `[0,5,0]`
However we also need to account for expressions such as:
```js
let var x := 3;
var y := 5;
var z := x*y
```
```haskell
declarationCompiler :: [Declaration] -> (VarEnv, [TAMInstr])
VarEnv :: [(Identifier, Address)]
```
NOTE: this can be defined with functions given in the `FunParser` library. Or using a `state monad`
### State Monad
$s_0 \rightarrow s_1 \rightarrow s_2 \rightarrow s_n$ for each change in state, there's a corresponding result generated.
$$
a_0 \quad\space\space\space a_1 \quad\space\space\space a_n
$$
- For each of these states, we need a variable environment and address
- For each of the results, we need to generate TAM instructions.
Example: $s_n$ could be your bank balance and $a_n$ could be the purchase history.
- In our case:
- States are VarEnv & next free address space for next variable
- Outputs are TAM instructions
We to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
```haskell
newtype ST st a = S (\st -> (a, st))
-- ST - state transformer
-- st - type of states
-- a - type of output/results
-- S - constructor
-- \st a function that takes a state and returns a value along with a new state
-- this is a general type definition with state type st and result type a
-- this is still just a type constructor, has to be applied to a type
instance Functor (ST st)
instance Applicative (ST st)
instance Monad (ST st)
--as we inherit the monad class, we can use do notation
```
```haskell
newtype ST st a = S (\st -> (a, st))
--type definition
ST Int
--type constructor
ST Int String
--type
```
```haskell
app :: ST st a -> st -> (a, st)
app (S f) x = f x
--applies the constructor to state x
```
```haskell
instance Functor (ST st) where
--fmap :: (a->b) -> ST st a -> ST st b
fmap g sta = S (\s -> let (x,s') = app sta s
in (g x, s'))
```
```haskell
instance Applicative (ST st) where
--pure :: a -> ST st a
pure x = S (\s -> (x,s))
--(<*>) :: (ST st (a -> b)) -> ST st a -> ST st b
stf <*> sta = S (\s -> let (f,s') = app stf s
(x,s'') = app sta s')
in (f x, s''))
```
```haskell
instance Monad (ST st) where
return = pure
-- (>>=) :: (ST st a) -> (a -> ST st b) -> ST st b
sta >>= f = S (\s -> let (x,s') = app sta s
(y,s'') = app (f x) s'
in (y,s''))
```
@@ -0,0 +1,164 @@
# Variable Environments
```haskell
type VarEnv = [(Identifier, StkAddress)]
-- String Int
address :: VarEnv -> Identifer -> StkAddress
address ve v = case lookup v ve of
Nothing -> error "variable not in enviroment"
Just a -> a
--Expr is AST of expressions
expCode :: VarEnv -> Expr -> [TAMInstr]
expCode ve (LitInteger x) = [LOADL x]
-- we must put variable value on top of the stack
expCode ve (Var v) = [LOAD (address ve v)]
```
How do we build a variable environment?
Every program begins with a sequence of variable declarations
```js
var x := 7;
var y := 3;
var z;
var w := x * y - 2
```
The parser will turn this into a list of AST for declarations
Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack.
We do this using the state monad
- We use as an underlying state the variable environment itself, as we build it sequentially
- We also keep the stack address as a state, where it keeps the next free address
```haskell
declsCode :: [Declarations] -> (VarEnv, [TAMInstr])
declsCode ds = let (tam,(ve,0a)) app (declsTAM ds) ([],0) --initial state
in (ve,tam)
declsTAM :: [Declarations] -> ST (VarEnv, StkAddress) [TAMInstr]
declsTAM [] = return []
declsTAM (d:ds) = do
td <- declTAM d
tds <- declsTAM ds
return (td++tds)
declTAM :: Declarations -> ST (VarEnv, StkAddress) [TAMInstr]
declTAM (VarDecl v) = do
(ve,a) <- stState
stUpdate ((v,a) : ve, a+1)
return [LOADL 0]
declTAM (VarInit v e) = do
(ve,a) <- stState
stUpdate ((v,a) : ve, a+1)
return (expCode ve e)
```
```shell
λ> parseAll declarations "var x:=7;var y:=3;var z;var w:=x*y-2"
[VarInit "x" (LitInteger 7), VarInit "y" (LitInteger 3), VarDecl "z", VarInit "w" (BinOp Subtraction (BinOp Multiplication (Var "x") (Var "y")) (LitInteger 2))]
λ> ds = parseAll declarations "var x:=7;var y:=3;var z;var w:=x*y-2"
λ> (ve,tam) = declsCode ds
λ> ve
[("w",3),("z",2),("y",1),("x",0)]
λ> tam
[LOADL 7, LOADL 3m LOADL 0, LOAD 0, LOAD 1, MUL, LOADL 2, SUB]
λ> execTAM [] tam
[19, 0, 3, 7]
```
## Designing ASTs for any grammar
- We turn every non-terminal of the grammar into a type of AST
- We turn every production of the non-terminal into a constructor of the type
Defining the grammar of TAM
```
command ::= identifier := expr
| if expr then command else command
| while expr do command
| getint ( identifier )
| printint ( expr )
| begin commands end
```
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminal
```haskell
data Command =
datatypes Identifier = String, Expr, Commands -- [Command]
```
Assign to every production one constructor for the data type.
This means we will have 6 constructors called `Assignment`, `IfThenElse`, `WhileDo`, `GetInt`, `PrintInt`, `BeginEnd`
```haskell
data Command = Assignment Identifier Expr
| IfThenElse Expr Command Command
| WhileDo Expr Command
| GetInt Identifer
| PrintInt Expr
| BeginEnd [Command]
type Commands = [Command]
--or
data Commands = SingleC Command
| MultipleC Command Commands
```
## Organising a Haskell Project
There are 6 Haskell modules, `Main.hs` is the entry point.
###### Defining a Module
```haskell
module <filename> where
import ...
--definitions
newtype ...
--functions
func :: a -> b
```
Note file name must start with a capital
When you import a module, can can use functions defined in the module
```haskell
data FileType = EXP | TAM
data Option = Trace | Run | Evaluate
main :: IO () --input output monad
```
this is the entry point, to compile
```shell
$ ghc Main.hs -o aec
$ ./aec arith_example.exp --evaluate
Evaluating Expression: 45
```
```haskell
stUpdate :: st -> ST st ()
stUpdate s = S (\_ -> ((), s))
stGet :: ST st st
stGet = S (\s -> (s,s))
stRevise :: (st -> st) -> ST st ()
stRevise f = stGet >>= stUpdate . f
```
+104
View File
@@ -0,0 +1,104 @@
# Compiling Branches
**Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output**
Before we could generate a list of instructions to be executed in sequence, now we need to implement code thats conditionally executed or executed multiple times.
```haskell
--Code for dealing with functions and commands
commCode :: VarEnv -> Command -> TAMProg
```
We will assume an if statement looks like this
IF $e$ THEN $c_1$ ELSE $c_2$
```python
if e then c1 else c2
IfThenElse e c1 c2
```
```haskell
expCode ve e
commCode ve c1
commCode ve c2
-- We dont want to execute both
```
- We can use `JUMPIFZ R1`, a branching function supplied by the TAM language.
- This means `commCode ve c1` & `commCode ve c2` need labels and a `JUMPA` after
```
MINI TRIANGLE PROGRAM
let var := 5
in
begin
if 1
then n := 6
else n := 7;
if 0
then n := 8
else n := 9;
end
```
```assembly
COMPILED VERSION
LOADL 5
LOAD 1 --if 1
JUMPIFZ "label1" --jump to else
LOAD 6 --load the number
STORE 0 --store 0 (stack[0] is 6 from line above) in the place of variable n, for other variables you would have to check the variable enviroment to get the stack address
JUMP "label2"
Label "label1"
LOAD 7
STORE 0
Label "label2"
LOAD 0
JUMPIFZ "label3"
```
![img](img/i.png)
### Generating Labels
Labels must **always** be **unique**.
This would require a global variable in our compiler to count the number of labels, haskell doesnt not allow global variables.
We can use the `stateMonad` instead.
```haskell
type LabelName = String
fresh :: ST Int LabelName
-- Whenever we call fresh, it generates a new label name
-- We can use do (because fresh is element of ST Monad)
fresh = do
n <- stGet --checks current state (which is num of labels)
stUpdate(n+1) --update number of labels
return ("#" : (show n)) -- # symbol to denote labels
-- show converts integer to string (fresh returns string)
commCode :: VarEnv -> Command -> ST Int [TAMInsrt]
-- TAMPrgm is interchangable with [TAMInstr]
commCode ve (IFTHENELSE e c1 c2) =
do l1 <- fresh
l2 <- fresh --generate the two labels needed for an if
let te = expCode ve e --the condition expression
tc1 <- commCode ve c1 --compile success branch
tc2 <- commCode ve c2 --compile else branch
return (te ++ [JUMPIFZ l1] ++ tc1 ++ [JUMP l2]
++ [Label l1] ++ tc2 ++ [Label l2])
--then return the tam instructions
```
**REMINDER**: `expCode` is a function that takes a variable environment `ve` and an expression `e` and generates a list of TAM instructions.
@@ -0,0 +1,88 @@
# Monad Revision
You can think of a monad as a container for a data type
If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype
One of the purposes of the `do` notation is to operate on the whole data structure by specify operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
$$
M_a=\{x_1, x_2, x_3,...\}
$$
```haskell
do x <- m
let y = x ** 2 + 7
return y
```
This extracts an element of type $a$ from $m$, squares and adds 7, and returns the new values as the data structure. Now $M$ is
$$
M_b=\{y_1, y_2, y_3, \ldots\}\\or\\M=\{x_1^2+7, x_2^2+7, x_3^2+7, \ldots\}
$$
The above can be written as a functor
```haskell
fmap (\x -> x**2+7) m
```
Monads have more functionality than functors though
If $x$ is an element of $a$ or $x :: a$
```haskell
x :: a
return x
-- we can also write
pure x
```
Monads can have containers within containers
Assume we have function `makeBlob` that maps every element of $a$ to an element of $M_b$
```haskell
makeBlob :: a -> Mb
makeBlob x1 = do x <- m
y <- makeBlob x
return y
-- this can be done instead with the bind operator
m >>= makeBlob
(>>=) :: Ma -> (a -> Mb) -> Mb
```
## The IO Monad
```haskell
square :: Int -> Int
square x = x*x
getInt :: IO Int
getInt = do putStrLn "Enter a number: "
s <- getLine -- getLine :: IO String
return (read s :: Int) --read :: String -> Int
squareIO :: IO Int
squareIO = do x <- getInt
let y <- square x
return y
-- as squareIO :: IO Int, returning y prints it out
squareIO :: IO () -- unit type, with only one element, also called ()
squareIO = do x <- getInt
let y <- square x
putStrLn("The square " ++ (show x) ++ " is " (show y))
return () --return unit type
-- in this case we dont even need return () as
-- putStrLn :: IO ()
--recursively asks for list unless 0 entered
getList :: IO [Int]
getList = do x <- getInt
if x == 0 then return []
else do
xs <- getList
return (x:xs)
```
Binary file not shown.

After

Width:  |  Height:  |  Size: 6.8 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 6.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 13 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 18 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 30 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 134 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 18 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 14 KiB