This commit is contained in:
John Gatward committed 2026-10-04 15:24:17 +01:00
1 parent d0f27f276b
commit d6f54d4ec2
103 files changed
+3663 -3779

No files matched your search

+8 -8
View File
@@ -1,6 +1,6 @@
# Compilers - COMP 3012
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to program written in a target programming language. A compiler is written in an **implementation language**.
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to a program written in a target programming language. A compiler is written in an **implementation language**.
![img](img/a.png)
@@ -10,21 +10,21 @@ An **interpreter** is a program that takes a source program and executes the pro
![img](img/b.png)
> NOTE: Java uses both. A java source program is compiled into byte code (by a compiler) which is then executed by an interpreter (called java virtual machine - JVM). JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
> NOTE: Java uses both. A Java source program is compiled into bytecode (by a compiler), which is then executed by an interpreter (called the Java virtual machine - JVM). The JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, where as converting IR to executable code is called **back end**.
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, whereas converting IR to executable code is called **back end**.
* The front end focuses on understand the source-language program.
* The back end focuses on mapping programs to the target machine
- The front end focuses on understanding the source-language program.
- The back end focuses on mapping programs to the target machine
![img](img/c.png)
* The front end, intermediate representation and the back end are all part of the compiler.
- The front end, intermediate representation and the back end are all part of the compiler.
IR is stored as an Abstract Syntax Tree **AST**.
The syntactic details needed for parsing the source program are represented in the structure of the tree.
The **IR** could be broken down into many sub-steps i.e. a IR1 could be created which is then ran through an optimiser to create IR2 which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
The **IR** could be broken down into many sub-steps, i.e. an IR1 could be created, which is then run through an optimiser to create IR2, which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
![img](img/d.png)
![img](img/d.png)
+7 -9
View File
@@ -28,7 +28,7 @@ exp -> exp + exp
-> 7 + (10 / 3) * (-2)
```
This grammar is **ambiguous**, this means one input expression could be generated in several different ways.
This grammar is **ambiguous**: this means one input expression could be generated in several different ways.
$$
5 - 4 \times 7
@@ -43,7 +43,7 @@ exp -> exp - exp
-> 5 - 4 * 7
```
However there is another way to derive this expression starting with `*`
However, there is another way to derive this expression starting with `*`
```haskell
exp -> exp * exp
@@ -52,7 +52,7 @@ exp -> exp * exp
-> 5 - 4 * 7
```
These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
![](img/e.png)
@@ -74,7 +74,7 @@ This grammar is unique (non-ambiguous)
## Semantics of Expressions
On the left hand side the $+$ is just a symbol, however on the right hand side it is an arithmetic sum operation.
On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
@@ -99,14 +99,12 @@ $[\![d_0 ]\!] = value(d_0)$
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
## Scanners and Parsers
![img](img/f.png)
Scanners take the source language as input and outputs a stream of tokens.
Scanners take the source language as input and output a stream of tokens.
A **token** is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis.
A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
The **grammar of tokens is always regular**, this means it can be generated and recognised by a DFA (deterministic finite automata).
The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).
@@ -2,8 +2,6 @@
In our parser - there's a lot of repeated code and a lot of cases.
Types of scanner and parser are very similar
```haskell
@@ -38,4 +36,3 @@ parseParenthesis = do symbol '('
symbol ')'
return t
```
+5 -8
View File
@@ -1,6 +1,6 @@
# Functor
Parsing an expression in parenthesis:
Parsing an expression in parentheses:
```haskell
parseP :: Parser AST
@@ -22,9 +22,9 @@ Before we write this sort of code, we need to understand `type classes` (especia
| String | Functor |
| | Monad |
**Eq**: typeclass equality; A type can only be typeclass equality if two like types can be compared
**Eq**: type class for equality; a type can only be in this type class if two values of that type can be compared
A type can be a *member* (instance) of a type class, meaning that if has the properties/functions that the class requires
A type can be a *member* (instance) of a type class, meaning that it has the properties/functions that the class requires
e.g. `Bool` is an instance of `Eq` and `Show`
@@ -61,7 +61,7 @@ newtype Parser a = P (String -> [a, String])
**Parser AST** is a type
Functor is a typeclass of which `parser` is an instance
Functor is a type class of which `parser` is an instance
##### Functor
@@ -95,7 +95,4 @@ fmap id = id -- identity
fmap (f . g) = fmap f . fmap g
```
Haskell doesn't enforce these rules however it is convention.
Haskell doesn't enforce these rules; however, following them is convention.
@@ -28,7 +28,7 @@ fmap2 :: (a -> b -> c) -> f a -> f b -> f c
fmap3 :: (a -> ... n) -> f a -> ... f n
```
`Functor f` can do `fmap1` however cannot do `fmap0` or `fmap2` etc.
`Functor f` can do `fmap1`; however, it cannot do `fmap0` or `fmap2` etc.
**Remember**: `a -> b -> c == a -> (b -> c)`
@@ -92,4 +92,3 @@ All parse does is apply a parser
Where `P` is the constructor
`parse ( P p ) = p`
+4 -13
View File
@@ -22,8 +22,6 @@ intORbin :: Parser Int
expr :: Parser AST
```
```
λ> parse (symbol "something") "nothing"
[]
@@ -59,8 +57,6 @@ instance Functor Parser where
in [(g x, src1)] )
```
```
λ> parse (fmap (+3) integer) "42 blah blah"
[(45, blah blah)]
@@ -75,7 +71,7 @@ instance Functor Parser where
*** Exception Non-exhaustive patterns
```
fixing `fmap`
Fixing `fmap`
```haskell
fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src])
@@ -152,11 +148,9 @@ pf <*> pa = P (\src -> [ (f x, src2) | (f,src1) <- parse pf src,
[(10201, ""), (25, "")]
```
### Monad Class of Parser
Monad class will facilitate the use of `do` notation.
The Monad class will facilitate the use of `do` notation.
```haskell
instance Monad Parser where
@@ -228,7 +222,7 @@ pa >>= fpb = P (\src -> [ r | (x,src1) <- parse pa src,
-- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true
```
What is the do notation and how is it connected to the bind function, we will show this by writing a simple parser
What is the `do` notation and how is it connected to the bind function? We will show this by writing a simple parser
```haskell
pairSum :: Parser Int
@@ -262,8 +256,6 @@ parse (symbol "number" >> integer) "number 9"
NOTE: >> is a non-dependant bind
```
```haskell
the grammer
--funApp ::= ( simpleFun integer )
@@ -402,7 +394,7 @@ string (c:cs) = do char c
[(' ',"hello")]
```
We have to fix leading white space causing failure
We have to fix leading whitespace causing failure
```haskell
space :: Parser ()
@@ -463,4 +455,3 @@ expr = do t1 <- mexpr
<|>
return t1)
```
+8 -8
View File
@@ -1,6 +1,6 @@
# Compiling Variables
A variable is identified by a alphanumeric string. We can store this as a list of pairs, with the variables identifier and its value.
A variable is identified by an alphanumeric string. We can store this as a list of pairs, with the variable's identifier and its value.
Variable Environment or VarEnv - `[(Identifier, Stack Address)]`
@@ -8,7 +8,7 @@ A stack address is an integer value that specifies where in the stack that varia
The bottom of the stack is indexed `0`.
Lets say our environment consists of 3 variables named x,y,z. It would look like:
Let's say our environment consists of 3 variables named x, y, z. It would look like:
`[("z",2), ("y",1), ("x",0)]`
@@ -18,13 +18,13 @@ Lets say our environment consists of 3 variables named x,y,z. It would look like
| y | 2 | 1 |
| z | 9 | 2 |
To get the value of a variable from the stackk, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
To get the value of a variable from the stack, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
`LOAD a` - copy address a to top of stack
`STORE a` - pop top of stack to address a
For example if `LOADL 2` is called, it will effect the stack in the following way:
For example, if `LOADL 2` is called, it will affect the stack in the following way:
| Variables | Stack (Values) | Index |
| :-------: | :------------: | :---: |
@@ -38,7 +38,7 @@ For example if `LOADL 2` is called, it will effect the stack in the following wa
expCode :: VarEnv -> Expr -> [TAMInst]
```
Before we just called the abstract syntax tree `AST` however with the extended grammar now we will have multiple ASTs, one for programs, one for commands, expressions. The AST for expressions we call `Expr`.
Before, we just called the abstract syntax tree `AST`; however, with the extended grammar, we will now have multiple ASTs: one for programs, one for commands and one for expressions. The AST for expressions we call `Expr`.
Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list.
@@ -87,10 +87,10 @@ $$
Example: $s_n$ could be your bank balance and $a_n$ could be the purchase history.
- In our case:
- States are VarEnv & next free address space for next variable
- Outputs are TAM instructions
- States are VarEnv & next free address space for next variable
- Outputs are TAM instructions
We to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
We need to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
```haskell
newtype ST st a = S (\st -> (a, st))
@@ -26,7 +26,7 @@ var z;
var w := x * y - 2
```
The parser will turn this into a list of AST for declarations
The parser will turn this into a list of ASTs for declarations
Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack.
@@ -91,7 +91,7 @@ command ::= identifier := expr
| begin commands end
```
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminal
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminals
```haskell
data Command =
@@ -134,7 +134,7 @@ func :: a -> b
Note file name must start with a capital
When you import a module, can can use functions defined in the module
When you import a module, you can use functions defined in the module
```haskell
data FileType = EXP | TAM
@@ -143,7 +143,7 @@ data Option = Trace | Run | Evaluate
main :: IO () --input output monad
```
this is the entry point, to compile
This is the entry point; to compile:
```shell
$ ghc Main.hs -o aec
@@ -161,4 +161,3 @@ stGet = S (\s -> (s,s))
stRevise :: (st -> st) -> ST st ()
stRevise f = stGet >>= stUpdate . f
```
+2 -2
View File
@@ -2,7 +2,7 @@
**Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output**
Before we could generate a list of instructions to be executed in sequence, now we need to implement code thats conditionally executed or executed multiple times.
Before, we could generate a list of instructions to be executed in sequence; now we need to implement code that's conditionally executed or executed multiple times.
```haskell
--Code for dealing with functions and commands
@@ -71,7 +71,7 @@ JUMPIFZ "label3"
Labels must **always** be **unique**.
This would require a global variable in our compiler to count the number of labels, haskell doesnt not allow global variables.
This would require a global variable in our compiler to count the number of labels; Haskell does not allow global variables.
We can use the `stateMonad` instead.
@@ -4,7 +4,7 @@ You can think of a monad as a container for a data type
If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype
One of the purposes of the `do` notation is to operate on the whole data structure by specify operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
One of the purposes of the `do` notation is to operate on the whole data structure by specifying operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
$$
M_a=\{x_1, x_2, x_3,...\}
@@ -41,7 +41,7 @@ pure x
Monads can have containers within containers
Assume we have function `makeBlob` that maps every element of $a$ to an element of $M_b$
Assume we have a function `makeBlob` that maps every element of $a$ to an element of $M_b$
```haskell
makeBlob :: a -> Mb
@@ -85,4 +85,4 @@ getList = do x <- getInt
else do
xs <- getList
return (x:xs)
```
```