Tidy up
This commit is contained in:
103 files changed
+3663
-3779
No files matched your search
@@ -1,6 +1,6 @@
|
||||
# Compilers - COMP 3012
|
||||
|
||||
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to program written in a target programming language. A compiler is written in an **implementation language**.
|
||||
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to a program written in a target programming language. A compiler is written in an **implementation language**.
|
||||
|
||||

|
||||
|
||||
@@ -10,21 +10,21 @@ An **interpreter** is a program that takes a source program and executes the pro
|
||||
|
||||

|
||||
|
||||
> NOTE: Java uses both. A java source program is compiled into byte code (by a compiler) which is then executed by an interpreter (called java virtual machine - JVM). JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
|
||||
> NOTE: Java uses both. A Java source program is compiled into bytecode (by a compiler), which is then executed by an interpreter (called the Java virtual machine - JVM). The JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
|
||||
|
||||
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, where as converting IR to executable code is called **back end**.
|
||||
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, whereas converting IR to executable code is called **back end**.
|
||||
|
||||
* The front end focuses on understand the source-language program.
|
||||
* The back end focuses on mapping programs to the target machine
|
||||
- The front end focuses on understanding the source-language program.
|
||||
- The back end focuses on mapping programs to the target machine
|
||||
|
||||

|
||||
|
||||
* The front end, intermediate representation and the back end are all part of the compiler.
|
||||
- The front end, intermediate representation and the back end are all part of the compiler.
|
||||
|
||||
IR is stored as an Abstract Syntax Tree **AST**.
|
||||
|
||||
The syntactic details needed for parsing the source program are represented in the structure of the tree.
|
||||
|
||||
The **IR** could be broken down into many sub-steps i.e. a IR1 could be created which is then ran through an optimiser to create IR2 which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
|
||||
The **IR** could be broken down into many sub-steps, i.e. an IR1 could be created, which is then run through an optimiser to create IR2, which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
|
||||
|
||||

|
||||

|
||||
@@ -28,7 +28,7 @@ exp -> exp + exp
|
||||
-> 7 + (10 / 3) * (-2)
|
||||
```
|
||||
|
||||
This grammar is **ambiguous**, this means one input expression could be generated in several different ways.
|
||||
This grammar is **ambiguous**: this means one input expression could be generated in several different ways.
|
||||
|
||||
$$
|
||||
5 - 4 \times 7
|
||||
@@ -43,7 +43,7 @@ exp -> exp - exp
|
||||
-> 5 - 4 * 7
|
||||
```
|
||||
|
||||
However there is another way to derive this expression starting with `*`
|
||||
However, there is another way to derive this expression starting with `*`
|
||||
|
||||
```haskell
|
||||
exp -> exp * exp
|
||||
@@ -52,7 +52,7 @@ exp -> exp * exp
|
||||
-> 5 - 4 * 7
|
||||
```
|
||||
|
||||
These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
|
||||
These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
|
||||
|
||||

|
||||
|
||||
@@ -74,7 +74,7 @@ This grammar is unique (non-ambiguous)
|
||||
|
||||
## Semantics of Expressions
|
||||
|
||||
On the left hand side the $+$ is just a symbol, however on the right hand side it is an arithmetic sum operation.
|
||||
On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
|
||||
|
||||
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
|
||||
|
||||
@@ -99,14 +99,12 @@ $[\![d_0 ]\!] = value(d_0)$
|
||||
|
||||
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
|
||||
|
||||
|
||||
|
||||
## Scanners and Parsers
|
||||
|
||||

|
||||
|
||||
Scanners take the source language as input and outputs a stream of tokens.
|
||||
Scanners take the source language as input and output a stream of tokens.
|
||||
|
||||
A **token** is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis.
|
||||
A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
|
||||
|
||||
The **grammar of tokens is always regular**, this means it can be generated and recognised by a DFA (deterministic finite automata).
|
||||
The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).
|
||||
@@ -2,8 +2,6 @@
|
||||
|
||||
In our parser - there's a lot of repeated code and a lot of cases.
|
||||
|
||||
|
||||
|
||||
Types of scanner and parser are very similar
|
||||
|
||||
```haskell
|
||||
@@ -38,4 +36,3 @@ parseParenthesis = do symbol '('
|
||||
symbol ')'
|
||||
return t
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Functor
|
||||
|
||||
Parsing an expression in parenthesis:
|
||||
Parsing an expression in parentheses:
|
||||
|
||||
```haskell
|
||||
parseP :: Parser AST
|
||||
@@ -22,9 +22,9 @@ Before we write this sort of code, we need to understand `type classes` (especia
|
||||
| String | Functor |
|
||||
| | Monad |
|
||||
|
||||
**Eq**: typeclass equality; A type can only be typeclass equality if two like types can be compared
|
||||
**Eq**: type class for equality; a type can only be in this type class if two values of that type can be compared
|
||||
|
||||
A type can be a *member* (instance) of a type class, meaning that if has the properties/functions that the class requires
|
||||
A type can be a *member* (instance) of a type class, meaning that it has the properties/functions that the class requires
|
||||
|
||||
e.g. `Bool` is an instance of `Eq` and `Show`
|
||||
|
||||
@@ -61,7 +61,7 @@ newtype Parser a = P (String -> [a, String])
|
||||
|
||||
**Parser AST** is a type
|
||||
|
||||
Functor is a typeclass of which `parser` is an instance
|
||||
Functor is a type class of which `parser` is an instance
|
||||
|
||||
##### Functor
|
||||
|
||||
@@ -95,7 +95,4 @@ fmap id = id -- identity
|
||||
fmap (f . g) = fmap f . fmap g
|
||||
```
|
||||
|
||||
Haskell doesn't enforce these rules however it is convention.
|
||||
|
||||
|
||||
|
||||
Haskell doesn't enforce these rules; however, following them is convention.
|
||||
@@ -28,7 +28,7 @@ fmap2 :: (a -> b -> c) -> f a -> f b -> f c
|
||||
fmap3 :: (a -> ... n) -> f a -> ... f n
|
||||
```
|
||||
|
||||
`Functor f` can do `fmap1` however cannot do `fmap0` or `fmap2` etc.
|
||||
`Functor f` can do `fmap1`; however, it cannot do `fmap0` or `fmap2` etc.
|
||||
|
||||
**Remember**: `a -> b -> c == a -> (b -> c)`
|
||||
|
||||
@@ -92,4 +92,3 @@ All parse does is apply a parser
|
||||
Where `P` is the constructor
|
||||
|
||||
`parse ( P p ) = p`
|
||||
|
||||
@@ -22,8 +22,6 @@ intORbin :: Parser Int
|
||||
expr :: Parser AST
|
||||
```
|
||||
|
||||
|
||||
|
||||
```
|
||||
λ> parse (symbol "something") "nothing"
|
||||
[]
|
||||
@@ -59,8 +57,6 @@ instance Functor Parser where
|
||||
in [(g x, src1)] )
|
||||
```
|
||||
|
||||
|
||||
|
||||
```
|
||||
λ> parse (fmap (+3) integer) "42 blah blah"
|
||||
[(45, blah blah)]
|
||||
@@ -75,7 +71,7 @@ instance Functor Parser where
|
||||
*** Exception Non-exhaustive patterns
|
||||
```
|
||||
|
||||
fixing `fmap`
|
||||
Fixing `fmap`
|
||||
|
||||
```haskell
|
||||
fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src])
|
||||
@@ -152,11 +148,9 @@ pf <*> pa = P (\src -> [ (f x, src2) | (f,src1) <- parse pf src,
|
||||
[(10201, ""), (25, "")]
|
||||
```
|
||||
|
||||
|
||||
|
||||
### Monad Class of Parser
|
||||
|
||||
Monad class will facilitate the use of `do` notation.
|
||||
The Monad class will facilitate the use of `do` notation.
|
||||
|
||||
```haskell
|
||||
instance Monad Parser where
|
||||
@@ -228,7 +222,7 @@ pa >>= fpb = P (\src -> [ r | (x,src1) <- parse pa src,
|
||||
-- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true
|
||||
```
|
||||
|
||||
What is the do notation and how is it connected to the bind function, we will show this by writing a simple parser
|
||||
What is the `do` notation and how is it connected to the bind function? We will show this by writing a simple parser
|
||||
|
||||
```haskell
|
||||
pairSum :: Parser Int
|
||||
@@ -262,8 +256,6 @@ parse (symbol "number" >> integer) "number 9"
|
||||
NOTE: >> is a non-dependant bind
|
||||
```
|
||||
|
||||
|
||||
|
||||
```haskell
|
||||
the grammer
|
||||
--funApp ::= ( simpleFun integer )
|
||||
@@ -402,7 +394,7 @@ string (c:cs) = do char c
|
||||
[(' ',"hello")]
|
||||
```
|
||||
|
||||
We have to fix leading white space causing failure
|
||||
We have to fix leading whitespace causing failure
|
||||
|
||||
```haskell
|
||||
space :: Parser ()
|
||||
@@ -463,4 +455,3 @@ expr = do t1 <- mexpr
|
||||
<|>
|
||||
return t1)
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Compiling Variables
|
||||
|
||||
A variable is identified by a alphanumeric string. We can store this as a list of pairs, with the variables identifier and its value.
|
||||
A variable is identified by an alphanumeric string. We can store this as a list of pairs, with the variable's identifier and its value.
|
||||
|
||||
Variable Environment or VarEnv - `[(Identifier, Stack Address)]`
|
||||
|
||||
@@ -8,7 +8,7 @@ A stack address is an integer value that specifies where in the stack that varia
|
||||
|
||||
The bottom of the stack is indexed `0`.
|
||||
|
||||
Lets say our environment consists of 3 variables named x,y,z. It would look like:
|
||||
Let's say our environment consists of 3 variables named x, y, z. It would look like:
|
||||
|
||||
`[("z",2), ("y",1), ("x",0)]`
|
||||
|
||||
@@ -18,13 +18,13 @@ Lets say our environment consists of 3 variables named x,y,z. It would look like
|
||||
| y | 2 | 1 |
|
||||
| z | 9 | 2 |
|
||||
|
||||
To get the value of a variable from the stackk, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
|
||||
To get the value of a variable from the stack, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
|
||||
|
||||
`LOAD a` - copy address a to top of stack
|
||||
|
||||
`STORE a` - pop top of stack to address a
|
||||
|
||||
For example if `LOADL 2` is called, it will effect the stack in the following way:
|
||||
For example, if `LOADL 2` is called, it will affect the stack in the following way:
|
||||
|
||||
| Variables | Stack (Values) | Index |
|
||||
| :-------: | :------------: | :---: |
|
||||
@@ -38,7 +38,7 @@ For example if `LOADL 2` is called, it will effect the stack in the following wa
|
||||
expCode :: VarEnv -> Expr -> [TAMInst]
|
||||
```
|
||||
|
||||
Before we just called the abstract syntax tree `AST` however with the extended grammar now we will have multiple ASTs, one for programs, one for commands, expressions. The AST for expressions we call `Expr`.
|
||||
Before, we just called the abstract syntax tree `AST`; however, with the extended grammar, we will now have multiple ASTs: one for programs, one for commands and one for expressions. The AST for expressions we call `Expr`.
|
||||
|
||||
Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list.
|
||||
|
||||
@@ -87,10 +87,10 @@ $$
|
||||
Example: $s_n$ could be your bank balance and $a_n$ could be the purchase history.
|
||||
|
||||
- In our case:
|
||||
- States are VarEnv & next free address space for next variable
|
||||
- Outputs are TAM instructions
|
||||
- States are VarEnv & next free address space for next variable
|
||||
- Outputs are TAM instructions
|
||||
|
||||
We to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
|
||||
We need to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
|
||||
|
||||
```haskell
|
||||
newtype ST st a = S (\st -> (a, st))
|
||||
|
||||
@@ -26,7 +26,7 @@ var z;
|
||||
var w := x * y - 2
|
||||
```
|
||||
|
||||
The parser will turn this into a list of AST for declarations
|
||||
The parser will turn this into a list of ASTs for declarations
|
||||
|
||||
Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack.
|
||||
|
||||
@@ -91,7 +91,7 @@ command ::= identifier := expr
|
||||
| begin commands end
|
||||
```
|
||||
|
||||
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminal
|
||||
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminals
|
||||
|
||||
```haskell
|
||||
data Command =
|
||||
@@ -134,7 +134,7 @@ func :: a -> b
|
||||
|
||||
Note file name must start with a capital
|
||||
|
||||
When you import a module, can can use functions defined in the module
|
||||
When you import a module, you can use functions defined in the module
|
||||
|
||||
```haskell
|
||||
data FileType = EXP | TAM
|
||||
@@ -143,7 +143,7 @@ data Option = Trace | Run | Evaluate
|
||||
main :: IO () --input output monad
|
||||
```
|
||||
|
||||
this is the entry point, to compile
|
||||
This is the entry point; to compile:
|
||||
|
||||
```shell
|
||||
$ ghc Main.hs -o aec
|
||||
@@ -161,4 +161,3 @@ stGet = S (\s -> (s,s))
|
||||
stRevise :: (st -> st) -> ST st ()
|
||||
stRevise f = stGet >>= stUpdate . f
|
||||
```
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output**
|
||||
|
||||
Before we could generate a list of instructions to be executed in sequence, now we need to implement code thats conditionally executed or executed multiple times.
|
||||
Before, we could generate a list of instructions to be executed in sequence; now we need to implement code that's conditionally executed or executed multiple times.
|
||||
|
||||
```haskell
|
||||
--Code for dealing with functions and commands
|
||||
@@ -71,7 +71,7 @@ JUMPIFZ "label3"
|
||||
|
||||
Labels must **always** be **unique**.
|
||||
|
||||
This would require a global variable in our compiler to count the number of labels, haskell doesnt not allow global variables.
|
||||
This would require a global variable in our compiler to count the number of labels; Haskell does not allow global variables.
|
||||
|
||||
We can use the `stateMonad` instead.
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ You can think of a monad as a container for a data type
|
||||
|
||||
If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype
|
||||
|
||||
One of the purposes of the `do` notation is to operate on the whole data structure by specify operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
|
||||
One of the purposes of the `do` notation is to operate on the whole data structure by specifying operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
|
||||
|
||||
$$
|
||||
M_a=\{x_1, x_2, x_3,...\}
|
||||
@@ -41,7 +41,7 @@ pure x
|
||||
|
||||
Monads can have containers within containers
|
||||
|
||||
Assume we have function `makeBlob` that maps every element of $a$ to an element of $M_b$
|
||||
Assume we have a function `makeBlob` that maps every element of $a$ to an element of $M_b$
|
||||
|
||||
```haskell
|
||||
makeBlob :: a -> Mb
|
||||
@@ -85,4 +85,4 @@ getList = do x <- getInt
|
||||
else do
|
||||
xs <- getList
|
||||
return (x:xs)
|
||||
```
|
||||
```
|
||||
Reference in new issue
Block a user