Add the rest of university notes
No files matched your search
@@ -1,3 +1,4 @@
|
|||||||
site/
|
site/
|
||||||
venv
|
venv
|
||||||
.venv
|
.venv
|
||||||
|
.idea
|
||||||
@@ -230,6 +230,7 @@ function getSharks() {
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
In relational algebra, you would instead write:
|
In relational algebra, you would instead write:
|
||||||
|
|
||||||
$$
|
$$
|
||||||
sharks = \sigma_{family =''Sharks''} (animals)
|
sharks = \sigma_{family =''Sharks''} (animals)
|
||||||
$$
|
$$
|
||||||
|
|||||||
@@ -0,0 +1,27 @@
|
|||||||
|
window.MathJax = {
|
||||||
|
tex: {
|
||||||
|
inlineMath: [["\\(", "\\)"]],
|
||||||
|
displayMath: [["\\[", "\\]"]],
|
||||||
|
processEscapes: true,
|
||||||
|
processEnvironments: true
|
||||||
|
},
|
||||||
|
options: {
|
||||||
|
ignoreHtmlClass: ".*|",
|
||||||
|
processHtmlClass: "arithmatex"
|
||||||
|
},
|
||||||
|
startup: {
|
||||||
|
typeset: false,
|
||||||
|
ready() {
|
||||||
|
MathJax.startup.defaultReady();
|
||||||
|
// Subscribe only after MathJax is ready, including on instant navigation.
|
||||||
|
MathJax.startup.promise.then(() => {
|
||||||
|
document$.subscribe(() => {
|
||||||
|
MathJax.startup.output.clearCache();
|
||||||
|
MathJax.typesetClear();
|
||||||
|
MathJax.texReset();
|
||||||
|
MathJax.typesetPromise();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
};
|
||||||
@@ -26,7 +26,7 @@ They have two parts: physical part and a social part
|
|||||||
|
|
||||||
Social structures are vital for these networks - think covid tracking networks
|
Social structures are vital for these networks - think covid tracking networks
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
Clouds have multiple layers
|
Clouds have multiple layers
|
||||||
|
|
||||||
@@ -57,11 +57,11 @@ This can be used to exchange warning and beacon messages via V2V (vehicle to veh
|
|||||||
|
|
||||||
### Fully autonomous Vehicles
|
### Fully autonomous Vehicles
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
Vehicles can connect to the cloud and share & request information to help other vehicles.
|
Vehicles can connect to the cloud and share & request information to help other vehicles.
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
An example of transient clouds - in this case vehicular clouds.
|
An example of transient clouds - in this case vehicular clouds.
|
||||||
|
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ One of the core features of a MANET node is the ability to autonomously connect
|
|||||||
* Typically routing is split into **route discovery** and **actual data transmission**.
|
* Typically routing is split into **route discovery** and **actual data transmission**.
|
||||||
* Nodes have to self organise in order to route.
|
* Nodes have to self organise in order to route.
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
(green boxes is route chosen)
|
(green boxes is route chosen)
|
||||||
|
|
||||||
@@ -43,7 +43,7 @@ The source has a limited range of nodes it can detect, it cannot send it direct
|
|||||||
|
|
||||||
Table showing all different protocols of MANETs
|
Table showing all different protocols of MANETs
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
### Delay/Disconnection Tolerance
|
### Delay/Disconnection Tolerance
|
||||||
|
|
||||||
|
|||||||
@@ -77,7 +77,7 @@ Communication is made possible in the network when intermediate nodes become **c
|
|||||||
* They allow mobile nodes that pass by to collect and leave data on them.
|
* They allow mobile nodes that pass by to collect and leave data on them.
|
||||||
* They contribute to increasing the frequency of node contacts and improve **delivery ratio** and **delivery delay**.
|
* They contribute to increasing the frequency of node contacts and improve **delivery ratio** and **delivery delay**.
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
## Categories of VANETs
|
## Categories of VANETs
|
||||||
|
|
||||||
|
|||||||
@@ -10,7 +10,7 @@
|
|||||||
* **The focus phase allows** each node to forward a copy of its messages to other potential nodes until the messages gets to its destination.
|
* **The focus phase allows** each node to forward a copy of its messages to other potential nodes until the messages gets to its destination.
|
||||||
* The protocol uses a single-copy utility based routing scheme to forward a copy of the message further.
|
* The protocol uses a single-copy utility based routing scheme to forward a copy of the message further.
|
||||||
* Forwarding decisions are made based on **timers** which record the times nodes come in communication range of each other.
|
* Forwarding decisions are made based on **timers** which record the times nodes come in communication range of each other.
|
||||||
* Node $$A$$ forwards message with destination $$D$$ to node $$B$$ , **if and only if** $$B$$ has a higher potential of delivering the message to $$D$$.
|
* Node $A$ forwards message with destination $D$ to node $B$ , **if and only if** $B$ has a higher potential of delivering the message to $D$.
|
||||||
|
|
||||||
#### SimBet
|
#### SimBet
|
||||||
|
|
||||||
|
|||||||
@@ -20,7 +20,7 @@ When deciding on the best carrier and the optimal number of messages, CAFREP dyn
|
|||||||
2. Predictive **node congestion** (node storage and in-network delays)
|
2. Predictive **node congestion** (node storage and in-network delays)
|
||||||
3. Predictive **ego network congestion**
|
3. Predictive **ego network congestion**
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
Each layer you go up, the more information is exchanged between the nodes.
|
Each layer you go up, the more information is exchanged between the nodes.
|
||||||
|
|
||||||
@@ -34,7 +34,7 @@ $$
|
|||||||
Ret(X) = B_c(X) - \sum^N_{i=1} \space M^i_{size}(X)
|
Ret(X) = B_c(X) - \sum^N_{i=1} \space M^i_{size}(X)
|
||||||
$$
|
$$
|
||||||
|
|
||||||
For a node $$X$$, it has buffer of size $$B_c(X)$$. When a message of size $$M^i_{size}$$ is sent to node $$X$$, it's buffer size is the total buffer minus the memory taken by the sum of all messages in the buffer.
|
For a node $X$, it has buffer of size $B_c(X)$. When a message of size $M^i_{size}$ is sent to node $X$, it's buffer size is the total buffer minus the memory taken by the sum of all messages in the buffer.
|
||||||
|
|
||||||
###### Node Receptiveness
|
###### Node Receptiveness
|
||||||
|
|
||||||
@@ -66,7 +66,7 @@ $$
|
|||||||
EN_{Ret}(X) = \frac{1}{N}\sum^N_{i=1}Ret(C_i(X))
|
EN_{Ret}(X) = \frac{1}{N}\sum^N_{i=1}Ret(C_i(X))
|
||||||
$$
|
$$
|
||||||
|
|
||||||
Gets the average of the retentiveness of node $$X$$ and it's neighbours $$c_i(X)$$
|
Gets the average of the retentiveness of node $X$ and it's neighbours $c_i(X)$
|
||||||
|
|
||||||
###### Ego Network Receptiveness
|
###### Ego Network Receptiveness
|
||||||
|
|
||||||
@@ -88,9 +88,11 @@ $$
|
|||||||
#### Contents of CAFREP Node
|
#### Contents of CAFREP Node
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
$$
|
$$
|
||||||
Replication\space rate = M \times \frac{TotalUtil(Y)}{TotalUtil(X) + TotalUtil(Y)}
|
Replication\space rate = M \times \frac{TotalUtil(Y)}{TotalUtil(X) + TotalUtil(Y)}
|
||||||
$$
|
$$
|
||||||
|
|
||||||
Total utility, changes constantly. The replication limit grows to take advantage of all available resources, and backs off when congestion increases.
|
Total utility, changes constantly. The replication limit grows to take advantage of all available resources, and backs off when congestion increases.
|
||||||
|
|
||||||
Social utility prevents replication at a high rate on free nodes that are not on the path to the destination.
|
Social utility prevents replication at a high rate on free nodes that are not on the path to the destination.
|
||||||
@@ -13,7 +13,7 @@
|
|||||||
* Application and content providers are independent of each other
|
* Application and content providers are independent of each other
|
||||||
* CDNs focus on web content distributions for major players
|
* CDNs focus on web content distributions for major players
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
**Important requirements for ICNs** (Information Centric Networks)
|
**Important requirements for ICNs** (Information Centric Networks)
|
||||||
|
|
||||||
@@ -79,7 +79,7 @@ Apart from routing protocols that use direct identifiers of nodes, networking ca
|
|||||||
|
|
||||||
##### Using Names in CCNs (Content Centric Networks)
|
##### Using Names in CCNs (Content Centric Networks)
|
||||||
|
|
||||||
- The hierarchical structure is used to do *longest match look-ups* which guarantees $$log(n)$$ state scaling for globally accessible data.
|
- The hierarchical structure is used to do *longest match look-ups* which guarantees $log(n)$ state scaling for globally accessible data.
|
||||||
- Although CCN names are longer than IP identifiers, their **explicit structure** allows look-ups as efficient as IP's.
|
- Although CCN names are longer than IP identifiers, their **explicit structure** allows look-ups as efficient as IP's.
|
||||||
|
|
||||||
### ICN Forwarding
|
### ICN Forwarding
|
||||||
|
|||||||
@@ -7,7 +7,7 @@ A Brief History of Networking
|
|||||||
- Wires are the dominant cost.
|
- Wires are the dominant cost.
|
||||||
- A *call* is not the conversation, its the **PATH** between two end-office line cards.
|
- A *call* is not the conversation, its the **PATH** between two end-office line cards.
|
||||||
- A *phone number* is not the name/address of the caller, its a **program** for the end-office switch fabric to build a path to the destination line card.
|
- A *phone number* is not the name/address of the caller, its a **program** for the end-office switch fabric to build a path to the destination line card.
|
||||||
- <img src="/lectures/acn/img/k.png" alt="switch board" style="zoom:50%;" />
|
- <img src="img/k.png" alt="switch board" style="zoom:50%;" />
|
||||||
- Path building is **non-local** and **encourages centralisation** and **monopoly**.
|
- Path building is **non-local** and **encourages centralisation** and **monopoly**.
|
||||||
- Calls fail is any element in the path fails so reliability goes down exponentially as the system scales up.
|
- Calls fail is any element in the path fails so reliability goes down exponentially as the system scales up.
|
||||||
- Data cannot flow until the path is set up so efficiency decreases with setup time.
|
- Data cannot flow until the path is set up so efficiency decreases with setup time.
|
||||||
@@ -49,7 +49,7 @@ CCN can run over and be run over anything e.g. IP.
|
|||||||
|
|
||||||
#### CCN Packets
|
#### CCN Packets
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
**Interest** - similar to HTTP `GET`
|
**Interest** - similar to HTTP `GET`
|
||||||
|
|
||||||
@@ -59,7 +59,7 @@ CCN can run over and be run over anything e.g. IP.
|
|||||||
|
|
||||||
Data packets are authenticated with digital signatures.
|
Data packets are authenticated with digital signatures.
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
#### CCN Forwarding
|
#### CCN Forwarding
|
||||||
|
|
||||||
@@ -91,6 +91,6 @@ In the current Internet, Quality of Service (QoS) Problems are highly localised
|
|||||||
|
|
||||||
Unlike IP, CCN is **local**, don't have queues and receivers have complete control
|
Unlike IP, CCN is **local**, don't have queues and receivers have complete control
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
Tree serves as transport state
|
Tree serves as transport state
|
||||||
@@ -4,7 +4,7 @@
|
|||||||
|
|
||||||
##### Interplanetary communication
|
##### Interplanetary communication
|
||||||
|
|
||||||
<img src="/lectures/acn/img/o.png" alt="DTN in space" style="zoom:50%;" />
|
<img src="img/o.png" alt="DTN in space" style="zoom:50%;" />
|
||||||
|
|
||||||
> **Characteristics**
|
> **Characteristics**
|
||||||
>
|
>
|
||||||
@@ -52,7 +52,7 @@
|
|||||||
>- High propagation delay
|
>- High propagation delay
|
||||||
>- Asymmetric data rate
|
>- Asymmetric data rate
|
||||||
>
|
>
|
||||||
>
|
>
|
||||||
>
|
>
|
||||||
>**Security**
|
>**Security**
|
||||||
>
|
>
|
||||||
@@ -127,7 +127,7 @@ Based on the *bundle* protocol
|
|||||||
* Access Control (only legit users with right permissions)
|
* Access Control (only legit users with right permissions)
|
||||||
* Limited protection from DoS attacks
|
* Limited protection from DoS attacks
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
- Payload Security Header is computed once at the source bundle agent, carried unchanged, and checked at the destination bundle agent (and possibly also security boundary bundle agents)
|
- Payload Security Header is computed once at the source bundle agent, carried unchanged, and checked at the destination bundle agent (and possibly also security boundary bundle agents)
|
||||||
|
|
||||||
|
|||||||
@@ -102,7 +102,7 @@ Precedence
|
|||||||
|
|
||||||
- 95% allocated already (440,000 netblocks)
|
- 95% allocated already (440,000 netblocks)
|
||||||
|
|
||||||
**IPv6** supports 128 bit address
|
**IPv6** supports 128-bit address
|
||||||
|
|
||||||
- Loads of addresses :white_check_mark:
|
- Loads of addresses :white_check_mark:
|
||||||
- Routing protocols need to ported :negative_squared_cross_mark:
|
- Routing protocols need to ported :negative_squared_cross_mark:
|
||||||
@@ -132,7 +132,6 @@ Because IPv6 did not magically solve address shortage problem and not all router
|
|||||||
|
|
||||||
###### Full Cone
|
###### Full Cone
|
||||||
|
|
||||||

|
|
||||||
```
|
```
|
||||||
ea:ep - NAT address : NAT port
|
ea:ep - NAT address : NAT port
|
||||||
```
|
```
|
||||||
@@ -141,17 +140,12 @@ When client receives packet from server 1 `da:dp`, the NAT translates the NAT ad
|
|||||||
|
|
||||||
###### Address Restricted Cone NAT
|
###### Address Restricted Cone NAT
|
||||||
|
|
||||||

|
|
||||||
In this case server 2 is not trusted and therefore any request will be dropped.
|
In this case server 2 is not trusted and therefore any request will be dropped.
|
||||||
|
|
||||||
###### Port Restricted Cone NAT
|
###### Port Restricted Cone NAT
|
||||||
|
|
||||||

|
|
||||||
|
|
||||||
If the router receives a packet from a bad IP or bad port, it will be dropped.
|
If the router receives a packet from a bad IP or bad port, it will be dropped.
|
||||||
|
|
||||||
###### Symmetric NAT
|
###### Symmetric NAT
|
||||||
|
|
||||||

|
|
||||||
|
|
||||||
Here the internal address is obfuscated from the external servers, same client can use different ports for different communications.
|
Here the internal address is obfuscated from the external servers, same client can use different ports for different communications.
|
||||||
@@ -41,7 +41,7 @@ DNS is a consistent namespace
|
|||||||
- Extract information from tree upon client requests
|
- Extract information from tree upon client requests
|
||||||
- `gethostbyname()`
|
- `gethostbyname()`
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
###### Root
|
###### Root
|
||||||
|
|
||||||
@@ -121,7 +121,7 @@ What happens when the resolver queries a server that doesn't know the answer? tw
|
|||||||
1. **Recursive** (optional)
|
1. **Recursive** (optional)
|
||||||
- Server generates a new query to the next server
|
- Server generates a new query to the next server
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
#### Load Balancing
|
#### Load Balancing
|
||||||
|
|
||||||
|
|||||||
@@ -14,7 +14,7 @@ Simplest possible paradigm
|
|||||||
- Wait for `ack(x)`
|
- Wait for `ack(x)`
|
||||||
- Transmit `seq(x+1)`
|
- Transmit `seq(x+1)`
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
This has really poor performance in high latency and uses high bandwidth (half the bandwidth is overhead (acknowledgements))
|
This has really poor performance in high latency and uses high bandwidth (half the bandwidth is overhead (acknowledgements))
|
||||||
|
|
||||||
@@ -88,12 +88,12 @@ RTT - round trip times
|
|||||||
- When the first `RTT` measurement is taken the sender sets the smoothed `RTT` (`SRTT`), `RTT` variance (`RTTVAR`) and `TIMEOUT` in the following way
|
- When the first `RTT` measurement is taken the sender sets the smoothed `RTT` (`SRTT`), `RTT` variance (`RTTVAR`) and `TIMEOUT` in the following way
|
||||||
- `SRTT = RTT`
|
- `SRTT = RTT`
|
||||||
- `RTTVAR = RTT/2`
|
- `RTTVAR = RTT/2`
|
||||||
- `TIMEOUT = `$\Mu\cdot$`SRTT + 4*RTTVAR`
|
- `TIMEOUT = `$\mu\cdot$`SRTT + 4*RTTVAR`
|
||||||
- Where $\Mu$ is a constant, which in this implementation is 1.08 (obtained experimentally)
|
- Where $\mu$ is a constant, which in this implementation is 1.08 (obtained experimentally)
|
||||||
- When subsequent `RTT` measurements are made the sender sets the `RTTVAR`, `SRTT`, TIMEOUT
|
- When subsequent `RTT` measurements are made the sender sets the `RTTVAR`, `SRTT`, TIMEOUT
|
||||||
- `RTTVAR`$= (1 - \frac{1}{4}) \times$`RTTVAR`$+ \frac14 \times |$`SRTT`$-$`RTT`$|$
|
- `RTTVAR`$= (1 - \frac{1}{4}) \times$`RTTVAR`$+ \frac14 \times |$`SRTT`$-$`RTT`$|$
|
||||||
- `SRTT`$= (-\frac18)\times$`SRTT`$+\frac18\times$`RTT`
|
- `SRTT`$= (-\frac18)\times$`SRTT`$+\frac18\times$`RTT`
|
||||||
- `TIMEOUT`$= \Mu\times$`SRTT`$+ 4\times$`RTTVAR`
|
- `TIMEOUT`$= \mu\times$`SRTT`$+ 4\times$`RTTVAR`
|
||||||
|
|
||||||
###### Packet loss rate calculation
|
###### Packet loss rate calculation
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,30 @@
|
|||||||
|
# Compilers - COMP 3012
|
||||||
|
|
||||||
|
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to program written in a target programming language. A compiler is written in an **implementation language**.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Notation of semantics of program $A$: $[\![A]\!]$
|
||||||
|
|
||||||
|
An **interpreter** is a program that takes a source program and executes the program.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
> NOTE: Java uses both. A java source program is compiled into byte code (by a compiler) which is then executed by an interpreter (called java virtual machine - JVM). JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
|
||||||
|
|
||||||
|
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, where as converting IR to executable code is called **back end**.
|
||||||
|
|
||||||
|
* The front end focuses on understand the source-language program.
|
||||||
|
* The back end focuses on mapping programs to the target machine
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
* The front end, intermediate representation and the back end are all part of the compiler.
|
||||||
|
|
||||||
|
IR is stored as an Abstract Syntax Tree **AST**.
|
||||||
|
|
||||||
|
The syntactic details needed for parsing the source program are represented in the structure of the tree.
|
||||||
|
|
||||||
|
The **IR** could be broken down into many sub-steps i.e. a IR1 could be created which is then ran through an optimiser to create IR2 which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
|
||||||
|
|
||||||
|

|
||||||
@@ -0,0 +1,112 @@
|
|||||||
|
# Arithmetic Grammar
|
||||||
|
|
||||||
|
## Syntax of Expressions
|
||||||
|
|
||||||
|
An expression can be defined as:
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
exp ::= int | exp + exp | exp - exp | exp * exp
|
||||||
|
| exp / exp | - exp | ( exp )
|
||||||
|
```
|
||||||
|
|
||||||
|
$$
|
||||||
|
7 + (10/3) \times (-2)
|
||||||
|
$$
|
||||||
|
|
||||||
|
Applying this to the above expression:
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
exp -> exp + exp
|
||||||
|
-> int + exp
|
||||||
|
-> 7 + exp
|
||||||
|
-> 7 + exp * exp
|
||||||
|
-> 7 + (exp) * exp
|
||||||
|
-> 7 + (exp / exp) * exp
|
||||||
|
-> 7 + (int / int) * exp
|
||||||
|
-> 7 + (10 / 3) * (-exp)
|
||||||
|
-> 7 + (10 / 3) * (-int)
|
||||||
|
-> 7 + (10 / 3) * (-2)
|
||||||
|
```
|
||||||
|
|
||||||
|
This grammar is **ambiguous**, this means one input expression could be generated in several different ways.
|
||||||
|
|
||||||
|
$$
|
||||||
|
5 - 4 \times 7
|
||||||
|
$$
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
exp -> exp - exp
|
||||||
|
-> int - exp
|
||||||
|
-> 5 - exp
|
||||||
|
-> 5 - exp * exp
|
||||||
|
...
|
||||||
|
-> 5 - 4 * 7
|
||||||
|
```
|
||||||
|
|
||||||
|
However there is another way to derive this expression starting with `*`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
exp -> exp * exp
|
||||||
|
-> exp - exp * exp
|
||||||
|
...
|
||||||
|
-> 5 - 4 * 7
|
||||||
|
```
|
||||||
|
|
||||||
|
These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
```haskell
|
||||||
|
exp ::= mexp | mexp + exp | mexp - exp
|
||||||
|
mexp ::= term | term * mexp | term / mexp --multiplicative expression
|
||||||
|
term ::= int | - term | ( exp )
|
||||||
|
|
||||||
|
exp -> mexp - exp
|
||||||
|
-> term - exp
|
||||||
|
-> int - exp
|
||||||
|
-> 5 - exp
|
||||||
|
-> 5 - mexp
|
||||||
|
-> 5 - term * mexp -> 5 - int * mexp -> 5 - 4 * mexp
|
||||||
|
-> 5 - 4 * term -> 5 - 4 * int -> 5 - 4 * 7
|
||||||
|
```
|
||||||
|
|
||||||
|
This grammar is unique (non-ambiguous)
|
||||||
|
|
||||||
|
## Semantics of Expressions
|
||||||
|
|
||||||
|
On the left hand side the $+$ is just a symbol, however on the right hand side it is an arithmetic sum operation.
|
||||||
|
|
||||||
|
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
|
||||||
|
|
||||||
|
$[\![ -exp]\!] = - [\![exp ]\!]$
|
||||||
|
|
||||||
|
$[\![ x]\!] = x$
|
||||||
|
|
||||||
|
$[\![(exp) ]\!] = [\![exp ]\!]$ - This is because parentheses change order of operations, not the operation itself.
|
||||||
|
|
||||||
|
Addition can be rewritten:
|
||||||
|
|
||||||
|
$$
|
||||||
|
[\![exp_1 + exp_2 ]\!] = +([\![exp_1 ]\!], [\![exp_2 ]\!])
|
||||||
|
$$
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
int ::= digit | int digit
|
||||||
|
digit ::= 0 | 1 | 2 | 3 ... | 9
|
||||||
|
```
|
||||||
|
|
||||||
|
$[\![d_0 ]\!] = value(d_0)$
|
||||||
|
|
||||||
|
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
## Scanners and Parsers
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Scanners take the source language as input and outputs a stream of tokens.
|
||||||
|
|
||||||
|
A **token** is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis.
|
||||||
|
|
||||||
|
The **grammar of tokens is always regular**, this means it can be generated and recognised by a DFA (deterministic finite automata).
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
# Triangle Abstract Machine
|
||||||
|
|
||||||
|
**TAM** instruction set
|
||||||
|
|
||||||
|
```assembly
|
||||||
|
LOADL (int)
|
||||||
|
NEG
|
||||||
|
ADD
|
||||||
|
SUB
|
||||||
|
MUL
|
||||||
|
DIV
|
||||||
|
```
|
||||||
|
|
||||||
|
TAM works on a stack of integers.
|
||||||
|
|
||||||
|
##### Executing a TAM program
|
||||||
|
|
||||||
|
```assembly
|
||||||
|
LOADL 7
|
||||||
|
ADD --adds top two numbers on the stack
|
||||||
|
LOADL 2
|
||||||
|
SUB -- note its 15-2
|
||||||
|
LOADL 4
|
||||||
|
DIV --integer division
|
||||||
|
```
|
||||||
|
|
||||||
|
The stack during this program:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\begin{bmatrix}
|
||||||
|
{8} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
\implies
|
||||||
|
\begin{bmatrix}
|
||||||
|
{7} \\
|
||||||
|
{8} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
\implies
|
||||||
|
\begin{bmatrix}
|
||||||
|
{15} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
\implies
|
||||||
|
\begin{bmatrix}
|
||||||
|
{2} \\
|
||||||
|
{15} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
\implies
|
||||||
|
\begin{bmatrix}
|
||||||
|
{13} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
\implies
|
||||||
|
\begin{bmatrix}
|
||||||
|
{4} \\
|
||||||
|
{13} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
\implies
|
||||||
|
\begin{bmatrix}
|
||||||
|
{3} \\
|
||||||
|
{5}
|
||||||
|
\end{bmatrix}
|
||||||
|
$$
|
||||||
|
|
||||||
|
## Compiler Complete Example
|
||||||
|
|
||||||
|
Program in **Arith**
|
||||||
|
|
||||||
|
```c
|
||||||
|
5 * ((8 + 7) - 2) / 4
|
||||||
|
```
|
||||||
|
|
||||||
|
**A**bstract **S**yntax **T**ree
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**TAM** program
|
||||||
|
|
||||||
|
```assembly
|
||||||
|
LOADL 5
|
||||||
|
LOADL 8
|
||||||
|
LOADL 7
|
||||||
|
ADD
|
||||||
|
LOADL 2
|
||||||
|
SUB
|
||||||
|
LOADL 4
|
||||||
|
DIV
|
||||||
|
MUL
|
||||||
|
```
|
||||||
|
|
||||||
|
## Implementing TAM in Haskell
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
module TAM where
|
||||||
|
|
||||||
|
data TamInstruction = LOADL Int
|
||||||
|
| ADD | SUB
|
||||||
|
| MUL | DIV
|
||||||
|
| NEG
|
||||||
|
deriving(Eq, Show)
|
||||||
|
type Stack = [Int]
|
||||||
|
|
||||||
|
execute :: [TamInstruction] -> Stack -> Stack
|
||||||
|
execute [] s = s --if stack empty, then return the stack
|
||||||
|
execute (LOADL n : tp) s = execute tp (n : s) --push n to top of stack
|
||||||
|
execute (ADD : tp) (a : b : s) = execute tp ((a+b):s) --push a+b
|
||||||
|
...
|
||||||
|
execute (DIV : tp) (a : b : s) = execute tp ((a`div`b):s)
|
||||||
|
```
|
||||||
|
|
||||||
|
Quicker way to write the execute function using `absOpToConcrOp`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
convOp :: TamInstruction -> Int -> Int -> Int
|
||||||
|
convOp ADD = (+)
|
||||||
|
convOp SUB = (-)
|
||||||
|
convOp MUL = (*)
|
||||||
|
convOp DIV = (`div`)
|
||||||
|
|
||||||
|
execute :: [TamInstruction] -> Stack -> Stack
|
||||||
|
execute [] s = s
|
||||||
|
execute (LOADL n : tp) s = execute tp (n : s)
|
||||||
|
execute (NEG : tp) (a : s) = execute tp (-a : s)
|
||||||
|
execute (op : tp) (a : b : s) = execute tp ((convOp op a b) : s)
|
||||||
|
```
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
# Functional Parsers
|
||||||
|
|
||||||
|
In our parser - there's a lot of repeated code and a lot of cases.
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
Types of scanner and parser are very similar
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
scanToken :: String -> Maybe (Token, String)
|
||||||
|
parseTerm :: [Token] -> Maybe (AST, [Token])
|
||||||
|
parseExp :: [Token] -> Maybe (AST, [Token])
|
||||||
|
|
||||||
|
general :: [c] -> Maybe (a, [c])
|
||||||
|
lessGeneral :: String -> Maybe (a, String)
|
||||||
|
|
||||||
|
--remember :t string :: [char]
|
||||||
|
-- [c] list of characters or tokens
|
||||||
|
|
||||||
|
Parser a :: String -> [(a, String)]
|
||||||
|
-- no maybe needed as failure is now returning an empty list
|
||||||
|
--The parser of type a, we can now define generic functions that operate on a given type => less repeated code
|
||||||
|
--This is an instance of a typeclass (monad yikes)
|
||||||
|
```
|
||||||
|
|
||||||
|
Do notation
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
Parser a = String -> [(a, String)]
|
||||||
|
|
||||||
|
symbol :: String -> Parser ()
|
||||||
|
-- no need to define result, as all it does it succeed or fail
|
||||||
|
|
||||||
|
exp :: Parser AST
|
||||||
|
parseParenthesis :: Parser AST
|
||||||
|
parseParenthesis = do symbol '('
|
||||||
|
t <- exp
|
||||||
|
symbol ')'
|
||||||
|
return t
|
||||||
|
```
|
||||||
|
|
||||||
@@ -0,0 +1,101 @@
|
|||||||
|
# Functor
|
||||||
|
|
||||||
|
Parsing an expression in parenthesis:
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
parseP :: Parser AST
|
||||||
|
parseP = do symbol '('
|
||||||
|
t <- exp
|
||||||
|
symbol ')'
|
||||||
|
return t
|
||||||
|
```
|
||||||
|
|
||||||
|
Before we write this sort of code, we need to understand `type classes` (especially `monads`)
|
||||||
|
|
||||||
|
## Types vs Typeclasses
|
||||||
|
|
||||||
|
| Types | Type classes |
|
||||||
|
| ------ | ------------ |
|
||||||
|
| Bool | Eq |
|
||||||
|
| Char | Show |
|
||||||
|
| AST | Num |
|
||||||
|
| String | Functor |
|
||||||
|
| | Monad |
|
||||||
|
|
||||||
|
**Eq**: typeclass equality; A type can only be typeclass equality if two like types can be compared
|
||||||
|
|
||||||
|
A type can be a *member* (instance) of a type class, meaning that if has the properties/functions that the class requires
|
||||||
|
|
||||||
|
e.g. `Bool` is an instance of `Eq` and `Show`
|
||||||
|
|
||||||
|
###### Is there a type that is **not** in `Eq`?
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
(\c -> c :: Int) == (\c -> c :: Int)
|
||||||
|
```
|
||||||
|
|
||||||
|
**ERROR**: No instance for `Eq(Int -> Int)`
|
||||||
|
|
||||||
|
Why?
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
f :: Int -> Int
|
||||||
|
g :: Int -> Int
|
||||||
|
```
|
||||||
|
|
||||||
|
Then `f == g` should be `fn == gn` for every n, the computer cannot do this (halting problem).
|
||||||
|
|
||||||
|
## Type Constructors
|
||||||
|
|
||||||
|
A type constructor takes a type to construct a new type.
|
||||||
|
|
||||||
|
`Maybe` - not a type but a type constructor
|
||||||
|
|
||||||
|
`Maybe String` - a type
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
newtype Parser a = P (String -> [a, String])
|
||||||
|
```
|
||||||
|
|
||||||
|
**Parser** is a type constructor
|
||||||
|
|
||||||
|
**Parser AST** is a type
|
||||||
|
|
||||||
|
Functor is a typeclass of which `parser` is an instance
|
||||||
|
|
||||||
|
##### Functor
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
class Functor f where
|
||||||
|
fmap :: (a -> b) -> fa -> fb
|
||||||
|
|
||||||
|
instance Functor Maybe where
|
||||||
|
fmap g (Just x) = Just (g x)
|
||||||
|
fmap g Nothing = Nothing -- fmap id = id
|
||||||
|
|
||||||
|
-- lists
|
||||||
|
instance Functor [] where
|
||||||
|
fmap g [] = []
|
||||||
|
fmap g (t:ts) = (g t) : fmap g ts
|
||||||
|
|
||||||
|
-- goal: write parser as a functor
|
||||||
|
newtype Parser a = P ( String -> [a, String] )
|
||||||
|
-- Need: fmap :: (a->b) -> Parser a -> Parser b
|
||||||
|
|
||||||
|
instance Functor Parser where
|
||||||
|
fmap g pa = -- parser pa
|
||||||
|
P (\str -> map (\(x,s) -> (gx,s))
|
||||||
|
parse pa str)
|
||||||
|
```
|
||||||
|
|
||||||
|
##### Rules of Functors
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
fmap id = id -- identity
|
||||||
|
fmap (f . g) = fmap f . fmap g
|
||||||
|
```
|
||||||
|
|
||||||
|
Haskell doesn't enforce these rules however it is convention.
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
# Applicative Functors
|
||||||
|
|
||||||
|
Types: `Bool`, `Int`, `Char`, `[Char] = String`
|
||||||
|
|
||||||
|
Type Constructors: `Maybe`, `[]`
|
||||||
|
|
||||||
|
(type) classes: `Eq`, `Show`, `Functor`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
newtype Parser a = P ( String -> [a, String])
|
||||||
|
parse :: Parser a -> String -> [(a, String)]
|
||||||
|
parse (P p) s = p s -- s's can be cancelled from both sides
|
||||||
|
|
||||||
|
instance Functor Parser where
|
||||||
|
-- fmap :: (a -> b) -> Parser a -> Parser b
|
||||||
|
fmap g pa = P (\s -> [(g x, s1) |
|
||||||
|
(x,s1) <- parse pa s])
|
||||||
|
```
|
||||||
|
|
||||||
|
Applicative - motivation
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
Functor f
|
||||||
|
fmap0 :: a -> f a
|
||||||
|
fmap1 :: (a -> b) -> f a -> f b
|
||||||
|
-- cannot do this with functors ie cannot deal with multiple parameters
|
||||||
|
fmap2 :: (a -> b -> c) -> f a -> f b -> f c
|
||||||
|
fmap3 :: (a -> ... n) -> f a -> ... f n
|
||||||
|
```
|
||||||
|
|
||||||
|
`Functor f` can do `fmap1` however cannot do `fmap0` or `fmap2` etc.
|
||||||
|
|
||||||
|
**Remember**: `a -> b -> c == a -> (b -> c)`
|
||||||
|
|
||||||
|
For `fmap2` we can use `fmap2 :: (a -> (b -> c)) -> f a -> f (a -> b)`
|
||||||
|
|
||||||
|
would need: `f(b -> c) -> f b -> f c`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
class Functor f => Applicative f where
|
||||||
|
pure :: a -> f a
|
||||||
|
(<*>) :: f (a -> b) -> f a -> f b
|
||||||
|
-- <*> infix operator
|
||||||
|
-- NOTE its f (a -> b) and not (a -> b) in fmap1
|
||||||
|
-- fmap1 not part of the applicative class
|
||||||
|
```
|
||||||
|
|
||||||
|
Writing `fmap3` in an applicative functor
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
fmap3 :: g x y z = (pure g) <*> x <*> y <*> z
|
||||||
|
```
|
||||||
|
|
||||||
|
##### Example Maybe
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Applicative Maybe where
|
||||||
|
-- pure :: a -> Maybe a
|
||||||
|
pure x = Just x
|
||||||
|
-- (<*>) :: Maybe (a -> b) -> Maybe a -> Maybe b
|
||||||
|
Just g <*> (Just x) = Just (g x)
|
||||||
|
_ <*> _ = Nothing
|
||||||
|
```
|
||||||
|
|
||||||
|
##### Example Lists
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Applicative [] where
|
||||||
|
-- pure :: a -> [a]
|
||||||
|
pure x = [x]
|
||||||
|
-- (<*>) :: [a -> b] -> [a] -> [b]
|
||||||
|
gs <*> xs = [g x | g <- gs, x <- xs]
|
||||||
|
```
|
||||||
|
|
||||||
|
##### Example Parser
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Applicative Parser where
|
||||||
|
-- pure :: a -> Parser a
|
||||||
|
-- newtype Parser a = P ( String -> [(a, String)] )
|
||||||
|
pure x = P (\s -> [(x,s)])
|
||||||
|
-- <*> :: Parser (a -> b) -> Parser a -> Parser b
|
||||||
|
pf <*> pa = P (\s -> [ (f x, s2) |
|
||||||
|
(f, s1) <- parse pf s,
|
||||||
|
(x, s2) <- parse pa s1)])
|
||||||
|
```
|
||||||
|
|
||||||
|
All parse does is apply a parser
|
||||||
|
|
||||||
|
`parse :: Parser a -> String -> [(a, String)]`
|
||||||
|
|
||||||
|
Where `P` is the constructor
|
||||||
|
|
||||||
|
`parse ( P p ) = p`
|
||||||
|
|
||||||
@@ -0,0 +1,466 @@
|
|||||||
|
### Functor Class of Parsers
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
newtype Parser a = P ( String -> [(a, String)] )
|
||||||
|
|
||||||
|
parse :: Parser a -> String -> [(a, String)]
|
||||||
|
parse (P f) src = f src
|
||||||
|
|
||||||
|
item :: Parser Char
|
||||||
|
item = P (\src -> case src of
|
||||||
|
[] -> []
|
||||||
|
(c:src') -> [(c,src')] )
|
||||||
|
|
||||||
|
symbol :: String -> Parser ()
|
||||||
|
|
||||||
|
integer :: Parser Int
|
||||||
|
|
||||||
|
binary :: Parser Int
|
||||||
|
|
||||||
|
intORbin :: Parser Int
|
||||||
|
|
||||||
|
expr :: Parser AST
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (symbol "something") "nothing"
|
||||||
|
[]
|
||||||
|
|
||||||
|
λ> parse (symbol "<=") "<= something nothing"
|
||||||
|
[((), "something nothing")]
|
||||||
|
NOTE: does nothing because all we have implemented for symbol is ()
|
||||||
|
|
||||||
|
λ> integer "123 blah blah"
|
||||||
|
[(123, "blah blah")]
|
||||||
|
|
||||||
|
λ> parse binary "101 blah"
|
||||||
|
[(5, "blah")]
|
||||||
|
|
||||||
|
λ> parse intORbin "101 blah"
|
||||||
|
[(101, "blah"), (5, "blah")]
|
||||||
|
|
||||||
|
λ> parse expr "1+2*3"
|
||||||
|
[(BinOp Addition (LitInteger 1) BinOp Multiplication (LitInteger 2) (LitInteger 3)), "")]
|
||||||
|
NOTE: expr defined in ArtihExpr
|
||||||
|
```
|
||||||
|
|
||||||
|
Defining the functor parser
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Functor Parser where
|
||||||
|
-- must not give type of fmap as it is already given in functor class
|
||||||
|
-- good practice to comment type
|
||||||
|
-- fmap :: (a -> b) -> Parser a -> Parser b
|
||||||
|
--first assume returns one value
|
||||||
|
-- doesnt fail, doesn't produce more than one result
|
||||||
|
fmap g pa = P (\src -> let [(x,src1)] = parse pa src
|
||||||
|
in [(g x, src1)] )
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (fmap (+3) integer) "42 blah blah"
|
||||||
|
[(45, blah blah)]
|
||||||
|
|
||||||
|
λ> parse (fmap evaluate expr) "1+2*3"
|
||||||
|
[(7,"")]
|
||||||
|
|
||||||
|
λ> parse (fmap (+3) integer) "42 blah blah"
|
||||||
|
*** Exception Non-exhaustive patterns
|
||||||
|
|
||||||
|
λ> parse (fmap (+3) intORbin) "101 blah"
|
||||||
|
*** Exception Non-exhaustive patterns
|
||||||
|
```
|
||||||
|
|
||||||
|
fixing `fmap`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src])
|
||||||
|
-- using list comprehension
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (fmap (+3) intORbin) "101 blah"
|
||||||
|
[(104, "blah"), (8, "blah")]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Applicative Class of Parsers
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Applicative Parser where
|
||||||
|
-- pure :: a -> Parser a
|
||||||
|
-- commenting type for good practice
|
||||||
|
pure x = P (\src -> [(x, src)])
|
||||||
|
|
||||||
|
-- (<*>) :: Parser (a -> b) -> Parser a -> Parser b
|
||||||
|
|
||||||
|
|
||||||
|
simpleFun :: Parser (Int -> Int)
|
||||||
|
-- parser the function "double" or "square"
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (fmap (\f -> f 3) simpleFun) "double blah"
|
||||||
|
[(6, "blah")]
|
||||||
|
|
||||||
|
a parser that returns a function as a result
|
||||||
|
λ> parse simpleFun "double blah blah"
|
||||||
|
parse simpleFun "double blah blah" :: [(Int -> Int, String)]
|
||||||
|
-- the function
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Applicative Parser where
|
||||||
|
-- pure :: a -> Parser a
|
||||||
|
-- commenting type for good practice
|
||||||
|
pure x = P (\src -> [(x, src)])
|
||||||
|
|
||||||
|
-- (<*>) :: Parser (a -> b) -> Parser a -> Parser b
|
||||||
|
pf <*> pa = P (\src -> let [(f,src1)] = parse pf src
|
||||||
|
[(x,src2)] = parse pa src
|
||||||
|
in [(f x, src2)] )
|
||||||
|
-- this works if the two parsers both give one, different result
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (simpleFun <*> integer) "double 7"
|
||||||
|
[(14, "")]
|
||||||
|
λ> parse (simpleFun <*> integer) "square 7"
|
||||||
|
[(49, "")]
|
||||||
|
λ> parse (simpleFun <*> integer) "cube 7"
|
||||||
|
*** Exception non-exhaustive pattern
|
||||||
|
|
||||||
|
λ> parse (simpleFun <*> intORbin) "square 101"
|
||||||
|
*** Exception non-exhaustive pattern
|
||||||
|
-- fails bc intORbin gives two results
|
||||||
|
```
|
||||||
|
|
||||||
|
Using list comprehension
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
pf <*> pa = P (\src -> [ (f x, src2) | (f,src1) <- parse pf src,
|
||||||
|
(x,src2) <- parse pa src1 ] )
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (simpleFun <*> integer) "cube 7"
|
||||||
|
[]
|
||||||
|
λ> parse (simpleFun <*> intORbin) "square 101"
|
||||||
|
[(10201, ""), (25, "")]
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
### Monad Class of Parser
|
||||||
|
|
||||||
|
Monad class will facilitate the use of `do` notation.
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Monad Parser where
|
||||||
|
-- return :: a -> Parser a
|
||||||
|
-- we dont have to define return as its automatically defined as
|
||||||
|
-- return = pure
|
||||||
|
--only method we need to define for the monad class is bind >>=
|
||||||
|
|
||||||
|
-- (>>=) :: Parser a -> (a -> Parser b) -> Parser b
|
||||||
|
pa >>= fpb = P (\src -> let [(x, src1)] = parse pa src
|
||||||
|
[(y, src2)] = parse (fpb x) src1
|
||||||
|
in [(y,src2)] )
|
||||||
|
|
||||||
|
checkNum :: Int -> Parser Bool
|
||||||
|
checkNum n = fmap (==n) integer
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (checkNum 7) " 7 blah blah"
|
||||||
|
[(True, "blah blah")]
|
||||||
|
|
||||||
|
λ> parse (checkNum 6) " 7 blah blah"
|
||||||
|
[(False, "blah blah")]
|
||||||
|
|
||||||
|
λ> parse (checkNum 7) " no blah blah"
|
||||||
|
[]
|
||||||
|
λ> parse (binary >>= checkNum) "101 5"
|
||||||
|
[(True, "")]
|
||||||
|
λ> parse (binary >>= checkNum) "101 6"
|
||||||
|
[(False, "")]
|
||||||
|
λ> parse (binary >>= checkNum) "no 101 6"
|
||||||
|
*** Exception non-exhaustive pattern
|
||||||
|
|
||||||
|
λ> parse (intORbin >>= checkNum) "101 6"
|
||||||
|
*** Exception non-exhaustive pattern
|
||||||
|
--cant cope with multiple values
|
||||||
|
```
|
||||||
|
|
||||||
|
Using list comprehension
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
pa >>= fpb = P (\src -> [ (y,src2) | (x,src1) <- parse pa src,
|
||||||
|
(y,src2) <- parse (fpb x) src1 ] )
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (binary >>= checkNum) "no 101 6"
|
||||||
|
[]
|
||||||
|
|
||||||
|
λ> parse (intORbin >>= checkNum) "110 6"
|
||||||
|
[(False,""), (True, "")]
|
||||||
|
-- false is 110 (base 10) != 6
|
||||||
|
-- true is 110 (base 2) == 6
|
||||||
|
```
|
||||||
|
|
||||||
|
Improving the definition further
|
||||||
|
|
||||||
|
As we unpack and repack `(y,src2)`, we can just call it `r` (result)
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
pa >>= fpb = P (\src -> [ r | (x,src1) <- parse pa src,
|
||||||
|
r <- parse (fpb x) src1 ] )
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (intORbin >>= checkNum) "113 113"
|
||||||
|
[(True,""), (True, "113 ")]
|
||||||
|
-- the integer part recognises 113 == 113
|
||||||
|
-- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true
|
||||||
|
```
|
||||||
|
|
||||||
|
What is the do notation and how is it connected to the bind function, we will show this by writing a simple parser
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
pairSum :: Parser Int
|
||||||
|
-- read (parse) an integer, bind it to a function, map it to another parser
|
||||||
|
pairSum = integer >>= \n -> integer >>= \m -> return (n+m)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse pairSum "3 8"
|
||||||
|
[(11, "")]
|
||||||
|
```
|
||||||
|
|
||||||
|
Rewriting `pairSum` with `do`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
pairSum :: Parser Int
|
||||||
|
-- apply integer and then put it into variable n
|
||||||
|
-- apply integer and bind to variable m
|
||||||
|
pairSum = do n <- integer
|
||||||
|
m <- integer
|
||||||
|
return (n+m)
|
||||||
|
--much cleaner & easier to understand
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
parse (symbol "number" >>= \u -> integer) "number 9"
|
||||||
|
[(9, "")]
|
||||||
|
parse (symbol "number" >> integer) "number 9"
|
||||||
|
[(9, "")]
|
||||||
|
|
||||||
|
NOTE: >> is a non-dependant bind
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
the grammer
|
||||||
|
--funApp ::= ( simpleFun integer )
|
||||||
|
-- will be a parser that returns an integer
|
||||||
|
funApp :: Parser Int
|
||||||
|
funApp = symbol '(' >> (simpleFun <*> integer) >>= \y -> symbol ')' >> return y
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse funApp "(double 5)"
|
||||||
|
[(10, "")]
|
||||||
|
```
|
||||||
|
|
||||||
|
Rewrite with `do`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
funApp = do symbol '('
|
||||||
|
f <- simpleFun
|
||||||
|
x <-integer
|
||||||
|
symbol ')'
|
||||||
|
return (f x)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Alternative Class of Parser
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Alternative Parser where
|
||||||
|
-- empty :: Parser a
|
||||||
|
empty = P (\src -> [])
|
||||||
|
|
||||||
|
-- (<|>) :: Parser a -> Parser a -> Parser a
|
||||||
|
p1 <|> p2 = P (\src -> case parse p1 src of
|
||||||
|
[] -> parse p2 src
|
||||||
|
rs -> rs)
|
||||||
|
-- if p1 fails, then parse with p2, else return result rs
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (symbol "abc" <|> symbol "acb") "abc"
|
||||||
|
[("abc", "")]
|
||||||
|
λ> parse (symbol "abc" <|> symbol "acb") "xyz"
|
||||||
|
[]
|
||||||
|
λ> parse (integer <|> binary) "1101"
|
||||||
|
[(1101,"")]
|
||||||
|
λ> parse (binary <|> integer) "1101"
|
||||||
|
[(13,"")]
|
||||||
|
-- will only apply p2 if p1 fails
|
||||||
|
λ> parse (binary <|> integer) "1201"
|
||||||
|
[(1,"201")]
|
||||||
|
-- binary successfully parses "1" and leaves "201"
|
||||||
|
```
|
||||||
|
|
||||||
|
Using parallel choice notation `<||>`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
(<||>) :: Parser a -> Parser a -> Parser a
|
||||||
|
p1 <||> p2 = P (\src -> parse p1 src ++ parse p2 src)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (binary <||> integer) "1101"
|
||||||
|
[(13, ""), (1101, "")]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Explaining the `FunParser.hs` library
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
satisfy :: Parser a -> (a -> Bool) -> Parser a
|
||||||
|
satisfy p cond = do x <- p
|
||||||
|
if (cond x) then return x
|
||||||
|
else empty
|
||||||
|
-- the way to denote failure is empty (from alternitve class)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (satisfy integer (>10)) "42"
|
||||||
|
[(42, "")]
|
||||||
|
λ> parse (satisfy integer (>10)) "9"
|
||||||
|
[]
|
||||||
|
```
|
||||||
|
|
||||||
|
Writing a satisfy function just for characters
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
sat :: (Char -> Bool) -> Parser Char
|
||||||
|
-- item parses 1 character
|
||||||
|
sat cond = satisfy item cond
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (sat isUpper) "a"
|
||||||
|
[]
|
||||||
|
λ> parse (sat isUpper) "A"
|
||||||
|
['A',""]
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
lower :: Parser Char
|
||||||
|
lower = sat isLower
|
||||||
|
|
||||||
|
upper :: Parser Char
|
||||||
|
upper = sat isUpper
|
||||||
|
|
||||||
|
digit :: Parser Char
|
||||||
|
digit = sat isDigit
|
||||||
|
|
||||||
|
--and so on for others like letter & alphaNumeric
|
||||||
|
|
||||||
|
char :: Char -> Parser Char
|
||||||
|
char c = sat (==c)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (char 'A') "not a captial a"
|
||||||
|
[]
|
||||||
|
λ> parse (char 'A') "A not a captial a"
|
||||||
|
['A'," not a capital a"]
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
string :: String -> Parser String
|
||||||
|
string [] = return [] --list as string is list of chars
|
||||||
|
string (c:cs) = do char c
|
||||||
|
string cs
|
||||||
|
return (c:cs)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (string "hello") "hello everybody"
|
||||||
|
[("hello", "everybody")]
|
||||||
|
λ> parse (string "hello") " hello everybody"
|
||||||
|
[]
|
||||||
|
λ> parse (sat isSpace) " hello"
|
||||||
|
[(' ',"hello")]
|
||||||
|
λ> parse (many (sat isSpace)) " hello"
|
||||||
|
[(' ',"hello")]
|
||||||
|
```
|
||||||
|
|
||||||
|
We have to fix leading white space causing failure
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
space :: Parser ()
|
||||||
|
-- a parser that succeeds or fails and does not return anything
|
||||||
|
space = do many (sat isSpace)
|
||||||
|
return ()
|
||||||
|
-- writing a parser to ignore white space
|
||||||
|
token :: Parser a -> Parser a
|
||||||
|
token p = do space
|
||||||
|
x <- p
|
||||||
|
space
|
||||||
|
return x
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (token (string "hello")) " hello everybody"
|
||||||
|
[("hello","everybody")]
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
symbol :: String -> Parser String
|
||||||
|
symbol = token (string s)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
λ> parse (symbol "hello") " hello everybody"
|
||||||
|
[("hello","everybody")]
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Defining parsers for arithmetic expressions
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
-- expr ::= mexpr + exp | mexpr - exp | mexpr
|
||||||
|
expr :: Parser AST
|
||||||
|
expr = do t1 <- mexpr
|
||||||
|
symbol '+'
|
||||||
|
t2 <- expr
|
||||||
|
return (BinOp Addition t1 t2)
|
||||||
|
<|>
|
||||||
|
do t1 <- mexpr
|
||||||
|
symbol '-'
|
||||||
|
t2 <- expr
|
||||||
|
return (BinOp Subtraction t1 t2)
|
||||||
|
<|>
|
||||||
|
mexpr
|
||||||
|
|
||||||
|
--we can optimise this grammer as all symbols start with mexpr
|
||||||
|
-- expr ::= mexpr ( + expr | - expr | empty)
|
||||||
|
expr :: Parser AST
|
||||||
|
expr = do t1 <- mexpr
|
||||||
|
(do symbol '+'
|
||||||
|
t2 <- expr
|
||||||
|
return (BinOp Addition t1 t2)
|
||||||
|
<|>
|
||||||
|
do symbol '-'
|
||||||
|
t2 <- expr
|
||||||
|
return (BinOp Subtraction t1 t2)
|
||||||
|
<|>
|
||||||
|
return t1)
|
||||||
|
```
|
||||||
|
|
||||||
@@ -0,0 +1,149 @@
|
|||||||
|
# Compiling Variables
|
||||||
|
|
||||||
|
A variable is identified by a alphanumeric string. We can store this as a list of pairs, with the variables identifier and its value.
|
||||||
|
|
||||||
|
Variable Environment or VarEnv - `[(Identifier, Stack Address)]`
|
||||||
|
|
||||||
|
A stack address is an integer value that specifies where in the stack that variable is contained. The bottom of the stack is reserved for variable values.
|
||||||
|
|
||||||
|
The bottom of the stack is indexed `0`.
|
||||||
|
|
||||||
|
Lets say our environment consists of 3 variables named x,y,z. It would look like:
|
||||||
|
|
||||||
|
`[("z",2), ("y",1), ("x",0)]`
|
||||||
|
|
||||||
|
| Variables | Stack (Values) | Index |
|
||||||
|
| :-------: | :------------: | :---: |
|
||||||
|
| x | 7 | 0 |
|
||||||
|
| y | 2 | 1 |
|
||||||
|
| z | 9 | 2 |
|
||||||
|
|
||||||
|
To get the value of a variable from the stackk, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
|
||||||
|
|
||||||
|
`LOAD a` - copy address a to top of stack
|
||||||
|
|
||||||
|
`STORE a` - pop top of stack to address a
|
||||||
|
|
||||||
|
For example if `LOADL 2` is called, it will effect the stack in the following way:
|
||||||
|
|
||||||
|
| Variables | Stack (Values) | Index |
|
||||||
|
| :-------: | :------------: | :---: |
|
||||||
|
| x | 7 | 0 |
|
||||||
|
| y | 2 | 1 |
|
||||||
|
| z | 9 | 2 |
|
||||||
|
| | … | |
|
||||||
|
| | 9 | |
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
expCode :: VarEnv -> Expr -> [TAMInst]
|
||||||
|
```
|
||||||
|
|
||||||
|
Before we just called the abstract syntax tree `AST` however with the extended grammar now we will have multiple ASTs, one for programs, one for commands, expressions. The AST for expressions we call `Expr`.
|
||||||
|
|
||||||
|
Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list.
|
||||||
|
|
||||||
|
## Declaration of Variables
|
||||||
|
|
||||||
|
```js
|
||||||
|
let var x; //no value given means initialised to 0
|
||||||
|
var y := 5 //note no semicolon
|
||||||
|
var z;
|
||||||
|
in ...
|
||||||
|
```
|
||||||
|
|
||||||
|
For the code above, we need to generate a VarEnv. The compiler needs to generate a variable environment and TAM code.
|
||||||
|
|
||||||
|
VarEnv: `[("z",2), ("y",1), ("x",0)]`
|
||||||
|
|
||||||
|
TAM code stack: `[0,5,0]`
|
||||||
|
|
||||||
|
However we also need to account for expressions such as:
|
||||||
|
|
||||||
|
```js
|
||||||
|
let var x := 3;
|
||||||
|
var y := 5;
|
||||||
|
var z := x*y
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
declarationCompiler :: [Declaration] -> (VarEnv, [TAMInstr])
|
||||||
|
VarEnv :: [(Identifier, Address)]
|
||||||
|
```
|
||||||
|
|
||||||
|
NOTE: this can be defined with functions given in the `FunParser` library. Or using a `state monad`
|
||||||
|
|
||||||
|
### State Monad
|
||||||
|
|
||||||
|
$s_0 \rightarrow s_1 \rightarrow s_2 \rightarrow s_n$ for each change in state, there's a corresponding result generated.
|
||||||
|
|
||||||
|
$$
|
||||||
|
a_0 \quad\space\space\space a_1 \quad\space\space\space a_n
|
||||||
|
$$
|
||||||
|
|
||||||
|
- For each of these states, we need a variable environment and address
|
||||||
|
|
||||||
|
- For each of the results, we need to generate TAM instructions.
|
||||||
|
|
||||||
|
Example: $s_n$ could be your bank balance and $a_n$ could be the purchase history.
|
||||||
|
|
||||||
|
- In our case:
|
||||||
|
- States are VarEnv & next free address space for next variable
|
||||||
|
- Outputs are TAM instructions
|
||||||
|
|
||||||
|
We to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
newtype ST st a = S (\st -> (a, st))
|
||||||
|
-- ST - state transformer
|
||||||
|
-- st - type of states
|
||||||
|
-- a - type of output/results
|
||||||
|
-- S - constructor
|
||||||
|
-- \st a function that takes a state and returns a value along with a new state
|
||||||
|
-- this is a general type definition with state type st and result type a
|
||||||
|
-- this is still just a type constructor, has to be applied to a type
|
||||||
|
instance Functor (ST st)
|
||||||
|
instance Applicative (ST st)
|
||||||
|
instance Monad (ST st)
|
||||||
|
--as we inherit the monad class, we can use do notation
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
newtype ST st a = S (\st -> (a, st))
|
||||||
|
--type definition
|
||||||
|
ST Int
|
||||||
|
--type constructor
|
||||||
|
ST Int String
|
||||||
|
--type
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
app :: ST st a -> st -> (a, st)
|
||||||
|
app (S f) x = f x
|
||||||
|
--applies the constructor to state x
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Functor (ST st) where
|
||||||
|
--fmap :: (a->b) -> ST st a -> ST st b
|
||||||
|
fmap g sta = S (\s -> let (x,s') = app sta s
|
||||||
|
in (g x, s'))
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Applicative (ST st) where
|
||||||
|
--pure :: a -> ST st a
|
||||||
|
pure x = S (\s -> (x,s))
|
||||||
|
--(<*>) :: (ST st (a -> b)) -> ST st a -> ST st b
|
||||||
|
stf <*> sta = S (\s -> let (f,s') = app stf s
|
||||||
|
(x,s'') = app sta s')
|
||||||
|
in (f x, s''))
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
instance Monad (ST st) where
|
||||||
|
return = pure
|
||||||
|
-- (>>=) :: (ST st a) -> (a -> ST st b) -> ST st b
|
||||||
|
sta >>= f = S (\s -> let (x,s') = app sta s
|
||||||
|
(y,s'') = app (f x) s'
|
||||||
|
in (y,s''))
|
||||||
|
```
|
||||||
@@ -0,0 +1,164 @@
|
|||||||
|
# Variable Environments
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
type VarEnv = [(Identifier, StkAddress)]
|
||||||
|
-- String Int
|
||||||
|
|
||||||
|
address :: VarEnv -> Identifer -> StkAddress
|
||||||
|
address ve v = case lookup v ve of
|
||||||
|
Nothing -> error "variable not in enviroment"
|
||||||
|
Just a -> a
|
||||||
|
--Expr is AST of expressions
|
||||||
|
expCode :: VarEnv -> Expr -> [TAMInstr]
|
||||||
|
expCode ve (LitInteger x) = [LOADL x]
|
||||||
|
-- we must put variable value on top of the stack
|
||||||
|
expCode ve (Var v) = [LOAD (address ve v)]
|
||||||
|
```
|
||||||
|
|
||||||
|
How do we build a variable environment?
|
||||||
|
|
||||||
|
Every program begins with a sequence of variable declarations
|
||||||
|
|
||||||
|
```js
|
||||||
|
var x := 7;
|
||||||
|
var y := 3;
|
||||||
|
var z;
|
||||||
|
var w := x * y - 2
|
||||||
|
```
|
||||||
|
|
||||||
|
The parser will turn this into a list of AST for declarations
|
||||||
|
|
||||||
|
Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack.
|
||||||
|
|
||||||
|
We do this using the state monad
|
||||||
|
|
||||||
|
- We use as an underlying state the variable environment itself, as we build it sequentially
|
||||||
|
- We also keep the stack address as a state, where it keeps the next free address
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
declsCode :: [Declarations] -> (VarEnv, [TAMInstr])
|
||||||
|
declsCode ds = let (tam,(ve,0a)) app (declsTAM ds) ([],0) --initial state
|
||||||
|
in (ve,tam)
|
||||||
|
|
||||||
|
declsTAM :: [Declarations] -> ST (VarEnv, StkAddress) [TAMInstr]
|
||||||
|
declsTAM [] = return []
|
||||||
|
declsTAM (d:ds) = do
|
||||||
|
td <- declTAM d
|
||||||
|
tds <- declsTAM ds
|
||||||
|
return (td++tds)
|
||||||
|
|
||||||
|
declTAM :: Declarations -> ST (VarEnv, StkAddress) [TAMInstr]
|
||||||
|
declTAM (VarDecl v) = do
|
||||||
|
(ve,a) <- stState
|
||||||
|
stUpdate ((v,a) : ve, a+1)
|
||||||
|
return [LOADL 0]
|
||||||
|
declTAM (VarInit v e) = do
|
||||||
|
(ve,a) <- stState
|
||||||
|
stUpdate ((v,a) : ve, a+1)
|
||||||
|
return (expCode ve e)
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
```shell
|
||||||
|
λ> parseAll declarations "var x:=7;var y:=3;var z;var w:=x*y-2"
|
||||||
|
[VarInit "x" (LitInteger 7), VarInit "y" (LitInteger 3), VarDecl "z", VarInit "w" (BinOp Subtraction (BinOp Multiplication (Var "x") (Var "y")) (LitInteger 2))]
|
||||||
|
|
||||||
|
λ> ds = parseAll declarations "var x:=7;var y:=3;var z;var w:=x*y-2"
|
||||||
|
|
||||||
|
λ> (ve,tam) = declsCode ds
|
||||||
|
λ> ve
|
||||||
|
[("w",3),("z",2),("y",1),("x",0)]
|
||||||
|
λ> tam
|
||||||
|
[LOADL 7, LOADL 3m LOADL 0, LOAD 0, LOAD 1, MUL, LOADL 2, SUB]
|
||||||
|
λ> execTAM [] tam
|
||||||
|
[19, 0, 3, 7]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Designing ASTs for any grammar
|
||||||
|
|
||||||
|
- We turn every non-terminal of the grammar into a type of AST
|
||||||
|
|
||||||
|
- We turn every production of the non-terminal into a constructor of the type
|
||||||
|
|
||||||
|
Defining the grammar of TAM
|
||||||
|
|
||||||
|
```
|
||||||
|
command ::= identifier := expr
|
||||||
|
| if expr then command else command
|
||||||
|
| while expr do command
|
||||||
|
| getint ( identifier )
|
||||||
|
| printint ( expr )
|
||||||
|
| begin commands end
|
||||||
|
```
|
||||||
|
|
||||||
|
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminal
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
data Command =
|
||||||
|
|
||||||
|
datatypes Identifier = String, Expr, Commands -- [Command]
|
||||||
|
```
|
||||||
|
|
||||||
|
Assign to every production one constructor for the data type.
|
||||||
|
|
||||||
|
This means we will have 6 constructors called `Assignment`, `IfThenElse`, `WhileDo`, `GetInt`, `PrintInt`, `BeginEnd`
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
data Command = Assignment Identifier Expr
|
||||||
|
| IfThenElse Expr Command Command
|
||||||
|
| WhileDo Expr Command
|
||||||
|
| GetInt Identifer
|
||||||
|
| PrintInt Expr
|
||||||
|
| BeginEnd [Command]
|
||||||
|
|
||||||
|
type Commands = [Command]
|
||||||
|
--or
|
||||||
|
data Commands = SingleC Command
|
||||||
|
| MultipleC Command Commands
|
||||||
|
```
|
||||||
|
|
||||||
|
## Organising a Haskell Project
|
||||||
|
|
||||||
|
There are 6 Haskell modules, `Main.hs` is the entry point.
|
||||||
|
|
||||||
|
###### Defining a Module
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
module <filename> where
|
||||||
|
import ...
|
||||||
|
--definitions
|
||||||
|
newtype ...
|
||||||
|
--functions
|
||||||
|
func :: a -> b
|
||||||
|
```
|
||||||
|
|
||||||
|
Note file name must start with a capital
|
||||||
|
|
||||||
|
When you import a module, can can use functions defined in the module
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
data FileType = EXP | TAM
|
||||||
|
data Option = Trace | Run | Evaluate
|
||||||
|
|
||||||
|
main :: IO () --input output monad
|
||||||
|
```
|
||||||
|
|
||||||
|
this is the entry point, to compile
|
||||||
|
|
||||||
|
```shell
|
||||||
|
$ ghc Main.hs -o aec
|
||||||
|
$ ./aec arith_example.exp --evaluate
|
||||||
|
Evaluating Expression: 45
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
stUpdate :: st -> ST st ()
|
||||||
|
stUpdate s = S (\_ -> ((), s))
|
||||||
|
|
||||||
|
stGet :: ST st st
|
||||||
|
stGet = S (\s -> (s,s))
|
||||||
|
|
||||||
|
stRevise :: (st -> st) -> ST st ()
|
||||||
|
stRevise f = stGet >>= stUpdate . f
|
||||||
|
```
|
||||||
|
|
||||||
@@ -0,0 +1,104 @@
|
|||||||
|
# Compiling Branches
|
||||||
|
|
||||||
|
**Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output**
|
||||||
|
|
||||||
|
Before we could generate a list of instructions to be executed in sequence, now we need to implement code thats conditionally executed or executed multiple times.
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
--Code for dealing with functions and commands
|
||||||
|
commCode :: VarEnv -> Command -> TAMProg
|
||||||
|
```
|
||||||
|
|
||||||
|
We will assume an if statement looks like this
|
||||||
|
|
||||||
|
IF $e$ THEN $c_1$ ELSE $c_2$
|
||||||
|
|
||||||
|
```python
|
||||||
|
if e then c1 else c2
|
||||||
|
IfThenElse e c1 c2
|
||||||
|
```
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
expCode ve e
|
||||||
|
commCode ve c1
|
||||||
|
commCode ve c2
|
||||||
|
-- We dont want to execute both
|
||||||
|
```
|
||||||
|
|
||||||
|
- We can use `JUMPIFZ R1`, a branching function supplied by the TAM language.
|
||||||
|
|
||||||
|
- This means `commCode ve c1` & `commCode ve c2` need labels and a `JUMPA` after
|
||||||
|
|
||||||
|
```
|
||||||
|
MINI TRIANGLE PROGRAM
|
||||||
|
|
||||||
|
let var := 5
|
||||||
|
in
|
||||||
|
begin
|
||||||
|
if 1
|
||||||
|
then n := 6
|
||||||
|
else n := 7;
|
||||||
|
if 0
|
||||||
|
then n := 8
|
||||||
|
else n := 9;
|
||||||
|
end
|
||||||
|
```
|
||||||
|
|
||||||
|
```assembly
|
||||||
|
COMPILED VERSION
|
||||||
|
|
||||||
|
LOADL 5
|
||||||
|
|
||||||
|
|
||||||
|
LOAD 1 --if 1
|
||||||
|
JUMPIFZ "label1" --jump to else
|
||||||
|
LOAD 6 --load the number
|
||||||
|
STORE 0 --store 0 (stack[0] is 6 from line above) in the place of variable n, for other variables you would have to check the variable enviroment to get the stack address
|
||||||
|
JUMP "label2"
|
||||||
|
Label "label1"
|
||||||
|
|
||||||
|
LOAD 7
|
||||||
|
STORE 0
|
||||||
|
|
||||||
|
Label "label2"
|
||||||
|
LOAD 0
|
||||||
|
JUMPIFZ "label3"
|
||||||
|
```
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Generating Labels
|
||||||
|
|
||||||
|
Labels must **always** be **unique**.
|
||||||
|
|
||||||
|
This would require a global variable in our compiler to count the number of labels, haskell doesnt not allow global variables.
|
||||||
|
|
||||||
|
We can use the `stateMonad` instead.
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
type LabelName = String
|
||||||
|
|
||||||
|
fresh :: ST Int LabelName
|
||||||
|
-- Whenever we call fresh, it generates a new label name
|
||||||
|
-- We can use do (because fresh is element of ST Monad)
|
||||||
|
|
||||||
|
fresh = do
|
||||||
|
n <- stGet --checks current state (which is num of labels)
|
||||||
|
stUpdate(n+1) --update number of labels
|
||||||
|
return ("#" : (show n)) -- # symbol to denote labels
|
||||||
|
-- show converts integer to string (fresh returns string)
|
||||||
|
|
||||||
|
commCode :: VarEnv -> Command -> ST Int [TAMInsrt]
|
||||||
|
-- TAMPrgm is interchangable with [TAMInstr]
|
||||||
|
commCode ve (IFTHENELSE e c1 c2) =
|
||||||
|
do l1 <- fresh
|
||||||
|
l2 <- fresh --generate the two labels needed for an if
|
||||||
|
let te = expCode ve e --the condition expression
|
||||||
|
tc1 <- commCode ve c1 --compile success branch
|
||||||
|
tc2 <- commCode ve c2 --compile else branch
|
||||||
|
return (te ++ [JUMPIFZ l1] ++ tc1 ++ [JUMP l2]
|
||||||
|
++ [Label l1] ++ tc2 ++ [Label l2])
|
||||||
|
--then return the tam instructions
|
||||||
|
```
|
||||||
|
|
||||||
|
**REMINDER**: `expCode` is a function that takes a variable environment `ve` and an expression `e` and generates a list of TAM instructions.
|
||||||
@@ -0,0 +1,88 @@
|
|||||||
|
# Monad Revision
|
||||||
|
|
||||||
|
You can think of a monad as a container for a data type
|
||||||
|
|
||||||
|
If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype
|
||||||
|
|
||||||
|
One of the purposes of the `do` notation is to operate on the whole data structure by specify operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
|
||||||
|
|
||||||
|
$$
|
||||||
|
M_a=\{x_1, x_2, x_3,...\}
|
||||||
|
$$
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
do x <- m
|
||||||
|
let y = x ** 2 + 7
|
||||||
|
return y
|
||||||
|
```
|
||||||
|
|
||||||
|
This extracts an element of type $a$ from $m$, squares and adds 7, and returns the new values as the data structure. Now $M$ is
|
||||||
|
|
||||||
|
$$
|
||||||
|
M_b=\{y_1, y_2, y_3, \ldots\}\\or\\M=\{x_1^2+7, x_2^2+7, x_3^2+7, \ldots\}
|
||||||
|
$$
|
||||||
|
|
||||||
|
The above can be written as a functor
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
fmap (\x -> x**2+7) m
|
||||||
|
```
|
||||||
|
|
||||||
|
Monads have more functionality than functors though
|
||||||
|
|
||||||
|
If $x$ is an element of $a$ or $x :: a$
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
x :: a
|
||||||
|
return x
|
||||||
|
-- we can also write
|
||||||
|
pure x
|
||||||
|
```
|
||||||
|
|
||||||
|
Monads can have containers within containers
|
||||||
|
|
||||||
|
Assume we have function `makeBlob` that maps every element of $a$ to an element of $M_b$
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
makeBlob :: a -> Mb
|
||||||
|
makeBlob x1 = do x <- m
|
||||||
|
y <- makeBlob x
|
||||||
|
return y
|
||||||
|
-- this can be done instead with the bind operator
|
||||||
|
m >>= makeBlob
|
||||||
|
(>>=) :: Ma -> (a -> Mb) -> Mb
|
||||||
|
```
|
||||||
|
|
||||||
|
## The IO Monad
|
||||||
|
|
||||||
|
```haskell
|
||||||
|
square :: Int -> Int
|
||||||
|
square x = x*x
|
||||||
|
|
||||||
|
getInt :: IO Int
|
||||||
|
getInt = do putStrLn "Enter a number: "
|
||||||
|
s <- getLine -- getLine :: IO String
|
||||||
|
return (read s :: Int) --read :: String -> Int
|
||||||
|
|
||||||
|
squareIO :: IO Int
|
||||||
|
squareIO = do x <- getInt
|
||||||
|
let y <- square x
|
||||||
|
return y
|
||||||
|
-- as squareIO :: IO Int, returning y prints it out
|
||||||
|
|
||||||
|
squareIO :: IO () -- unit type, with only one element, also called ()
|
||||||
|
squareIO = do x <- getInt
|
||||||
|
let y <- square x
|
||||||
|
putStrLn("The square " ++ (show x) ++ " is " (show y))
|
||||||
|
return () --return unit type
|
||||||
|
-- in this case we dont even need return () as
|
||||||
|
-- putStrLn :: IO ()
|
||||||
|
|
||||||
|
--recursively asks for list unless 0 entered
|
||||||
|
getList :: IO [Int]
|
||||||
|
getList = do x <- getInt
|
||||||
|
if x == 0 then return []
|
||||||
|
else do
|
||||||
|
xs <- getList
|
||||||
|
return (x:xs)
|
||||||
|
```
|
||||||
|
After Width: | Height: | Size: 6.8 KiB |
|
After Width: | Height: | Size: 6.3 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 13 KiB |
|
After Width: | Height: | Size: 18 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 134 KiB |
|
After Width: | Height: | Size: 18 KiB |
|
After Width: | Height: | Size: 14 KiB |
@@ -0,0 +1,196 @@
|
|||||||
|
# Cryptography
|
||||||
|
|
||||||
|
**Cryptology**
|
||||||
|
|
||||||
|
> “The science and art of writing and solving codes to hide the meaning of messages.”
|
||||||
|
|
||||||
|
**Symmetric**
|
||||||
|
|
||||||
|
> “Encryption methods in which both the encryption and decryption algorithms use the same key.”
|
||||||
|
|
||||||
|
**Asymmetric**
|
||||||
|
|
||||||
|
>“Methods which use separate, but related, private and public keys.”
|
||||||
|
|
||||||
|
**Protocols**
|
||||||
|
|
||||||
|
> “The application of cryptographic algorithms in secure systems.”
|
||||||
|
|
||||||
|
**Cryptanalysis**
|
||||||
|
|
||||||
|
> “The science and art of breaking cryptosystems.”
|
||||||
|
|
||||||
|
### Modern Cyptography (1970-)
|
||||||
|
|
||||||
|
**Fundamentally different** - a scientific and mathematical discipline
|
||||||
|
|
||||||
|
**Rigorously tested** - New approaches tested, justified through mathematical proofs and theory
|
||||||
|
|
||||||
|
**Extremely powerful** - Ciphers usually take milliseconds to use and lifetimes of the universe to break
|
||||||
|
|
||||||
|
**Wider uses** - including message integrity and authenticity
|
||||||
|
|
||||||
|
**Civilian use** - everyone benefits from cryptography now
|
||||||
|
|
||||||
|
## Ciphers
|
||||||
|
|
||||||
|
- Ciphers have been used for thousands of years
|
||||||
|
- Usually based around either transposition or substitution
|
||||||
|
|
||||||
|
#### Caesar Cipher
|
||||||
|
|
||||||
|
- An early substitution cipher, we replace each letter of plain text with a shifted letter $n$ letters away from the letter
|
||||||
|
- Therefore our key is an integer $-25\leq n \leq 25$
|
||||||
|
|
||||||
|
### Modular Arithmetic
|
||||||
|
|
||||||
|
- Modular arithmetic is a system of arithmetic for finite sets of integers
|
||||||
|
- Common sets include
|
||||||
|
- $\mathbb{N} = \{1,2,3,...\}$
|
||||||
|
- $\mathbb{Z} = \{..., -3, -2, -1, 0,1,2,3,...\}$
|
||||||
|
- Also $\mathbb{Q}, \mathbb{R}, \mathbb{C}$
|
||||||
|
- Cryptography is almost always interested in finite sets
|
||||||
|
- This is useful as it avoids overflow errors
|
||||||
|
- When we add or multiply two 1 byte binary digits, the result will always be 1 byte
|
||||||
|
|
||||||
|
|
||||||
|
###### Congruence
|
||||||
|
|
||||||
|
Let $a, r, m \in \mathbb{Z}$ and $m > 0$
|
||||||
|
|
||||||
|
$a \equiv r (mod\space m)$ if $\frac{m}{a-r}$
|
||||||
|
|
||||||
|
Check:
|
||||||
|
|
||||||
|
$a=12, m=7$
|
||||||
|
|
||||||
|
$a\equiv 5 (mod\space 7)$
|
||||||
|
|
||||||
|
$\frac{7}{12-5}$ :white_check_mark:
|
||||||
|
|
||||||
|
This can be rewritten as: $a = q\cdot m+r$
|
||||||
|
|
||||||
|
###### Equivalence Classes
|
||||||
|
|
||||||
|
- The sets of all integers **mod 5** form a series of equivalence classes
|
||||||
|
- All these numbers act the same in any modluo sum
|
||||||
|
|
||||||
|
For example
|
||||||
|
|
||||||
|
$74\cdot 62 - 47 (mod \space 5) \equiv 74\%5 \cdot 62\%5 - 47\%5$
|
||||||
|
|
||||||
|
Also works with exponentiation
|
||||||
|
|
||||||
|
$3^8\space (mod\space 7)$
|
||||||
|
|
||||||
|
$3^2 = 3\cdot 3 = 9 \equiv 2\space (mod\space 7)$
|
||||||
|
|
||||||
|
$3^4 = 3^2\cdot 3^2 = 2\cdot 2 \equiv 4\space (mod\space 7)$
|
||||||
|
|
||||||
|
$3^8 = 3^4 \cdot 3^4 = 4\cdot 4 = 16 \equiv 2 \space (mod\space 7)$
|
||||||
|
|
||||||
|
#### Integer Rings
|
||||||
|
|
||||||
|
- Modular arithmetic forms what in mathematics we would call a Ring
|
||||||
|
|
||||||
|
###### Ring Definition
|
||||||
|
|
||||||
|
The integer ring $\mathbb{Z}_m$ consists of:
|
||||||
|
|
||||||
|
1. The set $\mathbb{Z}_m = \{0, 1,\ldots m-1\}$
|
||||||
|
2. Two operations $+$ and $\cdot$ for all $a, b \in \mathbb{Z}_m$ such that:
|
||||||
|
1. $a+b \equiv c \space (mod\space m), (c\in \mathbb{Z})$
|
||||||
|
2. $a\cdot b \equiv d \space (mod\space m), (d\in \mathbb{Z})$
|
||||||
|
|
||||||
|
Any time you add or multiply any two numbers in the set, the result is always in the set. We use $\equiv$ instead of $=$ as it could be an intermediatary number e.g. 12 instead of 2.
|
||||||
|
|
||||||
|
##### Properties of Rings
|
||||||
|
|
||||||
|
- We can add or multiply any two numbers in the ring, and the result is in the ring
|
||||||
|
- It is closed
|
||||||
|
- Addition and multiplication are associative
|
||||||
|
- (a+b)+c = a + (b+c)
|
||||||
|
- There is a neutral element 0 for addition
|
||||||
|
- $a + 0 \equiv a\space mod \space m$
|
||||||
|
- The additive inverse always exists
|
||||||
|
- $a + (-a) = 0\space mod \space m$
|
||||||
|
- There is a neutral element for multiplication
|
||||||
|
- $a\cdot 1 \equiv a\space mod\space m$
|
||||||
|
- The multiplicative inverse exists for some but not all elements
|
||||||
|
- $a\cdot a^{-1} \equiv 1 \space mod \space m$
|
||||||
|
|
||||||
|
#### Modular Inversion
|
||||||
|
|
||||||
|
> In rings, the multiplicative inverse exists for some but not all elements
|
||||||
|
|
||||||
|
- Multiplicative inverses allow us to *divide* by a number
|
||||||
|
|
||||||
|
$$
|
||||||
|
\frac{b}{a} \equiv b \cdot a^{-1} \space (mod \space m)
|
||||||
|
$$
|
||||||
|
|
||||||
|
- Not all numbers in a ring have an inverse, you can determine whether one exists quite simply:
|
||||||
|
|
||||||
|
$$
|
||||||
|
gcd(a,m)=1
|
||||||
|
$$
|
||||||
|
|
||||||
|
Example
|
||||||
|
|
||||||
|
$3\cdot 9 \equiv 1 \space (mod\space 26)$
|
||||||
|
|
||||||
|
$5\cdot 9 \equiv 19 \space (mod\space 26)$
|
||||||
|
|
||||||
|
$19\cdot 3 \equiv 57 \equiv 5\space (mod\space 26)$
|
||||||
|
|
||||||
|
Here a=3 and b=5, we can *divide* by 19 to get back to 5.
|
||||||
|
|
||||||
|
#### Shift Cipher
|
||||||
|
|
||||||
|
We can formalise the shift cipher using modular arithmetic
|
||||||
|
|
||||||
|
Let $x, y, k \in \mathbb{Z}_{26}$
|
||||||
|
|
||||||
|
$$
|
||||||
|
e_k(x) = y \equiv x+k \space (mod \space 26) \\
|
||||||
|
d_k(y) = x \equiv y-k \space (mod \space 26)
|
||||||
|
$$
|
||||||
|
|
||||||
|
##### Frequency Analysis
|
||||||
|
|
||||||
|
- The frequency of occurrences of each character are very consistent
|
||||||
|
- The longer a cipher text is, the easier this becomes
|
||||||
|
|
||||||
|
#### Affine Cipher
|
||||||
|
|
||||||
|
We can extend the shift cipher into an affine cipher
|
||||||
|
|
||||||
|
Let $x,y,a,b \in \mathbb{Z}_{26}$
|
||||||
|
|
||||||
|
$$
|
||||||
|
e_k(x) = y \equiv a\cdot x+b\space (mod \space 26)\\
|
||||||
|
d_k(y) = x \equiv a^{-1}\cdot(y-b)\space (mod \space m)
|
||||||
|
$$
|
||||||
|
|
||||||
|
where $k=(a,b)$ and $gcd(a,26)=1$
|
||||||
|
|
||||||
|
This is a multiplication and a addition analagous to $y=mx+c$
|
||||||
|
|
||||||
|
In a Affine cipher, letters can be themselves
|
||||||
|
|
||||||
|
- The keyspace of an affine cipher
|
||||||
|
- a can be 0-25
|
||||||
|
- b can be 0-12
|
||||||
|
- 25*12=300
|
||||||
|
- More secure than a caesar cipher
|
||||||
|
|
||||||
|
Frequency analysis can still be used, in this case the columns will not only be shifted, but jumbled aswell.
|
||||||
|
|
||||||
|
- This is not hard to crack
|
||||||
|
|
||||||
|
#### The Vigenere Cipher
|
||||||
|
|
||||||
|
- An early stream cipher, the Vigenere cipher is a shift cipher with a running key
|
||||||
|
- Unlike caesar cipher, the key is repeated for as long as required.
|
||||||
|
- It is the equivalent to multiple interleaved Caesar ciphers
|
||||||
|
- Spreads outs occurrances of characters making frequency analysis hard.
|
||||||
@@ -0,0 +1,158 @@
|
|||||||
|
# Stream Ciphers
|
||||||
|
|
||||||
|
Stream ciphers encrypt bits one at a time, for as long as necessary.
|
||||||
|
|
||||||
|
Stream ciphers using modulo 2 addition
|
||||||
|
|
||||||
|
Let $x, y, s \in \{0,1\}$
|
||||||
|
|
||||||
|
**Encryption**: $e_{s_i} (x_i) = y_i \equiv x_i + s_i \space (mod\space 2)$
|
||||||
|
|
||||||
|
**Decryption**: $d_{s_i} (y_i) = x_i \equiv y_i + s_i \space (mod\space 2)$
|
||||||
|
|
||||||
|
Why does mod 2 work for both encryption and decryption?
|
||||||
|
|
||||||
|
$$
|
||||||
|
d_{s_i} (y_i) \equiv y_i + s_i \space (mod\space 2) \\
|
||||||
|
d_{s_i} (y_i) \equiv (x_i + s_i)+s_i \space (mod\space 2) \\
|
||||||
|
d_{s_i} (y_i) \equiv (x_i + 2s_i) \space (mod\space 2) \\
|
||||||
|
d_{s_i} (y_i) \equiv (x_i + 0\cdot s_i \space (mod\space 2) \\
|
||||||
|
d_{s_i} (y_i) \equiv x_i
|
||||||
|
$$
|
||||||
|
|
||||||
|
Note: 2 % 2 is 0, its like **xor**-ing twice.
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
#### Security of XOR
|
||||||
|
|
||||||
|
| $x_i$ | $s_i$ | $y_i$ |
|
||||||
|
| :---: | :---: | :---: |
|
||||||
|
| 0 | 0 | 0 |
|
||||||
|
| 0 | 1 | 1 |
|
||||||
|
| 1 | 0 | 1 |
|
||||||
|
| 1 | 1 | 0 |
|
||||||
|
|
||||||
|
When $y_i$ is 1, it could’ve been from the message or the key.
|
||||||
|
|
||||||
|
### Randomness
|
||||||
|
|
||||||
|
The security of a stream cipher depends entirely on the nature of the key stream
|
||||||
|
|
||||||
|
- If the stream is truly random, the output is truly random.
|
||||||
|
|
||||||
|
##### True Randomness
|
||||||
|
|
||||||
|
- True randomness is impossible to recreate except by chance
|
||||||
|
- coin flips
|
||||||
|
- Computer systems often use hardware sources for randomness
|
||||||
|
- Thermal or other noise
|
||||||
|
- Radioactive decay
|
||||||
|
- Clock drift
|
||||||
|
- Random timings of interrupts
|
||||||
|
|
||||||
|
##### Pseudo Randomness
|
||||||
|
|
||||||
|
- Generate a sequence of values based on a seed
|
||||||
|
- Usually the only requirement is statistical randomness
|
||||||
|
|
||||||
|
###### Linear Congruential Generator
|
||||||
|
|
||||||
|
Cs `rand()` function, this is a PRNG
|
||||||
|
|
||||||
|
$$
|
||||||
|
s_0 = 12345 \\
|
||||||
|
s_{i+1} \equiv 1103515245 \cdot s_i + 12345 \space (mod \space 2^{32})
|
||||||
|
$$
|
||||||
|
|
||||||
|
##### Cryptographically Secure Pseudo Randomness
|
||||||
|
|
||||||
|
- Is a PRNG whose output is unpredictable
|
||||||
|
- Given n bits of key stream, can we predict the next bit x?
|
||||||
|
|
||||||
|
$$
|
||||||
|
Pr[x=s_{n+1}] < 0.5 + \epsilon
|
||||||
|
$$
|
||||||
|
|
||||||
|
#### Unconditional Security
|
||||||
|
|
||||||
|
A crypto-system is **unconditional security** is unconditionally or information-theoretically secure if it cannot be broken, even with infinite computational resources.
|
||||||
|
|
||||||
|
**Perfect Secrecy**: The cipher-text should reveal no information about the plain text
|
||||||
|
|
||||||
|
$\forall_{m_0, m_1} \in M$ where $|m_0| = |m_1|$ and $\forall_c \in C$
|
||||||
|
|
||||||
|
$Pr[E(k,m_0) = c] = Pr[E(k,m_1) = c]$
|
||||||
|
|
||||||
|
The probability that $m_0$ encrypts to $c$ is the same as the probability of $m_1$ also encrypted to $c$
|
||||||
|
|
||||||
|
## One Time Pad
|
||||||
|
|
||||||
|
- Key stream generated by a TRNG
|
||||||
|
- The key stream is known only to the communicating parties
|
||||||
|
- Every key stream but $s_i$ is used only once
|
||||||
|
|
||||||
|
#### OTP has perfect Secrecy
|
||||||
|
|
||||||
|
**Proof**
|
||||||
|
|
||||||
|
$\forall m, c : Pr[E(k,m)=c] = \frac{\{k\in K|E(k,m)=c\}}{|K|}$
|
||||||
|
|
||||||
|
For every message, that encrypts to cipher text, the probability of m encrypting to c, is all the keys over all the messages
|
||||||
|
|
||||||
|
For OTP:
|
||||||
|
|
||||||
|
#$\{k\in K| E(k, m) = c\} = 1$
|
||||||
|
|
||||||
|
Because a key is only used once
|
||||||
|
|
||||||
|
because if $E(k,m)=c$ then $k=m \oplus c$
|
||||||
|
|
||||||
|
$\therefore Pr[E(k, m_0) = c] = Pr[E(k, m_1) = c]$
|
||||||
|
|
||||||
|
- Any plaintext is equally likely depending on the key
|
||||||
|
- This is an example where $M = C-K\space (mod\space 26)$
|
||||||
|
|
||||||
|
OTP is *not practical*:
|
||||||
|
|
||||||
|
- A 1GB file would need a 1GB key
|
||||||
|
- How are we transporting these keys & storing them
|
||||||
|
- If you ever reuse a key, the entire cipher is broken
|
||||||
|
|
||||||
|
## Modern Stream Ciphers
|
||||||
|
|
||||||
|
- Modern stream ciphers use an initial seed key to generate an infinite pseudo-random keystream
|
||||||
|
- Reusing keys catastrophically breaks the encryption
|
||||||
|
|
||||||
|
$$
|
||||||
|
M_1 \oplus K = C_1 \quad\quad M_2 \oplus K = C_2 \\
|
||||||
|
C_1 \oplus C_2 = (M_1 \oplus K) \oplus (M_2 \oplus K) \\
|
||||||
|
= (K \oplus K) \oplus M_1 \oplus M_2 \\
|
||||||
|
= 0 \oplus M_1 \oplus M_2 \\
|
||||||
|
= M_1 \oplus M_2 \\
|
||||||
|
$$
|
||||||
|
|
||||||
|
#### Crib Dragging
|
||||||
|
|
||||||
|
This involves guessing $M_1$, this can be a common message such as `HTTP` request.
|
||||||
|
|
||||||
|
This can be automated by checking $M_1$ over different parts of $M_2$.
|
||||||
|
|
||||||
|
- Stream ciphers use a *nonce* value to alter the keystream for a given key
|
||||||
|
- This allows us to create *different key-streams* for a given key
|
||||||
|
|
||||||
|
#### Number Used Once
|
||||||
|
|
||||||
|
- Numbers used once or *nonces* are vital for stream cipher security
|
||||||
|
- Instead of always using a unique key, the security requirement is you always use a unique (key + nonce) pair
|
||||||
|
- Nonces are not secret, they are public random seed for a key stream
|
||||||
|
|
||||||
|
### Could we use a LCG?
|
||||||
|
|
||||||
|
- LCG - Linear congruential generators
|
||||||
|
- Seed using some key, then
|
||||||
|
- $s_{i+1} \equiv A \cdot s_i + B \space (mod \space 2)$
|
||||||
|
- $s_i, A, B$ are $log_2m$ bits long
|
||||||
|
- This is trivial to break
|
||||||
|
- Given known plaintext $x_1, x_2, x_3$
|
||||||
|
- Calculate corresponding key $s_1, s_2, s_3$
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
# Modern Stream Ciphers
|
||||||
|
|
||||||
|
### Pseudo-randomness
|
||||||
|
|
||||||
|
- TRNGs - true Random Number Generator
|
||||||
|
- Not feasible at scale
|
||||||
|
- PRNGs - Pseudo Random Number Generator
|
||||||
|
- CSPRNGs - Cryptographically Secure Pseudo Random Number Generator
|
||||||
|
|
||||||
|
#### LFSRs
|
||||||
|
|
||||||
|
- A Linear-feedback Shift Register us a register if buts whose positions shift to the right
|
||||||
|
- Usually comprised of flip-flops, the last bit represents the output
|
||||||
|
|
||||||
|
(Where the squares at the bottom are flip-flops)
|
||||||
|
|
||||||
|
- If initialised to `000`, nothing happens as $0\oplus0 = 0$.
|
||||||
|
- Therefore, we have $2^n-1$ states
|
||||||
|
- Statistical randomness
|
||||||
|
- To add more randomness to the setup, we can add another (more) `xor` gate
|
||||||
|
- However, we have fewer states
|
||||||
|
|
||||||
|
$$
|
||||||
|
s_m \equiv s_{m-1}p_{m-1} + ... + s_1p_1 + s_0p_0\space (mod \space 2)\\
|
||||||
|
s_{m+1} \equiv s_{m}p_{m-1} + ... + s_2p_1 + s_1p_0\space (mod \space 2)
|
||||||
|
$$
|
||||||
|
|
||||||
|
- We usually represent m-bit LFSRs using polynomials of degree m.
|
||||||
|
- In general $P(x)=x^m + p_{m-1}x^{m-1} + ... + p_1x + p_0$
|
||||||
|
- LFSRs that have primitive polynoimials produce sequences of maximum length
|
||||||
|
- There are many and are easily computed
|
||||||
|
- $x^5 + x^2 + 1$ has 31 states
|
||||||
|
- $x^{10} + x^3 + 1$ has 1023
|
||||||
|
- $x^{85}+x^8+x^2+x+1$ has $10^{26}$ states
|
||||||
|
|
||||||
|
##### Attacking LFSRs
|
||||||
|
|
||||||
|
Suppose an attacker knows $2m-1$ plain text bits
|
||||||
|
|
||||||
|
**Step 1** Calculate key bits
|
||||||
|
|
||||||
|
$s_i \equiv y_i + x_i \space (mod\space 2), i=0,1,...2_{m-1}$
|
||||||
|
|
||||||
|
**Step 2** Reconstruct the LFSR
|
||||||
|
|
||||||
|
$s_m \equiv s_{m-1}p_{m-1} + ... + s_1p_1 + s_0p_0$
|
||||||
|
|
||||||
|
$s_{m+1} \equiv s_{m}p_{m-1} + ... + s_2p_1 + s_1p_0$
|
||||||
|
|
||||||
|
…
|
||||||
|
|
||||||
|
$s_{2m+1} \equiv s_{2m-1}p_{m} + ... + s_mp_1 + s_{m-1}p_0$
|
||||||
|
|
||||||
|
#### Trivium
|
||||||
|
|
||||||
|
- LFSRs are much more cryptographically secure if we combine more than one together in a non-linear way.
|
||||||
|
|
||||||
|
Trivium is 3 LFSR in a row
|
||||||
|
|
||||||
|
- Feedback between each with non-linear AND gates
|
||||||
|
- Initialises the LFSR with an 80-bit key and 80-bit random value
|
||||||
|
|
||||||
|
### ChaCha20
|
||||||
|
|
||||||
|
- ChaCha is a stream cipher written by Daniel Berstein
|
||||||
|
- A modification of a previous cipher, Salsa
|
||||||
|
- Very lightweight, using only `add`, `xor` and rotate operations
|
||||||
|
- One of two ciphers in `TLS 1.3`
|
||||||
|
- Dashes represent bit length
|
||||||
|
- Constants are not secret
|
||||||
|
- The block number can skip to anywhere
|
||||||
|
- Suppose someone skips ahead on a video stream, the cipher can skip unlike other synchronous stream ciphers
|
||||||
|
- Works well on low power devices, due to simplicity of encryption
|
||||||
|
- Once the input and the mixed words are added together it is hard to know what the starting thing was
|
||||||
|
- e.g. what two numbers have i added to make 100
|
||||||
|
|
||||||
|
ChaCha performs **20** rounds
|
||||||
|
|
||||||
|
- Alternates column and diagonal rounds
|
||||||
|
- Each round is 4 quarter rounds
|
||||||
|
|
||||||
|
#### Vulnerabilities
|
||||||
|
|
||||||
|
- Stream ciphers like ChaCha give us *confidentiality*, but *not integrity*
|
||||||
|
- Running a stream cipher by itself is not sufficient
|
||||||
@@ -0,0 +1,148 @@
|
|||||||
|
# Data Encryption Standard (DES)
|
||||||
|
|
||||||
|
### Pseudorandom Permutations
|
||||||
|
|
||||||
|
- A pseudorandom permutation is a function that cannot be distinguished from a random permutation
|
||||||
|
- Maps a set of values $\{0,1\}^n \times \{0,1\}^s \rightarrow \{0,1\}^n$ such that:
|
||||||
|
- For any key, the function F is a *bijection* (1:1)
|
||||||
|
- The key just changes the mapping
|
||||||
|
- There is an *efficient algorithm* to calculate $F(x)$ for all keys and all messages
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**Confusion**: Obscure the relationship between plaintext, key and ciphertext
|
||||||
|
|
||||||
|
- Often achieved through substitution operations
|
||||||
|
- Using lookup tables
|
||||||
|
|
||||||
|
**Diffusion**: Influence of each plaintext and key bit is distributed throughout the ciphertext
|
||||||
|
|
||||||
|
- Achieved via permutation
|
||||||
|
- Swapping or otherwise mixing bits/bytes
|
||||||
|
|
||||||
|
Shannon called a cipher like this a **product cipher**
|
||||||
|
|
||||||
|
### Feistal Network
|
||||||
|
|
||||||
|
- A Feistal Network is one mechanism used to create block ciphers
|
||||||
|
- Developed by Horst Feistal while he worked at IBM
|
||||||
|
- Underpins DES, GOST, Blowfish, Twofish and numerous others.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- To decrypt, we run the encrypted bits through the network again
|
||||||
|
|
||||||
|
##### A Single Feistal Round
|
||||||
|
|
||||||
|
- During each round, only half of the block is encrypted
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Very similar to a stream cipher.
|
||||||
|
|
||||||
|
###### Round $i$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
###### Round $i+1$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Note - the last round does a final swap so the left and right are in the correct places.
|
||||||
|
|
||||||
|
###### Decrypting
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Basically the `xor`s cancel themselves out, the most important part is choosing a good function $f$
|
||||||
|
|
||||||
|
#### About Feistal Networks
|
||||||
|
|
||||||
|
- 1 or 2 rounds is not sufficient
|
||||||
|
- Luby and Rackoff show that if $f$ is a cryptographically secure pseudorandom function then:
|
||||||
|
- 3 rounds are sufficient to make a pseudorandom permutation
|
||||||
|
- 4 rounds are sufficient to make a strong pseudorandom permutation
|
||||||
|
- Balanced Feistal networks
|
||||||
|
- L and R are equal sizes
|
||||||
|
- Unbalanced feistal networks
|
||||||
|
- L and R can be different sizes
|
||||||
|
- e.g. `skipjack`, `OAEP`
|
||||||
|
|
||||||
|
### DES
|
||||||
|
|
||||||
|
1972: NIST put out a call for a US standard for encryption
|
||||||
|
|
||||||
|
1974: IBM propose DES
|
||||||
|
|
||||||
|
1976: NIST accepts an altered version of DES following consultation with NSA
|
||||||
|
|
||||||
|
- Feistal network with 64-bit block size
|
||||||
|
- 56-bit key
|
||||||
|
- The most studied cipher in history
|
||||||
|
- Hasn’t been broken for over 46 years
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
This speeds up loading bits into registers
|
||||||
|
|
||||||
|
##### The F function
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
$S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits output based on lookup tables. The lookup tables for each s-box is different.
|
||||||
|
|
||||||
|
##### Expansion
|
||||||
|
|
||||||
|
- Adds *diffusion*
|
||||||
|
- Increases from 32 to 48 bits to match the round key
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
> Half the input bits are connected to two output positions
|
||||||
|
|
||||||
|
##### Substitution Boxes
|
||||||
|
|
||||||
|
- Add confusion
|
||||||
|
- The s-boxes map 6 bit inputs to 4-bit outputs
|
||||||
|
- There are 8 s-boxes in total, each is different
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- This s-box is not random, very carefully designed
|
||||||
|
- s-boxes need to be highly **non-linear**: $S(a) \oplus S(b) \neq S(a\oplus b)$
|
||||||
|
- This prevents simple systems of linear equations such as we saw in LFSRs.
|
||||||
|
- The formula needed to represent DES is too complicated
|
||||||
|
- Key design principles
|
||||||
|
1. No output bit should be too close to a linear combination of input bits
|
||||||
|
2. 1-bit change input should lead to at least 2-bits output
|
||||||
|
3. If you only change the 4 middle bits, each output must occur exactly once
|
||||||
|
4. If the first two bits are different but the last two are identical, the output must differ
|
||||||
|
5. For any non-zero difference in input, no more than 8 of the 32 inputs exhibiting this difference should share the same output difference
|
||||||
|
- We want to limit the number of predictable swaps
|
||||||
|
6. A collision (zero difference) is only possible for 3 adjacent s-boxes
|
||||||
|
|
||||||
|
##### Permutation
|
||||||
|
|
||||||
|
- At the end of $f()$ is a permuatation
|
||||||
|
- This moves bits between s-boxes on the next round
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Blue showing how $S_1$ output bits are diffused
|
||||||
|
|
||||||
|
#### The Avalanche Effect
|
||||||
|
|
||||||
|
If you input all 0s, we will see a random cipher text
|
||||||
|
|
||||||
|
However if we change one 0 to a 1, how does this effect the result.
|
||||||
|
|
||||||
|
- On average, if you change one (first) bit in $R$, one bit will change in the expansion
|
||||||
|
- Due to the way the s-boxes are setup, at least 2 of the 4 bits in the output will be different
|
||||||
|
- Now when the permutation happens, these two changes are spread to other s-boxes
|
||||||
|
- Now next round we’ll get 4 changes, then 8, then 16 …
|
||||||
|
|
||||||
|
For DES the worst case scenario when one bit is changed (with 5 rounds) is there will be an effect on every bit on the output.
|
||||||
@@ -0,0 +1,155 @@
|
|||||||
|
# Data Encryption Standard
|
||||||
|
|
||||||
|
#### Key Schedule
|
||||||
|
|
||||||
|
- The **DES** key schedule simply returns various permutations of $k$ as sub-keys
|
||||||
|
- $k_1, ... k_{16}$
|
||||||
|
|
||||||
|
##### PC-1
|
||||||
|
|
||||||
|
- Permutated Choice 1 (PC-1) selects 56 of the 64 bits
|
||||||
|
- The other ‘parity’ bits are discarded: DES only uses a 56-bit key
|
||||||
|
- Key bits are spread throughout the initial state of the key schedule
|
||||||
|
- Key bits 8, 16, 24,…64 are not used
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Left Rotation
|
||||||
|
|
||||||
|
- Left rotations (often written as `<<<`) represent a lift shift where the left most numbers wrap around to the right hand side
|
||||||
|
- In DES, each 28-bit block is rotated left by `<<<1` for rounds 1,2,9,16 and `<<<2` otherwise
|
||||||
|
- The total rotation is $4\cdot 1 + 12\cdot 2 = 28$ which means $C_0 = C_{16}$ and $D_0 = D_{16}$
|
||||||
|
- NOTE: $C_0$ or $D_0$ is not used
|
||||||
|
|
||||||
|
##### PC-2
|
||||||
|
|
||||||
|
- Permuted Choice 2 select 48 of the 56 bits to be used as a round key
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
###### Properties of the Key Schedules
|
||||||
|
|
||||||
|
- Is entirely permutation based
|
||||||
|
- Doesn’t use `xor`, addition or any other mixing operation
|
||||||
|
- Because $C_0 = C_{16}$ and $D_0 = D_{16}$ we don’t need to write seperate encrpt and decrypt functions
|
||||||
|
- Usful for writing implementations on low memory devices (smart cards)
|
||||||
|
|
||||||
|
### Breaking DES
|
||||||
|
|
||||||
|
- DES has a key length of 56-bits
|
||||||
|
- A brute force attack requires no knowledge of the cipher, only a pair $(x_0, y_0)$ of known plain and cipher text
|
||||||
|
|
||||||
|
$DES^{-1}k_i(y_0) = x_0$ for $i=0, 1, ... 2^{56}-1$
|
||||||
|
|
||||||
|
This would take minutes to hours on a cluster.
|
||||||
|
|
||||||
|
NOTE: $2^{56}-1$ is a very large number
|
||||||
|
|
||||||
|
#### Key Collisions
|
||||||
|
|
||||||
|
- For a 56-bit key but a 64-bit block is possible (though unlikely) a different key would work
|
||||||
|
- How likely is this to happen for a 1 bit key and an $n$ bit block cipher
|
||||||
|
- $\frac{2^l}{2^n}$ where $l$ is the length of the block and $n$ is the key length
|
||||||
|
- $\frac{2^{64}}{2^{56}} = 2^8$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- DES was first brute forced in 1997 and is no longer secure
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
(days on y axis)
|
||||||
|
|
||||||
|
#### Double Encryption
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Naive brute fource suggests $2^{56}\cdot 2^{56} = 2^{112}$ keyspace
|
||||||
|
- However using a meet-in-the middle attack this becomes trival.
|
||||||
|
- Step 1: Calculate encryptions of $x_1$ for all $k_{1...,i}$ and store intermediate values $Z_{1..,i}$
|
||||||
|
- Step 2: Calculate all decryptions of $y_1$ for all $k_{R, j}$ to find $Z_{R,i}$
|
||||||
|
- Step 3: Find any value of $Z_{R,j}$ matching existing $Z_L,i$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Meet-in-the-middle requires $2^{k+1}$ attemps rather than $2^{k\cdot 2}$
|
||||||
|
|
||||||
|
- This is much better than brute force, but doesn’t make it easy
|
||||||
|
- Trades off computation for storage - Petabytes for DES
|
||||||
|
- Assumes some kind of $O(1)$ for $Z_{L,I}$
|
||||||
|
|
||||||
|
## 3DES
|
||||||
|
|
||||||
|
- Triple DES uses three different keys
|
||||||
|
- Either `enc -> enc -> enc` or `enc -> dec -> enc`
|
||||||
|
- Often used in banking, smart cards and other payment systems
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
This prevents MITM attacks as one of the attacks will have to compute $2^{112}$ permutations
|
||||||
|
|
||||||
|
Why use `enc -> dec -> enc`?
|
||||||
|
|
||||||
|
This is for compatibility with legacy systems running DES.
|
||||||
|
|
||||||
|
This is why banking systems use 3DES as they already have the infrastructure for DES however 3DES is officially not recommended by NSA in 2016
|
||||||
|
|
||||||
|
## DES-X
|
||||||
|
|
||||||
|
- An alternative construction using a concept called **key-whitening**
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Theoretically this provides a seach space of $2^{k+2n}$ but meet-in-the-middle can be used here, as well as other more advanced attacks
|
||||||
|
- In practive securtity is $2^{k+n-m}$ where an attack has $2^m$ known plain texts
|
||||||
|
|
||||||
|
# Cryptanalysis
|
||||||
|
|
||||||
|
#### What is a break?
|
||||||
|
|
||||||
|
- In modern cryptography, a cipher is declared broken by essentially any attack that is more efficient than brute force
|
||||||
|
- For example, *differential cryptanalysis* requires $2^{47}$ operations on DES rather than $2^{56}$
|
||||||
|
- These are often academic breaks, rather than a practical security concern
|
||||||
|
- For example there is a *related key* attack on AES of $2^{99.5}$, compared to brute force of $2^{128}$
|
||||||
|
- Remember that a $2^{n-1}$ takes half the time $2^n$ does
|
||||||
|
|
||||||
|
##### Analytical Attacks
|
||||||
|
|
||||||
|
- Exploit some underlying structureal or mathematical weakness in a cipher
|
||||||
|
- e.g. meet in the middle attack
|
||||||
|
- Derivation of taps in LFSRs
|
||||||
|
|
||||||
|
##### Statistical Attacks
|
||||||
|
|
||||||
|
- Capture statistical patterns between input and output to recover key bits
|
||||||
|
- Differential cryptanalysis
|
||||||
|
- Linear cryptanalysis
|
||||||
|
|
||||||
|
###### Differential Cryptanalysis
|
||||||
|
|
||||||
|
- Different cryptanalysis is prehaps now the most important modern method for breaking block ciphers
|
||||||
|
- It is a **chosen plaintext** attack
|
||||||
|
- We aim to find predictable changes in output bits caused by known changes in the input bits
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Each of these s boxes has 4 bits, 16 possible values
|
||||||
|
- This means any input change $\Delta x$ should cause some change $\Delta y$ with probability $p=1/16$
|
||||||
|
- In a poor s-box, the likelihood might be much higher
|
||||||
|
- The sum input change resulting in some output change $(\Delta x, \Delta y)$ is called a **differential** and has some probability of occurring
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Tracing differentials through a cipher provides us with **differential characteristics** e.g.
|
||||||
|
- $(\Delta x, \Delta y) =$ (0x80, 0xA0) where $p \geq 2^{-3} = 1/8$
|
||||||
|
- These can be calculated by hand or using automated tools
|
||||||
|
- The attack then looks for these expected differentials as you manipulate sub-key bits
|
||||||
|
|
||||||
|
###### Resisting differential cryptanalysis
|
||||||
|
|
||||||
|
- S-boxes must be designed such that the probability of any pair $(\Delta x, \Delta y)$ is as low as possible
|
||||||
|
- AES has a maximum likelihood of a differential per s-box of $2^{-6}$
|
||||||
|
- This is because AES has such good diffusion
|
||||||
|
- More rounds make differentials even less likely
|
||||||
|
- Good permuation to involve more s-boxes is vital
|
||||||
|
- DES was specifically designed to resist this kind of attack
|
||||||
@@ -0,0 +1,144 @@
|
|||||||
|
# Finite Field Arithmetic
|
||||||
|
|
||||||
|
- A **finite field** is a set containing a finite number of elements
|
||||||
|
- This is sometimes called a *Galois Field*
|
||||||
|
- In a Galois field you can:
|
||||||
|
- Add
|
||||||
|
- Subtract
|
||||||
|
- Multiply
|
||||||
|
- Invert (divide)
|
||||||
|
- Fields are an extension of *groups* and related to *rings*
|
||||||
|
|
||||||
|
### Groups
|
||||||
|
|
||||||
|
A group is a set of elements $G$ together with an operation $\circ$ that combines two elements of $G$
|
||||||
|
|
||||||
|
> 1. The operation $\circ$ is **closed**
|
||||||
|
> - i.e. for all $a,b \in G$ then $a\circ b=c\in G$
|
||||||
|
> 2. The operation is associative
|
||||||
|
> - i.e. $a\circ(b\circ c) = (a\circ b)\circ c$ for all $a,b,c \in G$
|
||||||
|
> 3. There is an element $1\in G$ called a **neutral element** such that $a\circ 1 = 1\circ a = a$ for all $a\in G$
|
||||||
|
> 4. For each $a \in G$ there exists an element $a^{-1}\in G$ called the **inverse** of $a$ such that $a\circ a^{-1} = a^{-1}\circ a = 1$
|
||||||
|
> 5. A group $G$ is **abelian** (commutative) if $a\circ b = b \circ a$ for all $a,b\in G$
|
||||||
|
|
||||||
|
##### Example Group
|
||||||
|
|
||||||
|
- The set of integers $\mathbb{Z}_m = \{0,1,...m-1\}$ with the operation addition modulo m form a group with the neutral element 0
|
||||||
|
- Every element would have an inverse where $a + (-a) = 0$ mod m
|
||||||
|
- This group would not form a group with multiplication, as not all elements would have an inverse
|
||||||
|
- We wouldn’t have an inverse, we would need $5\times \frac15=1$ however $\frac15 \notin \mathbb{Z}$
|
||||||
|
|
||||||
|
### Fields
|
||||||
|
|
||||||
|
A field $F$ is a set of elements with the following properties
|
||||||
|
|
||||||
|
> 1. All elements of $F$ form an **additive group** with the group operation $+$ and the neutral element 0
|
||||||
|
> 2. All elements of $F$ except 0 form a multiplicative group with the group operation $\times$ and the neutral element 1
|
||||||
|
> 3. When the two group operations are mixed, the distributivity law holds.
|
||||||
|
> - i.e. for all $a,b,c \in F, a\cdot(b+c) = (a\cdot b) + (a\cdot c)$
|
||||||
|
|
||||||
|
##### Example Field
|
||||||
|
|
||||||
|
- The set of real numbers $\mathbb{R}$ is a field with neutral element 0 for addition and 1 for multiplication
|
||||||
|
- Every real number $a$ has a additive inverse $-a$
|
||||||
|
- Every non-zero number $a$ has a multiplicative inverse $\frac{1}{a}$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Finite Fields
|
||||||
|
|
||||||
|
> A finite field only exists if it has $p^m$ elements
|
||||||
|
>
|
||||||
|
> Where:
|
||||||
|
>
|
||||||
|
> - $p$ is a prime
|
||||||
|
> - $m$ is a positive integer
|
||||||
|
|
||||||
|
###### Examples
|
||||||
|
|
||||||
|
- There is a field with 11 elements: $GF(11)$
|
||||||
|
- There is a field with 256 elements: $GF(256)$ or $GF(2^8)$
|
||||||
|
- $GF(12)$ is not a finite field $(2^2 \cdot3)$
|
||||||
|
|
||||||
|
###### Prime and Extension Fields
|
||||||
|
|
||||||
|
When $m=1$ it creates a **prime field**
|
||||||
|
|
||||||
|
When $m>1$ it creates an **extension field**
|
||||||
|
|
||||||
|
### Prime Fields
|
||||||
|
|
||||||
|
- A prime field $GF(p)$ contains the integers $\{0,1,...p-1\}$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- These operations satisfy the properties of fields (*closure*)
|
||||||
|
|
||||||
|
##### Inversion in Prime Fields
|
||||||
|
|
||||||
|
$a \cdot a^{-1} \equiv 1 \space (mod \space p)$
|
||||||
|
|
||||||
|
- A modular inverse exists when $gcd(a,p) = 1$
|
||||||
|
- Because $p$ is prime, every number has a multiplicative inverse
|
||||||
|
- $gcd(a,p) = 1, \forall a \neq0 \in GF(p)$
|
||||||
|
- $a^{-1}$ can be calculated using the **extended Euclidean algorithm**
|
||||||
|
|
||||||
|
#### Extension Fields
|
||||||
|
|
||||||
|
- In prime fields, the elements are integers
|
||||||
|
- Elements in extension fields $GF(2^m)$ are polynomials of degree $m$
|
||||||
|
|
||||||
|
$a_{m-1}x^{m-1}, ..., a_1x + a_0 = A(x) \in GF(2^m)$
|
||||||
|
|
||||||
|
where $a_i \in GF(2) = \{0,1\}$
|
||||||
|
|
||||||
|
The coefficients of the polynomial are elements in $GF(2)$ the **sub-field**
|
||||||
|
|
||||||
|
##### Example $GF(2^3)$
|
||||||
|
|
||||||
|
- The field $GF(2^3)$, sometimes called $GF(8)$ is an extension field containing elements of the form: $A(x) = a_2 x^2 + a_1x^1 + a_0$
|
||||||
|
- Its often easier to simply write the coefficients $(a_2, a_1, a_0)$ e.g. 001 or 101
|
||||||
|
- $GF(2^3) = \{0, 1, x, x+1, x^2, x^2+1, x^2 + x, x^2 + x + 1\}$
|
||||||
|
- $|GF(2^3)| = 8$
|
||||||
|
|
||||||
|
#### Arithmetic in $GF(2^3)$
|
||||||
|
|
||||||
|
- Adding or subtracting two polynomials happens as expected, but adding the coefficients
|
||||||
|
- $A(x) = x^2 + x + 1$
|
||||||
|
- $B(x) = x^2 + 1$
|
||||||
|
- $A(x) + B(x) = (1+1)x^2 + (1)x + (1+1) = x$
|
||||||
|
- mod 2 is simply `xor`
|
||||||
|
- Addition and subtraction are identical
|
||||||
|
|
||||||
|
#### Multiplication in $GF(2^3)$
|
||||||
|
|
||||||
|
- $A(x) = x^2 + x + 1$
|
||||||
|
- $B(x) = x^2 + 1$
|
||||||
|
- $A(x) \cdot B(x) = (x^2 + x + 1)(x^2 + 1) = x^4 + x^3 + (1+1)x^2 + x + 1$
|
||||||
|
- $x^4 + x^3 + x + 1$ however this is **not in the field**
|
||||||
|
- The result must be reduced by the result modulo an **irreducible polynomial**
|
||||||
|
|
||||||
|
$$
|
||||||
|
A(x) \cdot B(x) = x^4 + x^3 + x + 1\space (mod \space x^3 + x + 1)
|
||||||
|
$$
|
||||||
|
|
||||||
|
- This means we have to do polynomial long division
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Inversion
|
||||||
|
|
||||||
|
- Inversion is performed in a similar way to prime fields, we find:
|
||||||
|
- $A(x) \cdot A^{-1}(x) \equiv 1 \space (mod \space P(x))$
|
||||||
|
- $A^{-1}(x)$ is calculated using the extended euclidean algorithm
|
||||||
|
|
||||||
|
### AES’ Finite Field
|
||||||
|
|
||||||
|
- AES uses the extension field $GF(2^8)$ for many of its operations
|
||||||
|
- Operations are the same as those in other $GF(2^m)$ fields, using the irreducible polynomial
|
||||||
|
|
||||||
|
$$
|
||||||
|
P(x) = x^8 + x^4 + x^3 + x + 1
|
||||||
|
$$
|
||||||
|
|
||||||
|
- As you might expect, these polynomials are typically represented as single bytes
|
||||||
@@ -0,0 +1,145 @@
|
|||||||
|
# Advanced Encryption Standard (AES)
|
||||||
|
|
||||||
|
- AES superseded DES as a standard in 2002
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Uses rounds of 4 layers and a final round of 3
|
||||||
|
- Bytes are represented as a 4x4 block called the *state*
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**Sub-Bytes** - similar to s-boxes in DES
|
||||||
|
|
||||||
|
**Shift Rows** - diffusion and permutation round
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
First row doesn’t move, second row is shifted to the left by 1, the third row is shifted two places to the left etc
|
||||||
|
|
||||||
|
Then, when the columns are mixed, this means the overall diffusion is extremely good
|
||||||
|
|
||||||
|
The last round doesn’t have a **mix column** step as its reversible and wouldn’t add additional security.
|
||||||
|
|
||||||
|
#### S-Box
|
||||||
|
|
||||||
|
- The AES s-box is based around the multiplicative inverse of 8-bit values in $GF(2^8)$
|
||||||
|
- This is strongly *non-linear* mapping
|
||||||
|
|
||||||
|
$$
|
||||||
|
A_i \cdot A_i^{-1} \equiv 1 \space (mod \space P(x)) \\
|
||||||
|
B'_i = \begin{cases}
|
||||||
|
0 \quad\quad\quad i=0 \\
|
||||||
|
A_i^{-1} \quad\space\space\space i > 0
|
||||||
|
\end{cases}
|
||||||
|
$$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Note: 0 maps to 0
|
||||||
|
- The inverses $B'_i$ then undergo an **affine transformation** to produce the final s-box
|
||||||
|
- This destroys any remaining mathematical structure
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Remember an affine transformation is a multiplication and addition by two constants (think of the affine cipher)
|
||||||
|
|
||||||
|
##### S-box Properties
|
||||||
|
|
||||||
|
- The s-box simply described, and is bijective, an invertible 1:1 mapping
|
||||||
|
- It has no fixed points
|
||||||
|
- i.e. no $A_i$ for which $S(A_i) = A_i$
|
||||||
|
- No inverse fixed points
|
||||||
|
- i.e. no $A_i$ for which $S(A_i) \oplus A_i = FF$
|
||||||
|
- Minimisation of the largest non-trivial correlation between linear combinations of input bits and linear combinations of output bits
|
||||||
|
- 0 is a non-trivial combination
|
||||||
|
- Minimisation of the largest non-trivial value in the `EXOR` table
|
||||||
|
- This stops differential cryptanalysis
|
||||||
|
|
||||||
|
#### AES Diffusion
|
||||||
|
|
||||||
|
Diffusion in AES consists of two layers:
|
||||||
|
|
||||||
|
1. Shift rows
|
||||||
|
2. Mix columns
|
||||||
|
|
||||||
|
Shift rows simply moves bytes around the block
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Mix Columns
|
||||||
|
|
||||||
|
- Performs a linear mixing of bytes within each column
|
||||||
|
- All the input bytes in a column influence all the output bytes
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Multiplying by `01` does not change the result
|
||||||
|
|
||||||
|
When multiplying by $x$, there’s a shortcut we can implement. We can set the equation equal to 0, and `xor` by $x^4 - x^3 - x - 1$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Key Schedule
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- The first round key used is just the key
|
||||||
|
- We then take $W[3]$ and put it through the $g$ function which just permutes it
|
||||||
|
- $g$ takes the word, shifts it one to the right and then passes it through the s-boxes
|
||||||
|
- We then `xor` it with $RC[i]$ which is just a constant value to ensure *something* changes
|
||||||
|
- Like for example if we had a bit stream of all 0s
|
||||||
|
|
||||||
|
### Implementation
|
||||||
|
|
||||||
|
1. All addition and subtractions are `xor`
|
||||||
|
|
||||||
|
2. Multiply by `01` has no effect
|
||||||
|
|
||||||
|
3. Multiplying by `02` (which is $x$) is simply a left shift followed by modular reduction
|
||||||
|
|
||||||
|
- Left shift multiplies by $x$
|
||||||
|
|
||||||
|
- If the original $x^7$ bit was set, then we must `xor` with `0x1B`
|
||||||
|
|
||||||
|
- ```java
|
||||||
|
// xtime
|
||||||
|
if ((a & 0x80) > 0) {
|
||||||
|
a = (a << 1) ^ 0x1b;
|
||||||
|
} else {
|
||||||
|
a <<= 1;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
4. Multiply by `03` ($x+1$) is simply `xtime(a) ^ a`
|
||||||
|
|
||||||
|
- Inverse multiplications are by `09`, `11`, `13`, `14`. these require either a more general function or lookup tables
|
||||||
|
|
||||||
|
- Consider the sum:
|
||||||
|
|
||||||
|
- $$
|
||||||
|
a = x^6 + x^4 + x^2 + 1 \\
|
||||||
|
b = x^7 + x^4 + x^2 + x \\
|
||||||
|
\therefore a\cdot b = a\cdot x^7 + a\cdot x^4 + a\cdot x^2 + a\cdot x
|
||||||
|
$$
|
||||||
|
|
||||||
|
- $$
|
||||||
|
a\curvearrowright a\cdot x \curvearrowright a\cdot x^2 \curvearrowright a\cdot x^3 \curvearrowright a\cdot x^4 \curvearrowright a\cdot x^5
|
||||||
|
$$
|
||||||
|
|
||||||
|
- Here in $a\cdot b$, $a$ is just being multiplied by various powers of $x$. This can be easily calculated by repeated multiplying $a$ by $x$.
|
||||||
|
|
||||||
|
- AES is very **fast in software** and pretty **fast in hardware**
|
||||||
|
|
||||||
|
- CPU instructions in AES-NI make AES much faster
|
||||||
|
|
||||||
|
- Much of the algorithm can be converted into a series of lookup tables
|
||||||
|
|
||||||
|
- **Trade off** between **speed** and **space**
|
||||||
|
|
||||||
|
- There are numerous cache-timing and other attacks possible
|
||||||
|
|
||||||
|
- Implementation must be constant time
|
||||||
|
- CPU instructions help mitigate this
|
||||||
|
|
||||||
|
- In general AES is much harder to implement safely than `ChaCha20`
|
||||||
@@ -0,0 +1,138 @@
|
|||||||
|
# Padding Oracle Attacks
|
||||||
|
|
||||||
|
#### Block Cipher Modes
|
||||||
|
|
||||||
|
- Most messages don’t come in convenient 128-bit block lengths
|
||||||
|
- We’ll need to run a block cipher repeatedly on consecutive blocks
|
||||||
|
- Why not use stream ciphers?
|
||||||
|
- Historically stream ciphgers have proven harder to implement
|
||||||
|
|
||||||
|
##### Padding
|
||||||
|
|
||||||
|
- ECB and some other modes require message length to be a multiple of the block size
|
||||||
|
- Public Key Cryptography Standards `PKCS7` is a common padding scheme:
|
||||||
|
1. Padding bytes are always added to the plaintext **before it is encrypted**
|
||||||
|
2. Each padding byte has a *value equal to the total number of padding bytes* that are added
|
||||||
|
3. The total number of padding bytes is **atleast one**
|
||||||
|
- 
|
||||||
|
- Note in this example there are 7 `7`s and 16 `16`s
|
||||||
|
- Note the bottom left example there is 1 `1`. This could be interpreted as 1 bytes of padding or some plaintext. This is why every block must contain at least one padding byte
|
||||||
|
|
||||||
|
### Electronic Code Book Mode (ECB)
|
||||||
|
|
||||||
|
- Just encrypt each block one after another
|
||||||
|
- This is quick as can be easily parallelised
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Weaknesses
|
||||||
|
|
||||||
|
- If $x_1$ and $x_3$ are the same, then $y_1$ and $y_3$ are also the same.
|
||||||
|
- ECB allows an attacker to infer information on the plaintext
|
||||||
|
- Consider a hypothetical bank transfer between two banks that use a fixed key, where the message format is roughly known
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- If we start splicing parts of messages together, we can send a legitimate looking message to the bank
|
||||||
|
- We could also use a chosen plaintext attack by requesting a bank transfer, finding our bank account info and splicing that with another message
|
||||||
|
- ECB divulges whenever messages or blocks are the same
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
> Here the RGB pixel data of this image has been encrypted using AES in ECB mode
|
||||||
|
|
||||||
|
### Deterministic vs Probabilistic Encryption
|
||||||
|
|
||||||
|
- An encryption scheme is **deterministic** if some plaintext is mapped to a fixed ciphertext if the key is unchanged
|
||||||
|
- ECB is deterministic, but most modern modes of operation of **probabilistic**
|
||||||
|
- Probabilistic encryption schemes add randomness to the encryption process to achieve a non-deterministic generation of the ciphertext
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- $r$ is not a secret
|
||||||
|
|
||||||
|
#### Cipher Block Chaining (CBC)
|
||||||
|
|
||||||
|
- `XOR` the output of each cipher block with the next input
|
||||||
|
- $IV$ - **Initialisation Vector**
|
||||||
|
- The initial random seed that randomises the whole stream
|
||||||
|
- If you encrypted the same plaintext later it will be different
|
||||||
|
- An attacker will be unable to tell if $y_1$ and $y_2$ are the same message but with different $IV$ or different messages with different $IV$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
$$
|
||||||
|
y_1 = e_k(x_1 \oplus IV) \\
|
||||||
|
y_i = e_k(x_i \oplus y_{i-1})
|
||||||
|
$$
|
||||||
|
|
||||||
|
##### CBC Decryption
|
||||||
|
|
||||||
|
- Similar to encryption, but now `XOR` takes place after decryption
|
||||||
|
- This is much easier to parallelise
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
$$
|
||||||
|
x_1 = d_k(y_1) \oplus IV \\
|
||||||
|
x_i = d_k(y_i) \oplus y_{i-1}
|
||||||
|
$$
|
||||||
|
|
||||||
|
- If we lost $y_1$ we would be unable to decrypt $y_2$
|
||||||
|
- We would be able to decrypt $y_3$ though
|
||||||
|
|
||||||
|
##### Weaknesses
|
||||||
|
|
||||||
|
- CBC was the primary method of encryption for many years
|
||||||
|
- Now it is less common
|
||||||
|
- 
|
||||||
|
- If you flip the first bit in $y_2$, the same bit is flipped for $x_3$
|
||||||
|
- Changing $y_2$ means $x_2$ no longer decrypts properly
|
||||||
|
|
||||||
|
### Padding Oracles
|
||||||
|
|
||||||
|
- Here, an **oracle** is a system we can query and it will tell us if, once decrypt, some text has **valid padding**
|
||||||
|
- A system is unlikely to tell you directly, but it might give away some clue
|
||||||
|
- Image an example `api` that receives a CBC encrypted authorisation token
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Padding Oracle Attacks
|
||||||
|
|
||||||
|
- Lets look at a single decryption block in CBC
|
||||||
|
- The attack is essentially the same for multiple blocks, just one at a time
|
||||||
|
- You attack the last block, which contains the padding
|
||||||
|
|
||||||
|
> The general strategy is to manipulate bits in the IV to find valid padding and recover $z_i$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Counter Mode (CTR)
|
||||||
|
|
||||||
|
- Encrypt a nonce + counter and use this to mask the plaintext with `XOR`
|
||||||
|
- This is very easily parallelised
|
||||||
|
- Each block is encrypted differently, avoiding the issues with ECB mode
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- We are now using our block cipher as a stream cipher
|
||||||
|
- The keystream generation (AES) is run through blocks
|
||||||
|
- Decrypting is super easy, just the reverse
|
||||||
|
|
||||||
|
### Galois Counter Mode
|
||||||
|
|
||||||
|
- Extends counter mode to add authenticity
|
||||||
|
- The sender definitely sent that message and it hasn’t been modified
|
||||||
|
- Very similar to ocunter mode, but **adds authentication tag**
|
||||||
|
- Uses multiplication in a Galois Finite field $GF(2^{128})$ modulo $x^{128} + x^7 + x^2 + x + 1$
|
||||||
|
- Extremely parallelsiable
|
||||||
|
- Robust to message modification
|
||||||
|
- Is now standard in `TLS1.3`
|
||||||
|
|
||||||
|

|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
# Public Key Mathematics
|
||||||
|
|
||||||
|
- Recap: in rings and fields, multiplicative inverse might exist such that
|
||||||
|
|
||||||
|
$$
|
||||||
|
a\cdot a^{-1} \equiv 1 \space (mod \space p)
|
||||||
|
$$
|
||||||
|
|
||||||
|
- Modular inverse exists when $gcd(a,p)=1$
|
||||||
|
- For prime fields, $gcd(a,p)=1, \forall a \neq 0 \in GF(p)$
|
||||||
|
|
||||||
|
#### Euclidean Algorithm
|
||||||
|
|
||||||
|
- The euclidean algorithm calculates the greatest common divisor of two numbers $gcd(r_0, r_1)$
|
||||||
|
- This is the largest number that divides both $r_0$ and $r_1$
|
||||||
|
- If $gcd(x,y)=1$ then $x$ and $y$ are **coprime** (sometimes called relatively prime)
|
||||||
|
- The Euclidean algorithm is based around the fact:
|
||||||
|
- $gcd(r_0, r_1) = gcd(r_1, r_0 - r_1)$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Computing $(x-y)\cdot gcd(r_0, r_1)$ is easier as its a smaller number
|
||||||
|
- Doing this repeatedly is slow, we can use $gcd(r_0,r_1) = gcd(r_1, r_0\space mod \space r_1)$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Example
|
||||||
|
|
||||||
|
$r_0 = 57 \\
|
||||||
|
r_1 = 12$
|
||||||
|
|
||||||
|
- At each step we convert $r_0$ and $r_1$ into the form $r_0=q\cdot r_1 + r_2$
|
||||||
|
|
||||||
|
$r_0=q\cdot r_1 + r_2 \\57=4\cdot 12 + 9\\ r_1=q\cdot r_2 + r_3 \\ 12=1\cdot 9 + 3 \\ 9 = 3\cdot 3 + 0$
|
||||||
|
|
||||||
|
- When the algorithm gets to 0, it is finished, therefore $gcd(57,12)=3$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Bezout’s Identity
|
||||||
|
|
||||||
|
- Bezout’s identity tells us that the greatest common divisor of two numbers can be expressed as the sum of multiples of these numbers
|
||||||
|
- $gcd(r_0,r_1) = s\cdot r_0 + t\cdot r_1$
|
||||||
|
- e.g. $gcd(99,20)=-1\cdot 99+5\cdot 20=1$
|
||||||
|
- $gcd(141,50)=11\cdot 141+-31\cdot 50=1$
|
||||||
|
|
||||||
|
##### Extended Euclidean Algorithm
|
||||||
|
|
||||||
|
- The extended euclidean algorithm calculates the $gcd(r_0,r_1)$ as normal, and in addition calculates $s$ and $t$.
|
||||||
|
|
||||||
|
| Euclidean Algorithm | Extended Euclidean Algorithm |
|
||||||
|
| ---------------------------------- | ------------------------------------------------------------ |
|
||||||
|
| $r_0=q_1\cdot r_1+r_2$ | $r_2=r_0-q_1\cdot r_1 \quad \rightarrow \quad r_2=s_2\cdot r_0-t_2\cdot r_1$ |
|
||||||
|
| $r_1=q_2\cdot r_2+r_3$ | $r_3=r_1-q_2\cdot r_2 \quad \rightarrow \quad r_3=s_3\cdot r_0-t_3\cdot r_1$ |
|
||||||
|
| $r_2=q_3\cdot r_3+r_4$ | $r_4=r_2-q_3\cdot r_3 \quad \rightarrow \quad r_4=s_4\cdot r_0-t_4\cdot r_1$ |
|
||||||
|
| … | … |
|
||||||
|
| $r_{l-2}=q_{l-1}\cdot r_{l-1}+r_l$ | $r_l=r_{l-2}-q_{l-1}\cdot r_{l-1} \quad \rightarrow \quad r_l=s_l\cdot r_0-t_l\cdot r_1$ |
|
||||||
|
| $r_{l-1}=q_{l}\cdot r_{l}+0$ | |
|
||||||
|
|
||||||
|
###### Example
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
###### Formula
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Modular Inverses
|
||||||
|
|
||||||
|
$$
|
||||||
|
\begin{split}
|
||||||
|
& a\cdot a^{-1} \equiv 1 \mod n \\
|
||||||
|
& gcd(n,a) = s\cdot n + t\cdot a = 1 \\
|
||||||
|
& s\cdot n + t\cdot a = 1 \\
|
||||||
|
& s\cdot 0 + t\cdot a \equiv 1 \space mod \space n \\
|
||||||
|
& t\cdot a \equiv 1 \space mod \space n \\
|
||||||
|
& t \equiv a^{-1} \space mod \space n
|
||||||
|
\end{split}
|
||||||
|
$$
|
||||||
|
|
||||||
|
Where $t$ is our multiplicative inverse
|
||||||
|
|
||||||
@@ -0,0 +1,165 @@
|
|||||||
|
# RSA
|
||||||
|
|
||||||
|
- Introduced in 1977 by Ron Rivest, Adi Shamir and Leonard Adleman
|
||||||
|
- The most popular public key algorithm in the world
|
||||||
|
- Solves an important problem that symmetric cryptography doesn’t
|
||||||
|
- RSA keys are normally `2084` or `4096` bits
|
||||||
|
- Security is built around the difficulty of *factoring large numbers*
|
||||||
|
|
||||||
|
### RSA Encryption
|
||||||
|
|
||||||
|
- Encryption performed by the *public key* can only be reversed using the *private key*
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### RSA Signatures
|
||||||
|
|
||||||
|
- The authenticity of signatures generated by the *private key* can be verified by the *public key*
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Euler Totient Function
|
||||||
|
|
||||||
|
- Integers $a$ and $m$ are *relatively prime* if they do not share a divisor (except 1)
|
||||||
|
- $gcd(a,m) = 1$
|
||||||
|
- The **Euler totient** $\Phi$ is the number of integers in $\mathbb{Z}_m = \{0,1,...m-1\}$ for which $gcd(a,m)=1$
|
||||||
|
- For example $\Phi(9)=6$ as:
|
||||||
|
- $gcd(1,9)=1$ :white_check_mark:
|
||||||
|
- $gcd(2,9)=1$ :white_check_mark:
|
||||||
|
- $gcd(3,9)=3$ ❌
|
||||||
|
- $gcd(4,9)=1$ :white_check_mark:
|
||||||
|
- $gcd(5,9)=1$ :white_check_mark:
|
||||||
|
- $gcd(6,9)=3$ ❌
|
||||||
|
- $gcd(7,9)=1$ :white_check_mark:
|
||||||
|
- $gcd(8,9)=1$ :white_check_mark:
|
||||||
|
|
||||||
|
###
|
||||||
|
|
||||||
|
#### Integer Factorisation
|
||||||
|
|
||||||
|
- Any integer can be expressed as the multiplication of a list of prime numbers
|
||||||
|
|
||||||
|
#### Calculating $\Phi(n)$
|
||||||
|
|
||||||
|
- The totient is much easier to calculate given the prime factorisation of $n$
|
||||||
|
|
||||||
|
$$
|
||||||
|
m = p_1^{e_1}\cdot p_2^{e_2} ... \cdot p_3^{e_3} \\
|
||||||
|
\Phi(n) = \prod^n_{i=1} (p_i^{e_i} - p_i^{e_i-1})
|
||||||
|
$$
|
||||||
|
|
||||||
|
##### $\Phi(p)$ for Primes
|
||||||
|
|
||||||
|
$$
|
||||||
|
\Phi(n) = \prod^n_{i=1} (p_i^{e_i} - p_i^{e_i-1}) \\
|
||||||
|
\Phi(n) = (p^1 - p_0) = (p-1)
|
||||||
|
$$
|
||||||
|
|
||||||
|
This is similar for semi-primes $n=p\cdot q$
|
||||||
|
|
||||||
|
$$
|
||||||
|
\Phi(n) = (p^1 - p_0) \cdot (q^1-q_0) = (p-1)(q-1)
|
||||||
|
$$
|
||||||
|
|
||||||
|
#### Fermat’s Little Theorem
|
||||||
|
|
||||||
|
- Fermat’s little theorem states that for some prime $p$, and any integer $a$:
|
||||||
|
- $a^{p-1} \equiv 1 \space (mod \space p)$
|
||||||
|
- Also note that $a^{p-1} = a\cdot a^{p-2} \equiv 1 \space (mod \space p)$
|
||||||
|
- Therefore $a^{p-2}$ is actually the inverse of $a\space (mod \space p)$
|
||||||
|
- It follows that $a^p \equiv p \space (mod \space p)$
|
||||||
|
|
||||||
|
#### Euler’s Theorem
|
||||||
|
|
||||||
|
- Generalisation of Fermat’s little theorem, not exclusive to primes
|
||||||
|
- $a^{\Phi(m)} \equiv 1 \space (mod \space m)$
|
||||||
|
- If $gcd(a,m)=1$
|
||||||
|
- This works for any integer ring $\mathbb{Z}_m$
|
||||||
|
- We can see that FLT is a special case of this
|
||||||
|
- $\Phi(p) = (p-1) \therefore a^{\Phi(p)} = a^{p-1} \equiv 1 \space (mod \space p)$
|
||||||
|
|
||||||
|
## RSA Key Generation
|
||||||
|
|
||||||
|
1. Choose two large primes, $p$ and $q$
|
||||||
|
2. Calculate the modulus $n=p\cdot q$
|
||||||
|
3. Calculate $\Phi(n) = (p-1)\cdot (q-1)$
|
||||||
|
4. Choose a value $e\in \{2, ..., \Phi(n) -1\}$ where $gcd(\Phi(n),e)=1$
|
||||||
|
5. Compute $d$ where $d\cdot e \equiv 1 \space (mod \space \Phi(n))$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
$d$ is very easy to calculate if you know $p$ and $q$
|
||||||
|
|
||||||
|
#### Example
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Encryption
|
||||||
|
|
||||||
|
- Now we have a public key $(3, 187)$ and private key $107$
|
||||||
|
- Encryption and decryption is performed by:
|
||||||
|
- $x^e \equiv y \space (mod \space n)$
|
||||||
|
- $y^d \equiv x \space (mod \space n)$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Proof
|
||||||
|
|
||||||
|
- We want to show that $(x^e)^d = x^{ed} \equiv x \space (mod \space n)$
|
||||||
|
- Let’s assume $gcd(x,n)=1$ So Euler’s theorem applies
|
||||||
|
- $e\cdot d=1\space (mod \space \Phi(n))$
|
||||||
|
- $\therefore e\cdot d = 1 + k\cdot \Phi(n)$
|
||||||
|
- $x^{e\cdot d} = x^{1+k\cdot \Phi(n)} = x\cdot x^{k+\Phi(n)}$
|
||||||
|
- $x\cdot (x^{\Phi(n)})^k=x\cdot(1)^k=x$
|
||||||
|
|
||||||
|
### Why is RSA Secure
|
||||||
|
|
||||||
|
- We’d like the message $x$ based on some ciphertext $y$, given the public key $e$:
|
||||||
|
- $y \equiv ?^d \space (mod \space n)$
|
||||||
|
- $x \equiv y^? \space (mod \space n)$
|
||||||
|
- It can be fairly easy to calculate $d$:
|
||||||
|
- $e\cdot d \equiv q \space (mod \space \Phi(n))$
|
||||||
|
- $\Phi(n) = (p-1)(q-1)$
|
||||||
|
- As an attacker we only have access to $e$ and $d$
|
||||||
|
|
||||||
|
### Exponentiation
|
||||||
|
|
||||||
|
$$
|
||||||
|
x^4 = x^2 \cdot x^2 \\
|
||||||
|
x^8 = x^4 \cdot x^4
|
||||||
|
$$
|
||||||
|
|
||||||
|
When calculating a exponent raised to a power of two, we can use previously calculated values.
|
||||||
|
|
||||||
|
##### Binary Exponentiation
|
||||||
|
|
||||||
|
Where we treat the exponent as a binary number
|
||||||
|
|
||||||
|
- We either square or multiply
|
||||||
|
|
||||||
|
$26=11010_2$
|
||||||
|
|
||||||
|
- Remember squaring is 1 bit shift to the left
|
||||||
|
- Multiplying is just adding $1$
|
||||||
|
|
||||||
|
$$
|
||||||
|
x^{101} \quad = \quad x^{1100101_2} \\
|
||||||
|
x\cdot x = x^2 \quad x^{10_2} \\
|
||||||
|
x^2 \cdot x = x^3 \quad x^{110_2}\\
|
||||||
|
x^3 \cdot x^3 = x^6 \quad x^{1100_2}\\
|
||||||
|
x^6 \cdot x^6 = x^{12} \quad x^{11000_2}\\
|
||||||
|
x^{12} \cdot x^{12} = x^{24} \quad x^{110000_2}\\
|
||||||
|
x^{24} \cdot x = x^{25} \quad x^{110001_2}\\
|
||||||
|
x^{25} \cdot x^{25} = x^{50} \quad x^{1100010_2}\\
|
||||||
|
x^{50} \cdot x^{50} = x^{100} \quad x^{11000100_2}\\
|
||||||
|
x^{100} \cdot x = x^{101} \quad x^{110001001_2}\\
|
||||||
|
$$
|
||||||
|
|
||||||
|
##### Computational Complexity
|
||||||
|
|
||||||
|
- What is the computational complexity of exponentiation?
|
||||||
|
- For a 2048 key:
|
||||||
|
- $X^{2^{2048}}$ - A ridiculously big number
|
||||||
|
- Where as using square and multiply
|
||||||
|
- $2048=T$ we need $\frac{3T}{2}$ calculations
|
||||||
|
|
||||||
@@ -0,0 +1,135 @@
|
|||||||
|
# Diffie-Hellman
|
||||||
|
|
||||||
|
- Two parties can jointly agree a *shared secret* over an *insecure channel*
|
||||||
|
- Mathematically, what we are doing is both calculating the same value, mod a prime $p$
|
||||||
|
- Remember $p$ is $\times 10^{600}$
|
||||||
|
- The parties separately compute the same key, rather than share it
|
||||||
|
|
||||||
|
### $\mathbb{Z}_n^*$
|
||||||
|
|
||||||
|
> The set $\mathbb{Z}_n^*$ consists of the integers $\{1,2,...,n-1\}$ for which $gcd(i,n)=1$
|
||||||
|
>
|
||||||
|
> This set forms an *abelian* group under multiplication modulo $n$. The identity element is 1
|
||||||
|
|
||||||
|
- In the majority of cases, we use a prime number as the modulus:
|
||||||
|
- $\mathbb{Z}_p^* = \{1,2,...,p-1\}$
|
||||||
|
|
||||||
|
**Group Cardinality** - The number of elements in that group
|
||||||
|
|
||||||
|
$$
|
||||||
|
|\mathbb{Z}_m^*| = p-1 \\
|
||||||
|
|\mathbb{Z}_m^*| = \Phi(n) \\
|
||||||
|
$$
|
||||||
|
|
||||||
|
- The security of ciphers often depend on the cardinality of the group
|
||||||
|
|
||||||
|
#### Cyclic Groups
|
||||||
|
|
||||||
|
- Lets consider group $\mathbb{Z}_{11}^*$
|
||||||
|
- Consider calculating powers of 3 in this group
|
||||||
|
|
||||||
|
$$
|
||||||
|
3^i \space (mod \space 11) \\
|
||||||
|
a^1=3\\
|
||||||
|
a^2=3\cdot 3 = 9 \\
|
||||||
|
a^3 = 27 \equiv 5 \\
|
||||||
|
a^4=a\cdot a^3=3\cdot 5 \equiv 4 \\
|
||||||
|
a^5=a\cdot a^4=3\cdot 4 \equiv 1
|
||||||
|
$$
|
||||||
|
|
||||||
|
- This pattern of $\{3,9,5,4,1\}$ repeats indefinitely
|
||||||
|
|
||||||
|
##### Order of an Element
|
||||||
|
|
||||||
|
> The order $ord(a)$ of an element $a$ of a group $(G, \circ)$ is the smallest positive integer $k$ such that:
|
||||||
|
>
|
||||||
|
> $a^k = \underbrace {a\circ a\circ ...\circ a}_{k\space times} =1$
|
||||||
|
>
|
||||||
|
> Where 1 is the neutral element of $G$
|
||||||
|
|
||||||
|
##### Another Cyclic Group
|
||||||
|
|
||||||
|
- What about $2^i$ in $\mathbb{Z}_{11}^*$
|
||||||
|
|
||||||
|
$$
|
||||||
|
2^i \space mod \space 11 \\
|
||||||
|
a^1 = 2 \\
|
||||||
|
a^2=4 \\
|
||||||
|
a^3=8 \\
|
||||||
|
a^4=5 \\
|
||||||
|
a^5=10 \\
|
||||||
|
a^6=9 \\
|
||||||
|
a^7=7 \\
|
||||||
|
a^8=3 \\
|
||||||
|
a^9=6 \\
|
||||||
|
a^{10}=1 \\
|
||||||
|
a^{11}=2 \\
|
||||||
|
a^{12}=4 \\
|
||||||
|
$$
|
||||||
|
|
||||||
|
- We have generated every value in this group before cycling back round
|
||||||
|
|
||||||
|
- A group that contains an element $g$ of maximum order is called a cyclic group
|
||||||
|
- Any element of maximum order is called a primitive root, or a generator
|
||||||
|
- $2$ is a generator of $\mathbb{Z}_{11}^* \quad ord(2)=10$
|
||||||
|
- 3 is not a generator $\mathbb{Z}_{11}^* \quad ord(3)=5$
|
||||||
|
|
||||||
|
##### Cyclic Subgroups
|
||||||
|
|
||||||
|
- For all primes, $(\mathbb{Z}_{11}^*, \cdot)$ is an *abelian finite cyclic group*
|
||||||
|
- Let $g \in G$ where $G$ is a cyclic group:
|
||||||
|
1. $g^{|G|}=1$
|
||||||
|
2. $ord(g)$ divides $|G|$
|
||||||
|
- These are called **cyclic subgroups**
|
||||||
|
- Orders of $\mathbb{Z}_{11}^*$
|
||||||
|
- 
|
||||||
|
- Note the neutral element generates an order of $1$
|
||||||
|
|
||||||
|
## Diffie-Hellman
|
||||||
|
|
||||||
|
1. Alice and Bob agree on a large prime $p$, and a generator $g$ that is a primitive root of $p$
|
||||||
|
2. Alice and Bob choose private numbers $a$ and $b$ at random in $\mathbb{Z}_p^*$
|
||||||
|
- Where $a\in \{1,2,...,p-1\}$
|
||||||
|
- and $b\in \{1,2,...,p-1\}$
|
||||||
|
3. Alice calculates $A=g^a\space mod \space p$ and sends $A$ publicly to Bob
|
||||||
|
4. Bob calculates $B=g^b\space mod \space p$ and sends $B$ pubicly to Alice
|
||||||
|
5. Alice computes $k_{ab}=B^a\space mod \space p$
|
||||||
|
6. Bob computes $k_{ab}=A^b\space mod \space p$
|
||||||
|
|
||||||
|
$$
|
||||||
|
B^a\space mod \space p = (g^b)^a = g^{ab}\space mod \space p \\
|
||||||
|
A^b\space mod \space p = (g^a)^b = g^{ab}\space mod \space p
|
||||||
|
$$
|
||||||
|
|
||||||
|
#### The Discrete Logarithm Problem
|
||||||
|
|
||||||
|
- Why is Diffie-Hellman so hard to break
|
||||||
|
- Consider $\mathbb{Z}^*_{10000079},\space g=3$
|
||||||
|
- Alice calculates $A=3^a\space mod \space 10000079 = 4675535$
|
||||||
|
- What is $a$?
|
||||||
|
- This is the discrete logarithm problem
|
||||||
|
|
||||||
|
**Brute Force** requires $O(|G|)$
|
||||||
|
|
||||||
|
**Shank’s Baby-Step Giant-Step** requires $O(\sqrt{|G|})$ and $\sim \sqrt{|G|}$ space
|
||||||
|
|
||||||
|
- Using 128 bits, this is $2^{64}$, which would need a cluster
|
||||||
|
|
||||||
|
**Pollard’s Rho** requires $O(\sqrt{|G|})$
|
||||||
|
|
||||||
|
**Pohlig-Hellman** is based on the prime factorisation of $|G|$
|
||||||
|
|
||||||
|
- The discrete log problem is solved mod each prime factor and the results combined using the Chinese remainder theorem
|
||||||
|
|
||||||
|
**Index calculus** directly attacks $\mathbb{Z}_p^*$ and is the reason Elliptic Curves is so much more efficient
|
||||||
|
|
||||||
|
##### Choosing Primes
|
||||||
|
|
||||||
|
- To avoid any unexpected small subgroup attacks, commonly used DH primes are **safe primes**
|
||||||
|
- A safe prime is a prime $p$ where $\frac{(p-1)}{2}$ is also a prime
|
||||||
|
- Consider the order of $\mathbb{Z}_p^*$ for a safe prime
|
||||||
|
- This will have two subgroups of order $p-1$ and $2$
|
||||||
|
- By choosing a generator of the **subgroup of large prime order**, we avoid attacks on small factors of the group order
|
||||||
|
- Basically this ensures the prime factorisation has one massive prime in it
|
||||||
|
|
||||||
|
|
||||||
@@ -0,0 +1,293 @@
|
|||||||
|
# Elliptic Curves
|
||||||
|
|
||||||
|
- We’d like to find another type of group and operation in which the discrete logarithm problem is hard
|
||||||
|
|
||||||
|
$$
|
||||||
|
ax^2+by^2=r^2
|
||||||
|
$$
|
||||||
|
|
||||||
|
- There are an infinite amount of solutions to this equation
|
||||||
|
- However if we restrict to only integers ($\mathbb{Z}$) and use mod, we have a finite set
|
||||||
|
|
||||||
|
- We define an elliptic curve over points in $\mathbb{Z}_p, \space p>3$
|
||||||
|
- Set of all pairs where:
|
||||||
|
- $y^2 \equiv x^3 + ax + b \space (mod \space p)$
|
||||||
|
- The neutral element is 0
|
||||||
|
- One requirement is:
|
||||||
|
- $4a^3 + 27b^2 \neq 0 \space (mod \space p)$
|
||||||
|
|
||||||
|
This is $y^2 \equiv x^3 -3x +3$ over $\mathbb{R}$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Notice the symmetry about the x axis, this is because we have a $y^2$ term meaning we have two solutions
|
||||||
|
|
||||||
|
- For a DLP problem, we need a cyclic group
|
||||||
|
- Elements within the group
|
||||||
|
- A group operation
|
||||||
|
- For ECs the elements are points on the curve
|
||||||
|
- The operation is point addition
|
||||||
|
|
||||||
|
#### Point Addition
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Point Doubling
|
||||||
|
|
||||||
|
$P + P = 2P$
|
||||||
|
|
||||||
|
- Here our line will be tangent to P
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Group Laws
|
||||||
|
|
||||||
|
In elliptic curves, to get $4P$, we can either do $P+3P$ or $2P+2P$
|
||||||
|
|
||||||
|
##### Group Properties
|
||||||
|
|
||||||
|
- Closed
|
||||||
|
- Any closed addition operation will end up somewhere on the curve
|
||||||
|
- Associative
|
||||||
|
- The order of calculations doesn’t matter
|
||||||
|
|
||||||
|
##### Point Addition Equations
|
||||||
|
|
||||||
|
- We can derive equations for this based on the equation for a line that intersects the curve in three places
|
||||||
|
- Given $y^3=x^3+ax+b$ and points:
|
||||||
|
- $P=(x_1,y_1)$
|
||||||
|
- $Q=(x_2, y_2)$
|
||||||
|
- line $y=s\cdot x + m$
|
||||||
|
- $(sx+m)^2 = x^3 + ax + b$
|
||||||
|
- $s^2x^2 + 2sxm + m^2 = x^3+ax+b$
|
||||||
|
- Plugging in $x_1, y_1, x_2, y_2$
|
||||||
|
- $P+Q=(x_3, y_3)$
|
||||||
|
- $x_3 = s^2 - x_1 - x_2$
|
||||||
|
- $y_3 = s(x_1 - x_3) - y_1$
|
||||||
|
|
||||||
|
$$
|
||||||
|
s = \cases{\frac{y_2-y_1}{x_2-x_1} \quad (mod\space p); P\neq Q\\{\frac{3x_1^2+a}{2y_1}}\quad (mod\space p); P=Q}
|
||||||
|
$$
|
||||||
|
|
||||||
|
###### Example
|
||||||
|
|
||||||
|
$y^2\equiv x^3+2x+2\space (mod \space 17)$
|
||||||
|
|
||||||
|
$(3,1)+(9,16)$
|
||||||
|
|
||||||
|
$$
|
||||||
|
s=\frac{16-1}{9-3}=\frac{15}{6}=15\cdot 6^{-1} \\
|
||||||
|
= 15\cdot 3 \mod{17} \\
|
||||||
|
= 11
|
||||||
|
$$
|
||||||
|
|
||||||
|
$$
|
||||||
|
x_3=11^2-3-9 \\
|
||||||
|
109 \space \mod{17} = 7 \\\\
|
||||||
|
y_3 = 11\cdot(3-7)-1=11\cdot 13 \mod{17} = 6 \\
|
||||||
|
$$
|
||||||
|
|
||||||
|
#### Inverses
|
||||||
|
|
||||||
|
The point reflected in the x axis is the inverse
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Neutral Element
|
||||||
|
|
||||||
|
$P-P=?$
|
||||||
|
|
||||||
|
$P+?=P$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
These are a pain as they don’t intersect the curve, we say they cross the curve at $\infty$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### The Point $\mathcal O$ at Infinity
|
||||||
|
|
||||||
|
- The point at infinity is the neutral element on a elliptic curve
|
||||||
|
- $P+(-P)=\mathcal O$
|
||||||
|
- $P+\mathcal O=P$
|
||||||
|
- In practice the point doesn’t have coordinates, and can’t be used within the normal formula
|
||||||
|
- $P=(x,y)$
|
||||||
|
- $-P=(x,-y)$
|
||||||
|
- When implementing, you have to detect when the x values are equal and y values are inverses $\mod p$
|
||||||
|
- e.g. $(7,6)+(7,11)$
|
||||||
|
- $\frac{y_2-y_1}{x_2-x_1}=\frac{-5}{0} = \mathcal O$
|
||||||
|
|
||||||
|
### Cyclic Groups
|
||||||
|
|
||||||
|
- The points on an elliptic curve including the neutral element $\mathcal O$ form a cyclic subgroup
|
||||||
|
- Under certain conditions all points for a cyclic group
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Given a curve $E$, a primitive root $P$, and a point $aP$, what is $a$?
|
||||||
|
- This is the elliptic curve discrete logarithm problem
|
||||||
|
|
||||||
|
$$
|
||||||
|
aP = \underbrace{P+P+...+P}_{a \space \mathrm {times}}
|
||||||
|
$$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
This is the graph modulus $p$
|
||||||
|
|
||||||
|
- Given a generator point, points on elliptic curves generate cyclic groups
|
||||||
|
- $y^2 \equiv x^3+2x+2 \mod 17$
|
||||||
|
- 
|
||||||
|
- Here the next two points is the point at infinity ($\mathcal O$) and then it loops back round to $(5,1)$
|
||||||
|
- Each cyclic group includes the point at infinity
|
||||||
|
|
||||||
|
## Elliptic Curve Discrete Logarithm
|
||||||
|
|
||||||
|
- We can construct a DLP in a very similar way to the modular exponentiation equivalent
|
||||||
|
- $aP = \underbrace{P+P+...+P}_{a \space\textrm{ times}} = A$
|
||||||
|
- Given points $P$ and $A$, find scalar value $a$
|
||||||
|
- Its important to remember the distinction between points on the curve, and integer values
|
||||||
|
- On elliptic curves, private keys such as $a$ are integers
|
||||||
|
- Generators and public keys are points
|
||||||
|
|
||||||
|
#### Group Cardinality
|
||||||
|
|
||||||
|
- The size of cyclic groups is very important to the security
|
||||||
|
- While easy to calculate for modular arithmetic, the number of points on a give elliptic curve is not so obvious
|
||||||
|
- You might imagine that a curve would have $2p+1$ points, in reality it is fewer than this
|
||||||
|
- This is closer to $p$
|
||||||
|
- Hasse’s theorem states that for a curve $E$ over a field $\mathbb{Z}_p$, the number of elements $\#E$ is bounded by:
|
||||||
|
- $\#E=p+1+\epsilon$
|
||||||
|
- where $|\epsilon| \leq 2\sqrt{p}$
|
||||||
|
|
||||||
|
##### #E
|
||||||
|
|
||||||
|
- A large #E is very important to prevent various attacks on ECDLP
|
||||||
|
- Calculating it exactly is hard, it can be done with Shoof’s algorithm
|
||||||
|
- Various properties of #E enable or restrict certain attacks
|
||||||
|
|
||||||
|
##### How Hard is ECDLP
|
||||||
|
|
||||||
|
- There are generic algorithms like **Polig-Hellman** that are applicable to any category of DLP
|
||||||
|
- Polig-Hellman requires $O(\sqrt{\#E})$ steps
|
||||||
|
- These are generic attacks mean curves and parameters should be chosen with care
|
||||||
|
- The most powerful attack on modular arithmetic based DLP is **index calculus**
|
||||||
|
- It is this attack that forces modular arithmetic based crypto-systems to use >2000 bit keys
|
||||||
|
- Index calculus does not work on elliptic curves so they only need to remain secure against generic attacks
|
||||||
|
|
||||||
|
#### Efficient Computation
|
||||||
|
|
||||||
|
- There is no nautral way of calculating $a\cdot P$
|
||||||
|
- Think back to binary exponentiation, square and multiply `->` double and add
|
||||||
|
|
||||||
|
| Decimal | Binary |
|
||||||
|
| ---------------- | ---------------- |
|
||||||
|
| $26_{10}\cdot P$ | $11010_2\cdot P$ |
|
||||||
|
| $1P$ | $1 \cdot P$ |
|
||||||
|
| $2P=1P+1P$ | $10\cdot P$ |
|
||||||
|
| $3P=2P+1P$ | $11\cdot P$ |
|
||||||
|
| $6P=3P+3P$ | $110\cdot P$ |
|
||||||
|
| $12P=6P+6P$ | $1100\cdot P$ |
|
||||||
|
| $12P+1P = 13P$ | $1101\cdot P$ |
|
||||||
|
| $26P = 13P+13P$ | $11010\cdot P$ |
|
||||||
|
|
||||||
|
### Elliptic Curve Diffie-Hellman (ECDH)
|
||||||
|
|
||||||
|
$$
|
||||||
|
E, \#E, G \\
|
||||||
|
\mathrm{Alice}: a\in \{1,2,...,\#E-1\} \\
|
||||||
|
\mathrm{Bob}: a\in \{1,2,...,\#E-1\} \\
|
||||||
|
$$
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
Alice takes point $G$ on the curve and add it to $a$: $A = a\cdot G$
|
||||||
|
|
||||||
|
Bob does the same: $B=b\cdot G$
|
||||||
|
|
||||||
|
Alice takes bob’s public key $k_{ab} = a\cdot B$
|
||||||
|
|
||||||
|
Bob does the same: $k_{ab}=b\cdot A$
|
||||||
|
|
||||||
|
$k_{ab} = a\cdot B = a \cdot (b \cdot G)=ab\cdot G$
|
||||||
|
|
||||||
|
$k_{ab} = b\cdot A = b \cdot (a \cdot G)=ab\cdot G$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### EC Structure
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Where each layer builds on the one beneath
|
||||||
|
|
||||||
|
## Implementation
|
||||||
|
|
||||||
|
#### Point Compression
|
||||||
|
|
||||||
|
- Since we know the formula for a given curve, we do not need to transport full $(x,y)$ coordinates
|
||||||
|
- Each point contains a unique $x$, and one or two $y$ where
|
||||||
|
- $y=\sqrt{x^3 + 2x + 2}\mod p$
|
||||||
|
- Most implementations will use the full $x$ value, and append a single bit representing a positive or negative y value
|
||||||
|
|
||||||
|
#### Projective Coordinates
|
||||||
|
|
||||||
|
- Some implementations adjust the formula for point addition to use projective coordinates $(x,y,z)$ rather than $(x,y)$
|
||||||
|
- The curve sits on the plane $z=1$
|
||||||
|
- Points at infinity $\mathcal O = (0,1,0)$
|
||||||
|
|
||||||
|
**Why?**
|
||||||
|
|
||||||
|
- Point addition in this system does not require a multiplicative inverse
|
||||||
|
|
||||||
|
#### Standard Curves
|
||||||
|
|
||||||
|
- The choice of curve parameters influences both security and efficiency of crypto-systems based around ECs
|
||||||
|
- Never use a randomly generated curve!
|
||||||
|
- The chances are the number of points we generate will have a subgroup susecpible to Polig-Hellmen
|
||||||
|
- Standard curves exist in various forms
|
||||||
|
- Varied equations
|
||||||
|
- Different implementation methods
|
||||||
|
- Different choices of prime
|
||||||
|
|
||||||
|
##### P-256
|
||||||
|
|
||||||
|
- Weierstrass curve $y^2\equiv x^3+ax+b\mod p$
|
||||||
|
- Very widely used
|
||||||
|
- One of the few curves in `TLS1.3` and NSA Suite B
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- $h$ is the cofactor, the size of the subgroup in $G$
|
||||||
|
- Because its 1 it means all the points are being generated
|
||||||
|
- If it was 2, only half of the points are being generated
|
||||||
|
|
||||||
|
##### secp256k1
|
||||||
|
|
||||||
|
- Koblitz curve $y^2\equiv x^3+7\mod p$
|
||||||
|
- Underpins Bitcoin digital signatures
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Curve25519
|
||||||
|
|
||||||
|
- Montgomery curve $y^2\equiv x^3 + 486662x^2+x\mod p$
|
||||||
|
- Primary alternative to `P-256`
|
||||||
|
- In `TLS1.3` and numerous other protocols
|
||||||
|
- The nature of this curve allows efficient multiplication using a Montgomery ladder, using only $X$ and $Z$
|
||||||
|
- This algorithm can compute numbers in constant time
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Curve448-Goldilocks
|
||||||
|
|
||||||
|
- Untwisted Edwards Curve $y^2+x^2\equiv 1 - 39081x^2y^2\mod p$
|
||||||
|
- 448 bit curve
|
||||||
|
- Primarily used within digital signatures as part of `Ed448`
|
||||||
|
- Edwards curve arithmetic mod this “goldilocks” prime is very efficient
|
||||||
|
|
||||||
|
#### Primary Applications
|
||||||
|
|
||||||
|
- Elliptic Curve Diffie Hellman
|
||||||
|
- DSA Signatures scheme, based on Elgamal signatures
|
||||||
|
- Similar schemes involving the alternative curves such as `Ed25519` and `Ed448`
|
||||||
@@ -0,0 +1,135 @@
|
|||||||
|
# Elgamal Encryption
|
||||||
|
|
||||||
|
#### Extending Diffie-Hellmen to Encryption
|
||||||
|
|
||||||
|
We could do is multiply the plain text by the key generated
|
||||||
|
|
||||||
|
$y\equiv x\cdot k_{ab}\mod p \rightarrow x\equiv y\cdot k_{ab}^{-1}$
|
||||||
|
|
||||||
|
### Elgamal
|
||||||
|
|
||||||
|
- Because this is public key encryption, we can make some efficiency savings by not sending all the information both ways every time
|
||||||
|
- If encryption is from Alice to Bob, Bob only needs to publish a public key once
|
||||||
|
- The scheme provides some other security benefits - **ephemeral keys**
|
||||||
|
|
||||||
|
#### Elgamal Key Generation
|
||||||
|
|
||||||
|
**Bob**:
|
||||||
|
|
||||||
|
1. Choose large prime $p$
|
||||||
|
2. Choose primitive element $g\in\mathbb{Z}^*_p$ or in a subgroup of $\mathbb{Z}^*_p$
|
||||||
|
3. Choose $k_{pr}=b\in\{1,2,...,p-1\}$
|
||||||
|
4. Compute $B \equiv g^b\mod p$
|
||||||
|
5. Publish public key $k_{pub}=(p,g,B)$
|
||||||
|
|
||||||
|
#### Elgamal Key Generation
|
||||||
|
|
||||||
|
**Alice**:
|
||||||
|
|
||||||
|
1. Choose $a\in \{1,2,...,p-1\}$
|
||||||
|
2. Compute ephemeral key
|
||||||
|
- $k_E\equiv g^a\mod p$
|
||||||
|
- Remember ephemeral means the key is generated every time communication happens
|
||||||
|
3. Compute masking key
|
||||||
|
- $k_M\equiv B^a\mod p$
|
||||||
|
4. Encrypt message $x\in\mathbb{Z}^*_p$
|
||||||
|
- $y\equiv x\cdot k_M\mod p$
|
||||||
|
5. Send $(k_E,y)$
|
||||||
|
|
||||||
|
#### Elgamal Decryption
|
||||||
|
|
||||||
|
1. Compute masking key
|
||||||
|
- $k_M\equiv k_E^b\mod p$
|
||||||
|
2. Decrypt message
|
||||||
|
- $x\equiv y\cdot k_M^{-1}\mod p$
|
||||||
|
|
||||||
|
### Computational Efficiency
|
||||||
|
|
||||||
|
To calculate bobs private key we use one exponentiation
|
||||||
|
|
||||||
|
Alice has to do two binary exponentiation to send a message to bob
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- Both the exponentiations during encryption can be pre-computed during down time
|
||||||
|
- We can also improve on the decryption step using Fermat’s little theorem
|
||||||
|
- Fermat’s Little Theorem: $a^{p-1}\equiv 1\mod p$
|
||||||
|
1. Compute $k_M=k_E^b\mod 67$
|
||||||
|
2. Compute $k_M^{-1}$
|
||||||
|
3. Decrypt $y=y\cdot k_M^{-1}\mod p$
|
||||||
|
|
||||||
|
#### Practicalities
|
||||||
|
|
||||||
|
- Elgamal is a probabilistic encryption scheme. It uses an ephemeral key pair $a$ and $k_E=g^a\mod p$
|
||||||
|
- Elgamal has a major weakness if you reuse an ephemeral key, and is also less efficient than simply using Diffie-Hellman than AES
|
||||||
|
- The other form of Elgamal is a scheme for digital signatures, variants of which are much more popular
|
||||||
|
|
||||||
|
### Elgamal Digital Signature
|
||||||
|
|
||||||
|
**Bob**:
|
||||||
|
|
||||||
|
1. Choose $g,p$
|
||||||
|
2. Choose $k_{pr}=b\in\{1,2,...,p-1\}$
|
||||||
|
3. Compute $k_{pub}=B\equiv g^b\mod p$
|
||||||
|
4. Publish public key $k_{pub}=(p,g,B)$
|
||||||
|
|
||||||
|
Then decide the ephemeral key $k\in\{1,2,...,p-2\}$ where $\gcd(k,p-1)=1$
|
||||||
|
|
||||||
|
- $r\equiv g^k\mod p$
|
||||||
|
- $s\equiv(m-b\cdot r)\cdot k^{-1}\mod p-1$
|
||||||
|
|
||||||
|
Bob sends the message, and $r$ and $s$
|
||||||
|
|
||||||
|
Alice to verify:
|
||||||
|
|
||||||
|
- $ver_{k_{pub}}(m,(r,s) =\\ g^m = B^rr^s\mod p$
|
||||||
|
|
||||||
|
#### Proof
|
||||||
|
|
||||||
|
Signature: $s=(m-b\cdot r)\cdot k^{-1}\mod p-1$
|
||||||
|
|
||||||
|
- $\therefore s\cdot k=x-b\cdot r \mod p-1$
|
||||||
|
- $\therefore x=b\cdot r+k\cdot s \mod p-1$
|
||||||
|
|
||||||
|
Then
|
||||||
|
|
||||||
|
- $g^x\equiv B^rr^s\equiv(g^b)^r(g^k)^s \mod p$
|
||||||
|
- $g^x\equiv g^{br}\cdot g^{ks}$
|
||||||
|
- $\therefore g^x\equiv g^{b\cdot r + k\cdot s}\mod p$
|
||||||
|
|
||||||
|
Recall: $a^{p-1}\equiv 1\mod p$ for some $m$
|
||||||
|
|
||||||
|
- $a^m\equiv a^{q\cdot(p-1)+r}\mod p$
|
||||||
|
- $\therefore a^m\equiv (a^q)^{(p-1)}\cdot a^r\mod p$
|
||||||
|
- $\therefore a^m\equiv 1\cdot a^r\mod p$
|
||||||
|
- So $a^m\equiv a^{m \mod p-1}\mod p$
|
||||||
|
|
||||||
|
> If exponents are equal $\mod p-1$, then terms are equal $\mod p$
|
||||||
|
|
||||||
|
#### Practicalities
|
||||||
|
|
||||||
|
- As with RSA it’s customary to hash the message and use $H(m)$ not $m$
|
||||||
|
- The combined message and signature $m, (r,s)$ is roughly 3 times the size of the prime $p$, which makes Elgamal signatures quite inefficient
|
||||||
|
- Note that the signature $s\equiv (m-b\cdot r)\cdot k^{-1}\mod p-1$ is calculated in a prime order subgroup of $\mathbb{Z}_p^*$
|
||||||
|
- Without hashing Elgamal is vulnerable to existential forgeries, and key recovery is possible if you reuse the ephemeral key $k$
|
||||||
|
|
||||||
|
### DSA
|
||||||
|
|
||||||
|
- Based on Elgamal, DSA was developed by NIST as an alternative to RSA
|
||||||
|
- Computed in a subgroup of prime order q, which is usually 160 bits
|
||||||
|
- This means the signature (r, s) is 320 bits
|
||||||
|
- Hashing is enforced by the algorithm, and a hash function must match the key size
|
||||||
|
- e.g. SHA-1 for 160-bit q, SHA-256 for 256 bit q
|
||||||
|
- Index calculus does not apply to the sub-group, so 160 bit DSA has a security of 80 bits
|
||||||
|
- In practice larger keys would be required now
|
||||||
|
|
||||||
|
#### ECDSA
|
||||||
|
|
||||||
|
- Identical to DSA, ECDSA operates on an elliptic curve over $\mathbb{Z}_p$ with the signature calculated over a subgroup of prime order $\#q$
|
||||||
|
- More efficient, does not require modulus of thousands of bits
|
||||||
|
- Security level is based on generic attacks against EC
|
||||||
|
- i.e $\sqrt{|\#q|}$
|
||||||
|
- Deterministic generation of $k$ is often used for safety (RFC 6979)
|
||||||
|
- This is where the ephemeral key isn’t random, it’s based off the hash of the message
|
||||||
|
- This is because reusing the ephemeral key is bad news
|
||||||
|
- Other variants like EdDSA using Edwards curves (Ed25519 / Ed448) exist
|
||||||
@@ -0,0 +1,183 @@
|
|||||||
|
# Digital Signatures
|
||||||
|
|
||||||
|
- A signature is proof of authenticity of the sender
|
||||||
|
- Verification is performed by checking the signature against a known signature
|
||||||
|
- Mostly works for the real world, not very robust
|
||||||
|
- This does not scale
|
||||||
|
|
||||||
|
#### Electronic Signature
|
||||||
|
|
||||||
|
- Create a binary signature and append this to any document
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
This is incredibly easy to forge, we need a cryptographic solution
|
||||||
|
|
||||||
|
- In many cases two parties will share a symmetric key $k$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Verification
|
||||||
|
|
||||||
|
To verify a message one must have the message and the signature
|
||||||
|
|
||||||
|
$$
|
||||||
|
(x,y)\rightarrow ver_k(x,y)=\begin{cases}\textrm{True; y is valid}\\\textrm{False; y is invalid}\end{cases}
|
||||||
|
$$
|
||||||
|
|
||||||
|
> **Non-repudiation**
|
||||||
|
>
|
||||||
|
> Symmetric keys for verification don’t work, because both parties have access to key $k$, either party can sign it.
|
||||||
|
>
|
||||||
|
> Bob needs to be able to prove that Alice and no one else signed the signature
|
||||||
|
>
|
||||||
|
> This requires using a private key
|
||||||
|
|
||||||
|
Symetric Signatures gives us:
|
||||||
|
|
||||||
|
**Authenticity**: The sender is confirmed as authentic - only Alice or Bob could have generated the signature
|
||||||
|
|
||||||
|
**Integrity**: The signature confirms the message hasn’t been altered - this is better than the real-world signature scheme
|
||||||
|
|
||||||
|
**Non-Repudiation**: We don’t have this - the symmetric key means that either Alice or Bob could have sent the message
|
||||||
|
|
||||||
|
### Pubic Key Signatures
|
||||||
|
|
||||||
|
- By using asymmetric cryptography we have non-repudiation.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### RSA Signatures
|
||||||
|
|
||||||
|
Notation:
|
||||||
|
|
||||||
|
- $m$ - message
|
||||||
|
- $s$ - signature
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Efficiency
|
||||||
|
|
||||||
|
Signing: $x^d\mod n$
|
||||||
|
|
||||||
|
Verification: $s^e\mod n$
|
||||||
|
|
||||||
|
- Signing and verification require one use of the *square and multiply* algorithm
|
||||||
|
- Efficiency depends on the exponents
|
||||||
|
- We often keep $e$ small
|
||||||
|
- $65537=2^{16}+1=10000000000001_2$
|
||||||
|
- This prioritises verification speed
|
||||||
|
|
||||||
|
##### Signature Forgeries
|
||||||
|
|
||||||
|
- A forgery is the ability to create a valid message / signature pair $(m,s)$ where $m$ hasn’t previously been signed by the legitimate signer
|
||||||
|
- For example replay attack using a previous $(m,s)$ wouldn’t count as a forgery
|
||||||
|
- As we cannot control the message contents
|
||||||
|
- Various severities of attack exist depending on the control over the message $m$
|
||||||
|
|
||||||
|
###### Existential Forgeries
|
||||||
|
|
||||||
|
- The attacker is able to create a valid message / signature pair $(m,s)$
|
||||||
|
- There are no constraints on $m$, it may well be entirely random
|
||||||
|
- $m$ does not need to be a valid message to be understood by a recipient
|
||||||
|
|
||||||
|
An attacker has access to Alice’s public key $(n,e)$
|
||||||
|
|
||||||
|
- They can calculate
|
||||||
|
- $s=\textrm{random}$
|
||||||
|
- $m' =s^e\mod n$
|
||||||
|
- It is trival to generate message and signature pairs based on an RSA public key
|
||||||
|
- Not very useful
|
||||||
|
|
||||||
|
###### Selective Forgeries
|
||||||
|
|
||||||
|
- The attacker is able to create a valid message / signature pair $(m,s)$ where they have selected $m$ in advanced
|
||||||
|
- $m$ may have some mathematical proprieties, or be all zeros etc
|
||||||
|
- It is a requirement that $m$ be fixed prior to the attack
|
||||||
|
|
||||||
|
###### Universal Forgeries
|
||||||
|
|
||||||
|
- The attacker can create a valid signature from any message $m$
|
||||||
|
- This is the strongest attack, and implies the previous attacks too
|
||||||
|
- In RSA, this would imply the attack has access to the private key
|
||||||
|
|
||||||
|
### Malleability
|
||||||
|
|
||||||
|
- RSA is also malleable: $RSA(m_1\cdot m_2)=RSA(m_1)\cdot RSA(m_2)$
|
||||||
|
- Given two messages $x_1, x_2$ and corresponding signatures $s_1,s_2$
|
||||||
|
- $(m_3,s_3)\equiv(m_1\cdot m_2, s_1\cdot s_2)(\mod m)$
|
||||||
|
- This is more control for an attacker than we would like to have for a signature scheme
|
||||||
|
- Malleability is a weakness of encryption with textbook RSA too
|
||||||
|
|
||||||
|
### Padding
|
||||||
|
|
||||||
|
- If we enforce rules about valid formatting on $m$, random messages produced by attackers are unlikely to pass
|
||||||
|
- 
|
||||||
|
- Likelihood of a successful forgery is $2^{-y}$
|
||||||
|
- Probability of last bit $2^{-1}$
|
||||||
|
- Probability of last 2 bits $2^{-2}$
|
||||||
|
- etc up to $y$
|
||||||
|
|
||||||
|
#### Hash-then-sign
|
||||||
|
|
||||||
|
- It is common to hash the message within any padding scheme
|
||||||
|
- $sig_{k_{prvA}}(x)\equiv H(x)^d \mod n$
|
||||||
|
- Verification recomputes the hash
|
||||||
|
- $ver_{k_{pubA}}(x,s)= s^e \mod n \equiv H(x)'$
|
||||||
|
- $H(x)\stackrel{?}{=}H(x)'$
|
||||||
|
- Existential forgeries are much harder
|
||||||
|
- You’d need a random message that’s also a valid hash
|
||||||
|
- Longer messages can be signed, the hash outputs a smaller message digest
|
||||||
|
|
||||||
|
##### PKCS v1.5
|
||||||
|
|
||||||
|
**P**ublic **K**ey **C**ryptography **S**tandards
|
||||||
|
|
||||||
|
- Modern padding schemes use hashing and padding for security
|
||||||
|
- Prevents existential forgeries, and attacks on small messages
|
||||||
|
- This is deterministic, the same message gives the same signature
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### RSASSA-PSS
|
||||||
|
|
||||||
|
**RSA** **S**ignature **S**cheme with **A**ppendix
|
||||||
|
|
||||||
|
- “with appendix” refers to any scheme that sends $(m,s)$ separately
|
||||||
|
- PKCS and similar schemes are deterministic
|
||||||
|
- The probabilistic signature scheme adds a random salt to the process, meaning repeated singatures on the same document produce different results
|
||||||
|
- Doesn’t effect security that much, some standards have gone back to a probabilistic approach
|
||||||
|
|
||||||
|
###### PSS Encoding
|
||||||
|
|
||||||
|
1. Hash message
|
||||||
|
2. Concatenate padding, hash and salt to create $M'$
|
||||||
|
3. Hash $M’$ into final hash $H$
|
||||||
|
4. Append padding to salt to create data block $DB$
|
||||||
|
5. Expand $H$ using $MGF$
|
||||||
|
6. Calculate $DB \oplus MGF(H)$ to create maskedDB
|
||||||
|
7. Output is maskedDB, $H$ and a constant `0xbc`
|
||||||
|
- `0xbc` is just a constant, no specific meaning other than formatting
|
||||||
|
8. Use RSA to calculate signature and send $(m,s)$ as normal
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
###### PSS Verifying
|
||||||
|
|
||||||
|
(if any of these steps fail, return false)
|
||||||
|
|
||||||
|
1. Use RSA public key to obtain unsigned signature
|
||||||
|
2. Check length and `0xbc` constant
|
||||||
|
3. Split signature into maskedDB and $H$
|
||||||
|
4. Calculate $MGF(H)$ and therefore $DB$
|
||||||
|
5. Check $DB$ padding `00 .. .. 00 1`
|
||||||
|
6. Extract salt from $DB$
|
||||||
|
7. Recreate $M'$ from padding, message and salt
|
||||||
|
8. Calculate $H(M')$
|
||||||
|
9. Verify $H(M')=H$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Nothing is faster than RSA verification, signing is slower
|
||||||
|
|
||||||
|
Its quick because of how 65537 is structured
|
||||||
@@ -0,0 +1,96 @@
|
|||||||
|
# Hash Functions
|
||||||
|
|
||||||
|
#### Multiple Signatures
|
||||||
|
|
||||||
|
- Could we simply split up a message and sign parts?
|
||||||
|
|
||||||
|
\
|
||||||
|
|
||||||
|
A lot of faff for signing large files
|
||||||
|
|
||||||
|
- An attacker can remove $s_{n-1}$ (or $s_{any}$) and it would still be valid
|
||||||
|
|
||||||
|
### Properties of Hash Functions
|
||||||
|
|
||||||
|
1. Any input length
|
||||||
|
2. Fixed output length
|
||||||
|
3. Pre-image resistance (one way)
|
||||||
|
4. Second pre-image resistance
|
||||||
|
- If we have a hashed message, we cannot find another message with the same hash
|
||||||
|
5. Collision resistance
|
||||||
|
|
||||||
|
#### Pre-image Resistance
|
||||||
|
|
||||||
|
- Hash functions must be one-way
|
||||||
|
- Given a hash of a message $H(x)$ it must be infeasible to calculate $x$
|
||||||
|
- Less applicable to digital signatures
|
||||||
|
- Crucial to password storage and key derivation
|
||||||
|
|
||||||
|
#### Second Pre-image Resistance
|
||||||
|
|
||||||
|
- Weak collision resistance
|
||||||
|
- Given a message $x_1$ and a hash of that message $H(x_1)$ it should be infeasible to find a second message $x_2$ such that $H(x_1)=H(x_2)$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Second pre-image attack
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Oscar finds a weak message (one of the messages is known ahead of time), he replaces the message $x_1$ with $x_2$. Now Oscar can send a signed message to Alice
|
||||||
|
|
||||||
|
#### Collision Resistance
|
||||||
|
|
||||||
|
- Strong collision resistance
|
||||||
|
- It is not possible to find **any** message pair $x_1, x_2$ such that $H(x_1)=H(x_2)$
|
||||||
|
- In practice, this is *much easier than finding a weak collision*
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Preventing Collisions
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Collision Attack
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### How Likely
|
||||||
|
|
||||||
|
**Second pre-image attacks**
|
||||||
|
|
||||||
|
- For a 256 bit hash with good random properties we might expect $2^{256}$ bit brute force before we find a collision with $x_1$
|
||||||
|
|
||||||
|
**Collision Attacks**
|
||||||
|
|
||||||
|
- There are many other possible collisions beyond those simply with $x_1$
|
||||||
|
|
||||||
|
### The Birthday Paradox
|
||||||
|
|
||||||
|
> What is the probability two people in this room share a birthday
|
||||||
|
|
||||||
|
- It is easier to first calculate the probability $P(n)$ that $n$ people do not share any birthdays:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\begin{align*}
|
||||||
|
P(2)&=(1-\frac{1}{365}) \\
|
||||||
|
P(3)&=(1-\frac{1}{365})\cdot (1-\frac{2}{365}) \\
|
||||||
|
P(n)&=(1-\frac{1}{365})\cdot (1-\frac{2}{365})\dots (1-\frac{n-1}{365})
|
||||||
|
\end{align*}
|
||||||
|
$$
|
||||||
|
|
||||||
|
- The probability of at least one collision is $1 – P(\textrm{no collision})$.
|
||||||
|
- The probability of a collision with only 23 people is ~50%!
|
||||||
|
- For 40 people it’s ~90%
|
||||||
|
- The same principle applies to hash functions, the more hashes computed, the more likely a collision becomes
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### The Birthday Attack
|
||||||
|
|
||||||
|
- The output of the hash must be long enough to avoid a birthday attack
|
||||||
|
- Given a hash function outputs $n$ bit hashes
|
||||||
|
- You will find a collision after approx $\sqrt{(2^n)}=2^{\frac n2}$ random attempts
|
||||||
|
- This means that your bit length needs to be double the size of your desired security margin
|
||||||
|
- `SHA-256` therefore offers equivalent security to `AES 128`
|
||||||
|
- left at `25:55`
|
||||||
@@ -0,0 +1,190 @@
|
|||||||
|
# Cryptographic Protocols
|
||||||
|
|
||||||
|
### Message Authentication Codes
|
||||||
|
|
||||||
|
- Provide integrity and authenticity - not confidentiality
|
||||||
|
- Protecting system files
|
||||||
|
- Ensuring messages haven’t been altered
|
||||||
|
- Calculate a keyed hash of the message, then append this to the end of the message
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### HMAC
|
||||||
|
|
||||||
|
- Double hashing in HMAC avoids length extension attacks
|
||||||
|
- $HMAC(k,m) = H((k\oplus opad) || H((k\oplus ipad)||m))$
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Authenticated Encryption (AEAD)
|
||||||
|
|
||||||
|
- It’s common to attach MACs to the end of ciphertext, that this is now usually built into ciphers as part of AEAD mode
|
||||||
|
- You’re often able to authenticate non-encrypted “associated” data too
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## Transport Layer Security
|
||||||
|
|
||||||
|
#### SSL/TLS
|
||||||
|
|
||||||
|
- TLS is a protocol that provides *authenticated* and *encrypted* sessions
|
||||||
|
- Secure Socket Layer (SSL) came first, then after `v3.0` it became TLS
|
||||||
|
- Transport Layer Security has two layers
|
||||||
|
1. The record layer
|
||||||
|
- Using established symmetric keys and other session info, will encrypt application packets, very like IPsec
|
||||||
|
2. The handshake layer
|
||||||
|
- Used to establish session keys, as well as authenticate either party - usually the server using a public key certificate
|
||||||
|
|
||||||
|
##### TLS Handshake
|
||||||
|
|
||||||
|
- The TLS handshake allows us to
|
||||||
|
- Establish the master secret
|
||||||
|
- Resume sessions
|
||||||
|
- Authenticate the identity of the server or client
|
||||||
|
- This is for TLS 1.2 - ECDHE_RSA
|
||||||
|
- Elliptic curve with Diffie-Hellman ephemeral with RSA
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**ClientHello**
|
||||||
|
|
||||||
|
```
|
||||||
|
Random nonce: f3bc12ad...
|
||||||
|
Supported Ciphers
|
||||||
|
{
|
||||||
|
TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256
|
||||||
|
TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256
|
||||||
|
TLS_ECDHE_ECDSA_WITH_AES_256_CBC_SHA
|
||||||
|
}
|
||||||
|
[Extensions]
|
||||||
|
[Session ID]
|
||||||
|
```
|
||||||
|
|
||||||
|
**ServerHello**
|
||||||
|
|
||||||
|
```Hell
|
||||||
|
//Pick maximum version client and server can both do
|
||||||
|
Version: 1.2
|
||||||
|
Random Number: 16cf43a...
|
||||||
|
|
||||||
|
//Server chooses the suite out of the ones listed in client hello
|
||||||
|
Suite: TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256
|
||||||
|
[Session ID]
|
||||||
|
```
|
||||||
|
Random nonce used to stop replay attacks
|
||||||
|
|
||||||
|
**Certificate**
|
||||||
|
|
||||||
|
The server sends its public-key certificate to the client
|
||||||
|
|
||||||
|
> **Client verification**:
|
||||||
|
>
|
||||||
|
> The client checks that the public key certificate is valid using a root certificate
|
||||||
|
|
||||||
|
**ServerKeyExchange**
|
||||||
|
|
||||||
|
```
|
||||||
|
Elliptic Curve Diffie-Hellman Parameters:
|
||||||
|
Named Curve: secp256r1 (0x0017)
|
||||||
|
DH Public Key: bG
|
||||||
|
```
|
||||||
|
|
||||||
|
Digital Signature calculated over the DH parameters
|
||||||
|
|
||||||
|
> **Authentication**:
|
||||||
|
>
|
||||||
|
> The client checks that the digital signature is valid
|
||||||
|
|
||||||
|
**[Certificate Request]**
|
||||||
|
|
||||||
|
Optional request for a certificate and singature from the client - only used in mutual TLS
|
||||||
|
|
||||||
|
Imagine two banks communicating where both parties need to prove their identity.
|
||||||
|
|
||||||
|
**ServerHelloDone**
|
||||||
|
|
||||||
|
Signals that there are no further messages to be sent
|
||||||
|
|
||||||
|
**ClientKeyExchange**
|
||||||
|
|
||||||
|
DH Public Key: aG
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**[Certificate]**
|
||||||
|
|
||||||
|
Optional client certificate, verified by the server using PKI
|
||||||
|
|
||||||
|
**[Certificate Verify]**
|
||||||
|
|
||||||
|
Digital signature computed over the bytes send in the handshake so far
|
||||||
|
|
||||||
|
**Change Cipher Spec**
|
||||||
|
|
||||||
|
Signals the change of cipher suite, in this case from no encryption to the agreed encryption
|
||||||
|
|
||||||
|
This can also be done when renewing keys
|
||||||
|
|
||||||
|
**Finished**
|
||||||
|
|
||||||
|
A MAC computed over all handshake messages. Verifies that server and client see the same messages.
|
||||||
|
|
||||||
|
Mitigates man-in-the-middle attacks
|
||||||
|
|
||||||
|
##### TLS 1.3
|
||||||
|
|
||||||
|
**Efficiency**
|
||||||
|
|
||||||
|
- Handshake shortened
|
||||||
|
- Change cipher spec removed
|
||||||
|
- Key exchange sent early in hello messages
|
||||||
|
|
||||||
|
**Security**
|
||||||
|
|
||||||
|
- All ciphers except AEAD removed
|
||||||
|
- Public key and key exchange separated from cipher suites
|
||||||
|
- Some handshake messages are encrypted
|
||||||
|
|
||||||
|
## Public Key Infrastructure
|
||||||
|
|
||||||
|
#### Why do we need PKI?
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
#### Digital Certificates
|
||||||
|
|
||||||
|
- If we want to use public key cryptography, we need *trust*
|
||||||
|
- We can use a trusted third party in order to *verify the ownership of a public key*
|
||||||
|
- Primarily managed through Public Key Infrastructure (PKI)
|
||||||
|
- Certificates usually held in `X509` format
|
||||||
|
|
||||||
|
###### Certificate Issuance
|
||||||
|
|
||||||
|
- A server has a public key that they want people to trust
|
||||||
|
- Using some subject details, the server creates a Certificate Signing Request (CSR)
|
||||||
|
- A Certification Authority (CA) uses this to create and sign a certificate
|
||||||
|
|
||||||
|
###### Certificate Use
|
||||||
|
|
||||||
|
- The server can supply signatures using the public key, backed by the certificate when requested (during the TLS handshake)
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
###### Chains of trust
|
||||||
|
|
||||||
|
- To verify the trust in `server.com` certificate, we need to examine the signing certificate
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
- In many cases, the chain involves multiple certificates
|
||||||
|
- Chains always end in a root certificate, located on your machine
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
##### Who manages the Root Certificates?
|
||||||
|
|
||||||
|
- Major OS vendors operate *root certificate programs*
|
||||||
|
- Apple for iOS and OS X
|
||||||
|
- Microsoft for Windows
|
||||||
|
- Mozilla maintains root certificate store
|
||||||
|
- Used in linux & firefox
|
||||||
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 60 KiB |
|
After Width: | Height: | Size: 85 KiB |
|
After Width: | Height: | Size: 36 KiB |
|
After Width: | Height: | Size: 51 KiB |
|
After Width: | Height: | Size: 74 KiB |
|
After Width: | Height: | Size: 91 KiB |
|
After Width: | Height: | Size: 61 KiB |
|
After Width: | Height: | Size: 99 KiB |
|
After Width: | Height: | Size: 261 KiB |
|
After Width: | Height: | Size: 71 KiB |
|
After Width: | Height: | Size: 140 KiB |
|
After Width: | Height: | Size: 94 KiB |
|
After Width: | Height: | Size: 147 KiB |
|
After Width: | Height: | Size: 46 KiB |
|
After Width: | Height: | Size: 41 KiB |
|
After Width: | Height: | Size: 49 KiB |
|
After Width: | Height: | Size: 51 KiB |
|
After Width: | Height: | Size: 14 KiB |
|
After Width: | Height: | Size: 39 KiB |
|
After Width: | Height: | Size: 19 KiB |
|
After Width: | Height: | Size: 21 KiB |
|
After Width: | Height: | Size: 32 KiB |
|
After Width: | Height: | Size: 32 KiB |
|
After Width: | Height: | Size: 26 KiB |
|
After Width: | Height: | Size: 16 KiB |
|
After Width: | Height: | Size: 57 KiB |
|
After Width: | Height: | Size: 218 KiB |
|
After Width: | Height: | Size: 14 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 14 KiB |
|
After Width: | Height: | Size: 44 KiB |
|
After Width: | Height: | Size: 26 KiB |
|
After Width: | Height: | Size: 27 KiB |
|
After Width: | Height: | Size: 64 KiB |
|
After Width: | Height: | Size: 78 KiB |
|
After Width: | Height: | Size: 50 KiB |
|
After Width: | Height: | Size: 16 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 22 KiB |
|
After Width: | Height: | Size: 225 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 17 KiB |
|
After Width: | Height: | Size: 17 KiB |
|
After Width: | Height: | Size: 37 KiB |
|
After Width: | Height: | Size: 54 KiB |
|
After Width: | Height: | Size: 63 KiB |
|
After Width: | Height: | Size: 57 KiB |
|
After Width: | Height: | Size: 45 KiB |