This commit is contained in:
John Gatward committed 2026-10-04 15:24:17 +01:00
1 parent d0f27f276b
commit d6f54d4ec2
103 files changed
+3663 -3779

No files matched your search

@@ -12,11 +12,11 @@ According to code 1.4 from the ACM code of ethics, computing professionals shoul
#### b)
Algorithmic bias is a series of systematic and repeatable errors, that over the course of the systems runtime, produces output that dis-proportionally discriminates against individuals and/or social groups. Selena Silva and Martin Kenny found 9 sources of algorithmic bias in their research paper, all of which capable of discriminating and producing bias
Algorithmic bias is a series of systematic and repeatable errors that, over the course of the system’s runtime, produces output that disproportionately discriminates against individuals and/or social groups. Selena Silva and Martin Kenny found 9 sources of algorithmic bias in their research paper, all of which are capable of discriminating and producing bias
Bias can be introduced in the development of a machine learning system. Training bias is where data used to train the algorithm may be unrepresentive or prejudiced, this can cause the system to unfairly associate one trait to another even though they have no effect on one another. This can be through the developers own bias by only including data sets representative to their own socitak group or through systemic bias where minority groups are under represented in national and global data sets. Developers can also introduce bias by including or excluding certain attributes. This is called algorithmic focus bias and developers must take variables supplied to the algorithm into careful consideration, evaluating why each variable needs to be included in the system. Similarly bias can arise from the way data is processed, for example this can be from weighting quantitative attributes higher than qualitative ones simply as quantitative data is easier to manipulate, this is called algorithmic processing bias. Non-transparency bias is where companies do not divulge or explain how they came to certain decisions, what their rationale was for different design choices. In the best case this can introduce bias in an unforeseen way as all the developers may come from similar social groups and in the worse case scenario developers can obstruct reviews of the algorithm, allowing discrimination to take place.
Bias can be introduced in the development of a machine learning system. Training bias is where data used to train the algorithm may be unrepresentative or prejudiced; this can cause the system to unfairly associate one trait with another even though they have no effect on one another. This can be through the developers’ own bias by only including data sets representative of their own social group or through systemic bias where minority groups are under-represented in national and global data sets. Developers can also introduce bias by including or excluding certain attributes. This is called algorithmic focus bias and developers must take variables supplied to the algorithm into careful consideration, evaluating why each variable needs to be included in the system. Similarly, bias can arise from the way data is processed, for example from weighting quantitative attributes higher than qualitative ones simply because quantitative data is easier to manipulate; this is called algorithmic processing bias. Non-transparency bias is where companies do not divulge or explain how they came to certain decisions or what their rationale was for different design choices. In the best case this can introduce bias in an unforeseen way as all the developers may come from similar social groups and in the worst-case scenario developers can obstruct reviews of the algorithm, allowing discrimination to take place.
Bias can also arise in the use of computing systems. Transfer context bias is where machine learning systems are used inappropriately. This can happen in job applications where credit checks are required or in justice systems where race needs to be explicitly stated. The assumption job performance correlates to wealth or criminal charges correlates to race is unfair and biased. Therefore the use of computer systems particularly in subjective use cases should be scrutinised to ensure the potential benefits outweigh the increased chance of discriminating or additional steps are taken after the system outputs to mitigate any potential harms. Similarly automation bias is where humans hold the output of a system in high regard and don’t question or apply additional thought. Computer systems used in subjective context such as justice systems should be treated as a second opinion or a statistical model and disregarded readily when an unsuitable result is returned. Consumer bias is where bias is introduced to the system via the training data. Humans are inherently flawed and biased and therefore extra care and additional review steps should be added to check the neutrality of the training data. Likewise feedback loop bias affects systems that learn from user behaviour, which again is prone to being discriminatory. This requires special attention has even when a system has been developed without bias, bias is introduced through the systems use lifetime. Lastly interpretation bias is where humans introduce bias from interpreting results from the algorithm. For example if the algorithm agrees with someones own bias, they might be more likely to give a more extreme verdict however if it opposes their own opinion, the result may be completely disregarded.
Bias can also arise in the use of computing systems. Transfer context bias is where machine learning systems are used inappropriately. This can happen in job applications where credit checks are required or in justice systems where race needs to be explicitly stated. The assumption that job performance correlates with wealth or criminal charges correlate with race is unfair and biased. Therefore, the use of computer systems particularly in subjective use cases should be scrutinised to ensure the potential benefits outweigh the increased chance of discriminating or additional steps are taken after the system outputs to mitigate any potential harms. Similarly, automation bias is where humans hold the output of a system in high regard and don’t question it or apply additional thought. Computer systems used in subjective contexts such as justice systems should be treated as a second opinion or a statistical model and disregarded readily when an unsuitable result is returned. Consumer bias is where bias is introduced to the system via the training data. Humans are inherently flawed and biased and therefore extra care and additional review steps should be added to check the neutrality of the training data. Likewise, feedback loop bias affects systems that learn from user behaviour, which again is prone to being discriminatory. This requires special attention as even when a system has been developed without bias, bias is introduced through the system’s useful lifetime. Lastly, interpretation bias is where humans introduce bias from interpreting results from the algorithm. For example, if the algorithm agrees with someone’s own bias, they might be more likely to give a more extreme verdict; however, if it opposes their own opinion, the result may be completely disregarded.
## Question 3
@@ -28,17 +28,16 @@ The application scope is worldwide, the regulation states “This Regulation app
#### b)
To enable proper data protection by design and default, a number of presets must be implemented.
To enable proper data protection by design and default, a number of presets must be implemented.
Firstly controllers must be transparent about how and why they are collecting and using data, how they use and share personal data and how data subjects can exercise their legal rights over data processing. This includes the right to: access, object, intervene, restrict, rectify, export and erase. This allows for data subjects to have full knowledge and control over their data and on top of this, recital 63 of GDPR states “where possible controller[s] should … provide remote access … with direct access to his or her personal data”. Controllers must also by default declare a valid legal basis for the processing. This ensures transparency as there is full disclosure of how data subjects legal rights are being maintained.
Firstly, controllers must be transparent about how and why they are collecting and using data, how they use and share personal data and how data subjects can exercise their legal rights over data processing. This includes the right to: access, object, intervene, restrict, rectify, export and erase. This allows for data subjects to have full knowledge and control over their data and on top of this, recital 63 of GDPR states “where possible controller[s] should … provide remote access … with direct access to his or her personal data”. Controllers must also by default declare a valid legal basis for the processing. This ensures transparency as there is full disclosure of how data subjects’ legal rights are being maintained.
Controllers must ensure their data processing operations are fair. This principle requires personal data should not be processed in ways that are unjustifiably detrimental, unexpected or misleading to the data subject. Fairness is especially prevalent in dealing with AI systems since these do not operate on predefined instructions written by humans, therefore controllers should be able to demonstrate fairness through the inputs and outputs of the system.
Controllers must explicitly state what the data collected on data subjects will be used for. These must be specific tasks and cannot be processed in way that doesn’t align with the initial reason given. This is called purpose limitation and prevents controllers from collecting as much data as possible for monetary gain or nefarious purposes. This allows data subjects to only give their data to controllers who’s vision aligns with their own.
Controllers must explicitly state what the data collected on data subjects will be used for. These must be specific tasks and the data cannot be processed in a way that doesn’t align with the initial reason given. This is called purpose limitation and prevents controllers from collecting as much data as possible for monetary gain or nefarious purposes. This allows data subjects to only give their data to controllers whose vision aligns with their own.
Controllers must practise data minimisation, this is a practice where the controller must review the data being asked and verifying all pieces of data are needed to meet the purposes for which they are being processed. This can also include the degree of identification, if the purpose is statistical this likely does not require any immediate identifying attributes. If continued identification is needed, data should be pseudonoymised to migrate damages caused from a data breach. Similarly data must be deleted once it has fulfilled it’s purpose. GDPR places no time limit on data storage of anonymised data however this can be reversed engineered and this data should be treated analogous to raw personal data.
Controllers must practise data minimisation; this is a practice where the controller must review the data being requested and verify that all pieces of data are needed to meet the purposes for which they are being processed. This can also include the degree of identification; if the purpose is statistical this likely does not require any immediate identifying attributes. If continued identification is needed, data should be pseudonymised to mitigate damage caused by a data breach. Similarly, data must be deleted once it has fulfilled its purpose. GDPR places no time limit on data storage of anonymised data; however, this can be reverse-engineered and this data should be treated analogously to raw personal data.
Controllers must also ensure data is accurate, and if not it is the controllers duty to rectify or erase mistakes immediately. This is important as data subjects could be relying on this data for employment, housing or other civic needs and not being able to obtain this could cause harm to the data subject and family.
Controllers must also ensure data is accurate, and if not it is the controller’s duty to rectify or erase mistakes immediately. This is important as data subjects could be relying on this data for employment, housing or other civic needs and not being able to obtain this could cause harm to the data subject and family.
Lastly controllers must put substantial measures in place to prevent unauthorised access, accidental loss and destruction or damage. Regular reviews should be conducted, testing security and inviting professional hackers to further test how the system stands up to new hacking methods.