Files
notes/docs/lectures/ethics/07_data_protection.md
T

21 KiB
Raw Blame History

Data Protection by Design and Default

DPbDD - Data Protection by Design and Default

What is Privacy
  • It is a state in which one is not observed or disturbed by other people
  • Or the state of being free from public attention
  • Or someone’s right to keep their personal matters and relationships secret
  • Or freedom from unauthorised intrusion
  • Or the right to make personal decisions regarding intimate matters
  • Or the right to lead one’s life in a manner that is reasonably secluded from public scrutiny

And so on

Privacy allows us to negotiate who we are and how we want to interact with the world around us, and is essential to who we are as human beings. It gives us a space to be ourselves without judgement, allows us to think freely without discrimination, and is essential to individual autonomy and the protection of human dignity.

https://privacyinternational.org/explainer/56/what-privacy

“No one shall be subjected to arbitrary interference with his privacy, family, home or correspondence, nor to attacks upon his honour and reputation. Everyone has the right to the protection of the law against such interference or attacks.” Article 12 of the UN declaration

Article 8.1 of the EU convention – the right to respect for private and family life

  1. Everyone has the right to respect for his private and family life, his home and his correspondence;
  2. There shall be no interference by a public authority with the exercise of this right except such as is in accordance with the law and is necessary in a democratic society in the interests of national security, public safety or the economic well-being of the country, for the prevention of disorder or crime, for the protection of health or morals, or for the protection of the rights and freedoms of others.

Privacy is a fundamental human right and underpins many other human rights including freedom of association and free speech.

It’s politically contentious status makes it an ethical imperative in professional computing and key to ensuring public confidence and trust.

That’s why the BCS and ACM include “respect for privacy” as a requirement in their ethics codes, and the IEEE has a separate Data Access and Use policy to align it with industry best practice and ensure compliance with international regulations including the European Union’s General Data Protection Regulation or GDPR

Informational Privacy

Informational privacy is a subset of general privacy concerns.

Warren and Brandeis argued that technology enabled harms to privacy including intrusion into one’s private life and affairs

  • public disclosure of embarrassing private facts
  • unwanted publicity
  • misuse of a person’s name or likeness for financial advantage.

Informational privacy is thus a concern with the protection of personal or private information from unauthorised disclosure and misuse. https://plato.stanford.edu/entries/privacy

Relevant Authority

From a UK perspective, GDPR still applies because it was adopted into UK law by the Data Protection Act 2018 and is enforced by the Information Commissioner’s Office or ICO, so it’s clearly relevant to BCS accreditation.

GDPR has global reach and violations may result in fines of up to 20 million euros or 20 up to 4 % of the total worldwide annual turnover of the preceding financial year, whichever is higher (Article 83)

Definitions

The data subject is a natural person, an individual who can be identified, directly or indirectly, by the personal data.

Personal data is any information relating to an identified or identifiable person (i.e., the ‘data subject’), either directly or indirectly. Personal data includes a bunch of technical information including such things as account handles, IP or MAC addresses, cookies, RFID frequencies, device fingerprints, etc.

  • The key point here is that personal data may not directly link to a data subject as say a passport might
  • But may relate indirectly to a person once the data has been procesed

Processing means any operation or set of operations which is performed on personal data or on sets of personal data, whether or not by automated means.

Processing includes:

  • collection
  • structuring
  • storage
  • alteration
  • retrieval
  • combination
  • adaptation
  • consultation
  • use, disclosure, dissemination, making available, restriction, erasure or destruction of personal data.

Basically, if you touch someone’s personal data in any way you are involved in processing it. Processing does not just mean the data is run through a computer in some way.

Similarly, processor does not refer to a CPU on a computer, but to the person, legal entity, public authority, agency or other body which processes personal data on behalf of the controller and may use computing to do so.

Controller means the person, legal entity, public authority, agency or other body which, alone or jointly with others, determines the purposes for which personal data will be processed and the means of processing them.

Data protection officer or DPO, who may be an employee of the controller or processor or an independent contractor who has expert knowledge of data protection law and must be consulted by the controller or processor in a timely manner in all issues which relate to the protection of personal data. A DPO must be appointed if a controller or processor’s core activities involve the processing of personal data on a large scale or involve large scale, regular and systematic monitoring of individuals.

GDPR

GDPR places specific legal requirements on controllers, which directly impact processors.

Article 23 of GDPR says that, “1. Taking into account the state of the art, the cost of implementation and the nature, scope, context and purposes of processing as well as the risks … posed by the processing, the controller shall, both at the time of the determination of the means for processing and at the time of the processing itself, implement appropriate technical and organisational measures … in an effective manner and … integrate the necessary safeguards into the processing in order to meet the requirements of this Regulation and protect the rights of data subjects. 2. The controller shall implement appropriate technical and organisational measures … by default …”

The European Data Protection Board or EDPD, which furnishes guidance on GDPR tells us that, “a ‘default’, as commonly defined in computer science, refers to the pre-existing or preselected value of a configurable setting that is assigned to a software application, computer program or device. Such settings are also called ‘presets’ or ‘factory presets’.” EDPB Guidelines

So the term implement by default in GDPR refers to the design of preset technical and organisational measures to ensure that data processing operations meet the requirements of GDPR and thus protects the legal rights of data subjects. We’ll take a look at what those presets are about shortly.

The controller is legally accountable for the choice of presets and implementing data protection by design and default. (Article 5 GDPR)

This means that the controller must be able to demonstrate to themselves, to data subjects and to supervisory authorities alike that the technical and organisational measures they have put in place are a) appropriate and b) effective in ensuring data protection by design and default.

Data Protection by Design and default

By default controllers must be transparent about how they collect, use and share personal data and how data subjects may exercise their legal rights over data processing.

These include:

  • the right to access any personal data held by the controller that relates to the data subject (Article 15)
  • to object to the processing of personal data (Article 21)
  • obtain human intervention when querying automated decisions (Article 22)
  • to restrict processing (Article 18)
  • to rectify inaccuracies (Article 16)
  • to export data in a commonly used and machine-readable format (Article 20)
  • to have data erased and be forgotten (Article 17).

Recital 63 which says, “Where possible, the controller should be able to provide remote access to a secure system which would provide the data subject with direct access to his or her personal data.”

So transparency is something that needs to built into systems in the long term and not simply be seen as a matter of appending documentation to their use.

The controller must also by default identify and declare a valid legal basis for the processing. Six legal grounds exist including:

  1. consent
  2. performance of a contract
  3. compliance with a legal obligation
  4. protecting vital interests
  5. carrying out a task in the public interest or official duty
  6. pursuing legitimate interests.

Fairness is an overarching principle of data protection, which requires that personal data should not be processed in ways that are unjustifiably detrimental, unexpected or misleading to the data subject.

Fairness is especially important with respect to data processing operations that rely on machine learning and AI, for as we saw in lecture 2 these technologies are responsible for widespread discrimination.

The controller must also ensure that data is only collected for specific, explicitly stated purposes and that data is not further processed in a manner that is incompatible with the purposes for which they were initially collected.

This is called purpose limitation. It means a controller cannot simply collect as much data as they like and do with it what they want. Data collection must be limited by default to specific purposes which are transparent to the data subject.

Data minimisation: the controller must ensure that data collection is limited to what is necessary to meet the purposes for which they are being processed.

Data minimisation requires that the controller verify whether the purposes can be achieved by processing less personal data, or having less detailed or aggregated personal data or without having to process personal data at all. Such verification should take place before any processing takes place, and be carried out at any during the processing lifecycle.

Data minimisation also refers to the degree of identification. If the purpose does not require the final set of data to refer to an individual (such as statistics) - then the controller should delete or anonymise personal data as soon as possible. If continued identification is needed for other processing activities, personal data should be pseudonymized to mitigate risks for the data subjects’ rights.

By default, the controller must limit the period for which personal data kept in a form which permits identification of data subjects are stored and retain data in such a form for no longer than is necessary to meet the purposes for which it has been collected.

No time periods are specified by GDPR, it all depends on the purposes for which the data was collected and the risks that attach to keeping the data in identifiable form. Anonymised data can be stored indefinitely, though risks of reverse engineering attach to pseudonymised data, which need to be mitigated if the data is to be retained for long periods.

By default, the controller must put technical and organisational measures in place to protect personal data against unauthorised access, accidental loss, destruction or damage, and to manage data breaches.

  • Regular reviews should be conducted to make sure it is being stored securely

Data Protection Impact Assessment

DPIA - Data Protection Impact Assessments

DPIAs are mandated by Article 35 GDPR, which says that “Where a type of processing in particular using new technologies … is likely to result in a high risk to the rights and freedoms of natural persons, the controller shall, prior to the processing, carry out an assessment of the impact of the envisaged processing operations on the protection of personal data.”

Article 35 goes on to say that a DPIA “shall in particular be required” in the case of automated processing, including profiling, and systems that produce decisions that have legal effects (e.g., which effect a person’s right to claim state benefits) or similarly significantly affect the natural person (e.g., by assessing their creditworthiness). This applies especially to machine learning and AI systems.

A DPIA is also required by law where large amounts of special category data are processed.

Special category data is data that reveal racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, bio-metric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation.

DPIAs are legally required for these areas of personal data processing, but they are generally recommended as “good practice” for any processing of personal data. https://ico.org.uk/for-organisations/guide-to-data-protection/guide-to-the-general-data-protection-regulation-gdpr/accountability-and-governance/data-protection-impact-assessments/

How to know when processing is high risk

There are 4 critieria specified in GDPR article 35

  1. The use of new technologies to process personal data
  2. Automated-decision making with legal or significant effect
  3. Processing of special category data
  4. Systematic monitoring of public spaces

There are additional criteria

  1. Evaluation or scoring, including profiling and predicting
  • especially of data concerning the data subject's performance at work, economic situation, health, personal preferences or interests, reliability or behavior, location or movements.
  • Examples of this are financial institutions that screen customers against a credit reference database
  1. The processing of sensitive data or data of a highly personal nature
    • Not only special categories of personal data, but also any data considered as sensitive as the term is commonly understood
    • e.g., data linked to household and private activities (such as electronic communications), or data that impact the exercise of a fundamental right (such as location data whose collection may impact freedom of movement), financial data, personal documents, personal information contained in life-logging applications, etc.
  2. The processing of personal data on a large scale
    • which is determined by the number of data subjects concerned
    • the volume of data and/or the range of different data items being processed
    • the duration or permanence of the data processing activity
    • the geographical extent of the processing activity
  3. Matching or combining datasets
    • data originating from two or more data processing operations performed for different purposes and/or by different data controllers in a way that would exceed the reasonable expectations of the data subject.
  4. Data is processed that relates to vulnerable data subjects
    • For example, children, employees, and vulnerable persons requiring special protection such as mentally ill persons, asylum seekers, the elderly, patients, etc.
    • Indeed any personal data where an imbalance in the relationship between the data subject and the controller can be identified and processing increases the power imbalance between them.
  5. Data processing that prevents data subjects from exercising a right, using a service or entering into a contract
    • This includes processing operations that permit, modify or refuse data subjects’ access to a service or entry into a contract.
    • An example of this is where a bank screens its customers against a credit reference database in order to decide whether to offer them a loan.

If a processing operation meets 2 of these criteria, then a DPIA is required by law.

Whats involved in carrying out a DPIA?

Step 1

Identify the need for a DPIA, which is done by applying the criteria we have just discussed.

Must be done before processing takes place

Step 2

Specify the nature of the processing including the source of the data

  • the nature of the data and its status (e.g., special category, sensitive, vulnerable, etc.)
  • how it will be collected, used, stored and deleted
  • the amount of data to be collected
  • the frequency and duration of collection and storage, and the geographical area covered
  • the flow of data and if it will be shared, how and with who
  • any types of processing that are identified as high risk.

Also involves specifying the purpose or purposes of the processing and what the controller wants to achieve by processing the data, including the intended effect on data subjects (if any), the benefits of the processing to the controller and more broadly.

Step 3

Is consider the need for consultation

  1. when and how the views of data subjects will be sought
  2. justifying why it is not appropriate to do so

Third & external parties need to be consulted to ensure data protection by design and default

Step 4

Accessing necessity and proportionality, which involves specifying how the processing will actually achieve the purpose and that there is no other way to achieve the same outcome.

  • the lawful basis for processing
  • how data minimisation and data quality will be ensured
  • how function creep will be prevented; what information will be given to data subjects and their rights will be supported
  • measures that will be taken to ensure processors are in compliance with DPbDD
  • how any international data transfers will be safeguarded.
Step 5

Identify and assess risks and involves identifying sources of risk and specifying

  1. risks to data subjects
  2. corporate risks
  3. compliance risks

and the potential impact of each.

Step 6

Identify and specify measures to mitigate the risks, including the options available to

  1. reduce risk
  2. eliminate risk
Step 7

Have the DPAI signed off and outcomes recorded. If the DPO’s advice is overruled, justification must be provided, as must the reasons for not abiding by consultation outcomes. A review date must also be specified for the DPIA and done so over the lifetime of a processing operation.

You cannot do a DPIA on your own. IBM’s Dave Whitelegg says you must have the following invovled

  • The developer lead or project manager, who is responsible for managing the DPIA process.
  • A data protection officer who must be consulted about and sign off on the DPIA process*.*
  • A security specialist who must verify that best practises are adopted throughout development.
  • A risk manager to advise on privacy risk management.
  • Project sponsors and business directors, who are accountable for privacy risks.
  • And where processing operations are developed for external organisations, who must be able to verify that the processing is compliant with GDPR.

Relevance of DPbDD and DPIA to computing

Now you might be tempted to think that data protection by design and default and DPIAs have little if anything to do with the actual business of developing computing systems.

However, we should not forget that documentation is a key part of the software engineering process – particularly requirements engineering – and that poor requirements specification is a primary source of computing failure.

As IBM’s Dave Whitelegg puts it, “GDPR privacy obligations should be documented as requirements within the requirements analysis phase of the Software Development Lifecycle or SDLC. The DPIA should be performed within the design phase of the SDLC. Then, further privacy risk verification should be conducted throughout the latter phases of the SDLC, to assure the privacy requirements are all achieved, and the design mitigates or eliminates privacy risks as intended.” Dave Whitelegg (2018) Application privacy by design

Privacy Engineering

Dave Whitelegg also tells us that while GDPR does not prescribe technical solutions, indeed it is in its own words “technologically neutral” (GDPR, recital 15), that nonetheless a “number of technical solutions that can be utilized to significantly enhance the protection of personal data” Dave Whitelegg (2018) Minimising application privacy risk.

OWASP’s Security Principles
  1. Data anonymisation methods include: nulling, deletion and redaction, which involves removing all direct and indirect identifier fields in a dataset,
    • removing names or postcodes.
  2. Substitution, which involves overwriting personal data identifier fields with fake personal data.
  3. Data masking, which involves substituting identifier field characters with a ‘mask’ character,
    • e.g., inserting X’s instead numbers on a credit card field.
  4. Scrambling / shuffling, which involves moving the contents of identifier fields around
    • e.g. moving surnames up or down.
  5. Aggregation / generalisation, which involves rendering data in statistical form.
  6. Hashing provides a method of pseudonymisation and involves using an algorithm to transform personal data fields into alphanumeric strings.
  7. Penetration testing is recommended to verify whether these methods enable reidentification in any actual case.

https://owasp.org/www-project-top-ten/

Privacy engineering may help you implement the presets and meet the requirements, but it is your ethical responsibility to know and respect the rules that pertain to professional work. You now know what rules you need to follow to respect people’s privacy and protect their data.