Posts tonen met het label bias. Alle posts tonen
Posts tonen met het label bias. Alle posts tonen

dinsdag 26 november 2019

How useful are A.I., Machine Learning and predictive algorithms for risk assessment and the automated decision making process?

Algorithms at the base of bias-driven results in profiling, automated decisions and risk assessments
Algorithms are the molecules of all forms of Artificial Intelligence (A.I). An algorithm can be described as a formula, a finite series that converts input data (for example, generated by commands in search engines, mouse clicks and visiting web pages) into "output", a certain result. To enable profiling of certain categories of people or phenomena, algorithms must be trained with data sets. The training of algorithms, Supervised Machine Learning, is still a human affair. The value assigned to the dataset by the researcher or client influences the outcome of the data process. 


Like the input of the data set, the outcome depends on prejudices. If the usual method of profiling is continued, a bias in the data sets and therefore in the profile or risk assessment is inevitable. Bias-driven profiling increases the risk of false positives, which is further enhanced by training with outdated (personal) data. Moreover, a weakness is attached to the application of algorithms in the data analysis process: algorithms are unable to describe causality, the relationship between cause and effect. Algorithms are only used to expose correlations between phenomena. That makes predictive algorithms unsuitable for testing expectations. 

I will now elaborate on the weaknesses of the Data Analysis process and Supervised and Unsupervised Machine Learning. I will explain why their weaknesses make these instruments unsuitable as methods for profiling and risk assessment (for example in the context of social security/social insurance fraud). Moreover, I will conclude as to why the outcome of the automated decision making process will be biased since algorithms are inherently biased.

(Big) Data analysis: 'garbage out' is necessary to prevent algorithms from being trained incorrectly
An analytical instrument must be used to separate random or unstructured data from relevant data. The amount of data must not only be reduced, but also refined to arrive at a more specific result. For this purpose use can be made of a so-called "Warehouse", a digital collection of data from various sources. To prevent predictive algorithms from being trained with outdated data and to reduce the risk of false positives, the data must be refreshed and limited in size, ie: "garbage out". Using this reference point, correlations between data can be discovered. Nothing is known about the duration of storing personal data in a Warehouse; on the basis of the aforementioned, it will probably be for an indefinite period of time, except for a regular refreshment. Data controllers should be cautious during this early phase of the data analysis process: if inaccurate data is stored and applied, the person under investigation is incorrectly designated as a 'suspect' or 'potential fraudster'.

Data mining, Supervised Machine Learning and Artificial Intelligence

An important step in discovering correlations between datasets is "Knowledge Discovery of Databases", or "data mining." Statistical Analysis System (SAS) defines data mining as "the process of searching for anomalies, patterns and correlations, in order to predict a certain outcome." The predecessor of data mining is "machine learning", a technique that involves training algorithms based on statistical data. Formulas are entered to develop algorithms, training sets of data are given as "input" and the result is provided as "output". Algorithms are instructed to make the connection between input and output and to evaluate themselves. The result of this feedback is used to identify patterns. This form of "supervised machine learning" is ideally suited to classify data: algorithms categorize data on the basis of pre-provided, labeled data sets and learn to "label" data, or assign a certain characteristic. If a picture of a blackjack is given as input and a similar picture with the title "blackjack" is provided as the output, the algorithms learn to classify pictures of blackjacks.

Unsupervised Machine Learning, unsuitable for profiling

Another form of Machine learning, "Unsupervised Machine Learning", lacks the example of labeled data sets. Algorithms generate patterns between unstructured data. Unsupervised Machine Learning is used to cluster unstructured data, in order to categorize these data without using certain labels. The algorithms, for example, assign all kinds of pictures of blackjacks to one location, but do not know what these weapons are called. This makes Unsupervised Machine Learning unsuitable for the process of profiling, in which not only relationships have to be generated, but also names and classifications (for example "fraud!") will have to be linked to a certain result, the outcome of the profiling process.


A subform of Machine Learning is Deep Learning, the discovery of complex patterns in large amounts of data through a layered neural structure. The distinguishing feature of deep learning is the need for substantial "computational power" to perform a complex task; one neural layer can consist of up to four hundred processors. Machine Learning and Deep Learning Fall can be regarded subsets of Artificial Intelligence (A.I), but note that A.I. is not synonymous with either Machine Learning or Deep Learning. Artificial Intelligence studies the ability of computers to perform complex tasks autonomously and to solve problems.
 
Conclusion: algorithms are trained by human prejudice and therefore judgmental by design
Profiling in the current form is characterized by 'supervised machine learning', the training of algorithms with pre-entered data sets. These data sets, therefore algorithms, express human judgments by design. Algorithms used for profiling are not intelligent in the sense that they can discover patterns themselves. A weakness in profiling is that (predictive) algorithms are not used to test hypotheses, but only to present correlations between data and phenomena. Investigating the causality of an event remains a human matter. 'Mining' large amounts of data is also not a suitable method for testing hypotheses. 
No significance may be attached to the results of profiling and data mining until further human intervention or intervention has been carried out by an appropriate method. "False positives" should be trained out of programs that use predictive algorithms. Since the data sets that train predictive algorithms are biased, the automated decision making process will be biased. Inherently, the outcome of the automated decision making process or risk assessment will be biased. 

Mercedes Bouter LL.M.

maandag 18 november 2019

Het 'black box'-karakter van geheime dataveillance & Big Data analyse en uw rechten (deel II)

Deel I: transparantie, uw concrete informatierechten en uw recht van inzage, rectificatie, wissing en beperking van gegevensverwerking 

Overzicht
1. Uw recht om niet te worden onderworpen aan geautomatiseerde verwerking en profilering;
2. Het recht om zich te verzetten tegen profilering, ongeacht geautomatiseerde besluitvorming;
3. 'Passende maatregelen' (appropriate safeguards) ter compensatie van de inbreuk op uw recht om niet geprofileerd te worden (art. 22 AVG)

1. Uw recht om niet te worden onderworpen aan geautomatiseerde verwerking en profilering
U hebt het recht niet te worden onderworpen aan een uitsluitend op geautomatiseerde verwerking, waaronder profilering, gebaseerd besluit waaraan voor u rechtsgevolgen zijn verbonden of dat u in aanmerkelijke mate treft (art. 22 AVG). 

De risicoprofielen die in SyRI worden opgemaakt vallen onder het bereik van art. 22 AVG. De verwerkingsverantwoordelijke kan slechts een gerechtvaardigd beroep doen op de uitzonderingen van art. 22 lid 2 AVG, indien de betrokkene toestemming heeft verleend of de betrokkene en de verwerkingsverantwoordelijke zulks zijn overeengekomen.

Een uitzondering is buiten die gevallen slechts mogelijk, indien een Unierechtelijke of lidstaatrechtelijke bepaling die op de verwerkingsverantwoordelijke van toepassing is, voorziet in passende maatregelen ter bescherming van de rechten en gerechtvaardigde belangen van de betrokkene die door de profilering wordt getroffen. De verwerkingsverantwoordelijke is slechts gerechtigd om te profileren, indien daarvoor een specifieke wettelijke bevoegdheidsregeling wordt verleend én concrete maatregelen ter bescherming van de rechten van de geprofileerde in de van toepassing zijnde regeling zijn opgenomen. Het Europees Comité heeft aangegeven welke maatregelen minimaal moeten worden getroffen door de verwerkingsverantwoordelijke.
 
2. Het recht om zich te verzetten tegen profilering, ongeacht geautomatiseerde besluitvorming
Het recht van betrokkene om geïnformeerd te worden op grond van art. 13 of 14 AVG in samenhang met art. 5 lid 1 AVG moet aldus worden uitgelegd, dat de betrokkene door de verwerkingsverantwoordelijke moet worden geïnformeerd over het feit dat hij/zij aan profilering wordt onderworpen én dat de betrokkene het recht heeft om zich te verzetten tegen profilering, ongeacht of geautomatiseerde besluitvorming op basis van profilering plaatsvindt (punt 60 van de preambule van de AVG; Guidelines on automated Individual decision-making and Profiling for the purposes of Regulation 2016/679, 6 feb. 2018, WP251rev.01, p. 17). 

Het recht van inzage (art. 15 AVG) houdt in dat de verwerkingsverantwoordelijke inzicht verschaft in de gegevens die als input zijn gebruikt om een profiel van de betrokkene op te maken; tevens dient de betrokkene toegang te hebben tot informatie over de hem betreffende profilering en de details van de categorisering waarin de betrokkene is ingedeeld. Uitzonderingen op deze concrete verplichtingen worden gemaakt voor zover het de handelsgeheimen en intellectuele eigendomsrechten van de verwerker betreft (Guidelines on automated Individual decision-making and Profiling for the purposes of Regulation 2016/679, 6 feb. 2018, WP251rev.01, p. 17).
  
3. 'Passende maatregelen' (appropriate safeguards) ter compensatie van de inbreuk op uw recht om niet geprofileerd te worden (art. 22 AVG) 
Het Europees Comité voor gegevensbescherming, European Data Protection Board, wijst op de gevaren die profilering met zich brengt. Fouten en/of bias in verzamelde en gedeelde data, alsmede fouten en bias in geautomatiseerde (besluitvormings)processen en instrumenten voor profilering resulteren in incorrecte classificaties en beoordelingen gebaseerd op inadequate projecties. De rechten en belangen van betrokkenen worden rechtstreeks getroffen. Het EDPB onderstreept de noodzaak tot transparantie. De betrokkene die aan geautomatiseerde besluitvorming of profilering is onderworpen, kan een beslissing slechts in twijfel trekken, als de verwerkingsverantwoordelijke inzichtelijk heeft gemaakt hoe de beslissing tot stand is gekomen en op welke basis de beslissing is gebaseerd. 

De verwerkingsverantwoordelijke is verplicht om de betrokkene specifieke informatie te verschaffen en te garanderen dat de betrokkene de besluitvorming in het kader van profilering aan kan vechten (punt 71 van de preambule AVG; Guidelines on automated Individual decision-making and Profiling for the purposes of Regulation 2016/679, 6 feb. 2018, WP251rev.01, p. 27). De verwerkingsverantwoordelijke is verplicht om de betrokkene: 

- te informeren over de activiteit waarvoor de data van de betrokkene worden verzameld en verwerkt;
- inzicht te verschaffen in de achterliggende logica van het instrument;
- uitleg te geven over de invloed op de belangen van en mogelijke gevolgen voor de betrokkene.

Het verschaffen van inzicht in de achterliggende instrument voor data-analyse, geautomatiseerde besluitvorming en profilering moet 'betekenisvol' zijn. Om te voorkomen dat de verwerker met onbegrijpelijke statistieken of berekeningen aan komt zetten om aan de plicht tot het verschaffen van inzicht in de logica te voldoen, heeft het EDPB aangegeven dat 'betekenisvolle informatie' over de toegepaste logica (algoritmen) onder meer betrekking heeft op:

- gegevens betreffende in het verleden geconstateerde fraude en betalingsachterstanden;
- de wijze van 'credit scoring': welke waarde wordt toegekend aan bepaalde gegevens, ofwel: in welke mate gegevens tot een bepaalde risicoscore leiden;
- officiële, publiekelijk bekende gegevens, zoals gegevens over insolventie en openbare gegevens over fraude;
- de garantie van de verwerkingsverantwoordelijke dat regelmatig wordt gecontroleerd en bewaakt dat de data-analyse en profilering eerlijk, adequaat en vrij van bias is.
(Guidelines on automated Individual decision-making and Profiling for the purposes of Regulation 2016/679, 6 feb. 2018, WP251rev.01, p. 24-26). 

In het volgende bericht, deel III, ga ik in op de wettelijke grondslag voor de rechtmatigheid van de verwerking van uw persoonsgegevens, de doelbindingseis en de eisen van proportionaliteit en subsidiariteit (noodzaak en dataminimalisatie)