Recommendations
- There is a need for more awareness of the limitations of algorithms and the extent to which they can do harm.
- There is a need for robust evaluation frameworks that do not stop at abstract evaluation metrics (such as accuracy and precision) but that consider other elements of how an algorithm can affect the lives of people.
- There is a need for mechanisms of effective algorithmic transparency, which is not mere access to source code but a degree of algorithmic explainability that enables humans to understand and challenge algorithmic decisions.
8.2 Context
Despite lacking “subjective” elements in their decisions, algorithms, particularly predictive modelling algorithms that are used to support decision-making, can be discriminatory. A more precise formulation follows.
Generic discrimination
This and following definitions are adapted from Lippert-Rasmussen [2013].
X discriminates against someone Y in relation to Z if:
- Y has property P and Z does not have property P (or X believes Y has property P and X believes Z does not have property P)
- X treats Y worse than s/he treats or would treat Z
- It is because Y has P (or because X believes Y has P) and Z does not have P (or because X believes Z does not have P) that X treats Y worse than Z In other words, generic discrimination is disadvantageous differential treatment.
Group discrimination
X group-discriminates against Y in relation to Z if:
-
X generically discriminates against Y in relation to Z 1 C. Castillo is partially funded by La Caixa project LCF/PR/PR16/11110009.
-
P is the property of belonging to a socially salient group
-
This makes people with P worse off relative to others or X is motivated by animosity towards people with P,
or by the belief that people with P are inferior
or should not intermingle with others
“A group is socially salient if perceived membership of it is important to the structure of social interactions across a wide range of social contexts” [Lippert-Rasmussen, 2013]
A socially salient group could be, for instance, gay people. A non-socially salient group could be, for instance, people with brown eyes.
Statistical discrimination
X statistically discriminates against Y in relation to Z if:
- X group-discriminates against Y in relation to Z
- P is statistically relevant (or X believes P is statistically relevant) For example:
- If an employer does not hire a highly-qualified woman because among his/her current employees, women have a higher probability of taking parental leave, then this employer is engaging in statistical discrimination.
- However, if an employer does not hire a highly-qualified woman because she has informed him/her that she intends to have a child and take parental leave, then this employer is engaging in non-statistical discrimination.
In statistical machine learning
An algorithm developed through statistical machine learning can statistically discriminate if we:
- Disregard intentions/animosity from the definition of group discrimination
- Understand the “statistically relevant” part of the definition as any information derived from training data. Below, we provide five examples of discrimination in algorithms.
Disparate impact. The model gives people with P a bad outcome more often, or in other terms, people with P experience a higher risk.
Let’s assume the following table:
Figure 1. Benefit for protected and unprotected groups.
Suppose:
"Protected group" ="people with disabilities"
"Benefit granted" = "getting a scholarship"
Intuitively, if:
a/n1, the risk that people with disabilities face of not getting a scholarship is much larger than c/n2, the risk that people without disabilities face of not getting a scholarship, then people with disabilities could claim they are being discriminated (see, e.g., Pedreschi et al. (2012))
Directly and indirectly discriminatory rules. A model associates attribute P to a bad outcome, or attribute Q which depends on P, to a bad outcome.
Directly discriminatory rules
Suppose that from a database of decisions made in the past, after applying an associations rule mining algorithm, we learn that gender = female ⇒ credit = no
P(gender=female, credit=no) / P(gender=female) > θ
This means we have found evidence of direct discrimination (Hajian et al. 2013).
Lack of calibration. The same output translates to different bad outcome probs. for P and not P.
A model lacks calibration if the probability of an actual bad outcome depends on the class and not only on the output.
For instance, a well-calibrated model for risk of recidivism should have the property that, for every recidivism score generated by the algorithm, the probability of recidivism for different groups if the same.
Disparate mistreatment / Lack of equal opportunity. The false positive rate of the bad outcome is higher for P than not P.
Suppose we have the following distributions of scores (taken from ProPublica’s research on COMPAS in 2016, https://github.com/propublica/compas-analysis).
Figure 2. COMPAS scores distribution.
In the left plot, we see that the average score given to white and black defendants are different. This does not mean immediately that there is discrimination, but in the right plot, we observe that if we consider only those who did not commit a new crime (i.e., non-recidivists), the scores are still different. This gap means disparate mistreatment, or lack of equal opportunity.
Unfair rankings. The model gives people with P a lower ranking.
Figure 3. Ranking comparison for different genres.
The results above are Top-10 results for job searches in XING (a recruitment site similar to LinkedIn), for selected professions: “Economist”, “Market Analyst”, and “Copywriter” (Zehlike et al., 2017). We observe that there is a difference in proportions between the top-10 and the top-40 in the three cases.
Algorithm-human interaction
Humans need to be able to receive explanations, and to correct outcomes.
Effective transparency does not mean source code, it means the human can understand and challenge the algorithmic decisions.
For instance, in RISCANVI (the method used to predict recidivism in Catalonia), experts can correct the assessment generated by the algorithm, indicating, for instance, that in their opinion the defendant has a higher/lower probability of recidivism that what the system generates.
8.3 Challenges
A personal opinion on transparency
- Many "customers" of algorithmic decision making systems do not value transparency. Transparency gives them insight on an algorithm and may generate doubts. Many want certainty, even if it is a false certainty.
- This is a perverse incentive for developers/providers, who may exaggerate their claims of accuracy.
- An additional problem is that numbers, plots, charts, suggest objectivity, so lack of mathematical literacy becomes problematic, as people tend to trust automated systems more than what they should. Challenges
- We need good evaluation frameworks. How can this be fixed? Improving mathematic literacy and using evaluation frameworks that integrate multiple dimensions In addition to accuracy: "dollars saved, lives preserved, time conserved, effort reduced, quality of living increased" [Wagstaff 2012] and respect to privacy, fairness, accountability, transparency.
- To the extent that algorithms can engage in disadvantageous differential treatment that leaves people of a socially salient group worse-off, based on statistical information, algorithms can discriminate.
- Current research looks at trade-offs of utility and fairness and at mechanisms for mitigating unfairness.