OER·harvester

← Back to the library
arXiv HTML resource

Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead

Black box machine learning models are currently being used for high stakes decision-making throughout society, causing problems throughout healthcare, criminal justice, and in other domains. People have hoped that creating methods for explaining these black box models will alleviate some of these problems, but trying to \textit{explain} black box models, rather than creating models that are \textit{interpretable} in…

Licence
SHARE_ALIKE CC-BY-SA-4.0
Authors
Cynthia Rudin
Published
2018-11-26 · arXiv
Language
en
Length
13437 words
Type
narrative text

Cites 1 work

inferred
Open original ↗

Appendix C Counterfactual Explanations

Some have argued that counterfactual explanations [37, e.g., see] are a way for black boxes to provide useful information while preserving secrecy of the global model. Counterfactual explanations, also called inverse classification, state a change in features that is sufficient (but not necessary) for the prediction to switch to another class (e.g., “If you reduced your debt by $5000 and increased your savings by $50% then you would have qualified for the loan you applied for”). This is important for recourse in certain types of decisions, meaning that the user could take an action to reverse a decision [61].

There are several problems with the argument that counterfactual explanations are sufficient. For loan applications, for instance, we would want the counterfactual explanation to provide the lowest cost action for the user to take, according to the user’s own cost metric. [See 35, for an example of lowest-cost counterfactual reasoning in product rankings]. In other words, let us say that there is more than one counterfactual explanation available (e.g., the first explanation is “If you reduced your debt by $5000 and increased your savings by $50% then you would have qualified for the loan you applied for” and the second explanation is “If you had gotten a job that pays $500 more per week, then you would have qualified for the loan”). In that case, the explanation shown to the user should be the easiest one for the user to actually accomplish. However, it is unclear in advance which explanation would be easier for the user to accomplish. In the credit example, perhaps it is easier for the user to save money rather than get a job or vice versa. In order to determine which explanation is the lowest cost for the user, we would need to elicit cost information for the user, and that cost information is generally very difficult to obtain; worse, the cost information could actually change as the user attempts to follow the policy provided by the counterfactual explanation (e.g., it turns out to be harder than the user thought to get a salary increase). For that reason it is unclear that counterfactual explanations would suffice for high stakes decisions. Additionally, counterfactual explanations of black boxes have many of the other pitfalls discussed throughout this paper.