Part III: Evaluation and regulation of algorithms
In this part, we address the following research questions:
— How should algorithms be evaluated in a research vs industrial context?
— Which are policy needs in terms of the usage of algorithms into real applications?
— Which is the research needed to support these policy needs?
11 Reality, requirements, regulation: Points of intersection with the machine-learning pipeline
Martha Larson
Radboud University
11.1 Statement
Intelligent systems such as search engines and recommender systems play an important role in mediating our consumption patterns and our decisions. We use such systems daily in order to find useful documents in the large amounts of content available online, and we perceive their ability to greatly exceed that of a single human searching by hand. Today’s search engines and recommender systems go beyond documents such as webpages or books to also provide users with multimedia content (e.g., images, music, videos) and items, services, and opportunities in both the online and offline worlds (e.g., places to live, places to eat, jobs to apply for).
At the core of search engines and recommender systems lie machine-learning algorithms. These algorithms can be considered recipes that take a large amount of data as input and provide predictions (i.e., a list of results or recommendations) as output. In contrast to conventional computer programs, it is not possible to know exactly what the output will be in a given situation.
Search engines and recommender systems are considered intelligent since they produce in a blink of an eye a result that would have, in previous eras, taken a long consultation with a reference librarian, or a detailed discussion with the clerk of a video rental store. We experience these systems as doing something that humans are good at doing, doing it faster and doing it in an environment containing seemingly unlimited information. However, how do we know that these systems are actually working the way we assume they are working?
In order to answer this question, we must evaluate the performance of AI systems. Upon first consideration, it seems easy to argue that it is impossible to evaluate search engines and recommender systems. As stated above, the very reason why we build such systems is because we find human effort alone would fail to find the results that AI is capable of generating. If we as humans are not able to find the “right answer” to our information needs, given a large collection of documents, how can we possibly know that what an AI system finds for us is optimal, or even correct? The issue is compounded by the problem of not knowing in advance the output of an AI system in a particular situation, due to the nature of machine learning. However, the dangers are clear: we cannot base our decisions and behavior on systems we are not sure we can trust to be doing what they are supposed to do. Giving up on evaluating AI systems is not an option.
The way to proceed is to return to engineering best practices. In particular, we must specify a set of requirements for our systems, and verify that they meet those requirements. Ultimately, the position that it is impossible to evaluate search engines and recommender system is untenable because adequate effort has not yet been devoted to establishing requirements and to evaluating systems. The history of engineering is a history of finding solutions for problems that initially appear impossible.
In order to come near the amount of effort it will take to apply engineering best practices to the evaluation of search engines and recommender systems it is important to admit the possibility that AI evaluation is actually more difficult than AI itself: If AI is smart, AI evaluation must be smarter. Human engineers created AI in the first place. There is no a priori reason to assume that they are not also able to develop procedures and processes that are actually “smarter than AI”. However, writing requirements and evaluating systems is a tedious and costly process, and must be recognized as such. If we are to insist on recommender systems and search engines that have been properly
evaluated and demonstrated to be performing according to specifications, then we must necessarily give up our assumption that AI is some sort of a shortcut or is necessarily giving us something for free.
The return to engineering best practices will mean not just focusing on system output, but writing requirements for, and evaluating, every step along the machine learning pipeline (See Figure 4).
Figure 4. Machine Learning (ML) Pipeline. Evaluation of AI with respect to requirements
and opportunities for regulation can be identified at every stage of the pipeline.
Formulating requirements for each stage of the pipeline requires considering the interplay of that stage with reality (i.e., with what the machine learning system attempts to capture). These stage-wise requirements provide detailed insight in what is happening inside an AI system. They can be seen as handles that allow us to get a grasp on the functioning of the system and to guide it toward desirable behavior and away from harmful behavior. The requirements represent an opportunity for regulation, which prevents machine learning from causing harm to individuals or society.
It is important to be aware that requirement specifications are in the interest of industry, and that grounding regulations in requirements has potential to support compliance. Companies are highly motivated to work according to engineering best practices. A well-specified set of requirements allows a company to focus on its core objectives, and be sure that the resources it is using are being directed to accomplish these objectives. The benefits of requirements to companies using recommender systems can be seen in concrete requirements documents, e.g., (CrowdRec project, 2014) and joint industry-academic visions on evaluation (Said et al., 2012). Requirements also allow companies to understand how they can adapt to regulations, including both achieving compliance and updating business models. For example, data regulations might make it difficult to compete with respect to collecting users’ personal data (“Data” stage in Error! Reference source not found.), freeing the company from the need to devote esources to this area, and allowing it to shift its attention to competing with respect to algorithms (“ML Algorithm” stage in Error! Reference source not found.). The ontribution of data protection law to the quality of data science R&D has been pointed out by Mireille Hildebrandt (2017), who states, “By requiring the specification of one or more legitimate purposes by the controller, data protection law unwittingly contributes to sustainable research designs that have a bigger chance of making good sense of the data than sloppy exploratory research that covers its tracks in the name of experimentation and the freedom to tinker.” Here, we point out that the positive impact of regulation promoting engineering best practices need not be unwitting.
Each stage of the machine-learning pipeline presents its own challenges, and for none of them is it easy to write requirements. For example, in the case of image search engines, training labels for image data (“Labels” stage) must incorporate the multiplicity of users’ perspectives on images (Larson et al., 2014)(van Miltenburg et al., 2017). Further, requirements for the stages may interact. For example, minimizing training data (“Data” stage) might speed up algorithms (“ML algorithms” stage) and promote user privacy (“Impact” stage) (Larson et al., 2017).
Finally, we point out that “ML Team” and “Impact” are shown in Error! Reference ource not found. with a highlight intended to draw attention to the fact that machine learning starts with people (the people on the ML team that creates systems) and ultimately impacts people. Keeping people central in the ML pipeline requires investing effort into education. Education should include engineering best practices, but also encourage people to understand themselves and their personal worth in an era in which
more and more of our lives are mediated by digital technology, cf. e.g., the thought of Jaron Lanier (2010). Engineering best practices have the potential to remind us that AI may appear magic, but it is the magic of a magician: we enjoy the tricks, but realize that they do not change physical world as we know it. With time and effort, we can also break AI down into its component stages, specify the requirements for each stage, and evaluate and regulate it using those requirements.
11.2 Challenges
- Putting as much effort into designing the requirements for AI, and evaluating AI as we put into creating AI in the first place: concretely, this effort includes building on existing benchmarking initiatives and starting new initiatives.
- Ensuring that the requirements for “Output” and “Impact” are consistent with fair treatment of user populations (bias elimination) and individuals (prohibiting harmful micro-targeting). This in turn means new requirements on “ML Algorithms”, but also on “Data” and “Labels”.
- Ensuring that the next generation of machine learning scientists is trained with the skills necessary to apply engineering best practices to AI evaluation, and are also motivated to do so.
References
— CrowdRec Project. 2014. Crowd-powered recommendation for continuous digital media access and exchange in social networks: Second Iteration Requirements. Deliverable 2.3.
— Hildebrandt, M. Privacy As Protection of the Incomputable Self: From Agnostic to Agonistic Machine Learning (December 3, 2017). Available at SSRN: https://ssrn.com/abstract=3081776
— Lanier, J. You Are Not a Gadget: A Manifesto. 2010. Alfred A. Knopf.
— Larson M., Melenhorst M., Menéndez M., Xu P. (2014) Using Crowdsourcing to Capture Complexity in Human Interpretations of Multimedia Content. In: Ionescu B., Benois-Pineau J., Piatrik T., Quénot G. (eds) Fusion in Computer Vision. Advances in Computer Vision and Pattern Recognition. Springer, Cham.
— Larson, M., Zito, A., Loni, B., Cremonesi, P. 2017. Towards Minimal Necessary Data: The Case for Analyzing Training Data Requirements of Recommender Algorithms. ACM RecSys Workshop on Responsible Recommendation (FATRec 2017). http://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1000&context=fatrec
— Said, A., Tikk, D., Shi, Y., Larson, M., Stumpf, K., Cremonesi, P. 2012. Recommender Systems Evaluation: A 3D Benchmark. ACM RecSys Workshop on Recommender Systems Evaluation: Beyond RMSE (RUE 2012). http://ceur-ws.org/Vol- 910/paper4.pdf
— van Miltenburg, E., Elliott, D., Vossen, P. 2017. Cross-linguistic differences and similarities in image descriptions. Proceedings of The 10th International Natural Language Generation conference, 21–30. http://www.aclweb.org/anthology/W17- 3503
12 Benchmarks and performance measures in artificial intelligence
Anders Jonsson
Universitat Pompeu Fabra
Artificial Intelligence (AI) is the field of computer science that studies the automatic generation of intelligent behaviour from a computational point of view. The term "intelligent behaviour" is usually defined in terms of how difficult it would be for a human to perform a given task (Russell, 2009). Computational problems that are historically considered part of AI include reasoning, knowledge discovery, planning, learning, natural language processing, perception and the ability to move and manipulate objects.
AI algorithms are mainly evaluated along two dimensions: 1) theoretical properties and performance guarantees; and 2) empirical performance.
Historically, the two dimensions carried similar weight, and most AI algorithms were published on the basis of little or no empirical support. This has changed dramatically in recent years, particularly with the advent of deep learning (Goodfellow et al., 2016). Today, empirical performance is a major factor in deciding whether a given AI algorithm is published, and theoretical analysis carries less weight. In fact, most deep learning algorithms come with no performance guarantees whatsoever.
Even though empirical performance is perhaps the most immediate measure of how well an AI algorithm works, theoretical analysis should not be overlooked as an evaluation metric. Performance guarantees come in many forms: an algorithm may eventually converge to the optimal performance level, or converge to a performance level that is within some bound of the optimal. Such an algorithm is guaranteed to always work well, no matter which task it is asked to solve. In contrast, an algorithm with no performance guarantees may perform very well on one task, but fail miserably on another.
An excessive focus on empirical performance provides researchers with strong incentives to boost the performance of their own algorithm relative to other algorithms, since this means their algorithm is more likely to be published. Researchers often make strong claims about their empirical results, such as having "solved" a particular domain or achieving human-level or superhuman-level performance. These claims should often be taken with a grain of salt, and independent verification and reproduction of empirical performance is becoming an increasingly vital task in order to establish the correctness of published results and determine whether they carry over to other domains.
Within the scope of empirical evaluation, there is also the question of precisely how the evaluation is carried out. In the case of stochastic algorithms that depend on random elements, a natural performance measure is the average performance across multiple trials, enhanced with a variance measure to test for robustness. However, researchers often publish the average of the K best trials, often without stating how many trials were carried out in total (Henderson et alt., 2017). The performance of deep learning algorithms is also highly dependent on other factors, such as the initial random seed, the values of hyperparameters, the network architecture, etc. Researchers often do not publish the details of how their algorithm was configured, making it harder to reproduce the work.
To apply AI algorithms in real-world domains it is necessary to establish much more stringent evaluation criteria. Ideally, these criteria should be adopted not only by researchers implementing AI algorithms in real-world domains, but by most or all researchers in AI. In addition, publishing source code and data would make independent verification and reproduction much easier. These measures would make published results much more trustworthy and make it easier to determine which algorithms hold the highest potential for real-world problems. Often, relatively basic statistics are sufficient to
strengthen the aggregate results of multiple trials, compare different alternative algorithms, handle unbalanced datasets, etc.
Another measure that typically strengthens published results is to establish a set of benchmark problems for a given domain. Sometimes benchmarks are accompanied by competitions in which algorithms square off against each other on a subset of the benchmark problems. Benchmarks are normally available to the public, and their purpose is to provide a much more unbiased system for comparing AI algorithms. The more benchmark problems exist and the more diverse they are, the more difficult it becomes to artificially boost the performance of a given AI algorithm. If benchmarks reflect the difficulty present in real-world problems, the performance of an algorithm on the benchmarks should have a higher chance of carrying over to other, similar domains.
As an illustration of the importance of benchmark problems, consider the problem of classifying objects in images. State-of-the-art algorithms for image classification have been evaluated for many years as part of the ImageNet Large Scale Visual Recognition Challenge, or ILSVRC (Russakovsky et al., 2015). From 2010 to 2015, the error rate of the winning algorithm at ILSVRC decreased steadily from 30% to 5%, which is similar to the error rate observed in humans. Since the datasets used for the competition are large and contain a lot of labelled examples for training and evaluation, and since many experiments have been carried out with humans acting as the classifier, many researchers conclude that state-of-the-art algorithms have become competitive with humans.
Benchmarks are not without problems, however. Sometimes benchmarks do not accurately reflect possible real-world problems of a given domain, which may lead researchers in the wrong direction. Sometimes so much computational power is required that only a few select companies or organizations have the computational resources available to solve all benchmark problems. The prospect of winning a competition may also cause researchers to implement algorithms that do not really advance the state-of- the-art. A common example are portfolios that are optimized to select between the most successful existing algorithms by deciding for example how many seconds of computational time should be allocated to each algorithm.
Comparing the performance of AI algorithms with that of humans is not only of academic interest but effectively determines when it may be beneficial to replace human expertise with algorithms. An important aspect that affects the quality of such a comparison is how easy it is to measure success in a given problem. At least part of the reason for the popularity of AI in games is that success is very easy to measure. In real-world problems success may be much harder to measure, however. Eventually, as AI algorithms become better at high-level reasoning, new performance measures will likely be needed since success is not only measured by the performance on a single isolated task, but performance across tasks and how well the algorithm can integrate information from different tasks.