OER·harvester

← Back to the library
arXiv PDF resource

Assessing the impact of machine intelligence on human behaviour: an interdisciplinary endeavour

This document contains the outcome of the first Human behaviour and machine intelligence (HUMAINT) workshop that took place 5-6 March 2018 in Barcelona, Spain. The workshop was organized in the context of a new research programme at the Centre for Advanced Studies, Joint Research Centre of the European Commission, which focuses on studying the potential impact of artificial intelligence on human behaviour. The works…

Licence
OPEN_NC CC-BY-NC-SA-4.0
Authors
Emilia Gómez, Carlos Castillo, Vicky Charisi, Verónica Dahl, Gustavo …
Published
2018-06-07 · arXiv
Language
en
Length
28859 words
Type
narrative text

Cites 4 works

inferred
Open ↗ Download Open original ↗

Part I: Human vs Machine Intelligence

In this part of the document, we address the following research questions:

— Which are the fundamental differences between human and machine intelligence?

— How do algorithms complement or replace human tasks now and how will they do this in the future?

— Will algorithms that take over some of our tasks affect the balance between human and machine intelligence?

We present a set of statements which address these questions from different disciplines and views, and provide a summary of discussions.

2 Unintuitive properties of deep neural networks

Joan Serrà
Telefónica Research

2.1 Statement

Deep neural networks are currently a hot topic, not only within both academia and industry, but also among society and the media. However, interestingly, the current success and practice of deep learning seems to be uncorrelated with its theoretical, more formal understanding. In particular, we find a number of unintuitive properties both in their design and operation that do not have yet an agreed explanation. Interestingly, some of these unintuitive properties are shared with humans, although current neural network approaches and the underlying mechanisms that lead to those properties do not resemble human mechanisms.

  • Neural networks can make dumb errors — neural networks can produce totally unexpected outputs from inputs with perceptually-irrelevant changes, which are commonly called adversarial examples. Humans can be also confused by ‘adversarial examples’: we all have seen images that we guessed were something (or a part of something) and later we were told they were not. However, the point here is that human adversarial examples do not correspond to those of neural networks because, in the latter case, they can be perceptually the same (Szegedy et al., 2014).
  • The solution space is unknown — as with many other machine learning algorithms, the training of neural networks proceeds by finding a combination of numbers, called network parameters or weights that yield the highest performance or, more properly, the minimum loss on some data. There are well known methodologies to find such a minimum for a few parameters with theoretical guarantees. However, deep neural networks are typically in the range of millions of parameters, for which a suitable combination that minimizes a certain loss must be found. The losses of current deep networks are non-convex, with multiple local minima and potentially many obstacles (Li et al., 2017).
  • Neural networks can easily memorize — recent work empirically shows that finite-sized networks can model any finite-sized data set, even if this is made of shuffled data, random data, or random labels (Zhang et al., 2017). This has the obvious implication that neural networks can remember any data seen during training, no matter the nature of that data. What is not so obvious is that, still, if the data is not totally random, neural networks are totally capable of extrapolating their memories to unseen cases and generalize. Doing so when the number of model parameters is several orders of magnitude larger than the number of training instances is what is intriguing and contradicts conventional machine learning wisdom.
  • Neural networks can be compressed — one can drastically reduce the number of parameters of a trained neural network and still maintain its performance on both seen and unseen data (Han et al., 2016). In some cases, the amount of pruning or compression is surprising: up to 100 times depending on the data set and network architecture. Besides practical considerations, the compressibility of networks poses several questions: Do we need a large network in the first place? Is there some architecture twist that combined with current minimum-finding algorithms allows to discover good parameter combinations for those small networks? Or is it just a matter of discovering new minimum-finding algorithms?
  • Learning is influenced by initialization and example order — As with human learning, current network learning depends on the order in which we present the examples. Practitioners know that different sample orderings yield different performances and, in particular, that early examples have more influence on the final accuracy (Erhan et al., 2010). Furthermore, it is now a classic trick to pre-train a

neural network in an unsupervised way or to transfer knowledge from a related task to benefit from additional sources. In addition, it is easy to show that random initializations can affect the final accuracy or, in the worst case, just prevent the network to learn at all.

  • Neural networks forget what they learn — This phenomenon is known as catastrophic forgetting or catastrophic interference (McCloskey & Cohen, 1989). Essentially, when a neural network that has been trained for a certain task is reused for learning a new task, it completely forgets how to perform the former. Beyond the philosophical objective of mimicking human learning and whereas machines should be able to do so or not, the problem of catastrophic forgetting has important consequences for the current development of systems that consider a large number of (potentially multimodal) tasks, and for those which aim towards a more general concept of intelligence. Some research is devoted to tackle catastrophic forgetting, but a general solution for a compact model is still not yet fully in place (Serrà et al., 2018).

2.2 Challenges

  • Adversarial examples are a big challenge right now. Perhaps we are not going to be able to solve the situation until some of the other the inner workings of neural networks are properly understood.
  • Generalization is another principal hurdle. Current statistical theories for generalization are based on quite old models (like support vector machines and the like) that do not represent the state-of-the-art in many machine learning tasks.
  • A flexible and straightforward solution to the problem of catastrophic forgetting, especially considering limited resources.

References

— Erhan, D., Bengio, Y., Courville, A., Manzagol, P.-A., Vincent, P., & Bengio, S. (2010). Why does unsupervised pre-training help deep learning? Journal of Machine Learning Research, 11, 625–660.

— Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: compressing deep neural networks with pruning, trained quantization and Huffman coding. In Proc. of the Int. Conf. on Learning Representations (ICLR).

— Li, H., Xu, Z., Taylor, G., & Goldstein, T. (2017). Visualizing the loss landscape of neural nets. ArXiv: 1719, 09913.

— McCloskey, M., & Cohen, N. (1989). Catastrophic interference in connectionist networks: the sequential learning problem. Psychology of Learning and Motivation, 24, 109–165.

— Serrà, J., Surís, D., Miron, M., & Karatzoglou, A. (2018). Overcoming catastrophic forgetting with hard attention to the task. ArXiv: 1801.01423.

— Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2014). Intriguing properties of neural networks. In Proc. of the Int. Conf. on Learning Representations (ICLR).

— Zhang, C., Bengio, S., Hardt, M., Recht, B., & Vinyals, O. (2017). Understanding deep learning requires rethinking generalization. In Proc. of the Int. Conf. on Learning Representations (ICLR).

3 Whole brain dynamics and model

Gustavo Deco
Universitat Pompeu Fabra

3.1 Statement

Whole-brain computational models aim to balance between complexity and realism in order to describe the most important features of the brain in vivo. This balance is extremely difficult to achieve because of the astronomical number of neurons and the underspecified connectivity at the neural level. Thus, the most successful whole-brain computational models have taken their lead from statistical physics where it has been shown that macroscopic physical systems obey laws that are independent of their mesoscopic constituents. The emerging collective macroscopic behaviour of brain models has been shown to depend only weakly on individual neuron behaviour (Breakspear and Jirsa, 2007). Thus, these models typically use mesoscopic top-down approximations of brain complexity with dynamical networks of local brain area attractor networks. The simplest models use basic neural mass or mean-field models to capture changes in mean firing rate, while the most advanced models use a dynamic mean field model derived from a proper reduction of a detailed spiking neuron model [see (Cabral et al., 2017; Deco and Kringelbach, 2014) and references therein for a review].

The link between anatomical structure and functional dynamics, introduced more than a decade ago (Jirsa et al., 2002), is at the heart of whole-brain network models. Structural connectivity data on the millimetre scale can be obtained in vivo by diffusion weighted/tensor imaging (DWI/DTI) combined with probabilistic tractography. The global dynamics of the whole-brain model results from the mutual interactions of local node dynamics coupled through the underlying empirical anatomical structural connectivity matrix. The structural matrix denotes the density of fibres between a pair of cortical areas ascertained from DTI-based tractography. Typically, the temporal dynamics of local brain areas in these models is taken to be either asynchronous (spiking models or their respective mean-field reduction) or oscillatory (Deco and Kringelbach, 2014).

Adding the temporal dimension to standard FC analysis paves new ways to characterize the switching behavior of resting-state activity. However, the best methodology to assess it is still under debate. The most commonly used strategy has been to calculate successive FC(t) matrices using a sliding-window. Recurrent FC configurations are then captured by applying unsupervised clustering to all the FC(t)s obtained over time. However, the sliding-window approach has limitations associated to the window size, which affects the temporal resolution and statistical validation. Recently, new methods have been proposed to calculate the FC(t) at a quasi- instantaneous level, namely Phase Coherence Connectivity or Multiplication of Temporal Derivatives, which allow for a higher temporal resolution with the caveat of being more susceptible to high-frequency noise fluctuations. To overcome this issue, we hereby propose to focus on the dominant FC pattern captured by the leading eigenvector of BOLD phase coherence matrices. The key idea of this task is to focus on spatiotemporal dynamical biomarkers instead of the classical static grand averaged biomarkers (e.g. FC): In concrete, both during task and at-rest, we will identify whole brain dynamical micro brain states, by clustering the dominant dynamic functional connectivity (FC) patterns captured by the leading eigenvector of those matrices. Recurrent FC patterns – or micro states – will be detected and characterized in terms of lifetime, probability of occurrence and switching profiles in link with the subjects’ performance on the Battery of behavioral tests, evolution of disease, and recovery.

This type of whole-brain modeling could be used for crucial translational

applications. The basic idea here is to exhaustively stimulate off-line a realistic subject specific fitted whole-brain model in order to detect which type and locus of stimulation is

more effective to reestablish a healthy dynamic of the whole brain (both under resting and task conditions) in order to expect that under that condition Hebbian learning will cause a meaningful recovery.

Thus, multimodal neuroimaging (DTI, fMRI) is essential for having patient specific tailored whole-brain models which can be studied exhaustively. Whole-brain models could be fitted in particular by the novel spatio-temporal dynamical features mentioned above which characterize the network dynamics in probability microstates space. In parallel, based on healthy control groups, we can characterize also those same features.

The idea is to discover, which kind of external stimulation (type and locus) would promote a transition from the patient specific affected probability microstate space to a healthy one. This study can be done exhaustively in off-line simulations. After that, one can try in vivo, with TMS, if those reestablishment of healthy spatio-temporal dynamics causes recovery.

References

— Breakspear, M. and Jirsa, V. K. (2007) Neuronal dynamics and brain connectivity.. In: Handbook of brain connectivity. pp. 3-64. Eds. V. K. Jirsa, A. R. McIntosh. Springer: Berlin Heidelberg, New York.

— Cabral, J., Kringelbach, M. L. and Deco, G. (2017) Functional connectivity dynamically evolves on multiple time-scales over a static structural connectome: Models and mechanisms. Neuroimage, in press

— Deco, G. and Kringelbach, M. L. (2014) Great Expectations: Using Whole-Brain Computational Connectomics for Understanding Neuropsychiatric Disorders. Neuron84, 892-905.

— Jirsa, V. K., Jantzen, K. J., Fuchs, A. and Kelso, J. A. S. (2002) Spatiotemporal forward solution of the EEG and MEG using network modeling. Medical Imaging, IEEE Transactions on 21, 493-504.

4 Extended minds and machines

Karina Vold

Leverhulme Centre for the Future of Intelligence and Faculty of Philosophy, University of Cambridge

4.1 Statement

While there are have been many astonishing feats by AI in the last decade—even in just the last year—there continue to be many fundamental differences between human and machine intelligence. For one, many of the headline-making accomplishments by AI have been in highly specialized domains, or in what experts call Artificial Narrow Intelligence (ANI). Humans possess a more general kind of intelligence. We must, after all, perform a wide-range of tasks in order to successfully navigate our complex environments. Specialists are working towards building Artificial General Intelligence (AGI)—machines capable of skilled performance in a wide range activities. But there exist fundamental differences between human and machine intelligence that may complicate this project.

One difference that is often pointed to between humans and machines, especially amongst philosophers, is consciousness. There is nothing it is like to be a machine. Machines can be damaged, but they do not feel pain (Dehaene, Lau, and Kouider 2017). There is no uncontroversial scientific or philosophical theory of consciousness but, given the central role of phenomenology in human life and the success of our species, it seems at least prima facie plausible that consciousness has a cognitive function (for dissenting opinions see Chalmers 1996; Jackson 1982). In this case, without consciousness computers may never have human-like intelligence.

Another fundamental difference is that human intelligence has a long evolutionary history. The physical world has put many constraints on human intelligence: our brains need to be small enough to fit through a human birth canal and light enough to be carried around on our necks. Furthermore, our biochemical processing speeds are slow, running on less power than a refrigerator lightbulb. This might have been useful when we had to conserve enough energy to scavenge for food, to build shelters, and to procreate, but computers do not have to do any of these things. Machines already process information more quickly and efficiently than the human brain (Reardon 2018) and all of their resources can be expended on one, very narrow task—hence their success in specialized domains.

These differences raise questions about whether we could ever build humanlike intelligence, as well as why we should want to build machines that mimic our own intellectual constraints, especially if it is possible to bypass human intelligence entirely and leapfrog into what experts call ‘superintelligence’, which would surpass humans in many or all cognitive tasks.

Humans have done remarkably well considering the constraints on our biological bodies. An increasingly popular set of views in philosophy of mind and cognition maintains that we have achieved cognitive success by finding ways of moving our thinking outside of our bodies. We created language, for example, a complex representational system that enables us to communicate ideas and build on them over time. We also created technologies, from pens and paper to smartphones, which allow us to augment our biological capacities and simplify the cognitive tasks our brains need to complete (Clark and Chalmers 1998). This points to yet another difference between human intelligence and machine intelligence: our cognitive functions crucially depend on our bodies, our environments, and our tools. Our intelligence, as it is sometimes put, is embodied, embedded, and extended. To attain human-like intelligence, machines too might need to be embodied (perhaps even in human-like forms), embedded in their environments, and able to extend their cognitive functions beyond their hardware through seamless integration with tools.

One concern that falls out of this idea that machines may move beyond their intended, or original, hardware base to make use of other tools to complete their tasks is whether us humans will become their tools of choice. We have already seen instances where the outcomes of algorithms have played a role in shaping human decision-making, e.g. in the application of risk-assessment algorithms in parole decision (Kehl, Guo, and Kessler

2017). Ng (2016) argues that any mental task that a typical person can do in less than one second of thought can be automated, which suggests that the more difficult tasks to automate will be precisely those that require us to reflect. Tasks that we can do in less than one second do not require consciousness (Kahneman 2011). Human consciousness may need to play a key reflective role in complementing algorithms, but will have to do so while being guarded from the biases that influence our ‘fasting thinking’ systems.

4.2 Challenges

  • Does it make sense to draw comparisons between human intelligence and machine intelligence?
  • Can these comparisons mislead us and even harm us? How can they be used to help humanity?
  • How is interaction with machines affecting human intelligence and cognitive capacities? Is it augmenting and enhancing our capacities (Savulescu and Bostrom 2009), is it simply changing our capacities (Carr 2010), or is it perhaps diminishing them?
  • How can humans rely on the suggested outcomes of algorithms without being unduly influenced by them?
  • How can we prevent humans from being used, nudged, or manipulated by machines? And when, if ever, is relinquishing control to machines for the best?

References

— Carr, N. (2010). The Shallows: How the internet is changing the way we think, read and remember. London: Atlantic Books.

— Clark, A. and Chalmers, D. J. (1998). “The Extended Mind.” Analysis 58: 7-19.

— Chalmers, D.J. (1996). The Conscious Mind: In search of a fundamental theory. Oxford University Press.

— Dehaene, S., H. Lau, and Kouider, S. (2017). “What is consciousness, and could machines have it?” Science Vol. 358.6362: 486-492.

— Jackson, F. (1982). “Epiphenomenal Qualia”, Philosophical Quarterly, 32: 127-36.

— Kahneman, D. (2011). Thinking, fast and slow. New York: Farrar, Straus and Giroux.

— Kehl, D., Guo, P., and Kessler, S. (2017). “Algorithms in the Criminal Justice System: Assessing the Use of Risk Assessments in Sentencing” Responsive Communities Initiative, Berkman Klein Center for Internet & Society, Harvard Law School.

— Ng, A. (2016). What Artificial Intelligence Can and Can’t do. Harvard Business Review.

— Reardon, S. (2018). “Artificial neurons compute faster than the human brain,” Nature, January 26, 2018. <https://www.nature.com/articles/d41586-018-01290-0>

— Savulescu, J. and Bostrom, N. (Eds.) (2009). Human Enhancement. Oxford University Press.

5 Slow and fast biases in decision making

Ruben Moreno-Bote
Universitat Pompeu Fabra

5.1 State of the art

The brain is the only known intelligent system in the whole universe. It consists of 100 billion neurons working together in intricate circuits to generate complex behaviour that allows adaptation and survival of the species. The study of the brain is important for clinical aspects, but also because it can inform and inspire new waves of AI. Also, knowing how it works will permit smoother interactions between humans and future ‘intelligent’ technologies.

An example of the benefit of studying the brain in AI is Deep Learning, which originated from early inspirations of theoreticians of the brain: the inspiration was that artificial neuronal networks consisting of interconnected non-linear units could represent increasingly complex ‘abstract’ variables. Thus, it is expected that new inspirations for AI will come from how the brain works.

It is interesting to observe that algorithms for AI could be designed to work at high performance in specific problems, but do not necessarily generalize well to new domains or even new datasets. However, this is part the strength of algorithmic AI, as it is possible to design systems that work very well under well-defined and constrained conditions. This is beyond the scope of human brains, which are not designed to work in repetitive and highly constrained or restricted conditions. Also, search algorithms are more efficient in some respects than the human brain, especially when dealing with vast amounts of data in digital format, but not necessarily so in other formats.

We think that it is important to understand human and complex animals’ behaviour in situations in which there is no a priori reason to expect that human behaviour can be worse than the algorithms designed by AI. For instance, in perceptual decision-making tasks, over-trained animals are expected to perform the task very efficiently, as it is typically observed. However, we observe biases from previous trials that affect performance. The presence of these biases is informative about how artificial the task is, and how much it deviates from its natural setting in which it was designed to operate. By knowing these biases, one can better compare human vs machine performance in situations in which experimental conditions can deviate from naturalistic environments.

5.2 Challenges

  • Define tasks in which a priori there is no reason that subjects can do it wrong, and yet it is shown that performance is far from optimal, or biases are shown, or wrong tendencies are shown.
  • Complementary, design tasks that are naturalistic and in which human behaviour is highly adapted and close to optimal. Then test performance of AI algorithms in these naturalistic settings.
  • How biases can be eradicated from behaviour? Can machines/algorithms help to correct for this? Recent work shows that blocking of certain areas improves performance, as if the knob for biases could be removed.