Search This Blog

Showing posts with label Philosophy of Cognitive Science. Show all posts
Showing posts with label Philosophy of Cognitive Science. Show all posts

Monday, 17 April 2017

A Brief Philosophy of Biology Assignment where I Subtly Attack Singer and Review Recent Work on Human Altruism

Singer and the Evolution of Altruism

Biological Assumptions in Peter Singer’s Ethics and Intuitions:
In this 2005 paper, Peter Singer uses decades-old research in theoretical evolutionary biology and some recent results in cognitive psychology and neuroscience to argue that we should, in general, be highly suspicious of our intuitions when engaging in moral reasoning, since they will tend to lead us astray from the impartial application of abstract moral principles in favour of philosophically indefensible ‘gut’ reactions.  
In the first section of the paper, Singer argues that modern science backs up many of the bleak descriptions of ‘human nature’ contained in the works of Enlightenment philosophers like David Hume and Thomas Hobbes. Singer begins the section by praising the following conjecture of Hume’s: “A man naturally loves his children better than his nephews, his nephews better than his cousins, his cousins better than strangers, where every thing else is equal” [Hume, 1739/1896: 251; Singer: 334]. Singer takes for granted that this is an accurate observation, and gestures towards the Hamiltonian/Dawkinsian, gene-centric kin selection theory of altruism as the explanation for it. Next, Singer begins to make some vaguely-scientifically-tethered claims about prehistory: à la Hobbes, Singer asserts that early human life was “more often a struggle for survival between different human beings” than a struggle against other species [335]. This leads him onto a (once more) unsourced account of the most general evolutionary theory for the origin of extra-genetic altruism, reciprocal altruism; in particular, he claims that “Many features of human morality could have grown out of simple reciprocal practices such as the mutual removal of parasites from awkward places” [336].
Finally, he refers to the empirical work of Jonathan Haidt on moral rationalisation, and Joshua Greene’s very well-publicised studies on the Trolley Problem, along with some neuroscientific studies about the link between abnormalities in the prefrontal cortex and “anti-social behaviour”, to make clear that there is a strong emotional element to the way we typically make moral judgments, and no sense in which our intuitions are based on consistent application of principle [339]. 
Singer uses all these considerations from disparate areas of scientific inquiry for a methodological conclusion about moral philosophy: specifically, he opposes his stance against John Rawls’ (fairly middle-ground) methodological approach to moral reasoning, the principle of “reflective equilibrium” pioneered in A Theory of Justice. Reflective equilibrium is, in Singer’s words, the principle that “where there is no inherently plausible theory that perfectly matches our initial moral judgments, we should modify either the theory, or the judgments, until we have an equilibrium between the two[344]. Singer disagrees with this because he thinks that it is often most rational often to ignore entirely our pre-theoretic judgments (intuitions) when we encounter a moral dilemma. He thinks that we should concern ourselves only with applying the “plausible theory” (he more or less ignores the complication that a moral theory wouldn’t be plausible in any sense if it ran totally against our moral emotions). Crucially, he thinks the science lends a lot of support to this methodological conclusion, since it shows that our intuitions are not useful “data” when it comes to the process of rigorously reasoning about what we ought to do, and lead us instead towards ethical ‘parochialisms’ and inconsistencies.

Two Scientific Papers Examining Evolutionary (including Multi-Level Selection)_Theories of Co-operation:
1.)    “Human Co-operation” (2013), by David Rand and Martin Nowak.
This paper is a discursive review of laboratory experiments and field studies of human behaviour by a Yale cognitive scientist (Rand) and a Harvard evolutionary biologist and mathematician (Nowak). It focusses on how the empirical evidence bears on the relative strength of the five mechanisms proposed to explain the evolution of human co-operation: direct reciprocity, indirect reciprocity, spatial selection, multilevel selection, and kin selection. The discussion in this paper relates directly to the vague scientific references made by Singer towards theoretical evolutionary biology, and his apparent belief that kin selection and direct reciprocity are the only games in town.
In their very diplomatic treatment of the evidence supporting each of the five mechanisms, Rand and Nowak seem to suggest that there is good reason to suppose that all five mechanisms have played a role in the evolution of human co-operation. Their analysis therefore doesn’t seem to support more hegemonic views: for example, Hume’s strong claim about the disproportionate power of familial affection (at least if “love” is interpreted in the broad sense that Singer himself seems to want to interpret it). In fact, as Rand and Nowak discuss in the section on multilevel selection, it has been experimentally shown that unrelated strangers can be readily induced to co-operate effectively with each other simply if they are told they now ‘belong to a team’ in competition with another [Rand and Nowak, 419-420].
Singer, of course, is aware of research on the power of group-loyalty and tribalism, as evinced by his references to Jonathan Haidt. However, he seemed unware (in 2005) that this well-studied phenomenon actually constitutes very good evidence for a kind of (probably cultural) group selection towards within-group co-operation, thus undermining purely gene-centric views about the evolution of human co-operation (which tend to downplay the existence of widespread human co-operation).

2.)    The evolution of extreme cooperation via shared dysphoric experiences” (2017), by Harvey Whitehouse et al (12 contributors).
This paper is a recent interdisciplinary research article in Nature in which the authors lay out a new mathematical model which they claim shows how “conditioning cooperation on previous shared experience can allow individually costly pro-group behavior to evolve”, with evidence provided in the form of the testing of the predictions of the model in a range of sample populations (including military veterans, college fraternity/sorority members, football fans, martial arts practitioners, and twins) [1]. Overall, they obtain strong support for the conclusion “that shared dysphoric experiences are a powerful mechanism for promoting pro-group behaviors which under certain conditions can be extremely costly to the individuals concerned” [6].
This finding doesn’t directly undermine Singer’s biological assumptions in any direct way, although the explanadum of the paper’s thesis (that humans sometimes engage in pro-group behaviours which are “extremely costly to the individuals concerned” (like suicide terrorism and fighting for ‘king and country’ in deadly conflicts [1])) is something that Singer’s human-nature-related assertions in Ethics and Intuitions completely miss out.  Furthermore, one subsidiary empirical conclusion of the article, used as evidence against the notion that self-sacrifice for the group can be fully explained by the so-called “psychological kin” phenomenon, does directly undermine one of Singer’s claims: namely, the Humean one.  As the authors point out early in the piece, a recent survey of participants in the Libyan uprising of 2011, thousands of whom died in combat, found that “frontline fighters were more likely to choose genetically unrelated fellow revolutionaries in preference to family as the group with which they are most fused” [2].

Singer would probably argue that, though these new results do impugn some parts of his generally vague scientific discussion in Ethics and Intuitions, they do not undermine the view that our moral “intuitions” can steer us away from a more rational and universalised morality. Indeed, the idea that powerful, large-scale co-operation tends to be group-based and highly context-dependent, typically reliant on extreme events and highly dependent on the existence of enemy groups, only lends support to the general thesis that we should try our best to transcend our more immediate moral emotions in favour of the application of general principles. On the other hand, it may be that Singer’s own distorted cognitive science has misled him as to the real-world plausibility of a co-operative and altruistic ethic which doesn’t rely psychologically on some kind of group identification. Singer may well overlook the extent to which even non-parochial moral systems rely on the commandeering of group-based moral sentiments; for example, achieving a successful ‘expansion of the circle’ may require the stigmatisation of those who don’t.







Bibliography
Rand, D.G. and Nowak, M.A. “Human cooperation”, Trends in Cognitive Sciences, Vol. 17: 8, August 2013, pp. 413-425.
[Google citations: 348; Journal Citation Reports lists its 2013 Impact Factor at 21.147.]

Singer, P.J. “Ethics and Intuitions”, The Journal of Ethics, Vol. 9: 3, October 2005, pp. 331-352.
[Google citations: 460; Impact Factor of journal: unavailable].

Whitehouse, H, et al. “The evolution of extreme cooperation via shared dysphoric experiences”, Nature: Scientific Reports, 7, Article number: 44292, February 2017.
[Google citations: 1; Impact Factor of journal: 5.228.]


Friday, 28 October 2016

AN ESSAY ON THE MYSTERIES OF THE MIND

Take Home Exercise for Philosophy of Mind and Cognition
(b)Critically evaluate the arguments for the Language of Thought and/or the Map Theory

In this essay, I examine the arguments for the several stances it is possible to take in the ‘debate’ (largely implicit) between the Language of Thought Hypothesis and the Map Theory of cognition. My ultimate conclusion is that, whilst so many fundamental issues are still extraordinarily unclear, there is one particular intermediate stance in this debate that has significantly more weight behind it than any other.
I first discuss how, for both hard-line connectionists (those who think that the brain undergoes no ‘classical computation’) and hard-line Fodorians (those who think the key parts of cognition will be explained only by classical computational models), the stakes between the “Language of Thought Hypothesis” and “Map Theory” of cognition are quite clear: if the LoTH is the right theory of all cognition, then the hard-line connectionists are completely wrong, and if complex cognition occurs without any real LoT, then the adherents of Fodor’s computational-representational theory are completely wrong. I argue that both these poles are probably wrong, but that the connectionist extreme is much less implausible than the Fodorian one. I secondly explore the intricate intermediate position that that all (or nearly all) cognition in non-human animals involves connectionist ‘software’, captured by “Map Theory”, and yet that the Language of Thought theory represents a classical ‘program’ run only by humans. This thesis has not been actually expounded in any literature I am aware of, although it seems to me to be the hypothesis that has the most weight behind it of all. As I argue, it is highly concordant with Noam Chomsky’s carefully considered hypothesis (since the mid-1990s) for the origin of the human language faculty/explanation for the “Great Leap Forward”; the notion that productive thought is bound up with productive language is supported by evidence from developmental psychology; the notion that the vast majority of biological ‘software’ is connectionist is strengthened by well-known considerations about the ‘hardware’ of the brain and the nature of evolution; and finally, despite Fodor’s many protestations, it seems to me that Prototype Theory is the best account of the nature of our concepts, and Prototype-structured concepts are better explained in terms of “Map Theory” than the “Language of Thought”.

In The Philosophy of Mind and Cognition, Frank Jackson and David Braddon-Mitchell are careful not to hitch the Map Theory to the connectionist programme in AI, or the Language of Thought to more ‘classical models of cognition’. They note that the truth is much more complex and that it may often be unclear whether a ‘classical’ program is being implemented in a connectionist substrate (indeed, they claim that there might be no “source code” for the brain at all) [Braddon-Mitchell & Jackson, 2007: 219]. Nevertheless, they do suggest a certain harmony between connectionism and the “Map Theory”, and between classical computation and Fodor’s LoT. This makes perfect sense, because a hard-line connectionist cannot possibly accept the LoTH, and the Map Theory is explicitly proposed as the one alternative to the LoTH.
Paul and Patricia Churchland, the influential reductive neuro-computationalists, are probably the most prominent proponents of what I’ve called ‘hard-line’ connectionism. For several decades, the central philosophical claim of these two neuro-philosophers has been that classical cognitive science is (and always has been) on the wrong track, because the brain’s hardware is connectionist[1] and we can only understand the brain’s ‘software’ by investigating the nature of the hardware, rather than constructing abstract, simple models of rules and representations which may have no “psychological reality”. In numerous articles and papers from the late 80s onwards, the Churchlands have argued that evidence from neurobiology[2] makes it clear that the Language of Thought (which they tend to lump together with ‘Folk Psychology’ as part of the one Fodorian package), even if apparently explanatory, simply doesn’t exist [P.S. Churchland, 1986; P.M. Churchland, 1989; P.S. & P.M. Churchland, 1990; P.S. Churchland & Sejnowski, 1990].
It is crucial to point out, of course, that there was always a key weakness in the Churchlands’ critiques of classical cognitive science and the LoTH – a weakness that would ultimately allow Fodor to evade their attacks without working up too much of a sweat at all. This key weakness was that the Churchlands didn’t attempt to provide a serious alternative, connectionist-based theory of thought (as opposed to mere detailed accounts of bio-chemical processes correlated to certain kinds of cognition) to the one they so roundly rejected. The absence of such a theory in the Churchlandian canon meant that it remained justified for Fodor to simply repeat the famous remark he made in his seminal 1975 work: that the LoTH is “the only game in town” [1975: 406]. And as Fodor and Pylyshyn pointed out in their influential 1988 critique of connectionism (which includes extensive reference to the Churchlands), without such an alternative account of the actual nature of thought, the neurobiological evidence marshalled in supposed contradiction of the LoTH cannot nullify the abductive argument, because the neurobiological details do not in themselves itself constitute an explanatory science of the mind [Fodor & Pylyshyn, 1988]. As Fodor and Pylyshyn emphasise in this same article, no connectionist can explain productivity and systematicity except by creating “Classical architecture” in connectionist models, yet these are the two key features of thought that form the basis of Fodor’s argument to begin with [Fodor & Pylyshyn, 1988: 33-40].
Fortunately for the connectionists, however, Frank Jackson and David Braddon-Mitchell’s “Map Theory” does provide such an alternative account of the ultimate nature of thought. Evidently, if such a theory is at all tenable, it lends significant credence to the hard-line connectionists, because the very fact of its existence destroys Fodor’s key abductive argument. And unfortunately for Fodor, it does indeed seem to me that the Map Theory is plausible as an alternative theory of thought – in fact, highly plausible. The idea of thought as involving highly structured, non-sentential, ‘nth’-dimensional map-like representations actually possesses a number of virtues over the LoTH: it is (at least intuitively) much more neatly reconcilable with connectionist networks and the ‘hardware’ of the brain, since a map is an inherently ‘distributed’ structure; it gives a far more plausible account than the LoTH of how it is that our beliefs about people, places and events can so fluidly change as we have new experiences and take in new information (our mental ‘maps’ are simply updated); it seems to explain memory and memory-retrieval in a much more plausible way than a Language of Thought, since it is not necessary to claim that we have a vast number of propositions stored in our memory which we can retrieve at will; it would appear to explain certain weaknesses we have in formal reasoning (for example, why content seems to matter to our ability to carry out tasks of reasoning with formally identical structure [Jackson, Braddon-Mitchell, 2007: 235]); it appears to have evolutionary considerations in its favour over the LoTH, since a Language of Thought is a highly elegant computational system which seems at odds with the haphazard makeshift nature of evolution (it seems unlikely that a LoT would be the kind of system to slowly crystallise, and it seems highly unlikely that our ancient Cambrian ancestors had a Language of Thought, whereas they might have had primitive ‘map-like’ representations); and it can account (seemingly) for “systematicity”, since the ‘places’ in the cognitive maps can be switched. All in all, therefore, the Churchlandian position is, in my view, massively strengthened by the “Map Theory”.
With that said, of course, the reason that I do not think that the hard-line connectionist position is likely to be right is that the Map Theory, even if it does seem to account for systematicity, it doesn’t quite answer Fodor’s productivity criterion. As a result, it cannot explain the apparent infinite generative capacity of human cognition – our ability to think an infinite array of discrete thoughts by finite (combinatorial) means. As I am about to suggest, however, it may very well be that only humans have this full-strength productivity. This would imply that the Map Theory might be a far more general account of biological cognition, and that the hard-line connectionist position might be the correct one for all creatures on Earth except us. As I’m about to argue, I think this is actually the most plausible stance of all.
Perhaps the major consideration in support of the view that only Homo sapiens has a LoT in Fodor’s sense (with a combinatorial syntax which allows for infinite productivity) is that it accords with Chomsky’s hypothesis about the evolution of the language faculty. Whilst undoubtedly a highly controversial hypothesis – disputed even by other notable generativists (for example, Ray Jackendoff and Steven Pinker [2005]) – the conjecture does have nontrivial theoretical and empirical backing.
Chomsky largely eschewed speculation about the evolution of the language faculty for the first few decades of the Universal Grammar research programme, but he has become quite vocal in espousing his ‘spandrelist’ hypothesis about the evolution of the “Faculty of Language” ever since he wrote The Minimalist Program in 1995. The central idea of Chomsky’s “Minimalist Program” is that it is possible to distil the ‘Universal Grammar’ into one basic computational operation called “Merge”, with two forms, “External” (for separate objects A and B) and “Internal” (for two objects where at least one object contains the other) [Chomsky, 1995]. This lends credence to this spandrelist hypothesis because of the elegant computational simplicity of such a system (“like a snowflake”), and the fact that this computational procedure subserves thought as fundamentally (or more fundamentally) as it subserves language.
The most rigorous defence of this evolutionary hypothesis appeared only recently, in a 2015 book Chomsky co-authored with the MIT computer scientist, Robert Berwick, entitled Why Only Us? : Language and Evolution. In this short work, Chomsky and Berwick specifically defend the thesis that this “Merge” operation first originated by a single mutation in a human individual around 80,000-100,000 years ago, which gave rise to recursion and hierarchical structure in that individual’s thoughts [Chomsky & Berwick, 2015]. Whilst they don’t invoke Fodor explicitly, it is clear that they see this as the moment the “productivity” property key to Fodor’s LoT actually emerged (for the first time in the history of life on the planet). They hold that this capacity would have had some selectional advantage, and would have spread throughout the population before being secondarily externalised [Chomsky & Berwick, 2015]. Much of the book is taken up by arguments in defence of an anti-adaptationist, ‘messy’ view of evolution as involving “stochastic effects” and both gradual and sudden developments, with more at work than simple natural selection in simple genetically varied populations. However, there are also a number of positive considerations adduced in favour of the hypothesis, many of them (in my view) highly compelling.
Firstly, recursion appears to be a property that is either absolutely present or absolutely absent, so seemingly couldn’t come about as a gradual adaptation [72]. Secondly, there is strong evidence (disputed by some) that no other species on the planet is capable of recursive communication – birdsong, they claim, never gets further than “linear chunking” or iteration of “motifs” [142]. Thirdly, the fact that human linguistic externalisation is modality-independent (sign language has the same level of syntactic complexity as spoken language, and is learnt as rapidly by children in the right environment) seems to constitute good evidence that the Faculty of Language was first a Faculty of Thought [74]. Fourthly, the sudden emergence of productivity around 80,000 years ago appears to be one of the best possible explanations for what Jarred Diamond famously called the “Great Leap Forward” [37]. Fifthly, the extent of language variation in the world can seemingly explained by the fact that the means of externalisation did not actually co-evolve with the core computational system, but are instead much more ancient systems being co-opted for the task [82]. Sixthly, the evolutionary timeframe of the change is certainly not ruled out by the reconstructed genomic evidence used to compare Homo sapiens with its ancestors and cousins [Chapter 4]. And, finally, there is a neuro-anatomical hypothesis for what actually happened to create the key “Merge” mutation – the dorsal and ventral “fiber tracts” linking different language-related areas of the brain formed a ring (which does not exist in birds) [Chapter 3].
There is also evidence from developmental psychology that lends support to the idea that language is bound up with a special form of thought that marks humans out as unique (and this naturally holds even if Chomsky and Berwick’s hypothesis is a fair way off the mark in terms of timeframe and suddenness of evolution). As Antoni Gomila notes in a critical review of Jerry Fodor’s 2008 book The Language of Thought Revisited, “there is no evidence for the systematicity and productivity of thought before the development of language, at the end of the second year” [Gomila, 2011: 151]. Gomila cites the child psychologist Elizabeth Spelke for this claim, whose experiments suggest that babies are born with modular packets of knowledge of particular contents (physical, numerical, biological, intentional), but that such knowledge does not admit of productive combination before language – instead it is context-dependent and encapsulated [Spelke, 2003]. It seems that the role of language in development is to create interaction between these different modules.
Another major virtue of this idea that the LoT is human-specific is that it means we can still maintain the benefits of a largely connectionist view of biological cognition (those benefits espoused by the Churchlands, in particular, general biological plausibility (in terms of physiology and evolution)), and the advantages of Map Theory as a general account of biological cognition. It only means that we have to say that humans may be the only species with a ‘classical program’.
Finally, adopting the Map Theory as an account of the majority of human thought also seems to work far better for explaining the fuzzy nature of human concepts, described by “Prototype Theory” in cognitive linguistics – the research programme pioneered by Eleanor Rosch’s studies about ‘categories’ in the 1970s. Map Theory would seem to suggest that concepts are not discretely structured, but form part of wider cognitive architectures. If one accepts the tight connection between connectionist software and Map Theory, then there is certainly reason to think that Map Theory is highly compatible with Prototype Theory since, as Gomila notes, connectionist models which are based on Prototype Theory have met with increasing success since the 1990s [2011: 149].
   
Overall, I think that the debate between Fodor’s Language of Thought Hypothesis and the “Map Theory” of cognition raise deep and interesting questions about the nature of the mind which are not currently soluble – although some possibilities seem distinctly more likely than others. I think that the LoTH is extremely unlikely to be a true theory of all cognition, and thus that the ‘hard-line’ Fodorians are seriously misled. I think that the Map Theory, most naturally joined with connectionist models, is far less implausible as a general account of the cognitive architecture of biological creatures – however I think it fails to account for productivity in humans. These considerations have ultimately led me to the view that the most likely hypothesis is that Homo sapiens is the only species to possess a Language of Thought, implemented as a ‘classical program’ on top of the connectionist hardware and software that all biological organisms probably possess.
Naturally, I acknowledge that this claim, despite the justifications I have given for it, is still an extremely speculative one. I am also aware that I have an aesthetic bias for this view: I think it is a beautiful notion that humans might be the only creature with a classical program that was key for our success in colonising the planet and reaching civilisation… So it might be totally wrong.





























Reference List

Braddon-Mitchell, D and Jackson, F. (2007), The Philosophy of Mind and Cognition, 2nd edition, Blackwell Publishing.

Chomsky, N. (1995), The Minimalist Program, The MIT Press, Cambridge, M.A.

Chomsky, N & Berwick, R. (2015), Why Only Us? : Language and Evolution, The MIT Press, Cambridge, M.A.

Churchland, P.M. (1989), A Neurocomputational Perspective: The Nature of Mind and the Structure of Science, The MIT Press, Cambridge, M.A.

Churchland, P.M. and P.S. (1990), Could a machine think? Scientific American 262 (1):32-37.

Churchland, P.S.  (1986), Neurophilosophy: Towards a Unified Science of Mind-Brain, The MIT Press, Cambridge, M.A.

Churchland, P.S and Sejnowski, T. (1990), “Neural representation and neural computation”, Philosophical Perspectives 4:343-382.

Fodor, J. (1975), The Language of Thought, Harvard University Press, Cambridge, M.A.

Fodor, J and Pylyshyn Z. (1988), “Connectionism and Cognitive Architecture: A Critical Analysis”, Cognition 28 (1-2):3-71.

Gomila, A. (2011), “The Language of Thought: Still a Game in Town?” Teorema: Revista Internacional De Filosofía, 30(1), 145-155.

Jackendoff, P and Pinker, S. (2005), “What's special about the human language faculty?” Cognition 95 (2).

Spelke, E. (2003), “What Makes Us Smart? Core Knowledge and Natural Language”, en Gentner, D. & Goldin-Meadow, D. (eds.): Language in Mind. Advances in the Study of Language and Thought, The MIT Press, Cambridge, M.A., 277-312.




[1] The most general and simplistic reason for thinking this is that sensory neurons can be easily analogised to the input nodes in a connectionist model, output nodes can be analogised to the motor neurons and the hidden nodes can be analogised to the web of neural connections in our nervous systems [Braddon-Mitchell & Jackson, 2007: 223])
[2] The most straightforward data included the “distributed representation” used by the brain (held to be strong counter-evidence of the existence of discrete symbols) and the fact of “graceful degradation” (certainly counter-evidence of digital hardware).

Friday, 21 October 2016

Just an Essay Justifying "Universal Grammar" in a non-technical way

21. How successful, in your view, is Chomsky’s attempt to interpret the theory of grammar as an investigation of a human biological capacity?

The research programme that Noam Chomsky began in linguistics, starting with the publication of his monograph, Syntactic Structures, in 1957, and properly expounded in his 1965 work Aspects of the Theory of Syntax, is not very well understood by the majority of its popular and even academic critics. Chomsky’s frequent failure to hedge his claims about the postulated “language organ” (and his penchant for overly strong pronouncements) has contributed to this failure of understanding, since few critics bother to understand how he explicates these terms in his methodological framework or where they sit in his metaphysics of mind. Of course, at the same time, he has himself ignored some issues in biology that ought to have compelled him to slightly change his tune on “innateness”. Nevertheless, in this essay, I claim that, if one does properly understand the methodological and philosophical foundations of the generative-grammar-based research programme into the “Universal Grammar”, two things are evident: 1.) that the programme has been about as successful as its progenitors hoped, if not more (with no degeneration, in Lakatos’ terms), and 2.) that the programme has been successful in the independent sense that it has led to genuine insights about the human mind.

One of the major things that critics of ‘UG’ are wont to gloss over is that, strictly speaking, there is no one “theory of Universal Grammar”. Instead, there is only a general ‘UG’-research programme based on the construction of generative grammars held to capture the innate implicit knowledge which is falsifiably (and each model ought to be falsifiable individually) held to be required for the achievement of human linguistic competence (the generative grammar in the human mind).[1] As this implies, the empirical support is (and could not be) uniform for ‘UG’, understood as the UG-research programme; instead, because the approaches to generative grammar that have been developed as part of this overall research programme since Chomsky’s exposition of a Transformational Grammar approach in Aspects to the Theory of Syntax (1965) vary in the level of ‘innate’ ‘knowledge’ they postulate, different evidence is required to falsify each of these approaches, and to confirm any one of them relative to the others. In particular, a stronger version of the famous “Poverty of the Stimulus” argument – one which draws on more evidence of linguistic universals, infant reliance on rules and hard-to-empirically-explain examples of “structure-dependence” –  are needed to support the early Transformational Grammar models than for the models proposed since the Principle and Parameters paradigm shift, with the Minimalist Program starting from the postulation that the generative system is simpler than a grammar (per se), and therefore probably being more prone to rejection by evidence of too many universals and such (this would, of course, be rejection relative to another model of generative grammar).
It is very important to acknowledge, of course, that there are some very sensible philosophical critics of the UG research-programme, like the Australian-born philosopher of biology, Fiona Cowie [1999], who see Chomsky’s very use of the word “innate” as problematic from a scientific viewpoint. One of Cowie’s motivations for her, in the end, quite moderate critique of “nativist” theories in What’s Within? is the fact that the word has no clear scientific interpretation. She shares this view with her fellow philosopher of biology, Paul Griffiths (of the University of Sydney), who in his 2002 paper “What is innateness?” points out that the word “innate” is no longer used in any field of biology except cognitive science, and argues that it ought to be discarded in cognitive science, too, since innateness is a “folk-biological” concept which yokes together three biologically separate notions: species-typicality, developmental fixity and intended design [Griffiths, 2002: 2].
Whilst I myself agree with this critique, I do not think it is in the least bit destructive of Chomsky’s research-programme, because I also think (and Griffiths does not explicitly deny) that the generativist talk of an “innate language faculty” can simply be substituted for one of two other phrases: a “developmentally canalised, species-typical language faculty” (in the case of Chomsky and his fellow non-adaptationists (spandrelists) about the evolution of the language faculty) or “an adaptive, environmentally canalised, species-typical language faculty” (in the case of Pinker, Jackendoff and the other adaptationists about the evolution of the language faculty).[2] One small criticism I have of Chomsky (along with fellow generativists like Charles Yang, Robert Berwick, Steven Pinker, Ray Jackendoff, etc) is that he hasn’t made this terminological alteration himself, still using the language of Cartesian or Humboldtian rationalism, of which he has always regarded his work a direct descendant [Chomsky, 1965, 1986]. However, since I don’t think that the use of this unscientific language is a significant problem for the UG-research programme, I will myself, in the rest of this essay, always type the folk term in single inverted commas while still assuming that I am defending Chomsky.
Past this terminological hurdle, the single biggest reason why I think that Chomsky’s research programme has been a success in both the senses I outlined in my introduction is empirical validation. It seems to me quite clear that the theory that Homo sapiens does have a developmentally canalised, species-typical language faculty has not been falsified over the decades, but is instead clearly still the best (indeed, the only) explanation for the empirical evidence of rapid language acquisition and human linguistic competence.
Despite the impression evoked by some critics, Chomsky’s generative research-programme did not begin as a scholastic, a priori enterprise, but was motivated by empirical considerations of a fairly fundamental kind. The reason that Chomsky’s famous 1957 monograph, Syntactic Structures, is often heralded as the founding document of modern cognitive science, despite not framing itself as a work of mentalistic investigation (and containing no argument for the existence of an ‘innate’ “language faculty”) is that Chomsky’s formal conclusions in SS directly entail the powerful scientific conclusion that at least one language (English) cannot be understood in behaviourist terms, and that English speakers must ‘possess’ (in some vague sense, leaving representational and acquisitional issues aside) what Chomsky would later call, in Aspects of the Theory of Syntax, “a system of generative processes” [1965: 4]. The most pivotal part of Chomsky’s monograph is his chapter 3 proof that a finite-state or Markov model (a model which involves the mono-directional chaining of words) is simply inadequate to generate all the grammatical sentences of English, and the implication, explored in the next two chapters (“Phrase Structure Grammar” and “Limitations of Phrase Structure Description”, in which he introduces the notion of a “transformation” to complement the inadequate phrase structure grammar), that an adequate generative grammar must be a hierarchical or syntax-based grammar of some kind.[3] Although the inadequacy of the finite-state model perhaps should have been obvious, it was no trifling result, because it refutes the strongest empiricist view about language production. If it is impossible to formally generate English sentences by a finite-state model, then it is also impossible to generate English sentences by the kind of finite-state model one might want to implement in a computer, or one might imagine existing in a brain: a complicated word-chain device relying on transitional probabilities. One instead needs generative principles, and this fact literally precludes the strongest empiricist understanding of language production.[4]
Evidently, behaviourists wanted to avoid talking about mentation at all, but the theory of language production found within works like Skinner’s Verbal Behavior (the subject of Chomsky’s famously destructive 1959 review) clearly rules out the possibility that human language is produced by a generative grammar. As Steven Pinker writes in The Language Instinct, the finite-state model is directly “congenial to stimulus-response theories: a stimulus elicits a spoken word as a response, then the speaker perceives his or her own response, which serves as the next stimulus, eliciting one out of several words as the next response, and so on” [1994: 93].
Of course, in Syntactic Structures itself, Chomsky uses his chapter 3 proof to motivate a more purely methodological conclusion: that linguistics ought to shift from static, structural description of corpora – the behaviourist methodology of American Structuralism[5] – towards the construction of the kind of descriptively adequate generative grammar he attempts to construct for English in SS (a transformational phrase structure grammar). Yet it is not hard to see how this led to the far bolder scientific shift expounded in Aspects of the Theory of Syntax. In order to make the case for an entirely new vision of linguistics in Aspects, Chomsky extended his insight about the necessity of a creative model for English sentence-production into all languages (on the strongly empirically based contention that no other human language can be adequately described by a finite state model either, supplemented by the empirically based assumption of strong human cognitive universality), and combined this with an argument in support of (essentially) 18th Century Rationalism. His new claim was that the aim of the discipline should be to construct generative grammars that are descriptively adequate for all natural languages, and explanatorily adequate as “universal grammars”. That is to say, linguists should attempt to describe the ‘innate’ linguistic competence or “system of generative processes” held to be necessary for the acquisition of any natural language by a child [1965: 6].
At this point, I should point out that, whilst I am claiming that Chomsky’s research-programme was empirically motivated at its inception, I am not claiming that all its empirical presuppositions were thoroughly confirmed in 1965. It clearly might have been the case that the evidence collected after 1965 strongly disconfirmed these empirical presuppositions, and thus rendered the research programme untenable. If it had turned out that there were significant cognitive group-differences in Homo sapiens – that some populations don’t have (and couldn’t have) language with any recursion (as Daniel Everett claimed to show, falsely[6])then that would be a kind of falsification of the universalist aspect of the programme (although, presumably, that wouldn’t sound the death-knell for the investigation of the linguistic competence of the human populations with the developmentally canalised language faculty). Similarly, if “Nim Chimpsky” (or, for that matter, some other animal) had proven capable of learning actual grammar, rather than mere sign-strings, then that would have been a blow to Chomsky’s pretty significant working assumption that the language faculty is unique to humans (and it would probably have forced him to attribute to the ‘language faculty’ a lower degree of developmental canalisation).
More interestingly, it might have been the case that the evidence gathered after 1965 strongly favoured the hypothesis that language acquisition occurs by means of solely domain-general cognitive processes, as Michael Tomasello’s “usage-based theory of language” holds (a theory of language which, despite its domain-generality, avoids getting wrecked on Chomsky’s SS proof about the inadequacy of Markov word-chains by positing that speakers learn entire grammatical constructions by means of a powerful “theory of mind” and contextual-awareness) [Tomasello, 2003]. If it had turned out that the condition known as “Specific Language Impairment” was in all cases actually just a misdiagnosed general cognitive impairment, or that all fully articulate older children diagnosed with autism or an autism-spectrum disorder have either been misdiagnosed or had full cognitive empathy in the critical period for acquisition, then that would lend strong support to a Tomasello-type theory over one which postulates a language faculty.
As it stands, however, things have not turned out this way. Thus, the UG-research programme has been completely justified in continuing to exist and grow.

 In more recent years, some subtler empirical objections to the UG research-programme have come from the AI community, which has had far more success with trained statistical models (for example, probabilistic context-free grammars) than categorical models (any of the ‘pure’ models created by generativists). The schism between these two worlds – one practical, one theoretical – came to the fore in 2011 after Chomsky made some highly disdainful comments about statistical approaches to “various linguistic problems” (accusing statistical modellers of doing ‘butterfly-collecting’ rather than “science”) at the Brains, Machines and Minds symposium held on MIT’s 150th Anniversary. Soon after, the Director of Research at Google, Peter Norvig, published an essay on Chomsky online in which he argued that the father of linguistics was entirely in the wrong, since probabilistic models have been far more successful in actual implementations than any categorical ones [Norvig, 2011]. Nevertheless, as Chomsky’s student Charles Yang has argued in response to similar critiques, whilst there is now “a good deal of evidence against the ‘triggering’ model of learning” (which, Yang happily concedes, deserves to be replaced by a “probabilistic model”, which is domain-general) “one needn’t, and shouldn’t, abandon the categorical theory of GRAMMAR”, which is domain-specific [Yang, 2007: 215]. In other words, it is still perfectly cogent to study the abstract generative processes, because, as Chomsky has always claimed, they represent the underlying linguistic competence and nothing more.

 There is one kind of critique of Chomsky’s research-programme which has nothing to do with the empirics of ‘UG’, but the metaphysics. Towards the end of her critique of ‘UG’ in What’s Within? Fiona Cowie brings up an objection of exactly this metaphysical kind. Tapping into a broader philosophical doubt many hard-line connectionists and ‘Churchlandian’ eliminativists have about the computational theory of mind in general, Cowie claims that it is deeply problematic that Chomsky cannot specify what the implicit “knowledge of language” [7] he theorises about actually is [Cowie, 1999: 274]. In particular, it is unclear, she claims, in what sense grammar actually could be “represented” [1999: 274].
Despite the seeming importance of this line of objection, Chomsky is simply a deflationist about this question of “representation”, and I think he is right to be so. The reality is that there is no possible theory of linguistic-competence – no possible explanatory scientific theory of language – other than one which posits ‘knowledge’ in the form of a system of “rules and representations” [Chomsky, 1980]. This means that Chomsky’s UG research-programme is the only possible scientific research programme into the faculty of language. The fact that, as Chomsky himself says in a reply to the philosopher Georges Rey (quoting Randy Gallistel) “we clearly do not understand how the nervous system computes,” or even “the foundations of its ability to compute,” should not put a halt to the only scientific investigation into the human language faculty, just as it shouldn’t put a halt to any other cognitive scientific studies (including in other species) [Chomsky in Chomsky and his Critics, 2003: 276]. As Chomsky says in that same reply, “surely no one expects that some isolable part of the organism is dedicated to digestion, or navigation, or language, or any other component that is singled out for investigation in any rational approach to the study of a complex system” [Chomsky, 2003: 276].

In summary, I believe that Chomsky’s attempt to interpret the theory of grammar as an investigation of a human biological capacity has been highly successful, according to any reasonable criteria for such things. The Universal Grammar-research programme he began in the early 1960s has enjoyed significant internal development while its core presuppositions have been strongly confirmed. This has, in turn, revealed to us important insights about the human mind.




[1] Of course, the “theory of Universal Grammar” might then be understood as the empirical presuppositions necessary for the tenability of this programme in general. However, across the history of the UG-research programme, it seems to me that the only constant empirical presuppositions are: 1.) human cognitive universality (no significant cognitive group-differences in Homo sapiens), and 2.) the existence, in all humans without severe impairment, of a developmentally canalised, domain-specific, computational ‘module’ (in the vaguest possible sense) which explains human “linguistic competence” (and possibly several other human abilities) for which we can construct generative models (not even grammars per se) by investigating the syntax of the world’s languages. It seems to me that many critics who mount general attacks on “Universal Grammar” believe they are attacking a stronger thesis than the conjunction of these two propositions.
[2] Strictly speaking, it would be best to rephrase Chomsky’s usage of “innate language faculty” with “a developmentally canalised, species-typical language faculty which originated as a spandrel but proved adaptive and was then selected for” [Chomsky, 2012: 14].  
[3] He doesn’t use the word “hierarchical” in SS.
[4] Of course, this in itself shows nothing about the ‘innateness’ (more properly, developmental canalization, etc) of the generative processes necessary for language production in an adult speaker. In itself, it also clearly doesn’t rule out the possibility of probabilistic language models more sophisticated than Markov chains (i.e. probabilistic generative grammars), or, arguably, Michael Tomasello’s usage-based theory of language [2003], which I’ll discuss later.  
[5] Of which one of the major proponents was Chomsky’s teacher and mentor, Zellig Harris.
[6] Pirahã does have recursion, and (more fundamentally) Piraha speakers can learn Portuguese, so Everett’s ‘argument’, such as it is, poses no problem for the research programme at all [Nevins, Pesetsky, Rodrigues: 2009].
[7] Which he has tried to re-term the “cognizance of language” in an attempt to stop the philosophical controversy.