Researching language learning: theories, evidence, claims
Epistemological assumptions and ILL research
'Governing laws'; the problem of generalizability
So we would agree fully with Cook's rejection of an 'unanswerable' question such as whether speech or writing is more important [in learners' progress towards an additional language. but we also sympathize with impatient practitioners or policy-makers who want to know where to invest their resources with a different group of learners working in different social and educational circumstances. Can the findings of a study based on a much more tightly specified research question such as the one Cook supplies be anything other than trivial? Are there any 'covering laws' or generalizable findings to be gleaned from ILL research? It is hard not to conclude that aggregated findings from studies of this kind seem not to produce very much consensus. This point is made baldly by Wardhaugh (1998) in an overview on the topic: 'Given all the thought and effort that has gone into so many years of such [SLA] research, one might well wonder ... why researchers still have such a poor understanding of how people learn a second language' (p.585). Other commentators are perhaps less damning, but recognition of divergent and often even contradictory findings is inescapable. In some commentaries, the problem is characterized as at least partially methodological, with inconsistencies represented as matters of technique, specificity or rigour, arising 'partly because of the lack of clear definitions and methods for the individual characteristics' (Lightbown and Spada 2001: 42). Chaudron (1988), in an overview of research conducted in 'second language classrooms', claims:
Since a number of the features reviewed here had conflicting findings across studies, and factors such as the identity of the speakers and listeners were not consistently controlled, it is evident that greater rigor and a well-defined research agenda are needed for future studies of LZ teacher talk. (p.89)
and:
... there has been little consistency throughout the classroom-oriented research in the choice of descriptors of task and activity types. Research in classrooms has been limited by not having an agreed-upon set of activity types ... , so little comparison was possible among studies.... Until there is greater uniformity, the research will be difficult to consolidate into immediate implications. (p.187)
More recent overviews of research in specific areas come to similar conclusions. In relation to research about input-based approaches to teaching grammar, Ellis (1999: 64, 73) states:
... it is becoming increasingly difficult to draw clear conclusions given the sheer amount of research now available, the problems of comparing results across studies, and the interactivity of the variables involved.... Design differences make it difficult to compare the results of these studies.
And in an overview of studies about the teaching of grammar in foreign language classrooms Mitchell (2000) concedes that'... it seems that we still lack a set of generally agreed principles, with clear empirical support, for the selection of grammar items which may merit explicit treatment in any "what works" program' (p.293); and 'applied linguists are not at present in a position to make firm research-based prescriptions about the detail of "what works" in FL grammar pedagogy' (p.296). Again, although mindful of the variations in contexts of teaching and learning, and therefore with somewhat more ambivalence than Chaudron displays, Mitchell calls for moves '... to increase agreement among researchers on what kinds of tests and resulting data will count as providing evidence of L2 learning' (pp.298-9). Other solutions often proposed are larger and/or longer studies (Wardhaugh 1998), more replication studies (Lightbown 2000), more detailed applications of sophisticated statistical procedures and so on.
In response to the inconclusiveness of findings, then, there are both calls for more studies and replications of existing ones, and depictions of the challenge of identifying and coping with large numbers of variables as largely technical matters. The summary by McDonough and McDonough (1997: 45) on this point questions whether all the factors involved in the areas to be researched can be taken into account: '[I]n most educational situations the list of possible confounding variables is so large, with some systematic and some unsystematic ones, that realistic and satisfactory control and counterbalance are nearly impossible'.
The issue goes deeper than this, however, and at this point we must introduce another objection to the importing of techniques developed principally for investigating inanimate objects into social domains. As social realists, we adhere to an essential distinction between people and things. This is expressed cogently by Hacking (1997: 15) (who has researched socio-medical problems such as schizophrenia), when he points out that:
while phenomena in the natural world are 'indifferent' to how they are labelled, human and social phenomena may be affected by the discussions of what they do, how they are labelled- they are thus of 'interactive' kind, and there is a 'looping' back into the object itself of how those involved understand it.
The 'variable' of a leaTIler's first language is much less stable than a variable such as temperature in a chemical experiment. This does not go unacknowledged in traditional approaches, as this observation by Long (1993: 235) illustrates: 'It will be important ... to be aware of the dangers inherent in importing criteria from the natural to the social sciences, since we are dealing with people, who can affect the systems and processes SLA theories seek to explain in ways physicists, for example, needed not worry about.' However, the implications are more far-reaching than this might suggest.
To the extent that 'subjects' have an active role in deciding how 'French' they wish to be, they may frustrate the quantifier's efforts to place them into the correct category, which points to a twofold problem. On the one hand is measurement of the language variety, and 'the tendency in SLA ... to reify language so that French, English and so on are treated unproblematically as homogenized "target languages" , (Roberts 2001: 109-10). On the other hand is social group membership. Rampton (2000: 101), for example, suggests that. in the light of developments in social theory 'whether it is age- or ethnicity-based, belonging to a group now seems a great deal less clear, less permanent and less omni-relevant than it did 15 years ago', and a similar point about age categories is made by Coupland (1997).
Discrepancies in how the researchers and the researched understand categories seem to have arisen in one of the studies cited above, when a proportion of the high-school students studied claimed greater proficiency in Spanish than was demonstrated in the tests they took, apparently because of their affective 'orientation towards maintenance of Spanish' (Hakuta and D'Andrea 1992: 77). This problem is reported as a methodological one: '... it appears that attitudinal orientation contaminates self-reported proficiency ... to a substantial degree' (p.95). The discourse in which such statements are made works to suggest similarities in the properties of the objects studied in the natural sciences, where one substance may be found to have 'contaminated' another in the laboratory, and the human subjects studied in social science, where people's values and aspirations 'contaminate' their assessment of their own language proficiency. Nevertheless, despite a discursively implied equation between human 'subjects' and physical objects ('[a] variable can be defined as an attribute of a person or of an object which "varies" from person to person or from object to object' (Hatch and Farhady 1982: 12, additional emphasis added), certain kinds of measurability turn out to be altogether more problematic- for what turn out to be sound theoretical reasons- than these methods textbooks would have us believe.
Critics of traditional ILL research have likewise problematized the way in which people are classified as Native Speakers and Nonnative Speakers, as though these too were given, absolute categories (e.g. Davies 1991; Firth and Wagner 1997; Rampton 1990), and similar problems inevitably recur in variables research which is designed in accordance with the dominant approach we have been describing. For example, in our own experience as supervisors of graduate students, we have had to respond to queries about whether a study of cross-cultural linguistic behavior (such as apologizing, for example) can include subjects who are 'English' on some criteria (such as nationality and place of birth), but seem to be untypical on others (such as skin color or residence outside the UK). Dilemmas such as these point to the danger of what Pawson (1989: 24) identifies as 'selective measurement due to the impregnation of observational categories by theoretical notions': in other words, the researcher has a preconceived concept, perhaps adequate for 'everyday' purposes (of 'Nonnative Speaker', or of 'Englishness'), but the process of scientific measurement is compromised by deploying this category as a variable without specifying the theoretical basis on which it rests.
Treating the world as though it comes in monadic, discrete, singular lumps omits consideration of the role of theory and conceptualization in the perceptual distinctions we make.... The reason why empiricist measurement is captured in such a critical loop is the lack of any notion of the role theory plays in understanding and justifying a particular regularity or model or law. (ibid.: 70-1)
Researchers with a more holistic orientation raise the possibility that this 'untying of the bundle' can never be achieved, although we believe that the full implications of this warrant further exploration. Lazaraton (1995: 465). for example, doubts whether the dominant approach to ILL research can ever generate the generalizations many see as desirable: 'Quantification of any set of data does not ensure generalizability to other contexts, nor does a large sample size.... In other words, generalizability is a serious problem in nearly all the research conducted in our field', while van Lier (1990: 38) claims that '[m]ost of our efforts at doing experiments or quasi-experiments, with all the attempted controls of variables and randomizations of treatment, may be doomed to failure (especially given the complexity of language learning processes)'.