Social categories and theoretical descriptions
A realist approach to theoretical descriptions
The epistemology we are advocating offers one way of defending a notion of social science without excluding the perspectives of the researched. This is because, and this is our third point, it is not methodologically prescriptive and will always seek to account for the interplay between people and the world in which they live. This entails acknowledging that a narrow sense of objective reality, one in which the irreducible subjective realities of human consciousness and being are left behind, is likely to be deficient. The tension this generates between the efforts to achieve an objective stand point by leaving a more subjective one behind involves a considerable epistemic risk (see Nagel 1986: 7), but in our view it is one worth taking.
One important methodological implication for sociolinguistic research seems to us to be as follows. Rather than defining the problem as the correlation of linguistic variation with variations in speakers' membership of (theoretically defined) social categories, researchers might start from the case. That is, the initial focus is on the language produced by speakers, and 'investigation becomes case not variable driven' (Williams 2000: 11). This kind of approach was adopted by Le Page and Tabouret-Keller when they collected data from a random sample of children who had immigrated to Britain from the Caribbean. They identified a series of linguistic variables and did not correlate these with pre-selected social categories, but subjected their findings to cluster analysis. Byrne (1998: 170) defines clusters as 'types, qualitative sets, which "emerge" from the application of computation to large multi-variate data sets'. From their results, Le Page and Tabouret Keller inferred varying self-identifications by the children with others in their 'multidimensional social space' (1985: 127).
Now as we have already explained, this account runs the risk of giving too great a weight to self-definition, although for some research purposes self-defined social categories may be the most appropriate. Alternatively, the researcher can follow Williams' (2000: 11) suggestions for those engaged in survey research, and defer the identification of categories until a later stage of the analysis:
The conjectural character of the data collection leads to a flexibility about both the definition and the measurement of 'variables', indeed the variables themselves are simply outcomes, or 'traces' (Byrne 2000) of yet unidentified (though possibly hypothesized) mechanisms. The only thing that we know is 'real' is the case itself and the operationalization of the variables is deferred to the identification of antecedent case characteristics.
Of course, the notion of a 'case' suggests that the researcher has already identified a phenomenon of which the object under scrutiny is itself an instance- or case- and this might seem to imply that the suspension of classification to a later stage of analysis is illusory. To be consistent in our argument, we would have to acknowledge that even the col lection of instances of speakers using language, and the identification of differences in the speech they produce- in other words, data collection approaches based on cases- are acts of interpretation and rely on some prior theorizing. This is why no director of a sociolinguistic project would be likely to send untrained researchers out with an instruction simply to 'record people talking': a lot of theoretically informed planning would underpin the collection of data. However, there is a difference between collecting 'cases' of linguistic production, on the one hand, and, on the other, deciding in advance which 'variables' are relevant, and which apply to each speaker.
We have highlighted some of the problems with social categories based on notions such as 'ethnicity', and suggested how a distinction between aggregates and collectives could be used in research into language variation. We will explore a further dimension of social identity and group identification, taking the applied linguistic topic of intercultural communication as our example. We will continue to develop the implications of our argument for approaches to research. In relation to the sociolinguistic issues, we shall suggest that analysts could allow for the possibilities not only that neither commonsense social categories, nor social scientific ones, will correlate categorically with the linguistic variables they have identified, but also that the mechanisms bringing about linguistic variation and change may be complex and multiplicative, rather than linear and additive.