Binding theory on minimalist assumptions
The Minimalist Program proposed in Chomsky (1995d) radically alters the foundations of syntactic theory by reformulating several fundamental theoretical constructs (e.g., involving phrase structure and transformations) as well as placing severe methodological restrictions on what tools and mechanisms might be employed in the construction of syntactic analyses. To a large extent, the reformulation of constructs is driven by methodological requirements—especially the assumption that all constructs must meet a criterion of conceptual necessity, the Ockham’s razor of the Minimalist Program. This paper attempts to sketch the ramifications of this and other assumptions of the Minimalist Program as they apply to a theory of binding. In particular, the discussion will focus on the effects of minimalist assumptions on the standard version of Binding Theory within the Principles and Parameters framework (Chomsky 1981; Chomsky & Lasnik 1993; Freidin 1994a). As in other areas of syntactic investigation, minimalist assumptions lead to a radical revision of the standard theory that has been in use for over a decade.
To begin, let us briefly review the standard theory in broad outline. It consists of the following three principles.

The relation bound is defined in terms of c-command and coindexation: one expression binds another if it c-commands the other and carries the same index. As formulated, the binding principles operate as output conditions (i.e., conditions on representations). Thus binding theory must specify to which level(s) of representation the binding principles apply. It must also define “local domain” for Principles A and B. A subsidiary question arises as to whether this definition is the same for both principles (cf. Chomsky & Lasnik 1993) or different (cf. Freidin 1986). Furthermore, binding theory must account for the fact that the three principles appear to be instantiated somewhat differently crosslinguistically. In the case of Principles A and B this may be due to parametric variation affecting the definition of local domain (see Yang 1983; Freidin 1992). Even for Principle C there appears to be some crosslinguistic variation which involves distinguishing domains in which the principle applies from those in which it does not (see Lasnik 1991).
The minimalist assumption that theoretical constructs must meet a criterion of conceptual necessity has a profound effect on the binding theory sketched above. The prime example discussed in Chomsky (1993) concerns levels of representation. The conceptual argument is crystal clear. The interface levels of Phonetic Form (PF) and Logical Form (LF) are required to account for how the computational system of human language CHL connects with other systems of the mind/brain involved in the production and perception of the physical signals of speech (including sign language) and in the translation of thought to language as well as language to thought. However, there is no such motivation for levels of D-structure and S-structure as characterized in previous work. Therefore, the postulation of such levels of representation is, under the Minimalist Program, illegitimate.2 This creates an immediate problem for any version of binding theory that proposes the application of binding principles at either level.
Consider for example the empirical argument discussed in Chomsky (1993) that Principle C applies at S-structure. The argument is based on the following evidence [Chomsky’s (23a–c)].

In (2a), because he c-commands John, the pronoun cannot take the name as its antecedent. In contrast, the pronoun in (2b) does not c-command the name and therefore the name may be interpreted as its antecedent. While the interpretation of (2c) is straightforward (the pronoun cannot take the name as its antecedent), its analysis is not. (2c) contains two wh-phrases, only one of which has moved from its grammatical function position to create the required quantifier/variable structure. Given the prohibition against vacuous quantification, the second wh-phrase must also move covertly at LF to create a quantifier/variable structure. If the entire phrase α adjoins to who, then in the resulting structure the name and pronoun will be in the same relation as in (2b). Since the name in (2c) cannot be interpreted as the antecedent of the pronoun as it can in (2b), this demonstrates that Principle C cannot apply at LF. However, there is another LF analysis for (2c) that doesn’t lead to this conclusion—namely, that only the quantifier how many gets fronted at LF. In this case the structural relation between the pronoun and the name remains at LF as we see it in (2c); hence there is no need to postulate a special level of S-structure at which Principle C can apply.
In the development of the Minimalist Program in chapter 4 of Chomsky (1995d), this line of analysis is motivated on economy grounds. Basically, since what drives movement is feature checking, it is assumed that what is being moved is a set of formal features. Overt movement involves pied piping of categories, for reasons that are not entirely clear. However, if economy requires that operations be minimal, then covert movement should apply only to features. Thus the covert quantificational movement of how many at LF would involve only the features on the quantifier, and not the remainder of the phrase α. The criterion of conceptual necessity argues in favor of this analysis over the alternative that requires the postulation of S-structure. In pursuing it, we discover that the alternative involved an unmotivated assumption—namely, that covert movement must involve categories.
The elimination of D-structure and S-structure as levels of representation has a salutary effect for a theory of binding. Without these levels, there is no possibility that either the three binding principles could apply at different levels within a single language or one or more principles could apply at different levels in different languages. Neither possibility was precluded in earlier versions of binding theory. Therefore the fact that they do not arise would have to be established via empirical argument, which as we have seen may be subject to unwarranted assumptions. The optimal situation given the minimalist perspective is when empirical arguments support conceptual arguments.
Limiting the application of binding principles to LF representations appears to require the adoption of the copy theory of movement transformations. In (3), for example, the pronoun cannot take the name as antecedent even though it does not c-command the name.

Assuming that Principle C is operative in such constructions, so that the antecedent relation in (3) is blocked for the same reason that it is blocked in (4), we are led to postulate an LF representation of (3) in which the pronoun c-commands the name.

The copying theory of movement would give us (5), which presumably would be translated into an LF representation along the lines of (6).

Without this kind of LF derivation, (3) might easily be construed as evidence that Principle C applies at a level of D-structure.
It is worth noting here that given the copy theory of movement operations in conjunction with the kind of deletion required to derive (6) from (5), one could still maintain that Move Category (i.e., Move α) applies covertly in (2c) but that the required deletion provides the same result as the Move Feature analysis. Therefore examples like (2c) do not provide empirical evidence for the Move Feature analysis as we might have otherwise expected.
Although the copying analysis is required if binding principles apply only at LF, it raises a potentially difficult problem for examples like (2b) where the moved wh-phrase includes a relative clause. Thus compare (7) to (3).

In (7) the name Alice may be construed as the antecedent of the pronoun, in contrast to (3) where it cannot. This requires that in the LF representation of (7), the relative clause does not show up in the position of the trace. Exactly how this to be achieved is not clear, nor is it clear exactly what LF representation of (7) would be. Under the copying analysis (7) could involve (8).

The derivation of the LF representation for (7) would involve some deletions— presumably pictures in the moved phrase and the quantifier how many in the copy. If we treat (7) on a par with (3) then the relative clause would be deleted in the moved phrase as well—yielding the wrong structure since Alice may be interpreted as the antecedent of the pronoun. This shows that there is an apparent asymmetry in the behavior of relative clauses and complements with respect to binding principles (cf. Freidin 1986, 1992, 1994a; Lebeaux 1988). How this is to be captured in an LF representation (7) seems problematic. Taking (6) as a model, (7) would presumably appear at LF as (9).

The problem with (9) is that x is a variable ranging over integers whereas the relative clause modifies pictures not an integer.
Rather than pursue this analysis further, let us consider a related set of facts that suggest that a Principle C analysis of these constructions is perhaps on the wrong track. If we substitute a copy of the name in (3) and (7) for the pronoun, we should presumably get the same results with respect to Principle C since it is the binding of the name that is at issue. Perhaps surprisingly, this turns out not to be the case.

In both (10) and (11) it is possible to interpret the two instances of the name Alice as referring to the same person. This is expected for (11) given that its LF representation is like that of (7) where the relative clause is not reconstructed in object position. The coreferential interpretation of (10) is, however, completely unexpected given that its LF representation would be parallel to that of (3)—i.e., (6), hence (12).

Under the standard theory the two names on the coreferential reading are in a binding relation which should be prohibited by Principle C. In (13) where no overt movement is involved this binding relation is prohibited.

The natural interpretation of (13) requires that there be two people named Alice.
The contrast between (3) and (10) is unexplained and apparently unexplainable under the standard theory. Moreover, the standard theory makes the wrong prediction for the interpretation of (10). Separating the pronoun/name case from the name/name case along the lines of Lasnik (1991), where Principle C is split into several conditions depending on the status of the binder, one of which states that an r-expression is pronoun-free (i.e., cannot be bound by a pronoun), eliminates the problem of having a principle apply in one case but fail to apply in a structurally identical case. However, we are still left with a serious problem for the residue of Principle C since it predicts the wrong interpretation for (10).
So far we have been discussing the application of the standard binding theory, specifically Principle C, at LF because that is where it would have to apply given the minimalist assumption that the only two levels of representation are the interface levels PF and LF. It can’t apply at PF given the further assumption that PF contains no structural information. “PF is a representation in universal phonetics, with no indication of syntactic elements or relations among them (X-bar structure, binding, government, etc.)” (Chomsky (1993:194)). Therefore, binding principles can only apply to LF representations. However, it is not clear that under minimalist assumptions the formulation of the binding principles in the standard theory is in fact conceptually motivated.
Consider first that fact that under the standard theory the definition of “bound” involves two nominal expressions in a c-command relation that are coindexed. Under minimalist assumptions, however, indices and similar devices are not available. Chomsky takes it as a natural condition “that outputs consist of nothing beyond properties of items of the lexicon (lexical features); in other words, that the interface levels consist of nothing more than arrangements of lexical features” (Chomsky 1995d:225), thereby meeting a condition of inclusiveness. Furthermore, he claims that:
A theoretical apparatus that takes indices seriously as entities, allowing them to figure in operations (percolation, matching, etc.), is questionable on more general grounds. Indices are basically the expression of a relationship, not entities in their own right. They should be replaceable without loss by a structural account of the relation they annotate. (Chomsky 1993:fn.53)
Obviously if we eliminate indices as a grammatical device, then binding theory must be recast in some other way since the standard theory is to a large extent a theory about the assignment of indices.
The alternative proposed in Chomsky & Lasnik (1993) (and adopted in Chomsky (1993)) involves replacing indexing procedures with interpretive procedures. As Chomsky & Lasnik note, the indexing procedures of the standard theory require interpretive procedures as well. By recasting the binding principles as interpretive procedures it is possible to dispense with the indexing procedures. Thus the binding principles in (1) become the interpretive procedures of (14), where D stands for the relevant local domain in which the procedure applies.

Under this proposal, the principles of binding are not conditions on representations. Rather, they assign certain interpretative relations among nominal expressions, and are thereby derivational in nature. Thus (14a) as an interpretive procedure does not account for cases where the interpretation cannot apply, e.g., (15).

That is, we need a further statement (16) to account for (15).

If (14a) is the only interpretive rule for anaphors, then the only possible antecedent will be a c-commanding phrase in D.8 In the case of (14b), we need no further condition to process the disjoint reference interpretation for pronouns. However, if (14b) is the only rule of pronoun interpretation in binding theory, then CHL does not account for the fact that sentences like (17) are ambiguous.

The pronoun and the name will not be interpreted as disjoint in reference by (14b), but that does not say whether they are coreferential or not. Thus we have returned in essence to Lasnik’s 1976 theory of pronominal coreference where the coreference possibility in (17) is not given by any rule of grammar.
Principle C of the interpretive theory (14c) does not fare any better than the standard theory version (1c) with respect to (10). It makes the same wrong prediction with respect to (12), the putative LF representation of (10). Moreover, the existence of such a rule of interpretation (or alternatively a condition on representations like (1c)) ought to be suspect on conceptual grounds. While both anaphors and pronouns act as anaphoric expressions—i.e., they stand in for some other nominal expression, r-expressions do not. Thus it seems inappropriate for that reason to treat them as if they could behave as anaphoric expressions and therefore must be interpreted as disjoint from c-commanding nominals.9 If we restrict our attention to the issue of antecedents for anaphoric expressions, then Principle C would be restricted to covering just examples like (18).

If so, then Principle C might be reformulated as the interpretive rule (19).

(19) is equivalent to the Lasnik (1991) principle that r-expressions be pronounfree, but without reference to “bound r-expressions,” which I am suggesting should be illegitimate on conceptual grounds. With the elimination of indexing, the issue of coreference between r-expressions should disappear. Presumably there is no need for a special grammatical mechanism to check pairs of r-expressions to determine whether they corefer or not.
At this point we still have no account for the fact that when two r-expressions are phonetically identical, there exists an interpretive option to treat them as having the same reference—which does not entail that one is anaphoric on the other. In some cases this option is realized (e.g., (10) and (11) above), in others it is prohibited (e.g., (13)). The fact that the difference depends on whether a c-command relation holds between the two r-expressions is suggestive that binding theory is somehow really involved, even though it is unclear how this could be achieved if binding theory is limited solely to anaphoric relations within sentences, where one expression stands in for another (its antecedent). The same reference option for r-expressions is an entirely different type of phenomenon, involving the assignment of reference to r-expressions which is presumably not part of CHL. Anaphoric relations, in contrast, concern the assignment of antecedents to anaphoric expressions (bound anaphors, pronouns, and pronominal epithets), a purely grammatical phenomenon. R-expressions involve word/world relations, while anaphoric expressions involve word/word relations. Thus on conceptual grounds alone it seems natural to separate the two cases.
The picture of binding theory on minimalist assumptions that is beginning to emerge seems very different from the standard theory. Instead of three conditions on indexing representations involving anaphors, pronouns, and r-expressions, we have three rules of interpretation—one for anaphors and two for pronouns—that involve the relations between anaphoric expressions and their antecedents. Furthermore, the rule for an anaphor specifies when a nominal can be interpreted as its antecedent, while the rules for a pronoun specify when a nominal expression cannot be its antecedent. The rule for anaphors can fail to apply or apply improperly yielding deviant structures. The rules for pronouns cannot because the only relation CHL specifies for a pronoun is disjoint reference and a pronoun, unlike an anaphor, does not require an antecedent in the sentence in which it occurs.
Recasting the principles of binding theory as rules of interpretation instead of conditions on representations avoids a potentially serious problem with respect to the minimalist assumption that all output conditions are external interface conditions (bare output conditions). First, unless parametric variation extends to bare output conditions (a totally unmotivated assumption at this point), it would be difficult to explain crosslinguistic variation for binding configurations documented in the literature. Moreover, it is far from clear how standard violations of Principles B and C could be construed in any real sense as violations of Full Interpretation (FI), the only candidate we presently have for a bare output condition (see Freidin 1997). With bound anaphors, however, it is possible to construe the failure of the interpretive rule (e.g., (15a–b)) as resulting in a violation of FI, if we can assume that an anaphor without an antecedent is assigned no referential interpretation. If the reference of nominal expressions is assigned outside of CHL, then FI will apply externally to CHL as well. This indicates that FI must operate as a bare output condition, since the failure to assign a referential interpretation occurs outside CHL.
At the conclusion of his survey of the history of modern binding theory (1989), Howard Lasnik writes:
…the developments explored here can best be seen not as a series of revolutionary upheavals in the study of anaphora, but rather the successive refinement of one basic approach, and one that has proven remarkably resilient. Given that BT has become the subject of intensive investigation, with new phenomena in previously unexplored languages being constantly brought to bear, and all this while old problems from familiar languages remain, further refinements, or even revolutionary upheavals, are inevitable. (Lasnik 1989:34)
From the preceding discussion of binding theory on minimalist assumptions it would seem that although some revolutionary upheaval may still be off in the future, the ground is certainly shifting so that our perspective appears to be undergoing a significant change.