segunda-feira, 4 de março de 2013

A Usage-Based Account of Constituency and Reanalysis


A Usage-Based Account of Constituency and Reanalysis
Constituent structure is considered to be the very foundation of linguistic competence and often considered to be innate, yet we show here that it is derivable from the domain-general processes of chunking and categorization. Using modern and diachronic corpus data, we show that the facts support a view of constituent structure as gradient (as would follow from its source in chunking and categorization) and subject to gradual changes over time. Usage factors (i.e., repetition) and semantic factors both influence chunking and categorization and, therefore, influence constituent structure.
(...)
in this article we propose to derive constituent structure from the domain-general processes of chunking and categorization within the storage network for language. Because language is a dynamic system, an important part of our argument will rest on the idea that constituent structure, like all of grammar, is constantly undergoing gradual change. Thus, structural reanalysis, as often discussed in the context of grammaticalization, will be pivotal to our argument and exposition.
We mean by 'structural reanalysis' a change in constituent structure, as when 'to' as an earlier allative or infinitive marker with a verb as its complement fuses with 'going' in the future expression 'be going to' ('going [to see]' > '[going to] see'). Indicators of reanalysis include changes in distribution and phonological changes ('going to' -> 'gonna').
Are such changes abrupt or gradual? In generative models of syntax (see, e.g., Lightfoot, 1979; Roberts & Roussou, 2003), structural reanalysis is necessarily abrupt, because it is held that a sequence of words has a unique, discrete constituent analysis.1 In this view, constituents are clearly defined and do not overlap; in a sequence such as going to VERB, to must be grouped either with the following verb, or with going, with no intermediate stages. 
(...)
However, because most linguistic change appears to be quite gradual, with slowly changing meanings and distributions and overlapping stages, a problem arises for a theory with discrete constituent structure. Evidence from the gradualness of change has led some researchers to doubt discrete categories and structures (Haspelmath, 1998; Hoffmann, 2005; Quirk, Greenbaum, Leech, & Svartvik, 1985).
Continuing from Bybee and Scheibman (1999), we join these researchers in proposing that constituent structure can change gradually. We take the view that it is altogether common even for an individual speaker to have nondiscrete syntactic representations for the same word sequence. Taking a complex systems-based perspective, we hold that syntactic structure is in fact much richer than the discrete constituency view would indicate. There are multiple overlapping and, at times, competing influences on the shape of units in the grammar, and these multiple factors have an ongoing effect on each speaker’s synchronic representations of syntactic structure. Specifically, syntactic constituents are subject to ongoing influence from general, abstract patterns in language, in addition to more localized, item-specific usage patterns. The foregoing perspective makes it possible that the same word sequence may be characterized by multiple constituent structures and that these structures have gradient strengths rather than discrete boundaries. Our position in this article is thus that constituency may change in a gradual fashion via usage, rather than via acquisition, and that structural reanalysis need not be abrupt.
(...) constituent structure is gradient, mutable, and emergent from domain-general processes (...) 
chunking and categorization together provide constituency analyses of phrases and utterances for speakers




Constituent Structure as Emergent From Chunking, Categorization, and Generalization
Bybee (2002, in press) discusses the nature of sequential learning and chunking as it applies to the formation of constituents. Because members of the same constituent appear in a linear sequence with some frequency, these items are subject to chunking, by which sequences of repeated behavior come to be stored and processed as a single unit.

A chunk is a unit of memory organization, formed by bringing together a set of already formed chunks in memory and welding them together into a larger unit. Chunking implies the ability to build up such structures recursively, thus leading to a hierarchical organization of memory. Chunking appears to be a ubiquitous feature of human memory.
Chunking occurs automatically as behaviors are repeated in the same order, whether they are motor activities such as driving a car or cognitive tasks such as memorizing a list. Repetition is the factor that leads to chunking, and chunking is the response that allows repeated behaviors to be accessed more quickly and produced more efficiently (Haiman, 1994). Chunking has been shown to be subject to The Power Law of Practice (Anderson, 1993), which stipulates that performance improves with practice, but the amount of improvement decreases as a function of increasing practice or frequency. Thus, once chunking occurs after several repetitions, further benefits or effects of repetition accrue much more slowly.

Changes in Constituent Structure
In grammaticalization it often happens that grammaticalizing expressions change their constituent structure. Thus, it is often said that grammaticalization is the reanalysis of a lexical item as a grammatical item. As Haspelmath (1998) pointed out, often the change can be thought of as a simple change in a category status. Thus, a verb becomes an auxiliary; a serial verb becomes an adposition or a complementizer; a noun becomes a conjunction or adposition. In some cases, however, shifts in constituent boundaries do occur; in particular, it is common to lose some internal constituent boundaries. A prominent example of such a change involves complex prepositions. Many complex prepositions start out as a sequence of two prepositional phrases (e.g., on top of NP) but evolve into a kind of intermediate structure in some analyses—the complex preposition—and eventually they can even develop further into simple prepositions, as has occurred with beside, behind, and among (Hopper & Traugott, 2003; Ko·nig & Kortmann, 1991; Svorou, 1994).
(...)
Given the above proposed model, Hay (2001) reasoned that if the complex unit is more frequent than its parts, it is more likely to be accessed as a unit, leading to the loss of analyzability that comes about through categorization. Applied to 'in spite of', we would predict that as the complex phrase becomes more frequent than the simple noun 'spite', it would also become more autonomous and less analyzable.


Conclusion

We have taken stock here of the traditional discrete constituency view that holds that a word sequence either has a holistic structure or a unique, nested hierarchical structure. The accounts we have examined ultimately reject usage as an indicator of constituent structure—discarding evidence from semantics and any usage data that might be countered by partial evidence from introspective syntactic tests. Such a conservative approach rejects even the possibility of finding evidence that particular sequences may have reached an intermediate stage of constituency. Moreover, the discrete constituency view would seem to hold that grammar is driven only by abstract syntactic generalizations and is immune to any gradual effects from item-specific usage patterns.
In contrast, as we do not take constituent structure as given innately, we do not give priority to syntactic tests. Rather we consider data from usage, semantics, and language change. Indeed, we have shown that chunking and categorization have semantic effects and change incrementally over time.
Moreover, in keeping with the theory of complex adaptive systems, we consider constituent structure to be emergent from the domain-general processes of chunking and categorization. Human minds track multiple factors related to constituency, and this complex processing correlates with a rich and dynamic structural representation for word sequences. In our model, constituency is the result of interacting influences that are both local and global in nature. The global influences that help shape constituents correspond to general patterns in usage. On the other hand, constituency may also be shaped locally by item-specific forces over time. If a sequence is consistently used in a particular context (with complex prepositions like in spite of as a case in point), that sequence will gradually form into a unit, overriding general patterns elsewhere in usage. In this regard, we embrace Bolinger’s early complex systems view of language as a “jerry-built” and heterogeneous structure that is also intricate and tightly organized (1976, p. 1). Rather than assuming that structure is given a priori via top-down blueprints, we agree with Bolinger (1976) and Hopper (1987) that structure emerges locally and is subject to ongoing revision, even while general patterns exhibit apparent stability.

Language is a complex adaptative system


Language has a fundamental social function. Processes of human interaction along with domain-general cognitive processes shape the structure and knowledge of language. Recent research in cognitive sciences has demonstrated that patterns of use strongly affect how language is acquired, is used, and changes. These processes are not independent of one another but are facets of the same complex adaptive system (CAS). Language as a CAS involves the following key features: The system consists of multiple agents (the speakers in the speech community) interacting with one another. The system is adaptive; that is, speakers' behavior is based on their past interactions, and current and past interactions together feed forward into future behavior. A speaker's behavior is the consequence of compering factors ranging from perceptual constraints to social motivations  The structure of language emerge from interrelated patterns of experience, social interaction, and cognitive mechanisms. The CAS approach reveals commonalities in many areas of language research, including first and second language acquisition, historical linguistics, psycholinguistics, language evolution, and computational modeling.


Introduction: Shared Assumptions

Language has a fundamentally social function. Processes of human interaction along with domain-general cognitive processes shape the structure and knowledge of language. Recent research across a variety of disciplines in the cognitive sciences has demonstrated that patterns of use strongly affect how language is acquired, is structured, is organized in cognition, and changes over time. However, there is mounting evidence that processes of language acquisition, use, and change are not independent of one another but are facets of the same system. We argue that this system is best constructed as a complex adaptive system (CAS). This system is radically different from the static system of grammatical principles characteristic of the widely held generativist approach. Instead, language as a CAS of dynamic usage and its experience involves the following key features: (a) The system consists of multiple agents (the speakers in the speech community) interacting with one another. (b) The system is adaptive; that is, speaker' behavior is based on their past interactions, and current and past interactions together feed forward into future behavior. (c) A speaker's behavior is the consequence of competing factors ranging from perceptual mechanics to social motivations. (d) The structure of language emerge from interrelated patterns of experience, social interaction, and cognitive processes.

The advantage of viewing language as a CAS is that it allows us to provide a unified account of seemingly unrelated linguistic phenomena. These phenomena include the following: variation at all level of linguistic organization; the probabilistic nature of linguistic behavior; continuous change within agents and across speech communities; the emergence of grammatical regularities from the interaction of agents in language use; and stage like transitions due to underlying nonlinear processes.


Language and Social Interaction

Language is shaped by human cognitive abilities such as categorization, sequential processing, and planning. However, it is more that their simple product. Such cognitive abilities do not require language; if we had only those abilities, we would not need to talk. Language is used for human social interaction, and so its origins and capacities are dependent on its role in our social life. To understand how language has evolved in the human lineage and why it has the properties we can observe today, we need to look at the combined effect of many interacting constraints, including the structure of thought processes, perceptual and motor biases, cognitive limitations, and socio-pragmatic factors.

Primates... reformed over time.




Usage-Based Grammar

We adopt here a usage-based theory of grammar in which the cognitive organization of language is based directly on experience with language. Rather than being an abstract set of rules or structures that are only indirectly related to experience with language, we see grammar as a network built up from the categorized instances of language use (Bybee, 2006; Hopper, 1987). The basic units of grammar are constructions, which are direct form-meaning pairings that range from the very specific (words or idioms) to the more general (passive construction, ditransitive construction), and from very small units (words with affixes, walked) to clause-level or even discourse-level units (Croft, 2001; Goldberg, 2003, 2006).

Because grammar is based on usage, it contains many details of cooccurrence as well as a record of the probabilities of occurrence and cooccurrence. The evidence for the impact of usage on cognitive organization includes the fact that language users are aware of specific instances of constructions that are conventionalized and the multiple ways in which frequency of use has an impact on structure. The latter include speed of access related to token frequency and resistance to regularization of high-frequency forms (Bybee, 1995, 2001, 2007); it also includes the role of probability in syntactic and lexical processing (Ellis, 2002; Jurafsky, 2003; MacDonald & Christiansen, 2002) and the strong role played by frequency of use in grammaticalization (Bybee, 2003).
A number of recent experimental studies (Saffran, Aslin, & Newport, 1996; Saffran, Johnson, Aslin, & Newport, 1999; Saffran & Wilson, 2003) show that both infants and adults track co-occurrence patterns and statistical regularities in artificial grammars. Such studies indicate that subjects learn patterns even when the utterance corresponds to no meaning or communicative intentions. Thus, it is not surprising that in actual communicative settings, the co-occurrence of words has an impact on cognitive representation. Evidence from multiple sources demonstrates that cognitive changes occur in response to usage and contribute to the shape of grammar. Consider the following three phenomena:

1. Speakers do not choose randomly from among all conceivable combinatorial possibilities when producing utterances. Rather there are conventional ways of expressing certain ideas (Sinclair, 1991). Pawley and Syder (1983) observed that “nativelike selection” in a language requires knowledge of expected speech patterns, rather than mere generative rules. A native English speaker might say I want to marry you, but would not say I want marriage with you or I desire you to become married to me, although these latter utterances do get the point across. Corpus analyses in fact verify that communication largely consists of prefabricated sequences, rather than an “open choice” among all available words (Erman & Warren, 2000). Such patterns could only exist if speakers were registering instances of co-occurring words, and tracking the contexts in which certain patterns are used.

2. Articulatory patterns in speech indicate that as words co-occur in speech, they gradually come to be retrieved as chunks. As one example, Gregory, Raymond, Bell, Fossler-Lussier, & Jurafsky (1999) find that the degree of reduction in speech sounds, such as word-final “flapping” of English [t], correlates with the “mutual information” between successive words (i.e., the probability that two words will occur together in contrast with a chance distribution) (see also Bush, 2001; Jurafsky, Bell, Gregory, & Raymond, 2001). A similar phenomenon happens at the syntactic level, where frequent word combinations become encoded as chunks that influence how we process sentences on-line (Ellis, 2008b; Ellis, Simpson-Vlach, & Maynard, 2008; Kapatsinski & Radicke, 2009; Reali & Christiansen, 2007a, 2007b).

3. Historical changes in language point toward a model in which patterns of co-occurrence must be taken into account. In sum, “items that are used together fuse together” (Bybee, 2002). For example, the English contracted forms (I’m, they’ll) originate from the fusion of co-occurring forms (Krug, 1998). Auxiliaries become bound to their more frequent collocate, namely the preceding pronoun, even though such developments run counter to a traditional, syntactic constituent analysis.

(...)
In the usage-based framework, we are interested in emergent generalizations across languages, specific patterns of use as contributors to change and as indicators of linguistic representations, and the cognitive underpinnings of language processing and change. Given these perspectives, the sources of data for usage-based grammar are greatly expanded over that of structuralist or generative grammar: Corpus-based studies of either synchrony or diachrony as well as experimental and modeling studies are considered to produce valid data for our understanding of the cognitive representation of language.





The Development of Grammar out of Language Use

The mechanisms that create grammar over time in languages have been identified as the result of intense study over the last 20 years (Bybee et al., 1994; Heine, Claudi, & H ̈ nnemeyer, 1991; Hopper & Traugott, 2003). In the history of well-documented languages it can be seen that lexical items within constructions can become grammatical items and loosely organized elements within and across clauses come to be more tightly joined. Designated “grammaticalization,” this process is the result of repetition across many speech events, during which sequences of elements come to be automatized as neuromotor routines, which leads to their phonetic reduction and certain changes in meaning (Bybee, 2003; Haiman, 1994). Meaning changes result from the habituation that follows from repetition, as well as from the effects of context. The major contextual effect comes from co-occurring elements and from frequently made inferences that become part of the meaning of the construction.

For example, the recently grammaticalized future expression in English be going to started out as an ordinary expression indicating that the subject is going somewhere to do something. In Shakespeare’s English, the construction had no special properties and occurred in all of the plays of the Bard (850,000 words) only six times. In current English, it is quite frequent, occurring in one small corpus of British English (350,000 words) 744 times. The frequency increase is made possible by changes in function, but repetition is also a factor in the changes that occur. For instance, it loses its sense of movement in space and takes on the meaning of “intention to do something,” which was earlier only inferred. With repetition also comes phonetic fusion and reduction, as the most usual present-day pronunciation of this phrase is (be) gonna. The component parts are no longer easily accessible.

The evidence that the process is essentially the same in all languages comes from a crosslinguistic survey of verbal markers and their diachronic sources in 76 unrelated languages (Bybee et al., 1994). (...)





First and Second Language Acquisition

Usage-based theories of language acquisition (Barlow & Kemmer, 2000) hold that we learn constructions while engaging in communication, through the “interpersonal communicative and cognitive processes that everywhere and always shape language” (Slobin, 1997). They have become increasingly influential in the study of child language acquisition (Goldberg, 2006; Tomasello, 2003). They have turned upside down the traditional generative assumptions of innate language acquisition devices, the continuity hypothesis, and top-down, rule-governed processing, replacing these with data-driven, emergent accounts of linguistic systematicities.

(...)

In summary, we have the following: (a) Usage leads to change: High-frequency use of grammatical functors causes their phonological erosion and homonymy. (b) Change affects perception: Phonologically reduced cues are hard to perceive. (c) Perception affects learning: Low-salience cues are difficult to learn, as are homonymous/polysemous constructions because of the low contingency of their form-function association. (d) Learning affects usage: (i) Where language is predominantly learned naturalistically by adults without any form-focus, a typical result is a Basic Variety of interlanguage, low in grammatical complexity but communicatively effective. Because usage leads to change, in cases in which the target language is not available from the mouths of L1 speakers, maximum contact languages learned naturalistically can thus simplify and lose grammatical intricacies. Alternatively, (ii) where there are efforts promoting formal accuracy, the attractor state of the Basic Variety can be escaped by means of dialectic forces, socially recruited, involving the dynamics of learner consciousness, form-focused attention, and explicit learning. Such influences promote language maintenance.




Modeling Usage-Based Acquisition and Change

In the various aspects of language considered here, it is always the case that form, user, and use are inextricably linked. However, such complex interactions are difficult to investigate in vivo. Detailed, dense longitudinal studies of language use and acquisition are rare enough for single individuals over a time course of months. Extending the scope to cover the community of language users, and the timescale to that for language evolution and change, is clearly not feasible. Thus, our corpus studies and psycholinguistic investigations try to sample and focus on times of most change and interactions of most significance. However, there are other ways to investigate how language might emerge and evolve as a CAS (complex adaptive system). A valuable tool featuring strongly in our methodology is mathematical or computational modeling.

Given the paucity of relevant data, one might imagine this to be of only limited use. We contend that this is not the case. Because we believe that many properties of language are emergent, modeling allows one to prove, at least in principle, that specific fundamental mechanisms can combine to produce some observed effect (Holland, 1995, 1998, 2006a, 2006b; Holland et al., 2005). Although this may also be possible through an entirely verbal argument, modeling provides additional quantitative information that can be used to locate and revise shortcomings. For example, a mathematical model constructed by Baxter et al. (2009) within a usage-based theory for new-dialect formation (Trudgill, 2004) was taken in conjunction with empirical data (Gordon et al., 2004) to show that although the model predicted a realistic dialect, its formation time was much longer than that observed. Another example comes from the work of Reali and Christiansen (2009), who demonstrated how the impact of cognitive constraints on sequential learning across many generations of learners could give rise to consistent word order regularities.

Modeling can also be informative about which mechanisms most strongly affect the emergent behavior and which have little consequence. To illustrate, let us examine our view that prior experience is a crucial factor affecting an individual speaker’s linguistic behavior. It is then natural to pursue this idea within an agent-based framework, in which different speakers may exhibit different linguistic behavior and may interact with different members of the community (as happens in reality). Even in simple models of imitation, the probability that a cultural innovation is adopted as a community norm, and the time taken to do so, is very strongly affected by the social network structure (Castellano, Fortunato, & Loreto, 2007, give a good overview of these models and their properties). This formal result thus provides impetus for the collection of high-quality social network data, as their empirical properties appear as yet poorly established. The few cases that have been discussed in the literature—for example, networks of movie co-stars (Watts & Strogatz, 1998), scientific collaborators (Newman, 2001), and sexually-active high school teens (Bearman, Moody, & Stovel, 2004)—do not have a clear relevance to language. We thus envisage a future in which formal modeling and empirical data collection mutually guide one another.

(...)

Above all, a usage-based model should provide insight into the frequencies of variants within the speech community. The rules for producing utterances should then be inducted from this information by general mechanisms. This approach contrasts with an approach that has speakers equipped with fixed, preexisting grammars.

(...)

Despite these observations, many details of the linguistic interactions remain unconstrained and one can ask whether having a model reproduce observed phenomena proves the specific set of assumptions that went into it. The answer is, of course, negative. However, greater confidence in the assumptions can be gained if a model based on existing data and theories makes new, testable predictions. In the event that a model contains ad hoc rules, one must, to be consistent with the view of language as a CAS, be able to show that these are emergent properties of more fundamental, general processes for which there is independent support.




Characteristics of Language as a Complex Adaptive System

We now highlight seven major characteristics of language as a CAS, which are consistent with studies in language change, language use, language acquisition, and computer modeling of these aspects.



Distributed Control and Collective Emergence

Language exists both in individuals (as idiolect) and in the community of users (as communal language). Language is emergent at these two distinctive but interdependent levels: An idiolect is emergent from an individual’s language use through social interactions with other individuals in the communal language, whereas a communal language is emergent as the result of the interaction of the idiolects. Distinction and connection between these two levels is a common feature in a CAS. Patterns at the collective level (such as bird flocks, fish schools, or economies) cannot be attributed to global coordination among individuals; the global pattern is emergent, resulting from long-term local interactions between individuals. Therefore, we need to identify the level of existence of a particular language phenomenon of interest. For example, language change is a phenomenon observable at the communal level; the mechanisms driving language change, such as production economy and frequency effects that result in phonetic reduction, may not be at work in every individual in the same way or at the same time. Moreover, functional or social mechanisms that lead to innovation in the early stages of language change need not be at work in later stages, as individuals later may acquire the innovation purely due to frequency when the innovation is established as the majority in the communal language. The actual process of language change is complicated and interwoven with a myriad of factors, and computer modeling provides a possible venue to look into the emergent dynamics (see, e.g., Christiansen & Chater, 2008, for further discussion).



Intrinsic Diversity

(...) Each idiolect is the product of the individual’s unique exposure and experiences of language use (Bybee, 2006). Sociolinguistics studies have revealed the large degree of orderly heterogeneity among idiolects (Weinreich, Labov, & Herzog, 1968), not only in their language use but also in their internal organization and representation (Dabrowska, 1997).



Perpetual Dynamics

Both communal language and idiolects are in constant change and reorganization. Languages are in constant flux, and language change is ubiquitous (Hopper, 1987). At the individual level, every instance of language use changes an idiolect’s internal organization (Bybee, 2006).



Adaptation Through Amplification and Competition of Factors

Complex adaptive systems generally consist of multiple interacting elements, which may amplify and/or compete with one another’s effects. Structure in complex systems tends to arise via positive feedback, in which certain factors perpetuate themselves, in conjunction with negative feedback, in which some constraint is imposed—for instance, due to limited space or resources (Camazine et al., 2001; Steels, 2006). Likewise in language, all factors interact and feed into one another. (...)



Nonlinearity and Phase Transitions

In complex systems, small quantitative differences in certain parameters often lead to phase transitions (i.e., qualitative differences). (...)



Sensitivity to and Dependence on Network Structure

Network studies of complex systems have shown that real-world networks are not random, as was initially assumed (Barab ́ si, 2002; Barbar ́ si & Albert, 1999; Watts & Strogatz, 1998), and that the internal structure and connectivity of the system can have a profound impact on system dynamics (Newman, 2001; Newman, Barab ́ si, & Watts, 2006). Similarly, linguistic interactions are not via random contacts; they are constrained by social networks. The social structure of language use and interaction has a crucial effect in the process of language change (Milroy, 1980) and language variation (Eckert, 2000), and the social structure of early humans must also have played important roles in language origin and evolution. An understanding of the social network structures that underlie linguistic interaction remains an important goal for the study of language acquisition and change. The investigation of their effects through computer and mathematical modeling is equally important (Baxter
et al., 2009).



Change Is Local

Complexity arises in systems via incremental changes, based on locally available resources, rather than via top-down direction or deliberate movement toward some goal (see, e.g., Dawkins, 1985). Similarly, in a complex systems framework, language is viewed as an extension of numerous domain-general cognitive capacities such as shared attention, imitation, sequential learning, chunking, and categorization (Bybee, 1998b; Ellis, 1996). Language is emergent from ongoing human social interactions, and its structure is fundamentally molded by the preexisting cognitive abilities, processing idiosyncrasies and limitations, and general and specific conceptual circuitry of the human brain. Because this has been true in every generation of language users from its very origin, in some formulations, language is said to be a form of cultural adaptation to the human mind, rather than the result of the brain adapting to process natural language grammar (Christiansen, 1994; Christiansen & Chater, 2008; Deacon, 1997; Schoenemann, 2005). These perspectives have consequences for how language is processed in the brain. Specifically, language will depend heavily on brain areas fundamentally linked to various types of conceptual understanding, the processing of social interactions, and pattern recognition and memory. It also predicts that so-called “language areas” should have more general, prelinguistic processing functions even in modern humans and, further, that the homologous areas of our closest primate relatives should also process information in ways that makes them predictable substrates for incipient language. Further, it predicts that the complexity of communication is to some important extent a function of social complexity. Given that social complexity is, in turn, correlated with brain size across primates, brain size evolution in early humans should give us some general clues about the evolution of language (Schoenemann, 2006). Recognizing language as a CAS allows us to understand change at all levels.





Conclusions

Cognition, consciousness, experience, embodiment, brain, self, human interaction, society, culture, and history are all inextricably intertwined in rich, complex, and dynamic ways in language. Everything is connected. Yet despite this complexity, despite its lack of overt government, instead of anarchy and chaos, there are patterns everywhere. Linguistic patterns are not preordained by God, genes, school curriculum, or other human policy. Instead, they are emergent -- synchronic patterns of linguistic organization at numerous levels (phonology, lexis, syntax, semantics, pragmatics, discourse, genre, etc.), dynamic patterns of usage, diachronic patterns of language change (linguistic cycles of grammaticalization, pidginization, creolization, etc.), ontogenetic developmental patterns in child language acquisition, global geopolitical patterns of language growth and decline, dominance and loss, and so forth. We cannot understand these phenomena unless we understand their interplay. The individual focus articles that follow in this special issue illustrate such interactions across a broad range of language phenomena, and they show how a CAS framework can guide future research and theory.

The Problem of Universals in Language


The Problem of Universals in Language
by Charles F. Hockett

1. Introduction

"The only useful generalizations about language are inductive generalizations" (Bloomfield, 1933, p.20). This admonition is clearly important, in the sense that we do not want to invert language universals, but to discover them, How to discover them is not so obvious. It would be fair to claim that the search is coterminous with the whole enterprise of linguistic in at least two ways. The first claim in which this claim is true is heuristic: we can never be sure, in any sort of linguistic study, that it will not reveal something of importance for the search. The second way in which the claim is plausible, if not automatically true, appears when we entertain one of the various possible definitions of linguistics as a branch of science: that branch devoted to the discovery of the place of human language in the universe. This definition leaves the field vague to the extent that the problem of linguistics remains unsolved. Only if, as is highly improbable, the problem were completely answered should we know exactly what linguistics is -- and at the same millennial moment there would cease to be any justification for the field. It is hard to discern any clear difference between "the search for language universals" and "the discovery of the place of the human language in the universe." They seem rather to be, respectively, a new-fangled and old-fashioned way of describing the same thing.


(...)


1.1 The assertion of a language universal must be founded on extrapolation as well as on empirical evidence.

Of course this is true in the trivial sense that we do not want to delay generalizing until we have full information on all the languages of the world. We should rather formulate generalizations as hypotheses, to be tested as new empirical information becomes available. But there is a deeper implication. If we had full information on all languages now spoken, there would remain languages recently extinct on which the information was inadequate. There is no point in imaging that we have adequate information also on these extinct languages, because that would be imagining the impossible. The universe seems to be so constructed that complete factual information is unattainable, at least in the sense that there are past events that have left only incomplete records. Surely we seek constantly to widen the empirical base for our generalizations; equally surely, we always want our generalizations to subsume some of the unobserved, and even some of the unobservable, along with all of the observed.



1.3 A feature can be widespread or even universal without being important

This is most easily shown by a trick. Suppose that all the languages of the world except English were to become extinct. Thereafter, any assertion true of English would also assent a (synchronic) language universal. Since languages no longer spoken may have lacked features we believe universal or widespread among those now spoken, mere frequency can hardly be a measure of importance.



1.4 The distinction between the universal and the merely widespread is not necessarily relevant

The reasoning is as for 1.3. Probably we all feel that the universality of certain features might be characterized as "accidental" -- they might just as well have turned out to be merely widespread. This does not tell us how to distinguish between the "accidentally" and the "essentially" universal. On the other hand, that which is empirically known to be merely widespread is thereby disqualified as an "essential" universal -- though careful study may show that it is symptomatic of one.



1.5 The search for universals cannot be usefully separated from the search for a meaningful taxonomy of languages.

(Here "taxonomy" refers to what might also be called "typology", not to genetic classification.) Suppose that some feature, believed to be important and universal, turns out to be lacking in a newly discovered language. The feature may still be important. To the extent that it is, its absence in the new language is a typological fact of importance about the language.

Conversely, if some feature is indeed universal, then it is taxonomically irrelevant.


Here is an example that illustrates both 1.4 and 1.5. It was at one time assumed that all languages distinguish between nouns and verbs -- by some suitable and sufficiently formal definition of those therms. One form of this assumption is that all languages have two distinct types of stems (in addition, possibly, to various other types), which by virtue of their behavior in inflection (if any) and in syntax can appropriately be labeled nouns and verbs. In this form, the generalization is rendered invalid by Nootka, where all inflectable stems have the same set of inflectional possibilities. The distinction between noun and verb at the level of stem is sufficiently widespread that its absence in Nootka is certainly worthy of typological note (1.5). But it turns out that even in Nootka something very much like the noun-verb contrast appears at the level of whole inflected words. Therefore, although Nootka forces the abandonment of the generalization in one form, it may still be that a modified form can be retained (1.4).

 The Port Royal Grammar constituted both a putative description of language universals and the basis of a taxonomy. The underlying assumption was that every language must provide, by one means or another, for all points in the grammatico-logical scheme described in the Grammar. Latin, of course, stood as the origin in this particular coordinate system. Any other language could be characterized typologically by listing the ways in which its machinery for satisfying the universal scheme deviated from that of Latin. This classical view in general grammar and in taxonomy has been set aside not because it is false in some logical sense but because it has proved clumsy for many languages: it tends to conceal differences that we have come to believe are important, and to reveal some that we now think are trivial.


Given a taxonomy, if we find that languages of the most diverse types nonetheless manifest some feature in common, that feature may be important. It is not apt to be, however, if it is an easily diffusible item. Thus the fact that many languages all over the world have phonetically similar words for 'mama' is more significant than a similar widespread general phonetic shape for 'tea'.

In allowing for diffusion, we must also take into consideration that even features that do not diffuse readily may spread from one language to others when the speakers of the languages go through a long period of intimate contact. This fact, if no other, would seem to render suspect any generalizations based solely on the languages of Western Europe. And it is true that some such generalizations are refused by the merest glance at an appropriate non-European language. But contrastive study based exclusively on European languages also has a merit: our knowledge of those languages is currently deeper and more detailed than our knowledge of languages elsewhere, so that generalizing hypotheses can also be deeper. They may be due for a longer wait before an appropriately broad survey can confirm or confute them, but they are valuable nonetheless.


1.9 A universal feature is more apt to be important if there are communicative systems, especially nonhuman ones, that do not share it.

It may seem peculiar at first to propose that we can learn more about human language by studying the communicative systems of other animals; but a moment's reflection is enough to show that we can only know what a thing is by also knowing what it is not. As long as we confine out investigations to human language, we constantly run the risk of mistaking an "accidental" universal for an "essential" one -- and we bypass the task of clearly defining the universe within which our generalization are intended to apply. Suppose, on the other  hand, that after discovering that a particular feature recurs in every language on which we have information, we find it lacking in some animal communicative system. In some cases, this might lead us to add the feature to our defining set of language. In any case, this seems to be one way of trying to avoid triviality in the assembling of our defining set.


1.10 The problem of language universals is not independent of our choice of assumptions and methodology in analyzing single languages.





5.2 Phonemes are not fruitful universals.

We can, indeed, speak quite validly of phonemes in the discussion of any language, but their status in the hierarchy of phonological units varies from one language to another, and also, to some extent, through varying preference or prejudice of analysts. The status of phonological components, on the other hand, is fixed once and for all by definition -- phonological components are the minimum (not further divisible) units of a phonological system. Given that all phonological patterning is hierarchical, the exact organization of the hierarchy, varying from one language to another, becomes a taxonomic consideration of importance, but not the basis of a generalization in the present context.

There are certain languages of the Caucasus (Kuipers, 1960) where one can, if one wishes, describe the phonological system in terms of perhaps a dozen phonological features organized into some seventy or eighty phonemes, which in turn occur in about twice that many syllables. Each syllable consists of one of the seventy-odd consonant phonemes, followed by one of the two vowel phonemes. It seems clear in such a case that the vowel "phonemes" are better regarded simply as two additional phonological features, so that a unit such as /ka/ is just a phoneme -- or, alternatively, that the term "phoneme" be discarded and one discuss the participation of features directly in syllables. Either way, one does not need both the term "phoneme" and the term "syllable". The case may be extreme, but it is real, and underscore the importance of the "anti-universal" given as 5.2.



5.5 Sound change is a universal. It is entailed by the basic design features of language, particularly by duality of patterning.

Note Conflict Fossils Histograms


Skewed Histograms

Paleontologists have done many computer simulations in order to gain a feel for the range of outcomes of random walks in biodiversity, with both equal and unequal rates of speciation and extinction. Some simulated groups expand to the point of swamping the computer's memory; other groups go extinct rather soon. Extinction of a group is most common when the group starts out small, just as a casino gambler is most likely to go broke quickly if he starts with a small stake, close to the absorbing boundary.

In evolution, a group such as a genus of family must, by definition, start with a singles species. For the fledgling group to survive, the founding species must speciate before it goes extinct. Because new evolutionary groups start small, they usually don't last long. This, in turn, yields an important facet of the history of life: most groups of species have life spans shorter than the average of all groups. Figure 3-2 shows a histogram of life spans of fossil genera. It has a skewed shape, with many short durations and only a few long ones.

The skewed (asymmetrical) shape of variation is typical of important biological properties germane to the extinction question. These include
- number of species in a genus
- life spans of species
- number of individuals in a species
- geographic ranges of species


http://sepkoski.com/sepvv02.cfm?XRPF=VVALEN

Distribution of Length of Lifespans of Genera
    Y-Axis: Number of Genera
    X-Axis: Lifespan of Genus


Fossils Histograms


Skewed Histograms

Paleontologists have done many computer simulations in order to gain a feel for the range of outcomes of random walks in biodiversity, with both equal and unequal rates of speciation and extinction. Some simulated groups expand to the point of swamping the computer's memory; other groups go extinct rather soon. Extinction of a group is most common when the group starts out small, just as a casino gambler is most likely to go broke quickly if he starts with a small stake, close to the absorbing boundary.

In evolution, a group such as a genus of family must, by definition, start with a singles species. For the fledgling group to survive, the founding species must speciate before it goes extinct. Because new evolutionary groups start small, they usually don't last long. This, in turn, yields an important facet of the history of life: most groups of species have life spans shorter than the average of all groups. Figure 3-2 shows a histogram of life spans of fossil genera. It has a skewed shape, with many short durations and only a few long ones.

The skewed (asymmetrical) shape of variation is typical of important biological properties germane to the extinction question. These include
- number of species in a genus
- life spans of species
- number of individuals in a species
- geographic ranges of species

In each category, the small "thing" is most abundant. Let me give another example. There are about 4,000 living species of mammals, grouped into about 1,000 genera. About half of these genera have only a single species, and about 15 percent have only two species. The numbers drop off smoothly (see Figure 3-3), so there are only a few genera with more than 25 species. The most speciose living mammal genus (a small insectivore) has about 160 species. The overall average is 4 species per genus (4,000/1,000), but because of the asymmetry, fully three-quarters of the genera have 1, 2, or 3 species and are thus below average.

Let me summarize a few of the foregoing points. For reasons that develop from both theory and observation -- some depending on the Gambler's Ruin problem -- we can make the following generalizations:

1. Most species and genera are short-lived (compared with the averages).
2. Most species have few individuals.
3. Most genera have few species.
4. Most species occupy small geographic areas.

Skewed variation is extremely common in nature. Strangely, however, most of us have been trained to believe that variation in natural phenomena is bell shaped, having just as many items above as below the average -- whether we are talking about heights or weights of people, wealth  or baseall averages. Nothing could be further from the truth.

Classic examples of skewed variations include incubation times of infectious diseases and life expectancies of cancer patients. In both, the majority of cases fall below the average, because the average is constructed by summing many short time intervals and a few long ones. It would be far better to use the median time -- that time exceeded by half the individuals.

Of course, bell-shaped curves (called normal or Gaussian distributions by mathematicians) do sometimes occur in nature. It is just that other shapes are more common. Statisticians wrestle with this problem because many of the best statistical tests are designed for the bell-shaped curve. Often they avoid the problem by transforming the raw data -- that is, distorting the scale of measurement so that they can treat the results as if they had a bell-shaped distribution. One such transformation that sometimes works is to convert all measurements to their logarithms (or even square roots). If the transformed numbers have a bell-shaped distribution, the analyst can proceed with tests that assume this shape.


Other Models

The Gambler's Ruin problem has led us to generalizations about species that are relevant to the extinction problem. However, many of the patters, especially the skewed distributions, can also be approximated by processes having nothing to do with gambling or biology.

Suppose you take a stick that is 100 inches long and break it at 25 random points -- not favoring the middle or any other part. When you are done, you will have 26 short sticks. Now, measure and count the short sticks and construct a histogram. The shape of variation in stick length will took very much like those I have shown for species and genera: a hump or spike to the left and a long tail extending to the right. Figure 3-4 shows the results of a computer simulation.

Variation in population size of cities in the United States shows a pattern like this, as do many other things we can measure or count. The so-called broken stick model is one of several that have been applied to these patterns, and many attempts have been made to find out which model makes most sense or fits the observations best. For the purposes of this book, the important thing is that many of the distributions are skewed. They are not even close to the symmetrical, bell-shaped curve that we have all heard about.

One lesson about extinction to be learned from this is that some plants and animals are much more likely, a priori, to go extinct than others. The majority of species living today have small populations and live in restricted geographic areas. There are the ones we rarely see. The abundant and widespread species are commonly seen but are surprisingly few in number. For this reason, it is possible to write useful field guides to mammals and insects in volumes of manageable size. It stands to reason that when things get environmentally tough, either biological or physically, the many rare species are the most vulnerable to extinction. So, when we say that a given extinction event eliminated 40 or 80 percent of biodiversity, we should also say which 40 or 80 percent. The significance of the event will depend heavily on whether the victims were abundant, cosmopolitan species or local endemics.

Memorandum Concerning Language Universals


Memorandum Concerning Language Universals
Greenberg, Osgood and Jenkins

Underlying the endless and fascinating idiosyncrasies of the world's languages there are uniformities of universal scope. Amid infinite diversity, all languages are, as it were, cut from the same pattern. Some interlinguistic similarities and identities have been formalized, others not, but working linguists are in many cases aware of them in some sense and use them as guides in their analyses of new languages.

Since language is at once both an aspect of individual behavior and an aspect of human culture, its universals provide both the major point of contact with underlying psychological principles (psycholinguistics) and the major source of implications for human culture in general (ethnolinguistics).

The tendency toward symmetry in the sound system of languages has psycholinguistic implications. The articulatory habits of speakers involved in the production of the phonemes consist of varied combinations of certain basic habits, those employed in the production of the features. This appears, for example, in language acquisition by the child. At the point in the development of the English-speaking child that he acquires the distinction between b and p based on voicing versus non-voicing  he simultaneously makes the distinctions between d and t, g and k, and other similar pairs. In other words, he has acquired the feature, voicing versus non-voicing, as a unit habit of motor differentiation.


Logical Structure of Universals
Unrestricted universals

These are characteristics possessed by all languages which are not merely definitional; that is, they are such that if a symbolic system did not possess them, we would still call it a language. Under this heading would be included not only such obvious universals as, for example, that all languages have vowels, but also those involving numerical limits, for example, that every language has at least two vowels. Also included are universally valid statements about the relative text or lexicon frequency of linguistic elements.


Universal implications
These always involve the relationship between two characteristics. It is asserted universally that if a language has a certain characteristic, it also has some other particular characteristic, but not vice versa. For example, if a language has a category of dual, it also has a category of plural but not necessarily vice versa.


Restricted equivalence
This is the case of mutual implication between characteristics which are not universal. For example, if a language has a lateral click, it always has a dental click and vice versa. Equivalences of more frequently appearing logically independent characteristics are difficult to find. They would be of great interest as indicating important necessary connections between empirically diverse properties of language.


Statistical universals
These are defined as follows: For any language a certain characteristic has a greater probability than some other (frequently its own negative). This includes 'near universals' in extreme cases. Only Quileute and a few neighboring Salishan languages among all the languages of the world lack nasal consonants. Hence we may say that, universally, the probability of a language having at least one nasal consonant is greater (in this instance far greater) than that it will lack nasal consonants.

terça-feira, 19 de fevereiro de 2013

corpus and databases

A list of a few corpus and databases that might be at hand...

1. Project Gutenberg
http://www.gutenberg.org/
is a volunteer effort to digitize and archive cultural works, to "encourage the creation and distribution of eBooks". As of February 2013, Project Gutenberg claimed over 42,000 items in its collection.

2. WordNet
http://wordnet.princeton.edu/wordnet/
is a large lexical database of English. Nouns, verbs, adjectives and adverbs are grouped into sets of cognitive synonyms (synsets), each expressing a distinct concept.

3. SWITCHBOARD
http://www.ldc.upenn.edu/Catalog/catalogEntry.jsp?catalogId=LDC97S62
http://groups.inf.ed.ac.uk/switchboard/index.html
http://www.ldc.upenn.edu/Catalog/readme_files/switchboard.readme.html
is a corpus of spontaneous conversations collected at Texas Instruments, it includes about 2430 conversations averaging 6 minutes in length; in other terms, over 240 hours of recorded speech, and about 3 million words of text, spoken by over 500 speakers of both sexes from every major dialect of American English.

4. CORPUS... the open parallel corpus
http://opus.lingfil.uu.se/
OPUS is a growing collection of translated texts from the web. In the OPUS project we try to convert and align free online data, to add linguistic annotation, and to provide the community with a publicly available parallel corpus. OPUS is based on open source products and the corpus is also delivered as an open content package. We used several tools to compile the current collection. All pre-processing is done automatically. No manual corrections have been carried out.

5. Google Ngram
http://books.google.com/ngrams
http://books.google.com/ngrams/datasets
Database of Ngram create from over 5 million books, in a time spam of 500 years, that were digitized by Google.

6. Wikipedia Data Dump
http://meta.wikimedia.org/wiki/Database_dump
Wikipedia content saved in XML format. Available in many languages.
Example of the XML format used: http://en.wikipedia.org/wiki/Special:Export/Moon_landing
Tools written in Perl to process MediaWiki dump files are available here: http://search.cpan.org/perldoc?Parse::MediaWikiDump