Showing posts with label Lupyan Dale. Show all posts
Showing posts with label Lupyan Dale. Show all posts

Thursday, 1 April 2010

Cultural Variation and Social Networks

Children learn language from exposure to speakers in their social network. This learning influences the input that will be given to the next generation. The learning biases that an individual has will influence the way the language changes over generations (Kirby, Dowman & Griffiths, 2007). However, language also plays a part in constructing and maintaining social networks. Recent studies have suggested that the structure of the social network also has an effect on the how a language evolves. Gong and Wang (2010) find that different network types influence the evolution of linguistic categories in an artificial categorisation game. Lupyan & Dale (2010) find that the amount of contact with other communities, and a community's spatial dispersion influences the morphological complexity of a language.

I was wondering whether bilingual communities have different social network structures to monolingual communities. Real social networks are very difficult to construct, so I wanted to use some online social networking sites. Twitter seemed like an obvious choice because of it's simple API, and also because 'following' someone has a genuine connection on a user's linguistic input.

I aquired some data for Twitter users. The data includes the number of followers (indegree) and people being followed (outdegree), the user's location, the number of status updates sent by a user and the amount of time since the last update. The last two features can be used to filter out people who are not active participants. The location information is optional and may be as specific as GPS coordinates or as general as a country, or even just a timezone. Following from this, communities were defined by country. Data mining techniques will be used to automatically assign users to countries.

Ideally, we would want the following statistics: Average Degree, Clustering Coefficient, Average shortest Path length. However, this requires information on the specific links between users. However, this requires more time and resources, so this data was not collected for this report. This is not a trivial point, however, because users can follow people in other countries.

Next, data on the linguistic variance is needed. As I showed in a recent post, estimating the amount of bilingualism is difficult. The best source of information is Ethnologue, but numbers of speakers are underestimated, probably due to inadequate data for small linguistic communities. I decided to use two measures of bilingualism: The number of languages spoken in a country and the percentage of the population of a country that speak the majority language. A country's nominal per-capita GDP and number of internet users is also taken into account.

Data was collected for about 31,000 users, about 25,000 of which was usable (there have been databases of up to 2.7 million users with over a billion connections between them). Although Twitter allows very many followers, for practical purposes I filtered out users with over 1,000 followers or following over 1,000 other users. This left data for 17,444 users ffrom 119 countries.

Initial results suggest a negative correlation between the indegree for each country and the percentage of the country's population who speak the majority language (using log indegree, t = 1.88, df = 117, p = 0.06).

Using a linear regression, the indegree and outdegree are significant predictors of linguistic variation, even when the effects of population size and access to the internet are partialled out (R-squared = 0.19, F(4,17439) = 1063, p <0.01; t =" -2.04," p =" 0.04;" t ="2.54," p ="0.01). This was based on data for 17,444 users ffrom 119 countries. Statistics for countries were taken from CIA factbook, 2010. The analysis revealed a negative correlation between linguistic variance and indegree, but a positive correlation between linguistic variance and outdegree. The same qualitative results were found by using the number of languages spoken in a country. However, there is a positive correlation between the number of languages and both indegree and outdegree. I'm not sure how to interpret this yet, or whether any of it makes any sense. In the meantime, here's a pretty uninformative map of the world, coloured by average number of Twitter friends. Darker countries have users with a higher average number of friends.


Lupyan G, & Dale R (2010). Language structure is partly determined by social structure. PloS one, 5 (1) PMID: 20098492
Kirby S, Dowman M, & Griffiths TL (2007). Innateness and culture in the evolution of language. Proceedings of the National Academy of Sciences of the United States of America, 104 (12), 5241-5 PMID: 17360393


Monday, 25 January 2010

Language Structure and Social Structure

Last week saw the publication of Lupyan & Dale (2010) (also discussed here). It's an analysis of languages from the World Atlas of Langauge Structures (WALS, quite fun to play around with), showing that languages spoken by more people tend to be less morphologically complex (fewer cases, fewer inflections, more of a tenancy to express things using separate words rather than with morphology). It's hypothesised that this is because a greater, more dispersed population will contain more second-language learners, therefore the language will tend to change to be easier to learn by adult learners, who seem to find morphology more difficult to learn than child learners.

The number of speakers in a language certainly varies a great deal. Lupyan & Dale emphasise this by pointing out that the median number of speakers in a language is 7,000 while the mean is 828,000. The number of second-language learners is also not a trivial influence in modern times. For example, only about 30% of English speakers are native speakers. Even more extreme is Malay with only 15% of speakers learning it as a first language.

In comparison, Siberian Yupik Eskimo has essentially no non-native speakers (from supporting data). In such communities, there is a greater pressure for the language to be learnable by children, therefore it retains morphological complexity.

The changes introduced by adults are likely to be acquired by children. Thus, Lupyan and Dale talk about exoteric and esoteric languages. The changes introduced by adults are likely to be acquired by children. Although the paper also takes into consideration the fact that children do not always learn from their parents:
It has been argued that there is no automatic transmission of the “mother tongue” from parents to offspring (47). For example, in a survey of 188 individuals in Senegal who listed Bambara as their native language, Bambara was the father’s native language in 16%, the mother’s in 19%, the native language of both parents in 26%, and the native language of neither parent in 39% (47).
Interestingly (for me at least), Welsh is pointed out in the regression graph and seems to fit the pattern - it has relatively few speakers and a high complexity score (calculated by "summing the number of features for which each language relies on lexical versus morphological coding and subtracting the total from zero"). While Welsh speakers were dispersed throughout Britain in the 8th century (and a Welsh speaking colony remains in Argentina), there are probably more child speakers than adult speakers currently. However, contact with English has introduced a bias for lexical forms over morphological forms (at least on my schoolyard - consider "Rydw i'n ysgrifennu" vs. "Ysgrifennaf"). Coupled with a pressure from language conservation groups to make Welsh more accessible to learners (see my post here), Welsh may indeed become less morphologically complicated.

Also interesting is the comparison with language data from Ethnologue, noting that the WALS seems to under-represent sub-Saharan Arfican languages (here). It appears that the Ethnologue is still the best source of data on population sizes, even though it appears not to be good enough to calculate the number of bilingual speakers (here).

One question is to what extent this phenomenon of exocentrism is a modern one. Large social groups and extensive travel only really took off in the last thousand years, so it's not clear if these dynamics can be scaled backwards in time to aid study of the evolution of language. It's been argued that human evolution allowed larger group sizes in comparison to chimpanzees (Isbell & Young, 1996), did this force communication to become more structured? Was language primarily lexically-based and used by adults?

Perhaps, however, it could be applied to animals - do songbirds living in larger, migratory flocks have simpler song morphology than those in smaller, more localised ones? The complexity of the domesticated finch is certainly more complex than its wild descendants (Honda & Okanoya, 1999). Could simply reducing the number of 'speakers' be an influence?

It's a fascinating article, and has already stimulated a lot of debate.



Lupyan G, Dale R (2010). Language Structure Is Partly Determined by Social Structure PLoS ONE, 5 (1) : 10.1371/journal.pone.0008559

Honda, E., & Okanoya, K. (1999). Acoustical and Syntactical Comparisons between Songs of the White-backed Munia (Lonchura striata) and Its Domesticated Strain, the Bengalese Finch (Lonchura striata var. domestica) Zoological Science, 16 (2), 319-326 DOI: 10.2108/zsj.16.319

Isbell, L. & Young (1996). The evolution of bipedalism in hominids and reduced group size in chimpanzees: alternative responses to decreasing resource availability Journal of Human Evolution, 30 (5), 389-397 DOI: 10.1006/jhev.1996.0034