Thursday, 24 June 2010

Cultural Induciton is Hard

Chater and Christiansen (2010) argue that culturally transmitted systems such as language are easier to learn than natural systems because they have adapted to learner's biases, so their intuitions will likely be correct. Being a speaker of one of the morphologically most complex languages in the world, I'm not so sure...

Maggie Tallerman gave a keynote speech at this year's EvoLang conference. A section of it used grammatically acceptable and unacceptable sentences in Welsh to illustrate the point. As a somewhat lapsed Welshspeaker, whose knowledge of Mutation was never great, I was a bit embarrassed to find I wasn't sure if the examples were correct. Last month, too, Mike Dowman gave a talk at Edinburgh University emphasising how impressive it is that we all make the same grammaticality judgements about sentences we have never seen before.

I've long wondered whether this is the case with Welsh mutation. Consonant mutation, or 'Treiglo' in Welsh, occurs in many Celtic languages and is a terrible affliction for the second language learner. In a number of (grammatical) contexts, the initial consonant of a word (nouns and verbs) changes to another. For instance, 'kitchen' in Welsh is 'cegin' [k3gIn], but 'his kitchen' is 'ei gegin' [g3gIn] and 'her kitchen' is 'ei chegin' [X3gIn].

There are three forms of mutuation - soft, nasal and aspirate. The Wikipedia page on Welsh Mutation gives a broad overview. The contexts they apply in are extensive, for example:

  • Nouns after the preposition 'in'
  • After imperatives.
  • After the personal pronoun (my)
  • Singular feminine nouns after the definite article (but not words beginning in ll or rh)
  • In the negative form of verbs in the Short Future Tense
  • Masculine nouns after 'three' and all nouns after 'six'

The rules of mutation in old Welsh were much simpler: It only occured for feminine nouns after the definate article. However, presumaply by a process of analogy, the 'rules' spread and became more complex. And here's the point I'm trying to make: Welsh morphology may be so complex, and subject to so much change by analogy, that there is very little agreement between people.

I'm not saying that mutuation is unsystematic. There's even an automatic mutuation checker online. However, it always annoys me when syntacticians cite some examples that they haven't actually gone out and tested.

Now, this may be an intuition I have from school. I grew up speaking Welsh, both my parents speak Welsh and I went to a Welsh-medium nursery, primary and secondary school where we were not allowed to speak English. Despite this, I'm not a confident speaker, especially after 7 years outside of Wales. I was particularly bad at mutation (although my English spelling was, and still is, equally as bad). We had it drilled into us with tables and excersises, but I still can't really do it properly. This is partly because of the minority langauge status of Welsh, and the fact that everybody spoke English as a form of rebellion. The influence of English has been felt in other areas of Welsh such as Subject-Verb order, too.

However, I've always felt guilty about not being able to speak the mothertongue properly. Then I became a linguist and found a way out: All along, my teachers had been prescribing language, and that prescription was a few generations old. In terms of language evolution, the language the children speak is the correct language (the descriptivist approach).

Long story short, I conducted my own grammaticality judgement experiment. I couldn't find a grammaticality judgement for Welsh mutation online, nor any information about how well learners pick it up (please send me links if you know of any!). Neither am I a trained syntactician, so I have no real idea how to do an experiment, nor do I have access to money to employ participants.

So I decided to do it in the form of a facebook quiz, using Quibblo. I found example sentences from an instructional pamphlet and took one from each major context. I then created alternative mutations for each sentence. Participants were presented with an English equivalent of each sentence, and asked to indicate which sentences they thought were correct. Participants could make more than one choice.

You can take the quiz here and view the results here.

12 people participanted, mainly schoolfriends since this was distributed via facebook. This is a good point rather than a bad point, since we're more likely to have been exposed to the same linguistic environment (and indeed been in frequent contact). However, many of these may be, like me, somewhat out of practice. On the other hand, 11 indicated they were 'fluent' and 1 'intermediate' speakers of welsh. All but one came from South Wales. All particpants chose only one answer in 17 out of 34 questions.

As it turned out, Quibblo wasn't a very good choice - it only records totals, not individual records of participant's choices. Anyway, here's some analysis.

For each setence, I worked out the average agreement. This is the likelyhood of any two people agreeing that at least one form was correct. For all sentences, the average agreement was 67.1% For sentences where participants only selected one answer, the average agreement was 60.9%.

Some sentences recieved 100% agreement - for the sentence 'the boy', all 12 participants chose the prescribed 'dy fachgen'. However, the sentence 'the girl' was split with two thirds going for the prescribed 'y ferch' and one third going for 'y merch'. On the other end of things, 9 participants chose 'dydd mawrth' for the meaning 'on tuesday', when the prescribed form is 'ddydd mawrth', which only 3 participants chose.

For the setences with more than two possible choices, the choices are spread. For the sentence 'I read a good book', 6 different options were selected with an average agreement of 54.5%. The worst agreement was for the sentence 'the sixth girl' with participants agreeing on average 25.5% of the time between 4 options. For this sentence, 12 participants chose 16 options, meaning that some participants thought at least two options were correct (I don't have the exact data on who chose what).

I put some tricky questions in to see what would happen. The first was designed to test whether adjacent adjectives should be mutated. That is, adjectives which follow a singular feminine noun mutate, but it's not clear whether a following adjective should too. Participants were given the sentence 'a big, tall, good girl' and given the option to mutate none, the first adjective, first and second adjectives and all adjectives. 3 participants chose to mutate only the first while 5 chose to mutate all adjectives (3 chose to mutate none and 1 chose to mutate two). The agreement was 24.2%.

The second tricky question involved loanwords. Nouns after a conjunction mutate, so participants were given the sentence 'gin and tonic' and the options 'jin a tonic' and 'jin a thonic'. Agreement was slightly better here at 84.8% in favour of the prescribed (and attested) mutated form, although one person thought both were correct.

The final one involved analogy and loanwords. I heard someone mutate 'chips', which has a [ tʃ ], which doesn't exist in Welsh to the voiced equivalent [ dʒ ]. There is no prescription here, but this makes perfect sense if mutation really does spread by analogy. Presented with the sentence 'a bag of chips', 5 participants voted for the unmutated variant and 10 for 'bag o jips' (83.3% agreement).

All in all, enough to make my old Welsh teachers weep. Of course, the sample is probably skewed and can't be verified and some might have looked the answer up etc. But part of the point of this is that, for very simple sentences, people should be choosing the same sentences.

In a forthcoming special issue of Cognitive Science (preview here), Nick Chater and Morton Christiansen argue that learning cultrually-transmitted systems is easier than learning about the natural world because cultural systems will be adapted towards a learner's biases. Therefore, a learner's intuitions and guesses are likely to be correct. That is, it's easier to co-ordinate your behaviour with other people than it is to be right about the world (an alternative name for the paper could have been 'Language Evolution: Specifically not the hardest problem in Science').

It's a great paper, and argues for my PhD thesis - that language acquisition should be looked at in the light of language evolution. However, cultrual induction may not be easier than learning about the natural world if everybody is doing something different. Consider the participants in my experiment: A child learning from them faces sources of cultural variants that not only contradict each other half of the time, but contradict themselves part of the time. At least mass and other physical attributes are Universal - gravity doesn't work differently in North Wales. However, since the data cultural learners are presented with comes from multiple people who themselves may have had different and non-overlapping sources of input, cultural learning may be pretty tricky after all.

So, there may be space in this topic for my PhD thesis: The social structure of cultural learners will have a huge impact on the ease of Cultural induction, and thus on the pressures and eventual forms of language.

Nick Chater & Morton H. Christiansen (2010). Language Acquisition Meets Language Evolution Cognitive Science

Tuesday, 25 May 2010

Evolutionary approaches to Bilingualism

I recently gave a talk at the University of Edinburgh LEL Postgraduate Conference. It was my first ever talk and it really forced me to figure out what I'm supposed to be studying! Here's a video of my talk:







Frank MC, Goodman ND, & Tenenbaum JB (2009). Using speakers' referential intentions to model early cross-situational word learning. Psychological science : a journal of the American Psychological Society / APS, 20 (5), 578-85 PMID: 19389131

Hunag, Y. (2009). Supporting Meaningful Social Networks Technical Report, ECS, University of Southampton
pdf

Healey, E. and Scarabela, B. (2009). Are children willing to accept two labels for one object? Proceedings of the Child Language Seminar. University of Reading.

Byers-Heinlein K, & Werker JF (2009). Monolingual, bilingual, trilingual: infants' language experience influences the development of a word-learning heuristic. Developmental science, 12 (5), 815-23 PMID: 19702772

Thursday, 13 May 2010

E-Coli, Linux and Language

A recent post on The Loom looks at a paper by Koon-Kiu Yang et al. which compares the hierarchical structures of the operating system Linux and the bacterium E-Coli. Really interesting analysis - and a good discussion on the blog.

I found it interesting that E-coli's structure is primarily lower-level 'workhorses' with relatively few master controllers. Linux on the other hand has a much larger percentage of high-level 'master' and 'middle manager' modules and reletively few 'workhorses'. Linux is designed while E-coli is evolved.

I’m wondering how linguistic systems would fit into this schema. What are the ‘workhorses’ and ‘master regulators’ of language? There are many more ‘low-level’ words that refer to things than ‘higher level’ syntactic structures. This would make it like e-coli.

On the other hand, there are relatively few ‘low level’ phonemes and very many ‘high level’ concepts. This would make it more like Linux.

Maybe language has more ‘middle managers’ than anything else?

Answering this may give an insight into how ‘designed’ language is, as opposed to ‘evolved’.


Yan, K., Fang, G., Bhardwaj, N., Alexander, R., & Gerstein, M. (2010). Comparing genomes to computer operating systems in terms of the topology and evolution of their regulatory control networks Proceedings of the National Academy of Sciences DOI: 10.1073/pnas.0914771107

Tuesday, 11 May 2010

Mutual Exclusivity biases in cross-situational learning: A comparison between monolingual and bilingual corpora

This report focuses on models of cross-situational learning and how current models compare when exposed to real monolingual and bilingual input. Several model types were evaluated against two transcribed videos of parent-child interaction, one being monolingual and the other being bilingual. Children have been shown to demonstrate a Mutual Exclusivity (ME) bias (Markman and Wachtel, 1988) during word learning. Frank et al. (2009) showed that their model also exhibited Mutual Exclusivity (ME) behaviour after learning from a monolingual corpus of contexts. The current study takes the same model but with bilingual input and asks whether the same behaviour is exhibited.

Method

Frank et al. (2009) provide a transcribed video of monolingual parent-child interaction coded for use in cross-situational learning. An equivalent bilingual corpus was looked for. The main criterion was a roughly equal number of utterances in both languages. The CHILDES database has suitable resources. A recording from a study by Yip and Matthews was selected (see CHILDES, 2010). The child in question was a native bilingual from birth. Her mother was a native Hong-Kong Cantonese speaker and her father was a native speaker of British English. She was 2;11 in the chosen recording. There are 967 utterances, 48% of which are Cantonese and 52% are English. The objects visible in the video were added to the transcription, along with the mapping between words and objects and the referential intentions of the speakers. The coding scheme was adopted from Frank et al. (2009). The code for the models was supplied by Frank et al. (from Frank,2010).

Results

The performance of different models were analysed in two ways. Firstly, the best estimated lexicon (word-object mappings) of each lexicon was evaluated against a gold-standard lexicon in terms of precision, recall and the resulting F-score.Secondly, the models were asked to guess the intended referent of each utterance-object context.Tables 1 and 2 show the lexicon results for the monolingual and bilingual corpora respectively. Frank et al.’s model returns the highest F-scores in both cases.This is largely due to an advantage in precision, likely stemming from the modeling of non-referential words. Frank et al.’s model returned a word-object mapping for the bilingual corpus with a precision of 0.31 and a recall of 0.27, giving an F-score of 0.29. This is lower than the score for the same model on the monolingual corpus. This could be due to the referential uncertainty (independent from amount of synonymy) in the bilingual corpus being higher.The results for the referential intentions, shown in Tables 3 and 4, have different trends. For the monolingual case, the precision of Frank et al.’s model allows it to outperform the other models. However, it performs relatively poorly in the bilingual case, with the Conditional Probability model performing best. However,all models perform with little precision and recall, suggesting that the task is harder. With more data, results might be different.




Mutual Exclusivity

After the model had processed the corpus, it was presented with a mutual exclusivity task and the relative likelihood of several interpretations were measured.In the task, the model was presented with a context with a new object (e.g. a dax) and a familiar object (a bird for the monolingual model and an orange for the bilingual model) and a new word (e.g. ”dax”). The probabilities were calculated for the model linking the new word with neither object (i.e. it considers the word non-referential), linking it with the new word, linking it with the old word or linking it with both. Figure 1 shows the results of the task with the results for a ’monolingual’ model for comparison.

The monolingual results are re-calculated for this study, so differ slightly from those reported in Frank et al. (2009).The results for the monolingual and bilingual models have the same trend - both rank the possible situations in the same order of likelihood. The most likely situation is the new word being linked to the new object, honouring mutual exclusivity. The second most likely situation is that the word refers to neither object.Intuitively, one would expect a bilingual to be more likely to consider that the new word was another word for the familiar object. Indeed, the bilingual model does consider this possibility relatively more likely than the monolingual model. However, the model still considers neither mapping to be more likely than an extra synonym. This may mean that, given an additional cue (e.g. pragmatic), the bilingual would be more ready to accept a synonymous interpretation. This is an empirical question.

The Prior Probability

The prior probability is simply the number of mappings in the hypothesised lexicon, modulated by a fixed parameter (alpha). This represents a preference for smaller lexicons. This means that a hypothesis which results in the lexicon with fewest mappings will receive the highest prior probability. With the default parameter (alpha = 7), the Mutually Exclusive preference (for DAX-dax) beats the preferences for the original mapping (map neither word to the unfamiliar object),both mappings and the mapping of the unfamiliar object with the familiar name.However, this ranking depends on the lexicon size bias (alpha) parameter. With a low alpha, the most likely mapping is the ME mapping. With a higher alpha, the most likely mapping is the original mapping (see Figure 2).

The same trend also exists between the preference for the ME mapping and both mappings, although the preference for both mappings does not overtake the preference for the ME mapping (see Figure 3).
The explanation is as follows: The original mapping receives a high prior probability because it doesn’t increase the size of the lexicon. However, the likelihood of experiencing a non-referential word is low, leading to a total probability that favours the ME mapping over the original. Assuming a larger lexicon (decreasing alpha), the relative increase in lexicon size is smaller, tipping the balance between the original and ME mapping preferences.Interestingly, the likelihood of choosing both mappings overtakes the original mapping when alpha is less than 1 (see figure 4).
That is, the likelihood of assuming both mappings increases when the prior is set to less than the number of word-object mappings in the lexicon. Such a setting makes sense for a bilingual (who have up to twice as many mappings as bilinguals) because it represents the number of concepts. Put another way, by compensating for the additional synonymy in bilingual input, the likelihood of assuming both mappings increases.The dependence of the ME experiment results on alpha is acknowledged by Frank et al.:

“Note that there is some parameter dependence in our models fit to the mutual exclusivity situation. Depending on the size of the corpus,it might be the case that the prior disadvantage of adding a word to the lexicon would not be outweighed by the increase in corpus likelihood caused by learning a new word. This fact makes a developmental prediction: in early development, when very few words are known,inferences about mutual exclusivity should be weaker.”
Supporting Information for Frank et al. (2009), p. 13.

This prediction is borne out in some studies (Merriman and Bowman, 1989; Frankand Poulin-Dubois, 2002; Merriman et al., 1993). However, Markman and Wachtel (1988) found that the ME constraint weakens over time, with older children showing less of a bias, while Deak et al. (2001) find no change.The issue here is the size of the lexicon. Bilingual children may know more words than monolinguals, but it may be more accurate to judge the lexicon size by the size of one language’s lexicon.The model does not provide a mechanism for modulating the lexicon size prior parameter during learning. Currently the prior is modulated by the alpha parameter and the number of mappings, meaning that adding new mappings is dis-preferred. Bilinguals will have a higher number of mappings, altering their prior probabilities. However, this does not lead to qualitative differences in the mutual exclusivity experiment.The motivation for modulating the prior by the number of mappings is mainly to simplify the model.

“We chose a prior probability distribution that favored parsimony,making lexicons exponentially less probable as they included more word-object pairings ... The choice of a simple prior puts most of the work of the model in the likelihood term ... hence, the likelihood term captures the learners assumptions about the structure of the learning task.”
Frank et al., 2009, p. 579

That is, the decision is driven by the statistical, computational approach to the formal problem rather than being psychologically motivated. Therefore, the interpretation is that mutual exclusivity behaviour stems from the child’s unwillingness to learn new signal-meaning mappings. This seems a little circular - children prefer not to extend mappings from familiar words to unfamiliar objects because they prefer not to extend mappings. It also seems to go against children’s obvious ability and motivation for learning new words and meanings. Several solutions which would make the prior more sensitive to the input involve incorporating the number of concepts, the number of words or the amount of synonymy (proportional to the number of words in the lexicon divided by the number of concepts). However, the nature of the model now changes - we are using it to test specific hypotheses about mutual exclusivity, judged against empirical data,rather than seeing if mutual exclusivity ’falls out’ of more basic assumptions.

Concept-based Prior

The mapping-based prior was biased towards a monolingual mode. The model was altered so that the prior was negatively related to the number of objects in the lexicon. This represents the number of concepts for which the child knows words. The model was run on the bilingual corpus and returned a lexicon with a precision of 0.05, a recall of 0.41 and a resulting F-score of 0.09. The model was also run on the monolingual corpus again, returning a precision of 0.05, a recall of0.79 and an F-score of 0.09. For both monolingual and bilingual corpora, the recall of this model is better than for a mapping-based prior, but the precision is much worse. That is, the model overestimates the number of word-concept mappings. In fact, the models accumulated many hundreds of word-concept mappings for tens of objects (Monolingual: 551 mappings for 22 objects and 419 words; Bilingual:641 mappings for 55 objects and 598 words). The models have failed to acquire a useful vocabulary.However, running the Mutual Exclusivity experiment again, the relative ranking of the preferences has changed. Although the ME mapping is still favoured, the next preferred interpretation is to make both mappings (rather than neither, see Figure 5). However, this difference is exhibited with both monolingual and bilingual input data. By neutralising the difference in the prior, the corpus likelihood now plays a bigger role, leading to a difference in the preferences.




How ’Monolingual’ is the Monolingual corpus?

Although the monolingual corpus is taken from a carer speaking one language, the lexicon the model learns contains synonymy. In fact, for the 15 objects it learned words for, 8 had more than one associated word. For half of these 8 objects, all synonyms were appropriate (e.g. ’bird’ and ’birdie’ to describe the object ’duck’),but half were not appropriate. In other words, the model accommodates synonymy.The original Mutual Exclusivity experiment in Frank et al. was done with the object ’bird’, which had one associated word. The ME experiment was applied for all words that the model learned from the monolingual corpus. There were no significant differences between the posterior probabilities for any of the situations (DAX-dax, Both etc.) for synonymous mappings versus non-synonymous mappings. This holds for both the original and the concept-based prior.

Conclusion

Frank et al.’s model can be used to model word learning in bilinguals. There are some quantitative differences in the ME behaviour of models run on monolingual and bilingual corpora. However, no qualitative differences were found. Even when the prior bias for minimising the number of mappings was neutralised, both models still preferred to map the new object with the new word.

Next Steps
The results are inconclusive, but may reflect the limited data. I suggest that synthetic corpora would make the dynamics more clear. Very simple cross-situational learning corpora could be created with varying amount of ’bilingualism’.

References

Frank MC, Goodman ND, & Tenenbaum JB (2009). Using speakers' referential intentions to model early cross-situational word learning. Psychological science : a journal of the American Psychological Society / APS, 20 (5), 578-85 PMID: 19389131

Byers-Heinlein K, & Werker JF (2009). Monolingual, bilingual, trilingual: infants' language experience influences the development of a word-learning heuristic. Developmental science, 12 (5), 815-23 PMID: 19702772

Deák GO, Yen L, & Pettit J (2001). By any other name: when will preschoolers produce several labels for a referent? Journal of child language, 28 (3), 787-804 PMID: 11797548

Frank, I., & Poulin-Dubois, D. (2002). Young monolingual and bilingual children's responses to violation of the Mutual Exclusivity Principle International Journal of Bilingualism, 6 (2), 125-146 DOI: 10.1177/13670069020060020201

Markman EM, & Wachtel GF (1988). Children's use of mutual exclusivity to constrain the meanings of words. Cognitive psychology, 20 (2), 121-57 PMID: 3365937

Merriman WE, & Bowman LL (1989). The mutual exclusivity bias in children's word learning. Monographs of the Society for Research in Child Development, 54 (3-4), 1-132 PMID: 2608077

Merriman WE, Marazita J, & Jarvis LH (1993). Four-year-olds' disambiguation of action and object word reference. Journal of experimental child psychology, 56 (3), 412-30 PMID: 8301246

Healey, E. and Scarabela, B. (2009). Are children willing to accept two labels for one object? Proceedings of the Child Language Seminar. University of Reading.