Showing posts with label language. Show all posts
Showing posts with label language. Show all posts

Monday, June 06, 2016

A methodology for analysis of pop lyrics - problems

Song lyrics are typically divided into blocks of text, with the blocks corresponding to musical sections of the song – verses, chorus, bridge. These typically consist of more than one sentence, and thus are analogous in rank scale to paragraphs. The first, and possibly most obvious, problem, in analysing song lyrics is repetition - the repetitiveness is one of the distinctive features of song lyrics. It has the effect of significantly increasing the word count without adding to the semantic content. It also has the potential to distort and bias word frequencies, particularly in the case of the most repetitive songs. Previous researchers have adopted different approaches to the issue of repetition. Kreyer andMukherjee chose to keep them all; Petrie,Pennebaker and Sivertsen chose to eliminate the third and subsequent occurences. (Incidentally, that last one is a really interesting paper - "A linguistic analysis of the Beatles".) In my project, I downloaded lyrics from lyric databases, and the files I downloaded generally included all repetitions, as they are written in such a way as to permit people to follow the song from beginning to end. However, I decided to produce a corpus in which I deleted blocks of text that were repeated unchanged in their entirety. Where there were changes, both blocks were kept. The effect of this was to reduce the average number of words per "song" from 343 to 220, a reduction of 35%. The justification that underlay this was that I was interested in exploring linguistic features. Having noted the scale of the repetition, there was little need to explore it further.

A second issue is that the language used in song lyrics often diverges from "standard English", both in terms of word choice and grammar. This means that it's necessary to come to a decision about how to write it down. Should I write "ooh" or "oooh" or "oooooooooh"? Should I stick with the official version, and end up with a range of different versions of a word that is functionally the same? I didn't come to an answer in my original work. I think the best approach is to attempt to standardise as far as possible, but record the extent to which changes were made.

Another issue is that of "definitiveness". The definitive versions of lyrics are likely to be found either on a band's website, or if not electronically, on an album sleeve notes. It is much more common for bands to make their lyrics available in this way than it was thirty years ago. The downside of this is that it takes substantially more research to get hold of the lyrics, compared with raiding a lyrics database. However, most such databases are "widely collaborative" enterprises, with people contributing lyrics as they see fit, which may or may not be subsequently corrected by other people if they contain errors. For my project, I used a specific lyric database as a starting point. However, if the lyrics were unconvincing, I would then check it against other databases or the band's own website if one was available.

Thursday, May 22, 2014

"Language versus literature" in the pulpit

I have nearly finished my degree in English Language and Literature. I have enjoyed pretty much all of it (though writing about the Benin Bronzes was pretty painful), and it is proving to be the jump-off point for lots of different reflections.

One is in relation to what happens in preaching, and Bible teaching. Frequently, teaching from the Bible can sound like literary analysis. The teacher takes a text, links it (apparently arbitrarily) with other texts, makes connections (apparently arbitrarily) with some of his own ideas and perspectives, makes (apparently arbitrary) assumptions about different aspects and shades of meaning, and draws (apparently arbitrary) conclusions. This is highly consonant with where we are culturally. From a literary point of view, there's a strong strand which says that meaning is not inherent in the text itself: it is imposed on the text by the listener/reader - hence, we can have black, or gay, or Marxist, or green readings of texts that apparently have little otherwise to do with those perspectives. But if one person derives a specific meaning from a text, it is quite possible that another person might derive a meaning which completely contradicts this. The effect of this understanding of the nature of the text and meaning is that any sense of authority of the text is completely undermined. The teacher explains a text - but this interpretation is just one amongst many; it only has force if you share his or her perspective; and if you don't, then you are free to ignore it. It raises the question of what exactly would be the point of Bible teaching - perhaps it's considered to be some shared existential experience which makes us part of the Christian community, but is not considered to have any real force.

However, this degree highlighted the fact that, in addition to the literature perspective to studying a text, there is also a language perspective. This was very interesting to come across - at various stages in the course, it became clear that the language approach was different. Writers on the language approach were reluctant to criticise their faculty co-members, but the divergence was clear. Firstly, they said, if you lose the idea of context, then you lose most of the meaning of a text. They talked in a Hallidayan way about register variables - field, tenor, mode. All of these have a bearing in understanding a text. And they said, with some deference to their colleagues, whilst different interpretations were possible, some were definitely preferable to others. In effect, whereas the literature approach puts the focus on the reader, the language approach places it back on the text and its purpose as originally written.

This will come as no surprise to Bible teachers from certain backgrounds. One of the thrusts in the Proclamation Trust approach, for example is to "take the listeners to Corinth". The literature approach takes words from 1 Corinthians, for example, disregards the context, and tries to go straight to understanding what it means to us. Proc Trust argue that to understand what it means to us, you need to understand what it meant to the people who heard it originally. Similarly, if a text was written as poetry (for example) then you don't try and interpret it as though it is a scientific treatise.

Or take the use of concordancing. This was introduced to us in E303, Grammar in Context. The idea is, if you want to understand the meaning and significance of a word, then look at how it is used elsewhere in the corpus. But this would be no surprise to those of us who have done Bible teaching. We are used to looking at how words are used throughout the corpus - so when we use the word "faith", for example, we know that we aren't using it in the modern, culturally-conditioned sense of "a leap in the dark". We don't only do this using one translation or version of the Bible, but refer to concordances in the original languages - Greek, Hebrew - to try and get closer to the actual meaning of the word. If we are using words in a way that is different from the way in which they were intended, then we are distorting the meaning.

What is the importance of all this? For Christians, we need to understand what the nature of Bible teaching is. It's not a subjective, literature approach, where meaning is all down to the reader/listener's interpretation. It's a language approach, where whilst we may not be able to fully unpack the meaning, we do accept that some meanings are more accurate than others. This further means that the message of the Bible is an objective matter - it's not something for people to take or leave, on the basis that someone else might interpret it differently. You may reject what a Bible teacher says - but if it has been faithfully explained, you are rejecting not an interpretation of the Bible, but the Bible itself.

Don Carson, in The Gagging of God: Christianity Confronts Pluralism, set about challenging what I have called "the literature perspective" and other ways in which postmodernism has altered our thought forms when it comes to understanding Christianity. But as far as I remember, he did not make reference to the fact that the language part of English faculties already assumes a greater role for objective meaning. It's not a simple question of "us against the world" - we have co-belligerents when it comes to epistemology.

Monday, November 25, 2013

Language stuff - types of sentence

Sentences function in different sorts of ways, and we can classify them accordingly. The most obvious type is declarative - this conveys information:

  • You are looking at the cat in the basket.
However, by reorganising the elements of the sentence, we can find the other sentence types. An interrogative sentence (question) is one which requests information.
  • Are you looking at the cat in the basket?
An imperative sentence (command) is one where the subject of the sentence is being instructed to do something.
  • Look at the cat in the basket.
Another class of sentence is exclamatory. Here, the sentence is intended to convey emotion through emphasis - so:
  • Awww! Look at the cat in the basket!
does not function as an imperative, although the words are the same as the previous example.

Word order and punctuation aren't sufficient to determine the type of sentence. For example, parents might say to their children:
  • Are you going to tidy up the floor?
in a way which acted as a command, rather than a question. Similarly, a declarative sentence can be used to ask a question through intonation. This might be represented in writing using a question mark, but the word order would be as for the first example above:
  • You are looking at the cat in the basket?
We can also distinguish major sentences from minor sentences. A minor sentence is an irregular sentence, in that it doesn't contain a finite verb (a process). These have various roles - here are some examples.
  • Yes.
  • Wow!
  • Hello?
Wikipedia notes that sentences consisting of a single word are called word sentences, and the words in these sentences are called sentence words.

Thursday, November 14, 2013

More song lyric language research

There's a pretty definite "north/south" divide apparent amongst some rock musicians. Following the research project I did for my OU module, I'm interested in seeing whether this is reflected linguistically in the lyrics. But I'd like some help ....

What I'm looking for are singers, and albums, that represent definitive "north" or "south" music. That is, music that self-consciously identifies itself as belonging to either "north" or "south". My plan would be then to create a language corpus from the lyrics of these songs, and carry out the sort of corpus analysis that is hinted about in my posts below about language stuff.

In a sense, I suspect that by "south" I may really mean London. The list I've been drawing up so far consists of:

  • South - Lily Allen, Madness, Blur, The Kinks
  • North - Oasis, The Housemartins/Beautiful South, The Smiths
Other ideas? And if you could pick one archetypal album from each band, which would it be? Your thoughts, please.

Friday, October 25, 2013

Agent de-emphasis and naturalism

In my previous post, I talked about three different ways in which English could be used to draw attention away from the subject of a verb - the agent that is carrying out a particular process. These are:
  • short passive verbs;
  • nominalisation;
  • ergative verbs.
I guess my aim in highlighting this is that I'd like to think that an awareness of this would become part of more people's critical thinking repertoire - "It was said..." By whom? "Research has shown..." Who did the research?

Naturalism is, according to the Oxford English Dictionary Online, "the idea or belief that only natural (as opposed to supernatural or spiritual) laws and forces operate in the world." It says that everything in the universe is the outcome of time and chance - the universe itself has no designer; the contents of the universe (including animals and us) don't have a designer either. This is a little bit problematic, because lots of things in the universe look designed. Richard Dawkins coined the term "designoid" to refer to complex objects which are neither simple nor, he believes, designed - or rather not designed by an intelligent agent.

Another way of thinking about naturalism is to talk about telos - a word that comes from Greek, meaning "ultimate purpose or aim". The universe of the naturalist is atelic - it has no ultimate purpose or aim. Specifically, evolution, to a naturalist, is atelic. Any particular outcome of the evolutionary process - whether it's humans, multicellular life or antibiotic resistance - isn't designed, it just happens to arise.

This causes problems when it comes to language use in the context of evolutionary processes. The sort of processes that change stuff in the world are material processes. I listed the possible participants in material processes as being "actor, goal, scope, attribute, client, recipient" - with the key ones being the actor (the participant carrying out the process) and the goal (the participant affected by the process). So:
  • Sam (participant, actor) eats (process, material) some sushi (participant, goal).
But when we come to considering evolutionary processes, grammar really struggles. In Darwin's Doubt, Stephen Meyer gives examples of the way in which neo-Darwinist writers use a "word salad" to make up scientific-sounding phrases effectively as "just-so stories" to explain how evolution must have occurred. But it's worthwhile looking at these phrases from a grammar point of view as well.

Meyer gives examples of people talking about exons being "recruited" or "donated". These are short passives - remember above that the short passive is a form that allows agent de-emphasis. So, who or what is the actor associated with these processes? The same with "radical change in the structure" - here we have a nominalisation ("change") - again, the question that is begged is who or what has changed the structure? The actor can only really be "evolution":

  • exons (participant, goal) were recruited (process, material) (by evolution - participant, actor, de-emphasised)
But evolution is not allowed telos. In other contexts, people would squirm if we talked about evolution "doing" something - evolution just happens. And yet, through agent de-emphasis, we can slip in the concept of evolution as the actor in material processes.

The effect of this is that neo-Darwinism smuggles in the idea, and the categories, of purposeful, telic activity through agent de-emphasis. I would suggest that this is misleading - it is difficult to talk about evolutionary processes as being blind and purposeless: however, it's also wrong to use purposeful categories for something which has been defined as purposeless. If it is impossible to work on the basis that evolution is genuinely atelic, then maybe this belief was wrong in the first place.


Thursday, October 24, 2013

Language stuff - process types

Verbs are "doing words". However, as I suggested in my discussion of lexical density, not every verb is a doing word all the time - sometimes verbs behave as function words. When a verb is a lexical verb - that is, when it's describing something that is actually happening, it can also be referred to as a process. In fact, we can divide clauses up into processes, participants and circumstances - and, with a clause being effectively an indivisible quantum of meaning, it usually has exactly one process.

Blerk again. What does that mean? Let's take some of the clauses above and break them down.
  • Verbs (participant) are (process) "doing words". (participant)
  • However (circumstance) not every verb (participant) is (process) a doing word (participant) all the time (circumstance)
  • sometimes (circumstance) verbs (participant) behave (process) as function words. (circumstance)
What happened to my suggestion, you may be asking? We can see, since it has exactly one process, that it is a clause itself:
  • as (conjunction) I (participant) suggested (process) in my discussion of lexical density (circumstance)
However, it doesn't make any sense without the context of the second clause above which surrounded it - hence it is a subordinate clause.

Also, why is "as function words" a circumstance, not a participant? Effectively the preposition followed by the noun is behaving like an adverb - it is describing how the verbs behave, not what they are.

There are different types of processes. In the OU course, we divided processes into five sorts:
  • material - a participant acts upon the material world or is acted upon in some way ("I ate sushi");
  • mental - processes of consciousness and cognition ("We thought it didn't matter");
  • verbal - processes of communication ("I told him so.");
  • relational - being, having, consisting of, locating ("He has no father.");
  • existential - indicating the existence of an entity ("There is a problem").
In grammatical terms, we can talk about subjects, direct and indirect objects and so forth. However, these different types of processes have been assigned different types of participants - it seems to make the whole thing pretty complicated, but in actual fact, when we reflect on what is going on in a sentence, the types of participant associated with a process help to clarify the sort of process we are looking at in some cases. This summary comes from here:

  • Material - actor, goal, scope, attribute, client, recipient
  • Mental - sensor, phenomenon
  • Verbal - sayer, receiver, verbiage
  • Relational - token, value
This provides us with a more comprehensive way of analysing processes.
  • Verbs (participant, token) are (process, relational) "doing words". (participant, value)
  • as (conjunction) (participant, sayer) suggested (process, verbal) to you (participant, receiver) in my discussion of lexical density (circumstance)
As the page I just linked to makes clear, it's also possible to go into more detail about different types of circumstance - but that's quite enough for one blog post!

Wednesday, October 23, 2013

Language stuff - modality

Prior to studying E303, my experience of modal verbs had basically come from learning foreign languages - most specifically, the verbs pouvoir, devoir and vouloir which we learnt in O-level French. I had never been given grammatical categories for the same things in English, although obviously I could see how il peut mapped onto he can; voulez-vous onto do you want, and so on. They work as forms of auxiliary verbs - that is, they don't function as the main process in a sentence.

There are two categories of modal verbs - epistemic, which are modal verbs that relate to the likelihood of something being true, and deontic, which are modal verbs relating to possibility or necessity of action. They can be ranked according to their strength - O'Halloran, in the E303 textbooks, offers the following scale of epistemic modal verbs, from strongest to weakest:
  • will
  • would
  • must (in the sense of "he must be there" - "surely he's there")
  • may
  • might
  • could
  • can
and of deontic modal verbs:
  • has to
  • must (in the sense of "he must do it" - "if he doesn't do it, he's doomed")
  • had better
  • ought
  • should
  • needs to
  • is supposed to
Modal verbs are used differently in different forms of discourse. If we consider conversation, we tend to hedge - that is, we tend to make statements less assertively than if we were writing them down. Strong modality tends to come across as being forceful, and thus rude. There are other means of toning down the modality - for example, by personalising statements - I don't think that's true or even I'm sure that's not true both have less strong modality than That's not true.

In song lyrics, the dominant epistemic modal verbs in the 33000 word corpus I constructed were:
  • will (also counting 'll, I'll, won't) (307 occurrences)
  • can (can't) (290 occurrences)
  • would (I'd, wouldn't) (101 occurrences)
  • could (couldn't) (59 occurrences)
The use of deontic modality is much less common, and the most common verbs were:
  • had to (have to, has to, got to) (51 occurrences)
  • should (33 occurrences)
  • need to (needed to, needs to) (18 occurrences)
  • must (15 occurrences)
The frequency of use of strong deontic modality was very similar to what is found in the fiction corpus. However, in fiction, the use of the verb "must" is much more common than it is in song lyrics, which lean much more on "need to" and "have to".

Tuesday, October 22, 2013

Language stuff - agent de-emphasis

Normally when we think about describing an event, we think in terms of who or what is actually doing it - that is, the agent.
David broke the plate.
We may, for various reasons, not wish to draw attention to the agent. English language allows us several options for doing this. The most obvious one is to use a "short passive":
The plate was broken. 
The passive voice is used, and the person who actually does the breaking is not specified. It is possible to include the agent when using the passive voice:
The plate was broken by David.
but there may be reasons for de-emphasising the agent by omitting it - for example, when the parents come downstairs to discover the reason for a loud noise, a child might choose to draw attention to the fact that the vase has been broken, rather than admitting that it was him rather than the dog that did it.

Another option for agent de-emphasis is the use of nominalisation. This is converting a verb into a noun. I have to come up with a more complicated sentence now, as nominalisation of "to break" will leave it without a process (verb).
David broke the plate. We glued it back together.
We can de-emphasise David by nominalising the verb:
After the breakage of the plate, we glued it back together.
Yes, I know, it's a little artificial, but hopefully you get the idea.

There's a third option, and this is to use what is known as an ergative verb. This is quite subtle. An ergative verb is one that can be transitive or intransitive - that is, it can either be used with a subject and object, or just with a subject - but the object when it is used transitively becomes the subject when it is used intransitively. Blerk. The easiest way of explaining this is to give an example.
The government closed the mines.
Here, "the government" is the subject of the verb, and "the mines" is the object - or to use different grammatical categories, "the government" is the actor and "the mines" is the goal. It's possible to write this using a short passive, as explained above:
The mines were closed (omitting "by the government").
The agent/actor/subject can be omitted, which means we don't need to mention who actually closed the mines. Or we can use a nominalisation, and talk about the closure of the mines - again, the agent disappears. But we have a third option, because "close" here is an ergative verb - that is, the object of the sentence above (the mines) actually becomes the subject if we use the verb intransitively. If we want to use this verb intransitively, then the sentence becomes:
The mines closed.
(rather than "The government closed.") Once again, the agent/actor has disappeared.

There are various reasons why agent de-emphasis might be considered desirable. Those of us who have been using computers for more than ten years probably remember earlier versions of Word for Windows nagging us about using the passive voice. In my case, it was because I was often writing about science - and an aspect of science writing is use of the passive voice - to de-emphasise the person actually doing the work. Using nominalised verbs allows the writer/speaker to increase the lexical density - that is, to convey more information in less space. This is valuable in media where word count and space is at a premium - like journalism.

More significantly, as the example above suggests, there may be political reasons for de-emphasising the agent. And I'd like to return to this in a future post ....

Monday, October 21, 2013

Language stuff - field, tenor, mode

One of the aspects of studying language that interested me is how recently much of linguistic theory has been developed. When I was studying computer science in the late 80s, many of the theoretical foundations were pretty new - Dijkstra's algorithm, for example, that we were taught about in my degree (which is now part of the Further Maths A Level syllabus!) was published in 1959. However, Michael Halliday's seminal book on language An Introduction to Functional Grammar was not even first published until 1985. Both Halliday and Noam Chomsky, perhaps the most famous linguistic theorist, are still alive.

The features of a use of language may be considered by considering its field, tenor and mode. The field is also referred to as the experiential metafunction. It is how language is used to make meaning about the world - in other words, the actual content of what is being communicated.

One might naively think that this was all that language was - the communication of information - but it is more subtle than that. Every language event takes place between one or more participants, and in addition to communicating information, language events are used as part of the process of enacting interpersonal relations. This is tenor, also referred to as the interpersonal metafunction. Additionally, language events can take place in many different forms (conversation, email, a sermon ...), and these are themselves largely detached from both field and tenor. This textual metafunction is the mode.

So, what does this mean in practice? Let's take this blog post.

  • Field - I am attempting to explain, in fairly simple terms, information about the theoretical use of language.
  • Tenor - I'm addressing unknown readers (who are you? Say hello!), but I'm writing in a fairly informal style - I'm assuming that the average reader will just have happened across this, and wants to read something that's engaging, friendly and not too heavy. Frankly, that's how I like communicating anyway, and since this is my space, really written for my own amusement, I guess I do what I want.
  • Mode - a blog post. Written language can be more planned and deliberate than spoken language, for a start - there's a definite structure, and I've assumed in writing it that people will start at the beginning and read through hopefully to the end!
You can imagine changing each of those individually would change the way in which language was used. For example - suppose (field) I was writing about something else, maybe a film review? Or suppose (tenor) I knew that the people reading this were children aged 12? Or suppose (mode) I was delivering this as a talk? Each of the metafunctions, then, has a bearing on how language is used, and this linguistic structure is something that has only really been described in the last 30 years or so.

Monday, October 14, 2013

Language stuff - lexical density

Words in a text can be divided into content words and function words. Content words refer to some object, action or other non-linguistic meaning. The sort of words that are content words are nouns, adjectives, adverbs and most verbs. These are "open classes" of words - that is, it's possible to invent more content words. If I said "The abdef ghijk was mnoping stuvly over there", even though you'd never come across a sentence like that, you'd probably conclude that "ghijk" was a thing (noun), of which you can get "abdef" ones (adjective) "to mnope" was a verb which it is possible to do "stuvly" (adverb). If someone feels like drawing a picture of this happening, I'll include it in this post!

Function words are words that have little meaning in themselves, but express grammatical relationships with other words in the sentence. In the sentence above, "the", "was" and "over there" are all function words. They include conjunctions, prepositions, modal and auxiliary verbs, and pronouns. These are all "closed classes" - that is, there isn't "space" in the English language to add new ones, Dr. Dan Streetmentioner notwithstanding. All the new words that come along are content words, not function words.

It is possible to work out the proportion of content words compared to the total number of words. This is the lexical density. Different sorts of texts will have different lexical densities. On our OU course, we used the Longman Student Grammar of Spoken and Written English. This is a descriptive grammar (in other words, it described how English was used, rather than saying how it ought to be used), and analysed four different styles of discourse. Based on large corpora, it gave the lexical density of different sorts of discourse as:

  • Conversation - 35%
  • Fiction - 47%
  • Academic prose - 51%
  • News - 54%
Conversation is low for several reasons. The first is that, unlike written discourse, in most conversations, there is a shared context. This means that it's possible to use pronouns to a greater extent than nouns, for example. Also, conversation is improvised to a greater extent than written discourse. This means that there are likely to be dysfluencies - such as hesitators and repetition - which have the effect of decreasing the lexical density. As part of the OU course, I did my own analysis of lyrics from pop records. Their lexical density turns out to be almost exactly the same as that of fiction.


Tuesday, October 08, 2013

Language stuff - corpus (pl. corpora)

A new(ish) tool for the systematic examination of the English language is the development of language corpora. These are collections of samples of English language into a "body", which can then be examined using computer programs.

According to Wikipedia, the first corpus used for language investigation was the Brown Corpus. It consisted of around a million words, gathered from about 500 samples of American English, and was used in the preparation of the benchmark work Computational Analysis of Present-Day American English by Kucera and Francis. This was as recently as 1967.

The size of corpora has increased with increasing computer power. The British National Corpus currently contains 100 million words. The Oxford English Corpus - used by the makers of Oxford Dictionaries amongst others - contains 2 billion words of English. The Cambridge English Corpus is a "multi-billion word" corpus. In addition to the texts that make up the corpora, the words they include can be tagged for parts of speech - for example, whether "love" as it appears in a text is being used as a noun ("His love was so great...") or a verb ("I love you"). A corpus can be examined with software called concordancing software. This will search for specific words, phrases or instances of grammar, and can do things like highlight words that are frequently collocated. This can be used to identify patterns in the language that might otherwise go unnoticed.

Corpora can be produced from particular classes of text - for example, transcribed conversations, newspaper articles, academic journals, fiction. For E303, the Open University undergraduate course that introduces corpus linguistics, we were provided with a 4 million word corpus, with a million words drawn from each of these classes. It's also possible to produce your own corpus. I created a corpus of pop song lyrics - only 33000 words or so, but still enough to look for trends and patterns of language use. There is software available that can tag a text with parts of speech - for example, CLAWS4. And for analysing the sofware, the AntConc concordancing software is freely available.

Monday, October 07, 2013

Language stuff - type/token ratio

I just finished the Open University module E303 - English Grammar in Context - it sounds pretty deadly, but I loved it. Language is inherent to who we are as human beings - we all communicate. And yet, it's only relatively recently that the resources have been available to examine language in a systematic, large-scale way. A lot of the underlying theory is actually newer than the computer science theory that felt pretty new when I was doing my first degree.

I've promised blog series before, and they rarely amount to much, but I'd like to see whether I can write about some of the ideas we covered, and maybe get across some of the reason that I found the material so fascinating.

The first concept is type/token ratio. "Tokens" are the number of words in a piece of text - if I do a word count, then it tells me the number of tokens. But not all of them are unique. The most common word "the" I have now used ten times so far in this text (don't hold me to that - it's likely to have been edited in a highly non-linear manner - but you get the idea). You can get some insight into a text by dividing the number of unique words by the total number of words, and expressing it as a percentage. So for the text up to the start of this sentence, there were 229 words and 134 types - giving a type/token ratio of 59%.

A couple of things about type/token ratio. The first is that as a piece of text gets longer, the type/token ratio is likely to fall. The number of words is clearly increasing, but the number of types is increasing more slowly - it's more likely that you will be using the same words again. What that means is that if you want to compare type/token ratio of two different texts, they need to be about the same size.

The next thing is that different sorts of text will have different type/token ratios, as they are a measure of the diversity of the vocabulary being used. For my final assignment, I looked at pop song lyrics. I had a database of around 34000 words, and this had a type/token ratio of just under 10%. I compared this with a slightly larger database of words from a work of fiction, and this had a higher type/token ratio - just over 12%. A slightly smaller database of words from transcribed conversations had a lower type/token ratio - about 6.5%.

One might assume that the language used in pop music was pretty narrow in its range. But it turns out that it is quite diverse - almost as diverse as fictional writing, and much more so than the sort of language that's used in everyday conversation.

Saturday, January 22, 2011

I think I disagree with my lecturers ...

... and the course hasn't even started yet!

The course is U211, Exploring the English Language, which is not technically due to start for another week or so. However, in a bid to get ahead, since I really don't think I'm going to have the 15-16 hours a week (!) that it claims I need, I've reached the section on accents, chapter 5 of the first book.

The focus in the course has been that no one variety of English should be privileged. That's the sense of the background reading - Crystal's "The Stories of English" emphasises the fact that the conventional narrative of the rise and rise of English disregards the fact that "standard English" is only one facet of the English language. Graddol's "English Next?" explores the issue that English is, in world terms, dominated by non-native speakers. And the opening chapters of the first book have been keen to emphasise that the prescriptive approach adopted to the language in spelling, grammar and pronunciation has only led to one of the expressions of English that we see today.

In discussing accents, however, I think the course goes too far. I am quite happy that in general, accents don't in themselves say anything about the intelligence of the speaker - I've known too many English speakers from all over the place to think otherwise. I'm also quite happy that RP is not a "non-accent".

However, in discussing accents, the focus has tended to be on aspects of pronunciation that are "neutral" - for example, glottal stops or dropping aitches ("ge' inside the 'ouse!"). There has been no discussion so far on the fact that a few aspects of accents quite often betray a level of ignorance of English, or illiteracy - for example, one feature that arises often is the use of "of" where "have" is correct in English (as in "I couldn't of"). The sense I get from the course is that it seeks, in quite a postmodern/pluralist way, to affirm all users of English, regardless of how the language is spoken. (In true postmodern fashion, of course, the language adopted for the course is itself standard English, and I strongly suspect that a response to the course that was not would be likely to raise eyebrows).

I'm trying to imagine how the writers of the course would respond to this issue. They might suggest that English is mutating so that "of" instead of "have" in this context will be considered acceptable usage. But if this is to take place, then sections of the rules of English relating to particles and verb tenses have to be basically disregarded, and in the fullness of time, this would be likely to tidily erode the comprehensibility of the language. They might suggest that there is a difference between an accent and an incorrect usage - but the emphasis hitherto has been that there is no "incorrect usage" - just different, and people need to swim with the tide in this regard. I'd be interested in hearing their thoughts on this.

Personally, for what it's worth, I think that whilst the prescriptive approach is wrong, and fails to take into account many valid expressions of English, the people who write the course are also wrong if they are saying that all expressions are equally valid. There is some discussion about the tone used for science writing, which has taken shape over the centuries, and the writers accept the requirements of the medium. More generally, whilst RP and Standard English have no right to a privileged position in the canon of English language beyond their usefulness as being most widely acceptable, I don't think that the substitution of varieties of English which undermine its ability to communicate can be regarded as progressive.

Wednesday, July 04, 2007

Frustration with French

My poor children are trying to learn French the National Curriculum way. This approach seems to rely fundamentally on the principle: "Don't mention grammar!" It's little wonder that English people are so bad at languages, because we aren't ever taught how they work.

So they learn lots of phrases and bits and pieces. You can say, "J'aime le sport. C'est génial!" from a pretty early stage. But nobody tells you that "j'aime" is part of the verb "aimer", which if you want to, you can do an awful lot more with. And even more frustratingly, nobody ever seems to say anymore: "être, avoir, faire, aller - these are four irregular verbs which will take a little while to learn, but once you have got to grips with them, you will actually be able to communicate a huge range of ideas." This represents about four lessons' worth of learning - but would do more to make people confident in basic French than three whole years of learning phrases about liking Coca Cola and going to the park to do sport tomorrow.

Of course, if you are immersed in French culture, then you learn fast - by trial and error - starting from a relatively small vocabulary. That's how children learn language at home. Think of the mistakes that we make as children - "brung" instead of "brought", "buyed" instead of "bought", "goed away" instead of "went away". The reason that those mistakes are made are because we are trying to apply rules we have intuitively picked up too simply to the language - not because we have no awareness of them. As we grow, our stock of rules grows, and we become more adept at using language.

But children learning a second language simply aren't in an environment where they have the ability to try out different rules they have intuitively worked out. So they just end up with the stock of a few hundred basic phrases, all individually learnt, all with no connection to each other. It's SOOOO much harder to learn that way, and it's hardly surprising that the whole lot has been dumped by the time the child is 17.

My recommendation for the government? It doesn't make any difference how young you start teaching a foreign language, if you aren't teaching it in a way that corresponds with how people's minds work. Literacy levels in primary schools have risen since the National Literacy Strategy was initiated - to the extent that many children go to secondary school with loads of English Language concepts that they simply don't need through GCSE. Unsurprisingly, children tend to "go backwards" educationally in year 7, which I think still hasn't been adapted to deal with the ground that is covered in Key Stage 1 and 2. If you want people to learn foreign languages, find better ways of teaching them. By all means teach vocabulary and phrases to primary children. But if you want secondary children to really learn languages, then for goodness' sake teach them how the languages work!

Saturday, April 14, 2007

Question about language and gender

Being British, my command of any other languages is pretty limited. I have a question about other languages, and I'd be interested in the insight of anybody who has a better command of something other than English.

In English, there are three genders (male, female, neuter) but male and female genders are only used for objects with a definite gender ("she" might be a girl, a woman, a ship or a female animal of some sort, but wouldn't normally be a city or an unspecified bee). "Cow" in English tends to be used to refer to any domestic bovine animal, including bulls, unless otherwise specified - which means that the word "cow" is basically a neuter gender word ("What is it? It's a cow!"). It can be a female gender word, but only usually if the cow(s) in question has been definitely identified as female.

However, in Spanish (to give an example), I was introduced to la vaca, which is a female word translated as "cow". Do they use this word in the same way as English people use the word "cow", despite it being a definitely female (and not neuter) word? Or, given the definite role that bulls play in Spanish culture, does la vaca definitely differentiate from el toro, and you have to plump for one or the other when you first see the creature?

And what about, say, cats? The Spanish for "the cat" is el gato. Again, that's not bad if you are (say) asking about a cat you just happen to have seen. But supposing you are referring to a cat you know to be female - your pet cat Phoebe, say? Are you still stuck with using male pronouns to refer to her, despite the fact that you know her to be female? Would you say Phoebe es un gato when she is actually female?

Does the same happen in other languages?