Introduction to Corpus Linguistics
1. Corpus Linguistics: Key Concepts
Corpus Linguistics is the study of language based on real-life language use through the analysis of corpora (collections of texts). It adopts a corpus-based approach to understand language patterns and usage (McEnery & Wilson 1996).
2. Sampling and Representativeness
- Language is infinite, so studying the entire language is impossible.
- We create a corpus as a sample that is maximally representative of the language variety under study.
- Representativeness means the sample includes the full range of variability in the population.
- Texts in a corpus are selected using explicit criteria to ensure they represent the language or its subset accurately.
- The corpus must be relevant and adequate to the research task to support valid claims based on statistical frequencies.
3. Reference Corpora
| Feature | Description |
|---|---|
| Size | Large, usually millions of words |
| Content | Wide range of texts covering different varieties of the language |
| Purpose | Provide thorough and comprehensive information about a language |
| Functionality | Serve as a comprehensive representation of the language |
4. Concordance
- A concordance is a list of all occurrences of a search pattern (word or phrase) in a corpus.
- It includes the immediate context around the search term, showing the node and its surrounding co-text or linguistic context.
> A corpus must be representative and relevant to the research question to yield meaningful linguistic insights.
Sampling and Representativeness
1. Sampling and Representativeness
KWIC (Key Word In Context)
- A KWIC display shows concordance lines with the node word (target word) centrally aligned.
- Each line presents the node word surrounded by its immediate context (words before and after).
- This format helps identify patterns of usage and meaning by focusing on the node's co-text.
Concordance and Concordance Analysis
- Purpose: To understand how a word is used and to infer its meaning through patterns.
- Process:
- Categorize occurrences of the node word.
- Identify regularities in its usage.
- Generalize from these patterns to describe typical behavior.
What to Look for in Concordances
- Recurrence of word combinations and structures in the node's co-text.
- Repeated patterns of co-selection (words that tend to appear together with the node).
- These patterns reveal syntagmatic relations (how words combine in sequences).
Reading Concordance Lines
- Scan vertically across lines to detect patterns of co-selection.
- Look for co-occurrence of lexical items and grammatical structures.
- This vertical reading highlights paradigmatic relations (choices available in similar contexts).
- Co-selection and recurrence provide a profile of the node word’s typical usage.
Co-selection and Recurrence
- Co-selection: The tendency of certain words or structures to appear together with the node.
- Recurrence: The repeated appearance of these co-selected items across concordance lines.
- Together, they form the basis for evaluating the lexical and grammatical profile of a word or expression.
Nature of Texts vs. Corpora
- Corpora provide large, authentic samples of language use, unlike isolated texts or dictionaries.
- Example: The phrase "in + adjective + context" shows that context is more often used as an adjunct (a modifying element) rather than in other syntactic positions.
Adjunct
- An adjunct is a non-essential element that adds information to a sentence (e.g., time, place, manner).
- Recognizing adjuncts in concordance lines helps understand the typical syntactic roles of words.
Key point: Concordance analysis through KWIC displays reveals patterns of word usage by focusing on co-occurrence and recurrence, enabling a detailed profile of lexical and grammatical behavior.
Reference Corpora and Concordance
1. Reference Corpora and Concordance
Reference corpora are large, structured collections of texts designed to represent a language or variety comprehensively. They serve as benchmarks for linguistic analysis, providing authentic data for studying language use.
Concordance is a tool or output that displays all occurrences of a word or phrase within a corpus, showing the immediate context. It helps analyze patterns of usage, collocations, and semantic prosody.
2. Key Concepts and Examples
-
Adjuncts: Phrases not essential to clause structure but adding extra meaning.
-
Semantic prosody: The tendency of a word to co-occur with positive or negative contexts.
- Example:
- Persistent + nouns → negative prosody (e.g., errors, opposition)
- Persistence (noun) → positive quality
- Example:
-
Thesauruses vs. Corpora:
Thesauruses list synonyms/antonyms; corpora provide real usage examples and context.
3. Meaning Differences in Similar Constructions
| Construction | Meaning Encoded | Example Usage |
|---|---|---|
| It is possible + to-infinitive | Ability | "It is possible to solve the problem." |
| It is possible + that-clause | Supposition or hypothesis | "It is possible that he will come." |
4. Verb + Preposition/Clause Patterns: Complain
| Pattern | Typical Complement | Meaning/Use | Translation Example |
|---|---|---|---|
| Complain of | Specific disease or noun | Specific complaint | "complains of a pain in her leg" |
| Complain about | General trouble or noun | General complaint | "complain about heavy traffic" |
| Complain that | Subordinate clause (situation) | Complaint about a situation | "complain that nothing works in Italy" |
5. Additional Notes on Complain
- Some Italian translations show subtle differences:
- "Fortunatamente non si lamentano vittime" → "Luckily no victims are claimed" (passive sense)
- "In Veneto si lamentano danni" → "Veneto area has been hardly hit" (damage reported)
- Informal use:
- "Come va?" — "Non mi lamento!" → "How are you?" — "Not bad!"
To retain:
Reference corpora provide authentic language data; concordances reveal word usage patterns and semantic prosody, crucial for understanding subtle meaning differences and collocational behavior.
KWIC Display and Concordance Analysis
1. KWIC Display and Concordance Analysis
Dictionaries vs. Concordances
- Dictionaries provide word meanings and definitions.
- Concordances reveal how words actually combine in real language use, showing patterns of lexical, grammatical, and textual co-occurrences.
2. What Concordance Analysis Provides
- Identifies probable language use rather than just possible or correct forms.
- Highlights frequent word combinations linked to specific registers, genres, or text types.
- Focuses on patterns and frequencies of language events.
- Detects repeated language events (frequent co-occurrences) and classifies them by parameters like word class or semantic similarity.
- Reveals patterns of co-selection and their functions, connecting form, meaning, and communicative purpose.
3. Example: "Eye" vs. "Eyes"
| Word | Meaning Focus | Typical Concordance Patterns |
|---|---|---|
| Eyes | Organ of sight | brown eyes, dark eyes, hazel eyes, grey eyes |
| Eye | Monitoring, critical examination | keep an eye on, turn a blind eye, with an eye for |
4. Core Principles of Concordances
- Context relevance: Meaning depends on the surrounding words.
- Recurrence relevance: Frequent repetition of patterns is significant.
- Concordances are best read vertically, scanning for patterns of co-selection and repetition rather than isolated examples.
Concordances reveal language patterns by combining the relevance of context and recurrence, enabling the study of actual language use beyond dictionary definitions.
Working with Concordances and Pattern Recognition
1. Concordances
A concordance is a list of occurrences of a word or phrase (called the node) within its immediate co-text (context). It helps identify repeated patterns around the node, revealing how it is typically used.
- Example: Concordance for the phrase "make up" shows its use in different contexts, such as:
- "make up for the lack of information..."
- "make up about 47% of PC users..."
Concordances allow comparison of usage patterns across languages or corpora, highlighting typical collocations and phrase structures.
2. Pattern Recognition in Corpus Linguistics
Pattern recognition involves identifying frequent and significant linguistic patterns in concordance lines or corpora.
- It focuses on repeated co-textual patterns around a node.
- Helps in understanding collocations, idiomatic expressions, and grammatical structures.
- Useful for translation studies, lexicography, and language teaching.
3. Key Points to Remember
| Concept | Definition / Role | Example / Note |
|---|---|---|
| Node | The target word or phrase in a concordance | "make up" |
| Co-text | The surrounding text around the node | Words before and after "make up" |
| Concordance | List of all occurrences of the node with co-text | Shows usage patterns |
| Pattern Recognition | Detecting frequent, meaningful patterns in co-text | Identifies collocations and idioms |
Concordances reveal how words are used in context by showing repeated patterns in their co-text, enabling pattern recognition essential for linguistic analysis.
Corpora versus Dictionaries and Thesauruses
1. Corpora versus Dictionaries and Thesauruses
Corpora are large, structured collections of real-world texts used to analyze language use in context, while dictionaries and thesauruses are curated reference tools providing definitions, synonyms, and lexical relations.
| Aspect | Corpora | Dictionaries and Thesauruses |
|---|---|---|
| Nature | Collections of authentic language data | Authoritative lexical references |
| Content | Actual usage examples, co-text, frequency data | Definitions, synonyms, antonyms, usage notes |
| Purpose | Empirical analysis of language patterns | Providing standardized meanings and lexical relations |
| Contextual information | Rich co-text showing how words combine and vary | Limited, often isolated word entries |
| Flexibility | Dynamic, can reflect language change and variation | Static, updated periodically |
a) Importance of Co-text in Corpora
- Co-text refers to the words immediately surrounding a target word or phrase.
- It is crucial for disambiguating meanings, distinguishing between literal and figurative uses, and identifying pragmatic nuances.
- Example: The verb phrase make up has different meanings depending on co-text:
- make up for the time lost (compensate)
- make up your mind (decide)
- Corpora reveal these patterns through concordance lines, showing frequent left and right collocates that clarify meaning.
Key insight: Even a few recurring co-textual elements on either side of a target word can sufficiently disambiguate its most frequent meanings and functions.
b) Corpora and Legibility
- Corpora provide authentic examples illustrating subtle semantic and pragmatic distinctions.
- For instance, the adjective legible appears in various contexts showing degrees of clarity:
- barely legible postmark vs. clearly legible label
- Such examples help understand gradations of meaning and usage in natural language.
> To distinguish meanings and pragmatic nuances effectively, analyzing co-text in large corpora is more reliable than relying solely on dictionary or thesaurus entries.
Classification Principles in Corpus Linguistics
1. Classification Principles in Corpus Linguistics
Corpus linguistics uses classification to analyze language by grouping words or expressions based on their usage patterns and meanings observed in real text samples.
2. Key Concepts
- Concordance: A tool that displays all occurrences of a word or phrase in a corpus, allowing the study of its context and usage.
- Concordance analysis: Examines how language is used in context to identify patterns, meanings, and distinctions between similar words.
3. Distinguishing Synonyms through Concordance
- Corpus data helps differentiate apparent synonyms by showing their typical contexts.
- Example: The words "readable" and "legible" appear similar but have distinct usage patterns:
- "Legible" typically refers to the clarity of handwriting or print (e.g., "the only legible thing was the note").
- "Readable" often describes text that is easy and pleasant to read (e.g., "a highly readable novel", "pleasantly readable books").
4. Practical Use of Classification in Corpus Linguistics
| Purpose | Description | Example |
|---|---|---|
| Identify conventional language | Understand typical language use and norms | Analyzing frequent collocations |
| Differentiate synonyms | Clarify subtle meaning differences | Comparing "readable" vs. "legible" usage |
| Observe language in context | Study how words function in real communication | Concordance lines showing usage patterns |
Concordance analysis is essential to classify language use by revealing contextual differences and conventional patterns, enabling precise semantic distinctions.
Collocation, Colligation, and Semantic Preference
Corpus linguistics studies how words combine, the associations they form, and the meanings that emerge from these combinations. To analyze language use and identify regular patterns, we apply classification principles that group linguistic elements based on shared features.
1. Key Classification Principles in Corpus Linguistics
| Principle | Definition | Focus |
|---|---|---|
| Collocation | The tendency of two or more words to occur close together. | Word + word |
| Colligation | The relationship between a word and a grammatical class of words it commonly associates with. | Word + grammatical class |
| Semantic Preference | The tendency of a word to co-occur with words from a particular semantic class. | Word + semantic class |
| Semantic Prosody | The typical "environment" or connotation (positive, negative, etc.) surrounding a pattern. | Contextual semantic environment |
2. Collocation
- Defined as the occurrence of two or more words within a short space of each other (Sinclair 1991: 170).
- It reflects the attraction between words, showing which words tend to co-occur frequently.
- Example: "strong tea" vs. "powerful tea" — "strong" collocates naturally with "tea."
3. Colligation
- Describes the relationship between a word and the grammatical class of words it associates with.
- Focuses on the grammatical environment around a word, such as the types of words or structures that commonly appear before or after it.
- Example: A verb might typically be followed by a noun phrase or an adverb.
4. Semantic Preference
- Similar to colligation but concerns the semantic class rather than grammatical class.
- It identifies the tendency of a word to co-occur with words belonging to a particular semantic category.
- Example: The verb "commit" often co-occurs with words related to crimes or actions (e.g., "commit a crime," "commit suicide").
5. Semantic Prosody
- Refers to the overall semantic environment in which a word or phrase tends to appear.
- This environment can carry a positive, negative, or neutral connotation, influencing the perceived meaning of the word.
- Example: Words like "cause" often appear in negative semantic prosodies (e.g., "cause problems," "cause damage").
To identify linguistic patterns, corpus linguistics relies on collocation, colligation, semantic preference, and semantic prosody as core classification principles.
Semantic Prosody
Semantic prosody differs from collocation and colligation by focusing on the pragmatic meaning or discoursal function that emerges from the typical contexts in which a word or phrase appears. It reflects how words influence each other's meanings beyond simple co-occurrence.
1. Key points on Semantic Prosody
- Definition: Semantic prosody is the "aura of meaning" that a word or phrase acquires due to its consistent presence in particular contexts.
- It involves the connotational colouring spreading beyond individual words to affect the overall meaning of expressions.
- Some words are so selective in their contexts that these contexts become part of their meaning.
2. Types of Semantic Prosody
| Type | Description | Example Contexts |
|---|---|---|
| Negative | Node word attracts collocates with strong negative semantic characteristics. | Words associated with failure, risk, or harm |
| Positive | Node word co-occurs with expressions referring to positive, desirable things. | Words linked to success, happiness, or benefit |
| Neutral | Node word co-occurs with both positive and negative collocates, balancing out the affective meaning. | Words used in mixed or neutral contexts |
3. Classification (Partington, 2004)
- Favourable (Positive): Pleasant or positive affective meaning.
- Neutral: Neither strongly positive nor negative.
- Unfavourable (Negative): Unpleasant or negative affective meaning.
Semantic prosody is the connotational meaning a word acquires from its habitual collocational environment, influencing how it is perceived beyond its dictionary definition.
Analysis of Concordance Lines
1. Analysis of Concordance Lines
Key concepts in concordance analysis:
-
Collocation: Identification of words that frequently co-occur with the target word, beyond random chance.
Example:
the + naked eye -
Colligation: The grammatical patterns or syntactic environments in which the word appears, such as specific prepositions or articles.
Example:
[preposition with/to/by] + the (definite article) + naked eye -
Semantic preference: The tendency of a word to co-occur with words sharing a related meaning or semantic field.
Example:
Words related to visibility (see, visible, look at) + [preposition] + [definite article] + naked eye -
Semantic prosody: The connotative or attitudinal meaning that spreads over a word’s environment, often positive or negative.
Example:
Words expressing difficulty in seeing (don’t, can’t, hardly, faint, difficult) + [visibility] + [preposition] + [definite article] + naked eye -
Neutral context: When no clear semantic prosody or connotation is detected, the instance is labeled as neutral.
2. What to examine in concordance lines?
| Aspect | Description |
|---|---|
| Collocates | Words that occur most frequently and significantly near the target word. |
| Chunks/Idioms | Recurring multi-word expressions or fixed phrases including the target word. |
| Syntactic restrictions | Typical syntactic patterns, such as preferred prepositions, clause positions, or tense/aspect constraints. |
| Semantic restrictions | Limitations on the semantic domain of the word’s usage (e.g., applied only to humans). |
| Semantic prosody | The overall positive or negative connotative environment surrounding the word. |
Semantic prosody reflects how a word’s meaning extends beyond its immediate collocates to the broader attitudinal or evaluative context it typically appears in.