Comparing Frequency and Dispersion Keywords: Effects of Variations in Target and Reference Corpora
Saved in:
| Title: | Comparing Frequency and Dispersion Keywords: Effects of Variations in Target and Reference Corpora |
|---|---|
| Language: | English |
| Authors: | Punjaporn Pojanapunya |
| Source: | rEFLections. 2025 32(3):1428-1446. |
| Availability: | King Mongkut's University of Technology Thonburi School of Liberal Arts. 126 Pracha Uthit Road, Bang Mod, Thung Khru, Bangkok, Thailand 10140. Tel: +66-2470-8756; Fax: +66-2428-3375; Web site: https://so05.tci-thaijo.org/index.php/reflections/index |
| Peer Reviewed: | Y |
| Page Count: | 19 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Information Retrieval, Word Frequency, Linguistics, Discourse Analysis, English, Generalizability Theory, Form Classes (Languages) |
| ISSN: | 1513-5934 2651-1479 |
| Abstract: | Dispersion keyword analysis, which identifies words that occur in significantly more texts in the target corpus than in the reference corpus, has recently been introduced as a more effective method than traditional frequency keyword analysis. Previous research has used this method to identify keywords within a target corpus, usually consisting of hundreds of texts, and used a much larger corpus as a reference. However, questions remain regarding its applicability for cases involving fewer texts and comparisons between smaller specific corpora. This study compares the top 100 frequency keywords and dispersion keywords identified under several conditions, which varied in terms of the number of texts in the target corpus (24, 100, and 200 texts) and the types of reference corpora used. Both methods identified unique and shared keywords; however, frequency keywords are found more frequent and widely dispersed not only within the target corpus but also in the reference corpus compared to dispersion ones, which are notably more relevant to the target corpus. The selection between frequency and dispersion methods and the relevance of frequency and dispersion keywords in research with differing focuses are discussed. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1498276 |
| Database: | ERIC |
| Abstract: | Dispersion keyword analysis, which identifies words that occur in significantly more texts in the target corpus than in the reference corpus, has recently been introduced as a more effective method than traditional frequency keyword analysis. Previous research has used this method to identify keywords within a target corpus, usually consisting of hundreds of texts, and used a much larger corpus as a reference. However, questions remain regarding its applicability for cases involving fewer texts and comparisons between smaller specific corpora. This study compares the top 100 frequency keywords and dispersion keywords identified under several conditions, which varied in terms of the number of texts in the target corpus (24, 100, and 200 texts) and the types of reference corpora used. Both methods identified unique and shared keywords; however, frequency keywords are found more frequent and widely dispersed not only within the target corpus but also in the reference corpus compared to dispersion ones, which are notably more relevant to the target corpus. The selection between frequency and dispersion methods and the relevance of frequency and dispersion keywords in research with differing focuses are discussed. |
|---|---|
| ISSN: | 1513-5934 2651-1479 |