Decomposing dependency analysis: revisiting the relation between annotation scheme and structure-based textual measures.
Saved in:
| Title: | Decomposing dependency analysis: revisiting the relation between annotation scheme and structure-based textual measures. |
|---|---|
| Authors: | Yih, Tsy1 (AUTHOR), Liu, Haitao2 (AUTHOR) |
| Source: | Digital Scholarship in the Humanities. Apr2025, Vol. 40 Issue 1, p400-418. 19p. |
| Subjects: | Relative clauses, Units of measurement, Digital humanities, Annotations, Buds |
| Abstract: | Standardized quantitative measurement of texts lies at the heart of digital approaches to humanities. Structure-based textual measures are known to be influenced by the choice of syntactic annotation schemes. Building on previous research, the present article further explores the relation between annotation schemes and the index of mean dependency distance (MDD) by comparing the treebanks of seventeen languages, respectively, within a tree representation (basic universal dependencies, BUD) and within a graphic representation (enhanced universal dependencies, EUD). Following the idea of decomposing annotation schemes into the combinations of analyses of specific constructions (coordinate structures, control constructions, and relative clauses), we design algorithms to identify them in the CoNLL-U format treebanks and explore their influences. It is found that the overall MDD of the EUD representation is statistically higher than that of BUD at corpus level, primarily affected by the coordinate structure due to its high frequency. At sentence level, all three constructions might contribute to either increased or decreased MDD, with stochastically intervening words and word order being two important determinants of the values of the measure. Finally, we propose and argue for the view that MDDs calculated under different annotation schemes should be regarded as different textual measures in nature. In sum, the present study provides another case study to deepen our understanding of the nature of syntactic annotation schemes and its relation with textual indices, which paves the way for standard measurement of texts in future humanities research. [ABSTRACT FROM AUTHOR] |
| Copyright of Digital Scholarship in the Humanities is the property of Oxford University Press / USA and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
Be the first to leave a comment!