Chinese ethnic minority book classification by large language models within CLC.
Saved in:
| Title: | Chinese ethnic minority book classification by large language models within CLC. |
|---|---|
| Authors: | Jia, Junzhi1 (AUTHOR) junzhij@163.com, Xu, Hui1 (AUTHOR) xuhui111@ruc.edu.cn, Guo, Yiqian1 (AUTHOR) guoyiqian@ruc.edu.cn, Gao, Wenjing2 (AUTHOR) 437189571@qq.com |
| Source: | Electronic Library. 2026, Vol. 44 Issue 2, p251-270. 20p. |
| Subject Terms: | *Library science, *Artificial intelligence, *Books, *Experimental design, *Information retrieval, *Comparative studies, Classification of books, Research funding, Natural language processing, Ethnology, Descriptive statistics, Minorities, Data analysis software |
| Geographic Terms: | China |
| Abstract: | Purpose: This study evaluates the capabilities and limitations of large language models (LLMs) in classifying Chinese ethic minority books under the scheme of Chinese Library Classification. Design/methodology/approach: A test collection of Chinese ethnic minority bibliographic records was constructed, and prompt engineering was used to compare the classification performance of DeepSeek-v3 and ChatGPT-4o under two input scenarios: "title + abstract" and "title only." By designing evaluation metrics that include accuracy, granularity and error-type analysis, this study systematically evaluates the performance differences between the models, diagnoses the causes of errors and proposes improvement strategies. Findings: Experimental results show both models performed well in the broad category classification of Chinese ethnic minority books, with DeepSeek-v3 exceeding 80% accuracy. Incorporating abstracts further improved accuracy and prompted longer, more detailed classification codes. However, accuracy declined for both as classification codes grew more specific. DeepSeek-v3 significantly outperformed ChatGPT-4o, achieving an overall accuracy of 40.78% and 33.50% with and without abstracts, respectively, while ChatGPT-4o remained below 6%. On the basis of classification error analysis, this study proposes improvements in classification system design, model capability enhancement and human–artificial intelligence (AI) collaboration to guide practical improvements in organizing ethnic minority resources. Originality/value: Combining librarianship, ethnography and artificial intelligence, this study is the first to compare the classification ability of different large-language models for Chinese ethnic minority books. It reveals cultural limitations in knowledge organization systems, identifies the "capability threshold" of LLMs in cultural context processing and establishes an empirical basis for developing culturally-aware AI governance frameworks. [ABSTRACT FROM AUTHOR] |
| Copyright of Electronic Library is the property of Emerald Publishing Limited and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Education Research Complete |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: ehh DbLabel: Education Research Complete An: 192696103 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Chinese ethnic minority book classification by large language models within CLC. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Jia%2C+Junzhi%22">Jia, Junzhi</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> junzhij@163.com</i><br /><searchLink fieldCode="AR" term="%22Xu%2C+Hui%22">Xu, Hui</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xuhui111@ruc.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Guo%2C+Yiqian%22">Guo, Yiqian</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> guoyiqian@ruc.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Gao%2C+Wenjing%22">Gao, Wenjing</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> 437189571@qq.com</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Electronic+Library%22">Electronic Library</searchLink>. 2026, Vol. 44 Issue 2, p251-270. 20p. – Name: Subject Label: Subject Terms Group: Su Data: *<searchLink fieldCode="DE" term="%22Library+science%22">Library science</searchLink><br />*<searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br />*<searchLink fieldCode="DE" term="%22Books%22">Books</searchLink><br />*<searchLink fieldCode="DE" term="%22Experimental+design%22">Experimental design</searchLink><br />*<searchLink fieldCode="DE" term="%22Information+retrieval%22">Information retrieval</searchLink><br />*<searchLink fieldCode="DE" term="%22Comparative+studies%22">Comparative studies</searchLink><br /><searchLink fieldCode="DE" term="%22Classification+of+books%22">Classification of books</searchLink><br /><searchLink fieldCode="DE" term="%22Research+funding%22">Research funding</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Ethnology%22">Ethnology</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Minorities%22">Minorities</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink> – Name: SubjectGeographic Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22China%22">China</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Purpose: This study evaluates the capabilities and limitations of large language models (LLMs) in classifying Chinese ethic minority books under the scheme of Chinese Library Classification. Design/methodology/approach: A test collection of Chinese ethnic minority bibliographic records was constructed, and prompt engineering was used to compare the classification performance of DeepSeek-v3 and ChatGPT-4o under two input scenarios: "title + abstract" and "title only." By designing evaluation metrics that include accuracy, granularity and error-type analysis, this study systematically evaluates the performance differences between the models, diagnoses the causes of errors and proposes improvement strategies. Findings: Experimental results show both models performed well in the broad category classification of Chinese ethnic minority books, with DeepSeek-v3 exceeding 80% accuracy. Incorporating abstracts further improved accuracy and prompted longer, more detailed classification codes. However, accuracy declined for both as classification codes grew more specific. DeepSeek-v3 significantly outperformed ChatGPT-4o, achieving an overall accuracy of 40.78% and 33.50% with and without abstracts, respectively, while ChatGPT-4o remained below 6%. On the basis of classification error analysis, this study proposes improvements in classification system design, model capability enhancement and human–artificial intelligence (AI) collaboration to guide practical improvements in organizing ethnic minority resources. Originality/value: Combining librarianship, ethnography and artificial intelligence, this study is the first to compare the classification ability of different large-language models for Chinese ethnic minority books. It reveals cultural limitations in knowledge organization systems, identifies the "capability threshold" of LLMs in cultural context processing and establishes an empirical basis for developing culturally-aware AI governance frameworks. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Electronic Library is the property of Emerald Publishing Limited and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=192696103 |
| RecordInfo | BibRecord: BibEntity: Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 20 StartPage: 251 Subjects: – SubjectFull: Library science Type: general – SubjectFull: Artificial intelligence Type: general – SubjectFull: Books Type: general – SubjectFull: Experimental design Type: general – SubjectFull: Information retrieval Type: general – SubjectFull: Comparative studies Type: general – SubjectFull: Classification of books Type: general – SubjectFull: Research funding Type: general – SubjectFull: Natural language processing Type: general – SubjectFull: Ethnology Type: general – SubjectFull: Descriptive statistics Type: general – SubjectFull: Minorities Type: general – SubjectFull: Data analysis software Type: general – SubjectFull: China Type: general Titles: – TitleFull: Chinese ethnic minority book classification by large language models within CLC. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Jia, Junzhi – PersonEntity: Name: NameFull: Xu, Hui – PersonEntity: Name: NameFull: Guo, Yiqian – PersonEntity: Name: NameFull: Gao, Wenjing IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Text: 2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 02640473 Numbering: – Type: volume Value: 44 – Type: issue Value: 2 Titles: – TitleFull: Electronic Library Type: main |
| ResultId | 1 |