Chinese ethnic minority book classification by large language models within CLC.

Saved in:
Bibliographic Details
Title: Chinese ethnic minority book classification by large language models within CLC.
Authors: Jia, Junzhi1 (AUTHOR) junzhij@163.com, Xu, Hui1 (AUTHOR) xuhui111@ruc.edu.cn, Guo, Yiqian1 (AUTHOR) guoyiqian@ruc.edu.cn, Gao, Wenjing2 (AUTHOR) 437189571@qq.com
Source: Electronic Library. 2026, Vol. 44 Issue 2, p251-270. 20p.
Subject Terms: *Library science, *Artificial intelligence, *Books, *Experimental design, *Information retrieval, *Comparative studies, Classification of books, Research funding, Natural language processing, Ethnology, Descriptive statistics, Minorities, Data analysis software
Geographic Terms: China
Abstract: Purpose: This study evaluates the capabilities and limitations of large language models (LLMs) in classifying Chinese ethic minority books under the scheme of Chinese Library Classification. Design/methodology/approach: A test collection of Chinese ethnic minority bibliographic records was constructed, and prompt engineering was used to compare the classification performance of DeepSeek-v3 and ChatGPT-4o under two input scenarios: "title + abstract" and "title only." By designing evaluation metrics that include accuracy, granularity and error-type analysis, this study systematically evaluates the performance differences between the models, diagnoses the causes of errors and proposes improvement strategies. Findings: Experimental results show both models performed well in the broad category classification of Chinese ethnic minority books, with DeepSeek-v3 exceeding 80% accuracy. Incorporating abstracts further improved accuracy and prompted longer, more detailed classification codes. However, accuracy declined for both as classification codes grew more specific. DeepSeek-v3 significantly outperformed ChatGPT-4o, achieving an overall accuracy of 40.78% and 33.50% with and without abstracts, respectively, while ChatGPT-4o remained below 6%. On the basis of classification error analysis, this study proposes improvements in classification system design, model capability enhancement and human–artificial intelligence (AI) collaboration to guide practical improvements in organizing ethnic minority resources. Originality/value: Combining librarianship, ethnography and artificial intelligence, this study is the first to compare the classification ability of different large-language models for Chinese ethnic minority books. It reveals cultural limitations in knowledge organization systems, identifies the "capability threshold" of LLMs in cultural context processing and establishes an empirical basis for developing culturally-aware AI governance frameworks. [ABSTRACT FROM AUTHOR]
Copyright of Electronic Library is the property of Emerald Publishing Limited and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
FullText Text:
  Availability: 0
Header DbId: ehh
DbLabel: Education Research Complete
An: 192696103
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Chinese ethnic minority book classification by large language models within CLC.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jia%2C+Junzhi%22">Jia, Junzhi</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> junzhij@163.com</i><br /><searchLink fieldCode="AR" term="%22Xu%2C+Hui%22">Xu, Hui</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xuhui111@ruc.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Guo%2C+Yiqian%22">Guo, Yiqian</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> guoyiqian@ruc.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Gao%2C+Wenjing%22">Gao, Wenjing</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> 437189571@qq.com</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Electronic+Library%22">Electronic Library</searchLink>. 2026, Vol. 44 Issue 2, p251-270. 20p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Library+science%22">Library science</searchLink><br />*<searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br />*<searchLink fieldCode="DE" term="%22Books%22">Books</searchLink><br />*<searchLink fieldCode="DE" term="%22Experimental+design%22">Experimental design</searchLink><br />*<searchLink fieldCode="DE" term="%22Information+retrieval%22">Information retrieval</searchLink><br />*<searchLink fieldCode="DE" term="%22Comparative+studies%22">Comparative studies</searchLink><br /><searchLink fieldCode="DE" term="%22Classification+of+books%22">Classification of books</searchLink><br /><searchLink fieldCode="DE" term="%22Research+funding%22">Research funding</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Ethnology%22">Ethnology</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Minorities%22">Minorities</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink>
– Name: SubjectGeographic
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22China%22">China</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Purpose: This study evaluates the capabilities and limitations of large language models (LLMs) in classifying Chinese ethic minority books under the scheme of Chinese Library Classification. Design/methodology/approach: A test collection of Chinese ethnic minority bibliographic records was constructed, and prompt engineering was used to compare the classification performance of DeepSeek-v3 and ChatGPT-4o under two input scenarios: "title + abstract" and "title only." By designing evaluation metrics that include accuracy, granularity and error-type analysis, this study systematically evaluates the performance differences between the models, diagnoses the causes of errors and proposes improvement strategies. Findings: Experimental results show both models performed well in the broad category classification of Chinese ethnic minority books, with DeepSeek-v3 exceeding 80% accuracy. Incorporating abstracts further improved accuracy and prompted longer, more detailed classification codes. However, accuracy declined for both as classification codes grew more specific. DeepSeek-v3 significantly outperformed ChatGPT-4o, achieving an overall accuracy of 40.78% and 33.50% with and without abstracts, respectively, while ChatGPT-4o remained below 6%. On the basis of classification error analysis, this study proposes improvements in classification system design, model capability enhancement and human–artificial intelligence (AI) collaboration to guide practical improvements in organizing ethnic minority resources. Originality/value: Combining librarianship, ethnography and artificial intelligence, this study is the first to compare the classification ability of different large-language models for Chinese ethnic minority books. It reveals cultural limitations in knowledge organization systems, identifies the "capability threshold" of LLMs in cultural context processing and establishes an empirical basis for developing culturally-aware AI governance frameworks. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Electronic Library is the property of Emerald Publishing Limited and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=192696103
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 251
    Subjects:
      – SubjectFull: Library science
        Type: general
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Books
        Type: general
      – SubjectFull: Experimental design
        Type: general
      – SubjectFull: Information retrieval
        Type: general
      – SubjectFull: Comparative studies
        Type: general
      – SubjectFull: Classification of books
        Type: general
      – SubjectFull: Research funding
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Ethnology
        Type: general
      – SubjectFull: Descriptive statistics
        Type: general
      – SubjectFull: Minorities
        Type: general
      – SubjectFull: Data analysis software
        Type: general
      – SubjectFull: China
        Type: general
    Titles:
      – TitleFull: Chinese ethnic minority book classification by large language models within CLC.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jia, Junzhi
      – PersonEntity:
          Name:
            NameFull: Xu, Hui
      – PersonEntity:
          Name:
            NameFull: Guo, Yiqian
      – PersonEntity:
          Name:
            NameFull: Gao, Wenjing
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: 2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 02640473
          Numbering:
            – Type: volume
              Value: 44
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Electronic Library
              Type: main
ResultId 1