From Data to Deployment: A Comprehensive Analysis of Risks in Large Language Model Research and Development.

Saved in:
Bibliographic Details
Title: From Data to Deployment: A Comprehensive Analysis of Risks in Large Language Model Research and Development.
Authors: Zhang, Tianshu1 (AUTHOR), Su, Ruidan1 (AUTHOR) suruidan@sjtu.edu.cn, Zhong, Anli1 (AUTHOR), Fang, Minwei1 (AUTHOR), Zhang, Yu-dong2 (AUTHOR) yudongzhang@ieee.org, Tian, Jiwei (AUTHOR) jiweitian@xjtu.edu.cn
Source: IET Information Security (Wiley-Blackwell). 6/23/2025, Vol. 2025, p1-13. 13p.
Subjects: Language models, Information processing, Research personnel, Language research, Risk assessment
Abstract: Large language models (LLMs) have evolved significantly, achieving unprecedented linguistic capabilities that underpin a wide range of AI applications. However, they also pose risks and challenges such as ethical concerns, bias and computational sustainability. How to balance the high performance in revolutionising information processing with the risks they pose is critical to their future development. LLM is a type of NLP model and many of the LLM risks are also risks that NLP has experienced in the past. We, therefore, summarise these risks, focusing more on the underlying understanding of these risks/technical tools, rather than simply describing their occurrence in LLM. In this paper, we first discuss and compare the current state of research on the four main risks in the process of developing LLMs: data, system, pretraining and inference, and then, try to summarise the rationale, complexity, prospects and challenges of the key issues and challenges in each phase. Finally, this review concludes with a discussion of the fundamental issues that should be of most concern and risk and that should be addressed in the early stages of modelling research, including the correlated issues of privacy preservation and countering attacks and model robustness. Based on the LLM research and development (R&D) process perspective, this review summarises the actual risks and provides guidance for research directions, with the aim of helping researchers to identify these risk points and technology directions worth investigating, as well as helping to establish a safe and efficient R&D process. [ABSTRACT FROM AUTHOR]
Copyright of IET Information Security (Wiley-Blackwell) is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 186137297
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: From Data to Deployment: A Comprehensive Analysis of Risks in Large Language Model Research and Development.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Tianshu%22">Zhang, Tianshu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Su%2C+Ruidan%22">Su, Ruidan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> suruidan@sjtu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Zhong%2C+Anli%22">Zhong, Anli</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Fang%2C+Minwei%22">Fang, Minwei</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhang%2C+Yu-dong%22">Zhang, Yu-dong</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> yudongzhang@ieee.org</i><br /><searchLink fieldCode="AR" term="%22Tian%2C+Jiwei%22">Tian, Jiwei</searchLink> (AUTHOR)<i> jiweitian@xjtu.edu.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IET+Information+Security+%28Wiley-Blackwell%29%22">IET Information Security (Wiley-Blackwell)</searchLink>. 6/23/2025, Vol. 2025, p1-13. 13p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Information+processing%22">Information processing</searchLink><br /><searchLink fieldCode="DE" term="%22Research+personnel%22">Research personnel</searchLink><br /><searchLink fieldCode="DE" term="%22Language+research%22">Language research</searchLink><br /><searchLink fieldCode="DE" term="%22Risk+assessment%22">Risk assessment</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Large language models (LLMs) have evolved significantly, achieving unprecedented linguistic capabilities that underpin a wide range of AI applications. However, they also pose risks and challenges such as ethical concerns, bias and computational sustainability. How to balance the high performance in revolutionising information processing with the risks they pose is critical to their future development. LLM is a type of NLP model and many of the LLM risks are also risks that NLP has experienced in the past. We, therefore, summarise these risks, focusing more on the underlying understanding of these risks/technical tools, rather than simply describing their occurrence in LLM. In this paper, we first discuss and compare the current state of research on the four main risks in the process of developing LLMs: data, system, pretraining and inference, and then, try to summarise the rationale, complexity, prospects and challenges of the key issues and challenges in each phase. Finally, this review concludes with a discussion of the fundamental issues that should be of most concern and risk and that should be addressed in the early stages of modelling research, including the correlated issues of privacy preservation and countering attacks and model robustness. Based on the LLM research and development (R&D) process perspective, this review summarises the actual risks and provides guidance for research directions, with the aim of helping researchers to identify these risk points and technology directions worth investigating, as well as helping to establish a safe and efficient R&D process. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IET Information Security (Wiley-Blackwell) is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=186137297
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1049/ise2/7358963
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 1
    Subjects:
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Information processing
        Type: general
      – SubjectFull: Research personnel
        Type: general
      – SubjectFull: Language research
        Type: general
      – SubjectFull: Risk assessment
        Type: general
    Titles:
      – TitleFull: From Data to Deployment: A Comprehensive Analysis of Risks in Large Language Model Research and Development.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zhang, Tianshu
      – PersonEntity:
          Name:
            NameFull: Su, Ruidan
      – PersonEntity:
          Name:
            NameFull: Zhong, Anli
      – PersonEntity:
          Name:
            NameFull: Fang, Minwei
      – PersonEntity:
          Name:
            NameFull: Zhang, Yu-dong
      – PersonEntity:
          Name:
            NameFull: Tian, Jiwei
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 23
              M: 06
              Text: 6/23/2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 17518709
          Numbering:
            – Type: volume
              Value: 2025
          Titles:
            – TitleFull: IET Information Security (Wiley-Blackwell)
              Type: main
ResultId 1