Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study.

Saved in:
Bibliographic Details
Title: Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study.
Authors: Gireesh, Haritha (AUTHOR), Shukla, Lekhansh (AUTHOR), Shivaprakash, Prakrithi (AUTHOR), Mukherjee, Animesh (AUTHOR), Chand, Prabhat (AUTHOR), Murthy, Pratima (AUTHOR)
Source: Drug & Alcohol Review. Jan2026, Vol. 45 Issue 1, p1-10. 10p.
Subjects: Language models, Natural language processing, Substance abuse, Data mining, Proofreading, Drug addiction, Medical records, Electronic health records
Abstract: Introduction: Electronic health records contain both structured and unstructured data, with unstructured clinical notes widely used in addiction psychiatry. Clinical notes have numerous errors and require proofreading to ensure accuracy and readability. This study evaluates natural language processing methods and adapts a Large Language Model (LLM) for proofreading clinical notes and extracting substance‐related information. Methods: We analysed clinical notes from a 5‐year addiction medicine electronic health record dataset (2018–2023), selecting 6500 notes. The proofreading task involved correcting spelling and expanding abbreviations, while information extraction identified the presence of substance use and quantified the time since last use. Annotations by a team of doctors and nurses provided the gold standard. Against this, we compared the performance of existing solutions, including LLMs, and adapted an LLM for these tasks. The final model (fine‐tuned LLAMA‐3.2‐3b) is also compared against a state‐of‐the‐art commercial model (Generative Pretrained Transformer‐4‐o), and a human‐preference experiment is done with masked raters choosing between model‐generated and human‐generated proofread versions. Results: Proofreading improved readability and decreased out‐of‐vocabulary words. LLM‐based solutions outperformed simpler approaches. The fine‐tuned model outperformed the Generative Pretrained Transformer‐4‐o on both tasks. Masked human evaluators chose model‐corrected clinical notes over the human‐corrected version in 62% of trials (p < 0.001). On the information extraction task, while the overall performance is satisfactory (Mean F1 0.99), it is poor on rarer substance classes like hallucinogens. Discussion and Conclusions: Fine‐tuned LLMs effectively standardised clinical notes and extracted structured information from addiction psychiatry records. Both these functionalities have important applications. Standardising improves the readability of clinical documentation and facilitates communication within and between interdisciplinary teams. Automated information extraction can decrease the burden on clinical staff, allow the creation of research cohorts from existing records and improve treatment outcomes by extracting critical information, such as 'time since last drink', which can be used to raise alerts. Even with limited computational resources, it is possible to adapt open‐source LLMs for bespoke tasks in the field of addiction psychiatry. Our proposed solution is a model that can be deployed on consumer‐grade servers, thus ensuring data privacy and security. [ABSTRACT FROM AUTHOR]
Copyright of Drug & Alcohol Review is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
FullText Text:
  Availability: 0
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 191183418
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study.
– Name: Author
  Label: Authors
  Group: Au
  Data: &lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Gireesh%2C+Haritha%22&quot;&gt;Gireesh, Haritha&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Shukla%2C+Lekhansh%22&quot;&gt;Shukla, Lekhansh&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Shivaprakash%2C+Prakrithi%22&quot;&gt;Shivaprakash, Prakrithi&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Mukherjee%2C+Animesh%22&quot;&gt;Mukherjee, Animesh&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Chand%2C+Prabhat%22&quot;&gt;Chand, Prabhat&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Murthy%2C+Pratima%22&quot;&gt;Murthy, Pratima&lt;/searchLink&gt; (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: &lt;searchLink fieldCode=&quot;JN&quot; term=&quot;%22Drug+%26+Alcohol+Review%22&quot;&gt;Drug &amp; Alcohol Review&lt;/searchLink&gt;. Jan2026, Vol. 45 Issue 1, p1-10. 10p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Language+models%22&quot;&gt;Language models&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Natural+language+processing%22&quot;&gt;Natural language processing&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Substance+abuse%22&quot;&gt;Substance abuse&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Data+mining%22&quot;&gt;Data mining&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Proofreading%22&quot;&gt;Proofreading&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Drug+addiction%22&quot;&gt;Drug addiction&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Medical+records%22&quot;&gt;Medical records&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Electronic+health+records%22&quot;&gt;Electronic health records&lt;/searchLink&gt;
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Introduction: Electronic health records contain both structured and unstructured data, with unstructured clinical notes widely used in addiction psychiatry. Clinical notes have numerous errors and require proofreading to ensure accuracy and readability. This study evaluates natural language processing methods and adapts a Large Language Model (LLM) for proofreading clinical notes and extracting substance‐related information. Methods: We analysed clinical notes from a 5‐year addiction medicine electronic health record dataset (2018–2023), selecting 6500 notes. The proofreading task involved correcting spelling and expanding abbreviations, while information extraction identified the presence of substance use and quantified the time since last use. Annotations by a team of doctors and nurses provided the gold standard. Against this, we compared the performance of existing solutions, including LLMs, and adapted an LLM for these tasks. The final model (fine‐tuned LLAMA‐3.2‐3b) is also compared against a state‐of‐the‐art commercial model (Generative Pretrained Transformer‐4‐o), and a human‐preference experiment is done with masked raters choosing between model‐generated and human‐generated proofread versions. Results: Proofreading improved readability and decreased out‐of‐vocabulary words. LLM‐based solutions outperformed simpler approaches. The fine‐tuned model outperformed the Generative Pretrained Transformer‐4‐o on both tasks. Masked human evaluators chose model‐corrected clinical notes over the human‐corrected version in 62% of trials (p &lt; 0.001). On the information extraction task, while the overall performance is satisfactory (Mean F1 0.99), it is poor on rarer substance classes like hallucinogens. Discussion and Conclusions: Fine‐tuned LLMs effectively standardised clinical notes and extracted structured information from addiction psychiatry records. Both these functionalities have important applications. Standardising improves the readability of clinical documentation and facilitates communication within and between interdisciplinary teams. Automated information extraction can decrease the burden on clinical staff, allow the creation of research cohorts from existing records and improve treatment outcomes by extracting critical information, such as &#39;time since last drink&#39;, which can be used to raise alerts. Even with limited computational resources, it is possible to adapt open‐source LLMs for bespoke tasks in the field of addiction psychiatry. Our proposed solution is a model that can be deployed on consumer‐grade servers, thus ensuring data privacy and security. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: &lt;i&gt;Copyright of Drug &amp; Alcohol Review is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder&#39;s express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.&lt;/i&gt; (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=191183418
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/dar.70059
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 10
        StartPage: 1
    Subjects:
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Substance abuse
        Type: general
      – SubjectFull: Data mining
        Type: general
      – SubjectFull: Proofreading
        Type: general
      – SubjectFull: Drug addiction
        Type: general
      – SubjectFull: Medical records
        Type: general
      – SubjectFull: Electronic health records
        Type: general
    Titles:
      – TitleFull: Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Gireesh, Haritha
      – PersonEntity:
          Name:
            NameFull: Shukla, Lekhansh
      – PersonEntity:
          Name:
            NameFull: Shivaprakash, Prakrithi
      – PersonEntity:
          Name:
            NameFull: Mukherjee, Animesh
      – PersonEntity:
          Name:
            NameFull: Chand, Prabhat
      – PersonEntity:
          Name:
            NameFull: Murthy, Pratima
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: Jan2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 09595236
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Drug & Alcohol Review
              Type: main
ResultId 1