Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study.
Saved in:
| Title: | Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study. |
|---|---|
| Authors: | Gireesh, Haritha (AUTHOR), Shukla, Lekhansh (AUTHOR), Shivaprakash, Prakrithi (AUTHOR), Mukherjee, Animesh (AUTHOR), Chand, Prabhat (AUTHOR), Murthy, Pratima (AUTHOR) |
| Source: | Drug & Alcohol Review. Jan2026, Vol. 45 Issue 1, p1-10. 10p. |
| Subjects: | Language models, Natural language processing, Substance abuse, Data mining, Proofreading, Drug addiction, Medical records, Electronic health records |
| Abstract: | Introduction: Electronic health records contain both structured and unstructured data, with unstructured clinical notes widely used in addiction psychiatry. Clinical notes have numerous errors and require proofreading to ensure accuracy and readability. This study evaluates natural language processing methods and adapts a Large Language Model (LLM) for proofreading clinical notes and extracting substance‐related information. Methods: We analysed clinical notes from a 5‐year addiction medicine electronic health record dataset (2018–2023), selecting 6500 notes. The proofreading task involved correcting spelling and expanding abbreviations, while information extraction identified the presence of substance use and quantified the time since last use. Annotations by a team of doctors and nurses provided the gold standard. Against this, we compared the performance of existing solutions, including LLMs, and adapted an LLM for these tasks. The final model (fine‐tuned LLAMA‐3.2‐3b) is also compared against a state‐of‐the‐art commercial model (Generative Pretrained Transformer‐4‐o), and a human‐preference experiment is done with masked raters choosing between model‐generated and human‐generated proofread versions. Results: Proofreading improved readability and decreased out‐of‐vocabulary words. LLM‐based solutions outperformed simpler approaches. The fine‐tuned model outperformed the Generative Pretrained Transformer‐4‐o on both tasks. Masked human evaluators chose model‐corrected clinical notes over the human‐corrected version in 62% of trials (p < 0.001). On the information extraction task, while the overall performance is satisfactory (Mean F1 0.99), it is poor on rarer substance classes like hallucinogens. Discussion and Conclusions: Fine‐tuned LLMs effectively standardised clinical notes and extracted structured information from addiction psychiatry records. Both these functionalities have important applications. Standardising improves the readability of clinical documentation and facilitates communication within and between interdisciplinary teams. Automated information extraction can decrease the burden on clinical staff, allow the creation of research cohorts from existing records and improve treatment outcomes by extracting critical information, such as 'time since last drink', which can be used to raise alerts. Even with limited computational resources, it is possible to adapt open‐source LLMs for bespoke tasks in the field of addiction psychiatry. Our proposed solution is a model that can be deployed on consumer‐grade servers, thus ensuring data privacy and security. [ABSTRACT FROM AUTHOR] |
| Copyright of Drug & Alcohol Review is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: pbh DbLabel: Psychology and Behavioral Sciences Collection An: 191183418 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Gireesh%2C+Haritha%22">Gireesh, Haritha</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Shukla%2C+Lekhansh%22">Shukla, Lekhansh</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Shivaprakash%2C+Prakrithi%22">Shivaprakash, Prakrithi</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Mukherjee%2C+Animesh%22">Mukherjee, Animesh</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Chand%2C+Prabhat%22">Chand, Prabhat</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Murthy%2C+Pratima%22">Murthy, Pratima</searchLink> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Drug+%26+Alcohol+Review%22">Drug & Alcohol Review</searchLink>. Jan2026, Vol. 45 Issue 1, p1-10. 10p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Substance+abuse%22">Substance abuse</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining%22">Data mining</searchLink><br /><searchLink fieldCode="DE" term="%22Proofreading%22">Proofreading</searchLink><br /><searchLink fieldCode="DE" term="%22Drug+addiction%22">Drug addiction</searchLink><br /><searchLink fieldCode="DE" term="%22Medical+records%22">Medical records</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+health+records%22">Electronic health records</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Introduction: Electronic health records contain both structured and unstructured data, with unstructured clinical notes widely used in addiction psychiatry. Clinical notes have numerous errors and require proofreading to ensure accuracy and readability. This study evaluates natural language processing methods and adapts a Large Language Model (LLM) for proofreading clinical notes and extracting substance‐related information. Methods: We analysed clinical notes from a 5‐year addiction medicine electronic health record dataset (2018–2023), selecting 6500 notes. The proofreading task involved correcting spelling and expanding abbreviations, while information extraction identified the presence of substance use and quantified the time since last use. Annotations by a team of doctors and nurses provided the gold standard. Against this, we compared the performance of existing solutions, including LLMs, and adapted an LLM for these tasks. The final model (fine‐tuned LLAMA‐3.2‐3b) is also compared against a state‐of‐the‐art commercial model (Generative Pretrained Transformer‐4‐o), and a human‐preference experiment is done with masked raters choosing between model‐generated and human‐generated proofread versions. Results: Proofreading improved readability and decreased out‐of‐vocabulary words. LLM‐based solutions outperformed simpler approaches. The fine‐tuned model outperformed the Generative Pretrained Transformer‐4‐o on both tasks. Masked human evaluators chose model‐corrected clinical notes over the human‐corrected version in 62% of trials (p < 0.001). On the information extraction task, while the overall performance is satisfactory (Mean F1 0.99), it is poor on rarer substance classes like hallucinogens. Discussion and Conclusions: Fine‐tuned LLMs effectively standardised clinical notes and extracted structured information from addiction psychiatry records. Both these functionalities have important applications. Standardising improves the readability of clinical documentation and facilitates communication within and between interdisciplinary teams. Automated information extraction can decrease the burden on clinical staff, allow the creation of research cohorts from existing records and improve treatment outcomes by extracting critical information, such as 'time since last drink', which can be used to raise alerts. Even with limited computational resources, it is possible to adapt open‐source LLMs for bespoke tasks in the field of addiction psychiatry. Our proposed solution is a model that can be deployed on consumer‐grade servers, thus ensuring data privacy and security. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Drug & Alcohol Review is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=191183418 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/dar.70059 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 10 StartPage: 1 Subjects: – SubjectFull: Language models Type: general – SubjectFull: Natural language processing Type: general – SubjectFull: Substance abuse Type: general – SubjectFull: Data mining Type: general – SubjectFull: Proofreading Type: general – SubjectFull: Drug addiction Type: general – SubjectFull: Medical records Type: general – SubjectFull: Electronic health records Type: general Titles: – TitleFull: Language Models for Standardising Clinical Notes and Information Extraction in Addiction Psychiatry—An Empirical Study. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Gireesh, Haritha – PersonEntity: Name: NameFull: Shukla, Lekhansh – PersonEntity: Name: NameFull: Shivaprakash, Prakrithi – PersonEntity: Name: NameFull: Mukherjee, Animesh – PersonEntity: Name: NameFull: Chand, Prabhat – PersonEntity: Name: NameFull: Murthy, Pratima IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Text: Jan2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 09595236 Numbering: – Type: volume Value: 45 – Type: issue Value: 1 Titles: – TitleFull: Drug & Alcohol Review Type: main |
| ResultId | 1 |