Beam angle optimization for radiotherapy using LLMs via reinforcement‐learning inspired iterative refinement.
Saved in:
| Title: | Beam angle optimization for radiotherapy using LLMs via reinforcement‐learning inspired iterative refinement. |
|---|---|
| Authors: | Cammarota, Sara1 (AUTHOR) sara.cammarota@uniroma2.eu, Ferrante, Matteo1 (AUTHOR), Carosi, Alessandra2 (AUTHOR), D'Angelillo, Rolando Maria3 (AUTHOR), Toschi, Nicola1,4 (AUTHOR) |
| Source: | Medical Physics. Feb2026, Vol. 53 Issue 2, p1-12. 12p. |
| Subjects: | Radiotherapy treatment planning, Reinforcement learning, Optimization algorithms, Radiotherapy, Language models, Therapeutics, Generative pre-trained transformers, Radiation doses |
| Abstract: | Background: Radiotherapy treatment planning (TP) aims to maximize radiation dose delivered to tumors while minimizing exposure to surrounding healthy tissues. Beam angle optimization (BAO) is a crucial component of TP, characterized by high dimensionality and non‐convexity, and is traditionally solved via heuristic or manual iterative approaches. These conventional methods are time‐consuming and often yield suboptimal solutions due to incomplete exploration of the vast solution space. Purpose: This study introduces a novel framework integrating a general‐purpose large language model (LLM) within a reinforcementlearning (RL)‐inspired iterative strategy to automate BAO in radiotherapy planning. Taking advantage of the inherent knowledge embedded in LLMs, the method uses visual and scalar feedback to produce clinically meaningful treatment plans without requiring any domain‐specific fine‐tuning or additional training. Methods: The proposed framework employs an off‐the‐shelf Generative Pre‐trained Transformer, GPT‐4 model (denoted GPT‐4o) in an inference‐only setting. At each iteration, GPT‐4o suggests a set of gantry angles, which are subsequently input into the MatRad software to generate a dose distribution. A scalar reward is computed from this distribution using a custom reward function designed to balance target dose conformity and sparing of organs‐at‐risk (OARs). This reward, along with the corresponding dose maps, serves as feedback for the LLM to iteratively refine its suggestions. The refinement process consists of distinct exploration and exploitation phases inspired by classical RL paradigms. We evaluated six configurations that varied in exploration duration and in the Computed Tomography (CT) slice inputs provided to the LLM (Single‐View vs. Multi‐View). Performance was benchmarked against a random‐angle selection baseline across three anatomical sites: prostate, head‐and‐neck, and liver. Results: Across the liver and head‐and‐neck cases, all LLM‐based configurations significantly outperformed the random baseline (p<0.05$p<0.05$). In the prostate scenario, most strategies demonstrated statistically significant improvements, except for the Multi‐View configurations with extended exploration phases (10 and 15 iterations). Rewards consistently increased during the exploitation phase, and the resulting dose–volume histograms and dose distributions exhibited improved conformity to target volumes with enhanced sparing of OARs. Notably, plans of clinically plausible quality were obtained within 20 iterative refinement steps in this proof‐of‐concept setting. Conclusions: This study demonstrates that general‐purpose LLMs, operating without specialized model training or fine‐tuning, can effectively serve as intelligent agents for automated radiotherapy TP, specifically addressing the BAO problem. This flexible and scalable framework has the potential to enhance clinical decision‐making workflows in radiotherapy. Future research directions include exploring more comprehensive and clinically nuanced reward functions and extending the methodology to other components of radiotherapy TP. [ABSTRACT FROM AUTHOR] |
| Copyright of Medical Physics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 191805140 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Beam angle optimization for radiotherapy using LLMs via reinforcement‐learning inspired iterative refinement. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Cammarota%2C+Sara%22">Cammarota, Sara</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> sara.cammarota@uniroma2.eu</i><br /><searchLink fieldCode="AR" term="%22Ferrante%2C+Matteo%22">Ferrante, Matteo</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Carosi%2C+Alessandra%22">Carosi, Alessandra</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22D'Angelillo%2C+Rolando+Maria%22">D'Angelillo, Rolando Maria</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Toschi%2C+Nicola%22">Toschi, Nicola</searchLink><relatesTo>1,4</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Medical+Physics%22">Medical Physics</searchLink>. Feb2026, Vol. 53 Issue 2, p1-12. 12p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Radiotherapy+treatment+planning%22">Radiotherapy treatment planning</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Optimization+algorithms%22">Optimization algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Radiotherapy%22">Radiotherapy</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Therapeutics%22">Therapeutics</searchLink><br /><searchLink fieldCode="DE" term="%22Generative+pre-trained+transformers%22">Generative pre-trained transformers</searchLink><br /><searchLink fieldCode="DE" term="%22Radiation+doses%22">Radiation doses</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Background: Radiotherapy treatment planning (TP) aims to maximize radiation dose delivered to tumors while minimizing exposure to surrounding healthy tissues. Beam angle optimization (BAO) is a crucial component of TP, characterized by high dimensionality and non‐convexity, and is traditionally solved via heuristic or manual iterative approaches. These conventional methods are time‐consuming and often yield suboptimal solutions due to incomplete exploration of the vast solution space. Purpose: This study introduces a novel framework integrating a general‐purpose large language model (LLM) within a reinforcementlearning (RL)‐inspired iterative strategy to automate BAO in radiotherapy planning. Taking advantage of the inherent knowledge embedded in LLMs, the method uses visual and scalar feedback to produce clinically meaningful treatment plans without requiring any domain‐specific fine‐tuning or additional training. Methods: The proposed framework employs an off‐the‐shelf Generative Pre‐trained Transformer, GPT‐4 model (denoted GPT‐4o) in an inference‐only setting. At each iteration, GPT‐4o suggests a set of gantry angles, which are subsequently input into the MatRad software to generate a dose distribution. A scalar reward is computed from this distribution using a custom reward function designed to balance target dose conformity and sparing of organs‐at‐risk (OARs). This reward, along with the corresponding dose maps, serves as feedback for the LLM to iteratively refine its suggestions. The refinement process consists of distinct exploration and exploitation phases inspired by classical RL paradigms. We evaluated six configurations that varied in exploration duration and in the Computed Tomography (CT) slice inputs provided to the LLM (Single‐View vs. Multi‐View). Performance was benchmarked against a random‐angle selection baseline across three anatomical sites: prostate, head‐and‐neck, and liver. Results: Across the liver and head‐and‐neck cases, all LLM‐based configurations significantly outperformed the random baseline (p<0.05$p<0.05$). In the prostate scenario, most strategies demonstrated statistically significant improvements, except for the Multi‐View configurations with extended exploration phases (10 and 15 iterations). Rewards consistently increased during the exploitation phase, and the resulting dose–volume histograms and dose distributions exhibited improved conformity to target volumes with enhanced sparing of OARs. Notably, plans of clinically plausible quality were obtained within 20 iterative refinement steps in this proof‐of‐concept setting. Conclusions: This study demonstrates that general‐purpose LLMs, operating without specialized model training or fine‐tuning, can effectively serve as intelligent agents for automated radiotherapy TP, specifically addressing the BAO problem. This flexible and scalable framework has the potential to enhance clinical decision‐making workflows in radiotherapy. Future research directions include exploring more comprehensive and clinically nuanced reward functions and extending the methodology to other components of radiotherapy TP. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Medical Physics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=191805140 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/mp.70258 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 1 Subjects: – SubjectFull: Radiotherapy treatment planning Type: general – SubjectFull: Reinforcement learning Type: general – SubjectFull: Optimization algorithms Type: general – SubjectFull: Radiotherapy Type: general – SubjectFull: Language models Type: general – SubjectFull: Therapeutics Type: general – SubjectFull: Generative pre-trained transformers Type: general – SubjectFull: Radiation doses Type: general Titles: – TitleFull: Beam angle optimization for radiotherapy using LLMs via reinforcement‐learning inspired iterative refinement. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Cammarota, Sara – PersonEntity: Name: NameFull: Ferrante, Matteo – PersonEntity: Name: NameFull: Carosi, Alessandra – PersonEntity: Name: NameFull: D'Angelillo, Rolando Maria – PersonEntity: Name: NameFull: Toschi, Nicola IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 02 Text: Feb2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 00942405 Numbering: – Type: volume Value: 53 – Type: issue Value: 2 Titles: – TitleFull: Medical Physics Type: main |
| ResultId | 1 |