Enhancing machine learning prediction of biomass pyrolysis yields through taxonomy-based preprocessing and rigorous model tuning.
Saved in:
| Title: | Enhancing machine learning prediction of biomass pyrolysis yields through taxonomy-based preprocessing and rigorous model tuning. |
|---|---|
| Authors: | Kristanto, Jonas1,2 (AUTHOR) kristanto.jonas@ugm.ac.id, Azis, Muhammad Mufti1,2 (AUTHOR), Purwono, Suryo1,3 (AUTHOR) |
| Source: | Biomass & Bioenergy. Feb2026, Vol. 205, pN.PAG-N.PAG. 1p. |
| Subjects: | Machine learning, Pyrolysis, Random forest algorithms, Forecasting, Boosting algorithms, Statistical models |
| Abstract: | Accurate prediction of biomass pyrolysis yields remains challenging because of the diversity of feedstocks and process conditions. This study developed machine learning models using a dataset of 1006 entries and introduced a preprocessing strategy based on feedstock taxonomy ranks. Linear regression, random forest, and gradient boosting algorithms were employed under systematic tuning and Monte Carlo cross validation to evaluate the effect of incorporating taxonomy ranks on model performance. Random forest achieved the most reliable performance, with mean test R2 values of 0.80, 0.78, and 0.61 for solid, liquid, and gas yields, respectively. The inclusion of taxonomy ranks improved predictive performance across all models, increasing the coefficient of determination by up to 60 % for linear regression and up to 6.4 % for random forest and gradient boosting, while also reducing average and relative errors consistently. Comparable or higher accuracy was obtained with lower computational demands and simpler model configurations. These results show that the inclusion of taxonomy ranks in preprocessing, together with cross-validation and model tuning, enhances generalization and ensures that the findings are reproducible in data-driven pyrolysis modeling. • Taxonomy ranks improved ML datasets for predicting biomass pyrolysis yields. • Missing values were imputed using averages from the same taxonomy rank. • RF and GB models were tuned through a stratified rigorous method. • RF and GB feature importances with LR modeling improved interpretability. • Temperature, volatile, carbon, and ash content drive pyrolysis yield outcomes. [ABSTRACT FROM AUTHOR] |
| Copyright of Biomass & Bioenergy is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 190825767 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Enhancing machine learning prediction of biomass pyrolysis yields through taxonomy-based preprocessing and rigorous model tuning. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Kristanto%2C+Jonas%22">Kristanto, Jonas</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> kristanto.jonas@ugm.ac.id</i><br /><searchLink fieldCode="AR" term="%22Azis%2C+Muhammad+Mufti%22">Azis, Muhammad Mufti</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Purwono%2C+Suryo%22">Purwono, Suryo</searchLink><relatesTo>1,3</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Biomass+%26+Bioenergy%22">Biomass & Bioenergy</searchLink>. Feb2026, Vol. 205, pN.PAG-N.PAG. 1p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Pyrolysis%22">Pyrolysis</searchLink><br /><searchLink fieldCode="DE" term="%22Random+forest+algorithms%22">Random forest algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Forecasting%22">Forecasting</searchLink><br /><searchLink fieldCode="DE" term="%22Boosting+algorithms%22">Boosting algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+models%22">Statistical models</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Accurate prediction of biomass pyrolysis yields remains challenging because of the diversity of feedstocks and process conditions. This study developed machine learning models using a dataset of 1006 entries and introduced a preprocessing strategy based on feedstock taxonomy ranks. Linear regression, random forest, and gradient boosting algorithms were employed under systematic tuning and Monte Carlo cross validation to evaluate the effect of incorporating taxonomy ranks on model performance. Random forest achieved the most reliable performance, with mean test R2 values of 0.80, 0.78, and 0.61 for solid, liquid, and gas yields, respectively. The inclusion of taxonomy ranks improved predictive performance across all models, increasing the coefficient of determination by up to 60 % for linear regression and up to 6.4 % for random forest and gradient boosting, while also reducing average and relative errors consistently. Comparable or higher accuracy was obtained with lower computational demands and simpler model configurations. These results show that the inclusion of taxonomy ranks in preprocessing, together with cross-validation and model tuning, enhances generalization and ensures that the findings are reproducible in data-driven pyrolysis modeling. • Taxonomy ranks improved ML datasets for predicting biomass pyrolysis yields. • Missing values were imputed using averages from the same taxonomy rank. • RF and GB models were tuned through a stratified rigorous method. • RF and GB feature importances with LR modeling improved interpretability. • Temperature, volatile, carbon, and ash content drive pyrolysis yield outcomes. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Biomass & Bioenergy is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=190825767 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.biombioe.2025.108524 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 1 StartPage: N.PAG Subjects: – SubjectFull: Machine learning Type: general – SubjectFull: Pyrolysis Type: general – SubjectFull: Random forest algorithms Type: general – SubjectFull: Forecasting Type: general – SubjectFull: Boosting algorithms Type: general – SubjectFull: Statistical models Type: general Titles: – TitleFull: Enhancing machine learning prediction of biomass pyrolysis yields through taxonomy-based preprocessing and rigorous model tuning. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Kristanto, Jonas – PersonEntity: Name: NameFull: Azis, Muhammad Mufti – PersonEntity: Name: NameFull: Purwono, Suryo IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 02 Text: Feb2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 09619534 Numbering: – Type: volume Value: 205 Titles: – TitleFull: Biomass & Bioenergy Type: main |
| ResultId | 1 |