A note on knowledge discovery and machine learning in digital soil mapping.

Saved in:
Bibliographic Details
Title: A note on knowledge discovery and machine learning in digital soil mapping.
Authors: Wadoux, Alexandre M. J.‐C.1 (AUTHOR) alexandre.wadoux@wur.nl, Samuel‐Rosa, Alessandro2 (AUTHOR), Poggio, Laura3 (AUTHOR), Mulder, Vera Leatitia1 (AUTHOR)
Source: European Journal of Soil Science. Mar2020, Vol. 71 Issue 2, p133-136. 4p.
Subjects: Digital soil mapping, Machine learning, Pattern recognition systems, Histosols
Abstract: In digital soil mapping, machine learning (ML) techniques are being used to infer a relationship between a soil property and the covariates. The information derived from this process is often translated into pedological knowledge. This mechanism is referred to as knowledge discovery. This study shows that knowledge discovery based on ML must be treated with caution. We show how pseudo‐covariates can be used to accurately predict soil organic carbon in a hypothetical case study. We demonstrate that ML methods can find relevant patterns even when the covariates are meaningless and not related to soil‐forming factors and processes. We argue that pattern recognition for prediction should not be equated with knowledge discovery. Knowledge discovery requires more than the recognition of patterns and successful prediction. It requires the pre‐selection and preprocessing of pedologically relevant environmental covariates and the posterior interpretation and evaluation of the recognized patterns. We argue that important ML covariates could serve the purpose of providing elements to postulate hypotheses about soil processes that, once validated through experiments, could result in new pedological knowledge. Highlights: We discuss the rationale of knowledge discovery based on the most important machine learning covariatesWe use pseudo‐covariates to predict topsoil organic carbon with random forestSoil organic carbon was accurately predicted in a hypothetical case studyPattern recognition by random forest should not be equated to knowledge discovery [ABSTRACT FROM AUTHOR]
Copyright of European Journal of Soil Science is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:In digital soil mapping, machine learning (ML) techniques are being used to infer a relationship between a soil property and the covariates. The information derived from this process is often translated into pedological knowledge. This mechanism is referred to as knowledge discovery. This study shows that knowledge discovery based on ML must be treated with caution. We show how pseudo‐covariates can be used to accurately predict soil organic carbon in a hypothetical case study. We demonstrate that ML methods can find relevant patterns even when the covariates are meaningless and not related to soil‐forming factors and processes. We argue that pattern recognition for prediction should not be equated with knowledge discovery. Knowledge discovery requires more than the recognition of patterns and successful prediction. It requires the pre‐selection and preprocessing of pedologically relevant environmental covariates and the posterior interpretation and evaluation of the recognized patterns. We argue that important ML covariates could serve the purpose of providing elements to postulate hypotheses about soil processes that, once validated through experiments, could result in new pedological knowledge. Highlights: We discuss the rationale of knowledge discovery based on the most important machine learning covariatesWe use pseudo‐covariates to predict topsoil organic carbon with random forestSoil organic carbon was accurately predicted in a hypothetical case studyPattern recognition by random forest should not be equated to knowledge discovery [ABSTRACT FROM AUTHOR]
ISSN:13510754
DOI:10.1111/ejss.12909