Evolutionary-scale enzymology enables exploration of a rugged catalytic landscape.
Saved in:
| Title: | Evolutionary-scale enzymology enables exploration of a rugged catalytic landscape. |
|---|---|
| Authors: | Muir, Duncan F., Asper, Garrison P. R., Notin, Pascal, Posner, Jacob A., Marks, Debora S., Keiser, Michael J., Pinney, Margaux M. |
| Source: | Science. 6/12/2025, Vol. 388 Issue 6752, p1-16. 16p. |
| Subjects: | Enzymology, Catalysis, Adenylate kinase, Mutagenesis, Statistical correlation |
| Abstract: | Quantitatively mapping enzyme sequence-catalysis landscapes remains a critical challenge in understanding enzyme function, evolution, and design. In this study, we leveraged emerging microfluidic technology to measure catalytic constants—kcat and KM—for hundreds of diverse orthologs and mutants of adenylate kinase (ADK). We dissected this sequence-catalysis landscape's topology, navigability, and mechanistic underpinnings, revealing catalytically heterogeneous neighborhoods organized by domain architecture. These results challenge long-standing hypotheses in enzyme adaptation, demonstrating that thermophilic enzymes are not universally slower than their mesophilic counterparts. Semisupervised models that combine our data with the rich sequence representations from large protein language models predict orthologous ADK-sequence catalytic parameters better than existing approaches. Our work demonstrates a promising strategy for dissecting sequence-catalysis landscapes across enzymatic evolution, opening previously unexplored avenues for enzyme engineering and functional prediction. Editor's summary: The catalytic activity of a particular enzyme can vary widely between orthologs in different species, depending on their environment and specific metabolic needs. But how do such differences evolve and relate to specific structural features? Using a high-throughput microfluidic system, Muir et al. assayed nearly 200 orthologs of the enzyme adenylate kinase and used the resulting data to build a landscape view of catalytic activity. There was minimal correlation between growth temperature and activity, and high activity peaks were widely distributed in the landscape and likely evolved independently. Current protein language models group enzymes by structure but fail in predicting the catalytic activity landscape; however, there is potential to train models using experimental activity data. —Michael A. Funk INTRODUCTION: Enzymes catalyze the reactions underlying virtually all biological processes, yet our understanding of how sequence variation translates to enzymatic function remains incomplete. Enzyme sequence-function relationships are often viewed as a "landscape," with mutational "walks" tracing paths across peaks and valleys in catalytic performance. Most experimental studies, however, focus on narrowly defined regions of sequence space through mutagenesis. In contrast, genomic databases contain widespread natural sequence variation that could alter catalytic parameters across diverse environments (e.g., temperature). Despite this wealth of natural sequence data, large-scale quantitative measurements of catalytic constants under consistent conditions remain scarce. This lack of data hinders the development of predictive models and leaves the global topology of sequence-catalysis landscapes underexplored. RATIONALE: In this study, we investigated the sequence-catalysis landscape of the model enzyme adenylate kinase (ADK) at the evolutionary scale, sampling broadly across bacterial and archaeal phylogeny. As adaptation to different temperature environments has been suggested to underlie differences in catalytic rates for ADK, we gathered hundreds of orthologs and mutants of ADK spanning the coldest and hottest environments on Earth. Adapting a high-throughput microfluidic platform, HT-MEK, we systematically measured the Michaelis-Menten parameters kcat (catalytic constant), KM (Michaelis constant), and kcat/KM (catalytic efficiency) for these ADK sequences under consistent conditions. We further analyzed the organization and traversability of this landscape and quantified how different temperature environments shape ADK function. We then evaluated how well unsupervised deep‐learning models—trained solely on protein sequence data—capture these empirical relationships and developed machine‐learning models trained on our kinetic dataset to predict the catalytic parameters of these naturally occurring ADKs. RESULTS: Our high-throughput kinetic data revealed that ADK orthologs vary in kcat by up to three orders of magnitude—even with conserved active sites and similar predicted structures. We address long-standing evolutionary hypotheses demonstrating that thermophilic enzymes are not universally slower than their mesophilic counterparts. We dissect the topology and navigability of this sequence-catalysis landscape, showing that it is rugged, with at least three global neighborhoods organized by distinct domain architectures. Stepwise point mutations and domain swaps show that this landscape remains navigable over long evolutionary timescales through path-dependent mechanisms. Finally, we show that an unsupervised protein language model organizes ADK sequence space by structure but not catalytic activity. kcat prediction models trained on the dataset collected herein outperform prior models trained only on public databases. CONCLUSION: Combining high-throughput kinetic measurements of natural enzyme sequences with machine learning, we charted an evolutionary-scale sequence-catalysis landscape for a ubiquitous enzyme family. Our data show that high catalytic activity can arise through multiple structural solutions and is not strictly limited by thermal adaptation, challenging assumptions about universal activity–stability trade-offs. Whereas protein language models excel at capturing broad structural features, our results emphasize that experimental annotations—especially for catalytic activity—are critical for building accurate sequence-to-function models. More broadly, coupling high‐throughput enzymology with machine and deep learning may uncover insights into enzyme evolution, guide the rational design of novel enzymes, and illuminate the fundamental constraints that shape protein function across the tree of life. Leveraging evolutionary-scale enzymology to map sequence-catalysis landscapes.: Sampling enzyme sequences across phylogeny and diverse environments reveals multiple evolutionary solutions to high catalytic activity arising within different domain architectures. Although protein language models capture sequence-encoded structural organization, the catalytic landscape remains highly rugged. [ABSTRACT FROM AUTHOR] |
| Copyright of Science is the property of American Association for the Advancement of Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Quantitatively mapping enzyme sequence-catalysis landscapes remains a critical challenge in understanding enzyme function, evolution, and design. In this study, we leveraged emerging microfluidic technology to measure catalytic constants—kcat and KM—for hundreds of diverse orthologs and mutants of adenylate kinase (ADK). We dissected this sequence-catalysis landscape's topology, navigability, and mechanistic underpinnings, revealing catalytically heterogeneous neighborhoods organized by domain architecture. These results challenge long-standing hypotheses in enzyme adaptation, demonstrating that thermophilic enzymes are not universally slower than their mesophilic counterparts. Semisupervised models that combine our data with the rich sequence representations from large protein language models predict orthologous ADK-sequence catalytic parameters better than existing approaches. Our work demonstrates a promising strategy for dissecting sequence-catalysis landscapes across enzymatic evolution, opening previously unexplored avenues for enzyme engineering and functional prediction. Editor's summary: The catalytic activity of a particular enzyme can vary widely between orthologs in different species, depending on their environment and specific metabolic needs. But how do such differences evolve and relate to specific structural features? Using a high-throughput microfluidic system, Muir et al. assayed nearly 200 orthologs of the enzyme adenylate kinase and used the resulting data to build a landscape view of catalytic activity. There was minimal correlation between growth temperature and activity, and high activity peaks were widely distributed in the landscape and likely evolved independently. Current protein language models group enzymes by structure but fail in predicting the catalytic activity landscape; however, there is potential to train models using experimental activity data. —Michael A. Funk INTRODUCTION: Enzymes catalyze the reactions underlying virtually all biological processes, yet our understanding of how sequence variation translates to enzymatic function remains incomplete. Enzyme sequence-function relationships are often viewed as a "landscape," with mutational "walks" tracing paths across peaks and valleys in catalytic performance. Most experimental studies, however, focus on narrowly defined regions of sequence space through mutagenesis. In contrast, genomic databases contain widespread natural sequence variation that could alter catalytic parameters across diverse environments (e.g., temperature). Despite this wealth of natural sequence data, large-scale quantitative measurements of catalytic constants under consistent conditions remain scarce. This lack of data hinders the development of predictive models and leaves the global topology of sequence-catalysis landscapes underexplored. RATIONALE: In this study, we investigated the sequence-catalysis landscape of the model enzyme adenylate kinase (ADK) at the evolutionary scale, sampling broadly across bacterial and archaeal phylogeny. As adaptation to different temperature environments has been suggested to underlie differences in catalytic rates for ADK, we gathered hundreds of orthologs and mutants of ADK spanning the coldest and hottest environments on Earth. Adapting a high-throughput microfluidic platform, HT-MEK, we systematically measured the Michaelis-Menten parameters kcat (catalytic constant), KM (Michaelis constant), and kcat/KM (catalytic efficiency) for these ADK sequences under consistent conditions. We further analyzed the organization and traversability of this landscape and quantified how different temperature environments shape ADK function. We then evaluated how well unsupervised deep‐learning models—trained solely on protein sequence data—capture these empirical relationships and developed machine‐learning models trained on our kinetic dataset to predict the catalytic parameters of these naturally occurring ADKs. RESULTS: Our high-throughput kinetic data revealed that ADK orthologs vary in kcat by up to three orders of magnitude—even with conserved active sites and similar predicted structures. We address long-standing evolutionary hypotheses demonstrating that thermophilic enzymes are not universally slower than their mesophilic counterparts. We dissect the topology and navigability of this sequence-catalysis landscape, showing that it is rugged, with at least three global neighborhoods organized by distinct domain architectures. Stepwise point mutations and domain swaps show that this landscape remains navigable over long evolutionary timescales through path-dependent mechanisms. Finally, we show that an unsupervised protein language model organizes ADK sequence space by structure but not catalytic activity. kcat prediction models trained on the dataset collected herein outperform prior models trained only on public databases. CONCLUSION: Combining high-throughput kinetic measurements of natural enzyme sequences with machine learning, we charted an evolutionary-scale sequence-catalysis landscape for a ubiquitous enzyme family. Our data show that high catalytic activity can arise through multiple structural solutions and is not strictly limited by thermal adaptation, challenging assumptions about universal activity–stability trade-offs. Whereas protein language models excel at capturing broad structural features, our results emphasize that experimental annotations—especially for catalytic activity—are critical for building accurate sequence-to-function models. More broadly, coupling high‐throughput enzymology with machine and deep learning may uncover insights into enzyme evolution, guide the rational design of novel enzymes, and illuminate the fundamental constraints that shape protein function across the tree of life. Leveraging evolutionary-scale enzymology to map sequence-catalysis landscapes.: Sampling enzyme sequences across phylogeny and diverse environments reveals multiple evolutionary solutions to high catalytic activity arising within different domain architectures. Although protein language models capture sequence-encoded structural organization, the catalytic landscape remains highly rugged. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 00368075 |
| DOI: | 10.1126/science.adu1058 |