How far can you go? Extrapolating values of catalytic activity from known protein landscapes in natural and directed evolution.

Saved in:
Bibliographic Details
Title: How far can you go? Extrapolating values of catalytic activity from known protein landscapes in natural and directed evolution.
Authors: Kell, Douglas B.1,2,3 (AUTHOR) dbk@liv.ac.uk, Roberts, Ivayla1 (AUTHOR)
Source: Chemical Society Reviews. 6/8/2026, Vol. 55 Issue 11, p6417-6452. 36p.
Subjects: Catalytic activity, Epistasis (Genetics), Protein engineering, Machine learning, Biocatalysis, Statistical models
Abstract: The number of possible variants representing the landscape of a protein sequence of length N residues, made of the standard unmodified proteinogenic amino acids, is 20N; its exhaustive experimental analysis is consequently intractable. Our focus is on the real and perceived shapes of different fitness landscapes. Epistasis refers to a phenomenon by which the 'best' amino acid at a given residue depends on the nature of the amino acid at one or more other residues. Because of epistasis, real protein landscapes display peaks representing local maxima in which weak mutation/strong-selection regimes can cause evolution to become trapped, leading to landscapes that are rugged. Fortunately, although they are necessarily somewhat rugged, such protein landscapes possess regularities that admit their modelling from more limited experimental data, using the methods of statistics and machine learning. We provide a variety of arguments that for typical proteins of length 300–500 residues some 105 or 106 examples, and in favourable cases even fewer, are likely sufficient to allow a reasonable initial modelling (and accurate predictive exploration) of the entire 20N landscape for properties such as kcat. The distribution of fitness effects (DFE) around an existing wild type is usually reasonably fitted statistically by a gamma distribution. However, we also survey modern ideas, especially extreme value theory, that allow extrapolation from the known, with a focus on methods – especially the Generalised Pareto Distribution – that provide means for generating the statistical likelihood of obtaining activities or fitnesses far greater than those observed in existing populations as measured with what are small numbers. These likelihoods typically decrease exponentially, as do the decreases in errors as a function of the size of the network and of the training data as found by deep neural network models as 'universal approximators'. This is entirely consistent with the large differences between the minuscule amount of available sequence-activity data, that are necessarily local in character, reflecting evolutionary contingency, and the overall distribution (20N, where N might usefully be decreased) that would be expected to contain examples that have much better properties than any observed thus far. This consequently requires careful choices of examples drawn from an extensive distribution (using active learning) for predictive modelling. For instance, a widespread view of a trade-off between catalytic activity and thermostability seems to follow directly from inadequate sampling. All of this has significant implications for the understanding, modelling, and optimisation of experiments in directed evolution and the biocatalysts they produce. [ABSTRACT FROM AUTHOR]
Copyright of Chemical Society Reviews is the property of Royal Society of Chemistry and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 194364688
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: How far can you go? Extrapolating values of catalytic activity from known protein landscapes in natural and directed evolution.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kell%2C+Douglas+B%2E%22">Kell, Douglas B.</searchLink><relatesTo>1,2,3</relatesTo> (AUTHOR)<i> dbk@liv.ac.uk</i><br /><searchLink fieldCode="AR" term="%22Roberts%2C+Ivayla%22">Roberts, Ivayla</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Chemical+Society+Reviews%22">Chemical Society Reviews</searchLink>. 6/8/2026, Vol. 55 Issue 11, p6417-6452. 36p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Catalytic+activity%22">Catalytic activity</searchLink><br /><searchLink fieldCode="DE" term="%22Epistasis+%28Genetics%29%22">Epistasis (Genetics)</searchLink><br /><searchLink fieldCode="DE" term="%22Protein+engineering%22">Protein engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Biocatalysis%22">Biocatalysis</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+models%22">Statistical models</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The number of possible variants representing the landscape of a protein sequence of length N residues, made of the standard unmodified proteinogenic amino acids, is 20N; its exhaustive experimental analysis is consequently intractable. Our focus is on the real and perceived shapes of different fitness landscapes. Epistasis refers to a phenomenon by which the 'best' amino acid at a given residue depends on the nature of the amino acid at one or more other residues. Because of epistasis, real protein landscapes display peaks representing local maxima in which weak mutation/strong-selection regimes can cause evolution to become trapped, leading to landscapes that are rugged. Fortunately, although they are necessarily somewhat rugged, such protein landscapes possess regularities that admit their modelling from more limited experimental data, using the methods of statistics and machine learning. We provide a variety of arguments that for typical proteins of length 300–500 residues some 105 or 106 examples, and in favourable cases even fewer, are likely sufficient to allow a reasonable initial modelling (and accurate predictive exploration) of the entire 20N landscape for properties such as kcat. The distribution of fitness effects (DFE) around an existing wild type is usually reasonably fitted statistically by a gamma distribution. However, we also survey modern ideas, especially extreme value theory, that allow extrapolation from the known, with a focus on methods – especially the Generalised Pareto Distribution – that provide means for generating the statistical likelihood of obtaining activities or fitnesses far greater than those observed in existing populations as measured with what are small numbers. These likelihoods typically decrease exponentially, as do the decreases in errors as a function of the size of the network and of the training data as found by deep neural network models as 'universal approximators'. This is entirely consistent with the large differences between the minuscule amount of available sequence-activity data, that are necessarily local in character, reflecting evolutionary contingency, and the overall distribution (20N, where N might usefully be decreased) that would be expected to contain examples that have much better properties than any observed thus far. This consequently requires careful choices of examples drawn from an extensive distribution (using active learning) for predictive modelling. For instance, a widespread view of a trade-off between catalytic activity and thermostability seems to follow directly from inadequate sampling. All of this has significant implications for the understanding, modelling, and optimisation of experiments in directed evolution and the biocatalysts they produce. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Chemical Society Reviews is the property of Royal Society of Chemistry and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=194364688
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1039/d5cs01387a
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 36
        StartPage: 6417
    Subjects:
      – SubjectFull: Catalytic activity
        Type: general
      – SubjectFull: Epistasis (Genetics)
        Type: general
      – SubjectFull: Protein engineering
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Biocatalysis
        Type: general
      – SubjectFull: Statistical models
        Type: general
    Titles:
      – TitleFull: How far can you go? Extrapolating values of catalytic activity from known protein landscapes in natural and directed evolution.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kell, Douglas B.
      – PersonEntity:
          Name:
            NameFull: Roberts, Ivayla
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 08
              M: 06
              Text: 6/8/2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 03060012
          Numbering:
            – Type: volume
              Value: 55
            – Type: issue
              Value: 11
          Titles:
            – TitleFull: Chemical Society Reviews
              Type: main
ResultId 1