タンパク質の進化を制限する要因を解明

ad

2026-03-31 沖縄科学技術大学院大学

沖縄科学技術大学院大学(OIST)などの国際研究チームは、大規模シミュレーションによりタンパク質進化を制限する主要因を解明した。理論上膨大な配列が存在し得る中で、実際に機能するタンパク質はごく一部に限られる。本研究では、既知タンパク質の配列空間を数理的に解析し、進化モデルと比較した結果、自然選択やエピスタシスよりも「進化の出発点(祖先タンパク質)」が多様性を強く制約することを示した。また初期進化ではDNA断片の組換えが重要な役割を果たした可能性も示唆された。これにより、生命起源の理解やAIによるタンパク質設計の限界と今後の拡張の必要性が明らかになった。

タンパク質の進化を制限する要因を解明
さまざまなタンパク質配列空間を抽象的に示した、縮尺が正確ではない視覚的表現。大きな箱は、すべてのアミノ酸配列の可能性(鎖の長さをLとした場合、約20L通りの組み合わせ)を表している。より小さな青い領域は、機能的なタンパク質を形成する配列を表しており、さらに小さな緑色の領域は、これまでに存在が確認されている配列を示している。金色の線はタンパク質の進化経路を表し、赤色の線は絶滅した進化経路(すなわち、現代の生物多様性には存在しないと考えられるもの)を表している。© イサコバほか

<関連情報>

共通祖先からの派生は、タンパク質配列空間の探索を制限する Descent from a common ancestor restricts exploration of protein sequence space

Lada H. Isakova, Elizaveta Streltsova, Olga O. Bochkareva, +1 , and Fyodor A. Kondrashov
Proceedings of the National Academy of Sciences  Published:March 31, 2026
DOI:https://doi.org/10.1073/pnas.2532018123

Significance

Are natural protein sequences representative of all possible sequences that are functional? The sequence space is immense but proteins have been evolving for billions of years, so much of the possible functional space may have already been explored. We find that because sequence evolution of homologous proteins starts from a single common ancestor, protein sequence diversity has been limited to an extreme degree and even 4 billion years of evolution was insufficient to explore the functional sequence space. Protein engineering models that learn only from natural protein sequences may be limited in their ability to predict sequences outside the explored sequence space and empirical approaches that explore the unnatural sequence space may be necessary to fully realize their potential.

Abstract

How functional protein sequences are distributed in sequence space is fundamentally important for evolutionary theory and protein design, particularly if a large diversity of protein functions are hidden in evolutionarily unexplored areas of the sequence space. However, this question is understudied in part because experimental and computational studies use extant sequences as a starting point to study sequence space. Here, we study whether extant sequences are representative of the entire functional sequence space. Across thousands of protein families from vertebrates and bacteria we calculate the dimensionality and the volume of sequence space occupied by extant homologs. We find that the observed dimensionality and volume of extant sequence space are minuscule, many orders of magnitude smaller than what we estimated using a model of protein evolution. Simulating sequence evolution we then quantify the impact of phylogeny, selection, and epistasis on restricting the evolutionary exploration of sequence space. We find that sequence evolution from a single common ancestor, or a single point of origin in sequence space, is by far the largest limiting factor that reduces the dimensionality and volume of extant sequence space. These results indicate that there are vast areas of functional sequence space that have not been explored in evolution because of the excessive restrictions on natural exploration of the protein sequence space imposed by the point of origin effect. We suggest that protein design methods that rely on extant sequences may be limited in their ability to discover truly novel functions.

細胞遺伝子工学
ad
ad
Follow
ad
タイトルとURLをコピーしました