科孊者たちが機械孊習を甚いお䜎分子化合物の前人未到の知芋を獲埗(Scientists use machine learning to gain unprecedented view of small molecules)

ad
ad

䜎分子を同定する新しいツヌルは、蚺断、創薬、基瀎研究などに圹立぀。 A new tool to identify small molecules offers benefits for diagnostics, drug discovery and fundamental research.

2022-12-20 フィンランド・アヌルト倧孊

 新しい機械孊習モデルは、科孊者が䜎分子を識別するのに圹立ち、医孊、創薬、環境化孊に応甚されたす。アヌルト倧孊ずルクセンブルク倧孊の研究者が開発したこのモデルは、数十の研究宀のデヌタを甚いお蚓緎され、䜎分子を同定するための最も正確なツヌルの䞀぀ずなりたした。

代謝物ず呌ばれる䜕千皮類もの䜎分子は、゚ネルギヌを茞送し、现胞情報を人䜓党䜓に䌝達しおいたす。代謝物は非垞に小さいため、血液サンプル分析では互いに区別するこずが困難です。しかし、これらの分子を特定するこずは、運動、栄逊、アルコヌルの䜿甚、代謝異垞が健康にどのように圱響するかを理解する䞊で重芁です。

代謝物の同定は、通垞、液䜓クロマトグラフィヌ質量分析法ず呌ばれる分離技術で質量ず保持時間を分析するこずによっお行われたす。この技術では、たずサンプルをカラムに通すこずで代謝物を分離し、その結果、枬定装眮での流速(保持時間)が異なりたす。次に質量分析蚈を甚いお、質量に応じお代謝物を分類し、同定䜜業を埮調敎したす。たた、タンデム質量分析法ず呌ばれる技術により、代謝物を现かく分解しお成分を分析するこずもできる。

このたび、Rousu教授の研究グルヌプは、䜎分子を同定するための新しい機械孊習モデルを開発した。これは最近『Nature Machine Intelligence』誌に掲茉された。

この新しいオヌプン゜ヌスのモデルは、研究コミュニティ党䜓に、䜎分子に぀いおの豊かな芋方を提䟛したす。糖尿病などの代謝異垞や癌を特定する方法の研究にも圹立぀でしょう」ず、Rousuは蚀う。

この新しいアプロヌチは、埓来の方法が盎面しおいた問題の1぀を゚レガントに回避しおいる。分子の保持時間は研究宀によっお異なるため、研究宀間でデヌタを比范するこずができないのだ。アヌルト倧孊の博士課皋に圚籍するEric Bachは、博士課皋での研究䞭に、この問題を解決する代替策を考え出した。

私たちの研究から、絶察的な保持時間は倉化しおも、保持順序は異なるラボによる枬定でも安定しおいるこずがわかりたした」ずBach氏は説明する。このため、代謝物に関する䞀般に公開されおいるすべおのデヌタを史䞊初めお統合し、機械孊習モデルに送り蟌むこずができたのです』。

䞖界䞭の数十の研究宀からのデヌタを取り蟌むこずで、機械孊習モデルは、立䜓化孊的倉異䜓ずしお知られる鏡像分子を識別するのに十分な粟床を持぀ようになったのです」。これたで、識別ツヌルは立䜓化孊的倉異䜓を芋分けるこずができなかったので、この新しい胜力は、創薬などの分野で新しい道を開くず期埅されおいたす。

<関連情報>

液䜓クロマトグラフィヌの保持順ずタンデム質量分析デヌタを甚いた䜎分子化合物の共同構造アノテヌション Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data

Eric Bach,Emma L. Schymanski &Juho Rousu

Nature Machine Intelligence  Published:19 December 2022

DOI:https://doi.org/10.1038/s42256-022-00577-2

科孊者たちが機械孊習を甚いお䜎分子化合物の前人未到の知芋を獲埗(Scientists use machine learning to gain unprecedented view of small molecules)

Abstract

Structural annotation of small molecules in biological samples remains a key bottleneck in untargeted metabolomics, despite rapid progress in predictive methods and tools during the past decade. Liquid chromatography–tandem mass spectrometry, one of the most widely used analysis platforms, can detect thousands of molecules in a sample, the vast majority of which remain unidentified even with best-of-class methods. Here we present LC-MS2Struct, a machine learning framework for structural annotation of small-molecule data arising from liquid chromatography–tandem mass spectrometry (LC-MS2) measurements. LC-MS2Struct jointly predicts the annotations for a set of mass spectrometry features in a sample, using a novel structured prediction model trained to optimally combine the output of state-of-the-art MS2 scorers and observed retention orders. We evaluate our method on a dataset covering all publicly available reversed-phase LC-MS2 data in the MassBank reference database, including 4,327 molecules measured using 18 different LC conditions from 16 contributors, greatly expanding the chemical analytical space covered in previous multi-MS2 scorer evaluations. LC-MS2Struct obtains significantly higher annotation accuracy than earlier methods and improves the annotation accuracy of state-of-the-art MS2 scorers by up to 106%. The use of stereochemistry-aware molecular fingerprints improves prediction performance, which highlights limitations in existing approaches and has strong implications for future computational LC-MS2 developments.

有機化孊・薬孊
ad
ad
Follow
ad
タむトルずURLをコピヌしたした