دورية أكاديمية

AliSim-HPC: parallel sequence simulator for phylogenetics.

التفاصيل البيبلوغرافية
العنوان: AliSim-HPC: parallel sequence simulator for phylogenetics.
المؤلفون: Ly-Trong N; School of Computing, College of Engineering, Computing and Cybernetics, Australian National University, Canberra, ACT 2600, Australia., Barca GMJ; School of Computing, College of Engineering, Computing and Cybernetics, Australian National University, Canberra, ACT 2600, Australia., Minh BQ; School of Computing, College of Engineering, Computing and Cybernetics, Australian National University, Canberra, ACT 2600, Australia.
المصدر: Bioinformatics (Oxford, England) [Bioinformatics] 2023 Sep 02; Vol. 39 (9).
نوع المنشور: Journal Article; Research Support, Non-U.S. Gov't
اللغة: English
بيانات الدورية: Publisher: Oxford University Press Country of Publication: England NLM ID: 9808944 Publication Model: Print Cited Medium: Internet ISSN: 1367-4811 (Electronic) Linking ISSN: 13674803 NLM ISO Abbreviation: Bioinformatics Subsets: MEDLINE
أسماء مطبوعة: Original Publication: Oxford : Oxford University Press, c1998-
مواضيع طبية MeSH: Software* , Computing Methodologies*, Phylogeny ; Computer Simulation ; Sequence Alignment
مستخلص: Motivation: Sequence simulation plays a vital role in phylogenetics with many applications, such as evaluating phylogenetic methods, testing hypotheses, and generating training data for machine-learning applications. We recently introduced a new simulator for multiple sequence alignments called AliSim, which outperformed existing tools. However, with the increasing demands of simulating large data sets, AliSim is still slow due to its sequential implementation; for example, to simulate millions of sequence alignments, AliSim took several days or weeks. Parallelization has been used for many phylogenetic inference methods but not yet for sequence simulation.
Results: This paper introduces AliSim-HPC, which, for the first time, employs high-performance computing for phylogenetic simulations. AliSim-HPC parallelizes the simulation process at both multi-core and multi-CPU levels using the OpenMP and message passing interface (MPI) libraries, respectively. AliSim-HPC is highly efficient and scalable, which reduces the runtime to simulate 100 large gap-free alignments (30 000 sequences of one million sites) from over one day to 11 min using 256 CPU cores from a cluster with six computing nodes, a 153-fold speedup. While the OpenMP version can only simulate gap-free alignments, the MPI version supports insertion-deletion models like the sequential AliSim.
Availability and Implementation: AliSim-HPC is open-source and available as part of the new IQ-TREE version v2.2.3 at https://github.com/iqtree/iqtree2/releases with a user manual at http://www.iqtree.org/doc/AliSim.
(© The Author(s) 2023. Published by Oxford University Press.)
References: Mol Biol Evol. 1995 Jul;12(4):546-57. (PMID: 7659011)
Mol Biol Evol. 2020 Dec 16;37(12):3632-3641. (PMID: 32637998)
Mol Biol Evol. 2020 May 1;37(5):1530-1534. (PMID: 32011700)
Mol Biol Evol. 2015 Jan;32(1):268-74. (PMID: 25371430)
Mol Biol Evol. 1994 Mar;11(2):261-77. (PMID: 8170367)
Bioinformatics. 2015 Aug 1;31(15):2577-9. (PMID: 25819675)
PLoS Comput Biol. 2024 Aug 5;20(8):e1012337. (PMID: 39102450)
Syst Biol. 2020 Mar 1;69(2):221-233. (PMID: 31504938)
Mol Biol Evol. 2009 Aug;26(8):1879-88. (PMID: 19423664)
Bioinformatics. 2023 Sep 2;39(9):. (PMID: 37669126)
Front Microbiol. 2022 Sep 07;13:787856. (PMID: 36160199)
J Mol Evol. 1994 Mar;38(3):305-9. (PMID: 8006998)
J Mol Evol. 1993 Feb;36(2):182-98. (PMID: 7679448)
PLoS Comput Biol. 2014 Apr 10;10(4):e1003537. (PMID: 24722319)
Bioinformatics. 2019 Nov 1;35(21):4453-4455. (PMID: 31070718)
Bioinformatics. 2004 Feb 12;20(3):407-15. (PMID: 14960467)
Bioinformatics. 2019 May 15;35(10):1771-1773. (PMID: 30321303)
Mol Biol Evol. 2020 Nov 1;37(11):3338-3352. (PMID: 32585030)
Mol Biol Evol. 2022 May 3;39(5):. (PMID: 35511713)
Mol Biol Evol. 1994 May;11(3):459-68. (PMID: 8015439)
PLoS Comput Biol. 2022 Apr 29;18(4):e1010056. (PMID: 35486906)
Comput Appl Biosci. 1997 Jun;13(3):235-8. (PMID: 9183526)
Bioinformatics. 2005 Nov 1;21 Suppl 3:iii31-8. (PMID: 16306390)
J Mol Evol. 1999 Nov;49(5):691-8. (PMID: 10552050)
تواريخ الأحداث: Date Created: 20230901 Date Completed: 20230929 Latest Revision: 20240926
رمز التحديث: 20240926
مُعرف محوري في PubMed: PMC10534053
DOI: 10.1093/bioinformatics/btad540
PMID: 37656933
قاعدة البيانات: MEDLINE
الوصف
تدمد:1367-4811
DOI:10.1093/bioinformatics/btad540