When using PhyloConstructor, please cite the PhyloConstructor GitHub repository together with the publications associated with the software tools used in your analysis.
The exact tools executed depend on the selected input sources and workflow parameters.
GitHub repository: https://github.com/dorinemerlat/phyloconstructor
Di Tommaso P, Chatzou M, Floden EW, Barja PP, Palumbo E, Notredame C. Nextflow enables reproducible computational workflows. Nat Biotechnol. 2017;35(4):316–319. doi: 10.1038/nbt.3820. PubMed PMID: 28398311.
The workflow overview figure included in this repository was initially generated using nf-metro and subsequently refined manually.
Project: https://github.com/nextflow-io/nf-metro
PhyloConstructor can retrieve genomic, proteomic and transcriptomic data from UniProt and NCBI resources. Please cite the databases from which data were obtained.
The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 2023;51(D1):D523–D531. doi: 10.1093/nar/gkac1052. PubMed PMID: 36408920.
O'Leary NA, Cox E, Holmes JB, et al. Exploring and retrieving sequence and metadata for species across the tree of life with NCBI Datasets. Sci Data. 2024;11:732. doi:10.1038/s41597-024-03571-y.
Leinonen R, Sugawara H, Shumway M; International Nucleotide Sequence Database Collaboration. The Sequence Read Archive. Nucleic Acids Res. 2011;39(Database issue):D19–D21. doi: 10.1093/nar/gkq1019. PubMed PMID: 21062823; PubMed Central PMCID: PMC3013647.
National Center for Biotechnology Information. Transcriptome Shotgun Assembly Sequence Database. Bethesda (MD): National Library of Medicine (US), National Center for Biotechnology Information. Available from: https://www.ncbi.nlm.nih.gov/genbank/tsa/.
Manni M, Berkeley MR, Seppey M, Simão FA, Zdobnov EM. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 2021;38(10):4647–4654. doi: 10.1093/molbev/msab199. PubMed PMID: 34320186; PubMed Central PMCID: PMC8476166.
When reporting BUSCO results, the lineage dataset and its version should also be stated.
Example:
BUSCO v6 was run using the
arthropoda_odb10lineage dataset.
These tools are used when PhyloConstructor constructs protein datasets from TSA transcriptomes or SRA RNA-seq reads.
National Center for Biotechnology Information. SRA Toolkit. Available from: https://github.com/ncbi/sra-tools.
Bushnell B. BBMap: a fast, accurate, splice-aware aligner. Lawrence Berkeley National Laboratory. Available from: https://sourceforge.net/projects/bbmap/.
Bushmanova E, Antipov D, Lapidus A, Prjibelski AD. rnaSPAdes: a de novo transcriptome assembler and its application to RNA-Seq data. GigaScience. 2019;8(9):giz100. doi: 10.1093/gigascience/giz100. PubMed PMID: 31510679; PubMed Central PMCID: PMC6751136.
The original SPAdes publication may also be cited:
Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD, Pyshkin AV, Sirotkin AV, Vyahhi N, Tesler G, Alekseyev MA, Pevzner PA. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 2012;19(5):455–477. doi: 10.1089/cmb.2012.0021. PubMed PMID: 22506599; PubMed Central PMCID: PMC3342519.
Haas BJ, Papanicolaou A. TransDecoder: find coding regions within transcripts. Available from: https://github.com/TransDecoder/TransDecoder.
Kuraku S, Zmasek CM, Nishimura O, Katoh K. aLeaves facilitates on-demand exploration of metazoan gene family trees on MAFFT sequence alignment server with enhanced interactivity. Nucleic Acids Research. 2013;41(W1):W22-W28. doi:10.1093/nar/gkt389.
Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009;25(15):1972–1973. doi: 10.1093/bioinformatics/btp348. PubMed PMID: 19505945; PubMed Central PMCID: PMC2712344.
Steenwyk JL, Buida TJ III, Li Y, Shen XX, Rokas A. PhyKIT: a broadly applicable UNIX shell toolkit for processing and analyzing phylogenomic data. Bioinformatics. 2021;37(16):2325–2331. doi: 10.1093/bioinformatics/btab096. PubMed PMID: 33892497; PubMed Central PMCID: PMC8370967.
Shen W, Le S, Li Y, Hu F. SeqKit: a cross-platform and ultrafast toolkit for FASTA/Q file manipulation. PLoS One. 2016;11(10):e0163962. doi: 10.1371/journal.pone.0163962. PubMed PMID: 27706213; PubMed Central PMCID: PMC5051824.
Li W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 2006;22(13):1658–1659. doi: 10.1093/bioinformatics/btl158. PubMed PMID: 16731699.
Dainat J, Hereñú D, LucileSol, Pascal-Git. NBISweden/AGAT: AGAT. Zenodo. doi: 10.5281/zenodo.3552717.
Minh BQ, Schmidt HA, Schrempf D, et al. IQ-TREE 3: phylogenomic inference software using complex evolutionary models. Molecular Biology and Evolution. 2025.
Kalyaanamoorthy S, Minh BQ, Wong TKF, von Haeseler A, Jermiin LS. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods. 2017;14(6):587–589. doi: 10.1038/nmeth.4285. PubMed PMID: 28481363; PubMed Central PMCID: PMC5453245.
Hoang DT, Chernomor O, von Haeseler A, Minh BQ, Vinh LS. UFBoot2: improving the ultrafast bootstrap approximation. Mol Biol Evol. 2018;35(2):518–522. doi: 10.1093/molbev/msx281. PubMed PMID: 29077904; PubMed Central PMCID: PMC5850222.
Mirarab S, Reaz R, Bayzid MS, Zimmermann T, Swenson MS, Warnow T. ASTRAL: genome-scale coalescent-based species tree estimation. Bioinformatics. 2014;30(17):i541–i548. doi: 10.1093/bioinformatics/btu462. PubMed PMID: 25161245; PubMed Central PMCID: PMC4147915.
Zhang C, Rabiee M, Sayyari E, Mirarab S. ASTRAL-III: polynomial time species tree reconstruction from partially resolved gene trees. BMC Bioinformatics. 2018;19(Suppl 6):153. doi: 10.1186/s12859-018-2129-y. PubMed PMID: 29745866; PubMed Central PMCID: PMC6001769.
Grüning B, Dale R, Sjödin A, Chapman BA, Rowe J, Tomkins-Tinch CH, Valieris R, Köster J; Bioconda Team. Bioconda: sustainable and comprehensive software distribution for the life sciences. Nat Methods. 2018;15(7):475–476. doi: 10.1038/s41592-018-0046-7. PubMed PMID: 29967506.
da Veiga Leprevost F, Grüning B, Aflitos SA, Röst HL, Uszkoreit J, Barsnes H, Vaudel M, Moreno P, Gatto L, Weber J, Bai M, Jimenez RC, Sachsenberg T, Pfeuffer J, Alvarez RV, Griss J, Nesvizhskii AI, Perez-Riverol Y. BioContainers: an open-source and community-driven framework for software standardization. Bioinformatics. 2017;33(16):2580–2582. doi: 10.1093/bioinformatics/btx192. PubMed PMID: 28379341; PubMed Central PMCID: PMC5870671.
Merkel D. Docker: lightweight Linux containers for consistent development and deployment. Linux Journal. 2014;2014(239):2. doi: 10.5555/2600239.2600241.
Docker images are used to build or distribute some software environments, although PhyloConstructor executions currently use the Singularity profile.
Kurtzer GM, Sochat V, Bauer MW. Singularity: scientific containers for mobility of compute. PLoS One. 2017;12(5):e0177459. doi: 10.1371/journal.pone.0177459. PubMed PMID: 28494014; PubMed Central PMCID: PMC5426675.
For a standard analysis based on downloaded proteomes and concatenation-based phylogenetic inference, cite at least:
- PhyloConstructor
- Nextflow
- the source databases used
- BUSCO
- MAFFT
- trimAl
- PhyKIT
- IQ-TREE
- ModelFinder
- UFBoot2
- ASTRAL or ASTRAL-III
- Singularity
- Bioconda and BioContainers
For an analysis including SRA-derived transcriptomes, additionally cite:
- Sequence Read Archive
- SRA Toolkit
- BBMap
- RNA-SPAdes
- TransDecoder