Machine learning enables scalable and systematic hierarchical virus taxonomy

Bolduc, Benjamin, Zablocki, Olivier, Turner, Dann, Jang, Ho Bin, Guo, Jiarong, Adriaenssens, Evelien M. ORCID: https://orcid.org/0000-0003-4826-5406, Dutilh, Bas E. and Sullivan, Matthew B. (2025) Machine learning enables scalable and systematic hierarchical virus taxonomy. Nature Biotechnology. ISSN 1087-0156

Full text not available from this repository. (Request a copy)

Abstract

Although virus ecogenomics has expanded access to and understanding of the virosphere, existing classification tools lack taxonomic resolution and are unable to scale to modern discovery-based datasets or classify previously unknown sequence space. Here we develop vConTACT3—a machine learning-based tool that improves scalability and accuracy of virus taxonomy. By optimizing gene-sharing thresholds and leveraging adaptive, realm-specific cut-offs, vConTACT3 expands classification to both eukaryote and prokaryote viruses for four of the six officially recognized realms, and establishes accurate hierarchical taxonomy from genus to order. Specifically, vConTACT3 achieves >95% agreement with official taxonomy for 35,545 and 13,524 public prokaryotic and eukaryotic virus genomes, respectively, to surpass vConTACT2 across most realms, while still uniquely classifying previously uncharacterized taxa, and doing so even faster. vConTACT3 application provides taxonomy assignments for tens of thousands of unclassified taxa rapidly, automatically and systematically; evaluates virus sequence space to reveal support for fewer taxonomic ranks than currently available and identifies taxonomically challenging areas across the virosphere.

Item Type: Article
Additional Information: Data availability: Data used for benchmarks (parameter optimizations), construction of databases and fine-tuning the pipeline are available from NCBI Virus RefSeq (v.218). Data used to test scalability, assess fragmentation and evaluate labeling stability are available from IMG/VR v.4.1 (December 2022 release). Databases used by vConTACT3 (as well as source files) are available via Zenodo at https://doi.org/10.5281/zenodo.10035619 and https://doi.org/10.5281/zenodo.10935513 (refs. 67,68). Code availability: vConTACT3 is available via Bitbucket at https://bitbucket.org/MAVERICLab/vcontact3 (ref. 69) as an installable Python package, as well as through Python package managers Anaconda (https://anaconda.org/) and Mamba (https://mamba.readthedocs.io/). Instructions for building an Apptainer container of vConTACT3 is available on Bitbucket, along with a definitions file. A comprehensive documentation site is available through https://vcontact3.readthedocs.io. Optimization benchmarks were performed using v.3.0.0b36 (‘beta’ v.36), fragmentation analyses using v.3.0.0b63 and label stability, v3.1.4. All other results (including tool comparisons) should be assumed as v.3.0.0.b36. Through all analyses, Python 3.10 was used. Data processing and analyses were conducted using numpy v.1.23.5, pandas v.2.1.1 and scipy v.1.10.1. Taxonomic parsing was handled by the ETE3 Toolkit, v.3.1.3. Statistical analyses were performed using scikit-learn v.1.2.2 and scikit-bio v.0.5.8. Data visualizations were created using matplotlib v.3.7.1, seaborn v.0.12.1 and UpSetPlot v.0.7.0. Networks were rendered through a combination of Cytoscape v.3.10.1, networkx v.3.1 and Python-igraph v.0.10.4. Gene predictions were done through pyprodigal v.2.3.0 and pyprodigal-gv v.0.3.1 and all sequence processing through biopython v.1.81. Protein clustering was done using MMSeqs2 v.14-7e284.
Faculty \ School: Faculty of Medicine and Health Sciences > Norwich Medical School
Related URLs:
Depositing User: LivePure Connector
Date Deposited: 16 Jul 2026 14:29
Last Modified: 16 Jul 2026 14:29
URI: https://ueaeprints.uea.ac.uk/id/eprint/103909
DOI: 10.1038/s41587-025-02946-9

Actions (login required)

View Item View Item