Generated October 6, 2026

Introduction

This tutorial demonstrates how to discover evidence of viral sequences within metagenomic data. After processing and assembling a metagenome, you can use vConTACT2 and VirSorter2 to identify and classify viral sequences, VirMatcher to predict virus-host relationships, and DRAMv to annotate viral genomes. The apps from the iVirus suite of tools were developed by the Sullivan lab at The Ohio State University and integrated into KBase with the support of the Microbes Persist: Systems Biology of the Soil Microbiome Science Focus Area (SFA). Example data in this tutorial comes from the from the Global Ocean Virome dataset. This tutorial Narrative builds off the Viral Analysis End-to-End pipeline Narrative built by the iVirus team in 2020 and adds new apps relevant to viral genomics researchers.

Tutorial Objectives

  • Assemble viral metagenome using raw reads
  • Identify viral sequences from assembly data
  • Classify predicted viral sequences

Table of contents:

  1. Scientific Background
  2. Importing read data from the SRA
  3. Assessing read quality with FastQC
  4. Quality control of reads with Trimmomatic
  5. Post-QC Read Quality Assesssment
  6. Assembling Metagenome with metaSPAdes
  7. Classify Viral Taxonomy with Kaiju
  8. Identifying viral genomes using VirSorter2
  9. Annotate Viral assemblies with DRAM-v
  10. Classifying Viral genomes using vConTACT2
  11. Post-vConTACT2 analysis
  12. vConTACT2 Quick Reference Guide
  13. Predict host-virus relationships with VirMatcher
  14. Conclusions
  15. Version History
  16. Feedback and Help

Scientific Background - why study viruses?!

Viruses are important for more than a few reasons!

The "most abundant biological entities on the planet"1,2, with 1031 virus-like particles3

3% (0-18%) of any microbial genome is really virus4,5

On the topic of composition, 8% of the human genome is viral6,7

Move 1029 genes per day, globally 8,9

Lyse between 20-40% of ocean microbes daily10

They steal metabolic genes and can encode key metabolic components (like photosynthesis!)11,12,13

Can infect other viruses (virophages)14

Viruses - not microbes - encode many of the toxins we think of as bacterial (bordtella, cholera, shiga, etc) 15

Fewer than 1% are culturable16,17

Viruses are often hidden in datasets (both viral and microbial) this guide will help you find them!

Good luck!

Importing read data from the SRA

The reads in this dataset were generated from the Global Ocean Virome. and deposited as ERR594369. This is also known as Tara Oceans Expedition Station 36 surface ("SRF"), a coastal area in the Indian Ocean (more specifically, Northwest Arabian Sea), taken from ~5 m depth.

Image of Tara Oceans Expedition Station 36
Figure of Station 36 location. Figure heavily modified from Roux et al (2016) Nature

To upload your own reads into KBase, you can use them drag and drop interface in the Data Panel. A full guide on data upload and download can be found at http://kbase.us/data-upload-download-guide/. For this dataset, we'll import our reads directly from SRA into KBase. Alternatively, you could go to the link (more below), download the file(s) from SRA directly to your computer, and then upload them into KBase: https://trace.ncbi.nlm.nih.gov/Traces/sra/?run=ERR594369

Import an SRA file from a web URL into your Narrative as a Reads data object.
This app completed without errors in 28m 54s.
Objects
Created Object Name Type Description
ERR594369_GlobalOceanViromes PairedEndLibrary Imported Reads
Links

Assessing read quality

Following import, always check the quality of the data going into your analysis. It's always a good idea to know what quality is [eventually] going into an assembly. To quote a populat CS phrase, "Garbage In, Garbage Out." Essentially, this means that if you put poor quality data into your analysis, you're going to get poor quality results.

A quality control application for high throughput sequence data.
This app completed without errors in 11m 48s.
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • ERR594369_GlobalOceanViromes_186849_73_1.rev_fastqc.zip - Zip file generated by fastqc that contains original images seen in the report
  • ERR594369_GlobalOceanViromes_186849_73_1.fwd_fastqc.zip - Zip file generated by fastqc that contains original images seen in the report

Quality control of reads

The FastQC of these enriched reads is already pretty good. You could get away with proceeding directly to assembly, but we'll trim reads to remove low-quality regions at the read ends and any residual adapters.

We'll trim using a popular read trimming tool, Trimmomatic . Defaults are okay, unless your read data has unique adapter or sequencing conditions.

Trim paired- or single-end Illumina reads with Trimmomatic.
This app completed without errors in 41m 16s.
Objects
Created Object Name Type Description
ERR594369_GlobalOceanViromes_trimmed.240719_paired PairedEndLibrary Trimmed Reads
ERR594369_GlobalOceanViromes_trimmed.240719_unpaired_fwd SingleEndLibrary Trimmed Unpaired Forward Reads
ERR594369_GlobalOceanViromes_trimmed.240719_unpaired_rev SingleEndLibrary Trimmed Unpaired Reverse Reads

Post-QC read assessment

Following read trimming, we'll check to ensure that 1) all adapters are removed (if they even existed in the 1st place) and 2) that the reads are of sufficient quality for assembly.

A quality control application for high throughput sequence data.
This app completed without errors in 10m 57s.
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • ERR594369_GlobalOceanViromes_trimmed.240719_paired_186849_76_1.fwd_fastqc.zip - Zip file generated by fastqc that contains original images seen in the report
  • ERR594369_GlobalOceanViromes_trimmed.240719_paired_186849_76_1.rev_fastqc.zip - Zip file generated by fastqc that contains original images seen in the report

Assembly with MetaSPAdes

With clean reads, we'll now assemble these reads into contigs using another popular assembler, MetaSPAdes . Of the assemblers currently available in the KBase ecosystem, I'd argue MetaSPAdes performs slightly better than MEGAHIT and better than IDBA-UD. For an excellent review of viral benchmarks regarding assemblers, see this PeerJ article .

Assemble metagenomic reads using the SPAdes assembler.
This app completed without errors in 2h 44m 53s.
Objects
Created Object Name Type Description
ERR594369_GlobalOceanViromes_metaSPAdes.240719 Assembly Assembled contigs
Summary
Assembly saved to: allenbh:narrative_1721147566557/ERR594369_GlobalOceanViromes_metaSPAdes.240719 Assembled into 16237 contigs. Avg Length: 5972.963355299625 bp. Contig Length Distribution (# of contigs -- min to max basepairs): 16018 -- 2000.0 to 40776.5 bp 139 -- 40776.5 to 79553.0 bp 41 -- 79553.0 to 118329.5 bp 21 -- 118329.5 to 157106.0 bp 8 -- 157106.0 to 195882.5 bp 4 -- 195882.5 to 234659.0 bp 2 -- 234659.0 to 273435.5 bp 0 -- 273435.5 to 312212.0 bp 2 -- 312212.0 to 350988.5 bp 2 -- 350988.5 to 389765.0 bp
Links

Classify Viral Taxonomy with Kaiju

Kaiju is a program for sensitive taxonomic classification of high-throughput sequencing reads from metagenomic whole genome sequencing or metatranscriptomics experiments. You can select from several databases to classify taxonomy. For this tutorial, we will run Kaiju using two viral databases

Allows users to perform taxonomic classification of shotgun metagenomic read data with Kaiju.
This app completed without errors in 23m 16s.
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • kaiju_classifications.zip
  • kaiju_summaries.zip
  • krona_data.zip
  • stacked_bar_abundance_plots_PNG+PDF.zip
Allows users to perform taxonomic classification of shotgun metagenomic read data with Kaiju.
This app completed without errors in 36m 44s.
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • kaiju_classifications.zip
  • kaiju_summaries.zip
  • krona_data.zip
  • stacked_bar_abundance_plots_PNG+PDF.zip

Identifying viral genomes using VirSorter2

Now we are ready to identify potential viral sequences sequences in our assembly with VirSorter2.

Overview of the VirSorter2 framework
Overview of the VirSorter2 framework from Guo, J., Bolduc, B., Zayed, A.A. et al. VirSorter2: a multi-classifier, expert-guided approach to detect diverse DNA and RNA viruses. Microbiome 9, 37 (2021)

VirSorter2 does not work as well for small, linear genomes. As a general rule of thumb, we'll filter out any contigs that are less than 5-kb in length OR any non-circular genomes less than 1.5-kb.

Viral groups:

  • dsDNAphage: double-stranded DNA phages aka Caudovirales aka bacteriophages
  • NCLDV: nucleocytoplasmic large DNA viruses
  • RNA: single-stranded RNA viruses
  • ssDNA: single-stranded DNA viruses
  • lavidaviridae: Virophages or "large virus dependent or associated"
Identifies viral sequences from viral and microbial metagenomes
This app completed without errors in 2h 51m 47s.
Objects
Created Object Name Type Description
ERR594369_GlobalOceanViromes_VirSorter2.dsDNAphage.240722 Assembly KBase.Assembly object from VirSorter2
Summary
Results from your VirSorter2 run. Above you'll find a report with the identified,*putative* virus genomes, and below, downloadable links to the results files and links to the KBase assembly object. For users who enabled DRAM-v compatibility, the shock ID is 01e60492-2e1e-4c90-a413-f4b730d3fd32
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • final-viral-boundary.tsv.tar.gz - Table with boundary information
  • final-viral-combined.fa.tar.gz - Viral sequences in FASTA format
  • final-viral-score.tsv.tar.gz - Table with scoring information

Annotate viral assemblies with DRAM-v

DRAM for vMAGs works by annotating viral genomes with a set of databases curated to the task, and integrating additional input from Virsorter. Note that you must start with a metagenomic assembly object and run the VirSorter KBase app. DRAM-v is run using the viral genome files along with the lshock ID from the KBase VirSorter Summary. The user is then given a tab delimited annotations file with all annotations from all databases for all genes in all genomes, with data on known and potential Auxiliary Metabolic Genes (AMGs). AMGs are virus-encoded microbial metabolic genes that allow metabolic reprogramming of the infected host.

Overview of DRAM and pipeline

Annotate vMAGs with DRAM and distill resulting annotations to create an interactive auxiliary metabolic gene summary. Use with the VirSorter KBase app.
This app completed without errors in 1h 45m 57s.
Summary
Here are the results from your DRAM run.
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • annotations.tsv - DRAM annotations in a tab separate table format
  • genes.fna - Genes as nucleotides predicted by DRAM with brief annotations
  • genes.faa - Genes as amino acids predicted by DRAM with brief annotations
  • genes.gff - GFF file of all DRAM annotations
  • trnas.tsv - Tab separated table of tRNAs as detected by tRNAscan-SE
  • genbank.tar.gz - Compressed folder of output genbank files
  • amg_summary.tsv - DRAM-v AMG summary table
  • vMAG_stats.tsv - DRAM-v vMAG statistics table

Classify annotated viral genomes with vConTACT2

Now that we've handled the minor details, we can take those called genes and use them in vConTACT2 .

vConTACT2 works by using a gene-sharing network to associate viral genomes. The more genes that are shared between two genomes, the higher the probability of those two genomes being phylogenetically related.

Image of vContact2's gene-sharing network
vConTACT2 virus classification. Figure heavily modified from Jang and Bolduc et al (2019) Nat. Biotech

So what's happening in the background? vConTACT2 will extract each viral genome and its associated gene predictions, and build the Gene2Genome table that underpins the whole analysis. Thankfully, vConTACT2+KBase generates this file in the background! For non-KBase users, this could be a challenge unless you let vConTACT2 handle everything.

There are a lot of options for vConTACT2. As a developer, there's a balance between giving enough options to allow for granular control of how the tool operates, and not over-burdening the user with options most are unlikely to change. In KBase, all the default options have been selected. There's no need to change anything - except if you want to use the most recent version of NCBI's Viral RefSeq. Often, users prefer to use the "older" version as that's what was used in the publication, so they're looking for consistency. If you'd like to use the most recent, then there might be very minor differences.

Viral cluster automatic cluster taxonomy
This app completed without errors in 1h 33m 36s.
Summary
Basic message to show in the report
Links
Files
These are only available in the live Narrative: https://narrative.kbase.us/narrative/186849
  • c1.ntw.tar.gz - ClusterONE network file suitable for import into Cytoscape or other graph tools
  • genome_by_genome_overview.csv - Final summary file generated directly by vConTACT2

A closer look at the vConTACT2 results

*NOTE - this section and screenshots are being updated! This content may not match the output of the vConTACT

After you've run vConTACT2, you'll get a table with ALL of the genomes in the analysis. The table can be a bit unwieldly as it contains a lot of rows and columns.

The easiest way to manage this (in KBase) is to use the filter function to find YOUR viral genomes. For example, all of our genomes contain NODE - it's a byproduct of the SPAdes assembler. By adding NODE in the appropriate filtering row under the column "Genome" will filter out all the reference genomes.

Image of NODE filtering on column
Filtering vConTACT2 results table using "NODE"

There are 2,638 viral genomes remaining. Now, how many of our viral genomes are clustered? Add Clustered to the "VC Status" column. There are 1,262 Clustered (or Clustered/Singletons) genomes. Not bad, not great - but this is actual data - not pretty "mock" data.

Image of NODE and VC Status filtering on column
Filtering vConTACT2 results table using "NODE" and "Clustered"

Now let's find out if our viral genomes are associated with any references. This is not fast through the table in KBase, but it is doable, unless you want to download the csv and do some data wrangling. What I do is sort the table by "VC" and scroll through, keeping track of the "Size" and count of the VC. If the "Size" of the VC is greater than the counts of the VC, make note of that VC.

After sorting by VC...

Image of NODE and VC Status filtering on column, then sorted
Sorting vConTACT2 results table after using "NODE" and "Clustered"

And finding some interesting clusters!

Image of NODE and VC Status filtering on column, then sorted
Finding VCs of interest after filtering, sorting, and comparing

VC_774

One of the first examples we encounter is VC_774. It has a VC size of 3, but only 1 member is seen. Remember - we still have the `NODE` filter on! Remove that NODE filter, revealing Cyanophage PSS2 and Synechococcus phage S-CBS2.

Image of VC 332
Revealing VC 332 without "NODE" filter on, revealing its members

This is excellent, as the Cyanophage and Synechococcus phage are incredibly common in the ocean. A literature search revealed that these sequences are indeed found at the station our SRA reads are derived!

Additional searching reveals at least 7 other VCs that include both reference data and environmental sequences:

  • VC_286
  • VC_331
  • VC_332
  • VC_333
  • VC_339
  • VC_461
  • VC_497
  • VC_506

And taking a quick peek into one of the above VCs...

Image of VC 497
VC 497 and its Pelagibacter phage member

A quick reference guide to vConTACT2

What is a "VC"?

A Viral Cluster is a unit of classification. vConTACT2 uses two terms - with subtle differences - in order to classify sequences.

A VC is a first-pass classification of viral genomes. This frequently represents a group of genomes within the same genus. An example is VC_115.

A VC Subcluster is a second-pass classification. It uses a pre-calculated distance calculation to refine the VCs. These are high-confidence, genus-level groupings. It is common for there to be no change between VC and VC Subcluster. An example is VC_115_0. The final value ("0") represents the subcluster within the original VC. If all members of a VC Subcluster have VC_xx_0, then there was no change. However, if there were further refinements, then the VC Subcluster would be: VC_115_0, VC_115_1, VC_115_2...

VC Status

How vConTACT2 describes how it classified a sequence

Clustered: Genomes "successfully" placed into a genus-level group alongside at least one other genome.

Singleton: Genome was not found to be related to any other genome in the dataset. Most likely reason? Very weak or no overlap with any other genes found on any other genome. How to fix? Add more related genomes.

Clustered/Singleton: Genome was initially clustered, but distance-based optimization identified its placement in the cluster as not genus-level. However, no other genomes were found to be within the same "subcluster" as this genome, and resulted in the genome being stranded, without another member. For these genomes, you can look at its "VC" to see distantly related members. So a viral genome that is Clustered/Singleton (VC_221_1 or VC_221_2) is related to other VC_221 members, but vConTACT2 does not have confidence that these genomes are related at the genus level.

As an aside, VC_221 contains 6 genomes. Four (4) are in VC_221_0, with the other two in VC_221_1 and VC_221_2. Both those genomes are related to the four members of 221_0, but not at the genus level.

Overlap (VC_NN/VC_XX): Genomes identified as sharing significant portions of its gene content with multiple VCs. In other words, vConTACT2 cannot confidently assign it to one OR the other VC. This is incredibly common for viral groups that undergo extensive recombination.

Outlier: Genomes clustered by ClusterONE (a tool used internally by vConTACT2) but were not strongly connected to the other VC members. It is common for these genomes to share a single gene or two to the other members of its closest VC, however the other members in that VC are likely be be sharing 20, 30 or 50+ genes. It is not only unlikely that that particular genome is related at the genus level, but could perhaps be a spurious shared gene and is unlikely to be related at anything lower than family or order.

Predict host-virus relationships with VirMatcher

A tool to predict virus-host relationships, leveraging a variety of bioinformatic methods, including; host CRISPR-spacers, integrated prophage, host tRNA genes, and k-mer signatures calculated by WIsH. The methodology is described in detail in Gregory et al, 2020.

Predicts host-virus matches
This app completed without errors in 2h 23m 26s.
Summary
Basic message to show in the report
Links

Conclusions

In summary, we've processed a viral metagenome from public reads available on SRA, identified contigs from the assembled sequence data as putative viruses using VirSorter2, and classified them in approximately genus-level clusters with vConTACT2. This analysis revealed 8 VCs where environmental sequence data was found associated with references, and we can have confidence that those sequences at related to those references at the genus level.

Additionally, we've seen a few larger clusters with no associations to reference sequences that could be further investigated using KBase tools. For example - align those genomes in those VCs, identify shared features, extract those features, and make a discovery about certain proteins found throughout your dataset. Or, go one step further and pull in JGI data and then compare against a variety of JGI datasets for global significance - all using existing KBase apps!

Thanks for following along and using these apps for your research!

For further reading:

  • Guo J, Bolduc B, Zayed AA, Varsani A, Dominguez-Huerta G, Delmont TO, et al. VirSorter2: a multi-classifier, expert-guided approach to detect diverse DNA and RNA viruses. Microbiome. 2021;9: 37. doi:10.1186/s40168-020-00990-y
  • VirSorter2 Paper
  • Shaffer M, Borton MA, McGivern BB, Zayed AA, La Rosa SL, Solden LM, et al. DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Research. 2020;48: 8883–8900. doi:10.1093/nar/gkaa621
  • DRAM-v Paper
  • Bin Jang H, Bolduc B, Zablocki O, Kuhn JH, Roux S, Adriaenssens EM, et al. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat Biotechnol. 2019;37: 632–639. doi:10.1038/s41587-019-0100-8
  • vConTACT2 Paper
  • Roux S, Enault F, Hurwitz BL, Sullivan MB. VirSorter: mining viral signal from microbial genomic data. PeerJ. 2015;3: e985. doi:10.7717/peerj.985
  • The original VirSorter paper
  • Bolduc B, Jang HB, Doulcier G, You Z-QZ, Roux S, Sullivan MB. vConTACT: an iVirus tool to classify double-stranded DNA viruses that infect Archaea and Bacteria. PeerJ. 2017;5: e3243. doi:10.7717/peerj.3243
  • The initial vConTACT paper describing the theory and background

  • Jang HB, Bolduc B, Zablocki O, Kuhn JH, Roux S, Adriaenssens EM, et al. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat Biotechnol. 2019;37: 632–639. doi:10.1038/s41587-019-0100-8
  • The significantly improved version of vConTACT that's faster, more accurate, and capable of handling larger datasets

Known Issues & Update History

Known Issues

  • VirSorter may generate a "Bad KBase object" error. Can be safely ignored
  • vConTACT2 can only work on a single annotated Genome and Assembly object

Update History (MM-DD-YYYY)

08-01-2024

11-03-2020

  • Revised introduction
  • Added navigation links to TOC and quick home buttons throughout
  • Adjusted some formatting to adhere better to KBase styling guidelines
  • Significantly expanded vConTACT2 results section
  • Added vConTACT2 quick reference guide
  • Expanded conclusion section

10-26-2020

  • Initial draft of tutorial built

Feedback & Helpdesk

Was this Narrative helpful? Please provide feedback and let us know: https://forms.gle/DRfvAQDxsv1LKr7FA

If you have a question about one of our apps, need to report a bug or have another system-related query, please join our Help Board and post a ticket. Learn about how to do this here: http://kbase.us/help-board.

import biokbase.narrative.clients as clients
ws = biokbase.narrative.clients.get('workspace')
ws.get_workspace_info({'id': '186849'})
Out[4]:
[186849,
 'allenbh:narrative_1721147566557',
 'allenbh',
 '2024-07-30T20:10:11+0000',
 93,
 'a',
 'r',
 'unlocked',
 {'narratorial': '1',
  'searchtags': 'narrative',
  'is_temporary': 'false',
  'narrative': '72',
  'narrative_nice_name': 'Viral Genomics in KBase',
  'cell_count': '31',
  'narratorial_description': 'Viral Genomics in KBase'}]
import biokbase.narrative.clients as clients
ws = biokbase.narrative.clients.get('workspace')

id = 186849
ws.alter_workspace_metadata({
        'wsi': {
            'id': id
        }, 
        'new': {
            'narratorial': '1',
            'narratorial_description': "Viral Genomics in KBase"
        }
    })

Apps

  1. Annotate and Distill Viral Assemblies with DRAM-v
    • DRAM source code
    • DRAM documentation
    • DRAM Tutorial
    • DRAM publication
  2. Assemble Reads with metaSPAdes - v3.15.3
    • Nurk S, Meleshko D, Korobeynikov A, Pevzner PA. metaSPAdes: a new versatile metagenomic assembler. Genome Res. 2017; 27:824 834. doi: 10.1101/gr.213959.116
    • Prjibelski A, Antipov D, Meleshko D, Lapidus A, Korobeynikov A. Using SPAdes De Novo Assembler. Curr Protoc Bioinformatics. 2020 Jun;70(1):e102. doi: 10.1002/cpbi.102.
  3. Assess Read Quality with FastQC - v0.12.1
    • FastQC source: Bioinformatics Group at the Babraham Institute, UK.
  4. Classify Taxonomy of Metagenomic Reads with Kaiju - v1.9.0
    • Chivian D, et al. Metagenome-assembled genome extraction and analysis from microbiomes using KBase. Nat Protoc. 2023 Jan;18(1):208-238. doi: 10.1038/s41596-022-00747-x
    • Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257. doi:10.1038/ncomms11257
    • Ondov BD, Bergman NH, Phillippy AM. Interactive metagenomic visualization in a Web browser. BMC Bioinformatics. 2011;12: 385. doi:10.1186/1471-2105-12-385
    • Kaiju Homepage:
    • Kaiju DBs from:
    • Github for Kaiju:
    • Krona homepage:
    • Github for Krona:
  5. Import SRA File as Reads From Web - v1.0.10
    • Arkin AP, Cottingham RW, Henry CS, Harris NL, Stevens RL, Maslov S, et al. KBase: The United States Department of Energy Systems Biology Knowledgebase. Nature Biotechnology. 2018;36: 566. doi: 10.1038/nbt.4163
  6. Trim Reads with Trimmomatic - v0.39
    • Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30: 2114 2120. doi:10.1093/bioinformatics/btu170
  7. vConTACT2 0.9.19
    • Bin Jang, H., Bolduc, B., Zablocki, O., Kuhn, J. H., Roux, S., Adriaenssens, E. M., Sullivan, M. B. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nature Biotechnology. 2019;37(6): 632 639. https://doi.org/10.1038/s41587-019-0100-8
  8. VirMatcher 0.3.3
    • Gregory, A. C. et al. The Gut Virome Database Reveals Age-Dependent Patterns of Virome Diversity in the Human Gut. Cell Host Microbe 28, 724-740.e8 (2020). https://doi.org/10.1016/j.chom.2020.08.003
  9. VirSorter2
    • Guo, J. et al. VirSorter2: a multi-classifier, expert-guided approach to detect diverse DNA and RNA viruses. Microbiome 9, 37 (2021). https://doi.org/10.1186/s40168-020-00990-y