VEE Lab Tools // Database · Live build Feb 2025
ArchaeaHQ

ArchaeaHQ

A quality-controlled, systematically curated reference database of 21,644 archaeal genomes across all four kingdoms — filtered, de-replicated, and linked with environmental and geographic metadata. Built to be explored and downloaded, not just read.

V1.0
arrow_back Back to Resources
Results · 01 — The Database

A High-Quality Archaeal Reference

Public archives grow fast but inconsistently — a mixture of pristine genomes, fragmented assemblies, and records with missing environmental data. ArchaeaHQ resolves that into a single analysis-ready set.

Abstract

Archaea are major players in biogeochemical cycles across diverse ecosystems, yet archaeal genomes remain underrepresented in the datasets used by popular computational biology tools. Here we present ArchaeaHQ, a quality-controlled, systematically curated reference database of 21,644 archaeal genomes compiled from 35,993 assemblies across all four archaeal kingdoms retrieved from NCBI: Methanobacteriati (Euryarchaeota), Thermoproteati (TACK), Nanobdellati (DPANN), and Promethearchaeati (Asgard). All genomes pass a standardized quality control requiring ≥70% completeness and ≤10% contamination. 44.2% achieve ≥90% completeness and 93.1% exhibit ≤5% contamination. ArchaeaHQ presents 16,199 metagenome-assembled genomes (74.8%) and 5,445 isolate genomes (25.2%). MAGs are classified into 17 ecologically meaningful categories, and 65.4% of genomes carry geographic metadata — an analysis-ready reference for metagenomic classification, biogeochemical and ecological studies, comparative genomics, and archaeal-specific tool development.

Full preprintRead the complete methods, results, and references on bioRxiv.
Read the manuscript north_east
How to cite

Bespiatykh, D., et al. (2026). ArchaeaHQ: A Quality-Controlled, Systematically Curated Reference Database of Archaeal Genomes. bioRxiv. https://doi.org/10.64898/2026.06.02.729493

Results · 02 — Curation

From 35,993 Assemblies to 21,644 Genomes

A multi-step pipeline applied contamination and completeness gates, then removed identity- and species-level redundancy — a 39.8% net reduction that keeps quality high while preserving the full taxonomic spectrum.

ArchaeaHQ database construction Sankey diagram

Fig. 1 — Sankey diagram of the sequential filtering steps from raw NCBI archaeal assemblies (35,993) to the final ArchaeaHQ collection (21,644). Colour bands represent the four archaeal kingdoms; grey flows are removed at each stage.

Unique species-level genomes retained per kingdom
Results · 03 — Genome Quality

Completeness & Contamination

Every genome passed the ≥70% completeness / ≤10% contamination gate. Filled points meet the MIMAG high-quality bar (≥90% complete, ≤5% contaminated). Toggle kingdoms in the legend.

Quality landscape

All 21,644 genomes · CheckM2 v1.0.1 · click legend to isolate a kingdom

Contamination (%)
Completeness (%)

High-quality genomes by kingdom

Share meeting ≥90% completeness & ≤5% contamination

Results · 04 — Environmental Diversity

17 Ecological Categories

The in-depth curation of environmental context is what sets ArchaeaHQ apart. Each genome is normalized into one of 17 categories from five metadata fields. Bars are stacked by kingdom — hover for detail.

Genomes per environment

Stacked by archaeal kingdom · 21,644 genomes

Results · 05 — Biogeography

Global Sampling Footprint

Geographic metadata are available for 14,168 genomes (65.4% of the database), enabling spatial analyses of archaeal biogeography. Georeferencing was verified from BioSample-level records during curation.

Global geographic distribution of ArchaeaHQ genomes

Fig. 2 — Distribution of the 14,168 georeferenced genomes. Each dot marks a sampling site, coloured by archaeal kingdom.

Top sampling locations

Genomes with a parseable country / ocean of origin

Georeferenced coverage per kingdom
The Database · Interactive

Genome Browser

Search and filter all 21,644 genomes by kingdom, environment, assembly type, completeness, and free text (species, accession, country). Accessions link directly to NCBI.

Click a column header to sort · tick rows to build a download set
No genomes selected Tick the checkboxes to build a set, then export the accessions to fetch the sequences from NCBI.
Results · 06 — In Context

How ArchaeaHQ Compares

Existing resources solve parts of the problem. ArchaeaHQ is designed specifically as a downloadable, quality-filtered, environmentally annotated archaeal reference set that stays compatible with NCBI-based workflows.

NCBI GenBank
Comprehensive archive
  • Most comprehensive assembly archive
  • Highly variable quality & completeness
  • No curated archaeal subset
  • Inconsistent environmental metadata
GTDB
Taxonomy framework · R09-RS220
  • Phylogenetically consistent taxonomy
  • ~Representative genomes only
  • Not an analysis-ready download set
  • Diverges from NCBI accessions
IMG/M
JGI · rich metadata
  • Large MAG collection
  • Rich environmental context
  • Not a precompiled download
  • Hard to use in CLI pipelines
ArchaeaHQ
This work
  • Quality-controlled & de-replicated
  • Fully downloadable, analysis-ready
  • Environmental data as a core feature
  • NCBI taxonomy & accession compatible
Methods

How It Was Built

The reproducible pipeline behind every genome in the database.

Use ArchaeaHQ in Your Research

ArchaeaHQ is freely available. Reach out to discuss collaborations, data access, or training in archaeal comparative genomics.

How to cite

Bespiatykh, D., et al. (2026). ArchaeaHQ: A Quality-Controlled, Systematically Curated Reference Database of Archaeal Genomes. bioRxiv. https://doi.org/10.64898/2026.06.02.729493