Curate reference tree: remove recombinant/frameshifted/low-quality sequences#465
Curate reference tree: remove recombinant/frameshifted/low-quality sequences#465nneune wants to merge 7 commits into
Conversation
TestingTry in Nextclade Web: ScienceTaxonomy, genome, and reference [click to expand]Coxsackievirus A16 (CVA16) belongs to genus Enterovirus, species Enterovirus A, family Picornaviridae -- a positive-sense single-stranded RNA virus and one of the major causative agents of hand, foot, and mouth disease (HFMD) in children. The CVA16 genome is approximately 7,400 nt and encodes a single polyprotein processed by viral proteases into structural proteins (VP4, VP2, VP3, VP1 in P1) and non-structural proteins (2A-2C in P2, 3A-3D in P3). VP1 is the standard molecular typing target for enteroviruses (Oberste et al., J Clin Microbiol 1999; Chen et al., PLoS ONE 2013). The dataset uses a Static Inferred Ancestor (SIA) as reference rather than the historical G-10 prototype strain (U05876.1), which diverges substantially from circulating strains. The upstream workflow constructs this ancestor by outgroup-rooting a CVA16 phylogeny, reconstructing the ingroup MRCA, and filling alignment gaps from the prototype reference (inference workflow). Lineage nomenclature and recombination [click to expand]CVA16 genotypes are defined by VP1 phylogeny. Chen et al. (2013) established genotypes A (prototype G-10) and B, with B subdivided into B1a, B1b, and B1c. A recent synthesis recognizes genotypes A, B, and D, with B1/B2 under B and B1a/B1b/B1c under B1 (Han et al., Virus Evol 2024). Hassel et al. (2017) identified clade D with intertype recombinant origin. Whole-genome classifications differ because recombination changes phylogenetic relationships outside VP1. Han et al. identified recombinant forms RF-A through RF-E from the 3D region, while another whole-genome analysis proposed G-a through G-e (Chu et al., Heliyon 2024). Both B1a and B1b contain sequences from other Enterovirus A donors in the 5'UTR and P2/P3 regions (Chen et al., 2013). The dataset's approach of removing singleton recombinants while retaining recurring circulating recombinant forms is standard curation practice. The upstream workflow explicitly classifies excluded accessions as recombinant outliers, frameshifted sequences, or UTR-associated outliers (curation commit). Epidemiological context and tree coverage [click to expand]B1a, B1b, and B1c remain epidemiologically relevant. B1b dominated a long-term mainland China dataset (Han et al., Virus Evol 2020), while B1a predominated among Thai CVA16 sequences through 2022 (Noisumdaeng and Puthavathana, Sci Rep 2023). In Hangzhou, B1c rose sharply during 2024 (Xu et al., Front Microbiol 2025). The curated tree spans 1997-2025 with 741 tips. Year distribution shows strong representation in 2008-2024, with 2024 having the most tips (111). The 2020 dip (2 tips) reflects reduced surveillance during the COVID-19 pandemic. One 2025 tip (PZ117058) extends coverage to recent circulation. The dataset is maintained by ENPEN (European Non-Polio Enterovirus Network), with the workflow at enterovirus-phylo/nextclade_a16. All sequence data is from GenBank; no restricted sources are identified. Blocking issuesCorrectness concerns worth addressing before merge. 🔴 H1. Nucleotide mutation labels emptied without documentation [click to expand]The The generated Nextclade uses these maps to label private mutations in the UI and tabular output. Labeled substitutions receive distinct QC weights to flag potential contamination, co-infection, or recombination (mutation-label documentation, private-mutation algorithm). Effect: all former clade-associated private mutations become unlabeled and receive weight 1 instead of 1.5. User-visible mutation annotations disappear, QC scores change, and the index metadata becomes misleading. Neither the PR description nor the CHANGELOG mention this change. Fix: regenerate Non-blocking issuesConvention drift and minor inconsistencies. Fix if time allows. 🟡 M1. Broken JSON indentation in pathogen.json [click to expand]Several sections have inconsistent indentation compared to the base version [src]:
Effect: valid JSON, but confusing for future manual edits. Likely artifacts of manual editing after deleting the large Fix: reformat with 🟡 M2. Recombinant clade nomenclature lacks explicit mapping [click to expand]The README states that recombinant forms C-F are "also referred to as B2, B3, and D" [src]. The tree contains clades A, B1, B1a, B1b, B1c, C, D, E, F, and the parent label RFs. Published nomenclatures are not directly interchangeable:
The current text does not establish how dataset clades C-F map to these systems. Fix: add a compact mapping table giving each dataset clade, the genomic region used to define it, the corresponding published designation, and a defining citation. If C-F are dataset-specific labels, state that explicitly. 🟡 M3. No `meta.bugs` or `meta["source code"]` URLs in pathogen.json [click to expand]Neither Fix: add 🔵 L1. "Xu et al." citation not identifiable from CHANGELOG [click to expand]The CHANGELOG references "Xu et al." without a publication year or title [src]. The upstream workflow's testing file Fix: replace with "Xu et al. (2025)." 🔵 L2. Trailing whitespace in README [click to expand]
Fix: remove trailing space and rebuild generated output. Clade distribution9 clades, 783 -> 741 tips [click to expand]
All clades retained. Removals concentrated in B1a and B1b. NotesClick to expand
|
Summary
Manually curated the dataset to improve reference tree quality: