[Rabies] Initialize Lyssavirus rabies all-clades community dataset #333
[Rabies] Initialize Lyssavirus rabies all-clades community dataset #333xonq wants to merge 7 commits into
Conversation
|
Thanks! Seems to be working As a dev I can only review the technical side. And I will let our scientists to check the sciency bits :) The virus is quite diverse it seems - lots of mutations. But this is probably expected. If you have an public repo where you prepare trees and other data for the dataset, it would be a great help to the users of your dataset if you add it to the readme. We typically use a boilerplate like this in Nextstrain datasets: But that's not mandatory. All looks good to me. Smooth work! |
|
Thank you. Rabies is indeed very diverse - I contemplated creating independent datasets for each clade, but the genotyping of this "all-clades" dataset has been sufficient for our SME partners. Additionally, there were issues with sub-clade metadata quality that limit the improvements more refined datasets may provide. I do not have a repository for tree building - I built the Nextclade dataset from the Nextstrain rabies build as a template, though I ended up deviating with the tree building methodology and metadata acquisition. The methodology is hopefully adequately documented for users in this PR's README. |
|
Thanks a lot for contributing this dataset! Overall, this looks very good. But I have a few suggestions to make it better.
|
|
hey @rneher, just wanted to reply and inform you that I cannot return to this to address your concerns until a later date. not sure when, but hopefully within the next several weeks. Thanks for your suggestions and my apologies for my ignorance to some of the standardized procedures. RE: alignment parameters: I'm not really certain how to systematically adjust these parameters - do you have specific recommendations/procedures to determine what parameters are more ideal, or do you suggest dragging and dropping the linked pathogen.json you sent? RE: Apart from the tree-building, the workflow was performed with AUGUR. With this in mind, do these steps deviate from Nextclade like you're suggesting?: Alignment:Tree building:performed as discussed in the README Refinement:Trait application:Nucleotide mutation calling:Translation:Clade mutation extraction (non-AUGUR):Clade mutation application:Export: |
|
Hey all, just chiming in to see if I could get some guidance on a couple points to move this forward:
|
|
Hi @xonq sorry for not responding earlier. I had read the first part of your message "... later day." and mentally shelved it under 'later'. Re alignment parameters: you can add these to the pathogen json. There is also a pre-set related, we recommend doing alignment and translation for your tree using nextclade rather than The latter is probably also responsible for mismatch of the amino acid mutations in the alignment/output and the reference tree. In the reference tree, CDS are named like |
|
Posting robotic review below. It mostly parrots what Richard already said, plus some low prio typos/citation oopsies (to be confirmed!). Nothing much to stress about
First review of TestingObservedWhole-genome Lyssavirus rabies genotyping dataset -- 3,296 tips, 48 clades, reference Blocking issues🔴 F1. Amino-acid CDS names in the tree do not match the genome annotation [click to expand]The reference tree and the genome annotation name the five CDS differently:
Coordinates match (L = 5418-11846 in both), only the names differ. Nextclade derives CDS names from the GFF3, so query AA mutations show up as Same mismatch Suggestions:
Non-blocking issues🟡 F2. README cites an unrelated paper (wrong Campbell 2022) [click to expand]The
That DOI resolves to a real but entirely unrelated paper on mental health of UK university students. The dataset's subclade labels ( Suggestions:
🟡 F3. No alignment tuning for a divergent virus [click to expand]
Suggestions:
🟡 F4. QC configuration is minimal compared with sibling datasets [click to expand]Only Suggestions:
🔵 F5. Local build path leaked into tree annotations [click to expand]Every entry in Suggestions:
🔵 F6. Boilerplate/description mismatch in README and dataset name [click to expand]
Suggestions:
🔵 F7. Citation formatting errors [click to expand]In Suggestions:
🔵 F8. Missing trailing newlines [click to expand]
Suggestions:
Clade distribution48 clades, 3,296 tips [click to expand]Coverage spans all major RABV clades. Dominant clades reflect sequencing effort (North American RAC-SK, Arctic, and SE-Asian lineages are heavily represented). Several README-included subclades are present but sparsely sampled (<3 tips), which limits confident assignment for those lineages.
Validation summaryValidation checks [click to expand]GFF3 annotation
Reference
Tree
Generated output
Registration
Nextclade CLI run
NotesClick to expand
BackgroundTaxonomy, genome, and clade system [click to expand]Taxonomy and reference strainRabies virus is the type species of genus Lyssavirus, family Rhabdoviridae, order Mononegavirales; non-segmented negative-sense ssRNA (Baltimore group V). ICTV binomial: Lyssavirus rabies (formerly "Rabies lyssavirus"); NCBI taxid 11292. The reference Genome organizationFive genes in conserved order 3'-N-P-M-G-L-5' (Wikipedia). The GFF3 labels P as "M1" and M as "M2" -- that's the Clade / lineage classificationRABV splits into bat-related and dog-related groups with different evolutionary dynamics (Troupin et al. 2016). Dog-related major clades: Cosmopolitan, Arctic, Asian, Africa-2, Africa-3, Indian subcontinent. Bat-related: bat and RAC-SK (raccoon-skunk). The finer subclade labels ( |
This pull request initializes a Lyssavirus rabies (rabies) Nextclade dataset with clade-subclade resolution. Created in collaboration with @kimandrews and with subject matter expertise/user input from Massachusetts Department of Public Health. Please review the README.md for information on dataset creation and citations.