Hi BAQLaVa team,
Thanks for developing BAQLaVa. I’m working through the viral trait enrichment analysis from the manuscript (Fig. 4) and would like to confirm an implementation detail.
Which ORF set was used for the Pfam/VFAM annotation?
The Methods state:
“We assigned molecular functions to BAQLaVa ORFs by annotating the ORFs to Pfams v37.0 and VFAMs (release 229) with HMMER v3.4…”
I assume “BAQLaVa ORFs” refers to the full ORF set distributed as BAQLaVa_ORFs_clean.faa, rather than the VGB-specific subset in utility_files/translated_protein_reference.txt. Could you confirm?
The two files clearly differ. For BAQ00000001, BAQLaVa_ORFs_clean.faa contains ORFs numbered contiguously from 1, whereas translated_protein_reference.txt retains only 38 non-contiguous entries (4, 5, 7, 8, 9, 12, 13, 20, …), consistent with the uniqueness and >200 nt filters applied when building the protein markers. Since the marker subset covers only ~50% of VGBs and is by construction depleted of proteins conserved across VGBs, the choice materially affects the trait annotation, so I’d like to use the same set you did. Even a one-line confirmation would be very helpful.
Thanks very much.
Best regards