# Clarifying "BAQLaVa ORFs" used for the manuscript's trait annotation

**URL:** <https://forum.biobakery.org/t/clarifying-baqlava-orfs-used-for-the-manuscripts-trait-annotation/9019>\
**Category:** BAQLaVa\
**Created:** [July 31, 2026, 4:16pm UTC](https://forum.biobakery.org/t/clarifying-baqlava-orfs-used-for-the-manuscripts-trait-annotation/9019 "2026-07-31T16:16:26Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![dsh82](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/dsh82/32/3712_2.png) [@dsh82](https://forum.biobakery.org/u/dsh82)\
**Post date:** [July 31, 2026, 4:16pm UTC](https://forum.biobakery.org/t/clarifying-baqlava-orfs-used-for-the-manuscripts-trait-annotation/9019/1 "2026-07-31T16:16:26Z")

</div>

Hi BAQLaVa team,

Thanks for developing BAQLaVa. I’m working through the viral trait enrichment analysis from the manuscript (Fig. 4) and would like to confirm an implementation detail.

**Which ORF set was used for the Pfam/VFAM annotation?**

The Methods state:

> “We assigned molecular functions to BAQLaVa ORFs by annotating the ORFs to Pfams v37.0 and VFAMs (release 229) with HMMER v3.4…”

I assume “BAQLaVa ORFs” refers to the full ORF set distributed as `BAQLaVa_ORFs_clean.faa`, rather than the VGB-specific subset in `utility_files/translated_protein_reference.txt`. Could you confirm?

The two files clearly differ. For `BAQ00000001`, `BAQLaVa_ORFs_clean.faa` contains ORFs numbered contiguously from 1, whereas `translated_protein_reference.txt` retains only 38 non-contiguous entries (4, 5, 7, 8, 9, 12, 13, 20, …), consistent with the uniqueness and \>200 nt filters applied when building the protein markers. Since the marker subset covers only ~50% of VGBs and is by construction depleted of proteins conserved across VGBs, the choice materially affects the trait annotation, so I’d like to use the same set you did. Even a one-line confirmation would be very helpful.

Thanks very much.

Best regards
