# Low percentage of aligned reads

**URL:** https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164
**Category:** HUMAnN
**Created:** [January 27, 2020, 8:53pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164 "2020-01-27T20:53:52Z")
**Posts on this page:** 18
**Page:** 1

<div class="post-metadata">

### Author: ![CHRISTINE\_OLSON](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/christine_olson/32/105_2.png) [@CHRISTINE\_OLSON](https://forum.biobakery.org/u/CHRISTINE_OLSON)
#### Post date: [January 27, 2020, 8:53pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/1 "2020-01-27T20:53:52Z")

</div>

Hello,

When running humann2, for some samples in our dataset we see a high percentage of unaligned reads (80-90%), and for others, the unaligned read count is closer to ~15% [example from Terminal below]. This is with running the Uniref50 database. I have seen at least one forum post where Eric ran through troubleshooting a similar problem with someone else ([https://groups.google.com/forum/m/#!msg/humann-users/vcGFdlBnZA0/kU5BHAYyBQAJ](https://groups.google.com/forum/m/#!msg/humann-users/vcGFdlBnZA0/kU5BHAYyBQAJ)), but I have not been successful with reducing the unaligned reads.

I am wondering whether you can help me troubleshoot what is happening with the samples with high % unaligned reads? Checking the quality using FastQC, I see no adapter sequences, I did deplete host DNA content using the databases for kneaddata, and checking FastQC does show a warning for duplicated sequences, but the totals to less than 0.5%. My understanding is that this is relatively low, so probably not what is driving the issue for this sample. Any advice you can give on parameters to change or check and how would be very useful. Thank you.

Christine

EXAMPLE RUN:

Running humann2 on this test sample:

Running metaphlan2.py …

Found g\_\_Blautia.s\_\_Ruminococcus\_gnavus : 20.85% of mapped reads

Found g\_\_Enterococcus.s\_\_Enterococcus\_avium : 19.59% of mapped reads

Found g\_\_Gammaretrovirus.s\_\_Murine\_osteosarcoma\_virus : 14.90% of mapped reads

Found g\_\_Anaerostipes.s\_\_Anaerostipes\_caccae : 12.52% of mapped reads

Found g\_\_Anaerostipes.s\_\_Anaerostipes\_unclassified : 6.49% of mapped reads

Found g\_\_Erysipelotrichaceae\_noname.s\_\_Clostridium\_innocuum : 5.45% of mapped reads

Found g\_\_Turicibacter.s\_\_Turicibacter\_unclassified : 4.48% of mapped reads

Found g\_\_Peptostreptococcaceae\_noname.s\_\_Clostridium\_difficile : 4.19% of mapped reads

Found g\_\_Erysipelotrichaceae\_noname.s\_\_Erysipelotrichaceae\_bacterium\_2\_2\_44A : 2.97% of mapped reads

Found g\_\_Betaretrovirus.s\_\_Mouse\_mammary\_tumor\_virus : 2.14% of mapped reads

Found g\_\_Lactococcus.s\_\_Lactococcus\_lactis : 2.04% of mapped reads

Found g\_\_Listeria.s\_\_Listeria\_monocytogenes : 1.30% of mapped reads

Found g\_\_Blautia.s\_\_Ruminococcus\_torques : 1.22% of mapped reads

Found g\_\_Clostridiaceae\_noname.s\_\_Clostridiaceae\_bacterium\_JC118 : 0.88% of mapped reads

Found g\_\_Turicibacter.s\_\_Turicibacter\_sanguinis : 0.50% of mapped reads

Found g\_\_Propionibacterium.s\_\_Propionibacterium\_acnes : 0.34% of mapped reads

Found g\_\_Corynebacterium.s\_\_Corynebacterium\_kroppenstedtii : 0.13% of mapped reads

Total species selected from prescreen: 17

Selected species explain 100.00% of predicted community composition

Creating custom ChocoPhlAn database …

Running bowtie2-build …

Running bowtie2 …

Total bugs from nucleotide alignment: 15

g\_\_Gammaretrovirus.s\_\_Murine\_osteosarcoma\_virus: 25 hits

g\_\_Erysipelotrichaceae\_noname.s\_\_Clostridium\_innocuum: 31300 hits

g\_\_Betaretrovirus.s\_\_Mouse\_mammary\_tumor\_virus: 13 hits

g\_\_Propionibacterium.s\_\_Propionibacterium\_acnes: 1059 hits

g\_\_Clostridiaceae\_noname.s\_\_Clostridiaceae\_bacterium\_JC118: 2754 hits

g\_\_Blautia.s\_\_Ruminococcus\_torques: 975 hits

g\_\_Peptostreptococcaceae\_noname.s\_\_Clostridium\_difficile: 33851 hits

g\_\_Listeria.s\_\_Listeria\_monocytogenes: 4844 hits

g\_\_Turicibacter.s\_\_Turicibacter\_sanguinis: 7739 hits

g\_\_Enterococcus.s\_\_Enterococcus\_avium: 62526 hits

g\_\_Corynebacterium.s\_\_Corynebacterium\_kroppenstedtii: 508 hits

g\_\_Erysipelotrichaceae\_noname.s\_\_Erysipelotrichaceae\_bacterium\_2\_2\_44A: 31950 hits

g\_\_Anaerostipes.s\_\_Anaerostipes\_caccae: 36787 hits

g\_\_Lactococcus.s\_\_Lactococcus\_lactis: 8110 hits

g\_\_Blautia.s\_\_Ruminococcus\_gnavus: 41868 hits

Total gene families from nucleotide alignment: 21047

Unaligned reads after nucleotide alignment: 87.8752111330 %

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [January 27, 2020, 10:29pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/2 "2020-01-27T22:29:28Z")

</div>

Agreed that the duplicated reads are likely not the issue here.

The example you pasted shows the % of reads unaligned after the pangenome (nucleotide) search stage but not the translated search stage. The latter is where I’d expect a lot of extra reads to map using the UniRef50 database. Can you confirm that the % unaligned rates you give at the start of your post correspond to the end of the HUMAnN2 workflow (i.e. after all mapping steps)?

In either case, one of the best ways to diagnose low mapping rates is to take ~10 reads from the final unmapped.fasta file and blast them online ([https://blast.ncbi.nlm.nih.gov/Blast.cgi](https://blast.ncbi.nlm.nih.gov/Blast.cgi)) to see if they hit new genomes / possible contaminants / etc.

---

<div class="post-metadata">

### Author: ![CHRISTINE\_OLSON](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/christine_olson/32/105_2.png) [@CHRISTINE\_OLSON](https://forum.biobakery.org/u/CHRISTINE_OLSON)
#### Post date: [February 3, 2020, 6:50pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/3 "2020-02-03T18:50:35Z")

</div>

Hello Eric,  
Thank you for the reply. Yes, we still see a high percentage of unaligned reads after all of the mapping steps.

We diagnosed the issue by blasting sequences as you suggested. A lot of the reads correspond to mouse DNA. So, it seems likely the host DNA depletion was not successful. I have previously depleted host DNA using the kneaddata suffix ::

kneaddata -i Raw\_Files/input1.fastq.gz -i Raw\_Files/input2.fastq.gz -o kneaddata\_out   
–reference-db kneaddata/knead\_mouse\_db

Is there another way to deplete host DNA or to troubleshoot this step? Thank you so much for your assistance!

Christine

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [February 5, 2020, 6:37pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/4 "2020-02-05T18:37:44Z")

</div>

Well in one sense that means that things are working well, since the mouse-derived reads are not being mapped by HUMAnN2.

I believe the default mouse database for KneadData is for a C57BL mouse. If you’re working with a different breed, you could make a custom database out of its genome (if available) to improve depletion, though I’d be surprised if this had a dramatic effect.

Was there anything in common among the mouse sequences you saw hit in the blast searches? Repeat regions, for example, can be hard to fully deplete since they are hard to get right even in model genomes.

---

<div class="post-metadata">

### Author: ![aimirza](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/aimirza/32/499_2.png) [@aimirza](https://forum.biobakery.org/u/aimirza)
#### Post date: [July 2, 2020, 4:30pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/5 "2020-07-02T16:30:27Z")

</div>

I am having a similar problem. Number of _mapped_ reads (nucleotides): range= 16%-70%, mean= 56%. I am using the latest humann version (3.0 alpha) via conda and UniRef90, human microbiome shotgun data. I ran blastn on 1% of the kneaddata QC reads (with removed repeats) from the sample with ~85% unmapped reads using the parameters; `blastn -query seq.fa -task megablast -db nt -evalue 1e-10 -best_hit_score_edge 0.05 -best_hit_overhang 0.25 -perc_identity 50 -max_target_seqs 1`. ~18% of the reads had hits and it seems that most of them are bacterial complete genomes.

Should I relax the parameters `--nucleotide-identity-threshold` and `--translated-identity-threshold`? `humann --help` says that the `--nucleotide-identity-threshold` default is 0.0, is that correct?

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [July 2, 2020, 4:49pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/6 "2020-07-02T16:49:41Z")

</div>

The default of 0.0 for nucleotide identity means “believe any hits reported by bowtie2.” Bowtie2 doesn’t explicitly filter on alignment identity, but it’s very rare for it to report alignments \<90% nucleotide ID. The filter allows you to post-process the bowtie2 hits to be more stringent (if desired).

Otherwise it looks like your blastn results are concordant with the low end of the mapping rates you’re seeing from HUMAnN 3.0. Maybe I’m misunderstanding something about your question? Were you suggesting that blastn was mapping reads that HUMAnN wasn’t?

---

<div class="post-metadata">

### Author: ![aimirza](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/aimirza/32/499_2.png) [@aimirza](https://forum.biobakery.org/u/aimirza)
#### Post date: [July 2, 2020, 5:03pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/7 "2020-07-02T17:03:35Z")

</div>

> [@franzosa](#):
>
> Were you suggesting that blastn was mapping reads that HUMAnN wasn’t?

My mistake. I ran blastn on the **unaligned** reads only.

Yes. I was trying to figure out why humann wasnt mapping a large proportion of my reads. I thought maybe because they are not bacterial reads or they are new genomes, but I ruled those out with my blastn search and it seems that humann3 is too stringent and missing reads, I dont know.

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [July 2, 2020, 5:24pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/8 "2020-07-02T17:24:27Z")

</div>

Gotcha. I’d guess that a good chunk of the 18% mapped by blastn and not HUMAnN are hitting non-protein-coding regions (RNAs and so forth). HUMAnN is only mapping to protein-coding genes, so you’ll definitely align more if working with full-length genomes.

As far as stringency, it’s probably the database coverage filters more than the identity filters that are preventing the rest of the reads from mapping. For example, if a single read aligns well to a 1 kb gene, it will be counted as “unmapped” since the gene itself is not sufficiently well-covered to be believable (in our view). We implemented this sort of coverage filter for translated search in v2.0 and extended it to nucleotide search in v3.0 (the filters are separately tunable). These filters dramatically reduce false positives compared to just giving each read its best hit, but you are welcome to turn them off (i.e. set to 0.0) if you want to maximize the number of mapped reads.

---

<div class="post-metadata">

### Author: ![aimirza](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/aimirza/32/499_2.png) [@aimirza](https://forum.biobakery.org/u/aimirza)
#### Post date: [July 2, 2020, 5:36pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/9 "2020-07-02T17:36:39Z")

</div>

Got it. Would tuning the coverage result in higher accuracy than using UniRef50 instead of UniREf9?

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [July 2, 2020, 5:51pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/10 "2020-07-02T17:51:53Z")

</div>

Using UniRef50 over UniRef90 will have a similar effect as lowering the identity threshold (i.e. you will find hits at more remote homology); it won’t rescue hits that are lost to the above-described coverage filters. I tend to stick with UniRef90 for anything human-associated and favor UniRef50 otherwise.

---

<div class="post-metadata">

### Author: ![aimirza](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/aimirza/32/499_2.png) [@aimirza](https://forum.biobakery.org/u/aimirza)
#### Post date: [July 3, 2020, 5:40pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/12 "2020-07-03T17:40:41Z")

</div>

The log below says 29514910 (83.44%) aligned 0 times.

For this sample (the sample that lost the most reads), it appears that the majority loss of reads during translation alignment is due to identity; 12,713,631 lost due to identity vs 7,218,428 lost due to query coverage or subject coverage 9,235,839. However, in the majority of the samples, most of the lost reads during translation are due to query or subject coverage, not identity.

On average, 30% of reads do not align after _translation_. Thats common right? One of the samples had nearly 60% reads lost after translation. Would you still suggest I change the parameters for all the samples or for just this one sample? If yes, which parameters?

> humann.utilities - DEBUG: b’35371160 reads; of these:\n 35371160 (100.00%) were unpaired; of these:\n 29514910 (83.44%) aligned 0 times\n 3151724 (8.91%) aligned exactly 1 time\n 2704526 (7.65%) aligned \>1 times\n16.56% overall alignment rate\n’
> 
> humann.humann - INFO: TIMESTAMP: Completed nucleotide alignment : 529 seconds  
> humann.utilities - DEBUG: Total alignments where percent identity is not a number: 0  
> humann.utilities - DEBUG: Total alignments where alignment length is not a number: 0  
> humann.utilities - DEBUG: Total alignments where E-value is not a number: 0  
> humann.utilities - DEBUG: Total alignments not included based on large e-value: 0  
> humann.utilities - DEBUG: Total alignments not included based on small percent identity: 0  
> humann.utilities - DEBUG: Total alignments not included based on small query coverage: 0  
> humann.search.blastx\_coverage - INFO: Total alignments without coverage information: 0  
> humann.search.blastx\_coverage - INFO: Total proteins in blastx output: 136134  
> humann.search.blastx\_coverage - INFO: Total proteins without lengths: 0  
> humann.search.blastx\_coverage - INFO: Proteins with coverage greater than threshold (50.0): 90546  
> humann.search.nucleotide - DEBUG: Total nucleotide alignments not included based on filtered genes: 231574  
> humann.search.nucleotide - DEBUG: Total nucleotide alignments not included based on small percent identities: 8  
> humann.search.nucleotide - DEBUG: Total nucleotide alignments not included based on query coverage threshold: 0  
> humann.search.nucleotide - DEBUG: Keeping sam file  
> humann.humann - INFO: TIMESTAMP: Completed nucleotide alignment post-processing : 1461 seconds  
> 05/29/2020 11:18:21 PM - humann.humann - INFO: Total bugs from nucleotide alignment: 44

> humann.humann - INFO: Total gene families from nucleotide alignment: 90546  
> humann.humann - INFO: Unaligned reads after nucleotide alignment: 84.0981296627 %

> humann.humann - INFO: TIMESTAMP: Completed translated alignment : 75698 seconds  
> humann.utilities - DEBUG: Total alignments where percent identity is not a number: 0  
> humann.utilities - DEBUG: Total alignments where alignment length is not a number: 0  
> humann.utilities - DEBUG: Total alignments where E-value is not a number: 0  
> Total alignments not included based on large e-value: 4072  
> humann.utilities - DEBUG: Total alignments not included based on small percent identity: 12713631  
> humann.utilities - DEBUG: Total alignments not included based on small query coverage: 7218428  
> humann.search.blastx\_coverage - INFO: Total alignments without coverage information: 0  
> humann.search.blastx\_coverage - INFO: Total proteins in blastx output: 716163  
> humann.search.blastx\_coverage - INFO: Total proteins without lengths: 0  
> humann.search.blastx\_coverage - INFO: Proteins with coverage greater than threshold (50.0): 107222  
> humann.utilities - DEBUG: Total alignments where percent identity is not a number: 0  
> humann.utilities - DEBUG: Total alignments where alignment length is not a number: 0  
> humann.utilities - DEBUG: Total alignments where E-value is not a number: 0  
> humann.utilities - DEBUG: Total alignments not included based on large e-value: 4072  
> humann.utilities - DEBUG: Total alignments not included based on small percent identity: 12713631  
> humann.utilities - DEBUG: Total alignments not included based on small query coverage: 7218428  
> humann.search.translated - DEBUG: Total translated alignments not included based on small subject coverage value: 9235839

> humann.humann - INFO: Total gene families after translated alignment: 188899  
> humann.humann - INFO: Unaligned reads after translated alignment: 59.2021777064 %

---

<div class="post-metadata">

### Author: ![dadahan](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/dadahan/32/414_2.png) [@dadahan](https://forum.biobakery.org/u/dadahan)
#### Post date: [July 3, 2020, 10:24pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/13 "2020-07-03T22:24:14Z")

</div>

Could the high % of unaligned reads after translated alignment also be due to you have novel species in your community - eg those not found in metaphlan2? Or is the translated unaligned % the % after mapping to the whole UniRef EC90 database (species agnostic)?

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [July 7, 2020, 7:54pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/14 "2020-07-07T19:54:17Z")

</div>

30% unmapped after translated search seems very reasonable (check out Fig. S3 of the HUMAnN 2.0 paper for reference). I would _not_ recommend varying the parameters sample-by-sample – this just seems like a pathway to unintended batch effects.

Are you confident that the samples have been fully host depleted? Perhaps the outlier sample has an excess of hidden contamination?

Otherwise you could, as an experiment, try running the outlier sample in UniRef50 mode to see if more reads align with relaxed translated search (as might be expected if the sample contains an abundant unknown species).

---

<div class="post-metadata">

### Author: ![huananfshi](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/huananfshi/32/415_2.png) [@huananfshi](https://forum.biobakery.org/u/huananfshi)
#### Post date: [September 23, 2020, 6:14pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/15 "2020-09-23T18:14:51Z")

</div>

Hi there,  
We also got very low mapping rates using humann2 for rat gut microbiome. (with 95%+ Unaligned reads after nucleotide alignment and 90%+ Unaligned reads after translated alignment against uniref90). We have used kneaddata to filter out the host sequences.  
I wonder if humann 3.0 also has a better performance in rat as well.  
Thanks!

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [October 2, 2020, 12:11pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/16 "2020-10-02T12:11:14Z")

</div>

I have not tested HUMAnN 3 on a rat microbiome myself, but it’s generally expected to explain more reads under default settings. Give it a try and let us know what you find. 🙂

---

<div class="post-metadata">

### Author: ![huananfshi](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/huananfshi/32/415_2.png) [@huananfshi](https://forum.biobakery.org/u/huananfshi)
#### Post date: [October 5, 2020, 2:53pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/17 "2020-10-05T14:53:24Z")

</div>

We just tested one of our samples. It resulted 94% unaligned reads after nucleotide alignment (95% for humann2). We used Uniref50 database instead of Uniref90 with default setting. the unaligned reads after translatedd alignment decreased to 48%, which is okay I think? The only potential issue is that the taxonomy from metaphlan3 is quite different from metaphlan2 taxonomy, with some taxa no longer there and abundance changes.

---

<div class="post-metadata">

### Author: ![okeydokey](https://avatars.discourse-cdn.com/v4/letter/o/d2c977/32.png) [@okeydokey](https://forum.biobakery.org/u/okeydokey)
#### Post date: [January 26, 2021, 6:10pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/18 "2021-01-26T18:10:31Z")

</div>

I’ve done something similar with mouse microbiome samples (~90% reads unaligned after translated alignment Uniref90). Metaphlan 3 was big issue for me, where the bug list was very different than paired 16s composition data for the same sample (16s = 30% Firmicutes, Metaphlan 3 = 1.5% Firmicutes). So, I ran Humann3 bypassing metaphlan (so it uses the entire chocophlan database) and was able to get down to ~70% reads unaligned after translated search. However, because to chocophlan database only includes taxonomy to the genus level, after bypassing metaphlan I can’t get a count of my sample’s taxonomy at the phylum level to see how it compares to paired 16s data. The metagenome functional data from Humann3 looked great, but if I couldn’t trust the taxonomic composition I wasn’t sure how representative the data was.

I ended up switching to a basic pipeline from the iMGMC mouse microbiome database, which resulted in a similar proportion of reads map but the bug composition at the metagenome level is significantly closer to the paired 16s data we have (16s = 30% Firmicutes, iMGMC mapping = 32% Firmicutes). The tools natively provided by iMGMC are not nearly as robust as Humann3, but I think the basic data output is more representative of my samples. I never tried uniref50 in Humann3 because of how off the taxonomy was, but perhaps that would have helped?

---

<div class="post-metadata">

### Author: ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)
#### Post date: [January 29, 2021, 3:00pm UTC](https://forum.biobakery.org/t/low-percentage-of-aligned-reads/164/19 "2021-01-29T15:00:18Z")

</div>

UniRef90 vs UniRef50 won’t make much of a difference if you’re doing most of your mapping against pangenomes (it just changes the “resolution” of the output to broader rather than more focused pangenomes).

UniRef50 makes more of a difference if you’re focusing on translated search, since in that case it allows the actual read-vs-protein alignments to be a lot more relaxed.

Once we release the updated `infer_taxonomy` script you should be able to make a guess about the phylum-level recruitment of the sample from any of these mapping options. Apologies for the delay there.
