# Removing human transcripts with polyA from RNA data

**URL:** <https://forum.biobakery.org/t/removing-human-transcripts-with-polya-from-rna-data/2410>\
**Category:** KneadData\
**Created:** [July 22, 2021, 11:21pm UTC](https://forum.biobakery.org/t/removing-human-transcripts-with-polya-from-rna-data/2410 "2021-07-22T23:21:54Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jbarlow](https://avatars.discourse-cdn.com/v4/letter/j/94ad74/32.png) [@jbarlow](https://forum.biobakery.org/u/jbarlow)\
**Post date:** [July 22, 2021, 11:21pm UTC](https://forum.biobakery.org/t/removing-human-transcripts-with-polya-from-rna-data/2410/1 "2021-07-22T23:21:54Z")

</div>

Hello,

I have some metatranscriptome samples that have human contamination. I used kneaddata with the human transcriptome as the decontaminant database (–reference-db human\_hg38\_refMrna). After downstream processing with humann I found that a large % of the reads were unaligned so I looked at the first 30 reads or so and found the majority have a large stretch of polyA at the end of the sequence. The front half of these sequences blasts to human. Is there a good way to filter out these sequences with kneaddata?

---

<div class="post-metadata">

**Author:** ![sagunmaharjann](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/sagunmaharjann/32/168_2.png) [@sagunmaharjann](https://forum.biobakery.org/u/sagunmaharjann)\
**Post date:** [July 30, 2021, 2:08pm UTC](https://forum.biobakery.org/t/removing-human-transcripts-with-polya-from-rna-data/2410/2 "2021-07-30T14:08:03Z")

</div>

Hi @jbarlow ,

Thank you for reaching out to bioBakery Lab. Can you confirm that you are using Kneaddata’s human\_transcriptome reference database `kneaddata_database --download human_transcriptome bowtie2 $DIR` please?

You could also try adding the blast results to the database (in the .faa file then build the index) as contaminants and decoys to see if it improves the performance?

Regards,  
Sagun

---

<div class="post-metadata">

**Author:** ![jbarlow](https://avatars.discourse-cdn.com/v4/letter/j/94ad74/32.png) [@jbarlow](https://forum.biobakery.org/u/jbarlow)\
**Post date:** [August 5, 2021, 5:35am UTC](https://forum.biobakery.org/t/removing-human-transcripts-with-polya-from-rna-data/2410/3 "2021-08-05T05:35:22Z")

</div>

Hi @sagunmaharjann ,

Thanks for following up on this. I can definitely confirm I was using the human\_transcriptome reference database from kneaddata. I ended up realizing I wasn’t doing adapter trimming correctly (needed to change the default to Truseq) and updated to the kneaddata 0.10 from pip instead of 0.7.4 from conda and then no longer had the issue. Not sure exactly what fixed the problem but all is good now!

Best,  
Jacob
