# Which reference DB should I use?

**URL:** <https://forum.biobakery.org/t/which-reference-db-should-i-use/3798>\
**Category:** KneadData\
**Created:** [July 4, 2022, 11:34am UTC](https://forum.biobakery.org/t/which-reference-db-should-i-use/3798 "2022-07-04T11:34:38Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kyoungmin\_Lee](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/kyoungmin_lee/32/1550_2.png) [@Kyoungmin\_Lee](https://forum.biobakery.org/u/Kyoungmin_Lee)\
**Post date:** [July 4, 2022, 11:34am UTC](https://forum.biobakery.org/t/which-reference-db-should-i-use/3798/1 "2022-07-04T11:34:38Z")

</div>

Hello everyone! I am a newbie for microbiome analysis. So I tried to practice general Metatranscriptomics  
analysis using the pipeline (1.Kneaddata, 2. Metaphlan3/HUMAnN/ or possible Kraken). But after I used Kneaddata, its output is weird. What have I missed?  
First of all, I downloaded SRR769427 using faster-dump. More detailed information is belowed  
 ![image](https://canada1.discourse-cdn.com/flex027/uploads/biobakery/original/2X/6/64243b3c7ee83cbd3751146627bceb3f48c2d7dd.png)  
Since it is human gut metatranscriptome,1) I downloaded DB using [$ kneaddata\_database --download human\_transcriptome bowtie2 $DIR]  
then 2) [kneaddata -i SRR769427\_1.fastq --i SRR769427\_2.fastq --reference-db $DATABASE --output $OUTPUT\_DIR --trimmomatic $TRIM\_DIR]  
then my outputs are

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/biobakery/original/2X/d/da3b7abb85de7cc43250afc9fd9fca94491de413.jpeg)  
The Kneaddata menual for pairends data says SRR\*\*\*.kneaddata\_paired\_1.fastq and SRR\*\*\*.kneaddata\_paired\_2.fastq are final outputs. But volume of mine are zero.  
Could you help me out what have I done wrong?  
My question is

1. Since the files that I want to analyze are human transcripomic data, shoud I use human transcriptome data as a DB in bowtie2 for removing human mRNA for metatranscriptome analysis ?

Many thankx in advance
