# Question about ribosomal\_RNA database

**URL:** <https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634>\
**Category:** KneadData\
**Created:** [June 25, 2020, 7:14pm UTC](https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634 "2020-06-25T19:14:39Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![helicam](https://avatars.discourse-cdn.com/v4/letter/h/6de8d8/32.png) [@helicam](https://forum.biobakery.org/u/helicam)\
**Post date:** [June 25, 2020, 7:14pm UTC](https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634/1 "2020-06-25T19:14:39Z")

</div>

Hello!  
I downloaded the SILVA ribosomal RNA database from your website (downloaded directly because kneaddata\_database failed for me), and I have a question. Looking at the index with bowtie2-inspect, it looks like the reference sequences contain only A/C/G and that the U’s present in the SILVA fasta files have been removed. Building a custom database using bowtie2-build with the downloaded FASTA files has the same problem (is it best to just convert the U’s to T’s in the SILVA fasta file?). My apologies if I’m missing something obvious - this is my first analysis with kneaddata, and my first analysis where I actually care about characterizing rRNA in an experiment.

In any case, thanks for developing this very useful tool!  
Michael

PS - I actually posted this question at [https://bitbucket.org/biobakery/kneaddata/issues](https://bitbucket.org/biobakery/kneaddata/issues), but I assume this is a better spot…

---

<div class="post-metadata">

**Author:** ![george-weingart](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/george-weingart/32/256_2.png) [@george-weingart](https://forum.biobakery.org/u/george-weingart)\
**Post date:** [June 30, 2020, 1:01am UTC](https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634/2 "2020-06-30T01:01:29Z")

</div>

Hello Michael !  
I looked at the “$KNEADDATA\_DB\_RIBOSOMAL\_RNA” database, looked at the fasta records used to create the Bowtie2 index and I see lots of "U"s, so can you elaborate?  
I am enclosing a screen print.  
Let me know…  
Best regards,  
George Weingart PhD  
Huttenhower Lab

 ![fasta_SILVA__rna_with_Us](https://canada1.discourse-cdn.com/flex027/uploads/biobakery/original/1X/2119f8c1be6ed87ac85f0869958bf6b0425b8e9d.png)

---

<div class="post-metadata">

**Author:** ![helicam](https://avatars.discourse-cdn.com/v4/letter/h/6de8d8/32.png) [@helicam](https://forum.biobakery.org/u/helicam)\
**Post date:** [July 2, 2020, 7:22am UTC](https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634/3 "2020-07-02T07:22:22Z")

</div>

Hi George,  
Thanks for your response. I agree that the original FASTA file for SILVA contains many U’s. The bowtie2 distribution includes an executable “bowtie2-inspect” which takes a Bowtie2 index an prints out the underlying FASTA sequence of the index. Running it on the pre-built kneaddata index (SILVA\_128\_LSUParc\_SSUParc\_ribosomal\_RNA.1.bt2l, etc.) yields the following:

% bowtie2-inspect SILVA\_128\_LSUParc\_SSUParc\_ribosomal\_RNA | head -6

> GCVF01000057.1.1978 Eukaryota;Amoebozoa;Thalassiosira rotula  
> CGAAAGACAACGAACCACAGAGCGGCCCCCGAAGCCCCAGGAAGCGGAACAAAAAAGACA  
> GGAAAGCGAAGAAGAAGCCGGGGAGAAACACCCGACACCAAACAAAAGGAAGAAGGAAAG  
> CAAAGAAACAGAAAACAAGCAGGGGCCAGGAAGCAGAACGGCGAGGGGAGAACCAAACGG  
> AAGAAAGGCCAAAAGACGCACAGAACCACAAAAGGGGACACAGACAGGGACGGGGCCAGG  
> AAGCGGAACCGCAAGGAGGGAACAACCACCAACCGAAGAAAGCCCGAAAAGGAGGCGCAA

You can see that this is precisely the start of the nucleotide sequence for GCVF01000057.1.1978, with all of the U’s omitted. So it isn’t clear to me that the index is correct for use in aligning our data. Perhaps it is a quirk of bowtie2-inspect, but it certainly seems odd.  
Best,  
Michael

---

<div class="post-metadata">

**Author:** ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)\
**Post date:** [July 9, 2020, 1:18pm UTC](https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634/4 "2020-07-09T13:18:07Z")

</div>

Thanks for pointing this out, @helicam. We confirmed in a recent evaluation that bowtie2 is quietly dropping the U characters from the SILVA sequences during index construction rather than treating them as T-equivalents (an unusual choice relative to other nucleotide-level aligners). We’ve now corrected this issue by manually replacing the Us with Ts prior to indexing, which substantially improves the performance of the SILVA database.

The updated database file is [available here](http://huttenhower.sph.harvard.edu/kneadData_databases/SILVA_128_LSUParc_SSUParc_ribosomal_RNA_v0.2.tar.gz). You can automatically update your local database using this command:

```auto
kneaddata_database --download ribosomal_RNA bowtie2 $DIR

```

I’ll pin a separate message about this to the KneadData board so that others users are aware of this important fix.

---

<div class="post-metadata">

**Author:** ![Duannai](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/duannai/32/1149_2.png) [@Duannai](https://forum.biobakery.org/u/Duannai)\
**Post date:** [December 11, 2021, 5:13am UTC](https://forum.biobakery.org/t/question-about-ribosomal-rna-database/634/5 "2021-12-11T05:13:19Z")

</div>

Sorry for re-opening this question.

Well, before indexing SILVA database, the Us have been replaced by Ts, as described in this link [Creating SILVA ribosomal\_RNA Database](https://github.com/biobakery/kneaddata#creating-a-bowtie2-database).

However, if the data is from metatranscriptomics, as we know RNA will be first converted to cDNA (may be strand-specific), and then be sequenced. Should the reference sequences in SILVA database be reverse complented first (AAUUCCGG→CCGGAATT) or still be replaced by Ts (AAUUCCGG → AATTCCGG)?

Thanks in advance!

Mort
