# Normalising the input reads of the samples

**URL:** <https://forum.biobakery.org/t/normalising-the-input-reads-of-the-samples/2061>\
**Category:** MetaPhlAn\
**Created:** [May 12, 2021, 6:22am UTC](https://forum.biobakery.org/t/normalising-the-input-reads-of-the-samples/2061 "2021-05-12T06:22:02Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![saras22](https://avatars.discourse-cdn.com/v4/letter/s/59ef9b/32.png) [@saras22](https://forum.biobakery.org/u/saras22)\
**Post date:** [May 12, 2021, 6:22am UTC](https://forum.biobakery.org/t/normalising-the-input-reads-of-the-samples/2061/1 "2021-05-12T06:22:02Z")

</div>

Hi Biobakery\_forum  
I have analyzed my shotgun metagenomic datasets(.fastq.gz) files of varying sizes with MetaPhlAn3. So, I have got different number of clades for different files. Now I my doubt is that are the read files normalized before analysis or I have to do it before analyzing? Shall I make all the files of same size?  
Also mention any method that you would recommend for normalizing the total reads for all the samples?  
for e.g I found the file having 10.3 Mbp total yield is giving 542 clades and file having 1.2 Mbp is giving 402 clades. Is it there any threshold after which the total yield would not matter?

Thanks in Advance  
Saraswati

---

<div class="post-metadata">

**Author:** ![fbeghini](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/fbeghini/32/81_2.png) [@fbeghini](https://forum.biobakery.org/u/fbeghini)\
**Post date:** [May 17, 2021, 9:09am UTC](https://forum.biobakery.org/t/normalising-the-input-reads-of-the-samples/2061/2 "2021-05-17T09:09:52Z")

</div>

You could try having a look at the profiles of the rarefied 10,3 M metagenome but if you run the analysis with `--unknown_estimation` you can obtain the profiles normalized by the total metagenome size.
