# Difference in sequencing depth

**URL:** <https://forum.biobakery.org/t/difference-in-sequencing-depth/8757>\
**Category:** MetaPhlAn\
**Created:** [February 5, 2026, 10:03am UTC](https://forum.biobakery.org/t/difference-in-sequencing-depth/8757 "2026-02-05T10:03:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![KSK](https://avatars.discourse-cdn.com/v4/letter/k/ac8455/32.png) [@KSK](https://forum.biobakery.org/u/KSK)\
**Post date:** [February 5, 2026, 10:03am UTC](https://forum.biobakery.org/t/difference-in-sequencing-depth/8757/1 "2026-02-05T10:03:21Z")

</div>

I have two batch of shotgun metagenomic sequences. Batch one has 20 samples with avg depth of 10 million sequences and batch 2 has 50 samples with avg depth of 25 million sequences. Will this difference in the sequencing depth cause a issue when I merge all the samples and do a diversity analysis (case vs control). Further is there a option like rarefying the sequences before diversity analysis like its there in qiime2?

Thank you

---

<div class="post-metadata">

**Author:** ![KSK](https://avatars.discourse-cdn.com/v4/letter/k/ac8455/32.png) [@KSK](https://forum.biobakery.org/u/KSK)\
**Post date:** [April 9, 2026, 8:43am UTC](https://forum.biobakery.org/t/difference-in-sequencing-depth/8757/2 "2026-04-09T08:43:11Z")

</div>

Kindly requesting some insights ..

---

<div class="post-metadata">

**Author:** ![Claudia\_Mengoni](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/claudia_mengoni/32/2871_2.png) [@Claudia\_Mengoni](https://forum.biobakery.org/u/Claudia_Mengoni)\
**Post date:** [April 9, 2026, 9:01am UTC](https://forum.biobakery.org/t/difference-in-sequencing-depth/8757/3 "2026-04-09T09:01:31Z")

</div>

Hi @KSK  
yes, sequencing depth is a technical confounder, to make sure you’re not introducing any bias you can use the subsampling options. This is extracted from the [documentation](https://github.com/biobakery/MetaPhlAn/wiki/MetaPhlAn-4.2) for metaphlan v4.2.\* : It is possible to subsample the reads before the MetaPhlAn run by passing the number of reads to use (which must be \< than the total number of reads of the sample) to `--subsampling`. In the following example, subsampling to 10,000 reads:

```auto
$ metaphlan metagenome.fastq --input_type fastq --subsampling 10000 -o profiled_metagenome_subsampled_10000.txt

```

Since MetaPhlAn 4.1.1, it is possible to use paired-end information during subsampling (above, paired-end reads would be treated as single-end, i.e., independent). For that, use `--subsampling_paired` instead:

```auto
metaphlan --subsampling_paired <N_PAIRED_READS> -1 <R1_FASTQ> -2 <R2_FASTQ> --input_type fastq --subsampling_out <SUBSAMPLED_READS_OUTPUT> -o <METAPHLAN_OUTPUT> --mapout <MAPOUT>

```

---

<div class="post-metadata">

**Author:** ![KSK](https://avatars.discourse-cdn.com/v4/letter/k/ac8455/32.png) [@KSK](https://forum.biobakery.org/u/KSK)\
**Post date:** [May 25, 2026, 5:21am UTC](https://forum.biobakery.org/t/difference-in-sequencing-depth/8757/4 "2026-05-25T05:21:15Z")

</div>

Hi, a follow up question: if I specify --subsampling 10000 in my command, does this means it will subsample 5000 reads from R1 and 5000 reads from R2? or it 10000 reads from R1 and 10000 reads from R2?

Thank you
