# HUMAnN 4 outputs and best practices for paired designs

**URL:** <https://forum.biobakery.org/t/humann-4-outputs-and-best-practices-for-paired-designs/8979>\
**Category:** MaAsLin\
**Created:** [June 22, 2026, 9:28am UTC](https://forum.biobakery.org/t/humann-4-outputs-and-best-practices-for-paired-designs/8979 "2026-06-22T09:28:32Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mguaita](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/mguaita/32/3668_2.png) [@mguaita](https://forum.biobakery.org/u/mguaita)\
**Post date:** [June 22, 2026, 9:28am UTC](https://forum.biobakery.org/t/humann-4-outputs-and-best-practices-for-paired-designs/8979/1 "2026-06-22T09:28:32Z")

</div>

Hello everyone,

I am currently working with HUMAnN 4 outputs to infer differential pathway abundances between two groups in a paired experimental design. Any guidance or standardized workflow for this specific scenario would be highly appreciated.

Which humann4 output is more appropriate for differential abundance testing? It is better to use Humann4 Raw Counts, that can have decimals and do break algorithms requiring strictly integer count data, or adjustedCPMs?

Which statistical method would you recommend for this setting? MaasLin3 raises an estimation error as there are less than 4 data points for each sample.

Thank you in advance for your time and your great tools!

---

<div class="post-metadata">

**Author:** ![WillNickols](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/willnickols/32/3223_2.png) [@WillNickols](https://forum.biobakery.org/u/WillNickols)\
**Post date:** [June 23, 2026, 2:49pm UTC](https://forum.biobakery.org/t/humann-4-outputs-and-best-practices-for-paired-designs/8979/2 "2026-06-23T14:49:13Z")

</div>

Hi,

I’d use the adjusted CPMs since those account for gene length and sequencing depth. If you then use the default log scaling and TSS normalization (though you’ll get the same results with and without TSS if you’re using CPM), everything should work fine.

Regarding the actual formula, can you provide a bit more information about what the experimental design is? Is it some sort of before/after design with 2 measurements per-person?

Will

---

<div class="post-metadata">

**Author:** ![mguaita](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/mguaita/32/3668_2.png) [@mguaita](https://forum.biobakery.org/u/mguaita)\
**Post date:** [June 25, 2026, 12:26pm UTC](https://forum.biobakery.org/t/humann-4-outputs-and-best-practices-for-paired-designs/8979/3 "2026-06-25T12:26:09Z")

</div>

> [@WillNickols](#):
>
> adjusted CPMs s

Hi,

Thank you for your reply! Yes, indeed it is a before/after treatment with 2 measurement per-person, that is just one measurement before and one measurement after.

Maria

---

<div class="post-metadata">

**Author:** ![WillNickols](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/willnickols/32/3223_2.png) [@WillNickols](https://forum.biobakery.org/u/WillNickols)\
**Post date:** [June 25, 2026, 3:05pm UTC](https://forum.biobakery.org/t/humann-4-outputs-and-best-practices-for-paired-designs/8979/4 "2026-06-25T15:05:12Z")

</div>

In that case, I’d use the formula `~ time + (1|subject)` (plus whatever else you care about) with the `small_random_effects=TRUE` parameter set. This’ll properly account for repeated subject sampling despite having only 2 samples per subject which would normally cause issues in random effects.
