# Humann output normalized by Metaphlan output

**URL:** <https://forum.biobakery.org/t/humann-output-normalized-by-metaphlan-output/6275>\
**Category:** HUMAnN\
**Created:** [December 1, 2023, 3:29pm UTC](https://forum.biobakery.org/t/humann-output-normalized-by-metaphlan-output/6275 "2023-12-01T15:29:33Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jeremy\_Tournayre](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/jeremy_tournayre/32/2088_2.png) [@Jeremy\_Tournayre](https://forum.biobakery.org/u/Jeremy_Tournayre)\
**Post date:** [December 1, 2023, 3:29pm UTC](https://forum.biobakery.org/t/humann-output-normalized-by-metaphlan-output/6275/1 "2023-12-01T15:29:33Z")

</div>

Hello,

Thank you for offering this tool!  
I have a question about the normalization of the Humann results by the Metaphlan results in case of paired metagenome/metatranscriptome.

In this documentation:

> **[GitHub - biobakery/humann: HUMAnN is the next generation of HUMAnN 1.0 (HMP...](https://github.com/biobakery/humann#workflows)**
>
> HUMAnN is the next generation of HUMAnN 1.0 (HMP Unified Metabolic Analysis Network). - GitHub - biobakery/humann: HUMAnN is the next generation of HUMAnN 1.0 (HMP Unified Metabolic Analysis Network).

It is written:

> **Analyzing a metatranscriptome with a paired metagenome.**
> 
> HUMAnN RNA-level outputs (e.g. transcript family abundance) **can** then be normalized by corresponding DNA-level outputs to quantify microbial expression independent of gene copy number.

The word “can” means that the data are not normalized, right?  
If I want to do this, what’s the best way?

I see that the pipeline “ASAIM MT” [https://f1000research.com/articles/10-103](https://f1000research.com/articles/10-103) have a tool to combine MetaPhlAn2 and HUMAnN2 Outputs :

> “This produces a table of functional terms and their abundances with the corresponding genus and species abundances for the taxa which contribute to said function via their expressed RNA sequences.”

More information on this in their tutorial here:

> **[Galaxy Training: Metatranscriptomics analysis using microbiome RNA-seq data](https://training.galaxyproject.org/topics/metagenomics/tutorials/metatranscriptomics/tutorial.html#combine-taxonomic-and-functional-information)**
>
> Metagenomics is a discipline that enables the genomic study of uncultured microorganisms

Sincerly,  
Jérémy Tournayre

---

<div class="post-metadata">

**Author:** ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)\
**Post date:** [December 8, 2023, 7:41pm UTC](https://forum.biobakery.org/t/humann-output-normalized-by-metaphlan-output/6275/2 "2023-12-08T19:41:43Z")

</div>

Our current recommendation for doing this is residualizing the RNA abundance against the DNA abundance. You can do this as part of a single model:

`RNA ~ DNA + treatment` (for example)

Or you can run `RNA ~ DNA` and then save the residual expression values, which will then be corrected for DNA copy number. We have a paper discussing this approach:

> **[Statistical approaches for differential expression analysis in...](https://pubmed.ncbi.nlm.nih.gov/34252963/)**
>
> Supplementary data are available at Bioinformatics online.

And some additional software links from here:

[http://huttenhower.sph.harvard.edu/mtx2021](http://huttenhower.sph.harvard.edu/mtx2021)

The paper+software also includes some recommendations for this problem if you DON’T have paired DNA data.
