# Interpretation of pathway abundances in light of pathways hierarchy

**URL:** <https://forum.biobakery.org/t/interpretation-of-pathway-abundances-in-light-of-pathways-hierarchy/2559>\
**Category:** HUMAnN\
**Created:** [September 4, 2021, 4:54pm UTC](https://forum.biobakery.org/t/interpretation-of-pathway-abundances-in-light-of-pathways-hierarchy/2559 "2021-09-04T16:54:48Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![efratmuller](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/efratmuller/32/998_2.png) [@efratmuller](https://forum.biobakery.org/u/efratmuller)\
**Post date:** [September 4, 2021, 4:54pm UTC](https://forum.biobakery.org/t/interpretation-of-pathway-abundances-in-light-of-pathways-hierarchy/2559/1 "2021-09-04T16:54:48Z")

</div>

Hi!  
I was wondering how should pathway abundances be interpreted when some of the pathways are nested in others?  
As an example, in an analysis I’m running I’ve got the following MetaCyc pathway abundances:  
_OANTIGEN-PWY … 1263.8116_  
_DTDPRHAMSYN-PWY … 2243.2606_  
_UDPNAGSYN-PWY … 962.1098_  
Now, _OANTIGEN-PWY_ is a super-pathway that is composed of the other two, _DTDPRHAMSYN-PWY_ and _UDPNAGSYN-PWY_. In the documentation (and [forum answers](https://forum.biobakery.org/t/pathway-abundance-and-cross-samples-comparison/2057)) it is recommended to turn abundances into relative abundances (i.e. divide by sample total). But unlike gene abundances or species abundances, here - these pathways aren’t independent entities, like in my example, no? Wouldn’t this kind of normalization distort comparisons between samples? Is there any specific recommendation about how to deal with these cases? (e.g. drop super-pathways somehow?)  
Many many thanks,  
Efrat

---

<div class="post-metadata">

**Author:** ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)\
**Post date:** [September 10, 2021, 5:57pm UTC](https://forum.biobakery.org/t/interpretation-of-pathway-abundances-in-light-of-pathways-hierarchy/2559/2 "2021-09-10T17:57:12Z")

</div>

In practice normalizing the pathways in this way seems to work OK. The resulting fractions are still proportional to the pathways’ copy numbers in the original sample, and the normalization corrects for differences in sequencing depth across samples.

Another approach is to normalize the gene family abundances (where no read mass is double-counted) from RPKs to CPMs and then run the normalized genes back through HUMAnN to directly compute pathway abundance in CPMs. The resulting pathway abundances will NOT sum to 1M in that case, but they will be corrected for sequencing depth, and this avoids any artifacts that might arise from sum-normalizing over overlapping pathways. (Recomputing pathways by providing gene family abundances as an input file is very fast.)
