Pathcoverage vs Pathabundance for filtering - some numbers where I don't understand how they can happen

Hi all,

I am quite new to working with HUMAnN-generated data; therefore my question: I wanted to do some filtering before further analysis of pathway data, to reduce the number of pathways to the ones that are most probably really present and in a minimum part of my cohort. For this, I thought to use the pathcoverage file, filtering all total pathways with a coverage > 0.1 in at least 10% of the samples. When I was looking into some specific pathways that were actually all filtered out by this step, I noticed something I do not understand. For one pathway (PWY-4984: urea cycle), for example, the coverage on the total pathway level is 0.0000 for most of the samples. This is also the case in some of the samples, where the pathcoverage on the species level has numbers like 0.6 (partly just for one species, partly for more). For the same pathway, the pathabundance file shows numbers like 4285.11092 or even higher on the total pathways level in most of the samples.

When reading through the forum, I found that some people experience something similar, and in answer, that the pathcoverage came from the first version and might not be accurate to use anymore. So my question is: Should I just ignore the pathcoverage file and filter directly on the pathabundance file? Or is it still better to filter on the pathcoverage file?

I also tried going down to passing all values >0 in at least 10% for total pathway coverage, but there is no big difference; instead of keeping 104 of 522 pathways, it’s then 109.

My reasoning was to filter only on the total pathway level and then apply the selected pathways also to the pathways at the species level.

Thank you in advance for your help!