# How to get count data (not RPK)?

**URL:** <https://forum.biobakery.org/t/how-to-get-count-data-not-rpk/8429>\
**Category:** HUMAnN\
**Created:** [August 25, 2025, 5:40am UTC](https://forum.biobakery.org/t/how-to-get-count-data-not-rpk/8429 "2025-08-25T05:40:57Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![bowornpol](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/bowornpol/32/3364_2.png) [@bowornpol](https://forum.biobakery.org/u/bowornpol)\
**Post date:** [August 25, 2025, 5:40am UTC](https://forum.biobakery.org/t/how-to-get-count-data-not-rpk/8429/1 "2025-08-25T05:40:57Z")

</div>

Hi,

I am using **HUMAnN v3.9** and have regrouped gene families to reactions using `regroup_table`. From my understanding, the output is in **RPK units** (Reads Per Kilobase).

I would like to perform **differential abundance analysis** with tools such as **DESeq2** or **edgeR** , but these methods require **raw count data** , not normalized units like RPK.

- Is there a way to obtain count data from HUMAnN v3.9, or convert RPK values back to counts?

- Do you have recommendations for approaches to use with HUMAnN v3.9 output?

- I also noticed that **HUMAnN v4 (alpha)** can produce count data—would you suggest using that instead if I want to apply DESeq2/edgeR?

Thanks for your help!

---

<div class="post-metadata">

**Author:** ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)\
**Post date:** [August 27, 2025, 7:12pm UTC](https://forum.biobakery.org/t/how-to-get-count-data-not-rpk/8429/2 "2025-08-27T19:12:47Z")

</div>

I think if you search the forum you’ll find some discussion on methods to back-calculate approximate counts from other forms of abundance data, but I wouldn’t really recommend doing that. Methods that want counts want TRUE counts, so you’re potentially misusing them (or at least not using them to their full potential) by approximating something that looks like a count.

We typically work with relative abundance units (e.g. RPKs sum-normalized to CPMs) and then analyzing them using linear models in MaAsLin as opposed to working with count-based models.

HUMAnN 4 _can_ output raw counts using a non-default normalization mode. We added this feature because it’s something that gets requested all the time, but it’s not a feature we use internally.

---

<div class="post-metadata">

**Author:** ![mguaita](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/mguaita/32/3668_2.png) [@mguaita](https://forum.biobakery.org/u/mguaita)\
**Post date:** [June 22, 2026, 5:45pm UTC](https://forum.biobakery.org/t/how-to-get-count-data-not-rpk/8429/3 "2026-06-22T17:45:34Z")

</div>

Hello, I have a question regarding this same topic. Humann v4 (alpha) returns decimals when using the parameter --count-normalization Counts. Is there a recommended way to transform this pseudo-counts to integer data to be used in count-based statistical models like ALDEx2? Or is it better to stick to adjustedCPMs for further downstream analysis?

Turns out MaAsLin is not an option for my data design.

Thank you,
