# Discrepancy in the number of non-zero samples in input file and output for Maaslin2

**URL:** <https://forum.biobakery.org/t/discrepancy-in-the-number-of-non-zero-samples-in-input-file-and-output-for-maaslin2/3267>\
**Category:** MaAsLin\
**Created:** [March 5, 2022, 12:27pm UTC](https://forum.biobakery.org/t/discrepancy-in-the-number-of-non-zero-samples-in-input-file-and-output-for-maaslin2/3267 "2022-03-05T12:27:33Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Claire](https://avatars.discourse-cdn.com/v4/letter/c/f6c823/32.png) [@Claire](https://forum.biobakery.org/u/Claire)\
**Post date:** [March 5, 2022, 12:27pm UTC](https://forum.biobakery.org/t/discrepancy-in-the-number-of-non-zero-samples-in-input-file-and-output-for-maaslin2/3267/1 "2022-03-05T12:27:33Z")

</div>

Hi all,

I’m using Maaslin2 to analyze the output from Metaphlan3. Below is the relative abundance table of the species and metadata.

[Species\_relab.txt](https://forum.biobakery.org/uploads/short-url/oUkl7551gg8xLYF2awTOztBUGg0.txt) (95.9 KB)

[Metadata.txt](https://forum.biobakery.org/uploads/short-url/h98iJ6Buda8VlhcAOnC3OnFyTXD.txt) (637 Bytes)

Then in R,

> relab ← read.table(“Species\_relab.txt”, header = TRUE, quote = “”, sep = “\t”, row.names = 1, stringsAsFactors = FALSE)
> 
> metadata ← read.table(“Metadata.txt”, header=TRUE, sep=“\t”, row.names=1, stringsAsFactors=FALSE)
> 
> fit\_data ← Maaslin2(  
> relab, metadata,‘output’, transform = “AST”,  
> fixed\_effects = c(‘Treatment’),  
> normalization = ‘NONE’,  
> standardize = FALSE)

Output for all the features,  
[all\_results.tsv](https://forum.biobakery.org/uploads/short-url/8dY7VYzczDbY1BoCUPcvZTGLMm0.tsv) (94.8 KB)

Taking Faecalibacterium prausnitzii as an example, the number of non-zero sample is only **6**. However, if checking the input Species\_relab.txt, the actual count of non-zero sample is **52**.

| metadata | feature | value | coef | stderr | N | N.not.0 | pval | qval |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Treatment | Faecalibacterium\_prausnitzii | B | -0.511171192 | 0.114657358 | 52 | **6** | 0.021009755 | 0.802131021 |
| Treatment | Faecalibacterium\_prausnitzii | C | -0.380650407 | 0.14503136 | 52 | **6** | 0.078688933 | 0.802131021 |

After checking all the 450 species, actually there are 126 discrepancies.  
I would like to ask why there is a discrepancy in the count of non-zero sample between the actual input file and output file? Any ideas would be highly appreciated.

Thank you!  
Claire

---

<div class="post-metadata">

**Author:** ![Kelsey\_Thompson](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/kelsey_thompson/32/65_2.png) [@Kelsey\_Thompson](https://forum.biobakery.org/u/Kelsey_Thompson)\
**Post date:** [March 9, 2022, 4:18pm UTC](https://forum.biobakery.org/t/discrepancy-in-the-number-of-non-zero-samples-in-input-file-and-output-for-maaslin2/3267/2 "2022-03-09T16:18:08Z")

</div>

Hi @Claire ,

The issue here is that you are on the scale 0-100 (relative abundance) for AST transformation the data needs to be in the 0-1 scale. So currently when MaAsLin AST transforms the data points in 0-1 are non-zero “transformed”, but anything above 1 is converting to a NaN. You should see that there were a lot of warnings after MaAsLin runs. We are currently working on getting MaAsLin to throw an error instead of letting AST transformation run when the underlying data isn’t in the 0-1 scale. Switching to a log transformation or converting your data frame to 0-1 - should solve the issues you are seeing.

Sorry for the confusion - I hope this helps!

Best,  
Kelsey

---

<div class="post-metadata">

**Author:** ![Bruno](https://avatars.discourse-cdn.com/v4/letter/b/c77e96/32.png) [@Bruno](https://forum.biobakery.org/u/Bruno)\
**Post date:** [September 2, 2022, 2:39am UTC](https://forum.biobakery.org/t/discrepancy-in-the-number-of-non-zero-samples-in-input-file-and-output-for-maaslin2/3267/3 "2022-09-02T02:39:09Z")

</div>

Hello,

If I’m not mistaken, entering relative abundances with very low values (e.g., 0.0003) cannot be detected by the model and counted in the results column N.not.zero, even if min\_abundance = 0, min\_prevalence = 0 and transform the data with AST (input\_data: scale 0-1). However, I believe that by introducing counts the model can account for all of them correctly.

Do you consider it acceptable to run the analysis with counts and then plot the raw data with relative abundance (%) so as not to lose information?

Thank you very much for your work.

All the best,  
Bruno
