# Some features missed by maaslin in untargeted metabolomics and other LM/CPLM issues

**URL:** <https://forum.biobakery.org/t/some-features-missed-by-maaslin-in-untargeted-metabolomics-and-other-lm-cplm-issues/6747>\
**Category:** MaAsLin\
**Created:** [March 1, 2024, 10:24pm UTC](https://forum.biobakery.org/t/some-features-missed-by-maaslin-in-untargeted-metabolomics-and-other-lm-cplm-issues/6747 "2024-03-01T22:24:28Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![siwook-hwang](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/siwook-hwang/32/2758_2.png) [@siwook-hwang](https://forum.biobakery.org/u/siwook-hwang)\
**Post date:** [March 1, 2024, 10:24pm UTC](https://forum.biobakery.org/t/some-features-missed-by-maaslin-in-untargeted-metabolomics-and-other-lm-cplm-issues/6747/1 "2024-03-01T22:24:28Z")

</div>

Hi there- thank you for creating this awesome tool.

I have untargeted metabolomics data that I am trying to work through.  
As is the nature of the dataset, it has a lot of zeros and roughly follows Tweedie distribution (and is continuous/ not count). My goal at this step is to filter out contaminants and internal standards by comparing my actual samples to method blanks (negative control). So I am running the following code:

Maaslin2(input\_data = df.eval.pos.SDA,  
input\_metadata = metabo.eval.metadata.pos.2,  
fixed\_effects = “type”,  
random\_effects = “block”,  
analysis\_method = “CPLM”,  
output = “metabo.pos.SDA\_output”,  
max\_significance = 0.05)

where “type” is a categorical variable with two levels: “sample” and “blanks”.

I have used both LM and CPLM. Both lead to some trouble-

LM works well for the most part, but also misses out on some obvious features that were identified by other methods:

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/biobakery/original/2X/5/53df1035b96e55fd010fe75a210e4e3a9ccbac76.png)  
As you can see “sample” has a lot more of this feature present while “blanks” have none. Maaslin passes this by while others (correctly) ID’s this feature as a real feature. I have been comparing with the results I get using a two part model: [SDA developed by Li et al 2018](https://bioconductor.org/packages/release/bioc/html/SDAMS.html).

Running CPLM, on the other hand results in all features having the equal pvalue of 1.

If you have any advice on how to approach this, I would appreciate it.
