# Humann3 chocophlan duplicate IDs

**URL:** <https://forum.biobakery.org/t/humann3-chocophlan-duplicate-ids/2482>\
**Category:** HUMAnN\
**Created:** [August 12, 2021, 8:51pm UTC](https://forum.biobakery.org/t/humann3-chocophlan-duplicate-ids/2482 "2021-08-12T20:51:48Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Billy\_Law](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/billy_law/32/85_2.png) [@Billy\_Law](https://forum.biobakery.org/u/Billy_Law)\
**Post date:** [August 12, 2021, 8:51pm UTC](https://forum.biobakery.org/t/humann3-chocophlan-duplicate-ids/2482/1 "2021-08-12T20:51:48Z")

</div>

Hi,  
I’m using the copy of chocophlan in my MT analysis pipeline, and I’m seeing an odd issue with the contents of chocophlan that came with HUMAnN3.

The issue is: There’s multiple copies of the same ID used in the entries.  
for example:  
655183\_\_B3WE58\_\_rpmF|k\_\_Bacteria.p\_\_Firmicutes.c\_\_Bacilli.o\_\_Lactobacillales.f\_\_Lactobacillaceae.g\_\_Lactobacillus.s\_\_Lactobacillus\_casei\_group|UniRef90\_B3WE58|UniRef50\_Q7C3P5|192les

This ID is used 3 times, with 3 different sequences.  
The produces a warning of duplicate IDs when I use samtools to parse my BWA run.

It was my understanding that we would get 1 unique ID per sequence. Is this no long the case?

---

<div class="post-metadata">

**Author:** ![franzosa](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/franzosa/32/3511_2.png) [@franzosa](https://forum.biobakery.org/u/franzosa)\
**Post date:** [August 16, 2021, 8:34pm UTC](https://forum.biobakery.org/t/humann3-chocophlan-duplicate-ids/2482/2 "2021-08-16T20:34:55Z")

</div>

There should be one gene sequence per UniRef90 per species. We may have ended up with duplicates here because this is a species group (a merging of independently defined species pangenomes). Any reads assigned to any of those sequences would be grouped by HUMAnN into the appropriate UniRef families.

If the non-unique names are causing issues outside of HUMAnN, you could modify the sequence headers to contain something unique, e.g. a prefix indicating the position of the sequence in the file, like `000123-`. Do not include this as a separate `|`ed field since that would throw off HUMAnN’s indexing of other information in the header.

---

<div class="post-metadata">

**Author:** ![Billy\_Law](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.biobakery.org/billy_law/32/85_2.png) [@Billy\_Law](https://forum.biobakery.org/u/Billy_Law)\
**Post date:** [August 16, 2021, 8:47pm UTC](https://forum.biobakery.org/t/humann3-chocophlan-duplicate-ids/2482/3 "2021-08-16T20:47:33Z")

</div>

I’ll give this a try.  
Thanks for the explanation, Eric!
