# Detecting identical or very similar samples in corpus

**URL:** https://discourse.flucoma.org/t/detecting-identical-or-very-similar-samples-in-corpus/2125
**Category:** Usage Questions
**Created:** [October 15, 2023, 2:34pm UTC](https://discourse.flucoma.org/t/detecting-identical-or-very-similar-samples-in-corpus/2125 "2023-10-15T14:34:34Z")
**Posts on this page:** 1
**Showing post:** 2

<div class="post-metadata">

### Author: ![rodrigo.constanzo](https://discourse.flucoma.org/user_avatar/discourse.flucoma.org/rodrigo.constanzo/32/12_2.png) [@rodrigo.constanzo](https://discourse.flucoma.org/u/rodrigo.constanzo)
#### Post date: [October 17, 2023, 9:42am UTC](https://discourse.flucoma.org/t/detecting-identical-or-very-similar-samples-in-corpus/2125/2 "2023-10-17T09:42:41Z")

</div>

@jamesbradbury work with FTIS may be of interest here:

[![](https://img.youtube.com/vi/IpD_XzW1Az4/maxresdefault.jpg "FluCoMa Plenary: James Bradbury, Finding Things In Stuff") ](https://www.youtube.com/watch?v=IpD_XzW1Az4)

[https://phd.jamesbradbury.net/tech/ftis/](https://phd.jamesbradbury.net/tech/ftis/)

This as well:  
[https://discourse.flucoma.org/t/segmentation-by-clustering](https://discourse.flucoma.org/t/segmentation-by-clustering)

Ultimately you’ll have to decide what criteria (and thresholds) you use for similarity, and I would imagine morphology/duration/time series being important factors especially if you want to remove sounds that sound identical.

When you say the variations in duration, do you mean that there may be a short file that sounds identical to a long file or is that just the lay-of-the-land in the corpus and you will be comparing similar duration files regardless?

---

_[View the full topic](https://discourse.flucoma.org/t/detecting-identical-or-very-similar-samples-in-corpus/2125)._
