# AudioGuide (talk)

**URL:** https://discourse.flucoma.org/t/audioguide-talk/550
**Category:** Interesting Links
**Created:** [July 6, 2020, 2:25pm UTC](https://discourse.flucoma.org/t/audioguide-talk/550 "2020-07-06T14:25:41Z")
**Posts on this page:** 1
**Showing post:** 3

<div class="post-metadata">

### Author: ![rodrigo.constanzo](https://discourse.flucoma.org/user_avatar/discourse.flucoma.org/rodrigo.constanzo/32/12_2.png) [@rodrigo.constanzo](https://discourse.flucoma.org/u/rodrigo.constanzo)
#### Post date: [July 7, 2020, 7:00pm UTC](https://discourse.flucoma.org/t/audioguide-talk/550/3 "2020-07-07T19:00:08Z")

</div>

First, thanks for the super detailed and thoughtful response!

It was interesting watching the video and hearing the nuts and bolts of your specific take and perspective on this, as it varies quite a bit from the (current) FluCoMa paradigm.

Thanks for the additional comments on the normalization stuff. I’m still getting my head around this aspect of things as it can get complex, particularly when MFCCs are in the mix.

> [@b.hackbarth](#):
>
> …the most important choice is whether to standardize the corpus and target’s descriptors together or separately.

That’s quite interesting.

I guess this makes the most sense in a “one off” context where you have a fixed target and a set corpus, since you can just normalize it as part of the query, but I wonder how this would fair with a stream of targets pouring in a real-time context, re-normalizing on a _per query_ basis.

> [@b.hackbarth](#):
>
> The thing that I think I like best about this approach is that is it feels creatively purposeful. Rather than asking for the best match on 40 dimensions, which tends to be impenetrable to the user (ditto for dimensional scaling), you dictate what you want and the order that you want those measurements to be considered. In my work I’ve found that there is no gold standard for measuring similarity, only what you’re interested in.

This is one of the toughest things to wrap my head around when dipping into the machine-learning side of things is that _penetrability_ evaporates almost instantly. Not a big deal when dealing with things like MFCCs or a high dimensional space, but there are still individual numbers (i.e. duration, loudness, etc…) that probably still _mean_ a lot.

At the moment [I’m trying to square that circle](https://discourse.flucoma.org/t/biasing-a-query/506/9) since the tools are built around a “match everything to everything” paradigm.

> [@b.hackbarth](#):
>
> There are lots of other interesting possibilities for hierarchical search functions. For instance, if a target seg’s noisiness is greater than 0.5, calculate similarity with descriptorN, otherwise use descriptorM.

I like this kind of conditional matching. @tremblap has done some conditional santizing where things that are below a certain loudness or have a spectral spread above a certain value are “dismissed” by the corpus creation process, but this could be very useful for querying varied input where things like pitch and/or confidence may be useless for certain targets as a way to just _skip_ that part of the query, rather than finding a way to sanitize the results, [which is not without its own problems](https://discourse.flucoma.org/t/data-sanitization-centroid-vs-nicol-loop/499).

> [@b.hackbarth](#):
>
> You’re correct - audioguide lets you match sounds using time-varying descriptor differences. And I do think that this is key to capturing morphological shape.

That’s great, and probably accounts for the _sound_ you get from AudioGuide, where things sound whole/complete (as opposed to granular/mosaicked).

I can’t think of how to do that in the FluCoMa context, as it on its face it would seem like a query per analysis frame or something like that. _OR_ just dumping the whole time series into a machine learning algorithm and letting it “sort itself out”. Presumably the time-series-ness would be reflected in the matching, but perhaps not explicitly, as it would be treated as any other distance relationship, rather than a hierarchical “container” for the rest of the querying to fall inside of.

As far as I understand it, the closest we have at the moment is having derivatives for any given value, which contains some kind of time varying information, though skewness/kurtosis can perhaps offer some idea as well. We don’t have vanilla linear regression (again, as far as I know).

> [@b.hackbarth](#):
>
> The most important thing with averaging spectral descriptors is to weight averages with linear amplitude. Are you guys doing this in fluid.bufstats~? If not, it should certainly be an option, if not the default.

At the moment, each statistic is an island. That is, you get seven stats (mean, standard deviation, skewness, kurtosis, and low/mid/high centiles), and then derivatives of these things. But each one is run on single data stream (typically a descriptor of some type, but since it’s buffer based it can happen on audio as well).

I suppose one could do this “manually”, but it would be quite tedious/messy since it would involve manually multiplying every sample in a buffer by a value, since all(ish) data types are buffers.

> [@b.hackbarth](#):
>
> I was originally doing this in a more computationally intense way with the mel spectrum. when the first segment was selected, its mel amplitudes were subtracted from the target’s amplitudes and target descriptors were recalculated on the residual mel spectrum. this only worked for mel-based descriptors like mel centroid

Did you abandon this approach due to complexity, or because of the limited usability? (i.e. only mel-based descriptors)

I’ve been working on some [real-time spectral compensation](https://discourse.flucoma.org/t/spectral-compensation/299/125) (e.g. using the (mel-band)-based spectral shape of the target, to apply a corresponding filter to the match to more closely have the two sound alike) so an approach like this might make sense since I’m already doing mel-band analysis of both the source and target anyways.

Presumably what follows below about the specifics of how to subtract and find remainders (based on loudness) would be the same when doing it _per_ mel-band?

> [@b.hackbarth](#):
>
> sound 1 frame 1 = centroid 1000, power = 0.01  
> sound 2 frame 1 = centroid 2000, power = 0.02  
> mixture frame 1: centroid 1666.66, power = 0.03

This makes more sense… And I understand what you mentioned above about weighing descriptors (in general) against their linear amplitude.

> [@b.hackbarth](#):
>
> My intuition tells me that this works best for approximating time varying descriptor mixtures, and will not work as well for sounds where descriptors have already been **averaged** into a single number.

Aaand the devil is in the details. So taking the means of spectral descriptors wouldn’t play so nice with this approach.

Thankfully for my most general use case I’m detail with tiny analysis windows (256 samples with `@fftsettings 256 64 512`), so the amount of smearing for so few frames is probably much less than what would happen across a file or segment that’s 1000ms+.

Either way, tons to think about, both in terms of things to test and apply, as well as some wish-list-y stuff for the FluCoMa tools.

---

_[View the full topic](https://discourse.flucoma.org/t/audioguide-talk/550)._
