# Dimensionality reduction (mixing the rational and irrational)

**URL:** https://discourse.flucoma.org/t/dimensionality-reduction-mixing-the-rational-and-irrational/534
**Category:** Pre-Release Toolbox2 New Ideas
**Created:** [June 28, 2020, 11:10pm UTC](https://discourse.flucoma.org/t/dimensionality-reduction-mixing-the-rational-and-irrational/534 "2020-06-28T23:10:21Z")
**Posts on this page:** 1
**Showing post:** 5

<div class="post-metadata">

### Author: ![rodrigo.constanzo](https://discourse.flucoma.org/user_avatar/discourse.flucoma.org/rodrigo.constanzo/32/12_2.png) [@rodrigo.constanzo](https://discourse.flucoma.org/u/rodrigo.constanzo)
#### Post date: [June 29, 2020, 11:10am UTC](https://discourse.flucoma.org/t/dimensionality-reduction-mixing-the-rational-and-irrational/534/5 "2020-06-29T11:10:01Z")

</div>

> [@weefuzzy](#):
>
> It’s a case of try-it-and-see. … It’s not at all uncommon to do PCA first and then follow with a non-linear reduction when trying to go from lots of input dimensions to very few.

I’ll try PCA and see how that fares. I think if I’m just in millisecond land, it would (intuitively) give results that made sense that way. If I include derivatives that would probably crumble without sanitization.

Curious about this PCA-\>non-linear workflow too. By this do you mean do PCA on something to bring it down to a smaller amount of dimensions and then to do MDS on that lower dimensional version?

> [@weefuzzy](#):
>
> > [@rodrigo.constanzo](#):
> >
> > I wanted to keep a sense of the difference in scales between the (currently random) entries,
> 
> This is the bit I’m not following. Keep a sense of the difference in scales for the purposes of weighting the dimensions’ relative importance in the reduction?
> 
> > [@](#):
> >
> > but I guess that would still be the case if I either standardized/normalized them anyways(?).
> 
> No: the normalisaing / standardising is done per dimension, explicitly to remove the difference in ranges. At this point it’s more productive to think about how the input data are _distributed_

I’m not really clear on things either, but say I have one file with a duration of 5000 and a time centroid of 2500, and another file with a duration of 400 and a time centroid of 200. By not sanitizing the data, I was hoping to maintain that difference in scale, rather than them having somewhat similar values (?) after sanitization.

More practically, I want the fact that the first sample has bigger values to be present in the ‘timeness’ metric that I can query later on, if, for example, I want to choose samples that are ‘timeness’-ier (yikes!).

> [@weefuzzy](#):
>
> They’re not so different!

In spirit I guess, but I’m still struggling with my normal use cases where I can find the nearest match for certain fields, and then some other criteria for other fields (ala [biasing](https://discourse.flucoma.org/t/biasing-a-query/506/9)).

---

_[View the full topic](https://discourse.flucoma.org/t/dimensionality-reduction-mixing-the-rational-and-irrational/534)._
