# The need for speed (fluid.dataset~ + fluid.datasetquery~)

**URL:** https://discourse.flucoma.org/t/the-need-for-speed-fluid-dataset-fluid-datasetquery/673
**Category:** Pre-Release Toolbox2 Usage
**Created:** [October 4, 2020, 10:19pm UTC](https://discourse.flucoma.org/t/the-need-for-speed-fluid-dataset-fluid-datasetquery/673 "2020-10-04T22:19:51Z")
**Posts on this page:** 1
**Showing post:** 24

<div class="post-metadata">

### Author: ![rodrigo.constanzo](https://discourse.flucoma.org/user_avatar/discourse.flucoma.org/rodrigo.constanzo/32/12_2.png) [@rodrigo.constanzo](https://discourse.flucoma.org/u/rodrigo.constanzo)
#### Post date: [October 10, 2020, 9:57am UTC](https://discourse.flucoma.org/t/the-need-for-speed-fluid-dataset-fluid-datasetquery/673/24 "2020-10-10T09:57:17Z")

</div>

> [@a.harker](#):
>
> So much a mystery to you that it’s not even involved in this case. The results are similar but not the same, but the change of computer makes direct comparison void - I’d expect if you tested on the same machine that the ratio would go from 6.8ish to 5.8ish, which is noticeable

On a semi interesting note here, both actually speed up with the fully separated timing. So the ratio is slightly better, it’s still only by a factor of 0.3

This is the “bad” timing method for 10k/8d:  
 ![Screenshot 2020-10-10 at 10.50.21 am](https://discourse.flucoma.org/uploads/default/original/1X/2d7cfa3119f02dce6d9f5e4e632cb4e95513898a.png)

This is the “good” timing method for 10k/8d:  
 ![Screenshot 2020-10-10 at 10.51.10 am](https://discourse.flucoma.org/uploads/default/original/1X/a1c2809b1b9fa9ae2630e33cb5ce42b1c3132b02.png)

> [@weefuzzy](#):
>
> 100k points, starting to pay off

I guess we’ll see in due time, but I’m curious if this holds true as you go smaller. Like does it only start becoming equal around 10k?

Perhaps there’ll be an algorithmic sweet spot that below a certain amount of points brute force is faster, then it gets into KDTree land, with perhaps something above that (that maybe isn’t super useful for corpora-level numbers).

I’m specifically thinking of an immediate/obvious use case for me being the [time-travel/prediction stuff](https://discourse.flucoma.org/t/regression-classification-regressification/547/59) where I want to find the nearest distance on a pre-trained set of data asap, before moving on to more wiggly/complex querying. At the moment my test corpus for this has been \<1000, but that will likely change when I build the proper version of that. It would still be in the few thousand range though, at most.

---

_[View the full topic](https://discourse.flucoma.org/t/the-need-for-speed-fluid-dataset-fluid-datasetquery/673)._
