# Updates on Test Data

**URL:** <https://forum.storj.io/t/updates-on-test-data/26034>\
**Category:** Announcements\
**Created:** [April 30, 2024, 3:36pm UTC](https://forum.storj.io/t/updates-on-test-data/26034 "2024-04-30T15:36:45Z")\
**Posts on this page:** 20\
**Page:** 29

<div class="post-metadata">

**Author:** ![Mitsos](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/mitsos/32/1125_2.png) [@Mitsos](https://forum.storj.io/u/Mitsos)\
**Post date:** [June 3, 2024, 10:15pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/644 "2024-06-03T22:15:25Z")

</div>

High CPU usage on this one (clocks boosting way up) + higher load.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 3, 2024, 10:20pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/645 "2024-06-03T22:20:25Z")

</div>

bestofn n=4 could be a bit too much. (I still owe you a descript what exactly it does)

In terms of throughput that one was a new record for the bandwidth optimized RS number we are currently running. In combination with one of the even faster RS numbers this could hit the target.

---

<div class="post-metadata">

**Author:** ![Toyoo](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/toyoo/32/1126_2.png) [@Toyoo](https://forum.storj.io/u/Toyoo)\
**Post date:** [June 3, 2024, 10:22pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/646 "2024-06-03T22:22:10Z")

</div>

> [@littleskunk](#):
>
> bestofn with n=2 was a bit better. I have seen 60MBit/s on my storage node for a few minutes. We could maybe increase n and see how that works.

Probably because other nodes started GC and started losing races.

---

<div class="post-metadata">

**Author:** ![ACarneiro](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/acarneiro/32/1129_2.png) [@ACarneiro](https://forum.storj.io/u/ACarneiro)\
**Post date:** [June 3, 2024, 10:22pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/647 "2024-06-03T22:22:23Z")

</div>

Yes, it trebled (at peak even quadrupled) the baseline I’ve been seeing for the test.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 3, 2024, 10:28pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/648 "2024-06-03T22:28:16Z")

</div>

Up next bestofn n=2 but this time we remove the ip subnet filter. Don’t get too exited. This is kind of impossible for production but since I get that question kind of in every meeting we can as well test it 😃

---

<div class="post-metadata">

**Author:** ![BrightSilence](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/brightsilence/32/6048_2.png) [@BrightSilence](https://forum.storj.io/u/BrightSilence)\
**Post date:** [June 3, 2024, 10:30pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/649 "2024-06-03T22:30:17Z")

</div>

Certainly seemed to push a lot more data. 🙂

 ![image](https://storj-s3.bcdn.literatehosting.net/original/3X/a/4/a40c6f32c7a95b5155caa7c176eb86bd63da343d.png)

IO wait has been an issue today since I restarted most of my nodes earlier today and filewalkers are running simultaneously with some GC going on as well.

 ![image](https://storj-s3.bcdn.literatehosting.net/original/3X/9/9/9999808cbd5ec7fcc2d88d560f8567a61889962c.png)

This is mostly unrelated to the new tests, though I did see it increase the issue during peaks. Nothing to be too concerned about though as it seemed to have been performing just fine for the most part.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 3, 2024, 11:01pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/650 "2024-06-03T23:01:15Z")

</div>

Disabling the ip subnet filter had almost no effect. Up next we will try bestofn n=3 to see if the results are closer to n=2 or n=4.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 3, 2024, 11:14pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/651 "2024-06-03T23:14:30Z")

</div>

While the test is running I can explain a bit what bestofn does. Current active RS settings are `16/20/30/38` With n=2 it would select 38\*2=76 nodes and order them by success rate picking only the 38 nodes with high success rate.

Downside is that with n=4 it kind of stops selecting nodes with a low success rate. It still selectes them from time to time but less frequent than the choiceofn selection would do. On my storage node I noticed an decrease in usage compared to n=2. That indicates n=4 is now missing out on the resources that the slow to resonable fast nodes still have to offer. However on the other side the total throughput was impressive. So as a storage node operator myself I would vote for a more balanced approach but I would also understand if my manager ignores that wish an picks the higher total throughput.

Thats why we are testing n=3 to get a sense of how it performance on these 2 aspects.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 3, 2024, 11:33pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/652 "2024-06-03T23:33:34Z")

</div>

n=3 was very close to the n=4 results. So somewhere between n=2 and n=3 is the most gain and increasing it further doesn’t help as much. It was also using my storage node again.

Up next we will keep n=3 but combine it with the fastest RS number `16/20/30/60` from previous tests. Technically `3*38!=3*60` so this isn’t a fair comparison. So later I might want to try n=1.9

---

<div class="post-metadata">

**Author:** ![Toyoo](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/toyoo/32/1126_2.png) [@Toyoo](https://forum.storj.io/u/Toyoo)\
**Post date:** [June 3, 2024, 11:40pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/653 "2024-06-03T23:40:17Z")

</div>

> [@littleskunk](#):
>
> I would also understand if my manager ignores that wish an picks the higher total throughput.

Not sure why wouldn’t this make sense. I said this already: Storj will work on a small number of well-connected data center nodes with lots of storage, there’s no need for small operators to exist. It would be actually worrisome if that wasn’t known by your manager.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 3, 2024, 11:57pm UTC](https://forum.storj.io/t/updates-on-test-data/26034/654 "2024-06-03T23:57:32Z")

</div>

> [@Toyoo](#):
>
> Storj will work on a small number of well-connected data center nodes with lots of storage, there’s no need for small operators to exist.

We have that. It is called the storj select network and currently performing worse than the public network. Decentralizations and the pure number of nodes just outperforms it.

---

<div class="post-metadata">

**Author:** ![Toyoo](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/toyoo/32/1126_2.png) [@Toyoo](https://forum.storj.io/u/Toyoo)\
**Post date:** [June 4, 2024, 12:02am UTC](https://forum.storj.io/t/updates-on-test-data/26034/655 "2024-06-04T00:02:42Z")

</div>

Yet your results for n=4 suggest you could get rid of probably half of the operators of the public network. My impression is that if you performed an experiment like: pick k nodes with the best success rate, and run the old node selection algorithm just on them, for k=500, 1000, 2000, then you’d outperform even the n=4 experiment for at least one of the k.

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 4, 2024, 12:08am UTC](https://forum.storj.io/t/updates-on-test-data/26034/656 "2024-06-04T00:08:59Z")

</div>

> [@Toyoo](#):
>
> Yet your results for n=4 suggest you could get rid of probably half of the operators of the public network.

I read the results totally different. The test results are showing that we have a decent number of fast nodes and bestofn can make use of that but it is missing out on the resources that the rest of the nodes still offer. This is the wrong node selection for the job. Even the slowest node out there will still bring additional resources to the party that an ideal node selection can utilize.

---

<div class="post-metadata">

**Author:** ![BrightSilence](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/brightsilence/32/6048_2.png) [@BrightSilence](https://forum.storj.io/u/BrightSilence)\
**Post date:** [June 4, 2024, 12:10am UTC](https://forum.storj.io/t/updates-on-test-data/26034/657 "2024-06-04T00:10:16Z")

</div>

You probably know better which test corresponds to which peak. But the last one definitely pushed most data to my node. Though with the higher number of initiated transfers that doesn’t necessarily mean I get to keep more data as I’m sure more got long tail cancelled as well.

 ![image](https://storj-s3.bcdn.literatehosting.net/original/3X/c/8/c87398f55ca6f7c76b8060b1995d2f41caf650c1.png)

It’s now easier to see the CPU impact as well as it seems most file walkers have finished. Though I still have one GC running.

 ![image](https://storj-s3.bcdn.literatehosting.net/original/3X/d/1/d128b818182888edae5742cc993e5709c1020234.png)  
It seems like that last test peaked normal CPU usage more than IO wait, which I wasn’t expecting.

> [@littleskunk](#):
>
> Current active RS settings are `16/20/30/38` With n=2 it would select 38\*2=76 nodes and order them by success rate picking only the 38 nodes with high success rate.

This would lead to a much larger number of nodes getting significantly less data. It’s clear that I’m not in that lower part from the traffic I saw, but that does make me worry a little more about distribution. Is that an aspect you are looking into for these tests? I’d say especially if you push these numbers too far (like maybe n=4) the chances of pieces ending up on more coordinated systems/locations goes up, possibly impacting durability risk.

---

<div class="post-metadata">

**Author:** ![Toyoo](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/toyoo/32/1126_2.png) [@Toyoo](https://forum.storj.io/u/Toyoo)\
**Post date:** [June 4, 2024, 12:10am UTC](https://forum.storj.io/t/updates-on-test-data/26034/658 "2024-06-04T00:10:29Z")

</div>

So I assume it’s useless to attempt to challenge you to perform this experiment? 😛

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 4, 2024, 12:10am UTC](https://forum.storj.io/t/updates-on-test-data/26034/659 "2024-06-04T00:10:42Z")

</div>

Now testing RS number `16/20/30/60` with bestofn n=1.9

---

<div class="post-metadata">

**Author:** ![littleskunk](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/littleskunk/32/63_2.png) [@littleskunk](https://forum.storj.io/u/littleskunk)\
**Post date:** [June 4, 2024, 12:22am UTC](https://forum.storj.io/t/updates-on-test-data/26034/660 "2024-06-04T00:22:01Z")

</div>

> [@BrightSilence](#):
>
> This would lead to a much larger number of nodes getting significantly less data. It’s clear that I’m not in that lower part from the traffic I saw, but that does make me worry a little more about distribution. Is that an aspect you are looking into for these tests?

There is a reason we tested everything else with the other RS setting that has a much shorter long tail. In fact it has the shortest long tail I could come up with without increasing the risk of upload failures too much.

The point of this test is to verify how it compares to earlier tests. It does show us if the node selection itself is able to utilize the available resources without the old long tail trick. Ideally we find a node selection that gets us maximum throughput with minimal resource overhead. So from time to time we have to test the excessive RS number to verify how good or bad the current node selection settings are.

---

<div class="post-metadata">

**Author:** ![Toyoo](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/toyoo/32/1126_2.png) [@Toyoo](https://forum.storj.io/u/Toyoo)\
**Post date:** [June 4, 2024, 12:24am UTC](https://forum.storj.io/t/updates-on-test-data/26034/661 "2024-06-04T00:24:24Z")

</div>

> [@littleskunk](#):
>
> So from time to time we have to test the excessive RS number to verify how good or bad the current node selection settings are.

How about automating these tests with some form of black box optimization? I did so ages ago with Postgres and Apache Druid using [hyperopt](https://hyperopt.github.io/hyperopt/), was quite helpful not to spend engineering time manually testing different hypotheses.

---

<div class="post-metadata">

**Author:** ![Roxor](https://storj.bcdn.literatehosting.com/letter_avatar/roxor/32/5_5575768a8748004e209b776fc1b2916d.png) [@Roxor](https://forum.storj.io/u/Roxor)\
**Post date:** [June 4, 2024, 12:41am UTC](https://forum.storj.io/t/updates-on-test-data/26034/662 "2024-06-04T00:41:54Z")

</div>

So… today slower nodes still get selected for the _opportunity_ to win upload races (even if they may fail often). But the recent highest-performance configs don’t even give them a chance? If they never get a chance: they’ll leave.

I see what @Toyoo is saying: a more winner-takes-all node selection could benefit paying clients. But to the detriment of the health of the SNO community. Because yes speed is a priority: but also a strong diversity in nodes to mitigate risk.

Like the fastest performance may be “send all client requests to @Th3Van”. But if the network atrophies to 500 fast nodes then you’re _screwed_ when he’s offline.

This is all interesting stuff! I get that Storj wants to tweak all the dials to see what’s _possible_. Then they’re going to have to make some tough business decision about what’s safe and realistic. Glad I’m not you! 😉

---

<div class="post-metadata">

**Author:** ![Mitsos](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/mitsos/32/1125_2.png) [@Mitsos](https://forum.storj.io/u/Mitsos)\
**Post date:** [June 4, 2024, 12:46am UTC](https://forum.storj.io/t/updates-on-test-data/26034/663 "2024-06-04T00:46:59Z")

</div>

I may be misunderstanding a few things, but doesn’t the /24 rule take care of spreading over the network? As far as I understand it, it just focuses on the fastest nodes _after_ spreading it out over different /24s. If a node is fast, it makes sense to get more data. A paying customer isn’t going to wait on a raspberrypi with an SD card (as the node’s storage) for his/her/its upload, imho.

[Previous page](https://forum.storj.io/t/updates-on-test-data/26034.md?page=28)

[Next page](https://forum.storj.io/t/updates-on-test-data/26034.md?page=30)
