Updates on Test Data

What a bloody clever idea! Not sure whether it will be easy or even necessary to implement, but seems like a great suggestion!
Fun times indeed! :smiley:

Will we receive an update on the results as well, or are they only for internal purposes? For instance, will information about the peak bandwidth and other tests conducted be shared? (I hope I didn’t miss any updates among the many messages.)

There will be if the customer then tries to download that repaired segment.

You’d have to be pretty unlucky, wouldn’t you? I mean, repair triggers before there are not enough segments to reconstruct a file, or did I not understand it right?

Seems like the node was overloaded and couldn’t handle the IO pressure. Could also result from disk failure.

The rest of the stacktrace is just the node forcefully shutting down from failing the readability test. You can increase the timeout (not recommended) with --storage2.monitor.verify-dir-readable-timeout or change the config to only log a warning when the readability/writability check fails with --storage2.monitor.verify-dir-warn-only.

However, the default behaviour is set to fail fast so you can quickly fix it and your node doesn’t get disqualified

Maybe I missed it, but is there a predetermined system/logic which determines the actually tested node? I mean the test traffic/load is directed to any given node, then the next one, then the next one and so on…? Does the test take place on each and every node? Or since it is from SLC satellite, the geographically close nodes are tested only?
I ask because I do see an increase in traffic, but far from so much that would really put the nodes to the test. My trafic is around 15Mbit/s and it is shared over 6 nodes (behind the same IP)…
I had much more than usual egress on saturday-sunday and 3 times more than usual ingress today. But nothing more, nothing else. I have a 1000/300 Mbit connection, so pleanty of room for more. The nodes are fine and happy even the smallest with the RPi3B+

Repair must be done quickly. The safety of the data is the priority, so as soon as all the segments are back to the standard number, the better. You can’t delay the repair. If this is tunable, than I will choose the fastest nodes, not the slowest. This is actualy the case now; the fastest nodes upload pieces first, and the rest get lost races.

5 to 10 seconds quickly, not less-than-a-second quickly, though… The slowest nodes would surely still be plenty fast enough?

That’s certainly a nice theory. However… for as long as the numbers have been on the dashboard… ‘healthy pieces’ min numbers have been around 50, and median around 65 (of 80). Repairing back to 80 is clearly on a best-effort/good-enough basis. Storj knows if you need 29: you don’t need to stress when you’re missing a few.

So repairs must only be done “fast enough”.

And repairs aren’t sequential: but they are almost perfectly scalable. If each action took twice as long Storj could certainly run twice as many repair jobs at the same time if they needed to.

I’m not saying the idea of a slow-node-tailored repair system is good or bad. Simply that “Repair must be done quickly” is not true: based on the numbers provided by the satellites that have been running for years.

I should have used more words. That data would end up on nodes that don’t do well under pressure. If the customer would then download that data, it would be downloaded from slower nodes as well. So there is an incentive to just upload less to nodes that don’t deal well with the pressure.

Repair running slowly might add to cost though if it means more repair workers are needed. But the speed of the repair itself is certainly not an issue.

That’s in line with what I see per IP.

This sounds pretty reasonable, esp. considering that this use case has a TTL at the time of the upload. Would only be logical that data thats only stored for a short time does not need that much replication.

Game plan for today:

  • Test 6 KB files to find out what the maximum throuput of the satellite DB is
  • Test 2-3 alternative RS numbers
  • Test an alternative node selection (requires a cherry pick and point release. Not sure if we can get that in time.)

Well you are pushing pretty hard on bandwidth which we get 0 payment for. I wouldnt complain if it was filling our disks but its not.

image

After all the changes… I think I just got back to the same stored-data I had May 1st. At least I’m not going backwards this month! :wink:

I see it like an investition , so Storj can betterly serve customers.
We are in it together for good and bad, if Storj needs that, i will provide whats needed, for theirs success is also mine.

I am. Apparently they cleared the old data from the saltlake satellite and I had about 7TB of it.

yeah, me too. Many nodes losing several TBs and only recovered 600gb with this testing… Anyways, I am doing the periodic maintenance earlier this year :smiley:

Fairly sedate test today. Nodes didn’t even break a sweat :wink:

This makes sense. If storj has to pay $1.5/tb/mo for a fast node, or $1.5/tb/mo for a slow node, why would they want data on the slow node?

Just tossing out ideas: but the best customer experience would come from always having many fast nodes available (and not full, rejecting ingress). Filling them with repair traffic, which no customer ever sees the performance of… would be a waste. So, since you still want to encourage slower nodes to stick around (since they always contribute raw capacity)… why not nudge repairs towards slower ones?

Like you said: every node gets the same $1.5/tb/mo. So why not configure things so there’s a greater chance fast ones always have space for those customers expecting fast uploads?

Not saying that’s a good reason, or bad reason… just a reason :slight_smile: Makes you think!