SNOs with lots of data... How?

As I’ve been entertaining myself with reading the restructuring threads tonight, a recurring question occurs to me. How do SNOs end up with hundreds of TB of data? I’ve been stuck at ~35% utilization for literally years. I have low latency 500/500 fiber internet, but only a single IP address. I would have expanded if my disks filled, but they just sit. Are my results really because my 9 nodes are behind a single IP address?

My success rates don't look terrible either (from a node at random).

========== AUDIT ==============
Critically failed: 0
Critical Fail Rate: 0.000%
Recoverable failed: 0
Recoverable Fail Rate: 0.000%
Successful: 2045
Success Rate: 100.000%
========== DOWNLOAD ===========
Failed: 9
Fail Rate: 0.003%
Canceled: 391
Cancel Rate: 0.110%
Successful: 355690
Success Rate: 99.888%
========== UPLOAD =============
Rejected: 0
Acceptance Rate: 100.000%
---------- accepted -----------
Failed: 0
Fail Rate: 0.000%
Canceled: 230
Cancel Rate: 1.028%
Successful: 22146
Success Rate: 98.972%
========== REPAIR DOWNLOAD ====
Failed: 0
Fail Rate: 0.000%
Canceled: 0
Cancel Rate: 0.000%
Successful: 19813
Success Rate: 100.000%
========== REPAIR UPLOAD ======
Failed: 0
Fail Rate: 0.000%
Canceled: 14
Cancel Rate: 0.587%
Successful: 2370
Success Rate: 99.413%
========== DELETE =============
Failed: 0
Fail Rate: 0.000%
Successful: 5
Success Rate: 100.000%

Of course. They are treated as 1 node.

By openly working around /24 rule and hosting a swarm of nodes on an industrial strength server farm:

Sure, of course, I understand. Does everyone show ~35% utilization? If the network was evenly distributed across every /24 with nodes of reasonable performance, we’d all show the same utilization?

Utilization will depend on the allocated size. Currently node levels out at 10-12TB – deletions balance out additions at about that size.

:rofl:

You’re absolutely right.

You’re saying that every /24 of nodes is getting around the same quantity of data. Theoretically, if I got 5 IP addresses split evenly between nodes, my 50TB might get filled?

Nobody knows. Now the choiceofn rule is in place, so it’s not guaranteed.

Yes, from my anecdotal experience on geo-distributed nodes, across western US.

Sorted by size:

  • The Node1 was created in 2022 (cable internet).
  • Node4 – 2024 (symmetric fiber).
  • Node 9 - early 2026 (also cable internet)

The 12TB seems to be a saturation point

Did you use a one connection channel?
Because if not, then it’s a normal distribution, I guess?

No, all except 2 and 9 are on the different hosts in different cities. 2 and 9 are on the same host/same IP.

Actually 2 and 9 also add up to 12! Nice!

Yes, I’m not saying it’s abnormal :). Works as designed.

Not exactly. Would fill bit more, up to ~ half per disk on average, just like rabbit showed.
So the additional IP’s alone won’t get Your disks filled, without reorganizing whole file system of Your machine to put more nodes on one disk, each with its own ip.

Which is a worse design. Thus I would support a @Vadim’s suggestion

This will punish you to point more than one node on the one disk. Run the next node not because you want more data, but to be able:

  • ditch of some to get some space back
  • distribute and minimize the failure

(i don’t approve that, just to be clear.
Aldo i once ran 5 nodes on 1 ultrastar disk for a year or two, but it was all 1 ip (no cheating), and the CPU had problem more than disk, coz it was some low-watt one)

Or You just need detection if some disk machine has nodes with different Ip’s to make sure not to send pieces of same file to them. Like anti-cheat in games.
I even wrote some guidance on how i would do that, but didn’t send You that.
Most of it a common sense, but some direct methods too i specified, could save You time.

Currently I don’t have an idea, how to achieve that without a /24 rule. Because in Select we have contracts, and distribute pieces in an equal way to allow the data to survive even if the one operator could plug-off. However, they at least are immune to a usual offline issues and have a NOC 24/7, also HA on every level and also they put their money to provide this.

Actually, I have one. The home DC owner can put a collateral equal to their equipment and/or 100% of their annual income to guarantee that they will be not a Byzantine actor, which they are. Just take a look, how easily they disposed /24 and 1 node - 1 disk rules.

This was suggested a number of times in the past. Example:

This is unrealistic. For someone storing several PBs of data, that is tens of thousands of dollars tied up as collateral. With zero guarantees Storj will still be around in a year, or that payouts won’t change again.
And again, such capital could go towards a better hardware, better connectivity etc.
We need to reduce SNOs expenses, not to create new ones.

The only way is to trust the existing SNOs and lift the /24 rule for the existing nodes only. And that’s it. That way nobody will be able to abuse anything, unless they were willing to spend even more money. Which I doubt, since even now what is left after paying the bills is barely worth it.

The smart contract might be implemented to support both sides. They are designed for these cases, where you do not need any arbitrage to get money if the condition is broken.

Actually, this is how the contract for Select operators is designed, the usual one, on paper. They will pay, if they break the contract. If they performed without issues, they will not pay this fine. Easy.

In the case of a smart contract yes, you need to prepay. But you will get everything back, if you didn’t broke it. Just don’t broke?

Alex,
Or You can keep /24 rule, but make sure the network strongly favors ingress in nodes without VPN’s.
VPNs have bigger pings (time to first byte).
A normal fiber connection is like 1-10ms. A VPN is up to 200ms
Not all VPN nodes are cheating, how ever this way, nodes circumventing /24 rule, will fill very very slow, and existing full nodes will start losing data more than gaining new. It will be healthier for the network, and just to nodes, that are using legit IPs.
In order to do so, the traffic should be directed first to those fastest nodes until they are almost full (eg. up to 80%, 90%) So until the fastest nodes have free space to fill, no VPN nodes should be chosen for new data.

Three birds with one stone:

  1. You need to take data from shady nodes, so they can’t threaten abrupt, non-GE exit.
  2. And so You have pieces spread firstly on nodes that don’t hide behind VPN’s.
    (VPNs are additional cost, such nodes are more likely to drop out, at any payout cut)
  3. VPN makes network slower, less performance for customers = less value = less customers, and we need more.

legit nodes will fill fast, so they can accept lower payout rates more willingly.
This would cut my nodes too.. or i can adapt by moving to legit, non-VPN /24 IPs!
Or i can just stay at one IP and pool all my disks, that would make sens if the traffic would really be distributed firstly to fast and unique IP’s

Staying on VPN won’t bring me any more benefits.

(Actually i don’t know how my router will handle this, because currently each node has its VPN, so a VPN absorbs many connections, its easy on my home router, but if all this many connection requests hit directly,… some people reported here on forum such problems once) So a VPN has this benefit, but network should prioritize non VPN for performance.
Actually, if 1 node per 1 IP per 1 router, the connections should not be that bad.
Actually i like that my ISP don’t see too much.
Well i mean it still allows VPN’s but makes them last in queue for data.
It will just improve performance and data distribution.
This can be combined with other ideas i found,
i just opened my storj2.txt from 2023…

A Byzantine actor is a term from distributed systems and computer science, referring to a participant in a network that may behave maliciously, unpredictably, or fail in arbitrary ways.

Currently I would say it is rather Storj who is a Byzantine Actor.

This is the real Storj problem:

If Storj would fill a 28TB disk in a year we would not have such a discussion I guess.
It would be great if Storj would spend their time on finding new customers for the global network for example by getting SOC2 for it.

If you remember the testing for the non Select customer Vivint, it has clearly showed the impact that a larger single customer can have can be massive up to the point where Storj had even to consider to operate surge nodes by themselves.

It’s really sad to see the community now arguing over the measly 58 PB we have instead of finding ways and customers to turn 58 PB into 3 EB.