Using multiple IPs on the same physical service

To be clear: I don’t have such a setup. And I’m don’t have the intention to create one. But thinking in the interest of the STORJ, night not ashtrays be my own interest in the short term. Although in the end, I hope it will benefit me as a single SNO.

That’s the whole point. I think SNOs are the biggest point of failure in the network. Even not necessarily their hardware. Often people coming along, who in the end don’t have the patience to wait for filling their nodes and held back period who are saying goodbye as soon as there are relative minor problems we all have seen along the way.

In what sense and under which circumstances? Total bandwidth, because they are outnumbered? Average latency? Uptime?

That only is a point, if you let it happen that more than one piece end up with the same SNO. Otherwise it wouldn’t even matter if five or ten node operators even having 1% of the network each, are updating their clusters.

Actually you’re underlining my point in the whole topic with this: the /24-rule actually doesn’t correlate all too much with independency.

No, I just say:

  • /24-rule might be a bad proxy voor independent nodes.
  • The ToS isn’t valid anymore, not aligned with previous rule / current practice and doesn’t actually target any measures to cheat that rule aside from no means to take enforce it.
  • They’re might be other means to consider especially on the level of SNOs, so for example benefiting SNOs who have same wallet in the whole network and are long committed. The total amount per wallet could be a could measurement for that.

It’s different in each case, we usually offer Storj Select to the companies which requires this certification, otherwise it is better to use Storj Public, especially if they are not in the regions where Storj Select is offered.
But usually the throughput is higher on Storj Public, if the customer is not in the same region with Storj Select operators. Sometimes even for some customers in US Storj Public works faster, it depends. So we usually always suggest to use Storj Public first.

Exactly. So I some kind of agree with your statement that we also need to include SNOs to the node selection, just it’s not technically easy. The email+wallet combination is easy to generate, so out of scope. I wouldn’t suggest to use that in the node selection alone.

It’s correlate. The Storj Select nodes are selected (sorry) differently than in Storj Public. But now we have an ability to set the selection dynamically for nodes.
The rule /24 still serve its purpose, but with combination of “choice of n” it also helps to distribute pieces better.

Even if I agree that it need to be updated, the quoted rule gives the ability to do not permit to alter the default network behavior, including /24 filter rule.
You will always get the same answer, independently of how much you want to break it or make invalid. It’s not allowed.

You may contact Storj to change the selection rules for your setup, if it matches requirements. Most of such setups doesn’t, thus the default rules are applied, including forbid to run multiple nodes with different IPs outside of /24 subnet on the same server.

The wallet is too easy to generate still, as an email address. They cannot be used in the node selection alone. All suggested combinations still allows to circumvent the node selection, they also cheapest from possible to bypass the rule.
We need also a strong incentive for SNO to always use the same for all nodes. It’s hard to convince, because they would know, that different combination would give them more data.

Your idea about dependency on the used storage sounds interesting, but unfortunately cannot be applied to everyone, otherwise the new node would not have enough traffic to be selected more often. And if we would change a coefficient to another case (to select empty nodes more often), it would again give an incentive to create millions of new nodes.

Nobody suggested something better, than we have now.

The terms of service so far have not really been enforced and have been a bit left behind (as we have seen with outdated terms).
With this I do not mean that therefore they are useless and should be disregarded, but rather that the terms of services are not the “final word” at the moment, atleast until they are kept up to date and are clear.
For now, experienced SNOs and Storj know how the network is designed to work and the reasons for it.

In the case of the /24 rule, its quite clear that the network has this rule as a “protection” method against storing multiple pieces in the same location, but I believe this rule no longer works as intended. Big SNOs (outside of select, which not everyone can join for obvious reasons) openly use multiple IPs because otherwise medium size operations would not even be feasable.

As Storj grows, there is the question of what is the target SNO profile. Multiple nodes per SNO, or multiple SNOs with single nodes. Generally SNOs tend to have multiple nodes, because while you can use what you have and host a single node, the uptime, bandwidth, and IOPS requirements might not be compesated by the revenue of a single drive.

Instead, I think its becoming clear that some SNOs provide a large chunk of capacity by having a lot of nodes, which in most cases means lots of nodes in a single physical place. As the capacity demand increases, most of the growth will probably come from these large operators, rather than loads of small operators (a bit of speculation from my part here). I believe Storj knows this as they have encouraged SNOs to add additional capacity to their systems.

The reason I bring this up, is that the current rule, if not taken advantage of by using multiple IPs, limits SNOs from expanding. I believe a rework of the rule could be advantageous for both SNOs and Storj.

One possible way of reworking this rule is to not longer limit the selection of nodes to a subnet. Node selection would randomly pick from all available nodes equally instead of using subnets, but would also prevent two nodes sharing subnet from being picked at the same time (this could be implemented in numerous ways, e.g. pick node sequantially and keep track of used subnets, e.g. 2 pick all nodes and check if there are multiple nodes of one subnet, if so, repick all expect one of the nodes that share subnet). The idea from this is that two nodes under the same subnet will have the same chance of being picked than two nodes in two subnets, but node selection will still ensure that the nodes dont store pieces of the same segment.

This "solution"comes with some advantages and disadvantages:

  • +Ensures network safety by not concentrating pieces of a segment in a single location.
  • +Encourages SNOs to grow capacity.
  • +Disencourages “cheating” as multiple IPs no longer provide much benefit.
  • -The proportion of data going towards smaller SNOs might decrease.
  • -Data may be concentrated a bit more (as these SNOs would probably get more ingress until they fill up, bestofN could partially mitigate this issue) in large capacity operations (keep in mind that redundancy would always be maintained).

I doubt anything will come out of this post by I rather provide my 2 cents.

Nothing to be sorry about, I think it’s just wise. The way of thinking from that point might differ, however. I would say considering their uptime, it’s not so bad to have two pieces of the same segment ending up in their nodes so now and then. But that’s not up to me to decide in the end.

We can go over the lack of definitions and no version management / explicit agreement to new versions again, but I won’t. Southe yourself in a non-juridical wat on this one, I would say :wink:

The idea essentially boils down to stratifying by subnet, but increasing (=multiplying) chance of a certain subnet being chosen by the amount of nodes in that certain subnet.

In that case, the number of nodes still does give node operators an opportunity to cheat. So instead of during up 1 node of 2X TB, it gives more selection chance by firing up 2 nodes of X TB.

Although the remainder of it sounds interesting. Because it also alleviates the concurrency for this in a subnet with many neighbors. Although those subnets might also turn out to be crowded by the more professional SNOs.

How do you mean?

That is a good point. One could argue that this can already be done by having multiple nodes in multiple subnets. Keep in mind that hosting multiple nodes in a single HDD will probably bring more issues than benefits (mostly due to IOPS contrains, and bestofN would remove bad performing nodes), therefore the “best” way to do this would be by purchasing more drives with less capacity.

Currently if you only operate in a single subnet, the biggest limit to your growth is ingress. If you want to deploy 100TB of capacity in a single subnet, you might as well wait a lifetime for it to fill (bit of sarcasm, the amount of time it would take to fill cannot be easily predicted).

While I do not know what the current “equilibrium” point of capacity is (ingress - deletions) for a subnet, its probably not very high, and reaching the maximum would take a very long time.

Therefore given that filling new capacity will keep taking longer and longer as you add more nodes to the subnet, it limits how much and how fast a single SNO can grow in a single subnet.

Indeed:

So yeah, since I got plenty drives laying around of 1-5TB I used them. So I got one subnet with about 20 nodes. Chance someone else will sign up successfully, is quite low. Although current traffic is really bigger then even was a year ago.

I see, indeed the most depending factor in the end becomes the amount of not rim-filled nodes in a certain subnet.

It only would be an incentive for a very starting node operator, say you set the lower limit to for example 2TB. People could be inclined to create multiple nodes.

But if you make sure, they can’t merge those nodes to one SNO in the future, and communicate this on beforehand people will quite soon run in a situation in which their assumed advantage will become a disadvantage as soon as the cumulative used size of the nodes grow over the lower limit. So it will be only a real advantage for very small SNOs, that already don’t pose a real threat to the network, because the low chance of getting two pieces of the same segment anyway.

Whether you need to do this with email address, wallet or SNO-key of any other kind is just for debate.

Each piece is unique, Erasure Coding allows to recreate an original segment using any n from k pieces.
So we usually talking about pieces of the same segment. If k-n+1 pieces of the segment would be hold by the one location, e.g. one server, then there is a very high chance to be unable to reconstruct a segment due to insufficient amount of remaining pieces if that location would be down for any reason - hardware issue, operator’s error, ISP failure, the local disaster with power outage or damage, etc.

In the Storj network we also have a minimum repair threshold m (n < m < k), when the repair job would be triggered, so in reality it could be enough to make the segment unrepairable even if the location holds even less pieces of the same segment m-n+1. The repairer wouldn’t have a chance to recover the segment, if that location would disappear with all held pieces.

We tested the removal of /24 subnet limit and discovered, that the distribution become worse. Not as bad as we expected, but noticeable worse than with /24 filter and it didn’t increase a throughput, so doesn’t make any sense to remove this rule now.

Thanks for the correction, I meant to say pieces of segments but accidentally said blobs of pieces. I have corrected the wording.

I believe you might have misunderstood my post. Im not advocating for the complete removal of the /24 rule, rather I am suggesting a rework of how node selection works while maintaining the 24 rule.

For a storage provider data integrity is paramount, therefore multiple pieces should never be stored together. That is why in my suggestion I say to enfore that only one piece of a segment may be stored together, but alter the selection instead.

Once again, the reason for this change is that currently “cheating” the system by hosting multiple nodes with different IPs in the same location is advategeous compared to sharing a subnet because otherwise running medium scale operations might as well be impossible due to the low ingress of a single subnet. This “cheating” also in turn breaks the /24 rule, making it such that multiple pieces may be stored in the same physical location.

My suggestion in turn eliminates the advantage that using multiple IPs provides, therefore disencouraging this form of “cheating” and enhancing network integrity. In addition, this eliminates the current “hard limit” that using a single subnet has (that having a lot of capacity in a single subnet would take an eternity to fill), encouraging SNOs to provide more capacity instead.

Yeah, a very likely situation. Like when one SNO/location has about 500+ subnets.

I did a very unlikely quick-and-dirty calculation, say:

  • every SNO has 5 nodes on average
  • 10% of the SNOs is cheating and have all nodes situated in different subnets, whether it may be VPN or a data center.
  • The ‘honest’ SNOs chare in 10% their subnet with others (making 1/3th of all selectable nodes ‘unhonest’)

Even in that case you must have more than 1000 subnets, to make a situation like this happen.

See:

So, it is really more about fairness than real threats.

Especially since I believe most SNOs are just set-and-hope-you-can-forget types.

In what sense ‘worse distribution’? If you define it as more pieces ending up in the same subnet, than that outcome was inevitable. Do we really have other measurements to see how it distributes over time?

Yeah, I understood your suggestion as first selecting all different subnets and then give them a weighing factor considering the amount of nodes are running within it. So a subnet with 1000 nodes is 1000 times more likely to be choosen than a subnet with one node within it.

We apparantly had the same intent.

I’m getting the feeling however, we’re wasting our time since it most probably will change as soon as the ToS will be aligned with current practices :wink:

My suggestion really doesnt have anything to do with the ToS. The fact is that the current implementation of the /24 rule indirectly encourages cheating and limits the growth of SNOs. Regardless of what the ToS are, at the moment it does not seem like Storj is actively attempting to limit workarounds of the /24 rule.

Instead, I was simply suggesting a rework to the current implementation of the /24 rule in a way that may benefit the network.

I know, I’m pointing only to the changeability of things as they are now.

Agree. Although I can’t judge them right nor wrong on this part, because I don’t have the data they have. But as far as I can see, cheating on a limited scale doesn’t really pose a threat to the network. Although for a good party by grace of the fact, they have 51 spare parts for every uploaded piece.

By I’ll take my own advice and won’t spend too much effort in it. I hope ‘they’ are wise and take good care of the network as a whole.

Of course, its not my intention either to judge if this is right or wrong, Im just pointing out what I believe to be some “flaws” in the current system and my suggestion about how it could be improved.

How the system is enforced and the effects of “cheating” is not something I have antyhing to say about.

I believe that it’s available only on the satellite’s side, unlikely we have some report to show. So only the person who performed a test can be sure, how evenly were distributed these pieces.
I can only share what I saw here on the forum, and I think @littleskunk did this test for sure.

With all respect I would disagree. ToS doesn’t prevent a Byzantine behavior as we know, so it’s better to implement something in the protocol instead. And we doing such improvements over the time.
So, I would hope that not only ToS would prevent from such behavior.

The obvious one example - do not run multiple nodes on the same disk. You can, of course, nothing will prevent you from doing this (except ToS, but you are sure that you can afford that), except nodes themselves - they are highly intensive IOPS-consuming processes, so such nodes will be always affected by each other, because they are trying to use the same limited resource - the disk in the same time. Especially worse when the server is restarted, they all will start in the same time.
Of course you may invent a bunch of scripting around that to start them one by one… however, it’s not free.

Or you could have 9 TB of RAM… Woot!

What’s this about? …?

Wow … sorry talk about off topic, my bad. LOL… I believe I was messaging somewhere else or rather, intending to…I do not recall that, must have imbibed too much last night. Today I started making a massive cluster shared ram drive. Defragging in ram is a miracle to watch.

I feel like maybe there’s room for something in-between the public network and Select. I’d say I have a setup that’s higher quality than usual - UPS, RAIDZ, ECC, SSD cache, PLP, all server-grade parts - but I’m certainly not going to be achieving SOC2 certification for a homelab any time soon. I’d say my node is probably a lot less likely to lose data than someone running an rpi node with a single HDD.

Or, and I realize this might not be the most popular option - increase holds a bit, so that people are incentivized to set up for the long term and not lose data.

domestic power system and internet will not making your node more reliable than raspberry pi, as one lightning strike can turn you off for several days.

This is still a single location for the purposes of considering natural disasters, you’re still a single operator for the purposes of considering the risk of human mistakes, and it’s still I guess a single ISP? The incremental value of your step being more stable is still smaller than just attracting two rpis operated independently. That’s how the probability calculus works.

SOC2 is valuable for Storj not because it certifies specific things to be high grade (it actually doesn’t!). It is valuable because it is a hard requirement for some customers.