Temporary Maintenance Mode for Storj Nodes

Not being paid by Storj to run a data centre. Merely renting out spare hard drive space.

Since I like to protect MY data, my nodes get the benefit of running on a redundant filesystem.

For a long time, the “recommended” way to run a node was a raspberry pi with no RAID. Probably with a single power supply (as while it is certainly possible to have more than one PSU for a raspberry it requires some electronics knowledge).

In my opinion, 5 hour reaction time would reaquire hiring staff, but Storj does not pay enough for that, unless I cheat and use VPNs to have many different public IPs. If you are always capable of fixing a problem within 5 hours, great. I don’t think many people can on account of them being at work, driving etc at the time.

If it’s suspension and not disqualification, then probably it does not really matter that much, it would stil suck and if this is advertised as “run on a raspbery or your NAS device”, then the expectation is unrealistic.

That is usually very hard to achieve in a residential setting. The government would not run a new line to my house or it would ask for so much money, I would have to run the node for couple centuries to get any ROI.

Why that specific number and not something else, say 99.2%, 99% or 98%? Why within a month and not within a year (which would give time to solve some uncommon major problem)?

I haven’t seen any recommendation to run nodes on raspberrypis. I’ve seen plenty of posts where “it works”. That’s not a recommendation.

Let me clarify something: I’m not saying as soon as you get to 5 hours and 1 second your node is erased from the network. I’m saying as soon as you get to 5 hours and 1 second, your node is suspended (=does not receive any more data) until it gets back to 4 hours, 59 minutes and 59 seconds of uptime. See below for rolling period comment.

There is nothing unrealistic about suspending a node for not adhering to the ToS. We all agreed to those terms, didn’t we?

Most houses (at least those that require more than X amperes (current) per single phase are required to have a three-phase set up. The easiest way to achieve A-B? Run one socket from L1 (the first phase) and the socket right next to it from L2 (the second phase). You can even go full paranoid mode and have a socket next to those two wired up to a generator.

Each of those phases is run directly from the electric company’s step-down transformer, on a separate cable, up to your house. Sure I can drink more than I should have one night, and crash into a utility pole, taking out both of those lines feeding your house. How long until the electricity company manages to fix this? A day? A week? For that once-in-a-lifetime event, I’m sure you can handle a suspended node for a bit.

Because that’s the number that gives us a 5 hour window during which the network can detect a node as being offline, start marking data on it as unavailable and possibly a candidate for repair. If a repair needs to be triggered, then it gets triggered.

Within a month because the network can’t wait for a 2d 13h 21m 39s window in order to mark a node as being offline. And don’t forget that you’ll need to wait for another full year for your node to fully recover from that.

Just to be clear: I’m not advocating for fully redundant infrastructure (multipath SAS disks all the way to redundant path power runs) nor am I advocating for cold-spare disks, ready to be swapped out at a moment’s notice, targetting five-9s of uptime. I’m advocating for nodes that adhere to the 99.3% uptime requirement.

This doesn’t give you A-B. Same power source, same power company, same fuse panel. Just different phase. A blackout will take out all phases anyway.

It took me literally less than a minute to find these two links:

Given other expectations to not buy hardware specifically for Storj, I can only infer that if a prospective operator had only an RPi, they were recommended to try it this way.

And when it fails, the entire three phase line fails. I don’t think I have seen a situation where one phase was missing (though I’m sure it happens) - it’s usually all or nothing. If there’s a short, it usually trips the main circuit breaker which cuts off all 3 phases. Same with an overload.

Why 5? Why not 6 or 24? As it is right now, nodes are only suspended after an extensive downtime and the network seems to be doing well, why not keep that?

And that is not that much different. I need spare parts, because it is not usually possible to get them in 5 hours (even if the failure occurs during working hours and I have a day off). The setup has to be redundant or at least there have to be cold spares for everything (routers, switches) because, again, 5 hours is not enough time for anything other than a quick replacement of a broken part and a reboot.

Storj network doesn’t really care if your node is available or not. If it can’t retrieve what it wants from your node, there is another 50+ nodes it can try.

The only person that loses is you. Whilst your node is down, you get no uploads, downloads, repairs, and you will likely receive some deletes when you come back online..

Onus is on you, to maintain whatever minimum uptime you desire. Obviously, higher uptime means better reputation = more $’s. (better chance at more $s anyway).

And for best results, live in a location that has lots of traffic, and cheap reliable power, cheap / fast internet.

regarding three-phases: yes two phases from the same company doesn’t give you true A-B. It gets you as close to it as you can get, in a residential setting. We are not talking about building a real datacenter for storj, everybody calm down.

I don’t like going off-topic, but here goes: In industrial setups (ie setups running mostly on three phases) there are situations where one phase fails. If you haven’t had the chance to see this happen, three phase motors start kicking and buzzing. There are protections against this, ie detect a low voltage on one phase and turn off all power to the motor to protect it.

If a short takes out your main breaker, I strongly suggest you find a new electrician. A short should never even trip the distinct circuit breaker, unless there isn’t any fuse between that circuit breaker and the load. And no, you don’t rely on the circuit breaker to protect you in case of a downstream short, you isolate the short as close to the load as humanly possible. The circuit breaker is there as a last line of defense against the wiring getting overwhelmed, in case there is a short in the building’s internal wiring (ie you are drilling a hole in the wall and hit that wiring). It’s not there to protect you (or the wiring) in case of an electric kettle shorting out when you try to boil some water.

And for the love of all that is holy: circuit breakers get stuck. You may not have seen this happen, but I have. When you need to protect against a short, you always rely on a fuse first. Technically speaking there isn’t any way for a fuse to not work as expected.

Yes I agree that nodes are suspended when they drop well below the 5 hours. What you didn’t write though, is that a node that is offline longer than 5 hours has its data marked as unavailable, until it comes back. In case a repair is needed for some of that data (during the time the node is offline), the data is repaired and removed from that node (as far as the satellite is concerned). This happens the next time the node gets a bloom after coming back online, and this is why you see a sharp drop in stored data after a long period of being offline.

Your drive failed on a Sunday. You turn off the node, take the drive offline and wait for a replacement. By the time the replacement arrives, and you had the chance to migrate your node to it, then start the node, a week has passed. You node is now back online and the uptime score starts recovering. You node should stay suspended until it gets back up to the 99.3% uptime, just to make sure that the new drive you put in isn’t going to fail in two days. Do you have any different suggestion? Shouldn’t nodes that are offline for hours and hours be penalized in any way?

@Toyoo a guide on how to install to particular device isn’t a “recommendation” to use that device. Here are the recommendations Step 1. Understand Prerequisites - Storj Docs . Does a raspberrypi fit those recommendations? Cool, use it.

As I said, I am sure it happens, just that I never saw it and a total failure happens way more often. It is more common to overload or short one phase and get the main breaker to trip disconnecting everything. At least before I upgraded the line to 20kW, now the main breaker is something like 42A so it’s hard to trip on overload, but before that I had a 10kW line, so it was really easy to trip the main breaker by putting too much load on one phase.

Last time I had a short, every breaker in line tripped. I’m pretty sure that’s supposed to happen. There are breakers with different curves, (C25 is slower than B25), but when a short happens it’s the electromagnetic coil that trips and they al trip at the same time (near instant).

Fuses are in devices, not in the panels. Some time ago there were no circuit breakers and everyone used fuses, but that was way less safe because when the fuse blows in the evening, people would just put wire on it. Everybody did that. If that blows, put a thicker wire. The first circuit breakers used the same sockets as fuses (so you could just screw it in), but now everyone uses DIN rail mounted breakers. I have talked to multiple electricians (both for home and for work) and nobody even suggested using regular fuses in the panels. It’s always breakers. So, I’m sure the code allows breaker-only installation.

I may have written that sentence incorrectly. Currently nodes are suspended only after they have been offline for way longer than 5 hours (i think it takes days to get suspended) and the network seems to work fine.

We are going off topic.

The fuses are not in the panels (not anymore, anyway). The fuses are always as close to the load as possible. In UK-based installations (ie not third world countries) there is the requirement for a fuse to be present in the plug you connect to the wall socket.

No, the breakers do not “all trip at the same time”. The fuse always blows first. Always. Even in third world countries that don’t require fuses in the plugs, there are fuses in the devices (as you have said). Don’t believe me? Open up any computer PSU (careful of the caps, the can store enough energy to hurt you). You’ll see the fuse right after the AC wires. It can be the normal wire one, or it can be the disc shaped one (the one that resets when you cut off all power). Then if the fault can’t be isolated by the fuse, the distinct circuit breaker trips. You know, the C25 (Type C, rated for 25A) you mentioned. Then (and only if everything else failed) the main breaker trips (ie if the main breaker is more sensitive, you correctly mentioned curves corresponding to different types, or if the short is upstream of a distinct circuit. That 25A rated one, go find its specs. You’ll see that it requires 3-5 times that current for “instant” trip (there is no such thing as instant in these scenarios, there is always delay). The main breaker is, you know, the 60A rated one. So, again, no they “don’t all trip at the same time”. Bonus point: Even if your main breaker got stuck and didn’t trip, the upstream fuse your electricity company requires you to install, will (and it’s the one they supply and seal shut when you have the installation inspected by them). Because as I have already said: there isn’t any technical scenario where a fuse will not work as expected. It’s the first and the last line of defense. The breakers are there to protect the electrical installation from catching fire, by you drilling holes in the walls (=to protect every step of the installation between the first and the last fuse).

No the in-panel fuses were not deprecated because people would put wires to bypass them. Fake? Lies? Read up where there is always a fuse waiting to protect you from your mistake, just upstream of where you are replacing that one. They were deprecated because people would stick their fingers in the naked electric contacts while trying to replace them (=instant death).

The network seems fine because the network considers any node offline longer than 5 hours as gone and start taking measures to protect the network (mark data unavailable, trigger repair if needed, as I said). The network isn’t relying on the node’s suspension to handle the abnormal situation. It’s relying on detecting the node as gone (>5 hours). Does that match the 99.3% uptime requirement? Yes it does actually, it’s the exact time (well, 6 minutes short) that 99.3% monthly uptime is needed.

I propose we get that changed. If a node is below 99.3% it needs to be suspended. If a node further falls below 50% it needs to be disqualified. You propose we don’t change anything. That’s OK, you can express your opinion and it will be a valued input to the discussion.

And going back on topic: since the network is behaving “fine” as is, there isn’t any reason to implement maintenance modes, correct? Because you know, that one week you need to replace that failed drive of yours, doesn’t actually have any negative impact on your node. The only negative impacts are when you cross the (around) two week mark. Your online score is below 60%, which does get the node suspended. If it doesn’t climb back up in a week (or in other words 3/4 of a month) then the node gets disqualified (allegedly).

I think that in that one time occurrence where you need to order and replace a drive immediately, 2 weeks to do so are more than enough. You even have one week extra in the current system. I propose we tighten that, just to keep you on your toes :smiley: .

Sorry if I’m being thick but…. has Storj started or even mentioned starting enforcing this 99.3% uptime requirement?

Last time I checked they were way more lenient than that and as far as I know this has had no deleterious effects on the network. Perhaps what we have is good enough and makes it a lot less “scary” for SNOs?

Seems like a “it’s not broken, no need to fix it” kind of situation…

Hasn’t changed as far as I’m aware. Suspension happens at 60% uptime. In my opinion the 99.3% SLA is set to make sure people who only plan for about 80% uptime don’t start nodes. As long as nodes generally stick to the SLA a one time drop isn’t going to hurt the network. But if people start caring for their nodes as if 60% is the SLA, they will need to start enforcing higher uptime requirements. So… as long as we all treat the 99.3% SLA as the guideline, we can enjoy more leniency if we fail to achieve that by exception. It’s not something you need to worry about if you run a stable always online system. Any unexpected failure leading to downtime will be survivable. Even if you can’t attend to it immediately. Just make sure you’re aiming for that 99.3%.

I have repaired a few PSUs and, yes, they have fuses. I have also encountered three failues where the PFC transsitor shorts, blows the fuse and trips the breaker. The complaint is usually “power went out and now the PC won’t turn on”. In order to avoid this type of failure, make sure the big high volage capacitor on the “hot” side of the PSU is high quality.

On an overload, a lower rated breaker will trip first. Yes. the C16 will trip before the C25 or the C45 outside. But a short is something like 1000A and will trip all of them at the same time.

There is no upstream fuse, not that I know of. There is a box outside (owned by the utility company, that’s why it’s outside, so they can check the meter any time they want), with the meter and a C45 (I do not remember exactly, but between 40 and 45, the line is supposed to be 20kW) breaker, then a cable goes to my house and it’s all breakers inside.

Before I upgraded it, the line was 10kW, so the main breaker was C25 (IIRC), inside the house breakers were C25 or C16 (depending on the wires), so it was really easy to trip the main breaker by overloading one phase (devices connected in different rooms, so from different 16A breakers and turning on the kettle and microwave while a miner is running would result in a trip). Later I got a big UPS that can distribute the load on the input and it solved the problem, even later I got the utility company to upgrade the line to 20kW which involved them replacing the breaker (as the cable was already big enough).

I recently got a generator wired up and there are no fuses there either, just breakers. Maybe the requirements are different in the UK compared to Lithuania. We do not have fuses in the plugs too, the current code is to protect outlets with 16A breakers (AFAIK). On older installations you may find a bigger breaker if there are lots of outlets connected to it, but now it is different.

The difference is that if the node comes back online, it is not suspended. Also, it’s 5 hours continuous, so if the node is offline for 2 hours today and 3 hours a week later, the data will not be marked as gone.

I don’t really see the need for suspension this fast. And yes, there have been many discussions about this and I always was of the opinion “if you say this is for regular people, do not expect datacenter reaction times”. I’m sure Storj Select node operators can react quickly.

With the times that they are now, correct. The only reason for a maintenance mode would be to inform the satellite that the node is going to be offline longer, so it does not need to wait the 5 hours to start repairing data.

If the network is running well as it is, then why tighten it? While I keep a spare drive, it’s on a shelf so I would need to be physically present to replace it and it may take longer than 5 hours.

Sure, that’s what I do, just that while I do not have frequent problems I am fully aware that I may have an infrequent problem that takes time to fix. Just like my reason for getting a generator - it’s not the frequent short term outages I am worried about (I have multiple UPSs for those), it’s the big outage once every few years that lasts many hours. Just because I have not had any outages in a while does not mean it is impossible to drain the oil from the transformer.

Last reply from me as far as the electricity stuff goes:

“That you know of” being the key phrase. What happens when you short something between the meter and your first breaker? The entire neighborhood goes down while we wait for someone to go and flip the breaker on the side of the power company? No. We blow the fuse that is upstream of your meter, just after the supply line comes in from the utility company. This is the fuse that is sealed shut with a tamper-proof seal when your electrical installation is inspected (see previous reply).

It’s not like that. The meter is in a completely separate box, outside my property and it is considered the property of the utility company, I cannot get inside it, other than to lok at the meter and to flip the breaker. The box also has a meter for my neighbor. Part of the box is transparent (so I can see the meter) and there is no fuse visible inside, just incoming cable, the meter, the C40 breaker and outgoing cable. There could be a fuse somewhere I couldn’t see. If there is, the fuse is rated for much higher current or is much slower than the breaker because a dead short inside my house trips every breaker in line, but does not blow the fuse.

My UPS has fuses though, rather expensive ones at that.

I wonder if this would be even necessary. There is an incentive to keep uptime high: more uploads, less files lost to repair. Would have to run some simulations, I guess, and data longevity is a factor here, but I have a feeling we don’t really need uptime SLA that much as long as getting to store data has a decent incentive.

Well a while back someone asked about turning the node off when they go to sleep in order to avoid the noise. You can probably stay above 60% that way, but if enough people start doing that and simply not caring about the other incentive, at some point they will have to enforce stricter rules. I’d like to stay ahead of that as a community and just aim for that 99.3% SLA, even if we sometimes incidentally can’t entirely make it.

The node uptime does not matter as much as the uptime seen by the downloader. Assuming you need at least 29 pieces out of 80 to reconstruct, you only need a node uptime probability of 50% to ensure that 99.5% of the time the end user will be able to download.

At 60% this is 99.999%.

Then just add some additional safety factor because these nodes are not necessarily independent.

I am however interested how many nodes in the network actually have a consistent uptime of 99.3% or higher.

I suspect you might be pleasantly surprised.

I have multiple nodes in a fairly unexciting setup with domestic power and domestic broadband and it’s rare for any of them to be below 99% for a very long period of time.

It does happen, of course (mostly at my dad’s because they just can’t seem to be able to leave things alone!), but I’ve found that if I don’t mess with them too much and just mostly leave them to do their thing they tend to be very stable :slightly_smiling_face:

I second that. Ironically, my 9 month streak of 100% uptime on the node in the closet on the patio was interrupted today by power company having to fix the pole after a tree fell on it. The outage lasted two hours, resulted in 90 min node downtime.