Ok got it. It will be a bit short notice. I expect just a few days between signing the contracts and the uploads hitting our nodes. So with the surge nodes we can buy time for the network to adopt to the new situation.
Yes. I am a node operator myself and fully agree on that. I have communicated that point to the stakeholders already. If we are lucky they might make a commitment to take the risk of no deal signed from the table. If the deals are getting signed we get a higher payout and if the deals are not getting signed we also get a higher payout. That would also allow us to start growing in advance. That commitment isn’t on me so lets wait and see how it will look like.
Could you please open a new thread for that and ask [@] elek how he did it? To my knowledge he managed to get a storage node down to a few minutes file walker with an inode cache. I don’t know the details but he can share his knowledge.
I have this exact same problem on some nodes only:
My trash is empty but its not updating the trash to the node, so at the moment is thinking i have 6.5 TB in trash. Im monitoring the logs now and lets see what it says. But nodes on the same storage does not have this issue. And filewalkers are finished successfully.
Thank you. I have a few old 6TB drives that I’ve spun up and got ready to start new nodes when you pull the trigger. If they fill up I’ll buy some decent SAS ones for more permanent use.
Will there be any changes to vetting in order to speed up the process of bringing new nodes online? I think you alluded to that earlier on but not sure if that is actually “a thing”.
No you don’t, that’s a different problem. I mean you might have the same problem as well in addition to that, but the dashboard won’t show you if you do.
I assume you have considered the surge nodes to also only be chosen for downloads if there’s not enough regular nodes? Whatever the place you will pick to host the surge nodes, this would reduce your egrees while maybe making SNOs a little bit happier (at least those who would not complain about their egrees).
It sounds like a lot of SNOs need to put their money where there mouth is. Lots of talk of “when I fill… I’ll expand”, so we’ll see if that’s true. Because that’s over 10% of online nodes: if they stop accepting ingress it’s definately significant.
Do you mean the Uptime on the dashboard? If it’s reset, then probably the node is restarted. Please check your logs (both - for storagenode and storagenode-updater), why is it restarted? The last message before a restart should have some explanation (unless you have had a power cut for a moment… but in this case dmesg or journalctl should show something).
Referring back to this, currently I have started seeing a large amount of the test data get garbage collected:
This might give insights to the issue you were diagnosing with the dashboard. To my untrained eye, it looks like the TTL data uploaded is not being correctly registered as “active” data, and therefore is being removed by GC before TTL kicks in.
I will probably expand by adding 20-30 nodes in the near future. Nothing close to 1000 of course but I think other people will add capacity as well. It takes a little bit of time though.
Right now I’m just trimming away some free space on existing nodes.
I still don’t understand why the dashboard can not provide reliable data. As a customer it is good to have good and right data. And we as node operator are in my opinion the same as paying customers.
I really don’t want to complain but the stats are almost every month wrong from the satellite. The public status page is inconsistent and so on. Why is that so? And how do we know if the satellite knows the right data? And why is it sending wrong/no data to the storage nodes if it knows the real value?
I really like the project and will continue to support it, but i wasn’t able to find clarification for that. All that I was finding is “The satellite knows the right data and you will get paid the right amount”.
SNOs are service providers: we don’t pay Storj like customers: we get paid. And payouts have been correct month-after-month. If we don’t feel we’re getting paid adequately: we stop providing our service (and find another project)
I have the same problem: 71 TB in trash and 166 TB used. Ext4, 2 nodes on a single HDD (I know that in the “new Storj reality” we shouldn’t use more than 1 node per HDD, but we’ve been here for years and we have what we have). It’s been a while since the last time my nodes completed GC on the US satellite, which is why I have had so much data that should be removed. What I really don’t understand is why the devs took so long to roll out version 1.105.4 to Docker until yesterday (this version fixes a critical bug with removing retain files when the node is restarted). It is obvious that many SNOs restarted their nodes while tests are running to adapt to increased data flow and lost retain files because of the bug. If there is someone like me with a lot of unremoved trash (I mean in used, not in the trash directory) and still not updated manually to 1.105.4 (I did it a couple of weeks ago), then it will be another month before their nodes clean themselves up. If I understood @littleskunk correctly, this may decrease network throughput.