I’m having a problem with my online status. As far as I can tell, the satellites haven’t been able to ping me since July 15. I’m getting these messages in the logs:
I deliberately didn’t change anything in my settings. Does anyone have any idea what might be causing this?
Since I unfortunately have a CGNAT, the node runs through a VPN.
Data is coming through, and the port is open, but a real ping times out. But as far as I know, it’s always been like that.
I’m also seeing something similar - all my 4 nodes have had a steadily decreasing online score for about two days now, yet Uptimerobot records no outage, GRC shields up show the ports as open, and piece up/downloads are happening at a normal rate with normal (for me) success rates
On my nodes, all four satellites consistently mark about 5% of audits offline, while the logs show satellite callback TCP timeouts. This is a satellite-side issue.
That pool is keyed by node ID, not TCP vs. QUIC or even the requested address. So a QUIC request can get a cached TCP connection, and a request to a new endpoint can get a connection to the old one.
That’s broken, and it’s where I’d look first if it was my job.
This started with the AWS routing outage last week, routing has been crazy since. They say it is fixed but its’s not.
“Last week (July 16, 2026), AWS suffered a 3.5-hour global CloudFront outage. The glitch caused widespread 5xx errors for CloudFront customers using VPC Origins, taking down websites like Hugging Face and the UK National Lottery. The issue was traced to a configuration routing failure and fully resolved by 4:18 AM PDT.”
It was worse than this note says, as many major government, bank and other businesses websites went down for many hours. This is in Europe I don’t know about other parts of the world.
The routing is still partly broken, for me in Sweden the online percentage for all my nodes has been slowly dropping since the outage, the same for the two nodes I have in Thailand.
It is nothing I can do more than watching the Online percentage slowly counting down on all nodes.
Are whole nodes affacted with the problem, or just some percentage of the nodes? Any idea? and does changing the ip address or reestablishing a connection with satellites somehow, may fix the problem? Anything we can do?
Since around the routing problems last week, all of my Storj nodes have been gradually losing Online Score on every satellite.
Current status:
Suspension: 100%
Audit: 100%
Online: between 91% and 94%, depending on the satellite.
In addition, my nodes have had virtually no uploads or downloads for several days. They appear to be reachable:
All required ports are open and reachable from the Internet.
I have verified external connectivity from multiple locations.
The nodes are running and responding.
However, the dashboard reports Offline and QUIC Misconfigured, and traffic has almost completely stopped.
I’m also seeing monitoring issues such as HTTP 500/502 errors from services like SiaScan, so I’m wondering if there is a broader networking or routing problem affecting connectivity between satellites and storage nodes.
I noticed declining online scores but no ingress/egress decline or quic errors. In fact I see highest ingress ever. I guess this is because many fast nodes are full now because of no bloom filters.
I’ll share some additional information in case it helps. My storage node is located in AS16276. Satus: Online, QUIC: OK. I reduced the storage size a few months ago, so “ingress” is 0, “egress” is within the normal state for my node.
This is mean that your node is definitely offline, you need to fix this issue. You may check your address with a port from your mobile phone via the mobile internet, if it doesn’t respond, your node is offline. http://node_external_address:node_port
No word from Storj on this ? Is it an routing issue or is it a storj issue ?
All my nodes in Sweden and my nodes in Thailand has slowly been ticking down from 100% to under 99% today (1% in one day) but no problem yet, it is a long way to 60% which I believe is the minimum and I get kicked out.
if I do a simple mtr but use TCP not icmp to one of the satellites,
It finds the way but twelve99.net seems to have issues, packet losses and extreme response times, maybe that causes some ‘node up’ checks to get lost? It can’t be all ‘node up’ checks getting lost because when it would be ticking down much faster and I would receive mail that my node is down.
If I do the same mtr test but to google on port 443 I have 0% loss and response times between 2ms - 4ms
PS: Since the AWS routing crash I noticed that my nodes in Thailand do not go east to connect to the US satellites, they go west with RTT around 330ms !