If this problem started with the move to bare metal, could something in the config there be the culprit? Undersized NIC buffers, or the conntrack table overflowing. Or upstream equipment dropping traffic for similar reasons? Might be a dirty fiber as well.
Yes, these are good candidates, but we have already checked all of these. As I wrote: the problem is under active investigation.
I saw this PR: https://review.dev.storj.tools/c/storj/storj/+/22617
It should fix the issue
seems it does , i already seen online status increase by couple 0.0x %
upd: prolly i was too fast to celebrate , its still dropping
Now I’m starting to worry too, all my nodes have dropped below 95% in the past few weeks.
are you using VPN or VPS to route traffic?
I have three nodes. One in my house with direct internet access, two in other apartments where the internet is under cgnat. Each of them accesses the internet through a separate vpn->vps provider. The containers running on the Synologys build a vpn tunnel and go to the internet with the vps ip.
The score of all three nodes is constantly decreasing.
This coloring code is a copy-paste of audit/suspension scores, however, the online score has different thresholds - below 60% - should be yellow (the node is suspended), below 25% - should be orange (close to disqualification), 0% - should be red (disqualification).
I have a problem since about the the 16th of Juli on all Satelites and on two different Hosts with separat Internet Connections.
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-08T00:00:00Z 86 86 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-08T12:00:00Z 46 46 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-09T00:00:00Z 58 58 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-09T12:00:00Z 48 48 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-10T00:00:00Z 54 54 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-10T12:00:00Z 104 104 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-11T00:00:00Z 60 60 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-11T12:00:00Z 55 55 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-12T00:00:00Z 56 56 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-12T12:00:00Z 61 61 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-13T00:00:00Z 101 101 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-13T12:00:00Z 39 39 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-14T00:00:00Z 52 52 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-14T12:00:00Z 101 62 39 61.39
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-15T00:00:00Z 45 45 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-15T12:00:00Z 75 75 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-16T00:00:00Z 85 73 12 85.88
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-16T12:00:00Z 40 39 1 97.5
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-17T00:00:00Z 62 60 2 96.77
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-17T12:00:00Z 106 102 4 96.23
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-18T00:00:00Z 59 57 2 96.61
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-18T12:00:00Z 95 89 6 93.68
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-19T00:00:00Z 81 76 5 93.83
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-19T12:00:00Z 48 48 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-20T00:00:00Z 50 49 1 98
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-20T12:00:00Z 76 74 2 97.37
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-21T00:00:00Z 97 93 4 95.88
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-21T12:00:00Z 65 60 5 92.31
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-22T00:00:00Z 62 57 5 91.94
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-22T12:00:00Z 44 42 2 95.45
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-23T00:00:00Z 75 69 6 92
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-23T12:00:00Z 90 83 7 92.22
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-24T00:00:00Z 74 71 3 95.95
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-24T12:00:00Z 55 52 3 94.55
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-25T00:00:00Z 64 62 2 96.88
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-25T12:00:00Z 50 50 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-26T00:00:00Z 79 74 5 93.67
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-26T12:00:00Z 74 70 4 94.59
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-27T00:00:00Z 58 57 1 98.28
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-27T12:00:00Z 51 51 0 100
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-28T00:00:00Z 73 71 2 97.26
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-28T12:00:00Z 88 83 5 94.32
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-29T00:00:00Z 63 58 5 92.06
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-29T12:00:00Z 101 94 7 93.07
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-30T00:00:00Z 95 89 6 93.68
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-30T12:00:00Z 118 110 8 93.22
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-31T00:00:00Z 81 78 3 96.3
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-07-31T12:00:00Z 57 54 3 94.74
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-01T00:00:00Z 63 61 2 96.83
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-01T12:00:00Z 51 47 4 92.16
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-02T00:00:00Z 90 85 5 94.44
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-02T12:00:00Z 64 62 2 96.88
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-03T00:00:00Z 51 49 2 96.08
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-03T12:00:00Z 60 56 4 93.33
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-04T00:00:00Z 82 76 6 92.68
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-04T12:00:00Z 174 163 11 93.68
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-05T00:00:00Z 138 130 8 94.2
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-05T12:00:00Z 142 135 7 95.07
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-06T00:00:00Z 159 153 6 96.23
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-06T12:00:00Z 139 132 7 94.96
1wFTAgs9DP5RSnCqKV1eLf6N9wtk4EAtmN5DpSxcs8EjT69tGE 2026-08-07T00:00:00Z 87 81 6 93.1
Does someone have the same Problem? The logs say that there is a ping timeout. But there is no reason for this as i can see.
sounds like: Dropping online status
BTW:
I don’t say they did it intentionally but… As a result of this bug my payment stats are down like 10..15%. ![]()
Unlikely. Payment estimates use average disk usage data, which hasn’t yet been fixed on the node side. storagenode: normalize missing storage usage for dashboard estimates by Aleksman4o · Pull Request #7806 · storj/storj · GitHub is ready; when the author rebases, it can be merged. Then it will be included in the next release.
Also, all statistics are reset at the beginning of the month, as usual, and this has been the case for several years now.
Do we know when will the fix be deployed?
Will online status be retroatively adjusted or well just have to wait for a month?
Usually - once a week, but sometimes there are point releases.
I usually checks by the tag.
I took a hash 73c543d796ea2ef5f557222e432ccd15268b05cd, then search for it on GitHub, it’s here contact: use WithForceDial for PingBack and PingMe dials · storj/storj@73c543d · GitHub
Then I take a look on tags, there are no tags, only main, so it’s not released yet.
Perhaps it would be released on the next week, like on 2026-08-09T17:00:00Z (don’t check the time, it’s not a promise, it’s an estimate
)
Dial-back handshakes fail 100% of the time when the satellite’s own source port is in 30000–32767 - 60/60 across 36,287 captured handshakes
Since 2026-07-18 one of my nodes (`1LdKrWaHJ4EjkjhNtQsaBVrpxAcPYFxFH2BATPuZRbGkgevHRm`, v1.155.9, public address `198.52.189.168:28967`) loses 3–7 check-ins a day. I captured 36,287 inbound dial-back handshakes (SYN/SYN-ACK only), and every failure shares one signature: a dial-back from your satellite hosts in 79.127.163.0/24 and 79.127.205.0/24 never completes when the satellite’s own TCP source port lands in 30000–32767 - 60 of 60 such flows failed, 0 of the other 36,227 did. The SYN arrives, my node answers in ~1 ms with the correct `ack`, and the satellite retransmits the same SYN 6–7 times as if it never saw the reply, then times out.
| Source | Source port | Flows | Failed | Rate |
|---|---|---|---|---|
| Storj sat (79.127.163.0/24 + 79.127.205.0/24) | 30000–32767 | 60 | 60 | 100.00% |
| Storj sat (79.127.163.0/24 + 79.127.205.0/24) | anything else | 5,068 | 0 | 0.00% |
| every other network | 30000–32767 | 133 | 0 | 0.00% |
| every other network | anything else | 31,026 | 0 | 0.00% |
(A flow counts as failed when my SYN-ACK is captured leaving intact but the peer keeps retransmitting the same SYN rather than completing or sending a RST.)
It splits by host, not by prefix. Within 79.127.205.0/24, three hosts - 79.127.205.225, 79.127.205.226 and 79.127.205.227 - never failed across 3,404 flows, and never once dialed from a port in 30000–32767. Their siblings (79.127.205.251 and 79.127.163.228 among them) failed at ~5%, exactly on the connections that did draw an in-range port: 45 of 907, all 45 failed. A ~5% draw at 100% failure is a ~5%/day check-in loss, which is an online score settling near 0.95 - where mine now sits, no longer falling.
| Satellite host | Flows | Failed | Rate |
|---|---|---|---|
| 79.127.205.225 | 1,126 | 0 | 0.00% |
| 79.127.205.226 | 1,214 | 0 | 0.00% |
| 79.127.205.227 | 1,064 | 0 | 0.00% |
| 79.127.205.236 | 61 | 1 | 1.64% |
| 79.127.205.241 | 74 | 1 | 1.35% |
| 79.127.205.242 | 75 | 4 | 5.33% |
| 79.127.205.249 | 135 | 7 | 5.19% |
| 79.127.205.250 | 155 | 7 | 4.52% |
| 79.127.205.251 | 126 | 9 | 7.14% |
| 79.127.163.228 | 92 | 7 | 7.61% |
| 79.127.163.229 | 89 | 4 | 4.49% |
| 79.127.163.230 | 100 | 5 | 5.00% |
The top three (.225/.226/.227) are the ones that also never drew an in-range source port; every other host both drew from that range and failed exactly on those draws.
My side is ruled out: other networks dialing from the same range are fine (133 flows, 0 failures), the path on my end is plain L2, and my own 223 NodePorts aren’t involved (57 of the 60 failures hit ports where I run no service). It also isn’t `73c543d`'s pooled-dialer bug - these are fresh SYNs retransmitted 6–7×, not a reused dead connection.
The lead, and why I’m posting: 30000–32767 is the default Kubernetes NodePort range (`–service-node-port-range`). If the failing hosts are k8s nodes whose `net.ipv4.ip_local_port_range` dips below 32768, ~5% of their outbound connections would draw a source port that also matches a local NodePort, and the return SYN-ACK - arriving with that as its destination port - could be DNAT’d away from the socket in `SYN_SENT` (kubernetes/kubernetes #111823 and #111697). An ephemeral range like `10000 65535` gives 4.98%, close to the observed 5.0%. That fits it being per-host, invisible from outside, and unaffected by `WithForceDial`.
I could easily be wrong here - I can’t see your side. But if it’s quick, comparing `sysctl net.ipv4.ip_local_port_range` on a failing host (say 79.127.205.251) with a clean one (79.127.205.225), and against `–service-node-port-range` if those are k8s nodes, might settle it. If it pans out the fix would just be a sysctl, and it’d help every node those hosts dial. Either way I’m glad to keep digging from my end - happy to share the pcaps, the per-flow census, or run a timed capture against any window and satellite you name.
Thanks the detailed analyses. It’s really helpful.
For the records: In you table there are satellite and edge gateway servers as well. 79.127.205.251 is satellite and 79.127.205.225 is gateway. So you also see the traffic of the real upload/downloads.
I already checked net.ipv4.ip_local_port_range (I had some very bad errors on Hetzner, years ago, due to similar problems):
sysctl net.ipv4.ip_local_port_range
net.ipv4.ip_local_port_range = 32768 60999
This looks good to me, but your analyses certainly helps. Doing more tcpdump on our side.
Additional info: I can reproduce the problem on our side. During the last week I executed multiple nodespacescan (connecting to each storagenode manually), and analyzing data.
FTR: We are starting reset online scores + removing DQ flags from nodes which were disqualified during last 2 weeks due to low online score.
(As there are many layers of caching, this can be bumpy, but eventually it will happen).

