It seems that the new feature “save-state-resume GC filewalker” isn’t functioning as expected

And it seems that it is progressing:

https://review.dev.storj.io/c/storj/storj/+/12806?tab=comments

Overall it looks good to me.

@elek

That would make it better with frequent restarts. Even if one round is finished, we may not need to re-calculate the usage if it was calculated recently…
Why is there no setting to control the frequency?

I see following cases:

  • Frequent restarts by SNO: If a node operator is testing a setting or something like that and restarts the node several time in a short period, I do not see the need for the filewalker to run each time.
    This could be handled by a delayed start option. But the biggest painpoint was indeed that it would start from the beginning every time and this should go away. At least it would resume and come to an end at some point.

  • Errors: When there is an error during filewalking it is questionable if it should restart. Chance is that the error repeats. However it would depend on the kind of error. If it is simply something like memory starvation or slow disk, I don’t see a reason why it should not restart and resume calculating. As it would not start from beginning even if it fails again it should move forward and come to an and at some point as well.

  • Successful completion: When it has completed successfully there is normally no urgent need to run it again soon. AFAIK once we have current data this gets cached and updated by the actual up- and downloads. So a configurable setting when to run again after the last successful completion would help to prevent unnecessary filewalker runs. Maybe it would be enough to run it once a week only if we have actual data from a previous run.