# It seems that the new feature “save-state-resume GC filewalker” isn’t functioning as expected

**URL:** <https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874>\
**Category:** troubleshooting\
**Tags:** filewalker\
**Created:** [April 16, 2024, 5:47pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874 "2024-04-16T17:47:00Z")\
**Posts on this page:** 20\
**Page:** 3

<div class="post-metadata">

**Author:** ![pdeline06](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/pdeline06/32/1129_2.png) [@pdeline06](https://forum.storj.io/u/pdeline06)\
**Post date:** [June 10, 2024, 5:59pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/42 "2024-06-10T17:59:14Z")

</div>

“save-state-resume feature for used space filewalker” (resume from the last point it stopped) - is it already working correctly?

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 10, 2024, 6:20pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/43 "2024-06-10T18:20:44Z")

</div>

As far as I see it, it has not been merged:

[https://review.dev.storj.io/c/storj/storj/+/12806](https://review.dev.storj.io/c/storj/storj/+/12806)

It seems new features are getting added:

[https://review.dev.storj.io/c/storj/storj/+/13421/2](https://review.dev.storj.io/c/storj/storj/+/13421/2)

---

<div class="post-metadata">

**Author:** ![pdeline06](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/pdeline06/32/1129_2.png) [@pdeline06](https://forum.storj.io/u/pdeline06)\
**Post date:** [June 10, 2024, 7:36pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/44 "2024-06-10T19:36:08Z")

</div>

> [@jammerdan](#):
>
> As far as I see it, it has not been merged:
> 
> [https://review.dev.storj.io/c/storj/storj/+/12806](https://review.dev.storj.io/c/storj/storj/+/12806)

but build verify is finished successfully…need waiting for merge?

> [@jammerdan](#):
>
> It seems new features are getting added:
> 
> [https://review.dev.storj.io/c/storj/storj/+/13421/2](https://review.dev.storj.io/c/storj/storj/+/13421/2)

This makes sense if 1 point is working

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 11, 2024, 4:59am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/45 "2024-06-11T04:59:45Z")

</div>

> [@pdeline06](#):
>
> This makes sense if 1 point is working

I wish they would have released it before adding new features.  
I am waiting desperately for this feature for weeks now for my nodes where I had to turn the filewalker off because it would not finish.  
I have no information about the size on them currently.

---

<div class="post-metadata">

**Author:** ![elek](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/elek/32/9189_2.png) [@elek](https://forum.storj.io/u/elek)\
**Post date:** [June 11, 2024, 7:43am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/46 "2024-06-11T07:43:21Z")

</div>

> [@pdeline06](#):
>
> but build verify is finished successfully…need waiting for merge?

The process:

1. build-verify → this is like a quick (?) smoketest
2. two :+2: reviews
3. build-premerge (reamaining, long-runnint integration test)
4. merge

---

<div class="post-metadata">

**Author:** ![elek](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/elek/32/9189_2.png) [@elek](https://forum.storj.io/u/elek)\
**Post date:** [June 11, 2024, 7:56am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/47 "2024-06-11T07:56:40Z")

</div>

> [@jammerdan](#):
>
> I wish they would have released it before adding new features.  
> I am waiting desperately for this feature for weeks now for my nodes where I had to turn the filewalker off because it would not finish.  
> I have no information about the size on them currently.

I totally agree with you, we need this feature, but I couldn’t resist to add my wider notes 😉

1. The problem with filewalkers is retrieving the size and last modification time. With ext4 file system, directory entries contain only the file names. An additional read of the inode is required for getting the size and last modification time of the files.

2. Assuming you have 10M pieces, it’s at least 10M additional IO operation. With 200 IOPS/sec, it’s 10M / 200 / 60 / 60 = 13 hours to complete

3. This particular feature (save-state-resume GC file walker) can help to survive restarts, but won’t solve fully the problem

4. We write pieces files only once. It should be possible to cache size/modification time information

5. If you have enough memory, OS can also cache. I have experience with a server with 126Gb memory. Because majority of the ram is not used, Linux is happily uses 66Gb to cache inodes → the filewalker can finishes in 5-20 minutes, even if I have ten millions of files. (That’s the easiest workaround. If you have memory, just put it to the server)

6. @littleskunk reported similar results to use RAM (and SSD) for ZFS cache.

7. But let’s say you don’t have enough ram. I am doing some experiments, and I found that a simple db based cache still can help (49m walking for 14M pieces with hot cache. With cold cache it was 15h, without cache it was 13h)

My personal opinion is that walkers should be improved (in addition to save-state-resume), and make them faster with caching OS level size/creation time.

---

<div class="post-metadata">

**Author:** ![Roxor](https://storj.bcdn.literatehosting.com/letter_avatar/roxor/32/5_5575768a8748004e209b776fc1b2916d.png) [@Roxor](https://forum.storj.io/u/Roxor)\
**Post date:** [June 11, 2024, 8:34am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/48 "2024-06-11T08:34:11Z")

</div>

> [@elek](#):
>
> My personal opinion is that walkers should be improved (in addition to save-state-resume), and make them faster with caching OS level size/creation time.

Are those IO constraints related to the max-space-per-node recommendation ([24TB](https://docs.storj.io/node/get-started/prerequisites))? Like if you have 3-4mil small files per-TB… at some point the filewalker-type housekeeping tasks just have too many files to check all the time?

---

<div class="post-metadata">

**Author:** ![elek](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/elek/32/9189_2.png) [@elek](https://forum.storj.io/u/elek)\
**Post date:** [June 11, 2024, 9:23am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/49 "2024-06-11T09:23:23Z")

</div>

> [@Roxor](#):
>
> Are those IO constraints related to the max-space-per-node recommendation ([24TB](https://docs.storj.io/node/get-started/prerequisites))?

It’s a different type of recommendation, I guess. With 24TB you may have other limits as well.

This problem depends on the disk IO + file system. You can be a happy owner of a 24TB with good disk (for example with RAID0, you may have better IOPS bandwidth). Or using ZFS + SSD cache.

I have also seen dozens of storagenodes running on a single 1.7TB NVMe disk (QA satellite only for testing), without any io pressure.

I would re-phrase recommendation like this:

1. don’t use more than 24TB for one Storagenode instance (as written in that page). May work, but it’s not really tested…
2. as a rule of thumb: use one SN process per disk
3. maximum space should depend on the max IO speed. (If you can use only slow HDDs, it can be better to use more, but smaller disks)
4. If you have big and slow disks, it’s better to have huge memory (at least for ext4)

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 12, 2024, 2:49am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/50 "2024-06-12T02:49:22Z")

</div>

> [@elek](#):
>
> but I couldn’t resist to add my wider notes

If it helps with restarts better, fine. Also making it configurable is good. I had my own thoughts on that [here](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/27) and [here](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/29) , I don’t know if @clement is aware of them.

But also it would be important to get the base stop-resume feature out and add the improved features later.  
It is important to get to a state to receive reliable information from the node which is currently not the case in many areas and different reasons unfortunately:

- Used space not correct
- Trash folder not updated
- Bandwidth display not correct

Then we see glitches on satellites, avg. used values not consistent etc.  
I know some of it has already been fixed but it the fixes did not yet arrive on Docker nodes.

Currently the situation is that the used space filewalker would not complete on some nodes which means that it continuously restarts and retries from the start. This is why I need the feature to get it done and the used space updated at least once.  
I agree that it does not really solve the underlying problem that is too much IOPS needed for those operations. That is also why I made my previous suggestion not to always repeat the filewalker on restart when it is not necessary but make it configurable when it runs.

> [@elek](#):
>
> If you have enough memory, OS can also cache.

When you are referring to inode cache do you mean `vfs_cache_pressure`? I see that it can influence the tendency what the OS caches but I don’t know how useful is that. Is it better to cache inodes or pieces for the customer? As I understood it, the used-space filewalker is basically a one time thing at least as long as the node is running so basically “wasting” cache for that instead for pieces data to serve to customers sounds wrong. But maybe caching inodes also helps with the pieces for customers, I don’t know, maybe…

> [@elek](#):
>
> But let’s say you don’t have enough ram. I am doing some experiments, and I found that a simple db based cache still can help

Of course if there are ways to cache better or more, this would be good.

> [@elek](#):
>
> We write pieces files only once. It should be possible to cache size/modification time information

At least it sounds like this is something that could be doable, to not rely on inode as we basically have all the data from the moment when a piece gets written and to store this information for filewalker use. Maybe this information could be even moved to a different disk or even ramdisk then, freeing the data disk from inode reads for that purpose altogether.

> [@elek](#):
>
> My personal opinion is that walkers should be improved (in addition to save-state-resume), and make them faster with caching OS level size/creation time.

Together with database I could not agree more.

---

<div class="post-metadata">

**Author:** ![Alexey](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/alexey/32/41_2.png) [@Alexey](https://forum.storj.io/u/Alexey)\
**Post date:** [June 12, 2024, 6:09am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/51 "2024-06-12T06:09:39Z")

</div>

> [@jammerdan](#):
>
> Maybe this information could be even moved to a different disk or even ramdisk then, freeing the data disk from inode reads for that purpose altogether.

This is already possible for zfs for example, exactly what you suggest. It wouldn’t be a ramdisk directly, but RAM (+SSD?) used to cache this metadata.  
It’s partially possible for ext4 too:

> [@Metadata cache ext4 (RAM maxed)](https://forum.storj.io/t/metadata-cache-ext4-ram-maxed/26216/5):
>
> I don’t think you can directly control what content is cached in RAM in ext4 as you can do in ZFS. Indirectly you can tune the vfs\_cache\_pressure parameter. As mentioned earlier someone used LVM to achieve storing metadata on SSD: [Ext4 speedup by storing metadata and data on separate devices](https://linux-ext4.vger.kernel.narkive.com/66JGZzf0/ext4-speedup-by-storing-metadata-and-data-on-separate-devices)

---

<div class="post-metadata">

**Author:** ![Mitsos](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/mitsos/32/1125_2.png) [@Mitsos](https://forum.storj.io/u/Mitsos)\
**Post date:** [June 12, 2024, 6:19am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/52 "2024-06-12T06:19:55Z")

</div>

Cache isn’t just used once, it’s used every time the file(metadata) is needed. That could be every time the GCs run for example.

There is no point in caching pieces for customers. Resources are better spent on caching information about the millions of pieces that the node is storing, speeding up any internal proccess (used space, GC, trash-cleanup).

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 12, 2024, 6:24am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/53 "2024-06-12T06:24:10Z")

</div>

> [@Mitsos](#):
>
> Cache isn’t just used once, it’s used every time the file(metadata) is needed. That could be every time the GCs run for example.

Right, I forgot about the other filewalkers.

> [@Mitsos](#):
>
> There is no point in caching pieces for customers.

What `vfs_cache_pressure` value do you suggest?

---

<div class="post-metadata">

**Author:** ![Mitsos](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/mitsos/32/1125_2.png) [@Mitsos](https://forum.storj.io/u/Mitsos)\
**Post date:** [June 12, 2024, 6:42am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/54 "2024-06-12T06:42:01Z")

</div>

> [@jammerdan](#):
>
> What `vfs_cache_pressure` value do you suggest?

I don’t change it for my systems.

---

<div class="post-metadata">

**Author:** ![elek](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/elek/32/9189_2.png) [@elek](https://forum.storj.io/u/elek)\
**Post date:** [June 12, 2024, 6:52pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/55 "2024-06-12T18:52:58Z")

</div>

> [@Mitsos](#):
>
> I don’t change it for my systems.

Neither me, but without changing, I surprised that Linux did exactly what I wish:

Having lot’s of unused memory:

 ![image](https://storj-s3.bcdn.literatehosting.net/original/3X/d/5/d5bf2f5a209bd2b04e8e2c6fa068fff717e2673f.png)

And significant part of the cache (yellow lines) spent to cache the inodes (ext4\_inode\_cache):

 ![image](https://storj-s3.bcdn.literatehosting.net/original/3X/6/e/6e18edae7b2c93ecb0eff0137f0845cfdcf0ed43.png)

(first is from `htop` output, second is from `slabtop`)

---

<div class="post-metadata">

**Author:** ![Mitsos](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/mitsos/32/1125_2.png) [@Mitsos](https://forum.storj.io/u/Mitsos)\
**Post date:** [June 12, 2024, 6:58pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/56 "2024-06-12T18:58:29Z")

</div>

When in htop: press F2 (setup) \> down arrow to meters \> right arrow until you get to memory \> press space. it should be like this:  
 ![Screenshot_20240612_215737](https://storj-s3.bcdn.literatehosting.net/original/3X/b/b/bb852fa701d31c70f39e865819347c3a5b608f2f.jpeg)

This is the result:  
 ![Screenshot_20240612_215808](https://storj-s3.bcdn.literatehosting.net/original/3X/f/e/fe45964e3373e3d26adeb4fb10d4425487960c82.jpeg)

Edit: because discourse is doing its thing where replied-to gets evaporated into thin air, @elek

---

<div class="post-metadata">

**Author:** ![Alexey](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/alexey/32/41_2.png) [@Alexey](https://forum.storj.io/u/Alexey)\
**Post date:** [June 13, 2024, 7:34am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/57 "2024-06-13T07:34:34Z")

</div>

2 posts were split to a new topic: [Discourse vape out topics](https://forum.storj.io/t/discourse-vape-out-topics/26611)

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 13, 2024, 7:32am UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/58 "2024-06-13T07:32:25Z")

</div>

> [@elek](#):
>
> I totally agree with you, we need this feature,

I see things are going to be discussed: [Storage node performance for filewalker is still extremely slow on large nodes · Issue #6998 · storj/storj · GitHub](https://github.com/storj/storj/issues/6998)

Please do not delay this feature any further even if you might agree on other solutions to improve the filewalker situation.

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 13, 2024, 12:02pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/59 "2024-06-13T12:02:09Z")

</div>

> [@elek](#):
>
> but won’t solve fully the problem

Here is one problem I have now as I had to turn off the regular filewalker:

> [@Overriding available space?](https://forum.storj.io/t/overriding-available-space/26621):
>
> I had to disable the filewalker on this node as it would not finish. Now of course the used space is not correct and I cannot let it correct it. The node believes it is full, however according to df there it 1.4 TB free. It does not accept any more uploads. But I cannot raise the available space due to this change I guess: as it is already at max space. So no matter what I enter in the docker run command, it remains at the max disk size. So what can I do, is there a way to override this?

I should be gone with the save-state-resume feature working.

---

<div class="post-metadata">

**Author:** ![jammerdan](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/jammerdan/32/1121_2.png) [@jammerdan](https://forum.storj.io/u/jammerdan)\
**Post date:** [June 15, 2024, 1:47pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/60 "2024-06-15T13:47:10Z")

</div>

> [@elek](#):
>
> The problem with filewalkers is retrieving the size and last modification time. With ext4 file system, directory entries contain only the file names. An additional read of the inode is required for getting the size and last modification time of the files.

What if we add that to the filename?  
Something like `piecename.sj1.year-month-day` e.g.: Taking a random piece from one of my nodes: `xwx3njdry2vhdthddq23773yfwcm4tv73qdpv65o7topw3pepq.sj1.2024-02-07`.  
Then we have everything when we have the files name. We could even add the size:

```auto
xwx3njdry2vhdthddq23773yfwcm4tv73qdpv65o7topw3pepq.sj1.2024-02-07.2319872

```

With some simple text manipulation operations you have all information about the file. I just tested, ext4 does accept such a filename.

---

<div class="post-metadata">

**Author:** ![Toyoo](https://storj.bcdn.literatehosting.com/user_avatar/forum.storj.io/toyoo/32/1126_2.png) [@Toyoo](https://forum.storj.io/u/Toyoo)\
**Post date:** [June 15, 2024, 8:54pm UTC](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874/61 "2024-06-15T20:54:39Z")

</div>

> [@jammerdan](#):
>
> Something like `piecename.sj1.year-month-day` e.g.: Taking a random piece from one of my nodes: `xwx3njdry2vhdthddq23773yfwcm4tv73qdpv65o7topw3pepq.sj1.2024-02-07`.

Then you make downloads much slower. You can’t “just” open a file, you have to scan the whole subdirectory each time to figure out what are the possible file names for a given piece ID.

[Previous page](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874.md?page=2)

[Next page](https://forum.storj.io/t/it-seems-that-the-new-feature-save-state-resume-gc-filewalker-isn-t-functioning-as-expected/25874.md?page=4)
