[HN Gopher] I built Foyer: a Rust hybrid cache that slashes S3 l...
___________________________________________________________________
I built Foyer: a Rust hybrid cache that slashes S3 latency
Author : Sheldon_fun
Score : 153 points
Date : 2025-09-23 16:25 UTC (4 days ago)
(HTM) web link (medium.com)
(TXT) w3m dump (medium.com)
| Eikon wrote:
| I use Foyer in-memory (not hybrid) in ZeroFS [0] and had a great
| experience with it.
|
| The only quirk I've experienced is that in-memory and hybrid
| modes don't share the same invalidation behavior. In hybrid mode,
| there's no way to await a value being actually discarded after
| deletion, while in-memory mode shows immediate deletion.
|
| [0] https://github.com/Barre/ZeroFS
| oulipo2 wrote:
| Interesting! What would be a typical use-case of ZeroFS? could
| I use this to store my Immich and Jellyfin data on S3 so I
| don't need disk?
| Eikon wrote:
| That should work!
| ofek wrote:
| The Jellyfin metadata would certainly be a fit but what
| about streaming video content i.e. sequential reads of
| large files with random access?
| Eikon wrote:
| If you have the network that matches, it should be
| perfectly fine.
| dexterdog wrote:
| If you don't mind paying about a dollar to stream one of your
| own movies as well as a couple of bucks per year to store it.
| Eikon wrote:
| You don't have to use AWS S3, any compatible implementation
| will work.
| mmastrac wrote:
| This is interesting. I would be curious to try a setup where I
| keep a local hybrid cache and transition blocks to deep storage
| for long-term archival via S3 rules.
|
| Some napkin math suggests this could be a few dollars a month
| to keep a few TB of precious data nearline.
|
| Restore costs are pricy but hopefully this is something that's
| only hit in case of true disaster. Are there any techniques for
| reducing egress on restore?
| Eikon wrote:
| The easy way to achieve that, is using ZeroFS NBD server with
| ZFS L2ARC (L2ARC with local storage and "main" pool on
| ZeroFS).
| nitishr wrote:
| I haven't used it yet but I have been looking for something like
| this for a long time. Kudos!
| jmpman wrote:
| How does this compare to S3 Mountpoint with caching?
| huntaub wrote:
| S3 Mountpoint is exposing a POSIX-like file system abstraction
| for you to use with your file-based applications. Foyer appears
| to be a library that helps your application coordinate access
| to S3 (with a cache), for applications that don't need files
| and you can change the code for.
| import wrote:
| Very curious about comparison between the rclone etc.. s3 caching
| vs this one.
| Mizza wrote:
| Has Medium stopped working on Firefox for anybody else? Once the
| page is finished loading, it stops responding to scroll events.
| tomrod wrote:
| I avoid medium where possible.
|
| If I could pipe text content to my terminal with confidence, I
| would.
| lukax wrote:
| Maybe you have an ad-blocker that just hides the popup but does
| not restore scrolling (scrolling is usually prevented when
| popups are visible)
| osigurdson wrote:
| Firefox has a lot of weird little pop up ads these days. It
| seems like this is a very recent phenominon. Is this actually
| Firefox doing this or some kind of plug-in accidentally
| installed?
| duttish wrote:
| Hm, I haven't seen that. Perhaps it's worth reviewing your
| plugins
| osigurdson wrote:
| Thanks! I think it might have been notifications from
| futurism.com. I don't remember visiting that site or
| allowing notifications (on purpose anyway).
| micw wrote:
| Same here. Meanwhile I close a link/page as soon as I realize
| it's on medium.
| speed_spread wrote:
| Have you tried using reader mode?
| stevekemp wrote:
| No idea, if I see a medium link I just ignore it. Substack is
| heading the same way for me too, it seems to be self-promotion,
| shallow-takes, and spam more than anything real.
| secondcoming wrote:
| Seems ok for me on Firefox 143.0.1
| overhead4075 wrote:
| The page loads a "subscribe to author" modal pretty quickly
| after the page loads. You may have partially blocked it, so you
| won't see the modal but it still prevents scroll.
| boldlybold wrote:
| Same. Hit escape shortly after the page loads to stop loading
| whatever modal is likely blocking scroll. I don't see the modal
| so it's likely blocked by ublock, but still stops scroll.
| mindreframer wrote:
| Here is a non-medium article with the same content:
| https://risingwave.com/blog/the-case-for-hybrid-cache-for-ob...
| mystifyingpoi wrote:
| Sounds exactly like AWS Storage Gateway, how does it compare?
| huntaub wrote:
| Storage Gateway is an appliance that you connect multiple
| instances to, this appears to be a library that you use in your
| program to coordinate caching for that process.
| alongub wrote:
| Foyer is great!
| winter_blue wrote:
| Does S3 really have that high of a latency? So high that ---- if
| you run a static file several in an EC2, would that be faster
| than S3?
| reese_john wrote:
| S3 has a low-latency offering[0] which promises single digit
| millisecond latency, I'm surprised not to see it mentioned.
|
| [0]: https://aws.amazon.com/s3/storage-classes/express-one-
| zone/
| huntaub wrote:
| These are, effectively, different use cases. You want to use
| (and pay for) Express One Zone in situations in which you
| need the same object reused from _multiple instances_
| repeatedly, while it looks like this on-disk or in-memory
| cache is for when you may want the same file repeatedly used
| from the _same instance_.
| manquer wrote:
| Is it the same instance ? Rising wave (and similar tools
| )are designed to run in production on a lot of distributed
| compute nodes for processing data , serving/streaming
| queries and running control panes .
|
| Even for any single query it will likely run on multiple
| nodes with distributed workers gathering and processing
| data from storage layer, that is whole idea behind
| MapReduce after all.
| artursapek wrote:
| Also, aren't most people putting Cloudfront in front of S3
| anyway?
| hobofan wrote:
| For CDN use-cases yes, but not for DB storage-compute
| separation use-cases as described here.
| huntaub wrote:
| Yes, definitely. S3 has a time to first byte of 50-150ms
| (depending on how lucky you are). If you're serving from memory
| that goes to ~0, and if you're serving from disk, that goes to
| 0.2-1ms.
|
| It will depend on your needs though, since some use cases won't
| want to trade off the scalability of S3's ability to serve
| arbitrary amounts of throughput.
| manquer wrote:
| In that case you run the proxy service load balanced to get
| desired throughput or run a sidecar/process in each compute
| instance where data is needed .
|
| You are limited anyway by the network capacity of the
| instance you are fetching the data from .
| k_bx wrote:
| Interesting to compare against ZeroFs
| shikhar wrote:
| Foyer is a great open source contribution from RisingWave
|
| We built a S3 read-through cache service for s2.dev so that
| multiple clients could share a Foyer hybrid cache with key
| affinity, https://github.com/s2-streamstore/cachey
| erikcw wrote:
| This looks really useful! Am I correct that there isn't an S3
| compatible API, just the "fetch" API?
|
| Being able to set an S3 client's endpoint to proxy traffic
| straight through this would be quite useful.
| riquito wrote:
| > Cost is reduced because far fewer requests hit S3
|
| I wonder. Given how cheap are S3 GET requests you need a massive
| number of requests to make provisioning and maintaining the cache
| server cheaper than the alternative.
| jitl wrote:
| I wonder to what degree it's actually necessary to explicitly
| manage memory spilling to disk like this. Want unified interface
| over non durable memory + disk? There already is one: memory with
| swap.
|
| Materialize.com switched from explicit disk cache management to
| "just use swap" and saw substantial performance improvement.
| https://materialize.com/blog/scaling-beyond-memory/
|
| To get good performance from this strategy your memory layout
| already needs to be optimized to pages boundaries etc to have
| good sympathy for underlying swap system, and you can explicitly
| force pages to stay in memory with a few syscalls.
| shikhar wrote:
| TIL https://kubernetes.io/blog/2025/03/25/swap-linux-
| improvement...
| hamandcheese wrote:
| > There already is one: memory with swap.
|
| Actually, there's already two. The other one: just read from
| disk (and let OS manage caches).
| cowsandmilk wrote:
| I prefer this abstraction as it is more widely supported
| (I've had to deploy to hosts that intentionally kill when you
| swap) and results in development assuming that the access may
| be at disk speed. When you rely on swap, I often see
| developers assuming everything is accessible at memory speed
| and then surprised when swap causes sudden degradation.
| freeloyer wrote:
| How does it compare to cachelib
| shikhar wrote:
| Essentially CacheLib in Rust
|
| > foyer draws inspiration from Facebook/CacheLib, a highly-
| regarded hybrid cache library written in C++, and ben-
| manes/caffeine, a popular Java caching library, among other
| projects.
|
| https://github.com/foyer-rs/foyer
| philip1209 wrote:
| Distributed Chroma, the open-source project backing Chroma Cloud,
| uses Foyer extensively:
|
| https://github.com/chroma-core/chroma/blob/2cb5c00d2e97ef449...
| ComputerGuru wrote:
| I think the article could use more on the cache invalidation and
| write-through (?) behavior. Are updates to the same file batched
| or written back to S3 immediately? Do you do anything with write
| conflicts, which one wins?
___________________________________________________________________
(page generated 2025-09-27 23:00 UTC)