[HN Gopher] Deduplicating a 10.4 TiB game preservation archive (...
       ___________________________________________________________________
        
       Deduplicating a 10.4 TiB game preservation archive (WIP)
        
       Hi folks,  I am working on a game preservation project, where the
       data set holds 10.4 TiB. It contains 1044 earlier versions of a
       single game in a multitude of different languages, architectures
       and stages of development. As you can guess, that means extreme
       redundancy.  The goals are: - bring the size down - retain good
       read speed (for further processing/reversing) - easy sharable
       format - lower end machines can use it  My choice fell on the BTRFS
       filesystem, since it provides advanced features for deduplication,
       which is not as resource hungry as ZFS. Once the data is processed,
       it no longer requires a lot of system resources.  In the first
       round of deduplication, I used "jdupes -rQL" (yes, I know what -Q
       does) to replace exact copies of files in different directories via
       hardlinks to minimize data and metadata. This got it down to
       roughly 874 GiB already, out of which 866 GiB are MPQ files. That's
       99,08%... everything besides is a drop in the bucket.  For those
       uninitiated: this is an archive format. Representing it as a
       pseudo-code struct it looks something like this { header, files[],
       hash_table[], block_table[] } Compression exists, but it is applied
       to each file individually. This means the same file is compressed
       the same way in different MPQ archives, no matter the offset it
       happens to be in.  What is throwing a wrench into my plans of
       further data deduplication are the following points: - the order of
       files seems not to be deterministic when MPQ files were created (at
       least I picked that up somewhere) - altered order of elements
       (files added or removed at the start) causes shifts in file offsets
       I thought for quite some time about this, and I think the smartest
       way forward is, that I manually hack apart the file into multiple
       extents at specific offsets. Thus the file would contain of an
       extent for: - the header - each file individually - the hash table
       - the block table It will increase the size for each file of
       course, because of wasted space at the end of the last block in
       each extent. But it allows for sharing whole extents between
       different archives (and extracted files of it), as long as the file
       within is content-wise the same, no matter the exact offset. The
       second round of deduplication will then be whole extents via
       duperemove, which should cut down the size dramatically once more.
       This is where I am hanging right now: I don't know how to pull it
       off on a technical level. I already was crawling through
       documentation, googling, asking ChatGPT and fighting it's
       hallucinations, but so far I wasn't very successful in finding
       leads (probably need to perform some ioctl calls).  From what I
       imagine, there are probably two ways to do this: - rewrite the file
       with a new name in the intended extent layout, delete the original
       and rename the new one to take it's place - rewrite the extent
       layout of an already existing file, without bending over backwards
       like described above  I need is a reliable way to, without chances
       of the filesystem optimizing away my intended layout, while I write
       it. The best case scenario for a solution would be a call, which
       takes a file/inode and a list of offsets, and then reorganizes it
       into that extents. If something like this does not exist, neither
       through btrfs-progs, nor other third party applications, I would be
       up for writing a generic utility like described above. It would
       enable me to solve my problem, and others to write their own custom
       dedicated deduplicaton software for their specific scenario.  If
       YOU - can guide me into the right direction - give me hints how to
       solve this - tell me about the right btrfs communities where I can
       talk about it - brainstorm ideas I would be eternally grateful :)
       This is not a call for YOU to solve my problem, but for some
       guidance, so I can do it on my own.  I think that BTRFS is superb
       for deduplicated archives, and it can really shine, if you can give
       it a helping hand.
        
       Author : DrFrugal
       Score  : 7 points
       Date   : 2024-12-17 20:00 UTC (3 hours ago)
        
       | brudgers wrote:
       | Deduplication is a poor archival strategy. Storage is cheap.
       | 10TiB is a couple of thousand dollars with reasonable redundancy.
       | 
       | Write Once is how to manage an archive.
       | 
       | Displaying a curated subset is a good interface. Good luck.
        
         | DrFrugal wrote:
         | this was not helpful at all, and i think you also did not read
         | the goals of this project
        
           | brudgers wrote:
           | I assumed the goal was archiving games for preservation.
           | 
           | If the goal is algorithmic erasure of data rather than
           | preservation of games, then "archive" might again create
           | confusion like the type I probably have.
           | 
           | If there is a strong business case for deduplication, I
           | recommend hiring a consultant with expertise and experience
           | in the problem.
           | 
           | To be clear search is the only way to identify redundancy. If
           | you have search, then redundancy is not a problem.
        
       | moonshadow565 wrote:
       | Lookup FICLONERANGE ioctl
        
         | DrFrugal wrote:
         | this is either a very big coincidence, or you are in the
         | datamining discord as well. the original archive i base my
         | project on uses RMAN to store everything :D --- thanks for the
         | hint about the FICLONERANGE ioctl... it seems to be fine
         | grained enough to allow me deduplicate on arbitrary offsets,
         | not just whole blocks. will give it a go.
        
       | sillystuff wrote:
       | [delayed]
        
       ___________________________________________________________________
       (page generated 2024-12-17 23:01 UTC)