[HN Gopher] Deduplicating a 10.4 TiB game preservation archive (...
___________________________________________________________________
Deduplicating a 10.4 TiB game preservation archive (WIP)
Hi folks, I am working on a game preservation project, where the
data set holds 10.4 TiB. It contains 1044 earlier versions of a
single game in a multitude of different languages, architectures
and stages of development. As you can guess, that means extreme
redundancy. The goals are: - bring the size down - retain good
read speed (for further processing/reversing) - easy sharable
format - lower end machines can use it My choice fell on the BTRFS
filesystem, since it provides advanced features for deduplication,
which is not as resource hungry as ZFS. Once the data is processed,
it no longer requires a lot of system resources. In the first
round of deduplication, I used "jdupes -rQL" (yes, I know what -Q
does) to replace exact copies of files in different directories via
hardlinks to minimize data and metadata. This got it down to
roughly 874 GiB already, out of which 866 GiB are MPQ files. That's
99,08%... everything besides is a drop in the bucket. For those
uninitiated: this is an archive format. Representing it as a
pseudo-code struct it looks something like this { header, files[],
hash_table[], block_table[] } Compression exists, but it is applied
to each file individually. This means the same file is compressed
the same way in different MPQ archives, no matter the offset it
happens to be in. What is throwing a wrench into my plans of
further data deduplication are the following points: - the order of
files seems not to be deterministic when MPQ files were created (at
least I picked that up somewhere) - altered order of elements
(files added or removed at the start) causes shifts in file offsets
I thought for quite some time about this, and I think the smartest
way forward is, that I manually hack apart the file into multiple
extents at specific offsets. Thus the file would contain of an
extent for: - the header - each file individually - the hash table
- the block table It will increase the size for each file of
course, because of wasted space at the end of the last block in
each extent. But it allows for sharing whole extents between
different archives (and extracted files of it), as long as the file
within is content-wise the same, no matter the exact offset. The
second round of deduplication will then be whole extents via
duperemove, which should cut down the size dramatically once more.
This is where I am hanging right now: I don't know how to pull it
off on a technical level. I already was crawling through
documentation, googling, asking ChatGPT and fighting it's
hallucinations, but so far I wasn't very successful in finding
leads (probably need to perform some ioctl calls). From what I
imagine, there are probably two ways to do this: - rewrite the file
with a new name in the intended extent layout, delete the original
and rename the new one to take it's place - rewrite the extent
layout of an already existing file, without bending over backwards
like described above I need is a reliable way to, without chances
of the filesystem optimizing away my intended layout, while I write
it. The best case scenario for a solution would be a call, which
takes a file/inode and a list of offsets, and then reorganizes it
into that extents. If something like this does not exist, neither
through btrfs-progs, nor other third party applications, I would be
up for writing a generic utility like described above. It would
enable me to solve my problem, and others to write their own custom
dedicated deduplicaton software for their specific scenario. If
YOU - can guide me into the right direction - give me hints how to
solve this - tell me about the right btrfs communities where I can
talk about it - brainstorm ideas I would be eternally grateful :)
This is not a call for YOU to solve my problem, but for some
guidance, so I can do it on my own. I think that BTRFS is superb
for deduplicated archives, and it can really shine, if you can give
it a helping hand.
Author : DrFrugal
Score : 7 points
Date : 2024-12-17 20:00 UTC (3 hours ago)
| brudgers wrote:
| Deduplication is a poor archival strategy. Storage is cheap.
| 10TiB is a couple of thousand dollars with reasonable redundancy.
|
| Write Once is how to manage an archive.
|
| Displaying a curated subset is a good interface. Good luck.
| DrFrugal wrote:
| this was not helpful at all, and i think you also did not read
| the goals of this project
| brudgers wrote:
| I assumed the goal was archiving games for preservation.
|
| If the goal is algorithmic erasure of data rather than
| preservation of games, then "archive" might again create
| confusion like the type I probably have.
|
| If there is a strong business case for deduplication, I
| recommend hiring a consultant with expertise and experience
| in the problem.
|
| To be clear search is the only way to identify redundancy. If
| you have search, then redundancy is not a problem.
| moonshadow565 wrote:
| Lookup FICLONERANGE ioctl
| DrFrugal wrote:
| this is either a very big coincidence, or you are in the
| datamining discord as well. the original archive i base my
| project on uses RMAN to store everything :D --- thanks for the
| hint about the FICLONERANGE ioctl... it seems to be fine
| grained enough to allow me deduplicate on arbitrary offsets,
| not just whole blocks. will give it a go.
| sillystuff wrote:
| [delayed]
___________________________________________________________________
(page generated 2024-12-17 23:01 UTC)