[HN Gopher] Show HN: ToplingDB - A Persistent Key-Value Store fo...
       ___________________________________________________________________
        
       Show HN: ToplingDB - A Persistent Key-Value Store for External
       Storage
        
       As the creator of TerarkDB (acquired by ByteDance in 2019), I have
       developed ToplingDB in recent years.  ToplingDB is forked from
       RocksDB, where we have replaced almost all components with more
       efficient alternatives(db_bench shows ToplingDB is about ~8x faster
       than RocksDB):  * MemTable: SkipList is replaced by CSPP(Crash Safe
       Parallel Patricia trie), which is 8x faster.  * SST:
       BlockBasedTable is replaced by ToplingZipTable, implemented by
       searchable compression algo, it is very small and fast, typically
       less than 1ms per lookup:                 * Keys/Indexes are
       compressed   using NestLoudsTrie(a multi-layer nesting LOUDS
       succinct trie).            * Values in a SST are compressed
       together with better zip ratio than zstd, and can unzip by a single
       value at 1GB/sec.            * BlockCache is no longer needed,
       double caching(BlockCache & PageCache) is avoided       Other
       hotspots are also improved:  * Flush MemTable to L0 is omited,
       greatly reducing write amp and is very friendly for large(GB)
       MemTable                 * MemTable   serves as the index of Key to
       "value position in WAL log"            * Since WAL file content
       almost always in page cache, thus value content can be efficiently
       accessed by mmap            * When Flush happens, MemTable is
       dumpped as an SST and WAL is treated as a blob file              *
       CSPP MemTable use integer index instead of physical pointers, thus
       in-memory format is exactly same with in-file format       * Prefix
       cache for searching candidate SSTs and prefix cache for scanning by
       iterators                 * Caching fixed len key prefix into an
       array, binary search it as an uint array       * Distributed
       compaction(superior replacement to rocksdb remote compaction)
       * Gracefully support MergeOperator, CompactionFilter,
       PropertiesCollector...            * Out of the box, development
       efforts are significantly reduced            * Very easy to share
       compaction service on spot instances for many DB nodes       Useful
       Bonus Feature:  * Config by json/yaml: can config almost all
       features  * Optional embeded WebView: show db structures in web
       browser, refreshing pages like animation  * Online update db
       configs by http  MySQL integration, ToplingDB has integrated into
       MySQL by MyTopling, which is forked from MyRocks with great
       improvements, like improvements of ToplingDB on RocksDB:  *
       WBWI(WriteBatchWithIndex): like MemTable, SkipList is replace with
       CSPP, 20x faster(speedup is more than MemTable).  * LockManager &
       LockTracker: 10x faster  * Encoding & Decoding: 5x faster  * Others
       ....  MyRocks has many disadvantages compared to InnoDB, while
       MyTopling outperforms InnoDB at almost all aspect - excluding
       feature differences.  We have create ~100 PRs for RocksDB, in which
       ~40 were accepted. Our PRs are mostly "small" changes, since big
       changes are not likely accepted.  ToplingDB has been deployed in
       numerous production environments.  Welcome every one using
       ToplingDB & MyTopling, and discuss in
       https://github.com/topling/toplingdb/discussions
        
       Author : rockeetterark
       Score  : 57 points
       Date   : 2025-07-01 10:07 UTC (12 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | ChocolateGod wrote:
       | I'm confused what makes this cloud native?
        
         | dboreham wrote:
         | It has an embedded http server?
        
         | faizshah wrote:
         | From what I gather it has an embedded http control plane,
         | yaml/json config for plugins, prometheus integration, and
         | distributed compaction workers on separate, potentially
         | serverless, hosts.
        
       | andybak wrote:
       | This is failing my "Can I figure out what the hell it is in 60
       | seconds?" test.
       | 
       | Sometimes that means I'm just not the target market. I do do web
       | dev (among other things) so that doesn't seem to be the case at
       | first glance?
        
         | faizshah wrote:
         | It's RocksDB but faster because data can be searched while
         | still compressed allowing you to load more records in less
         | cache/ram leading to up to 10x performance of RocksDB. It adds
         | an embedded http control plane as well as supporting other
         | extensions like MyRocks (MySQL) and Todis (redis
         | compatibility).
         | 
         | Or at least thats what I got from it correct me if I am wrong
         | rockeet.
        
       | alexpadula wrote:
       | Very extensive, great work on TerarkDB and Topling!
        
       | dangoodmanUT wrote:
       | Without better (english) docs it will be hard to get adoption,
       | unfortunately. 8x perf gain over rocksdb is... a lot... unless
       | you're poking at particularly bad metrics.
        
       | absoluteunit1 wrote:
       | For the laymen folks reading this - what are the ideal use cases
       | for this?
        
         | nbf_1995 wrote:
         | Like RocksDB from which this appears to be forked, the primary
         | usage is as a storage engine for other applications/databases.
         | Compared to rocksdb, it seems like ToplingDB has added more
         | facilities to better support distributed use-cases.
         | 
         | Some databases that utilize RocksDB for their storage engine:
         | https://kvrocks.apache.org/ - Redis/ValKey compatible
         | distributed database with disk persistence via RockDB.
         | https://github.com/pingcap/tidb - MySQL compatible distributed
         | database. Mentioned elsewhere in this thread.
         | https://github.com/tikv/tikv - Distributed, transactional, key
         | value store. Originally by the same company as TiDB.
         | 
         | In theory you could use it as an in-process KV store similar to
         | how SQLite provides an in process sql database, but the api is
         | far from ergonomic for that use case.
        
           | absoluteunit1 wrote:
           | Ah I see! Thanks for explanation :)
        
       | alex7o wrote:
       | What does it have to do with external storage in this context,
       | does it mean S3. Initially I thought it is a db for thumb drives?
        
       | ozgrakkurt wrote:
       | Would be really interesting to have faster compilation and more
       | simplicity (auto tuning parameters etc.) compared to rocksdb. In
       | my experience rocksdb performance is very good and it is reliable
       | but it is a pain to integrate into the build process and has too
       | many configurations
        
       | esafak wrote:
       | A distributed KV-store plus a relational layer makes it a
       | competitor to NewSQL databases like TiDB, which is also based on
       | Facebook's RocksDB.
       | 
       | It doesn't look like it's very actively developed:
       | https://github.com/topling/toplingdb/pulse/monthly
       | 
       | To the OP who's developing it: I suggest polishing your README.
       | Provide a simple installation tutorial, maybe a trial offering
       | like tidbcloud.com, and comparative benchmark results, since you
       | advertise your performance.
        
         | jauntywundrkind wrote:
         | It's quite active. They just aren't using GitHub pull requests
         | in their workflow, which is what GitHub Pulse measures.
         | https://github.com/topling/toplingdb/commits/memtable_as_log...
        
       ___________________________________________________________________
       (page generated 2025-07-01 23:02 UTC)