[HN Gopher] Quickwit 0.8: Indexing and Search at Petabyte Scale
       ___________________________________________________________________
        
       Quickwit 0.8: Indexing and Search at Petabyte Scale
        
       Author : vvoyer
       Score  : 104 points
       Date   : 2024-03-19 14:56 UTC (4 days ago)
        
 (HTM) web link (quickwit.io)
 (TXT) w3m dump (quickwit.io)
        
       | dracyr wrote:
       | Never had the chance to use Quickwit at a $DAYJOB (yet?), but I
       | really appreciate the fact that it scales down quite well too.
       | Currently running it on my homelab, after a number of small
       | annoyances using Loki in a single-node cluster, and it's been
       | working very well with very reasonable resource usage.
       | 
       | I also decide to use Tantivy (the rust library powering/written
       | by Quickwit) for my own bookmarking search tool by embedding it
       | in Elixir, and the API and docs have been quite pleasant to work
       | with. Hats of to the team, looking forward to what's coming next!
        
         | francoismassot wrote:
         | Some companies are using it with AWS Lambda to scale to 0.
        
         | tecleandor wrote:
         | Ah Loki, I wanted to try it at my homelab bit it wasn't as
         | simple as it says. Now I wanted to try Zincsearch or
         | Openobserve. Have you tried that?
        
           | netingle wrote:
           | > it wasn't as simple as it says
           | 
           | mind elaborating? we built loki for some pretty massive scale
           | but I've always tried to make it work at super small scale
           | to. what went wrong?
        
           | pranay01 wrote:
           | You might want to have a look at SigNoz [1] as well. We have
           | also published some perf benchmark wrt Elastic & Loki [2] and
           | have some cool features like logs pipeline for manipulating
           | logs before ingestion
           | 
           | [1] https://github.com/signoz/signoz [2]
           | https://signoz.io/blog/logs-performance-benchmark/
        
           | bbkane wrote:
           | I use OpenObserve and I quite enjoy it
        
           | mdaniel wrote:
           | in case it matters to others,
           | https://github.com/openobserve/openobserve/tree/v0.7.0 is the
           | last Apache2 licensed copy before they went AGPL with 0.7.1
           | 
           | https://github.com/openobserve/openobserve/blob/v0.7.0/.env..
           | .. is some "onoz" for me, but just recently someone submitted
           | https://github.com/aenix-io/etcd-operator to the CNCF sandbox
           | so maybe things have gotten better around keeping that PoS
           | alive
        
       | up2isomorphism wrote:
       | 13.4GB/s with 200x6 vcpus, gives 11MB/s per core, it is good but
       | hard to say impressive.
        
         | francoismassot wrote:
         | Building the inverted index is quite CPU-intensive, and we are
         | also merging index files called "splits".
        
           | kikimora wrote:
           | I never being able to understand why log indexing has to
           | build inverted index. Decent columnar store with partitioning
           | by date should be enough to quickly filter gigabytes of logs.
        
             | nh2 wrote:
             | Because you want to find all occurrences of "error abc123"
             | over the last year, immediately?
        
             | fulmicoton wrote:
             | Quickwit co-founder here... I actually agree. For a few
             | GBs, done right, columnar works fine AND is cost efficient.
             | 
             | After all, it does not matter much if a log search query
             | answers in 300ms or 1s. However, there are use cases where
             | a few GB just does not cut it.
             | 
             | The tale saying that you can always prune your dataset
             | using timestamp and tags is simply not always valid.
        
         | fulmicoton wrote:
         | What is your frame of reference?
        
           | dist1ll wrote:
           | Per-core store bandwidth is at least 14GB/s on Zen3, 35GB/s
           | for non-temporal stores. Parsing JSON can be done at +2GB/s.
           | 
           | It's very healthy to take maximum bandwidth limits into
           | consideration when reasoning about performance. For instance,
           | for temporal stores, the bottlenecks you see are due to RAM
           | latency and memory parallelism, because of the write-
           | allocate. The load/store uarch can actually retire _way_ more
           | data from SIMD registers.
           | 
           | So there's already some headroom for CPU-bound tasks. For
           | instance 11MB/s is very slow for JIT baseline compiler. But
           | if your particular problem demands arbitrary random access
           | that exceed L3 regularly, maybe that speed is justified.
        
             | fulmicoton wrote:
             | What we do is CPU bound and we are not just parsing JSON
             | here.
             | 
             | The largest work we do is building an inverted index.
             | Oversimplified, it is equivalent to this:
             | inverted_index = defaultdict(list)       for (doc_id,
             | doc_json) in enumerate(doc_jsons):         c =
             | json.loads(payload)         for (field, field_text) in
             | c.items():           for (position, token) in enumerate():
             | inverted_index[token].push((doc, position))
             | 
             | serialize_in_compressed_way_that_allows_lookup(inverted_ind
             | ex)
             | 
             | You can implement it in a couple of hours in the language
             | of your choice to get a proper baseline.
             | 
             | I am sure we can still improve our indexing throughput...
             | but I have never seen any search engine indexing as fast as
             | tantivy.
             | 
             | If someone knows a project I should know of, I'd be
             | genuinely keen on learning from it.
        
               | dist1ll wrote:
               | I'm curious, what is your frame of reference with regards
               | to maximum speed of building inverted indices? Like, what
               | is the maximum throughput you'd expect for this type of
               | task, and what is your reasoning for it?
        
       | halvorbo wrote:
       | Amazing to see how far Tantiviy has come. Remember using and
       | making some smaller contributions to this 3 years ago - slop to
       | phrase queries for example. Curious how the design has changed to
       | enable large scale production usage.
        
         | francoismassot wrote:
         | Thanks! Quickwit is the distributed engine built on top of
         | tantivy, we basically separated compute and storage for search,
         | I wrote this blog post to introduce the architecture:
         | https://quickwit.io/blog/quickwit-101
         | 
         | PS: it's tantivy!!!
        
         | fulmicoton wrote:
         | Very valuable contribution!
        
       | godber wrote:
       | We did some experimentation with quickwit about a year ago,
       | writing about 1m docs/second of data into it for several months.
       | It worked well and was pretty straight forward to learn and
       | operate. If we didn't also manage our own S3/Ceph it might be a
       | big win, once feature complete. It's definitely worth a look.
        
       | arisudesu wrote:
       | musl support would be highly appreciated.
        
       ___________________________________________________________________
       (page generated 2024-03-23 23:02 UTC)