[HN Gopher] Quickwit 0.8: Indexing and Search at Petabyte Scale
___________________________________________________________________
Quickwit 0.8: Indexing and Search at Petabyte Scale
Author : vvoyer
Score : 104 points
Date : 2024-03-19 14:56 UTC (4 days ago)
(HTM) web link (quickwit.io)
(TXT) w3m dump (quickwit.io)
| dracyr wrote:
| Never had the chance to use Quickwit at a $DAYJOB (yet?), but I
| really appreciate the fact that it scales down quite well too.
| Currently running it on my homelab, after a number of small
| annoyances using Loki in a single-node cluster, and it's been
| working very well with very reasonable resource usage.
|
| I also decide to use Tantivy (the rust library powering/written
| by Quickwit) for my own bookmarking search tool by embedding it
| in Elixir, and the API and docs have been quite pleasant to work
| with. Hats of to the team, looking forward to what's coming next!
| francoismassot wrote:
| Some companies are using it with AWS Lambda to scale to 0.
| tecleandor wrote:
| Ah Loki, I wanted to try it at my homelab bit it wasn't as
| simple as it says. Now I wanted to try Zincsearch or
| Openobserve. Have you tried that?
| netingle wrote:
| > it wasn't as simple as it says
|
| mind elaborating? we built loki for some pretty massive scale
| but I've always tried to make it work at super small scale
| to. what went wrong?
| pranay01 wrote:
| You might want to have a look at SigNoz [1] as well. We have
| also published some perf benchmark wrt Elastic & Loki [2] and
| have some cool features like logs pipeline for manipulating
| logs before ingestion
|
| [1] https://github.com/signoz/signoz [2]
| https://signoz.io/blog/logs-performance-benchmark/
| bbkane wrote:
| I use OpenObserve and I quite enjoy it
| mdaniel wrote:
| in case it matters to others,
| https://github.com/openobserve/openobserve/tree/v0.7.0 is the
| last Apache2 licensed copy before they went AGPL with 0.7.1
|
| https://github.com/openobserve/openobserve/blob/v0.7.0/.env..
| .. is some "onoz" for me, but just recently someone submitted
| https://github.com/aenix-io/etcd-operator to the CNCF sandbox
| so maybe things have gotten better around keeping that PoS
| alive
| up2isomorphism wrote:
| 13.4GB/s with 200x6 vcpus, gives 11MB/s per core, it is good but
| hard to say impressive.
| francoismassot wrote:
| Building the inverted index is quite CPU-intensive, and we are
| also merging index files called "splits".
| kikimora wrote:
| I never being able to understand why log indexing has to
| build inverted index. Decent columnar store with partitioning
| by date should be enough to quickly filter gigabytes of logs.
| nh2 wrote:
| Because you want to find all occurrences of "error abc123"
| over the last year, immediately?
| fulmicoton wrote:
| Quickwit co-founder here... I actually agree. For a few
| GBs, done right, columnar works fine AND is cost efficient.
|
| After all, it does not matter much if a log search query
| answers in 300ms or 1s. However, there are use cases where
| a few GB just does not cut it.
|
| The tale saying that you can always prune your dataset
| using timestamp and tags is simply not always valid.
| fulmicoton wrote:
| What is your frame of reference?
| dist1ll wrote:
| Per-core store bandwidth is at least 14GB/s on Zen3, 35GB/s
| for non-temporal stores. Parsing JSON can be done at +2GB/s.
|
| It's very healthy to take maximum bandwidth limits into
| consideration when reasoning about performance. For instance,
| for temporal stores, the bottlenecks you see are due to RAM
| latency and memory parallelism, because of the write-
| allocate. The load/store uarch can actually retire _way_ more
| data from SIMD registers.
|
| So there's already some headroom for CPU-bound tasks. For
| instance 11MB/s is very slow for JIT baseline compiler. But
| if your particular problem demands arbitrary random access
| that exceed L3 regularly, maybe that speed is justified.
| fulmicoton wrote:
| What we do is CPU bound and we are not just parsing JSON
| here.
|
| The largest work we do is building an inverted index.
| Oversimplified, it is equivalent to this:
| inverted_index = defaultdict(list) for (doc_id,
| doc_json) in enumerate(doc_jsons): c =
| json.loads(payload) for (field, field_text) in
| c.items(): for (position, token) in enumerate():
| inverted_index[token].push((doc, position))
|
| serialize_in_compressed_way_that_allows_lookup(inverted_ind
| ex)
|
| You can implement it in a couple of hours in the language
| of your choice to get a proper baseline.
|
| I am sure we can still improve our indexing throughput...
| but I have never seen any search engine indexing as fast as
| tantivy.
|
| If someone knows a project I should know of, I'd be
| genuinely keen on learning from it.
| dist1ll wrote:
| I'm curious, what is your frame of reference with regards
| to maximum speed of building inverted indices? Like, what
| is the maximum throughput you'd expect for this type of
| task, and what is your reasoning for it?
| halvorbo wrote:
| Amazing to see how far Tantiviy has come. Remember using and
| making some smaller contributions to this 3 years ago - slop to
| phrase queries for example. Curious how the design has changed to
| enable large scale production usage.
| francoismassot wrote:
| Thanks! Quickwit is the distributed engine built on top of
| tantivy, we basically separated compute and storage for search,
| I wrote this blog post to introduce the architecture:
| https://quickwit.io/blog/quickwit-101
|
| PS: it's tantivy!!!
| fulmicoton wrote:
| Very valuable contribution!
| godber wrote:
| We did some experimentation with quickwit about a year ago,
| writing about 1m docs/second of data into it for several months.
| It worked well and was pretty straight forward to learn and
| operate. If we didn't also manage our own S3/Ceph it might be a
| big win, once feature complete. It's definitely worth a look.
| arisudesu wrote:
| musl support would be highly appreciated.
___________________________________________________________________
(page generated 2024-03-23 23:02 UTC)