[HN Gopher] Zpdf: PDF text extraction in Zig - 5x faster than MuPDF
___________________________________________________________________
Zpdf: PDF text extraction in Zig - 5x faster than MuPDF
Author : lulzx
Score : 80 points
Date : 2025-12-30 19:57 UTC (3 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| lulzx wrote:
| I built a PDF text extraction library in Zig that's significantly
| faster than MuPDF for text extraction workloads.
|
| ~41K pages/sec peak throughput.
|
| Key choices: memory-mapped I/O, SIMD string search, parallel page
| extraction, streaming output. Handles CID fonts, incremental
| updates, all common compression filters.
|
| ~5,000 lines, no dependencies, compiles in <2s.
|
| Why it's fast: - Memory-mapped file I/O (no read
| syscalls) - Zero-copy parsing where possible - SIMD-
| accelerated string search for finding PDF structures -
| Parallel extraction across pages using Zig's thread pool -
| Streaming output (no intermediate allocations for extracted text)
|
| What it handles: - XRef tables and streams (PDF
| 1.5+) - Incremental PDF updates (/Prev chain) -
| FlateDecode, ASCII85, LZW, RunLength decompression - Font
| encodings: WinAnsi, MacRoman, ToUnicode CMap - CID fonts
| (Type0, Identity-H/V, UTF-16BE with surrogate pairs)
| tveita wrote:
| What kind of performance are you seeing with/without SIMD
| enabled?
|
| From https://github.com/Lulzx/zpdf/blob/main/src/main.zig it
| looks like the help text cites an unimplemented "-j" option to
| enable multiple threads.
|
| There is a "--parallel" option, but that is only implemented
| for the "bench" command.
| lulzx wrote:
| I have now made parallel by default and added an option to
| enable multiple threads.
|
| I haven't tested without SIMD.
| cheshire_cat wrote:
| You've released quite a few projects lately, very impressive.
|
| Are you using LLMs for parts of the coding?
|
| What's your work flow when approaching a new project like this?
| littlestymaar wrote:
| > Are you using LLMs for parts of the coding?
|
| I can't talk about the code, but the readme and commit
| messages are most likely LLM-generated.
|
| And when you take into account that the first commit happened
| just three hours ago, it feels like the entire project has
| been vibe coded.
| Neywiny wrote:
| Hard disagree. Initial commit was 6k LOC. Author could've
| spent years before committing. Ill advised but not
| impossible.
| littlestymaar wrote:
| Why would you make Claude write your commit message for a
| commit you've spent years working on though?
| lulzx wrote:
| Claude Code.
| jeffbee wrote:
| What's fast about mmap?
| jonstewart wrote:
| What's the fidelity like compared to tika?
| lulzx wrote:
| The accuracy difference is marginal (1-2%) but the speed
| difference is massive.
| agentifysh wrote:
| excellent stuff what makes zig so fast
| observationist wrote:
| Not being slow - they compile straight to bytecode, they aren't
| interpreted, and have aggressive, opinionated optimizations
| baked in by default, so it's even faster than compiled c (under
| default conditions.)
|
| Contrasted with python, which is interpreted, has a clunky
| runtime, minimal optimizations, and all sorts of choices that
| result in slow, redundant, and also slow, performance.
|
| The price for performance is safety checks, redundancy, how
| badly wrong things can go, and so on.
|
| A good compromise is luajit - you get some of the same
| aggressive optimizations, but in an interpreted language, with
| better-than-c performance but interpreted language convenience,
| access to low level things that can explode just as
| spectacularly as with zig or c, but also a beautiful language.
| agentifysh wrote:
| will add this to the list, now learning new languages is less
| of a barrier with LLMs
| Zambyte wrote:
| Zig is safer than C under default conditions, not faster. By
| default does a lot of illegal behavior safety checking, such
| as array and slice bounds checking, numeric overflow
| checking, and invalid union access checking. These features
| are disabled by certain (non default) build modes, or
| explicitly disabled at a per scope level.
|
| It may be easier to write code that runs faster in Zig than
| in C under similar build optimization levels, because writing
| high performance C code looks a lot like writing idiomatic
| Zig code. The Zig standard library offers a lot of structures
| like hash maps, SIMD primitives, and allocators with
| different performance characteristics to better fit a given
| use-case. C application code often skips on these things
| simply because it is a lot more friction to do in C than in
| Zig.
| AndyKelley wrote:
| It makes your development workflow smooth enough that you have
| the time and energy to do stuff like all the bullet points
| listed in https://news.ycombinator.com/item?id=46437289
| mpeg wrote:
| very nice, it'd be good to see a feature comparison as when I use
| mupdf it's not really just about speed, but about the level of
| support of all kinds of obscure pdf features, and good level of
| accuracy of the built-in algorithms for things like handling two-
| column pages, identifying paragraphs, etc.
|
| the licensing is a huge blocker for using mupdf in non-OSS tools,
| so it's very nice to see this is MIT
|
| python bindings would be good too
| lulzx wrote:
| added a comparison, will improve further.
| https://github.com/Lulzx/zpdf?tab=readme-ov-file#comparison-...
|
| also, added python bindings.
| mpeg wrote:
| thanks, claude, I guess haha
|
| as others have commented, I think while this is a nice
| portfolio piece, I would worry about its longevity as a vibe
| coded project
| odie5533 wrote:
| Now we just need Python bindings so I can use it in my trash
| language of choice.
| lulzx wrote:
| added python bindings!
| hiq wrote:
| Were you working on it already, or did it take you less than
| 17 minutes to commit https://github.com/Lulzx/zpdf/commit/9f5
| a7b70eb4b53672c0e4d8... ?
| littlestymaar wrote:
| - First commit 3hours ago.
|
| - commit message: LLM-generated.
|
| - README: LLM-generated.
|
| I'm not convinced that projects vibe coded over the evening
| deserve the HN front page...
|
| Edit: and of course the author's blog is also full of AI slop...
|
| 2026 hasn't even started I already hate it.
| kingkongjaffa wrote:
| Wait, but why?
|
| If it's really better than what we had before, what does it
| matter how it was made? It's literally hacked together with the
| tools of the day (LLMs) isn't that the very hacker ethos?
| Patching stuff together that works in a new and useful way.
|
| 5x speed improvements on pdf text extraction might be great for
| some applications I'm not aware of, I wouldn't just dismiss it
| out of hand because the author used $robot to write the code.
|
| Presumably the thought to make the thing in the first place and
| decide what features to add and not add was more important than
| how the code is generated?
| forgotpwd16 wrote:
| 74910,74912c187768,187779 < [Example 1: If you want to use
| the code conversion facetcodecvt_utf8to output tocouta UTF-8
| multibyte sequence < corresponding to a wide string, but
| you don't want to alter the locale forcout, you can write
| something like:\237 D.27.21954
| \251ISO/IECN4950wstring_convert<std::codecvt_utf8<wchar_t>>
| myconv; < std::string mbstring =
| myconv.to_bytes\050L"Hello\134n"\051; --- > >
| [Example 1: If you want to use the code conversion facet
| codecvt_utf8 to output to cout a UTF-8 multibyte sequence >
| corresponding to a wide string, but you don't want to alter the
| locale for cout, you can write something like: > > SS
| D.27.2 > 1954 > > (c) ISO/IEC > N4950
| > > wstring_convert<std::codecvt_utf8<wchar_t>> myconv;
| > std::string mbstring = myconv.to_bytes(L"Hello\n");
|
| Is indeed faster but output is messier.
___________________________________________________________________
(page generated 2025-12-30 23:00 UTC)