[HN Gopher] Zpdf: PDF text extraction in Zig - 5x faster than MuPDF
       ___________________________________________________________________
        
       Zpdf: PDF text extraction in Zig - 5x faster than MuPDF
        
       Author : lulzx
       Score  : 80 points
       Date   : 2025-12-30 19:57 UTC (3 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | lulzx wrote:
       | I built a PDF text extraction library in Zig that's significantly
       | faster than MuPDF for text extraction workloads.
       | 
       | ~41K pages/sec peak throughput.
       | 
       | Key choices: memory-mapped I/O, SIMD string search, parallel page
       | extraction, streaming output. Handles CID fonts, incremental
       | updates, all common compression filters.
       | 
       | ~5,000 lines, no dependencies, compiles in <2s.
       | 
       | Why it's fast:                 - Memory-mapped file I/O (no read
       | syscalls)       - Zero-copy parsing where possible       - SIMD-
       | accelerated string search for finding PDF structures       -
       | Parallel extraction across pages using Zig's thread pool       -
       | Streaming output (no intermediate allocations for extracted text)
       | 
       | What it handles:                 - XRef tables and streams (PDF
       | 1.5+)       - Incremental PDF updates (/Prev chain)       -
       | FlateDecode, ASCII85, LZW, RunLength decompression       - Font
       | encodings: WinAnsi, MacRoman, ToUnicode CMap       - CID fonts
       | (Type0, Identity-H/V, UTF-16BE with surrogate pairs)
        
         | tveita wrote:
         | What kind of performance are you seeing with/without SIMD
         | enabled?
         | 
         | From https://github.com/Lulzx/zpdf/blob/main/src/main.zig it
         | looks like the help text cites an unimplemented "-j" option to
         | enable multiple threads.
         | 
         | There is a "--parallel" option, but that is only implemented
         | for the "bench" command.
        
           | lulzx wrote:
           | I have now made parallel by default and added an option to
           | enable multiple threads.
           | 
           | I haven't tested without SIMD.
        
         | cheshire_cat wrote:
         | You've released quite a few projects lately, very impressive.
         | 
         | Are you using LLMs for parts of the coding?
         | 
         | What's your work flow when approaching a new project like this?
        
           | littlestymaar wrote:
           | > Are you using LLMs for parts of the coding?
           | 
           | I can't talk about the code, but the readme and commit
           | messages are most likely LLM-generated.
           | 
           | And when you take into account that the first commit happened
           | just three hours ago, it feels like the entire project has
           | been vibe coded.
        
             | Neywiny wrote:
             | Hard disagree. Initial commit was 6k LOC. Author could've
             | spent years before committing. Ill advised but not
             | impossible.
        
               | littlestymaar wrote:
               | Why would you make Claude write your commit message for a
               | commit you've spent years working on though?
        
           | lulzx wrote:
           | Claude Code.
        
         | jeffbee wrote:
         | What's fast about mmap?
        
         | jonstewart wrote:
         | What's the fidelity like compared to tika?
        
           | lulzx wrote:
           | The accuracy difference is marginal (1-2%) but the speed
           | difference is massive.
        
       | agentifysh wrote:
       | excellent stuff what makes zig so fast
        
         | observationist wrote:
         | Not being slow - they compile straight to bytecode, they aren't
         | interpreted, and have aggressive, opinionated optimizations
         | baked in by default, so it's even faster than compiled c (under
         | default conditions.)
         | 
         | Contrasted with python, which is interpreted, has a clunky
         | runtime, minimal optimizations, and all sorts of choices that
         | result in slow, redundant, and also slow, performance.
         | 
         | The price for performance is safety checks, redundancy, how
         | badly wrong things can go, and so on.
         | 
         | A good compromise is luajit - you get some of the same
         | aggressive optimizations, but in an interpreted language, with
         | better-than-c performance but interpreted language convenience,
         | access to low level things that can explode just as
         | spectacularly as with zig or c, but also a beautiful language.
        
           | agentifysh wrote:
           | will add this to the list, now learning new languages is less
           | of a barrier with LLMs
        
           | Zambyte wrote:
           | Zig is safer than C under default conditions, not faster. By
           | default does a lot of illegal behavior safety checking, such
           | as array and slice bounds checking, numeric overflow
           | checking, and invalid union access checking. These features
           | are disabled by certain (non default) build modes, or
           | explicitly disabled at a per scope level.
           | 
           | It may be easier to write code that runs faster in Zig than
           | in C under similar build optimization levels, because writing
           | high performance C code looks a lot like writing idiomatic
           | Zig code. The Zig standard library offers a lot of structures
           | like hash maps, SIMD primitives, and allocators with
           | different performance characteristics to better fit a given
           | use-case. C application code often skips on these things
           | simply because it is a lot more friction to do in C than in
           | Zig.
        
         | AndyKelley wrote:
         | It makes your development workflow smooth enough that you have
         | the time and energy to do stuff like all the bullet points
         | listed in https://news.ycombinator.com/item?id=46437289
        
       | mpeg wrote:
       | very nice, it'd be good to see a feature comparison as when I use
       | mupdf it's not really just about speed, but about the level of
       | support of all kinds of obscure pdf features, and good level of
       | accuracy of the built-in algorithms for things like handling two-
       | column pages, identifying paragraphs, etc.
       | 
       | the licensing is a huge blocker for using mupdf in non-OSS tools,
       | so it's very nice to see this is MIT
       | 
       | python bindings would be good too
        
         | lulzx wrote:
         | added a comparison, will improve further.
         | https://github.com/Lulzx/zpdf?tab=readme-ov-file#comparison-...
         | 
         | also, added python bindings.
        
           | mpeg wrote:
           | thanks, claude, I guess haha
           | 
           | as others have commented, I think while this is a nice
           | portfolio piece, I would worry about its longevity as a vibe
           | coded project
        
       | odie5533 wrote:
       | Now we just need Python bindings so I can use it in my trash
       | language of choice.
        
         | lulzx wrote:
         | added python bindings!
        
           | hiq wrote:
           | Were you working on it already, or did it take you less than
           | 17 minutes to commit https://github.com/Lulzx/zpdf/commit/9f5
           | a7b70eb4b53672c0e4d8... ?
        
       | littlestymaar wrote:
       | - First commit 3hours ago.
       | 
       | - commit message: LLM-generated.
       | 
       | - README: LLM-generated.
       | 
       | I'm not convinced that projects vibe coded over the evening
       | deserve the HN front page...
       | 
       | Edit: and of course the author's blog is also full of AI slop...
       | 
       | 2026 hasn't even started I already hate it.
        
         | kingkongjaffa wrote:
         | Wait, but why?
         | 
         | If it's really better than what we had before, what does it
         | matter how it was made? It's literally hacked together with the
         | tools of the day (LLMs) isn't that the very hacker ethos?
         | Patching stuff together that works in a new and useful way.
         | 
         | 5x speed improvements on pdf text extraction might be great for
         | some applications I'm not aware of, I wouldn't just dismiss it
         | out of hand because the author used $robot to write the code.
         | 
         | Presumably the thought to make the thing in the first place and
         | decide what features to add and not add was more important than
         | how the code is generated?
        
       | forgotpwd16 wrote:
       | 74910,74912c187768,187779       < [Example 1: If you want to use
       | the code conversion facetcodecvt_utf8to output tocouta UTF-8
       | multibyte sequence       < corresponding to a wide string, but
       | you don't want to alter the locale forcout, you can write
       | something like:\237 D.27.21954
       | \251ISO/IECN4950wstring_convert<std::codecvt_utf8<wchar_t>>
       | myconv;       < std::string mbstring =
       | myconv.to_bytes\050L"Hello\134n"\051;       ---       >       >
       | [Example 1: If you want to use the code conversion facet
       | codecvt_utf8 to output to cout a UTF-8 multibyte sequence       >
       | corresponding to a wide string, but you don't want to alter the
       | locale for cout, you can write something like:       >       > SS
       | D.27.2       > 1954       >       > (c) ISO/IEC       > N4950
       | >       > wstring_convert<std::codecvt_utf8<wchar_t>> myconv;
       | > std::string mbstring = myconv.to_bytes(L"Hello\n");
       | 
       | Is indeed faster but output is messier.
        
       ___________________________________________________________________
       (page generated 2025-12-30 23:00 UTC)