[HN Gopher] Zpdf: PDF text extraction in Zig
       ___________________________________________________________________
        
       Zpdf: PDF text extraction in Zig
        
       Author : lulzx
       Score  : 207 points
       Date   : 2025-12-30 19:57 UTC (1 days ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | lulzx wrote:
       | I built a PDF text extraction library in Zig that's significantly
       | faster than MuPDF for text extraction workloads.
       | 
       | ~41K pages/sec peak throughput.
       | 
       | Key choices: memory-mapped I/O, SIMD string search, parallel page
       | extraction, streaming output. Handles CID fonts, incremental
       | updates, all common compression filters.
       | 
       | ~5,000 lines, no dependencies, compiles in <2s.
       | 
       | Why it's fast:                 - Memory-mapped file I/O (no read
       | syscalls)       - Zero-copy parsing where possible       - SIMD-
       | accelerated string search for finding PDF structures       -
       | Parallel extraction across pages using Zig's thread pool       -
       | Streaming output (no intermediate allocations for extracted text)
       | 
       | What it handles:                 - XRef tables and streams (PDF
       | 1.5+)       - Incremental PDF updates (/Prev chain)       -
       | FlateDecode, ASCII85, LZW, RunLength decompression       - Font
       | encodings: WinAnsi, MacRoman, ToUnicode CMap       - CID fonts
       | (Type0, Identity-H/V, UTF-16BE with surrogate pairs)
        
         | tveita wrote:
         | What kind of performance are you seeing with/without SIMD
         | enabled?
         | 
         | From https://github.com/Lulzx/zpdf/blob/main/src/main.zig it
         | looks like the help text cites an unimplemented "-j" option to
         | enable multiple threads.
         | 
         | There is a "--parallel" option, but that is only implemented
         | for the "bench" command.
        
           | lulzx wrote:
           | I have now made parallel by default and added an option to
           | enable multiple threads.
           | 
           | I haven't tested without SIMD.
        
         | cheshire_cat wrote:
         | You've released quite a few projects lately, very impressive.
         | 
         | Are you using LLMs for parts of the coding?
         | 
         | What's your work flow when approaching a new project like this?
        
           | littlestymaar wrote:
           | > Are you using LLMs for parts of the coding?
           | 
           | I can't talk about the code, but the readme and commit
           | messages are most likely LLM-generated.
           | 
           | And when you take into account that the first commit happened
           | just three hours ago, it feels like the entire project has
           | been vibe coded.
        
             | Neywiny wrote:
             | Hard disagree. Initial commit was 6k LOC. Author could've
             | spent years before committing. Ill advised but not
             | impossible.
        
               | littlestymaar wrote:
               | Why would you make Claude write your commit message for a
               | commit you've spent years working on though?
        
               | Neywiny wrote:
               | 1. Be not good at or a fan of git when coding
               | 
               | 2. Be not good at or a fan of git when committing
               | 
               | Not sure what the disconnect is.
               | 
               | Now _if_ it were vibecoded, I wouldn 't be surprised. But
               | benefit of the doubt
        
               | Jach wrote:
               | We're well beyond benefit of the doubt these days. If it
               | looks like a duck... For me there wasn't any doubt, the
               | author's first top comment here was evidence enough, then
               | seeing the readme + random code + random commit message,
               | it's all obvious LLM-speak to me.
               | 
               | I don't particularly care, though, and I'm more positive
               | about LLMs than negative even if I don't (yet?) use them
               | very much. I think it's hilarious that a few people asked
               | for Python bindings and then bam, done, and one person is
               | like "..wha?" Yes, LLMs can do that sort of grunt work
               | now! How cool, if kind of pointless. Couldn't the cycles
               | have just been spent on trying to make muPDF better?
               | Though I see they're in C and AGPL, I suppose either is
               | motivation enough to do a rewrite instead. (This is MIT
               | Licensed though it's still unclear to me how 100% or even
               | large-% vibe-coded code deserves any copyright
               | protection, I think all such should generally be under
               | the Unlicense/public domain.)
               | 
               | If the intent of "benefit of the doubt" is to reduce
               | people having a freak out over anyone who dares use these
               | tools, I get that.
        
               | lulzx wrote:
               | I have updated the licence to WTFPL.
               | 
               | I'll try my best to make it a really good one!
        
               | littlestymaar wrote:
               | > I have updated the licence to WTFPL.
               | 
               | You still have no basis in claiming copyright protection
               | hence you cannot set a license on that code.
               | 
               | Instead of the WTFPL you should just write a disclaimer
               | that due to being machine generated and devoid of
               | creating work, the work is not protected by copyright and
               | free to be used without any license.
        
               | lulzx wrote:
               | hasn't world moved on from these things already?
        
           | lulzx wrote:
           | Claude Code.
        
         | jeffbee wrote:
         | What's fast about mmap?
        
           | rishabhaiover wrote:
           | it allows the program to reference memory without having to
           | manage it in the heap space. it would make the program faster
           | in a memory managed language, otherwise it would reduce the
           | memory footprint consumed by the program.
        
             | jeffbee wrote:
             | You mean it converts an expression like `buf[i]` into a
             | baroque sequence of CPU exception paths, potentially
             | involving a trap back into the kernel.
        
               | rishabhaiover wrote:
               | I don't fully understand the under the hood mechanics of
               | mmap, but I can sense that you're trying to convey that
               | mmap shouldn't be used a blanket optimization technique
               | as there are tradeoffs in terms of page fault overheads
               | (being at the mercy of OS page cache mechanics)
        
               | jibal wrote:
               | I think he's conveying that he doesn't know what he's
               | talking about. buf[i] generates the same code regardless
               | of whether mmap is being used. The first access to a page
               | will cause a trap that loads the page into memory, but
               | this is also true if the memory is read into.
        
               | StilesCrisis wrote:
               | Tradeoffs such as "if an I/O error occurs, the program
               | immediately segfaults." Also, I doubt you're I/O bound to
               | the point where mmap noticeably better than read, but I
               | guess it's fine for an experiment.
        
               | jibal wrote:
               | An I/O error on a mmapped file causes a SIGBUS, which the
               | program can catch and report.
               | 
               | And I/O bound programs are I/O bound whereas programs
               | that aren't, aren't, so it really isn't meaningful to
               | talk about whether "you" are I/O bound to the point that
               | it's significant--maybe you are, maybe you aren't. I
               | agree about experimentation.
        
           | kennethallen wrote:
           | Two big advantages:
           | 
           | You avoid an unnecessary copy. Normal read system call gets
           | the data from disk hardware into the kernel page cache and
           | then copies it into the buffer you provide in your process
           | memory. With mmap, the page cache is mapped directly into
           | your process memory, no copy.
           | 
           | All running processes share the mapped copy of the file.
           | 
           | There are a lot of downsides to mmap: you lose explicit error
           | handling and fine-grained control of when exactly I/O
           | happens. Consult the classic article on why sophisticated
           | systems like DBMSs do not use mmap:
           | https://db.cs.cmu.edu/mmap-cidr2022/
        
             | saidinesh5 wrote:
             | This is a very interesting link. I didn't expect mmap to be
             | less performant than read() calls.
             | 
             | I now wonder which use cases would mmap suit better - if
             | any...
             | 
             | > All running processes share the mapped copy of the file.
             | 
             | So something like building linkers that deal with read only
             | shared libraries "plugins" etc ..?
        
               | squirrellous wrote:
               | One reason to use shared memory mmap is to ensure that
               | even if your process crashes, the memory stays intact.
               | Another is to communicate between different processes.
        
             | commandersaki wrote:
             | _you lose explicit error handling_
             | 
             | I've never had to use mmap but this is always been the
             | issue in my head. If you're treating I/O as memory pages,
             | what happens when you read a page and it needs to "fault"
             | by reading the backing storage but the storage fails to
             | deliver? What can be said at that point, or does the
             | program crash?
        
             | nextaccountic wrote:
             | > Consult the classic article on why sophisticated systems
             | like DBMSs do not use mmap: https://db.cs.cmu.edu/mmap-
             | cidr2022/
             | 
             | Sqlite does (or can optionally use mmap). How come?
             | 
             | Is sqlite with mmap less reliable or anything?
        
               | jeffbee wrote:
               | I know that the spirit of HN will strike me down for
               | this, but sqlite is not a "sophisticated system". It
               | assumes the hardware is lawful neutral. Real hardware is
               | chaotic. Sqlite has a good reputation because it is very
               | easy to use. In fact this is the same reason programmers
               | like mmap: it is a hell of a shortcut.
        
         | jonstewart wrote:
         | What's the fidelity like compared to tika?
        
           | lulzx wrote:
           | The accuracy difference is marginal (1-2%) but the speed
           | difference is massive.
        
         | DannyBee wrote:
         | FWIW - mupdf is simply not fast. I've done lots of pdf indexing
         | apps, and mupdf is by far the slowest and least able to open
         | valid pdfs when it came to text extraction. It also takes
         | _tons_ of memory.
         | 
         | a better speed comparison would either be multi-process pdfium
         | (since pdfium was forked from foxit before multi-thread
         | support, you can't thread it), multi-threaded foxit, or
         | something like syncfusion (which is quite fast and supports
         | multiple threads). Or even single thread pdfium vs single
         | thread your-code.
         | 
         | These were always the fastest/best options. I can (and do)
         | achieve 41k pages/sec or better on these options.
         | 
         | The other thing it doesn't appear you mention is whether you
         | handle putting the words in reading order (IE how they appear
         | on the page), or only stream order (which varies in its
         | relation to apperance order) .
         | 
         | If it's only stream order, sure, that's really fast to do. But
         | also not anywhere near as helpful as reading order, which is
         | what other text-extraction engines do.
         | 
         | Looking at the code, it looks like the code to do reading order
         | exists, but is not what is being benchmarked or used by
         | default?
         | 
         | If so, this is really comparing apples and oranges.
        
         | littlestymaar wrote:
         | > I built
         | 
         | You didn't. Claude did. Like it did write this comment.
         | 
         | And you didn't even bother testing it before submitting, which
         | is insulting to everyone.
        
           | lulzx wrote:
           | tools are tools.
        
       | agentifysh wrote:
       | excellent stuff what makes zig so fast
        
         | observationist wrote:
         | Not being slow - they compile straight to bytecode, they aren't
         | interpreted, and have aggressive, opinionated optimizations
         | baked in by default, so it's even faster than compiled c (under
         | default conditions.)
         | 
         | Contrasted with python, which is interpreted, has a clunky
         | runtime, minimal optimizations, and all sorts of choices that
         | result in slow, redundant, and also slow, performance.
         | 
         | The price for performance is safety checks, redundancy, how
         | badly wrong things can go, and so on.
         | 
         | A good compromise is luajit - you get some of the same
         | aggressive optimizations, but in an interpreted language, with
         | better-than-c performance but interpreted language convenience,
         | access to low level things that can explode just as
         | spectacularly as with zig or c, but also a beautiful language.
        
           | agentifysh wrote:
           | will add this to the list, now learning new languages is less
           | of a barrier with LLMs
        
           | Zambyte wrote:
           | Zig is safer than C under default conditions, not faster. By
           | default does a lot of illegal behavior safety checking, such
           | as array and slice bounds checking, numeric overflow
           | checking, and invalid union access checking. These features
           | are disabled by certain (non default) build modes, or
           | explicitly disabled at a per scope level.
           | 
           | It may be easier to write code that runs faster in Zig than
           | in C under similar build optimization levels, because writing
           | high performance C code looks a lot like writing idiomatic
           | Zig code. The Zig standard library offers a lot of structures
           | like hash maps, SIMD primitives, and allocators with
           | different performance characteristics to better fit a given
           | use-case. C application code often skips on these things
           | simply because it is a lot more friction to do in C than in
           | Zig.
        
           | jibal wrote:
           | > they compile straight to bytecode
           | 
           | machine code, not https://en.wikipedia.org/wiki/Bytecode
           | 
           | > The price for performance is safety checks
           | 
           | In Zig, non-ReleaseFast build modes have significant safety
           | checks.
           | 
           | > luajit ... with better-than-c performance
           | 
           | No.
        
         | AndyKelley wrote:
         | It makes your development workflow smooth enough that you have
         | the time and energy to do stuff like all the bullet points
         | listed in https://news.ycombinator.com/item?id=46437289
        
           | forgotpwd16 wrote:
           | >you have the time and energy to do stuff like all the bullet
           | points listed
           | 
           | Don't disagree but in specific case, per the author, project
           | was made via Claude Code. Although could as well be that Zig
           | is better as LLM target. Noticed many new vibe projects
           | decide to use Zig as target.
        
       | mpeg wrote:
       | very nice, it'd be good to see a feature comparison as when I use
       | mupdf it's not really just about speed, but about the level of
       | support of all kinds of obscure pdf features, and good level of
       | accuracy of the built-in algorithms for things like handling two-
       | column pages, identifying paragraphs, etc.
       | 
       | the licensing is a huge blocker for using mupdf in non-OSS tools,
       | so it's very nice to see this is MIT
       | 
       | python bindings would be good too
        
         | lulzx wrote:
         | added a comparison, will improve further.
         | https://github.com/Lulzx/zpdf?tab=readme-ov-file#comparison-...
         | 
         | also, added python bindings.
        
           | mpeg wrote:
           | thanks, claude, I guess haha
           | 
           | as others have commented, I think while this is a nice
           | portfolio piece, I would worry about its longevity as a vibe
           | coded project
        
             | chanbam wrote:
             | If he made something legitimately useful, who cares how?
        
               | littlestymaar wrote:
               | It seems that he didn't even test it before submitting
               | though...
               | 
               | The author has created 30 new projects on github, in half
               | a dozen different programming language, over the past
               | month alone, and he also happen to have an LLM-generated
               | blog. I think it's fair to say it's not "legitimately
               | useful" except as a way for the author to fill his resume
               | as he's looking for a job.
               | 
               | This kind of behavior is toxic.
        
               | mpeg wrote:
               | Exactly this, I like to give the benefit of the doubt to
               | people but pushing huge chunks of code this quickly shows
               | the whole thing is vibe coded
               | 
               | I actually don't mind LLM generated code when it's been
               | manually reviewed, but this and a quick look through
               | other submissions makes me realise the author is simply
               | trying to pad their resume with OSS projects. Respect the
               | hustle, but it shows a lack of respect for other's time
               | to then submit it to show HN
        
               | lulzx wrote:
               | Fair point. I won't submit here again until I've put in
               | the work to make something that respects people's time to
               | evaluate it. Lesson learned. :)
        
       | odie5533 wrote:
       | Now we just need Python bindings so I can use it in my trash
       | language of choice.
        
         | lulzx wrote:
         | added python bindings!
        
           | hiq wrote:
           | Were you working on it already, or did it take you less than
           | 17 minutes to commit https://github.com/Lulzx/zpdf/commit/9f5
           | a7b70eb4b53672c0e4d8... ?
        
             | qeternity wrote:
             | Claude Code.
        
               | littlestymaar wrote:
               | + not testing the output.
        
       | littlestymaar wrote:
       | - First commit 3hours ago.
       | 
       | - commit message: LLM-generated.
       | 
       | - README: LLM-generated.
       | 
       | I'm not convinced that projects vibe coded over the evening
       | deserve the HN front page...
       | 
       | Edit: and of course the author's blog is also full of AI slop...
       | 
       | 2026 hasn't even started I already hate it.
        
         | kingkongjaffa wrote:
         | Wait, but why?
         | 
         | If it's really better than what we had before, what does it
         | matter how it was made? It's literally hacked together with the
         | tools of the day (LLMs) isn't that the very hacker ethos?
         | Patching stuff together that works in a new and useful way.
         | 
         | 5x speed improvements on pdf text extraction might be great for
         | some applications I'm not aware of, I wouldn't just dismiss it
         | out of hand because the author used $robot to write the code.
         | 
         | Presumably the thought to make the thing in the first place and
         | decide what features to add and not add was more important than
         | how the code is generated?
        
           | utopiah wrote:
           | > If it's really better than what we had before
           | 
           | That's a very big if. The whole point is that what we had
           | before was made slowly. This was made quickly. In itself it's
           | not better but what it typically means is hours and hours of
           | testing. Going through painful problems that highlight
           | idiosyncrasies of the problem space. Things that are really
           | weird and specific to whatever the tool is trying to address.
           | 
           | In such cases we can be expect that with very little time
           | very few things were tested and tested properly (including a
           | comment mentioned how tests were also generated). "We" the
           | audience of potentially interested users have then to do that
           | work (as plenty did commenting on that post).
           | 
           | IMHO what you bring forward is precisely that :
           | 
           | - can the new "solution" actually pass ALL the tests the
           | previous one did? More?
           | 
           | This should be brought to the top and the actual compromises
           | can then be understood, "we" can then decide if it's "better"
           | for our context. In some cases faster with lossy output is
           | actually better, in others absolutely not. The difference
           | between the new and the old solutions isn't binary and have
           | no visibility on that is what makes such a process nothing
           | more than yet another showcase that LLMs can indeed produce
           | "something" that is absolutely boring while consuming a TON
           | of resources, including our own attention.
           | 
           | TL;DR: there should be test "harness" made by 3rd parties (or
           | from well known software it is the closest too) that an LLM
           | generated piece of code should pass before being actually
           | compared.
        
             | utopiah wrote:
             | related https://news.ycombinator.com/item?id=46437688
        
         | dmytrish wrote:
         | ...and it does not work. I tried it on ~10 random pdfs,
         | including very simple ones (e.g. a hello world from typst), it
         | segfaults on every single one.
        
           | forgotpwd16 wrote:
           | Tried few and works. Maybe you've older or newer Zig version
           | than whatever project targets. (Mine is 0.15.2.)
        
             | dmytrish wrote:
             | ~/c/t/s/zpdf (main)> zig version        0.15.2
             | 
             | Sky is blue, water is wet, slop does not work.
        
         | ncgl wrote:
         | Using Ai isn't lazier than your regurgitated dismissal, to be
         | fair.
        
           | littlestymaar wrote:
           | Using AI is not necessarily lazy.
           | 
           | Using AI lazily is a problem though. Writing code has never
           | been the most important part of software development, making
           | sure that the code does what the user needs is what takes
           | most of the time. But from the github issues and the comment
           | here from the few who have tested the tool, it looka like the
           | author didn't even test the AI output on real PDF.
           | 
           | If you use AI to build in 3 month something that would have
           | taken a year without it, then cool. But here we're talking
           | about someone who's spending 2-3 hours every other day
           | building a new fake software project to pad his resume. This
           | isn't something anyone should endorse.
        
       | forgotpwd16 wrote:
       | 74910,74912c187768,187779       < [Example 1: If you want to use
       | the code conversion facetcodecvt_utf8to output tocouta UTF-8
       | multibyte sequence       < corresponding to a wide string, but
       | you don't want to alter the locale forcout, you can write
       | something like:\237 D.27.21954
       | \251ISO/IECN4950wstring_convert<std::codecvt_utf8<wchar_t>>
       | myconv;       < std::string mbstring =
       | myconv.to_bytes\050L"Hello\134n"\051;       ---       >       >
       | [Example 1: If you want to use the code conversion facet
       | codecvt_utf8 to output to cout a UTF-8 multibyte sequence       >
       | corresponding to a wide string, but you don't want to alter the
       | locale for cout, you can write something like:       >       > SS
       | D.27.2       > 1954       >       > (c) ISO/IEC       > N4950
       | >       > wstring_convert<std::codecvt_utf8<wchar_t>> myconv;
       | > std::string mbstring = myconv.to_bytes(L"Hello\n");
       | 
       | Is indeed faster but output is messier. And doesn't handle
       | Unicode in contrast to mutool that does. (Probably also explains
       | the big speed boost.)
        
         | lulzx wrote:
         | fixed.
        
           | TZubiri wrote:
           | Lol, but there's 100 competitors in the PDF text extraction
           | space, some are multi million dollar industries: AWS
           | textract, ABBY PDFreader, PDFBox, I think you may be
           | underestimating the challenge here.
        
           | forgotpwd16 wrote:
           | Yeah, sorry for confusion. When said Unicode, meant foreign
           | text rather (just) the unescaped symbols, e.g. Greek. At one
           | random Greek textbook[0], zpdf output is (extract | head
           | -15):                 01F9020101FC020401F9020301FB02070205020
           | 800030209020701FF01F90203020901F9012D020A0201020101FF01FB01FE
           | 0208        0200012E0219021802160218013202120222 0209021D0212
           | 021D012E013202200222000301FA021A0220021C022002160213012E02220
           | 00F000301F90206012C            020301FF02000205020101FC020901
           | F90003020001F9020701F9020E020802000205020A        01FC028C021
           | 3021B022002230221021800030200012E021902180216021201320221021A
           | 012E00030209021D0212021D012E013202200222000301FA021A0220021C0
           | 22002160213012E0222000F000301F90206012C              0200020D
           | 02030208020901F90203020901FF0203020502080003012B020001F9012B0
           | 20001F901FA0205020A01FD01FE0208
           | 020201300132012E012F021A012F0210021B013202200221012E0222 0209
           | 021D0212021D012E013202200222000301FA021A0220021C0220021602130
           | 12E0222000F000301F90206012C
           | 
           | This for entire book. Mutool extracts the text just fine.
           | 
           | [0]: https://repository.kallipos.gr/handle/11419/15087
        
             | lulzx wrote:
             | sorry, I haven't yet figured out non-latin with tounicode
             | references.
        
             | lulzx wrote:
             | works now!
             | 
             | ALEKsANDROS TRIANTAPhULLIDES Kathegetes Tmematos Biologias,
             | APTh                    NIKOLETA KARAISKOU
             | Epikoure Kathegetria Tmematos Biologias, APTh
             | KONSTANTINOS GKAGKABOUZES          Metadidaktoras Tmematos
             | Biologias, APTh
             | Gonidiomata          Dome, Leitourgia kai Epharmoges
        
               | forgotpwd16 wrote:
               | Nice! Speed wasn't even compromised. Still 5x when
               | benching. Also saw now there's page with tool compiled to
               | wasm. Cool.
        
               | lulzx wrote:
               | thanks! :)
        
         | TZubiri wrote:
         | In my experience with parsing PDFs, speed has never been an
         | issue, it has always been a matter of quality.
        
           | DetroitThrow wrote:
           | I tried a small PDF and got a memory error. It's definitely
           | much faster than MuPDF on that file.
        
             | littlestymaar wrote:
             | "The fastest PDF extractor is the one that crashes at the
             | beginning of the file" or something.
        
       | amkharg26 wrote:
       | Impressive performance gains! 5x faster than MuPDF is
       | significant, especially for applications processing large volumes
       | of PDFs. Zig's memory safety without garbage collection overhead
       | makes it ideal for this kind of performance-critical work.
       | 
       | I'm curious about the trade-offs mentioned in the comments
       | regarding Unicode handling. For document analysis pipelines (like
       | extracting text from technical documentation or research papers),
       | robust Unicode support is often critical.
       | 
       | Would be interesting to see benchmarks on different PDF types -
       | academic papers with equations, scanned documents with OCR
       | layers, and complex layouts with tables. Performance can vary
       | wildly depending on the document structure.
        
         | polyaniline wrote:
         | What memory safety?
        
           | Retr0id wrote:
           | (the comment was written by an llm bot)
        
       | nullorempty wrote:
       | Tomorrow's headlines
       | 
       | fpdf
       | 
       | jpdf
       | 
       | cpdf
       | 
       | cpppdf
       | 
       | bfpdf
       | 
       | ppdf
       | 
       | ...
       | 
       | opdf
        
       | pm2222 wrote:
       | What's the format that's perhaps free, easy to parse and render?
       | Build one please.
        
       | fainpul wrote:
       | These vibe coded tests are terrible:
       | 
       | https://github.com/Lulzx/zpdf/blob/main/python/tests/test_zp...
        
         | lulzx wrote:
         | this is more like a quick test for python bindings, the zig
         | files have tests within them for broad range of things.
        
       | xvilka wrote:
       | Test it on major PDF corpora[1]
       | 
       | [1] https://github.com/pdf-association/pdf-corpora
        
       | manmal wrote:
       | Is there the possibility to hook in OCR for text blocks flattened
       | into an image, maybe with some callback? That's my biggest gripe
       | with dealing with PDFs.
        
       | ceving wrote:
       | The spacing issue isn't working quite right yet.
       | zpdf extract texbook.pdf | grep -m1 Stanford         DONALD E.
       | KNUTHStanford UniversityIllustrations by
        
       ___________________________________________________________________
       (page generated 2025-12-31 23:01 UTC)