[HN Gopher] Parsing PDFs (and more) in Elixir using Rust
       ___________________________________________________________________
        
       Parsing PDFs (and more) in Elixir using Rust
        
       Author : bustylasercanon
       Score  : 25 points
       Date   : 2025-01-29 21:05 UTC (1 hours ago)
        
 (HTM) web link (www.chriis.dev)
 (TXT) w3m dump (www.chriis.dev)
        
       | joshchernoff wrote:
       | FYI: your preview image from the html header meta tag is broken.
        
         | bustylasercanon wrote:
         | Thanks! I need to fix that
        
       | cpursley wrote:
       | I've been thinking a lot about how to accomplish various RAG
       | things in Elixir (for LLM applications). PDF is one of the
       | missing pieces, so glad to see work here. The really tricky part
       | is not just parsing out the text (you can just call the pdftotext
       | unix command line utility for that), but accurately pulling out
       | things like complex tables, etc in a way that could be
       | chunked/post processed in a useful way. I'd love to see something
       | like Unstructured or Marker but in Rust (i.e., fast) that Elixir
       | could NIF out to it. And maybe some kind of hybrid system that
       | uses open llm models with vision capabilities. Ref:
       | 
       | - https://github.com/Unstructured-IO/unstructured
       | 
       | - https://github.com/VikParuchuri/marker
        
         | cpursley wrote:
         | Well derp, I should have read the linked extractous repo. This
         | looks like the extract solution I've been after (see what I did
         | there).
        
         | vikp wrote:
         | Hey, I'm the author of marker - thanks for sharing. Most of the
         | processing time is model inference right now. I've been
         | retraining some models lately onto new architectures to improve
         | speed (layout, tables, LaTeX OCR).
         | 
         | We recently integrated gemini flash (via the --use_llm flag),
         | which maybe moves us towards the "hybrid system" you mentioned.
         | Hoping to add support for other APIs soon, but focusing on
         | improving quality/speed now.
         | 
         | Happy to chat if anyone wants to talk about the difficulties of
         | parsing PDFs, or has feedback - email in profile.
        
       ___________________________________________________________________
       (page generated 2025-01-29 23:00 UTC)