[HN Gopher] Show HN: Defuddle, an HTML-to-Markdown alternative t...
       ___________________________________________________________________
        
       Show HN: Defuddle, an HTML-to-Markdown alternative to Readability
        
       Defuddle is an open-source JS library I built to parse and extract
       the main content and metadata from web pages. It can also return
       the content as Markdown.  I built Defuddle while working on
       Obsidian Web Clipper[1] (also MIT-licensed) because Mozilla's
       Readability[2] appears to be mostly abandoned, and didn't work well
       for many sites.  It's still very much a work in progress, but I
       thought I'd share it today, in light of the announcement that
       Mozilla is shutting down Pocket. This library could be helpful to
       anyone building a read-it-later app.  Defuddle is also available as
       a CLI:  https://github.com/kepano/defuddle-cli  [1]
       https://github.com/obsidianmd/obsidian-clipper  [2]
       https://github.com/mozilla/readability
        
       Author : kepano
       Score  : 378 points
       Date   : 2025-05-22 21:40 UTC (1 days ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | busymom0 wrote:
       | In the playground, after I enter a url, I can't seem to figure
       | out how to submit it to fetch the url? I tried pressing the
       | return key on iOS keyboard but it didn't do anything. Am I
       | missing something?
        
         | kepano wrote:
         | The input is there to test the url option -- which I admit is a
         | bit confusing, so I have removed it for now. I haven't found a
         | good and free way to proxy requests from a GitHub page (yet).
        
       | rcarmo wrote:
       | The Python analogues seem to be well maintained. I did my own
       | implementation of the Readability algorithm years ago and dropped
       | it in favor them, and I have a few scrapers going strong with
       | regular updates.
        
         | kepano wrote:
         | Are there any in particular you can recommend?
        
           | khimaros wrote:
           | not parent, but this one looks maintained
           | https://github.com/buriy/python-readability
        
       | fkfyshroglk wrote:
       | For those not in the know:
       | [Readability](https://github.com/mozilla/readability)
        
       | billconan wrote:
       | Are you using ai models behind the scenes? I saw Gemini and
       | others in the code. I am asking mainly to understand the cost of
       | using yours vs. readability. Thank!
        
         | kepano wrote:
         | No it's all rules-based. I think the code you're referring to
         | is "extractors", which are website-specific rules that I'm
         | working on to standardize the output from sites with comments
         | threads (e.g. HN, Reddit) and conversational chats (ChatGPT,
         | Claude, Gemini).
        
           | pugio wrote:
           | I would love something which reliably extracted a markdown
           | back/forth from all the main LLM providers. I tried
           | `defuddle` on a shared Gemini URL and it returned nothing but
           | the "Sign In" link. Maybe I'm using your extractor wrong? How
           | are you managing to get the rendered conversation HTML?
        
             | bambax wrote:
             | I think most LLM APIs return markdown and the conversion
             | md->html happens after; so if you query the API directly
             | you get markdown "for free".
        
       | tmpfs wrote:
       | Interesting as I was researching this recently and certainly not
       | impressed with the quality of the Readability implementations in
       | various languages. Although Readability.js was clearly the best,
       | it being Javascript didn't suit my project.
       | 
       | In the end I found the python trifatura library to extract the
       | best quality content with accurate meta data.
       | 
       | You might want to compare your implementation to trifatura to see
       | if there is room for improvement.
        
         | fabmilo wrote:
         | reference to the library:
         | https://trafilatura.readthedocs.io/en/latest/
         | 
         | for the curious: Trafilatura means "extrusion" in Italian.
         | 
         | | This method creates a porous surface that distinguishes pasta
         | trafilata for its extraordinary way of holding the sauce.
         | search maccheroni trafilati vs maccheroni lisci :)
         | 
         | (btw I think you meant trafilatura not trifatura)
        
           | thm wrote:
           | Been using it since day one but development has stalled quite
           | a bit since 2.0.0.
        
         | acrophobic wrote:
         | > ...it being Javascript didn't suit my project.
         | 
         | If you're using Go, I maintain Go ports of Readability[0] and
         | Trafilatura[1]. They're actively maintained, and for
         | Trafilatura, the extraction performance is comparable to the
         | Python version.
         | 
         | [0]: https://github.com/go-shiori/go-readability
         | 
         | [1]: https://github.com/markusmobius/go-trafilatura
        
           | breadchris wrote:
           | this is what i came here to see, thanks!
        
         | winddude wrote:
         | It's a bit old, but I bench marked a number of the web
         | extraction tools years ago, https://github.com/Nootka-io/wee-
         | benchmarking-tool, resiliparse-plain was my clear winner at the
         | time.
        
       | creakingstairs wrote:
       | I was just looking at obsidian web-clipper's source code because
       | I've been quite impressed at its markdown conversion results and
       | came across Defuddle in there. I'll be using for my bespoke read-
       | it-later/ knowledge-base app, so thank you in advance :D
        
       | input_sh wrote:
       | A bit off-topic, but I'm very excited to see the launch of Bases!
       | I've obsessively followed the roadmap for like a year awaiting
       | this day and have been frequently disappointed to still see it
       | stuck somewhere under "planned".
       | 
       | Not that I didn't already implement a read-it-later solution with
       | Obsidian+Dataview, but this definitely makes things simpler!
        
         | jeanlucas wrote:
         | Didn't it release just some days ago?
        
         | sn0n wrote:
         | Bases?
        
           | input_sh wrote:
           | https://help.obsidian.md/bases
           | 
           | Note that I'm using a preview (catalyst) version, it will
           | reach stable soon. I'm assuming kepano will submit it here
           | then.
        
       | inhumantsar wrote:
       | can confirm that readability seems to be on life support. I used
       | it slurp, an obsidian plugin which serves the same basic purpose
       | as web clipper, and always had a hard time getting PRs reviewed
       | and merged.
       | 
       | i started working on my own alternative but life (and web
       | clipper) derailed the work.
       | 
       | it's funny. somehow slurp keeps gaining new users even though web
       | clipper exists. so i might have to refactor it to use your
       | library sometime soon even though I don't use slurp myself
       | anymore.
        
       | 90s_dev wrote:
       | Neat. With ~3 more lines of code, you could get a URL and render
       | it in simpler HTML and be a full fledged replacement.
        
       | khaki54 wrote:
       | seems pretty much perfect including obsidian clipper. Thanks!
        
       | shrinks99 wrote:
       | I've been super happy with Obsidian Web Clipper! It's worked
       | really well for me with the one exception of importing publish
       | dates (which is more than forgivable !)
        
       | jeanlucas wrote:
       | Obsidian Web Clipper is a great tool to turn chatGPT
       | conversations in markdown, or to just print it (believe me, it is
       | a user case)
        
         | T0Bi wrote:
         | I just ask ChatGPT to provide the summary or whatever I need as
         | a markdown file.
        
         | emaro wrote:
         | Not sure about other clients, but Kagi Assistant directly
         | offers to save a conversation as Markdown. Using Obsidian's
         | web-clipper is a good idea too though.
        
         | kouru225 wrote:
         | Is that a paid plugin?
        
       | acrophobic wrote:
       | Is Mozilla's Readability really abandoned? The latest release
       | (v0.6.0) is just 2 months ago, and its maintainer (Gijs) is
       | pretty active on responding issues.
        
         | khasan222 wrote:
         | That codebase definitely leaves much to be desired, I've
         | already had to fork it for work in order to fix some bugs.
         | 
         | 1 such bug, find a foreign language with commas in between
         | numbers instead of periods, like Dutch(I think), and a lot of
         | prices on the page. It'll think all the numbers are relevant
         | text.
         | 
         | And of course I tried to open a pr and get it merged, but they
         | require tests, and of course the tests don't work on the page
         | Im testing. It's just very snafu imho
        
           | fabrice_d wrote:
           | This seems to be https://github.com/mozilla/readability/pull/
           | 853#issuecomment... and I think their expectations are pretty
           | reasonable.
        
             | khasan222 wrote:
             | Meh, maybe I'm standing too close to the problem, Idk. It
             | is always frustrating trying to use a tool, and it not work
             | though. I know it's free and all, but then I feel like
             | helping people make good contributions is paramount in
             | maintaining and fixing bugs.
             | 
             | Clearly the comma thing is a bug, it's the lack of wanting
             | to fix it actually that is a bit disheartening, and why I
             | think it is a deadish repo
        
       | Tsarp wrote:
       | Been using the obsidian clipper since it was out and this is a
       | really neat. The per website profile based extraction is awesome.
       | 
       | Even if you are not a obsidian user, the markdown extraction
       | quality is the most reliable Ive seen.
        
         | audessuscest wrote:
         | thanks for the tip!
        
       | jonplackett wrote:
       | Does anyone know why readers don't work for some websites where
       | it looks like they should - ie normal article with lots of text.
       | 
       | You just get a completely white page (on the iPhone reader).
       | Usually it's a news website.
       | 
       | Is this the website intentionally obscuring the content to ensure
       | they can serve their ads? If so how do they go about it?
        
         | miki123211 wrote:
         | Cookie and "we care about your privacy" banners are often the
         | cause here, especially if you're in the EU / UK / possibly
         | California[1].
         | 
         | On some websites, those are just modals that obscure the
         | content, something that reader mode can usually deal with just
         | fine, but on others, they're implemented as redirects or
         | rendered server-side.
         | 
         | If reader mode doesn't work, dismiss those first and try again.
        
       | revskill wrote:
       | Interesting that Markdown does not support form element.
        
       | Andr2Andr wrote:
       | Serious question - who and why would be using this tool? What is
       | the use case? In other comments I have only seen exporting
       | ChatGPT conversations to md
        
         | degosuke wrote:
         | I use LogSeq a lot - and having the option to scrape a website
         | with only the text in MD seems like a great fit.
        
         | rollcat wrote:
         | This is a library, not a tool. You can use it for a number of
         | purposes:
         | 
         | - Providing "reader mode" for your visitors
         | 
         | - Using it in a browser extension to add reader mode
         | 
         | - Scrapping
         | 
         | - Plugging it into a [reverse] proxy that automatically removes
         | unnecessary bloat from pages, for e.g. easier access on retro
         | hardware <https://web.archive.org/web/20240621144514/https://hu
         | mungus....> (archive.org link, because the website goes down
         | regularly)
        
       | timdeve wrote:
       | Looks good, I'm gonna try to swap readability in my RSS reader
       | with this.
       | 
       | And with Pocket going away I might have to add save it later to
       | it...
        
       | ulrischa wrote:
       | I have build something similar:https://devkram.de/markydown but
       | with php. Easy for self hosting
        
       | ioma8 wrote:
       | Tried it on some webpages, doesnt work well.
        
       | severusdd wrote:
       | This is very cool! Given how messy and busy many websites have
       | become, we really need a robust markdown converter that lets
       | readers focus on reading the content. Nice to see something
       | stepping up where Readability left off.
       | 
       | Thank you for picking up this work :-)
        
       | ricardonunez wrote:
       | I'll give it a try. I'm not happy with my current setup for
       | markdown to HTML on the wysiwyg editor I'm using, this may
       | provide better results if I go with my own tool bar and editor.
        
       | binarymax wrote:
       | Really nice work. I appreciate the example with JSDOM as that's
       | exactly how I use readability, and this looks like a nice drop-in
       | replacement.
       | 
       | Question: How did you validate this? You say it works better than
       | readability but I don't see any tests or datasets in the repo to
       | evaluate accuracy or coverage. Would it be possible to share that
       | as well?
        
         | kepano wrote:
         | Currently I am relying on manual testing and user feedback, but
         | yes, I'd like to add tests.
         | 
         | Defuddle works quite differently from Readability. Readability
         | tends to be overly conservative and tends to remove useful
         | content because it tests blocks to find the beginning and end
         | of the "main" content.
         | 
         | Defuddle is able to run multiple passes and detect if it
         | returned no content to try and expand its results. It also uses
         | a greater variety of techniques to clean the content -- for
         | example, by using a page's mobile styles to detect content that
         | can be hidden.
         | 
         | Lastly, Defuddle is not only extracting the content but also
         | standardizing the output (which Readability doesn't do). For
         | example footnotes and code blocks all aim to output a single
         | format, whereas Readability keeps the original DOM intact.
        
       | ahsd1 wrote:
       | Cool. Im looking for something similar but for stripping
       | signatures and boilerplate disclaimers from html email. Could
       | this work for that?
        
       | infogulch wrote:
       | Since it's written in javascript is there any chance it could be
       | packaged as a bookmarklet?
        
       | miketromba wrote:
       | Excellent work. A modern alternative to readability was much
       | needed. This is especially useful for building clean web context
       | for LLMs. Thanks for open-sourcing this!
        
         | elcritch wrote:
         | I found LLMs are really good at taking a web page and
         | transforming it to markdown. Well rather commercial LLMs like
         | Claude and Gemini are.
         | 
         | Unfortunately I tried a bunch of hugging face mode on a I could
         | run on my MacBook and all of them ignored my prompts despite
         | trying every variation I could think of. Half the time they
         | just tried summarizing it and describing what JavaScript was.
         | :/
        
       ___________________________________________________________________
       (page generated 2025-05-23 23:01 UTC)