[HN Gopher] Minifying HTML for GPT-4o: Remove all the HTML tags
       ___________________________________________________________________
        
       Minifying HTML for GPT-4o: Remove all the HTML tags
        
       Author : edublancas
       Score  : 50 points
       Date   : 2024-09-05 13:51 UTC (1 days ago)
        
 (HTM) web link (blancas.io)
 (TXT) w3m dump (blancas.io)
        
       | giancarlostoro wrote:
       | I wonder if this is due to some template engines looking
       | minimalist like that. I think maybe Pug?
       | 
       | https://github.com/pugjs/pug?tab=readme-ov-file#syntax
       | 
       | It is whitespace sensitive though, but essentially looks like
       | that. I doubt this is the only unique template engine like this
       | though.
        
       | cj wrote:
       | Related article from 4 days ago (with comments on scraping,
       | specifically discussing removing HTML tags)
       | 
       | https://news.ycombinator.com/item?id=41428274
       | 
       | Edit: looks like it's actually the same author
        
       | throwup238 wrote:
       | I don't think that Mercury Prize table is a representative
       | example because each column has an obviously unique structure
       | that the LLM can key in on: (year) (Single Artist/Album pair)
       | (List of Artist/Album pairs) (image) (citation link)
       | 
       | I think a much better test would be something like "List of
       | elements by atomic properties" [1] that has a lot of adjacent
       | numbers in a similar range and overlapping first/last column
       | types. However, the danger with that table might be easy for the
       | LLM to infer just from the element names since they're well known
       | physical constants. The table of counties by population density
       | might be less predictable [2] or list of largest cities [3]
       | 
       | The test should be repeated with every available sorting function
       | too, to see if that causes any new errors.
       | 
       | [1]
       | https://en.wikipedia.org/wiki/List_of_elements_by_atomic_pro...
       | 
       | [2]
       | https://en.wikipedia.org/wiki/List_of_countries_and_dependen...
       | 
       | [3] https://en.wikipedia.org/wiki/List_of_largest_cities#List
        
         | edublancas wrote:
         | thanks a lot for the feedback! you're right, this is much
         | better input data. I'll re-run the code with these tables!
        
         | cal85 wrote:
         | Good points. But I feel like even with the cities article it
         | could still 'cheat' by recognising what the data is supposed to
         | be and filling in the blanks. Does it even need to be real
         | though? What about generating a fake article to use as a test
         | so it can't possibly recognise the contents? You could even get
         | GPT to generate it, just give it the 'Largest cities' HTML and
         | tell it to output identical HTML but with all the names and
         | statistics changed randomly.
        
         | curl-up wrote:
         | Additionally, using any Wiki page is misleading, as LLMs have
         | seen their format many times during training, and can probably
         | reproduce the original HTML from the stripped version fairly
         | well.
         | 
         | Instead, using some random, messy, scattered-with-spam site
         | would be a much more realistic test environment.
        
       | topaz0 wrote:
       | Is .8 or .9 considered good enough accuracy for something as
       | simple as this?
        
         | edublancas wrote:
         | I'd say how much is good enough highly depends on your use
         | case. For something that still has to be reviewed by a human, I
         | think even .7 is great; if you're planning to automate
         | processes end-to-end, I'd aim for higher than .95
        
         | LunaSea wrote:
         | Well, when "simply" extracting the core text of an article is a
         | task where most solutions (rule-based, visual, traditional
         | classifiers and LLMs) rarely score above 0.8 in precision on
         | datasets with a variety of websites and / or multilingual
         | pages, I would consider that not too bad.
        
         | moralestapia wrote:
         | Yes, because the prompt is simple as well.
         | 
         | Chain of thought or some similar strategies (I hate that they
         | have their own name and like a paper and authors, lol) can help
         | you push that 0.9 to a 0.95-0.99.
        
       | yawnxyz wrote:
       | I found that reducing html down to markdown using turndown or
       | https://github.com/romansky/dom-to-semantic-markdown works well;
       | 
       | if you want the AI to be able to select stuff, give it cheerio or
       | jQuery access to navigate through the html document;
       | 
       | if you need to give tags, classes, and ids to the llm, I use an
       | html-to-pug converter like https://www.npmjs.com/package/html2pug
       | which strips a lot of text and cuts costs. I don't think LLMs are
       | particularly trained on pug content though so take this with a
       | grain of salt
        
         | rcarmo wrote:
         | Hmmm. That's interesting. I wish there was a Node-RED node for
         | the first library (I can always import the library directly and
         | build my own subflow, but since I have cheerio for Node-RED and
         | use it for paring down input to LLMs already...)
        
       | ravedave5 wrote:
       | ChatGPT is clearly trained on wikipedia, is there any concern
       | about its knowledge from there polluting the responses? Seems
       | like it would be better to try against data it didn't potentially
       | already know.
        
       | IncreasePosts wrote:
       | Isn't GPT-4o multimodal? Shouldn't I be able to just feed in an
       | image of the rendered HTML, instead of doing work to strip tags
       | out?
        
         | spencerchubb wrote:
         | it is theoretically possible, but the results and bandwidth
         | would be worse. sending an image that large would take a lot
         | longer than sending text
        
       | CharlieDigital wrote:
       | I roughly came to the same conclusion a few months back and wrote
       | a simple, containerized, open source general purpose scraper for
       | use with GPT using Playwright in C# and TypeScript that's fairly
       | easy to deploy and use with GPT function calling[0]. My
       | observation was that using `document.body.innerText` was
       | sufficient for GPT to "understand" the page and
       | `document.body.innerText` preserves some whitespace in Firefox
       | (and I think Chrome).
       | 
       | I use more or less this code as a starting point for a variety of
       | use cases and it seems to work just fine for my use cases
       | (scraping and processing travel blogs which tend to have pretty
       | consistent layouts/structures).
       | 
       | Some variations can make this better by adding logic to look for
       | the `main` content and ignore `nav` and `footer` (or variants
       | thereof whether using semantic tags or CSS selectors) and taking
       | only the `innerText` from the main container.
       | 
       | [0] https://github.com/CharlieDigital/playwright-scrape-api
        
       | beepbooptheory wrote:
       | You step back and realize: we are thinking about how to best
       | remove _some_ symbols from documents that not a moment ago we
       | were deciding certainly needed to be in there, all to feed a
       | certain kind of symbol machine which has seen all the symbols
       | before anyway, all so we don 't pay as much cents for the symbols
       | we know or think we need.
       | 
       | If I was not a human but some other kind of being suspended above
       | this situation, with no skin in the game so to speak, it would
       | all seem so terribly inefficient... But as fleshy mortal I do
       | understand how we got here.
        
       | cpursley wrote:
       | What I do is convert to markdown, that way you still get some
       | semantic structure. Even built an Elixir library for this:
       | https://github.com/agoodway/html2markdown
        
         | bearjaws wrote:
         | Seems to be the most common method I've seen, it makes sense
         | given how well LLMs understand markdown.
        
       ___________________________________________________________________
       (page generated 2024-09-06 23:00 UTC)