[HN Gopher] Minifying HTML for GPT-4o: Remove all the HTML tags
___________________________________________________________________
Minifying HTML for GPT-4o: Remove all the HTML tags
Author : edublancas
Score : 50 points
Date : 2024-09-05 13:51 UTC (1 days ago)
(HTM) web link (blancas.io)
(TXT) w3m dump (blancas.io)
| giancarlostoro wrote:
| I wonder if this is due to some template engines looking
| minimalist like that. I think maybe Pug?
|
| https://github.com/pugjs/pug?tab=readme-ov-file#syntax
|
| It is whitespace sensitive though, but essentially looks like
| that. I doubt this is the only unique template engine like this
| though.
| cj wrote:
| Related article from 4 days ago (with comments on scraping,
| specifically discussing removing HTML tags)
|
| https://news.ycombinator.com/item?id=41428274
|
| Edit: looks like it's actually the same author
| throwup238 wrote:
| I don't think that Mercury Prize table is a representative
| example because each column has an obviously unique structure
| that the LLM can key in on: (year) (Single Artist/Album pair)
| (List of Artist/Album pairs) (image) (citation link)
|
| I think a much better test would be something like "List of
| elements by atomic properties" [1] that has a lot of adjacent
| numbers in a similar range and overlapping first/last column
| types. However, the danger with that table might be easy for the
| LLM to infer just from the element names since they're well known
| physical constants. The table of counties by population density
| might be less predictable [2] or list of largest cities [3]
|
| The test should be repeated with every available sorting function
| too, to see if that causes any new errors.
|
| [1]
| https://en.wikipedia.org/wiki/List_of_elements_by_atomic_pro...
|
| [2]
| https://en.wikipedia.org/wiki/List_of_countries_and_dependen...
|
| [3] https://en.wikipedia.org/wiki/List_of_largest_cities#List
| edublancas wrote:
| thanks a lot for the feedback! you're right, this is much
| better input data. I'll re-run the code with these tables!
| cal85 wrote:
| Good points. But I feel like even with the cities article it
| could still 'cheat' by recognising what the data is supposed to
| be and filling in the blanks. Does it even need to be real
| though? What about generating a fake article to use as a test
| so it can't possibly recognise the contents? You could even get
| GPT to generate it, just give it the 'Largest cities' HTML and
| tell it to output identical HTML but with all the names and
| statistics changed randomly.
| curl-up wrote:
| Additionally, using any Wiki page is misleading, as LLMs have
| seen their format many times during training, and can probably
| reproduce the original HTML from the stripped version fairly
| well.
|
| Instead, using some random, messy, scattered-with-spam site
| would be a much more realistic test environment.
| topaz0 wrote:
| Is .8 or .9 considered good enough accuracy for something as
| simple as this?
| edublancas wrote:
| I'd say how much is good enough highly depends on your use
| case. For something that still has to be reviewed by a human, I
| think even .7 is great; if you're planning to automate
| processes end-to-end, I'd aim for higher than .95
| LunaSea wrote:
| Well, when "simply" extracting the core text of an article is a
| task where most solutions (rule-based, visual, traditional
| classifiers and LLMs) rarely score above 0.8 in precision on
| datasets with a variety of websites and / or multilingual
| pages, I would consider that not too bad.
| moralestapia wrote:
| Yes, because the prompt is simple as well.
|
| Chain of thought or some similar strategies (I hate that they
| have their own name and like a paper and authors, lol) can help
| you push that 0.9 to a 0.95-0.99.
| yawnxyz wrote:
| I found that reducing html down to markdown using turndown or
| https://github.com/romansky/dom-to-semantic-markdown works well;
|
| if you want the AI to be able to select stuff, give it cheerio or
| jQuery access to navigate through the html document;
|
| if you need to give tags, classes, and ids to the llm, I use an
| html-to-pug converter like https://www.npmjs.com/package/html2pug
| which strips a lot of text and cuts costs. I don't think LLMs are
| particularly trained on pug content though so take this with a
| grain of salt
| rcarmo wrote:
| Hmmm. That's interesting. I wish there was a Node-RED node for
| the first library (I can always import the library directly and
| build my own subflow, but since I have cheerio for Node-RED and
| use it for paring down input to LLMs already...)
| ravedave5 wrote:
| ChatGPT is clearly trained on wikipedia, is there any concern
| about its knowledge from there polluting the responses? Seems
| like it would be better to try against data it didn't potentially
| already know.
| IncreasePosts wrote:
| Isn't GPT-4o multimodal? Shouldn't I be able to just feed in an
| image of the rendered HTML, instead of doing work to strip tags
| out?
| spencerchubb wrote:
| it is theoretically possible, but the results and bandwidth
| would be worse. sending an image that large would take a lot
| longer than sending text
| CharlieDigital wrote:
| I roughly came to the same conclusion a few months back and wrote
| a simple, containerized, open source general purpose scraper for
| use with GPT using Playwright in C# and TypeScript that's fairly
| easy to deploy and use with GPT function calling[0]. My
| observation was that using `document.body.innerText` was
| sufficient for GPT to "understand" the page and
| `document.body.innerText` preserves some whitespace in Firefox
| (and I think Chrome).
|
| I use more or less this code as a starting point for a variety of
| use cases and it seems to work just fine for my use cases
| (scraping and processing travel blogs which tend to have pretty
| consistent layouts/structures).
|
| Some variations can make this better by adding logic to look for
| the `main` content and ignore `nav` and `footer` (or variants
| thereof whether using semantic tags or CSS selectors) and taking
| only the `innerText` from the main container.
|
| [0] https://github.com/CharlieDigital/playwright-scrape-api
| beepbooptheory wrote:
| You step back and realize: we are thinking about how to best
| remove _some_ symbols from documents that not a moment ago we
| were deciding certainly needed to be in there, all to feed a
| certain kind of symbol machine which has seen all the symbols
| before anyway, all so we don 't pay as much cents for the symbols
| we know or think we need.
|
| If I was not a human but some other kind of being suspended above
| this situation, with no skin in the game so to speak, it would
| all seem so terribly inefficient... But as fleshy mortal I do
| understand how we got here.
| cpursley wrote:
| What I do is convert to markdown, that way you still get some
| semantic structure. Even built an Elixir library for this:
| https://github.com/agoodway/html2markdown
| bearjaws wrote:
| Seems to be the most common method I've seen, it makes sense
| given how well LLMs understand markdown.
___________________________________________________________________
(page generated 2024-09-06 23:00 UTC)