[HN Gopher] Show HN: OCR Benchmark Focusing on Automation
___________________________________________________________________
Show HN: OCR Benchmark Focusing on Automation
OCR/Document extraction field has seen lot of action recently with
releases like Mixtral OCR, Andrew Ng's agentic document processing
etc. Also there are several benchmarks for OCR, however all testing
for something slightly different which make good comparison of
models very hard. To give an example, some models like mixtral-ocr
only try to convert a document to markdown format. You have to use
another LLM on top of it to get the final result. Some VLM's
directly give structured information like key fields from documents
like invoices, but you have to either add business rules on top of
it or use some LLM as a judge kind of system to get sense of which
output needs to be manually reviewed or can be taken as correct
output. No benchmark attempts to measure the actual rate of
automation you can achieve. We have tried to solve this problem
with a benchmark that is only applicable for documents/usecases
where you are looking for automation and its trying to measure that
end to end automation level of different models or systems. We
have collected a dataset that represents documents like invoices
etc which are applicable in processes where automation is needed vs
are more copilot in nature where you would need to chat with
document. Also have annotated these documents and published the
dataset and repo so it can be extended. Here is writeup:
https://nanonets.com/automation-benchmark Dataset:
https://huggingface.co/datasets/nanonets/nn-auto-bench-ds Github:
https://github.com/NanoNets/nn-auto-bench Looking for suggestions
on how this benchmark can be improved further.
Author : prats226
Score : 27 points
Date : 2025-03-12 20:49 UTC (2 days ago)
(HTM) web link (nanonets.com)
(TXT) w3m dump (nanonets.com)
| constantinum wrote:
| There was a discussion on this benchmark https://getomni.ai/ocr-
| benchmark couple of weeks ago here >
| https://news.ycombinator.com/item?id=43118514
| prats226 wrote:
| Yeah this new benchmark is kind of inspired by these existing
| benchmarks, things that are missing wrt automation
| themanmaran wrote:
| Love to see another benchmark! We published the OmniAI OCR
| benchmark the other week. Thanks for adding us to the list.
|
| One question on the "Automation" score in the results, is this a
| function of extraction accuracy vs the accuracy of the LLM's
| "confidence score". I noticed the "accuracy" column was very
| tightly grouped (between 79 & 84%) but the automation score was
| way more variable.
|
| And side note: is there an open source Mistral benchmark for
| their latest OCR model? I know they claimed it was 95% accurate,
| but it looks that was based on an internal evaluation.
| prats226 wrote:
| Automation is combination of both, accuracy and accuracy of
| confidence scores.
|
| Good way to think about automation is recall at high precision
| which is what you need for true automation where you don't
| worry about documents that are very likely to have correct
| results and focus on manually correcting documents likely to
| have errors.
|
| The reason accuracies are tighly grouped but not the automation
| is because these models are trained to be accurate but not
| necessarily predictable, where there is no real way to get
| confidence score calibrated.
|
| Couldn't find the benchmark mistral used as well
| criddell wrote:
| How do these compare to traditional commercial and open source
| OCR tools? What about things like the Apple Vision APIs?
| sumedh wrote:
| > Apple Vision APIs
|
| Can we even OCR pdfs using Apple Vision APIs
| prats226 wrote:
| Don't think you can. And also there is big difference in
| plain old OCR, which is just getting all text out from image
| and document processing which is can you only get the
| relevant information in a good structure that can be directly
| pushed into a database.
| kapitalx wrote:
| Great list! I'll definitely run your benchmark against Doctly.ai
| (our PDF-to-Markdown service) specially as we publish our
| workflow service, to see how we stack up.
|
| One thing I've noticed in many benchmarks, though, is the
| potential for bias. I'm actually working on a post about this
| issue, so it's top of mind for me. For example, in the omni
| benchmark, the ground truth expected a specific order for heading
| information--like logo, phone number, and customer details. While
| this data was all located near the top of the document, the exact
| ordering felt subjective. Should the model prioritize horizontal
| or vertical scanning? Since the ground truth was created by the
| company running the benchmark, their model naturally scored the
| highest for maintaining the same order as the ground-truth.
|
| However, this approach penalized other LLMs for not adhering to
| the "correct" order, even though the order itself was arguably
| arbitrary. This kind of bias can skew results and make it harder
| to evaluate models fairly. I'd love to see benchmarks that
| account for subjectivity or allow for multiple valid
| interpretations of document structure.
|
| Did you run into this when looking at the benchmarks?
|
| On a side note, Doctly.ai leverages multiple LLMs to evaluate
| documents, and runs a tournament with a judge for each page to
| get the best data (this is only on the Precision Ultra
| selection).
| prats226 wrote:
| Bias wrt ordering is a great point. What we consider structured
| information in this benchmark is irrespective of how its
| presentation (Order, format etc), it should be directly
| comparable. So the benchmark does that it into account.
|
| Example is if you are only converting lets say an invoice into
| markdown, you can introduce bias wrt ordering etc. But if the
| task is to find out invoice number, total amount, number of
| line items with headers like price, amount, description, in
| that case you can compare two outputs without a lot of bias. Eg
| even if columns are interchanged, you will still get the same
| metric.
| kapitalx wrote:
| Exactly. You still have to be explicit in order to remove
| bias. Either by sorting the keys, or looking up specific
| keys. For arrays, I would say order still matters. For
| example when you capture a list of invoice items, you should
| maintain order.
| themanmaran wrote:
| Hey I wrote the Omni benchmark. I think you might be misreading
| the methodology on our side. Order on page does not matter in
| our accuracy scoring. In fact we are only scoring on JSON
| extraction as a measurement of accuracy. Which is order
| independent.
|
| We chose this method for all the same reasons you highlight.
| Text similarity based measurements are very subject to bias,
| and don't correlate super well with accuracy. I covered the
| same concepts in the "The case against text-similarity"[1]
| section of our writeup.
|
| [1] https://getomni.ai/ocr-benchmark
| bn-l wrote:
| In your pricing example I see $6.27 for a 10 page document. That
| is extremely expensive.
| prats226 wrote:
| Can share link? Maybe some kind of mistake.
___________________________________________________________________
(page generated 2025-03-14 23:00 UTC)