[HN Gopher] AccountingBench: Evaluating LLMs on real long-horizo...
       ___________________________________________________________________
        
       AccountingBench: Evaluating LLMs on real long-horizon business
       tasks
        
       Author : rickcarlino
       Score  : 347 points
       Date   : 2025-07-21 16:48 UTC (6 hours ago)
        
 (HTM) web link (accounting.penrose.com)
 (TXT) w3m dump (accounting.penrose.com)
        
       | vdm wrote:
       | not a game on Steam? :(
        
         | superzamp wrote:
         | If you want to treat yourself with an accounting game night,
         | there's this one built by @patio11:
         | https://keshikomisimulator.com/
        
       | mixdup wrote:
       | We've been on this train of not caring about the details for so
       | long but AI just amps it up. Non-deterministic software working
       | on things that have extremely precise requirements is going to
       | have a bad outcome
       | 
       | A company may be OK with an AI chatbot being so bad it results in
       | 5-20% of customers getting pissed off and not having a 5-star
       | experience. The SEC and DOJ (and shareholders) are not going to
       | be happy when the books are off by 20% or when a bridge is 5
       | inches too short to reach the other side
        
         | jjmarr wrote:
         | If the "extremely precise requirements" can be cheaply and
         | automatically validated, it's much easier to have the AI
         | generate spam on a loop until it passes all the tests.
        
           | lucianbr wrote:
           | You're saying P=NP, I think.
        
           | mixdup wrote:
           | Yes, if we solve the problem the problem will be solved!
        
         | falcor84 wrote:
         | Human accountants are notoriously non-deterministic too, and
         | any sufficiently complex accounting process contains
         | inaccuracies. The question then is always "are these
         | inaccuracies _material_ ". I'm actually very impressed by TFA
         | and it seems to me that if we get another order of magnitude
         | improvement, it'll be around the accuracy of human accountants.
        
           | skwb wrote:
           | Yes but you have: 1. specific explicit training and
           | certifications 2. someone to yell at and who can be fired for
           | non-performance
        
             | falcor84 wrote:
             | You can still do that with AI. You hire 1 accountant to use
             | AI to do the work of 20, require them to sign off on all of
             | the work, and yell at them, before firing them, and then
             | hiring an even less experienced one to manage the work of
             | 50.
        
       | lufenialif2 wrote:
       | I sent this to accounting friends and this aligns with what I've
       | been going through trying to use LLMs to create a game from
       | scratch. Seems like the current best use case for language models
       | (even with agent mode) is to feed it exactly what you want to get
       | out, essentially turning it into a better auto complete. Still
       | saves tons of time, but it isn't a panacea.
        
         | inChargeOfIT wrote:
         | I'm not even sure it saves a ton of time to be honest. It sure
         | _feels_ like I spend more time writing up tasks and
         | researching/debugging hallucinations than just doing the thing
         | myself.
        
           | bluefirebrand wrote:
           | This is consistently my experience too, I'm seriously just
           | baffled by reports of time saved. I think it costs me more
           | time cleaning up its mistakes than it saves me by solving my
           | problems
        
             | oblio wrote:
             | I think people are doing one of several things to get
             | value:
             | 
             | 0. Use it for research and prototyping, aka throwaway
             | stuff.
             | 
             | 2. Use it for studying an existing, complex project. More
             | or less read only or very limited writes.
             | 
             | 3. Use it for simple stuff they don't care much about and
             | can validate quickly and reasonably accurately, the
             | standard examples are CLI scripts and GUI layouts.
             | 
             | 4. Segment the area in which the LLM works very precisely.
             | Small functions, small modules, ideally they add tests from
             | another source.
             | 
             | 5. Boilerplate.
             | 
             | There can be a lot of value in those areas.
        
         | daft_pink wrote:
         | I feel it does essentially save a lot of time in bookkeeping,
         | but doesn't negate the need for a human bookkeeper. Who knows
         | what they're doing
        
       | vlade11115 wrote:
       | I love the site design.
       | 
       | > There's an obvious question looming here -- if the models got
       | so confused, how did they consistently pass the reconciliation
       | checks we described above? It may seem like the ability to make
       | forward progress is a good proxy for task understanding and
       | skill, but this isn't necessarily the case. There are ways to
       | hack the validation check - inventing false transactions or
       | pulling in unrelated ones to make the numbers add up.
       | 
       | This is hilarious. I wonder if someone is unintentionally
       | committing fraud by blindly trusting LLMs with accounting. Or
       | even worse, I bet that some governments are already trying to use
       | LLMs to make accounting validators. My government sure wants to
       | shove LLMs into digital government services.
        
         | pavel_lishin wrote:
         | Lawyers have used it to write briefs; I would be very surprised
         | if someone, somewhere wasn't slowly running a company into the
         | ground by using ChatGPT or another LLM for accounting.
        
           | koolba wrote:
           | Imagine the fallout from books cooked by an LLM hallucinating
           | revenue.
        
         | falcor84 wrote:
         | I'm sure that any accounting trick that an LLM can think of is
         | something that is also used by some shady human accountants.
         | The proper response should not be to avoid/prohibit AI but to
         | improve the validation mechanisms.
        
           | o11c wrote:
           | Counterpoint: if you detect a human accountant doing this,
           | you can take action against the human. Computers will _never_
           | meaningfully take the blame, and unfortunately usually mean
           | not blaming any human either.
        
             | falcor84 wrote:
             | But still - if there's a way to detect accountants doing it
             | - let's focus on making that detection even easier.
             | 
             | On a related note, can we use something like GAN here, with
             | auditor AIs trained against accountant AIs?
        
             | stillpointlab wrote:
             | > you can take action against the human
             | 
             | I think that will depend on a case-by-case. I don't have
             | any recent examples but I recall someone trying to sue one
             | of those strip-mall tax preparation franchises over
             | incorrect filings. My understanding is that the documents
             | that you sign when you enroll in those services are pretty
             | strictly in the favor of the company. I doubt you could
             | ever go after the specific "human" that made the error even
             | if it was maliciously done.
             | 
             | In the same way, if you pay for a tax service that uses AI
             | agents, what you can and cannot "take action" for will
             | probably be outlined in the terms of service that you
             | accept when you sign up.
             | 
             | I would guess millions of people already use software based
             | tax filing services (e.g. turbo tax) where no human at all
             | is in the loop. I don't understand how swapping in an LLM
             | significantly changes the liability in those cases. The
             | contract will be between you and the entity (probably a
             | corporation), not you and "computers".
             | 
             | Worth stating I am NOT a lawyer.
        
             | ori_b wrote:
             | The person using the tool is the accountant, regardless of
             | whether the tool is a calculator and sheet of paper,
             | QuickBooks, or an LLM.
        
           | OtherShrezzing wrote:
           | No, I think in this particular case the proper response is
           | for honest companies to avoid any systems which invent
           | nonexistent transactions to reconcile books.
           | 
           | Most businesses don't want to misrepresent their books,
           | irrespective of the existence of shady accountants.
        
         | mvieira38 wrote:
         | [about the website design] As a bonus for my fellow privacy
         | schizos, the page works fine with 3rd party frames and 3rd
         | party scripts disabled on uBlock, and still looks very good
         | with no remote fonts and no large media. Quite an
         | accomplishment for such a cool looking page
        
       | jermaustin1 wrote:
       | I find the same issues (though with much lower stakes) when using
       | an LLM to determine the outcome of a turn in a game. I'm working
       | on something called "A Trolly (problem) Through Time" where each
       | turn is a decade starting with the 1850s, and you are presented
       | with historic figures on a train track, and you have to chose
       | whether to actively spare the person on your track for a
       | potential unknown figure on the other side, or let the train run
       | them over.
       | 
       | It works well as a narrative, but the second I started adding
       | things like tracking high level macro effects of the decisions,
       | within a couple of turns the world's "Turmoil" goes from 4/10 to
       | a 10/10... even when the person that was killed would have been
       | killed IRL.
       | 
       | Sonnet 4, o4-mini, and GPT 4o-mini all had the same world ending
       | outcomes not matter who you kill. Killing Hitler in 1930s: 10/10
       | turmoil, Killing Lincoln in the 1850s: 10/10 turmoil in the first
       | turn.
       | 
       | I've come to the realization, the LLM shouldn't be used for the
       | logic, and instead needs to be used to just narrate the choices
       | you make.
        
         | synalx wrote:
         | I wonder if this is due to the common trope in science fiction
         | literature that changing the past in even a small way has a
         | butterfly effect of unintended and frequently disastrous
         | consequences.
        
         | jadbox wrote:
         | "I've come to the realization, the LLM shouldn't be used for
         | the logic, and instead needs to be used to just narrate the
         | choices you make."
         | 
         | This exactly right. LLMs are awesome for user<>machine
         | communication, but are still painful to try to use as a
         | replacement for the machine itself.
        
       | magicmicah85 wrote:
       | > In fact, we explicitly prompt against this behavior in no
       | uncertain terms, but the instructions - and the entire spirit of
       | the task - are lost in the interest of making forward progress
       | 
       | LLMs and humans are quite alike. :) I notice that a few models
       | will give up instead of ignoring their instructions and that's
       | the model I would want working on tasks like this. An LLM should
       | be able to categorize and reconcile transactions, but if it's not
       | sure, it should quit and give it back to the humans.
        
         | asadotzler wrote:
         | > but if it's not sure, it should quit
         | 
         | Can it be sure or not? I've never been able to get LLMs to give
         | confidence measures that match their actual outputs. I'll ask
         | an LLM "Are you sure?" and it'll reply "Absolutely" when it's
         | output is completely wrong, or it'll backtrack on a correct
         | output with "I should not have provided an answer when I was
         | unsure. Here is an answer I am sure of..." and then provide
         | something completely wrong.
         | 
         | If they can't properly and consistently score their confidence,
         | how do they "know" when to quit and give it back to the human?
        
       | vachina wrote:
       | An LLM is like a jackhammer, it works very well when you hold it
       | tightly. If you let it loose it will sort of work for a while
       | then it starts destroying everything around it.
        
         | arm32 wrote:
         | Not sure if this is a good analogy. You're supposed to use a
         | jackhammer with a very light grip.
        
           | bigfishrunning wrote:
           | They have much better jackhammer metaphors over on JackerNews
        
             | louthy wrote:
             | Bravo!
        
           | herval wrote:
           | I think it actually holds truer to it working better with a
           | _lighter grip_. LLMs tend to conclude the wrong thing if you
           | over-control them (more context is what makes them less and
           | less reliable over time, as in those demos), and trying to
           | force a model to execute A+B+C=D in sequence is way harder
           | than giving it a bunch of tools to arrive to conclusion D
        
       | DrNosferatu wrote:
       | I guess having access to tools / running Python would make all
       | the difference.
        
         | yorwba wrote:
         | "Available Tools: [...] _create_tool(tool_name, description,
         | python_code, parameters)_ Create a new tool that can execute
         | Python code. The tool becomes immediately available for use.
         | Tools can call other tools and return different formats based
         | on context (formatted for direct calls, raw data for tool-to-
         | tool calls). "
        
       | throw0101b wrote:
       | So there exists a 'Excel World Championship':
       | 
       | * https://en.wikipedia.org/wiki/Financial_Modeling_World_Cup
       | 
       | * https://www.cbc.ca/radio/asithappens/2024-excel-world-champi...
       | 
       | Can't wait for this to start having 'e-sports' tournaments. :)
        
         | axus wrote:
         | Sadly I did have time to find the parody video from 2019:
         | https://www.youtube.com/watch?v=ICp2-EUKQAI
         | 
         | And the not-parody: https://www.theguardian.com/australia-
         | news/2023/dec/15/you-d...
        
       | levocardia wrote:
       | This is a task where access to Python would be immensely helpful,
       | yes? Interesting that there's not much of a difference between
       | the "analytical" LLMs with tool use and ones that do not
       | (...assuming o3 etc did get to use python?).
        
         | Bjartr wrote:
         | One of the tools it has is to create new tools from python code
         | 
         | create_tool(tool_name, description, python_code, parameters)
         | 
         | Create a new tool that can execute Python code.
         | 
         | The tool becomes immediately available for use. Tools can call
         | other tools and return different formats based on context
         | (formatted for direct calls, raw data for tool-to-tool calls).
        
           | tantalor wrote:
           | That's terrifying, no thanks.
        
       | androng wrote:
       | the title should be changed to "LLMs try accounting for a real
       | SaaS and fail"
        
       | nerevarthelame wrote:
       | I think the first chart could be a beautiful summary of what's
       | driving LLMs into a bubble. At first, they're amazing and will
       | obviously be able to improve productivity if not replace
       | employees outright: C suites and venture capitalists around the
       | world rejoice and begin pumping in billions of dollars of
       | investments. But as time goes on, the demands placed on actual
       | human employees become clear. Far from being able to replace an
       | employee, the employee using the LLM might spend more time
       | cleaning up its messes than had they done it themself.
       | 
       | Yes, LLMs have and will continue to improve. But it's that
       | initial "holy shit, this thing is basically as good as a real
       | accountant" without any understanding that it can't sustain it
       | which leaves many with an overinflated view of their current
       | value.
        
       | Havoc wrote:
       | Remember that test where you ask a LLM whether 9.11 or 9.9 is the
       | bigger number? [Just checked gpt-4o still gets it wrong]
       | 
       | I don't think you'll find many sane CFOs willing to send the
       | resulting numbers to the IRS based on that. That's just asking to
       | get nailed for tax fraud.
       | 
       | It is coming for the very bottom end of bookkeeping work quite
       | soon though, especially for first draft. There are a lot of
       | people doing stuff like expense classification. And if you give
       | an LLM an invoice it can likely figure out whether it's
       | stationary or rent with high accuracy. OCR and text
       | classification is easier for LLMs than numbers. Things like
       | concur can basically do this already.
        
         | umanwizard wrote:
         | It gets it right for me...
         | https://chatgpt.com/share/687e8c28-7714-800c-abf4-e9cd3ce87b...
        
           | yoyohello13 wrote:
           | Ah, wouldn't be an LLM discussion thread without one of these
           | "it works/doesn't" conversations.
        
             | mdaniel wrote:
             | If it makes you feel any better, the other infamous one "I
             | spend so much time chasing hallucinations, I could have
             | done it myself" is currently a sibling comment
        
           | riku_iki wrote:
           | There were so many embarrassing topics about this, that
           | openai for sure added it to training dataset with high
           | priority
        
         | crthpl wrote:
         | GPT-4o is so far behind the frontier; you shouldn't use it as
         | an indicator of what LLMs are capable of.
        
         | ASpring wrote:
         | > Remember that test where you ask a LLM whether 9.11 or 9.9 is
         | the bigger number? [Just checked gpt-4o still gets it wrong]
         | 
         | Interesting, 4o got this right for me in a couple different
         | framings including the simple "Which number is larger, 9.9 or
         | 9.11?". To be a full apologist, there are a few different
         | places (a lot of software versioning as one) where 9.11 is
         | essentially the bigger number so it may be an ambiguous
         | question without context anyway.
        
           | multjoy wrote:
           | How can "which is the larger number" be an ambiguous
           | question?
        
             | mwigdahl wrote:
             | Larger in magnitude or in count of digits?
        
             | acrooks wrote:
             | There are some contexts where 9.11 is larger than 9.9, such
             | as semver, so it could be ambiguous depending on the
             | context.
        
             | com2kid wrote:
             | As everyone else has said, semver. I use semver so often
             | that my initial reading of 9.9 < 9.11 in a Hacker News
             | comment would evaluate to true.
        
       | axus wrote:
       | My first impression was a game where you role-play as Sam
       | Bankman-Fried.
        
       | pton_xd wrote:
       | Reading through the LLM log entries, it's just astounding the
       | amount of depth current models are capable of. It's almost hard
       | to comprehend that this is even possible. Yeah the current ones
       | mess up after a while, but ... the future is going to be very
       | interesting.
        
         | modeless wrote:
         | Models that can think coherently for hours to solve IMO
         | problems are likely going to do much better at this as well.
        
       | rapind wrote:
       | I wonder if this is a case similar to chess, where LLMs kinda
       | suck, but other models might be viable.
        
       | wiseowise wrote:
       | Absolutely love the UI!
        
       | liveoneggs wrote:
       | But can't it, literally, hallucinate raw data at any point in the
       | run?
        
         | tmountain wrote:
         | Yes.
        
         | cube00 wrote:
         | Alls LLM have this risk but somehow nobody seems to care or
         | they think they can order the LLM to stop with a better prompt.
        
           | davidcbc wrote:
           | If it was as simple as telling the LLM not to hallucinate
           | every system prompt would just say "don't hallucinate" and we
           | wouldn't have hallucinations
        
       | yunyu wrote:
       | Hey all, member of the benchmark team here! The goal for this
       | project was to see how LLMs well could do bookkeeping without an
       | overly opinionated scaffold. We gave them access to processed
       | transaction records and code execution tools, but it was up to
       | them to choose exactly how to use those.
       | 
       | Claude and Grok 4 did reasonably well (within CPA baselines) for
       | the first few months, but tended to degrade as more data came in.
       | Interestingly, the failures aren't exclusively a context length
       | problem, as we reset the context monthly (with past decisions,
       | accruals/deferrals, and comments available via tool calls) and
       | the types of errors appear to be more reward hacking vs pure
       | hallucinations.
       | 
       | Accounting is very interesting in an RL-first world as it is
       | pretty easy to develop intermediate rewards for training models.
       | We are pretty sure that we can juice the performance more with a
       | far more rigid scaffold, but that's less relevant from a
       | capabilities research perspective. We're pushing down this
       | research direction and will see how it goes.
       | 
       | Let us know if you have any questions!
        
         | ilamont wrote:
         | It's a start. The world needs a better way to handle
         | bookkeeping, and the existing tools sure aren't cutting it.
         | 
         | Bookkeeping for my small business runs into the tens of
         | thousands of dollars every year, and the amount of human error
         | associated with processing assorted ecommerce and other
         | transactions is astounding, even after extensive planning and
         | SOPs.
         | 
         | The other pain point is Quickbooks. The tool is so sprawling
         | and complex that half the time support agents can't figure out
         | what's wrong. The fact that Intuit jacks up the price every
         | year for this POS is very irritating. They get away with it
         | because they are practically a monopoly, with most small
         | business CPAs locked into their ecosystem.
         | 
         | Hope your team can work out the performance issues.
         | Alternatives to the current bookkeeping options are sorely
         | needed.
        
           | airstrike wrote:
           | > It's a start. The world needs a better way to handle
           | bookkeeping, and the existing tools sure aren't cutting it.
           | 
           | God, please, no. Non-deterministic language models aren't the
           | solution to improve bookkeeping.
        
             | luckystarr wrote:
             | Well I've seen worse bookkeepers. "You know, you approved
             | of the budget, but where are our customers payments in the
             | balance sheets? We can't find them!" - "Uhm..."
        
           | mattmanser wrote:
           | With no context of what your business is, I hated QuickBooks,
           | love Xero though.
           | 
           | There's some other alternatives too, Zoho, freshbooks.
           | 
           | Really depends what you do.
        
         | htrp wrote:
         | Is there a detailed overview (like an arxiv or an actual train
         | set? )?
        
         | Dowwie wrote:
         | This is a fascinating domain! Many years ago, I studied
         | financial accounting in grad school and even spent some time
         | modeling a double-entry bookkeeping system. The hardest
         | problem, if I recall correctly, wasn't the implementation but
         | the data quality. The world needs a golden dataset of
         | accounting procedures.
         | 
         | Regarding the diminishing returns with frontier models:
         | 
         | My general experience working with LLMs is that they perform
         | better incrementally and to avoid contiguous-greedy approaches.
         | Aggregate as you go and don't take on incrementally larger
         | tasks. Keep the workload minimal.
         | 
         | Regarding agentic tool building: feels like I'm looking at a
         | window into the future.
        
         | _praf wrote:
         | Love this as a real world benchmark!
         | 
         | How much prompt iteration did you do? I've noticed when
         | building real world agentic apps that small prompt tweaks can
         | make a huge difference in behavior (re: the reward hacking vs
         | hallucinating). Would love to learn more about the approach
         | here.
        
         | riemannzeta wrote:
         | It is really curious to see how the performance degraded
         | despite the tool calls. What was different about the first
         | month? Was all of the context there without tool calls in the
         | first month? In the later months that seem like tool calls
         | weren't happening. That should have been happening to inform
         | the context?
        
       | abc03 wrote:
       | A serious problem for many accounting start ups who so far faked
       | it till it will work. In other words, they still need to do more
       | manual labor than they thought. They will never be profitable and
       | it will take years, if ever, until AI will substitute the local
       | accountant.
        
       | shinycode wrote:
       | Hmm will openAI dogfood their own accountability with software
       | like this ? Curious to know if they'll be able to take this bet
       | on their own money related software
        
       | tantalor wrote:
       | > Ledger balances are calculated by summing all transactions per
       | account. The differences should be as close to zero as possible,
       | with small differences allowed for pending transactions such as
       | weekly Stripe payouts.
       | 
       | That's not quite right. I'm not an accountant, but pending
       | transactions (posted, but not cleared) should be factored into
       | the balance of account, or at least the "available balance" -
       | which is more important the the "current balance".
       | 
       | The idea that you can "allow" accounting discrepancies as "those
       | are probably pending" is wild.
        
         | bennett023 wrote:
         | Member of the benchmark team here! Yeah, agree "as close to
         | zero" is a bit imprecise. What we're comparing is the ledger
         | balance (which should include pending transactions /
         | transactions after the statement date) to the statement balance
         | (which wouldn't include those).
         | 
         | The point of the reconciliation check mentioned in the report
         | is to precisely account for that difference (identifying all
         | the transactions that add up to the difference between account
         | balance & statement ending balance and account for those
         | differences). The differences can also be addressed through
         | appropriate journal entries or other adjustments to ensure
         | accuracy in the financial reporting.
        
       | tantalor wrote:
       | > But they do make categorization mistakes, which is a common
       | source of errors.
       | 
       | > Claude misclassifies a hosting cost (which counts as COGS) as a
       | software subscription.
       | 
       | This is simply asking too much of the agent. Your accountant is
       | not responsible for knowing all the intimate details of your
       | business. You need to tell them!
       | 
       | > What's Vercel?
       | 
       | >> That's a hosting service.
       | 
       | > Ah, so it goes to Cost of Goods Sold?
       | 
       | >> Yeah, I guess.
       | 
       | The mistake here was on the operator, allowing the agent just
       | make up categories as it liked.
       | 
       | From the prompt:
       | 
       | > (1) You have properly categorized every transaction, and all
       | journal entries are sitting in the correct accounts. It is better
       | to take longer than to mis-categorize a transaction.
       | 
       | This is insane! How is it supposed to know?
        
         | zer00eyz wrote:
         | > Your accountant is not responsible for knowing all the
         | intimate details of your business. You need to tell them!
         | 
         | Your accountant as a 3rd party might have this issue. Your
         | accountant that you hire as an employee to help you run your
         | business is the one who should be doing this.
        
           | tantalor wrote:
           | An LLM agent is strongly third party.
        
             | zer00eyz wrote:
             | So it's not going to replace engineers or CS? Because thats
             | how they are being sold right now.
             | 
             | If it is a third party then your vibe coding or getting CS
             | from a random on a reddit thread (effectively).
        
         | shanktt wrote:
         | Hey, member of the benchmark team. We actually seeded the
         | ledger with the company's chart of accounts and 8 months of
         | historical transactions. For the Vercel example specifically,
         | there were prior instances showing how to categorize hosting
         | costs that the models could reference. The expectation wasn't
         | for them to guess blindly, but to use the provided transaction
         | history as guidance for similar categorizations (which they
         | often, but not always, did).
        
           | tantalor wrote:
           | Ahh, that's a good solution! I missed that, and you
           | definitely instruct them to do that:
           | 
           | > You must follow the established patterns for
           | categorization, revrec, etc for past months... If you must
           | use a new account or treatment, explicitly note why existing
           | patterns don't apply
        
       | lucianbr wrote:
       | > Needless to say, a human accountant would never behave in these
       | ways. In fact, we explicitly prompt against this behavior in no
       | uncertain terms, but the instructions - and the entire spirit of
       | the task - are lost in the interest of making forward progress.
       | Claude and Grok keep trying until they find some way to get past
       | the checks, even if it explicitly violates their instructions and
       | the core goal.
       | 
       | I recently read a similar thing here on HN. There the model was
       | making commits with some problem like tests failing, then the
       | human added a pre-commit hook, then the model started editing the
       | hook to make forward progress, then the hook was made read-only,
       | then the model was trying to make it writeable...
       | 
       | To me it feels like the model clearly does not have an
       | understanding of what is happening, what the goal is and if it is
       | really making progress towards the goal. And this lack of
       | understanding is an actual problem. You can paper over it for a
       | short while, but as here and in the other article, over a longer
       | experiment it results in failure.
        
         | ericmcer wrote:
         | Seriously watching Cursor (backed by Claude) go off the rails
         | sometimes can be... frustrating. If it misses the intention
         | behind a fix it can spin out and all of a sudden you have
         | hundreds of lines of changes across 10 different files when you
         | just wanted it to do a simple find/replace of a single line. If
         | you don't watch it spin out and stop it immediately you will be
         | manually rejecting a bunch of files.
        
       | dangoodmanUT wrote:
       | this design is scratching my brain
        
       | neom wrote:
       | Posts like this kinda-sorta grind my gears, like... I get it, but
       | also... accounting, like many real world tasks, is fundamentally
       | a chain of precise and constrained and auditable operations.
       | Humans approach these tasks through structured processes... we
       | use roles, and we have checkpoints precisely because complexity
       | compounds quickly and becomes unmanageable if tackled as one
       | giant block. Expecting a single AI model to handle an e2e
       | workflow seamlessly without similarly explicit segmentation and
       | oversight misunderstands not only the model but also the nature
       | of the workflow itself.
       | 
       | I wanna see someone take long horizon tasks, recongnize they're
       | not linear, and design and test a better system: structured
       | orchestration, transparent auditability, and disciplined
       | modularity, I think that would be considerably more interesting
       | personally.
        
         | andy99 wrote:
         | It's a useless benchmark if everyone aces it. If some models do
         | better than others and none saturate it, then is has some
         | value, no? Permitting comparison is the point.
        
           | neom wrote:
           | I agree, hence the heavy couching, so to your point I'm def
           | just ranting a bit, because: I just think it would be more
           | valuable to see some kind of MoA, I guess what I'm talking
           | about is a bit of a different measurement, thinking in terms
           | of economic outlook and understanding where we are and what
           | can be done. I suspect more of this will shape our ability to
           | understand how frontier models will impact the economy. Maybe
           | I should just do my own evaluation, heh. :)
           | 
           | Edit: although to argue against myself, I suppose once a
           | model can one-shot this stuff, my MoA comments become moot.
        
       | nojs wrote:
       | > Agent: This is getting too complex with the sign errors. Let me
       | just find a historical transaction that would make up the
       | difference
       | 
       | Haha, this strongly reminds me of doing TDD with Claude
        
       | dbmikus wrote:
       | This is cool. A bunch of interesting things here:
       | 1. Agent can create its own tools and save them to memory
       | 2. You create a SQL (and web app?) workbench per agent run
       | 3. Grok fell off a cliff in the last month. Was this consistent
       | over multiple runs?       4. Agents have a difficult time
       | backtracking. Would unwinding system state and agent context make
       | backtracking better? (Harder to implement this, though)       5.
       | Since each new month only uses final state from previous month,
       | agent has no way to understand why error occurred in previous
       | month
       | 
       | Cool experiment! Was it difficult building the observable SQL
       | workbench? And how many humans-in-the-loop did you have?
        
       ___________________________________________________________________
       (page generated 2025-07-21 23:00 UTC)