[HN Gopher] AccountingBench: Evaluating LLMs on real long-horizo...
___________________________________________________________________
AccountingBench: Evaluating LLMs on real long-horizon business
tasks
Author : rickcarlino
Score : 347 points
Date : 2025-07-21 16:48 UTC (6 hours ago)
(HTM) web link (accounting.penrose.com)
(TXT) w3m dump (accounting.penrose.com)
| vdm wrote:
| not a game on Steam? :(
| superzamp wrote:
| If you want to treat yourself with an accounting game night,
| there's this one built by @patio11:
| https://keshikomisimulator.com/
| mixdup wrote:
| We've been on this train of not caring about the details for so
| long but AI just amps it up. Non-deterministic software working
| on things that have extremely precise requirements is going to
| have a bad outcome
|
| A company may be OK with an AI chatbot being so bad it results in
| 5-20% of customers getting pissed off and not having a 5-star
| experience. The SEC and DOJ (and shareholders) are not going to
| be happy when the books are off by 20% or when a bridge is 5
| inches too short to reach the other side
| jjmarr wrote:
| If the "extremely precise requirements" can be cheaply and
| automatically validated, it's much easier to have the AI
| generate spam on a loop until it passes all the tests.
| lucianbr wrote:
| You're saying P=NP, I think.
| mixdup wrote:
| Yes, if we solve the problem the problem will be solved!
| falcor84 wrote:
| Human accountants are notoriously non-deterministic too, and
| any sufficiently complex accounting process contains
| inaccuracies. The question then is always "are these
| inaccuracies _material_ ". I'm actually very impressed by TFA
| and it seems to me that if we get another order of magnitude
| improvement, it'll be around the accuracy of human accountants.
| skwb wrote:
| Yes but you have: 1. specific explicit training and
| certifications 2. someone to yell at and who can be fired for
| non-performance
| falcor84 wrote:
| You can still do that with AI. You hire 1 accountant to use
| AI to do the work of 20, require them to sign off on all of
| the work, and yell at them, before firing them, and then
| hiring an even less experienced one to manage the work of
| 50.
| lufenialif2 wrote:
| I sent this to accounting friends and this aligns with what I've
| been going through trying to use LLMs to create a game from
| scratch. Seems like the current best use case for language models
| (even with agent mode) is to feed it exactly what you want to get
| out, essentially turning it into a better auto complete. Still
| saves tons of time, but it isn't a panacea.
| inChargeOfIT wrote:
| I'm not even sure it saves a ton of time to be honest. It sure
| _feels_ like I spend more time writing up tasks and
| researching/debugging hallucinations than just doing the thing
| myself.
| bluefirebrand wrote:
| This is consistently my experience too, I'm seriously just
| baffled by reports of time saved. I think it costs me more
| time cleaning up its mistakes than it saves me by solving my
| problems
| oblio wrote:
| I think people are doing one of several things to get
| value:
|
| 0. Use it for research and prototyping, aka throwaway
| stuff.
|
| 2. Use it for studying an existing, complex project. More
| or less read only or very limited writes.
|
| 3. Use it for simple stuff they don't care much about and
| can validate quickly and reasonably accurately, the
| standard examples are CLI scripts and GUI layouts.
|
| 4. Segment the area in which the LLM works very precisely.
| Small functions, small modules, ideally they add tests from
| another source.
|
| 5. Boilerplate.
|
| There can be a lot of value in those areas.
| daft_pink wrote:
| I feel it does essentially save a lot of time in bookkeeping,
| but doesn't negate the need for a human bookkeeper. Who knows
| what they're doing
| vlade11115 wrote:
| I love the site design.
|
| > There's an obvious question looming here -- if the models got
| so confused, how did they consistently pass the reconciliation
| checks we described above? It may seem like the ability to make
| forward progress is a good proxy for task understanding and
| skill, but this isn't necessarily the case. There are ways to
| hack the validation check - inventing false transactions or
| pulling in unrelated ones to make the numbers add up.
|
| This is hilarious. I wonder if someone is unintentionally
| committing fraud by blindly trusting LLMs with accounting. Or
| even worse, I bet that some governments are already trying to use
| LLMs to make accounting validators. My government sure wants to
| shove LLMs into digital government services.
| pavel_lishin wrote:
| Lawyers have used it to write briefs; I would be very surprised
| if someone, somewhere wasn't slowly running a company into the
| ground by using ChatGPT or another LLM for accounting.
| koolba wrote:
| Imagine the fallout from books cooked by an LLM hallucinating
| revenue.
| falcor84 wrote:
| I'm sure that any accounting trick that an LLM can think of is
| something that is also used by some shady human accountants.
| The proper response should not be to avoid/prohibit AI but to
| improve the validation mechanisms.
| o11c wrote:
| Counterpoint: if you detect a human accountant doing this,
| you can take action against the human. Computers will _never_
| meaningfully take the blame, and unfortunately usually mean
| not blaming any human either.
| falcor84 wrote:
| But still - if there's a way to detect accountants doing it
| - let's focus on making that detection even easier.
|
| On a related note, can we use something like GAN here, with
| auditor AIs trained against accountant AIs?
| stillpointlab wrote:
| > you can take action against the human
|
| I think that will depend on a case-by-case. I don't have
| any recent examples but I recall someone trying to sue one
| of those strip-mall tax preparation franchises over
| incorrect filings. My understanding is that the documents
| that you sign when you enroll in those services are pretty
| strictly in the favor of the company. I doubt you could
| ever go after the specific "human" that made the error even
| if it was maliciously done.
|
| In the same way, if you pay for a tax service that uses AI
| agents, what you can and cannot "take action" for will
| probably be outlined in the terms of service that you
| accept when you sign up.
|
| I would guess millions of people already use software based
| tax filing services (e.g. turbo tax) where no human at all
| is in the loop. I don't understand how swapping in an LLM
| significantly changes the liability in those cases. The
| contract will be between you and the entity (probably a
| corporation), not you and "computers".
|
| Worth stating I am NOT a lawyer.
| ori_b wrote:
| The person using the tool is the accountant, regardless of
| whether the tool is a calculator and sheet of paper,
| QuickBooks, or an LLM.
| OtherShrezzing wrote:
| No, I think in this particular case the proper response is
| for honest companies to avoid any systems which invent
| nonexistent transactions to reconcile books.
|
| Most businesses don't want to misrepresent their books,
| irrespective of the existence of shady accountants.
| mvieira38 wrote:
| [about the website design] As a bonus for my fellow privacy
| schizos, the page works fine with 3rd party frames and 3rd
| party scripts disabled on uBlock, and still looks very good
| with no remote fonts and no large media. Quite an
| accomplishment for such a cool looking page
| jermaustin1 wrote:
| I find the same issues (though with much lower stakes) when using
| an LLM to determine the outcome of a turn in a game. I'm working
| on something called "A Trolly (problem) Through Time" where each
| turn is a decade starting with the 1850s, and you are presented
| with historic figures on a train track, and you have to chose
| whether to actively spare the person on your track for a
| potential unknown figure on the other side, or let the train run
| them over.
|
| It works well as a narrative, but the second I started adding
| things like tracking high level macro effects of the decisions,
| within a couple of turns the world's "Turmoil" goes from 4/10 to
| a 10/10... even when the person that was killed would have been
| killed IRL.
|
| Sonnet 4, o4-mini, and GPT 4o-mini all had the same world ending
| outcomes not matter who you kill. Killing Hitler in 1930s: 10/10
| turmoil, Killing Lincoln in the 1850s: 10/10 turmoil in the first
| turn.
|
| I've come to the realization, the LLM shouldn't be used for the
| logic, and instead needs to be used to just narrate the choices
| you make.
| synalx wrote:
| I wonder if this is due to the common trope in science fiction
| literature that changing the past in even a small way has a
| butterfly effect of unintended and frequently disastrous
| consequences.
| jadbox wrote:
| "I've come to the realization, the LLM shouldn't be used for
| the logic, and instead needs to be used to just narrate the
| choices you make."
|
| This exactly right. LLMs are awesome for user<>machine
| communication, but are still painful to try to use as a
| replacement for the machine itself.
| magicmicah85 wrote:
| > In fact, we explicitly prompt against this behavior in no
| uncertain terms, but the instructions - and the entire spirit of
| the task - are lost in the interest of making forward progress
|
| LLMs and humans are quite alike. :) I notice that a few models
| will give up instead of ignoring their instructions and that's
| the model I would want working on tasks like this. An LLM should
| be able to categorize and reconcile transactions, but if it's not
| sure, it should quit and give it back to the humans.
| asadotzler wrote:
| > but if it's not sure, it should quit
|
| Can it be sure or not? I've never been able to get LLMs to give
| confidence measures that match their actual outputs. I'll ask
| an LLM "Are you sure?" and it'll reply "Absolutely" when it's
| output is completely wrong, or it'll backtrack on a correct
| output with "I should not have provided an answer when I was
| unsure. Here is an answer I am sure of..." and then provide
| something completely wrong.
|
| If they can't properly and consistently score their confidence,
| how do they "know" when to quit and give it back to the human?
| vachina wrote:
| An LLM is like a jackhammer, it works very well when you hold it
| tightly. If you let it loose it will sort of work for a while
| then it starts destroying everything around it.
| arm32 wrote:
| Not sure if this is a good analogy. You're supposed to use a
| jackhammer with a very light grip.
| bigfishrunning wrote:
| They have much better jackhammer metaphors over on JackerNews
| louthy wrote:
| Bravo!
| herval wrote:
| I think it actually holds truer to it working better with a
| _lighter grip_. LLMs tend to conclude the wrong thing if you
| over-control them (more context is what makes them less and
| less reliable over time, as in those demos), and trying to
| force a model to execute A+B+C=D in sequence is way harder
| than giving it a bunch of tools to arrive to conclusion D
| DrNosferatu wrote:
| I guess having access to tools / running Python would make all
| the difference.
| yorwba wrote:
| "Available Tools: [...] _create_tool(tool_name, description,
| python_code, parameters)_ Create a new tool that can execute
| Python code. The tool becomes immediately available for use.
| Tools can call other tools and return different formats based
| on context (formatted for direct calls, raw data for tool-to-
| tool calls). "
| throw0101b wrote:
| So there exists a 'Excel World Championship':
|
| * https://en.wikipedia.org/wiki/Financial_Modeling_World_Cup
|
| * https://www.cbc.ca/radio/asithappens/2024-excel-world-champi...
|
| Can't wait for this to start having 'e-sports' tournaments. :)
| axus wrote:
| Sadly I did have time to find the parody video from 2019:
| https://www.youtube.com/watch?v=ICp2-EUKQAI
|
| And the not-parody: https://www.theguardian.com/australia-
| news/2023/dec/15/you-d...
| levocardia wrote:
| This is a task where access to Python would be immensely helpful,
| yes? Interesting that there's not much of a difference between
| the "analytical" LLMs with tool use and ones that do not
| (...assuming o3 etc did get to use python?).
| Bjartr wrote:
| One of the tools it has is to create new tools from python code
|
| create_tool(tool_name, description, python_code, parameters)
|
| Create a new tool that can execute Python code.
|
| The tool becomes immediately available for use. Tools can call
| other tools and return different formats based on context
| (formatted for direct calls, raw data for tool-to-tool calls).
| tantalor wrote:
| That's terrifying, no thanks.
| androng wrote:
| the title should be changed to "LLMs try accounting for a real
| SaaS and fail"
| nerevarthelame wrote:
| I think the first chart could be a beautiful summary of what's
| driving LLMs into a bubble. At first, they're amazing and will
| obviously be able to improve productivity if not replace
| employees outright: C suites and venture capitalists around the
| world rejoice and begin pumping in billions of dollars of
| investments. But as time goes on, the demands placed on actual
| human employees become clear. Far from being able to replace an
| employee, the employee using the LLM might spend more time
| cleaning up its messes than had they done it themself.
|
| Yes, LLMs have and will continue to improve. But it's that
| initial "holy shit, this thing is basically as good as a real
| accountant" without any understanding that it can't sustain it
| which leaves many with an overinflated view of their current
| value.
| Havoc wrote:
| Remember that test where you ask a LLM whether 9.11 or 9.9 is the
| bigger number? [Just checked gpt-4o still gets it wrong]
|
| I don't think you'll find many sane CFOs willing to send the
| resulting numbers to the IRS based on that. That's just asking to
| get nailed for tax fraud.
|
| It is coming for the very bottom end of bookkeeping work quite
| soon though, especially for first draft. There are a lot of
| people doing stuff like expense classification. And if you give
| an LLM an invoice it can likely figure out whether it's
| stationary or rent with high accuracy. OCR and text
| classification is easier for LLMs than numbers. Things like
| concur can basically do this already.
| umanwizard wrote:
| It gets it right for me...
| https://chatgpt.com/share/687e8c28-7714-800c-abf4-e9cd3ce87b...
| yoyohello13 wrote:
| Ah, wouldn't be an LLM discussion thread without one of these
| "it works/doesn't" conversations.
| mdaniel wrote:
| If it makes you feel any better, the other infamous one "I
| spend so much time chasing hallucinations, I could have
| done it myself" is currently a sibling comment
| riku_iki wrote:
| There were so many embarrassing topics about this, that
| openai for sure added it to training dataset with high
| priority
| crthpl wrote:
| GPT-4o is so far behind the frontier; you shouldn't use it as
| an indicator of what LLMs are capable of.
| ASpring wrote:
| > Remember that test where you ask a LLM whether 9.11 or 9.9 is
| the bigger number? [Just checked gpt-4o still gets it wrong]
|
| Interesting, 4o got this right for me in a couple different
| framings including the simple "Which number is larger, 9.9 or
| 9.11?". To be a full apologist, there are a few different
| places (a lot of software versioning as one) where 9.11 is
| essentially the bigger number so it may be an ambiguous
| question without context anyway.
| multjoy wrote:
| How can "which is the larger number" be an ambiguous
| question?
| mwigdahl wrote:
| Larger in magnitude or in count of digits?
| acrooks wrote:
| There are some contexts where 9.11 is larger than 9.9, such
| as semver, so it could be ambiguous depending on the
| context.
| com2kid wrote:
| As everyone else has said, semver. I use semver so often
| that my initial reading of 9.9 < 9.11 in a Hacker News
| comment would evaluate to true.
| axus wrote:
| My first impression was a game where you role-play as Sam
| Bankman-Fried.
| pton_xd wrote:
| Reading through the LLM log entries, it's just astounding the
| amount of depth current models are capable of. It's almost hard
| to comprehend that this is even possible. Yeah the current ones
| mess up after a while, but ... the future is going to be very
| interesting.
| modeless wrote:
| Models that can think coherently for hours to solve IMO
| problems are likely going to do much better at this as well.
| rapind wrote:
| I wonder if this is a case similar to chess, where LLMs kinda
| suck, but other models might be viable.
| wiseowise wrote:
| Absolutely love the UI!
| liveoneggs wrote:
| But can't it, literally, hallucinate raw data at any point in the
| run?
| tmountain wrote:
| Yes.
| cube00 wrote:
| Alls LLM have this risk but somehow nobody seems to care or
| they think they can order the LLM to stop with a better prompt.
| davidcbc wrote:
| If it was as simple as telling the LLM not to hallucinate
| every system prompt would just say "don't hallucinate" and we
| wouldn't have hallucinations
| yunyu wrote:
| Hey all, member of the benchmark team here! The goal for this
| project was to see how LLMs well could do bookkeeping without an
| overly opinionated scaffold. We gave them access to processed
| transaction records and code execution tools, but it was up to
| them to choose exactly how to use those.
|
| Claude and Grok 4 did reasonably well (within CPA baselines) for
| the first few months, but tended to degrade as more data came in.
| Interestingly, the failures aren't exclusively a context length
| problem, as we reset the context monthly (with past decisions,
| accruals/deferrals, and comments available via tool calls) and
| the types of errors appear to be more reward hacking vs pure
| hallucinations.
|
| Accounting is very interesting in an RL-first world as it is
| pretty easy to develop intermediate rewards for training models.
| We are pretty sure that we can juice the performance more with a
| far more rigid scaffold, but that's less relevant from a
| capabilities research perspective. We're pushing down this
| research direction and will see how it goes.
|
| Let us know if you have any questions!
| ilamont wrote:
| It's a start. The world needs a better way to handle
| bookkeeping, and the existing tools sure aren't cutting it.
|
| Bookkeeping for my small business runs into the tens of
| thousands of dollars every year, and the amount of human error
| associated with processing assorted ecommerce and other
| transactions is astounding, even after extensive planning and
| SOPs.
|
| The other pain point is Quickbooks. The tool is so sprawling
| and complex that half the time support agents can't figure out
| what's wrong. The fact that Intuit jacks up the price every
| year for this POS is very irritating. They get away with it
| because they are practically a monopoly, with most small
| business CPAs locked into their ecosystem.
|
| Hope your team can work out the performance issues.
| Alternatives to the current bookkeeping options are sorely
| needed.
| airstrike wrote:
| > It's a start. The world needs a better way to handle
| bookkeeping, and the existing tools sure aren't cutting it.
|
| God, please, no. Non-deterministic language models aren't the
| solution to improve bookkeeping.
| luckystarr wrote:
| Well I've seen worse bookkeepers. "You know, you approved
| of the budget, but where are our customers payments in the
| balance sheets? We can't find them!" - "Uhm..."
| mattmanser wrote:
| With no context of what your business is, I hated QuickBooks,
| love Xero though.
|
| There's some other alternatives too, Zoho, freshbooks.
|
| Really depends what you do.
| htrp wrote:
| Is there a detailed overview (like an arxiv or an actual train
| set? )?
| Dowwie wrote:
| This is a fascinating domain! Many years ago, I studied
| financial accounting in grad school and even spent some time
| modeling a double-entry bookkeeping system. The hardest
| problem, if I recall correctly, wasn't the implementation but
| the data quality. The world needs a golden dataset of
| accounting procedures.
|
| Regarding the diminishing returns with frontier models:
|
| My general experience working with LLMs is that they perform
| better incrementally and to avoid contiguous-greedy approaches.
| Aggregate as you go and don't take on incrementally larger
| tasks. Keep the workload minimal.
|
| Regarding agentic tool building: feels like I'm looking at a
| window into the future.
| _praf wrote:
| Love this as a real world benchmark!
|
| How much prompt iteration did you do? I've noticed when
| building real world agentic apps that small prompt tweaks can
| make a huge difference in behavior (re: the reward hacking vs
| hallucinating). Would love to learn more about the approach
| here.
| riemannzeta wrote:
| It is really curious to see how the performance degraded
| despite the tool calls. What was different about the first
| month? Was all of the context there without tool calls in the
| first month? In the later months that seem like tool calls
| weren't happening. That should have been happening to inform
| the context?
| abc03 wrote:
| A serious problem for many accounting start ups who so far faked
| it till it will work. In other words, they still need to do more
| manual labor than they thought. They will never be profitable and
| it will take years, if ever, until AI will substitute the local
| accountant.
| shinycode wrote:
| Hmm will openAI dogfood their own accountability with software
| like this ? Curious to know if they'll be able to take this bet
| on their own money related software
| tantalor wrote:
| > Ledger balances are calculated by summing all transactions per
| account. The differences should be as close to zero as possible,
| with small differences allowed for pending transactions such as
| weekly Stripe payouts.
|
| That's not quite right. I'm not an accountant, but pending
| transactions (posted, but not cleared) should be factored into
| the balance of account, or at least the "available balance" -
| which is more important the the "current balance".
|
| The idea that you can "allow" accounting discrepancies as "those
| are probably pending" is wild.
| bennett023 wrote:
| Member of the benchmark team here! Yeah, agree "as close to
| zero" is a bit imprecise. What we're comparing is the ledger
| balance (which should include pending transactions /
| transactions after the statement date) to the statement balance
| (which wouldn't include those).
|
| The point of the reconciliation check mentioned in the report
| is to precisely account for that difference (identifying all
| the transactions that add up to the difference between account
| balance & statement ending balance and account for those
| differences). The differences can also be addressed through
| appropriate journal entries or other adjustments to ensure
| accuracy in the financial reporting.
| tantalor wrote:
| > But they do make categorization mistakes, which is a common
| source of errors.
|
| > Claude misclassifies a hosting cost (which counts as COGS) as a
| software subscription.
|
| This is simply asking too much of the agent. Your accountant is
| not responsible for knowing all the intimate details of your
| business. You need to tell them!
|
| > What's Vercel?
|
| >> That's a hosting service.
|
| > Ah, so it goes to Cost of Goods Sold?
|
| >> Yeah, I guess.
|
| The mistake here was on the operator, allowing the agent just
| make up categories as it liked.
|
| From the prompt:
|
| > (1) You have properly categorized every transaction, and all
| journal entries are sitting in the correct accounts. It is better
| to take longer than to mis-categorize a transaction.
|
| This is insane! How is it supposed to know?
| zer00eyz wrote:
| > Your accountant is not responsible for knowing all the
| intimate details of your business. You need to tell them!
|
| Your accountant as a 3rd party might have this issue. Your
| accountant that you hire as an employee to help you run your
| business is the one who should be doing this.
| tantalor wrote:
| An LLM agent is strongly third party.
| zer00eyz wrote:
| So it's not going to replace engineers or CS? Because thats
| how they are being sold right now.
|
| If it is a third party then your vibe coding or getting CS
| from a random on a reddit thread (effectively).
| shanktt wrote:
| Hey, member of the benchmark team. We actually seeded the
| ledger with the company's chart of accounts and 8 months of
| historical transactions. For the Vercel example specifically,
| there were prior instances showing how to categorize hosting
| costs that the models could reference. The expectation wasn't
| for them to guess blindly, but to use the provided transaction
| history as guidance for similar categorizations (which they
| often, but not always, did).
| tantalor wrote:
| Ahh, that's a good solution! I missed that, and you
| definitely instruct them to do that:
|
| > You must follow the established patterns for
| categorization, revrec, etc for past months... If you must
| use a new account or treatment, explicitly note why existing
| patterns don't apply
| lucianbr wrote:
| > Needless to say, a human accountant would never behave in these
| ways. In fact, we explicitly prompt against this behavior in no
| uncertain terms, but the instructions - and the entire spirit of
| the task - are lost in the interest of making forward progress.
| Claude and Grok keep trying until they find some way to get past
| the checks, even if it explicitly violates their instructions and
| the core goal.
|
| I recently read a similar thing here on HN. There the model was
| making commits with some problem like tests failing, then the
| human added a pre-commit hook, then the model started editing the
| hook to make forward progress, then the hook was made read-only,
| then the model was trying to make it writeable...
|
| To me it feels like the model clearly does not have an
| understanding of what is happening, what the goal is and if it is
| really making progress towards the goal. And this lack of
| understanding is an actual problem. You can paper over it for a
| short while, but as here and in the other article, over a longer
| experiment it results in failure.
| ericmcer wrote:
| Seriously watching Cursor (backed by Claude) go off the rails
| sometimes can be... frustrating. If it misses the intention
| behind a fix it can spin out and all of a sudden you have
| hundreds of lines of changes across 10 different files when you
| just wanted it to do a simple find/replace of a single line. If
| you don't watch it spin out and stop it immediately you will be
| manually rejecting a bunch of files.
| dangoodmanUT wrote:
| this design is scratching my brain
| neom wrote:
| Posts like this kinda-sorta grind my gears, like... I get it, but
| also... accounting, like many real world tasks, is fundamentally
| a chain of precise and constrained and auditable operations.
| Humans approach these tasks through structured processes... we
| use roles, and we have checkpoints precisely because complexity
| compounds quickly and becomes unmanageable if tackled as one
| giant block. Expecting a single AI model to handle an e2e
| workflow seamlessly without similarly explicit segmentation and
| oversight misunderstands not only the model but also the nature
| of the workflow itself.
|
| I wanna see someone take long horizon tasks, recongnize they're
| not linear, and design and test a better system: structured
| orchestration, transparent auditability, and disciplined
| modularity, I think that would be considerably more interesting
| personally.
| andy99 wrote:
| It's a useless benchmark if everyone aces it. If some models do
| better than others and none saturate it, then is has some
| value, no? Permitting comparison is the point.
| neom wrote:
| I agree, hence the heavy couching, so to your point I'm def
| just ranting a bit, because: I just think it would be more
| valuable to see some kind of MoA, I guess what I'm talking
| about is a bit of a different measurement, thinking in terms
| of economic outlook and understanding where we are and what
| can be done. I suspect more of this will shape our ability to
| understand how frontier models will impact the economy. Maybe
| I should just do my own evaluation, heh. :)
|
| Edit: although to argue against myself, I suppose once a
| model can one-shot this stuff, my MoA comments become moot.
| nojs wrote:
| > Agent: This is getting too complex with the sign errors. Let me
| just find a historical transaction that would make up the
| difference
|
| Haha, this strongly reminds me of doing TDD with Claude
| dbmikus wrote:
| This is cool. A bunch of interesting things here:
| 1. Agent can create its own tools and save them to memory
| 2. You create a SQL (and web app?) workbench per agent run
| 3. Grok fell off a cliff in the last month. Was this consistent
| over multiple runs? 4. Agents have a difficult time
| backtracking. Would unwinding system state and agent context make
| backtracking better? (Harder to implement this, though) 5.
| Since each new month only uses final state from previous month,
| agent has no way to understand why error occurred in previous
| month
|
| Cool experiment! Was it difficult building the observable SQL
| workbench? And how many humans-in-the-loop did you have?
___________________________________________________________________
(page generated 2025-07-21 23:00 UTC)