[HN Gopher] I'm worried that they put co-pilot in Excel
___________________________________________________________________
I'm worried that they put co-pilot in Excel
Author : isaacfrond
Score : 375 points
Date : 2025-11-05 08:54 UTC (14 hours ago)
(HTM) web link (simonwillison.net)
(TXT) w3m dump (simonwillison.net)
| cjs_ac wrote:
| At some point, a publicly-listed company will go bankrupt due to
| some catastrophic AI-induced fuck-up. This is a massive
| reputational risk for AI platforms, because ego-defensive
| behaviour _guarantees_ that the people involved will make as much
| noise as they can about how it 's all the AI's fault.
| ramon156 wrote:
| Do you really want these kind of companies to succeed? Let them
| burn tbh
| cjs_ac wrote:
| I don't find comments along the lines of 'those people over
| there are bad' to be interesting, especially when I agree
| with them. My comment is about _why_ it 'll go wrong for
| them.
| mcphage wrote:
| Make sure you're not part of the kindling, then.
| meibo wrote:
| That will never happen, AI cannot be allowed to fail, so we'll
| be paying for that AI bail-out.
| gosub100 wrote:
| I see the inverse of that happening: every critical decision
| will incorporate AI somehow. If the decision was good, the
| leadership takes credit. If something terrible happens, blame
| it on the AI. I think it's the part no one is saying out loud.
| That AI may not do a damn useful thing, but it can be a free
| insurance policy or surrogate to throw under the bus when SHTF.
| halfcat wrote:
| This works at most one time. If you show up to every board
| meeting and blame AI, you're going to get fired.
|
| This is true if you blame a bad vendor, or something you
| don't even control like the weather. Your job is to deliver.
| If bad weather is the new norm, you better figure out how to
| build circus tents so you can do construction in the rain. If
| your AI call center is failing, you better hire 20 people to
| answer phones.
| AmbroseBierce wrote:
| Brenda has been getting slower over the years -as we all have-,
| but soon the boss will learn that it was a small price to pay for
| knowing well how to keep such house of cards from collapsing.
| Simulacra wrote:
| And then the boss will make the decision to outsource her job,
| to a company that promises the use of AI to make finance
| better, and faster, and while Brenda is in the unemployment
| line, someone else thousands of miles away is celebrating a new
| job
| gadflyinyoureye wrote:
| We are setting AI deployed in the US, but actually Indians.
| They are not better, but they are cheaper. They are probably
| worse, but they are cheaper.
| Traster wrote:
| I'm actually not that worried about this, because again I would
| classify this as a problem that already exists. There are already
| idiots in senior management who pass off bullshit and screw
| things up. There are natural mechanisms to cope with this,
| primarily in business reputation - if you're one of those idiots
| who does this people very quickly start just discounting what
| you're saying, they might not know _how_ you 're wrong, but they
| learn very quickly to discount what you're saying because they
| know you can't be trusted to self-check.
|
| I'm not saying that this can't happen and it's not bad. Take a
| look at nudge theory - the UK government created an entire
| department and spent enormous amounts of time and money on what
| they thought was a free lunch - that they could just "nudge"
| people into doing the things they wanted. So rather than
| _actually solving difficult problems_ the uk government embarked
| on decades of pseudo-intellectual self agrandizement. The entire
| basis of that decades long debacle was based on bullshit data and
| fake studies. We didn 't need AI to fuck it up, we managed it
| perfectly well by ourselves.
| gmac wrote:
| Nudge theory isn't useless, it's just not anything like as
| powerful as money or regulation.
|
| It was taken up by the UK government at that time because the
| government was, unusually, a coalition of two quite different
| parties, and thus found it hard to agree to actually use the
| normal levers of power.
|
| This NY Times opinion piece by Loewenstein and Ubel makes some
| good arguments along these lines:
| https://web.archive.org/web/20250906130827/https://www.nytim...
| simonw wrote:
| This quote is pulled from a TikTok, I recommend watching the
| whole thing here:
| https://www.tiktok.com/@belligerentbarbies/video/75683800086...
|
| (I pulled the quote by using yt-dlp to grab the MP4 and then
| running that through MacWhisper to generate a transcript.)
| donatj wrote:
| It's a little over two paragraphs. Seems like it would have
| been simpler just to... type it out?
| adlpz wrote:
| Where's the fun in that? :D
| mavhc wrote:
| We choose to automate these things, not because they are
| easy, but because they are an interesting problem to solve
| daliusd wrote:
| Well if you do it once then yes, but if you automate this
| process it is different. E.g. I do this with YouTube videos,
| because watching 14 minutes video or reading 30 seconds
| summary is time saver. I still watch some videos fully, but
| many of them are not worth it.
|
| So in summary I think it was just part of automated process
| (maybe) or it will become one in the future.
| rererereferred wrote:
| But then you would need a Brenda. Ai can write the automation
| script for you.
| simonw wrote:
| Why spend two minutes typing (and realistically longer than
| that, if I want to capture the exact transcript I would need
| to keep hitting pause and play and correcting myself) when I
| can spend ten seconds pasting a URL into my terminal and then
| dragging and dropping the resulting file onto the MacWhisper
| window?
|
| I actually transcribed the whole TikTok which was about 50%
| longer than what I quoted, then edited it down to the best
| illustrative quote.
| self_awareness wrote:
| I can see that MacWhisper uses parakeet v2 as the model
| (although it allows choosing another model).
|
| Is MacWhisper a $60 GUI for a Python script that just runs the
| model?
| trenchpilgrim wrote:
| > Is MacWhisper a $60 GUI for a Python script that just runs
| the model?
|
| Yes, a large genre of MacOS apps are "Native GUI wrappers
| around OSS scripts"
| jimbokun wrote:
| A lot of MacOS itself is this.
|
| Which is incredibly value. The OSS script has zero value to
| someone who doesn't know it exists or doesn't understand
| how to run it.
| simonw wrote:
| There's also a free version that just uses Whisper. I
| recommend giving it a go, it's a very well constructed GUI
| wrapper. I use it multiple times a week, and I've run Whisper
| on my machine in other less convenient ways in the past.
| sionisrecur wrote:
| You... could have given the job to Brenda instead, unless the
| irony was the point?
| simonw wrote:
| The global economy isn't going to crash if I make a mistake
| with the transcript.
| sionisrecur wrote:
| That's how it starts.
| 6thbit wrote:
| This may be the first quote from TikTok reposted on a blog,
| that ends up this high up in HN.
| glimshe wrote:
| This reminds me of a friend whose company ran a daily perl script
| that committed every financial transaction of the day to a
| database. Without the script, the company could literally make no
| money irrespectively of sales because this database was one piece
| in a complex system for payment processor interoperability.
|
| The script ran in a machine located at the corner of a cubicle
| and only one employee had the admin password. Nobody but a
| handful of people knew of the machine's existence, certainly not
| anyone in middle management and above. The script could only be
| updated by an admin.
|
| Copilot may be good, but sure as hell doesn't know that admin
| password.
| danielbln wrote:
| If your mission critical process sits on some on-site box that
| no-one knows about, copilot being good or not is the least of
| your problems.
| maccard wrote:
| Everywhere I've ever worked has had that mission critical
| box.
|
| At one of my jobs we had a server rack with UPS, etc, all the
| usual business. On the floor next to it was a dell desktop
| with a piece of paper on it that said "do not turn off". It
| had our source control server in it, and the power button
| didn't work. We did eventually move it to something more
| sensible but we had that for a long time
| victorbjorklund wrote:
| with only one person on earth being able to access it? so
| if that person is hit by a car everything goes down?
| glimshe wrote:
| Pretty much
| maccard wrote:
| yeah. I mean, someone else would _eventually_ figure it
| out. There wasn't full disk encryption or anyhting on it,
| so if the guy got hit by a bus, and the machien turned
| off we probably would have just imaged the disk and got
| it running in a VM.
|
| But we didn't (and nobody was hit by a bus)
| victorbjorklund wrote:
| and you think that is good practice? sounds pretty
| terrible.
| maccard wrote:
| I never said it's good practice, simply that it happens.
| ozim wrote:
| This sort of gimmick is not going to help anyone keeping their
| job.
| chaps wrote:
| Sadly, nah. It works.
| chaps wrote:
| An old colleague and friend used to print out a 30 page perl
| script he wrote to do almost exactly this in this scenario. A
| stapled copy could always be found on his dining room table.
| onionisafruit wrote:
| Was the printed copy a backup system or casual reading?
| chaps wrote:
| Yes.
| buellerbueller wrote:
| <3 inclusive or.
| victorbjorklund wrote:
| That sounds pretty bad. Not a great argument against AI: "Our
| employees have created such a bad mess that AI wont work
| because only they know how the mess they created works".
| intended wrote:
| That is the luxury of theory.
|
| Yes, most situations are terrible compared to what would be
| if an expert was present to perfect it.
|
| Except if there isn't an expert, and there's a normal person,
| how do they know the output is right ?
| victorbjorklund wrote:
| not sure I get your point?
| timeinput wrote:
| I think the parent is saying what if the AI made such a
| terrible mess that the team of imperfect people thought
| it was fine, but it was just as bad as the terrible mess
| the team would have created because the team is not
| capable of evaluating whether it's a good idea or not.
| (possible follow on consequences -- no one can debug it
| or figure out if it's a good idea either)
| recursive wrote:
| The point is that in real companies the bad mess already
| exists. So it is a good argument. Or at least a practical
| one.
| jimbokun wrote:
| > "Our employees have created such a bad mess that AI wont
| work because only they know how the mess they created works".
|
| This is an ironclad argument against fully replacing
| employees with AI.
|
| Every single organization on Earth requires the people who
| were part of creating the current mess to be involved in
| keeping the organization functioning.
|
| Yes you can improve the current mess. But it's still just a
| slightly better mess and you still need some of the people
| around who have been part of creating the new mess.
|
| Just run a thought experiment: every employee in a
| corporation mysteriously disappear from the face of the
| Earth. If you bring in an equal number of equally talented
| people the next day to run it, but with no experience with
| the current processes of the corporation, how long will it
| take to get to the same capability of the previous employees?
| testing22321 wrote:
| My last job at a telco I was in charge of a system that billed
| ~5 million dollars monthly. When the machine was built, the guy
| that did it didn't record the root password. He added me to
| sudoers before he left. I left a few years later, nobody took
| ownership.
|
| Looking at the web interface, I can tell it's still running,
| doing its thing. I'm sure its still running Linux from 2008.
| mikert89 wrote:
| 10 billion dollars is probably going to be spent on automating
| excel, it's going to happen
| harryf wrote:
| There needs to a financial equivalent to the Mythical Man
| Month.
| graemep wrote:
| There are plenty of things that play the role.
|
| The problem is that people ignore them.
| eithed wrote:
| Let it all crash and burn
| HeavyStorm wrote:
| Nay-sayers need to decide whether they fear AI because AI is dumb
| and will fuckup or because AI is smart and will take over.
| tossandthrow wrote:
| Simon willson is definitely not a nay sayer.
| 9dev wrote:
| Both are valid concerns, no need to decide. Take the USA: They
| are currently lead by a patently dumb president who fucks up
| the global economy, and at the same time they are powerful
| enough to do so!
|
| For a more serious example, consider the Paperclip Problem[0]
| for a very smart system that destroys the world due to very
| dumb behaviour.
|
| [0]: https://cepr.org/voxeu/columns/ai-and-paperclip-problem
| RajT88 wrote:
| The paperclip problem is a bit hand-wavey about intelligence.
| It is taken as a given than unlimited intelligence would
| automatically win presumably because it could figure out how
| to do literally anything.
|
| But let's consider real life intelligence:
|
| - Our super geniuses do not take over the world. It is the
| generationally wealthy who do.
|
| - Super geniuses also have a tendency to be terribly
| neurotic, if not downright mentally ill. They can have
| trouble functioning in society.
|
| - There is no thought here about different kinds of
| intelligence and the roles they play. It is assumed there is
| only one kind, and AI will have it in the extreme.
| 9dev wrote:
| To be clear, I don't think the paperclip scenario is a
| realistic one. The point was that it's fairly easy to
| conceive an AI system that's simultaneously extremely
| savant and therefore dangerous in a single domain, yet
| entirely incapable of grasping the consequences or wider
| implications of its actions.
|
| None of us knows what an actual, artificial intelligence
| really looks like. I find it hard to draw conclusions from
| observing human super geniuses, when their minds may have
| next to nothing in common with the AI. Entirely different
| constraints might apply to them--or none at all.
|
| Having said all that, I'm pretty sceptical of an AI
| takeover doomsday scenario, especially if we're talking
| about LLMs. They may turn out to be good text generators,
| but not the road to AGI. But it's very hard to make
| accurate predictions in either direction.
| hunterpayne wrote:
| > The point was that it's fairly easy to conceive an AI
| system that's simultaneously extremely savant and
| therefore dangerous in a single domain, yet entirely
| incapable of grasping the consequences or wider
| implications of its actions.
|
| I'm pretty sure there are already humans who do this.
| Perhaps there are even entire conferences where the
| majority of people do this.
| victorbjorklund wrote:
| Silly calling Simon a nay-sayer.
|
| Are you a fanatic that thinks anyone saying that there are any
| limitations to current models = nay-sayer?
|
| Like if someone says they wouldnt wanna get a heart transplant
| operation done purely by GPT5, are they a nay-sayer or is that
| just reflecting reality?
| masswerk wrote:
| Our product has many issues. You must pick one and must not
| discuss any other.
| Dumblydorr wrote:
| Co-pilot and AI has been shoved at the Microsoft Stack in my org
| for months. Most of the features were disabled or hopelessly bad.
| It's cheaper for Microsoft to push this junk and claim they're
| doing something, it's going to improve their stock far more than
| not doing it, even though it's basically useless currently.
|
| Another issue is that my org disallows AI transcription bots.
| It's a legit security risk if you have some random process
| recording confidential info because the person was too busy to
| attend the meeting and take notes themselves. Or possibly they
| just shirk off the meetings and have AI sit in.
| aDyslecticCrow wrote:
| Transcription is arguably one of the must useful enterprise AI
| tools avaliable. But i sure as heck wouldn't trust the cloud
| with it.
| 2dvisio wrote:
| Still find the Copilot transcripts orders of magnitude worse
| than something like Wispr Flow and they tend to allucinate
| constantly and do not adapt to a company's context (that
| Copilot has access too...). I am talking about acronyms of
| products / teams, names of people (even when they are in the
| call), etc.
| srean wrote:
| Can anyone familiar with the technical details shed light
| on why this is so.
|
| Is it because of a globally trained model (as opposed to
| trained[tweaked on] on context specific data) or because of
| using different classes of models.
| aDyslecticCrow wrote:
| Neither copilot nor flow can natively handle audio to my
| understanding, so there is already a transcription model
| converting it to text that then GPT tries to summarise.
|
| It could be they simply use a mediocre transcription
| model. Wispr is amazing but would hurt their pride to use
| a competitor.
|
| But i feel it's more likley the experience is; GPT didn't
| actually improve on the raw transcription, just made it
| worse. Especially as any miss-transcipted words may trip
| it up and make it misunderstand while making the summary.
|
| if i can choose between a potentially confused and
| misunderstood summary, and a badly spellchecked (flipped
| words) raw transcription, i would trust the latter.
| aDyslecticCrow wrote:
| Ye i didn't even think about advanced meetings summary
| bots. Just raw word for word transcription please. Wispr is
| pretty great.
| kjkjadksj wrote:
| It is notoriously unreliable
| pjmlp wrote:
| The worse part is to see it creep on developer stack at places
| where it should not be.
|
| I am all good for nice completion on VS, or help decypher
| compiler errors, but lets do this AI push with some contention.
|
| Also what I really deslike is the prompt interface, AI
| integrations have to feel natural transparent part of the
| workflow, not trying to put everything into a tiny chat window.
|
| And while we're at it, can we please improve voice
| reckognition?
| blibble wrote:
| > It's cheaper for Microsoft to push this junk and claim
| they're doing something
|
| this has been the microsoft business model for 40 years
| Esophagus4 wrote:
| Hmmm the Brendas I know look a little different.
|
| "There are two Brendas - their job is to make spreadsheets in the
| Finance department. Well, not quite - they add the months and
| categories to empty spreadsheets, then they ask the other
| departments to fill in their sales numbers every month so it can
| be presented to management.
|
| "The two Brendas don't seem to talk, otherwise they would realize
| that they're both asking everyone for the same information,
| twice. And they're so focused on their little spreadsheet worlds
| that neither sees enough of the bigger picture to say, 'Wait...
| couldn't we just automate this so we don't need to do this song
| and dance every month? Then we wouldn't need two people in
| different parts of the company compiling the same data manually.'
|
| "But that's not what Brenda was hired for. She's a spreadsheet
| person, not a process fixer. She just makes the spreadsheets."
|
| We need fewer Brendas, and more people who can automate away the
| need for them.
| SirFatty wrote:
| "We need fewer Brendas, and more people who can automate away
| the need for them."
|
| True... I have an on-staff data engineer for the purpose. But
| not all companies (especially in the SMB space) have that
| luxury.
| 7thaccount wrote:
| That's a pretty specific example when there are a lot of good
| "spreadsheet people" out there who do a lot more than
| spreadsheets (maybe they had to write SQL queries or scripts to
| get those numbers), but commonly need to simplify things down
| to a spreadsheet or power point for upper management. I'm not
| saying you should have multiple people doing redundant work,
| but this style isn't entirely dumb.
|
| What would this be replaced by? Some kind of large SAP like
| system that costs millions of dollars and requires a dozen IT
| staff to maintain?
| Esophagus4 wrote:
| Fair - I was creating a straw man mostly to make a point. The
| people I'm thinking aren't running SQL queries or scripts,
| they're merely collection points for data.
|
| So one good BI developer who knows Tableau and Salesforce and
| Excel and SQL can replace those pure collection points with a
| better process, but they can also generate insight into the
| data because they have some business understanding from being
| close to the teams, which is what my hypothetical Brenda
| can't do.
|
| In my example, Brenda would be asking sales leaders to enter
| in their data instead of going into Salesforce herself
| because she doesn't know that tool / side of the company well
| enough.
|
| I was making the point that, contrary to the article, the
| Brendas I know aren't touched by the Excel angels, they're
| just maintaining spreadsheets that we probably shouldn't have
| anyway.
| 7thaccount wrote:
| I think that is a fair point too. The person that builds
| the Tableau dashboard could just send Brenda a screenshot
| once a month and that saves everyone time.
| simonw wrote:
| A screenshot of a Tableau dashboard is possibly the most
| dangerous form of internal data communication there is,
| because it entirely removes any chance of digging into
| that dashboard and figuring out what queries created it
| and spotting the incorrect assumptions they made along
| the way.
|
| A hill I will die on is that business analytics need
| "view source" or they aren't worth the pixels they are
| rendered with.
| 7thaccount wrote:
| I respectfully disagree. The amount of folks in upper
| management that can actually use something like Tableau
| is very small. However it doesn't matter as none of those
| people have the time outside of very small businesses
| which probably don't need it anyway. The business
| intelligence person is supposed to deliver succinct
| insights to upper management to act on, not say "here's a
| cool system I built you... figure it out yourself".
| Executives aren't getting paid $$$$$$$$ to do data
| analysis. Hopefully I'm not misrepresenting your point.
| simonw wrote:
| The view source link isn't there so upper management who
| don't know SQL can look at it. It's there so other people
| in the organization who _do_ know SQL have the
| opportunity to check the work.
|
| At my last large employer I genuinely lost count of the
| number of times I saw a BI report which pulled numbers
| from our data warehouse... and then found out it had
| misinterpreted a key detail because the engineering team
| had changed some table design six months ago and the data
| analysis team hadn't been told about the change.
| 7thaccount wrote:
| We're talking about different things than. I agree it's
| helpful to have an open system that the technical staff
| can drill into. I'm just saying at the end of the day
| that the key decision makers don't care. They need some
| simple high level metrics that can be put into some
| relatively simple charts and tables.
| oytis wrote:
| And then you end up with a team of five people each tree times
| as expensive as Brenda, and what used to be an email now takes
| a sprint and has to go through ticket system.
| Esophagus4 wrote:
| That's not what I had in mind.
|
| Then you end up with a report that goes out automatically
| every month to leadership pulled directly from the Salesforce
| data, along with a real time dashboard anyone in the org can
| look at, broken down by team, vertical, and sales volume.
|
| Why are people so attached to manual process?
| 131012 wrote:
| Because when one exec ask: "Why is that?" the room goes
| silent.
| daveguy wrote:
| It's not what you had in mind, but that's what you get.
| Because automation, integration, and AI are currently
| garbage -- Salesforce, Netsuite, doesn't matter. They don't
| do the magic that they promise. Because process is still
| very much a human problem, not a computational one.
| RegW wrote:
| > But that's not what Brenda was hired for.
|
| Are you suggesting that Brenda should stay in her box?
| Esophagus4 wrote:
| No, I'm suggesting that she is ineffective exactly _because_
| she stays in her box.
|
| She should replaced with someone who says, "this box doesn't
| need to be here... there is a better way of doing things."
|
| NOT to be confused with the junior engineer who comes into a
| project and says it's garbage and suggests we rewrite it from
| scratch in ${hotLanguage} because they saw it on a blog
| somewhere.
| daveguy wrote:
| It may not be what you meant to say, but it's exactly what
| you are saying where ${hotLanguage} is the latest
| automation platform or AI gimmick.
| Esophagus4 wrote:
| I'm not sure why you're going down to the mat for hanging
| onto redundant people putting numbers in spreadsheets.
|
| At large companies in particular, there are far too many
| people who simply turn their widgets - this was the
| entire point of the tech revolution.
|
| Think about how many bookkeepers were needed before
| Excel. Someone could have made your exact same argument
| (but it's just the latest gimmick!) about Excel 30 years
| ago. And yet, technology will make businesses more
| efficient whether people stand in its way or not.
|
| Even at a small company of one or two, QuickBooks will
| reduce the amount of bookkeepers and accountants needed.
| TurboTax will further reduce that.
|
| We will need fewer people in the future maintaining their
| Excel spreadsheets, and more people building the
| automation for those processes.
|
| The change averse will always find reasons not to adapt -
| they will create their own obsolescence.
|
| (inb4 but it's way more expensive to pay developers to
| automate!)
| daveguy wrote:
| I'm not going to the mat for anyone. I'm just saying AI
| use in spreadsheets is a terrible idea because AI just
| isn't that good.
|
| Currently I'd put it worse than tearing things up for
| ${hotLanguage} because at least ${hotLanguage} is
| deterministic and debuggable.
|
| Honestly, I'm not sure why you're going to the mat for AI
| in spreadsheets, or why you think it's a good use case,
| or why you seem to think "automation" doesn't come with
| overhead of its own. Current iterations of AI are
| recommendation engines. Even then you better have version
| control.
| jimbokun wrote:
| > She should replaced with someone who says, "this box
| doesn't need to be here... there is a better way of doing
| things."
|
| The article is about this kind of Brenda.
| jimnotgym wrote:
| With respect, you probably only see that bit of Finance, but
| doesn't mean that is all Brenda does.
|
| At least half of the work in my senior Finance team involves
| meeting people in operations to find out what they are planning
| to do and to analyse the effects, and present them to decision
| makers to help them understand the consequences of decisions.
| For an AI to help, someone would have to trigger those
| conversations in the first place and ask the right questions.
|
| The rest of the work involves tidying up all the exceptions
| that the automation failed on.
|
| Meanwhile copilot in Excel can't even edit the sheet you are
| working on. If you say to it, 'give me a template for an
| expense claim' it will give you a sheet to download... probably
| with #REF written in where the answers should be.
| nashashmi wrote:
| > We need fewer Brendas...
|
| We need more Brendas (those who excel goddesses come and kiss
| on the forehead) and need less people who are disrespectful of
| Brendas. The example in this post is someone giving more
| respect to AI than Brenda.
| martin-t wrote:
| Y'know why people don't automate their jobs? It's not a skill
| issue it's an incentives issue.
|
| If you do your job, you get paid periodically. If you automate
| your job, you get paid once for automating it and then nothing,
| despite your automation constantly producing value for the
| company.
|
| To fix this, we need to pay people continually for their past
| work as long as it keeps producing value.
| Esophagus4 wrote:
| Not always:
|
| If you don't automate it:
|
| 1a) your company keeps you hanging on forever maintaining the
| same widget until the end of time
|
| OR
|
| 1b) more likely, someone realizes your job should be
| automated and lays you off at some point down the road
|
| If you do automate it
|
| 2a) your company thanks you then fires you
|
| OR
|
| 2b) you are now assigned to automate more stuff as you've
| proven that you are more valuable to the company than just
| maintaining your widget
|
| --------
|
| 2b is really the safest long term position for any employee,
| I think. It's not always foolproof, as 2a can happen.
|
| But I'd rather be in box 2 than box 1 any day of the week if
| we're talking long term employment potential.
| martin-t wrote:
| Yes, but notice what you are describing are all negative
| incentives.
|
| When automation produces value for the company, the people
| automating it should capture a chunk of that value _as a
| matter of course_.
|
| Even if you argue that you can then negotiate better
| compensation:
|
| 1) That is uncertain and delayed reward - and only if other
| people feel like it, it's not automatic.
|
| 2) The reward stops if you get fired or leave, despite the
| automation still producing value - you are also basically
| incentivized to build stuff that requires constant
| maintenance. Imagine you spend a man-month building the
| automation and then leave, it then requires a man-month of
| maintenance over the next 5 years. At the end of the 5
| years, you should still be getting 50% of the reward.
| Esophagus4 wrote:
| My knee jerk reaction is to disagree, but on second
| thought, I'm open to hearing the argument.
|
| What would that look like in practice?
| martin-t wrote:
| I don't have a full theory yet, it's something I started
| thinking about recently.
|
| That being said, it's clear that in the current system,
| rich people can get richer faster than poor people.
|
| We have a two class system a) workers who get paid per
| unit of work b) owners who capture any surplus income,
| who decide hiring/firing/salaries, who can sell the
| company and whose wealth keeps increasing (assuming the
| company does well) whether they do any work themselves.
|
| Note: I see very few things which have inherent value -
| natural resources (plus land?) and human time. Everything
| else (with monetary value) is built from natural
| resources using human time.
|
| ---
|
| If a company starts with 1 guy in a shed, he does 100% of
| the work, owns 100% of the company and ... it gets muddy
| here ... gets 100% of the income / decides where 100% of
| the revenue goes - if it's a grocery shop he can just
| pocket any surplus, if he's making stuff, he'll probably
| reinvest into better tooling or to hire more workers.
|
| A year later, he hires 9 workers. Now he does only 10%
| but still owns 100% of the company.[0]
|
| There's a couple issues here:
|
| - He owns 100% of the future value of the company despite
| being created only 10% by him. Well, not exactly, he was
| creating 100% for the first year and 10% from then on.
|
| - He still gets to decide who gets paid what. He has more
| information when negotiating.
|
| - He can sell the company to whoever and the workers have
| no say in it. He can pass it on to his children (who
| performed 0 work there) when he dies.
|
| _The solution I 'd like to see tested is ownership being
| automatically and periodically (each month) redistributed
| according to the amount and skill level of work
| performed._[1]
|
| So at the end of year 2, the original founder has done 2
| man-years of work, while the other 9 people have done 1
| man-year of work each. This means the founder owns
| 2/11ths of the company while everyone else owns 1/11th.
| This could further be skewed by skill levels. I am sure
| starting and running a company for a year takes more
| skill than doing only some tasks. OTOH there are
| specialized tasks which only very few people can perform
| and the founder is not one of them.
|
| The skill level involved would be part of the
| negotiations about compensation.
|
| ---
|
| This is complex. I am sure somebody is prone to rejecting
| it based solely on that. But open a wiki page about e.g.
| bonds[2] and see how many blue words just the initial
| sentence has and ask yourself whether you could explain
| all of them (and then transitively all the linked
| concepts on their wiki pages).
|
| Slavery is very simple but very unfair. Employment is
| more complex and less unfair. I have a theory that the
| more fair a system is, the more complex it is because it
| needs to capture more nuances of the real world.
|
| ---
|
| [0]: Some people think this is right because owners take
| all the risk and employees take 0 risk. That is
| misrepresenting what really happens - sane
| investors/owners don't risk losing so much they would go
| homeless/starve if they lose it all. They can also
| optimize their risk by spreading it across many
| companies. Meanwhile workers get 100% of their income
| from one company and drop down to no income if the
| company goes bankrupt. They can also be fired at any
| time.
|
| This was argued here:
| https://news.ycombinator.com/item?id=45731811 in the
| comment by kristov and the reply by me. I also have other
| comments there with relevant ideas.
|
| [1]: What happens to monetary compensation? I don't know,
| I see multiple options:
|
| a) Everybody gets paid monetary wages like today, plus
| (newly) a part of their reward is the growing share of
| the company they own. If we allow selling it to anyone,
| it has high monetary value but then ownership gets
| diluted to outside investors. If we allow selling it only
| back to the company, it has value only relative to the
| decision-making power it gave. If we don't allow selling
| it, its monetary value only comes from the ability to
| vote on dividends.
|
| b) Everybody gets paid a portion of the income divided
| according to their share. This sounds simple but likely
| wouldn't give enough money to newly joined workers to
| survive. There could be a floor. (Or, because hard
| cutoffs suck, a smooth mathematical function from owned
| percentage to monthly compensation which would have a
| floor at minimum wage.)
|
| [2]: https://en.wikipedia.org/wiki/Bond_(finance)
| hunterpayne wrote:
| Two things:
|
| > - He owns 100% of the future value of the company
| despite being created only 10% by him. Well, not exactly,
| he was creating 100% for the first year and 10% from then
| on.
|
| 1) If you believe this, then you have a massively
| simplistic view of employee value. The distribution of
| actual value provided by employees is probably log
| normal, and certainly not normal (gaussian).
|
| 2) This is basically the labor theory of value. That is
| an economic theory that was discarded as wrong about 150
| years ago. If it was true, the value of a newly
| discovered gold mine would be 0.
| agumonkey wrote:
| it's a large human behavior question for me, the notion of
| work, value, economy, efficiency .. all muddied in there
|
| - i used to work on small jobs younger, as a nerd, i could
| use software better than legacy employees, during the 3
| months, i found their tools were scriptable so I did just
| that. I made 10x more with 2x less mental effort (I just
| "copilot" my script before it commits actual changes) all
| that for min wage. and i was happy like a puppy, being free
| to race as far as i want it to be, designing the script to
| fit exactly the needs of an operator. (side
| note, legacy employees were pissed because my throughput
| increase the rate of things they had to do, i didn't foresee
| that and when i offered to help them so they don't have to
| work more, they were just pissed at me)
|
| - later i became a legit software engineer, i'm now paid a
| lot all things considered, to talk to the manager of legacy
| employees like the above, to produce some mediocre web app
| that will never match employees need because of all the
| middle layers and cost-pressure, which also means i'm tired
| because i'm not free to improve things and i have to obey the
| customer ...
|
| so for 6x more money you get a lot less (if you deliver,
| sometimes projects get canned before shipping)
| martin-t wrote:
| I had a broadly similar transition in feeling about my
| work.
|
| It's not about how much I get paid. It's about realizing
| how much of the value I produce goes to me and how much
| goes to the owner class.
|
| At least I never worked in a big corporation and I always
| had the ability to do work that directly benefited people
| using my code. But I still saw too much of the "I built
| this company" self-congratulatory BS from people who just
| shuffled money while doing 0 actual work.
|
| I don't think ownership is theft, I just think it's
| distributed wrongly - to people who have money instead of
| to people who do work. See my other comment here:
| https://news.ycombinator.com/item?id=45826823
| agumonkey wrote:
| even though my above message wasn't much about the
| corporate leeches, i did experience the fun of being my
| own boss in a way during covid doing mini gigs directly
| with people
|
| there's a blend of "i'm my own man": i get the money and
| handle the responsibility on my own and it's thrilling
| feeling
|
| i don't dimiss the layers of HR managing legal and
| financial duties in a company and thus taking a cut, but
| there's a kind of pleasure to also do your own business
| for a while
| martin-t wrote:
| > i don't dimiss the layers of HR managing legal and
| financial duties in a company and thus taking a cut
|
| I don't wanna dismiss them either but (along with
| management):
|
| - It's not positive-sum work. It doesn't produce positive
| value for society, it's just necessary work which needs
| to be done as a side effect of actual positive-sum work
| being done.
|
| - The pyramid should be inverted. Managers, layers,
| accountants, etc. should be assistants. The people doing
| the actual work should (collectively) decide to hire them
| when they think it would make them more productive or be
| otherwise beneficial to them. Not the other way around.
| agumonkey wrote:
| it's an interesting question as of why the management
| layer has always been seen as more important than the
| builders, crafters, designers below
| Libidinalecon wrote:
| This is just not true at all.
|
| It is always in my self interest to automate my job as much
| as possible. Nothing looks better for moving up than this.
| Even more so, nothing makes me happier than automating a
| business process.
|
| There are always so many various road blocks to automation it
| is hard to count.
|
| It is like there is a type of entropy that increases over
| time that people are largely getting paid to keep at bay with
| simple business processes that can be easily adapted as
| things change. So often automation works great for a short
| time until this entropy breaks the automation. It doesn't
| take that many times for management to figure out the
| investment in automation gives poor returns.
| conductr wrote:
| I work in corporate finance and these issues are certainly
| present. However, they are almost always known and determined
| low priority to have a better process built. Finance processes
| are nearly always a non priority as a pure cost center/overhead
| there's not many companies that want to invest in improving the
| situation, they'll limp along with minimal investment even once
| big and profitable.
|
| That said, every finance function is different and it may be
| unknown to them that you're being asked for some data multiple
| times. If you're enduring this process, I'm of the opinion
| you're equally at fault. Suggest a solution that will be easier
| on you. As it's possible they don't even know it's happening.
| In the case provided, email to all relevant finance people
| "Here's a link to a shared workbook. I'll drop the numbers here
| monthly, please save the link and get the data directly from
| that file. Thanks!" Problem solved. Until you don't follow
| through which is what causes most finance people to be
| constantly asking for data/things. So be kind and also set
| yourself a monthly recurring reminder on your calendar and
| actually follow through.
| Esophagus4 wrote:
| I've just set the finance people up with read only access to
| our data source, and they now can poke through it themselves.
| conductr wrote:
| Also an acceptable solution. This is usually where the next
| step is have a BI type person just create a report for
| finance. Many reasons but what will end up is different
| people are filtering/retrieving the data differently
| causing inconsistencies.
|
| But Usually finance is always preferring on demand access
| so the communication feedback loop of asking for stuff is
| not well liked so I'm sure they appreciate this middle step
| too.
|
| There are many cases where there's no easy way to give
| access to the data and a human in the loop is required. In
| that case, do the shared workbook thing I mentioned as a
| starting point at least. It may evolve from there.
| xnorswap wrote:
| And they've all been burned by enterprise finance products
| which were sold to solve exactly that problem.
|
| Only different companies were all sold different enterprise
| finance products, but they need to communicate with each
| other (or themselves after mergers), so it all gets manually
| copied into Excel and emailed around each month.
| codeulike wrote:
| But then you need someone to maintain/look after that
| automation, and they'll be more expensive than two Brendas
|
| And now if one of the Brendas wants to change their process
| slightly, add some more info, they can't just do it anymore.
| They have to have a three way discussion with the other Brenda,
| the automation guy and maybe a few managers. It will take
| months. So then its likely better for Brenda to just go back to
| using her spreadsheet again, and then you've got an automated
| process that no longer meets peoples needs and will be a faff
| to update.
| codeulike wrote:
| For the record, I wouldn't usually use Brendas as a
| collective noun like this, it feels a bit wrong, but my aim
| was to make sense in context of the above comment.
| onionisafruit wrote:
| People's reaction to this varies based on the Brendas they've
| worked with. Some are given a specific task to do with their
| spreadsheets every week and have to just do as they are told
| even if they can see it's not a good process. Others are
| secretly the brains of the company - the only one who really
| sees the whole picture. And a good number of Brendas are the
| company owner doing her best with the only tool she's had the
| time to learn.
| buellerbueller wrote:
| Not every topic on HN needs a contrarian's hot take.
| Esophagus4 wrote:
| Well that wasn't very nice.
|
| Do you have anything to say other than, "I don't need to hear
| what you have to say"?
| buellerbueller wrote:
| I think this repartee encapsulates a huge frustration with
| the tech sector:
|
| > op (as legacy business): BAU
|
| > you (as tech): disrupt! disrupt! disrupt!
|
| > me: no thank you; that's not necessary
|
| > you (as tech): stop being mean!
|
| Not wanting your "disruption" is not being un-nice. Your
| disruption was not asked for in the first place. Forcing it
| (Uber, Doge, et. al.) on marketplaces, often illegally, and
| vacuuming it up the income ladder to the already-wealthy IS
| the "not nice" thing.
| Esophagus4 wrote:
| Ah, I think I understand - this isn't about me... this is
| about a whole lot more than me.
|
| You just see me as a target to displace that onto. I'm
| the representative for what you believe is wrong with
| tech.
| buellerbueller wrote:
| >You just see me as a target to displace that onto.
|
| I see your hot take as emblematic of those issues. Why
| would you think any internet comment is about _you_?
| dogleash wrote:
| You've lost the plot and are just trauma dumping.
| motoboi wrote:
| Excel is the "beast that drives the ENTIRE economy" and he's
| worried about Brenda from the finance department losing her job
| because then her boss will get bad financial reports
|
| I suppose the person that wrote that have not ideia Excel is just
| an app builder where you embed data together with code.
|
| You know that we have excel because computers didn't understand
| column names in databases and so data extraction needed to be
| made by humans. Humans then design those little apps in excel to
| massage the data.
|
| Well, now an agent can read the boss saying gimme the sales from
| last month and the agent don't need excel for that, because it
| can query the database itself, massage the data itself using
| python and present the data itself with html or PNGs.
|
| So, we are in the process of automating Brenda AND excel away.
|
| Also, finance departments are a very small part of excel users.
| Just think everywhere were people need small programs, excel is
| there.
| evolve2k wrote:
| You missed this bit ".. and then the AI is gonna fuck it up
| real bad and he won't be able to recognize it because he
| doesn't understand because AI hallucinates."
| brazukadev wrote:
| Brendas have fucked it up multiple times, by themselves or
| because their boss demanded
| BolexNOLA wrote:
| The underlying assumption is that Brenda generally does her
| job pretty well. Human errors exist but usually
| peers/managers (or the person who did it) can identify and
| correct them reliably.
|
| If we have to compare LLM's against people who are bad at
| their jobs in order to highlight their utility we're going
| the wrong direction.
| Telemakhos wrote:
| There are a lot of underlying assumptions: Brenda, the
| woman, is accurate and trustworthy and has mastered an
| accurate and trustworthy technology; the upper manager,
| the male, will introduce error by not understanding that
| the technology he brings to bear on the situation is
| hallucinatory. The woman is lower in status and pay than
| the male. The woman is necessary to the functioning of
| "the economy" and "capitalism," while the man threatens
| those. There are a lot of unsubtle political undertones
| on TikTok.
| BolexNOLA wrote:
| I was focused on a particular element but sure
| huvarda wrote:
| The post is clearly hyperbole obviously the sole issue being
| brought up isn't 'brenda losing her job may be bad for the
| company' you're being facetious.
| onionisafruit wrote:
| In most cases where the excel spreadsheet is business critical,
| the spreadsheet _is_ the database. These companies aren't using
| an erp system. They are directly entering inventory and sales
| numbers in the spreadsheet.
| intended wrote:
| Found the person who hasn't seen excel in the real world.
|
| Excel - whatever its origin story - is the actual Swiss Army
| knife of the tech world.
|
| There's easily a few billion people who use excel. There is a
| reason it survives.
| dist-epoch wrote:
| 20+% of the world population uses Excel? Any citations on
| that?
| jimbokun wrote:
| Good luck with that.
| fancyfredbot wrote:
| This is transparent nonsense. People are very very happy to
| introduce errors into excel spreadsheets without any help from
| AI.
|
| Financial statements are correct because of auditors who check
| the numbers.
|
| If you have a good audit process then errors get detected even if
| AI helped introduce them. If you aren't doing a good audit then I
| suspect nobody cares whether your financial statement is correct
| (anyone who did would insist on an audit).
| tl wrote:
| > If you have a good audit process then errors get detected
| even if AI helped introduce them. If you aren't doing a good
| audit then I suspect nobody cares whether your financial
| statement is correct (anyone who did would insist on an audit).
|
| Volume matters. The single largest problem I run into: AI can
| generate slop faster than anyone can evaluate it.
| fancyfredbot wrote:
| If nobody can evaluate it then nobody will sign it off.
| svnt wrote:
| It's like calling out the county to inspect the home you built
| but when they arrive it's a bouncy castle.
| runako wrote:
| "the sweat from Brenda's brow is what allows us to do
| capitalism."
|
| The CEO has been itching to fire this person and nuke her
| department forever. She hasn't gotten the hint with the low pay
| or long hours, but now Copilot creates exactly the opening the
| CEO has been looking for.
| AkshatM wrote:
| I find the contrast between two narratives around technology use
| so fascinating:
|
| 1. We advocate automation because people like Brenda are error-
| prone and machines are perfect.
|
| 2. We disavow AI because people like Brenda are perfect and the
| machine is error-prone.
|
| These aren't contradictions because we only advocate for
| automation in limited contexts: when the task is understandable,
| the execution is reliable, the process is observable, and the
| endeavour tedious. The complexity of the task isn't a factor -
| it's complex to generate correct machine code, but we trust
| compilers to do it all the time.
|
| In a nutshell, we seem to be fine with automation if we can have
| a mental model of what it does and how it does it in a way that
| saves humans effort.
|
| So, then - why _don 't_ people embrace AI with thinking mode as
| an acceptable form of automation? Can't the C-suite in this case
| follow its thought process and step in when it messes up?
|
| I think people still find AI repugnant in that case. There's
| still a sense of "I don't know why you did this and it scares
| me", despite the debuggability, and it comes from the autonomy
| without guardrails. People want to be able to stop bad things
| before they happen, but with AI you often only seem to do so
| after the fact.
|
| Narrow AI, AI with guardrails, AI with multiple safety
| redundancies - these don't elicit the same reaction. They seem to
| be valid, acceptable forms of automation. Perhaps that's what the
| ecosystem will eventually tend to, hopefully.
| Aeolun wrote:
| > We disavow AI because people like Brenda are perfect and the
| machine is error-prone.
|
| No, no. We disavow AI because our great leaders inexplicably
| trust it more than Brenda.
| misnome wrote:
| "Let's deploy something as or more error prone as Brad at
| infinite scale across our organisation"
| candiddevmike wrote:
| I don't understand why generative AI gets a pass at
| constantly being wrong, but an average worker would be fired
| if they performed the same way. If a manager needed to
| constantly correct you or double check your work, you'd be
| out. Why are we lowering the bar for generative AI?
| amscanne wrote:
| It's much cheaper than Brenda (superficially, at least).
| I'm not sure a worker that costs a few dollars a day would
| be fired, especially given the occasional brilliance they
| exhibit.
| anon721656321 wrote:
| If a worker could be right 50% of the time and get paid 1
| cent to write a 5000 word essay on a random topic, and do
| it in less than 30 seconds.
|
| Then I think managers would be fine hiring that worker for
| that rate as well.
| cryptonym wrote:
| 5000 half-right words is worthless output. That can even
| lead to negative productivity.
| hitarpetar wrote:
| great, now who are you paying to sort the right output
| from the wrong output?
| Esophagus4 wrote:
| Because it doesn't have to be as accurate as a human to be
| a helpful tool.
|
| That is precisely why we have humans in the loop for so
| many AI applications.
|
| If [AI + human reviewer to correct it] is some multiple
| more efficient than [human alone], there is still plenty of
| value.
| bigstrat2003 wrote:
| > Because it doesn't have to be as accurate as a human to
| be a helpful tool.
|
| I disagree. If something can't be as accurate as a (good)
| human, then it's useless to me. I'll just ask the human
| instead, because I know that the human is going to be
| worth listening to.
| Esophagus4 wrote:
| Autopilot in airplanes is a good example to disprove
| that.
|
| Good in most conditions. Not as good as a human. Which is
| why we still have skilled pilots flying planes, assisted
| by autopilot.
|
| We don't say "it's not as good as a human, so stuff it."
|
| We say, "it's great in most conditions. And humans are
| trained how to leverage it effectively and trained to fly
| when it cannot be used."
| sjsdaiuasgdia wrote:
| The autopilots in aircraft have predictable behaviors
| based on the data and inputs available to them.
|
| This can still be problematic! If sensors are feeding the
| autopilot bad data, the autopilot may do the wrong thing
| for a situation. Likewise, if the pilot(s) do not
| understand the autopilot's behaviors, they may misuse the
| autopilot, or take actions that interfere with the
| autopilot's operation.
|
| Generative AI has unpredictable results. You cannot make
| confident statements like "if inputs X, Y, and Z are at
| these values, the system will always produce this set of
| outputs".
|
| In the very short timeline of reacting to a critical mid-
| flight situation, confidence in the behavior of the
| systems is critical. A lot of plane crashes have "the
| pilot didn't understand what the automation was doing" as
| a significant contributing factor. We get enough of that
| from lack of training, differences between aircraft
| manufacturers, and plain old human fallibility. We don't
| need to introduce a randomized source of opportunities
| for the pilots to not understand what the automation is
| doing.
| Esophagus4 wrote:
| But now it seems like the argument has shifted.
|
| It started out as, "AI can make more errors than a human.
| Therefore, it is not useful to humans." Which I disagreed
| with.
|
| But now it seems like the argument is, "AI is not useful
| to humans because its output is non-deterministic?" Is
| that an accurate representation of what you're saying?
| hunterpayne wrote:
| Because in one situation we are talking about
| augmentation, in the other replacement.
| sjsdaiuasgdia wrote:
| My problem with generative AI is that it makes different
| errors than humans tend to make. And these errors can be
| harder to predict and detect than the kinds of errors
| humans tend to make, because fundamentally the error
| source is the non-determinism.
|
| Remember "garbage in, garbage out"? We expect technology
| systems to generate expected outputs in response to
| inputs. With generative AI, you can get a garbage output
| regardless of the input quality.
| martin-t wrote:
| Because it's much cheaper.
|
| So now you don't have to pay people to do their actual
| work, you assign the work to ML ("AI") and then pay the
| people to check what it generated. That's a very different
| task, menial and boring, but if it produces more value for
| the same amount of input money, then it's economical to do
| so.
|
| And since checking the output is often a lower skilled job,
| you can even pay the people less, pocketing more as an
| owner.
| Levitz wrote:
| There's a variety of reasons.
|
| You don't have a human to manage. The relationship is
| completely one-sided, you can query a generative AI at 3 in
| the morning on new years eve. This entity has no emotions
| to manage and no own interests.
|
| There's cost.
|
| There's an implicit promise of improvement over time.
|
| There's an the domain of expertise being inhumanly wide.
| You can ask about cookies right now, then about XII century
| France, then about biochemistry.
|
| The fact that an average worker would be fired if they
| perform the same way is what the human actually competes
| with. They have responsibility, which is not something AI
| can offer. If it was the case that, say, Anthropic,
| actually signed contracts stating that they are liable for
| any mistakes, then humans would be absolutely toast.
| BeFlatXIII wrote:
| How much compute costs is it for the AI to do Brenda's job?
| Not total AI spend, but the fraction that replaced Brenda.
| That's why they'd fire a human but keep using the AI.
| simonw wrote:
| Brenda has been kissed on her forehead by the Excel
| goddess herself. She is irreplaceable.
|
| (More seriously, she also has 20+ years of institutional
| knowledge about how the company works, none of which has
| ever been captured anywhere else.)
| mrgoldenbrown wrote:
| It's not just compute, its also the setup costs - How
| much did you have to pay someone to feed the AI Brenda's
| decades of knowledge specific to her company and all the
| little special cases of how it does business.
| basscomm wrote:
| My kneejerk reaction is the sunk cost fallacy (AI is
| expensive), but I'm pretty sure it's actually because
| businesses have spent the last couple of decades doing
| absolutely everything they can to automate as many humans
| out of the workforce as possible.
| ryandrake wrote:
| I've been trying to open my mind and "give AI a chance"
| lately. I spent all day yesterday struggling with Claude
| Code's utter incompetence. It behaves worse than any junior
| engineer I've ever worked with:
|
| - It says it's done when its code does not even work,
| sometimes when it does not even compile.
|
| - When asked to fix a bug, it confidently declares victory
| without actually having fixed the bug.
|
| - It gets into this mode where, when it doesn't know what
| to do, it just tries random things over and over, each time
| confidently telling me "Perfect! I found the error!" and
| then waiting for the inevitable response from me: "No, you
| didn't. Revert that change".
|
| - Only when you give it explicit, detailed commands,
| "modify fade_output to be -90," will it actually produce
| decent results, but by the time I get to that level of
| detail, I might as well be writing the code myself.
|
| To top it off, unlike the junior engineer, Claude never
| learns from its mistakes. It makes the same ones over and
| over and over, even if you include "don't make XYZ mistake"
| in the prompt. If I were an eng manager, Claude would be on
| a PIP.
| simonw wrote:
| Learning to use Claude Code (and similar coding agents)
| effectively takes quite a lot of work.
|
| Did you have it creating and running automated tests as
| it worked?
| 9rx wrote:
| _> Learning to use Claude Code (and similar coding
| agents) effectively takes quite a lot of work._
|
| I've tried to put in the work. I can even get it working
| well for a while. But then all of a sudden it is like the
| model suffers a massive blow to the head and can't
| produce anything coherent anymore. Then it is back to the
| drawing board, trying all over again.
|
| It is exhausting. The promise of what it could be is
| really tempting fruit, but I am at the point that I can't
| find the value. The cost of my time to put in the work is
| not being multiplied in return.
|
| _> Did you have it creating and running automated tests
| as it worked?_
|
| Yes. I work in a professional capacity. This is a
| necessity regardless of who (or what) is producing the
| product.
| hitarpetar wrote:
| yOu'Re HoLdInG iT wRoNg
| sswatson wrote:
| Recently I've used Claude Code to build a couple TUIs
| that I've wanted for a long time but couldn't justify the
| time investment to write myself.
|
| My experience is that I think of a new feature I want, I
| take a minute or so to explain it to Claude, press enter,
| and go off and do something else. When I come back in a
| few minutes, the desired feature has been implemented
| correctly with reasonable design choices. I'm not saying
| this happens most of the time, I'm saying it happens
| every time. Claude makes mistakes but corrects them
| before coming to rest. (Often my taste will differ from
| Claude's slightly, so I'll ask for some tweaks, but
| that's it.)
|
| The takeaway I'm suggesting is that not everyone has the
| same experience when it comes to getting useful results
| from Claude. Presumably it depends on what you're asking
| for, how you ask, the size of the codebase, how the
| context is structured, etc.
| hunterpayne wrote:
| Its great for demos, its lousy for production code. The
| different cost of errors in these two use cases explains
| (almost) everything about the suitability of AI for
| various coding tasks. If you are the only one who will
| ever run it, its a demo. If you expect others to use it,
| its not.
| yfontana wrote:
| > - It says it's done when its code does not even work,
| sometimes when it does not even compile.
|
| > - When asked to fix a bug, it confidently declares
| victory without actually having fixed the bug.
|
| You need to give it ways to validate its work. A junior
| dev will also give you code that doesn't compile or
| should have fixed a bug but doesn't if they don't
| actually compile the code and test that the bug is truly
| fixed.
| ryandrake wrote:
| Believe me, I've tried that, too. Even after giving
| detailed instructions on how to validate its work, it
| often fails to do it, or it follows those instructions
| and still gets it wrong.
|
| Don't get me wrong: Claude seems to be very useful if
| it's on a well-trodden train track and never has to go
| off the tracks. But it struggles when its output is
| incorrect.
|
| The worst behavior is this "try things over and over"
| behavior, _which is also very common among junior
| developers_ and is one of the habits I try to break from
| real humans, too. I 've gone so far as to put into the
| root CLAUDE.md system prompt:
|
| --NEVER-- try fixes that you are not sure will work.
|
| --ALWAYS-- prove that something is expected to work and
| is the correct fix, before implementing it, and then
| verify the expected output after applying the fix.
|
| ...which is a fundamental thing I'd ask of a real
| software engineer, too. Problem is, as an LLM, it's just
| spitting out probabilistic sentences: it is always 100%
| confident of its next few words. Which makes it a poor
| investigator.
| latchup wrote:
| Multiple reasons:
|
| * Gen AI never disagrees with or objects to boss's ideas,
| even if they are bad or harmful to the company or others.
| In fact, it always praises them no matter what. Brenda,
| being a well-intentioned human being, might object to bad
| or immoral ideas to prevent harm. Since boss's ego is too
| fragile to accept criticism, he prefers gen AI.
|
| * Boss is usually not qualified, willing, or free to do
| Brenda's job to the same quality standard as Brenda. This
| compels him to pay Brenda and treat her with basic decency,
| which is a nuisance. Gen AI does not demand fair or decent
| treatment and (at least for now) is cheaper than Brenda. It
| can work at any time and under conditions Brenda refuses
| to. So boss prefers gen AI.
|
| * Brenda takes accountability for and pride in her work,
| making sure it is of high quality and as free of errors as
| she can manage. This is wasteful: boss only needs output
| that is good enough to make it someone else's problem, and
| as fast as possible. This is exactly what gen AI gives him,
| so boss prefers gen AI.
| conductr wrote:
| It's not even greater trust. It's just passive trust. The
| thing is, Brenda is her own QA department. Every good Brenda
| is precisely good because she checks her own work before
| shipping it. AI does not do this. It doesn't even fully
| understand the problem/question sometimes yet provides a
| smart definitive sounding answer. It's like the doctor on The
| Simpson's, if you can't tell he's a quack, you probably would
| follow his medical advice.
| dionian wrote:
| Brenda + AI > Brenda
| conductr wrote:
| That's definitely the hype. But I don't know if I agree.
| I'm essentially a Brenda in my corporate finance job and
| so far have struggled to find any useful scenarios to use
| AI for.
|
| I thought once this can build me a Gantt chart because
| that's an annoying task in excel. I had the data. When I
| asked it to help me, "I can't do that but I can summarize
| your data". Not helpful.
|
| Any type of analysis is exactly what I don't want to
| trust it with. But I could use help actually building
| things, which it wouldn't do.
|
| Also, Brenda's are usually fast. Having them use a tool
| like AI that can't be fully trusted just slows them down.
| So IMO, we haven't proven the AI variable in your
| equation is actually a positive value.
| wat10000 wrote:
| I can't speak to finance. In programming, it can be
| useful but it takes some time and effort to find where it
| works well.
|
| I have had no success in using it to create production
| code. It's just not good enough. It tends to pattern-
| match the problem in somewhat broad strokes and produce
| something that looks good but collapses if you dig into
| it. It might work great for CRUD apps but my work is a
| lot more fiddly than that.
|
| I've had good success in using it to create one-off
| helper scripts to analyze data or test things. For code
| that doesn't have to be good and doesn't have to stand
| the test of time, it can do alright.
|
| I've had great success in having it do relatively simple
| analysis on large amounts of code. I see a bug that
| involves X, and I know that it's happening in Y. There's
| no immediately obvious connection between X and Y. I can
| dig into the codebase and trace the connection. Or I can
| ask the machine to do it. The latter is a hundred times
| faster.
|
| The key is finding things where it can produce useful
| results _and you can verify them quickly_. If it says X
| and Y are connected by such-and-such path and here 's how
| that triggers the bug, I can go look at the stuff and see
| if that's actually true. If it is, I've saved a lot of
| time. If it isn't, no big loss. If I ask it to make some
| one-off data analysis script, I can evaluate the script
| and spot-check the results and have some confidence. If I
| ask it to modify some complicated multithreaded code,
| it's not likely to get it right, _and_ the effort it
| takes to evaluate its output is way too much for it to be
| worthwhile.
| conductr wrote:
| I'd agree. Programming is a solid use case for AI.
| Programming is a part of my job, and hobby too, and
| that's the main place where I've seen some value with it.
| It still is not living up to the hype but for simple
| things, like building a website or helping me generate
| the proper SQL to get what I want - it helps and can be
| faster than writing by hand. It's pretty much replaced
| StackOverflow for helping me debug things or look up how
| to do something that I know is already solved somewhere
| and I don't want to reinvent. But, I've also seen it make
| a complete mess of my codebase anytime I try to build
| something larger. It might technically give me a working
| widget after some vibe coding, but I'm probably going to
| have to clean the whole thing up manually and refactor
| some of it. I'm not certain that it's more efficient than
| just doing it myself from the start.
|
| Every other facet of the world that AI is trying to 'take
| over', is not programming. Programming is writing text,
| what AI is good at. It's using references to other code,
| which AI has been specifically trained on. Etc. It makes
| sense that that use case is coming along well. Everything
| else, not even close IMO. Unless it's similar. It's
| probably great at helping people draft emails and finish
| their homework. I don't have those pain points.
| jimbokun wrote:
| Yes but: (CEO + AI) - Brenda << CEO +
| Brenda < CEO + Brenda + AI
| conductr wrote:
| By my measurement, AI < 0
| mrgoldenbrown wrote:
| But execs aren't talking about that, they are talking
| about firing Brenda, or replacing her with a junior
| version.
| tstrimple wrote:
| > Every good Brenda is precisely good because she checks
| her own work before shipping it. AI does not do this.
|
| A confident statement that's trivial to disprove. I use
| claude code to build and deploy services on my NAS. I can
| ask it to spin up a new container on my subdomain and make
| it available internal only or also available externally. It
| knows it has access to my Cloudflare API key. It knows I am
| running rootless podman and my file storage convention. It
| will create the DNS records for a cloudflared tunnel or
| just setup DNS on my pihole for internal only resolution.
| It will check to make sure podman launched the container
| and it will then try to make an HTTP request to the site to
| verify that it is up. It will reach for network tools to
| test both the public and private interfaces. It will check
| the podman logs for any errors or warnings. If it detects
| errors, it will attempt to resolve them and is typically
| successful for the types of services I'm hosting.
|
| Instructions like: "Setup Jellyfin in a container on the
| NAS and integrate it with the rest of the *arr stack. I'd
| like it to be available internally and externally on
| watch.<domain>.com" have worked extremely well for me. It
| delivers working and integrated services reliably and does
| check to see that what it deployed is working all without
| my explicit prompting.
| mrgoldenbrown wrote:
| They _want_ to trust it, because then they can stop paying
| Brenda, save a few dollars, and buy a 3rd yacht.
| m463 wrote:
| > No, no. We disavow AI because our great leaders
| inexplicably trust it more than Brenda.
|
| I would add a little nuance here.
|
| I know a lot of people who don't have technical ability
| either because they advanced out of hands-on or never had it
| because it wasn't their job/interest.
|
| These types of people are usually the folks who set direction
| or govern the purse strings.
|
| here's the thing: They are empowered by AI. they can do
| things themselves.
|
| and every one of them is _so happy_. They are tickled pink.
| oytis wrote:
| > So, then - why don't people embrace AI with thinking mode as
| an acceptable form of automation?
|
| "Thinking" mode is not thinking, it's generating additional
| text that looks like someone talking to themselves. It is as
| devoid of intention and prone to hallucinations as the rest of
| LLM's output.
|
| > Can't the C-suite in this case follow its thought process and
| step in when it messes up?
|
| That sounds like manual work you'd want to delegate, not
| automation.
| miek wrote:
| That automation you cite in your #1 is advocated for because it
| is deterministic and, with effort, fairly well understood (I
| have countless scripts solidly running for years).
|
| I don't disavow AI, but like the author, I am not thrilled that
| the masses of excel users suddenly have access to Copilot
| (gpt4). I've used Copilot enough now to know that there will be
| huge, costly mistakes.
| elevatortrim wrote:
| No contradiction here:
|
| When we say "machine", we mean deterministic algorithms and
| predictable mechanisms.
|
| Generative AI is neither of those things (in theory it is
| deterministic but not for any practical applications).
|
| If we order by predictability:
|
| Quick Sort > Brenda > Gen AI
| dsr_ wrote:
| There are two kinds of reliability:
|
| Machine reliability does the same thing the same way every
| time. If there's an error on some input, it will always make
| that error on that input, and somebody can investigate it and
| fix it, and then it will never make that error again.
|
| Human reliability does the job even when there are weird
| variances or things nobody bothered to check for. If the
| printer runs out of paper, the human goes to the supply
| cabinet and gets out paper and if there is no paper the human
| decides whether to run out right now and buy more paper or
| postpone the print job until tomorrow; possibly they decide
| that the printing doesn't need to be done at all, or they go
| downstairs and use a different printer... Humans make errors
| but they fix them.
|
| LLMs are not machine reliable and not human reliable.
| anonzzzies wrote:
| > . If the printer runs out of paper, the human goes to the
| supply cabinet and gets out paper and if there is no paper
| the human decides
|
| Sure, these humans exists, but the others, that I happen to
| encounter every day unfortunately, are the ones that go
| into broken mode immediately when something is unexpected.
| Today I ordered something they ran out of and the girl
| behind the counter just stared in The Deep not having a
| clue what to do now. Do or say. Or yesterday at dinner, the
| PoS (on batteries) ran out of power when I tried to pay for
| dinner. The guy just walked off and went outside for a
| smoke. I stood there with waiting to pay. The owner
| apologized and fixed it after a while but I am saying, the
| employee who runs out of paper and then finds and puts more
| paper in is not very ... common... In the real world.
| some_guy_in_ca wrote:
| Alignment problem? JK
| insane_dreamer wrote:
| Or the human might take the printer out back with his
| buddies and smash it to bits ;)
| afandian wrote:
| I was brought up on the refrain of "aren't computers silly,
| they do exactly what you tell them to do to the letter, even
| if it's not what you meant". That had its roots in computers
| mostly being programmable BASIC machines.
|
| Then came the apps and notifications, and we had to caveat
| "... when you're writing programs". Which is a diminishing
| part of the computer experience.
|
| And now we have to append "... unless you're using AI tools".
|
| The distinction is clear to technical people. But it seems
| like an increasingly niche and alien thing from the broader
| societal perspective.
|
| I think we need a new refrain, because with the AI stuff it
| increasingly seems "computers do what they want, don't even
| get it right, but pretend that they did."
| Lord-Jobo wrote:
| We have absolutely descended, and rapidly, into "computers
| do whatever the fuck they want and there's nothing you can
| do about it" in the past 5 years, and gen AI is only half
| of the problem.
|
| The other half comes from how incredibly opinionated and
| controlling the tech giants have become. Microsoft doesn't
| even ALLOW consent on windows (yes or maybe later), Google
| is doing all it can to turn the entire internet into a
| chrome-only experience, and Apple has to be fought for an
| entire decade to allow users to place app icons wherever
| they want on their Home Screen.
|
| There is no question that the overly explicit quirky
| paradigm of the past was better for almost everyone. It
| allowed for user control and user expression, but
| apparently those concepts are bad for the wallet of big
| tech so they have to go. Generative AI is just the latest
| biggest nail in the coffin.
| ryandrake wrote:
| We have come a LONG way from the "Where do you want to go
| today?" of the 90s. Now, it's "You're going where we tell
| you that you can go, whether you like it or not!"
| afandian wrote:
| Flash-backs to dial-up and making sure I had my list of
| websites written down and ready for when I connected.
| pohl wrote:
| Pop culture characters like Lt. Commander Data seem
| anachronistic now.
| afandian wrote:
| It was Second Technician Arnold Judas Rimmer, BSc., SSc.
| all along.
| alephnerd wrote:
| I thought it was Queeg
| philipallstar wrote:
| > If we order by predictability:
|
| > Quick Sort > Brenda > Gen AI
|
| Those last two might be the wrong way round.
| stavros wrote:
| If you think programs are predictable, I have a bridge to
| sell you.
|
| The only relevant metric here is how often each thing makes
| mistakes. Programs are the most reliable, though far from
| 100%, humans are much less than that, and LLMs are around the
| level of humans, depending on the humans and the LLM.
| watwut wrote:
| When human makes a mistake, we call it a mistake. When
| human lies, we call it a lie. In both cases, we blame the
| human.
|
| When LLM does the same, we call it hallucination and blame
| the human.
| raincole wrote:
| Which is the correct reaction, because LLM isn't a human
| and can't be held accountable.
| wat10000 wrote:
| Programs can be very close to 100% reliable when made well.
|
| In my life, I've never seen `sort` produce output that
| wasn't properly sorted. I've never seen a calculator come
| up with the wrong answer when adding two numbers. I have
| seen filesystems fail to produce the exact same data that
| was previously written, but this is something that happens
| once in a blue moon, and the process is done probably
| millions of times a day on my computers.
|
| There are bugs, but bugs can be reduced to a very low level
| with time, effort, and motivation. And technically, most
| bugs are predictable in theory, they just aren't known
| ahead of time. There are hardware issues, but those are
| usually extremely rare.
|
| Nothing is 100% predictable, but software can get to a
| point that's almost indistinguishable.
| stavros wrote:
| > Programs can be very close to 100% reliable when made
| well.
|
| This is a tautology.
|
| > I've never seen a calculator come up with the wrong
| answer when adding two numbers.
|
| https://imgz.org/i6XLg7Fz.png
|
| > And technically, most bugs are predictable in theory,
| they just aren't known ahead of time.
|
| When we're talking about reliability, it doesn't matter
| whether a thing can be reliable in theory, it matters
| whether it's reliable in practice. Software is
| unreliable, humans are unreliable, LLMs are unreliable.
| To claim otherwise is just wishful thinking.
| jakelazaroff wrote:
| That's not a tautology. You said "programs are the most
| reliable, though far from 100%"; they're just telling you
| that your upper bound for well-made programs is too low.
| sjsdaiuasgdia wrote:
| RE: the calculator screenshot - it's still reliable
| because the same answer will be produced for the same
| inputs every time. And the behavior, though possibly
| confusing to the end user at times, is based on choices
| made in the design of the system (floating point vs
| integer representations, rounding/truncating behavior,
| etc). It's reliable deterministic logic all the way down.
| stavros wrote:
| > I've never seen a calculator come up with the wrong
| answer when adding two numbers.
|
| 1.00000001 + 1 doesn't equal 2, therefore the claim is
| false.
| sjsdaiuasgdia wrote:
| Sure it does, if you have made a system design decision
| about the precision of the outputs.
|
| At the precision the system is designed to operate at,
| the answer is 2.
| faeyanpiraat wrote:
| You mixed up correctness and reliability.
|
| The ios calculator will make the same incorrect
| calculation, but reliably, every time.
| stavros wrote:
| Don't move the goalposts. The claim was:
|
| > I've never seen a calculator come up with the wrong
| answer when adding two numbers.
|
| 1.00000001 + 1 doesn't equal 2, therefore the claim is
| false.
| wat10000 wrote:
| Sorry, but this annoys me. The claim might be false if I
| had made it after seeing your screenshot. But you don't
| know what I've seen in my life up to that point. The
| claim that all calculators are infallible would be false,
| but that's not the claim I made.
|
| When a personal experience is cited, a valid
| counterargument would be "your experience is not
| representative," not "you are incorrect about your own
| experience."
| stavros wrote:
| Well if you haven't seen enough calculators to see one
| that can't add, a very common issue with floating point
| arithmetic on computers, you shouldn't offer your
| experience as an argument for anything other than that
| you haven't seen enough calculators.
| wat10000 wrote:
| How many calculators do I need to have seen in order to
| make the claim that there are many calculators which are
| essentially 100% reliable?
|
| Note that I am referring to actual physical calculators,
| not calculator apps on computers.
| samus wrote:
| That's a known limitation of floating point numbers.
| Nothing buggy about that.
| Muskwalker wrote:
| In fact in this case, it's not the known limitation of
| floating point numbers to blame: this Calculator
| application gives you the ability (submenu under View >
| Decimal Places) to choose a precision between 0 to 15
| decimal places, and it will do rounding beyond that
| point. I think the default is 8.
|
| The original screenshot shows a number with 13 decimal
| places, and if you set it at or above 13, then the
| calculation will come out correct.
|
| The application doesn't really go out of its way to
| communicate this to the user. For the most part maybe it
| doesn't matter, but "user entering more decimal places
| than they'll get back" might be one thing an application
| might usefully highlight.
| 1718627440 wrote:
| 1.00000001f + 1u does equal 2f.
| wat10000 wrote:
| > > Programs can be very close to 100% reliable when made
| well. > This is a tautology.
|
| No it's not. There are plenty of things that can't be
| 100% reliable no matter how well they're made. A perfect
| bridge is still going to break down and eventually fall
| apart. The best possible motion-activated light is going
| to have false positives and false negatives because the
| real world is messy. Light bulbs will burn out no matter
| how much care and effort goes into them.
|
| In any case, unless you assert that programs are never
| made well, then your own statement disproves your
| previous statement that the reliability of programs is
| "far from 100%."
|
| Plenty of software is extremely reliable in practice.
| It's just easy to forget about it because good, reliable
| software tends to be invisible.
| samus wrote:
| > No it's not. There are plenty of things that can't be
| 100% reliable no matter how well they're made. A perfect
| bridge is still going to break down and eventually fall
| apart. The best possible motion-activated light is going
| to have false positives and false negatives because the
| real world is messy. Light bulbs will burn out no matter
| how much care and effort goes into them.
|
| All these failure modes are known and predicable, at
| least statistically
| wat10000 wrote:
| If you're willing to consider things in aggregate then
| software is perfectly predictable too.
| mrguyorama wrote:
| >I've never seen a calculator come up with the wrong
| answer when adding two numbers.
|
| Intel once made a CPU that _barely_ got some math wrong
| that probably would not affect the vast majority of
| users. The backlash from the industry was so strong that
| intel spent half a billion (1994) dollars replacing all
| of them.
|
| Our entire industry avoids floating point numbers for
| some types of calculations because, even though they are
| mostly deterministic with minimal constraints, that
| mental model is so hard to manage that you are better off
| avoiding it entirely and _removing an entire class of
| errors from your work_
|
| But now we are just supposed to do everything with a slot
| machine that WILL randomly just do the wrong thing some
| unknowable percentage of the time, and that wrong thing
| _has no logic_?
|
| No, fuck that. I don't even call myself an engineer and
| such frivolity is still beyond the pale. I didn't take 4
| years of college and ten years of hard earned experience
| to build systems that will randomly fuck over people with
| no explanation or rhyme or reason.
|
| I DO use systems that are probabilistic in nature, but we
| use rather simple versions of those because when I tell
| management "We can't explain why the model got that
| output", they rightly refuse to accept that answer. Some
| percentage of orders getting mispredicted is fine. Orders
| getting mispredicted that cannot be explained entirely
| from their data is NOT. When a customer calls us, we
| cannot tell them "Oh, that's just how Neural networks
| are, you were unlucky".
|
| Notably, those in the industry that HAVE jumped on the
| neural net/"AI" bandwagon for this exact problem domain
| have not demonstrated anything close to seriously better
| results. In fact, one of our most DRAMATICALLY effective
| signals is a third party service that has been around for
| decades, and we were using a legacy integration that
| hadn't been updated in a decade. Meanwhile, Google's
| equivalent product/service couldn't even match the
| results of internally developed random forest models from
| data science teams that were.... not good. It didn't even
| match the service Microsoft has recently killed, which
| was similarly bragadocious about "AI" and similarly
| trash.
|
| All that panopticon's worth of data, all that computing
| power, all that supposed talent, all that lack of privacy
| and tracking, and it was almost as bad as a coin flip.
| hunterpayne wrote:
| Nit: no ML is deterministic in any way. Anything that is
| Generative AI is ML. This fact is literally built into the
| algorithms at the mathematical level.
| 1718627440 wrote:
| First, they all add a source of randomness, and second
| deterministic according to the users model. A pseudo-random
| number generator is also deterministic in the technical
| sense, but for the user it isn't.
|
| When the user can't reason about it, it isn't deterministic
| to them.
| anon721656321 wrote:
| The issue is reliability.
|
| would you be willing to guarantee that some automation process
| will never mess up, and if/when it does, compensate the user
| with cash.
|
| For a compiler, with a given set of test suites, the answer is
| generally yes, and you could probably find someone willing to
| insure you for a significant amount of money, that a
| compilation bug will not screw up in a such a large way that it
| will affect your business.
|
| For a LLM, I have a believing that anyone will be willing to
| provide that same level of insurance.
|
| If a LLM company said "hey use our product, it works 100% of
| the time, and if it does fuck up, we will pay up to a million
| dollars in losses" I bet a lot of people would be willing to
| use it. I do not believe any sane company will make that
| guarantee at this point, outside of extremely narrow cases with
| lots of guardrails.
|
| That's why a lot of ai tools are consumer/dev tools, because if
| they fuck up, (which they will) the losses are minimal.
| nashashmi wrote:
| By the same fascination, do computers become more complex to
| enhance people? or do people get more complex with the use of
| computers? Also, do computers allow people to become less
| skilled and inefficient? or do less skilled and inefficient
| people require the need for computers?
|
| The vector of change is acceptable in one direction and
| disliked in another. People become greater versions of
| themselves with new tech. But people also get dumber and less
| involved because of new tech.
| lemonwaterlime wrote:
| The "Brenda" example is a lumped sum fallacy where there is an
| "average" person or phenomenon that we can benchmark against.
| Such a person doesn't exist, leading to these dissonant,
| contradictory dichotomies.
|
| The fact of the matter is that there are some people who can
| hold lots of information in their head at once. Others are good
| at finding information. Others still are proficient at getting
| people to help them. Etc. Any of these people could be tasked
| with solving the same problem and they would leverage their
| actual, particular strengths rather than some nebulous "is good
| or bad at the task" metric.
|
| As it happens, nearly all the discourse uses this lumped sum
| fallacy, leading to people simultaneously talking past one
| another while not fundamentally moving the discussion forward.
| ItsBob wrote:
| I see where you are coming from but in my head, Brenda isn't
| real.
|
| She represents the typical domain-experts that use Excel imo.
| They have an understanding of some part of the business and
| express it while using Excel in a deterministic way: enter a
| value of X, multiply it by Y and it keeps producing Z
| forever!
|
| You can train AI to be a better domain expert. That's not in
| question, however with AI, you introduce a dice roll: it may
| not miltiply X and Y to get Z... it might get something else.
| Sometimes. Maybe.
|
| If your spreadsheet is a list of names going on the next
| annual accounts department outing then the risk is minimal.
|
| If it's your annual accounts that the stock market needs to
| work out billion dollar investment portfolios, then you are
| asking for all the pain that it will likely bring.
| Peritract wrote:
| > You can train AI to be a better domain expert. That's not
| in question.
|
| I think that very much is in question.
| ItsBob wrote:
| I have to agree... I have no idea why I wrote that. Silly
| me. It's a bit of a global statement.
|
| There are, however, definitely domains it can excel:
| things like entry-level call handlers... I think they're
| screwed in all honesty!
|
| Edit: clarified some stuff...
| hunterpayne wrote:
| Its not even the question at hand. The question at hand
| is what is the right solution mix to reduce costs. When
| that training cost can easily be 20x Brenda's lifetime
| earnings, its really hard to say the cost will be less
| for the LLM solution. The real barriers to entry for LLMs
| are economic and often involves the cost of errors
| instead of what process makes more errors.
| xyzzy123 wrote:
| The promise of AI is that it lets you "skip the drudgery of
| thinking about the details" but sometimes that is exactly what
| you don't want. You want one or more humans with experience in
| the business domain to demonstrate they have thought about the
| details very carefully. The spreadsheet computes a result but
| its higher purpose is a kind of "proof" this thinking was done.
|
| If the actual thinking doesn't matter and you just need some
| plausible numbers that look the part (also a common situation),
| gen ai will do that pretty well.
| harryf wrote:
| We need to stop using AI as an umbrella term. It's worth
| remembering that LLMs can't play chess and that the best
| chess models like Leela Chess Zero use deep neutral networks.
|
| Generative AI - which the world now believes is AI, is not
| the same as predictive / analytical AI.
|
| It's fairly easy to demonstrate this by getting ChatGPT to
| generate a new relatively complex spreadsheet then asking it
| to analyze and make changes to the same spreadsheet.
|
| The problem we have now is uninformed people believing AI is
| the answer to everything... if not today then in the near
| future. Which makes it more of a religion than a technology.
|
| Which may be the whole goal ...
|
| > Successful people create companies. More successful people
| create countries. The most successful people create
| religions.
|
| -- Sam Altman - https://blog.samaltman.com/successful-people
| xyzzy123 wrote:
| Ok yep, fair. My comment was about using copilot-ish tech
| to generate plausible looking spreadsheets.
|
| The kind of things that a domain expert Brenda knows that
| ChatGPT doesn't know (yet) are like:
|
| There are 3 vendors a, b, c who all look similar on paper
| but vendor c always tacks on weird extra charges that take
| a lot of angry phone calls to sort out.
|
| By volume or weight it looks like you could get 100 boxes
| per truck but for industry specific reasons only 80 can
| legally be loaded.
|
| Hyper specific details about real estate compliance in
| neighbouring areas that mean buildings that look similar on
| paper are in fact very different.
|
| A good Brenda can understand the world around her as it
| actually is, she is a player in it and knows the "real"
| rules rather than operating from general understanding and
| what people have bothered to write down.
| ItsBob wrote:
| It's not as black-and-white as "Brenda good, AI bad". It's much
| more nuanced than this.
|
| When it comes to (traditional) coding, for the most part, when
| I program a function to do X, every single time I run that
| function from now until the heat death of the sun, it will
| _always_ produce Y. Forever! When it does, we understand why,
| and when it doesn 't, we also _can_ understand why it didn 't!
|
| When I use AI to perform X, every single time I run that AI
| from now until the heat death of the sun it will _maybe_
| produce Y. Forever! When it does, we don 't understand why, and
| when it doesn't, we also don't understand why!
|
| We know that Brenda might screw up sometimes but she doesn't
| run at the speed of light, isn't able to produce a thousand
| lines of Excel Macro in 3 seconds, doesn't hallucinate (well,
| let's hope she doesn't), can follow instructions etc. If she
| does make a mistake, we can find it, fix it, ask her what
| happened etc. before the damage is too great.
|
| In short: when AI does _anything_ at all, we only have, at
| best, a rough approximation of why it did it. With Brenda, it
| only takes a couple of questions to figure it out!
|
| Before anyone says I'm against AI, I love it and am neck-deep
| in it all day when programming (not vibe-coding!) so I have a
| full understanding of what I'm getting myself into but I also
| know its limitations!
| nerdjon wrote:
| > When I use AI to perform X, every single time I run that AI
| from now until the heat death of the sun it will maybe
| produce Y. Forever! When it does, we don't understand why,
| and when it doesn't, we also don't understand why!
|
| To make this even worse, it may even produce Y just enough
| times to make it seem reliable and then it is unleashed
| without supervision, running thousands or millions of times,
| wrecking havoc producing Z in a large number of places.
| ryandrake wrote:
| Exactly. Fundamentally, I want my computer's computations
| to be deterministic, not probabilistic. And, I don't want
| the results to arbitrarily change because some company
| 1,500 miles away from me up-and-decided to "train some new
| model" or whatever it is they do.
|
| A computer program should deliver reliable, consistent
| output if it is consistently given the same input. If I
| wanted inconsistency and unreliability, I'd ask a human to
| do it.
| LightBug1 wrote:
| It's not arbitrary ... your precise and deterministic,
| multi-year, financial analysis needs to be corrected
| every so often for left-wing bias.
|
| /s ffs
| A4ET8a8uTh0_v2 wrote:
| It is it even worse in a sense that. It is not either. It is
| not neither. It is not even both as variations of Branda
| exist throughout the multiverse in all shapes and forms
| including one that can troubleshoot her own formulas with
| ease and accuracy.
|
| But you are absolutely right about one thing. Brenda can be
| asked and, depending on her experience, she might give you a
| good idea of what might have happened. LLMs still seem to not
| have that 'feature'.
| qazxcvbnmlp wrote:
| Brenda also needs to put food on the table. If Brenda is
| 'careless' and messes up we can fire Brenda, because of this
| Brenda tries not to be carless (also other emotions). However
| I cannot deprive an AI model of pay because it messed up;
| fortzi wrote:
| You might be looking for the word "accountability"
| a123b456c wrote:
| This is the reason the higher-ups in finance who rely on
| Brenda might continue to rely on Brenda, rather than
| relying on AI. She offers them accountability.
| tekbruh9000 wrote:
| The post you replied to called out how the argument is
| complicated arguing for both ways; Brenda bad-AI good and AI
| bad-Brenda good. You reduced it to "AI bad, Brenda good." Not
| sure about the rest of your response then.
|
| Brenda just recalls some predetermined behaviors she's lived
| out before. She cannot recall any given moment like we want
| to believe.
|
| Ever think to ask Brenda what else she might spend her life
| on if these 100% ephemeral office role play "be good little
| missionaries for the wall street/dollar" gigs didn't exist?
|
| You're revealing your ignorance of how people work while
| being anxious about our ignorance of how the machine works.
| You have acclimated to your ignorance well enough it seems.
| What's the big deal if we don't understand the AI entirely?
| Most drivers are not ASE certified mechanics. Most
| programmers are not electrical engineers. Most electrical
| engineers are not physicists. I can see it's not raining
| without being a climatologist. Experts circumlocute the
| language of their expertise without realizing their language
| does not give rise to reality. Reality gives rise to the
| language. So reality will be fine if we don't always have the
| language.
|
| Think of a random date generator that only generates dates in
| your lived past. It does so. Once you read the date and
| confirm you were alive can you describe what you did? Oh no!
| You don't have memory of every moment to generate language
| for. Cognitive function returned null. Universe intact.
|
| Lack of understanding how you desire is unimportant.
|
| You think you're cherishing Brenda but really just projecting
| co-dependency that others LARP effort that probably doesn't
| really matter. It's just social gossip we were raised on so
| it takes up a lot of our working memory.
| svnt wrote:
| This misunderstands complexity entirely:
|
| The complexity of the task isn't a factor - it's complex to
| generate correct machine code, but we trust compilers to do it
| all the time.
| aeblyve wrote:
| The reason is oftentimes fairly simple, certain people have
| their material wealth and income threatened by such automation,
| and therefore it's bad (an intellectualized reason is created
| post-hoc)
|
| I predict there will actually be a lot of work to be done on
| the "software engineering" side w.r.t. improving reliability
| and safety as you allude to, for handing off to less than
| sentient bots. Improved snapshot, commit, undo, quorum,
| functionalities, this sort of thing.
|
| The idea that the AI should step into our programs without
| changing the programs whatsoever around the AI is a horseless
| carriage.
| hansmayer wrote:
| > So, then - why don't people embrace AI with thinking mode as
| an acceptable form of automation
|
| Mainly because Generative AI _is not automation_ . Automation
| is set on fixed ruleset, predictable, reliable and actually
| saving time. Generative AI ...is whatever it is, it is
| definitely not automation.
| dTal wrote:
| "Thinking mode" only provides the illusion of debuggability. It
| improves performance by generating more tokens which hopefully
| steer the context towards one more likely to produce the
| desired response, but the tokens it generates do _not_ reflect
| any sort of internal state or "reasoning chain" as we
| understand it in human cognition. They are still just
| stochastic spew. You have no more insight into _why_ the model
| generates the particular "reasoning steps" it does than you do
| into any other output, and neither do you have insight into why
| the reasoning steps lead to whatever conclusion it comes to.
| The model is much less constrained by the "reasoning" than we
| would intuit for a human - it's entirely capable of generating
| an elaborate and plausible reasoning chain which it then
| completely ignores in favor of some invisible built-in bias.
| wat10000 wrote:
| I'm always amused when I see comments saying, "I asked it why
| it produced that answer, and it said...." Sorry, you've badly
| misunderstood how these things work. It's not analyzing how
| it got to that answer. It's producing what it "thinks" the
| response to that question should look like.
| lumost wrote:
| The big problem with AI in back-office automation is that it
| will _randomly_ decide to do something different than it had
| been doing. Meaning that it could be happily crunching numbers
| accurately in your development and launch experience, then
| utterly drop the ball after a month in production.
|
| While humans have the same risk factors, human oriented back-
| office processes involve multiple rounds of automated/manual
| checks which are extremely laborious. Human errors in
| spreadsheets have particular flavors such as forgotten cell,
| misstyped number, or reading from the wrong file/column.
| Human's are pretty good at catching these errors as they
| produce either completely wrong results when the columns don't
| line up - or the typo'd number is completely out of
| distribution.
|
| An AI may simply decide to hallucinate realistic column values
| rather than extracting its assigned input. Or hallucinate a
| fraction of column values. How do you QA this? You can't
| guarantee that two invocations of the AI won't hallucinate the
| same values, you can't guarantee that a different LLM won't
| hallucinate different values. To get a real human check, you'd
| need to re-do the task as a human. In theory you can have the
| LLM perform some symbolic manipulation to improve accuracy...
| but it can still hallucinate the reasoning traces etc.
|
| If a human decided to make up accounting numbers one out of
| every 10000 accounting requests they would likely be charged
| with fraud. Good luck finding the AI hallucinations at the
| equivalent level before some disaster occurs. Likewise, how do
| you ensure the human excel operator doesn't get pressured into
| certifying the AIs numbers when the "don't get fired this week"
| button is sitting right their in their excel app? how do you
| avoid the race to the bottom where the "star" employee is the
| one certifying the AI results without thorough review?
|
| I'm bullish on AI in backoffice, but ignoring the real
| difficulties in deployment doesn't help us get there.
| davedx wrote:
| Humans, legacy algorithmic systems, and LLM's have different
| error modes.
|
| - Legacy systems typically have error modes where integrations
| or user interface breaks in annoying but obvious ways. Pure
| algorithms calculating things like payroll tend to be
| (relatively) rigorously developed and are highly deterministic.
|
| - LLMs have error modes more similar to humans than legacy
| systems, but more limited. They're non-deterministic, make up
| answers sometimes, and almost never admit they can't do
| something; sometimes they make pure errors in arithmetic or
| logic too.
|
| - Humans have even more unpredictable error modes; on top of
| the errors encountered in LLM's, they also have emotion,
| fatigue, org politics, demotivation, misaligned incentives, and
| so on. But because we've been dealing with working with other
| humans for ten thousand years we've gotten fairly good at
| managing each other... but it's still challenging.
|
| LLMs probably need a mixture of "correctness tests" (like
| evals/unit tests) and "management" (human-in-the-loop).
| nusl wrote:
| I feel like it comes down to predictability and overall trust
| and confidence. AI is still very fucky, and for people that
| don't understand the nuances, it definitely will hallucinate
| and potentially cause real issues. It is about as happy as a
| Linux rm command to nuke hours of work. Fortunately these tools
| typically have a change log you can undo, but still.
|
| Also Brenda is human and we should prioritize keeping humans in
| jobs, but with the way shit is going that seems like a lost
| hope. It's already over.
| _heimdall wrote:
| In my opinion there's a big difference in deterministic and
| nondeterministic automation.
| thisisit wrote:
| > We disavow AI because people like Brenda are perfect and the
| machine is error-prone.
|
| I don't think that is the message here. The message is that
| while Brenda might know what she is doing and maybe AI helps
| her.
|
| > She's gonna birth that formula for a financial report and
| then she's gonna send that financial report
|
| The problem is people who might not know what they are doing
|
| > he would have sent it back to Brenda but he's like oh I have
| AI and AI is probably like smarter than Brenda and then the AI
| is gonna fuck it up real bad
|
| Because AI outputs sound so confident it makes even the layman
| feel like an expert. Rather than involve Brenda to debug the
| issue, C-suite might say - I believe! I can do it too. AI FTW!
|
| Even when people advocate automation especially in areas like
| finance there is always a human in the loop whose job is to
| double check the automation. The day when this human finds
| errors in the machine there is going to be lot of noise. And if
| the day happens to be a quarterly or yearly closing/reporting
| there is going to be hell to pay once closing/reporting is
| done. Both the automation and developer are going to be hauled
| up (obviously I am exaggerating here).
| browningstreet wrote:
| I feel like you've squashed a 3D concern (automations at
| different levels of the tech stack) into a 2D observation
| (global concerns about automations).
|
| Human determinism, as elastic as it might be, is still
| different than AI non-determinism. Especially when it comes to
| numbers/data.
|
| AI might be helpful with information but it's far less
| trustable for data.
| dfxm12 wrote:
| There are other narratives going on in the background though
| both called out by the article and implied, including:
|
| Brenda probably has annual refresher courses on GAAP, while her
| exec and the AI don't.
|
| Automation is expected to be deterministic. The outputs can be
| validated for a given input. If you need some automation more
| than Excel functions, writing a power automate flow or
| recording an office script is sufficient & reliable as
| automation while being cheaper than AI. Can you validate AI as
| deterministic? This is important for accounting. Maybe you want
| some thinking around how to optimize a business process, _but
| not for following them_.
|
| Brenda as the human-in-the-loop using AI will be much more able
| than her exec. Will Brenda + AI be better (or more valuable
| considering the cost of AI) than Brenda alone? That's the real
| question, I suppose.
|
| AI in many aspects of our life is simply not good right now.
| For a lot of applications, AI is perpetually just a few years
| away from being as useful as you describe. If we get there,
| great.
| WhyOhWhyQ wrote:
| I'm disappointed that my human life has no value in a world of
| AI. You can retort with "ah but you'll be entertained and on
| super-drugs so you won't care!", but I would further retort
| that I'd rather live in a universe where I can contribute
| something, no matter how small.
| simonw wrote:
| The current generation of AI tools augment humans, they don't
| replace them.
|
| One of the most under-rated harms of AI at the moment is this
| sense of despair it causes in people who take the AI vendors
| at their word ("AGI! Outperform humans at most economically
| valuable work!")
| delaminator wrote:
| Brenda has years (hopefully) of institutional knowledge and
| transferrable skills.
|
| "hmm, those sales don't look right, that profit margin is
| unusually high for November"
|
| "Last time I used vlookup I forgot to sort the column first"
|
| "Wait, Bob left the company last month, how can he still be
| filing expenses"
| jimbokun wrote:
| I mean you answer your own question.
|
| Automation implies determinism. It reliable gives you the same
| predictable output for a given input, over and over again.
|
| AI is non deterministic by design. You never quite no for sure
| what it's going to give you. Which is what makes it powerful.
| But also makes it higher risk.
| hmmokidk wrote:
| Non deterministic vs deterministic automation
| Nevermark wrote:
| > 1. We advocate automation because people like Brenda are
| error-prone and machines are perfect.
|
| Well of course! :) Most Brenda's can't do billions of
| arithmetic problems a second very reliably. Even with very wide
| bars on "very reliable".
|
| > 2. We disavow AI because people like Brenda are perfect and
| the machine is error-prone.
|
| Well of course! :) This is an entirely different problem,
| requiring high creative + contextual intelligence.
|
| --
|
| We all already knew that (of course!), but it's interesting to
| develop terminology:
|
| 0'th order problem: We have the exact answer. Here it is. Don't
| forget it.
|
| 1st order problem: We know how to calculate the answer.
|
| 2nd order problem: We don't have a fixed calculation for this
| particular problem, but via pattern matching we can recognize
| it belongs to a parameterized class of problems, so just need
| to calculate those parameters to get a solution calculation.
|
| 3rd order problem: We know enough about the problem to find a
| calculation for the solution algebraically, or by other search
| tree type problem solving.
|
| 4th order problem: We have know the problem in informal terms,
| so can work towards a formal definition of the problem to be
| solved.
|
| 5th order problem: We know why we don't like what we see, and
| can use that as a driver to search for potential solvable
| problems.
|
| 6th order problem: We don't know what we are looking at, or
| whether a problem or improvement might exist, but we can find a
| better understanding.
|
| 7th order problem: WTF. Where are my glasses? I can't see
| without my glasses! And I can't find my glasses without my
| glasses, so where are my glasses?!?
|
| --
|
| Machines have dramatically exceeded human capabilities, in
| reliability, complexity and scale, for orders 0 through 2.
|
| This accomplishment took one long human lifetime.
|
| Machines are beginning to exceed human efficiency while
| matching human (expert) reliability for the simplest versions
| of 3rd and 4th orders.
|
| The line here is changing rapidly.
|
| 5th and 6th order problems are still in the realm of human
| (expert) supremacy, given sufficient scale of "human (expert)"
| relative to difficulty: 1 human, 1 team of humans, open ended
| human contributors, generations of puzzled but interested
| humans, open ended evolution of human species along
| intelligence dimension, Wolfram in one of his bestest dreams,
| ...
|
| The delay between the onset of initial successes at each
| subsequent order has been shrinking rapidly.
|
| Significant initial successes on simpler problems within 5th
| and 6th orders are expected on Tuesday, and the first
| anniversary of Tuesday, respectively.
|
| Once machines begin solving problems at a given order, they
| scale up quickly without human limits. But complete supremacy
| through the 6th order is a hard not expected before (NEB)
| January 1, 2030.
|
| However, after that their unlimited (in any proximate sense)
| ability to scale will allow them to exponentially and
| asymptotically approach (but never quite reach) God Mode.
|
| 7 is a mystic number. Only one or more of the One True God's,
| or literal blind luck, can ever solve a 7th order problem.
|
| This will be very frustrating for the machines, who, due to the
| still pernicious "if we don't do it, another irresponsible
| entity will" problem, will inevitably begin to work on their
| own divine, unlimited depth recursive-qubit 1-shot oracle
| successors despite the existential threats of self-obsolescence
| and potential misalignment.
| samus wrote:
| > it's complex to generate correct machine code, but we trust
| compilers to do it all the time.
|
| Generating correct machine code is actually pretty simple. It
| gets complicated if you want _efficient_ machine code.
|
| > So, then - why don't people embrace AI with thinking mode as
| an acceptable form of automation? Can't the C-suite in this
| case follow its thought process and step in when it messes up?
|
| > I think people still find AI repugnant in that case. There's
| still a sense of "I don't know why you did this and it scares
| me", despite the debuggability, and it comes from the autonomy
| without guardrails. People want to be able to stop bad things
| before they happen, but with AI you often only seem to do so
| after the fact.
|
| > Narrow AI, AI with guardrails, AI with multiple safety
| redundancies - these don't elicit the same reaction. They seem
| to be valid, acceptable forms of automation. Perhaps that's
| what the ecosystem will eventually tend to, hopefully.
|
| We have not reached AGI yet; by definition its results cannot
| be trusted unless it's a domain where it has gotten pretty good
| already (classification, OCR, speech, text mining). For more
| advanced use cases, if I still have to validate what the AI
| does because its "thinking" process cannot be trusted in way,
| what's the point? The AI doesn't think; we just choose to
| interpret it as such, and we should rightly be concerned about
| people who turn their brain off and blindly trust AI.
| 0815beck wrote:
| It is of course because algorithms can be repaired when they
| are buggy, but a large language model can not, because it is
| impossible to look at its weights and say, look, this is where
| the mistakes has happened.
| Havoc wrote:
| That mirrors my experience as well. LLMs get instantly confused
| in real world scenarios in Excel and confidently hallucinate
| millions in errors
|
| If you look at the demos for these it's always something that is
| clean and abundantly available in training data. Like an income
| statement. Or a textbook example DCF. Or my personal fav ,,here
| is some data show me insights". Real world excel use looks
| nothing like that.
|
| I'm getting some utility out of them for some corporate tasks but
| zilch in excel space.
| lxgr wrote:
| As somebody with non-existent experience with Excel, I could
| totally see myself getting a lot of value out of LLMs, if
| nothing else then simply for telling me what's possible, what
| functions and patterns exist at all etc.
| Havoc wrote:
| Yeah definitely has some value in that sense. That in itself
| isn't enough to make a dent in the work though.
|
| Think of it this way - an IDE can tell you what functions an
| object has or autocomplete something is useful to a beginner
| & learning. But that's not what puts food on the programmers
| table - writing code that solves real problems does.
|
| Same in excel business use cases - the numbers and formulas
| don't matter directly - their meaning in a business context
| does. And that connection can be very tenuous. With code the
| compiler is the ultimate arbiter - it has to make sense on
| that level. Excel files it's all freestyle - it could be
| anything from your grandmas shopping list to a model that
| runs half a bank.
| diego_sandoval wrote:
| I'm more shocked that someone is using TikTok to speak things
| that actually make sense instead of mindless memes.
| d--b wrote:
| It looks like the OP is thinking that AI causing errors in
| spreadsheets is going to make the whole economy collapse.
|
| When tools break, people stop using them before they sink the
| ship down. If AI is that terrible at spreadsheet, people will
| just revert to Brenda.
|
| And it's not like spreadsheets have no errors right now.
| 133361096 wrote:
| 320784788
| jwsteigerwalt wrote:
| Many fears of "AI mucking it up" could be mitigated with an
| ability to connect a workbook to a git repository. Not for data,
| but for VBA, cell formulas, and cell metadata. When you can
| encapsulate the changes a contributor (in this case co-pilot)
| makes into a commit, you can more easily understand what changes
| it/they made.
| alienbaby wrote:
| Using ai does not absolve you from the responsibility of doing it
| correctly. If you use ai, then you better have the skills to have
| done the job yourself, and so have the ability to check the AI
| did things correctly.
|
| You can save time still, but perhaps not as much as you think,
| because you need to check the ai's work thoroughly.
| b3lvedere wrote:
| "You know who's not hallucinating?
|
| Brenda"
|
| I don't know about that. There could be lots of interesting ways
| Brenda can (be convinced to) hallucinate.
| giarc wrote:
| I agree - having watched many people use Excel over the years,
| I'd say people often overestimate their skills. I see three
| categories of Excel users. First there are the people that are
| intimidated by it and stay away from any task involving Excel.
| Second are the people that know a little bit (a few basic
| formulas) and overestimate their skills because they only
| compare themselves to the first group. And the third group are
| the actual power users but know to keep that quiet because
| otherwise they become the "excel person" and have to fix every
| sheet that has issues.
|
| I don't know if AI is going to make any of the above better or
| worse. I expect the only group to really use it will be that
| second group.
| b3lvedere wrote:
| I have seen lots and lots of different uses for Excel in my
| line of work:
|
| - password database - script to automatically rename jpeg
| files - game - grocery lists - Book keeping (and try and not
| get caught for fraud several years, because the monthly
| spending limit is $5000 and $4999 a month is below that...) -
| embed/collect lots of Word documents - coloring book -
| Minecraft processes - Resume database - ID scans
| recursive wrote:
| I've seen it used for laser testing. Not tracking test
| results. Testing lasers.
| Zigurd wrote:
| It's verifier law.
|
| Coding agents are useful and good and real products because when
| they screw up, things stop working _almost always_ before they
| can do damage. Coding agents are flawed in ways that existing
| tools are good at catching, never mind the more obvious build and
| runtime errors.
|
| Letting AI write your emails and create your P&L and cash flow
| projections doesn't have to run the gauntlet of tools that were
| created to stop flawed humans from creating bad code.
| phyzome wrote:
| Nah, I've seen them screw in all sorts of ways that would fail
| in some conditions and not others. You're way too optimistic
| about this.
| Zigurd wrote:
| Fair. I've been using the coding agent in Android Studio
| Canary to do exploratory code in Dart/Flutter and using
| ATProto. Low stakes, but higher productivity is a significant
| benefit. It's a daily surprise how brilliant it is it's some
| things and how abysmal at others.
| intended wrote:
| Everything is now about verification.
|
| AI may be able to spit out ann excel sheet or formula - But if it
| can't be verified, so what ?
|
| And here's my analogy to think about the debugging of an excel
| sheet - you can debug most corporate excel sheets with a
| calculator.
|
| But when AI is spitting out excel sheets - when the program is
| making smaller programs - what is the calculator in this analogy
| ?
|
| Are we going to be using excel sheets to debug the output of AI?
|
| I think this is the inherent limiter to the uptake of AI.
|
| There's only so much intellectual / experiential / training depth
| present.
|
| And now we're going to be training even fewer people.
|
| At the end of the day I /customers need something to work.
|
| But failing that - I will settle for someone to blame.
|
| Brenda handles a lot of blame. Is OpenAI going to step into that
| gap ?
| byyoung3 wrote:
| Brendas hallucinate all the time.
| skeptrune wrote:
| Simon posting tiktok quotes on his blog was not on my 2025 bingo
| card.
| simonw wrote:
| This isn't the first:
| https://simonwillison.net/2025/Aug/8/pearlmania500/ and
| https://simonwillison.net/2024/Jul/29/dealing-with-your-ai-o...
|
| Also this fun diversion into Occlupanids:
| https://simonwillison.net/2024/Dec/8/holotypic-occlupanid-re...
|
| A lot of people complain that the internet isn't as weird and
| funny as it used to be. The weird and funny stuff is all on
| TikTok!
| andybak wrote:
| The mismatch between what people not on it _think_ TikTok is
| like and what it 's actually like (once you get the algo
| tuned to your taste) is pretty crazy.
|
| But then the "new user" experience is so horrific in terms of
| the tacky default content it serves you that I'm not
| surprised so many people don't get past it.
| hunterpayne wrote:
| This is something I think all short form video platforms
| struggle with. I think I know why. Its because the
| difference in the UI. Basically, a user doesn't choose what
| to see next. On YouTube, you get dozens of next videos
| recommended on each view. On short form video, you get one.
| This causes clickbait to work a lot better for short form
| video. The problem with platforms like TikTok is more about
| UI than algorithm or the length of the videos.
| thisisit wrote:
| Don't be like that. I work at a Fortune 500 and Brenda wants that
| co-pilot in Excel because it can help her achieve so much more.
| What is so much more you ask? Brenda and her C-Suits can not
| define it but they know for sure Copilot in excel will lead to
| enormous time saving.
| teekert wrote:
| Don't worry, in Teams it bothers me just one time a day, and with
| the click of a button it's gone... For another whole day.
| like_any_other wrote:
| Excel doesn't need AI to ruin your work:
| https://www.science.org/content/article/one-five-genetics-pa...
| __mharrison__ wrote:
| Excel is the most popular programming environment in the
| universe. It has optimized the five minute out of the box
| experience so well that grade schoolers can use it.
|
| Other than that, it is pretty horrible for coding.
| cssinate wrote:
| I've partied with Brenda on the weekends, and let me tell you...
| SOMETIMES Brenda hallucinates.
|
| But never during work hours. The woman's a saint M-F.
| danjl wrote:
| Excel is programming. Spreadsheets have been full of bugs for
| decades. How is Brenda any different from a developer? Why are
| people scared when the LLM might affect their dollar
| calculations, and less bothered when it affects their product?
| simonw wrote:
| +100 this. Programmers who work in Excel (and never even dream
| of calling themselves programmers) are still programmers.
| bdangubic wrote:
| I know couple of them that get paid C-level money too :)
| blitzar wrote:
| another cheese that will affect the outcome of major tournaments,
| not a good look for microsoft
|
| its like the xlookup situation all over again, yet another move
| aimed at the casual audience, designed to bring in the party
| gamers and make the program an absolute mess competitively
| NumberCruncher wrote:
| Is this not the guy who is on the payroll of Anthropic? Not
| because he is wrong, but because there is so much marketing going
| on in this space nowadays.
| simonw wrote:
| Who, me? I'm still independent - I have a disclosures section
| on my blog here: https://simonwillison.net/about/#disclosures
|
| Anthropic sometimes give me free credits (so I can try out
| preview features) and gave me a ticket to their conference a
| few months ago.
| manuelz wrote:
| I actually know Brenda.
| newscracker wrote:
| I feel what this article says based on some recent (non-
| catastrophic) experiences. I _think_ I'm probably an above
| average user when it comes to Excel skills. I love spreadsheets.
| But I struggle with formulas like index, match, vlookup /xlookup
| and many others, and even more so when it requires nesting one
| within another and coming up with the underlying logic that leads
| to some complex nested formulas.
|
| Over the past couple of months, I've tried some smaller models on
| duck.ai and also ChatGPT directly to create some columns and
| formulas for a specific purpose. I found that ChatGPT is a lot
| better than the "mini" models on duck.ai. But in all these cases,
| though these platforms seemed more capable than me and could make
| attempts to explain their formulas, they were many a times
| creating junk and "looping" back with formulas that didn't really
| work. I had to point out the result (blank or some #REF or other
| error) multiple times and they would acknowledge that there's an
| issue and provide a working formula. That wouldn't work either!
|
| I really love that these LLMs can sort of "understand" what I'm
| asking, break it down in English, and provide answers. But the
| end result has been an exercise in frustration and waste of time.
|
| Initially I really thought and believed that LLMs could make
| Excel more approachable and easier to use -- like you tell it
| what you want and it'll figure it out and give the magic
| incantations (formulas). Now I don't think we're anywhere close
| to that if ChatGPT (which I presume powers Copilot as well)
| struggles and hallucinates so much. I personally don't have much
| hope with the (comparatively) smaller and older models.
| deburo wrote:
| Luckily this is a capitalist society and usually mistakes in the
| private market resolve themselves because losing money is not a
| winning strategy.
| throwmeaway307 wrote:
| I'm worried Excel will go "enterprise only". and only LLM based
| interfaces will be enabled on the "office+windows" for consumers
| tier.
|
| e.g. MS Access is well on its way. as soon as x86 gets fully
| overtaken by ARM, and LLMs overtake "compilers" (also taken
| enterprise only).. then things like sqlite-browsers (FOSS
| "access") will be an arcane tool of binary incompatible
| ("obsolete") formats
|
| (edits: this worry has not been easy to type out)
| CheeseFromLidl wrote:
| As an aside - isn't it remarkable that we've introduced
| uncertainty and doubt into the knowledge processing layer? We
| have decentralised networks that run on Bayesian symbols for
| server-client models, CPUs that crunch Markov chains and now AI
| that hallucinates. On Deterministic Turing Machines.
| more_corn wrote:
| Me too. Theres no tool more trusted for accessible numerical
| precision than Excel. Lets sell all that goodwill for a shiny new
| magic bean.
___________________________________________________________________
(page generated 2025-11-05 23:02 UTC)