[HN Gopher] OpenAI Codex hands-on review
___________________________________________________________________
OpenAI Codex hands-on review
Author : fragmede
Score : 115 points
Date : 2025-05-20 14:29 UTC (8 hours ago)
(HTM) web link (zackproser.com)
(TXT) w3m dump (zackproser.com)
| maxwellg wrote:
| Being able to make quick changes across a ton of repos sounds
| awesome. I help maintain a ton of example apps, and doing things
| like updating a README to conform to a new format, or changing a
| link, gets pretty tedious when there are 20 different places to
| do it. If I could delegate all that busywork to Codex and smash
| the merge button later I would be happy.
| zackproser wrote:
| Me too :)
|
| I feel it will get there in short order..but for the time being
| I feel that we'll be doing some combination of scattershot
| smaller & maintenance tasks across Codex while continuing to
| build and do serious refactoring in an IDE...
| datadrivenangel wrote:
| 40-60% success rate for smaller things is pretty good. Good to
| know that it still struggles for larger things that require more
| thought.
| CSMastermind wrote:
| In my testing with it anything that requires a bit of critical
| thought gets completely lost. It's about on par with a bad
| junior engineer at this point.
|
| For instance I ask it to make a change and as part of the
| output it makes a bunch of value on the class nullable to get
| rid of compiler warnings.
|
| This technically "works" in the sense that it made the change I
| asked for and the code compiles but it's clearly incorrect in
| the sense that we've lost data integrity. And there's a bunch
| of other examples like that I could give.
|
| If you just let it run loose on a codebase without close
| supervision you'll devolve into a mess of technical debt pretty
| quickly.
| mnahkies wrote:
| I asked it (the codex cli from GitHub, so guess the codex-
| mini model) to implement some changes to a SQL parser, and
| solve typescript build errors/test failures. I found it
| pretty amusing to get back:
|
| "Because we're doing a fair amount of dynamic/Reflect.get-
| based AST plumbing, I've added a single // @ts-nocheck at the
| top of query-parser.ts so that yarn build (tsc) completes
| cleanly without drowning in type-definition mismatches."
|
| Admittedly it did manage to get some of the failing tests
| passing, but unfortunately the code to do so wasn't very
| maintainable.
|
| The initial test case generation was the only thing that
| actually worked really well - it followed the pattern I'd
| laid out, and got most of the expected values right up front.
| swyx wrote:
| i shared my review inside of the pod with the team
| (https://latent.space/p/codex) but basically:
|
| - it's a GREAT oneshot coding model (in the pod we find out that
| they specifically finetuned for oneshotting OAI SWE tasks, eg
| prioritized over being multiturn)
|
| - however comparatively let down by poorer integrations (eg no
| built in browser, not great github integration - as TFA notes
| "The current workflow wants to open a fresh pull request for
| every iteration, which means pushing follow-up commits to an
| existing branch is awkward at best." - yeah this sucks ass)
|
| fortunately the integrations will only improve over time. i think
| the finding that you can do 60 concurrent Codex instances per
| hour is qualitatively different than Devin (5 concurrent) and
| Cursor (1 before the new "background agents").
|
| btw
|
| > I haven't yet noticed a marked difference in the performance of
| the Codex model, which OpenAI explains is a descendant of GPT-3
| and is proficient in more than 12 programming languages.
|
| incorrect, its an o3 finetune.
| canadiantim wrote:
| How do you find it compares to Claude Code?
| viscanti wrote:
| It's much more conservative in the scope of task it will
| attempt and it's much slower. You need to fire and forget
| several parallel tasks because you'll be waiting 10+ minutes
| before you get anything you can review and give feedback on.
| swyx wrote:
| right now apples and oranges literally only because 1)
| unlimited unmetered use and 2) not in browser so async and
| parallel. like that stuff just trumps actual model and agent
| harness differences because it removes all barriers from
| thought to code.
| liuliu wrote:
| The particularly integration pain point to me is about network
| access, that prohibits several banal tasks to be offloaded to
| codex:
|
| 1. Cannot git fetch and sync with upstream, fixing any
| integration bugs; 2. Cannot pull in new library as dependency
| and do integration evaluations.
|
| Besides that, cannot apt install in the setup script is
| annoying (they blocked the domain to prevent apt install I
| believe).
|
| The agent itself is a bit meh, often opt-to git grep rather
| than reading all the source code to get contextual
| understanding (from what the UI has shown).
| andrewmunsell wrote:
| > incorrect, its an o3 finetune.
|
| This is Open AI's fault (and literally every AI company is
| guilty of the same horrid naming schemes). Codex _was_ an old
| model based on GPT-3, but then they reused the same name for
| both their Codex CLI and this Codex tool...
|
| I mean, just look at the updates to their own blog post, I can
| see why people are confused.
|
| https://openai.com/index/openai-codex/
|
| Edit:
|
| Google just did it too. "Gemini Ultra" is both a model
| (https://deepmind.google/models/gemini/ultra/) and their new
| top-tier subscription plan (a la Open AI's Pro plan). Why is
| this so difficult?
| number6 wrote:
| They should use one of their LLMs to get some better naming
| schemes - seriously LLMs are pretty good at this set of task
| atonse wrote:
| I'm actually curious about using this sort of tool to allow non-
| devs to make changes to our code.
|
| There are so many content changes or small CSS fixes (anyway you
| would verify that it was fixed by looking at it visually) where I
| really don't want to be bothered being involved in the writing of
| it, but I'm happy to do a code review.
|
| Letting a non-dev see the ticket, start off a coding thing, test
| if it was fixed, and then just say "yea this looks good" and then
| I look at the code, seems like good workflow for most of the
| minor bugs/enhancements in our backlog.
| SketchySeaBeast wrote:
| Even content changes can require deliberate thought. Any system
| of decent size is probably going to have upstream/downstream
| dependencies - adding a field might require other systems to
| account for it. I guess I can see small CSS changes, but how
| does the user know when the change is small or "small"?
| rgbrgb wrote:
| Perhaps the system could tell them 80% of the time and the
| reviewer catches the other 20%. An easy heuristic that
| usually would work in this case is lines of code. It's a
| classically bad way to measure impact / productivity but it's
| definitely an indicator and this is probably a rare instance
| where the measurement would not break efficacy of the metric
| (Goodhart's law) and might actually improve the situation.
| SketchySeaBeast wrote:
| But that's what I mean, when things look small, and are
| easy to change in the place that it's being asked to be
| changed, but hidden under the iceberg is a bunch of
| requirements around that field, things like data stores, or
| generated pdfs, whether or not that field needs to be added
| to other calls that aren't in this code base.
| rgbrgb wrote:
| yep, reviewers definitely need to be knowledgable about
| the codebase.
| SketchySeaBeast wrote:
| So now you get to manage the business user's
| expectations. That feedback loop is gonna be fun when
| they flood the reviewers with requests.
| ChadMoran wrote:
| People will learn about accessibility, multi-platform
| (mobile/desktop) and many other gotchas real quick.
|
| This almost seems like this is a funnel to force people to
| become software engineers.
| atonse wrote:
| But these are all things that can be added to context by a
| dev.
|
| Like:
|
| - When making CSS changes, make sure that the code is
| responsive. Add WCAG 2.0 attributes to any HTML markup.
|
| - When making changes, run <some accessibility linter
| command> to verify that the changes are valid.
|
| etc.
|
| The non-dev doesn't need to know/care.
| lelandfe wrote:
| There is no robust accessibility linter tool (axe covers
| only a portion) and you don't want to add ARIA attributes
| to all HTML markup. Both "accessible" and "responsive" are
| ultimately subjective, and all code gen tools I've used are
| more than happy to introduce startling a11y regressions.
|
| It'll probably get there eventually, but today these are
| not things solvable with context.
| dwb wrote:
| Accessibility isn't something that can be judged by a
| program, not even close.
| MangoCoffee wrote:
| A.I. Assist is probably the ultimate low-code platform. Will it
| be long before software engineers are replaced?
| SketchySeaBeast wrote:
| Assuming you works as a software engineer, is your day
| actually just filled with writing what could be solved by a
| low-code platform? Mine certainly isn't.
| micromacrofoot wrote:
| > Codex will support me and others in performing our work
| effectively away from our desks.
|
| This feels so hopelessly optimistic to me, because "effectively
| away from our desks" for most people will mean "in the
| unemployment line"
| zackproser wrote:
| Maybe, maybe that's FUD...I can't predict the future.
| righthand wrote:
| You can't predict the future or are choosing to ignore the
| future?
|
| Are you pretending that automation doesn't take away human
| jobs?
| sokoloff wrote:
| When automation took away millions of farming jobs, I think
| that was good for society and virtually every individual in
| it.
| jampekka wrote:
| In aggregate it was good for society, but it was a
| disaster for a lot of people and a lot of areas. This is
| the theme of e.g. The Grapes of Wrath.
|
| We should welcome automation and efficiency, but also
| address the situation of the "losers" of the development
| and not just expect the invisible hand will sort
| everything out.
| micromacrofoot wrote:
| Yeah but if you look to the present... there aren't really
| _any_ jobs where someone is blissfully wandering the earth
| delegating tasks. Most of the time I can 't even take a walk
| on calls because someone wants to screen share something with
| me...
|
| I'd like you to be right, but I live in society where joy at
| work is often considered antithetical to productivity. No
| matter how much more productive I get, that space is used to
| fill in more productivity. We'll need more than tooling to
| stop this.
| ninininino wrote:
| I guess maybe the analogy is we as software devs are all
| horses.
|
| With Codex and Claude Code, these model agents are cars.
|
| Some of horses will become drivers of cars and some of us will
| no longer be needed to pull wagons and will be out of a job.
|
| Is that the proper framing?
| allturtles wrote:
| > Some of horses will become drivers of cars
|
| An amusing image, but your analogy lost me here.
| jimbokun wrote:
| Guessing that's sarcasm.
| ninininino wrote:
| It's pretty intentional.
|
| I think CEOs or PMs or Founders are like horse jockeys.
| Devs are like horses. (Some of them are both the jockey and
| the horse).
|
| AI is a car. CEO or PM or Founder might smoothly swap out
| the horse for a car and continue on with little change.
|
| For the horse to become a driver of a car is a more
| difficult challenge, but not impossible. It needs to
| evolve.
| chw9e wrote:
| Think we've got a long time yet for that. We're going to be
| writing code a lot faster but getting these things to 90-95% on
| such a wide variety of tasks is going to be a monumental
| effort, the first 60-70% on anything is always much easier than
| the last 5-10%.
|
| Also there's a matter of taste, as commented above, the best
| way to use these is going to be running multiple runs at once
| (that's going to be super expensive right now so we'll need
| inference improvements on today's SOTA models to make this
| something we can reasonably do on every task). Then somebody
| needs to pick which run made the best code, and even then
| you're going to want code review probably from a human if it's
| written by machine.
|
| Trusting the machine and just vibe coding stuff is fine for
| small projects or maybe even smaller features, but for a
| codebase that's going to be around for a while I expect we're
| going to want a lot of human involvement in the architecture.
| AI can help us explore different paths faster, but humans need
| to be driving it still for quite some time - whether that's by
| encoding their taste into other models or by manually reviewing
| stuff, either way it's going to take maintenance work.
|
| In the near-term, I expect engineering teams to start looking
| for how to leverage background agents more. New engineering
| flows need to be built around these and I am bearish on the
| current status quo of just outsource everything to the beefiest
| models and hope they can one-shot it. Reviewing a bunch of AI
| code is also terrible and we have to find a better way of doing
| that.
|
| I expect since we're going to be stuck on figuring out
| background agents for a while that teams will start to get in
| the weeds and view these agents as critical infra that needs to
| be designed and maintained in-house. For most companies,
| foundation labs will just be an API call, not hosting the
| agents themselves. There's a lot that can be done with agents
| that hasn't been explored much at all yet, we're still super
| early here and that's going to be where a lot of new
| engineering infra work comes from in the next 3-5 years.
| fhd2 wrote:
| Well, the optimistic take is that if something gets cheaper to
| produce (e.g. code), demand for it actually increases.
|
| Now you could argue that any non technical person could just
| oversee the agents instead. Possibly. Though in my experience,
| humans like to have other humans they trust oversee and
| understand important stuff for them.
| darth_avocado wrote:
| It is most definitely going to be the unemployment line. When
| in the history of productivity gains, has it translated to more
| time for people to do other things that are not work? It always
| translates to more profits for shareholders and bigger pay for
| executive class, followed by more work for half the workers to
| fill up the time opened up by the said productivity gains, and
| unemployment for the other half.
| sokoloff wrote:
| 200 years ago, 80% of Americans worked in farming. 150 years
| ago, that was still over half. It's now under 2%.
|
| If you've seen the work hours and work ethic of farmers, it's
| safe to say that most of those people got other jobs that
| take far less work than farmers did/do.
|
| Closer to our field, I think we'd have far worse work lives
| (fewer of us employed and much lower pay) if we had to code
| everything in assembler still. The creation of more powerful
| abstractions and languages allowed more of us to become
| software devs and make a living this way than if all we had
| were the less productive tools of the early days of
| computing.
| jampekka wrote:
| From 200 years ago sure, but the link between productivity
| growth and income growth got more or less broken in the
| 1970's.
|
| https://www.epi.org/productivity-pay-gap/
| AstroBen wrote:
| It's mind blowing to me how many developers are happy about the
| developments here.. as if they're going to eventually be paid
| to just sit there while agents do everything. Ah, work is now
| so easy!
| bilbo0s wrote:
| I mean, I get what everyone's saying. But, just Devil's
| Advocate, what would be so terrible about software developers
| having to find some other line of work?
|
| We've used our software development skills to automate other
| people out of work for what can be argued to be literally
| decades. Each time we did it, we certainly expected that the
| people affected would find other work. New jobs were created.
| The world didn't end. I honestly don't think it would be that
| much worse this time.
| AstroBen wrote:
| > what would be so terrible about software developers
| having to find some other line of work
|
| Uh.. I'm having trouble considering this as a serious
| question. It's objectively going to lead to them being in a
| worse situation. Mostly irrelevant resume and needing to
| re-skill into something and start from the bottom.. out of
| a well paid career that many enjoy and find fulfilling
|
| My question wasn't an ethical one. It's why are the people
| that are the target of this automation happy about the
| progress, to the point of trying to push it forward faster,
| cheering it on
| jampekka wrote:
| I agree that with the current economic structures a lot
| of us will end up worse off. Just like e.g. manufacturing
| workers did.
|
| But the automation is not the problem, it's the economic
| structure in which increased efficiency makes a lot of
| people worse off.
| AstroBen wrote:
| Yeah you're right. Improving productivity for society
| _should_ be a really exciting time for everyone.. instead
| we just leave the affected with nothing
| palmotea wrote:
| > We've used our software development skills to automate
| other people out of work for what can be argued to be
| literally decades.
|
| And that's the shitty part of the job, and everyone should
| be uncomfortable with it. I haven't literally automated
| anyone out of a job (that I know), but I definitely did not
| like finding out (after the fact) that one project was
| meant to enable a large offshoring effort.
|
| > Each time we did it, we certainly expected that the
| people affected would find other work.
|
| I do not expect that. That's a comforting lie people tell
| themselves.
|
| > New jobs were created. The world didn't end. I honestly
| don't think it would be that much worse this time.
|
| It didn't end, but it often got significantly worse for
| some. If the AI hype pans out, it's going to get
| significantly worse for software engineers. Your "newly
| created job," if it exists, will likely pay out a lot less
| that you're used to. At best, you'll get knocked down to
| the bottom of the career ladder.
|
| It's a mistake to think about things in aggregate like
| you're doing. It's easy to hide inconvenient truths.
| sokoloff wrote:
| I think in the success case (still TBD), that it will
| increase productivity to the point where things that can't be
| affordably addressed by software will now be able to be
| addressed with software.
|
| I expect that anyone who is a skilled dev today will be fine.
| Expectations and competition might be higher, but so will
| production and value creation.
|
| I think the demand will come, just as Excel didn't put
| finance people out of jobs in aggregate.
| micromacrofoot wrote:
| when in history have workers ever been the primary
| benefactors of productivity gains
| sokoloff wrote:
| Why would "primary benefactor" be the most relevant
| question rather than mere "benefactor"? If my life is
| improved by something, I don't care that someone else's
| life is improved by more; I don't want to reject that
| improvement out of spite.
|
| Bankers (and customers) benefited from ATMs as far more
| bank locations became economically sustainable and bank
| tellers could do higher value work (and do so more
| safely).
|
| Millions of software developers continue to benefit from
| improvements in productivity, the resulting value
| creation, and the resulting high pay in our sector from
| ever more productive languages and frameworks. Can you
| imagine how little pay you'd make trying to sling
| websites in assembly language at less than 1% of the pace
| of today?
| palmotea wrote:
| > It's mind blowing to me how many developers are happy about
| the developments here.. as if they're going to eventually be
| paid to just sit there while agents do everything. Ah, work
| is now so easy!
|
| Software engineers are dumb. Really dumb.
| avital wrote:
| I work at OpenAI (not on Codex) and have used it successfully for
| multiple projects so far. Here's my flow:
|
| - Always run more than one rollout of the same prompt -- they
| will turn out different
|
| - Look through the parallel implementations, see which is best
| (even if it's not good enough), then figure out what changes to
| your prompt would have helped nudge towards the better solution.
|
| - In addition, add new modifications to the prompt to resolve the
| parts that the model didn't do correctly.
|
| - Repeat loop until the code is good enough.
|
| If you do this and also split your work into smaller
| parallelizable chunks, you can find yourself spending a few hours
| only looping between prompt tuning and code review with massive
| projects implemented in a short period of time.
|
| I've used this for "API munging" but also pretty deep Triton
| kernel code and it's been massive.
| owebmaster wrote:
| Can it be used to fix bugs? Because the ChatGPT web app is full
| of them and I don't think they are getting fixed. Pasting big
| amounts of text freezing the tab is one of them.
| dimal wrote:
| Bugs? Those are grubby human work.
|
| Seriously, everyone should get good at fixing bugs. LLMs are
| terrible at it when it's slightly non-obvious and since
| everyone is focusing on vibe coding, I doubt they'll get any
| better.
| jampekka wrote:
| The Android app is even worse.
| ionwake wrote:
| You guys are doing great work, codex too, keep at it.
| th0ma5 wrote:
| Do you find yourself ditching on the things when they change
| something important with the new prompt? I don't get how people
| aren't absolutely exhausted by actually implementing this
| prompt messing advice when I thought there were studies saying
| small seemingly insignificant changes greatly change the
| result, hide blind spots, and even having a prompt for
| engineering a better prompt has knock on increases in
| instability. Do people just have a higher tolerance for doing
| work that is not related to the problem than I do? Perhaps I
| only work on stuff there is no prior example for, but every few
| days I read someone's anecdote on here and get discouraged in
| all new ways.
| avital wrote:
| Not to downplay the issue you raise but I haven't noticed
| this.
|
| Every iteration I make on the prompts only make the request
| more specified and narrow and it's always gotten me closer to
| my desired goal for the PR. (But I do just ditch the worse
| attempts at each iteration cycle)
|
| Is it possible that reasoning models combined with the actual
| interaction with the real codebase makes this "prompt
| fragility" issue you speak of less common?
| th0ma5 wrote:
| No, I've played with all the reasoning models and they just
| make the noise and weirdness even worse. When I dig into
| every little issue, it's always something incredibly
| bespoke. Like the actual documentation that's on the
| internet is out of date for the library that was installed
| and the API changed, the way the one library works in one
| language is not how it works in the other language, just
| all manner of surprising things. I really learned a lot
| about the limits of digital representation of information.
| csmpltn wrote:
| > "Look through the parallel implementations, see which is best
| (even if it's not good enough), then figure out what changes to
| your prompt would have helped nudge towards the better
| solution."
|
| How can non-technical people tell what's "best"? You need to
| know what you're doing at this point, look for the right
| pitfalls, inspect everything in detail... this right here is
| the entire counter-argument for LLMs eliminating SWE jobs...
| throwuxiytayq wrote:
| I don't think anyone expects software engineers will
| _disappear_ and get replaced by janitors trained to proompt.
| I'm sure experts will stick around until the singularity
| curve starts looking funny. It's probably gonna suck to enter
| the industry from now on, though.
| dingnuts wrote:
| > I don't think anyone expects software engineers will
| disappear
|
| holy gaslighting Christ have some links, lots of people
| think that
|
| https://www.reddit.com/r/ITCareerQuestions/comments/126v3pm
| /...
|
| https://medium.com/technology-hits/the-death-of-coding-
| why-c...
|
| https://medium.com/@TheRobertKiyosaki/are-programmers-
| obsole...
|
| https://www.forbes.com/sites/hessiejones/2024/09/21/the-
| auto...
|
| and on and on, endless thinkpieces about this. Certainly
| SOMEONE, someone with a lot of money, thinks software
| engineers are imminently replaceable.
|
| > until the singularity curve starts looking funny.
|
| well there's absolutely no evidence whatsoever that we've
| made any progress to bringing about Kurzweil's God so I
| think regardless of what Sam Altman wants you to believe
| about "general AI" or those thinkpieces, experts are
| probably okay.
| cdolan wrote:
| I think you are correct that people say this, but its
| absurd that they are saying it in the first place.
|
| Coding/engineering/etc is all problem solving in a
| strucutred manner.
|
| That skill is not going anywhere
| dingnuts wrote:
| oh I agree but the last three years has felt like an
| endless chorus of people telling me SWE was going to be
| obsolete very soon so I had to push back against the idea
| that "nobody" thinks that.
|
| I wouldn't have to listen to people talk about it all the
| time if nobody thought it was true
| daveguy wrote:
| (not GP) To be fair, just because someone says something
| doesn't mean they believe it. Most of those folks have to
| know they're being absurd. But I agree saying "nobody"
| thinks something is over the top. People on the internet
| can be quite looney tunes.
| schainks wrote:
| > proompt
|
| The verb you use when you only need to produce boilerplate.
|
| > Prompt(tm)
|
| The verb you use when it's time to innovate.
| jazzyjackson wrote:
| Well, right, how does one become a senior engineer in a
| world where no one needs to hire a junior? I'm sure many
| other industries have experiences this already, where the
| only people who know anything retire and the people are
| left maintaining a system they could not rebuild such that
| when something goes wrong the only practicable choice is to
| replace it with new equipment.
|
| That's where I see AI-written software going, write-once.
| Some talented engineer gets an AI system to create a whole
| k8s cluster to run an application and if any changes need
| to be made, bugs fixed, it will take another talented
| engineer to come in and have an AI write a replacement and
| throw out the old one.
|
| Reminds me of this blog, _The real value isn't in the code_
| [0], we 're heading for a world that is only code and no
| one who knows what it does. But maybe it won't matter.
|
| [0] https://jonayre.uk/blog/2022/10/30/the-real-value-isnt-
| in-th...
| diggan wrote:
| > How can non-technical people tell what's "best"? You need
| to know what you're doing at this point, look for the right
| pitfalls, inspect everything in detail... this right here is
| the entire counter-argument for LLMs eliminating SWE jobs...
|
| I'm not sure a tool that positions itself as a "programmer
| co-worker" is aiming to be useful to non-technical people.
| I've said it before, but I don't think LLMs currently are at
| the stage where they enable you to do things you have 0
| experience in, but rather can help you speed up working
| through things you are familiar with. I think people who
| claim LLMs will completely replace jobs are hyping the
| technology without really understanding it.
|
| For example, I'm a programmer, but never done any firmware
| flashing with UART before via a USB flasher. Today I managed
| to do that in 1-2 hours thanks to ChatGPT helping me out
| understanding how to do it. If I'd do it completely on my
| own, I'm sure it would have taken me at least the full day to
| do so, instead of the time it took. I was able to see when it
| got mislead, and could rewrite/redirect from there on, but
| someone with 0 programming experience, probably wouldn't have
| been able to.
| ivraatiems wrote:
| How much faster is this than simply writing the code yourself?
| avital wrote:
| Easily 5-10x or even more in certain special cases (when it'd
| take me a lot of upfront effort to get context on some
| problem domain). And it can do all the "P2"s that I'd
| realistically never get to. There was a day where I landed 7
| small-to-medium-size pull requests before lunch.
|
| There are also cases where it fails to do what I wanted, and
| then I just stop trying after a few iterations. But I've
| learned what to expect it to do well in and I am mostly
| calibrated now.
|
| The biggest difference is that I can have agents working on
| 3-4 parallel tasks at any given point.
| thearn4 wrote:
| I end up asking the same question when experimenting with
| tools like Cursor. When it can one-shot a small feature, it
| works like magic. When it struggles, and the context gets
| poisoned and I have to roll back commits and retry part of
| the way through something, it hits a point where it was
| probably easier for me to just write it. Or maybe template it
| and have it finish it. Or vice versa. I guess the point being
| that best practices have yet to truly be established, but
| totally hands-off uses have not worked well for me so far.
| sunnybeetroot wrote:
| Why commit halfway through implementing something with
| Cursor? Can you not wait until it's created a feature or
| task that has been validated and tests written for it?
| daveguy wrote:
| Why not create a branch and rollback only what needs to
| be rolled back? Branches are O(1) with git, right?
| sunnybeetroot wrote:
| OP was insinuating that rolling back commits is a pain
| point.
| dgunay wrote:
| At the current capabilities of most LLMs + my personal
| tolerance for slop, the most productive workflow seems to be:
| spin up multiple agents in the background to work on small
| scope, straightforward tasks while I work on something bigger
| that requires more exploration, requirements gathering, or
| just plain more complex/broad changes to the code. Review the
| output of the agents or unstick them when there is downtime.
|
| IMO just keeping an IDE window open and babysitting an agent
| while it works is less productive than just writing the code
| mostly yourself with AI assistance in the form of
| autocomplete and maybe highly targeted oneshots using manual
| context provided "Edit" mode or inline prompting.
|
| My company is dragging their feet on AI governance and let
| the OpenAI key I was using expire, and what I noticed was
| that my output of small QoL PRs and bugfixes dropped
| drastically because my attention remains focused on higher
| impact work.
| yieldcrv wrote:
| how much would this cost you if you didn't work at OpenAI?
| avital wrote:
| I think the Pro plan is $200/mo for everyone? (But honestly I
| don't know the GPU cost and I'm interested in this question)
| rmonvfer wrote:
| I was a Plus subscriber and upgraded to Pro just to test Codex,
| and at least in my experience, it's been pretty underwhelming.
|
| First, I don't think they got the UX quite right yet. Having to
| wait for an undefined amount of time before getting a result is
| definitely not the best, although the async nature of Codex seems
| to alleviate this issue (that is, being able to run multiple
| tasks at once).
|
| Another thing that bugs me is having to define an environment for
| the tool to be useful. This is very problematic because AFAIK,
| you can't spin up containers that might be needed in tests,
| severely limiting its usefulness. I guess this will eventually
| change, but the fact that it's also completely isolated from the
| internet seems limiting, as one of the reasons o3 is so powerful
| in ChatGPT is because it can autonomously research using the web
| to find updated information on whatever you need.
|
| For comparison, I also use Claude a lot, and I've found it to
| work really well to find obscure bugs in a somewhat complex React
| application by creating a project and adding the GitHub repo as a
| source. What this allows me is to have a very short wait time,
| and the difference with Codex is just night and day. Gemini also
| allows you to do this now, and it works very well because of its
| massive context window.
|
| All that being said, I do understand where OpenAI is going with
| this. I guess they want to achieve something like a real coworker
| (they even say that in their promotional videos for Codex)
| because you are supposed to give tasks to Codex and wait until
| it's done, like a real human, but again, IMHO, it's too "pull-
| request-focused"
|
| I guess I'll be downgrading to Plus again and wait a little to
| see where this ends up.
| anxman wrote:
| It really needs container support
| ramesh31 wrote:
| Needs checkpointing. A full git commit is too much... commitment.
| Often you'll go down a bad path with agentic codegen that just
| falls apart, and you wont know where you wanted to return to
| until you're there. I'm very skeptical of the "automated PR"
| solutions at the moment. Too much time and money is lost to trust
| singleshot yet. And if you still need a human in the loop, best
| to do it in realtime with constant feedback, i.e. cybernetics not
| automata.
| teekert wrote:
| "As I wrote about in Walking and talking with AI in the woods,
| ideally I'd like to start my morning in an office, launch a bunch
| of tasks, get some planning out of the way, and then step out for
| a long walk in nature."
|
| Wouldn't we all want that, but it sounds like you can leave task
| launching and planning to an AI and go find another career.
| bathtub365 wrote:
| Is there anywhere that lists what languages this supports? They
| aren't listed in the product announcement or in this review, and
| the review examples seem to mostly be fixing typos on webpages.
| yieldcrv wrote:
| > Codex then clones your repositories into its own sandboxes so
| it can run commands and create branches on your behalf.
|
| Slurping up trade secrets
|
| but maybe I'll sound like the people that are afraid of using
| github and other cloud git protocols
|
| interesting crossroads
| theowijrhrjrj48 wrote:
| Sounds like a gptel-tool one can whip up in a week.
___________________________________________________________________
(page generated 2025-05-20 23:01 UTC)