[HN Gopher] O3 Turns Pro
       ___________________________________________________________________
        
       O3 Turns Pro
        
       Author : jsnider3
       Score  : 143 points
       Date   : 2025-06-17 14:49 UTC (8 hours ago)
        
 (HTM) web link (thezvi.substack.com)
 (TXT) w3m dump (thezvi.substack.com)
        
       | iLoveOncall wrote:
       | > My experience so far is that waiting a long time is annoying,
       | sufficiently annoying that you often won't want to wait.
       | 
       | My solution for this has been to use non-reasoning models, and so
       | far in 90% of the situations I have received the exact same
       | results from both.
        
         | jasonjmcghee wrote:
         | On the complete other end of the spectrum, I found deep
         | research (whether it's actually performing searches or not) to
         | be a significant upgrade in quality. But you need to be cool
         | with having to wait 15-30 minutes. It's certainly not for
         | everything, but definitely worth trying.
         | 
         | It tends to output significantly longer and more detailed
         | output. So when you want that kind of thing- works well.
         | Especially if you need up to date stuff or want to find related
         | sources.
        
           | joshstrange wrote:
           | Deep research is very cool, no doubt, but run it on a problem
           | space you are familiar with and you will see the
           | shortcomings.
           | 
           | Anytime I do my own "deep" research I like to then throw the
           | same problem at OpenAI and see how well it fares. Often it
           | misses things or gets things subtly wrong. The results look
           | impressive so it's easy to fool people and I'm not saying the
           | results are useless, I've absolutely gotten value out of it,
           | but I don't love using it for anything I actually care about.
        
             | bcrosby95 wrote:
             | I view the results more as a starting point than an end
             | unto itself. For that I think it's pretty useful.
        
               | matwood wrote:
               | Same, it will pull enough sources together that I end up
               | with an idea of where to go next.
        
               | joshstrange wrote:
               | Absolutely, I agree it's useful as a starting point,
               | sometimes it's all I need (if it's low-stakes and I just
               | wanted a bit more data). I was just cautioning "trusting"
               | it completely, since it's very easy to fall into that
               | trap (I've done it).
        
       | vessenes wrote:
       | I'm using Pro. It's definitely a "hand it to the team and have
       | them schedule a meeting to get back to me" speed tool. But, it
       | "feels" better to me than o3, and significantly better than
       | gemini/claude for that use case. I do trust it more on
       | confabulations; my current trust hierarchy would be o3-pro -> o3
       | -> gemini -> claude opus -> (a bunch of stuff) -> 4o.
       | 
       | That said, I'd like this quality with a relatively quick tool
       | using model; I'm not sure what else I'd want to call it "AGI" at
       | that point.
        
         | bananapub wrote:
         | what do you trust it to do?
         | 
         | the only example uses I see written about on HN appear to
         | basically be Substack users asking o3 marketing questions and
         | then writing substack posts about it, and a smattering of vague
         | posts about debugging.
        
           | vessenes wrote:
           | Long form research reporting.
           | 
           | Example: Pull together a list of the top 20 startups funded
           | in Germany this year, valuation, founder and business model.
           | Estimate which is most likely to want to take on private
           | equity investment from a lower mid market US PE fund, as well
           | as which would be most suitable taking into consideration
           | their business model, founders and market; write an approach
           | letter in english and in german aimed at getting a meeting.
           | make sure that it's culturally appropriate for german startup
           | founders.
           | 
           | I have no idea what the output of this query would be by the
           | way, but it's one I would trust to get right on
           | 
           | * the list of startups
           | 
           | * the letter and its cultural sensitivity
           | 
           | * broad strokes of what the startup is doing
           | 
           | Stuff I'd "trust but verify" would be
           | 
           | * Names of the founders
           | 
           | * Size of company and target market
           | 
           | Stuff I'd double check / keep my own counsel on
           | 
           | * Suitability and why (note that o3 pro is def. better at
           | this than o3 which is already not bad; it has some genuinely
           | novel and good ideas, but often misses things.)
        
             | leptons wrote:
             | This is all stuff I would expect an LLM to "hallucinate"
             | about. Every bit of it.
        
               | steveklabnik wrote:
               | These tools can search the web to find this kind of data,
               | and show you what they searched. Double checking is
               | essential because hallucinations are still possible, but
               | it's not like in the past where it would just try to make
               | up the data from its training set. That said, it also may
               | find bad data and give you a summary of that, which isn't
               | a direct hallucination, but can still be inaccurate. This
               | is why checking the sources is helpful too.
        
               | majormajor wrote:
               | I wouldn't expect it to hallucinate but how do you
               | evaluate it's ability to distinguish spam from good info?
               | I.e. the "the first four pages of google results is all
               | crap nowdays" problem.
        
               | steveklabnik wrote:
               | By looking at the pages it looked at and deciding for
               | yourself, just like you would with a web search you
               | invoked yourself. I've generally found it to use
               | trustworthy stuff like Stack Overflow, Wikipedia, and
               | university websites. But I also haven't used it in this
               | way that much or for very serious things. I'd imagine
               | more obscure questions are more likely to end up
               | involving less trustworthy sites.
        
               | leptons wrote:
               | >By looking at the pages it looked at and deciding for
               | yourself, just like you would with a web search you
               | invoked yourself.
               | 
               | Or you could just cut out the middleman(bot) and just do
               | the search yourself, since you're going to have to anyway
               | to verify what the "AI" wrote. It's just all so stupid
               | that society is rushing towards this iffy-at-best
               | technology when we still need to do the same work anyway
               | to verify it isn't bullshitting us. Ugh, I hate this
               | timeline.
        
               | thelock85 wrote:
               | I recently tried a version of this landscape analysis
               | within a space I understand very well (CA college access
               | nonprofits) and was shocked at how few organizations were
               | named, let alone described in detail. Even worse, the
               | scope and reach of the named orgs were pretty off the
               | mark. My best guess is that they were the SEO winners of
               | the past.
        
               | vessenes wrote:
               | Well you'd be wrong in this case: Deep research will
               | trigger a series of web searches first then reach out to
               | tooling for follow ups as needed; most of the facts will
               | be grounded in the sources it finds.
               | 
               | With no deep research - agreed; too recent to believe
               | info is accurately stored in the model weights.
        
             | bananapub wrote:
             | why would you trust it to get any of that right? things
             | like "top 20 startups in Germany" sound hard to determine.
             | 
             | how do you validate all of that is actually correct?
        
               | jazzyjackson wrote:
               | A lot of stuff doesn't need to be accurate, it just needs
               | to be enough information to act on.
               | 
               | Like how there's a ton of psychics, tarot and palm
               | readers around Wall St.
        
               | bananapub wrote:
               | That's fine, but no one - not Sam Altman, not the fans on
               | HN - are promoting them as $120/million token
               | clairvoyants, they're claiming they are srs bzns "iq
               | maxxing" research tools.
               | 
               | If OP had suggested that they were just medium-quality
               | nonsense generators I would have just agreed and not
               | replied.
        
           | lovich wrote:
           | I've been using it in my job search by handing it stuff like
           | the hn whose hiring threads, giving it a list of criteria i
           | care about, and have it scour those posts for matching jobs,
           | and then chase down all the companies posting and see if they
           | have anything on their corporate site matching my
           | descriptions.
           | 
           | Then I have it take those matches and try and chase down the
           | hiring manager based on public info.
           | 
           | I did it at first just to see if it was possible, but I am
           | getting direct emails that have been accurate a handful of
           | times and I never would have gotten that on my own
        
             | bananapub wrote:
             | This is a good data point - I guess another dimension is
             | incompleteness-tolerance. An LLM is absolutely going to
             | miss some but for your case that doesn't matter very much.
             | 
             | Thank you!
        
           | lukeschlather wrote:
           | I don't think it's necessarily a question of trust, it's a
           | question of cost/benefit, and I can apply this just as much
           | to myself. I have been using a lot more SQL queries lately
           | when I use ChatGPT, because I trust it pretty well to write
           | gnarly queries with subqueries and CASE statements. Things
           | that I wouldn't write myself because it's not worth the time
           | to make the query correct, but ChatGPT can do it in seconds.
           | 
           | I had an example where o1 really wowed me - something I don't
           | want to post on the internet because I want to use it to test
           | models. In that case I was thinking through a problem where I
           | had made an incorrect mathematical assumption. I explained my
           | reasoning to o1 and it was able to point out the flaw in my
           | reasoning, along with some examples mathematical expressions
           | that disproved my thinking.
           | 
           | The funny thing in this case it basically functioned as a
           | rubber duck. When it started producing a response I had
           | deduced essentially what it told me - but it was pretty nice
           | to see the detailed reasoning with examples that might've
           | taken me a few more minutes to work out. And I never would've
           | produced a little report explaining in detail why I was
           | wrong, I would've just adjusted my thinking. Having the
           | report was helpful.
        
         | Y_Y wrote:
         | Are you setting the "reasoning effort"? I find going from the
         | default (medium) to high makes a big difference on coding tasks
         | for openai reasoning models.
        
           | pas wrote:
           | what/how does that work internally?
        
         | qwertox wrote:
         | What are you using it for? It's not like that wouldn't matter.
         | 
         | With coding using anything is always a hit and miss, so I
         | prefer to have faster models where I can throw away the chat if
         | it turns into an idiot.
         | 
         | Would I wait 15 minutes for a transcription from Python to Rust
         | if I don't know what the result will be? No.
         | 
         | Would I wait 15 minutes if I'd be a mathematician working on
         | some kind of proof? Probably yes.
        
           | AaronAPU wrote:
           | I feed most of my questions/code to 4o, Gemini, o3-pro (in
           | that order). By the time I've read through 4o, Gemini is
           | ready. Etc.
           | 
           | It's the progressive jpg download of 2025. You can short
           | circuit after the first model which gives a good enough
           | response.
        
             | plufz wrote:
             | How do you reason about the energy consumption/climate
             | impact of feeding the same question to three models? Im not
             | saying there is a clear answer here, would just be
             | interesting to hear your thinking.
        
               | dfsegoat wrote:
               | It's a tough question and I do things the same way.
               | 
               | I feel like we are in awkward phase of: "We know this has
               | severe environmental impact - but we need to know if
               | these tools are actually going to be useful and worth
               | adopting..." - so it seems like just keeping the
               | environmental question at the forefront will be important
               | as things progress.
        
               | true_religion wrote:
               | How much energy does an AI model use during inferencing
               | versus a human being?
               | 
               | This is a rhetorical question.
               | 
               | Sure we aren't capturing every last externality, but
               | optimization of large systems should be pushed toward the
               | creators and operators of those systems. Customers
               | shouldn't have to validate environmental impact every
               | time they spend 0.05 dollars to use a machine.
        
               | kridsdale1 wrote:
               | I actually did the math on this last year some time. For
               | gpt4 or so. Attempted to derive a per-user energy use
               | value. Based on known data LLM training used many
               | hundreds of times the energy use of agriculture and
               | transport costs to feed a human to do equivalent mental
               | work. Inference was much lower. But the climate critique
               | of AI doesn't distinguish.
        
               | blharr wrote:
               | 100x more inefficient than a human _in only food_ is
               | pretty efficient. Consider that humans in the developed
               | world spend far more in energy on heating /AC,
               | transportation, housing, lawn care, refrigeration,
               | washers and dryers, etc, and an LLM can probably be
               | several factors more efficient.
               | 
               | I don't really understand the critique of GPT-4 in
               | particular. GPT-4 cost >$100 Million to train. But likely
               | less than 1 billion. Even if they pissed out $100 million
               | in pure greenhouse gases, that'd be a drop in the bucket
               | compared to, say 1/1000 of the US military's
               | contributions
        
               | ben_w wrote:
               | That sounds on the low side?
               | 
               | Does that "hundreds" include the cost of training one
               | human to do the work, or enough humans to do the full
               | range of tasks that an LLM can do? It's not like-for-like
               | unless it's the full range of capabilities.
               | 
               | Given the training gets amortised over all uses until the
               | model becomes obsolete (IDK, let's say 9 months?), I'd
               | say details like this do matter -- while I want the
               | creation to be climate friendly just in its own right
               | anyway, once it's made, greater or lesser use does very
               | little:
               | 
               | As a rough guess, let's say that any given extra use of a
               | model is roughly equivalent to turning API costs into kWh
               | of electricity. So, at energy cost of $0.1/kWh,
               | GPT-4.1-mini is currently about 62,500 tokens per kWh.
               | 
               | IDK the typical speed of human thought (and it probably
               | doesn't map well to tokens), but for the sake of a rough
               | guide, I think most people reading a book of that length
               | would take something around 3 hours? Which means if the
               | models burn electricity at about 333 W, they equal the
               | performance (speed) of a human, whose biological
               | requirements are on average 100 W... except 100 W is what
               | you get from dividing 2065 kcal by 24h, and humans not
               | only sleep, but object to working all waking hours 7 days
               | a week, so those 3 hours of wall-clock time come with
               | about 9 hours of down-time (40 hour work week/(7 days
               | times 24 hours/day) ~= 1/4), making the requirements for
               | 3 hours work into 12 hours of calories, or the equivalent
               | of 400 W.
               | 
               | But that's for _reading_ a book. Humans could easily
               | spend months writing a book that size, so an AI model
               | good enough to _write_ 62,500 useful tokens could easily
               | be (2 months * 2065 kcal /day = 144 kWh), at $0.1/kWh
               | around $14.4, or $230/megatoken price range, and still
               | more energy efficient than a human doing the same task.
               | 
               | I've not tried o3*, but I have tried o1, and I don't
               | think o1 can write a book-sized artefact that's worth
               | reading. But well architected code isn't a single
               | monolith function with global state like a book can be,
               | you can break everything down usefully and if one piece
               | doesn't fit the style of the rest it isn't the end of the
               | world, so it may be fine for code.
               | 
               | * I need to "verify my organisation", but also I'm a solo
               | nerd right now, not an organisation... if they'd say I'm
               | good, then that verification seems not very important?
        
               | themanmaran wrote:
               | The same way you might reason about the climate impact of
               | having a youtube video on in the background I expect.
        
               | AaronAPU wrote:
               | I don't have nearly a luxurious enough life for that to
               | be a blip on my radar of concerns.
        
               | naming_the_user wrote:
               | I don't think about it at all. When they cost money we'll
               | care more about it.
               | 
               | Until then, the choice is being made by the entities
               | funding all of this.
        
               | omikun wrote:
               | Likely how you reason about driving to the beach or
               | flying to a vacation destination. Or playing a game in 4k
               | high quality with ray tracing turned on.
        
             | phillco wrote:
             | Do you have any specific tooling for querying all three at
             | once beyond just copy paste?
        
         | achierius wrote:
         | Ideally it should be able to do things outside of the realm of
         | programming with strong reliability (at least as strong as
         | human experts), as well as be able to pick up new skills and
         | learn new facts dynamically.
        
         | IamLoading wrote:
         | The time o3 pro takes is so annoying. I still need some time to
         | get used to that.
        
         | JamesBarney wrote:
         | I haven't tested o3-pro yet enough to have a good hierarchy of
         | confabulation.
         | 
         | I use AI a lot to double check my code via a code review what
         | I've found is
         | 
         | Gemini - really good at contextual reasoning. Doesn't
         | confabulate bugs that don't exist. Is really good at finding
         | issues related to large context. (this method calls this
         | method, and it does it with a value that could be this)
         | 
         | Sonnet/Opus - Seems to be the more creative. More likely to
         | confabulate bugs that don't exist, but also most likely to
         | catch a bug o3 and gemini missed.
         | 
         | o3 - Somewhere in the middle
        
         | agambrahma wrote:
         | Hmm, Gemini + O3 > Claude-Opus for ... what kinds of things?
        
       | snissn wrote:
       | I've found throw the problem at 3 o3 pros and have another one
       | evaluate and synthesize works really well
        
         | ActionHank wrote:
         | So like, a whole forest of trees per query is what we're saying
         | here?
        
           | LeafItAlone wrote:
           | Ideally just a few split atoms
        
         | kridsdale1 wrote:
         | Now You're Playing With Agent Power!
        
       | franze wrote:
       | I use Claude Code a lot. A lot lot. I make it do Atomic Git
       | commits for me. When it gets stuck and instead of just saying so
       | starts to refactor half of the codebase, I jump back to commit
       | where the issue first appeared and get a summary of the involved
       | files. Those in full text (not files) into o3 pro. And you can be
       | sure it finds the issue or gives a direction where the issue does
       | not appear. Would love o3-pro as am MCP so whenever Claude Code
       | goes on a "lets refactor everything" coding spree it just asks o3
       | pro.
        
         | jgalt212 wrote:
         | > When it gets stuck and instead of just saying so starts to
         | refactor half of the codebase
         | 
         | That's pretty scary.
        
           | franze wrote:
           | Atomic Commits.
           | 
           | I put this into Claude.md and need to remind it every other
           | hour. But yeah, you need to jump back every few hours or so.
        
             | nevertoolate wrote:
             | Can you give an example what claude works on autonomously
             | for hours? I only use the chat, maybe I'm just not
             | prompting well, but I throw away almost everything claude
             | writes and solve it in significantly less lines of code
             | using the proper abstractions.
        
               | ActionHank wrote:
               | Yeah, so far, I've only seen cases where the work is
               | extremely simple and using pervasively used libraries and
               | solutions to create widely implemented solutions. Add
               | something a little out there and things start to unravel.
        
               | franze wrote:
               | currently i am coding a node/react/ts firebase app that
               | allows dynamic multiagent workflows to automate content
               | workflows (a workflow.json defines call this model and
               | the pass this part of the output of that model to that
               | model and then combine it with this model to do that)
               | 
               | my setup is claude code in yolo mode with playwright MCP
               | + browser MCP (to do stuff in the logged i firebase web
               | interface) plus search enabled.
               | 
               | the prototype was developed via firebase studio until i
               | reached a dead end there, then i used claude code to rip
               | out firebase genkit and hooked in google-genai, openai,
               | ...
               | 
               | the whole codebase goes into google gemini studio (caus
               | the million token window) to write tickets, more tickets
               | and even more tickets.
               | 
               | claude code then has the job to implemt these tickets
               | (create a detailed tasklist for each ticket first) and
               | then code it until done. end of each tasklist is a
               | working playwright end to end test with verified output.
               | 
               | and atomic commits.
               | 
               | i hooked anydesk to my computer so i can check i at some
               | point to tell to to continue or to read Claude.md again
               | (the meta instructions which basically tells it to not to
               | fallbacks, mock data or cheat in amy other way.)
               | 
               | ever fourth ticket is refactoring for sinplicity and
               | documentation.
               | 
               | the tickets mist be updated before each commit and moved
               | to the do done folder only when 100 tested ok.
               | 
               | so yeah, when i wale up in the morning either magic
               | happend and the tockets are all done. or it got stuck and
               | refactores half the codebase. in that case it works for
               | an hoor to go over all git commits to find out where it
               | went wrong.
               | 
               | what i need are multiple coding agent which challenge
               | each other at crucial points.
        
               | throw234234234 wrote:
               | I have to ask a probably naive question - after the
               | initial boilerplate/scaffolding is this actually any
               | faster than just typing in the code you want? Or using
               | the standard AI flow before these long task agents? It
               | feels like you are juggling and bouncing async tools,
               | doubling back on output, and constant trial and error to
               | get things working.
               | 
               | I'm sure lots of code is being generated, but I do wonder
               | about the effectiveness ratio of it when I read comments
               | like above. Like there is a sweet spot after initial
               | scaffold where its easier just to express yourself in
               | code?
        
         | BeetleB wrote:
         | Sounds like you're doing the equivalent of Aider's architect
         | mode (use one model for the reasoning, and another for the code
         | changes).
         | 
         | I would encourage you to try it. It's generally (much) cheaper
         | doing stuff in Aider, but if you're paying a monthly
         | subscription and using it a lot, Claude Code may be cheaper...
        
       | A_D_E_P_T wrote:
       | Chat just isn't the best format for something that takes 15-20
       | minutes (on average) to come up with a response. Email would
       | unironically be better. Send a very long and detailed prompt,
       | like a business email, and get a response back whenever it's
       | ready. Then you can refine the prompt in another email, etc.
       | 
       | But I should note that o3-pro has been getting faster for me
       | lately. At first every damn thing, however simple, took 15+
       | minutes. Today I got a few answers back within 5 minutes.
        
       | starik36 wrote:
       | I've tried o3 Pro for my use cases (parsing emails in the legal
       | profession) and didn't have better results than the non pro.
       | 
       | In fact, o1-preview has given me more consistently correct
       | results than any other model. But it's being sunset next month so
       | I have to move to o3.
        
         | AaronAPU wrote:
         | IMO 4o is much better at people-parsing. The reasoning models
         | o1-pro / o3-pro are really good at writing code and solving
         | algorithmic problems.
        
           | resters wrote:
           | what is people parsing?
        
             | starik36 wrote:
             | The email from the lawyer might mention lots of names. Who
             | are the plaintiffs, who are defendants, their attorneys,
             | assistants, or insurance adjusters. The model parses out
             | who is who and connects names to titles to email addresses.
        
               | resters wrote:
               | Interesting, that's what I thought it meant, but didn't
               | realize it was a term of art.
        
             | AaronAPU wrote:
             | Things like inferring the meaning of "people parsing" when
             | it isn't explicitly defined but can be implied by context.
             | 
             | Not strict rational A+B=C, nuance.
        
           | starik36 wrote:
           | I've tried it with various models. And 4o is really good
           | given that it returns data at least 10 times faster. But if
           | you ask it to fill out a Json document, o3 (or other
           | reasoning models) is still better, more correct and
           | predictable. Or at least, better enough to justify waiting a
           | minute for the API call to return vs 3-5 seconds.
        
         | ActionHank wrote:
         | Out of interest, how widespread would you say this usage is
         | amongst your peers in the legal profession?
        
           | starik36 wrote:
           | ChatGPT is pretty widespread. The only obstacle in the past
           | was the fear that confidential documents might be used for
           | training. OpenAI fixed that with a business account type that
           | guarantees no training.
           | 
           | As far as usage of API for business processes (like document
           | processing) - I can't say.
        
             | AtlasBarfed wrote:
             | Guarantees ...... How?
             | 
             | You should assume Facebook level morality.
        
               | starik36 wrote:
               | I hear you. But the in-house lawyer read and approved the
               | SLA. So all asses are covered!
        
       | swyx wrote:
       | > Arena has gotten quite silly if treated as a comprehensive
       | measure (as in Gemini 2.5 Flash is rated above o3)
       | 
       | > The problem with o3-pro is that it is slow.
       | 
       | well maybe Arena is not that silly then. poorly argued/organized
       | article.
        
       | jes5199 wrote:
       | I haven't tried pro yet but just yesterday I asked O3 to review a
       | file and I saw a message in the chain-of-thought like "it's going
       | to be hard to give a comprehensive answer within the time limit"
       | so now I'm tempted
        
       | rotcev wrote:
       | I use O3-pro not as a coding model, but as a strategic assistant.
       | For me, the long delay between responses makes the model
       | unsuitable for coding workflows, however, it is actually a
       | feature when it comes to getting answers to hard questions
       | impacting my (or my friend's/family's) day to day life.
        
       | b0a04gl wrote:
       | when o3 pricing dropped 80%, most wrote the entire model family
       | off as a downgrade (including me). but usage patterns flipped
       | people finally ran real tasks through it. it's one of the few
       | that holds state across fragmented prompts without collapsing
       | context. used it to audit a messy auth flow spread over 6
       | services. didn't shortcut, didn't hallucinate edge cases. slow,
       | but deliberate. in kahneman terms, it runs system 2 by default.
       | many still benchmark on token speed, missing what actually
       | matters
        
         | lysecret wrote:
         | This feels very Ai generated.
        
           | mettamage wrote:
           | Some people write in similar ways yea. I've also been accused
           | of writing as an AI.
           | 
           | But we're still human mate.
           | 
           | Stop discriminating or actually solve the problem. I've had
           | enough of this attitude.
        
             | cshimmin wrote:
             | almost as though the AIs were trained on a corpus of text
             | written by... humans
        
           | b0a04gl wrote:
           | yes im agi by the way
        
             | kridsdale1 wrote:
             | hi agi we've been trying so hard to find you
        
           | motoxpro wrote:
           | I would say the opposite. Unless the person has a lot of
           | custom instructions going on. Getting sentences like "but
           | usage patterns flipped people finally ran real tasks through
           | it." seem like it would take some amount of work.
        
           | SkyPuncher wrote:
           | Feels like a lot of software engineers I work with (including
           | myself at times).
           | 
           | Short, concise statements that don't necessarily string
           | together sequentially. However, they still aggregate to a
           | holistic, meaningful thought. No that much different that how
           | a lot of code is written.
        
           | gala8y wrote:
           | Actually, it does not.
        
         | lubujackson wrote:
         | I have been using o3 almost exclusively in Cursor now for my
         | "vibe coding" project. I was able to get to a point with faster
         | models before hitting a thrashing problem of forgetting about
         | structure/not updating types/no using right types/ignoring
         | existing functions, etc. Even when providing specific context.
         | o3 rarely hits those issues and can happily implement a fully
         | feature without breaking anything that touches multiple files.
         | Speed is definitely an issue, but much less hassle on the back
         | side.
        
       | metalrain wrote:
       | "'take your profits' in quality versus quantity is up to you."
       | 
       | As mainly AI invester not AI user, I think profitability is great
       | importance. It has been race to top so far, soon we see race to
       | the bottom.
        
         | resters wrote:
         | Right! We are in a sense lucky to be getting access to actual
         | state-of-the-art models. Soon the actual model may be kept
         | internal and the customers will get "good enough for solid ROI"
         | distilled versions that can be hosted profitably.
        
       | boole1854 wrote:
       | Here are my own anecdotes from using o3-pro recently.
       | 
       | My primary use cases where I am willing to wait 10-20 minutes for
       | an answer from the "big slow" model (o3-pro) is code reviews of
       | large amounts of code. I have been comparing results on this task
       | from the three models above.
       | 
       | Oddly, I see many cases where each model will surface issues that
       | the other two miss. In previous months when running this test
       | (e.g., Claude 3.7 Sonnet vs o1-pro vs earlier Gemini), that
       | wasn't the case. Back then, the best model (o1-pro) would almost
       | always find all the issues that the other models found. But now
       | it seems they each have their own blindspots (although they are
       | also all better than the previous generation of models).
       | 
       | With that said, I am seeing Claude Opus 4 (w/extended thinking)
       | be distinctly worse at missing problems which o3-pro and Gemini
       | find. It seems fairly consistent that Opus will be the worst out
       | of the three (despite sometimes noticing things the others do
       | not).
       | 
       | Whether o3-pro or Gemini 2.5 Pro is better is less clear. o3-pro
       | will report _more_ issues, but it also has a tendency to
       | confabulate problems. My workflow involves providing the model
       | with a diff of all changes, plus the full contents of the files
       | that were changed. o3-pro seems to have a tendency to imagine and
       | report problems in the files that were not provided to it. It
       | also has an odd new failure mode, which is very consistent: it
       | gets confused by the fact that I provide both the diff and the
       | full file contents. It  "sees" parts of the same code twice and
       | will usually report that there has accidentally been some code
       | duplicated. Base o3 does this as well. None of the other models
       | get confused in that way, and I also do not remember seeing that
       | failure mode with o1-pro.
       | 
       | Nevertheless, it seems o3-pro can sometimes find real issues that
       | Gemini 2.5 Pro and Opus 4 cannot more often than vice versa.
       | 
       | Back in the o1-pro days, it was fairly straightforward in my
       | testing for this use case that o1-pro was simply better across
       | the board. Now with o3-pro compared particularly with Gemini 2.5
       | Pro, it's no longer clear whether the bonus of occasionally
       | finding a problem that Gemini misses is worth the trouble of (1)
       | waiting _way_ longer for an answer and (2) sifting through more
       | false positives.
       | 
       | My other common code-related use case is actually writing code.
       | Here, Claude Code (with Opus 4) is amazing and has replaced all
       | my other use of coding models, including Cursor. I now code
       | almost exclusively by peer programming with Claude Code, allowing
       | it to be the code writer while I oversee and review. The OpenAI
       | competitor to Claude Code, called Codex CLI, feels distinctly
       | undercooked. It has a recurring problem where it seems to
       | "forget" that it is an agent that needs to go ahead and edit
       | files, and it will instead start to offer me suggestions about
       | how I can make the change. It also hallucinates running commands
       | on a regular basis (e.g., I tell it to commit the changes we've
       | done, and outputs that it has done so, but it has not.)
       | 
       | So where will I spend my $200 monthly model budget? Answer:
       | Claude, for nearly unlimited use of Claude Code. For highly
       | complex tasks, I switch to Gemini 2.5 Pro, which is still free in
       | AI Studio. If I can wait 10+ minutes, I may hand it to o3-pro.
       | But once my ChatGPT Pro subscription expires this month, I may
       | either stop using o3-pro altogether, or I may occasionally use it
       | as a second opinion by paying on-demand through the API.
        
         | JamesBarney wrote:
         | > With that said, I am seeing Claude Opus 4 (w/extended
         | thinking) be distinctly worse at missing problems which o3-pro
         | and Gemini find. It seems fairly consistent that Opus will be
         | the worst out of the three (despite sometimes noticing things
         | the others do not).
         | 
         | I've found the same thing. That claude is more likely miss a
         | bug than o3 or gemini but more likely to catch something o3 and
         | gemini missed. If I had to pick one model I'd pick o3 or
         | gemini, but if I had to pick a second model I'd pick opus.
         | 
         | It's also seems to have a much higher false positive rate where
         | as gemini seems to have the lowest false positive rate.
         | 
         | Basically o3 and gemini are better, but also more correlated
         | which gives opus a lot of value.
        
         | throwdbaaway wrote:
         | For the code review use case, maybe can try to create the diff
         | with something like `git diff -U99999`, and then send only the
         | diff.
        
       | highfrequency wrote:
       | What is the difference between o3 pro and deep research? From a
       | glance, both seem to take 10-15mins to respond and use o3 as the
       | base model.
        
       ___________________________________________________________________
       (page generated 2025-06-17 23:01 UTC)