[HN Gopher] Improving GPT-5.6 Sol in ChatGPT, expanding GPT-5.6 ...
       ___________________________________________________________________
        
       Improving GPT-5.6 Sol in ChatGPT, expanding GPT-5.6 Luna access for
       free users
        
       Author : tedsanders
       Score  : 132 points
       Date   : 2026-08-06 17:02 UTC (6 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | tosh wrote:
       | free unlimited luna is a pretty badass move
       | 
       | luna is very good
        
         | skybrian wrote:
         | It's also pretty cheap if you're paying for it (for coding).
         | But I don't quite trust it for anything complicated, so I often
         | use Terra.
        
           | redox99 wrote:
           | Terra is awful and not cheap enough to make up for it. I'd
           | suggest using luna or sol.
        
             | Sammi wrote:
             | So Luna is more awful but it's OK because it's cheaper?
        
               | redox99 wrote:
               | Yes. You're better off running Luna Max or Sol Medium and
               | never touching Terra.
        
               | LaurensBER wrote:
               | I've had great success with using Luna and having
               | DeepSeek 4 Flash check Luna's output. Oh My Pi has a mode
               | built-in that does this automatically ("advisor" mode).
               | Deepseek only interrupts when it spots an issue so it
               | doesn't slow Luna down.
               | 
               | Both models are cheap enough that I can run 4 sessions at
               | the same time without running out of the 20 USD codex and
               | 10 USD Opencode plan. I've burned through almost a
               | billion tokens this week and I've done some pretty big
               | refactors as well.
               | 
               | I have a Claude Max subscription but I've barely been
               | using it because of the many issues they've had this
               | week.
        
             | stuartq wrote:
             | Terra is far from awful. I've not used Luna enough to judge
             | whether it's significantly better than Luna, but at least
             | via GitHub Copilot, Terra is leaps ahead of Sonnet 5.
        
         | causal wrote:
         | I guess I should give it another shot because I had pretty bad
         | experience with Luna when it first came out. Stuff Sonnet knew
         | better.
        
         | timpera wrote:
         | Any improvement to the ChatGPT free plan is really nice. It's
         | easy to forget that most people have never used a SOTA model,
         | and only think of their experience with GPT-4o or Google
         | Search's AI Overviews when asked about AI.
        
         | msq22 wrote:
         | Were free users subject to some limits? I've never run into any
         | limits.
        
           | timpera wrote:
           | It used to be 10 messages every 5 hours using GPT-5, then
           | unlimited 4o-mini. More recently, it went down to ~5 messages
           | per day on 5.5 Instant, then unlimited on 5.5 mini.
        
       | jauntywundrkind wrote:
       | At first I thought the Sol updates was perhaps trying to help
       | with some complaints of Sol burning through tokens, complaints
       | that have prompted some new data points on https://codex-
       | resets.com/ .
       | 
       | But seeing the graphic with the visual weather report: that makes
       | me think that is not the goal at all. :)
        
         | laweijfmvo wrote:
         | What a bizarre example of a more direct answer. Any human would
         | simply say "No, it doesn't rain here in summer."
         | 
         | Even after identifying the 0% chance of rain, it still drags
         | the conversation on and on and on
        
           | taikahessu wrote:
           | Would you like to know more? Seriously though, of course,
           | it's tuned for maximum engagement, not maximum efficiency. I
           | wonder how long this engagement dopamine circus can last...
           | too long apparently.
        
           | sunaookami wrote:
           | GPT models were RLHF'd to death, they will never give a
           | final, direct answer. Every release since the GPT-4o
           | catastrophy is like this, it's so tiresome. They need a
           | complete reset before it can become actually usable again. Or
           | not as maximum engagement seems to be their goal.
        
       | colingauvin wrote:
       | They must really be feeling the commoditization pressure. I'm not
       | sure what the way out of this is, ChatGPT and Claude are still
       | good products, but they are not necessarily premium products
       | anymore.
       | 
       | I expect a few things to happen in the next year:
       | 
       | 1) Exclusive MCP server deals/API integrations
       | 
       | 2) Significant switch to B2B marketing, even moreso than we've
       | seen before, with API interfaces being paid and chat-client
       | interfaces becoming more and more free, perhaps just with limits
       | more on integrations or data visualization/analysis
       | 
       | 3) US restrictions on B2B contracts with non-US hosted models
       | that do any sort of contracting with the government
       | 
       | Obviously there's a bunch of stuff I'm not foreseeing. But it
       | really does feel like the bottom of the market is collapsing into
       | free. I assume OpenAI and Anthropic think their next generation
       | of models will restore their halo tier status and that the cash
       | burn is justified to just get there, but this has to really mess
       | up IPO plans.
        
         | redox99 wrote:
         | I'm not sure I agree
         | 
         | 1) Back then, even as a free user you'd be able to use the
         | strongest model (even if with tight limits). Now, you _need_ to
         | pay to use Sol, and you _need_ to pay to use Opus or Fable. It
         | does seem fairly premium in that sense. Idk about 5.6 Luna, but
         | the previous Instant was really bad, even for very casual
         | users. It would hallucinate non stop.
         | 
         | 2) When $100 and $200 per month plans launched, they were
         | received as outrageous even here. Nowadays they are pretty
         | common among power users.
        
           | colingauvin wrote:
           | >1) Back then, even as a free user you'd be able to use the
           | strongest model (even if with tight limits). Now, you need to
           | pay to use Sol, and you need to pay to use Opus or Fable. It
           | does seem fairly premium in that sense. Idk about 5.6 Luna,
           | but the previous Instant was really bad, even for very casual
           | users. It would hallucinate non stop.
           | 
           | This is kind of what I'm saying though. Bottom has fallen
           | out, differentiation is just can you be much more premium
           | than the competition. Currently that remains unanswered.
           | 
           | EDIT: I'm basing this off the assumption that for chat,
           | premium is not a point of differentiation at all. For
           | coding/analysis, it is.
        
             | davidguetta wrote:
             | especially when 99% tasks don't need fable, and certainly
             | not the next model
        
           | user43928 wrote:
           | I also scoffed at the ChatGPT $200 Pro plan back then.
           | 
           | Back when coding for me still meant copy-paste from the web
           | version, it was only worth the $20/month for me.
           | 
           | They only added the $100 Pro plan in April during GPT 5.4
           | times.
           | 
           | Today I happily pay $400/month for Codex and Claude Code.
        
             | thejazzman wrote:
             | Don't say the last part out loud or it will be $800.
        
       | ignoramous wrote:
       | Every week, 1 billion people turn to ChatGPT for everything from
       | quick questions and web searches to planning, research, advice,
       | and complex decisions.
       | 
       | Guess, Google's _AI Mode_ is chipping away at their consumers (I
       | know I haven 't used Chat in a long, long while for 'quick
       | questions and web searches' after OpenAI did away with "think"
       | which I _always_ use). The money-minting office  & coding market
       | Anthropic has cornered is hyper-competitive at both the frontier
       | & low-cost ends. OpenAI is reactive [0] and seems right up
       | against it, despite the strength of its excellent models.
       | 
       | [0] Won't put it past OpenAI (and/or Google) to open weight
       | larger models!
        
         | skybrian wrote:
         | $20/month also gives you API access for coding, in any coding
         | agent. I keep hitting the weekly limit but it's a good deal
         | while it lasts.
        
         | porridgeraisin wrote:
         | Did they do away with think? I think now you have to do it with
         | /think
        
         | drivebyhooting wrote:
         | Google's AI has been very glitchy for me lately. I used to
         | reserve chatGPT for serious work and Gemini for daily
         | personalized unimportant things. But now I switched completely
         | to ChatGPT and resigned myself to their
         | memories/personalization.
        
       | firasd wrote:
       | I think it's a misread to think the default ChatGPT model
       | switching to GPT 5.6 Luna is some sort of desperation move. Keep
       | in mind that Claude .ai never had this extreme stratification
       | between the frontier models and the free tier (Sonnet is
       | available to free users with rate limits).
       | 
       | So 5.6 Luna is just their next version of what they used to call
       | 5.5 instant tier
       | 
       | And 5.x instant models were never much to write home about anyway
       | so the default ChatGPT free model hasn't been particularly
       | distinctive since 4o
        
       | ElijahLynn wrote:
       | I can't wait to never see a reasoning button ever again. Why do I
       | have to reason about what reasoning level to use?
        
         | timpera wrote:
         | I think it's nice to be able to make the model reason for
         | dozens of minutes when you want to go deep on a topic, even if
         | the router thinks it's an easy question.
        
         | skybrian wrote:
         | Some people want quick results. Some people want it to keep
         | searching for a new math proof overnight without giving up, and
         | they have money to burn.
         | 
         | It seems like giving it a time limit or a budget in dollars
         | would be clearer, though?
         | 
         | Or, keep searching until I come back to the computer and ask
         | about progress.
        
           | pllbnk wrote:
           | The _Intelligence_ part of AGI should be able to guide the
           | user through that without all the knobs.
        
             | vanuatu wrote:
             | intelligence is not omniscience though
        
               | stymaar wrote:
               | Asking questions to clarify user intent is a very low bar
               | for intelligence. A bar that all SOTA models fail
               | consistently at though. (It's both funny and legit
               | infuriating when Opus, after having made a dozen wild
               | assumptions without checking with you, then comes back
               | with a request for clarification on some mundane topic).
        
             | minimaxir wrote:
             | OpenAI tried auto-routing with the initial GPT-5 release
             | and it was immediately clear why that was a bad idea.
        
           | dbbk wrote:
           | I don't understand this at all. Whenever I ask Gemini 3.1 Pro
           | Extended, or Claude 5 Max something in chat, the most I ever
           | wait is maybe 30 seconds. Is that really so bad?
        
             | oceanplexian wrote:
             | "Wait 30 seconds" as a concept has been totally
             | incompatible with the web, smartphones, etc for about 20
             | years now.
        
         | awakeasleep wrote:
         | Because your incentives are opposed to the provider's
         | incentives
        
         | redox99 wrote:
         | Because the model can't read your mind and know if you want a
         | quick answer, or an hour long deep dive.
        
           | Jtarii wrote:
           | Then it should just ask the user what they want if its
           | unclear from the context.
        
             | 2sk21 wrote:
             | Exactly! I posted much the same comment in another thread
             | and there were lots of huffy complaints that amounted to
             | "you're prompting it wrong"
        
             | Sammi wrote:
             | That's what the reasoning slider is for!
        
             | redox99 wrote:
             | So instead of just getting an answer, I have to wait until
             | it asks me, and I have to type back a response? Extremely
             | annoying.
        
         | egorfine wrote:
         | I absolutely need instant mode as this is what I use 90% of the
         | time.
         | 
         | Sadly it's not available anymore on the updated desktop app
         | (formerly Codex) and the previous desktop app (formerly
         | ChatGPT) is abandoned.
        
         | nojs wrote:
         | Auto-effort and similarly auto model routing suffer from a
         | halting problem sort of issue: you don't reliably know if a
         | request is complex unless you use a complex model to make the
         | decision.
        
         | miki123211 wrote:
         | I have the opposite problem. I'm not well-calibrated on when
         | I'd want _lower_ reasoning than what 's available to me (and
         | how to compare that to lower-tier models). OpenAI now has Luna,
         | Terra and Sol, each at Low, Medium, High and Xhigh, with
         | Pro/Ultra depending on harness and plan. That's ~15 possible
         | combinations of model and reasoning level, and there isn't a
         | satisfactory explanation of which one you want for any
         | particular task.
         | 
         | I feel that work is basically split into two tiers, hard (which
         | requires a good model and lots of reasoning by definition) and
         | relatively easy (which won't consume much of my limits despite
         | a great model and reasoning, so I may just as well keep it on
         | Sol High).
        
       | kingstnap wrote:
       | Its always fun to try to read between the lines here to speculate
       | why they are doing this.
       | 
       | Maybe Luna efficiency gain was actually significant enough that
       | putting all the free users and giving them super generous limits
       | makes sense.
       | 
       | They might be doing this to improve the messaging of AI among
       | causal users since right now there is a huge amount of datacenter
       | backlash in the US due to AI grievances.
       | 
       | Maybe they have too much excess capacity or they really want to
       | juice token numbers and market share on their dashboards for
       | marketing.
       | 
       | I also wonder if being given access to an actually a decent model
       | like luna with actual thinking budget instead of brainless
       | "instant" modes will start to make causal users understand the
       | real capabilities of these models.
        
         | deanc wrote:
         | It's also going to be more training data for them.
        
         | ToValueFunfetti wrote:
         | Is this new behavior from them? I haven't looked at their free
         | chat offering in a minute, but I thought they always had them
         | close behind paid tier, often with essentially equal products
         | that made it weird for them to sell the paid tier for chat.
        
           | mkozlows wrote:
           | No, free tier was total garbage with GPT-5.
        
           | kingstnap wrote:
           | You can't even pick 5.6 luna on the chat app with a paid
           | subscription. It just gives you sol (or use older models)
           | with what seems like basically as much usage as you want. And
           | sol is considerably smarter than luna.
           | 
           | All of this is of pretty minor importance though. You can't
           | read as many tokens as a subcription can produce so more chat
           | is not the value add nor super important.
           | 
           | I mean there are literally so many providers for free chat if
           | you are willing to use several seperate apps.
           | 
           | The real value in these subs is using codex cli, much like
           | the real point of anthropic subs is using claude code.
           | Because agentic work actually does require a lot of tokens.
        
         | simianwords wrote:
         | > Maybe Luna efficiency gain was actually significant enough
         | that putting all the free users and giving them super generous
         | limits makes sense.
         | 
         | Definitely this. The recent 80% discount was a reaction to
         | Deepseek's update so that they still position near the
         | frontier. My theory: Luna has always had a much higher
         | efficiency. You do know that the model didn't get faster after
         | the discount?
        
         | planb wrote:
         | Let's speculate: they want luna to be their first model "on
         | silicon" and need as much test data as possible before
         | finalizing the design.
        
       | porridgeraisin wrote:
       | I dont pay for a chatgpt subscription, but sometimes I did use
       | the web app for throwaway questions. GPT 5.5 Instant or whatever
       | it was that they had was absolutely horrendous. Never answered a
       | question straight and was pedantic in a way even a redditor
       | wouldn't be. So I dropped it and just opened my paid coding agent
       | for everything. Grok.com is quite good now with grok 4.5 though
       | and I find myself using that often. Hopefully luna will be
       | similar.
        
       | ilaksh wrote:
       | > Our mission is to ensure that artificial general intelligence
       | benefits all of humanity. We're introducing updates to ChatGPT
       | that improve everyday conversations while expanding access for
       | Free users.
       | 
       | This clearly implies that they believe ChatGPT models are AGI and
       | are now willing to say it out loud.
       | 
       | Which I think is a fair interpretation of the term. They are
       | general purpose intelligence in that you can get help from them
       | about almost anything. They are not like narrow single purpose AI
       | models.
       | 
       | I don't think we need that term to mean "can completely emulate a
       | human" or "can do every task any human on earth can do as well as
       | them".
       | 
       | It also needs to be differentiated from ASI with godlike powers
       | many times greater than human.
        
         | kkoncevicius wrote:
         | If we could show the current models to someone like Alan
         | Turing, I am sure he would conclude that we have AGI.
        
           | stymaar wrote:
           | And after half an hour using it he'd just admit that his test
           | was way too simple as these models are still way too dumb
        
             | klibertp wrote:
             | I think of the Turing test as one of the starting lines,
             | along with image recognition ("a summer break project for a
             | group of grad students" resisted being solved for decades).
             | 
             | It _is_ a huge leap. Now we can start talking about
             | "intelligence" at all - we really couldn't before. That
             | we're still hovering barely above the starting line is a
             | separate matter (also worth noting, of course).
        
         | kubb wrote:
         | > This clearly implies that they believe ChatGPT models are AGI
         | and are now willing to say it out loud.
         | 
         | Well, the models are smart enough to point out why this is
         | wrong.
        
         | andai wrote:
         | >It also needs to be differentiated from ASI with godlike
         | powers many times greater than human.
         | 
         | Wasn't there a thread the other day about how very few humans
         | on earth can understand the new math proofs?
         | 
         | Although "with sufficient study" vs "not even with unlimited
         | study" are probably worth distinguishing there.
        
           | ilaksh wrote:
           | Which brings up another point in that we should have a term
           | that distinguished between super intelligence in some
           | capacity and godlike super intelligence.
        
         | taytus wrote:
         | "This clearly implies that they believe ChatGPT models are AGI
         | and are now willing to say it out loud."
         | 
         | I cannot believe the comments I read on HN nowadays. ZERO
         | critical thinking.
        
       | aniceperson wrote:
       | a m-dash in the title? whoa things are degrading fast
        
       | heaney-555 wrote:
       | Giving free ChatGPT users access to reasoning (the 'Think'
       | toggle) will have a broader impact on the world than every new
       | paid model and coding agent combined.
        
         | daemonologist wrote:
         | Going from 5.5 Instant (which was _noticeably_ bad) to 5.6 Luna
         | is a big jump as well. OpenAI is probably the most prominent
         | among the general public - an advantage in some respects but
         | they 're giving away a lot of free inference and thus have to
         | use a pretty small model to do it.
        
           | johnsmith1840 wrote:
           | I really don't see how they're going to be google here. The
           | free tier is dominated by verticle integration and the
           | platform that people use. Long term I imagine google wins the
           | bottom of the market and I'd be suprised if they lost.
        
             | Tostino wrote:
             | People tell these models _everything_. You don 't think a
             | pretty unethical company can figure out a way to monetize
             | that?
        
             | tokioyoyo wrote:
             | Ads.
        
               | stymaar wrote:
               | Ads work when you can server $.01 worth of ads to a user
               | for $.0001 of server cost. I fail to see how you can make
               | it work when doing LLM inference which is significantly
               | costlier than web search.
        
               | tokioyoyo wrote:
               | In about 4 months, OpenAI's fundraiser decks will leak,
               | which will include their current ad revenue.
        
               | kristofferR wrote:
               | 90+% of queries are probably so common/evergreen that, if
               | a cheaper model made all the different language varients
               | and ways of asking the question into one query, they
               | could be cached quite effectively.
        
               | virgildotcodes wrote:
               | I recently saw some discussion about the CPC for personal
               | injury lawyers being ~$200 on Google ads. Seems insane to
               | me, but I'm totally disconnected from the advertising
               | world. That said, depending on the conversion rates
               | between chats and clicks, the math could be favorable for
               | OAI.
        
               | famouswaffles wrote:
               | > I fail to see how you can make it work when doing LLM
               | inference which is significantly costlier than web
               | search.
               | 
               | The median LLM query isn't significantly costlier than
               | web search.
        
         | gavinray wrote:
         | I'm not sure people fully comprehend the trickle-down pop-
         | culture/zeitgeist effects that LLM's are having/are going to
         | have on humanity.
         | 
         | Because everyone now outsources much of their thinking and
         | researching to LLM's, our collective culture + brain is shaped
         | in a cyclical manner by using them.
         | 
         | It's the mechanical homogenization of culture and groupthink.
        
         | thorum wrote:
         | Free users have had access to reasoning for a while. o4-mini
         | and the initial GPT-5 launch both included reasoning modes for
         | free users.
         | 
         | They took away the button a few months ago and are now putting
         | it back.
        
           | heaney-555 wrote:
           | Technically yes, but the free usage equated to single-digit
           | messages per day before the auto-downgrade. With GPT-5 it was
           | just 1 message per day.
           | 
           | Perhaps I should have said "proper access".
        
         | in-silico wrote:
         | Does the free tier really not have access to reasoning models?
         | 
         | That would explain a lot of the terrible AI/LLM takes online.
        
           | heaney-555 wrote:
           | The free tier of ChatGPT is powered by GPT-5.5 Instant, a
           | non-reasoning version of GPT 5.5.
        
       | simianwords wrote:
       | Does 5.6 Sol finally have an "instant" form? Is that the change?
        
         | redox99 wrote:
         | Yes
        
       | simonw wrote:
       | Nothing in the ChatGPT model release notes yet:
       | https://help.openai.com/en/articles/9624314-model-release-no...
       | 
       | (This is a subtle nudge at anyone from OpenAI who reads this to
       | make sure they get updated.)
       | 
       | OpenAI have a model called "chat-latest" - I wonder if that's
       | running this new model yet:
       | https://developers.openai.com/api/docs/models/chat-latest
       | 
       | It's described as "points to the latest Instant model currently
       | used in ChatGPT" - so presumably that's "GPT-5.6 Instant" in the
       | app.
        
         | firasd wrote:
         | There's no 5.6 instant I think -- 5.6 Luna is gonna be the new
         | instant tier model
        
           | simonw wrote:
           | No, Luna is the new free model.
           | 
           | https://gist.github.com/simonw/aae4febd3c6f7bc5b7811857edb3c.
           | .. has screenshots that still show "Instant" as an option for
           | ChatGPT Chat... but not for ChatGPT Work.
        
             | firasd wrote:
             | Hmm
             | 
             | It's hard to understand this .. like sure we can select
             | instant but is there an actual model called 5.6 instant?
             | Like is it on LM Arena and OpenRouter or available via API
             | etc
             | 
             | 5.5 instant is definitely A Thing it's even name checked in
             | this OAI post
        
               | simonw wrote:
               | It's SO hard to understand this. I couldn't confidently
               | explain it at all.
        
       | OsamaJaber wrote:
       | Expanding free access is mostly an inference cost
       | 
       | Serving cheaply at that scale means routing, batching, and cache
       | hits, not a better model :D
        
       | Squarex wrote:
       | Why pay for ChatGPT Go then?
        
       | sunaookami wrote:
       | This is actually a downgrade for free users since currently it
       | uses GPT-5.5 for a few messages before it drops you down to
       | GPT-5.5-mini. Now it always uses a model worse than Mini (Luna is
       | nano-equivalent, "It roughly corresponds to the nano model tier
       | used in earlier GPT-5 families."
       | https://developers.openai.com/api/docs/models/gpt-5.6-luna ). I
       | guess it's a bit better with Thinking though. They should use
       | Terra for a few messages first before dropping down to Luna. And
       | image inputs are still limited.
        
       | saithound wrote:
       | While their math results are impressive, vibe coding their own
       | web UIs with their subpar design models is really going to
       | backfire if their plan is to attract new users with better free
       | model offerings.
       | 
       | The Aug 6 update has forced the entry box to auto-format Markdown
       | in an attempt to imitate Claude. The implementation is buggy and
       | even simple copy-and-paste has gone entirely haywire. They also
       | forgot to leave a switch to turn the confounded autoformatting
       | thing off.
       | 
       | Chat mode in general is currently crawling with more UX bugs than
       | a porch screen in summer.
        
       | applfanboysbgon wrote:
       | > Our mission is to ensure that artificial general intelligence
       | benefits all of humanity. We're introducing updates to ChatGPT
       | that improve everyday conversations while expanding access for
       | Free users.
       | 
       | My mission is world conquest. I'm writing a comment on an HN
       | thread.
       | 
       | No, those two clauses have no relation whatsoever. I just felt
       | like saying the first sentence because it sounded cool.
        
       | kgeist wrote:
       | >avoid extra detail when it does not help
       | 
       | I wonder if they actually do it to optimize inference. I maintain
       | a corporate AI server and one of the tricks to reduce the load
       | was to modify the system prompt to be as terse as possible so the
       | average response completes faster and requests queue up less
       | often.
        
       ___________________________________________________________________
       (page generated 2026-08-07 00:01 UTC)