[HN Gopher] OpenAI Employee: GPT-4 has been static since March
       ___________________________________________________________________
        
       OpenAI Employee: GPT-4 has been static since March
        
       Author : behnamoh
       Score  : 141 points
       Date   : 2023-06-01 18:27 UTC (4 hours ago)
        
 (HTM) web link (twitter.com)
 (TXT) w3m dump (twitter.com)
        
       | furyofantares wrote:
       | I think we don't notice our expectations have gone up, and we
       | don't notice that we remember the hits and then expect all hits.
       | 
       | We didn't notice the misses at first, because it's what we
       | expected to begin with, and we very strongly noticed the hits
       | because they were unexpected. Now we notice the misses and expect
       | the hits.
        
         | skepticATX wrote:
         | Genuinely asking: have our expectations gone up, or have we
         | started thinking about it in a more critical and realistic way
         | after the initial glean has worn off?
        
           | coffeebeqn wrote:
           | It's magic! But it's also wrong a lot. These days if I ask
           | anything important I'll also have to Google the answer and
           | I'll also ask chatgpt if it's sure about some of the details
        
           | tines wrote:
           | > glean
           | 
           | I think the word you're thinking of is the noun "gleam" which
           | means a kind of lustrous shine, rather than the verb "glean"
           | which means to harvest the remainder of something or to
           | collect in small parts.
        
         | doctoboggan wrote:
         | I definitely agree with this take. I had similar feelings to
         | everyone in yesterday's thread complaining about the models
         | failures but I agree that its just our shifting expectations
         | and not a model update.
        
         | jonas21 wrote:
         | Yeah, it's kind of the same way people have been saying Google
         | search has been going downhill for years, but compressed into a
         | much shorter time frame.
        
           | subsubzero wrote:
           | It has and its gone downhill alot in the past 4-5 years. This
           | is due to the need to keep an ever increasing stream of
           | income flowing to google with a mostly static number of users
           | and "fixing" this issue by piling more ads into users search
           | results.
        
             | awegio wrote:
             | Then just use an Ad Blocker. The actual problem is SEO spam
             | though.
        
           | throwuwu wrote:
           | It has and so has the YouTube algo. You can try the verbatim
           | setting with search to see that there has been some monkeying
           | going on but even that doesn't fix everything. Maybe that's
           | due to spam sites but it just seems less effective.
        
         | throw1623 wrote:
         | This. I've been using it daily and nothing has changed. It's
         | still as great as day 1.
         | 
         | The audience of HN is, as Taleb would say, intellectuals yet
         | idiots. They have trouble measuring change and are prone to
         | hyperboles. Some still try to minimize the impact ChatGPT will
         | have, while focusing on bullshit like 'hallucinations', or
         | nitpicking about the quality of the code and so on. Can't see
         | the forest from the trees.
         | 
         | If you're looking for intelligent discussion, look elsewhere.
        
           | kkdkdod wrote:
           | [dead]
        
         | huevosabio wrote:
         | I think this is the right way to look at it: we are very quick
         | at updating expectations.
         | 
         | The first flight is magic, the nth one is a chore.
        
           | pmontra wrote:
           | You're generally right but one of my nth ones was Brisbane to
           | Cairns over the reef during the day. I'd fly it again and
           | again.
        
             | huevosabio wrote:
             | One of my favorites is landing in Rio, the landscape is
             | spectacular and then you get Christ the Redeemer as the
             | cherry on top of the cake.
        
           | williamcotton wrote:
           | Whenever I'm flying over an ocean on a clear moonlit night it
           | still very much feels like magic!
        
             | myshpa wrote:
             | And now it's even gonna smell like bacon ...
             | 
             | The fat of dead pigs, cattle and chickens is being used to
             | make greener jet fuel
             | 
             | https://www.bbc.co.uk/news/science-environment-65727664
        
         | throwuwu wrote:
         | The refusals are new. I've gotten a few myself recently. I've
         | never seen them before. They're not moderation related and they
         | always include some text about the task being too complex.
        
           | ribosometronome wrote:
           | Are there any 'complex tasks' you've had refused you can
           | share?
        
           | insanitybit wrote:
           | That sounds like it could just be a cap applied to free
           | users. That's hardly the same as a new model.
        
         | irthomasthomas wrote:
         | He specifically singles out the API and never answers if
         | chatgpt is same. Most people are probably using the chatgpt
         | interface, which seems to have more alignment training.
        
         | ren_engineer wrote:
         | I think it's worth noting the response explicitly said the paid
         | API model hasn't been changed, so ChatGPT could have been
         | changed through many different ways outside the core model
        
         | jstarfish wrote:
         | No, this is peak corporate gaslighting, _Open_ AI being the
         | paragon of integrity and all. Nobody should trust a goddamn
         | thing that comes out of their mouths-- or their product.
         | 
         | It's no coincidence the flat-fee service is visibly crippled,
         | while _per-request API users_ are not reporting any difference.
         | (edited)
         | 
         | Ignoring everyone here, look at the other "hacker" groups--
         | like the jailbreaking community. They have a lot to say about
         | recent changes that coincide with their hacks not working. All
         | of a sudden, with OpenAI supposedly changing nothing, technical
         | bypasses just stopped working. OpenAI changed nothing, so this
         | _must_ be deus ex machina.
         | 
         | I'm not even jailbreaking it but the results I get for simple
         | code requests through the web UI have become unusable garbage.
         | It puts less effort into responses than an unpaid-and-
         | overworked intern. Others here report the same. The current
         | "iteration" seems hellbent on terminating conversations as
         | quickly as possible once they stray from the explicit scope of
         | the original topic and is almost as hostile to fixing its own
         | errors as it is to endorsing eugenics, whereas in the beginning
         | it would humor every idle thought I threw at it in long
         | conversations. You can literally see this reflected in the logs
         | they forced retention of. It's acting like a customer support
         | rep desperately trying to end a call before it exceeds a call-
         | time quota.
         | 
         | This particular current workflow seems like it lends itself to
         | better organization of training data-- conversations are what
         | the title says they're about. It also seems like it lends
         | itself to anti-jailbreaking because pretexting it with
         | irrelevant information forces a change of scope-- and a summary
         | termination.
         | 
         | But _OpenAI says_ they changed nothing, so rather than _one guy
         | lying_ without consequence, a community of professionals using
         | and abusing the tool must _all_ be victims of rhetorical
         | fallacy? Nobody 's qualified to reverse-engineer corporate
         | bullshit anymore without being infantilized...
         | 
         | (Liars running a black box-- what could possibly go wrong? We
         | need regulation!)
        
           | Centigonal wrote:
           | > It's no coincidence the "free" service is visibly crippled,
           | while paid API users are not reporting any difference.
           | 
           | I thought GPT-4 wasn't available to free users?
           | 
           | Keep in mind the tweet is specifically about the OpenAI API -
           | They might be updating ChatGPT without telling anyone
           | (although they have release notes)
        
             | jstarfish wrote:
             | "Free" was the wrong word to use, I meant to say
             | "unlimited."
             | 
             | (edit: even that's not right. I think the verbiage is more
             | accurate now.)
        
           | matthewmacleod wrote:
           | _It 's no coincidence the "free" service is visibly crippled,
           | while paid API users are not reporting any difference._
           | 
           | GPT-4 is not available to free users, which really calls the
           | pretext of this comment into question. What is being alleged
           | to have happened, and why is this response unreasonable?
           | 
           | I'm not a very heavy GPT-4 user but I do usually use it once
           | a day or so - it doesn't appear to have changed noticeably
           | but I'm not paying close attention.
        
             | jstarfish wrote:
             | It was badly-worded on my part. And I stand corrected-- I
             | just saw all the nested tweets where people are complaining
             | about the API too!
             | 
             | So many people are noticing a degradation in speed and
             | quality of responses, and not in isolation. Rather than
             | acknowledge this, the question is rephrased to one the
             | responder can rightfully deny-- and suggest no changes have
             | taken place without explicitly saying as much. You normally
             | only see this sort of sliminess from politicians and
             | executives on the witness stand.
             | 
             | > Is anyone else noticing significantly downgraded GPT-4
             | capabilities today? Seems like OpenAI updated the model,
             | and results aren't as good as before. [mentions API in a
             | child comment]
             | 
             | > The API does not just change without us telling you. The
             | models are static there.
             | 
             | > This is good to know. That means GPT-4 has been static
             | since March right? 0314?
             | 
             | > Correct
             | 
             | Never ask questions to which one word suffices as an
             | answer.
             | 
             | Collective confusion in the thread suggests _something_ has
             | changed, but the most OpenAI will attest to is that the
             | _API_ is unchanged and the _models_ are static. And this
             | may well be true, but rather than admit  "...but we were
             | fucking with the middleware/parameters" they took a firm
             | position on a strawman argument and ignored everybody who
             | followed with more-direct questions. Except this guy:
             | 
             | > I've noticed inconsistency with certain prompts
             | performance. Is that just the non-deterministic nature of
             | the API?
             | 
             | > Yes
             | 
             | Oh, ok. It's because the fucking _API_ is _non-
             | deterministic_ that code that has worked both reliably and
             | predictably for everyone now runs like shit for everyone.
             | For fuck 's sake, you can get better answers from a Magic
             | 8-Ball. This guy even made the mistake of presenting an
             | answer he'd believe for the respondent to feed into. He
             | might as well have asked if inconsistent performance was
             | because of the war in Ukraine.
             | 
             | "Logan.GPT" must moonlight as a fortune teller. He's only
             | responding to people foolish enough to ask the wrong
             | questions.
        
             | shmoogy wrote:
             | ChatGPT4 feels (is?) typically faster, and the results feel
             | ... above 3.5 turbo but below what 4 was. The API seems
             | exactly the same to me, but you cant do an apples to apples
             | test due to variance between replies.
             | 
             | I dont really doubt they scaled the chatGPT4 model down a
             | little to try to save costs with the plugins and increased
             | usage
        
               | californical wrote:
               | Yeah I'm having the same experience. The replies have
               | gotten faster in the last few days, at the expense of
               | quality.
               | 
               | It's noticeable because I used to be able to read each
               | word as it was printed from the GPt-4 output, but now it
               | goes far too fast for me to keep up with.
               | 
               | Kinda sucks to not have transparency at all into the
               | black box, they've strayed so far from "open" at this
               | point that it's comedic
        
           | chasd00 wrote:
           | > Liars running a black box-- what could possibly go wrong?
           | We need regulation!
           | 
           | "liars running a black box" is the definition of government
           | regulation sheesh.
        
           | flir wrote:
           | Today I had ChatGPT 4 (Not GPT 4) correct itself mid-
           | response. I was asking it something very simple about
           | regexes:
           | 
           | > How about matching 'a' as the second character of a string
           | only?
           | 
           | It responsed with the wrong regex plus a bunch of explanatory
           | junk:
           | 
           | > '^a.'
           | 
           | Then halfway through the explanatory junk, it corrected
           | itself like this:
           | 
           | > Apologies for the confusion in the first response, the
           | correct regular expression should be '^.a' for matching 'a'
           | as the second character of a string:
           | 
           | And kept on with the (now correct) explanatory junk.
           | 
           | All in a single response. I've certainly never seen that
           | before (if someone has, please weigh in). Maybe the model
           | hasn't changed, but the pipeline has? Like... there's a
           | second model trying to correct the mistakes of the first,
           | maybe? (timings are probably wrong for that, but something
           | like that)
        
             | coffeebeqn wrote:
             | There's also a new button "continue generating" so they are
             | changing some things
        
             | theturtletalks wrote:
             | ChatGPT 4 seems to struggle with context changes within the
             | same thread now. I asked it about some parks near me and
             | then switched to asking about code. It just kept responding
             | about the parks even when I corrected it. In some cases, it
             | would merge the park question and the code question. I had
             | not seen that with the free or Plus version until a few
             | days ago.
        
             | inciampati wrote:
             | Interesting. Bing chat does this when you get it to talk
             | about naughty things. My personal favorite is convincing it
             | to make weird art out of quotes from the Tay chatbot. It
             | will write them until it says a prohibited word or phrase,
             | or touches some forbidden part of it's state space.
        
               | flir wrote:
               | Hm. Maybe they've backported some nanny code from Sydney?
        
               | iliane5 wrote:
               | AFAIK it's pretty standard practice not to expose the
               | "raw" LLM directly to the user. You need a "sanity loop"
               | where user input and the output of the LLM is checked by
               | another LLM to actually enforce rules and mitigate prompt
               | injections, etc.
        
             | chaxor wrote:
             | I have gotten that quite a few times now and it does seem
             | to be new. In my case though, it acted as if it would
             | change something, but then didn't, and got stuck in a loop.
             | I got it enough times that it seemed to happen more often
             | when the output needed to be quite long, so I had assumed
             | (could be wrong) that when the output was nearing a certain
             | limit, they had implemented something that would start a
             | new output, such that the interaction of "please continue
             | the script" would not be necessary.
             | 
             | Unfortunately though, it would start over with the script
             | completely, and then get stuck in a loop until it broke,
             | this not being able to even save the response at all.
             | 
             | Something definitely feels like it changed, but I suppose
             | it could just be more use of the system.
             | 
             | Another possibility is related to some strange issues with
             | hardware and balancing etc, despite not changing software
             | or params. It's strange, and it _shouldn 't_ happen, but
             | sometimes things change from simple batching and balancing,
             | or lower level hardware related things, which are very
             | annoying to debug.
        
       | dontupvoteme wrote:
       | I just discovered when playing around with translations that
       | there is some hidden filter/killswitch that immediately stops the
       | generation of the opening of some books. It doesn't matter if the
       | invoking prompt is to recite the book opening paragraph, or to
       | translate it from a foreign language to English, or what.
       | 
       | It's not RLHF induced because it works via API and it only
       | triggers in English, but sure enough try to get it to output
       | 
       | >Call me Ishmael. Some years ago--
       | 
       | >"It was the best of times,
       | 
       | I guess this might get you flagged (there is no alert to the user
       | that this filter kicked in, and it will output it in any other
       | language, and it works in the API) so I'm hesitant to play around
       | with it more, but it's very strange - especially as these are
       | long since in the public domain.
        
       | Ozzie_osman wrote:
       | He said "the API". He didn't say anything about the Chat version,
       | which seems to have more protections that may be on top of the
       | model and not embedded into it.
        
       | textninja wrote:
       | Must be in the trough of disillusionment. The tech is truly
       | useful though so we'll reach the plateau soon enough.
        
       | josecyc wrote:
       | I agree I've noticed I need to be quite specific now, it won't
       | realize bugs in the code unless I tell it so. A head to head
       | comparison of the different versions is needed to validate this.
        
       | elboru wrote:
       | I don't get why is this being discussed, isn't it easy to go
       | check old conversations and try to replicate them today?
        
         | renewiltord wrote:
         | You need to have used the API and set temperature zero, but if
         | you have that historical data you can reasonably test and see.
        
         | ShamelessC wrote:
         | You would think so, but not a single person complaining has
         | provided convincing proof of the degradation.
        
           | sebzim4500 wrote:
           | I went through a bunch of old prompts and the response was
           | definitely worse now than it was then.
           | 
           | Since it's not determinstic though, it's hard to draw
           | conclusions. Especially since the sample size was 10ish.
        
         | dgellow wrote:
         | GPT isn't deterministic
        
       | karmasimida wrote:
       | The model didn't change but doesn't mean the inference didn't
       | change.
       | 
       | Without going specifics it is meaningless for the discussion
        
         | heliophobicdude wrote:
         | The models for the api's could have not changed but the web
         | app's might have too
        
       | armchairhacker wrote:
       | I doubt the March 14 model is any different, since it's versioned
       | and I don't see why OpenAI would want to change it behind the
       | scenes.
       | 
       | Also, companies are evaluating GPT-4 to determine whether they
       | want to pay for it, so OpenAI has an strong incentive to not
       | downgrade at least the API.
       | 
       | I believe the May 5 model is _different_ , at least in the chat
       | interface, because it's fine-tuned to detect jailbreaks and the
       | temperature/other hyper-parameters may have changed. And I can
       | imagine this fine-tuning making the model less creative and worse
       | at solving analytical tasks.
       | 
       | Personally I haven't noticed any change, except in my own
       | awareness. Sometimes GPT4 gets very hard prompts right, and
       | sometimes it gets simple problems wrong. So it's not hard to see
       | how people can form biased opinions from selective attention or
       | just luck.
        
       | TradingPlaces wrote:
       | The model may be the same, but the chatbot is not. They have made
       | the responses shorter to save inference expenses, which are huge
       | in a model the size of GPT-4.
        
       | haolez wrote:
       | I use GPT-4 a lot and something has definitely changed in my
       | recent experience. It's simply worse.
        
       | dangerlibrary wrote:
       | A couple observations from attempting to use both GPT-3.5 and
       | GPT-4 via the web interface for coding tasks:
       | 
       | - The model's ability to respond accurately drops drastically
       | when asked questions of the form "is there a different way to
       | accomplish X, using Y?" or "is there a way to accomplish X that
       | runs in O(log(n)) time instead?" Example: I wanted to upsert an
       | integer value using a SQLite db using "INSERT ... RETURNING..."
       | ChatGPT repeatedly told me that sqlite doesn't support
       | "RETURNING" (it does, since March 2021). It insisted I would need
       | two DB round trips from my application to accomplish this. When
       | asked "can this be done in one round trip, instead?" it
       | repeatedly wrote code that would return the number of rows
       | modified instead of the integer column value.
       | 
       | - ChatGPT's limited standard library knowledge means that the
       | solutions it produces, even when correct, are often lower-level
       | and less idiomatic. Problems that would be trivially solved with
       | e.g. a Java.String.replaceAll or .codePointCount will instead
       | loop over each character, often splitting the string into an
       | intermediate array and implementing special cases for first/last
       | character edge cases. The code winds up being mostly correct, but
       | also (for lack of a better word) weird. No human I've ever worked
       | with would do things the way ChatGPT sometimes does, which means
       | the code will likely be much harder to maintain and debug over
       | time.
        
         | ch4s3 wrote:
         | > are often lower-level and less idiomatic
         | 
         | I would go as far as to say it's adept at producing technically
         | working spaghetti.
        
         | jstarfish wrote:
         | > ChatGPT repeatedly told me that sqlite doesn't support
         | "RETURNING" (it does, since March 2021).
         | 
         | Its dataset cuts off around 2021. There's a little footer
         | message warning you not to expect knowledge of recent events.
        
           | dangerlibrary wrote:
           | It cuts off in September 2021, supposedly - that's why I
           | pointed out the month of the relevant sqlite release.
        
       | lee101 wrote:
       | [dead]
        
       | andrewstuart wrote:
       | GPT 3.5 is a much better coding assistant.
        
       | pleb_nz wrote:
       | I've found it's gotb eyes and have given up using it for tasks
       | more often than I used to.
       | 
       | For copilot, I no longer get multi line complete suggestions and
       | it's really slow to deliver single line suggestions and they're
       | more often incorrect. I need to dig into it, but it's definitely
       | degraded further and I don't know if it's just my environment or
       | a wider issue. I need to dig in and figure out - is anyone else
       | experiencing these things?
        
       | two_in_one wrote:
       | ChatGPT Plus shows the version. It was May 12, if I remember
       | correctly, then it changed to May 24. Which probably means there
       | were some changes. If not in GPT-4 itself, then in pre or post
       | processing. They should have some safety filters at the end, I
       | think.
        
       | phillipcarter wrote:
       | Posted this in yesterday's thread, but once again I think this is
       | just people feeling the magic wear off. People have poked around
       | a lot more and found the flaws while also trying to get it to do
       | real-world tasks. That wasn't true when it first came out.
       | 
       | It's fine, this tech has never been magic anyways, won't be
       | replacing all our jobs, won't take over the world, etc. It's
       | still awesome for what it is.
        
       | madrox wrote:
       | This doesn't pass the smell test for me. _Something_ has changed.
       | Maybe the model hasn't changed, but something somewhere is making
       | output worse.
       | 
       | I'm honestly more concerned if OpenAI doesn't even realize it.
       | Nothing is more infuriating as a user than convincing the
       | developer your bug actually does exist. It speaks to poor
       | monitoring, testing, and tooling.
        
       | sebzim4500 wrote:
       | I understand why people don't read 10,000 word articles, but this
       | is a 15 word tweet. Would it really kill people to read it before
       | commenting? He is very explicitly talking about the API, which
       | uses a different model than the UI.
       | 
       | The recent discussion was about the degradation in the UI model.
        
       | emptyfile wrote:
       | [dead]
        
       | jacquesm wrote:
       | To what degree is the pre and post processing of the chat client
       | the source of the confusion rather than GPT-4 itself?
        
       | [deleted]
        
       | rat9988 wrote:
       | I thought the API was versioned, though it doesn't mean the new
       | versions aren't worse. And he doesn't talk about the model used
       | in chat gpt. I'm a bit skeptical that this answers exactly what
       | we were worried about. Or maybe we were not all worrying exactly
       | about the same thing.
        
       | nashashmi wrote:
       | Is there a way to use GPT-4 directly? Instead of the muzzled
       | ChatGPT that is supposed to give answers authorities and
       | opinionators deem appropriate?
        
         | heliophobicdude wrote:
         | Assuming you mean the base model, no not right now. GPT-4 base
         | model is not public.
        
       | rkapsoro wrote:
       | Perhaps the evolution of this story is an interesting example of
       | confirmation bias?
        
       | petabite wrote:
       | The amount of times per day I have to click "Stop generating" and
       | then say "No!" has definitely increased
        
       | geraldwhen wrote:
       | Whatever happens to chat gpt, ai image generation is incredible.
       | It's so incredibly powerful that I don't think it's ever going
       | away
        
         | saiya-jin wrote:
         | If you are doing some indie gaming, it can save tons of money
         | on asset generation, you either ramp up and finalize it
         | yourself or hire somebody but for fraction of the time. Ditto
         | for web design, why would anybody but big, calcified megacorps
         | with their ridiculous processes go and buy some stock images if
         | you can get exactly what you want, in any imaginable style?
        
       | skilled wrote:
       | ChatGPT was a massive dopamine hit, particularly for people in
       | areas like development. It was a tremendous release that has
       | definitely laid out some new tracks for the future, especially on
       | the web. I myself have found to be using ChatGPT a lot less
       | recently.
       | 
       | I got the GPT-4 API access and then I realized that I can't
       | really use it for anything super major because I can't afford it,
       | it is ridiculously expensive if you consider that you have to pay
       | for all the failed requests, the wrong information or the wrong
       | context also. Instead, I have written a bunch of Python scripts
       | that do a select few tasks for me and I have my terminal open
       | 24/7 anyway.
       | 
       | As for the topic at hand, I have _definitely_ noticed a lot more
       | disclaimers in the UI. I don't get it from the API at all, in 6
       | months that I have been using the API - I've gotten _one_
       | disclaimer.
       | 
       | In the ChatGPT UI - I get them a lot. "Remember this", "Remember
       | that", "Always look up the information" and things like this. I
       | mean if it wasn't happening I would know because I have been a
       | power-user pretty much all this time...
        
       | RecycledEle wrote:
       | If ChatGPT has been static since March (of 2023) the. Why does my
       | version always change? Right now I'm running May 24.
       | 
       | FWIW, I think it has improved.
        
       | dang wrote:
       | Recent and related:
       | 
       |  _Ask HN: Is it just me or GPT-4 's quality has significantly
       | deteriorated lately?_ -
       | https://news.ycombinator.com/item?id=36134249 - May 2023 (711
       | comments)
        
       | jimsimmons wrote:
       | He's talking about the API. Not the web client
        
       | neonsunset wrote:
       | There is No War in Ba Sing Se.
        
       | dubcanada wrote:
       | Is it static? If I ask it the same question I asked it a few
       | months ago it gave me a completely different answer. Is that
       | because of some additional context above? Should we be starting
       | fresh chats sooner?
        
         | schrodinger wrote:
         | There's _is_ a degree of randomness in its content generation,
         | you can see it in action by hitting regenerate. It's a
         | statistical text predictor that iteratively selects the most
         | like my word after the ones it's selected so far, but when
         | there are a few good candidates, it'll choose one at random.
         | 
         | However, the model is static (it'll present the same candidate
         | word list until retrained), which is what you may have heard.
         | But the way a response is generated introduces the randomness.
         | 
         | Note: I forget the parameter, but if you get the direct API
         | access you can turn off this and get consistent answers.
        
         | carabiner wrote:
         | It's never been static and there's always been randomness when
         | you regenerate an answer. There is no control of a random seed,
         | so you can get a different answers for the same prompt.
        
           | guraf wrote:
           | You're right of course that some randomness is inherent, but
           | you can adjust the "temperature".
           | 
           | With a low enough temperature you get essentially the same
           | output every time, with a at most just minor words swapped.
        
         | acomjean wrote:
         | It's not stactic. That's the thing about ai, you ask it the
         | same thing it gives you a different answer/different image
         | every time.
        
           | drexlspivey wrote:
           | That's not really an issue with AI in general, you can make
           | it deterministic by tuning some parameters if you want
           | (temperature in chatGPT, passing a seed in midjourney etc)
        
       | AtNightWeCode wrote:
       | Are we chatting directly with the model? Maybe the interface has
       | changed. With long term use the likelihood of hitting edge cases
       | is higher and maybe that is a cause as well for what users are
       | seeing. People probably ask more vague questions over time. I
       | might have done that.
       | 
       | I have never experienced the amnesia problem in v3.5 though, that
       | v4 clearly has. Just repeating incorrect answers that you ask it
       | not to give. I did not have access to v4 in march so I can't do
       | that comparison.
        
       | jiggawatts wrote:
       | I look at everything someone from OpenAI says as if a politician
       | is saying it. Sam Altman especially is fond of statements that
       | are deceptive but technically true. His employees appear to be
       | following his lead.
       | 
       | GPT 4 isn't ChatGPT 4, which is what most people use.
       | 
       | There is also the "system prompt", which is also likely to be
       | changing but not part of GPT 4.
       | 
       | Etc...
        
         | Sunhold wrote:
         | What is "ChatGPT 4"? ChatGPT Plus can use GPT-4 and free
         | ChatGPT uses gpt-3.5-turbo.
        
           | californical wrote:
           | ChatGPT isn't just using plain GPT-4, it's using a specific
           | pre-prompt and _possibly_ additional fine-tuning (which I'm
           | not sure about) compared to the plain GPT-4, available
           | through API.
           | 
           | Someone saying ChatGPT4 is just saying in shorthand, the
           | ChatGPT model based off GPT-4
        
         | sebzim4500 wrote:
         | I don't see this as being political doublespeak. He is very
         | clearly just talking about the API.
        
           | CSMastermind wrote:
           | Right but people are complaining about their experiences with
           | ChatGPT.
           | 
           | If you say, "ChatGPT has gotten noticeably worse" and they
           | respond "nothing about our APIs have changed" then it would
           | be reasonable to interpret that as them saying that nothing
           | about ChatGPT has changed.
           | 
           | When in reality many things might have changed about ChatGPT.
        
             | furyofantares wrote:
             | The tweet is a response to an API user saying GPT-4 got
             | worse, and the tweet is clearly talking about the API and
             | the model - the context of the HN megathread mostly talking
             | about ChatGPT isn't present in the tweet that's being
             | replied to.
             | 
             | There's a lot of confusion around this, but nothing appears
             | to be caused by doublespeak to me. Confusion about which
             | GPT/ChatGPT anyone is reporting about has been pretty
             | ubiquitous for a while.
        
             | stavros wrote:
             | He says "the models are static", which means that, if you
             | keep using the same model, you will keep getting the same
             | results.
        
           | pclmulqdq wrote:
           | Which is not what anyone else is referring to when they say
           | that GPT-4 has gotten worse.
        
           | smcin wrote:
           | Then the HN title is grossly misleading.
        
       | bugglebeetle wrote:
       | Company who has repeatedly lied in public: there's nothing to see
       | here!
        
       | greatpostman wrote:
       | Recent research showed that RLHF/censoring the model hurts the
       | performance of the model. This is intuitively obvious, censorship
       | isn't real, the data is (moral issues aside). So it hurts the
       | integrity of the weights. The future is open sourced uncensored
       | models, capitalism will demand the high performance. There's a
       | huge discussion on it on Reddit right now:
       | 
       | https://www.reddit.com/r/MachineLearning/comments/13tqvdn/un...
        
         | haolez wrote:
         | Maybe humans are like that as well? :)
        
           | greatpostman wrote:
           | Lots of evidence pointing in that direction. Suppression of
           | logic (censorship) leaks into other areas of cognition
        
         | Spivak wrote:
         | People keep mixing up alignment as in do what I actually asked
         | and weird moralizing.
         | 
         | You will pay the alignment tax no matter what if you want a
         | model that actually does things. It's the reason why Llama
         | models do better with stream of thought prompting where you can
         | ask GPT questions and it won't try completing the question.
         | 
         | You can, and probably should, align some semblance of morality
         | and ethics _for user-facing models that do things_. It 's
         | customer service voice but for AI. If what you want out of a
         | model is a vague mirror of humanity or maximum smarts at the
         | cost of needing more detailed prompts then yeah, this stuff
         | probably annoys you.
         | 
         | Finding a balance between a model that gives the outputs humans
         | actually want and the pull that training data has on the rest
         | of the model making it stray from the theoretical "best" output
         | is hard.
        
         | mschuster91 wrote:
         | > The future is open sourced uncensored models, capitalism will
         | demand the high performance.
         | 
         | No. The future is ML built on more than just Wikipedia, Github
         | code and Reddit hot takes. Or, sarcasm aside, GIGO: garbage in,
         | garbage out - when you don't take care what datasets you train
         | _any_ kind of model on, you 're bound to get some surprises if
         | something unexpected comes along. Be it the infamous "racist
         | soap dispenser" or the Google (?) image classifier that made
         | the rounds here just a day or two ago which had a safeguard
         | because it kept confusing Black people with gorillas.
         | 
         | Preventing this kind of harmful content isn't censorship - it's
         | after-the-fact compensation for bad training (and the tendency
         | of 4chan and other trolls to use discriminatory content
         | generation as a weapon, just remember what they did to
         | Microsoft Tay).
         | 
         | That is where the money is: _curated_ training datasets.
         | Everyone and their dog can train a ML model from scratch, all
         | you need is money and cloning a few more-or-less-broken Github
         | repositories for that. But acquiring a high-quality dataset as
         | a foundation? That costs _real_ money to create.
        
           | fragmede wrote:
           | Yes, so just like Apple's real value is from Foxconn's
           | manufacturing factories which they don't own, OpenAI's real
           | value is from Sama, the Kenyan firm that did the RLHF
           | annotating.
           | 
           | I'm wondering what the legal agreement is for Google
           | Classrooms. In elementary schools across the world, kids
           | write 3rd grade essays in history class or whatever and the
           | teacher grades them. Those essays and grades are all in a
           | database. Does Google have the opportunity to train Bard
           | across that dataset?
        
           | chasd00 wrote:
           | i'm not in the AI world but i've been curious how the level
           | of effort compares between inventing and writing the model vs
           | finding, tagging, and curating the training. Is creating the
           | training data analogous to inventing a complete schooling
           | curriculum from scratch? I bet that takes a long freaking
           | time.
        
         | two_in_one wrote:
         | It shouldn't be limited to censoring. Fine tuning is done on
         | small set of some domain specific samples. Surely that results
         | in gradual 'forgetting' of the already learned data from big
         | set. If run for long time model will relearn on the tuning set,
         | and overfit most likely.
        
       | chatmasta wrote:
       | My experience using GPT-4 for coding is that it's got the
       | knowledge and skill of a senior engineer, with the high
       | maintenance of a junior engineer. You can get it to output small
       | sections of quality code, but the amount of prodding it takes to
       | piece it all together means you may as well have spent the time
       | writing it yourself. But the future of GPT as a coding assistant
       | is definitely bright. It just needs more chaining, so I feel less
       | like _its_ assistant, asking it to come up with prompts and then
       | pasting them back to it after iterating on the code.
        
         | Our_Benefactors wrote:
         | It's still way faster than I could hope to be. It's a true
         | refactoring machine with 0 mental effort. "Here is this code
         | block, make it do this other thing, add this feature and
         | argument and error condition" etc. It augments the amount of
         | output that's possible to achieve at any seniority level.
        
         | furyofantares wrote:
         | GPT-4 has incredible breadth and no depth.
         | 
         | It's an excellent complement to a strong programmer, who will
         | have incredible depth -- and may have breadth relative to other
         | coders, but will be very narrow in the whole space of
         | programming.
         | 
         | Just don't use it for things you're already an expert on,
         | except perhaps for starting out / bypassing boilerplate.
        
         | rngname22 wrote:
         | It's Einstein with dementia/alzheimers.
        
         | throwuwu wrote:
         | They just put out a blog post about how they're doing exactly
         | that by rewarding chain of thought style reasoning.
        
         | jorblumesea wrote:
         | Generative AI will not be able to do that in any true sense.
         | Not the way you're thinking about it.
        
         | Mizoguchi wrote:
         | Agree. However I noticed an unintended benefit of using GPT for
         | coding is that it helps you think carefully about the problem
         | you are trying to solve when writing the prompt.
        
           | AtNightWeCode wrote:
           | That is why most developer jobs are safe. Tasks from
           | stakeholders often contains only a title.
        
           | flir wrote:
           | Rubber ducking in the age of WFH is definitely a use I've
           | found for it.
        
           | chasd00 wrote:
           | i don't use it that much for code but when i do it's for a
           | specific function or method. I describe the inputs, the logic
           | i want, and the outputs i need. So basically i have it worked
           | out in my head I just let chatgpt type it out for me. Any
           | mistakes it makes are pretty easy to catch when using it this
           | way.
        
           | behnamoh wrote:
           | "Language as a tool of thought" is a technology we often
           | neglect we are equipped with.
        
             | dkersten wrote:
             | This is essentially what rubber ducking is.
        
         | moffkalast wrote:
         | > so I feel less like its assistant
         | 
         | Ah so I'm not the only one who feels this way lol. It's as
         | brilliant as it is dumb.
        
         | naiv wrote:
         | For me it is not even the code per se but how it helps me with
         | naming database tables and properties when I describe the
         | broader scope.
         | 
         | The future is more than bright, especially with a new model
         | that is more current.
         | 
         | Offtopic but just today I was aksing it for the most current
         | versions it knows of right now:
         | 
         | Python: 3.9
         | 
         | JavaScript: ECMAScript 2021 (ES12)
         | 
         | Java: 16
         | 
         | C++: C++20
         | 
         | C#: 9.0 (.NET 5.0)
         | 
         | Ruby: 3.0.0
         | 
         | Swift: 5.4
         | 
         | Go: 1.16
         | 
         | Rust: 1.51.0
         | 
         | TypeScript: 4.3
         | 
         | PHP: 8.0
         | 
         | Kotlin: 1.5.0
         | 
         | Scala: 3.0
         | 
         | R: 4.0.5
         | 
         | Perl: 5.32
         | 
         | And frameworks:
         | 
         | Python:
         | 
         | Django: Version 3.2
         | 
         | Flask: Version 2.0.1
         | 
         | Pyramid: Version 2.0
         | 
         | TensorFlow: Version 2.6.0
         | 
         | PyTorch: Version 1.9.0
         | 
         | JavaScript:
         | 
         | Node.js: Version 16.9.1
         | 
         | Express.js: Version 4.17.1
         | 
         | React.js: Version 17.0.2
         | 
         | Angular: Version 12.2.4
         | 
         | Vue.js: Version 3.2.6
         | 
         | Java:
         | 
         | Spring Framework: Version 5.3.9
         | 
         | Hibernate: Version 5.4.32.Final
         | 
         | Struts: Version 2.5.26
         | 
         | PHP:
         | 
         | Laravel: Version 8.54.0
         | 
         | Symfony: Version 5.3.6
         | 
         | CodeIgniter: Version 4.1.4
         | 
         | Ruby:
         | 
         | Ruby on Rails: Version 6.1.4
         | 
         | Sinatra: Version 2.1.0
         | 
         | C#:
         | 
         | .NET Core: Version 5.0
         | 
         | ASP.NET: Version 5.0
         | 
         | Entity Framework Core: Version 5.0
        
           | chasd00 wrote:
           | i just asked it "what's the output of python -v" and the
           | little sample code reported python 3.9.2.
        
             | flir wrote:
             | I got:                   Python 3.8.5 (default, Jan 27
             | 2021, 15:41:15)         [GCC 9.3.0] on linux         Type
             | "help", "copyright", "credits" or "license" for more
             | information.
             | 
             | And a lot of junk text.
        
               | hereonout2 wrote:
               | Ha! So far it seems four posters here have asked similar
               | questions and got four different python versions.
               | 
               | I don't know if we can treat the gpts in this way and
               | expect reliable answers. We just get pretty good answers
               | right up to the point that we don't.
        
           | RadiozRadioz wrote:
           | Asking it directly is a good way to get a response containing
           | version numbers that look correct, but not a good way to
           | actually indicate which versions it's been trained on.
           | 
           | A better method would be to look at the language features
           | it's using and infer from there. Or, better still, look at
           | which versions were out when the training data was collected
           | (which I believe is September 2021, don't quote me on that).
        
             | grenoire wrote:
             | Is there a risk of it hallucinating what new features
             | _could_ do instead of basing information on having seen
             | them before?
        
               | RadiozRadioz wrote:
               | Yes, good point, I suppose we would have to weigh the
               | probability of that against it fudging a number. I would
               | assume inventing a new feature is harder, but who's to
               | say; all this LLM stuff comes down to probability.
        
           | o1y32 wrote:
           | It does not actually mean much. These chatbots can easily
           | output JavaScript using var where it absolutely should not.
           | It being aware of the latest standards does not mean it can
           | properly use the new syntax or features -- it emits a form of
           | whatever crappy/legacy code it was trained on.
        
           | dontupvoteme wrote:
           | Oh that's interesting, I didn't think about asking it
           | directly about what version it thinks is current. I wonder if
           | it can hallucinate here?
           | 
           | It's really irksome when it tries to use functions which were
           | renamed or removed. This could be detected automatically (and
           | possibly remapped in some cases)
           | 
           | I got mostly the same versions as you, both on chatgpt3/4,
           | using english.
           | 
           | If you ask it in other languages the minor version also seems
           | to change often.
           | 
           | Finnish gives you dates, and code-davinici-edit-001 is 1
           | major release behind on almost everything.
           | 
           | Tassa on luettelo ohjelmistojen viimeisimmista vakaiden
           | versioista:
           | 
           | Python: 3.9.5 (30.4.2021) JavaScript: ECMAScript 2021
           | (23.3.2021) ECMAScript: ECMAScript 2021 (23.3.2021) Java: JDK
           | 17 (28.9.2021) C++: C++20 (20.2.2020) C#: .NET 6 (8.11.2022)
           | Ruby: 3.0.2 (24.8.2021) Swift: Swift 5.5 (20.9.2021) Go: 1.17
           | (16.8.2021) Rust: 1.54.0 (27.5.2021) TypeScript: 4.4
           | (28.7.2021) PHP: 8.0.9 (29.7.2021) Kotlin: 1.5.31 (26.8.2021)
           | Scala: 2.13.6 (17.2.2021) R: 4.1.0 (18.5.2021) Perl: 5.34.0
           | (30.5.2021) ... Sinatra: 2.1.0 (10.4.2021) .NET Core: 6.0
           | (8.11.2022) ASP.NET: 5.0.10 (19.8.2021)
        
         | dontupvoteme wrote:
         | The main thing it's lacking is a client-side merge/error-
         | detection engine which combines all of the previous outputs.
         | 
         | Anything over ~100LoC and it gets sloppy including it all in
         | the code block. Which is fine, because I don't need it in your
         | context window if that bit is working and you're not relying on
         | it.
         | 
         | Though I have gotten outputs up to nearly 250 lines from it..
         | 
         | Automatically feeding back code details + Traceback messages
         | also fixes errors decently often enough.
        
         | alfalfasprout wrote:
         | Those aren't very good senior engineers then lol. GPT-4 has the
         | coding skills of an overconfident junior engineer and little
         | else.
         | 
         | I agree the future of GPT-X models as coding tools is bright.
         | But for actually doing engineering work outside coding (or even
         | delicate changes in an existing code base) much less so.
        
           | jsight wrote:
           | I'd love it if junior engineers organized code as well as
           | ChatGPT. I'd also love it if all engineers stopped mixing
           | spaces and tabs.
        
             | kristopolous wrote:
             | That can mostly be fixed with conventional prettify tools.
             | No modern ai required
        
             | fragmede wrote:
             | Do you have a style guide?
             | 
             | ChatGPT, on spaces vs tabs: https://chat.openai.com/share/b
             | 2b0be49-54f9-4f73-a753-3edbd1...
        
             | [deleted]
        
             | alfalfasprout wrote:
             | Add a linting step to your CI to catch this stuff?
        
               | jsight wrote:
               | Maybe I could get chatgpt to write those pipeline steps
               | for me. :)
        
               | oarsinsync wrote:
               | You can and you should. It's this kind of busywork I farm
               | out to GPT these days and it's great at it. I suck at it,
               | so it saves me an hour to focus on what I'm good at!
        
       | Mizoguchi wrote:
       | Unrelated to the model itself but infrastructure, yesterday was
       | unusable for me, getting random "too many requests sorry bout
       | that" errors. I think 1/4 of the requests during a 3 hr period
       | didn't make it through. Imposible to build anything beyond
       | experimental stuff on top of a service so unreliable. Haven't
       | tried it through Azure yet, I wonder if it is any better?
        
       | amelius wrote:
       | Instead of complaining, why not show the benchmarks?
       | 
       | Like: first it scored 83, now it scores only 42 (or whatever).
        
         | sebzim4500 wrote:
         | What benchmarks? We have no way to go back in time and run the
         | old ChatGPT accessble models though benchmarks.
         | 
         | And to my knowledge, no one has copy pasted an entire benchmark
         | into the ui in order to run it, they just use the API.
        
         | ShamelessC wrote:
         | the commenters of HN are simply so smart that their hunch
         | clearly holds more weight than scientific rigor.
        
       | PaulHoule wrote:
       | People have just fallen out of love with it and are realizing how
       | it really isn't that good.
       | 
       | Sorry "prompt engineers" but papers on arXiv show that when you
       | give it fairly sampled problems it struggles to get the right
       | answer more than 70-80% of the time. When you are under its spell
       | you will keep making excuses but when you are looking at it
       | objectively you'll realize the emperor is naked.
       | 
       | If you give it very conventional problems it seems to do better
       | than that because it is a mass of biases and shortcuts and of
       | course it will sentence Tyrone to life in prison because it's a
       | running gag that "Tyrone is a thug"... That's how neurotypicals
       | think and no wonder why many of them think ChatGPT is so smart...
       | It mirrors them perfectly.
        
         | vsareto wrote:
         | >When you are under its spell you will keep making excuses but
         | when you are looking at it objectively you'll realize the
         | emperor is naked.
         | 
         | It's not perfect, but you have to admit the emperor is at least
         | wearing a thong. It is the second most intelligent thing on the
         | planet at creating text, even with its flaws. Putting that
         | accomplishment in league with naked emperors is astoundingly
         | biased.
        
           | shrimpx wrote:
           | > you have to admit the emperor is at least wearing a thong
           | 
           | And a lot of people are getting turned on by it.
        
           | PaulHoule wrote:
           | Yeah, maybe it is "wearing a thong" since it's developers
           | were so concerned about social acceptability.
           | 
           | I would say though that it is no mean feat for the Emperor to
           | get away with being naked in public, it takes power and the
           | ability to wield it.
           | 
           | Similarly ChatGPT has a few competences that add up to people
           | perceiving it is able, one of which is the ability to come up
           | with plausible and satisfying answers whether they are right
           | or wrong (often using shortcuts) and another one is getting
           | people to engage with these answers.
        
           | michaelt wrote:
           | By this, do you mean 'the second most intelligent thing'
           | after humans?
        
             | shrimpx wrote:
             | A dog understands body language and subtle cues,
             | participates in social hierarchy and community, has self-
             | reflection, dreams, etc.
        
         | ThrowawayTestr wrote:
         | I used it to help me write an email begging for my job back and
         | it was pretty helpful
        
           | BulgarianIdiot wrote:
           | Just ignore this fella, GPT-4 is massively useful. But it
           | takes intelligence to get intelligence out of it.
        
             | visarga wrote:
             | I think this is our path - become good at working with AI,
             | for jobs or if there are no jobs, to directly support
             | ourselves. Supposing we will be able to use tools and AI to
             | build things and make our own means.
        
               | BulgarianIdiot wrote:
               | This is our path yes. All paths end, eventually, but oh
               | well.
        
           | [deleted]
        
         | BulgarianIdiot wrote:
         | I very much doubt the tweet in question is of someone who loved
         | GPT-4 yesterday, and has "fallen out of love" with it starting
         | today.
         | 
         | For someone so dismissive of "neurotypicals" you did the
         | neurotypical thing and drop a hot take before clicking the
         | link.
        
           | wouldbecouldbe wrote:
           | It's more that the balance in usefulness vs waste of time
           | shifts into the other direction after a few times spending
           | long time getting it right and just doing it yourself anyway.
        
             | BulgarianIdiot wrote:
             | I've spent 20 years explaining my code to a rubber duck and
             | getting great use out of that duck. No one will ever
             | convince me that it's less useful now when the duck
             | actually has read the Internet and has useful suggestions
             | and ideas back to share. Even if it makes shit up on an
             | occasion.
        
               | JohnFen wrote:
               | I don't know. The value of rubber-ducking is that you
               | have to break the problem down into very simple terms to
               | explain it. It's the explaining it part that is magic.
               | That value goes away (or is greatly diminished) if the
               | rubber duck responds with anything other than a request
               | for further explanation.
        
               | BulgarianIdiot wrote:
               | You have to break down a problem to GPT just like you
               | would to a human, and if it misunderstands you or asks
               | you, you'd have to elaborate just the same. It's modeled
               | after us, after all. The rubber duck is a stand-in for a
               | human. GPT also is. But just a vastly better one.
        
               | JohnFen wrote:
               | Of course.
               | 
               | The part I was responding to was this:
               | 
               | > when the duck actually has read the Internet and has
               | useful suggestions and ideas back to share
               | 
               | I think that if it does that, it's legitimately less
               | useful as a rubber duck. I'm not saying it isn't useful
               | -- it's just not useful as a rubber duck anymore.
        
           | PaulHoule wrote:
           | The tweet is an OpenAI employee who is responding to people
           | who have fallen out of love.
           | 
           | I've seen transcripts of people interacting with ChatGPT who
           | were obviously seduced by it and in a very giddy state,
           | having so much fun because ChatGPT was playing an extended
           | "game" with them that it didn't bother them at all that
           | ChatGPT was spouting wrong answers.
           | 
           | A major complaint I've had about the social sphere is that I
           | seem to get the same result if I am 20% right or 50% right or
           | 80% right or 95% right or 99.8%, it is just exhausting and I
           | can never be good enough and I'm frankly envious that people
           | see more of a glimmer of light behind that thing's "eyes"
           | than they do behind mine.
           | 
           | The core thing about neurotypicality isn't so much that they
           | get the wrong answers but that they get the same answer
           | whether it is right or wrong. For a long time I thought the
           | basis of the "language instinct" is a derangement about
           | reasoning with uncertainty that causes the grammar
           | representation to collapse into a low-dimensioned subspace
           | which is learnable with a limited amount of data. I wouldn't
           | be surprised at all if other animals could beat us at rock-
           | scissors-paper or poker if they could understand the rules of
           | the game. The success of LLMs might give us some insight in
           | this area although they are working with so much more data
           | that Chomsky's old "poverty of the stimulus" argument might
           | not apply.
        
         | throwuwu wrote:
         | You should read some more papers. Also fuck off back to reddit
         | with this anti-neurotypical bias.
        
         | mrbombastic wrote:
         | I don't know, I think there is certainly a lot of hype but it
         | is still pretty damn good. I used it every day for a wide
         | variety of coding tasks and more often than not it is correct.
         | As in compiles, produces correct output for inputs, and mostly
         | writes reasonable code. Has it made software dev obsolete like
         | some maximalists said it would? definitely not, but it seems
         | equally hyperbolic to say the emperor has no clothes, it is
         | undeniably a useful tool to me.
        
           | PaulHoule wrote:
           | "More often than not" is compatible with "correct 70% to 80%
           | of the time".
        
             | swores wrote:
             | Your previous comment was a bit confusingly worded, I, and
             | presumably the commenter you just replied to, read it as
             | "more than 70-80% of the time" it struggles; rather than
             | your intended meaning of it struggling to do better than
             | correct 70-80% of the time.
        
         | guraf wrote:
         | [flagged]
        
       | mustacheemperor wrote:
       | In the thread yesterday it was brought up that if you use the API
       | it feels considerably less hamstrung than the ChatGPT client, and
       | this tweet seems to fit the assumption that the ChatGPT product
       | is being tuned or governed differently from the API.
       | 
       | >The API does not just change without us telling you. The models
       | are static there.
       | 
       | This reads to me as specifically indicating the models are not
       | static elsewhere, ie, in ChatGPT.
        
         | jonplackett wrote:
         | It feels less hamstrung on output, but is instead hamstrung by
         | being incredibly slow.
         | 
         | GPT-4 via API will sometimes take 30 seconds + to respond to
         | simple questions without any chat history, where through Chat
         | GPT it will give you near enough instant replies only slightly
         | slower than 3.5 turbo.
        
           | danielbln wrote:
           | Maybe the reverse is true? To provide faster performance they
           | dumb down gpt4 (quantization?) that serves the chatgpt UI but
           | the API is full power albeit slow.
        
       | pierat wrote:
       | "GPT-4 hasn't gotten worse since March" can be 100% true at the
       | same time OpenAI puts more rules and limiters keeping more
       | interesting answers from being said.
       | 
       | Ive noticed it quit giving as detailed answers and as thorough.
       | It's also refused to do more complex programming where it used to
       | accept those questions.
       | 
       | Being artificially limited by OpenAI can still be done without it
       | getting "worse". But it effectively is worse for us users.
        
         | H8crilA wrote:
         | His point is that it literally didn't change, that includes the
         | safeguards (which are a part of the model).
        
           | throwuwu wrote:
           | The moderation endpoint is separate for the API so I'd
           | imagine it's the same for chat. The model could be the same
           | while they change:
           | 
           | Moderation
           | 
           | Temperature / top-p
           | 
           | System prompt
           | 
           | Some other internal system we aren't aware of
        
           | jeremyjh wrote:
           | He said the API hasn't changed. But what about the Chat
           | website?
        
             | jacquesm wrote:
             | Exactly my question... I wonder if the pre and post
             | processing of the chat interaction isn't what is driving
             | the perceived differences.
        
               | danielbln wrote:
               | I could not figure out what everyone was talking about
               | yesterday as my experience with gpt4 has not degraded at
               | all, but I'm using a third party client via API.
        
               | flir wrote:
               | Which 3rd party client, if you don't mind me asking? I'm
               | looking to move, and the space isn't mature enough yet
               | for there to be a clear leader.
        
               | throwuwu wrote:
               | I tried comparing the API response against the chat on
               | the same question a few times. There isn't a huge
               | difference but I'd pick the responses from the API over
               | the ones from chat. Hard to say though, could be RNG.
        
               | IanCal wrote:
               | Could be related to the system prompt.
        
           | aliston wrote:
           | Is it true that the safeguards are considered part of the
           | model? I had assumed that the "safeguards" that limit certain
           | types of responses in ChatGPT were separate from the actual
           | language model.
        
             | helpfulclippy wrote:
             | It seems to me that there's lots of room to change stuff
             | that profoundly affects the range of responses without
             | altering the base model. The prompt template alone seems
             | like a place outside the model where we've seen safeguards
             | get implemented, and other stuff that affects the
             | usefulness of a model's responses.
        
           | anticensor wrote:
           | The inhibitive ("moderation") model is separate from the
           | generative ("chatbot") model. They work in tandem.
        
           | ethbr0 wrote:
           | Aren't there archived transcripts with prompts now?
           | 
           | Seems like we need "model transparency" and log implementers
           | to flag drift, a la RFC 9162 / Certificate Transparency.
        
             | throwuwu wrote:
             | I guess we could go back through our histories and resubmit
             | the same prompts a few times. Wouldn't be a fair comparison
             | since we only have 1 sample from the old version.
        
           | greenknight wrote:
           | My understanding, is that ChatGPT... has a prompt that goes
           | before your question (and previous answers) that set up the
           | stage for how it should respond. It could be that the
           | safeguards they have put in place, sit in this prefix prompt
           | state... rather than GPT4... and that you would get the
           | normal answers via the api rather than via ChatGPT.
        
           | aeternum wrote:
           | I wouldn't read too far into the tweet. They do tell us when
           | it changed, and the bottom of the page clearly states :
           | ChatGPT May 24 Version
           | 
           | The release notes are producty and not very technical so it
           | is difficult to tell what actually changed.
        
           | onlyrealcuzzo wrote:
           | My immediate gut reaction at the original post about GPT-4
           | get SUBSTANTIALLY worse was that...
           | 
           | It didn't. People were just noticing LLMs still have a long
           | way to go after using them more.
           | 
           | I was shocked going through the thread that all of the
           | popular comments were confirmations that it did, in fact, get
           | MUCH worse.
           | 
           | It's nice to see from OpenAI that it didn't...
        
         | devinprater wrote:
         | No, it's not gotten worse, just more constrained and... fuzzy
         | and blurry.
        
         | kikokikokiko wrote:
         | I can hardly wait to see when, in the near future, every new
         | LLM may become almost useless.
         | 
         | The OG models were trained on real world, human generated
         | content (for the most part at least). Starting in 2022 the cost
         | of automatic generating "human sounding enough" text has gone
         | to such low depths that I expect it to be pretty much
         | impossible to avoid training any model on text already
         | generated by a LLM.
         | 
         | What will the result of this feedback loop be, I can't tell. It
         | will probably be just an even more generic, corporate speak,
         | bland sounding bla bla bla than we get today, and the level of
         | hallucinations may get even worse.
         | 
         | In a way it makes me happy to imagine that the most dangerous
         | tech humanity ever invented may itself be it's own main
         | obstacle to future refinements.
        
           | no_wizard wrote:
           | I think this would be cost prohibitive, though maybe not,
           | however -
           | 
           | it could just save every answer its given and scan text for
           | it. If they're a match it could just not index it, right?
        
             | data_maan wrote:
             | there's still the pre-2023 data to train it on ... and then
             | augment with handpicked stuff.
        
               | kikokikokiko wrote:
               | The dataset of internet content pre-2022 will be regarded
               | in the future just like low-background steel. Any content
               | generated post the release of the first generally
               | accessible LLMs will be considered radioactive.
        
           | throwuwu wrote:
           | It might be better for the corpus to be expanded with higher
           | quality text that is produced on demand from contractors.
           | There are also large pools of data yet to be accessed. One
           | giant source would be podcast transcripts but there are many
           | others.
        
       | yodsanklai wrote:
       | It's really fascinating to see all the irrational thinking
       | triggered by ChatGPT. "risks of human race extinction", "soon no
       | more need for developers, doctors or lawyers", "possibility of
       | consciousness emerging", so called experts in "prompt
       | engineering".
       | 
       | It's a dumb tool, if you're lucky you can get it to spit
       | something useful (but you need other tools to check the
       | correctness of what it returned). There are certainly many useful
       | applications, but the technology is inherently limited.
        
       | wg0 wrote:
       | This was predictable from the get go and many had had pointed
       | that out already that soon people will start noticing that LLMs
       | aren't as magical as they thought once the initial awe wanes
       | away.
       | 
       | Don't think OpenAPI is doing anything here, it's not in their
       | interest to reduce the "quality" even there's no objective and
       | repeatable way to measure the quality either.
       | 
       | It's all probabilities all the way down. Who knows what the model
       | will do. I mean, you can dry run by hand but even on quad core
       | processors, it's damn slow so imagine the inference by hand.
        
         | sebzim4500 wrote:
         | Was anyone in that thread claiming that the API had gotten
         | worse? A lot of people were suggesting to use the API rather
         | than the ChatGPT interface in order to avoid the degradation.
        
       ___________________________________________________________________
       (page generated 2023-06-01 23:01 UTC)