[HN Gopher] OpenAI Employee: GPT-4 has been static since March
___________________________________________________________________
OpenAI Employee: GPT-4 has been static since March
Author : behnamoh
Score : 141 points
Date : 2023-06-01 18:27 UTC (4 hours ago)
(HTM) web link (twitter.com)
(TXT) w3m dump (twitter.com)
| furyofantares wrote:
| I think we don't notice our expectations have gone up, and we
| don't notice that we remember the hits and then expect all hits.
|
| We didn't notice the misses at first, because it's what we
| expected to begin with, and we very strongly noticed the hits
| because they were unexpected. Now we notice the misses and expect
| the hits.
| skepticATX wrote:
| Genuinely asking: have our expectations gone up, or have we
| started thinking about it in a more critical and realistic way
| after the initial glean has worn off?
| coffeebeqn wrote:
| It's magic! But it's also wrong a lot. These days if I ask
| anything important I'll also have to Google the answer and
| I'll also ask chatgpt if it's sure about some of the details
| tines wrote:
| > glean
|
| I think the word you're thinking of is the noun "gleam" which
| means a kind of lustrous shine, rather than the verb "glean"
| which means to harvest the remainder of something or to
| collect in small parts.
| doctoboggan wrote:
| I definitely agree with this take. I had similar feelings to
| everyone in yesterday's thread complaining about the models
| failures but I agree that its just our shifting expectations
| and not a model update.
| jonas21 wrote:
| Yeah, it's kind of the same way people have been saying Google
| search has been going downhill for years, but compressed into a
| much shorter time frame.
| subsubzero wrote:
| It has and its gone downhill alot in the past 4-5 years. This
| is due to the need to keep an ever increasing stream of
| income flowing to google with a mostly static number of users
| and "fixing" this issue by piling more ads into users search
| results.
| awegio wrote:
| Then just use an Ad Blocker. The actual problem is SEO spam
| though.
| throwuwu wrote:
| It has and so has the YouTube algo. You can try the verbatim
| setting with search to see that there has been some monkeying
| going on but even that doesn't fix everything. Maybe that's
| due to spam sites but it just seems less effective.
| throw1623 wrote:
| This. I've been using it daily and nothing has changed. It's
| still as great as day 1.
|
| The audience of HN is, as Taleb would say, intellectuals yet
| idiots. They have trouble measuring change and are prone to
| hyperboles. Some still try to minimize the impact ChatGPT will
| have, while focusing on bullshit like 'hallucinations', or
| nitpicking about the quality of the code and so on. Can't see
| the forest from the trees.
|
| If you're looking for intelligent discussion, look elsewhere.
| kkdkdod wrote:
| [dead]
| huevosabio wrote:
| I think this is the right way to look at it: we are very quick
| at updating expectations.
|
| The first flight is magic, the nth one is a chore.
| pmontra wrote:
| You're generally right but one of my nth ones was Brisbane to
| Cairns over the reef during the day. I'd fly it again and
| again.
| huevosabio wrote:
| One of my favorites is landing in Rio, the landscape is
| spectacular and then you get Christ the Redeemer as the
| cherry on top of the cake.
| williamcotton wrote:
| Whenever I'm flying over an ocean on a clear moonlit night it
| still very much feels like magic!
| myshpa wrote:
| And now it's even gonna smell like bacon ...
|
| The fat of dead pigs, cattle and chickens is being used to
| make greener jet fuel
|
| https://www.bbc.co.uk/news/science-environment-65727664
| throwuwu wrote:
| The refusals are new. I've gotten a few myself recently. I've
| never seen them before. They're not moderation related and they
| always include some text about the task being too complex.
| ribosometronome wrote:
| Are there any 'complex tasks' you've had refused you can
| share?
| insanitybit wrote:
| That sounds like it could just be a cap applied to free
| users. That's hardly the same as a new model.
| irthomasthomas wrote:
| He specifically singles out the API and never answers if
| chatgpt is same. Most people are probably using the chatgpt
| interface, which seems to have more alignment training.
| ren_engineer wrote:
| I think it's worth noting the response explicitly said the paid
| API model hasn't been changed, so ChatGPT could have been
| changed through many different ways outside the core model
| jstarfish wrote:
| No, this is peak corporate gaslighting, _Open_ AI being the
| paragon of integrity and all. Nobody should trust a goddamn
| thing that comes out of their mouths-- or their product.
|
| It's no coincidence the flat-fee service is visibly crippled,
| while _per-request API users_ are not reporting any difference.
| (edited)
|
| Ignoring everyone here, look at the other "hacker" groups--
| like the jailbreaking community. They have a lot to say about
| recent changes that coincide with their hacks not working. All
| of a sudden, with OpenAI supposedly changing nothing, technical
| bypasses just stopped working. OpenAI changed nothing, so this
| _must_ be deus ex machina.
|
| I'm not even jailbreaking it but the results I get for simple
| code requests through the web UI have become unusable garbage.
| It puts less effort into responses than an unpaid-and-
| overworked intern. Others here report the same. The current
| "iteration" seems hellbent on terminating conversations as
| quickly as possible once they stray from the explicit scope of
| the original topic and is almost as hostile to fixing its own
| errors as it is to endorsing eugenics, whereas in the beginning
| it would humor every idle thought I threw at it in long
| conversations. You can literally see this reflected in the logs
| they forced retention of. It's acting like a customer support
| rep desperately trying to end a call before it exceeds a call-
| time quota.
|
| This particular current workflow seems like it lends itself to
| better organization of training data-- conversations are what
| the title says they're about. It also seems like it lends
| itself to anti-jailbreaking because pretexting it with
| irrelevant information forces a change of scope-- and a summary
| termination.
|
| But _OpenAI says_ they changed nothing, so rather than _one guy
| lying_ without consequence, a community of professionals using
| and abusing the tool must _all_ be victims of rhetorical
| fallacy? Nobody 's qualified to reverse-engineer corporate
| bullshit anymore without being infantilized...
|
| (Liars running a black box-- what could possibly go wrong? We
| need regulation!)
| Centigonal wrote:
| > It's no coincidence the "free" service is visibly crippled,
| while paid API users are not reporting any difference.
|
| I thought GPT-4 wasn't available to free users?
|
| Keep in mind the tweet is specifically about the OpenAI API -
| They might be updating ChatGPT without telling anyone
| (although they have release notes)
| jstarfish wrote:
| "Free" was the wrong word to use, I meant to say
| "unlimited."
|
| (edit: even that's not right. I think the verbiage is more
| accurate now.)
| matthewmacleod wrote:
| _It 's no coincidence the "free" service is visibly crippled,
| while paid API users are not reporting any difference._
|
| GPT-4 is not available to free users, which really calls the
| pretext of this comment into question. What is being alleged
| to have happened, and why is this response unreasonable?
|
| I'm not a very heavy GPT-4 user but I do usually use it once
| a day or so - it doesn't appear to have changed noticeably
| but I'm not paying close attention.
| jstarfish wrote:
| It was badly-worded on my part. And I stand corrected-- I
| just saw all the nested tweets where people are complaining
| about the API too!
|
| So many people are noticing a degradation in speed and
| quality of responses, and not in isolation. Rather than
| acknowledge this, the question is rephrased to one the
| responder can rightfully deny-- and suggest no changes have
| taken place without explicitly saying as much. You normally
| only see this sort of sliminess from politicians and
| executives on the witness stand.
|
| > Is anyone else noticing significantly downgraded GPT-4
| capabilities today? Seems like OpenAI updated the model,
| and results aren't as good as before. [mentions API in a
| child comment]
|
| > The API does not just change without us telling you. The
| models are static there.
|
| > This is good to know. That means GPT-4 has been static
| since March right? 0314?
|
| > Correct
|
| Never ask questions to which one word suffices as an
| answer.
|
| Collective confusion in the thread suggests _something_ has
| changed, but the most OpenAI will attest to is that the
| _API_ is unchanged and the _models_ are static. And this
| may well be true, but rather than admit "...but we were
| fucking with the middleware/parameters" they took a firm
| position on a strawman argument and ignored everybody who
| followed with more-direct questions. Except this guy:
|
| > I've noticed inconsistency with certain prompts
| performance. Is that just the non-deterministic nature of
| the API?
|
| > Yes
|
| Oh, ok. It's because the fucking _API_ is _non-
| deterministic_ that code that has worked both reliably and
| predictably for everyone now runs like shit for everyone.
| For fuck 's sake, you can get better answers from a Magic
| 8-Ball. This guy even made the mistake of presenting an
| answer he'd believe for the respondent to feed into. He
| might as well have asked if inconsistent performance was
| because of the war in Ukraine.
|
| "Logan.GPT" must moonlight as a fortune teller. He's only
| responding to people foolish enough to ask the wrong
| questions.
| shmoogy wrote:
| ChatGPT4 feels (is?) typically faster, and the results feel
| ... above 3.5 turbo but below what 4 was. The API seems
| exactly the same to me, but you cant do an apples to apples
| test due to variance between replies.
|
| I dont really doubt they scaled the chatGPT4 model down a
| little to try to save costs with the plugins and increased
| usage
| californical wrote:
| Yeah I'm having the same experience. The replies have
| gotten faster in the last few days, at the expense of
| quality.
|
| It's noticeable because I used to be able to read each
| word as it was printed from the GPt-4 output, but now it
| goes far too fast for me to keep up with.
|
| Kinda sucks to not have transparency at all into the
| black box, they've strayed so far from "open" at this
| point that it's comedic
| chasd00 wrote:
| > Liars running a black box-- what could possibly go wrong?
| We need regulation!
|
| "liars running a black box" is the definition of government
| regulation sheesh.
| flir wrote:
| Today I had ChatGPT 4 (Not GPT 4) correct itself mid-
| response. I was asking it something very simple about
| regexes:
|
| > How about matching 'a' as the second character of a string
| only?
|
| It responsed with the wrong regex plus a bunch of explanatory
| junk:
|
| > '^a.'
|
| Then halfway through the explanatory junk, it corrected
| itself like this:
|
| > Apologies for the confusion in the first response, the
| correct regular expression should be '^.a' for matching 'a'
| as the second character of a string:
|
| And kept on with the (now correct) explanatory junk.
|
| All in a single response. I've certainly never seen that
| before (if someone has, please weigh in). Maybe the model
| hasn't changed, but the pipeline has? Like... there's a
| second model trying to correct the mistakes of the first,
| maybe? (timings are probably wrong for that, but something
| like that)
| coffeebeqn wrote:
| There's also a new button "continue generating" so they are
| changing some things
| theturtletalks wrote:
| ChatGPT 4 seems to struggle with context changes within the
| same thread now. I asked it about some parks near me and
| then switched to asking about code. It just kept responding
| about the parks even when I corrected it. In some cases, it
| would merge the park question and the code question. I had
| not seen that with the free or Plus version until a few
| days ago.
| inciampati wrote:
| Interesting. Bing chat does this when you get it to talk
| about naughty things. My personal favorite is convincing it
| to make weird art out of quotes from the Tay chatbot. It
| will write them until it says a prohibited word or phrase,
| or touches some forbidden part of it's state space.
| flir wrote:
| Hm. Maybe they've backported some nanny code from Sydney?
| iliane5 wrote:
| AFAIK it's pretty standard practice not to expose the
| "raw" LLM directly to the user. You need a "sanity loop"
| where user input and the output of the LLM is checked by
| another LLM to actually enforce rules and mitigate prompt
| injections, etc.
| chaxor wrote:
| I have gotten that quite a few times now and it does seem
| to be new. In my case though, it acted as if it would
| change something, but then didn't, and got stuck in a loop.
| I got it enough times that it seemed to happen more often
| when the output needed to be quite long, so I had assumed
| (could be wrong) that when the output was nearing a certain
| limit, they had implemented something that would start a
| new output, such that the interaction of "please continue
| the script" would not be necessary.
|
| Unfortunately though, it would start over with the script
| completely, and then get stuck in a loop until it broke,
| this not being able to even save the response at all.
|
| Something definitely feels like it changed, but I suppose
| it could just be more use of the system.
|
| Another possibility is related to some strange issues with
| hardware and balancing etc, despite not changing software
| or params. It's strange, and it _shouldn 't_ happen, but
| sometimes things change from simple batching and balancing,
| or lower level hardware related things, which are very
| annoying to debug.
| dontupvoteme wrote:
| I just discovered when playing around with translations that
| there is some hidden filter/killswitch that immediately stops the
| generation of the opening of some books. It doesn't matter if the
| invoking prompt is to recite the book opening paragraph, or to
| translate it from a foreign language to English, or what.
|
| It's not RLHF induced because it works via API and it only
| triggers in English, but sure enough try to get it to output
|
| >Call me Ishmael. Some years ago--
|
| >"It was the best of times,
|
| I guess this might get you flagged (there is no alert to the user
| that this filter kicked in, and it will output it in any other
| language, and it works in the API) so I'm hesitant to play around
| with it more, but it's very strange - especially as these are
| long since in the public domain.
| Ozzie_osman wrote:
| He said "the API". He didn't say anything about the Chat version,
| which seems to have more protections that may be on top of the
| model and not embedded into it.
| textninja wrote:
| Must be in the trough of disillusionment. The tech is truly
| useful though so we'll reach the plateau soon enough.
| josecyc wrote:
| I agree I've noticed I need to be quite specific now, it won't
| realize bugs in the code unless I tell it so. A head to head
| comparison of the different versions is needed to validate this.
| elboru wrote:
| I don't get why is this being discussed, isn't it easy to go
| check old conversations and try to replicate them today?
| renewiltord wrote:
| You need to have used the API and set temperature zero, but if
| you have that historical data you can reasonably test and see.
| ShamelessC wrote:
| You would think so, but not a single person complaining has
| provided convincing proof of the degradation.
| sebzim4500 wrote:
| I went through a bunch of old prompts and the response was
| definitely worse now than it was then.
|
| Since it's not determinstic though, it's hard to draw
| conclusions. Especially since the sample size was 10ish.
| dgellow wrote:
| GPT isn't deterministic
| karmasimida wrote:
| The model didn't change but doesn't mean the inference didn't
| change.
|
| Without going specifics it is meaningless for the discussion
| heliophobicdude wrote:
| The models for the api's could have not changed but the web
| app's might have too
| armchairhacker wrote:
| I doubt the March 14 model is any different, since it's versioned
| and I don't see why OpenAI would want to change it behind the
| scenes.
|
| Also, companies are evaluating GPT-4 to determine whether they
| want to pay for it, so OpenAI has an strong incentive to not
| downgrade at least the API.
|
| I believe the May 5 model is _different_ , at least in the chat
| interface, because it's fine-tuned to detect jailbreaks and the
| temperature/other hyper-parameters may have changed. And I can
| imagine this fine-tuning making the model less creative and worse
| at solving analytical tasks.
|
| Personally I haven't noticed any change, except in my own
| awareness. Sometimes GPT4 gets very hard prompts right, and
| sometimes it gets simple problems wrong. So it's not hard to see
| how people can form biased opinions from selective attention or
| just luck.
| TradingPlaces wrote:
| The model may be the same, but the chatbot is not. They have made
| the responses shorter to save inference expenses, which are huge
| in a model the size of GPT-4.
| haolez wrote:
| I use GPT-4 a lot and something has definitely changed in my
| recent experience. It's simply worse.
| dangerlibrary wrote:
| A couple observations from attempting to use both GPT-3.5 and
| GPT-4 via the web interface for coding tasks:
|
| - The model's ability to respond accurately drops drastically
| when asked questions of the form "is there a different way to
| accomplish X, using Y?" or "is there a way to accomplish X that
| runs in O(log(n)) time instead?" Example: I wanted to upsert an
| integer value using a SQLite db using "INSERT ... RETURNING..."
| ChatGPT repeatedly told me that sqlite doesn't support
| "RETURNING" (it does, since March 2021). It insisted I would need
| two DB round trips from my application to accomplish this. When
| asked "can this be done in one round trip, instead?" it
| repeatedly wrote code that would return the number of rows
| modified instead of the integer column value.
|
| - ChatGPT's limited standard library knowledge means that the
| solutions it produces, even when correct, are often lower-level
| and less idiomatic. Problems that would be trivially solved with
| e.g. a Java.String.replaceAll or .codePointCount will instead
| loop over each character, often splitting the string into an
| intermediate array and implementing special cases for first/last
| character edge cases. The code winds up being mostly correct, but
| also (for lack of a better word) weird. No human I've ever worked
| with would do things the way ChatGPT sometimes does, which means
| the code will likely be much harder to maintain and debug over
| time.
| ch4s3 wrote:
| > are often lower-level and less idiomatic
|
| I would go as far as to say it's adept at producing technically
| working spaghetti.
| jstarfish wrote:
| > ChatGPT repeatedly told me that sqlite doesn't support
| "RETURNING" (it does, since March 2021).
|
| Its dataset cuts off around 2021. There's a little footer
| message warning you not to expect knowledge of recent events.
| dangerlibrary wrote:
| It cuts off in September 2021, supposedly - that's why I
| pointed out the month of the relevant sqlite release.
| lee101 wrote:
| [dead]
| andrewstuart wrote:
| GPT 3.5 is a much better coding assistant.
| pleb_nz wrote:
| I've found it's gotb eyes and have given up using it for tasks
| more often than I used to.
|
| For copilot, I no longer get multi line complete suggestions and
| it's really slow to deliver single line suggestions and they're
| more often incorrect. I need to dig into it, but it's definitely
| degraded further and I don't know if it's just my environment or
| a wider issue. I need to dig in and figure out - is anyone else
| experiencing these things?
| two_in_one wrote:
| ChatGPT Plus shows the version. It was May 12, if I remember
| correctly, then it changed to May 24. Which probably means there
| were some changes. If not in GPT-4 itself, then in pre or post
| processing. They should have some safety filters at the end, I
| think.
| phillipcarter wrote:
| Posted this in yesterday's thread, but once again I think this is
| just people feeling the magic wear off. People have poked around
| a lot more and found the flaws while also trying to get it to do
| real-world tasks. That wasn't true when it first came out.
|
| It's fine, this tech has never been magic anyways, won't be
| replacing all our jobs, won't take over the world, etc. It's
| still awesome for what it is.
| madrox wrote:
| This doesn't pass the smell test for me. _Something_ has changed.
| Maybe the model hasn't changed, but something somewhere is making
| output worse.
|
| I'm honestly more concerned if OpenAI doesn't even realize it.
| Nothing is more infuriating as a user than convincing the
| developer your bug actually does exist. It speaks to poor
| monitoring, testing, and tooling.
| sebzim4500 wrote:
| I understand why people don't read 10,000 word articles, but this
| is a 15 word tweet. Would it really kill people to read it before
| commenting? He is very explicitly talking about the API, which
| uses a different model than the UI.
|
| The recent discussion was about the degradation in the UI model.
| emptyfile wrote:
| [dead]
| jacquesm wrote:
| To what degree is the pre and post processing of the chat client
| the source of the confusion rather than GPT-4 itself?
| [deleted]
| rat9988 wrote:
| I thought the API was versioned, though it doesn't mean the new
| versions aren't worse. And he doesn't talk about the model used
| in chat gpt. I'm a bit skeptical that this answers exactly what
| we were worried about. Or maybe we were not all worrying exactly
| about the same thing.
| nashashmi wrote:
| Is there a way to use GPT-4 directly? Instead of the muzzled
| ChatGPT that is supposed to give answers authorities and
| opinionators deem appropriate?
| heliophobicdude wrote:
| Assuming you mean the base model, no not right now. GPT-4 base
| model is not public.
| rkapsoro wrote:
| Perhaps the evolution of this story is an interesting example of
| confirmation bias?
| petabite wrote:
| The amount of times per day I have to click "Stop generating" and
| then say "No!" has definitely increased
| geraldwhen wrote:
| Whatever happens to chat gpt, ai image generation is incredible.
| It's so incredibly powerful that I don't think it's ever going
| away
| saiya-jin wrote:
| If you are doing some indie gaming, it can save tons of money
| on asset generation, you either ramp up and finalize it
| yourself or hire somebody but for fraction of the time. Ditto
| for web design, why would anybody but big, calcified megacorps
| with their ridiculous processes go and buy some stock images if
| you can get exactly what you want, in any imaginable style?
| skilled wrote:
| ChatGPT was a massive dopamine hit, particularly for people in
| areas like development. It was a tremendous release that has
| definitely laid out some new tracks for the future, especially on
| the web. I myself have found to be using ChatGPT a lot less
| recently.
|
| I got the GPT-4 API access and then I realized that I can't
| really use it for anything super major because I can't afford it,
| it is ridiculously expensive if you consider that you have to pay
| for all the failed requests, the wrong information or the wrong
| context also. Instead, I have written a bunch of Python scripts
| that do a select few tasks for me and I have my terminal open
| 24/7 anyway.
|
| As for the topic at hand, I have _definitely_ noticed a lot more
| disclaimers in the UI. I don't get it from the API at all, in 6
| months that I have been using the API - I've gotten _one_
| disclaimer.
|
| In the ChatGPT UI - I get them a lot. "Remember this", "Remember
| that", "Always look up the information" and things like this. I
| mean if it wasn't happening I would know because I have been a
| power-user pretty much all this time...
| RecycledEle wrote:
| If ChatGPT has been static since March (of 2023) the. Why does my
| version always change? Right now I'm running May 24.
|
| FWIW, I think it has improved.
| dang wrote:
| Recent and related:
|
| _Ask HN: Is it just me or GPT-4 's quality has significantly
| deteriorated lately?_ -
| https://news.ycombinator.com/item?id=36134249 - May 2023 (711
| comments)
| jimsimmons wrote:
| He's talking about the API. Not the web client
| neonsunset wrote:
| There is No War in Ba Sing Se.
| dubcanada wrote:
| Is it static? If I ask it the same question I asked it a few
| months ago it gave me a completely different answer. Is that
| because of some additional context above? Should we be starting
| fresh chats sooner?
| schrodinger wrote:
| There's _is_ a degree of randomness in its content generation,
| you can see it in action by hitting regenerate. It's a
| statistical text predictor that iteratively selects the most
| like my word after the ones it's selected so far, but when
| there are a few good candidates, it'll choose one at random.
|
| However, the model is static (it'll present the same candidate
| word list until retrained), which is what you may have heard.
| But the way a response is generated introduces the randomness.
|
| Note: I forget the parameter, but if you get the direct API
| access you can turn off this and get consistent answers.
| carabiner wrote:
| It's never been static and there's always been randomness when
| you regenerate an answer. There is no control of a random seed,
| so you can get a different answers for the same prompt.
| guraf wrote:
| You're right of course that some randomness is inherent, but
| you can adjust the "temperature".
|
| With a low enough temperature you get essentially the same
| output every time, with a at most just minor words swapped.
| acomjean wrote:
| It's not stactic. That's the thing about ai, you ask it the
| same thing it gives you a different answer/different image
| every time.
| drexlspivey wrote:
| That's not really an issue with AI in general, you can make
| it deterministic by tuning some parameters if you want
| (temperature in chatGPT, passing a seed in midjourney etc)
| AtNightWeCode wrote:
| Are we chatting directly with the model? Maybe the interface has
| changed. With long term use the likelihood of hitting edge cases
| is higher and maybe that is a cause as well for what users are
| seeing. People probably ask more vague questions over time. I
| might have done that.
|
| I have never experienced the amnesia problem in v3.5 though, that
| v4 clearly has. Just repeating incorrect answers that you ask it
| not to give. I did not have access to v4 in march so I can't do
| that comparison.
| jiggawatts wrote:
| I look at everything someone from OpenAI says as if a politician
| is saying it. Sam Altman especially is fond of statements that
| are deceptive but technically true. His employees appear to be
| following his lead.
|
| GPT 4 isn't ChatGPT 4, which is what most people use.
|
| There is also the "system prompt", which is also likely to be
| changing but not part of GPT 4.
|
| Etc...
| Sunhold wrote:
| What is "ChatGPT 4"? ChatGPT Plus can use GPT-4 and free
| ChatGPT uses gpt-3.5-turbo.
| californical wrote:
| ChatGPT isn't just using plain GPT-4, it's using a specific
| pre-prompt and _possibly_ additional fine-tuning (which I'm
| not sure about) compared to the plain GPT-4, available
| through API.
|
| Someone saying ChatGPT4 is just saying in shorthand, the
| ChatGPT model based off GPT-4
| sebzim4500 wrote:
| I don't see this as being political doublespeak. He is very
| clearly just talking about the API.
| CSMastermind wrote:
| Right but people are complaining about their experiences with
| ChatGPT.
|
| If you say, "ChatGPT has gotten noticeably worse" and they
| respond "nothing about our APIs have changed" then it would
| be reasonable to interpret that as them saying that nothing
| about ChatGPT has changed.
|
| When in reality many things might have changed about ChatGPT.
| furyofantares wrote:
| The tweet is a response to an API user saying GPT-4 got
| worse, and the tweet is clearly talking about the API and
| the model - the context of the HN megathread mostly talking
| about ChatGPT isn't present in the tweet that's being
| replied to.
|
| There's a lot of confusion around this, but nothing appears
| to be caused by doublespeak to me. Confusion about which
| GPT/ChatGPT anyone is reporting about has been pretty
| ubiquitous for a while.
| stavros wrote:
| He says "the models are static", which means that, if you
| keep using the same model, you will keep getting the same
| results.
| pclmulqdq wrote:
| Which is not what anyone else is referring to when they say
| that GPT-4 has gotten worse.
| smcin wrote:
| Then the HN title is grossly misleading.
| bugglebeetle wrote:
| Company who has repeatedly lied in public: there's nothing to see
| here!
| greatpostman wrote:
| Recent research showed that RLHF/censoring the model hurts the
| performance of the model. This is intuitively obvious, censorship
| isn't real, the data is (moral issues aside). So it hurts the
| integrity of the weights. The future is open sourced uncensored
| models, capitalism will demand the high performance. There's a
| huge discussion on it on Reddit right now:
|
| https://www.reddit.com/r/MachineLearning/comments/13tqvdn/un...
| haolez wrote:
| Maybe humans are like that as well? :)
| greatpostman wrote:
| Lots of evidence pointing in that direction. Suppression of
| logic (censorship) leaks into other areas of cognition
| Spivak wrote:
| People keep mixing up alignment as in do what I actually asked
| and weird moralizing.
|
| You will pay the alignment tax no matter what if you want a
| model that actually does things. It's the reason why Llama
| models do better with stream of thought prompting where you can
| ask GPT questions and it won't try completing the question.
|
| You can, and probably should, align some semblance of morality
| and ethics _for user-facing models that do things_. It 's
| customer service voice but for AI. If what you want out of a
| model is a vague mirror of humanity or maximum smarts at the
| cost of needing more detailed prompts then yeah, this stuff
| probably annoys you.
|
| Finding a balance between a model that gives the outputs humans
| actually want and the pull that training data has on the rest
| of the model making it stray from the theoretical "best" output
| is hard.
| mschuster91 wrote:
| > The future is open sourced uncensored models, capitalism will
| demand the high performance.
|
| No. The future is ML built on more than just Wikipedia, Github
| code and Reddit hot takes. Or, sarcasm aside, GIGO: garbage in,
| garbage out - when you don't take care what datasets you train
| _any_ kind of model on, you 're bound to get some surprises if
| something unexpected comes along. Be it the infamous "racist
| soap dispenser" or the Google (?) image classifier that made
| the rounds here just a day or two ago which had a safeguard
| because it kept confusing Black people with gorillas.
|
| Preventing this kind of harmful content isn't censorship - it's
| after-the-fact compensation for bad training (and the tendency
| of 4chan and other trolls to use discriminatory content
| generation as a weapon, just remember what they did to
| Microsoft Tay).
|
| That is where the money is: _curated_ training datasets.
| Everyone and their dog can train a ML model from scratch, all
| you need is money and cloning a few more-or-less-broken Github
| repositories for that. But acquiring a high-quality dataset as
| a foundation? That costs _real_ money to create.
| fragmede wrote:
| Yes, so just like Apple's real value is from Foxconn's
| manufacturing factories which they don't own, OpenAI's real
| value is from Sama, the Kenyan firm that did the RLHF
| annotating.
|
| I'm wondering what the legal agreement is for Google
| Classrooms. In elementary schools across the world, kids
| write 3rd grade essays in history class or whatever and the
| teacher grades them. Those essays and grades are all in a
| database. Does Google have the opportunity to train Bard
| across that dataset?
| chasd00 wrote:
| i'm not in the AI world but i've been curious how the level
| of effort compares between inventing and writing the model vs
| finding, tagging, and curating the training. Is creating the
| training data analogous to inventing a complete schooling
| curriculum from scratch? I bet that takes a long freaking
| time.
| two_in_one wrote:
| It shouldn't be limited to censoring. Fine tuning is done on
| small set of some domain specific samples. Surely that results
| in gradual 'forgetting' of the already learned data from big
| set. If run for long time model will relearn on the tuning set,
| and overfit most likely.
| chatmasta wrote:
| My experience using GPT-4 for coding is that it's got the
| knowledge and skill of a senior engineer, with the high
| maintenance of a junior engineer. You can get it to output small
| sections of quality code, but the amount of prodding it takes to
| piece it all together means you may as well have spent the time
| writing it yourself. But the future of GPT as a coding assistant
| is definitely bright. It just needs more chaining, so I feel less
| like _its_ assistant, asking it to come up with prompts and then
| pasting them back to it after iterating on the code.
| Our_Benefactors wrote:
| It's still way faster than I could hope to be. It's a true
| refactoring machine with 0 mental effort. "Here is this code
| block, make it do this other thing, add this feature and
| argument and error condition" etc. It augments the amount of
| output that's possible to achieve at any seniority level.
| furyofantares wrote:
| GPT-4 has incredible breadth and no depth.
|
| It's an excellent complement to a strong programmer, who will
| have incredible depth -- and may have breadth relative to other
| coders, but will be very narrow in the whole space of
| programming.
|
| Just don't use it for things you're already an expert on,
| except perhaps for starting out / bypassing boilerplate.
| rngname22 wrote:
| It's Einstein with dementia/alzheimers.
| throwuwu wrote:
| They just put out a blog post about how they're doing exactly
| that by rewarding chain of thought style reasoning.
| jorblumesea wrote:
| Generative AI will not be able to do that in any true sense.
| Not the way you're thinking about it.
| Mizoguchi wrote:
| Agree. However I noticed an unintended benefit of using GPT for
| coding is that it helps you think carefully about the problem
| you are trying to solve when writing the prompt.
| AtNightWeCode wrote:
| That is why most developer jobs are safe. Tasks from
| stakeholders often contains only a title.
| flir wrote:
| Rubber ducking in the age of WFH is definitely a use I've
| found for it.
| chasd00 wrote:
| i don't use it that much for code but when i do it's for a
| specific function or method. I describe the inputs, the logic
| i want, and the outputs i need. So basically i have it worked
| out in my head I just let chatgpt type it out for me. Any
| mistakes it makes are pretty easy to catch when using it this
| way.
| behnamoh wrote:
| "Language as a tool of thought" is a technology we often
| neglect we are equipped with.
| dkersten wrote:
| This is essentially what rubber ducking is.
| moffkalast wrote:
| > so I feel less like its assistant
|
| Ah so I'm not the only one who feels this way lol. It's as
| brilliant as it is dumb.
| naiv wrote:
| For me it is not even the code per se but how it helps me with
| naming database tables and properties when I describe the
| broader scope.
|
| The future is more than bright, especially with a new model
| that is more current.
|
| Offtopic but just today I was aksing it for the most current
| versions it knows of right now:
|
| Python: 3.9
|
| JavaScript: ECMAScript 2021 (ES12)
|
| Java: 16
|
| C++: C++20
|
| C#: 9.0 (.NET 5.0)
|
| Ruby: 3.0.0
|
| Swift: 5.4
|
| Go: 1.16
|
| Rust: 1.51.0
|
| TypeScript: 4.3
|
| PHP: 8.0
|
| Kotlin: 1.5.0
|
| Scala: 3.0
|
| R: 4.0.5
|
| Perl: 5.32
|
| And frameworks:
|
| Python:
|
| Django: Version 3.2
|
| Flask: Version 2.0.1
|
| Pyramid: Version 2.0
|
| TensorFlow: Version 2.6.0
|
| PyTorch: Version 1.9.0
|
| JavaScript:
|
| Node.js: Version 16.9.1
|
| Express.js: Version 4.17.1
|
| React.js: Version 17.0.2
|
| Angular: Version 12.2.4
|
| Vue.js: Version 3.2.6
|
| Java:
|
| Spring Framework: Version 5.3.9
|
| Hibernate: Version 5.4.32.Final
|
| Struts: Version 2.5.26
|
| PHP:
|
| Laravel: Version 8.54.0
|
| Symfony: Version 5.3.6
|
| CodeIgniter: Version 4.1.4
|
| Ruby:
|
| Ruby on Rails: Version 6.1.4
|
| Sinatra: Version 2.1.0
|
| C#:
|
| .NET Core: Version 5.0
|
| ASP.NET: Version 5.0
|
| Entity Framework Core: Version 5.0
| chasd00 wrote:
| i just asked it "what's the output of python -v" and the
| little sample code reported python 3.9.2.
| flir wrote:
| I got: Python 3.8.5 (default, Jan 27
| 2021, 15:41:15) [GCC 9.3.0] on linux Type
| "help", "copyright", "credits" or "license" for more
| information.
|
| And a lot of junk text.
| hereonout2 wrote:
| Ha! So far it seems four posters here have asked similar
| questions and got four different python versions.
|
| I don't know if we can treat the gpts in this way and
| expect reliable answers. We just get pretty good answers
| right up to the point that we don't.
| RadiozRadioz wrote:
| Asking it directly is a good way to get a response containing
| version numbers that look correct, but not a good way to
| actually indicate which versions it's been trained on.
|
| A better method would be to look at the language features
| it's using and infer from there. Or, better still, look at
| which versions were out when the training data was collected
| (which I believe is September 2021, don't quote me on that).
| grenoire wrote:
| Is there a risk of it hallucinating what new features
| _could_ do instead of basing information on having seen
| them before?
| RadiozRadioz wrote:
| Yes, good point, I suppose we would have to weigh the
| probability of that against it fudging a number. I would
| assume inventing a new feature is harder, but who's to
| say; all this LLM stuff comes down to probability.
| o1y32 wrote:
| It does not actually mean much. These chatbots can easily
| output JavaScript using var where it absolutely should not.
| It being aware of the latest standards does not mean it can
| properly use the new syntax or features -- it emits a form of
| whatever crappy/legacy code it was trained on.
| dontupvoteme wrote:
| Oh that's interesting, I didn't think about asking it
| directly about what version it thinks is current. I wonder if
| it can hallucinate here?
|
| It's really irksome when it tries to use functions which were
| renamed or removed. This could be detected automatically (and
| possibly remapped in some cases)
|
| I got mostly the same versions as you, both on chatgpt3/4,
| using english.
|
| If you ask it in other languages the minor version also seems
| to change often.
|
| Finnish gives you dates, and code-davinici-edit-001 is 1
| major release behind on almost everything.
|
| Tassa on luettelo ohjelmistojen viimeisimmista vakaiden
| versioista:
|
| Python: 3.9.5 (30.4.2021) JavaScript: ECMAScript 2021
| (23.3.2021) ECMAScript: ECMAScript 2021 (23.3.2021) Java: JDK
| 17 (28.9.2021) C++: C++20 (20.2.2020) C#: .NET 6 (8.11.2022)
| Ruby: 3.0.2 (24.8.2021) Swift: Swift 5.5 (20.9.2021) Go: 1.17
| (16.8.2021) Rust: 1.54.0 (27.5.2021) TypeScript: 4.4
| (28.7.2021) PHP: 8.0.9 (29.7.2021) Kotlin: 1.5.31 (26.8.2021)
| Scala: 2.13.6 (17.2.2021) R: 4.1.0 (18.5.2021) Perl: 5.34.0
| (30.5.2021) ... Sinatra: 2.1.0 (10.4.2021) .NET Core: 6.0
| (8.11.2022) ASP.NET: 5.0.10 (19.8.2021)
| dontupvoteme wrote:
| The main thing it's lacking is a client-side merge/error-
| detection engine which combines all of the previous outputs.
|
| Anything over ~100LoC and it gets sloppy including it all in
| the code block. Which is fine, because I don't need it in your
| context window if that bit is working and you're not relying on
| it.
|
| Though I have gotten outputs up to nearly 250 lines from it..
|
| Automatically feeding back code details + Traceback messages
| also fixes errors decently often enough.
| alfalfasprout wrote:
| Those aren't very good senior engineers then lol. GPT-4 has the
| coding skills of an overconfident junior engineer and little
| else.
|
| I agree the future of GPT-X models as coding tools is bright.
| But for actually doing engineering work outside coding (or even
| delicate changes in an existing code base) much less so.
| jsight wrote:
| I'd love it if junior engineers organized code as well as
| ChatGPT. I'd also love it if all engineers stopped mixing
| spaces and tabs.
| kristopolous wrote:
| That can mostly be fixed with conventional prettify tools.
| No modern ai required
| fragmede wrote:
| Do you have a style guide?
|
| ChatGPT, on spaces vs tabs: https://chat.openai.com/share/b
| 2b0be49-54f9-4f73-a753-3edbd1...
| [deleted]
| alfalfasprout wrote:
| Add a linting step to your CI to catch this stuff?
| jsight wrote:
| Maybe I could get chatgpt to write those pipeline steps
| for me. :)
| oarsinsync wrote:
| You can and you should. It's this kind of busywork I farm
| out to GPT these days and it's great at it. I suck at it,
| so it saves me an hour to focus on what I'm good at!
| Mizoguchi wrote:
| Unrelated to the model itself but infrastructure, yesterday was
| unusable for me, getting random "too many requests sorry bout
| that" errors. I think 1/4 of the requests during a 3 hr period
| didn't make it through. Imposible to build anything beyond
| experimental stuff on top of a service so unreliable. Haven't
| tried it through Azure yet, I wonder if it is any better?
| amelius wrote:
| Instead of complaining, why not show the benchmarks?
|
| Like: first it scored 83, now it scores only 42 (or whatever).
| sebzim4500 wrote:
| What benchmarks? We have no way to go back in time and run the
| old ChatGPT accessble models though benchmarks.
|
| And to my knowledge, no one has copy pasted an entire benchmark
| into the ui in order to run it, they just use the API.
| ShamelessC wrote:
| the commenters of HN are simply so smart that their hunch
| clearly holds more weight than scientific rigor.
| PaulHoule wrote:
| People have just fallen out of love with it and are realizing how
| it really isn't that good.
|
| Sorry "prompt engineers" but papers on arXiv show that when you
| give it fairly sampled problems it struggles to get the right
| answer more than 70-80% of the time. When you are under its spell
| you will keep making excuses but when you are looking at it
| objectively you'll realize the emperor is naked.
|
| If you give it very conventional problems it seems to do better
| than that because it is a mass of biases and shortcuts and of
| course it will sentence Tyrone to life in prison because it's a
| running gag that "Tyrone is a thug"... That's how neurotypicals
| think and no wonder why many of them think ChatGPT is so smart...
| It mirrors them perfectly.
| vsareto wrote:
| >When you are under its spell you will keep making excuses but
| when you are looking at it objectively you'll realize the
| emperor is naked.
|
| It's not perfect, but you have to admit the emperor is at least
| wearing a thong. It is the second most intelligent thing on the
| planet at creating text, even with its flaws. Putting that
| accomplishment in league with naked emperors is astoundingly
| biased.
| shrimpx wrote:
| > you have to admit the emperor is at least wearing a thong
|
| And a lot of people are getting turned on by it.
| PaulHoule wrote:
| Yeah, maybe it is "wearing a thong" since it's developers
| were so concerned about social acceptability.
|
| I would say though that it is no mean feat for the Emperor to
| get away with being naked in public, it takes power and the
| ability to wield it.
|
| Similarly ChatGPT has a few competences that add up to people
| perceiving it is able, one of which is the ability to come up
| with plausible and satisfying answers whether they are right
| or wrong (often using shortcuts) and another one is getting
| people to engage with these answers.
| michaelt wrote:
| By this, do you mean 'the second most intelligent thing'
| after humans?
| shrimpx wrote:
| A dog understands body language and subtle cues,
| participates in social hierarchy and community, has self-
| reflection, dreams, etc.
| ThrowawayTestr wrote:
| I used it to help me write an email begging for my job back and
| it was pretty helpful
| BulgarianIdiot wrote:
| Just ignore this fella, GPT-4 is massively useful. But it
| takes intelligence to get intelligence out of it.
| visarga wrote:
| I think this is our path - become good at working with AI,
| for jobs or if there are no jobs, to directly support
| ourselves. Supposing we will be able to use tools and AI to
| build things and make our own means.
| BulgarianIdiot wrote:
| This is our path yes. All paths end, eventually, but oh
| well.
| [deleted]
| BulgarianIdiot wrote:
| I very much doubt the tweet in question is of someone who loved
| GPT-4 yesterday, and has "fallen out of love" with it starting
| today.
|
| For someone so dismissive of "neurotypicals" you did the
| neurotypical thing and drop a hot take before clicking the
| link.
| wouldbecouldbe wrote:
| It's more that the balance in usefulness vs waste of time
| shifts into the other direction after a few times spending
| long time getting it right and just doing it yourself anyway.
| BulgarianIdiot wrote:
| I've spent 20 years explaining my code to a rubber duck and
| getting great use out of that duck. No one will ever
| convince me that it's less useful now when the duck
| actually has read the Internet and has useful suggestions
| and ideas back to share. Even if it makes shit up on an
| occasion.
| JohnFen wrote:
| I don't know. The value of rubber-ducking is that you
| have to break the problem down into very simple terms to
| explain it. It's the explaining it part that is magic.
| That value goes away (or is greatly diminished) if the
| rubber duck responds with anything other than a request
| for further explanation.
| BulgarianIdiot wrote:
| You have to break down a problem to GPT just like you
| would to a human, and if it misunderstands you or asks
| you, you'd have to elaborate just the same. It's modeled
| after us, after all. The rubber duck is a stand-in for a
| human. GPT also is. But just a vastly better one.
| JohnFen wrote:
| Of course.
|
| The part I was responding to was this:
|
| > when the duck actually has read the Internet and has
| useful suggestions and ideas back to share
|
| I think that if it does that, it's legitimately less
| useful as a rubber duck. I'm not saying it isn't useful
| -- it's just not useful as a rubber duck anymore.
| PaulHoule wrote:
| The tweet is an OpenAI employee who is responding to people
| who have fallen out of love.
|
| I've seen transcripts of people interacting with ChatGPT who
| were obviously seduced by it and in a very giddy state,
| having so much fun because ChatGPT was playing an extended
| "game" with them that it didn't bother them at all that
| ChatGPT was spouting wrong answers.
|
| A major complaint I've had about the social sphere is that I
| seem to get the same result if I am 20% right or 50% right or
| 80% right or 95% right or 99.8%, it is just exhausting and I
| can never be good enough and I'm frankly envious that people
| see more of a glimmer of light behind that thing's "eyes"
| than they do behind mine.
|
| The core thing about neurotypicality isn't so much that they
| get the wrong answers but that they get the same answer
| whether it is right or wrong. For a long time I thought the
| basis of the "language instinct" is a derangement about
| reasoning with uncertainty that causes the grammar
| representation to collapse into a low-dimensioned subspace
| which is learnable with a limited amount of data. I wouldn't
| be surprised at all if other animals could beat us at rock-
| scissors-paper or poker if they could understand the rules of
| the game. The success of LLMs might give us some insight in
| this area although they are working with so much more data
| that Chomsky's old "poverty of the stimulus" argument might
| not apply.
| throwuwu wrote:
| You should read some more papers. Also fuck off back to reddit
| with this anti-neurotypical bias.
| mrbombastic wrote:
| I don't know, I think there is certainly a lot of hype but it
| is still pretty damn good. I used it every day for a wide
| variety of coding tasks and more often than not it is correct.
| As in compiles, produces correct output for inputs, and mostly
| writes reasonable code. Has it made software dev obsolete like
| some maximalists said it would? definitely not, but it seems
| equally hyperbolic to say the emperor has no clothes, it is
| undeniably a useful tool to me.
| PaulHoule wrote:
| "More often than not" is compatible with "correct 70% to 80%
| of the time".
| swores wrote:
| Your previous comment was a bit confusingly worded, I, and
| presumably the commenter you just replied to, read it as
| "more than 70-80% of the time" it struggles; rather than
| your intended meaning of it struggling to do better than
| correct 70-80% of the time.
| guraf wrote:
| [flagged]
| mustacheemperor wrote:
| In the thread yesterday it was brought up that if you use the API
| it feels considerably less hamstrung than the ChatGPT client, and
| this tweet seems to fit the assumption that the ChatGPT product
| is being tuned or governed differently from the API.
|
| >The API does not just change without us telling you. The models
| are static there.
|
| This reads to me as specifically indicating the models are not
| static elsewhere, ie, in ChatGPT.
| jonplackett wrote:
| It feels less hamstrung on output, but is instead hamstrung by
| being incredibly slow.
|
| GPT-4 via API will sometimes take 30 seconds + to respond to
| simple questions without any chat history, where through Chat
| GPT it will give you near enough instant replies only slightly
| slower than 3.5 turbo.
| danielbln wrote:
| Maybe the reverse is true? To provide faster performance they
| dumb down gpt4 (quantization?) that serves the chatgpt UI but
| the API is full power albeit slow.
| pierat wrote:
| "GPT-4 hasn't gotten worse since March" can be 100% true at the
| same time OpenAI puts more rules and limiters keeping more
| interesting answers from being said.
|
| Ive noticed it quit giving as detailed answers and as thorough.
| It's also refused to do more complex programming where it used to
| accept those questions.
|
| Being artificially limited by OpenAI can still be done without it
| getting "worse". But it effectively is worse for us users.
| H8crilA wrote:
| His point is that it literally didn't change, that includes the
| safeguards (which are a part of the model).
| throwuwu wrote:
| The moderation endpoint is separate for the API so I'd
| imagine it's the same for chat. The model could be the same
| while they change:
|
| Moderation
|
| Temperature / top-p
|
| System prompt
|
| Some other internal system we aren't aware of
| jeremyjh wrote:
| He said the API hasn't changed. But what about the Chat
| website?
| jacquesm wrote:
| Exactly my question... I wonder if the pre and post
| processing of the chat interaction isn't what is driving
| the perceived differences.
| danielbln wrote:
| I could not figure out what everyone was talking about
| yesterday as my experience with gpt4 has not degraded at
| all, but I'm using a third party client via API.
| flir wrote:
| Which 3rd party client, if you don't mind me asking? I'm
| looking to move, and the space isn't mature enough yet
| for there to be a clear leader.
| throwuwu wrote:
| I tried comparing the API response against the chat on
| the same question a few times. There isn't a huge
| difference but I'd pick the responses from the API over
| the ones from chat. Hard to say though, could be RNG.
| IanCal wrote:
| Could be related to the system prompt.
| aliston wrote:
| Is it true that the safeguards are considered part of the
| model? I had assumed that the "safeguards" that limit certain
| types of responses in ChatGPT were separate from the actual
| language model.
| helpfulclippy wrote:
| It seems to me that there's lots of room to change stuff
| that profoundly affects the range of responses without
| altering the base model. The prompt template alone seems
| like a place outside the model where we've seen safeguards
| get implemented, and other stuff that affects the
| usefulness of a model's responses.
| anticensor wrote:
| The inhibitive ("moderation") model is separate from the
| generative ("chatbot") model. They work in tandem.
| ethbr0 wrote:
| Aren't there archived transcripts with prompts now?
|
| Seems like we need "model transparency" and log implementers
| to flag drift, a la RFC 9162 / Certificate Transparency.
| throwuwu wrote:
| I guess we could go back through our histories and resubmit
| the same prompts a few times. Wouldn't be a fair comparison
| since we only have 1 sample from the old version.
| greenknight wrote:
| My understanding, is that ChatGPT... has a prompt that goes
| before your question (and previous answers) that set up the
| stage for how it should respond. It could be that the
| safeguards they have put in place, sit in this prefix prompt
| state... rather than GPT4... and that you would get the
| normal answers via the api rather than via ChatGPT.
| aeternum wrote:
| I wouldn't read too far into the tweet. They do tell us when
| it changed, and the bottom of the page clearly states :
| ChatGPT May 24 Version
|
| The release notes are producty and not very technical so it
| is difficult to tell what actually changed.
| onlyrealcuzzo wrote:
| My immediate gut reaction at the original post about GPT-4
| get SUBSTANTIALLY worse was that...
|
| It didn't. People were just noticing LLMs still have a long
| way to go after using them more.
|
| I was shocked going through the thread that all of the
| popular comments were confirmations that it did, in fact, get
| MUCH worse.
|
| It's nice to see from OpenAI that it didn't...
| devinprater wrote:
| No, it's not gotten worse, just more constrained and... fuzzy
| and blurry.
| kikokikokiko wrote:
| I can hardly wait to see when, in the near future, every new
| LLM may become almost useless.
|
| The OG models were trained on real world, human generated
| content (for the most part at least). Starting in 2022 the cost
| of automatic generating "human sounding enough" text has gone
| to such low depths that I expect it to be pretty much
| impossible to avoid training any model on text already
| generated by a LLM.
|
| What will the result of this feedback loop be, I can't tell. It
| will probably be just an even more generic, corporate speak,
| bland sounding bla bla bla than we get today, and the level of
| hallucinations may get even worse.
|
| In a way it makes me happy to imagine that the most dangerous
| tech humanity ever invented may itself be it's own main
| obstacle to future refinements.
| no_wizard wrote:
| I think this would be cost prohibitive, though maybe not,
| however -
|
| it could just save every answer its given and scan text for
| it. If they're a match it could just not index it, right?
| data_maan wrote:
| there's still the pre-2023 data to train it on ... and then
| augment with handpicked stuff.
| kikokikokiko wrote:
| The dataset of internet content pre-2022 will be regarded
| in the future just like low-background steel. Any content
| generated post the release of the first generally
| accessible LLMs will be considered radioactive.
| throwuwu wrote:
| It might be better for the corpus to be expanded with higher
| quality text that is produced on demand from contractors.
| There are also large pools of data yet to be accessed. One
| giant source would be podcast transcripts but there are many
| others.
| yodsanklai wrote:
| It's really fascinating to see all the irrational thinking
| triggered by ChatGPT. "risks of human race extinction", "soon no
| more need for developers, doctors or lawyers", "possibility of
| consciousness emerging", so called experts in "prompt
| engineering".
|
| It's a dumb tool, if you're lucky you can get it to spit
| something useful (but you need other tools to check the
| correctness of what it returned). There are certainly many useful
| applications, but the technology is inherently limited.
| wg0 wrote:
| This was predictable from the get go and many had had pointed
| that out already that soon people will start noticing that LLMs
| aren't as magical as they thought once the initial awe wanes
| away.
|
| Don't think OpenAPI is doing anything here, it's not in their
| interest to reduce the "quality" even there's no objective and
| repeatable way to measure the quality either.
|
| It's all probabilities all the way down. Who knows what the model
| will do. I mean, you can dry run by hand but even on quad core
| processors, it's damn slow so imagine the inference by hand.
| sebzim4500 wrote:
| Was anyone in that thread claiming that the API had gotten
| worse? A lot of people were suggesting to use the API rather
| than the ChatGPT interface in order to avoid the degradation.
___________________________________________________________________
(page generated 2023-06-01 23:01 UTC)