[HN Gopher] The consequences of generative AI for online knowled...
___________________________________________________________________
The consequences of generative AI for online knowledge communities
Author : mooreds
Score : 50 points
Date : 2024-07-31 16:19 UTC (6 hours ago)
(HTM) web link (www.nature.com)
(TXT) w3m dump (www.nature.com)
| dzink wrote:
| TLDR: People start asking more sophisticated questions on Stack
| Overflow and far less basic questions. There is fear of fewer
| novice people participating.
|
| What if junior people are able to use LLMs as sparring partners
| and learn without wasting time of more senior engineers? That
| leaves more time for online communities to discuss things that
| are on the edge of what's known that need worked on.
| JimDabell wrote:
| It also means that juniors with actually difficult questions
| are more likely to get them answered because less expert time
| is being wasted on regurgitating what is already spelt out in
| the documentation and tutorials.
| ssalka wrote:
| Just my initial thoughts from the abstract:
|
| > [A]ctivity in Reddit communities shows no evidence of decline,
| suggesting the importance of social fabric as a buffer against
| the community-degrading effects of LLMs.
|
| This seems rather disingenuous? People go to StackOverflow
| primarily to find answers to technical questions, usually
| questions where there is 1 or a small number of valid answers.
| Reddit, on the other hand, is a place where people often go to
| share their own opinions/posts/etc, or go to browse not looking
| for anything in particular. The decline in SO viewership makes
| sense with the rise of ChatGPT, but I wouldn't expect that to
| cause any equivalent decline from Reddit. If anything, this looks
| to me like evidence to the contrary: that LLMs _don 't_ degrade
| communities.
| kjkjadksj wrote:
| It makes a lot more sense to have an llm marketing on reddit
| than stackoverflow. You can't exactly shill a product while
| making a technical answer on stackoverflow unless your name is
| Ole Tang. With that in mind, the analysis fall short on
| measuring contamination of the reddit community with LLM
| content and if this is masking a human exodus.
| ssalka wrote:
| Fair point, more research is needed.
| JimDabell wrote:
| > We observe significant declines in both website visits and
| question volumes at Stack Overflow, particularly around topics
| where ChatGPT excels. By contrast, activity in Reddit communities
| shows no evidence of decline, suggesting the importance of social
| fabric as a buffer against the community-degrading effects of
| LLMs.
|
| This equates _declining activity_ with _degraded community_ ,
| which I don't think is a valid assumption to make.
|
| Taking Stack Overflow as an example, LLMs _should_ be decimating
| question volumes there. That's the whole point of LLMs! To answer
| questions! But _which_ questions are being eliminated? The
| boring, rote ones that can be answered by a glance at the
| documentation? Or the interesting ones that need expert knowledge
| and insight? My guess is that LLMs displace the majority of the
| former but less of the latter. So what's left once you take the
| lazy spam questions away from Stack Overflow? I would say that if
| LLMs can act as first-line support for those questions instead of
| expecting humans to answer them manually, the result will be a
| healthier community that more experts are willing to participate
| in. And sure enough, the paper reports:
|
| > upon ChatGPT's release, a systematic rise began to take place,
| such that users were increasingly likely to be more established,
| older accounts. The implication of this result is that newer user
| accounts became systematically less likely to participate in the
| Stack Overflow community after ChatGPT became available. Figure 6
| depicts the effects, indicating that questions exhibited a
| systematic rise in complexity following the release of ChatGPT.
|
| That looks like an improvement to me, not degradation.
| doe_eyes wrote:
| I think the mistake you might be making here is that some level
| of routine engagement and mentorship is needed to sustain the
| community. If I'm an expert, but I can only say something
| useful once every six months when some super-niche question
| drops, I'm just gonna stop visiting the site because of the
| negative feedback loop of going there and not finding anything
| to do 99% of the time.
|
| The other extreme - drowning in "noob" questions - is also
| harmful, which is why you end up with FAQs and rules. But there
| are human dynamics that underpin all these communities, and
| LLMs are undoubtedly destroying the incentives to participate.
| bilater wrote:
| I just read the abstract and wanted to add that a critical
| difference between subreddits and Stack Overflow is that you go
| to the latter for an objective answer (even though you might not
| always get one), whereas subreddits often provide subjective
| content, like advice or reviews.
|
| I think this kind of opinionated community will still thrive
| because LLMs are great for quick answers but don't yet provide
| strong opinions, and likely won't for a while due to regulations.
| In that way, having a close-knit human community with strong
| opinions on a subject will become even more valuable.
| lovethevoid wrote:
| Nothing about LLMs prevent them from writing what looks like
| strong opinions. Actually that's what I assume most bots that
| utilize LLMs on social platforms are doing, adhering to a
| "style" and writing "strong opinions" on a product.
|
| Here's one I asked chatGPT to generate about a hypothetical
| video game (related to a real one as source for what to
| reference):
|
| > hypothetical video game is sooo fggin trash. like, the
| parkour is meh, combat is a complete joke, and the whole game
| just feels off. seriously, what a waste of time.
|
| People only take safety in opinionated communities with the
| assumption everyone else is a human being offering their
| sincere take, and due to this assumption it's also much easier
| to get your bot to "fit in", pushing whatever narrative you
| want or letting it go wild.
| janalsncm wrote:
| The value of opinions comes from the fact that they were
| written by humans.
|
| For example, I'm sure you could get an LLM to write about how
| delicious peanut butter and pickles is as a combination, but
| most people do not agree.
| kjkjadksj wrote:
| But there is no tell on reddit who is a human without
| strong opinion biases, who is a human who is uninformed or
| otherwise biased, who is a human paid to market opinion or
| product, or who is a bot written to market opinion or
| content. Right now the only value is in the sense of the
| emporer not having clothes, of people assuming most
| comments are the first case when they really are more
| likely to be the last case.
|
| Reddit doesn't even pass the eye test. Look at
| /r/programming. over 6 million subscribed accounts, but the
| activity is like barely anyone is there. The top three
| posts of the last 24hrs currently have 72, 218, and 39
| comments respectively, and after that engagement falls off
| to like 0-5 comments a post. So we have around 350 comments
| a day on a board supposedly subscribed by 6 million account
| holders. The comment making rate of this community is
| apparently 0.006%, and a good deal of that is liable to be
| bot traffic.
|
| Dead internet theory is probably right. I should pack up
| and leave with the rest too I guess.
| lovethevoid wrote:
| Remember the pineapple and pizza trend? It's pretty old
| now, but a lot of people were buying it to try it out
| because of what people were saying online.
|
| Now that we have very accessible LLM, that same trend can
| be artificially created with ease. You can dominate entire
| subreddits with artificial content, and nobody cares as
| long as it isn't a giant glaring outlier. You can make
| people think a certain way, act a certain way, buy certain
| things, just as long as you seem human enough to pass the
| sniff test and that a large enough group of "people" were
| doing so.
|
| These opinionated communities will not survive much longer
| without implementing more severe human ID verification
| measures. You can bet against that, but I see the writing
| on the wall in my smaller public communities. They already
| don't trust much that isn't coming from someone in real
| time video.
| kjkjadksj wrote:
| I think its more that the shills have been on reddit for so
| long that people are blind to the marketing techniques they are
| being exposed to. You can probably get an LLM to spit out some
| boilerplate that passes the snuff for a reddit comment already.
| True for HN as well, I have seen a user mention they used a
| language model to reformulate their comment before posting as
| some security measure.
|
| In either case LLM are really a steam shovel technological
| moment for the shill marketing world. What used to take active
| posting and constant maintenance of your bank of canned
| responses can now be done with one operator running a script
| that simulates all these active posters with enough variance to
| not have the responses seem to be from the same bank of
| material. You could dominate an entire subreddit, have 1000 of
| your accounts to 1 user putting out enough AI derived signal
| that even the few real users who post start regurgitating it.
| Havoc wrote:
| My gut feeling is that lemmy (and mastodon) will be where the
| quality goes to hide.
|
| Also it's not just knowledge - Insta is rapidly turning into an
| AI trashfire. AI "digital creators" are taking other people's
| photos, face swapping another face (AI or another human not sure)
| and posting it as their own. Very Deja vu vibes sometimes if you
| see both versions.
| TeeMassive wrote:
| I found myself using Phind a whole lot more than Google or Stack
| Overflow. I find what I'm looking for very easily. I don't have
| to ask the same question in 10 different manners to look what I'm
| looking for, it does that for me. I don't have to deal with the
| community and snobbish power tripping moderators. I get the
| direct sources the answer is based on. It looks up the
| documentation for me.
|
| This is a real usage of AI that is there to stay IMO.
| lacoolj wrote:
| This is going to have ripple effects across the development
| ecosystem in the coming years. It probably won't be obvious at
| first, but devs that use LLMs to figure things out vs those who
| don't will slowly become more divided, both in opinion and
| knowledge/experience level. Garbage in/garbage out is very much a
| thing. I suspect the affect of recursively training models on
| other models is going to become a similar comparison here
| shortly.
|
| Hiring just became much harder
| internet101010 wrote:
| Given that Python is a an extremely popular first language, it
| makes sense that Python, pandas, SQL, etc. saw the largest drop
| and average account age went up. From what I have seen on
| programming subreddits the threads are mostly memes rather than
| QA; to me the fact that they remain unaffected tells me the
| communities are fine.
|
| While you do need a constant funnel of new users to sustain a
| community, I think the end result of this will be a higher
| quality funnel that will eventually create an environment of
| people solving niche problems.
|
| All that to say that I now do a first pass through a tuned model
| and only go to stackoverflow as a last resort. I skip Medium
| search results because they are mostly people shilling their
| company.
___________________________________________________________________
(page generated 2024-07-31 23:01 UTC)