[HN Gopher] The consequences of generative AI for online knowled...
       ___________________________________________________________________
        
       The consequences of generative AI for online knowledge communities
        
       Author : mooreds
       Score  : 50 points
       Date   : 2024-07-31 16:19 UTC (6 hours ago)
        
 (HTM) web link (www.nature.com)
 (TXT) w3m dump (www.nature.com)
        
       | dzink wrote:
       | TLDR: People start asking more sophisticated questions on Stack
       | Overflow and far less basic questions. There is fear of fewer
       | novice people participating.
       | 
       | What if junior people are able to use LLMs as sparring partners
       | and learn without wasting time of more senior engineers? That
       | leaves more time for online communities to discuss things that
       | are on the edge of what's known that need worked on.
        
         | JimDabell wrote:
         | It also means that juniors with actually difficult questions
         | are more likely to get them answered because less expert time
         | is being wasted on regurgitating what is already spelt out in
         | the documentation and tutorials.
        
       | ssalka wrote:
       | Just my initial thoughts from the abstract:
       | 
       | > [A]ctivity in Reddit communities shows no evidence of decline,
       | suggesting the importance of social fabric as a buffer against
       | the community-degrading effects of LLMs.
       | 
       | This seems rather disingenuous? People go to StackOverflow
       | primarily to find answers to technical questions, usually
       | questions where there is 1 or a small number of valid answers.
       | Reddit, on the other hand, is a place where people often go to
       | share their own opinions/posts/etc, or go to browse not looking
       | for anything in particular. The decline in SO viewership makes
       | sense with the rise of ChatGPT, but I wouldn't expect that to
       | cause any equivalent decline from Reddit. If anything, this looks
       | to me like evidence to the contrary: that LLMs _don 't_ degrade
       | communities.
        
         | kjkjadksj wrote:
         | It makes a lot more sense to have an llm marketing on reddit
         | than stackoverflow. You can't exactly shill a product while
         | making a technical answer on stackoverflow unless your name is
         | Ole Tang. With that in mind, the analysis fall short on
         | measuring contamination of the reddit community with LLM
         | content and if this is masking a human exodus.
        
           | ssalka wrote:
           | Fair point, more research is needed.
        
       | JimDabell wrote:
       | > We observe significant declines in both website visits and
       | question volumes at Stack Overflow, particularly around topics
       | where ChatGPT excels. By contrast, activity in Reddit communities
       | shows no evidence of decline, suggesting the importance of social
       | fabric as a buffer against the community-degrading effects of
       | LLMs.
       | 
       | This equates _declining activity_ with _degraded community_ ,
       | which I don't think is a valid assumption to make.
       | 
       | Taking Stack Overflow as an example, LLMs _should_ be decimating
       | question volumes there. That's the whole point of LLMs! To answer
       | questions! But _which_ questions are being eliminated? The
       | boring, rote ones that can be answered by a glance at the
       | documentation? Or the interesting ones that need expert knowledge
       | and insight? My guess is that LLMs displace the majority of the
       | former but less of the latter. So what's left once you take the
       | lazy spam questions away from Stack Overflow? I would say that if
       | LLMs can act as first-line support for those questions instead of
       | expecting humans to answer them manually, the result will be a
       | healthier community that more experts are willing to participate
       | in. And sure enough, the paper reports:
       | 
       | > upon ChatGPT's release, a systematic rise began to take place,
       | such that users were increasingly likely to be more established,
       | older accounts. The implication of this result is that newer user
       | accounts became systematically less likely to participate in the
       | Stack Overflow community after ChatGPT became available. Figure 6
       | depicts the effects, indicating that questions exhibited a
       | systematic rise in complexity following the release of ChatGPT.
       | 
       | That looks like an improvement to me, not degradation.
        
         | doe_eyes wrote:
         | I think the mistake you might be making here is that some level
         | of routine engagement and mentorship is needed to sustain the
         | community. If I'm an expert, but I can only say something
         | useful once every six months when some super-niche question
         | drops, I'm just gonna stop visiting the site because of the
         | negative feedback loop of going there and not finding anything
         | to do 99% of the time.
         | 
         | The other extreme - drowning in "noob" questions - is also
         | harmful, which is why you end up with FAQs and rules. But there
         | are human dynamics that underpin all these communities, and
         | LLMs are undoubtedly destroying the incentives to participate.
        
       | bilater wrote:
       | I just read the abstract and wanted to add that a critical
       | difference between subreddits and Stack Overflow is that you go
       | to the latter for an objective answer (even though you might not
       | always get one), whereas subreddits often provide subjective
       | content, like advice or reviews.
       | 
       | I think this kind of opinionated community will still thrive
       | because LLMs are great for quick answers but don't yet provide
       | strong opinions, and likely won't for a while due to regulations.
       | In that way, having a close-knit human community with strong
       | opinions on a subject will become even more valuable.
        
         | lovethevoid wrote:
         | Nothing about LLMs prevent them from writing what looks like
         | strong opinions. Actually that's what I assume most bots that
         | utilize LLMs on social platforms are doing, adhering to a
         | "style" and writing "strong opinions" on a product.
         | 
         | Here's one I asked chatGPT to generate about a hypothetical
         | video game (related to a real one as source for what to
         | reference):
         | 
         | > hypothetical video game is sooo fggin trash. like, the
         | parkour is meh, combat is a complete joke, and the whole game
         | just feels off. seriously, what a waste of time.
         | 
         | People only take safety in opinionated communities with the
         | assumption everyone else is a human being offering their
         | sincere take, and due to this assumption it's also much easier
         | to get your bot to "fit in", pushing whatever narrative you
         | want or letting it go wild.
        
           | janalsncm wrote:
           | The value of opinions comes from the fact that they were
           | written by humans.
           | 
           | For example, I'm sure you could get an LLM to write about how
           | delicious peanut butter and pickles is as a combination, but
           | most people do not agree.
        
             | kjkjadksj wrote:
             | But there is no tell on reddit who is a human without
             | strong opinion biases, who is a human who is uninformed or
             | otherwise biased, who is a human paid to market opinion or
             | product, or who is a bot written to market opinion or
             | content. Right now the only value is in the sense of the
             | emporer not having clothes, of people assuming most
             | comments are the first case when they really are more
             | likely to be the last case.
             | 
             | Reddit doesn't even pass the eye test. Look at
             | /r/programming. over 6 million subscribed accounts, but the
             | activity is like barely anyone is there. The top three
             | posts of the last 24hrs currently have 72, 218, and 39
             | comments respectively, and after that engagement falls off
             | to like 0-5 comments a post. So we have around 350 comments
             | a day on a board supposedly subscribed by 6 million account
             | holders. The comment making rate of this community is
             | apparently 0.006%, and a good deal of that is liable to be
             | bot traffic.
             | 
             | Dead internet theory is probably right. I should pack up
             | and leave with the rest too I guess.
        
             | lovethevoid wrote:
             | Remember the pineapple and pizza trend? It's pretty old
             | now, but a lot of people were buying it to try it out
             | because of what people were saying online.
             | 
             | Now that we have very accessible LLM, that same trend can
             | be artificially created with ease. You can dominate entire
             | subreddits with artificial content, and nobody cares as
             | long as it isn't a giant glaring outlier. You can make
             | people think a certain way, act a certain way, buy certain
             | things, just as long as you seem human enough to pass the
             | sniff test and that a large enough group of "people" were
             | doing so.
             | 
             | These opinionated communities will not survive much longer
             | without implementing more severe human ID verification
             | measures. You can bet against that, but I see the writing
             | on the wall in my smaller public communities. They already
             | don't trust much that isn't coming from someone in real
             | time video.
        
         | kjkjadksj wrote:
         | I think its more that the shills have been on reddit for so
         | long that people are blind to the marketing techniques they are
         | being exposed to. You can probably get an LLM to spit out some
         | boilerplate that passes the snuff for a reddit comment already.
         | True for HN as well, I have seen a user mention they used a
         | language model to reformulate their comment before posting as
         | some security measure.
         | 
         | In either case LLM are really a steam shovel technological
         | moment for the shill marketing world. What used to take active
         | posting and constant maintenance of your bank of canned
         | responses can now be done with one operator running a script
         | that simulates all these active posters with enough variance to
         | not have the responses seem to be from the same bank of
         | material. You could dominate an entire subreddit, have 1000 of
         | your accounts to 1 user putting out enough AI derived signal
         | that even the few real users who post start regurgitating it.
        
       | Havoc wrote:
       | My gut feeling is that lemmy (and mastodon) will be where the
       | quality goes to hide.
       | 
       | Also it's not just knowledge - Insta is rapidly turning into an
       | AI trashfire. AI "digital creators" are taking other people's
       | photos, face swapping another face (AI or another human not sure)
       | and posting it as their own. Very Deja vu vibes sometimes if you
       | see both versions.
        
       | TeeMassive wrote:
       | I found myself using Phind a whole lot more than Google or Stack
       | Overflow. I find what I'm looking for very easily. I don't have
       | to ask the same question in 10 different manners to look what I'm
       | looking for, it does that for me. I don't have to deal with the
       | community and snobbish power tripping moderators. I get the
       | direct sources the answer is based on. It looks up the
       | documentation for me.
       | 
       | This is a real usage of AI that is there to stay IMO.
        
       | lacoolj wrote:
       | This is going to have ripple effects across the development
       | ecosystem in the coming years. It probably won't be obvious at
       | first, but devs that use LLMs to figure things out vs those who
       | don't will slowly become more divided, both in opinion and
       | knowledge/experience level. Garbage in/garbage out is very much a
       | thing. I suspect the affect of recursively training models on
       | other models is going to become a similar comparison here
       | shortly.
       | 
       | Hiring just became much harder
        
       | internet101010 wrote:
       | Given that Python is a an extremely popular first language, it
       | makes sense that Python, pandas, SQL, etc. saw the largest drop
       | and average account age went up. From what I have seen on
       | programming subreddits the threads are mostly memes rather than
       | QA; to me the fact that they remain unaffected tells me the
       | communities are fine.
       | 
       | While you do need a constant funnel of new users to sustain a
       | community, I think the end result of this will be a higher
       | quality funnel that will eventually create an environment of
       | people solving niche problems.
       | 
       | All that to say that I now do a first pass through a tuned model
       | and only go to stackoverflow as a last resort. I skip Medium
       | search results because they are mostly people shilling their
       | company.
        
       ___________________________________________________________________
       (page generated 2024-07-31 23:01 UTC)