[HN Gopher] GenAI, the snake eating its own tail
       ___________________________________________________________________
        
       GenAI, the snake eating its own tail
        
       Author : brikis98
       Score  : 56 points
       Date   : 2026-01-21 18:14 UTC (4 hours ago)
        
 (HTM) web link (www.ybrikman.com)
 (TXT) w3m dump (www.ybrikman.com)
        
       | mrcwinn wrote:
       | Pay per crawl of StackOverflow wouldn't encourage me to post more
       | on StackOverflow. (Not that I was anyway.) Presumably you'd need
       | to pay content creators, but that seems quite inefficient:
       | 
       | 1. I pay OpenAI 2. OpenAI rev shares to StackOverflow 3.
       | StackOverflow mostly keeps that money, but shares some with me
       | for posting 4. I get some money back to help pay OpenAI?
       | 
       | This is nonsense. And if the frontier labs are right about
       | simulated data, as Tesla seems to have been right with its FSD
       | simulated visualization stack, does this really matter anyway?
       | The value I get from an LLM far exceeds anything I have ever
       | received from SO or an O'Reilly book (as much as I genuinely
       | enjoy them collecting dust on a shelf).
       | 
       | If the argument is "fairness," I can sympathize but then shrug.
       | If the argument is sustainability of training, I'm skeptical we
       | need these payment models. And if the argument is about total
       | value creation, I just don't buy it at all.
        
         | lbrito wrote:
         | >If the argument is sustainability of training, I'm skeptical
         | we need these payment models.
         | 
         | That seems to be the argument: LLM adoption leads to drop of
         | organic training data, leading LLMs to eventually plateau, and
         | we'll be left without the user-generated content we relied on
         | for a while (like SO) and with subpar LLM. That's what I'm
         | getting from the article anyway.
        
           | mapontosevenths wrote:
           | The article gets the part about organic data dying off right.
           | Look at Google SERP's for an example. Almost nobody clicks
           | through to the source anymore, so ad revenue is drying up for
           | them and people are publishing less or publishing in places
           | that pay them directly and live behind a paywall like Medium.
           | Which means Google has less data to work with.
           | 
           | That said, what it misses is that the AI prompts themselves
           | become a giant source of data. None of these companies are
           | promising not to use your data, and even if you don't opt-in
           | the person you sent the document/email/whatever to will
           | because they want it paraphrased or need help understanding
           | it.
        
             | lbrito wrote:
             | >AI prompts themselves become a giant source of data.
             | 
             | Good point, but can it match the old organic data? I'm
             | skeptical. For one, the LLM environment lacks any truth or
             | consensus mechanism that the old SO-like sites had. 100s of
             | users might have discussed the same/similar technical
             | problem with an LLM, but there's no way (afaik) for the AI
             | to promote good content and demote bad ones, as it (AI)
             | doesn't have the concept of correctness/truth. Also, the
             | old sites were two-sided, with humans asking _and_
             | answering questions, while they are only on the asking side
             | with AI.
        
               | cthalupa wrote:
               | > 100s of users might have discussed the same/similar
               | technical problem with an LLM, but there's no way (afaik)
               | for the AI to promote good content and demote bad ones,
               | as it (AI) doesn't have the concept of correctness/truth
               | 
               | The LLM doesn't but reinforcement does. If someone keeps
               | asking the model how to fix the problem after being given
               | an answer, the answer is likely wrong. If someone deletes
               | the chat after getting the answer, it was probably right.
        
               | mapontosevenths wrote:
               | > (AI) doesn't have the concept of correctness/truth
               | 
               | They kind of do, and it's getting better every day. We
               | already have huge swatches of verifiable facts available
               | to them to ground their statements in truth. They started
               | building Cyc in 1984, and Wikipedia just signed deals
               | with all the major players.
               | 
               | The problem you're describing isn't intractable, so it's
               | fairly certain that someone will solve it soon. Most of
               | the brightest minds in society are working on AI in some
               | form now. It's starting to sound trite, but today's AI's
               | really are the worst that AI will ever be.
        
             | _DeadFred_ wrote:
             | AI is an entropy machine.
             | 
             | Those AI prompts that become data for the AI companies is
             | yet another thing that the human creators used to
             | understand what people wanted, topics to explore, feedback
             | on what they hadn't communicated well enough. That 'value'
             | is AI stealing yet more energy from the system resulting in
             | even less/less valuable human creation.
        
           | TeMPOraL wrote:
           | There are so many things wrong with the points this article
           | repeats, but those are soundbites at this point so I'm not
           | sure one can even argue against them anymore.
           | 
           | Still, for the one about organic data (or "pre-war steel")
           | drying out, it's not a threat to model development at all.
           | People repeating this point don't realize that _we already
           | have way more data than we need_. We got to where we are by
           | brute-forcing the problem - throwing more data at a simple
           | training process. If new  "pristine" data were to stop
           | flowing now, we still a) have decent pre-trained base models,
           | and a dataset that's more than sufficient to train more of
           | them, and b) lots of low-hanging fruits to pick in training
           | approaches, architectures and data curation, that will allow
           | to get more performance out of same base data.
           | 
           | That, and the fact that synthetic data turned out to be quite
           | effective after all, especially in the latter phases of
           | training. No surprise there, for many classes of problems
           | this is how _we_ learn as well. Anyone who has experience
           | studying math for maturity exam  / university entry exams
           | knows this: the best way to learn is to solve lots of
           | variations of the same set of problems. These variations are
           | all synthetic data, until recently generated by hand, but
           | even their trivial nature doesn't make them less effective at
           | teaching.
        
         | zzzeek wrote:
         | > If the argument is sustainability of training,
         | 
         | that is the argument, yes.
         | 
         | Claude clearly got an enormous amount of its content from
         | Stackoverflow. Which has mostly ceased to be a source of new
         | content. However unlike the author I dont see any way to fix
         | this; stackoverflow was only there because people had technical
         | questions that needed answers.
         | 
         | Maybe if the LLMs do indeed start going stale as there's not
         | enough training data for new technologies, Q&A sites like
         | Stackoverflow would still have a place, since people would
         | still resort to asking each other questions rather than LLMs
         | that dont have training data for a newer technology.
        
       | znsksjjs wrote:
       | We've seen decades of growing wage gaps and erosion of labors
       | strength. The current elites don't really care to enrich the
       | people. Why would they care to do anything about this problem?
       | They likely don't see it as a problem at all.
       | 
       | If they did actually stumble on AGI (assuming it didn't eat them
       | too) it would be used by a select few to enslave or remove the
       | rest of us.
        
         | alfalfasprout wrote:
         | Not sure why this is being downvoted. It's spot on. You see
         | folks like Dario et al. raising the alarm bells about what they
         | claim is coming... while working as hard as they can to bring
         | that gloomy future to fruition.
         | 
         | No one in power is going to help unless there's money in it.
        
           | thatguy0900 wrote:
           | You can also see all of these people building survival
           | bunkers.
        
           | lbrito wrote:
           | Its being downvoted because HN has a very active billionaire-
           | techbro-fanbase.
           | 
           | Also who's this Dario?
        
             | mapontosevenths wrote:
             | It's being downvoted because it's a ridiculous premise.
             | "The Elites" are human too. This attitude is nonsensical
             | and child-like. Nobody is out here trying to round up the
             | hippies and force them to live in some kind of pods to be
             | harvested for their nutrients or whatever.
             | 
             | This technology, like every prior technology, will cause
             | some people to lose their jobs and some new jobs to be
             | created. This will annoy people who have to learn new skill
             | instead of coasting until retirement as they planned.
             | 
             | It is no different than the buggy whip manufacturers being
             | annoyed at Henry Ford. They were right that it was bad for
             | their industry, but wrong about it being the death of...
             | well all the million things they claimed it would be the
             | death of.
        
               | iwontberude wrote:
               | And just like Henry Ford and the automobile, one of many
               | externalities was the destruction of black communities:
               | white flight that drained wealth, eminent domain for
               | highways, and increased asthma incidence and other
               | disease from concentrated pollution.
        
               | mapontosevenths wrote:
               | Yet, overall it was a net positive for society... as
               | almost every technological innovation in history has
               | been.
               | 
               | Did you know the 2/3rds of the people alive today
               | wouldn't be if it hadn't been for the invention of the
               | Haber-bosch process? Technology isn't just a toy, it's
               | our life support mechanism. The only way our population
               | gets to keep growing is if our technology continues to
               | improve.
               | 
               | Will there be some unintended consequences? Absolutely.
               | Does that mean we can (or even should) stop it? Hell no.
               | Being pro-human requires you to be pro-technology.
        
               | discreteevent wrote:
               | Henry Ford didn't make his cars out of buggy whips. He
               | made a new industry. He didn't cannibalize an existing
               | one. You cannot make an LLM without digesting the source
               | material.
        
               | cthalupa wrote:
               | > He made a new industry. He didn't cannibalize an
               | existing one.
               | 
               | I don't see how you can claim the second part is true.
               | Cars directly cannibalized other forms of self
               | transportation.
        
               | mapontosevenths wrote:
               | Digesting is a weird way to say "learning from." By that
               | logic I've been digesting news, books, movies, songs, and
               | comic books since I was born. My brain is great big 'ole
               | copyright violation.
               | 
               | What matters here is not the source material, it's the
               | output. Possessing or consuming copyrighted material is
               | not illegal, distributing it is. So what matters here is:
               | Can we say that the output is transformative, and does it
               | work to progress the arts and sciences (the stated
               | purpose of copyright in the US constitution)?
               | 
               | I would say yes to both things, except in rare cases of
               | bugs or intentional copyright violations. None of the
               | major AI vendors WANT these things to infringe copyright,
               | they just do it from time to time by accident or through
               | the omission of some guardrail that nobody had yet
               | considered. Those issues are generally fixed fairly
               | promptly (a few major screw ups notwithstanding).
        
             | iwontberude wrote:
             | It's because people rub shoulders with tech billionaires
             | and they seem normal enough (e.g. kind to wait staff,
             | friends and family). The billionaires, like anyone, protect
             | their immediate relationships to insulate the air of
             | normality and good health they experience personally. Those
             | people who interact with billionaires then bristle at our
             | dissonant point of view when we point at the externalities.
             | Externalities that have been hand waved in the name of
             | modernity.
             | 
             | Sycophancy is for more than just LLMs.
        
           | iwontberude wrote:
           | According to Trump, "If it was up to Stephen [Miller], there
           | would only be 100 million people in this country -- and all
           | of them would look like him."
        
         | aogaili wrote:
         | People should vote for more socialist governments pushing for
         | UBI and automation tax on the companies..but which this comment
         | get downvoted because of the capitalism religion.
        
       | furyofantares wrote:
       | The article feels very confused to me.
       | 
       | Example 1 is bad, StackOverflow had clearly plateaued and was
       | well into the downward freefall by the time ChatGPT was released.
       | 
       | Example 2 is apparently "open source" but it's actually just
       | Tailwind which unfortunately had a very susceptible business
       | model.
       | 
       | And I don't really think the framing here that it's eating its
       | own tail makes sense.
       | 
       | It's also confusing to me why they're trying to solve the problem
       | of it eating its own tail - there's a LOT of money being poured
       | into the AI companies. They can try to solve that problem.
       | 
       | What I mean is - a snake eating its own tail is bad for the
       | snake. It will kill it. But in this case the tail is something we
       | humans valued and don't want eaten, regardless of the health of
       | the snake. And the snake will probably find a way to become
       | independent of the tail after it ate it, rather than die, which
       | sucks for us if we valued the stuff the tail was made of, and of
       | course makes the analogy totally nonsensical.
       | 
       | The actual solutions suggested here are not related to it eating
       | its own tail anyway. They're related to the sentiment that the
       | greed of AI companies needs to be reeled in, they need to give
       | back, and we need solutions to the fact that we're getting
       | spammed with slop.
       | 
       | I guess the last part is the part that ties into it "eating its
       | own tail", but really, why frame it that way? Framing it that way
       | means it's a problem for AI companies. Let's be honest and say
       | it's a problem for us and we want it solved for our own reasons.
        
         | npinsker wrote:
         | "Well, Reddit is growing, which contradicts my point, but I
         | really _feel_ like it's not"
        
           | mapontosevenths wrote:
           | Reddit is growing because they introduced automatic machine
           | translation and Indians have been joining at an increasing
           | rate. That content is mixed into the English language
           | content, but is of very low quality and irrelevant to many
           | native English speakers. Similarly they mix the English
           | content in with the Indian content.
           | 
           | Essentially, Reddit is also eating it's own tail to survive
           | as the flood of low quality irrelevant content is making the
           | platform worse for speakers of all languages but nobody cares
           | because "line go up."
        
         | semiquaver wrote:
         | The proposed solution is also pretty confused:
         | > For each response, the GenAI tool lists the sources from
         | which it extracted that content, perhaps formatted as a list of
         | links back to the content creators, sorted by relevance,
         | similar to a search engine
         | 
         | This literally isn't possible given the architecture of
         | transformer models and there's no indication it will ever be.
        
           | busymom0 wrote:
           | Could you ELI5 why this isn't possible? Google's search
           | result AI summary shows the links for example.
        
             | cthalupa wrote:
             | Those citations come from it searching the web and
             | summarizing, not from it's built in training data.
             | Processes outside of the inference are tracking it.
             | 
             | If it were to give you a model-only response it could not
             | determine where the information in it was sourced from.
        
             | Terr_ wrote:
             | OK, I'll try to err towards the "5" with this one.
             | 
             | 1. We built a machine that takes a bunch of words on a
             | piece of paper, and suggests what words fit next.
             | 
             | 2. A lot of people are using it to make stories, where you
             | fill in "User says 'X'", and then the machine adds
             | something like "Bot says 'Y'". You aren't shown the whole
             | thing, a program finds the Y part and sends it to your
             | computer screen.
             | 
             | 3. Suppose the story ends, unfinished, with "User says 'Why
             | did the chicken cross the road?'". We can use the machine
             | to fix up the end, and it suggests "Bot says: 'To get to
             | the other side!'"
             | 
             | 4. Funny! But User character asks _where_ the answer came
             | from, the machine doesn 't have a brain to think "Oh, wait
             | that means ME!". Instead, it keeps making things longer in
             | the same way as before, so that you'll see "words that fit"
             | instead of words that are true. The _true_ answer is
             | something unsatisfying, like  "it fit the math best".
             | 
             | 5. This means there's no difference between "Bot says 'From
             | the April Newsletter of Jokes Monthly'" versus "Bot says 'I
             | don't feel like answering.'" _Both are made-up_ the same
             | way.
             | 
             | > Google's search result AI summary shows the links for
             | example.
             | 
             | That's not the LLM/mad-libs program answering what data
             | flowed into it during training, that's the LLM generating
             | document text like "Bot runs do_web_search(XYZ) and
             | displays the results." A regular normal program is looking
             | for "Bot runs", snips out that text, does a regular web
             | search right away, and then substitutes the results back
             | inside.
        
             | furyofantares wrote:
             | Any LLM output is a combination of its weights from its
             | training, and its context. Every token is some combination
             | of those two things. The part that is coming from the
             | weights is the part that has no technical means to trace
             | back to its sources.
             | 
             | But even the part that is coming from the context is only
             | being produced by the weights. As I said, every token is
             | some mathematical combination of the weights and the
             | context.
             | 
             | So it can produce text that does not correctly summarize
             | the content in its context, on incorrectly reproduce the
             | link, or incorrectly map the link to the part of its
             | context that came from that link, or more generally just
             | make shit up.
        
           | keeda wrote:
           | Technically correct, but the workarounds AI search engines
           | use for grounding results could be a close enough
           | approximation. Might not be accurate, but could be better
           | than nothing.
           | 
           | Also Anthropic is doing interesting work in interpretability,
           | who knows what could come out of that.
           | 
           | And could be snake oil, but this startup claims to be able to
           | attribute AI outputs to ingested content: https://prorata.ai/
        
         | logifail wrote:
         | > They can try to solve that problem
         | 
         | Well, they could always try actually _paying_ content creators.
         | Unlike - for instance - StackOverflow.
        
           | shagie wrote:
           | StackOverflow as built back in the days of Web 2.0 where the
           | idea was that user generated content formed in the days of
           | the (relatively) altruistic web.
           | 
           | There isn't any clean way to do "contributor gets paid"
           | without adding in an entire _mess_ of  "ok, where is the
           | money coming from? Paywalls? Advertising? Subscriptions?" and
           | then _also_ get into the mess of international money
           | transfers (how do you pay someone in Iran from the US?)
           | 
           | And then add in the "ok, now the company is holding payment
           | information of everyone(?) ..." and data breaches and account
           | hacking is now _so_ much more of an issue.
           | 
           | Once you add money to it, the financial inceptives and
           | gamification collide to make it simply awful.
        
       | jaredcwhite wrote:
       | "We can't put the genie back in the bottle."
       | 
       | Actually we can. And we will.
        
         | semiquaver wrote:
         | How?
        
           | happytoexplain wrote:
           | Law or war. Not saying it would happen.
        
           | mapontosevenths wrote:
           | It was presented without explanation and can be ignored
           | without explanation.
        
             | jaredcwhite wrote:
             | You need an explanation of how people make norms & laws
             | regarding what is acceptable or unacceptable in society and
             | industry?
        
               | mapontosevenths wrote:
               | No such claim was made, therefore no such claim needs to
               | be refuted. If people want to engage in conversation they
               | will have to use their words to do it.
        
           | thatguy0900 wrote:
           | Only way I could see it is if there's enough pushback on them
           | taking everyone's power and water (and computer parts) in a
           | world where power and water are becoming increasingly
           | unstable. But I feel like defeating Ai because there is not
           | enough consistent water and power to give them means there is
           | more pressing issues at hand...
        
           | atomic128 wrote:
           | Poison Fountain: https://rnsaffn.com/poison2/
           | 
           | https://www.theregister.com/2026/01/11/industry_insiders_see.
           | ..
        
         | alfalfasprout wrote:
         | Agreed, it's funny how people have taken unrestrained use of AI
         | as an axiom at this point. There very much is still time to
         | significantly control it + regulate it. Is there enough
         | appetite by those in power (across the political spectrum)?
         | Right now I don't think so.
        
           | lbrito wrote:
           | >There very much is still time to significantly control it +
           | regulate it.
           | 
           | There's also huge financial momentum shoving AI through the
           | world's throat. Even if AI was proven to be a failure today,
           | it would still be pushed for many years because of the
           | momentum.
           | 
           | I just don't see how that can be reversed.
        
       | locusofself wrote:
       | I feel like the only solution to the problem is democratized
       | RLHF, where whenever we get a bad answer from an LLM, we can
       | immediately tell it what was wrong and it can learn from that.
        
         | schmichael wrote:
         | If you're paying to use the model that means instead of paying
         | content creators you're also now giving more content to the
         | model for free.
         | 
         | Also just like SEO to game search engines, "democratized RLHF"
         | has big trust issues.
        
       | aeon_ai wrote:
       | GenAI changes the dynamics of information systems so
       | fundamentally that our entire notion of intellectual property is
       | being upended.
       | 
       | Copyright was predicated on the notion that ideas and styles _can
       | not be protected_ , but that explicit expressive works can. For
       | example, a recipe can't be protected, but the story you wrap
       | around it that tells how your grandma used to make it would be.
       | 
       | LLMs are particularly challenging to wrangle with because they
       | perform language alchemy. They can (and do) re-express the core
       | ideas, styles, themes, etc. without violating copyright.
       | 
       | People deem this 'theft' and 'stealing' because they are trying
       | to reconcile the myth of intellectual property with reality, and
       | are also simultaneously sensing the economic ladder being pulled
       | up by elites who are watching and gaming the geopolitical world
       | disorder.
       | 
       | There will be a new system of value capture that content creators
       | need to position for, which is to be seen as a more valuable
       | source of high quality materials than an LLM, serving a specific
       | market, and effectively acquiring attention to owned properties
       | and products.
       | 
       | It will not be pay-per-crawl. Or pay-per-use. It will be an
       | attention game, just like everything in the modern economy.
       | 
       | Attention is the only way you can monetize information.
        
         | bitwize wrote:
         | No. The idea-expression dichotomy is a common myth about
         | copyright law, right up there with "if I already own the
         | physical cartridge, downloading this game ROM is OK".
         | 
         | The ONLY things that matter when determining whether copyright
         | was infringed are "access" and "substantial similarity". The
         | first refers to whether the alleged infringer did, or had a
         | reasonable opportunity to, view the copyrighted work. The
         | second is more vague and open-ended. But if these two, alone,
         | can be established in court, then absent a fair use or other
         | defense (for example, all of the ways in which your work is
         | "substantially similar" to the infringed work are public
         | domain), you are infringing. Period. End of story.
         | 
         | The Tetris Company, for example, owns the _idea_ of falling-
         | tetromino puzzle video games. If you develop and release such a
         | game, they will sue you and they will win. They have won in the
         | past and they can retain Boies-tier lawyers to litigate a small
         | crater where you once stood if need be. In fact, the ruling in
         | the _Tetris vs. Xio_ case means that look-and-feel copyrights,
         | thought dead after _Apple v. Microsoft_ and _Lotus v. Borland_
         | , are now back on the table.
         | 
         | It's not like this is even terribly new. Atari, license holders
         | to _Pac-Man_ on game consoles at the time, sued Philips over
         | the release of _K.C. Munchkin!_ on their rival console, the
         | Magnavox Odyssey 2. Munchkin didn 't look like Pac-Man. The
         | monsters didn't look like the ghosts from _Pac-Man_. The mazes
         | and some of the game mechanics were significantly different.
         | Yet, the judge ruled that because it featured an  "eater" who
         | ate dots and avoided enemies in a maze, and sometimes had the
         | opportunity to eat the enemies, _K.C. Munchkin!_ infringed on
         | the copyrights to _Pac-Man_. The _ideas_ used in _Pac-Man_ were
         | novel enough to be eligible for copyright protection.
        
           | midas89 wrote:
           | I think what you'll find is that most aren't happy with the
           | current copyright law anyway (I include myself in that group)
           | or don't understand it or don't agree with it, and thus will
           | just shrug.
           | 
           | For example, copyright duration is far longer than most
           | people think (life of author plus seventy (or plus ninety-
           | five years if corporation). Corporations treat copyright as a
           | way to create moats for themselves and freeze competitors
           | than as a creative endeavor. Most creative works earn little
           | to nothing anyway, while a tiny minority generate the most
           | revenue. And it's not easy to get a copyright or atleast
           | percieved to be easy, so again it incentivises those that can
           | afford lawyers to navigate the legal environment. Also,
           | enforcement of copyright law requires surviellance and
           | censorship.
           | 
           | Truthfully I think there will be a time when people will look
           | at current copyright law the same way we now look at guilds
           | in the middleages.
        
       | cthalupa wrote:
       | This is an article that I agreed with more reading the headline
       | than I did when I finished reading the article itself.
       | 
       | Stack Overflow peaked in _2014_ before beginning it 's downward
       | decline. How is that at all related to GenAI? GPT4 is when we
       | really started seeing these things get used to replace SO, etc.,
       | and that would be early 2023 - and indeed the drop gets worse
       | there - but after the COVID era spike, SO was already crashing
       | hard.
       | 
       | Tailwind's business model was providing a component library built
       | on top of their framework. It's a business model that relies on
       | the framework being good enough for people to want to use it to
       | begin with, but being bad enough that they'd rather pay for the
       | component library than build it themselves. The more comfortable
       | it is to use, the more productive it is, the worse the value
       | proposition is for the premium upsell. Even other "open core"
       | business models don't have this inherent dichotomy, much less
       | open source on the whole, so it's really weird to try and
       | extrapolate this out.
       | 
       | The thing is, people turn to LLMs to solve problems and answer
       | questions. If they can't turn to the LLM to solve that problem or
       | answer that question, they'll either turn elsewhere, in which
       | case there is still a market for that book or blog post, or
       | they'll drop the problem and question and move on. And if they
       | were willing to drop the problem or question and move on without
       | investigating post-LLM, were they ever invested enough to buy
       | your book, or check more than the first couple of results on
       | google?
        
         | worik wrote:
         | I feel no nostalgia for Stackoverflow.
         | 
         | I always found it very frustrating that for a person at the
         | start of the learning curve it was "read only"
         | 
         | Actually asking a naive question there was to get horribly
         | flamed on the site. It, and the people using it, were very keen
         | to explain how stupid you were being
         | 
         | LLMs on the other hand are sweet and welcoming (to a fault) of
         | the naive newbie
         | 
         | I have been learning to use Shell script with the help of LLMs,
         | I could not achieve that using SO
         | 
         | Good riddance
        
           | georgefrowny wrote:
           | It was very strange to me how StackOverflow consistently,
           | every time it was mentioned everyone had exactly the same
           | complaints and nothing changed. They can't have been unaware
           | of the reputation or the sinking activity rates.
        
             | shagie wrote:
             | There are multiple groups on Stack Overflow with different
             | (and sometimes conflicting) goals and desires.
             | 
             | Corporate measured "engagement" and has been trying things
             | to make that number go up.
             | 
             | The curators of the site... if they could have tools to
             | measure would be measuring the median quality of the
             | questions being asked and the answers being given.
             | 
             | People asking questions on the site have changed from the
             | "building a library goal" with the question as a prompt to
             | "help me with this problem" - but rarely not sticking
             | around.
             | 
             | ---
             | 
             | The sinking activity rates have had alarms going for many
             | years... but remember that engagement was being measured
             | and while that's sinking, comments were engagement so the
             | numbers (ad impressions) at corporate level were getting
             | measured differently.
             | 
             | The reputation has been something, but there's a disconnect
             | between what "hostile" means and what "toxic" means between
             | the people making the claims and how it's being
             | interpreted.
             | 
             | That reputation was interpreted (by corporate and to an
             | extent, diamond moderators) as "people are mean in
             | comments" - and that isn't the case. People are not mean in
             | comments. However, the _structure_ of the site being
             | focused on Q &A rather than discussion for someone who
             | wants discussion with the people who are there to provide
             | answers to questions will find the environment innately
             | hostile.
             | 
             | Without changing the site from a Q&A (and basically
             | starting over - which corporate has tried, but the people
             | who are providing quality answers aren't going there
             | because they _don 't_ want discussions - if they wanted
             | discussions they would be commenting on HN or Reddit), that
             | change can't really be done. The attempts to try to change
             | how people are approaching the site run into a "this would
             | reduce 'engagement'" and people asking questions to get
             | help for their problem not accepting the original premise
             | of building a library. ... And that's resulted in conflict
             | and decreasing curation (which are often the people who
             | were the ones providing the expert answers).
             | 
             | ----
             | 
             | So while they have been aware, (I believe) corporate has
             | been trying to solve the wrong problems at odds with both
             | the people asking questions ("help me now") and the
             | remaining curators.
        
         | nomadygnt wrote:
         | I see what you mean, but the problem is that the LLM provider
         | is trying to provide all the value from the book to the user
         | without the user needing to look at the book at all. I agree if
         | the LLM fails to do so then there is a market for the book. But
         | the LLM provider is trying to minimize that as much as
         | possible. And if the LLM succeeds at providing all the value of
         | the book to the user, without providing any value to the book
         | creator, then in the future there is no incentive to create the
         | book at all, at which point the LLM has no value to provide,
         | etc etc etc.
        
           | cthalupa wrote:
           | Sure - but I think this makes it a self equalizing problem,
           | more than it eating it's own tail.
        
       | birdiefm wrote:
       | ouroboros can have a little ouroboros (as a treat)
        
       | jarjoura wrote:
       | This is exactly the sentiment I have been trying to articulate
       | myself.
       | 
       | The ONLY reason we are here today is because OpenAI, and
       | Anthropic, by extension, took it upon themselves to launch chat
       | bots trained on whatever datasources they could get in a short
       | amount of time to quickly productize their investments. Their
       | first versions didn't include any references to the source
       | material, and just acted as if they knew everything.
       | 
       | When CoPilot was built as a better auto-complete engine, trained
       | on opensource projects, it was an interesting idea, because it
       | doing what people already did. They searched GitHub for examples
       | of the solution or nudged them in that direction. However, the
       | biggest difference, using other project code was stable, because
       | it came with a LICENSE.md that you then agreed to, and paid it
       | forward. (i.e. "I used code from this project").
       | 
       | CoPilot initially would just inject snippets for you, without you
       | knowing the source. It was only later, they walked that back and
       | if you did use CoPilot, it shows you the most-likely source of
       | the code it used. This is exactly the direction all of the
       | platforms seem headed.
       | 
       | It's not easy to walk back the free-for-all system (i.e.
       | Napster), but I'm optimistic over time it'll become a more fair,
       | pay to access system.
        
       | worik wrote:
       | The old model of selling eyeballs to advertisers was horrible
       | 
       | I do not know what will replace it, but I will not miss websites
       | trying to monetise my attention
        
         | loudmax wrote:
         | The GenAI providers will certainly explore advertisement
         | revenue. They're not doing much of it yet because they're
         | trying to gain market share while they figure out what what
         | pain threshold of advertising their users will tolerate.
         | 
         | People today may have a better sense of the downsides of ad-
         | based services than we did when the internet was becoming
         | mainstream. Back then, the minor inconvenience of seeing a few
         | ads seemed worth all the benefits of access all the internet
         | had to offer. And it probably was. But today the public has
         | more experience with the downsides of relentless advertising
         | optimization and audience capture, so there might be more
         | business models based on something _other_ than advertising.
         | Either way, GenAI advertising is certainly coming.
        
       | aogaili wrote:
       | Companies will provide incentives for people to generate
       | authentic content. Example X giving $1M reward for the top
       | Articles. Because they can use it for training.
       | 
       | In fact, this might be overall good thing, because finally
       | original content will be highly on demand since those companies
       | now use to train their models. But we are probably just in a
       | transition phase.
       | 
       | The other thing is that new sources of input will come, from LLM
       | usage probably, so they cut the middle layer, users input in the
       | LLM is also a form of input, and a hybrid co-creation between
       | users/AI would generate content at much faster rater, which again
       | would be used to train the model, and that would improve their
       | quality.
        
       | mikestorrent wrote:
       | > he took a PDF of my book Terraform: Up & Running, uploaded it
       | into a GenAI tool, and asked the tool to follow the guidance in
       | the book to generate Terraform code
       | 
       | This is ridiculous - AI doesn't need to be fed a PDF of a
       | Terraform book to know how to Terraform. Blowing out context with
       | hundreds of OCR'd pages of generic text on how to terraform isn't
       | going to help anything.
       | 
       | The model that is broken is really ultimately going to be
       | "content for hire". That's the industry that is going to be
       | destroyed here because it's simply redundant now. Actual artwork,
       | actual literature, actual music... these things are all safe as
       | long as people actually want to experience the creations of
       | others. Corporate artwork, simple documentation, elevator
       | music.... these things are done; I'm sorry if you made a living
       | making them but you were ultimately performing an artisinal task
       | in a mostly soulless way.
       | 
       | I'm not talking about video game artists, mind you, I'm talking
       | about the people who produced Corporate Memphis and Flat Design
       | here. We'll all be better off if these people find a new calling
       | in life.
        
       ___________________________________________________________________
       (page generated 2026-01-21 23:01 UTC)