[HN Gopher] Show HN: Slop or not - can you tell AI writing from ...
       ___________________________________________________________________
        
       Show HN: Slop or not - can you tell AI writing from human in
       everyday contexts?
        
       I've been building a crowd-sourced AI detection benchmark. Two
       responses to the same prompt -- one from a real human (pre-2022,
       provably pre prevalence of AI slop on the internet), one generated
       by AI. You pick the slop. Three wrong and you're out.  The dataset:
       16K human posts from Reddit, Hacker News, and Yelp, each paired
       with AI generations from 6 models across two providers (Anthropic
       and OpenAI) at three capability tiers. Same prompt, length-matched,
       no adversarial coaching -- just the model's natural voice with
       platform context. Every vote is logged with model, tier, source,
       response time, and position.  Early findings from testing: Reddit
       posts are easy to spot (humans are too casual for AI to mimic), HN
       is significantly harder.  I'll be releasing the full dataset on
       HuggingFace and I'll publish a paper if I can get enough data via
       this crowdsourced study.  If you play the HN-only mode, you're
       helping calibrate how detectable AI is on here specifically.  Would
       love feedback on the pairs -- are any trivially obvious? Are some
       genuinely hard?
        
       Author : eigen-vector
       Score  : 5 points
       Date   : 2026-03-12 21:53 UTC (1 hours ago)
        
 (HTM) web link (slop-or-not.space)
 (TXT) w3m dump (slop-or-not.space)
        
       | lucastonelli wrote:
       | Hey, congratulations on the final product. It even feels fun.
       | Some are really hard, but some feel blatantly obvious. I don't
       | know why though. I guess it's just because the way we communicate
       | feels off when compared to AI, some times.
        
         | eigen-vector wrote:
         | Thanks for checking it out! The obvious ones are (hopefully)
         | weaker models :) but yes my experience has been unless you're
         | engaging with human written content consistently the line
         | really blurs easily.
        
       | SsgMshdPotatoes wrote:
       | Nice idea! Em dashes were giveaways for AIs and typos for human,
       | at least in the ones I did, so those are at least trivial. So
       | might have to do some filtering at least for those.
       | 
       | Some were hard though, yeah (at least if not looking longer than
       | 5-10 seconds). Btw, it seemed more logical to me to just see a
       | green/red card when you click, i.e. right choice or wrong choice.
       | Getting red for the correct answer confused me a bit (but this
       | might just be me).
        
         | lucastonelli wrote:
         | The coloring is a fair point. I was some times confused if I
         | got the right or the wrong one XD
        
         | eigen-vector wrote:
         | Thanks for checking it out! The color signal is useful
         | feedback. Let me think about it and rework!
         | 
         | Yeah there are some very obvious tells, but the models that are
         | most capable are very good at writing like human.
         | 
         | Especially when the human responses for reddit or HN prompts
         | were presumably made after reading the content of the article
         | or the post; whilw the model is simply going off of the title.
        
         | SsgMshdPotatoes wrote:
         | Also for example this one has a giveaway for the human case:
         | "There are lots of great people here at /r/personalfinance"
         | (actually, not sure if that is a giveaway, that was my guess,
         | but depends on how the model was prompted, I guess). And human
         | ones often seem to have two spaces sometimes instead of one,
         | idk why. If you want to get a serious dataset, maybe you could
         | use this one to find all the flaws and perfect it, and then try
         | to get a real dataset from the next one? People will be more
         | eager to help too if they've seen you designed it all very
         | carefully. (Or you could filter the results from this one to
         | make it a good dataset if you get lots of responses.)
        
           | eigen-vector wrote:
           | You'd be surprised at the nuances we tend to miss :)
           | 
           | This time around I prompted the models not necessarily to be
           | adversarial - i didn't ask them to try and fool the reader.
           | But i gave them contextual info - something to the effect of
           | "you're a user posting on hacker news"
        
       ___________________________________________________________________
       (page generated 2026-03-12 23:01 UTC)