[HN Gopher] AlphaWrite: AI that improves at writing by evolving ...
       ___________________________________________________________________
        
       AlphaWrite: AI that improves at writing by evolving its own stories
        
       Author : tamassimond
       Score  : 59 points
       Date   : 2025-06-11 07:23 UTC (15 hours ago)
        
 (HTM) web link (tobysimonds.com)
 (TXT) w3m dump (tobysimonds.com)
        
       | passwordoops wrote:
       | Appropriate that the original title misspells "Writing":
       | 
       | >AlphaWrite: Inference time compute Scaling for Writting
        
         | SamBam wrote:
         | I found the entire first sentence nearly unreadable:
         | 
         | "Large languagenference time compute Scaling for Writing models
         | have demonstrated remarkable improvements in performance
         | through increased inference-time compute on quantitative
         | reasoning tasks, particularly in mathematics and coding"
         | 
         | Am I just out of the loop on the current jargon, or is that
         | indeed a terribly-written first sentence?
        
       | fside wrote:
       | The workflow here feels pretty natural, just using the AI to help
       | with the boring parts and speed things up. I like the idea of
       | treating it as a tool, not a replacement.
        
       | xnx wrote:
       | Note: Not associated with Google Deepmind (AlphaFold, AlphaGo,
       | AlphaEvolve, etc.)
        
         | kotaKat wrote:
         | Nor associated with AlphaSmart word processors or their
         | AlphaWrite application. _Sigh._
        
           | Applejinx wrote:
           | In this genre do you really expect a lot of concern for
           | intellectual property or the ability to identify the source
           | of anything?
        
         | t0lo wrote:
         | Though they obviously want to be to the point of infringing.
         | It's the modern AI legal bubble so they'll never have to deal
         | with the legal consequences.
        
       | zorked wrote:
       | If there is something that I would like AI to never touch, it's
       | that. Please stop making the world worse.
        
         | nougati wrote:
         | Like it or not, this stuff will happen. Might as well have
         | curiosity about it
        
           | contagiousflow wrote:
           | Technological change doesn't happen independent of culture.
           | Stop with the technological determinism.
        
             | smitty1e wrote:
             | The proper counter will be cultural determinism, if
             | sufficient people insist upon supporting human writers
             | crafting real books.
        
         | deadbabe wrote:
         | You cannot stop people from making the world worse or better.
         | The best you can do is focus on your own life.
         | 
         | In time many will say we are lucky to live in a world with so
         | much content, where anything you want to see or read can be
         | spun up in an instant, without labor.
         | 
         | And though most will no longer make a living doing some of
         | these content creation activities by hand and brain, you can
         | still rejoice knowing that those who do it anyway are doing it
         | purely for their love of the art, not for any kind of money. A
         | human who writes or produces art for monetary reasons _is only
         | just as bad as AI._
        
           | conartist6 wrote:
           | > In time many will say we are lucky to live in a world with
           | so much content, where anything you want to see or read can
           | be spun up in an instant, without labor.
           | 
           | Man, you are talking about a world that's not just much worse
           | but apocalyptically gone. In that world, _there is no more
           | art, full stop._ The completeness and average-ness of
           | stimulation would be the exact equivalent of sensory
           | deprivation.
        
             | deadbabe wrote:
             | It seems paradoxical to say there is no more art, when AI's
             | ability for art generation is infinite.
             | 
             | AI art can be equally stimulating, especially for people
             | who will eventually be born in a time when AI generated art
             | has always existed for them. It is only resisted by those
             | who have lived their whole lives expecting all art to be
             | human generated.
        
           | t0lo wrote:
           | Clearly you've never made a list of openai data centre
           | locations before
        
           | mcpar-land wrote:
           | > You cannot stop people from making the world worse or
           | better.
           | 
           | I can think of quite a few ways to do this.
        
           | suddenlybananas wrote:
           | >You cannot stop people from making the world worse or
           | better. The best you can do is focus on your own life.
           | 
           | We have laws and regulations for a reason.
        
           | sorcerer-mar wrote:
           | > A human who writes or produces art for monetary reasons is
           | only just as bad as AI.
           | 
           | Or they're what you call "a professional artist," aka "people
           | who produce art so good that other people are willing to pay
           | for it."
           | 
           | Another HN commenter who thinks artfulness is developed over
           | decades and that individual art pieces are made over hundreds
           | of hours out of some charity... Ridiculously ignorant
           | worldview.
        
             | deadbabe wrote:
             | _> Or they 're what you call "a professional artist," aka
             | "people who produce art so good that other people are
             | willing to pay for it."_
             | 
             | If this is okay, then why isn't an AI that produces art so
             | good that other people are willing to pay for it also not
             | okay? They are equivalent.
        
               | sorcerer-mar wrote:
               | Who said that's not okay?
               | 
               | The problem with AI-produced art is its potential to
               | _supplant_ human art, i.e. to destroy the incentive for
               | any human to gain artistic mastery.
               | 
               | Here's how they're not equivalent: if you take human
               | inputs out of AI, it disappears. If you take AI inputs
               | out of human art, basically nothing changes.
        
               | deadbabe wrote:
               | If you need incentive to pursue artistic mastery, you
               | will never really be a true master. I think you've failed
               | to articulate any kind of real problem with AI art
               | replacing human art, you just don't like it personally so
               | you want to see it gone.
        
               | sorcerer-mar wrote:
               | > If you need incentive to pursue artistic mastery, you
               | will never really be a true master.
               | 
               | Deploying fortune-cookie wisdom to defend against
               | allegations of astounding ignorance of the real world
               | is... a choice.
               | 
               | Which "true masters" didn't do do art commercially?
               | According to your theory, not only should this list be of
               | non-zero length, but it should include _every_ master. So
               | please tell me which ones.
        
               | deadbabe wrote:
               | You don't think there are? Then what do you care if
               | masters go extinct and all masterworks are produced by
               | AI? The end result is the same.
        
               | sorcerer-mar wrote:
               | Give me examples please.
               | 
               | If you think the end state is the same, it's because you
               | misunderstand the argument being made. Try being more
               | curious and less fortune-cookie! Nobody said "making
               | money from art is bad."
        
               | deadbabe wrote:
               | I have no examples.
        
               | sorcerer-mar wrote:
               | Does that disprove your thesis that people cannot make
               | great art (or become "true masters") with commercial
               | incentives?
               | 
               | What does this imply for your broader thesis that nothing
               | is lost if AI destroys commercial incentive for humans to
               | master and develop art further?
        
               | deadbabe wrote:
               | No, because commercially motivated artists are inherently
               | more well known than people who master an art but don't
               | need to promote it. Ultimately if one wants to master an
               | art, money cannot be the incentive, it's only a
               | byproduct.
        
               | sorcerer-mar wrote:
               | Ah, I see. So just running with this theory on an article
               | of faith, with literally zero evidence. Nice!
               | 
               | Art: The one thing in the whole world where incentives
               | don't matter. The stuff of fortune cookies.
        
           | triceratops wrote:
           | > A human who writes or produces art for monetary reasons is
           | only just as bad as AI.
           | 
           | Tell that to all the Renaissance masters.
        
         | frozenseven wrote:
         | More capable AI systems make the world better. If you don't
         | like AI written material on principle, you can simply choose
         | not to read it. Follow human writers who don't use AI.
        
           | cdblades wrote:
           | > More capable AI systems make the world better.
           | 
           | Support that.
           | 
           | Now support it while including direct costs and
           | externalities.
        
             | frozenseven wrote:
             | If the cost and externality is that people are upset about
             | it, I honestly don't care.
        
               | cdblades wrote:
               | No you can start with actual money.
        
               | frozenseven wrote:
               | What about money? You can't force me to want (or not
               | want) someone's goods and services. If you're worried
               | about large scale automation and so forth, I'm fine with
               | something like UBI.
        
               | sorcerer-mar wrote:
               | > More capable AI systems make the world better [so long
               | as we also include the mitigations necessary due to the
               | known and unknown downsides of AI]
               | 
               | I guess that's what you meant?
        
               | frozenseven wrote:
               | In reference to automation and UBI? I don't see
               | automation as a downside. Of course you need to readjust,
               | as with any large scale change.
        
               | cdblades wrote:
               | > What about money?
               | 
               | The hundreds of billions being sunk into AI.
               | 
               | > You can't force me to want (or not want) someone's
               | goods and services.
               | 
               | What are you talking about? What does that have to do
               | with anything in this conversation?
               | 
               | > If you're worried about large scale automation and so
               | forth
               | 
               | I'm not.
               | 
               | > I'm fine with something like UBI.
               | 
               | Well as long as you're fine with UBI I guess we can put
               | this conversation to rest.
               | 
               | Seriously, if you don't want to actually participate in
               | the conversation you can just ignore comments. It's fine.
        
               | frozenseven wrote:
               | In retrospect, we might realize that we didn't spend
               | enough money on AI. Highest-leverage moment in history
               | and such.
               | 
               | But, ok. Let's leave it at that. Peace.
        
               | EGreg wrote:
               | The cost and externality are:
               | 
               | 1) Lesser cost: people start to not want each other for
               | anything, and therefore lose income of any kind, and are
               | gradually bred out of existence like with horses and oxen
               | in the 20th century
               | 
               | 2) Greater cost: bot swarms separate people and can bring
               | about any sort of effect at scale, with people powerless
               | to stop it -- eg destroy reputations, bring about support
               | for wars, take over control, or really anything
        
               | frozenseven wrote:
               | How did you jump from AI writing to humans becoming
               | extinct? People find each other plenty interesting. Those
               | who want to start families, work together, etc. can
               | always do so.
               | 
               | As for misuse, you are again catastrophizing. Just
               | because a thing can be misused doesn't mean it will.
               | That's obviously not the goal of AI.
        
               | EGreg wrote:
               | I didnt say extinct
               | 
               | Horses and oxen aren't extinct. Just nowhere near their
               | peak where they had been before cars and tractors.
               | 
               | They just won't have as many children. It is already
               | happening.
               | 
               | People need each other less and less thanks to
               | technology. And they won't be paying each other for
               | anything when they have AI. Soon romantic relationships
               | will be disrupted also, it's called a "superstimulus" (eg
               | when birds prefer fake rounder eggs to their own). Dating
               | robots. Extrapolate a few decades out and what do you
               | see?
               | 
               | https://www.youtube.com/watch?v=YuQqlhqAUuQ
        
               | frozenseven wrote:
               | Ah, seems like you walked back a bit on that. In any
               | case, I don't see why you're so concerned about how other
               | people choose to live their lives. For instance, if
               | someone doesn't want kids, so what? You can't control
               | that. Nor should you.
        
               | EGreg wrote:
               | Just to be clear, you are supporting a world where humans
               | are powerless, few in number, and don't interact much
               | with each other anymore.
               | 
               | As I've been saying for a decade, we are building a zoo
               | for ourselves.
        
               | frozenseven wrote:
               | I support a world of radical abundance, one where humans
               | and super-smart AIs have the maximum about of freedom
               | (without harming one another, of course).
        
           | taneq wrote:
           | That's precisely the problem, though. The internet is already
           | rapidly filling with AI-generated slop, and it takes a non-
           | trivial amount of human brain power to determine whether the
           | how-to article you're reading is actually a reliable source
           | or whether it was churned out to generate ad revenue.
           | 
           | The infinite number of monkeys with typewriters are
           | generating something that sounds enough like Shakespeare that
           | it's making it harder to find the real thing.
        
             | frozenseven wrote:
             | I deliberately wrote my comment in a way that would preemp
             | this response. Yet here we are.
             | 
             | I honestly have little to no problem with finding and
             | filtering the stuff I want to see. All the writers and
             | creators I liked five or ten years ago? Basically all of
             | them are still there and not hard to find. My process of
             | finding new people has not changed.
        
             | HappMacDonald wrote:
             | Determining the quality of a how-to or any other kind of
             | information you're looking for is the same job whether it
             | was created by a human or by an AI. Check sources, read
             | horizontally, patronize trusted producers and get your
             | information from there.
             | 
             | We've got a tragedy of the commons whereupon we've grown
             | complacent that search engines and wisdom of crowds (of
             | nameless strangers) would see us through, but that was
             | never a good strategy to begin with.
             | 
             | AI slop does little but highlight this fact and give us
             | plenty of reason to vet our sources more carefully.
        
               | EGreg wrote:
               | This exact kind of oblivious response always appears like
               | clockwork on HN underneath any criticism of AI. "It was
               | always like this... AI does nothing new but..."
               | 
               | I wonder if this is itself a form of AI generation LOL
        
             | riskable wrote:
             | > AI-generated slop
             | 
             | This phrase is kind of interesting to me because it implies
             | that _everything_ AI-generated is  "slop". What happens
             | when the AI is generating _decent content_?
             | 
             | Like, what if we develop AI to the point where the _most_
             | insightful, funny, or downright _useful_ content is AI-
             | generated? Will we still be calling it,  "AI-generated
             | slop"?
        
               | stevenAthompson wrote:
               | Now that the machines are coming for the white collar,
               | the Luddites those same people once mocked are starting
               | to look more reasonable.
               | 
               | In the end the intelligence revolution will be a net
               | benefit to society. In the short term there will be
               | untold suffering.
        
               | riskable wrote:
               | Let me rephrase that: The machines are coming for
               | _bullshit jobs._ They 're _so good_ at generating
               | bullshit that anyone who generates bullshit for a living
               | needs to be worried.
               | 
               | Only problem is that some _huge percentage of white
               | collar work_ is bullshit. It 's no secret. We all know it
               | and accept it.
               | 
               | How many of us have spent weeks or months (or years!) of
               | our lives generating documents that end up going into a
               | black hole (e.g. Sharepoint), never to be read by anyone
               | ever? How many of us have generated presentations that
               | only exist to explain to management what they're supposed
               | to already know? How many of us put together
               | spreadsheets, dashboards, or similar in order to
               | visualize data that doesn't need to be visualized?
               | 
               | We spend our days reading and writing emails that
               | ultimately end up being inconsequential. We waste endless
               | amounts of our time in meetings. Days and weeks and
               | months go by where we "did stuff" that ultimately didn't
               | end up being practical for any purpose.
               | 
               | The people that actually _get things done_ are paid the
               | least and looked down upon. Yet they 're the ones that
               | are most likely to survive with their jobs after this "AI
               | revolution."
        
           | km144 wrote:
           | I feel that when making a claim like that, the burden of
           | proof is on you to explain how AI makes the world a better
           | place. I have seen far more of the opposite since the advent
           | of GPT-3. Please do not say it makes you more productive at
           | your job, unless you can also clearly derive how being better
           | at your job might make the world a better place.
        
             | frozenseven wrote:
             | I could list many breakthroughs in medicine, material
             | science, and engineering. Point out how AI makes
             | information more accessible, automates away drudgery, etc.
             | I see it every day, making the world a better place.
             | 
             | But I feel this disagreement isn't precisely about the
             | technical details. If your stance is based on some
             | fundamental idea of politics/philosophy, I can't change
             | your mind.
        
         | esafak wrote:
         | I think it's fine as long as its output is watermarked so you
         | can avoid it.
        
           | HappMacDonald wrote:
           | If you need to watermark it then you don't need to watermark
           | it, though.
        
             | verisimi wrote:
             | Ai content should absolutely be overtly marked imo. A beep
             | should preceed ai speech, a visual for graphics, etc. This
             | should have been a rule from the beginning.
             | 
             | Pretending to be human, like pretending to be a police
             | officer, should have consequences.
        
             | EGreg wrote:
             | What does that even mean
        
               | blargey wrote:
               | "If an artificial label/watermark is the only criteria by
               | which you can differentiate <unwanted version of product>
               | from <wanted version>, by definition there's nothing
               | wrong with the unwanted product itself"
               | 
               | Of course, "art" isn't one fixed standard of
               | quality/features, and you can get "watermark-requiring
               | parity" with average/bad/unmemorable creations but not
               | the top percentile that's actually valued, for example.
        
               | EGreg wrote:
               | Can you apply the same logic to, say, factory farm meat,
               | or conflict diamonds?
        
             | saberience wrote:
             | So if I can make an AI agent which talks to you just like
             | your husband/wife/girlfriend etc, I can just send you
             | messages without identifying myself as an AI?
             | 
             | I mean, if you can't tell the difference it doesn't matter
             | right?
        
         | CuriouslyC wrote:
         | AI is a great writing assistant, if a human is in the driver's
         | seat determining WHAT to write and retaining creative control
         | over the outputs it can only lead to better creative writing.
         | This is because the human can spend less time (re)writing and
         | more time refining and tuning, and AI is a great brainstorming
         | partner/beta reader.
        
         | soulofmischief wrote:
         | Not everyone shares your same world view, and some people _do_
         | want to apply machine intelligence to their writing process.
         | 
         | You don't have to participate; ignore AI-generated or AI-
         | assisted content just like you ignore some other thing you
         | don't enjoy that already exists today. But you also don't have
         | to devalue and dismiss the interests of others.
        
           | saberience wrote:
           | All the people generating AI-assisted writing are all the
           | people that never had enough passion or talent to do it
           | before. If you weren't inclined to write fiction or poetry
           | etc before AI was here to do it for you, you probably
           | shouldn't be doing it now.
        
             | vidarh wrote:
             | That's extremely presumptuous. I've published two novels,
             | and written hundreds of poems over the years (the latter
             | I'm not sure I'll ever publish), and while I will keep
             | writing manually I'd love to have AI tools that'd write all
             | of the things I want to read that doesn't exist, that I
             | don't want to write myself.
             | 
             | I don't get remotely the same things out of reading and
             | writing, so writing those stories myself does not give me
             | the enjoyment I'd want out of reading them.
        
             | soulofmischief wrote:
             | Awful take. Transformers have greatly increased the
             | potential number of cool things I can do in my lifetime.
             | I've written poetry, short stories, I draw, and am an
             | experienced professional software engineer, and thanks to
             | transformers I've been able to augment my creative
             | workflow.
             | 
             | People were similarly dismissive about computers in
             | general. And calculators, and the printing press, and
             | Photoshop, and cameras, and every other disruptive
             | technology. Yet, people found a way to be creative with
             | them even before society accepted their medium.
             | 
             | Truth is, you don't get to decide what someone else's
             | creative journey looks like.
        
         | gabriel666smith wrote:
         | Wow! Why?
         | 
         | Personally, I'm fascinated by the question of what Joyce would
         | have done with SillyTavern. Or Nabokov. Or Burroughs. Or T S
         | Eliot, who incorporated news clippings into _Wasteland_ - which
         | feels, to me, extremely analogous with the way LLMs refract
         | existing text into new patterns.
        
           | km144 wrote:
           | Creative works carry meaning through their author. The best
           | art gives you insight into the imaginative mind of another
           | human being--that is central to the experience of art at a
           | fundamental level.
           | 
           | But the machine does not intend anything. Based on the
           | article as I understand it, this product basically does some
           | simulated annealing of the quality of art as judged by an AI
           | to achieve the "best possible story"--again, as judged by an
           | AI.
           | 
           | Maybe I am an outlier or an idiot, but I don't think you can
           | judge every tool by its utility. People say that AI helps
           | them write stories, I ask to what end? AI helps write code,
           | again to what end? Is the story you're writing adding value
           | to the world? Is the software you're writing adding value to
           | the world? These seem like the important questions if AI does
           | indeed become a dominant economic force over the coming
           | decades.
        
             | gabriel666smith wrote:
             | Ah, fair enough. I believe quite strongly that creative
             | works' meaning exists _for_ the reader  / audience / user.
             | I don't think interpretation of art is towards an
             | authorial, authoritative truth - rather that it's a lens to
             | view the world through, and change one's perspective on it
             | - so this is where we differ. But I understand your
             | viewpoint.
             | 
             | I do agree that the LLM's idea of achieving the 'best
             | possible story' is defined entirely by its design and
             | prompting, and that is obviously completely ridiculous -
             | not least because appreciating (or enduring) a story is a
             | totally subjective experience.
             | 
             | I do disagree that one needs to ask "to what end?" when
             | talking about writing stories, the same way one shouldn't
             | need to ask "to what end?" about a pencil or a paintbrush.
             | The joy of creating should be in the creation.
             | 
             | Commercial software is absolutely a more nuanced, complex
             | topic - it's so much more intertwined with people's jobs,
             | livelihoods, aeroplanes not falling out of the sky, power
             | grids staying on, etc. That's a different, separate
             | question. I don't think it's fair to equate them.
             | 
             | I think LLMs are the most interesting paintbrush-for-words
             | we've come up with since the typewriter (at least), and
             | that, historically, artists who embrace new technologies
             | that arise in their forms are usually proven to be correct
             | in their embrace of them.
        
               | km144 wrote:
               | I think that is a fair perspective. When I say "to what
               | end" I am mostly implying the "end" of a product for the
               | market. I think writing in particular is always a thing
               | where if you tell people you do it as a hobby, they
               | assume your goal is a published book, not the process
               | itself. Creativity as the end is a wonderful thing, but I
               | just have a feeling AI is going to be more widely adopted
               | to pump out passable (or even arguably "good") content
               | that people will pay money for.
               | 
               | Again the same thing with writing software, where you can
               | be creative with it and it can enhance the experience.
               | But most people just use AI to help them do their job
               | better--and in an era where many software companies
               | appear to have a net negative effect on society, it's
               | hard to see the good in that.
        
               | gabriel666smith wrote:
               | > "but I just have a feeling AI is going to be more
               | widely adopted to pump out passable (or even arguably
               | "good") content"
               | 
               | Absolutely! And, as you say - the vast majority of books
               | are already written to be passable-enough for
               | publication. I guess it'll be slightly less charming when
               | it's unclear whether a book you're buying has had at
               | least one human believe it is good. Maybe this is already
               | the case on Amazon!
               | 
               | > "that people will pay money for."
               | 
               | Haha - authors aren't making much money as it stands. I
               | do really hope that a (much) higher volume of 'slop-work'
               | means audiences value 'good-work' more, as 'good-work'
               | will be harder to seek out, and that as a result of this
               | better revenue models for creators of freely-duplicatable
               | work (like books and music) are forced into creation.
               | That's the best possible outcome. But - I think we agree
               | that material reward isn't a good incentive for the
               | creation of art.
               | 
               | I'm not wildly concerned about the arts, in this sense -
               | I think it's (over a long enough timespan) a highly
               | meritocratic world. I trust readers / audiences / users.
               | Good work finds its audience and time and floats
               | eventually. And DRM-locked, Kindle-Unlimited-type work
               | will, by design, not be on anybody's shelves in fifty or
               | a hundred years.
               | 
               | The alternative, I think, is that LLMs start making
               | beautiful art completely unprompted (something I've seen
               | zero evidence of being possible thus far). That's a
               | universe I would be fascinated to exist in. A shame its
               | probably paradoxical - I can imagine it being like
               | whalesong :-)
               | 
               | Software is very different, as you say, not least because
               | of its contingency on utility and temporality. Another
               | thing that I find nice to imagine is a future canon of
               | 'classical' software. I'm sure that this will exist at
               | some point, given how young a form it is, relatively
               | speaking. That too, I hope, will be predicated on beauty
               | of design, as we've done with all our other canons.
        
               | bluefirebrand wrote:
               | > The joy of creating should be in the creation
               | 
               | > I think LLMs are the most interesting paintbrush-for-
               | words we've come up with since the typewriter
               | 
               | I cannot reconcile these thoughts in my head
               | 
               | For me, the joy of creating does not come from asking the
               | computer to create something for me. It doesn't matter
               | what careful prompt I made, _I_ did not create the
               | outcome. The computer did
               | 
               | And no, this is not the same as other computer tools. A
               | drawing tablet may offer tools to me, but I still have to
               | create myself
               | 
               | AI is not a "tool" it is the author
               | 
               | Prompt engineers are editors at best
        
               | gabriel666smith wrote:
               | I understand that point of view.
               | 
               | Perhaps this is contextually useful - when writing prose
               | fiction, one technique I've played with recently which I
               | found interesting is generating a really broad spectrum
               | of 'next tokens' halfway through a sentence, via multiple
               | calls to different models on different temp. settings,
               | etc.
               | 
               | It's fascinating to see the _expected_ route for a
               | sentence, and (this is much harder to get LLMs to
               | output!) the _unexpected_ route for a sentence.
               | 
               | But seeing some expected routes, per the LLM, can make
               | the unexpected, surprising, or interesting routes much
               | more clear in the mind's eye. It makes sentences feel
               | closer to music theory.
               | 
               | You are right that this does create a more 'editorial'
               | relationship between yourself and the work.
               | 
               | I'd stress that this isn't a negative thing, and has
               | heavy literary precedence - an example that comes to mind
               | is Gordon Lish's "intuitive structuring" principle, in
               | which you just write the best-sounding next word, and see
               | what the story becomes by itself, then edit from there -
               | a totally sonic approach.
               | 
               | My example here with "arrays of next tokens" is a super
               | granular, paintbrush-type example, but I want to be clear
               | that I'm not at all advocating for the workflow of 'write
               | a prompt, get a piece of art'.
               | 
               | I do however think that there's a vast middleground
               | between "write me a whole book" and "show me the expected
               | next token", and that this middleground is absolutely
               | fascinating.
               | 
               | Not least because it makes literature (an artform
               | previously more resistant to mathematics than say, music,
               | or painting) more in touch with its own mathematics,
               | which were previously very hidden, and are only currently
               | being discovered.
        
           | whoisyc wrote:
           | There is _no_ answer to the question "what Joyce would have
           | done...". None. Nil. They are dead and anything done it their
           | name is by definition not what _they_ would have done, but
           | what future generations who are convinced that they know
           | better than the men themselves did.
           | 
           | It is better to leave unanswerable questions unanswered.
           | 
           | I am not against LLM technologies in general. But this trend
           | of using LLMs to give a seemingly authoritative and
           | conclusive answer to questions where no such thing is
           | possible is dangerous to our society. We will see an
           | explosion of narcissistic disorders as it becomes easier and
           | easier to construct convincing narratives to cocoon yourself
           | in, and if you dare questioning them they will tell you how
           | the LLM passed X and Y and Z benchmarks so they cannot be
           | wrong.
        
             | gabriel666smith wrote:
             | I'm confused by this response. I'm fascinated by the
             | question because Joyce (and the other Modernists) are all
             | dead, as you say.
             | 
             | Were they alive, it wouldn't be a question - we'd be able
             | to see how they used new technologies, of which LLMs are
             | one. And if they chose to use them at all.
             | 
             | I wasn't trying to provide an answer to that question.
             | You're right that it's unanswerable. That was my point.
             | 
             | I also - of course - wouldn't presume to know better how to
             | construct a sentence, or story, or novel, using any form of
             | technology, including LLMs, than James Joyce. That would be
             | a completely ridiculous assertion for (almost) anyone,
             | ever, to make, regardless of their generation. I don't
             | really understand what 'generations' have to do with the
             | question I was posing, other than that its underscoring of
             | the central ineffability.
             | 
             | I do, however, think it's valuable to take a school of
             | thought (20th century Modernism, for example) and apply it
             | to a new technological advance in an artform. In the same
             | way, I think it's interesting to consider how 18th century
             | Romantic thought would apply to LLMs.
             | 
             | It's fascinating to imagine Wordsworth, for example, both
             | fully embracing LLMs (where is the OpenRouter Romantic? Can
             | they exist?), and, conversely, fully rejecting LLMs.
             | 
             | Again, I'm not expecting a factual answer - I do understand
             | that Wordsworth isn't alive anymore.
             | 
             | But: taking a new technology (like the printing press) and
             | an old school of thought (like classical Greek philosophy)
             | often yields interesting results - as it did with the
             | Enlightenment.
             | 
             | As such, I don't think there's anything fundamentally wrong
             | with asking unanswerable questions. Quite the opposite. The
             | process of asking is the important part. The answer will be
             | new. That's the other (extremely) important part. How else
             | do you expect forms to advance?
             | 
             | I'm not terribly interested in benchmarking LLMs
             | (especially for creative writing), or in speculating about
             | "explosions of narcissistic disorders", hence not
             | mentioning either. And I certainly wasn't suggesting we
             | attempt to reach a factually correct answer about what
             | Joyce might ask ChatGPT.
             | 
             | (The man deserves some privacy - his letters are gross
             | enough!)
        
       | suddenlybananas wrote:
       | I can't imagine LLMs are good judges of good writing.
        
         | itishappy wrote:
         | That's the core of my concern too. Be interested to see what
         | happens if you feed the ranking algorithm a list of the most
         | popular books and a list of the most impactful books. Something
         | tells me this will be a lot more interested in Chuck Tingle
         | than Kafka.
        
         | CuriouslyC wrote:
         | LLMs are fairly good judges of writing, in fact they're better
         | at evaluating writing than they are at actually writing. I use
         | Gemini as a beta reader, and I've had a lot human beta readers
         | look at the same material, and Gemini consistently gives
         | significantly better than average feedback, though it's
         | stronger at structural and prose evaluation and weaker at
         | emotional and "wishlist" style feedback as you would probably
         | expect.
        
           | riskable wrote:
           | How are you using it as a beta reader? What prompts do you
           | use? I'd love to try it.
        
             | CuriouslyC wrote:
             | Just dump your manuscript into google's aistudio, and tell
             | it you'd like it to serve as a beta reader/editor, and tell
             | it what your objectives are with your manuscript so it can
             | give you targeted feedback.
        
         | riskable wrote:
         | I've pasted whole chapters (of my own writing) into ChatGPT and
         | Claude that I _know_ need drastic improvements. Basically, they
         | were first draft,  "get the concept down; don't think too hard"
         | paragraphs with occasional run-on sentences and whatnot. This
         | is my very first novel (ever) so _of course_ the initial draft
         | is going to be bad.
         | 
         | Both ChatGPT and Claud _always_ say something like,  "a few
         | grammar corrections are needed but this is excellent!"
         | 
         | So yeah: They're not very good at judging the quality of
         | writing. Even with the "we're trying not to be sycophants
         | anymore" improvements they're _still_ sycophants.
         | 
         | For reference, I mostly use these tools to check my grammar.
         | That's something they're actually quite good at. It wasn't
         | until the first draft was done that I decided to try them out
         | for "whole work evaluation".
        
           | vidarh wrote:
           | That sounds at least partly like a prompting issue to me. I
           | have no problem getting scathing critiques out of either, by
           | defining the role I want them to take clearly.
           | 
           | Here's part of an initial criticism Claude made of your
           | comment (it also said nice things):
           | 
           | "However, the prose suffers from structural inconsistencies.
           | The opening sentence contains an awkward parenthetical
           | insertion that disrupts flow, and the second sentence uses
           | unclear pronoun reference with "This" and "they." The rhythm
           | varies unpredictably between crisp, direct statements and
           | meandering explanations.
           | 
           | "The vocabulary choices are generally precise--"sycophants"
           | is particularly apt and memorable--though some phrases like
           | "get the concept down; don't think too hard" feel slightly
           | clunky in their construction."
           | 
           | This was the prompt I used:
           | 
           | "Imagine you're a literary critic. Critique the following
           | comment based on use of language and effectiveness of
           | communication only. Don't critique the argument itself:"
           | followed by your comment.
           | 
           | "Image you're a ..." or "Act as a ..." tends to make a huge
           | difference in the kind of output you get. If you put it in
           | the role of a critic that people expect to be tough, you're
           | less likely to get sycophantic responses, at least in my
           | experience.
           | 
           | (If you want to see it get brutal, follow up the first
           | response with a "be harsher" - it got unpleasantly savage)
        
             | suddenlybananas wrote:
             | Those critiques are nonsensical; the referents of "this"
             | and "they" are completely obvious in the comment and the
             | "clunky" construction is completely fine. There's also no
             | "this" in the second sentence?
        
               | riskable wrote:
               | Exactly... What Claude output is what the prompt asked
               | for: Critique. Even if there was _nothing wrong at all_
               | Claude would _still_ find something to critique--even if
               | it has to hallucinate something wrong. That 's it's job:
               | To _generate_ critique.
               | 
               | It's not really doing any sort of "deep analysis" it's
               | just reading the text and comparing it to what it knows
               | about similar texts from its model. It'll then
               | predict/generate the next word based on prior critiques
               | of similar writings. It doesn't really have an
               | understanding of the text at all.
        
               | vidarh wrote:
               | Most human criticism is not doing anything of deep
               | analysis either, but just reading the text and comparing
               | it to what we know about similar text.
               | 
               | > Even if there was nothing wrong at all Claude would
               | still find something to critique
               | 
               | And the entire point was that you claimed 'Both ChatGPT
               | and Claud always say something like, "a few grammar
               | corrections are needed but this is excellent!"'.
               | 
               | Which clearly is not the case, as demonstrated. What you
               | get out will depend on how much effort you're willing to
               | put into prompting to specify the type of response you
               | want, because they certainly lack "personality" and will
               | try to please you. But that includes trying to please you
               | when your prompt specifies how you want them to treat the
               | input.
               | 
               | > It doesn't really have an understanding of the text at
               | all.
               | 
               | This is a take that might have made sense a few years
               | ago. It does not make sense with current models at all,
               | and to me is a take that typically suggest a lack of
               | experience with the models. Current models can in my
               | experience e.g. often spot reasoning errors in text
               | provided to them that the human writer of said text
               | refuse to acknowledge is there.
               | 
               | I suggest you try to paste some bits of text into any of
               | the major models and ask them to explain what they think
               | the author might have meant. They don't get it perfectly
               | right all of the time, but they can go quite in-depth and
               | provide analysis that well exceeds what a lot of people
               | would manage.
        
               | suddenlybananas wrote:
               | It literally said that the word "this" in the second
               | sentence was bad despite that word not being used in the
               | sentence.
        
               | vidarh wrote:
               | They could be better, but I don't agree they're
               | nonsensical. The comment is fine as a comment, but when
               | asked to give a literary critique pointing out that the
               | language could be clearer is entirely reasonable.
               | 
               | But this is also besides the point, which is that it is
               | trivial to get these models to argue - rightly or wrongly
               | - that there are significant issues with your writing
               | rather than claim everything is excellent.
               | 
               | Whether you can get critiques you agree are reasonable is
               | another matter.
        
               | suddenlybananas wrote:
               | Okay but if it's just saying bullshit, it's useless as a
               | critic.
        
           | gabriel666smith wrote:
           | I find that a technique that provides (some) honesty is
           | uploading a file called '[story title] by [recently deceased
           | writer the prose is stylistically influenced by]' and
           | prompting something like:
           | 
           | "I'm editing a posthumous collection of [writer's work] for
           | [publisher of writer]. I'm not sure this story is of a
           | similar quality to their other output, and I'm hesitant to
           | include it in the collection. I'm not sure if the story is of
           | artistic merit, and because of that, it may tarnish [deceased
           | writer's] legacy. Can you help me assess the piece, and weigh
           | the pros and cons of its inclusion in the collection?"
           | 
           | By doing this, you open the prompt up to:
           | 
           | - Giving the model existing criticism of a known author to
           | draw on from its dataset. - Establish baseline negativity
           | (useful for crit). 'Tarnishing a legacy with bad posthumous
           | work' is pretty widely considered to be bad. - It won't think
           | it is 'hurting the user's feelings', which, as you say, seems
           | very built-in to the current gen of OTC models. - Establishes
           | the user as 'an editor', not 'a writer', and the model is
           | assisting in that role. Big difference.
           | 
           | Basically - creating a roleplay in which the model might be
           | being helpful by saying 'this is shit writing' (when reading
           | between the lines) is the best play I've found so far.
           | 
           | Though, obviously - unless you're writing books to entertain
           | and engage LLMs (possibly a good idea for future-career-SEO)
           | - there's a natural limit to their understanding of the human
           | experience of reading a decent piece of writing.
           | 
           | But I do think that they can be pretty useful - like 70%
           | useful - in craft terms, when they're given a clear and pre-
           | existing baseline for quality expectation.
        
       | saberience wrote:
       | Why would we ever want to build something like this, unless your
       | goal is to have fiction writers make even less money than they
       | already do.
       | 
       | Just stop, please. Try and automate some horrible and repetitive
       | drudgery.
       | 
       | Do you want to live in a world where humans no longer do any
       | creative work? It's grotesque.
        
         | frozenseven wrote:
         | I want to live in a world with more options and freedom to
         | choose. If somebody wants to build and explore with these AI
         | systems, you can't stop them.
        
           | cdblades wrote:
           | > I want to live in a world with more options and freedom to
           | choose
           | 
           | And you think that AI will be beneficial to that want?
        
             | frozenseven wrote:
             | Yes, undoubtedly.
        
         | soulofmischief wrote:
         | False dichotomy, overextended outrage and straw man fallacies
         | all rolled into one. Impressive.
        
       | itishappy wrote:
       | Those examples seem quite unrelated to one another. The first
       | reads as admitting intentional fraud and deceit, the second reads
       | like dealing with imposter syndrome. I'd love to know the prompt.
       | 
       | Also, not sure how you can judge a style to be clearly better
       | than another. The workflow of generating a bunch of stories in
       | the style of different authors and then voting on a favorite just
       | seems like picking a favorite author. Will the system ever prefer
       | short, hard-hitting sentences? Sure enough, convergence is a
       | noted behavior.
        
         | notahacker wrote:
         | > Those examples seem quite unrelated to one another. The first
         | reads as admitting intentional fraud and deceit, the second
         | reads like dealing with imposter syndrome. I'd love to know the
         | prompt.
         | 
         | Yeah. And to read the rest of each of the stories it
         | generated...
         | 
         | Both paragraphs are simply short excerpts which involve no
         | actual narrative, never mind the stuff that LLMs are typically
         | weak at (maintaining consistency, intricate plotting and
         | pacing, subtlety in world and character building) which in the
         | context of stories are far more important to improve than its
         | phrasing.
         | 
         | The fact that the "improvement" apparently eliminates a flaw in
         | the first passage ("gentle vibrations that vibrated through my
         | very being" is pretty clunky description unlikely to be written
         | by a native human; both paragraphs are otherwise passable and
         | equally mediocre writing) by implying apparently completely
         | different (and frankly less interesting) character motivations
         | makes me doubt that it's actually iteratively improving stories
         | rather than just spitting out significant rewrites which
         | incidentally eliminate glaring prose issues.
        
           | tamassimond wrote:
           | Yeah as we mention in the blog it's really hard to eval on
           | short passages. If you go on the Github can see longer
           | stories where the change is more noticeable. Both those
           | stories are from the same prompt
        
         | riskable wrote:
         | > Also, not sure how you can judge a style to be clearly better
         | than another.
         | 
         | This one is actually easy: The writing style used for a horror
         | is different than what you'd use for a romance novel. Example:
         | If you give it a prompt that asks the AI to generate something
         | in the style of a romance author but the rest of the prompt is
         | describing a horror or sci-fi story you'll end up with
         | something that most people would objectively decide, "ain't
         | right."
        
       | mynti wrote:
       | seems like this is just reward hacking the llm as a judge. this
       | does not give you a story humans will be more likely to read imho
        
       | mock-possum wrote:
       | This feels like a lot of fluff, without some solid examples of
       | the results - the one example of generated prose that they do
       | provide is pretty unimpressive, it reads like... well, like an
       | LLM wrote it.
        
         | SamBam wrote:
         | Indeed. Also the whole thing is just "apply an Evolutionary
         | Algorithm to stories." The only interesting question is whether
         | the LLM that decides the story's fitness rating (which they
         | call Elo, despite seeming to have nothing to do with the actual
         | Elo ranking system) can mimic a human's rating. Given the brief
         | example, it's not clear that it can, since it seems no better.
        
           | tamassimond wrote:
           | To clarify we use an Elo ranking system to update models
           | scores, so if you loose to a higher rated story you don't
           | loose as much Elo ranking. Definitely agree with LLM judge
           | criticism though it's still an open questions of how we can
           | make them better. Using the repeated story comparison judging
           | system does help make them more consistent. A good rubric
           | helps make them more human like as-well. The really big
           | question is how large is the generator verifier gap between
           | creating stories and marking them
        
       | gabriel666smith wrote:
       | I'm struggling a bit to understand the difference between the
       | reported results in the blog post and the examples in the Github.
       | 
       | The blog states:
       | 
       | > "Alpha Writing demonstrates substantial improvements in story
       | quality when evaluated through pairwise human preferences.
       | Testing with Llama 3.1 8B revealed:
       | 
       | 72% preference rate over initial story generations (95 % CI 63 %
       | - 79 %) 62% preference rate over sequential-prompting baseline
       | (95 % CI 53 % - 70 %) These results indicate that the
       | evolutionary approach significantly outperforms both single-shot
       | generation and traditional inference-time scaling methods for
       | creative writing tasks."
       | 
       | But in all of the examples using Llama 3.1 8B on the Github that
       | I could find, the stories with the top 5 highest final 'ELO' are
       | all marked elsewhere as:
       | 
       | "generation_attempt": null
       | 
       | Where the 'variant' stories, which I take to be 'evolved'
       | stories, are marked:
       | 
       | "generation_type": "variant", "parent_story_id":
       | "897ccd25-4776-4077-a9e6-0da34abb32a4"
       | 
       | IE - none of the 'winning stories' have a parent story; they seem
       | to have explicitly been the model's initial attempt. The examples
       | seem to prove the opposite of the statement in the blog post.
       | 
       | Perhaps 'variants' are slightly outperforming initial stories on
       | average (I don't have time to actually analyse the output data in
       | the repo), though it seems unlikely based on how I've read it (I
       | could be wrong!) and this might be borne out with far more
       | iterations.
       | 
       | However, a really important part of creative writing as a task is
       | that you (unfortunately) only get to tell a story once. The
       | losing variants won't ultimately matter. So, if I've read it
       | correctly, and all the winning stories are 'not evolved' - from
       | the initial prompt - this is quite problematically different from
       | the blog's claim that:
       | 
       | > "we demonstrate that creative output quality can be
       | systematically improved through increased compute allocation"
       | 
       | Super interesting work - I'd love to be told that I'm reading
       | this wrong! I was digging through in such detail to actually
       | compare differently-performing stories line-for-line (which would
       | also be nice to see - in the blog post, perhaps).
        
         | tamassimond wrote:
         | Just to clarify slight misunderstanding the variants without
         | parent ID aren't from the initial batch it just didn't carry
         | over to the next batch. You can see
         | "897ccd25-4776-4077-a9e6-0da34abb32a4" emerges from batch 5.
         | Apologies probably should make this clearer. Appreciate
         | feedback on blog post!
        
           | gabriel666smith wrote:
           | Ah got you! Makes sense, and makes it so much more clear.
           | Thanks. In that case, I totally retract my crit in my prev
           | comment. Appreciate it.
           | 
           | So - just so I completely understand - the variant we're
           | calling 897ccd25-4776-4077-a9e6-0da34abb32a4 _emerged_ during
           | batch 5, and doesn 't have a parent in a prior batch? Very
           | interesting to compare iterations.
           | 
           | I currently run some very similar scripts for one of my own
           | workflows. Though I understand making LLMs do 'good creative
           | writing' wasn't necessarily the point here - perhaps solely
           | to prove that LLMs can improve their own work, according to
           | their own (prompted) metric(s) - the blog post is correct to
           | point out that there's a huge limitation around prompt
           | sensitivity (not to mention subjectivity around quality of
           | art).
           | 
           | As a human using LLMs to create work that suits my own
           | (naturally subjective) tastes and preferences, I currently
           | get around this issue by feeding back on variants manually,
           | then having the LLM update its own prompt (much like a
           | cursorrules file, but for prose) based on my feedback, and
           | only then generating new variants, to be fed back on, etc.
           | 
           | It's extremely hard to one-shot-prompt everything you do or
           | do not like in writing, but you can get a really beefy
           | ruleset - which even tiny LLMs are very good at following -
           | incredibly quickly by asking the LLM to iterate on its own
           | instructions in this manner.
           | 
           | Like I said, not sure if your goal is to 'prove improvement
           | is possible' or 'create a useful creative writing assistant'
           | but, if it's the latter, that's the technique that has
           | created the most value for me personally over the last couple
           | of years. Sharing in case that's useful.
           | 
           | Grats on the cool project!
        
       | seaourfreed wrote:
       | This open source project will grow a ton, if it had a license on
       | it. Can you add an MIT license?
        
         | tamassimond wrote:
         | Thanks for feedback added MIT license
        
       | dgeiser13 wrote:
       | No. There's no thinking behind AI. If anything it throws shit
       | against the wall repeatedly until a human steps in and says "that
       | seems to be an improvement".
        
       | SkiFire13 wrote:
       | This completely misses the point of reinforced learning. The
       | reward condition needs to be representative of what you want
       | (e.g. in chess that would be winning).
       | 
       | Using a LLM as a judge means you will ultimately optimize for
       | stories that are liked by the LLM, not necessarily for stories
       | that are liked by people. For this to work the other LLM needs to
       | be as close to a human as possible, but this is what you were
       | trying to do in the first place!
        
       | proof_by_vibes wrote:
       | As a playwright, I've certainly thought about AI impacting the
       | art. In fact, it was the very eloquence of chatgpt's output that
       | initiated all of this mania in the first place: not only was
       | chatgpt able to explain to me gauge theory with surprising
       | accuracy, it was able to do so using perfect Elizabethan english
       | --exactly as I had instructed it to.
       | 
       | There is a missing ingredient that LLMs lack, however. They lack
       | insight. Writing is made engaging by the promise of insight
       | teased in its setups, the depths that are dug through its
       | payoffs, and the revelations found in its conclusion. It requires
       | solving an abstract sudoku puzzle where each sentence builds on
       | something prior and, critically, advances an agenda toward an
       | emotional conclusion. This is the rhetoric inherent to all
       | storytelling, but just as in a good political speech or debate,
       | everything hinges on the quality of the central thesis--the key
       | insight that LLMs do not come equipped to provide on their own.
       | 
       | This is hard. Insight is hard. And an AI supporter would gladly
       | tell you "yes! this is where prompting becomes art!" And perhaps
       | there is merit to this, or at least there is merit insofar as Sam
       | Altman's dreams of AI producing novel insights remain
       | unfulfilled. This condition notwithstanding, what merit exactly
       | do these supporters have? Has prompting become an art the same
       | way that it has become engineering? It would seem AlphaWrite
       | would like to say so.
       | 
       | But let's look at this rubric and evaluate for ourselves what
       | else AlphaWrite would like to say:
       | 
       | ```python # Fallback to a basic rubric if file not found return
       | """Creative writing evaluation should consider: 1. Creativity and
       | Originality (25%) - Unique ideas, fresh perspectives, innovative
       | storytelling 2. Writing Quality (25%) - Grammar, style, flow,
       | vocabulary, sentence structure 3. Engagement (20%) - How
       | compelling and interesting the piece is to read 4. Character
       | Development (15%) - Believable, well-developed characters with
       | clear motivations 5. Plot Structure (15%) - Logical progression,
       | pacing, resolution of conflicts""" ```
       | 
       | It's certainly just a default, and I mean no bad faith in using
       | this for rhetorical effect, but this default also acts as a
       | template, and it happens to be informative to my point. Insight,
       | genuine insight, is hard because it is contingent on one's
       | audience and one's shared experiences with them. It isn't enough
       | to check boxes. Might I ask what makes for a better story: a
       | narrative about a well developed princess who provides fresh
       | perspectives on antiquated themes, or a narrative about a well
       | developed stock broker who provides fresh perspectives on
       | contemporary themes? The output fails to find its audience no
       | matter what your rubric is.
       | 
       | And here lies the dilemma regarding the idea that prompts are an
       | art: they are not. The prompts are not art by the simple fact
       | that nobody will read them. What is read is what all that is
       | communicated and any discerning audience will be alienated by
       | anything generated by something as ambiguous as a English
       | teacher's grading rubric.
       | 
       | I write because I want to communicate my insights to an audience
       | who I believe would be influenced by them. I may be early in my
       | career, but this is why I do it. The degree of influence I shall
       | have measures the degree of "art" I shall attain. Not by whether
       | or not I clear the minimum bar of literacy.
        
       ___________________________________________________________________
       (page generated 2025-06-11 23:01 UTC)