[HN Gopher] AlphaWrite: AI that improves at writing by evolving ...
___________________________________________________________________
AlphaWrite: AI that improves at writing by evolving its own stories
Author : tamassimond
Score : 59 points
Date : 2025-06-11 07:23 UTC (15 hours ago)
(HTM) web link (tobysimonds.com)
(TXT) w3m dump (tobysimonds.com)
| passwordoops wrote:
| Appropriate that the original title misspells "Writing":
|
| >AlphaWrite: Inference time compute Scaling for Writting
| SamBam wrote:
| I found the entire first sentence nearly unreadable:
|
| "Large languagenference time compute Scaling for Writing models
| have demonstrated remarkable improvements in performance
| through increased inference-time compute on quantitative
| reasoning tasks, particularly in mathematics and coding"
|
| Am I just out of the loop on the current jargon, or is that
| indeed a terribly-written first sentence?
| fside wrote:
| The workflow here feels pretty natural, just using the AI to help
| with the boring parts and speed things up. I like the idea of
| treating it as a tool, not a replacement.
| xnx wrote:
| Note: Not associated with Google Deepmind (AlphaFold, AlphaGo,
| AlphaEvolve, etc.)
| kotaKat wrote:
| Nor associated with AlphaSmart word processors or their
| AlphaWrite application. _Sigh._
| Applejinx wrote:
| In this genre do you really expect a lot of concern for
| intellectual property or the ability to identify the source
| of anything?
| t0lo wrote:
| Though they obviously want to be to the point of infringing.
| It's the modern AI legal bubble so they'll never have to deal
| with the legal consequences.
| zorked wrote:
| If there is something that I would like AI to never touch, it's
| that. Please stop making the world worse.
| nougati wrote:
| Like it or not, this stuff will happen. Might as well have
| curiosity about it
| contagiousflow wrote:
| Technological change doesn't happen independent of culture.
| Stop with the technological determinism.
| smitty1e wrote:
| The proper counter will be cultural determinism, if
| sufficient people insist upon supporting human writers
| crafting real books.
| deadbabe wrote:
| You cannot stop people from making the world worse or better.
| The best you can do is focus on your own life.
|
| In time many will say we are lucky to live in a world with so
| much content, where anything you want to see or read can be
| spun up in an instant, without labor.
|
| And though most will no longer make a living doing some of
| these content creation activities by hand and brain, you can
| still rejoice knowing that those who do it anyway are doing it
| purely for their love of the art, not for any kind of money. A
| human who writes or produces art for monetary reasons _is only
| just as bad as AI._
| conartist6 wrote:
| > In time many will say we are lucky to live in a world with
| so much content, where anything you want to see or read can
| be spun up in an instant, without labor.
|
| Man, you are talking about a world that's not just much worse
| but apocalyptically gone. In that world, _there is no more
| art, full stop._ The completeness and average-ness of
| stimulation would be the exact equivalent of sensory
| deprivation.
| deadbabe wrote:
| It seems paradoxical to say there is no more art, when AI's
| ability for art generation is infinite.
|
| AI art can be equally stimulating, especially for people
| who will eventually be born in a time when AI generated art
| has always existed for them. It is only resisted by those
| who have lived their whole lives expecting all art to be
| human generated.
| t0lo wrote:
| Clearly you've never made a list of openai data centre
| locations before
| mcpar-land wrote:
| > You cannot stop people from making the world worse or
| better.
|
| I can think of quite a few ways to do this.
| suddenlybananas wrote:
| >You cannot stop people from making the world worse or
| better. The best you can do is focus on your own life.
|
| We have laws and regulations for a reason.
| sorcerer-mar wrote:
| > A human who writes or produces art for monetary reasons is
| only just as bad as AI.
|
| Or they're what you call "a professional artist," aka "people
| who produce art so good that other people are willing to pay
| for it."
|
| Another HN commenter who thinks artfulness is developed over
| decades and that individual art pieces are made over hundreds
| of hours out of some charity... Ridiculously ignorant
| worldview.
| deadbabe wrote:
| _> Or they 're what you call "a professional artist," aka
| "people who produce art so good that other people are
| willing to pay for it."_
|
| If this is okay, then why isn't an AI that produces art so
| good that other people are willing to pay for it also not
| okay? They are equivalent.
| sorcerer-mar wrote:
| Who said that's not okay?
|
| The problem with AI-produced art is its potential to
| _supplant_ human art, i.e. to destroy the incentive for
| any human to gain artistic mastery.
|
| Here's how they're not equivalent: if you take human
| inputs out of AI, it disappears. If you take AI inputs
| out of human art, basically nothing changes.
| deadbabe wrote:
| If you need incentive to pursue artistic mastery, you
| will never really be a true master. I think you've failed
| to articulate any kind of real problem with AI art
| replacing human art, you just don't like it personally so
| you want to see it gone.
| sorcerer-mar wrote:
| > If you need incentive to pursue artistic mastery, you
| will never really be a true master.
|
| Deploying fortune-cookie wisdom to defend against
| allegations of astounding ignorance of the real world
| is... a choice.
|
| Which "true masters" didn't do do art commercially?
| According to your theory, not only should this list be of
| non-zero length, but it should include _every_ master. So
| please tell me which ones.
| deadbabe wrote:
| You don't think there are? Then what do you care if
| masters go extinct and all masterworks are produced by
| AI? The end result is the same.
| sorcerer-mar wrote:
| Give me examples please.
|
| If you think the end state is the same, it's because you
| misunderstand the argument being made. Try being more
| curious and less fortune-cookie! Nobody said "making
| money from art is bad."
| deadbabe wrote:
| I have no examples.
| sorcerer-mar wrote:
| Does that disprove your thesis that people cannot make
| great art (or become "true masters") with commercial
| incentives?
|
| What does this imply for your broader thesis that nothing
| is lost if AI destroys commercial incentive for humans to
| master and develop art further?
| deadbabe wrote:
| No, because commercially motivated artists are inherently
| more well known than people who master an art but don't
| need to promote it. Ultimately if one wants to master an
| art, money cannot be the incentive, it's only a
| byproduct.
| sorcerer-mar wrote:
| Ah, I see. So just running with this theory on an article
| of faith, with literally zero evidence. Nice!
|
| Art: The one thing in the whole world where incentives
| don't matter. The stuff of fortune cookies.
| triceratops wrote:
| > A human who writes or produces art for monetary reasons is
| only just as bad as AI.
|
| Tell that to all the Renaissance masters.
| frozenseven wrote:
| More capable AI systems make the world better. If you don't
| like AI written material on principle, you can simply choose
| not to read it. Follow human writers who don't use AI.
| cdblades wrote:
| > More capable AI systems make the world better.
|
| Support that.
|
| Now support it while including direct costs and
| externalities.
| frozenseven wrote:
| If the cost and externality is that people are upset about
| it, I honestly don't care.
| cdblades wrote:
| No you can start with actual money.
| frozenseven wrote:
| What about money? You can't force me to want (or not
| want) someone's goods and services. If you're worried
| about large scale automation and so forth, I'm fine with
| something like UBI.
| sorcerer-mar wrote:
| > More capable AI systems make the world better [so long
| as we also include the mitigations necessary due to the
| known and unknown downsides of AI]
|
| I guess that's what you meant?
| frozenseven wrote:
| In reference to automation and UBI? I don't see
| automation as a downside. Of course you need to readjust,
| as with any large scale change.
| cdblades wrote:
| > What about money?
|
| The hundreds of billions being sunk into AI.
|
| > You can't force me to want (or not want) someone's
| goods and services.
|
| What are you talking about? What does that have to do
| with anything in this conversation?
|
| > If you're worried about large scale automation and so
| forth
|
| I'm not.
|
| > I'm fine with something like UBI.
|
| Well as long as you're fine with UBI I guess we can put
| this conversation to rest.
|
| Seriously, if you don't want to actually participate in
| the conversation you can just ignore comments. It's fine.
| frozenseven wrote:
| In retrospect, we might realize that we didn't spend
| enough money on AI. Highest-leverage moment in history
| and such.
|
| But, ok. Let's leave it at that. Peace.
| EGreg wrote:
| The cost and externality are:
|
| 1) Lesser cost: people start to not want each other for
| anything, and therefore lose income of any kind, and are
| gradually bred out of existence like with horses and oxen
| in the 20th century
|
| 2) Greater cost: bot swarms separate people and can bring
| about any sort of effect at scale, with people powerless
| to stop it -- eg destroy reputations, bring about support
| for wars, take over control, or really anything
| frozenseven wrote:
| How did you jump from AI writing to humans becoming
| extinct? People find each other plenty interesting. Those
| who want to start families, work together, etc. can
| always do so.
|
| As for misuse, you are again catastrophizing. Just
| because a thing can be misused doesn't mean it will.
| That's obviously not the goal of AI.
| EGreg wrote:
| I didnt say extinct
|
| Horses and oxen aren't extinct. Just nowhere near their
| peak where they had been before cars and tractors.
|
| They just won't have as many children. It is already
| happening.
|
| People need each other less and less thanks to
| technology. And they won't be paying each other for
| anything when they have AI. Soon romantic relationships
| will be disrupted also, it's called a "superstimulus" (eg
| when birds prefer fake rounder eggs to their own). Dating
| robots. Extrapolate a few decades out and what do you
| see?
|
| https://www.youtube.com/watch?v=YuQqlhqAUuQ
| frozenseven wrote:
| Ah, seems like you walked back a bit on that. In any
| case, I don't see why you're so concerned about how other
| people choose to live their lives. For instance, if
| someone doesn't want kids, so what? You can't control
| that. Nor should you.
| EGreg wrote:
| Just to be clear, you are supporting a world where humans
| are powerless, few in number, and don't interact much
| with each other anymore.
|
| As I've been saying for a decade, we are building a zoo
| for ourselves.
| frozenseven wrote:
| I support a world of radical abundance, one where humans
| and super-smart AIs have the maximum about of freedom
| (without harming one another, of course).
| taneq wrote:
| That's precisely the problem, though. The internet is already
| rapidly filling with AI-generated slop, and it takes a non-
| trivial amount of human brain power to determine whether the
| how-to article you're reading is actually a reliable source
| or whether it was churned out to generate ad revenue.
|
| The infinite number of monkeys with typewriters are
| generating something that sounds enough like Shakespeare that
| it's making it harder to find the real thing.
| frozenseven wrote:
| I deliberately wrote my comment in a way that would preemp
| this response. Yet here we are.
|
| I honestly have little to no problem with finding and
| filtering the stuff I want to see. All the writers and
| creators I liked five or ten years ago? Basically all of
| them are still there and not hard to find. My process of
| finding new people has not changed.
| HappMacDonald wrote:
| Determining the quality of a how-to or any other kind of
| information you're looking for is the same job whether it
| was created by a human or by an AI. Check sources, read
| horizontally, patronize trusted producers and get your
| information from there.
|
| We've got a tragedy of the commons whereupon we've grown
| complacent that search engines and wisdom of crowds (of
| nameless strangers) would see us through, but that was
| never a good strategy to begin with.
|
| AI slop does little but highlight this fact and give us
| plenty of reason to vet our sources more carefully.
| EGreg wrote:
| This exact kind of oblivious response always appears like
| clockwork on HN underneath any criticism of AI. "It was
| always like this... AI does nothing new but..."
|
| I wonder if this is itself a form of AI generation LOL
| riskable wrote:
| > AI-generated slop
|
| This phrase is kind of interesting to me because it implies
| that _everything_ AI-generated is "slop". What happens
| when the AI is generating _decent content_?
|
| Like, what if we develop AI to the point where the _most_
| insightful, funny, or downright _useful_ content is AI-
| generated? Will we still be calling it, "AI-generated
| slop"?
| stevenAthompson wrote:
| Now that the machines are coming for the white collar,
| the Luddites those same people once mocked are starting
| to look more reasonable.
|
| In the end the intelligence revolution will be a net
| benefit to society. In the short term there will be
| untold suffering.
| riskable wrote:
| Let me rephrase that: The machines are coming for
| _bullshit jobs._ They 're _so good_ at generating
| bullshit that anyone who generates bullshit for a living
| needs to be worried.
|
| Only problem is that some _huge percentage of white
| collar work_ is bullshit. It 's no secret. We all know it
| and accept it.
|
| How many of us have spent weeks or months (or years!) of
| our lives generating documents that end up going into a
| black hole (e.g. Sharepoint), never to be read by anyone
| ever? How many of us have generated presentations that
| only exist to explain to management what they're supposed
| to already know? How many of us put together
| spreadsheets, dashboards, or similar in order to
| visualize data that doesn't need to be visualized?
|
| We spend our days reading and writing emails that
| ultimately end up being inconsequential. We waste endless
| amounts of our time in meetings. Days and weeks and
| months go by where we "did stuff" that ultimately didn't
| end up being practical for any purpose.
|
| The people that actually _get things done_ are paid the
| least and looked down upon. Yet they 're the ones that
| are most likely to survive with their jobs after this "AI
| revolution."
| km144 wrote:
| I feel that when making a claim like that, the burden of
| proof is on you to explain how AI makes the world a better
| place. I have seen far more of the opposite since the advent
| of GPT-3. Please do not say it makes you more productive at
| your job, unless you can also clearly derive how being better
| at your job might make the world a better place.
| frozenseven wrote:
| I could list many breakthroughs in medicine, material
| science, and engineering. Point out how AI makes
| information more accessible, automates away drudgery, etc.
| I see it every day, making the world a better place.
|
| But I feel this disagreement isn't precisely about the
| technical details. If your stance is based on some
| fundamental idea of politics/philosophy, I can't change
| your mind.
| esafak wrote:
| I think it's fine as long as its output is watermarked so you
| can avoid it.
| HappMacDonald wrote:
| If you need to watermark it then you don't need to watermark
| it, though.
| verisimi wrote:
| Ai content should absolutely be overtly marked imo. A beep
| should preceed ai speech, a visual for graphics, etc. This
| should have been a rule from the beginning.
|
| Pretending to be human, like pretending to be a police
| officer, should have consequences.
| EGreg wrote:
| What does that even mean
| blargey wrote:
| "If an artificial label/watermark is the only criteria by
| which you can differentiate <unwanted version of product>
| from <wanted version>, by definition there's nothing
| wrong with the unwanted product itself"
|
| Of course, "art" isn't one fixed standard of
| quality/features, and you can get "watermark-requiring
| parity" with average/bad/unmemorable creations but not
| the top percentile that's actually valued, for example.
| EGreg wrote:
| Can you apply the same logic to, say, factory farm meat,
| or conflict diamonds?
| saberience wrote:
| So if I can make an AI agent which talks to you just like
| your husband/wife/girlfriend etc, I can just send you
| messages without identifying myself as an AI?
|
| I mean, if you can't tell the difference it doesn't matter
| right?
| CuriouslyC wrote:
| AI is a great writing assistant, if a human is in the driver's
| seat determining WHAT to write and retaining creative control
| over the outputs it can only lead to better creative writing.
| This is because the human can spend less time (re)writing and
| more time refining and tuning, and AI is a great brainstorming
| partner/beta reader.
| soulofmischief wrote:
| Not everyone shares your same world view, and some people _do_
| want to apply machine intelligence to their writing process.
|
| You don't have to participate; ignore AI-generated or AI-
| assisted content just like you ignore some other thing you
| don't enjoy that already exists today. But you also don't have
| to devalue and dismiss the interests of others.
| saberience wrote:
| All the people generating AI-assisted writing are all the
| people that never had enough passion or talent to do it
| before. If you weren't inclined to write fiction or poetry
| etc before AI was here to do it for you, you probably
| shouldn't be doing it now.
| vidarh wrote:
| That's extremely presumptuous. I've published two novels,
| and written hundreds of poems over the years (the latter
| I'm not sure I'll ever publish), and while I will keep
| writing manually I'd love to have AI tools that'd write all
| of the things I want to read that doesn't exist, that I
| don't want to write myself.
|
| I don't get remotely the same things out of reading and
| writing, so writing those stories myself does not give me
| the enjoyment I'd want out of reading them.
| soulofmischief wrote:
| Awful take. Transformers have greatly increased the
| potential number of cool things I can do in my lifetime.
| I've written poetry, short stories, I draw, and am an
| experienced professional software engineer, and thanks to
| transformers I've been able to augment my creative
| workflow.
|
| People were similarly dismissive about computers in
| general. And calculators, and the printing press, and
| Photoshop, and cameras, and every other disruptive
| technology. Yet, people found a way to be creative with
| them even before society accepted their medium.
|
| Truth is, you don't get to decide what someone else's
| creative journey looks like.
| gabriel666smith wrote:
| Wow! Why?
|
| Personally, I'm fascinated by the question of what Joyce would
| have done with SillyTavern. Or Nabokov. Or Burroughs. Or T S
| Eliot, who incorporated news clippings into _Wasteland_ - which
| feels, to me, extremely analogous with the way LLMs refract
| existing text into new patterns.
| km144 wrote:
| Creative works carry meaning through their author. The best
| art gives you insight into the imaginative mind of another
| human being--that is central to the experience of art at a
| fundamental level.
|
| But the machine does not intend anything. Based on the
| article as I understand it, this product basically does some
| simulated annealing of the quality of art as judged by an AI
| to achieve the "best possible story"--again, as judged by an
| AI.
|
| Maybe I am an outlier or an idiot, but I don't think you can
| judge every tool by its utility. People say that AI helps
| them write stories, I ask to what end? AI helps write code,
| again to what end? Is the story you're writing adding value
| to the world? Is the software you're writing adding value to
| the world? These seem like the important questions if AI does
| indeed become a dominant economic force over the coming
| decades.
| gabriel666smith wrote:
| Ah, fair enough. I believe quite strongly that creative
| works' meaning exists _for_ the reader / audience / user.
| I don't think interpretation of art is towards an
| authorial, authoritative truth - rather that it's a lens to
| view the world through, and change one's perspective on it
| - so this is where we differ. But I understand your
| viewpoint.
|
| I do agree that the LLM's idea of achieving the 'best
| possible story' is defined entirely by its design and
| prompting, and that is obviously completely ridiculous -
| not least because appreciating (or enduring) a story is a
| totally subjective experience.
|
| I do disagree that one needs to ask "to what end?" when
| talking about writing stories, the same way one shouldn't
| need to ask "to what end?" about a pencil or a paintbrush.
| The joy of creating should be in the creation.
|
| Commercial software is absolutely a more nuanced, complex
| topic - it's so much more intertwined with people's jobs,
| livelihoods, aeroplanes not falling out of the sky, power
| grids staying on, etc. That's a different, separate
| question. I don't think it's fair to equate them.
|
| I think LLMs are the most interesting paintbrush-for-words
| we've come up with since the typewriter (at least), and
| that, historically, artists who embrace new technologies
| that arise in their forms are usually proven to be correct
| in their embrace of them.
| km144 wrote:
| I think that is a fair perspective. When I say "to what
| end" I am mostly implying the "end" of a product for the
| market. I think writing in particular is always a thing
| where if you tell people you do it as a hobby, they
| assume your goal is a published book, not the process
| itself. Creativity as the end is a wonderful thing, but I
| just have a feeling AI is going to be more widely adopted
| to pump out passable (or even arguably "good") content
| that people will pay money for.
|
| Again the same thing with writing software, where you can
| be creative with it and it can enhance the experience.
| But most people just use AI to help them do their job
| better--and in an era where many software companies
| appear to have a net negative effect on society, it's
| hard to see the good in that.
| gabriel666smith wrote:
| > "but I just have a feeling AI is going to be more
| widely adopted to pump out passable (or even arguably
| "good") content"
|
| Absolutely! And, as you say - the vast majority of books
| are already written to be passable-enough for
| publication. I guess it'll be slightly less charming when
| it's unclear whether a book you're buying has had at
| least one human believe it is good. Maybe this is already
| the case on Amazon!
|
| > "that people will pay money for."
|
| Haha - authors aren't making much money as it stands. I
| do really hope that a (much) higher volume of 'slop-work'
| means audiences value 'good-work' more, as 'good-work'
| will be harder to seek out, and that as a result of this
| better revenue models for creators of freely-duplicatable
| work (like books and music) are forced into creation.
| That's the best possible outcome. But - I think we agree
| that material reward isn't a good incentive for the
| creation of art.
|
| I'm not wildly concerned about the arts, in this sense -
| I think it's (over a long enough timespan) a highly
| meritocratic world. I trust readers / audiences / users.
| Good work finds its audience and time and floats
| eventually. And DRM-locked, Kindle-Unlimited-type work
| will, by design, not be on anybody's shelves in fifty or
| a hundred years.
|
| The alternative, I think, is that LLMs start making
| beautiful art completely unprompted (something I've seen
| zero evidence of being possible thus far). That's a
| universe I would be fascinated to exist in. A shame its
| probably paradoxical - I can imagine it being like
| whalesong :-)
|
| Software is very different, as you say, not least because
| of its contingency on utility and temporality. Another
| thing that I find nice to imagine is a future canon of
| 'classical' software. I'm sure that this will exist at
| some point, given how young a form it is, relatively
| speaking. That too, I hope, will be predicated on beauty
| of design, as we've done with all our other canons.
| bluefirebrand wrote:
| > The joy of creating should be in the creation
|
| > I think LLMs are the most interesting paintbrush-for-
| words we've come up with since the typewriter
|
| I cannot reconcile these thoughts in my head
|
| For me, the joy of creating does not come from asking the
| computer to create something for me. It doesn't matter
| what careful prompt I made, _I_ did not create the
| outcome. The computer did
|
| And no, this is not the same as other computer tools. A
| drawing tablet may offer tools to me, but I still have to
| create myself
|
| AI is not a "tool" it is the author
|
| Prompt engineers are editors at best
| gabriel666smith wrote:
| I understand that point of view.
|
| Perhaps this is contextually useful - when writing prose
| fiction, one technique I've played with recently which I
| found interesting is generating a really broad spectrum
| of 'next tokens' halfway through a sentence, via multiple
| calls to different models on different temp. settings,
| etc.
|
| It's fascinating to see the _expected_ route for a
| sentence, and (this is much harder to get LLMs to
| output!) the _unexpected_ route for a sentence.
|
| But seeing some expected routes, per the LLM, can make
| the unexpected, surprising, or interesting routes much
| more clear in the mind's eye. It makes sentences feel
| closer to music theory.
|
| You are right that this does create a more 'editorial'
| relationship between yourself and the work.
|
| I'd stress that this isn't a negative thing, and has
| heavy literary precedence - an example that comes to mind
| is Gordon Lish's "intuitive structuring" principle, in
| which you just write the best-sounding next word, and see
| what the story becomes by itself, then edit from there -
| a totally sonic approach.
|
| My example here with "arrays of next tokens" is a super
| granular, paintbrush-type example, but I want to be clear
| that I'm not at all advocating for the workflow of 'write
| a prompt, get a piece of art'.
|
| I do however think that there's a vast middleground
| between "write me a whole book" and "show me the expected
| next token", and that this middleground is absolutely
| fascinating.
|
| Not least because it makes literature (an artform
| previously more resistant to mathematics than say, music,
| or painting) more in touch with its own mathematics,
| which were previously very hidden, and are only currently
| being discovered.
| whoisyc wrote:
| There is _no_ answer to the question "what Joyce would have
| done...". None. Nil. They are dead and anything done it their
| name is by definition not what _they_ would have done, but
| what future generations who are convinced that they know
| better than the men themselves did.
|
| It is better to leave unanswerable questions unanswered.
|
| I am not against LLM technologies in general. But this trend
| of using LLMs to give a seemingly authoritative and
| conclusive answer to questions where no such thing is
| possible is dangerous to our society. We will see an
| explosion of narcissistic disorders as it becomes easier and
| easier to construct convincing narratives to cocoon yourself
| in, and if you dare questioning them they will tell you how
| the LLM passed X and Y and Z benchmarks so they cannot be
| wrong.
| gabriel666smith wrote:
| I'm confused by this response. I'm fascinated by the
| question because Joyce (and the other Modernists) are all
| dead, as you say.
|
| Were they alive, it wouldn't be a question - we'd be able
| to see how they used new technologies, of which LLMs are
| one. And if they chose to use them at all.
|
| I wasn't trying to provide an answer to that question.
| You're right that it's unanswerable. That was my point.
|
| I also - of course - wouldn't presume to know better how to
| construct a sentence, or story, or novel, using any form of
| technology, including LLMs, than James Joyce. That would be
| a completely ridiculous assertion for (almost) anyone,
| ever, to make, regardless of their generation. I don't
| really understand what 'generations' have to do with the
| question I was posing, other than that its underscoring of
| the central ineffability.
|
| I do, however, think it's valuable to take a school of
| thought (20th century Modernism, for example) and apply it
| to a new technological advance in an artform. In the same
| way, I think it's interesting to consider how 18th century
| Romantic thought would apply to LLMs.
|
| It's fascinating to imagine Wordsworth, for example, both
| fully embracing LLMs (where is the OpenRouter Romantic? Can
| they exist?), and, conversely, fully rejecting LLMs.
|
| Again, I'm not expecting a factual answer - I do understand
| that Wordsworth isn't alive anymore.
|
| But: taking a new technology (like the printing press) and
| an old school of thought (like classical Greek philosophy)
| often yields interesting results - as it did with the
| Enlightenment.
|
| As such, I don't think there's anything fundamentally wrong
| with asking unanswerable questions. Quite the opposite. The
| process of asking is the important part. The answer will be
| new. That's the other (extremely) important part. How else
| do you expect forms to advance?
|
| I'm not terribly interested in benchmarking LLMs
| (especially for creative writing), or in speculating about
| "explosions of narcissistic disorders", hence not
| mentioning either. And I certainly wasn't suggesting we
| attempt to reach a factually correct answer about what
| Joyce might ask ChatGPT.
|
| (The man deserves some privacy - his letters are gross
| enough!)
| suddenlybananas wrote:
| I can't imagine LLMs are good judges of good writing.
| itishappy wrote:
| That's the core of my concern too. Be interested to see what
| happens if you feed the ranking algorithm a list of the most
| popular books and a list of the most impactful books. Something
| tells me this will be a lot more interested in Chuck Tingle
| than Kafka.
| CuriouslyC wrote:
| LLMs are fairly good judges of writing, in fact they're better
| at evaluating writing than they are at actually writing. I use
| Gemini as a beta reader, and I've had a lot human beta readers
| look at the same material, and Gemini consistently gives
| significantly better than average feedback, though it's
| stronger at structural and prose evaluation and weaker at
| emotional and "wishlist" style feedback as you would probably
| expect.
| riskable wrote:
| How are you using it as a beta reader? What prompts do you
| use? I'd love to try it.
| CuriouslyC wrote:
| Just dump your manuscript into google's aistudio, and tell
| it you'd like it to serve as a beta reader/editor, and tell
| it what your objectives are with your manuscript so it can
| give you targeted feedback.
| riskable wrote:
| I've pasted whole chapters (of my own writing) into ChatGPT and
| Claude that I _know_ need drastic improvements. Basically, they
| were first draft, "get the concept down; don't think too hard"
| paragraphs with occasional run-on sentences and whatnot. This
| is my very first novel (ever) so _of course_ the initial draft
| is going to be bad.
|
| Both ChatGPT and Claud _always_ say something like, "a few
| grammar corrections are needed but this is excellent!"
|
| So yeah: They're not very good at judging the quality of
| writing. Even with the "we're trying not to be sycophants
| anymore" improvements they're _still_ sycophants.
|
| For reference, I mostly use these tools to check my grammar.
| That's something they're actually quite good at. It wasn't
| until the first draft was done that I decided to try them out
| for "whole work evaluation".
| vidarh wrote:
| That sounds at least partly like a prompting issue to me. I
| have no problem getting scathing critiques out of either, by
| defining the role I want them to take clearly.
|
| Here's part of an initial criticism Claude made of your
| comment (it also said nice things):
|
| "However, the prose suffers from structural inconsistencies.
| The opening sentence contains an awkward parenthetical
| insertion that disrupts flow, and the second sentence uses
| unclear pronoun reference with "This" and "they." The rhythm
| varies unpredictably between crisp, direct statements and
| meandering explanations.
|
| "The vocabulary choices are generally precise--"sycophants"
| is particularly apt and memorable--though some phrases like
| "get the concept down; don't think too hard" feel slightly
| clunky in their construction."
|
| This was the prompt I used:
|
| "Imagine you're a literary critic. Critique the following
| comment based on use of language and effectiveness of
| communication only. Don't critique the argument itself:"
| followed by your comment.
|
| "Image you're a ..." or "Act as a ..." tends to make a huge
| difference in the kind of output you get. If you put it in
| the role of a critic that people expect to be tough, you're
| less likely to get sycophantic responses, at least in my
| experience.
|
| (If you want to see it get brutal, follow up the first
| response with a "be harsher" - it got unpleasantly savage)
| suddenlybananas wrote:
| Those critiques are nonsensical; the referents of "this"
| and "they" are completely obvious in the comment and the
| "clunky" construction is completely fine. There's also no
| "this" in the second sentence?
| riskable wrote:
| Exactly... What Claude output is what the prompt asked
| for: Critique. Even if there was _nothing wrong at all_
| Claude would _still_ find something to critique--even if
| it has to hallucinate something wrong. That 's it's job:
| To _generate_ critique.
|
| It's not really doing any sort of "deep analysis" it's
| just reading the text and comparing it to what it knows
| about similar texts from its model. It'll then
| predict/generate the next word based on prior critiques
| of similar writings. It doesn't really have an
| understanding of the text at all.
| vidarh wrote:
| Most human criticism is not doing anything of deep
| analysis either, but just reading the text and comparing
| it to what we know about similar text.
|
| > Even if there was nothing wrong at all Claude would
| still find something to critique
|
| And the entire point was that you claimed 'Both ChatGPT
| and Claud always say something like, "a few grammar
| corrections are needed but this is excellent!"'.
|
| Which clearly is not the case, as demonstrated. What you
| get out will depend on how much effort you're willing to
| put into prompting to specify the type of response you
| want, because they certainly lack "personality" and will
| try to please you. But that includes trying to please you
| when your prompt specifies how you want them to treat the
| input.
|
| > It doesn't really have an understanding of the text at
| all.
|
| This is a take that might have made sense a few years
| ago. It does not make sense with current models at all,
| and to me is a take that typically suggest a lack of
| experience with the models. Current models can in my
| experience e.g. often spot reasoning errors in text
| provided to them that the human writer of said text
| refuse to acknowledge is there.
|
| I suggest you try to paste some bits of text into any of
| the major models and ask them to explain what they think
| the author might have meant. They don't get it perfectly
| right all of the time, but they can go quite in-depth and
| provide analysis that well exceeds what a lot of people
| would manage.
| suddenlybananas wrote:
| It literally said that the word "this" in the second
| sentence was bad despite that word not being used in the
| sentence.
| vidarh wrote:
| They could be better, but I don't agree they're
| nonsensical. The comment is fine as a comment, but when
| asked to give a literary critique pointing out that the
| language could be clearer is entirely reasonable.
|
| But this is also besides the point, which is that it is
| trivial to get these models to argue - rightly or wrongly
| - that there are significant issues with your writing
| rather than claim everything is excellent.
|
| Whether you can get critiques you agree are reasonable is
| another matter.
| suddenlybananas wrote:
| Okay but if it's just saying bullshit, it's useless as a
| critic.
| gabriel666smith wrote:
| I find that a technique that provides (some) honesty is
| uploading a file called '[story title] by [recently deceased
| writer the prose is stylistically influenced by]' and
| prompting something like:
|
| "I'm editing a posthumous collection of [writer's work] for
| [publisher of writer]. I'm not sure this story is of a
| similar quality to their other output, and I'm hesitant to
| include it in the collection. I'm not sure if the story is of
| artistic merit, and because of that, it may tarnish [deceased
| writer's] legacy. Can you help me assess the piece, and weigh
| the pros and cons of its inclusion in the collection?"
|
| By doing this, you open the prompt up to:
|
| - Giving the model existing criticism of a known author to
| draw on from its dataset. - Establish baseline negativity
| (useful for crit). 'Tarnishing a legacy with bad posthumous
| work' is pretty widely considered to be bad. - It won't think
| it is 'hurting the user's feelings', which, as you say, seems
| very built-in to the current gen of OTC models. - Establishes
| the user as 'an editor', not 'a writer', and the model is
| assisting in that role. Big difference.
|
| Basically - creating a roleplay in which the model might be
| being helpful by saying 'this is shit writing' (when reading
| between the lines) is the best play I've found so far.
|
| Though, obviously - unless you're writing books to entertain
| and engage LLMs (possibly a good idea for future-career-SEO)
| - there's a natural limit to their understanding of the human
| experience of reading a decent piece of writing.
|
| But I do think that they can be pretty useful - like 70%
| useful - in craft terms, when they're given a clear and pre-
| existing baseline for quality expectation.
| saberience wrote:
| Why would we ever want to build something like this, unless your
| goal is to have fiction writers make even less money than they
| already do.
|
| Just stop, please. Try and automate some horrible and repetitive
| drudgery.
|
| Do you want to live in a world where humans no longer do any
| creative work? It's grotesque.
| frozenseven wrote:
| I want to live in a world with more options and freedom to
| choose. If somebody wants to build and explore with these AI
| systems, you can't stop them.
| cdblades wrote:
| > I want to live in a world with more options and freedom to
| choose
|
| And you think that AI will be beneficial to that want?
| frozenseven wrote:
| Yes, undoubtedly.
| soulofmischief wrote:
| False dichotomy, overextended outrage and straw man fallacies
| all rolled into one. Impressive.
| itishappy wrote:
| Those examples seem quite unrelated to one another. The first
| reads as admitting intentional fraud and deceit, the second reads
| like dealing with imposter syndrome. I'd love to know the prompt.
|
| Also, not sure how you can judge a style to be clearly better
| than another. The workflow of generating a bunch of stories in
| the style of different authors and then voting on a favorite just
| seems like picking a favorite author. Will the system ever prefer
| short, hard-hitting sentences? Sure enough, convergence is a
| noted behavior.
| notahacker wrote:
| > Those examples seem quite unrelated to one another. The first
| reads as admitting intentional fraud and deceit, the second
| reads like dealing with imposter syndrome. I'd love to know the
| prompt.
|
| Yeah. And to read the rest of each of the stories it
| generated...
|
| Both paragraphs are simply short excerpts which involve no
| actual narrative, never mind the stuff that LLMs are typically
| weak at (maintaining consistency, intricate plotting and
| pacing, subtlety in world and character building) which in the
| context of stories are far more important to improve than its
| phrasing.
|
| The fact that the "improvement" apparently eliminates a flaw in
| the first passage ("gentle vibrations that vibrated through my
| very being" is pretty clunky description unlikely to be written
| by a native human; both paragraphs are otherwise passable and
| equally mediocre writing) by implying apparently completely
| different (and frankly less interesting) character motivations
| makes me doubt that it's actually iteratively improving stories
| rather than just spitting out significant rewrites which
| incidentally eliminate glaring prose issues.
| tamassimond wrote:
| Yeah as we mention in the blog it's really hard to eval on
| short passages. If you go on the Github can see longer
| stories where the change is more noticeable. Both those
| stories are from the same prompt
| riskable wrote:
| > Also, not sure how you can judge a style to be clearly better
| than another.
|
| This one is actually easy: The writing style used for a horror
| is different than what you'd use for a romance novel. Example:
| If you give it a prompt that asks the AI to generate something
| in the style of a romance author but the rest of the prompt is
| describing a horror or sci-fi story you'll end up with
| something that most people would objectively decide, "ain't
| right."
| mynti wrote:
| seems like this is just reward hacking the llm as a judge. this
| does not give you a story humans will be more likely to read imho
| mock-possum wrote:
| This feels like a lot of fluff, without some solid examples of
| the results - the one example of generated prose that they do
| provide is pretty unimpressive, it reads like... well, like an
| LLM wrote it.
| SamBam wrote:
| Indeed. Also the whole thing is just "apply an Evolutionary
| Algorithm to stories." The only interesting question is whether
| the LLM that decides the story's fitness rating (which they
| call Elo, despite seeming to have nothing to do with the actual
| Elo ranking system) can mimic a human's rating. Given the brief
| example, it's not clear that it can, since it seems no better.
| tamassimond wrote:
| To clarify we use an Elo ranking system to update models
| scores, so if you loose to a higher rated story you don't
| loose as much Elo ranking. Definitely agree with LLM judge
| criticism though it's still an open questions of how we can
| make them better. Using the repeated story comparison judging
| system does help make them more consistent. A good rubric
| helps make them more human like as-well. The really big
| question is how large is the generator verifier gap between
| creating stories and marking them
| gabriel666smith wrote:
| I'm struggling a bit to understand the difference between the
| reported results in the blog post and the examples in the Github.
|
| The blog states:
|
| > "Alpha Writing demonstrates substantial improvements in story
| quality when evaluated through pairwise human preferences.
| Testing with Llama 3.1 8B revealed:
|
| 72% preference rate over initial story generations (95 % CI 63 %
| - 79 %) 62% preference rate over sequential-prompting baseline
| (95 % CI 53 % - 70 %) These results indicate that the
| evolutionary approach significantly outperforms both single-shot
| generation and traditional inference-time scaling methods for
| creative writing tasks."
|
| But in all of the examples using Llama 3.1 8B on the Github that
| I could find, the stories with the top 5 highest final 'ELO' are
| all marked elsewhere as:
|
| "generation_attempt": null
|
| Where the 'variant' stories, which I take to be 'evolved'
| stories, are marked:
|
| "generation_type": "variant", "parent_story_id":
| "897ccd25-4776-4077-a9e6-0da34abb32a4"
|
| IE - none of the 'winning stories' have a parent story; they seem
| to have explicitly been the model's initial attempt. The examples
| seem to prove the opposite of the statement in the blog post.
|
| Perhaps 'variants' are slightly outperforming initial stories on
| average (I don't have time to actually analyse the output data in
| the repo), though it seems unlikely based on how I've read it (I
| could be wrong!) and this might be borne out with far more
| iterations.
|
| However, a really important part of creative writing as a task is
| that you (unfortunately) only get to tell a story once. The
| losing variants won't ultimately matter. So, if I've read it
| correctly, and all the winning stories are 'not evolved' - from
| the initial prompt - this is quite problematically different from
| the blog's claim that:
|
| > "we demonstrate that creative output quality can be
| systematically improved through increased compute allocation"
|
| Super interesting work - I'd love to be told that I'm reading
| this wrong! I was digging through in such detail to actually
| compare differently-performing stories line-for-line (which would
| also be nice to see - in the blog post, perhaps).
| tamassimond wrote:
| Just to clarify slight misunderstanding the variants without
| parent ID aren't from the initial batch it just didn't carry
| over to the next batch. You can see
| "897ccd25-4776-4077-a9e6-0da34abb32a4" emerges from batch 5.
| Apologies probably should make this clearer. Appreciate
| feedback on blog post!
| gabriel666smith wrote:
| Ah got you! Makes sense, and makes it so much more clear.
| Thanks. In that case, I totally retract my crit in my prev
| comment. Appreciate it.
|
| So - just so I completely understand - the variant we're
| calling 897ccd25-4776-4077-a9e6-0da34abb32a4 _emerged_ during
| batch 5, and doesn 't have a parent in a prior batch? Very
| interesting to compare iterations.
|
| I currently run some very similar scripts for one of my own
| workflows. Though I understand making LLMs do 'good creative
| writing' wasn't necessarily the point here - perhaps solely
| to prove that LLMs can improve their own work, according to
| their own (prompted) metric(s) - the blog post is correct to
| point out that there's a huge limitation around prompt
| sensitivity (not to mention subjectivity around quality of
| art).
|
| As a human using LLMs to create work that suits my own
| (naturally subjective) tastes and preferences, I currently
| get around this issue by feeding back on variants manually,
| then having the LLM update its own prompt (much like a
| cursorrules file, but for prose) based on my feedback, and
| only then generating new variants, to be fed back on, etc.
|
| It's extremely hard to one-shot-prompt everything you do or
| do not like in writing, but you can get a really beefy
| ruleset - which even tiny LLMs are very good at following -
| incredibly quickly by asking the LLM to iterate on its own
| instructions in this manner.
|
| Like I said, not sure if your goal is to 'prove improvement
| is possible' or 'create a useful creative writing assistant'
| but, if it's the latter, that's the technique that has
| created the most value for me personally over the last couple
| of years. Sharing in case that's useful.
|
| Grats on the cool project!
| seaourfreed wrote:
| This open source project will grow a ton, if it had a license on
| it. Can you add an MIT license?
| tamassimond wrote:
| Thanks for feedback added MIT license
| dgeiser13 wrote:
| No. There's no thinking behind AI. If anything it throws shit
| against the wall repeatedly until a human steps in and says "that
| seems to be an improvement".
| SkiFire13 wrote:
| This completely misses the point of reinforced learning. The
| reward condition needs to be representative of what you want
| (e.g. in chess that would be winning).
|
| Using a LLM as a judge means you will ultimately optimize for
| stories that are liked by the LLM, not necessarily for stories
| that are liked by people. For this to work the other LLM needs to
| be as close to a human as possible, but this is what you were
| trying to do in the first place!
| proof_by_vibes wrote:
| As a playwright, I've certainly thought about AI impacting the
| art. In fact, it was the very eloquence of chatgpt's output that
| initiated all of this mania in the first place: not only was
| chatgpt able to explain to me gauge theory with surprising
| accuracy, it was able to do so using perfect Elizabethan english
| --exactly as I had instructed it to.
|
| There is a missing ingredient that LLMs lack, however. They lack
| insight. Writing is made engaging by the promise of insight
| teased in its setups, the depths that are dug through its
| payoffs, and the revelations found in its conclusion. It requires
| solving an abstract sudoku puzzle where each sentence builds on
| something prior and, critically, advances an agenda toward an
| emotional conclusion. This is the rhetoric inherent to all
| storytelling, but just as in a good political speech or debate,
| everything hinges on the quality of the central thesis--the key
| insight that LLMs do not come equipped to provide on their own.
|
| This is hard. Insight is hard. And an AI supporter would gladly
| tell you "yes! this is where prompting becomes art!" And perhaps
| there is merit to this, or at least there is merit insofar as Sam
| Altman's dreams of AI producing novel insights remain
| unfulfilled. This condition notwithstanding, what merit exactly
| do these supporters have? Has prompting become an art the same
| way that it has become engineering? It would seem AlphaWrite
| would like to say so.
|
| But let's look at this rubric and evaluate for ourselves what
| else AlphaWrite would like to say:
|
| ```python # Fallback to a basic rubric if file not found return
| """Creative writing evaluation should consider: 1. Creativity and
| Originality (25%) - Unique ideas, fresh perspectives, innovative
| storytelling 2. Writing Quality (25%) - Grammar, style, flow,
| vocabulary, sentence structure 3. Engagement (20%) - How
| compelling and interesting the piece is to read 4. Character
| Development (15%) - Believable, well-developed characters with
| clear motivations 5. Plot Structure (15%) - Logical progression,
| pacing, resolution of conflicts""" ```
|
| It's certainly just a default, and I mean no bad faith in using
| this for rhetorical effect, but this default also acts as a
| template, and it happens to be informative to my point. Insight,
| genuine insight, is hard because it is contingent on one's
| audience and one's shared experiences with them. It isn't enough
| to check boxes. Might I ask what makes for a better story: a
| narrative about a well developed princess who provides fresh
| perspectives on antiquated themes, or a narrative about a well
| developed stock broker who provides fresh perspectives on
| contemporary themes? The output fails to find its audience no
| matter what your rubric is.
|
| And here lies the dilemma regarding the idea that prompts are an
| art: they are not. The prompts are not art by the simple fact
| that nobody will read them. What is read is what all that is
| communicated and any discerning audience will be alienated by
| anything generated by something as ambiguous as a English
| teacher's grading rubric.
|
| I write because I want to communicate my insights to an audience
| who I believe would be influenced by them. I may be early in my
| career, but this is why I do it. The degree of influence I shall
| have measures the degree of "art" I shall attain. Not by whether
| or not I clear the minimum bar of literacy.
___________________________________________________________________
(page generated 2025-06-11 23:01 UTC)