[HN Gopher] Drawing a childrens story with DALL-E 2
       ___________________________________________________________________
        
       Drawing a childrens story with DALL-E 2
        
       Author : nutanc
       Score  : 58 points
       Date   : 2022-06-19 13:37 UTC (9 hours ago)
        
 (HTM) web link (twitter.com)
 (TXT) w3m dump (twitter.com)
        
       | YeGoblynQueenne wrote:
       | That's a really great idea! If we train our kids from a young
       | enough age to not notice the incoherence in content generated by
       | systems like DALL-E they will probably never learn to notice it.
       | Then, when they're adults, they won't mind it! And we can all
       | pretend it's not even there!
       | 
       | (e.g. why does the little girl keep changing hair, face, clothes,
       | age, ....)
        
         | renewiltord wrote:
         | Or perhaps they'll learn the concept / instance separation
         | earlier because they will comprehend the notion of an image
         | depicting something without needing continuity of characters to
         | express it.
        
           | YeGoblynQueenne wrote:
           | Oh, I'm sure it's possible to understand something and still
           | be annoyed by it.
           | 
           | For example, an incoherent story. I think that will piss off
           | most people.
        
         | ffhhj wrote:
         | The first rule of the utopia is not questioning the utopia ;)
        
         | [deleted]
        
       | jcims wrote:
       | Hey nutanc did you create the text or is that generated as well?
       | 
       | I had to mess around with the prompt a little but GPT-3 just did
       | this:
       | 
       | Prompt:                   Hi i'm Elsie, a first grader going to
       | school for the first time.           This is a six panel comic
       | speaking in the first person about how my day went.
       | I'll include a suggestion for the illustrator to draw each frame
       | in square brackets:
       | 
       | Generated text:                   [Elsie is standing in front of
       | her house with her backpack on, looking a little nervous.]
       | Hi, I'm Elsie. Today is my first day of school. I'm a little
       | nervous, but I'm excited too.                  [Elsie is walking
       | to school with her mother.]                  I walk to school
       | with my mom. We talk about what my day might be like.
       | [Elsie is standing in front of her school with her mother, who is
       | giving her a hug and kiss goodbye.]                  My mom gives
       | me a hug and a kiss, and then I'm off!                  [Elsie is
       | in her classroom, sitting at her desk.]                  I sit at
       | my desk and the teacher starts talking. I learn a lot of new
       | things today.                  [Elsie is standing in front of her
       | house again, this time with a big smile on her face.]
       | I had a great first day of school! I can't wait to go back
       | tomorrow.
       | 
       | Any time i mess with this stuff i get lost trying different
       | variations. Here's a few more (# marks the prompt)
       | https://pastebin.com/deC0RxDb
        
         | renewiltord wrote:
         | Remarkable. One could imagine the Primer from The Diamond Age
         | actually existing in the near future.
        
         | nutanc wrote:
         | The text is picked up from
         | https://storyweaver.org.in/stories/98-rani-s-first-day-at-sc...
         | for this example. But I have experimented with GPT3 generating
         | content and Dalle generating the image. For example here,
         | https://twitter.com/nutanc/status/1538372995764617216
        
           | blacksmith_tb wrote:
           | As the parent of a dyslexic kid, ALL CAPS is a sadistic
           | choice for something intended for children to read...
        
             | jcims wrote:
             | Can you expand on this a little? I've never heard of it.
             | Also have you seen the bionic font?
             | 
             | https://lithub.com/will-this-bionic-font-help-you-read-
             | faste...
             | 
             | Curious if that helps
        
           | [deleted]
        
         | dejobaan wrote:
         | Not the author, but tangentially relatedly, using GPT-3 and
         | Midjourney together is a bunch of fun. Here's a Penny Arcade
         | comic strip generated using both:
         | https://docs.google.com/document/d/1xfluFTKoMM5Avm9Nfkwe8pbp...
         | 
         | Not quite ready to replace comic artists.
        
           | jcims wrote:
           | I did this one with the same combo a little while back:
           | 
           | https://docs.google.com/presentation/d/e/2PACX-1vT4XWNx2SdEg.
           | ..
           | 
           | Lesson learned is that if you don't prompt it with a premise
           | it will probably copy it from something in the training data.
           | In this case the words are novel but the theme comes from a
           | book of the same name.
        
       | glutamate wrote:
       | How do you ensure a continuity of the main characters with a
       | DALL-E generated comic? The example here doesn't seem to do this
       | well. There characters are different in each frame
        
         | capableweb wrote:
         | I don't have access so not sure it would work, but maybe the
         | author could be more specific with the prompt to get a
         | character that looks more similar for each frame.
         | 
         | "Girl with brown hair, small nose, blue eyes, white shirt and
         | blue skirt goes inside school" for example.
        
           | nutanc wrote:
           | I tried this. Every time a new girl pops up :)
        
             | bredren wrote:
             | If it was run enough times you would get similar enough
             | looking girls which could be used!
             | 
             | Just need to train it on that girl somehow?
             | 
             | This is the infinite universes thing right?
        
             | capableweb wrote:
             | Hm, that's sad. Maybe in the future you'll be able to
             | assign specific results into variables (like saving one
             | result of this avatar you generated as
             | $SCARED_OF_SCHOOL_GIRL) and enforce that avatar to be used
             | in the future, would help in these cases.
        
       | jonahx wrote:
       | The pictures, taken one by one, are impressive.
       | 
       | Is it possible to tell DALL-E to use the same girl in each panel?
       | That is the only detail preventing me from being convinced this
       | isn't a real children's book.
        
         | gwern wrote:
         | The DALL-E API doesn't give you access to the latents or model,
         | so editing/consistency is much more difficult than it has to
         | be.
         | 
         | If you wanted to work within the current feature set, I think
         | inpainting or 'uncropping' is probably the way to go. For
         | example, generate a character until you get one that you like;
         | generate a bunch of variants of that one face in various
         | positions/angles/emotions (either through the variation feature
         | or by reprompting); now, cut the faces or figures out of each
         | one, and use them for each illustration as inpainting or uncrop
         | with a text prompt describing that scene.
         | 
         | So you'd generate your target girl, get a bunch of her faces
         | (uncertain, sad, facing left, facing right), cut them out for
         | an outline, and then inpaint/uncrop: "Shawna felt sad to be
         | away from home and Mom all day"+[sad-face.jpg]+"a brown-haired
         | young white girl in a gingham dress standing alone in a
         | colorful cozy kindergarten classroom"; and so on for each
         | scene.
        
         | nutanc wrote:
         | I tried a lot, but thats one shortcoming. You can try editing
         | and redrawing. Some success is there, but scope for improvement
         | is there.
        
           | bredren wrote:
           | Seems like this would seriously enhance this product.
           | 
           | As a comparison, stock photos and video of people are often
           | available as a scene showing the same folks in a bunch of
           | scenarios and angles.
           | 
           | You kind of need this to build something of any complexity
           | out of the art.
        
         | monkeydust wrote:
         | Just tried, I was curious. Answer is yes but not very well.
         | 
         | So you can generate a picture say 'picture of a boy playing
         | with a ball' then go into edit mode (think this is new) and
         | erase ball, then change prompt to 'picture of a boy playing
         | with toy car'.
         | 
         | It will keep the original elements of boy not erased and put in
         | a toy car. The character is kept the same but doesn't work as
         | well as you might think as the body, face pose are the same.
         | 
         | Still, you can see this getting better over time.
         | 
         | Only got access this morning, hard not be impressed.
        
         | grumbel wrote:
         | It can work when the character itself is in the training set,
         | e.g. "Homer Simpson" can give excellent results:
         | 
         | https://old.reddit.com/r/dalle2/comments/vfi8lj/homer_simpso...
         | 
         | For original characters, it's tricky, you have to use a simple
         | design that can be describe in text, in this example it's the
         | prominent red hoodie and dark pants that make it look like the
         | same character across pictures:
         | 
         | https://www.youtube.com/watch?v=J_laffNOoQw
         | 
         | Another problem is that DALL*E has a hard time telling
         | different characters apart, e.g. requesting Captain America and
         | Iron Man in the same image will give you two characters that
         | have attributes of both Captain America and Iron Man mixed
         | together:
         | 
         | https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/a39...
        
         | GistNoesis wrote:
         | I don't know if it's possible yet, but the technology behind
         | (guided diffusion), allows it. But you will probably have to
         | train it with a dataset that shows multiple associated images
         | together.
         | 
         | Denoising diffusions have been successfully applied to produce
         | coherent videos.
         | 
         | In fact, there are even some benefits to train on video because
         | images share information between them, which make context
         | extraction easier.
         | 
         | Counterintuitively, because of the shared information between
         | frames, their computations can also be partially shared, it can
         | be almost free to generate multiple frames at the same time : A
         | pair of RGB image can be considered as a 6D image, a 5-frame
         | sequence a 15D image, and the neural network can learn to
         | compress the various image channel together. The subsequent
         | neural network layers can be kept 256D (for example) the same
         | size as they were for a single frame.
         | 
         | Of course this trick won't work if the pair of image are too
         | different but you can always stack images side by side to have
         | a bigger film-roll image and you use a standard attention
         | mechanism to attend to various parts.
        
       | ffhhj wrote:
       | Seems useful to create quick sketches for an artist to re-draw
       | the scenes so they keep a given style. I can create 3D models
       | from sketches, but don't have much creativity to compose the
       | scenes, a tool like DALL-E would give me a big kick-start.
        
         | rebuilder wrote:
         | I'm not entirely sure about that. It feels like saying you
         | could just use an AI tool to sketch out a piece of music and
         | then have a composer fix it up.
         | 
         | I suspect it will turn out to be much more work for the artist
         | to find what it is you actually liked about the generated
         | output and rework the pictures based on that, rather than the
         | usual process where the client describes what they need and the
         | artist uses their experience and understanding of the human
         | mind to decide how to represent that.
         | 
         | I guess I'm saying that the art in "art" is not the superficial
         | skill of making pretty pictures, but the concurrently honed
         | skill of making meaningful choices. I'm not sure you can just
         | bypass learning one without compromising the other.
        
       | manimino wrote:
       | It ought to be possible to guide a diffusion-based model to make
       | several panels about a single, consistent, AI-invented character.
       | 
       | It can already generate several independent images containing
       | known characters that are in the training data (Obama, Pikachu,
       | etc.)
        
         | jcims wrote:
         | I'm a total layperson in this area but there used to be a
         | concept of 'fine tuning' with language models (may still be).
         | My understanding of those is that you could inject additional
         | training data that overlays(?) the existing model in some way
         | to help direct the output.
         | 
         | In this case it seems that you could fine-tune the generative
         | image model with previous frames to provide that continuity.
         | After all that's what we do when we read the panel, we
         | instantly store the previous one in memory so that we can
         | actually recognize the difference in the next panel.
        
       | mistrial9 wrote:
       | as a trained artist - I find this cancerous. Is it not obvious
       | that some extreme computer users are explicitly going into the
       | most sensitive and therefore fragile aspects of human life.. like
       | a thrill seeker addict going for more and more intensity.. major
       | no
        
         | whateveracct wrote:
         | HN is full of Idea Guys and Engineers - two clades of
         | businesspeople who see art skill as an unnecessary moat & cost
         | center.
         | 
         | I predict DALL-E art to be on the level of those freaky auto-
         | generated kids' videos for a long time. People see those as
         | successful too. Millions of Views with very little Cost.
        
           | mistrial9 wrote:
           | regarding graphic design and art skills on the PC, my
           | recollection of Microsoft competing with Apple was that the
           | frat guys and outsourcers at MSFT specifically outsourced any
           | artwork to get it as cheaply as possible, and actively mocked
           | expensive graphic designers, in much the same way they
           | actively mocked business rivals -- loudly jeering.
        
         | Tao3300 wrote:
         | > cancerous
         | 
         | That's the perfect word for this, because we have to ask what
         | this would look like if it metastisized.
         | 
         | Algorithm illustrates the book -> Algorithm writes the book ->
         | Algorithm generates the audio to read the book to your kids for
         | you -> Algorithm mines trends and engagement stats to decide
         | which books it will write.
         | 
         | It's the total annihilation of literature, and all the way more
         | creative people get cut out and the bottom line pads the
         | pockets of the publishers who own it. Humanities without
         | humanity, i.e. nothing.
        
         | [deleted]
        
         | ALittleLight wrote:
         | I'm sure professional drivers don't like the idea of self-
         | driving cars either.
        
           | omnicognate wrote:
           | There are strong reasons to want to eliminate drivers from
           | cars. I'm very optimistic about the overall impact of it on
           | society and the environment if it actually ends up happening.
           | 
           | Children's book illustrators, though, really? How do we
           | benefit if we make it impossible for that to be a living? Is
           | there any benefit to you or your children if they look at
           | pictures generated from a machine learning algorithm based on
           | past images instead of by a human artist that was paid for
           | their creation? Is there any benefit to artists? Is there any
           | benefit to society as a whole, or to any individuals except
           | the owners of the algorithms generating the pictures and the
           | publishers who save a relatively minor [1] cost?
           | 
           | Personally I doubt ML-illustrated children's books would
           | catch on as anything more than a novelty simply because
           | generating illustrations in this way seems so dystopian,
           | impersonal and tacky, but the way people react to new
           | technology is something I find very hard to predict.
           | 
           | [1] Typically between PS6,500 and PS8,500 per book according
           | to https://www.peopleofpublishing.com/post/how-much-does-a-
           | chil...
        
             | ALittleLight wrote:
             | I know multiple people who have ideas for children's books
             | but don't really want to actually produce one in the
             | current system. Easy to understand why people don't follow
             | through creating children's books - it's hard and not very
             | lucrative. Creating a children's book is usually a money
             | losing proposition - most books aren't published and most
             | published books don't sell very well.
             | 
             | Imagine taking the cost of illustrating a children's book
             | from 10k USD (plus who knows how much time) to paying a 10
             | dollar per month fee to create unlimited illustrations?
             | What would that enable? How many more children's books
             | would get produced? How much more happiness for children
             | and would be authors?
             | 
             | My mother used to tell us stories about characters she had
             | made up. With a little more time, if this technology
             | existed, she could have illustrated her stories and shown
             | them to us. Or, if this technology gets invented tomorrow,
             | I could create generate the illustrations from what I
             | remember of my mom's stories and share it with her. Why
             | wouldn't I want this? Because professional illustrators
             | would be devalued?
             | 
             | Let's not forget the children who lack interesting content
             | or concerned parents who would create it for them. These
             | kids could benefit from a reddit or imgur of children's
             | books and get an infinite scroll of content. If software
             | becomes good enough the infinite scroll could be generated
             | automatically.
        
               | mistrial9 wrote:
               | pass the extra lucre with those clouds, Lucas
        
           | mistrial9 wrote:
           | no - not the same .. what you say is machinery utilitarian,
           | for pay
        
             | ALittleLight wrote:
             | I am neither a professional driver nor a professional
             | artist but I have paid both for their services. In both
             | cases it's an exchange, money to take something somewhere
             | or money for the illustrations I want. They both, from my
             | perspective as a user, could be replaced by software that
             | performs well.
             | 
             | I'm reading your comment as saying that there is some sine
             | qua non about art that professional driving lacks - but I
             | think that may just be your bias as an artist. If software
             | produces results that are as good or better than yours,
             | what is lost by replacing artists with software?
             | 
             | I'm also not writing this just to dig at professional
             | artists and drivers. I think my career too is in the
             | process of being replaced by software. I have doubts and
             | misgivings about whether this will be entirely good, but it
             | seems clear that it is happening and that most (all?)
             | professions are or will soon be in a similar state.
        
               | mistrial9 wrote:
               | >what is lost by replacing artists with software ?
               | 
               | computers are useless, they can only provide answers
        
         | npc12345 wrote:
         | Srry but youre the horse carriage maker in the 1920s, adapt.
        
           | Tao3300 wrote:
           | This is different. You're talking about technology. I'm sure
           | there were artistically beautiful horse carriages, but they
           | are largely a functional thing. Illustration and writing are
           | themselves art. Uniquely human culture.
        
       | guerrilla wrote:
       | Ummm... looks as bad as you'd expect? What is there to talk
       | about?
        
         | tyingq wrote:
         | Human captions on the generated pictures are much better.
         | 
         | https://pbs.twimg.com/media/FVomoVbXoAI5pCB?format=jpg
        
         | baq wrote:
         | yeah exactly, I wouldn't want to read this book to my children.
         | 
         | it _is_ an alpha version of the future hence good content for
         | hacker news, not necessarily something that should go into a
         | real children 's book ;)
        
         | tartoran wrote:
         | Umm.. another somewhat lucrative profession killed by tech?
        
           | djmips wrote:
           | You're hedging about it being lucrative. I can tell you it's
           | not. Also this is pretty terrible. It'll need a lot of
           | improvement.
        
             | tartoran wrote:
             | Ok, even if it wasn't very lucrative it most likely to
             | become even less so.
        
           | fullshark wrote:
           | We don't say that here, we say it's another somewhat
           | lucrative profession "democratized."
        
         | djmips wrote:
         | I'm glad it's not just me that looks at this kind of thing and
         | thinks it's awful.
        
       | awillen wrote:
       | Naturally the first comments here are criticisms that the
       | characters change from frame to frame, but I think that's worth
       | putting aside, since any rational person will understand that
       | systems like DALL-E will have the ability to maintain a level of
       | continuity between images in the near future (if it's not already
       | there and just not exposed).
       | 
       | Besides that, this is pretty good - certainly a few things that
       | don't seem quite ideal (the fifth frame doesn't seem to capture
       | the meaning of the text especially well), but enough to make you
       | think that in not too long, it is fully plausible that DALL-E
       | will be fully able to act as an illustrator for this use case. On
       | the one hand, pretty exciting. On the other, certainly harrowing
       | for children's book illustrators.
       | 
       | I wonder how long it will be before I can use this for my
       | marketing emails. I sell dog treats, and I have a fairly simple
       | template with an image at the top. That's almost always a photo
       | of my products and/or my dogs. How long before I can just ask for
       | an image from DALL-E for something like "Dog sitting next to
       | grill in back yard with American flags and other patriotic
       | decorations" for the top of my Fourth of July email? How long
       | before I can feed it a picture of my products and have it
       | generate photorealistic images of dogs eating them, thus
       | replacing the photographer I use for product shoots? It's an
       | exciting prospect for me as a small business owner - lots of time
       | and money saved in an area where I'm not an expert - but
       | definitely pretty scary for people who create visual imagery of
       | any kind.
        
         | nutanc wrote:
         | Dalle2 is really good for stock marketing content. Your use
         | case is certainly feasible. Especially with the edit feature we
         | can make dogs eat your product. If you can DM me a product
         | image I can show a POC :)
        
         | synu wrote:
         | Maybe you can set it aside as a tech demo, but the ask to
         | please show this to your children feels somehow creepy to me.
         | They are still so young and impressionable, and providing them
         | with perceptibly incoherent stories and imagery just seems..
         | off.
        
           | capableweb wrote:
           | > providing them with perceptibly incoherent stories and
           | imagery just seems.. off
           | 
           | Not sure it really matters, half of all children books are
           | filled of incoherent stories and imagery, but I think it's
           | mostly adults who notice that.
        
         | jfoster wrote:
         | I think there's currently some barriers in doing that:
         | 
         | 1. OpenAI claim ownership of the content produced.
         | 
         | 2. The content policy forbids what you're describing.
         | (commercial use is ruled out)
         | 
         | Content policy: https://labs.openai.com/policies/content-policy
         | 
         | Sharing & publication policy:
         | https://openai.com/api/policies/sharing-publication/
        
       | random_upvoter wrote:
       | This is very impressive, technically. Surely looks like AI is
       | going to be the death of mediocre artists.
        
       | kazinator wrote:
       | Oh goodie, here comes the deluge of low-effort children's books.
        
         | capableweb wrote:
         | Yeah, because that would definitely be new! There are already a
         | TON of low-effort children's book, I'm not sure if DALL-E
         | generated ones would be a improvement or not, I've definitely
         | seen many that are worse than this.
        
       | ducktective wrote:
       | I suggested a similar thing and couple of people dropped in with
       | more specialized knowledge on this:
       | 
       | https://news.ycombinator.com/item?id=31425690
        
       | TheMagicHorsey wrote:
       | I wish there was a program that allowed you to draw some input
       | pictures/shapes, and the program takes that input and creates a
       | picture in a given style. That would be a gamechanger for
       | storytelling.
        
       | TheMagicHorsey wrote:
       | I wish there was a program that allowed you to draw some input
       | pictures/shapes, and the program takes that input and creates a
       | picture in a given style. That would be a gamechanger for
       | storytelling.
        
         | yojo wrote:
         | NVIDIA did something like that for landscapes:
         | https://www.nvidia.com/en-us/studio/canvas/
        
       | ALittleLight wrote:
       | I was just fantasizing about how to do stuff like this. In my
       | fantasy there were two DALL-E-like models. One for generating
       | characters and one for generating scenes and a system for adding
       | a generated character to a scene.
       | 
       | The user describes a character to the character model and gets
       | candidate illustrations. The user can then save a character by
       | giving it a name and use the name to generate variations of the
       | same character. I'm imagining that either you just have a lot of
       | sliders to vary the vector that the character is drawn from or
       | maybe it's possible to combine ideas (e.g. Character X +
       | running).
       | 
       | Then the user just illustrates the scenes they want, the
       | characters, combines them together, and gets a useful image.
        
       ___________________________________________________________________
       (page generated 2022-06-19 23:02 UTC)