[HN Gopher] Drawing a childrens story with DALL-E 2
___________________________________________________________________
Drawing a childrens story with DALL-E 2
Author : nutanc
Score : 58 points
Date : 2022-06-19 13:37 UTC (9 hours ago)
(HTM) web link (twitter.com)
(TXT) w3m dump (twitter.com)
| YeGoblynQueenne wrote:
| That's a really great idea! If we train our kids from a young
| enough age to not notice the incoherence in content generated by
| systems like DALL-E they will probably never learn to notice it.
| Then, when they're adults, they won't mind it! And we can all
| pretend it's not even there!
|
| (e.g. why does the little girl keep changing hair, face, clothes,
| age, ....)
| renewiltord wrote:
| Or perhaps they'll learn the concept / instance separation
| earlier because they will comprehend the notion of an image
| depicting something without needing continuity of characters to
| express it.
| YeGoblynQueenne wrote:
| Oh, I'm sure it's possible to understand something and still
| be annoyed by it.
|
| For example, an incoherent story. I think that will piss off
| most people.
| ffhhj wrote:
| The first rule of the utopia is not questioning the utopia ;)
| [deleted]
| jcims wrote:
| Hey nutanc did you create the text or is that generated as well?
|
| I had to mess around with the prompt a little but GPT-3 just did
| this:
|
| Prompt: Hi i'm Elsie, a first grader going to
| school for the first time. This is a six panel comic
| speaking in the first person about how my day went.
| I'll include a suggestion for the illustrator to draw each frame
| in square brackets:
|
| Generated text: [Elsie is standing in front of
| her house with her backpack on, looking a little nervous.]
| Hi, I'm Elsie. Today is my first day of school. I'm a little
| nervous, but I'm excited too. [Elsie is walking
| to school with her mother.] I walk to school
| with my mom. We talk about what my day might be like.
| [Elsie is standing in front of her school with her mother, who is
| giving her a hug and kiss goodbye.] My mom gives
| me a hug and a kiss, and then I'm off! [Elsie is
| in her classroom, sitting at her desk.] I sit at
| my desk and the teacher starts talking. I learn a lot of new
| things today. [Elsie is standing in front of her
| house again, this time with a big smile on her face.]
| I had a great first day of school! I can't wait to go back
| tomorrow.
|
| Any time i mess with this stuff i get lost trying different
| variations. Here's a few more (# marks the prompt)
| https://pastebin.com/deC0RxDb
| renewiltord wrote:
| Remarkable. One could imagine the Primer from The Diamond Age
| actually existing in the near future.
| nutanc wrote:
| The text is picked up from
| https://storyweaver.org.in/stories/98-rani-s-first-day-at-sc...
| for this example. But I have experimented with GPT3 generating
| content and Dalle generating the image. For example here,
| https://twitter.com/nutanc/status/1538372995764617216
| blacksmith_tb wrote:
| As the parent of a dyslexic kid, ALL CAPS is a sadistic
| choice for something intended for children to read...
| jcims wrote:
| Can you expand on this a little? I've never heard of it.
| Also have you seen the bionic font?
|
| https://lithub.com/will-this-bionic-font-help-you-read-
| faste...
|
| Curious if that helps
| [deleted]
| dejobaan wrote:
| Not the author, but tangentially relatedly, using GPT-3 and
| Midjourney together is a bunch of fun. Here's a Penny Arcade
| comic strip generated using both:
| https://docs.google.com/document/d/1xfluFTKoMM5Avm9Nfkwe8pbp...
|
| Not quite ready to replace comic artists.
| jcims wrote:
| I did this one with the same combo a little while back:
|
| https://docs.google.com/presentation/d/e/2PACX-1vT4XWNx2SdEg.
| ..
|
| Lesson learned is that if you don't prompt it with a premise
| it will probably copy it from something in the training data.
| In this case the words are novel but the theme comes from a
| book of the same name.
| glutamate wrote:
| How do you ensure a continuity of the main characters with a
| DALL-E generated comic? The example here doesn't seem to do this
| well. There characters are different in each frame
| capableweb wrote:
| I don't have access so not sure it would work, but maybe the
| author could be more specific with the prompt to get a
| character that looks more similar for each frame.
|
| "Girl with brown hair, small nose, blue eyes, white shirt and
| blue skirt goes inside school" for example.
| nutanc wrote:
| I tried this. Every time a new girl pops up :)
| bredren wrote:
| If it was run enough times you would get similar enough
| looking girls which could be used!
|
| Just need to train it on that girl somehow?
|
| This is the infinite universes thing right?
| capableweb wrote:
| Hm, that's sad. Maybe in the future you'll be able to
| assign specific results into variables (like saving one
| result of this avatar you generated as
| $SCARED_OF_SCHOOL_GIRL) and enforce that avatar to be used
| in the future, would help in these cases.
| jonahx wrote:
| The pictures, taken one by one, are impressive.
|
| Is it possible to tell DALL-E to use the same girl in each panel?
| That is the only detail preventing me from being convinced this
| isn't a real children's book.
| gwern wrote:
| The DALL-E API doesn't give you access to the latents or model,
| so editing/consistency is much more difficult than it has to
| be.
|
| If you wanted to work within the current feature set, I think
| inpainting or 'uncropping' is probably the way to go. For
| example, generate a character until you get one that you like;
| generate a bunch of variants of that one face in various
| positions/angles/emotions (either through the variation feature
| or by reprompting); now, cut the faces or figures out of each
| one, and use them for each illustration as inpainting or uncrop
| with a text prompt describing that scene.
|
| So you'd generate your target girl, get a bunch of her faces
| (uncertain, sad, facing left, facing right), cut them out for
| an outline, and then inpaint/uncrop: "Shawna felt sad to be
| away from home and Mom all day"+[sad-face.jpg]+"a brown-haired
| young white girl in a gingham dress standing alone in a
| colorful cozy kindergarten classroom"; and so on for each
| scene.
| nutanc wrote:
| I tried a lot, but thats one shortcoming. You can try editing
| and redrawing. Some success is there, but scope for improvement
| is there.
| bredren wrote:
| Seems like this would seriously enhance this product.
|
| As a comparison, stock photos and video of people are often
| available as a scene showing the same folks in a bunch of
| scenarios and angles.
|
| You kind of need this to build something of any complexity
| out of the art.
| monkeydust wrote:
| Just tried, I was curious. Answer is yes but not very well.
|
| So you can generate a picture say 'picture of a boy playing
| with a ball' then go into edit mode (think this is new) and
| erase ball, then change prompt to 'picture of a boy playing
| with toy car'.
|
| It will keep the original elements of boy not erased and put in
| a toy car. The character is kept the same but doesn't work as
| well as you might think as the body, face pose are the same.
|
| Still, you can see this getting better over time.
|
| Only got access this morning, hard not be impressed.
| grumbel wrote:
| It can work when the character itself is in the training set,
| e.g. "Homer Simpson" can give excellent results:
|
| https://old.reddit.com/r/dalle2/comments/vfi8lj/homer_simpso...
|
| For original characters, it's tricky, you have to use a simple
| design that can be describe in text, in this example it's the
| prominent red hoodie and dark pants that make it look like the
| same character across pictures:
|
| https://www.youtube.com/watch?v=J_laffNOoQw
|
| Another problem is that DALL*E has a hard time telling
| different characters apart, e.g. requesting Captain America and
| Iron Man in the same image will give you two characters that
| have attributes of both Captain America and Iron Man mixed
| together:
|
| https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/a39...
| GistNoesis wrote:
| I don't know if it's possible yet, but the technology behind
| (guided diffusion), allows it. But you will probably have to
| train it with a dataset that shows multiple associated images
| together.
|
| Denoising diffusions have been successfully applied to produce
| coherent videos.
|
| In fact, there are even some benefits to train on video because
| images share information between them, which make context
| extraction easier.
|
| Counterintuitively, because of the shared information between
| frames, their computations can also be partially shared, it can
| be almost free to generate multiple frames at the same time : A
| pair of RGB image can be considered as a 6D image, a 5-frame
| sequence a 15D image, and the neural network can learn to
| compress the various image channel together. The subsequent
| neural network layers can be kept 256D (for example) the same
| size as they were for a single frame.
|
| Of course this trick won't work if the pair of image are too
| different but you can always stack images side by side to have
| a bigger film-roll image and you use a standard attention
| mechanism to attend to various parts.
| ffhhj wrote:
| Seems useful to create quick sketches for an artist to re-draw
| the scenes so they keep a given style. I can create 3D models
| from sketches, but don't have much creativity to compose the
| scenes, a tool like DALL-E would give me a big kick-start.
| rebuilder wrote:
| I'm not entirely sure about that. It feels like saying you
| could just use an AI tool to sketch out a piece of music and
| then have a composer fix it up.
|
| I suspect it will turn out to be much more work for the artist
| to find what it is you actually liked about the generated
| output and rework the pictures based on that, rather than the
| usual process where the client describes what they need and the
| artist uses their experience and understanding of the human
| mind to decide how to represent that.
|
| I guess I'm saying that the art in "art" is not the superficial
| skill of making pretty pictures, but the concurrently honed
| skill of making meaningful choices. I'm not sure you can just
| bypass learning one without compromising the other.
| manimino wrote:
| It ought to be possible to guide a diffusion-based model to make
| several panels about a single, consistent, AI-invented character.
|
| It can already generate several independent images containing
| known characters that are in the training data (Obama, Pikachu,
| etc.)
| jcims wrote:
| I'm a total layperson in this area but there used to be a
| concept of 'fine tuning' with language models (may still be).
| My understanding of those is that you could inject additional
| training data that overlays(?) the existing model in some way
| to help direct the output.
|
| In this case it seems that you could fine-tune the generative
| image model with previous frames to provide that continuity.
| After all that's what we do when we read the panel, we
| instantly store the previous one in memory so that we can
| actually recognize the difference in the next panel.
| mistrial9 wrote:
| as a trained artist - I find this cancerous. Is it not obvious
| that some extreme computer users are explicitly going into the
| most sensitive and therefore fragile aspects of human life.. like
| a thrill seeker addict going for more and more intensity.. major
| no
| whateveracct wrote:
| HN is full of Idea Guys and Engineers - two clades of
| businesspeople who see art skill as an unnecessary moat & cost
| center.
|
| I predict DALL-E art to be on the level of those freaky auto-
| generated kids' videos for a long time. People see those as
| successful too. Millions of Views with very little Cost.
| mistrial9 wrote:
| regarding graphic design and art skills on the PC, my
| recollection of Microsoft competing with Apple was that the
| frat guys and outsourcers at MSFT specifically outsourced any
| artwork to get it as cheaply as possible, and actively mocked
| expensive graphic designers, in much the same way they
| actively mocked business rivals -- loudly jeering.
| Tao3300 wrote:
| > cancerous
|
| That's the perfect word for this, because we have to ask what
| this would look like if it metastisized.
|
| Algorithm illustrates the book -> Algorithm writes the book ->
| Algorithm generates the audio to read the book to your kids for
| you -> Algorithm mines trends and engagement stats to decide
| which books it will write.
|
| It's the total annihilation of literature, and all the way more
| creative people get cut out and the bottom line pads the
| pockets of the publishers who own it. Humanities without
| humanity, i.e. nothing.
| [deleted]
| ALittleLight wrote:
| I'm sure professional drivers don't like the idea of self-
| driving cars either.
| omnicognate wrote:
| There are strong reasons to want to eliminate drivers from
| cars. I'm very optimistic about the overall impact of it on
| society and the environment if it actually ends up happening.
|
| Children's book illustrators, though, really? How do we
| benefit if we make it impossible for that to be a living? Is
| there any benefit to you or your children if they look at
| pictures generated from a machine learning algorithm based on
| past images instead of by a human artist that was paid for
| their creation? Is there any benefit to artists? Is there any
| benefit to society as a whole, or to any individuals except
| the owners of the algorithms generating the pictures and the
| publishers who save a relatively minor [1] cost?
|
| Personally I doubt ML-illustrated children's books would
| catch on as anything more than a novelty simply because
| generating illustrations in this way seems so dystopian,
| impersonal and tacky, but the way people react to new
| technology is something I find very hard to predict.
|
| [1] Typically between PS6,500 and PS8,500 per book according
| to https://www.peopleofpublishing.com/post/how-much-does-a-
| chil...
| ALittleLight wrote:
| I know multiple people who have ideas for children's books
| but don't really want to actually produce one in the
| current system. Easy to understand why people don't follow
| through creating children's books - it's hard and not very
| lucrative. Creating a children's book is usually a money
| losing proposition - most books aren't published and most
| published books don't sell very well.
|
| Imagine taking the cost of illustrating a children's book
| from 10k USD (plus who knows how much time) to paying a 10
| dollar per month fee to create unlimited illustrations?
| What would that enable? How many more children's books
| would get produced? How much more happiness for children
| and would be authors?
|
| My mother used to tell us stories about characters she had
| made up. With a little more time, if this technology
| existed, she could have illustrated her stories and shown
| them to us. Or, if this technology gets invented tomorrow,
| I could create generate the illustrations from what I
| remember of my mom's stories and share it with her. Why
| wouldn't I want this? Because professional illustrators
| would be devalued?
|
| Let's not forget the children who lack interesting content
| or concerned parents who would create it for them. These
| kids could benefit from a reddit or imgur of children's
| books and get an infinite scroll of content. If software
| becomes good enough the infinite scroll could be generated
| automatically.
| mistrial9 wrote:
| pass the extra lucre with those clouds, Lucas
| mistrial9 wrote:
| no - not the same .. what you say is machinery utilitarian,
| for pay
| ALittleLight wrote:
| I am neither a professional driver nor a professional
| artist but I have paid both for their services. In both
| cases it's an exchange, money to take something somewhere
| or money for the illustrations I want. They both, from my
| perspective as a user, could be replaced by software that
| performs well.
|
| I'm reading your comment as saying that there is some sine
| qua non about art that professional driving lacks - but I
| think that may just be your bias as an artist. If software
| produces results that are as good or better than yours,
| what is lost by replacing artists with software?
|
| I'm also not writing this just to dig at professional
| artists and drivers. I think my career too is in the
| process of being replaced by software. I have doubts and
| misgivings about whether this will be entirely good, but it
| seems clear that it is happening and that most (all?)
| professions are or will soon be in a similar state.
| mistrial9 wrote:
| >what is lost by replacing artists with software ?
|
| computers are useless, they can only provide answers
| npc12345 wrote:
| Srry but youre the horse carriage maker in the 1920s, adapt.
| Tao3300 wrote:
| This is different. You're talking about technology. I'm sure
| there were artistically beautiful horse carriages, but they
| are largely a functional thing. Illustration and writing are
| themselves art. Uniquely human culture.
| guerrilla wrote:
| Ummm... looks as bad as you'd expect? What is there to talk
| about?
| tyingq wrote:
| Human captions on the generated pictures are much better.
|
| https://pbs.twimg.com/media/FVomoVbXoAI5pCB?format=jpg
| baq wrote:
| yeah exactly, I wouldn't want to read this book to my children.
|
| it _is_ an alpha version of the future hence good content for
| hacker news, not necessarily something that should go into a
| real children 's book ;)
| tartoran wrote:
| Umm.. another somewhat lucrative profession killed by tech?
| djmips wrote:
| You're hedging about it being lucrative. I can tell you it's
| not. Also this is pretty terrible. It'll need a lot of
| improvement.
| tartoran wrote:
| Ok, even if it wasn't very lucrative it most likely to
| become even less so.
| fullshark wrote:
| We don't say that here, we say it's another somewhat
| lucrative profession "democratized."
| djmips wrote:
| I'm glad it's not just me that looks at this kind of thing and
| thinks it's awful.
| awillen wrote:
| Naturally the first comments here are criticisms that the
| characters change from frame to frame, but I think that's worth
| putting aside, since any rational person will understand that
| systems like DALL-E will have the ability to maintain a level of
| continuity between images in the near future (if it's not already
| there and just not exposed).
|
| Besides that, this is pretty good - certainly a few things that
| don't seem quite ideal (the fifth frame doesn't seem to capture
| the meaning of the text especially well), but enough to make you
| think that in not too long, it is fully plausible that DALL-E
| will be fully able to act as an illustrator for this use case. On
| the one hand, pretty exciting. On the other, certainly harrowing
| for children's book illustrators.
|
| I wonder how long it will be before I can use this for my
| marketing emails. I sell dog treats, and I have a fairly simple
| template with an image at the top. That's almost always a photo
| of my products and/or my dogs. How long before I can just ask for
| an image from DALL-E for something like "Dog sitting next to
| grill in back yard with American flags and other patriotic
| decorations" for the top of my Fourth of July email? How long
| before I can feed it a picture of my products and have it
| generate photorealistic images of dogs eating them, thus
| replacing the photographer I use for product shoots? It's an
| exciting prospect for me as a small business owner - lots of time
| and money saved in an area where I'm not an expert - but
| definitely pretty scary for people who create visual imagery of
| any kind.
| nutanc wrote:
| Dalle2 is really good for stock marketing content. Your use
| case is certainly feasible. Especially with the edit feature we
| can make dogs eat your product. If you can DM me a product
| image I can show a POC :)
| synu wrote:
| Maybe you can set it aside as a tech demo, but the ask to
| please show this to your children feels somehow creepy to me.
| They are still so young and impressionable, and providing them
| with perceptibly incoherent stories and imagery just seems..
| off.
| capableweb wrote:
| > providing them with perceptibly incoherent stories and
| imagery just seems.. off
|
| Not sure it really matters, half of all children books are
| filled of incoherent stories and imagery, but I think it's
| mostly adults who notice that.
| jfoster wrote:
| I think there's currently some barriers in doing that:
|
| 1. OpenAI claim ownership of the content produced.
|
| 2. The content policy forbids what you're describing.
| (commercial use is ruled out)
|
| Content policy: https://labs.openai.com/policies/content-policy
|
| Sharing & publication policy:
| https://openai.com/api/policies/sharing-publication/
| random_upvoter wrote:
| This is very impressive, technically. Surely looks like AI is
| going to be the death of mediocre artists.
| kazinator wrote:
| Oh goodie, here comes the deluge of low-effort children's books.
| capableweb wrote:
| Yeah, because that would definitely be new! There are already a
| TON of low-effort children's book, I'm not sure if DALL-E
| generated ones would be a improvement or not, I've definitely
| seen many that are worse than this.
| ducktective wrote:
| I suggested a similar thing and couple of people dropped in with
| more specialized knowledge on this:
|
| https://news.ycombinator.com/item?id=31425690
| TheMagicHorsey wrote:
| I wish there was a program that allowed you to draw some input
| pictures/shapes, and the program takes that input and creates a
| picture in a given style. That would be a gamechanger for
| storytelling.
| TheMagicHorsey wrote:
| I wish there was a program that allowed you to draw some input
| pictures/shapes, and the program takes that input and creates a
| picture in a given style. That would be a gamechanger for
| storytelling.
| yojo wrote:
| NVIDIA did something like that for landscapes:
| https://www.nvidia.com/en-us/studio/canvas/
| ALittleLight wrote:
| I was just fantasizing about how to do stuff like this. In my
| fantasy there were two DALL-E-like models. One for generating
| characters and one for generating scenes and a system for adding
| a generated character to a scene.
|
| The user describes a character to the character model and gets
| candidate illustrations. The user can then save a character by
| giving it a name and use the name to generate variations of the
| same character. I'm imagining that either you just have a lot of
| sliders to vary the vector that the character is drawn from or
| maybe it's possible to combine ideas (e.g. Character X +
| running).
|
| Then the user just illustrates the scenes they want, the
| characters, combines them together, and gets a useful image.
___________________________________________________________________
(page generated 2022-06-19 23:02 UTC)