[HN Gopher] Photorealistic Video Generation with Diffusion Models
       ___________________________________________________________________
        
       Photorealistic Video Generation with Diffusion Models
        
       Author : smusamashah
       Score  : 125 points
       Date   : 2023-12-11 18:00 UTC (5 hours ago)
        
 (HTM) web link (walt-video-diffusion.github.io)
 (TXT) w3m dump (walt-video-diffusion.github.io)
        
       | youssefabdelm wrote:
       | better but... still sucks... cant wait til it works
        
       | londons_explore wrote:
       | Is this another case where the architecture is trivial and
       | obvious, but the data and compute requirements are obscene?
        
         | j45 wrote:
         | While this interests me, a lot of this stuff started obscene
         | and became not obscene.
         | 
         | A line of sight, however disappointing is still better than no
         | line of sight.
         | 
         | Maybe it will turn the open source efforts onto these things
         | that are announced.
        
         | VikingCoder wrote:
         | 1. All the data in the entire world
         | 
         | 2. ???
         | 
         | 3. Profit
        
         | 6510 wrote:
         | They did wrote: "Images in this section is not generated by our
         | model."
        
       | thequadehunter wrote:
       | No code..as usual :/
        
         | sschueller wrote:
         | Looks like a pretty standard comfyui svd workflow.
         | 
         | https://github.com/thecooltechguy/ComfyUI-Stable-Video-Diffu...
        
           | GaggiX wrote:
           | This paper is not about SVD (Stable Video Diffusion).
        
           | peab wrote:
           | that's a different model, by stability AI.
           | 
           | Right now everyone is coming out with video models, and
           | they're all very similar.
           | 
           | Stability, Meta, Google, Nvidia, tencent, Alibaba all have
           | released either demos or code for Stable Diffusion based
           | video generation. They're all fairly similar but slightly
           | different architecture.
        
       | ur-whale wrote:
       | Amazing.
       | 
       | The fact that we can synthesize stuff like this from text makes
       | me believe the real world obeys some laws of physics we still
       | haven't figured out.
       | 
       | Specifically, the fact that when we interpolate in embedding
       | space, the interpolated point still represents something almost
       | always physically plausible (not always though, witness the
       | surfing cat that doesn't get wet underwater).
        
         | thfuran wrote:
         | But the embedding space is deliberately and laboriously
         | constructed such that it represents realistic things at as many
         | points as possible. That's what the training is, right?
        
         | renegade-otter wrote:
         | Eh? It's automated statistical learning. Yes, it's very cool,
         | but it is far from AGI, let alone magic.
         | 
         | I am playing around with SD _right now_ - pulling sleepless
         | nights because it 's so fascinating, but I do get to see the
         | patterns.
         | 
         | Generating unreal, magical stuff is actually a lot of fun, but
         | if you want it to generate something that _makes sense_ - good
         | luck. It is really hard.
         | 
         | I can get Elizabeth Olsen into a Halo suit in a war zone on an
         | alien world, but getting her to sit in a car like a normal
         | human being has been impossible. It's always some twisted
         | horror show. The predictive model does not know what to do with
         | it and how it all fits. That's after you clean up all the
         | anatomical nightmare fuel like the fingers situation.
         | 
         | This is why most of the SD artwork you see are magical elf
         | girls in the foreground or some unearthly monsters.
         | 
         | A well-defined subject in the foreground (a human-ish one), it
         | will do VERY well. If you really want to get your imagination
         | going with every little detail being near-perfect, you will be
         | busy.
         | 
         | Oh, and the text - it obviously cannot do text. It's in the
         | right place, but it looks funny.
         | 
         | This is amazing tech, but it's not about to go sentient and
         | kill me at night.
        
           | echelon wrote:
           | No no, I feel the OP is onto something.
           | 
           | > obeys some laws of physics we still haven't figured out.
           | 
           | That statement might have been over the top, but I think
           | they're right with respect to the information. Spatial logic,
           | optics, signal processing, reasoning. It's beginning to feel
           | like there's an interesting representation that cross-cuts
           | these spaces in a powerful, generalizable way.
        
           | chankstein38 wrote:
           | Agreed. I've generated tens of thousands of images with AI
           | between Midjourney, Dall-E, and Stable Diffusion and yeah.
           | The patterns almost ruin it for me. It's super interesting
           | and a really cool tech but yeah at this point I can pretty
           | much look at a picture and recognize it's AI generated just
           | by the "vibe." Like you're saying, it's always this over
           | focused single-subject image with a generic background.
           | 
           | Like you said, as well, it's impossible to do some of the
           | stuff that seems super simple but based in reality where
           | completely dream-like images are generally pretty automatic.
           | It's cool and, when I have another project that I need it
           | for, I'll go nuts with it again but I kind of lost interest
           | in just generating images to experiment because the
           | experimentation stopped netting interesting results. It was
           | just always the same stuff.
        
         | einpoklum wrote:
         | > The fact that we can synthesize stuff like this from text
         | 
         | We can't synthesize stuff like that from text. It's synthesized
         | from a(n obscenely large) combination of annotated still images
         | and video clips. The text guides the synthesis (or perhaps
         | better to call it reconstitution) process.
        
       | plastic3169 wrote:
       | Seems gruel to even say this as the tech is pure magic, but sad
       | that these videos often get that ultra processed stockfootage
       | look. I would be more impressed to just have bland everyday
       | videos with muted colors, shaky cam and uninteresting people
       | doing everyday stuff.
        
         | nuz wrote:
         | Loras can often achieve that look. No weights though but that
         | would have likely worked to skew it in the 'amateur video'
         | direction.
        
         | eurekin wrote:
         | I imagine some kind of service could pop up, exactly for that.
         | For pictures, we already have:
         | 
         | https://magnific.ai/
         | 
         | Which seems to have been endorsed by one of the best retouchers
         | I know of (https://www.youtube.com/@pratiknaikedu)
        
           | jacobolus wrote:
           | What they are asking for is to have these machine-generated
           | "photorealistic" outputs look more like real photos/videos,
           | not to turn every normal looking photo into a lurid dream-
           | scene monstrosity with completely made up details.
        
       | sschueller wrote:
       | If you tweak the outputs you can get much better results but
       | takes time and resources. You are also quite limited to the
       | duration at this point.
       | 
       | Here is an example I made yesterday, probably around 5 minutes of
       | computing on my 3080ti: https://imgur.com/a/4dPv1hC
        
         | eurekin wrote:
         | Could you create a writeup (or record screen, or screenshot
         | relevant config) how to reproduce that effect?
        
           | sschueller wrote:
           | I will do a post on my blog when I get a chance. Only got it
           | working yesterday after a lot of trial and error mostly due
           | to the lack of GPU power.
           | 
           | It's pretty much the standard comfyui svd workflow with extra
           | upscaling.
        
         | yieldcrv wrote:
         | reminds me of an ad for an FKK sauna club in central europe
         | 
         | this would be good enough for promos of those!
        
         | nuz wrote:
         | This is about a different model than stable diffusion. You're
         | using SVD I'm guessing, the weights here aren't released.
        
       | kleiba wrote:
       | OT: re "a spaceship destroyer approaching Earth" (example in the
       | first row of the second set) - I cannot for the life of me
       | generate an SD image of a two-person space ship (a shuttle, but
       | more Star Wars X-Wing than USS Enterprise shuttle) flying towards
       | the viewer.
       | 
       | Any helpful tips?
        
         | gaogao wrote:
         | Sketch it out and then ControlNet img2img
        
       | einpoklum wrote:
       | I like how the destroyer is the love child of the Millenium
       | Falcon and the front of a BSG battlestar.
       | 
       | Also can't say no to a cat video (even if it doesn't really
       | remind of Van Gogh).
        
       | dylan604 wrote:
       | I was chatting with a studio manager at a post f/x house when he
       | received a demo reel that immediately was relocated to the trash
       | without viewing it. I called him on it about not even looking at
       | it before rejecting it. He had me take it out of the trash to
       | play it. Before it stared, he told me exactly what would be on
       | the tape and the order of the clips, which is exactly what was on
       | the tape. He said it was the fall semester with such and such
       | teacher at the school, and every submission was exactly the same.
       | They were just the class assignments made into a demo reel.
       | 
       | Everyone of these generative AI demo pages feels exactly like
       | that. Nothing new other than we've done the same thing as someone
       | else type of releases.
       | 
       | Now, I'm willing to say I'm jaded, so if this represents
       | something we haven't already seen umpteengillion times already,
       | where did they bury the lede? It's all still uncanny valley. The
       | wheels on do not look real. The wake from the viking's vacuum
       | does not look real. I'm not sure what the corgi is walking on,
       | but it doesn't look like the ground.
        
         | echelon wrote:
         | That studio manager will be out of a job in five years if he
         | doesn't get on the bandwagon and evolve.
         | 
         | This is the film to digital transition. It's going to be rough
         | for a bit, then take over everything all at once.
         | 
         | The upside for these editors is that they'll be able to be
         | single-person studios and pursue their own vision without
         | meddling from studios or massive capital, equipment, and
         | personnel needs.
        
           | dylan604 wrote:
           | Ha! I took a phone call back when I was in the shiny round
           | disc business that said based on the prices we were quoting,
           | that we'd be out of business in a year because Apple released
           | DVD Studio Pro. Instead, we gained even more business from
           | people that tried to have their projects done in DVDSP, but
           | then ran into road blocks making it impossible to do the job.
           | They'd then come back to us, and receive an even higher quote
           | since now there is even less time before their deadlines. cie
           | la vie.
           | 
           | You say out of business. I say, mediocrity all the way down
           | from here.
        
             | echelon wrote:
             | > You say out of business. I say, mediocrity all the way
             | down from here.
             | 
             | Feels "Old man yells at cloud." There will be lots of
             | garbage, for sure, but more people than ever before will
             | have access to creation.
             | 
             | Becoming a director means getting access to a very limited
             | playing field. Very few domestic productions get greenlit
             | per year due to the intense financial, logistical, and
             | personnel requirements. Very few people are able to make
             | the opportunity cost sacrifice to learn the tools and
             | people network.
             | 
             | For all the people that write fan fiction, web comics, or
             | have a DeviantArt, Itch.io, or Newgrounds, this is going to
             | give them access to their own personal Miyazaki/Scorsese
             | tooling.
             | 
             | It's going to be awesome seeing passionate people finally
             | have access.
        
               | dylan604 wrote:
               | > Feels "Old man yells at cloud."
               | 
               | It much more like, old man has sat through many a product
               | demo and seen the brochures, but what's in the box rarely
               | matches. If this is what's being demoed as the examples
               | they've culled together, the tech is lacking.
        
         | youssefabdelm wrote:
         | Agree... they'll definitely evolve and get better, we're in the
         | floppy disk era, but so far they suck. Nothing but weird
         | motion, a shitty echo of mainstream / stock footage.
         | 
         | I also get the sense that in the early days where data curation
         | matters, research teams need to hire curators that can siphon
         | good from bad examples. E.g. don't train on shit stock footage,
         | pick up a book on fine art photographers, train on that shit +
         | good films and you'll get a pretty good model.
         | 
         | In the future I suspect curation will matter less and less
         | because in the limit you should train on literally the entire
         | internet.
         | 
         | I'm definitely excited for where it might go though. If it ever
         | develops to a sophisticated point, which it most certainly
         | will, we'll be able to do stuff with it that is just impossible
         | by typical production means.
        
         | Keyframe wrote:
         | Still a long way until clients will be able to play "pixel
         | peeping" and "what's wrong with this shot" games with AI.
         | Chores like matte/roto and matchmove are closer to danger.
         | Distant mattes, procedural fill (even textures) will be
         | supercharged with it though.
        
           | dylan604 wrote:
           | a lot of these examples look like bad photoshop examples to
           | me though.
        
         | golergka wrote:
         | > Everyone of these generative AI demo pages feels exactly like
         | that.
         | 
         | Except you don't have to hire anyone and you get results in two
         | minutes, not two days later.
        
           | chankstein38 wrote:
           | Truth but, to be fair, you could've done that with the other
           | countless AI Video Generator papers/techs/releases we've
           | seen.
        
             | yreg wrote:
             | So what's the complaint here? That researchers research?
             | 
             | These folks wanted to make their own thing so they made
             | their own thing. Eventually someone will make a thing
             | that's far better than anything we have today.
        
         | chankstein38 wrote:
         | To add, the very first video, the parrot doesn't have an eye. A
         | few videos down the line, the horse's legs meld together and
         | split again as it's walking.
         | 
         | Agreed, this is no more impressive than Stable Video or any of
         | the other things like this we've seen.
        
       | system2 wrote:
       | I would love to see something longer than half a second, one day.
       | This is the 50th video demo and it is still the same. Extremely
       | short blurry gifs, always single angle and something morphing.
       | Ai's video generation progress seems to be not going forward.
        
       ___________________________________________________________________
       (page generated 2023-12-11 23:01 UTC)