[HN Gopher] Photorealistic Video Generation with Diffusion Models
___________________________________________________________________
Photorealistic Video Generation with Diffusion Models
Author : smusamashah
Score : 125 points
Date : 2023-12-11 18:00 UTC (5 hours ago)
(HTM) web link (walt-video-diffusion.github.io)
(TXT) w3m dump (walt-video-diffusion.github.io)
| youssefabdelm wrote:
| better but... still sucks... cant wait til it works
| londons_explore wrote:
| Is this another case where the architecture is trivial and
| obvious, but the data and compute requirements are obscene?
| j45 wrote:
| While this interests me, a lot of this stuff started obscene
| and became not obscene.
|
| A line of sight, however disappointing is still better than no
| line of sight.
|
| Maybe it will turn the open source efforts onto these things
| that are announced.
| VikingCoder wrote:
| 1. All the data in the entire world
|
| 2. ???
|
| 3. Profit
| 6510 wrote:
| They did wrote: "Images in this section is not generated by our
| model."
| thequadehunter wrote:
| No code..as usual :/
| sschueller wrote:
| Looks like a pretty standard comfyui svd workflow.
|
| https://github.com/thecooltechguy/ComfyUI-Stable-Video-Diffu...
| GaggiX wrote:
| This paper is not about SVD (Stable Video Diffusion).
| peab wrote:
| that's a different model, by stability AI.
|
| Right now everyone is coming out with video models, and
| they're all very similar.
|
| Stability, Meta, Google, Nvidia, tencent, Alibaba all have
| released either demos or code for Stable Diffusion based
| video generation. They're all fairly similar but slightly
| different architecture.
| ur-whale wrote:
| Amazing.
|
| The fact that we can synthesize stuff like this from text makes
| me believe the real world obeys some laws of physics we still
| haven't figured out.
|
| Specifically, the fact that when we interpolate in embedding
| space, the interpolated point still represents something almost
| always physically plausible (not always though, witness the
| surfing cat that doesn't get wet underwater).
| thfuran wrote:
| But the embedding space is deliberately and laboriously
| constructed such that it represents realistic things at as many
| points as possible. That's what the training is, right?
| renegade-otter wrote:
| Eh? It's automated statistical learning. Yes, it's very cool,
| but it is far from AGI, let alone magic.
|
| I am playing around with SD _right now_ - pulling sleepless
| nights because it 's so fascinating, but I do get to see the
| patterns.
|
| Generating unreal, magical stuff is actually a lot of fun, but
| if you want it to generate something that _makes sense_ - good
| luck. It is really hard.
|
| I can get Elizabeth Olsen into a Halo suit in a war zone on an
| alien world, but getting her to sit in a car like a normal
| human being has been impossible. It's always some twisted
| horror show. The predictive model does not know what to do with
| it and how it all fits. That's after you clean up all the
| anatomical nightmare fuel like the fingers situation.
|
| This is why most of the SD artwork you see are magical elf
| girls in the foreground or some unearthly monsters.
|
| A well-defined subject in the foreground (a human-ish one), it
| will do VERY well. If you really want to get your imagination
| going with every little detail being near-perfect, you will be
| busy.
|
| Oh, and the text - it obviously cannot do text. It's in the
| right place, but it looks funny.
|
| This is amazing tech, but it's not about to go sentient and
| kill me at night.
| echelon wrote:
| No no, I feel the OP is onto something.
|
| > obeys some laws of physics we still haven't figured out.
|
| That statement might have been over the top, but I think
| they're right with respect to the information. Spatial logic,
| optics, signal processing, reasoning. It's beginning to feel
| like there's an interesting representation that cross-cuts
| these spaces in a powerful, generalizable way.
| chankstein38 wrote:
| Agreed. I've generated tens of thousands of images with AI
| between Midjourney, Dall-E, and Stable Diffusion and yeah.
| The patterns almost ruin it for me. It's super interesting
| and a really cool tech but yeah at this point I can pretty
| much look at a picture and recognize it's AI generated just
| by the "vibe." Like you're saying, it's always this over
| focused single-subject image with a generic background.
|
| Like you said, as well, it's impossible to do some of the
| stuff that seems super simple but based in reality where
| completely dream-like images are generally pretty automatic.
| It's cool and, when I have another project that I need it
| for, I'll go nuts with it again but I kind of lost interest
| in just generating images to experiment because the
| experimentation stopped netting interesting results. It was
| just always the same stuff.
| einpoklum wrote:
| > The fact that we can synthesize stuff like this from text
|
| We can't synthesize stuff like that from text. It's synthesized
| from a(n obscenely large) combination of annotated still images
| and video clips. The text guides the synthesis (or perhaps
| better to call it reconstitution) process.
| plastic3169 wrote:
| Seems gruel to even say this as the tech is pure magic, but sad
| that these videos often get that ultra processed stockfootage
| look. I would be more impressed to just have bland everyday
| videos with muted colors, shaky cam and uninteresting people
| doing everyday stuff.
| nuz wrote:
| Loras can often achieve that look. No weights though but that
| would have likely worked to skew it in the 'amateur video'
| direction.
| eurekin wrote:
| I imagine some kind of service could pop up, exactly for that.
| For pictures, we already have:
|
| https://magnific.ai/
|
| Which seems to have been endorsed by one of the best retouchers
| I know of (https://www.youtube.com/@pratiknaikedu)
| jacobolus wrote:
| What they are asking for is to have these machine-generated
| "photorealistic" outputs look more like real photos/videos,
| not to turn every normal looking photo into a lurid dream-
| scene monstrosity with completely made up details.
| sschueller wrote:
| If you tweak the outputs you can get much better results but
| takes time and resources. You are also quite limited to the
| duration at this point.
|
| Here is an example I made yesterday, probably around 5 minutes of
| computing on my 3080ti: https://imgur.com/a/4dPv1hC
| eurekin wrote:
| Could you create a writeup (or record screen, or screenshot
| relevant config) how to reproduce that effect?
| sschueller wrote:
| I will do a post on my blog when I get a chance. Only got it
| working yesterday after a lot of trial and error mostly due
| to the lack of GPU power.
|
| It's pretty much the standard comfyui svd workflow with extra
| upscaling.
| yieldcrv wrote:
| reminds me of an ad for an FKK sauna club in central europe
|
| this would be good enough for promos of those!
| nuz wrote:
| This is about a different model than stable diffusion. You're
| using SVD I'm guessing, the weights here aren't released.
| kleiba wrote:
| OT: re "a spaceship destroyer approaching Earth" (example in the
| first row of the second set) - I cannot for the life of me
| generate an SD image of a two-person space ship (a shuttle, but
| more Star Wars X-Wing than USS Enterprise shuttle) flying towards
| the viewer.
|
| Any helpful tips?
| gaogao wrote:
| Sketch it out and then ControlNet img2img
| einpoklum wrote:
| I like how the destroyer is the love child of the Millenium
| Falcon and the front of a BSG battlestar.
|
| Also can't say no to a cat video (even if it doesn't really
| remind of Van Gogh).
| dylan604 wrote:
| I was chatting with a studio manager at a post f/x house when he
| received a demo reel that immediately was relocated to the trash
| without viewing it. I called him on it about not even looking at
| it before rejecting it. He had me take it out of the trash to
| play it. Before it stared, he told me exactly what would be on
| the tape and the order of the clips, which is exactly what was on
| the tape. He said it was the fall semester with such and such
| teacher at the school, and every submission was exactly the same.
| They were just the class assignments made into a demo reel.
|
| Everyone of these generative AI demo pages feels exactly like
| that. Nothing new other than we've done the same thing as someone
| else type of releases.
|
| Now, I'm willing to say I'm jaded, so if this represents
| something we haven't already seen umpteengillion times already,
| where did they bury the lede? It's all still uncanny valley. The
| wheels on do not look real. The wake from the viking's vacuum
| does not look real. I'm not sure what the corgi is walking on,
| but it doesn't look like the ground.
| echelon wrote:
| That studio manager will be out of a job in five years if he
| doesn't get on the bandwagon and evolve.
|
| This is the film to digital transition. It's going to be rough
| for a bit, then take over everything all at once.
|
| The upside for these editors is that they'll be able to be
| single-person studios and pursue their own vision without
| meddling from studios or massive capital, equipment, and
| personnel needs.
| dylan604 wrote:
| Ha! I took a phone call back when I was in the shiny round
| disc business that said based on the prices we were quoting,
| that we'd be out of business in a year because Apple released
| DVD Studio Pro. Instead, we gained even more business from
| people that tried to have their projects done in DVDSP, but
| then ran into road blocks making it impossible to do the job.
| They'd then come back to us, and receive an even higher quote
| since now there is even less time before their deadlines. cie
| la vie.
|
| You say out of business. I say, mediocrity all the way down
| from here.
| echelon wrote:
| > You say out of business. I say, mediocrity all the way
| down from here.
|
| Feels "Old man yells at cloud." There will be lots of
| garbage, for sure, but more people than ever before will
| have access to creation.
|
| Becoming a director means getting access to a very limited
| playing field. Very few domestic productions get greenlit
| per year due to the intense financial, logistical, and
| personnel requirements. Very few people are able to make
| the opportunity cost sacrifice to learn the tools and
| people network.
|
| For all the people that write fan fiction, web comics, or
| have a DeviantArt, Itch.io, or Newgrounds, this is going to
| give them access to their own personal Miyazaki/Scorsese
| tooling.
|
| It's going to be awesome seeing passionate people finally
| have access.
| dylan604 wrote:
| > Feels "Old man yells at cloud."
|
| It much more like, old man has sat through many a product
| demo and seen the brochures, but what's in the box rarely
| matches. If this is what's being demoed as the examples
| they've culled together, the tech is lacking.
| youssefabdelm wrote:
| Agree... they'll definitely evolve and get better, we're in the
| floppy disk era, but so far they suck. Nothing but weird
| motion, a shitty echo of mainstream / stock footage.
|
| I also get the sense that in the early days where data curation
| matters, research teams need to hire curators that can siphon
| good from bad examples. E.g. don't train on shit stock footage,
| pick up a book on fine art photographers, train on that shit +
| good films and you'll get a pretty good model.
|
| In the future I suspect curation will matter less and less
| because in the limit you should train on literally the entire
| internet.
|
| I'm definitely excited for where it might go though. If it ever
| develops to a sophisticated point, which it most certainly
| will, we'll be able to do stuff with it that is just impossible
| by typical production means.
| Keyframe wrote:
| Still a long way until clients will be able to play "pixel
| peeping" and "what's wrong with this shot" games with AI.
| Chores like matte/roto and matchmove are closer to danger.
| Distant mattes, procedural fill (even textures) will be
| supercharged with it though.
| dylan604 wrote:
| a lot of these examples look like bad photoshop examples to
| me though.
| golergka wrote:
| > Everyone of these generative AI demo pages feels exactly like
| that.
|
| Except you don't have to hire anyone and you get results in two
| minutes, not two days later.
| chankstein38 wrote:
| Truth but, to be fair, you could've done that with the other
| countless AI Video Generator papers/techs/releases we've
| seen.
| yreg wrote:
| So what's the complaint here? That researchers research?
|
| These folks wanted to make their own thing so they made
| their own thing. Eventually someone will make a thing
| that's far better than anything we have today.
| chankstein38 wrote:
| To add, the very first video, the parrot doesn't have an eye. A
| few videos down the line, the horse's legs meld together and
| split again as it's walking.
|
| Agreed, this is no more impressive than Stable Video or any of
| the other things like this we've seen.
| system2 wrote:
| I would love to see something longer than half a second, one day.
| This is the 50th video demo and it is still the same. Extremely
| short blurry gifs, always single angle and something morphing.
| Ai's video generation progress seems to be not going forward.
___________________________________________________________________
(page generated 2023-12-11 23:01 UTC)