[HN Gopher] Efficient high-resolution image synthesis with linea...
___________________________________________________________________
Efficient high-resolution image synthesis with linear diffusion
transformer
Author : Vt71fcAqt7
Score : 129 points
Date : 2024-10-16 14:56 UTC (8 hours ago)
(HTM) web link (nvlabs.github.io)
(TXT) w3m dump (nvlabs.github.io)
| smusamashah wrote:
| > (e.g. Flux-12B), being 20 times smaller and 100+ times faster
| in measured throughput. Moreover, Sana-0.6B can be deployed on a
| 16GB laptop GPU, taking less than 1 second to generate a 1024 x
| 1024 resolution image.
| echelon wrote:
| Image models are going to be widely available. They'll probably
| be a dime a dozen soon. It's great that an increasing number of
| models are going open, because these are the ecosystems that will
| grow.
|
| 3D models (sculpts, texture, retopo, etc.) are following a
| similar trend and trajectory.
|
| Open video models are lagging behind by several years. While
| CogVideo and Pyramid are promising, video models are petabyte
| scale and so much more costly to build and train.
|
| I'm hoping video becomes free and cheap, but it's looking like we
| might be waiting a while.
|
| Major kudos to all of the teams building and training open source
| models!
| cube2222 wrote:
| This looks like quite a huge breakthrough, unless I'm missing
| something?
|
| ~25x faster performance than Flux-dev, while offering comparable
| quality in benchmarks. And visually the examples (surely cherry-
| picked, but still) look great!
|
| Especially since with GenAI the best way to get good results is
| to just generate a large amount of them and pick the best (imo).
| Performance like this will make that much easier/faster/cheaper.
|
| Code is unfortunately "(Coming soon)" for now. Can't wait to play
| with it!
| liuliu wrote:
| If you read closer to the benchmark, it seems to be slightly
| worse than FLUX [dev] on prompt adherence and quality. However,
| the best is to evaluate the result oneself, and the track-
| record of PixArt Sigma (from the same author?) is pretty good!
| Archit3ch wrote:
| If you generate 25x more images, you can afford to cherry-pick.
| cube2222 wrote:
| It would be interesting to have benchmarks that take this
| into account (maybe they already do or I'm misunderstanding
| how those benchmarks work). I.e. when comparing quality
| between two different models of vastly different performance,
| you could be doing best-of-n in the faster model.
| Vt71fcAqt7 wrote:
| That sounds like it could be an intiresting metric. Worth
| noting that there is a difference between an algorithmic
| "best of n" selection (via eg. an FID score) vs. manual
| cherry picking which takes more factors into account such
| as user preference and also takes time to evaluate, which
| is what GP was suggesting.
| cube2222 wrote:
| Yeah I'd likely just pick the best scoring one (that is,
| the pick is made by the evaluation tool, not the model) -
| to simulate "whatever the receiver deemed best for what
| they wanted".
| Lerc wrote:
| That transfers computer time to user time. It's great when
| you want variations, less so when you want precision and
| consistency. Picking the best image tires the brain quite
| quickly, you have to take into account the at a glance
| quality without it overriding the detail quality.
|
| I'd be curious to see how a vision model would go if it were
| finetuned to select the best image match to a given criteria.
|
| It's possible that you could do O1 style training to build a
| final stage auto-cherrypicker.
| Lerc wrote:
| > _This looks like quite a huge breakthrough, unless I 'm
| missing something?_
|
| Looking at their methodology, it seems like it's more of an
| accumulation of existing good ideas into the one model.
|
| If it performs as well as they say, perhaps you can say the
| breakthrough is discovering just how much can be gained by
| combining recent advances.
|
| It's sitting on just the edge of sounding too good to be true
| to me. I will certainly be pleased if it holds up to scrutiny.
| lpasselin wrote:
| This comes from the same group as the EfficientViT model. A few
| months ago, their EfficientViT model was the only modern and
| small ViT style model I could find that had raw pytorch code
| available. No dependencies to the shitty framework and libraries
| that other ViT are using.
| henning wrote:
| Trained on stolen copyrighted work? Or fairly licensed? Not that
| AI bros give a shit about the law or treating people fairly.
| david-gpu wrote:
| Do you believe that human artists should pay license fees for
| all the art that they have ever seen, studied or drawn
| inspiration from? Whether graphic artists, writers or what have
| you.
| kadoban wrote:
| Human artists get in copyright trouble if the spam out a copy
| of something they studied and sell it. The businesses using
| AI artists do not seem to.
| ClassyJacket wrote:
| Image generation models don't do that either
| david-gpu wrote:
| Artists who think that their copyright has been infringed
| upon are free to sue, just as they do when the alleged
| plagiarist is a human. I fail to see the difference.
| ClassyJacket wrote:
| This argument is only fair if you also think human artists
| should be banned, from birth, from ever looking at any other
| art. After all that would be training on stolen copyrighted
| work.
___________________________________________________________________
(page generated 2024-10-16 23:00 UTC)