[HN Gopher] Efficient high-resolution image synthesis with linea...
       ___________________________________________________________________
        
       Efficient high-resolution image synthesis with linear diffusion
       transformer
        
       Author : Vt71fcAqt7
       Score  : 129 points
       Date   : 2024-10-16 14:56 UTC (8 hours ago)
        
 (HTM) web link (nvlabs.github.io)
 (TXT) w3m dump (nvlabs.github.io)
        
       | smusamashah wrote:
       | > (e.g. Flux-12B), being 20 times smaller and 100+ times faster
       | in measured throughput. Moreover, Sana-0.6B can be deployed on a
       | 16GB laptop GPU, taking less than 1 second to generate a 1024 x
       | 1024 resolution image.
        
       | echelon wrote:
       | Image models are going to be widely available. They'll probably
       | be a dime a dozen soon. It's great that an increasing number of
       | models are going open, because these are the ecosystems that will
       | grow.
       | 
       | 3D models (sculpts, texture, retopo, etc.) are following a
       | similar trend and trajectory.
       | 
       | Open video models are lagging behind by several years. While
       | CogVideo and Pyramid are promising, video models are petabyte
       | scale and so much more costly to build and train.
       | 
       | I'm hoping video becomes free and cheap, but it's looking like we
       | might be waiting a while.
       | 
       | Major kudos to all of the teams building and training open source
       | models!
        
       | cube2222 wrote:
       | This looks like quite a huge breakthrough, unless I'm missing
       | something?
       | 
       | ~25x faster performance than Flux-dev, while offering comparable
       | quality in benchmarks. And visually the examples (surely cherry-
       | picked, but still) look great!
       | 
       | Especially since with GenAI the best way to get good results is
       | to just generate a large amount of them and pick the best (imo).
       | Performance like this will make that much easier/faster/cheaper.
       | 
       | Code is unfortunately "(Coming soon)" for now. Can't wait to play
       | with it!
        
         | liuliu wrote:
         | If you read closer to the benchmark, it seems to be slightly
         | worse than FLUX [dev] on prompt adherence and quality. However,
         | the best is to evaluate the result oneself, and the track-
         | record of PixArt Sigma (from the same author?) is pretty good!
        
         | Archit3ch wrote:
         | If you generate 25x more images, you can afford to cherry-pick.
        
           | cube2222 wrote:
           | It would be interesting to have benchmarks that take this
           | into account (maybe they already do or I'm misunderstanding
           | how those benchmarks work). I.e. when comparing quality
           | between two different models of vastly different performance,
           | you could be doing best-of-n in the faster model.
        
             | Vt71fcAqt7 wrote:
             | That sounds like it could be an intiresting metric. Worth
             | noting that there is a difference between an algorithmic
             | "best of n" selection (via eg. an FID score) vs. manual
             | cherry picking which takes more factors into account such
             | as user preference and also takes time to evaluate, which
             | is what GP was suggesting.
        
               | cube2222 wrote:
               | Yeah I'd likely just pick the best scoring one (that is,
               | the pick is made by the evaluation tool, not the model) -
               | to simulate "whatever the receiver deemed best for what
               | they wanted".
        
           | Lerc wrote:
           | That transfers computer time to user time. It's great when
           | you want variations, less so when you want precision and
           | consistency. Picking the best image tires the brain quite
           | quickly, you have to take into account the at a glance
           | quality without it overriding the detail quality.
           | 
           | I'd be curious to see how a vision model would go if it were
           | finetuned to select the best image match to a given criteria.
           | 
           | It's possible that you could do O1 style training to build a
           | final stage auto-cherrypicker.
        
         | Lerc wrote:
         | > _This looks like quite a huge breakthrough, unless I 'm
         | missing something?_
         | 
         | Looking at their methodology, it seems like it's more of an
         | accumulation of existing good ideas into the one model.
         | 
         | If it performs as well as they say, perhaps you can say the
         | breakthrough is discovering just how much can be gained by
         | combining recent advances.
         | 
         | It's sitting on just the edge of sounding too good to be true
         | to me. I will certainly be pleased if it holds up to scrutiny.
        
       | lpasselin wrote:
       | This comes from the same group as the EfficientViT model. A few
       | months ago, their EfficientViT model was the only modern and
       | small ViT style model I could find that had raw pytorch code
       | available. No dependencies to the shitty framework and libraries
       | that other ViT are using.
        
       | henning wrote:
       | Trained on stolen copyrighted work? Or fairly licensed? Not that
       | AI bros give a shit about the law or treating people fairly.
        
         | david-gpu wrote:
         | Do you believe that human artists should pay license fees for
         | all the art that they have ever seen, studied or drawn
         | inspiration from? Whether graphic artists, writers or what have
         | you.
        
           | kadoban wrote:
           | Human artists get in copyright trouble if the spam out a copy
           | of something they studied and sell it. The businesses using
           | AI artists do not seem to.
        
             | ClassyJacket wrote:
             | Image generation models don't do that either
        
             | david-gpu wrote:
             | Artists who think that their copyright has been infringed
             | upon are free to sue, just as they do when the alleged
             | plagiarist is a human. I fail to see the difference.
        
         | ClassyJacket wrote:
         | This argument is only fair if you also think human artists
         | should be banned, from birth, from ever looking at any other
         | art. After all that would be training on stolen copyrighted
         | work.
        
       ___________________________________________________________________
       (page generated 2024-10-16 23:00 UTC)