[HN Gopher] Muse: Text-to-Image Generation via Masked Generative...
       ___________________________________________________________________
        
       Muse: Text-to-Image Generation via Masked Generative Transformers
        
       Author : jasondavies
       Score  : 64 points
       Date   : 2023-01-15 10:51 UTC (2 days ago)
        
 (HTM) web link (muse-model.github.io)
 (TXT) w3m dump (muse-model.github.io)
        
       | mikemoka wrote:
       | Here is an available implementation:
       | 
       | https://github.com/lucidrains/muse-maskgit-pytorch
        
         | rafaelero wrote:
         | Hopefully stability.ai will train this model and release it
         | open-source. It's much easier to train it than Stable Diffusion
         | after all.
        
       | andybak wrote:
       | Sigh. At this point - if I can't try it out, then I don't really
       | care. It's just a tease.
        
       | pr337h4m wrote:
       | Would stuff like DreamBooth and textual inversion be usable with
       | transformer models like this one?
       | 
       | https://dreambooth.github.io/ https://textual-
       | inversion.github.io/
        
       | CuriouslyC wrote:
       | Fidelity on the output isn't great, but the coherence (assuming
       | the examples weren't massively cherry-picked) seems very good.
       | Given the number of parameters this should be able to run on end-
       | user machines, and in theory this could be fine tuned to produce
       | better looking output than stable diffusion/etc.
       | 
       | What this model does more than anything else is demonstrate we're
       | still in the early stages of generative models, and we can expect
       | a lot of progress from architectural improvements over the next
       | decade (in addition to the progress in compute and data that
       | we're already counting on).
        
       | Garlef wrote:
       | It'd be interesting to see some results where the training set
       | has higher artistic quality (and how this model influences the
       | "house style"). The output does not look great when compared to
       | what other (trained) models deliver.
       | 
       | But the promise of a big efficieny gain will be an incentive for
       | companies like midjourney to give it a go with their data.
        
       | seydor wrote:
       | More amazement . I wonder where this field will end up. Cute
       | animal and nature images are nice but have limited real-life use
       | (i mean, we have to accept that visual media ends after everyone
       | can be an artist). I wonder when we 'll start interfacing
       | language models with robotics to do some real-life work
        
         | activatedgeek wrote:
         | We are starting to see this kind of work already, e.g.
         | https://code-as-policies.github.io
        
       | kleiba wrote:
       | Please stop teasing and post the link to your free trial web
       | interface. Please?
        
         | azinman2 wrote:
         | Most google ML work is just papers; correct me if I'm wrong.
         | Some models have made their way to hugging face like T5 but I
         | don't think any have a web interface.
        
       ___________________________________________________________________
       (page generated 2023-01-17 23:02 UTC)