[HN Gopher] Meissonic, High-Resolution Text-to-Image Synthesis o...
___________________________________________________________________
Meissonic, High-Resolution Text-to-Image Synthesis on consumer
graphics cards
Author : jinqueeny
Score : 36 points
Date : 2024-10-14 17:33 UTC (5 hours ago)
(HTM) web link (arxiv.org)
(TXT) w3m dump (arxiv.org)
| fngjdflmdflg wrote:
| >Meissonic, with just 1B parameters, offers comparable or
| superior 1024x1024 high-resolution, aesthetically pleasing images
| while being able to run on consumer-grade GPUs with only 8GB VRAM
| without the need for any additional model optimizations.
| Moreover, Meissonic effortlessly generates images with solid-
| color backgrounds, a feature that usually demands model fine-
| tuning or noise offset adjustments in diffusion models.
|
| This looks really cool. Also nice to see another architecture
| being used for image generation besides diffusion. It seems like
| every NLP problem can be solved with transformers now: text
| generation/understanding, image generation/understanding,
| translation, OCR. Perhaps llama 4/5 will have image generation as
| well. eidt: llama 3.2 already has image editing, they probably
| just don't want to release an image generator for other reasons.
| jensenbox wrote:
| The images in the PDF are amazing.
___________________________________________________________________
(page generated 2024-10-14 23:01 UTC)