[HN Gopher] Building an agentic image generator that improves it...
___________________________________________________________________
Building an agentic image generator that improves itself
Hey HN! We recently graduated from YC, and have been building
customer personas for large e-commerce companies. We recently
expanded into the image generation space, and have been working on
research about how to automatically improve the quality of
generated images.
Author : palashshah
Score : 51 points
Date : 2025-05-21 13:12 UTC (9 hours ago)
(HTM) web link (simulate.trybezel.com)
(TXT) w3m dump (simulate.trybezel.com)
| average_r_user wrote:
| Quite interesting, do you have some documentation of your
| platform and capabilities? Your landing page is quite synthetic
| palashshah wrote:
| hey! we're working with an initial set of customers, and plan
| to launch full capabilities soon. stay tuned :)
| ramesh31 wrote:
| This is a wonderful writeup of building a simple agentic system
| in general. What OP describes is more or less the bare minimum
| you should be doing at this point to get good (consistent)
| results from an LLM; single-shot prompting is a thing of the
| past.
| palashshah wrote:
| appreciate the compliment! yep, it's definitely necessary and
| is the bare minimum for building image generation systems in
| production.
| shmoogy wrote:
| I'm surprised you landed on using o3 as the judge - we found it
| way too expensive. I use llm as a judge for generating color
| variations of products, definitely hoping for some improvements -
| it can be brutal to get non hallucinated features along with
| proper final rendering.
| omneity wrote:
| Have you tried open weights vision models such as Qwen VL,
| MiniCPM, PaliGemma...?
|
| I'm also curious how usable are simpler vision models such as
| Florence in case you explored this direction.
| palashshah wrote:
| we're currently in the process of doing this. i think
| something that could potentially work is to iterate upon the
| initial image composition / structure using cheaper models,
| and then upscale at the end. this way you're saving on that
| iteration cost, but eventually land on a higher-scale image.
| elif wrote:
| This is great and provides a good starting point for any similar
| efforts.
|
| However I think the temptation to lean all tasks on AI is perhaps
| a little naive if not lazy.
|
| For mask generation, there is really not much reason to use AI.
| In this example, simple stochastic blob detection, a trivial
| function you could get from openCV or ask a college sophomore to
| write would generate much better quality masks.
| palashshah wrote:
| totally agreed here. i think my goal primarily with the mask
| generation was to test out how effective openai's capabilities
| were.
|
| we're currently working on pipelines that limit the the
| involvement of AI to various tasks. for example, when
| generating an ad there's usually logo, some banner text, and
| background image.
|
| we can use gpt-image-1 to generate the background image,
| another LLM to identify the coordinates of where we place the
| logo, and just add the logo onto the image. this is just one
| example!
| jackphilson wrote:
| Why do you agree? I think we should outsource as much as we
| can to abstraction. We've been doing it forever.
| mentalgear wrote:
| Again another example of "the unreasonable effectiveness of LLMs
| in a loop". At with time, the tasks for loop become bigger and
| more complex, until we find ourselves "outlooped" at least job
| wise.
| ramoz wrote:
| Nice retrospective but I guess this process is no longer needed
| as model's get better; esp as they start enabling features like
| consistent subjects. Seems like a lot of overhead to correct text
| for inspirational images, but I can imagine you need to always
| present some form of _quality_ to your clients.
|
| Feel like control nets and some minimal photoshop work would've
| been better.
| palashshah wrote:
| totally. it got to a point where most of the text generated in
| our images was incorrect, and so it wasn't a great look showing
| that to our clients.
|
| we're actually working on some form of what you described where
| we take images generated from LLMs + add consistent logos
| discretely rather than generatively.
| abshkbh wrote:
| Palash this is a great post, I learnt a lot as an image gen noob!
| Keep writing more :)
| palashshah wrote:
| this is incredible to hear! i plan to keep writing on a weekly
| basis, and will be posting them on twitter.
| t_mann wrote:
| I was kind of hoping this would be in the 'Dreambooth mold' of
| finetuning open weights models. I have used that with some
| success some ~2 years ago, does anyone know what improvements
| there have been in that direction since Dreambooth?
| zahlman wrote:
| It's frankly amazing to me that "ask another LLM to evaluate the
| image" actually produces useful feedback that results in actual
| improvement from the first LLM.
|
| But then, I guess it's not much different of an idea from the
| earlier use of GANs, or of telling LLMs to "stop hallucinating",
| etc.
| palashshah wrote:
| totally. the way i think about it (purely based on intuition)
| is that asking an LLM to do understanding + image generation is
| too complex for it to be effective. if we separate out the
| tasks into discrete steps, the evaluation becomes better, and
| the generation simply becomes instruction following.
___________________________________________________________________
(page generated 2025-05-21 23:01 UTC)