[HN Gopher] Infinite Images and the Latent Camera
___________________________________________________________________
Infinite Images and the Latent Camera
Author : bryanrasmussen
Score : 37 points
Date : 2022-05-06 07:04 UTC (15 hours ago)
(HTM) web link (mirror.xyz)
(TXT) w3m dump (mirror.xyz)
| amelius wrote:
| I typed: "a painting of a bridge, giving me the
| satisfaction of having painted it myself"
|
| No such thing appeared.
| 0xrisk wrote:
| author of the piece here! I think these tools are best
| understood as a step in an artist workflow that people will use
| to varying degree. Many painters use photography or sampled
| images to audition compositions, for example. Musicians use
| samples to spark an idea or set a palette for them to create
| within. The satisfaction of creating on a blank canvas won't
| change, but most artists use tools to help them realize what
| they want to create, and these are incredibly powerful to that
| end.
|
| I think that typing to create imagery is cool, but also agree
| that the ability to instantly create convincing images (with
| those images being considered the final work) is not the most
| exciting/rewarding part, at least when those images are
| isolated as singular works to consider(like in painting) and
| not part of a greater narrative. This perspective was part of
| our motivation to frame the subject differently!
| nh23423fefe wrote:
| Seems trite. Like why not, "a painting that I could sell for
| millions"
|
| Honestly, you could come up with some painting you would be
| proud of. But you'd actually have to engage with the tool and
| make something.
|
| The verb paint is open to interpretation.
| TrevorJ wrote:
| What's interesting here is that as a content creator, it really
| does feel as if this externalizes the 'latent space' that exists
| in my own head. 99% of my time/skill is used in the pursuit of
| getting the vision of something that exists only in my mind, into
| a state where it can be shared with others in the physical world.
|
| Collaborating with an external latent space and teasing certain
| viewpoints out of it is an interesting dance, and something that
| I really look forward to participating in as soon as I can get
| access to the tools.
|
| I'm very curious about where the capacity for iterative
| refinement is/will be in the future as well. "Ok, that looks
| great, now can you try it with a bit more green?" etc.
| axg11 wrote:
| Thought provoking write up. I think with GPT-3/Co-pilot and
| DALL-E 2 we're seeing the advent of a new computing interface and
| information paradigm. We're very early.
|
| The old/existing way is: humans produce content (writing, images,
| videos, music) using software and then use the internet to host
| and distribute the content. As the volume of information
| increased, aggregators of content became more important. Google
| helps you find anything, but primarily text and images, YouTube
| indexes and recommends medium length videos, etc.
|
| A very common problem pattern is that you find multiple sources
| that are _nearly_ what you want, but not quite. For example you
| combine two stock photos in photoshop (with some editing) to
| produce a final image.
|
| The future way is: large foundational models that act as the
| aggregators of information _and_ the editing tools. Co-pilot is
| the best current example we have. Once the performance of Co-
| pilot becomes good enough, I will never have to leave VSCode +
| Co-pilot, replacing the previous pattern of Google +
| documentation + Stackoverflow. The indexers and aggregators will
| be made obsolete. This is exacerbated by the fact that Co-pilot
| is able to do more than just aggregating the correct
| Stackoverflow examples. Co-pilot can do the customization and
| translation of the code examples into your specific use case.
|
| These large/foundational models are nowhere near perfect yet.
| Once the performance is good enough to make existing workflows
| easier, we'll see rapid adoption in every domain. Recommendation
| algorithms will look entirely different in 10 years time. The
| 2030 version of YouTube won't be serving users the closest match
| video from its library, it will be generating custom edits on-
| the-fly tailored to the user.
| 0xrisk wrote:
| totally agree, part of why we wrote this was attempting to
| frame DALL-E as an early step towards something much bigger,
| rather than overly focussing on its current application!
| bckr wrote:
| Unfortunately, I can still never look at these blog posts for
| more than a few seconds without becoming disturbed. Any time an
| animal (human or otherwise) is included in the image, it's
| terrifying--especially the hands/hooves/paws. Granted, I'm more
| phobic of monstrosities than the average person, but still. The
| technology has some distance to cross before I can call its
| output beautiful.
___________________________________________________________________
(page generated 2022-05-06 23:02 UTC)