[HN Gopher] Infinite Images and the Latent Camera
       ___________________________________________________________________
        
       Infinite Images and the Latent Camera
        
       Author : bryanrasmussen
       Score  : 37 points
       Date   : 2022-05-06 07:04 UTC (15 hours ago)
        
 (HTM) web link (mirror.xyz)
 (TXT) w3m dump (mirror.xyz)
        
       | amelius wrote:
       | I typed:                   "a painting of a bridge, giving me the
       | satisfaction of having painted it myself"
       | 
       | No such thing appeared.
        
         | 0xrisk wrote:
         | author of the piece here! I think these tools are best
         | understood as a step in an artist workflow that people will use
         | to varying degree. Many painters use photography or sampled
         | images to audition compositions, for example. Musicians use
         | samples to spark an idea or set a palette for them to create
         | within. The satisfaction of creating on a blank canvas won't
         | change, but most artists use tools to help them realize what
         | they want to create, and these are incredibly powerful to that
         | end.
         | 
         | I think that typing to create imagery is cool, but also agree
         | that the ability to instantly create convincing images (with
         | those images being considered the final work) is not the most
         | exciting/rewarding part, at least when those images are
         | isolated as singular works to consider(like in painting) and
         | not part of a greater narrative. This perspective was part of
         | our motivation to frame the subject differently!
        
         | nh23423fefe wrote:
         | Seems trite. Like why not, "a painting that I could sell for
         | millions"
         | 
         | Honestly, you could come up with some painting you would be
         | proud of. But you'd actually have to engage with the tool and
         | make something.
         | 
         | The verb paint is open to interpretation.
        
       | TrevorJ wrote:
       | What's interesting here is that as a content creator, it really
       | does feel as if this externalizes the 'latent space' that exists
       | in my own head. 99% of my time/skill is used in the pursuit of
       | getting the vision of something that exists only in my mind, into
       | a state where it can be shared with others in the physical world.
       | 
       | Collaborating with an external latent space and teasing certain
       | viewpoints out of it is an interesting dance, and something that
       | I really look forward to participating in as soon as I can get
       | access to the tools.
       | 
       | I'm very curious about where the capacity for iterative
       | refinement is/will be in the future as well. "Ok, that looks
       | great, now can you try it with a bit more green?" etc.
        
       | axg11 wrote:
       | Thought provoking write up. I think with GPT-3/Co-pilot and
       | DALL-E 2 we're seeing the advent of a new computing interface and
       | information paradigm. We're very early.
       | 
       | The old/existing way is: humans produce content (writing, images,
       | videos, music) using software and then use the internet to host
       | and distribute the content. As the volume of information
       | increased, aggregators of content became more important. Google
       | helps you find anything, but primarily text and images, YouTube
       | indexes and recommends medium length videos, etc.
       | 
       | A very common problem pattern is that you find multiple sources
       | that are _nearly_ what you want, but not quite. For example you
       | combine two stock photos in photoshop (with some editing) to
       | produce a final image.
       | 
       | The future way is: large foundational models that act as the
       | aggregators of information _and_ the editing tools. Co-pilot is
       | the best current example we have. Once the performance of Co-
       | pilot becomes good enough, I will never have to leave VSCode +
       | Co-pilot, replacing the previous pattern of Google +
       | documentation + Stackoverflow. The indexers and aggregators will
       | be made obsolete. This is exacerbated by the fact that Co-pilot
       | is able to do more than just aggregating the correct
       | Stackoverflow examples. Co-pilot can do the customization and
       | translation of the code examples into your specific use case.
       | 
       | These large/foundational models are nowhere near perfect yet.
       | Once the performance is good enough to make existing workflows
       | easier, we'll see rapid adoption in every domain. Recommendation
       | algorithms will look entirely different in 10 years time. The
       | 2030 version of YouTube won't be serving users the closest match
       | video from its library, it will be generating custom edits on-
       | the-fly tailored to the user.
        
         | 0xrisk wrote:
         | totally agree, part of why we wrote this was attempting to
         | frame DALL-E as an early step towards something much bigger,
         | rather than overly focussing on its current application!
        
       | bckr wrote:
       | Unfortunately, I can still never look at these blog posts for
       | more than a few seconds without becoming disturbed. Any time an
       | animal (human or otherwise) is included in the image, it's
       | terrifying--especially the hands/hooves/paws. Granted, I'm more
       | phobic of monstrosities than the average person, but still. The
       | technology has some distance to cross before I can call its
       | output beautiful.
        
       ___________________________________________________________________
       (page generated 2022-05-06 23:02 UTC)