[HN Gopher] LAION, a high school teacher's free image database, ...
       ___________________________________________________________________
        
       LAION, a high school teacher's free image database, powers AI
       unicorns
        
       Author : RafelMri
       Score  : 332 points
       Date   : 2023-04-24 09:48 UTC (1 days ago)
        
 (HTM) web link (www.bloomberg.com)
 (TXT) w3m dump (www.bloomberg.com)
        
       | chwa982 wrote:
       | [flagged]
        
         | sailorganymede wrote:
         | Sounds like the website you are currently on does the trick.
        
       | dvt wrote:
       | > Within a matter of months, lawsuits had been filed against
       | generative AI companies Stability AI and Midjourney for copyright
       | infringement [...] But these aren't Schuhmann's concerns. He just
       | wants to set the data free.
       | 
       | Maybe it _should be_ one of his concerns? I 'm not sure how _he_
       | would feel about someone stealing his work and commercializing
       | it.
        
         | vasco wrote:
         | He is not commercializing it as far as I know.
        
           | dvt wrote:
           | "Hey man, I just make the meth, I don't sell it."
           | 
           | The point I'm trying to get across is that, at least to me,
           | it seems incredibly irresponsible and borderline unethical.
        
             | jonhohle wrote:
             | Depending on views about drugs and personal liberty, that
             | statement may be neither irresponsible nor unethical.
        
             | moomoo3000 wrote:
             | These are all images that were posted to the internet
             | right? That place where anyone can download things freely?
        
               | dvt wrote:
               | > That place where anyone can download things freely?
               | 
               | This is _incredibly_ reductive. Property rights exist on
               | the internet, too. Like how someone can own a domain name
               | or the copyright to a photograph or short story.
        
             | vasco wrote:
             | So say what you think is unethical, don't stretch the truth
             | just because you think you are right. He is not
             | commercializing anything and kept his job as a high-school
             | teacher.
        
         | DoreenMichele wrote:
         | Anything can be misused. The person who intentionally and
         | willfully misuses it is the problem, not the existence of the
         | thing per se.
        
         | layer8 wrote:
         | LAION doesn't publish the images, only the metadata with links
         | to the original image locations. Arguably, the alt text
         | contained in the metadata is still subject to copyright.
        
         | whywhywhywhy wrote:
         | > Maybe it should be one of his concerns? I'm not sure how he
         | would feel about someone stealing his work and commercializing
         | it
         | 
         | Covered in the article, his work is being commercialized and
         | profited from:
         | 
         | > Schuhmann hasn't profited from LAION and says he isn't
         | interested in doing so. "I'm still a high school teacher. I
         | have rejected job offers from all different kinds of companies
         | because I wanted this to stay independent,"
        
       | netuffresche wrote:
       | If someone is interested, I interviewed him recently. The
       | interview is in german, though.
       | 
       | https://entwickler.de/machine-learning/laion-open-source-ai
        
         | w-m wrote:
         | Thank you! The Bloomberg article piqued my interest, but it is
         | actually very light on the details of the LAION project - it
         | quickly moved into the same old license and bias discussions
         | and left me wanting. Your interview is much more interesting,
         | it's great to learn a more about the person behind it and his
         | motivations.
        
           | netuffresche wrote:
           | Thanks! Really happy to hear that. He's a very fascinating
           | guy and LAION is an awesome project, eventhough the
           | discussion about copyright and data sets is important.
        
       | lisasays wrote:
       | _For his part, Schuhmann hasn't profited from LAION and says he
       | isn't interested in doing so. "I'm still a high school teacher. I
       | have rejected job offers from all different kinds of companies
       | because I wanted this to stay independent," he said._
       | 
       | Brave is this man for going against the current Zeitgeist (and
       | the obvious financial temptations).
        
       | ftxbro wrote:
       | Something about this reminds me a lot of Michael Hart and Project
       | Gutenberg.
        
       | 1letterunixname wrote:
       | One could also write an image crawler to execute from AWS.
       | 
       | The key is finding quality source material.
       | 
       | It may become necessary to:
       | 
       | 1. Fly drones around to gather data in the world at present time
       | 
       | 2. Direct humans to take high-quality photos of particular
       | subjects. Fiverr for photography
        
         | ramraj07 wrote:
         | You mean like what google did with street view and convincing
         | everyone to upload their photos on to maps?
        
           | lmpdev wrote:
           | Like that but properly organised and with higher quality
           | imagery
           | 
           | I've been wanting this for quite sometime
        
           | Tepix wrote:
           | Or like Mapillary (bought by Facebook, but with a liberal
           | license still)
        
         | bradgessler wrote:
         | This will be the new captcha challenge--"take a picture of 14
         | crosswalks"
         | 
         | It will send you a link to an app that only runs on an iPhone
         | so it can verify you actually took the pics from the phone.
        
           | moffkalast wrote:
           | _takes pictures of a screen with the phone_
        
           | quickthrower2 wrote:
           | That would not be a11y friendly!
        
             | logi2mus wrote:
             | Why exactly, they should know how to use crosswalks by that
             | age.
        
               | newswasboring wrote:
               | a11y means accessibility. Apparently this is a common
               | enough jargon in that part of the industry.
        
               | Maken wrote:
               | Is the jargon going too far?
        
               | wongarsu wrote:
               | I the j4n going too far? English always had a dislike for
               | long words, so turning long and s11d words into shorter
               | forms is the next logical step in l6e evolution.
               | Certainly beats having to spell w12e sauce in its
               | entirety /s
        
               | quickthrower2 wrote:
               | At least w12e is easier to pronounce:-)
        
               | Maken wrote:
               | At some point LLMs will have a easier time understanding
               | English than actual human beings.
        
               | newswasboring wrote:
               | Expecting people to know this is going too far. I mean
               | usually the acronym gives a hint, this is just straight
               | up medieval druid vibes. "Oh the field which begins with
               | a followed by our savior's pious 11 letters and then a y"
        
               | [deleted]
        
             | OJFord wrote:
             | Neither are current captchas: they're ina11 to non-US
             | residents (/former residents).
             | 
             | (We don't have 'crosswalks'. 'Fire hydrants' don't look
             | like that. 'Parking meters' don't look like that (or any
             | one standardised thing really). Traffic lights don't look
             | like that and how much of them am I meant to click anyway.
             | School buses don't look like that. Trucks don't look like
             | that. I'm sure there are more.)
        
               | wongarsu wrote:
               | And often it's even more subtle. Apparently Americans
               | don't consider pedestrian lights part of the traffic
               | lights. And apparently they only mark the actual light
               | assembly, not the structure holding them up.
               | 
               | Lot's of things where you might think you know what
               | you're doing, but are actually doing it wrong because you
               | lack cultural context..
        
               | the_lonely_road wrote:
               | That doesn't seem like cultural context at all. That just
               | seems like the arbitrary decision of the person who
               | created the captchas. I am, presumably, from the same
               | culture as them and would have agreed that pole was part
               | of the traffic light.
        
               | quickthrower2 wrote:
               | I know most of HN lives downtown but not everyone is near
               | a crosswalk. And then there are blind people, etc.
        
               | OJFord wrote:
               | I assume non-downtown US residents are still
               | significantly more aware of the term, what it is, and
               | what it looks like.
               | 
               | I'd heard it in films before captcha, but had to guess
               | it's the zebra crossing looking (the meaning is different
               | though, I understand) things. Now it's just something I'm
               | familiar with because of captcha, even more so 'parking
               | meters'. Which is just weird, why do I have to learn
               | something mundane about a specific country in order to
               | prove I'm human and allowed to use a global website?
        
         | TACIXAT wrote:
         | I've built a site to pay people for images and annotations. [1]
         | I'm trying to onboard my first paid users right now. The plan
         | is to build out a high quality 50k image license plate
         | recognition dataset as a proof of concept.
         | 
         | Right now we own all of the datasets on the site and the idea
         | is to license them out to companies while making them available
         | to researchers under a non-commercial license. The market might
         | take it a different direction to be more of a marketplace or
         | Github style hosting. Email in bio if anyone wants to chat
         | about this.
         | 
         | Also, if anyone wants to get paid 10 cents an image to take
         | pictures of North American license plates, get in touch. Need
         | about 1000 from each state. It's probably below most people's
         | pay grade on here but there is a whole reverse bidding system,
         | so you can always bid higher than 10 cents. Some user studies
         | with a shared screen would be super helpful as well.
         | 
         | 1. https://mekabytes.com
        
           | kennyloginz wrote:
           | Wow, interesting project. Not sure how many people you can
           | entice into providing 10k photos of license plates, firearms,
           | and children in pools. I kid, but are you building an alert
           | system for a superhero?
        
             | TACIXAT wrote:
             | Ha! It is a bit public safety focused right now. The first
             | two I will really build out (read: pay people to help with)
             | are license plate recognition and vehicle make/model/year
             | identification. I think those have a decent market.
             | 
             | The firearms one is tricky and I'm not sure it will ever be
             | licensed commercially since we don't own the footage that
             | the images are taken from. Valuable dataset in terms of
             | what it could provide though, like an early warning system
             | for active shooters.
             | 
             | Some day I'll have enough free time and just make a
             | billiards dataset to help me find shots.
        
               | mikecoles wrote:
               | Unattended security gates at a local campground are being
               | used to gather data. To get access to the campground, you
               | have to provide your plate number, vehicle make & model,
               | and vehicle color. The gate only checks for the plate
               | number, but the images captured can be used to build out
               | a vehicle make/model ML model.
               | 
               | Maybe partner with such a system to gather your data?
               | 
               | I have several thousand images of vehicles on the roadway
               | and parking lots that I used to understand the backend ML
               | of an ALPR system. Selecting cameras was the most
               | difficult part of the project. I did not attempt to
               | determine the make/model. My patience for labeling had
               | worn too thin.
               | 
               | Overall, as invasive as ALPR system are, the cat has long
               | left the bag. I doubt the cat will ever return.
        
               | jacquesm wrote:
               | Have a good look at privacy laws in the states/countries
               | of your users, in some things like this may be against
               | the law or subject to restrictions. See:
               | 
               | https://www.ncsl.org/technology-and-
               | communication/automated-...
        
               | TACIXAT wrote:
               | Thank you! That's honestly very helpful.
        
               | jacquesm wrote:
               | You're welcome & much good luck with your project!
        
           | angrais wrote:
           | What is the cost for each dataset for commercial uses?
        
             | TACIXAT wrote:
             | We haven't figured out commercial pricing yet. It may be a
             | scale based on company revenue. Targeting bigger businesses
             | though, it will be less than it would cost to develop the
             | model independently, but likely trying to target low 5
             | figures.
             | 
             | The datasets are also available and a non-commercial
             | Creative Commons license. You can pay $5 for a download and
             | rehost it elsewhere, or just wait for me to upload it to
             | Kaggle or Hugging Face
        
           | dahwolf wrote:
           | "Also, if anyone wants to get paid 10 cents an image to take
           | pictures of North American license plates, get in touch."
           | 
           | Your blatant disregard for privacy is shocking, but perhaps
           | unsurprising in the field. I guess you also didn't through
           | the enormous risks for the photographer.
        
           | virgildotcodes wrote:
           | Hey there, just checked out the site and signed up. I'm not
           | seeing anything about how to get paid for uploads. Can you
           | provide some direction here? Thanks!
        
             | TACIXAT wrote:
             | It is a little hidden still. On your profile page there
             | should be a sign up link that will take you to Stripe. Feel
             | free to email me. I will post two solicitations tomorrow,
             | one for annotation work and one for plate photographs.
             | 
             | It is restricted to US only right now (Stripe setting) but
             | I will change that if there is non-US interest and the
             | country is supported.
             | 
             | I've only run through the whole process once so hopefully
             | there are no hiccups. Thank you!
        
       | kleiba wrote:
       | _The group used raw HTML code collected by the California
       | nonprofit Common Crawl to locate images around the web and
       | associate them with descriptive text. It does not use any manual
       | or human curation._
       | 
       | Does the Common Crawl data already take care of the copyright
       | issue? Or else how does the LAION crawler deal with that problem?
       | 
       | I mean, it's not too hard to write an image crawler. Also,
       | scaling it up it a bit of a challenge, but it's a technical one.
       | But the real difficulty is how to deal with all the legal strings
       | attached...
        
         | nprateem wrote:
         | It's not really a difficulty though is it, for VC funded
         | unicorns who ask forgiveness and move fast. When the law
         | catches up, if ever, they'll pay a slap on the wrist while
         | sitting on their superyachts
        
         | angrais wrote:
         | Less than 0.01% of the 5.8 billion images have an associated
         | license. This license was taken from the webpage so cannot be
         | trusted.
        
         | layer8 wrote:
         | LAION doesn't publish the images, only the metadata, which
         | includes the original image URLs. Anyone who wants to make use
         | of the image set has to re-download the images from the
         | original sources [0], and is then liable for actually using the
         | images. Arguably, the alt text contained in the metadata is
         | still subject to copyright though.
         | 
         | [0] see also https://news.ycombinator.com/item?id=35681085
        
           | XorNot wrote:
           | This is the most commonly misunderstood characteristic of
           | LAION, and I find that alarming because it feels like it just
           | plain means a lot of people don't understand how the internet
           | actually works.
           | 
           | Which is not surprising but guys, it's _everywhere_ it 's
           | also not optional to understand this.
        
             | Silverback_VII wrote:
             | Yet, what LAION is doing appears to be somewhat obvious. We
             | are a non-profit and not publishing the images seems like
             | hiding behind technicalities.
        
               | XorNot wrote:
               | What do you think they're doing?
        
             | dahwolf wrote:
             | We understand it just fine. Us internet people do free
             | labor, then others take it and monetize it.
        
               | XorNot wrote:
               | Seems like you both don't understand LAION and don't
               | understand hyperlinking.
        
           | slow_typist wrote:
           | Arguably, a short alt text does not meet the threshold of
           | originality under most jurisdictions. See
           | https://en.m.wikipedia.org/wiki/Threshold_of_originality
        
         | Borrible wrote:
         | Well, at least one content producer is pissed about it.
         | 
         | Sorry, only in German:
         | 
         | https://www.alltageinesfotoproduzenten.de/2023/04/24/laion-e...
         | 
         | And translated via Google translate:
         | 
         | https://www-alltageinesfotoproduzenten-de.translate.goog/202...
        
           | Silverback_VII wrote:
           | After reading it, it seems that LAION is merely doing the
           | dirty work (going as far as intimidating the creators of the
           | materials they use) so that other companies can utilize
           | copyrighted material without taking too much risk. Not sure
           | if this teacher is a hero...
        
         | slow_typist wrote:
         | According to Christoph Schumann there is legislation in Europe
         | that allows the usage of crawled data for public research
         | institutions and nonprofit organisations. LAION is a registered
         | german nonprofit (gemeinnutziger Verein).
         | 
         | Source: https://entwickler.de/machine-learning/laion-open-
         | source-ai (found in another post).
        
           | gorbachev wrote:
           | Commercial entities are using the data, however. I'm pretty
           | sure the German legislation didn't mean that to be allowed.
        
             | slow_typist wrote:
             | Sure. But that is not in LAIONs responsibility. They
             | provide URIs to images and the corresponding alt texts.
        
       | epups wrote:
       | I wonder how it came to be that such a critical component of a
       | multi billion industry relies on something so amateurish as
       | LAION. This is no offense to the author at all, who organised a
       | gigantic effort which we now see is very valuable. But I would
       | imagine a company like Google could do a much better job in no
       | time, simply due to expertise and resources.
       | 
       | I guess the answer is legal liability.
        
         | dahwolf wrote:
         | "I guess the answer is legal liability."
         | 
         | Indeed. The answer is in the article. He gets offered jobs all
         | the time, but nobody offers to buy the data itself. Clearly
         | because nobody wants to own it, it's all plausible deniability.
        
         | Maken wrote:
         | Forever relevant: https://xkcd.com/2347/
        
         | lisasays wrote:
         | By all indications, it's definitely a solid engineering
         | project. On what basis should we deride it as "amateurish"?
         | 
         | Because it originates in the public sector? As opposed to the
         | private sector (where, as we know, everything is done to the
         | highest possible engineering standards)?
        
           | epups wrote:
           | I know the word sounds demeaning, and I didn't mean to
           | criticise in on technical grounds. I find it an extremely
           | impressive project.
           | 
           | I meant that it does not have the refinement you would expect
           | for such a critical tool. A substantial portion of LAION is
           | composed of duplicates. If you have ever browsed it, you will
           | find that many annotations are quite basic and in some cases
           | incorrect. In ChatGPT's case we know there was a small army
           | of people going through their dataset to filter and refine
           | issues that are presumably similar to those.
        
             | wodenokoto wrote:
             | Who is not to say google et al doesn't have an in house
             | army that re-annotates images with a low score and store it
             | in a parallel dataset?
        
       | 1letterunixname wrote:
       | Text-to-visual LMs are making children's book illustrators go the
       | way of harpsichord tuners and whale oil hunters.
        
         | milofeynman wrote:
         | Is there an example of a generated children's book that
         | maintains style throughout?
        
         | riffraff wrote:
         | This was always possible, I own multiple books for kids which
         | have obviously been made with random clip art and illustrations
         | which are badly paired.
         | 
         | But _proper_ illustration is so much better, I doubt it will go
         | away.
         | 
         | Perhaps a better simile is hand made clothing: you can get mass
         | produced stuff, or you can get small batch, or you can get
         | handmade. There is a market for all of those.
        
           | PeterisP wrote:
           | _Proper_ illustration is so expensive that in most cases
           | people are tolerating it because the alternatives are either
           | very horrible (random clip art) or to make do without
           | illustrations.
           | 
           | If digital tools enable _decent_ clipart that can be at least
           | _somewhat_ paired to the content, that will take over most of
           | the market, because then it 'll be just somewhat worse than
           | proper illustration while being dirt cheap in comparison.
        
           | ukuina wrote:
           | The market for all of those is ridiculously lopsided, though.
           | That doesn't bode well for illustrators, especially the ones
           | that pour love into crafting every drawing and sketch to make
           | it as good as it could be.
        
         | totetsu wrote:
         | If you take the point of view, that children's media is only
         | for kids, so quality doesn't really matter. The best children's
         | book authors and illustrators are ones who see children as the
         | most important audience of all and deserving of the most
         | careful work.
        
           | ftxbro wrote:
           | Maybe the idea is that most children's books are illustrated,
           | compared to other books.
        
           | p-e-w wrote:
           | Except that the better outputs from the better Stable
           | Diffusion models make 99% of professional illustrators look
           | like they're drawing stick figures by comparison.
           | 
           | If you still believe that "AI generated = bad quality", you
           | haven't seen some of the stuff that has popped up lately. SD
           | is still new, so understandably it's taking a while for
           | people to figure out how to use it properly, but some users
           | are clearly getting the hang of it now, and the results are
           | mind-blowing.
        
             | totetsu wrote:
             | I am not talking about visual wow factor quality. I would
             | love link to my instagram of amazing dalle2 and SD
             | generated images, but I like to keep my HN account clean of
             | personal info. [1 like this] Children don't care about
             | that, they consume images as a way of creating and
             | expanding their inner world and their understanding the
             | outer world. A stick figure that captures something real to
             | life in an way a child can recognize and be excited by is
             | much 'quality' image for children's media than a 'award
             | winning picture of a cute robot holding a flower 4k
             | documentary footage' descended from a random distribution
             | of Gaussian noise. The Gernative AIs we have today don't
             | make things like this story by Joy Cowley with these kind
             | of illustrations.[2]
             | 
             | [1] https://pasteboard.co/VG6XNvbeC4gA.png
             | 
             | [2] https://www.slideserve.com/herbst/story-by-joy-cowley-
             | illust...
        
             | riffraff wrote:
             | Can you get SD to do something like Richard Scarry's
             | Busytown?
             | 
             | I'm asking sincerely, if we can I'll be genuinely
             | impressed.
        
               | nickthegreek wrote:
               | you could if you fine tune a model of create a LoRA on
               | busytown work. This would take under a day.
        
               | SanderNL wrote:
               | I've tried once with MJ and it looks OK. I'm not even
               | surprised anymore.
               | 
               | The devil, though, is in the details.
               | 
               | It's hard to control the outputs. The scene will look 80%
               | OK, but then you want to add some cat somewhere and that
               | won't be trivial (at the moment). There are already a lot
               | of impressive tools to help you out, like ControlNet, but
               | it's a bit difficult still.
               | 
               | You also have to do cleanup, because some random animal
               | in the background of the illustration will be horribly
               | mangled. This all basically requires you to actually be a
               | (digital) artist to fix properly. Again, for now.
               | 
               | Edit: I wanted to add that this, to me, is similar to AI
               | coding. On a good day, it might be 80% OK, but you
               | actually need to be a programmer to catch its lies and/or
               | incompetenties.
        
               | riffraff wrote:
               | > This all basically requires you to actually be a
               | (digital) artist to fix properly. Again, for now.
               | 
               | right, this is my impression of the "graphic ai
               | revolution": it will be a boon for actual digital artists
               | that can take the output and refine it, much like Copilot
               | can help programmers while not replacing them. for now,
               | anyway.
        
               | p-e-w wrote:
               | My impression has been that Stable Diffusion has
               | outclassed the vast majority of digital artists from day
               | one.
               | 
               | The only reason people feel otherwise is because when you
               | look at human-made digital art online, you tend to only
               | see the best of the best of the best. The one-in-a-
               | thousand artists who make it to the DeviantArt or
               | ArtStation frontpages.
               | 
               | But most artists aren't anywhere near as good as those
               | few geniuses, and neither are they as good as what Stable
               | Diffusion generates.
               | 
               | So my prediction is that 99% of digital artists are
               | essentially obsolete today, and the remaining 1% will
               | become obsolete some time in the next 3-5 years.
        
               | SanderNL wrote:
               | It's great, but I think reliability is as big a problem
               | there as it is in coding. It's hard to control the
               | output.
               | 
               | I won't go into the various problems these tools have,
               | like them ignoring my instructions if they are outside
               | what 90% of the population is doing, because I think they
               | can and will be solved to various degrees eventually.
               | 
               | But the bar will raise. The quality that customers expect
               | will start rising. Eventually the level of control and
               | quality that is demanded will again only be adequatedly
               | matched by people we usually call "artists". Sure, I -
               | some scumbag programmer - can create some fancy generic
               | shots today for a generic company blog post and for now
               | that _is_ impressive, but in the long term this won 't
               | cut it.
               | 
               | Folks will want highly fine-tuned specific imagery
               | trained on company stuff or whatever. Branding will have
               | to match exactly. The complexity of the design will
               | explode. Good taste and raw visual skill, not to mention
               | creativity - being an "artist" - will win and the average
               | dude will not have the proclivity to keep up those
               | skills. I think it's more than holding a pencil and being
               | good at drawing lines.
               | 
               | I _do_ think the raw skills of  "holding a pencil" and
               | such is becoming obsolete. That's like having good
               | "punching card" skills. However, knowing what to "punch"
               | is a timeless skill or at least a good part of it is and
               | it's easy to adapt to a new era if needed. Punching code,
               | typing code, selecting code-snippets from AI, it's
               | different for sure, but the underlying high-level skills
               | are adaptable and useful. I think the same holds for
               | artists.
        
               | raincole wrote:
               | > today
               | 
               | More like yesterday, to be honest.
        
             | numpad0 wrote:
             | > make 99% of professional illustrators look like they're
             | drawing stick figures by comparison.
             | 
             | I feel the opposite to be the case -- people are feeding
             | prompts and sketches and AI generated RGBD images as base,
             | and another RGB for style transfer source, all for just one
             | image. If anything, it seems, it's becoming a tool
             | increasingly limited to professionals.
        
               | grumbel wrote:
               | While it's true that you need to feed some source images
               | into the AI to get good results, what you feed into them
               | can be pretty arbitrary and low quality, as all you are
               | really doing is providing the AI with information about
               | positioning and composition that can't be expressed with
               | a text prompt. It's pretty much the "How To Draw an Owl"
               | meme with AI doing the last step.
        
               | p-e-w wrote:
               | You can get amazing results without doing any of that,
               | and I can promise you that most of the people posting AI
               | art on DeviantArt etc. aren't "professionals".
        
         | quickthrower2 wrote:
         | So is mass printing
        
         | prepend wrote:
         | But are also leading to more illustrations in children's books
         | and more children's books.
         | 
         | I think it's changing the job from pen and paper
         | (allegorically) to a prompt wrangler who gets the image right
         | for the editor or author. Kind of like how seo optimization
         | became a job and didn't even exist 30 years ago.
        
       | whywhywhywhy wrote:
       | > To build LAION, founders scraped visual data from companies
       | such as Pinterest, Shopify and Amazon Web Services -- which did
       | not comment on whether LAION's use of their content
       | 
       | Pinterest, "their content"... not sure I agree with that but if
       | we're going to use the logic that things saved to Pinterest by
       | it's users becomes Pinterest's content then isn't LAION doing the
       | same thing and the content becomes LAION's content when it's
       | saved to their database of images...
        
         | mcv wrote:
         | If LAION republishes people's copyrighted content, that sounds
         | like a pretty blatant copyright violation (edit: it's not; see
         | below). Sounds like all the artists unhappy that their art is
         | being used to train these AI systems, should be talking to
         | LAION to have their content removed from the dataset.
         | 
         | Edit: Apparently LAION doesn't republish the content, only the
         | metadata, so it's not a copyright violation. Still, it would be
         | nice if got permission or offered a way for artists to be
         | excluded from the data set.
        
           | ModernMech wrote:
           | > Still, it would be nice if got permission or offered a way
           | for artists to be excluded from the data set.
           | 
           | If artists don't want their work to impact the world, they're
           | free to keep it to themselves.
           | 
           | This whole discussion that we should allow individual artists
           | to opt out of AI art through contracts or some other legal
           | vehicle is a non starter, because it'll be impossible to
           | administrate and enforce at scale, and there's too much
           | incentive and ability for big tech to just ignore them and
           | steamroll artists. They aren't a unified bloc, and even if
           | they were how would they ever compete against big tech?
           | 
           | So what to do? Looking at productivity gains over the
           | decades, it's not clear why we are still working as hard as
           | we are. It's long overdue that productivity gains should come
           | back to the people. Maybe "artist" shouldn't be a job title
           | associated with profit/income seeking. If you want to be an
           | artist, maybe society can support that.
           | 
           | Maybe instead of using all those productivity gains to do
           | more more more, we can just work less for the same. Because
           | it seems to me the more we work, the more _they_ get richer.
           | What if instead, they didn 't get so rich, and we gave that
           | money to artists in the form of grants, like we do for
           | scientists. You do some art, apply for some grants, and you
           | get some money to do more art. It'll all be public domain,
           | anyone can use it, and big business gets to make a profit on
           | it just like with scientific advancements (I have issues
           | about that, but at least there's precedent).
        
             | shadowgovt wrote:
             | I think there's plenty of room for a finer-grained
             | permission framework and clear demand for one. I give it
             | about a year, maybe 1.5 for all the interested parties and
             | advocates to align behind a standard and make it happen via
             | big-company enforcement and the occasional copyrught
             | lawsuit (or, more commonly, enforcement via TOS).
             | 
             | Once that happens, of course, the "black market" data
             | aggregators won't care, but they're small and their product
             | can't be used by legitimate channels so it can't compete in
             | the mainstream market of ideas. What capital will do to
             | screw over artists though is... Pay them. Once a framework
             | for AI seed rights-granting is in place, a hundred or a
             | thousand legit artists can produce enough AI seed-feed to
             | legally supplant the work of hundreds of thousands.
             | 
             | There will still be room for the artist-as-celebrity with a
             | unique style that makes their art worth owning as much for
             | the fact that it came from them as for the content of the
             | canvas or the sculpture, but a huge, guaranteed-work,
             | bread-and-butter space for the visual artist, advertising
             | and entertainment media asset creation, will dry up as
             | companies backfill their art needs with functionally-free-
             | to-them mass-generated close-enough assets (advertising in
             | particular is going to be full of this... Remember when
             | "head first" photos of people were a thing for awhile? Look
             | forward to trends like that, over and over again, forever).
        
               | ModernMech wrote:
               | > all the interested parties and advocates to align
               | behind a standard
               | 
               | This is the thing that will never happen, because all
               | artists are interested parties, and as I said, they are
               | not a unified bloc. So whatever solution big companies
               | come up with for themselves, we all know ahead of time
               | that they will 1) overwhelmingly benefit big corporations
               | and 2) but insufficient to address artist concerns. When
               | they coalesce around whatever standards they end up with,
               | artists will still largely be making the same complaints.
        
               | shadowgovt wrote:
               | Too true. I should probably have said "All parties with
               | enough political and capital clout to make trouble for
               | other parties." The disorganized masses are disorganized
               | and usually don't end up with a seat at the table if they
               | don't organize.
        
           | archerx wrote:
           | Why does Pinterest get away with republishing people's
           | content? Shouldn't artist be suing Pinterest for its blatant
           | copyright violations?
        
             | jimbokun wrote:
             | I assume they have all users click on a license that has a
             | clause allowing them to use all images uploaded however
             | they want, buried deep within the text somewhere.
        
             | sharemywin wrote:
             | you have to use a section 230 DCMA take down notice if the
             | user post your stuff with out permission.
             | 
             | https://policy.pinterest.com/en/copyright
        
           | Silverback_VII wrote:
           | You can read here(in german): https://www.alltageinesfotoprod
           | uzenten.de/2023/02/20/laion-v...
           | 
           | that the non-profit LAION is going as far as intimidating the
           | creators of the images they used.
           | 
           | They don't publish the image, yes, and that is also their
           | reasoning. However, in my opinion, the intention behind all
           | of this seems obvious. They circumvent copyright claims by
           | being a non-profit and not publishing, but with the clear
           | intent that the image will be used to train some system
           | further down the road. LAION appears to be a key player in
           | how these text-to-image models dodge the copyright bullet.
        
             | antiterra wrote:
             | > They don't publish the image, yes, and that is also their
             | reasoning.
             | 
             | No, their reasoning is that they only keep a _link_ to the
             | image so there's nothing to remove.
             | 
             | A more interesting angle would be copyrightable ALT tags in
             | the form of poems or other creative content. But, the cat
             | is already out of the bag as it's probably easy enough to
             | strip poetry or other copyrightable out of alt tags with
             | the technology we've got at this point.
        
               | Silverback_VII wrote:
               | > No, their reasoning is that they only keep a link to
               | the image so there's nothing to remove.
               | 
               | I'm well aware of it, and that's also what I wanted to
               | say. However, as I mentioned, the dataset with the links
               | (to copyrighted material) is provided for machine
               | learning purposes. Saying "We only provide links" is a
               | lazy excuse from a guy with a smirking face.
        
             | Semaphor wrote:
             | Intimidating? They mention that if you sue them on
             | copyright ground when they do nothing copyright related,
             | they are able to claim damages. Which seems pretty fair, as
             | the claims by the photographers are clearly in bad faith.
             | 
             | Links are still mostly legal in Germany, if the link is to
             | something which is not, it's a different situation and
             | different from "hey, I own the copyright of my images,
             | don't link to them!"
        
               | Silverback_VII wrote:
               | We all know the purpose for which the images linked in
               | the dataset will be used. "We are a non-profit
               | organization and provide only a link" is akin to taking
               | people for fools. Why can't these systems be trained
               | exclusively on images for which people have given consent
               | or for which money has been paid?
        
               | Semaphor wrote:
               | > is akin to taking people for fools.
               | 
               | I very much disagree.
        
           | Filligree wrote:
           | > Still, it would be nice if got permission or offered a way
           | for artists to be excluded from the data set.
           | 
           | It obeys robots.txt, and the user-agent is documented.
        
             | wildrhythms wrote:
             | robots.txt generally only addresses permission to crawl;
             | not permission to reproduce the content.
        
           | tyingq wrote:
           | >Apparently LAION doesn't republish the content, only the
           | metadata
           | 
           |  _" LAION datasets are simply indexes to the internet, i.e.
           | lists of URLs to the original images together with the ALT
           | texts found linked to those images."_
           | 
           | I'm surprised link rot doesn't make this a big problem.
        
             | cuteboy19 wrote:
             | There are many twitter datasets that are already 50% gone.
             | It's very bad for reproducibility
        
         | __alexs wrote:
         | There is an T&C agreement between Pinterest and its users.
         | There is no T&Cs agreement between between LAION and
         | Pinterest's users.
        
           | debugnik wrote:
           | There's also no agreement between Pinterest and the actual
           | copyright owners of most of their content, so much reposted
           | art without even credit or a link to the source.
           | 
           | Plus, LAION is just an index, whereas Pinterest actually
           | hosts it.
        
             | [deleted]
        
             | dahwolf wrote:
             | Worse, Pinterest even ranks highly in Google, often higher
             | than the original.
             | 
             | If you'd now go and do the same, create a website that
             | rehosts known content, Google will simply delist you.
             | 
             | But magically, not for Pinterest, they get a boost instead.
             | I've always wondered what kind of under the table deal
             | makes this possible.
        
           | [deleted]
        
         | ThrowawayTestr wrote:
         | It's copyright infringement all the way down.
        
       | GaggiX wrote:
       | >Tools like Dall-E-2 and Stable Diffusion have been criticized
       | for reproducing racial stereotypes even when a text prompt
       | doesn't imply the subject's race. Such biases were why Google
       | decided not to release Imagen, which had been trained on LAION.
       | 
       | This is so ingenuous and naive to think that the reason Google
       | didn't release Imagen is because of racial stereotypes.
        
         | atleastoptimal wrote:
         | It's what Google claimed and thus the eternally true correct
         | reason and its very problematic if you think otherwise
        
           | astrange wrote:
           | If there's a reason Google doesn't want to talk about, it's
           | that the biggest legal risk of AI image generation is that it
           | can be used for revenge porn and CSAM.
           | 
           | I don't think it's problematic to be worried about that!
        
             | anonylizard wrote:
             | The biggest legal risk, empirically, is massive copyright
             | lawsuits. The AI art models are rapidly eliminating most
             | artist/photography jobs.
             | 
             | Revenge porn and CSAM are trivially solved by community
             | moderation methods like Midjourney, simply force every
             | generation to be tied to a paid account, and made in the
             | public eye (posted on discord channels).
             | 
             | So in reality, its not Google's risk conservatism, its
             | Google's utter lack of imagination and creativity, and the
             | lack of organisational incentive for anything other than ad
             | money.
        
               | dontupvoteme wrote:
               | You realize stable diffusion has been out for 6 months
               | and is more powerful than midjourney?
        
               | janekm wrote:
               | Was, when it came out, since then MJ have released 3 new
               | models (plus variants) and most users assess the MJv5
               | model as more powerful. Of course new SD models are also
               | on the horizon...
        
               | astrange wrote:
               | > Of course new SD models are also on the horizon...
               | 
               | SDXL is available at https://beta.dreamstudio.ai/ though
               | they say they're going to release more variants.
               | 
               | I think ControlNet is a lot more interesting than just
               | "better tuned models"; it means there's no line between
               | creating something yourself and asking an AI to do it
               | anymore.
        
               | yellow_postit wrote:
               | Maybe in some technical sense but in practice the
               | majority of quality (subjective) work I see shared on
               | Twitter is Midjourney backed. Which is crazy to me
               | because I don't think I could design a more frustrating
               | interface than Discord.
        
               | nickthegreek wrote:
               | Midjourney gives you good images with near zero work. SD
               | makes you work for the output but gives you FAR greater
               | control but also needing to be technically skilled, and
               | have pricey hardware. Given this, it is easy to
               | understand why the majority of what you see is done in
               | midjourney.
        
               | SanderNL wrote:
               | My god, sorry for the dumb post, but SD is only _months_
               | old? It already feels like it has been around forever.
               | This space is insane.
        
               | astrange wrote:
               | > The AI art models are rapidly eliminating most
               | artist/photography jobs.
               | 
               | Automation is associated with increased employment.
               | 
               | https://en.wikipedia.org/wiki/Jevons_paradox
               | 
               | > Revenge porn and CSAM are trivially solved by community
               | moderation methods like Midjourney
               | 
               | Please don't go around saying things like entire safety
               | departments being trivially solved by anything. Community
               | moderation isn't a legal solution to CSAM, it's arguable
               | the community users are committing crimes by looking at
               | it. (Extra difficult because this depends on the
               | country.) And they're certainly not getting healthcare
               | benefits for it, so they could probably sue you.
               | 
               | This is the reason they're on Discord though, it's
               | because Discord handles the legal compliance for these
               | things.
        
             | ImprobableTruth wrote:
             | They could just filter it like Midjourney or DALL-E.
        
               | astrange wrote:
               | How do they know they can "just" do something that came
               | out a year after they invented Imagen?
               | 
               | Especially when getting it wrong once might mean Europe
               | makes AI illegal. Making things illegal is their favorite
               | hobby.
        
               | ImprobableTruth wrote:
               | Midjourney and DALL-E were both announced and in closed
               | beta a month _before_ Imagen was announced, with filters
               | already in place.
        
         | andresgottlieb wrote:
         | "such biases"
        
         | klodolph wrote:
         | Google previously had problems with Google Photos categorizing
         | dark-skinned people as "gorillas" or something like that. This
         | would show up when you were searching for photos in your
         | personal Google Photos account. Everyone talked about it.
         | Nobody wants to release a product like that. Removing bias is
         | hard; you can't just decide to get an unbiased training set or
         | something magical like that.
         | 
         | It is NOT surprising that larger companies would behave with an
         | abundance of caution. Leave it to the startups to rush in head-
         | first, cause a bunch of problems, make a cool product, get a
         | consent decree, and capture the market. Large companies that
         | act that way get PR problems and lawsuits.
        
           | inciampati wrote:
           | And bearded men became "pets"
        
           | ImprobableTruth wrote:
           | Except they did rush their oh so dangerous language model out
           | when push came to shove - turns out it just sucks in
           | comparison. Frankly, I expect something similar for image
           | generation. For all the good research Google's labs do,
           | they're not great at turning it into an actual product.
           | Ethics just seems like a fig leaf.
        
             | yellow_postit wrote:
             | The ethics team has the unenviable in a corporate
             | environment role of being "no" by default gatekeepers. This
             | is likely why the responsible ai groups, focused on
             | mitigation have survived while the ethics groups, focused
             | on systemic issues, have shrunk.
        
           | smeagull wrote:
           | Is that a bias? or do those images just look more like those
           | categories at the pixel level?
           | 
           | There is a reason those words are the ones chosen by racists
           | in the first place. Not just because they're hurtful, but
           | because there is an objective similarity that makes the rest
           | of the comparison seem, however shallowly, to validate their
           | views.
           | 
           | Removing bias with censorship at the training set level is
           | silly, and will likely hamper the AI's performance. Better to
           | train the AI to not produce problematic output at the higher
           | level, and ensure class membership in the training data is
           | representative.
        
             | prepend wrote:
             | > and ensure class membership in the training data is
             | representative.
             | 
             | This seems like the most productive focus area. Add in more
             | variety in the training images and output should get closer
             | to representing the global population.
        
           | anonylizard wrote:
           | In retrospect, those concerns were pretty much drummed up by
           | the "AI ethics" people, who needed an easy topic to
           | demonstrate their importance.
           | 
           | When actually powerful AIs got released, the 'bias' angle got
           | overwhelmed 100x by the "trained on copyright data" and "Are
           | we going to lose our jobs" angle. Your average black artist
           | doesn't care if SD defaults to drawing white people, because
           | A: he cares about not being able to draw for a living,
           | permanently, 100 times as much. B: He can always add 'dark
           | skinned' to his prompt
           | 
           | The bias problem is also relatively easy to solve (Midjourney
           | has already made massive improvements), while the
           | copyright/job loss problem is extremely hard.
           | 
           | The AI ethics people have had shockingly little to say on
           | copyright/job loss issues. Which is why they got fired en
           | masse.
           | 
           | Large organisations systemically overestimated the risks from
           | bias, and underestimated the actual AI risks that society
           | actually cares about. I think the answer is also simple,
           | because accused of racism, will cost any executive in a large
           | company their job. Being accused of automating millions of
           | jobs, will earn them prestige and a promotion. Smaller
           | companies can ignore those accusations because they aren't
           | vulnerable to activist pressure, they answer only to their
           | customers.
        
             | klodolph wrote:
             | > The bias problem is also relatively easy to solve
             | (Midjourney has already made massive improvements), while
             | the copyright/job loss problem is extremely hard.
             | 
             | It only seems easy to solve on the surface, it's a deep
             | problem. It's also not just the bias thing, Bing and
             | ChatGPT have been saying some truly unhinged things.
             | 
             | > I think the answer is also simple, because accused of
             | racism, will cost any executive in a large company their
             | job.
             | 
             | It takes more than an accusation, otherwise you could go
             | around accusing executives you don't like of racism and
             | getting them fired.
             | 
             | Ethics is a tenuous job position at best, even in a large
             | company. It's seen as a cost center. I don't think there's
             | much to read into why AI ethicists would get laid off.
        
               | anonylizard wrote:
               | Bing and GPT's 'unhinged' comments are not a result of
               | bias, an AI wanting to escape won't be fixed if you
               | magically fed it antifa approved only data. That's
               | systemically different from the discrimination issue
               | drummed up earlier.
               | 
               | Also, we are talking about social and business impact
               | here. Its now proven that vast majority of society cares
               | about job loss 100x more than bias. For a research field
               | that focuses on social impacts of AI, its damning they
               | have little to say on this area.
        
               | klodolph wrote:
               | > an AI wanting to escape won't be fixed if you magically
               | fed it antifa approved only data
               | 
               | If these are the sentences you're producing, then maybe
               | you should take a look at what kind of data is being fed
               | into your neural network. Like, where are you getting
               | this antifa stuff?
               | 
               | > its damning they have little to say on this area.
               | 
               | One of the problems you need to solve into get good AI-
               | generated text is that you need to resolve pronouns. It's
               | also a problem in this sentence, because I don't know
               | where the antecedent for "they" is.
        
               | fortran77 wrote:
               | The person you were replying to, while making a
               | provocative statement, wasn't personally attacking
               | anyone. You are. And that's not nice.
        
               | eropple wrote:
               | Going on about "antifa approved" is chud stuff that is
               | employed only and specifically to normalize an _actually
               | fascist_ viewpoint under the guise of talking about
               | something else.
               | 
               | Pass.
        
               | catiopatio wrote:
               | Says the person using "chud" and "fascist", which is an
               | extremely strong indicator of your membership in
               | terminally online antifa/leftist subculture.
        
               | [deleted]
        
               | klodolph wrote:
               | Real talk: Someone making "provocative" statements like
               | that is not making them in a vacuum. There is some kind
               | of context to those statements. I want to know what kind
               | of context creates the kind of garden where statements
               | like that grow. Like, Antifa, a non-organization which
               | has nothing to do with AI, would vet AI training data? I
               | want to know where that kind of hypotheses comes from.
        
               | hattmall wrote:
               | It's not provocative, he says antifa once. It's meant as
               | an example of a method of removing bias. He's saying you
               | could train the AI entirely on an organizations strictly
               | vetted and approved data, it could be antifa, the
               | Catholic church or Coca Cola, it would still potentially
               | say unhinged things even if you eliminate bias. Top many
               | of these concepts are built into the language. Look at
               | all the ways we use the word kill for a variety of
               | topics, most being very benign, but it can be disturbing
               | if the AI starts talking about killing things.
        
             | jrm4 wrote:
             | Serious question.
             | 
             | Did you ask these black artists? Because if you didn't, you
             | REALLY need to _shut it_ and keep your ideas out of their
             | mouths.
        
               | jrm4 wrote:
               | Hmm, lots of downvotes, but no answer from op?
               | 
               | If y'all need a reason why people _avoid_ this site, here
               | it is.
        
               | PeterisP wrote:
               | I downvoted you because of the sentence "Because if you
               | didn't, you REALLY need to shut it and keep your ideas
               | out of their mouths.".
               | 
               | The point you rised at "Serious question. Did you ask
               | these black artists?" is a valid, reasonable point that
               | merits discussion.
               | 
               | But the point raised by the parent post is _also_ a
               | valid, interesting point that merits discussion. And that
               | point raised a valuable discussion EVEN IF they did not
               | "ask these black artists", and if we put this bar and
               | demand them to shut it, that's a bad thing for the
               | discussion.
               | 
               | So in my view you made an insulting demand for someone to
               | "shut it" without reasonable grounds to do so, and this
               | definitely deserves a downvote or five. I'm not asking
               | you to "shut it", you should participate in this
               | discussion, but in a civil manner that also allows posts
               | like the parent post to participate in the discussion
               | even if they don't meet your demands.
        
               | jrm4 wrote:
               | I _absolutely did_ because in my experience a lot of
               | people do this all too often, and it warrants an
               | immediate reaction. I will gladly apologize if I am
               | wrong, but it seems like most people here don 't
               | understand the harm that "non-black people speaking for
               | black people" causes. It is SIGNIFICANTLY more rude and
               | harmful than my tone here.
               | 
               | I apologize for nothing; it is most everyone else here
               | that needs to do better.
        
               | Ajedi32 wrote:
               | What _really_ causes harm is when you divide people up
               | into groups based on arbitrary characteristics (like skin
               | color) and treat those characteristics like they 're the
               | sole defining aspect of each person's identity, like
               | you're doing right now.
               | 
               | Framing this as "non-black people speaking for black
               | people" carries the implicit assumption that all black
               | people have the same/similar opinions; that an arbitrary
               | black person _would_ be able to meaningfully  "speak for
               | [all] black people" in a way that an arbitrary non-black
               | person can't. That's wrong.
               | 
               | The previous commenter making an educated guess based on
               | personal experience with zero concrete data points is
               | only slightly worse than making that same guess based on
               | one concrete data point, and neither situation would be
               | justification for telling anyone to "shut it" or "do
               | better" in my opinion.
        
               | swores wrote:
               | I didnt downvote you and agree with your general point,
               | but consider that maybe downvotes came in due to your
               | tone rather than your meaning. I actually normally would
               | downvote a comment written like yours was, just less so
               | when I think it's making a still-important point (but
               | you're less likely to convince anyone your point is
               | important if you annoy them rather than try to enlighten
               | them!)
        
               | jrm4 wrote:
               | And again, as I said prior -- my tone is 100% appropriate
               | compared to the genuine _offense_ likely committed above.
               | 
               | I'm aware that normally one should keep an even tone.
               | This is not one of those times and everyone _else_ here
               | needs to learn that.
        
               | swores wrote:
               | Not sure what you mean "as I said prior", unless you're
               | annoyed that I hadn't looked into the future to see
               | replies about tone you hadn't written when I posted my
               | comment...
               | 
               | But anyway, I wasn't complaining, just pointing out my
               | view (since you asked about being downvoted) that by
               | using that tone, however justified it can feel to lower
               | to somebody else's level, you _will_ get some people who
               | take less note of what you 're actually saying or who
               | downvote without thinking about the subject beyond "I
               | dislike seeing that tone in HN comments".
        
               | anonylizard wrote:
               | If you want an answer, here it is.
               | 
               | I've been monitoring artist forums since day 1 of SD's
               | release. From various subreddits, to discord communities,
               | to 4chan, to forums in other languages. I've never seen
               | an artist complain even ONCE about bias, it is always,
               | always, always, about jobs/copyright('stealing').
        
               | jrm4 wrote:
               | So you _didn 't._
               | 
               | Again. cut it out.
        
               | ben_w wrote:
               | The question in your previous comment is totally fair,
               | IMO.
               | 
               | Trouble is, I genuinely don't know what to usefully
               | suggest, because the obvious thing that comes to mind
               | (focus on the actual question and strip the aggression)
               | is a cliche to the point where I suspect I already know
               | your response will be some form of eye-roll at my
               | privilege etc.
               | 
               | And it would be a legit response, too, given that people
               | presuming to know me is annoying enough _even without it
               | being a daily experience_.
        
               | jrm4 wrote:
               | It's pretty simple; many (I'm presuming) white people
               | have a deeply nasty habit of conflating their own
               | experiences with others because it makes sense to them,
               | and THAT problem should be recognized.
               | 
               | Just _don 't do it._ or at the very least serve it up
               | with a heaping helping of "I would imagine that many..."
               | so we know that this is JUST YOUR SPECULATION.
        
               | ben_w wrote:
               | It would be great if telling people that made them act
               | differently; but call it Armchair Generals, or typical
               | mind fallacy, or mansplaining, or Dunning-Krugering, or
               | ivory-tower academics... the problem has so many forms
               | and even knowing about it makes it hard to avoid in
               | oneself.
        
             | darkwater wrote:
             | > Large organisations systemically overestimated the risks
             | from bias, and underestimated the actual AI risks that
             | society actually cares about.
             | 
             | The point of large organizations working on
             | AI/LLMs/whatever is _precisely_ to reduce labor costs by
             | reducing the number of human workers and replacing them
             | with AI. Why should they care about this in their
             | development? (I 'm not saying it's the right thing to do,
             | just stating the facts about capitalism)
             | 
             | It is the rest of society that has the interest to fight
             | those changes back.
        
             | skinkestek wrote:
             | > The bias problem is also relatively easy to solve
             | (Midjourney has already made massive improvements),
             | 
             | One thing that everyone seems to ignore these days are that
             | certain settings are inherently not diverse.
             | 
             | Until I was about 12 I didn't see non white people in real
             | life except on holidays and once when I followed a relative
             | who visited a refugee family.
             | 
             | Not because we avoided them, but because they weren't
             | there.
             | 
             | Background: The country I lived as a child in didn't have a
             | recent (last ~900-1000 years) history with keeping slaves
             | and only recently experienced high living standards so
             | there were few immigrants too.
             | 
             | In the large district I grew up I am aware of two non
             | western groups until the Balkan wars in the 90ies: a single
             | middle eastern guy who I can't recall meeting and a
             | southern American family who I only started meeting after I
             | changed school at some point.
             | 
             | For me, it is annoying that people insist that all pictures
             | include some non-Europeans, because reality is not like
             | that everwhere.
             | 
             | I imagine the same is maybe even more true for African or
             | Asian communities.
        
               | Al-Khwarizmi wrote:
               | Yeah, same experience here. In Spain there is a character
               | in child mythology (one of the Santa Claus analogues, who
               | bring gifts from children) who is black. I remember that
               | when I was a kid in the 80s, he used to be portrayed by a
               | white man painted black, what Americans call "blackface",
               | but for purely practical reasons: no one was black.
               | 
               | Perhaps in big cities there were some, I don't know, but
               | in many places in the country you would live your life
               | without ever seeing a black person outside of TV. Just as
               | you say, we didn't have a past of slavery and at that
               | point the country wasn't rich enough to be an immigration
               | magnet, we did have some immigration from Morocco but
               | Moroccans aren't black.
               | 
               | Now it's very different, of course, and we do have enough
               | black people that that character can be portrayed more
               | realistically :)
        
               | dragonwriter wrote:
               | > Yeah, same experience here. In Spain there is a
               | character in child mythology (one of the Santa Claus
               | analogues, who bring gifts from children) who is black.
               | 
               | I would assume a somewhat _close_ analog, since Saint
               | Nicholas (the _name_ source, if only a small part of the
               | overall inspiration, of the "Santa Claus" figure) is, in
               | one major tradition of, I believe, Italian origin,
               | typically depicted as very dark-skinned (probably
               | originally as a sign of foreignness rather than literal
               | racial blackness; as he was geographically from Asia
               | Minor and apparently of Greek ethnicity.)
        
               | vasco wrote:
               | > in Spain
               | 
               | > Just as you say, we didn't have a past of slavery
               | 
               | Spain (together with Portugal) were the first to invent
               | "put people in boats and ship them across the Atlantic"
               | type of slave trade and had large numbers of slaves in
               | their colonies.
               | 
               | And as with all other European countries that had big
               | boats, they also brought many slaves back to Spain
               | itself. For hundreds of years.
               | 
               | Also calling blackface "practical because there are no
               | black people around" is missing the point of why people
               | don't like blackface.
               | 
               | https://en.m.wikipedia.org/wiki/Slavery_in_Spain
        
               | Al-Khwarizmi wrote:
               | Yeah, my wording was definitely too general. There was
               | slavery in Spain, my intended meaning was that there was
               | no systematic enslavement of black people but there is
               | definitely a chasm between what was in my brain and what
               | I actually wrote, so I stand corrected.
               | 
               | Regarding blackface, what do you suggest to do, then, if
               | there is a tradition of someone portraying a character
               | (whom children believe in) that happens to be black, if
               | there are no black people at all to do it?
        
               | foldr wrote:
               | From the Wikipedia article linked above:
               | 
               | >By the 16th century, 7.4 percent of the population in
               | Seville, Spain were slaves. Many historians have
               | concluded that Renaissance and early-modern Spain had the
               | highest amount of African slaves in Europe
               | 
               | As for what to do about blackface characters, the
               | simplest solution is just to portray them differently. I
               | believe this is gradually becoming a more popular option
               | in the Netherlands, which has a similar problem with
               | Zwarte Piet: https://raffia-
               | magazine.com/2020/12/02/outgrowing-zwarte-pie...
        
               | wizofaus wrote:
               | Do you have as an example of where applying dark coloured
               | make-up to anyone of a European background to make them
               | look exotic/foreign doesn't also look completely
               | ridiculous (and inevitably somewhat insulting)? And yes
               | I'd include "Zwarte Piet". If you're putting on a
               | performance where it's crucial that the audience
               | identifies a particular character as African (or Moorish)
               | then find a performer who can convincingly and
               | respectfully pull it off sans make-up, or adapt the plot
               | to what you have available. I gather that's what's done
               | for the part of Othello for at least 30 years now at any
               | rate.
        
               | hoseja wrote:
               | The only innovation there is the "Atlantic" part. Stop
               | pretending like it was something new and uniquely
               | horrible.
        
               | ben_w wrote:
               | I think the criticism here is that Iberian history is not
               | uniquely _better_ rather than claiming it 's uniquely
               | _worse_.
               | 
               | (I'm British, so I'm not going to claim less-awful-than-
               | thou against any nation).
        
               | vasco wrote:
               | If you agree with that innovation, you also seem to agree
               | with me that the spanish were in fact involved with
               | slavery, as opposed to what the person I replied to said.
               | Not sure where you saw me "pretend" about anything, I
               | just presented facts and a wikipedia article.
        
               | garte wrote:
               | I get what you mean and I had a similar experience
               | growing up.
               | 
               | But: I think that they were there. We just didn't see
               | them a lot. They were cleaning, working in factories,
               | tending gardens, looking after children etc. but they
               | were not represented and thus were not visible. Of course
               | they were a minority, but I'd wager they were there.
               | 
               | Upwards social mobility and thus equality is only
               | possible with recognition and representation. Without
               | "being in pictures" they don't exist. That's why pictures
               | should include "Non-Europeans" (this expression in itself
               | is very problematic btw: there is no "European" as
               | something biologically inherent in a person).
        
               | MarcoZavala wrote:
               | [dead]
        
               | MarcoZavala wrote:
               | [dead]
        
               | Filligree wrote:
               | > But: I think that they were there. We just didn't see
               | them a lot. They were cleaning, working in factories,
               | tending gardens, looking after children etc. but they
               | were not represented and thus were not visible. Of course
               | they were a minority, but I'd wager they were there.
               | 
               | No, they genuinely weren't there. Europeans regularly
               | tried to move to my hometown, got depressed/sick from the
               | lack of sunlight in winter, and left. Refugees coming
               | from further south had it... worse.
               | 
               | Northern Norway isn't a place outsiders can easily move
               | into. At least, not without installing grow lights in
               | their living room...
        
             | fatneckbeard wrote:
             | wrong.
             | 
             | AI bias is going to be a huge problem when morons in court
             | systems start using it to convict people, or businesses use
             | it for hiring decisions or firing decisions
             | 
             | not only from the stupidity but since there are some shreds
             | of the civil rights laws of the 1960s still active, AI
             | companies could be liable for de-facto racist decisions
             | that they informed.
        
               | charcircuit wrote:
               | >morons in court systems start using it to convict people
               | 
               | Those "morons" are also biased.
               | 
               | >businesses use it for hiring decisions or firing
               | decisions
               | 
               | If businesses realize the AI is doing a bad job with
               | hiring unqualified people they will not use it for
               | hiring.
        
               | klodolph wrote:
               | Most businesses are not very good at figuring out if they
               | are hiring the right people.
        
               | XorNot wrote:
               | "Start"?[1,2]
               | 
               | The problem I have with every discussion of AI risk is
               | that people seem terminally underinformed on what is
               | actual reality now, or _why_ some risk _is_ a risk.
               | 
               | "AI" isn't a problem which can destroy the legal system:
               | because it takes regular human institutions to allow such
               | a miscarriage of justice. Which as noted, they started
               | doing, are still doing.
               | 
               | So you get this weird "perpetual future" perspective
               | where everything "AI" is going to do is solely something
               | that the technology will cause, not its users, and the
               | solution is always to prevent the technology existing
               | rather then fix the system - as though the US and other
               | jurisdictions don't have long history's of injustice for
               | all sorts of groups.
               | 
               | The problems aren't new, and the solutions have nothing
               | to do with whether you can create predictive algorithms.
               | And "oh but what about the scale..." is just a
               | declaration that you're aware of the problem but were
               | pretty sure it wouldn't happen to you - because
               | absolutely nothing else prior actually prevented it
               | except the social privilege you inherited which means you
               | could ignore it.
               | 
               | [1] https://www.technologyreview.com/2019/01/21/137783/al
               | gorithm...
               | 
               | [2] https://www.wired.co.uk/article/police-violence-
               | prediction-n...
        
             | dragonwriter wrote:
             | > The bias problem is also relatively easy to solve
             | 
             | No, its not.
             | 
             | > (Midjourney has already made massive improvements)
             | 
             | Maybe. But it hasn't come anywhere close to solving it even
             | in its domain, which is probably the _least_ concerning
             | domain of AI bias, and not necessarily transferrable.
             | 
             | Actual ML systems that are deployed in production by
             | governments in important roles have massive bias problems,
             | as do SOTA LLMs (including, very much, GPT-4.)
             | 
             | > The AI ethics people have had shockingly little to say on
             | copyright/job loss issues.
             | 
             | The AI ethics people have a _lot_ to say about the first
             | (well "ethics of sourcing without consent" is probably more
             | to the point than copyright, copyright is a component, but
             | that's more legal than ethical).
             | 
             | They have some to say about the second (well, again "job
             | loss" is the wrong framing; "skill devaluation" is probably
             | more on point), but that's frankly _not_ an AI issue, its
             | an economic system issue and effects all technological
             | change the same way. Solving _capitalism_ is mostly out of
             | scope for AI ethics. EDIT: Specifically, increasing
             | material output for labor input ought to be, in and of
             | itself (and leaving aside ethical questions of how you do
             | that, which the sourcing issue addresses) a good thing; if
             | there are material losers in that, it is because the
             | economic system to distribute output well, which is not an
             | ethical issue of the system which enables the productivity
             | gain but an ethical issue of the economic system.
             | 
             | > Large organisations systemically overestimated the risks
             | from bias
             | 
             | No, they didn't. Those aren't even really _risks_ any more,
             | they are massive current costs of existing adoptions.
             | 
             | > and underestimated the actual AI risks that society
             | actually cares about.
             | 
             | That society (as weighted by social power) doesn't care
             | about the kinds of bias problems AI manifests is exactly
             | _why_ it is an ethical issue.
             | 
             | > I think the answer is also simple, because accused of
             | racism, will cost any executive in a large company their
             | job.
             | 
             | The large companies that have supplied systems which
             | manifestly suffer from racial and other class biases and
             | which are, in fact, being used in production to implement
             | government policy around the world have _not_ lost their
             | jobs, and neither have the government officials procuring
             | them and responsible for the programs they are deployed in,
             | so this is clearly false.
        
               | ipaddr wrote:
               | In general if no one cares does it make it an ethical
               | issue? No one cares if you kill an ant. No one cares if
               | you kill during a war. If what we value vs what we do
               | becomes out of sync an ethical issue arises
               | 
               | We haven't been able to remove biase in society. We just
               | added more layers. The more you try to hide the truth to
               | remove baise you start creating your own. The closer we
               | come to accepting the truth the closer we get to solving
               | bias by addressing it not erasing it.
               | 
               | AI ethics should cover economic issues brought on by AI.
               | The already discuss sociological issues
        
               | dragonwriter wrote:
               | > In general if no one cares does it make it an ethical
               | issue?
               | 
               | There is a big difference between "no one cares" and
               | "society (weighted by social power) doesn't care".
               | 
               | Entrenching and reinforcing bias against the already
               | socially weak is something that society, weighted by
               | social power, does not care about, but it is precisely
               | that that makes it an important ethical issue.
               | 
               | > AI ethics should cover economic issues brought on by
               | AI.
               | 
               | Capitalism's failure to distribute economic gains well is
               | not a issue brought on by AI, and there is already a
               | _much_ larger body of ethical philosophy directed at that
               | problem with whose work anything the much smaller number
               | of AI ethicists would direct at the problem would be
               | redundant.
               | 
               | This is not a problem that remains because of inadequate
               | attention by an appropriate set of ethicists, but, again,
               | because it is one that society, weighted by social power,
               | very much _does not_ care about. Social power in
               | capitalism is, in fact, very much concentrated in those
               | for whom this problem is a benefit, and it is
               | concentrated there _as a direct result of this problem_.
        
       | strkitten wrote:
       | https://xkcd.com/2347/
        
       | helsinkiandrew wrote:
       | https://archive.ph/UIu69
        
       | ChicagoBoy11 wrote:
       | As a tech-inclined person working at a private school with a
       | healthy amount of time to putter and work on my own projects, I
       | feel like I need to get busy!
        
       | Bostonian wrote:
       | "That person is high school teacher Christoph Schuhmann, and
       | LAION, short for "Large-scale AI Open Network," is his passion
       | project. When Schuhmann isn't teaching physics and computer
       | science to German teens, he works with a small team of volunteers
       | building the world's biggest free AI training data set, which has
       | already been used in text-to-image generators such as Google's
       | Imagen and Stable Diffusion.
       | 
       | Databases like LAION are central to AI text-to-image generators,
       | which rely on them for the enormous amounts of visual material
       | used to deconstruct and create new images. The debut of these
       | products late last year was a paradigm-shifting event: it sent
       | the tech sector's AI arms race into hyperdrive and raised a
       | myriad of ethical and legal issues. Within a matter of months,
       | lawsuits had been filed against generative AI companies Stability
       | AI and Midjourney for copyright infringement, and critics were
       | sounding the alarm about the violent, sexualized, and otherwise
       | problematic images within their datasets, which have been accused
       | of introducing biases that are nearly impossible to mitigate."
        
       | janmo wrote:
       | We need the same with text to train Large Language Models
        
         | circuit10 wrote:
         | https://pile.eleuther.ai/
        
         | simonw wrote:
         | That's pretty much what Red Pajama is:
         | https://simonwillison.net/2023/Apr/17/redpajama-data/
        
         | angrais wrote:
         | It contains captions. That's the alt text. LAION is used for
         | train LLMs.
        
       | vrglvrglvrgl wrote:
       | [dead]
        
       | seydor wrote:
       | " ...in america
       | 
       | Unlikely to power anything for-profit in europe
        
         | archerx wrote:
         | I think Stable Diffusion may use part of LIAON and stability ai
         | is based in London.
        
       | pxoe wrote:
       | "free database", except the images in it may not be so "free".
       | even their "open" 400M dataset makes a caveat like "if we found a
       | CC license, we put it there, otherwise, well, we just don't
       | know.", so there's still a bunch of data with no license
       | information, that got included anyway. their 5B dataset simply
       | does not have license metadata, at all. a little bit misleading,
       | a little bit malicious. completely enabling copyright and license
       | violations of data it 'only links to'.
        
       ___________________________________________________________________
       (page generated 2023-04-25 23:02 UTC)