[HN Gopher] Llama and ChatGPT Are Not Open-Source
       ___________________________________________________________________
        
       Llama and ChatGPT Are Not Open-Source
        
       Author : webmaven
       Score  : 33 points
       Date   : 2023-07-27 21:27 UTC (1 hours ago)
        
 (HTM) web link (spectrum.ieee.org)
 (TXT) w3m dump (spectrum.ieee.org)
        
       | aaur0 wrote:
       | I wrote a small piece sharing my views on the topic
       | https://medium.com/@anandandbeyond/rethinking-open-source-fo...
        
         | version_five wrote:
         | And posted the link to your subscribers only medium page twice
         | in this thread
        
       | smoldesu wrote:
       | Previously:
       | 
       | https://news.ycombinator.com/item?id=36820122
       | 
       | https://news.ycombinator.com/item?id=36806536
       | 
       | https://news.ycombinator.com/item?id=36791671
       | 
       | https://news.ycombinator.com/item?id=36783019
        
         | version_five wrote:
         | To be fair, the title is the same but this article is really
         | about a study of the openness characteristics of different AI
         | models, which is new. This link had the comparison of different
         | models: https://opening-up-chatgpt.github.io/
        
       | littlestymaar wrote:
       | Can we please stop using terminology related to code for
       | something that's _not_ code?
        
         | hprotagonist wrote:
         | sometimes it's hard to see a difference.
        
         | aaur0 wrote:
         | Exactly - I wrote something on similar lines here :
         | https://medium.com/@anandandbeyond/rethinking-open-source-fo...
        
       | jehb wrote:
       | Of note, if you're interested in helping participate in the
       | discussion of what an open source AI _would_ actually look like,
       | the Open Source Initiative is looking for your help:
       | 
       | https://opensource.org/deepdive/
        
         | version_five wrote:
         | In there somewhere in there where they are asking for input? It
         | presents itself that way but it seems the only way to "help" is
         | to propose a presentation under their call for speakers:
         | https://sessionize.com/deepdiveai
         | 
         | Is there another way to contribute?
        
       | bick_nyers wrote:
       | I personally do not want the companies to release training data
       | (at least for a while) because then it gives people leverage to
       | neuter it.
       | 
       | I don't want a sanitized LLM, and I don't have $60M lying around
       | to train my own.
       | 
       | Copyrighted material, sexual content, political opinions, throw
       | it all in and release it please!
       | 
       | Yes, reducing bias in the models is a noble goal, but introducing
       | new bias and blindspots to do it is a no-no.
       | 
       | Maybe I just got added to a list somewhere for having this
       | opinion.
        
         | kristopolous wrote:
         | We really need to somehow separate a bias towards accuracy as
         | distinct from some bias towards say, a sports team.
         | 
         | Using everything would be like taking a bunch of students final
         | exams and then claiming the most common answers are the correct
         | ones.
         | 
         | This isn't how expertise and accuracy works. Most things worth
         | doing are not only genuinely hard and complicated but something
         | that only a minority subset of accomplished people can do
         | consistently well in.
         | 
         | Listening to everybody and incorporating their thoughts is only
         | going to lead to wrong answers.
         | 
         | Being selective is the key here unless you genuinely want say,
         | answers about to space to involve aliens and UFOs - because way
         | more people believe in that then there are qualified PhD
         | astrophysicists in the world.
         | 
         | Similarly, way more people believe in vaccine conspiracy
         | theories then there are people with significant viral
         | epidemiology backgrounds.
         | 
         | This pattern is true in every field.
        
           | bick_nyers wrote:
           | Why is it a binary choice? "Most students would answer X, due
           | to this common misconception about Y".
           | 
           | As we have seen from history, there is not often an absolute
           | truth to questions, only clusters of truths. We want our LLM
           | to be able to perform reasoning, mathematics, and science,
           | but expecting absolute truths in anything outside of those
           | fields is a bit much. Wikipedia often takes a good approach
           | here and strikes this balance well. You can represent the
           | "crazies" and show their reasoning, and then draw attention
           | to how it is commonly refuted. You don't get this if you just
           | omit the "crazies" to begin with.
        
         | version_five wrote:
         | 100% agree. People that want training data released mostly just
         | want to find something to attack, it has nothing to do with
         | transparency.
        
         | threeseed wrote:
         | > reducing bias in the models is a noble goal
         | 
         | It's also a necessary goal in order for these models to be more
         | broadly adopted.
         | 
         | We've seen too many examples of bias in the training data set
         | manifesting in ways that actively discriminate against people.
         | Which is unethical and in many places illegal.
         | 
         | And having copyrighted material and sexual content in your
         | model will simply open you up to lawsuits as is happening right
         | now between authors and OpenAI. Not sure that is a position
         | most startups want to be in.
        
           | bick_nyers wrote:
           | College students don't want to be criminally charged with
           | theft for pirating an $800 college textbook either. I'm not
           | giving business advice, I'm giving humanity and knowledge
           | propagation advice.
           | 
           | As I mentioned, reducing bias is good so long as it doesn't
           | introduce more bias elsewhere. The ultimate goal of course
           | being a 100% bias free model.
        
           | version_five wrote:
           | Nah, these are foundation models. Companies want to be able
           | to put in guard rails that are applicable to their
           | application, not start with a model lobotomized according to
           | US tech company values. The censorship is about telling
           | people how to think, like it always is, not for the good of
           | the people using the models.
        
             | bick_nyers wrote:
             | Yup. Advance the foundation models as far as possible to
             | create a representation of the internet/human experience
             | and place safeguards on top.
             | 
             | Don't want your LLM to be used to create erotic fanfiction?
             | Instead of gutting the LLM, just put a detector on the
             | query and answers feeding into the LLM. Thankfully, with a
             | non-neutered model you have access to a tool than can be
             | used to perform such detection...
        
             | threeseed wrote:
             | > US tech company values
             | 
             | Which do broadly align with the values across most of the
             | world.
             | 
             | But if you want to build something that has different
             | values then go ahead.
             | 
             | But you can't expect companies to be complicit in doing
             | something which is unethical or illegal.
        
               | [deleted]
        
             | porantz wrote:
             | [dead]
        
       ___________________________________________________________________
       (page generated 2023-07-27 23:01 UTC)