[HN Gopher] Samsung workers made a major error by using ChatGPT
       ___________________________________________________________________
        
       Samsung workers made a major error by using ChatGPT
        
       Author : deesep
       Score  : 43 points
       Date   : 2023-04-05 17:46 UTC (5 hours ago)
        
 (HTM) web link (www.techradar.com)
 (TXT) w3m dump (www.techradar.com)
        
       | withinrafael wrote:
       | Confusing article. It appears the company discovered employees
       | were pasting confidential information into ChatGPT and are
       | assuming that data is now comprised given OpenAI policies stating
       | conversations are periodically reviewed and used for training.
       | The data doesn't appear to be accessible to the public directly.
        
         | voytec wrote:
         | It says at the beginning, that Samsung allowed employees to do
         | so:
         | 
         | >> The company allowed engineers at its semiconductor arm to
         | use the AI writer to help fix problems with their source code.
         | 
         | > The data doesn't appear to be accessible to the public
         | directly.
         | 
         | Allegedly (reddit), some random chats were recently listed in
         | people's chat lists[1]
         | 
         | [1] https://news.ycombinator.com/item?id=35236660
        
           | navanchauhan wrote:
           | Not allegedly, they actually were.[0]
           | 
           | [0] https://openai.com/blog/march-20-chatgpt-outage
        
           | withinrafael wrote:
           | I forgot about that incident, good call! It's not
           | inconceivable that one of these chats got reported back to
           | Samsung.
        
       | brundolf wrote:
       | I think the market is going to explode (if it hasn't already) for
       | on-prem, or at least private, LLMs on par with ChatGPT. This
       | could be served by companies building their own, or by open-
       | source projects, or by OpenAI or OpenAI's competitors
       | 
       | As a side-effect, this feels like a bright spot in the
       | potentially authoritarian trajectory that AI could take as labor
       | becomes less and less valuable. It encourages development of LLMs
       | that compete with the current default option and can be run on
       | more and more limited hardware. Enterprises might even want
       | separate departments, or separate individuals, to be able to run
       | their own models to prevent leakage
        
         | eachro wrote:
         | Open source is a race to the bottom. Seems like the only
         | obvious winner then is people selling the shovels aka NVIDIA.
        
           | esafak wrote:
           | More like a race to the top: you're forgetting all the
           | applications that will be built on top of these open source
           | models.
        
             | eachro wrote:
             | I'm not so sure. Does anyone pay for pytorch, numpy,
             | tensorflow? In the matter of weeks we've seen llama.cpp,
             | alpaca.cpp released to the public. Barrier to entry in this
             | market is quickly going to zero.
        
       | bob1029 wrote:
       | How did they gain access to ChatGPT from their offices?
       | 
       | I worked in the ATX factory about a decade ago and the network
       | was _very_ locked-down at the time. You can 't even get your
       | phone into the building without a security guard doing things to
       | it. Taking basic stuff like paper in/out is also disallowed.
       | 
       | I would have expected a total ban on personal computing devices
       | leaving the parking lot if this happened during my time there.
        
       | timcavel wrote:
       | [dead]
        
       | fullsend wrote:
       | [dead]
        
       | ftxbro wrote:
       | The article says "now in the wild after being leaked" but then it
       | says "the data is impossible to retrieve as it is now stored on
       | the servers belonging to OpenAI." So did the source code leak out
       | of OpenAI into the wild, or are they saying that OpenAI itself is
       | "the wild"? As far as I see from the article, it's not accessible
       | to the general public.
        
         | brundolf wrote:
         | If ChatGPT is trained on the data and ChatGPT is accessible to
         | the general public, then the data may as well be accessible to
         | the general public
        
           | sterlind wrote:
           | it seems unlikely to me that ChatGPT is directly trained on
           | chat data. if it is, we should see it know information past
           | its knowledge cutoff. afaik that hasn't happened.
           | 
           | I assume the chat logs are instead training a reward model,
           | which itself is then used as the reward function during RLHF
           | training.
        
             | bigyikes wrote:
             | These models have a very long lead time before they're
             | released to the public. Maybe GPT-5 is being trained on
             | ChatGPT logs. I'm not sure we'd be able to detect if this
             | was happening.
        
       ___________________________________________________________________
       (page generated 2023-04-05 23:02 UTC)