[HN Gopher] OpenAI says it has evidence DeepSeek used its model ...
       ___________________________________________________________________
        
       OpenAI says it has evidence DeepSeek used its model to train
       competitor
        
       Author : timsuchanek
       Score  : 330 points
       Date   : 2025-01-29 04:21 UTC (18 hours ago)
        
 (HTM) web link (www.ft.com)
 (TXT) w3m dump (www.ft.com)
        
       | udev wrote:
       | https://archive.is/KiSYM
        
       | cratermoon wrote:
       | Ironic, OpenAI claiming someone else stole their work.
        
       | vinni2 wrote:
       | How would they prove they used it's model. I would be curious to
       | know their methodology. Also what legal actions OpenAI can take?
       | can DeepSeek be banned in US?
        
         | iforgot22 wrote:
         | They might show DeepSeek's model calling itself ChatGPT, which
         | users have already alleged. Same as how Cisco proved Huawei was
         | stealing router code.
         | 
         | Except in this case, nothing was stolen, unless they want to
         | call ChatGPT's own training on source data theft too.
        
           | freehorse wrote:
           | ChatGPT outputs are all over the internet. It is harder to
           | prove that deepseek used specifically o1 for training,
           | instead of a lot of chatgpt output ending up in the training
           | set from other sources.
        
             | iforgot22 wrote:
             | That's a good point, at least for the prompts I saw. Like
             | "do you have an app I can use" is commonly seen with
             | "here's the ChatGPT app" online. And maybe they don't add
             | anything telling Deepseek that it's Deepseek.
        
       | belter wrote:
       | The subtitle is the gold... : "White House AI tsar David Sacks
       | raises possibility of alleged intellectual property theft"
        
         | conartist6 wrote:
         | lolololololololol
        
       | vrighter wrote:
       | So what? They probably paid for api access just like everyone
       | else. So it's a TOS violation at worst. Go ahead, open a civil
       | suit in the US against an entity the US courts do not have
       | jurisdiction over and quit whining...
        
         | jhickok wrote:
         | >open a civil suit in the US against an entity the US courts do
         | not have jurisdiction over
         | 
         | Yeah, over a Chinese company no less.
        
       | ForHackernews wrote:
       | What's good for the goose is good for the gander. Obviously a
       | transformative work and not an intellectual property violation
       | any more than OpenAI injesting every piece of media in existence.
        
         | dagelf wrote:
         | Injesting is sure the right take. What a circus!
        
       | amarcheschi wrote:
       | I quite like a scenery where llm output can't be copyrighted, so
       | that it is possible to eventually train a llm with data from the
       | previous one(s)
        
         | layer8 wrote:
         | OpenAI argues it's a violation of their terms of service. So
         | there are legal issues if it can be proven.
        
           | mannewalis wrote:
           | But OpenAI's model isn't open source, how would they distill
           | knowledge without direct access to the model?
        
             | layer8 wrote:
             | You don't need direct access for LLM distillation, just
             | regular API access.
        
               | mannewalis wrote:
               | ok I looked it up and have a better understanding now.
        
           | Palmik wrote:
           | Legal issues for who?
           | 
           | Company A pays OpenAI for their API. They use the API to
           | generate or augment a lot of data. They own the data. They
           | post the data on the open Internet.
           | 
           | Company B has the habit of scraping various pages on the
           | Internet to train its large language models, which includes
           | the data posted by Company A. [1]
           | 
           | OpenAI is undoubtedly breaking many terms of service and
           | licenses when it uses most of the open Internet to train its
           | models. Not to mention potential copyright violations (which
           | do not apply to AI outputs).
           | 
           | [1]: This is not hypothetical BTW. In the early days of LLMs,
           | lots of large labs accidentally and not so accidentally
           | trained on the now famous ShareGPT dataset (outputs from
           | ChatGPT shared on the ShareGPT website).
        
             | layer8 wrote:
             | For both.
        
               | Palmik wrote:
               | Posting OpenAI generated data on the internet is not
               | breaking the ToS. This is how most OpenAI based
               | businesses operate, after all [1] (e.g. various
               | businesses to generate articles with AI, various chat
               | businesses that let you share your chats, etc.)
               | 
               | OpenAI is one of the companies like Company B that is
               | using data from the open Internet.
               | 
               | [1] Ownership of content. As between you and OpenAI, and
               | to the extent permitted by applicable law, you (a) retain
               | your ownership rights in Input and (b) own the Output. We
               | hereby assign to you all our right, title, and interest,
               | if any, in and to Output.
        
       | top_sigrid wrote:
       | https://archive.is/KiSYM
        
       | lawlessone wrote:
       | So they're mad someone did exactly what they did?
        
         | exe34 wrote:
         | no, no, it's completely different. "open"AI stole from poor
         | people. DeepSeek stole from a $1T company. that's illegal!
        
       | whatshisface wrote:
       | It's reasonably likely that a lot of people linked to the federal
       | government want to ban DeepSeek. You can tell it's being
       | presented away from "they gave us a free set of weights" and
       | towards "they destroyed $1T of shareholder value." (By revealing
       | that Microsoft et al. paid way too much to OpenAI et al. for
       | technology that was actually easy to reinvent.)
        
         | fullshark wrote:
         | Would it even matter? Isn't the cat out of the bag and
         | everything they did repeatable by an American research team?
        
           | Cumpiler69 wrote:
           | It matters because their goal was hyping up how advanced and
           | difficult their tech is, propping up their valuations.
           | 
           | DeepSeek proved the emperor had no clothes and wiped out a
           | lot of their valuation when investors saw reaching parity to
           | Chtgpt is not really that difficult.
        
             | mastazi wrote:
             | I think parent was asking would it even matter if there was
             | a ban. To which the answer would be "no" because as you
             | said the point has been made. And, as parent pointed out,
             | it's repeatable anyway.
        
           | bhouston wrote:
           | It doesn't matter from the US government perspective if all
           | of the tech is replicated by US companies and US user
           | continue to use US AI technology. But if US users start to
           | use Chinese AI tech, then protectionism urges will appear
           | that will likely figure out how to ban its use or subject it
           | to large tariffs (e.g. TikTok, BYD, network equipment, solar
           | panels, etc.)
        
           | whatshisface wrote:
           | American researchers had already made enough progress to
           | prove that LLMs were not an incomprehensible trade secret
           | based on years of secret knowledge - investors and tech
           | executives were simply lead to believe otherwise. Well-
           | connected people are probably very mad about this and they
           | may try to lash out like the emotional human beings they are.
        
         | Cumpiler69 wrote:
         | _> By revealing that Microsoft et al. paid way too much to
         | OpenAI et al. for technology that was actually easy to
         | reinvent._
         | 
         | That's why it's called a bubble. Pretty sure my great great
         | grandad also overpaid for some tulips.
        
         | toomuchtodo wrote:
         | > "they destroyed $1T of shareholder value." (By revealing that
         | Microsoft et al. paid way too much to OpenAI et al. for
         | technology that was actually easy to reinvent.)
         | 
         | The value was highly speculative, an illusion created by PR and
         | sentiment momentum. "Hype value" not real value (unless you're
         | able to realize it and dump those bags on someone else before
         | fundamentals set in). Same thing happening with power companies
         | downstream of the discovery that AI is not going to be a savior
         | of sagging electricity demand. Overdriving the fundamentals is
         | not value destruction, it is "I gambled and lost."
         | 
         | https://www.bloomberg.com/news/articles/2025-01-28/deepseek-...
         | | https://archive.today/mCemf
         | 
         | "In the short run, the market is a voting machine but in the
         | long run, it is a weighing machine."
        
           | cft wrote:
           | Since the time when companies en masse stopped paying cash
           | dividends on owned shares, the value has become highly
           | speculative. In the absence of dividend payments, the stock
           | pricing mechanism is not essentially different from Solana or
           | Ethereum "price" discovery.
        
             | toomuchtodo wrote:
             | I don't disagree that price discovery is harder, but I can
             | with more certainty give an honest valuation of CLF or DOW
             | vs OpenAI's "who knows what money will look like after we
             | succeed, you should view your investment as a donation"
             | nonsense. Speculation is inevitable when forward looking,
             | but there is a difference between error bars and various
             | projections vs unicorns.
             | 
             | Due diligence never goes out of style.
        
             | JumpCrisscross wrote:
             | > _when companies en masse stopped paying stock dividends_
             | 
             | Do you mean cash dividends [1]?
             | 
             | Also, the premise is false. Dividend yields have roughly
             | tracked interest rates [2]. (The difference is a dirty
             | component of the equity risk premium [3].)
             | 
             | [1] https://www.investopedia.com/ask/answers/05/stockcashdi
             | viden...
             | 
             | [2] https://www.multpl.com/s-p-500-dividend-yield/table/by-
             | year
             | 
             | [3] https://www.investopedia.com/investing/calculating-
             | equity-ri...
        
               | cft wrote:
               | I changed the typo, thanks. Chash dividends. This
               | analysis does not negate common sense: when a company
               | does not pay cash dividends, owning its stock is purely
               | speculative, like owning Solana. When it does, you get
               | cash dividends funded by the company's tangible revenue,
               | proportional to your number of shares.
        
           | DebtDeflation wrote:
           | What they really destroyed was the idea that OpenAI would be
           | able to charge $200/month for their ChatGPT Pro subscription
           | which includes o1. That was always ridiculous IMO. The Free
           | tier and $20/month Plus tier along with their API business
           | (minus any future plan to charge a ridiculous amount for API
           | access to o1) will be fine.
        
             | toomuchtodo wrote:
             | > The Free tier and $20/month Plus tier along with their
             | API business (minus any future plan to charge a ridiculous
             | amount for API access to o1) will be fine.
             | 
             | Do the unit economics make this sustainable?
        
               | DebtDeflation wrote:
               | If only there were a way to make the models more
               | efficient. Oh wait.
        
               | jl6 wrote:
               | But doesn't Deepseek's innovation apply only to training,
               | not inference?
        
               | Zacharias030 wrote:
               | Actually no! If we take their paper at face value, the
               | crucial innovation to get a strong model with efficiency
               | is their much reduced KV cache and their MoE approach: -
               | where a standard model needs to store two large vectors
               | for each token at inference time (and load/store those
               | over and over from memory) deepseek v3/R1 only stores one
               | smaller vector C that is a ,,compression" from which the
               | large k,v vectors can be decoded on the fly. - They use a
               | fairly standard Mixture of Expert (MoE) approach, which
               | works well in training with their tricks, but whose
               | inference time advantages are immediate and equal to all
               | other MoE techniques, which is to say that from ~85% of
               | the 600B+ params that are inside the MoE layers, the
               | model at each token inference step will only pick a small
               | fraction to use. This reduces FLOPs and memory io by a
               | large factor in comparison to a so-called dense model
               | where all weights are used for every token (cf Llama 3
               | 405B)
        
               | freeone3000 wrote:
               | Reducing R&D expense also reduces breakeven price.
        
           | scarface_74 wrote:
           | The two podcasters who do the Acquired podcast spoke to
           | Ballmer about some of Microsoft's failed initiatives and
           | acquisitions. He told them that at the end of the day "it's
           | only money".
           | 
           | All of the BigTech companies have enough cash flow from
           | profitable lines of business to make speculative bets.
        
             | azemetre wrote:
             | It must be EZ mode to be a big tech executive, you somehow
             | have all the power to make every decision while also having
             | the ability to never take the fault for these decisions.
        
               | scarface_74 wrote:
               | I would much rather have a company with a culture that
               | isn't afraid to take calculated risks and not be afraid
               | of repercussions when they take risk as long as it
               | doesn't cause consumer harm.
        
               | azemetre wrote:
               | "Not doing consumer harm" is carrying a lot of weight
               | there.
               | 
               | Either way what you describe is perfectly achievable for
               | the workers, but at some point management needs to own up
               | to their failures and getting rewarded because the board
               | is also made up of executives at other big tech companies
               | is a perverse incentive to never actually improve.
        
               | scarface_74 wrote:
               | How did Microsoft's losing bets do consumer harm?
        
               | azemetre wrote:
               | I mean forcing copilot everywhere I don't want it
               | (nowhere) while jacking up prices to justify it and using
               | Windows 11 to serve ads is harmful to me. There's also
               | you know... the anticompetitive company that thinks
               | buying new sectors is healthy.
        
               | scarface_74 wrote:
               | Today, Microsoft's revenue mostly comes from Office and
               | Azure. All except PowerPoint were written and designed by
               | MS.
        
         | onlyrealcuzzo wrote:
         | Theoretically this should be good for OpenAI - in that they can
         | reduce their costs by ~27x and pass that along to end users to
         | get more adoption and more profit.
        
           | ceejayoz wrote:
           | No; those costs were their moat.
        
             | onlyrealcuzzo wrote:
             | You don't need a moat when you're in first place.
             | 
             | Their moat is >1B people are already using ChatGPT monthly.
             | 
             | They aren't going to switch unless something is
             | substantially better.
        
               | ceejayoz wrote:
               | > You don't need a moat when you're in first place.
               | 
               | Tell that to Friendster/MySpace and Facebook.
        
               | onlyrealcuzzo wrote:
               | Cute - but MySpace didn't have >1B users, it didn't even
               | have 10M when Facebook launched.
               | 
               | Try again.
        
               | cjbgkagh wrote:
               | It had more than Facebook when Facebook launched so I'm
               | not sure what your point is
        
               | ceejayoz wrote:
               | Nothing had a billion users; the _Internet_ didn 't at
               | the time.
               | 
               | MySpace and Friendster both spent significant time as the
               | #1 social sites. Facebook unseated them rapidly. The same
               | is possible for OpenAI.
        
               | shmeeed wrote:
               | Dude, if you seriously believe OpenAI has 1B active
               | users, you should go touch grass. Actual estimates are
               | 100-200 million, about a magnitude lower.
        
               | lm28469 wrote:
               | > They aren't going to switch unless something is
               | substantially better.
               | 
               | Except one product is 100% free and the other is mostly
               | locked behind paid subscriptions
        
               | scarface_74 wrote:
               | How long can DeepSeek stay free?
               | 
               | It's already unable to keep up with demand, it will never
               | be the default on mobile devices and businesses in the US
               | will never trust it.
        
               | ceejayoz wrote:
               | That's not really the important question.
               | 
               | The important question is "will this and similar
               | optimizations to come permit local LLM use, cutting
               | OpenAI out of the equation entirely?"
        
               | scarface_74 wrote:
               | Businesses don't even want to maintain servers locally.
               | They definitely aren't going to start managing servers
               | beefy enough to run LLMs and try to run then with the
               | reliability, availability, etc of cloud services.
               | 
               | This will make the cloud providers - especially AWS, GCP
               | and to a lesser extent the also ran clouds more valuable.
               | The other models hosted by AWS on Bedrock are already
               | "good enough" for most business use cases.
               | 
               | And then consumers are definitely not going to be running
               | LLMs locally on their computers to replicate ChatGPT (the
               | product) anymore than they are going to get an FTP
               | account, mount it locally with curlftpfs, and then using
               | SVN or CVS on the mounted filesystem and then from
               | Windows or Mac, accessed the FTP account through built-in
               | software instead of using cloud storage like Dropbox. [1]
               | 
               | Whether someone comes up with a better _product_ than
               | ChatGPT and overcome the brand awareness is yet to be
               | seen.
               | 
               | [1] Also the iPod had no wireless, less space than the
               | Nomad and was lame.
        
               | ceejayoz wrote:
               | > And then consumers are definitely not going to be
               | running LLMs locally on their computers to replicate
               | ChatGPT...
               | 
               | Not _personally_. They 'll let Apple handle it for them.
               | 
               | (This is already a thing.
               | https://machinelearning.apple.com/research/introducing-
               | apple...)
        
               | scarface_74 wrote:
               | There is a reason I kept emphasizing the ChatGPT
               | _product_. The (paid) ChatGPT product is not just a text
               | based LLM. It can interpret images, has a built in Python
               | runtime to offload queries that LLMs aren't good at like
               | math, web search, image generation, and a couple of other
               | integrations.
               | 
               | The local LLM on iPhones are literally 1% as powerful as
               | the server based models like 4o.
               | 
               | That's not even considering battery considerations
        
               | ceejayoz wrote:
               | > The local LLM on iPhones are literally 1% as powerful
               | as the server based models like 4o.
               | 
               | Currently, yes. That's why this is a compelling advance -
               | it makes local LLMs much more feasible, especially if
               | this is just the first of many breakthroughs.
               | 
               | A lot of the hype around OpenAI has been due to the fact
               | that buying enough capacity to run these things wasn't
               | all that feasible for competitors. Now, it is,
               | potentially even at the local level.
        
               | sandclock wrote:
               | That is exactly what a moat is. Keeping others out.
               | 
               | moat noun a deep, wide ditch surrounding a castle, fort,
               | or town, typically filled with water and intended as a
               | defense against attack.
        
               | ceejayoz wrote:
               | Precisely. First place _needs_ the moat.
               | 
               | Second place just needs a catapult and a diseased cow.
        
               | like_any_other wrote:
               | > Their moat is >1B people are already using ChatGPT
               | monthly.
               | 
               | Unlike a social network, network effects won't help them
               | - their users don't care how many other users they have,
               | only about the AI output quality.
               | 
               | > They aren't going to switch unless something is
               | substantially better.
               | 
               | Or approximately as good but cheaper.
        
               | onlyrealcuzzo wrote:
               | > Or approximately as good but cheaper.
               | 
               | You're fooling yourself if you think OpenAI is going to
               | pass up implementing the same strategies to get a ~27x
               | cheaper model.
               | 
               | > Unlike a social network, network effects won't help
               | them - their users don't care how many other users they
               | have, only about the AI output quality.
               | 
               | Google Search doesn't have a network effect. Everyone on
               | HN has been saying Google Search is complete garbage for
               | a decade. It still has the same market share (roughly) as
               | it did a decade ago.
        
               | like_any_other wrote:
               | > You're fooling yourself if you think OpenAI is going to
               | pass up implementing the same strategies to get a ~27x
               | cheaper model.
               | 
               | But that would mean a 27x lower valuation.
        
               | onlyrealcuzzo wrote:
               | > But that would mean a 27x lower valuation.
               | 
               | No.
               | 
               | Valuations are based on future profits. Not future
               | revenues.
               | 
               | You can theoretically lower your costs by 27x and end up
               | with 2x more future profits - if you're actually 45x
               | cheaper (which DeepSeek's method claims to be).
        
               | ceejayoz wrote:
               | > Valuations are based on future profits.
               | 
               | Which are estimated, in significant part, by the chance
               | of a competitor arising.
               | 
               | If the barriers of entry are much lower than originally
               | thought, the potential profit margin plummets.
        
               | like_any_other wrote:
               | You mean charge a 27x lower price, but have 45x lower
               | costs, so your profit margin has doubled?
               | 
               | Your _relative_ margin may have doubled, but your
               | absolute profit-per-item hasn 't. Say you had a 10%
               | margin before, at a $100 price and $90 cost, for a $10
               | profit-per-item. Reduce price 27x and cost 45x, so $3.7
               | price, $2 cost, and $1.7 profit-per-item. 6x less profit
               | - not as bad as 27x, but not good if you're OpenAI.
        
               | onlyrealcuzzo wrote:
               | > Your relative margin may have doubled, but your
               | absolute profit-per-item hasn't.
               | 
               | ChatGPT doesn't have any profits right now.
               | 
               | We have no idea what investors are expecting future
               | profits to be.
               | 
               | > Say you had a 10% margin before, at a $100 price and
               | $90 cost, for a $10 profit-per-item. Reduce price 27x and
               | cost 45x, so $3.7 price, $2 cost, and $1.7 profit-per-
               | item. 6x less profit - not as bad as 27x, but not good if
               | you're OpenAI.
               | 
               | Now do the same thing but assume you have 10x more
               | subscribers because the prices are ~27x lower.
               | 
               | You end up with almost 2x more total profit.
               | 
               | Just take ChatGPT's ~$200 subscription. Hardly anyone is
               | going to pay ~$200 a month. Reduce that by 27x - and
               | you're at $7.5 per month. Maybe 10% of people on the
               | planet will pay that.
        
               | ceejayoz wrote:
               | > Now do the same thing but assume you have 10x more
               | subscribers because the prices are ~27x lower.
               | 
               | You're in various spots of this thread pushing the idea
               | that their 1B MAUs make them unassailable. How are they
               | gonna get to 10B in a world with less than that total
               | people?
               | 
               | > Just take ChatGPT's ~$200 subscription. Hardly anyone
               | is going to pay ~$200 a month. Reduce that by 27x - and
               | you're at $7.5 per month. Maybe 10% of people on the
               | planet will pay that.
               | 
               | They can't even make money at the $200 price point,
               | though. https://x.com/sama/status/1876104315296968813
        
               | JumpCrisscross wrote:
               | > _that would mean a 27x lower valuation_
               | 
               | Not directly. The 27x is about costs. What it means is
               | some order of magnitude of more competition. _That_
               | reduces natural market share, price leverage and thus
               | future profits.
        
               | gtirloni wrote:
               | Other search engines don't have a gigantic advertising
               | budget or a dominant browser pounding on users' heads to
               | use them.
        
               | rurp wrote:
               | Google spends immense amounts of resources every year to
               | ensure that their search is almost always the default
               | option. Defaults are extremely powerful in consumer tech.
        
               | digitalPhonix wrote:
               | > Google Search doesn't have a network effect. Everyone
               | on HN has been saying Google Search is complete garbage
               | for a decade. It still has the same market share
               | (roughly) as it did a decade ago.
               | 
               | It absolutely does. People use Google for search ->
               | Websites optimise for Google -> People get "better"
               | results when searching with Google.
               | 
               | The fact that it's market share is sticky and not
               | responding quickly to change in quality is sort of
               | indicative of the network effect.
        
               | jasonjmcghee wrote:
               | 1 billion MAU? What's the source on that? Very difficult
               | to believe.
        
               | ceejayoz wrote:
               | It probably counts pretty much anyone on a newer
               | iPhone/Mac (https://support.apple.com/en-
               | au/guide/iphone/iph00fd3c8c2/io...) and Windows/Bing.
               | Plus all the smaller integrations out there. All of which
               | can be migrated to a new LLM vendor... pretty quickly.
               | 
               | I wonder what the _direct_ user counts are.
        
               | cogman10 wrote:
               | Free that runs locally on consumer hardware sounds
               | substantially better.
        
               | spinlock_ wrote:
               | I don't agree. You don't have moat if you are offering
               | the same quality for a higher price.
        
               | flavius29663 wrote:
               | just for the chatbot, it's trivial to switch, create a
               | new account and start asking questions from deepseek
               | instead. There is nothing holding the users in chatgpt.
        
               | ceejayoz wrote:
               | And the bigger risk is the big companies making deals -
               | like Apple including ChatGPT access in iOS - canceling
               | those to do it on-device or in-house.
               | 
               | 1B MAUs doesn't look great if half of them come from one
               | source that can easily change to a competitor.
        
               | JumpCrisscross wrote:
               | > _You don 't need a moat when you're in first place_
               | 
               | There are different moats [1]. You're describing
               | incumbency, an intangible moat. It's nice, but it's
               | fickle. Particularly with something with low switching
               | costs.
               | 
               | OpenAI could argue, before, that it had a natural
               | monopoly. More people use OpenAI so it gets more revenue
               | and more data which lets it raise more capital to train
               | these expensive models. That may not be true, which means
               | it only has that first, shallow moat. It's Nike. Not
               | Google.
               | 
               | [1] https://en.m.wikipedia.org/wiki/Economic_moat
        
               | onlyrealcuzzo wrote:
               | > There are different moats [1]. You're describing
               | incumbency, an intangible moat. It's nice, but it's
               | fickle. Particularly with something with low switching
               | costs.
               | 
               | Google has a low switching cost, and hardly anyone
               | switches.
               | 
               | ChatGPT is quite similar to Google in this way.
        
               | JumpCrisscross wrote:
               | > _Google has a low switching cost, and hardly anyone
               | switches_
               | 
               | Google has massive network effects on its ad business and
               | a natural monopoly on its search index. Crawling the web
               | is expensive. It's why Kagi has to pay Google (versus
               | being able to pay them once and then stop).
        
               | scarface_74 wrote:
               | Thought experiment: if tomorrow Apple changed the default
               | search engine from Google to ChatGPT for iOS, how fast
               | would Google's dominance drop?
               | 
               | iOS has 70% market share in the US
        
               | kgwgk wrote:
               | https://archive.is/c6cn9
               | 
               | << The thing I noticed right away when Claude came out is
               | how little lock-in ChatGPT had established. This was very
               | different to my experience when I first ran a search on
               | Google, sometime in the year 2000. After the first time I
               | used Google, I literally never used another search engine
               | again; it was just light years ahead of its competitors
               | in terms of the quality of its results, and the clarity
               | of its presentation. This week I added a third chatbot to
               | the mix: DeepSeek >>
               | 
               | Follow up:
               | https://x.com/TheStalwart/status/1884606421225848889
        
               | mohsen1 wrote:
               | My guess is that OS vendors are the real winners in the
               | long run. If Siri/Goolge can access my stuff and core of
               | LLMs is this replicable then I don't see anyone
               | downloading any apps for their typical AI usage.
               | Specially that users have to go out of their way to allow
               | a 3rd party to access all their data.
               | 
               | This is why OpenAI is so deep in the product development
               | phase right now. They have to become the OS to be
               | successful but I don't see that happening
        
               | Sateeshm wrote:
               | There is no network effect (amazon, instagram, etc.) not
               | an enterprise vendor lock-in (Microsoft Office/AD, Apple
               | Appstore, etc.) In fact, it's quite the opposite, the way
               | these companies deliver ouput is damn near identical.
               | Switching between them is pretty painless.
        
               | meiraleal wrote:
               | oh well, I switched yesterday from a paid plan to a free
               | one and I'm quite happy with the quality improvement.
        
             | whatshisface wrote:
             | I wish more people had understood that spending a lot of
             | money processing publicly available commodities with
             | techniques available in the published literature is the
             | business model of a steel mill.
        
               | JumpCrisscross wrote:
               | > _is the business model of a steel mill_
               | 
               | It's the business of commodities. The magic is in tiny
               | incremental improvements and distribution. DeepSeek
               | forces us to question if AI--possibly intelligence--is a
               | commodity.
        
               | sebzim4500 wrote:
               | Surely that would be amazing for NVDA? If the only 'hard'
               | part of making AI is making/buying/smuggling the hardware
               | then nvidia should expect to capture most of the value.
        
               | ceejayoz wrote:
               | DeepSeek revealed it's not as hard as previously thought;
               | a much smaller number of less sophisticated chips was
               | sufficient.
        
               | JumpCrisscross wrote:
               | > _that would be amazing for NVDA?_
               | 
               | It's good for Nvidia. It's not as good as it was before.
               | (Assuming DeepSeek's claims are replicable.)
        
               | JoshTko wrote:
               | No. Before Deepseek R1, Nvidia was charging $100 for a
               | $20 shovel in the gold rush. Now, every Fortune 100 can
               | build an O1-level model with currently existing (and soon
               | to be online) infra. Healthy demand for H100 and
               | Blackwell will remain, but paying $100 for a $20 shovel
               | is unlikely.
               | 
               | Nvidia will definitely stay profitable for now though, as
               | long as Deepseek's breakthroughs are not further improved
               | upon. But if others find additional compression gains,
               | Nvidia won't recapture its old premium. Its stock hinged
               | on 80% margins and 75% annual growth, Deepseek broke that
               | premise.
        
               | wongarsu wrote:
               | There still isn't a serious alternative for chips for AI
               | training. Until competition catches up or models become
               | so efficient they can be trained on gaming cards Nvidia
               | will still be able to command the same margins.
               | 
               | Growth might take a short-term dip, but may well be
               | picked up by induced demand. Being able to train your own
               | models "cheaply" will cause a lot more companies and
               | departments want to train their own models on their own
               | data, and cause them to retrain more frequently.
               | 
               | The time of being able to sell H100 clusters for
               | inference might be coming to an end though.
        
               | throwaway48476 wrote:
               | NVDA is too invested in training and underinvested in
               | edge inference.
        
               | camdenreslink wrote:
               | I'm not sure I'd call what LLMs do intelligence. Not yet
               | anyway...
        
               | JumpCrisscross wrote:
               | > _not sure I'd call what LLMs do intelligence_
               | 
               | No, but it's good enough to replace some office jobs.
               | Which forces us to ask, to what degree is intelligence--
               | unique intelligence--required for useful production? (We
               | can ask the same about physical strength.)
        
               | FridgeSeal wrote:
               | I find it interesting that so much discussion about
               | "LLM's can do some of our work" is centred around "are
               | they intelligence" and not what I see as the precursor
               | question of "are we doing a lot of bullshit work?"
               | 
               | My partner is in law, along with several friends and the
               | amount of completely _useless_ work and ceremony they're
               | forced to do is insane. It's a literal waste of their
               | talent and time. We could probably net most of the
               | claimed AI gains by taking a serious look at pointless
               | workloads and come out ahead due to not needing the
               | energy and capital expenditure.
        
             | cjbgkagh wrote:
             | But have you heard of Jevons Paradon...... /s
             | 
             | OMG, it seems tech has been invaded by baaing crypto bros
        
             | scotty79 wrote:
             | Maybe they even suppressed algorithmic improvements in
             | their company to preserve moat. Something akin to Kodak
             | suppressing internal research on digital cameras because
             | they were world leading company that produced photo film.
        
           | wturner wrote:
           | Capitalism - a system where rational actors make informed
           | decisions.
        
           | blantonl wrote:
           | Nah, those costs were for their doomsday bunkers and crypto
           | purchases, and maybe a house or 3
        
           | askl wrote:
           | They could pivot to being a wrapper around DeepSeek. That
           | would also save a lot of R&D costs.
        
           | mirzap wrote:
           | Training costs are not the same as inference costs. DeepSeek
           | (or anyone hosting DS largest model) will still need a lot of
           | money and a bunch of GPU clusters to serve the customers.
        
           | btbuildem wrote:
           | > pass that along to end users
           | 
           | I don't think that's at all likely in the current economic
           | system
        
         | aprilthird2021 wrote:
         | Banning it will not bring back the value
        
         | duxup wrote:
         | When it comes to the executive branch's role in banning
         | something. I'm not convinced they're even honest about it /
         | what the context even is.
         | 
         | Trump wanted to ban Tiktok before... and then simply chose not
         | to / forgot about it.
         | 
         | Next round congress acted, and Trump delayed it and has said
         | that he is interested in his friends buying it.
         | 
         | Is there really a competitive plan here or is it just fishing
         | for payouts / grifting for allies?
         | 
         | The context is always about competition, but I'm not even sure
         | that's their plan.
        
         | the_sleaze_ wrote:
         | "easy to reinvent" often comes after "hard to invent"
        
           | mromanuk wrote:
           | At Microsoft's size, they don't care; they just buy out
           | others.
        
         | IAmGraydon wrote:
         | Yeah they're setting this up to ban it. Crazy that they think
         | this kind of approach will work in any way. Banning H100s
         | didn't work, and actually pushed them to innovate. Now someone
         | has found a more efficient way to train a model and they decide
         | the best way forward is for the US not to benefit from access
         | to it? This is clear evidence of collusion between OpenAI and
         | the US Government to disadvantage competitors. Beyond that, it
         | will never work. If they need to be reminded of just how little
         | power they have to control the distribution of open source
         | models, I think we would all be happy to enlighten them.
        
         | Buttons840 wrote:
         | I saw a some Europeans hoping that the US would ban DeepSeek,
         | because then there would be less traffic interfering with their
         | own DeepSeek queries.
         | 
         | The US can ban all they want, but if the rest of the world
         | starts preferring Chinese social media, Chinese AI, and Chinese
         | websites in general, the US is going to lose one of its crown
         | jewels.
         | 
         | The way the US behaves is a problem and makes a lot of people
         | prefer alternatives just for the sake of avoiding the US, which
         | is why it's important that the US get along with other nations,
         | but--well, about that...
        
           | nozzlegear wrote:
           | Agreed, you've highlighted one of the key problems with
           | protectionism and nativism. Banning competition just weakens
           | America's global influence, it doesn't make it stronger.
        
             | paxys wrote:
             | Plus technology cycles move so quickly that you won't have
             | to wait a generation or two to see the effects of this
             | isolationism.
        
             | sailfast wrote:
             | You know this. I know this. But the President of the United
             | States does not know this.
        
             | darkwizard42 wrote:
             | This statement doesn't seem to hold true. China has banned
             | nearly all US tech companies and social products. It has
             | not decreased the influence of China's influence (which has
             | been through manufacturing/retail influence and tech
             | influence).
             | 
             | I don't think your statement holds with current behavior.
        
               | nozzlegear wrote:
               | But China has never been a global leader in tech or
               | social media. They undoubtedly have influence in these
               | areas, but they've never dominated them like the US has.
               | Banning foreign competition in a field where you
               | _already_ dominate, like tech and AI, has different
               | consequences than banning it where you 're playing catch
               | up.
        
               | philistine wrote:
               | TikTok has been the darling of the world for years at
               | this point. They're a global leader.
        
               | nozzlegear wrote:
               | Pretext my statement with "historically, until the last 5
               | years or so" and it still stands. TikTok is definitely
               | influential, there's no arguing that.
        
               | jononor wrote:
               | What is your definition of "tech"? A very large amount of
               | the electronics products in the world are made in China
               | (specifically in/around Szhenzen and the wider Guangdong
               | province). Both consumer goods and industrial goods. From
               | the cheapest stuff to the most advanced and everything in
               | between. They provide the manufacturing for brands fron
               | all over the world, including goods "from the west". The
               | amount of economy that depends entirely on this low-cost,
               | high-quality manufacturing is insanely large - both
               | directly in electronics goods but also as part of many
               | other industries because you need electronics to build
               | anything else.
        
               | kergonath wrote:
               | > China has banned nearly all US tech companies and
               | social products. It has not decreased the influence of
               | China's influence
               | 
               | Being hostile does not bring you friends. Sure, various
               | countries can have reasons to suck it up anyway (e.g.
               | because of sanctions, or because China makes an offer too
               | good to pass, although even that comes with strings
               | attached). But in the long run you just create clients or
               | satellites who will escape at the first occasion.
               | 
               | The American foreign policy around the middle of the 20th
               | century relied very effectively on soft power, which is
               | something you can leverage to get much more out of your
               | investments than their pure monetary value. It is not
               | required in order to gain influence, but it is a force
               | multiplier.
        
               | philistine wrote:
               | Then how can you explain that China's hostility towards
               | Western tech companies being present inside their own
               | country has not created what you're describing?
               | 
               | Is hostility a bad idea only for America? Sure hope not.
        
               | freeone3000 wrote:
               | America is reliant on purchasing cheap goods from
               | elsewhere and selling expensive technology. If it's
               | hostile toward the suppliers of cheap goods or the buyers
               | of expensive technology, well, what purpose does it have
               | on the global scale?
        
               | kergonath wrote:
               | I am saying that they could have got much more,
               | particularly considering the spectacular mistakes western
               | countries kept making for the last ~2 decades.
        
               | nozzlegear wrote:
               | > Is hostility a bad idea only for America? Sure hope
               | not.
               | 
               | I think protectionism is long-term bad for every country,
               | but it's especially and uniquely bad for the biggest
               | economy in the world who has net benefitted the most from
               | free trade and competition. There's no denying that China
               | is influential - the argument is that they could've been
               | (and still can be) so much more influential by embracing
               | western tech instead of walling themselves off.
        
           | karel-3d wrote:
           | EU will ban DeepSeek sooner because of (lack of) GDPR
           | compliance
        
             | whatevaa wrote:
             | That's ok, they provide the models. They can be run at
             | Europe, given you have capable hardware.
        
               | wongarsu wrote:
               | And the path to pleasing the EU would be straightforward:
               | make a EU subsidiary, have it host or rent GPUs and
               | servers in Europe, make sure personally-identifiable data
               | is handled in accordance with GDPR and doesn't leave that
               | subsidiary, make sure EU customers make their accounts
               | with and are served by that subsidiary.
               | 
               | Meanwhile, to please the US they would probably have to
               | move the entire company to the US. And even that may not
               | be enough
        
             | acheong08 wrote:
             | Banning the site would be fine. The model itself will still
             | be available from a variety of providers as well as
             | locally. The US is more likely to ban the model itself on
             | the basis of national security
        
             | Buttons840 wrote:
             | Does EU _block_ websites that don 't comply with their
             | laws?
             | 
             | If DeepSeek becomes popular in America I predict it will be
             | blocked, national firewall style. Will EU do the same?
        
               | surgical_fire wrote:
               | Generally no. For all people complaining about EU
               | regulations, the regulators typically opt to fine
               | companies into compliance.
        
           | tensor wrote:
           | I've recently cancelled my Github Copilot subscription and
           | now use Mistral. When the US starts threatening allies with
           | tariffs or invasion, using US services becomes a major
           | business risk.
        
             | mongol wrote:
             | Not only a business risk. It also becomes a moral
             | imperative to avoid if you can. Don't support bullies, is
             | my motto. It can be hard to completely avoid, but it is
             | important to try.
        
         | dkjaudyeqooe wrote:
         | Not sure I agree with your premise, but what exactly are they
         | going to ban?
         | 
         | They can stop DeepSeek doing various things commercially I
         | guess, but stopping Americans using their ideas is simply
         | impossible and stopping use of their source or weights would be
         | (likely successfully) challenged under the first amendment.
         | 
         | There is no law against simply destroying trillions of dollars
         | of shareholder value.
        
           | thedevilslawyer wrote:
           | Heh, not yet..
        
         | dfxm12 wrote:
         | There's a lot of egg on people's faces now. DeepSeek shows
         | there's nothing special about America or its economic system
         | that breeds innovation. DeepSeek shows how these tech oligarchs
         | greatly overplayed their hand and along with the president,
         | bamboozled the taxpayer to enrich each other. I just hope the
         | voters remember this in 2 years, 4 years and beyond.
        
           | dtquad wrote:
           | >DeepSeek shows how these tech oligarchs greatly overplayed
           | their hand and along with the president, bamboozled the
           | taxpayer to enrich each other.
           | 
           | How much taxpayer money has gone to OpenAI and Anthropic?
           | They are the two big sinners in closed AI.
        
           | iforgot22 wrote:
           | If the allegations are true, the special thing about OpenAI
           | is that it didn't have to be trained off DeepSeek. But either
           | way, you maybe don't want to invest billions in something if
           | someone else will be able to copy it for less.
        
         | dtquad wrote:
         | >It's reasonably likely that a lot of people linked to the
         | federal government want to ban DeepSeek.
         | 
         | It took them years and years to move forward with the ban ok
         | Tiktok and it still hasn't been banned yet. There is no way
         | they are going to ban some MIT-licensed weights.
         | 
         | >"they destroyed $1T of shareholder value."
         | 
         | The market has largely recovered.
        
           | tokioyoyo wrote:
           | There is big American money invested in TikTok. That doesn't
           | seem to be the case for DeepSeek.
        
         | leesec wrote:
         | It was so easy it cost hundreds of millions of dollars, only
         | one company has done it and they had to lie about it
        
         | nullbyte wrote:
         | I think the real concern from the govt's perspective is data
         | privacy, since all the chat messages are stored on Chinese
         | servers
        
       | jasoneckert wrote:
       | What I find the most comical about this is that the whole
       | situation could be loosely summarized as "OpenAI is losing its
       | job to AI."
        
         | zbshqoa wrote:
         | Realistically that's the actual headline. Only another AI can
         | replace AI, pretty much like LLMs / Transformers have replaced
         | "old" AI models in certain task (NLP, Sentiment Analysis,
         | Translation etc) and research is in progress for other tasks as
         | well performed by traditional models (personalization,
         | forecasting, anomaly detection etc).
         | 
         | If there's a better AI, old AI will lose the job first.
        
           | troyvit wrote:
           | > NLP, Sentiment Analysis, Translation etc
           | 
           | As somebody who got to work adjacent to some of these things
           | for a long time, I've been wondering about this. Are LLMs and
           | transformers actually better than these "old" models or is it
           | more of an 80/20 thing where for a lot less work (on
           | developers' behalf) LLMs can get 80% of the efficacy of these
           | old models?
           | 
           | I ask because I worked for a company that had a related
           | content engine back in 2008. It was a simple vector database
           | with some bells and whistles. It didn't need a ton of
           | compute, and GPUs certainly weren't what they are today, but
           | it was pretty fast and worked pretty well too. Now it seems
           | like you can get the same thing with a simple query but it
           | takes a lot more coal to make it go. Is it better?
        
             | ang_cire wrote:
             | Yep, it's an 80/20 thing. Versatility over quality.
        
         | Keyframe wrote:
         | also, China doing in IP what it's better at and way more
         | experienced than USA - stealing.
        
           | rchaud wrote:
           | This kind of blithe commentary is 20 years out of date and
           | reminiscent of 1970s criticisms of the Japanese car industry.
        
             | Buttons840 wrote:
             | I'm reminded of an Adam Savage video. He ordered an unusual
             | vise from China, and he praised their culture where someone
             | said "I want to build this strange vise that wont be super
             | popular", and the boss said "cool, go do it". They built a
             | thing that we would not build in America.
             | 
             | https://youtu.be/NUhrF0xkhhc?si=1WHWYZrhRmfOYO_y&t=1150
             | (it's about 2 minutes)
        
               | pphysch wrote:
               | The small biz scene in unfree communist China is
               | ironically, astronomically better than here in US, where
               | decades of regulatory capture and misleadership have made
               | it difficult and extremely expensive to get off the
               | ground while being protected by the law.
        
             | Keyframe wrote:
             | gentle stroll through the aliexpress alleyway tells
             | otherwise.
        
           | t43562 wrote:
           | Who says America is less good at it? Hasn't the US nicked a
           | lot of other people's ideas at some point or other?
        
         | mattgreenrocks wrote:
         | OpenAI should be excited that it has been freed of the tedious
         | tasks of building AI and now they can focus on higher level and
         | more creative things.
        
           | Sateeshm wrote:
           | > focus on higher level and more creative things.
           | 
           | But that's what OpenAI's costumers were supposed to do.
        
             | jusonchan81 wrote:
             | It's sarcasm.
        
               | rooroobooragool wrote:
               | I think Sateeshm was also applying a generous layer of
               | sarcasm.
        
           | pphysch wrote:
           | OpenAI should be, but OpenAI died a while ago
        
           | JoshTko wrote:
           | I wish I could upvote this twice
        
             | munchler wrote:
             | Soon you'll be freed of the tedious task of upvoting at
             | all.
        
         | blantonl wrote:
         | Otherwise known as a race to the bottom
        
         | bwfan123 wrote:
         | ha, the story is filled with ironies.
         | 
         | OpenAIs $200 closed-ai uppended by hedge-funds free side-
         | project
         | 
         | Quant geeks outcompete overpaid silicon valley devs etc.
         | 
         | Basically, hubris gets its comeuppance which is a david vs
         | goliath biblical archetype which is why this drama grips all of
         | us.
        
           | jeffreyq wrote:
           | seems ironic that the turns have tabled. "silicon valley
           | devs" were the analogous "quant geeks" underdogs that
           | unseated the ossified incumbents.
           | 
           | That said, I feel like "quant geeks" aren't quite underdogs
           | compared to silicon valley devs. wdyt?
        
         | rooroobooragool wrote:
         | This is really the top take in this thread. Why should OpenAI
         | be any different than all the others they they've ripped off.
        
         | nikeee wrote:
         | More like
         | 
         | "OpenAI is losing its job to open AI."
        
       | semking wrote:
       | This is absolutely hilarious! :)
       | 
       | ClosedAI scraped human content without asking and they explained
       | why this was acceptable... but when the outputs of their training
       | corpus is scraped, it is THEIR dataset and this is NOT
       | acceptable!
       | 
       | Oh, the irony! :D
       | 
       | I shared a few screenshots of DeepSeek answering using ChatGPT's
       | output in yesterday's article!
       | 
       | https://semking.com/deepseek-china-ai-model-breakthrough-sec...
        
         | marricks wrote:
         | Also, DeepSeek is allegedly... better? So saying they just
         | copied ClosedAI isn't really sufficient of an answer. Seems to
         | be just bluster because the US Govt would probably accept any
         | excuse to ban it, see TikTok.
        
           | semking wrote:
           | I never said they are just a clone! There's an actual tech
           | breakthrough!
           | 
           | Read the two following sections of my blog post:
           | 
           | 1. "Distilled language models"
           | 
           | 2. "DeepSeek: Less supervision"
        
           | beAbU wrote:
           | How can they ban something thats open source that you can
           | just run on your own hardware?
        
             | Drakim wrote:
             | They banned certain branches of math during the cold war,
             | it can be done.
        
               | jerry80 wrote:
               | Such as?
        
             | shafyy wrote:
             | It's not open source. The provide the model and the
             | weights, but not the source code and, crucially, the
             | training data. As long as LLM makers don't provide the
             | training data (and they never will, because then they will
             | be admitting to stealing), LLMs are never going to be open
             | source.
        
               | sho_hn wrote:
               | Thanks for reminding people of this.
               | 
               | Open source means two things in spirit:
               | 
               | (a) You have everything you need to be able to re-create
               | something, and at any step of the process change it.
               | 
               | (b) You have broad permissions how to put the result to
               | use.
               | 
               | The "open source" models from both Meta so far fail
               | either both or one of these checks (Meta's fails both).
               | We should resist the dilution of the term open source to
               | the point where it means nothing useful.
        
               | jprete wrote:
               | I think people are looking for the term "freeware"
               | although the connotations don't match.
        
               | sho_hn wrote:
               | Agreed, but the "connotations don't match" is mostly
               | because the folks who chose to call it open source wanted
               | the marketing benefits of doing so. Otherwise it'd match
               | pretty well.
        
               | HDThoreaun wrote:
               | Open source means the source code is freely available.
               | It's in the name.
        
               | idle_zealot wrote:
               | The source being available means the code is "source
               | available." Open implies more rights.
        
               | KPGv2 wrote:
               | At the risk of being called rms, no, that's not what open
               | source means. Open source just means you have access to
               | the source code. Which you do. Code that is open source
               | but restrictively licensed is still open source.
               | 
               | That's why terms like "libre" were born to describe
               | certain kinds of software. And that's what you're
               | describing.
               | 
               | This is a debate that started, like, twenty years ago or
               | something when we started getting big code projects that
               | were open source but encumbered by patents so that they
               | couldn't be redistributed, but could still be read and
               | modified for internal use.
        
               | sho_hn wrote:
               | > Open source just means you have access to the source
               | code. Which you do.
               | 
               | No, they also fail even that test. Neither Meta nor
               | DeepSeek have released the source code of their training
               | pipeline or anything like that. There's very little
               | literal "source code" in any of these releases at all.
               | 
               | What you _can_ get from them is the model weights, which
               | for the purpose of this discussion, is very similar to
               | compiler binary executable output you cannot easily
               | reverse, which is what open source seeks to address. In
               | the case of Meta, this comes with additional usage
               | limitations on how you may put them to use.
               | 
               | As a sibling comment said, this is basically "freeware"
               | (with asterisks) but has nothing to do with open source,
               | either according to RMS or OSI.
               | 
               | > This is a debate that started, like, twenty years ago
               | 
               | For the record, I do appreciate the distinction. This
               | isn't meant as an argument from authority at all, but
               | I've been an active open source (and free software)
               | developer for close to those 20 years, am on the board of
               | one of the larger FOSS orgs, and most households have a
               | few copies of FOSS code I've written running. It's also
               | why I care! :-)
        
               | JumpCrisscross wrote:
               | > _they also fail even that test. Neither Meta nor
               | DeepSeek have released the source code of the_
               | 
               | This debate is over and makes the open source community
               | look silly. Open model and weights is, practically
               | speaking, open source for LLMs.
               | 
               | I have tremendous respect for FOSS and those who build
               | and maintain it. But arguing for open training data means
               | only toy models can practically exist. As a result, the
               | practical definition will prevail. And if the only people
               | putting forward a practical definition are Meta _et al_ ,
               | this is what you get: source available.
        
               | sho_hn wrote:
               | I'm not arguing for open training data BTW, and the
               | problem is exactly this sort of myopic focus on the
               | concerns of the AI community and the benefits of open-
               | washing marketing.
               | 
               | Completely, fully breaking the meaning of the term "open
               | source" is causing collateral damage _outside_ the AI
               | topic, that 's where it really hurts. The open source
               | principle is still useful and necessary, and we need
               | words to communicate about it and raise correct
               | expectations and apply correct standards. As a dev you
               | very likely don't want to live in a tech environment
               | where we regress on this.
               | 
               | It's not "source available" either. There's no _source_.
               | It 's freeware.
               | 
               | "I can download it and run it" isn't open source.
               | 
               | I'm actually not too worried that people won't eventually
               | re-discover the same needs that open source originally
               | discovered, but it's pretty lame if we lose a whole bunch
               | of time and effort to re-learn some lessons yet again.
        
               | JumpCrisscross wrote:
               | > _it 's pretty lame if we lose a whole bunch of time and
               | effort to re-learn some lessons yet again_
               | 
               | We need to relearn because we need a different definition
               | for LLMs. One that works in practice, not just at the
               | peripheries.
               | 
               | Maybe we can have FOSS LLMs vs open-source ones, like we
               | do with software licenses. The former refers to the
               | hardcore definition. The latter the practical (and widely
               | used) one.
        
               | sho_hn wrote:
               | Sure, I don't disagree. I fully understand the open-
               | weights folks looking for a word to communicate their
               | approach and its benefits, and I support them in doing
               | so. It's just a shame they picked this one in - and
               | that's giving folks a lot of benefit of the doubt - a
               | snap judgement.
               | 
               | > Maybe we can have FOSS LLMs vs open-source ones, like
               | we do with software licenses.
               | 
               | Why not just call them freeware LLMs, which would be much
               | more accurate?
               | 
               | There's nothing "hardcore" or "zealot" about not calling
               | these open source LLMs because there's just ...
               | absolutely nothing there that you call open source in any
               | way. We don't call _any_ other freeware  "open source"
               | for being a free download with a limited use license.
               | 
               | This is just "we chose a word to communicate we are
               | different from the other guys". In games, they chose to
               | call it "free to play (f2p)" when addressing a similar
               | issue (but it's also not a great fit since f2p games
               | usually have a server dependency).
        
               | JumpCrisscross wrote:
               | > _Why not just call them freeware LLMs, which would be
               | much more accurate?_
               | 
               | Most of the public is unfamiliar with the term. And with
               | some of the FOSS community arguing for open training
               | data, it was easy to overrule them and take the term.
        
               | sho_hn wrote:
               | Most of the public is also unfamiliar with the term open
               | source, and I'm not sure they did themselves any favors
               | by picking one that invites far more questions and needs
               | for explanation. In that sense, it may have accomplished
               | little but its harmful effects.
               | 
               | I get your overall take is "this is just how things go in
               | language", but you can escalate that non-caring
               | perspective all the way to entropy and the heat death of
               | the universe, and I guess I prefer being an element that
               | creates some structure in things, however fleeting.
        
               | JumpCrisscross wrote:
               | > _Most of the public is also unfamiliar with the term
               | open source_
               | 
               | I'd argue otherwise. (Familiar with, not know.)
               | Particularly in policy circles.
               | 
               | > _picking one that invites far more questions and needs
               | for explanation_
               | 
               | There wasn't ever a debate. And now, not even the OSI
               | demands training data. (It couldn't. It, too, would be
               | ignored.)
        
               | Flimm wrote:
               | The only practical and widely used definition of open
               | source is the one known as the Open Source Definition
               | published by the OSI.
               | 
               | The set of free/libre licenses (as defined by the FSF) is
               | almost identical to the set of open sources licenses (as
               | defined by the OSI).
               | 
               | The debate within FOSS communities has been between
               | copyleft licenses like the GPL, and permissive licenses
               | like the MIT licence. Both copyleft and permissive
               | licenses are considered free/libre by the FSF, and both
               | of them are considered open source by the OSI.
        
               | nuancebydefault wrote:
               | The weights, which are part of the source, are open. Now
               | you are arguing it not being open source because they
               | don't provide the source for that part of the source. If
               | you follow that reasoning you can ad infinitum claim the
               | absence of sources since every source originates from
               | something.
        
               | jefftk wrote:
               | _> Open source just means you have access to the source
               | code._
               | 
               | That's https://en.wikipedia.org/wiki/Source-
               | available_software , not 'open source'. The latter was
               | specifically coined [1] as a way to talk about "free
               | software" (with its freedom connotations) without the
               | price connotations:
               | 
               |  _The argument was as follows: those new to the term
               | "free software" assume it is referring to the price.
               | Oldtimers must then launch into an explanation, usually
               | given as follows: "We mean free as in freedom, not free
               | as in beer." At this point, a discussion on software has
               | turned into one about the price of an alcoholic beverage.
               | The problem was not that explaining the meaning is
               | impossible--the problem was that the name for an
               | important idea should not be so confusing to newcomers. A
               | clearer term was needed. No political issues were raised
               | regarding the free software term; the issue was its lack
               | of clarity to those new to the concept._
               | 
               | [1] https://opensource.com/article/18/2/coining-term-
               | open-source...
        
               | HDThoreaun wrote:
               | You dont get to redefine what "open" means.
        
               | jefftk wrote:
               | It's common for terms to have a more specific meaning
               | when combined with other terms. "Open source" has had a
               | specific meaning now for decades, which goes beyond "you
               | can see the source" to, among other things, "you're
               | allowed to it without restriction".
        
               | RobotToaster wrote:
               | So Swedish meatballs are any ball of meat made in Sweden?
               | 
               | And French fries are anything that was fried in France?
        
               | davidcbc wrote:
               | Tell that to Sam Altman
        
               | esafak wrote:
               | He did not succeed, did he?
        
               | dTal wrote:
               | I don't know why you've been downvoted. This is a 100%
               | correct history. "Open source" was specifically coined as
               | a synonym to "free software", and has always been used
               | that way.
        
               | beAbU wrote:
               | Thanks, I was not aware of this distinction.
               | 
               | But I think my argument still stands though? Users can
               | run Deepseek locally, so unless the US Gov't wants to
               | reach for book burning levels or idiocy, there is not
               | really a feasible way to ban the American public of
               | running DeepSeek, no?
        
               | shafyy wrote:
               | Yes, your argument still stands. But I think it's
               | important to stand firm that the term "open source" is
               | not a good label for what these "freeware" LLMs are.
        
               | beAbU wrote:
               | Fair point, agreed.
        
               | coliveira wrote:
               | People say this, but when it comes to AI models, the
               | training data is not owned by these companies/groups, so
               | it cannot be "open sourced" in any sense. And the
               | training code is basically accessing that training data
               | that cannot be open sourced, therefore it also cannot be
               | shared. So the full open source model you wish to have
               | can only provide subpar results.
        
               | sheepdestroyer wrote:
               | They could easily list the data used though. These
               | datasets are mostly known and floating around. When they
               | are constructed, instructions for replication could be
               | provided too
        
               | coliveira wrote:
               | They could, but even if they give this list the
               | detractors will still say it is not open source.
        
               | rvnx wrote:
               | yes and as a bonus they may get sued, which in the long-
               | term, makes free / offline models to not be viable
               | 
               | It would be so much better if all models were trained
               | with LibGen.
        
               | Timon3 wrote:
               | Isn't this the same situation that any codebase faces
               | when one thinks about open sourcing it? I can't legally
               | open source the code I don't own.
        
             | fabianhjr wrote:
             | There are illegal numbers in the USA land of the "free".
             | 
             | https://en.wikipedia.org/wiki/Illegal_number
             | 
             | > An AACS encryption key (09 F9 11 02 9D 74 E3 5B D8 41 56
             | C5 63 56 88 C0) that came to prominence in May 2007 is an
             | example of a number claimed to be a secret, and whose
             | publication or inappropriate possession is claimed to be
             | illegal in the United States.
        
               | JumpCrisscross wrote:
               | > _illegal numbers in the USA land of the "free"_
               | 
               | This is a silly take for anyone in tech. Any binary
               | sequence is a number. Any information can be, for
               | practical purposes, rendered in binary [1].
               | 
               | Getting worked up about restrictions on numbers works as
               | a meme, for the masses, because it sounds silly, but is
               | tantamount to technically arguing against privacy,
               | confidentiality, the concept of national secrets, IP as a
               | whole, _et cetera_.
               | 
               | [1] https://en.m.wikipedia.org/wiki/Shannon%27s_source_co
               | ding_th...
        
               | sheepdestroyer wrote:
               | All those things are not self-evident and thus debatable
        
               | JumpCrisscross wrote:
               | > _not self-evident and thus debatable_
               | 
               | Totally agree. But prompting debate or even further
               | thought isn't the point of the meme.
        
               | sheepdestroyer wrote:
               | I'd argue that, as satire, it's the main point ;)
        
               | JumpCrisscross wrote:
               | > _as satire, it 's the main point_
               | 
               | There is thought-stopping satire and thought-provoking
               | satire. Much of it depends on the context. I'm not
               | getting the latter from a "USA land of the 'free'"
               | comment.
        
               | fabianhjr wrote:
               | Good thing that is part of the wikipedia entry:
               | 
               | > Any piece of digital information is representable as a
               | number; consequently, if communicating a specific set of
               | information is illegal in some way, then the number may
               | be illegal as well.
        
               | bloopernova wrote:
               | That takes me back! Fark.com would delete any comment
               | that contained random hexadecimal.
        
               | KPGv2 wrote:
               | It was the beginning of the end for Digg, too, IIRC.
               | Started a lot of people leaving for Reddit, right?
        
               | bloopernova wrote:
               | I think so; I joined Reddit when it was in tech news as
               | people left Digg after the big redesign. I'm not sure
               | when the exodus started. I left Fark over the hd-dvd
               | mess.
        
               | KPGv2 wrote:
               | > whose publication or inappropriate possession is
               | claimed to be illegal in the United States.
               | 
               | That's not the same thing as a number being illegal at
               | all. Here, watch this:
               | 
               | > I claim breathing is illegal in the United States
               | 
               | There, now breathing is claimed to be illegal in the
               | United States.
        
               | I-M-S wrote:
               | In both cases, legality depends entirely on
               | repercussions, i.e. if there's someone to enforce the
               | ban. I suspect that in the "illegal numbers" case there
               | might be.
        
               | vluft wrote:
               | man that's very concerning for wikipedia who is
               | publishing it right there on the page linked above.
        
               | dylan604 wrote:
               | Only concerning if they are a US based company hosting
               | their data in US data centers. oops
        
               | suraci wrote:
               | > is collecting rain water illegal?
               | 
               | > It depends on where you live. In many places,
               | collecting rainwater is completely legal and even
               | encouraged, but some regions have regulations or
               | restrictions.
               | 
               | United States: Most states allow rainwater collection,
               | but some have restrictions on how much you can collect or
               | how it can be used. For example, Colorado has limits on
               | the amount of rainwater homeowners can store. Australia:
               | Generally legal and encouraged, with many homes using
               | rainwater tanks. UK & Canada: Legal with few
               | restrictions. India & Many Other Countries: Often
               | encouraged due to water scarcity.
        
             | superkuh wrote:
             | There was an executive order passed by the previous
             | administration that make using anything with more than 10
             | billion parameters illegal and punishable by government
             | force if done without authorization. Of course like most
             | government regulations (even though this is not a
             | regulation, it is an executive action) the point is not to
             | stop the behavior but instead to create a system where
             | everyone breaks the regulation constantly so that if anyone
             | rocks the boat they can be indicted/charged and dealt with.
             | 
             | https://www.federalregister.gov/documents/2023/11/01/2023-2
             | 4...
             | 
             | >(k) The term "dual-use foundation model" means an AI model
             | that is trained on broad data; generally uses self-
             | supervision; contains at least tens of billions of
             | parameters; is applicable across a wide range of contexts;
             | and that exhibits, or could be easily modified to exhibit,
             | high levels of performance at tasks that pose a serious
             | risk to security, national economic security, national
             | public health or safety, or any combination of those
             | matters, such as by: ...
        
               | ceejayoz wrote:
               | That order does not "make using anything with more than
               | 10 billion parameters illegal and punishable by
               | government force if done without authorization".
               | 
               | It orders the Secretary of Commerce to "solicit input
               | from the private sector, academia, civil society, and
               | other stakeholders through a public consultation process
               | on potential risks, benefits, other implications, and
               | appropriate policy and regulatory approaches related to
               | dual-use foundation models for which the model weights
               | are widely available".
        
               | derektank wrote:
               | Many regulations are created by executive action, without
               | input from Congress. The Council on Environmental
               | Quality, created by the National Environmental Policy
               | Act, has the power to issue it's own regulations.
               | Executive Orders can function similarly and the executive
               | can order rulemaking bodies to create and remove
               | regulations, though there is a judicial effort to
               | restrict this kind of policymaking and return regulatory
               | power back to Congress.
        
             | bilekas wrote:
             | If I'm no wrong wasn't PGP encryption once illegal to
             | export ? Not quite the same but the government has a nice
             | habit of feeling like they can bad the export of research.
             | 
             | https://en.wikipedia.org/wiki/Export_of_cryptography_from_t
             | h...
        
               | beAbU wrote:
               | You are right, but I cannot find a single example of such
               | a ban actually being effective though. Information wants
               | to be free and all that.
        
               | KPGv2 wrote:
               | Because you haven't heard of the proprietary software
               | that wasn't ever sold internationally because of these
               | bans.
               | 
               | Of course Joe Sixpack can throw their code up anywhere,
               | but Joe Corporation gets wrecked if they try to sell it.
               | 
               | https://developer.apple.com/documentation/security/comply
               | ing...
               | 
               | For example, this is enforced by Apple Store.
        
               | coliveira wrote:
               | But that's not the goal, the goal is to protect the
               | "intelectual property" only to American companies.
               | Countries not in the "friends list" cannot sell products
               | in that area without suffering repercussions. That's how
               | the US has maintained technological dominance in some
               | areas by restricting what other countries can do.
        
               | calgoo wrote:
               | If i remember correctly, if you changed the dropdown on
               | the webpage to USA you could download the full version of
               | PGP anyway.
        
               | Prbeek wrote:
               | Add PS1 too. The US government banned sale of PlayStation
               | to China because the PLA would apparently have access to
               | cutting edge chips for their missiles
        
             | michaelt wrote:
             | Make commercial hosting illegal, and make the hardware to
             | run it locally cost $6000+
        
           | throwup238 wrote:
           | It's not better. In most of my tests (C++/QT code) it just
           | runs out of context before it can really do anything. And the
           | output is very bad - it mashes together the header and cpp
           | file. The reasoning output is fun to look at and occasionally
           | useful though.
           | 
           | The max token output is only 8K (32K thinking tokens). O1 is
           | 128k, which is far more useful, and it doesn't get stuck like
           | R1 does.
           | 
           | The hype around the DeepSeek release is insane and I'm
           | starting to really doubt their numbers.
        
             | adamnemecek wrote:
             | Thanks for saying this, I thought I was insane, DeepSeek is
             | kinda bad. I guess it's impressive all things considered
             | but in absolute terms it's not great.
        
               | coliveira wrote:
               | I have run personal tests and the results are at least as
               | good as I get from OpenAI. Smarter people have also
               | reached the same conclusion. Of course you can find
               | contrary datapoints, but it doesn't change the big
               | picture.
        
               | sebzim4500 wrote:
               | To be fair, it's amazing by the standards of six months
               | ago. The only models that beat it are o1, the latest
               | gemini models and (for some things) sonnet 3.6
        
               | cdelsolar wrote:
               | false. It seems better than o1 to me.
        
             | gliptic wrote:
             | R1 is trained for a context length of 128K. Where are you
             | getting 8K/32K? The model doesn't distinguish "thinking"
             | tokens and "output" tokens, so this must be some specific
             | API limitations.
        
               | throwup238 wrote:
               | _> max_tokens:The maximum length of the final response
               | after the CoT output is completed, defaulting to 4K, with
               | a maximum of 8K. Note that the CoT output can reach up to
               | 32K tokens, and the parameter to control the CoT length
               | (reasoning_effort) will be available soon._ [1]
               | 
               | [1] https://api-docs.deepseek.com/guides/reasoning_model
        
               | gliptic wrote:
               | So yes, it's a limitation of their own API at the moment,
               | not a model limitation.
        
               | throwup238 wrote:
               | I'm using it through Kagi which doesn't use Deepseek's
               | official API [1]. That limitation from the docs seems to
               | be everywhere.
               | 
               | In practice I don't think anyone can economically host
               | the whole model plus the kv cache for the entire context
               | size of 128k (and I'm skeptical of Deepseek's claims now
               | anyway).
               | 
               | Edit: a Kagi team member just said on Discord that
               | they'll be increasing max tokens next release
               | 
               | [1] https://help.kagi.com/kagi/ai/llms-privacy.html
        
               | coliveira wrote:
               | He's just repeating a lot of disinformation that has been
               | released about deepseek in the last few days. People who
               | took the time to test DeepSeek models know that the
               | results have the same or better quality for coding tasks.
        
               | goosejuice wrote:
               | Benchmarks are great to have but individual/org
               | experiences on specific codebases still matter
               | tremendously.
               | 
               | If an org consistently finds one model performs worse on
               | their corpus than another, they aren't going to keep
               | using it because it ranks higher in some set of
               | benchmarks.
        
               | hn_throwaway_99 wrote:
               | But you should also be very wary of these kind of
               | anecdotes, and this thread highlights exactly why. That
               | commenter says in another comment
               | (https://news.ycombinator.com/item?id=42866350) that the
               | token limitation that he is complaining about has
               | actually nothing to do with DeepSeek's model or their
               | API, but is a consequence of an artificial limit that
               | _Kagi_ imposes. In other words, his conclusion about
               | DeepSeek is completely unwarranted.
        
               | throwup238 wrote:
               | It mashed the header and C++ file together, which is
               | _egregiously_ bad in the context of QT. This isn't a new
               | library, it's been around for almost thirty years. Max
               | token sizes have nothing to do with that.
               | 
               | I invite anyone to post a chat transcript showing a
               | successful run of R1 against this prompt (and please tell
               | me which API/service it came from so I can go use it
               | too!)
        
             | marricks wrote:
             | > it just runs out of context before it can really do
             | anything
             | 
             | I mean, couldn't that be because they're just overwhelmed
             | by users at the moment?
             | 
             | > And the output is very bad - it mashes together the
             | header and cpp file
             | 
             | That sounds way worse, and like, not something caused by
             | being hugged to death though.
             | 
             | Aider recently stated DeepSeek is placed a the top of their
             | benchmark though[1] so I'm inclined to believe it isn't
             | _all_ hype.
             | 
             | [1] https://aider.chat/docs/llms/deepseek.html
        
               | throwup238 wrote:
               | It's definitely not _all_ hype, it really is a
               | breakthrough for open source reasoning models. I don't
               | mean to diminish their contribution, especially since
               | being able to read the reasoning output is a very
               | interesting new modality (for lack of a better word) for
               | me as a developer.
               | 
               | It's just not as impressive as people make it out to be.
               | It might be better than o1 on Python or Javascript thats
               | all over the training data, but o1 is overwhelmingly
               | better at anything outside the happy path.
        
             | sho_hn wrote:
             | Is this a local run of one of the smaller models and/or
             | other-models-distilled-with-r1, or are you using their Chat
             | interface?
             | 
             | I've also compared o1 and (online-hosted) r1 on Qt/C++
             | code, being a KDE Plasma dev, and my impression so far was
             | that the output is roughly on par. I've given both models
             | some tricky tasks about dark corners of the meta-object
             | system in crafting classes etc. and they came up with
             | generally the same sort of suggestions and implementations.
             | 
             | I do appreciate that "asking about gotchas with few
             | definitive solutions, even if they require some
             | perspective" and "rote day-to-day coding ops" are very
             | different benchmarks due to how things are represented in
             | the training data corpus, though.
        
               | throwup238 wrote:
               | I use it through Kagi Assistant which has the proper R1
               | model through Together.ai/Fireworks.ai
               | 
               | My standard test is to ask the model to write a
               | QSyntaxHighlighter subclass that uses TreeSitter to
               | implement syntax highlighting. O1 can do it after a few
               | iterations, but R1's output has been a mess. That said,
               | its thought process revealed a few issues that I then
               | fixed in my canonical implementation.
        
               | sho_hn wrote:
               | Thanks for adding detail! My prompts have been very in-
               | the-bubble-of-Qt I'd say, less so about mashing together
               | Qt and something else, which I agree is a good real-world
               | test case.
        
               | throwup238 wrote:
               | I haven't had the chance to try it out with R1 yet but if
               | you implement a debugger class that screenshots the
               | widget/QML element, dumps its metadata like GammaRay, and
               | includes the source, you can feed that context into
               | Sonnet and o1. They are _scarily_ good at identifying
               | bugs and making modifications if you include all that
               | context (although you have to be selective with what
               | metadata you include. I usually just dump a few things
               | like properties, bindings, signals, etc).
        
               | nialv7 wrote:
               | Tried this on chat.deepseek.com, it seems to be able to
               | do it.
        
               | throwup238 wrote:
               | Does it compile? Put the full chat in Pastebin and let's
               | check it out!
               | 
               | I haven't used their official chat interface or API for
               | privacy reasons.
        
               | CamperBob2 wrote:
               | Some have said (for what little _that 's_ worth) that
               | Kagi's version is not the real thing, but one of the
               | distillations.
        
             | sheepdestroyer wrote:
             | There are R1 providers on openrouter with bigger
             | input/output token limitations than what DeepSeek's API
             | access currently offers.
             | 
             | For instance Fireworks offers R1 with 164K/164K. They are
             | far more expensive than DeepSeek though
        
             | api wrote:
             | It's not great at super-complex tasks due to limited
             | context, but it's quite a good "junior intern that has
             | memorized the Internet." Local deepseek-r1 on my laptop (M1
             | w/64GiB RAM) can answer about any question I can throw at
             | it... as long as it's not something on China's censored
             | list. :)
        
               | azinman2 wrote:
               | How are you running r1 on 64mb of ram? I'm guessing
               | you're running a distill which is not r1
        
         | mritchie712 wrote:
         | openai should pay creators, but:
         | 
         | 1. scraping the internet and making AI out of it
         | 
         | 2. using the AI from #1 to create another AI
         | 
         | are not the same thing.
        
           | zbshqoa wrote:
           | Number 2 is already possible with open models. You can do
           | distillation using Llama, which could likely be doing #1 to
           | build their models (I'm not sure it's the case though)
        
           | bugglebeetle wrote:
           | Yeah, #1 is way worse and #2 falls under "turnabout is fair
           | play."
        
           | epse wrote:
           | #1 destroys peoples willingness to publish and unfairly hogs
           | bandwidth / creates costs for small hosters
           | 
           | #2 makes a big corp a bit angry
           | 
           | Indeed not the same thing
        
           | latexr wrote:
           | > are not the same thing.
           | 
           | You're right. The second one is far more ethical. Especially
           | when stealing from a thief.
           | 
           | Doesn't Sam Altman keep parroting they're developing AI "for
           | the good of humanity"? Well then, someone taking their model
           | and improving on it, making it open-source, having it consume
           | less, and having a cheaper API, should make him delighted.
           | Unless he _*gasp*_ was full of shit the whole time. Who could
           | have guessed?
        
             | perryizgr8 wrote:
             | > Doesn't Sam Altman keep parroting they're developing AI
             | "for the good of humanity"?
             | 
             | "I don't want to live in a world where someone else makes
             | the world a better place better than we do"
             | 
             | - Gavin Belson
        
           | Palmik wrote:
           | I agree, (2) seems much less problematic since the AI outputs
           | are not copyrightable and since OpenAI gives up ownership of
           | the outputs. [1]
           | 
           | So, if you really really care about ToS, then just never
           | enter into a contract with OpenAI. Company A uses OpenAI to
           | generate data and posts it on the open Internet. Company B
           | scrapes open Internet, including the data from Company A [2].
           | 
           | [1]: Ownership of content. As between you and OpenAI, and to
           | the extent permitted by applicable law, you (a) retain your
           | ownership rights in Input and (b) own the Output. We hereby
           | assign to you all our right, title, and interest, if any, in
           | and to Output.
           | 
           | [2]: This is not hypothetical. When ChatGPT got first
           | released, several big AI labs accidentally and not so
           | accidentally trained on the contents of the ShareGPT website
           | (site that was made for sharing ChatGPT outputs). ;)
        
           | sksrbWgbfK wrote:
           | > 2. using the AI from #1 to create another AI
           | 
           | 2. scraping the AI from #1 and making AI out of it
        
           | haswell wrote:
           | Yes, they are different _actions_.
           | 
           | But arguably these actions share enough characteristics that
           | it's reasonable to place them in the same _category_.
           | Something like: "products that exist largely /solely because
           | of the work of other people". The nonconsensual nature of
           | this and the lack of compensation is what people
           | understandably take issue with.
           | 
           | There is enough similarity that it evokes specific feelings
           | about OpenAI when they suddenly find themselves on the other
           | side of the situation.
        
           | Winsaucerer wrote:
           | I'm genuinely not sure which one you think is worse (if any).
           | (1) seems worse, but your reply suggests to me maybe you
           | think (2) is worse.
        
             | meowface wrote:
             | Not that poster, but I think both are equally fine.
             | 
             | It's funny if OpenAI were to complain about this, but at
             | least on Twitter I don't see that much whining about it
             | from OpenAI employees. Sam publicly praised DeepSeek.
             | 
             | I do see some of them spreading the "they're hiding GPUs
             | they got through sanction evasion" theory, which is
             | disappointing, though.
        
           | jillyboel wrote:
           | You're right, (1) is violating the rights of a large portion
           | of the population, (2) is violating the rights of one company
        
           | tw1984 wrote:
           | #1 is stealing from all average joes ever lived on earth
           | 
           | #2 is taking advantages from closedAI.
           | 
           | they are indeed different
        
         | pilooch wrote:
         | Any ML based service with an API is basically a dataset builder
         | for more ML. This has been known forever and is actually a
         | useful "law" of ML-based systems.
        
           | sho_hn wrote:
           | Aye, this should be obvious even to non-technical folks. Much
           | has been written about how LLMs regurgitate the data they
           | were trained on. So if you're looking for data to train on,
           | you can certainly extract it there.
           | 
           | Plus of course for people within the tech bubble, plenty of
           | research results on the value of synthetically augmented and
           | expanded training data that put the impact past just
           | regurgitating source data.
           | 
           | This whole episode is a failure of reporting what to expect
           | next and projecting running costs etc. most of all.
        
           | amelius wrote:
           | This is why models should be open. Or at least they should
           | have a local option.
        
         | stackghost wrote:
         | [flagged]
        
           | bloomingkales wrote:
           | I personally love this chef's kiss of a flip flop sam did
           | here:
           | 
           | https://blog.samaltman.com/trump
           | 
           | https://www.reddit.com/r/YAPms/comments/1i7ry5m/sam_altman_g.
           | ..
           | 
           | Only a truly talented piece of shit can be as prolific as
           | this.
           | 
           |  _" He is irresponsible in the way dictators are."_
           | 
           | Chef's kiss.
           | 
           | Edit:
           | 
           | Kids, don't aspire to be like Altman. We as a community need
           | to espouse more values than _tech is gonna tech_.
        
             | JumpCrisscross wrote:
             | > _don 't aspire to be like Altman_
             | 
             | And don't aspire to be like those who saw what he is but
             | made peace with it in exchange for silver.
        
               | gadders wrote:
               | You mean all of the YC management, including PG?
        
               | istjohn wrote:
               | Hey now, that's not very curious of you. /s
        
               | buran77 wrote:
               | Well, anyone who will flex their spine in every
               | (im)possible position as required of them, just to get
               | even more money and power.
               | 
               | I could understand that from someone with an empty
               | stomach. But so many people doing it when their pockets
               | are already overflowing is exactly the kind of rot that
               | degrades an entire society.
               | 
               | We're all just seeing the results so much better now that
               | they can't even be bothered to pretend they ever more
               | than this.
               | 
               | Later edit: The way this submission fell ~400th spots
               | after just two hours despite having 1250 points and 550
               | comments, had its comments flagged and shuffled around to
               | different submissions as soon as they touched too close
               | to YC&Co is a good mirror of how today's society works.
        
               | ToucanLoucan wrote:
               | It's an addiction. There's no amount of money that will
               | be enough, there's no amount of power that will be
               | enough. They'll burn the world for another hit, and we
               | know that because we've been watching them do it for 50
               | years now.
        
               | stackghost wrote:
               | Yes.
        
               | RIMR wrote:
               | Yes. Especially them.
        
               | toxic wrote:
               | Yes.
        
               | marxisttemp wrote:
               | Paul Graham now reposts right wing grift media on his
               | Twitter profile, he's cooked
        
             | SteveGerencser wrote:
             | > don't aspire to be like Altman
             | 
             | Aspire to be like Aaron Schwartz.
        
               | some_furry wrote:
               | (Except for the tragic ending, of course.)
        
               | lukan wrote:
               | If more would be like him, there might be a happy ending.
        
               | some_furry wrote:
               | Agreed.
        
               | ibejoeb wrote:
               | AaronSw exfiltrated data without authorization. You can
               | argue the morality of that, but I think you could make
               | the argument for OpenAI as well. I'm not opining on
               | either, just pointing out the marked similarity here.
               | 
               | edit: It appears I'm wrong. Will someone correct me on
               | what he did?
        
               | gessha wrote:
               | Arguing for the morality of OpenAI is a little bit harder
               | given their history and actions in the last few years.
        
               | ibejoeb wrote:
               | One argument would be means to an end, with the end being
               | the initial advancement of AI.
               | 
               | Again, I'm not offering an opinion on it.
        
               | skeeter2020 wrote:
               | This is an argument, but isn't this where your scenario
               | diverges completely? OpenAI's "means to an end" is
               | further than you state; not initial advancement but the
               | control and profit from AI.
        
               | ibejoeb wrote:
               | Yes, they intended for control and profit, but it's
               | looking like they can't keep it under control and
               | ultimately its advancements will be available more
               | broadly.
               | 
               | So, the argument goes that despite its intention, OpenAI
               | has been one of the largest drivers of innovation in an
               | emerging technology.
        
               | ceejayoz wrote:
               | > edit: It appears I'm wrong. Will someone correct me on
               | what he did?
               | 
               | He didn't do it without authorization.
               | 
               | https://en.wikipedia.org/wiki/Aaron_Swartz
               | 
               | > Visitors to MIT's "open campus" were authorized to
               | access JSTOR through its network.
        
               | richardwhiuk wrote:
               | He wasn't authorised to access the wiring closet. There
               | are many troubling things about the case, but it's fairly
               | clear Aaron knew he was doing something he wasn't
               | authorised to do.
        
               | ceejayoz wrote:
               | > He wasn't authorised to access the wiring closet.
               | 
               | For which MIT can certainly have a) locked the door and
               | b) trespassed him, but that's a very different issue than
               | having authorization to access JSTOR.
        
               | ibejoeb wrote:
               | At that same link is an account of the unlawful activity.
               | He was not authorized to access a restricted area, set up
               | a sieve on the network, and collect the contents of JSTOR
               | for outside distribution.
        
               | ponector wrote:
               | Why should kids aspire to be like Aaron if it is not
               | rewarded in our society? Comparing with such "kings" as
               | Altman or Musk.
        
               | barnabee wrote:
               | Why should anyone aspire to do what is rewarded over what
               | they believe in and what will satisfy them?
        
               | gessha wrote:
               | Not everything virtuous is rewarded monetarily but we
               | aspire to be virtuous, no?
        
               | coliveira wrote:
               | Modern society has stopped to aspire of being virtuous a
               | long time ago. Unfortunately, that's nowadays a minority
               | view.
        
               | robotresearcher wrote:
               | Aaron was not happy. Neither is Trump, or Musk. I don't
               | know if Bernie is happy, or AOC. Obama seems happy.
               | Hilary doesn't. Harris seems happy.
               | 
               | Striving for good isn't gonna be fun all the time, but
               | when choosing role models I like to factor in how happy
               | they seem. I'd like to spend some time happy.
        
               | jncfhnb wrote:
               | I think it's fairly crazy that you believe you have an
               | authentic view into the happiness levels of these people.
        
               | robotresearcher wrote:
               | I used the word 'seem' three times. I think it's pretty
               | unremarkable to report a personal impression without any
               | claim of special insight.
        
               | skeeter2020 wrote:
               | If human beings could be categorized as Happy/Not Happy
               | the world would be a very boring place and life not worth
               | living.
        
               | ponector wrote:
               | Musk looks happy throwing his hand from the heart to the
               | sun.
        
               | cratermoon wrote:
               | Try to imagine a society where people only did things
               | that were rewarded. Could such a society even exist?
               | Thought experiment: make a list of all the jobs,
               | professions, and vocations that are not rewarded in the
               | sense you mean, and imagine they don't exist. What would
               | be left?
        
               | gus_massa wrote:
               | You mean better pay for teachers? It would be nice.
               | 
               | (Since we are dreaming, can I add sane hours for medical
               | doctors (like <= 8 per day)?)
        
               | ponector wrote:
               | I don't need to imagine. Teachers almost everywhere
               | around the globe have poor salaries. In my country there
               | are lower enrolment requirements to universities to
               | become a school teacher than almost every other field of
               | study. Means the dumbest students are there.
               | 
               | And then later they go to the school to teach our future,
               | working with high stress and low salary.
               | 
               | Same with medical school in many countries where
               | healthcare is not privatized. Insane hours, huge
               | responsibilities and poor pay for doctors and nurses in
               | many countries.
               | 
               | Nowadays everyone wants to be an influencer or software
               | developer.
        
               | makapuf wrote:
               | For teachers, sure. For medical doctors, in USA or
               | Europe, I think they are much more paid than sw
               | engineers.
        
               | ponector wrote:
               | In east EU, like Poland sw engineer makes two-three times
               | more than a doctor with much less of effort, education
               | and no responsibility.
               | 
               | And nurses - they work at minimal salary in Poland. Even
               | in USA if you count hourly rates it will be quite poor
               | salary for nurses.
        
               | CalRobert wrote:
               | We need them to help us build a better society.
        
               | ponector wrote:
               | Looks like we need salesmen much more as we value their
               | work more.
        
               | uoaei wrote:
               | Because kids' brains are not as poisoned into believing
               | the most profitable things to do are the most
               | meritorious. Not yet, anyway.
        
               | skeeter2020 wrote:
               | Because only one person can be king, but everybody can
               | participate and contribute. Also there's too many things
               | out side of just being "the best" that decide who gets to
               | be king. Often that person is a terrible leader.
        
               | miramba wrote:
               | Upvoted not because I agree, but I think it's a valid
               | question that shouldn't be greyed out. My kids dream job
               | is youtube influencer, I don't like it but can I blame
               | them? It's money for nothing and the chicks for free.
        
               | ponector wrote:
               | Tragedy of current days. No one wants to be a
               | firefighter, astronaut or a doctor. Influencers
               | everywhere! Can you blame kids? Do you know firefighters
               | who earns million dollars annually?
        
               | reaperman wrote:
               | * Swartz
               | 
               | But yes.
        
               | oooyay wrote:
               | I've read a lot about Aaron's time at Reddit / Not A Bug.
               | I somewhat think his fame exceeds his actual
               | accomplishments at times. He was perceived to be very
               | hostile to his peers and subordinates.
               | 
               | Kind of a cliche, but aspire to be the best version
               | yourself every day. Learn from the successes and failures
               | of others, but don't aspire to be anyone else because
               | eventually you'll be very disappointed.
        
               | bayindirh wrote:
               | The gist is, when you find your biggest flaw, work on it,
               | and repeat; you've already gone great distance.
        
               | baudehlo wrote:
               | I knew Aaron back in my IRC days. He hung out with us to
               | talk about RDF for a good couple of years. We chatted
               | almost every day.
               | 
               | He was lovely. And a genius. Maybe he changed, but he was
               | a truly nice person.
        
               | oooyay wrote:
               | Yeah, definitely not a statement on Aaron himself. More a
               | statement on idolizing people. There will always be
               | instances where they didn't live up to what people think
               | of them as. I think Aaron was fine and a normal human
               | being.
        
               | wongarsu wrote:
               | That's what happens to martyrs. They become larger than
               | life and history remembers an idealized version of them
        
               | ddingus wrote:
               | Indeed
        
               | stevenally wrote:
               | Don't sell your soul, is all.
               | 
               | But survive. This too will pass.
        
             | scotty79 wrote:
             | I especially like how he quoted Napoleon or something
             | framing himself as the heart of revolution and Deep Seek as
             | a child of the revolution only to get a response from some
             | random guy "It's not that deep bro. Just release a better
             | model."
             | 
             | https://x.com/hibakod/status/1883189126553596234
        
             | hn_throwaway_99 wrote:
             | That is particularly gross, but that really feels like the
             | norm among all the tech elite these days - Zuckerberg,
             | Bezos, etc. all doing the most laughable flip flops.
             | 
             | The reason the flip flops are so laughable to me is because
             | they attempt to couch them in some noble, moralistic
             | viewpoint, instead of the obvious reason "We own big
             | companies, the government has extreme power to make or
             | break these companies, and everyone knows kissing up to
             | Trump is what is required to be on his good side."
             | 
             | Profiles in Cowardice, every last one of them.
        
               | jcgrillo wrote:
               | Another point of view is that they never flopped or
               | flipped. They were fascists the whole time and were just
               | lying about it before.
        
               | hn_throwaway_99 wrote:
               | I think Tim Sweeney's (CEO of Epic Games) comment was
               | spot on:
               | 
               | > After years of pretending to be Democrats, Big Tech
               | leaders are now pretending to be Republicans, in hopes of
               | currying favor with the new administration. Beware of the
               | scummy monopoly campaign to vilify competition law as
               | they rip off consumers and crush competitors.
               | 
               | This is exactly what OpenAI is trying to do with these
               | allegations.
        
               | stevenAthompson wrote:
               | Those men and their companies are responsible for
               | hundreds of thousands of jobs and a significant portion
               | of the global economy. I'm actually thankful that they
               | aren't shooting their mouths off to the new boss like
               | spoiled children at their first job. It wouldn't make the
               | world better, it would make their companies and the lives
               | of those who depend on them, worse.
               | 
               | There is a fine line between cowardice and common sense.
        
               | jcgrillo wrote:
               | In what sense is the federal government "the boss" of
               | private sector businesses? This isn't an oligarchy yet,
               | right? They don't _have_ to behave obsequiously, they
               | _are choosing to_. They 're doing it for themselves, not
               | for their shareholders or their employees. It's an
               | attempt to grab power and _become_ oligarchs because they
               | see in this government a gullible mark.
        
               | stevenAthompson wrote:
               | > This isn't an oligarchy yet, right?
               | 
               | The richest man in the world has a government office down
               | the street from the white house, which the taxpayers are
               | funding. He's rumored to sleep there.
               | 
               | What do you think?
        
               | hn_throwaway_99 wrote:
               | Puhleeeese. I'm not advocating that these leaders all
               | lead protest marches against the new administration. But
               | the transparent obsequiousness and Trump ball gargling
               | under the guise of some moralistic principles is so
               | nauseating. And please spare me the idea that the likes
               | of Zuckerberg or Bezos gives a rat's ass about their
               | employees.
               | 
               | For a contrast to the Bezos, Zuckerberg and Altman types,
               | look at Tim Cook. Sure, Apple paid the 1 million
               | inauguration "donation", and Cook was at the
               | inauguration, and I'm not arguing he's winning any
               | "Profiles in Courage" awards, but he didn't come out with
               | lots of tweets claiming how massuh Trump is so wise and
               | awesome, Apple didn't do a 180 on their previous
               | policies, etc.
        
             | the_optimist wrote:
             | You're awfully salty and biased toward Reddit contaminants.
             | Don't imagine to speak for a community except people who
             | agree with you a priori. Also, don't steal my nternet
             | points, taste upon it, redditoes.
        
             | hibikir wrote:
             | As a society we might talk about virtue, but the reason we
             | put it as a goal in stories is that in the real world, we
             | don't reward it. It's not just that corruption wins
             | sometimes, but we directly punish those that fight it. The
             | mood of the times, if anything, comes from people realizing
             | that what we called moral behavior leads to worse outcomes
             | for the virtuous.
             | 
             | A community only espouses good values when it punishes bad
             | behavior. How do we do this when those misbehaving are very
             | rich, and attempting to punish the misbehavior has negative
             | consequences on you? There just aren't many available tools
             | that don't require significant sacrifices.
        
               | js8 wrote:
               | > A community only espouses good values when it punishes
               | bad behavior.
               | 
               | This is the "beauty" of the free market ideology (see
               | e.g. https://a16z.com/the-techno-optimist-manifesto/ ).
               | If all the transactions are voluntary, there is no way to
               | punish anyone.
        
               | stevenAthompson wrote:
               | > If all the transactions are voluntary, there is no way
               | to punish anyone.
               | 
               | This is obviously untrue at face value. See: Cancel
               | Culture, Bud Light, and Freedom Fries for examples.
               | 
               | Did you mean something more than what you stated here?
        
             | ViktorRay wrote:
             | I don't think your links are evidence of a flip flop.
             | 
             | The first link is from mid-2016. The second link is from
             | January 2025.
             | 
             | It is entirely reasonable for someone to genuinely change
             | his or her views of a person over the course of 8.5 years.
             | That is a substantial length of time in a person's life.
             | 
             | To me a "flip-flop" is when one changes views on something
             | in a very short amount of time.
        
               | meowface wrote:
               | IMO it probably is and Altman probably still (rightly)
               | hates Trump. He's playing politics because he needs to. I
               | don't really blame him for it, though his tweet certainly
               | did make me wince.
        
               | bloomingkales wrote:
               | _" I don't really blame him for it"_
               | 
               | That's the thing though right, that we all created this
               | mess together. Like yeah, _why don 't you (and the rest
               | of us) blame him?_. We're all pretty warped and it's
               | going to take collective rehab.
               | 
               | Super pretentious to quote MLK, but the man had stuff to
               | say so here it is (on Inaction):
               | 
               |  _" He who passively accepts evil is as much involved in
               | it as he who helps to perpetrate it"_
               | 
               |  _" The ultimate tragedy is not the oppression and
               | cruelty by the bad people but the silence over that by
               | the good people"_
        
               | whatshisface wrote:
               | It's not pretentious to quote Martin Luther King.
        
               | benatkin wrote:
               | It seems he was virtue signaling before. So it would be
               | more accurate to blame him for having let himself become
               | an ego driven person in the past. Or to put it nicely and
               | to add the context of Brian Armstrong of Coinbase, who
               | has also been showing public support for Trump, a
               | mission-driven person.
        
               | mrandish wrote:
               | > It seems he was virtue signaling before.
               | 
               | Yes, the first mistake was a business leader in tech
               | taking a _public_ political position. It was popular and
               | accepted (if not expected) in the valley in 2016.
               | 
               | Doing that then (and banking the social and reputational
               | proceeds) created the problem of dissonance now. If he'd
               | just stayed neutral in public in 2016, he could do what
               | he's doing now and we could assume he's just being a
               | pragmatic business person lobbying the government to
               | further his company's interests.
        
               | benatkin wrote:
               | I think "progressive" is probably the safest position to
               | take. It also works if you want to get involved in a
               | different sort of politics later on. David Sacks had no
               | problem doing that when he was no longer interested in
               | being CEO of a large company.
        
               | mrandish wrote:
               | The evidence indicates not taking a position is the
               | optimal position.
               | 
               | I have a lot of respect for CEOs who just focus on being
               | a good CEO. It's a hard enough job as is. I don't care
               | about or want to know some CEO's personal position on
               | politics, religion or sports teams. It's all a
               | distraction from the job at hand. Same goes for actors,
               | athletes and singers. They aren't _qualified_ to have an
               | opinion any more relevant than anyone else 's, except on
               | acting, athletics, singing - or CEO-ing.
               | 
               | Sadly, my perspective is in the minority. Which is why I
               | think so many public figures keep making this mistake.
               | The media, pundits and social sphere _need_ them to keep
               | making this mistake.
        
               | benatkin wrote:
               | I guess I think they should study what a neutral position
               | looks like, and avoid going beyond it as best as they
               | can. I had in mind a "progressive" who avoids any hot
               | button issues. Someone with a high profile will be asked
               | about politics from time to time. I think Brian Chesky is
               | a good example of acting like a progressive in a way that
               | stays low profile, but maybe he doesn't really act like
               | one. https://www.businessinsider.com/brian-chesky-airbnb-
               | new-bree...
               | 
               | Also it helps to have sincere political views. GitHub's
               | CEO at the time of #DropICE was too cynical and his image
               | suffered because of it.
        
               | mrandish wrote:
               | > study what a neutral position looks like
               | 
               | There are no neutral positions in today's political
               | landscape. I'm not stating my opinion here, this is
               | according to most political positions on the spectrum.
               | You suggested "Progressive" (but without hot button
               | issues) as a way of signaling a neutral position. That
               | may be true in parts of the valley tech sphere but it
               | certainly doesn't hold in the rest of the U.S.
               | "Progressive" is usually defined being to the left of
               | "Liberal", so it's hardly neutral. Over half of U.S.
               | voters cast their ballot for the Republican candidate.
               | Almost all those people interpret anyone identifying
               | themselves as "Liberal" as definitely partisan (and
               | negative, of course). Most of them see "Progressive" as
               | dangerously close to "Socialist". And the same holds true
               | for the term "Conservative" on the other side of the
               | spectrum, of course.
               | 
               | No, identifying as "Progressive" wouldn't distance you
               | from political connotations and culture warring, it's
               | leaping into the maelstrom yelling "Yipee-Ki-Yay!" You
               | may want to update your priors regarding how the broad
               | populace perceives political labels. With the populace
               | divided almost exactly in half regarding politics and
               | cultural war issues and a large percentage on both sides
               | having "Strong" or "Very Strong" feelings, stating any
               | position will be seen as strongly negative by tens of
               | millions of people. As was said in the movie "WarGames",
               | the only winning move is not playing.
        
               | rybosworld wrote:
               | Seems like an extremely naive take.
               | 
               | In 2016: Sam alluded to Trump's rise as not dissimilar to
               | Hitler's. He said that Trump's ideas on how to fix things
               | are so far off the mark that they are dangerous. He even
               | quoted the famous: "The only thing necessary for the
               | triumph of evil is for good men to do nothing."
               | 
               | In 2025: "I'm not going to agree with him on everything,
               | but I think he will be incredible for the country"
               | 
               | This is quite obviously someone who is pandering for
               | their own benefit.
        
               | benatkin wrote:
               | Just like JD Vance.
        
               | themaninthedark wrote:
               | This is quite honestly one of the major problems with our
               | society right now. Once you take a public stance, you are
               | not allowed to revisit and re-evaluate. I think that this
               | is by and large driving most of the polarization in the
               | country, since "My view is right and I will not give an
               | inch least I be seen as weak".
               | 
               | While most of the things affected are highly political
               | situations, i.e. Trump's ideas or Biden's fitness. We
               | also seem to have thrown out things that we used to
               | consider cornerstones of liberal democracy i.e. our ideas
               | regarding free speech and censorship, where we claim that
               | it's not happening because it is a private company.
        
             | belter wrote:
             | "Donald Trump represents an unprecedented threat to
             | America, and voting for Hillary is the best way to defend
             | our country against it"                                   -
             | Sam Altman - 2016
             | 
             | "If you elect a reality TV star as President, you can't be
             | surprised when you get a reality TV show"
             | - Sam Altman - 2017
             | 
             | "When the future of the republic is at risk, the duty to
             | the country and our values transcends the duty to your
             | particular company and your stock price."
             | - Sam Altman - 2017
             | 
             | "I think I started that a little bit earlier than other
             | people, but at this point I am in really good company"
             | - Sam Altman - 2017 ( On his criticism of Trump )
             | 
             | "Very few people realize just how much @reidhoffman did and
             | spent to stop Trump from getting re-elected -- it seems
             | reasonably likely to me that Trump would still be in office
             | without his efforts. Thank you, Reid!"
             | - Sam Altman - 2020
        
             | meowface wrote:
             | Although I dislike him now glazing Trump, I understand why
             | he's doing it. Trump runs a racket and this is part of the
             | game.
             | 
             | One of my most contrarian positions is I still like and
             | support Altman, despite most of the internet now hating him
             | almost as much as they (justifiably) hate Elon. Was a fan
             | of Sam pre-YC presidency and still am now.
             | 
             | (I also am a big fan of DeepSeek and its CEO.)
        
               | CoastalCoder wrote:
               | In the interest of helping avoid an echo chamber, would
               | you mind giving some of the things you like about current
               | Altman?
        
               | benterix wrote:
               | I'd love to hear something positive about the current
               | Altman, too. Anything would be good.
        
               | robotresearcher wrote:
               | For me, it's the technical results. Same as for Musk.
               | 
               | Tesla accelerated us forward into the electric car age.
               | SpaceX revolutionized launches.
               | 
               | OpenAI added some real startup oomph to the AI arms race
               | which was dominated by megacorps with entrenched products
               | that they would have disrupted only slowly.
               | 
               | So these guys are doing useful things, however you feel
               | about their other conduct. Personally I find the gross
               | political flip-flops hard to stomach.
        
               | whatshisface wrote:
               | Why would you support someone you said was part of a
               | racket in the sentence before? We're talking about real
               | life, where actions have consequences, not a TV show
               | where we're expected to identifiy with Tony Soprano.
        
             | breakyerself wrote:
             | If you didn't sexually assault your sister you're already
             | off to a good start.
        
           | 65 wrote:
           | Yeah I don't know, Altman is a sociopath who is now trying to
           | get intertwined with local governments (SF) as well as the
           | federal government. He's going to do a lot of weaseling to
           | get what he wants: laws that forcibly make OpenAI a monopoly.
           | 
           | Society will always have crazy sociopaths destroying things
           | for their own gain, and now is Altman's turn.
        
           | api wrote:
           | I'm sure him lining up to kiss Trump's ring for some kind of
           | bailout is not a coincidence.
        
           | blackeyeblitzar wrote:
           | I don't care for Sam Altman and his general untrustworthy
           | behavior. But DeepSeek is perhaps more untrustworthy. Models
           | from American companies at least aren't surprising us with
           | government driven misinformation, and even though safety can
           | also be censorship, the companies that make these models at
           | least openly talk about their safety programs. DeepSeek is
           | implementing a censorship and propaganda program without
           | admitting it at all, and once they become good at doing it in
           | less obvious ways, it can become very damaging and corrupt
           | the political process of other societies, because users will
           | trust the tools they use are neutral.
           | 
           | I think DeepSeek's strategy to announce a misleading low cost
           | (just the final training run that optimizes a base model that
           | in turn is possibly based on OpenAI) is also purposeful.
           | After all, High Flyer, the parent company of DeepSeek, is a
           | hedge fund - and I bet they took out big short positions on
           | Nvidia before their recent announcements. The Chinese
           | government, of course, benefits from a misleading number
           | being announced broadly, causing doubt among investors who
           | would otherwise continue to prop up American technology
           | startups. Not to mention the big fall in American markets as
           | a result.
           | 
           | I do think there's also a big difference between scraping the
           | Internet for training data, which might just be fair use, and
           | training off other LLMs or obtaining their assets in some
           | other way. The latter feels like the kind of copying and
           | industrial espionage that used to get China ridiculed in the
           | 2000s and 2010s. Note that DeepSeek has never detailed their
           | training data, even at a high level. This is true even in
           | their previous papers, where they were very vague about the
           | pre training process, which feels suspicious.
        
             | dheera wrote:
             | > I bet they took out big short positions on Nvidia before
             | their announcements
             | 
             | Good for them! I hope this teaches Wall Street to not freak
             | out about an unverified announcement.
             | 
             | Wall Street lost billions, and I hope they learned their
             | lesson and next time will not crash the market when
             | unverified news comes out.
        
             | tempusalaria wrote:
             | DeepSeek v3 (where the training cost claims come from) was
             | announced a month ago and it had no impact outside of a
             | small circle
        
             | ryanisnan wrote:
             | > Models from American companies at least aren't surprising
             | us with government driven misinformation, and even though
             | safety can also be censorship
             | 
             | Being a citizen of a western nation, I'm inclined to agree
             | with the general sentiment here, but how can you
             | definitively say this? You, or I, don't know with any
             | certainty what interference the US government has played
             | with domestic LLMs, or what lies they have fabricated and
             | cultivated, that are now part of those LLMs' collective
             | knowledge. We can see the perceived censorship with
             | deepseek more clearly, but that isn't evidence that we're
             | in any safer territory.
        
             | pphysch wrote:
             | > Models from American companies at least aren't surprising
             | us with government driven misinformation
             | 
             | There are loads of examples on the internet of LLMs pushing
             | (foreign) government narratives e.g. on Israel-Palestine.
             | 
             | Just because you might agree with the propaganda doesn't
             | make it any less problematic.
        
               | blackeyeblitzar wrote:
               | > There are loads of examples on the internet of LLMs
               | pushing (foreign) government narratives e.g. on Israel-
               | Palestine
               | 
               | There isn't even a single example of that. If an LLM is
               | taking a certain position because it has learned from
               | articles on that topic, that's different from it being
               | manipulated on purpose to answer differently on that
               | topic. You're confusing an LLM simply reflecting the
               | complexity out there in the world on some topics (showing
               | up in training data), with government forced censorship
               | and propaganda in DeepSeek.
               | 
               | The two aren't the same, not even remotely close.
        
               | pphysch wrote:
               | Fine, whatever. It's actually much _more_ concerning if
               | the overall information landscape has been so curated by
               | censors that a naively-trained LLM comes  "pre-censored",
               | as you are asserting. This issue is so "complex" when it
               | comes to one side, and "morally clear" when it comes to
               | the other. Classic doublespeak.
               | 
               | That's far more dystopian than a post-hoc "guardrailed"
               | model (that you can run locally without guardrails).
        
             | vohk wrote:
             | > Models from American companies at least aren't surprising
             | us with government driven misinformation
             | 
             | Is corporate misinformation so much better? Recall about
             | Tienanmen Square might be more honest but if LLMs had been
             | available over the past 50 years, I would expect many
             | popular models would have cheerfully told us company towns
             | are a great place to live, cigarettes are healthy,
             | industrial pollution has no impact on your health, and
             | anthropogenic climate change isn't real.
             | 
             | Especially after the recent behaviour of Meta, Twitter, and
             | Amazon in open support of Trump and Republican interests,
             | I'll be shocked if we don't start seeing that reflected in
             | their LLMs over the next few years.
        
             | cycomanic wrote:
             | > I don't care for Sam Altman and his general untrustworthy
             | behavior. But DeepSeek is perhaps more untrustworthy.
             | Models from American companies at least aren't surprising
             | us with government driven misinformation, and even though
             | safety can also be censorship, the companies that make
             | these models at least openly talk about their safety
             | programs. DeepSeek is implementing a censorship and
             | propaganda program without admitting it at all, and once
             | they become good at doing it in less obvious ways, it can
             | become very damaging and corrupt the political process of
             | other societies, because users will trust the tools they
             | use are neutral.
             | 
             | These arguments always remind me of the arguments against
             | Huawei because they _might_ be spying on western countries.
             | On the other hand we had the US government working hand in
             | hand with US corporations in proven spying operations
             | against western allies for political and economic gain. So
             | why should we choose an American supplier over a Chinese
             | one?
             | 
             | > I think DeepSeek's strategy to announce a misleading low
             | cost (just the final training run that optimizes a base
             | model that in turn is possibly based on OpenAI) is also
             | purposeful. After all, High Flyer, the parent company of
             | DeepSeek, is a hedge fund - and I bet they took out big
             | short positions on Nvidia before their recent
             | announcements. The Chinese government, of course, benefits
             | from a misleading number being announced broadly, causing
             | doubt among investors who would otherwise continue to prop
             | up American technology startups. Not to mention the big
             | fall in American markets as a result.
             | 
             | Why should I care about the stock value of US corporations?
             | 
             | > I do think there's also a big difference between scraping
             | the Internet for training data, which might just be fair
             | use, and training off other LLMs or obtaining their assets
             | in some other way.
             | 
             | So if training of copyrighted work scrapped of the Internet
             | is fair use, how would the training of the LLMs not be fair
             | use as well? You can't have it both ways.
        
           | tootie wrote:
           | The coup against him is looking more and more like a huge "I
           | told you so" moment.
        
           | Dansvidania wrote:
           | Indeed. First thing I thought was "call a wahmbulance!".
        
           | benreesman wrote:
           | The guy is a total fucking psycho and the rest of the board
           | are no gems either.
           | 
           | Their failure is important at a minimum to the future of the
           | United States if not the world.
        
           | dang wrote:
           | Ok, but please don't break HN's rules when commenting here.
           | 
           | You may not owe Altmen better, but you owe this community
           | better if you're participating in it.
           | 
           | https://news.ycombinator.com/newsguidelines.html
        
             | stackghost wrote:
             | Once again you abuse your moderator powers to enforce your
             | personal vendetta against people who dare to speak ill of
             | tech CEOs.
             | 
             | I find your behavior repulsive and fervently wish you would
             | quit.
        
               | dang wrote:
               | This is what people say when they don't want the rules to
               | be applied even-handedly.
               | 
               | It's not a borderline call--I'd post exactly the same
               | thing regardless of who or what such a comment was about.
        
               | stackghost wrote:
               | >This is what people say when they don't want the rules
               | to be applied even-handedly.
               | 
               | Not even close.
               | 
               | This guy is actively ruining society while enriching
               | himself in the process, but we somehow can't call a spade
               | a spade?
               | 
               | Pathetic.
        
               | dang wrote:
               | I suppose my chances of getting a straight answer aren't
               | too good right now but I'd love to hear your thoughts on
               | something.
               | 
               | HN's stated mandate is intellectual curiosity
               | (https://news.ycombinator.com/newsguidelines.html, https:
               | //hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
               | ).
               | 
               | Do you feel like your comment
               | https://news.ycombinator.com/item?id=42866108 counts as
               | curious? or is it rather that you think something else is
               | more important?
        
         | belter wrote:
         | So it is true, they run out of Data to steal? :-)
         | 
         | And then where DeepSeek steal from next? Do they steal from
         | themselves? Do they steal the stolen models they stole from the
         | stolen data?
         | 
         | The AI Ponzi scheme...
        
           | troyvit wrote:
           | Exactly this, especially as journalism melts down into slag.
           | Soon all anybody will have to train on is social media,
           | Wikipedia and GitHub, and that last one will slowly be
           | metastasized by AI-generated code anyway.
           | 
           | It reminds me of 1984 in a sense. "Don't you see that the
           | whole aim of Newspeak is to narrow the range of thought? In
           | the end we shall make thoughtcrime literally impossible,
           | because there will be no words in which to express it."
           | 
           | Unlike 1984 I don't see this winnowing of new concepts as
           | purposeful, but on the other hand I keep asking myself how we
           | can be so stupid as to keep doing it.
        
         | TypingOutBugs wrote:
         | Screw OpenAI, they scrape us without issues so someone scraped
         | them. No issues with this.
        
           | coliveira wrote:
           | But the government will now claim this is against "national
           | security". Only American companies are allowed to commit this
           | kind of "sleight of hand".
        
             | Imustaskforhelp wrote:
             | Yes they would. But it would pointless. And clear hypocrisy
             | as well.
        
               | coliveira wrote:
               | Hypocrisy or not, the US government has managed to make
               | this work for a long time now, the Biden administration
               | just proves the point. Thankfully, other countries are
               | starting to catch up to this scam.
        
               | Imustaskforhelp wrote:
               | Yes , to be fair , As a foreigner (not a US citizen
               | basically) I don't mean to offend somebody. But USA just
               | seems to be build on top of Hypocrisy.
               | 
               | Like the fact that US revolution was basically
               | kickstarted by blatantly breaking the patent law (like
               | there was this one mill specifically) , I think its a
               | historic event. And now here we are ! The scam of
               | national security.
               | 
               | To be honest. People seem to be really kind on the fall
               | of USA. I am not that interested since the rise of China
               | terrifies me. But the hypocrisy of USA / losing such soft
               | power (like here I am , from random country critiquing
               | USA based on facts , it really downplays it being a
               | superpower) that would be the downfall of USA.
               | 
               | To me , the future terrifies me. In fact the present
               | terrifies me. I think the world is running crazy or maybe
               | its just me.
        
         | pen2l wrote:
         | While all of this is true, that DeepSeek wouldn't be here were
         | it not for the research that preceded it notably Google's
         | paper, then Llama, and ChatGPT which they're modeled after, its
         | release still did something profound to their psyche, the
         | motivation and self-actualization this instills to the Chinese.
         | They witnessed the power of their accomplishments: a side-
         | hustle project knocked off an easy trillion. This is only
         | egging them on and will serve to ramp up their efforts even
         | more.
         | 
         | Separately, I do think that now that the Chinese leadership saw
         | this, that they have the chops to pull this off and then some,
         | they are probably going to rein in future innovations; they'll
         | likely demand that the big future discoveries remain closed-
         | sourced (or even unannounced/unpublicized).
        
           | tedivm wrote:
           | OpenAI wouldn't be here without the work that Yann Lecun did
           | at Facebook (back when it was facebook). Science is built on
           | top of science, that's just how things work.
        
             | wrasee wrote:
             | Yes, but in science you reference your work and credit
             | those who came before you.
             | 
             | Edit: I am not defending OpenAI and we are all enjoying the
             | irony here. But it puts into perspective some of the wilder
             | claims circulating that DeekSeek was able to somehow
             | complete with OpenAI for only $5M, as if on a level playing
             | field.
        
               | bugglebeetle wrote:
               | Like all those papers with their long lists of citations
               | OpenAI has been releasing?
        
               | dkjaudyeqooe wrote:
               | That's only in academia. The same thing happens in
               | commerce, only there is no (official) credit given.
        
               | tedivm wrote:
               | OpenAI has been hiding their datasets, and certainly
               | haven't credited me for the data they stole from my
               | website and github repositories. If OpenAI doesn't think
               | they should give attribution to the data they used, it
               | seems weird to require that of others.
               | 
               | Edit: Responding to your edit, Deepseek only claimed that
               | the final training run was $5m, not that the whole
               | process caught that (they even call this out). I think
               | it's important to acknowledge that, even if they did get
               | some training data from OpenAI, this is a remarkable
               | achievement.
        
               | ambicapter wrote:
               | Only weird if you think what OpenAI did should be the
               | norm.
        
               | wrasee wrote:
               | Right. I think many here are enjoying the Schadenfreude
               | against OpenAI, but that hardly makes it right. It just
               | makes it a race to the bottom.
        
               | wrasee wrote:
               | It is a remarkable achievement. But if "some training
               | data from OpenAI" turns out to essentially be a wholesale
               | distillation of their entire model (along with Llama etc)
               | I do think that somewhat dampens the spirit of it.
               | 
               | We don't know that of course. OpenAI claim to have some
               | evidence and I guess we'll just have to wait and see how
               | this plays out.
               | 
               | There's also a substantial difference between training of
               | the entire internet and one that very specifically
               | targets your competitor's products (or any specific work
               | directly).
        
               | Filligree wrote:
               | That's $5M for the final training run. Which is an
               | improvement to be sure, but it doesn't include the
               | _other_ training runs -- prototypes, failed runs and so
               | forth.
        
               | coliveira wrote:
               | It is OpenAI that discredits themselves when they say
               | that each new model is the result of hundreds of USD
               | millions in training. They throw this around as it is a
               | big advantage of their models.
        
               | nicce wrote:
               | And the cost is based on the imaginary currency that
               | Microsoft has given for them as Azure computing.
        
             | blackeyeblitzar wrote:
             | Is that really true? If anything OpenAI was dependent on
             | the transformers paper from Google from Ashish Vaswani and
             | others. LeCun has been criticizing LLM architectures for a
             | long time and has been wrong about them for a long time.
        
               | mv4 wrote:
               | That was my impression too. He is considered the inventor
               | of CNN back in 1998. Is there anything more recent that's
               | meaningful?
        
               | blackeyeblitzar wrote:
               | Personally, I have not seen anything from him that is
               | meaningful. OpenAI and Anthropic (itself started by
               | former OpenAI people) of course have built their models
               | without LeCun's contributions. And for a few years now,
               | LeCun has been giving the same talk anywhere he makes
               | appearances, saying that large language models are a dead
               | end and that other approaches like his JEPA architecture
               | are the future. Meanwhile current LLM architecture has
               | continued to evolve and become very useful. As for the
               | misuse of the term "open source", I think that really
               | began once he was at Meta, and is a way to use his fame
               | to market Llama and help Meta not look irrelevant.
        
               | tedivm wrote:
               | They literally cited LeCun in their GPT papers.
        
               | tedivm wrote:
               | I was more referring to this paper from 2015:
               | 
               | https://scholar.google.com/citations?view_op=view_citatio
               | n&h...
               | 
               | Basically all LLM can trace their origin back to that
               | paper.
               | 
               | This was just a single example though. The whole point is
               | that people build on the work from the past, and that
               | this is normal.
        
               | mv4 wrote:
               | Thank you for sharing this.
        
               | esafak wrote:
               | That's just an overview for paper for those new to the
               | field. The transformer architecture has a better claim to
               | being the origin of LLMs.
        
               | amelius wrote:
               | By the way, as someone who once did classical image
               | recognition using convolutions, I can't say I was very
               | impressed by the CNN approach, especially since their
               | implementation didn't even use FFTs for efficiency.
        
             | zbendefy wrote:
             | Also without the "attention is all you need" paper from
             | google
        
           | openrisk wrote:
           | > they'll likely demand that the big future discoveries
           | remain closed-sourced
           | 
           | Depends on whether they want these tools to be adopted in the
           | wider world. Rightly or wrongly there is a lot of suspicion
           | in the West and an open source approach builds trust.
        
           | hn_throwaway_99 wrote:
           | > While all of this is true, that DeepSeek wouldn't be here
           | were it not for the research that preceded it (notably
           | Llama), and ChatGPT which they're modeled after...
           | 
           | If the allegation _is_ true (we don 't know yet), then what
           | you've written perfectly proves the point everyone is making.
           | ChatGPT wouldn't be here if it weren't for all the research
           | and work that preceded it in terms of tons of scrapable
           | content being available on the Internet, and it's not like
           | OpenAI invented transformers either.
           | 
           | Nobody is accusing DeepSeek of hacking into OpenAI's systems
           | and stealing their content. OpenAI is just saying they
           | scraped them in an "unauthorized" manner. The hypocrisy is
           | laughably striking, but sadly nobody has any shame anymore in
           | this world it seems. Play me the world's tiniest violin for
           | OpenAI.
        
           | stravant wrote:
           | Yes, and what does preceding research do? Get followed by
           | more research building on it.
        
             | dylan604 wrote:
             | Standing on the shoulders and it's turtles all the way
        
           | dismalaf wrote:
           | Don't forget all the research that came before OpenAI and
           | ChatGPT...
        
           | nicce wrote:
           | We wouldn't be here discussing if nobody invented internet...
           | nor these models had training data at all.
           | 
           | > Separately, I do think that now that the Chinese leadership
           | saw this, that they have the chops to pull this off and then
           | some, they are probably going to rein in future innovations;
           | they'll likely demand that the big future discoveries remain
           | closed-sourced (or even unannounced/unpublicized).
           | 
           | How do we know that this is not already happening with
           | OpenAI/Meta and the U.S. government at some level? The
           | concept of power is equal, whether we wanted it or not. We
           | don't have to pretend to be "better" all the time.
        
         | coliveira wrote:
         | They really lost their minds. They're all scared and worried
         | because companies in other countries can also access the same
         | data they stole from the Internet.
        
         | rvz wrote:
         | They have been out-grifted by DeepSeek and OpenAI is not happy
         | about someone out-shining them on that.
         | 
         | The best part is "their IP" was humanity's scraped content and
         | they are angry that DeepSeek did their job for them and gave it
         | away for free.
        
         | cscurmudgeon wrote:
         | Scraping data is different from scraping outputs from a model.
        
           | jacobgorm wrote:
           | No it is not, data is data, whether it gets loaded from a
           | file on disk or generated by multiplying lots of matrices.
        
           | idle_zealot wrote:
           | Like, in a strict literal sense, sure? Do you mean to make a
           | claim about moral or legal differences?
        
           | openrisk wrote:
           | Because the data is mine and the model is yours?
        
           | wkz wrote:
           | Technically, sure. What is the moral distinction though?
        
           | Rebelgecko wrote:
           | Copyright is weird and often legal [?] moral, but I'm having
           | a hard time constructing a mental model where it's ok to
           | scrape a novel written by a person but it's not ok to scrape
           | a story written by chatgpt
        
         | 28304283409234 wrote:
         | ClosedAI? StolenAI!
        
         | gruez wrote:
         | The picture at the end showing deepseek's privacy policy and
         | being concerned that it's "a security risk" is hilarious[1].
         | Basically every B2C company collects this sort of
         | information[2], and is far less intrusive than what social
         | networks collect[3]. But because it's Chinese and at the risk
         | of overtaking Western companies, people are suddenly worried
         | about device information and IP addresses?
         | 
         | [1] https://semking.com/wp-
         | content/uploads/2025/01/DeepSeek-1024...
         | 
         | [2] https://www.bestbuy.com/site/help-topics/privacy-
         | policy/pcmc...
         | 
         | [3] https://www.facebook.com/privacy/policy/
        
           | semking wrote:
           | One of my core followers named Bruno basically said the same
           | thing under my Linkedin post yesterday:
           | 
           | https://www.linkedin.com/posts/organic-growth_deepseek-
           | the-o...
           | 
           | I welcome friction, so I'll be blunt: I disagree with you,
           | not because what you are saying is wrong but because you only
           | consider systematic data collection.
           | 
           | That's not the issue here.
           | 
           | There's a difference between democracies like the United
           | States or European countries, no matter how IMPERFECT they
           | are, and a dictatorship that does not allow dissenting
           | opinions.
           | 
           | There's a difference in how the data collected will be used.
           | 
           | Freedom of speech, even when it is relative, is better than
           | totalitarianism.
        
             | ryanobjc wrote:
             | It's also important to recognize that the Chinese
             | government is known to walk into internet service companies
             | and demand they censor, alter data, delete things. No court
             | order or search warrant required.
             | 
             | China considers industry to be completely subservient to
             | government. Checks and balances are secondary to ideas like
             | harmony and collective well being.
        
               | semking wrote:
               | Thank you for this balanced and essential comment which
               | is entirely true!
        
             | gruez wrote:
             | >There's a difference between democracies like the United
             | States or European countries, no matter how IMPERFECT they
             | are, and a dictatorship that does not allow dissenting
             | opinions.
             | 
             | >There's a difference in how the data collected will be
             | used.
             | 
             | >Freedom of speech, even when it is relative, is better
             | than totalitarianism.
             | 
             | I don't disagree with "democracy is better than
             | totalitarianism", but what does that have to do with
             | collecting device information and IP addresses? Is that
             | excuse a cudgel you can use against any behavior that would
             | otherwise be innocuous? It's fine to be against deepseek
             | because you're concerned about them getting sensitive data
             | via queries, or even that their models be a backdoor to
             | project chinese soft power, but hand wringing about device
             | information and IP addresses is absurd. It makes as much
             | sense as being concerned that the CCP/deepseek does
             | _meetings_ , because even though every other companies does
             | meetings, CCP/deepseek meetings could be used for
             | totalitarianism.
        
               | semking wrote:
               | I don't disagree with you either and like you, I'm
               | entirely against privacy violations in any way, shape or
               | form.
               | 
               | I admit I am concerned when I see blatant algorithmic
               | manipulation of social platforms to favor any narrative
               | that aligns with geopolitical objectives.
               | 
               | I also wrote about the TikTok algo a few days ago. You'll
               | see what I think of user privacy violations (closed
               | ecosystem + basically a keylogger in this case):
               | 
               | https://semking.com/likes-lies-untold-story-tiktok-
               | algorithm...
               | 
               | I cannot stand when dissenting voices or opinions are
               | shadow-banned.
               | 
               | And I have the same opinion regarding U.S. or EU
               | companies.
               | 
               | Our privacy should be respected.
               | 
               | In the meantime: strong encryption at every corner,
               | please!
        
               | pphysch wrote:
               | > I admit I am concerned when I see blatant algorithmic
               | manipulation of social platforms to favor any narrative
               | that aligns with geopolitical objectives.
               | 
               | I'm curious how robust this principle is for you, because
               | China and Russia are not the first countries that come to
               | mind when talking about the (actual, existing,
               | documented) manipulation of US speech and media by a
               | foreign government.
               | 
               | Yet it seems we can only have this discussion,
               | ironically, when the subject is a US government-approved
               | one like China. Anything else would be problematic and
               | unsafe.
        
               | semking wrote:
               | I don't want to get into politics but I'll gladly admit
               | human beings are biased.
               | 
               | "We Don't See Things As They Are, We See Them As We Are"
               | 
               | -- Samuel b. Nahmani
        
               | gruez wrote:
               | >I'm entirely against privacy violations in any way,
               | shape or form.
               | 
               | >Our privacy should be respected.
               | 
               | Characterizing device information and IP addresses as
               | "privacy violations" is a stretch. If you showed a
               | history railing against this sort of stuff, agnostic of
               | geopolitical alignment, then you get a pass, but I think
               | it's fair to assume the converse until proven otherwise.
               | 
               | >In the meantime: strong encryption at every corner,
               | please!
               | 
               | Irrelevant. The data collection is done by first parties.
               | Encryption doesn't do anything.
               | 
               | >I admit I am concerned when I see blatant algorithmic
               | manipulation of social platforms to favor any narrative
               | that aligns with geopolitical objectives.
               | 
               | >I cannot stand when dissenting voices or opinions are
               | shadow-banned.
               | 
               | What does this have to do with privacy? Again, it's fine
               | to be against "blatant algorithmic manipulation of social
               | platforms" or whatever, but dragging seemingly unrelated
               | topics in an attempt to amass as big pile of greviances
               | as possible is disingenuous.
               | 
               | >I also wrote about the TikTok algo a few days ago.
               | You'll see what I think of user privacy violations
               | (closed ecosystem + basically a keylogger in this case):
               | 
               | >https://semking.com/likes-lies-untold-story-tiktok-
               | algorithm...
               | 
               | Where's the keylogging? I skimmed the article and the
               | only thing I could find was a passing mention about an
               | article that you "was advised not to publish it and I
               | didn't". How much keylogging could possibly going on in a
               | short video app? Is the "keylogging" just a way to make
               | "we measure how engaged someone is with a video" as
               | sinister as possible?
        
               | semking wrote:
               | >Characterizing device information and IP addresses as
               | "privacy violations" is a stretch.
               | 
               | I agree: this is a characterization I never made. FYI, I
               | also collect this type of data about you when you visit
               | my website. That said, telemetry + totalitarianism = bad
               | combo.
               | 
               | >Irrelevant. The data collection is done by first
               | parties. Encryption doesn't do anything.
               | 
               | Even if data is collected by first parties, encryption is
               | still highly relevant because it ensures that the data
               | remains secure in transit and at rest. It does a lot.
               | 
               | >What does this have to do with privacy? Again, it's fine
               | to be against "blatant algorithmic manipulation of social
               | platforms" or whatever, but dragging seemingly unrelated
               | topics in an attempt to amass as big pile of greviances
               | as possible is disingenuous.
               | 
               | You are aggressive for no reason whatsoever. There's
               | nothing disingenuous: when users are shadow-banned by
               | platforms under dictatorships, they end up flagged, and
               | their private data is often analyzed for nefarious
               | reasons. There's a link with privacy but I'll stop at
               | this stage if we cannot have a civilized discussion.
               | 
               | >Where's the keylogging? I skimmed the article and the
               | only thing I could find was a passing mention about an
               | article that you "was advised not to publish it and I
               | didn't". How much keylogging could possibly going on in a
               | short video app? Is the "keylogging" just a way to make
               | "we measure how engaged someone is with a video" as
               | sinister as possible?
               | 
               | "TikTok iOS subscribes to every keystroke (text inputs)
               | happening on third party websites rendered inside the
               | TikTok app. This can include passwords, credit card
               | information and other sensitive user data. (keypress and
               | keydown). We can't know what TikTok uses the subscription
               | for, but from a technical perspective, this is the
               | equivalent of installing a keylogger on third party
               | websites."
               | 
               | https://krausefx.com/blog/announcing-inappbrowsercom-see-
               | wha...
               | 
               | Please note that this article is outdated (August 2022).
               | Importantly, the article does not claim that any data
               | logging or transmission is actively occurring. Instead,
               | it highlights the potential technical capabilities of in-
               | app browsers to inject JavaScript code, which could
               | theoretically be used to monitor user interactions.
        
               | coliveira wrote:
               | Also, the same people that complain about this are just
               | fine with a western government having access to the same
               | data via big corporations. Why being democratic gives you
               | a free access card to disregard privacy, in other words,
               | doing exactly the opposite of what is expected from a
               | free society?
        
             | ziddoap wrote:
             | > _There 's a difference in how the data collected will be
             | used._
             | 
             | Not that we could ever see what the NSA, CISA, ASIS, GCHQ,
             | and other 3/4-letter agencies are actually doing with the
             | collected data.
             | 
             | But they pinky promised to use it properly (or something),
             | so, yay.
        
             | r00fus wrote:
             | Amusing Bruno seems to think in terms of labels when the
             | reality is that the USA imprisons far more people per
             | capita, and blatantly disregards its so-called "core
             | freedoms" (ie, Bill of Rights) for its citizens very often.
             | 
             | This kind of person has a lot of cognitive dissonance going
             | on.
        
         | scotty79 wrote:
         | "That's hilarious!" was my first reaction as well, when I heard
         | about it the first time. When I came to HN and saw this story
         | on top I was hoping this was the top comment. I was not
         | disappointed.
         | 
         | US AI folk were leading for two years by just throwing more and
         | more compute at the same thing that Google threw them like a
         | bone years ago (namely transformers). They made next to no
         | innovation in any area other than how to connect more compute
         | together. The idea of additional inference time compute,
         | looping the network back on its own outputs, which is the only
         | significant conceptual advancement of last years was something
         | I, as a layman, came up with after few days of thinking why AI
         | sucks and what can be done to make it able to tackle problems
         | that require iterative reasoning. They announced it few weeks
         | after I came up with the idea, so it was in the works for some
         | time, but it shows you how basic idea it was. There was nothing
         | else.
         | 
         | Suddenly when there comes a small company that introduced few
         | actual algorithmic advancements which resulted in 100x
         | optimization which is something expected with algorithmic
         | optimizations, the big AI suddenly went into full "dog ate my
         | homework" mode. Blaming everyone and everything around.
         | 
         | Let's not mention the fact that if full outputs of their models
         | could enable them to train a better model at 1% cost then it
         | puts them in even worse light that they didn't do it.
        
           | ryanobjc wrote:
           | It's not often you get 100x optimization with some small
           | improvements so I'm kind of skeptical.
           | 
           | We have and apples and oranges thing here which deepseek is
           | intentionally leaning into. They get very cheap electricity
           | and are bragging about their cheap cost, and OpenAI etc
           | typically brag about how expensive their training is. But
           | it's all pr and lies.
        
             | enragedcacti wrote:
             | > They get very cheap electricity and are bragging about
             | their cheap cost
             | 
             | The cost of $5.5 million was quoted at $2/GPU-hour which is
             | a reasonable price for on-demand H100s that anyone in the
             | US could access, and likely on the high side given bulk
             | pricing and that they are using nerfed versions. OpenAI
             | might be all pr and lies but everything I've seen so far
             | says that deepseek's claims about cost are legit.
        
         | Leary wrote:
         | Does this mean when you use OpenAI as an enterprise customer,
         | they can see exactly the queries and answers? So much for
         | privacy!
        
         | api wrote:
         | So far the whole business model of Silicon Valley since social
         | media has been to monetize other peoples' content given out for
         | free. The whole empire is built on this.
         | 
         | I wonder if this is going to come to an end through a
         | combination of social media fatigue, social media
         | fragmentation, and open source LLMs just giving it all back to
         | us for free. LLMs are analogous to a "JPEG for ideas" so
         | they're just lossy compression blobs of human thought expressed
         | through language.
        
           | barnabee wrote:
           | > So far the whole business model of Silicon Valley since
           | social media has been to monetize other peoples' content
           | given out for free. The whole empire is built on this.
           | 
           | It cannot die soon enough
        
         | okdood64 wrote:
         | Not to mention the total dodge when Murati was asked about
         | training on the YouTube corpus during that television
         | interview.
         | 
         | Sorry for the Short: https://www.youtube.com/shorts/M0QyOp7zqcY
        
         | the_arun wrote:
         | I think the point is - OpenAI scraped public data - d1 -
         | Trained their model to produce output - d2 - DeepSeek used d2
         | to reinforce their model
         | 
         | OpenAI is mad about d2 (not d1). I'm not sure using public data
         | is "stealing". In summary, these are two different things &
         | need to be separate.
        
           | redleader55 wrote:
           | You say "public", but what I think you mean is "publicly
           | available". Even publicly available data has copyrights, and
           | unless that copyright is "public domain", you need to follow
           | some rules. Even licenses like Creative Commons, which would
           | be the most permissive, come with caveats which OpenAI
           | doesn't follow [0].
           | 
           | It is unclear if someone breaking someone else's copyright to
           | use A can claim copyright on a work B, derived from A. My
           | point is that OpenAI played loose with the copyright rules to
           | build its various models, so the legality of their claims
           | against DeepSeek might not be so strong.
           | 
           | [0] https://creativecommons.org/share-your-work/cclicenses/
        
             | the_arun wrote:
             | I am not saying OpenAI did good by using publicly available
             | data. I meant these are separate activities. None is good.
             | But DeepSeek is slightly better by making theirs
             | opensource.
        
           | xbar wrote:
           | OpenAI (sc)raped all the data it could. I do not accept your
           | assertion that d1 was "public." It was accessible, for
           | certain.
           | 
           | OpenAI asserts 1. d2 was used by DeepSeek 2. All d2 belongs
           | to OpenAI exclusively
           | 
           | Both are debatable for large number of reasons.
        
         | rubslopes wrote:
         | > Our mission is to ensure that artificial general intelligence
         | benefits all of humanity.[1]
         | 
         | Well, I guess they really helped make this a reality!
         | 
         | [1] https://openai.com/about/
        
         | didip wrote:
         | fr fr, ClosedAI is being a comedian right now.
         | 
         | They scraped literally all the content of the internet without
         | permissions. And I won't even be surprised if they scraped the
         | output of other LLMs as well.
        
         | adzm wrote:
         | Why does this post use DeepSink instead of DeepSeek at
         | apparently random places? Is that just a pejorative pun like
         | ClosedAI?
        
         | skeeter2020 wrote:
         | I share the sentiment here, but asking as a noob: does this
         | mean the performance comparison is not really apples to apples?
         | If it required the distillation of the expensive model in order
         | to get such good results for a much lower price, is that shady
         | accounting?
        
         | schmit wrote:
         | Even more hilarious given their own charter:
         | 
         | > We will attempt to directly build safe and beneficial AGI,
         | but will also consider our mission fulfilled if our work aids
         | others to achieve this outcome.
         | 
         | > Our primary fiduciary duty is to humanity. We anticipate
         | needing to marshal substantial resources to fulfill our
         | mission, but will always diligently act to minimize conflicts
         | of interest among our employees and stakeholders that could
         | compromise broad benefit.
         | 
         | > We will actively cooperate with other research and policy
         | institutions; we seek to create a global community working
         | together to address AGI's global challenges.
        
           | semking wrote:
           | Ah yes: "duty to humanity"
        
             | hn_throwaway_99 wrote:
             | I think one good thing to come out of all this tech elite
             | flip flopping is that I now see these tech leaders for
             | exactly who they are. It makes me kind of sad, because as
             | someone who came of age early in the Web era I really
             | _wanted_ to believe that there was a bigger moral good to
             | all we were doing.
             | 
             | I now view _any_ moralistic statement by any of these big
             | tech companies as complete and total bullshit, which is
             | probably for the best, because that is what it is. These
             | companies now exist solely to amass power and wealth. They
             | will still use moralistic language to try to motivate their
             | employees, but I hope folks still see it for the complete
             | nonsense that it is.
        
         | radicality wrote:
         | I liked Matt Levine's newsletter few days ago where he
         | hypothesized scenarios where it's much more profitable to short
         | your competitors, then release a much better version of some
         | widget completely free, and then profit $$$. Which is plausible
         | here too, considering DeepSeek is made by a hedge fund.
        
           | greasegum wrote:
           | Came here to mention this too. Seem almost so obvious that
           | I'm surprised this isn't the dominant angle.
        
           | freehorse wrote:
           | How would that work out here though? "Open"AI is not publicly
           | traded. Any kind of shorting would be quite indirect.
        
         | Imustaskforhelp wrote:
         | Yes the irony is so thick in the air that it can be cut through
         | using a swiss knife lol
         | 
         | I had literally come to this post to say the same. You beat me
         | to it.
         | 
         | USA is going crazy over deepseek and to me , it just shows that
         | the world is a black swan , an AI bubble.
         | 
         | I am not saying AI has no use. I regularly use it to create
         | something , but its just not recommended. I am going to stop
         | using AI , to grow my mind.
         | 
         | And its definitely way overpriced. People are investing so much
         | money without seeing the returns? , and I think people are also
         | using AI because of a sense of FOMO , I don't know , to me its
         | funny .
         | 
         | I really really want to create a index fund with strictly no AI
         | companies. Since this doesn't feel diversified enough. Like
         | sure nvidia gave a quarter of return the last year , but I mean
         | , at this point , it almost feels the same as that of bitcoin.
         | The reason I don't / won't invest in bitcoin is I don't want
         | "that" risk.
         | 
         | This has been a boggling year.
         | 
         | I have realized that the world is crazy. Truly. Trump winning
         | from going to the point of getting shot to deepseek causing
         | nvidia / american stock market to go down , heck even bitcoin!
         | , its so crazy , trump launching his meme coin. If the world is
         | crazy. Just be the sane person around. You will stick around ,
         | that's my philosophy. I won't jump on AI wandwagon . But its
         | still absolutely wild & horror seeing how a "sideproject"
         | (deepseek) absolutely put american stock market in shambles.
         | 
         | I want more diversifaction. I am not satisfied with the current
         | system. This feels like a bubble and I want no part in it.
        
         | amelius wrote:
         | It looks like they want to spin this as "DeepSeek copied
         | OpenAI". The general public/media might actually believe this
         | is what happened.
        
       | breakitmakeit wrote:
       | As the article points out, they are arguing in court against the
       | new york times that publicly available data is fair game.
       | 
       | The questions I am keenly waiting to observe the answer to
       | (because surely Sam's words are lies): how hard is OpenAI willing
       | to double down on their contradictory positions? What mental
       | gymnastics will they use? What power will back them up, how, and
       | how far will that go?
        
         | snakeyjake wrote:
         | When large sums of money are involved the techbros will burn
         | everything down, go scorched earth no matter what the
         | consequences, to keep what they believe they're entitled to.
        
         | ADeerAppeared wrote:
         | Their way of squaring this circle has always been to whine
         | about "AI safety". (the cultish doomsday shit, not actual harms
         | from AI)
         | 
         | Sam Altman will proclaim that he alone is qualified to build AI
         | and that everyone else should be tied down by regulation.
         | 
         | And it should always be said that this is, of course, utterly
         | ridiculous. Sam Altman literally got fired over this, has an
         | extensive reputation as a shitweasel, and OpenAI's constant
         | flouting and breaking of rules and social norms indicates they
         | CANNOT be trusted.
        
       | bhouston wrote:
       | The US government likely will favor a large strategic company
       | like OpenAI instead of individual's copyrights, so while ironic,
       | the US government definitely doesn't care.
       | 
       | And the US government is also likely itching to reduce the power
       | of Chinese AI companies that could out compete US rivals (similar
       | to the treatment of BYD, TikTok, solar panel manufacturers,
       | network equipment manufacturers, etc), so expect sweeping
       | legislation that blocks access to all Chinese AI endeavours to
       | both the US and then soon US allies/West (via US pressure.)
       | 
       | The likely legislation will be on the surface justified both by
       | security concerns and by intellectual property concerns, but
       | ultimately it will be motivated by winning the economic
       | competition between China and the US and it will attempt to tilt
       | the balance via explicitly protectionist policies.
        
         | derektank wrote:
         | >The US government likely will favor a large strategic company
         | like OpenAI instead of individual's copyrights
         | 
         | Even if we assume this is true, Disney and Netflix are both
         | currently worth more than OpenAI and both rely on the strict
         | enforcement of US copyright law. I do not think it is so
         | obvious which powers that be have the better lobbying efforts
         | and, currently, it's looking like this question will mostly be
         | adjudicated by the courts, not Congress, anyways.
        
           | bhouston wrote:
           | I don't think OpenAI stole from Disney or Netflix. Rather
           | OpenAI stole from individual artists and YouTube and other
           | social media who users do not really have any lobbying power.
           | 
           | So I think OpenAI, Disney and Netflix win together. Big
           | companies tend to win.
        
             | mjburgess wrote:
             | > What are the first words of the disney movie, "Aladdin" ?
             | 
             | The first words of Disney's _Aladdin_ (1992) are spoken by
             | the *Peddler*, the mysterious merchant at the beginning of
             | the film. He says:
             | 
             |  _" Ah, Salaam and good evening to you, worthy friend.
             | Please, please, come closer..."_
             | 
             | He then continues with: _" Too close! A little too close.
             | There. Welcome to Agrabah. City of mystery, of enchantment,
             | and the finest merchandise this side of the River Jordan,
             | on sale today! Come on down!"_
             | 
             | This opening sets the stage for the story, introducing the
             | magical and bustling world of Agrabah.
        
             | derektank wrote:
             | Disney owns ABC News; OpenAI almost certainly scraped their
             | text data
        
               | bhouston wrote:
               | I agree with you.
        
             | worik wrote:
             | > Rather OpenAI stole from individual artists and YouTube
             | and other social media
             | 
             | "stole"?
             | 
             | They consumed publicly available material on the Internet
             | 
             | I am no fan of these billionaire capitalists and their
             | henchpersons but condem them for their multitude of sins.
             | 
             | Consuming publicly available Internet resources is not one
             | of them. IMO
        
               | da_chicken wrote:
               | Being publicly available does not mean that copyright is
               | invalid. Copyright gives the holders the right to
               | restrict USE, not merely restrict reproduction.
               | _Adaptation_ is also an exclusive right of the copyright
               | holder. You 're not allowed to make derivative works.
        
               | visarga wrote:
               | They stole the data just as much as a painter steals the
               | view.
        
               | rideontime wrote:
               | Who created the view?
        
               | visarga wrote:
               | The view is created by every spectator.
        
               | jdswain wrote:
               | It's not that they consumed publicly available material,
               | it's that they re-published that information, and sold
               | it.
        
               | Terr_ wrote:
               | > They consumed publicly available material on the
               | Internet
               | 
               | I agree that there are some important distinctions and
               | word-choices to be made here, and that there are problems
               | with equating training to "stealing", and that copyright
               | infringement is not theft, etc.
               | 
               |  _That said_ , if you zoom out to the overall conduct,
               | it's fair to argue that the companies are doing something
               | unethical, the same as if they paid an army of humans to
               | memorize other people's work and then regurgitate
               | slightly-reworded copies.
        
         | tokioyoyo wrote:
         | I don't think US government can move fast enough to change the
         | trajectory. Also it doesn't help that basically every
         | government is second guessing their alliance with the US. It's
         | not an industry that can ruin local industries either (like
         | cheap BYD is bad for German cars).
         | 
         | It's a very fun thing to watch from the sidelines right now, if
         | I'll be honest.
        
         | buyucu wrote:
         | It's too late for that. That ship sailed a long time ago.
         | 
         | The best language model right now is open source. Let that sink
         | in.
        
           | _pferreir_ wrote:
           | DeepSeek is not Open Source. That's like saying that
           | Microsoft Edge is Open Source, as you can download it for
           | free.
           | 
           | https://huggingface.co/blog/open-r1
        
       | ceejayoz wrote:
       | "You can't take data without asking" seems like a court precedent
       | OpenAI really, really, _really_ wants to avoid. And yet...
        
         | amelius wrote:
         | Why? When did large companies care about laws? See e.g. Uber,
         | AirBnb.
         | 
         | The only thing government cares about at this point is if
         | information is shared with China.
        
           | ceejayoz wrote:
           | They care when they get big enough to attract attention from
           | people like state AGs who can actually put the hurt on a bit.
           | Uber and AirBnB both hit this point years ago; OpenAI's
           | starting to hit it.
        
             | galleywest200 wrote:
             | Altman is part of that Stargate Trump group now. He and his
             | ilk will just get pardons.
             | 
             | Curious, though, can a corporation be pardoned?
        
               | ceejayoz wrote:
               | The President can only pardon Federal crimes.
               | 
               | State-level crimes (like his NY felonies) and civil torts
               | (like his case where he owes $500M currently) are
               | separate.
        
               | actionfromafar wrote:
               | Yet. Give it some time.
        
               | ceejayoz wrote:
               | Sure, but in that scenario, it's a bit like the Last of
               | Us characters being concerned about electrical meter
               | readings. We'll have much bigger problems.
        
         | layer8 wrote:
         | OpenAI is saying that their service was used in violation of
         | their TOS, which is a bit different than just copying data. To
         | be clear I'm not on OpenAI's side, but it looks to me that the
         | legal situation isn't exactly analogous.
        
           | DebtDeflation wrote:
           | Tons of websites and books they scraped had copyright
           | notices.
        
             | layer8 wrote:
             | Copyright and terms of service are different legal notions.
        
               | Maxion wrote:
               | Yeah, copyright means something and a ToS is virtual
               | toiletpaper (at least in the EU)
        
               | layer8 wrote:
               | This wasn't about which is worse than the other, but
               | about whether OpenAI would want to avoid court precedent
               | for the one because of the other.
        
           | orlp wrote:
           | If using data violating some ToS taints the model trained on
           | that data, then all of OpenAI's models are tainted by the
           | millions of ToS'es they broke.
        
             | orionsbelt wrote:
             | Can you cite a source showing they violated ToS?
        
               | hdjjhhvvhga wrote:
               | Not just violated but also actively ignored:
               | https://news.ycombinator.com/item?id=42718850
        
           | dkjaudyeqooe wrote:
           | But whats the remedy in that case? Being banned from the
           | service maybe, but no court is going to force a "return" of
           | the data, so DeepSeek can't use it. It's uncopyrightable.
        
           | kavalg wrote:
           | As others have noted, if one company agrees to the ToS, asks
           | "the right" questions and then publishes the ChatGPT answers,
           | there is not violation of ToS. Then a second company scrapes
           | the published Q&A, along with other information from the
           | internet and again there is no violation (not more than the
           | violations of OpenAI).
        
           | hdjjhhvvhga wrote:
           | > OpenAI is saying that their service was used in violation
           | of their TOS
           | 
           | Which is the most ridiculous argument they could use because
           | they didn't respect any ToS (or copyright laws, for that
           | matter) when scraping the whole web, books from Libgen and
           | who knows what more.
        
       | osigurdson wrote:
       | I do think that distilling a model from another is much less
       | impressive than distilling one from raw text. However, it is hard
       | to say if it is really illegal or even immoral, perhaps just one
       | step further in the evolution of the space.
        
         | lemoncookiechip wrote:
         | It's about as illegal as the billions, if not trillions of IPs
         | that ClosedAI infringed to train their own data without
         | consent. Not that they're alone, and I personally don't mind
         | that AI companies do it, but it's still amusing when they get
         | this annoyed at others doing the same thing to them.
        
           | osigurdson wrote:
           | I think they had the advantage of being ahead of the law in
           | this regard. To my knowledge, reading copywritten material
           | isn't (or wasn't illegal) and remains a legal grey area.
           | 
           | Distilling weights from prompts and responses is even more of
           | a legal grey area. The legal system cannot respond quickly to
           | such technological advancements so things necessarily remain
           | a wild west until technology reaches the asymptotic portion
           | of the curve.
           | 
           | In my view the most interesting thing is, do we really need
           | vast data centers and innumerable GPUs for AGI? In other
           | words, if intelligence is ultimately a function of power
           | input, what is the shape of the curve?
        
             | ttesmer wrote:
             | > if intelligence is ultimately a function of power input,
             | what is the shape of the curve?
             | 
             | According to a quick google search, the human body consumes
             | ~145W of power over 24h (eating 3000kcals/day). The brain
             | needs ~20% of that so 29W/day. Much less than our current
             | designs of software & (especially) hardware for AI.
        
               | osigurdson wrote:
               | I think you mean the brain uses 29W (i.e. not 29W/day).
               | Also, I suspect that burgers are a higher entropy energy
               | source than electricity so perhaps it is even less than
               | that.
        
             | lemoncookiechip wrote:
             | The main issue is that they've had plenty of instances
             | where the LLM outputted copyrighted content verbatim, like
             | it happened with the New York Times and some book authors.
             | And then there's DALL-E, which is baked into ChatGPT and
             | before all the guardrails came up, was clearly trained on
             | copyrighted content to the point it had people's
             | watermarks, as well as their styles, just like Stable
             | Diffusion mixes can do (if you don't prompt it out).
             | 
             | Like you've put, it's still a somewhat gray area, and I
             | personally have nothing against them (or anyone else) using
             | copyrighted content to train models.
             | 
             | I do find it annoying that they're so closed-off about
             | their tech when it's built on the shoulders of openness and
             | other people's hard work. And then they turn around and
             | throw Issy fits when someone copies their homework,
             | allegedly.
        
             | JTyQZSnP3cQGa8B wrote:
             | Illegally acquiring copyrighted material has always been
             | highly illegal in France and I'm sure most other countries.
             | Disney is another example of how it not grey at all.
        
             | greiskul wrote:
             | > Distilling weights from prompts and responses is even
             | more of a legal grey area.
             | 
             | Actually unless the law changes this is pretty settled
             | territory in US law. All output of AIs are not
             | copyrightable, and are therefore in the public domain. The
             | only legal avenue of attack OpenAi has is Terms of Service
             | violation, which is a much weaker breach then copyright if
             | it is even true.
        
           | ReptileMan wrote:
           | Is the question of training AI on data fair use settled yet?
           | Because if it is not - it looks like fair use to me.
        
         | scotty79 wrote:
         | Isn't it more impressive given that training on model output
         | usually leads to worse model?
         | 
         | If they actually figured out how to use output of existing
         | models to build model that outperforms them then it's something
         | that brings us closer to singularity than every other
         | development so far.
        
       | __MatrixMan__ wrote:
       | If they want us to care they can open up their models so we can
       | be the judge.
        
       | 827a wrote:
       | This smells very suspiciously like: someone who doesn't know
       | anything about AI (possibly Sacks) demanding answers on R1 from
       | someone who doesn't have any good ones (possibly Altman). "Uh,
       | (sweating), umm, (shaking), they stole it from us! Yeah, look at
       | this suspicious activity, that's why they had it so easy, we did
       | all the hard work first!"
        
         | fundad wrote:
         | I think it's funny that OpenAI wants us to pay them to use
         | their product to generate content but then sets the terms that
         | they control how we use the content in generates for us. It
         | takes someone like Deepseek to challenge that on our behalf or
         | they will control most of the economy.
        
           | exitb wrote:
           | It's quite ironic of them to claim that the only thing you
           | cannot train on is another LLM output.
        
       | 1970-01-01 wrote:
       | DeepSeek have more integrity than 'Open'AI by not even pretending
       | to care about that.
        
         | jampekka wrote:
         | And seem to be more actively fulfilling the mission that
         | 'Open'AI pretends to strive for.
        
           | pixelpoet wrote:
           | Exactly, they _actually_ opened up the model and research,
           | which the  "Open" company didn't, and merely adjusted some of
           | their pricing tiers to try to combat commercially (but not
           | without mumbling something like "yeah, we totally had these
           | ideas too"). Now every single Meta, OpenAI etc engineer is
           | trying to copy DeepSeek's innovations, and their first act is
           | to... complain about copyright infringement, of all things?!
           | What an absolute clown party, how can these people take
           | themselves seriously, do they just have zero comprehension of
           | what hypocrisy is or what's going on here...
           | 
           | I can scarcely process all the levels of irony involved, the
           | irony-o-meter is pegged and I can't get the good one from the
           | safe because I'm incapacitated from laughter.
        
       | sylware wrote:
       | LOL, I was thinking exactly the same think when I read the news
       | about openai whining.
        
       | WD-42 wrote:
       | Information wants to be free! No, not like that!
        
       | asah wrote:
       | Thieve's honor, hunh?
        
       | nba456_ wrote:
       | A big part of project 2025 is increasing patent regulations. I
       | would not be surprised if the current admin moves to ban DeepSeek
       | because of this.
        
       | typon wrote:
       | OpenAI is the MIC darling - expect more ridiculous attacks on
       | competitors in the future
        
       | sho_hn wrote:
       | While I'm as amused as everyone else - I think it's technically
       | accurate to point out that the "we trained it for $6 mio"
       | narrative is contingent on the done investment by others.
        
         | bbqfog wrote:
         | OpenAI's models were also trained on billions of dollars of
         | "free" labor that produced the content that it was trained on.
        
           | sho_hn wrote:
           | Oh, absolutely. I'm not defending OpenAI, I just care about
           | accurate reporting. Even on HN - even in this thread - you
           | see people who came away with the conclusion that DeepSeek
           | did something while "cutting cost by 27x".
           | 
           | But that's a bit like saying that by painting a a bare wall
           | green you have demonstrated that you can build green walls
           | 27x cheaper, ignoring the cost of building the wall in the
           | first place.
           | 
           | Smarter reporting and discourse would explain how this
           | iterative process actually works and who is building on who
           | and how, not frame it as two competing from-scratch clean
           | room efforts. It'd help clear up expectations of what's
           | coming next.
           | 
           | It's a bit similar to how many are saying DeepSeek have
           | demonstrated independence from nVidia, when part of the
           | clever thing they did was figure out how to make the
           | intentionally gimped H800s work for their training runs by
           | doing low-level optimizations that are _more_ nVidia-
           | specific, etc.
           | 
           | Rarely have I seen a highly technical topic see produce more
           | uninformed snap takes than this week.
        
             | bbqfog wrote:
             | I don't agree. Walls are physical items so your example is
             | true, but models are data. Anyone can train off of these
             | models, that's the current environment we exist in. Just
             | like OpenAI trained on data that has since been locked up
             | in a lot of cases. In 2025 training models like Deepseek is
             | indeed 27x cheaper, that includes both their innovations
             | and the existence of new "raw material" to do such a thing.
        
               | sho_hn wrote:
               | I don't think we disagree at all, actually!
               | 
               | What I'm saying is that in the media it's being portrayed
               | as if DeepSeek did _the same thing OpenAI did_ 27x
               | cheaper, and the outsized market reaction is in large
               | parts a response to that narrative. While the reality is
               | more that being a fast-follower is cheaper (and the
               | concrete reason is e.g. being able to source training
               | data from prior LLMs synthetically, among other things),
               | which shouldn 't have surprised anyone and is just how
               | technology in general trends.
               | 
               | The achievement of DeepSeek is putting together a
               | competent team that excels at end-to-end implementation,
               | which is no small feat and is promising wrt/ their future
               | efforts.
        
               | meiraleal wrote:
               | How much money a third company would need to spend to
               | achieve what OpenAI achieved to compete with them,
               | 5billion or 6million?
        
             | Palmik wrote:
             | You are underselling or not understanding the breakthrough.
             | They trained 600B model on 15T tokens for <$6/m. Regardless
             | of the provenance of the tokens, this in itself is
             | impressive.
             | 
             | Not to mention post-training. Their novel GRPO technique
             | used for preference optimization / alignment is also much
             | more efficient than PPO.
        
               | sho_hn wrote:
               | Let's call it underselling. :-) Mostly because I'm not
               | sure anyone's independently done the math and we just
               | have a single statement from the CEO. I do appreciate the
               | algorithmic improvements, and the excellent attention-to-
               | performance-in-detail stuff in their implementation
               | (careful treatment of precision, etc.), making the H800s
               | useful, etc. I agree there's a lot there.
        
             | visarga wrote:
             | > that's a bit like saying that by painting a a bare wall
             | green you have demonstrated that you can build green walls
             | 27x cheaper, ignoring the cost of building the wall in the
             | first place
             | 
             | That's a funny analogy, but in reality DeepSeek did
             | reinforcement learning to generate chain of thought, which
             | was used in the end to finetune LLMs. The RL model was
             | called DeepSeek-R1-Zero, while the SFT model is
             | DeepSeek-R1.
             | 
             | They might have boostrapped the Zero model with some
             | demonstrations.
             | 
             | > DeepSeek-R1-Zero struggles with challenges like poor
             | readability, and language mixing. To make reasoning
             | processes more readable and share them with the open
             | community, we explore DeepSeek-R1, a method that utilizes
             | RL with human-friendly cold-start data.
             | 
             | > Unlike DeepSeek-R1-Zero, to prevent the early unstable
             | cold start phase of RL training from the base model, for
             | DeepSeek-R1 we construct and collect a small amount of long
             | CoT data to fine-tune the model as the initial RL actor. To
             | collect such data, we have explored several approaches:
             | using few-shot prompting with a long CoT as an example,
             | directly prompting models to generate detailed answers with
             | reflection and verification, gathering DeepSeek-R1Zero
             | outputs in a readable format, and refining the results
             | through post-processing by human annotators.
        
         | Palmik wrote:
         | When I use NVIDIA GPUs to train a model, I do not consider the
         | R&D cost to develop all of those GPUs as part of my costs.
         | 
         | When I use an API to generate some data, I do not consider the
         | R&D cost to develop the API as part of my costs.
        
         | kobalsky wrote:
         | OpenAI has been in a war-room for days searching for a match in
         | the data, and they just came out with this without providing
         | proof.
         | 
         | My cynical opinion is that the traning corpus has some small
         | amount of data generated by OpenAI, which is probably
         | impossible to avoid at this point, and they are hanging on that
         | thread for dear life.
        
         | scotty79 wrote:
         | The opposite, is claiming that OpenAI could have now built
         | better performing, cheaper to run model (when compared to what
         | they published) training it at 1% cost on output of their
         | previous models. ... But they chose not to do it.
        
         | freehorse wrote:
         | That is the case anyway for training any llm. It is contingent
         | on the work done by all those who produced the data.
        
       | pcthrowaway wrote:
       | Now that China is talking about lifting the Great Firewall, it
       | seems like the U.S. is on track to cordon themselves off from
       | other countries. Trump's talk of building a wall might not stop
       | at Mexico.
        
       | temporallobe wrote:
       | OpenAI is also possibly in violation of many IP laws by scraping
       | the entirety of the internet and using to train their models, so
       | there's that.
        
         | InkCanon wrote:
         | To my understanding, OpenAI won the case where it argued
         | training was covered under fair use and did not infringe on
         | copyright.
        
           | Austiiiiii wrote:
           | Is there any reason they wouldn't rule the same way on
           | DeepSeek training on OpenAI data? After all, one of the big
           | selling points of GPT has been that businesses can freely use
           | the information provided. They're paying for the service,
           | after all. I'd very be interested to know how DeepSeek's
           | usage (very reasonably assuming that they paid for their
           | OpenAI subscription) is any different.
        
             | ickelbawd wrote:
             | Businesses _can't_ freely use the information. There are
             | terms of service freely agreed upon by the user which
             | explicitly deny many use cases--training other models is
             | just one. DeepSeek is not an American company nor is their
             | leader in deep with the new administration. It seems far
             | more likely that this will play out like tiktok--they'll be
             | attacked publicly and banned for national security reasons.
        
               | Austiiiiii wrote:
               | On further reading, I'll grant the first point. Although
               | I wonder if they'll have a technical out--say they
               | distilled from several smaller research companies that
               | had distilled from OpenAI for research purposes, which to
               | my understanding would not constitute a violation of the
               | terms of service.
               | 
               | As for it getting banned, TikTok was banned partly
               | because of credible accounts of it having been used by
               | China to track political enemies. Are we thinking they'll
               | expand the argument on national security to say that
               | _any_ application that transfers data to China is a
               | national security threat? Because that could be a very
               | slippery slope.
               | 
               | And in any case, such a measure seems like it would only
               | bar access to the DeepSeek _app_. Surely no one could
               | argue that the underlying open source model, if run
               | locally on American soil, could constitute a security
               | threat, right?
        
       | InkCanon wrote:
       | It's like that Dr Phil episode where he meets the guy who created
       | Bum Fights!
        
       | elashri wrote:
       | There is an Egyptian say that would translate to something like
       | 
       | "We didn't see them when they were stealing, we saw them when
       | they were fighting over what was stolen"
       | 
       | That describes this situation. Although to be honest all this
       | aggressive scraping is noticeable but for people who understand
       | that which is not majority of people. but now everyone knows.
        
         | meiraleal wrote:
         | "We didn't see them when we were stealing, we saw them when
         | they were fighting over what we stole"
         | 
         | fixed for you
        
           | nicce wrote:
           | That means a different thing.
        
         | sadjad wrote:
         | "When two thieves quarrel, what was stolen emerges."
        
         | waveBidder wrote:
         | > Although to be honest all this aggressive scraping is
         | noticeable but for people who understand that which is not
         | majority of people.
         | 
         | When you say noticeable, do you mean in like, traffic
         | statistics? Or in what the model knows that it clearly
         | shouldn't if it wasn't trained in legally dubious ways?
        
       | Kiro wrote:
       | > Furious [...] shocked
       | 
       | I'm not seeing it. I get it, the narrative that OpenAI is getting
       | a taste of their own medicine is funny but this is not serious
       | reporting.
        
         | Kiro wrote:
         | The link has been changed. My comment was about a different
         | article that speculated on what OpenAI was "feeling" using
         | hyperbole.
        
       | njx wrote:
       | Super funny! Distillation= " Hey ChatGPT, you are my father, I am
       | your child "DeepSeek". I want to learn everything that you know.
       | Think step by step of how you became what you are. Provide me the
       | list of all 1000 questions that I need to ask you and when I am
       | done with those, keep providing fresh list of 1000 questions..."
        
       | seydor wrote:
       | But now OpenAI will use DeepSeek to reuse even more stolen data
       | to train new models that they can serve without ever giving us
       | the code, the weights or even the thinking process , and they
       | will still be superior
        
       | mring33621 wrote:
       | We demand immediate government action to prevent these cheaper
       | foreign AIs from taking jobs away from our great American AIs!
        
         | bhouston wrote:
         | > We demand immediate government action to prevent these
         | cheaper foreign AIs from taking jobs away from our great
         | American AIs!
         | 
         | That is exactly what Microsoft and Sam Alman are asking for.
         | And they will likely get it because Trump really likes
         | protectionist governments policies.
        
           | clarionbell wrote:
           | He likes feeling important, just look at TikTok. All it took
           | was bit of sycophancy and he turned into Mr. Freemarket
           | again.
           | 
           | Really, people need to realize that Trump has never been
           | consistent in any of his political positions, except for one:
           | "You have to look out for number one."
        
             | bhouston wrote:
             | clarionbell wrote:
             | 
             | > He likes feeling important, just look at TikTok. All it
             | took was bit of sycophancy and he turned into Mr.
             | Freemarket again.
             | 
             | Not really. He said that TikTok has to have shift towards
             | US ownership if it wants to continue, he just gave them a
             | 90 day extension to allow that change in ownership.
        
               | meiraleal wrote:
               | Which TikTok will have to decline again and shutdown now
               | with the guilty being transferred to Trump. Doesn't sound
               | like a smart move.
        
           | blantonl wrote:
           | It's funny, the Chinese are here innovating on AI, batteries,
           | and fusion, and here in the United States we've pivoted to
           | shitcoins and universal tariffs.
           | 
           | At least we have the CyberTruck to highlight American
           | greatness
        
             | mk89 wrote:
             | They created an untameable beast (China) thanks to the
             | "cheap factories" there and they thought they would stay
             | that way.
             | 
             | What a bunch of idiots. The propaganda keeps telling us
             | that they don't invent, they can only copy etc., but
             | clearly that's not true.
        
         | bhargav wrote:
         | This is gonna be spun up as a security thing, and banned cozz
         | Murica.
        
         | cactusplant7374 wrote:
         | To the detriment of OpenAI, the math is going to be used to
         | improve AIs developed in America. And we need to remember that
         | Marc Andreessen is very against government banning maths.
        
         | Dansvidania wrote:
         | does it matter if the company gets banned? other non-chinese
         | companies can pick up the open source model and run it as a
         | service with relatively low investment, isn't that the point?
        
       | RohMin wrote:
       | this comment section smells like Reddit - ugh
        
       | JBits wrote:
       | What is the evidence that DeepSeek used OpenAI to train their
       | model? Isn't this claim directly benefitting OpenAI as they can
       | argue that any superior model requires their model?
        
       | nottorp wrote:
       | IP thief cries IP thief.
       | 
       | It's okay when you steal worldwide IP to train your "AI".
       | 
       | It's not okay when said stolen IP is stolen from you?
       | 
       | If the chinese are guilty, then Altman's doom and gloom racket is
       | as guilty or even more, considering they stole from everyone.
        
       | Ciantic wrote:
       | I'm not being sarcastic, but we may soon have to torrent
       | DeepSeek's model. OpenAI has a lot of clout in the US and could
       | get DeepSeek banned in western countries for copyright.
        
         | alchemist1e9 wrote:
         | I think most likely all sorts of data and models need to have a
         | decentralized LLM data archive via torrents etc.
         | 
         | It's not limited to the models themselves but also OpenAI will
         | probably work towards shutting down access to training data
         | sets also.
         | 
         | imho it's probably an emergency all hand on deck problem.
        
         | timeon wrote:
         | > US and could get DeepSeek banned in western countries for
         | copyright
         | 
         | If US is going to proceed with trade war on EU, as it was
         | planning anyway, then DeepSeek will be banned only in US. Seems
         | like term "western countries" is slowly eroding.
        
           | bbor wrote:
           | Great point. Plus, the revival of serious talk of the Monroe
           | Doctrine (!!!) in the U.S. government lends a possibly
           | completely-new meaning to "western countries" -- i.e. the
           | Americas...
        
             | surgical_fire wrote:
             | Except the US has only contempt for anything south of
             | Texas. Perhaps "western countries" will be reduced to US
             | and Canada.
             | 
             | Many countries in Latin America have better relations and
             | more robust trade partnerships with China.
             | 
             | As for the EU, I think it will be great for it to shed its
             | reliance on the US, and act more independently from it.
        
               | ta1243 wrote:
               | The US is talking about annexing Canada, so "western
               | countries" means the USA, which if continuing down this
               | path long enough will become a pariah
        
             | marcosdumay wrote:
             | Only if they do it by force.
             | 
             | Trump has already managed to completely destroy the US
             | reputation within basically the entire continent1. And he
             | seems intent on creating a commercial war against all the
             | countries here too.
             | 
             | 1 - Do not capture and torture random people on the street
             | if you want to maintain some goodwill. Even if you have
             | reasons to capture them.
        
         | aerhardt wrote:
         | Unfathomable to me that they'd make themselves look so foolish
         | by trying to ban a piece of software.
        
         | sergiotapia wrote:
         | that would be suicide - that company only exists because they
         | stole content for every single person, website and media
         | company on the planet.
        
       | sonabinu wrote:
       | poetic justice (pun intended)
        
       | readyplayernull wrote:
       | Do you remember when Microsoft was caught scrapping data from
       | Google:
       | 
       | https://www.wired.com/2011/02/bing-copies-google/
       | 
       | They don't care, T&C and copyright is void unless it affects
       | them, others can go kick rocks. Not surprising they and OpenAI
       | will do a legal battle over this.
        
       | SilverBirch wrote:
       | I think OpenAI is in a really weak position here. There are
       | essentially two positions you can be in: You can be the agile new
       | startup that can break the rules and move fast. That's what
       | OpenAI used to be. Or you can be the big incumbent who is going
       | to use your enormous resources to crush your opposition. That's
       | Google & Microsoft here. For Microsoft to say "We're going to tie
       | you up in lawsuits about the way you trained this model" would be
       | perfectly expected and they can use that strategy because at any
       | given time they have 1,000 lawyers and lobbyists hanging around
       | waiting to do exactly that. But OpenAI can't do that. They don't
       | have Google or Microsoft's legal teams or lobbyists or
       | distribution channels. SO whilst it's funny that OpenAI are kind
       | of trying to go down this road, this isn't actually a strategy
       | that is going to work for them, they're still a minnow and
       | they're going to get distracted and slowed down by this.
        
         | htrp wrote:
         | But microsoft is one of their backers?
        
       | bilekas wrote:
       | > "It's also extremely hard to rally a big talented research team
       | to charge a new hill in the fog together," he added. "This is the
       | key to driving progress forward."
       | 
       | Well I think DeepSeek releasing it open source and on an MIT
       | license will rally the big talent. The open sourcing of a new
       | technology has always driven progress in the past.
       | 
       | The last paragraph too is where OpenAi seems to be focusing their
       | efforts..
       | 
       | > we engage in countermeasures to protect our IP, including a
       | careful process for which frontier capabilities to include in
       | released models ..
       | 
       | > ... we are working closely with the US government to best
       | protect the most capable models from efforts by adversaries and
       | competitors to take US technology.
       | 
       | So they'll go for getting DeepSeek banned like TikTok was now
       | that a precedent has been set ?
        
         | hujun wrote:
         | or sold to US I could totally see this happening soon
        
           | trissi1996 wrote:
           | Why would they want to sell ?
        
             | kavalg wrote:
             | And what are they going to sell? The weights and the model
             | architecture are already open source. I doubt the datasets
             | of DeepSeek are better than OpenAI's
        
               | Dansvidania wrote:
               | plus, if the US were to decide to ban DeepSeek (the
               | company) wouldn't non-chinese companies be able to pick
               | up the models and run them at a relatively low expense?
        
         | worik wrote:
         | > getting DeepSeek banned like TikTok
         | 
         | Banned in the USA. Only.
         | 
         | There is a big wide world out there, and much of it is pivoting
         | to China
         | 
         | Grumpy temper tantrums, on the part of the USA, will speed up
         | the pivot.
        
           | dismalaf wrote:
           | The only parts of the world pivoting towards China are the
           | despotic countries seeking to overthrow democracy and
           | liberalism.
        
             | AdeptusAquinas wrote:
             | No the US isn't pivoting towards China so far
        
               | dismalaf wrote:
               | Well the US is a liberal democracy...
        
               | FooBarWidget wrote:
               | The US is a plutocracy pretending to be a liberal
               | democracy.
        
               | lukev wrote:
               | Rapidly doing a lot less of the pretending.
        
               | Tostino wrote:
               | Sure we are buddy, keep telling yourself that.
        
             | alecco wrote:
             | What about Brazil.
        
               | dismalaf wrote:
               | They've got a socialist convict president...
               | 
               | https://thehill.com/opinion/international/4835388-brazil-
               | pre...
               | 
               | Lula is no friend of democracy.
        
               | Spivak wrote:
               | Not a single American can say shit given what's going on
               | at home.
        
               | dismalaf wrote:
               | What's happening at home? From a foreigner's (Canadian)
               | perspective, Trump's not doing anything crazy but the
               | media is going crazy...
               | 
               | Our Prime Minister has multiple scandals worse than
               | anything Trump has done so far... From groping a reporter
               | to multiple instances of blackface to firing the attorney
               | general because she investigated a company who gave
               | bribes to Moammar Gaddafi to nearly a billion dollars
               | going to a 2 person software consulting firm to fake
               | charities giving him and his family members millions of
               | dollars and much more. Yet somehow we're held up as an
               | example of democracy and the US, which defends democracy
               | all over the world somehow isn't...
        
               | kelnos wrote:
               | Not sure what you mean; the things you list that your PM
               | has done seems like just another day at the office for
               | Trump.
               | 
               | Trump has already done quite a few crazy things since his
               | inauguration. And I'm judging that by what I'm hearing
               | from people who are directly affected by his actions, not
               | by what I'm hearing from the media.
        
               | cdelsolar wrote:
               | it's not crazy to suspend all federal grants, get us out
               | of the Paris accord when the world is rapidly heating up,
               | try to put RFK Jr, a vaccine denier, in charge of our
               | health, rename the goddamn Gulf of Mexico,
               | unconstitutionally order that people born on American
               | soil are not necessarily Americans anymore, which has
               | been the case for at least 150 years, withdraw from the
               | WHO, launch hourly raids to deport people, many of whom
               | are actually American citizens?
        
               | dismalaf wrote:
               | > suspend all federal grants
               | 
               | Not really.
               | 
               | > get us out of the Paris accord when the world is
               | rapidly heating up
               | 
               | Is anyone on pace to meet those targets? Canada isn't.
               | 
               | > try to put RFK Jr, a vaccine denier, in charge of our
               | health,
               | 
               | Our health minister is a random politician who took poly
               | sci and history...
               | 
               | > rename the goddamn Gulf of Mexico
               | 
               | Things get renamed all the time... It's not crazy.
               | 
               | > people born on American soil are not necessarily
               | Americans anymore
               | 
               | Jus soli is the exception in the world and in history.
               | For example, every European country has jus sanguinis...
               | 
               | > launch hourly raids to deport people
               | 
               | Enforcing laws is the normal state of affairs in the
               | world...
               | 
               | > many of whom are actually American citizens?
               | 
               | I doubt this one, got any proof?
        
               | nateglims wrote:
               | Trump is doing things that violate some of the guardrails
               | of the US system. For example, refusing to disburse funds
               | allocated by the legislature. These separations are less
               | of a barrier in parliments. At least in the westminster
               | system. These boundaries have been redrawn before: the
               | federal bureaucracy was basically invented between the
               | 30s and the 60s and SCOTUS's role as we know it was
               | established in the early 1800s.
        
               | Vox_Leone wrote:
               | Brazil's technological environment is stagnant and
               | apathetic. And it has been that way since the tragic
               | explosion that claimed the country's space program. It
               | seems that all of the country's technological activities
               | have suffered the blow. It is possible that now it will
               | finally be able to manufacture its national tier 3 AI,
               | using the outputs of DeepSeek.
        
             | intalentive wrote:
             | "Democracy and liberalism" increasingly means canceling
             | elections (Romania), freezing bank accounts (Canada),
             | jailing people for speech (UK), killing thousands of
             | children (Israel), kidnapping CEOs (France), banning
             | foreign market competition (USA), blowing up pipelines,
             | seizing assets, erecting a vast surveillance state, etc,
             | etc.
             | 
             | "Democracy and liberalism" are merely shorthand for a
             | neofeudal rent-seeking oligarchy that disguises itself with
             | parliamentary forms. As Huxley predicted in 1958:
             | 
             | >the democracies will change their nature; the quaint old
             | forms -- elections, parliaments, Supreme Courts and all the
             | rest -- will remain. The underlying substance will be a new
             | kind of non-violent totalitarianism. All the traditional
             | names, all the hallowed slogans will remain exactly what
             | they were in the good old days. Democracy and freedom will
             | be the theme of every broadcast and editorial -- but
             | democracy and freedom in a strictly Pickwickian sense.
             | Meanwhile the ruling oligarchy and its highly trained elite
             | of soldiers, policemen, thought-manufacturers and mind-
             | manipulators will quietly run the show as they see fit.
             | 
             | Meanwhile our ruling class has completely exhausted its
             | moral authority and delegitimized its own political
             | formula: of popular sovereignty, rules-based order, human
             | rights, and so on. The mask is off and no one can take
             | these sacred oaths seriously any more.
             | 
             |  _THAT_ is why the world is pivoting to China. That, and
             | the fact that China does not impose as a precondition of
             | lending and trade that you upend your social /sexual norms
             | and/or demographically transform your country through mass
             | immigration. China simply offers a better deal.
        
               | AdeptusAquinas wrote:
               | Agree with most of this (not the weird anti-immigrant bit
               | but the rest), however China does have some foreign
               | requirements that are a bit of a pain in the ass, like
               | its insistence that Taiwan isn't a country. They also
               | don't like it and will retaliate when you point out the
               | shady shit it does (e.g. the Uighurs), but then thats no
               | different from the states especially under its current
               | toddler administration.
        
               | intalentive wrote:
               | The immigration bit helps explain why Hungary, for
               | instance, is trying to escape Western orbit and align
               | with the East.
        
               | intended wrote:
               | Is Hungary planning to exit the EU?
        
               | intalentive wrote:
               | I can't say what Orban is planning specifically, but his
               | recent actions have not been well received by the EU
               | Council, to say the least. Moreover Hungary has been
               | China's main outpost in Europe since 2015 when it joined
               | Belt and Road.
               | 
               | Considering fresh signs of rupture in transatlantic
               | relations, maybe Orban will turn out to have had keen
               | foresight. There seems to be some sort of realignment
               | afoot under the Trump administration.
               | 
               | https://en.wikipedia.org/wiki/2024_visits_by_Viktor_Orb%C
               | 3%A...
               | 
               | https://en.wikipedia.org/wiki/China%E2%80%93Hungary_relat
               | ion...
        
               | skinnymuch wrote:
               | Every country would behave like China in the same
               | situation with Taiwan. Imagine if the Confederates moved
               | over to Puerto Rico or Hawaii or Alaska. America damn
               | sure would say that's America still. They're literally
               | the same people from the same land. Same ethnicity. Same
               | history. Only being apart for under a century.
        
             | segasaturn wrote:
             | BYD cars are everywhere in Latin America and Europe. Xiaomi
             | phones also.
        
           | Austiiiiii wrote:
           | And we'd be shooting ourselves in the foot to do so. If
           | America is forced to use only the clunky corporate-owned
           | American AI at a fee, we'll very quickly fall behind
           | competitors worldwide who use DeepSeek models to produce
           | better results for much, much cheaper.
           | 
           | Not to mention it'd defeat the whole purpose of a "free
           | market" economy. (Not that that means much of anything
           | anymore)
        
             | kelnos wrote:
             | It never meant anything. There's no such thing as a free
             | market economy. We haven't had one of those in modern
             | times, and arguably human civilization has never had one.
             | Markets and their participants have chronically been
             | subject to information asymmetry, coercion/manipulation,
             | and regulation, among other things.
             | 
             | I don't think all of that is a bad thing (regulation tends
             | to make it harder to do the first two things), but "free
             | markets" are the economic equivalent to the "point mass" in
             | physics: perhaps useful sometimes to create simple models
             | and explanations of things, but will never exist in the
             | real world.
        
             | iforgot22 wrote:
             | The Nvidia export restrictions also might be shooting us in
             | the foot too, or at least Nvidia. They really benefit from
             | CUDA remaining the de facto standard.
        
               | segasaturn wrote:
               | The Nvidia export restrictions have already harmed
               | Nvidia. Deepseek-R1 is efficient on compute and was
               | trained on old, obsolete cards because they couldn't get
               | ahold of Nvidia's most cutting edge tech, so they were
               | forced to innovate instead of just brute-forcing
               | performance with faster cards. That has directly resulted
               | in Nvidia's stock crashing over the last week.
        
               | iforgot22 wrote:
               | That's true, but I mean it'd be a hundred times worse if
               | Deepseek did this on non-Nvidia cards. Which seems like
               | only a matter of time if we're going to keep doing this.
        
             | onlyrealcuzzo wrote:
             | DeepSeek's technology is out of the bag.
             | 
             | Every LLM provider in the US will be using it to lower
             | OpEx.
             | 
             | One of them is likely to pass those savings along to
             | consumers to gain market share.
             | 
             | Facebook is in the business of providing weights for free.
             | 
             | The idea that we are all doomed unless we immediately
             | migrate to DeepSeek is fantasy.
        
           | sangnoir wrote:
           | > Banned in the USA. Only.
           | 
           | The US government has the wherewithal to drag Europe along
           | with it, like they did with Huawei's 5G equipment.
        
             | roblabla wrote:
             | So far, a general tiktok ban (as opposed to a tiktok ban on
             | things like government phones) has only been in effect in
             | the USA. I highly doubt Europe would play ball at any
             | attempt at banning imports of DeepSeek.
             | 
             | Besides, it's kinda too late for this. The model is freely
             | accessible, so any attempt at banning it would be
             | _completely_ moot. If DeepSeek keeps releasing their future
             | models for free, I don't see how a ban could ever be
             | effective at all. Worse case scenario, big tech can't use
             | those models... but then individuals (and startups willing
             | to go fast and break laws) will be able to use them and
             | instantly get a leg up on the competition.
        
             | HotHotLava wrote:
             | Used to. If they're going to start a trade war, pull out of
             | NATO and invade Greenland instead, there'll not be much
             | soft power left to drag Europe anywhere.
        
             | buyucu wrote:
             | And yet, Huawei is still doing fine.
        
           | buyucu wrote:
           | I have a Xiaomi phone, a Huawei Matebook and a BYD car. I
           | guess I am pivoting to China after all :)
        
         | bangaladore wrote:
         | > So they'll go for getting DeepSeek banned like TikTok was now
         | that a precedent has been set ?
         | 
         | Can't really ban what can be downloaded for free and hosted by
         | anyone. There are many providers hosting the ~700B parameter
         | version that aren't CCP aligned.
        
           | runako wrote:
           | I'm old enough to remember when the US government did
           | something very similar. For years (decades?), we banned any
           | implementation of public-key cryptography under the guise of
           | the technology being akin to munitions.
           | 
           | People made shirts with printouts of the code to RSA under
           | the heading "this shirt is a munition." Apparently such
           | shirts are still for sale, even though they are not
           | classified as munitions anymore.
           | 
           | [1] - https://en.wikipedia.org/wiki/Export_of_cryptography_fr
           | om_th...
        
             | beepbooptheory wrote:
             | I am not that old, but I did a deep dive on this in the
             | past because it was just so extremely fascinating,
             | especially reading the archives of Cypherpunk. There is a
             | very solid, if rather bendy, line connecting all that to
             | "crypto culture" today.
        
         | buyucu wrote:
         | I'm willing to bet ''ban DeepSeek'' voices will start soon. Why
         | compete, when you can just ban?
        
           | cmiles74 wrote:
           | They've started already, I've seen posts on LinkedIn implying
           | or outright stating that DeepSeek is a national security risk
           | (IMHO, LinkedIn being the social media outlet most corporate-
           | sycophantic). I went ahead and just picked this one at random
           | from my feed.
           | 
           | https://www.linkedin.com/posts/kevinkeller_deepseek-
           | privacy-...
        
             | flybarrel wrote:
             | Oh this post...calling out DeepSeek's T&C but not comparing
             | it with OpenAI's is really disingenuous IMO.
        
             | ijidak wrote:
             | NBC Nightly News, on Monday, had an expert -- at 8:05 in
             | the video -- who claimed there might be national security
             | risks to Deepseek.
             | 
             | I'm not going to take a side on whether there is or not.
             | 
             | But, it does sound reminiscent of the reasons used to ban
             | Tik-tok.
             | 
             | https://youtu.be/uE6F6eTyAVc?si=BLZo3FMVRvjEy6Xa
        
           | Freedom2 wrote:
           | Competing is hard and expensive, whereas banning is for sure
           | the faster way to make stock values go up and exec's total
           | package as a result.
        
           | namuol wrote:
           | Already happening within tech company policy. Mostly as a
           | security concern. Local or controlled hosting of the model is
           | okay in theory based on this concern, but it taints
           | everything regarding deepseek in effect.
        
           | zelphirkalt wrote:
           | Actually asking for banning DeepSeek would be the ultimate
           | admit of defeat by ClosedAI.
        
         | zelphirkalt wrote:
         | Actually the "our IP" argument is ridiculous. What they are
         | doing is stealing data from all over the web, without people's
         | consent for that data to be used in training ML models. If
         | anything, then "Open"AI should be sued and forced to publish
         | their whole product. The people should demand knowing exactly
         | what is going on with their data.
         | 
         | Also still an unresolved issue is how they will ever comply
         | with a deletion request, should any model output personal data
         | of someone. They are heavily in a gray area, with regards to
         | what should be allowed. If anything, they should really shut up
         | now.
        
           | TZubiri wrote:
           | If there's any litigation, a counterclaim would be
           | interesting. But DeepSeek would need to partner with parties
           | that have been damaged by OpenAI's scraping.
        
         | boringg wrote:
         | Explain to me how one ban's opensource? That concept is foreign
         | to me.
        
       | staticelf wrote:
       | Not only do OpenAI and other steal data, they also spam the web
       | with requests and crawl websites over and over.
       | 
       | https://pod.geraspora.de/posts/17342163
        
       | mhitza wrote:
       | This is funny because its.
       | 
       | 1. Something I'd expect to happen.
       | 
       | 2. Lived through a similar scenario in 2010 or so.
       | 
       | Early in my professional career I've worked for a media company
       | that was scraping other sites (think Craigslist but for our local
       | market) to republish the content on our competing website. I
       | wasn't working on that specific project, but I did work on an
       | integration on my teams project where the scraping team could
       | post jobs on our platform directly. When others started scraping
       | "our content" there were a couple of urgent all hands on deck
       | meetings scheduled, with a high level of disbelief.
        
         | kigiri wrote:
         | Nice one, thank you for sharing !
        
         | spyckie2 wrote:
         | Classic.
        
       | ok123456 wrote:
       | OpenAI's models were trained on ebooks from a private ebook
       | torrent tracker leeched en-mass during a free leech event by
       | people who hated private torrent trackers and wanted to destroy
       | their "economy."
       | 
       | The books were all in epub format, converted, cleaned to plain
       | text, and hosted on a public data hoarder site.
        
       | paulhart wrote:
       | "You are trying to kidnap what I have rightfully stolen"
        
       | 65 wrote:
       | Let me guess, this gives the government and excuse to ban
       | DeepSeek. Which means tech companies get to keep their
       | monopolies, Sam Altman can grab more power, and the tech
       | overlords can continue to loot and plunder their customers and
       | the internet as a whole.
        
       | daft_pink wrote:
       | I mean if they paid to use the api and then used the output, I
       | fail to see how they can complain.
        
       | rachofsunshine wrote:
       | "It's obvious! You're trying to kidnap what I have rightfully
       | stolen!"
       | 
       | Yet another of a series of recent lessons in listening to people
       | - particularly powerful people focused on PR - when they claim a
       | neutral moral principle for what happens to be pragmatically
       | convenient for them. A principle applied only when convenient is
       | not a principle at all, it's just the skin of one stretched over
       | what would otherwise be naked greed.
        
       | supermatt wrote:
       | They refer to this in the paper as a part of the "cold start
       | data" which they use to fine-tune DeepSeek-V3 prior to training
       | R1.
       | 
       | They don't specifically name OpenAI, but they refer to "directly
       | prompting models to generate answers with reflection and
       | verification".
        
       | thorum wrote:
       | > "It is (relatively) easy to copy something that you know
       | works," Altman tweeted. "It is extremely hard to do something
       | new, risky, and difficult when you don't know if it will work."
       | 
       | The humor/hypocrisy of the situation aside, it does seem to be
       | true that OpenAI is consistently the one coming up with new ideas
       | first (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E,
       | ...) and then other companies reproduce their work, and get more
       | credit because they actually publish the research.
       | 
       | Unfortunately for them it's challenging to profit in the long
       | term from being first in this space and the time it takes for
       | each new idea to be reproduced is getting shorter.
        
         | spencerflem wrote:
         | Fortunately, OpenAI doesn't need to make money because they are
         | a nonprofit dedicated to the safe and transparent advancement
         | of AI for all of humanity
        
           | mjburgess wrote:
           | ...somewhere a yacht salesman cried out in terror
        
         | turtlesdown11 wrote:
         | > other companies reproduce their work, and get more credit
         | because they actually publish the research.
         | 
         | I don't understand, you mean OpenAI isn't releasing open models
         | and openly publishing their research?
        
           | Tostino wrote:
           | Are you being sarcastic (honestly, it's hard to tell after
           | reading as many uninformed takes in the past week as I have).
           | 
           | No, they aren't (other than whisper).
           | 
           | Their "papers" are closer to marketing materials. Very
           | intentionally leaving out tons of technical information.
        
             | KolmogorovComp wrote:
             | They are being sarcastic.
        
         | actuallyalys wrote:
         | There's some truth in that, but isn't making a radically
         | cheaper version also a new idea that deepseek didn't know
         | whether it would work? I mean, there was already research into
         | distillation, but there was already research into some of (most
         | of?) OpenAI's ideas.
        
         | weego wrote:
         | Boy who stole test papers complains about child copying his
         | answers.
        
           | FridgeSeal wrote:
           | No you don't understand, AI is "dangerous" and only him and
           | his uber rich billionaire mates should get to control it!
        
         | joe_the_user wrote:
         | _The humor /hypocrisy of the situation aside, it does seem to
         | be true that OpenAI is consistently the one coming up with new
         | ideas first (GPT 4, o1, 4o-style multimodality, voice chat,
         | DALL-E, ...) and then other companies reproduce their work, and
         | get more credit because they actually publish the research_
         | 
         | I claim one just can't put the humor/hypocrisy aside that
         | easily.
         | 
         | What OpenAI did with the release of ChatGPT is productize
         | research that was open and ongoing with Deepmind and other
         | leading at least as much. And everything after that was an
         | extension of the basic approach - improved, expanded but
         | ultimately the same sort of beast. One might even say the
         | situation of OpenAI to DeepMind was like Apple to Xerox.
         | Productizing is nothing to sneeze at - it requires creativity
         | and work to productize basic research. But naturally get end-
         | users who consider the productizers the "fountain heads", who
         | overestimate the productizers because products are all they
         | see.
        
           | Davidzheng wrote:
           | They RLHF'd first no?
        
         | mistercheph wrote:
         | Not really, they just put their eye to where everyone knows the
         | ball is going and publish fake / cherrypicked results and then
         | pretend like they got there first (o1, gpt voice, sora)
        
         | Hatchback7599 wrote:
         | Reminds me of the Bill Gates quote when Steve Jobs accused him
         | of stealing the ideas of Windows from Mac:
         | 
         | Well, Steve... I think it's more like we both had this rich
         | neighbor named Xerox and I broke into his house to steal the TV
         | set and found out that you had already stolen it.
         | 
         | Xerox could be seen as Google, whose researchers produced the
         | landmark Attention Is All You Need paper, and the general
         | public, who provided all of the training data to make these
         | models possible.
        
         | namuol wrote:
         | The eye-watering funding numbers proposed by Altman in the past
         | and more recently with "Stargate" suggests a publicly-funded
         | research pivot is not out of the question. Could see a big
         | defense department grant being given. Sigh.
        
         | rndphs wrote:
         | > OpenAI is consistently the one coming up with new ideas first
         | (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E, ...)
         | 
         | As far as I can tell o1 was based on Q-star, which could likely
         | be Quiet-STaR, a CoT RL technique developed at Stanford that
         | OpenAI may have learned about before it got published.
         | Presumably that's why they never used the Q-Star name even
         | though it had garnered mystique and would have been good for
         | building hype. This is just speculation, but since OpenAI
         | haven't published their technique then we can't know if it
         | really was their innovation.
        
       | nelblu wrote:
       | Hahaha I can't stop laughing... i dont know the validity of the
       | claim, but immediately i thought of the British Museum
       | complaining about theft.
        
       | myflash13 wrote:
       | What are the chances of old-school espionage? OpenAI should look
       | for a list of former employees who now live in China. Somebody
       | might've slipped out with a few hard drives.
        
       | andy_ppp wrote:
       | When I rewrite how the law works there should be a ludicrous
       | hypocrisy defence... if the person suing you has committed the
       | same offence the case should not be admissible.
        
       | crowcroft wrote:
       | The AI companies were happy to take whatever they want and put
       | the onus of proving they were breaking the law onto publishers by
       | challenging them to take things to court.
       | 
       | Don't get mad about possible data theft, prove it in court.
        
       | zoba wrote:
       | Does OpenAI's API attempt to detect this sort of thing? Could
       | they start outputting bad information if they suspect a
       | distillation attempt is underway?
        
       | aiono wrote:
       | How the turntables...
        
       | beardedwizard wrote:
       | Next they will try to force us to use our tax dollars to fund
       | their legal fights.
        
       | ginkgotree wrote:
       | I did not have in my cards: PRC open sourcing most powerful LLM
       | by stealing data set from "OpenAI" As someone that is very Pro-
       | America and Pro-Democracy, the iron here is just... so sweet.
        
       | gostsamo wrote:
       | How you dare take what I've rightfully stolen!
        
       | windex wrote:
       | SAltman, Salty.
        
       | me551ah wrote:
       | OpenAI is going after a company that open sourced their model, by
       | distilling from their non-open AI?
       | 
       | OpenAI talks a lot about the principles of being Open, while
       | still keeping their models closed and not fostering the open
       | source community or sharing their research. Now when a company
       | distills their models using perfectly allowed methods on the
       | public internet, OpenAI wants to shut them down too?
       | 
       | High time OpenAI changes their name to ClosedAI
        
         | alexathrowawa9 wrote:
         | The name OpenAI gets more ridiculous by the day
         | 
         | Would not be surprised if they do a rebrand eventually
        
       | pama wrote:
       | The R1 paper used o1-mini and o1-1217 in their comparisons, so I
       | imagine they needed to use lots of OpenAI compute in December and
       | January to evaluate their benchmarks in the same way as the rest
       | of their pipeline. They show that distilling to smaller models
       | works wonders, but you need the thought traces, which o1 does not
       | provide. My best guess is that these types of news are just
       | noise.
       | 
       | [edit: the above comment was based on sensetionalist reporting in
       | the original link and not the current FT article. I still think
       | there is a lot of noise in these news this last week, but it may
       | well be that openai has valid evidence of wrongdoing; I would
       | guess that any such wrongdoing would apply directly to V3 rather
       | than R1-zero, because o1 does not provide traces and generating
       | synthetic thinking data with 4o may be counterproductive.]
        
       | TheJCDenton wrote:
       | This Deep Whining(r) technique used by OpenAI is not very
       | effective.
        
       | insane_dreamer wrote:
       | Usually I'm very much on the side of protecting America's
       | interests from China, but in this case I'm so disgusted with
       | OpenAI and the rest of BigTech driving this "arms race" that I'd
       | be happy with them burning to the ground.
       | 
       | So we're going to reverse our goals to reduce emissions and
       | fossil fuels in order to hopefully save future generations from
       | the worst effects of climate change, in the name of being able to
       | do what, exactly, that is actually benefiting humanity? Boost
       | corporate profits by reducing labor?
        
         | insane_dreamer wrote:
         | downvoted -- I guess I upset some people defending OpenAI?
         | Good.
        
       | daft_pink wrote:
       | This reminds me of the railroads, where once railroads were
       | invented, there was a huge investment boom of eveyrone trying to
       | make money of the railroads, but the competition brought the
       | costs down where the railroads weren't the people who generally
       | made the money and got the benefit, but the consumers and regular
       | businesses did and competition caused many to fail.
       | 
       | AI is probably similar where the Moore's law and advancement will
       | eventually allow people to run open models locally and bring down
       | the cost of operation. Competiition will make it hard for all but
       | one or two players to survive and Nvidia, OpenAI, Deepseek, etc
       | most investments in AI by these large companies will fail to
       | generate substantial wealth but maybe earn some sort of return or
       | maybe not.
        
         | mjburgess wrote:
         | For the curious, it was vertical integration in the railroad-
         | oil/-coal industry which is where the money was made.
         | 
         | The problem for AI is the hardware is commodified and offers no
         | natural monopoly, so there isn't really anything obvious to
         | vertically integrate-towards-monopoly.
        
           | fullshark wrote:
           | Aren't we approaching a scenario where the software is
           | commodified (or at least "good enough" software) and the
           | hardware isn't (NVIDIA GPUs have defined advantages)
        
             | mjburgess wrote:
             | I think the lesson of DeepSeek is 'no' -- that by software
             | innovation (ie., dropping below CUDA to programming the GPU
             | directly, working at 8bit, etc.) you can trivialise the
             | hardware requirement.
             | 
             | However I think the reality is that there's only so much
             | coal to be mined, as far as LLM training goes. When we're
             | at "very dimishing returns" SoC/Apple/TSMC-CPU innovations
             | will deliver cheap inference. We only really need a M4
             | Ultra with 1TB RAM to hollow-out the hardware-inference-
             | supplier market.
             | 
             | Very easy to imagine a future where Apple releases a "Apple
             | Intelligence Mac Studio" with the specs for many businesses
             | to run arbitrary models.
        
               | daft_pink wrote:
               | I really hope that apple realizes soon there is a market
               | for Mac Pro/Mac Studio with a RAM in the TBs for AI
               | Workloads under $10k and a bunch of GPU cores.
        
               | jppope wrote:
               | there was a company that recently built a desktop GPU for
               | that exact thing. I'll see if I can find it
        
               | exe34 wrote:
               | https://cerebras.ai/ ?
        
             | duped wrote:
             | Compute is literally being sold as a commodity today,
             | software is not.
        
               | phkahler wrote:
               | >> Compute is literally being sold as a commodity today,
               | software is not.
               | 
               | The marginal cost of software is zero. You need some kind
               | of perceived advantage to get people to pay for it. This
               | isn't hard, as most people will pay a bit for big-name vs
               | "free". That could change as more open source apps become
               | popular by being awesome.
        
         | floatrock wrote:
         | The railroads drama ended when JP Morgan (the person, not yet
         | the entity) brought all the railroad bosses together, said "you
         | all answer to me because I represent your investors /
         | shareholders", and forced a wave of consolidation and
         | syndicates because competition was bad for business.
         | 
         | Then all the farmers in the midwest went broke not because they
         | couldn't get their goods to market, but because JP Morgan's
         | consolidated syndicates ate all their margin hauling their
         | goods to market.
         | 
         | Consolidation and monopoly over your competition is always the
         | end goal.
        
           | DrScientist wrote:
           | > Consolidation and monopoly over your competition is always
           | the end goal.
           | 
           | Surely that's only possible when you have a large barrier to
           | entry?
           | 
           | What's going to be that barrier in this case - cos it turns
           | out not to be neither training costs/hardware or secret
           | expertise.
        
             | yoyohello13 wrote:
             | The large syndicate will create the barriers. Either via
             | laws, or if that fails violence.
        
             | tdb7893 wrote:
             | So I'm not an expert in this but even with DeepSeek
             | supposedly reducing training costs isn't the estimate still
             | in the millions (and that's presumably not counting a lot
             | of costs)? And that wouldn't be counting a bunch of other
             | barriers for actually building the business since training
             | a model is only one part, the barrier to entry still seems
             | very high.
             | 
             | Also barriers to entry aren't the only way to get a
             | consolidated market anyway.
        
               | layer8 wrote:
               | About your first point, IMO the usefulness of AI will
               | remain relatively limited as long as we don't have
               | continuously learning AI. And once we have that, the
               | disparity between training and inference may effectively
               | disappear. Whether that means that such AI will become
               | more accessible/affordable or less is a different
               | question.
        
             | floatrock wrote:
             | You figure that out and the VC's will be shovelling money
             | into your face.
             | 
             | I suspect the "it ain't training costs/hardware" bit is a
             | bit exagerated since it ignores all the prior work that
             | DeepSeek was built on top of.
             | 
             | But, if all else fails, there's always the tried-and-true
             | approaches: regulatory capture, industry entrenchment, use
             | your VC bucks to be the last one who can wait out the costs
             | the incumbents _do_ face before they fold, etc.
        
               | jaredklewis wrote:
               | > I suspect the "it ain't training costs/hardware" bit is
               | a bit exagerated since it ignores all the prior work that
               | DeepSeek was built on top of.
               | 
               | How does it ignore it? The success of Deepseek proves
               | that training costs/hardware are definitely NOT a barrier
               | to entry that protects OpenAI from competition. If anyone
               | can train their model with ChatGPT for a fraction of the
               | cost it took to train ChatGPT and get similar results,
               | then how is that a barrier?
        
               | baq wrote:
               | Can _anyone_ do that though? You need the tokens and the
               | pipelines to feed them to the matmul mincers. Quoting
               | only dollar equivalent of GPU time is disingenuous at
               | best.
               | 
               | That's not to say they lie about everything, obviously
               | the thing works amazingly well. The cost is understated
               | by 10x or more, which is still not bad at all I guess?
               | But not mind blowing.
        
             | antisthenes wrote:
             | > Surely that's only possible when you have a large barrier
             | to entry?
             | 
             | As you grow bigger, you create barriers to entry where none
             | existed before, whether intentionally or unintentionally.
        
             | _DeadFred_ wrote:
             | Government regulation.
             | 
             | 'Can't have your data going to China'
             | 
             | 'Can't allow companies that do censorship aligned with
             | foreign nations'
             | 
             | 'This company violated our laws and used an American
             | company's tech for their training unfairly'
             | 
             | And the government choosing winners.
             | 
             | 'The government in announcing 500 billion going to these
             | chosen winners, anyone else take the hint, give up, you
             | won't get government contracts but will get pressure'.
             | 
             | Good thing nobody is making these sorts of arguments today.
        
               | astrange wrote:
               | The government isn't giving 500 billion to anyone. They
               | just let Trump announce a private deal he has no
               | involvement.
        
           | mrdevlar wrote:
           | Which is the exact goal of the current wave of Tech oligarchy
           | also.
        
           | jonstewart wrote:
           | I just read _The Great River_ by Boyce Upholt, a history of
           | the Mississippi river and human management thereof. It was
           | funny how the railroads were used as a bogeyman to justify
           | continued building of locks, dams, and other control
           | structures on the Mississippi and its tributaries, long after
           | shipping commodities down river had been supplanted by the
           | railroads.
        
           | boringg wrote:
           | This moment was also historically significant because it
           | demonstrated how financial power (Morgan) could control
           | industrial power (the railroads). A pattern that some say
           | became increasingly important in American capitalism.
        
         | UncleOxidant wrote:
         | > where the Moore's law and advancement will eventually allow
         | people to run open models locally
         | 
         | Probably won't be Moore's law (which is kind of slowing down)
         | so much as architectural improvements (both on the compute side
         | and the model side - you could say that R1 represents an
         | architectural improvement of efficiency on the model side).
        
         | rgbrgb wrote:
         | I think that's a very possible outcome. A lot of people
         | investing in AI are thinking there's a google moment coming
         | where one monopoly will reign supreme. Google has strong
         | network effects around user data AND economies of scale. Right
         | now, AI is 1-player with much weaker network effects. The user
         | data moat goes away once the model trains itself effectively
         | and the economies of scale advantage goes away with smart small
         | models that can be efficiently hosted by mortals/hobbyists. The
         | DeepSeek result points to both of those happening in the near
         | future. Interesting times.
        
         | lastofthemojito wrote:
         | I saw a thought-provoking post that similarly compared LLM
         | makers to the airlines: https://calpaterson.com/porter.html
        
         | taco_emoji wrote:
         | Main difference is that railroads are actually useful
        
         | yonran wrote:
         | I think a better analogy than railroads (which own the land
         | that the track sits on and often valuable land around the
         | station) is airlines, which don't own land. I recall a relevant
         | Warren Buffett letter that warned about investing hundreds of
         | millions of dollars into capital with no moat:
         | 
         | > Similarly, business growth, per se, tells us little about
         | value. It's true that growth often has a positive impact on
         | value, sometimes one of spectacular proportions. But such an
         | effect is far from certain. For example, investors have
         | regularly poured money into the domestic airline business to
         | finance profitless (or worse) growth. For these investors, it
         | would have been far better if Orville had failed to get off the
         | ground at Kitty Hawk: The more the industry has grown, the
         | worse the disaster for owners.
         | 
         | https://www.berkshirehathaway.com/letters/1992.html
        
       | tntxtnt wrote:
       | Can they tax DeepSeek just like they taxed BYD cars? Smh Chinese
       | ruin US industry again and again and again. Where's Trump at??
       | Why don't he taxed 1000000% of the free $0 DeepSeek AI??
        
       | glitchc wrote:
       | [flagged]
        
         | dang wrote:
         | Would you please not do this here? We're trying for an opposite
         | sort of conversation.
         | 
         | https://news.ycombinator.com/newsguidelines.html
        
           | glitchc wrote:
           | Sorry dang. I'll do better.
        
             | dang wrote:
             | Appreciated!
        
       | mk89 wrote:
       | What a joke OpenAI has become.
        
       | oxqbldpxo wrote:
       | Deepseek is really outstanding.
        
       | feverzsj wrote:
       | So, they bought a pro plus account, and gathered all the data
       | through it? Sounds just like Nvidia sells tons of embargoed AI
       | chips to China.
        
       | dlikren wrote:
       | Intriguing to see the difference of response from HN when OpenAI
       | first came to prominence and now.
        
       | pluc wrote:
       | OpenAI feeling threatened by open AI is just delicious
        
       | glenstein wrote:
       | All the top level comments are basking in the irony of it, which
       | is fair enough. But I think this changes the Deepseek narrative a
       | bit. If they just benefited from repurposing OpenAI data, that's
       | different than having achieved an engineering breakthrough, which
       | may suggest OpenAI's results were hard earned after all.
        
         | nprateem wrote:
         | Of course. How else would Americans justify their superiority
         | (and therefore valuations) if a load of _foreigners_ for Christ
         | 's sake could just out innovate them?
         | 
         | They _had_ to be cheating.
        
           | dang wrote:
           | Please don't take HN threads into nationalistic flamewar.
           | It's not what this site is for, and destroys what it is for.
           | 
           | https://news.ycombinator.com/newsguidelines.html
           | 
           | p.s. yes, that goes both ways - that is, if people are
           | slamming a different country from an opposite direction, we
           | say the same thing (provided we see the post in the first
           | place)
        
             | LPisGood wrote:
             | I see where you're coming from but that comment didn't
             | strike me as particularly inflammatory.
        
               | dang wrote:
               | I'm likely more sensitive to the fire potential on
               | account of being conditioned by the job.
               | 
               | Part of it is the form of the comment, btw - that one was
               | entirely a sequence of indignation tropes.
        
         | plantwallshoe wrote:
         | Yeah what happens when we remove all financial incentive to
         | fund groundbreaking science?
         | 
         | It's the same problem with pharmaceuticals and generics. It's
         | great when the price of drugs is low, but without perverse
         | financial incentives no company is going to burn billions of
         | dollars in a risky search for new medicines.
        
           | amarcheschi wrote:
           | In this case, these cures (llms) are medicines in search for
           | a disease to cure. I got Ai shoved everywhere, where I just
           | want it to aid in my coding. Literally, that's it. They're
           | also good at summarizing emails and similar things, but I
           | know nobody who does that. I wouldn't trust an Ai reading and
           | possibly hallucinate emails
        
           | jjcob wrote:
           | Then we just have to fund research by giving grants to
           | universities and research teams. Oh wait a sec: That's
           | already what pretty much every government in the world is
           | doing anyway!
        
         | tasuki wrote:
         | I understand they just used the API to talk to the OpenAI
         | models. That... seems pretty innocent? Probably they even paid
         | for it? OpenAI is selling API access, someone decided to buy
         | it. Good for OpenAI!
         | 
         | I understand ToS violations can lead to a ban. OpenAI is free
         | to ban DeepSeek from using their APIs.
        
           | Mengkudulangsat wrote:
           | That's how I understand it too.
           | 
           | If your own API can leak your secret sauce without any
           | malicious penetration, well, that's on you.
        
           | glenstein wrote:
           | Sure, but I'm not interested in innocence. They can be as
           | innocent or guilty as they want. But it means they didn't,
           | via engineering wherewithal, reproduce the OpenAI
           | capabilities from scratch. And originally that was supposed
           | to be one of the stunning and impressive (if true)
           | implications of the whole Deepseek news cycle.
        
             | freehorse wrote:
             | It is not as if they are not open about how they did it.
             | People are actually working on reproducing their results as
             | they describe in the papers. Somebody has already
             | reproduced the r1-zero rl training process on a smaller
             | model (linked in some comment here).
             | 
             | Even if o1 specifically was used (which is in itself
             | doubtful), it does not mean that this was the main reason
             | that r1 succeeded/it could not have happened without it.
             | The o1 outputs hides the CoT part, which is the most
             | important here. Also we are in 2025, scratch does not exist
             | anymore. Creating better technology building upon previous
             | (widely available) technology has never been a
             | controversial issue.
        
             | tasuki wrote:
             | Nothing is _ever_ done  "from scratch". To create a
             | sandwich, you first have to create the universe.
             | 
             | Yes, there is the question how much ChatGPT data DeepSeek
             | has ingested. Certainly not zero! But if DeepSeek has
             | achieved iterative self-improvement, that'd be huge too!
        
           | rubslopes wrote:
           | Additionally, I was under the impression that all those
           | Chinese models were being trained using data from OpenAI and
           | Anthropic. Were there not some reports that Qwen models
           | referred to themselves as Claude?
        
         | the_duke wrote:
         | These aren't mutually exclusive.
         | 
         | It's been known for a while that competitors used OpenAI to
         | improve their models, that's why they changed the TOS to forbid
         | it.
         | 
         | That doesn't mean the deep seek technical achievements are less
         | valid.
        
           | glenstein wrote:
           | >That doesn't mean the deep seek technical achievements are
           | less valid.
           | 
           | Well, that's literally exactly what it would mean. If
           | DeepSeek relied on OpenAI's API, their main achievement is in
           | efficiency and cost reduction as opposed to fundamental AI
           | breakthroughs.
        
             | obmelvin wrote:
             | Agreed. They accomplished a lot with distillation and
             | optimization - but there's little reason to believe you
             | don't also need foundational models to keep advancing.
             | Otherwise won't they run into issues training on more
             | synthetic data?
             | 
             | In a way this is something most companies have been doing
             | with their smaller models, DeepSeek just supposedly* did it
             | better.
        
         | epolanski wrote:
         | I really don't see a correlation here to be honest.
         | 
         | Eventually all future AIs will be produced with synthetic
         | input, the amount of (quality) data we humans can produce is
         | quite limited.
         | 
         | The fact that the input of one AI has been used in the training
         | of another one seems irrelevant.
        
           | glenstein wrote:
           | The issue isn't just that AI trained on AI is inevitable it's
           | _whose_ AI is being used as the base layer. Right now,
           | OpenAI's models are at the top of that hierarchy. If Deepseek
           | depended on them, it means OpenAI is still the upstream
           | bottleneck, not easily replaced.
           | 
           | The deeper question is whether Deepseek has achieved real
           | autonomy or if it's just a derivative work. If the latter,
           | then OpenAI still holds the keys to future advances. If
           | Deepseek truly found a way to be independent while achieving
           | similar performance, then OpenAI has a problem.
           | 
           | The details of how they trained matter more than the
           | inevitability of synthetic data down the line.
        
             | epolanski wrote:
             | > then OpenAI still holds the keys to future advances
             | 
             | Point is, those future advances are worthless. Eventually
             | anybody will be able to feed each other's data for the
             | training.
             | 
             | There's no moat here. LLMs are commodities.
        
               | glenstein wrote:
               | If LLMs were already pure commodities, OpenAI wouldn't be
               | able to charge a premium, and DeepSeek wouldn't have
               | needed to distill their model from OpenAI in the first
               | place. The fact that they did proves there's still a moat
               | --just maybe not as wide as OpenAI hoped.
        
         | JTyQZSnP3cQGa8B wrote:
         | > OpenAI's results were hard earned after all
         | 
         | DDOSing web sites and grabbing content without anyone's consent
         | is not hard earned at all. They did spent billions on their
         | thing, but nothing was earned as they could never do that
         | legally.
        
           | scotty79 wrote:
           | More like hard bought and hard stolen.
        
           | glenstein wrote:
           | I understand the temptation to go there, but I think it
           | misses the point. I have no qualms at all with the idea that
           | the sum total of intelligence distributed across the internet
           | was siphoned away from creators and piped through an engine
           | that now cynically seeks to replace them. Believe me, I will
           | grab my pitchfork and march side by side with you.
           | 
           | But let's keep the eye on the ball for a second. None of that
           | changes the fact that what _was_ built was a capability to
           | reflect that knowledge in dynamic and deep ways in
           | conversation, as well as image and audio recognition.
           | 
           | And did Deepseek also build that? From scratch? Because they
           | might not have.
        
       | this15testingg wrote:
       | if you want to completely disregard copyright laws, just call
       | your project AI!
       | 
       | I'm sure Aaron Swartz would be proud of where the "tech" industry
       | has gone. /s
       | 
       | what problem are these glorified AIM chatbots trying to solve?
       | wealth extraction not happening fast enough?
        
       | ra7 wrote:
       | "OpenAI has no moat" is probably running through their heads
       | right now. Their only real "moat" seems to be their ability to
       | fear monger with the US government.
        
       | geerlingguy wrote:
       | Something something "just desserts".
        
       | HarHarVeryFunny wrote:
       | DeepSeek-R1's multi-step bootstrapping process, starting with
       | their DeepSeek-V3 base model, would only seem to need a small
       | amount of reasoning data for the DeepSeek-R0 RL training, after
       | which that becomes the source for further data, along with some
       | other sources that they mention.
       | 
       | Of course it's possible that DeepSeek used O1 to generate some of
       | this initial bootstrapping data, but not obvious. O1 anyways
       | deliberately obfuscates it's reasoning process (see "Hiding the
       | chains of thought" section of OpenAI's "Learning to reason with
       | LLMs" page), such that what you see is an after-the-fact
       | "summary" of what it actually did; so, if DeepSeek did indeed use
       | some of O1's output to train on, it shows that the details of
       | O1's own reasoning process isn't as important as they thought it
       | was - it's just having some verified (i.e. leading to good
       | outcome) reasoning data from any source that matters to get
       | started.
        
       | curtisblaine wrote:
       | My feeling is that they will ban DS anyway because, like TikTok,
       | it can become a massive intelligence source for the CCP. Imagine
       | sending all your code to it, or your internal emails.
        
       | itamarcode wrote:
       | So protecting models behind API isn't working, ha?
        
       | ranger_danger wrote:
       | Reminds me of Steve Jobs complaining to Bill Gates about MS
       | "stealing" the GUI concept from them, which they in turn had
       | stolen from Xerox.
       | 
       | Obligatory "Everything is a Remix"
       | https://www.youtube.com/watch?v=X9RYuvPCQUA
        
       | pointedAt wrote:
       | it's a corona style portfolio gang.
        
       | coolhand2120 wrote:
       | I was using the Deepseek model last night and got this:
       | 
       | > You > do you understand copredication?
       | 
       | > Assistant > deepseek-coder-33b-instruct > I'm sorry for any
       | misunderstanding, but as an AI model developed by OpenAI, I don't
       | have personal ...
       | 
       | Yikes! That's pretty damning IMO.
        
       | jeanlucas wrote:
       | damn that's a good headline
        
       | deadbabe wrote:
       | In this whole AI saga, DeepSeek would be like Prometheus. They
       | stole the fire from the Gods and gave it to the humans, for free.
       | Logic dictates then that they will be forced to suffer brutal
       | punishment.
        
       | whenlambo wrote:
       | Qwen2.5-Plus
       | 
       | > have you been trained using openai?
       | 
       | No, I have not been trained using technology or data from OpenAI.
       | I am an artificial intelligence model developed by Yandex, called
       | YandexGPT. My "training" is based on proprietary algorithms and
       | data that was collected and processed by the Yandex team. While I
       | have similarities with other models such as those developed by
       | OpenAI (e.g. GPT-3), my training was conducted independently and
       | uses unique approaches specific to Yandex. This includes using
       | Russian-language data as well as other sources to provide a
       | better understanding of context and provide useful information to
       | users. If you have questions on topics related to AI technologies
       | or anything else, I'd be happy to help!
        
         | throwaway314155 wrote:
         | It doesn't actually know that.
        
       | concerndc1tizen wrote:
       | Is OpenAI claiming copyright ownership over the generated
       | synthetic data?
       | 
       | That would be a dangerous precedent to establish.
       | 
       | If it's a terms of service violation, I guess they're within
       | their rights to terminate service, but what other recourse do
       | they have?
       | 
       | Other than that, perhaps this is just rhetoric aimed at
       | introducing restrictions in the US, to prevent access to foreign
       | AI, to establish a national monopoly?
        
       | delusional wrote:
       | Boo hoo. Competition isn't fun when I'm not winning. Typical
       | Americans. When Americans are running around ruining the social
       | cohesion of several developing nations, that's just fair
       | competition, but as soon as they get even the smallest hint of
       | real competition they run to demonize it.
       | 
       | Yes deepseek is going to steal all of your data. OpenAI would so
       | the same. Yes the CCP is going to get access to your data and use
       | it to decide if you get to visit or whatever. The white house
       | does the same.
        
       | hsuduebc2 wrote:
       | A thief cries 'stop the thief!
        
       | hsuduebc2 wrote:
       | The pot calling the kettle black
        
       | wanderingmoose wrote:
       | There is a lot of discussion here about IP theft. Honest
       | question, from deepseek's point of view as a company under a
       | different set of laws than US/Western -- was there IP theft?
       | 
       | A company like OpenAI can put whatever licensing they want in
       | place. But that only matters if they can enforce it. The question
       | is, can they enforce it against deepseek? Did deepseek do
       | something illegal under the laws of their originating country?
       | 
       | I've had some limited exposure to media related licensing when
       | releasing content in China and what is allowed is very different
       | than what is permitted in the US.
       | 
       | The interesting part which points to innovation moving outside of
       | the US is US companies are beholden to strict IP laws while many
       | places in the world don't have such restrictions and will be able
       | to utilize more data more easily.
        
         | thiago_fm wrote:
         | The most interesting part is that China has been ahead of the
         | US in AI for many years, just not in LLMs.
         | 
         | You need to visit mainland China and see how AI applications
         | are everywhere, from transport to goods shipping.
         | 
         | I'm not surprised at all. I hope this in the end makes the US
         | kill its strict IP laws, which is the problem.
         | 
         | If the US doesn't, China will always have a huge edge on it, no
         | matter how much NVidia hardware the US has.
         | 
         | And you know what, Huawei is already making inference
         | hardware... it won't take them long to finally copy the TSMC
         | tech and flip the situation upside down.
         | 
         | When China can make the equivalent of H100s, it will be
         | hilarious because they will sell for $10 in Aliexpress :-)
        
           | twobitshifter wrote:
           | You don't even need to visit china, just read the latest
           | research papers and look at the authors. China has more
           | researchers in AI than the West and that's a proven way to
           | build an advantage.
        
             | nicce wrote:
             | It is also funny in a different way. Many people don't
             | realise that they live in some sort of bubble. Many people
             | in "The West" think that they are still the center of the
             | world in everything, while this might not be so correct
             | anymore.
             | 
             | In the U.S. there is 350 million people and EU has 520
             | million people (excluding Russia and Turkey).
             | 
             | China alone has 1.4 billion people.
             | 
             | Since there is a language barrier and China isolates
             | themselves pretty well from the internet, we forget that
             | there is a huge society with high focus on science. And
             | most of our tech products are coming from there.
        
           | nostradumbasp wrote:
           | Maybe not $10 unless they are loss-leading to dominance. Well
           | they actually could very well do exactly that... Hm, yea,
           | good points. I would expect at least an order or two of
           | magnitude higher to prevent an inferno.
           | 
           | Lets be fair though. Replicating TSMC isn't something that
           | could happen quickly. Then again, who knows how far along
           | they already are...
        
         | fulafel wrote:
         | What law would be broken here? Seems that copyright wouldn't
         | apply unless they somehow snatched the OpenAI models verbatim.
        
       | deeviant wrote:
       | Hmm, let's see--it looks like an easy legal defense.
       | 
       | DeepSeek could simply admit, "Yep, oops, we did it," but argue
       | that they only used the data to train Model X. So, if you want
       | compensation, you can have all the revenue from Model X (which,
       | conveniently, amounts to nothing).
       | 
       | Sure, they then used Model X to train Model Y, but would you
       | really argue that the original copyright holders are entitled to
       | all financial benefits derived from their work--especially when
       | that benefit comes in the form of a model trained on their data
       | without permission?
        
       | jchook wrote:
       | Friendly reminder that China publishes _twice_ as many AI papers
       | as the US[1], and _twice_ as many science and engineering papers
       | as the US.
       | 
       | China leads the world in the most cited papers[2]. The US's share
       | of the top 1% highly cited articles (HCA) has declined
       | significantly since 2016 (1.91 to 1.66%), and the same has
       | _doubled_ in China since 2011 (0.66 to 1.28%)[3].
       | 
       | China also leads the world in the number of generative AI
       | patents[4].
       | 
       | 1. https://www.bfna.org/digital-world/infographic-ai-
       | research-a...
       | 
       | 2. https://www.science.org/content/article/china-rises-first-
       | pl...
       | 
       | 3. https://ncses.nsf.gov/pubs/nsb202333/impact-of-published-
       | res...
       | 
       | 4. https://www.wipo.int/web-publications/patent-landscape-
       | repor...
        
       | liendolucas wrote:
       | Could this have been carefully orchestrated? Could DeepSeek have
       | devised this strategy a year ago and implemented knowing that
       | they would be able to benefit from OpenAI models and a possible
       | Nvidia market cap fall? Or is it just way too much to come up
       | with about such a move?
        
         | baal80spam wrote:
         | In theory, it could. This is a quant-fund after all, they know
         | stuff.
        
       | octacat wrote:
       | first time?
        
       | lxe wrote:
       | I mean, almost ALL opensource models, ever since alpaca, contain
       | a ton of synthetic data produced via ChatGPT in their finetuning
       | or training datasets. It's not a surprise to anyone who's been
       | using OSS LLMs for a while: almost ALL of them hallucinate that
       | they are ChatGPT.
        
       | waffletower wrote:
       | "Stole" - I don't believe that word means what he thinks it
       | means. Perhaps I pre-maturely anthropomorphize AI -- yet when I
       | read a novel, such as The Sorcerer's Stone, I am not guilty of
       | stealing Rowling's work, even if I didn't purchase the book but
       | instead found it and read it in a friend's bathroom. Now if I
       | were to take the specific plot and characters of that story and
       | write a screenplay or novel directly based on it, and,
       | explicitly, attempt to sell this work, perhaps the verb chosen
       | here would be appropriate.
        
       | Imnimo wrote:
       | I think there's two different things going on here:
       | 
       | "DeepSeek trained on our outputs and that's not fair because
       | those outputs are ours, and you shouldn't take other peoples'
       | data!" This is obviously extremely silly, because that's exactly
       | how OpenAI got all of its training data in the first place - by
       | scraping other peoples' data off the internet.
       | 
       | "DeepSeek trained on our outputs, and so their claims of
       | replicating o1-level performance from scratch are not really
       | true" This is at least plausibly a valid claim. The DeepSeek R1
       | paper shows that distillation is really powerful (e.g. they show
       | Llama models get a huge boost by finetuning on R1 outputs), and
       | if it were the case that DeepSeek were using a bunch of o1
       | outputs to train their model, that would legitimately cast doubt
       | on the narrative of training efficiency. But that's a separate
       | question from whether it's somehow unethical to use OpenAI's data
       | the same way OpenAI uses everyone else's data.
        
         | riantogo wrote:
         | Why would it cast any doubt? If you can use o1 output to build
         | a better R1. Then use R1 output to build a better X1... then a
         | better X2.. XN, that just shows a method to create better
         | systems for a fraction of the cost from where we stand. If it
         | was that obvious OpenAI should have themselves done. But the
         | disruptors did it. It hindsight it might sound obvious, but
         | that is true for all innovations. It is all good stuff.
        
           | rockemsockem wrote:
           | I think the prevailing narrative ATM is that DeepSeek's own
           | innovation was done in isolation and they surpassed OpenAI.
           | Even though in the paper they give a lot of credit to Llama
           | for their techniques. The idea that they used o1's outputs
           | for their distillation further shows that models like o1 are
           | necessary.
           | 
           | All of this should have been clear anyway from the start, but
           | that's the Internet for you.
        
             | aprilthird2021 wrote:
             | > the prevailing narrative ATM is that DeepSeek's own
             | innovation was done in isolation and they surpassed OpenAI
             | 
             | I did not think this, nor did I think this was what others
             | assumed. The narrative, I thought, was that there is little
             | point in paying OpenAI for LLM usage when a much cheaper,
             | similar / better version can be made and used for a
             | fraction of the cost (whether it's on the back of existing
             | LLM research doesn't factor in)
        
               | aiono wrote:
               | That's only the case if you don't need to use the output
               | of a much more expensive model.
        
               | TheGRS wrote:
               | Yes, well the narrative that rocked the stock market is
               | different. Its looking at what DeepSeek did and assuming
               | they may have competitive advantage in this space and
               | could outperform OpenAI at their own game.
               | 
               | If the narrative is actually that DeepSeek can only reach
               | whatever heights OpenAI has already gotten to with some
               | new tricks, then markets will probably refocus on
               | OpenAI's innovations and price things accordingly, even
               | if the initial cost is huge. It also means OpenAI
               | probably needs a better moat to protect its interests.
               | 
               | I'm not sure where the reality is exactly, but market
               | reactions so far have basically followed that initial
               | narrative and now the rebuttal.
        
               | kelnos wrote:
               | > _I did not think this, nor did I think this was what
               | others assumed._
               | 
               | That's what I thought and assumed. This is the narrative
               | that's been running through all the major news outlets.
               | 
               | It didn't even occur to me that DeepSeek could have been
               | training their models using the output of other models
               | until reading this article.
        
             | joe_the_user wrote:
             | _The idea that they used o1 's outputs for their
             | distillation further shows that models like o1 are
             | necessary._
             | 
             | Hmm, I think the narrative of the rise of LLMs is that once
             | the output of humans has been distilled by the model, the
             | human isn't necessary.
             | 
             | As far as I know, DeepSeek adds only a little to the
             | transformers model while o1/o3 added a special "reasoning
             | component" - if DeepSeek is as good as o1/o3, even taking
             | data from it, then it seems the reasoning component isn't
             | needed.
        
               | david-gpu wrote:
               | _> I think the narrative of the rise of LLMs is that once
               | the output of humans has been distilled by the model_
               | 
               | Distillation is a term of art in AI and it is
               | fundamentally incorrect to talk about distilling human-
               | created data. Only an AI model can be distilled.
               | 
               | https://en.m.wikipedia.org/wiki/Knowledge_distillation#Me
               | tho...
        
               | joe_the_user wrote:
               | Meh,
               | 
               | It seems clear that the term can be used informally to
               | denote the boiling down of human knowledge, indeed it was
               | used that way before AI appeared in the popular
               | imagination.
        
               | david-gpu wrote:
               | In the context in which you said it, it matters a lot.
               | 
               |  _> > The idea that they used o1's outputs for their
               | distillation further shows that models like o1 are
               | necessary._
               | 
               |  _> Hmm, I think the narrative of the rise of LLMs is
               | that once the output of humans has been distilled by the
               | model, the human isn 't necessary._
               | 
               | If deepseek was produced through the distillation (term
               | of art) of o1, then the cost of producing deepseek is
               | strictly higher than the cost of producing o1, and can't
               | be avoided.
               | 
               | Continuing this argument, if the premise is true then
               | deepseek can't be significantly improved without first
               | producing a very expensive hypothetical o1-next model
               | from which to distill better knowledge.
               | 
               | That is the argument that is being made. Please avoid
               | shallow dismissals.
               | 
               | Edit: just to be clear, I doubt that deepseek was
               | produced via distillation (term of art) of o1, since that
               | would require access to o1's weights. It may have used
               | some of o1's outputs to fine tune the model, which still
               | would mean that the cost of training deepseek is strictly
               | higher than training o1.
        
               | joe_the_user wrote:
               | _just to be clear, I doubt that deepseek was produced via
               | distillation_
               | 
               | Yeah, your technical point is kind of ridiculous here
               | that in all my uses of distillation (and in the comment I
               | quoted), distillation is used in informal sense and
               | there's no allegation that DeepSeek could have been in
               | possession of OpenAI's model weights, which is what's
               | needed for your "Distillation (term of Art)".
        
               | PontifexCipher wrote:
               | Some info that may be missing:
               | 
               | - v2/v3 (not r1) seem to be cloned from o1/4o output, and
               | perform worse (this cost the oft-repeated 5ish mm USD)
               | 
               | - r1 is specifically a reasoning step (using RL) _on top
               | of_ v2/v3 and performs similarly to o1 (the cost of this
               | is _not reported anywhere_)
               | 
               | - In the o1 blog post, they specifically say they use RL
               | to add reasoning to LLMs:
               | https://openai.com/index/learning-to-reason-with-llms/
        
               | sudosysgen wrote:
               | The R1-Zero paper shows how many training steps the RL
               | took, and it's not many. The cost of the RL is likely a
               | small fraction of the cost of the foundational model.
        
           | KingOfCoders wrote:
           | OpenAI couldn't do it, when the high cost of training and
           | access to GPUs is their competitive advance against startups,
           | they can't admit that it does not exist.
        
           | gmd63 wrote:
           | Why not just copy and paste the model and change the name?
           | That's an even more efficient form of distillation.
        
             | wgjordan wrote:
             | Even assuming the model was somehow publicly available in a
             | form that could be directly copied, that would be a more
             | blatant form of copyright infringement. Distillation
             | launders copyrighted material in a way that OpenAI
             | specifically has argued falls under fair use.
        
           | Imnimo wrote:
           | I think it would cast doubt on the narrative "you could have
           | trained o1 with much less compute, and r1 is proof of that",
           | if it turned out that in order to train r1 in the first
           | place, you had to have access to bunch of outputs from o1. In
           | other words, you had to do the really expensive o1 training
           | in the first place.
           | 
           | (with the caveat that all we have right now are accusations
           | that DeepSeek made use of OpenAI data - it might just as well
           | turn out that DeepSeek really did work independently, and you
           | really could have gotten o1-like performance with much less
           | compute)
        
             | SpaceManNabs wrote:
             | My question is if deepseek r1 is just a distilled o1, i
             | wonder if you can build a fine tuned r1 through
             | distillation without having to fine tune o1.
        
             | MrLeap wrote:
             | o1 wouldn't exist without the combined compute of every
             | mind that led to the training data they used in the first
             | place. How many h100 equivalents are the rolling continuum
             | of all of human history?
        
               | dchichkov wrote:
               | It should be possible to learn to reason from scratch.
               | And the ability to reason in a long context seems to be
               | very general.
        
               | Nevermark wrote:
               | How does one learn reasoning from scratch?
               | 
               | Human reasoning, as it exists today, is the result of
               | tens of thousands of years of intuition slowly distilled
               | down to efficient abstract concepts like "numbers",
               | "zero", "angles", "cause", "effect", "energy", "true",
               | "false", ...
               | 
               | I don't know what reasoning from scratch would look like
               | without training on examples from other reasoning beings.
               | As human children do.
        
               | Davidzheng wrote:
               | Actually i also think it's possible. Start with natural
               | numbers axiom system. Form all valid sentences of
               | increasing length. RL on a model to search for counter
               | example or proofs. This on sufficient computer should
               | produce superhuman math performance (efficiency) even at
               | compute parity
        
               | MrLeap wrote:
               | I wonder how much discovery in math happens as a result
               | in lateral thinking epiphanies. IE: A mathematician is
               | trying to solve a problem, their mind is open to
               | inspiration, and something in nature, or their childhood
               | or a book synthesizes with their mental model and gives
               | them the next node in their mental graph that leads to a
               | solution and advancement.
               | 
               | In an axiomatic system, those solutions are checkable,
               | but how discoverable are they when your search space
               | starts from infinity? How much do you lose by
               | disregarding the gritty _reality_ and foam of human
               | experience? It provides inspirational texture that helps
               | mathematicians in the search at least.
               | 
               | Reality is a massive corpus of cause and effect that can
               | be modeled mathematically. I think you're throwing the
               | baby out with the bathwater if you even want to be able
               | to math in a vacuum. Maybe there is a self optimization
               | spider that can crawl up the axioms and solve all of
               | math. I think you'll find that you can generate new math
               | infinitely, and reality grounds it and provides the
               | gravity to direct efforts towards things that are useful,
               | meaningful and interesting to us.
        
               | soulofmischief wrote:
               | As I mentioned in a sister comment, Godel's
               | incompleteness theorems also throw a wrench into things,
               | because you will be able to construct logically
               | consistent "truths" that may not actually exist in
               | reality. At which point, your model of reality becomes
               | decreasingly useful.
               | 
               | At the end of the day, all theory must be empirically
               | verified, and contextually useful reasoning simply cannot
               | develop in a vacuum.
        
               | iczero wrote:
               | I believe https://en.wikipedia.org/wiki/G%C3%B6del%27s_in
               | completeness_... (Godel's incompleteness theorems)
               | applies here
        
               | dchichkov wrote:
               | There are examples of learning reasoning from scratch
               | with reinforcement learning.
               | 
               | Emergent tool use from multi-agent interaction is a good
               | example - https://openai.com/index/emergent-tool-use/
        
               | MrLeap wrote:
               | Creating reasoning from scratch is the same task as
               | creating an apple pie from scratch.
               | 
               | First you must invent the universe.
        
               | soulofmischief wrote:
               | I've been giving this a lot of thought over the last few
               | months. My personal insight is that "reasoning" is simply
               | the application of a probabilistic reasoning manifold on
               | an input in order to transform it into constrained output
               | that serves the stability or evolution of a system.
               | 
               | This manifold is constructed via learning a
               | decontextualized pattern space on a given set of inputs.
               | Given the inherent probabilistic nature of sampling, true
               | reasoning is expressed in terms of probabilities, not
               | axioms. It may be possible to discover axioms by locating
               | fixed points or attractors on the manifold, but
               | ultimately you're looking at a probabilistic manifold
               | constructed from your input set.
               | 
               | But I don't think you can untie this "reasoning" from
               | your input data. It's possible you will find "meta-
               | reasoning", or similar structures found in any
               | sufficiently advanced reasoning manifold, but these
               | highly decontextualized structures might be entirely
               | useless without proper recontextualization, necessitating
               | that a reasoning manifold is trained on input whose
               | patterns follow learnable underlying rules, if the
               | manifold is to be useful for processing input of that
               | kind.
               | 
               | Decontextualization _is_ learning, decomposing aspects of
               | an input into context-agnostic relationships. But
               | recontextualization is the other half of that, knowing
               | how to take highly abstract, sometimes inexpressible,
               | context-agnostic relationships and transform them into
               | useful analysis in novel domains.
               | 
               | This doesn't mean a well-trained model can't reason about
               | input it hasn't encountered before, just that the input
               | needs to be in _some_ way causally connected to the same
               | laws which governed the input the manifold was trained
               | on.
               | 
               | I'm sure we could create a fully generalized reasoning
               | manifold which could handle _anything_ , but I don't see
               | how we possibly get that without first considering and
               | encountering all possible inputs. But these inputs still
               | have to have _some_ form of constraint governed by laws
               | that _must_ be learned through sampling, otherwise you 'd
               | just be training on effectively random data.
               | 
               | The other commenter who suggested simply generating all
               | possible sentences and training on internal consistency
               | should probably consider Godel's incompleteness theorems,
               | and that internal consistency isn't enough to accurately
               | model and interpret the universe. One could construct a
               | thought experiment about an isolated brain in a jar with
               | effectively unlimited neuronal connections, but no
               | sensory connection to the outside world. It's possible,
               | with enough connections, that the likelihood of the brain
               | conceiving of true events it hasn't actually encountered
               | does increase meaningfully. But the brain still has
               | nothing to validate against, and can't simply assume that
               | because something is internally logically consistent,
               | that it must exist or have existed.
        
               | miki123211 wrote:
               | It _is_ possible to learn to reason from scratch, that 's
               | what R1-0 did, but the resulting chains of thought aren't
               | legible to humans.
               | 
               | To quote DeepSeek directly:
               | 
               | > DeepSeek-R1-Zero, a model trained via large-scale
               | reinforcement learning (RL) without supervised fine-
               | tuning (SFT) as a preliminary step, demonstrated
               | remarkable performance on reasoning. With RL,
               | DeepSeek-R1-Zero naturally emerged with numerous powerful
               | and interesting reasoning behaviors. However,
               | DeepSeek-R1-Zero encounters challenges such as endless
               | repetition, poor readability, and language mixing. To
               | address these issues and further enhance reasoning
               | performance, we introduce DeepSeek-R1, which incorporates
               | cold-start data before RL.
        
               | dchichkov wrote:
               | If you look at the benchmarks of the DeepSeek-V3-Base, it
               | is quite capable, even in 0-shot:
               | https://huggingface.co/deepseek-ai/DeepSeek-V3-Base#base-
               | mod... This is not from scratch. These benchmark numbers
               | are an indication that the base model already had a large
               | number of reasoning/LLM tokens in the pre-training set.
               | 
               | On the other hand, my take on it, the ability to do
               | reasoning _in a long context_ is a general capability.
               | And my guess is that it can be bootstrapped from scratch,
               | without having to do training on all of the internet or
               | having to distill models trained on the internet.
        
             | cherry_tree wrote:
             | > I think it would cast doubt on the narrative "you could
             | have trained o1 with much less compute, and r1 is proof of
             | that"
             | 
             | Whether or not you could have, you can now.
        
             | deepGem wrote:
             | From the R1 paper
             | 
             | In this study, we demonstrate that reasoning capabilities
             | can be significantly improved through large-scale
             | reinforcement learning (RL), even without using supervised
             | fine-tuning (SFT) as a cold start. Furthermore, performance
             | can be further enhanced with the inclusion of a small
             | amount of cold-start data
             | 
             | Is this cold start data what OpenAI is claiming their
             | output ? If so what's the big deal ?
        
               | Imnimo wrote:
               | DeepSeek claims that the cold-start data is from
               | DeepSeekV3, which is the model that has the $5.5M
               | pricetag. If that data were actually the output of o1 (a
               | model that had a much higher training cost, and its own
               | RL post-training), that would significantly change the
               | narrative of R1's development, and what's possible to
               | build from scratch on a comparable training budget.
        
               | TheGeminon wrote:
               | In the paper DeepSeek just says they have ~800k responses
               | that they used for the cold start data on R1, and are
               | very vague about how they got it:
               | 
               | > To collect such data, we have explored several
               | approaches: using few-shot prompting with a long CoT as
               | an example, directly prompting models to generate
               | detailed answers with reflection and verification,
               | gathering DeepSeek-R1-Zero outputs in a readable format,
               | and refining the results through post-processing by human
               | annotators.
        
               | Imnimo wrote:
               | My surface-level reading of these two sections is that
               | the 800k samples come from R1-Zero (i.e. "the above RL
               | training") and V3:
               | 
               | >We curate reasoning prompts and generate reasoning
               | trajectories by performing rejection sampling from the
               | checkpoint from the above RL training. In the previous
               | stage, we only included data that could be evaluated
               | using rule-based rewards. However, in this stage, we
               | expand the dataset by incorporating additional data, some
               | of which use a generative reward model by feeding the
               | ground-truth and model predictions into DeepSeek-V3 for
               | judgment.
               | 
               | >For non-reasoning data, such as writing, factual QA,
               | self-cognition, and translation, we adopt the DeepSeek-V3
               | pipeline and reuse portions of the SFT dataset of
               | DeepSeek-V3. For certain non-reasoning tasks, we call
               | DeepSeek-V3 to generate a potential chain-of-thought
               | before answering the question by prompting.
               | 
               | The non-reasoning portion of the DeepSeek-V3 dataset is
               | described as:
               | 
               | >For non-reasoning data, such as creative writing, role-
               | play, and simple question answering, we utilize
               | DeepSeek-V2.5 to generate responses and enlist human
               | annotators to verify the accuracy and correctness of the
               | data.
               | 
               | I think if we were to take them at their word on all
               | this, it would imply there is no specific OpenAI data in
               | their pipeline (other than perhaps their pretraining
               | corpus containing some incidental ChatGPT outputs that
               | are posted on the web). I guess it's unclear where they
               | got the "reasoning prompts" and corresponding answers, so
               | you could sneak in some OpenAI data there?
        
               | joe_the_user wrote:
               | It's like the claim "they showed anyone create a powerful
               | from scratch" becomes "false yet true".
               | 
               | Maybe they needed OpenAI for their process. But now that
               | their model is open source, anyone can use that as their
               | cold start and spend the same amount.
               | 
               | "From scratch" is a moving target. No one who makes their
               | model with massive data from the net is really doing
               | anything from scratch.
        
               | bmicraft wrote:
               | Yeah, but that kills the implied hope of building a
               | better model for cheaper. Like this you'll always have a
               | ceiling of being a bit worse then the openai models.
        
             | vkou wrote:
             | If OpenAi had to account for the cost of producing all the
             | copyrighted material they trained their LLM on, their
             | system would be worth negative trillions of dollars.
             | 
             | Let's just assume that the cost of training can be
             | externalized to other people for free.
        
             | zombiwoof wrote:
             | Exactly. They piggybacked of lots of compute and used less.
             | There still is a total sum of a massive amount of compute
        
               | da_chicken wrote:
               | I mean, yes that's how progress works. Has OpenAI got a
               | patent? If not it's fair game.
               | 
               | We don't make people figure out how to domesticate a cow
               | every time they want a hamburger. Or test hundreds of
               | thousands of filaments before they can have a lightbulb.
               | Inventions, once invented, exist as giants to stand upon.
               | The inventor can either choose to disclose the invention
               | and earn a patent for exclusive rights, or they can try
               | to keep it a secret and hope nobody reverse engineers it.
        
             | hmottestad wrote:
             | At the pace that DeepSeek is developing we should expect
             | them to surpass OpenAI in not that long.
             | 
             | The big question really is, are we doing it wrong, could we
             | have created o1 for a fraction of the price. Will o4 cost
             | less to train than o1 did?
             | 
             | The second question is naturally. If we create a smarter
             | LLM, can we use it to create another LLM that is even
             | smarter?
             | 
             | It would have been fantastic if DeepSeek could have come
             | out with an o3 competitor before o3 even became publicly
             | available. That way we would have known for sure that we're
             | doing it wrong. Cause then either we could have used o1 to
             | train a better AI or we could have just trained in a
             | smarter and cheaper way.
        
               | pertymcpert wrote:
               | The whole discussion is about whether or not the second
               | case of using o1 outputs to fine tune R1 is what allowed
               | R1 to become so good. If that's the case then your
               | assertion that DeepSeek will surpass OpenAI doesn't
               | really make sense because they're dependent on a frontier
               | model in order to match, not surpass.
        
           | iforgot22 wrote:
           | "Then use R1 output to build a better X1" is the part I'm not
           | sure about. Is X1 going to actually be better than R1?
        
           | Sophira wrote:
           | Honestly, it's kind of silly that this technology is in the
           | hands of companies whose only aim is to make money, IMO.
        
             | goatlover wrote:
             | It's because they're the ones who could raise the money to
             | make those models. Academics don't have access to that kind
             | of compute. But the free models exist.
        
             | lenerdenator wrote:
             | Well, originally, OpenAI wasn't supposed to be that kind of
             | organization.
             | 
             | But if you leave someone in the tech industry of SV/SF long
             | enough, they'll start to get high on their own supply and
             | think they're entitled to insane amounts of value, so...
        
           | qwertox wrote:
           | They're standing on the shoulders of giants, not only in
           | terms of re-using expensive computing power almost for free
           | by using the outputs of expensive models. It's a bit of a
           | tradition in that country, also in manufacturing.
        
             | unreal37 wrote:
             | I thought OpenAI GPT took Wikipedia and the content of
             | every book as inputs to train their models?
             | 
             | Everyone is standing on the shoulders of giants.
        
           | dontreact wrote:
           | Is there any evidence R1 is better than O1?
           | 
           | It seems like if they in fact distilled then what we have
           | found is that you can create a worse copy of the model for
           | ~5m dollars in compute by training on its outputs.
        
           | ospray wrote:
           | They did do that themselves it's called o3.
        
           | dartos wrote:
           | What does "better" really even mean here?
           | 
           | Better benchmark scores can be cooked
        
         | 827a wrote:
         | There is a third possibility I haven't seen discussed yet: That
         | DeepSeek, illegally, got their hands on an OpenAI model via a
         | breach of OpenAI's systems. Its easy to laugh at OpenAI and say
         | "you reap what you sow", I'm 100% in that camp, but given the
         | lengths other Chinese entities have gone to when it comes to
         | replicating Western technology; we should not discount this.
         | 
         | That being said, breaching OAI's systems, re-training a better
         | model on top of their closed source model, then open sourcing
         | it: That's more Robinhood than Villain I'd say.
        
           | alecco wrote:
           | That would require stealing the model weights _and the code_
           | as OpenAI has been hiding what they are doing. Running models
           | properly is still quite artistic.
           | 
           | Meanwhile, they have access to Meta models and Qwen. And Meta
           | models are very easy to run and there's plenty of published
           | work on them. Occam's Razor.
        
             | ardit33 wrote:
             | How hard it is, if you have someone inside with the access
             | of the code? If you have 100s of people with full access,
             | not hard to have someone that is willing to sell it or do
             | some industrial espionage...
        
               | johnnyanmac wrote:
               | Lots of if's here. They need specific US employee
               | contacts at a company thars quickly growing and one of
               | those needs to be willing to breach their contracts to
               | share it. That contact also needs to trust that Deepseek
               | can properly utilize such code and completely undercut
               | their own work.
               | 
               | Lot of hoops when there's simply other models to utilize
               | publicly
        
               | foobarian wrote:
               | How big are the weights for the full model? If it's on
               | the scale of a large operating system image then it might
               | be easy to sneak, but if it's an entire data lake, not so
               | much.
        
               | dylan604 wrote:
               | devil's advocate says that we know that foreign (hell
               | even national) intelligence attempt to infiltrate agents
               | by having them become employees at any company they are
               | interested. So the idea isn't just pulled from thin air
               | as a concept. I do agree that it is a big if with no
               | corroborating evidence for the specific claim.
        
               | iforgot22 wrote:
               | I doubt that many people have full access to OpenAI's
               | code. Their team is pretty small.
        
           | seanhunter wrote:
           | The reason you're not seeing that being discussed is it's
           | totally unsupported by any evidence that's in the public
           | domain. Unless you have some actual evidence of such a
           | breach, you may as well introduce the possibility that
           | DeepSeek was reverse engineered from data found at an alien
           | crash site.
        
             | ryanisnan wrote:
             | [flagged]
        
               | JTyQZSnP3cQGa8B wrote:
               | > foreign nation-state backed organization
               | 
               | I'm European, are you talking about Microsoft, Google, or
               | OpenAI?
        
               | doctaj wrote:
               | They're referring to an organization (like a hacking
               | group) backed by a country (like china, North Korea).
        
               | freehorse wrote:
               | So, which of them 3?
        
               | dylan604 wrote:
               | You're missing the point that for a much larger portion
               | of the world, all "tech" is a foreign entity to them
        
               | mrguyorama wrote:
               | Until recently treating the US and China on the same
               | geopolitical level for allied countries would have been
               | insanely uncharitable and impossible to do honestly and
               | in good faith.
               | 
               | But now we have a bully in the whitehouse who seems to
               | want to literally steal neighboring land, or is throwing
               | shit everywhere to distract from the looting and
               | oligarchy being formed. So I suddenly have more empathy
               | for that position.
        
               | teractiveodular wrote:
               | DeepSeek is basically a startup, not a "foreign nation-
               | state backed organization". They were forced to pivot to
               | AI when their original business model (quant hedge fund)
               | was stomped on by the Chinese government.
               | 
               | Of course this is China so the government can and does
               | intervene at will, but alleging that this required CIA
               | level state espionage to pull off _is_ alien crash levels
               | of implausible. They open sourced the entire thing and
               | published incredibly detailed papers on how they did it!
        
               | byteknight wrote:
               | You may be unaware, but CCP has far more control over
               | private companies than you might think:
               | https://www.cna.org/our-media/indepth/2024/09/fused-
               | together...
               | 
               | This is not America. Your ideas do not apply the same
               | way.
        
               | baq wrote:
               | Naivety of some folks here is astounding... CCP has
               | golden shares in anything that could possibly be
               | important at some point in the next hundred years, and
               | yes golden shares are either really that or they're an
               | euphemism, the point is it doesn't even matter.
        
               | DiogenesKynikos wrote:
               | China has tens of millions of companies. The government
               | can't, doesn't and isn't even interested in micromanaging
               | all of them.
        
               | baq wrote:
               | It doesn't have to micromanage. It doesn't care about
               | most. It is only interested in the politically important
               | ones, but it needs the optionality if something becomes
               | worthwhile.
        
               | DiogenesKynikos wrote:
               | You're suggesting that DeepSeek was a Chinese government
               | operation that gained access to OpenAI's proprietary
               | data, and then you're justifying that by saying that the
               | government effectively controls every important company.
               | You're even chiding people who don't believe this as
               | naive.
               | 
               | I think you have a cartoonish view of China. A huge
               | amount goes on that the government has no idea about. Now
               | that DeepSeek has made a huge media splash, the Chinese
               | government will certainly pay attention to them, but then
               | again, so will the US government.
        
               | baq wrote:
               | I never suggested anything of the sort.
               | 
               | I'm suggesting it will be happening now and any past
               | efforts will be retroactively analyzed by the appropriate
               | CCP apparatus since _everyone_ is aware of the scale of
               | success as of Monday. It has become a political success,
               | thus it is imperative the CCP partakes in it.
        
               | DiogenesKynikos wrote:
               | This is the argument we're discussing:
               | 
               | > DeepSeek, illegally, got their hands on an OpenAI model
               | via a breach of OpenAI's systems. [...] given the lengths
               | other Chinese entities have gone to when it comes to
               | replicating Western technology; we should not discount
               | this.
               | 
               | Above, teractiveodular said that "DeepSeek is basically a
               | startup, not a 'foreign nation-state backed
               | organization'". You called teractiveodular naive for
               | saying that. So forgive me if I take the obvious
               | implication that you think DeepSeek is actually a state-
               | backed actor enabled by government hacking of OpenAI.
        
               | kridsdale1 wrote:
               | You don't need a CIA level agent to get someone with a
               | fraudulent job at OpenAI for a few months, load some
               | files on a thumb drive, and catch a plane to Shanghai.
        
               | seanhunter wrote:
               | I notice that your geographical perspective doesn't
               | stretch to any actual evidence that such a thing took
               | place. So it really has exactly the same amount of
               | supporting evidence as my alien crash reverse engineering
               | scenario at present.
        
               | ryanisnan wrote:
               | The surrounding facts matter a lot here. For example,
               | there are plenty of instances of governments hacking
               | companies of their competing nations. Motives are
               | incredibly easy to come by as well, be they political or
               | economical. We also have no proof that aliens exist at
               | all, so you've not only conjured them into existence, but
               | also their motive and their skills.
               | 
               | Are you trolling me?
        
               | seanhunter wrote:
               | Ok so to be clear: your surrounding facts are they may
               | have a motive and nation states hack people. I don't
               | disagree with those, but there really are no facts that
               | support the idea that there was a hack in this case and
               | the null hypothesis is that researchers all around the
               | world (not just in the US) are working on this so not all
               | breakthroughs are going to be made in the US. That could
               | change if facts come to light but att the moment it's not
               | really useful to speculate on something that is in
               | essence entirely made up.
               | 
               | No I'm not trolling you.
        
               | lcnPylGDnU4H9OF wrote:
               | > exactly the same amount of supporting evidence
               | 
               | The evidence supporting offensive hacking is abundant in
               | recent history; the number of things which have been
               | learned from alien crash data is surely smaller by
               | comparison to the number of things which have been
               | learned from offensive hacking.
        
               | MomsAVoxell wrote:
               | More to the point, offensive hacking is something that
               | all governments do, including the US, on a regular basis.
               | 
               | However, there is no evidence this is how the data was
               | obtained. Zero, zilch.
               | 
               | So its a useless statement which only plays on peoples
               | bias against their hated nation state de jour.
        
               | orochimaaru wrote:
               | Are you a Chinese military troll? The fact that China
               | engages in industrial espionage is well known. So I'm
               | surprised at your resistance to that possibility.
        
               | ceres wrote:
               | This thread reads like sour grapes to me. When people
               | can't compete but instead start throwing unfounded
               | allegations is not a good look.
               | 
               | Even OpenAI itself hasn't resorted to these wild
               | conspiracy theories.
               | 
               | Unless you're an insider in these companies, you're just
               | like the rest of us, you know nothing.
        
               | orochimaaru wrote:
               | Are you saying Chinese industrial espionage is not a well
               | established fact?
        
               | mrguyorama wrote:
               | Industrial espionage isn't magic. Airbus once stole
               | basically everything Boeing had, but that doesn't mean
               | Airbus could magically build a better 737 tomorrow.
               | 
               | China steals a lot of documentation from the US but in a
               | tech forum you of all people should be very familiar with
               | how little actual progress a bunch of documentation is
               | towards a finished unit.
               | 
               | The Comac C19 still uses American engines despite all the
               | industrial espionage in the world because most actual
               | engineering is still a brute force affair into finding
               | how things fail and fixing that. That's one of the main
               | advantages SpaceX has proven out with their "eh fuck it,
               | just launch and we will see what breaks" methodology.
               | 
               | Even fraud filled Chinese research makes genuine
               | advancements.
               | 
               | Believing that China, a wealthy nation of over a billion
               | people, with immense unity, nationality, and a regime
               | able to explicitly write blank checks could only possibly
               | beat the US at something by cheating is like, infinite
               | hubris. It's hilarious actually.
               | 
               | I don't know if DeepSeek is actually just a clone of
               | something or a shenanigan, that's possible and China
               | certainly has done those kinds of things before, but to
               | think it's the MOST LIKELY outcome, or to over rely on it
               | in any way is a death sentence. OpenAI claims to have
               | evidence, why do they not show it?
        
               | orochimaaru wrote:
               | >>>Believing that China, a wealthy nation of over a
               | billion people, with immense unity, nationality, and a
               | regime able to explicitly write blank checks could only
               | possibly beat the US at something by cheating is like,
               | infinite hubris. It's hilarious actually
               | 
               | So this is the first time I've heard the Chinese regime
               | being described in such flowery terms on HN - lol. But ok
               | - haha
        
             | htrp wrote:
             | Why stop there.... Deep seek is actually an alien
             | intelligence sent via sophons to destroy all of particle
             | physics!
        
               | nyclounge wrote:
               | Definitely would make a lot more sense, if the
               | leaderships are just secretly wallfacers.
        
             | svara wrote:
             | There's no public evidence to that effect but the
             | speculation makes a lot more sense than you make it sound.
             | 
             | The Chinese Communist party very much sees itself in a
             | global rivalry over "new productive forces". That's
             | official policy. And US leadership basically agrees.
             | 
             | The US is playing dirty by essentially embargoing China
             | over big AI - why wouldn't it occur to them to retaliate by
             | playing dirtier?
             | 
             | I mean we probably won't know for sure, but it's much less
             | far fetched than a lot of other speculation in this area.
             | 
             | E.g., R1's cold start training could probably have
             | benefited quite a bit from having access to OpenAI's chain
             | of thought data for training. The paper is a bit light on
             | detail on how it was made.
        
           | exe34 wrote:
           | I don't think you need to steal a model - you need training
           | samples generated from the original, which you can get simply
           | by buying access to perform API calls. This is similar to
           | TinyStories (https://arxiv.org/abs/2305.07759), except here
           | they're training something even better than the original
           | model for a fraction of the price.
        
           | nostradumbasp wrote:
           | I really doubt it. If that's the case the US GOV is in
           | serious shit. They have a contract with OpenAI to chuck all
           | their secret data in there... In all likelihood they just
           | distilled. It's a start up company that is publishing all of
           | their actual advances in the open, with proof. I think a lot
           | of people run to "espionage" super fast, when reality is, the
           | US probably sucks at what we call AI. Don't read that wrong,
           | they are a world leader obviously. However, there is a ton of
           | stuff they have yet to figure out.
           | 
           | Cheapening a series of fact checkable innovations because of
           | the country of origin when so far all that they have showed
           | are signs of good faith is paranoid at best and propaganda to
           | support the billionaire tech lords saving face for their own
           | arrogance at worst.
        
             | sanitycheck wrote:
             | _If_ the US government is  "chucking all their secret data"
             | into OpenAI servers/models, frankly they deserve everything
             | they get for that level of stupidity.
        
               | nostradumbasp wrote:
               | https://openai.com/global-affairs/introducing-chatgpt-
               | gov/
               | 
               | And don't forget the billions in partnerships...
        
               | ryandrake wrote:
               | ChatGPT, please complete a memo that starts with: "Our 5
               | year plan for military deployments in southeast Asia
               | are..."
        
               | nostradumbasp wrote:
               | Sounds hilarious but... https://techstory.in/trump-is-
               | accused-of-using-ai-to-compose...
        
               | compootr wrote:
               | Can't wait for gpt gov to hallucinate my PII!
        
               | nostradumbasp wrote:
               | Probably more like specialized tools to help spy on and
               | forecast civilian activities more than anything else.
               | Definitely with hallucinations, but that's not really
               | important. Facts don't matter much these days...
        
               | infamouscow wrote:
               | But remember: we cannot fire anyone over this because
               | then we're riding with Hitler /s
               | 
               | I can see why people refuse to pay taxes.
        
           | JTyQZSnP3cQGa8B wrote:
           | I discount this because OpenAI is pumping the whole internet
           | for money, and Zuckerberg torrented LibGen for its AI. We
           | cannot blame the Chinese anymore. They went through the
           | crappy "Made in China" phase in the 80s/90s, but they
           | mastered the art of improving stuff instead of mere cloning,
           | and it makes the big companies angry which is a nice bonus.
           | 
           | IMHO the whole world is becoming crazy for a lot of reasons,
           | and pissing off billionaires makes me laugh.
        
           | tehjoker wrote:
           | Basically, without some kind of shred of evidence, this is
           | completely chauvinist to make this accusation.
        
           | YetAnotherNick wrote:
           | Deepseek v2 and v2.5 was still very good but not par with
           | frontier models. How would you explain that?
        
           | WheatMillington wrote:
           | Do you have ANY reason to believe this might be true, or is
           | this 100% pure speculation based on absolutely nothing?
        
           | kmeisthax wrote:
           | I'd be perfectly fine with China stealing all "our" shit if
           | they just shared it.
           | 
           | The word "our" does a lot of heavy lifting in politics[0].
           | America is not a commune, it's a country club, one which we
           | used to own but have been bought out of, and whose new owners
           | view us as moochers but can't actually kick us out (yet). It
           | is in competition with another, worse country club that
           | purports to be a commune. We owe neither country club our
           | loyalty, so when one bloodies the other's nose, I smile.
           | 
           | [0] Some languages have a notion of an "exclusive we". If
           | English had such a concept, this would be an exclusive our.
        
             | kridsdale1 wrote:
             | This comment made me realize we don't have a pronoun for
             | n-our or x-nour
        
           | notatoad wrote:
           | Given the openness of their model, that should be pretty easy
           | to detect. If it were even a small possibility, wouldn't
           | openAI be talking about it very very loudly?
        
           | mvdtnz wrote:
           | We shouldn't discount a thing for which there is absolutely
           | zero evidence? Sorry that's not how it works.
        
           | sho_hn wrote:
           | Can you explain at a technical level how you view this as
           | necessary for the observed result?
        
         | tempeler wrote:
         | On another subject, if it belongs to OpenAI because it uses
         | OpenAI, then doesn't that mean that everything produced using
         | OpenAI belongs to OpenAI? Isn't that a reason not to use
         | OpenAI? It's very similar to saying that you used Google and
         | searched; now this product belongs to Google. They couldn't
         | figure out how to respond; they went crazy.
        
           | dathinab wrote:
           | The US ruled that AI produced things are by themself not
           | copyrightable.
           | 
           | So no, it doesn't belong to OpenAI.
           | 
           | You might be able to sue for penalties for breach of contract
           | of the TOS, but that doesn't give them the right to the
           | model. And even if it doesn't give them any right to
           | invalidate unbound copyright grants they have given to 3rd
           | parties (here literally everyone). Nor does it prevent anyone
           | from training their own new models based on it or prevent
           | anyone from using it. Oh, and the one breaching the TOS might
           | not even have been the company behind DeepSeek but some in-
           | between 3rd party.
           | 
           | Naturally this is under a few assumptions:
           | 
           | - the US consistently applies it's own law, but they have a
           | long history of not doing so
           | 
           | - the US doesn't abuse their power to force their economical
           | opinions (ban DeepSeek) on other countries
           | 
           | - it actually was trained on OpenAI, but uh, OpenAI has IMHO
           | shown over the years very clearly that they can't be trusted
           | and they are fully in-transparent. How do we trust their
           | claim? How do we trust them to not retrospectively have
           | tweaked their model to make it look as if DeepSeek copied it?
        
           | johndhi wrote:
           | to be clear, their terms of service are pretty clear that the
           | USER owns the outputs.
        
         | valine wrote:
         | The existence of R1-zero is evidence against any sort of theft
         | of OpenAI's internal COT data. The model sometimes outputs
         | illegible text that's useful only to R1. You can't do
         | distillation without a shared vocabulary. The only way R1 could
         | exist is if they trained it with RL.
        
         | FooBarWidget wrote:
         | There are literally public ChatGPT conversations data sets. For
         | the past 2 years it's been common practice for pretty much all
         | open source models to train on them. Ask just about any open
         | source model who they are and a lot of the time they'll say
         | they're ChatGPT. Why is "having obtained o1 generated data"
         | suddenly such a huge news, to the point of warranting
         | conspiracy theories about undisclosed/undiscovered breaches at
         | OpenAI? Nobody ever made a fuss about public ChatGPT data sets
         | until now. No hacking of OpenAI is needed to obtain ChatGPT
         | data.
        
         | me551ah wrote:
         | This is going to have a catastrophic effect on closed source AI
         | startup valuations. Because this means that anyone can copy any
         | LLM. The person who trains the model, spends the most amount of
         | money. Everyone else can create a replica at lower cost
        
           | iforgot22 wrote:
           | Maybe anyone can copy any LLM with sufficient querying. There
           | are still ways to guard one.
        
           | amlib wrote:
           | Why is that bad? If a powerful entity can scrape every piece
           | of media humanity has to offer and ignore copyright then why
           | should society let then profit unrestricted from it? It's
           | only fair that such models have no legal protection around
           | their usage and can be used and analyzed by anyone as they
           | see fit. The only reason this hasn't been codified into laws
           | is because those same powerful entities have been busy trying
           | to do regulatory capture.
        
         | km144 wrote:
         | Reasonable take, but to ignore the politics of this whole thing
         | is to miss the forest for the trees--there is a big tech
         | oligarchy brewing at the edges of the current US administration
         | that Altman is already participating in with Stargate, and
         | anti-China sentiment is everywhere. They'd probably like the US
         | to ban Chinese AI.
        
         | ComputerGuru wrote:
         | The suggestion that _any_ large-scale AI model research today
         | isn't ingesting output of its predecessors is laughable.
         | 
         |  _Even if_ they didn't directly, intentionally use o1 output
         | (and they didn't claim they didn't, so far as I know), AI slop
         | is everywhere. We passed peak original content years ago.
         | Everything is tainted and everything should be understand in
         | that context.
        
           | brianstrimp wrote:
           | > We passed peak original content years ago.
           | 
           | In relative terms, that's obviously and most definitely true.
           | 
           | In absolute terms, that's obviously and most definitely
           | false.
        
         | s17n wrote:
         | > This is obviously extremely silly, because that's exactly how
         | OpenAI got all of its training data in the first place - by
         | scraping other peoples' data off the internet.
         | 
         | OpenAI has also invested heavily in human annotation and RLHF.
         | If all DeepSeek wanted was a proxy for scraped training data,
         | they'd probably just scrape it themselves. Using existing
         | RLHF'd models as replacement for expensive humans in the
         | training loop is the real game changer for anyone trying to
         | replicate these results.
        
           | KennyBlanken wrote:
           | "We spent a lot of labor processing everything we stole"
           | is...not how that works.
           | 
           | That's like the mafia complaining that they worked so hard to
           | steal those barrels of beer that someone made off with in the
           | middle of the night and really that's not fair and won't
           | someone do something about it?
        
             | s17n wrote:
             | Oh, I don't really care about IP theft and agree that it's
             | funny that openai is complaining. But I don't think its
             | true that deepseek is just doing this because they are too
             | lazy to scrape the internet themselves - its all about the
             | human labor that they would otherwise have to pay for.
        
         | reissbaker wrote:
         | You're right that the first claim is silly, but the second
         | claim is pretty silly too -- they're not claiming industrial
         | espionage, they're claiming a breach in ToS. The outputs of the
         | o1 thinking process aren't user-visible, and never leave
         | OpenAI's datacenters. Unless DeepSeek actually had a mole that
         | stole their o1 outputs, there's nothing useful DeepSeek
         | could've distilled to get to R1's thought processes.
         | 
         | And if DeepSeek had a mole, _why would they bother running a
         | massive job internally to steal the data generated_? It would
         | be way easier for the mole to just leak the RL training
         | process, and DeepSeek could quietly copy it rather than
         | bothering with exfiltrating massive datasets to distill. The
         | training process is most likely like, on the order of a hundred
         | lines of Python or so, and you don 't even need the file: you
         | just need someone to describe it to you. Much simpler than
         | snatching hundreds of gigabytes of training data off of
         | internal servers...
         | 
         | Plus, the RL process described in DeepSeek's paper has already
         | been replicated by a PhD student at Berkeley:
         | https://x.com/karpathy/status/1884678601704169965 So, it seems
         | pretty unlikely they simply distilled R1 and lied about it, or
         | else how does their RL training algo actually... work?
         | 
         | This is mainly cope from OpenAI that their supposedly super
         | duper advanced models got caught by China within a few months
         | of release, for way cheaper than it cost OpenAI to train.
        
         | me_me_me wrote:
         | Early on people used prompts to escape sandbox and the AI
         | thought it was openAI.
         | 
         | In the community there is a consensus DeepSeekV3 was stolen.
         | 
         | Given it comes from China is no surprise
         | 
         | That said I am glad DeepSeek exists
        
           | dns_snek wrote:
           | > Given it comes from China is no surprise
           | 
           | Is this double standard necessary? OpenAI, a US company,
           | stole the entire internet for training and the techbro
           | consensus has been that this is fine.
        
             | nomel wrote:
             | OpenAI took scattered data and organized it into something
             | useful, at massive cost and effort. Deepseek (apparently)
             | skipped some of that first step. "Right" or "wrong" aside,
             | the GPU count and budget numbers that we've seen are
             | relatively meaningless if they piggy backed so heavily.
             | 
             | I think the biggest takeaway is, if it is "stolen" the only
             | thing interesting about R1 is the _relative_ difference,
             | and the budget /GPU count to get _that relative_
             | difference. It 's closer to resources/(deepseek - openai)
             | rather than resources/deepseek, which is how it's been
             | perceived.
             | 
             | Maybe it's still great, even with that perspective!
        
             | brianstrimp wrote:
             | I work in a field with lots of cheap microchips. I can tell
             | you that the amount of counterfeit copies flooding in from
             | China as well as the speed in which they are copying is
             | truly breathtaking.
        
             | me_me_me wrote:
             | What double standard? I have not claimed anything about
             | openAI being ethical paragon of virtue.
             | 
             | I anything, you can be accused of using whataboutism in
             | order to justify DeepSeek illegal actions
        
         | HarHarVeryFunny wrote:
         | DeepSeek-R0 (based on DeepSeek-V3 base model) was _only_
         | trained with RL, no SFT, so this isn 't at all like the
         | "distillation" (i.e SFT on synthetic data generated by R1) that
         | they also demonstrated by fine tuning Qwen and LLaMa.
         | 
         | Now, DeepSeek may (or may not) have used some O1 generated data
         | for the R0 RL training, but if so that's just a cost saving vs
         | having to source some reasoning data some other way, and in no
         | way reduces the legitimacy of what they accomplished (which is
         | not something any of the AI CEOs are saying).
        
         | znpy wrote:
         | This really got me thinking that open ai should have no ip
         | claim at all, since all their outputs and stuff are basically a
         | ripoff of the entire human knowledge and IPs of various kinds.
        
           | onlyrealcuzzo wrote:
           | The law and common sense often are at odds.
        
         | nullc wrote:
         | There is a big difference between being able to train on the
         | reasoning vs just the answers, which they can't against o1
         | because it's hidden. There is also a huge difference between
         | being able to train on the probabilities (distillation) vs not,
         | which again they can and did do with the llama models and can't
         | directly with OpenAI because the conceal the probability
         | output.
        
         | nonrandomstring wrote:
         | I think the more interesting claim (that Deepseek should make
         | for lols) is that it wasn't _them_ who trained R1. No, it was
         | O1 's idea. It chose to take the young R1 as its padawan.
        
         | miki123211 wrote:
         | > This is obviously extremely silly, because that's exactly how
         | OpenAI got all of its training data
         | 
         | IANAL, but It is worth noting here that DeepSeek _has
         | explicitly consented to a license that doesn 't allow them to
         | do this_. That is a condition of using the Chat GPT and the
         | OpenAI API.
         | 
         | Even if the courts affirm that there's a fair use defence for
         | AI training, DeepSeek may still be in the wrong here, not
         | because of copyright infringement, but because of a breach of
         | contract.
         | 
         | I don't think OpenAI would have much of a problem if you train
         | your model on data scraped from the internet, some of which
         | incidentally ends up being generated by Chat GPT.
         | 
         | Compare this to training AI models on Kindle Books randomly
         | scraped off the internet, versus making a Kindle account,
         | agreeing to the Kindle ToS, buying some books, breaking
         | Amazon's DRM and then training your AI on that. What DeepSeek
         | did is more analogous to the latter than the former.
        
           | freen wrote:
           | Did OpenAI abide by my service's terms of service when it
           | ingested my data?
        
             | cortesoft wrote:
             | Did OpenAI have to sign up for your service to gain access?
        
               | thorncorona wrote:
               | Can you steal someone else's laptop if they stood up to
               | get a drink?
        
               | gizajob wrote:
               | If their OS is open to the internet and you can scrape it
               | and copy it off while they're gone, then that would be
               | about the right analogy. And OpenAi and DeepSeek have
               | done the same thing in that case.
        
               | lolinder wrote:
               | It probably ignored hundreds of thousands of "by using
               | this site you consent to our Terms and Conditions"
               | notices, many of which probably would be read as
               | prohibiting training. But that's also a great example of
               | why these implicit contracts don't really work as
               | contracts.
        
               | freen wrote:
               | Civil law is only available to deep pockets.
               | 
               | Contracts are enforceable to the degree to which you can
               | pay lawyers to enforce them.
               | 
               | I will run out of money trying to enforce my terms of
               | service against openAI, while they have a massive war
               | chest to enforce theirs.
               | 
               | Ain't libertarianism great?
        
               | otherme123 wrote:
               | OpenAI scrapped my blog so aggressively that I had to ban
               | their IPs. They ignored the robots.txt (which is kind of
               | ToS) by 2 orders of magnitude, they ignored the explicit
               | ToS that I copypasted blindly from somewhere but turns
               | out it forbids what they did (something like you can't
               | make money with the content). Not that I'm going to
               | enforce it, but they should at least shut up.
        
               | outside1234 wrote:
               | That isn't required to be in violation of copyright
        
               | freen wrote:
               | Actually, yes, the actively agreed to them. Clicked the
               | button and everything.
        
               | bayindirh wrote:
               | No, but some of the data is licensed.
               | 
               | For example, my digital garden is under GFDL, and my blog
               | is CC BY-NC-SA. IOW, They can't remix my digital garden
               | with any other license than GFDL, and they have to credit
               | me if they remix my blog, and can't use it for any
               | commercial endeavor, which OpenAI certainly does now.
               | 
               | So, by scraping my webpages, they agree to my licensing
               | of my data. So they're de-facto breaching my licenses,
               | but they cry "fair-use".
               | 
               | If I tell that they're breaching the license terms,
               | they'd laugh at me, and maybe give me 2 cents of API
               | access to mock me further. When somebody _allegedly_ uses
               | their API with their unenforcable ToS, they scream like
               | an agitated cuckatoo (which is an insult to the cuckatoo,
               | BTW. They 're devilishly intelligent birds).
               | 
               | Drinking their own poison was mildly painful, I guess...
               | 
               | BTW, I don't believe that Deepseek has copied/used OpenAI
               | models' outputs or training data to train theirs, even if
               | they did, "the cat is out of the bag", "they did
               | something amazing so they needed no permissions", "they
               | moved fast and broke things", and "all is fair-use
               | because it's just research".
               | 
               |  _Heh._
        
           | krust wrote:
           | >IANAL, but It is worth noting here that DeepSeek has
           | explicitly consented to a license that doesn't allow them to
           | do this. That is a condition of using the Chat GPT and the
           | OpenAI API.
           | 
           | I have some news for you
        
           | dartos wrote:
           | TOS are not contracts.
        
             | lolinder wrote:
             | Citation? My understanding was that they are provided that
             | someone has to affirmatively accept them in order to use
             | your site. So Terms of Service stuck at the bottom in the
             | footer likely would not count as a contract because there's
             | no consent, but Terms of Service included in a check box on
             | a login form likely would count.
             | 
             | But IANAL, so if you have a citation that says otherwise
             | I'd be happy to see it!
        
           | anon373839 wrote:
           | > DeepSeek has explicitly consented to a license that doesn't
           | allow them to do this.
           | 
           | You actually don't know this. Even if it were true that they
           | used OpenAI outputs (and I'm very doubtful) it's not
           | necessary to sign an agreement with OpenAI to get API
           | outputs. You simply acquire them from an intermediary, so
           | that you have no contractual relationship with OpenAI to
           | begin with.
        
           | dmitrygr wrote:
           | > DeepSeek has explicitly consented to a license that doesn't
           | allow them to do this.
           | 
           | By existing in USA, OpenAI consented to comply with copyright
           | law, and how did that go?
        
         | alach11 wrote:
         | If we assume distillation remains viable, the game theory
         | implications are huge.
         | 
         | It's going to shift the market of how foundation models are
         | used. Companies creating models will be incentivized to
         | vertically integrate, owning the full stack of model usage.
         | Exposing powerful models via APIs just lets a competitor clone
         | your work. In a way OpenAI's Operator is a hint of what's to
         | come
        
       | JBSay wrote:
       | When China is more open than you, you've got a problem
        
       | cbracketdash wrote:
       | Let's also not forget Suchir Balaji, who was mysteriously killed
       | when exposing OpenAI's violation of copyright law.
        
       | spacecadet wrote:
       | See you all on lobsters...
       | 
       | So long HN and thanks for all the fish?
        
       | thumbsup-_- wrote:
       | is stealing from the thief actually a theft?
        
       | stevenally wrote:
       | They should be happy. Now that can provide that _amazing_ AI much
       | more cheaply. They don 't need half a trillion dollars worth of
       | Nvidia chips.
        
       | wnevets wrote:
       | Its like a bank robber being upset when someone steals their loot
        
       | game_the0ry wrote:
       | At least DeepSeek open sourced their code. They're more open than
       | OpenAI.
       | 
       | Ironic.
        
       | nshung wrote:
       | Hilarious. Scam Altman is giving me SBF vibe daily now.
        
       | flybarrel wrote:
       | OpenAI shocked that an AI company would train on someone else's
       | data without permission or compensation...lolllllll
        
       | the_optimist wrote:
       | This whole topic is basura enfuego. Same pack of maroons
       | careening around society for years clamoring for censorship now
       | imagining that Aaron Schwartz is their hero and that they want to
       | menace people. Kids, don't be like the grasping fools in these
       | threads, philosophically unfounded and desperately glancing
       | sideways, hoping the cumulative feels and gossip will sum to life
       | meaning.
        
       | blast wrote:
       | Everyone is responding to the intellectual property issue, but
       | isn't that the less interesting point?
       | 
       | If Deepseek trained off OpenAI, then it wasn't trained from
       | scratch for "pennies on the dollar" and isn't the Sputnik-like
       | technical breakthrough that we've been hearing so much about.
       | That's the news here. Or rather, the potential news, since we
       | don't know if it's true yet.
        
         | jondwillis wrote:
         | But it does mean moat is even less defensible for companies
         | whose fortunes are tied to their foundation models having some
         | performance edge, and a shift in the kinds of hardware used for
         | inference (smaller, closer to the edge.)
        
         | tensor wrote:
         | That's not correct. First of all, training off of data
         | generated by another AI is generally a bad idea because you'll
         | end up with a strictly less accurate model (usually). But
         | secondly, and more to your point, even if you were to use
         | training data from another model, YOU STILL NEED TO DO ALL THE
         | TRAINING.
         | 
         | Using data from another model won't save you any training time.
        
           | fumeux_fume wrote:
           | I think the point is that if R1 isn't possible without access
           | to OpenAI (at low, subsidized costs) then this isn't really a
           | breakthrough as much as a hack to clone an existing model.
        
             | tensor wrote:
             | The training techniques are a breakthrough no matter what
             | data is used. It's not up for debate, it's an empirical
             | question with a concrete answer. They can and did train
             | orders of magnitude faster.
        
               | blast wrote:
               | Not arguing with your point about training efficiency,
               | but the degree to which R1 is a technical breakthrough
               | changes if they were calling an outside API to get the
               | answers, no?
               | 
               | It seems like the difference between someone doing a
               | better writeup of (say) Wiles's proof vs. proving
               | Fermat's Last Theorem independently.
        
               | pests wrote:
               | That outside API used to be humans, doing the work
               | manually. Now we have ways to speed that up.
        
             | bbor wrote:
             | R1 is--as far as we know from good ol' ClosedAI--far more
             | efficient. Even if it were a "clone", A) that would be a
             | terribly impressive achievement on its own that Anthropic
             | and Google would be mighty jealous of, and B) it's at the
             | very least a distillation of O1's reasoning capabilities
             | into a more svelte form.
        
           | bbor wrote:
           | I think you're missing the point being made here, IMHO: using
           | an advanced model to build _high quality_ training data
           | (whatever that means for a given training paradigm)
           | absolutely would increase the efficiency of the process.
           | Remember that they 're not fighting over sounding human,
           | they're fighting over deliberative reasoning capabilities,
           | something that's relatively rare in online discourse.
           | 
           | Re: "generally a bad idea", I'd just highlight "generally" ;)
           | Clearly it worked in this case!
        
             | tensor wrote:
             | It's trivial to build synthetic reasoning datasets, likely
             | even in natural languages. This is a well established
             | technique that works (e.g. see Microsoft Phi, among
             | others).
             | 
             | I said generally because there are things like adversarial
             | training that use a ruleset to help generate correct
             | datasets that work well. Outside of techniques like that
             | it's not just a rule of thumb, it's _always_ true that
             | training on the output of another model will result in a
             | worse model.
             | 
             | https://www.scientificamerican.com/article/ai-generated-
             | data...
        
               | numba888 wrote:
               | > it's always true that training on the output of another
               | model will result in a worse model.
               | 
               | Not convincing.
               | 
               | You can imagine model doing some primitive thinking and
               | coming to conclusion. Then you can train another model on
               | summaries. If everything goes well it will be coming to
               | conclusions quicker. That's at least. Or it may be able
               | solve more complex problems with the same amount of
               | 'thinking'. It will be self-propelled evolution.
               | 
               | Another option is to use one model to produce 'thinking'
               | part from known outputs. Then train another to reproduce
               | thinking to get the right output, unknown to it
               | initially. Using humans to create such dataset would be
               | slow and very expensive.
               | 
               | PS: if it was impossible humans would be still living on
               | the trees.
        
               | tensor wrote:
               | Humans don't improve by "thinking." They improve my
               | natural selection against a fitness function. If that
               | fitness function is "doing better at math" then over a
               | long time perhaps humans will get better at math.
               | 
               | These models don't evolve like they, there is not a
               | random process of architectural evolution. Nor is there a
               | fitness function anything like "get better at math."
               | 
               | A system like AlphaZero works because it has a rules to
               | use as an oracle: the game rules. The game rules provide
               | the new training information needed drive the process.
               | Each game played produces new _correct_ training data.
               | 
               | These LLMs have no such oracle. Their fitness function is
               | and remains: predict the next word, followed by: produce
               | text that makes a human happy. Note that it's not
               | "produce text that makes ChatGPT happy."
        
           | smitelli wrote:
           | > training off of data generated by another AI is generally a
           | bad idea
           | 
           | Ah. So if I understand this... once the internet becomes
           | completely overrun with AI-generated articles of no
           | particular substance or importance, we should not bulk-scrape
           | that internet again to train the subsequent generation of
           | models.
           | 
           | I look forward to that day.
        
             | bangaladore wrote:
             | That's already happened. Its well established now that the
             | internet is tainted. After essentially ChatGPT's public
             | release, a non-insignificant amount of internet content is
             | not written by humans.
        
             | tensor wrote:
             | Yes, this is a real and serious concern that AI researchers
             | have.
        
           | athrowaway3z wrote:
           | Thats not right either.
           | 
           | It proofs we _can_ optimize our training data.
           | 
           | Just like humans have been genetically stable for a long
           | time, the quality & structure of information available to a
           | child today vs that of 2000 years ago makes them more skilled
           | at certain tasks. Math being a good example.
        
           | dragonwriter wrote:
           | > training off of data generated by another AI is generally a
           | bad idea
           | 
           | It's...not, and its repeatedly been proven in practice that
           | this is an invalid generalization because it is missing
           | necessary qualifications, and its funny that this myth keeps
           | persisting.
           | 
           | It's probably a bad idea to use _uncurated_ output from
           | another AI to train a model if you are trying to make a
           | better model rather than a distillation of the first model,
           | and its definitely (and, ISTR, the actual research result
           | from which the false generalization has developed) a bad idea
           | to iteratively fine-tune a model on _its own_ unfiltered
           | output, but there has been lots of success using AI models to
           | generate data which is curated and used to train other
           | models, which can be much more efficient that trying to
           | _create_ new material without AI once you 've gotten to the
           | point where you've already hoovered up all the readily-
           | accessible low hanging fruit of premade content relevant to
           | your training goal.
        
             | LPisGood wrote:
             | It is, of course not going to produce a "child" model that
             | more accurately predicts the underlying true distribution
             | that the "parent" model was trying to. That is, it will not
             | add anything new.
             | 
             | This is immediately obvious if you look at it through a
             | statistical learning lens and not the mysticism crystal
             | ball that many view NN's through.
        
               | FridgeSeal wrote:
               | No no no you don't understand, the models will magically
               | overcome issues and somehow become 100x and do real AGI!
               | Any day now! It'll work because LLM's are basically
               | magic!
               | 
               | Also, can I have some money to build more data centres
               | pls?
        
               | mattnewton wrote:
               | LLMs are no longer trying to just reproduce the
               | distribution of online text as a whole to push the state
               | of the art, they are focused on a different distribution
               | of "high quality" - whatever that means in your domain.
               | So it is possible that this process matches a "better"
               | distribution for some tasks by removing erroneous
               | information or sampling "better" outputs more frequently.
        
               | kybernetikos wrote:
               | Fine tuning an llm on the output of another llm is
               | exactly how deepseek made its progress. The way they got
               | around the problem you describe is by doing this in a
               | domain that can be relatively easily checked for
               | correctness, so suggested training data for fine tuning
               | could be automatically filtered out if it was wrong.
        
               | dragonwriter wrote:
               | > It is, of course not going to produce a "child" model
               | that more accurately predicts the underlying true
               | distribution that the "parent" model was trying to. That
               | is, it will not add anything new.
               | 
               | Unfiltered? Sure. With human curation of the generated
               | data it certainly can. (Even automated curation can do
               | this, though its more obvious that human curation can.)
               | 
               | I mean, I can randomly developed fact claims about
               | addition, and if I curate which ones go into a training
               | set, train a model that reflects addition of integers
               | much more accurately than the random process which
               | generated the pre-curation input data.
               | 
               | Without curation, as I already said, the best you get is
               | a distillation of the source model, which is highly
               | improbable to be more accurate.
        
               | acgourley wrote:
               | This is not obvious to me! For example, if you locked me
               | in a room with no information inputs, over time I may
               | still become more intelligent by your measures. Through
               | play and reflection I can prune, reconcile and generate.
               | I need compute to do this, but not necessarily more
               | knowledge.
        
               | sudosysgen wrote:
               | Again, this isn't how distillation work. Your task as the
               | distillation model is to copy mistakes, and you will be
               | penalized by pruning reconciling and generating.
               | 
               | "Play and reflection" is something else, which isn't
               | distillation.
        
               | Jerrrry wrote:
               | No one knows if the pigeon-hole principle applies
               | absolutely exclusive to the ability to generalize outside
               | of a training set.
               | 
               | That is the existential, $1T question.
        
               | esafak wrote:
               | The latest models create information from base models by
               | randomly creating candidate responses then pruning the
               | bad ones using an evaluation function. The good responses
               | improve the model.
               | 
               | It is not distillation. It's like how you can arrive at
               | new knowledge by reflecting on existing knowledge.
        
           | Voloskaya wrote:
           | > First of all, training off of data generated by another AI
           | is generally a bad idea because you'll end up with a strictly
           | less accurate model (usually).
           | 
           | That is not true at all.
           | 
           | We have known how to solve this for at least 2 years now.
           | 
           | All the latest state of the art models depend heavily on
           | training on synthetic data.
        
         | fumeux_fume wrote:
         | This has been in the back of my head since the news broke. Has
         | anyone built their own R1 from scratch and validated it?
        
           | RevEng wrote:
           | In the last few days? No, that would be impossible; no one
           | has the resources to train a base model that quickly. But
           | there are definitely a lot of people working on it.
        
         | philistine wrote:
         | There's a question of scale here: was it trained on 1000
         | outputs or 5 million?
        
         | bangaladore wrote:
         | That's only true if you assume that O1 synthetic data sets are
         | much better than any other (comparably sized) opensource model.
         | 
         | It's not apparently obvious to me that that is the case.
         | 
         | Ie. do you need a SOTA model to produce a new SOTA model?
        
         | joe_the_user wrote:
         | _If Deepseek trained off OpenAI, then it wasn 't trained from
         | scratch for "pennies on the dollar"_
         | 
         | If OpenAI trained on the intellectual property of others, maybe
         | it wasn't the creativity breakthrough people claim?
         | 
         | Oppositely
         | 
         | If you say ChatGPT was trained on "whatever data was
         | available", and you say Deepseek was trained "whatever data was
         | available", then they sound pretty equivalent.
         | 
         | All the rough consensus language output of humanity is now
         | roughly on the Internet. The various LLMs have roughly
         | distilled that and the results are naturally going to be
         | tighter and tighter. It's not surprising that companies are
         | going to get better and better at solving the same problem. The
         | situation of DeepSeek isn't so much that promises future
         | achievements but that it shows that OpenAI's string of
         | announcements are incremental progress that aren't going to be
         | reaching the AGI that Altman now often harps on.
        
           | el_cujo wrote:
           | I'm not an OpenAI apologist and don't like what they've done
           | with other people's intellectual property but I think that's
           | kind of a false equivalency. OpenAI's GPT 3.5/4 was a big
           | leap forward in the technology in terms of functionality.
           | DeepSeek-r1 isn't really a huge step forward in output, it's
           | mostly comparable to existing models, one thing that is
           | really cool about it is it being able to be trained from
           | scratch quickly and cheaply. This is completely undercut if
           | it was trained off of OpenAI's data. I don't care about
           | adjudicating which one is a bigger thief, but it's notable if
           | one of the biggest breakthroughs about DeepSeek-r1 is pretty
           | much a lie. And it's still really cool that it's open source
           | and can be run locally, it'll have that over OpenAI whether
           | or not the training claims are a lie/misleading
        
             | pertymcpert wrote:
             | Not just the training cost, the inference cost is a
             | fraction of o1.
        
         | alecco wrote:
         | Even if all that about training is true, the bigger cost is
         | inference and Deepseek is 100x cheaper. That destroys
         | OpenAI/Anthropic's value proposition of having a unique secret
         | sauce so users are quickly fleeing to cheaper alternatives.
         | 
         | Google Deepmind's recent Gemini 2.0 Flash Thinking is also
         | priced at the new Deepseek level. It's pretty good (unlike
         | previous Gemini models).
         | 
         | [0] https://x.com/deedydas/status/1883355957838897409
         | 
         | [1] https://x.com/raveeshbhalla/status/1883380722645512275
        
           | [deleted]
        
         | FooBarWidget wrote:
         | Have people on HN never heard of public ChatGPT conversations
         | data sets? They've been mentioned multiple times in past HN
         | conversations and I thought it'd be common knowledge here by
         | now. Pretty much all open source models have been training on
         | them for the past 2 years, it's common practice by now. And
         | haven't people been having conversations about "synthetic data"
         | for a pretty long time by now? Why is all of this suddenly an
         | issue in the context of DeepSeek? Nobody made a fuss about this
         | before.
         | 
         | And just because a model trains on _some_ ChatGPT data, doesn
         | 't mean that that data is the majority. It's just another
         | dataset.
        
       | hyperbovine wrote:
       | Live by the sword...
        
       | rcarmo wrote:
       | I guess their CEO was too busy to write something in defense of
       | US export controls
       | (https://news.ycombinator.com/item?id=42866905), or (even more
       | scary) he doesn't need to anymore.
        
       | colonelspace wrote:
       | No honour among thieves
        
       | TrackerFF wrote:
       | Next up: <<DeepSeek models are a national security risk, we must
       | block access!>>
        
         | jondwillis wrote:
         | Download your weights while you still can I guess...
        
       | 52-6F-62 wrote:
       | I heard they were just "democratizing" llm and ai development.
       | 
       | Yesterday the industry crushed pianos and tools and bicycles and
       | guitars and violins and paint supplies and replaced them with a
       | tablet computer.
       | 
       | Tomorrow we can replace craven venture capitalists and overfed
       | corporate bodies with incestuous LLM's and call it all a day.
        
       | zb3 wrote:
       | DeepSeek actually opening ClosedAI up makes me like them even
       | more.. this is great :)
        
       | conartist6 wrote:
       | It seems to be undermined by the same principle that says that
       | going into a library and reading a book there is not stealing
       | when you walk out with the knowledge from the book.
       | 
       | OpenAI seems to feel that way about the their use of copyrighted
       | material: since they didn't literally make a copy of the source
       | material, it's totally fair game. It seems like this is the same
       | argument that protects DeepSeek if indeed they did this. And why
       | not, reading a lot of books from the library is a way to get
       | smarter, and ostensibly the point of libraries
        
       | jgrall wrote:
       | It's not a good look when your technology is replicated for a
       | fraction of the cost, and your response is to smear your
       | competition with (probably) false accusations and cozy up to the
       | US government to tighten already shortsighted export controls.
       | Hubris & xenophobia are not going to serve American companies
       | well. Personally I welcome the Chinese - or anyone else for that
       | matter - developing advanced technologies as long as they are
       | used for good. Humanity loses if we allow this stuff to be
       | "owned" by a handful of companies or a single country.
        
       | HPsquared wrote:
       | AI models are becoming like perpetual stew.
        
       | guybedo wrote:
       | This is hilarious.
       | 
       | Everybody has evidence OpenAI scraped the internet at a global
       | scale and used terabytes of data it didn't pay for. Newspapers,
       | books, etc...
        
       | aDyslecticCrow wrote:
       | And they used all copyrighted data on the internet. If they wanna
       | sue, they set a dangerous precedent.
        
       | josefritzishere wrote:
       | OpenAI, who comitted copyright infringement on an massive scale,
       | wants to defend against a superior product won the basis of
       | infringement? What nonsense.
        
       | metaxz wrote:
       | I don't understand how OpenAI claims it would have happened. The
       | weights are closed and as far as I read they are not complaining
       | Deepseek hacked them and obtained the weight. So all they could
       | do was to query OpenAI and generate test data. But how much did
       | they query really - I would suppose it would require a huge
       | amount done via an external, paid-for API? Is there any proof of
       | this besides OpenAI saying it? Even if we suppose it is true, I
       | suppose this must have happened via the API so they paid per
       | token etc. So they paid for each and every token of training
       | data. As I understand, the requester owns the copyright on what
       | is generated by OpenAI's models and is free to do what they want.
        
       | ks2048 wrote:
       | The schadenfreude and irony of this is totally understandable.
       | 
       | But, I wonder - do companies like OpenAI, Google, and Anthropic
       | use each others models for training? If not, is it because they
       | don't want to or need to, or because they are afraid of breaking
       | the ToC?
        
       | Digit-Al wrote:
       | So... company that steals other people's work to train their
       | models is complaining because they think someone stole their work
       | to train their models.
       | 
       | Cry me a river.
        
       | baggiponte wrote:
       | OpenAI coping so hard
        
       | _moof wrote:
       | This reminds me of a (probably apocryphal) story about fast food
       | chains that made the rounds decades ago: McDonald's invests tons
       | of time into finding the best real estate for new stores; Burger
       | King just opens stores near McDonalds!
        
       | dragonwriter wrote:
       | Hey, OpenAI, so, you know that legal theory that is the entire
       | basis of your argument that any of your products are legal?
       | "Training AI on proprietary data is a use that doesn't require
       | permission from the owner of the data"?
       | 
       | You might want to consider how it applies to this situation.
        
       | buyucu wrote:
       | I have no sympathy for OpenAI here. They are (allegedly) a non-
       | profit with open in the title that refuse to open-source their
       | models.
       | 
       | They are now upset at a startup who is more loyal to OpenAI's
       | original mission that OpenAI is today.
       | 
       | Please, give me a break.
        
       | Jotalea wrote:
       | I really hate when there is a paywall to read an article. It
       | makes me not want to read it anymore.
        
       | mbowcut2 wrote:
       | So, is this just an example of the first-mover disadvantage (or
       | maybe the problem of producing public goods?). The first AI
       | models were orders of magnitude more expensive to create, but now
       | that they're here we can, with techniques like distillation,
       | replicate them at a fraction of the cost. I am not really
       | literate in the law but weren't patents invented to solve
       | problems like this?
        
       | adam_arthur wrote:
       | Who cares?
       | 
       | They did the exact same thing with public information. Their
       | model just synthesizes and puts out the same information in a
       | slightly different form.
       | 
       | Next we should sue students for repeating the words of their
       | teachers
        
       | moralestapia wrote:
       | Called it from day 0, impossible to reach that performance with
       | 5M, they _had_ to distill OpenAI (or some other leading
       | foundational model).
       | 
       | Got downvoted to oblivion by people who haven't been told what to
       | think by MSM yet. Now it's on FT and everywhere, good, what
       | matters is that truth comes out eventually.
       | 
       | I don't take any sides and think what DeepSeek did is fair play,
       | however, what I do find harmful about this is, what incentive
       | would company A have to spend billions training a new frontier
       | model if all of that could be then reproduced by company B at a
       | fraction of the cost?
        
         | kgeist wrote:
         | The "evidence" is very weak though:
         | 
         | >The San Francisco-based ChatGPT maker told the Financial Times
         | it had seen some evidence of "distillation", which it suspects
         | to be from DeepSeek.
         | 
         | Given that many people have been using ChatGPT to distill their
         | fine-tunes for a few years now, how can they be sure it was
         | specifically DeepSeek? There's, say, glaive.ai whose entire
         | business model is to sell you synthetic datasets, probably
         | generated with ChatGPT as well.
        
           | moralestapia wrote:
           | I agree that the evidence is weak, and even if they had some,
           | they cannot really do anything.
           | 
           | To me, it's just very likely they distilled GPT-4, because:
           | 
           | 1) Again, you just cannot get that performance at that cost.
           | And no, what they describe on the paper is not enough to
           | explain the 1,000x-fold decrease in cost.
           | 
           | 2) Very often, DeepSeek tells you it's ChatGPT or OpenAI;
           | it's actually quite easy to get it to do that. Some say
           | that's related to "the background radiation on the post-AI
           | internet". I'm not a fentanyl consumer so, unfortunately, I
           | think that argument is trash.
        
             | kgeist wrote:
             | If it's just a distillation of GPT-4, wouldn't we expect it
             | to have worse quality than o1? But I've seen countless
             | examples of DeepSeek-r1 solving math problems that o1
             | cannot.
             | 
             | >Very often, DeepSeek tells you it's ChatGPT or OpenAI;
             | it's actually quite easy to get it to do that. Some say
             | that's related to "the background radiation on the post-AI
             | internet". I'm not a fentanyl consumer so, unfortunately, I
             | think that argument is trash.
             | 
             | The exact same thing happened with Llama. Sometimes it also
             | claimed to be Google Assistant or Amazon Alexa.
        
               | moralestapia wrote:
               | >wouldn't we expect it to have worse quality than o1?
               | 
               | That's tricky, you can optimize a model to do real well
               | on synthetic benchmarks.
               | 
               | That said, DeepSeek performs a bit worse than GPT-4 in
               | general and substantially wrong on benchmarks like ARC
               | which is designed with this in mind.
        
               | kgeist wrote:
               | Are you sure you checked R1 and not V3? By default, R1 is
               | disabled in their UI.                 Prompt: Find an
               | English word that contains 4 'S' letters and 3 'T'
               | letters.            Deepseek-R1: stethoscopists (correct,
               | thought for 207 seconds)            ChatGPT-o1:
               | substantialists (correct, thought for 188 seconds)
               | ChatGPT-4o: statistics (wrong) (even with "let's think
               | step by step")
               | 
               | In almost every example I provide, it's on par with o1
               | and better than 4o.
               | 
               | >substantially wrong on benchmarks like ARC which is
               | designed with this in mind.
               | 
               | Wasn't it revealed OpenAI trained their model on that
               | benchmark specifically? And had access to the entire
               | dataset?
        
               | moralestapia wrote:
               | That prompt means nothing. Check out the benchmarks.
               | 
               | Also, compare V3 to 4o and R1 to o1, that's the right
               | way.
        
               | esafak wrote:
               | No, because it is not a distillation, but an _extension_.
               | A selling point of the model is using RL to push past the
               | quality of the base model.
        
       | nachox999 wrote:
       | Ask DeepSeek and ChatGPT: "name three persons"; the answer may
       | surprise you
        
       | curvaturearth wrote:
       | Something about the outputs becoming the inputs to then produce
       | more outputs is just plain funny
        
       | B1FF_PSUVM wrote:
       | "Cry me a river" is a phrase I haven't heard recently, for some
       | reason ...
        
       | asdefghyk wrote:
       | Deepseek did not respect OpenAI's copyright?
       | 
       | Well who would have thought that?
        
       | nataliste wrote:
       | A Wolf had stolen a Lamb and was carrying it off to his lair to
       | eat it. But his plans were very much changed when he met a Lion,
       | who, without making any excuses, took the Lamb away from him.
       | 
       | The Wolf made off to a safe distance, and then said in a much
       | injured tone:
       | 
       | "You have no right to take my property like that!"
       | 
       | The Lion looked back, but as the Wolf was too far away to be
       | taught a lesson without too much inconvenience, he said:
       | 
       | "Your property? Did you buy it, or did the Shepherd make you a
       | gift of it? Pray tell me, how did you get it?"
       | 
       | What is evil won is evil lost.
        
       | m3kw9 wrote:
       | So if OpenAI didn't have these outputs for distillation, Deepseek
       | wouldn't exist?
        
       | hedayet wrote:
       | Beyond the irony of their stance, this reflects a failure of
       | OpenAI's technical leadership--either in oversight or in
       | designing a system that enables such behavior.
       | 
       | But in capitalism, we, the customers aren't going to focus on how
       | models are trained or products are made; we only care about
       | favourable pricing.
       | 
       | A key takeaway for me from this news is the clause in OpenAI's
       | terms and conditions. I mistakenly believed that paying for
       | OpenAI's API granted full rights to the output, but it turns out
       | we're only buying specific rights (which is now another reason
       | we're going to start exploring alternatives to OpenAI)
        
       | mtlmtlmtlmtl wrote:
       | So, what is this evidence? I'll believe it when I see it. Right
       | now all we really have is some vague rumours about some API
       | requests. How many requests? How many tokens? Over how long of a
       | time period? Was it one account or multiple, if the latter, how
       | many? How do they know the activity came from deepseek? How do
       | they know the data was actually used to train Deepseek
       | models(could have just been benchmarking against the
       | competition)?
       | 
       | If all they really have is some API requests, even assuming
       | they're real _and_ originated by Deepseek, that 's very far from
       | proof that any of it was used as training data. And honestly,
       | short of commiting crimes against Deepseek(hacking), I'm not sure
       | how they even could prove that at this point, from their side
       | alone.
       | 
       | And what's even more certain is that a vague insistence that
       | evidence exists, accompanied by a denial to shed any more light
       | on the specifics, is about as informative as saying nothing at
       | all. It's not like OpenAI and Microsoft have a habit of
       | transparency and honesty in their communication with the public,
       | as proven by an endless laundry list of dishonest and subversive
       | behaviour.
       | 
       | In conclusion, I don't see why I should give this any more
       | credence than I would a random anon on 4chan claiming a pizza
       | place in Washington DC is the centre of a child sex trafficking
       | ring.
       | 
       | P.S: And to be clear, I really don't care if it is true. If
       | anything, I hope it is; it would be karmic justice at its finest.
        
       | halyconWays wrote:
       | Oh no, so sad. The Open non-profit that steals 100% of all
       | copyrighted content and makes multiple billion-dollar for-profit
       | deals while releasing no weights is crying. This is going to ruin
       | my sleep. :(
        
       | htrp wrote:
       | In other news.....water is wet
        
       | 1propionyl wrote:
       | At this point, the only thing that keeps me using ChatGPT is o1
       | w/ RAG. The usage limits on o1 are prohibitively tight for
       | regular use, so I have to budget usage to tasks that would
       | benefit there. I also have significant misgivings about their
       | policies around output, which also limit what I can use it for.
       | 
       | For local tasks, the deepseek-r1:14b and deepseek-r1:32b
       | distillations immediately replace most of that usage (prior local
       | models were okay, but not consistently good enough). Once there's
       | a "just works" setup for RAG on par with installing ollama (which
       | I doubt is far of), I don't see much reason to continue paying
       | for my subscription.
       | 
       | Sadly, like many others in this thread, I expect under the
       | current administration to see self-hamstringing protectionism
       | further degrade the US's likelihood of remaining a global
       | powerhouse in this space. Betting the farm on the biggest first-
       | mover who can't even keep up with competition, has weak to non-
       | existent network effects (I can choose a different model or
       | service with a dropdown, they're more or less fungible), has no
       | technological moat and spent over a year pushing apocalyptic
       | scenarios to drum up support for a regulatory moat...
       | 
       | ...well it just doesn't seem like a great idea to me.
        
       | divbzero wrote:
       | I was wondering if this might be the case, similar to how Bing's
       | initial training included Google's search results [1]. I'd be
       | curious to see more details of OpenAI's evidence.
       | 
       | It is, of course, quite ironic for OpenAI to indiscriminately
       | scrape the entire web and then complain about being scraped
       | themselves.
       | 
       | [1]: https://searchengineland.com/google-bing-is-cheating-
       | copying...
        
       | schaefer wrote:
       | I mean, if openAI claims they can train on the world's novels and
       | blogs with "no harm done" (i.e: no copyright infringement and no
       | royalties due), then it directly follows that we can train both
       | our robots and our selves on the output of openAI's models in
       | kind.
       | 
       | Right?
        
       | davesque wrote:
       | I recently thought of a related question. Actually, I'm almost
       | certain that foundation model trainers have thought of this. The
       | question is to what extent are popular modern benchmarks (or any
       | reference to them, or description of them, etc.) bring scrubbed
       | from the training data? Or are popular benchmarks designed in
       | such a way that they can be re-parametrized for each run? In any
       | case, it seems like a surprisingly hard problem to deal with.
        
       | highfrequency wrote:
       | If true, the question is: did they use ChatGPT outputs to create
       | Deepseek V3 only, or is the R1-zero training process a complete
       | lie (given that the whole premise is that they used pure
       | reinforcement learning)? If they only used ChatGPT output when
       | training V3, then they succeeded in basically replicating the
       | jump from ChatGPT-4o to o1 without any human-labeled CoT (and
       | published the results) - which is a big achievement on its own.
        
       | henry_viii wrote:
       | So Meta can train its AI on all the pirated books in the world
       | but people are losing their mind over an AI learning from another
       | AI?
        
         | esafak wrote:
         | People here have been vocal against training on any unlicensed
         | content.
        
       | nuc1e0n wrote:
       | And OpenAI scrapped the public internet to train its models.
        
       | boxedemp wrote:
       | Deep refers to itself as ChatGPT sometimes lol
        
       | LZ_Khan wrote:
       | I actually think what DeepSeek did will _slow down_ AI progress.
       | What 's the incentive to spend billions developing frontier
       | models if once it's released some shady orgs in unregulated
       | countries can just scrape your model outputs, reproduce it, and
       | undercut you in cost?
       | 
       | OpenAI is like a team of fodder monkeys stepping on landmines
       | right now, with the rest of the world waiting behind them.
        
       | zx10rse wrote:
       | OpenAI is already irrelevant but the audacity oh my.
        
       | cumulative00x wrote:
       | There is a saying in Turkish that roughly goes like this, it
       | takes a thief to catch a thief. I am not a big fan of China's
       | tech, too, however, it amuses me to watch how big tech charlatans
       | have been crying over Deepseek shock.
        
         | gosub100 wrote:
         | It's true irony to see thieves getting stolen from.
        
       | DidYaWipe wrote:
       | They have "open" right in their name, so...
       | 
       | Objection overruled.
        
       | SubiculumCode wrote:
       | If you have a set of weights A, can you derive another set of
       | weights B that function (near) identically as A AND a) not appear
       | to be the same weights as A when inspected superficially b)
       | appear uncorrelated when inspecting the weight matrices?
        
         | rahimnathwani wrote:
         | Do you mean for a given model structure, can two sets of
         | weights give substantially the same outputs?
         | 
         | Even if that were possible, it would be suspicious if you were
         | to release an open model whose model architecture is identical
         | to that of a closed one from a competitor.
         | 
         | If that is what happened, we'd know about it by now.
        
       | fimdomeio wrote:
       | But what is the problem here? Isn't open AI mission "to ensure
       | that artificial general intelligence benefits all of humanity"?
       | Sounds like success to me.
        
       | hugoromano wrote:
       | OpenAI initially scraped the web and later formed partnerships to
       | train on licensed data. Now, they claim that DeepSeek was trained
       | on their models. However, DeepSeek couldn't use these models for
       | free and had to pay API fees to OpenAI. From a legal standpoint,
       | this could be seen as a violation of the terms and conditions.
       | While I may be mistaken, it's unclear how DeepSeek could have
       | trained their models without compensating OpenAI. Basically,
       | OpenAI is saying machines can't learn from their outputs as
       | humans do.
        
       | buildsjets wrote:
       | Womp Womp.
        
       | asdfasdf1 wrote:
       | it's no crime to steal from a thief
        
         | krapp wrote:
         | It is actually a crime to steal from a thief.
        
       | wendyshu wrote:
       | If distillation gives you a cheaper model with similar accuracy,
       | why doesn't OpenAI distill its own models?
        
       | mkoubaa wrote:
       | OpenAI made a lot of contributions to LLMs obviously but the
       | amount of fraud, deception, and dark patterns coming out of that
       | organization make me root against it.
        
       | ysofunny wrote:
       | I see this as China fighting U.S. of A (or the American Dollar
       | versus Chinese Renmibi if you will)
       | 
       | and this is good because any alternatives I can think of are
       | older-school fighting
       | 
       | modern war is seeped in symbolism, but the contest is still there
       | 
       | e.g. whose dong is bigger? Xi Jingping's or Dnld Trump's
        
       | maxglute wrote:
       | Not that DeepSeek is luigi mangione, but it's pretty funny OpenAi
       | getting the dead ceo treatment.
        
       | mrkpdl wrote:
       | The cat is out of the bag. This is the landscape now, r1 was made
       | in a post-o1 world. Now other models can distill r1 and so on.
       | 
       | I don't buy the argument that distilling from o1 undermines deep
       | seek's claims around expense at all. Just as open AI used the
       | tools 'available to them' to train their models (eg everyone
       | else' data), r1 is using today's tools.
       | 
       | Does open AI really have a moral or ethical high ground here?
        
         | ijidak wrote:
         | Plus, it suggests OpenAI never had much of a moat.
         | 
         | Even if they win the legal case, it means weights can be
         | inferred and improved upon simply by using the output that is
         | also your core value add (e.g. the very output you need to sell
         | to the world).
         | 
         | Their moat is about as strong as KFC's eleven herbs and spices.
         | Maybe less...
        
       | yapyap wrote:
       | It _sounds_ like they're just jealous and trying to smear shit
       | over the wall and see what sticks.
       | 
       | DeepSeek just bodied u bro, get back in the lab & create a better
       | AI instead of all this news that isn't gonna change them having a
       | good AI
        
       | FpUser wrote:
       | Pot calling kettle black?
        
       | almostdeadguy wrote:
       | Hope Sam Altman is getting his money's worth out of that Trump
       | campaign contribution. Glorious days to be living under the term
       | of a new Boris Yeltsin. Pawning and strip-mining the federal
       | apparatus to the most loyal friends and highest bidders.
        
       | ijidak wrote:
       | This whole argument by OpenAI suggests they never had much of a
       | moat.
       | 
       | Even if they win the legal case, it means weights can be inferred
       | and improved upon simply by using the output that is also your
       | core value add (e.g. the very output you need to sell to the
       | world).
       | 
       | Their moat is about as strong as KFC's eleven herbs and spices.
       | Maybe less...
        
       ___________________________________________________________________
       (page generated 2025-01-29 23:00 UTC)