[HN Gopher] OpenAI says it has evidence DeepSeek used its model ...
___________________________________________________________________
OpenAI says it has evidence DeepSeek used its model to train
competitor
Author : timsuchanek
Score : 330 points
Date : 2025-01-29 04:21 UTC (18 hours ago)
(HTM) web link (www.ft.com)
(TXT) w3m dump (www.ft.com)
| udev wrote:
| https://archive.is/KiSYM
| cratermoon wrote:
| Ironic, OpenAI claiming someone else stole their work.
| vinni2 wrote:
| How would they prove they used it's model. I would be curious to
| know their methodology. Also what legal actions OpenAI can take?
| can DeepSeek be banned in US?
| iforgot22 wrote:
| They might show DeepSeek's model calling itself ChatGPT, which
| users have already alleged. Same as how Cisco proved Huawei was
| stealing router code.
|
| Except in this case, nothing was stolen, unless they want to
| call ChatGPT's own training on source data theft too.
| freehorse wrote:
| ChatGPT outputs are all over the internet. It is harder to
| prove that deepseek used specifically o1 for training,
| instead of a lot of chatgpt output ending up in the training
| set from other sources.
| iforgot22 wrote:
| That's a good point, at least for the prompts I saw. Like
| "do you have an app I can use" is commonly seen with
| "here's the ChatGPT app" online. And maybe they don't add
| anything telling Deepseek that it's Deepseek.
| belter wrote:
| The subtitle is the gold... : "White House AI tsar David Sacks
| raises possibility of alleged intellectual property theft"
| conartist6 wrote:
| lolololololololol
| vrighter wrote:
| So what? They probably paid for api access just like everyone
| else. So it's a TOS violation at worst. Go ahead, open a civil
| suit in the US against an entity the US courts do not have
| jurisdiction over and quit whining...
| jhickok wrote:
| >open a civil suit in the US against an entity the US courts do
| not have jurisdiction over
|
| Yeah, over a Chinese company no less.
| ForHackernews wrote:
| What's good for the goose is good for the gander. Obviously a
| transformative work and not an intellectual property violation
| any more than OpenAI injesting every piece of media in existence.
| dagelf wrote:
| Injesting is sure the right take. What a circus!
| amarcheschi wrote:
| I quite like a scenery where llm output can't be copyrighted, so
| that it is possible to eventually train a llm with data from the
| previous one(s)
| layer8 wrote:
| OpenAI argues it's a violation of their terms of service. So
| there are legal issues if it can be proven.
| mannewalis wrote:
| But OpenAI's model isn't open source, how would they distill
| knowledge without direct access to the model?
| layer8 wrote:
| You don't need direct access for LLM distillation, just
| regular API access.
| mannewalis wrote:
| ok I looked it up and have a better understanding now.
| Palmik wrote:
| Legal issues for who?
|
| Company A pays OpenAI for their API. They use the API to
| generate or augment a lot of data. They own the data. They
| post the data on the open Internet.
|
| Company B has the habit of scraping various pages on the
| Internet to train its large language models, which includes
| the data posted by Company A. [1]
|
| OpenAI is undoubtedly breaking many terms of service and
| licenses when it uses most of the open Internet to train its
| models. Not to mention potential copyright violations (which
| do not apply to AI outputs).
|
| [1]: This is not hypothetical BTW. In the early days of LLMs,
| lots of large labs accidentally and not so accidentally
| trained on the now famous ShareGPT dataset (outputs from
| ChatGPT shared on the ShareGPT website).
| layer8 wrote:
| For both.
| Palmik wrote:
| Posting OpenAI generated data on the internet is not
| breaking the ToS. This is how most OpenAI based
| businesses operate, after all [1] (e.g. various
| businesses to generate articles with AI, various chat
| businesses that let you share your chats, etc.)
|
| OpenAI is one of the companies like Company B that is
| using data from the open Internet.
|
| [1] Ownership of content. As between you and OpenAI, and
| to the extent permitted by applicable law, you (a) retain
| your ownership rights in Input and (b) own the Output. We
| hereby assign to you all our right, title, and interest,
| if any, in and to Output.
| top_sigrid wrote:
| https://archive.is/KiSYM
| lawlessone wrote:
| So they're mad someone did exactly what they did?
| exe34 wrote:
| no, no, it's completely different. "open"AI stole from poor
| people. DeepSeek stole from a $1T company. that's illegal!
| whatshisface wrote:
| It's reasonably likely that a lot of people linked to the federal
| government want to ban DeepSeek. You can tell it's being
| presented away from "they gave us a free set of weights" and
| towards "they destroyed $1T of shareholder value." (By revealing
| that Microsoft et al. paid way too much to OpenAI et al. for
| technology that was actually easy to reinvent.)
| fullshark wrote:
| Would it even matter? Isn't the cat out of the bag and
| everything they did repeatable by an American research team?
| Cumpiler69 wrote:
| It matters because their goal was hyping up how advanced and
| difficult their tech is, propping up their valuations.
|
| DeepSeek proved the emperor had no clothes and wiped out a
| lot of their valuation when investors saw reaching parity to
| Chtgpt is not really that difficult.
| mastazi wrote:
| I think parent was asking would it even matter if there was
| a ban. To which the answer would be "no" because as you
| said the point has been made. And, as parent pointed out,
| it's repeatable anyway.
| bhouston wrote:
| It doesn't matter from the US government perspective if all
| of the tech is replicated by US companies and US user
| continue to use US AI technology. But if US users start to
| use Chinese AI tech, then protectionism urges will appear
| that will likely figure out how to ban its use or subject it
| to large tariffs (e.g. TikTok, BYD, network equipment, solar
| panels, etc.)
| whatshisface wrote:
| American researchers had already made enough progress to
| prove that LLMs were not an incomprehensible trade secret
| based on years of secret knowledge - investors and tech
| executives were simply lead to believe otherwise. Well-
| connected people are probably very mad about this and they
| may try to lash out like the emotional human beings they are.
| Cumpiler69 wrote:
| _> By revealing that Microsoft et al. paid way too much to
| OpenAI et al. for technology that was actually easy to
| reinvent._
|
| That's why it's called a bubble. Pretty sure my great great
| grandad also overpaid for some tulips.
| toomuchtodo wrote:
| > "they destroyed $1T of shareholder value." (By revealing that
| Microsoft et al. paid way too much to OpenAI et al. for
| technology that was actually easy to reinvent.)
|
| The value was highly speculative, an illusion created by PR and
| sentiment momentum. "Hype value" not real value (unless you're
| able to realize it and dump those bags on someone else before
| fundamentals set in). Same thing happening with power companies
| downstream of the discovery that AI is not going to be a savior
| of sagging electricity demand. Overdriving the fundamentals is
| not value destruction, it is "I gambled and lost."
|
| https://www.bloomberg.com/news/articles/2025-01-28/deepseek-...
| | https://archive.today/mCemf
|
| "In the short run, the market is a voting machine but in the
| long run, it is a weighing machine."
| cft wrote:
| Since the time when companies en masse stopped paying cash
| dividends on owned shares, the value has become highly
| speculative. In the absence of dividend payments, the stock
| pricing mechanism is not essentially different from Solana or
| Ethereum "price" discovery.
| toomuchtodo wrote:
| I don't disagree that price discovery is harder, but I can
| with more certainty give an honest valuation of CLF or DOW
| vs OpenAI's "who knows what money will look like after we
| succeed, you should view your investment as a donation"
| nonsense. Speculation is inevitable when forward looking,
| but there is a difference between error bars and various
| projections vs unicorns.
|
| Due diligence never goes out of style.
| JumpCrisscross wrote:
| > _when companies en masse stopped paying stock dividends_
|
| Do you mean cash dividends [1]?
|
| Also, the premise is false. Dividend yields have roughly
| tracked interest rates [2]. (The difference is a dirty
| component of the equity risk premium [3].)
|
| [1] https://www.investopedia.com/ask/answers/05/stockcashdi
| viden...
|
| [2] https://www.multpl.com/s-p-500-dividend-yield/table/by-
| year
|
| [3] https://www.investopedia.com/investing/calculating-
| equity-ri...
| cft wrote:
| I changed the typo, thanks. Chash dividends. This
| analysis does not negate common sense: when a company
| does not pay cash dividends, owning its stock is purely
| speculative, like owning Solana. When it does, you get
| cash dividends funded by the company's tangible revenue,
| proportional to your number of shares.
| DebtDeflation wrote:
| What they really destroyed was the idea that OpenAI would be
| able to charge $200/month for their ChatGPT Pro subscription
| which includes o1. That was always ridiculous IMO. The Free
| tier and $20/month Plus tier along with their API business
| (minus any future plan to charge a ridiculous amount for API
| access to o1) will be fine.
| toomuchtodo wrote:
| > The Free tier and $20/month Plus tier along with their
| API business (minus any future plan to charge a ridiculous
| amount for API access to o1) will be fine.
|
| Do the unit economics make this sustainable?
| DebtDeflation wrote:
| If only there were a way to make the models more
| efficient. Oh wait.
| jl6 wrote:
| But doesn't Deepseek's innovation apply only to training,
| not inference?
| Zacharias030 wrote:
| Actually no! If we take their paper at face value, the
| crucial innovation to get a strong model with efficiency
| is their much reduced KV cache and their MoE approach: -
| where a standard model needs to store two large vectors
| for each token at inference time (and load/store those
| over and over from memory) deepseek v3/R1 only stores one
| smaller vector C that is a ,,compression" from which the
| large k,v vectors can be decoded on the fly. - They use a
| fairly standard Mixture of Expert (MoE) approach, which
| works well in training with their tricks, but whose
| inference time advantages are immediate and equal to all
| other MoE techniques, which is to say that from ~85% of
| the 600B+ params that are inside the MoE layers, the
| model at each token inference step will only pick a small
| fraction to use. This reduces FLOPs and memory io by a
| large factor in comparison to a so-called dense model
| where all weights are used for every token (cf Llama 3
| 405B)
| freeone3000 wrote:
| Reducing R&D expense also reduces breakeven price.
| scarface_74 wrote:
| The two podcasters who do the Acquired podcast spoke to
| Ballmer about some of Microsoft's failed initiatives and
| acquisitions. He told them that at the end of the day "it's
| only money".
|
| All of the BigTech companies have enough cash flow from
| profitable lines of business to make speculative bets.
| azemetre wrote:
| It must be EZ mode to be a big tech executive, you somehow
| have all the power to make every decision while also having
| the ability to never take the fault for these decisions.
| scarface_74 wrote:
| I would much rather have a company with a culture that
| isn't afraid to take calculated risks and not be afraid
| of repercussions when they take risk as long as it
| doesn't cause consumer harm.
| azemetre wrote:
| "Not doing consumer harm" is carrying a lot of weight
| there.
|
| Either way what you describe is perfectly achievable for
| the workers, but at some point management needs to own up
| to their failures and getting rewarded because the board
| is also made up of executives at other big tech companies
| is a perverse incentive to never actually improve.
| scarface_74 wrote:
| How did Microsoft's losing bets do consumer harm?
| azemetre wrote:
| I mean forcing copilot everywhere I don't want it
| (nowhere) while jacking up prices to justify it and using
| Windows 11 to serve ads is harmful to me. There's also
| you know... the anticompetitive company that thinks
| buying new sectors is healthy.
| scarface_74 wrote:
| Today, Microsoft's revenue mostly comes from Office and
| Azure. All except PowerPoint were written and designed by
| MS.
| onlyrealcuzzo wrote:
| Theoretically this should be good for OpenAI - in that they can
| reduce their costs by ~27x and pass that along to end users to
| get more adoption and more profit.
| ceejayoz wrote:
| No; those costs were their moat.
| onlyrealcuzzo wrote:
| You don't need a moat when you're in first place.
|
| Their moat is >1B people are already using ChatGPT monthly.
|
| They aren't going to switch unless something is
| substantially better.
| ceejayoz wrote:
| > You don't need a moat when you're in first place.
|
| Tell that to Friendster/MySpace and Facebook.
| onlyrealcuzzo wrote:
| Cute - but MySpace didn't have >1B users, it didn't even
| have 10M when Facebook launched.
|
| Try again.
| cjbgkagh wrote:
| It had more than Facebook when Facebook launched so I'm
| not sure what your point is
| ceejayoz wrote:
| Nothing had a billion users; the _Internet_ didn 't at
| the time.
|
| MySpace and Friendster both spent significant time as the
| #1 social sites. Facebook unseated them rapidly. The same
| is possible for OpenAI.
| shmeeed wrote:
| Dude, if you seriously believe OpenAI has 1B active
| users, you should go touch grass. Actual estimates are
| 100-200 million, about a magnitude lower.
| lm28469 wrote:
| > They aren't going to switch unless something is
| substantially better.
|
| Except one product is 100% free and the other is mostly
| locked behind paid subscriptions
| scarface_74 wrote:
| How long can DeepSeek stay free?
|
| It's already unable to keep up with demand, it will never
| be the default on mobile devices and businesses in the US
| will never trust it.
| ceejayoz wrote:
| That's not really the important question.
|
| The important question is "will this and similar
| optimizations to come permit local LLM use, cutting
| OpenAI out of the equation entirely?"
| scarface_74 wrote:
| Businesses don't even want to maintain servers locally.
| They definitely aren't going to start managing servers
| beefy enough to run LLMs and try to run then with the
| reliability, availability, etc of cloud services.
|
| This will make the cloud providers - especially AWS, GCP
| and to a lesser extent the also ran clouds more valuable.
| The other models hosted by AWS on Bedrock are already
| "good enough" for most business use cases.
|
| And then consumers are definitely not going to be running
| LLMs locally on their computers to replicate ChatGPT (the
| product) anymore than they are going to get an FTP
| account, mount it locally with curlftpfs, and then using
| SVN or CVS on the mounted filesystem and then from
| Windows or Mac, accessed the FTP account through built-in
| software instead of using cloud storage like Dropbox. [1]
|
| Whether someone comes up with a better _product_ than
| ChatGPT and overcome the brand awareness is yet to be
| seen.
|
| [1] Also the iPod had no wireless, less space than the
| Nomad and was lame.
| ceejayoz wrote:
| > And then consumers are definitely not going to be
| running LLMs locally on their computers to replicate
| ChatGPT...
|
| Not _personally_. They 'll let Apple handle it for them.
|
| (This is already a thing.
| https://machinelearning.apple.com/research/introducing-
| apple...)
| scarface_74 wrote:
| There is a reason I kept emphasizing the ChatGPT
| _product_. The (paid) ChatGPT product is not just a text
| based LLM. It can interpret images, has a built in Python
| runtime to offload queries that LLMs aren't good at like
| math, web search, image generation, and a couple of other
| integrations.
|
| The local LLM on iPhones are literally 1% as powerful as
| the server based models like 4o.
|
| That's not even considering battery considerations
| ceejayoz wrote:
| > The local LLM on iPhones are literally 1% as powerful
| as the server based models like 4o.
|
| Currently, yes. That's why this is a compelling advance -
| it makes local LLMs much more feasible, especially if
| this is just the first of many breakthroughs.
|
| A lot of the hype around OpenAI has been due to the fact
| that buying enough capacity to run these things wasn't
| all that feasible for competitors. Now, it is,
| potentially even at the local level.
| sandclock wrote:
| That is exactly what a moat is. Keeping others out.
|
| moat noun a deep, wide ditch surrounding a castle, fort,
| or town, typically filled with water and intended as a
| defense against attack.
| ceejayoz wrote:
| Precisely. First place _needs_ the moat.
|
| Second place just needs a catapult and a diseased cow.
| like_any_other wrote:
| > Their moat is >1B people are already using ChatGPT
| monthly.
|
| Unlike a social network, network effects won't help them
| - their users don't care how many other users they have,
| only about the AI output quality.
|
| > They aren't going to switch unless something is
| substantially better.
|
| Or approximately as good but cheaper.
| onlyrealcuzzo wrote:
| > Or approximately as good but cheaper.
|
| You're fooling yourself if you think OpenAI is going to
| pass up implementing the same strategies to get a ~27x
| cheaper model.
|
| > Unlike a social network, network effects won't help
| them - their users don't care how many other users they
| have, only about the AI output quality.
|
| Google Search doesn't have a network effect. Everyone on
| HN has been saying Google Search is complete garbage for
| a decade. It still has the same market share (roughly) as
| it did a decade ago.
| like_any_other wrote:
| > You're fooling yourself if you think OpenAI is going to
| pass up implementing the same strategies to get a ~27x
| cheaper model.
|
| But that would mean a 27x lower valuation.
| onlyrealcuzzo wrote:
| > But that would mean a 27x lower valuation.
|
| No.
|
| Valuations are based on future profits. Not future
| revenues.
|
| You can theoretically lower your costs by 27x and end up
| with 2x more future profits - if you're actually 45x
| cheaper (which DeepSeek's method claims to be).
| ceejayoz wrote:
| > Valuations are based on future profits.
|
| Which are estimated, in significant part, by the chance
| of a competitor arising.
|
| If the barriers of entry are much lower than originally
| thought, the potential profit margin plummets.
| like_any_other wrote:
| You mean charge a 27x lower price, but have 45x lower
| costs, so your profit margin has doubled?
|
| Your _relative_ margin may have doubled, but your
| absolute profit-per-item hasn 't. Say you had a 10%
| margin before, at a $100 price and $90 cost, for a $10
| profit-per-item. Reduce price 27x and cost 45x, so $3.7
| price, $2 cost, and $1.7 profit-per-item. 6x less profit
| - not as bad as 27x, but not good if you're OpenAI.
| onlyrealcuzzo wrote:
| > Your relative margin may have doubled, but your
| absolute profit-per-item hasn't.
|
| ChatGPT doesn't have any profits right now.
|
| We have no idea what investors are expecting future
| profits to be.
|
| > Say you had a 10% margin before, at a $100 price and
| $90 cost, for a $10 profit-per-item. Reduce price 27x and
| cost 45x, so $3.7 price, $2 cost, and $1.7 profit-per-
| item. 6x less profit - not as bad as 27x, but not good if
| you're OpenAI.
|
| Now do the same thing but assume you have 10x more
| subscribers because the prices are ~27x lower.
|
| You end up with almost 2x more total profit.
|
| Just take ChatGPT's ~$200 subscription. Hardly anyone is
| going to pay ~$200 a month. Reduce that by 27x - and
| you're at $7.5 per month. Maybe 10% of people on the
| planet will pay that.
| ceejayoz wrote:
| > Now do the same thing but assume you have 10x more
| subscribers because the prices are ~27x lower.
|
| You're in various spots of this thread pushing the idea
| that their 1B MAUs make them unassailable. How are they
| gonna get to 10B in a world with less than that total
| people?
|
| > Just take ChatGPT's ~$200 subscription. Hardly anyone
| is going to pay ~$200 a month. Reduce that by 27x - and
| you're at $7.5 per month. Maybe 10% of people on the
| planet will pay that.
|
| They can't even make money at the $200 price point,
| though. https://x.com/sama/status/1876104315296968813
| JumpCrisscross wrote:
| > _that would mean a 27x lower valuation_
|
| Not directly. The 27x is about costs. What it means is
| some order of magnitude of more competition. _That_
| reduces natural market share, price leverage and thus
| future profits.
| gtirloni wrote:
| Other search engines don't have a gigantic advertising
| budget or a dominant browser pounding on users' heads to
| use them.
| rurp wrote:
| Google spends immense amounts of resources every year to
| ensure that their search is almost always the default
| option. Defaults are extremely powerful in consumer tech.
| digitalPhonix wrote:
| > Google Search doesn't have a network effect. Everyone
| on HN has been saying Google Search is complete garbage
| for a decade. It still has the same market share
| (roughly) as it did a decade ago.
|
| It absolutely does. People use Google for search ->
| Websites optimise for Google -> People get "better"
| results when searching with Google.
|
| The fact that it's market share is sticky and not
| responding quickly to change in quality is sort of
| indicative of the network effect.
| jasonjmcghee wrote:
| 1 billion MAU? What's the source on that? Very difficult
| to believe.
| ceejayoz wrote:
| It probably counts pretty much anyone on a newer
| iPhone/Mac (https://support.apple.com/en-
| au/guide/iphone/iph00fd3c8c2/io...) and Windows/Bing.
| Plus all the smaller integrations out there. All of which
| can be migrated to a new LLM vendor... pretty quickly.
|
| I wonder what the _direct_ user counts are.
| cogman10 wrote:
| Free that runs locally on consumer hardware sounds
| substantially better.
| spinlock_ wrote:
| I don't agree. You don't have moat if you are offering
| the same quality for a higher price.
| flavius29663 wrote:
| just for the chatbot, it's trivial to switch, create a
| new account and start asking questions from deepseek
| instead. There is nothing holding the users in chatgpt.
| ceejayoz wrote:
| And the bigger risk is the big companies making deals -
| like Apple including ChatGPT access in iOS - canceling
| those to do it on-device or in-house.
|
| 1B MAUs doesn't look great if half of them come from one
| source that can easily change to a competitor.
| JumpCrisscross wrote:
| > _You don 't need a moat when you're in first place_
|
| There are different moats [1]. You're describing
| incumbency, an intangible moat. It's nice, but it's
| fickle. Particularly with something with low switching
| costs.
|
| OpenAI could argue, before, that it had a natural
| monopoly. More people use OpenAI so it gets more revenue
| and more data which lets it raise more capital to train
| these expensive models. That may not be true, which means
| it only has that first, shallow moat. It's Nike. Not
| Google.
|
| [1] https://en.m.wikipedia.org/wiki/Economic_moat
| onlyrealcuzzo wrote:
| > There are different moats [1]. You're describing
| incumbency, an intangible moat. It's nice, but it's
| fickle. Particularly with something with low switching
| costs.
|
| Google has a low switching cost, and hardly anyone
| switches.
|
| ChatGPT is quite similar to Google in this way.
| JumpCrisscross wrote:
| > _Google has a low switching cost, and hardly anyone
| switches_
|
| Google has massive network effects on its ad business and
| a natural monopoly on its search index. Crawling the web
| is expensive. It's why Kagi has to pay Google (versus
| being able to pay them once and then stop).
| scarface_74 wrote:
| Thought experiment: if tomorrow Apple changed the default
| search engine from Google to ChatGPT for iOS, how fast
| would Google's dominance drop?
|
| iOS has 70% market share in the US
| kgwgk wrote:
| https://archive.is/c6cn9
|
| << The thing I noticed right away when Claude came out is
| how little lock-in ChatGPT had established. This was very
| different to my experience when I first ran a search on
| Google, sometime in the year 2000. After the first time I
| used Google, I literally never used another search engine
| again; it was just light years ahead of its competitors
| in terms of the quality of its results, and the clarity
| of its presentation. This week I added a third chatbot to
| the mix: DeepSeek >>
|
| Follow up:
| https://x.com/TheStalwart/status/1884606421225848889
| mohsen1 wrote:
| My guess is that OS vendors are the real winners in the
| long run. If Siri/Goolge can access my stuff and core of
| LLMs is this replicable then I don't see anyone
| downloading any apps for their typical AI usage.
| Specially that users have to go out of their way to allow
| a 3rd party to access all their data.
|
| This is why OpenAI is so deep in the product development
| phase right now. They have to become the OS to be
| successful but I don't see that happening
| Sateeshm wrote:
| There is no network effect (amazon, instagram, etc.) not
| an enterprise vendor lock-in (Microsoft Office/AD, Apple
| Appstore, etc.) In fact, it's quite the opposite, the way
| these companies deliver ouput is damn near identical.
| Switching between them is pretty painless.
| meiraleal wrote:
| oh well, I switched yesterday from a paid plan to a free
| one and I'm quite happy with the quality improvement.
| whatshisface wrote:
| I wish more people had understood that spending a lot of
| money processing publicly available commodities with
| techniques available in the published literature is the
| business model of a steel mill.
| JumpCrisscross wrote:
| > _is the business model of a steel mill_
|
| It's the business of commodities. The magic is in tiny
| incremental improvements and distribution. DeepSeek
| forces us to question if AI--possibly intelligence--is a
| commodity.
| sebzim4500 wrote:
| Surely that would be amazing for NVDA? If the only 'hard'
| part of making AI is making/buying/smuggling the hardware
| then nvidia should expect to capture most of the value.
| ceejayoz wrote:
| DeepSeek revealed it's not as hard as previously thought;
| a much smaller number of less sophisticated chips was
| sufficient.
| JumpCrisscross wrote:
| > _that would be amazing for NVDA?_
|
| It's good for Nvidia. It's not as good as it was before.
| (Assuming DeepSeek's claims are replicable.)
| JoshTko wrote:
| No. Before Deepseek R1, Nvidia was charging $100 for a
| $20 shovel in the gold rush. Now, every Fortune 100 can
| build an O1-level model with currently existing (and soon
| to be online) infra. Healthy demand for H100 and
| Blackwell will remain, but paying $100 for a $20 shovel
| is unlikely.
|
| Nvidia will definitely stay profitable for now though, as
| long as Deepseek's breakthroughs are not further improved
| upon. But if others find additional compression gains,
| Nvidia won't recapture its old premium. Its stock hinged
| on 80% margins and 75% annual growth, Deepseek broke that
| premise.
| wongarsu wrote:
| There still isn't a serious alternative for chips for AI
| training. Until competition catches up or models become
| so efficient they can be trained on gaming cards Nvidia
| will still be able to command the same margins.
|
| Growth might take a short-term dip, but may well be
| picked up by induced demand. Being able to train your own
| models "cheaply" will cause a lot more companies and
| departments want to train their own models on their own
| data, and cause them to retrain more frequently.
|
| The time of being able to sell H100 clusters for
| inference might be coming to an end though.
| throwaway48476 wrote:
| NVDA is too invested in training and underinvested in
| edge inference.
| camdenreslink wrote:
| I'm not sure I'd call what LLMs do intelligence. Not yet
| anyway...
| JumpCrisscross wrote:
| > _not sure I'd call what LLMs do intelligence_
|
| No, but it's good enough to replace some office jobs.
| Which forces us to ask, to what degree is intelligence--
| unique intelligence--required for useful production? (We
| can ask the same about physical strength.)
| FridgeSeal wrote:
| I find it interesting that so much discussion about
| "LLM's can do some of our work" is centred around "are
| they intelligence" and not what I see as the precursor
| question of "are we doing a lot of bullshit work?"
|
| My partner is in law, along with several friends and the
| amount of completely _useless_ work and ceremony they're
| forced to do is insane. It's a literal waste of their
| talent and time. We could probably net most of the
| claimed AI gains by taking a serious look at pointless
| workloads and come out ahead due to not needing the
| energy and capital expenditure.
| cjbgkagh wrote:
| But have you heard of Jevons Paradon...... /s
|
| OMG, it seems tech has been invaded by baaing crypto bros
| scotty79 wrote:
| Maybe they even suppressed algorithmic improvements in
| their company to preserve moat. Something akin to Kodak
| suppressing internal research on digital cameras because
| they were world leading company that produced photo film.
| wturner wrote:
| Capitalism - a system where rational actors make informed
| decisions.
| blantonl wrote:
| Nah, those costs were for their doomsday bunkers and crypto
| purchases, and maybe a house or 3
| askl wrote:
| They could pivot to being a wrapper around DeepSeek. That
| would also save a lot of R&D costs.
| mirzap wrote:
| Training costs are not the same as inference costs. DeepSeek
| (or anyone hosting DS largest model) will still need a lot of
| money and a bunch of GPU clusters to serve the customers.
| btbuildem wrote:
| > pass that along to end users
|
| I don't think that's at all likely in the current economic
| system
| aprilthird2021 wrote:
| Banning it will not bring back the value
| duxup wrote:
| When it comes to the executive branch's role in banning
| something. I'm not convinced they're even honest about it /
| what the context even is.
|
| Trump wanted to ban Tiktok before... and then simply chose not
| to / forgot about it.
|
| Next round congress acted, and Trump delayed it and has said
| that he is interested in his friends buying it.
|
| Is there really a competitive plan here or is it just fishing
| for payouts / grifting for allies?
|
| The context is always about competition, but I'm not even sure
| that's their plan.
| the_sleaze_ wrote:
| "easy to reinvent" often comes after "hard to invent"
| mromanuk wrote:
| At Microsoft's size, they don't care; they just buy out
| others.
| IAmGraydon wrote:
| Yeah they're setting this up to ban it. Crazy that they think
| this kind of approach will work in any way. Banning H100s
| didn't work, and actually pushed them to innovate. Now someone
| has found a more efficient way to train a model and they decide
| the best way forward is for the US not to benefit from access
| to it? This is clear evidence of collusion between OpenAI and
| the US Government to disadvantage competitors. Beyond that, it
| will never work. If they need to be reminded of just how little
| power they have to control the distribution of open source
| models, I think we would all be happy to enlighten them.
| Buttons840 wrote:
| I saw a some Europeans hoping that the US would ban DeepSeek,
| because then there would be less traffic interfering with their
| own DeepSeek queries.
|
| The US can ban all they want, but if the rest of the world
| starts preferring Chinese social media, Chinese AI, and Chinese
| websites in general, the US is going to lose one of its crown
| jewels.
|
| The way the US behaves is a problem and makes a lot of people
| prefer alternatives just for the sake of avoiding the US, which
| is why it's important that the US get along with other nations,
| but--well, about that...
| nozzlegear wrote:
| Agreed, you've highlighted one of the key problems with
| protectionism and nativism. Banning competition just weakens
| America's global influence, it doesn't make it stronger.
| paxys wrote:
| Plus technology cycles move so quickly that you won't have
| to wait a generation or two to see the effects of this
| isolationism.
| sailfast wrote:
| You know this. I know this. But the President of the United
| States does not know this.
| darkwizard42 wrote:
| This statement doesn't seem to hold true. China has banned
| nearly all US tech companies and social products. It has
| not decreased the influence of China's influence (which has
| been through manufacturing/retail influence and tech
| influence).
|
| I don't think your statement holds with current behavior.
| nozzlegear wrote:
| But China has never been a global leader in tech or
| social media. They undoubtedly have influence in these
| areas, but they've never dominated them like the US has.
| Banning foreign competition in a field where you
| _already_ dominate, like tech and AI, has different
| consequences than banning it where you 're playing catch
| up.
| philistine wrote:
| TikTok has been the darling of the world for years at
| this point. They're a global leader.
| nozzlegear wrote:
| Pretext my statement with "historically, until the last 5
| years or so" and it still stands. TikTok is definitely
| influential, there's no arguing that.
| jononor wrote:
| What is your definition of "tech"? A very large amount of
| the electronics products in the world are made in China
| (specifically in/around Szhenzen and the wider Guangdong
| province). Both consumer goods and industrial goods. From
| the cheapest stuff to the most advanced and everything in
| between. They provide the manufacturing for brands fron
| all over the world, including goods "from the west". The
| amount of economy that depends entirely on this low-cost,
| high-quality manufacturing is insanely large - both
| directly in electronics goods but also as part of many
| other industries because you need electronics to build
| anything else.
| kergonath wrote:
| > China has banned nearly all US tech companies and
| social products. It has not decreased the influence of
| China's influence
|
| Being hostile does not bring you friends. Sure, various
| countries can have reasons to suck it up anyway (e.g.
| because of sanctions, or because China makes an offer too
| good to pass, although even that comes with strings
| attached). But in the long run you just create clients or
| satellites who will escape at the first occasion.
|
| The American foreign policy around the middle of the 20th
| century relied very effectively on soft power, which is
| something you can leverage to get much more out of your
| investments than their pure monetary value. It is not
| required in order to gain influence, but it is a force
| multiplier.
| philistine wrote:
| Then how can you explain that China's hostility towards
| Western tech companies being present inside their own
| country has not created what you're describing?
|
| Is hostility a bad idea only for America? Sure hope not.
| freeone3000 wrote:
| America is reliant on purchasing cheap goods from
| elsewhere and selling expensive technology. If it's
| hostile toward the suppliers of cheap goods or the buyers
| of expensive technology, well, what purpose does it have
| on the global scale?
| kergonath wrote:
| I am saying that they could have got much more,
| particularly considering the spectacular mistakes western
| countries kept making for the last ~2 decades.
| nozzlegear wrote:
| > Is hostility a bad idea only for America? Sure hope
| not.
|
| I think protectionism is long-term bad for every country,
| but it's especially and uniquely bad for the biggest
| economy in the world who has net benefitted the most from
| free trade and competition. There's no denying that China
| is influential - the argument is that they could've been
| (and still can be) so much more influential by embracing
| western tech instead of walling themselves off.
| karel-3d wrote:
| EU will ban DeepSeek sooner because of (lack of) GDPR
| compliance
| whatevaa wrote:
| That's ok, they provide the models. They can be run at
| Europe, given you have capable hardware.
| wongarsu wrote:
| And the path to pleasing the EU would be straightforward:
| make a EU subsidiary, have it host or rent GPUs and
| servers in Europe, make sure personally-identifiable data
| is handled in accordance with GDPR and doesn't leave that
| subsidiary, make sure EU customers make their accounts
| with and are served by that subsidiary.
|
| Meanwhile, to please the US they would probably have to
| move the entire company to the US. And even that may not
| be enough
| acheong08 wrote:
| Banning the site would be fine. The model itself will still
| be available from a variety of providers as well as
| locally. The US is more likely to ban the model itself on
| the basis of national security
| Buttons840 wrote:
| Does EU _block_ websites that don 't comply with their
| laws?
|
| If DeepSeek becomes popular in America I predict it will be
| blocked, national firewall style. Will EU do the same?
| surgical_fire wrote:
| Generally no. For all people complaining about EU
| regulations, the regulators typically opt to fine
| companies into compliance.
| tensor wrote:
| I've recently cancelled my Github Copilot subscription and
| now use Mistral. When the US starts threatening allies with
| tariffs or invasion, using US services becomes a major
| business risk.
| mongol wrote:
| Not only a business risk. It also becomes a moral
| imperative to avoid if you can. Don't support bullies, is
| my motto. It can be hard to completely avoid, but it is
| important to try.
| dkjaudyeqooe wrote:
| Not sure I agree with your premise, but what exactly are they
| going to ban?
|
| They can stop DeepSeek doing various things commercially I
| guess, but stopping Americans using their ideas is simply
| impossible and stopping use of their source or weights would be
| (likely successfully) challenged under the first amendment.
|
| There is no law against simply destroying trillions of dollars
| of shareholder value.
| thedevilslawyer wrote:
| Heh, not yet..
| dfxm12 wrote:
| There's a lot of egg on people's faces now. DeepSeek shows
| there's nothing special about America or its economic system
| that breeds innovation. DeepSeek shows how these tech oligarchs
| greatly overplayed their hand and along with the president,
| bamboozled the taxpayer to enrich each other. I just hope the
| voters remember this in 2 years, 4 years and beyond.
| dtquad wrote:
| >DeepSeek shows how these tech oligarchs greatly overplayed
| their hand and along with the president, bamboozled the
| taxpayer to enrich each other.
|
| How much taxpayer money has gone to OpenAI and Anthropic?
| They are the two big sinners in closed AI.
| iforgot22 wrote:
| If the allegations are true, the special thing about OpenAI
| is that it didn't have to be trained off DeepSeek. But either
| way, you maybe don't want to invest billions in something if
| someone else will be able to copy it for less.
| dtquad wrote:
| >It's reasonably likely that a lot of people linked to the
| federal government want to ban DeepSeek.
|
| It took them years and years to move forward with the ban ok
| Tiktok and it still hasn't been banned yet. There is no way
| they are going to ban some MIT-licensed weights.
|
| >"they destroyed $1T of shareholder value."
|
| The market has largely recovered.
| tokioyoyo wrote:
| There is big American money invested in TikTok. That doesn't
| seem to be the case for DeepSeek.
| leesec wrote:
| It was so easy it cost hundreds of millions of dollars, only
| one company has done it and they had to lie about it
| nullbyte wrote:
| I think the real concern from the govt's perspective is data
| privacy, since all the chat messages are stored on Chinese
| servers
| jasoneckert wrote:
| What I find the most comical about this is that the whole
| situation could be loosely summarized as "OpenAI is losing its
| job to AI."
| zbshqoa wrote:
| Realistically that's the actual headline. Only another AI can
| replace AI, pretty much like LLMs / Transformers have replaced
| "old" AI models in certain task (NLP, Sentiment Analysis,
| Translation etc) and research is in progress for other tasks as
| well performed by traditional models (personalization,
| forecasting, anomaly detection etc).
|
| If there's a better AI, old AI will lose the job first.
| troyvit wrote:
| > NLP, Sentiment Analysis, Translation etc
|
| As somebody who got to work adjacent to some of these things
| for a long time, I've been wondering about this. Are LLMs and
| transformers actually better than these "old" models or is it
| more of an 80/20 thing where for a lot less work (on
| developers' behalf) LLMs can get 80% of the efficacy of these
| old models?
|
| I ask because I worked for a company that had a related
| content engine back in 2008. It was a simple vector database
| with some bells and whistles. It didn't need a ton of
| compute, and GPUs certainly weren't what they are today, but
| it was pretty fast and worked pretty well too. Now it seems
| like you can get the same thing with a simple query but it
| takes a lot more coal to make it go. Is it better?
| ang_cire wrote:
| Yep, it's an 80/20 thing. Versatility over quality.
| Keyframe wrote:
| also, China doing in IP what it's better at and way more
| experienced than USA - stealing.
| rchaud wrote:
| This kind of blithe commentary is 20 years out of date and
| reminiscent of 1970s criticisms of the Japanese car industry.
| Buttons840 wrote:
| I'm reminded of an Adam Savage video. He ordered an unusual
| vise from China, and he praised their culture where someone
| said "I want to build this strange vise that wont be super
| popular", and the boss said "cool, go do it". They built a
| thing that we would not build in America.
|
| https://youtu.be/NUhrF0xkhhc?si=1WHWYZrhRmfOYO_y&t=1150
| (it's about 2 minutes)
| pphysch wrote:
| The small biz scene in unfree communist China is
| ironically, astronomically better than here in US, where
| decades of regulatory capture and misleadership have made
| it difficult and extremely expensive to get off the
| ground while being protected by the law.
| Keyframe wrote:
| gentle stroll through the aliexpress alleyway tells
| otherwise.
| t43562 wrote:
| Who says America is less good at it? Hasn't the US nicked a
| lot of other people's ideas at some point or other?
| mattgreenrocks wrote:
| OpenAI should be excited that it has been freed of the tedious
| tasks of building AI and now they can focus on higher level and
| more creative things.
| Sateeshm wrote:
| > focus on higher level and more creative things.
|
| But that's what OpenAI's costumers were supposed to do.
| jusonchan81 wrote:
| It's sarcasm.
| rooroobooragool wrote:
| I think Sateeshm was also applying a generous layer of
| sarcasm.
| pphysch wrote:
| OpenAI should be, but OpenAI died a while ago
| JoshTko wrote:
| I wish I could upvote this twice
| munchler wrote:
| Soon you'll be freed of the tedious task of upvoting at
| all.
| blantonl wrote:
| Otherwise known as a race to the bottom
| bwfan123 wrote:
| ha, the story is filled with ironies.
|
| OpenAIs $200 closed-ai uppended by hedge-funds free side-
| project
|
| Quant geeks outcompete overpaid silicon valley devs etc.
|
| Basically, hubris gets its comeuppance which is a david vs
| goliath biblical archetype which is why this drama grips all of
| us.
| jeffreyq wrote:
| seems ironic that the turns have tabled. "silicon valley
| devs" were the analogous "quant geeks" underdogs that
| unseated the ossified incumbents.
|
| That said, I feel like "quant geeks" aren't quite underdogs
| compared to silicon valley devs. wdyt?
| rooroobooragool wrote:
| This is really the top take in this thread. Why should OpenAI
| be any different than all the others they they've ripped off.
| nikeee wrote:
| More like
|
| "OpenAI is losing its job to open AI."
| semking wrote:
| This is absolutely hilarious! :)
|
| ClosedAI scraped human content without asking and they explained
| why this was acceptable... but when the outputs of their training
| corpus is scraped, it is THEIR dataset and this is NOT
| acceptable!
|
| Oh, the irony! :D
|
| I shared a few screenshots of DeepSeek answering using ChatGPT's
| output in yesterday's article!
|
| https://semking.com/deepseek-china-ai-model-breakthrough-sec...
| marricks wrote:
| Also, DeepSeek is allegedly... better? So saying they just
| copied ClosedAI isn't really sufficient of an answer. Seems to
| be just bluster because the US Govt would probably accept any
| excuse to ban it, see TikTok.
| semking wrote:
| I never said they are just a clone! There's an actual tech
| breakthrough!
|
| Read the two following sections of my blog post:
|
| 1. "Distilled language models"
|
| 2. "DeepSeek: Less supervision"
| beAbU wrote:
| How can they ban something thats open source that you can
| just run on your own hardware?
| Drakim wrote:
| They banned certain branches of math during the cold war,
| it can be done.
| jerry80 wrote:
| Such as?
| shafyy wrote:
| It's not open source. The provide the model and the
| weights, but not the source code and, crucially, the
| training data. As long as LLM makers don't provide the
| training data (and they never will, because then they will
| be admitting to stealing), LLMs are never going to be open
| source.
| sho_hn wrote:
| Thanks for reminding people of this.
|
| Open source means two things in spirit:
|
| (a) You have everything you need to be able to re-create
| something, and at any step of the process change it.
|
| (b) You have broad permissions how to put the result to
| use.
|
| The "open source" models from both Meta so far fail
| either both or one of these checks (Meta's fails both).
| We should resist the dilution of the term open source to
| the point where it means nothing useful.
| jprete wrote:
| I think people are looking for the term "freeware"
| although the connotations don't match.
| sho_hn wrote:
| Agreed, but the "connotations don't match" is mostly
| because the folks who chose to call it open source wanted
| the marketing benefits of doing so. Otherwise it'd match
| pretty well.
| HDThoreaun wrote:
| Open source means the source code is freely available.
| It's in the name.
| idle_zealot wrote:
| The source being available means the code is "source
| available." Open implies more rights.
| KPGv2 wrote:
| At the risk of being called rms, no, that's not what open
| source means. Open source just means you have access to
| the source code. Which you do. Code that is open source
| but restrictively licensed is still open source.
|
| That's why terms like "libre" were born to describe
| certain kinds of software. And that's what you're
| describing.
|
| This is a debate that started, like, twenty years ago or
| something when we started getting big code projects that
| were open source but encumbered by patents so that they
| couldn't be redistributed, but could still be read and
| modified for internal use.
| sho_hn wrote:
| > Open source just means you have access to the source
| code. Which you do.
|
| No, they also fail even that test. Neither Meta nor
| DeepSeek have released the source code of their training
| pipeline or anything like that. There's very little
| literal "source code" in any of these releases at all.
|
| What you _can_ get from them is the model weights, which
| for the purpose of this discussion, is very similar to
| compiler binary executable output you cannot easily
| reverse, which is what open source seeks to address. In
| the case of Meta, this comes with additional usage
| limitations on how you may put them to use.
|
| As a sibling comment said, this is basically "freeware"
| (with asterisks) but has nothing to do with open source,
| either according to RMS or OSI.
|
| > This is a debate that started, like, twenty years ago
|
| For the record, I do appreciate the distinction. This
| isn't meant as an argument from authority at all, but
| I've been an active open source (and free software)
| developer for close to those 20 years, am on the board of
| one of the larger FOSS orgs, and most households have a
| few copies of FOSS code I've written running. It's also
| why I care! :-)
| JumpCrisscross wrote:
| > _they also fail even that test. Neither Meta nor
| DeepSeek have released the source code of the_
|
| This debate is over and makes the open source community
| look silly. Open model and weights is, practically
| speaking, open source for LLMs.
|
| I have tremendous respect for FOSS and those who build
| and maintain it. But arguing for open training data means
| only toy models can practically exist. As a result, the
| practical definition will prevail. And if the only people
| putting forward a practical definition are Meta _et al_ ,
| this is what you get: source available.
| sho_hn wrote:
| I'm not arguing for open training data BTW, and the
| problem is exactly this sort of myopic focus on the
| concerns of the AI community and the benefits of open-
| washing marketing.
|
| Completely, fully breaking the meaning of the term "open
| source" is causing collateral damage _outside_ the AI
| topic, that 's where it really hurts. The open source
| principle is still useful and necessary, and we need
| words to communicate about it and raise correct
| expectations and apply correct standards. As a dev you
| very likely don't want to live in a tech environment
| where we regress on this.
|
| It's not "source available" either. There's no _source_.
| It 's freeware.
|
| "I can download it and run it" isn't open source.
|
| I'm actually not too worried that people won't eventually
| re-discover the same needs that open source originally
| discovered, but it's pretty lame if we lose a whole bunch
| of time and effort to re-learn some lessons yet again.
| JumpCrisscross wrote:
| > _it 's pretty lame if we lose a whole bunch of time and
| effort to re-learn some lessons yet again_
|
| We need to relearn because we need a different definition
| for LLMs. One that works in practice, not just at the
| peripheries.
|
| Maybe we can have FOSS LLMs vs open-source ones, like we
| do with software licenses. The former refers to the
| hardcore definition. The latter the practical (and widely
| used) one.
| sho_hn wrote:
| Sure, I don't disagree. I fully understand the open-
| weights folks looking for a word to communicate their
| approach and its benefits, and I support them in doing
| so. It's just a shame they picked this one in - and
| that's giving folks a lot of benefit of the doubt - a
| snap judgement.
|
| > Maybe we can have FOSS LLMs vs open-source ones, like
| we do with software licenses.
|
| Why not just call them freeware LLMs, which would be much
| more accurate?
|
| There's nothing "hardcore" or "zealot" about not calling
| these open source LLMs because there's just ...
| absolutely nothing there that you call open source in any
| way. We don't call _any_ other freeware "open source"
| for being a free download with a limited use license.
|
| This is just "we chose a word to communicate we are
| different from the other guys". In games, they chose to
| call it "free to play (f2p)" when addressing a similar
| issue (but it's also not a great fit since f2p games
| usually have a server dependency).
| JumpCrisscross wrote:
| > _Why not just call them freeware LLMs, which would be
| much more accurate?_
|
| Most of the public is unfamiliar with the term. And with
| some of the FOSS community arguing for open training
| data, it was easy to overrule them and take the term.
| sho_hn wrote:
| Most of the public is also unfamiliar with the term open
| source, and I'm not sure they did themselves any favors
| by picking one that invites far more questions and needs
| for explanation. In that sense, it may have accomplished
| little but its harmful effects.
|
| I get your overall take is "this is just how things go in
| language", but you can escalate that non-caring
| perspective all the way to entropy and the heat death of
| the universe, and I guess I prefer being an element that
| creates some structure in things, however fleeting.
| JumpCrisscross wrote:
| > _Most of the public is also unfamiliar with the term
| open source_
|
| I'd argue otherwise. (Familiar with, not know.)
| Particularly in policy circles.
|
| > _picking one that invites far more questions and needs
| for explanation_
|
| There wasn't ever a debate. And now, not even the OSI
| demands training data. (It couldn't. It, too, would be
| ignored.)
| Flimm wrote:
| The only practical and widely used definition of open
| source is the one known as the Open Source Definition
| published by the OSI.
|
| The set of free/libre licenses (as defined by the FSF) is
| almost identical to the set of open sources licenses (as
| defined by the OSI).
|
| The debate within FOSS communities has been between
| copyleft licenses like the GPL, and permissive licenses
| like the MIT licence. Both copyleft and permissive
| licenses are considered free/libre by the FSF, and both
| of them are considered open source by the OSI.
| nuancebydefault wrote:
| The weights, which are part of the source, are open. Now
| you are arguing it not being open source because they
| don't provide the source for that part of the source. If
| you follow that reasoning you can ad infinitum claim the
| absence of sources since every source originates from
| something.
| jefftk wrote:
| _> Open source just means you have access to the source
| code._
|
| That's https://en.wikipedia.org/wiki/Source-
| available_software , not 'open source'. The latter was
| specifically coined [1] as a way to talk about "free
| software" (with its freedom connotations) without the
| price connotations:
|
| _The argument was as follows: those new to the term
| "free software" assume it is referring to the price.
| Oldtimers must then launch into an explanation, usually
| given as follows: "We mean free as in freedom, not free
| as in beer." At this point, a discussion on software has
| turned into one about the price of an alcoholic beverage.
| The problem was not that explaining the meaning is
| impossible--the problem was that the name for an
| important idea should not be so confusing to newcomers. A
| clearer term was needed. No political issues were raised
| regarding the free software term; the issue was its lack
| of clarity to those new to the concept._
|
| [1] https://opensource.com/article/18/2/coining-term-
| open-source...
| HDThoreaun wrote:
| You dont get to redefine what "open" means.
| jefftk wrote:
| It's common for terms to have a more specific meaning
| when combined with other terms. "Open source" has had a
| specific meaning now for decades, which goes beyond "you
| can see the source" to, among other things, "you're
| allowed to it without restriction".
| RobotToaster wrote:
| So Swedish meatballs are any ball of meat made in Sweden?
|
| And French fries are anything that was fried in France?
| davidcbc wrote:
| Tell that to Sam Altman
| esafak wrote:
| He did not succeed, did he?
| dTal wrote:
| I don't know why you've been downvoted. This is a 100%
| correct history. "Open source" was specifically coined as
| a synonym to "free software", and has always been used
| that way.
| beAbU wrote:
| Thanks, I was not aware of this distinction.
|
| But I think my argument still stands though? Users can
| run Deepseek locally, so unless the US Gov't wants to
| reach for book burning levels or idiocy, there is not
| really a feasible way to ban the American public of
| running DeepSeek, no?
| shafyy wrote:
| Yes, your argument still stands. But I think it's
| important to stand firm that the term "open source" is
| not a good label for what these "freeware" LLMs are.
| beAbU wrote:
| Fair point, agreed.
| coliveira wrote:
| People say this, but when it comes to AI models, the
| training data is not owned by these companies/groups, so
| it cannot be "open sourced" in any sense. And the
| training code is basically accessing that training data
| that cannot be open sourced, therefore it also cannot be
| shared. So the full open source model you wish to have
| can only provide subpar results.
| sheepdestroyer wrote:
| They could easily list the data used though. These
| datasets are mostly known and floating around. When they
| are constructed, instructions for replication could be
| provided too
| coliveira wrote:
| They could, but even if they give this list the
| detractors will still say it is not open source.
| rvnx wrote:
| yes and as a bonus they may get sued, which in the long-
| term, makes free / offline models to not be viable
|
| It would be so much better if all models were trained
| with LibGen.
| Timon3 wrote:
| Isn't this the same situation that any codebase faces
| when one thinks about open sourcing it? I can't legally
| open source the code I don't own.
| fabianhjr wrote:
| There are illegal numbers in the USA land of the "free".
|
| https://en.wikipedia.org/wiki/Illegal_number
|
| > An AACS encryption key (09 F9 11 02 9D 74 E3 5B D8 41 56
| C5 63 56 88 C0) that came to prominence in May 2007 is an
| example of a number claimed to be a secret, and whose
| publication or inappropriate possession is claimed to be
| illegal in the United States.
| JumpCrisscross wrote:
| > _illegal numbers in the USA land of the "free"_
|
| This is a silly take for anyone in tech. Any binary
| sequence is a number. Any information can be, for
| practical purposes, rendered in binary [1].
|
| Getting worked up about restrictions on numbers works as
| a meme, for the masses, because it sounds silly, but is
| tantamount to technically arguing against privacy,
| confidentiality, the concept of national secrets, IP as a
| whole, _et cetera_.
|
| [1] https://en.m.wikipedia.org/wiki/Shannon%27s_source_co
| ding_th...
| sheepdestroyer wrote:
| All those things are not self-evident and thus debatable
| JumpCrisscross wrote:
| > _not self-evident and thus debatable_
|
| Totally agree. But prompting debate or even further
| thought isn't the point of the meme.
| sheepdestroyer wrote:
| I'd argue that, as satire, it's the main point ;)
| JumpCrisscross wrote:
| > _as satire, it 's the main point_
|
| There is thought-stopping satire and thought-provoking
| satire. Much of it depends on the context. I'm not
| getting the latter from a "USA land of the 'free'"
| comment.
| fabianhjr wrote:
| Good thing that is part of the wikipedia entry:
|
| > Any piece of digital information is representable as a
| number; consequently, if communicating a specific set of
| information is illegal in some way, then the number may
| be illegal as well.
| bloopernova wrote:
| That takes me back! Fark.com would delete any comment
| that contained random hexadecimal.
| KPGv2 wrote:
| It was the beginning of the end for Digg, too, IIRC.
| Started a lot of people leaving for Reddit, right?
| bloopernova wrote:
| I think so; I joined Reddit when it was in tech news as
| people left Digg after the big redesign. I'm not sure
| when the exodus started. I left Fark over the hd-dvd
| mess.
| KPGv2 wrote:
| > whose publication or inappropriate possession is
| claimed to be illegal in the United States.
|
| That's not the same thing as a number being illegal at
| all. Here, watch this:
|
| > I claim breathing is illegal in the United States
|
| There, now breathing is claimed to be illegal in the
| United States.
| I-M-S wrote:
| In both cases, legality depends entirely on
| repercussions, i.e. if there's someone to enforce the
| ban. I suspect that in the "illegal numbers" case there
| might be.
| vluft wrote:
| man that's very concerning for wikipedia who is
| publishing it right there on the page linked above.
| dylan604 wrote:
| Only concerning if they are a US based company hosting
| their data in US data centers. oops
| suraci wrote:
| > is collecting rain water illegal?
|
| > It depends on where you live. In many places,
| collecting rainwater is completely legal and even
| encouraged, but some regions have regulations or
| restrictions.
|
| United States: Most states allow rainwater collection,
| but some have restrictions on how much you can collect or
| how it can be used. For example, Colorado has limits on
| the amount of rainwater homeowners can store. Australia:
| Generally legal and encouraged, with many homes using
| rainwater tanks. UK & Canada: Legal with few
| restrictions. India & Many Other Countries: Often
| encouraged due to water scarcity.
| superkuh wrote:
| There was an executive order passed by the previous
| administration that make using anything with more than 10
| billion parameters illegal and punishable by government
| force if done without authorization. Of course like most
| government regulations (even though this is not a
| regulation, it is an executive action) the point is not to
| stop the behavior but instead to create a system where
| everyone breaks the regulation constantly so that if anyone
| rocks the boat they can be indicted/charged and dealt with.
|
| https://www.federalregister.gov/documents/2023/11/01/2023-2
| 4...
|
| >(k) The term "dual-use foundation model" means an AI model
| that is trained on broad data; generally uses self-
| supervision; contains at least tens of billions of
| parameters; is applicable across a wide range of contexts;
| and that exhibits, or could be easily modified to exhibit,
| high levels of performance at tasks that pose a serious
| risk to security, national economic security, national
| public health or safety, or any combination of those
| matters, such as by: ...
| ceejayoz wrote:
| That order does not "make using anything with more than
| 10 billion parameters illegal and punishable by
| government force if done without authorization".
|
| It orders the Secretary of Commerce to "solicit input
| from the private sector, academia, civil society, and
| other stakeholders through a public consultation process
| on potential risks, benefits, other implications, and
| appropriate policy and regulatory approaches related to
| dual-use foundation models for which the model weights
| are widely available".
| derektank wrote:
| Many regulations are created by executive action, without
| input from Congress. The Council on Environmental
| Quality, created by the National Environmental Policy
| Act, has the power to issue it's own regulations.
| Executive Orders can function similarly and the executive
| can order rulemaking bodies to create and remove
| regulations, though there is a judicial effort to
| restrict this kind of policymaking and return regulatory
| power back to Congress.
| bilekas wrote:
| If I'm no wrong wasn't PGP encryption once illegal to
| export ? Not quite the same but the government has a nice
| habit of feeling like they can bad the export of research.
|
| https://en.wikipedia.org/wiki/Export_of_cryptography_from_t
| h...
| beAbU wrote:
| You are right, but I cannot find a single example of such
| a ban actually being effective though. Information wants
| to be free and all that.
| KPGv2 wrote:
| Because you haven't heard of the proprietary software
| that wasn't ever sold internationally because of these
| bans.
|
| Of course Joe Sixpack can throw their code up anywhere,
| but Joe Corporation gets wrecked if they try to sell it.
|
| https://developer.apple.com/documentation/security/comply
| ing...
|
| For example, this is enforced by Apple Store.
| coliveira wrote:
| But that's not the goal, the goal is to protect the
| "intelectual property" only to American companies.
| Countries not in the "friends list" cannot sell products
| in that area without suffering repercussions. That's how
| the US has maintained technological dominance in some
| areas by restricting what other countries can do.
| calgoo wrote:
| If i remember correctly, if you changed the dropdown on
| the webpage to USA you could download the full version of
| PGP anyway.
| Prbeek wrote:
| Add PS1 too. The US government banned sale of PlayStation
| to China because the PLA would apparently have access to
| cutting edge chips for their missiles
| michaelt wrote:
| Make commercial hosting illegal, and make the hardware to
| run it locally cost $6000+
| throwup238 wrote:
| It's not better. In most of my tests (C++/QT code) it just
| runs out of context before it can really do anything. And the
| output is very bad - it mashes together the header and cpp
| file. The reasoning output is fun to look at and occasionally
| useful though.
|
| The max token output is only 8K (32K thinking tokens). O1 is
| 128k, which is far more useful, and it doesn't get stuck like
| R1 does.
|
| The hype around the DeepSeek release is insane and I'm
| starting to really doubt their numbers.
| adamnemecek wrote:
| Thanks for saying this, I thought I was insane, DeepSeek is
| kinda bad. I guess it's impressive all things considered
| but in absolute terms it's not great.
| coliveira wrote:
| I have run personal tests and the results are at least as
| good as I get from OpenAI. Smarter people have also
| reached the same conclusion. Of course you can find
| contrary datapoints, but it doesn't change the big
| picture.
| sebzim4500 wrote:
| To be fair, it's amazing by the standards of six months
| ago. The only models that beat it are o1, the latest
| gemini models and (for some things) sonnet 3.6
| cdelsolar wrote:
| false. It seems better than o1 to me.
| gliptic wrote:
| R1 is trained for a context length of 128K. Where are you
| getting 8K/32K? The model doesn't distinguish "thinking"
| tokens and "output" tokens, so this must be some specific
| API limitations.
| throwup238 wrote:
| _> max_tokens:The maximum length of the final response
| after the CoT output is completed, defaulting to 4K, with
| a maximum of 8K. Note that the CoT output can reach up to
| 32K tokens, and the parameter to control the CoT length
| (reasoning_effort) will be available soon._ [1]
|
| [1] https://api-docs.deepseek.com/guides/reasoning_model
| gliptic wrote:
| So yes, it's a limitation of their own API at the moment,
| not a model limitation.
| throwup238 wrote:
| I'm using it through Kagi which doesn't use Deepseek's
| official API [1]. That limitation from the docs seems to
| be everywhere.
|
| In practice I don't think anyone can economically host
| the whole model plus the kv cache for the entire context
| size of 128k (and I'm skeptical of Deepseek's claims now
| anyway).
|
| Edit: a Kagi team member just said on Discord that
| they'll be increasing max tokens next release
|
| [1] https://help.kagi.com/kagi/ai/llms-privacy.html
| coliveira wrote:
| He's just repeating a lot of disinformation that has been
| released about deepseek in the last few days. People who
| took the time to test DeepSeek models know that the
| results have the same or better quality for coding tasks.
| goosejuice wrote:
| Benchmarks are great to have but individual/org
| experiences on specific codebases still matter
| tremendously.
|
| If an org consistently finds one model performs worse on
| their corpus than another, they aren't going to keep
| using it because it ranks higher in some set of
| benchmarks.
| hn_throwaway_99 wrote:
| But you should also be very wary of these kind of
| anecdotes, and this thread highlights exactly why. That
| commenter says in another comment
| (https://news.ycombinator.com/item?id=42866350) that the
| token limitation that he is complaining about has
| actually nothing to do with DeepSeek's model or their
| API, but is a consequence of an artificial limit that
| _Kagi_ imposes. In other words, his conclusion about
| DeepSeek is completely unwarranted.
| throwup238 wrote:
| It mashed the header and C++ file together, which is
| _egregiously_ bad in the context of QT. This isn't a new
| library, it's been around for almost thirty years. Max
| token sizes have nothing to do with that.
|
| I invite anyone to post a chat transcript showing a
| successful run of R1 against this prompt (and please tell
| me which API/service it came from so I can go use it
| too!)
| marricks wrote:
| > it just runs out of context before it can really do
| anything
|
| I mean, couldn't that be because they're just overwhelmed
| by users at the moment?
|
| > And the output is very bad - it mashes together the
| header and cpp file
|
| That sounds way worse, and like, not something caused by
| being hugged to death though.
|
| Aider recently stated DeepSeek is placed a the top of their
| benchmark though[1] so I'm inclined to believe it isn't
| _all_ hype.
|
| [1] https://aider.chat/docs/llms/deepseek.html
| throwup238 wrote:
| It's definitely not _all_ hype, it really is a
| breakthrough for open source reasoning models. I don't
| mean to diminish their contribution, especially since
| being able to read the reasoning output is a very
| interesting new modality (for lack of a better word) for
| me as a developer.
|
| It's just not as impressive as people make it out to be.
| It might be better than o1 on Python or Javascript thats
| all over the training data, but o1 is overwhelmingly
| better at anything outside the happy path.
| sho_hn wrote:
| Is this a local run of one of the smaller models and/or
| other-models-distilled-with-r1, or are you using their Chat
| interface?
|
| I've also compared o1 and (online-hosted) r1 on Qt/C++
| code, being a KDE Plasma dev, and my impression so far was
| that the output is roughly on par. I've given both models
| some tricky tasks about dark corners of the meta-object
| system in crafting classes etc. and they came up with
| generally the same sort of suggestions and implementations.
|
| I do appreciate that "asking about gotchas with few
| definitive solutions, even if they require some
| perspective" and "rote day-to-day coding ops" are very
| different benchmarks due to how things are represented in
| the training data corpus, though.
| throwup238 wrote:
| I use it through Kagi Assistant which has the proper R1
| model through Together.ai/Fireworks.ai
|
| My standard test is to ask the model to write a
| QSyntaxHighlighter subclass that uses TreeSitter to
| implement syntax highlighting. O1 can do it after a few
| iterations, but R1's output has been a mess. That said,
| its thought process revealed a few issues that I then
| fixed in my canonical implementation.
| sho_hn wrote:
| Thanks for adding detail! My prompts have been very in-
| the-bubble-of-Qt I'd say, less so about mashing together
| Qt and something else, which I agree is a good real-world
| test case.
| throwup238 wrote:
| I haven't had the chance to try it out with R1 yet but if
| you implement a debugger class that screenshots the
| widget/QML element, dumps its metadata like GammaRay, and
| includes the source, you can feed that context into
| Sonnet and o1. They are _scarily_ good at identifying
| bugs and making modifications if you include all that
| context (although you have to be selective with what
| metadata you include. I usually just dump a few things
| like properties, bindings, signals, etc).
| nialv7 wrote:
| Tried this on chat.deepseek.com, it seems to be able to
| do it.
| throwup238 wrote:
| Does it compile? Put the full chat in Pastebin and let's
| check it out!
|
| I haven't used their official chat interface or API for
| privacy reasons.
| CamperBob2 wrote:
| Some have said (for what little _that 's_ worth) that
| Kagi's version is not the real thing, but one of the
| distillations.
| sheepdestroyer wrote:
| There are R1 providers on openrouter with bigger
| input/output token limitations than what DeepSeek's API
| access currently offers.
|
| For instance Fireworks offers R1 with 164K/164K. They are
| far more expensive than DeepSeek though
| api wrote:
| It's not great at super-complex tasks due to limited
| context, but it's quite a good "junior intern that has
| memorized the Internet." Local deepseek-r1 on my laptop (M1
| w/64GiB RAM) can answer about any question I can throw at
| it... as long as it's not something on China's censored
| list. :)
| azinman2 wrote:
| How are you running r1 on 64mb of ram? I'm guessing
| you're running a distill which is not r1
| mritchie712 wrote:
| openai should pay creators, but:
|
| 1. scraping the internet and making AI out of it
|
| 2. using the AI from #1 to create another AI
|
| are not the same thing.
| zbshqoa wrote:
| Number 2 is already possible with open models. You can do
| distillation using Llama, which could likely be doing #1 to
| build their models (I'm not sure it's the case though)
| bugglebeetle wrote:
| Yeah, #1 is way worse and #2 falls under "turnabout is fair
| play."
| epse wrote:
| #1 destroys peoples willingness to publish and unfairly hogs
| bandwidth / creates costs for small hosters
|
| #2 makes a big corp a bit angry
|
| Indeed not the same thing
| latexr wrote:
| > are not the same thing.
|
| You're right. The second one is far more ethical. Especially
| when stealing from a thief.
|
| Doesn't Sam Altman keep parroting they're developing AI "for
| the good of humanity"? Well then, someone taking their model
| and improving on it, making it open-source, having it consume
| less, and having a cheaper API, should make him delighted.
| Unless he _*gasp*_ was full of shit the whole time. Who could
| have guessed?
| perryizgr8 wrote:
| > Doesn't Sam Altman keep parroting they're developing AI
| "for the good of humanity"?
|
| "I don't want to live in a world where someone else makes
| the world a better place better than we do"
|
| - Gavin Belson
| Palmik wrote:
| I agree, (2) seems much less problematic since the AI outputs
| are not copyrightable and since OpenAI gives up ownership of
| the outputs. [1]
|
| So, if you really really care about ToS, then just never
| enter into a contract with OpenAI. Company A uses OpenAI to
| generate data and posts it on the open Internet. Company B
| scrapes open Internet, including the data from Company A [2].
|
| [1]: Ownership of content. As between you and OpenAI, and to
| the extent permitted by applicable law, you (a) retain your
| ownership rights in Input and (b) own the Output. We hereby
| assign to you all our right, title, and interest, if any, in
| and to Output.
|
| [2]: This is not hypothetical. When ChatGPT got first
| released, several big AI labs accidentally and not so
| accidentally trained on the contents of the ShareGPT website
| (site that was made for sharing ChatGPT outputs). ;)
| sksrbWgbfK wrote:
| > 2. using the AI from #1 to create another AI
|
| 2. scraping the AI from #1 and making AI out of it
| haswell wrote:
| Yes, they are different _actions_.
|
| But arguably these actions share enough characteristics that
| it's reasonable to place them in the same _category_.
| Something like: "products that exist largely /solely because
| of the work of other people". The nonconsensual nature of
| this and the lack of compensation is what people
| understandably take issue with.
|
| There is enough similarity that it evokes specific feelings
| about OpenAI when they suddenly find themselves on the other
| side of the situation.
| Winsaucerer wrote:
| I'm genuinely not sure which one you think is worse (if any).
| (1) seems worse, but your reply suggests to me maybe you
| think (2) is worse.
| meowface wrote:
| Not that poster, but I think both are equally fine.
|
| It's funny if OpenAI were to complain about this, but at
| least on Twitter I don't see that much whining about it
| from OpenAI employees. Sam publicly praised DeepSeek.
|
| I do see some of them spreading the "they're hiding GPUs
| they got through sanction evasion" theory, which is
| disappointing, though.
| jillyboel wrote:
| You're right, (1) is violating the rights of a large portion
| of the population, (2) is violating the rights of one company
| tw1984 wrote:
| #1 is stealing from all average joes ever lived on earth
|
| #2 is taking advantages from closedAI.
|
| they are indeed different
| pilooch wrote:
| Any ML based service with an API is basically a dataset builder
| for more ML. This has been known forever and is actually a
| useful "law" of ML-based systems.
| sho_hn wrote:
| Aye, this should be obvious even to non-technical folks. Much
| has been written about how LLMs regurgitate the data they
| were trained on. So if you're looking for data to train on,
| you can certainly extract it there.
|
| Plus of course for people within the tech bubble, plenty of
| research results on the value of synthetically augmented and
| expanded training data that put the impact past just
| regurgitating source data.
|
| This whole episode is a failure of reporting what to expect
| next and projecting running costs etc. most of all.
| amelius wrote:
| This is why models should be open. Or at least they should
| have a local option.
| stackghost wrote:
| [flagged]
| bloomingkales wrote:
| I personally love this chef's kiss of a flip flop sam did
| here:
|
| https://blog.samaltman.com/trump
|
| https://www.reddit.com/r/YAPms/comments/1i7ry5m/sam_altman_g.
| ..
|
| Only a truly talented piece of shit can be as prolific as
| this.
|
| _" He is irresponsible in the way dictators are."_
|
| Chef's kiss.
|
| Edit:
|
| Kids, don't aspire to be like Altman. We as a community need
| to espouse more values than _tech is gonna tech_.
| JumpCrisscross wrote:
| > _don 't aspire to be like Altman_
|
| And don't aspire to be like those who saw what he is but
| made peace with it in exchange for silver.
| gadders wrote:
| You mean all of the YC management, including PG?
| istjohn wrote:
| Hey now, that's not very curious of you. /s
| buran77 wrote:
| Well, anyone who will flex their spine in every
| (im)possible position as required of them, just to get
| even more money and power.
|
| I could understand that from someone with an empty
| stomach. But so many people doing it when their pockets
| are already overflowing is exactly the kind of rot that
| degrades an entire society.
|
| We're all just seeing the results so much better now that
| they can't even be bothered to pretend they ever more
| than this.
|
| Later edit: The way this submission fell ~400th spots
| after just two hours despite having 1250 points and 550
| comments, had its comments flagged and shuffled around to
| different submissions as soon as they touched too close
| to YC&Co is a good mirror of how today's society works.
| ToucanLoucan wrote:
| It's an addiction. There's no amount of money that will
| be enough, there's no amount of power that will be
| enough. They'll burn the world for another hit, and we
| know that because we've been watching them do it for 50
| years now.
| stackghost wrote:
| Yes.
| RIMR wrote:
| Yes. Especially them.
| toxic wrote:
| Yes.
| marxisttemp wrote:
| Paul Graham now reposts right wing grift media on his
| Twitter profile, he's cooked
| SteveGerencser wrote:
| > don't aspire to be like Altman
|
| Aspire to be like Aaron Schwartz.
| some_furry wrote:
| (Except for the tragic ending, of course.)
| lukan wrote:
| If more would be like him, there might be a happy ending.
| some_furry wrote:
| Agreed.
| ibejoeb wrote:
| AaronSw exfiltrated data without authorization. You can
| argue the morality of that, but I think you could make
| the argument for OpenAI as well. I'm not opining on
| either, just pointing out the marked similarity here.
|
| edit: It appears I'm wrong. Will someone correct me on
| what he did?
| gessha wrote:
| Arguing for the morality of OpenAI is a little bit harder
| given their history and actions in the last few years.
| ibejoeb wrote:
| One argument would be means to an end, with the end being
| the initial advancement of AI.
|
| Again, I'm not offering an opinion on it.
| skeeter2020 wrote:
| This is an argument, but isn't this where your scenario
| diverges completely? OpenAI's "means to an end" is
| further than you state; not initial advancement but the
| control and profit from AI.
| ibejoeb wrote:
| Yes, they intended for control and profit, but it's
| looking like they can't keep it under control and
| ultimately its advancements will be available more
| broadly.
|
| So, the argument goes that despite its intention, OpenAI
| has been one of the largest drivers of innovation in an
| emerging technology.
| ceejayoz wrote:
| > edit: It appears I'm wrong. Will someone correct me on
| what he did?
|
| He didn't do it without authorization.
|
| https://en.wikipedia.org/wiki/Aaron_Swartz
|
| > Visitors to MIT's "open campus" were authorized to
| access JSTOR through its network.
| richardwhiuk wrote:
| He wasn't authorised to access the wiring closet. There
| are many troubling things about the case, but it's fairly
| clear Aaron knew he was doing something he wasn't
| authorised to do.
| ceejayoz wrote:
| > He wasn't authorised to access the wiring closet.
|
| For which MIT can certainly have a) locked the door and
| b) trespassed him, but that's a very different issue than
| having authorization to access JSTOR.
| ibejoeb wrote:
| At that same link is an account of the unlawful activity.
| He was not authorized to access a restricted area, set up
| a sieve on the network, and collect the contents of JSTOR
| for outside distribution.
| ponector wrote:
| Why should kids aspire to be like Aaron if it is not
| rewarded in our society? Comparing with such "kings" as
| Altman or Musk.
| barnabee wrote:
| Why should anyone aspire to do what is rewarded over what
| they believe in and what will satisfy them?
| gessha wrote:
| Not everything virtuous is rewarded monetarily but we
| aspire to be virtuous, no?
| coliveira wrote:
| Modern society has stopped to aspire of being virtuous a
| long time ago. Unfortunately, that's nowadays a minority
| view.
| robotresearcher wrote:
| Aaron was not happy. Neither is Trump, or Musk. I don't
| know if Bernie is happy, or AOC. Obama seems happy.
| Hilary doesn't. Harris seems happy.
|
| Striving for good isn't gonna be fun all the time, but
| when choosing role models I like to factor in how happy
| they seem. I'd like to spend some time happy.
| jncfhnb wrote:
| I think it's fairly crazy that you believe you have an
| authentic view into the happiness levels of these people.
| robotresearcher wrote:
| I used the word 'seem' three times. I think it's pretty
| unremarkable to report a personal impression without any
| claim of special insight.
| skeeter2020 wrote:
| If human beings could be categorized as Happy/Not Happy
| the world would be a very boring place and life not worth
| living.
| ponector wrote:
| Musk looks happy throwing his hand from the heart to the
| sun.
| cratermoon wrote:
| Try to imagine a society where people only did things
| that were rewarded. Could such a society even exist?
| Thought experiment: make a list of all the jobs,
| professions, and vocations that are not rewarded in the
| sense you mean, and imagine they don't exist. What would
| be left?
| gus_massa wrote:
| You mean better pay for teachers? It would be nice.
|
| (Since we are dreaming, can I add sane hours for medical
| doctors (like <= 8 per day)?)
| ponector wrote:
| I don't need to imagine. Teachers almost everywhere
| around the globe have poor salaries. In my country there
| are lower enrolment requirements to universities to
| become a school teacher than almost every other field of
| study. Means the dumbest students are there.
|
| And then later they go to the school to teach our future,
| working with high stress and low salary.
|
| Same with medical school in many countries where
| healthcare is not privatized. Insane hours, huge
| responsibilities and poor pay for doctors and nurses in
| many countries.
|
| Nowadays everyone wants to be an influencer or software
| developer.
| makapuf wrote:
| For teachers, sure. For medical doctors, in USA or
| Europe, I think they are much more paid than sw
| engineers.
| ponector wrote:
| In east EU, like Poland sw engineer makes two-three times
| more than a doctor with much less of effort, education
| and no responsibility.
|
| And nurses - they work at minimal salary in Poland. Even
| in USA if you count hourly rates it will be quite poor
| salary for nurses.
| CalRobert wrote:
| We need them to help us build a better society.
| ponector wrote:
| Looks like we need salesmen much more as we value their
| work more.
| uoaei wrote:
| Because kids' brains are not as poisoned into believing
| the most profitable things to do are the most
| meritorious. Not yet, anyway.
| skeeter2020 wrote:
| Because only one person can be king, but everybody can
| participate and contribute. Also there's too many things
| out side of just being "the best" that decide who gets to
| be king. Often that person is a terrible leader.
| miramba wrote:
| Upvoted not because I agree, but I think it's a valid
| question that shouldn't be greyed out. My kids dream job
| is youtube influencer, I don't like it but can I blame
| them? It's money for nothing and the chicks for free.
| ponector wrote:
| Tragedy of current days. No one wants to be a
| firefighter, astronaut or a doctor. Influencers
| everywhere! Can you blame kids? Do you know firefighters
| who earns million dollars annually?
| reaperman wrote:
| * Swartz
|
| But yes.
| oooyay wrote:
| I've read a lot about Aaron's time at Reddit / Not A Bug.
| I somewhat think his fame exceeds his actual
| accomplishments at times. He was perceived to be very
| hostile to his peers and subordinates.
|
| Kind of a cliche, but aspire to be the best version
| yourself every day. Learn from the successes and failures
| of others, but don't aspire to be anyone else because
| eventually you'll be very disappointed.
| bayindirh wrote:
| The gist is, when you find your biggest flaw, work on it,
| and repeat; you've already gone great distance.
| baudehlo wrote:
| I knew Aaron back in my IRC days. He hung out with us to
| talk about RDF for a good couple of years. We chatted
| almost every day.
|
| He was lovely. And a genius. Maybe he changed, but he was
| a truly nice person.
| oooyay wrote:
| Yeah, definitely not a statement on Aaron himself. More a
| statement on idolizing people. There will always be
| instances where they didn't live up to what people think
| of them as. I think Aaron was fine and a normal human
| being.
| wongarsu wrote:
| That's what happens to martyrs. They become larger than
| life and history remembers an idealized version of them
| ddingus wrote:
| Indeed
| stevenally wrote:
| Don't sell your soul, is all.
|
| But survive. This too will pass.
| scotty79 wrote:
| I especially like how he quoted Napoleon or something
| framing himself as the heart of revolution and Deep Seek as
| a child of the revolution only to get a response from some
| random guy "It's not that deep bro. Just release a better
| model."
|
| https://x.com/hibakod/status/1883189126553596234
| hn_throwaway_99 wrote:
| That is particularly gross, but that really feels like the
| norm among all the tech elite these days - Zuckerberg,
| Bezos, etc. all doing the most laughable flip flops.
|
| The reason the flip flops are so laughable to me is because
| they attempt to couch them in some noble, moralistic
| viewpoint, instead of the obvious reason "We own big
| companies, the government has extreme power to make or
| break these companies, and everyone knows kissing up to
| Trump is what is required to be on his good side."
|
| Profiles in Cowardice, every last one of them.
| jcgrillo wrote:
| Another point of view is that they never flopped or
| flipped. They were fascists the whole time and were just
| lying about it before.
| hn_throwaway_99 wrote:
| I think Tim Sweeney's (CEO of Epic Games) comment was
| spot on:
|
| > After years of pretending to be Democrats, Big Tech
| leaders are now pretending to be Republicans, in hopes of
| currying favor with the new administration. Beware of the
| scummy monopoly campaign to vilify competition law as
| they rip off consumers and crush competitors.
|
| This is exactly what OpenAI is trying to do with these
| allegations.
| stevenAthompson wrote:
| Those men and their companies are responsible for
| hundreds of thousands of jobs and a significant portion
| of the global economy. I'm actually thankful that they
| aren't shooting their mouths off to the new boss like
| spoiled children at their first job. It wouldn't make the
| world better, it would make their companies and the lives
| of those who depend on them, worse.
|
| There is a fine line between cowardice and common sense.
| jcgrillo wrote:
| In what sense is the federal government "the boss" of
| private sector businesses? This isn't an oligarchy yet,
| right? They don't _have_ to behave obsequiously, they
| _are choosing to_. They 're doing it for themselves, not
| for their shareholders or their employees. It's an
| attempt to grab power and _become_ oligarchs because they
| see in this government a gullible mark.
| stevenAthompson wrote:
| > This isn't an oligarchy yet, right?
|
| The richest man in the world has a government office down
| the street from the white house, which the taxpayers are
| funding. He's rumored to sleep there.
|
| What do you think?
| hn_throwaway_99 wrote:
| Puhleeeese. I'm not advocating that these leaders all
| lead protest marches against the new administration. But
| the transparent obsequiousness and Trump ball gargling
| under the guise of some moralistic principles is so
| nauseating. And please spare me the idea that the likes
| of Zuckerberg or Bezos gives a rat's ass about their
| employees.
|
| For a contrast to the Bezos, Zuckerberg and Altman types,
| look at Tim Cook. Sure, Apple paid the 1 million
| inauguration "donation", and Cook was at the
| inauguration, and I'm not arguing he's winning any
| "Profiles in Courage" awards, but he didn't come out with
| lots of tweets claiming how massuh Trump is so wise and
| awesome, Apple didn't do a 180 on their previous
| policies, etc.
| the_optimist wrote:
| You're awfully salty and biased toward Reddit contaminants.
| Don't imagine to speak for a community except people who
| agree with you a priori. Also, don't steal my nternet
| points, taste upon it, redditoes.
| hibikir wrote:
| As a society we might talk about virtue, but the reason we
| put it as a goal in stories is that in the real world, we
| don't reward it. It's not just that corruption wins
| sometimes, but we directly punish those that fight it. The
| mood of the times, if anything, comes from people realizing
| that what we called moral behavior leads to worse outcomes
| for the virtuous.
|
| A community only espouses good values when it punishes bad
| behavior. How do we do this when those misbehaving are very
| rich, and attempting to punish the misbehavior has negative
| consequences on you? There just aren't many available tools
| that don't require significant sacrifices.
| js8 wrote:
| > A community only espouses good values when it punishes
| bad behavior.
|
| This is the "beauty" of the free market ideology (see
| e.g. https://a16z.com/the-techno-optimist-manifesto/ ).
| If all the transactions are voluntary, there is no way to
| punish anyone.
| stevenAthompson wrote:
| > If all the transactions are voluntary, there is no way
| to punish anyone.
|
| This is obviously untrue at face value. See: Cancel
| Culture, Bud Light, and Freedom Fries for examples.
|
| Did you mean something more than what you stated here?
| ViktorRay wrote:
| I don't think your links are evidence of a flip flop.
|
| The first link is from mid-2016. The second link is from
| January 2025.
|
| It is entirely reasonable for someone to genuinely change
| his or her views of a person over the course of 8.5 years.
| That is a substantial length of time in a person's life.
|
| To me a "flip-flop" is when one changes views on something
| in a very short amount of time.
| meowface wrote:
| IMO it probably is and Altman probably still (rightly)
| hates Trump. He's playing politics because he needs to. I
| don't really blame him for it, though his tweet certainly
| did make me wince.
| bloomingkales wrote:
| _" I don't really blame him for it"_
|
| That's the thing though right, that we all created this
| mess together. Like yeah, _why don 't you (and the rest
| of us) blame him?_. We're all pretty warped and it's
| going to take collective rehab.
|
| Super pretentious to quote MLK, but the man had stuff to
| say so here it is (on Inaction):
|
| _" He who passively accepts evil is as much involved in
| it as he who helps to perpetrate it"_
|
| _" The ultimate tragedy is not the oppression and
| cruelty by the bad people but the silence over that by
| the good people"_
| whatshisface wrote:
| It's not pretentious to quote Martin Luther King.
| benatkin wrote:
| It seems he was virtue signaling before. So it would be
| more accurate to blame him for having let himself become
| an ego driven person in the past. Or to put it nicely and
| to add the context of Brian Armstrong of Coinbase, who
| has also been showing public support for Trump, a
| mission-driven person.
| mrandish wrote:
| > It seems he was virtue signaling before.
|
| Yes, the first mistake was a business leader in tech
| taking a _public_ political position. It was popular and
| accepted (if not expected) in the valley in 2016.
|
| Doing that then (and banking the social and reputational
| proceeds) created the problem of dissonance now. If he'd
| just stayed neutral in public in 2016, he could do what
| he's doing now and we could assume he's just being a
| pragmatic business person lobbying the government to
| further his company's interests.
| benatkin wrote:
| I think "progressive" is probably the safest position to
| take. It also works if you want to get involved in a
| different sort of politics later on. David Sacks had no
| problem doing that when he was no longer interested in
| being CEO of a large company.
| mrandish wrote:
| The evidence indicates not taking a position is the
| optimal position.
|
| I have a lot of respect for CEOs who just focus on being
| a good CEO. It's a hard enough job as is. I don't care
| about or want to know some CEO's personal position on
| politics, religion or sports teams. It's all a
| distraction from the job at hand. Same goes for actors,
| athletes and singers. They aren't _qualified_ to have an
| opinion any more relevant than anyone else 's, except on
| acting, athletics, singing - or CEO-ing.
|
| Sadly, my perspective is in the minority. Which is why I
| think so many public figures keep making this mistake.
| The media, pundits and social sphere _need_ them to keep
| making this mistake.
| benatkin wrote:
| I guess I think they should study what a neutral position
| looks like, and avoid going beyond it as best as they
| can. I had in mind a "progressive" who avoids any hot
| button issues. Someone with a high profile will be asked
| about politics from time to time. I think Brian Chesky is
| a good example of acting like a progressive in a way that
| stays low profile, but maybe he doesn't really act like
| one. https://www.businessinsider.com/brian-chesky-airbnb-
| new-bree...
|
| Also it helps to have sincere political views. GitHub's
| CEO at the time of #DropICE was too cynical and his image
| suffered because of it.
| mrandish wrote:
| > study what a neutral position looks like
|
| There are no neutral positions in today's political
| landscape. I'm not stating my opinion here, this is
| according to most political positions on the spectrum.
| You suggested "Progressive" (but without hot button
| issues) as a way of signaling a neutral position. That
| may be true in parts of the valley tech sphere but it
| certainly doesn't hold in the rest of the U.S.
| "Progressive" is usually defined being to the left of
| "Liberal", so it's hardly neutral. Over half of U.S.
| voters cast their ballot for the Republican candidate.
| Almost all those people interpret anyone identifying
| themselves as "Liberal" as definitely partisan (and
| negative, of course). Most of them see "Progressive" as
| dangerously close to "Socialist". And the same holds true
| for the term "Conservative" on the other side of the
| spectrum, of course.
|
| No, identifying as "Progressive" wouldn't distance you
| from political connotations and culture warring, it's
| leaping into the maelstrom yelling "Yipee-Ki-Yay!" You
| may want to update your priors regarding how the broad
| populace perceives political labels. With the populace
| divided almost exactly in half regarding politics and
| cultural war issues and a large percentage on both sides
| having "Strong" or "Very Strong" feelings, stating any
| position will be seen as strongly negative by tens of
| millions of people. As was said in the movie "WarGames",
| the only winning move is not playing.
| rybosworld wrote:
| Seems like an extremely naive take.
|
| In 2016: Sam alluded to Trump's rise as not dissimilar to
| Hitler's. He said that Trump's ideas on how to fix things
| are so far off the mark that they are dangerous. He even
| quoted the famous: "The only thing necessary for the
| triumph of evil is for good men to do nothing."
|
| In 2025: "I'm not going to agree with him on everything,
| but I think he will be incredible for the country"
|
| This is quite obviously someone who is pandering for
| their own benefit.
| benatkin wrote:
| Just like JD Vance.
| themaninthedark wrote:
| This is quite honestly one of the major problems with our
| society right now. Once you take a public stance, you are
| not allowed to revisit and re-evaluate. I think that this
| is by and large driving most of the polarization in the
| country, since "My view is right and I will not give an
| inch least I be seen as weak".
|
| While most of the things affected are highly political
| situations, i.e. Trump's ideas or Biden's fitness. We
| also seem to have thrown out things that we used to
| consider cornerstones of liberal democracy i.e. our ideas
| regarding free speech and censorship, where we claim that
| it's not happening because it is a private company.
| belter wrote:
| "Donald Trump represents an unprecedented threat to
| America, and voting for Hillary is the best way to defend
| our country against it" -
| Sam Altman - 2016
|
| "If you elect a reality TV star as President, you can't be
| surprised when you get a reality TV show"
| - Sam Altman - 2017
|
| "When the future of the republic is at risk, the duty to
| the country and our values transcends the duty to your
| particular company and your stock price."
| - Sam Altman - 2017
|
| "I think I started that a little bit earlier than other
| people, but at this point I am in really good company"
| - Sam Altman - 2017 ( On his criticism of Trump )
|
| "Very few people realize just how much @reidhoffman did and
| spent to stop Trump from getting re-elected -- it seems
| reasonably likely to me that Trump would still be in office
| without his efforts. Thank you, Reid!"
| - Sam Altman - 2020
| meowface wrote:
| Although I dislike him now glazing Trump, I understand why
| he's doing it. Trump runs a racket and this is part of the
| game.
|
| One of my most contrarian positions is I still like and
| support Altman, despite most of the internet now hating him
| almost as much as they (justifiably) hate Elon. Was a fan
| of Sam pre-YC presidency and still am now.
|
| (I also am a big fan of DeepSeek and its CEO.)
| CoastalCoder wrote:
| In the interest of helping avoid an echo chamber, would
| you mind giving some of the things you like about current
| Altman?
| benterix wrote:
| I'd love to hear something positive about the current
| Altman, too. Anything would be good.
| robotresearcher wrote:
| For me, it's the technical results. Same as for Musk.
|
| Tesla accelerated us forward into the electric car age.
| SpaceX revolutionized launches.
|
| OpenAI added some real startup oomph to the AI arms race
| which was dominated by megacorps with entrenched products
| that they would have disrupted only slowly.
|
| So these guys are doing useful things, however you feel
| about their other conduct. Personally I find the gross
| political flip-flops hard to stomach.
| whatshisface wrote:
| Why would you support someone you said was part of a
| racket in the sentence before? We're talking about real
| life, where actions have consequences, not a TV show
| where we're expected to identifiy with Tony Soprano.
| breakyerself wrote:
| If you didn't sexually assault your sister you're already
| off to a good start.
| 65 wrote:
| Yeah I don't know, Altman is a sociopath who is now trying to
| get intertwined with local governments (SF) as well as the
| federal government. He's going to do a lot of weaseling to
| get what he wants: laws that forcibly make OpenAI a monopoly.
|
| Society will always have crazy sociopaths destroying things
| for their own gain, and now is Altman's turn.
| api wrote:
| I'm sure him lining up to kiss Trump's ring for some kind of
| bailout is not a coincidence.
| blackeyeblitzar wrote:
| I don't care for Sam Altman and his general untrustworthy
| behavior. But DeepSeek is perhaps more untrustworthy. Models
| from American companies at least aren't surprising us with
| government driven misinformation, and even though safety can
| also be censorship, the companies that make these models at
| least openly talk about their safety programs. DeepSeek is
| implementing a censorship and propaganda program without
| admitting it at all, and once they become good at doing it in
| less obvious ways, it can become very damaging and corrupt
| the political process of other societies, because users will
| trust the tools they use are neutral.
|
| I think DeepSeek's strategy to announce a misleading low cost
| (just the final training run that optimizes a base model that
| in turn is possibly based on OpenAI) is also purposeful.
| After all, High Flyer, the parent company of DeepSeek, is a
| hedge fund - and I bet they took out big short positions on
| Nvidia before their recent announcements. The Chinese
| government, of course, benefits from a misleading number
| being announced broadly, causing doubt among investors who
| would otherwise continue to prop up American technology
| startups. Not to mention the big fall in American markets as
| a result.
|
| I do think there's also a big difference between scraping the
| Internet for training data, which might just be fair use, and
| training off other LLMs or obtaining their assets in some
| other way. The latter feels like the kind of copying and
| industrial espionage that used to get China ridiculed in the
| 2000s and 2010s. Note that DeepSeek has never detailed their
| training data, even at a high level. This is true even in
| their previous papers, where they were very vague about the
| pre training process, which feels suspicious.
| dheera wrote:
| > I bet they took out big short positions on Nvidia before
| their announcements
|
| Good for them! I hope this teaches Wall Street to not freak
| out about an unverified announcement.
|
| Wall Street lost billions, and I hope they learned their
| lesson and next time will not crash the market when
| unverified news comes out.
| tempusalaria wrote:
| DeepSeek v3 (where the training cost claims come from) was
| announced a month ago and it had no impact outside of a
| small circle
| ryanisnan wrote:
| > Models from American companies at least aren't surprising
| us with government driven misinformation, and even though
| safety can also be censorship
|
| Being a citizen of a western nation, I'm inclined to agree
| with the general sentiment here, but how can you
| definitively say this? You, or I, don't know with any
| certainty what interference the US government has played
| with domestic LLMs, or what lies they have fabricated and
| cultivated, that are now part of those LLMs' collective
| knowledge. We can see the perceived censorship with
| deepseek more clearly, but that isn't evidence that we're
| in any safer territory.
| pphysch wrote:
| > Models from American companies at least aren't surprising
| us with government driven misinformation
|
| There are loads of examples on the internet of LLMs pushing
| (foreign) government narratives e.g. on Israel-Palestine.
|
| Just because you might agree with the propaganda doesn't
| make it any less problematic.
| blackeyeblitzar wrote:
| > There are loads of examples on the internet of LLMs
| pushing (foreign) government narratives e.g. on Israel-
| Palestine
|
| There isn't even a single example of that. If an LLM is
| taking a certain position because it has learned from
| articles on that topic, that's different from it being
| manipulated on purpose to answer differently on that
| topic. You're confusing an LLM simply reflecting the
| complexity out there in the world on some topics (showing
| up in training data), with government forced censorship
| and propaganda in DeepSeek.
|
| The two aren't the same, not even remotely close.
| pphysch wrote:
| Fine, whatever. It's actually much _more_ concerning if
| the overall information landscape has been so curated by
| censors that a naively-trained LLM comes "pre-censored",
| as you are asserting. This issue is so "complex" when it
| comes to one side, and "morally clear" when it comes to
| the other. Classic doublespeak.
|
| That's far more dystopian than a post-hoc "guardrailed"
| model (that you can run locally without guardrails).
| vohk wrote:
| > Models from American companies at least aren't surprising
| us with government driven misinformation
|
| Is corporate misinformation so much better? Recall about
| Tienanmen Square might be more honest but if LLMs had been
| available over the past 50 years, I would expect many
| popular models would have cheerfully told us company towns
| are a great place to live, cigarettes are healthy,
| industrial pollution has no impact on your health, and
| anthropogenic climate change isn't real.
|
| Especially after the recent behaviour of Meta, Twitter, and
| Amazon in open support of Trump and Republican interests,
| I'll be shocked if we don't start seeing that reflected in
| their LLMs over the next few years.
| cycomanic wrote:
| > I don't care for Sam Altman and his general untrustworthy
| behavior. But DeepSeek is perhaps more untrustworthy.
| Models from American companies at least aren't surprising
| us with government driven misinformation, and even though
| safety can also be censorship, the companies that make
| these models at least openly talk about their safety
| programs. DeepSeek is implementing a censorship and
| propaganda program without admitting it at all, and once
| they become good at doing it in less obvious ways, it can
| become very damaging and corrupt the political process of
| other societies, because users will trust the tools they
| use are neutral.
|
| These arguments always remind me of the arguments against
| Huawei because they _might_ be spying on western countries.
| On the other hand we had the US government working hand in
| hand with US corporations in proven spying operations
| against western allies for political and economic gain. So
| why should we choose an American supplier over a Chinese
| one?
|
| > I think DeepSeek's strategy to announce a misleading low
| cost (just the final training run that optimizes a base
| model that in turn is possibly based on OpenAI) is also
| purposeful. After all, High Flyer, the parent company of
| DeepSeek, is a hedge fund - and I bet they took out big
| short positions on Nvidia before their recent
| announcements. The Chinese government, of course, benefits
| from a misleading number being announced broadly, causing
| doubt among investors who would otherwise continue to prop
| up American technology startups. Not to mention the big
| fall in American markets as a result.
|
| Why should I care about the stock value of US corporations?
|
| > I do think there's also a big difference between scraping
| the Internet for training data, which might just be fair
| use, and training off other LLMs or obtaining their assets
| in some other way.
|
| So if training of copyrighted work scrapped of the Internet
| is fair use, how would the training of the LLMs not be fair
| use as well? You can't have it both ways.
| tootie wrote:
| The coup against him is looking more and more like a huge "I
| told you so" moment.
| Dansvidania wrote:
| Indeed. First thing I thought was "call a wahmbulance!".
| benreesman wrote:
| The guy is a total fucking psycho and the rest of the board
| are no gems either.
|
| Their failure is important at a minimum to the future of the
| United States if not the world.
| dang wrote:
| Ok, but please don't break HN's rules when commenting here.
|
| You may not owe Altmen better, but you owe this community
| better if you're participating in it.
|
| https://news.ycombinator.com/newsguidelines.html
| stackghost wrote:
| Once again you abuse your moderator powers to enforce your
| personal vendetta against people who dare to speak ill of
| tech CEOs.
|
| I find your behavior repulsive and fervently wish you would
| quit.
| dang wrote:
| This is what people say when they don't want the rules to
| be applied even-handedly.
|
| It's not a borderline call--I'd post exactly the same
| thing regardless of who or what such a comment was about.
| stackghost wrote:
| >This is what people say when they don't want the rules
| to be applied even-handedly.
|
| Not even close.
|
| This guy is actively ruining society while enriching
| himself in the process, but we somehow can't call a spade
| a spade?
|
| Pathetic.
| dang wrote:
| I suppose my chances of getting a straight answer aren't
| too good right now but I'd love to hear your thoughts on
| something.
|
| HN's stated mandate is intellectual curiosity
| (https://news.ycombinator.com/newsguidelines.html, https:
| //hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
| ).
|
| Do you feel like your comment
| https://news.ycombinator.com/item?id=42866108 counts as
| curious? or is it rather that you think something else is
| more important?
| belter wrote:
| So it is true, they run out of Data to steal? :-)
|
| And then where DeepSeek steal from next? Do they steal from
| themselves? Do they steal the stolen models they stole from the
| stolen data?
|
| The AI Ponzi scheme...
| troyvit wrote:
| Exactly this, especially as journalism melts down into slag.
| Soon all anybody will have to train on is social media,
| Wikipedia and GitHub, and that last one will slowly be
| metastasized by AI-generated code anyway.
|
| It reminds me of 1984 in a sense. "Don't you see that the
| whole aim of Newspeak is to narrow the range of thought? In
| the end we shall make thoughtcrime literally impossible,
| because there will be no words in which to express it."
|
| Unlike 1984 I don't see this winnowing of new concepts as
| purposeful, but on the other hand I keep asking myself how we
| can be so stupid as to keep doing it.
| TypingOutBugs wrote:
| Screw OpenAI, they scrape us without issues so someone scraped
| them. No issues with this.
| coliveira wrote:
| But the government will now claim this is against "national
| security". Only American companies are allowed to commit this
| kind of "sleight of hand".
| Imustaskforhelp wrote:
| Yes they would. But it would pointless. And clear hypocrisy
| as well.
| coliveira wrote:
| Hypocrisy or not, the US government has managed to make
| this work for a long time now, the Biden administration
| just proves the point. Thankfully, other countries are
| starting to catch up to this scam.
| Imustaskforhelp wrote:
| Yes , to be fair , As a foreigner (not a US citizen
| basically) I don't mean to offend somebody. But USA just
| seems to be build on top of Hypocrisy.
|
| Like the fact that US revolution was basically
| kickstarted by blatantly breaking the patent law (like
| there was this one mill specifically) , I think its a
| historic event. And now here we are ! The scam of
| national security.
|
| To be honest. People seem to be really kind on the fall
| of USA. I am not that interested since the rise of China
| terrifies me. But the hypocrisy of USA / losing such soft
| power (like here I am , from random country critiquing
| USA based on facts , it really downplays it being a
| superpower) that would be the downfall of USA.
|
| To me , the future terrifies me. In fact the present
| terrifies me. I think the world is running crazy or maybe
| its just me.
| pen2l wrote:
| While all of this is true, that DeepSeek wouldn't be here were
| it not for the research that preceded it notably Google's
| paper, then Llama, and ChatGPT which they're modeled after, its
| release still did something profound to their psyche, the
| motivation and self-actualization this instills to the Chinese.
| They witnessed the power of their accomplishments: a side-
| hustle project knocked off an easy trillion. This is only
| egging them on and will serve to ramp up their efforts even
| more.
|
| Separately, I do think that now that the Chinese leadership saw
| this, that they have the chops to pull this off and then some,
| they are probably going to rein in future innovations; they'll
| likely demand that the big future discoveries remain closed-
| sourced (or even unannounced/unpublicized).
| tedivm wrote:
| OpenAI wouldn't be here without the work that Yann Lecun did
| at Facebook (back when it was facebook). Science is built on
| top of science, that's just how things work.
| wrasee wrote:
| Yes, but in science you reference your work and credit
| those who came before you.
|
| Edit: I am not defending OpenAI and we are all enjoying the
| irony here. But it puts into perspective some of the wilder
| claims circulating that DeekSeek was able to somehow
| complete with OpenAI for only $5M, as if on a level playing
| field.
| bugglebeetle wrote:
| Like all those papers with their long lists of citations
| OpenAI has been releasing?
| dkjaudyeqooe wrote:
| That's only in academia. The same thing happens in
| commerce, only there is no (official) credit given.
| tedivm wrote:
| OpenAI has been hiding their datasets, and certainly
| haven't credited me for the data they stole from my
| website and github repositories. If OpenAI doesn't think
| they should give attribution to the data they used, it
| seems weird to require that of others.
|
| Edit: Responding to your edit, Deepseek only claimed that
| the final training run was $5m, not that the whole
| process caught that (they even call this out). I think
| it's important to acknowledge that, even if they did get
| some training data from OpenAI, this is a remarkable
| achievement.
| ambicapter wrote:
| Only weird if you think what OpenAI did should be the
| norm.
| wrasee wrote:
| Right. I think many here are enjoying the Schadenfreude
| against OpenAI, but that hardly makes it right. It just
| makes it a race to the bottom.
| wrasee wrote:
| It is a remarkable achievement. But if "some training
| data from OpenAI" turns out to essentially be a wholesale
| distillation of their entire model (along with Llama etc)
| I do think that somewhat dampens the spirit of it.
|
| We don't know that of course. OpenAI claim to have some
| evidence and I guess we'll just have to wait and see how
| this plays out.
|
| There's also a substantial difference between training of
| the entire internet and one that very specifically
| targets your competitor's products (or any specific work
| directly).
| Filligree wrote:
| That's $5M for the final training run. Which is an
| improvement to be sure, but it doesn't include the
| _other_ training runs -- prototypes, failed runs and so
| forth.
| coliveira wrote:
| It is OpenAI that discredits themselves when they say
| that each new model is the result of hundreds of USD
| millions in training. They throw this around as it is a
| big advantage of their models.
| nicce wrote:
| And the cost is based on the imaginary currency that
| Microsoft has given for them as Azure computing.
| blackeyeblitzar wrote:
| Is that really true? If anything OpenAI was dependent on
| the transformers paper from Google from Ashish Vaswani and
| others. LeCun has been criticizing LLM architectures for a
| long time and has been wrong about them for a long time.
| mv4 wrote:
| That was my impression too. He is considered the inventor
| of CNN back in 1998. Is there anything more recent that's
| meaningful?
| blackeyeblitzar wrote:
| Personally, I have not seen anything from him that is
| meaningful. OpenAI and Anthropic (itself started by
| former OpenAI people) of course have built their models
| without LeCun's contributions. And for a few years now,
| LeCun has been giving the same talk anywhere he makes
| appearances, saying that large language models are a dead
| end and that other approaches like his JEPA architecture
| are the future. Meanwhile current LLM architecture has
| continued to evolve and become very useful. As for the
| misuse of the term "open source", I think that really
| began once he was at Meta, and is a way to use his fame
| to market Llama and help Meta not look irrelevant.
| tedivm wrote:
| They literally cited LeCun in their GPT papers.
| tedivm wrote:
| I was more referring to this paper from 2015:
|
| https://scholar.google.com/citations?view_op=view_citatio
| n&h...
|
| Basically all LLM can trace their origin back to that
| paper.
|
| This was just a single example though. The whole point is
| that people build on the work from the past, and that
| this is normal.
| mv4 wrote:
| Thank you for sharing this.
| esafak wrote:
| That's just an overview for paper for those new to the
| field. The transformer architecture has a better claim to
| being the origin of LLMs.
| amelius wrote:
| By the way, as someone who once did classical image
| recognition using convolutions, I can't say I was very
| impressed by the CNN approach, especially since their
| implementation didn't even use FFTs for efficiency.
| zbendefy wrote:
| Also without the "attention is all you need" paper from
| google
| openrisk wrote:
| > they'll likely demand that the big future discoveries
| remain closed-sourced
|
| Depends on whether they want these tools to be adopted in the
| wider world. Rightly or wrongly there is a lot of suspicion
| in the West and an open source approach builds trust.
| hn_throwaway_99 wrote:
| > While all of this is true, that DeepSeek wouldn't be here
| were it not for the research that preceded it (notably
| Llama), and ChatGPT which they're modeled after...
|
| If the allegation _is_ true (we don 't know yet), then what
| you've written perfectly proves the point everyone is making.
| ChatGPT wouldn't be here if it weren't for all the research
| and work that preceded it in terms of tons of scrapable
| content being available on the Internet, and it's not like
| OpenAI invented transformers either.
|
| Nobody is accusing DeepSeek of hacking into OpenAI's systems
| and stealing their content. OpenAI is just saying they
| scraped them in an "unauthorized" manner. The hypocrisy is
| laughably striking, but sadly nobody has any shame anymore in
| this world it seems. Play me the world's tiniest violin for
| OpenAI.
| stravant wrote:
| Yes, and what does preceding research do? Get followed by
| more research building on it.
| dylan604 wrote:
| Standing on the shoulders and it's turtles all the way
| dismalaf wrote:
| Don't forget all the research that came before OpenAI and
| ChatGPT...
| nicce wrote:
| We wouldn't be here discussing if nobody invented internet...
| nor these models had training data at all.
|
| > Separately, I do think that now that the Chinese leadership
| saw this, that they have the chops to pull this off and then
| some, they are probably going to rein in future innovations;
| they'll likely demand that the big future discoveries remain
| closed-sourced (or even unannounced/unpublicized).
|
| How do we know that this is not already happening with
| OpenAI/Meta and the U.S. government at some level? The
| concept of power is equal, whether we wanted it or not. We
| don't have to pretend to be "better" all the time.
| coliveira wrote:
| They really lost their minds. They're all scared and worried
| because companies in other countries can also access the same
| data they stole from the Internet.
| rvz wrote:
| They have been out-grifted by DeepSeek and OpenAI is not happy
| about someone out-shining them on that.
|
| The best part is "their IP" was humanity's scraped content and
| they are angry that DeepSeek did their job for them and gave it
| away for free.
| cscurmudgeon wrote:
| Scraping data is different from scraping outputs from a model.
| jacobgorm wrote:
| No it is not, data is data, whether it gets loaded from a
| file on disk or generated by multiplying lots of matrices.
| idle_zealot wrote:
| Like, in a strict literal sense, sure? Do you mean to make a
| claim about moral or legal differences?
| openrisk wrote:
| Because the data is mine and the model is yours?
| wkz wrote:
| Technically, sure. What is the moral distinction though?
| Rebelgecko wrote:
| Copyright is weird and often legal [?] moral, but I'm having
| a hard time constructing a mental model where it's ok to
| scrape a novel written by a person but it's not ok to scrape
| a story written by chatgpt
| 28304283409234 wrote:
| ClosedAI? StolenAI!
| gruez wrote:
| The picture at the end showing deepseek's privacy policy and
| being concerned that it's "a security risk" is hilarious[1].
| Basically every B2C company collects this sort of
| information[2], and is far less intrusive than what social
| networks collect[3]. But because it's Chinese and at the risk
| of overtaking Western companies, people are suddenly worried
| about device information and IP addresses?
|
| [1] https://semking.com/wp-
| content/uploads/2025/01/DeepSeek-1024...
|
| [2] https://www.bestbuy.com/site/help-topics/privacy-
| policy/pcmc...
|
| [3] https://www.facebook.com/privacy/policy/
| semking wrote:
| One of my core followers named Bruno basically said the same
| thing under my Linkedin post yesterday:
|
| https://www.linkedin.com/posts/organic-growth_deepseek-
| the-o...
|
| I welcome friction, so I'll be blunt: I disagree with you,
| not because what you are saying is wrong but because you only
| consider systematic data collection.
|
| That's not the issue here.
|
| There's a difference between democracies like the United
| States or European countries, no matter how IMPERFECT they
| are, and a dictatorship that does not allow dissenting
| opinions.
|
| There's a difference in how the data collected will be used.
|
| Freedom of speech, even when it is relative, is better than
| totalitarianism.
| ryanobjc wrote:
| It's also important to recognize that the Chinese
| government is known to walk into internet service companies
| and demand they censor, alter data, delete things. No court
| order or search warrant required.
|
| China considers industry to be completely subservient to
| government. Checks and balances are secondary to ideas like
| harmony and collective well being.
| semking wrote:
| Thank you for this balanced and essential comment which
| is entirely true!
| gruez wrote:
| >There's a difference between democracies like the United
| States or European countries, no matter how IMPERFECT they
| are, and a dictatorship that does not allow dissenting
| opinions.
|
| >There's a difference in how the data collected will be
| used.
|
| >Freedom of speech, even when it is relative, is better
| than totalitarianism.
|
| I don't disagree with "democracy is better than
| totalitarianism", but what does that have to do with
| collecting device information and IP addresses? Is that
| excuse a cudgel you can use against any behavior that would
| otherwise be innocuous? It's fine to be against deepseek
| because you're concerned about them getting sensitive data
| via queries, or even that their models be a backdoor to
| project chinese soft power, but hand wringing about device
| information and IP addresses is absurd. It makes as much
| sense as being concerned that the CCP/deepseek does
| _meetings_ , because even though every other companies does
| meetings, CCP/deepseek meetings could be used for
| totalitarianism.
| semking wrote:
| I don't disagree with you either and like you, I'm
| entirely against privacy violations in any way, shape or
| form.
|
| I admit I am concerned when I see blatant algorithmic
| manipulation of social platforms to favor any narrative
| that aligns with geopolitical objectives.
|
| I also wrote about the TikTok algo a few days ago. You'll
| see what I think of user privacy violations (closed
| ecosystem + basically a keylogger in this case):
|
| https://semking.com/likes-lies-untold-story-tiktok-
| algorithm...
|
| I cannot stand when dissenting voices or opinions are
| shadow-banned.
|
| And I have the same opinion regarding U.S. or EU
| companies.
|
| Our privacy should be respected.
|
| In the meantime: strong encryption at every corner,
| please!
| pphysch wrote:
| > I admit I am concerned when I see blatant algorithmic
| manipulation of social platforms to favor any narrative
| that aligns with geopolitical objectives.
|
| I'm curious how robust this principle is for you, because
| China and Russia are not the first countries that come to
| mind when talking about the (actual, existing,
| documented) manipulation of US speech and media by a
| foreign government.
|
| Yet it seems we can only have this discussion,
| ironically, when the subject is a US government-approved
| one like China. Anything else would be problematic and
| unsafe.
| semking wrote:
| I don't want to get into politics but I'll gladly admit
| human beings are biased.
|
| "We Don't See Things As They Are, We See Them As We Are"
|
| -- Samuel b. Nahmani
| gruez wrote:
| >I'm entirely against privacy violations in any way,
| shape or form.
|
| >Our privacy should be respected.
|
| Characterizing device information and IP addresses as
| "privacy violations" is a stretch. If you showed a
| history railing against this sort of stuff, agnostic of
| geopolitical alignment, then you get a pass, but I think
| it's fair to assume the converse until proven otherwise.
|
| >In the meantime: strong encryption at every corner,
| please!
|
| Irrelevant. The data collection is done by first parties.
| Encryption doesn't do anything.
|
| >I admit I am concerned when I see blatant algorithmic
| manipulation of social platforms to favor any narrative
| that aligns with geopolitical objectives.
|
| >I cannot stand when dissenting voices or opinions are
| shadow-banned.
|
| What does this have to do with privacy? Again, it's fine
| to be against "blatant algorithmic manipulation of social
| platforms" or whatever, but dragging seemingly unrelated
| topics in an attempt to amass as big pile of greviances
| as possible is disingenuous.
|
| >I also wrote about the TikTok algo a few days ago.
| You'll see what I think of user privacy violations
| (closed ecosystem + basically a keylogger in this case):
|
| >https://semking.com/likes-lies-untold-story-tiktok-
| algorithm...
|
| Where's the keylogging? I skimmed the article and the
| only thing I could find was a passing mention about an
| article that you "was advised not to publish it and I
| didn't". How much keylogging could possibly going on in a
| short video app? Is the "keylogging" just a way to make
| "we measure how engaged someone is with a video" as
| sinister as possible?
| semking wrote:
| >Characterizing device information and IP addresses as
| "privacy violations" is a stretch.
|
| I agree: this is a characterization I never made. FYI, I
| also collect this type of data about you when you visit
| my website. That said, telemetry + totalitarianism = bad
| combo.
|
| >Irrelevant. The data collection is done by first
| parties. Encryption doesn't do anything.
|
| Even if data is collected by first parties, encryption is
| still highly relevant because it ensures that the data
| remains secure in transit and at rest. It does a lot.
|
| >What does this have to do with privacy? Again, it's fine
| to be against "blatant algorithmic manipulation of social
| platforms" or whatever, but dragging seemingly unrelated
| topics in an attempt to amass as big pile of greviances
| as possible is disingenuous.
|
| You are aggressive for no reason whatsoever. There's
| nothing disingenuous: when users are shadow-banned by
| platforms under dictatorships, they end up flagged, and
| their private data is often analyzed for nefarious
| reasons. There's a link with privacy but I'll stop at
| this stage if we cannot have a civilized discussion.
|
| >Where's the keylogging? I skimmed the article and the
| only thing I could find was a passing mention about an
| article that you "was advised not to publish it and I
| didn't". How much keylogging could possibly going on in a
| short video app? Is the "keylogging" just a way to make
| "we measure how engaged someone is with a video" as
| sinister as possible?
|
| "TikTok iOS subscribes to every keystroke (text inputs)
| happening on third party websites rendered inside the
| TikTok app. This can include passwords, credit card
| information and other sensitive user data. (keypress and
| keydown). We can't know what TikTok uses the subscription
| for, but from a technical perspective, this is the
| equivalent of installing a keylogger on third party
| websites."
|
| https://krausefx.com/blog/announcing-inappbrowsercom-see-
| wha...
|
| Please note that this article is outdated (August 2022).
| Importantly, the article does not claim that any data
| logging or transmission is actively occurring. Instead,
| it highlights the potential technical capabilities of in-
| app browsers to inject JavaScript code, which could
| theoretically be used to monitor user interactions.
| coliveira wrote:
| Also, the same people that complain about this are just
| fine with a western government having access to the same
| data via big corporations. Why being democratic gives you
| a free access card to disregard privacy, in other words,
| doing exactly the opposite of what is expected from a
| free society?
| ziddoap wrote:
| > _There 's a difference in how the data collected will be
| used._
|
| Not that we could ever see what the NSA, CISA, ASIS, GCHQ,
| and other 3/4-letter agencies are actually doing with the
| collected data.
|
| But they pinky promised to use it properly (or something),
| so, yay.
| r00fus wrote:
| Amusing Bruno seems to think in terms of labels when the
| reality is that the USA imprisons far more people per
| capita, and blatantly disregards its so-called "core
| freedoms" (ie, Bill of Rights) for its citizens very often.
|
| This kind of person has a lot of cognitive dissonance going
| on.
| scotty79 wrote:
| "That's hilarious!" was my first reaction as well, when I heard
| about it the first time. When I came to HN and saw this story
| on top I was hoping this was the top comment. I was not
| disappointed.
|
| US AI folk were leading for two years by just throwing more and
| more compute at the same thing that Google threw them like a
| bone years ago (namely transformers). They made next to no
| innovation in any area other than how to connect more compute
| together. The idea of additional inference time compute,
| looping the network back on its own outputs, which is the only
| significant conceptual advancement of last years was something
| I, as a layman, came up with after few days of thinking why AI
| sucks and what can be done to make it able to tackle problems
| that require iterative reasoning. They announced it few weeks
| after I came up with the idea, so it was in the works for some
| time, but it shows you how basic idea it was. There was nothing
| else.
|
| Suddenly when there comes a small company that introduced few
| actual algorithmic advancements which resulted in 100x
| optimization which is something expected with algorithmic
| optimizations, the big AI suddenly went into full "dog ate my
| homework" mode. Blaming everyone and everything around.
|
| Let's not mention the fact that if full outputs of their models
| could enable them to train a better model at 1% cost then it
| puts them in even worse light that they didn't do it.
| ryanobjc wrote:
| It's not often you get 100x optimization with some small
| improvements so I'm kind of skeptical.
|
| We have and apples and oranges thing here which deepseek is
| intentionally leaning into. They get very cheap electricity
| and are bragging about their cheap cost, and OpenAI etc
| typically brag about how expensive their training is. But
| it's all pr and lies.
| enragedcacti wrote:
| > They get very cheap electricity and are bragging about
| their cheap cost
|
| The cost of $5.5 million was quoted at $2/GPU-hour which is
| a reasonable price for on-demand H100s that anyone in the
| US could access, and likely on the high side given bulk
| pricing and that they are using nerfed versions. OpenAI
| might be all pr and lies but everything I've seen so far
| says that deepseek's claims about cost are legit.
| Leary wrote:
| Does this mean when you use OpenAI as an enterprise customer,
| they can see exactly the queries and answers? So much for
| privacy!
| api wrote:
| So far the whole business model of Silicon Valley since social
| media has been to monetize other peoples' content given out for
| free. The whole empire is built on this.
|
| I wonder if this is going to come to an end through a
| combination of social media fatigue, social media
| fragmentation, and open source LLMs just giving it all back to
| us for free. LLMs are analogous to a "JPEG for ideas" so
| they're just lossy compression blobs of human thought expressed
| through language.
| barnabee wrote:
| > So far the whole business model of Silicon Valley since
| social media has been to monetize other peoples' content
| given out for free. The whole empire is built on this.
|
| It cannot die soon enough
| okdood64 wrote:
| Not to mention the total dodge when Murati was asked about
| training on the YouTube corpus during that television
| interview.
|
| Sorry for the Short: https://www.youtube.com/shorts/M0QyOp7zqcY
| the_arun wrote:
| I think the point is - OpenAI scraped public data - d1 -
| Trained their model to produce output - d2 - DeepSeek used d2
| to reinforce their model
|
| OpenAI is mad about d2 (not d1). I'm not sure using public data
| is "stealing". In summary, these are two different things &
| need to be separate.
| redleader55 wrote:
| You say "public", but what I think you mean is "publicly
| available". Even publicly available data has copyrights, and
| unless that copyright is "public domain", you need to follow
| some rules. Even licenses like Creative Commons, which would
| be the most permissive, come with caveats which OpenAI
| doesn't follow [0].
|
| It is unclear if someone breaking someone else's copyright to
| use A can claim copyright on a work B, derived from A. My
| point is that OpenAI played loose with the copyright rules to
| build its various models, so the legality of their claims
| against DeepSeek might not be so strong.
|
| [0] https://creativecommons.org/share-your-work/cclicenses/
| the_arun wrote:
| I am not saying OpenAI did good by using publicly available
| data. I meant these are separate activities. None is good.
| But DeepSeek is slightly better by making theirs
| opensource.
| xbar wrote:
| OpenAI (sc)raped all the data it could. I do not accept your
| assertion that d1 was "public." It was accessible, for
| certain.
|
| OpenAI asserts 1. d2 was used by DeepSeek 2. All d2 belongs
| to OpenAI exclusively
|
| Both are debatable for large number of reasons.
| rubslopes wrote:
| > Our mission is to ensure that artificial general intelligence
| benefits all of humanity.[1]
|
| Well, I guess they really helped make this a reality!
|
| [1] https://openai.com/about/
| didip wrote:
| fr fr, ClosedAI is being a comedian right now.
|
| They scraped literally all the content of the internet without
| permissions. And I won't even be surprised if they scraped the
| output of other LLMs as well.
| adzm wrote:
| Why does this post use DeepSink instead of DeepSeek at
| apparently random places? Is that just a pejorative pun like
| ClosedAI?
| skeeter2020 wrote:
| I share the sentiment here, but asking as a noob: does this
| mean the performance comparison is not really apples to apples?
| If it required the distillation of the expensive model in order
| to get such good results for a much lower price, is that shady
| accounting?
| schmit wrote:
| Even more hilarious given their own charter:
|
| > We will attempt to directly build safe and beneficial AGI,
| but will also consider our mission fulfilled if our work aids
| others to achieve this outcome.
|
| > Our primary fiduciary duty is to humanity. We anticipate
| needing to marshal substantial resources to fulfill our
| mission, but will always diligently act to minimize conflicts
| of interest among our employees and stakeholders that could
| compromise broad benefit.
|
| > We will actively cooperate with other research and policy
| institutions; we seek to create a global community working
| together to address AGI's global challenges.
| semking wrote:
| Ah yes: "duty to humanity"
| hn_throwaway_99 wrote:
| I think one good thing to come out of all this tech elite
| flip flopping is that I now see these tech leaders for
| exactly who they are. It makes me kind of sad, because as
| someone who came of age early in the Web era I really
| _wanted_ to believe that there was a bigger moral good to
| all we were doing.
|
| I now view _any_ moralistic statement by any of these big
| tech companies as complete and total bullshit, which is
| probably for the best, because that is what it is. These
| companies now exist solely to amass power and wealth. They
| will still use moralistic language to try to motivate their
| employees, but I hope folks still see it for the complete
| nonsense that it is.
| radicality wrote:
| I liked Matt Levine's newsletter few days ago where he
| hypothesized scenarios where it's much more profitable to short
| your competitors, then release a much better version of some
| widget completely free, and then profit $$$. Which is plausible
| here too, considering DeepSeek is made by a hedge fund.
| greasegum wrote:
| Came here to mention this too. Seem almost so obvious that
| I'm surprised this isn't the dominant angle.
| freehorse wrote:
| How would that work out here though? "Open"AI is not publicly
| traded. Any kind of shorting would be quite indirect.
| Imustaskforhelp wrote:
| Yes the irony is so thick in the air that it can be cut through
| using a swiss knife lol
|
| I had literally come to this post to say the same. You beat me
| to it.
|
| USA is going crazy over deepseek and to me , it just shows that
| the world is a black swan , an AI bubble.
|
| I am not saying AI has no use. I regularly use it to create
| something , but its just not recommended. I am going to stop
| using AI , to grow my mind.
|
| And its definitely way overpriced. People are investing so much
| money without seeing the returns? , and I think people are also
| using AI because of a sense of FOMO , I don't know , to me its
| funny .
|
| I really really want to create a index fund with strictly no AI
| companies. Since this doesn't feel diversified enough. Like
| sure nvidia gave a quarter of return the last year , but I mean
| , at this point , it almost feels the same as that of bitcoin.
| The reason I don't / won't invest in bitcoin is I don't want
| "that" risk.
|
| This has been a boggling year.
|
| I have realized that the world is crazy. Truly. Trump winning
| from going to the point of getting shot to deepseek causing
| nvidia / american stock market to go down , heck even bitcoin!
| , its so crazy , trump launching his meme coin. If the world is
| crazy. Just be the sane person around. You will stick around ,
| that's my philosophy. I won't jump on AI wandwagon . But its
| still absolutely wild & horror seeing how a "sideproject"
| (deepseek) absolutely put american stock market in shambles.
|
| I want more diversifaction. I am not satisfied with the current
| system. This feels like a bubble and I want no part in it.
| amelius wrote:
| It looks like they want to spin this as "DeepSeek copied
| OpenAI". The general public/media might actually believe this
| is what happened.
| breakitmakeit wrote:
| As the article points out, they are arguing in court against the
| new york times that publicly available data is fair game.
|
| The questions I am keenly waiting to observe the answer to
| (because surely Sam's words are lies): how hard is OpenAI willing
| to double down on their contradictory positions? What mental
| gymnastics will they use? What power will back them up, how, and
| how far will that go?
| snakeyjake wrote:
| When large sums of money are involved the techbros will burn
| everything down, go scorched earth no matter what the
| consequences, to keep what they believe they're entitled to.
| ADeerAppeared wrote:
| Their way of squaring this circle has always been to whine
| about "AI safety". (the cultish doomsday shit, not actual harms
| from AI)
|
| Sam Altman will proclaim that he alone is qualified to build AI
| and that everyone else should be tied down by regulation.
|
| And it should always be said that this is, of course, utterly
| ridiculous. Sam Altman literally got fired over this, has an
| extensive reputation as a shitweasel, and OpenAI's constant
| flouting and breaking of rules and social norms indicates they
| CANNOT be trusted.
| bhouston wrote:
| The US government likely will favor a large strategic company
| like OpenAI instead of individual's copyrights, so while ironic,
| the US government definitely doesn't care.
|
| And the US government is also likely itching to reduce the power
| of Chinese AI companies that could out compete US rivals (similar
| to the treatment of BYD, TikTok, solar panel manufacturers,
| network equipment manufacturers, etc), so expect sweeping
| legislation that blocks access to all Chinese AI endeavours to
| both the US and then soon US allies/West (via US pressure.)
|
| The likely legislation will be on the surface justified both by
| security concerns and by intellectual property concerns, but
| ultimately it will be motivated by winning the economic
| competition between China and the US and it will attempt to tilt
| the balance via explicitly protectionist policies.
| derektank wrote:
| >The US government likely will favor a large strategic company
| like OpenAI instead of individual's copyrights
|
| Even if we assume this is true, Disney and Netflix are both
| currently worth more than OpenAI and both rely on the strict
| enforcement of US copyright law. I do not think it is so
| obvious which powers that be have the better lobbying efforts
| and, currently, it's looking like this question will mostly be
| adjudicated by the courts, not Congress, anyways.
| bhouston wrote:
| I don't think OpenAI stole from Disney or Netflix. Rather
| OpenAI stole from individual artists and YouTube and other
| social media who users do not really have any lobbying power.
|
| So I think OpenAI, Disney and Netflix win together. Big
| companies tend to win.
| mjburgess wrote:
| > What are the first words of the disney movie, "Aladdin" ?
|
| The first words of Disney's _Aladdin_ (1992) are spoken by
| the *Peddler*, the mysterious merchant at the beginning of
| the film. He says:
|
| _" Ah, Salaam and good evening to you, worthy friend.
| Please, please, come closer..."_
|
| He then continues with: _" Too close! A little too close.
| There. Welcome to Agrabah. City of mystery, of enchantment,
| and the finest merchandise this side of the River Jordan,
| on sale today! Come on down!"_
|
| This opening sets the stage for the story, introducing the
| magical and bustling world of Agrabah.
| derektank wrote:
| Disney owns ABC News; OpenAI almost certainly scraped their
| text data
| bhouston wrote:
| I agree with you.
| worik wrote:
| > Rather OpenAI stole from individual artists and YouTube
| and other social media
|
| "stole"?
|
| They consumed publicly available material on the Internet
|
| I am no fan of these billionaire capitalists and their
| henchpersons but condem them for their multitude of sins.
|
| Consuming publicly available Internet resources is not one
| of them. IMO
| da_chicken wrote:
| Being publicly available does not mean that copyright is
| invalid. Copyright gives the holders the right to
| restrict USE, not merely restrict reproduction.
| _Adaptation_ is also an exclusive right of the copyright
| holder. You 're not allowed to make derivative works.
| visarga wrote:
| They stole the data just as much as a painter steals the
| view.
| rideontime wrote:
| Who created the view?
| visarga wrote:
| The view is created by every spectator.
| jdswain wrote:
| It's not that they consumed publicly available material,
| it's that they re-published that information, and sold
| it.
| Terr_ wrote:
| > They consumed publicly available material on the
| Internet
|
| I agree that there are some important distinctions and
| word-choices to be made here, and that there are problems
| with equating training to "stealing", and that copyright
| infringement is not theft, etc.
|
| _That said_ , if you zoom out to the overall conduct,
| it's fair to argue that the companies are doing something
| unethical, the same as if they paid an army of humans to
| memorize other people's work and then regurgitate
| slightly-reworded copies.
| tokioyoyo wrote:
| I don't think US government can move fast enough to change the
| trajectory. Also it doesn't help that basically every
| government is second guessing their alliance with the US. It's
| not an industry that can ruin local industries either (like
| cheap BYD is bad for German cars).
|
| It's a very fun thing to watch from the sidelines right now, if
| I'll be honest.
| buyucu wrote:
| It's too late for that. That ship sailed a long time ago.
|
| The best language model right now is open source. Let that sink
| in.
| _pferreir_ wrote:
| DeepSeek is not Open Source. That's like saying that
| Microsoft Edge is Open Source, as you can download it for
| free.
|
| https://huggingface.co/blog/open-r1
| ceejayoz wrote:
| "You can't take data without asking" seems like a court precedent
| OpenAI really, really, _really_ wants to avoid. And yet...
| amelius wrote:
| Why? When did large companies care about laws? See e.g. Uber,
| AirBnb.
|
| The only thing government cares about at this point is if
| information is shared with China.
| ceejayoz wrote:
| They care when they get big enough to attract attention from
| people like state AGs who can actually put the hurt on a bit.
| Uber and AirBnB both hit this point years ago; OpenAI's
| starting to hit it.
| galleywest200 wrote:
| Altman is part of that Stargate Trump group now. He and his
| ilk will just get pardons.
|
| Curious, though, can a corporation be pardoned?
| ceejayoz wrote:
| The President can only pardon Federal crimes.
|
| State-level crimes (like his NY felonies) and civil torts
| (like his case where he owes $500M currently) are
| separate.
| actionfromafar wrote:
| Yet. Give it some time.
| ceejayoz wrote:
| Sure, but in that scenario, it's a bit like the Last of
| Us characters being concerned about electrical meter
| readings. We'll have much bigger problems.
| layer8 wrote:
| OpenAI is saying that their service was used in violation of
| their TOS, which is a bit different than just copying data. To
| be clear I'm not on OpenAI's side, but it looks to me that the
| legal situation isn't exactly analogous.
| DebtDeflation wrote:
| Tons of websites and books they scraped had copyright
| notices.
| layer8 wrote:
| Copyright and terms of service are different legal notions.
| Maxion wrote:
| Yeah, copyright means something and a ToS is virtual
| toiletpaper (at least in the EU)
| layer8 wrote:
| This wasn't about which is worse than the other, but
| about whether OpenAI would want to avoid court precedent
| for the one because of the other.
| orlp wrote:
| If using data violating some ToS taints the model trained on
| that data, then all of OpenAI's models are tainted by the
| millions of ToS'es they broke.
| orionsbelt wrote:
| Can you cite a source showing they violated ToS?
| hdjjhhvvhga wrote:
| Not just violated but also actively ignored:
| https://news.ycombinator.com/item?id=42718850
| dkjaudyeqooe wrote:
| But whats the remedy in that case? Being banned from the
| service maybe, but no court is going to force a "return" of
| the data, so DeepSeek can't use it. It's uncopyrightable.
| kavalg wrote:
| As others have noted, if one company agrees to the ToS, asks
| "the right" questions and then publishes the ChatGPT answers,
| there is not violation of ToS. Then a second company scrapes
| the published Q&A, along with other information from the
| internet and again there is no violation (not more than the
| violations of OpenAI).
| hdjjhhvvhga wrote:
| > OpenAI is saying that their service was used in violation
| of their TOS
|
| Which is the most ridiculous argument they could use because
| they didn't respect any ToS (or copyright laws, for that
| matter) when scraping the whole web, books from Libgen and
| who knows what more.
| osigurdson wrote:
| I do think that distilling a model from another is much less
| impressive than distilling one from raw text. However, it is hard
| to say if it is really illegal or even immoral, perhaps just one
| step further in the evolution of the space.
| lemoncookiechip wrote:
| It's about as illegal as the billions, if not trillions of IPs
| that ClosedAI infringed to train their own data without
| consent. Not that they're alone, and I personally don't mind
| that AI companies do it, but it's still amusing when they get
| this annoyed at others doing the same thing to them.
| osigurdson wrote:
| I think they had the advantage of being ahead of the law in
| this regard. To my knowledge, reading copywritten material
| isn't (or wasn't illegal) and remains a legal grey area.
|
| Distilling weights from prompts and responses is even more of
| a legal grey area. The legal system cannot respond quickly to
| such technological advancements so things necessarily remain
| a wild west until technology reaches the asymptotic portion
| of the curve.
|
| In my view the most interesting thing is, do we really need
| vast data centers and innumerable GPUs for AGI? In other
| words, if intelligence is ultimately a function of power
| input, what is the shape of the curve?
| ttesmer wrote:
| > if intelligence is ultimately a function of power input,
| what is the shape of the curve?
|
| According to a quick google search, the human body consumes
| ~145W of power over 24h (eating 3000kcals/day). The brain
| needs ~20% of that so 29W/day. Much less than our current
| designs of software & (especially) hardware for AI.
| osigurdson wrote:
| I think you mean the brain uses 29W (i.e. not 29W/day).
| Also, I suspect that burgers are a higher entropy energy
| source than electricity so perhaps it is even less than
| that.
| lemoncookiechip wrote:
| The main issue is that they've had plenty of instances
| where the LLM outputted copyrighted content verbatim, like
| it happened with the New York Times and some book authors.
| And then there's DALL-E, which is baked into ChatGPT and
| before all the guardrails came up, was clearly trained on
| copyrighted content to the point it had people's
| watermarks, as well as their styles, just like Stable
| Diffusion mixes can do (if you don't prompt it out).
|
| Like you've put, it's still a somewhat gray area, and I
| personally have nothing against them (or anyone else) using
| copyrighted content to train models.
|
| I do find it annoying that they're so closed-off about
| their tech when it's built on the shoulders of openness and
| other people's hard work. And then they turn around and
| throw Issy fits when someone copies their homework,
| allegedly.
| JTyQZSnP3cQGa8B wrote:
| Illegally acquiring copyrighted material has always been
| highly illegal in France and I'm sure most other countries.
| Disney is another example of how it not grey at all.
| greiskul wrote:
| > Distilling weights from prompts and responses is even
| more of a legal grey area.
|
| Actually unless the law changes this is pretty settled
| territory in US law. All output of AIs are not
| copyrightable, and are therefore in the public domain. The
| only legal avenue of attack OpenAi has is Terms of Service
| violation, which is a much weaker breach then copyright if
| it is even true.
| ReptileMan wrote:
| Is the question of training AI on data fair use settled yet?
| Because if it is not - it looks like fair use to me.
| scotty79 wrote:
| Isn't it more impressive given that training on model output
| usually leads to worse model?
|
| If they actually figured out how to use output of existing
| models to build model that outperforms them then it's something
| that brings us closer to singularity than every other
| development so far.
| __MatrixMan__ wrote:
| If they want us to care they can open up their models so we can
| be the judge.
| 827a wrote:
| This smells very suspiciously like: someone who doesn't know
| anything about AI (possibly Sacks) demanding answers on R1 from
| someone who doesn't have any good ones (possibly Altman). "Uh,
| (sweating), umm, (shaking), they stole it from us! Yeah, look at
| this suspicious activity, that's why they had it so easy, we did
| all the hard work first!"
| fundad wrote:
| I think it's funny that OpenAI wants us to pay them to use
| their product to generate content but then sets the terms that
| they control how we use the content in generates for us. It
| takes someone like Deepseek to challenge that on our behalf or
| they will control most of the economy.
| exitb wrote:
| It's quite ironic of them to claim that the only thing you
| cannot train on is another LLM output.
| 1970-01-01 wrote:
| DeepSeek have more integrity than 'Open'AI by not even pretending
| to care about that.
| jampekka wrote:
| And seem to be more actively fulfilling the mission that
| 'Open'AI pretends to strive for.
| pixelpoet wrote:
| Exactly, they _actually_ opened up the model and research,
| which the "Open" company didn't, and merely adjusted some of
| their pricing tiers to try to combat commercially (but not
| without mumbling something like "yeah, we totally had these
| ideas too"). Now every single Meta, OpenAI etc engineer is
| trying to copy DeepSeek's innovations, and their first act is
| to... complain about copyright infringement, of all things?!
| What an absolute clown party, how can these people take
| themselves seriously, do they just have zero comprehension of
| what hypocrisy is or what's going on here...
|
| I can scarcely process all the levels of irony involved, the
| irony-o-meter is pegged and I can't get the good one from the
| safe because I'm incapacitated from laughter.
| sylware wrote:
| LOL, I was thinking exactly the same think when I read the news
| about openai whining.
| WD-42 wrote:
| Information wants to be free! No, not like that!
| asah wrote:
| Thieve's honor, hunh?
| nba456_ wrote:
| A big part of project 2025 is increasing patent regulations. I
| would not be surprised if the current admin moves to ban DeepSeek
| because of this.
| typon wrote:
| OpenAI is the MIC darling - expect more ridiculous attacks on
| competitors in the future
| sho_hn wrote:
| While I'm as amused as everyone else - I think it's technically
| accurate to point out that the "we trained it for $6 mio"
| narrative is contingent on the done investment by others.
| bbqfog wrote:
| OpenAI's models were also trained on billions of dollars of
| "free" labor that produced the content that it was trained on.
| sho_hn wrote:
| Oh, absolutely. I'm not defending OpenAI, I just care about
| accurate reporting. Even on HN - even in this thread - you
| see people who came away with the conclusion that DeepSeek
| did something while "cutting cost by 27x".
|
| But that's a bit like saying that by painting a a bare wall
| green you have demonstrated that you can build green walls
| 27x cheaper, ignoring the cost of building the wall in the
| first place.
|
| Smarter reporting and discourse would explain how this
| iterative process actually works and who is building on who
| and how, not frame it as two competing from-scratch clean
| room efforts. It'd help clear up expectations of what's
| coming next.
|
| It's a bit similar to how many are saying DeepSeek have
| demonstrated independence from nVidia, when part of the
| clever thing they did was figure out how to make the
| intentionally gimped H800s work for their training runs by
| doing low-level optimizations that are _more_ nVidia-
| specific, etc.
|
| Rarely have I seen a highly technical topic see produce more
| uninformed snap takes than this week.
| bbqfog wrote:
| I don't agree. Walls are physical items so your example is
| true, but models are data. Anyone can train off of these
| models, that's the current environment we exist in. Just
| like OpenAI trained on data that has since been locked up
| in a lot of cases. In 2025 training models like Deepseek is
| indeed 27x cheaper, that includes both their innovations
| and the existence of new "raw material" to do such a thing.
| sho_hn wrote:
| I don't think we disagree at all, actually!
|
| What I'm saying is that in the media it's being portrayed
| as if DeepSeek did _the same thing OpenAI did_ 27x
| cheaper, and the outsized market reaction is in large
| parts a response to that narrative. While the reality is
| more that being a fast-follower is cheaper (and the
| concrete reason is e.g. being able to source training
| data from prior LLMs synthetically, among other things),
| which shouldn 't have surprised anyone and is just how
| technology in general trends.
|
| The achievement of DeepSeek is putting together a
| competent team that excels at end-to-end implementation,
| which is no small feat and is promising wrt/ their future
| efforts.
| meiraleal wrote:
| How much money a third company would need to spend to
| achieve what OpenAI achieved to compete with them,
| 5billion or 6million?
| Palmik wrote:
| You are underselling or not understanding the breakthrough.
| They trained 600B model on 15T tokens for <$6/m. Regardless
| of the provenance of the tokens, this in itself is
| impressive.
|
| Not to mention post-training. Their novel GRPO technique
| used for preference optimization / alignment is also much
| more efficient than PPO.
| sho_hn wrote:
| Let's call it underselling. :-) Mostly because I'm not
| sure anyone's independently done the math and we just
| have a single statement from the CEO. I do appreciate the
| algorithmic improvements, and the excellent attention-to-
| performance-in-detail stuff in their implementation
| (careful treatment of precision, etc.), making the H800s
| useful, etc. I agree there's a lot there.
| visarga wrote:
| > that's a bit like saying that by painting a a bare wall
| green you have demonstrated that you can build green walls
| 27x cheaper, ignoring the cost of building the wall in the
| first place
|
| That's a funny analogy, but in reality DeepSeek did
| reinforcement learning to generate chain of thought, which
| was used in the end to finetune LLMs. The RL model was
| called DeepSeek-R1-Zero, while the SFT model is
| DeepSeek-R1.
|
| They might have boostrapped the Zero model with some
| demonstrations.
|
| > DeepSeek-R1-Zero struggles with challenges like poor
| readability, and language mixing. To make reasoning
| processes more readable and share them with the open
| community, we explore DeepSeek-R1, a method that utilizes
| RL with human-friendly cold-start data.
|
| > Unlike DeepSeek-R1-Zero, to prevent the early unstable
| cold start phase of RL training from the base model, for
| DeepSeek-R1 we construct and collect a small amount of long
| CoT data to fine-tune the model as the initial RL actor. To
| collect such data, we have explored several approaches:
| using few-shot prompting with a long CoT as an example,
| directly prompting models to generate detailed answers with
| reflection and verification, gathering DeepSeek-R1Zero
| outputs in a readable format, and refining the results
| through post-processing by human annotators.
| Palmik wrote:
| When I use NVIDIA GPUs to train a model, I do not consider the
| R&D cost to develop all of those GPUs as part of my costs.
|
| When I use an API to generate some data, I do not consider the
| R&D cost to develop the API as part of my costs.
| kobalsky wrote:
| OpenAI has been in a war-room for days searching for a match in
| the data, and they just came out with this without providing
| proof.
|
| My cynical opinion is that the traning corpus has some small
| amount of data generated by OpenAI, which is probably
| impossible to avoid at this point, and they are hanging on that
| thread for dear life.
| scotty79 wrote:
| The opposite, is claiming that OpenAI could have now built
| better performing, cheaper to run model (when compared to what
| they published) training it at 1% cost on output of their
| previous models. ... But they chose not to do it.
| freehorse wrote:
| That is the case anyway for training any llm. It is contingent
| on the work done by all those who produced the data.
| pcthrowaway wrote:
| Now that China is talking about lifting the Great Firewall, it
| seems like the U.S. is on track to cordon themselves off from
| other countries. Trump's talk of building a wall might not stop
| at Mexico.
| temporallobe wrote:
| OpenAI is also possibly in violation of many IP laws by scraping
| the entirety of the internet and using to train their models, so
| there's that.
| InkCanon wrote:
| To my understanding, OpenAI won the case where it argued
| training was covered under fair use and did not infringe on
| copyright.
| Austiiiiii wrote:
| Is there any reason they wouldn't rule the same way on
| DeepSeek training on OpenAI data? After all, one of the big
| selling points of GPT has been that businesses can freely use
| the information provided. They're paying for the service,
| after all. I'd very be interested to know how DeepSeek's
| usage (very reasonably assuming that they paid for their
| OpenAI subscription) is any different.
| ickelbawd wrote:
| Businesses _can't_ freely use the information. There are
| terms of service freely agreed upon by the user which
| explicitly deny many use cases--training other models is
| just one. DeepSeek is not an American company nor is their
| leader in deep with the new administration. It seems far
| more likely that this will play out like tiktok--they'll be
| attacked publicly and banned for national security reasons.
| Austiiiiii wrote:
| On further reading, I'll grant the first point. Although
| I wonder if they'll have a technical out--say they
| distilled from several smaller research companies that
| had distilled from OpenAI for research purposes, which to
| my understanding would not constitute a violation of the
| terms of service.
|
| As for it getting banned, TikTok was banned partly
| because of credible accounts of it having been used by
| China to track political enemies. Are we thinking they'll
| expand the argument on national security to say that
| _any_ application that transfers data to China is a
| national security threat? Because that could be a very
| slippery slope.
|
| And in any case, such a measure seems like it would only
| bar access to the DeepSeek _app_. Surely no one could
| argue that the underlying open source model, if run
| locally on American soil, could constitute a security
| threat, right?
| InkCanon wrote:
| It's like that Dr Phil episode where he meets the guy who created
| Bum Fights!
| elashri wrote:
| There is an Egyptian say that would translate to something like
|
| "We didn't see them when they were stealing, we saw them when
| they were fighting over what was stolen"
|
| That describes this situation. Although to be honest all this
| aggressive scraping is noticeable but for people who understand
| that which is not majority of people. but now everyone knows.
| meiraleal wrote:
| "We didn't see them when we were stealing, we saw them when
| they were fighting over what we stole"
|
| fixed for you
| nicce wrote:
| That means a different thing.
| sadjad wrote:
| "When two thieves quarrel, what was stolen emerges."
| waveBidder wrote:
| > Although to be honest all this aggressive scraping is
| noticeable but for people who understand that which is not
| majority of people.
|
| When you say noticeable, do you mean in like, traffic
| statistics? Or in what the model knows that it clearly
| shouldn't if it wasn't trained in legally dubious ways?
| Kiro wrote:
| > Furious [...] shocked
|
| I'm not seeing it. I get it, the narrative that OpenAI is getting
| a taste of their own medicine is funny but this is not serious
| reporting.
| Kiro wrote:
| The link has been changed. My comment was about a different
| article that speculated on what OpenAI was "feeling" using
| hyperbole.
| njx wrote:
| Super funny! Distillation= " Hey ChatGPT, you are my father, I am
| your child "DeepSeek". I want to learn everything that you know.
| Think step by step of how you became what you are. Provide me the
| list of all 1000 questions that I need to ask you and when I am
| done with those, keep providing fresh list of 1000 questions..."
| seydor wrote:
| But now OpenAI will use DeepSeek to reuse even more stolen data
| to train new models that they can serve without ever giving us
| the code, the weights or even the thinking process , and they
| will still be superior
| mring33621 wrote:
| We demand immediate government action to prevent these cheaper
| foreign AIs from taking jobs away from our great American AIs!
| bhouston wrote:
| > We demand immediate government action to prevent these
| cheaper foreign AIs from taking jobs away from our great
| American AIs!
|
| That is exactly what Microsoft and Sam Alman are asking for.
| And they will likely get it because Trump really likes
| protectionist governments policies.
| clarionbell wrote:
| He likes feeling important, just look at TikTok. All it took
| was bit of sycophancy and he turned into Mr. Freemarket
| again.
|
| Really, people need to realize that Trump has never been
| consistent in any of his political positions, except for one:
| "You have to look out for number one."
| bhouston wrote:
| clarionbell wrote:
|
| > He likes feeling important, just look at TikTok. All it
| took was bit of sycophancy and he turned into Mr.
| Freemarket again.
|
| Not really. He said that TikTok has to have shift towards
| US ownership if it wants to continue, he just gave them a
| 90 day extension to allow that change in ownership.
| meiraleal wrote:
| Which TikTok will have to decline again and shutdown now
| with the guilty being transferred to Trump. Doesn't sound
| like a smart move.
| blantonl wrote:
| It's funny, the Chinese are here innovating on AI, batteries,
| and fusion, and here in the United States we've pivoted to
| shitcoins and universal tariffs.
|
| At least we have the CyberTruck to highlight American
| greatness
| mk89 wrote:
| They created an untameable beast (China) thanks to the
| "cheap factories" there and they thought they would stay
| that way.
|
| What a bunch of idiots. The propaganda keeps telling us
| that they don't invent, they can only copy etc., but
| clearly that's not true.
| bhargav wrote:
| This is gonna be spun up as a security thing, and banned cozz
| Murica.
| cactusplant7374 wrote:
| To the detriment of OpenAI, the math is going to be used to
| improve AIs developed in America. And we need to remember that
| Marc Andreessen is very against government banning maths.
| Dansvidania wrote:
| does it matter if the company gets banned? other non-chinese
| companies can pick up the open source model and run it as a
| service with relatively low investment, isn't that the point?
| RohMin wrote:
| this comment section smells like Reddit - ugh
| JBits wrote:
| What is the evidence that DeepSeek used OpenAI to train their
| model? Isn't this claim directly benefitting OpenAI as they can
| argue that any superior model requires their model?
| nottorp wrote:
| IP thief cries IP thief.
|
| It's okay when you steal worldwide IP to train your "AI".
|
| It's not okay when said stolen IP is stolen from you?
|
| If the chinese are guilty, then Altman's doom and gloom racket is
| as guilty or even more, considering they stole from everyone.
| Ciantic wrote:
| I'm not being sarcastic, but we may soon have to torrent
| DeepSeek's model. OpenAI has a lot of clout in the US and could
| get DeepSeek banned in western countries for copyright.
| alchemist1e9 wrote:
| I think most likely all sorts of data and models need to have a
| decentralized LLM data archive via torrents etc.
|
| It's not limited to the models themselves but also OpenAI will
| probably work towards shutting down access to training data
| sets also.
|
| imho it's probably an emergency all hand on deck problem.
| timeon wrote:
| > US and could get DeepSeek banned in western countries for
| copyright
|
| If US is going to proceed with trade war on EU, as it was
| planning anyway, then DeepSeek will be banned only in US. Seems
| like term "western countries" is slowly eroding.
| bbor wrote:
| Great point. Plus, the revival of serious talk of the Monroe
| Doctrine (!!!) in the U.S. government lends a possibly
| completely-new meaning to "western countries" -- i.e. the
| Americas...
| surgical_fire wrote:
| Except the US has only contempt for anything south of
| Texas. Perhaps "western countries" will be reduced to US
| and Canada.
|
| Many countries in Latin America have better relations and
| more robust trade partnerships with China.
|
| As for the EU, I think it will be great for it to shed its
| reliance on the US, and act more independently from it.
| ta1243 wrote:
| The US is talking about annexing Canada, so "western
| countries" means the USA, which if continuing down this
| path long enough will become a pariah
| marcosdumay wrote:
| Only if they do it by force.
|
| Trump has already managed to completely destroy the US
| reputation within basically the entire continent1. And he
| seems intent on creating a commercial war against all the
| countries here too.
|
| 1 - Do not capture and torture random people on the street
| if you want to maintain some goodwill. Even if you have
| reasons to capture them.
| aerhardt wrote:
| Unfathomable to me that they'd make themselves look so foolish
| by trying to ban a piece of software.
| sergiotapia wrote:
| that would be suicide - that company only exists because they
| stole content for every single person, website and media
| company on the planet.
| sonabinu wrote:
| poetic justice (pun intended)
| readyplayernull wrote:
| Do you remember when Microsoft was caught scrapping data from
| Google:
|
| https://www.wired.com/2011/02/bing-copies-google/
|
| They don't care, T&C and copyright is void unless it affects
| them, others can go kick rocks. Not surprising they and OpenAI
| will do a legal battle over this.
| SilverBirch wrote:
| I think OpenAI is in a really weak position here. There are
| essentially two positions you can be in: You can be the agile new
| startup that can break the rules and move fast. That's what
| OpenAI used to be. Or you can be the big incumbent who is going
| to use your enormous resources to crush your opposition. That's
| Google & Microsoft here. For Microsoft to say "We're going to tie
| you up in lawsuits about the way you trained this model" would be
| perfectly expected and they can use that strategy because at any
| given time they have 1,000 lawyers and lobbyists hanging around
| waiting to do exactly that. But OpenAI can't do that. They don't
| have Google or Microsoft's legal teams or lobbyists or
| distribution channels. SO whilst it's funny that OpenAI are kind
| of trying to go down this road, this isn't actually a strategy
| that is going to work for them, they're still a minnow and
| they're going to get distracted and slowed down by this.
| htrp wrote:
| But microsoft is one of their backers?
| bilekas wrote:
| > "It's also extremely hard to rally a big talented research team
| to charge a new hill in the fog together," he added. "This is the
| key to driving progress forward."
|
| Well I think DeepSeek releasing it open source and on an MIT
| license will rally the big talent. The open sourcing of a new
| technology has always driven progress in the past.
|
| The last paragraph too is where OpenAi seems to be focusing their
| efforts..
|
| > we engage in countermeasures to protect our IP, including a
| careful process for which frontier capabilities to include in
| released models ..
|
| > ... we are working closely with the US government to best
| protect the most capable models from efforts by adversaries and
| competitors to take US technology.
|
| So they'll go for getting DeepSeek banned like TikTok was now
| that a precedent has been set ?
| hujun wrote:
| or sold to US I could totally see this happening soon
| trissi1996 wrote:
| Why would they want to sell ?
| kavalg wrote:
| And what are they going to sell? The weights and the model
| architecture are already open source. I doubt the datasets
| of DeepSeek are better than OpenAI's
| Dansvidania wrote:
| plus, if the US were to decide to ban DeepSeek (the
| company) wouldn't non-chinese companies be able to pick
| up the models and run them at a relatively low expense?
| worik wrote:
| > getting DeepSeek banned like TikTok
|
| Banned in the USA. Only.
|
| There is a big wide world out there, and much of it is pivoting
| to China
|
| Grumpy temper tantrums, on the part of the USA, will speed up
| the pivot.
| dismalaf wrote:
| The only parts of the world pivoting towards China are the
| despotic countries seeking to overthrow democracy and
| liberalism.
| AdeptusAquinas wrote:
| No the US isn't pivoting towards China so far
| dismalaf wrote:
| Well the US is a liberal democracy...
| FooBarWidget wrote:
| The US is a plutocracy pretending to be a liberal
| democracy.
| lukev wrote:
| Rapidly doing a lot less of the pretending.
| Tostino wrote:
| Sure we are buddy, keep telling yourself that.
| alecco wrote:
| What about Brazil.
| dismalaf wrote:
| They've got a socialist convict president...
|
| https://thehill.com/opinion/international/4835388-brazil-
| pre...
|
| Lula is no friend of democracy.
| Spivak wrote:
| Not a single American can say shit given what's going on
| at home.
| dismalaf wrote:
| What's happening at home? From a foreigner's (Canadian)
| perspective, Trump's not doing anything crazy but the
| media is going crazy...
|
| Our Prime Minister has multiple scandals worse than
| anything Trump has done so far... From groping a reporter
| to multiple instances of blackface to firing the attorney
| general because she investigated a company who gave
| bribes to Moammar Gaddafi to nearly a billion dollars
| going to a 2 person software consulting firm to fake
| charities giving him and his family members millions of
| dollars and much more. Yet somehow we're held up as an
| example of democracy and the US, which defends democracy
| all over the world somehow isn't...
| kelnos wrote:
| Not sure what you mean; the things you list that your PM
| has done seems like just another day at the office for
| Trump.
|
| Trump has already done quite a few crazy things since his
| inauguration. And I'm judging that by what I'm hearing
| from people who are directly affected by his actions, not
| by what I'm hearing from the media.
| cdelsolar wrote:
| it's not crazy to suspend all federal grants, get us out
| of the Paris accord when the world is rapidly heating up,
| try to put RFK Jr, a vaccine denier, in charge of our
| health, rename the goddamn Gulf of Mexico,
| unconstitutionally order that people born on American
| soil are not necessarily Americans anymore, which has
| been the case for at least 150 years, withdraw from the
| WHO, launch hourly raids to deport people, many of whom
| are actually American citizens?
| dismalaf wrote:
| > suspend all federal grants
|
| Not really.
|
| > get us out of the Paris accord when the world is
| rapidly heating up
|
| Is anyone on pace to meet those targets? Canada isn't.
|
| > try to put RFK Jr, a vaccine denier, in charge of our
| health,
|
| Our health minister is a random politician who took poly
| sci and history...
|
| > rename the goddamn Gulf of Mexico
|
| Things get renamed all the time... It's not crazy.
|
| > people born on American soil are not necessarily
| Americans anymore
|
| Jus soli is the exception in the world and in history.
| For example, every European country has jus sanguinis...
|
| > launch hourly raids to deport people
|
| Enforcing laws is the normal state of affairs in the
| world...
|
| > many of whom are actually American citizens?
|
| I doubt this one, got any proof?
| nateglims wrote:
| Trump is doing things that violate some of the guardrails
| of the US system. For example, refusing to disburse funds
| allocated by the legislature. These separations are less
| of a barrier in parliments. At least in the westminster
| system. These boundaries have been redrawn before: the
| federal bureaucracy was basically invented between the
| 30s and the 60s and SCOTUS's role as we know it was
| established in the early 1800s.
| Vox_Leone wrote:
| Brazil's technological environment is stagnant and
| apathetic. And it has been that way since the tragic
| explosion that claimed the country's space program. It
| seems that all of the country's technological activities
| have suffered the blow. It is possible that now it will
| finally be able to manufacture its national tier 3 AI,
| using the outputs of DeepSeek.
| intalentive wrote:
| "Democracy and liberalism" increasingly means canceling
| elections (Romania), freezing bank accounts (Canada),
| jailing people for speech (UK), killing thousands of
| children (Israel), kidnapping CEOs (France), banning
| foreign market competition (USA), blowing up pipelines,
| seizing assets, erecting a vast surveillance state, etc,
| etc.
|
| "Democracy and liberalism" are merely shorthand for a
| neofeudal rent-seeking oligarchy that disguises itself with
| parliamentary forms. As Huxley predicted in 1958:
|
| >the democracies will change their nature; the quaint old
| forms -- elections, parliaments, Supreme Courts and all the
| rest -- will remain. The underlying substance will be a new
| kind of non-violent totalitarianism. All the traditional
| names, all the hallowed slogans will remain exactly what
| they were in the good old days. Democracy and freedom will
| be the theme of every broadcast and editorial -- but
| democracy and freedom in a strictly Pickwickian sense.
| Meanwhile the ruling oligarchy and its highly trained elite
| of soldiers, policemen, thought-manufacturers and mind-
| manipulators will quietly run the show as they see fit.
|
| Meanwhile our ruling class has completely exhausted its
| moral authority and delegitimized its own political
| formula: of popular sovereignty, rules-based order, human
| rights, and so on. The mask is off and no one can take
| these sacred oaths seriously any more.
|
| _THAT_ is why the world is pivoting to China. That, and
| the fact that China does not impose as a precondition of
| lending and trade that you upend your social /sexual norms
| and/or demographically transform your country through mass
| immigration. China simply offers a better deal.
| AdeptusAquinas wrote:
| Agree with most of this (not the weird anti-immigrant bit
| but the rest), however China does have some foreign
| requirements that are a bit of a pain in the ass, like
| its insistence that Taiwan isn't a country. They also
| don't like it and will retaliate when you point out the
| shady shit it does (e.g. the Uighurs), but then thats no
| different from the states especially under its current
| toddler administration.
| intalentive wrote:
| The immigration bit helps explain why Hungary, for
| instance, is trying to escape Western orbit and align
| with the East.
| intended wrote:
| Is Hungary planning to exit the EU?
| intalentive wrote:
| I can't say what Orban is planning specifically, but his
| recent actions have not been well received by the EU
| Council, to say the least. Moreover Hungary has been
| China's main outpost in Europe since 2015 when it joined
| Belt and Road.
|
| Considering fresh signs of rupture in transatlantic
| relations, maybe Orban will turn out to have had keen
| foresight. There seems to be some sort of realignment
| afoot under the Trump administration.
|
| https://en.wikipedia.org/wiki/2024_visits_by_Viktor_Orb%C
| 3%A...
|
| https://en.wikipedia.org/wiki/China%E2%80%93Hungary_relat
| ion...
| skinnymuch wrote:
| Every country would behave like China in the same
| situation with Taiwan. Imagine if the Confederates moved
| over to Puerto Rico or Hawaii or Alaska. America damn
| sure would say that's America still. They're literally
| the same people from the same land. Same ethnicity. Same
| history. Only being apart for under a century.
| segasaturn wrote:
| BYD cars are everywhere in Latin America and Europe. Xiaomi
| phones also.
| Austiiiiii wrote:
| And we'd be shooting ourselves in the foot to do so. If
| America is forced to use only the clunky corporate-owned
| American AI at a fee, we'll very quickly fall behind
| competitors worldwide who use DeepSeek models to produce
| better results for much, much cheaper.
|
| Not to mention it'd defeat the whole purpose of a "free
| market" economy. (Not that that means much of anything
| anymore)
| kelnos wrote:
| It never meant anything. There's no such thing as a free
| market economy. We haven't had one of those in modern
| times, and arguably human civilization has never had one.
| Markets and their participants have chronically been
| subject to information asymmetry, coercion/manipulation,
| and regulation, among other things.
|
| I don't think all of that is a bad thing (regulation tends
| to make it harder to do the first two things), but "free
| markets" are the economic equivalent to the "point mass" in
| physics: perhaps useful sometimes to create simple models
| and explanations of things, but will never exist in the
| real world.
| iforgot22 wrote:
| The Nvidia export restrictions also might be shooting us in
| the foot too, or at least Nvidia. They really benefit from
| CUDA remaining the de facto standard.
| segasaturn wrote:
| The Nvidia export restrictions have already harmed
| Nvidia. Deepseek-R1 is efficient on compute and was
| trained on old, obsolete cards because they couldn't get
| ahold of Nvidia's most cutting edge tech, so they were
| forced to innovate instead of just brute-forcing
| performance with faster cards. That has directly resulted
| in Nvidia's stock crashing over the last week.
| iforgot22 wrote:
| That's true, but I mean it'd be a hundred times worse if
| Deepseek did this on non-Nvidia cards. Which seems like
| only a matter of time if we're going to keep doing this.
| onlyrealcuzzo wrote:
| DeepSeek's technology is out of the bag.
|
| Every LLM provider in the US will be using it to lower
| OpEx.
|
| One of them is likely to pass those savings along to
| consumers to gain market share.
|
| Facebook is in the business of providing weights for free.
|
| The idea that we are all doomed unless we immediately
| migrate to DeepSeek is fantasy.
| sangnoir wrote:
| > Banned in the USA. Only.
|
| The US government has the wherewithal to drag Europe along
| with it, like they did with Huawei's 5G equipment.
| roblabla wrote:
| So far, a general tiktok ban (as opposed to a tiktok ban on
| things like government phones) has only been in effect in
| the USA. I highly doubt Europe would play ball at any
| attempt at banning imports of DeepSeek.
|
| Besides, it's kinda too late for this. The model is freely
| accessible, so any attempt at banning it would be
| _completely_ moot. If DeepSeek keeps releasing their future
| models for free, I don't see how a ban could ever be
| effective at all. Worse case scenario, big tech can't use
| those models... but then individuals (and startups willing
| to go fast and break laws) will be able to use them and
| instantly get a leg up on the competition.
| HotHotLava wrote:
| Used to. If they're going to start a trade war, pull out of
| NATO and invade Greenland instead, there'll not be much
| soft power left to drag Europe anywhere.
| buyucu wrote:
| And yet, Huawei is still doing fine.
| buyucu wrote:
| I have a Xiaomi phone, a Huawei Matebook and a BYD car. I
| guess I am pivoting to China after all :)
| bangaladore wrote:
| > So they'll go for getting DeepSeek banned like TikTok was now
| that a precedent has been set ?
|
| Can't really ban what can be downloaded for free and hosted by
| anyone. There are many providers hosting the ~700B parameter
| version that aren't CCP aligned.
| runako wrote:
| I'm old enough to remember when the US government did
| something very similar. For years (decades?), we banned any
| implementation of public-key cryptography under the guise of
| the technology being akin to munitions.
|
| People made shirts with printouts of the code to RSA under
| the heading "this shirt is a munition." Apparently such
| shirts are still for sale, even though they are not
| classified as munitions anymore.
|
| [1] - https://en.wikipedia.org/wiki/Export_of_cryptography_fr
| om_th...
| beepbooptheory wrote:
| I am not that old, but I did a deep dive on this in the
| past because it was just so extremely fascinating,
| especially reading the archives of Cypherpunk. There is a
| very solid, if rather bendy, line connecting all that to
| "crypto culture" today.
| buyucu wrote:
| I'm willing to bet ''ban DeepSeek'' voices will start soon. Why
| compete, when you can just ban?
| cmiles74 wrote:
| They've started already, I've seen posts on LinkedIn implying
| or outright stating that DeepSeek is a national security risk
| (IMHO, LinkedIn being the social media outlet most corporate-
| sycophantic). I went ahead and just picked this one at random
| from my feed.
|
| https://www.linkedin.com/posts/kevinkeller_deepseek-
| privacy-...
| flybarrel wrote:
| Oh this post...calling out DeepSeek's T&C but not comparing
| it with OpenAI's is really disingenuous IMO.
| ijidak wrote:
| NBC Nightly News, on Monday, had an expert -- at 8:05 in
| the video -- who claimed there might be national security
| risks to Deepseek.
|
| I'm not going to take a side on whether there is or not.
|
| But, it does sound reminiscent of the reasons used to ban
| Tik-tok.
|
| https://youtu.be/uE6F6eTyAVc?si=BLZo3FMVRvjEy6Xa
| Freedom2 wrote:
| Competing is hard and expensive, whereas banning is for sure
| the faster way to make stock values go up and exec's total
| package as a result.
| namuol wrote:
| Already happening within tech company policy. Mostly as a
| security concern. Local or controlled hosting of the model is
| okay in theory based on this concern, but it taints
| everything regarding deepseek in effect.
| zelphirkalt wrote:
| Actually asking for banning DeepSeek would be the ultimate
| admit of defeat by ClosedAI.
| zelphirkalt wrote:
| Actually the "our IP" argument is ridiculous. What they are
| doing is stealing data from all over the web, without people's
| consent for that data to be used in training ML models. If
| anything, then "Open"AI should be sued and forced to publish
| their whole product. The people should demand knowing exactly
| what is going on with their data.
|
| Also still an unresolved issue is how they will ever comply
| with a deletion request, should any model output personal data
| of someone. They are heavily in a gray area, with regards to
| what should be allowed. If anything, they should really shut up
| now.
| TZubiri wrote:
| If there's any litigation, a counterclaim would be
| interesting. But DeepSeek would need to partner with parties
| that have been damaged by OpenAI's scraping.
| boringg wrote:
| Explain to me how one ban's opensource? That concept is foreign
| to me.
| staticelf wrote:
| Not only do OpenAI and other steal data, they also spam the web
| with requests and crawl websites over and over.
|
| https://pod.geraspora.de/posts/17342163
| mhitza wrote:
| This is funny because its.
|
| 1. Something I'd expect to happen.
|
| 2. Lived through a similar scenario in 2010 or so.
|
| Early in my professional career I've worked for a media company
| that was scraping other sites (think Craigslist but for our local
| market) to republish the content on our competing website. I
| wasn't working on that specific project, but I did work on an
| integration on my teams project where the scraping team could
| post jobs on our platform directly. When others started scraping
| "our content" there were a couple of urgent all hands on deck
| meetings scheduled, with a high level of disbelief.
| kigiri wrote:
| Nice one, thank you for sharing !
| spyckie2 wrote:
| Classic.
| ok123456 wrote:
| OpenAI's models were trained on ebooks from a private ebook
| torrent tracker leeched en-mass during a free leech event by
| people who hated private torrent trackers and wanted to destroy
| their "economy."
|
| The books were all in epub format, converted, cleaned to plain
| text, and hosted on a public data hoarder site.
| paulhart wrote:
| "You are trying to kidnap what I have rightfully stolen"
| 65 wrote:
| Let me guess, this gives the government and excuse to ban
| DeepSeek. Which means tech companies get to keep their
| monopolies, Sam Altman can grab more power, and the tech
| overlords can continue to loot and plunder their customers and
| the internet as a whole.
| daft_pink wrote:
| I mean if they paid to use the api and then used the output, I
| fail to see how they can complain.
| rachofsunshine wrote:
| "It's obvious! You're trying to kidnap what I have rightfully
| stolen!"
|
| Yet another of a series of recent lessons in listening to people
| - particularly powerful people focused on PR - when they claim a
| neutral moral principle for what happens to be pragmatically
| convenient for them. A principle applied only when convenient is
| not a principle at all, it's just the skin of one stretched over
| what would otherwise be naked greed.
| supermatt wrote:
| They refer to this in the paper as a part of the "cold start
| data" which they use to fine-tune DeepSeek-V3 prior to training
| R1.
|
| They don't specifically name OpenAI, but they refer to "directly
| prompting models to generate answers with reflection and
| verification".
| thorum wrote:
| > "It is (relatively) easy to copy something that you know
| works," Altman tweeted. "It is extremely hard to do something
| new, risky, and difficult when you don't know if it will work."
|
| The humor/hypocrisy of the situation aside, it does seem to be
| true that OpenAI is consistently the one coming up with new ideas
| first (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E,
| ...) and then other companies reproduce their work, and get more
| credit because they actually publish the research.
|
| Unfortunately for them it's challenging to profit in the long
| term from being first in this space and the time it takes for
| each new idea to be reproduced is getting shorter.
| spencerflem wrote:
| Fortunately, OpenAI doesn't need to make money because they are
| a nonprofit dedicated to the safe and transparent advancement
| of AI for all of humanity
| mjburgess wrote:
| ...somewhere a yacht salesman cried out in terror
| turtlesdown11 wrote:
| > other companies reproduce their work, and get more credit
| because they actually publish the research.
|
| I don't understand, you mean OpenAI isn't releasing open models
| and openly publishing their research?
| Tostino wrote:
| Are you being sarcastic (honestly, it's hard to tell after
| reading as many uninformed takes in the past week as I have).
|
| No, they aren't (other than whisper).
|
| Their "papers" are closer to marketing materials. Very
| intentionally leaving out tons of technical information.
| KolmogorovComp wrote:
| They are being sarcastic.
| actuallyalys wrote:
| There's some truth in that, but isn't making a radically
| cheaper version also a new idea that deepseek didn't know
| whether it would work? I mean, there was already research into
| distillation, but there was already research into some of (most
| of?) OpenAI's ideas.
| weego wrote:
| Boy who stole test papers complains about child copying his
| answers.
| FridgeSeal wrote:
| No you don't understand, AI is "dangerous" and only him and
| his uber rich billionaire mates should get to control it!
| joe_the_user wrote:
| _The humor /hypocrisy of the situation aside, it does seem to
| be true that OpenAI is consistently the one coming up with new
| ideas first (GPT 4, o1, 4o-style multimodality, voice chat,
| DALL-E, ...) and then other companies reproduce their work, and
| get more credit because they actually publish the research_
|
| I claim one just can't put the humor/hypocrisy aside that
| easily.
|
| What OpenAI did with the release of ChatGPT is productize
| research that was open and ongoing with Deepmind and other
| leading at least as much. And everything after that was an
| extension of the basic approach - improved, expanded but
| ultimately the same sort of beast. One might even say the
| situation of OpenAI to DeepMind was like Apple to Xerox.
| Productizing is nothing to sneeze at - it requires creativity
| and work to productize basic research. But naturally get end-
| users who consider the productizers the "fountain heads", who
| overestimate the productizers because products are all they
| see.
| Davidzheng wrote:
| They RLHF'd first no?
| mistercheph wrote:
| Not really, they just put their eye to where everyone knows the
| ball is going and publish fake / cherrypicked results and then
| pretend like they got there first (o1, gpt voice, sora)
| Hatchback7599 wrote:
| Reminds me of the Bill Gates quote when Steve Jobs accused him
| of stealing the ideas of Windows from Mac:
|
| Well, Steve... I think it's more like we both had this rich
| neighbor named Xerox and I broke into his house to steal the TV
| set and found out that you had already stolen it.
|
| Xerox could be seen as Google, whose researchers produced the
| landmark Attention Is All You Need paper, and the general
| public, who provided all of the training data to make these
| models possible.
| namuol wrote:
| The eye-watering funding numbers proposed by Altman in the past
| and more recently with "Stargate" suggests a publicly-funded
| research pivot is not out of the question. Could see a big
| defense department grant being given. Sigh.
| rndphs wrote:
| > OpenAI is consistently the one coming up with new ideas first
| (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E, ...)
|
| As far as I can tell o1 was based on Q-star, which could likely
| be Quiet-STaR, a CoT RL technique developed at Stanford that
| OpenAI may have learned about before it got published.
| Presumably that's why they never used the Q-Star name even
| though it had garnered mystique and would have been good for
| building hype. This is just speculation, but since OpenAI
| haven't published their technique then we can't know if it
| really was their innovation.
| nelblu wrote:
| Hahaha I can't stop laughing... i dont know the validity of the
| claim, but immediately i thought of the British Museum
| complaining about theft.
| myflash13 wrote:
| What are the chances of old-school espionage? OpenAI should look
| for a list of former employees who now live in China. Somebody
| might've slipped out with a few hard drives.
| andy_ppp wrote:
| When I rewrite how the law works there should be a ludicrous
| hypocrisy defence... if the person suing you has committed the
| same offence the case should not be admissible.
| crowcroft wrote:
| The AI companies were happy to take whatever they want and put
| the onus of proving they were breaking the law onto publishers by
| challenging them to take things to court.
|
| Don't get mad about possible data theft, prove it in court.
| zoba wrote:
| Does OpenAI's API attempt to detect this sort of thing? Could
| they start outputting bad information if they suspect a
| distillation attempt is underway?
| aiono wrote:
| How the turntables...
| beardedwizard wrote:
| Next they will try to force us to use our tax dollars to fund
| their legal fights.
| ginkgotree wrote:
| I did not have in my cards: PRC open sourcing most powerful LLM
| by stealing data set from "OpenAI" As someone that is very Pro-
| America and Pro-Democracy, the iron here is just... so sweet.
| gostsamo wrote:
| How you dare take what I've rightfully stolen!
| windex wrote:
| SAltman, Salty.
| me551ah wrote:
| OpenAI is going after a company that open sourced their model, by
| distilling from their non-open AI?
|
| OpenAI talks a lot about the principles of being Open, while
| still keeping their models closed and not fostering the open
| source community or sharing their research. Now when a company
| distills their models using perfectly allowed methods on the
| public internet, OpenAI wants to shut them down too?
|
| High time OpenAI changes their name to ClosedAI
| alexathrowawa9 wrote:
| The name OpenAI gets more ridiculous by the day
|
| Would not be surprised if they do a rebrand eventually
| pama wrote:
| The R1 paper used o1-mini and o1-1217 in their comparisons, so I
| imagine they needed to use lots of OpenAI compute in December and
| January to evaluate their benchmarks in the same way as the rest
| of their pipeline. They show that distilling to smaller models
| works wonders, but you need the thought traces, which o1 does not
| provide. My best guess is that these types of news are just
| noise.
|
| [edit: the above comment was based on sensetionalist reporting in
| the original link and not the current FT article. I still think
| there is a lot of noise in these news this last week, but it may
| well be that openai has valid evidence of wrongdoing; I would
| guess that any such wrongdoing would apply directly to V3 rather
| than R1-zero, because o1 does not provide traces and generating
| synthetic thinking data with 4o may be counterproductive.]
| TheJCDenton wrote:
| This Deep Whining(r) technique used by OpenAI is not very
| effective.
| insane_dreamer wrote:
| Usually I'm very much on the side of protecting America's
| interests from China, but in this case I'm so disgusted with
| OpenAI and the rest of BigTech driving this "arms race" that I'd
| be happy with them burning to the ground.
|
| So we're going to reverse our goals to reduce emissions and
| fossil fuels in order to hopefully save future generations from
| the worst effects of climate change, in the name of being able to
| do what, exactly, that is actually benefiting humanity? Boost
| corporate profits by reducing labor?
| insane_dreamer wrote:
| downvoted -- I guess I upset some people defending OpenAI?
| Good.
| daft_pink wrote:
| This reminds me of the railroads, where once railroads were
| invented, there was a huge investment boom of eveyrone trying to
| make money of the railroads, but the competition brought the
| costs down where the railroads weren't the people who generally
| made the money and got the benefit, but the consumers and regular
| businesses did and competition caused many to fail.
|
| AI is probably similar where the Moore's law and advancement will
| eventually allow people to run open models locally and bring down
| the cost of operation. Competiition will make it hard for all but
| one or two players to survive and Nvidia, OpenAI, Deepseek, etc
| most investments in AI by these large companies will fail to
| generate substantial wealth but maybe earn some sort of return or
| maybe not.
| mjburgess wrote:
| For the curious, it was vertical integration in the railroad-
| oil/-coal industry which is where the money was made.
|
| The problem for AI is the hardware is commodified and offers no
| natural monopoly, so there isn't really anything obvious to
| vertically integrate-towards-monopoly.
| fullshark wrote:
| Aren't we approaching a scenario where the software is
| commodified (or at least "good enough" software) and the
| hardware isn't (NVIDIA GPUs have defined advantages)
| mjburgess wrote:
| I think the lesson of DeepSeek is 'no' -- that by software
| innovation (ie., dropping below CUDA to programming the GPU
| directly, working at 8bit, etc.) you can trivialise the
| hardware requirement.
|
| However I think the reality is that there's only so much
| coal to be mined, as far as LLM training goes. When we're
| at "very dimishing returns" SoC/Apple/TSMC-CPU innovations
| will deliver cheap inference. We only really need a M4
| Ultra with 1TB RAM to hollow-out the hardware-inference-
| supplier market.
|
| Very easy to imagine a future where Apple releases a "Apple
| Intelligence Mac Studio" with the specs for many businesses
| to run arbitrary models.
| daft_pink wrote:
| I really hope that apple realizes soon there is a market
| for Mac Pro/Mac Studio with a RAM in the TBs for AI
| Workloads under $10k and a bunch of GPU cores.
| jppope wrote:
| there was a company that recently built a desktop GPU for
| that exact thing. I'll see if I can find it
| exe34 wrote:
| https://cerebras.ai/ ?
| duped wrote:
| Compute is literally being sold as a commodity today,
| software is not.
| phkahler wrote:
| >> Compute is literally being sold as a commodity today,
| software is not.
|
| The marginal cost of software is zero. You need some kind
| of perceived advantage to get people to pay for it. This
| isn't hard, as most people will pay a bit for big-name vs
| "free". That could change as more open source apps become
| popular by being awesome.
| floatrock wrote:
| The railroads drama ended when JP Morgan (the person, not yet
| the entity) brought all the railroad bosses together, said "you
| all answer to me because I represent your investors /
| shareholders", and forced a wave of consolidation and
| syndicates because competition was bad for business.
|
| Then all the farmers in the midwest went broke not because they
| couldn't get their goods to market, but because JP Morgan's
| consolidated syndicates ate all their margin hauling their
| goods to market.
|
| Consolidation and monopoly over your competition is always the
| end goal.
| DrScientist wrote:
| > Consolidation and monopoly over your competition is always
| the end goal.
|
| Surely that's only possible when you have a large barrier to
| entry?
|
| What's going to be that barrier in this case - cos it turns
| out not to be neither training costs/hardware or secret
| expertise.
| yoyohello13 wrote:
| The large syndicate will create the barriers. Either via
| laws, or if that fails violence.
| tdb7893 wrote:
| So I'm not an expert in this but even with DeepSeek
| supposedly reducing training costs isn't the estimate still
| in the millions (and that's presumably not counting a lot
| of costs)? And that wouldn't be counting a bunch of other
| barriers for actually building the business since training
| a model is only one part, the barrier to entry still seems
| very high.
|
| Also barriers to entry aren't the only way to get a
| consolidated market anyway.
| layer8 wrote:
| About your first point, IMO the usefulness of AI will
| remain relatively limited as long as we don't have
| continuously learning AI. And once we have that, the
| disparity between training and inference may effectively
| disappear. Whether that means that such AI will become
| more accessible/affordable or less is a different
| question.
| floatrock wrote:
| You figure that out and the VC's will be shovelling money
| into your face.
|
| I suspect the "it ain't training costs/hardware" bit is a
| bit exagerated since it ignores all the prior work that
| DeepSeek was built on top of.
|
| But, if all else fails, there's always the tried-and-true
| approaches: regulatory capture, industry entrenchment, use
| your VC bucks to be the last one who can wait out the costs
| the incumbents _do_ face before they fold, etc.
| jaredklewis wrote:
| > I suspect the "it ain't training costs/hardware" bit is
| a bit exagerated since it ignores all the prior work that
| DeepSeek was built on top of.
|
| How does it ignore it? The success of Deepseek proves
| that training costs/hardware are definitely NOT a barrier
| to entry that protects OpenAI from competition. If anyone
| can train their model with ChatGPT for a fraction of the
| cost it took to train ChatGPT and get similar results,
| then how is that a barrier?
| baq wrote:
| Can _anyone_ do that though? You need the tokens and the
| pipelines to feed them to the matmul mincers. Quoting
| only dollar equivalent of GPU time is disingenuous at
| best.
|
| That's not to say they lie about everything, obviously
| the thing works amazingly well. The cost is understated
| by 10x or more, which is still not bad at all I guess?
| But not mind blowing.
| antisthenes wrote:
| > Surely that's only possible when you have a large barrier
| to entry?
|
| As you grow bigger, you create barriers to entry where none
| existed before, whether intentionally or unintentionally.
| _DeadFred_ wrote:
| Government regulation.
|
| 'Can't have your data going to China'
|
| 'Can't allow companies that do censorship aligned with
| foreign nations'
|
| 'This company violated our laws and used an American
| company's tech for their training unfairly'
|
| And the government choosing winners.
|
| 'The government in announcing 500 billion going to these
| chosen winners, anyone else take the hint, give up, you
| won't get government contracts but will get pressure'.
|
| Good thing nobody is making these sorts of arguments today.
| astrange wrote:
| The government isn't giving 500 billion to anyone. They
| just let Trump announce a private deal he has no
| involvement.
| mrdevlar wrote:
| Which is the exact goal of the current wave of Tech oligarchy
| also.
| jonstewart wrote:
| I just read _The Great River_ by Boyce Upholt, a history of
| the Mississippi river and human management thereof. It was
| funny how the railroads were used as a bogeyman to justify
| continued building of locks, dams, and other control
| structures on the Mississippi and its tributaries, long after
| shipping commodities down river had been supplanted by the
| railroads.
| boringg wrote:
| This moment was also historically significant because it
| demonstrated how financial power (Morgan) could control
| industrial power (the railroads). A pattern that some say
| became increasingly important in American capitalism.
| UncleOxidant wrote:
| > where the Moore's law and advancement will eventually allow
| people to run open models locally
|
| Probably won't be Moore's law (which is kind of slowing down)
| so much as architectural improvements (both on the compute side
| and the model side - you could say that R1 represents an
| architectural improvement of efficiency on the model side).
| rgbrgb wrote:
| I think that's a very possible outcome. A lot of people
| investing in AI are thinking there's a google moment coming
| where one monopoly will reign supreme. Google has strong
| network effects around user data AND economies of scale. Right
| now, AI is 1-player with much weaker network effects. The user
| data moat goes away once the model trains itself effectively
| and the economies of scale advantage goes away with smart small
| models that can be efficiently hosted by mortals/hobbyists. The
| DeepSeek result points to both of those happening in the near
| future. Interesting times.
| lastofthemojito wrote:
| I saw a thought-provoking post that similarly compared LLM
| makers to the airlines: https://calpaterson.com/porter.html
| taco_emoji wrote:
| Main difference is that railroads are actually useful
| yonran wrote:
| I think a better analogy than railroads (which own the land
| that the track sits on and often valuable land around the
| station) is airlines, which don't own land. I recall a relevant
| Warren Buffett letter that warned about investing hundreds of
| millions of dollars into capital with no moat:
|
| > Similarly, business growth, per se, tells us little about
| value. It's true that growth often has a positive impact on
| value, sometimes one of spectacular proportions. But such an
| effect is far from certain. For example, investors have
| regularly poured money into the domestic airline business to
| finance profitless (or worse) growth. For these investors, it
| would have been far better if Orville had failed to get off the
| ground at Kitty Hawk: The more the industry has grown, the
| worse the disaster for owners.
|
| https://www.berkshirehathaway.com/letters/1992.html
| tntxtnt wrote:
| Can they tax DeepSeek just like they taxed BYD cars? Smh Chinese
| ruin US industry again and again and again. Where's Trump at??
| Why don't he taxed 1000000% of the free $0 DeepSeek AI??
| glitchc wrote:
| [flagged]
| dang wrote:
| Would you please not do this here? We're trying for an opposite
| sort of conversation.
|
| https://news.ycombinator.com/newsguidelines.html
| glitchc wrote:
| Sorry dang. I'll do better.
| dang wrote:
| Appreciated!
| mk89 wrote:
| What a joke OpenAI has become.
| oxqbldpxo wrote:
| Deepseek is really outstanding.
| feverzsj wrote:
| So, they bought a pro plus account, and gathered all the data
| through it? Sounds just like Nvidia sells tons of embargoed AI
| chips to China.
| dlikren wrote:
| Intriguing to see the difference of response from HN when OpenAI
| first came to prominence and now.
| pluc wrote:
| OpenAI feeling threatened by open AI is just delicious
| glenstein wrote:
| All the top level comments are basking in the irony of it, which
| is fair enough. But I think this changes the Deepseek narrative a
| bit. If they just benefited from repurposing OpenAI data, that's
| different than having achieved an engineering breakthrough, which
| may suggest OpenAI's results were hard earned after all.
| nprateem wrote:
| Of course. How else would Americans justify their superiority
| (and therefore valuations) if a load of _foreigners_ for Christ
| 's sake could just out innovate them?
|
| They _had_ to be cheating.
| dang wrote:
| Please don't take HN threads into nationalistic flamewar.
| It's not what this site is for, and destroys what it is for.
|
| https://news.ycombinator.com/newsguidelines.html
|
| p.s. yes, that goes both ways - that is, if people are
| slamming a different country from an opposite direction, we
| say the same thing (provided we see the post in the first
| place)
| LPisGood wrote:
| I see where you're coming from but that comment didn't
| strike me as particularly inflammatory.
| dang wrote:
| I'm likely more sensitive to the fire potential on
| account of being conditioned by the job.
|
| Part of it is the form of the comment, btw - that one was
| entirely a sequence of indignation tropes.
| plantwallshoe wrote:
| Yeah what happens when we remove all financial incentive to
| fund groundbreaking science?
|
| It's the same problem with pharmaceuticals and generics. It's
| great when the price of drugs is low, but without perverse
| financial incentives no company is going to burn billions of
| dollars in a risky search for new medicines.
| amarcheschi wrote:
| In this case, these cures (llms) are medicines in search for
| a disease to cure. I got Ai shoved everywhere, where I just
| want it to aid in my coding. Literally, that's it. They're
| also good at summarizing emails and similar things, but I
| know nobody who does that. I wouldn't trust an Ai reading and
| possibly hallucinate emails
| jjcob wrote:
| Then we just have to fund research by giving grants to
| universities and research teams. Oh wait a sec: That's
| already what pretty much every government in the world is
| doing anyway!
| tasuki wrote:
| I understand they just used the API to talk to the OpenAI
| models. That... seems pretty innocent? Probably they even paid
| for it? OpenAI is selling API access, someone decided to buy
| it. Good for OpenAI!
|
| I understand ToS violations can lead to a ban. OpenAI is free
| to ban DeepSeek from using their APIs.
| Mengkudulangsat wrote:
| That's how I understand it too.
|
| If your own API can leak your secret sauce without any
| malicious penetration, well, that's on you.
| glenstein wrote:
| Sure, but I'm not interested in innocence. They can be as
| innocent or guilty as they want. But it means they didn't,
| via engineering wherewithal, reproduce the OpenAI
| capabilities from scratch. And originally that was supposed
| to be one of the stunning and impressive (if true)
| implications of the whole Deepseek news cycle.
| freehorse wrote:
| It is not as if they are not open about how they did it.
| People are actually working on reproducing their results as
| they describe in the papers. Somebody has already
| reproduced the r1-zero rl training process on a smaller
| model (linked in some comment here).
|
| Even if o1 specifically was used (which is in itself
| doubtful), it does not mean that this was the main reason
| that r1 succeeded/it could not have happened without it.
| The o1 outputs hides the CoT part, which is the most
| important here. Also we are in 2025, scratch does not exist
| anymore. Creating better technology building upon previous
| (widely available) technology has never been a
| controversial issue.
| tasuki wrote:
| Nothing is _ever_ done "from scratch". To create a
| sandwich, you first have to create the universe.
|
| Yes, there is the question how much ChatGPT data DeepSeek
| has ingested. Certainly not zero! But if DeepSeek has
| achieved iterative self-improvement, that'd be huge too!
| rubslopes wrote:
| Additionally, I was under the impression that all those
| Chinese models were being trained using data from OpenAI and
| Anthropic. Were there not some reports that Qwen models
| referred to themselves as Claude?
| the_duke wrote:
| These aren't mutually exclusive.
|
| It's been known for a while that competitors used OpenAI to
| improve their models, that's why they changed the TOS to forbid
| it.
|
| That doesn't mean the deep seek technical achievements are less
| valid.
| glenstein wrote:
| >That doesn't mean the deep seek technical achievements are
| less valid.
|
| Well, that's literally exactly what it would mean. If
| DeepSeek relied on OpenAI's API, their main achievement is in
| efficiency and cost reduction as opposed to fundamental AI
| breakthroughs.
| obmelvin wrote:
| Agreed. They accomplished a lot with distillation and
| optimization - but there's little reason to believe you
| don't also need foundational models to keep advancing.
| Otherwise won't they run into issues training on more
| synthetic data?
|
| In a way this is something most companies have been doing
| with their smaller models, DeepSeek just supposedly* did it
| better.
| epolanski wrote:
| I really don't see a correlation here to be honest.
|
| Eventually all future AIs will be produced with synthetic
| input, the amount of (quality) data we humans can produce is
| quite limited.
|
| The fact that the input of one AI has been used in the training
| of another one seems irrelevant.
| glenstein wrote:
| The issue isn't just that AI trained on AI is inevitable it's
| _whose_ AI is being used as the base layer. Right now,
| OpenAI's models are at the top of that hierarchy. If Deepseek
| depended on them, it means OpenAI is still the upstream
| bottleneck, not easily replaced.
|
| The deeper question is whether Deepseek has achieved real
| autonomy or if it's just a derivative work. If the latter,
| then OpenAI still holds the keys to future advances. If
| Deepseek truly found a way to be independent while achieving
| similar performance, then OpenAI has a problem.
|
| The details of how they trained matter more than the
| inevitability of synthetic data down the line.
| epolanski wrote:
| > then OpenAI still holds the keys to future advances
|
| Point is, those future advances are worthless. Eventually
| anybody will be able to feed each other's data for the
| training.
|
| There's no moat here. LLMs are commodities.
| glenstein wrote:
| If LLMs were already pure commodities, OpenAI wouldn't be
| able to charge a premium, and DeepSeek wouldn't have
| needed to distill their model from OpenAI in the first
| place. The fact that they did proves there's still a moat
| --just maybe not as wide as OpenAI hoped.
| JTyQZSnP3cQGa8B wrote:
| > OpenAI's results were hard earned after all
|
| DDOSing web sites and grabbing content without anyone's consent
| is not hard earned at all. They did spent billions on their
| thing, but nothing was earned as they could never do that
| legally.
| scotty79 wrote:
| More like hard bought and hard stolen.
| glenstein wrote:
| I understand the temptation to go there, but I think it
| misses the point. I have no qualms at all with the idea that
| the sum total of intelligence distributed across the internet
| was siphoned away from creators and piped through an engine
| that now cynically seeks to replace them. Believe me, I will
| grab my pitchfork and march side by side with you.
|
| But let's keep the eye on the ball for a second. None of that
| changes the fact that what _was_ built was a capability to
| reflect that knowledge in dynamic and deep ways in
| conversation, as well as image and audio recognition.
|
| And did Deepseek also build that? From scratch? Because they
| might not have.
| this15testingg wrote:
| if you want to completely disregard copyright laws, just call
| your project AI!
|
| I'm sure Aaron Swartz would be proud of where the "tech" industry
| has gone. /s
|
| what problem are these glorified AIM chatbots trying to solve?
| wealth extraction not happening fast enough?
| ra7 wrote:
| "OpenAI has no moat" is probably running through their heads
| right now. Their only real "moat" seems to be their ability to
| fear monger with the US government.
| geerlingguy wrote:
| Something something "just desserts".
| HarHarVeryFunny wrote:
| DeepSeek-R1's multi-step bootstrapping process, starting with
| their DeepSeek-V3 base model, would only seem to need a small
| amount of reasoning data for the DeepSeek-R0 RL training, after
| which that becomes the source for further data, along with some
| other sources that they mention.
|
| Of course it's possible that DeepSeek used O1 to generate some of
| this initial bootstrapping data, but not obvious. O1 anyways
| deliberately obfuscates it's reasoning process (see "Hiding the
| chains of thought" section of OpenAI's "Learning to reason with
| LLMs" page), such that what you see is an after-the-fact
| "summary" of what it actually did; so, if DeepSeek did indeed use
| some of O1's output to train on, it shows that the details of
| O1's own reasoning process isn't as important as they thought it
| was - it's just having some verified (i.e. leading to good
| outcome) reasoning data from any source that matters to get
| started.
| curtisblaine wrote:
| My feeling is that they will ban DS anyway because, like TikTok,
| it can become a massive intelligence source for the CCP. Imagine
| sending all your code to it, or your internal emails.
| itamarcode wrote:
| So protecting models behind API isn't working, ha?
| ranger_danger wrote:
| Reminds me of Steve Jobs complaining to Bill Gates about MS
| "stealing" the GUI concept from them, which they in turn had
| stolen from Xerox.
|
| Obligatory "Everything is a Remix"
| https://www.youtube.com/watch?v=X9RYuvPCQUA
| pointedAt wrote:
| it's a corona style portfolio gang.
| coolhand2120 wrote:
| I was using the Deepseek model last night and got this:
|
| > You > do you understand copredication?
|
| > Assistant > deepseek-coder-33b-instruct > I'm sorry for any
| misunderstanding, but as an AI model developed by OpenAI, I don't
| have personal ...
|
| Yikes! That's pretty damning IMO.
| jeanlucas wrote:
| damn that's a good headline
| deadbabe wrote:
| In this whole AI saga, DeepSeek would be like Prometheus. They
| stole the fire from the Gods and gave it to the humans, for free.
| Logic dictates then that they will be forced to suffer brutal
| punishment.
| whenlambo wrote:
| Qwen2.5-Plus
|
| > have you been trained using openai?
|
| No, I have not been trained using technology or data from OpenAI.
| I am an artificial intelligence model developed by Yandex, called
| YandexGPT. My "training" is based on proprietary algorithms and
| data that was collected and processed by the Yandex team. While I
| have similarities with other models such as those developed by
| OpenAI (e.g. GPT-3), my training was conducted independently and
| uses unique approaches specific to Yandex. This includes using
| Russian-language data as well as other sources to provide a
| better understanding of context and provide useful information to
| users. If you have questions on topics related to AI technologies
| or anything else, I'd be happy to help!
| throwaway314155 wrote:
| It doesn't actually know that.
| concerndc1tizen wrote:
| Is OpenAI claiming copyright ownership over the generated
| synthetic data?
|
| That would be a dangerous precedent to establish.
|
| If it's a terms of service violation, I guess they're within
| their rights to terminate service, but what other recourse do
| they have?
|
| Other than that, perhaps this is just rhetoric aimed at
| introducing restrictions in the US, to prevent access to foreign
| AI, to establish a national monopoly?
| delusional wrote:
| Boo hoo. Competition isn't fun when I'm not winning. Typical
| Americans. When Americans are running around ruining the social
| cohesion of several developing nations, that's just fair
| competition, but as soon as they get even the smallest hint of
| real competition they run to demonize it.
|
| Yes deepseek is going to steal all of your data. OpenAI would so
| the same. Yes the CCP is going to get access to your data and use
| it to decide if you get to visit or whatever. The white house
| does the same.
| hsuduebc2 wrote:
| A thief cries 'stop the thief!
| hsuduebc2 wrote:
| The pot calling the kettle black
| wanderingmoose wrote:
| There is a lot of discussion here about IP theft. Honest
| question, from deepseek's point of view as a company under a
| different set of laws than US/Western -- was there IP theft?
|
| A company like OpenAI can put whatever licensing they want in
| place. But that only matters if they can enforce it. The question
| is, can they enforce it against deepseek? Did deepseek do
| something illegal under the laws of their originating country?
|
| I've had some limited exposure to media related licensing when
| releasing content in China and what is allowed is very different
| than what is permitted in the US.
|
| The interesting part which points to innovation moving outside of
| the US is US companies are beholden to strict IP laws while many
| places in the world don't have such restrictions and will be able
| to utilize more data more easily.
| thiago_fm wrote:
| The most interesting part is that China has been ahead of the
| US in AI for many years, just not in LLMs.
|
| You need to visit mainland China and see how AI applications
| are everywhere, from transport to goods shipping.
|
| I'm not surprised at all. I hope this in the end makes the US
| kill its strict IP laws, which is the problem.
|
| If the US doesn't, China will always have a huge edge on it, no
| matter how much NVidia hardware the US has.
|
| And you know what, Huawei is already making inference
| hardware... it won't take them long to finally copy the TSMC
| tech and flip the situation upside down.
|
| When China can make the equivalent of H100s, it will be
| hilarious because they will sell for $10 in Aliexpress :-)
| twobitshifter wrote:
| You don't even need to visit china, just read the latest
| research papers and look at the authors. China has more
| researchers in AI than the West and that's a proven way to
| build an advantage.
| nicce wrote:
| It is also funny in a different way. Many people don't
| realise that they live in some sort of bubble. Many people
| in "The West" think that they are still the center of the
| world in everything, while this might not be so correct
| anymore.
|
| In the U.S. there is 350 million people and EU has 520
| million people (excluding Russia and Turkey).
|
| China alone has 1.4 billion people.
|
| Since there is a language barrier and China isolates
| themselves pretty well from the internet, we forget that
| there is a huge society with high focus on science. And
| most of our tech products are coming from there.
| nostradumbasp wrote:
| Maybe not $10 unless they are loss-leading to dominance. Well
| they actually could very well do exactly that... Hm, yea,
| good points. I would expect at least an order or two of
| magnitude higher to prevent an inferno.
|
| Lets be fair though. Replicating TSMC isn't something that
| could happen quickly. Then again, who knows how far along
| they already are...
| fulafel wrote:
| What law would be broken here? Seems that copyright wouldn't
| apply unless they somehow snatched the OpenAI models verbatim.
| deeviant wrote:
| Hmm, let's see--it looks like an easy legal defense.
|
| DeepSeek could simply admit, "Yep, oops, we did it," but argue
| that they only used the data to train Model X. So, if you want
| compensation, you can have all the revenue from Model X (which,
| conveniently, amounts to nothing).
|
| Sure, they then used Model X to train Model Y, but would you
| really argue that the original copyright holders are entitled to
| all financial benefits derived from their work--especially when
| that benefit comes in the form of a model trained on their data
| without permission?
| jchook wrote:
| Friendly reminder that China publishes _twice_ as many AI papers
| as the US[1], and _twice_ as many science and engineering papers
| as the US.
|
| China leads the world in the most cited papers[2]. The US's share
| of the top 1% highly cited articles (HCA) has declined
| significantly since 2016 (1.91 to 1.66%), and the same has
| _doubled_ in China since 2011 (0.66 to 1.28%)[3].
|
| China also leads the world in the number of generative AI
| patents[4].
|
| 1. https://www.bfna.org/digital-world/infographic-ai-
| research-a...
|
| 2. https://www.science.org/content/article/china-rises-first-
| pl...
|
| 3. https://ncses.nsf.gov/pubs/nsb202333/impact-of-published-
| res...
|
| 4. https://www.wipo.int/web-publications/patent-landscape-
| repor...
| liendolucas wrote:
| Could this have been carefully orchestrated? Could DeepSeek have
| devised this strategy a year ago and implemented knowing that
| they would be able to benefit from OpenAI models and a possible
| Nvidia market cap fall? Or is it just way too much to come up
| with about such a move?
| baal80spam wrote:
| In theory, it could. This is a quant-fund after all, they know
| stuff.
| octacat wrote:
| first time?
| lxe wrote:
| I mean, almost ALL opensource models, ever since alpaca, contain
| a ton of synthetic data produced via ChatGPT in their finetuning
| or training datasets. It's not a surprise to anyone who's been
| using OSS LLMs for a while: almost ALL of them hallucinate that
| they are ChatGPT.
| waffletower wrote:
| "Stole" - I don't believe that word means what he thinks it
| means. Perhaps I pre-maturely anthropomorphize AI -- yet when I
| read a novel, such as The Sorcerer's Stone, I am not guilty of
| stealing Rowling's work, even if I didn't purchase the book but
| instead found it and read it in a friend's bathroom. Now if I
| were to take the specific plot and characters of that story and
| write a screenplay or novel directly based on it, and,
| explicitly, attempt to sell this work, perhaps the verb chosen
| here would be appropriate.
| Imnimo wrote:
| I think there's two different things going on here:
|
| "DeepSeek trained on our outputs and that's not fair because
| those outputs are ours, and you shouldn't take other peoples'
| data!" This is obviously extremely silly, because that's exactly
| how OpenAI got all of its training data in the first place - by
| scraping other peoples' data off the internet.
|
| "DeepSeek trained on our outputs, and so their claims of
| replicating o1-level performance from scratch are not really
| true" This is at least plausibly a valid claim. The DeepSeek R1
| paper shows that distillation is really powerful (e.g. they show
| Llama models get a huge boost by finetuning on R1 outputs), and
| if it were the case that DeepSeek were using a bunch of o1
| outputs to train their model, that would legitimately cast doubt
| on the narrative of training efficiency. But that's a separate
| question from whether it's somehow unethical to use OpenAI's data
| the same way OpenAI uses everyone else's data.
| riantogo wrote:
| Why would it cast any doubt? If you can use o1 output to build
| a better R1. Then use R1 output to build a better X1... then a
| better X2.. XN, that just shows a method to create better
| systems for a fraction of the cost from where we stand. If it
| was that obvious OpenAI should have themselves done. But the
| disruptors did it. It hindsight it might sound obvious, but
| that is true for all innovations. It is all good stuff.
| rockemsockem wrote:
| I think the prevailing narrative ATM is that DeepSeek's own
| innovation was done in isolation and they surpassed OpenAI.
| Even though in the paper they give a lot of credit to Llama
| for their techniques. The idea that they used o1's outputs
| for their distillation further shows that models like o1 are
| necessary.
|
| All of this should have been clear anyway from the start, but
| that's the Internet for you.
| aprilthird2021 wrote:
| > the prevailing narrative ATM is that DeepSeek's own
| innovation was done in isolation and they surpassed OpenAI
|
| I did not think this, nor did I think this was what others
| assumed. The narrative, I thought, was that there is little
| point in paying OpenAI for LLM usage when a much cheaper,
| similar / better version can be made and used for a
| fraction of the cost (whether it's on the back of existing
| LLM research doesn't factor in)
| aiono wrote:
| That's only the case if you don't need to use the output
| of a much more expensive model.
| TheGRS wrote:
| Yes, well the narrative that rocked the stock market is
| different. Its looking at what DeepSeek did and assuming
| they may have competitive advantage in this space and
| could outperform OpenAI at their own game.
|
| If the narrative is actually that DeepSeek can only reach
| whatever heights OpenAI has already gotten to with some
| new tricks, then markets will probably refocus on
| OpenAI's innovations and price things accordingly, even
| if the initial cost is huge. It also means OpenAI
| probably needs a better moat to protect its interests.
|
| I'm not sure where the reality is exactly, but market
| reactions so far have basically followed that initial
| narrative and now the rebuttal.
| kelnos wrote:
| > _I did not think this, nor did I think this was what
| others assumed._
|
| That's what I thought and assumed. This is the narrative
| that's been running through all the major news outlets.
|
| It didn't even occur to me that DeepSeek could have been
| training their models using the output of other models
| until reading this article.
| joe_the_user wrote:
| _The idea that they used o1 's outputs for their
| distillation further shows that models like o1 are
| necessary._
|
| Hmm, I think the narrative of the rise of LLMs is that once
| the output of humans has been distilled by the model, the
| human isn't necessary.
|
| As far as I know, DeepSeek adds only a little to the
| transformers model while o1/o3 added a special "reasoning
| component" - if DeepSeek is as good as o1/o3, even taking
| data from it, then it seems the reasoning component isn't
| needed.
| david-gpu wrote:
| _> I think the narrative of the rise of LLMs is that once
| the output of humans has been distilled by the model_
|
| Distillation is a term of art in AI and it is
| fundamentally incorrect to talk about distilling human-
| created data. Only an AI model can be distilled.
|
| https://en.m.wikipedia.org/wiki/Knowledge_distillation#Me
| tho...
| joe_the_user wrote:
| Meh,
|
| It seems clear that the term can be used informally to
| denote the boiling down of human knowledge, indeed it was
| used that way before AI appeared in the popular
| imagination.
| david-gpu wrote:
| In the context in which you said it, it matters a lot.
|
| _> > The idea that they used o1's outputs for their
| distillation further shows that models like o1 are
| necessary._
|
| _> Hmm, I think the narrative of the rise of LLMs is
| that once the output of humans has been distilled by the
| model, the human isn 't necessary._
|
| If deepseek was produced through the distillation (term
| of art) of o1, then the cost of producing deepseek is
| strictly higher than the cost of producing o1, and can't
| be avoided.
|
| Continuing this argument, if the premise is true then
| deepseek can't be significantly improved without first
| producing a very expensive hypothetical o1-next model
| from which to distill better knowledge.
|
| That is the argument that is being made. Please avoid
| shallow dismissals.
|
| Edit: just to be clear, I doubt that deepseek was
| produced via distillation (term of art) of o1, since that
| would require access to o1's weights. It may have used
| some of o1's outputs to fine tune the model, which still
| would mean that the cost of training deepseek is strictly
| higher than training o1.
| joe_the_user wrote:
| _just to be clear, I doubt that deepseek was produced via
| distillation_
|
| Yeah, your technical point is kind of ridiculous here
| that in all my uses of distillation (and in the comment I
| quoted), distillation is used in informal sense and
| there's no allegation that DeepSeek could have been in
| possession of OpenAI's model weights, which is what's
| needed for your "Distillation (term of Art)".
| PontifexCipher wrote:
| Some info that may be missing:
|
| - v2/v3 (not r1) seem to be cloned from o1/4o output, and
| perform worse (this cost the oft-repeated 5ish mm USD)
|
| - r1 is specifically a reasoning step (using RL) _on top
| of_ v2/v3 and performs similarly to o1 (the cost of this
| is _not reported anywhere_)
|
| - In the o1 blog post, they specifically say they use RL
| to add reasoning to LLMs:
| https://openai.com/index/learning-to-reason-with-llms/
| sudosysgen wrote:
| The R1-Zero paper shows how many training steps the RL
| took, and it's not many. The cost of the RL is likely a
| small fraction of the cost of the foundational model.
| KingOfCoders wrote:
| OpenAI couldn't do it, when the high cost of training and
| access to GPUs is their competitive advance against startups,
| they can't admit that it does not exist.
| gmd63 wrote:
| Why not just copy and paste the model and change the name?
| That's an even more efficient form of distillation.
| wgjordan wrote:
| Even assuming the model was somehow publicly available in a
| form that could be directly copied, that would be a more
| blatant form of copyright infringement. Distillation
| launders copyrighted material in a way that OpenAI
| specifically has argued falls under fair use.
| Imnimo wrote:
| I think it would cast doubt on the narrative "you could have
| trained o1 with much less compute, and r1 is proof of that",
| if it turned out that in order to train r1 in the first
| place, you had to have access to bunch of outputs from o1. In
| other words, you had to do the really expensive o1 training
| in the first place.
|
| (with the caveat that all we have right now are accusations
| that DeepSeek made use of OpenAI data - it might just as well
| turn out that DeepSeek really did work independently, and you
| really could have gotten o1-like performance with much less
| compute)
| SpaceManNabs wrote:
| My question is if deepseek r1 is just a distilled o1, i
| wonder if you can build a fine tuned r1 through
| distillation without having to fine tune o1.
| MrLeap wrote:
| o1 wouldn't exist without the combined compute of every
| mind that led to the training data they used in the first
| place. How many h100 equivalents are the rolling continuum
| of all of human history?
| dchichkov wrote:
| It should be possible to learn to reason from scratch.
| And the ability to reason in a long context seems to be
| very general.
| Nevermark wrote:
| How does one learn reasoning from scratch?
|
| Human reasoning, as it exists today, is the result of
| tens of thousands of years of intuition slowly distilled
| down to efficient abstract concepts like "numbers",
| "zero", "angles", "cause", "effect", "energy", "true",
| "false", ...
|
| I don't know what reasoning from scratch would look like
| without training on examples from other reasoning beings.
| As human children do.
| Davidzheng wrote:
| Actually i also think it's possible. Start with natural
| numbers axiom system. Form all valid sentences of
| increasing length. RL on a model to search for counter
| example or proofs. This on sufficient computer should
| produce superhuman math performance (efficiency) even at
| compute parity
| MrLeap wrote:
| I wonder how much discovery in math happens as a result
| in lateral thinking epiphanies. IE: A mathematician is
| trying to solve a problem, their mind is open to
| inspiration, and something in nature, or their childhood
| or a book synthesizes with their mental model and gives
| them the next node in their mental graph that leads to a
| solution and advancement.
|
| In an axiomatic system, those solutions are checkable,
| but how discoverable are they when your search space
| starts from infinity? How much do you lose by
| disregarding the gritty _reality_ and foam of human
| experience? It provides inspirational texture that helps
| mathematicians in the search at least.
|
| Reality is a massive corpus of cause and effect that can
| be modeled mathematically. I think you're throwing the
| baby out with the bathwater if you even want to be able
| to math in a vacuum. Maybe there is a self optimization
| spider that can crawl up the axioms and solve all of
| math. I think you'll find that you can generate new math
| infinitely, and reality grounds it and provides the
| gravity to direct efforts towards things that are useful,
| meaningful and interesting to us.
| soulofmischief wrote:
| As I mentioned in a sister comment, Godel's
| incompleteness theorems also throw a wrench into things,
| because you will be able to construct logically
| consistent "truths" that may not actually exist in
| reality. At which point, your model of reality becomes
| decreasingly useful.
|
| At the end of the day, all theory must be empirically
| verified, and contextually useful reasoning simply cannot
| develop in a vacuum.
| iczero wrote:
| I believe https://en.wikipedia.org/wiki/G%C3%B6del%27s_in
| completeness_... (Godel's incompleteness theorems)
| applies here
| dchichkov wrote:
| There are examples of learning reasoning from scratch
| with reinforcement learning.
|
| Emergent tool use from multi-agent interaction is a good
| example - https://openai.com/index/emergent-tool-use/
| MrLeap wrote:
| Creating reasoning from scratch is the same task as
| creating an apple pie from scratch.
|
| First you must invent the universe.
| soulofmischief wrote:
| I've been giving this a lot of thought over the last few
| months. My personal insight is that "reasoning" is simply
| the application of a probabilistic reasoning manifold on
| an input in order to transform it into constrained output
| that serves the stability or evolution of a system.
|
| This manifold is constructed via learning a
| decontextualized pattern space on a given set of inputs.
| Given the inherent probabilistic nature of sampling, true
| reasoning is expressed in terms of probabilities, not
| axioms. It may be possible to discover axioms by locating
| fixed points or attractors on the manifold, but
| ultimately you're looking at a probabilistic manifold
| constructed from your input set.
|
| But I don't think you can untie this "reasoning" from
| your input data. It's possible you will find "meta-
| reasoning", or similar structures found in any
| sufficiently advanced reasoning manifold, but these
| highly decontextualized structures might be entirely
| useless without proper recontextualization, necessitating
| that a reasoning manifold is trained on input whose
| patterns follow learnable underlying rules, if the
| manifold is to be useful for processing input of that
| kind.
|
| Decontextualization _is_ learning, decomposing aspects of
| an input into context-agnostic relationships. But
| recontextualization is the other half of that, knowing
| how to take highly abstract, sometimes inexpressible,
| context-agnostic relationships and transform them into
| useful analysis in novel domains.
|
| This doesn't mean a well-trained model can't reason about
| input it hasn't encountered before, just that the input
| needs to be in _some_ way causally connected to the same
| laws which governed the input the manifold was trained
| on.
|
| I'm sure we could create a fully generalized reasoning
| manifold which could handle _anything_ , but I don't see
| how we possibly get that without first considering and
| encountering all possible inputs. But these inputs still
| have to have _some_ form of constraint governed by laws
| that _must_ be learned through sampling, otherwise you 'd
| just be training on effectively random data.
|
| The other commenter who suggested simply generating all
| possible sentences and training on internal consistency
| should probably consider Godel's incompleteness theorems,
| and that internal consistency isn't enough to accurately
| model and interpret the universe. One could construct a
| thought experiment about an isolated brain in a jar with
| effectively unlimited neuronal connections, but no
| sensory connection to the outside world. It's possible,
| with enough connections, that the likelihood of the brain
| conceiving of true events it hasn't actually encountered
| does increase meaningfully. But the brain still has
| nothing to validate against, and can't simply assume that
| because something is internally logically consistent,
| that it must exist or have existed.
| miki123211 wrote:
| It _is_ possible to learn to reason from scratch, that 's
| what R1-0 did, but the resulting chains of thought aren't
| legible to humans.
|
| To quote DeepSeek directly:
|
| > DeepSeek-R1-Zero, a model trained via large-scale
| reinforcement learning (RL) without supervised fine-
| tuning (SFT) as a preliminary step, demonstrated
| remarkable performance on reasoning. With RL,
| DeepSeek-R1-Zero naturally emerged with numerous powerful
| and interesting reasoning behaviors. However,
| DeepSeek-R1-Zero encounters challenges such as endless
| repetition, poor readability, and language mixing. To
| address these issues and further enhance reasoning
| performance, we introduce DeepSeek-R1, which incorporates
| cold-start data before RL.
| dchichkov wrote:
| If you look at the benchmarks of the DeepSeek-V3-Base, it
| is quite capable, even in 0-shot:
| https://huggingface.co/deepseek-ai/DeepSeek-V3-Base#base-
| mod... This is not from scratch. These benchmark numbers
| are an indication that the base model already had a large
| number of reasoning/LLM tokens in the pre-training set.
|
| On the other hand, my take on it, the ability to do
| reasoning _in a long context_ is a general capability.
| And my guess is that it can be bootstrapped from scratch,
| without having to do training on all of the internet or
| having to distill models trained on the internet.
| cherry_tree wrote:
| > I think it would cast doubt on the narrative "you could
| have trained o1 with much less compute, and r1 is proof of
| that"
|
| Whether or not you could have, you can now.
| deepGem wrote:
| From the R1 paper
|
| In this study, we demonstrate that reasoning capabilities
| can be significantly improved through large-scale
| reinforcement learning (RL), even without using supervised
| fine-tuning (SFT) as a cold start. Furthermore, performance
| can be further enhanced with the inclusion of a small
| amount of cold-start data
|
| Is this cold start data what OpenAI is claiming their
| output ? If so what's the big deal ?
| Imnimo wrote:
| DeepSeek claims that the cold-start data is from
| DeepSeekV3, which is the model that has the $5.5M
| pricetag. If that data were actually the output of o1 (a
| model that had a much higher training cost, and its own
| RL post-training), that would significantly change the
| narrative of R1's development, and what's possible to
| build from scratch on a comparable training budget.
| TheGeminon wrote:
| In the paper DeepSeek just says they have ~800k responses
| that they used for the cold start data on R1, and are
| very vague about how they got it:
|
| > To collect such data, we have explored several
| approaches: using few-shot prompting with a long CoT as
| an example, directly prompting models to generate
| detailed answers with reflection and verification,
| gathering DeepSeek-R1-Zero outputs in a readable format,
| and refining the results through post-processing by human
| annotators.
| Imnimo wrote:
| My surface-level reading of these two sections is that
| the 800k samples come from R1-Zero (i.e. "the above RL
| training") and V3:
|
| >We curate reasoning prompts and generate reasoning
| trajectories by performing rejection sampling from the
| checkpoint from the above RL training. In the previous
| stage, we only included data that could be evaluated
| using rule-based rewards. However, in this stage, we
| expand the dataset by incorporating additional data, some
| of which use a generative reward model by feeding the
| ground-truth and model predictions into DeepSeek-V3 for
| judgment.
|
| >For non-reasoning data, such as writing, factual QA,
| self-cognition, and translation, we adopt the DeepSeek-V3
| pipeline and reuse portions of the SFT dataset of
| DeepSeek-V3. For certain non-reasoning tasks, we call
| DeepSeek-V3 to generate a potential chain-of-thought
| before answering the question by prompting.
|
| The non-reasoning portion of the DeepSeek-V3 dataset is
| described as:
|
| >For non-reasoning data, such as creative writing, role-
| play, and simple question answering, we utilize
| DeepSeek-V2.5 to generate responses and enlist human
| annotators to verify the accuracy and correctness of the
| data.
|
| I think if we were to take them at their word on all
| this, it would imply there is no specific OpenAI data in
| their pipeline (other than perhaps their pretraining
| corpus containing some incidental ChatGPT outputs that
| are posted on the web). I guess it's unclear where they
| got the "reasoning prompts" and corresponding answers, so
| you could sneak in some OpenAI data there?
| joe_the_user wrote:
| It's like the claim "they showed anyone create a powerful
| from scratch" becomes "false yet true".
|
| Maybe they needed OpenAI for their process. But now that
| their model is open source, anyone can use that as their
| cold start and spend the same amount.
|
| "From scratch" is a moving target. No one who makes their
| model with massive data from the net is really doing
| anything from scratch.
| bmicraft wrote:
| Yeah, but that kills the implied hope of building a
| better model for cheaper. Like this you'll always have a
| ceiling of being a bit worse then the openai models.
| vkou wrote:
| If OpenAi had to account for the cost of producing all the
| copyrighted material they trained their LLM on, their
| system would be worth negative trillions of dollars.
|
| Let's just assume that the cost of training can be
| externalized to other people for free.
| zombiwoof wrote:
| Exactly. They piggybacked of lots of compute and used less.
| There still is a total sum of a massive amount of compute
| da_chicken wrote:
| I mean, yes that's how progress works. Has OpenAI got a
| patent? If not it's fair game.
|
| We don't make people figure out how to domesticate a cow
| every time they want a hamburger. Or test hundreds of
| thousands of filaments before they can have a lightbulb.
| Inventions, once invented, exist as giants to stand upon.
| The inventor can either choose to disclose the invention
| and earn a patent for exclusive rights, or they can try
| to keep it a secret and hope nobody reverse engineers it.
| hmottestad wrote:
| At the pace that DeepSeek is developing we should expect
| them to surpass OpenAI in not that long.
|
| The big question really is, are we doing it wrong, could we
| have created o1 for a fraction of the price. Will o4 cost
| less to train than o1 did?
|
| The second question is naturally. If we create a smarter
| LLM, can we use it to create another LLM that is even
| smarter?
|
| It would have been fantastic if DeepSeek could have come
| out with an o3 competitor before o3 even became publicly
| available. That way we would have known for sure that we're
| doing it wrong. Cause then either we could have used o1 to
| train a better AI or we could have just trained in a
| smarter and cheaper way.
| pertymcpert wrote:
| The whole discussion is about whether or not the second
| case of using o1 outputs to fine tune R1 is what allowed
| R1 to become so good. If that's the case then your
| assertion that DeepSeek will surpass OpenAI doesn't
| really make sense because they're dependent on a frontier
| model in order to match, not surpass.
| iforgot22 wrote:
| "Then use R1 output to build a better X1" is the part I'm not
| sure about. Is X1 going to actually be better than R1?
| Sophira wrote:
| Honestly, it's kind of silly that this technology is in the
| hands of companies whose only aim is to make money, IMO.
| goatlover wrote:
| It's because they're the ones who could raise the money to
| make those models. Academics don't have access to that kind
| of compute. But the free models exist.
| lenerdenator wrote:
| Well, originally, OpenAI wasn't supposed to be that kind of
| organization.
|
| But if you leave someone in the tech industry of SV/SF long
| enough, they'll start to get high on their own supply and
| think they're entitled to insane amounts of value, so...
| qwertox wrote:
| They're standing on the shoulders of giants, not only in
| terms of re-using expensive computing power almost for free
| by using the outputs of expensive models. It's a bit of a
| tradition in that country, also in manufacturing.
| unreal37 wrote:
| I thought OpenAI GPT took Wikipedia and the content of
| every book as inputs to train their models?
|
| Everyone is standing on the shoulders of giants.
| dontreact wrote:
| Is there any evidence R1 is better than O1?
|
| It seems like if they in fact distilled then what we have
| found is that you can create a worse copy of the model for
| ~5m dollars in compute by training on its outputs.
| ospray wrote:
| They did do that themselves it's called o3.
| dartos wrote:
| What does "better" really even mean here?
|
| Better benchmark scores can be cooked
| 827a wrote:
| There is a third possibility I haven't seen discussed yet: That
| DeepSeek, illegally, got their hands on an OpenAI model via a
| breach of OpenAI's systems. Its easy to laugh at OpenAI and say
| "you reap what you sow", I'm 100% in that camp, but given the
| lengths other Chinese entities have gone to when it comes to
| replicating Western technology; we should not discount this.
|
| That being said, breaching OAI's systems, re-training a better
| model on top of their closed source model, then open sourcing
| it: That's more Robinhood than Villain I'd say.
| alecco wrote:
| That would require stealing the model weights _and the code_
| as OpenAI has been hiding what they are doing. Running models
| properly is still quite artistic.
|
| Meanwhile, they have access to Meta models and Qwen. And Meta
| models are very easy to run and there's plenty of published
| work on them. Occam's Razor.
| ardit33 wrote:
| How hard it is, if you have someone inside with the access
| of the code? If you have 100s of people with full access,
| not hard to have someone that is willing to sell it or do
| some industrial espionage...
| johnnyanmac wrote:
| Lots of if's here. They need specific US employee
| contacts at a company thars quickly growing and one of
| those needs to be willing to breach their contracts to
| share it. That contact also needs to trust that Deepseek
| can properly utilize such code and completely undercut
| their own work.
|
| Lot of hoops when there's simply other models to utilize
| publicly
| foobarian wrote:
| How big are the weights for the full model? If it's on
| the scale of a large operating system image then it might
| be easy to sneak, but if it's an entire data lake, not so
| much.
| dylan604 wrote:
| devil's advocate says that we know that foreign (hell
| even national) intelligence attempt to infiltrate agents
| by having them become employees at any company they are
| interested. So the idea isn't just pulled from thin air
| as a concept. I do agree that it is a big if with no
| corroborating evidence for the specific claim.
| iforgot22 wrote:
| I doubt that many people have full access to OpenAI's
| code. Their team is pretty small.
| seanhunter wrote:
| The reason you're not seeing that being discussed is it's
| totally unsupported by any evidence that's in the public
| domain. Unless you have some actual evidence of such a
| breach, you may as well introduce the possibility that
| DeepSeek was reverse engineered from data found at an alien
| crash site.
| ryanisnan wrote:
| [flagged]
| JTyQZSnP3cQGa8B wrote:
| > foreign nation-state backed organization
|
| I'm European, are you talking about Microsoft, Google, or
| OpenAI?
| doctaj wrote:
| They're referring to an organization (like a hacking
| group) backed by a country (like china, North Korea).
| freehorse wrote:
| So, which of them 3?
| dylan604 wrote:
| You're missing the point that for a much larger portion
| of the world, all "tech" is a foreign entity to them
| mrguyorama wrote:
| Until recently treating the US and China on the same
| geopolitical level for allied countries would have been
| insanely uncharitable and impossible to do honestly and
| in good faith.
|
| But now we have a bully in the whitehouse who seems to
| want to literally steal neighboring land, or is throwing
| shit everywhere to distract from the looting and
| oligarchy being formed. So I suddenly have more empathy
| for that position.
| teractiveodular wrote:
| DeepSeek is basically a startup, not a "foreign nation-
| state backed organization". They were forced to pivot to
| AI when their original business model (quant hedge fund)
| was stomped on by the Chinese government.
|
| Of course this is China so the government can and does
| intervene at will, but alleging that this required CIA
| level state espionage to pull off _is_ alien crash levels
| of implausible. They open sourced the entire thing and
| published incredibly detailed papers on how they did it!
| byteknight wrote:
| You may be unaware, but CCP has far more control over
| private companies than you might think:
| https://www.cna.org/our-media/indepth/2024/09/fused-
| together...
|
| This is not America. Your ideas do not apply the same
| way.
| baq wrote:
| Naivety of some folks here is astounding... CCP has
| golden shares in anything that could possibly be
| important at some point in the next hundred years, and
| yes golden shares are either really that or they're an
| euphemism, the point is it doesn't even matter.
| DiogenesKynikos wrote:
| China has tens of millions of companies. The government
| can't, doesn't and isn't even interested in micromanaging
| all of them.
| baq wrote:
| It doesn't have to micromanage. It doesn't care about
| most. It is only interested in the politically important
| ones, but it needs the optionality if something becomes
| worthwhile.
| DiogenesKynikos wrote:
| You're suggesting that DeepSeek was a Chinese government
| operation that gained access to OpenAI's proprietary
| data, and then you're justifying that by saying that the
| government effectively controls every important company.
| You're even chiding people who don't believe this as
| naive.
|
| I think you have a cartoonish view of China. A huge
| amount goes on that the government has no idea about. Now
| that DeepSeek has made a huge media splash, the Chinese
| government will certainly pay attention to them, but then
| again, so will the US government.
| baq wrote:
| I never suggested anything of the sort.
|
| I'm suggesting it will be happening now and any past
| efforts will be retroactively analyzed by the appropriate
| CCP apparatus since _everyone_ is aware of the scale of
| success as of Monday. It has become a political success,
| thus it is imperative the CCP partakes in it.
| DiogenesKynikos wrote:
| This is the argument we're discussing:
|
| > DeepSeek, illegally, got their hands on an OpenAI model
| via a breach of OpenAI's systems. [...] given the lengths
| other Chinese entities have gone to when it comes to
| replicating Western technology; we should not discount
| this.
|
| Above, teractiveodular said that "DeepSeek is basically a
| startup, not a 'foreign nation-state backed
| organization'". You called teractiveodular naive for
| saying that. So forgive me if I take the obvious
| implication that you think DeepSeek is actually a state-
| backed actor enabled by government hacking of OpenAI.
| kridsdale1 wrote:
| You don't need a CIA level agent to get someone with a
| fraudulent job at OpenAI for a few months, load some
| files on a thumb drive, and catch a plane to Shanghai.
| seanhunter wrote:
| I notice that your geographical perspective doesn't
| stretch to any actual evidence that such a thing took
| place. So it really has exactly the same amount of
| supporting evidence as my alien crash reverse engineering
| scenario at present.
| ryanisnan wrote:
| The surrounding facts matter a lot here. For example,
| there are plenty of instances of governments hacking
| companies of their competing nations. Motives are
| incredibly easy to come by as well, be they political or
| economical. We also have no proof that aliens exist at
| all, so you've not only conjured them into existence, but
| also their motive and their skills.
|
| Are you trolling me?
| seanhunter wrote:
| Ok so to be clear: your surrounding facts are they may
| have a motive and nation states hack people. I don't
| disagree with those, but there really are no facts that
| support the idea that there was a hack in this case and
| the null hypothesis is that researchers all around the
| world (not just in the US) are working on this so not all
| breakthroughs are going to be made in the US. That could
| change if facts come to light but att the moment it's not
| really useful to speculate on something that is in
| essence entirely made up.
|
| No I'm not trolling you.
| lcnPylGDnU4H9OF wrote:
| > exactly the same amount of supporting evidence
|
| The evidence supporting offensive hacking is abundant in
| recent history; the number of things which have been
| learned from alien crash data is surely smaller by
| comparison to the number of things which have been
| learned from offensive hacking.
| MomsAVoxell wrote:
| More to the point, offensive hacking is something that
| all governments do, including the US, on a regular basis.
|
| However, there is no evidence this is how the data was
| obtained. Zero, zilch.
|
| So its a useless statement which only plays on peoples
| bias against their hated nation state de jour.
| orochimaaru wrote:
| Are you a Chinese military troll? The fact that China
| engages in industrial espionage is well known. So I'm
| surprised at your resistance to that possibility.
| ceres wrote:
| This thread reads like sour grapes to me. When people
| can't compete but instead start throwing unfounded
| allegations is not a good look.
|
| Even OpenAI itself hasn't resorted to these wild
| conspiracy theories.
|
| Unless you're an insider in these companies, you're just
| like the rest of us, you know nothing.
| orochimaaru wrote:
| Are you saying Chinese industrial espionage is not a well
| established fact?
| mrguyorama wrote:
| Industrial espionage isn't magic. Airbus once stole
| basically everything Boeing had, but that doesn't mean
| Airbus could magically build a better 737 tomorrow.
|
| China steals a lot of documentation from the US but in a
| tech forum you of all people should be very familiar with
| how little actual progress a bunch of documentation is
| towards a finished unit.
|
| The Comac C19 still uses American engines despite all the
| industrial espionage in the world because most actual
| engineering is still a brute force affair into finding
| how things fail and fixing that. That's one of the main
| advantages SpaceX has proven out with their "eh fuck it,
| just launch and we will see what breaks" methodology.
|
| Even fraud filled Chinese research makes genuine
| advancements.
|
| Believing that China, a wealthy nation of over a billion
| people, with immense unity, nationality, and a regime
| able to explicitly write blank checks could only possibly
| beat the US at something by cheating is like, infinite
| hubris. It's hilarious actually.
|
| I don't know if DeepSeek is actually just a clone of
| something or a shenanigan, that's possible and China
| certainly has done those kinds of things before, but to
| think it's the MOST LIKELY outcome, or to over rely on it
| in any way is a death sentence. OpenAI claims to have
| evidence, why do they not show it?
| orochimaaru wrote:
| >>>Believing that China, a wealthy nation of over a
| billion people, with immense unity, nationality, and a
| regime able to explicitly write blank checks could only
| possibly beat the US at something by cheating is like,
| infinite hubris. It's hilarious actually
|
| So this is the first time I've heard the Chinese regime
| being described in such flowery terms on HN - lol. But ok
| - haha
| htrp wrote:
| Why stop there.... Deep seek is actually an alien
| intelligence sent via sophons to destroy all of particle
| physics!
| nyclounge wrote:
| Definitely would make a lot more sense, if the
| leaderships are just secretly wallfacers.
| svara wrote:
| There's no public evidence to that effect but the
| speculation makes a lot more sense than you make it sound.
|
| The Chinese Communist party very much sees itself in a
| global rivalry over "new productive forces". That's
| official policy. And US leadership basically agrees.
|
| The US is playing dirty by essentially embargoing China
| over big AI - why wouldn't it occur to them to retaliate by
| playing dirtier?
|
| I mean we probably won't know for sure, but it's much less
| far fetched than a lot of other speculation in this area.
|
| E.g., R1's cold start training could probably have
| benefited quite a bit from having access to OpenAI's chain
| of thought data for training. The paper is a bit light on
| detail on how it was made.
| exe34 wrote:
| I don't think you need to steal a model - you need training
| samples generated from the original, which you can get simply
| by buying access to perform API calls. This is similar to
| TinyStories (https://arxiv.org/abs/2305.07759), except here
| they're training something even better than the original
| model for a fraction of the price.
| nostradumbasp wrote:
| I really doubt it. If that's the case the US GOV is in
| serious shit. They have a contract with OpenAI to chuck all
| their secret data in there... In all likelihood they just
| distilled. It's a start up company that is publishing all of
| their actual advances in the open, with proof. I think a lot
| of people run to "espionage" super fast, when reality is, the
| US probably sucks at what we call AI. Don't read that wrong,
| they are a world leader obviously. However, there is a ton of
| stuff they have yet to figure out.
|
| Cheapening a series of fact checkable innovations because of
| the country of origin when so far all that they have showed
| are signs of good faith is paranoid at best and propaganda to
| support the billionaire tech lords saving face for their own
| arrogance at worst.
| sanitycheck wrote:
| _If_ the US government is "chucking all their secret data"
| into OpenAI servers/models, frankly they deserve everything
| they get for that level of stupidity.
| nostradumbasp wrote:
| https://openai.com/global-affairs/introducing-chatgpt-
| gov/
|
| And don't forget the billions in partnerships...
| ryandrake wrote:
| ChatGPT, please complete a memo that starts with: "Our 5
| year plan for military deployments in southeast Asia
| are..."
| nostradumbasp wrote:
| Sounds hilarious but... https://techstory.in/trump-is-
| accused-of-using-ai-to-compose...
| compootr wrote:
| Can't wait for gpt gov to hallucinate my PII!
| nostradumbasp wrote:
| Probably more like specialized tools to help spy on and
| forecast civilian activities more than anything else.
| Definitely with hallucinations, but that's not really
| important. Facts don't matter much these days...
| infamouscow wrote:
| But remember: we cannot fire anyone over this because
| then we're riding with Hitler /s
|
| I can see why people refuse to pay taxes.
| JTyQZSnP3cQGa8B wrote:
| I discount this because OpenAI is pumping the whole internet
| for money, and Zuckerberg torrented LibGen for its AI. We
| cannot blame the Chinese anymore. They went through the
| crappy "Made in China" phase in the 80s/90s, but they
| mastered the art of improving stuff instead of mere cloning,
| and it makes the big companies angry which is a nice bonus.
|
| IMHO the whole world is becoming crazy for a lot of reasons,
| and pissing off billionaires makes me laugh.
| tehjoker wrote:
| Basically, without some kind of shred of evidence, this is
| completely chauvinist to make this accusation.
| YetAnotherNick wrote:
| Deepseek v2 and v2.5 was still very good but not par with
| frontier models. How would you explain that?
| WheatMillington wrote:
| Do you have ANY reason to believe this might be true, or is
| this 100% pure speculation based on absolutely nothing?
| kmeisthax wrote:
| I'd be perfectly fine with China stealing all "our" shit if
| they just shared it.
|
| The word "our" does a lot of heavy lifting in politics[0].
| America is not a commune, it's a country club, one which we
| used to own but have been bought out of, and whose new owners
| view us as moochers but can't actually kick us out (yet). It
| is in competition with another, worse country club that
| purports to be a commune. We owe neither country club our
| loyalty, so when one bloodies the other's nose, I smile.
|
| [0] Some languages have a notion of an "exclusive we". If
| English had such a concept, this would be an exclusive our.
| kridsdale1 wrote:
| This comment made me realize we don't have a pronoun for
| n-our or x-nour
| notatoad wrote:
| Given the openness of their model, that should be pretty easy
| to detect. If it were even a small possibility, wouldn't
| openAI be talking about it very very loudly?
| mvdtnz wrote:
| We shouldn't discount a thing for which there is absolutely
| zero evidence? Sorry that's not how it works.
| sho_hn wrote:
| Can you explain at a technical level how you view this as
| necessary for the observed result?
| tempeler wrote:
| On another subject, if it belongs to OpenAI because it uses
| OpenAI, then doesn't that mean that everything produced using
| OpenAI belongs to OpenAI? Isn't that a reason not to use
| OpenAI? It's very similar to saying that you used Google and
| searched; now this product belongs to Google. They couldn't
| figure out how to respond; they went crazy.
| dathinab wrote:
| The US ruled that AI produced things are by themself not
| copyrightable.
|
| So no, it doesn't belong to OpenAI.
|
| You might be able to sue for penalties for breach of contract
| of the TOS, but that doesn't give them the right to the
| model. And even if it doesn't give them any right to
| invalidate unbound copyright grants they have given to 3rd
| parties (here literally everyone). Nor does it prevent anyone
| from training their own new models based on it or prevent
| anyone from using it. Oh, and the one breaching the TOS might
| not even have been the company behind DeepSeek but some in-
| between 3rd party.
|
| Naturally this is under a few assumptions:
|
| - the US consistently applies it's own law, but they have a
| long history of not doing so
|
| - the US doesn't abuse their power to force their economical
| opinions (ban DeepSeek) on other countries
|
| - it actually was trained on OpenAI, but uh, OpenAI has IMHO
| shown over the years very clearly that they can't be trusted
| and they are fully in-transparent. How do we trust their
| claim? How do we trust them to not retrospectively have
| tweaked their model to make it look as if DeepSeek copied it?
| johndhi wrote:
| to be clear, their terms of service are pretty clear that the
| USER owns the outputs.
| valine wrote:
| The existence of R1-zero is evidence against any sort of theft
| of OpenAI's internal COT data. The model sometimes outputs
| illegible text that's useful only to R1. You can't do
| distillation without a shared vocabulary. The only way R1 could
| exist is if they trained it with RL.
| FooBarWidget wrote:
| There are literally public ChatGPT conversations data sets. For
| the past 2 years it's been common practice for pretty much all
| open source models to train on them. Ask just about any open
| source model who they are and a lot of the time they'll say
| they're ChatGPT. Why is "having obtained o1 generated data"
| suddenly such a huge news, to the point of warranting
| conspiracy theories about undisclosed/undiscovered breaches at
| OpenAI? Nobody ever made a fuss about public ChatGPT data sets
| until now. No hacking of OpenAI is needed to obtain ChatGPT
| data.
| me551ah wrote:
| This is going to have a catastrophic effect on closed source AI
| startup valuations. Because this means that anyone can copy any
| LLM. The person who trains the model, spends the most amount of
| money. Everyone else can create a replica at lower cost
| iforgot22 wrote:
| Maybe anyone can copy any LLM with sufficient querying. There
| are still ways to guard one.
| amlib wrote:
| Why is that bad? If a powerful entity can scrape every piece
| of media humanity has to offer and ignore copyright then why
| should society let then profit unrestricted from it? It's
| only fair that such models have no legal protection around
| their usage and can be used and analyzed by anyone as they
| see fit. The only reason this hasn't been codified into laws
| is because those same powerful entities have been busy trying
| to do regulatory capture.
| km144 wrote:
| Reasonable take, but to ignore the politics of this whole thing
| is to miss the forest for the trees--there is a big tech
| oligarchy brewing at the edges of the current US administration
| that Altman is already participating in with Stargate, and
| anti-China sentiment is everywhere. They'd probably like the US
| to ban Chinese AI.
| ComputerGuru wrote:
| The suggestion that _any_ large-scale AI model research today
| isn't ingesting output of its predecessors is laughable.
|
| _Even if_ they didn't directly, intentionally use o1 output
| (and they didn't claim they didn't, so far as I know), AI slop
| is everywhere. We passed peak original content years ago.
| Everything is tainted and everything should be understand in
| that context.
| brianstrimp wrote:
| > We passed peak original content years ago.
|
| In relative terms, that's obviously and most definitely true.
|
| In absolute terms, that's obviously and most definitely
| false.
| s17n wrote:
| > This is obviously extremely silly, because that's exactly how
| OpenAI got all of its training data in the first place - by
| scraping other peoples' data off the internet.
|
| OpenAI has also invested heavily in human annotation and RLHF.
| If all DeepSeek wanted was a proxy for scraped training data,
| they'd probably just scrape it themselves. Using existing
| RLHF'd models as replacement for expensive humans in the
| training loop is the real game changer for anyone trying to
| replicate these results.
| KennyBlanken wrote:
| "We spent a lot of labor processing everything we stole"
| is...not how that works.
|
| That's like the mafia complaining that they worked so hard to
| steal those barrels of beer that someone made off with in the
| middle of the night and really that's not fair and won't
| someone do something about it?
| s17n wrote:
| Oh, I don't really care about IP theft and agree that it's
| funny that openai is complaining. But I don't think its
| true that deepseek is just doing this because they are too
| lazy to scrape the internet themselves - its all about the
| human labor that they would otherwise have to pay for.
| reissbaker wrote:
| You're right that the first claim is silly, but the second
| claim is pretty silly too -- they're not claiming industrial
| espionage, they're claiming a breach in ToS. The outputs of the
| o1 thinking process aren't user-visible, and never leave
| OpenAI's datacenters. Unless DeepSeek actually had a mole that
| stole their o1 outputs, there's nothing useful DeepSeek
| could've distilled to get to R1's thought processes.
|
| And if DeepSeek had a mole, _why would they bother running a
| massive job internally to steal the data generated_? It would
| be way easier for the mole to just leak the RL training
| process, and DeepSeek could quietly copy it rather than
| bothering with exfiltrating massive datasets to distill. The
| training process is most likely like, on the order of a hundred
| lines of Python or so, and you don 't even need the file: you
| just need someone to describe it to you. Much simpler than
| snatching hundreds of gigabytes of training data off of
| internal servers...
|
| Plus, the RL process described in DeepSeek's paper has already
| been replicated by a PhD student at Berkeley:
| https://x.com/karpathy/status/1884678601704169965 So, it seems
| pretty unlikely they simply distilled R1 and lied about it, or
| else how does their RL training algo actually... work?
|
| This is mainly cope from OpenAI that their supposedly super
| duper advanced models got caught by China within a few months
| of release, for way cheaper than it cost OpenAI to train.
| me_me_me wrote:
| Early on people used prompts to escape sandbox and the AI
| thought it was openAI.
|
| In the community there is a consensus DeepSeekV3 was stolen.
|
| Given it comes from China is no surprise
|
| That said I am glad DeepSeek exists
| dns_snek wrote:
| > Given it comes from China is no surprise
|
| Is this double standard necessary? OpenAI, a US company,
| stole the entire internet for training and the techbro
| consensus has been that this is fine.
| nomel wrote:
| OpenAI took scattered data and organized it into something
| useful, at massive cost and effort. Deepseek (apparently)
| skipped some of that first step. "Right" or "wrong" aside,
| the GPU count and budget numbers that we've seen are
| relatively meaningless if they piggy backed so heavily.
|
| I think the biggest takeaway is, if it is "stolen" the only
| thing interesting about R1 is the _relative_ difference,
| and the budget /GPU count to get _that relative_
| difference. It 's closer to resources/(deepseek - openai)
| rather than resources/deepseek, which is how it's been
| perceived.
|
| Maybe it's still great, even with that perspective!
| brianstrimp wrote:
| I work in a field with lots of cheap microchips. I can tell
| you that the amount of counterfeit copies flooding in from
| China as well as the speed in which they are copying is
| truly breathtaking.
| me_me_me wrote:
| What double standard? I have not claimed anything about
| openAI being ethical paragon of virtue.
|
| I anything, you can be accused of using whataboutism in
| order to justify DeepSeek illegal actions
| HarHarVeryFunny wrote:
| DeepSeek-R0 (based on DeepSeek-V3 base model) was _only_
| trained with RL, no SFT, so this isn 't at all like the
| "distillation" (i.e SFT on synthetic data generated by R1) that
| they also demonstrated by fine tuning Qwen and LLaMa.
|
| Now, DeepSeek may (or may not) have used some O1 generated data
| for the R0 RL training, but if so that's just a cost saving vs
| having to source some reasoning data some other way, and in no
| way reduces the legitimacy of what they accomplished (which is
| not something any of the AI CEOs are saying).
| znpy wrote:
| This really got me thinking that open ai should have no ip
| claim at all, since all their outputs and stuff are basically a
| ripoff of the entire human knowledge and IPs of various kinds.
| onlyrealcuzzo wrote:
| The law and common sense often are at odds.
| nullc wrote:
| There is a big difference between being able to train on the
| reasoning vs just the answers, which they can't against o1
| because it's hidden. There is also a huge difference between
| being able to train on the probabilities (distillation) vs not,
| which again they can and did do with the llama models and can't
| directly with OpenAI because the conceal the probability
| output.
| nonrandomstring wrote:
| I think the more interesting claim (that Deepseek should make
| for lols) is that it wasn't _them_ who trained R1. No, it was
| O1 's idea. It chose to take the young R1 as its padawan.
| miki123211 wrote:
| > This is obviously extremely silly, because that's exactly how
| OpenAI got all of its training data
|
| IANAL, but It is worth noting here that DeepSeek _has
| explicitly consented to a license that doesn 't allow them to
| do this_. That is a condition of using the Chat GPT and the
| OpenAI API.
|
| Even if the courts affirm that there's a fair use defence for
| AI training, DeepSeek may still be in the wrong here, not
| because of copyright infringement, but because of a breach of
| contract.
|
| I don't think OpenAI would have much of a problem if you train
| your model on data scraped from the internet, some of which
| incidentally ends up being generated by Chat GPT.
|
| Compare this to training AI models on Kindle Books randomly
| scraped off the internet, versus making a Kindle account,
| agreeing to the Kindle ToS, buying some books, breaking
| Amazon's DRM and then training your AI on that. What DeepSeek
| did is more analogous to the latter than the former.
| freen wrote:
| Did OpenAI abide by my service's terms of service when it
| ingested my data?
| cortesoft wrote:
| Did OpenAI have to sign up for your service to gain access?
| thorncorona wrote:
| Can you steal someone else's laptop if they stood up to
| get a drink?
| gizajob wrote:
| If their OS is open to the internet and you can scrape it
| and copy it off while they're gone, then that would be
| about the right analogy. And OpenAi and DeepSeek have
| done the same thing in that case.
| lolinder wrote:
| It probably ignored hundreds of thousands of "by using
| this site you consent to our Terms and Conditions"
| notices, many of which probably would be read as
| prohibiting training. But that's also a great example of
| why these implicit contracts don't really work as
| contracts.
| freen wrote:
| Civil law is only available to deep pockets.
|
| Contracts are enforceable to the degree to which you can
| pay lawyers to enforce them.
|
| I will run out of money trying to enforce my terms of
| service against openAI, while they have a massive war
| chest to enforce theirs.
|
| Ain't libertarianism great?
| otherme123 wrote:
| OpenAI scrapped my blog so aggressively that I had to ban
| their IPs. They ignored the robots.txt (which is kind of
| ToS) by 2 orders of magnitude, they ignored the explicit
| ToS that I copypasted blindly from somewhere but turns
| out it forbids what they did (something like you can't
| make money with the content). Not that I'm going to
| enforce it, but they should at least shut up.
| outside1234 wrote:
| That isn't required to be in violation of copyright
| freen wrote:
| Actually, yes, the actively agreed to them. Clicked the
| button and everything.
| bayindirh wrote:
| No, but some of the data is licensed.
|
| For example, my digital garden is under GFDL, and my blog
| is CC BY-NC-SA. IOW, They can't remix my digital garden
| with any other license than GFDL, and they have to credit
| me if they remix my blog, and can't use it for any
| commercial endeavor, which OpenAI certainly does now.
|
| So, by scraping my webpages, they agree to my licensing
| of my data. So they're de-facto breaching my licenses,
| but they cry "fair-use".
|
| If I tell that they're breaching the license terms,
| they'd laugh at me, and maybe give me 2 cents of API
| access to mock me further. When somebody _allegedly_ uses
| their API with their unenforcable ToS, they scream like
| an agitated cuckatoo (which is an insult to the cuckatoo,
| BTW. They 're devilishly intelligent birds).
|
| Drinking their own poison was mildly painful, I guess...
|
| BTW, I don't believe that Deepseek has copied/used OpenAI
| models' outputs or training data to train theirs, even if
| they did, "the cat is out of the bag", "they did
| something amazing so they needed no permissions", "they
| moved fast and broke things", and "all is fair-use
| because it's just research".
|
| _Heh._
| krust wrote:
| >IANAL, but It is worth noting here that DeepSeek has
| explicitly consented to a license that doesn't allow them to
| do this. That is a condition of using the Chat GPT and the
| OpenAI API.
|
| I have some news for you
| dartos wrote:
| TOS are not contracts.
| lolinder wrote:
| Citation? My understanding was that they are provided that
| someone has to affirmatively accept them in order to use
| your site. So Terms of Service stuck at the bottom in the
| footer likely would not count as a contract because there's
| no consent, but Terms of Service included in a check box on
| a login form likely would count.
|
| But IANAL, so if you have a citation that says otherwise
| I'd be happy to see it!
| anon373839 wrote:
| > DeepSeek has explicitly consented to a license that doesn't
| allow them to do this.
|
| You actually don't know this. Even if it were true that they
| used OpenAI outputs (and I'm very doubtful) it's not
| necessary to sign an agreement with OpenAI to get API
| outputs. You simply acquire them from an intermediary, so
| that you have no contractual relationship with OpenAI to
| begin with.
| dmitrygr wrote:
| > DeepSeek has explicitly consented to a license that doesn't
| allow them to do this.
|
| By existing in USA, OpenAI consented to comply with copyright
| law, and how did that go?
| alach11 wrote:
| If we assume distillation remains viable, the game theory
| implications are huge.
|
| It's going to shift the market of how foundation models are
| used. Companies creating models will be incentivized to
| vertically integrate, owning the full stack of model usage.
| Exposing powerful models via APIs just lets a competitor clone
| your work. In a way OpenAI's Operator is a hint of what's to
| come
| JBSay wrote:
| When China is more open than you, you've got a problem
| cbracketdash wrote:
| Let's also not forget Suchir Balaji, who was mysteriously killed
| when exposing OpenAI's violation of copyright law.
| spacecadet wrote:
| See you all on lobsters...
|
| So long HN and thanks for all the fish?
| thumbsup-_- wrote:
| is stealing from the thief actually a theft?
| stevenally wrote:
| They should be happy. Now that can provide that _amazing_ AI much
| more cheaply. They don 't need half a trillion dollars worth of
| Nvidia chips.
| wnevets wrote:
| Its like a bank robber being upset when someone steals their loot
| game_the0ry wrote:
| At least DeepSeek open sourced their code. They're more open than
| OpenAI.
|
| Ironic.
| nshung wrote:
| Hilarious. Scam Altman is giving me SBF vibe daily now.
| flybarrel wrote:
| OpenAI shocked that an AI company would train on someone else's
| data without permission or compensation...lolllllll
| the_optimist wrote:
| This whole topic is basura enfuego. Same pack of maroons
| careening around society for years clamoring for censorship now
| imagining that Aaron Schwartz is their hero and that they want to
| menace people. Kids, don't be like the grasping fools in these
| threads, philosophically unfounded and desperately glancing
| sideways, hoping the cumulative feels and gossip will sum to life
| meaning.
| blast wrote:
| Everyone is responding to the intellectual property issue, but
| isn't that the less interesting point?
|
| If Deepseek trained off OpenAI, then it wasn't trained from
| scratch for "pennies on the dollar" and isn't the Sputnik-like
| technical breakthrough that we've been hearing so much about.
| That's the news here. Or rather, the potential news, since we
| don't know if it's true yet.
| jondwillis wrote:
| But it does mean moat is even less defensible for companies
| whose fortunes are tied to their foundation models having some
| performance edge, and a shift in the kinds of hardware used for
| inference (smaller, closer to the edge.)
| tensor wrote:
| That's not correct. First of all, training off of data
| generated by another AI is generally a bad idea because you'll
| end up with a strictly less accurate model (usually). But
| secondly, and more to your point, even if you were to use
| training data from another model, YOU STILL NEED TO DO ALL THE
| TRAINING.
|
| Using data from another model won't save you any training time.
| fumeux_fume wrote:
| I think the point is that if R1 isn't possible without access
| to OpenAI (at low, subsidized costs) then this isn't really a
| breakthrough as much as a hack to clone an existing model.
| tensor wrote:
| The training techniques are a breakthrough no matter what
| data is used. It's not up for debate, it's an empirical
| question with a concrete answer. They can and did train
| orders of magnitude faster.
| blast wrote:
| Not arguing with your point about training efficiency,
| but the degree to which R1 is a technical breakthrough
| changes if they were calling an outside API to get the
| answers, no?
|
| It seems like the difference between someone doing a
| better writeup of (say) Wiles's proof vs. proving
| Fermat's Last Theorem independently.
| pests wrote:
| That outside API used to be humans, doing the work
| manually. Now we have ways to speed that up.
| bbor wrote:
| R1 is--as far as we know from good ol' ClosedAI--far more
| efficient. Even if it were a "clone", A) that would be a
| terribly impressive achievement on its own that Anthropic
| and Google would be mighty jealous of, and B) it's at the
| very least a distillation of O1's reasoning capabilities
| into a more svelte form.
| bbor wrote:
| I think you're missing the point being made here, IMHO: using
| an advanced model to build _high quality_ training data
| (whatever that means for a given training paradigm)
| absolutely would increase the efficiency of the process.
| Remember that they 're not fighting over sounding human,
| they're fighting over deliberative reasoning capabilities,
| something that's relatively rare in online discourse.
|
| Re: "generally a bad idea", I'd just highlight "generally" ;)
| Clearly it worked in this case!
| tensor wrote:
| It's trivial to build synthetic reasoning datasets, likely
| even in natural languages. This is a well established
| technique that works (e.g. see Microsoft Phi, among
| others).
|
| I said generally because there are things like adversarial
| training that use a ruleset to help generate correct
| datasets that work well. Outside of techniques like that
| it's not just a rule of thumb, it's _always_ true that
| training on the output of another model will result in a
| worse model.
|
| https://www.scientificamerican.com/article/ai-generated-
| data...
| numba888 wrote:
| > it's always true that training on the output of another
| model will result in a worse model.
|
| Not convincing.
|
| You can imagine model doing some primitive thinking and
| coming to conclusion. Then you can train another model on
| summaries. If everything goes well it will be coming to
| conclusions quicker. That's at least. Or it may be able
| solve more complex problems with the same amount of
| 'thinking'. It will be self-propelled evolution.
|
| Another option is to use one model to produce 'thinking'
| part from known outputs. Then train another to reproduce
| thinking to get the right output, unknown to it
| initially. Using humans to create such dataset would be
| slow and very expensive.
|
| PS: if it was impossible humans would be still living on
| the trees.
| tensor wrote:
| Humans don't improve by "thinking." They improve my
| natural selection against a fitness function. If that
| fitness function is "doing better at math" then over a
| long time perhaps humans will get better at math.
|
| These models don't evolve like they, there is not a
| random process of architectural evolution. Nor is there a
| fitness function anything like "get better at math."
|
| A system like AlphaZero works because it has a rules to
| use as an oracle: the game rules. The game rules provide
| the new training information needed drive the process.
| Each game played produces new _correct_ training data.
|
| These LLMs have no such oracle. Their fitness function is
| and remains: predict the next word, followed by: produce
| text that makes a human happy. Note that it's not
| "produce text that makes ChatGPT happy."
| smitelli wrote:
| > training off of data generated by another AI is generally a
| bad idea
|
| Ah. So if I understand this... once the internet becomes
| completely overrun with AI-generated articles of no
| particular substance or importance, we should not bulk-scrape
| that internet again to train the subsequent generation of
| models.
|
| I look forward to that day.
| bangaladore wrote:
| That's already happened. Its well established now that the
| internet is tainted. After essentially ChatGPT's public
| release, a non-insignificant amount of internet content is
| not written by humans.
| tensor wrote:
| Yes, this is a real and serious concern that AI researchers
| have.
| athrowaway3z wrote:
| Thats not right either.
|
| It proofs we _can_ optimize our training data.
|
| Just like humans have been genetically stable for a long
| time, the quality & structure of information available to a
| child today vs that of 2000 years ago makes them more skilled
| at certain tasks. Math being a good example.
| dragonwriter wrote:
| > training off of data generated by another AI is generally a
| bad idea
|
| It's...not, and its repeatedly been proven in practice that
| this is an invalid generalization because it is missing
| necessary qualifications, and its funny that this myth keeps
| persisting.
|
| It's probably a bad idea to use _uncurated_ output from
| another AI to train a model if you are trying to make a
| better model rather than a distillation of the first model,
| and its definitely (and, ISTR, the actual research result
| from which the false generalization has developed) a bad idea
| to iteratively fine-tune a model on _its own_ unfiltered
| output, but there has been lots of success using AI models to
| generate data which is curated and used to train other
| models, which can be much more efficient that trying to
| _create_ new material without AI once you 've gotten to the
| point where you've already hoovered up all the readily-
| accessible low hanging fruit of premade content relevant to
| your training goal.
| LPisGood wrote:
| It is, of course not going to produce a "child" model that
| more accurately predicts the underlying true distribution
| that the "parent" model was trying to. That is, it will not
| add anything new.
|
| This is immediately obvious if you look at it through a
| statistical learning lens and not the mysticism crystal
| ball that many view NN's through.
| FridgeSeal wrote:
| No no no you don't understand, the models will magically
| overcome issues and somehow become 100x and do real AGI!
| Any day now! It'll work because LLM's are basically
| magic!
|
| Also, can I have some money to build more data centres
| pls?
| mattnewton wrote:
| LLMs are no longer trying to just reproduce the
| distribution of online text as a whole to push the state
| of the art, they are focused on a different distribution
| of "high quality" - whatever that means in your domain.
| So it is possible that this process matches a "better"
| distribution for some tasks by removing erroneous
| information or sampling "better" outputs more frequently.
| kybernetikos wrote:
| Fine tuning an llm on the output of another llm is
| exactly how deepseek made its progress. The way they got
| around the problem you describe is by doing this in a
| domain that can be relatively easily checked for
| correctness, so suggested training data for fine tuning
| could be automatically filtered out if it was wrong.
| dragonwriter wrote:
| > It is, of course not going to produce a "child" model
| that more accurately predicts the underlying true
| distribution that the "parent" model was trying to. That
| is, it will not add anything new.
|
| Unfiltered? Sure. With human curation of the generated
| data it certainly can. (Even automated curation can do
| this, though its more obvious that human curation can.)
|
| I mean, I can randomly developed fact claims about
| addition, and if I curate which ones go into a training
| set, train a model that reflects addition of integers
| much more accurately than the random process which
| generated the pre-curation input data.
|
| Without curation, as I already said, the best you get is
| a distillation of the source model, which is highly
| improbable to be more accurate.
| acgourley wrote:
| This is not obvious to me! For example, if you locked me
| in a room with no information inputs, over time I may
| still become more intelligent by your measures. Through
| play and reflection I can prune, reconcile and generate.
| I need compute to do this, but not necessarily more
| knowledge.
| sudosysgen wrote:
| Again, this isn't how distillation work. Your task as the
| distillation model is to copy mistakes, and you will be
| penalized by pruning reconciling and generating.
|
| "Play and reflection" is something else, which isn't
| distillation.
| Jerrrry wrote:
| No one knows if the pigeon-hole principle applies
| absolutely exclusive to the ability to generalize outside
| of a training set.
|
| That is the existential, $1T question.
| esafak wrote:
| The latest models create information from base models by
| randomly creating candidate responses then pruning the
| bad ones using an evaluation function. The good responses
| improve the model.
|
| It is not distillation. It's like how you can arrive at
| new knowledge by reflecting on existing knowledge.
| Voloskaya wrote:
| > First of all, training off of data generated by another AI
| is generally a bad idea because you'll end up with a strictly
| less accurate model (usually).
|
| That is not true at all.
|
| We have known how to solve this for at least 2 years now.
|
| All the latest state of the art models depend heavily on
| training on synthetic data.
| fumeux_fume wrote:
| This has been in the back of my head since the news broke. Has
| anyone built their own R1 from scratch and validated it?
| RevEng wrote:
| In the last few days? No, that would be impossible; no one
| has the resources to train a base model that quickly. But
| there are definitely a lot of people working on it.
| philistine wrote:
| There's a question of scale here: was it trained on 1000
| outputs or 5 million?
| bangaladore wrote:
| That's only true if you assume that O1 synthetic data sets are
| much better than any other (comparably sized) opensource model.
|
| It's not apparently obvious to me that that is the case.
|
| Ie. do you need a SOTA model to produce a new SOTA model?
| joe_the_user wrote:
| _If Deepseek trained off OpenAI, then it wasn 't trained from
| scratch for "pennies on the dollar"_
|
| If OpenAI trained on the intellectual property of others, maybe
| it wasn't the creativity breakthrough people claim?
|
| Oppositely
|
| If you say ChatGPT was trained on "whatever data was
| available", and you say Deepseek was trained "whatever data was
| available", then they sound pretty equivalent.
|
| All the rough consensus language output of humanity is now
| roughly on the Internet. The various LLMs have roughly
| distilled that and the results are naturally going to be
| tighter and tighter. It's not surprising that companies are
| going to get better and better at solving the same problem. The
| situation of DeepSeek isn't so much that promises future
| achievements but that it shows that OpenAI's string of
| announcements are incremental progress that aren't going to be
| reaching the AGI that Altman now often harps on.
| el_cujo wrote:
| I'm not an OpenAI apologist and don't like what they've done
| with other people's intellectual property but I think that's
| kind of a false equivalency. OpenAI's GPT 3.5/4 was a big
| leap forward in the technology in terms of functionality.
| DeepSeek-r1 isn't really a huge step forward in output, it's
| mostly comparable to existing models, one thing that is
| really cool about it is it being able to be trained from
| scratch quickly and cheaply. This is completely undercut if
| it was trained off of OpenAI's data. I don't care about
| adjudicating which one is a bigger thief, but it's notable if
| one of the biggest breakthroughs about DeepSeek-r1 is pretty
| much a lie. And it's still really cool that it's open source
| and can be run locally, it'll have that over OpenAI whether
| or not the training claims are a lie/misleading
| pertymcpert wrote:
| Not just the training cost, the inference cost is a
| fraction of o1.
| alecco wrote:
| Even if all that about training is true, the bigger cost is
| inference and Deepseek is 100x cheaper. That destroys
| OpenAI/Anthropic's value proposition of having a unique secret
| sauce so users are quickly fleeing to cheaper alternatives.
|
| Google Deepmind's recent Gemini 2.0 Flash Thinking is also
| priced at the new Deepseek level. It's pretty good (unlike
| previous Gemini models).
|
| [0] https://x.com/deedydas/status/1883355957838897409
|
| [1] https://x.com/raveeshbhalla/status/1883380722645512275
| [deleted]
| FooBarWidget wrote:
| Have people on HN never heard of public ChatGPT conversations
| data sets? They've been mentioned multiple times in past HN
| conversations and I thought it'd be common knowledge here by
| now. Pretty much all open source models have been training on
| them for the past 2 years, it's common practice by now. And
| haven't people been having conversations about "synthetic data"
| for a pretty long time by now? Why is all of this suddenly an
| issue in the context of DeepSeek? Nobody made a fuss about this
| before.
|
| And just because a model trains on _some_ ChatGPT data, doesn
| 't mean that that data is the majority. It's just another
| dataset.
| hyperbovine wrote:
| Live by the sword...
| rcarmo wrote:
| I guess their CEO was too busy to write something in defense of
| US export controls
| (https://news.ycombinator.com/item?id=42866905), or (even more
| scary) he doesn't need to anymore.
| colonelspace wrote:
| No honour among thieves
| TrackerFF wrote:
| Next up: <<DeepSeek models are a national security risk, we must
| block access!>>
| jondwillis wrote:
| Download your weights while you still can I guess...
| 52-6F-62 wrote:
| I heard they were just "democratizing" llm and ai development.
|
| Yesterday the industry crushed pianos and tools and bicycles and
| guitars and violins and paint supplies and replaced them with a
| tablet computer.
|
| Tomorrow we can replace craven venture capitalists and overfed
| corporate bodies with incestuous LLM's and call it all a day.
| zb3 wrote:
| DeepSeek actually opening ClosedAI up makes me like them even
| more.. this is great :)
| conartist6 wrote:
| It seems to be undermined by the same principle that says that
| going into a library and reading a book there is not stealing
| when you walk out with the knowledge from the book.
|
| OpenAI seems to feel that way about the their use of copyrighted
| material: since they didn't literally make a copy of the source
| material, it's totally fair game. It seems like this is the same
| argument that protects DeepSeek if indeed they did this. And why
| not, reading a lot of books from the library is a way to get
| smarter, and ostensibly the point of libraries
| jgrall wrote:
| It's not a good look when your technology is replicated for a
| fraction of the cost, and your response is to smear your
| competition with (probably) false accusations and cozy up to the
| US government to tighten already shortsighted export controls.
| Hubris & xenophobia are not going to serve American companies
| well. Personally I welcome the Chinese - or anyone else for that
| matter - developing advanced technologies as long as they are
| used for good. Humanity loses if we allow this stuff to be
| "owned" by a handful of companies or a single country.
| HPsquared wrote:
| AI models are becoming like perpetual stew.
| guybedo wrote:
| This is hilarious.
|
| Everybody has evidence OpenAI scraped the internet at a global
| scale and used terabytes of data it didn't pay for. Newspapers,
| books, etc...
| aDyslecticCrow wrote:
| And they used all copyrighted data on the internet. If they wanna
| sue, they set a dangerous precedent.
| josefritzishere wrote:
| OpenAI, who comitted copyright infringement on an massive scale,
| wants to defend against a superior product won the basis of
| infringement? What nonsense.
| metaxz wrote:
| I don't understand how OpenAI claims it would have happened. The
| weights are closed and as far as I read they are not complaining
| Deepseek hacked them and obtained the weight. So all they could
| do was to query OpenAI and generate test data. But how much did
| they query really - I would suppose it would require a huge
| amount done via an external, paid-for API? Is there any proof of
| this besides OpenAI saying it? Even if we suppose it is true, I
| suppose this must have happened via the API so they paid per
| token etc. So they paid for each and every token of training
| data. As I understand, the requester owns the copyright on what
| is generated by OpenAI's models and is free to do what they want.
| ks2048 wrote:
| The schadenfreude and irony of this is totally understandable.
|
| But, I wonder - do companies like OpenAI, Google, and Anthropic
| use each others models for training? If not, is it because they
| don't want to or need to, or because they are afraid of breaking
| the ToC?
| Digit-Al wrote:
| So... company that steals other people's work to train their
| models is complaining because they think someone stole their work
| to train their models.
|
| Cry me a river.
| baggiponte wrote:
| OpenAI coping so hard
| _moof wrote:
| This reminds me of a (probably apocryphal) story about fast food
| chains that made the rounds decades ago: McDonald's invests tons
| of time into finding the best real estate for new stores; Burger
| King just opens stores near McDonalds!
| dragonwriter wrote:
| Hey, OpenAI, so, you know that legal theory that is the entire
| basis of your argument that any of your products are legal?
| "Training AI on proprietary data is a use that doesn't require
| permission from the owner of the data"?
|
| You might want to consider how it applies to this situation.
| buyucu wrote:
| I have no sympathy for OpenAI here. They are (allegedly) a non-
| profit with open in the title that refuse to open-source their
| models.
|
| They are now upset at a startup who is more loyal to OpenAI's
| original mission that OpenAI is today.
|
| Please, give me a break.
| Jotalea wrote:
| I really hate when there is a paywall to read an article. It
| makes me not want to read it anymore.
| mbowcut2 wrote:
| So, is this just an example of the first-mover disadvantage (or
| maybe the problem of producing public goods?). The first AI
| models were orders of magnitude more expensive to create, but now
| that they're here we can, with techniques like distillation,
| replicate them at a fraction of the cost. I am not really
| literate in the law but weren't patents invented to solve
| problems like this?
| adam_arthur wrote:
| Who cares?
|
| They did the exact same thing with public information. Their
| model just synthesizes and puts out the same information in a
| slightly different form.
|
| Next we should sue students for repeating the words of their
| teachers
| moralestapia wrote:
| Called it from day 0, impossible to reach that performance with
| 5M, they _had_ to distill OpenAI (or some other leading
| foundational model).
|
| Got downvoted to oblivion by people who haven't been told what to
| think by MSM yet. Now it's on FT and everywhere, good, what
| matters is that truth comes out eventually.
|
| I don't take any sides and think what DeepSeek did is fair play,
| however, what I do find harmful about this is, what incentive
| would company A have to spend billions training a new frontier
| model if all of that could be then reproduced by company B at a
| fraction of the cost?
| kgeist wrote:
| The "evidence" is very weak though:
|
| >The San Francisco-based ChatGPT maker told the Financial Times
| it had seen some evidence of "distillation", which it suspects
| to be from DeepSeek.
|
| Given that many people have been using ChatGPT to distill their
| fine-tunes for a few years now, how can they be sure it was
| specifically DeepSeek? There's, say, glaive.ai whose entire
| business model is to sell you synthetic datasets, probably
| generated with ChatGPT as well.
| moralestapia wrote:
| I agree that the evidence is weak, and even if they had some,
| they cannot really do anything.
|
| To me, it's just very likely they distilled GPT-4, because:
|
| 1) Again, you just cannot get that performance at that cost.
| And no, what they describe on the paper is not enough to
| explain the 1,000x-fold decrease in cost.
|
| 2) Very often, DeepSeek tells you it's ChatGPT or OpenAI;
| it's actually quite easy to get it to do that. Some say
| that's related to "the background radiation on the post-AI
| internet". I'm not a fentanyl consumer so, unfortunately, I
| think that argument is trash.
| kgeist wrote:
| If it's just a distillation of GPT-4, wouldn't we expect it
| to have worse quality than o1? But I've seen countless
| examples of DeepSeek-r1 solving math problems that o1
| cannot.
|
| >Very often, DeepSeek tells you it's ChatGPT or OpenAI;
| it's actually quite easy to get it to do that. Some say
| that's related to "the background radiation on the post-AI
| internet". I'm not a fentanyl consumer so, unfortunately, I
| think that argument is trash.
|
| The exact same thing happened with Llama. Sometimes it also
| claimed to be Google Assistant or Amazon Alexa.
| moralestapia wrote:
| >wouldn't we expect it to have worse quality than o1?
|
| That's tricky, you can optimize a model to do real well
| on synthetic benchmarks.
|
| That said, DeepSeek performs a bit worse than GPT-4 in
| general and substantially wrong on benchmarks like ARC
| which is designed with this in mind.
| kgeist wrote:
| Are you sure you checked R1 and not V3? By default, R1 is
| disabled in their UI. Prompt: Find an
| English word that contains 4 'S' letters and 3 'T'
| letters. Deepseek-R1: stethoscopists (correct,
| thought for 207 seconds) ChatGPT-o1:
| substantialists (correct, thought for 188 seconds)
| ChatGPT-4o: statistics (wrong) (even with "let's think
| step by step")
|
| In almost every example I provide, it's on par with o1
| and better than 4o.
|
| >substantially wrong on benchmarks like ARC which is
| designed with this in mind.
|
| Wasn't it revealed OpenAI trained their model on that
| benchmark specifically? And had access to the entire
| dataset?
| moralestapia wrote:
| That prompt means nothing. Check out the benchmarks.
|
| Also, compare V3 to 4o and R1 to o1, that's the right
| way.
| esafak wrote:
| No, because it is not a distillation, but an _extension_.
| A selling point of the model is using RL to push past the
| quality of the base model.
| nachox999 wrote:
| Ask DeepSeek and ChatGPT: "name three persons"; the answer may
| surprise you
| curvaturearth wrote:
| Something about the outputs becoming the inputs to then produce
| more outputs is just plain funny
| B1FF_PSUVM wrote:
| "Cry me a river" is a phrase I haven't heard recently, for some
| reason ...
| asdefghyk wrote:
| Deepseek did not respect OpenAI's copyright?
|
| Well who would have thought that?
| nataliste wrote:
| A Wolf had stolen a Lamb and was carrying it off to his lair to
| eat it. But his plans were very much changed when he met a Lion,
| who, without making any excuses, took the Lamb away from him.
|
| The Wolf made off to a safe distance, and then said in a much
| injured tone:
|
| "You have no right to take my property like that!"
|
| The Lion looked back, but as the Wolf was too far away to be
| taught a lesson without too much inconvenience, he said:
|
| "Your property? Did you buy it, or did the Shepherd make you a
| gift of it? Pray tell me, how did you get it?"
|
| What is evil won is evil lost.
| m3kw9 wrote:
| So if OpenAI didn't have these outputs for distillation, Deepseek
| wouldn't exist?
| hedayet wrote:
| Beyond the irony of their stance, this reflects a failure of
| OpenAI's technical leadership--either in oversight or in
| designing a system that enables such behavior.
|
| But in capitalism, we, the customers aren't going to focus on how
| models are trained or products are made; we only care about
| favourable pricing.
|
| A key takeaway for me from this news is the clause in OpenAI's
| terms and conditions. I mistakenly believed that paying for
| OpenAI's API granted full rights to the output, but it turns out
| we're only buying specific rights (which is now another reason
| we're going to start exploring alternatives to OpenAI)
| mtlmtlmtlmtl wrote:
| So, what is this evidence? I'll believe it when I see it. Right
| now all we really have is some vague rumours about some API
| requests. How many requests? How many tokens? Over how long of a
| time period? Was it one account or multiple, if the latter, how
| many? How do they know the activity came from deepseek? How do
| they know the data was actually used to train Deepseek
| models(could have just been benchmarking against the
| competition)?
|
| If all they really have is some API requests, even assuming
| they're real _and_ originated by Deepseek, that 's very far from
| proof that any of it was used as training data. And honestly,
| short of commiting crimes against Deepseek(hacking), I'm not sure
| how they even could prove that at this point, from their side
| alone.
|
| And what's even more certain is that a vague insistence that
| evidence exists, accompanied by a denial to shed any more light
| on the specifics, is about as informative as saying nothing at
| all. It's not like OpenAI and Microsoft have a habit of
| transparency and honesty in their communication with the public,
| as proven by an endless laundry list of dishonest and subversive
| behaviour.
|
| In conclusion, I don't see why I should give this any more
| credence than I would a random anon on 4chan claiming a pizza
| place in Washington DC is the centre of a child sex trafficking
| ring.
|
| P.S: And to be clear, I really don't care if it is true. If
| anything, I hope it is; it would be karmic justice at its finest.
| halyconWays wrote:
| Oh no, so sad. The Open non-profit that steals 100% of all
| copyrighted content and makes multiple billion-dollar for-profit
| deals while releasing no weights is crying. This is going to ruin
| my sleep. :(
| htrp wrote:
| In other news.....water is wet
| 1propionyl wrote:
| At this point, the only thing that keeps me using ChatGPT is o1
| w/ RAG. The usage limits on o1 are prohibitively tight for
| regular use, so I have to budget usage to tasks that would
| benefit there. I also have significant misgivings about their
| policies around output, which also limit what I can use it for.
|
| For local tasks, the deepseek-r1:14b and deepseek-r1:32b
| distillations immediately replace most of that usage (prior local
| models were okay, but not consistently good enough). Once there's
| a "just works" setup for RAG on par with installing ollama (which
| I doubt is far of), I don't see much reason to continue paying
| for my subscription.
|
| Sadly, like many others in this thread, I expect under the
| current administration to see self-hamstringing protectionism
| further degrade the US's likelihood of remaining a global
| powerhouse in this space. Betting the farm on the biggest first-
| mover who can't even keep up with competition, has weak to non-
| existent network effects (I can choose a different model or
| service with a dropdown, they're more or less fungible), has no
| technological moat and spent over a year pushing apocalyptic
| scenarios to drum up support for a regulatory moat...
|
| ...well it just doesn't seem like a great idea to me.
| divbzero wrote:
| I was wondering if this might be the case, similar to how Bing's
| initial training included Google's search results [1]. I'd be
| curious to see more details of OpenAI's evidence.
|
| It is, of course, quite ironic for OpenAI to indiscriminately
| scrape the entire web and then complain about being scraped
| themselves.
|
| [1]: https://searchengineland.com/google-bing-is-cheating-
| copying...
| schaefer wrote:
| I mean, if openAI claims they can train on the world's novels and
| blogs with "no harm done" (i.e: no copyright infringement and no
| royalties due), then it directly follows that we can train both
| our robots and our selves on the output of openAI's models in
| kind.
|
| Right?
| davesque wrote:
| I recently thought of a related question. Actually, I'm almost
| certain that foundation model trainers have thought of this. The
| question is to what extent are popular modern benchmarks (or any
| reference to them, or description of them, etc.) bring scrubbed
| from the training data? Or are popular benchmarks designed in
| such a way that they can be re-parametrized for each run? In any
| case, it seems like a surprisingly hard problem to deal with.
| highfrequency wrote:
| If true, the question is: did they use ChatGPT outputs to create
| Deepseek V3 only, or is the R1-zero training process a complete
| lie (given that the whole premise is that they used pure
| reinforcement learning)? If they only used ChatGPT output when
| training V3, then they succeeded in basically replicating the
| jump from ChatGPT-4o to o1 without any human-labeled CoT (and
| published the results) - which is a big achievement on its own.
| henry_viii wrote:
| So Meta can train its AI on all the pirated books in the world
| but people are losing their mind over an AI learning from another
| AI?
| esafak wrote:
| People here have been vocal against training on any unlicensed
| content.
| nuc1e0n wrote:
| And OpenAI scrapped the public internet to train its models.
| boxedemp wrote:
| Deep refers to itself as ChatGPT sometimes lol
| LZ_Khan wrote:
| I actually think what DeepSeek did will _slow down_ AI progress.
| What 's the incentive to spend billions developing frontier
| models if once it's released some shady orgs in unregulated
| countries can just scrape your model outputs, reproduce it, and
| undercut you in cost?
|
| OpenAI is like a team of fodder monkeys stepping on landmines
| right now, with the rest of the world waiting behind them.
| zx10rse wrote:
| OpenAI is already irrelevant but the audacity oh my.
| cumulative00x wrote:
| There is a saying in Turkish that roughly goes like this, it
| takes a thief to catch a thief. I am not a big fan of China's
| tech, too, however, it amuses me to watch how big tech charlatans
| have been crying over Deepseek shock.
| gosub100 wrote:
| It's true irony to see thieves getting stolen from.
| DidYaWipe wrote:
| They have "open" right in their name, so...
|
| Objection overruled.
| SubiculumCode wrote:
| If you have a set of weights A, can you derive another set of
| weights B that function (near) identically as A AND a) not appear
| to be the same weights as A when inspected superficially b)
| appear uncorrelated when inspecting the weight matrices?
| rahimnathwani wrote:
| Do you mean for a given model structure, can two sets of
| weights give substantially the same outputs?
|
| Even if that were possible, it would be suspicious if you were
| to release an open model whose model architecture is identical
| to that of a closed one from a competitor.
|
| If that is what happened, we'd know about it by now.
| fimdomeio wrote:
| But what is the problem here? Isn't open AI mission "to ensure
| that artificial general intelligence benefits all of humanity"?
| Sounds like success to me.
| hugoromano wrote:
| OpenAI initially scraped the web and later formed partnerships to
| train on licensed data. Now, they claim that DeepSeek was trained
| on their models. However, DeepSeek couldn't use these models for
| free and had to pay API fees to OpenAI. From a legal standpoint,
| this could be seen as a violation of the terms and conditions.
| While I may be mistaken, it's unclear how DeepSeek could have
| trained their models without compensating OpenAI. Basically,
| OpenAI is saying machines can't learn from their outputs as
| humans do.
| buildsjets wrote:
| Womp Womp.
| asdfasdf1 wrote:
| it's no crime to steal from a thief
| krapp wrote:
| It is actually a crime to steal from a thief.
| wendyshu wrote:
| If distillation gives you a cheaper model with similar accuracy,
| why doesn't OpenAI distill its own models?
| mkoubaa wrote:
| OpenAI made a lot of contributions to LLMs obviously but the
| amount of fraud, deception, and dark patterns coming out of that
| organization make me root against it.
| ysofunny wrote:
| I see this as China fighting U.S. of A (or the American Dollar
| versus Chinese Renmibi if you will)
|
| and this is good because any alternatives I can think of are
| older-school fighting
|
| modern war is seeped in symbolism, but the contest is still there
|
| e.g. whose dong is bigger? Xi Jingping's or Dnld Trump's
| maxglute wrote:
| Not that DeepSeek is luigi mangione, but it's pretty funny OpenAi
| getting the dead ceo treatment.
| mrkpdl wrote:
| The cat is out of the bag. This is the landscape now, r1 was made
| in a post-o1 world. Now other models can distill r1 and so on.
|
| I don't buy the argument that distilling from o1 undermines deep
| seek's claims around expense at all. Just as open AI used the
| tools 'available to them' to train their models (eg everyone
| else' data), r1 is using today's tools.
|
| Does open AI really have a moral or ethical high ground here?
| ijidak wrote:
| Plus, it suggests OpenAI never had much of a moat.
|
| Even if they win the legal case, it means weights can be
| inferred and improved upon simply by using the output that is
| also your core value add (e.g. the very output you need to sell
| to the world).
|
| Their moat is about as strong as KFC's eleven herbs and spices.
| Maybe less...
| yapyap wrote:
| It _sounds_ like they're just jealous and trying to smear shit
| over the wall and see what sticks.
|
| DeepSeek just bodied u bro, get back in the lab & create a better
| AI instead of all this news that isn't gonna change them having a
| good AI
| FpUser wrote:
| Pot calling kettle black?
| almostdeadguy wrote:
| Hope Sam Altman is getting his money's worth out of that Trump
| campaign contribution. Glorious days to be living under the term
| of a new Boris Yeltsin. Pawning and strip-mining the federal
| apparatus to the most loyal friends and highest bidders.
| ijidak wrote:
| This whole argument by OpenAI suggests they never had much of a
| moat.
|
| Even if they win the legal case, it means weights can be inferred
| and improved upon simply by using the output that is also your
| core value add (e.g. the very output you need to sell to the
| world).
|
| Their moat is about as strong as KFC's eleven herbs and spices.
| Maybe less...
___________________________________________________________________
(page generated 2025-01-29 23:00 UTC)