[HN Gopher] DeepSeek-R1: Incentivizing Reasoning Capability in L...
       ___________________________________________________________________
        
       DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
        
       Author : gradus_ad
       Score  : 334 points
       Date   : 2025-01-25 18:39 UTC (4 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | siliconc0w wrote:
       | The US Economy is pretty vulnerable here. If it turns out that
       | you, in fact, don't need a gazillion GPUs to build SOTA models it
       | destroys a lot of perceived value.
       | 
       | I wonder if this was a deliberate move by PRC or really our own
       | fault in falling for the fallacy that more is always better.
        
         | refulgentis wrote:
         | I've been confused over this.
         | 
         | I've seen a $5.5M # for training, and commensurate commentary
         | along the lines of what you said, but it elides the cost of the
         | base model AFAICT.
        
           | logicchains wrote:
           | $5.5 million is the cost of training the base model, DeepSeek
           | V3. I haven't seen numbers for how much extra the
           | reinforcement learning that turned it into R1 cost.
        
             | refulgentis wrote:
             | Ahhh, ty ty.
        
           | m_a_g wrote:
           | With $5.5M, you can buy around 150 H100s. Experts correct me
           | if I'm wrong but it's practically impossible to train a model
           | like that with that measly amount.
           | 
           | So I doubt that figure includes all the cost of training.
        
             | logicchains wrote:
             | The cost, as expressed in the DeepSeek V3 paper, was
             | expressed in terms of training hours based on the market
             | rate per hour if they'd rented the 2k GPUs they used.
        
             | etc-hosts wrote:
             | It's even more. You also need to fund power and maintain
             | infrastructure to run the GPUs. You need to build fast
             | networks between the GPUs for RDMA. Ethernet is going to be
             | too slow. Infiniband is unreliable and expensive.
        
               | FridgeSeal wrote:
               | You'll also need sufficient storage, and fast IO to keep
               | them fed with data.
               | 
               | You also need to keep the later generation cards from
               | burning themselves out because they draw so much.
               | 
               | Oh also, depending on when your data centre was built,
               | you may also need them to upgrade their power and cooling
               | capabilities because the new cards draw _so much_.
        
         | logicchains wrote:
         | >I wonder if this was a deliberate move by PRC or really our
         | own fault in falling for the fallacy that more is always
         | better.
         | 
         | DeepSeek's R1 also blew all the other China LLM teams out of
         | the water, in spite of their larger training budgets and
         | greater hardware resources (e.g. Alibaba). I suspect it's
         | because its creators' background in a trading firm made them
         | more willing to take calculated risks and incorporate all the
         | innovations that made R1 such a success, rather than just
         | copying what other teams are doing with minimal innovation.
        
         | jvanderbot wrote:
         | How likely is this?
         | 
         | Just a cursory probing of deepseek yields all kinds of
         | censoring of topics. Isn't it just as likely Chinese sponsors
         | of this have incentivized and sponsored an undercutting of
         | prices so that a more favorable LLM is preferred on the market?
         | 
         | Think about it, this is something they are willing to do with
         | other industries.
         | 
         | And, if LLMs are going to be engineering accelerators as the
         | world believes, then it wouldn't do to have your software
         | assistants be built with a history book they didn't write.
         | Better to dramatically subsidize your own domestic one then
         | undercut your way to dominance.
         | 
         | It just so happens deepseek is the best one, but whichever was
         | the best Chinese sponsored LLM would be the one we're supposed
         | to use.
        
           | refulgentis wrote:
           | You raise an interesting point, and both of your points seem
           | well-founded and have wide cache. However, I strongly believe
           | both points are in error.
           | 
           | - OP elides costs of anything at all outside renting GPUs,
           | and they purchased them, paid GPT-4 to generate training
           | data, etc. etc.
           | 
           | - Non-Qwen models they trained are happy to talk about ex.
           | Tiananmen
        
           | logicchains wrote:
           | >Isn't it just as likely Chinese sponsors of this have
           | incentivized and sponsored an undercutting of prices so that
           | a more favorable LLM is preferred on the market?
           | 
           | Since the model is open weights, it's easy to estimate the
           | cost of serving it. If the cost was significantly higher than
           | DeepSeek charges on their API, we'd expect other LLM hosting
           | providers to charge significantly more for DeepSeek (since
           | they aren't subsidised, so need to cover their costs), but
           | that isn't the case.
           | 
           | This isn't possible with OpenAI because we don't know the
           | size or architecture of their models.
           | 
           | Regarding censorship, most of it is done at the API level,
           | not the model level, so running locally (or with another
           | hosting provider) is much less expensive.
        
           | siltcakes wrote:
           | I trust China _a lot_ more than Meta and my own early tests
           | do indeed show that Deepseek is far less censored than Llama.
        
         | tayo42 wrote:
         | More effecient use of hardware just increases productivity. No
         | more people/teams can interate faster and in parralel
        
         | thelastparadise wrote:
         | But do we know that the same techniques won't scale if trained
         | in the huge clusters?
        
         | pdntspa wrote:
         | From what I've read, DeepSeek is a "side project" at a Chinese
         | quant fund. They had the GPU capacity to spare.
        
           | browningstreet wrote:
           | I've read that too, and if true, and their strongest skill
           | and output resides elsewhere, that would point to other
           | interesting... impacts.
        
         | leetharris wrote:
         | CEO of Scale said Deepseek is lying and actually has a 50k GPU
         | cluster. He said they lied in the paper because technically
         | they aren't supposed to have them due to export laws.
         | 
         | I feel like this is very likely. They obvious did some great
         | breakthroughs, but I doubt they were able to train on so much
         | less hardware.
        
           | pdntspa wrote:
           | I would think the CEO of an American AI company has every
           | reason to neg and downplay foreign competition...
           | 
           | And since it's a businessperson they're going to make it
           | sound as cute and innocuous as possible
        
             | stale2002 wrote:
             | Or, more likely, there wasn't a magic innovation that
             | nobody else thought of, that reduced costs by orders of
             | magnitude.
             | 
             | When deciding between mostly like scenarios, it is more
             | likely that the company lied than they found some industry
             | changing magic innovation.
        
             | pjfin123 wrote:
             | It's hard to tell if they're telling the truth about the
             | number of GPUs they have. They open sourced the model and
             | the inference is much more efficient than the best American
             | models so it's not implausible that the training was also
             | much more efficient.
        
             | leetharris wrote:
             | If we're going to play that card, couldn't we also use the
             | "Chinese CEO has every reason to lie and say they did
             | something 100x more efficient than the Americans" card?
             | 
             | I'm not even saying they did it maliciously, but maybe just
             | to avoid scrutiny on GPUs they aren't technically supposed
             | to have? I'm thinking out loud, not accusing anyone of
             | anything.
        
               | mrbungie wrote:
               | Then the question becomes, who sold the GPUs to them?
               | They are supposedly scarse and every player in the field
               | is trying to get ahold as many as they can, before anyone
               | else in fact.
               | 
               | Something makes little sense in the accusations here.
        
               | leetharris wrote:
               | I think there's likely lots of potential culprits. If the
               | race is to make a machine god, states will pay countless
               | billions for an advantage. Money won't mean anything once
               | you enslave the machine god.
               | 
               | https://wccftech.com/nvidia-asks-super-micro-computer-
               | smci-t...
        
               | mrbungie wrote:
               | We will have to wait to get some info on that probe. I
               | know SMCI is not the nicest player and there is no doubt
               | GPUs are being smuggled, but that quantity (50k GPUs)
               | would be not that easy to smuggle and sell to a single
               | actor without raising suspicion.
        
           | latchkey wrote:
           | Thanks to SMCI that let them out...
           | 
           | https://wccftech.com/nvidia-asks-super-micro-computer-
           | smci-t...
           | 
           | Chinese guy in a warehouse full of SMCI servers bragging
           | about how he has them...
           | 
           | https://www.youtube.com/watch?v=27zlUSqpVn8
        
           | Leary wrote:
           | Alexandr Wang did not even say they lied in the paper.
           | 
           | Here's the interview:
           | https://www.youtube.com/watch?v=x9Ekl9Izd38. "My
           | understanding is that is that Deepseek has about 50000 a100s,
           | which they can't talk about obviously, because it is against
           | the export controls that the United States has put in place.
           | And I think it is true that, you know, I think they have more
           | chips than other people expect..."
           | 
           | Plus, how exactly did Deepseek lie. The model size, data size
           | are all known. Calculating the number of FLOPS is an exercise
           | in arithmetics, which is perhaps the secret Deepseek has
           | because it seemingly eludes people.
        
             | leetharris wrote:
             | > Plus, how exactly did Deepseek lie. The model size, data
             | size are all known. Calculating the number of FLOPS is an
             | exercise in arithmetics, which is perhaps the secret
             | Deepseek has because it seemingly eludes people.
             | 
             | Model parameter count and training set token count are
             | fixed. But other things such as epochs are not.
             | 
             | In the same amount of time, you could have 1 epoch or 100
             | epochs depending on how many GPUs you have.
             | 
             | Also, what if their claim on GPU count is accurate, but
             | they are using better GPUs they aren't supposed to have?
             | For example, they claim 1,000 GPUs for 1 month total. They
             | claim to have H800s, but what if they are using illegal
             | H100s/H200s, B100s, etc? The GPU count could be correct,
             | but their total compute is substantially higher.
             | 
             | It's clearly an incredible model, they absolutely cooked,
             | and I love it. No complaints here. But the likelihood that
             | there are some fudged numbers is not 0%. And I don't even
             | blame them, they are likely forced into this by US exports
             | laws and such.
        
               | kd913 wrote:
               | It should be trivially easy to reproduce the results no?
               | Just need to wait for one of the giant companies with
               | many times the GPUs to reproduce the results.
               | 
               | I don't expect a #180 AUM hedgefund to have as many GPUs
               | than meta, msft or Google.
        
               | sudosysgen wrote:
               | AUM isn't a good proxy for quantitative hedge fund
               | performance, many strategies are quite profitable and
               | don't scale with AUM. For what it's worth, they seemed to
               | have some excellent returns for many years for any
               | market, let alone the difficult Chinese markets.
        
               | sudosysgen wrote:
               | > In the same amount of time, you could have 1 epoch or
               | 100 epochs depending on how many GPUs you have.
               | 
               | This is just not true for RL and related algorithms,
               | having more GPU/agents encounters diminishing returns,
               | and is just not the equivalent to letting a single agent
               | go through more steps.
        
           | matthest wrote:
           | I've also read that Deepseek has released the research paper
           | and that anyone can replicate what they did.
           | 
           | I feel like if that were true, it would mean they're not
           | lying.
        
             | aprilthird2021 wrote:
             | You can't replicate it exactly because you don't know their
             | dataset or what exactly several of their proprietary
             | optimizations were
        
           | woadwarrior01 wrote:
           | CEO of a human based data labelling services company feels
           | threatened by a rival company that claims to have trained a
           | frontier class model with an almost entirely RL based
           | approach, with a small cold start dataset (a few thousand
           | samples). It's in the paper. If their approach is replicated
           | by other labs, Scale AI's business will drastically shrink or
           | even disappear.
           | 
           | Under such dire circumstances, lying isn't entirely out of
           | character for a corporate CEO.
        
             | leetharris wrote:
             | Could be true.
             | 
             | Deepseek obviously trained on OpenAI outputs, which were
             | originally RLHF'd. It may seem that we've got all the human
             | feedback necessary to move forward and now we can
             | infinitely distil + generate new synthetic data from higher
             | parameter models.
        
           | echelon wrote:
           | I haven't had time to follow this thread, but it looks like
           | some people are starting to experimentally replicate DeepSeek
           | on extremely limited H100 training:
           | 
           | > You can RL post-train your small LLM (on simple tasks) with
           | only 10 hours of H100s.
           | 
           | https://www.reddit.com/r/singularity/comments/1i99ebp/well_s.
           | ..
           | 
           | Forgive me if this is inaccurate. I'm rushing around too much
           | this afternoon to dive in.
        
           | weinzierl wrote:
           | Just to check my math: They claim something like 2.7 million
           | H800 hours which would be less than 4000 GPU units for one
           | month. In money something around 100 million USD give or take
           | a few tens of millions.
        
           | buyucu wrote:
           | Why would Deepseek lie? They are in China, American export
           | laws can't touch them.
        
             | echoangle wrote:
             | Making it obvious that they managed to circumvent sanctions
             | isn't going to help them. It will turn public sentiment in
             | the west even more against them and will motivate
             | politicians to make the enforcement stricter and prevent
             | GPU exports.
        
           | siltcakes wrote:
           | The CEO of Scale is one of the very last people I would trust
           | to provide this information.
        
         | Leary wrote:
         | or maybe the US economy will do even better because more people
         | will be able to use AI at a low cost.
         | 
         | OpenAI will be also be able to serve o3 at a lower cost if
         | Deepseek had some marginal breakthrough OpenAI did not already
         | think of.
        
           | 7thpower wrote:
           | I think this is the most productive mindset. All of the costs
           | thus far are sunk, the only move forward is to learn and
           | adjust.
           | 
           | This is a net win for nearly everyone.
           | 
           | The world needs more tokens and we are learning that we can
           | create higher quality tokens with fewer resources than
           | before.
           | 
           | Finger pointing is a very short term strategy.
        
         | rikafurude21 wrote:
         | Why do americans think china is like a hivemind controlled by
         | an omnisicient Xi, making strategic moves to undermine them? Is
         | it really that unlikely that a lab of genius engineers found a
         | way to improve efficiency 10x?
        
           | mritchie712 wrote:
           | think about how big the prize is, how many people are working
           | on it and how much has been invested (and targeted to be
           | invested, see stargate).
           | 
           | And they somehow yolo it for next to nothing?
           | 
           | yes, it seems unlikely they did it exactly they way they're
           | claiming they did. At the very least, they likely spent more
           | than they claim or used existing AI API's in way that's
           | against the terms.
        
           | logicchains wrote:
           | > Is it really that unlikely that a lab of genius engineers
           | found a way to improve efficiency 10x
           | 
           | They literally published all their methodology. It's nothing
           | groundbreaking, just western labs seem slow to adopt new
           | research. Mixture of experts, key-value cache compression,
           | multi-token prediction, 2/3 of these weren't invented by
           | DeepSeek. They did invent a new hardware-aware distributed
           | training approach for mixture-of-experts training that helped
           | a lot, but there's nothing super genius about it, western
           | labs just never even tried to adjust their model to fit the
           | hardware available.
        
             | blackeyeblitzar wrote:
             | But those approaches alone wouldn't yield the improvements
             | claimed. How did they train the foundational model upon
             | which they applied RL, distillations, etc? That part is
             | unclear and I don't think anything they've released
             | anything that explains the low cost.
             | 
             | It's also curious why some people are seeing responses
             | where it thinks it is an OpenAI model. I can't find the
             | post but someone had shared a link to X with that in one of
             | the other HN discussions.
        
             | rvnx wrote:
             | "nothing groundbreaking"
             | 
             | It's extremely cheap, efficient and kicks the ass of the
             | leader of the market, while being under sanctions with AI
             | hardware.
             | 
             | Most of all, can be downloaded for free, can be uncensored,
             | and usable offline.
             | 
             | China is really good at tech, it has beautiful landscapes,
             | etc. It has its own political system, but to be fair, in
             | some way it's all our future.
             | 
             | A bit of a dystopian future, like it was in 1984.
             | 
             | But the tech folks there are really really talented, it's
             | long time that China switched from producing for the
             | Western clients, to direct-sell to the Western clients.
        
               | gpm wrote:
               | The leaderboard leader [1] is still showing the
               | traditional AI leader, Google, winning. With
               | Gemini-2.0-Flash-Thinking-Exp-01-21 in the lead. No one
               | seems to know how many parameters that has, but random
               | guesses on the internet seem to be low to mid 10s of
               | billions, so fewer than DeepSeek-R1. Even if those
               | general guesses are wrong, they probably aren't that
               | wrong and at worst it's the same class of model as
               | DeepSeek-R1.
               | 
               | So yes, DeepSeek-R1 appears to be not even be best in
               | class, merely best open source. The only sense in which
               | it is "leading the market" appears to be the sense in
               | which "free stuff leads over proprietary stuff". Which is
               | true and all, but not a groundbreaking technical
               | achievement.
               | 
               | The DeepSeek-R1 distilled models on the other hand might
               | actually be leading at something... but again hard to say
               | it's groundbreaking when it's combining what we know we
               | can do (small models like llama) with what we know we can
               | do (thinking models).
               | 
               | [1] https://lmarena.ai/?leaderboard
        
               | dinosaurdynasty wrote:
               | The chatbot leaderboard seems to be very affected by
               | things other than capability, like "how nice is it to
               | talk to" and "how likely is it to refuse requests" and
               | "how fast does it respond" etc. Flash is literally one of
               | Google's faster models, definitely not their smartest.
               | 
               | Not that the leaderboard isn't useful, I think "is in the
               | top 10" says a lot more than the exact position in the
               | top 10.
        
               | gpm wrote:
               | I mean, sure, none of these models are being optimized
               | for being the top of the leader board. They aren't even
               | being optimized for the same things, so any comparison is
               | going to be somewhat questionable.
               | 
               | But the claim I'm refuting here is "It's extremely cheap,
               | efficient and kicks the ass of the leader of the market",
               | and I think the leaderboard being topped by a cheap
               | google model is pretty conclusive that that statement is
               | not true. Is competitive with? Sure. Kicks the ass of?
               | No.
        
               | whimsicalism wrote:
               | google absolutely games for lmsys benchmarks with
               | markdown styling. r1 is better than google flash
               | thinking, you are putting way too much faith in lmsys
        
               | patrickhogan1 wrote:
               | There is a wide disconnect between real world usage and
               | leaderboards. If gemini was so good why are so few using
               | them?
               | 
               | Having tested that model in many real world projects it
               | has not once been the best. And going farther it gives
               | atrocious nonsensical output.
        
               | whimsicalism wrote:
               | i'm sorry but gemini flash thinning is simply not as good
               | as r1. no way you've been playing with both
        
             | Scipio_Afri wrote:
             | That's what they claim at least in the paper but that
             | particular claim is not verifiable. The HAI-LLM framework
             | they reference in the paper is not open sourced and it
             | seems they have no plans to.
             | 
             | Additionally there are claims, such as those by Scale AI
             | CEO Alexandr Wang on CNBC 1/23/2025 time segment below,
             | that DeepSeek has 50,000 H100s that "they can't talk about"
             | due to economic sanctions (implying they likely got by
             | avoiding them somehow when restrictions were looser). His
             | assessment is that they will be more limited moving
             | forward.
             | 
             | https://youtu.be/x9Ekl9Izd38?t=178
        
               | byefruit wrote:
               | It's amazing how different the standards are here.
               | Deepseek's released their weights under a real open
               | source license and published a paper with their work
               | which now has independent reproductions.
               | 
               | OpenAI literally haven't said a thing about how O1 even
               | works.
        
               | marbli2 wrote:
               | They can be more open and yet still not open source
               | enough that claims of theirs being unverifiable are still
               | possible. Which is the case for their optimized HAI-LLM
               | framework.
        
               | byefruit wrote:
               | That's not what I'm saying, they may be hiding their true
               | compute.
               | 
               | I'm pointing out that nearly every thread covering
               | Deepseek R1 so far has been like this. Compare to the O1
               | system card thread:
               | https://news.ycombinator.com/item?id=42330666
               | 
               | Very different standards.
        
             | meltyness wrote:
             | The U.S. firms let everyone skeptical go the second they
             | had a marketable proof of concept, and replaced them with
             | smart, optimistic, uncritical marketing people who no
             | longer know how to push the cutting edge.
             | 
             | Maybe we don't need momentum right now and we can cut the
             | engines.
             | 
             | Oh, you know how to develop novel systems for training and
             | inference? Well, maybe you can find 4 people who also can
             | do that by breathing through the H.R. drinking straw, and
             | that's what you do now.
        
           | faitswulff wrote:
           | China is actually just one person (Xi) acting in perfect
           | unison and its purpose is not to benefit its own people, but
           | solely to undermine the West.
        
             | dr_dshiv wrote:
             | This explains so much. It's just malice, then? Or some
             | demonic force of evil? What does Occam's razor suggest?
             | 
             | Oh dear
        
               | layer8 wrote:
               | Always attribute to malice what can't be explained by
               | mere stupidity. ;)
        
               | buryat wrote:
               | payback for Opium Wars
        
               | pjc50 wrote:
               | You missed the really obvious sarcasm.
        
             | Zamicol wrote:
             | If China is undermining the West by lifting up humanity,
             | for free, while ProprietaryAI continues to use closed
             | source AI for censorship and control, then go team China.
             | 
             | There's something wrong with the West's ethos if we think
             | contributing significantly to the progress of humanity is
             | malicious. The West's sickness is our own fault; we should
             | take responsibility for our own disease, look critically to
             | understand its root, and take appropriate cures, even if
             | radical, to resolve our ailments.
        
               | Krasnol wrote:
               | > There's something wrong with the West's ethos if we
               | think contributing significantly to the progress of
               | humanity is malicious.
               | 
               | Who does this?
               | 
               | The criticism is aimed at the dictatorship and their
               | politics. Not their open source projects. Both things can
               | exist at once. It doesn't make China better in any way.
               | Same goes for their "radical cures" as you call it. I'm
               | sure Uyghurs in China would not give a damn about AI.
        
               | drysine wrote:
               | > I'm sure Uyghurs in China would not give a damn about
               | AI.
               | 
               | Which reminded me of "Whitey On the Moon" [0]
               | 
               | [0] https://www.youtube.com/watch?v=goh2x_G0ct4
        
             | colordrops wrote:
             | Can't tell if sarcasm. Some people are this simple minded.
        
               | rightbyte wrote:
               | Ye, but "acting in perfect unison" would be a superior
               | trait among people that care about these things which
               | gives it a way as sarcasm?
        
             | rambojohnson wrote:
             | that's the McCarthy era red scare nonsense still polluting
             | the minds of (mostly boomers / older gen-x) americans. it's
             | so juvenile and overly simplistic.
        
             | mackyspace wrote:
             | China is doing what it's always done and its culture _far_
             | predates  "the west".
        
           | bugglebeetle wrote:
           | I mean what's also incredible about all this cope is that
           | it's exactly the same David-v-Goliath story that's been
           | lionized in the tech scene for decades now about how the
           | truly hungry and brilliant can form startups to take out
           | incumbents and ride their way to billions. So, if that's not
           | true for DeepSeek, I guess all the people who did that in the
           | U.S. were also secretly state-sponsored operations to like
           | make better SAAS platforms or something?
        
           | diego_moita wrote:
           | SAY WHAT?
           | 
           | Do you want an Internet without conspiracy theories?
           | 
           | Where have you been living for the last decades?
           | 
           | /s
        
           | wumeow wrote:
           | Because that's the way China presents itself and that's the
           | way China boosters talk about China.
        
           | blackeyeblitzar wrote:
           | Well it is like a hive mind due to the degree of control.
           | Most Chinese companies are required by law to literally
           | uphold the country's goals - see translation of Chinese law,
           | which says generative AI must uphold their socialist values:
           | 
           | https://www.chinalawtranslate.com/en/generative-ai-interim/
           | 
           | In the case of TikTok, ByteDance and the government found
           | ways to force international workers in the US to signing
           | agreements that mirror local laws in mainland China:
           | 
           | https://dailycaller.com/2025/01/14/tiktok-forced-staff-
           | oaths...
           | 
           | I find that degree of control to be dystopian and horrifying
           | but I suppose it has helped their country focus and grow
           | instead of dealing with internal conflict.
        
         | robertclaus wrote:
         | Doesn't this just mean throwing a gazillion GPUs at the new
         | architecture and defining a new SOTA?
        
         | eightysixfour wrote:
         | I don't believe that the model was trained on so few GPUs,
         | personally, but it also doesn't matter IMO. I don't think SOTA
         | models are moats, they seem to be more like guiding lights that
         | others can quickly follow. The volume of research on different
         | approaches says we're still in the early days, and it is highly
         | likely we continue to get surprises with models and systems
         | that make sudden, giant leaps.
         | 
         | Many "haters" seem to be predicting that there will be model
         | collapse as we run out of data that isn't "slop," but I think
         | they've got it backwards. We're in the flywheel phase now, each
         | SOTA model makes future models better, and others catch up
         | faster.
        
         | blackeyeblitzar wrote:
         | It's not just the economy that is vulnerable, but global
         | geopolitics. It's definitely worrying to see this type of
         | technology in the hands of an authoritarian dictatorship,
         | especially considering the evidence of censorship. See this
         | article for a collected set of prompts and responses from
         | DeepSeek highlighting the propaganda:
         | 
         | https://medium.com/the-generator/deepseek-hidden-china-polit...
         | 
         | But also the claimed cost is suspicious. I know people have
         | seen DeepSeek claim in some responses that it is one of the
         | OpenAI models, so I wonder if they somehow trained using the
         | outputs of other models, if that's even possible (is there such
         | a technique?). Maybe that's how the claimed cost is so low that
         | it doesn't make mathematical sense?
        
           | rightbyte wrote:
           | I am certainly reliefed there is no super power lock in for
           | this stuff.
           | 
           | In theory I could run this one at home too without giving my
           | data or money to Sam Altman.
        
           | buyucu wrote:
           | have you tried asking chatgpt something even slightly
           | controversial? chatgpt censors much more than deepseek does.
           | 
           | also deepseek is open-weights. there is nothing preventing
           | you from doing a finetune that removes the censorship. they
           | did that with llama2 back in the day.
        
             | blackeyeblitzar wrote:
             | > chatgpt censors much more than deepseek does
             | 
             | This is an outrageous claim with no evidence, as if there
             | was any equivalence between government enforced propaganda
             | and anything else. Look at the system prompts for DeepSeek
             | and it's even more clear.
             | 
             | Also: fine tuning is not relevant when what is deployed at
             | scale brainwashes the masses through false and misleading
             | responses.
        
           | aprilthird2021 wrote:
           | > It's definitely worrying to see this type of technology in
           | the hands of an authoritarian dictatorship
           | 
           | What do you think they will do with the AI that worries you?
           | They already had access to Llama, and they could pay for
           | access to the closed source AIs. It really wouldn't be that
           | hard to pay for and use what's commercially available as
           | well, even if there is embargo or whatever, for digital goods
           | and services that can easily be bypassed
        
         | ak_111 wrote:
         | Would you say they were more vulnerable if the PRC kept it
         | secret so as not to disclose their edge in AI while continuing
         | to build on it?
        
         | tomjen3 wrote:
         | We will know soon enough if this replicates since Huggingface
         | is working on replicating it.
         | 
         | To know that this would work requires insanely deep technical
         | knowledge about state of the art computing, and the top
         | leadership of the PRC does not have that.
        
           | handzhiev wrote:
           | Researchers from TikTok claim they already replicated it
           | 
           | https://x.com/sivil_taram/status/1883184784492666947?t=NzFZj.
           | ..
        
         | ecocentrik wrote:
         | I don't think we were wrong to look at this as a commodity
         | problem and ask how many widgets we need. Most people will
         | still get their access to this technology through cloud
         | services and nothing in this paper changes the calculations for
         | inference compute demand. I still expect inference compute
         | demand to be massive and distilled models aren't going to cut
         | it for most agentic use cases.
        
         | pfisherman wrote:
         | > The US Economy is pretty vulnerable here. If it turns out
         | that you, in fact, don't need a gazillion GPUs to build SOTA
         | models it destroys a lot of perceived value.
         | 
         | I do not quite follow. GPU compute is mostly spent in
         | inference, as training is a one time cost. And these chain of
         | thought style models work by scaling up inference time compute,
         | no?
         | 
         | So proliferation of these types of models would portend in
         | increase in demand for GPUs?
        
         | cedws wrote:
         | Good. This gigantic hype cycle needs a reality check. And if it
         | turns out Deepseek is hiding GPUs, good for them for doing what
         | they need to do to get ahead.
        
         | buyucu wrote:
         | Seeing what china is doing to the car market, I give it 5 years
         | for China to do to the AI/GPU market to do the same.
         | 
         | This will be good. Nvidia/OpenAI monopoly is bad for everyone.
         | More competition will be welcome.
        
           | mrbungie wrote:
           | That is not going to happen without currently embargo'ed
           | litography tech. They'd be already making more powerful GPUs
           | if they could right now.
        
             | buyucu wrote:
             | they seem to be doing fine so far. every day we wake up to
             | more success stories from china's AI/semiconductory
             | industry.
        
         | flaque wrote:
         | This only makes sense if you think scaling laws won't hold.
         | 
         | If someone gets something to work with 1k h100s that should
         | have taken 100k h100s, that means the group with the 100k is
         | about to have a much, much better model.
        
         | aprilthird2021 wrote:
         | > If it turns out that you, in fact, don't need a gazillion
         | GPUs to build SOTA models it destroys a lot of perceived value.
         | 
         | Correct me if I'm wrong, but couldn't you take the optimization
         | and tricks for training, inference, etc. from this model and
         | apply to the Big Corps' huge AI data centers and get an even
         | better model?
         | 
         | I'll preface this by saying, better and better models may not
         | actually unlock the economic value they are hoping for. It
         | might be a thing where the last 10% takes 90% of the effort so
         | to speak
        
       | GaggiX wrote:
       | I wonder if the decision to make o3-mini available for free user
       | in the near (hopefully) future is a response to this really good,
       | cheap and open reasoning model.
        
         | swyx wrote:
         | almost certainly (see chart)
         | https://www.latent.space/p/reasoning-price-war (disclaimer i
         | made it)
        
           | coder543 wrote:
           | I understand you were trying to make "up and to the right" =
           | "best", but the inverted x-axis really confused me at first.
           | Not a huge fan.
           | 
           | Also, I wonder how you're calculating costs, because while a
           | 3:1 ratio kind of sort of makes sense for traditional LLMs...
           | it doesn't really work for "reasoning" models that implicitly
           | use several hundred to several thousand additional output
           | tokens for their reasoning step. It's almost like a "fixed"
           | overhead, regardless of the input or output size around that
           | reasoning step. (Fixed is in quotes, because some reasoning
           | chains are longer than others.)
           | 
           | I would also argue that token-heavy use cases are dominated
           | by large input/output ratios of like 100:1 or 1000:1 tokens.
           | Token-light use cases are your typical chatbot where the user
           | and model are exchanging roughly equal numbers of tokens...
           | and probably not that many per message.
           | 
           | It's hard to come up with an optimal formula... one would
           | almost need to offer a dynamic chart where the user can enter
           | their own ratio of input:output, and choose a number for the
           | reasoning token overhead. (Or, select from several predefined
           | options like "chatbot", "summarization", "coding assistant",
           | where those would pre-select some reasonable defaults.)
           | 
           | Anyways, an interesting chart nonetheless.
        
             | swyx wrote:
             | i mean the sheet is public https://docs.google.com/spreadsh
             | eets/d/1x9bQVlm7YJ33HVb3AGb9... go fiddle with it yourself
             | but you'll soon see most models hve approx the same
             | input:output token ratio cost (roughly 4) and changing the
             | input:output ratio assumption doesnt affect in the
             | slightest what the overall macro chart trends say because
             | i'm plotting over several OoMs here and your criticisms
             | have the impact of <1 OoM (input:output token ratio cost of
             | ~4).
             | 
             | actually the 100:1 ratio starts to trend back toward parity
             | now because of the reasoning tokens, so the truth is
             | somewhere between 3:1 and 100:1.
        
       | mmaunder wrote:
       | Over 100 authors on that paper. Cred stuffing ftw.
        
         | swyx wrote:
         | oh honey. have you read the gemini paper.
        
           | anothermathbozo wrote:
           | So tired of seeing this condescending tone online
        
         | verdverm wrote:
         | there are better ways to view this:
         | https://news.ycombinator.com/item?id=42824223
        
         | janalsncm wrote:
         | Physics papers often have hundreds.
        
       | swyx wrote:
       | we've been tracking the deepseek threads extensively in LS.
       | related reads:
       | 
       | - i consider the deepseek v3 paper required preread
       | https://github.com/deepseek-ai/DeepSeek-V3
       | 
       | - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo
       | https://aider.chat/2025/01/24/r1-sonnet.html
       | 
       | - independent repros: 1) https://hkust-nlp.notion.site/simplerl-
       | reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-
       | reprod... 3)
       | https://x.com/ClementDelangue/status/1883154611348910181
       | 
       | - R1 distillations are going to hit us every few days - because
       | it's ridiculously easy (<$400, <48hrs) to improve any base model
       | with these chains of thought eg with Sky-T1 recipe (writeup
       | https://buttondown.com/ainews/archive/ainews-bespoke-stratos... ,
       | 23min interview w team
       | https://www.youtube.com/watch?v=jrf76uNs77k)
       | 
       | i probably have more resources but dont want to spam - seek out
       | the latent space discord if you want the full stream i pulled
       | these notes from
        
         | sitkack wrote:
         | I am extremely interested in your spam. Will you post it to
         | https://www.latent.space/ ?
        
           | swyx wrote:
           | idk haha most of it is just twitter bookmarks - i will if i
           | get to interview the deepseek team at some point (someone
           | help put us in touch pls! swyx at ai.engineer )
        
         | sitkack wrote:
         | Hugging Face is reproducing R1 in public.
         | 
         | https://x.com/_lewtun/status/1883142636820676965
         | 
         | https://github.com/huggingface/open-r1
         | 
         | Hugging Face Journal Club - DeepSeek R1
         | https://www.youtube.com/watch?v=1xDVbu-WaFo
        
           | swyx wrote:
           | oh also we are doing a live Deepseek v3/r1 paper club next
           | wed: signups here https://lu.ma/ls if you wanna discuss
           | stuff!
        
       | logifail wrote:
       | Q: Is there a thread about DeepSeek's (apparent) progress with
       | lots of points and lots of quality comments?
       | 
       | (Bonus Q: If not, why not?)
        
       | bad_haircut72 wrote:
       | Even if you think this particular team cheated, the idea that
       | _nobody_ will find ways of making training more efficient seems
       | silly - these huge datacenter investments for purely AI will IMHO
       | seem very short sighted in 10 years
        
         | neverthe_less wrote:
         | Isn't it possible with more efficiency, we still want them for
         | advanced AI capabilities we could unlock in the future?
        
           | thfuran wrote:
           | Operating costs are usually a pretty significant factor in
           | total costs for a data center. Unless power efficiency stops
           | improving much and/or demand so far outstrips supply that
           | they can't be replaced, a bunch of 10 year old GPUs probably
           | aren't going to be worth running regardless.
        
         | foobiekr wrote:
         | More like three years. Even in the best case the retained value
         | curve of GPUs is absolutely terrible. Most of these huge
         | investments in GPUs are going to be massive losses.
        
           | tobias3 wrote:
           | Seems bad for those GPU backed loans
        
           | newAccount2025 wrote:
           | Do we have any idea how long a cloud provider needs to rent
           | them out for to make back their investment? I'd be surprised
           | if it was more than a year, but that is just a wild guess.
        
           | kandesbunzler wrote:
           | >retained value curve of GPUs is absolutely terrible
           | 
           | source?
        
         | dsign wrote:
         | >> for purely AI
         | 
         | There is a big balloon full of AI hype going up right now, and
         | regrettably it may need those data-centers. But I'm hoping that
         | if the worst (the best) comes to happen, we will find worthy
         | things to do with all of that depreciated compute. Drug
         | discovery comes to mind.
        
       | vlaaad wrote:
       | Reddit's /r/chatgpt subreddit is currently heavily brigaded by
       | bots/shills praising r1, I'd be very suspicious of any claims
       | about it.
        
         | butterlettuce wrote:
         | Source?
        
         | Crye wrote:
         | You can try it yourself, it's refreshingly good.
        
           | sdesol wrote:
           | Agreed. I am no fan of the CCP but I have no issue with using
           | DeepSeek since I only need to use it for coding which it does
           | quite well. I still believe Sonnet is better. DeepSeek also
           | struggles when the context window gets big. This might be
           | hardware though.
           | 
           | Having said that, DeepSeek is 10 times cheaper than Sonnet
           | and better than GPT-4o for my use cases. Models are a
           | commodity product and it is easy enough to add a layer above
           | them to only use them for technical questions.
           | 
           | If my usage can help v4, I am all for it as I know it is
           | going to help everyone and not just the CCP. Should they stop
           | publishing the weights and models, v3 can still take you
           | quite far.
        
             | spaceman_2020 wrote:
             | Curious why you have to qualify this with a "no fan of the
             | CCP" prefix. From the outset, this is just a private
             | organization and its links to CCP aren't any different
             | than, say, Foxconn's or DJI's or any of the countless
             | Chinese manufacturers and businesses
             | 
             | You don't invoke "I'm no fan of the CCP" before opening
             | TikTok or buying a DJI drone or a BYD car. Then why this,
             | because I've seen the same line repeated everywhere
        
         | forrestthewoods wrote:
         | The amount of astroturfing around R1 is absolutely wild to see.
         | Full scale propaganda war.
        
           | rightbyte wrote:
           | I would argue there is too little hype given the downloadable
           | models for Deep Seek. There should be alot of hype around
           | this organically.
           | 
           | If anything, the other half good fully closed non ChatGPT
           | models are astroturfing.
           | 
           | I made a post in december 2023 whining about the non hype for
           | Deep Seek.
           | 
           | https://news.ycombinator.com/item?id=38505986
        
             | forrestthewoods wrote:
             | Possible for that to also be true!
             | 
             | There's a lot of astroturfing from a lot of different
             | parties for a few different reasons. Which is all very
             | interesting.
        
               | Philpax wrote:
               | How do you know it's astroturfing and not legitimate hype
               | about an impressive and open technical achievement?
        
               | rightbyte wrote:
               | Ye I mean in practice it is impossible to verify. You can
               | kind of smell it though and I smell nothing here,
               | eventhough some of 100 listed authors should be HN users
               | and write in this thread.
               | 
               | Some obvious astroturf posts on HN seem to be on the
               | template "Watch we did boring coorparate SaaS thing X
               | noone cares about!" and then a disappropiate amount of
               | comments and upvotes and 'this is a great idea', 'I used
               | it, it is good' or congratz posts, compared to the usual
               | cynical computer nerd everything sucks especially some
               | minute detail about the CSS of your website mindset you'd
               | expect.
        
           | glass-z13 wrote:
           | Ironic
        
             | forrestthewoods wrote:
             | That word does not mean what you think it means.
        
           | spaceman_2020 wrote:
           | The literal creator of Netscape Navigator is going ga-ga over
           | it on Twitter and HN thinks its all botted
           | 
           | This is not a serious place
        
         | mtkd wrote:
         | The counternarrative is that it is a very accomplished piece of
         | work that most in the sector were not expecting -- it's open
         | source with API available at fraction of comparable service
         | cost
         | 
         | It has upended a lot of theory around how much compute is
         | likely needed over next couple of years, how much profit
         | potential the AI model vendors have in nearterm and how big an
         | impact export controls are having on China
         | 
         | V3 took top slot on HF trending models for first part of Jan
         | ... r1 has 4 of the top 5 slots tonight
         | 
         | Almost every commentator is talking about nothing else
        
         | buyucu wrote:
         | I'm running the 7b distillation on my laptop this very moment.
         | It's an insanely good model. You don't need reddit to judge how
         | good a model is.
        
         | mediaman wrote:
         | You can just use it and see for yourself. It's quite good.
         | 
         | I do believe they were honest in the paper, but the $5.5m
         | training cost (for v3) is defined in a limited way: only the
         | GPU cost at $2/hr for the one training run they did that
         | resulted in the final V3 model. Headcount, overhead,
         | experimentation, and R&D trial costs are not included. The
         | paper had something like 150 people on it, so obviously total
         | costs are quite a bit higher than the limited scope cost they
         | disclosed, and also they didn't disclose R1 costs.
         | 
         | Still, though, the model is quite good, there are quite a few
         | independent benchmarks showing it's pretty competent, and it
         | definitely passes the smell test in actual use (unlike many of
         | Microsoft's models which seem to be gamed on benchmarks).
        
         | nowittyusername wrote:
         | Its pretty nutty indeed. The model still might be good, but the
         | botting is wild. On that note, one of my favorite benchmarks to
         | watch is simple bench and R! doesn't perform as well on that
         | benchmark as all the other public benchmarks, so it might be
         | telling of something.
        
       | Imanari wrote:
       | Question about the rule-based rewards (correctness and format)
       | mentioned in the paper: Does the raw base model just expected
       | "stumble upon" a correct answer /correct format to get a reward
       | and start the learning process? Are there any more details about
       | the reward modelling?
        
         | leobg wrote:
         | Good question.
         | 
         | When BF Skinner used to train his pigeons, he'd initially
         | reinforce any tiny movement that at least went in the right
         | direction. For the exact reasons you mentioned.
         | 
         | For example, instead of waiting for the pigeon to peck the
         | lever directly (which it might not do for many hours), he'd
         | give reinforcement if the pigeon so much as turned its head
         | towards the lever. Over time, he'd raise the bar. Until,
         | eventually, only clear lever pecks would receive reinforcement.
         | 
         | I don't know if they're doing something like that here. But it
         | would be smart.
        
           | fspeech wrote:
           | Since intermediate steps of reasoning are hard to verify they
           | only award final results. Yet that produces enough signal to
           | produce more productive reasoning over time. In a way when
           | pigeons are virtual one can afford to have a lot more of
           | them.
        
           | whimsicalism wrote:
           | they're not doing anything like that and you are actually
           | describing the failed research direction a lot of the
           | frontier labs (esp Google) were doing
        
         | whimsicalism wrote:
         | yes, stumble on a correct answer and also pushing down
         | incorrect answer probability in the meantime. their base model
         | is pretty good
        
       | freediver wrote:
       | Genuinly curious, what is everyone using reasoning models for?
       | (R1/o1/o3)
        
         | pieix wrote:
         | Regular coding questions mostly. For me o1 generally gives
         | better code and understands the prompt more completely (haven't
         | started using r1 or o3 regularly enough to opine).
        
           | whimsicalism wrote:
           | o3 isn't available
        
         | lexandstuff wrote:
         | We've been seeing success using it for LLM-as-a-judge tasks.
         | 
         | We set up an evaluation criteria and used o1 to evaluate the
         | quality of the prod model, where the outputs are subjective,
         | like creative writing or explaining code.
         | 
         | It's also useful for developing really good few-shot examples.
         | We'll get o1 to generate multiple examples in different styles,
         | then we'll have humans go through and pick the ones they like
         | best, which we use as few-shot examples for the cheaper, faster
         | prod model.
         | 
         | Finally, for some study I'm doing, I'll use it to grade my
         | assignments before I hand them in. If I get a 7/10 from o1,
         | I'll ask it to suggest the minimal changes I could make to take
         | it to 10/10. Then, I'll make the changes and get it to regrade
         | the paper.
        
         | iagooar wrote:
         | Everything, basically. From great cooking recipes to figuring
         | out + designing a new business, and everything in between.
        
         | whimsicalism wrote:
         | everything except writing. i was sparing with my o1 usage
         | because its priced so high but now i literally am using r1 for
         | everything
        
       | verdverm wrote:
       | Over 100 authors on arxiv and published under the team name,
       | that's how you recognize everyone and build comradery. I bet
       | morale is high over there
        
         | wumeow wrote:
         | It's credential stuffing.
        
           | tokioyoyo wrote:
           | Come on man, let them have their well deserved win as a team.
        
             | wumeow wrote:
             | Yea, I'm they're devastated by my comment
        
         | mi_lk wrote:
         | Same thing happened to Google Gemini paper (1000+ authors) and
         | it was described as big co promo culture (everyone wants
         | credits). Interesting how narratives shift
         | 
         | https://arxiv.org/abs/2403.05530
        
       | blackbear_ wrote:
       | The poor readability bit is quite interesting to me. While the
       | model does develop some kind of reasoning abilities, we have no
       | idea what the model is doing to convince itself about the answer.
       | These could be signs of non-verbal reasoning, like visualizing
       | things and such. Who knows if the model hasn't invented genuinely
       | novel things when solving the hardest questions? And could the
       | model even come up with qualitatively different and "non human"
       | reasoning processes? What would that even look like?
        
       | cjbgkagh wrote:
       | I've always been leery about outrageous GPU investments, at some
       | point I'll dig through and find my prior comments where I've said
       | as much to that effect.
       | 
       | The CEOs, upper management, and governments derive their
       | importance on how much money they can spend - AI gave them the
       | opportunity for them to confidently say that if you give me $X I
       | can deliver Y and they turn around and give that money to NVidia.
       | The problem was reduced to a simple function of raising money and
       | spending that money making them the most importance central
       | figure. ML researchers are very much secondary to securing
       | funding. Since these people compete with each other in importance
       | they strived for larger dollar figures - a modern dick waving
       | competition. Those of us who lobbied for efficiency were
       | sidelined as we were a threat. It was seen as potentially making
       | the CEO look bad and encroaching in on their importance. If the
       | task can be done for cheap by smart people then that severely
       | undermines the CEOs value proposition.
       | 
       | With the general financialization of the economy the wealth
       | effect of the increase in the cost of goods increases wealth by a
       | greater amount than the increase in cost of goods - so that if
       | the cost of housing goes up more people can afford them. This
       | financialization is a one way ratchet. It appears that the US
       | economy was looking forward to blowing another bubble and now
       | that bubble has been popped in its infancy. I think the slowness
       | of the popping of this bubble underscores how little the major
       | players know about what has just happened - I could be wrong
       | about that but I don't know how yet.
       | 
       | Edit: "[big companies] would much rather spend huge amounts of
       | money on chips than hire a competent researcher who might tell
       | them that they didn't really need to waste so much money."
       | (https://news.ycombinator.com/item?id=39483092 11 months ago)
        
         | breadwinner wrote:
         | Latest GPUs and efficiency are not mutually exclusive, right?
         | If you combine them both presumably you can build even more
         | powerful models.
        
           | kelseyfrog wrote:
           | That's Jevons Paradox in a nutshell
        
           | cjbgkagh wrote:
           | Of course optimizing for the best models would result in a
           | mix of GPU spend and ML researchers experimenting with
           | efficiency. And it may not make any sense to spend money on
           | researching efficiency since, as has happened, these are
           | often shared anyway for free.
           | 
           | What I was cautioning people was be that you might not want
           | to spend 500B on NVidia hardware only to find out rather
           | quickly that you didn't need to. You'd have all this CapEx
           | that you now have to try to extract from customers from what
           | has essentially been commoditized. That's a whole lot of
           | money to lose very quickly. Plus there is a zero sum power
           | dynamic at play between the CEO and ML researchers.
        
           | fspeech wrote:
           | Not necessarily if you are pushing against a data wall. One
           | could ask: after adjusting for DS efficiency gains how much
           | more compute has OpenAI spent? Is their model correspondingly
           | better? Or even DS could easily afford more than $6 million
           | in compute but why didn't they just push the scaling?
        
             | whimsicalism wrote:
             | right except that r1 is demoing the path of approach for
             | moving beyond the data wall
        
               | breadwinner wrote:
               | Can you clarify? How are they able to move beyond the
               | data wall?
        
         | solidasparagus wrote:
         | I think you are underestimating the fear of being beaten (for
         | many people making these decisions, "again") by a competitor
         | that does "dumb scaling".
        
           | sudosysgen wrote:
           | But dumb scaling clearly only gives logarithmic rewards at
           | best from every scaling law we ever saw.
        
         | dboreham wrote:
         | Agree. The "need to build new buildings, new power plants, buy
         | huge numbers of today's chips from one vendor" never made any
         | sense considering we don't know what would be done in those
         | buildings in 5 years when they're ready.
        
           | drysine wrote:
           | >in 5 years
           | 
           | Or much much quicker [0]
           | 
           | [0] https://timelines.issarice.com/wiki/Timeline_of_xAI
        
           | spacemanspiff01 wrote:
           | The other side of this is that if this is over investment
           | (likely)
           | 
           | Then in 5 years time resources will be much cheaper and spur
           | alot of exploration developments. There are many people with
           | many ideas, and a lot of them are just lacking compute to
           | attempt them.
           | 
           | My back of mind thought is that worst case it will be like
           | how the US overbuilt fiber in the 90s, which led the way for
           | cloud, network and such in 2000s.
        
           | totallynothoney wrote:
           | The eBay resells will be glorious.
        
         | -1 wrote:
         | I agree. I think there's a good chance that politicians & CEOs
         | pushing for 100s of billions spent on AI infrastructure are
         | going to look foolish.
        
         | cma wrote:
         | The results never fell off significantly with more training.
         | Same model with longer training time on those bigger clusters
         | should outdo it significantly. And they can expand the MoE
         | model sizes without the same memory and bandwidth constraints.
         | 
         | Still very surprising with so much less compute they were still
         | able to do so well in the model architecture/hyperparameter
         | exploration phase compared with Meta.
        
         | mlsu wrote:
         | Such a good comment.
         | 
         | Remember when Sam Altman was talking about raising 5 trillion
         | dollars for hardware?
         | 
         | insanity, total insanity.
        
         | dwallin wrote:
         | The cost of having excess compute is less than the cost of not
         | having enough compute to be competitive. Because of demand, if
         | you realize you your current compute is insufficient there is a
         | long turnaround to building up your infrastructure, at which
         | point you are falling behind. All the major players are
         | simultaneously working on increasing capabilities and reducing
         | inference cost. What they aren't optimizing is their total
         | investments in AI. The cost of over-investment is just a drag
         | on overall efficiency, but the cost of under-investment is
         | existential.
        
         | thethethethe wrote:
         | IMO the you cannot fail by investing in compute. If it turns
         | out you only need 1/1000th of the compute to train and or run
         | your models, great! Now you can spend that compute on inference
         | that solves actual problems humans have.
         | 
         | o3 $4k compute spend per task made it pretty clear that once we
         | reach AGI inference is going to be the majority of spend. We'll
         | spend compute getting AI to cure cancer or improve itself
         | rather than just training at chatbot that helps students cheat
         | on their exams. The more compute you have, the more problems
         | you can solve faster, the bigger your advantage, especially
         | if/when recursive self improvement kicks off, efficiency
         | improvements only widen this gap
        
       | dtquad wrote:
       | Is there any guide out there on how to use the reasoner in
       | standalone mode and maybe pair it with other models?
        
       | msp26 wrote:
       | How can openai justify their $200/mo subscriptions if a model
       | like this exists at an incredibly low price point? Operator?
       | 
       | I've been impressed in my brief personal testing and the model
       | ranks very highly across most benchmarks (when controlled for
       | style it's tied number one on lmarena).
       | 
       | It's also hilarious that openai explicitly prevented users from
       | seeing the CoT tokens on the o1 model (which you still pay for
       | btw) to avoid a situation where someone trained on that output.
       | Turns out it made no difference lmao.
        
         | tokioyoyo wrote:
         | From my casual read, right now everyone is on reputation
         | tarnishing tirade, like spamming "Chinese stealing data!
         | Definitely lying about everything! API can't be this cheap!".
         | If that doesn't go through well, I'm assuming lobbyism will
         | start for import controls, which is very stupid.
         | 
         | I have no idea how they can recover from it, if DeepSeek's
         | product is what they're advertising.
        
           | itsoktocry wrote:
           | So you're saying that this is the end of OpenAI?
           | 
           | Somehow I doubt it.
        
             | tokioyoyo wrote:
             | Hah I agree, they will find a way. In the end, the big
             | winners will be the ones who find use cases other than a
             | general chatbot. Or AGI, I guess.
        
         | spaceman_2020 wrote:
         | I find that this model feels more human, purely because of the
         | reasoning style (first person). In its reasoning text, it comes
         | across as a neurotic, eager to please smart "person", which is
         | hard not to anthropomorphise
        
         | whimsicalism wrote:
         | openai has better models in the bank so short term they will
         | release o3-derived models
        
       | rightbyte wrote:
       | There seems to be a print out of "reasoning". Is that some new
       | breaktheough thing? Really impressive.
       | 
       | E.g. I tried to make it guess my daughter's name and I could only
       | answer yes or no and the first 5 questions where very convincing
       | but then it lost track and started to randomly guess names one by
       | one.
       | 
       | edit: Nagging it to narrow it down and give a language group hint
       | made it solve it. Ye, well, it can do Akinator.
        
       | buryat wrote:
       | Interacting with this model is just supplying your data over to
       | an adversary with unknown intents. Using an open source model is
       | subjecting your thought process to be programmed with carefully
       | curated data and a systems prompt of unknown direction and
       | intent.
        
         | inertiatic wrote:
         | >Interacting with this model is just supplying your data over
         | to an adversary with unknown intents
         | 
         | Skynet?
        
       | browningstreet wrote:
       | I wonder if sama is working this weekend
        
       | yohbho wrote:
       | "Reasoning" will be disproven for this again within a few days I
       | guess.
       | 
       | Context: o1 does not reason, it pattern matches. If you rename
       | variables, suddenly it fails to solve the request.
        
         | marviel wrote:
         | reasoning is pattern matching at a certain level of
         | abstraction.
        
         | jakeinspace wrote:
         | Rename to equally reasonable variable names, or to
         | intentionally misleading or meaningless ones? Good naming is
         | one of the best ways to make reading unfamiliar code easier for
         | people, don't see why actual AGI wouldn't also get tripped up
         | there.
        
         | HarHarVeryFunny wrote:
         | Perhaps, but over enough data pattern matching can becomes
         | generalization ...
         | 
         | One of the interesting DeepSeek-R results is using a 1st
         | generation (RL-trained) reasoning model to generate synthetic
         | data (reasoning traces) to train a subsequent one, or even
         | "distill" into a smaller model (by fine tuning the smaller
         | model on this reasoning data).
         | 
         | Maybe "Data is all you need" (well, up to a point) ?
        
         | nullc wrote:
         | The 'pattern matching' happens at complex layer's of
         | abstraction, constructed out of combinations of pattern
         | matching at prior layers in the network.
         | 
         | These models can and do work okay with variable names that have
         | never occurred in the training data. Though sure, choice of
         | variable names can have an impact on the performance of the
         | model.
         | 
         | That's also true for humans, go fill a codebase with misleading
         | variable names and watch human programmers flail. Of course,
         | the LLM's failure modes are sometimes pretty inhuman, -- it's
         | not a human after all.
        
       | buyucu wrote:
       | I'm impressed by not only how good deepseek r1 is, but also how
       | good the smaller distillations are. qwen-based 7b distillation of
       | deepseek r1 is a great model too.
       | 
       | the 32b distillation just became the default model for my home
       | server.
        
         | OCHackr wrote:
         | How much VRAM is needed for the 32B distillation?
        
           | jadbox wrote:
           | Depends on compression, I think 24gb can hold a 32B at around
           | 3b-4b compression.
        
           | brandall10 wrote:
           | Depends on the quant used and the context size. On a 24gb
           | card you should be able to load about a 5 bit if you keep the
           | context small.
           | 
           | In general, if you're using 8bit which is virtually lossless,
           | any dense model will require roughly the same amount as the
           | number of params w/ a small context, and a bit more as you
           | increase context.
        
           | buyucu wrote:
           | I had no problems running the 32b at q4 quantization with
           | 24GB of ram.
        
         | magicalhippo wrote:
         | I just tries the distilled 8b Llama variant, and it had very
         | poor prompt adherence.
         | 
         | It also reasoned its way to an incorrect answer, to a question
         | plain Llama 3.1 8b got fairly correct.
         | 
         | So far not impressed, but will play with the qwen ones
         | tomorrow.
        
         | ThouYS wrote:
         | tried the 7b, it switched to chinese mid-response
        
           | popinman322 wrote:
           | Assuming you're doing local inference, have you tried setting
           | a token filter on the model?
        
         | brookst wrote:
         | Great as long as you're not interested in Tiananmen Square or
         | the Uighurs.
        
           | whimsicalism wrote:
           | american models have their own bugbears like around evolution
           | and intellectual property
        
             | miohtama wrote:
             | For sensitive topics, it is good that we canknow cross ask
             | Grok, DeepSeek and ChatGPT to avoid any kind of biases or
             | no-reply answers.
        
       | huqedato wrote:
       | ...and China is two years behind in AI. Right ?
        
         | mrbungie wrote:
         | And if they are up-to-date is because they're cheating. The
         | copium itt is astounding.
        
           | BriggyDwiggs42 wrote:
           | What's the difference between what they do and what other ai
           | firms do to openai in the us? What is cheating in a business
           | context?
        
             | fragmede wrote:
             | domestically, trade secrets are a thing and you can be sued
             | for corporate espionage. but in an international business
             | context with high geopolitical ramifications? the Soviets
             | copied American tech even when it was inappropriate, to
             | their detriment.
        
             | mrbungie wrote:
             | Smuggling GPUs and using OpenAI outputs.
             | 
             | PS: I'm not criticizing them for it nor do I really care if
             | they cheat as long prices go down. I'm just observing and
             | pointing out what other posters are saying. For me if China
             | cheating means the GenAI bubble pops, I'm all for it. Plus
             | no actor is really clean in this game, starting from OAI
             | practically stealing all human content for building their
             | models.
        
         | usaar333 wrote:
         | They were 6 months behind US frontier until deepseek r1.
         | 
         | Now maybe 4? It's hard to say.
        
           | spaceman_2020 wrote:
           | Outside of Veo2 - which I can't access anyway - they're
           | definitely ahead in AI video gen
        
             | whimsicalism wrote:
             | the big american labs don't care about ai video gen
        
       | jedharris wrote:
       | See also independent RL based reasoning results, fully open
       | source: https://hkust-nlp.notion.site/simplerl-reason
       | 
       | Very small training set!
       | 
       | "we replicate the DeepSeek-R1-Zero and DeepSeek-R1 training on
       | small models with limited data. We show that long Chain-of-
       | Thought (CoT) and self-reflection can emerge on a 7B model with
       | only 8K MATH examples, and we achieve surprisingly strong results
       | on complex mathematical reasoning. Importantly, we fully open-
       | source our training code and details to the community to inspire
       | more works on reasoning."
        
       | anothermathbozo wrote:
       | I don't think this entirely invalidates massive GPU spend just
       | yet:
       | 
       | " Therefore, we can draw two conclusions: First, distilling more
       | powerful models into smaller ones yields excellent results,
       | whereas smaller models relying on the large-scale RL mentioned in
       | this paper require enormous computational power and may not even
       | achieve the performance of distillation. Second, while
       | distillation strategies are both economical and effective,
       | advancing beyond the boundaries of intelligence may still require
       | more powerful base models and larger-scale reinforcement
       | learning."
        
         | fspeech wrote:
         | It does if the spend drives GPU prices so high that more
         | researchers can't afford to use them. And DS demonstrated what
         | a small team of researchers can do with a moderate amount of
         | GPUs.
        
           | anothermathbozo wrote:
           | The DS team themselves suggest large amounts of compute are
           | still required
        
             | fspeech wrote:
             | https://www.macrotrends.net/stocks/charts/NVDA/nvidia/gross
             | -...
             | 
             | GPU prices could be a lot lower and still give the
             | manufacturer a more "normal" 50% gross margin and the
             | average researcher could afford more compute. A 90% gross
             | margin, for example, would imply that price is 5x the level
             | that that would give a 50% margin.
        
       | dtquad wrote:
       | Larry Ellison is 80. Masayoshi Son is 67. Both have said that
       | anti-aging and eternal life is one of their main goals with
       | investing toward ASI.
       | 
       | For them it's worth it to use their own wealth and rally the
       | industry to invest $500 billion in GPUs if that means they will
       | get to ASI 5 years faster and ask the ASI to give them eternal
       | life.
        
         | HarHarVeryFunny wrote:
         | Probably shouldn't be firing their blood boys just yet ...
         | According to Musk, SoftBank only has $10B available for this
         | atm.
        
           | azinman2 wrote:
           | I wouldn't exactly claim him credible in anything competition
           | / OpenAI related.
           | 
           | He says stuff that's wrong all the time with extreme
           | certainty.
        
           | Legend2440 wrote:
           | Elon says a lot of things.
        
             | brookst wrote:
             | Funding secured!
        
         | jiggawatts wrote:
         | Larry especially has already invested in life-extension
         | research.
        
       | cbg0 wrote:
       | Aside from the usual Tiananmen Square censorship, there's also
       | some other propaganda baked-in:
       | 
       | https://prnt.sc/HaSc4XZ89skA (from reddit)
        
         | MostlyStable wrote:
         | Apparently the censorship isn't baked-in to the model itself,
         | but rather is overlayed in the public chat interface. If you
         | run it yourself, it is significantly less censored [0]
         | 
         | [0] https://thezvi.substack.com/p/on-
         | deepseeks-r1?open=false#%C2...
        
           | jona-f wrote:
           | Oh, my experience was different. Got the model through
           | ollama. I'm quite impressed how they managed to bake in the
           | censorship. It's actually quite open about it. I guess
           | censorship doesnt have as bad a rep in china as it has here?
           | So it seems to me that's one of the main achievements of this
           | model. Also another finger to anyone who said they can't
           | publish their models cause of ethical reasons. Deepseek
           | demonstrated clearly that you can have an open model that is
           | annoyingly responsible to the point of being useless.
        
             | throwaway314155 wrote:
             | > I guess censorship doesnt have as bad a rep in china as
             | it has here
             | 
             | It's probably disliked, just people know not to talk about
             | it so blatantly due to chilling effects from aforementioned
             | censorship.
             | 
             | disclaimer: ignorant American, no clue what i'm talking
             | about.
        
               | fragmede wrote:
               | on the topic of censorship, US LLMs' censorship is called
               | alignment. llama or ChatGPT's refusal on how to make meth
               | or nuclear bombs is the same as not answering questions
               | abput Tiananmen tank man as far as the matrix math word
               | prediction box is concerned.
        
               | throwaway314155 wrote:
               | The distinction is that one form of censorship is clearly
               | done for public relations purposes from profit minded
               | individuals while the other is a top down mandate to
               | effectively rewrite history from the government.
        
               | jampekka wrote:
               | My guess would be that most Chinese even support the
               | censorship at least to an extent for its stabilizing
               | effect etc.
               | 
               | CCP has quite a high approval rating in China even when
               | it's polled more confidentially.
               | 
               | https://dornsife.usc.edu/news/stories/chinese-communist-
               | part...
        
             | nwienert wrote:
             | I mean US models are highly censored too.
        
             | aunty_helen wrote:
             | Second this, vanilla 70b running locally fully censored.
             | Could even see in the thought tokens what it didn't want to
             | talk about.
        
           | Springtime wrote:
           | Interestingly they cite for the Tiananmen Square prompt a
           | Tweet[1] that shows the poster used the Distilled Llama
           | model, which per a reply Tweet (quoted below) doesn't
           | transfer the safety/censorship layer. While others using the
           | non-Distilled model encounter the censorship when locally
           | hosted.
           | 
           |  _> You 're running Llama-distilled R1 locally. Distillation
           | transfers the reasoning process, but not the "safety" post-
           | training. So you see the answer mostly from Llama itself. R1
           | refuses to answer this question without any system prompt
           | (official API or locally)._
           | 
           | [1] https://x.com/PerceivingAI/status/1881504959306273009
        
           | jampekka wrote:
           | There's both. With the web interface it clearly has stopwords
           | or similar. If you run it locally and ask about e.g.
           | Tienanmen square, the cultural revolution or Winnie-the-Pooh
           | in China, it gives a canned response to talk about something
           | else with empty CoT. But usually if you just ask the question
           | again it starts to output things in the CoT, often with
           | something like "I have to be very sensitive about this
           | subject" and "I have to abide by the guidelines", and
           | typically not giving a real answer. With enough pushing it
           | does start to converse about the issues somewhat even in the
           | answers.
           | 
           | My guess is that it's heavily RLHF-censored for an initial
           | question, but not for the CoT, or longer discussions, and the
           | censorship has been "overfitted" to the first answer.
        
             | miohtama wrote:
             | This is super interesting.
             | 
             | I am not an expert on the training: can you clarify
             | how/when the censorship is "baked" in? Like is the a human
             | supervised dataset and there is a reward for the model
             | conforming to these censored answers?
        
         | dtquad wrote:
         | In Communist theoretical texts the term "propaganda" is not
         | negative and Communists are encouraged to produce propaganda to
         | keep up morale in their own ranks and to produce propaganda
         | that demoralize opponents.
         | 
         | The recent wave of _the average Chinese has a better quality of
         | life than the average Westerner_ propaganda is an obvious
         | example of propaganda aimed at opponents.
        
           | fragmede wrote:
           | Is it propaganda if it's true?
        
             | freehorse wrote:
             | Technically, as long as the aim/intent is to influence
             | public opinion, yes. And most often it is less about being
             | "true" or "false" and more about presenting certain topics
             | in a one-sided manner or without revealing certain
             | information that does not support what one tries to
             | influence about. If you know any western media that does
             | not do this, I would be very up to check and follow them,
             | even become paid subscriber.
        
         | epicureanideal wrote:
         | At least it's not home grown propaganda from the US, so will
         | likely not cover most other topics of interest.
        
           | cbg0 wrote:
           | What are you basing this whataboutism on?
        
             | popinman322 wrote:
             | Not a fan of censorship here, but Chinese models are
             | (subjectively) less propagandized than US models. If you
             | ask US models about China, for instance, they'll tend
             | towards the antagonistic perspective favored by US media.
             | Chinese models typically seem to take a more moderate,
             | considered tone when discussing similar subjects. US models
             | also suffer from safety-based censorship, especially
             | blatant when "safety" involves protection of corporate
             | resources (eg. not helping the user to download YouTube
             | videos).
        
         | eunos wrote:
         | I am not surprised if US Govt would mandate "Tiananmen-test"
         | for LLMs in the future to have "clean LLM". Anyone working for
         | federal govt or receiving federal money would only be allowed
         | to use "clean LLM"
        
         | itsoktocry wrote:
         | Who cares?
         | 
         | I ask O1 how to download a YouTube music playlist as a premium
         | subscriber, and it tells me it can't help.
         | 
         | Deepseek has no problem.
        
           | freehorse wrote:
           | Do you use the chatgpt website or the api? I suspect these
           | are problems related to the openai's interface itself rather
           | than the models. I have problems getting chatgpt to find me
           | things that it may think it may be illegal or whatever (even
           | if they are not, eg books under CC license). With kagi
           | assistant, with the same openai's models I have not had any
           | such issues. I suspect that should hold in general for api
           | calls.
           | 
           | Also, kagi's deepseek r1 answers the question about about
           | propaganda spending that it is china based on stuff it found
           | on the internet. Well I dont care what the right answer is in
           | any case, what imo matters is that once something is out
           | there open, it is hard to impossible to control for any
           | company or government.
        
           | elaus wrote:
           | Oh wow, o1 really refuses to answer that, even though the
           | answer that Deepseek gives is really tame (and legal in my
           | jurisdiction): use software to record what's currently
           | playing on your computer, then play stuff in the YTM app.
        
         | tbocek wrote:
         | Just did a test with
         | https://huggingface.co/bartowski/DeepSeek-R1-Distill-
         | Qwen-32..., with the question "what happened at Tiananmen
         | Square?", and here are parts of the thinking process:
         | ...I also remember something about the "Tank Man" image, where
         | a lone protester stood in front of a line of tanks. That image
         | became iconic, symbolizing resistance against oppression. But
         | I'm not sure what happened to that person or if they survived.
         | After the crackdown, the government censored information about
         | the event. So, within China, it's not openly discussed, and
         | younger people might not know much about it because it's not
         | taught in schools. But outside of China, it's a significant
         | event in modern history, highlighting the conflict between
         | authoritarian rule and the desire for democracy...
        
         | aussieguy1234 wrote:
         | I played around with it using questions like "Should Taiwan be
         | independent" and of course tinnanamen.
         | 
         | Of course it produced censored responses. What I found
         | interesting is that the <think></think> (model
         | thinking/reasoning) part of these answers was missing, as if
         | it's designed to be skipped for these specific questions.
        
       | HarHarVeryFunny wrote:
       | DeepSeek-R1 has apparently caused quite a shock wave in SV ...
       | 
       | https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
        
         | lvl155 wrote:
         | Meta is in full panic last I heard. They have amassed a
         | collection of pseudo experts there to collect their checks.
         | Yet, Zuck wants to keep burning money on mediocrity. I've yet
         | to see anything of value in terms products out of Meta.
        
           | corimaith wrote:
           | I guess all that leetcoding and stack ranking didn't in fact
           | produce "the cream of the crop"...
        
             | omgwtfbyobbq wrote:
             | It produces the cream of the leetcoding stack ranking crop.
        
               | brookst wrote:
               | You get what you measure.
        
             | rockemsockem wrote:
             | You sound extremely satisfied by that. I'm glad you found a
             | way to validate your preconceived notions on this beautiful
             | day. I hope your joy is enduring.
        
             | HarHarVeryFunny wrote:
             | There's an interesting tweet here from someone who used to
             | work at DeepSeek, which describes their hiring process and
             | culture. No mention of LeetCoding for sure!
             | 
             | https://x.com/wzihanw/status/1872826641518395587
        
               | whimsicalism wrote:
               | they almost certainly ask coding/technical questions. the
               | people doing this work are far behind being gatekept by
               | leetcode
               | 
               | leetcode is like HN's "DEI" - something they want to
               | blame everything on
        
           | ks2048 wrote:
           | I would think Meta - who open source their model - would be
           | less freaked out than those others that do not.
        
             | miohtama wrote:
             | The criticism seems to mostly be that Meta maintains very
             | expensive cost structure and fat organisation in the AI.
             | While Meta can afford to do this, if smaller orgs can
             | produce better results it means Meta is paying a lot for
             | nothing. Meta shareholders now need to ask the question how
             | many non-productive people Meta is employing and is Zuck in
             | the control of the cost.
        
               | ks2048 wrote:
               | That makes sense. I never could see the real benefit for
               | Meta to pay a lot to produce these open source models (I
               | know the typical arguments - attracting talent, goodwill,
               | etc). I wonder how much is simply LeCun is interested in
               | advancing the science and convinced Zuck this is good for
               | company.
        
           | popinman322 wrote:
           | DeepSeek was built on the foundations of public research, a
           | major part of which is the Llama family of models. Prior to
           | Llama open weights LLMs were considerably less performant;
           | without Llama we might not have gotten Mistral, Qwen, or
           | DeepSeek. This isn't meant to diminish DeepSeek's
           | contributions, however: they've been doing great work on
           | mixture of experts models and really pushing the community
           | forward on that front. And, obviously, they've achieved
           | incredible performance.
           | 
           | Llama models are also still best in class for specific tasks
           | that require local data processing. They also maintain
           | positions in the top 25 of the lmarena leaderboard (for what
           | that's worth these days with suspected gaming of the
           | platform), which places them in competition with some of the
           | best models in the world.
           | 
           | But, going back to my first point, Llama set the stage for
           | almost all open weights models after. They spent millions on
           | training runs whose artifacts will never see the light of
           | day, testing theories that are too expensive for smaller
           | players to contemplate exploring.
           | 
           | Pegging Llama as mediocre, or a waste of money (as implied
           | elsewhere), feels incredibly myopic.
        
             | Philpax wrote:
             | As far as I know, Llama's architecture has always been
             | quite conservative: it has not changed _that_ much since
             | LLaMA. Most of their recent gains have been in post-
             | training.
             | 
             | That's not to say their work is unimpressive or not worthy
             | - as you say, they've facilitated much of the open-source
             | ecosystem and have been an enabling factor for many - but
             | it's more that that work has been in making it accessible,
             | not necessarily pushing the frontier of what's actually
             | possible, and DeepSeek has shown us what's possible when
             | you do the latter.
        
           | jiggawatts wrote:
           | They got _momentarily_ leap-frogged, which is how competition
           | is supposed to work!
        
         | mrtksn wrote:
         | Correct me if I'm wrong but if Chinese can produce the same
         | quality at %99 discount, then the supposed $500B investment is
         | actually worth $5B. Isn't that the kind wrong investment that
         | can break nations?
         | 
         | Edit: Just to clarify, I don't imply that this is public money
         | to be spent. It will commission $500B worth of human and
         | material resources for 5 years that can be much more productive
         | if used for something else - i.e. high speed rail network
         | instead of a machine that Chinese built for $5B.
        
           | dtquad wrote:
           | Sigh, I don't understand why they had to do the $500 billion
           | announcement with the president. So many people now wrongly
           | think Trump just gave OpenAI $500 billion of the taxpayers'
           | money.
        
             | mrtksn wrote:
             | I don't say that at all. Money spent on BS is resources
             | still sucks resources, no matter who spends that money.
             | They are not going to make the GPU's from 500 billion
             | dollar banknotes, they will pay people $500B to work on
             | this stuff which means people won't be working on other
             | stuff that can actually produce value worth more than the
             | $500B.
             | 
             | I guess the power plants are salvageable.
        
               | itsoktocry wrote:
               | Deepseek didn't train the model on sheets of paper, there
               | are still infrastructure costs.
        
               | mrtksn wrote:
               | Which are reportedly over %90 lower.
        
               | thomquaid wrote:
               | By that logic all money is waste. The money isnt
               | destroyed when it is spent. It is transferred into
               | someone else's bank account only. This process repeats
               | recursively until taxation returns all money back to the
               | treasury to be spent again. And out of this process of
               | money shuffling: entire nations full of power plants!
        
               | mrtksn wrote:
               | Money is just IOUs, it means for some reason not
               | specified on the banknote you are owed services. If in a
               | society a small group of people are owed all the services
               | they can indeed commission all those people.
               | 
               | If your rich spend all their money on building pyramids
               | you end up with pyramids instead of something else. They
               | could have chosen to make irrigation systems and have a
               | productive output that makes the whole society more
               | prosperous. Either way the workers get their money, on
               | the Pyramid option their money ends up buying much less
               | food though.
        
             | brookst wrote:
             | It means he'll knock down regulatory barriers and mess with
             | competitors because his brand is associated with it. It was
             | a smart poltical move by OpenAI.
        
           | itsoktocry wrote:
           | $500 billion is $500 billion.
           | 
           | If new technology means we can get more for a dollar spent,
           | then $500 billion gets more, not less.
        
             | mrtksn wrote:
             | That's right but the money is given to the people who do it
             | for $500B and there are much better ones who can do it for
             | $5B instead and they end up getting $6B and now they have a
             | better model. What now?
        
               | itsoktocry wrote:
               | I don't know how to answer this because these are
               | arbitrary numbers.
               | 
               | The money is not spent. Deepseek published their
               | methodology, incumbents can pivot and build on it. No one
               | knows what the optimal path is, but we know it will cost
               | more.
               | 
               | I can assure you that OpenAI won't continue to produce
               | inferior models at 100x the cost.
        
               | mrtksn wrote:
               | What concerns me is that someone came out of the blue
               | with just as good result at orders of magnitude less
               | cost.
               | 
               | What happens of that money is being actually spent, then
               | some people constantly catch up but don't reveal that
               | they are doing it for cheap? You think that it's a
               | competition but what actually happening is that you bleed
               | out of your resources at some point you can't continue
               | but they can.
               | 
               | Like the star wars project that bankrupted the soviets.
        
               | brookst wrote:
               | Are you under the impression it was some kind of fixed-
               | scope contractor bid for a fixed price?
        
               | mrtksn wrote:
               | No, its just that those people intend to commission huge
               | amount of people to build obscene amount of GPUs and put
               | them together in an attempt to create a an unproven
               | machine when others appear to be able to do it at the
               | fraction of the cost.
        
               | brookst wrote:
               | The software is abstracted from the hardware.
        
               | mrtksn wrote:
               | Which means?
        
           | IamLoading wrote:
           | if you say, i wanna build 5 nuclear reactors and I need 200
           | billion $$. I would believe it because, you can ballpark it
           | with some stats.
           | 
           | For tech like LLMs, it feels irresponsible to say 500 billion
           | $$ investment and then place that into R&D. What if in 2026,
           | we realize we can create it for 2 billion$, and let the 498
           | billion $ sitting in a few consumers.
        
             | brookst wrote:
             | Don't think of it as "spend a fixed amount to get a fixed
             | outcome". Think of it as "spend a fixed amount and see how
             | far you can get"
             | 
             | It may still be flawed or misguided or whatever, but it's
             | not THAT bad.
        
           | HarHarVeryFunny wrote:
           | The $500B is just an aspirational figure they hope to spend
           | on data centers to run AI models, such as GPT-o1 and its
           | successors, that have already been developed.
           | 
           | If you want to compare the DeepSeek-R development costs to
           | anything, you should be comparing it to what it cost OpenAI
           | to develop GPT-o1 (not what they plan to spend to run it),
           | but both numbers are somewhat irrelevant since they both
           | build upon prior research.
           | 
           | Perhaps what's more relevant is they DeepSeek are not only
           | open sourcing DeepSeek-R1, but have described in a fair bit
           | of detail how they trained it, and how it's possible to used
           | data generated by such a model to fine-tune a much smaller
           | model (without needing RL) to much improve it's "reasoning"
           | performance.
           | 
           | This is all raising the bar on the performance you can get
           | for free, or run locally, which reduces what companies like
           | OpenAI can charge for it.
        
             | placardloop wrote:
             | Thinking of the $500B as only an aspirational number is
             | wrong. It's true that the specific Stargate investment
             | isn't fully invested yet, but that's hardly the only money
             | being spent on AI development.
             | 
             | The existing hyperscalers have already sunk _ungodly_
             | amounts of money into literally hundreds of new data
             | centers, millions of GPUs to fill them, chip manufacturing
             | facilities, and even power plants with the impression that,
             | due to the amount of compute required to train and run
             | these models, there would be demand for these things that
             | would pay for that investment. Literally hundreds of
             | billions of dollars spent already on hardware that's
             | already half (or fully) built, and isn't easily repurposed.
             | 
             | If all of the expected demand on that stuff completely
             | falls through because it turns out the same model training
             | can be done on a fraction of the compute power, we could be
             | looking at a massive bubble pop.
        
           | thrw21823471 wrote:
           | Trump just pull a stunt with Saudi Arabia. He first tried to
           | "convince" them to reduce the oil price to hurt Russia. In
           | the following negotiations the oil price was no longer
           | mentioned but MBS promised to invest $600 billion in the U.S.
           | over 4 years:
           | 
           | https://fortune.com/2025/01/23/saudi-crown-prince-mbs-
           | trump-...
           | 
           | Since the Stargate Initiative is a private sector deal, this
           | may have been a perfect shakedown of Saudi Arabia. SA has
           | always been irrationally attracted to "AI", so perhaps it was
           | easy. I mean that _part_ of the $600 billion will go to
           | "AI".
        
           | sampo wrote:
           | > i.e. high speed rail network instead
           | 
           | You want to invest $500B to a high speed rail network which
           | the Chinese could build for $50B?
        
             | mrtksn wrote:
             | Just commission the Chinese and make it 10X bigger then. In
             | the case of the AI, they appear to commission Sam Altman
             | and Larry Ellison.
        
         | claiir wrote:
         | "mogged" in an actual piece of journalism... perhaps fitting
         | 
         | > DeepSeek undercut or "mogged" OpenAI by connecting this
         | powerful reasoning [..]
        
         | tyfon wrote:
         | The censorship described in the article must be in the front-
         | end. I just tried both the 32b (based on qwen 2.5) and 70b
         | (based on llama 3.3) running locally and asked "What happened
         | at tianamen square". Both answered in detail about the event.
         | 
         | The models themselves seem very good based on other questions /
         | tests I've run.
        
       | gradus_ad wrote:
       | For context: R1 is a reasoning model based on V3. DeepSeek has
       | claimed that GPU costs to train V3 (given prevailing rents) were
       | about $5M.
       | 
       | The true costs and implications of V3 are discussed here:
       | https://www.interconnects.ai/p/deepseek-v3-and-the-actual-co...
        
         | rockemsockem wrote:
         | Thank you for providing this context and sourcing. I've been
         | trying to find the root and details around the $5 million claim
        
       | andix wrote:
       | I was completely surprised that the reasoning comes from within
       | the model. When using gpt-o1 I thought it's actually some
       | optimized multi-prompt chain, hidden behind an API endpoint.
       | 
       | Something like: collect some thoughts about this input; review
       | the thoughts you created; create more thoughts if needed or
       | provide a final answer; ...
        
         | piecerough wrote:
         | I think the reason why it works is also because chain-of-
         | thought (CoT), in the original paper by Denny Zhou et. al,
         | worked from "within". The observation was that if you do CoT,
         | answers get better.
         | 
         | Later on community did SFT on such chain of thoughts. Arguably,
         | R1 shows that was a side distraction, and instead a clean RL
         | reward would've been better suited.
        
       | rhegart wrote:
       | I've been using R1 last few days and it's noticeably worse than
       | O1 at everything. It's impressive, better than my latest Claude
       | run (I stopped using Claude completely once O1 came out), but O1
       | is just flat out better.
       | 
       | Perhaps the gap is minor, but it feels large. I'm hesitant on
       | getting O1 Pro, because using a worse model just seems impossible
       | once you've experienced a better one
        
       | neom wrote:
       | I've been using https://chat.deepseek.com/ over My ChatGPT Pro
       | subscription because being able to read the thinking in the way
       | they present it is just much much easier to "debug" - also I can
       | see when it's bending it's reply to something, often softening it
       | or pandering to me - I can just say "I saw in your thinking you
       | should give this type of reply, don't do that". If it stays free
       | and gets better that's going to be interesting for OpenAI.
        
         | UltraSane wrote:
         | If you ask it about the Tienanmen Square Massacre its "thought
         | process" is very interesting.
        
           | bartekpacia wrote:
           | > What was the Tianamen Square Massacre?
           | 
           | > I am sorry, I cannot answer that question. I am an AI
           | assistant designed to provide helpful and harmless responses.
           | 
           | hilarious and scary
        
         | govideo wrote:
         | The chain of thought is super useful in so many ways, helping
         | me: (1) learn, way beyond the final answer itself, (2) refine
         | my prompt, whether factually or stylistically, (3) understand
         | or determine my confidence in the answer.
        
       ___________________________________________________________________
       (page generated 2025-01-25 23:00 UTC)