Post BAPMwKoGxb4hSJMZFI by ariadne@social.treehouse.systems
(DIR) More posts by ariadne@social.treehouse.systems
(DIR) Post #BAPKpGXIFhAMGUvGbo by ariadne@social.treehouse.systems
0 likes, 0 repeats
RE: https://social.treehouse.systems/@ariadne/117270693101928179i took a few minutes to quickly read through the essay dario amodei (CEO of anthropic) dropped over the weekend.my takeaway from it is that "pacing the frontier" is more likely a convenient narrative for a problem the frontier labs are inevitably going to run into anyway: as these monolithic models scale, the physical and economic difficulty of scaling them grows with them.RT: https://social.treehouse.systems/users/ariadne/statuses/117270693101928179
(DIR) Post #BAPLECn1JmLtsyAelE by ariadne@social.treehouse.systems
0 likes, 0 repeats
the scaling laws themselves will likely continue to hold just fine, but the bottleneck is increasingly the infrastructure required to realize the next point on the scaling curve.for a concrete example: a 1-trillion-parameter dense model at 16-bit precision needs about 2 TB of memory just to hold its weights. that's roughly 25 nvidia H100 accelerators just to load one instance of the model.and once a single inference has to cross dozens of accelerators, memory bandwidth and chip-to-chip communication become part of the latency budget too.at that point, “make it bigger” is no longer mainly a model-design problem. it is a hardware, networking, power, and datacenter problem.
(DIR) Post #BAPLd1RycAxxxNA57w by ariadne@social.treehouse.systems
0 likes, 0 repeats
the growth in infrastructure requirements just to scale models dramatically changes what frontier scaling looks like over time.instead of a relatively smooth progression where each new generation mostly waits on another training run, you get something much more stair-stepped similar to the intel tick-tock model:new accelerator generation (tick) → build out enough capacity → scale the model → optimize around the new bottlenecks (tock) → plateau → wait for the next hardware generation.in other words, the pace of frontier capability starts getting set less by model architecture and more by nvidia, tsmc, hbm supply, networking, power delivery and datacenter construction.
(DIR) Post #BAPLkWavDfPsZlvI80 by dvshkn@social.treehouse.systems
0 likes, 0 repeats
@ariadne Yeah, Dario is already on record complaining about how volatile it is provisioning the next year's compute budget. Too little and your company dies, too much and your company and possibly also the US economy also dies. I do feel like this must've been in the back of his mind at least somewhat as he started this most recent PR blitz.
(DIR) Post #BAPLufjs4lJdpfwou8 by mcc@mastodon.social
1 likes, 0 repeats
@ariadne also has implications for "but what if the model starts making copies of itself???" hypothetical
(DIR) Post #BAPLzomjry9bKwk7A8 by ariadne@social.treehouse.systems
1 likes, 1 repeats
and this is where dario's new "pacing" narrative becomes useful.if frontier development is going to become naturally gated by hardware generations and datacenter buildout anyway, then a slower release cadence no longer has to look like the scaling story is running out of steam.instead, you can frame that cadence as intentional: we are slowing down because we are "responsible", because we need time for evaluations and safety work, because the frontier must be paced.that is a much better story to tell investors, who are eagerly waiting for a credible path to ROI at this point, than "the next meaningful scaling step has to wait for another generation of accelerators and a few gigawatts of infrastructure."
(DIR) Post #BAPMEfqpMtfNyKjsu0 by ariadne@social.treehouse.systems
0 likes, 0 repeats
and the timing for this essay drop matters here: anthropic is preparing to go public.an IPO means asking public-market investors to underwrite a company whose future value depends heavily on the assumption that frontier capabilities will continue advancing, while the cost and physical difficulty of producing each new generation keeps increasing.if the actual development cadence is about to become more visibly constrained by hardware and infrastructure, you need a story for why that is happening."we have reached the point where responsible AI development requires us to pace the frontier" is an extremely good story, even if the real reason for pacing is due to physics: they simply cannot build the necessary infrastructure at the current scaling pace.
(DIR) Post #BAPMTfHXla24YDNq1Q by ariadne@social.treehouse.systems
0 likes, 0 repeats
this is basically the inversion i keep coming back to:you see, there is no scaling wall. instead, the frontier labs are simply "pacing themselves."that is a much friendlier way to explain the transition from rapid monolithic scaling to a hardware-constrained cadence, especially ahead of an IPO.and people will absolutely eat that framing up. we live in a world where a lot of otherwise technical people have started treating token prediction as though it were touching the divine.that credulity scares me, especially when the same labs are taking shortcuts on safety while asking everyone else to interpret their constraints as wisdom.
(DIR) Post #BAPMwKoGxb4hSJMZFI by ariadne@social.treehouse.systems
0 likes, 0 repeats
if this interpretation is right, it should make some pretty clear predictions.frontier capability jumps should increasingly cluster around new accelerator generations, higher-bandwidth memory, better interconnects, and new datacenter capacity coming online.between those jumps, we should see more emphasis on optimization, complaining about distillation from competitors, etc.in other words, the interesting question is no longer just "does scaling work?"it is "how much machine do you need to buy to realize the next point on the curve?"and the answer is always: "as big of a machine as possible."
(DIR) Post #BAPNNFRH2NMucIRGc4 by ariadne@social.treehouse.systems
0 likes, 0 repeats
this is also why i think the more interesting path forward is composable intelligence, not infinitely larger monoliths.a concrete example would be a system where one model handles planning, another does code or math, a retrieval system pulls in external knowledge, a verifier checks the result, and a memory layer carries state across the whole process.none of those components has to be the largest possible model. the capability comes from how they are composed.in other words, we do not need to be wasting all of humanity's resources on scaling monoliths, nor do i think doing so will yield the AGI the entire industry is betting on.the scaling wall does not mean progress stops, instead it means that new opportunities to pursue more economical directions of research emerge.
(DIR) Post #BAPOPrGTTN1cSsCCx6 by oblomov@sociale.network
0 likes, 0 repeats
@ariadne this was kind of the idea behind (some of the) expert models
(DIR) Post #BAPOWhgLV1eYltYr8y by ariadne@social.treehouse.systems
0 likes, 0 repeats
@oblomov problem is MoEs are still monoliths, and the labs cannot afford to have idle GPUs, so they punish specialization in the training process :(
(DIR) Post #BAPOliWycqLJkRL1IO by oblomov@sociale.network
0 likes, 0 repeats
@ariadne eh, I think the latter is more of an issue than the former. It would be possible to define MoEs in a composable and distributed way, but research is focused on centralization because centralization is more efficient and gives power to whoever controls the system 8-/
(DIR) Post #BAPVYBdkIiV7ea01dA by ariadne@social.treehouse.systems
0 likes, 0 repeats
@oblomov MoEs are a far less efficient way of doing composability though
(DIR) Post #BAPcBaaq3izM3xoMrI by ariadne@social.treehouse.systems
0 likes, 1 repeats
@subrealz but my point is Amodei does not actually *want* a pause. he is being forced, by the laws of physics and economics, to slow down.
(DIR) Post #BAPcUORFTS8NrsIAnA by ariadne@social.treehouse.systems
0 likes, 1 repeats
@subrealz and this distinction, that he is being *forced* to slow down is important, because his framing of this slow down as "pacing" implies agency he actually does not have.in other words, calling that "pacing" is branding a constraint as a choice.
(DIR) Post #BAPj9DiwoZOQEeQDM8 by be_far@social.treehouse.systems
0 likes, 0 repeats
@ariadne lying through their teeth because reality isn’t marketable? Sounds like the AI industry alright, at least in my opinion
(DIR) Post #BAQDOpYnQ5ZEYE27M0 by crzwdjk@mastodon.social
0 likes, 0 repeats
@ariadne Apparently the next big thing is 8-bit, 6-bit, and even 4-bit floating point formats, I guess to cram in ever more parameters into limited RAM.
(DIR) Post #BAQyBqmEHSlAPCd32G by Zenie@piaille.fr
0 likes, 0 repeats
@ariadne I don't believe it's about pacing at all. It's about the AI companies working together to overcome political roadblocks and remove guard rails to their growth and monopolistic goals.
(DIR) Post #BAVzH4flmKNPQ5DXfc by Okanogen@mastodon.social
0 likes, 0 repeats
@ariadne The humble brag that "please slow us down! We are too smart and will create a world-killing Super AI in ten years!" is also a great tale to tell investors.
(DIR) Post #BAWNUkHqlc0j9G55Si by Okanogen@mastodon.social
0 likes, 0 repeats
@ariadne #AI is #PataPhysics: "The science of imaginary solutions" monitarized.